← Back to Index
Daily Research Digest

arXiv Papers

2026-08-25
571
Papers
4
Categories
571
Translated
收藏清单 0
机器人学 (Robotics)
80
cs.RO / 1 / 2608.21380

RoboShape: Information-Theoretic Point Cloud Representations for Privacy-Aware Robot Perception

RoboShape:用于隐私意识机器人感知的信息论点云表示
Baser, Oguzhan, Sozen, Mirac, Kale, Kaan, Chinchali, Sandeep, Vishwanath, Sriram
Abstract
With the increased adoption of robotic agents operating in human environments by scanning and sharing 3D representations (e.g., for fleet learning, cloud-based planning, or collaborative mapping), collected point clouds reveal not just the objects in a scene but also sensitive spatial context, such as room function or information that occupants never consented to disclose. Traditional point cloud encoders offer no principled control over this: either all is preserved, or none. Hence, we introduce RoboShape, an information theory guided compression head following the frozen {\tt Sonata} encoder. We project voxel-level embeddings using the Donsker-Varadhan formulation of mutual information (MI). Specifically, we maximize the MI between embeddings and object-level understanding while minimizing it for private attributes. RoboShape leads to 87.5\% smaller embeddings that retain 98.7\% of object classification utility while collapsing sensitive attribute predictions by 39.3\% across the three real-world indoor LiDAR datasets. Its privacy-preserving embeddings are cheaper to transmit over the network or to train a model for any downstream tasks. We release the RoboShape codebase to give the robotics community a practical, encoder-agnostic tool for building perception pipelines that are compact, privacy-aware, and deployment-ready.
Chinese Translation
随着机器人代理在扫描和共享三维表示(例如,用于车队学习、基于云的规划或协作映射)的人类环境中的应用日益增加,收集到的点云不仅揭示了场景中的物体,还暴露了敏感的空间上下文,例如房间功能或居住者未曾同意披露的信息。传统的点云编码器对这一问题没有原则性的控制:要么全部保留,要么全部丢弃。因此,我们提出了RoboShape,这是一个基于信息理论指导的压缩头,跟随冻结的{ t Sonata}编码器。我们使用Donsker-Varadhan互信息(MI)公式对体素级嵌入进行投影。具体而言,我们在最大化嵌入与物体级理解之间的MI的同时,最小化与私密属性之间的MI。RoboShape的嵌入体积缩小了87.5\%,同时保留了98.7\\%的物体分类效用,而在三个真实世界的室内LiDAR数据集上,敏感属性预测减少了39.3\\%。其隐私保护的嵌入在网络上传输或训练下游任务模型时成本更低。我们发布了RoboShape代码库,为机器人社区提供一个实用的、与编码器无关的工具,以构建紧凑、隐私意识且准备部署的感知管道。
cs.RO / 2 / 2608.21387

Multimodal-Language-Model-Driven Interaction and Companionship for Service Robots in Elderly-Care Facilities

多模态语言模型驱动的服务机器人在养老院的互动与陪伴
Liu, Ching-Chieh, Vu, Cong-Thanh, Liu, Yen-Chen
Abstract
Service robots are increasingly deployed in elderly-care facilities to alleviate caregiver workload and enhance the quality of daily care. However, most existing studies focus on isolated service functions and lack integrated capabilities for continuous companionship, natural interaction, and safety monitoring. In this paper, we present an intelligent companion robot system that unifies active visual human-following, real-time LLM-driven speech interaction for intent understanding and task execution, and VLM-based safety monitoring for fall detection and abnormal posture assessment. The perception layer ensures robust human tracking and uses an active gimbal to maintain the user in view during occlusions or abrupt movements. At the interaction layer, a Large Language Model interprets spoken requests and maps them to robot actions, enabling escorting and semantic navigation. Simultaneously, a VLM-based safety agent continuously analyzes visual observations to detect fall-related or abnormal postures and triggers emergency responses when necessary. Experimental results demonstrate the system's ability to reliably follow and interact with humans, while effectively detecting potential falls to ensure user safety.
Chinese Translation
服务机器人在养老院的应用日益增多,以减轻护理人员的工作负担并提升日常护理的质量。然而,现有研究大多集中于孤立的服务功能,缺乏持续陪伴、自然互动和安全监测的综合能力。本文提出了一种智能陪伴机器人系统,该系统统一了主动视觉人跟踪、基于实时大型语言模型(LLM)的语音互动以理解意图和执行任务,以及基于视觉语言模型(VLM)的安全监测以进行跌倒检测和异常姿势评估。感知层确保稳健的人体追踪,并使用主动云台在遮挡或突发运动期间保持用户在视野内。在互动层,大型语言模型解读口头请求并将其映射到机器人动作,从而实现陪同和语义导航。同时,基于视觉语言模型的安全代理持续分析视觉观察,以检测与跌倒相关或异常的姿势,并在必要时触发紧急响应。实验结果表明,该系统能够可靠地跟随和与人互动,同时有效检测潜在的跌倒,以确保用户安全。
cs.RO / 3 / 2608.21388

Gimbal-Based Human Tracking for Companion Robots Using Continual Learning

基于云台的人类跟踪方法用于伴侣机器人,采用持续学习
Vu, Cong-Thanh, Liu, Ching-Chieh, Liu, Yen-Chen
Abstract
Reliable and continuous human tracking is essential for natural human-robot interaction, particularly for companion robots. However, many existing approaches rely on wearable tags or fixed cameras with limited fields of view, which reduces system flexibility and often causes tracking failures when the target moves outside the sensing range. In this paper, we present a human tracking approach based on a gimbal-mounted camera integrated into a mobile robot. By actively controlling the gimbal mechanism, the camera can dynamically adjust its viewing direction to maintain the target within the field of view, even under substantial relative motion between the robot and the human. Furthermore, a continual learning strategy is applied to the person re-identification (ReID) task to adapt to changes in appearance and environmental conditions during long-term tracking. Experimental results demonstrate that the proposed system significantly improves the stability and continuity of human tracking, enables real-time re-identification, and provides responsive feedback for reliable tracking of human motion from walking to running. User studies further indicate that the proposed approach enhances user comfort by eliminating the need for wearable tags.
Chinese Translation
可靠且连续的人类跟踪对于自然的人机交互至关重要,尤其是对于伴侣机器人。然而,许多现有的方法依赖于可穿戴标签或固定摄像头,这些设备的视野有限,降低了系统的灵活性,并且在目标移动到感知范围之外时,常常导致跟踪失败。本文提出了一种基于云台安装摄像头的人类跟踪方法,该摄像头集成在移动机器人中。通过主动控制云台机制,摄像头能够动态调整其视角,以保持目标在视野范围内,即使在机器人与人之间存在显著的相对运动。此外,我们在长期跟踪过程中对人物重识别(ReID)任务应用了持续学习策略,以适应外观和环境条件的变化。实验结果表明,所提出的系统显著提高了人类跟踪的稳定性和连续性,实现了实时重识别,并为可靠跟踪人类运动(从步行到跑步)提供了响应反馈。用户研究进一步表明,所提出的方法通过消除对可穿戴标签的需求,增强了用户的舒适感。
cs.RO / 4 / 2608.21390

On the Optimized Use of Non-Orthonormality Constraints for the Quasi-Static INS Alignment of Autonomous Underwater and Surface Vehicles

优化非正交约束在自主水下和地面车辆准静态惯性导航系统对准中的应用
Durao, Carlos Renato C., Silva, Felipe O., Klein, Itzik, Cavalcanti, Vinıcius M. G. B., Frutuoso, Adriano, de Barros, Ettore A., Farrell, Jay A.
Abstract
Inertial navigation systems are specialized navigation apparatuses that equip almost all autonomous underwater and surface vehicles. They require precise initial alignment, i.e., determination of their initial attitude, which is typically achieved: (a) in quasi-static conditions (whenever possible); and (b) in two stages: Coarse Alignment (CA), using methods like TRI-axis Attitude Determination (TRIAD), and Fine Alignment (FA), via Zero Velocity Update (ZVU)-based Extended Kalman Filtering (EKF). However, conventional methods suffer from slow convergence and limited bias estimability. In response, this paper introduces: (a) an optimized version of a recently proposed CA method, namely, TRIAD with Coarse Bias Estimation (TRIAD-CBE); and (b) a novel FA EKF observation model that incorporates Non-Orthonormality (NON) error constraints derived from TRIAD, directly linking these errors to the inertial sensor biases. As validated through extensive Monte Carlo (MC) simulations, as well as real-world experiments using two Inertial Measurement Units (IMUs) of different grades, our approaches substantially accelerate the convergence of misalignment and bias estimates (from minutes to seconds), while maintaining accuracy/precision comparable to traditional techniques.
Chinese Translation
惯性导航系统是几乎所有自主水下和地面车辆所配备的专业导航设备。它们需要精确的初始对准,即确定其初始姿态,这通常通过以下方式实现:(a)在准静态条件下(尽可能);(b)分两阶段进行:粗对准(Coarse Alignment, CA),使用如三轴姿态确定(TRIAD)等方法,以及精对准(Fine Alignment, FA),通过基于零速度更新(Zero Velocity Update, ZVU)的扩展卡尔曼滤波(Extended Kalman Filtering, EKF)。然而,传统方法存在收敛速度慢和偏差估计能力有限的问题。为此,本文提出:(a)一种最近提出的CA方法的优化版本,即带粗偏差估计的三轴姿态确定(TRIAD with Coarse Bias Estimation, TRIAD-CBE);(b)一种新颖的FA EKF观测模型,该模型结合了源自TRIAD的非正交性(Non-Orthonormality, NON)误差约束,直接将这些误差与惯性传感器偏差联系起来。通过广泛的蒙特卡洛(Monte Carlo, MC)仿真以及使用两种不同等级的惯性测量单元(Inertial Measurement Units, IMUs)的实际实验验证,我们的方法显著加快了错位和偏差估计的收敛速度(从分钟缩短到秒),同时保持了与传统技术相当的准确性和精度。
cs.RO / 5 / 2608.21395

ODG-NoMaD: Overhead-Camera Direction-Guided NoMaD

ODG-NoMaD:基于顶视摄像头的方向引导NoMaD
Bastian, Blossom Treesa, Shetty, Keerthi S., Kolachalam, Manish, Malhotra, Rani, Dutta, Ashish
Abstract
NoMaD [31] is a learned vision-navigation policy that unifies goal-conditioned navigation and exploration in a single goal-masked diffusion policy. In an unseen environment, however - where neither a goal image nor a topological map is available - it can only explore undirectedly, wandering without global awareness. We present ODG-NoMaD, which gives NoMaD's exploration mode a global sense of where to proceed, without retraining the policy. An overhead depth camera is used once on deployment to build an occupancy map and plan a global path, which is segmented to yield a desired heading; a per-frame traversability map from the robot's onboard depth then refines this into a collision-free direction. The gradient of a cosine direction cost is injected into the final denoising steps, rotating sampled trajectories toward this direction while preserving the multimodality of exploration. In simulated office environments with and without random obstacles, ODG-NoMaD reduces the residual distance to the target by up to an order of magnitude over unguided exploration, outperforms the point-goal cost guidance of NaviDiffusor [37], and is the only configuration that remains collision-free on every trial.
Chinese Translation
NoMaD [31] 是一种学习的视觉导航策略,统一了目标条件下的导航与探索,形成了单一的目标掩蔽扩散策略。然而,在一个未知环境中——即没有目标图像或拓扑地图可用的情况下——它只能进行无方向的探索,漫无目的地游荡而缺乏全局意识。我们提出了ODG-NoMaD,它为NoMaD的探索模式提供了全球性的前进方向,而无需重新训练策略。在部署时,使用一次顶视深度摄像头构建占用地图并规划全局路径,然后对其进行分段以获得所需的航向;机器人的车载深度传感器生成的每帧可通行性地图进一步将其细化为无碰撞方向。余弦方向成本的梯度被注入到最终的去噪步骤中,旋转采样轨迹朝向该方向,同时保持探索的多模态性。在有无随机障碍物的模拟办公室环境中,ODG-NoMaD将目标的残余距离减少了一个数量级,优于NaviDiffusor [37] 的点目标成本引导,并且是唯一在每次试验中保持无碰撞的配置。
cs.RO / 6 / 2608.21400

Active Interaction-Aware Model Predictive Path Integral via Ego-Conditioned Generative Predictions

基于自我条件生成预测的主动交互感知模型预测路径积分
Mustafa, Khaled A., Bouzidi, Mohamed-Khalil, Schlauch, Christian, Gazar, Ahmad, Klein, Nadja, Reichardt, Joerg, Alonso-Mora, Javier
Abstract
Dense traffic is inherently interactive. The ego vehicle and surrounding agents continuously influence each other's reactions, making "what-if" reasoning essential for safe and efficient driving. To enable such an active interaction-aware behavior, we propose a planning framework that integrates an ego-conditioned generative autoregressive prediction model within Model Predictive Path Integral (MPPI) control. The generative prediction model outputs stochastic, multi-modal predictions of surrounding agents conditioned on each of the ego's considered future actions. A nested sampling scheme enables tractable evaluation of expected cost and collision risk under the induced distribution. This formulation allows the ego to actively probe how different candidate actions shape the interaction outcomes and to identify actions that reduce ambiguity in uncertain interactions. Closed-loop simulations demonstrate improved safety and efficiency compared to conventional predict-then-plan and passive interaction-aware approaches.
Chinese Translation
密集交通本质上是互动的。自我车辆与周围代理之间不断相互影响反应,使得“假设”推理对于安全和高效驾驶至关重要。为了实现这种主动的交互感知行为,我们提出了一种规划框架,该框架将自我条件生成自回归预测模型集成到模型预测路径积分(MPPI)控制中。生成预测模型输出基于自我车辆考虑的未来动作的随机多模态预测。嵌套采样方案使得在诱导分布下对期望成本和碰撞风险进行可处理的评估成为可能。这种公式化允许自我车辆主动探测不同候选动作如何塑造交互结果,并识别减少不确定交互中模糊性的动作。闭环仿真表明,与传统的先预测后规划和被动交互感知方法相比,安全性和效率得到了改善。
cs.RO / 7 / 2608.21402

Selective Cross-View Consistency for World Action Models: Held-Out Viewpoint Robustness Without Test-Time Camera Information

世界动作模型的选择性跨视图一致性:无需测试时相机信息的保留视点鲁棒性
Huang, Bingqi, Wei, Bingchuan, Cai, Yingkai, Wang, Zhaokui
Abstract
World action models (WAMs) jointly denoise future video frames and robot actions, and the video prior is expected to generalize their control. Camera viewpoint change remains one of their hardest perturbation axes. We study a question specific to this model class: when training with same-state cross-view image pairs, on which output coordinates should a consistency loss be imposed? The WAM denoising target mixes view-covariant coordinates, namely the predicted future scene, with view-invariant coordinates, namely the action chunk, future proprioception, and value. We show that consistency applied to the covariant block is provably harmful, shrinking legitimate view-specific content to a fraction $1/(1+4\lambda)$ of its true value, and we verify this shrinkage law in controlled experiments. Selective cross-view consistency (SCVC) therefore constrains only the invariant block, requires no camera labels, extrinsics, depth, or view synthesis at training or test time, and leaves the deployment interface unchanged. We introduce a carve-and-hold-out evaluation protocol on the LIBERO-Plus camera track that separates a distribution-matched ceiling from genuine interpolation and extrapolation to held-out viewpoints, with a matched pair-trained control isolating the effect of the consistency term from pair exposure. On held-out orbital viewpoints beyond the training envelope, SCVC improves closed-loop success over the matched control by 12.2 points (95% CI [7.4, 17.0]; +15.5, CI [11.7, 19.4], under an independent second seed) -- an effect two further camera axes replicate -- while interpolation within the envelope shows no gain in either seed (-1.2 and -4.3 points) and in-distribution competence is preserved (-0.6, -0.2). We also report a cross-backbone audit showing that published camera-robustness numbers are confounded by wrist-camera pose stability.
Chinese Translation
世界动作模型(WAMs)共同去噪未来视频帧和机器人动作,视频先验被期望能够推广其控制能力。相机视点变化仍然是它们面临的最困难的扰动轴之一。我们研究了一个特定于该模型类别的问题:在使用同状态跨视图图像对进行训练时,应该在输出坐标的哪些位置施加一致性损失?WAM去噪目标将视图协变坐标(即预测的未来场景)与视图不变坐标(即动作块、未来本体感知和价值)混合。我们证明,对协变块施加一致性是有害的,它将合法的视图特定内容缩小到其真实值的 $1/(1+4eta)$,并在控制实验中验证了这一缩小规律。因此,选择性跨视图一致性(SCVC)仅约束不变块,不需要相机标签、外部参数、深度或视图合成,在训练或测试时保持部署接口不变。我们在LIBERO-Plus相机轨道上引入了一种切割与保留的评估协议,该协议将分布匹配的上限与真正的插值和外推到保留视点进行分离,匹配对训练的控制隔离了一致性项的影响与对的曝光。在超出训练范围的保留轨道视点上,SCVC在闭环成功率上比匹配控制提高了12.2个百分点(95% CI [7.4, 17.0]; +15.5, CI [11.7, 19.4], 在独立的第二种种子下)——这一效果在另外两个相机轴上得到了复制——而在范围内的插值则没有任何增益(-1.2和-4.3个百分点),且在分布内的能力得以保持(-0.6, -0.2)。我们还报告了一项跨骨干网络审计,显示已发布的相机鲁棒性数据受到腕部相机姿态稳定性的干扰。
cs.RO / 8 / 2608.21404

Tolerance-Dependent Inspection Disagreement Between a Fixed CMM and a Portable Articulated-Arm CMM

固定坐标测量机与便携式关节臂坐标测量机之间的公差依赖性检验差异
Ahsan, Md Manjurul, Samadi, Hamidreza, Raman, Shivakumar
Abstract
Fixed coordinate measuring machines (CMMs) and portable articulated-arm CMMs are often assigned to the same inspection task, but their nominal accuracy specifications do not show whether a change of instrument will preserve the disposition of a part. The question is not simply how far the two results differ, but whether that difference crosses the tolerance boundary. We examined this issue with recorded measurements of cylindrical, cubic, and spherical features under nominal 20 {\deg}C and 30 {\deg}C conditions. Repeated records and two roughness profiles without sufficient acquisition information were removed, leaving six dimensional and four form profiles. For each dimensional feature, the distances of the two system means from nominal define the exact tolerance interval in which the systems receive opposite direct labels. The fixed-CMM stream was approximately 11.2 {\mu}m higher than the articulated-arm stream at both conditions. All four form profiles fell on opposite sides of the recorded 10 {\mu}m upper limit. The dimensional disagreement intervals also overlapped strongly; their mean widths were 6.573 {\mu}m at 20 {\deg}C and 4.995 {\mu}m at 30 {\deg}C. The results clarify why an average difference between instruments is not, by itself, a measure of substitution risk. The proposed tolerance map identifies the feature-tolerance combinations for which instrument choice can change the recorded inspection label and, therefore, where a controlled equivalence study and a task-specific uncertainty budget are needed before substitution.
Chinese Translation
固定坐标测量机(CMM)和便携式关节臂CMM通常被分配到相同的检验任务中,但它们的名义精度规格并未显示更换仪器是否会保持零件的状态。问题不仅在于两个结果之间的差异有多大,还在于这种差异是否超出了公差边界。我们通过在名义温度为20°C和30°C条件下记录的圆柱形、立方体和球形特征的测量数据来研究这个问题。去除了重复记录和缺乏足够采集信息的两个粗糙度轮廓,最终保留了六个尺寸特征和四个形状轮廓。对于每个尺寸特征,两个系统均值与名义值的距离定义了系统获得相反直接标签的确切公差区间。在这两种条件下,固定CMM的测量结果比关节臂CMM高出约11.2微米。所有四个形状轮廓均位于记录的10微米上限的两侧。尺寸差异区间也有很强的重叠;在20°C时其平均宽度为6.573微米,在30°C时为4.995微米。结果澄清了为什么仪器之间的平均差异本身并不能作为替代风险的衡量标准。所提出的公差图识别了仪器选择可能改变记录检验标签的特征-公差组合,因此,在替代之前需要进行受控等效研究和特定任务的不确定性预算。
cs.RO / 9 / 2608.21407

Mamba-based Selective State Space Modeling Improves the Accuracy-Complexity Tradeoff of SmolVLA Vision-Language-Action Experts

基于Mamba的选择性状态空间建模提升SmolVLA视觉-语言-动作专家的准确性与复杂度权衡
Mohsen, Farida, Elkaffash, Thowayba, Qazani, Mohammad Reza Chalak, Mabrok, Mohamed, Meskin, Nader, Safa, Ali
Abstract
Vision-language-action (VLA) models face a crucial tradeoff between their task success rate and the policy-call frequency. Executing a single action per inference ($N=1$) enables accurate robot control but comes at the cost of huge compute time overheads, making real-time implementation infeasible. On the other hand, executing longer action horizons before replanning ($N\gg1$) reduces compute complexity, but inevitably degrades the system's success rate. In order to improve the VLA accuracy-complexity tradeoff, this paper investigates Mamba's selective state-space modeling as an alternative to causal self-attention within the action expert of the popular SmolVLA model, widely used as a reference model for its highly accurate yet low complexity nature. We evaluate both the Mamba- and Transformer-based experts on the widely-adopted LIBERO benchmark suites across three execution horizons $N\!\in\!\{1,25,50\}$, respectively corresponding to high, moderate and low compute complexities. Our results remarkably show that the advantage of the Mamba expert increases with the execution horizon, indicating significant success retention under long execution horizons $N = 50$ and $N = 25$. When $N = 50$ actions are executed before replanning (i.e., corresponding to feasible real-time deployment), the Mamba expert outperforms the Transformer baseline by $7.8\%$. In addition, when $N = 25$ actions are executed before replanning, our Mamba expert outperforms the Transformer baseline by $3.7\%$. Finally, under per-action replanning ($N=1$), our Mamba variant matches the Transformer-based mean success rate while significantly reducing the overall model parameter complexity by $24\%$ thanks to Mamba's compute-efficient nature.
Chinese Translation
视觉-语言-动作(VLA)模型在任务成功率与策略调用频率之间面临关键权衡。每次推理执行单一动作($N=1$)能够实现精确的机器人控制,但代价是巨大的计算时间开销,导致实时实现不可行。另一方面,在重新规划前执行较长的动作序列($N\gg1$)可以降低计算复杂度,但不可避免地降低系统的成功率。为改善VLA的准确性与复杂度权衡,本文研究了Mamba的选择性状态空间建模,作为流行的SmolVLA模型动作专家中因果自注意力(causal self-attention)的替代方案。SmolVLA因其高准确率且低复杂度被广泛用作参考模型。我们在广泛采用的LIBERO基准套件上,针对三个执行时长$N\in\{1,25,50\}$(分别对应高、中、低计算复杂度)评估了基于Mamba和Transformer的专家模型。结果显著表明,随着执行时长的增加,Mamba专家的优势愈发明显,表明在长执行时长$N=50$和$N=25$下成功率保持显著。具体而言,当在重新规划前执行$N=50$个动作(即可行的实时部署场景)时,Mamba专家比Transformer基线高出7.8%的成功率;当执行$N=25$个动作时,Mamba专家比Transformer基线高出3.7%。最后,在每动作重新规划($N=1$)的情况下,我们的Mamba变体在匹配Transformer基线平均成功率的同时,得益于Mamba计算效率高的特性,整体模型参数复杂度降低了24%。
cs.RO / 10 / 2608.21410

Position: Robot Privacy as Embodied Boundary Work. Connecting Capabilities, Contexts, and Design Responses in Everyday Robotics

位置:机器人隐私作为具身边界工作。在日常机器人中连接能力、情境和设计响应
He, Liwen, Zhang, Shuning, Zhang, Chengwen, Yi, Xin, Yu, Chun, Jeung, Jihong, Tong, Xin
Abstract
Robots are increasingly entering everyday environments where privacy is shaped not only by data practices, but also by spatial, bodily, social, and relational boundaries. Their embodied capabilities allow them to reshape these boundaries through situated action, challenging privacy framings centered on data flows, interface settings, or one-time consent. Prior work has examined robot privacy through sensing, data collection, telepresence, transparency, consent, bystander awareness, and multi-stakeholder governance. Building on this work, we propose embodied boundary privacy as a capability-by-context framing for examining how physically present robots may reshape privacy boundaries in situated interaction. Specifically, this framing organizes privacy risks across seven robot capabilities and five deployment contexts, asking how embodied capabilities enable boundary crossings and how situated contexts shape who is affected, how these crossings are interpreted, and when they become contested. We use this perspective to outline design and research implications for embodied privacy mechanisms, including boundary checkpoints, viewpoint-aware sensing control, remote-presence disclosure, object- and body-level access rules, constraints on socially persuasive privacy influence, and local interruption rights. We encourage HRI research, design, and governance to treat robot movement, orientation, proximity, object access, remote presence, and social expression as privacy-relevant actions whose meaning depends on context.
Chinese Translation
机器人正越来越多地进入日常环境,在这些环境中,隐私不仅受到数据实践的影响,还受到空间、身体、社会和关系边界的塑造。它们的具身能力使得机器人能够通过情境行动重塑这些边界,挑战以数据流、界面设置或一次性同意为中心的隐私框架。先前的研究通过传感、数据收集、远程呈现、透明度、同意、旁观者意识和多方治理等方面考察了机器人隐私。在此基础上,我们提出了具身边界隐私作为一种能力-情境框架,用于研究物理存在的机器人如何在情境互动中重塑隐私边界。具体而言,该框架将隐私风险组织在七种机器人能力和五种部署情境中,探讨具身能力如何促进边界跨越,以及情境如何影响受影响者、这些跨越的解读方式以及何时引发争议。我们利用这一视角概述了具身隐私机制的设计和研究启示,包括边界检查点、视角感知的传感控制、远程存在披露、物体和身体级别的访问规则、对社会劝说性隐私影响的限制以及本地干预权。我们鼓励人机交互(HRI)研究、设计和治理将机器人运动、方向、接近度、物体访问、远程存在和社会表达视为与隐私相关的行为,其意义依赖于情境。
cs.RO / 11 / 2608.21411

Social Graph Mamba: Forecasting Pedestrian Movements Based on Social Context

社交图Mamba:基于社交背景的行人运动预测
Nguyen, Hong-Son, Liu, Yen-Chen
Abstract
Forecasting pedestrian motion has always been fundamental for autonomous navigation in crowded environments. While attention-based methods achieve strong performance, they suffer from quadratic computational complexity in modeling social interactions, limiting scalability. Additionally, the existing methods often achieve high accuracy on prediction benchmarks at the individual level, but fail to fully capture the natural movement behaviors of crowds in real-world scenarios, particularly group structures. In this study, we propose Social Graph Mamba (SGM), a novel architecture that replaces attention-based social reasoning with Selective State Space Models (SSMs) operating on dynamically constructed interaction graphs. SGM introduces a dynamic interaction graph with social triplet factorization to decompose crowd interactions sequentially, and a community-aware module to effectively discover group structures via differentiable MinCut optimization and conditions both the embedding space and multi-modal decoder on group membership. Our experiments on standard benchmarks (ETH/UCY, SDD) demonstrate competitive performance with linear sequence complexity compared to quadratic attention-based methods. We further validate SGM in physical robot experiments by integrating predicted trajectories into a Social Force Model (SFM) for real-world implementation.
Chinese Translation
行人运动预测一直是拥挤环境中自主导航的基础。尽管基于注意力的方法在性能上表现出色,但在建模社交互动时存在二次计算复杂度的问题,限制了其可扩展性。此外,现有方法在个体层面的预测基准上通常能达到高准确率,但未能充分捕捉现实场景中人群的自然运动行为,特别是群体结构。在本研究中,我们提出了社交图Mamba(Social Graph Mamba,SGM),这是一种新颖的架构,用选择性状态空间模型(Selective State Space Models,SSMs)替代基于注意力的社交推理,后者在动态构建的互动图上运行。SGM引入了一个动态互动图,通过社交三元组分解逐步分解人群互动,并通过可微分的最小割优化有效发现群体结构,同时对嵌入空间和多模态解码器施加群体成员资格的条件。我们在标准基准(ETH/UCY,SDD)上的实验表明,与基于注意力的二次方法相比,SGM在序列复杂度上具有线性竞争性能。我们进一步通过将预测轨迹整合到社交力模型(Social Force Model,SFM)中,在物理机器人实验中验证了SGM的实际应用。
cs.RO / 12 / 2608.21414

RiskWorld: Object-Centric Latent World Modeling for Autonomous Driving Risk Identification

RiskWorld:面向自主驾驶风险识别的对象中心潜在世界建模
Li, Jingzheng, Ge, Yufei, Mao, Qianren, Chen, Zhijun, Li, Bing, Peng, Xingyu, Zhang, Baochang, Liu, Xianglong
Abstract
Autonomous driving risk identification aims to determine which observed object is likely to become safety-critical to the ego vehicle. Existing approaches typically predict scene-level accidents, infer risk objects indirectly from ego behavior, or apply geometric checks after trajectory forecasting, without directly using predicted ego--object relations for risk-source localization. We propose RiskWorld, an object-centric latent world model that identifies risk from the imagined evolution of each candidate relative to the ego vehicle. RiskWorld combines pretrained predictive video representations with structured ego--object histories, contextualizes observed interactions, and rolls relation-aware object states into the future using RSSM-style latent dynamics. It decodes the rollout into object-level risk scores, supported by auxiliary future-relation and temporal-risk predictions. Inference uses only observations up to the current time, while logged futures provide training supervision. On RiskBench, RiskWorld achieves the best overall F1 of 63.0\% and the lowest false-alarm rate of 2.1\%. Further analyses show that the learned rollout captures the evolution of object-level risk before critical events, while RiskWorld's selections preserve planning-critical information under filtered observation.
Chinese Translation
自主驾驶风险识别旨在确定哪些观察到的对象可能对自我车辆的安全构成威胁。现有方法通常预测场景级事故,间接推断风险对象通过自我行为,或在轨迹预测后应用几何检查,而未直接利用预测的自我-对象关系进行风险源定位。我们提出了RiskWorld,一种对象中心的潜在世界模型,通过想象每个候选对象相对于自我车辆的演变来识别风险。RiskWorld结合了预训练的预测视频表示与结构化的自我-对象历史,情境化观察到的交互,并使用RSSM风格的潜在动态将关系感知的对象状态向未来滚动。它将滚动解码为对象级风险评分,并辅以辅助的未来关系和时间风险预测。推断仅使用截至当前时间的观察,而记录的未来提供训练监督。在RiskBench上,RiskWorld实现了最佳的整体F1值63.0%和最低的误报率2.1%。进一步分析表明,学习到的滚动捕捉了关键事件之前对象级风险的演变,而RiskWorld的选择在过滤观察下保留了规划关键的信息。
cs.RO / 13 / 2608.21416

Operational digital twin clinics enable task-based evaluation of embodied AI

操作性数字双胞胎诊所实现了基于任务的具身人工智能评估
Wu, Xinyuan, Zhang, Jingrao, Xu, Mengdi, Chu, Henry K., He, Mingguang, Shi, Danli
Abstract
Embodied artificial intelligence (AI) must be tested in the clinical environments where it will operate, but building realistic, robot-testable settings is costly and difficult to scale. Here we show that routine clinic images can be transformed into operational digital twins for task-based evaluation of embodied AI. Using 39 ophthalmic clinic scenes, we converted single photographs into editable, simulator-ready environments and assessed reconstruction quality, room-scale geometry, mesh grounding, multi-robot feasibility, perturbation sensitivity and closed-loop policy performance. The reconstructed scenes preserved workspace structure, while local editing enabled controlled device reconfiguration. Device meshes, collision proxies and semantic anchors converted visual reconstructions into contact-aware simulation scenes. Across three robot embodiments, shared task targets showed different patterns of reachability and contact feasibility. Small device translations and rotations produced task-specific changes in contact margins that were not captured by visual similarity alone. Digital-twin trajectories also supported local policy learning and closed-loop evaluation. These findings establish operational validity as a key principle for clinical digital twins and provide an intermediate layer between offline development and physical deployment of embodied AI in healthcare.
Chinese Translation
具身人工智能(AI)必须在其将要操作的临床环境中进行测试,但构建现实的、可供机器人测试的环境既昂贵又难以扩展。在此,我们展示了如何将常规诊所图像转化为用于具身AI基于任务评估的操作性数字双胞胎。通过使用39个眼科诊所场景,我们将单张照片转换为可编辑的、适合模拟器使用的环境,并评估了重建质量、房间规模几何、网格基础、多机器人可行性、扰动敏感性和闭环策略性能。重建的场景保留了工作空间结构,而局部编辑则实现了受控设备重新配置。设备网格、碰撞代理和语义锚点将视觉重建转换为感知接触的模拟场景。在三种机器人具身中,共享任务目标显示出不同的可达性和接触可行性模式。小范围的设备平移和旋转产生了任务特定的接触边界变化,而这些变化仅通过视觉相似性无法捕捉。数字双胞胎轨迹还支持局部策略学习和闭环评估。这些发现确立了操作有效性作为临床数字双胞胎的关键原则,并为具身AI在医疗保健中的离线开发与物理部署之间提供了一个中间层。
cs.RO / 14 / 2608.21420

Evaluating Human and LLM-Generated Thematic Analysis in HRI for Vulnerable Populations: A Comparative and Ethical Analysis

评估人类与大型语言模型生成的主题分析在脆弱群体人机交互中的应用:比较与伦理分析
Markelius, Alva, Dogan, Fethiye Irmak, Bailey, Julie, Gunes, Hatice
Abstract
Thematic analysis (TA) has long been regarded as an inherently human, reflexive, and interpretive process. However, the extent to which LLM-generated TA is appropriate for Human-Robot Interaction (HRI) research involving vulnerable populations remains largely unexamined and raises critical questions about validity and ethics, particularly in sensitive research contexts. This paper presents a comparative study of human- and LLM-generated TA in an HRI context with a focus on vulnerable populations. We evaluate both objective and semantic agreement between human- and LLMgenerated themes, and examine whether observed divergences reflect systematic interpretive patterns with ethical significance. Our analysis investigates whether LLM-generated TA risks marginalising or misrepresenting the experiences of vulnerable participants, with implications for researchers employing LLM-assisted TA in HRI.
Chinese Translation
主题分析(Thematic Analysis, TA)长期以来被视为一种固有的人类、反思性和解释性的过程。然而,LLM(大型语言模型)生成的主题分析在涉及脆弱群体的人机交互(Human-Robot Interaction, HRI)研究中的适用性仍然未得到充分检验,并引发了关于有效性和伦理的关键问题,尤其是在敏感的研究背景下。本文呈现了一项在人机交互背景下对人类与LLM生成的主题分析的比较研究,重点关注脆弱群体。我们评估了人类与LLM生成主题之间的客观和语义一致性,并考察观察到的差异是否反映出具有伦理意义的系统性解释模式。我们的分析探讨了LLM生成的主题分析是否存在边缘化或误表述脆弱参与者经历的风险,这对采用LLM辅助主题分析的人机交互研究者具有重要影响。
cs.RO / 15 / 2608.21433

The Setting of IMU Parameters in Kalman Filtering-based Information Fusion

基于卡尔曼滤波的信息融合中IMU参数的设置
Hu, Qiang, Zou, Yanhua, Huo, Shuaiyi, Ge, Haibo, Ouyang, Wei
Abstract
The setting or tuning of specifications for the inertial measurement unit (IMU) is tricky in sensor fusion. The underneath conundrum is caused by the fact that the working condition of IMU is more complex than the stationary calibration scenario. Since the noises and biases instabilities calibrated under static condition cannot accommodate other cases, the effective tuning of IMU parameters largely hinges on the experience or profound understanding of the system. In the current work, the setting method of IMU parameters based on Allan variance calibration is delved into within the Kalman filtering framework. Specifically, the relationship between the power sepctral density and Allan variance is leveraged in formulating the process uncertainty in continuous-time filtering. Two typical IMU-based sensor fusion systems are considered to show the feasibility and effectiveness of this parameter setting process.
Chinese Translation
在传感器融合中,惯性测量单元(IMU)规格的设置或调整是一项复杂的任务。这一困境的根源在于IMU的工作条件比静态校准场景更为复杂。由于在静态条件下校准的噪声和偏差不稳定性无法适应其他情况,因此IMU参数的有效调整在很大程度上依赖于对系统的经验或深刻理解。在当前的研究中,探讨了基于艾伦方差(Allan variance)校准的IMU参数设置方法,并将其置于卡尔曼滤波框架内。具体而言,利用功率谱密度与艾伦方差之间的关系来制定连续时间滤波中的过程不确定性。考虑了两个典型的基于IMU的传感器融合系统,以展示该参数设置过程的可行性和有效性。
cs.RO / 16 / 2608.21440

Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics

Geo-VLA:通过地图语义的内化实现几何感知的视觉-语言-动作规划
Chen, Ran, Ren, Jiaxing, Zhang, Zhikun, Hou, Yunhao, Zhuo, Junbao, Zou, Bochao
Abstract
Vision-language-action (VLA) models have advanced end-to-end autonomous driving by leveraging foundation models for semantic reasoning and long-tail generalization. However, their planning performance remains limited in complex driving environments because image-only representations inadequately capture planning-relevant road geometry and topology. In this paper, we propose Geo-VLA, a plug-and-play framework that enhances VLA models by learning geometry-aware visual representations. During training, Geo-VLA internalizes geometric map semantics to strengthen road-structure representations, while requiring no HD maps or additional lane information during inference. To support this approach, we introduce Geo-QA, a geometry-focused question-answering dataset that injects road geometry into vision-language representations through contrastive learning and instruction tuning. Experiments on NAVSIM v1 demonstrate that Geo-VLA consistently improves VLA planners with distinct action-generation architectures, achieving 92.1 PDMS and establishing a new state-of-the-art among single-camera VLA planners.
Chinese Translation
视觉-语言-动作(VLA)模型通过利用基础模型进行语义推理和长尾泛化,推动了端到端自主驾驶的发展。然而,由于仅依赖图像的表示无法充分捕捉与规划相关的道路几何和拓扑信息,其规划性能在复杂驾驶环境中仍然有限。本文提出了Geo-VLA,一个即插即用的框架,通过学习几何感知的视觉表示来增强VLA模型。在训练过程中,Geo-VLA内化几何地图语义,以增强道路结构的表示,同时在推理时不需要高清地图或额外的车道信息。为支持这一方法,我们引入了Geo-QA,一个以几何为中心的问题回答数据集,通过对比学习和指令调优将道路几何信息注入视觉-语言表示。NAVSIM v1上的实验表明,Geo-VLA在具有不同动作生成架构的VLA规划器中始终提升性能,达到了92.1 PDMS,并在单摄像头VLA规划器中建立了新的最先进水平。
cs.RO / 17 / 2608.21441

Constructing Predictive Surgical Path for AI-based Capsulorhexis Skill Transfer

构建基于人工智能的囊袋切开技能转移的预测手术路径
Ahmadi, Mohammad Javad, Taghirad, Hamid D.
Abstract
Automated training of surgeons is one of the most crucial factors that significantly minimize surgical training risks and expenses. With recent advances in artificial intelligence (AI) knowledge and available data from various surgeries, AI's involvement in surgical training is becoming very promising. It is recommended that at the early stages of AI development, it interferes in the surgery as a third agent alongside the trainer. As trust in AI increases, this process will lead to an AI agent acting as a trainer in the future. The first phase in which AI can intervene in the training process is to suggest an improved surgical path to the trainer. A platform must be constructed in the first step, to accomplish this task and to enhance the movement path of trainee surgeons. This paper introduces this platform along with an annotated capsulorhexis surgery dataset called the ARAS-Farabi dataset. In this research, a deep convolutional neural network is pre-trained with JIGSAWS and ARAS-Farabi surgical datasets that can extract surgical skill characteristics from surgery tool tip motion data. The proposed platform develops a reference model from the feature space of an expert surgeon's movement trajectory and proposes an improved path to enhance the skill of a novice surgeon. An optimization with two loss functions is utilized to create a path that raises the skill level of the novice surgeon's path while simultaneously predicting and preserving his/her intent. The results of this study reveal that, with the assistance of an AI agent, the trainee surgeon's movement path can be enhanced by at least 20 percent while maintaining his intentional objective. In addition to the recommended deep network, various tangible indicators have also been developed in this research to verify the level of trainee improvement.
Chinese Translation
外科医生的自动化培训是显著降低外科培训风险和费用的关键因素之一。随着人工智能(AI)知识的最新进展以及来自各种手术的可用数据,AI在外科培训中的参与变得非常有前景。建议在AI发展的早期阶段,它作为第三方介入手术,与培训师并肩工作。随着对AI信任的增加,这一过程将导致未来AI代理作为培训师的角色。AI可以介入培训过程的第一阶段是向培训师建议改进的手术路径。为实现这一目标,必须首先构建一个平台,以增强培训外科医生的运动路径。本文介绍了这一平台,并提供了一个名为ARAS-Farabi的数据集,包含注释的囊袋切开手术数据。在本研究中,深度卷积神经网络使用JIGSAWS和ARAS-Farabi外科数据集进行预训练,能够从手术工具尖端运动数据中提取外科技能特征。所提出的平台从专家外科医生的运动轨迹特征空间中开发参考模型,并提出改进路径以提升新手外科医生的技能。利用两个损失函数的优化来创建一条路径,既提高新手外科医生的技能水平,又同时预测和保留其意图。本研究的结果表明,在AI代理的帮助下,培训外科医生的运动路径可以提高至少20%,同时保持其意图目标。除了推荐的深度网络外,本研究还开发了各种可量化指标,以验证培训改进的水平。
cs.RO / 18 / 2608.21533

Model-Free Adaptive Parameter Tuning for Efficient Multi-Robot Warehouse Operations

无模型自适应参数调优用于高效的多机器人仓库操作
Tokekar, Pratap, Benosman, Mouhacine, Chandan, Rahul, Barbosa, Alexandre Ormiga Galvao, Caldara, Michael, Durham, Joseph W.
Abstract
Robotic Fulfillment Centers (FCs) store inventory on shelves (pods) arranged in dense blocks. Retrieving a target pod that is buried deep in a block requires moving obstructing pods out of the way (i.e., digout). Multi-robot planners use parameterized cost functions to control digout behavior, producing a spectrum of strategies: at one extreme, obstructing pods are sent to other blocks (using more robots in travel lanes); at the other, pods are shuffled within the block (avoiding lane congestion but increasing extraction time). Each point on this spectrum has different downstream consequences for floor congestion and throughput. The optimal operating point depends on the specific facility configuration and shifts with operational conditions such as varying station demand and congestion patterns, making offline tuning impractical. We present an adaptive parameter tuning framework based on Extremum Seeking Control (ESC) that continuously adjusts planner parameters in response to measured throughput. ESC performs model-free optimization by perturbing parameters with sinusoidal dither signals and correlating perturbations with performance changes to estimate gradients, making it robust to the multi-minute delayed effects and credit assignment challenges inherent in large FC operations. Simulation studies demonstrate that the adaptive policy improves upon fixed policies across several conditions. We observe an improvement in throughput by an average of 5.0% across map and robot fleet size variations, and by 8.4% under dynamic operating conditions. This work eliminates manual parameter provisioning and enables real-time adaptation, providing a self-tuning paradigm for FC storage operations.
Chinese Translation
机器人履行中心(Fulfillment Centers, FCs)将库存存储在密集块状排列的货架(pods)上。检索深埋在块中的目标货架需要将阻碍的货架移开(即,挖掘)。多机器人规划器使用参数化成本函数来控制挖掘行为,产生一系列策略:在一端,阻碍的货架被送往其他块(在运输通道中使用更多机器人);在另一端,货架在块内进行重新排列(避免通道拥堵但增加提取时间)。这一系列策略的每个点对地面拥堵和吞吐量都有不同的下游影响。最佳操作点依赖于特定的设施配置,并随着操作条件的变化而变化,例如不同的站点需求和拥堵模式,使得离线调优变得不切实际。我们提出了一种基于极值寻求控制(Extremum Seeking Control, ESC)的自适应参数调优框架,该框架根据测量的吞吐量持续调整规划器参数。ESC通过对参数施加正弦扰动信号进行无模型优化,并将扰动与性能变化相关联以估计梯度,从而使其对大型FC操作中固有的多分钟延迟效应和信用分配挑战具有鲁棒性。仿真研究表明,自适应策略在多种条件下优于固定策略。我们观察到在地图和机器人车队规模变化下,吞吐量平均提高了5.0%,在动态操作条件下提高了8.4%。这项工作消除了手动参数配置,实现了实时适应,为FC存储操作提供了一种自调节范式。
cs.RO / 19 / 2608.21550

GOLEM: Modular Humanoid Autonomy Towards Electric Vehicle Battery Disassembly

GOLEM:面向电动汽车电池拆解的模块化类人自主系统
Conway, Max, Xie, William, Devaraj, Allen, Zhang, Yutong, Pudasaini, Niraj, Feit, Mateo, Abid, Adam, Allen, Zachary, Liu, Chen, Tan, Xuan, Lavering, Jensen, Chen, Jason, Antieau, Lyle, Von Pischke, Anthony, Roncone, Alessandro, Sunberg, Zachary, Correll, Nikolaus
Abstract
Disassembling end-of-life electric vehicle (EV) battery packs is dull and dangerous work, performed almost entirely by humans. We present GOLEM (Generalized Open Library of Embodied Modules), an end-to-end, open-source system architecture for EV battery disassembly with the Unitree H1-2 humanoid robot in which walking, manipulation, dynamic stability, navigation, and spatial memory are independent modules with abstract interfaces, so that methods are easily developed, interchanged, and compared. GOLEM is deployed as a Docker-based ROS 2 abstraction in which MuJoCo and IsaacLab digital twins expose interfaces matching the physical robot. GOLEM's composability and per-module customization enable development and demonstration of humanoid EV battery disassembly, from simulation to reality. GOLEM provides fair comparison between humanoid modules, enabling evaluation as a capability ladder, in which one module is characterized at a time and added as a rung: LiDAR-inertial navigation places the robot within 13.0cm of a 6m goal; a learned standing controller recovers from external disturbances that sampling-based lower-body MPC does not; and grasping loosened fasteners from a real Hyundai Ioniq 5 pack degrades from 97% tethered to 87% free-standing to 37% under navigation-induced pose variance. Source code is available at the project page https://golem-humanoid.github.io
Chinese Translation
拆解报废电动汽车(EV)电池组是一项枯燥且危险的工作,几乎完全由人类完成。我们提出了GOLEM(通用化具身模块开放库),这是一个端到端的开源系统架构,用于与Unitree H1-2类人机器人一起进行电动汽车电池拆解。在该架构中,行走、操控、动态稳定性、导航和空间记忆都是独立模块,具有抽象接口,从而使得方法的开发、互换和比较变得更加容易。GOLEM作为基于Docker的ROS 2抽象进行部署,其中MuJoCo和IsaacLab数字双胞胎暴露出与物理机器人匹配的接口。GOLEM的可组合性和每个模块的定制化使得从仿真到现实的类人电动汽车电池拆解的开发和演示成为可能。GOLEM提供了类人模块之间的公平比较,能够将评估作为能力阶梯,其中每次只对一个模块进行特征描述并作为一个阶梯添加:LiDAR惯性导航使机器人在6米目标内的距离为13.0厘米;一个学习的站立控制器能够从外部干扰中恢复,而基于采样的下肢模型预测控制(MPC)则无法做到;从真实的现代Ioniq 5电池组中抓取松动的紧固件的成功率从97%(有绳)降至87%(独立站立),在导航引起的姿态变化下降至37%。源代码可在项目页面 https://golem-humanoid.github.io 获取。
cs.RO / 20 / 2608.21554

Model-Based Reinforcement Learning for Heterogeneous Multi-Robot Task Assignment Under Distribution Shifts

基于模型的强化学习在分布变化下的异构多机器人任务分配
Garces, Daniel, Castro, Sara, Haimovich, Adrian, Crowe, Byron, Gil, Stephanie
Abstract
Heterogeneous multi-robot service systems must assign requests to compatible robots, construct feasible schedules, and adapt as new tasks arrive online. Historical data can help anticipate future demand, but relying too heavily on inaccurate predictions can degrade performance under distribution shifts. We develop a prediction-aware adaptive rollout framework for heterogeneous multi-robot task assignment with scheduled and real-time requests. The problem is formulated as a finite-horizon stochastic dynamic program incorporating robot-task compatibility, ordered service requirements, routing constraints, service windows, and end-of-horizon return requirements. The proposed policy evaluates current assignments using sampled future request scenarios while restricting immediate commitments to requests already observed. To enable online use, the framework combines pruned candidate controls, wait actions, and an interaction-aware base policy for efficient future-cost estimation. Robustness to forecast error is provided by adaptively reweighting predicted requests based on recent prediction mismatch and selectively re-optimizing assigned but unstarted requests. We also introduce a historical-data-driven procedure for selecting the heterogeneous fleet composition before deployment. In a case study using real nursing-task requests from hospital inpatient floors, the proposed approach achieves near-complete service and reduces serviced-request wait times relative to reactive, token-passing, prediction-positioning, and myopic greedy baselines, with the largest improvements in tail-delay metrics.
Chinese Translation
异构多机器人服务系统必须将请求分配给兼容的机器人,构建可行的调度,并在新任务在线到达时进行适应。历史数据可以帮助预测未来需求,但过于依赖不准确的预测可能会在分布变化下降低性能。我们开发了一种预测感知的自适应回滚框架,用于处理异构多机器人任务分配,包括计划请求和实时请求。该问题被表述为一个有限时间范围的随机动态规划,结合了机器人-任务兼容性、服务顺序要求、路线约束、服务时间窗口和时间结束回报要求。所提出的策略使用采样的未来请求场景评估当前分配,同时限制对已观察请求的即时承诺。为了支持在线使用,该框架结合了修剪的候选控制、等待动作和交互感知的基础策略,以实现高效的未来成本估计。通过根据近期预测不匹配自适应地重新加权预测请求,并选择性地重新优化已分配但未开始的请求,提供了对预测误差的鲁棒性。我们还引入了一种基于历史数据的程序,用于在部署前选择异构车队组成。在使用来自医院住院楼的真实护理任务请求的案例研究中,所提出的方法实现了几乎完全的服务,并相对于反应式、令牌传递、预测定位和短视贪婪基线减少了服务请求的等待时间,在尾延迟指标上取得了最大的改善。
cs.RO / 21 / 2608.21572

Betting for Sim-to-Real Performance Certificates

针对模拟到现实性能证书的投注
Chen, Yujia, Weng, Bowen
Abstract
Consider a typical test of a robot system: one observes a sequence of outcomes concerning some aspect of interest (crash or no crash, tracking error, time to completion), and reports a mean (crash risk, average error, mean time to completion) and, more importantly, an interval guaranteed to contain that mean at a prescribed confidence, referred to as a performance certificate. Given expensive real-world trials, the sample size is therefore small, and the certificate is often loose. Now consider the same procedure, except that before each real outcome is revealed, the operator ``peeks'' at a large bank of simulated results, and places a bet on where the real outcome will land. As the real outcomes settle the bets, the operator gains or loses wealth. One's ``trust'' over simulators also shifts within the portfolio. This paper develops that idea into a sim-to-real betting certificate framework with three contributions: (i) An algorithm that links a scalable bank of simulators to effective bets, and the accumulated betting wealth to the certificate. (ii) A proof that the returned certificate is anytime valid, covering the true mean with the prescribed probability, using any simulator bank. (iii) The guaranteed wealth-regret bounds yield configuration principles for the proposed algorithm and simulator bank design to deliver tight certificates. Experiments across synthetic distributions and real-world robot tests, covering both replayed standardized testing outcomes and online runtime evaluation, show the proposed method narrows the certificate by $51.6\%\pm16\%$ against classic and state-of-the-art baselines, and by $32.26\%\pm8\%$ in the extremely limited-sample regime ($\leq30$ samples).
Chinese Translation
考虑一个典型的机器人系统测试:观察与某个感兴趣方面相关的一系列结果(碰撞或未碰撞、跟踪误差、完成时间),并报告一个均值(碰撞风险、平均误差、平均完成时间)以及更重要的一个区间,保证在规定的置信度下包含该均值,称为性能证书。由于现实世界的试验成本高昂,样本量通常较小,因此证书往往较为宽松。现在考虑相同的程序,只是在每个真实结果揭示之前,操作员“窥视”一大批模拟结果,并对真实结果的落点进行投注。随着真实结果的确定,投注结果将影响操作员的财富。操作员对模拟器的“信任”也会在投资组合中发生变化。本文将这一思想发展为一个模拟到现实的投注证书框架,提出了三项贡献:(i)一个将可扩展的模拟器库与有效投注链接起来的算法,以及将累积的投注财富与证书关联起来;(ii)一个证明,返回的证书在任何时刻都是有效的,以规定的概率覆盖真实均值,使用任何模拟器库;(iii)保证的财富-遗憾界限为所提算法和模拟器库设计提供了配置原则,以提供紧凑的证书。在合成分布和真实机器人测试中的实验,涵盖了重放的标准化测试结果和在线运行时评估,显示所提方法相较于经典和最先进的基线,证书缩小了$51.6\% ext{±}16\\%$,在样本极为有限的情况下($ ext{≤}30$样本)缩小了$32.26\\% ext{±}8\\%$。
cs.RO / 22 / 2608.21592

Force/Torque-Based Kinematic Adaptation for Robotic Manipulation Tasks

基于力/扭矩的机器人操作任务运动学适应
Henshaw, Carl Glen
Abstract
Contact-rich robotic manipulation requires an accurate model of the kinematic relationship between a robot's joints and the task features it senses. This relationship is rarely known exactly: it changes with each tool the robot picks up and shifts, sometimes almost instantaneously, as contact modes change --- especially for multi-fingered hands that make and break contact at points that are not exactly prescribed, as in full-hand grasping. This paper develops an adaptive scheme that estimates that relationship online, using only joint-angle sensing and a wrist-mounted force/torque sensor, with no exteroceptive measurement of the tool tip. We derive a provably stable kinematic update law that identifies the kinematics of an unknown tool from force/torque feedback alone, and prove stability of both the rigid case and the case with a compliance controller as an inner loop. We show that identification is confined to the directions the motion excites --- so that, for example, a tool's length is unobservable under a rigid insertion push, while a compliant loop's passive yielding partially excites it; and that with a second-order admittance the compliant certificate holds unconditionally in continuous time. We also pose the combined control and estimation problem as a Quadratic Program (QP): the formulation yields the prediction term of the update law exactly but, instructively, cannot reproduce the tracking adaptation term. We validate the scheme in simulation on a peg-in-hole insertion. This work is the first step in a research program aimed at factoring manipulation learning into a task policy which can be learned in isolation of the robot, for instance by reinforcement learning, and an adaptive kinematic component that adapts online to the particular robot, hand, or tool in use.
Chinese Translation
接触丰富的机器人操作需要准确的模型来描述机器人关节与其感知的任务特征之间的运动学关系。这种关系很少被准确知晓:它会随着机器人拾取和移动的每个工具而变化,有时在接触模式变化时几乎是瞬时的——尤其是在多指手的情况下,接触点并不是完全规定的,如在全手抓取中。本文开发了一种自适应方案,在线估计这种关系,仅使用关节角度传感和安装在手腕上的力/扭矩传感器,而无需对工具尖端的外部测量。我们推导出一个可证明稳定的运动学更新法则,该法则仅通过力/扭矩反馈识别未知工具的运动学,并证明了刚性情况和具有合规控制器作为内环的情况的稳定性。我们展示了识别仅限于运动激发的方向——例如,在刚性插入推力下,工具的长度是不可观察的,而合规环的被动屈服部分激发了它;并且在二阶导纳下,合规证书在连续时间内无条件成立。我们还将控制与估计的组合问题表述为一个二次规划(Quadratic Program, QP):该公式精确地给出了更新法则的预测项,但有启发性的是,无法重现跟踪适应项。我们在插销入孔的仿真中验证了该方案。这项工作是一个研究计划的第一步,旨在将操作学习分解为可以独立于机器人学习的任务策略,例如通过强化学习,以及一个在线适应于特定机器人、手或工具的运动学组件。
cs.RO / 23 / 2608.21620

Why Personalization Matters: Cross-Subject Challenges in EMG-IMU-based HRI Activity Recognition

个性化的重要性:基于EMG-IMU的人机交互活动识别中的跨主题挑战
Carminati, Ruan Rithelle Chagas de Faria, Braglia, Giovanni, Biagiotti, Luigi, Rohrich, Ronnier Frates, de Oliveira, Andre Schneider, Hartmann, Mikael Nedel, Lazzaretti, André Eugenio
Abstract
This paper investigates wearable-based recognition of human activities and gestures to support Human-Robot Interaction (HRI) in object-handover and assembly-like scenarios. Electromyography (EMG) and Inertial Measurement Unit (IMU) signals were collected using a Myo armband, culminating in a novel dataset introduced as MAGIC-HRI (Multimodal Activity, Gesture and Intention Collection) with a large taxonomy of 53 movement classes, including Brazilian Sign Language (LIBRAS) numbers (0-9), hand gestures, object/tool handover actions (pick up/give/hold), tool-manipulation tasks, and generic assembly/idle motions, collected from 11 participants with 10 samples per class (530 samples per participant). Signals are segmented by detecting muscle activation via an EMG energy envelope, then processed using sliding windows; time- and frequency-domain features are extracted. Multiple classical classifiers are tuned via cross-validated grid search, with Random Forest as the strongest baseline. A Leave-One-Subject-Out (LOSO) protocol reveals a large generalization gap, indicating substantial subject dependence. A personalized adaptation experiment suggests that injecting a small number of samples from a new user can markedly improve recognition. Overall, the study contributes a broad, HRI-driven multimodal dataset, a rigorous evaluation emphasizing generalization, and practical evidence that personalization is likely required for robust deployment in practical HRI.
Chinese Translation
本文研究了基于可穿戴设备的人类活动和手势识别,以支持在人机交互(HRI)中的物体交接和组装场景。使用Myo臂带收集了肌电图(EMG)和惯性测量单元(IMU)信号,最终形成了一个新颖的数据集MAGIC-HRI(多模态活动、手势和意图收集),该数据集包含53个运动类别的大型分类法,包括巴西手语(LIBRAS)数字(0-9)、手势、物体/工具交接动作(拾取/给予/保持)、工具操作任务以及通用的组装/闲置动作,数据来自11名参与者,每个类别收集10个样本(每位参与者530个样本)。通过检测肌肉激活的EMG能量包络对信号进行分段,然后使用滑动窗口进行处理;提取时间域和频率域特征。通过交叉验证网格搜索调整多个经典分类器,其中随机森林作为最强基线。留一主题外(LOSO)协议显示出较大的泛化差距,表明存在显著的主题依赖性。个性化适应实验表明,从新用户中注入少量样本可以显著提高识别效果。总体而言,本研究贡献了一个广泛的、以HRI驱动的多模态数据集,强调泛化的严格评估,以及个性化在实际HRI中强健部署所需的实证证据。
cs.RO / 24 / 2608.21631

OpenSCvx: An Open-Source Modular and Extensible Nonlinear Trajectory Planning Package

OpenSCvx:一个开源的模块化和可扩展的非线性轨迹规划包
Hayner, Christopher R., Norris, Griffin J., Spada, Fabio, Uzun, Samet, Mittal, Avi, Acıkmese, Behcet, Leung, Karen
Abstract
Trajectory optimization computes dynamically feasible motions that enable autonomous systems to accomplish complex tasks while satisfying operational and environmental constraints. This tutorial presents OpenSCvx, an open-source Python framework that bridges the gap between high-level problem specification and efficient numerical optimization. Rather than requiring users to derive solver-specific mathematical formulations, OpenSCvx provides a symbolic modeling interface that automatically constructs and solves trajectory optimization problems from modular descriptions of objectives, dynamics, and constraints. Beyond simplifying problem formulation, OpenSCvx supports (i) continuous-time constraint modeling, (ii) temporal and logical specifications, (iii) automatic vectorization for scalable and batched optimization, and (iv) a modular architecture that enables new algorithms, models, and solver backends to be incorporated with minimal effort. These capabilities allow researchers and practitioners to rapidly prototype, solve, and extend state-of-the-art trajectory optimization methods.
Chinese Translation
轨迹优化计算动态可行的运动,使自主系统能够在满足操作和环境约束的情况下完成复杂任务。本教程介绍了OpenSCvx,一个开源的Python框架,它弥合了高层次问题规范与高效数值优化之间的差距。OpenSCvx不要求用户推导特定求解器的数学公式,而是提供了一个符号建模接口,能够从目标、动态和约束的模块化描述中自动构建和解决轨迹优化问题。除了简化问题的表述外,OpenSCvx还支持(i)连续时间约束建模,(ii)时间和逻辑规范,(iii)用于可扩展和批量优化的自动向量化,以及(iv)一个模块化架构,使得新的算法、模型和求解器后端可以以最小的努力被纳入。这些功能使研究人员和从业者能够快速原型化、解决和扩展最先进的轨迹优化方法。
cs.RO / 25 / 2608.21676

Lifelong Robot Recomposition via Persistent Categorical Modeling for Unified Task-Driven Co-Design, Verification, and Planning

通过持久性类别建模实现终身机器人重组以统一任务驱动的共同设计、验证和规划
Swanbeck, Steven, Pryor, Mitch
Abstract
Robotic systems are traditionally designed and deployed in static configurations, with assumptions made at design-time becoming immutable constraints during runtime. This design-then-deploy paradigm produces performant systems under narrow operating conditions, but renders robots brittle when qualities of themselves, their tasks, or their environments unexpectedly change. We address this challenge with a compositional framework that formalizes robotic systems as abstract circuits within a strict symmetric monoidal category, in which design and runtime composition of hardware, software, and behavior are synthesized simultaneously via an SMT-based solver, with monoidal functors projecting the system into lifecycle-specific views and free symbolic variables simultaneously solving for parameters and entire component specifications within larger compositions. This persistent model also supports queries a long-lived system needs beyond plan existence across its entire lifecycle, including mapping Pareto fronts over candidate compositions, diagnosing why a composition has become infeasible, finding its minimal restoration, and reconfiguring with limited change to the deployed system. We evaluate against official implementations of optimal numeric, stream-based, and SMT-based planners all measured onboard a deployed robot and demonstrate the approach end-to-end in a search-and-rescue scenario in which the robot recognizes when it has become unfit and synthesizes and assumes new holistic configurations to restore operation. We release our solver and supporting software open-source.
Chinese Translation
机器人系统传统上是在静态配置中设计和部署的,设计时所做的假设在运行时变成不可变的约束。这种先设计后部署的范式在狭窄的操作条件下产生高性能系统,但当机器人自身、其任务或环境的特性意外变化时,导致机器人变得脆弱。我们通过一个组合框架来应对这一挑战,该框架将机器人系统形式化为严格对称单元范畴中的抽象电路,其中硬件、软件和行为的设计与运行时组合通过基于 SMT 的求解器同时合成,单元函子将系统投影到生命周期特定的视图中,自由符号变量同时求解更大组合中的参数和整个组件规格。该持久性模型还支持长期系统在整个生命周期中所需的查询,超越计划存在的范围,包括对候选组合的帕累托前沿进行映射、诊断组合为何变得不可行、找到其最小恢复方案,以及在对部署系统进行有限更改的情况下重新配置。我们与官方实现的最优数值、基于流的和基于 SMT 的规划器进行评估,所有评估均在已部署的机器人上进行,并在一个搜索与救援场景中展示了该方法的端到端应用,其中机器人识别到其已变得不适合,并合成和假设新的整体配置以恢复操作。我们将我们的求解器和支持软件开源发布。
cs.RO / 26 / 2608.21685

In-Situ Reconstruction of the International Space Station Using 3D Gaussian Splatting and Astrobee

使用3D高斯点云和Astrobee对国际空间站进行原位重建
Kim, Hudson, Soussan, Ryan, Coltin, Brian, Kam, Jordan
Abstract
This article presents a novel 3D reconstruction and mapping of the interior of the International Space Station (ISS) using 3D Gaussian Splatting (3DGS). Using existing grayscale images from the Astrobee free-flying robot dataset, we construct a full 3D splat of the ISS' Kib\=o or Japanese Experiment Module (JEM). 3DGS has in recent years shown promise in providing novel view synthesis of scenes captured from many images or videos, this article applies this approach to human spaceflight systems. We compare our 3DGS architecture to existing methods such as Nerfacto and TensoRF and show that reconstruction improves the state-of-the-art in both scene quality and rendering speed. We show that with as little as 500 in-situ images, a high-fidelity map can be constructed using Astrobee's Navigation Camera (NavCam) during free-flight in the JEM. These reconstructions could enable free-flyers to rapidly create and update interior maps for intra-vehicular habitats like the ISS.
Chinese Translation
本文提出了一种新颖的国际空间站(ISS)内部3D重建和映射方法,采用3D高斯点云(3D Gaussian Splatting, 3DGS)。利用来自Astrobee自由飞行机器人数据集的现有灰度图像,我们构建了ISS的Kibō或日本实验模块(Japanese Experiment Module, JEM)的完整3D点云。近年来,3DGS在从多张图像或视频中捕获场景的视图合成方面显示出良好的前景,本文将该方法应用于人类航天系统。我们将我们的3DGS架构与现有方法如Nerfacto和TensoRF进行了比较,结果表明重建在场景质量和渲染速度上均提升了当前的技术水平。我们展示了仅使用500张原位图像,就可以利用Astrobee的导航相机(Navigation Camera, NavCam)在JEM的自由飞行中构建高保真地图。这些重建结果可以使自由飞行器快速创建和更新国际空间站等内部栖息地的地图。
cs.RO / 27 / 2608.21699

Towards insect-like distributed proprioception in actuators and appendages for flapping-wing insect-scale aerial robots

朝着类昆虫的分布式本体感觉在拍打翅膀的昆虫尺度空中机器人中的应用
Hedrick, Alexander, Gupta, Arvind, Jayaram, Kaushik
Abstract
Modern flapping-wing insect-scale air vehicles display agility similar to that of their insect counterparts; however, these impressive maneuvers are only possible with off-board sensors like optical tracking cameras. In this manuscript, we introduce two embedded proprioceptive sensors for insect-scale aerial robots: thin film piezoelectric polymers integrated directly into a driving actuator and a pitching hinge which track stroke and pitch angle, respectively. We fabricate the aforementioned size-agnostic mechanically intelligent structures (sensor-actuator, sensor-flexure) using laminate stack fabrication methods. Chirp experiments with our sensors integrated into an insect-size flapping-wing robot show accurate tracking of stroke (RMSE = 0.44 deg) and pitch (RMSE = 2.44 deg) angles in the relevant frequency range. As the first step towards demonstrating the utility of these sensors for enabling numerous onboard autonomy applications, including closed-loop wingbeat control and sensor fusion with existing insect-scale sensor suites for more accurate proprioception and localization, we show one application for each sensor. The proprioceptive hinge enables collision detection, reducing the chance of permanent damage if the robot's wing collides with an object. The proprioceptive actuator enables asynchronous flapping, which is hypothesized to increase adaptability and efficiency in insects and robots alike. A microrobot equipped with our proprioceptive actuator allows us to test these hypotheses with potential for improving flapping aerial robot performance. We foresee proprioceptive sensors having an important role in progressing both the fields of insect-scale aerial robots and robo-physics due to the bio-inspired nature and high integration level of our sensors.
Chinese Translation
现代拍打翅膀的昆虫尺度空中飞行器展现出与其昆虫同类相似的灵活性;然而,这些令人印象深刻的机动性仅依赖于诸如光学跟踪摄像头等外部传感器。在本文中,我们介绍了两种用于昆虫尺度空中机器人的嵌入式本体感觉传感器:直接集成到驱动执行器中的薄膜压电聚合物和分别跟踪拍击角度和俯仰角度的俯仰铰链。我们采用层压堆叠制造方法制造了上述尺寸无关的机械智能结构(传感器-执行器,传感器-柔性结构)。与昆虫尺寸的拍打翅膀机器人集成的传感器的啁啾实验显示,在相关频率范围内,拍击(均方根误差 RMSE = 0.44 度)和俯仰(均方根误差 RMSE = 2.44 度)角度的跟踪精度很高。作为展示这些传感器在实现多种机载自主应用(包括闭环拍打控制和与现有昆虫尺度传感器套件的传感器融合以实现更准确的本体感觉和定位)中的实用性的第一步,我们展示了每种传感器的一个应用。本体感觉铰链能够实现碰撞检测,降低机器人翅膀与物体碰撞时造成永久性损坏的风险。本体感觉执行器则能够实现异步拍打,这被假设为提高昆虫和机器人在适应性和效率方面的能力。配备本体感觉执行器的微型机器人使我们能够测试这些假设,并有潜力改善拍打空中机器人的性能。我们预见到本体感觉传感器将在推动昆虫尺度空中机器人和机器人物理学领域的发展中发挥重要作用,因其生物启发的特性和高集成度。
cs.RO / 28 / 2608.21735

Safety-Critical Bilateral Teleoperation for Omnidirectional Aerial Manipulation Using Force-Sensorless Haptic Feedback

基于无力传感器触觉反馈的全向空中操作的安全关键双向遥操作框架
Kim, Yubin, Lee, Jinwoo, You, Yongjun, Kim, H. Jin, Byun, Jeonghyun
Abstract
This paper presents a safety-critical bilateral teleoperation framework for omnidirectional aerial manipulators that integrates visual and force-sensorless haptic wrench feedback. Unlike existing approaches that either rely on onboard force/torque sensors or use model-dependent wrench estimates, which may become unreliable under model uncertainties or induce unintended feedback during free-flight, our method implements a hierarchical safety filter based on control barrier functions to avoid such limitations. The safety filter, being the key contribution, explicitly accounts for tracking errors arising from physical interaction between the aerial manipulator and its surroundings while enforcing thrust limits, a factor overlooked despite its critical importance for flight safety. This safety filter adjusts the command from the operator to ensure safe and stable aerial manipulation and avoid motor saturation. The adjustment made by the filter is mapped to haptic feedback, which is intuitive to the operator and conveys information on physical interaction and impending motor saturation. By actual experiments with a hexarotor-based omnidirectional aerial manipulator, we demonstrate that the proposed method avoids haptic feedback during free-flight, provides directionally consistent feedback under physical interaction, and can be operated for diverse manipulative tasks. Moreover, an ablation study further shows that the saturation filter improves interaction stability by explicitly preventing motor saturation and informing the operator of corrective actions.
Chinese Translation
本文提出了一种针对全向空中操纵器的安全关键双向遥操作框架,该框架集成了视觉和无力传感器的触觉扭矩反馈。与现有方法依赖于机载力/扭矩传感器或使用依赖模型的扭矩估计(在模型不确定性下可能变得不可靠或在自由飞行中引发意外反馈)不同,我们的方法实施了一种基于控制障碍函数的分层安全过滤器,以避免这些局限性。安全过滤器作为关键贡献,明确考虑了空中操纵器与其周围环境之间物理交互所产生的跟踪误差,同时强制执行推力限制,这一因素在飞行安全中至关重要但常常被忽视。该安全过滤器调整操作者的指令,以确保安全和稳定的空中操作,并避免电机饱和。过滤器所做的调整被映射到触觉反馈中,这对操作者直观,并传达有关物理交互和即将发生的电机饱和的信息。通过与基于六旋翼的全向空中操纵器的实际实验,我们证明了所提出的方法在自由飞行中避免触觉反馈,在物理交互下提供方向一致的反馈,并能够用于多种操作任务。此外,消融研究进一步表明,饱和过滤器通过明确防止电机饱和并告知操作者纠正措施来提高交互稳定性。
cs.RO / 29 / 2608.21740

CounterAlign: Counterfactual Supervision for Vision-Language-Action Models

CounterAlign:用于视觉-语言-动作模型的反事实监督
Kondoh, Haru, Ota, Kei, Kanezaki, Asako, Wu, Yueh-Hua
Abstract
Vision-Language-Action (VLA) models are typically trained with behavior cloning (BC) on expert demonstrations. However, BC provides only positive supervision for expert actions, without explicit negative supervision indicating which actions are instruction-inconsistent or otherwise inappropriate. Reinforcement learning (RL) can provide such corrective signals, but often relies on externally specified rewards or curated non-expert data, both of which are costly to obtain in robotics. We show that offline RL for VLA models need not rely on curated non-expert trajectories: successful expert demonstrations alone can be transformed into dense corrective supervision through instruction relabeling. Specifically, by pairing expert actions with mismatched alternative instructions, we synthesize counterfactual instruction-observation-action tuples from the dataset and combine them with adversarial discriminator training to learn an instruction-grounded reward model for offline RL, without collecting additional rollouts or annotations. On the robustness-focused LIBERO-PRO benchmark, our method improves robustness to object position and task perturbations over a strong state-of-the-art baseline. It also outperforms competitive baselines in real-robot experiments on the TX-G2 (compatible with AGIBot G2). More broadly, our results suggest that, for data-constrained VLA learning, extracting denser supervision from each demonstration can complement collecting additional data.
Chinese Translation
视觉-语言-动作(VLA)模型通常通过行为克隆(BC)在专家演示上进行训练。然而,BC 仅提供对专家动作的正向监督,而没有明确的负向监督来指示哪些动作与指令不一致或不适当。强化学习(RL)可以提供这种纠正信号,但通常依赖于外部指定的奖励或策划的非专家数据,这在机器人领域获取成本高昂。我们展示了 VLA 模型的离线 RL 不必依赖策划的非专家轨迹:成功的专家演示可以通过指令重新标注转化为密集的纠正监督。具体而言,通过将专家动作与不匹配的替代指令配对,我们从数据集中合成反事实的指令-观察-动作元组,并将其与对抗性鉴别器训练结合,以学习用于离线 RL 的指令基础奖励模型,而无需收集额外的回合或注释。在以稳健性为重点的 LIBERO-PRO 基准上,我们的方法在物体位置和任务扰动方面提高了对强大最先进基线的稳健性。在 TX-G2(与 AGIBot G2 兼容)的真实机器人实验中,它也优于竞争基线。更广泛地说,我们的结果表明,对于数据受限的 VLA 学习,从每个演示中提取更密集的监督可以补充收集额外数据。
cs.RO / 30 / 2608.21778

Vision Guided Target Conditioned Control for Autonomous Excavation

基于视觉引导的目标条件控制用于自主挖掘
Zhao, Shuai, Pan, Ji-An, Li, Junwei, Tang, Xun, Xi, Fansen, Xu, Qing, Li, Keqiang, Wang, Jianqiang
Abstract
Autonomous excavation requires an intelligent control system that can convert spatial work intent into coordinated bucket motion under contact-rich soil interaction. This paper presents a target-conditioned intelligent control framework for autonomous excavation in a physics-based deformable-soil simulation workflow. An image-aligned target mask serves as a visual spatial command for the desired digging region, while a mask-conditioned Action Chunking Transformer maps multi-view RGB observations, proprioception, and the target mask to temporally extended joystick commands. To reduce target-ignoring behavior, demonstrations are organized with paired-condition supervision, where the same or closely matched scene is demonstrated with different target masks and corresponding action chunks. The framework is evaluated through both a diagnostic manipulation task and an excavation simulation benchmark with single-scoop and sequential pile-clearing protocols. In manipulation, target success is 4\% for no-condition ACT, 63\% for non-paired mask-conditioned ACT, and 96\% for paired-condition mask-conditioned ACT. In sequential pile clearing, paired-condition mask-conditioned ACT removes 76.8\% of the pile versus 27.4\% and 15.7\% for the two baselines, with 91.0\% human-normalized efficiency. The results show that visual target conditioning, paired demonstration structure, and action-chunk control form a practical cyber-physical simulation pipeline for excavator automation.
Chinese Translation
自主挖掘需要一个智能控制系统,能够将空间工作意图转换为在接触丰富的土壤交互下的协调铲斗运动。本文提出了一种针对自主挖掘的目标条件智能控制框架,该框架应用于基于物理的可变形土壤仿真工作流程。图像对齐的目标掩模作为期望挖掘区域的视觉空间指令,而掩模条件的动作分块变换器(Action Chunking Transformer)将多视角RGB观测、自我感知和目标掩模映射为时间扩展的操纵杆指令。为了减少忽视目标的行为,演示通过配对条件监督进行组织,其中相同或相似的场景使用不同的目标掩模和相应的动作分块进行演示。该框架通过诊断操作任务和单铲及顺序清堆协议的挖掘仿真基准进行评估。在操作中,无条件ACT的目标成功率为4\%,非配对掩模条件ACT为63\%,配对条件掩模条件ACT为96\%。在顺序清堆中,配对条件掩模条件ACT去除了76.8\%的堆积物,而两个基线分别为27.4\%和15.7\%,人类标准化效率为91.0\%。结果表明,视觉目标条件、配对演示结构和动作分块控制形成了一个实用的网络物理仿真管道,用于挖掘机自动化。
cs.RO / 31 / 2608.21879

Vision-Guided Morphing Quadcopter for Multi-Geometry Payload Transport through Narrow Passages

视觉引导的变形四旋翼无人机用于通过狭窄通道运输多几何形状的有效载荷
Sahu, Aashish, Hari, Shriram, Kumar, R. Prasanth
Abstract
Aerial payload transport using multirotor unmanned aerial vehicles is challenging because payload geometry, contact interaction, grasp stability, flight control, and narrow-passage traversal are strongly coupled during pickup and transport. Object-specific grippers often cannot adapt their footprint or grasp geometry when the payload shape or passage width changes. This paper presents a vision-guided morphing quadcopter for multi-geometry payload transport through narrow passages. The proposed platform uses four hybrid arm-leg structures that function as both landing supports and grasping members. A centrally placed actuator drives a tendon-based morphing mechanism, enabling all four arms to synchronously retract or expand for object grasping, footprint reduction, and post-transport release. Onboard vision estimates the payload geometry and passage width, while endpoint force feedback is used to confirm grasp contact during payload engagement. A phase-wise mission planner, PID-based flight stabilization, and morphology-adaptive grasp controller are implemented in a MuJoCo simulation environment. The framework is evaluated using box, cylindrical, and spherical payloads, representing flat-faced, rolling-curved, and fully curved contact conditions. Across the three cases, the simulated system completes the pickup-transport-release sequence with a maximum RMS position error of 0.31 m, a final drop-zone error below 0.18 m, a compact grasp footprint of 0.09-0.21 m2, and a footprint reduction of 75.0-89.7 percent. The results demonstrate that a single-actuator morphing quadcopter can adapt its grasp footprint for the transport of payloads with different geometries while reducing its overall footprint for narrow-passage traversal.
Chinese Translation
使用多旋翼无人机进行空中有效载荷运输具有挑战性,因为在拾取和运输过程中,有效载荷几何形状、接触交互、抓取稳定性、飞行控制和狭窄通道穿越之间存在强耦合。特定物体的抓取器通常无法在有效载荷形状或通道宽度变化时适应其接触面积或抓取几何形状。本文提出了一种视觉引导的变形四旋翼无人机,用于通过狭窄通道运输多几何形状的有效载荷。所提出的平台使用四个混合臂腿结构,既作为着陆支撑,又作为抓取部件。中央放置的执行器驱动基于腱的变形机制,使四个臂能够同步收缩或扩展,以实现物体抓取、接触面积减小和运输后释放。机载视觉系统估计有效载荷几何形状和通道宽度,同时使用末端力反馈确认有效载荷接触。在MuJoCo仿真环境中实现了分阶段任务规划器、基于PID的飞行稳定化和形态自适应抓取控制器。该框架使用盒形、圆柱形和球形有效载荷进行评估,代表平面接触、滚动曲面和完全曲面接触条件。在这三种情况下,仿真系统以最大均方根位置误差0.31米完成拾取-运输-释放序列,最终投放区误差低于0.18米,紧凑的抓取接触面积为0.09-0.21平方米,接触面积减少75.0-89.7%。结果表明,单执行器变形四旋翼无人机能够适应不同几何形状有效载荷的抓取接触面积,同时在狭窄通道穿越时减少其整体接触面积。
cs.RO / 32 / 2608.21894

An Interpretable Deep Learning Framework for Material Perception and Classification from Multisensory Tactile Data

一种可解释的深度学习框架用于多感官触觉数据的材料感知与分类
Zou, Li, Hogendoorn, Dave, Vardar, Yasemin
Abstract
Human tactile perception relies on complex multisensory cues. Yet the relationship between tactile signals and perceptual representations remains poorly understood, limiting the integration of touch in digital environments and human-like robotic perception. To address this gap, we developed a computational framework comprising three interconnected deep learning models that map multisensory touch data to material perception, without relying on hand-crafted features. The models represent progressively different routes from tactile signals to material class: from low-level interaction signals to perceptual attribute distributions (Model 1), from predicted attribute distributions to material classification (Model 2), and directly from tactile signals to material categories, bypassing intermediate representations (Model 3). By combining deep learning with Integrated Gradients, the framework achieved high accuracy while offering interpretability, revealing which sensory modalities most strongly drive its decisions. Our results show that deep learning can approach near-perfect material classification when unconstrained by intermediate perceptual stages, but matching human-like performance is harder once those stages are modeled explicitly. Notably, thermal cues emerged as particularly informative across all models, providing robust signals for material differentiation. The results offer a computational account of how tactile signals lead to material perception and show how interpretable deep learning can both approach human-level performance and reveal cues that robotic and haptic systems need to incorporate.
Chinese Translation
人类的触觉感知依赖于复杂的多感官线索。然而,触觉信号与感知表征之间的关系仍然不够清晰,这限制了触觉在数字环境和类人机器人感知中的整合。为了解决这一问题,我们开发了一个计算框架,该框架由三个相互关联的深度学习模型组成,能够将多感官触觉数据映射到材料感知上,而无需依赖手工设计的特征。这些模型代表了从触觉信号到材料类别的逐步不同路径:从低级交互信号到感知属性分布(模型1),从预测的属性分布到材料分类(模型2),以及直接从触觉信号到材料类别,绕过中间表征(模型3)。通过将深度学习与集成梯度(Integrated Gradients)相结合,该框架在提供可解释性的同时实现了高准确率,揭示了哪些感官模态对其决策影响最大。我们的结果表明,当不受中间感知阶段的限制时,深度学习可以接近完美的材料分类,但一旦这些阶段被明确建模,匹配类人表现就变得更加困难。值得注意的是,热感线索在所有模型中都表现出特别的信息量,为材料区分提供了稳健的信号。这些结果为触觉信号如何导致材料感知提供了计算解释,并展示了可解释的深度学习如何既能接近人类水平的表现,又能揭示机器人和触觉系统需要整合的线索。
cs.RO / 33 / 2608.21899

CIDER: Continual Interactive Distillation for Embodied Reinforcement Learning

CIDER:用于具身强化学习的持续互动蒸馏
Li, Houlin, Xu, Minghui, Xu, Guo, Du, Xuan, Yan, Xiaohan, Wang, Chun, Yan, Yuxiang, Yang, Shukai, Liu, Yongcheng, Shan, Wei, Yao, Maoqing
Abstract
Human-in-the-loop real-world reinforcement learning enables rapid acquisition of effective robotic manipulation policies for individual tasks, often within tens of minutes. Yet it remains unclear how to extend this paradigm to continual learning, where a single policy must acquire new skills without losing previously learned behaviors. Existing real-world continual learning methods do not explicitly constrain prior behaviors, leading to severe catastrophic forgetting. We introduce Continual Interactive Distillation for Embodied Reinforcement Learning (CIDER), a continual reinforcement learning framework that freezes the accumulated historical policy as a teacher before learning each new task and interleaves task learning with distillation-based retention. We further introduce gradient routing to separate the gradients used for acquiring new tasks from those used for preserving prior behaviors. We evaluate our method with a single shared actor on six real-world household and industrial manipulation tasks. Interactive Distillation maintains high measured success on previously learned tasks across our six-task real-robot sequence while acquiring each new task in 10 to 20 minutes, whereas every baseline forgets at least one previous task. Additional ablations reveal the key design choices that govern the tradeoff between stability and plasticity in real-world continual reinforcement learning.
Chinese Translation
人机协作的现实世界强化学习能够在短短几十分钟内快速获取针对单个任务的有效机器人操作策略。然而,如何将这一范式扩展到持续学习仍然不明确,其中单一策略必须在不失去先前学习行为的情况下获取新技能。现有的现实世界持续学习方法并未明确约束先前的行为,导致严重的灾难性遗忘。我们提出了用于具身强化学习的持续互动蒸馏(CIDER),这是一种持续强化学习框架,在学习每个新任务之前,将累积的历史策略冻结为教师,并将任务学习与基于蒸馏的保留交替进行。我们进一步引入梯度路由,以将用于获取新任务的梯度与用于保留先前行为的梯度分开。我们在六个现实世界的家庭和工业操作任务上使用单个共享演员评估我们的方法。互动蒸馏在我们的六任务真实机器人序列中保持了对先前学习任务的高成功率,同时在10到20分钟内获取每个新任务,而每个基线方法至少遗忘一个先前任务。额外的消融实验揭示了在现实世界持续强化学习中影响稳定性与可塑性之间权衡的关键设计选择。
cs.RO / 34 / 2608.22008

Stakeholder Insights for Designing In-Home Social Robots for Dementia Disorientation Detection and Caregiver-Aware Intervention

为设计家庭社交机器人以检测痴呆症患者的定向障碍和关注照顾者的干预提供利益相关者洞察
Akinrintoyo, Emmanuel, Salomons, Nicole
Abstract
Disorientation is a common and distressing experience for people living with dementia. It often manifests as confusion about time, place, or personal context. These episodes can increase anxiety, agitation, and safety risks, especially for persons with dementia (PwDs) who live independently at home. While assistive technologies have explored reminders, monitoring, and activity support, little research has explored how socially assistive robots can support the detection and management of disorientation in everyday living contexts. Hence, disorientation detection and intervention remain under-examined as socio-technical challenges. We conducted 14 semi-structured interviews with dementia caregivers and practitioners, including family and professional caregivers, occupational therapists, mental health practitioners, dementia nurse practitioners, and well-being and technology leads. The findings reveal that recurrent and fluctuating. It often emerges through behavioural cues such as repeated questioning, inappropriate activity timing and disrupted daily routines. Caregivers described orientation as emotionally charged, and direct correction may increase distress. The participants were generally receptive to robotic assistance when framed as supportive rather than corrective. Based on the insights, we identify essential design implications for socially assistive robots that provide context-aware orientation support, integrate into daily routines and support caregivers through timely escalation.
Chinese Translation
定向障碍是生活在痴呆症患者中常见且令人痛苦的经历。它通常表现为对时间、地点或个人背景的困惑。这些情节可能增加焦虑、烦躁和安全风险,尤其是对于那些独立居住在家中的痴呆症患者(PwDs)。尽管辅助技术已经探索了提醒、监测和活动支持,但很少有研究探讨社交辅助机器人如何在日常生活环境中支持定向障碍的检测和管理。因此,定向障碍的检测和干预仍然作为社会技术挑战未得到充分研究。我们与14位痴呆症照顾者和从业者进行了半结构化访谈,包括家庭和专业照顾者、职业治疗师、心理健康从业者、痴呆症护士从业者以及福祉和技术负责人。研究结果表明,定向障碍是反复和波动的,通常通过行为线索表现出来,如重复提问、不当的活动时机和日常生活的中断。照顾者描述定向障碍为情感充沛的过程,直接纠正可能增加痛苦。参与者普遍对将机器人辅助视为支持而非纠正持开放态度。基于这些洞察,我们确定了社交辅助机器人设计的关键启示,这些机器人能够提供上下文感知的定向支持,融入日常生活,并通过及时升级支持照顾者。
cs.RO / 35 / 2608.22028

Design of a Human-Assistance Robot System with Contextual Action Recognition

具有情境动作识别的人机协作机器人系统设计
Ergogo, Amanuel, Zielińska, Teresa
Abstract
This paper presents a conceptual design for a proactive human assisting robot system capable of recognizing human activities and responding proactively. The system leverages contextual human activity recognition to interpret human actions across diverse contexts, while behavior trees are utilized to define dynamic and interpretable robot behaviors. We outline the system architecture, incorporating contextual human action recognition (HAR), behavior trees (BTs), and ROS, using the Spot robot platform as a representative example. We explain how HAR enables the robot to provide proactive assistance, discuss its limitations, and introduce methodologies for contextual HAR to address these limitations, thereby enhancing the robot's decision-making in complex human activity scenarios.
Chinese Translation
本文提出了一种主动人机协作机器人系统的概念设计,该系统能够识别人的活动并主动做出响应。该系统利用情境人类活动识别(HAR)来解释不同情境下的人类动作,同时采用行为树(BTs)来定义动态且可解释的机器人行为。我们概述了系统架构,结合了情境人类动作识别、行为树和机器人操作系统(ROS),并以Spot机器人平台作为代表性示例。我们解释了HAR如何使机器人提供主动帮助,讨论了其局限性,并介绍了应对这些局限性的情境HAR方法,从而增强机器人在复杂人类活动场景中的决策能力。
cs.RO / 36 / 2608.22033

DELTA: Deformable Elevation-Based Local Terrain Attention Encoder for Sparse-Terrain Quadrupedal Locomotion

DELTA:基于可变形高程的局部地形注意力编码器用于稀疏地形四足运动
Park, Sanghyun, Jung, Moonkyu, Hwangbo, Jemin
Abstract
Stable quadrupedal locomotion on sparse terrain requires selecting state-relevant terrain evidence for precise foot placement. Model-based foothold planners provide precise foothold selection but rely heavily on explicit model assumptions. Recent attention-based map encoding (AME) studies show that end-to-end reinforcement learning (RL) can learn implicit foothold guidance. However, the computational cost of dense AME encoding grows with map resolution, limiting its scalability to fine-grained sparse terrain. We propose DELTA, a Deformable Elevation-Based Local Terrain Attention encoder. DELTA predicts state-conditioned sampling locations, forms terrain evidence tokens from adaptive local elevation patches, and attends only to a fixed-size token set. With fixed sampling and patch settings, DELTA's encoder cost is independent of map resolution. Experiments show that DELTA achieves final traversal performance comparable to AME at the standard resolution while improving learning efficiency. This fixed encoder cost enables the use of higher-resolution terrain maps, improving traversal on fine-grained sparse terrain. DELTA also demonstrates strong generalization to unseen mixed evaluation courses composed of continuous and discrete terrain elements. Beyond simulation, DELTA demonstrates successful sim-to-real transfer on RAIBO2. Analysis of the learned sampling offsets and attention weights shows that DELTA samples steppable regions and attends to terrain evidence relevant to future touchdowns without foothold labels or attention supervision.
Chinese Translation
在稀疏地形上实现稳定的四足运动需要选择与状态相关的地形证据以确保精确的足部放置。基于模型的足部规划器提供了精确的足部选择,但严重依赖于明确的模型假设。最近的基于注意力的地图编码(AME)研究表明,端到端的强化学习(RL)能够学习隐式的足部引导。然而,密集的AME编码的计算成本随着地图分辨率的增加而增长,限制了其在细粒度稀疏地形上的可扩展性。我们提出了DELTA,一种基于可变形高程的局部地形注意力编码器。DELTA预测状态条件下的采样位置,从自适应局部高程补丁中形成地形证据标记,并仅关注固定大小的标记集。通过固定的采样和补丁设置,DELTA的编码器成本与地图分辨率无关。实验表明,DELTA在标准分辨率下实现的最终穿越性能可与AME相媲美,同时提高了学习效率。这种固定的编码器成本使得可以使用更高分辨率的地形地图,从而改善在细粒度稀疏地形上的穿越表现。DELTA还展示了对由连续和离散地形元素组成的未见混合评估课程的强泛化能力。超越仿真,DELTA在RAIBO2上成功实现了仿真到现实的转移。对学习到的采样偏移和注意力权重的分析表明,DELTA能够采样可步行区域,并关注与未来着陆相关的地形证据,而无需足部标签或注意力监督。
cs.RO / 37 / 2608.22035

Ludi${}_{\scriptscriptstyle 0.1}$: An Agentic System for Socially Intelligent Robots

Ludi${}_{ ext{0.1}}$: 一种用于社会智能机器人的自主系统
Chung, Wooseong, Cong, William, Dworakowski, Jakub, Ewer, Ethan, Guntara, Tri Wahyu, Jeong, Yeonwoo, Jiang, Tianchong, Kim, Chaewon, Kim, Hyunseo, Kim, Jinwoo, Kim, Jinyeon, Kim, Yea-Seul, Kunde, Jack, Lee, Kangwook, Lee, Sangheon, Nowak, Robert, Roh, Junha
Abstract
Robot foundation models have substantially advanced perception and control, but natural human-robot collaboration requires more than executing isolated commands. A robot must recognize ambiguity, maintain context across turns, communicate its intentions, and revise ongoing behavior as the user's intent changes. We present $\scriptstyle\mathsf{Ludi}_{\scriptscriptstyle 0.1}$, an agentic system for socially intelligent robots that integrates interactive speech, multimodal reasoning, memory, navigation, and learned manipulation. Its decision-making core is a fine-tuned vision-language model trained on multi-turn interaction traces spanning ambiguous requests, clarifications, corrections, interruptions, mixed social and task dialogue, and multi-step tasks. A purpose-built harness manages the model-tool interaction loop, while specialized navigation and manipulation policies execute physical skills. Ludi${}_{\scriptscriptstyle 0.1}$ demonstrates a practical path toward fluid human-robot collaboration today while producing the multimodal interaction traces needed to develop a more deeply integrated foundation model for robots and people.
Chinese Translation
机器人基础模型在感知和控制方面取得了显著进展,但自然的人机协作不仅仅需要执行孤立的指令。机器人必须能够识别模糊性,在对话中保持上下文,传达其意图,并在用户意图变化时调整其行为。我们提出了$ ext{Ludi}_{ ext{0.1}}$,这是一个用于社会智能机器人的自主系统,集成了互动语音、多模态推理、记忆、导航和学习的操作。其决策核心是一个经过精细调优的视觉-语言模型,训练于涵盖模糊请求、澄清、纠正、中断、混合社交与任务对话以及多步骤任务的多轮交互轨迹。一个专门构建的工具管理模型与工具之间的交互循环,而专门的导航和操作策略则执行物理技能。Ludi${}_{ ext{0.1}}$展示了实现流畅人机协作的实际路径,同时生成了开发更深度集成的机器人与人类基础模型所需的多模态交互轨迹。
cs.RO / 38 / 2608.22067

Inferring Action from Future Latent State for Robotic Manipulation

从未来潜在状态推断机器人操作的动作
Lei, Fenghao, Huang, Zhixiong, Yang, Long, Chen, Jiabao, Cheng, Jie, Huang, Peilin, Fu, Han, Li, Zhuo, Ren, Xiaoxue
Abstract
World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions. We argue that video generation is an unnecessary intermediate objective for world-action modeling. For robotic manipulation, the goal of a world model is not to reproduce how the world looks at every intermediate moment, but to predict the state that the world will reach after an action is executed. The intermediate frames only describe the visual transition between physical states, which consumes substantial model capacity and computation, but do not directly specify the physical outcome that the robot action is intended to produce. In this paper, we propose DELE-w0.5, which infers robot actions from predicted future states without relying on video generation. Concretely, DELE-w0.5 infers the action sequence from its corresponding compact future latent state. The future latent state captures the action-relevant physical outcome of robot interaction and serves as an explicit bridge between world modeling and action generation. The core design principle of DELE-w0.5 is to model how the physical world changes under robot actions, rather than how its visual appearance evolves frame by frame. This formulation removes the high-dimensional visual redundancy introduced by dense video representations, and it therefore enables cheaper training and low-latency inference. Across 480 real-robot trials on four long-horizon manipulation tasks, our DELE-w0.5 achieves the best performance among all compared policies, attaining 62.5 overall full-task success and 81.3 macro ordered-stage progress, outperforming the strongest baseline by 47.5 and 30.7 percentage points, respectively.
Chinese Translation
世界-动作模型(World-Action Models, WAMs)基于视频生成骨干构建机器人控制,联合预测密集的未来视觉轨迹和机器人动作。我们认为视频生成并不是世界-动作建模的必要中介目标。对于机器人操作而言,世界模型的目标并不是在每个中间时刻重现世界的外观,而是预测在执行某个动作后世界将达到的状态。中间帧仅描述物理状态之间的视觉过渡,这消耗了大量的模型容量和计算资源,但并没有直接指定机器人动作所期望产生的物理结果。在本文中,我们提出了DELE-w0.5,它从预测的未来状态中推断机器人动作,而无需依赖视频生成。具体而言,DELE-w0.5从其对应的紧凑未来潜在状态推断动作序列。未来潜在状态捕捉了机器人交互的动作相关物理结果,并作为世界建模与动作生成之间的显式桥梁。DELE-w0.5的核心设计原则是建模物理世界在机器人动作下如何变化,而不是其视觉外观如何逐帧演变。这一表述消除了由密集视频表示引入的高维视觉冗余,因此使得训练成本更低、推断延迟更小。在四个长时间操作任务的480次真实机器人试验中,我们的DELE-w0.5在所有比较策略中表现最佳,整体全任务成功率达到62.5,宏观有序阶段进展为81.3,分别比最强基线高出47.5和30.7个百分点。
cs.RO / 39 / 2608.22093

EndoNav: Semantic-to-Geometric Grounding for Language-Guided Robotic Endoscopic Examination

EndoNav:语言引导下的机器人内窥镜检查的语义到几何基础
Mao, Jecia Z. Y., Ishida, Hisashi, Jung, Kathryn, Ishii, Masaru, Taylor, Russell H., Sahu, Manish
Abstract
Minimally invasive procedures performed within confined anatomical spaces depend on continuous endoscopic visualization. Current robotic endoscope systems can stabilize or reposition an endoscope, but they do not possess relevant context to provide effective visualization assistance. We present EndoNav, an anatomy-grounded natural-language framework that translates high-level surgeon commands into autonomous endoscopic visualization behaviors within patient-specific sinonasal anatomy. Spoken surgeon commands are transcribed and interpreted by an endoscopic viewpoint agent conditioned on a patient-specific anatomical scene representation. Rather than generating robot motion directly, the viewpoint agent generates structured visualization objectives that are converted into target viewpoints and inspection trajectories, which are then executed through geometry-constrained endoscope motion planning and joint-space control. We evaluate EndoNav using a structured three-pass sinus examination across three CT-derived anatomical models. For one cadaveric specimen, autonomous visualization is compared with sinus examinations performed by two resident surgeons. EndoNav achieved mean visualization IoUs of 87.04% and 84.37% relative to the two surgeon examinations, compared with an inter-surgeon IoU of 87.44%, while recovering 92.91% and 93.20% of surgeon-observed anatomical surfaces, respectively. These results demonstrate the feasibility of grounding high-level anatomical commands into patient-specific geometric objectives and translating them into anatomically constrained robotic visualization behaviors.
Chinese Translation
在有限的解剖空间内进行的微创手术依赖于持续的内窥镜可视化。目前的机器人内窥镜系统能够稳定或重新定位内窥镜,但缺乏提供有效可视化辅助的相关上下文。我们提出了EndoNav,这是一个基于解剖的自然语言框架,将高级外科医生命令转化为患者特定的鼻窦解剖中的自主内窥镜可视化行为。外科医生的口头命令被转录并由一个内窥镜视角代理进行解释,该代理以患者特定的解剖场景表示为条件。该视角代理并不是直接生成机器人运动,而是生成结构化的可视化目标,这些目标被转换为目标视角和检查轨迹,然后通过几何约束的内窥镜运动规划和关节空间控制执行。我们使用结构化的三次鼻窦检查在三个CT衍生的解剖模型上评估EndoNav。对于一个尸体标本,自主可视化与两名住院外科医生进行的鼻窦检查进行了比较。EndoNav相对于两名外科医生的检查实现了87.04%和84.37%的平均可视化交并比(IoU),而两名外科医生之间的IoU为87.44%,同时分别恢复了92.91%和93.20%的外科医生观察到的解剖表面。这些结果证明了将高级解剖命令基础于患者特定几何目标并将其转化为解剖约束的机器人可视化行为的可行性。
cs.RO / 40 / 2608.22100

Contact-Rich Robotic Manipulation in Construction via Zero-Shot Learning: A Diffusion Policy-Guided Adaptive Control

通过零样本学习实现建筑中的接触丰富机器人操作:一种扩散策略引导的自适应控制
Ibrahimov, Roman, Mozaffari, Salma, Adel, Arash
Abstract
Construction robotics and automation offer promising means of improving productivity, alleviating workforce shortages, and reducing workers' exposure to physically demanding tasks. However, reliable contact-rich robotic assembly remains challenging under tight tolerances, fabrication inaccuracies, and uncertain contact dynamics. To address this challenge, we present a framework coupling diffusion policies trained on simulation-generated pose and force/torque data with an L1-inspired adaptive controller that corrects policy-predicted actions online to compensate for unmodeled contact dynamics. We benchmark the framework against baselines in timber joinery, pipe fitting, and sequential full-scale truss assembly. It achieves 100% success on single-task assemblies and 90-100% success across sequential truss assembly subtasks, with lower, more stable contact forces than the baselines. By enabling zero-shot sim-to-real transfer for force-aware contact-rich assembly, the framework reduces costly, labor-intensive real-world data collection for policy training and advances scalable, robust automation of multistage assembly, motivating extension to broader contact-rich manipulation tasks in construction.
Chinese Translation
建筑机器人和自动化提供了提高生产力、缓解劳动力短缺以及减少工人接触体力劳动任务的有希望的手段。然而,在严格公差、制造不准确性和不确定的接触动态下,可靠的接触丰富机器人组装仍然具有挑战性。为了解决这一挑战,我们提出了一个框架,将在模拟生成的姿态和力/扭矩数据上训练的扩散策略与一种受L1启发的自适应控制器相结合,该控制器在线修正策略预测的动作,以补偿未建模的接触动态。我们在木材连接、管道配件和顺序全尺度桁架组装中对该框架进行了基准测试。它在单任务组装中实现了100%的成功率,在顺序桁架组装子任务中实现了90-100%的成功率,并且接触力比基准更低且更稳定。通过实现针对力感知接触丰富组装的零样本仿真到现实转移,该框架减少了政策训练所需的昂贵且劳动密集的现实数据收集,并推动了多阶段组装的可扩展、稳健的自动化,激励其扩展到建筑中更广泛的接触丰富操作任务。
cs.RO / 41 / 2608.22149

Meta-Ctrl: Guaranteed Plan Generation by Decoupling Syntactic and Semantic Constraints

Meta-Ctrl:通过解耦语法和语义约束实现保证的计划生成
Yidou-Weng, Gwen, Sun, Edward, Ma, Tianyi, Dogan, Metin Alp, Wang, Benjie, Peng, Allen, Broeck, Guy Van den, Cui, Yuchen
Abstract
LLMs generate fluent plans for robots but routinely violate the syntactic and se8mantic constraints they must satisfy to execute, and existing remedies trade formal guarantees against plan quality: soft methods (affordance scoring, grounded decoding) give no guarantee, while symbolic planners (LLM+P) discard the LM's commonsense. We propose \textbf{Meta-Ctrl}, a constrained-decoding framework that guarantees the encoded constraints while preserving the base LM's plan quality. Meta-Ctrl introduces \emph{meta-tokens}---a compact vocabulary of grounded actions---enforcing syntax at the token level and semantics (preconditions, goals, ordering) at the action level, an exact factorization that cuts the memory of constrained decoding from over 107TB to under 2GB. With it, a small open-weight LM becomes competitive where it otherwise sits at the bottom of the leaderboard: on WAH-NL under the LoTa-Bench protocol it reaches the highest reported subgoal success rate, exceeding GPT-4's, with consistent gains across the Embodied Agent Interface. We further demonstrate it on a real tabletop robot, where every generated plan satisfies its preconditions and goals by construction. Project website: https://meta-ctrlg.github.io/.
Chinese Translation
大型语言模型(LLMs)为机器人生成流畅的计划,但通常违反执行所需满足的语法和语义约束,而现有的解决方案在形式保证和计划质量之间进行权衡:软方法(如赋能评分、基于语境的解码)没有保证,而符号规划器(LLM+P)则舍弃了语言模型的常识。我们提出了 extbf{Meta-Ctrl},一个约束解码框架,能够在保证编码约束的同时保持基础语言模型的计划质量。Meta-Ctrl引入了 extit{meta-tokens}——一种紧凑的基础动作词汇——在标记级别强制执行语法,在动作级别强制执行语义(前提条件、目标、顺序),这种精确的因式分解将约束解码的内存从超过107TB减少到不足2GB。借助此框架,一个小型的开放权重语言模型在原本处于排行榜底部的情况下变得具有竞争力:在LoTa-Bench协议下的WAH-NL上,它达到了最高报告的子目标成功率,超过了GPT-4,并在具身代理接口上实现了一致的增益。我们进一步在一个真实的桌面机器人上演示了这一点,其中每个生成的计划在构造上都满足其前提条件和目标。项目网站:https://meta-ctrlg.github.io/
cs.RO / 42 / 2608.22187

BehaviorWorldGen: Closing the Loop between Action Models and World Simulators via Controllable Behavior-Aware Structured World Generation

行为世界生成:通过可控的行为感知结构化世界生成闭合动作模型与世界模拟器之间的循环
Wang, Jiaqi, Zhang, Zhuo, Guan, Haining, Zhou, Tingguang, Cui, Haowen, Zhu, Zhongyang, Zheng, Yulong, Wang, ChuanYe, Chen, Xuefeng, Yang, Zhen, Deng, Tianchen, Tan, Feiyang, Zhou, Hangning, Dai, Bo, Shen, Lixia, Chen, Xiwu, Wang, Xiyang, Zhu, Jiajun
Abstract
Modern driving action models are increasingly improved in a self-improvement loop, where a learned world simulator imagines future observations and the resulting data is fed back to refine the action model. However, the bottleneck of this loop lies in the simulators' inability to generate behaviorally plausible responses by surrounding agents, making generated data both unrealistic in interaction and imbalanced in distribution. We introduce BehaviorWorldGen, a framework that closes the loop between action models and world simulators through controllable behavior-aware structured world generation. Its core component is BehaviorFlow, a meta-action-conditioned traffic-flow model that injects interpretable behavior controls and jointly generates multi-agent rollouts. BehaviorFlow realizes the specified agent behaviors while allowing surrounding vehicles to respond to the ego and to one another. The resulting rollouts are rendered by a world simulator into realistic multi-view observations, which are paired with corrected interaction-aware trajectories for action-model refinement. Since BehaviorWorldGen uses structured trajectories as the interface between its modules, it is compatible with diverse action models and world simulators. Experiments on world generation, scene extrapolation, and policy refinement demonstrate consistent improvements, with the largest benefits concentrated on difficult interactive scenarios.
Chinese Translation
现代驾驶动作模型在自我改进循环中不断提升,其中学习到的世界模拟器想象未来的观察结果,并将生成的数据反馈用于优化动作模型。然而,这一循环的瓶颈在于模拟器无法生成周围代理的行为上合理的反应,使得生成的数据在交互上既不真实又在分布上不平衡。我们提出了BehaviorWorldGen,一个通过可控的行为感知结构化世界生成来闭合动作模型与世界模拟器之间循环的框架。其核心组件是BehaviorFlow,一个元动作条件的交通流模型,注入可解释的行为控制,并共同生成多代理的滚动轨迹。BehaviorFlow实现了指定的代理行为,同时允许周围车辆对自我和彼此做出反应。生成的滚动轨迹由世界模拟器渲染为真实的多视角观察,并与经过修正的交互感知轨迹配对,以便于动作模型的优化。由于BehaviorWorldGen使用结构化轨迹作为其模块之间的接口,因此它与多种动作模型和世界模拟器兼容。在世界生成、场景外推和策略优化的实验中,表现出一致的改进,最大的收益集中在困难的交互场景上。
cs.RO / 43 / 2608.22278

DreamMimic: Learning Visuomotor Whole-Body Loco-Manipulation via World Model

DreamMimic:通过世界模型学习视觉运动全身步态操控
Yin, Jie, Lai, Xingyu
Abstract
Vision-based whole-body loco-manipulation on humanoid robots is challenging due to partial observability, contact-rich dynamics, and the difficulty of learning long-horizon behaviors from high-dimensional visual inputs. We present \href{https://github.com/DreamMimic/DreamMimic}{DreamMimic}, a framework that distills privileged teacher policies into vision-based humanoid controllers via world-model-assisted distillation. Instead of using a Dreamer-style RSSM for planning, we repurpose it to learn predictive latent dynamics that serve as both a representation space and an action-conditioned multi-step supervision signal, while exposing compact predictive features to the student policy to reduce long-term drift. Beyond standard reconstruction objectives for proprioceptive and visual observations, we add auxiliary prediction heads for privileged state, contact, object state, and reward estimation. These heads provide additional supervision related to agent--object interaction and task progress, encouraging the latent representation to retain signals that are useful for contact-rich loco-manipulation. We further introduce Performance-Conditioned Guidance (PCG), a reward-driven adaptive distillation schedule that computes performance scores for both teacher and student to dynamically balance guidance and exploration. PCG prevents both premature teacher annealing and excessive teacher interference in challenging visual settings. Experiments on OMOMO and BEHAVE show improved tracking-based loco-manipulation performance over strong vision-based baselines, without exposing online privileged interaction states to the student at deployment. Qualitative simulations further examine morphology and simulator changes. These results suggest that world models can provide a useful mechanism for stabilizing visual policy distillation in contact-rich humanoid behaviors.
Chinese Translation
基于视觉的类人机器人全身步态操控面临部分可观测性、接触丰富的动态以及从高维视觉输入中学习长时间行为的困难。我们提出了DreamMimic,一个通过世界模型辅助蒸馏将特权教师策略提炼为基于视觉的类人控制器的框架。我们不使用Dreamer风格的RSSM进行规划,而是将其重新利用以学习预测潜在动态,这些动态既作为表示空间,又作为动作条件的多步监督信号,同时向学生策略暴露紧凑的预测特征,以减少长期漂移。除了对本体感知和视觉观测的标准重建目标外,我们还为特权状态、接触、物体状态和奖励估计添加了辅助预测头。这些预测头提供了与代理-物体交互和任务进展相关的额外监督,鼓励潜在表示保留对接触丰富的步态操控有用的信号。我们进一步引入了性能条件引导(Performance-Conditioned Guidance, PCG),这是一种基于奖励的自适应蒸馏调度,它为教师和学生计算性能评分,以动态平衡引导和探索。PCG防止了在挑战性视觉环境中教师过早退火和过度干预。我们在OMOMO和BEHAVE上的实验显示,在不向学生暴露在线特权交互状态的情况下,基于跟踪的步态操控性能优于强大的基于视觉的基线。定性模拟进一步考察了形态和模拟器的变化。这些结果表明,世界模型可以为稳定接触丰富类人行为的视觉策略蒸馏提供有用机制。
cs.RO / 44 / 2608.22294

Beyond Instance Slots: Semantically Rich World Models for Physical Interaction Planning

超越实例槽:用于物理交互规划的语义丰富世界模型
Cheng, Juntao, Wang, Jingkai, Shen, Yijun, Chen, Xiansheng, Yu, Zhiwei
Abstract
World models for physical interaction are typically trained to predict future observations or latent features; however, a planning-oriented model must answer a fundamentally different question: whether a candidate action produces a task-consistent future while preserving essential relations.Monolithic state representations obscure the underlying entities, while standard instance-level object slots merely identify \emph{what} is present without specifying \emph{what role} each entity plays in the task context. To bridge this gap, we present the Semantically Rich World Model (SR-WM), a task-conditioned world model structured around five functional roles: gripper, target, goal, relation, and phase.Within SR-WM, a visual entity encoder extracts soft entity hypotheses from pretrained patch features, allowing segmentation masks to serve as optional proposal priors without mandating them as required state representations or inference inputs.A role binder subsequently maps these hypotheses to task-specific roles, while an action-conditioned dynamics model predicts role transitions alongside fine-grained semantics, including grasp/contact, predicate establishment, relation preservation, fixture state, and phase change.Crucially, this unified role state grounds downstream multi-candidate action generation, stage-aware reranking, and violation-aware suffix resampling.Our comprehensive evaluation protocol spans all four LIBERO simulation suites, cross-suite transfer, perception diagnostics, and action-sensitivity analysis.Ultimately, this formulation transforms object-centric prediction into a semantic interface linking visual dynamics with planning-oriented decision making.
Chinese Translation
物理交互的世界模型通常被训练用于预测未来观察或潜在特征;然而,面向规划的模型必须回答一个根本不同的问题:候选动作是否能够产生一个与任务一致的未来,同时保持基本关系的完整性。单一状态表示模糊了潜在实体,而标准的实例级对象槽仅仅识别出 extit{什么}存在,而未指定每个实体在任务上下文中扮演的 extit{什么角色}。为了解决这一问题,我们提出了语义丰富世界模型(Semantically Rich World Model, SR-WM),这是一个围绕五个功能角色构建的任务条件世界模型:抓取器、目标、目标状态、关系和阶段。在SR-WM中,视觉实体编码器从预训练的补丁特征中提取软实体假设,使得分割掩码可以作为可选的提议先验,而不强制要求它们作为必要的状态表示或推理输入。角色绑定器随后将这些假设映射到任务特定的角色,而基于动作的动态模型则预测角色转变及其细粒度语义,包括抓取/接触、谓词建立、关系保持、夹具状态和阶段变化。至关重要的是,这种统一的角色状态为下游多候选动作生成、阶段感知重排序和违规感知后缀重采样奠定了基础。我们的综合评估协议涵盖了所有四个LIBERO模拟套件、跨套件转移、感知诊断和动作敏感性分析。最终,这一表述将以对象为中心的预测转化为一个语义接口,将视觉动态与面向规划的决策制定联系起来。
cs.RO / 45 / 2608.22296

TONAV: Task-Oriented Navigation and Action-Velocity Chunk Learning for Articulated Object Quadrupedal Mobile Manipulation

TONAV:面向任务的导航与动作速度块学习用于关节物体四足移动操控
Lin, Haoran, Yang, Mingyu, Qi, Pengfei, Chen, Kehan, Diao, Qiang, Zeng, Liangji, Chen, Wenrui, Wang, Yaonan, Yang, Kailun
Abstract
Quadruped mobile manipulation requires two tightly coupled capabilities: reaching manipulation-ready configurations and maintaining stable contact throughout articulated-object interaction. However, existing methods often terminate navigation near the target, leaving a gap between reachability and manipulation readiness, while tracking lag, motion jitter, and contact instability limit continuous interaction. To address these challenges, we present TONAV, a unified framework integrating task-oriented navigation with action-velocity chunk learning. First, we introduce a position-velocity-coupled teleoperation framework that explicitly captures motion dynamics to improve master-follower consistency and collect smooth, temporally consistent demonstrations. Next, task-oriented navigation leverages vision-language reasoning to decompose high-level instructions into executable subgoals and adaptively refine the robot base toward a manipulation-ready configuration. Finally, action-velocity chunk learning jointly models joint positions and their temporal transitions under velocity supervision, enabling smooth and stable sustained-contact manipulation. Real-world experiments across diverse articulated-object tasks demonstrate that TONAV achieves higher success rates in both task-oriented navigation and complete mobile manipulation, mitigating the navigation-manipulation gap and improving continuous-contact interaction. The project page is at https://haochen611.github.io/TONAV.
Chinese Translation
四足移动操控需要两个紧密耦合的能力:达到适合操控的配置以及在与关节物体交互过程中保持稳定接触。然而,现有方法通常在目标附近终止导航,导致可达性与操控准备之间存在差距,同时跟踪延迟、运动抖动和接触不稳定限制了连续交互。为了解决这些挑战,我们提出了TONAV,一个将面向任务的导航与动作速度块学习相结合的统一框架。首先,我们引入一个位置-速度耦合的遥操作框架,明确捕捉运动动态,以改善主从一致性并收集平滑、时间一致的演示。接下来,面向任务的导航利用视觉-语言推理将高层指令分解为可执行的子目标,并自适应地调整机器人底座以达到适合操控的配置。最后,动作速度块学习在速度监督下联合建模关节位置及其时间过渡,实现平滑且稳定的持续接触操控。在多种关节物体任务的实际实验中,TONAV在面向任务的导航和完整移动操控中均实现了更高的成功率,减小了导航与操控之间的差距,并改善了连续接触交互。项目页面可访问 https://haochen611.github.io/TONAV。
cs.RO / 46 / 2608.22301

The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction

模仿者游戏:超越动作预测的机器人模仿能力基准测试
Zhou, Xunzhe, Cai, Yiyang, Wang, Fengyi, Ju, Ran, Ren, Hanxiang, Liu, Ruizhe, Zhang, Yu, Luo, Qian, Chen, Feng, Zhou, Pei, Ma, Yi, Yang, Yanchao
Abstract
Humans imitate at the level of intent: given a demonstration, we infer its goal and carry it out with whatever tools, objects, and layouts are at hand. Current robot policies instead learn observation-to-action mappings from visual inputs and language instructions, without explicitly inferring the demonstrated task. Learning from human video thus remains largely trajectory-level: models can replay motions in near-identical scenes, but still struggle to imitate what the demonstrator intends rather than merely what they do. We introduce The Imitator Game, a four-level benchmark (L0-L3) that progressively widens the gap between the human demonstration and the robot's own scene, isolating where trajectory replay ceases to suffice and task understanding becomes necessary. We pair it with IG-10K, the largest environment-aligned paired human-robot dataset to date and the only one instantiated across all four levels in both real and simulated settings (20,000+ paired episodes, 50+ tasks, 6 domains), and Imitator Arena, an open platform for blind A/B human evaluation. Across nine state-of-the-art models, performance is stable from L0 to L2 but collapses at L3, identifying functional substitution - achieving the same intent through a different object affordance - as the decisive barrier to intent-level imitation. Human-video-conditioned models outperform caption-conditioned ones, yet every model falls below 13% zero-shot success on unseen tasks; fine-tuning IG-10K-pretrained models with only $10$ paired human-robot demonstrations yields large gains that grow with pretraining scale. The project website and access to Imitator Arena are available at https://imitator-game.github.io.
Chinese Translation
人类在意图层面进行模仿:在给定演示的情况下,我们推断其目标,并利用手头的工具、物体和布局来实现该目标。目前的机器人策略则是从视觉输入和语言指令中学习观察到动作的映射,而未明确推断所演示的任务。因此,从人类视频学习仍然主要停留在轨迹层面:模型可以在几乎相同的场景中重放动作,但仍然难以模仿演示者的意图,而不仅仅是他们所做的事情。我们引入了模仿者游戏,这是一个四级基准测试(L0-L3),逐步扩大人类演示与机器人自身场景之间的差距,隔离出轨迹重放何时不足以满足需求,以及任务理解何时变得必要。我们将其与IG-10K配对,这是迄今为止最大的环境对齐的人机配对数据集,也是唯一一个在真实和模拟环境中跨越所有四个级别实例化的数据集(超过20,000个配对情节,50多个任务,6个领域),以及模仿者竞技场,一个用于盲目A/B人类评估的开放平台。在九个最先进的模型中,性能在L0到L2之间稳定,但在L3时崩溃,识别出功能替代——通过不同的物体可用性实现相同意图——是意图层面模仿的决定性障碍。人类视频条件下的模型优于标题条件下的模型,但每个模型在未见任务上的零-shot成功率均低于13%;仅用10个配对的人机演示对IG-10K预训练模型进行微调,带来了显著的收益,并随着预训练规模的增加而增长。项目网站和模仿者竞技场的访问链接可在 https://imitator-game.github.io 找到。
cs.RO / 47 / 2608.22326

GCS-Bridging: Restoring Connectivity of Disconnected Convex Sets for Graph-of-Convex-Sets Motion Planning

GCS-桥接:恢复图形凸集运动规划中断连凸集的连通性
Zhou, Xiaokai, Cao, Baoshi, Liu, Yang, Sun, Kui, Ma, Boyu, Wang, Zhengpu, Xie, Zongwu
Abstract
Graph-of-Convex-Sets (GCS)-based trajectory optimization represents collision-free regions in configuration space as a finite collection of convex sets and directly performs collision-free trajectory planning over these sets, substantially simplifying the planning process. However, existing GCS-based trajectory planning methods generally assume sufficient connectivity among the convex regions and do not explicitly address cases in which the start and goal regions belong to different connected components of the initial GCS map. To address this limitation, we propose GCS-Bridging, which reconnects disconnected convex regions through collision-free point paths followed by convex region inflation, thereby recovering the feasibility of otherwise disconnected GCS planning problems. Extensive simulations across multiple IRIS-related algorithms and scenarios demonstrate that GCS-Bridging restores missing start-to-goal connectivity in the initial GCS map with a 99.8% success rate. In addition, a hardware experiment on a single-arm Franka platform in a real-world scenario with initially disconnected start and goal regions validates the effectiveness of the proposed method in practical motion planning. Project website: https://zhouxk1997.github.io/GCS_Bridging/
Chinese Translation
基于图形凸集(Graph-of-Convex-Sets,GCS)的轨迹优化将配置空间中的无碰撞区域表示为有限的凸集集合,并直接在这些集合上执行无碰撞轨迹规划,从而显著简化了规划过程。然而,现有的基于GCS的轨迹规划方法通常假设凸区域之间具有足够的连通性,并未明确解决起始区域和目标区域属于初始GCS图的不同连通分量的情况。为了解决这一限制,我们提出了GCS-桥接,通过无碰撞点路径重新连接断开的凸区域,随后进行凸区域膨胀,从而恢复原本断开的GCS规划问题的可行性。在多个与IRIS相关的算法和场景下进行的大量仿真实验表明,GCS-桥接以99.8%的成功率恢复了初始GCS图中缺失的起始到目标的连通性。此外,在一个实际场景中,针对初始断开的起始和目标区域在单臂Franka平台上的硬件实验验证了所提方法在实际运动规划中的有效性。项目网站:https://zhouxk1997.github.io/GCS_Bridging/
cs.RO / 48 / 2608.22398

MotionDLO: Hybrid Event- and Frame-Based Tracking of Deformable Linear Objects

MotionDLO:可变形线性物体的混合事件与帧基跟踪
Hartmann, Annalena, Ajithkumar, Priyamvada, Bründl, Patrick, Franke, Jörg
Abstract
Reliably tracking moving deformable linear objects (DLOs) while simultaneously ensuring robustness, accuracy, and temporally consistent state estimation remains a fundamental challenge in robot perception. We introduce MotionDLO, a real-time tracking framework specifically designed to overcome these limitations in temporal continuity and latency. The method exploits the high temporal resolution and sparsity of event-based cameras and combines segmentation with the Coherent Point Drift (CPD) algorithm under the principles of Motion Coherence Theory. This integration enables temporally consistent shape estimation while maintaining a low computational overhead. Existing event-based tracking methods are typically computationally efficient but exhibit reduced accuracy compared to frame-based approaches, or alternatively compromise event sparsity to achieve competitive performance. To resolve this trade-off, we propose a hybrid event- and frame-based tracking architecture that preserves the complementary strengths of both sensing modalities. The event stream ensures high-frequency motion updates, while frame-based information stabilizes spatial accuracy and object identity. We demonstrate that the proposed framework reliably associates DLO instances across video sequences, enabling robust perception for robotic manipulation tasks. Experimental results validate real-time performance at 12 ms update rates and accurate shape tracking with an point-to-curve error as measurement of accuracy of up to 0.43 mm, supporting dynamic path adaptation during manipulation. The source code and demonstration datasets are publicly available.
Chinese Translation
在机器人感知中,可靠地跟踪移动的可变形线性物体(DLO)并同时确保鲁棒性、准确性和时间一致的状态估计仍然是一个基本挑战。我们提出了MotionDLO,一个实时跟踪框架,专门设计用于克服时间连续性和延迟方面的限制。该方法利用事件摄像头的高时间分辨率和稀疏性,并结合基于Motion Coherence Theory的分割与一致点漂移(Coherent Point Drift, CPD)算法。这种集成使得在保持低计算开销的同时,实现时间一致的形状估计。现有的基于事件的跟踪方法通常计算效率高,但与基于帧的方法相比,准确性较低,或者为了实现竞争性能而妥协事件稀疏性。为了解决这一权衡,我们提出了一种混合事件与帧基的跟踪架构,保留了两种传感模式的互补优势。事件流确保高频率的运动更新,而基于帧的信息则稳定了空间准确性和物体身份。我们证明了所提出的框架能够可靠地关联视频序列中的DLO实例,从而为机器人操作任务提供鲁棒的感知。实验结果验证了在12毫秒更新率下的实时性能,以及准确的形状跟踪,其点到曲线的误差作为准确性度量可达到0.43毫米,支持在操作过程中动态路径适应。源代码和演示数据集已公开。
cs.RO / 49 / 2608.22403

LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models

LD4WAM:从人类视频中学习潜在动态以构建世界行动模型
Shen, Zhenhao, Liang, Jiaqi, Lu, Jasper, Jiang, Feng, Wang, Yuran, Wei, Chuanbo, Liu, Jiayi, Yang, Jianchun, Yu, Qize, You, Jiadi, Hao, Ce, He, Guanqi, Xie, Chen, Wu, Ruihai
Abstract
Human video is playing an increasingly central role in training World Action Models (WAMs), owing to its diversity and low collection cost relative to teleoperated robot data. However, most WAMs learn from such video only by predicting pixel-level future frames, giving dynamics that are not directly actionable, whereas motion retargeting recovers directly actionable actions but leaves a large visual gap across embodiments. We therefore propose motion-aligned latent dynamics as an embodiment-agnostic representation to bridge video priors and low-level actions. We further present LD4WAM, which pairs a Latent Dynamics Model trained with semantic reconstruction and real motion alignment with a World Dynamics Action Model built as a mixture-of-transformers (MoT), which preserves full future-video generation and uses learnable queries to distill these latent dynamics from generated futures for action conditioning. Pretrained on our curated unified dataset of over 5{,}000 hours of human and robot data, LD4WAM performs strongly in RoboTwin simulation and on real robots equipped with both grippers and dexterous hands, while generalizing well to unseen objects and backgrounds.
Chinese Translation
人类视频在训练世界行动模型(World Action Models, WAMs)中扮演着越来越重要的角色,这主要得益于其多样性和相较于遥控机器人数据的低收集成本。然而,大多数WAM仅通过预测像素级的未来帧来学习这些视频,从而产生的动态并不能直接用于行动,而运动重定向则恢复了可直接执行的动作,但在不同表现形式之间留下了较大的视觉差距。因此,我们提出了运动对齐的潜在动态作为一种与表现形式无关的表示,以桥接视频先验和低级动作。我们进一步提出了LD4WAM,它将一个通过语义重建和真实运动对齐训练的潜在动态模型与一个构建为变换器混合模型(mixture-of-transformers, MoT)的世界动态行动模型相结合,该模型保留了完整的未来视频生成,并使用可学习的查询从生成的未来中提炼这些潜在动态以进行行动条件化。LD4WAM在我们精心策划的统一数据集上进行了预训练,该数据集包含超过5000小时的人类和机器人数据,在RoboTwin仿真和配备夹持器及灵巧手的真实机器人上表现出色,同时在未见过的物体和背景上也能很好地泛化。
cs.RO / 50 / 2608.22419

Robust Bimanual Vision-Language-Action Models via Embarrassingly Simple Modality Masking

通过极其简单的模态屏蔽构建稳健的双手视觉-语言-动作模型
Cheng, Dongzhou, Li, Ziang, Zhou, Yixiao, Li, Haojuan, Zhang, Jinghao, Lei, Lei, Dong, Minjing, Gui, Jie, Wang, Jiaqi
Abstract
Query-based Vision-Language-Action (VLA) models offer low-latency inference that is attractive for bimanual robotic manipulation, but we observe that they can still exhibit discontinuous actions and execution failures in complex dual-arm tasks. We hypothesize that unstable multi-view and language fusion is one contributing factor in these failures, often coinciding with attention spreading to distracting regions. To improve robustness, we introduce the Modality Masking Mechanism (M3), an embarrassingly simple, training-only strategy that requires no architectural changes or large-scale robot pretraining. M3 stochastically masks subsets of modality channels during training, exposing the policy to controlled partial observations and encouraging it to rely less on distracting cues and more on evidence that remains reliable. We evaluate M3 on ten bimanual tasks from RoboTwin 2.0 and on three long-horizon real-world tasks. Compared with the Adapter baseline, M3 improves average success by 21.7% in the Clean setting and 11.4% in Clean2Rand, where policies are trained on clean demonstrations and evaluated on randomized scenes, while also improving averaged real-world full-task success by over 30%. These results suggest that structured training-time masking is a practical way to improve the robustness of query-based VLA policies for bimanual manipulation.
Chinese Translation
基于查询的视觉-语言-动作(VLA)模型提供了低延迟的推理,这对于双手机器人操作具有吸引力,但我们观察到它们在复杂的双臂任务中仍可能表现出不连续的动作和执行失败。我们假设不稳定的多视角和语言融合是导致这些失败的一个因素,通常伴随着注意力分散到干扰区域。为了提高稳健性,我们引入了模态屏蔽机制(Modality Masking Mechanism, M3),这是一种极其简单的仅训练策略,无需架构更改或大规模机器人预训练。M3在训练过程中随机屏蔽模态通道的子集,使策略暴露于受控的部分观察中,并鼓励其减少对干扰线索的依赖,更多地依赖于可靠的证据。我们在RoboTwin 2.0的十个双手任务和三个长时间真实世界任务上评估了M3。与适配器基线相比,M3在Clean设置中将平均成功率提高了21.7%,在Clean2Rand中提高了11.4%,其中策略在干净演示上训练并在随机场景中评估,同时在真实世界全任务的平均成功率上也提高了超过30%。这些结果表明,结构化的训练时屏蔽是一种实用的方法,可以提高基于查询的VLA策略在双手操作中的稳健性。
cs.RO / 51 / 2608.22449

EMPIRE: Explicit Manipulation Planning as a Learnable Intermediate Representation for Egocentric Hand-Motion Forecasting

EMPIRE:作为可学习中间表示的显式操控规划用于自我中心手部动作预测
Wang, Wen, Hou, Ruibing, Chang, Hong, Shan, Shiguang, Chen, Xilin
Abstract
Forecasting dexterous hand motions from egocentric observations is fundamental to intelligent interactive systems. Existing VLM-based methods typically map observations directly to future motions, overlooking the underlying manipulation process that governs hand-object interactions. Moreover, end-to-end optimization couples manipulation learning with motion synthesis, causing motion-generation gradients to interfere with the pre-learned manipulation-aware representations. To overcome these limitations, we propose EMPIRE, a two-stage framework that introduces Explicit Manipulation Planning as an Intermediate Representation for Egocentric hand-motion forecasting. Stage I: Learn to Plan. EMPIRE first learns explicit manipulation plans from multimodal context to capture the progression of hand-object interactions. Stage II: Learn to Act. A motion generator synthesizes future bimanual hand motions conditioned on frozen planner representations, preventing motion-generation gradients from affecting manipulation planning. To support our method, we further construct EMPIRE-651K, a bimanual hand-motion forecasting dataset comprising 650,910 training windows across 111 tasks, each paired with an explicit per-hand manipulation plan. Under identical training and evaluation protocols, EMPIRE achieves state-of-the-art forecasting accuracy, with an MPJPE of 84.53 mm and a finger-relative error of 38.97mm. We release the code and dataset at https://github.com/wangwen-banban/EMPIRE.
Chinese Translation
从自我中心观察中预测灵巧的手部动作是智能交互系统的基础。现有的基于视觉语言模型(VLM)的方法通常直接将观察映射到未来的动作,忽视了支配手-物体交互的基本操控过程。此外,端到端优化将操控学习与动作合成耦合,导致动作生成梯度干扰预先学习的操控感知表示。为克服这些局限性,我们提出了EMPIRE,一个两阶段框架,引入显式操控规划作为自我中心手部动作预测的中间表示。第一阶段:学习规划。EMPIRE首先从多模态上下文中学习显式操控计划,以捕捉手-物体交互的进展。第二阶段:学习行动。动作生成器在冻结的规划器表示的条件下合成未来的双手动作,防止动作生成梯度影响操控规划。为了支持我们的方法,我们进一步构建了EMPIRE-651K,一个包含650,910个训练窗口的双手动作预测数据集,涵盖111个任务,每个任务都配有一个显式的每手操控计划。在相同的训练和评估协议下,EMPIRE实现了最先进的预测精度,MPJPE为84.53毫米,手指相对误差为38.97毫米。我们将在https://github.com/wangwen-banban/EMPIRE发布代码和数据集。
cs.RO / 52 / 2608.22496

A Unified Neural-Aided Alignment and Calibration Method for AUVs

一种统一的神经辅助对齐与标定方法用于自主水下航行器
Damari, Guy, Yampolsky, Zeev, Klein, Itzik
Abstract
Autonomous underwater vehicles (AUVs) rely on the fusion of inertial navigation systems (INS) and Doppler velocity logs (DVL) for accurate navigation. Before deployment, this fusion requires a DVL initialization pipeline consisting of two stages: alignment, which estimates the rotation between the INS and DVL frames, and calibration, which estimates the DVL error terms. Conventionally, both stages are solved with model-based algorithms that demand complex vehicle maneuvers, surface-level satellite reference measurements, and simplified error models, making initialization time-consuming, trajectory-dependent, and sensitive to sensor quality. In this work, we propose a fully neural- aided DVL initialization pipeline that replaces both stages with two complementary neural networks: ResAlignNet for alignment and DCNet for calibration. The unified pipeline operates in situ on a single nearly constant-velocity trajectory and uses the same inputs as the model-based baseline. Using real-world data recorded across five distinct sensor error-term combinations, the proposed pipeline reduces the velocity root mean squared error by an average of 68.7% over the model-based baseline, using only 25s of data for initialization.
Chinese Translation
自主水下航行器(AUV)依赖于惯性导航系统(INS)与多普勒速度计(DVL)的融合以实现精确导航。在部署之前,这种融合需要一个由两个阶段组成的DVL初始化流程:对齐阶段用于估计INS与DVL框架之间的旋转,标定阶段用于估计DVL误差项。传统上,这两个阶段通过基于模型的算法解决,这些算法要求复杂的航行器机动、地面卫星参考测量和简化的误差模型,使得初始化过程耗时、依赖轨迹,并对传感器质量敏感。在本研究中,我们提出了一种完全神经辅助的DVL初始化流程,用两个互补的神经网络替代这两个阶段:ResAlignNet用于对齐,DCNet用于标定。该统一流程在几乎恒定速度的单一轨迹上原位运行,并使用与基于模型的基线相同的输入。利用在五种不同传感器误差项组合下记录的真实数据,所提出的流程将速度均方根误差平均降低了68.7%,仅使用25秒的数据进行初始化。
cs.RO / 53 / 2608.22507

What is the effect of running-specific prostheses on long jumps? Optimization-based prediction and analysis using biomechanical models

跑步专用假肢对跳远的影响是什么?基于优化的预测与生物力学模型分析
Emonds, Anna Lena, Funken, Johannes, Potthast, Wolfgang, Mombaur, Katja
Abstract
Long jumpers with below the knee amputation (BKA) that take off from their running-specific prosthesis (RSP) improved performances significantly over the last years. The long jump biomechanics differs compared to athletes without BKA and the question arises whether the spring-like properties of the RSP facilitate achieving long jumping distances. The aim of this work is to propose a long jump model for athletes with and without BKA, to evaluate it and to apply it for comparing long jump motions with and without RSP. We establish rigid multi-body system models of one athlete with and one athlete without below the knee amputation (BKA). Long jump motions are computed by solving a specific optimal control problem (OCP) with constraints enforcing a physically correct dynamics, both for motion reconstruction or motion synthesis. With the proposed long jump model, we are able to compute realistic long jump motions. We discuss the causes of differences in measured long jumps and show directions for eliminating them. For both athletes, the synthesized solutions reveal potential for performance improvement. The jumping distance of the athlete without BKA is 64cm (6.9%) longer than the one of the athlete with BKA in the synthesized solutions.
Chinese Translation
近年来,使用跑步专用假肢(RSP)的下肢膝下截肢运动员在跳远方面的表现显著提升。跳远的生物力学与没有膝下截肢的运动员有所不同,因此产生了一个问题,即RSP的弹簧特性是否有助于实现更远的跳远距离。本研究的目的是为有无膝下截肢的运动员提出一个跳远模型,评估该模型,并应用于比较使用和不使用RSP的跳远动作。我们建立了一个包含一名膝下截肢运动员和一名非截肢运动员的刚性多体系统模型。通过解决一个特定的最优控制问题(OCP),在约束条件下强制实现物理上正确的动态,我们计算了跳远动作,无论是运动重建还是运动合成。利用所提出的跳远模型,我们能够计算出现实的跳远动作。我们讨论了测量跳远差异的原因,并展示了消除这些差异的方向。对于两名运动员,合成的解决方案显示出提升表现的潜力。在合成的解决方案中,非膝下截肢运动员的跳跃距离比膝下截肢运动员长64厘米(6.9%)。
cs.RO / 54 / 2608.22591

WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning

WorldToken:面向机器人模仿学习的时间优先序列建模
Yang, Chunkai, Yang, Andong, Gao, Chao
Abstract
Robot policies receive heterogeneous observations at each decision step, yet sequence models differ in how they organize these inputs over time. We introduce WorldToken, a time-first policy instantiation that fuses multiview images, proprioception, and task conditioning within each policy timestep into one world token. A causal temporal Transformer models the resulting world-token sequence, and a diffusion action head generates action chunks. On 23 RoboCasa tasks, an 85.3M-parameter policy trained from scratch apart from a frozen pretrained CLIP text encoder achieves 59.45% mean closed-loop success using 2,900 generated demonstrations per task. A complete factorial sweep over five dataset sizes, five model sizes, and two training seeds shows consistent gains from additional target-domain data and diminishing returns beyond moderate model size. Under same-checkpoint history truncation, reducing visible history to one or two policy timesteps lowers closed-loop success for all 50 RoboCasa policies. On RMBench Blocks Ranking, reducing visible history from 146 to 8 seconds lowers evaluator success from 95% to 28%, while an exploratory extended rollout sustains the reference swap sequence for over 850 seconds. These results establish the empirical feasibility of the complete WorldToken instantiation and characterize its data-scaling and temporal-context behavior under the tested recipes. They do not establish superiority over alternative sequence organizations or isolate which components of the complete implementation drive the observed performance.
Chinese Translation
机器人策略在每个决策步骤接收异构观察,但序列模型在如何组织这些输入随时间变化方面存在差异。我们提出了WorldToken,这是一种时间优先的策略实例,它将多视角图像、本体感知和任务条件融合在每个策略时间步中形成一个世界令牌。因果时间变换器对生成的世界令牌序列进行建模,而扩散动作头生成动作块。在23个RoboCasa任务中,一个从头开始训练的85.3M参数策略(除冻结的预训练CLIP文本编码器外)在每个任务中使用2,900个生成演示达到了59.45%的平均闭环成功率。对五个数据集大小、五个模型大小和两个训练种子的完整因子扫描显示,额外的目标领域数据带来了持续的收益,而在适中的模型大小之后收益递减。在相同检查点历史截断下,将可见历史减少到一个或两个策略时间步降低了所有50个RoboCasa策略的闭环成功率。在RMBench Blocks Ranking上,将可见历史从146秒减少到8秒使评估者的成功率从95%降至28%,而探索性扩展回合则将参考交换序列维持了超过850秒。这些结果确立了完整WorldToken实例化的经验可行性,并表征了其在测试方案下的数据扩展和时间上下文行为。它们并未确立相对于其他序列组织的优越性,也未隔离出完整实现中驱动观察到的性能的具体组件。
cs.RO / 55 / 2608.22629

Enhancing Sim2Real Transfer for Torque-Controlled Robots through Real2Sim Dynamics Estimation and Reinforcement Learning

通过Real2Sim动力学估计和强化学习增强扭矩控制机器人Sim2Real转移
Bargellini, Davide, Pasquali, Alex, Govoni, Andrea, Zanella, Riccardo, Palli, Gianluca
Abstract
Transferring reinforcement learning policies from simulation to Real-World robots remains a major challenge, particularly when dealing with low-level torque control, where even small modelling inaccuracies can lead to unstable or unsafe behaviours. In this work, we propose a Real2Sim2Real pipeline that improves Sim2Real transfer for torque-controlled robotic arms by combining trajectory matching, parameter optimization via genetic algorithms, and domain randomization. Using the 7-DOF Franka Emika Panda robot, we first identify friction, inertia, and gravity compensation parameters by minimizing the error between real and simulated joint trajectories. These calibrated dynamics are then used to train a TQC-based reinforcement learning agent in simulation. The trained policy is evaluated in both Gazebo and MuJoCo environments, and finally deployed on the real robot. Our results demonstrate a significant improvement in tracking accuracy and policy robustness after parameter tuning, with smooth policy transfer from simulation to the Real-World across multiple target-reaching tasks. This work highlights the effectiveness of accurate physical modelling in enabling stable and generalizable torque-based reinforcement learning policies.
Chinese Translation
将强化学习策略从仿真转移到真实机器人仍然是一个主要挑战,特别是在处理低级扭矩控制时,甚至微小的建模不准确都可能导致不稳定或不安全的行为。在本研究中,我们提出了一种Real2Sim2Real管道,通过结合轨迹匹配、基于遗传算法的参数优化和领域随机化,来改善扭矩控制机器人臂的Sim2Real转移。使用7自由度的Franka Emika Panda机器人,我们首先通过最小化真实和仿真关节轨迹之间的误差来识别摩擦、惯性和重力补偿参数。这些校准后的动力学随后用于在仿真中训练基于TQC的强化学习代理。训练后的策略在Gazebo和MuJoCo环境中进行评估,最终部署到真实机器人上。我们的结果表明,在参数调优后,跟踪精度和策略鲁棒性显著提高,且在多个目标到达任务中实现了从仿真到真实世界的平滑策略转移。本研究强调了准确物理建模在实现稳定和可推广的基于扭矩的强化学习策略中的有效性。
cs.RO / 56 / 2608.22657

Physical Agentic AI: An Architecture for Orchestrating a Robot Crew with LLMs

物理自主人工智能:一种利用大型语言模型协调机器人团队的架构
Liu, Xinyuan, Sadikoglu, Eren, Chatterjee, Riana, Senanayake, Ransalu
Abstract
Agentic AI frameworks interpret open-ended task goals and decompose them into multi-step plans. Richer information about embodiment-specific capabilities, physical preconditions, and cross-robot coordination improves grounding, but does not eliminate infeasible, mistimed, or unsafe physical actions. Physical robot crews therefore require an explicit architectural interface between semantic planning and execution, where every planned action is verified against robot capabilities, system state, and workflow constraints before actuation. This paper introduces Physical Agentic AI, a framework for skill-grounded robot agent orchestration, in which each robot exposes a typed library of executable skills while a foundation model planner decomposes a task into phases and assigns each phase to a robot-skill pair. A Robot Orchestration layer exposes the skill library, robot state, named locations, and workflow contracts to a non-actuating Mission Planner, while a deterministic Robot Orchestrator validates and authorizes one skill at a time. We evaluate on a drone-UGV search-and-dispatch mission, where every mission in every condition is executed live in Gazebo, and on a humanoid-quadruped transportation task using hardware-equivalent skill interfaces plus two physical trials on a Unitree G1 and Go2. Varying planner knowledge and runtime enforcement independently, we find that retrieval raises skill grounding from 51% to 96% yet leaves informed planners dispatching 23-29% of faulted steps. Per-dispatch enforcement reduces false dispatch to 0% with no false blocks, and a held-plan ablation confirms that the gate, not plan variation, is responsible. Live execution makes the difference physical: without enforcement all eight injected faults crossed the orchestration boundary and six produced robot motion; with enforcement all eight were refused before motion.
Chinese Translation
自主人工智能框架能够解释开放式任务目标,并将其分解为多步骤计划。关于具身特定能力、物理前提条件和跨机器人协调的更丰富信息可以改善基础,但并不能消除不可行、时机不当或不安全的物理动作。因此,物理机器人团队需要在语义规划和执行之间建立明确的架构接口,在此接口中,每个计划的动作都需在执行前与机器人能力、系统状态和工作流程约束进行验证。本文介绍了物理自主人工智能(Physical Agentic AI),这是一个基于技能的机器人代理协调框架,其中每个机器人都暴露出可执行技能的类型库,而基础模型规划器则将任务分解为多个阶段,并将每个阶段分配给一个机器人-技能对。机器人协调层向非执行的任务规划器暴露技能库、机器人状态、命名位置和工作流程合同,而确定性机器人协调器则一次验证并授权一个技能。我们在无人机-地面无人车搜索与调度任务中进行了评估,在Gazebo中实时执行每个条件下的每个任务,并在类人四足机器人运输任务中使用硬件等效的技能接口以及在Unitree G1和Go2上进行的两次物理试验。通过独立变化规划者知识和运行时强制执行,我们发现检索将技能基础提高了51%至96%,但仍然让有信息的规划者调度了23-29%的错误步骤。每次调度的强制执行将错误调度减少到0%,且没有错误阻塞,而保持计划的消融实验确认是门控,而非计划变异,负责。实时执行使差异变得物理化:在没有强制执行的情况下,所有八个注入的故障都越过了协调边界,其中六个导致了机器人的运动;而在强制执行下,所有八个故障在运动前都被拒绝。
cs.RO / 57 / 2608.22671

Exact Finite-Length Theory of Uniform Car Parking: Spatial Laws, Absorption, and Aggregation

均匀停车的精确有限长度理论:空间法则、吸收与聚合
Kumar, Ganesh P
Abstract
The uniform car-parking process is the one-dimensional random sequential adsorption of unit cars on a segment of finite length $s$: cars arrive at uniformly random positions and park wherever they fit, until no gap admits another. This paper develops the exact finite-$s$ theory. The joint density of the parked positions is resolved into jamming cells, on each of which it is a rational function, and evaluated by a subset recursion in $O(2^n n)$ operations; the marginal and gap order statistics are obtained as hyperlogarithms whose weight is fixed by the number of coordinates integrated out; and the absorption count and the aggregate quantities are treated through the integral equation descending from R\'enyi.
Chinese Translation
均匀停车过程是在有限长度 $s$ 的区间上对单位汽车进行的一维随机顺序吸附:汽车以均匀随机的位置到达,并在适合的地方停车,直到没有空隙可以再停放其他汽车。本文发展了精确的有限-$s$ 理论。停放位置的联合密度被分解为拥堵单元,在每个单元上它是一个有理函数,并通过 $O(2^n n)$ 次操作的子集递归进行评估;边际和间隙顺序统计量作为超对数函数获得,其权重由被积分的坐标数量固定;吸收计数和聚合量通过源自 Rényi 的积分方程进行处理。
cs.RO / 58 / 2608.22675

VikPath: A Vision Kansformer Framework for Effective Obstacle Avoidance in Self-Supervised Pathfinding

VikPath:一种用于自监督路径规划的有效障碍物避让的视觉变换框架
Wang, Junyao, Xu, Yulin, Faruque, Mohammad Abdullah Al
Abstract
Pathfinding is a fundamental problem in artificial intelligence and autonomous systems. Traditional heuristic-based algorithms, such as A*, rely on predefined heuristic functions to guide the search process. Although effective in structured environments, their search efficiency can degrade substantially in complex, obstacle-rich scenarios, where handcrafted heuristics may provide limited guidance. Recent studies have explored learning-based approaches to improve pathfinding efficiency; however, most existing methods rely on supervised learning and require labels generated by conventional planners or obtained through manual annotation. As a result, their performance is inherently influenced by the quality of the underlying supervision and may degrade when the labeling heuristics fail to capture complex environmental structures. Moreover, existing methods primarily optimize for path length while paying limited attention to obstacle clearance and trajectory smoothness, which can lead to paths that are difficult or unsafe to execute in real-world environments. To address these limitations, we propose $\Design$, a self-supervised pathfinding framework that jointly considers obstacle proximity and path smoothness. At its core, our novel \textit{Vision Kansformer} module learns representations of obstacle distributions without relying on labeled trajectories, enabling the model to better adapt to complex environments. We further introduce a sharp-turn penalty to encourage smoother and more practically executable paths. Extensive experiments demonstrate that, compared with state-of-the-art (SOTA) approaches, $\Design$ achieves an average of 3.28\% greater obstacle clearance and 87.07\% lower inference latency while maintaining smooth path generation.
Chinese Translation
路径规划是人工智能和自主系统中的一个基本问题。传统的基于启发式的方法,如 A* 算法,依赖于预定义的启发式函数来指导搜索过程。尽管在结构化环境中效果显著,但在复杂的障碍物丰富场景中,它们的搜索效率可能显著下降,因为手工设计的启发式方法可能提供有限的指导。近期的研究探索了基于学习的方法以提高路径规划效率;然而,大多数现有方法依赖于监督学习,并需要由传统规划器生成的标签或通过人工标注获得的标签。因此,它们的性能本质上受到基础监督质量的影响,当标注启发式未能捕捉复杂环境结构时,性能可能会下降。此外,现有方法主要优化路径长度,而对障碍物间距和轨迹平滑度关注有限,这可能导致在现实环境中难以或不安全执行的路径。为了解决这些局限性,我们提出了 $ ext{Design}$,一种自监督路径规划框架,联合考虑障碍物接近度和路径平滑性。我们的新颖的 extit{Vision Kansformer} 模块在不依赖标记轨迹的情况下学习障碍物分布的表示,使模型能够更好地适应复杂环境。我们进一步引入了急转弯惩罚,以鼓励生成更平滑和更具可执行性的路径。大量实验表明,与最先进的方法(SOTA)相比,$ ext{Design}$ 实现了平均 3.28 ext{%} 的障碍物间距提升和 87.07 ext{%} 的推理延迟降低,同时保持路径生成的平滑性。
cs.RO / 59 / 2608.22678

RACO: Reliability-Aware Coarse-Goal Optimization for Inspection-Oriented UAV Vision-Language Navigation

RACO:面向检查的无人机视觉-语言导航的可靠性意识粗目标优化
Wang, Sen, Sun, Yiming, He, Jiaxuan, Zhu, Pengfei
Abstract
UAV vision-language navigation (UAV-VLN) is commonly evaluated as goal reaching, but inspection-oriented deployment requires the agent to stop within a valid inspection region and avoid falsely confirming visually or semantically similar distractors. This requirement exposes a key weakness in existing coarse-to-fine UAV-VLN policies: the coarse goal predicted before local refinement is often treated as reliable, although it may drift toward plausible but incorrect object regions and limit the ability of the local stage to recover. To systematically evaluate this problem, we introduce LG-UVI, an object-centric inspection evaluation setting derived from CityNav/CityRefer. LG-UVI extends standard UAV-VLN episodes with target objects, hard distractors, type-aware inspection regions, and diagnostics for inspection-region arrival and object-level confirmation. To address this inspection-oriented setting, we further propose RACO, a reliability-aware adaptive coarse-to-fine navigation framework. Instead of treating the predicted coarse goal as a fixed waypoint, RACO views it as a runtime hypothesis and uses object-level candidate anchors to check and correct coarse localization before Stage 1 and at the Stage 1-to-Stage 2 boundary. RACO also applies scale-adaptive terminal refinement to handle terminal near-miss cases using runtime-observable geometric and anchor-based evidence. Under a unified online evaluation protocol, RACO improves SR over the reproduced HETT baseline by 9.53 and 7.98 percentage points on validation-unseen and test-unseen, respectively. It also improves inspection-region arrival and reduces false verification risk, showing that coarse-goal reliability optimization is an effective complement to existing coarse-to-fine UAV-VLN policies.
Chinese Translation
无人机视觉-语言导航(UAV-VLN)通常以目标到达作为评估标准,但面向检查的部署要求代理在有效的检查区域内停止,并避免错误确认视觉或语义上相似的干扰物。这一要求暴露了现有粗到细UAV-VLN策略的一个关键弱点:在局部细化之前预测的粗目标通常被视为可靠,尽管它可能偏向于合理但不正确的物体区域,从而限制了局部阶段的恢复能力。为了系统性地评估这个问题,我们引入了LG-UVI,这是一个基于CityNav/CityRefer派生的以物体为中心的检查评估设置。LG-UVI通过目标物体、困难干扰物、类型感知的检查区域以及检查区域到达和物体级确认的诊断扩展了标准的UAV-VLN情节。为了解决这一面向检查的设置,我们进一步提出了RACO,一个可靠性意识的自适应粗到细导航框架。RACO将预测的粗目标视为运行时假设,而不是将其视为固定的航点,并使用物体级候选锚点在第一阶段之前及第一阶段与第二阶段的边界处检查和修正粗定位。RACO还应用了尺度自适应终端细化,以利用运行时可观察的几何和基于锚点的证据处理终端近失案例。在统一的在线评估协议下,RACO在验证未见和测试未见上分别提高了9.53和7.98个百分点的成功率(SR),并改善了检查区域到达率,降低了错误验证风险,表明粗目标可靠性优化是对现有粗到细UAV-VLN策略的有效补充。
cs.RO / 60 / 2608.22701

Physics Filtering Favors the Generalization of Robot Learning

物理过滤促进机器人学习的泛化能力
Jia, Jindou, Han, Shixuan, Wang, Meng, Li, Gen, Yang, Zihan, Zhou, Sicheng, Guo, Kexin, Yang, Jianfei, Yu, Xiang, Wang, Wei, Guo, Lei
Abstract
Living organisms exhibit extraordinary adaptability to unseen environments through their intrinsic physical structures and lifelong feedback-driven learning. Endowing robots with comparable generalization is critical for reliable operation in the real world. While recent approaches attempt to improve generalization by scaling training data, such strategies remain impractical for robotics, where collecting real-world demonstrations at the scale of large language models is prohibitively costly and slow. Contrary to this reliance on massive datasets, we show that robots can generalize effectively under dynamics uncertainties even with limited training data by leveraging a feedback mechanism, namely PhyFilter, that corrects learning outputs with physics-filtered learning residuals. PhyFilter operates as a lightweight, model-agnostic module whose parameters can be automatically optimized through an auto-learning algorithm, eliminating manual tuning and enabling seamless integration with diverse robot policies. We validate PhyFilter across four representative robotic systems, demonstrating that it enables quadruped robots to generalize to unseen terrains, payload variations, and speed ranges; drones to flight under unseen wind disturbances; aerial manipulators to achieve centimeter-level in-air capture despite wind and mass uncertainties; and acceleration differentiators to remain robust with distribution shift. These results show that physics-filtered feedback can serve as a powerful alternative to massive data scaling.
Chinese Translation
生物体通过其内在的物理结构和终身反馈驱动的学习,展现出对未知环境的卓越适应能力。赋予机器人类似的泛化能力对于在现实世界中可靠操作至关重要。尽管近期的方法试图通过扩大训练数据来提高泛化能力,但这种策略在机器人领域仍然不切实际,因为以大型语言模型的规模收集真实世界的示范成本高昂且缓慢。与对大规模数据集的依赖相反,我们展示了机器人能够在动态不确定性下有效泛化,即使在有限的训练数据条件下,通过利用一种反馈机制——物理过滤器(PhyFilter),该机制通过物理过滤的学习残差来修正学习输出。PhyFilter作为一个轻量级、与模型无关的模块运行,其参数可以通过自学习算法自动优化,从而消除手动调优,并实现与多种机器人策略的无缝集成。我们在四个具有代表性的机器人系统中验证了PhyFilter,证明它使四足机器人能够泛化到未知地形、负载变化和速度范围;无人机能够在未知风扰动下飞行;空中操控器能够在风和质量不确定性下实现厘米级的空中捕获;加速度分离器在分布变化下保持稳健。这些结果表明,物理过滤反馈可以作为大规模数据扩展的强大替代方案。
cs.RO / 61 / 2608.22799

Reproducible Vision-Guided 6-DoF Robotic Manipulator with a Mixed Stepper-Driver Architecture and Browser-Native Control

可重复的视觉引导6自由度机器人操纵器,采用混合步进驱动架构和浏览器原生控制
Perera, Lasan, Priyadarshana, Deneth, Pitiwaduge, Dulana, Dinujaya, Isitha, Colambage, Mokshan
Abstract
We present the NeuralNexus Arm, an open, low-cost 6-DOF robotic manipulator built by an undergraduate engineering team, together with the design decisions and debugging experience needed to reproduce it. The arm is driven by a single STM32H743 microcontroller on a custom printed circuit board (PCB) and combines two stepper-driver strategies on one controller: push-pull 3.3 V step/direction outputs for onboard TMC2209 drivers on the three wrist joints, and open-drain outputs for external CL57T and DM542 drivers on the three high-torque proximal joints. We describe the mechanical design, mixed-driver electronics, interrupt-driven firmware, a MATLAB/Simscape-based inverse-kinematics pipeline, a browser-native control interface using the Web Serial API, and a lightweight vision pipeline for object localisation and autonomous pick-and-place tasks. We also document non-obvious hardware and firmware failure modes encountered during the transition from a development board to the custom PCB as reproducibility guidance. All design files and firmware are released openly. The platform actuates all six axes under coordinated control at a 2 kHz update rate and executes both manual and pre-recorded motions from the browser interface.
Chinese Translation
我们介绍了NeuralNexus Arm,这是一个由本科工程团队构建的开放式、低成本的6自由度机器人操纵器,并分享了重现该设备所需的设计决策和调试经验。该机械臂由一个单独的STM32H743微控制器驱动,安装在定制的印刷电路板(PCB)上,并在一个控制器上结合了两种步进驱动策略:为三个腕关节上的TMC2209驱动器提供推拉3.3 V步进/方向输出,以及为三个高扭矩近端关节上的外部CL57T和DM542驱动器提供开漏输出。我们描述了机械设计、混合驱动电子学、中断驱动固件、基于MATLAB/Simscape的逆运动学管道、使用Web Serial API的浏览器原生控制界面,以及用于物体定位和自主抓取与放置任务的轻量级视觉管道。我们还记录了在从开发板过渡到定制PCB过程中遇到的非显而易见的硬件和固件故障模式,以作为可重复性指导。所有设计文件和固件均已公开发布。该平台以2 kHz的更新频率协调控制所有六个轴,并从浏览器界面执行手动和预录制的运动。
cs.RO / 62 / 2608.22800

Triplet2Track: A Hierarchical System with Object-Centric Representations for Reliable Long-Horizon Manipulation

Triplet2Track:一种具有以对象为中心表示的分层系统,用于可靠的长时间操作
Liu, Jianxiang, Zhang, Gaojing, Wen, Chuan, Liu, Qipeng, Zhao, Yuxuan, Guo, Ning, Lian, Wenzhao
Abstract
Ensuring reliability in uncertain environments remains difficult for long-horizon robotic manipulation. End-to-end VLA models are data-heavy and opaque, making diagnosis and verification difficult. Hierarchical pipelines are more interpretable, but their plans are often weakly grounded in observations, weakly aligned with low-level actions, and computed without online feedback, leading to open-loop behavior and hallucinations. To address these issues, we introduce the Triplet-to-Track System (TTS), a closed-loop long-horizon imitation learning system that uses human videos to reduce reliance on robot-collected data. TTS represents high-level subgoals as instance-grounded triplets, translates them into continuous track priors for execution, and monitors task progress from observations for online replanning. Across diverse real-world long-horizon tasks, TTS achieves a 74.8\% average success rate and supports object-level and compositional generalization.
Chinese Translation
在不确定环境中确保长时间机器人操作的可靠性仍然困难。端到端的VLA模型数据量大且不透明,使得诊断和验证变得困难。分层管道更具可解释性,但它们的计划往往与观察结果的基础联系较弱,与低级动作的对齐程度较低,并且在没有在线反馈的情况下进行计算,导致开放式行为和幻觉。为了解决这些问题,我们提出了Triplet-to-Track系统(TTS),这是一种闭环的长时间模仿学习系统,利用人类视频减少对机器人收集数据的依赖。TTS将高层次子目标表示为实例基础的三元组,将其转换为执行的连续轨迹先验,并从观察中监测任务进展以进行在线重新规划。在各种真实世界的长时间任务中,TTS实现了74.8%的平均成功率,并支持对象级和组合泛化。
cs.RO / 63 / 2608.22869

UniMem: Unifying Multimodal Memory and Control for Vision-Language-Action Models

UniMem:统一多模态记忆与控制的视觉-语言-行动模型
Osterberg, Lars, Wang, Maggie, Schwager, Mac
Abstract
While Vision-Language-Action (VLA) models have leveraged internet-scale pretraining and task-focused finetuning to achieve strong performance on long-horizon tasks, they often struggle with non-Markovian tasks that require memory. Existing approaches to memory typically involve additional Vision-Language-Models (VLMs) for long-term memory management, introducing a memory bottleneck and a fractured training pipeline. Conditioning on multiple historical frames can provide the VLA with access to more descriptive features of past scenes, but can degrade performance if frames are chosen at arbitrary, fixed intervals. To address these limitations, we present UniMem, a framework that unifies high-level, multimodal memory and low-level control under one backbone. UniMem employs an event classifier for memory updates, a keyframe encoder for dense spatial memory, and a keyframe caching technique to minimize overhead during policy rollouts. We evaluate UniMem across five simulation and four hardware tasks targeting sequential and spatial memory, demonstrating that our unified, single-model system outperforms fixed-interval image sampling baselines (93.4% vs. 68.2%) in simulation and hierarchical baselines (80.0% vs. 43.5%) in hardware, while offering faster inference and a simple training pipeline for easy adoption. Project website: https://losterberg3.github.io/unimem-vla/
Chinese Translation
尽管视觉-语言-行动(VLA)模型利用互联网规模的预训练和任务聚焦的微调在长时间任务上取得了良好的表现,但它们在处理需要记忆的非马尔可夫任务时常常面临困难。现有的记忆处理方法通常涉及额外的视觉-语言模型(VLM)用于长期记忆管理,这引入了记忆瓶颈和断裂的训练流程。基于多个历史帧进行条件处理可以为VLA提供对过去场景更具描述性的特征的访问,但如果帧是以任意固定间隔选择的,则可能会降低性能。为了解决这些局限性,我们提出了UniMem,一个在同一主干下统一高层多模态记忆和低层控制的框架。UniMem采用事件分类器进行记忆更新,关键帧编码器用于密集空间记忆,以及关键帧缓存技术以最小化策略执行过程中的开销。我们在五个仿真任务和四个硬件任务上评估了UniMem,针对顺序和空间记忆,结果表明我们的统一单模型系统在仿真中优于固定间隔图像采样基线(93.4%对68.2%),在硬件中优于层次基线(80.0%对43.5%),同时提供更快的推理速度和简单的训练流程,便于采用。项目网站:https://losterberg3.github.io/unimem-vla/
cs.RO / 64 / 2608.22896

SuperMap: A Spatio-Temporal SLAM System for Visual-Language Navigation

SuperMap:一种用于视觉-语言导航的时空SLAM系统
Zhao, Shibo, Chen, Guofei, Zhu, Honghao, Li, Zhiheng, Yao, Changwei, Zantout, Nader, Kim, Seungchan, Wang, Wenshan, Zhang, Ji, Scherer, Sebastian
Abstract
Robotic navigation in human environments requires a spatio-temporal semantic representation that can rec- oncile open-vocabulary perception with long-term environmental changes. While foundation models provide strong zero-shot recognition, their predictions are intermittent and view-dependent, and naively integrating them into mapping pipelines leads to identity drift and stale semantics over time. We present SuperMap, a 4D spatio-temporal mapping framework for language-guided navigation that integrates high-frequency geometric SLAM with asynchronous open-vocabulary perception. Our core contribution is a consistency-driven mapping engine that combines 3D-aware instance association/re-activation with a principled existence-and-label confidence update to maintain stable object identities and prune outdated map content under occlusions and scene changes. SuperMap produces a queryable 4D scene-graph representation that interfaces naturally with Vision-Language Models by supporting compositional queries over object semantics, relations, We demonstrate SuperMap on benchmarks and real robots, including dynamic scenes with appearance/disappearance and relocation, and provide ablations and runtime analysis. We release the full system as open-source to provide the community with a deployable baseline for open-vocabulary spatio-temporal mapping. Project website: superodometry.com/supermap.
Chinese Translation
在人工环境中,机器人导航需要一种时空语义表示,以调和开放词汇感知与长期环境变化。尽管基础模型提供了强大的零样本识别能力,但它们的预测是间歇性的且依赖于视角,简单地将其整合到映射管道中会导致身份漂移和语义过时。我们提出了SuperMap,一个用于语言引导导航的4D时空映射框架,结合了高频几何SLAM与异步开放词汇感知。我们的核心贡献是一个一致性驱动的映射引擎,它结合了3D感知的实例关联/再激活以及原则性的存在与标签置信度更新,以维持稳定的物体身份,并在遮挡和场景变化下修剪过时的地图内容。SuperMap生成一个可查询的4D场景图表示,自然地与视觉-语言模型接口,支持对物体语义、关系的组合查询。我们在基准测试和真实机器人上展示了SuperMap,包括动态场景中的出现/消失和重新定位,并提供了消融实验和运行时分析。我们将完整系统作为开源发布,以为社区提供一个可部署的开放词汇时空映射基线。项目网站:superodometry.com/supermap。
cs.RO / 65 / 2608.22976

Privileged Critic Training Enables Sensor-Free Thruster Fault Adaptation in End-to-End RL

特权评论训练使无传感器推进器故障适应在端到端强化学习中成为可能
Castan, Ricard Marsal I, Olivares-Méndez, Miguel A.
Abstract
Fault-tolerant navigation for thruster-actuated robots requires online adaptation to failures that are neither binary nor fully observable: thrusters may degrade continuously, fail dead, or jam stuck-open. Classical fault detection pipelines require dedicated sensors unavailable at deployment; oracle controllers that observe the true failure state are equally impractical. We show that privileged critic training is sufficient for sensor-free fault adaptation: giving the PPO value function access to the true degradation state dgt during training, while the actor receives only standard task observations, shapes a policy that compensates for failures at deployment without any dedicated fault sensing. We propose RAFT (Recurrent Asymmetric Fault Tolerant), a policy with recurrent memory trained with a privileged asymmetric critic. Evaluated on a floating-platform robot (8 thrusters, 1 reaction wheel) under up to four simultaneous thruster failures, RAFT achieves 70.2% success at four concurrent failures, closing 84% of the gap from a failure-naive baseline (4.8%) to an oracle policy that sees the full degradation state at deployment (82.4%). All code, checkpoints, and data are open-source.
Chinese Translation
推进器驱动机器人的容错导航需要对既非二元又非完全可观察的故障进行在线适应:推进器可能会持续退化、完全失效或卡住在开启状态。传统的故障检测流程需要在部署时不可用的专用传感器;观察真实故障状态的神谕控制器同样不切实际。我们展示了特权评论训练足以实现无传感器的故障适应:在训练期间,给予PPO(Proximal Policy Optimization)价值函数访问真实退化状态dgt的权限,而演员仅接收标准任务观察,从而塑造出一种在部署时能够补偿故障的策略,而无需任何专用的故障传感。我们提出了RAFT(Recurrent Asymmetric Fault Tolerant),一种使用特权不对称评论器训练的具有递归记忆的策略。在一个浮动平台机器人(8个推进器,1个反应轮)上进行评估,在最多四个同时推进器故障的情况下,RAFT在四个并发故障下实现了70.2%的成功率,缩小了从一个对故障不敏感的基线(4.8%)到一个在部署时能够看到完整退化状态的神谕策略(82.4%)的84%的差距。所有代码、检查点和数据均为开源。
cs.RO / 66 / 2608.22983

CSymPlan: Certified Symbolic Planning and Control for High-DOF Manipulators

CSymPlan:高自由度机械臂的认证符号规划与控制
Narendra, Aditya, Saini, Ashok Kumar, Anand, Mahathi, Khaled, Mahmoud, Abu-Dakka, Fares J., Swikir, Abdalla
Abstract
Robot manipulators are commonly engineered around a decoupled motion-generation stack: a planner computes a collision-free path and a lower-level controller tracks the resulting reference. This separation is computationally convenient, but it can produce references that are difficult to execute under actuator limits, tracking error, model mismatch, and small obstacle clearances. We present CSymPlan, a certified symbolic planning and control framework for high-DOF manipulators with two complementary implementations: an offline implementation that precomputes certified reach-avoid feedback policies for known workspaces; and an online implementation that synthesizes or updates symbolic policies at runtime from changing task and perception information using parallelization. The offline implementation reduces the manipulator dynamics to a sampled perturbed double-integrator model in operational space through feedback linearization, treats torque-realization errors, modeling inaccuracies, and measurement uncertainty as bounded disturbances, and refines the synthesized symbolic policy to the Franka FR3 through a quantization--lookup--torque realization pipeline. The online implementation uses the same abstraction and refinement interface, but replaces the precomputed policy table with a runtime pFaces request--synthesis--execution loop. In randomized simulated benchmarks and perception-driven Franka FR3 experiments, both implementations complete reach-avoid tasks with zero safety violations; whenever no certified action exists, the robot holds, replans, or stops safely instead of executing an uncertified command.
Chinese Translation
机器人机械臂通常围绕解耦的运动生成框架进行设计:规划器计算无碰撞路径,而低层控制器跟踪生成的参考路径。这种分离在计算上是方便的,但可能会产生在执行时受到驱动器限制、跟踪误差、模型不匹配和小障碍物间隙影响的难以执行的参考。我们提出了CSymPlan,一个针对高自由度机械臂的认证符号规划与控制框架,具有两种互补的实现方式:一种离线实现,预计算已知工作空间的认证避让反馈策略;另一种在线实现,利用并行化从变化的任务和感知信息中合成或更新符号策略。离线实现通过反馈线性化将机械臂的动力学简化为操作空间中的采样扰动双积分模型,将扭矩实现误差、建模不准确性和测量不确定性视为有界干扰,并通过量化-查找-扭矩实现管道对合成的符号策略进行细化,应用于Franka FR3。在线实现使用相同的抽象和细化接口,但用运行时的pFaces请求-合成-执行循环替换了预计算的策略表。在随机化的模拟基准测试和基于感知的Franka FR3实验中,两种实现均以零安全违规完成避让任务;每当不存在认证动作时,机器人会安全地保持、重新规划或停止,而不是执行未认证的命令。
cs.RO / 67 / 2608.22990

InstructMove: A Text-Indispensable Benchmark for Instruction-Following Manipulation

InstructMove:一种对文本不可或缺的指令跟随操作基准
Zhao, Mengao, Li, Ziang, Huang, Chaodong, Ma, Mengchen, Jiang, Haoyi, Jin, Yiwei, Wang, Xinjie, Du, Yun, Lin, Xuewu, Ding, Taojun, Xie, Hongyu, Jiang, Jackson, Yu, Chunlei, Zhang, Kaihua, Huang, Lichao, Liu, Liu, Lin, Tianwei, Su, Zhizhong
Abstract
Vision-language-action (VLA) models have made general-purpose robot manipulation increasingly plausible by conditioning robot actions on natural-language instructions. A key test of such generality is whether policies actually follow language instructions. Yet many manipulation benchmarks leave this ability underdetermined: the intended object or destination is often visually salient or uniquely feasible, allowing policies to succeed without grounding the instruction. We argue that instruction-following evaluation should be text-indispensable: multiple actions should be visually and physically plausible, while only one should be consistent with the language instruction. We introduce InstructMove, a text-indispensable benchmark for instruction-following manipulation. InstructMove instantiates this principle in pick-and-place scenes with semantic distractors, decomposing instruction following into category identification, attribute discrimination, spatial reasoning, and compositional pick-and-place. InstructMove supports a train-eval protocol with InstructMove training data and held-out evaluation tasks, with additional diagnostics for language dependence. Experiments with representative VLA policies show that InstructMove provides a controlled testbed for diagnosing visual shortcuts and that InstructMove simulation data can improve real-world instruction-following manipulation performance. Code: https://github.com/HorizonRobotics/RoboOrchardSim
Chinese Translation
视觉-语言-行动(VLA)模型通过将机器人动作与自然语言指令相结合,使通用机器人操作变得愈加可行。这种通用性的一个关键测试是策略是否真正遵循语言指令。然而,许多操作基准未能充分测试这一能力:预期的对象或目标通常在视觉上显著或唯一可行,使得策略能够在没有依据指令的情况下成功。我们认为,指令跟随评估应该是对文本不可或缺的:多个动作应在视觉和物理上都是合理的,而只有一个动作应与语言指令一致。我们提出了InstructMove,这是一种对文本不可或缺的指令跟随操作基准。InstructMove在具有语义干扰物的取放场景中实现了这一原则,将指令跟随分解为类别识别、属性区分、空间推理和组合取放。InstructMove支持一种训练-评估协议,包含InstructMove训练数据和保留的评估任务,并提供额外的语言依赖性诊断。与代表性VLA策略的实验表明,InstructMove提供了一个控制的测试平台,用于诊断视觉捷径,并且InstructMove模拟数据可以提高现实世界中的指令跟随操作性能。代码链接:https://github.com/HorizonRobotics/RoboOrchardSim
cs.RO / 68 / 2608.23000

Free-Energy-Gated Plasticity for Real-Time Online Motor Learning in Physical Human--Robot Interaction

基于自由能门控可塑性的实时在线运动学习在物理人机交互中的应用
Sawada, Hiroki, Tani, Jun
Abstract
Fully online embodied learning requires synaptic adaptation to acquire new behaviors while preserving previously learned dynamics during ongoing interaction. We extend the Predictive-Coding-inspired Variational Recurrent Neural Network (PV-RNN) to continuously adapt its synaptic weights and propose Free-Energy-Gated Plasticity (FEGP), which regulates the effective learning rate according to variational free energy. In real-time physical human--robot interaction, a randomly initialized network acquired three cyclic motor patterns without offline pretraining, replay, or task-boundary signals, with all three patterns emerging in autonomous rollouts. Controlled experiments over ten randomized teaching streams and five network initializations per stream showed that FEGP substantially improved repertoire coverage and retention of previously acquired patterns after they left the recent observation window. Neither a constant learning rate matched to the gate's time-averaged effective rate nor replay of the same gain values with disrupted temporal organization reproduced these improvements. These results indicate that the temporal allocation of plasticity relative to model--environment mismatch, rather than simply its average magnitude or distribution, is critical for maintaining previously acquired behaviors during continued online learning.
Chinese Translation
完全在线的具身学习需要突触适应,以在持续交互中获取新行为的同时保持先前学习的动态。我们扩展了受预测编码启发的变分递归神经网络(PV-RNN),使其能够持续调整突触权重,并提出了自由能门控可塑性(FEGP),该方法根据变分自由能调节有效学习率。在实时的物理人机交互中,一个随机初始化的网络在没有离线预训练、重放或任务边界信号的情况下,成功获取了三种循环运动模式,所有三种模式均在自主回放中出现。对十个随机教学流和每个流五次网络初始化的控制实验表明,FEGP显著提高了模式覆盖范围和对先前获取模式的保留,即使这些模式已离开最近的观察窗口。无论是与门的时间平均有效率相匹配的恒定学习率,还是以破坏时间组织的方式重放相同增益值,都未能再现这些改进。这些结果表明,相对于模型与环境的不匹配,塑性的时间分配而非其平均幅度或分布,对于在持续在线学习过程中维持先前获取的行为至关重要。
cs.RO / 69 / 2608.23040

RoboRacer Arena: Scaling High-Fidelity Autonomous Racing in Isaac Sim

RoboRacer竞技场:在Isaac Sim中扩展高保真自主赛车
Clement, Mihaela-Larisa, Poks, Agnes, Bartocci, Ezio
Abstract
RoboRacer offers a standardized platform for research using 1:10-scale autonomous vehicles, but the variety of available tracks hinders the process of acquiring policies. Although existing occupancy-grid simulators allow for the quick addition of new maps, they fail to include physical contact, while 3D simulators require each circuit to be implemented as a separate asset, thus limiting their scalability. In order to overcome this issue, we have developed RoboRacer Arena, a system that creates 3D racing environments directly from occupancy maps. Our method starts by using a flood fill algorithm to extract the drivable corridors and to identify the track boundaries, which are then used to establish the barriers. A distance field is calculated to define the collision boundaries. The track surfaces, collision properties, and materials are assembled into a USD stage, which allows for the automated and reproducible generation of the environment in Isaac Sim. The input maps can be obtained from SLAM sessions, from rescaled Formula 1 circuits, or from natural-language descriptions. When the input is based on natural language, we use Gemma 4 31B to generate a track specification without specifying any coordinates or geometry. To guarantee consistency and reproducibility, we apply geometric screening, procedural generation, and raster-level validation. The simulation environments are initialized in a time range of 1.18 to 2.48 seconds, with the initialization time increasing linearly as the raster size increases. In 30 matched trials involving 10 tracks and 3 seeds, 21 maps were generated and all passed validation. RoboRacer Arena currently contains 130 tracks and supports the generation of tracks from natural language. In benchmark tests, the system attains 8,707 vehicle-steps per second when using 256 parallel rigid-body vehicles, excluding the time taken for rendering and policy execution.
Chinese Translation
RoboRacer提供了一个标准化的平台,用于使用1:10比例的自主车辆进行研究,但可用赛道的多样性阻碍了策略的获取过程。尽管现有的占用网格模拟器允许快速添加新地图,但它们未能包含物理接触,而3D模拟器则要求每个赛道作为单独资产实现,从而限制了其可扩展性。为了解决这一问题,我们开发了RoboRacer竞技场,一个直接从占用地图创建3D赛车环境的系统。我们的方法首先使用洪水填充算法提取可行驶的走廊并识别赛道边界,然后利用这些边界建立障碍物。计算距离场以定义碰撞边界。赛道表面、碰撞属性和材料被组装成一个USD舞台,从而实现了在Isaac Sim中环境的自动化和可重复生成。输入地图可以通过SLAM会话、重新缩放的一级方程式赛道或自然语言描述获得。当输入基于自然语言时,我们使用Gemma 4 31B生成赛道规范,而无需指定任何坐标或几何形状。为了保证一致性和可重复性,我们应用几何筛选、程序生成和光栅级验证。模拟环境的初始化时间范围为1.18到2.48秒,初始化时间随着光栅大小的增加而线性增加。在涉及10条赛道和3个种子的30次匹配试验中,生成了21张地图,并且全部通过了验证。RoboRacer竞技场目前包含130条赛道,并支持从自然语言生成赛道。在基准测试中,该系统在使用256个并行刚体车辆时达到了每秒8,707个车辆步骤,未计算渲染和策略执行所需的时间。
cs.RO / 70 / 2608.23068

Switched Turn-based Adaptive Source Seeking Strategy using Estimation and Information-driven Direction of Improvement

基于切换的自适应源寻求策略:利用估计和信息驱动的改进方向
Banerjee, Shubhra, Ghosh, Satadal
Abstract
Source seeking arises in applications such as gas leak localization, radiation monitoring, and environmental surveillance, where the origin of an unknown signal field must be estimated from spatial measurements. In practice, the source location is not directly observable and must be inferred from noisy scalar measurements collected during motion.In robotic source seeking, estimation and motion are closely linked: measurements improve the source estimate, while the chosen trajectory affects the quality of future measurements.Existing loop-based geometric strategies generate feasible motion but do not explicitly use estimation uncertainty to regulate direction updates.This paper presents a loop-based source-seeking framework that combines Extended Kalman Filter (EKF) estimation with Fisher Information Matrix (FIM)-based direction selection. The source estimate is updated during motion, and the heading is changed at loop boundaries using both estimation uncertainty and predicted information gain. A measurement-based stopping condition is used to detect convergence without requiring prior knowledge of the source location.Simulation results under stationary and moving source scenarios demonstrate improved tracking performance and reduced estimation error compared to purely information-driven or estimate-driven strategies.
Chinese Translation
源寻求出现在气体泄漏定位、辐射监测和环境监控等应用中,在这些应用中,必须从空间测量中估计未知信号场的来源。在实际操作中,源位置不可直接观察,必须通过在运动过程中收集的噪声标量测量进行推断。在机器人源寻求中,估计与运动密切相关:测量改善源估计,而所选择的轨迹影响未来测量的质量。现有的基于循环的几何策略产生可行的运动,但并未明确利用估计的不确定性来调节方向更新。本文提出了一种基于循环的源寻求框架,该框架结合了扩展卡尔曼滤波器(Extended Kalman Filter, EKF)估计与基于费舍尔信息矩阵(Fisher Information Matrix, FIM)的方向选择。在运动过程中更新源估计,并在循环边界处利用估计不确定性和预测信息增益来改变航向。使用基于测量的停止条件来检测收敛,而无需事先了解源位置。在静态和移动源场景下的仿真结果表明,与纯粹的信息驱动或估计驱动策略相比,跟踪性能得到了改善,估计误差得到了降低。
cs.RO / 71 / 2608.23100

Shaping the Evolutionary Dynamics of Robot Morphology via Adaptive Control Learning

通过自适应控制学习塑造机器人形态的进化动态
Song, Junru, Yang, Yang, Xu, Yaqing, Wen, Ying, Peng, Wei, Li, Guozhen, Zhou, Wei'en, Yao, Wen
Abstract
Robot co-design via bi-level optimization couples within-lifetime controller learning for fitness evaluation with cross-generational morphological evolution. Prior work has established that well-adapted morphology facilitates faster control learning, a property termed morphological intelligence. Yet how control learning reciprocally shapes morphological evolution remains unexplored. This paper examines both directions for a holistic account of brain-body interplay. We first show that morphological contributions to control learning decouple into two orthogonal dimensions. We formalize the convergence speed as morphological intelligence and identify the performance ceiling as a complementary quantity termed true potential. A concise functional relation is then established to jointly characterize both quantities from individual learning curves, which, when aggregated at the population level, capture evolutionary profiles. Through extensive experiments on simulated voxel-based soft robots, we reveal that premature fitness evaluation systematically underestimates true potential and biases selection towards fast learners. This restricts design space exploration, compromising both optimization efficiency and morphological diversity. Notably, the widely recognized morphological Baldwin effect emerges as an artifact of this bias rather than a general evolutionary tendency. We therefore propose AdaControl, which monitors disproportionate selection for morphological intelligence during evolution and allocates minimally sufficient control learning for unbiased fitness evaluation. With AdaControl, a simple genetic algorithm rivals state-of-the-art generative-model-based co-design methods in discovering diverse high-performing designs while cutting computation by up to 80% versus exhaustive control.
Chinese Translation
机器人共设计通过双层优化将生命周期内的控制器学习与跨代形态进化相结合。先前的研究表明,适应良好的形态能够促进更快的控制学习,这一特性被称为形态智能。然而,控制学习如何相互影响形态进化仍未得到探索。本文从整体上考察了大脑与身体之间的相互作用。我们首先展示了形态对控制学习的贡献可以解耦为两个正交维度。我们将收敛速度形式化为形态智能,并将性能上限识别为一个互补量,称为真实潜力。然后建立了一个简洁的函数关系,以共同表征这两个量,从个体学习曲线中提取,当在群体层面聚合时,捕捉进化特征。通过对模拟体素基础软机器人进行广泛实验,我们揭示了过早的适应性评估系统性地低估了真实潜力,并使选择偏向于快速学习者。这限制了设计空间的探索,妨碍了优化效率和形态多样性。值得注意的是,广泛认可的形态巴尔德温效应实际上是这种偏见的产物,而非一种普遍的进化趋势。因此,我们提出了AdaControl,它在进化过程中监控形态智能的失衡选择,并分配最小足够的控制学习以进行无偏的适应性评估。使用AdaControl,一个简单的遗传算法在发现多样化高性能设计方面与最先进的基于生成模型的共设计方法相媲美,同时计算量相比于全面控制减少了多达80%。
cs.RO / 72 / 2608.23138

Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation

Pointing-VLA:用于视觉-语言-动作操作的类型化空间基础接口
Chen, Xiwen, Li, Zelin, Zhou, Zhiruo, Chen, Huiming, Wang, Chenwei, Zhu, Xiaojun
Abstract
Vision-language-action (VLA) models often expose spatial grounding through autoregressive text coordinates or opaque action tokens, creating brittle interfaces between multimodal reasoning and robot execution. We present Pointing-VLA, a typed hidden-state spatial readout built on Embodied-R1. Geometry-specific heads predict normalized points, object-functional grounding (OFG) heatmaps, and visual trajectories without serializing geometry as text. For the evaluated Bridge/WidowX and physical pick-place deployments, an explicit execution contract assigns PICK to source-conditioned OFG and PLACE to Pointing, providing direct stage-aligned spatial targets. Pointing-VLA achieves SOTA performance on Bridge/WidowX, averaging 72.9\% across the evaluated four-task set without Bridge-specific finetuning under collision-enabled CuRobo execution. Pointing and OFG show complementary strengths across native and cross-dataset evaluations. The OFG/contact readout transfers to NORA-1.5, preserving or improving success while reducing recorded controller time by more than 20$\times$; typed heads are also 6.68--6.90$\times$ faster than Embodied-R1 text decoding on a shared external suite. When integrated as spatial guidance for a $\pi_{0.5}$ action policy, Pointing-VLA raises autonomous real-robot success from 52.7\% to 80.7\% across three visual contexts. These results establish typed spatial readouts as an efficient, inspectable interface between embodied reasoning and robot execution.
Chinese Translation
视觉-语言-动作(VLA)模型通常通过自回归文本坐标或不透明的动作标记暴露空间基础,导致多模态推理与机器人执行之间的脆弱接口。我们提出了Pointing-VLA,一种基于Embodied-R1构建的类型化隐状态空间读取。几何特定头部预测归一化点、对象功能基础(OFG)热图和视觉轨迹,而无需将几何序列化为文本。在评估的Bridge/WidowX和物理拾取-放置部署中,明确的执行合同将PICK分配给源条件的OFG,将PLACE分配给Pointing,从而提供直接的阶段对齐空间目标。Pointing-VLA在Bridge/WidowX上实现了SOTA性能,在评估的四任务集上平均达到72.9\%,在启用碰撞的CuRobo执行下,无需针对Bridge的特定微调。Pointing和OFG在本地和跨数据集评估中显示出互补的优势。OFG/接触读取转移到NORA-1.5,成功率保持或提高,同时记录的控制器时间减少了20倍以上;类型化头部在共享外部套件上的解码速度也比Embodied-R1快6.68到6.90倍。当作为$ ext{π}_{0.5}$动作策略的空间指导集成时,Pointing-VLA将自主真实机器人成功率从52.7 ext{%}提高到80.7 ext{%},覆盖三个视觉上下文。这些结果确立了类型化空间读取作为具身推理与机器人执行之间高效、可检查的接口。
cs.RO / 73 / 2608.23163

Spinning Quadrotor: Hover Thrust Augmentation with Passive Lifting Surfaces

旋转四旋翼:利用被动升力表面增强悬停推力
Parkala, Aniketh, Kandath, Harikumar
Abstract
Conventional multirotor aerial vehicles actively suppress yaw rotation during hover, expending power to maintain a fixed heading despite the fact that yaw regulation is not required for force balance or altitude control. This paper challenges that paradigm by proposing a spinning quadrotor architecture that intentionally operates at a sustained yaw rate, converting power traditionally spent on yaw regulation into useful aerodynamic effects. A dynamic model of the spinning quadrotor is developed, analysis for low Re range is conducted to choose an airfoil for lifting surfaces. Preliminary hardware tests show a 22% reduction in thrust required. These findings suggest that intentional yaw rotation, rather than being suppressed, can be exploited as a design mechanism for efficient and robust multirotor flight.
Chinese Translation
传统的多旋翼空中飞行器在悬停时主动抑制偏航旋转,消耗能量以维持固定航向,尽管偏航调节并不是力平衡或高度控制所必需的。本文挑战了这一范式,提出了一种旋转四旋翼架构,该架构故意以持续的偏航速率运行,将传统上用于偏航调节的能量转化为有用的气动效应。本文开发了旋转四旋翼的动态模型,并对低雷诺数范围进行了分析,以选择升力表面的翼型。初步硬件测试表明所需推力减少了22%。这些发现表明,故意的偏航旋转不仅可以被抑制,还可以作为高效和稳健的多旋翼飞行设计机制加以利用。
cs.RO / 74 / 2608.23204

Guided Riemannian Optimization (GuRO): Bridging Model Predictive Control and Decision Transformers

引导黎曼优化(GuRO):连接模型预测控制与决策变换器
Abdi, Hossein, Dash, Satya Prakash, Sun, Mingfei
Abstract
Decision-making in high-dimensional, nonlinear systems remains a central challenge in robotics. While model-based methods like Model Predictive Control (MPC) offer sample efficiency and interpretability, their performance degrades when the dynamics model is inaccurate or long-horizon predictions are required. Conversely, model-free reinforcement learning (RL) learns policies directly from interaction but suffers from high sample complexity and unstable optimization. Recent advances in sequence modeling have inspired transformer-based decision-making frameworks that can unify MPC and RL, but their training typically faces significant optimization challenges due to highly non-convex loss landscapes. In this work, we propose a novel framework that integrates MPC with RL in a sequence decision-making framework and leverages a curvature-aware optimization to efficiently tackle non-convex loss landscapes. MPC provides predictions of locally optimal trajectories that guide the decision transformer, removing the need for extensive offline pretraining. To address the slow and unstable convergence of traditional optimizers, we train the policy in a Riemannian parameter space using an efficient Riemannian (curvature-aware) method, leading to faster and more robust optimization. We evaluate our framework on high-dimensional quadruped control tasks and demonstrate consistent improvements over strong baselines, including TRPO, SAC, and Online Decision Transformer, achieving higher returns and faster convergence.
Chinese Translation
在高维非线性系统中进行决策仍然是机器人技术中的一个核心挑战。虽然基于模型的方法如模型预测控制(MPC)提供了样本效率和可解释性,但当动态模型不准确或需要长时间预测时,其性能会下降。相反,基于模型的强化学习(RL)直接从交互中学习策略,但面临高样本复杂性和不稳定优化的问题。最近在序列建模方面的进展激发了基于变换器的决策框架,这些框架可以统一MPC和RL,但它们的训练通常面临由于高度非凸损失景观而导致的显著优化挑战。在本研究中,我们提出了一个新颖的框架,将MPC与RL集成在序列决策框架中,并利用曲率感知优化有效应对非凸损失景观。MPC提供局部最优轨迹的预测,指导决策变换器,从而消除了对大量离线预训练的需求。为了解决传统优化器收敛缓慢和不稳定的问题,我们在黎曼参数空间中使用高效的黎曼(曲率感知)方法训练策略,从而实现更快和更稳健的优化。我们在高维四足控制任务上评估了我们的框架,并展示了相较于强基线(包括TRPO、SAC和在线决策变换器)的持续改进,取得了更高的回报和更快的收敛速度。
cs.RO / 75 / 2608.23224

Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation

仅在需要时思考:用于视觉-语言-动作操作的选择性慢路径干预的提示权限控制
Zhou, Zhiruo, Li, Zelin, Chen, Xiwen, Li, Jiazhuo, Wang, Chenwei, Chen, Huiming, Zhu, Xiaojun
Abstract
Retrieval can efficiently and effectively augment a frozen vision--language--action (VLA) policy without retraining, yet retrieved text becomes a control intervention once it enters the executed prompt. In a matched audit, raw appended text reduces mean success from 92.47\% to 3.00\%, while meaningful and length-matched meaningless appends both fail on all 500 states. This result identifies \emph{prompt-form collapse}: changing the instruction form, rather than adding useful semantics, can dominate execution. We introduce TOWN-VLA (Think Only When Needed), a prompt-authority interface that separates candidate generation from permission to alter the policy input. A fixed compatibility rule authorizes a canonical compact instruction; otherwise, the interface restores the original Base prompt exactly. Across 900 audited routes, every route follows this contract: 525 routes recover Base with matching hashes, and all 375 authorized prompts preserve the task signature. On a matched $4\times7$ LIBERO-Plus evaluation with 10{,}030 episodes per method, success rises from 69.5\% to 73.1\% ($+362$ episodes; 95\% CI 1.89--5.45 points), improving on six perturbation axes and all four suites. On a physical PiPER arm with a frozen \pizerofive{} checkpoint, success rises from 52.7\% to 78.7\% over 150 trials per method ($p=3.16\times10^{-6}$). Prompt authority is enforceable for a frozen controller; oracle-free admission calibration is the next deployment target.
Chinese Translation
检索可以高效且有效地增强一个冻结的视觉-语言-动作(VLA)策略,而无需重新训练,但一旦检索的文本进入执行提示,它就成为控制干预。在一次匹配审计中,原始附加文本将平均成功率从92.47\%降低到3.00\%,而有意义的和长度匹配的无意义附加文本在所有500个状态上均失败。这个结果识别了 extit{提示形式崩溃}:改变指令形式,而不是添加有用的语义,可能主导执行。我们引入了TOWN-VLA(仅在需要时思考),一个提示权限接口,它将候选生成与修改策略输入的权限分开。一个固定的兼容性规则授权一个规范的简洁指令;否则,接口将原始Base提示完全恢复。在900条审计路径中,每条路径都遵循这一契约:525条路径通过匹配哈希恢复Base,所有375个授权提示保留了任务特征。在与10,030个每种方法的回合进行的匹配$4 imes7$ LIBERO-Plus评估中,成功率从69.5\\%上升到73.1\\%(增加362个回合;95\\%置信区间1.89--5.45点),在六个扰动轴和所有四个套件上都有所改善。在一个冻结的 extit{pizerofive}检查点的物理PiPER臂上,成功率在每种方法的150次试验中从52.7\\%上升到78.7\\%($p=3.16 imes10^{-6}$)。对于一个冻结的控制器,提示权限是可执行的;无oracle的入场校准是下一个部署目标。
cs.RO / 76 / 2608.23304

Design of a Biomimetic Joint-Covering Skin with Tissue-Like Structure to Enhance Proprioception in a Musculoskeletal Humanoid

设计一种仿生关节覆盖皮肤,具有类组织结构以增强肌肉骨骼类人形态的本体感觉
Miki, Akihiro, Hasegawa, Shun, Ribayashi, Yoshimoto, Kawaharazuka, Kento, Okada, Kei
Abstract
Proprioception in musculoskeletal humanoids is typically estimated primarily from muscle sensing, while the role of cutaneous deformation around joints remains insufficiently explored. In biological systems, mechanoreceptors distributed within soft tissue complement muscle feedback and support reliable joint state estimation. This study presents the design of a biomimetic joint-covering skin with a tissue-like layered structure that integrates pressure- and stretch-sensitive elements within the joint-covering tissue. The proposed skin is implemented on the musculoskeletal humanoid Musashi-W, and its independent proprioceptive capability as well as its integration with muscle sensing are evaluated. Experimental results show that the proposed skin alone achieves joint angle estimation with an average error of approximately 3 degrees. Furthermore, integration with muscle sensing improves estimation accuracy. Owing to its joint-covering structure, the skin may mechanically mitigate the influence of external disturbances on the muscles, and the integration of multiple modalities suggests the possibility of contributing to the identification of external stimuli that are difficult to interpret using muscle sensing alone. This work presents a design methodology for biomimetic joint-covering skin and demonstrates that such tissue-structured skin can serve as an effective approach for extending proprioceptive systems in musculoskeletal humanoids.
Chinese Translation
肌肉骨骼类人形态的本体感觉通常主要通过肌肉传感来估计,而关节周围皮肤变形的作用尚未得到充分探索。在生物系统中,分布在软组织中的机械感受器补充了肌肉反馈,支持可靠的关节状态估计。本研究提出了一种具有类组织分层结构的仿生关节覆盖皮肤的设计,该皮肤在关节覆盖组织中集成了压力和拉伸敏感元件。所提出的皮肤应用于肌肉骨骼类人形态Musashi-W,并评估了其独立的本体感觉能力及其与肌肉传感的集成。实验结果表明,单独使用所提出的皮肤可以实现关节角度估计,平均误差约为3度。此外,与肌肉传感的集成提高了估计精度。由于其关节覆盖结构,该皮肤可能在机械上减轻外部干扰对肌肉的影响,并且多模态的集成表明其有助于识别难以仅通过肌肉传感解释的外部刺激的可能性。本研究提出了一种仿生关节覆盖皮肤的设计方法,并证明这种类组织结构的皮肤可以作为扩展肌肉骨骼类人形态本体感觉系统的有效方法。
cs.RO / 77 / 2608.23320

ROS2SmolVLA: Enabling Small Vision-Language-Action Models for Integration into Industrial-Grade Lightweight Robots

ROS2SmolVLA:为工业级轻量级机器人集成小型视觉-语言-动作模型提供支持
Mandischer, Nils, Böckmann, Noah, Holl, Ludwig, Mikelsons, Lars
Abstract
Industrial demand changes the paradigms of production. Due to smaller batch sizes and more variations in products, companies face a growing challenge to adopt more adaptive production systems. In particular, robot-based automation is usually static and fails to respond to constantly changing processes. Vision-Language-Action (VLA) Models are a promising opportunity to mitigate this challenge by generating robot actions based on the observed system state. However, current research either focuses on large models that cannot be computed on premise, creating compliance and security challenges, or use lab-grade robot hardware that obscures exploitation in real industrial settings. In this work, we adapt Hugging Face's SmolVLA for Universal Robots lightweight robots. Further, we release the open-source repository ROS2SmolVLA that implements an interface for ROS 2 to SmolVLA, and makes it applicable for industrial-grade hardware. By this, we allow a lenient adoption into lab and industrial environments. We validate the functionality of SmolVLA for a Universal Robots UR10e using a pick-and-place task and give implementation guidelines. Our findings support that SmolVLA is a well-suited option for small-sized tasks that need to be computed on premise.
Chinese Translation
工业需求改变了生产范式。由于批量规模更小和产品种类更多,公司面临着采用更具适应性的生产系统的日益挑战。特别是,基于机器人的自动化通常是静态的,无法应对不断变化的过程。视觉-语言-动作(VLA)模型为缓解这一挑战提供了一个有前景的机会,通过根据观察到的系统状态生成机器人动作。然而,目前的研究要么专注于无法在现场计算的大型模型,从而带来合规性和安全性挑战,要么使用实验室级机器人硬件,这使得在真实工业环境中的应用变得困难。在本研究中,我们将Hugging Face的SmolVLA适配于Universal Robots的轻量级机器人。此外,我们发布了开源库ROS2SmolVLA,该库实现了ROS 2与SmolVLA的接口,使其适用于工业级硬件。通过这一方式,我们允许在实验室和工业环境中宽松地采用该技术。我们使用拾取和放置任务验证了SmolVLA在Universal Robots UR10e上的功能,并提供了实施指南。我们的研究结果支持SmolVLA是适合需要在现场计算的小型任务的良好选择。
cs.RO / 78 / 2608.23354

OptiSight: Bridging Semantic Reasoning and Geometric Control for Embodied Navigation

OptiSight:连接语义推理与几何控制的具身导航
Avan, Alperen, Sanchez-Riera, Jordi
Abstract
Autonomous indoor navigation requires both semantic understanding and precise geometric control. We propose OptiSight, a hybrid framework that combines Vision-Language Model reasoning with deterministic visual servoing through a finite-state Chain-of-Thought architecture. Grounded-SAM localizes open-vocabulary targets, while camera projection geometry converts visual observations into navigation commands without requiring dense mapping. The VLM is queried only at key decision points, reducing computational overhead while geometric control handles continuous navigation. Experiments in AI Habitat demonstrate reliable zero-shot navigation across diverse indoor scenarios, including obstacle avoidance and semantic ambiguity, while operating within an 8~GB VRAM budget. The source code is available at https://github.com/avanalperen/OptiSight-Python-Multimodal-CoT-for-Visual-Reasoning.
Chinese Translation
自主室内导航既需要语义理解,又需要精确的几何控制。我们提出了OptiSight,一个混合框架,结合了视觉-语言模型(Vision-Language Model)推理与通过有限状态链式思维(Chain-of-Thought)架构的确定性视觉伺服。Grounded-SAM能够定位开放词汇目标,而相机投影几何将视觉观察转换为导航指令,无需密集映射。视觉-语言模型仅在关键决策点进行查询,从而减少计算开销,而几何控制则处理连续导航。在AI Habitat中的实验表明,OptiSight在多样的室内场景中实现了可靠的零-shot导航,包括避障和语义模糊,同时在8~GB的显存预算内运行。源代码可在 https://github.com/avanalperen/OptiSight-Python-Multimodal-CoT-for-Visual-Reasoning 获取。
cs.RO / 79 / 2608.23452

Reward-Free Continual Adaptation for Resilient Space Robots

无奖励持续适应的弹性空间机器人
Orsula, Andrej, Olivares-Mendez, Miguel, Martinez, Carol
Abstract
Space robots operate in extreme environments where hardware degradation can critically compromise traditional control strategies. While continual reinforcement learning offers a promising mechanism for online adaptation, it inherently requires access to a reward signal during deployment. However, precise reward computation in space is often infeasible due to the lack of external tracking systems and the overall complexity of the environment. To address the challenge of unobservable rewards, we introduce a reward-free continual learning framework that leverages latent-state world models. By pre-training a model-based agent across diverse simulations, the world model learns a robust predictor of the reward structure within its latent space. Upon deployment to an environment with severe hardware degradation, we freeze the observation encoder and reward predictor to update only the transition dynamics of the world model through unsupervised rollouts. By training the policy entirely on imagined trajectories generated by this updated world model, the agent adapts to altered dynamics without receiving new rewards. We demonstrate our approach across simulated planetary traversal, orbital navigation, and precision assembly tasks subjected to severe morphological failures.
Chinese Translation
空间机器人在极端环境中操作,硬件退化可能严重影响传统控制策略。尽管持续强化学习提供了一种有前景的在线适应机制,但其在部署过程中本质上需要访问奖励信号。然而,由于缺乏外部跟踪系统和环境的整体复杂性,在太空中精确计算奖励往往是不可行的。为了解决不可观察奖励的挑战,我们提出了一种无奖励持续学习框架,该框架利用潜在状态世界模型。通过在多样化的模拟中对基于模型的智能体进行预训练,世界模型学习到其潜在空间中奖励结构的稳健预测器。在部署到一个存在严重硬件退化的环境时,我们冻结观察编码器和奖励预测器,仅通过无监督的滚动更新世界模型的转移动态。通过完全基于这个更新后的世界模型生成的想象轨迹来训练策略,智能体在没有接收新奖励的情况下适应改变的动态。我们在模拟的行星穿越、轨道导航和精密组装任务中展示了我们的方法,这些任务经历了严重的形态失效。
cs.RO / 80 / 2608.23478

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

有意行动:为视觉-语言-行动模型提炼行为意图
Lee, Sangoh, Mo, Sangwoo, Han, Wook-Shin
Abstract
Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which motor command was demonstrated while leaving implicit the local objective served by the behavior under the instruction. Future-based supervision enriches action learning with frames, latent observations, trajectories, or motion representations, but these signals capture particular realizations of what may happen rather than the shared semantic objective of the forthcoming behavior. We propose Intention Distillation (INDI), which distills behavior-level intent into the action decoder. During training, a frozen teacher VLM interprets a demonstrated segment from the current observation, instruction, coarse action summary, and corresponding execution video. From its standard inputs, the deployed VLA recovers the resulting multimodal intent representation at an intermediate decoder layer and uses it to organize action prediction together with representations of how the behavior unfolds and what it achieves. On SimplerEnv-Bridge, INDI improves GR00T-N1.7 from 64.3% to 84.7%, and on RoboCasa Kitchen it improves the controlled GR00T-N1.7 baseline from 64.1% to 70.3%, with consistent gains on $\pi_{0.5}$ across both benchmarks. In real-world tasks, INDI improves average success from 62.0% to 68.7%, with gains of up to 12.0 pp on longer-horizon tasks. Further analyses show that the recovered latent is used by the decoder, captures behavior objective and execution progress, and organizes downstream predictions in an objective-dependent manner. These results show that action decoders benefit from explicitly modeling the semantic objective of the behavior they generate.
Chinese Translation
视觉-语言-行动(VLA)模型能够将多模态上下文转化为机器人动作,但其动作解码器仍主要通过行为克隆进行训练。这种方式监督了所展示的运动指令,同时隐含了行为在指令下所服务的局部目标。基于未来的监督通过框架、潜在观察、轨迹或运动表示丰富了动作学习,但这些信号捕捉的是可能发生的特定实现,而非即将发生的行为的共享语义目标。我们提出了意图提炼(Intention Distillation,INDI),该方法将行为级别的意图提炼到动作解码器中。在训练过程中,一个冻结的教师视觉语言模型(VLM)解释当前观察、指令、粗略动作摘要和相应执行视频中的展示片段。通过其标准输入,部署的VLA在中间解码器层恢复结果的多模态意图表示,并利用该表示来组织动作预测,同时结合行为展开的方式和所实现的目标。在SimplerEnv-Bridge上,INDI将GR00T-N1.7的性能从64.3%提升至84.7%;在RoboCasa Kitchen上,将受控的GR00T-N1.7基线从64.1%提升至70.3%,在两个基准测试中,$ ext{π}_{0.5}$均表现出一致的提升。在现实世界任务中,INDI将平均成功率从62.0%提升至68.7%,在较长时间任务中提升幅度达到12.0个百分点。进一步分析表明,恢复的潜在表示被解码器使用,捕捉了行为目标和执行进展,并以目标依赖的方式组织下游预测。这些结果表明,动作解码器从明确建模其生成行为的语义目标中获益。
计算机视觉 (Computer Vision)
185
cs.CV / 1 / 2608.21422

Topology of a Smile: Persistent Homology in Dental Imaging

微笑的拓扑:牙科成像中的持久同调
Dahlmeier, Leon, Kališnik, Sara, Mehl, Albert, Rieck, Bastian
Abstract
CBCT (Cone Beam Computed Tomography) scans provide detailed three-dimensional images, widely used in dentistry for diagnostic and treatment planning tasks. While invaluable, analyzing and documenting these scans is labor-intensive, prompting efforts to automate key steps like the classification and segmentation of anatomical structures to identify tooth types and associated pathologies. In this article, we propose an approach to automation that leverages persistent homology, a framework from topological data analysis that studies the shape of data by identifying features like connected components, holes, and voids across multiple scales. Persistent homology, together with a support vector machine, allows us to classify teeth in a CBCT scan and to perform diagnostics. Our method advances the state of the art, reaching average accuracy scores of 97.67% for tooth-labeling and 96.77% for diagnostic tasks, outperforming a CNN trained on the same data with accuracy of 70.27% and 86.67%, respectively.
Chinese Translation
锥束计算机断层扫描(CBCT)提供了详细的三维图像,广泛应用于牙科的诊断和治疗规划任务。尽管其价值不可估量,但分析和记录这些扫描过程劳动密集,促使我们努力自动化关键步骤,如解剖结构的分类和分割,以识别牙齿类型及相关病理。在本文中,我们提出了一种利用持久同调的自动化方法,持久同调是拓扑数据分析中的一个框架,通过识别多个尺度上的连通分量、孔洞和空隙等特征来研究数据的形状。持久同调结合支持向量机,使我们能够对CBCT扫描中的牙齿进行分类并进行诊断。我们的方法推动了该领域的技术进步,牙齿标记的平均准确率达到97.67%,诊断任务的准确率为96.77%,均优于在相同数据上训练的卷积神经网络(CNN),其准确率分别为70.27%和86.67%。
cs.CV / 2 / 2608.21424

EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing

EditStream:一个统一的自回归框架用于交互式视频生成与编辑
Zhou, Yuqian, Zhou, Zhenghong, Wu, Zongze, Smith, Cameron, Zhang, Richard, Luo, Jiebo, Shechtman, Eli, Lin, Zhe
Abstract
Interactive video generation and editing are becoming increasingly important for creative design. In this report, we introduce EditStream: a unified framework for interactive video generation and editing. EditStream unifies multiple video creation and manipulation tasks within a single DiT-based model through flexible task-specific conditioning, and further transforms it into a fast, few-step autoregressive model for efficient streaming. It supports Text-to-Video, Image-to-Video, Video-to-Video, Editing Propagation, Reference-guided Video Editing, and Camera Pose Change, enabling flexible control over video generation, transformation, and editing within one system. To make the unified model practical for interactive use, we develop a two-stage distillation approach that combines Velocity Moment Matching (VMM) with autoregressive unrolling. VMM matches conditional velocity moments at student-reached intermediate states to preserve generation quality and motion, while unrolling exposes the student to its own autoregressive predictions to improve temporal stability. Together, they alleviate common challenges in few-step autoregressive video generation, including over-saturation, degraded motion, temporal instability, and complex training. EditStream provides a practical and scalable solution that bridges high-quality diffusion-based video models with interactive creative workflows.
Chinese Translation
交互式视频生成与编辑在创意设计中变得越来越重要。在本报告中,我们介绍了EditStream:一个用于交互式视频生成与编辑的统一框架。EditStream通过灵活的任务特定条件,将多个视频创作与操作任务统一在一个基于DiT(Denoising Transformer)的模型中,并进一步将其转化为一个快速的、少步的自回归模型,以实现高效流媒体处理。它支持文本到视频(Text-to-Video)、图像到视频(Image-to-Video)、视频到视频(Video-to-Video)、编辑传播(Editing Propagation)、参考引导视频编辑(Reference-guided Video Editing)和摄像机姿态变化(Camera Pose Change),使得在一个系统内对视频生成、转换和编辑进行灵活控制成为可能。为了使统一模型在交互式使用中更具实用性,我们开发了一种两阶段蒸馏方法,将速度矩匹配(Velocity Moment Matching, VMM)与自回归展开相结合。VMM在学生模型达到的中间状态下匹配条件速度矩,以保持生成质量和运动,而展开则使学生模型接触其自身的自回归预测,以提高时间稳定性。两者共同缓解了少步自回归视频生成中的常见挑战,包括过饱和、运动退化、时间不稳定性和复杂训练。EditStream提供了一个实用且可扩展的解决方案,架起了高质量扩散基础视频模型与交互式创意工作流程之间的桥梁。
cs.CV / 3 / 2608.21425

Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation

人类感知的对齐:用于视频生成的校准分布奖励学习
Zhai, Nai-Xin, Cheng, Weihua, Yu, Dexu, Gu, Yikai, Du, Hanwen, Fu, Junchen, Huang, Chenxi, Song, Yingwei, Ma, Liyuan Lillian, Ran, Yang, Li, Youhua, Ni, Yongxin
Abstract
Video generation is central to AI-powered content creation. Aligning generated videos with human preferences is a key criterion for evaluating generation quality. Despite significant progress in visual quality, three key challenges remain. First, the reliability of reward signals is constrained by the quality of human preference data, which is often affected by subjective noise and bias. Second, standard scalar reward models collapse multi-aspect human preferences into a single value, leading to the loss of dynamic trade-offs across multiple preference dimensions. Third, in policy optimization, the widely adopted KL divergence imposes primarily local constraints and may fail to capture the global structure of human preferences. To address these challenges, we propose a unified preference-aware learning framework for video generation. First, we introduce elite-guided filtering to calibrate preference data and construct reliable supervision for reward model training. We then model video quality as a multidimensional reward distribution to capture the uncertainty inherent in human preferences, and use the Wasserstein distance to align the learned reward distribution with the empirical human preference distribution. Finally, we introduce Wasserstein-based distributional alignment into GRPO, guiding policy optimization to better match the global structure of human preferences over videos. Experiments on reward modeling and video generation demonstrate that our approach improves the reliability of reward signals and the perceptual consistency of generated videos. Our code is available at https://github.com/alignhs26/ahs.
Chinese Translation
视频生成是人工智能驱动内容创作的核心。将生成的视频与人类偏好对齐是评估生成质量的关键标准。尽管在视觉质量上取得了显著进展,但仍然面临三大挑战。首先,奖励信号的可靠性受到人类偏好数据质量的限制,而这些数据常常受到主观噪声和偏见的影响。其次,标准的标量奖励模型将多方面的人类偏好压缩为单一值,导致在多个偏好维度之间动态权衡的丧失。第三,在策略优化中,广泛采用的KL散度主要施加局部约束,可能无法捕捉人类偏好的全局结构。为了解决这些挑战,我们提出了一种统一的偏好感知学习框架用于视频生成。首先,我们引入精英引导过滤来校准偏好数据,并为奖励模型训练构建可靠的监督。然后,我们将视频质量建模为多维奖励分布,以捕捉人类偏好中固有的不确定性,并使用Wasserstein距离将学习到的奖励分布与经验人类偏好分布对齐。最后,我们将基于Wasserstein的分布对齐引入GRPO,指导策略优化更好地匹配视频的人类偏好的全局结构。在奖励建模和视频生成的实验中,我们的方法提高了奖励信号的可靠性和生成视频的感知一致性。我们的代码可在 https://github.com/alignhs26/ahs 获取。
cs.CV / 4 / 2608.21426

AI Visual Inspection for Garment Production

服装生产中的人工智能视觉检测
Kong, Ray Wai Man, Ning, Ding, Kong, Theodore Ho Tin
Abstract
The garment manufacturing industry is under increasing pressure to improve product quality, reduce costs, and accelerate digital transformation toward Industry 4.0. One of the most challenging quality-control activities is sewing-line inspection, where defects such as broken stitches and skipped stitches are difficult to detect consistently through manual inspection. Human-based inspection is often affected by fatigue, subjective judgement, and inconsistent performance, resulting in defect leakage, rework, and reduced production efficiency. This study presents the development and validation of an Artificial Intelligence (AI)-based visual inspection system for garment sewing-line quality control. The system utilizes Convolutional Neural Networks (CNNs) to detect sewing defects and was initially trained using black fabric and black sewing thread samples. Experimental testing was conducted on black, red, dark green, light blue, silver, and fluorescent yellow fabrics. The results demonstrated successful detection of jump sewing-line defects on black, red, and dark green materials, while performance limitations were observed for broken sewing-line defects and fabrics with significantly different visual characteristics, including light blue, silver, and fluorescent yellow colours. These findings indicate that model accuracy is strongly influenced by the diversity of training data and the ability to generalize across different fabric and thread colours.
Chinese Translation
服装制造行业面临着提高产品质量、降低成本以及加速向工业4.0数字化转型的日益压力。其中,缝合线检查是最具挑战性的质量控制活动之一,手动检查难以一致地检测到如断线和漏缝等缺陷。基于人工的检查往往受到疲劳、主观判断和表现不一致的影响,导致缺陷漏检、返工和生产效率降低。本研究提出了一种基于人工智能(AI)的视觉检测系统,用于服装缝合线质量控制的开发和验证。该系统利用卷积神经网络(CNN)来检测缝合缺陷,并最初使用黑色面料和黑色缝线样本进行训练。在黑色、红色、深绿色、浅蓝色、银色和荧光黄色面料上进行了实验测试。结果表明,该系统成功检测到了黑色、红色和深绿色材料上的跳缝缺陷,而在断缝缺陷和具有显著不同视觉特征的面料(包括浅蓝色、银色和荧光黄色)上则观察到了性能限制。这些发现表明,模型的准确性受到训练数据多样性和在不同面料及缝线颜色上泛化能力的强烈影响。
cs.CV / 5 / 2608.21427

Few-Shot Cross-Dataset Adaptation for Tuberculosis Detection Using DenseNet

基于 DenseNet 的肺结核检测的少样本跨数据集适应
Biswas, Bidhan, Sohag, Shahadat Hossain, Ashab, Nabil, Kundu, Soumit Kumar, Parvez, Saif Mahmud
Abstract
Tuberculosis (TB) is one of the most common and dangerous bacterial ailments. Every year, it causes a large number of deaths worldwide. Although many deep learning models can detect tuberculosis from chest X-rays quite accurately, severe domain shift across datasets makes the task challenging. Different imaging protocols, patient demographics, and equipment across domains make the task of generalization difficult. In real-world settings, a model may perform well on one dataset but show a noticeable drop in performance when tested on another. In this work, we address this domain adaptation challenge through a few-shot scaling study. A controlled cross-dataset evaluation is presented in this paper using TBX11K as the source domain and the Mendeley TB dataset as the target domain. It is investigated how varying the number of target samples affects model performance under three training regimes: frozen backbone adaptation, full fine-tuning of a source-pretrained DenseNet121 model, and training from scratch. The results indicate that the model can perform well even with limited data and can achieve 98.36\% accuracy with just 75 labeled samples per class. The adaptation curves demonstrate how fine-tuning effectively mitigates domain shift. These findings establish full fine-tuning of pretrained models as a highly effective and practical strategy for mitigating domain shift in low-resource clinical deployment scenarios.
Chinese Translation
肺结核(TB)是最常见和危险的细菌疾病之一。每年,它在全球造成大量死亡。尽管许多深度学习模型能够相当准确地从胸部 X 光片中检测肺结核,但不同数据集之间的严重领域转移使得这一任务变得具有挑战性。不同的成像协议、患者人口统计特征和设备使得泛化任务变得困难。在实际环境中,一个模型在一个数据集上表现良好,但在另一个数据集上测试时可能会显著下降。在本研究中,我们通过少样本扩展研究来解决这一领域适应挑战。本文展示了一个受控的跨数据集评估,使用 TBX11K 作为源域,Mendeley TB 数据集作为目标域。我们研究了在三种训练模式下,目标样本数量的变化如何影响模型性能:冻结主干适应、对源预训练的 DenseNet121 模型进行全面微调,以及从头开始训练。结果表明,即使在数据有限的情况下,模型也能够表现良好,并且仅使用每类 75 个标记样本就能达到 98.36\% 的准确率。适应曲线展示了微调如何有效缓解领域转移。这些发现确立了对预训练模型进行全面微调作为在低资源临床部署场景中缓解领域转移的高度有效和实用的策略。
cs.CV / 6 / 2608.21429

Measuring Gender Representation in Animated Films

测量动画电影中的性别表现
Bamman, David, Cooper, Allison, Rubio, Ruby Alvarez, Kushihashi, Reina, Mar, Madison
Abstract
Animated films--often developed with an audience of children in mind--are an important vector for enculturation, and empirical work that has examined the representation of gender at scale in these films has largely focused on counting the gender composition of the cast rather than deploying a more fine-grained instrument (such as assessing the visibility of those characters in overall screentime). In this work, we develop a computational pipeline for recognizing animated characters in these films, and use it to test several hypotheses about gender representation in a corpus of 224 popular animated movies. We find that while the overall representation of female characters in animated films largely tracks with those of live-action films (over the period 1980-2025), we see stark differences between the representation of human characters (much greater representation among women and girls) and non-humans (largely male). Contrary to past work on Disney, we do not see female characters declining in antagonist roles in animated films, and characters who are women and girls are much more likely to share scenes together than their live action contemporaneous counterparts.
Chinese Translation
动画电影通常是以儿童观众为目标的重要文化传播媒介,而对这些电影中性别表现的实证研究主要集中在统计演员性别构成上,而非采用更细致的工具(例如评估这些角色在整体屏幕时间中的可见性)。在本研究中,我们开发了一种计算管道,用于识别这些电影中的动画角色,并利用该管道测试关于224部流行动画电影中性别表现的若干假设。我们的研究发现,尽管动画电影中女性角色的整体表现与真人电影(1980-2025年期间)大致相符,但我们观察到人类角色(女性和女孩的表现更为突出)与非人类角色(主要为男性)之间存在显著差异。与以往关于迪士尼的研究相反,我们没有发现动画电影中女性角色在反派角色中的比例下降,且女性和女孩角色更有可能共同出现在场景中,而不是与其同时期的真人电影角色相比。
cs.CV / 7 / 2608.21431

Boosting Knowledge-based Visual Question Answering with Structured Context Reasoning

基于结构化上下文推理的知识驱动视觉问答增强
Liu, Qiyou, Zhang, Yong, Luo, Jianjie, Yang, Zhenguo, Yu, Yi
Abstract
Knowledge-based Visual Question Answering aims to answer questions about an image by integrating external knowledge with visual and textual information. Recent approaches often rely on in-context learning to prompt Large Language Models (LLMs) with multimodal context in a zero-shot or few-shot manner. However, we observe that directly concatenating heterogeneous visual descriptions and retrieved knowledge into long, unstructured prompts often degrades reasoning performance, due to both excessive irrelevant context and the lack of explicit relational structure. In this paper, we propose an LLM-based Structured Context Reasoning (SCoRe) framework that infers both explicit and implicit relationships for prediction. SCoRe consists of three stages: Context Acquisition, which generates diverse visual notes and retrieves explicit knowledge via an efficient two-stage multimodal retrieval strategy; Context Selection, which filters relevant visual, explicit, and implicit knowledge using LLM-guided selection; and Context Compression, which performs Relational Logic Distillation (RLD) to transform raw text into explicit entity-relation triplets. These relational triplets serve as a concise and structured prompt for final answer prediction. Extensive experiments on the OK-VQA and A-OKVQA benchmarks demonstrate that SCoRe consistently outperforms state-of-the-art methods.
Chinese Translation
知识驱动的视觉问答旨在通过整合外部知识与视觉及文本信息来回答关于图像的问题。近期方法通常依赖于上下文学习,以零样本或少样本方式通过多模态上下文提示大型语言模型(LLMs)。然而,我们观察到,直接将异构的视觉描述和检索到的知识拼接成冗长且无结构的提示,往往会因包含过多无关上下文且缺乏显式关系结构而降低推理性能。本文提出了一种基于LLM的结构化上下文推理框架(Structured Context Reasoning,SCoRe),该框架同时推断显式和隐式关系以进行预测。SCoRe包含三个阶段:上下文获取,通过高效的两阶段多模态检索策略生成多样化的视觉笔记并检索显式知识;上下文选择,利用LLM引导筛选相关的视觉、显式及隐式知识;上下文压缩,执行关系逻辑蒸馏(Relational Logic Distillation,RLD),将原始文本转化为显式的实体-关系三元组。这些关系三元组作为简洁且结构化的提示用于最终答案预测。在OK-VQA和A-OKVQA基准上的大量实验表明,SCoRe持续优于现有最先进方法。
cs.CV / 8 / 2608.21438

DesignAgent3D: Interactive 3D Scene Editing via Designer-like Multimodal Reasoning

DesignAgent3D:通过设计师般的多模态推理进行交互式3D场景编辑
Liu, Xiujin, Yang, Tianyu, Zhao, Yilun, Zhang, Xiangliang
Abstract
Text guided 3D scene editing provides an intuitive interface for modifying reconstructed environments, but remains difficult because natural language design requests are often semantically underspecified and must be grounded in cluttered 3D scenes. Existing methods typically formulate the task as one-shot conditional generation from a single prompt, failing to resolve ambiguous user intents or achieve precise spatial grounding. Consequently, they suffer from severe object localization drift, tracking failure under occlusions, and the notorious multi-view "sticker effect." To overcome these limitations, we present DesignAgent3D, an interactive multimodal agentic framework that reformulates 3D scene editing as a designer-like Plan-Perceive-Act paradigm. The agent first plans by interacting with the user to clarify underspecified design goals, then perceives by grounding the intended edit to specific objects or regions in the 3D scene, and finally acts by applying controlled visual modifications while preserving scene consistency. The edits are further integrated into the underlying 3D representation, supporting persistent and multi-view consistent novel-view rendering. Extensive experiments across both NeRF and 3D Gaussian Splatting backbones demonstrate that DesignAgent3D significantly outperforms state-of-the-art baselines, delivering superior semantic intent alignment, impeccable spatial localization accuracy, and high-fidelity multi-view consistency.
Chinese Translation
文本引导的3D场景编辑为修改重建环境提供了直观的接口,但由于自然语言设计请求通常语义不明确且必须在杂乱的3D场景中进行定位,这一过程仍然困难。现有方法通常将任务表述为从单一提示进行一次性条件生成,未能解决模糊的用户意图或实现精确的空间定位。因此,它们面临严重的物体定位漂移、遮挡下的跟踪失败以及臭名昭著的多视角“贴纸效应”。为克服这些局限性,我们提出了DesignAgent3D,一个交互式多模态代理框架,将3D场景编辑重新表述为设计师般的计划-感知-行动(Plan-Perceive-Act)范式。该代理首先通过与用户互动来澄清不明确的设计目标进行规划,然后通过将预期的编辑与3D场景中的特定物体或区域进行定位来进行感知,最后通过应用受控的视觉修改来保持场景一致性。编辑结果进一步集成到基础的3D表示中,支持持久且多视角一致的新视图渲染。在NeRF和3D高斯溅射(3D Gaussian Splatting)基础上进行的广泛实验表明,DesignAgent3D显著优于最先进的基线,提供了卓越的语义意图对齐、无可挑剔的空间定位精度和高保真的多视角一致性。
cs.CV / 9 / 2608.21439

WorldMind: Decoupled Game World Model for State-Aware NPC Behavior

WorldMind:用于状态感知NPC行为的解耦游戏世界模型
Deng, Zhiyang, Zhang, Boran, Chen, Danze, Jin, Yeying
Abstract
Game world models have recently demonstrated promising capabilities in generating visually coherent and action-controllable gameplay videos. However, non-player character (NPC) behavior in existing models is either implicitly entangled with video generation or explicitly prescribed through external control signals. Consequently, a game world model has to jointly understand the state, plan the NPC's response and render its visual outcome, limiting its ability to produce responsive and state-aware NPC behavior. The challenge lies in the lack of an explicit interface for state-grounded decision-making. To this end, we introduce WorldMind, to our knowledge the first decoupled framework for state-aware NPC behavior in game world models. WorldMind separates interactive world modeling into four layers: an Understanding Layer that constructs a compact state from generated frames; a Decision Layer that reasons over the compact state to plan the NPC's next action; a Control Layer that translates the actions into temporally aligned conditions; and a Generation Layer that synthesizes their visual outcomes. By reconnecting layers in a closed interaction loop, WorldMind grounds NPC behavior in the evolving game state. We further introduce BOSS-140K, a dataset of gameplay videos paired with rich internal game states, together with an agent that automates the collection at scale. Experiments on BOSS-140K demonstrate reliable compact state reconstruction and mechanics-grounded planning, with WorldMind preferred over the baselines in approximately 70% of pairwise comparisons for its more tactically appropriate and coherent NPC behavior. Project page: https://teawhite.cn/worldmind_projectpage/
Chinese Translation
游戏世界模型最近在生成视觉连贯且可控的游戏视频方面展现了良好的能力。然而,现有模型中的非玩家角色(NPC)行为要么隐式地与视频生成纠缠在一起,要么通过外部控制信号显式规定。因此,游戏世界模型必须共同理解状态、规划NPC的反应并渲染其视觉结果,这限制了其生成响应性和状态感知NPC行为的能力。挑战在于缺乏一个明确的接口用于基于状态的决策。为此,我们提出了WorldMind,据我们所知,这是第一个用于游戏世界模型中状态感知NPC行为的解耦框架。WorldMind将交互式世界建模分为四个层次:理解层(Understanding Layer)从生成的帧构建紧凑状态;决策层(Decision Layer)在紧凑状态上进行推理以规划NPC的下一步动作;控制层(Control Layer)将动作转换为时间对齐的条件;生成层(Generation Layer)合成它们的视觉结果。通过在闭环交互中重新连接各层,WorldMind将NPC行为与不断变化的游戏状态相结合。我们进一步引入了BOSS-140K,这是一个与丰富内部游戏状态配对的游戏视频数据集,并提供一个能够大规模自动收集数据的代理。对BOSS-140K的实验展示了可靠的紧凑状态重建和基于机制的规划,WorldMind在约70%的成对比较中优于基线,因其生成了更具战术适应性和连贯性的NPC行为。项目页面:https://teawhite.cn/worldmind_projectpage/
cs.CV / 10 / 2608.21443

Text-Guided Visual Dependency Graph Learning with Cross-Modal Attention Priors

基于文本引导的视觉依赖图学习与跨模态注意力先验
Wang, Fei, Zhang, Yutong, Ye, Yang, Chen, Jinxian, Wenshuai, Wang, Wang, Xiong
Abstract
Estimating interpretable conditional-dependence structures from multimodal visual-linguistic features remains largely unexplored. We propose CM-GLasso (Cross-Modal Graphical Lasso), a framework that bridges vision-language representation learning and sparse Gaussian Graphical Models. CM-GLasso introduces three key components: (i) a text visualization strategy that renders class-attribute descriptions as images and processes them through the same SigLIP-2 vision encoder as natural images, yielding prototype-indexed patch-level attention footprints in a shared feature coordinate system; (ii) a cross-attention distillation mechanism that condenses high-dimensional patches into a small set of semantic graph nodes, whose attention-footprint similarities yield cross-modal structural priors for non-uniform L1 penalization; (iii) a joint ADMM formulation that estimates shared and class-specific precision components within a single convex objective, avoiding the need to first estimate and then decompose separate class-wise graphs. The learned sparse graph topologies directly support a parameter-free, precision-based classification rule and a lightweight topology-aware segmentation head. Extensive experiments on eight benchmarks demonstrate that CM-GLasso achieves competitive or superior performance compared with strong feature-based and task-specific baselines. Under the matched controlled protocol, it attains the highest average classification accuracy (91.97%) and the highest segmentation mIoU among the controlled baselines on VOC (74.75%) and ADE20K (64.01%), while also yielding explicit sparse conditional-dependence graphs with common-specific decomposition.
Chinese Translation
从多模态视觉-语言特征中估计可解释的条件依赖结构仍然在很大程度上未被探索。我们提出了CM-GLasso(跨模态图形套索),这是一个将视觉-语言表示学习与稀疏高斯图形模型相结合的框架。CM-GLasso引入了三个关键组件:(i)一种文本可视化策略,将类别属性描述呈现为图像,并通过与自然图像相同的SigLIP-2视觉编码器进行处理,从而在共享特征坐标系统中生成原型索引的补丁级注意力足迹;(ii)一种跨注意力蒸馏机制,将高维补丁浓缩为一小组语义图节点,其注意力足迹相似性为非均匀L1惩罚提供了跨模态结构先验;(iii)一种联合ADMM(交替方向乘子法)公式,在单一凸目标内估计共享和类别特定的精度组件,避免了首先估计然后分解单独类别图的需要。学习到的稀疏图拓扑直接支持无参数、基于精度的分类规则和轻量级拓扑感知分割头。在八个基准上的广泛实验表明,CM-GLasso在与强特征基础和任务特定基线的比较中实现了具有竞争力或更优的性能。在匹配控制协议下,它在VOC(91.97%)上获得了最高的平均分类准确率,以及在控制基线中在VOC(74.75%)和ADE20K(64.01%)上获得了最高的分割mIoU,同时还产生了具有共同特异分解的明确稀疏条件依赖图。
cs.CV / 11 / 2608.21445

ViTexSZ: Heterogeneous Vision-Text Knowledge Distillation for EEG Seizure Detection

ViTexSZ:用于脑电图癫痫发作检测的异构视觉-文本知识蒸馏
Liu, Chenxi, Li, Mingzhao, Liu, Yicong, Miao, Hao, Zhang, Hongyuan, Chen, Ziyi, Meng, Gaofeng
Abstract
Automated seizure detection from electroencephalography (EEG) is essential for continuous neurological monitoring, particularly for subclinical epileptic seizures that may exhibit only subtle electrographic changes. Existing time-series methods are often designed for fixed EEG channel configurations, thereby limiting their applicability to heterogeneous EEG recordings with irregular channel layouts. Although visual and language modeling offer promising alternatives, aligning heterogeneous EEG representations with clinical semantics remains challenging. We introduce ViTexSZ, a heterogeneous Vision-Text knowledge distillation framework for EEG seizure detection. ViTexSZ converts EEG recordings into structured waveform images and introduces a query-based multi-channel alignment module that maps source-dependent visual features into a unified token space. A heterogeneous teacher further integrates the aligned EEG representations with clinical prompts through a multimodal large language model, associating high-level clinical semantics with seizure-related evidence. Vision-text knowledge distillation then transfers the teacher representations to a lightweight student during detection. Experiments on four EEG seizure datasets demonstrate the generalizability of ViTexSZ across both subclinical and general seizure detection scenarios, achieving the highest accuracy on all datasets and relative improvements of up to 12.9% over the second-best baselines, showing its effectiveness.
Chinese Translation
自动化的脑电图(EEG)癫痫发作检测对于持续的神经监测至关重要,尤其是对于可能仅表现出微妙电图变化的亚临床癫痫发作。现有的时间序列方法通常是为固定的EEG通道配置设计的,因此限制了它们在具有不规则通道布局的异构EEG记录中的适用性。尽管视觉和语言建模提供了有前景的替代方案,但将异构EEG表示与临床语义对齐仍然具有挑战性。我们提出了ViTexSZ,一种用于EEG癫痫发作检测的异构视觉-文本知识蒸馏框架。ViTexSZ将EEG记录转换为结构化波形图像,并引入基于查询的多通道对齐模块,将源依赖的视觉特征映射到统一的标记空间。一个异构教师进一步通过多模态大型语言模型将对齐的EEG表示与临床提示整合,将高级临床语义与癫痫相关证据关联起来。随后,视觉-文本知识蒸馏在检测过程中将教师表示转移到轻量级学生模型上。在四个EEG癫痫发作数据集上的实验表明,ViTexSZ在亚临床和一般癫痫发作检测场景中的普适性,所有数据集上均实现了最高准确率,并相对于第二好的基线提高了多达12.9%,显示了其有效性。
cs.CV / 12 / 2608.21447

BIMScript: Material-Aware Structured Scene Programs for BIM Ingestion

BIMScript:面向BIM导入的材料感知结构场景程序
Naikade, Prakash Kondibhau, Moeslund, Thomas B., Møgelmose, Andreas
Abstract
Structured-language models such as SceneScript reconstruct a scene as a short program of parametric commands, an inherently editable and semantically explicit representation. We ask three questions that stand between such models and their most compelling application, automated ingestion of existing buildings into BIM tools, studied here on synthetic scans: \emph{what} is the scene made of, \emph{how fast} can it be produced, and \emph{exactly where} is each element. BIMScript answers all three within one grammar. First, we extend the layout language with per-element \emph{material} and \emph{condition} attributes, supervised by a vision-language-model material-passport corpus we build over 100k synthetic scenes (1.9M pseudo-labeled elements), and route image appearance to the material tokens through a lifted-feature point encoder. Second, we show that autoregressive decoding of these programs is dominated not by compute but by kernel-launch and host-synchronization overhead, and remove it with an output-exact CUDA-graph decoder (1.9 vs 6.4\,ms/step, $3.4\times$) plus a grammar-parallel, tolerance-verified draft-and-verify scheme that exploits the deterministic entity schema. Third, we address the model's 5cm token-grid granularity with training-free geometric snapping and a hybrid discrete--continuous decoder head that regresses a sub-bin offset, and measure how much of the residual error each recovers. Because each command maps one-to-one onto a native Revit object, we validate direct ingestion into a BIM authoring tool end to end with a working add-in and its IFC4 export, and the same program's language form is designed to support LLM-driven, sustainability-aware reasoning over the built asset.
Chinese Translation
结构语言模型如SceneScript通过一系列参数化命令的短程序重建场景,这是一种固有可编辑且语义明确的表示方式。我们提出三个问题,这些问题在此类模型与其最具吸引力的应用——将现有建筑自动导入BIM工具之间存在障碍,本文在合成扫描数据上进行了研究: extit{场景由什么构成}, extit{生产速度有多快},以及 extit{每个元素的确切位置}。BIMScript在一个语法中回答了这三个问题。首先,我们通过构建一个包含10万个合成场景(190万个伪标记元素)的视觉-语言模型材料护照语料库,扩展了布局语言,增加了每个元素的 extit{材料}和 extit{条件}属性,并通过提升特征点编码器将图像外观路由到材料标记。其次,我们展示了这些程序的自回归解码主要受制于内核启动和主机同步的开销,而非计算能力,并通过输出精确的CUDA图解码器(每步1.9毫秒对比6.4毫秒,$3.4 imes$)以及利用确定性实体模式的语法并行、容忍验证的草拟与验证方案消除了这一开销。第三,我们通过无训练的几何快照和一个混合离散-连续解码头来解决模型的5厘米标记网格粒度问题,该解码头回归一个子箱偏移,并测量每个恢复的残差误差。由于每个命令一一映射到原生Revit对象,我们通过一个工作插件及其IFC4导出验证了BIM创作工具的直接导入,且同一程序的语言形式旨在支持基于LLM的、关注可持续性的建筑资产推理。
cs.CV / 13 / 2608.21450

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering

超越视觉相似性:面向知识基础视觉问答的实体对齐检索
Xu, Hangrui, Wu, Zhengxian, Yu, Yunyao, Chen, Zhuohong, Cong, Rui, Deng, Xiangwen, Liu, Zhifang, Jiao, Peng, Wang, Haoqian
Abstract
Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involving long-tail entities. However, existing retrieval pipelines predominantly employ CLIP-style dual encoders, which prioritize surface-level visual similarity over entity-level semantic alignment. This paradigm often fails when semantically identical concepts exhibit large visual variations or when distinct entities appear visually similar. To address this, we propose KBMR, the first MLLM-based embedding retriever tailored for KB-VQA. Leveraging the robust autoregressive capabilities of MLLMs, KBMR maps images into a semantic space that better preserves concept identity. To tackle the challenge of noisy supervision in Wikipedia-scale retrieval, we introduce an MLLM-based semantic discriminator that generates continuous entity-consistency weights. These weights guide a novel continuous semantic distillation objective, enabling effective hard negative sampling and soft supervision beyond rigid binary labels. Extensive experiments demonstrate that KBMR significantly outperforms CLIP baselines, yielding up to a 14.7% improvement in retrieval Recall@1 and a 9.4% gain in end-to-end VQA accuracy. Code is available at https://github.com/realHarryX/KBMR.
Chinese Translation
知识基础视觉问答(KB-VQA)依赖于检索外部信息以回答涉及长尾实体的查询。然而,现有的检索流程主要采用CLIP风格的双编码器,这些编码器优先考虑表层视觉相似性,而非实体层面的语义对齐。这种范式在语义上相同的概念表现出较大视觉差异或当不同实体在视觉上相似时,往往会失败。为了解决这一问题,我们提出了KBMR,这是首个针对KB-VQA量身定制的基于多模态大语言模型(MLLM)的嵌入检索器。KBMR利用MLLM强大的自回归能力,将图像映射到一个更好地保留概念身份的语义空间。为了解决维基百科规模检索中的噪声监督问题,我们引入了一种基于MLLM的语义鉴别器,生成连续的实体一致性权重。这些权重指导一种新颖的连续语义蒸馏目标,使得有效的困难负样本采样和超越严格二元标签的软监督成为可能。大量实验表明,KBMR显著优于CLIP基线,在检索Recall@1上提高了多达14.7%,在端到端VQA准确性上提升了9.4%。代码可在 https://github.com/realHarryX/KBMR 获取。
cs.CV / 14 / 2608.21454

Multi-Scale Fruit Capsules: Dilated Convolutions and Dynamic Routing for In-the-Wild Explainable Fruit Recognition

多尺度果实胶囊:扩张卷积与动态路由用于野外可解释的果实识别
Chattoraj, Subhankar, Pratiher, Sawon, Das, Samiran, Konik, Hubert
Abstract
The same fruit appears in a bunch, unpicked, peeled, bagged in plastic, or sliced on a dish, so automated fruit classification in the wild (AFCW) must absorb wide intra- class and narrow inter-class variability in shape, size, colour and texture. Convolutional networks route information through pooling, which discards the pose and location of the region of interest and therefore generalises poorly across these presentations. We propose FruitCapsNet, a capsule network whose Fruit Capsules replace the standard convolutional front end with dilated convolutions: the receptive field grows exponentially at constant parameter cost, so each capsule encodes multi-scale context before dynamic routing resolves part whole spatial agreement. Hyper-parameters, including the dilation factor, are selected by Bayesian optimisation rather than grid search. On three public datasets (SMP, FruitsGB, Fruits-360) and a new 19-class, 10,639-image in-the-wild dataset (PD-19), FruitCapsNet exceeds ten fine-tuned transfer-learning backbones at one-third the depth, with the largest margin (+2.7% over the nearest competitor) on the hardest set. Grad-CAM saliency propagated from the DigitCaps layer shows that the improvement comes from attributing decisions to whole-fruit regions rather than to object edges, giving post-hoc evidence that the gain is not a dataset artefact.
Chinese Translation
同一种果实可能以一束的形式出现,未采摘、去皮、用塑料袋包装或切片放在盘子上,因此,野外自动果实分类(AFCW)必须吸收形状、大小、颜色和纹理上的广泛类内变异和狭窄类间变异。卷积网络通过池化路由信息,这会丢弃感兴趣区域的姿态和位置,因此在这些表现形式上泛化效果较差。我们提出了FruitCapsNet,这是一种胶囊网络,其果实胶囊用扩张卷积替代了标准卷积前端:感受野以恒定参数成本呈指数增长,因此每个胶囊在动态路由解决部分与整体空间一致性之前编码了多尺度上下文。超参数,包括扩张因子,通过贝叶斯优化选择,而不是网格搜索。在三个公共数据集(SMP、FruitsGB、Fruits-360)和一个新的19类、10,639图像的野外数据集(PD-19)上,FruitCapsNet在深度仅为三分之一的情况下超越了十个经过微调的迁移学习骨干网络,在最困难的数据集上具有最大的优势(比最近的竞争对手高出+2.7%)。从DigitCaps层传播的Grad-CAM显著性图显示,改进来自于将决策归因于整个果实区域而非物体边缘,提供了后验证据,表明这一提升不是数据集伪影。
cs.CV / 15 / 2608.21455

Tomatoes, Potatoes, and Onions: Questioning the Need for Faces in Face Presentation Attack Detection

番茄、土豆和洋葱:质疑面部呈现攻击检测中面孔的必要性
Ozgur, Guray, Boutros, Fadi, Damer, Naser
Abstract
Face presentation attack detection (PAD) is traditionally formulated as a face-specific problem, although many of the visual artifacts introduced by print, replay, and recapture processes are not inherently tied to facial appearance. In this work, we investigate whether transferable PAD representations can be learned without using faces during downstream PAD training. To this end, we introduce TPO, a controlled face-free presentation attack dataset consisting of bona fide, print, and replay recordings of, almost randomly chosen, tomatoes, potatoes, and onions acquired under protocols that closely mirror conventional face PAD datasets. Using a foundation-model-based PAD architecture, we demonstrate that a detector trained on TPO achieves an average AUC of 92.70% across four standard cross-dataset face PAD benchmarks, outperforming training on synthetic faces and remaining competitive with models trained on real face datasets. Conversely, models trained on face PAD datasets transfer consistently above chance to TPO, suggesting that the learned representations capture characteristics of the presentation process rather than object semantics. Furthermore, incorporating TPO into conventional face PAD training consistently improves cross-dataset performance under fixed optimization budgets, indicating that face-free data provides complementary information rather than simply additional training samples. Finally, representation and frequency analyses provide further evidence that transferable PAD representations cannot be explained by a single spectral artifact but instead encode richer presentation cues shared across object categories. Together, these results provide empirical evidence that transferable presentation attack representations can be learned independently of facial content, opening new opportunities for privacy-preserving and identity-independent PAD development.
Chinese Translation
面部呈现攻击检测(PAD)传统上被视为一个面部特定的问题,尽管许多由打印、重播和重新捕捉过程引入的视觉伪影并不固有地与面部外观相关。在本研究中,我们探讨是否可以在下游PAD训练中不使用面孔的情况下学习可转移的PAD表示。为此,我们引入了TPO,一个受控的无面孔呈现攻击数据集,包含真实、打印和重播记录的几乎随机选择的番茄、土豆和洋葱,这些记录是在与传统面部PAD数据集密切相似的协议下获得的。使用基于基础模型的PAD架构,我们证明了在TPO上训练的检测器在四个标准跨数据集面部PAD基准测试中达到了92.70%的平均AUC,优于在合成面孔上训练,并与在真实面孔数据集上训练的模型保持竞争力。相反,在面部PAD数据集上训练的模型在TPO上的转移表现始终高于偶然水平,表明学习到的表示捕捉的是呈现过程的特征,而不是物体语义。此外,将TPO纳入传统面部PAD训练在固定优化预算下始终改善跨数据集性能,表明无面孔数据提供了互补信息,而不仅仅是额外的训练样本。最后,表示和频率分析提供了进一步证据,表明可转移的PAD表示无法通过单一的光谱伪影来解释,而是编码了跨物体类别共享的更丰富的呈现线索。这些结果共同提供了实证证据,表明可转移的呈现攻击表示可以独立于面部内容进行学习,为隐私保护和身份独立的PAD开发开辟了新的机会。
cs.CV / 16 / 2608.21457

CLSC DETR: Reliable Candidate Ranking via Cross Layer Geometric Support for UAV Small Object Detection

CLSC DETR:通过跨层几何支持实现无人机小物体检测的可靠候选排名
Lin, Junyan
Abstract
Unmanned aerial vehicle (UAV) object detection is critical for applications such as target search, where accurate detection of small objects in complex aerial scenes remains challenging. The limited spatial extent, dense distribution, and frequent occlusion of small objects make reliable candidate ranking particularly difficult. Existing Detection Transformer (DETR) based methods improve ranking by estimating localization quality from individual queries and incorporating it into classification scores. However, a single query often lacks sufficient geometric evidence for small objects with weak boundary cues, resulting in unreliable quality estimation and unstable ranking. To address this limitation, we propose Cross Layer Local Support and Consistency Calibration for DETR, termed CLSC DETR. Specifically, the Cross Layer Local Support module establishes correspondences between final layer queries and intermediate layer candidates to aggregate complementary geometric evidence for more reliable localization quality estimation, while the Classification and Localization Consistency Calibration module adaptively adjusts classification scores according to localization quality and classification reliability to improve candidate ranking. Experiments show that CLSC DETR improves AP and AP$_{75}$ over the baseline by 1.5\% and 2.0\% on VisDrone, respectively, while achieving consistent improvements on UAVDT.
Chinese Translation
无人机(UAV)物体检测对于目标搜索等应用至关重要,但在复杂的空中场景中,准确检测小物体仍然具有挑战性。小物体的有限空间范围、密集分布和频繁遮挡使得可靠的候选排名特别困难。现有基于检测变换器(DETR)的方法通过从单个查询中估计定位质量并将其纳入分类分数来改善排名。然而,单个查询通常缺乏足够的几何证据来处理边界线索较弱的小物体,导致质量估计不可靠和排名不稳定。为了解决这一限制,我们提出了用于DETR的跨层局部支持与一致性校准,称为CLSC DETR。具体而言,跨层局部支持模块在最终层查询和中间层候选之间建立对应关系,以聚合互补的几何证据,从而实现更可靠的定位质量估计,而分类与定位一致性校准模块则根据定位质量和分类可靠性自适应调整分类分数,以改善候选排名。实验表明,CLSC DETR在VisDrone上相较于基线提高了1.5 ext{%}和2.0 ext{%}的AP和AP$_{75}$,同时在UAVDT上也实现了一致的改进。
cs.CV / 17 / 2608.21460

FigmaTrace: Capturing Creative Nuances in Human Figma Design Workflows

FigmaTrace:捕捉人类 Figma 设计工作流程中的创意细微差别
Deshpande, Darshan, Fujinuma, Yoshinari, Markiewicz, Martyna, Bansal, Devanshu, Jain, Shivani, Saban, Nicholas, Maheshwari, Chirag, Kannappan, Anand
Abstract
Vision Language Models have recently shown improvements in several objective and verifiable domains such as object detection but continue to underperform on subjective and creative design tasks. A major contributor to this performance gap is the lack of high quality human workflow data that captures a diverse set of preferences and decisions that make human experts good at design tasks. In this work, we first define a unique, expert curated taxonomy of design skills and best practices which we further expand into a set of 126 open ended, subjective, long horizon tasks. Built on top of this and expert solutions, our dataset FigmaTrace contains over 200 hours of human captured video data converted into 3469 design trajectories using a novel design phase-based method. We use our dataset to train four models and show that training on FigmaTrace leads to a performance improvement comparable to frontier closed models such as \textsc{Claude-Opus-5} and \textsc{GPT-5.6-Sol} on four out of distribution agentic GUI environments. We further perform a useful ablation to attribute these performance improvements to a design phase-based video to trajectory conversion which outperforms prior length-based conversion approaches. Finally, we perform a qualitative analysis on the best performing \textsc{Qwen3.8-27B} outputs to better correlate performance improvements to FigmaTrace's trends. We open source our dataset and the best model for the community.
Chinese Translation
视觉语言模型最近在多个客观和可验证的领域(如物体检测)中表现出改善,但在主观和创意设计任务上仍然表现不佳。造成这一性能差距的一个主要原因是缺乏高质量的人类工作流程数据,这些数据能够捕捉到多样化的偏好和决策,使人类专家在设计任务中表现出色。在本研究中,我们首先定义了一种独特的、由专家策划的设计技能和最佳实践的分类法,并进一步扩展为一组126个开放式、主观、长期的任务。基于此及专家解决方案,我们的数据集 FigmaTrace 包含超过200小时的人类捕获视频数据,通过一种新颖的基于设计阶段的方法转换为3469条设计轨迹。我们利用该数据集训练了四个模型,并展示了在 FigmaTrace 上训练能带来与前沿封闭模型(如 extsc{Claude-Opus-5} 和 extsc{GPT-5.6-Sol})相当的性能提升,适用于四个超出分布的自主图形用户界面环境。我们进一步进行了有益的消融实验,将这些性能提升归因于基于设计阶段的视频到轨迹转换,该方法优于先前基于长度的转换方法。最后,我们对表现最佳的 extsc{Qwen3.8-27B} 输出进行了定性分析,以更好地关联性能提升与 FigmaTrace 的趋势。我们将数据集和最佳模型开源,供社区使用。
cs.CV / 18 / 2608.21464

Complexity Induction: Compositional Generalization via Structured Label Distortion

复杂性引导:通过结构化标签扭曲实现组合泛化
Abramov, Aleksandr
Abstract
We demonstrate that structured distortion of training data - which we term complexity induction - can induce compositional generalization in a standard CNN classifier without architectural modification. Using synthetic images of colored geometric shapes, we encode classes as flat string labels (e.g., "red-circle") with no explicit attribute decomposition, and exclude certain color-shape combinations from training entirely. We apply two distortion methods derived from Jaccard string similarity between class names: mixed labels (soft target distributions encoding inter-class overlap) and expanded dataset (false training samples with structurally motivated incorrect labels). Both methods induce the ability to predict unseen class combinations, and act at different levels: mixed labels activate the classifier for unseen combinations by exploiting the CNN's natural embedding structure, while expanded training improves the embedding factorization itself. A control with random (unstructured) false labels confirms that the effect depends on the structure of the distortion, not on noise per se. These results suggest that structured complication of training signals can influence both the internal organization of learned representations and their compositional interpretation - a principle that may underlie the role of natural language in cognitive development.
Chinese Translation
我们展示了训练数据的结构化扭曲——我们称之为复杂性引导——可以在不修改架构的情况下诱导标准卷积神经网络(CNN)分类器实现组合泛化。使用彩色几何形状的合成图像,我们将类别编码为平面字符串标签(例如,“红色-圆形”),没有明确的属性分解,并完全排除某些颜色-形状组合的训练。我们应用了两种基于类别名称之间的 Jaccard 字符串相似度衍生的扭曲方法:混合标签(编码类别间重叠的软目标分布)和扩展数据集(具有结构性动机的错误标签的虚假训练样本)。这两种方法都能够诱导预测未见类别组合的能力,并在不同层面上发挥作用:混合标签通过利用 CNN 的自然嵌入结构激活分类器以处理未见组合,而扩展训练则改善了嵌入因式分解本身。与随机(无结构)错误标签的对照实验确认了这一效果依赖于扭曲的结构,而非噪声本身。这些结果表明,训练信号的结构化复杂性可以影响学习表征的内部组织及其组合解释——这一原理可能是自然语言在认知发展中所起作用的基础。
cs.CV / 19 / 2608.21468

3D Point Cloud from Close-Range Photogrammetry for Defect Characterisation of Rubberised Concrete

基于近距离摄影测量的三维点云用于橡胶混凝土缺陷特征化
Liu, Jiacheng, Alnahhal, Mohammed, Hajimohammadi, Ailar, Barsanti, Sara Gonizzi, Wang, Jinling, Kalantari, Mohsen
Abstract
While three-dimensional (3D) point clouds are widely used in civil engineering, mainstream LiDAR systems such as Terrestrial Laser Scanning (TLS) are physically constrained to laboratory environments. Since their laser spot size typically exceeds the width of microcracks, the beam physically bridges over voids, rendering TLS unsuitable for fine-scale defect analysis. Alternatively, close-range photogrammetry utilising Structure-from-Motion (SfM) and Multi-View Stereo (MVS) algorithms offers a solution for testing highly tortuous materials, and its utility at fine-scale remains underexplored. This study adapts photogrammetric workflows specifically for rubberised concrete (RuC), a sustainable composite exhibiting high ductility and complex fracture morphologies. High-resolution image sets were captured using a Canon DSLR and an iPhone 16 to generate dense 3D models. Comparisons revealed that the DSLR-based reconstruction achieved sub-millimetre resolution, demonstrating superior performance for fine-scale surface monitoring. An RGB-guided crack extraction method was developed to enhance the identification of surface defects and isolate potential crack areas from the background. The extracted crack regions were visually distinguishable and provided a well-structured geometrical representation of defect morphology. Furthermore, a Pre and Post-Test deformation analysis was conducted to quantify surface displacement across testing stages. The results confirm that this close-range photogrammetry workflow is a flexible, high-resolution alternative to LiDAR for surface inspection and deformation monitoring of specimens in laboratory settings. Ultimately, this approach establishes a robust geometric baseline for future automated 3D feature characterisation and material performance evaluation.
Chinese Translation
三维(3D)点云在土木工程中被广泛应用,但主流的激光雷达(LiDAR)系统,如地面激光扫描(TLS),在物理上受到实验室环境的限制。由于其激光光斑尺寸通常超过微裂缝的宽度,光束会物理性地跨越空隙,使得TLS不适合进行细尺度缺陷分析。相对而言,利用运动结构(Structure-from-Motion, SfM)和多视角立体(Multi-View Stereo, MVS)算法的近距离摄影测量为测试高度曲折的材料提供了解决方案,而其在细尺度上的应用仍未得到充分探索。本研究专门为橡胶混凝土(Rubberised Concrete, RuC)适配了摄影测量工作流程,这是一种具有高延展性和复杂断裂形态的可持续复合材料。使用佳能单反相机和iPhone 16捕获高分辨率图像集,以生成密集的3D模型。比较结果显示,基于单反相机的重建达到了亚毫米级别的分辨率,展现了在细尺度表面监测中的优越性能。开发了一种RGB引导的裂缝提取方法,以增强表面缺陷的识别,并将潜在裂缝区域从背景中分离。提取的裂缝区域在视觉上可区分,并提供了缺陷形态的良好结构化几何表示。此外,进行了测试前后变形分析,以量化测试阶段的表面位移。结果确认该近距离摄影测量工作流程是LiDAR在实验室环境中进行表面检查和变形监测的灵活、高分辨率替代方案。最终,该方法为未来自动化的3D特征特征化和材料性能评估建立了稳健的几何基线。
cs.CV / 20 / 2608.21486

EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment

EXPL-FR:通过视觉-语言对齐解释人脸识别模型
Ozgur, Guray, Tamyapar, Mustafa Efe, Damer, Naser, Boutros, Fadi
Abstract
Deep face recognition (FR) models reach near-saturated accuracy but remain opaque: a practitioner cannot ask which semantic attributes a similarity score relied upon. EXPL-FR answers this inside the FR model's own embedding space. A lightweight adapter aligns a vision-language model's (VLM) image encoder with the frozen FR space, trained on face images alone and never on text. Because the VLM's encoders share one space, the same adapter applies to the text encoder, turning 978 attribute prompts in 22 categories, also extendable, into FR-space anchors at no extra cost. We do not assume this transfer works: a face-verification protocol measures it, and an ablation changing only the adapter isolates its contribution. Not every concept survives, because an FR model earns its invariances by discarding the factors it must verify identities across. A label-free detectability measure compares each concept's separability in FR space against the VLM space, and the 100 most detectable form the model's readable semantic signature, which separates identities better than the full vocabulary. We cover four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations. We benchmark attribute-level auditing under three supervision settings, human labels (current practice), VLM pseudo-labels, and our fully prompt-driven audit, against real verification behavior. With no labels, the prompt-driven audit ranks four FR models by their measured per-ethnicity RFW errors and ranks controlled attribute changes by their true verification cost.
Chinese Translation
深度人脸识别(FR)模型已达到近乎饱和的准确率,但仍然不透明:从业者无法询问相似性得分依赖于哪些语义属性。EXPL-FR 在 FR 模型自身的嵌入空间中回答了这个问题。一个轻量级适配器将视觉-语言模型(VLM)的图像编码器与冻结的 FR 空间对齐,该空间仅在面部图像上训练,而从未在文本上训练。由于 VLM 的编码器共享一个空间,同样的适配器适用于文本编码器,将 978 个属性提示(涵盖 22 个类别,且可扩展)转化为 FR 空间的锚点,且无需额外成本。我们并不假设这种转移有效:一个人脸验证协议对此进行了测量,而仅更改适配器的消融实验则隔离了其贡献。并非每个概念都能存活下来,因为 FR 模型通过丢弃必须验证身份的因素来获得其不变性。一个无标签的可检测性度量将每个概念在 FR 空间的可分性与 VLM 空间进行比较,前 100 个最可检测的概念形成模型的可读语义特征,这些特征在身份分离上优于完整词汇。我们涵盖了四个 FR 主干网络和两个 VLM 编码器,EXPL-FR 不需要访问架构,并支持身份级、每图像和差异性解释。我们在三种监督设置下对属性级审计进行了基准测试,包括人类标签(当前实践)、VLM 伪标签和我们完全基于提示的审计,针对真实的验证行为。没有标签的情况下,基于提示的审计根据测量的每个种族 RFW 错误对四个 FR 模型进行排名,并根据其真实验证成本对受控属性变化进行排名。
cs.CV / 21 / 2608.21487

TASSO: TAsk-Specific Subspace Optimization for Continual Learning of Vision-Language Models

TASSO:针对视觉-语言模型的任务特定子空间优化的持续学习
Sun, Chang, Barbato, Francesco, Caligiuri, Matteo, Zanuttigh, Pietro
Abstract
Vision-Language Models (VLMs) exhibit strong zero-shot capabilities, making them an attractive solution for continual learning across diverse tasks. However, during continual adaptation, both catastrophic forgetting and zero-shot degradation occur, severely degrading performance. In this paper, we introduce TASSO, a new paradigm that efficiently preserves the latent space geometry while ensuring network plasticity. We achieve this with two complementary techniques: subspace learning and geometry-aware knowledge distillation. Specifically, we first learn a sequence of task-specific low-rank projectors, which we use to project the latent representations before optimizing cross-entropy. Secondly, we employ a geodesic-distance-based loss that distills knowledge from the previous-task model while effectively preserving the latent space geometry. These design choices not only avoid unnecessary parameter updates along the full embedding dimensions but also improve learning by focusing on task-specific manifolds. Moreover, the geometry-aware distillation provides strong regularization and significantly reduces both catastrophic forgetting and zero-shot degradation throughout the continual learning sequence. Experimental results with the CLIP vision language model in the multi-domain task incremental and class incremental learning benchmarks demonstrate clear improvements over state-of-the-art methods in mitigating forgetting and preserving zero-shot capabilities.
Chinese Translation
视觉-语言模型(VLMs)展现出强大的零-shot能力,使其成为在多样任务中进行持续学习的有吸引力的解决方案。然而,在持续适应过程中,灾难性遗忘和零-shot性能下降同时发生,严重影响性能。本文提出了TASSO,一种新的范式,能够有效地保持潜在空间几何形状,同时确保网络的可塑性。我们通过两种互补技术实现这一目标:子空间学习和几何感知知识蒸馏。具体而言,我们首先学习一系列任务特定的低秩投影器,用于在优化交叉熵之前对潜在表示进行投影。其次,我们采用基于测地距离的损失,从前一个任务模型中蒸馏知识,同时有效地保持潜在空间的几何形状。这些设计选择不仅避免了在完整嵌入维度上不必要的参数更新,还通过关注任务特定的流形来改善学习。此外,几何感知蒸馏提供了强有力的正则化,并显著减少了整个持续学习序列中的灾难性遗忘和零-shot性能下降。使用CLIP视觉语言模型在多领域任务增量和类别增量学习基准上的实验结果表明,在减轻遗忘和保持零-shot能力方面,相较于最先进的方法有明显的改善。
cs.CV / 22 / 2608.21529

DamageScope: Vision-Language Retrieval at Scale for Disaster Damage Assessment from Satellite Imagery

DamageScope:基于视觉-语言检索的灾害损害评估卫星图像大规模处理
Rajendran, Ravi K., Debnath, Biplob, Sankaradas, Murugan, Chakradhar, Srimat T.
Abstract
Timely and accurate assessment of property damage is critical following natural disasters. Traditional on-site inspections are labor-intensive, costly, and often pose safety risks. Advances in satellite imagery and vision-language models (VLMs) enable scalable remote damage assessment; however, integrating VLMs into large-scale Earth observation pipelines presents challenges in computational efficiency, data organization, and information retrieval. To address these challenges, we present DamageScope, a retrieval-augmented framework that combines satellite imagery with Vision-Language Models (VLMs) and Large Language Models (LLMs) to automate property damage analysis. Built on a Retrieval-Augmented Generation (RAG) framework, DamageScope extracts structured visual representations from satellite imagery to support interactive natural language queries for damage assessment. To address scalability, we introduce a novel multi-vector embedding-based clustering algorithm that outperforms traditional single-vector embedding approaches while reducing indexing time by up to 14x. Furthermore, a dual-store data architecture minimizes LLM API calls, reducing both operational cost and response latency by up to approximately 3x. By effectively balancing scalability and operational efficiency, DamageScope provides a robust and practical solution for real-world damage assessment tasks.
Chinese Translation
在自然灾害发生后,及时和准确的财产损害评估至关重要。传统的现场检查劳动密集、成本高昂,并且往往存在安全风险。卫星图像和视觉-语言模型(VLMs)的进步使得可扩展的远程损害评估成为可能;然而,将VLMs整合到大规模地球观测管道中面临计算效率、数据组织和信息检索等挑战。为了解决这些挑战,我们提出了DamageScope,这是一种检索增强框架,将卫星图像与视觉-语言模型(VLMs)和大型语言模型(LLMs)结合,以实现财产损害分析的自动化。DamageScope基于检索增强生成(RAG)框架,从卫星图像中提取结构化视觉表示,以支持损害评估的交互式自然语言查询。为了解决可扩展性问题,我们引入了一种新颖的多向量嵌入聚类算法,该算法在性能上优于传统的单向量嵌入方法,同时将索引时间缩短了多达14倍。此外,双存储数据架构最小化了LLM API调用,将运营成本和响应延迟减少了约3倍。通过有效平衡可扩展性和运营效率,DamageScope为现实世界的损害评估任务提供了一个强大而实用的解决方案。
cs.CV / 23 / 2608.21543

presto: Efficient, Training-free, and Open-world Object Placement via Imaginary Search

presto:高效、无训练且开放世界的物体放置方法通过想象搜索
Ding, Weixuan, Liu, Shang, Pei, Hanyu, Liu, Zeyan
Abstract
Object placement is critical in image composition, requiring spatially and semantically coherent positioning of objects within diverse scenes. Existing approaches typically rely on hand-crafted rules or supervised learning on limited datasets, which restricts their generalization and interpretability, especially in open-world scenarios involving novel objects and scenes. In this work, we reformulate open-world object placement as a heuristic search task guided by reasoning from a Multimodal Large Language Model (MLLM). We introduce \textsf{presto}, a zero-shot, training-free framework that operates within an imaginary action space to iteratively refine object position and scale. Our coarse-to-fine search strategy ensures fast convergence, and we evaluate two decision-making variants: Metric-guided Selection and MLLM-as-a-judge. Experiments across multiple benchmarks show that \textsf{presto}~achieves state-of-the-art performance, particularly in previously unseen, open-world settings. Human studies further reveal that the MLLM-as-a-judge variant produces more perceptually coherent placements than metric-driven approaches, highlighting a gap between standard evaluation metrics and human visual judgment.
Chinese Translation
物体放置在图像构图中至关重要,需要在多样场景中对物体进行空间和语义上的一致定位。现有方法通常依赖于手工设计的规则或在有限数据集上的监督学习,这限制了它们的泛化能力和可解释性,特别是在涉及新物体和场景的开放世界场景中。在本研究中,我们将开放世界物体放置重新表述为一个由多模态大型语言模型(Multimodal Large Language Model, MLLM)引导的启发式搜索任务。我们提出了 extsf{presto},一个零-shot、无训练的框架,操作在一个想象的动作空间中,迭代地优化物体的位置和尺度。我们的粗到细搜索策略确保了快速收敛,并评估了两种决策变体:度量引导选择和MLLM作为评判者。跨多个基准的实验表明, extsf{presto}在以前未见的开放世界设置中实现了最先进的性能。人类研究进一步表明,MLLM作为评判者的变体产生的放置在感知上更为一致,相较于基于度量的方法,突显了标准评估指标与人类视觉判断之间的差距。
cs.CV / 24 / 2608.21571

Extending the Horizon of Early Diagnosis: Lung Cancer Prediction with Vision Transformers

扩展早期诊断的视野:基于视觉变换器的肺癌预测
Kotevska, Olivera, Goethert, Ian, McGee, Michael, Mahbub, Maria, Wilkinson, Sean R., Yip, Rowena, Selvan, Myvizhi Esai, Gumus, Zeynep H., Henschke, Claudia, Klein, Robert J., Morales, Providencia, Aguayo, Samuel M, Danciu, Ioana, Chandrashekar, Mayanka
Abstract
Lung cancer remains a leading cause of cancer-related mortality worldwide, and early diagnosis is critical for improving survival. However, early-stage malignancies can be subtle on chest X-rays, creating challenges for radiologists. This study evaluates Vision Transformers (ViTs) for predicting lung cancer one to two years before clinical diagnosis. We analyzed 259,361 chest X-rays from 91,020 imaging studies at the Jamaica Plains VA Hospital in Boston, MA. The dataset showed extreme class imbalance, approximately 1:150 cancer to non-cancer, which was addressed using hybrid under- and over-sampling and class-weighted loss optimization. Three ViT configurations were evaluated: a model trained from scratch, an ImageNet-pretrained model, and a Corona-pretrained model fine-tuned on the lung cancer dataset. Transfer learning improved performance, with pretrained models exceeding the scratch baseline by 6-10 percentage points in AUC and about 10-12 percent in balanced accuracy. ImageNet-pretrained models showed the most stable overall performance, while Corona-pretrained models achieved higher sensitivity in some settings but greater variability. Moderate resampling ratios, including 1:1 undersampling and 1.5:2 oversampling, provided favorable trade-offs between sensitivity, precision, and computational efficiency, reducing runtime by up to 70 percent without major performance loss. These findings demonstrate the potential of ViTs for early lung cancer risk prediction from routine chest X-rays. Although performance remains below clinical deployment thresholds, the results support further development of ViT-based triage systems to flag high-risk patients for earlier evaluation.
Chinese Translation
肺癌仍然是全球癌症相关死亡的主要原因,早期诊断对提高生存率至关重要。然而,早期恶性肿瘤在胸部X光片上可能表现得非常微妙,这给放射科医生带来了挑战。本研究评估了视觉变换器(Vision Transformers, ViTs)在临床诊断前一至两年预测肺癌的效果。我们分析了来自马萨诸塞州波士顿的牙买加平原退伍军人医院的259,361张胸部X光片,涉及91,020个影像研究。数据集显示出极端的类别不平衡,癌症与非癌症的比例约为1:150,为此我们采用了混合的欠采样和过采样方法以及类别加权损失优化。我们评估了三种ViT配置:从头开始训练的模型、经过ImageNet预训练的模型,以及在肺癌数据集上微调的Corona预训练模型。迁移学习提高了性能,预训练模型在AUC上比从头开始训练的基线高出6-10个百分点,在平衡准确率上高出约10-12个百分点。经过ImageNet预训练的模型表现出最稳定的整体性能,而Corona预训练模型在某些设置下实现了更高的灵敏度,但波动性更大。适度的重采样比例,包括1:1的欠采样和1.5:2的过采样,在灵敏度、精确度和计算效率之间提供了良好的权衡,将运行时间减少了多达70%,而性能损失不大。这些发现展示了ViTs在从常规胸部X光片中进行早期肺癌风险预测的潜力。尽管性能仍低于临床部署的阈值,但结果支持进一步开发基于ViT的分诊系统,以标记高风险患者以便进行更早的评估。
cs.CV / 25 / 2608.21595

Perturb the Thought, Not the Pixels: Latent-Space Rollout Diversification for Reinforcement Learning of Vision-Language Models

扰动思维,而非像素:用于视觉-语言模型强化学习的潜在空间展开多样化
Jerge, Michael, Pelczar, Joseph, Downes, Justin
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning ability of vision-language models (VLMs), and diversifying the rollouts within each optimization group amplifies its gains. Existing approaches diversify through decoding temperature or pixel-space image distortion; we ask whether the perturbation belongs in the model's latent space instead. We introduce Noise-Contrastive GRPO (NC-GRPO), which injects scale-calibrated Gaussian noise into the last hidden layer of the prompt-encoding pass for half of each rollout group, branching those rollouts from a displaced departure state. Branches that reach the answer despite the displacement are reinforced over those derailed by it, converting sensitivity at the branch point into policy-gradient signal; the objective, reward, and inference protocol are untouched. On Qwen2.5-VL-7B trained on Geometry3K, NC-GRPO significantly improves out-of-domain mathematical reasoning over vanilla GRPO across five held-out benchmarks (pooled McNemar $p \le 0.001$) while also improving in-domain accuracy and hallucination robustness -- the latter an axis on which image-space noise regresses even while posting a larger OOD average on perception-heavy benchmarks. Mechanism ablations indicate that independent stochastic diversity, not noise budget or direction, is the active ingredient, and a noise-scale study exposes a dial between reasoning specialization and general capability. NC-GRPO is designed to be modality-agnostic and integrates into a standard RLVR pipeline as a ~50-line change to the inference engine.
Chinese Translation
具有可验证奖励的强化学习(RLVR)提高了视觉-语言模型(VLMs)的推理能力,而在每个优化组内多样化展开则进一步放大了其收益。现有方法通过解码温度或像素空间图像失真来实现多样化;我们提出是否可以将扰动引入模型的潜在空间。我们引入了噪声对比GRPO(NC-GRPO),该方法在每个展开组的一半中向提示编码传递的最后隐藏层注入经过尺度校准的高斯噪声,从而使这些展开从一个偏移的起始状态分支。尽管存在偏移,仍然能够到达答案的分支相较于被偏移扰乱的分支获得了更多的强化,将分支点的敏感性转化为策略梯度信号;目标、奖励和推理协议保持不变。在基于Geometry3K训练的Qwen2.5-VL-7B上,NC-GRPO在五个保留基准测试中显著提高了超出领域的数学推理能力(合并McNemar $p le 0.001$),同时也提高了领域内的准确性和幻觉鲁棒性——后者是图像空间噪声在感知重的基准测试中表现出更大OOD平均值时却回归的一个维度。机制消融实验表明,独立的随机多样性,而非噪声预算或方向,是活跃成分,而噪声规模研究揭示了推理专业化与通用能力之间的调节。NC-GRPO旨在实现模态无关,并作为对推理引擎约50行代码的更改集成到标准RLVR管道中。
cs.CV / 26 / 2608.21636

Semantic Slots for Video Object-Centric Learning

面向视频对象中心学习的语义槽
Sabri, Khalil, Bilodeau, Guillaume-Alexandre, Saunier, Nicolas, Bouachir, Wassim
Abstract
Video Object-Centric Learning (OCL) has traditionally focused on refining the encoder architecture to ensure temporal consistency. In this paper, we argue that the primary bottleneck lies in the decoder. We show that traditional decoders force slots to be spatially anchored, hindering their ability to adapt to motion. We propose SemanticSlots, which uses a Transformer-based decoder that leverages image context, relieving slots from encoding boundary precision and spatial location. This allows slots to function as semantic queries that are inherently object position invariant, retrieving matching features rather than memorizing coordinates. More importantly, this property allows slots computed from a single frame to decompose subsequent video frames, eliminating the need for complex temporal predictors or auxiliary temporal losses. Results on YouTube-VIS show that SemanticSlots improves upon VideoSAUR by 31 points in mBO and outperforms current state-of-the-art methods by 21 points, achieving 86.6% ARI and 62.8% mBO.
Chinese Translation
视频对象中心学习(OCL)传统上专注于优化编码器架构以确保时间一致性。在本文中,我们认为主要瓶颈在于解码器。我们展示了传统解码器迫使槽在空间上固定,阻碍了它们适应运动的能力。我们提出了SemanticSlots,它使用基于Transformer的解码器,利用图像上下文,减轻了槽在编码边界精度和空间位置上的负担。这使得槽能够作为固有对象位置不变的语义查询,检索匹配特征而不是记忆坐标。更重要的是,这一特性使得从单帧计算的槽能够分解后续视频帧,消除了对复杂时间预测器或辅助时间损失的需求。在YouTube-VIS上的结果表明,SemanticSlots在mBO上比VideoSAUR提高了31个百分点,并且比当前最先进的方法提高了21个百分点,达到了86.6%的ARI和62.8%的mBO。
cs.CV / 27 / 2608.21659

SketchFlow: Zero-Shot Vector Sketch Generation via GMM Prior Flow in CLIP Latent Space

SketchFlow:通过 GMM 先验流在 CLIP 潜在空间中实现零-shot 向量草图生成
Zhou, Jin, Yang, Hongliang, Xu, Pengfei, Huang, Hui
Abstract
Vector sketches remain one of the most concise and immediate mediums for abstract human expression. However, generating high-quality vector strokes that exhibit human-like drawing styles remains an open challenge due to the severe scarcity of fine-grained, high-quality text-to-sketch paired data. Existing text-conditioned generation methods often rely on unstable, time-consuming optimization or struggle to generalize to unseen categories in a zero-shot manner. To address these limitations, we present SketchFlow, a novel generative framework rooted in Optimal Transport (OT) theory and flow matching. By leveraging pre-trained CLIP models to bypass labor-intensive image-level text annotations, we formulate cross-modal alignment as a continuous mapping problem directly within the CLIP latent space. To bridge the inevitable modality gap between discrete text concepts and continuous sketch features, we first inject noise into discrete category embeddings to construct a continuous Gaussian Mixture Model (GMM) prior. We then utilize an Optimal Transport Conditional Flow Matching (OT-CFM) model to learn a deterministic vector field mapping from this continuous GMM prior to the target sketch feature distribution. Finally, a Hybrid Diffusion Decoder, fusing 1D U-Net and Transformer architectures, is designed to decode these features into fast and high-fidelity stroke trajectories. Extensive experiments demonstrate that SketchFlow substantially outperforms existing baselines in visual quality and adherence to natural human drawing styles. Furthermore, our geometry-preserving framework demonstrates promising local zero-shot synthesis for prompts beyond the QuickDraw training vocabulary, including unseen concept labels and semantic modifiers, while enabling smooth, continuous semantic interpolation between distinct concepts. Source code is available at: https://github.com/QiuHong-1202/SketchFlow.
Chinese Translation
向量草图仍然是抽象人类表达中最简洁和直接的媒介之一。然而,由于高质量的文本到草图配对数据的严重匮乏,生成展现人类绘画风格的高质量向量笔画仍然是一个开放的挑战。现有的文本条件生成方法往往依赖不稳定且耗时的优化,或者在零-shot 方式下难以推广到未见类别。为了解决这些局限性,我们提出了 SketchFlow,这是一种基于最优传输(Optimal Transport, OT)理论和流匹配的新型生成框架。通过利用预训练的 CLIP 模型来绕过劳动密集型的图像级文本注释,我们将跨模态对齐公式化为 CLIP 潜在空间内的连续映射问题。为了弥合离散文本概念与连续草图特征之间不可避免的模态差距,我们首先向离散类别嵌入注入噪声,以构建连续的高斯混合模型(Gaussian Mixture Model, GMM)先验。然后,我们利用最优传输条件流匹配(Optimal Transport Conditional Flow Matching, OT-CFM)模型,从这个连续的 GMM 先验学习一个确定性的向量场映射到目标草图特征分布。最后,设计了一种混合扩散解码器,融合了一维 U-Net 和 Transformer 架构,将这些特征解码为快速且高保真的笔画轨迹。大量实验表明,SketchFlow 在视觉质量和遵循自然人类绘画风格方面显著优于现有基线。此外,我们的几何保持框架在 QuickDraw 训练词汇之外的提示上展示了有希望的局部零-shot 合成,包括未见的概念标签和语义修饰符,同时实现了不同概念之间的平滑、连续的语义插值。源代码可在以下网址获取:https://github.com/QiuHong-1202/SketchFlow。
cs.CV / 28 / 2608.21697

Emotion Intensity Matters: Generating Realistic Expressions in Virtual Humans with CVAEs

情感强度的重要性:利用条件变分自编码器生成虚拟人类的真实表情
Peres, Vitor Miguel Xavier, Volpato, Lara, Scnheider, Gabriel Ferri, Musse, Soraia Raupp
Abstract
Generating expressive facial behavior in virtual humans (VHs) remains a central challenge in affective computing and character animation. This paper presents a novel approach based on Conditional Variational Autoencoders (CVAEs), trained on real human facial expression data, to synthesize controllable emotional expressions at varying intensities. Using a dataset comprising six basic emotions represented at two intensity levels (low and high), we train a CVAE model to generate synthetic facial expression data while preserving semantic consistency with real human expressions. Despite the limited amount of training data (only 7,680 facial expression samples), the proposed approach learns meaningful latent representations and generates coherent emotional variations. Our method enables control over emotional intensity, making it suitable for animating virtual characters without requiring actor performances or manual artistic intervention. Our research aimed to evaluate whether the method (CVAE) preserves the characteristics associated with the different intensity levels present in the dataset. Results show that the proposed model preserves key expressive characteristics across intensity levels while supporting generalization across emotional intensity levels, contributing to the creation of emotionally expressive virtual characters from relatively small datasets.
Chinese Translation
在情感计算和角色动画中,生成虚拟人类(VHs)的表现性面部行为仍然是一个核心挑战。本文提出了一种基于条件变分自编码器(CVAEs)的新方法,该方法在真实人类面部表情数据上进行训练,以合成可控的情感表达,且具有不同的强度。我们使用一个包含六种基本情感的数据库,这些情感在两个强度水平(低和高)上进行表示,训练一个CVAEs模型以生成合成的面部表情数据,同时保持与真实人类表情的语义一致性。尽管训练数据量有限(仅有7,680个面部表情样本),但所提出的方法能够学习有意义的潜在表示,并生成连贯的情感变化。我们的方法使得对情感强度的控制成为可能,适合于动画虚拟角色,而无需演员表演或手动艺术干预。我们的研究旨在评估该方法(CVAEs)是否保留了数据集中不同强度水平相关的特征。结果表明,所提出的模型在不同强度水平上保留了关键的表现特征,同时支持情感强度水平的泛化,为从相对较小的数据集中创建情感丰富的虚拟角色做出了贡献。
cs.CV / 29 / 2608.21710

StereoDiffuer: Diffusion-based Progressive Geometry Modeling with Saliency Attention Perception for Stereo Matching

StereoDiffuer:基于扩散的渐进几何建模与显著性注意感知用于立体匹配
Li, Bohan
Abstract
With the advance of deep neural networks, the quality of disparity maps obtained through stereo matching has steadily improved. However, existing stereo matching methods still struggle to preserve fine-grained geometric details, resulting in blurred edges and over-smoothed predictions in challenging regions. To address these limitations, we propose StereoDiffuer, an iterative diffusion-based stereo matching framework that explicitly models geometric details and progressively refines disparity estimates. The framework incorporates a Saliency Attention Perception (SAP) module to extract salient geometric cues, including object boundaries, thin structures, and sharp edges. Confidence-guided SAP features are combined with the initial disparity estimate to condition an iterative denoising diffusion process, which corrects residual disparity errors and restores geometric details suppressed during cost-volume regularization and upsampling. Experimental results on the Scene Flow and KITTI benchmarks demonstrate the effectiveness of the proposed framework and its competitive performance relative to the compared stereo matching methods.
Chinese Translation
随着深度神经网络的发展,通过立体匹配获得的视差图质量稳步提高。然而,现有的立体匹配方法仍然难以保留细粒度的几何细节,导致在复杂区域出现模糊的边缘和过度平滑的预测。为了解决这些局限性,我们提出了StereoDiffuer,一种迭代的基于扩散的立体匹配框架,该框架明确建模几何细节并逐步细化视差估计。该框架结合了显著性注意感知(Saliency Attention Perception, SAP)模块,以提取显著的几何线索,包括物体边界、细结构和锐利边缘。基于置信度的SAP特征与初始视差估计相结合,以条件化迭代去噪扩散过程,从而纠正残余视差误差并恢复在代价体积正则化和上采样过程中被抑制的几何细节。在Scene Flow和KITTI基准上的实验结果证明了所提框架的有效性及其相对于比较的立体匹配方法的竞争性能。
cs.CV / 30 / 2608.21713

The Plan, Not the Decoder: Diagnosing and Repairing Compositional Failure in Reasoning-Augmented Text-to-Image Generation

规划,而非解码器:诊断和修复推理增强的文本到图像生成中的组合失败
Gonuguntla, Ashritha
Abstract
Reasoning-augmented text-to-image models such as GoT-R1 emit an explicit textual plan - object names, attributes, and bounding boxes - before generating image tokens. When such a model fails a compositional prompt, is the plan wrong, or is the plan right and the decoder unfaithful? Because the plan is machine-readable it can be edited before decoding, which makes the two separable. We first validate the ruler. Swapping the two bounding boxes inside the model's own chain demonstrably flips the generated layout: detector-based accuracy falls 0.75 -> 0.48 (p<1e-3), while a widely used VQA-based spatial metric rises. A five-rater human study agrees with the detector on 81% of items and with the VQA judge on 57%. All spatial results therefore use geometric scoring. Under sound measurement the decoder is a faithful executor: 94% of generated layouts realize the planned relation, and object-box binding survives reordering of the plan's object segments. The planner is the bottleneck. It writes wrong relations for phrasing-dependent reasons - 98% accuracy on "left" against 54% on "right" for semantically identical layouts, a raster-order bias we isolate with a mention-order control - and cluttered geometry that the decoder faithfully reproduces. Editing the plan therefore fixes the image without retraining: symbolic verification with resampling gives +5.0 points (p<1e-3), minimal in-place repair +6.0 (p=.02), rewriting only box geometry +10.7 (p<1e-4), and replacing the plan outright +13.3 (p=1e-4). Gains are indifferent to the plan's prose style and to its likelihood under the planner, but not to its geometry. Modular planner-decoder designs are therefore viable, provided the plan is internally consistent: box-text contradictions induce object duplication and identity fusion. We release the plan-fidelity evaluation protocol, all plans, and 12k generated images.
Chinese Translation
推理增强的文本到图像模型,如 GoT-R1,在生成图像标记之前会发出明确的文本计划——对象名称、属性和边界框。当这样的模型未能满足组合提示时,问题出在计划上,还是计划是正确的而解码器不忠实?由于计划是机器可读的,因此可以在解码之前进行编辑,这使得两者可以分开。我们首先验证了这一点。在模型自身链内交换两个边界框显著改变了生成的布局:基于检测器的准确率从 0.75 降至 0.48 (p<1e-3),而广泛使用的基于 VQA 的空间指标则上升。五位评审的人工研究在 81% 的项目上与检测器一致,在 57% 的项目上与 VQA 评审一致。因此,所有空间结果均使用几何评分。在合理的测量下,解码器是一个忠实的执行者:94% 的生成布局实现了计划中的关系,且对象-框绑定在计划的对象段重新排序后仍然有效。规划者是瓶颈。由于依赖于措辞的原因,它写出了错误的关系——在语义上相同的布局中,“左” 的准确率为 98%,而“右” 的准确率为 54%,这是我们通过提及顺序控制所隔离的光栅顺序偏差——以及解码器忠实再现的杂乱几何。因此,编辑计划可以在不重新训练的情况下修复图像:符号验证与重采样给出 +5.0 分 (p<1e-3),最小的就地修复 +6.0 (p=.02),仅重写框几何 +10.7 (p<1e-4),以及完全替换计划 +13.3 (p=1e-4)。收益与计划的文体风格及其在规划者下的可能性无关,但与其几何形状有关。因此,模块化的规划者-解码器设计是可行的,前提是计划内部一致:框-文本矛盾会导致对象重复和身份融合。我们发布了计划保真度评估协议、所有计划和 12,000 张生成图像。
cs.CV / 31 / 2608.21748

Calibrate What You SHIP: Post-Selection Risk Control for Verifier-Guided Text-to-Image Generation

校准您所发布的内容:基于验证者指导的文本到图像生成的后选择风险控制
Yin, Xuanhua, Mao, Shunqi, Guo, Wei, Xu, Chuanzhi, Cai, Weidong
Abstract
Verifier-guided text-to-image systems increasingly use test-time search to select, refine, or stop among multiple candidates, yet release thresholds are often calibrated on individual images. This creates a candidate-to-policy calibration mismatch: search changes both which prompts receive an output and which candidate is released, so candidate-level risk control need not imply control of released-output risk. We formalize this estimand shift through prompt reweighting and within-prompt selection, and introduce SHIP, Selection-aware Held-out calibration of Inference Policies. SHIP runs or replays the complete deployed policy on held-out prompts, evaluates the image it actually releases using an independent target judge, and selects the most permissive threshold whose risk upper bound satisfies a prescribed budget. For replayable policies with a prespecified threshold grid, simultaneous confidence control provides finite-sample validity. Experiments across fixed, sequential, and adaptive T2I inference procedures show that policy-level calibration recovers lower-risk operating points while exposing policy-dependent tradeoffs among risk, coverage, and compute. On GenEval2 with FLUX at N=16, a pooled-candidate threshold yields released risk 0.310, whereas SHIP reduces it to 0.162. Across 200 cached-stream splits, the fixed-grid certificate has no target crossing. Reliable inference-time scaling therefore requires calibrating the output distribution induced by the complete deployed policy.
Chinese Translation
基于验证者指导的文本到图像系统越来越多地使用测试时搜索来选择、细化或停止多个候选项,但发布阈值通常是在单个图像上进行校准的。这导致了候选项与策略之间的校准不匹配:搜索改变了哪些提示会产生输出以及哪个候选项被发布,因此候选级别的风险控制并不一定意味着对发布输出风险的控制。我们通过提示重加权和提示内选择形式化了这种估计量的转变,并引入了SHIP(Selection-aware Held-out calibration of Inference Policies)。SHIP在保留的提示上运行或重放完整的已部署策略,使用独立的目标评估者评估其实际发布的图像,并选择满足规定预算的风险上限的最宽松阈值。对于具有预先指定阈值网格的可重放策略,同时置信控制提供有限样本有效性。在固定、顺序和自适应的文本到图像推理程序中的实验表明,策略级校准恢复了较低风险的操作点,同时暴露了风险、覆盖率和计算之间的策略依赖权衡。在GenEval2上,使用FLUX且N=16时,汇总候选阈值导致发布风险为0.310,而SHIP将其降低至0.162。在200个缓存流拆分中,固定网格证书没有目标交叉。因此,可靠的推理时间扩展需要校准由完整已部署策略引起的输出分布。
cs.CV / 32 / 2608.21754

Fidelity-Diversity-Consistency (FDC): Data Pruning for Remote Sensing Change Detection

保真性-多样性-一致性 (FDC):遥感变化检测的数据剪枝
Zhu, Dongyao, Vatsavai, Ranga Raju
Abstract
Despite the success of data pruning (DP) in reducing training data sizes and improving downstream model performance in classification and segmentation tasks, its potential in remote sensing change detection remains unexplored. For the first time, we benchmark six representative DP methods across building- and forest-change datasets, CNN- and transformer-based models, and three pruning budgets, and show that existing baselines yield no reliable advantage over random selection. Notably, even the strongest evaluated baseline, Feature Diversity, is matched or exceeded by $\sim$33\% of randomly sampled subsets. To understand the underlying mechanism, we conduct a systematic regression study over 540 randomly sampled data subsets, characterizing each with four descriptors covering label statistics, image diversity, and feature-space geometry. Random Forest models show that \emph{change distribution fidelity} is the most prominent factor in determining the quality of change detection data subsets, a property absent from the existing pruning literature. Our analyses further show that pixel-wise image diversity and label-feature consistency are secondary factors. We translate these findings into Fidelity-Diversity-Consistency (FDC), a simple two-stage pruning method that shows consistent improvements over existing baselines across change detection benchmarks and backbones, especially at lower pruning ratios. Code is available at \href{https://github.com/ddydyd32/fidelity-diversity-consistency}{https://github.com/ddydyd32/fidelity-diversity-consistency}.
Chinese Translation
尽管数据剪枝 (DP) 在减少训练数据量和提高分类与分割任务中下游模型性能方面取得了成功,但其在遥感变化检测中的潜力仍未得到探索。首次,我们对六种具有代表性的 DP 方法在建筑和森林变化数据集、基于 CNN 和变换器的模型以及三种剪枝预算下进行了基准测试,结果表明现有基线在随机选择上并未表现出可靠的优势。值得注意的是,即使是评估的最强基线——特征多样性,也被约 33% 的随机抽样子集所匹敌或超越。为了理解其背后的机制,我们对 540 个随机抽样的数据子集进行了系统的回归研究,使用四个描述符对每个子集进行了特征化,这些描述符涵盖了标签统计、图像多样性和特征空间几何。随机森林模型表明, extit{变化分布保真性} 是决定变化检测数据子集质量的最重要因素,而这一特性在现有的剪枝文献中并不存在。我们的分析进一步表明,像素级图像多样性和标签-特征一致性是次要因素。我们将这些发现转化为保真性-多样性-一致性 (FDC),这是一种简单的两阶段剪枝方法,在变化检测基准和主干模型中相较于现有基线显示出一致的改进,尤其是在较低的剪枝比例下。代码可在 exttt{https://github.com/ddydyd32/fidelity-diversity-consistency} 获取。
cs.CV / 33 / 2608.21762

Learning to Look Again: Loss-Gap Supervision for Free-form Crop Routing in Vision-Language Models

重新审视学习:视觉-语言模型中自由形式作物路径的损失-间隙监督
Zhu, Jinchang, Fu, Rong, Ding, Yi, Wu, Chenghao, Liu, Ying, Yang, Menglin
Abstract
Vision-language models (VLMs) fail many detail-centric questions for a concrete reason: the answer is visible in the image, yet lost after the image is compressed into a low-resolution global view. Allocating more visual tokens to every query improves some OCR and document cases, but it spends computation indiscriminately and can disturb tasks that rely on global context. We propose GapSight, a framework for learning visual re-reading: a VLM first takes a global glance, then selectively returns to a free-form region when the question calls for local evidence. The supervision comes from the target model's own failure signal. Offline, we compare answer loss or multiple-choice option margin under a global-only view and candidate crop-augmented views; crops that improve the target answer become model-specific review labels. A lightweight free-form crop router distills these labels into a one-shot inference policy that predicts whether to review, expected utility, and a continuous crop box from the global state. Across LLaVA-1.5-7B, InternVL2.5-8B, and Qwen2-VL-2B-Instruct, GapSight improves the Base no-zoom baseline on six benchmarks spanning OCR, documents, charts, infographics, VStarBench, and MME-RealWorld-Lite. On InternVL2.5-8B, GapSight raises the six-benchmark average from 52.25 to 64.29, above CropVLM (57.16), ViCrop (55.84), and ZoomRefine (54.43). Mechanism analyses show that the router rescues concrete wrong answers, adapts its action rate by task, and forms a favorable token-performance profile. These results position loss-gap supervision as a practical route to teaching VLMs when and where to look again.
Chinese Translation
视觉-语言模型(VLMs)在许多细节导向的问题上表现不佳,其原因在于:答案在图像中可见,但在图像被压缩为低分辨率全局视图后丢失。为每个查询分配更多的视觉标记可以改善一些OCR和文档案例,但这无差别地消耗计算资源,并可能干扰依赖于全局上下文的任务。我们提出了GapSight,一个学习视觉重读的框架:VLM首先进行全局观察,然后在问题需要局部证据时选择性地返回到自由形式区域。监督来自目标模型自身的失败信号。在离线阶段,我们比较了在仅全局视图下的答案损失或多选项边际与候选裁剪增强视图下的表现;那些改善目标答案的裁剪成为模型特定的复审标签。一个轻量级的自由形式裁剪路由器将这些标签提炼为一种单次推理策略,该策略预测是否需要复审、预期效用以及来自全局状态的连续裁剪框。在LLaVA-1.5-7B、InternVL2.5-8B和Qwen2-VL-2B-Instruct上,GapSight在涵盖OCR、文档、图表、信息图、VStarBench和MME-RealWorld-Lite的六个基准测试中改善了基础无缩放基线。在InternVL2.5-8B上,GapSight将六个基准的平均值从52.25提高到64.29,超过了CropVLM(57.16)、ViCrop(55.84)和ZoomRefine(54.43)。机制分析表明,路由器能够挽救具体的错误答案,按任务调整其动作率,并形成有利的标记-性能特征。这些结果将损失-间隙监督定位为教导VLM何时以及何地重新审视的实用途径。
cs.CV / 34 / 2608.21764

LiteEvent-AE: Lightweight Autoencoder for Event-Based Vision on Low-Latency Energy-Constrained Edge Devices

LiteEvent-AE:低延迟能量受限边缘设备上的轻量级自编码器用于事件驱动视觉
Islam, Riadul, Mule, Joey, Challagundla, Dhandeep, Rizvi, Shahmir, Carson, Sean, Saini, Rachit
Abstract
Event-based vision has emerged as a promising paradigm for energy-aware artificial intelligence (AI), offering sparse, low-latency visual signals that reduce redundant data processing and support sustainable edge computing. However, the asynchronous and noise-prone nature of event streams creates challenges for conventional deep learning models, which are often too computationally intensive for low-power embedded platforms. This work presents a compact and configurable event-driven autoencoder that efficiently compresses neuromorphic data while preserving essential spatiotemporal structure for downstream inference. The architecture integrates lightweight convolutional encoding with robust performance under adaptive event thresholding and a minimal classifier head, enabling substantial reductions in computational cost without degrading recognition fidelity. Extensive evaluations on the Smart Event Face Dataset (SEFD) and Event-Based Crossing Dataset (EBCD) show that the proposed framework achieves competitive or superior accuracy compared to YOLOv9 while requiring up to 35.6$\times$ fewer parameters. To assess real-world sustainability, the model is deployed on resource-constrained hardware: a Raspberry Pi 4B and a NVIDIA Jetson Nano. On NVIDIA Jetson Nano, it delivers real-time throughput of 44.8 FPS. On a Raspberry Pi 4B CPU, the 50\% autoencoder classifier consumes 16.19 J for the evaluated inference workload, corresponding to approximately 726.3$\times$ lower energy consumption than YOLOv9 under the same evaluation protocol. These results demonstrate the potential of compact event-driven models to advance environmentally conscious, low-power AI systems for high-speed perception in autonomous, mobile, and embedded computing environments.
Chinese Translation
基于事件的视觉已成为一种有前景的能量感知人工智能(AI)范式,提供稀疏、低延迟的视觉信号,从而减少冗余数据处理并支持可持续的边缘计算。然而,事件流的异步和噪声易感特性给传统深度学习模型带来了挑战,这些模型通常对低功耗嵌入式平台来说计算负担过重。本研究提出了一种紧凑且可配置的事件驱动自编码器,能够高效压缩神经形态数据,同时保留下游推理所需的基本时空结构。该架构将轻量级卷积编码与在自适应事件阈值下具有强大性能的最小分类器头相结合,实现了在不降低识别精度的情况下显著降低计算成本。在Smart Event Face Dataset (SEFD) 和 Event-Based Crossing Dataset (EBCD) 上的广泛评估表明,所提出的框架在准确性上与YOLOv9相当或更优,同时所需参数量减少了多达35.6倍。为了评估其在现实世界中的可持续性,该模型被部署在资源受限的硬件上:Raspberry Pi 4B和NVIDIA Jetson Nano。在NVIDIA Jetson Nano上,它实现了44.8 FPS的实时吞吐量。在Raspberry Pi 4B CPU上,50%的自编码器分类器在评估的推理工作负载中消耗了16.19焦耳,约为YOLOv9在相同评估协议下的726.3倍更低的能量消耗。这些结果展示了紧凑型事件驱动模型在推动环保、低功耗AI系统方面的潜力,以实现自主、移动和嵌入式计算环境中的高速感知。
cs.CV / 35 / 2608.21776

SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling

SpatialDiff:通过隐式空间建模实现3D感知的物体运动
Liu, Zheng, He, Zijian, He, Huiguo, Zhong, Weizhi, Tang, Yejun, Yang, Huan, Gai, Kun, Li, Guanbin
Abstract
Recent advances in image editing allow impressive manipulation of objects, existing methods still struggle to handle spatial movement in complex scenes, such as objects span different depth layers or are partially occluded. Most image editing methods focus solely on prior information from 2D datasets, emphasizing planar features while lacking support for spatial structures. Even approaches that incorporate explicit positional information fail to capture true 3D spatial relationships, thus limiting accurate object movement in complex scenes. In this paper, we present SpatialDiff, a method that effectively captures 3D spatial structures, enabling precise and consistent object movements in complex scenes. Our core innovations are twofold: (1) Implicit 3D Spatial Modeling, which introduces 3D prior knowledge and enables the model to internally build a comprehensive understanding of the three-dimensional spatial structure; and (2) Global Spatial Supervision, which constrains the latent spatial features to enable the model to perceive changes in object spatial positions caused by editing operations. Experimental results demonstrate that our method significantly improves the accuracy and fidelity of spatial movement in complex scenes.
Chinese Translation
近期图像编辑的进展使得物体的操控变得令人印象深刻,但现有方法在处理复杂场景中的空间运动时仍然面临挑战,例如物体跨越不同的深度层或部分被遮挡。大多数图像编辑方法仅关注来自2D数据集的先前信息,强调平面特征,而缺乏对空间结构的支持。即使是那些结合了显式位置信息的方法,也未能捕捉到真实的3D空间关系,从而限制了在复杂场景中物体运动的准确性。本文提出了SpatialDiff,一种有效捕捉3D空间结构的方法,使得在复杂场景中实现精确且一致的物体运动成为可能。我们的核心创新有两个方面:(1)隐式3D空间建模,它引入了3D先验知识,使模型能够内部构建对三维空间结构的全面理解;(2)全局空间监督,它约束潜在的空间特征,使模型能够感知因编辑操作而导致的物体空间位置变化。实验结果表明,我们的方法显著提高了复杂场景中空间运动的准确性和保真度。
cs.CV / 36 / 2608.21784

DefaultShift: Auditing Semantic Default Shift in Accelerated Text-to-Image Models

DefaultShift:加速文本到图像模型中的语义默认偏移审计
Yin, Xuanhua, Xu, Chuanzhi, Mao, Shunqi, Guo, Wei, Cai, Weidong
Abstract
Few-step text-to-image models increasingly replace slower generators, yet acceleration can silently change distributions over unspecified attributes even when individual outputs remain plausible and aligned. We call these distributions semantic defaults and their change under replacement semantic default shift. Existing quality, preference, and diversity evaluations do not test whether a replacement preserves its reference model's semantic defaults. We introduce DefaultShift, a paired audit that labels repeated samples with closed semantic vocabularies, measures probability-mass movement, and separates interpretable ranking from confirmatory cross-fit inference. Across 14 reference and replacement pairs, adjusted color discrepancies range from 0.054 to 0.303 with recipe-specific directions. A 1,000-image human audit reproduces the ordering. We further introduce DefaultShift-Select, an offline calibration method that reduces human-measured shift by 10.3 percent to 35.1 percent across Turbo, DMD2, and FLUX without material quality loss. Under balanced evaluation, selected data recover 4.3 accuracy points and 7.5 worst-group points over uncalibrated replacement data. DefaultShift makes semantic preservation under acceleration measurable and actionable.
Chinese Translation
少步文本到图像模型越来越多地取代较慢的生成器,然而加速可能在未指定属性上悄然改变分布,即使单个输出仍然合理且一致。我们称这些分布为语义默认值,其在替换下的变化为语义默认偏移。现有的质量、偏好和多样性评估并未测试替换是否保留其参考模型的语义默认值。我们引入了DefaultShift,这是一种配对审计,使用封闭的语义词汇对重复样本进行标记,测量概率质量的移动,并将可解释的排名与确认性交叉拟合推断分开。在14对参考和替换模型中,调整后的颜色差异范围从0.054到0.303,具有特定的配方方向。通过对1,000幅图像进行人工审计,重现了排序。我们进一步引入DefaultShift-Select,这是一种离线校准方法,在Turbo、DMD2和FLUX中将人类测量的偏移减少了10.3%至35.1%,且没有实质性的质量损失。在平衡评估下,选定的数据比未校准的替换数据恢复了4.3个准确度点和7.5个最差组点。DefaultShift使得在加速下的语义保留可测量且可操作。
cs.CV / 37 / 2608.21786

HP-UniIF: Hierarchical Prompt Learning for Unified Image Fusion

HP-UniIF:用于统一图像融合的层次化提示学习
Xu, Xingxin, Zhao, Siqi, Li, Xin, Yao, Xinjie, Sun, Yiming, Zhu, Pengfei
Abstract
General image fusion seeks to integrate complementary information from multiple source images, yet real-world applications often require a single system to support heterogeneous fusion, degradation restoration, and task-oriented perception simultaneously. Existing unified frameworks struggle with these orthogonal objectives, resulting in entangled representations and degraded performance across subtasks. We propose HP-UniIF, a unified vision framework that leverages diffusion priors to bridge heterogeneous fusion, visual restoration, and downstream perception. To address the limited adaptability of diffusion models to domain-, degradation-, and task-level objectives within one pipeline, HP-UniIF introduces a depth-wise hierarchical conditional modulation strategy that decouples these objectives across network stages. Task prompt modulation at bottleneck layers adapts the backbone to different fusion paradigms, the degradation prompt router at shallow layers injects degradation-aware constraints for local restoration, and the application prompt bank at decoding stages aligns generation with downstream tasks. This hierarchical design enables HP-UniIF to produce visually faithful results while preserving task-relevant semantics. Extensive experiments across multiple fusion tasks, diverse degradations, and various downstream applications demonstrate the superior performance of HP-UniIF.
Chinese Translation
一般图像融合旨在整合来自多个源图像的互补信息,但现实世界的应用通常需要一个单一系统同时支持异构融合、退化恢复和任务导向的感知。现有的统一框架在处理这些正交目标时表现不佳,导致表示混淆和子任务性能下降。我们提出了HP-UniIF,一个统一的视觉框架,利用扩散先验来连接异构融合、视觉恢复和下游感知。为了解决扩散模型在一个管道中对领域、退化和任务级目标的适应性有限的问题,HP-UniIF引入了一种深度层次条件调制策略,在网络阶段之间解耦这些目标。瓶颈层的任务提示调制使主干网络适应不同的融合范式,浅层的退化提示路由器为局部恢复注入了退化感知约束,而解码阶段的应用提示库则使生成与下游任务对齐。这种层次化设计使HP-UniIF能够在保持任务相关语义的同时产生视觉上真实的结果。在多个融合任务、各种退化和不同下游应用中的广泛实验表明,HP-UniIF具有优越的性能。
cs.CV / 38 / 2608.21796

SAFE-G: Structure-aware Faithful Evidence-guided Generation for Knowledge-based Visual Question Answering

SAFE-G:结构感知的可信证据引导生成用于基于知识的视觉问答
Shu, Long, Liu, Shuochen, Chen, Wei, Lin, Junda, Zheng, Zhi, Hou, Huijun, Xu, Tong
Abstract
Knowledge-based Visual Question Answering (KB-VQA) aims to answer queries that necessitate reasoning over external knowledge sources beyond the visual content. Typically, current methods fuse multimodal features to retrieve external information, subsequently leveraging Multimodal Large Language Models (MLLMs) to derive answers from the retrieved evidence. However, these methods often struggle to capture structural associations within complex contexts to effectively filter noise. Furthermore, they frequently fail to ensure that the reasoning process remains strictly faithful to the retrieved evidence. To address these challenges, we propose SAFE-G, a Structure-Aware Faithful Evidence-guided Generation framework, which enables precise evidence localization and trustworthy reasoning. Specifically, we first employ a coarse-grained hybrid search fusing visual and textual modalities to recall candidate documents, and subsequently implement a structure-aware fine-grained graph retrieval that captures structural dependencies to filter noise and pinpoint precise evidence. Moreover, we introduce a reinforcement learning (RL) strategy with an evidence-grounded reward that assigns credit to correct answers only when the selected evidence is correct. This strict alignment constraint compels the model to anchor its response in the retrieved context, effectively enhancing its capability to locate evidence via multimodal features and perform faithful reasoning. Extensive experiments on the Encyclopedic-VQA and InfoSeek benchmarks demonstrate that SAFE-G outperforms prior methods by a margin of 8.9% and 3.5%, substantially enhancing the overall reasoning accuracy. Our source code is publicly available at: https://github.com/MINE-USTC/SAFE-G.
Chinese Translation
基于知识的视觉问答(KB-VQA)旨在回答需要对超越视觉内容的外部知识源进行推理的查询。通常,当前的方法融合多模态特征以检索外部信息,随后利用多模态大型语言模型(MLLMs)从检索到的证据中推导答案。然而,这些方法往往难以捕捉复杂上下文中的结构关联,从而有效过滤噪声。此外,它们常常无法确保推理过程严格忠实于所检索的证据。为了解决这些挑战,我们提出了SAFE-G,一个结构感知的可信证据引导生成框架,能够实现精确的证据定位和可靠的推理。具体而言,我们首先采用粗粒度的混合搜索,融合视觉和文本模态以召回候选文档,随后实施结构感知的细粒度图检索,捕捉结构依赖关系以过滤噪声并精确定位证据。此外,我们引入了一种强化学习(RL)策略,采用基于证据的奖励机制,仅在所选证据正确时给予正确答案以信用。这一严格的对齐约束迫使模型将其响应锚定在检索到的上下文中,有效增强其通过多模态特征定位证据和进行可信推理的能力。在Encyclopedic-VQA和InfoSeek基准上的大量实验表明,SAFE-G的表现优于先前的方法,提升幅度分别为8.9%和3.5%,显著提高了整体推理准确性。我们的源代码已公开,地址为:https://github.com/MINE-USTC/SAFE-G。
cs.CV / 39 / 2608.21804

FlashReg: GPU-Accelerated 3-Clique Point Cloud Registration for Real-Time Correspondence-to-Pose Estimation

FlashReg:用于实时对应到姿态估计的GPU加速三团点云配准
Yu, Ziyang, Li, Xiang, Chang, Qiong, Miyazaki, Jun
Abstract
Graph-based point cloud registration achieves high robustness by identifying geometrically consistent correspondence sets, but constructing second-order compatibility graphs and enumerating candidate cliques remain compute- and memory-intensive. This work presents FlashReg, a GPU-oriented correspondence-to-pose estimator that avoids materializing the dense scored second-order graph. Its Fast First- and Second-Order Graph (FFSOG) construction builds a capacity-bounded sparse second-order graph directly from the binary first-order graph. A dataflow-optimized three-node clique (3-clique) search then selects pivots from compact per-row candidate pools and enumerates triples through sorted sparse-neighborhood intersections. Across indoor and outdoor benchmarks, FlashReg reduces correspondence-to-pose latency by 2--3x relative to TurboReg at comparable registration recall, while using about 50% of its peak allocated tensor memory on an embedded GPU. These results make FlashReg suitable as a high-throughput registration backend within onboard perception pipelines.
Chinese Translation
基于图的点云配准通过识别几何一致的对应集实现了高鲁棒性,但构建二阶兼容性图和枚举候选团仍然计算和内存密集。本文提出了FlashReg,一种面向GPU的对应到姿态估计器,避免了物化稠密评分的二阶图。其快速一阶和二阶图(Fast First- and Second-Order Graph,FFSOG)构建直接从二进制一阶图构建一个容量受限的稀疏二阶图。经过数据流优化的三节点团(3-clique)搜索从紧凑的每行候选池中选择枢轴,并通过排序的稀疏邻域交集枚举三元组。在室内和室外基准测试中,FlashReg相对于TurboReg在可比的配准召回率下将对应到姿态的延迟减少了2-3倍,同时在嵌入式GPU上使用了约50%的峰值分配张量内存。这些结果使得FlashReg适合作为机载感知管道中的高吞吐量配准后端。
cs.CV / 40 / 2608.21813

Through the Schr\"odinger Bridge: Benchmarking Antemortem Image Restoration from Postmortem Autolysis to Enhance Forensic Diagnostics

通过薛定谔桥:基于尸后自溶的生前影像恢复基准测试以增强法医学诊断
Hao, Shuang, Yue, Jiacheng, Zhao, Yaxuan, Wang, Fan, Ma, Jianhua, Huang, Erwen, Lian, Chunfeng
Abstract
Forensic histopathology, essential for determining cause of death and disease diagnosis, is severely impeded by postmortem autolysis, i.e., an irreversible, stochastic degradation process that distorts tissue morphology and introduces diagnostic subjectivity, thereby underscoring the value of restoring autolyzed images to a diagnostically plausible, pre-autolysis state for improving objectivity in forensic practice. This restoration task is fundamentally challenging due to the large, non-deterministic morphological changes caused by autolysis and the infeasibility of pixel-wise paired data, which invalidates assumptions underlying supervised and cycle/structure-consistent unpaired translation methods. To address this, we formalize forensic histopathology autolysis restoration as a new task: under unpaired supervision, transform postmortem images with severe autolysis into diagnostically meaningful ``antemortem'' representations. We contribute AutoPath, the first homologous yet unpaired dataset for this problem, constructed by splitting specimens into adjacent tissue blocks---one processed immediately, the other exposed to induce autolysis---yielding nearly ten thousand $10\times$ patches from 69 cases with varying liver conditions. We further frame the problem as a Schr\"odinger Bridge between the autolyzed and non-autolyzed distributions, offering a principled approach to modeling stochastic, severe morphological degradation. Critically, we demonstrate the misalignment of generic image-level generative metrics (e.g., FID) with diagnostic utility and propose a forensically grounded, slide-level diagnostic distribution consistency evaluation. Overall, this work establishes a reproducible benchmark (encompassing task definition, a real-world dataset, and an evaluation methodology) toward rigorous and practically meaningful progress in autolysis restoration for forensic pathology.
Chinese Translation
法医学组织病理学对于确定死亡原因和疾病诊断至关重要,但受到尸后自溶的严重影响,即一种不可逆转的随机降解过程,它扭曲了组织形态并引入了诊断的主观性,因此强调了将自溶图像恢复到诊断上合理的生前状态以提高法医学实践客观性的价值。由于自溶引起的大规模非确定性形态变化以及像素级配对数据的不可行性,这一恢复任务在根本上是具有挑战性的,这使得监督学习和循环/结构一致的无配对翻译方法的基本假设失效。为了解决这个问题,我们将法医学组织病理学自溶恢复形式化为一项新任务:在无配对监督下,将严重自溶的尸后图像转化为具有诊断意义的“生前”表征。我们贡献了AutoPath,这是针对该问题的第一个同源但无配对的数据集,通过将标本分割为相邻的组织块构建而成——一个立即处理,另一个暴露以诱导自溶——从69个不同肝脏状况的案例中获得近一万个$10 imes$补丁。我们进一步将问题框架化为自溶与非自溶分布之间的薛定谔桥,提供了一种建模随机严重形态降解的原则性方法。关键的是,我们展示了通用图像级生成指标(例如,FID)与诊断效用之间的不一致,并提出了一种基于法医学的幻灯片级诊断分布一致性评估。总体而言,这项工作建立了一个可重复的基准(涵盖任务定义、真实世界数据集和评估方法),以推动法医学病理学中自溶恢复的严格且具有实际意义的进展。
cs.CV / 41 / 2608.21819

PatchGate: Narrowing the Verbalization Gap with Intrinsic Object Inventories in Frozen Vision-Language Models

PatchGate:通过冻结的视觉-语言模型中的内在对象库缩小语言表达差距
Ko, Jihyung, Jung, Eunji, Kim, Hyeongsub, Lee, Ziseok, Cho, Jae Won, Jo, Sanghyun, Kim, Kyungsu
Abstract
Reliable image captioning in Vision-Language Models (VLMs) requires captions to be both precise and complete, avoiding unsupported object mentions while covering visible objects. Existing training-free methods primarily address the former requirement, suppressing unsupported object words by intervening on model-predicted mentions during generation. Because they operate only on objects the model is already likely to mention, visible objects omitted from the output remain difficult to recover. We propose PatchGate, a training-free framework that extracts prompt-free object evidence intrinsic to a frozen VLM before generation and uses it to narrow the gap between an intrinsic object set and final object mentions. In the first stage, Visual Evidence eXtraction (VEX) reads patch-level lexical evidence from the latter half of LM decoder layers and constructs an image-conditioned object set without any task prompt. In the second stage, Visual-Evidence Inclusion-Exclusion Decoding (VIED) uses this object evidence to calibrate decoding logits, promoting evidence-supported but under-verbalized objects and suppressing weakly supported but over-verbalized objects. On AMBER, PatchGate improves both sides of object-level reliability, increasing visible-object coverage from 49.4 to 56.0 (+13.4%) and reducing object hallucination by lowering CHAIR from 7.5 to 6.6 (-12.0%), without external detectors or fine-tuning and with one extra forward pass.
Chinese Translation
在视觉-语言模型(VLMs)中,可靠的图像描述要求描述既要准确又要完整,避免不支持的对象提及,同时覆盖可见对象。现有的无训练方法主要解决前者的要求,通过在生成过程中干预模型预测的提及来抑制不支持的对象词。由于这些方法仅针对模型已经可能提及的对象,因此输出中遗漏的可见对象仍然难以恢复。我们提出了PatchGate,这是一种无训练框架,在生成之前提取冻结的VLM中与提示无关的内在对象证据,并利用这些证据缩小内在对象集与最终对象提及之间的差距。在第一阶段,视觉证据提取(Visual Evidence eXtraction, VEX)从语言模型解码器层的后半部分读取补丁级词汇证据,并在没有任何任务提示的情况下构建图像条件的对象集。在第二阶段,视觉证据包含-排除解码(Visual-Evidence Inclusion-Exclusion Decoding, VIED)利用这些对象证据来校准解码logits,促进证据支持但表达不足的对象,并抑制证据支持弱但表达过多的对象。在AMBER数据集上,PatchGate提高了对象级可靠性的两个方面,将可见对象覆盖率从49.4提高到56.0(+13.4%),并通过将CHAIR的对象幻觉率从7.5降低到6.6(-12.0%)来减少对象幻觉,无需外部检测器或微调,并且只需额外一次前向传递。
cs.CV / 42 / 2608.21828

Towards Alias-Free 4D Gaussian Representations with Motion-Aware Filtering

朝着无别名的4D高斯表示与运动感知滤波的方向
Dhiman, Ankit, Kathare, Kunal A, Vignesh, Pranav, Boregowda, Lokesh R, Radhakrishnan, Venkatesh Babu
Abstract
Novel-view synthesis of dynamic scenes, crucial for AR/VR applications, remains a challenging problem. Recent methods adapt representations like 3D Gaussian Splatting (3DGS) and Neural Radiance Fields (NeRF) for dynamic scenes by incorporating time as the fourth dimension (4D representations). These 4D representations still suffer from aliasing artifacts, especially when generating novel views from divergent viewpoints (zoom-in/zoom-out operations). While using 3D smoothing filters like those proposed in Mip-Splatting might seem like a possible solution, they fail to account for local motion and also exhibit aliasing. To address this, we propose a motion-aware 3D smoothing filter specifically designed for 4D representations. Our approach adapts the filter strength based on local motion information, effectively mitigating aliasing without compromising rendering quality. This is achieved by estimating the joint density function of time and focal-to-depth ratio using a non-parametric estimation method. During inference, we sample from this joint distribution to determine the appropriate smoothing filter. This flexible strategy can be integrated with various 4D representations. Our evaluations on standard datasets demonstrate superior performance compared to state-of-the-art methods.
Chinese Translation
动态场景的新视角合成对于增强现实/虚拟现实(AR/VR)应用至关重要,但仍然是一个具有挑战性的问题。最近的方法通过将时间作为第四维度(4D表示)来调整如3D高斯点云(3D Gaussian Splatting, 3DGS)和神经辐射场(Neural Radiance Fields, NeRF)等表示,以适应动态场景。然而,这些4D表示在从不同视点(放大/缩小操作)生成新视角时仍然遭受别名伪影。虽然使用像Mip-Splatting中提出的3D平滑滤波器似乎是一个可能的解决方案,但它们未能考虑局部运动,并且也会出现别名现象。为了解决这个问题,我们提出了一种专门为4D表示设计的运动感知3D平滑滤波器。我们的方法根据局部运动信息调整滤波强度,有效减轻别名现象而不影响渲染质量。这是通过使用非参数估计方法估计时间与焦距深度比的联合密度函数来实现的。在推理过程中,我们从这个联合分布中采样,以确定适当的平滑滤波器。这种灵活的策略可以与各种4D表示集成。我们在标准数据集上的评估显示出比最先进的方法更优越的性能。
cs.CV / 43 / 2608.21837

Towards Bitstream-corrupted Harsh Visual Understanding: Through Bitstream Language Modeling as Robust Semantic Priors

面向比特流损坏的恶劣视觉理解:通过比特流语言建模作为鲁棒语义先验
Huang, Chaoran, Li, Fangcheng, Liu, Tianyi, Liu, Wenyang, Wu, Kejun
Abstract
Bitstream-corrupted Harsh Visual Understanding (BcHVU) aims to understand harshly degraded videos originally decoded from a severely corrupted bitstream in real-world multimedia communication. The ill-posed nature of BcHVU poses a major challenge for existing vision models, as even subtle bitstream corruption can lead to irreversible pixel distortion and significant semantic loss. To address these challenges in BcHVU, we propose Bitstream Language Modeling as Robust Semantic Priors (BLMSP), a framework for learning and injecting bitstream-native semantic cues. Our proposed BLMSP framework learns to extract bitstream-native semantic cues by bitstream language modeling, and leverages them as priors by injecting into off-the-shelf vision models of BcHVU tasks. Specifically, we present a Video Bitstream Byte Model (VBBM) that integrates byte-level modeling and cross-codec semantic distillation, enabling it to interpret robust semantics from byte sequences in multiple corrupted bitstream formats. The learned bitstream semantics are leveraged as robust priors and fused into BcHVU model backbones for improving the quality of video restoration, captioning, and human pose estimation. To train BLMSP, we construct a large-scale multi-source Corrupted-bitstream Harsh-video Paired (CHP) dataset containing 607k corrupted bitstream segments and 287k paired harsh video clips. Extensive experimental results show that the learned bitstream priors improve video restoration, captioning, and human pose estimation by 2.51 dB in PSNR, 0.20 in CIDEr, and 0.18 in [email protected] on average, respectively. These results demonstrate that corrupted bitstream can serve as robust semantic priors in solving pixel distortion and semantic loss in BcHVU.
Chinese Translation
比特流损坏的恶劣视觉理解(BcHVU)旨在理解从严重损坏的比特流中解码的恶劣降级视频,这在现实世界的多媒体通信中具有重要意义。BcHVU的病态特性对现有视觉模型构成了重大挑战,因为即使是微小的比特流损坏也可能导致不可逆的像素失真和显著的语义丧失。为了解决BcHVU中的这些挑战,我们提出了比特流语言建模作为鲁棒语义先验(BLMSP),这是一个学习和注入比特流原生语义线索的框架。我们提出的BLMSP框架通过比特流语言建模学习提取比特流原生语义线索,并将其作为先验注入到现成的BcHVU任务视觉模型中。具体而言,我们提出了一种视频比特流字节模型(VBBM),它集成了字节级建模和跨编解码器语义蒸馏,使其能够从多种损坏的比特流格式中的字节序列中解释鲁棒语义。学习到的比特流语义被作为鲁棒先验利用,并融合到BcHVU模型主干中,以提高视频恢复、字幕生成和人体姿态估计的质量。为了训练BLMSP,我们构建了一个大规模多源的损坏比特流恶劣视频配对(CHP)数据集,包含607k个损坏的比特流片段和287k个配对的恶劣视频剪辑。大量实验结果表明,学习到的比特流先验在视频恢复、字幕生成和人体姿态估计方面分别提高了2.51 dB的PSNR、0.20的CIDEr和0.18的[email protected]。这些结果表明,损坏的比特流可以作为解决BcHVU中像素失真和语义丧失的鲁棒语义先验。
cs.CV / 44 / 2608.21839

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

FIRM-Video:在评分前进行检查以实现可靠的文本到视频奖励建模
Zhang, Peiyuan, Zhao, Xiangyu, Liu, Hongbo, Hu, Xiaoxing, Liu, Mingxin, Ma, Shuran, Shen, Yunhang, Hu, Jian, Gao, Haihan, Cao, Haoyu, Yang, Xue
Abstract
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.
Chinese Translation
可靠的奖励模型对于文本到视频的评估和对齐至关重要。然而,评估准确性与推理效率之间的权衡对训练监督的质量提出了很高的要求。现有的方法通常依赖于具有固定评分标准的整体评审者或开放式推理,导致检查不全面、理由不真实和归因混乱。我们提出了FIRM-Video,一个基于检查前评分原则的统一检查表驱动数据构建框架:构建特定维度的检查表,针对时间视觉证据验证每个标准,并仅聚合经过验证的决策。在指令遵循方面,FIRM-Video将提示分解为加权原子要求;在世界一致性方面,它构建了基于可见实体和动作的提示校准、目标特定的检查;在感知质量方面,它应用了一种通用的视觉缺陷分类法。经过验证的标准和评分进一步转化为自然语言分析,以实现端到端的奖励建模。随后,我们构建了FIRM-Video-90K,包含来自29,348个视频的88,044个特定维度实例,并引入FIRM-Video-Bench,涵盖250个视频的750个逐点人类注释。基于Qwen3-VL的FIRM-Video-8B在FIRM-Video-Bench上实现了最佳的整体平均绝对误差(MAE),同时在三个视频生成器的最佳8次抽样中持续提供最高的VBench总分、质量分和语义分。
cs.CV / 45 / 2608.21847

BC-IHV: Conditioning the Color Space for Stable Rectified-Flow Low-Light Enhancement

BC-IHV:为稳定的校正流低光增强调节色彩空间
Ai, Yi, Chen, Zheng, Cai, Yuanhao, Zhang, Yulun, Yang, Xiaokang
Abstract
Low-light image enhancement (LLIE) must correct ambiguous exposure without overwriting structure already supported by the input. Generative transport can model exposure ambiguity; however, its flexibility may also alter observable geometry and chromatic content. Moreover, fixed invertible color coordinates are usually treated only as representations, although their inverse mappings reshape the RGB-domain gradients received by the enhancement network. To address these issues, we propose Structure-Anchored Rectified Flow (SA-RF), which maintains correspondence through separate chromaticity/intensity stems, a scale-matched condition pyramid, and HybridAda. HybridAda assigns location-specific retrieval to spatial cross-attention and global exposure modulation to pooled AdaLN. We further introduce BC-IHV, a learnable Box--Cox polar color space whose analytically invertible intensity mapping controls the inverse-gradient dynamic range through a single exponent. This allows the representation to balance dark-range expansion and gradient conditioning instead of adopting a fixed linear or logarithmic law. Experiments on three LOL benchmarks, blind image-quality evaluation, and cross-dataset tests demonstrate consistent reconstruction and perceptual advantages over the sota. Controlled studies further support the effectiveness of both the proposed framework and color representation.
Chinese Translation
低光图像增强(LLIE)必须在不覆盖输入中已支持的结构的情况下纠正模糊的曝光。生成传输可以建模曝光模糊性;然而,它的灵活性也可能改变可观察的几何形状和色彩内容。此外,固定的可逆色彩坐标通常仅被视为表示,尽管它们的逆映射会重塑增强网络接收到的RGB域梯度。为了解决这些问题,我们提出了结构锚定校正流(SA-RF),通过独立的色度/强度支撑、尺度匹配的条件金字塔和HybridAda保持对应关系。HybridAda将位置特定的检索分配给空间交叉注意力,并将全局曝光调制分配给池化的AdaLN。我们进一步引入BC-IHV,一种可学习的Box-Cox极坐标色彩空间,其解析可逆的强度映射通过单个指数控制逆梯度动态范围。这使得表示能够平衡暗区扩展和梯度调节,而不是采用固定的线性或对数法则。在三个LOL基准、盲图像质量评估和跨数据集测试中的实验表明,与现有最优方法相比,具有一致的重建和感知优势。控制研究进一步支持所提出的框架和色彩表示的有效性。
cs.CV / 46 / 2608.21849

GaussVid: Sparse-View Gaussian Splatting with 3D-Aware Video Diffusion Priors

GaussVid:基于稀疏视图的高斯点云与3D感知视频扩散先验
Liu, Xinhui, Wang, Can, Jiang, Wei, Wang, Wei, Xu, Dong
Abstract
3D Gaussian Splatting (3DGS) has achieved remarkable success in novel view synthesis; however, reconstructions under sparse views often exhibit noticeable artifacts. While recent video diffusion models provide strong spatio-temporal priors for 3DGS restoration, directly fine-tuning them for restoration is suboptimal, as they lack awareness of the underlying multi-camera geometry, resulting in multi-view inconsistencies. In this work, we propose a novel 3D-aware video restoration framework designed to enhance the quality of sparse 3DGS reconstruction. Specifically, we construct a large-scale 3DGS video dataset to enable specialized fine-tuning. To bridge the gap between 2D video generation and 3D multi-view constraints, we introduce a camera-conditioned geometric prior. By using the first and last frames as boundary anchors and encoding the corresponding camera relationships, we explicitly inject spatial structure into the video generation pipeline. This boundary-anchored, camera-aware prior guides the network toward geometrically grounded restoration that remains coherent across viewpoints. Extensive experiments show that, among video-prior restoration methods, our approach attains the best pixel- and structure-level fidelity (PSNR/SSIM) and improves multi-view consistency, while remaining competitive in perceptual quality (LPIPS).
Chinese Translation
3D高斯点云(3DGS)在新视图合成中取得了显著成功;然而,在稀疏视图下的重建往往会出现明显的伪影。尽管最近的视频扩散模型为3DGS重建提供了强大的时空先验,但直接对其进行微调以进行重建并不是最优选择,因为它们缺乏对基础多摄像头几何的认知,导致多视图不一致。在本研究中,我们提出了一种新颖的3D感知视频重建框架,旨在提高稀疏3DGS重建的质量。具体而言,我们构建了一个大规模的3DGS视频数据集,以便进行专业的微调。为了弥合2D视频生成与3D多视图约束之间的差距,我们引入了一种基于摄像机条件的几何先验。通过使用第一帧和最后一帧作为边界锚点,并编码相应的摄像机关系,我们将空间结构显式地注入到视频生成管道中。这种基于边界锚定的、感知摄像机的先验引导网络朝向几何基础的重建,确保在不同视点之间保持一致性。大量实验表明,在视频先验重建方法中,我们的方法在像素和结构级别的保真度(PSNR/SSIM)方面达到了最佳,同时改善了多视图一致性,并在感知质量(LPIPS)上保持竞争力。
cs.CV / 47 / 2608.21854

Frame-Level Evaluation in Weakly Supervised Video Anomaly Detection Mostly Measures Video-Level Ranking

弱监督视频异常检测中的帧级评估主要测量视频级排名
Song, Inpyo, Lee, Jangwon
Abstract
Weakly supervised video anomaly detectors are trained with video-level labels but are commonly evaluated as temporal localizers using Micro-AUROC or AP over pooled test frames. Because these metrics compare frames from different videos, a detector can score well by separating videos without accurately ordering moments within them. We exactly decompose Micro-AUROC by video identity into Within-AUROC for temporal ordering within videos and Cross-AUROC for comparisons across videos. Across ShanghaiTech, XD-Violence, and UCF-Crime, only 0.071-0.388% of comparisons between anomalous and normal frames occur within the same video. When both classes remain distributed across V videos, this share decreases as O(1/V), a benchmark property we call temporal dilution. We train anomaly video binary classifiers under the same video-level supervision and repeat each video score across all frames. These video-constant outputs reach 81.40-97.18 Micro-AUROC despite having no within-video variation. Across 72 controlled runs, replacing every frame score with its video mean preserves a median 98.6% of the Micro-AUROC margin above chance. The same empirical pattern holds for author-released outputs and for XD-Violence under its official AP evaluation. A detector can therefore achieve a high pooled score even when it assigns the same score to every moment within each video.
Chinese Translation
弱监督视频异常检测器使用视频级标签进行训练,但通常作为时间定位器进行评估,使用Micro-AUROC或在汇总测试帧上计算的AP。由于这些指标比较来自不同视频的帧,因此检测器可以通过分离视频而在其中未能准确排序时获得良好的评分。我们准确地将Micro-AUROC按视频身份分解为Within-AUROC(用于视频内的时间排序)和Cross-AUROC(用于视频间的比较)。在ShanghaiTech、XD-Violence和UCF-Crime数据集中,仅有0.071-0.388%的异常帧与正常帧的比较发生在同一视频内。当两个类别在V个视频中分布时,这一比例随着O(1/V)的减少而下降,这一基准特性我们称之为时间稀释。我们在相同的视频级监督下训练异常视频二分类器,并在所有帧中重复每个视频的得分。这些视频常量输出在没有视频内变化的情况下达到了81.40-97.18的Micro-AUROC。在72次受控实验中,将每个帧的得分替换为其视频均值保持了高达98.6%的Micro-AUROC边际中位数,超过了随机水平。同样的经验模式适用于作者发布的输出以及在其官方AP评估下的XD-Violence。因此,检测器即使在每个视频内为每个时刻分配相同的得分时,也能获得高的汇总评分。
cs.CV / 48 / 2608.21869

GuardPaint:SpeculativeSafetyDecodingforText-to-ImageGeneration

GuardPaint:用于文本到图像生成的推测安全解码
Dhoot, Shreyash, Dhiman, Paras, Naqvi, Arsh Abbas, Dutta, Aranbi, Chadha, Aman, Jain, Vinija, Das, Amitava
Abstract
Text-to-image (T2I) diffusion models offer powerful visual generation, but their controllability creates a critical safety challenge: adversarial prompts can steer the denoising trajectory toward policy-violating content such as explicit nudity or graphic violence. Existing safeguards mostly act before generation through prompt filtering or after generation through image classification, leaving the diffusion process itself unguarded and often yielding only refusal rather than safe visual repair. We introduce GuardPaint, a speculative decoding framework for safe T2I generation that intervenes inside the diffusion trajectory without modifying the base model. A lightweight auditor monitors intermediate images, localizes unsafe regions, and triggers surgical inpainting repair only where needed. Candidate repairs are generated by a policy-aligned inpainter and selected through a guarded tournament that accepts edits only when they improve policy compliance while preserving prompt fidelity and perceptual quality. Across five jailbreak families SneakPrompt, MMA, PGJ, DACA, and RABell and UNet/flow-matching models including SD~1.5, SDXL, SD~3.5, and FLUX.1-dev. GuardPaint reduces attack success and harmful generations with minimal degradation to image quality, prompt fidelity, and benign behavior. Content warning: This paper contains examples involving nudity and violence that some readers may find disturbing, distressing, or offensive.
Chinese Translation
文本到图像(T2I)扩散模型提供了强大的视觉生成能力,但其可控性带来了一个关键的安全挑战:对抗性提示可以将去噪轨迹引导至违反政策的内容,例如露骨的裸体或血腥暴力。现有的安全措施主要在生成之前通过提示过滤或在生成之后通过图像分类来实施,这使得扩散过程本身未受到保护,通常只能拒绝而无法进行安全的视觉修复。我们提出了GuardPaint,这是一种用于安全T2I生成的推测解码框架,它在不修改基础模型的情况下干预扩散轨迹。一个轻量级审计器监控中间图像,定位不安全区域,并仅在必要时触发精确的修复。候选修复由与政策对齐的修复器生成,并通过一个受保护的比赛进行选择,只有在改善政策合规性同时保持提示忠实度和感知质量的情况下,才接受编辑。在五个越狱家族SneakPrompt、MMA、PGJ、DACA和RABell以及包括SD~1.5、SDXL、SD~3.5和FLUX.1-dev在内的UNet/流匹配模型中,GuardPaint在对图像质量、提示忠实度和良性行为造成最小降级的情况下,减少了攻击成功率和有害生成。内容警告:本文包含涉及裸体和暴力的示例,可能会让某些读者感到不安、痛苦或冒犯。
cs.CV / 49 / 2608.21878

ViSMoE: Visual-Aware Sparse Mixture-of-Experts for Embodied Referring Expression Grounding

ViSMoE:面向视觉的稀疏专家混合模型用于具身指称表达定位
Feng, Shuo, Li, Piji
Abstract
Embodied Referring Expression Grounding is the task of enabling an agent to navigate in real environments and to localize a remote object based on natural language instructions. In this scenario, the agent needs to select one view for navigation at each step and identify a specific object among all candidate objects at the destination. However, most of the previous approaches fail to distinguish between views and objects, instead processing them using the vanilla vision encoder, which results in ambiguous representations of both views and objects. To address the above issues, we propose ViSMoE, which equips sparse Mixture-of-Experts with a visual-aware routing policy for the embodied agent. This framework processes different types of visual information specifically, resulting in discriminative visual representations for both views and objects. Experimental results on REVERIE and SOON datasets demonstrate that ViSMoE outperforms the previous state-of-the-art methods, showing the superiority of our proposed method.
Chinese Translation
具身指称表达定位是使代理能够在真实环境中导航并根据自然语言指令定位远程物体的任务。在这种情况下,代理需要在每一步选择一个视图进行导航,并在目的地的所有候选物体中识别特定物体。然而,大多数先前的方法未能区分视图和物体,而是使用普通的视觉编码器进行处理,这导致视图和物体的表示模糊。为了解决上述问题,我们提出了ViSMoE,该模型为具身代理配备了具有视觉感知的稀疏专家混合模型路由策略。该框架专门处理不同类型的视觉信息,从而为视图和物体生成具有区分性的视觉表示。在REVERIE和SOON数据集上的实验结果表明,ViSMoE优于先前的最先进方法,展示了我们提出的方法的优越性。
cs.CV / 50 / 2608.21881

Region-Weighted Losses and Model Fusion for Cross-Modal PET Attenuation Correction

区域加权损失与模型融合用于跨模态PET衰减校正
Nguyen, Khoa Tuan, Vankerschaver, Joris, De Neve, Wesley
Abstract
We describe our approach to the Big Cross-Modal Attenuation Correction (BIC-MAC) challenge, which asks for a pseudo-CT in Hounsfield Units to be synthesized from Non-Attenuation-Corrected PET (NAC-PET), DIXON MRI and a topogram, and scores both the pseudo-CT and the Attenuation-Corrected PET (AC-PET) reconstructed from it. Three ideas carried our improvements over the organizers' 3D U-Net baseline. The loss matters more than the architecture: we compute the $L_1$ error in the Carney attenuation-coefficient ($\mu$) space that the CT metric itself uses, weighted by anatomical region. Only once that loss was in place did the unregistered DIXON MRI work as extra input channels. A fixed convex combination of two independently trained models then beat both of its members on three of the four metrics and ranks first overall on the public validation leaderboard.
Chinese Translation
我们描述了我们在大规模跨模态衰减校正(BIC-MAC)挑战中的方法,该挑战要求从非衰减校正的PET(NAC-PET)、DIXON MRI和一幅顶图合成出以亨斯菲尔德单位表示的伪CT,并对伪CT及其重建的衰减校正PET(AC-PET)进行评分。我们的改进基于三个思路,超越了组织者提供的3D U-Net基线。损失函数比架构更为重要:我们在CT指标所使用的Carney衰减系数($6 ext{μ}$)空间中计算加权的$L_1$误差,权重由解剖区域决定。只有在该损失函数建立后,未配准的DIXON MRI才作为额外输入通道发挥作用。然后,一个固定的凸组合由两个独立训练的模型组成,在四个指标中的三个上超越了这两个模型,并在公共验证排行榜上排名第一。
cs.CV / 51 / 2608.21883

VIG: Visual Information Gain as a Reward Signal for Multimodal Chain-of-Thought Compression

VIG:作为多模态思维链压缩奖励信号的视觉信息增益
Luo, Wen, Yi, Xiaohan, Huang, Xiaotao, Huang, Liqun
Abstract
Multimodal large reasoning models often rely on long Chain-of-Thought (CoT) traces in which a substantial fraction of tokens, such as repeated visual descriptions, self-reflection, and other visually-disengaged filler, inflate inference cost without contributing to the answer. Existing CoT compression methods optimize output length but never measure whether a reasoning token is actually grounded in the image. We propose \textbf{VIG} (Visual Information Gain), an information-theoretic GRPO reward that scores each reasoning token by how much the image reduces its predictive uncertainty. VIG is computed online from two forward passes of the same policy, one with and one without the image, so no reference chains, external annotations, or auxiliary reward models are needed. Across six main multimodal reasoning benchmarks and three Qwen3-VL-Thinking model sizes (2B/4B/8B), plus an additional R1-Onevision-Bench evaluation on 8B, VIG consistently improves the accuracy--efficiency trade-off, supporting our central claim: \emph{efficient multimodal reasoning emerges from raising visual information density, where every reasoning token earns its place by anchoring to the image, rather than from imposing a length budget.} Our source code is available at https://github.com/chaser682/vig.
Chinese Translation
多模态大型推理模型通常依赖于长的思维链(Chain-of-Thought, CoT)轨迹,其中相当一部分令牌(如重复的视觉描述、自我反思和其他与视觉无关的填充内容)增加了推理成本,却未对答案作出贡献。现有的 CoT 压缩方法优化输出长度,但从未衡量推理令牌是否真正与图像相关。我们提出了 extbf{VIG}(视觉信息增益),这是一种信息论的 GRPO 奖励,通过图像减少预测不确定性来为每个推理令牌打分。VIG 是通过同一策略的两次前向传递在线计算的,一次包含图像,一次不包含,因此不需要参考链、外部注释或辅助奖励模型。在六个主要的多模态推理基准和三个 Qwen3-VL-Thinking 模型规模(2B/4B/8B)上,以及在 8B 上的额外 R1-Onevision-Bench 评估中,VIG 一贯改善了准确性与效率的权衡,支持我们的核心论点: extit{高效的多模态推理源于提高视觉信息密度,每个推理令牌通过与图像的锚定来获得其位置,而不是通过施加长度预算。}我们的源代码可在 https://github.com/chaser682/vig 获取。
cs.CV / 52 / 2608.21885

Pixel-Space Diffusion via Observation Operators

通过观测算子进行像素空间扩散
Guo, Shaojie, Ma, Lichen, Tong, Haoyang, He, Yu, Guo, Zipeng, Liu, Xiaoan, Yan, Feng, Guo, Yu, Wang, Fei, Huang, Junshi, Wang, Yan
Abstract
Pixel-space diffusion models directly model image distributions but remain difficult to optimize. Recent methods alleviate this challenge through target reparameterization, while still relying on a fixed clean-image target throughout denoising. Through empirical analysis, we identify a scale-time mismatch: image structures become predictable from coarse to fine as noise decreases, whereas existing models are forced to predict the full image even under high noise, resulting in low-SNR gradients that hinder optimization. To resolve this mismatch, we propose Observation Operator Diffusion, a unified framework that aligns both the supervision trajectory and feature refinement with the intrinsic recovery order of image structures. Specifically, we replace fixed full-image supervision along the standard flow path with a time-indexed observation trajectory that evolves from coarse structures to the full image during denoising. This trajectory is instantiated with a family of Gaussian-Lanczos operators at varying observation scales, yielding a path-consistent training objective. We further introduce GL-CoDA, a decoder that injects scale-specific Gaussian-Lanczos observations across decoding stages for coarse-to-fine feature refinement. Extensive experiments show that the proposed approach converges substantially faster while consistently improving generation quality, achieving an FID of 1.52 on ImageNet-256.
Chinese Translation
像素空间扩散模型直接建模图像分布,但优化仍然困难。近期的方法通过目标重参数化缓解了这一挑战,但在去噪过程中仍依赖于固定的干净图像目标。通过实证分析,我们识别出一个尺度-时间不匹配的问题:随着噪声的减少,图像结构从粗到细变得可预测,而现有模型被迫在高噪声下预测完整图像,导致低信噪比(SNR)梯度,妨碍了优化。为了解决这一不匹配,我们提出了观测算子扩散(Observation Operator Diffusion),这是一个统一框架,旨在将监督轨迹和特征细化与图像结构的内在恢复顺序对齐。具体而言,我们用一个时间索引的观测轨迹替代了标准流路径中的固定全图监督,该轨迹在去噪过程中从粗糙结构演变到完整图像。该轨迹通过在不同观测尺度下使用一系列高斯-兰茨(Gaussian-Lanczos)算子来实现,从而产生一个路径一致的训练目标。我们进一步引入了GL-CoDA,一个解码器,在解码阶段注入尺度特定的高斯-兰茨观测,以实现从粗到细的特征细化。大量实验表明,所提方法收敛速度显著更快,同时持续提高生成质量,在ImageNet-256上达到了1.52的FID。
cs.CV / 53 / 2608.21893

A Scalable Vector Graphics Latent Space

可扩展矢量图形的潜在空间
Zini, Leonardo, Frigieri, Elia, Baraldi, Lorenzo
Abstract
Scalable Vector Graphics are a fundamental medium for resolution-independent visual content, yet the deep learning community lacks a continuous, dense, and invertible latent space for vector representations, the kind of foundational building block that Variational Autoencoders and their descendants have long provided for raster images. We introduce SLS (SVG Latent Space), a Transformer-based autoencoder that learns compact dense representations of individual SVG paths, the atomic visual elements from which any SVG image can be composed. By modeling SVG commands, coordinate data, and visual properties within a unified BPE-based token vocabulary, SLS learns fixed-size latent representations that jointly capture structure and appearance, and can be decoded back into valid, style-consistent SVG paths with high fidelity. The resulting embedding space is robust, invertible, and structured: embeddings lie on a unit hypersphere, enabling efficient similarity search, composition, and downstream conditioning through simple vector-space operations. Finally, we demonstrate that SLS generalizes across diverse tasks reducing their FLOPs by over 150 times compared to token-based approaches, and establishing a general-purpose latent foundation for vector graphics research.
Chinese Translation
可扩展矢量图形(Scalable Vector Graphics,SVG)是实现分辨率无关视觉内容的基础媒介,然而深度学习领域尚缺乏针对矢量表示的连续、稠密且可逆的潜在空间,这类基础构件长期以来已被变分自编码器(Variational Autoencoders)及其衍生模型为光栅图像所提供。本文提出了SLS(SVG Latent Space),一种基于Transformer的自编码器,能够学习单个SVG路径的紧凑稠密表示,SVG路径是构成任意SVG图像的原子视觉元素。通过在统一的基于BPE(Byte Pair Encoding)的令牌词汇表中建模SVG命令、坐标数据及视觉属性,SLS学习到固定尺寸的潜在表示,能够联合捕捉结构与外观特征,并能高保真地解码回有效且风格一致的SVG路径。所得嵌入空间具有鲁棒性、可逆性和结构化特性:嵌入分布于单位超球面上,支持通过简单的向量空间操作实现高效的相似性搜索、组合及下游条件控制。最后,我们展示了SLS在多样化任务上的泛化能力,其计算量(FLOPs)相比基于令牌的方法降低了150倍以上,确立了面向矢量图形研究的通用潜在基础。
cs.CV / 54 / 2608.21913

Entity-Constrained CBCT Retrieval for Low-Resource Dental Record Completion

基于实体约束的CBCT检索用于低资源牙科病历补全
Nguyen, Nhi Ngoc-Yen, Nguyen, Thai, Tuan, Kiet Huynh Cao, Pham, Huy-Hieu
Abstract
Completing dental records from cone-beam computed tomography (CBCT) is difficult when annotation is scarce and individual clinical fields are supported by different types of evidence. MMDental Task 3 requires seven-field record completion from only 50 labeled CBCT cases and scores the correctness of structured FDI positions and ICD codes; consequently, a visually plausible retrieved record can still be harmful when it introduces an unsupported entity. We propose Entity-Constrained CBCT-Guided Retrieval (ECCR), a parameter-free framework that separates evidence availability from evidence authority. A corpus-derived prior first supplies the complete record. A frozen 3D encoder retrieves image-conditioned Diagnosis evidence, which is appended only if it does not expand the prior FDI or ICD entity set, so the asserted entity set is invariant by construction. On public validation, ECCR reaches a weighted score of 0.3134, improving on both full-record multimodal retrieval (0.2237) and a static text-only prior (0.2915); the guard blocks 63.3% of retrieved candidates, each of which would otherwise have injected an FDI position or ICD code absent from the prior. On the final test evaluation, ECCR obtains 11.37 of a 97.4-point attainable maximum, securing second place overall. The result indicates that, in an extreme low-resource setting, controlling what multimodal evidence is allowed to modify can be more reliable than transferring an entire retrieved record.
Chinese Translation
在标注稀缺且各临床领域依赖不同类型证据的情况下,从锥束计算机断层扫描(CBCT)完成牙科病历具有较大难度。MMDental任务3要求仅通过50例带标签的CBCT病例完成七字段病历补全,并对结构化的FDI位置和ICD编码的正确性进行评分;因此,即使检索到的病历在视觉上合理,若引入了不支持的实体,仍可能产生负面影响。本文提出了一种参数无关的框架——基于实体约束的CBCT引导检索(Entity-Constrained CBCT-Guided Retrieval,ECCR),该框架将证据的可用性与证据的权威性分离。首先,利用语料库导出的先验提供完整病历;随后,冻结的三维编码器检索基于图像的诊断证据,且仅当该证据不会扩展先验中的FDI或ICD实体集合时才予以附加,从而确保所断言的实体集合在构造上保持不变。在公开验证集上,ECCR达到加权得分0.3134,优于全病历多模态检索(0.2237)和静态文本先验(0.2915);该机制阻止了63.3%的检索候选项,这些候选项本会引入先验中不存在的FDI位置或ICD编码。在最终测试评估中,ECCR获得了97.4分满分中的11.37分,整体排名第二。结果表明,在极端低资源环境下,控制允许哪些多模态证据进行修改比直接迁移整个检索病历更为可靠。
cs.CV / 55 / 2608.21926

AirAlign: Geometry-Aware Relative Pose Alignment for UAV Last-Meter Navigation

AirAlign:面向几何感知的无人机最后一米导航相对姿态对齐
Zhou, Jinyi, Feng, Shuo, Wu, Yufei, Li, Piji
Abstract
Unmanned aerial vehicle (UAV) navigation in modern low-altitude environments requires more accurate pose alignment in the final approach stage for target information acquisition or manipulation, making "last-meter" navigation increasingly important. However, severe viewpoint and appearance variations make this task challenging. To tackle this problem, we propose AirAlign, a framework for RGB-only image-pair relative pose alignment for UAVs. AirAlign uses a pretrained visual geometry reconstruction model as the backbone to extract geometry-aware features from source-target image pairs. In addition, to better utilize the limited training data, we split the training set into multiple scene-disjoint folds for unseen cross-validation and model selection. During inference, the predictions of the selected models are averaged to form the ensemble output of the overall framework. Experiments on the PairUAV challenge at the ACMMM 2026 Workshop on UAVs in Multimedia demonstrate the effectiveness and robustness of our method, while comprehensive ablation studies validate the contribution of each component.
Chinese Translation
在现代低空环境中,无人机(UAV)导航在最终接近阶段需要更准确的姿态对齐,以便获取或操作目标信息,使得“最后一米”导航变得越来越重要。然而,严重的视角和外观变化使得这一任务具有挑战性。为了解决这个问题,我们提出了AirAlign,一个针对无人机的仅基于RGB图像对的相对姿态对齐框架。AirAlign使用预训练的视觉几何重建模型作为骨干网络,从源-目标图像对中提取几何感知特征。此外,为了更好地利用有限的训练数据,我们将训练集划分为多个场景不重叠的折叠,以进行未见交叉验证和模型选择。在推理过程中,选定模型的预测结果被平均,以形成整体框架的集成输出。在2026年ACMMM无人机多媒体研讨会的PairUAV挑战赛上的实验验证了我们方法的有效性和鲁棒性,而全面的消融研究则验证了每个组件的贡献。
cs.CV / 56 / 2608.21937

C$^2$Path: Class-Conditional Pathway Decoupling for Vision-Language Incremental Object Detection

C$^2$Path:面向视觉-语言增量目标检测的类别条件路径解耦
Xu, Lecheng, Shao, Feifei, Ye, Ouyangzi, Wang, Zhen, Li, Lin, Li, Kexin, Wang, Zhao, Huang, Changqin
Abstract
Incremental Object Detection (IOD) aims to enable detectors to continuously learn novel categories while preserving previously acquired knowledge. However, existing methods suffer from two forms of \textbf{class knowledge coupling}: class boundary erosion induced by shared parameter updates and class representation entanglement arising from mixed feature encoding. We argue that effective incremental learning requires class-specific computational pathways that enable isolated parameter updates and separated class-wise injection. To this end, we propose \textbf{C$^2$Path}, a class-conditional pathway decoupling framework for vision-language incremental object detection that leverages token-level class cues to establish dedicated and updatable computational pathways for different categories. Specifically, C$^2$Path introduces a category expert library and a class-conditional decoupling module. The expert library consists of learnable low-rank computational nodes that capture category-specific knowledge, while the decoupling module generates class-aware routing signals to dynamically compose \textit{ClassLoRA} adapters from these experts, thereby forming class-specific computational pathways for isolated updates and separated injection across categories. Extensive experiments on COCO 2017 under multiple incremental learning settings demonstrate that C$^2$Path consistently outperforms state-of-the-art methods, providing an effective and scalable solution for continual category expansion in vision-language detectors.
Chinese Translation
增量目标检测(IOD)旨在使检测器能够在保留先前获得知识的同时持续学习新类别。然而,现有方法存在两种形式的 extbf{类别知识耦合}:由共享参数更新引起的类别边界侵蚀和由混合特征编码引起的类别表示纠缠。我们认为,有效的增量学习需要类别特定的计算路径,以实现独立的参数更新和分离的类别注入。为此,我们提出了 extbf{C$^2$Path},一个面向视觉-语言增量目标检测的类别条件路径解耦框架,该框架利用标记级类别线索为不同类别建立专用且可更新的计算路径。具体而言,C$^2$Path引入了一个类别专家库和一个类别条件解耦模块。专家库由可学习的低秩计算节点组成,捕捉类别特定知识,而解耦模块生成类别感知的路由信号,以动态组合来自这些专家的 extit{ClassLoRA}适配器,从而形成类别特定的计算路径,以实现跨类别的独立更新和分离注入。在多个增量学习设置下对COCO 2017的广泛实验表明,C$^2$Path始终优于最先进的方法,为视觉-语言检测器中的持续类别扩展提供了有效且可扩展的解决方案。
cs.CV / 57 / 2608.21948

Sparse Multi-Stage Expert-Agent Routing for Complex Clinical Reasoning

复杂临床推理的稀疏多阶段专家-代理路由
Xiang, Sike, Chen, Shuang, sun, Qian, Cheng, Jia, Wei, Yusi, Atapour-Abarghouei, Amir
Abstract
Complex clinical reasoning requires models to update diagnostic hypotheses as new evidence emerges and to coordinate different medical specialities under limited consultation resources. Existing LLM-based clinical reasoning systems typically perform single-pass prediction or rely on fixed multi-agent workflows, making expert participation either static or unnecessarily exhaustive. We propose Sparse Multi-Stage Expert-Agent Routing, a language-based clinical reasoning framework that models diagnosis as a stage-wise routing process. Given progressively available clinical evidence derived from multiple modalities, the framework maintains an evolving case state and adaptively activates a sparse set of medical expert agents, supported by expert-specific memory across stages. To evaluate free-text diagnostic conclusions beyond surface similarity, we further introduce ClinFEScore, a fact-aware semantic evaluation protocol for clinical reasoning outputs. On reconstructed multi-stage cases from MAC and AgentClinic-NEJM, our framework reduces the average number of activated experts from 17.0 to 3.0 whilst maintaining strong fact-level diagnostic quality. On 200 real-world hospital MDT cases, ClinFEScore correlates strongly with clinician judgements (Spearman's $\rho=0.81$; Pearson's $r=0.87$), whilst our method achieves 91.5\% clinician-verified diagnostic accuracy with approximately five expert-agent/LLM calls per case. These results support sparse stage-wise coordination as an efficient and clinically relevant approach to LLM-based clinical reasoning.
Chinese Translation
复杂的临床推理要求模型在新证据出现时更新诊断假设,并在有限的咨询资源下协调不同的医学专业。现有的基于大型语言模型(LLM)的临床推理系统通常执行单次预测或依赖固定的多代理工作流程,使得专家参与要么是静态的,要么是不必要的冗长。我们提出了稀疏多阶段专家-代理路由,这是一种基于语言的临床推理框架,将诊断建模为阶段性路由过程。该框架根据来自多种模态的逐步可用的临床证据,维护不断演变的案例状态,并自适应地激活一组稀疏的医学专家代理,支持跨阶段的专家特定记忆。为了评估自由文本诊断结论超越表面相似性,我们进一步引入了ClinFEScore,这是一种关注事实的语义评估协议,用于临床推理输出。在从MAC和AgentClinic-NEJM重建的多阶段案例中,我们的框架将激活的专家平均数量从17.0减少到3.0,同时保持强大的事实级诊断质量。在200个真实世界医院多学科团队(MDT)案例中,ClinFEScore与临床医生的判断高度相关(Spearman's $ ho=0.81$;Pearson's $r=0.87$),而我们的方法在每个案例中实现了91.5\%的临床医生验证的诊断准确率,平均每个案例约有五次专家代理/LLM调用。这些结果支持稀疏阶段性协调作为一种高效且与临床相关的LLM基础临床推理方法。
cs.CV / 58 / 2608.21967

Trustworthy Visual Quality Inspection under Data Scarcity in Manufacturing

在制造业数据稀缺下的可信视觉质量检测
Sapoutzoglou, Panagiotis, Ribaira, Jessy, Kanounnikoff, Martin, Tijsma, Bas, Geiß, Christian, Pateraki, Maria
Abstract
Automated visual inspection in manufacturing aims to replace slow and inconsistent manual checks, but its economic value depends on whether its decisions can be trusted enough to automate routine inspection while reserving human expertise for ambiguous cases. In production-line settings, defective samples are scarce, since the process is optimized to produce good parts, which limits any learning-based inspector trained on real data alone. Compounding this, defect decisions emitted as hard labels with no confidence estimate carry an asymmetric cost: a false reject wastes a good product, while a false accept may increase the risk of undetected defects progressing through the production process. We address both problems by mitigating data scarcity through the generation of synthetic defective samples with a diffusion model, and meeting the need for confidence-aware decisions with a Bayesian classifier that defers ambiguous units to human review rather than misclassifying them. These components are embedded in a staged pipeline of successive, complementary checks. We evaluate how synthetic augmentation affects classification and localization on a test set of real defects, and examine the system's trustworthiness at three points: the decision, the synthetic data, and the pipeline structure. This work-in-progress reports preliminary results suggesting that diffusion-generated defects, combined with uncertainty-aware classification, can lower the cost of reaching a trustworthy, deployable inspection model under data scarcity.
Chinese Translation
制造业中的自动视觉检测旨在取代缓慢且不一致的人工检查,但其经济价值依赖于其决策是否足够可信,以便在保留人类专业知识用于模糊案例的同时实现常规检查的自动化。在生产线环境中,缺陷样本稀少,因为该过程被优化以生产良品,这限制了仅基于真实数据训练的学习型检测器的能力。更为复杂的是,以硬标签形式发出的缺陷决策没有置信度估计,带来了不对称的成本:错误拒绝会浪费良好的产品,而错误接受则可能增加未被发现的缺陷在生产过程中继续存在的风险。我们通过使用扩散模型生成合成缺陷样本来缓解数据稀缺问题,并通过贝叶斯分类器满足对置信度感知决策的需求,将模糊单元交由人工审核,而不是错误分类。这些组件嵌入在一个分阶段的管道中,进行连续的互补检查。我们评估合成增强如何影响真实缺陷测试集上的分类和定位,并在三个方面检验系统的可信度:决策、合成数据和管道结构。这项正在进行的工作报告了初步结果,表明扩散生成的缺陷结合不确定性感知分类可以降低在数据稀缺情况下实现可信、可部署的检测模型的成本。
cs.CV / 59 / 2608.21970

Retrieval-Augmented Visual Prompting: Guiding Foundation Models in Two-Photon Imaging

检索增强视觉提示:引导基础模型在双光子成像中的应用
Calcagno, Salvatore, Finocchiaro, Marco, Bellitto, Giovanni, Giordano, Daniela, Spampinato, Concetto, Salanitri, Federica Proietto
Abstract
Two-photon calcium imaging presents a challenging setting for foundation models: image appearance varies substantially across recordings and experimental conditions, annotations are scarce, and rapid adaptation is often needed. Rather than adapting model weights through fine-tuning, we ask whether a foundation model can be guided at inference time by injecting external visual memory directly into its input. We implement this idea with SAM 3 and introduce Retrieval-Augmented Visual Prompting (RAVP), a framework in which each target tile is augmented with a retrieved annotated exemplar whose bounding box is used as a concept prompt. RAVP turns retrieval into a form of visual prompting and enables adaptation through input design alone. We study multiple exemplar selection strategies, including fluorescence-guided heuristics and a lightweight recall predictor trained to estimate which exemplar is most informative for a target tile. Experiments on the Allen Brain Observatory show that exemplar-augmented inference consistently strengthens zero-shot neuron detection and instance segmentation. Ablation studies further show that a single carefully selected exemplar is more effective than prompting with multiple retrieved examples. These results position inference-time visual memory injection as a simple and effective alternative to parameter adaptation for foundation models in specialized biomedical imaging.
Chinese Translation
双光子钙成像为基础模型提供了一个具有挑战性的环境:图像外观在不同记录和实验条件下变化显著,注释稀缺,且通常需要快速适应。我们提出的问题是,是否可以通过在推理时将外部视觉记忆直接注入模型输入来引导基础模型,而不是通过微调来调整模型权重。我们使用 SAM 3 实现了这一想法,并引入了检索增强视觉提示(RAVP)框架,其中每个目标图块都通过检索到的带注释示例进行增强,其边界框用作概念提示。RAVP 将检索转化为一种视觉提示形式,并仅通过输入设计实现适应。我们研究了多种示例选择策略,包括荧光引导启发式方法和一个轻量级回忆预测器,该预测器经过训练以估计哪个示例对目标图块最具信息量。在艾伦脑观察所的实验表明,增强示例的推理在零样本神经元检测和实例分割中始终增强了效果。消融研究进一步表明,单个精心选择的示例比使用多个检索示例进行提示更为有效。这些结果将推理时的视觉记忆注入定位为基础模型在专业生物医学成像中参数适应的简单有效替代方案。
cs.CV / 60 / 2608.21972

Improved denoising diffusion probabilistic models with efficient non-diagonal covariance modeling

改进的去噪扩散概率模型与高效的非对角协方差建模
Xia, Rui, Das, Ayan, Artemev, Artem, Zhang, Andi, Hennequin, Guillaume, Bernacchia, Alberto
Abstract
The sampling process of Denoising Diffusion Probabilistic Models (DDPMs) can be accelerated by leveraging second-order information in the form of approximations to the denoising posterior covariance -- allowing samples of acceptable quality to be produced in fewer but larger sampling steps. Previous attempts at using such information have used drastic (e.g.\ diagonal) simplifications of the covariance. These do not do justice to the peculiar statistical structure of natural images, which exhibit strong non-diagonal correlations between pixels and color channels, and a slow-decaying power-law frequency spectrum. Here, we develop a novel covariance model that captures these features. Our Kronecker-DCT (K-DCT) model uses a Kronecker-factored decomposition of inter-color covariances and spatial covariances modeled in the frequency domain using the Discrete Cosine Transform (DCT). The use of the DCT reduces the computational complexity from quadratic to log-linear, resulting in negligible computational and memory overhead in each denoising step. By learning K-DCT-structured amortizations of the denoising posterior covariance using pre-trained score models on CIFAR-10, Celeb-A, ImageNet and LSUN datasets, we show improved performance compared to previous SOTA denoising samplers, both in terms of FID and likelihoods, especially in the regime of few denoising steps.
Chinese Translation
去噪扩散概率模型(Denoising Diffusion Probabilistic Models, DDPMs)的采样过程可以通过利用二阶信息(以对去噪后验协方差的近似形式)来加速,从而在更少但更大的采样步骤中生成可接受质量的样本。之前尝试使用此类信息的方法采用了剧烈的(例如,对角)简化协方差。这些简化未能充分反映自然图像的特殊统计结构,后者在像素和颜色通道之间表现出强烈的非对角相关性,以及缓慢衰减的幂律频谱。在此,我们开发了一种新颖的协方差模型,以捕捉这些特征。我们的 Kronecker-DCT(K-DCT)模型使用了对颜色间协方差和在频域中使用离散余弦变换(DCT)建模的空间协方差的 Kronecker 分解。使用 DCT 将计算复杂度从二次降低到对数线性,导致每个去噪步骤中的计算和内存开销微乎其微。通过使用在 CIFAR-10、Celeb-A、ImageNet 和 LSUN 数据集上预训练的评分模型学习 K-DCT 结构的去噪后验协方差的摊销,我们展示了与之前的最先进(SOTA)去噪采样器相比,在 FID 和似然性方面的性能提升,尤其是在少量去噪步骤的情况下。
cs.CV / 61 / 2608.22003

Close Shortcut Wins Long: Seeking Diverse and Stable Generators for Data-Free Knowledge Distillation

缩短捷径赢得长远:寻求多样化和稳定的无数据知识蒸馏生成器
Lyu, Kailin, Zhang, Zherui, Dong, Junhao, Fu, Kexue, Pang, Weiguang, Xu, Rongtao, Wang, Qizheng, Wu, Di, Kwoh, Chee-Keong, Gao, Longxiang, Xu, Shibiao, Wang, Changwei, Hao, Ce, Zhang, Yu
Abstract
Data-Free Knowledge Distillation (DFKD) preserves privacy by transferring knowledge without real data access. However, existing generator-based DFKD methods suffer from over-reliance on teacher preferences and pattern collapse, exhibiting "generative shortcut learning" in the frequency domain: dependent on specific frequency components and frequency positions, resulting in inconsistent synthetic image quality and class diversity. In this paper, we propose a CSWL framework aimed at introducing insights from the frequency domain perspective to improve generator diversity and training stability to Close the phenomenon of Shortcut learning to Win in the Longer term. To address the issue of generative shortcut learning, we introduce frequency-domain augmentation at the feature level, encouraging the generator to attend to the full frequency spectrum and thereby suppress shortcut learning behavior. To tackle training instability, we propose a Cross-Stage Frequency Reconstruction (CSFR) auxiliary task, which implicitly constructs an Exponential Moving Average (EMA) mechanism to promote long-term optimization and stability. Extensive experiments, including downstream tasks and various image recognition datasets at multiple resolutions, validate the effectiveness of CSWL in improving both diversity and stability from the frequency view.
Chinese Translation
无数据知识蒸馏(Data-Free Knowledge Distillation, DFKD)通过在没有真实数据访问的情况下转移知识来保护隐私。然而,现有的基于生成器的DFKD方法过于依赖教师偏好和模式崩溃,在频域中表现出“生成捷径学习”:依赖于特定的频率成分和频率位置,导致合成图像质量和类别多样性不一致。本文提出了一个CSWL框架,旨在从频域视角引入见解,以提高生成器的多样性和训练稳定性,从而缩短捷径学习现象,以赢得长期效果。为了解决生成捷径学习的问题,我们在特征层面引入频域增强,鼓励生成器关注完整的频谱,从而抑制捷径学习行为。为了解决训练不稳定性,我们提出了一种跨阶段频率重构(Cross-Stage Frequency Reconstruction, CSFR)辅助任务,隐式构建了一个指数移动平均(Exponential Moving Average, EMA)机制,以促进长期优化和稳定性。大量实验,包括下游任务和多分辨率的各种图像识别数据集,验证了CSWL在从频率视角提高多样性和稳定性方面的有效性。
cs.CV / 62 / 2608.22039

ORBIT++: Benchmarking SfM in the Wild with 360{\deg} Video

ORBIT++:在复杂环境中基于360°视频的结构光束(SfM)基准测试
Sabour, Sara, Jin, Linyi, Tucker, Richard, Hertz, Amir, Brubaker, Marcus, Saxena, Saurabh, Hur, Junhwa, Tagliasacchi, Andrea, Sun, Deqing, Fleet, David J., Szeliski, Richard, Snavely, Noah
Abstract
Structure-from-Motion (SfM) is a cornerstone of 3D perception, yet current methods often fail when applied to complex videos involving challenging camera motions or dynamic scenes. Compounding the problem, the field lacks reliable ground-truth benchmarks for such difficult scenarios, making it hard to gauge real-world progress or to pinpoint where improvements are most needed. To address this gap, we introduce a new benchmark for evaluating camera pose estimation. Our key insight is to leverage online panoramic 360{\deg} video as a source of data from which to construct challenging clips, while still enabling robust ground-truth trajectory recovery. The panoramic nature of these videos provides richer visual context for tracking camera motion, even when parts of the view are affected by blur, motion, or dynamic objects. After tracking camera motion across full 360{\deg} videos, we crop and reproject selected portions to generate perspective-view clips that serve as our benchmark, called ORBIT. Experiments show that COLMAP, as well as recent optimization-based and feed-forward SfM methods struggle to accurately estimate camera poses on our benchmark. Hence, ORBIT provides a valuable testbed where researchers can meaningfully measure progress on truly challenging, real-world SfM problems.
Chinese Translation
结构光束(SfM)是三维感知的基石,但当前的方法在处理涉及复杂相机运动或动态场景的复杂视频时往往失败。更糟糕的是,该领域缺乏可靠的真实基准来评估这些困难场景,使得评估现实世界进展或确定最需要改进的地方变得困难。为了解决这一问题,我们引入了一种新的基准,用于评估相机姿态估计。我们的关键见解是利用在线全景360°视频作为数据来源,从中构建具有挑战性的剪辑,同时仍然能够实现稳健的真实轨迹恢复。这些视频的全景特性为跟踪相机运动提供了更丰富的视觉上下文,即使视图的某些部分受到模糊、运动或动态物体的影响。在对完整的360°视频进行相机运动跟踪后,我们裁剪并重新投影选定部分,以生成作为我们基准的透视视图剪辑,称为ORBIT。实验表明,COLMAP以及最近的基于优化和前馈的SfM方法在我们的基准上难以准确估计相机姿态。因此,ORBIT提供了一个有价值的测试平台,研究人员可以在此平台上有意义地衡量在真正具有挑战性的现实世界SfM问题上的进展。
cs.CV / 63 / 2608.22054

Robust Global Structure-from-Motion via View Graph Pruning

通过视图图修剪实现鲁棒的全局运动重建
Xu, Jiamin, Yao, Lixing, Dai, Weichen, Gu, Renshu, Zhu, Zunjie, Xu, Weiwei, Xu, Gang
Abstract
Structure-from-Motion (SfM) aims to estimate camera poses and reconstruct 3D structures from a collection of unordered images. Compared with incremental SfM, global SfM achieves better scalability by jointly estimating camera poses based on a view graph constructed from pairwise correspondences. However, its performance is highly sensitive to erroneous edges caused by visually ambiguous matches, which may lead to incorrect camera registration and reconstruction artifacts. In this work, we propose a subgraph-guided view graph pruning framework for robust global SfM. Our key idea is to exploit the internal consistency of reliable subgraphs to identify and remove unreliable connections. Specifically, we first partition the view graph into locally consistent subgraphs and perform global SfM within each subgraph to obtain reliable camera poses. We then apply RANSAC-based edge pruning across subgraphs to remove inconsistent edges, and finally perform global SfM on the refined view graph. Extensive experiments on ambiguous, sequential, and unordered image datasets demonstrate that our method improves the robustness of global SfM under challenging conditions. Further evaluation with neural rendering shows that the improved camera estimation leads to higher-quality novel view synthesis results.
Chinese Translation
运动重建(Structure-from-Motion, SfM)旨在从一组无序图像中估计相机姿态并重建三维结构。与增量式SfM相比,全球SfM通过基于成对对应关系构建的视图图联合估计相机姿态,从而实现更好的可扩展性。然而,它的性能对由视觉模糊匹配引起的错误边缘高度敏感,这可能导致相机配准不正确和重建伪影。在本工作中,我们提出了一种基于子图引导的视图图修剪框架,以实现鲁棒的全局SfM。我们的关键思想是利用可靠子图的内部一致性来识别和删除不可靠的连接。具体而言,我们首先将视图图划分为局部一致的子图,并在每个子图内执行全局SfM以获得可靠的相机姿态。然后,我们在子图之间应用基于RANSAC的边缘修剪,以去除不一致的边缘,最后在精炼的视图图上执行全局SfM。在模糊、顺序和无序图像数据集上的大量实验表明,我们的方法在具有挑战性的条件下提高了全局SfM的鲁棒性。进一步的神经渲染评估显示,改进的相机估计导致了更高质量的新视图合成结果。
cs.CV / 64 / 2608.22064

Competitive Memory Readout for Robust Video Object Segmentation: 2nd Place Technical Report for the MOSEv2 Track of the 8th LSVOS Challenge

鲁棒视频目标分割的竞争记忆读出:第八届LSVOS挑战赛MOSEv2赛道的第二名技术报告
Gao, Mingqi, Li, Sijie, Han, Jungong
Abstract
We present our solution for the MOSEv2 track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge at ECCV 2026. The challenge evaluates robust video object segmentation under complex temporal dynamics, including long-term occlusion, disappearance and reappearance, large appearance changes, and strong interference from visually similar objects. Our method builds on SAM~3 and focuses on its memory readout. Standard target-only memory retrieval can confuse the annotated target with same-class non-target objects because such distractors are represented only implicitly as background. Our method introduces Competitive Memory Readout, which explicitly incorporates same-class competitor evidence when retrieving target information from memory. To prevent excessive suppression of weak or reappearing targets, we further apply a lightweight adaptive restoration rule after competition. The resulting system retains the original SAM~3 tracking pipeline while improving target identity preservation in challenging videos. Our submission achieves 66.20 on the primary challenge score and ranks 2nd in the MOSEv2 track.
Chinese Translation
我们提出了针对2026年ECCV第八届大规模视频目标分割(LSVOS)挑战赛MOSEv2赛道的解决方案。该挑战评估在复杂时间动态下的鲁棒视频目标分割,包括长期遮挡、消失与重现、大幅外观变化以及来自视觉上相似物体的强干扰。我们的方法基于SAM~3,重点关注其记忆读出。标准的仅目标记忆检索可能会将标注目标与同类非目标物体混淆,因为这些干扰物仅作为背景隐式表示。我们的方法引入了竞争记忆读出(Competitive Memory Readout),在从记忆中检索目标信息时明确纳入同类竞争者证据。为了防止对弱或重现目标的过度抑制,我们在竞争后进一步应用了一种轻量级自适应恢复规则。最终的系统保留了原始的SAM~3跟踪流程,同时在具有挑战性的视频中改善了目标身份的保持。我们的提交在主要挑战评分中获得了66.20分,并在MOSEv2赛道中排名第二。
cs.CV / 65 / 2608.22066

ADMIL: Attention-Distilled Multiple Instance Learning for Selective Foundation Model Inference in Pathology

ADMIL:用于病理学选择性基础模型推理的注意力蒸馏多实例学习
Stothers, Duncan, Wu, Ren-Chin, Lotter, William
Abstract
Attention-based multiple instance learning (ABMIL) using pathology foundation model embeddings is effective for slide-level tasks, but exhaustive inference requires applying a large image encoder to every foreground tile despite the subsequent attention distribution often concentrating over a small subset of informative regions. We introduce ADMIL (Attention-Distilled Multiple Instance Learning), a selective-compute framework that distills an ABMIL teacher's attention into a lightweight tile-selection model, PriorNet. Using an EfficientNet architecture, PriorNet learns the teacher attention distribution from raw tile pixels with KL divergence; at inference, it scores the foreground pool, selects the top-K tiles, and invokes the expensive foundation model only on that subset before a selected-bag ABMIL student predicts the slide label. Across BRACS, PANDA, and CAMELYON16, ADMIL matches full-teacher headline performance at K=4, 8, and 128 tiles, respectively, avoiding >98% of foundation model (Virchow2) tile embeddings and model inference FLOPs. Random and teacher-attention oracle controls show that this result depends on task-relevant selection rather than tile-count reduction alone. Quantitative and qualitative analyses suggest that PriorNet recovers the teacher's tile ordering with high fidelity while focusing on task-relevant morphological regions. ADMIL shows that nearly all expensive tile encodings can be removed without sacrificing slide-level performance, providing a potential path for more efficient deployment in clinical settings where latency and compute costs are key considerations.
Chinese Translation
基于注意力的多实例学习(ABMIL)利用病理基础模型嵌入在幻灯片级任务中表现出色,但全面推理需要对每个前景图块应用大型图像编码器,尽管随后的注意力分布通常集中在少量信息丰富的区域上。我们提出了ADMIL(注意力蒸馏多实例学习),这是一种选择性计算框架,将ABMIL教师的注意力蒸馏为轻量级图块选择模型PriorNet。使用EfficientNet架构,PriorNet通过KL散度从原始图块像素中学习教师的注意力分布;在推理时,它对前景池进行评分,选择前K个图块,并仅在该子集上调用昂贵的基础模型,然后由选定的包ABMIL学生预测幻灯片标签。在BRACS、PANDA和CAMELYON16数据集上,ADMIL在K=4、8和128个图块时分别匹配全教师的性能,避免了超过98%的基础模型(Virchow2)图块嵌入和模型推理FLOPs。随机和教师注意力oracle控制显示,这一结果依赖于与任务相关的选择,而不仅仅是图块数量的减少。定量和定性分析表明,PriorNet以高保真度恢复了教师的图块排序,同时关注与任务相关的形态区域。ADMIL表明,几乎所有昂贵的图块编码都可以去除,而不牺牲幻灯片级性能,为在延迟和计算成本是关键考虑因素的临床环境中更高效的部署提供了潜在路径。
cs.CV / 66 / 2608.22072

Spiking Neural Networks for Energy-Efficient Object Detection in Forward-Looking Sonar Imagery

用于前视声纳图像中能效目标检测的脉冲神经网络
Frank, Gwenevere, Cauwenberghs, Gert
Abstract
Autonomous underwater vehicles (AUVs) are increasingly important tools in industries ranging from research, to energy, to defense. AUVs are power-constrained platforms operating in remote environments with fixed battery capacities, where propulsion competes with compute and sensors for power over lengthy mission durations. AUVs frequently operate in dark or turbid waters where optical sensing is of limited value, and rely on sonar as their primary sensing modality. Convolutional neural networks (CNNs) are the state-of-the-art solution for object detection in forward-looking sonar imagery, but are energy expensive (e.g. YOLOv8m: 322 mJ/inference). Spiking neural networks (SNNs) rely on binary spike activations and thus sparse accumulate-only operations, allowing them to be remarkably energy efficient, particularly when paired with dedicated neuromorphic hardware. The sparse, high-contrast structure of forward-looking sonar (FLS) returns is structurally matched to spike coding in a way that optical imagery is not. No prior work has assessed the suitability of SNNs for object detection in FLS imagery. SpikeYOLO, a fully spiking network trained with surrogate gradients, was benchmarked against state-of-the-art CNN baselines on three FLS object detection datasets. Key results: SpikeYOLO T=2 achieves 3.3$\times$ lower theoretical compute energy on UATD (97 vs 322 mJ) at competitive accuracy (0.529 [email protected]:0.95 vs. YOLOv8m's 0.575); SpikeYOLO matches YOLOv8m on [email protected] and outperforms YOLO-SONAR and Fast R-CNN baselines on the sparse Marine-Debris-FLS dataset at 4.4$\times$ lower energy; SpikeYOLO demonstrates superior robustness to multiplicative speckle noise (3.0% degradation at $\sigma{=}0.4$ vs. 8.9% for YOLOv8m), outperforming YOLOv8m outright at $\sigma{=}0.6$, directly relevant to real-world FLS deployment.
Chinese Translation
自主水下航行器(AUV)在从研究到能源再到国防等多个行业中越来越重要。AUV是受电力限制的平台,在固定电池容量的偏远环境中运行,其推进与计算和传感器在长时间任务中争夺电力。AUV通常在黑暗或浑浊的水域中操作,光学传感的价值有限,因此依赖声纳作为其主要传感方式。卷积神经网络(CNN)是前视声纳图像中目标检测的最先进解决方案,但其能耗较高(例如,YOLOv8m: 322 mJ/推理)。脉冲神经网络(SNN)依赖于二进制脉冲激活,因此采用稀疏累积操作,使其在能效上表现出色,尤其是在配备专用神经形态硬件时。前视声纳(FLS)返回的稀疏高对比度结构在结构上与脉冲编码相匹配,而光学图像则不然。此前没有研究评估SNN在FLS图像中进行目标检测的适用性。SpikeYOLO是一个完全脉冲的网络,使用代理梯度进行训练,并在三个FLS目标检测数据集上与最先进的CNN基准进行了对比。关键结果:SpikeYOLO T=2在UATD上实现了3.3倍更低的理论计算能耗(97 vs 322 mJ),且准确率具有竞争力(0.529 [email protected]:0.95 vs. YOLOv8m的0.575);SpikeYOLO在[email protected]上与YOLOv8m相当,并在稀疏的Marine-Debris-FLS数据集上以4.4倍更低的能耗超越YOLO-SONAR和Fast R-CNN基准;SpikeYOLO在乘法散斑噪声下表现出更强的鲁棒性(在σ=0.4时降级3.0% vs. YOLOv8m的8.9%),在σ=0.6时直接超越YOLOv8m,这与实际FLS部署密切相关。
cs.CV / 67 / 2608.22082

When More References Hurt: Contamination-Aware DINOv2 Memory Banks for Few-Shot Steel Defect Detection

更多参考文献反而有害:面向少样本钢铁缺陷检测的污染意识DINOv2记忆库
Kalantari, Hannaneh, Khoramdel, Javad
Abstract
Patch-memory anomaly detectors assume that their reference bank is normal, an assumption that is difficult to guarantee when additional industrial images are unverified. We study whether a few trusted normal images can safely recover useful normal patches from such references without defect masks. Starting from the DINOv2 patch-memory formulation used by AnomalyDINO, we score candidate patches by distance to a clean seed bank, discard the most suspicious 20%, merge the retained patches with the seed, and enforce a fixed budget by greedy coreset selection. On Severstal, naive additional references contain 9.46% anomalous patches; the proposed trim rejects 78.1\% of them and reduces residual contamination to 2.59%. At an equal 51,200-patch development budget, the proposed bank reaches 0.1084 AUPRC versus 0.0950 for naive expansion, 0.0952 for random removal, and 0.1030 for eight clean images. Injecting only 0.5\% anomalous patches into a clean bank reduces AUPRC from 0.1030 to 0.0759. On all five completed held-out pairs, the proposed bank improves over naive expansion, with a mean gain of 0.0142 AUPRC. Reference purity is therefore a first-order design variable, and unverified images are useful only when their contribution is filtered explicitly.
Chinese Translation
补丁记忆异常检测器假设其参考库是正常的,这一假设在额外的工业图像未经验证时难以保证。我们研究了是否可以通过少量可信的正常图像在没有缺陷掩膜的情况下安全地从这些参考中恢复有用的正常补丁。从AnomalyDINO使用的DINOv2补丁记忆公式出发,我们通过与干净种子库的距离对候选补丁进行评分,丢弃最可疑的20%,将保留的补丁与种子合并,并通过贪婪的核心集选择强制执行固定预算。在Severstal数据集上,天真的额外参考包含9.46%的异常补丁;所提出的修剪方法拒绝了78.1%的异常补丁,并将残余污染降低至2.59%。在相同的51,200补丁开发预算下,所提出的记忆库达到了0.1084的AUPRC,而天真的扩展为0.0950,随机移除为0.0952,八张干净图像为0.1030。仅将0.5%的异常补丁注入干净库中,AUPRC从0.1030降至0.0759。在所有五对完成的保留对中,所提出的记忆库相较于天真的扩展有了改善,平均增益为0.0142 AUPRC。因此,参考纯度是一个一阶设计变量,未经验证的图像只有在其贡献被明确过滤时才是有用的。
cs.CV / 68 / 2608.22096

Three-Phase Scribble-Adaptive Curriculum Learning for autoPETV Grand Challenge

三相涂鸦自适应课程学习用于 autoPETV 大挑战
Zhang, Libo
Abstract
This report describes Libo Zhang's algorithmic solution to autoPETV Grand Challenge on interactive lesion segmentation in whole-body PET/CT. Interaction is encoded as two additional input channels that rasterize the accumulated foreground and background scribbles, and a residual-encoder U-Net of about 140 million parameters is trained with a three-phase curriculum over 4000 epochs: the network first learns fully automatic segmentation with silent interaction channels, then observes ground-truth-derived scribbles under randomly sampled visibility modes, and finally adapts to its own mistakes through online simulation of up to five error-driven correction steps. Training draws on 1811 autoPET and DeepPSMA studies, and the submission ensembles the best and final checkpoints of five folds by logit averaging. In interactive five-fold cross-validation with six interaction steps, the final checkpoints reach a mean AUC-Dice of 3.836 and a mean AUC-DMM of 3.869, improving monotonically in every fold, with roughly half of the total gain delivered by the first corrective scribble. Our code and trained model checkpoints are available on https://github.com/Libo1023/autoPETV-Curriculum.
Chinese Translation
本报告描述了 Libo Zhang 针对 autoPETV 大挑战在全身 PET/CT 中进行交互式病灶分割的算法解决方案。交互被编码为两个额外的输入通道,这些通道对累积的前景和背景涂鸦进行光栅化,并且一个约有 1.4 亿参数的残差编码器 U-Net 在 4000 个周期内通过三阶段课程进行训练:网络首先在无交互通道的情况下学习完全自动化的分割,然后在随机采样的可见性模式下观察基于真实值的涂鸦,最后通过在线模拟多达五个基于错误的修正步骤来适应自身的错误。训练基于 1811 个 autoPET 和 DeepPSMA 研究,并且提交结果通过对五个折叠的最佳和最终检查点进行对数平均来集成。在包含六个交互步骤的交互式五折交叉验证中,最终检查点的平均 AUC-Dice 达到 3.836,平均 AUC-DMM 达到 3.869,并且在每个折叠中单调提高,其中大约一半的总增益来自第一次修正涂鸦。我们的代码和训练模型检查点可在 https://github.com/Libo1023/autoPETV-Curriculum 获取。
cs.CV / 69 / 2608.22102

Learning Implicit Constitutive Laws for Dynamic 3D Gaussian Splatting from Monocular Videos

从单目视频中学习动态3D高斯点云的隐式本构定律
Liu, Xiaoyang, Han, Kai
Abstract
We present GCA (Gaussian Constitutive Alignment), a framework for learning implicit constitutive laws from monocular dynamic video of deformable objects represented by 3D Gaussians. Given a static multi-view scan for geometric initialization, our method learns intrinsic physical dynamics solely from a single fixed-viewpoint video of the moving object. Existing implicit methods often suffer from local minima under noisy supervision and lack physical interpretability, while explicit approaches rely on predefined constitutive equations, limiting generalizability and becoming unstable in monocular settings. To address these challenges, our framework unifies LoRA-based adaptation with two key alignment modules. First, we propose Rank-based Depth-Geometric Anchors (RDGA) to establish robust geometric constraints from monocular dynamic observations via scale-invariant rank-based depth alignment, reducing the reliance on unreliable pixel-level color supervision. Second, a Constitutive Prior Regularizer (CPR) integrates classical constitutive models as soft differentiable priors, regularizing the optimization while preserving the flexibility of implicit modeling---even when the actual material is absent from the hypotheses. Extensive experiments on synthetic, real-to-sim, and real-world datasets demonstrate that GCA outperforms existing methods, achieving 48% lower Chamfer Distance than the strongest baseline on synthetic benchmarks while remaining robust under monocular supervision.
Chinese Translation
我们提出了GCA(高斯本构对齐)框架,用于从单目动态视频中学习由3D高斯表示的可变形物体的隐式本构定律。在几何初始化的静态多视图扫描的基础上,我们的方法仅通过一段固定视角的移动物体视频学习内在的物理动态。现有的隐式方法往往在噪声监督下遭遇局部最小值,并且缺乏物理可解释性,而显式方法依赖于预定义的本构方程,限制了其通用性,并在单目设置中变得不稳定。为了解决这些挑战,我们的框架将基于LoRA的适应与两个关键对齐模块结合在一起。首先,我们提出了基于秩的深度几何锚点(RDGA),通过尺度不变的基于秩的深度对齐,从单目动态观察中建立稳健的几何约束,减少对不可靠的像素级颜色监督的依赖。其次,本构先验正则化器(CPR)将经典本构模型作为软可微先验整合,正则化优化的同时保持隐式建模的灵活性——即使在假设中缺乏实际材料的情况下。对合成、真实到仿真和真实世界数据集的广泛实验表明,GCA的表现优于现有方法,在合成基准上实现了比最强基线低48%的Chamfer距离,同时在单目监督下保持稳健性。
cs.CV / 70 / 2608.22116

Vehicle speed dataset for the major European road network derived from Sentinel-2 imagery, 2022-2026

基于Sentinel-2影像的主要欧洲公路网络车辆速度数据集,2022-2026
Adamiak, Maciej, Fendrich, Sascha, Psotta, Julian, Zipf, Alexander
Abstract
The dataset provides individual vehicle speed observations on European E-roads: motorways, trunk roads, primary and secondary roads, as tagged in OpenStreetMap as e-road, for the years 2022-2026. Speeds are derived from Copernicus Sentinel-2 Level-2A satellite optical imagery using a processing pipeline that exploits the short, well-characterized acquisition delays between the blue (B02_10m), green (B03_10m), and red bands (B04_10m) of the Sentinel-2 push-broom instrument. A moving vehicle appears at slightly displaced positions in the three bands, forming a moving echo. The detected displaced intensity peaks are linked into per-vehicle trajectories through a prediction-and-matching procedure. The resulting displacements are converted into ground speeds using publicly accessible inter-band time delays. Each record contains the trajectory geometry, per-channel displacements and headings, internal quality indicators, the estimated speed, the acquisition timestamp, and the source Sentinel-2 product identifier. The dataset is distributed as GeoPackage files, with one record per detected vehicle, and can support studies of traffic patterns, speed behavior, transport modeling, and the calibration of road network attributes at a continental scale.
Chinese Translation
该数据集提供了2022-2026年间欧洲E路(高速公路、干线公路、主要和次要道路)上单个车辆的速度观测数据,这些道路在OpenStreetMap中被标记为e-road。速度数据源自Copernicus Sentinel-2 Level-2A卫星光学影像,采用处理管道利用了Sentinel-2推扫仪的蓝色(B02_10m)、绿色(B03_10m)和红色(B04_10m)波段之间短暂且特征明确的采集延迟。移动中的车辆在三个波段中呈现出略微位移的位置,形成移动回声。通过预测与匹配程序,将检测到的位移强度峰值链接成每辆车的轨迹。随后,利用公开可获取的波段间时间延迟将得到的位移转换为地面速度。每条记录包含轨迹几何、每通道位移和航向、内部质量指标、估计速度、采集时间戳以及源Sentinel-2产品标识符。该数据集以GeoPackage文件的形式分发,每个记录对应一个检测到的车辆,可支持交通模式、速度行为、运输建模以及在大陆尺度上公路网络属性的校准研究。
cs.CV / 71 / 2608.22131

TRACE: Artifact-Robust Statistical Shape Modeling from Imperfect Surface Scans - A Case Study in Craniosynostosis 3D Photography

TRACE:从不完美表面扫描中提取抗伪影的统计形状建模 - 颅缝早闭症三维摄影的案例研究
Bhandari, Sanjay, Khan, Nawazish, Novotna, Alzbeta, Jeong, Tiffany, Bowman, Loretta, Hernandez, Michael, Somorin, Tobi, Govani, Viraj, Goldstein, Jesse, Elhabian, Shireen
Abstract
Craniosynostosis severity analysis increasingly relies on statistical shape models (SSMs) to quantify cranial morphology, but most existing workflows depend on computed tomography or heavily curated three-dimensional (3D) photographs. Raw clinical 3D photographs provide a radiation-free and repeatable alternative, yet often contain shoulders, hands, hair, clothing, scanner noise, and incomplete boundaries that corrupt correspondences. We introduce the Template-constrained Robust Artifact-aware Correspondence Estimation (TRACE) framework, an unsupervised method for constructing SSMs directly from artifact-contaminated clinical 3D head photographs. TRACE predicts sparse anatomically corresponding head-surface control points from the raw point cloud, refines them through a coarse-to-fine Surface-Aware Deformation cascade, and uses thin-plate spline warping to deform a clean template mesh into a subject-specific head reconstruction. This template-constrained formulation keeps dense correspondences on clinically relevant head anatomy while suppressing non-head artifacts. The correspondence module is decoupled from the point-cloud encoder, enabling the same deformation pipeline to be paired with different backbones, including PointNet, DGCNN, and Point Transformer V3. Across all backbones, TRACE substantially improves surface sampling, topology preservation, and shape-model quality over prior SSM methods, providing a scalable foundation for photograph-based craniosynostosis shape analysis and a framework that may extend to other artifact-contaminated surface scans when an appropriate clean template is available.
Chinese Translation
颅缝早闭症的严重程度分析日益依赖统计形状模型(SSMs)来量化颅骨形态,但大多数现有工作流程依赖于计算机断层扫描或经过精心整理的三维(3D)照片。原始临床3D照片提供了一种无辐射且可重复的替代方案,但通常包含肩膀、手、头发、衣物、扫描噪声和不完整边界,这些因素会干扰对应关系。我们提出了模板约束的鲁棒伪影感知对应估计(TRACE)框架,这是一种无监督的方法,旨在直接从受伪影污染的临床3D头部照片中构建SSMs。TRACE从原始点云中预测稀疏的解剖对应头面控制点,通过粗到细的表面感知变形级联对其进行精细化,并使用薄板样条变形将干净的模板网格变形为特定于个体的头部重建。这种模板约束的形式保持了与临床相关的头部解剖结构的密集对应关系,同时抑制非头部伪影。对应模块与点云编码器解耦,使得相同的变形管道可以与不同的骨干网络配对,包括PointNet、DGCNN和Point Transformer V3。在所有骨干网络中,TRACE显著改善了表面采样、拓扑保持和形状模型质量,相较于先前的SSM方法,为基于照片的颅缝早闭症形状分析提供了可扩展的基础,并为其他受伪影污染的表面扫描提供了一个框架,当有适当的干净模板可用时。
cs.CV / 72 / 2608.22174

When Does Visual Generation Help Visual Understanding in Unified Multimodal Models?

视觉生成何时有助于统一多模态模型中的视觉理解?
Zhu, Yubo, Kan, Zhehan, Yang, Jingyi, Chen, Miaolin, Xing, Jinbo, Zhu, Kai, Wang, Zijian, Zhong, Sheng, Tong, Wei
Abstract
Unified multimodal models (UMMs) can perform both understanding and generation, raising a central question: can visual generation improve understanding? Existing evaluations provide mixed evidence, but confound task difficulty, reasoning paradigms, and the closed-loop interaction between generation and understanding. We introduce VGAU-Diag, a fine-grained evaluation framework for vision generation-assisted understanding. It stratifies samples by difficulty, enables unified evaluation of multiple reasoning paradigms, and uses Oracle-Assisted Reference Protocols. Our analysis shows that generated visual aids help on easier instances but become unreliable as reasoning complexity increases. Oracle-assisted diagnosis further reveals that the main bottleneck often lies on the visual-understanding side rather than the visual-generation side, as current UMMs struggle to leverage even faithful visual aids. We also show that effective visual generation should target visual-understanding bottlenecks rather than add more reasoning steps, and identify a three-stage transition from task-irrelevant noise, to misleading plausible guidance, and finally to useful assistance. These findings would be useful to guide the development of better UMMs.The code is available at https://github.com/zyb1029/VGAU-Diag.
Chinese Translation
统一多模态模型(UMMs)能够同时进行理解和生成,这引发了一个核心问题:视觉生成能否改善理解?现有评估提供了混合证据,但混淆了任务难度、推理范式以及生成与理解之间的闭环互动。我们引入了VGAU-Diag,一个用于视觉生成辅助理解的细粒度评估框架。该框架根据难度对样本进行分层,能够统一评估多种推理范式,并使用Oracle-Assisted Reference Protocols(Oracle辅助参考协议)。我们的分析表明,生成的视觉辅助在较简单的实例中有帮助,但随着推理复杂性的增加,其可靠性下降。Oracle辅助诊断进一步揭示,主要瓶颈往往在视觉理解方面,而非视觉生成方面,因为当前的UMMs在利用甚至真实的视觉辅助时都面临困难。我们还展示了有效的视觉生成应针对视觉理解瓶颈,而不是增加更多的推理步骤,并识别出从任务无关的噪声,到误导性的合理指导,最终到有用的辅助的三阶段转变。这些发现将有助于指导更好的UMMs的开发。代码可在https://github.com/zyb1029/VGAU-Diag获取。
cs.CV / 73 / 2608.22183

VERDICT: Agreement Beats Pixel-Space Verification in Real-Document OCSR

VERDICT:协议优于像素空间验证在真实文档的光学化学结构识别中
Guan, Yani, Dong, Dengpan, Luo, Shuang, Wei, Zi, Han, Joah, Hannah, Dan, Zhang, Yumin, Hu, Qichao, Xu, Kang
Abstract
Optical Chemical Structure Recognition (OCSR) converts 2D molecular depictions in the published literature into SMILES, and is increasingly important for constructing large-scale chemical training datasets. Automation at that scale requires identifying unreliable predictions in the absence of ground truth. Three families of label-free signals were compared on $263$ ACS journal depictions with verified ground truth: model confidence, re-rendering similarity, and agreement among recognizers. Pixel-space re-rendering performed little better than chance (AUROC $0.547$, $95\%$ CI $[0.465,0.629]$), and an oracle-tuned threshold on it reduced correct labels per image from $0.745$ to $0.205$. Agreement among four architecturally distinct recognizers instead reached an AUROC of $0.916$ ($[0.880,0.952]$). The two-of-four rule accepted $81.7\%$ of images at $88.8\%$ precision, the three-of-four rule $52.1\%$ at $98.5\%$. The same pattern held on CLEF-IP, UOB, and USPTO. This distinction is obscured on synthetic benchmarks, where re-rendered predictions naturally resemble their inputs. A substance filter removed $2{,}193$ false agreements on wildcards and R-group fragments, after which the three-of-four rule rejected all $68$ generic depictions. VERDICT was then applied to PMC Open Access, producing $6{,}146$ structure labels for $4{,}833$ molecules; chemist adjudication of $400$ released labels in two independent samples yielded precisions of $0.995$ for the three-of-four tier and $0.958$ for the two-of-four tier. VERDICT therefore enables validated labels for multimodal molecular databases linking structure images, machine-readable representations, and source-publication information. In SES AI's Molecular Universe platform, VERDICT further serves as an image-based interface for searching and retrieving molecular records.
Chinese Translation
光学化学结构识别(Optical Chemical Structure Recognition, OCSR)将已发表文献中的二维分子图示转换为SMILES,并在构建大规模化学训练数据集方面变得越来越重要。这种规模的自动化需要在缺乏真实标签的情况下识别不可靠的预测。我们比较了三类无标签信号在263个具有验证真实标签的ACS期刊图示上的表现:模型置信度、重新渲染相似性和识别者之间的协议。像素空间的重新渲染表现仅略好于随机(AUROC 0.547,95% CI [0.465,0.629]),而对其进行的oracle调优阈值将每幅图像的正确标签数从0.745减少到0.205。四个架构上不同的识别者之间的协议则达到了AUROC 0.916([0.880,0.952])。两者中的四个规则在88.8%的精度下接受了81.7%的图像,三者中的四个规则则在98.5%的精度下接受了52.1%。在CLEF-IP、UOB和USPTO上也观察到了相同的模式。这一区别在合成基准上被掩盖,因为重新渲染的预测自然与其输入相似。一个物质过滤器去除了2,193个关于通配符和R-组片段的错误协议,之后三者中的四个规则拒绝了所有68个通用图示。VERDICT随后应用于PMC开放获取,为4,833个分子生成了6,146个结构标签;对400个释放标签的化学家裁定在两个独立样本中获得了三者中的四个层次0.995和两者中的四个层次0.958的精度。因此,VERDICT使得多模态分子数据库能够链接结构图像、机器可读表示和来源出版信息的验证标签。在SES AI的分子宇宙平台上,VERDICT进一步作为一个基于图像的界面,用于搜索和检索分子记录。
cs.CV / 74 / 2608.22193

SAM3Dual: A 3rd Place Solution to the MOSEv2 Track, 8th LSVOS Challenge

SAM3Dual:第八届大规模视频目标分割挑战赛MOSEv2赛道的第三名解决方案
Kim, JeongRae, Kim, Chaehyun, Lim, Changwon
Abstract
We present SAM3Dual, our third-place solution to the MOSEv2 track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge at ECCV 2026. SAM3Dual is a training-free inference extension of pretrained SAM 3 that explicitly separates temporal memory into a short-term branch for recent observations and a long-term branch for interval-sampled historical representations. The two memory responses are combined using a deterministic sequence-relative fusion schedule and conservatively modulated by the previous-frame object confidence. All pretrained SAM 3 parameters remain frozen, requiring no task-specific training, fine-tuning, test-time training, or online parameter optimization. The complete system achieved an official J&F score of 64.37 and ranked third in the MOSEv2 track. This result highlights the potential of reorganizing temporal memory entirely at inference time to obtain competitive long-term VOS performance while preserving the pretrained model.
Chinese Translation
我们提出了SAM3Dual,这是我们在2026年ECCV第八届大规模视频目标分割(LSVOS)挑战赛MOSEv2赛道上的第三名解决方案。SAM3Dual是预训练的SAM 3的无训练推理扩展,明确将时间记忆分为短期分支(用于最近观察)和长期分支(用于间隔采样的历史表示)。这两种记忆响应通过确定性的序列相对融合调度进行组合,并由前一帧的目标置信度进行保守调节。所有预训练的SAM 3参数保持不变,无需特定任务的训练、微调、测试时训练或在线参数优化。完整系统在MOSEv2赛道上获得了官方J&F分数64.37,并排名第三。该结果突显了在推理时完全重组时间记忆以获得具有竞争力的长期视频目标分割性能的潜力,同时保留预训练模型。
cs.CV / 75 / 2608.22217

UR$^{2}$-MLLM: Uncertainty-aware Revisit Reasoning in Multimodal Large Language Models for Radiology Report Generation

UR$^{2}$-MLLM:用于放射学报告生成的多模态大语言模型中的不确定性感知重访推理
Chen, Yucheng, Yu, Yang, Zhou, Jiazhou, Shi, Yufei, Lan, Yongying, Zhang, Yichi, Li, Liyi, Yeo, Si Yong
Abstract
Radiologists generate diagnostic reports through iterative and selective revisiting of suspicious regions to refine their interpretations. Recent multimodal large language models (MLLMs) for radiology report generation (RRG) have shifted from text-only reasoning toward a ``Thinking-with-Images'' paradigm, incorporating visual evidence into the reasoning process. However, existing methods provide static visual evidence without a dynamic revisit mechanism during reasoning, neglecting how radiologists re-examine uncertain observations. To this end, we propose an Uncertainty-aware Revisit Reasoning MLLM (UR$^{2}$-MLLM) framework that dynamically revisits uncertain regions during reasoning for RRG. UR$^{2}$-MLLM is first equipped with uncertainty perception by training on an uncertainty-aware dataset. We then construct a multimodal reasoning trajectory dataset together with a detect-and-copy mechanism, which guides when and where to revisit. Finally, a visual grounding reward refines this behavior through reinforcement learning, aligning the revisited regions with corresponding anatomical structures. Experiments on MIMIC-CXR and IU-Xray show that UR$^{2}$-MLLM achieves state-of-the-art performance, highlighting the value of uncertainty-aware visual revisit reasoning for reliable and clinically aligned report generation.
Chinese Translation
放射科医生通过对可疑区域的迭代和选择性重访来生成诊断报告,以完善他们的解读。最近用于放射学报告生成(RRG)的多模态大语言模型(MLLMs)已从仅文本推理转向“与图像思考”(Thinking-with-Images)范式,将视觉证据纳入推理过程。然而,现有方法在推理过程中提供静态的视觉证据,而没有动态重访机制,忽视了放射科医生如何重新审视不确定的观察结果。为此,我们提出了一种不确定性感知重访推理的多模态大语言模型框架(UR$^{2}$-MLLM),该框架在RRG推理过程中动态重访不确定区域。UR$^{2}$-MLLM首先通过在不确定性感知数据集上进行训练来具备不确定性感知。然后,我们构建了一个多模态推理轨迹数据集,并结合检测与复制机制,指导何时以及在哪里重访。最后,通过强化学习,视觉定位奖励优化了这一行为,使重访区域与相应的解剖结构对齐。在MIMIC-CXR和IU-Xray上的实验表明,UR$^{2}$-MLLM实现了最先进的性能,突显了不确定性感知视觉重访推理在可靠和临床对齐的报告生成中的价值。
cs.CV / 76 / 2608.22238

Hyper^2: Unleashing Hyperbolic Geometry's Full Potential via Dual-Space Consistency

Hyper^2:通过双空间一致性释放双曲几何的全部潜力
Zheng, Guantian, Xu, Haiyang, Gao, Tianyu
Abstract
HyperbolicCD pioneered hyperbolic geometry for point cloud completion by replacing the Euclidean Chamfer distance with arcosh(1+alpha||x-y||^2), but the reported gains are modest (3-7% Chamfer reduction across SeedFormer, PointAttN and PMP-Net backbones on PCN and ShapeNet-55). We argue the bottleneck lies elsewhere: the loss is hyperbolic but the encoder it back-propagates through is Euclidean, so the position-dependent supervision of the loss is averaged away by the chain rule before it reaches the parameters. We call this a cross-geometry mismatch, and make it testable through two model-agnostic indicators, feature-loss correlation r_FL and effective gradient utilisation u_G. On an SVDFormer backbone trained with HyperbolicCD's loss alone we measure (r_FL, u_G) = (0.68, 39%). We propose Hyper^2, a dual-space consistency framework that extends HyperbolicCD by reusing the identical arcosh(1+alpha d^2) functional form as a positional bias on the refinement attention (a hyperbolic distance encoding), paired with HyperbolicCD's hyperbolic Chamfer loss under a single shared curvature alpha. Both operators are O(N log N) scalar non-linearities on Euclidean distances and together add only ~1.6% FLOPs over SVDFormer. Hyper^2 delivers -22.9% Chamfer on ShapeNet-55 over SVDFormer (well above the 13.2% linear sum of the -12.0% loss-only and -1.2% encoding-only single-space ablations) and -37.5% on the 21 unseen ShapeNet-34 categories. The two indicators remain essentially flat for any single-space configuration but jump together to (0.95, 87%) only when both encoder and loss are hyperbolic, supporting the claim that geometric consistency across encoder and loss, rather than either operator alone, is what enables hyperbolic supervision in point cloud completion. Code is available at https://github.com/Ethan-Zheng136/Hyper-2.
Chinese Translation
HyperbolicCD通过用arcosh(1+alpha||x-y||^2)替代欧几里得Chamfer距离,开创了双曲几何在点云补全中的应用,但报告的增益有限(在PCN和ShapeNet-55上,SeedFormer、PointAttN和PMP-Net骨干网的Chamfer减少幅度为3-7%)。我们认为瓶颈在于其他地方:损失是双曲的,但其反向传播的编码器是欧几里得的,因此损失的依赖位置的监督在到达参数之前被链式法则平均掉。我们称之为跨几何不匹配,并通过两个与模型无关的指标进行测试,即特征损失相关性r_FL和有效梯度利用率u_G。在仅使用HyperbolicCD损失训练的SVDFormer骨干网中,我们测得(r_FL, u_G) = (0.68, 39%)。我们提出了Hyper^2,一个双空间一致性框架,通过在细化注意力(一个双曲距离编码)上重用相同的arcosh(1+alpha d^2)函数形式,扩展了HyperbolicCD,并结合了HyperbolicCD的双曲Chamfer损失,采用单一共享曲率alpha。这两个算子在欧几里得距离上都是O(N log N)的标量非线性,并且总共仅增加约1.6%的FLOPs。Hyper^2在ShapeNet-55上相较于SVDFormer实现了-22.9%的Chamfer减少(远高于-12.0%损失单独和-1.2%编码单独的单空间消融的线性和13.2%),在21个未见的ShapeNet-34类别上实现了-37.5%。这两个指标在任何单空间配置下基本保持平稳,但只有当编码器和损失都是双曲时,它们才一起跃升至(0.95, 87%),支持了编码器和损失之间的几何一致性,而非单个算子,使得点云补全中的双曲监督成为可能。代码可在https://github.com/Ethan-Zheng136/Hyper-2获取。
cs.CV / 77 / 2608.22263

Training-Free VLM Personalization via Calibrated Residual Decoding

无训练的视觉语言模型个性化通过校准残差解码
Yu, Jiaao, Ma, Yujian, Hu, Xianming, Wang, Pengran, Li, Ang
Abstract
Vision-language models can be personalized in a training-free manner by directly providing user profiles, preferences, or visual references at inference time, without updating model parameters. However, direct personalized prompting does not guarantee that the model will reliably exploit such evidence. The predictive distribution under the positive user profile often mixes two sources: personalized signals genuinely supported by the current profile, and the model's generic visual or linguistic priors. As a result, from the positive-profile response alone, it is difficult to determine whether a high-confidence answer is supported by the user profile or merely reflects the model's default preference. To address this problem, we propose a training-free calibrated residual decoding framework. Given the same image and question, we construct three evidence conditions: a positive profile , a counterfactual profile , and an empty profile . Our method keeps the prediction under as the anchored base, and explicitly estimates the marginal contribution of personalization from score differences across the three conditions. We further introduce normalized-entropy-based uncertainty calibration, allowing the strength of personalized enhancement to adapt to the reliability of the residual signal. Experiments on MMPB, YoLLaVA, and MyVLM show that the proposed method improves personalized multimodal understanding without fine-tuning, with consistent gains on identity-sensitive visual personalization tasks. Additional analysis shows that entropy calibration stabilizes residual decoding when the contrastive personalization signal is uncertain.
Chinese Translation
视觉语言模型可以通过在推理时直接提供用户档案、偏好或视觉参考,以无训练的方式进行个性化,而无需更新模型参数。然而,直接的个性化提示并不能保证模型会可靠地利用这些证据。在积极用户档案下的预测分布通常混合了两个来源:由当前档案真实支持的个性化信号,以及模型的通用视觉或语言先验。因此,仅凭积极档案的响应,很难确定高置信度的答案是由用户档案支持的,还是仅仅反映了模型的默认偏好。为了解决这个问题,我们提出了一种无训练的校准残差解码框架。在给定相同图像和问题的情况下,我们构建了三种证据条件:积极档案、反事实档案和空档案。我们的方法将预测保持在作为锚定基准的情况下,并明确估计个性化的边际贡献,基于三个条件之间的得分差异。我们进一步引入基于归一化熵的置信度校准,使个性化增强的强度能够适应残差信号的可靠性。在MMPB、YoLLaVA和MyVLM上的实验表明,所提出的方法在不进行微调的情况下提高了个性化的多模态理解,并在身份敏感的视觉个性化任务上取得了一致的提升。额外分析表明,当对比个性化信号不确定时,熵校准能够稳定残差解码。
cs.CV / 78 / 2608.22272

GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets

GAN-Diff:将预训练的WGAN-GP特征与条件扩散U-Net相结合
Ahmed, Saif, Galib, Ashadulla Hil, Antu, S. M. Riaz Rahman, Dhrubo, Ahmed Faizul Haque, Pramanik, Souvik, Qayum, Mohammad Abdul, Sajjad, Mohsin, Khan, Mohammad Ashrafuzzaman
Abstract
Generative adversarial networks (GANs) can provide efficient image generation, while diffusion models offer high-quality image restoration but require iterative sampling. This paper presents a hybrid GAN-guided diffusion framework that uses a pretrained Wasserstein GAN with gradient penalty (WGAN-GP) as a feature prior for conditional diffusion-based image restoration. Intermediate features from the frozen WGAN-GP generator are incorporated into a diffusion U-Net through cross-attention and remain fixed during the DDIM sampling process. The framework is evaluated on two restoration tasks, Gaussian denoising and 2Xsuper-resolution, using CelebA face images. During development, several sources of instability were identified and addressed, including adversarial learning-rate imbalance, inappropriate diffusion initialization, excessive corruption, and insufficient parameter averaging. The resulting framework consistently improves the quality of both degraded and low-resolution images. In particular, it improves denoising performance by 4.40 dB in PSNR and super-resolution performance by 3.70 dB over their respective input baselines. These results demonstrate the potential of a frozen GAN feature prior to guide diffusion models toward stable and effective image restoration.
Chinese Translation
生成对抗网络(GANs)能够高效地生成图像,而扩散模型则提供高质量的图像恢复,但需要迭代采样。本文提出了一种混合的GAN引导扩散框架,该框架使用预训练的带梯度惩罚的Wasserstein GAN(WGAN-GP)作为条件扩散图像恢复的特征先验。来自冻结的WGAN-GP生成器的中间特征通过交叉注意力被纳入扩散U-Net,并在DDIM采样过程中保持固定。该框架在两个恢复任务上进行了评估,即高斯去噪和2倍超分辨率,使用CelebA人脸图像。在开发过程中,识别并解决了多个不稳定因素,包括对抗学习率不平衡、不当的扩散初始化、过度损坏和参数平均不足。最终的框架持续提高了退化和低分辨率图像的质量。特别是,它在PSNR上提高了去噪性能4.40 dB,在超分辨率性能上提高了3.70 dB,相较于各自的输入基线。这些结果展示了冻结的GAN特征先验引导扩散模型朝向稳定和有效的图像恢复的潜力。
cs.CV / 79 / 2608.22279

OVIBench: Benchmarking Online Video Question Answering under Interruption

OVIBench:中断下在线视频问答的基准测试
Liu, Naiming, Wu, Zhiheng, Wang, Shuning, Zhang, Tie, Liu, Bowen, Wang, Tong
Abstract
Recent vision language models (VLMs) have achieved strong progress in video understanding. However, most existing video QA research and benchmarks still follow an offline, single-round paradigm, overlooking realistic interactions where users may interrupt the model during answer generation. To address this gap, we formulate the task of Online Video Question Answering under Interruption and introduce OVIBench, the first standardized benchmark for evaluating VLMs in this setting. OVIBench categorizes interruptions into three types: Cancellation, False Trigger, Correction and supports both open-ended and multiple-choice evaluations. To enable large-scale and reproducible testing, we develop an offline simulation protocol that reproduces interruption during generation under a unified temporal setup, together with a multi-dimensional metric suite for assessing interruption understanding and response generation. Experiments demonstrate that OVIBench effectively distinguishes models' interruption-handling abilities, especially in following correction requests. Finally, we construct a train set OVI-Train for interruption-aware fine-tuning. Models fine-tuned on this dataset achieve significant gains on OVIBench, validating the effectiveness of our benchmark and data design. OVIBench, OVI-Train, and the evaluation code will be released.
Chinese Translation
近期的视觉语言模型(VLMs)在视频理解方面取得了显著进展。然而,大多数现有的视频问答研究和基准仍然遵循离线的单轮范式,忽视了用户在答案生成过程中可能中断模型的现实交互。为了解决这一问题,我们提出了中断下在线视频问答的任务,并引入了OVIBench,这是第一个用于在此设置下评估VLMs的标准化基准。OVIBench将中断分为三种类型:取消、误触发和纠正,并支持开放式和多项选择的评估。为了实现大规模和可重复的测试,我们开发了一种离线模拟协议,在统一的时间设置下重现生成过程中的中断,同时提供了一套多维度的评估指标,用于评估中断理解和响应生成。实验表明,OVIBench能够有效区分模型的中断处理能力,特别是在遵循纠正请求方面。最后,我们构建了一个用于中断感知微调的训练集OVI-Train。基于该数据集微调的模型在OVIBench上取得了显著提升,验证了我们基准和数据设计的有效性。OVIBench、OVI-Train及评估代码将被公开发布。
cs.CV / 80 / 2608.22289

DECO: Depth-Guided Co-Visibility Reasoning for Low-Altitude UAV Visual Localization

DECO:用于低空无人机视觉定位的深度引导共视推理
Ye, Yibin, Teng, Xichao, Chen, Shuo, Song, Xiaokai, Guan, Dongdong, Yu, Qifeng, Li, Zhang
Abstract
Unmanned aerial vehicles (UAVs) increasingly require robust visual localization in GNSS-denied environments. A common solution estimates UAV poses by matching keypoints between UAV images and geo-tagged orthographic reference maps derived from satellite or aerial imagery, followed by Perspective-\(n\)-Point (PnP) pose solving. However, such reference maps mainly record top-down surfaces such as roofs and ground planes, while vertical structures such as facades and walls are often compressed or missing. Consequently, many visually distinctive keypoints in low-altitude UAV images have no valid counterparts in the reference map, leading to redundant matches and inaccurate pose estimation. To address this issue, we propose DECO, a DEpth-guided CO-visibility reasoning framework for low-altitude UAV visual localization. DECO uses monocular depth priors to infer local surface geometry and estimate co-visible regions between UAV images and the reference map. Based on this prior, a Geometry-Saliency Coupled Co-visibility Score is introduced to jointly consider geometric co-visibility and detector saliency for keypoint ranking. In this way, DECO retains keypoints that are both visually distinctive and geometrically co-visible, improving feature matching and PnP-based pose estimation. Extensive experiments demonstrate that DECO achieves superior localization performance and can be integrated with different depth models, feature detectors, and matchers. The source code will be available at https://github.com/UAV-AVL/DECO.
Chinese Translation
无人驾驶飞行器(UAV)在无GNSS环境中越来越需要稳健的视觉定位。一个常见的解决方案是通过匹配无人机图像与来自卫星或航拍图像的地理标记正射参考图之间的关键点来估计无人机姿态,随后进行透视- 点(PnP)姿态求解。然而,这些参考图主要记录了如屋顶和地面等自上而下的表面,而垂直结构如立面和墙壁往往被压缩或缺失。因此,许多在低空无人机图像中视觉上独特的关键点在参考图中没有有效的对应点,导致冗余匹配和不准确的姿态估计。为了解决这个问题,我们提出了DECO,一个用于低空无人机视觉定位的深度引导共视推理框架。DECO利用单目深度先验推断局部表面几何形状,并估计无人机图像与参考图之间的共视区域。在此基础上,引入几何-显著性耦合共视评分,以联合考虑几何共视性和检测器显著性进行关键点排序。通过这种方式,DECO保留了既视觉上独特又几何上共视的关键点,从而改善特征匹配和基于PnP的姿态估计。大量实验表明,DECO实现了卓越的定位性能,并且可以与不同的深度模型、特征检测器和匹配器集成。源代码将发布在 https://github.com/UAV-AVL/DECO。
cs.CV / 81 / 2608.22299

Targeted Iterative Filtering

针对性迭代滤波
Åström, Freddie, Felsberg, Michael, Baravdish, George, Lundström, Claes
Abstract
The assessment of image denoising results depends on the respective application area, i.e. image compression, still-image acquisition, and medical images require entirely different behavior of the applied denoising method. In this paper we propose a novel, nonlinear diffusion scheme that is derived from a linear diffusion process in a value space determined by the application. We show that application-driven linear diffusion in the transformed space compares favorably with existing nonlinear diffusion techniques.
Chinese Translation
图像去噪结果的评估依赖于各自的应用领域,即图像压缩、静态图像采集和医学图像需要应用不同的去噪方法行为。在本文中,我们提出了一种新颖的非线性扩散方案,该方案源自于在由应用决定的值空间中的线性扩散过程。我们展示了在变换空间中以应用驱动的线性扩散与现有的非线性扩散技术相比具有良好的效果。
cs.CV / 82 / 2608.22300

Self-Calibrating Dense Displacement Fields for Reliable Co-Registration of Large Optical Satellite Imagery

自校准密集位移场用于大规模光学卫星影像的可靠配准
Sun, Shoukun, Wang, Zhe, Salati, Sanaz, Zhang, Jiyin, Wang, Hui, Ma, Xiaogang
Abstract
Co-registration underlies nearly every multi-temporal and multi-sensor use of optical satellite imagery, and operational products still carry documented offsets well above the fraction-of-a-pixel scale at which change detection, time series, and data fusion degrade. Real image pairs differ along several axes at once (sensor response, scene content, viewing geometry, resolution, mosaic seams), and the last of these is not a single global motion. Existing tools embed a motion model and constants tuned to their development data; a pair that fits is registered precisely, while one that does not either fails to match or returns a result wrong by tens of pixels with no failure reported. Learned matchers add a GPU requirement and carry no accuracy guarantee outside their training distribution. We present SCDF (self-calibrating displacement fields), a training-free, GPU-free estimator whose motion model is the dense per-pixel displacement field itself, so no scene motion falls outside the model. A single predict--measure--filter loop runs over a resolution pyramid: the accumulated field predicts where each patch of the moving image falls in the reference, RootSIFT matching and a correlation pass measure the displacement there to sub-pixel precision, and filters whose thresholds are all calibrated on the image pair itself decide what survives. One configuration, with no per-dataset tuning, processes full $8192^2$ scenes on a single CPU core. On 584 constructed-ground-truth pairs built from real Sentinel-2, Landsat-8/9, and NAIP imagery, against seven classical baselines and two zero-shot pretrained matchers, SCDF registers every pair with zero failures, reduces the best baseline's real-pair median end-point error from 6.83 to 4.17m, and cuts its 90th percentile from 17.8 to 7.77m.
Chinese Translation
配准是几乎所有多时相和多传感器光学卫星影像应用的基础,而现有的操作产品仍然存在明显高于像素分数级别的偏差,这会导致变化检测、时间序列和数据融合的效果下降。真实的图像对在多个维度上存在差异(传感器响应、场景内容、视角几何、分辨率、马赛克接缝),而最后一个因素并不是单一的全局运动。现有工具嵌入了运动模型和针对其开发数据调优的常数;适合的图像对能够精确配准,而不适合的图像对则要么无法匹配,要么返回数十个像素的错误结果而不报告失败。学习型匹配器增加了对GPU的需求,并且在其训练分布之外不提供准确性保证。我们提出了SCDF(自校准位移场),这是一种无训练、无GPU的估计器,其运动模型就是密集的逐像素位移场,因此没有场景运动会超出模型范围。一个单一的预测-测量-滤波循环在分辨率金字塔上运行:累积的位移场预测移动图像的每个补丁在参考图像中的位置,RootSIFT匹配和相关性传递在此处以亚像素精度测量位移,所有阈值都在图像对本身上进行校准的滤波器决定了哪些结果得以保留。一个配置,无需针对每个数据集的调优,能够在单个CPU核心上处理完整的$8192^2$场景。在584对基于真实Sentinel-2、Landsat-8/9和NAIP影像构建的地面真实数据对上,与七个经典基线和两个零样本预训练匹配器相比,SCDF以零失败的方式注册每一对,减少了最佳基线的真实对中位端点误差从6.83米降至4.17米,并将其90百分位数从17.8米降至7.77米。
cs.CV / 83 / 2608.22302

On Tensor-Based PDEs and their Corresponding Variational Formulations with Application to Color Image Denoising

基于张量的偏微分方程及其对应的变分形式在彩色图像去噪中的应用
Åström, Freddie, Baravdish, George, Felsberg, Michael
Abstract
The case when a partial differential equation (PDE) can be considered as an Euler-Lagrange (E-L) equation of an energy functional, consisting of a data term and a smoothness term is investigated. We show the necessary conditions for a PDE to be the E-L equation for a corresponding functional. This energy functional is applied to a color image denoising problem and it is shown that the method compares favorably to current state-of-the-art color image denoising techniques.
Chinese Translation
研究了当偏微分方程(PDE)可以视为能量泛函的欧拉-拉格朗日(E-L)方程的情况,该能量泛函由数据项和光滑性项组成。我们展示了一个PDE成为对应泛函的E-L方程的必要条件。该能量泛函被应用于彩色图像去噪问题,并且结果表明该方法与当前最先进的彩色图像去噪技术相比具有良好的效果。
cs.CV / 84 / 2608.22313

Adapting Dense Vision-Language Relationships for Multi-label Classification with Partial Label

适应稠密视觉-语言关系以进行部分标签的多标签分类
Chen, Cheng, Zhao, Yifan, Li, Jia
Abstract
Learning multi-label image classification with incomplete annotations is a challenging task that has been widely studied for its superior trade-off between high efficiency and less labor consumption on large-scale datasets. Predominant methods rely on strong prior assumptions to recover the missing semantics from partial annotations. However, these statistic priors suffer from unstable semantic mistakes and thus lead to catastrophic overfitting. Toward this end, we propose a Language-driven Dense Semantic Adaptor (LDSA) that excavates prior-adaptive relationships from multimodal pretrained CLIP models. In our approach, the densely contrastive adaptor is first proposed to construct dense visual contrastive constraints, transferring the task-specific knowledge to visual domains. We then propose a language-driven interactive decoder with the help of class-specific prompt tuning, which adapts language proxies with visual domains. With the collaborative learning of proposed modules, experimental results demonstrate our proposed LDSA achieves a new state of the art on public multi-label classification benchmarks, and interpretable analyses reveal that our LDSA discovers implicit semantic relationships with the prior-adaptive learning scheme.
Chinese Translation
学习具有不完整注释的多标签图像分类是一项具有挑战性的任务,因其在大规模数据集上高效性与低劳动消耗之间的优越权衡而受到广泛研究。主要方法依赖于强先验假设,以从部分注释中恢复缺失的语义。然而,这些统计先验存在不稳定的语义错误,从而导致灾难性的过拟合。为此,我们提出了一种语言驱动的稠密语义适配器(Language-driven Dense Semantic Adaptor, LDSA),该适配器从多模态预训练的CLIP模型中挖掘先验自适应关系。在我们的方法中,首先提出了稠密对比适配器,以构建稠密的视觉对比约束,将任务特定知识转移到视觉领域。然后,我们提出了一种语言驱动的交互解码器,借助类别特定的提示调优,适配语言代理与视觉领域。通过所提模块的协同学习,实验结果表明我们提出的LDSA在公共多标签分类基准上达到了新的最先进水平,且可解释性分析揭示我们的LDSA通过先验自适应学习方案发现了隐含的语义关系。
cs.CV / 85 / 2608.22314

On the Choice of Tensor Estimation for Corner Detection, Optical Flow and Denoising

关于角点检测、光流和去噪的张量估计选择
Åström, Freddie, Felsberg, Michael
Abstract
Many image processing methods such as corner detection, optical flow and iterative enhancement make use of image tensors. Generally, these tensors are estimated using the structure tensor. In this work we show that the gradient energy tensor can be used as an alternative to the structure tensor in several cases. We apply the gradient energy tensor to common image problem applications such as corner detection, optical flow and image enhancement. Our experimental results suggest that the gradient energy tensor enables real-time tensor-based image enhancement using the graphical processing unit (GPU) and we obtain 40% increase of frame rate without loss of image quality.
Chinese Translation
许多图像处理方法,如角点检测、光流和迭代增强,利用图像张量。通常,这些张量是通过结构张量进行估计的。在本研究中,我们展示了在某些情况下,梯度能量张量可以作为结构张量的替代方案。我们将梯度能量张量应用于常见的图像问题,如角点检测、光流和图像增强。我们的实验结果表明,梯度能量张量能够利用图形处理单元(GPU)实现实时的基于张量的图像增强,并且在不损失图像质量的情况下,帧率提高了40%。
cs.CV / 86 / 2608.22323

MedReaMM: Evaluating Large Multimodal Models on Expert-Level Clinical Diagnostic Synthesis

MedReaMM:评估大型多模态模型在专家级临床诊断综合中的表现
Wei, Lai, Chen, Yuchao, Cao, Zhenbiao, Zhang, Xiaojin, Wei, Zhongyu, Wang, Bangting, Chen, Wei, Bai, Xiang
Abstract
The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing interest. However, existing benchmarks largely focus on textual reasoning or isolated visual question-answering (VQA) tasks, lacking holistic integration of clinical narratives and medical imaging, and thus failing to assess the multimodal diagnostic synthesis capability central to expert clinical judgment. To bridge this gap, we introduce MedReaMM, a benchmark specifically designed to evaluate models' ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm. Constructed from case reports sourced from top-tier medical journals and curated clinical case databases, MedReaMM comprises 625 expert-validated cases with an average of 2.79 medical images per case and a total of 1,042 standardized diagnoses annotated with ICD-11 codes. These cases predominantly represent rare, atypical, or multi-system presentations that demand expert-level evidence integration beyond routine pattern recognition. We evaluate 23 Large Multimodal Models (LMMs) and find that most achieve diagnostic accuracy scores below 50%, underscoring a substantial gap in multimodal diagnostic synthesis capability. Further analysis reveals that medical knowledge proficiency, medical image understanding, and evidence integration are all highly correlated with diagnostic performance.
Chinese Translation
大型语言模型(LLMs)在诊断决策中的应用引起了越来越多的关注。然而,现有的基准测试主要集中在文本推理或孤立的视觉问答(VQA)任务上,缺乏临床叙事和医学影像的整体整合,因此未能评估对专家临床判断至关重要的多模态诊断综合能力。为填补这一空白,我们引入了MedReaMM,这是一个专门设计的基准,用于评估模型在完整信息范式下综合异质临床证据的能力,该证据包括详细的病史和多种医学影像,以形成准确的鉴别诊断。MedReaMM基于来自顶级医学期刊和策划的临床案例数据库的病例报告构建,包含625个经过专家验证的病例,每个病例平均有2.79幅医学影像,总共注释了1,042个带有ICD-11代码的标准化诊断。这些病例主要代表稀有、非典型或多系统表现,要求超出常规模式识别的专家级证据整合。我们评估了23个大型多模态模型(LMMs),发现大多数模型的诊断准确率低于50%,突显了多模态诊断综合能力的显著差距。进一步分析表明,医学知识熟练度、医学影像理解和证据整合与诊断表现高度相关。
cs.CV / 87 / 2608.22338

AcroMELD: Recovering Interactive PDF Forms with Structure-Aware Graph Set Transformers

AcroMELD:利用结构感知图集变换器恢复交互式PDF表单
Abramov, Samuel
Abstract
Interactive PDF form fields are often absent from documents that visually resemble forms, leaving users unable to enter data without printing or external editing tools. Detecting the missing widgets is difficult because a field may be indicated by several overlapping cues, born-digital PDFs expose useful but incomplete drawing structure, and dense pages can contain hundreds of fields. We introduce AcroMELD (AcroForm Multi-source Evidence Linking Decoder), a 39.4M-parameter detector that combines a high-resolution visual transformer with label-free PDF primitives. Its 896-query set comprises 384 visual proposals, 384 structure-seeded proposals, and 128 learned recovery queries. Four graph-set layers exchange information over geometry-biased sparse neighborhoods and cross-attend to PDF structure. A learned same-field relation links co-referent candidates, while a localization-quality head is trained on the containment-aware overlap used by the downstream recovery decision. We define a hash-bound evaluation protocol with disjoint development, calibration, internal-test, and quarantined external-holdout roles. The sealed, single-seed candidate reaches native containment micro-$F_1$ 0.9344 on the internal test and 0.8477 on the one-shot external holdout (95% PDF-cluster bootstrap interval [0.8339, 0.8605]). This passes the registered historical FFGBT-v8 reference by 0.0186 absolute $F_1$. Under the stricter external adapter, however, performance is 0.7786 IoU-$0.5$ $F_1$ and 0.2900 COCO mAP, below a locally evaluated CommonForms-L reference; the signature class receives no prediction at the selected threshold. Thus the result supports the registered operational gate while exposing substantial domain and rare-class limitations.
Chinese Translation
交互式PDF表单字段在视觉上类似于表单的文档中常常缺失,导致用户无法在不打印或使用外部编辑工具的情况下输入数据。检测缺失的控件非常困难,因为一个字段可能由多个重叠的线索指示,原生数字PDF展示了有用但不完整的绘图结构,而密集的页面可能包含数百个字段。我们提出了AcroMELD(AcroForm多源证据链接解码器),这是一种具有3940万个参数的检测器,结合了高分辨率视觉变换器和无标签的PDF原语。其896个查询集包括384个视觉提议、384个结构种子提议和128个学习的恢复查询。四个图集层在几何偏向的稀疏邻域中交换信息,并交叉关注PDF结构。一个学习的同字段关系将共同指代的候选项连接起来,而一个定位质量头则在下游恢复决策所使用的包含感知重叠上进行训练。我们定义了一个哈希绑定评估协议,具有不重叠的开发、校准、内部测试和隔离的外部保留角色。密封的单种子候选在内部测试中达到原生包含微-$F_1$ 0.9344,在一次性外部保留中达到0.8477(95% PDF集群自助区间[0.8339, 0.8605])。这比注册的历史FFGBT-v8参考高出0.0186个绝对$F_1$。然而,在更严格的外部适配器下,性能为0.7786 IoU-$0.5$ $F_1$和0.2900 COCO mAP,低于本地评估的CommonForms-L参考;在所选阈值下,签名类别没有任何预测。因此,结果支持注册的操作门,同时暴露出相当大的领域和稀有类别的局限性。
cs.CV / 88 / 2608.22341

TransHands: Repurposing Human Pose Encoders as Hand Pose Encoders

TransHands:将人体姿态编码器重新用于手部姿态编码器
Piccioli, Milo, Amprimo, Gianluca, Ferraris, Claudia, Olmo, Gabriella
Abstract
Lifting 3D hand poses from 2D monocular representations remains challenging due to the limited availability of large-scale, diverse 3D-annotated hand datasets, in contrast to the abundance of human body motion data. We address this limitation by transferring motion representations learned from large body pose corpora to the hand domain. We introduce TransHands, a backbone-agnostic transfer learning framework that enables pre-trained human motion encoders to be effectively adapted for 3D hand pose estimation from 2D pose inputs. Rather than training hand-specific biomechanical models from scratch, TransHands combines a two-stage training and fine-tuning strategy with a lightweight hand-specific input adaptation module that aligns hand kinematics with the representation space learned for full-body motion. We evaluate TransHands across four state-of-the-art motion modeling architectures, including transformer-based, graph-based, and frequency- domain models. Results demonstrate that motion priors learned from body pose data transfer consistently across architectures, yielding consistent accuracy gains, strong cross-domain generalization, particularly in challenging egocentric settings, and applicability for downstream tasks in real-world contexts.
Chinese Translation
从2D单目表示中提取3D手部姿态仍然具有挑战性,这主要是由于大规模、多样化的3D标注手部数据集的稀缺,而与丰富的人体运动数据形成对比。我们通过将从大型人体姿态语料库中学习到的运动表示转移到手部领域来解决这一限制。我们提出了TransHands,一个与主干网络无关的迁移学习框架,使得预训练的人体运动编码器能够有效地适应从2D姿态输入进行3D手部姿态估计。TransHands并不是从头开始训练特定于手部的生物力学模型,而是结合了两阶段的训练和微调策略,以及一个轻量级的手部特定输入适配模块,该模块将手部运动学与为全身运动学习的表示空间对齐。我们在四种最先进的运动建模架构上评估TransHands,包括基于变换器的、基于图的和频域模型。结果表明,从身体姿态数据中学习到的运动先验在不同架构间一致地转移,带来了稳定的准确性提升,强大的跨领域泛化能力,尤其是在具有挑战性的自我中心设置中,以及在现实世界上下游任务中的适用性。
cs.CV / 89 / 2608.22344

Fast and Compact 3D Gaussian Splatting with Polarized Opacity Prior

快速紧凑的3D高斯点云渲染与极化透明度先验
Wang, Zi-Ming, Duan, Kai-Wen, Huang, Kowei, Sugimoto, Akihiro, Lai, Shang-Hong
Abstract
3D Gaussian Splatting (3DGS) achieves state-of-the-art rendering quality at real-time speeds but suffers from "model bloat" - a large number of redundant, low-opacity Gaussians that inflate memory usage and training costs. This inefficiency stems from the standard "densify-then-prune" paradigm, which expands the model aggressively before relying on pruning to achieve compactness. To mitigate this problem, we present an efficient training framework that builds an intrinsically compact representation, replacing the conventional densify-then-prune cycle. Our method leverages a synergistic design: an L2 reconstruction loss to provide error-proportional gradients that stabilize optimization, and a novel Polarized Opacity Prior (POP) to actively manage the Gaussian population. POP steers informative primitives toward full opacity and uninformative ones toward transparency, enabling natural pruning and accelerating rendering through Early Ray Termination. Experiments on three public datasets demonstrate that our approach consistently achieves accelerated 3DGS training with significantly fewer Gaussians while maintaining comparable visual reconstruction quality. These results show that the proposed framework provides a simple and effective path toward fast and inherently compact 3DGS training.
Chinese Translation
3D高斯点云渲染(3DGS)在实时速度下实现了最先进的渲染质量,但面临“模型膨胀”问题——大量冗余的低透明度高斯分布导致内存使用和训练成本增加。这种低效源于标准的“先稠密再修剪”范式,该范式在依赖修剪实现紧凑性之前,过度扩展模型。为了解决这个问题,我们提出了一种高效的训练框架,构建内在紧凑的表示,替代传统的稠密-修剪循环。我们的方法利用了一种协同设计:L2重建损失提供与误差成比例的梯度以稳定优化,以及一种新颖的极化透明度先验(Polarized Opacity Prior, POP)来主动管理高斯分布的数量。POP将信息丰富的原始元素引导至完全不透明,而将信息贫乏的元素引导至透明,从而实现自然修剪并通过早期光线终止加速渲染。在三个公共数据集上的实验表明,我们的方法在保持可比视觉重建质量的同时,始终实现了加速的3DGS训练,并显著减少了高斯数量。这些结果表明,所提出的框架为快速且内在紧凑的3DGS训练提供了一条简单有效的路径。
cs.CV / 90 / 2608.22346

Multimodal examination answer data with expert-designed Outcome-Based Education rubrics for criterion-level assessment

基于专家设计的结果导向教育评分标准的多模态考试答案数据用于标准级评估
SM, Jahangir Alam, Syfullah, Md Khalid, Ahmed, Saad, Mou, Munira Akter, Rahman, A K Z Rasel, Rahman, A. K. M. Masudur, Ali, Mohammed Sowket
Abstract
This data article describes a multimodal collection of scanned examination answers paired with expert-designed Outcome-Based Education (OBE) grading metadata. The collection contains 485 answer submissions from 415 consenting students at four academic institutions. Eight faculty contributors supplied examination materials covering nine subjects and 12 distinct question templates. Each answer-level item links a scanned PDF to a randomized identifier, subject label, question, model answer, criterion definitions, performance-level descriptions, criterion marks, and a total mark. The 12 rubrics contain 47 criteria in total. The scans retain realistic academic content, including handwriting, printed text, equations, tables, code, figures, sketches, and diagrams. CamScanner, Adobe Scan, and conventional scanners contributed variation in illumination, contrast, orientation, compression, and resolution. Diverse handwriting, crossed-out work, revised calculations, and inserted corrections add further visual variability for robustness and generalization studies. Preparation involved heterogeneous-source consolidation, label and text standardization, score validation, identifier randomization, filename randomization, and JSON-to-PDF integrity checks. An answer-level audit confirmed 485 unique identifiers, 485 unique PDF filenames, agreement between each total mark and its criterion-mark sum, and scores within the applicable rubric maximum. The data can support rubric-aware automated evaluation, multimodal document understanding, criterion-level feedback, score prediction, and privacy-aware OBE assessment research. Access is restricted to research use and is available from the corresponding author upon reasonable request.
Chinese Translation
本文数据文章描述了一种多模态的扫描考试答案集合,配有专家设计的结果导向教育(Outcome-Based Education, OBE)评分元数据。该集合包含来自四所学术机构的415名同意参与的学生提交的485份答案。八位教师贡献者提供了涵盖九个学科和12种不同问题模板的考试材料。每个答案级项目将扫描的PDF与随机标识符、学科标签、问题、模型答案、标准定义、表现级别描述、标准分数和总分链接。12个评分标准共包含47个标准。扫描保留了真实的学术内容,包括手写、打印文本、方程式、表格、代码、图形、草图和图解。CamScanner、Adobe Scan和传统扫描仪在照明、对比度、方向、压缩和分辨率上提供了变化。多样的手写、划掉的工作、修订的计算和插入的更正进一步增加了视觉变异性,以增强稳健性和泛化研究。准备工作涉及异构来源的整合、标签和文本标准化、分数验证、标识符随机化、文件名随机化和JSON到PDF的完整性检查。答案级审计确认了485个唯一标识符、485个唯一PDF文件名、每个总分与其标准分数之和的一致性,以及分数在适用评分标准的最大值范围内。该数据可以支持基于评分标准的自动评估、多模态文档理解、标准级反馈、分数预测和隐私意识的OBE评估研究。访问权限仅限于研究使用,并可根据合理请求从通讯作者处获得。
cs.CV / 91 / 2608.22359

Pre-Decoding Acoustic Triage for Budgeted Vision-Language Captioning of Untrimmed Egocentric Video

预算化视觉-语言字幕生成的预解码声学分流
Jalayer, Masoud, Li, Changyi, Xiao, Yu
Abstract
Automatically analyzing hours-long egocentric video is increasingly essential for progress monitoring, quality control, and safety in logistics, construction, and manufacturing. Yet current pipelines that process short, fixed-size windows with a vision-language model (VLM) are prohibitively expensive because cost scales with the number of model calls. To reduce this cost, prior work proposes triage policies to select which windows merit a VLM invocation. However, these policies either sample uniformly or rank windows using visual features, which ironically requires the video decoding that the budget constraints are meant to avoid. We propose audio-first triage: select windows using the lightest modality, scored before any video frame is decoded, so the approach composes naturally with token compression or quantization. The novelty lies in the objective, not the representation: rather than a per-frame sound-event detector, we train the selector to trigger once per action. This objective shift improves action coverage by 4.0-10.8 percentage points across all evaluated call rates, using frozen AudioSet-pretrained features without domain-specific sound-event labels. Using fewer than half of the available calls, the triage cuts 9-20% of VLM calls at matched coverage on EPIC-KITCHENS-100 (EK-100), surpasses uniform sampling through the mid-range on Ego4D over 247 clips, and outperforms two recent visual keyframe selectors. Code, the reference implementation and every results file this manuscript reads are at https://github.com/masjalayer/PreDecoding-AcousticTriage.
Chinese Translation
自动分析数小时的自我中心视频在物流、建筑和制造业的进度监控、质量控制和安全性方面变得越来越重要。然而,当前使用视觉-语言模型(VLM)处理短固定大小窗口的流程成本过高,因为成本与模型调用次数成正比。为了降低这一成本,之前的研究提出了分流策略,以选择哪些窗口值得进行VLM调用。然而,这些策略要么均匀采样,要么使用视觉特征对窗口进行排名,这反而需要进行视频解码,而预算限制正是为了避免这一点。我们提出了音频优先的分流:使用最轻的模态选择窗口,在解码任何视频帧之前进行评分,因此该方法可以自然地与令牌压缩或量化结合。其创新之处在于目标,而非表示:我们训练选择器在每个动作中触发一次,而不是每帧的声音事件检测器。这一目标的转变在所有评估的调用率中提高了4.0-10.8个百分点的动作覆盖率,使用冻结的AudioSet预训练特征而无需领域特定的声音事件标签。在使用不到一半的可用调用次数的情况下,该分流在EPIC-KITCHENS-100(EK-100)上以匹配的覆盖率减少了9-20%的VLM调用,在247个片段的Ego4D上超越了中等范围的均匀采样,并且优于两个最近的视觉关键帧选择器。代码、参考实现及本文引用的每个结果文件均可在https://github.com/masjalayer/PreDecoding-AcousticTriage获取。
cs.CV / 92 / 2608.22366

When Do VLMs Help Arabic Manuscript OCR? A Cross-Dataset Study

视觉语言模型何时有助于阿拉伯手稿的光学字符识别?跨数据集研究
Farazi, Moshiur, Alam, Firoj, Maaradji, Abderrahmane, Maamar, Zakaria, Mubarak, Hamdy, Zaghouani, Wajdi
Abstract
Vision-language models (VLMs) are increasingly being used for document understanding, yet their role in Arabic and Islamic manuscript recognition remains underexplored. To address such a gap in this paper, we evaluate traditional OCR, general-purpose VLMs, Arabic-specialized VLMs, and OCR-conditioned VLM correction across eight Arabic text datasets spanning historical manuscripts, aged printed books, clean print, multi-domain documents, and handwriting. The results show that no single approach dominates across setups. On line-level historical manuscripts, VLMs are close to Tesseract; on page-level manuscript images, they perform better; and in several settings, an OCR-conditioned corrector improves over both standalone OCR and standalone VLMs. The central finding is an OCR-prior recoverability principle: OCR conditioning helps when the OCR output remains visually and textually recoverable, providing anchors that the VLM can refine against the image. It improves recognition on aged print, clean print, mixed-domain Arabic, and some Naskh manuscripts, but degrades performance when the prior is script-mismatched or systematically misleading, as in Maghribi manuscripts and realistic student handwriting. Additional diagnostics show that Arabic VLM-OCR is sensitive to diacritics, preprocessing, generation budget, and repetition loops. These findings support an adaptive OCR-VLM workflow that routes pages according to script, OCR-prior recoverability, length diagnostics, and failure-mode indicators.
Chinese Translation
视觉语言模型(VLMs)在文档理解中的应用日益增多,但它们在阿拉伯和伊斯兰手稿识别中的作用仍然未被充分探讨。为了解决这一空白,本文评估了传统光学字符识别(OCR)、通用VLM、阿拉伯专用VLM以及基于OCR的VLM校正,涵盖了八个阿拉伯文本数据集,这些数据集包括历史手稿、老旧印刷书籍、清晰印刷、多领域文档和手写文本。结果显示,没有单一方法在所有设置中占主导地位。在行级历史手稿中,VLM的表现接近Tesseract;在页级手稿图像中,它们的表现更佳;在多个设置中,基于OCR的校正器在独立OCR和独立VLM之上有所提升。核心发现是OCR先验可恢复性原则:当OCR输出在视觉和文本上可恢复时,OCR条件化有助于提供VLM可以针对图像进行细化的锚点。它在老旧印刷、清晰印刷、混合领域的阿拉伯文本以及一些Naskh手稿中提高了识别率,但在先验与脚本不匹配或系统性误导的情况下(如Maghribi手稿和真实的学生手写文本)则会降低性能。额外的诊断表明,阿拉伯VLM-OCR对发音符号、预处理、生成预算和重复循环敏感。这些发现支持一种自适应的OCR-VLM工作流程,根据脚本、OCR先验可恢复性、长度诊断和故障模式指标对页面进行路由。
cs.CV / 93 / 2608.22368

DiD It in 87 Minutes: A Label-Free Softmax-to-Linear Adaptation of Vision Transformers for Object Detection

87分钟内完成:一种无标签的Softmax到线性适配的视觉变换器在目标检测中的应用
Qin, Huaiyuan, Goenawan, Gabriel James, Lin, Zihang, Yang, Muli, Zhu, Hongyuan
Abstract
While linear attention is a compelling mechanism for high-resolution object detection due to its reduced cost for global token mixing, converting the Softmax-attention ViT backbone of a trained detector into a linear-attention one is not a trivial drop-in replacement. Directly swapping the attention operator leads to severe performance degradation, and generic label-free distillation, though effective for classification, often fails on detection tasks. We argue that the central challenge is \textit{detector-interface preservation}: the converted backbone must reproduce the exact feature tensors expected by the fixed downstream detector, rather than merely imitating internal Softmax hidden states. To address this, we introduce Detector-Interface Distillation (DiD), a label-free conversion method that exclusively trains the linear-attention backbone by aligning detector-facing interface tensors with those of a frozen Softmax teacher. On DOTA-v1.5, DiD substantially outperforms established baselines and matches supervised, fully trained linear models. Adaptation completes in roughly 87 minutes on 4 GPUs, and the linearized backbone cuts inference latency by ~62% and peak memory by ~49%. We hope our findings offer the community a simple, label-free route to reusing trained Softmax detectors as efficient linear ones, and encourage interface-aware objectives in future architecture-conversion work.
Chinese Translation
线性注意力由于其在全局标记混合中的成本降低,成为高分辨率目标检测的一个引人注目的机制。然而,将训练好的检测器的Softmax注意力ViT主干转换为线性注意力并不是一个简单的替换。直接替换注意力操作符会导致严重的性能下降,而通用的无标签蒸馏虽然在分类任务中有效,但在检测任务中往往失败。我们认为核心挑战在于 extit{检测器接口的保持}:转换后的主干必须重现固定下游检测器所期望的确切特征张量,而不仅仅是模仿内部的Softmax隐藏状态。为了解决这个问题,我们提出了检测器接口蒸馏(Detector-Interface Distillation, DiD),这是一种无标签的转换方法,专门通过将面向检测器的接口张量与冻结的Softmax教师的张量对齐来训练线性注意力主干。在DOTA-v1.5上,DiD显著超越了既有基线,并与监督的、完全训练的线性模型相匹配。适配过程在4个GPU上大约完成于87分钟,线性化的主干将推理延迟减少了约62%,峰值内存减少了约49%。我们希望我们的发现为社区提供了一条简单的无标签路径,以将训练好的Softmax检测器重用为高效的线性检测器,并鼓励未来架构转换工作中关注接口的目标。
cs.CV / 94 / 2608.22370

LiST: Local-Simplex Test-Time LoRA Fusion

LiST:局部单纯形测试时 LoRA 融合
Shao, Yihua, Li, Jia, Chen, Siyu, Luo, Xinyu, Liu, Yang, Chen, Kecheng, Long, Xinwei, Zhu, Lingyu, Zeng, Fanhu, Wang, Maolin, Yan, Ziyang, Guo, Jingcai, Tang, Hao, Sebe, Nicu, Wang, Zhenyi
Abstract
Task-specific LoRA adapters offer a modular way to specialize large language and vision-language models. However, existing adapter composition methods are mostly static and cannot adapt to individual test inputs. To address these issues, we propose \textbf{LiST}, a label-free test-time LoRA fusion framework that converts an existing LoRA bank into a target-conditioned local simplex and searches sample-specific fusion weights at inference time. LiST builds joint task representations from LoRA parameter anchors and prompt-level behavior vectors, retrieves neighboring adapters as a local search space, and performs branch-preserving fusion without updating the backbone or adapters. Candidate weights are selected by a prompt-level energy with prior, geometric, and stochastic-consistency constraints, and are deployed only when they pass a safe acceptance rule. Otherwise, LiST falls back to a target-conditioned prior. Experiments on multimodal and language benchmarks show that LiST outperforms static LoRA merging and conventional test-time adaptation baselines, while preserving task-specific adapter utility and improving robustness on unseen tasks.
Chinese Translation
任务特定的 LoRA 适配器为大型语言和视觉-语言模型提供了一种模块化的专业化方式。然而,现有的适配器组合方法大多是静态的,无法适应个别测试输入。为了解决这些问题,我们提出了 extbf{LiST},一种无标签的测试时 LoRA 融合框架,它将现有的 LoRA 库转换为目标条件的局部单纯形,并在推理时搜索样本特定的融合权重。LiST 从 LoRA 参数锚点和提示级行为向量构建联合任务表示,检索邻近适配器作为局部搜索空间,并在不更新主干或适配器的情况下执行保持分支的融合。候选权重通过具有先验、几何和随机一致性约束的提示级能量进行选择,并仅在它们通过安全接受规则时才被部署。否则,LiST 将回退到目标条件的先验。在多模态和语言基准上的实验表明,LiST 超越了静态 LoRA 合并和传统测试时适应的基线,同时保持了任务特定适配器的效用,并提高了在未见任务上的鲁棒性。
cs.CV / 95 / 2608.22465

M$^3$ISR: A Multi-Modal Multi-View Benchmark for 3D/4D Gaussian Splatting and Feedforward Compression

M$^3$ISR:一个用于3D/4D高斯喷溅和前馈压缩的多模态多视角基准
Liu, Xinhui, Liu, Lei, Chen, Zhenghao, Zhou, Lebin, Wang, Wei, Jiang, Wei
Abstract
High-fidelity free-viewpoint video (FVV) and interactive rendering increasingly rely on explicit Gaussian representations, yet practical deployment remains constrained by representation size, dynamic updates, and computational cost. Existing multi-view video benchmarks provide valuable real-captured content, but they make it difficult to isolate the effects of controlled camera geometry, representation efficiency, and temporal redundancy. We introduce M$^3$ISR, a controlled synthetic benchmark for 3D and 4D Gaussian Splatting (3DGS/4DGS). The benchmark contains 25 scenes from five indoor and outdoor scene groups, two camera/motion configurations, six synchronized 1080p views, and dense ground-truth annotations including RGB, camera parameters, depth, semantic and instance segmentation, and static--dynamic masks. The shared-center camera design intentionally isolates angular view variation and enables controlled evaluation of novel-view synthesis and representation efficiency. We organize M$^3$ISR into five complementary tracks covering 3DGS synthesis, 4DGS synthesis, 4DGS streaming, 3DGS compression, and 4DGS compression. Representative baseline results show small differences in static reconstruction quality but substantial differences in representation storage, while the evaluated streaming methods exhibit substantially higher reported training or reconstruction cost than the corresponding offline dynamic reconstruction baselines. We further define feedforward compression tasks for 3DGS and 4DGS and provide reference rate--distortion formulations and preliminary baseline evaluations. The benchmark is intended as a controlled and complementary testbed for systematic study of Gaussian-based FVV reconstruction, compression, and streaming.
Chinese Translation
高保真自由视点视频(FVV)和交互式渲染越来越依赖于显式高斯表示,然而实际部署仍受到表示大小、动态更新和计算成本的限制。现有的多视角视频基准提供了有价值的真实捕获内容,但难以隔离受控相机几何、表示效率和时间冗余的影响。我们引入了M$^3$ISR,这是一个用于3D和4D高斯喷溅(3DGS/4DGS)的受控合成基准。该基准包含来自五个室内和室外场景组的25个场景,两个相机/运动配置,六个同步的1080p视图,以及包括RGB、相机参数、深度、语义和实例分割以及静态-动态掩膜在内的密集真实标注。共享中心相机设计有意隔离角度视图变化,并能够对新视图合成和表示效率进行受控评估。我们将M$^3$ISR组织为五个互补的轨道,涵盖3DGS合成、4DGS合成、4DGS流媒体、3DGS压缩和4DGS压缩。代表性的基线结果显示静态重建质量差异较小,但表示存储差异显著,而评估的流媒体方法在报告的训练或重建成本上显著高于相应的离线动态重建基线。我们进一步定义了3DGS和4DGS的前馈压缩任务,并提供了参考的速率-失真公式和初步基线评估。该基准旨在作为一个受控和互补的测试平台,用于系统研究基于高斯的FVV重建、压缩和流媒体。
cs.CV / 96 / 2608.22485

HeatTok: Enhancing Remote Sensing Image Understanding via Thermodiffusion-based Tokenization

HeatTok:通过热扩散基础的标记化增强遥感图像理解
Yan, Yingying, Tang, Jiaqi, Wei, Wei, Wang, Qianzhou, Wu, Jinjian, Geng, Botong, Chen, Jianmin, Xia, Yuyang, Zhang, Lei
Abstract
Current visual tokenizers in Multimodal Large Language Models (MLLMs) predominantly rely on patch-based partitioning, which causes severe semantic mixture and object fragmentation in remote sensing imagery due to the irregular contours of geo-objects. Moreover, existing adaptive methods struggle to extract precise object-level tokens and lack dedicated geometric positional encodings for irregular regions. In this paper, we propose HeatTok, a semantic-aware tokenizer driven by thermodiffusion aggregation. Inspired by the physical principles of heat conduction, HeatTok adaptively merges adjacent homogeneous regions to generate semantically independent, object-aligned irregular tokens. To enable MLLMs to perceive these irregular shapes, we design the Gaussian Multimodal Rotary Positional Embedding (G-MRoPE), which models token spatial distributions via 2D Gaussians and explicitly injects center, scale, and orientation cues. Extensive evaluations on the VRSBench and EarthVQA datasets demonstrate that HeatTok effectively preserves object-level semantic integrity and achieves state-of-the-art performance under a reasonable token budget. The code is available: https://github.com/YingyingYan1/HeatTok.
Chinese Translation
当前多模态大语言模型(MLLMs)中的视觉标记器主要依赖于基于补丁的分割,这导致由于地物的不规则轮廓,遥感图像中出现严重的语义混合和物体碎片化。此外,现有的自适应方法在提取精确的物体级标记方面存在困难,并且缺乏针对不规则区域的专用几何位置编码。本文提出了HeatTok,一种由热扩散聚合驱动的语义感知标记器。HeatTok受到热传导物理原理的启发,自适应地合并相邻的同质区域,以生成语义独立、与物体对齐的不规则标记。为了使MLLMs能够感知这些不规则形状,我们设计了高斯多模态旋转位置嵌入(G-MRoPE),它通过二维高斯模型标记的空间分布,并明确注入中心、尺度和方向线索。在VRSBench和EarthVQA数据集上的广泛评估表明,HeatTok有效地保持了物体级语义完整性,并在合理的标记预算下实现了最先进的性能。代码可在此获取:https://github.com/YingyingYan1/HeatTok。
cs.CV / 97 / 2608.22500

Learning Sample-wise Rank-aware Interpolation Weights for Composed Visual Data Retrieval

学习样本级排名感知插值权重用于组合视觉数据检索
Jeong, Boseung, Park, Taegyu, Kwon, Donghyeon, Cho, Hyunsouk, Kwak, Suha
Abstract
At the heart of composed visual data retrieval is the fusion of a reference visual input and a textual modification into a single query. While current state-of-the-art methods utilize multimodal large language models for this fusion, their complexity introduces prohibitive querytime latency, limiting their scalability. We instead revisit the efficacy of simple linear interpolation within an embedding space, and introduce SRAIN, the first framework that dynamically predicts query-specific interpolation weights. The key challenge lies in the fact that the quality of an interpolation weight should be measured by the interpolated embedding's discriminability from negatives as well as its proximity to true targets; this makes collecting and predicting optimal weights intractable. We overcome this bottleneck through two key innovations: batch-wise rank-aware weight estimation during training, and a compact memory bank that synthesizes hard negatives during inference. SRAIN achieves the best in composed video retrieval and matches the current state of the art in composed image retrieval, all while substantially reducing querytime latency compared to MLLM-based alternatives.
Chinese Translation
组合视觉数据检索的核心在于将参考视觉输入与文本修改融合为单一查询。尽管当前最先进的方法利用多模态大型语言模型进行这种融合,但其复杂性导致查询时间延迟过高,限制了其可扩展性。我们重新审视了在嵌入空间内简单线性插值的有效性,并引入了SRAIN,这是第一个动态预测查询特定插值权重的框架。关键挑战在于插值权重的质量应通过插值嵌入与负样本的可区分性以及与真实目标的接近度来衡量;这使得收集和预测最优权重变得困难。我们通过两个关键创新克服了这一瓶颈:在训练过程中进行批量排名感知权重估计,以及在推理过程中合成困难负样本的紧凑记忆库。SRAIN在组合视频检索中取得了最佳效果,并在组合图像检索中与当前最先进的方法相匹配,同时显著减少了与基于MLLM的替代方案相比的查询时间延迟。
cs.CV / 98 / 2608.22516

TRACE: Temporal Retrieval with Anchored and Convergent Evidence for Long-Horizon Video Understanding

TRACE:基于锚定和汇聚证据的长时视频理解的时间检索
Liu, Pengyiang, Niu, Junbo, Hu, Xiaoyang, Shi, Zhongyue, Wang, Zitian, Huang, Linjiang, Liu, Si
Abstract
A long-video answer is evidence-supported only when the frames decoded from the video cover every event the answer depends on. Existing evaluations score final-answer correctness or predicted evidence intervals, but the frames a method decodes before answering are rarely audited, so correct answers can still rest on incomplete observation. We introduce VES-Bench, a 600-question benchmark of Temporal Ordering and Event Counting items over 348 public long videos. Each item carries a jointly necessary set of evidence intervals, letting us audit at three strictness levels whether a method's decoded frames cover every one of them. We also propose TRACE, a training-free agent that grounds answers in raw visual clips, builds an evidence bundle round by round, and stops only when the answer stabilises as the bundle grows and a final pass over the same clips returns the same answer. Under a same-backbone audit, TRACE answers 50.7% of questions correctly with at least two decoded frames inside every evidence interval, at 98.7 frames per question: over 10 points above uniform decoding at 128 frames (40.2%), and within 2.6 points of uniform decoding at 256 frames at 0.39x its frame cost, while reaching the highest answer accuracy in the audit (63.5%). TRACE also stays competitive on Video-MME (86.1), LVBench (75.6), and LongVideoBench (75.1).
Chinese Translation
长视频答案只有在解码自视频的帧覆盖答案所依赖的每个事件时,才是有证据支持的。现有评估主要评分最终答案的正确性或预测的证据区间,但在回答之前,一个方法解码的帧很少被审计,因此正确答案仍可能基于不完整的观察。我们引入了 VES-Bench,这是一个包含 600 个时间排序和事件计数项目的基准,覆盖 348 个公共长视频。每个项目都携带一组共同必要的证据区间,使我们能够在三个严格程度上审计一个方法解码的帧是否覆盖了每一个证据区间。我们还提出了 TRACE,一个无需训练的代理,它在原始视觉片段中定位答案,逐轮构建证据包,只有在答案随着包的增长而稳定,并且对相同片段的最终遍历返回相同答案时才停止。在相同骨干网络的审计下,TRACE 在每个证据区间内至少有两个解码帧的情况下,正确回答了 50.7% 的问题,平均每个问题解码 98.7 帧:比在 128 帧下的均匀解码(40.2%)高出 10 个点,并且在 256 帧下的均匀解码(0.39 倍的帧成本)内相差仅 2.6 个点,同时在审计中达到了最高的答案准确率(63.5%)。TRACE 在 Video-MME(86.1)、LVBench(75.6)和 LongVideoBench(75.1)上的表现也具有竞争力。
cs.CV / 99 / 2608.22521

VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation

VISTA:视觉自回归生成的测试时组合对齐
Shahabadi, Hossein, Sepasian, Niki, Baghshah, Mahdieh Soleymani
Abstract
Visual autoregressive (VAR) models have emerged as a fast, high-quality alternative to diffusion for text-to-image generation, but like diffusion models they exhibit persistent compositional failures, producing images that violate the attribute bindings and spatial relations specified in the prompt. While a rich line of test-time alignment methods has developed for diffusion, no comparable approach exists for next-scale VAR generation, whose stateful, discrete, multi-resolution sampling process makes existing techniques inapplicable. We close this gap with \textbf{VISTA} (\textbf{Vi}sual Autoregressive \textbf{S}emantic \textbf{T}est-time \textbf{A}lignment), the first gradient-based test-time alignment framework for next-scale autoregressive image generation. Built on Infinity, VISTA intervenes directly in the generation process, optimizing intermediate representations through the frozen transformer to steer visual predictions toward compositional constraints, without modifying model parameters or requiring additional training. VISTA introduces the mechanisms needed to make such optimization stable across scales, together with an extensible objective space that any differentiable constraint on cross-attention can plug into. Across two benchmarks and two model scales, VISTA improves every targeted compositional category, raising the mean targeted score by nearly 20\% on a 2B backbone and almost 6\% on an 8B backbone, with the largest gains on spatial relations. Image quality is preserved: an independent preference model VISTA never optimizes scores its outputs nearly 20\% higher. Notably, the 2B model with VISTA surpasses a backbone four times its size, indicating that a substantial part of the compositional gap between model scales is recoverable at test time.
Chinese Translation
视觉自回归(VAR)模型已成为文本到图像生成中快速、高质量的替代方案,但与扩散模型一样,它们也存在持续的组合失败,生成的图像违反了提示中指定的属性绑定和空间关系。尽管针对扩散模型已经发展出丰富的测试时对齐方法,但对于下一尺度的VAR生成尚无可比拟的方法,因为其状态性、离散性和多分辨率采样过程使现有技术无法适用。我们通过 extbf{VISTA}( extbf{Vi}sual extbf{A}utoregressive extbf{S}emantic extbf{T}est-time extbf{A}lignment)填补了这一空白,这是首个基于梯度的测试时对齐框架,旨在下一尺度的自回归图像生成。VISTA建立在Infinity之上,直接干预生成过程,通过冻结的变换器优化中间表示,以引导视觉预测朝向组合约束,而无需修改模型参数或额外训练。VISTA引入了在不同尺度上实现此类优化所需的机制,以及一个可扩展的目标空间,任何可微分的交叉注意力约束都可以接入。在两个基准和两个模型尺度上,VISTA改善了每个目标组合类别,使得在2B主干上的平均目标分数提高了近20 extperthousand,8B主干上提高了近6 extperthousand,空间关系的提升尤为显著。图像质量得以保持:一个独立的偏好模型在未优化的情况下,其输出分数比VISTA高出近20 extperthousand。值得注意的是,使用VISTA的2B模型超越了一个四倍于其规模的主干,表明模型尺度之间的组合差距在测试时有相当一部分是可以恢复的。
cs.CV / 100 / 2608.22526

RS$^3$-Prune: Read-Sparse, Store-Sparse Token Pruning for Video Object Segmentation

RS$^3$-Prune:用于视频目标分割的读稀疏、存储稀疏的标记修剪
Mandal, Avilasha, Shashikumar, Sarvesh
Abstract
We introduce RS$^3$-Prune, a training-free token-pruning recipe that instantiates as a small set of inference time hooks atop existing video object segmentation (VOS) networks. Modern VOS models have converged on a common, expensive design: an image encoder produces a dense token grid for every frame, and a memory bank accumulates these tokens across all previously processed frames to condition future predictions. As a video grows longer, the resulting token budget governs both per-frame latency and peak GPU memory. Hence these models break on use cases such as --- long-form video or real-time deployment on memory-bounded accelerators. In this work we argue that the right axis along which to compress memory-bank VOS is the token budget itself. RS$^3$-Prune operates in two precise locations within an arbitrary memory-bank VOS pipeline: at the boundary between the image encoder and the memory-attention readout, where we restrict the queries that participate in the cross-frame attention to only a small, geometrically informed subset; and at the boundary between the memory encoder and the memory bank, where we restrict which tokens are ever permitted to enter the bank to those that lie within the object's spatial extent. Over various established benchmarks, RS$^3$-Prune delivers up to $38.8\%$ FPS speedup and reduces $13.1\%$ peak memory usage, while preserving a competitive $\mathcal{J}$&$\mathcal{F}$ compared to the unmodified VOS networks.
Chinese Translation
我们提出了RS$^3$-Prune,这是一种无训练的标记修剪方法,它在现有的视频目标分割(VOS)网络之上实现了一小组推理时间钩子。现代VOS模型已经趋向于一种共同的、昂贵的设计:图像编码器为每一帧生成一个稠密的标记网格,而内存库则在所有先前处理的帧中累积这些标记,以便为未来的预测提供条件。随着视频变得更长,结果的标记预算决定了每帧的延迟和峰值GPU内存。因此,这些模型在一些用例中表现不佳,例如——长格式视频或在内存受限的加速器上的实时部署。在这项工作中,我们认为压缩内存库VOS的正确方向是标记预算本身。RS$^3$-Prune在任意内存库VOS管道中的两个精确位置操作:在图像编码器与内存注意力读取之间的边界,我们限制参与跨帧注意力的查询仅为一个小的、几何信息驱动的子集;以及在内存编码器与内存库之间的边界,我们限制进入内存库的标记仅为位于对象空间范围内的标记。在各种已建立的基准测试中,RS$^3$-Prune实现了高达$38.8\%$的FPS加速,并减少了$13.1\\%$的峰值内存使用,同时在与未修改的VOS网络相比时保持了竞争力的$ ext{J}$&$ ext{F}$。
cs.CV / 101 / 2608.22532

SymmAdapt: Symmetrical Flow Matching for Source-Free Domain Adaptation in Medical Image Segmentation

SymmAdapt:用于医学图像分割的无源领域适应的对称流匹配
Grossman, Tal, Cahan, Noa, Greenspan, Hayit
Abstract
Domain shift across imaging modalities and acquisition sites remains a significant barrier to the clinical deployment of segmentation models. Source-free unsupervised domain adaptation (SFUDA) addresses this by adapting a pretrained model to an unlabeled target domain without requiring access to sensitive source data. We introduce a novel SFUDA framework built on Symmetrical Flow Matching, a unified generative model that segments an input image and synthesizes a source-like image from a mask within the same learned flow. By initializing inference from a domain-agnostic Gaussian origin, the model preserves structural consistency across domains and grounds predictions in learned anatomy rather than shifted texture statistics. Our pipeline leverages this symmetry to generate reliable pseudo-labels and corresponding source-like synthetic images from unlabeled target data, creating a generative replay buffer that anchors source knowledge during a generative self-training stage that fine-tunes on a joint set of real target and synthetic source-like images. We evaluate on abdominal multi-organ and cardiac segmentation, covering cross-modality MRI<->CT shifts, and multi-site prostate segmentation. Our approach outperforms SFUDA baselines and is competitive with conventional UDA methods.
Chinese Translation
成像模态和采集地点之间的领域转移仍然是分割模型临床应用的一个重大障碍。无源无监督领域适应(SFUDA)通过将预训练模型适应于无标签目标领域,而无需访问敏感的源数据来解决这一问题。我们提出了一种基于对称流匹配的新型SFUDA框架,这是一种统一的生成模型,能够对输入图像进行分割,并从同一学习流中的掩模合成源样图像。通过从与领域无关的高斯起点初始化推理,该模型在不同领域之间保持结构一致性,并将预测基于学习到的解剖结构,而不是转移的纹理统计。我们的流程利用这种对称性,从无标签目标数据生成可靠的伪标签和相应的源样合成图像,创建一个生成重放缓冲区,在生成自我训练阶段锚定源知识,该阶段在真实目标和合成源样图像的联合集上进行微调。我们在腹部多脏器和心脏分割上进行了评估,涵盖了跨模态MRI<->CT转移和多地点前列腺分割。我们的方法优于SFUDA基线,并与传统的无监督领域适应(UDA)方法具有竞争力。
cs.CV / 102 / 2608.22586

Vision-Language Models for Occupational Physical Exposure Assessment: Estimating External Hand Forces in Manual Material Handling Tasks from RGB Video

用于职业身体暴露评估的视觉-语言模型:从RGB视频中估计手部外力在手动物料搬运任务中的应用
Rajabi, Mohammad Sadra, Ojelade, Aanuoluwapo, Kim, Sunwook, Nussbaum, Maury A.
Abstract
External hand forces are important inputs to biomechanical analyses of occupational physical exposure and injury risk, yet continuous force measurements during manual material handling (MMH) typically requires instrumented objects or specialized sensing. We evaluated a vision-language model (VLM)-based pipeline that combines task-specific textual cues, visual representations, and known box mass to estimate dynamic, triaxial, bilateral external hand forces from RGB video. Thirty-five healthy young adults performed five MMH tasks involving lifting, carrying, pushing, and pulling with box masses of 6, 9, and 12 kg. The pipeline used text-guided localization of participant and handled-object regions of interest (ROIs), pretrained vision-transformer feature extraction, and transformer-based temporal regression. Performance was evaluated using leave-one-subject-out validation across seven camera-view conditions (three single-view and four multi-view conditions) and four ROI strategies. Overall, root mean square error was ~4.7-5.6 N for the horizontal and mediolateral force components and ~10.6-11.0 N for the vertical component. Including the handled object as a second ROI generally improved force estimation, with some of the largest benefits under single-camera conditions, whereas pixel-level segmentation provided little additional improvement. Multi-camera capture provided the clearest benefit for peak-force estimation, particularly for the vertical component, whereas differences in overall frame-level error among camera configurations were comparatively modest. These findings demonstrate the feasibility of estimating continuous, bilateral, directional hand-force estimates from RGB video and known load mass without requiring sensors on the worker or handled objects as model inputs, supporting the development of more scalable occupational physical exposure and risk assessments.
Chinese Translation
手部外力是职业身体暴露和伤害风险生物力学分析的重要输入,然而在手动物料搬运(MMH)过程中进行连续的力测量通常需要仪器化物体或专用传感器。我们评估了一种基于视觉-语言模型(VLM)的流程,该流程结合了特定任务的文本线索、视觉表征和已知的箱体质量,以从RGB视频中估计动态的三轴双侧手部外力。三十五名健康年轻成年人执行了五个MMH任务,涉及6、9和12公斤的箱体质量,包括提起、搬运、推和拉。该流程使用文本引导的参与者和处理物体的兴趣区域(ROI)定位、预训练的视觉变换器特征提取以及基于变换器的时间回归。通过在七个摄像机视角条件下(包括三个单视角和四个多视角条件)以及四种ROI策略进行留一法验证来评估性能。总体而言,水平和内外侧力分量的均方根误差约为4.7-5.6 N,垂直分量的均方根误差约为10.6-11.0 N。将处理物体作为第二个ROI通常改善了力的估计,尤其在单摄像机条件下效果显著,而像素级分割提供的额外改进有限。多摄像机捕捉在峰值力估计中提供了最明显的好处,特别是对于垂直分量,而不同摄像机配置之间的整体帧级误差差异相对较小。这些发现表明,从RGB视频和已知负载质量中估计连续的双侧方向手部力是可行的,而无需在工人或处理物体上使用传感器作为模型输入,支持更具可扩展性的职业身体暴露和风险评估的发展。
cs.CV / 103 / 2608.22617

AI-based worker guidance in assembly and disassembly operations using multimodal ego/exo-centric data capture and structured task knowledge

基于人工智能的装配和拆卸操作中的工人指导:利用多模态自我/外部中心数据捕获和结构化任务知识
Chavan, Vivek, Krüger, Jörg
Abstract
Assembly and disassembly processes rely on expert knowledge that is difficult to document, reuse, and transfer. This paper presents a data-centric approach for extracting structured task knowledge from expert demonstrations using egocentric and exocentric recordings. Temporal and multimodal information from video and narration is jointly encoded to derive structured task representations that enable procedural documentation and context-aware worker guidance. The approach is evaluated on a real-world disassembly case study, demonstrating that video-based representations capture procedural structure and execution context beyond static image-based methods. The results highlight the potential of egocentric video understanding for repair, training, and circular manufacturing applications. Project website: https://indego-assistant.github.io/
Chinese Translation
装配和拆卸过程依赖于难以记录、重用和转移的专家知识。本文提出了一种以数据为中心的方法,通过自我中心和外部中心的录音,从专家演示中提取结构化任务知识。视频和叙述中的时间和多模态信息被联合编码,以推导出结构化任务表示,从而实现程序文档编制和上下文感知的工人指导。该方法在一个真实的拆卸案例研究中进行了评估,结果表明,基于视频的表示能够捕捉程序结构和执行上下文,超越了静态图像方法。结果突显了自我中心视频理解在维修、培训和循环制造应用中的潜力。项目网站:https://indego-assistant.github.io/
cs.CV / 104 / 2608.22637

OmniCAD: A Large-Scale Benchmark for 3D Spatial Reasoning in Robotics Assemblies

OmniCAD:用于机器人装配中3D空间推理的大规模基准测试
Wang, Mingjia, Lu, Taiting, Dong, Ziwei, Bei, Sisong, Zeng, Jingying, Liu, Runze, Lin, Kaiyuan, Pan, Hongxing, Zhang, Kai, Hou, Yizheng, Zheng, Yangshoudu, Guo, Chenchen, Meng, Weiyuan, Lyu, Shubin, Zheng, Zhijun, Wang, Dexu, Bai, Xinyu, Qian, Shurui, Zhangzixin, Pan, Mengyu, Shi, Guoliang, Ma, Ling, Yang, Yifan, He, Qi, Chen, Yi-Chao, Jin, Yincheng, Chen, Sung-Liang, Gowda, Mahanth
Abstract
Recent vision-language models (VLMs) show strong capabilities in robotic perception and spatial reasoning, yet their ability to reason about complex mechanical assemblies remains underexplored. We introduce OmniCAD, a large-scale benchmark for assembly-aware 3D spatial reasoning across diverse industrial systems, including robotic mechanisms, automotive components, aerospace structures, and agricultural machinery. OmniCAD contains 25k mechanical assemblies, with an average of 12 parts per assembly and 21 types of mate relationships. Each assembly includes a human-verified ground-truth 3D model and renderings from 20 viewpoints. The benchmark evaluates three capabilities: (1) component-level 3D spatial reasoning, requiring prediction of part positions and orientations; (2) part-to-part relational reasoning, requiring identification of mating relationships and assembly constraints; and (3) tool-augmented agentic reasoning, where models iteratively select viewpoints, inspect visual evidence, and refine predictions. Experiments show that current VLMs struggle with industrial assembly reasoning, often producing inaccurate poses, invalid mating relationships, part interpenetration, and degraded performance as assembly complexity increases. We will open-source the benchmark, evaluation code, and tool interfaces to support research on accurate, physically valid, and scalable 3D assembly reasoning.
Chinese Translation
最近的视觉-语言模型(VLMs)在机器人感知和空间推理方面展现出强大的能力,但它们在复杂机械装配方面的推理能力仍然未得到充分探索。我们推出了OmniCAD,这是一个针对多种工业系统(包括机器人机制、汽车组件、航空航天结构和农业机械)的装配感知3D空间推理的大规模基准测试。OmniCAD包含25,000个机械装配,每个装配平均有12个零件和21种配合关系。每个装配都包括经过人工验证的真实3D模型和来自20个视角的渲染图。该基准测试评估三种能力:(1)组件级3D空间推理,要求预测零件的位置和方向;(2)零件间关系推理,要求识别配合关系和装配约束;(3)工具增强的代理推理,模型需迭代选择视角、检查视觉证据并细化预测。实验表明,当前的VLM在工业装配推理方面表现不佳,常常产生不准确的姿态、无效的配合关系、零件相互穿透,并且随着装配复杂性的增加,性能下降。我们将开源该基准测试、评估代码和工具接口,以支持对准确、物理有效和可扩展的3D装配推理的研究。
cs.CV / 105 / 2608.22655

Multiple View Neural Regression of a Facial Shape Model

面部形状模型的多视角神经回归
Li, Xiang
Abstract
Creating re-topologized 3D facial meshes is essential for high-quality facial animation but remains labor-intensive and time-consuming. This dissertation explores more efficient approaches for capturing production-ready facial meshes through: (1) the development of VarIS, a custom light sphere for capturing high-resolution stereo geometry and reflectance maps; (2) analysis of camera parameters affecting automatic 2D and 3D landmarking; (3) synthetic-data methods for training neural face regression; and (4) techniques for improving neural multi-view face-shape regression. While VarIS enables photorealistic face capture, its operational and processing costs motivate a more scalable approach. A deep learning framework is therefore proposed to directly predict re-topologized facial meshes from synthetic multiview images generated with Visage Craft, an in-house physically based rendering system using an Appearance 3D Morphable Model (A3DMM). The system produces standardized meshes ready for rigging and animation with minimal human supervision. Results show that incorporating accurate camera intrinsics and extrinsics improves landmark accuracy and geometric consistency, while 3D landmark regularization further improves reconstruction quality.
Chinese Translation
创建重新拓扑的3D面部网格对于高质量面部动画至关重要,但仍然是一个劳动密集且耗时的过程。本论文探讨了通过以下方式捕捉生产就绪的面部网格的更有效方法:(1) 开发VarIS,一种用于捕捉高分辨率立体几何和反射率图的定制光球;(2) 分析影响自动2D和3D标记的相机参数;(3) 用于训练神经面部回归的合成数据方法;(4) 改进神经多视角面部形状回归的技术。虽然VarIS能够实现逼真的面部捕捉,但其操作和处理成本促使我们寻求更具可扩展性的方法。因此,提出了一种深度学习框架,直接从使用基于物理的渲染系统Visage Craft生成的合成多视角图像中预测重新拓扑的面部网格,该系统使用外观3D可变形模型(Appearance 3D Morphable Model, A3DMM)。该系统生成标准化网格,准备进行绑定和动画,且对人工监督的需求极小。结果表明,准确的相机内参和外参的结合提高了标记的准确性和几何一致性,而3D标记正则化进一步提升了重建质量。
cs.CV / 106 / 2608.22665

Hyperbolic Hierarchical Clustering for Visual Representation Learning

用于视觉表示学习的双曲层次聚类
Wei, Jianan, Chen, Guikun, Weng, Zhiyuan, Guo, Chunchao, Wang, Yujia, Wang, Wenguan
Abstract
We investigate the token mixer in vision backbones by revisiting clustering, one of the most classic approaches in machine learning. An effective token mixer is a fundamental component of modern vision backbones like vision Transformers, facilitating information exchange between image patches. Mainstream token mixers, which rely on convolution, attention, MLP, or their hybrids, primarily focus on navigating the trade-off between accuracy and computational cost. However, a significant drawback of these methods is their black-box nature; their encoding process is opaque and lacks interpretability. Diverging from these opaque designs, we introduce ClusterMixer, a transparent token mixer that is grounded in a clustering paradigm and interpretable by design. ClusterMixer explicitly formulates the token mixing process through a hierarchical clustering mechanism. To model the natural, tree-like relationships inherent in visual data, the clustering is performed in hyperbolic space, which is well-suited for embedding hierarchies with low distortion. Building on this innovation, we present HCFormer, a new backbone architecture that integrates ClusterMixer with a series of meticulously designed clustering strategies to ensure robust performance across tasks. Extensive experiments demonstrate that HCFormer consistently outperforms its counterparts across diverse tasks, including image classification, object detection, instance segmentation, and semantic segmentation. Considering its transparency and efficacy, we hope HCFormer can facilitate a paradigm shift toward interpretable backbones.
Chinese Translation
我们通过重新审视聚类这一机器学习中最经典的方法之一,研究视觉骨干网络中的标记混合器。有效的标记混合器是现代视觉骨干网络(如视觉变换器,Vision Transformers)的基本组成部分,促进了图像块之间的信息交换。主流的标记混合器依赖于卷积、注意力机制、多层感知器(MLP)或它们的混合,主要集中在准确性与计算成本之间的权衡。然而,这些方法的一个显著缺点是其黑箱特性;它们的编码过程不透明,缺乏可解释性。与这些不透明的设计不同,我们引入了ClusterMixer,这是一种基于聚类范式的透明标记混合器,设计上具有可解释性。ClusterMixer通过层次聚类机制明确地制定了标记混合过程。为了建模视觉数据中固有的自然树状关系,聚类在双曲空间中进行,这非常适合嵌入低失真层次结构。在此创新基础上,我们提出了HCFormer,一种新的骨干架构,将ClusterMixer与一系列精心设计的聚类策略相结合,以确保在各项任务中的稳健性能。大量实验表明,HCFormer在图像分类、目标检测、实例分割和语义分割等多种任务中始终优于其对手。考虑到其透明性和有效性,我们希望HCFormer能够促进向可解释骨干网络的范式转变。
cs.CV / 107 / 2608.22679

Contextrast++: Robust Multi-Scale Contextual Contrastive Learning for Semantic Segmentation

Contextrast++:用于语义分割的鲁棒多尺度上下文对比学习
Sung, Changki, Lim, Hyungtae, Kim, Wanhee, Seo, Youngwoo, Myung, Hyun
Abstract
Semantic segmentation has rapidly advanced with deep learning; however, challenges remain in effectively capturing local and global contexts as well as addressing the long-tailed distribution problem. To tackle these issues, we present Contextrast++, a robust contrastive learning method for semantic segmentation that improves multi-scale feature integration and mitigates class imbalance issues. Our method consists of two key components: 1) contextual contrastive learning (CCL) and 2) boundary-aware negative (BANE) sampling. CCL includes three subcomponents: adaptive fusion module, pixel-to-anchor (PA) loss, and anchor-to-anchor (AA) loss. The adaptive fusion module dynamically balances local and global feature integration, resulting in a more context-aware representation. While the PA loss leverages the fused multi-scale features to improve feature representation learning, the AA loss focuses on addressing the long-tailed distribution problem by utilizing a memory bank that stores a fixed number of class-balanced representative anchors. Meanwhile, BANE sampling enhances segmentation precision by selecting hard negatives from misclassified boundary regions, which refines fine-grained details during contrastive learning. As verified in extensive experiments using public datasets, we demonstrate that Contextrast++ substantially improves semantic segmentation performance over existing contrastive learning-based state-of-the-art approaches, while introducing no additional computational overhead during inference.
Chinese Translation
语义分割在深度学习的推动下迅速发展;然而,在有效捕捉局部和全局上下文以及解决长尾分布问题方面仍然面临挑战。为了解决这些问题,我们提出了Contextrast++,一种用于语义分割的鲁棒对比学习方法,旨在改善多尺度特征集成并缓解类别不平衡问题。我们的方法由两个关键组件组成:1)上下文对比学习(CCL)和2)边界感知负样本(BANE)采样。CCL包括三个子组件:自适应融合模块、像素到锚点(PA)损失和锚点到锚点(AA)损失。自适应融合模块动态平衡局部和全局特征的集成,从而生成更具上下文感知的表示。PA损失利用融合的多尺度特征来改善特征表示学习,而AA损失则通过利用存储固定数量类别平衡代表性锚点的记忆库,专注于解决长尾分布问题。同时,BANE采样通过从误分类的边界区域选择困难负样本来提高分割精度,这在对比学习过程中细化了细粒度细节。通过在公共数据集上进行的大量实验验证,我们展示了Contextrast++在语义分割性能上显著优于现有基于对比学习的最先进方法,并且在推理过程中没有引入额外的计算开销。
cs.CV / 108 / 2608.22690

MorphoCLIP: Text-Supervised Contrastive Learning for Perturbation Matching in Cell Painting Images

MorphoCLIP:用于细胞绘画图像扰动匹配的文本监督对比学习
Ilyosbekov, Sukhrobbek, Gajjar, Shubham, Jin, Rongfei
Abstract
Cell Painting microscopy captures how cells change after a chemical or genetic perturbation. Connecting these images to the perturbations that produced them could make large imaging screens easier to search and interpret, but the task remains difficult because biological effects are subtle and technical variation is substantial. We introduce MorphoCLIP, a contrastive model that links Cell Painting profiles with text descriptions of compounds, CRISPR knockouts, and ORF overexpressions. The model keeps its vision and language backbones frozen and trains only a compact cross-channel module and projection layers, so it can be trained on a single consumer GPU. On held-out CPJUMP1 data, MorphoCLIP searches in both directions: from a cell image to its perturbation description and from a description to matching cell images. In both cases, a correct match appears among the top ten results much more often than expected by chance. Adding a replicate-alignment loss makes profiles from repeated experiments more consistent, although this improvement does not yet translate into reliable gene-compound matching. Gene-aware labels and plate correction also show no consistent retrieval benefit. These findings suggest that text supervision can help organize chemical and genetic Cell Painting data. Matching compounds with genetic perturbations, however, remains an open problem.
Chinese Translation
细胞绘画显微镜捕捉了细胞在化学或遗传扰动后如何变化。将这些图像与产生它们的扰动相连接,可以使大规模成像筛选更易于搜索和解释,但由于生物效应微妙且技术变异显著,这一任务仍然困难。我们提出了MorphoCLIP,这是一种对比模型,将细胞绘画特征与化合物、CRISPR基因敲除和ORF过表达的文本描述联系起来。该模型保持其视觉和语言骨干网络不变,仅训练一个紧凑的跨通道模块和投影层,因此可以在单个消费级GPU上进行训练。在保留的CPJUMP1数据上,MorphoCLIP在两个方向上进行搜索:从细胞图像到其扰动描述,以及从描述到匹配的细胞图像。在这两种情况下,正确匹配的结果出现在前十名结果中的频率远高于随机预期。添加重复对齐损失使重复实验的特征更加一致,尽管这一改进尚未转化为可靠的基因-化合物匹配。基因感知标签和板块校正也未显示出一致的检索优势。这些发现表明,文本监督可以帮助组织化学和遗传细胞绘画数据。然而,将化合物与遗传扰动匹配仍然是一个未解决的问题。
cs.CV / 109 / 2608.22692

Hybrid Generative-Discriminative Object Placement

混合生成-判别对象放置
Zhou, Siyuan, Niu, Li
Abstract
As an important operation of image composition, object placement aims to predict the plausible placement (location, scale) for the inserted foreground object. Previous object placement methods can be divided into generative methods and discriminative methods, both of which cannot balance efficiency and effectiveness well. In this work, we propose a semi-generative method in the middle ground between them. In particular, we assign uniformly distributed anchors on the background. Then, we fuse foreground and background features to predict the rationality score for each anchor and predict plausible placement sets for positive anchors. Extensive experiments on the OPA dataset show that our method can strike a good balance between efficiency and effectiveness.
Chinese Translation
作为图像合成的重要操作,对象放置旨在预测插入前景对象的合理放置(位置、尺度)。之前的对象放置方法可以分为生成方法和判别方法,但两者在效率和有效性之间的平衡都不理想。在本研究中,我们提出了一种介于两者之间的半生成方法。具体而言,我们在背景上分配均匀分布的锚点。然后,我们融合前景和背景特征,以预测每个锚点的合理性评分,并为正锚点预测合理的放置集。在 OPA 数据集上的大量实验表明,我们的方法能够在效率和有效性之间取得良好的平衡。
cs.CV / 110 / 2608.22723

LoViF 2026 The First Challenge on Unified Removal of Raindrops and Reflections: Methods and Results

LoViF 2026 第一次统一去除雨滴和反射的挑战:方法与结果
He, Zewei, Tong, Xi, Chen, Yu, Liu, Xingyu, Li, Xin, Wang, Zepeng, Hu, Jiagao, Li, Fuhao, Chen, Yuxuan, Wang, Fei, Zhou, Daiguo, Yi, Minmin, Zhang, Chuanrui, Zhang, Liwen, Jeong, Yeongjin, Cho, Hyunjin, Lee, Jiwon, Kim, Minsang, Soh, Jae Woong, Jiang, Jin-Hui, Jian, Rong-Lin, Hsu, Chih-Chung, Oh, Youngjin, Kwon, Junhyeong, Park, Junyoung, Park, Jae Hyun, Lee, Sung Ju, Cho, Nam Ik, Shukla, Vishwajeet, Baurai, Himanshu, Zhang, Zhiqi, Jiang, Kui, Yu, Zhaocheng, Li, Runzhe, Fan, Dawei, Li, Hao, Zhang, Zhanshuo, Ji, Fan, Li, Jiangmeng, Tang, Xiongxin, Xu, Fanjiang, Sun, Shangquan, Duong, Anh-Kiet, Gomez-Krämer, Petra, Carozza, Jean-Michel, Zhang, Ruibo, Hong, Dexiang, Liu, Xinyan, Tang, Shengeng, Chen, Weidong, Weng, Tzu-Hsuan, Sun, Min-Te
Abstract
This workshop paper comprehensively reviews the First Challenge on Unified Removal of Raindrops and Reflections. The challenge aims to address a frequently encountered practical problem in the field of autonomous driving, i.e., raindrop-reflection composite degradation on rainy days. This competition attracted 149 registered participants and received 12 valid final submissions with corresponding fact sheets, significantly contributing to the progress of unified removal of raindrops and reflections. All the methods are developed and evaluated on our real-shot RainDrop and ReFlection (RDRF) dataset. A detailed analysis of the submitted methods and corresponding results is provided in this report, which highlights effective approaches and provides interesting insights for future research.
Chinese Translation
本文综述了第一次统一去除雨滴和反射的挑战。该挑战旨在解决自动驾驶领域中一个常见的实际问题,即雨天雨滴与反射的复合降解。此次竞赛吸引了149名注册参与者,并收到了12份有效的最终提交及相应的事实表,显著推动了雨滴与反射的统一去除技术的发展。所有方法均在我们的真实拍摄的雨滴与反射(RainDrop and ReFlection, RDRF)数据集上进行开发和评估。本文提供了对提交方法及其相应结果的详细分析,突出了有效的方法并为未来研究提供了有趣的见解。
cs.CV / 111 / 2608.22740

Seeing the Unseen: Semantic-in-Gaussian for Sparse-View 3D Generalization

洞察未见:用于稀疏视角三维泛化的语义高斯表示(Semantic-in-Gaussian)
Bai, Zeyang, Wang, Yunpeng, Wang, Yunbiao, Xiao, Jun
Abstract
Generalizable 3D Gaussian Splatting (G-3DGS) has emerged as a promising approach for novel view synthesis undersparse-view settings. However, existing frameworks remain restricted by pixel-aligned Gaussian estimation, whichstruggles in partially observed or occluded regions and often leads to incomplete surfaces or structural collapse. Toaddress these challenges, we propose SeeU (Seeing the Unseen), a novel G-3DGS framework. We frame its core design asSemantic-in-Gaussian: semantic-conditioned refinement in Gaussian space. Specifically, we introduce a Cross-viewEntropy-Aware (CEA) module that aggregates multi-view semantic and geometric cues into compact embeddings. Theseembeddings guide the Conditional Gaussian Transformer, which applies residual updates to coarse Gaussians, helpingrecover under-constrained regions of partially observed structures while preserving surface consistency. Comprehensiveexperiments on multiple benchmarks demonstrate that SeeU consistently improves rendering quality and structuralcompleteness while retaining efficient feed-forward inference. Especially under challenging extrapolation settings,SeeU achieves an average improvement of 2.44 dB in PSNR compared to recent SOTA G-3DGS methods.
Chinese Translation
可泛化的三维高斯点云渲染(Generalizable 3D Gaussian Splatting,G-3DGS)作为一种在稀疏视角条件下进行新视角合成的有前景方法,近年来备受关注。然而,现有框架仍受限于基于像素对齐的高斯估计,这在部分观测或遮挡区域表现不佳,常导致表面不完整或结构塌陷。为解决这些挑战,我们提出了SeeU(Seeing the Unseen)——一种新颖的G-3DGS框架。其核心设计被定义为Semantic-in-Gaussian:即在高斯空间中进行语义条件的细化。具体而言,我们引入了一个跨视角熵感知(Cross-view Entropy-Aware,CEA)模块,该模块将多视角的语义与几何线索聚合为紧凑的嵌入表示。这些嵌入引导条件高斯变换器(Conditional Gaussian Transformer)对粗糙高斯进行残差更新,帮助恢复部分观测结构中受约束不足的区域,同时保持表面一致性。在多个基准测试上的全面实验表明,SeeU在提升渲染质量和结构完整性的同时,依然保持了高效的前向推理能力。尤其在具有挑战性的外推设置下,SeeU相比近期最先进的G-3DGS方法,PSNR平均提升了2.44 dB。
cs.CV / 112 / 2608.22757

Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

Object-Uni:一种面向对象的空间理解与可控生成的统一模型
Tan, Mining, Wang, Yinuo, Zhou, Ziqi, Quan, Weize, Li, Sifei, Chen, Jingdong, Zheng, DanDan, Wang, Libin, Dong, Weiming
Abstract
Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects in natural language, but they struggle to precisely represent continuous object poses and generate geometrically consistent images under target viewpoints. To mitigate this, we propose \emph{Object-Uni}, a unified model for object-centric spatial understanding and controllable generation. Specifically, we formulate object-centric spatial intelligence as a unified problem connecting pose perception, spatial reasoning, pose-conditioned generation, and object-centric novel view synthesis. We treat object pose as an explicit geometric variable shared by understanding and generation, rather than merely a prediction label or control signal. To make pose usable by multimodal large language models, we propose a viewpoint-based orientation abstraction that maps orientation into structured viewpoint descriptions while preserving continuous geometric supervision. We further construct an object-centric spatial benchmark (UniSpatial-80K) and train a unified model with an object-token-grounded pose anchor to associate each instance with its pose state. Experiments show that our model improves object-level pose understanding and pose-controllable generation, moving unified models from describing objects toward manipulating spatial states.
Chinese Translation
视觉理解与生成的统一模型取得了快速进展,但仍然缺乏理解和操控对象实例空间状态的能力。现有模型能够用自然语言描述对象,但在精确表示连续对象姿态和在目标视点下生成几何一致的图像方面存在困难。为此,我们提出了 extit{Object-Uni},一种面向对象的空间理解与可控生成的统一模型。具体而言,我们将面向对象的空间智能形式化为一个统一问题,连接姿态感知、空间推理、姿态条件生成和面向对象的新视图合成。我们将对象姿态视为理解与生成共享的显式几何变量,而不仅仅是预测标签或控制信号。为了使姿态能够被多模态大型语言模型使用,我们提出了一种基于视点的方向抽象,将方向映射为结构化的视点描述,同时保留连续的几何监督。我们进一步构建了一个面向对象的空间基准(UniSpatial-80K),并训练了一个统一模型,利用对象标记锚点将每个实例与其姿态状态关联起来。实验表明,我们的模型提高了对象级姿态理解和姿态可控生成的能力,使统一模型从描述对象转向操控空间状态。
cs.CV / 113 / 2608.22760

ByteAction: Byte-space Action Recognition Foundation Model

ByteAction:字节空间动作识别基础模型
Li, Fangcheng, Yu, Zhen, Wu, Kejun, Liu, Qiong, Yang, You
Abstract
Byte-space Action Recognition (BAR) aims to recognize human actions directly from compressed image bitstreams without any pixel decoding. By operating entirely in byte space, BAR is inherently independent of file integrity and pixel-level reconstruction, making it naturally applicable to privacy-sensitive scenarios and robust against bitstream corruption. In this paper, we propose ByteAction, a BAR foundation model that achieves accurate action recognition on corrupted image bitstreams. ByteAction follows a dual-view byte-level recognition framework. It constructs weakly and strongly corrupted bitstream views, which are augmented by Bitstream Pattern Augmentation (BPA) and encoded with a shared ByteFormer backbone. The model is optimized with both classification and corruption consistency objectives. Specifically, we propose Bitstream Pattern Augmentation (BPA), which reshapes one-dimensional byte sequences into two-dimensional byte matrix and applies region-level erasure to encourage the model to learn robust cross-region byte dependencies. We further propose a Corruption Consistency Training strategy that constrains the model to maintain stable predictions across different corruption severities through bidirectional KL divergence. Experiments on the image bitstream from Stanford40, PPMI, and PASCAL VOC 2012 Action demonstrate that ByteAction achieves state-of-the-art corruption robustness across all scenarios while maintaining competitive intact bitstream performance.
Chinese Translation
字节空间动作识别(BAR)旨在直接从压缩的图像比特流中识别人的动作,而无需任何像素解码。通过完全在字节空间中操作,BAR 本质上独立于文件完整性和像素级重建,使其自然适用于隐私敏感场景,并对比特流损坏具有鲁棒性。本文提出了 ByteAction,一个 BAR 基础模型,能够在损坏的图像比特流上实现准确的动作识别。ByteAction 遵循双视图字节级识别框架。它构建了弱损坏和强损坏的比特流视图,这些视图通过比特流模式增强(BPA)进行增强,并使用共享的 ByteFormer 主干进行编码。该模型通过分类和损坏一致性目标进行优化。具体而言,我们提出了比特流模式增强(BPA),它将一维字节序列重塑为二维字节矩阵,并应用区域级擦除以鼓励模型学习鲁棒的跨区域字节依赖关系。我们进一步提出了一种损坏一致性训练策略,通过双向 KL 散度约束模型在不同损坏严重性下保持稳定的预测。对 Stanford40、PPMI 和 PASCAL VOC 2012 动作的图像比特流的实验表明,ByteAction 在所有场景中实现了最先进的损坏鲁棒性,同时保持了竞争性的完整比特流性能。
cs.CV / 114 / 2608.22773

LagrangeGS: Non-Conservative Lagrangian System on Dynamic 3D Gaussian Splatting

LagrangeGS:动态3D高斯点云上的非保守拉格朗日系统
Sato, Shogo, Kaneko, Takuhiro, Takeda, Shoichiro, Shimada, Tomoyasu, Inoue, Riku, Murasaki, Kazuhiko, Tanida, Ryuichi
Abstract
Dynamic 3D Gaussian Splatting (3DGS) achieves photorealistic reconstruction of time-varying scenes, and recent physics-aware extensions improve extrapolation by explicitly predicting velocity fields. However, these extensions merely fit vector fields to visual deformations without satisfying Lagrangian mechanics, leading to three major issues: (i) physically inconsistent trajectories, (ii) lack of time-reversibility, and (iii) geometric collapse during long-term extrapolation. In this paper, we propose LagrangeGS, which formulates dynamic 3DGS as a non-conservative Lagrangian system. While this Lagrangian formulation fundamentally solves (i), a direct application of general LNNs to dynamic 3DGS requires a large velocity-Hessian inversion for millions of Gaussian particles. To overcome this computational bottleneck, we approximate the velocity-Hessian as an identity matrix, decoupling particle dynamics for computational tractability. For (ii), we restrict the non-conservative forces to be explicitly time independent, enabling consistent backward integration. Finally, to address (iii), we introduce local rigid alignment that regularizes particle trajectories. Extensive evaluations on dynamic scene benchmarks demonstrate that LagrangeGS enables stable long-term extrapolation, consistent time reversal, and counterfactual physics-based editing without retraining.
Chinese Translation
动态3D高斯点云(3DGS)实现了对时间变化场景的照片级真实重建,最近的物理感知扩展通过显式预测速度场来改善外推。然而,这些扩展仅仅是将向量场拟合到视觉变形上,而未能满足拉格朗日力学,导致三个主要问题:(i)物理上不一致的轨迹,(ii)缺乏时间可逆性,以及(iii)在长期外推过程中几何崩溃。本文提出了LagrangeGS,将动态3DGS公式化为一个非保守拉格朗日系统。虽然这种拉格朗日公式从根本上解决了(i)的问题,但将一般的LNN直接应用于动态3DGS需要对数百万个高斯粒子进行大规模的速度-海森矩阵反演。为了解决这一计算瓶颈,我们将速度-海森矩阵近似为单位矩阵,从而解耦粒子动力学以提高计算的可行性。针对(ii),我们限制非保守力显式地与时间无关,从而实现一致的反向积分。最后,为了解决(iii),我们引入局部刚性对齐来规范粒子轨迹。在动态场景基准上的广泛评估表明,LagrangeGS能够实现稳定的长期外推、一致的时间反转,以及无需重新训练的反事实基于物理的编辑。
cs.CV / 115 / 2608.22780

Can We Perform Online RL for Image Editing without Editing Rewards?

我们能否在没有编辑奖励的情况下进行图像编辑的在线强化学习?
Ma, Qichao, Cheng, Jikang, Liang, Ling, Yu, Zhaofei, Huang, Tiejun, Yan, Renye
Abstract
Reinforcement learning (RL) enables direct preference optimization for image editing through editing-specific rewards, which remain less developed due to costly triplet supervision and complex task-dependent calibration. In contrast, text-to-image (T2I) generation benefits from a mature and diverse reward ecosystem spanning semantic alignment, aesthetics, realism, glyph shape, and other visual preferences. Extending this ecosystem to image editing would substantially broaden the range of visual preferences accessible to RL-based optimization, prompting the central question: \emph{Can We Perform Image Editing RL without Editing Rewards?} In this paper, we argue that the standard image editing dimensions have potential to be mapped to the T2I reward space: image quality can transfer directly, prompt following can be aligned through a description of the desired visual state, and reference consistency admits a coarse semantic conversion by encoding the source content to preserve. However, editing instructions specify relative changes, whereas T2I rewards require self-contained target descriptions; moreover, semantically valid captions from generic vision-language models may be incompatible with the frozen reward. Hence, we further introduce Lever-Edit, a two-stage framework that learns a reward-aligned captioner for counterfactual target descriptions, freezes it, and optimizes the editing policy solely with the transferred T2I reward. Experiments show competitive editing alignment and source preservation against editing-reward-based fine-tuning, while outperforming intuitive transfer baselines.
Chinese Translation
强化学习(RL)通过特定于编辑的奖励实现图像编辑的直接偏好优化,但由于昂贵的三元组监督和复杂的任务依赖校准,这一领域的发展仍然较为滞后。相比之下,文本到图像(T2I)生成受益于成熟且多样的奖励生态系统,涵盖语义对齐、美学、真实感、字形形状及其他视觉偏好。将这一生态系统扩展到图像编辑将大大拓宽可供基于RL的优化访问的视觉偏好的范围,从而引发中心问题: extit{我们能否在没有编辑奖励的情况下进行图像编辑的强化学习?} 在本文中,我们论证了标准图像编辑维度有潜力映射到T2I奖励空间:图像质量可以直接转移,遵循提示可以通过对所需视觉状态的描述进行对齐,而参考一致性则通过编码要保留的源内容进行粗略的语义转换。然而,编辑指令指定的是相对变化,而T2I奖励要求自包含的目标描述;此外,通用视觉-语言模型生成的语义有效标题可能与冻结的奖励不兼容。因此,我们进一步提出了Lever-Edit,一个两阶段框架,学习一个与奖励对齐的标题生成器用于反事实目标描述,冻结该生成器,并仅使用转移的T2I奖励优化编辑策略。实验结果显示,与基于编辑奖励的微调相比,Lever-Edit在编辑对齐和源内容保留方面表现出竞争力,同时超越了直观转移基线。
cs.CV / 116 / 2608.22785

OmicSync: Reliability-Aware Spatial Multi-Omics Clustering with Evidence-Constrained LLM Reasoning

OmicSync:基于证据约束大语言模型推理的可靠性感知空间多组学聚类
Sadia, Rabeya Tus, Ye, Qiang, Cheng, Qiang
Abstract
Spatial multi-omics technologies jointly profile gene expression, surface proteins, and histology at each tissue spot, yet most spatial domain discovery methods provide only cluster assignments, without indicating assignment reliability, modality contributions, or why a domain decision should be trusted. We present OmicSync, a reliability-aware spatial multi-omics framework that couples unsupervised domain clustering with evidence-constrained LLM reasoning using model-derived per-spot signals, including assignment confidence, epistemic routing uncertainty, and modality-routing weights. These signals are converted into structured evidence dictionaries and used to generate standard, stepwise, counterfactual, contrastive, and uncertainty-focused explanations. OmicSync integrates a KAN-GCN backbone with spatial encoding, cross-modal fusion, uncertainty-aware routing, cell-type supervision, and missing-modality imputation. We further introduce OmicSync-R, which closes the reasoning-clustering loop by using automatically computed reasoning-quality scores as REINFORCE rewards, allowing reasoning coherence to shape the latent structure without backpropagating through the language model. Across four 10x CytAssist FFPE spatial proteomics benchmarks, OmicSync achieves the best average rank on Human Tonsil (1.44), Glioblastoma (1.78), and Tonsil Add-on (1.22), and second-best on Human Breast Cancer (2.33). OmicSync-R further improves ARI on Human Breast Cancer from 45.73 to 46.72 and outperforms existing methods on six of nine clustering metrics. Together, OmicSync and OmicSync-R enable reliability-aware, spot-level auditable spatial domain discovery guided by evidence-constrained reasoning.
Chinese Translation
空间多组学技术能够在每个组织点同时分析基因表达、表面蛋白和组织学信息,但大多数空间域发现方法仅提供聚类分配结果,未能指示分配的可靠性、模态贡献或为何该域决策值得信赖。我们提出了OmicSync,一种可靠性感知的空间多组学框架,将无监督域聚类与基于模型生成的每点信号(包括分配置信度、认知路由不确定性和模态路由权重)的证据约束大语言模型(LLM)推理相结合。这些信号被转换为结构化的证据字典,用于生成标准、分步、反事实、对比和不确定性聚焦的解释。OmicSync整合了具有空间编码、跨模态融合、不确定性感知路由、细胞类型监督和缺失模态插补功能的KAN-GCN骨干网络。我们进一步引入了OmicSync-R,通过使用自动计算的推理质量评分作为REINFORCE奖励,闭合推理与聚类的循环,使推理一致性能够塑造潜在结构,而无需通过语言模型反向传播。在四个10x CytAssist FFPE空间蛋白质组学基准测试中,OmicSync在人类扁桃体(1.44)、胶质母细胞瘤(1.78)和扁桃体附加数据集(1.22)上取得最佳平均排名,在人类乳腺癌(2.33)上排名第二。OmicSync-R进一步将人类乳腺癌的ARI从45.73提升至46.72,并在九项聚类指标中的六项优于现有方法。综上,OmicSync及OmicSync-R实现了基于证据约束推理指导的可靠性感知、可审计的点级空间域发现。
cs.CV / 117 / 2608.22789

GuidedFlow: An Attention-Guided Framework for Anomaly Detection in Additive Manufacturing

GuidedFlow:一种基于注意力引导的增材制造异常检测框架
Paul, Sosmita, Roy, Krishna
Abstract
Additive Manufacturing (AM) plays a vital role in the ongoing industrial revolution. However, quality control remains crucial and challenging due to printing defects or potential cyber-physical intrusions. Image or video-based anomaly detection is a key effort towards addressing these challenges. Various approaches have been explored in this domain, including reconstruction-based, embedding-based, and flow-based methods. Though normalizing flow-based methods address some of the core challenges of unforeseen defects and generalization while maintaining detection performance, existing approaches struggle with tiny/stringing defects common in 3D printing. In a small-data setting, this poses a limitation in generalization. To address these limitations, we propose \textbf{GuidedFlow}, a novel attention-guided normalizing flow model for anomaly detection and localization. GuidedFlow employs a pre-trained ResNet model, fine-tuned on the domain dataset. An attention-guided spatial and temporal flow framework models the dynamics across multiple scales and frames. A Spatio-Temporal Attention Network (SAN) enables the flow model to prioritize relevant contextual cues from input frames. We evaluate GuidedFlow on our AM3D-AD dataset, consisting of benign and anomalous real 3D printed object images and videos. We also conduct a comparative study using the MVTec-AD industrial image anomaly detection dataset. Experimental results demonstrate that GuidedFlow outperforms most of the state-of-the-art models with enhanced detection accuracy and AUROC.
Chinese Translation
增材制造(Additive Manufacturing, AM)在当前的工业革命中发挥着至关重要的作用。然而,由于打印缺陷或潜在的网络物理入侵,质量控制仍然至关重要且具有挑战性。基于图像或视频的异常检测是应对这些挑战的关键努力。在这一领域,已经探索了多种方法,包括基于重建、基于嵌入和基于流的方法。尽管归一化流(normalizing flow)方法在保持检测性能的同时解决了一些不可预见缺陷和泛化的核心挑战,但现有方法在处理3D打印中常见的微小/拉丝缺陷时仍然存在困难。在小数据环境下,这限制了泛化能力。为了解决这些局限性,我们提出了 extbf{GuidedFlow},一种新颖的基于注意力引导的归一化流模型,用于异常检测和定位。GuidedFlow采用预训练的ResNet模型,并在领域数据集上进行微调。一个基于注意力引导的时空流框架建模了多个尺度和帧之间的动态。时空注意力网络(Spatio-Temporal Attention Network, SAN)使流模型能够优先考虑输入帧中的相关上下文线索。我们在包含良性和异常真实3D打印物体图像及视频的AM3D-AD数据集上评估了GuidedFlow。我们还使用MVTec-AD工业图像异常检测数据集进行了比较研究。实验结果表明,GuidedFlow在检测准确性和AUROC方面优于大多数最先进的模型。
cs.CV / 118 / 2608.22795

VersaDB: A High-Performance AI Storage Database for Unifying Mutimodal Datasets

VersaDB:一种高性能的人工智能存储数据库,用于统一多模态数据集
Wang, Cong, Liu, Zelin, Zhang, Yang Luo Ran, Guo, Zhijian, Zhang, Hui, Yu, Fan, Cao, Yanfei, Gu, Naijie, Yu, Jun
Abstract
The AI field has been rapidly developing, leading to the emergence of a large number of AI training datasets of various types. These datasets contain different modalities, including text, images, audio, etc., and may come in various data storage formats. With the advancement of AI hardware, AI computation units like GPUs, TPUs, and NPUs can greatly accelerate the training speed of AI models, which in turn increases the demand for faster data processing. When using existing AI processing frameworks to handle datasets with different modalities and storage formats, processing speeds may be suboptimal due to issues such as data layout and the way users handle the data. Therefore, using a unified database to store multiple data formats can better manage and optimize data access. In this paper, we introduce VersaDB, a database designed specifically for AI datasets with various modalities. We implemented a page-based storage system, separating structured and unstructured data. Additionally, we generated B+ tree-based index files to accelerate data access. VersaDB supports automatic sharding and maintains a hierarchical metadata management system, with corresponding metadata maintained at the page, shard, and global levels, forming the foundation for the efficient operation of the database. We also focused on ease of use by providing APIs for directly converting datasets into VersaDB, as well as APIs for converting popular AI data storage formats (e.g., CSV, TFRecord, .bin) into VersaDB.Our experiments show that using VersaDB can achieve up to 5.35x acceleration and maintain consistent performance across different parallelism levels.
Chinese Translation
人工智能领域正在迅速发展,导致出现大量各种类型的人工智能训练数据集。这些数据集包含不同的模态,包括文本、图像、音频等,并可能采用各种数据存储格式。随着人工智能硬件的进步,GPU、TPU和NPU等人工智能计算单元可以大幅加速人工智能模型的训练速度,从而增加对更快数据处理的需求。在使用现有的人工智能处理框架处理具有不同模态和存储格式的数据集时,由于数据布局和用户处理数据的方式等问题,处理速度可能并不理想。因此,使用统一的数据库来存储多种数据格式可以更好地管理和优化数据访问。在本文中,我们介绍了VersaDB,这是一种专门为具有多种模态的人工智能数据集设计的数据库。我们实现了一种基于页面的存储系统,将结构化数据和非结构化数据分开。此外,我们生成了基于B+树的索引文件,以加速数据访问。VersaDB支持自动分片,并维护一个分层的元数据管理系统,相应的元数据在页面、分片和全局级别上进行维护,为数据库的高效运行奠定基础。我们还关注易用性,提供了将数据集直接转换为VersaDB的API,以及将流行的人工智能数据存储格式(例如CSV、TFRecord、.bin)转换为VersaDB的API。我们的实验表明,使用VersaDB可以实现高达5.35倍的加速,并在不同的并行级别上保持一致的性能。
cs.CV / 119 / 2608.22819

Direct, Parallel, or Sequential? A Comparative Study of Training-Free Multi-Subject Image-to-Video Generation

直接、并行还是顺序?无训练多主体图像到视频生成的比较研究
Qi, Yanliang, Chen, Kexi, Ye, Muchao, Ni, Haomiao
Abstract
Text-conditioned image-to-video (I2V) generation has advanced rapidly, yet generating videos with multiple subjects remains challenging. A model must simultaneously preserve the appearance of each subject, assign distinct motions, and maintain coherent spatial and temporal interactions. This paper presents a systematic study of three representative paradigms for training-free multi-subject I2V generation: direct, parallel, and sequential generation. Direct generation applies a pretrained I2V model to the complete reference image and prompt, requiring all subjects and motions to be synthesized jointly. Parallel and sequential generation instead decompose the reference image and prompt into subject-specific visual and textual conditions. Parallel generation synthesizes each subject independently and subsequently composes the resulting videos, reducing the complexity of each generation step at the cost of weaker inter-subject context. Sequential generation first synthesizes a background video and then progressively introduces individual subjects. This preserves accumulated scene context but introduces sensitivity to subject ordering and error propagation. We empirically evaluate the three paradigms across diverse multi-subject scenes, comparing appearance preservation, motion fidelity, temporal consistency, and inter-subject coherence, while also characterizing their distinct failure modes. Our findings reveal the strengths and limitations of each paradigm and offer practical insights for designing controllable multi-subject video generation systems.
Chinese Translation
基于文本的图像到视频(I2V)生成技术发展迅速,但生成包含多个主体的视频仍然具有挑战性。模型必须同时保留每个主体的外观,分配不同的运动,并保持一致的空间和时间交互。本文系统地研究了三种无训练多主体I2V生成的代表性范式:直接生成、并行生成和顺序生成。直接生成将预训练的I2V模型应用于完整的参考图像和提示,要求所有主体和运动共同合成。相对而言,并行生成和顺序生成将参考图像和提示分解为主体特定的视觉和文本条件。并行生成独立合成每个主体,然后组合生成的视频,降低了每个生成步骤的复杂性,但代价是弱化了主体间的上下文关系。顺序生成则首先合成背景视频,然后逐步引入各个主体。这种方法保留了累积的场景上下文,但对主体顺序和错误传播敏感。我们在多样的多主体场景中对这三种范式进行了实证评估,比较了外观保留、运动保真度、时间一致性和主体间一致性,同时还描述了它们各自的失败模式。我们的研究结果揭示了每种范式的优缺点,并为设计可控的多主体视频生成系统提供了实用的见解。
cs.CV / 120 / 2608.22821

SiZeUp: Fast 3D Proxy from Aerial Images via Depth Ordinal Loss

SiZeUp:通过深度序数损失从航空图像快速构建3D代理模型
Zhou, Wenjun, Li, Yunshan, Zhu, Qiaoyu, Xiong, Weidan, Zhang, Hao, Cohen-Or, Daniel, Huang, Hui
Abstract
We present SiZeUp, a fast and scalable approach for constructing large-scale 3D urban proxy models directly from calibrated oblique aerial imagery. Our method adopts a height-from-footprint representation, reducing 3D building abstraction to a low-dimensional optimization problem in which building footprints are extruded by a single height parameter. To enable efficient and robust height estimation, we introduce an ordinal depth consistency loss that enforces agreement between the relative depth ordering of rendered proxies and depth priors predicted by a monocular depth model. This is realized through a differentiable renderer that maps parametric building proxies into multi-view depth images, allowing gradients to be propagated from depth supervision to building heights. Our ordinal formulation produces stable optimization in practice and avoids explicit feature matching or dense point cloud reconstruction. Rather than relying on metric depth, which can be unreliable under monocular scale ambiguity, our ordinal depth consistency loss operates on relative depths, providing a more reliable signal across views. Combined with an efficient dynamic view selection, our approach achieves a 23-52$\times$ speedup over state-of-the-art proxy reconstruction pipelines while maintaining comparable proxy-level coverage and volume consistency, making it well suited for large-scale urban modeling tasks.
Chinese Translation
我们提出了SiZeUp,一种快速且可扩展的方法,用于直接从校准的倾斜航空图像构建大规模3D城市代理模型。我们的方法采用基于轮廓的高度表示,将3D建筑抽象简化为一个低维优化问题,其中建筑轮廓通过单一的高度参数进行挤出。为了实现高效且稳健的高度估计,我们引入了一种序数深度一致性损失,强制渲染的代理与单目深度模型预测的深度先验之间的相对深度排序一致。这是通过一个可微分的渲染器实现的,该渲染器将参数化建筑代理映射到多视图深度图像中,从而允许从深度监督向建筑高度传播梯度。我们的序数形式在实践中产生了稳定的优化,并避免了显式特征匹配或密集点云重建。与依赖于度量深度(在单目尺度模糊下可能不可靠)不同,我们的序数深度一致性损失在相对深度上操作,提供了跨视图更可靠的信号。结合高效的动态视图选择,我们的方法在保持可比的代理级覆盖率和体积一致性的同时,相较于最先进的代理重建管道实现了23-52倍的加速,使其非常适合大规模城市建模任务。
cs.CV / 121 / 2608.22828

VeCAS: Vessel-Focused Contrast-Free Angiogram Synthesis for Vascular Interventions

VeCAS:针对血管介入的以血管为中心的无对比剂血管造影合成
Huang, De-Xing, Wang, Chen-Yu, Liang, Hao, Zhou, Xiao-Hu, Gui, Mei-Jiang, Xiang, Tian-Yu, Zhang, Qin-Yi, Wang, Chen, Xie, Xiao-Liang, Liu, Shi-Qi, Liu, Ming-Yuan, Wang, Zhen-Chang, Hou, Zeng-Guang
Abstract
X-ray angiography relies on iodinated contrast agents to visualize vascular structures during image-guided interventions. However, contrast administration carries risks of adverse events, motivating the development of contrast-free alternatives. Generating X-ray angiograms directly from non-contrast X-ray images offers a potential solution, but existing approaches remain limited by (i) insufficient control over vascular localization and (ii) inefficient modeling of redundant background content. To address these challenges, we propose VeCAS, a two-stage vessel-focused contrast-free angiogram synthesis framework that separates vascular structure localization from angiographic appearance synthesis. In Stage I, a discriminative model localizes vascular structures in non-contrast X-ray images, while cross-modality latent distillation transfers vessel-sensitive knowledge from X-ray angiograms during training. In Stage II, a vessel-focused inpainting model synthesizes angiographic appearance within the localized vascular regions while preserving the non-vascular background. Experiments on an in-house lower-limb vascular intervention dataset show that VeCAS outperforms the comparison methods in terms of vascular structural fidelity and image quality. Visual Turing tests and physician assessments indicate the perceptual realism of the synthesized angiograms. In addition, robotic guidewire navigation experiments in vascular phantoms show that VeCAS guidance reduces the time to target by 41.4% and the number of operation steps by 40.7% compared with non-contrast guidance. Together, these results suggest the potential of VeCAS to serve as ``meta contrast agent'' for vascular interventions.
Chinese Translation
X射线血管造影依赖于含碘对比剂在图像引导介入过程中可视化血管结构。然而,对比剂的使用存在不良事件的风险,这促使了无对比剂替代方案的发展。直接从无对比X射线图像生成X射线血管造影提供了一种潜在的解决方案,但现有方法受限于(i)对血管定位的控制不足和(ii)冗余背景内容建模效率低下。为了解决这些挑战,我们提出了VeCAS,一个两阶段的以血管为中心的无对比剂血管造影合成框架,该框架将血管结构定位与血管造影外观合成分离。在第一阶段,判别模型在无对比X射线图像中定位血管结构,同时跨模态潜在蒸馏在训练过程中从X射线血管造影中转移对血管敏感的知识。在第二阶段,专注于血管的修复模型在定位的血管区域内合成血管造影外观,同时保留非血管背景。在我们自建的下肢血管介入数据集上的实验表明,VeCAS在血管结构保真度和图像质量方面优于对比方法。视觉图灵测试和医生评估表明合成的血管造影在感知上具有真实感。此外,在血管模型中的机器人导丝导航实验显示,与无对比指导相比,VeCAS指导将目标时间缩短了41.4%,操作步骤减少了40.7%。综合来看,这些结果表明VeCAS有潜力作为血管介入的“元对比剂”。
cs.CV / 122 / 2608.22861

Following Motion for Sequential Modeling in Video Frame Interpolation

基于运动的序列建模在视频帧插值中的应用
Park, Jaehyun, Cho, Nam Ik
Abstract
State Space Models (SSMs) have surfaced as a promising architecture in Video Frame Interpolation (VFI), as they can capture long-range dependencies with linear computational complexity. However, their predefined scanning order limits their effectiveness in modeling the dynamic motion trajectories inherent in VFI problems. To tackle this challenge, we propose Motion-Guided Mamba for Video Frame Interpolation (MGMVFI), an adaptation of the selective state space model tailored explicitly for VFI. MGMVFI introduces Motion-Guided Serialization (MGS), which leverages optical flow to define a motion-adaptive 1D input order for the SSM. This aligns the causal state updates with semantically related tokens, enabling motion-consistent feature propagation, particularly for large and dynamic motions. Additionally, to mitigate the unreliable feature representations caused by inaccurate optical flow estimates, we introduce contextual synthesis that utilizes the surrounding spatial context for robust inter-frame feature synthesis. These components are seamlessly integrated within our tailored Mamba architecture, which also employs a lightweight refinement block to enhance local detail reconstruction at a reduced computational cost. Extensive experiments on standard VFI benchmarks demonstrate that MGMVFI achievesstate-of-the-artperformance,particularly on complex and dynamic motions, thereby establishing a new direction for sequence modeling in video interpolation.
Chinese Translation
状态空间模型(SSMs)作为视频帧插值(VFI)中一种有前景的架构,能够以线性计算复杂度捕捉长程依赖关系。然而,它们预定义的扫描顺序限制了其在建模VFI问题中固有的动态运动轨迹的有效性。为了解决这一挑战,我们提出了运动引导的Mamba视频帧插值(MGMVFI),这是针对VFI专门调整的选择性状态空间模型的适配。MGMVFI引入了运动引导序列化(MGS),利用光流定义运动自适应的一维输入顺序,以便将因果状态更新与语义相关的标记对齐,从而实现运动一致的特征传播,特别是在大规模和动态运动情况下。此外,为了减轻因光流估计不准确而导致的不可靠特征表示,我们引入了上下文合成,利用周围的空间上下文进行稳健的帧间特征合成。这些组件在我们定制的Mamba架构中无缝集成,该架构还采用轻量级的细化模块,以降低计算成本并增强局部细节重建。在标准VFI基准上的大量实验表明,MGMVFI在复杂和动态运动方面实现了最先进的性能,从而为视频插值中的序列建模开辟了新的方向。
cs.CV / 123 / 2608.22866

Toward Sub-1 kB Identity-Preserving Face Compression: A Benchmark of Codecs, a Custom Learned Codec, and Studies of Resolution, Demographic Fairness, Recompression, and Adversarial Robustness

朝向亚1 kB身份保留人脸压缩:编解码器基准测试、自定义学习编解码器,以及分辨率、人口公平性、重压缩和对抗鲁棒性的研究
Hurtik, Petr, Sochor, Jakub
Abstract
Storing face images under a hard sub-kilobyte budget, as required for identity documents, smart-card biometrics and bandwidth-constrained verification, forces a codec to discard most of the signal while keeping what a face matcher actually reads: identity. Generic codecs optimize pixel fidelity, not the embedding distances that drive verification, so which codec, resolution and setting best preserve identity at 1024 bytes or less, and how that degrades at 512, is unclear. We benchmark ten general and face-specific codecs across resolutions, byte budgets, two datasets (controlled Color FERET, in-the-wild AI-Solutions-KK) and four anchor face matchers, with a fourteen-model ViT and CNN roster confirming the ranking is backbone-invariant. We then train a custom identity-preserving codec that hits the byte budget exactly via binary search over a frozen gain table, and run four studies: resolution, demographic fairness, recompression, and no-box adversarial robustness. Sub-kilobyte identity preservation is feasible, but which codec to deploy depends entirely on the budget. At 1024 bytes and the 112 px working resolution the problem is close to solved: modern codecs hold Color FERET equal-error rate under 0.35 percent on the ArcFace anchor. At 512 bytes the field re-sorts: AVIF, HEIF, JPEG XL and legacy JPEG collapse to 28 to 98 percent false-non-match rate at FMR 1e-4, while WebP, JPEG-AI and our byte-budgeted learned codecs stay out of that band, with 24.3 percent for WebP against 6.9 percent for our accurate variant in the wild. That re-sort, not the 1024-byte ranking, is the operational result: a codec chosen at 1 kB is not the codec to deploy at half that.
Chinese Translation
在身份文件、智能卡生物识别和带宽受限的验证中,存储人脸图像的硬性亚千字节预算迫使编解码器丢弃大部分信号,同时保留人脸匹配器实际读取的内容:身份。通用编解码器优化像素保真度,而不是驱动验证的嵌入距离,因此在1024字节或更少的情况下,哪种编解码器、分辨率和设置能最好地保留身份,以及在512字节时如何退化,尚不清楚。我们在不同分辨率、字节预算、两个数据集(受控的Color FERET和野外的AI-Solutions-KK)以及四个锚定人脸匹配器上,对十种通用和人脸特定的编解码器进行了基准测试,使用十四个模型的ViT和CNN阵容确认排名与骨干网络无关。然后,我们训练了一个自定义的身份保留编解码器,通过对冻结增益表的二分搜索精确达到字节预算,并进行了四项研究:分辨率、人口公平性、重压缩和无框对抗鲁棒性。亚千字节身份保留是可行的,但选择哪种编解码器完全取决于预算。在1024字节和112 px工作分辨率下,问题接近解决:现代编解码器在ArcFace锚定下保持Color FERET的等错误率低于0.35%。在512字节时,领域重新排序:AVIF、HEIF、JPEG XL和传统JPEG在FMR 1e-4下崩溃至28%到98%的假非匹配率,而WebP、JPEG-AI和我们的字节预算学习编解码器则保持在该范围之外,WebP的假非匹配率为24.3%,而我们在野外的准确变体为6.9%。这种重新排序,而非1024字节的排名,是操作结果:在1 kB选择的编解码器并不是在其一半时部署的编解码器。
cs.CV / 124 / 2608.22879

Large-Small Model Collaboration for Zero-Shot Surgical Phase Recognition

大模型与小模型协作实现零-shot外科阶段识别
Zhang, Yiyi, Zheng, Ying, Fan, Wenxin, Zhu, Yu, Yuan, Yuchen, Zhao, Litao, Li, Zheng, Heng, Pheng-Ann
Abstract
Task-specific lightweight models for surgical phase recognition excel at capturing temporal dynamics but generalize poorly under domain shift. Conversely, surgical foundation models (FMs) offer superior transferability via large-scale pretraining, yet their lack of explicit temporal modeling often yields temporally inconsistent predictions, leading to degraded performance. To exploit the complementary strengths of both paradigms, we propose \textbf{La}rge-\textbf{S}mall \textbf{T}emporal adaptation (\textbf{LaST}), a novel large-small collaborative framework that enables zero-shot adaptation to unseen clinical domains. In LaST, the FM initiates the pipeline by generating frame-level phase priors that serve as initial weak supervision. To effectively utilize these noisy phase priors, we introduce an iterative temporal refinement scheme that integrates dynamic quality control to filter reliable predictions and dual-model cross-learning to mitigate confirmation bias. Simultaneously, the lightweight model leverages its intrinsic temporal modeling ability to progressively correct inconsistent predictions and enhance overall accuracy across iterations. At the end, a cycle replay strategy is employed to close the loop: the refined, more accurate predictions are utilized as upgraded supervision signals for the subsequent iterations, fostering a self-reinforcing evolution of both label quality and model capability. Extensive experiments demonstrate that LaST achieves robust adaptation to unseen domains for zero-shot surgical phase recognition, outperforming the baseline (PeskaVLP) by 24.85\%-43.17\% in accuracy and even surpassing fully supervised linear probing and several state-of-the-art few-shot approaches. Codes will be released at https://github.com/YIYIZH/LaST.
Chinese Translation
针对外科阶段识别的任务特定轻量模型在捕捉时间动态方面表现优异,但在领域迁移时泛化能力较差。相反,外科基础模型(FMs)通过大规模预训练提供了更好的可迁移性,但其缺乏明确的时间建模,常常导致时间上不一致的预测,从而降低性能。为了利用这两种范式的互补优势,我们提出了 extbf{La}rge- extbf{S}mall extbf{T}emporal adaptation( extbf{LaST}),一种新颖的大-小协作框架,能够实现对未见临床领域的零-shot适应。在LaST中,基础模型启动管道,通过生成帧级阶段先验作为初始弱监督。为了有效利用这些噪声阶段先验,我们引入了一种迭代时间精炼方案,结合动态质量控制以过滤可靠预测,并通过双模型交叉学习来减轻确认偏差。同时,轻量模型利用其内在的时间建模能力,逐步纠正不一致的预测,并在迭代中提高整体准确性。最后,采用循环重放策略来闭合循环:经过精炼的更准确预测被用作后续迭代的升级监督信号,促进标签质量和模型能力的自我强化演变。大量实验表明,LaST在零-shot外科阶段识别中对未见领域实现了稳健的适应性,准确率超越基线(PeskaVLP)24.85%-43.17%,甚至超过了完全监督的线性探测和几种最先进的少样本方法。代码将发布在 https://github.com/YIYIZH/LaST。
cs.CV / 125 / 2608.22883

FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding

FOVEA:面向需求的视觉证据适应以实现缓存友好的多模态推测解码
Zhu, Hengjie, Wu, Dayan, Zhang, Zihao, Liu, Xinze, Yu, Jingxuan, Fu, Peng, Lin, Zheng, Wang, Weiping, Wang, Ding
Abstract
Multimodal speculative decoding accelerates vision-language models by allowing a lightweight draft model to propose candidate tokens for parallel verification by a larger target model. Existing methods typically condition the drafter on a fixed visual interface, such as a predefined visual-token budget or a static compressed representation. However, our controlled visual-budget analysis shows that visual demand varies substantially across tasks and decoding stages, which means more visual input is not always beneficial. Actually, insufficient evidence may weaken visual grounding, while excessive context adds overhead and may disrupt drafting. We propose FOVEA (Focused On-demand Visual Evidence Adaptation), a cache-friendly approach that builds a reusable visual memory and dynamically retrieves a bounded subset for a draft state. A cumulative-mass rule determines both how many and which entries are selected. The selected entries are aggregated into a visual readout and fused with the current draft hidden state through a lightweight gated residual correction. Rather than inserting visual tokens into the autoregressive context, the correction modifies only the representation passed to the language-model head. Experiments across multiple vision-language backbones and multimodal benchmarks show that FOVEA improves draft acceptance and end-to-end decoding speed, achieving up to $2.13\times$ speedup over autoregressive decoding. These results demonstrate that state-conditioned evidence retrieval is an effective alternative to reusing a fixed visual representation throughout multimodal generation.
Chinese Translation
多模态推测解码通过允许轻量级草稿模型提出候选标记,以便由更大目标模型进行并行验证,从而加速视觉-语言模型。现有方法通常将草稿模型的条件固定在一个固定的视觉接口上,例如预定义的视觉标记预算或静态压缩表示。然而,我们的受控视觉预算分析表明,视觉需求在不同任务和解码阶段之间存在显著差异,这意味着更多的视觉输入并不总是有利的。实际上,证据不足可能削弱视觉基础,而过多的上下文则增加了开销并可能干扰草稿生成。我们提出了FOVEA(面向需求的视觉证据适应),这是一种缓存友好的方法,构建可重用的视觉记忆,并动态检索一个有限的子集用于草稿状态。累积质量规则决定了选择多少条目以及选择哪些条目。所选条目被聚合为视觉读出,并通过轻量级门控残差修正与当前草稿隐藏状态融合。与将视觉标记插入自回归上下文不同,修正仅修改传递给语言模型头的表示。在多个视觉-语言骨干网络和多模态基准测试中的实验表明,FOVEA提高了草稿接受率和端到端解码速度,相较于自回归解码实现了高达$2.13 imes$的加速。这些结果表明,状态条件下的证据检索是整个多模态生成过程中重用固定视觉表示的有效替代方案。
cs.CV / 126 / 2608.22885

DRAgent: Discriminative Reasoning Agent for Referring Expression Segmentation

DRAgent:用于指称表达分割的区分推理代理
Qi, Yujie, Zhang, Luyan
Abstract
Referring Expression Segmentation (RES) aims to generate a pixel-level mask for the object specified by a language expression. Recent methods based on multimodal large language models (MLLMs) often rely on one-pass coordinate prediction for visual localization, which serializes continuous spatial locations as discrete text tokens and may lead to localization bias and alignment errors. To address these issues, we propose DRAgent, an MLLM-driven discriminative reasoning (DR) framework for RES. Instead of requiring the MLLM to generate localization coordinates, DRAgent first constructs a detector-generated candidate space and then uses the MLLM as a visual-semantic target discriminator. Specifically, the MLLM performs reliable target selection among potential distractors through a two-stage DR mechanism, which first screens high-recall candidates and then performs instance-wise verification. The selected target box is subsequently used as a spatial prompt for a foundation segmentation model to produce the final pixel-level mask. Furthermore, we construct a self-consistency-filtered reasoning-chain data pipeline for LoRA-based fine-tuning, providing more reliable supervision for enhancing the MLLM's discriminative reasoning capability. Experiments demonstrate that DRAgent achieves competitive performance on RefCOCO, RefCOCO+, and RefCOCOg.
Chinese Translation
指称表达分割(RES)旨在为语言表达指定的对象生成像素级掩码。基于多模态大型语言模型(MLLM)的近期方法通常依赖于一次性坐标预测进行视觉定位,这将连续的空间位置序列化为离散的文本标记,可能导致定位偏差和对齐错误。为了解决这些问题,我们提出了DRAgent,一种基于MLLM的区分推理(DR)框架用于RES。DRAgent并不要求MLLM生成定位坐标,而是首先构建一个检测器生成的候选空间,然后将MLLM用作视觉-语义目标鉴别器。具体而言,MLLM通过两阶段的DR机制在潜在干扰项中进行可靠的目标选择,首先筛选高召回候选项,然后进行实例级验证。所选目标框随后被用作基础分割模型的空间提示,以生成最终的像素级掩码。此外,我们构建了一个自一致性过滤的推理链数据管道用于基于LoRA的微调,为增强MLLM的区分推理能力提供了更可靠的监督。实验表明,DRAgent在RefCOCO、RefCOCO+和RefCOCOg上取得了竞争性的性能。
cs.CV / 127 / 2608.22888

NemoSplat: Feed-Forward 4D Gaussian Splatting for Media-Aware Underwater Reconstruction

NemoSplat:面向媒体感知的前馈4D高斯点云重建
Guo, Xiaopeng, Tse, Wai Chung, Zhu, Yipeng, Zhang, Hanwen, Huang, Huajian, Yeung, Sai-Kit
Abstract
Reconstructing photorealistic scenes in unconstrained underwater environments remains challenging due to severe media-induced light scattering and unpredictable dynamic objects. Recent feed-forward visual foundation models have demonstrated remarkable capabilities in generalized novel view synthesis and tracking. However, when directly applied to aquatic videos, optical attenuation and motion interference fatally corrupt their feature aggregation, leading to severe tracking and reconstruction failures. To overcome these limitations, we present NemoSplat, the first feed-forward 4D Gaussian Splatting framework tailored for media-aware dynamic reconstruction directly from uncalibrated marine videos. Beyond providing robust estimations of camera poses and dense scene depth, we devise a Promptable Dynamic Disentangler that utilizes a confidence-aware fusion strategy of learned dynamic probabilities and optional semantic text priors, effectively isolating massive transient entities. Furthermore, to counteract visual degradation, a Media-Aware Gaussian Predictor is formulated to jointly estimate intrinsic 3D Gaussian attributes alongside physical media parameters, rendering pristine scene appearance in a single forward pass. Additionally, we introduce a large-scale underwater dataset with massive dynamic elements to facilitate training and evaluation. Extensive experiments on our dataset demonstrate that NemoSplat achieves state-of-the-art tracking accuracy and high-fidelity rendering.
Chinese Translation
在不受限制的水下环境中重建逼真的场景仍然面临挑战,主要由于严重的介质引起的光散射和不可预测的动态物体。最近的前馈视觉基础模型在广义新视图合成和跟踪方面表现出了显著的能力。然而,当这些模型直接应用于水下视频时,光学衰减和运动干扰严重破坏了特征聚合,导致跟踪和重建的失败。为克服这些限制,我们提出了NemoSplat,这是第一个专为媒体感知动态重建而设计的前馈4D高斯点云框架,能够直接从未校准的海洋视频中进行重建。除了提供稳健的相机姿态和密集场景深度估计外,我们还设计了一种可提示的动态解耦器,利用学习到的动态概率和可选的语义文本先验的置信度感知融合策略,有效地隔离大量瞬态实体。此外,为了抵消视觉退化,我们提出了一种媒体感知高斯预测器,能够联合估计内在的3D高斯属性和物理介质参数,从而在单次前向传递中呈现清晰的场景外观。此外,我们还引入了一个大规模的水下数据集,包含大量动态元素,以促进训练和评估。在我们的数据集上进行的广泛实验表明,NemoSplat实现了最先进的跟踪精度和高保真渲染。
cs.CV / 128 / 2608.22906

AquaFlow: A Monocular Gaussian Splatting SLAM for Underwater Streaming Reconstruction

AquaFlow:一种用于水下流式重建的单目高斯点云SLAM
Xu, Yingxiang, Ren, Kerui, Guo, Wenqi, Jiang, Changjian, Lu, Tao, Xu, Linning, Yu, Mulin
Abstract
Recent monocular 3D Gaussian Splatting (3DGS) streaming reconstruction methods have achieved impressive performance by balancing reconstruction quality and efficiency. However, extending these frameworks to underwater scenes remains challenging due to severe visual degradation, such as light attenuation and scattering, which degrades camera pose tracking and distorts scene geometry. To address these challenges, we propose AquaFlow, a monocular Gaussian Splatting streaming reconstruction framework for efficient and high-fidelity underwater reconstruction. Specifically, AquaFlow fine-tunes a 3D vision foundation model on large-scale underwater data for robust pose and pointmap estimation, and introduces a medium-guided incremental Gaussian initialization strategy for streaming mapping. Furthermore, we develop a streaming-compatible hybrid scene representation that integrates structured, distance-conditioned neural Gaussians with a physics-inspired optical model to compensate for underwater image formation effects, enabling accurate scene reconstruction. We evaluate AquaFlow on a comprehensive dataset of 62 diverse underwater trajectories, collected from both public benchmarks and in-the-wild web videos across various scales. Extensive experiments demonstrate that AquaFlow achieves state-of-the-art tracking and rendering performance, reducing average localization error by 13.2% and improving PSNR by 4.74 dB compared to WaterSplat-SLAM.
Chinese Translation
近期的单目3D高斯点云(3DGS)流式重建方法在重建质量与效率之间取得了令人瞩目的平衡。然而,由于光衰减和散射等严重的视觉退化,将这些框架扩展到水下场景仍然面临挑战,这会降低相机姿态跟踪的精度并扭曲场景几何形状。为了解决这些挑战,我们提出了AquaFlow,一种用于高效且高保真水下重建的单目高斯点云流式重建框架。具体而言,AquaFlow在大规模水下数据上微调了3D视觉基础模型,以实现稳健的姿态和点图估计,并引入了一种中等引导的增量高斯初始化策略用于流式映射。此外,我们开发了一种兼容流式处理的混合场景表示,结合了结构化的、距离条件的神经高斯模型与受物理启发的光学模型,以补偿水下图像形成效应,从而实现准确的场景重建。我们在一个包含62条多样化水下轨迹的综合数据集上评估了AquaFlow,这些轨迹来自公共基准测试和各种规模的野外网络视频。大量实验表明,AquaFlow在跟踪和渲染性能上达到了最先进的水平,与WaterSplat-SLAM相比,平均定位误差降低了13.2%,PSNR提高了4.74 dB。
cs.CV / 129 / 2608.22914

Results of the 1st Asynchronous CASTLE Challenge at the Joint Egocentric Vision Workshop in Conjunction with CVPR 2026

2026年CVPR会议联合自我中心视觉研讨会首届异步CASTLE挑战赛的结果
Rossetto, Luca, Bailer, Werner, Gurrin, Cathal, Healy, Graham, Khan, Omar Shahbaz, Rudinac, Stevan, Schöffmann, Klaus, Tran, Allie
Abstract
This report summarizes the contributions and results of the 1st Asynchronous CASTLE Challenge at the Joint Egocentric Vision Workshop in conjunction with CVPR 2026.
Chinese Translation
本报告总结了2026年CVPR会议联合自我中心视觉研讨会首届异步CASTLE挑战赛的贡献和结果。
cs.CV / 130 / 2608.22926

Motion-Based Tokenization for Cross-Dataset Egocentric Gaze Modeling

基于运动的跨数据集自我中心注视建模的标记化
Maquiling, Virmarie, Cai, Zhuojiang, Kasneci, Enkelejda
Abstract
Gaze is increasingly used as an input signal for vision and multimodal models, yet no consensus exists on how to represent it across datasets. Raw traces preserve detail but are noisy and device-dependent, while coarse event labels are easy to model but can discard local motion structure. We formulate event-aligned, fixed-horizon angular displacement as an interpretable, event-conditioned motion vocabulary and compare it with event-only, spatial, absolute-angle, learned vector-quantized, and continuous representations. To assess transfer alongside target predictability and token collapse, our evaluation combines next-token prediction with target-domain regret, low-order target references, paired bootstrap, order sensitivity, motif overlap, and frozen structural probes. In an event-aligned headset benchmark, angular-motion tokens have lower target-domain regret than frozen-codebook VQ tokens in one transfer direction, while the reverse direction is inconclusive. The probes reveal complementary representation properties, and event-only tokens show that low perplexity can retain little motion information. On a third egocentric dataset, a matched comparison of I-VT, native, and frame-span interfaces shows that event construction materially changes transfer: native events have the lowest regret into EGTEA, while frame-span events have zero motif overlap and fail severely as a source. Motion-based tokenization therefore provides a compact representation for event-aligned egocentric gaze streams, while the evaluation identifies how target predictability and event construction shape cross-dataset conclusions.
Chinese Translation
注视越来越多地被用作视觉和多模态模型的输入信号,但在不同数据集之间如何表示注视尚无共识。原始轨迹保留了细节,但噪声大且依赖于设备,而粗略的事件标签易于建模,但可能会丢失局部运动结构。我们将事件对齐的固定视野角位移表述为可解释的、事件条件的运动词汇,并将其与仅基于事件的空间绝对角度、学习的向量量化和连续表示进行比较。为了评估转移能力、目标可预测性和标记崩溃,我们的评估将下一个标记预测与目标领域的悔恨、低阶目标参考、配对自助法、顺序敏感性、主题重叠和冻结结构探针结合在一起。在一个事件对齐的头戴式设备基准测试中,角运动标记在一个转移方向上比冻结代码本向量量化标记具有更低的目标领域悔恨,而反向方向则不确定。探针揭示了互补的表示特性,仅基于事件的标记显示低困惑度可能保留很少的运动信息。在第三个自我中心数据集中,I-VT、原生和帧跨度接口的匹配比较显示事件构建实质性改变了转移:原生事件在EGTEA中的悔恨最低,而帧跨度事件的主题重叠为零,作为源的表现严重失败。因此,基于运动的标记化为事件对齐的自我中心注视流提供了一种紧凑的表示,同时评估识别了目标可预测性和事件构建如何塑造跨数据集的结论。
cs.CV / 131 / 2608.22937

Quality Inspection of Printed Circuit Board Pin Insertion via Semantic Segmentation and Board-Level Feature Extraction

基于语义分割和板级特征提取的印刷电路板引脚插入质量检测
Rabeneck, Nils, Kiunke, André, Hoess, Nicole, Mauerer, Wolfgang
Abstract
Quality control during printed circuit board (PCB) assembly is a critical step in ensuring reliable electronic products. Detecting misaligned pins during or after pin insertion remains a particularly challenging inspection task. This paper presents an automated defect detection method for identifying incorrectly inserted pins on PCBs. The proposed pipeline combines semantic segmentation using a U-Net architecture with contour-based feature extraction and logistic regression for board-level pass/fail classification. Segmentation masks are used to derive contour representations of individual pins, from which board-level features -such as average contour size- are extracted and used to train a logistic regression classifier. We evaluate the method on two datasets: an industrial collection of real-world PCB images, and a publicly available PCB pin-inspection dataset with substantially different visual characteristics. To assess the effectiveness of the proposed approach, a comparison against PatchCore, an anomaly detection technique new to be applied to pin inspection, as well as instance segmentation-based pin detection is made. The developed method achieved Area Under the Receiver Operating Characteristic Curve (ROC-AUC) values of 0.990 on a random test set split from the industrial data and 1.000 on the public dataset indicating strong separation between pass and fail boards. The results indicate that the proposed approach is a promising candidate for automated pin inspection in industrial environments and achieves strong performance on datasets with substantially different visual characteristics after dataset-specific training.
Chinese Translation
在印刷电路板(PCB)组装过程中,质量控制是确保电子产品可靠性的关键步骤。在引脚插入过程中或之后检测引脚对齐不当仍然是一项特别具有挑战性的检查任务。本文提出了一种自动化缺陷检测方法,用于识别PCB上插入不正确的引脚。所提出的流程结合了使用U-Net架构的语义分割、基于轮廓的特征提取以及用于板级合格/不合格分类的逻辑回归。分割掩膜用于推导单个引脚的轮廓表示,从中提取板级特征,例如平均轮廓大小,并用于训练逻辑回归分类器。我们在两个数据集上评估该方法:一个是工业收集的真实PCB图像,另一个是具有显著不同视觉特征的公开PCB引脚检测数据集。为了评估所提方法的有效性,我们与PatchCore(一种新应用于引脚检测的异常检测技术)以及基于实例分割的引脚检测进行了比较。所开发的方法在从工业数据中随机拆分的测试集上达到了0.990的接收者操作特征曲线下面积(ROC-AUC)值,在公共数据集上达到了1.000,表明合格和不合格电路板之间有很强的区分性。结果表明,所提方法是工业环境中自动化引脚检测的有希望的候选方案,并且在经过特定数据集训练后,在具有显著不同视觉特征的数据集上实现了强大的性能。
cs.CV / 132 / 2608.22950

WADE: A Reasoning-Annotated Benchmark for Multi-Instance Floating-Waste Grounding with Compact Vision-Language Models

WADE:一种用于多实例漂浮废物定位的推理注释基准,适用于紧凑型视觉-语言模型
Shuvo, Md. Asaduzzaman, Farabi, Ahsan, Minhaz, Md. Abdul Ahad, Hasan, Mahedi, Khandaker, Israt, Shanto, Ibrahim Khalil, Kabir, Muhammad Nomani
Abstract
Floating waste in inland waterways threatens aquatic ecosystems and requires timely monitoring under cluttered, multi-object conditions. Existing aquatic-waste datasets provide limited geographic coverage, sparse multi-instance annotations, and little supervision beyond boxes and labels. Compact vision-language models (VLMs) therefore remain insufficiently evaluated for jointly localizing, classifying, counting, and explaining floating waste. We introduce WADE, a reasoning-annotated benchmark containing 2,167 images from rural Bangladesh, 13,608 bounding boxes, and ten waste categories. Each annotation is associated with class-level recognition rules covering visual cues, likely confusions, and discriminative features. We evaluate six VLMs under zero-shot, two-shot, reasoning-guided, and fine-tuned settings using detection, counting, and hallucination metrics. For resource-efficient adaptation, we jointly fine-tune Qwen3-VL-2B on boxes, labels, and reasoning chains using QLoRA. Fine-tuning increases recall from 0.0248 to 0.2339 and F1 from 0.0257 to 0.2163, while reducing image-level hallucination from 0.6836 to 0.0883. However, over three-quarters of instances remain undetected, establishing WADE as a challenging benchmark for dense floating-waste grounding with compact VLMs.
Chinese Translation
内陆水道中的漂浮废物威胁水生生态系统,需要在杂乱的多物体条件下进行及时监测。现有的水生废物数据集在地理覆盖范围、稀疏的多实例注释以及超出边界框和标签的监督方面都存在局限。因此,紧凑型视觉-语言模型(VLMs)在联合定位、分类、计数和解释漂浮废物方面的评估仍然不足。我们提出了WADE,一个推理注释基准,包含来自孟加拉国农村的2,167张图像、13,608个边界框和十个废物类别。每个注释都与涵盖视觉线索、可能的混淆和区分特征的类别级识别规则相关联。我们在零样本、双样本、推理引导和微调设置下,使用检测、计数和幻觉指标评估了六个VLMs。为了资源高效的适应,我们使用QLoRA对Qwen3-VL-2B在边界框、标签和推理链上进行了联合微调。微调使得召回率从0.0248提高到0.2339,F1值从0.0257提高到0.2163,同时将图像级幻觉从0.6836降低到0.0883。然而,超过四分之三的实例仍未被检测到,这使得WADE成为一个针对紧凑型VLMs的密集漂浮废物定位的挑战性基准。
cs.CV / 133 / 2608.22959

WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans

WildHandBench:一个挑战多语言大型模型和人类的手写文本理解基准
Zhang, Jun, Zhao, Qiao, Cui, Cheng, Qu, Jianying, Sun, Zhongkai, Yang, Jianwen, Zhou, Changda, Liu, ZhuoXin, Han, Shubin
Abstract
While the top model on OmniDocBench now reaches 96.34% overall on printed-document parsing, the ability of current models to handle challenging handwritten documents remains largely uncharacterized. Existing benchmarks focus on isolated text or formulas, overlook handwritten tables and real-world degradation, and report aggregate accuracy without explaining why models fail. We present WildHandBench, a benchmark containing 500 handwritten documents across three structures (free text, tables, formulas), four languages, and nine real-world scenarios. We introduce a Prior-Driven Error (PDE) metric that quantifies whether errors originate from language priors rather than visual evidence. Evaluating 18 state-of-the-art models together with calibrated human baselines, we find: (1) the best model achieves only 71.85% overall; (2) humans outperform all models yet the gap is narrow (77.09% vs. 71.85%); and (3) model errors are qualitatively different from human errors -- 63-91% of model errors are prior-driven versus only 49% for humans, exposing systematic reliance on language priors that conventional accuracy metrics cannot capture.
Chinese Translation
尽管当前在OmniDocBench上的最佳模型在打印文档解析中达到了96.34%的整体准确率,但现有模型处理具有挑战性的手写文档的能力仍然在很大程度上未被表征。现有基准主要集中在孤立的文本或公式上,忽视了手写表格和现实世界的退化,并报告了整体准确率而未解释模型失败的原因。我们提出了WildHandBench,一个包含500份手写文档的基准,涵盖三种结构(自由文本、表格、公式)、四种语言和九种现实世界场景。我们引入了一种优先驱动错误(Prior-Driven Error, PDE)指标,用于量化错误是否源于语言先验而非视觉证据。通过评估18个最先进的模型以及经过校准的人类基线,我们发现:(1)最佳模型的整体准确率仅为71.85%;(2)人类的表现优于所有模型,但差距较小(77.09%对71.85%);(3)模型错误与人类错误在质上有所不同——63-91%的模型错误是由先验驱动的,而人类仅为49%,这揭示了对语言先验的系统性依赖,而传统的准确率指标无法捕捉这一点。
cs.CV / 134 / 2608.22965

Simplified Cross-Modal Calibration for Heterogeneous Event-RGB Stereo Systems

异构事件-RGB立体系统的简化跨模态标定
Hessenthaler, Nico, Müller, Adam T., Stache, Nicolaj C.
Abstract
Accurate extrinsic calibration between event-based and frame-based cameras remains a practical bottleneck for heterogeneous stereo systems. Existing approaches often require sensor or target motion, precise synchronization, or computationally expensive event-to-image reconstruction. We propose a simple, motion-free cross-modal calibration framework that uses a temporally modulated, blended ChArUco target presented on standard consumer displays. By alternating between the original pattern and a partially blended version, the target reliably triggers events while remaining continuously observable to a frame-based camera, avoiding blank frames and reducing synchronization constraints to a coarse, trigger-based alignment. We discretize events into frames coarsely aligned with the RGB images, apply lightweight denoising, and perform ChArUco-based intrinsic and stereo extrinsic calibration. Extensive experiments assess robustness to blending opacity, display brightness, external illumination, viewing angle, and handheld acquisition. Compared to the strongest motion-based reference (E2Calib + Kalibr) and a non-motion-based reference (Plasberg et al.), our approach reduces the mean reprojection error by $44\%$ and $6\%$, respectively, while substantially simplifying the calibration procedure. Finally, we demonstrate practical utility in a robotic eye-to-hand calibration case study, showing consistent transformations and stable downstream geometric measurements even under partial occlusions. Code is publicly available at https://github.com/nhessenthaler/simple-evrgb-cal.
Chinese Translation
事件驱动和帧驱动相机之间的精确外部标定仍然是异构立体系统的一个实际瓶颈。现有的方法通常需要传感器或目标运动、精确同步,或计算开销较大的事件到图像重建。我们提出了一种简单的无运动跨模态标定框架,该框架使用在标准消费显示器上呈现的时间调制混合ChArUco目标。通过在原始图案和部分混合版本之间交替,目标可靠地触发事件,同时对帧驱动相机保持持续可观察性,从而避免空白帧并将同步约束减少到粗略的基于触发的对齐。我们将事件离散化为与RGB图像粗略对齐的帧,应用轻量级去噪,并执行基于ChArUco的内部和立体外部标定。大量实验评估了对混合不透明度、显示亮度、外部照明、视角和手持采集的鲁棒性。与最强的基于运动的参考(E2Calib + Kalibr)和一个非运动基准(Plasberg等)相比,我们的方法分别将平均重投影误差降低了44%和6%,同时大大简化了标定过程。最后,我们在一个机器人眼手标定案例研究中展示了实际应用,显示出一致的变换和稳定的下游几何测量,即使在部分遮挡的情况下也能保持稳定。代码已公开发布于 https://github.com/nhessenthaler/simple-evrgb-cal.
cs.CV / 135 / 2608.22972

Optimize Surgical Triplet Recognition: A Knowledge-Driven Mixture-of-Experts Solution

优化外科三元组识别:一种知识驱动的专家混合解决方案
Zhang, Yiyi, Yuan, Yuchen, Zheng, Ying, Pei, Jialun, Li, Jinpeng, Li, Zheng, Heng, Pheng-Ann
Abstract
Surgical action triplet recognition constitutes a critical task in context-aware robot-assisted surgery, facilitating automatic surgical action perception by identifying instrument, verb, target, and their association. However, existing works struggle to analyze such complex surgical scenes due to three main issues: (1) component-level optimization conflicts caused by entangled feature spaces, (2) category-level optimization conflicts arising from severe data imbalance, and (3) lack of domain knowledge guidance that limits model interpretability and robustness. To address these challenges, we propose a Mixture-of-Experts-guided Co-Optimization (\textit{MoeCo}) framework powered by knowledge-driven learning. Within the co-optimization pipeline, to first mitigate component-level conflicts, we introduce a component-tailored adapter that disentangles task-specific features across spatial-temporal regimes, facilitating effective component specialization. Next, we develop a coordinated gradient learning strategy to handle category-level conflicts, which adaptively rebalances positive-negative gradients to enhance the perception of rare categories. Notably, inspired by surgical domain expertise, we introduce a knowledge-driven mixture-of-experts mechanism that dynamically integrates multimodal large language model-guided knowledge via activated experts, thereby enriching the co-optimization pipeline with more expressive and robust representations. Extensive experiments on the public CholecT45 and CholecT50 datasets confirm the effectiveness of the proposed co-optimization pipeline and the superiority of dynamic priors integration via the knowledge-driven mixture-of-experts mechanism.
Chinese Translation
外科动作三元组识别是上下文感知机器人辅助手术中的一项关键任务,通过识别工具、动词、目标及其关联,促进自动外科动作感知。然而,现有研究在分析复杂的外科场景时面临三大主要问题:(1)由于特征空间交织导致的组件级优化冲突,(2)由于严重的数据不平衡引发的类别级优化冲突,以及(3)缺乏领域知识指导,限制了模型的可解释性和鲁棒性。为了解决这些挑战,我们提出了一种基于知识驱动学习的专家混合引导共同优化框架(Mixture-of-Experts-guided Co-Optimization, extit{MoeCo})。在共同优化流程中,为了首先缓解组件级冲突,我们引入了一种组件定制适配器,该适配器在时空域中解耦任务特定特征,从而促进有效的组件专业化。接下来,我们开发了一种协调梯度学习策略来处理类别级冲突,该策略自适应地重新平衡正负梯度,以增强对稀有类别的感知。值得注意的是,受到外科领域专业知识的启发,我们引入了一种知识驱动的专家混合机制,通过激活专家动态整合多模态大型语言模型引导的知识,从而丰富共同优化流程,使其具备更具表现力和鲁棒性的表示能力。在公共数据集CholecT45和CholecT50上的大量实验验证了所提共同优化流程的有效性,以及通过知识驱动的专家混合机制动态先验整合的优越性。
cs.CV / 136 / 2608.22996

ENCORE: Entropy-Guided Cropping and Attention Regularization for Robust Vision--Language Understanding

ENCORE:基于熵的裁剪和注意力正则化用于鲁棒的视觉-语言理解
Sun, Yuanhao, Ji, Huawei, Ding, Jiaxin, Fu, Luoyi, Wang, Xinbing
Abstract
Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution sub-images, compromising object integrity in lightweight VLMs. Existing methods only focus on the visual modality and fail to dynamically preserve the integrity of prompt-relevant regions, limiting performance. In this work, we observe that the early-layer image-text entropy of cross-modal attention strongly correlates with answer grounding quality and task accuracy. Building on this finding, we propose \textbf{ENCORE}, an entropy-guided framework with two components: At inference, an \textbf{Entropy-based Cropping Strategy} (ECS) evaluates a small set of candidate crops and selects the one with minimal entropy, preserving contiguous regions relevant to the prompt. At training, \textbf{Entropy Regularization Training} (ERT) augments next-token prediction with an entropy term that sharpens attention on key visual tokens while down-weighting irrelevant ones. Experiments on ten VQA benchmarks show that ENCORE, fine-tuning only 0.14\% of parameters, achieves an average 1.43\% accuracy gain and state-of-the-art performance among recent 2B-parameter VLMs. Our code is released in https://github.com/baokou-fw2/ENCORE.
Chinese Translation
视觉-语言模型(VLMs)在多样的视觉-语言任务中表现良好,但基于变换器的视觉编码器将图像分割为固定分辨率的子图像,从而在轻量级VLMs中妥协了物体的完整性。现有方法仅关注视觉模态,未能动态保持与提示相关区域的完整性,限制了性能。在本研究中,我们观察到跨模态注意力的早期层图像-文本熵与答案定位质量和任务准确性之间存在强相关性。基于这一发现,我们提出了 extbf{ENCORE},一个基于熵的框架,包含两个组件:在推理阶段, extbf{基于熵的裁剪策略}(ECS)评估一小组候选裁剪,并选择熵最小的裁剪,保留与提示相关的连续区域。在训练阶段, extbf{熵正则化训练}(ERT)通过熵项增强下一个标记的预测,聚焦于关键视觉标记,同时降低与之无关的标记的权重。在十个视觉问答基准上的实验表明,ENCORE仅微调0.14 ext{%}的参数,平均实现1.43 ext{%}的准确率提升,并在最近的2B参数VLMs中达到最先进的性能。我们的代码已发布在https://github.com/baokou-fw2/ENCORE。
cs.CV / 137 / 2608.23011

Coarse Indexing, Fine Evidence: Decoupling Temporal Granularity in Long-Video RAG

粗略索引,精细证据:长视频 RAG 中时间粒度的解耦
Jin, Zhe, Lin, Zhimin, Zheng, Bin, Fang, Junhua, Yang, Huihua
Abstract
Graph-based retrieval-augmented generation (RAG) provides a scalable paradigm for long-video understanding, but existing systems typically inherit a fixed temporal granularity from video segmentation when constructing their retrieval index. We argue that this design unnecessarily couples indexing granularity with evidence granularity: coarse representations can often suffice for locating relevant temporal regions, while fine-grained evidence remains important for downstream reasoning. We propose \textbf{Density-Aware Graph Construction (DAGC)}, a training-free approach that decouples a query-independent coarse retrieval index from the original fine-grained evidence space. DAGC constructs a compact, density-adaptive graph index by merging visually redundant neighboring chunks, while preserving mappings to the original temporal units. Retrieved coarse regions are subsequently expanded back to the original chunk granularity for fine-grained evidence refinement and answer generation. Experiments on MLVU, VideoMME, and LongVideoBench show that DAGC retains only about 40--50\% of the original graph nodes and achieves $1.3$--$1.7\times$ end-to-end wall-clock acceleration while preserving approximately 99\% of the original QA performance. The gains transfer across different LVLM backbones and video RAG pipelines, suggesting that long-video RAG need not maintain the same temporal granularity for indexing and evidence reasoning.
Chinese Translation
基于图的检索增强生成(RAG)为长视频理解提供了一种可扩展的范式,但现有系统在构建检索索引时通常继承自视频分割的固定时间粒度。我们认为这种设计不必要地将索引粒度与证据粒度耦合:粗略表示通常足以定位相关的时间区域,而细粒度证据在下游推理中仍然重要。我们提出了 extbf{密度感知图构建(DAGC)},这是一种无训练的方法,能够将与查询无关的粗略检索索引与原始细粒度证据空间解耦。DAGC通过合并视觉上冗余的相邻块构建一个紧凑的、密度自适应的图索引,同时保留与原始时间单元的映射。检索到的粗略区域随后被扩展回原始块粒度,以进行细粒度证据的精炼和答案生成。在 MLVU、VideoMME 和 LongVideoBench 上的实验表明,DAGC 仅保留约 40% 至 50% 的原始图节点,并实现了 $1.3$ 至 $1.7 imes$ 的端到端时钟加速,同时保留了约 99% 的原始 QA 性能。这些增益在不同的 LVLM 主干和视频 RAG 流水线中都能迁移,表明长视频 RAG 不必在索引和证据推理中保持相同的时间粒度。
cs.CV / 138 / 2608.23012

Misanthrope: A Privacy-Preserving Keypoint Detector

Misanthrope:一种隐私保护的关键点检测器
Vultaggio, Francesco, Djindjic, Predrag, Gerke, Markus, Tschiatschek, Sebastian, Fanta-Jende, Phillipp
Abstract
Image matching is a core component of applications such as Simultaneous Localization and Mapping (SLAM), Visual Localization, and Structure from Motion (SfM). However, the local image features central to this task are vulnerable to inversion attacks, which enable adversaries to reconstruct privacy-sensitive scene content from local features. These attacks pose a particular threat in distributed computing scenarios where the pre-computed features leave edge devices to be processed by remote servers. In this work, we introduce Misanthrope, a novel privacy-preserving keypoint detector trained through self-distillation to avoid detecting keypoints on people---a predominant source of privacy-sensitive content in most localization scenarios---thus mitigating inversion attacks at the source rather than through post-hoc obfuscation. We demonstrate how inverted images from traditional feature detection pipelines can be used to detect and re-identify people in the scene, while Misanthrope is able to mitigate these attacks. Furthermore, Misanthrope maintains image matching performance on par with the state of the art and even surpasses it in challenging settings where people act as distractors, such as phototourism and in-the-wild odometry. On the Image Matching Challenge 2021 Phototourism test set, Misanthrope is the top-performing sparse feature extractor in 7 out of 9 scenes. We make our model and its evaluation script available here: https://github.com/fratopa/misanthrope
Chinese Translation
图像匹配是同时定位与地图构建(SLAM)、视觉定位和运动重建(SfM)等应用的核心组成部分。然而,任务中核心的局部图像特征容易受到反演攻击,这使得攻击者能够从局部特征重建隐私敏感的场景内容。这些攻击在分布式计算场景中尤其具有威胁,因为预先计算的特征会从边缘设备传输到远程服务器进行处理。在本研究中,我们介绍了Misanthrope,一种新颖的隐私保护关键点检测器,通过自我蒸馏进行训练,以避免检测到人类关键点——在大多数定位场景中隐私敏感内容的主要来源——从而在源头减轻反演攻击,而不是通过事后模糊处理来解决。我们展示了传统特征检测管道中的反演图像如何被用来检测和重新识别场景中的人,而Misanthrope能够减轻这些攻击。此外,Misanthrope在图像匹配性能上与最先进的技术相当,甚至在如摄影旅游和野外测距等具有挑战性的环境中超越了它。在2021年图像匹配挑战赛的摄影旅游测试集中,Misanthrope在9个场景中的7个场景中表现最佳。我们在此提供我们的模型及其评估脚本:https://github.com/fratopa/misanthrope
cs.CV / 139 / 2608.23014

AnaDiffusion: Anatomically CompositionalLatent Diffusion for Controllable 3D Brain MRI Generation

AnaDiffusion:用于可控3D脑MRI生成的解剖组成潜在扩散模型
Han, Huiwen, Liu, Lulin, Liu, Bangya, Cai, Yuanhao, Chen, Nuo, Wang, Xiaoqing, Xie, Ziqian, You, Chenyu, Ji, Shuiwang, Zhi, Degui, Fan, Zhiwen
Abstract
3D brain MRI generation has made significant advances in medical imaging, simulation, and controllable anatomical analysis. However, existing generative models typically synthesize 3D volumes monolithically, often overlooking regional anatomical structures and limiting local controllability. To address these limitations, we introduce AnaDiffusion, an anatomically compositional latent diffusion framework that factorizes the generation process into distinct, anatomically meaningful regions, followed by part-to-whole assembly and global refinement. Our approach first trains part diffusion models to capture local structural priors. We then inject an assembled anatomical composite of the parts into the whole-brain latent representation and continue denoising. This mechanism enables the model to resolve global context while preserving the injected anatomy. As a result, AnaDiffusion produces both explicit part assets and a globally coherent volume, thereby enabling controllable part editing without requiring subject-specific dense segmentation maps at inference time while maintaining consistent part-to-whole brain structure. On the subject-disjoint ADNI test split, AnaDiffusion achieves the lowest FID across the whole brain, left and right hemispheres, cerebellar-brainstem complex, and seam regions. It also achieves the best cerebellar and second-best ventricular and brainstem absolute Cohen's d values among the evaluated methods. In localized editing experiments, paired MS-SSIM demonstrates high target transfer and off-target preservation, supporting controllable part replacement with minimal unintended anatomical alterations.
Chinese Translation
3D脑MRI生成在医学成像、模拟和可控解剖分析方面取得了显著进展。然而,现有的生成模型通常以整体方式合成3D体积,往往忽视区域解剖结构,限制了局部可控性。为了解决这些局限性,我们提出了AnaDiffusion,一种解剖组成的潜在扩散框架,它将生成过程分解为不同的、具有解剖意义的区域,然后进行部分到整体的组装和全局优化。我们的方法首先训练部分扩散模型以捕捉局部结构先验。然后,我们将组装的解剖复合体注入整体脑的潜在表示中,并继续去噪。这一机制使模型能够在保留注入解剖结构的同时解析全局上下文。因此,AnaDiffusion同时生成明确的部分资产和全局一致的体积,从而实现可控的部分编辑,而无需在推理时依赖于特定个体的密集分割图,同时保持一致的部分到整体脑结构。在与受试者不重叠的ADNI测试集上,AnaDiffusion在整个脑、左半球、右半球、小脑-脑干复合体和缝合区域中实现了最低的FID值。它还在评估的方法中取得了最佳的小脑和第二好的脑室及脑干绝对Cohen's d值。在局部编辑实验中,配对的MS-SSIM显示出高目标转移和非目标保留,支持可控的部分替换,且对解剖结构的意外改变最小。
cs.CV / 140 / 2608.23024

When the Edit Changes the Patient: Measuring Identity Preservation in Counterfactual Retinal Images

当编辑改变患者:在反事实视网膜图像中测量身份保留
Posada, Andrea, Karbole, Wenke, Doan, Bach Ngoc, Weers, Alexander, Abdolrahimzadeh, Solmaz, Patsiamanidi, Maria, Haider, Kahkashan, Khare, Vaishali, Rueckert, Daniel, Lotery, Andrew, Sivaprasad, Sobha, Menten, Martin J.
Abstract
Counterfactual medical image generation aims to modify an existing image to reflect a hypothetical scenario in which certain characteristics of the imaged subject are altered, while keeping their identity fixed. Most existing works repurpose established image editing methods, which do not directly supervise identity preservation. Instead, they assume that identity is implicitly preserved by anchoring generation to the source image. This assumption is rarely tested and may fail in domains where biometric cues are subtle, such as retinal optical coherence tomography (OCT). In this work, we explicitly measure identity preservation for three groups of text-conditioned editing methods - source-anchored, structured-prompt, and paired-training - using referee classifiers, embedding alignment scores, and a blind reader study. We find that all methods produce high-quality OCT images with comparable editing success, yet their identity preservation differs markedly. Source-anchored editing frequently alters the depicted subject, while paired-training preserves it best. We argue that future work on medical counterfactual generation must explicitly measure and report identity preservation alongside image realism and editing success.
Chinese Translation
反事实医学图像生成旨在修改现有图像,以反映一种假设场景,其中成像对象的某些特征被改变,同时保持其身份不变。大多数现有工作重新利用已建立的图像编辑方法,这些方法并未直接监督身份保留。相反,它们假设通过将生成锚定于源图像,身份会被隐式保留。这一假设很少经过验证,并且在生物特征线索微妙的领域(如视网膜光学相干断层扫描(OCT))中可能会失败。在本研究中,我们使用裁判分类器、嵌入对齐分数和盲读者研究,明确测量三组文本条件编辑方法的身份保留——源锚定、结构化提示和配对训练。我们发现所有方法都能生成高质量的OCT图像,编辑成功率相当,但它们的身份保留差异显著。源锚定编辑经常改变所描绘的对象,而配对训练则能最好地保留身份。我们认为,未来的医学反事实生成工作必须明确测量并报告身份保留,同时考虑图像真实感和编辑成功率。
cs.CV / 141 / 2608.23065

Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia

文化时刻基准:评估东南亚视频文化推理与基础
Satar, Burak, Ma, Zhixin, Yu-Tong, Cheng, Tran, Huy Hoang, Nguyen, Phuong Anh, Ngo, Chong-Wah
Abstract
Cultural understanding in video means more than recognizing what is visible; it requires grasping the symbolic and temporal significance of cultural concepts. We decompose this into three abilities: naming what a concept symbolizes, visually recognizing it on video, and locating its sub-events in time. Existing video-cultural benchmarks tend to test what is seen, collapsing these three abilities into a single score that hides the bottleneck. We introduce the Cultural Moment Benchmark (CMB): 306 expert-curated concepts from seven countries in Southeast Asia across five categories. We evaluate each concept through three stages, one per ability. Given a description, Stage 1 (S1) selects from four candidate concept names, Stage 2 (S2) selects from four candidate video moments, and Stage 3 (S3) predicts the start and end times of the moment in a video. To keep each stage focused on a distinct ability, we use three design choices: semantic-similarity distractors (S1, S2), unlabeled video moments (S2), and free-form localization on a different example video (S3). Across six vision-language models, failure modes vary by ability and modality. i) Even the strongest closed-source models score below 30% when all three stages must be correct; ii) The three abilities do not fully cascade: naming a concept correctly helps half the models recognize it on video, but recognizing it has little effect on locating the sub-event in time; iii) Audio is complementary, redundant, or distracting depending on the concept, more often distracting in non-Latin-script countries; removing both audio and subtitles hurts Games and Music the most. Our 14-rater human study shows that even Expert raters score below chance on concepts from a neighboring country, indicating that CMB requires country-specific cultural knowledge. CMB acts as a diagnostic harness, attributing failures to a specific ability or modality.
Chinese Translation
视频中的文化理解不仅仅是识别可见的内容;它还需要把握文化概念的象征性和时间意义。我们将其分解为三种能力:命名一个概念所象征的内容、在视频中视觉识别该内容,以及在时间上定位其子事件。现有的视频文化基准往往测试可见的内容,将这三种能力合并为一个单一的评分,从而掩盖了瓶颈。我们引入了文化时刻基准(Cultural Moment Benchmark, CMB):来自东南亚七个国家的306个专家策划的概念,涵盖五个类别。我们通过三个阶段评估每个概念,每个阶段对应一种能力。给定一个描述,阶段1(Stage 1, S1)从四个候选概念名称中选择,阶段2(Stage 2, S2)从四个候选视频时刻中选择,阶段3(Stage 3, S3)预测视频中时刻的开始和结束时间。为了使每个阶段专注于特定的能力,我们采用了三种设计选择:语义相似性干扰项(S1, S2)、未标记的视频时刻(S2)以及在不同示例视频上自由形式的定位(S3)。在六种视觉-语言模型中,失败模式因能力和模态而异。i) 即使是最强的闭源模型在所有三个阶段都必须正确时得分低于30%;ii) 三种能力并未完全级联:正确命名一个概念有助于一半模型在视频中识别它,但识别它对在时间上定位子事件的影响很小;iii) 音频的作用因概念而异,可能是互补的、冗余的或分散注意力的,在非拉丁文字国家更常分散注意力;去除音频和字幕对游戏和音乐的影响最大。我们的14名评分者的人类研究表明,即使是专家评分者在邻国的概念上得分也低于随机水平,表明CMB需要特定国家的文化知识。CMB作为一种诊断工具,将失败归因于特定的能力或模态。
cs.CV / 142 / 2608.23074

Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?

基础不是知识:视觉语言模型(VLMs)在空间推理中是否需要物体定位?
Liu, Xiwei, Li, Yulong, Zhuang, Xinlin, Li, Xuhui, Lu, Zhixiang, Yang, Haolin, Razzak, Imran, Xie, Yutong
Abstract
Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understood. It is underexplored whether spatial reasoning internally requires precise objects localization, or can bypass explicit localization through global layout cues. In this work, we investigate two representative model families, LLaVA-1.5 and Qwen2.5-VL, using a suite of mechanistic interpretability tools, including token ablation, layer-wise probing, attention knockout, and causal mediation analysis. We find that spatial relation prediction follows a staged grounding-to-reasoning process in which object-aligned tokens establish coarse target-reference anchors, while precise bounding-box boundaries are not required. Positional information becomes decodable before relation decisions emerge, and a small set of attention heads mediates the causal effects of both localization and spatial reasoning. The two tasks share early grounding-related processing but ultimately rely on partially distinct specialized pathways. Through rigorous experiments, we provide a token-, layer-, and head-level account of how VLMs transform object grounding into spatial relations, showing that knowing where objects are is not equivalent to knowing how they relate.
Chinese Translation
视觉语言模型(VLMs)能够回答空间问题,但将物体定位与空间推理连接起来的机制仍然不甚清楚。目前尚未深入探讨空间推理是否内部需要精确的物体定位,或是否可以通过全局布局线索绕过显式定位。在本研究中,我们使用一系列机制可解释性工具,包括标记消融、层级探测、注意力剔除和因果中介分析,研究了两个代表性模型家族,LLaVA-1.5和Qwen2.5-VL。我们发现,空间关系预测遵循一个分阶段的基础到推理过程,其中物体对齐的标记建立了粗略的目标参考锚点,而精确的边界框边界并不是必需的。位置信息在关系决策出现之前就可以被解码,并且一小组注意力头介导了定位和空间推理的因果效应。这两个任务共享早期的基础相关处理,但最终依赖于部分不同的专业路径。通过严格的实验,我们提供了一个标记、层级和头级的解释,说明VLMs如何将物体定位转化为空间关系,表明知道物体的位置并不等同于知道它们之间的关系。
cs.CV / 143 / 2608.23090

Loopy: Seamless Video Loop Generation via Anchored Looping Shift of Positional Embedding

Loopy:通过锚定循环位移的定位嵌入实现无缝视频循环生成
Dong, Haotian, Wang, Wenjing, Li, Chen, Lyu, Jing, Wang, Xin, Lin, Di
Abstract
Looping videos are essential for practical applications such as web graphics, game development, and social media. However, existing approaches typically fail to generate high-quality looping videos due to the neglect of how video generation models perceive temporal order and how this relates to the looping behavior. In this work, we are the first to reveal that position embedding at different attention layers within DiT exhibits varying levels of positional control, with the most pronounced layer acting as an anchor. We formulate this anchored layer as the reference point of the looping video, offering strong contextual priors for the remaining layers to facilitate the generation of seamless and coherent video content. Based on this insight, we propose an anchored position embedding shifting strategy that applies layer-specific shift lengths according to each layer's temporal control effect, effectively transforming DiT's temporal perception from a straight line to a circle. Leveraging this strategy, we develop a general framework, Loopy, for high-quality looping video generation, supporting both RGB and RGBA videos, while also enabling advanced AIGC features such as identity control and style transfer. Experiments demonstrate that our approach significantly improves temporal consistency and visual fidelity in generated looping videos. The released model is available on our website: https://donghaotian123.github.io/Loopy.
Chinese Translation
循环视频在网页图形、游戏开发和社交媒体等实际应用中至关重要。然而,现有方法通常未能生成高质量的循环视频,因为忽视了视频生成模型如何感知时间顺序以及这与循环行为的关系。在本研究中,我们首次揭示了DiT中不同注意力层的位置信息嵌入展现出不同程度的位置信控,最显著的层作为锚点。我们将这一锚定层构建为循环视频的参考点,为其余层提供强大的上下文先验,以促进无缝且连贯的视频内容生成。基于这一洞察,我们提出了一种锚定位置嵌入位移策略,根据每层的时间控制效果应用特定的位移长度,有效地将DiT的时间感知从直线转变为圆形。借助这一策略,我们开发了一个通用框架Loopy,用于高质量循环视频生成,支持RGB和RGBA视频,同时还启用身份控制和风格迁移等高级AIGC功能。实验表明,我们的方法显著提高了生成循环视频的时间一致性和视觉保真度。发布的模型可在我们的网站上获取:https://donghaotian123.github.io/Loopy。
cs.CV / 144 / 2608.23102

Training-Free Pseudo-Fusion for Composed Image Retrieval with Diffusion Models and Multimodal Large Language Models

无训练伪融合用于扩展图像检索的扩散模型与多模态大型语言模型
Xu, Fan, Leiva, Luis A.
Abstract
Composed Image Retrieval (CIR) is an emerging paradigm in content-based image retrieval that enables users to formulate compositional queries by combining a reference image with an auxiliary modality, usually text-based. This approach supports fine-grained search where the target image shares structural elements with the user-provided image while incorporating the modifications specified by the auxiliary text. Conventional CIR methods rely on multimodal fusion to combine visual and textual features into a joint query embedding, which requires training modules that align composed queries with the targets. In this work, we propose PeFuse (for pseudo-fusion), a training-free framework that leverages pretrained Diffusion Models and Multimodal Large Language Models to bridge modalities via generative conversion. We introduce two novel strategies: uni-directional and bi-directional conversion, which convert CIR into four single-modality retrieval problems. These methods reformulate CIR as either intra-modal or cross-modal single-query retrieval tasks, bypassing the need for dedicated task-specific training. Extensive experiments on standard benchmarks demonstrate that converting CIR into text-to-image retrieval tasks is more effective than alternative conversion strategies, achieving competitive or superior performance compared with state-of-the-art methods, while maintaining high flexibility thanks to replaceable components of the conversion pipeline. These results highlight the effectiveness of the pseudo-fusion paradigm for zero-shot CIR. Our code is publicly available at: https://github.com/StevenXuf/PeFuse4CIR.
Chinese Translation
扩展图像检索(CIR)是基于内容的图像检索中一种新兴范式,使用户能够通过将参考图像与辅助模态(通常是基于文本的)结合来构建组合查询。这种方法支持细粒度搜索,其中目标图像与用户提供的图像共享结构元素,同时结合了辅助文本所指定的修改。传统的CIR方法依赖于多模态融合,将视觉和文本特征结合成一个联合查询嵌入,这需要训练模块来对齐组合查询与目标。在本研究中,我们提出了PeFuse(伪融合),这是一个无训练框架,利用预训练的扩散模型和多模态大型语言模型通过生成转换来连接模态。我们引入了两种新策略:单向转换和双向转换,将CIR转换为四个单模态检索问题。这些方法将CIR重新表述为内部模态或跨模态单查询检索任务,绕过了专门任务特定训练的需求。在标准基准上的大量实验表明,将CIR转换为文本到图像检索任务比其他转换策略更有效,达到了与最先进方法相媲美或更优的性能,同时由于转换管道的可替换组件保持了高度灵活性。这些结果突显了伪融合范式在零-shot CIR中的有效性。我们的代码已公开可用,网址为:https://github.com/StevenXuf/PeFuse4CIR。
cs.CV / 145 / 2608.23136

Bridge Damage Detection from Low-Light UAV Imagery via Degradation-Aware Mixture-of-Experts Enhancement

基于降质感知混合专家增强的低光照无人机影像桥梁损伤检测
Wang, Hu, Pu, Hongxu, Hu, Zhiqi, Lin, Fangzhou, Wang, Wang
Abstract
Poor illumination obscures small, low-contrast defects in UAV bridge imagery, reducing the reliability and operational flexibility of automated inspection. This paper investigates whether degradation-aware image restoration can improve bridge damage detection under low-light conditions and transfer from synthetic degradations to real inspection scenes. We propose DaL- MoE, a detector-agnostic restoration front end trained with an ISP-aware low-light synthesis pipeline and equipped with degradation-aware guidance estimation and complementary experts for noise suppression, color adjustment, and structural-detail recovery. On paired synthetic data, DaL-MoE achieves 23.12 dB PSNR and 0.8482 SSIM, increasing YOLOv11m box mAP50 from 0.3097 to 0.4923 and mask mAP50 from 0.2281 to 0.3529. On real low-light UAV imagery without paired normal-light references, sim-to-real evaluation shows improved defect visibility and more complete detections than direct inference on raw low-light inputs. Future work will develop low-light-aware bridge damage detectors with stronger cross-scene generalization across bridge sites, imaging conditions, and illumination levels.
Chinese Translation
低光照条件下,照明不足会掩盖无人机桥梁影像中的小型低对比度缺陷,从而降低自动检测的可靠性和操作灵活性。本文研究了降质感知图像恢复是否能够改善低光照条件下的桥梁损伤检测,并探讨其从合成降质到真实检测场景的迁移能力。我们提出了DaL-MoE,这是一种与检测器无关的恢复前端,使用ISP感知的低光照合成管道进行训练,并配备降质感知引导估计和补充专家,以实现噪声抑制、颜色调整和结构细节恢复。在配对的合成数据上,DaL-MoE达到了23.12 dB的PSNR和0.8482的SSIM,使YOLOv11m的框架mAP50从0.3097提高到0.4923,掩膜mAP50从0.2281提高到0.3529。在没有配对正常光照参考的真实低光照无人机影像上,模拟到真实的评估显示出比直接在原始低光输入上推理更好的缺陷可见性和更完整的检测结果。未来的工作将开发具有更强跨场景泛化能力的低光照桥梁损伤检测器,以适应不同桥梁地点、成像条件和照明水平。
cs.CV / 146 / 2608.23137

A Simulator-Grounded Framework For Constructing Verifiable Muscle-Grounded QA From 3D Tongue Meshes

基于模拟器的框架用于从3D舌头网格构建可验证的肌肉驱动问答
Eum, Seungho, Park, Unsang
Abstract
Existing articulatory corpora based on real-time MRI and electromagnetic articulography capture tongue shape and motion but do not provide traceable labels for the muscle-driven process that generated an observed configuration. We introduce a simulator-grounded data-construction framework and instantiate it as 3DTongueQA. Controlled 11-dimensional muscle activations are mapped to fixed-topology tongue meshes with the ArtiSynth Badin finite-element model, converted into structured biomechanical records, and rendered as deterministic QA on muscle state, geometry, and target-directed change. We screen 295,157 configurations, retain 295,115 valid meshes, and construct 891,156 QA records per language. Language naturalization changes only surface form and is verified against the source records; English and Korean instantiations demonstrate construction-level portability. A swappable SpiralNet++--Qwen3-8B baseline reaches 62.9 $\pm$ 9.2 Muscle EM, 74.0 $\pm$ 0.2 Value Accuracy, and 65.9 $\pm$ 4.7 Direction EM, while mismatching the paired mesh reduces Muscle EM to 2.2; a dataset-leakage-controlled anchor-held-out model retains 80.4--98.6\% of the full-inventory scores on unseen anchors. Task-specific structured readouts further reach 88.7 $\pm$ 0.7 Muscle EM and 93.3 $\pm$ 1.0 Direction EM. These complementary results show that the constructed supervision supports both efficient structured prediction and heterogeneous natural-language QA rather than being tied to a particular decoder architecture.
Chinese Translation
现有的基于实时MRI和电磁发音描记法的发音语料库捕捉舌头的形状和运动,但未提供可追溯的标签以标识生成观察到的配置的肌肉驱动过程。我们引入了一种基于模拟器的数据构建框架,并将其实例化为3DTongueQA。控制的11维肌肉激活映射到固定拓扑的舌头网格,使用ArtiSynth Badin有限元模型转换为结构化的生物力学记录,并呈现为关于肌肉状态、几何形状和目标导向变化的确定性问答。我们筛选了295,157个配置,保留了295,115个有效网格,并为每种语言构建了891,156个问答记录。语言自然化仅改变表面形式,并与源记录进行验证;英语和韩语的实例展示了构建级别的可移植性。可更换的SpiralNet++--Qwen3-8B基线模型达到了62.9 ± 9.2的肌肉EM,74.0 ± 0.2的价值准确率,以及65.9 ± 4.7的方向EM,而配对网格的不匹配将肌肉EM降低至2.2;一个控制数据集泄漏的锚定模型在未见锚点上保留了80.4--98.6%的完整库存分数。任务特定的结构化读出进一步达到了88.7 ± 0.7的肌肉EM和93.3 ± 1.0的方向EM。这些互补结果表明,构建的监督支持高效的结构化预测和异构自然语言问答,而不是依赖于特定的解码器架构。
cs.CV / 147 / 2608.23140

MIVIFI: Bridging Perspective and Fisheye Domains for Training Multi-View Fisheye Image Generation Models

MIVIFI:桥接视角和鱼眼域以训练多视角鱼眼图像生成模型
Neuwirth-Trapp, Matthias, Altunbas, Begüm, Wang, Jiayi, Xia, Yan, Bieshaar, Maarten, Huang, Xinyu, Cremers, Daniel
Abstract
Achieving 360{\deg} coverage is critical for the visual perception systems of autonomous vehicles. Fisheye cameras offer a cost-effective solution by enabling full surround coverage with as few as two sensors. However, existing multi-view fisheye datasets are limited, and synthesizing rare corner cases typically requires computationally expensive 3D simulations, hindering the training. While generative models have achieved significant success in standard perspective imagery, their application to wide-angle distortion remains unexplored. In this work, we formally introduce the novel problem of multi-view fisheye image generation conditioned on volumetric semantic representations and present two distinct methods. We first propose SyntheOcc-FE, which adapts the SyntheOcc architecture to fisheye data. While effective, this method is constrained by the scarcity of fisheye datasets, which limits its generalization. To overcome these limitations, we propose our second method, MIVIFI (multi-view fisheye), which leverages cross-domain learning with Equirectangular Projections. By bridging the gap between dataset domains using KITTI-360 fisheye images alongside nuScenes multi-view standard images, our approach enables high-fidelity manipulation of scene content. This framework enables the structural modification of semantic occupancy inputs to introduce or eliminate specific actors and facilitates the rendering of diverse meteorological conditions and illumination scenarios absent in the limited fisheye datasets. Quantitative and qualitative experiments demonstrate that our methods achieve robust photorealistic multi-view fisheye image generation and highlight the specific advantages of our cross-domain strategy for handling data scarcity.
Chinese Translation
实现360°覆盖对于自主车辆的视觉感知系统至关重要。鱼眼相机通过仅使用两个传感器即可实现全面覆盖,提供了一种具有成本效益的解决方案。然而,现有的多视角鱼眼数据集有限,合成稀有角落案例通常需要计算成本高昂的3D模拟,阻碍了训练。尽管生成模型在标准透视图像中取得了显著成功,但其在广角畸变中的应用仍未被探索。在本研究中,我们正式引入了基于体积语义表示的多视角鱼眼图像生成这一新问题,并提出了两种不同的方法。我们首先提出了SyntheOcc-FE,该方法将SyntheOcc架构调整为适应鱼眼数据。尽管有效,但该方法受到鱼眼数据集稀缺的限制,限制了其泛化能力。为克服这些限制,我们提出了第二种方法MIVIFI(多视角鱼眼),该方法利用等距投影进行跨域学习。通过使用KITTI-360鱼眼图像与nuScenes多视角标准图像之间的桥接,我们的方法实现了场景内容的高保真操控。该框架使得对语义占用输入的结构性修改成为可能,以引入或消除特定的参与者,并促进渲染在有限鱼眼数据集中缺失的多样气象条件和光照场景。定量和定性实验表明,我们的方法实现了稳健的照片级真实感多视角鱼眼图像生成,并突出了我们跨域策略在处理数据稀缺方面的具体优势。
cs.CV / 148 / 2608.23142

How Merge-Tolerant Are Vision Transformers for Wheat Phenotyping?

视觉变换器在小麦表型分析中的合并容忍度如何?
Ravé, Simon, Rasti, Pejman, Rousseau, David
Abstract
Vision-based wheat phenotyping requires repeated measurements under deployment constraints, from growth-stage recognition to wheat-head counting and organ segmentation. Plain Vision Transformers (ViTs) provide a common architecture for these tasks, but quadratic attention limits high-throughput and edge inference. Training-free token merging is attractive because it can be inserted into trained models without retraining. We provide a systematic benchmark of ToMe and Mutual Pair Merging across growth-stage classification, wheat-head detection, and wheat-organ segmentation, measuring task quality, throughput, token count, and peak GPU memory, with additional Raspberry Pi 5 measurements. The benchmark reveals a clear hierarchy: classification is highly merge-tolerant, while detection and segmentation are constrained by repeated instances, thin organs, dense boundaries, reconstruction, and runtime overhead. Optimized attention backends can erase apparent speedups, so deployment value must be profiled on the target runtime rather than inferred from token count.
Chinese Translation
基于视觉的小麦表型分析需要在部署约束下进行重复测量,从生长阶段识别到小麦穗计数和器官分割。普通的视觉变换器(ViTs)为这些任务提供了一个通用架构,但二次注意力限制了高通量和边缘推理。无训练的令牌合并具有吸引力,因为它可以在不重新训练的情况下插入到已训练的模型中。我们对ToMe和互配合并在生长阶段分类、小麦穗检测和小麦器官分割中的系统基准进行了评估,测量了任务质量、吞吐量、令牌数量和峰值GPU内存,并附加了Raspberry Pi 5的测量结果。基准测试揭示了明显的层次结构:分类具有很高的合并容忍度,而检测和分割则受到重复实例、细小器官、密集边界、重建和运行时开销的限制。优化的注意力后端可能会消除明显的加速,因此部署价值必须在目标运行时进行分析,而不是仅仅通过令牌数量推断。
cs.CV / 149 / 2608.23143

An end-to-end-trained vision-language model for native-language prostate pathology report generation

端到端训练的视觉-语言模型用于本土语言前列腺病理报告生成
Grashei, Christian, Gülhan, Fabian, Legnar, Maximilian, Stögbauer, Fabian, Weis, Cleo-Aron, Mogler, Carolin, Schüffler, Peter
Abstract
Prostate cancer is among the most frequently diagnosed malignancies worldwide, and structured reporting of each biopsy core burdens pathologists. Existing tools frame this as classification, leaving pathologists to assemble coherent reports, while many slide-level vision-language models rely on English-centric encoders that transfer poorly to other clinical languages. We present a slide-level framework generating prostate biopsy reports that is language-independent by construction: tokenizer and model are trained from scratch, demonstrated here in German. To address paired-data scarcity, an automated pipeline uses a locally deployed large language model to split composite reports into core-specific image-text pairs, yielding 17,344 pairs from 2,402 historical cases without manual annotation. Evaluated for clinical attributes rather than linguistic similarity, the model achieves 96.2% F1 for malignancy detection and 65.2% for Gleason grading, competitive with an FDA-cleared classifier. Grading is further validated on three external cohorts with latent-space augmentation. Institutions can thus train native-language reporting models on their own archives.
Chinese Translation
前列腺癌是全球最常被诊断的恶性肿瘤之一,而每个活检核心的结构化报告给病理学家带来了负担。现有工具将此问题框架化为分类,导致病理学家需要自行组装连贯的报告,同时许多基于幻灯片的视觉-语言模型依赖于以英语为中心的编码器,难以有效转移到其他临床语言。我们提出了一种幻灯片级框架,生成语言独立的前列腺活检报告:分词器和模型从零开始训练,这里以德语为例。为了解决配对数据稀缺的问题,自动化管道利用本地部署的大型语言模型将复合报告拆分为核心特定的图像-文本对,从2402个历史案例中生成17344对,无需手动标注。该模型在临床属性评估中表现优异,恶性肿瘤检测的F1值达到96.2%,Gleason分级的F1值为65.2%,与FDA批准的分类器相当。分级进一步在三个外部队列中通过潜在空间增强进行了验证。因此,机构可以在自己的档案上训练本土语言报告模型。
cs.CV / 150 / 2608.23173

BenthicFlow: Generating Extensible Underwater Environments via Flow Matching

BenthicFlow:通过流匹配生成可扩展的水下环境
Figueira, Joaquín, Lendering, Camile, Gonzalez-Hernandez, Manfred, D'Amicantonio, Giacomo, Akdag, Erkut, Bondarev, Egor
Abstract
Computer vision applications for 3D scene understanding in underwater environments remain challenging due to the lack of high-quality 3D data and the inability of surface-trained models to generalize to underwater scenes. To address this challenge, an emerging trend is to employ generative models to close the data domain gap. However, existing methods assemble large scenes by stitching independently generated tiles post hoc with separately trained models, while demonstrating heterogeneous landscapes only within individual survey sites. We introduce BenthicFlow, a unified framework based on a single conditional flow-matching model that jointly generates aligned textures and depth maps. A MultiDiffusion-inspired sampling procedure reconciles overlapping windows throughout the generative trajectory, enabling spatially extensible RGBD mosaics without a separate stitching model. The generated mosaics are subsequently lifted into explicit 3D benthic environments using surface-aligned Gaussian surfels. Experiments across geographically distinct survey sites demonstrate that BenthicFlow preserves site-specific appearance while generating coherent, large-scale 3D scenes that closely match the target distributions. Code and trained models are available at https://github.com/jacomof/BenthicFlow.
Chinese Translation
由于缺乏高质量的三维数据以及在水面训练的模型难以推广至水下场景,计算机视觉在水下环境三维场景理解方面仍面临挑战。为了解决这一问题,近年来兴起的趋势是采用生成模型以弥合数据域差异。然而,现有方法通常通过后期拼接独立生成的图块,并使用分别训练的模型组装大型场景,且仅在单个调查点展示异质景观。我们提出了BenthicFlow,一种基于单一条件流匹配(conditional flow-matching)模型的统一框架,能够联合生成对齐的纹理和深度图。受MultiDiffusion启发的采样过程在生成轨迹中协调重叠窗口,实现无需额外拼接模型的空间可扩展RGBD拼接图。生成的拼接图随后通过与表面对齐的高斯表面元(Gaussian surfels)提升为显式的三维底栖环境。在多个地理位置不同的调查点上的实验表明,BenthicFlow在生成与目标分布高度匹配的连贯大规模三维场景的同时,保持了场地特有的外观特征。代码和训练模型可在https://github.com/jacomof/BenthicFlow获取。
cs.CV / 151 / 2608.23175

Neighbor-Aware View Synthesis for Restoring Missing Views in Light-Field Camera Arrays

邻域感知视图合成用于恢复光场相机阵列中的缺失视图
Goel, Sakshi, Goyal, Ayush, Venkatesh, K S, Jerripothula, Koteswar Rao
Abstract
In light-field (LF) imaging systems, dense spatial sampling from a camera array enables powerful post-capture capabilities such as refocusing and depth estimation. However, real-world LF capture is often affected by hardware malfunctions, where one or more cameras in the array fail, leading to missing sub-aperture images and degraded reconstruction quality. This paper addresses the problem of defective or missing view restoration in light-field camera arrays. We propose a novel generative framework that synthesizes the absent views by exploiting information from a carefully selected subset of neighboring cameras. These selected images, along with a positional encoding map indicating both their locations and the desired target view, are fed into a conditional Generative Adversarial Network (cGAN) trained to generate the missing viewpoint in a geometrically consistent manner. Extensive experiments on synthetic and real-world LF datasets demonstrate that our method produces visually plausible and photometrically accurate reconstructions, outperforming baselines for view interpolation both quantitatively and qualitatively. The proposed framework thus offers a robust and efficient solution for fault-tolerant light-field image acquisition.
Chinese Translation
在光场(LF)成像系统中,来自相机阵列的密集空间采样使得强大的后期捕获能力成为可能,如重新聚焦和深度估计。然而,现实世界中的光场捕获常常受到硬件故障的影响,其中一个或多个相机发生故障,导致缺失子孔径图像和重建质量下降。本文解决了光场相机阵列中缺陷或缺失视图恢复的问题。我们提出了一种新颖的生成框架,通过利用从精心选择的邻近相机子集获取的信息来合成缺失的视图。这些选定的图像以及一个指示其位置和所需目标视图的位置信息编码图被输入到一个条件生成对抗网络(cGAN)中,该网络经过训练以几何一致的方式生成缺失的视点。在合成和真实世界的光场数据集上进行的广泛实验表明,我们的方法在视觉上产生可信且光度上准确的重建,在定量和定性上均优于视图插值的基线。因此,所提出的框架为容错光场图像采集提供了一种稳健且高效的解决方案。
cs.CV / 152 / 2608.23189

EchoWM: Open and Enterable Omnimodal World Models

EchoWM:开放且可进入的全模态世界模型
Zhang, Songchun, Li, Yaowei, Zhuang, Junhao, Jin, Weiyang, Wang, Haoyu, Lu, Xin, Sun, Yilang, Zhang, Shiyi, Li, Haoran, Ma, Xiaoxiao, Li, Yuming, Liu, Yijun, Su, Yaofeng, Ma, Yanwen, Wu, Haoyu, Su, Zihan, Ma, Yue, Zhang, Lvmin, Huang, Haoyang, Xue, Zeyue, Rao, Anyi, Duan, Nan
Abstract
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data. To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation. Extensive evaluations show that \model achieves strong trajectory following and high visual quality on public world-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech over long-horizon generation.
Chinese Translation
我们提出了EchoWM,一种用于可进入生成媒体的全模态世界模型,它能够响应连续导航,同时生成720p视频、环境声音、音乐和语音。我们围绕相机意图组织交互:在第一人称场景中,它指定观察者的运动,而在第三人称场景中,相机与角色的动态则是从数据中学习的,无需视图特定的控制器。离散命令和连续姿态被映射到一个共享的度量尺度相对6自由度轨迹上,数据集级别的校准保持了异构数据间的运动幅度。为了共同学习音频-视觉生成和轨迹控制,我们构建了一个互补的数据引擎,并采用渐进式训练,随后进行自回归后训练以实现长时间生成。大量评估表明, extit{model}在公共世界模型基准测试中实现了强大的轨迹跟踪和高视觉质量,支持第一人称和第三人称在不同主题间的交互,并在长时间生成中保持环境声音和语音的同步。
cs.CV / 153 / 2608.23190

Toward a Foundation Plug-and-Play Prior for Computed Tomography Reconstruction via a Multimodal Diffusion Model

基于多模态扩散模型的计算机断层扫描重建的可插拔先验基础
Duba-Sullivan, Haley, Fernandez-Zelaia, Patxi, Rahman, Obaidullah, Ziabari, Amirkoushyar
Abstract
Computed tomography (CT) throughput is limited by scan time, which grows with both the number of projections acquired and the detector integration time for each. Reconstructing high-quality volumes from sparse-view or low-dose measurements therefore depends on an informative prior, typically a neural network trained for one specific scan setting and retrained whenever the modality, geometry, or material changes. We investigate whether a single diffusion model trained across several imaging domains can instead serve as a prior for many CT problems simultaneously. We evaluate the proposed method using the same frozen model on three datasets that differ in modality, beam geometry, material, and degradation type, spanning flaw analysis in additively manufactured metal parts imaged with cone-beam X-ray CT and concrete microstructure imaged with parallel-beam neutron CT. Our proposed method out-performs analytic reconstructions in all three cases, providing a step toward a reusable foundation prior for heterogeneous CT reconstruction problems.
Chinese Translation
计算机断层扫描(CT)的吞吐量受到扫描时间的限制,而扫描时间随着采集的投影数量和每个投影的探测器积分时间的增加而增长。因此,从稀疏视图或低剂量测量中重建高质量体积依赖于一个信息丰富的先验,通常是为特定扫描设置训练的神经网络,并在每次模态、几何或材料变化时重新训练。我们研究了是否可以使用一个在多个成像领域训练的单一扩散模型作为许多CT问题的先验。我们使用相同的冻结模型在三个数据集上评估所提出的方法,这些数据集在模态、光束几何、材料和降解类型上存在差异,涵盖了使用锥束X射线CT成像的增材制造金属部件的缺陷分析和使用平行束中子CT成像的混凝土微观结构。我们提出的方法在所有三个案例中均优于解析重建,为异构CT重建问题提供了可重用的基础先验的一个步骤。
cs.CV / 154 / 2608.23206

Learning Spherical Occupancy Profiles for Multi-View 3D Reconstruction and Generation

学习球形占用轮廓用于多视角三维重建与生成
Tsai, YiHsuan
Abstract
We study spherical occupancy profiles-the ray-wise occupancy probability profiles P(r) = T(r) o(r) distilled from multi-view 3D Gaussian reconstructions-as a unified intermediate representation for both discriminative and generative 3D reconstruction from images. On a 999-object subset of Google Scanned Objects with 48 turntable views each, we train (i) a discriminative per-ray decoder that injects global view-averaged and ray-specific image evidence into a FiLM-conditioned profile head, reaching median soft depth error 0.035 (normalized) on an independent 90-object test split, and (ii) a generative pipeline built on a profile VAE and a latent diffusion model, which supports unconditional sampling that matches the reconstruction manifold and image-conditioned multi-solution reconstruction whose per-object solution spread is quantifiable and tunable via classifier-free guidance. We further analyze the morphology of predicted profiles: post-hoc power sharpening and a learned sharpening target both recover ground-truth profile width without degrading depth, exposing a monotonic width-peak frontier in the L1-per-ray loss family and motivating a principled redefinition of morphology gates. Real-photo validation on two DTU scenes confirms the pipeline transfers to non-synthetic input. Our results suggest that ray-wise occupancy profiles offer a compact, learned, and uncertainty-aware interface between multi-view reconstruction and generative priors.
Chinese Translation
我们研究球形占用轮廓——从多视角三维高斯重建中提炼出的按光线划分的占用概率轮廓 P(r) = T(r) o(r)——作为一种统一的中间表示,用于从图像中进行判别性和生成性的三维重建。在一个包含999个对象的Google扫描对象子集上,每个对象有48个转盘视图,我们训练了(i)一个判别性每光线解码器,该解码器将全局视图平均和光线特定的图像证据注入到FiLM条件的轮廓头中,在独立的90对象测试集上达到了中位数软深度误差0.035(归一化),以及(ii)一个基于轮廓变分自编码器(profile VAE)和潜在扩散模型的生成管道,支持无条件采样,匹配重建流形,并进行图像条件的多解重建,其每个对象的解的分布是可量化和可调的,通过无分类器引导进行调节。我们进一步分析了预测轮廓的形态:后处理的功率增强和学习的增强目标都能在不降低深度的情况下恢复真实轮廓宽度,揭示了L1每光线损失家族中的单调宽度-峰值边界,并促使对形态门的原则性重新定义。在两个DTU场景上的真实照片验证确认了该管道能够转移到非合成输入。我们的结果表明,按光线划分的占用轮廓提供了一种紧凑、学习的、并且具有不确定性意识的接口,连接多视角重建与生成先验。
cs.CV / 155 / 2608.23213

Bee Detection and Tracking at Hive Entrance using YOLO11 and ByteTrack

基于YOLO11和ByteTrack的蜂巢入口蜜蜂检测与追踪
Nguyen, Thi Thu Thao, Reschke, Johannes
Abstract
This work presents an automatic bee entrance monitoring system based on YOLO11 transfer learning and the ByteTrack tracking algorithm. The study investigates the influence of data augmentation, backbone freezing, and tracker parameter optimization on the detection and counting of small, fast-moving bees. The detector with progressive backbone unfreezing strategy achieved about 97.0% precision and 98.7% mAP50, while providing more stable convergence than full fine-tuning. Experiments also showed that light augmentation outperformed heavy augmentation. For tracking, ByteTrack parameters were optimized to improve trajectory continuity under low-confidence detections. On an independent 25 FPS side-view video, the optimized YOLO11-ByteTrack system correctly counted 43 of 47 incoming bees (91.5%) and 7 of 30 outgoing bees (23.3%). Error analysis showed that most counting errors were caused by missed detections due to rapid bee motion and motion blur, while tracking failures became less frequent after parameter optimization. Overall, the results indicate that moderate augmentation, progressive backbone unfreezing, and ByteTrack tuning improve the reliability of automatic bee entrance monitoring under realistic recording conditions.
Chinese Translation
本研究提出了一种基于YOLO11迁移学习和ByteTrack追踪算法的自动蜜蜂入口监测系统。研究探讨了数据增强、主干网络冻结和追踪器参数优化对小型快速移动蜜蜂的检测与计数的影响。采用逐步解冻主干网络策略的检测器实现了约97.0%的精确率和98.7%的mAP50,同时提供了比完全微调更稳定的收敛性。实验还表明,轻度增强优于重度增强。在追踪方面,优化了ByteTrack参数,以提高低置信度检测下的轨迹连续性。在一个独立的25 FPS侧视视频中,优化后的YOLO11-ByteTrack系统正确计数了47只入巢蜜蜂中的43只(91.5%)和30只出巢蜜蜂中的7只(23.3%)。误差分析显示,大多数计数错误是由于蜜蜂快速运动和运动模糊导致的漏检,而在参数优化后,追踪失败的频率减少。总体而言,结果表明适度的增强、逐步解冻主干网络和ByteTrack调优提高了在现实录制条件下自动蜜蜂入口监测的可靠性。
cs.CV / 156 / 2608.23215

BenthicDINO: Physics-Informed Self-Distillation for View-Invariant Side-Scan Sonar Representations

BenthicDINO:基于物理知识的自蒸馏方法用于视角不变的侧扫声纳表示学习
Hamoda, Taqi, Rajani, Hayat, Gracias, Nuno
Abstract
Automated perception in side-scan sonar (SSS) imagery is severely hindered by physical acoustic artifacts, resulting in representations that inextricably mix intrinsic seabed reflectivity with transient viewing geometries. Existing self-supervised learning (SSL) frameworks rely on augmentations designed for natural images, failing to account for acoustic degradation and explicitly enforce view-invariance. To address this gap, we introduce a physics-informed self-distillation framework built upon the DINOv3 architecture utilizing a ConvNeXt-v2-Tiny backbone to maximize data efficiency. The proposed methodology enforces view-invariance through two primary mechanisms: physically motivated augmentations that simulate speckle noise, range-dependent attenuation, and radiometric miscalibration; and a Hilbert-Schmidt Independence Criterion (HSIC) penalty that explicitly decouples learned dense patch features from physical viewing parameters. Furthermore, we propose a dense, hierarchical feature fusion strategy across all four network stages to preserve fine-grained sediment details alongside deep semantic abstractions. Extensive evaluation demonstrates that the framework natively groups complex benthic topographies into stable, noise-free semantic clusters without relying on manual annotations. During supervised downstream tasks on the S3Seg dataset, the fused representations exhibited exceptional data efficiency, achieving 96% of its absolute peak performance using only 10% of the available annotated data, ultimately reaching a mean Intersection over Union (mIoU) of 71.4% and an overall accuracy of 86.5%.
Chinese Translation
侧扫声纳(SSS)图像中的自动感知严重受限于物理声学伪影,导致表示无法区分内在的海床反射率与瞬时的视角几何。现有的自监督学习(SSL)框架依赖于为自然图像设计的数据增强方法,未能考虑声学退化且未明确实现视角不变性。为填补这一空白,我们提出了一种基于物理知识的自蒸馏框架,构建于DINOv3架构之上,采用ConvNeXt-v2-Tiny骨干网络以最大化数据效率。该方法通过两大机制实现视角不变性:一是物理驱动的数据增强,模拟斑点噪声、距离相关衰减及辐射计误校准;二是引入Hilbert-Schmidt独立性准则(HSIC)惩罚项,显式解耦学习到的密集图像块特征与物理视角参数。此外,我们提出了一种跨越网络四个阶段的密集层次特征融合策略,以在保留细粒度沉积物细节的同时兼顾深层语义抽象。大量评估表明,该框架能够原生地将复杂的海底地形聚类为稳定且无噪声的语义簇,无需依赖人工标注。在S3Seg数据集上的监督下游任务中,融合表示展现出卓越的数据效率,仅用10%的标注数据即达到绝对峰值性能的96%,最终实现了71.4%的平均交并比(mIoU)和86.5%的整体准确率。
cs.CV / 157 / 2608.23234

MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge

基于MLLM的音频引导视频目标分割:第八届LSVOS挑战赛MeViS-Audio赛道的第三名报告
Shi, Liangtao, Xie, Jinxia, Hu, Xiantao, Liu, Ting
Abstract
In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language Models (MLLMs) with SAM-based segmentation models. We decompose the task into several stages and identify suitable foundation models for each stage. Without introducing additional model training or task-specific fine-tuning, our approach leverages the strong multimodal reasoning capabilities of MLLMs to model text-visual correspondence and employs SAM-based models for accurate object mask generation. The proposed framework demonstrates the effectiveness of leveraging foundation models for audio-guided video segmentation and achieves competitive performance in the MeViS-Audio Track of the 8th LSVOS Challenge.
Chinese Translation
在本技术报告中,我们提出了一种无训练框架用于音频引导的视频目标分割,该框架将多模态大型语言模型(Multimodal Large Language Models, MLLMs)与基于SAM的分割模型相结合。我们将任务分解为多个阶段,并为每个阶段确定合适的基础模型。在不引入额外模型训练或任务特定微调的情况下,我们的方法利用MLLMs强大的多模态推理能力来建模文本与视觉之间的对应关系,并采用基于SAM的模型进行准确的目标掩膜生成。所提出的框架展示了利用基础模型进行音频引导视频分割的有效性,并在第八届LSVOS挑战赛的MeViS-Audio赛道中取得了具有竞争力的表现。
cs.CV / 158 / 2608.23238

Mover360: Controllable Object Manipulation in 360{\deg} Panoramic Images

Mover360:360°全景图像中的可控物体操控
Zhong, Haoyi, Zhang, Fang-Lue, Chalmers, Andrew, Rhee, Taehyun
Abstract
We present Mover360, a controllable object manipulation framework for 360{\deg} images. Unlike perspective images, 360{\deg} images in equirectangular projection (ERP) exhibit horizontal wrap-around, latitude-dependent distortion, and global scene continuity, which makes object-level edits difficult for existing perspective editors to produce and for users to specify. To address this, Mover360 centers on object Translation (relocating a specified object within an existing panorama) while supporting reference-guided Insert and Remove as auxiliary tasks. Its interface unifies point-, bbox-, and mask-guided control by encoding each task into a fixed prompt and a compact, ERP-aligned instruction map. In the default point mode, a single click relocates an object, allowing the model to infer a plausible size, support, and illumination using panoramic context and an auxiliary depth condition. Structurally, Mover360 is a lightweight adaptation of a pretrained diffusion transformer. To generate paired supervision, we construct a UE5 data-generation pipeline with surface-aware object placement and randomized illumination, yielding large-scale paired data and a dual-domain benchmark of synthetic and real panoramas with ground truth for all three tasks. Across both test domains and two evaluation protocols, Mover360 outperforms strong baselines for perspective editing, insertion, and inpainting in reconstruction fidelity, semantic consistency, and distributional quality. Code and our benchmark dataset are available at https://zhonghaoyi.github.io/Mover360/.
Chinese Translation
我们提出了Mover360,一个用于360°图像的可控物体操控框架。与透视图像不同,360°图像在等矩形投影(ERP)中表现出水平环绕、纬度依赖的失真和全局场景连续性,这使得现有透视编辑器在进行物体级编辑时面临困难,用户也难以指定。为了解决这一问题,Mover360专注于物体的平移(在现有全景中重新定位指定物体),同时支持参考引导的插入和删除作为辅助任务。其界面通过将每个任务编码为固定提示和紧凑的ERP对齐指令图,统一了点、边界框和掩码引导控制。在默认的点模式下,单击一次即可重新定位一个物体,模型能够利用全景上下文和辅助深度条件推断出合理的大小、支持和照明。在结构上,Mover360是一个轻量级的预训练扩散变换器的改编。为了生成配对监督,我们构建了一个UE5数据生成管道,采用表面感知的物体放置和随机照明,生成大规模配对数据以及合成和真实全景的双域基准测试,所有三个任务都有真实值。在两个测试域和两个评估协议中,Mover360在重建保真度、语义一致性和分布质量方面超越了强基线,表现出色。代码和我们的基准数据集可在https://zhonghaoyi.github.io/Mover360/获取。
cs.CV / 159 / 2608.23249

Semantic Reconstruction and 3-D Detection via Learned Multi-Pair Fusion in RF Imaging

通过学习的多对融合实现语义重建和三维检测在射频成像中的应用
Rezaei, Amir, Pan, Wen-Xin, Caire, Giuseppe
Abstract
We consider a multistatic radio-frequency imaging problem with anisotropy, in which the reflection from a point depends on the positions of the transmit (Tx) and receive (Rx) arrays. The goal is to label the voxels of a field of view by a finite set of semantic classes and to group them into object instances. For the image formation of each Tx--Rx pair we apply a standard inverse-problem solver, and we feed the resulting per-pair reconstructions into a trained three-dimensional (3-D) U-Net that performs the fusion implicitly and the per-voxel classification explicitly. On a controlled, under-determined multistatic setup, we consider the following image formation methods: back-projection (BP) and the least absolute shrinkage and selection operator (LASSO) from a single deterministic snapshot, and incoherent BP and group-LASSO from multiple fading snapshots. For each imaging method we train a separate U-Net that fuses the six Tx--Rx pairs (its input channels) and assigns each voxel a probability vector over the classes. Taking the most probable class gives a labeled volume---the semantic reconstruction. Object instances and their oriented bounding boxes then follow by geometric post-processing (clustering and principal-component analysis). Across a wide range of signal-to-noise ratio, the semantic reconstruction (scored against ground truth by segmentation intersection-over-union) and the resulting 3-D detection degrade far more gracefully than the classical intensity reconstruction: the detection in particular stays reliable well into noise levels at which that reconstruction has dissolved. Because real scenes contain objects of classes the network was not trained on, we add an explicit unknown class trained by outlier exposure, which labels held-out novel objects as unknown instead of mislabeling them as a known class by reconstructed shape.
Chinese Translation
我们考虑一个具有各向异性的多静态射频成像问题,其中来自一个点的反射取决于发射(Tx)和接收(Rx)阵列的位置。目标是通过有限的语义类别集合对视场中的体素进行标记,并将其分组为对象实例。对于每个 Tx-Rx 对的图像形成,我们应用标准的逆问题求解器,并将得到的每对重建结果输入到一个经过训练的三维(3-D)U-Net 中,该网络隐式地进行融合并显式地进行每个体素的分类。在一个受控的欠定多静态设置中,我们考虑以下图像形成方法:从单个确定性快照的反向投影(BP)和最小绝对收缩和选择算子(LASSO),以及从多个衰落快照的非相干 BP 和组 LASSO。对于每种成像方法,我们训练一个单独的 U-Net,该网络融合六个 Tx-Rx 对(其输入通道),并为每个体素分配一个类别的概率向量。选择最可能的类别得到一个标记体积——即语义重建。然后,通过几何后处理(聚类和主成分分析)获得对象实例及其定向边界框。在广泛的信噪比范围内,语义重建(通过分割交并比与真实值对比评分)及其产生的三维检测的降级远比经典的强度重建更加平滑:特别是检测在噪声水平较高时仍然保持可靠,而此时重建已失效。由于真实场景中包含网络未训练过的类别对象,我们通过异常值曝光添加了一个显式的未知类别,该类别将未见的新对象标记为未知,而不是通过重建形状错误标记为已知类别。
cs.CV / 160 / 2608.23253

E2S-Pruner: Progressive Two-Stage Evidence Fusion for Visual Token Pruning in Vision-Language Models

E2S-Pruner:视觉语言模型中视觉标记修剪的渐进式两阶段证据融合
Qian, Taoyu, Wang, Qi, Shi, Daqian, Jiang, Yuanhao, Gao, Shang, Yu, Hualong
Abstract
Vision-language models typically encode an image into hundreds of visual tokens, incurring substantial inference latency and GPU memory overhead. Existing pruning methods largely rely on attention scores and directly aggregate outputs across attention heads and network layers, making it difficult to characterize evidential uncertainty and conflict. We propose E2S-Pruner, a progressive two-stage evidence-fusion framework for visual token pruning that requires no auxiliary model, trainable parameters, or fine-tuning. In the first stage, E2S-Pruner treats each attention head as an independent evidence source, estimates its reliability from evidence clarity and inter-head consistency, and represents each visual token using three states: important, unimportant, and uncertain. In the second stage, Dempster--Shafer evidence theory is used to quantify inter-layer conflict and fuse complementary evidence from multiple network layers. We further introduce a spatial novelty constraint that promotes coverage of distinct image regions and prevents the retained tokens from concentrating in a few locally salient areas. On LLaVA-1.5-7B, E2S-Pruner retains 98.0%, 96.8%, and 90.6% of the aggregate performance when the average numbers of retained visual tokens are 192, 128, and 64, respectively, while improving throughput by 1.96x and 2.09x under the 128-token and 64-token settings. Experiments on Qwen2-VL-7B further demonstrate cross-model generalization. Code is available at https://github.com/taoyu-qian/E2S-Pruner.git.
Chinese Translation
视觉语言模型通常将图像编码为数百个视觉标记,这会导致显著的推理延迟和GPU内存开销。现有的修剪方法主要依赖于注意力分数,并直接在注意力头和网络层之间聚合输出,这使得表征证据的不确定性和冲突变得困难。我们提出了E2S-Pruner,一种渐进式两阶段证据融合框架,用于视觉标记修剪,该框架不需要辅助模型、可训练参数或微调。在第一阶段,E2S-Pruner将每个注意力头视为独立的证据源,根据证据的清晰度和头间一致性来估计其可靠性,并使用三种状态表示每个视觉标记:重要、不重要和不确定。在第二阶段,使用Dempster-Shafer证据理论来量化层间冲突,并融合来自多个网络层的互补证据。我们进一步引入了一种空间新颖性约束,以促进对不同图像区域的覆盖,并防止保留的标记集中在少数局部显著区域。在LLaVA-1.5-7B上,当保留的视觉标记平均数量分别为192、128和64时,E2S-Pruner分别保留了98.0%、96.8%和90.6%的整体性能,同时在128标记和64标记设置下将吞吐量提高了1.96倍和2.09倍。对Qwen2-VL-7B的实验进一步证明了跨模型的泛化能力。代码可在https://github.com/taoyu-qian/E2S-Pruner.git获取。
cs.CV / 161 / 2608.23258

Progressively Learning Heterogeneous Skills in a Unified Latent Space

在统一潜在空间中逐步学习异构技能
Zhang, Yue-Yi, Gong, Ming, He, Linpu, Zheng, Wei-Shi, Zhao, Zhilin
Abstract
We propose HetSkills, a novel framework designed to progressively learn heterogeneous skills within a unified latent space for physics-based character control. The core idea is to treat this latent space as a shared executable interface, enabling seamless integration of skills learned from diverse data sources, supervision forms, and tasks. HetSkills begins by learning a tracking skill that establishes a strong foundation in motion control and creates a shared motion decoder, which can be reused across tasks without the need for retraining or separate controllers. To prevent the text-to-motion skill from exploiting shortcut pathways instead of learning language semantics, we introduce motion intuition distillation to ground text-to-motion generation in language semantics and a task-guidance module that dynamically adjusts actions based on high-level language instructions. This enables HetSkills to preserve natural motion while continuously expanding its skill repertoire, making it highly adaptable for long-horizon tasks. Experimental results demonstrate the effectiveness in motion tracking, text-to-motion generation, motion completion, and downstream task adaptation, achieving impressive success rates even under challenging conditions.
Chinese Translation
我们提出了HetSkills,一个新颖的框架,旨在在统一潜在空间中逐步学习异构技能,以实现基于物理的人物控制。其核心思想是将这个潜在空间视为一个共享的可执行接口,从而实现从不同数据源、监督形式和任务中学习的技能的无缝集成。HetSkills首先学习一种跟踪技能,为运动控制奠定坚实基础,并创建一个共享的运动解码器,该解码器可以在不同任务中重复使用,而无需重新训练或单独的控制器。为了防止文本到运动技能利用捷径路径而不是学习语言语义,我们引入了运动直觉蒸馏,将文本到运动生成与语言语义相结合,并设计了一个任务引导模块,根据高层语言指令动态调整动作。这使得HetSkills能够保持自然运动,同时不断扩展其技能库,使其在长时间任务中具有高度适应性。实验结果表明,在运动跟踪、文本到运动生成、运动补全和下游任务适应方面的有效性,即使在挑战性条件下也取得了令人印象深刻的成功率。
cs.CV / 162 / 2608.23268

Dual-Grained Agent Memory and Shapley Context Attribution for Multimodal Agentic Learner

双粒度智能体记忆与沙普利上下文归因用于多模态智能学习者
Wang, Jieke, Shen, Tiancheng, Yang, Yibo, Yang, Ming-Hsuan
Abstract
Frontier multimodal large language models (MLLMs) deliver impressive perception yet still falter on scientific and mathematical reasoning. Parameter-level adaptation is unavailable for closed-weight or on-device backbones, and stateless prompting forfeits any compounding benefit from problems already solved. We propose \textbf{DG-Mem}, a dual-grained agentic memory framework that augments a frozen MLLM with a non-parametric, externally stored memory built once from training-time rollouts and consulted read-only at test time. Motivated by the Complementary Learning Systems (CLS) account of human memory, DG-Mem factors its store into an instance-grounded exemplar memory and a category-level schema memory of IF-THEN rules, with a transient reflection store mediating their construction so that schemas are synthesized only from abstract reflections, never from exemplar text. Two design choices distinguish DG-Mem: an online concept categorizer that grows the category space incrementally during training rather than committing to a predefined taxonomy, and a Shapley context attribution procedure that decomposes correctness across the entire retrieved rule set and yields a per-rule utility that re-weights retrieval at test time. The pipeline introduces no gradient updates and is deployable on closed-weight or on-device backbones. Across MathVista, MMMU, and MMMU-Pro on four open-weight and proprietary backbones (Qwen3.5-27B, Qwen3.5-122B-A10B, GPT-5-Nano, Gemini-3-Flash), DG-Mem improves consistently over no-memory and competitive memory baselines.
Chinese Translation
前沿的多模态大型语言模型(MLLMs)展现出令人印象深刻的感知能力,但在科学和数学推理方面仍然存在不足。对于封闭权重或设备端骨干网络,参数级适应不可用,而无状态提示则放弃了从已解决问题中获得的任何累积收益。我们提出了 extbf{DG-Mem},一种双粒度智能体记忆框架,它通过一个非参数的、外部存储的记忆来增强一个冻结的MLLM,该记忆在训练时从回滚中构建,并在测试时以只读方式进行查询。受到人类记忆的互补学习系统(Complementary Learning Systems, CLS)理论的启发,DG-Mem将其存储分为基于实例的示例记忆和类别级的IF-THEN规则模式记忆,且一个瞬态反思存储介导它们的构建,以便模式仅从抽象反思中合成,而不是从示例文本中提取。DG-Mem的两个设计选择使其与众不同:一个在线概念分类器,在训练过程中逐步扩展类别空间,而不是承诺于预定义的分类法;以及一个沙普利上下文归因程序,它对整个检索规则集的正确性进行分解,并产生每条规则的效用,从而在测试时重新加权检索。该流程不引入梯度更新,并可在封闭权重或设备端骨干网络上部署。在MathVista、MMMU和MMMU-Pro的四个开放权重和专有骨干网络(Qwen3.5-27B、Qwen3.5-122B-A10B、GPT-5-Nano、Gemini-3-Flash)上,DG-Mem在没有记忆和竞争记忆基线的情况下持续改善。
cs.CV / 163 / 2608.23279

Spatiotemporally Decoupled Autoregressive Diffusion Model for Human Motion Generation

时空解耦自回归扩散模型用于人类运动生成
Yang, Chengqun, Xu, Liang, Li, Yanping, Liu, Fulong, Gao, Jingnan, Zeng, Weili, Yan, Yichao
Abstract
Text-driven human motion synthesis has made substantial development with two core modules of motion representation and generative architecture. For representation, Vector Quantization (VQ)-based methods compress motion data into discrete tokens while latent-based models operate directly in continuous space. However, both of these representations exhibit significant limitations. VQ-based methods suffer from inherent information loss, which compromises the quality, diversity, and generalization of generated motions, while continuous representation on holistic whole-body motion hinders part-level flexibility. For architecture, diffusion and autoregressive diffusion models have demonstrated their superiority, yet the fine-grained controllability over individual body parts is also limited. Thus, we propose a unified spatiotemporally decoupled framework named DeMoDiff, which jointly redesigns representation and architecture. To enhance representation extraction capabilities and offer greater part-level controllability, we present a spatial-temporal VAE that encodes each body joint rather than compressing the whole-body motion into a single latent space. Then, we incorporate spatial-temporal masking and attention mechanisms into an autoregressive diffusion generator, achieving both generative capability and controllable editability. Extensive experiments on the HumanML3D and KIT-ML datasets demonstrate that our model achieves state-of-the-art reconstruction performance and compelling motion generation results. Moreover, our framework demonstrates strong temporal and spatial editing capabilities, further validating its effectiveness. Our project page: https://rex0191.github.io/DeMoDiff/
Chinese Translation
基于文本的人类运动合成在运动表示和生成架构两个核心模块上取得了显著进展。在表示方面,基于向量量化(Vector Quantization, VQ)的方法将运动数据压缩为离散的标记,而基于潜变量的模型则直接在连续空间中操作。然而,这两种表示方式都存在显著的局限性。基于VQ的方法由于固有的信息损失,影响了生成运动的质量、多样性和泛化能力,而对整体全身运动的连续表示则限制了局部的灵活性。在架构方面,扩散模型和自回归扩散模型已显示出其优越性,但对个体身体部位的精细控制能力也有限。因此,我们提出了一种统一的时空解耦框架,命名为DeMoDiff,该框架共同重新设计了表示和架构。为了增强表示提取能力并提供更大的部件级控制能力,我们提出了一种时空变分自编码器(VAE),该编码器对每个身体关节进行编码,而不是将整个身体运动压缩到单一的潜在空间中。接着,我们将时空掩蔽和注意力机制融入自回归扩散生成器,实现了生成能力和可控编辑性的结合。在HumanML3D和KIT-ML数据集上的大量实验表明,我们的模型在重建性能和运动生成结果上达到了最先进的水平。此外,我们的框架展现了强大的时间和空间编辑能力,进一步验证了其有效性。我们的项目页面:https://rex0191.github.io/DeMoDiff/
cs.CV / 164 / 2608.23290

Spotter: Efficient Urban Visual Localization via Geo-Referenced Facade Landmarks in GPS-Degraded Environments

Spotter:在GPS信号衰减环境中通过地理参考建筑立面标志实现高效城市视觉定位
Valls, Antoni, Sanchez-Riera, Jordi
Abstract
Accurate visual localization on robotic and wearable platforms remains challenging in dense urban environments. Existing methodologies typically rely on GPS for absolute positioning, yet GPS signals frequently degrade in urban canyons due to multipath propagation. Consequently, standard solutions like visual odometry suffer from unmitigated drift over time, while map-matching techniques struggle to acquire the reliable GPS priors they need, on top of being too computationally heavy for real-time edge execution. To address these limitations, we propose Spotter, a robuts and real-time visual localization framework that uses building facades as a reliable source of global geo-reference, while retaining the capability to integrate GPS signals when available. In an offline stage, Spotter processes Google Street View panoramas by semantically segmenting facades and pairing multi-view stereo depth with cartographic data to build a compact metric database. At runtime, query images are matched via a cascaded retrieval and geometric verification pipeline to recover fine-grained global camera localization. We benchmark Spotter on a newly collected dataset of pedestrian sequences acquired with wearable smart glasses across several districts of Barcelona. Experimental results show that Spotter outperforms odometry-based baselines and achieves localization accuracy comparable to state-of-the-art map-based methods while operating at significantly higher frame rates.
Chinese Translation
在密集城市环境中,机器人和可穿戴平台的准确视觉定位仍然面临挑战。现有的方法通常依赖GPS进行绝对定位,但由于多径传播,GPS信号在城市峡谷中经常衰减。因此,标准解决方案如视觉里程计随着时间的推移会遭遇不可避免的漂移,而地图匹配技术则难以获取所需的可靠GPS先验信息,并且在实时边缘执行时计算负担过重。为了解决这些局限性,我们提出了Spotter,一个稳健且实时的视觉定位框架,利用建筑立面作为可靠的全球地理参考源,同时在可用时保留集成GPS信号的能力。在离线阶段,Spotter通过对Google街景全景进行语义分割和将多视角立体深度与制图数据配对,构建一个紧凑的度量数据库。在运行时,查询图像通过级联检索和几何验证管道进行匹配,以恢复精细的全球相机定位。我们在一个新收集的行人序列数据集上对Spotter进行了基准测试,该数据集是通过可穿戴智能眼镜在巴塞罗那的多个区域获取的。实验结果表明,Spotter的性能优于基于里程计的基线,并且在显著更高的帧率下实现了与最先进的基于地图的方法相当的定位精度。
cs.CV / 165 / 2608.23295

What Memory Composition Does Not Tell Us About Anomaly Detection

记忆组合未能揭示的异常检测
Chae, Joongwon, Wang, Runming, Qin, Peiwu
Abstract
Memory-based anomaly detectors store nominal training patches and score test patches against this memory. A patch selected for coverage therefore becomes a nor- mal reference without a separate check that geometric rarity makes it safe to trust. We probe this coupling with sparse training contamination. Under fixed representa- tions and memory budgets, we compare random, medoid, local, and global coverage selectors. We then use CLEANCON, an out-of-bag cross-image support gate that changes candidate-image eligibility while fixing the representation, absolute mem- ory size, builder, and inference rule. Global coverage strongly over-represents sparse contamination. CLEANCON reduces final-memory contamination to approx- imately zero and increases category-macro P-AP in all 12 matched comparisons. Yet along a retention sweep, the lowest-contamination memory does not attain the highest P-AP; performance continues to improve while contamination rises. Mem- ory contamination therefore does not order the resulting memories by P-AP
Chinese Translation
基于记忆的异常检测器存储正常的训练样本,并根据这些记忆对测试样本进行评分。因此,选定的覆盖样本成为正常参考,而没有单独检查几何稀有性是否使其值得信赖。我们通过稀疏训练污染来探讨这种耦合。在固定的表示和记忆预算下,我们比较了随机、质心、局部和全局覆盖选择器。然后,我们使用 CLEANCON,这是一种袋外跨图像支持门,它在固定表示、绝对记忆大小、构建器和推理规则的同时改变候选图像的资格。全局覆盖在很大程度上过度代表了稀疏污染。CLEANCON 将最终记忆污染减少到接近零,并在所有12个匹配比较中提高了类别宏 P-AP。然而,在保留扫描中,最低污染的记忆并未达到最高的 P-AP;性能在污染上升的同时继续改善。因此,记忆污染并未按 P-AP 对结果记忆进行排序。
cs.CV / 166 / 2608.23299

What Remains Normal? Clean Images Miss Useful Near-Defect Normal Patches for Anomaly Detection

什么仍然是正常的?干净图像错过了用于异常检测的有用近缺陷正常补丁
Chae, Joongwon, Wang, Runming, Qin, Peiwu
Abstract
Memory-based anomaly detectors store nominal training patches and score test patches against this memory. A patch selected for coverage therefore becomes a nor- mal reference without a separate check that geometric rarity makes it safe to trust. We probe this coupling with sparse training contamination. Under fixed representa- tions and memory budgets, we compare random, medoid, local, and global coverage selectors. We then use CLEANCON, an out-of-bag cross-image support gate that changes candidate-image eligibility while fixing the representation, absolute mem- ory size, builder, and inference rule. Global coverage strongly over-represents sparse contamination. CLEANCON reduces final-memory contamination to approx- imately zero and increases category-macro P-AP in all 12 matched comparisons. Yet along a retention sweep, the lowest-contamination memory does not attain the highest P-AP; performance continues to improve while contamination rises. Mem- ory contamination therefore does not order the resulting memories by P-AP.Code is publicly available at https://github.com/jw-chae/cleancon.
Chinese Translation
基于记忆的异常检测器存储正常训练补丁,并根据该记忆对测试补丁进行评分。因此,选择用于覆盖的补丁成为正常参考,而没有单独检查几何稀有性是否使其值得信赖。我们通过稀疏训练污染来探讨这种耦合。在固定的表示和记忆预算下,我们比较了随机、质心、局部和全局覆盖选择器。然后,我们使用CLEANCON,一种超出袋交叉图像支持门,它在固定表示、绝对记忆大小、构建器和推理规则的同时改变候选图像的资格。全局覆盖强烈过度表示稀疏污染。CLEANCON将最终记忆污染减少到接近零,并在所有12个匹配比较中提高了类别宏P-AP。然而,在保留扫描中,最低污染的记忆并未达到最高的P-AP;性能在污染上升的同时继续改善。因此,记忆污染并未按P-AP对结果记忆进行排序。代码可在https://github.com/jw-chae/cleancon公开获取。
cs.CV / 167 / 2608.23302

Grounding Free-Form Instructions for Fashion Complementary Image Generation

基于自由形式指令的时尚互补图像生成
Attimonelli, Matteo, Pomo, Claudio, De Bellis, Alessandro, Danese, Danilo, Jannach, Dietmar, Di Noia, Tommaso
Abstract
Fashion complementary image generation (CIG) aims to create garments that stylistically match a seed item based on user intent, making it a natural multimodal grounding problem where models must interpret language in visual context. Existing CIG benchmarks rely on rigid template prompts (e.g., "a photo of a skirt"), failing to reflect natural user queries and obscuring model behavior across levels of linguistic specificity. We introduce fashion complementary image generation with free-form instructions, a multimodal language-grounding setting where a model generates a compatible garment from a seed image and a natural-language instruction. To this end, we enrich three CIG benchmarks with low-, medium-, and high-specificity instructions generated by a vision-language model and validated by human annotators. We instantiate the task with StyleFlow, a Rectified Flow Matching model that jointly conditions on the seed image and instruction within a single multimodal transformer. Across image quality metrics, catalog-alignment analysis, ablations, and human evaluation, StyleFlow consistently produces instruction-aligned and stylistically coherent garments while reducing architectural complexity and inference cost relative to auxiliary-module approaches.
Chinese Translation
时尚互补图像生成(CIG)旨在根据用户意图创建与种子物品在风格上相匹配的服装,这使其成为一个自然的多模态基础问题,模型必须在视觉上下文中解读语言。现有的CIG基准依赖于固定模板提示(例如,“一条裙子的照片”),未能反映自然用户查询,并且模糊了模型在语言特异性各个层次上的行为。我们引入了基于自由形式指令的时尚互补图像生成,这是一个多模态语言基础设置,其中模型从种子图像和自然语言指令生成兼容的服装。为此,我们通过视觉-语言模型生成并由人工注释者验证的低、中、高特异性指令,丰富了三个CIG基准。我们使用StyleFlow,一个整流流匹配模型,在单个多模态变换器中共同条件化种子图像和指令。通过图像质量指标、目录对齐分析、消融实验和人工评估,StyleFlow始终生成与指令对齐且在风格上连贯的服装,同时相较于辅助模块方法降低了架构复杂性和推理成本。
cs.CV / 168 / 2608.23329

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

超越视频的思考:统一视频推理与开放世界视频智能体的深度研究
Liu, Wenqi, Ma, Shijie, Wang, Yunxiao, Liu, Meng, Su, Qile, Liu, Han, Hou, Bohan, Zheng, Xuanyu, Liu, Changyi, Zhang, Tianke, Fan, Haonan, Jiang, Kaiyu, Li, Yingxin, Chen, Jiankang, Wang, Xu, Wen, Bin, Gao, Tingting, Li, Han, Yin, Jianhua, Wei, Yinwei, Song, Xuemeng
Abstract
Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perception and Deep Research supports multi-step information seeking, the two capabilities are typically developed in isolation. We introduce VideoRover, a unified Video Deep Research framework that iteratively coordinates video cropping, multimodal search, and webpage browsing. Given a video-question pair, VideoRover uses each tool result to select the next action, so localized video clips guide external retrieval and retrieved evidence triggers further video inspection and verification. To develop this capability, we construct an automated data curation pipeline, producing 26K verified SFT trajectories and 3K challenging RL instances. We also introduce VideoRover-Bench, a benchmark stratified by video duration and research difficulty. Experiments on VideoDR and VideoRover-Bench show that our VideoRover-8B-RL achieves performance comparable to proprietary models in the direct-answer setting without tool use while outperforming larger open-source models equipped with the same tool suite. Ablation studies and training dynamics further validate the complementary roles of active video grounding, external retrieval, and long-horizon reinforcement learning.
Chinese Translation
开放世界视频理解通常要求模型定位稀疏的视觉证据,并获取视频及其参数记忆中缺失的外部知识。尽管“视频思维”(Thinking-with-Videos)能够实现主动的时间感知,而“深度研究”(Deep Research)支持多步骤的信息获取,但这两种能力通常是孤立发展的。我们提出了VideoRover,一个统一的视频深度研究框架,能够迭代协调视频裁剪、多模态搜索和网页浏览。给定一个视频-问题对,VideoRover利用每个工具的结果选择下一步行动,从而使局部视频片段引导外部检索,而检索到的证据又触发进一步的视频检查和验证。为了开发这一能力,我们构建了一个自动化数据策划管道,生成了26K个经过验证的SFT轨迹和3K个具有挑战性的RL实例。我们还推出了VideoRover-Bench,一个按视频时长和研究难度分层的基准测试。对VideoDR和VideoRover-Bench的实验表明,我们的VideoRover-8B-RL在不使用工具的直接回答设置中达到了与专有模型相当的性能,同时在配备相同工具套件的情况下超越了更大的开源模型。消融研究和训练动态进一步验证了主动视频定位、外部检索和长时间强化学习的互补作用。
cs.CV / 169 / 2608.23330

IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning

IntentQA:通过认知上下文推理实现视频中的意图问答
Li, Jiapeng, Wei, Ping, Han, Wenjuan, Zhu, Song-Chun, Fan, Lifeng
Abstract
Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions (often termed the "dark matter" of social intelligence). To bridge the gap between visual observation and intent reasoning, we introduce a novel task, IntentQA, and contribute a large-scale VideoQA dataset specifically tailored for this purpose. However, recognizing that standard metrics may overestimate capabilities due to dataset biases, we go beyond simple accuracy to rigorously evaluate model robustness. We augment the benchmark by generating five distinct contrast sets via Large Language Models (LLMs) and introducing a "Contrast Performance Decline" metric. We propose the X-CaVIR (eXplainable Context-aware Video Intent Reasoning) framework, which leverages three types of "Cognitive Context" to enhance video analysis: i) Situational Context via a cross-modal Video Query Language (VQL) module, ii) Contrastive Context via a Contrastive Learning module, and iii) Commonsense Context via a Commonsense Reasoning module. Crucially, to overcome the opacity of traditional black-box models, we refine the integration of LLMs within X-CaVIR by employing a transparent pipeline that synergizes video captions with VQA model outputs. This approach not only improves performance by effectively utilizing rich commonsense knowledge but also renders the reasoning process explicitly interpretable. Extensive experiments demonstrate the effectiveness of our components, the superiority of X-CaVIR over state-of-the-art baselines, and its stability against perturbations on the contrast sets.
Chinese Translation
视频理解要求智能体超越对视觉事实的单纯识别,理解人类行为背后的潜在意图(通常被称为社会智能的“暗物质”)。为了弥合视觉观察与意图推理之间的差距,我们提出了一项新任务——IntentQA,并贡献了一个专门为此目的量身定制的大规模VideoQA数据集。然而,我们认识到标准指标可能由于数据集偏差而高估能力,因此我们超越简单的准确性,严格评估模型的鲁棒性。我们通过大型语言模型(LLMs)生成五个不同的对比集来增强基准测试,并引入了“对比性能下降”指标。我们提出了X-CaVIR(可解释的上下文感知视频意图推理)框架,该框架利用三种类型的“认知上下文”来增强视频分析:i)通过跨模态视频查询语言(VQL)模块实现的情境上下文,ii)通过对比学习模块实现的对比上下文,以及iii)通过常识推理模块实现的常识上下文。至关重要的是,为了克服传统黑箱模型的不可解释性,我们通过采用一个透明的管道,将视频字幕与VQA模型输出相结合,优化了LLMs在X-CaVIR中的整合。这种方法不仅通过有效利用丰富的常识知识提高了性能,还使推理过程变得明确可解释。大量实验表明我们各组件的有效性,X-CaVIR在性能上优于最先进的基线,并且在对比集的扰动下表现出稳定性。
cs.CV / 170 / 2608.23336

Can Coding Agents Build Robust Baselines? A Skill-Based Approach for Automating the Medical Imaging Model-Development Pipeline

编码代理能否构建稳健的基线?一种基于技能的医学影像模型开发流程自动化方法
Moris, Eugenia, Orlando, José Ignacio
Abstract
Developing competitive deep learning baselines for medical imaging remains a highly iterative process requiring literature review, implementation, experimentation, and expert refinement. Existing automation approaches typically optimize isolated components, such as architecture search or hyperparameter tuning, rather than the complete baseline development process. We present an agentic AI Scientist workflow that combines literature-guided reasoning, automated code generation, and hypothesis-driven experimentation to generate competitive baseline models for medical imaging challenges. The framework is evaluated on four public benchmarks spanning segmentation, classification, and detection. Across all tasks, the Experimentation Pipeline consistently improves validation performance, achieving competitive leaderboard results, including 6th place on both PUMA tracks (15 teams) and 31st place on MILK10k (125 teams). On MIDOG25, the resulting model also demonstrates strong domain generalization across scanners, tumor types, and species. Using the same workflow across all challenges without task-specific redesign, we demonstrate that skill-based, literature-guided agentic workflows can substantially reduce the engineering effort required to develop competitive medical imaging baselines.
Chinese Translation
开发具有竞争力的医学影像深度学习基线仍然是一个高度迭代的过程,需要文献回顾、实现、实验和专家优化。现有的自动化方法通常优化孤立的组件,如架构搜索或超参数调优,而不是完整的基线开发过程。我们提出了一种代理人工智能科学家工作流程,结合了文献指导的推理、自动代码生成和假设驱动的实验,以生成针对医学影像挑战的竞争性基线模型。该框架在四个公共基准上进行了评估,涵盖了分割、分类和检测。在所有任务中,实验管道始终提高了验证性能,取得了竞争性的排行榜结果,包括在PUMA两个赛道上获得第六名(15支队伍)和在MILK10k上获得第31名(125支队伍)。在MIDOG25上,所得到的模型还展示了在扫描仪、肿瘤类型和物种之间的强域泛化。通过在所有挑战中使用相同的工作流程而无需特定任务的重新设计,我们证明了基于技能的、文献指导的代理工作流程可以显著减少开发竞争性医学影像基线所需的工程努力。
cs.CV / 171 / 2608.23343

Controllable blind deblurring with diffusion models

基于扩散模型的可控盲去模糊
Salah, Imane Si, Cribelier, Emile, Veit, Thomas, Hauser, Wolf, Leclaire, Arthur
Abstract
Image acquisition with a camera involves several degradations due to the optical system, sensor, or low-level processing steps. We address blind deblurring in professional photography: we aim to invert unknown isotropic blur without knowledge of the degradation kernel.For such inverse problems,where some high-frequency information is lost, it is challenging to use generative models to produce details that are both photo-realistic and faithful to the input. We propose SuperSharpen, a diffusion-based blind deblurring method offering explicit control over restoration strength through a blur measure. We compare two conditioning strategies: a ControlNet-style adapter on a frozen backbone, and full finetuning of the diffusion prior. Our experiments show that finetuning achieves better fidelity with fewer hallucinated details. We validate our approach on synthetic and real-world blur, demonstrating improved perceptual quality and controllable restoration strength.
Chinese Translation
使用相机进行图像采集时,由于光学系统、传感器或低级处理步骤,图像会遭受多种退化。我们关注专业摄影中的盲去模糊问题:我们的目标是逆转未知的各向同性模糊,而无需了解退化核。对于这种逆问题,由于某些高频信息的丢失,使用生成模型来生成既真实又忠实于输入的细节是具有挑战性的。我们提出了SuperSharpen,这是一种基于扩散的盲去模糊方法,通过模糊度量提供对恢复强度的明确控制。我们比较了两种条件策略:在冻结主干网络上的ControlNet风格适配器,以及对扩散先验的完全微调。我们的实验表明,微调在保持更高保真度的同时,减少了幻觉细节的出现。我们在合成和真实世界的模糊图像上验证了我们的方法,展示了改善的感知质量和可控的恢复强度。
cs.CV / 172 / 2608.23363

DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

DF-MoE:通过多模态稀疏专家混合实现可泛化的深伪检测
Hondru, Vlad, Croitoru, Florinel Alin, Georgescu, Iuliana, Koepke, A. Sophia, Ionescu, Radu Tudor
Abstract
Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize across deepfake generation methods. We conjecture that overfitting can be mitigated by extracting multiple high-level cues from the available audio and visual modalities via pre-trained models. We therefore assemble a wide variety of pre-trained models to extract features that encode mouth movements, face parsing, facial expressions, head pose, gaze tracking, heart rate, audio emotion and speech activity. We further integrate both unimodal and multimodal cues via a Mixture-of-Experts (MoE) backbone to detect deepfakes. We perform in-domain and cross-domain experiments on five benchmarks for deepfake detection (MAVOS-DD, AVLips, PolyGlotFake, BioDeepAV, FakeAVCeleb) to compare our framework (DF-MoE) with state-of-the-art methods. Our results indicate that DF-MoE obtains superior deepfake detection results, surpassing all competing methods. We release our code at https://github.com/vladhondru25/DF-MoE.
Chinese Translation
音视频深伪检测是一个积极研究的主题,其中一个主要挑战是开发能够在不同深伪生成方法之间泛化的检测器。我们推测,通过利用预训练模型从可用的音频和视觉模态中提取多个高层次线索,可以减轻过拟合。因此,我们组装了多种预训练模型,以提取编码嘴部运动、面部解析、面部表情、头部姿态、注视追踪、心率、音频情感和语音活动的特征。我们进一步通过混合专家(Mixture-of-Experts, MoE)骨干网络整合单模态和多模态线索,以检测深伪。我们在五个深伪检测基准(MAVOS-DD、AVLips、PolyGlotFake、BioDeepAV、FakeAVCeleb)上进行领域内和跨领域实验,将我们的框架(DF-MoE)与最先进的方法进行比较。我们的结果表明,DF-MoE在深伪检测方面取得了优越的结果,超越了所有竞争方法。我们在 https://github.com/vladhondru25/DF-MoE 发布了我们的代码。
cs.CV / 173 / 2608.23383

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

长时间音视频生成用于持久故事和互动世界
Duan, Nan, Huang, Haoyang, Jin, Weiyang, Li, Haoran, Li, Yaowei, Li, Yuming, Liu, Yijun, Lu, Xin, Ma, Xiaoxiao, Ma, Yanwen, Su, Yaofeng, Sun, Yilang, Wang, Haoyu, Xue, Zeyue, Zhang, Songchun, Zhuang, Junhao
Abstract
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.
Chinese Translation
视频生成正在从孤立的剪辑向长篇叙事和互动世界发展,这要求模型能够保持角色身份、遵循用户控制,并在长时间的生成过程中保持稳定。我们提出了JoyAI-Echo-1.5,一个统一的音视频生成系统,具有两个专门构建的变体。长视频变体引入了可组合的跨镜头记忆,能够聚合来自多个先前镜头的视觉证据和从语音过滤的全镜头音频中提取的说话者线索,从而在文本、图像和记忆条件的灵活组合中实现持久的角色外观和声音身份。世界模型变体将异构导航输入转换为校准的度量6自由度(6-DoF)相机轨迹,并通过几何感知的条件路径注入这些轨迹,从而实现跨灵活视角的控制器无关交互。为了支持高效的长时间生成,我们将双向音视频骨干网络转变为一个因果的少步生成器,采用渐进式教师强迫和自生成回滚的短期和长期自梯度强迫。实验结果在这两种设置中均表现出强劲的性能。JoyAI-Echo-1.5在跨镜头一致性、视觉质量、文本对齐和语音保真度方面相较于现有的长视频基准取得了改进。其世界模型变体在WBench上排名第一,平均得分为81.7,并在SANA-WM-Bench上实现了领先的视觉质量和长时间持久性。这些结果表明,记忆、几何控制和回滚感知训练为生成连贯的故事和持续演变的互动世界提供了实用的基础。项目页面:https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/
cs.CV / 174 / 2608.23405

MomADv2: Reliable Temporal Memory for End-to-End Autonomous Driving

MomADv2:用于端到端自主驾驶的可靠时间记忆
Song, Ziying, Zhang, Shengkai, Liu, Lin, Wu, Peiliang, Yang, Lei, Xu, Dongyang, Sun, Bin, Wang, Li, Xu, Shaoqing, Jia, Caiyan, Luo, Yadan
Abstract
Long-horizon planning is critical for safe autonomous driving in complex scenarios. Existing methods improve planning continuity with temporal memory, but such memory may become invalid and mislead decisions when the driving command changes. Thus, selectively leveraging useful history while suppressing command-inconsistent memory remains a key challenge. To address this issue, we propose MomADv2, a reliable state-space memory framework for long-horizon end-to-end autonomous driving. At its core, MomADv2 introduces a Selective State-Space Planning Memory Query Module, which filters historical planning queries based on temporal continuity and command consistency, selects planning modes relevant to the current command, and models the evolution of planning intentions through a selective state-space mechanism. To further alleviate local trajectory deviations and error accumulation in long-horizon planning, we design a Flow-Matching Trajectory Residual Refiner. It learns a continuous residual correction field from the refined planning output to the expert trajectory, enabling fine-grained trajectory refinement while preserving the stability of anchor-based planning. Extensive experiments on closed-loop NAVSIM and Bench2Drive, as well as open-loop nuScenes, demonstrate that MomADv2 improves long-horizon planning consistency and reduces the average collision rate by 15.6% over MomAD under 6-second planning.
Chinese Translation
长时间规划对于在复杂场景中安全的自主驾驶至关重要。现有方法通过时间记忆提高规划的连续性,但当驾驶指令发生变化时,这种记忆可能失效并误导决策。因此,如何选择性地利用有用的历史信息,同时抑制与指令不一致的记忆,仍然是一个关键挑战。为了解决这个问题,我们提出了MomADv2,一个用于长时间端到端自主驾驶的可靠状态空间记忆框架。MomADv2的核心是引入了选择性状态空间规划记忆查询模块,该模块根据时间连续性和指令一致性过滤历史规划查询,选择与当前指令相关的规划模式,并通过选择性状态空间机制建模规划意图的演变。为了进一步减轻长时间规划中的局部轨迹偏差和误差累积,我们设计了流匹配轨迹残差精炼器。它从精炼的规划输出到专家轨迹学习一个连续的残差修正场,使得在保持基于锚点的规划稳定性的同时,实现精细的轨迹精炼。在闭环NAVSIM和Bench2Drive以及开环nuScenes上的大量实验表明,MomADv2在6秒规划下提高了长时间规划的一致性,并将平均碰撞率降低了15.6%。
cs.CV / 175 / 2608.23410

Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers

基于下一尺度变换器的人脸照片级真实感新视图合成
Stella, Federico, Jiang, Fei, Jiang, Zhongshi, Barzelay, Zohar, Garbin, Emanuel, Jourabloo, Amin, Ge, Liuhao
Abstract
Photorealistic novel view synthesis of people remains challenging at high spatial resolutions and across multiple target cameras, where preserving identity, fine appearance details, and geometric coherence is critical. We build on the next-scale autoregressive paradigm and adapt it for human-centric view synthesis by enabling higher image resolutions, multi-view outputs and stronger cross-view consistency in a single forward pass. We train on a synthetic dataset of human faces spanning diverse identities and apparel. Contrary to diffusion models, this paradigm does not need 2D pre-training and, thanks to its next-scale architecture, it benefits from lower-resolution, general-purpose pre-trainings, with the full-sized purpose-specific images being used only in the last training stages. This enables our architecture to converge with a smaller amount of purpose-specific training data, allowing us to use a smaller but more realistic training dataset. The resulting model produces sharp and realistic views, with the option to synthesize multiple novel viewpoints simultaneously for improved agreement across views. Empirically, we observe gains in perceptual fidelity and cross-view coherence on human subjects, demonstrating that next-scale autoregression is an effective backbone for scalable, multi-output human view synthesis. We also couple our pipeline with an existing transformer-based model for pixel-aligned 3D gaussian lifting from multi-view facial inputs, resulting in accurate and photorealistic 3D models of human faces.
Chinese Translation
在高空间分辨率和多个目标相机下,人脸的照片级真实感新视图合成仍然具有挑战性,其中保持身份、细致外观细节和几何一致性至关重要。我们基于下一尺度自回归范式,并将其适应于以人为中心的视图合成,通过在单次前向传播中实现更高的图像分辨率、多视角输出和更强的跨视角一致性。我们在一个涵盖多样身份和服装的合成人脸数据集上进行训练。与扩散模型不同,这一范式不需要二维预训练,并且得益于其下一尺度架构,能够利用低分辨率的通用预训练,只有在最后的训练阶段才使用全尺寸的特定目的图像。这使得我们的架构能够以更少的特定目的训练数据收敛,从而使我们能够使用更小但更真实的训练数据集。最终模型生成清晰且真实的视图,并可以同时合成多个新视角,以提高视图之间的一致性。从经验上看,我们观察到在人类对象上的感知保真度和跨视角一致性都有所提升,证明了下一尺度自回归是可扩展的多输出人类视图合成的有效基础。我们还将我们的管道与现有的基于变换器的模型结合,用于从多视角面部输入进行像素对齐的三维高斯提升,最终生成准确且照片级真实感的人脸三维模型。
cs.CV / 176 / 2608.23432

Image-Conditioned Diffusion Models for Quality Assurance of Organ-at-Risk Segmentations in Radiotherapy

基于图像条件的扩散模型在放射治疗中风险器官分割质量保证中的应用
Dronne, Clea, Clark, Catharine H, Loizeau, Xavier, Miles, Elizabeth, Hoskin, Peter, McClelland, Jamie R
Abstract
Accurate organ-at-risk segmentation is essential for radiotherapy planning, but reviewing segmentations is time-consuming and subjective. We investigate normative modelling for segmentation error detection in head-and-neck CT, comparing a VAE framework with an image-conditioned segmentation diffusion model. Models were evaluated on RADCURE brainstem and spinal cord segmentations using simulated boundary and width perturbations. Error detection was assessed using the Dice similarity coefficient and the Distance to Agreement (DTA) between the input and reconstructed segmentations. While both models detected some simulated errors, regional DTA showed that the diffusion model localised subtle boundary errors more consistently. These results support image-conditioned diffusion reconstruction as a promising framework for localised, anatomy-aware segmentation QA.
Chinese Translation
准确的风险器官分割对于放射治疗规划至关重要,但审查分割结果既耗时又主观。我们研究了在头颈部CT中进行分割错误检测的规范建模,比较了变分自编码器(VAE)框架与图像条件的分割扩散模型。模型在RADCURE脑干和脊髓分割上进行了评估,使用模拟的边界和宽度扰动。通过Dice相似系数和输入与重建分割之间的一致性距离(DTA)评估错误检测。尽管两个模型都检测到了一些模拟错误,但区域DTA显示扩散模型更一致地定位了细微的边界错误。这些结果支持图像条件的扩散重建作为一种有前景的框架,用于局部、解剖意识的分割质量保证。
cs.CV / 177 / 2608.23435

Towards Comprehensive Basketball Understanding

迈向全面的篮球理解
Hu, Yirong, Rao, Jiayuan, Zhang, Yu, Di, Shangzhe, Xie, Weidi
Abstract
Understanding a basketball game requires recognizing events, localizing actions, identifying players, and relating these to structured game knowledge. Existing benchmarks primarily evaluate these abilities one at a time, leaving the interactions among these abilities under-explored. We introduce BasketballBench, a multimodal benchmark comprising 7,980 questions across ten tasks in text, image, and video. It is built from the 2025-2026 NBA season and includes official playby-play, rosters and profiles for 530 active players, and 2,501 possession-level broadcast clips. We further propose BasketballSkills, an agent that composes eight basketball-specific perception and retrieval tools under four reusable skills that specify tool order, evidence bindings, and stopping conditions. Experiments show that current MLLMs struggle particularly on questions requiring the integration of multiple capabilities, whereas BasketballSkills outperforms them, highlighting the effectiveness of explicitly composing domain-specific capabilities for comprehensive basketball understanding.
Chinese Translation
理解一场篮球比赛需要识别事件、定位动作、识别球员,并将这些与结构化的比赛知识关联起来。现有的基准测试主要一次评估这些能力中的一种,导致这些能力之间的相互作用未得到充分探索。我们引入了BasketballBench,这是一个多模态基准,包含跨越十个任务的7,980个问题,涵盖文本、图像和视频。该基准基于2025-2026 NBA赛季构建,包括官方的逐场比赛记录、530名活跃球员的名单和个人资料,以及2,501个基于回合的广播剪辑。我们进一步提出了BasketballSkills,一个代理,组合了八种特定于篮球的感知和检索工具,基于四种可重用的技能,指定工具顺序、证据绑定和停止条件。实验表明,当前的多模态大语言模型(MLLMs)在需要整合多种能力的问题上表现尤为困难,而BasketballSkills的表现优于它们,突显了明确组合领域特定能力在全面篮球理解中的有效性。
cs.CV / 178 / 2608.23479

Geometry-Driven Opti-Acoustic Co-Registration and View-Invariant Reflectivity Mapping for Side-Scan Sonar

基于几何驱动的光声共配准与视角不变的侧扫声纳反射率映射
Hamoda, Taqi, Gracias, Nuno
Abstract
Side-Scan Sonar (SSS) is a primary modality for large-scale underwater mapping, yet automated perception and cross-modal alignment are severely bottlenecked by acoustic complexities such as speckle noise, shadows, and extreme viewpoint dependencies. Traditional handcrafted descriptors and modern deep learning matchers fail to bridge the physical domain gap between optical and acoustic imagery without 3D geometric constraints. To overcome these limitations, we propose a novel geometry-driven framework for pixel-level opti-acoustic co-registration and view-invariant reflectivity mapping. Our method utilizes Structure-from-Motion (SfM) to reconstruct a dense 3D seafloor mesh, acting as a geometric anchor between the visual and acoustic domains. We introduce a First Bottom Return (FBR) extraction algorithm to dynamically correct non-linear altitude drift caused by uncalibrated SfM reconstruction. Furthermore, we apply an inverse Lambertian model and a dual-Gaussian weighting function to isolate the intrinsic seabed reflectivity, effectively neutralizing slant-range propagation loss and geometric view-dependence. By deterministically associating these isolated acoustic properties with optical pixels, our pipeline generates highly accurate, strictly co-registered multi-modal datasets. This automated, physics-guided approach eliminates the need for manual annotation and paves the way for advanced self-supervised learning in benthic habitat mapping.
Chinese Translation
侧扫声纳(Side-Scan Sonar, SSS)是大规模水下制图的主要手段,但由于声学复杂性,如斑点噪声、阴影和极端视角依赖性,自动感知和跨模态对齐受到严重瓶颈。传统的手工描述符和现代深度学习匹配器在没有三维几何约束的情况下,无法弥合光学图像与声学图像之间的物理领域差距。为克服这些限制,我们提出了一种新颖的基于几何驱动的像素级光声共配准与视角不变反射率映射框架。我们的方法利用运动结构重建(Structure-from-Motion, SfM)技术重建密集的三维海底网格,作为视觉和声学领域之间的几何锚点。我们引入了一种首次底部回波(First Bottom Return, FBR)提取算法,以动态校正由未校准的SfM重建引起的非线性高度漂移。此外,我们应用逆兰伯特模型和双高斯加权函数来隔离内在的海底反射率,有效中和倾斜范围传播损失和几何视角依赖性。通过确定性地将这些隔离的声学特性与光学像素关联,我们的流程生成高度准确、严格共配准的多模态数据集。这种自动化的、以物理为导向的方法消除了手动标注的需要,为底栖栖息地映射中的高级自监督学习铺平了道路。
cs.CV / 179 / 2608.23486

GeoWAM: Visual Geometry World Action Models for Autonomous Driving

GeoWAM:用于自动驾驶的视觉几何世界动作模型
Lu, Yiren, Ye, Xin, Liu, Jiaming, Yao, Jin, Chen, Yi-chung, Merino, Liam, Kurra, Dhruva Dixith, Cai, Min, Lampo, Tom, Yin, Yu, Guo, Danhua, Yaman, Burhan
Abstract
World action models (WAMs) have recently gained increasing attention as a framework for jointly modeling scene evolution and ego actions in autonomous driving. Most existing WAMs learn scene dynamics in pixel space by combining a video-generation backbone for future-observation prediction with an action head for ego-trajectory prediction. Pixels, however, provide only an indirect representation of these dynamics: they entangle geometry and motion with appearance, texture, and illumination, forcing the model to infer three-dimensional transformations from two-dimensional observations. We argue that geometry, represented by point clouds, offers a more natural state space for driving because it explicitly captures spatial structure and the rigid and non-rigid transformations that govern scene evolution while directly aligning with the space in which driving actions are executed. Building on this insight, we introduce \textbf{GeoWAM}, a visual geometry world action model for autonomous driving. Rather than predicting future images, GeoWAM is pretrained to forecast future scene geometry, yielding representations that jointly encode spatial structure and temporal evolution. A geometry-conditioned action head then leverages these learned geometric dynamics to predict future ego trajectories. Extensive open-loop and closed-loop evaluations show that visual geometry world modeling yields substantially stronger driving policies than image-based alternatives, establishing future-geometry prediction as an effective pretraining objective for autonomous driving.
Chinese Translation
世界动作模型(WAMs)近年来作为一个框架,越来越受到关注,用于联合建模自动驾驶中的场景演变和自我动作。大多数现有的WAMs通过结合视频生成主干网络进行未来观察预测和动作头进行自我轨迹预测,在像素空间中学习场景动态。然而,像素仅提供了这些动态的间接表示:它们将几何和运动与外观、纹理和光照纠缠在一起,迫使模型从二维观察中推断三维变换。我们认为,点云所表示的几何为驾驶提供了一个更自然的状态空间,因为它明确捕捉了空间结构以及支配场景演变的刚性和非刚性变换,同时直接与执行驾驶动作的空间对齐。基于这一见解,我们引入了 extbf{GeoWAM},一种用于自动驾驶的视觉几何世界动作模型。GeoWAM并不是预测未来图像,而是预训练以预测未来场景几何,从而产生共同编码空间结构和时间演变的表示。然后,几何条件的动作头利用这些学习到的几何动态来预测未来的自我轨迹。广泛的开放环路和闭环评估表明,视觉几何世界建模产生的驾驶策略显著优于基于图像的替代方案,确立了未来几何预测作为自动驾驶有效的预训练目标。
cs.CV / 180 / 2608.23499

SVD-Based Typicality Maps for Out-of-Distribution Detection in Vision Transformers

基于奇异值分解的典型性图用于视觉变换器中的分布外检测
Sartor, Aldo Sean, Rosa, Leandro de Souza, Enttsel, Andriy, Mangia, Mauro, Rovatti, Riccardo
Abstract
We present a method for analyzing the internal representations of Vision Transformers (ViTs) exploiting the geometry of their learned parameters. Each affine layer's weight matrix is factored via Singular Value Decomposition (SVD), and activations are projected onto the leading right singular vectors to obtain compact, layer-intrinsic representations. A class-conditional density model is then fitted at each layer, producing per-class \emph{typicality scores} that are stacked across depth into \emph{typicality maps}: two-dimensional summaries of how class-specific evidence evolves through the network. From these maps, we derive two post-hoc scores for Out-Of-Distribution (OOD) detection: a \emph{Prototype Alignment Score} (PAS), measuring agreement with class reference prototype patterns, and a \emph{Multi-Layer Soft Voting} (MLSV) score, capturing cross-layer consensus without stored prototypes. On ViT-B/16 fine-tuned on CIFAR-100, the proposed scores achieve competitive detection performance without retraining or OOD exposure.
Chinese Translation
我们提出了一种分析视觉变换器(Vision Transformers, ViTs)内部表征的方法,该方法利用其学习参数的几何特性。通过奇异值分解(Singular Value Decomposition, SVD)对每个仿射层的权重矩阵进行分解,并将激活值投影到主要的右奇异向量上,以获得紧凑的层内表征。然后在每个层上拟合一个类条件密度模型,生成每类的 extit{典型性得分},并在深度上堆叠形成 extit{典型性图}:这些图是对类特定证据在网络中演变的二维总结。从这些图中,我们推导出两个后验得分用于分布外(Out-Of-Distribution, OOD)检测: extit{原型对齐得分}(Prototype Alignment Score, PAS),用于测量与类参考原型模式的一致性,以及 extit{多层软投票}(Multi-Layer Soft Voting, MLSV)得分,用于捕捉跨层共识而无需存储原型。在对CIFAR-100进行微调的ViT-B/16上,所提出的得分在不重新训练或接触OOD的情况下实现了竞争性的检测性能。
cs.CV / 181 / 2608.23503

Action-Aligned Retrieval with Pairwise Multimodal Reranking for Text-Based Person Anomaly Search

基于动作对齐的文本人异常搜索的成对多模态重排序
Nguyen, Thanh-Khoi, Vo, Thanh-Nhan, Nguyen, Trong-Thuan, Tran, Minh-Triet
Abstract
Text-based person anomaly search requires distinguishing individuals based on fine-grained, context-dependent behaviors rather than mere appearance. Existing methods struggle to capture these context-conditioned actions, frequently relying on isolated skeletal geometry, discarding raw query details during reformulation, or utilizing absolute pointwise scoring for multimodal verification. To address these limitations, we propose \textbf{ActPair}, a unified three-stage coarse-to-fine framework that combines action-aligned retrieval with pairwise multimodal reranking to bridge the pose-semantic gap. First, we fine-tune a vision-language model (VLM) with an action-aligned multi-task objective that encourages the representations to encode action-discriminative semantics. Second, we perform parallel late-fusion retrieval using the original query and a large language model (LLM)-generated context-grounded rewrite, retaining complementary details from both semantic views. Finally, we propose an efficient off-the-shelf reranking module that leverages a pivot-promote algorithm to perform direct pairwise visual comparisons, mitigating residual spatial and compositional ambiguities without the prohibitive inference costs of exhaustive evaluation. Extensive experiments demonstrate that our framework achieves the best results among the compared methods on the Pedestrian Anomaly Behavior (PAB) public test and transfers effectively to an unseen, non-anomaly-specific dataset.
Chinese Translation
基于文本的人异常搜索需要根据细粒度、依赖于上下文的行为来区分个体,而不仅仅是外观。现有方法难以捕捉这些依赖于上下文的动作,常常依赖于孤立的骨架几何,忽视了在重构过程中原始查询的细节,或使用绝对的逐点评分进行多模态验证。为了解决这些局限性,我们提出了 extbf{ActPair},一个统一的三阶段粗到细框架,结合了动作对齐检索与成对多模态重排序,以弥合姿态与语义之间的差距。首先,我们通过一个动作对齐的多任务目标对视觉-语言模型(VLM)进行微调,鼓励其表示编码动作区分的语义。其次,我们使用原始查询和大型语言模型(LLM)生成的上下文基础重写进行并行晚融合检索,保留来自两个语义视角的互补细节。最后,我们提出了一种高效的现成重排序模块,利用枢轴促进算法进行直接成对视觉比较,减轻残余的空间和组合模糊性,而无需耗费大量计算成本进行全面评估。大量实验表明,我们的框架在行人异常行为(PAB)公共测试中实现了与比较方法中最佳的结果,并有效地迁移到一个未见过的、非异常特定的数据集上。
cs.CV / 182 / 2608.23518

Investigating Relational Reasoning in VLMs

探究视觉语言模型中的关系推理
Geetha, Adhithya Laxman Ravi Shankar, Rakhmasari, Aulia Kharis, Ramzan, Haleema, Yap, Xander
Abstract
Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or simply employ shortcuts such as language cues or priors. To investigate this, we use the Qwen3-VL-4B (Bai et al., 2025), a modern VLM, to decode how visual information is encoded across depths. For this, we propose a synthetic dataset of simple geometric shapes for controlled analysis, along with queries crafted to precisely test language cues. Furthermore, the dataset is modified to test causal reliance on visual evidence. Our results show that current VLMs combine genuine visual reasoning with shortcut strategies primarily rooted in language cues.
Chinese Translation
视觉语言模型(VLMs)在视觉推理任务中表现出色,但尚不清楚它们是否真正理解视觉关系,还是仅仅依赖于语言提示或先验知识等捷径。为此,我们使用现代视觉语言模型 Qwen3-VL-4B(Bai et al., 2025)来解码视觉信息在不同深度中的编码方式。为进行控制分析,我们提出了一个简单几何形状的合成数据集,并设计了精确测试语言提示的查询。此外,该数据集还经过修改,以测试对视觉证据的因果依赖。我们的结果表明,当前的视觉语言模型将真实的视觉推理与主要基于语言提示的捷径策略相结合。
cs.CV / 183 / 2608.23531

Predicting Multiple Clinical Outcomes Related to Functional Recovery and Social Isolation Among Older Adults After Lower-Limb Fracture or Hip Replacement

预测老年人在下肢骨折或髋关节置换术后与功能恢复和社会孤立相关的多重临床结果
Ray, Santosh, Mishra, Pratik K., Abedi, Ali, Chu, Charlene H., Ahmad, Amir, Khan, Shehroz S.
Abstract
Older adults recovering after lower-limb fracture or hip replacement may experience complex recovery trajectories. Most of the time, these clinical aspects are studied in isolation, masking their joint impact on recovery. This study used the MAISON-LLF dataset, which contains multimodal sensor and clinical assessment data from 18 older adults recovering in the community after lower-limb fracture or hip replacement. Participants were monitored for up to eight weeks, corresponding to a maximum of 1,008 participant-days of sensor monitoring. Forty-six daily features were extracted from indoor motion, acceleration, step count, heart rate, out-of-home mobility, and sleep data. Five clinical outcomes were assessed every two weeks: the Social Isolation Scale, Oxford Hip Score, Oxford Knee Score, Timed Up and Go test, and 30-second Chair Stand test. We utilize an inherent relationship between multi-modal sensor data and different clinical scores and formulate it as a multi-output regression problem. We tested various machine learning and deep learning single- and multi-output regression algorithms to predict these scores simultaneously. The results showed that predicting clinical scores jointly was better than separately. The tabular DL multi-output regressor, NODE, gave a remarkable performance of MSE=3.96 and MAE=1.02 in comparison to other multi- and single-output regressors. The SHAP feature analysis further showed the importance of including multimodal sensors to provide a good estimate of patients' recovery trajectory. This work may support the simultaneous assessment of functional recovery and social engagement among community-dwelling older adults and ultimately help improve their care and quality of life.
Chinese Translation
老年人在下肢骨折或髋关节置换术后恢复过程中可能经历复杂的恢复轨迹。大多数情况下,这些临床方面被孤立研究,掩盖了它们对恢复的共同影响。本研究使用了MAISON-LLF数据集,该数据集包含18名在社区中恢复的老年人的多模态传感器和临床评估数据,数据来源于下肢骨折或髋关节置换术后的恢复。参与者的监测时间最长可达八周,相当于最多1,008个参与者日的传感器监测。从室内运动、加速度、步数、心率、户外活动和睡眠数据中提取了46个每日特征。每两周评估五个临床结果:社会孤立量表、牛津髋关节评分、牛津膝关节评分、定时起立走测试和30秒椅子站立测试。我们利用多模态传感器数据与不同临床评分之间的内在关系,将其构建为一个多输出回归问题。我们测试了多种机器学习和深度学习的单输出和多输出回归算法,以同时预测这些评分。结果表明,联合预测临床评分的效果优于单独预测。表格深度学习多输出回归模型NODE在与其他多输出和单输出回归模型比较时,表现出显著的性能,均方误差为3.96,平均绝对误差为1.02。SHAP特征分析进一步显示了包含多模态传感器以提供患者恢复轨迹良好估计的重要性。本研究可能支持社区老年人功能恢复和社会参与的同时评估,并最终帮助改善他们的护理和生活质量。
cs.CV / 184 / 2608.23549

FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors

FixAnything:通过视频生成先验进行三维一致性渲染精细化
Vuong, Khiem, Ramanan, Deva, Narasimhan, Srinivasa
Abstract
Rendering views using 3D scene representations such as Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), meshes, or even point clouds produces artifacts when input views are sparse or target views lie far from the input. Recent work mitigates these artifacts using diffusion-based generative priors, but is specialized to individual representations and require custom architectures or extensive retraining. We present FixAnything, a single model for fixing a wide range of rendering artifacts. It does so by repurposing a pretrained video generative model, leveraging its implicit multi-view priors with only minimal modification and lightweight finetuning. Our key insight is that even noisily-rendered sequences preserve camera motion and coarse scene structure, allowing cleanup to be formulated as video-to-video translation. To control what scene structure should be preserved, we introduce a binary mask denoting the clean pixels, enabling the model to anchor its output to high-quality inputs (e.g. training views) while refining the rest. To encourage FixAnything to produce 3D-consistent renderings that support downstream reconstruction, we use camera pose accuracy (recovered via structure-from-motion) as a reward signal for direct preference optimization (DPO). Across four distinct 3D representations, FixAnything consistently improves rendering quality with lightweight finetuning, demonstrating that a single generalist video prior can replace multiple specialist refinement pipelines. The simplicity of the framework enables immediate adoption of stronger future video models without architectural redesign.
Chinese Translation
使用高斯溅射(Gaussian Splatting, 3DGS)、神经辐射场(Neural Radiance Fields, NeRF)、网格或甚至点云等三维场景表示进行渲染时,当输入视图稀疏或目标视图远离输入时,会产生伪影。近期的研究通过基于扩散的生成先验来减轻这些伪影,但这些方法专门针对单一表示,并且需要定制架构或大量重新训练。我们提出了FixAnything,这是一个用于修复广泛渲染伪影的单一模型。它通过重新利用一个预训练的视频生成模型,利用其隐式的多视图先验,仅需进行最小修改和轻量级微调。我们的关键见解是,即使是噪声渲染的序列也能保留相机运动和粗略的场景结构,从而使清理过程可以被表述为视频到视频的转换。为了控制应保留的场景结构,我们引入了一个二进制掩码,用于标示干净的像素,使模型能够将其输出锚定于高质量输入(例如训练视图),同时精细化其余部分。为了鼓励FixAnything生成支持下游重建的三维一致性渲染,我们使用相机姿态精度(通过运动结构恢复)作为直接偏好优化(Direct Preference Optimization, DPO)的奖励信号。在四种不同的三维表示中,FixAnything通过轻量级微调持续提高渲染质量,证明了一个通用的视频先验可以替代多个专业的精细化流程。该框架的简单性使得未来更强的视频模型能够立即采用,而无需重新设计架构。
cs.CV / 185 / 2608.23563

EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings

EG-ARSA:一种基于专家的开放模型,用于低资源环境下的视觉道路安全审计
Chowdhury, Md Thamed Bin Zaman, Hossain, Moazzem
Abstract
Road traffic injuries remain a major challenge in low- and middle-income countries, where proactive road safety auditing is limited by incomplete crash records, shortages of qualified auditors, and the high cost of large-scale field inspections. To address this problem, we propose Expert-Grounded Distillation (EGD), a novel artificial intelligence framework that transfers institutional road safety expertise into a compact vision-language model for scalable visual road safety auditing. The key innovation is a quantified expert-grounding stage in which the teacher vision-language model is calibrated against authoritative field audits. Large-scale annotation is permitted only after the teacher reaches substantial agreement with expert risk assessments (Cohen's kappa = 0.74). The calibrated teacher then generates structured supervision that is distilled into an 8-billion-parameter student vision-language model using Low-Rank Adaptation and a single leakage-free prompt. We also introduce Bangladesh Road Safety Audit (BD-ARSA), the first open, expert-grounded Bangladeshi visual road safety audit dataset containing 21,947 image-audit records with near-national coverage, and Expert-Grounded Road Safety Auditor (EG-ARSA), the first vision-language model developed specifically for this task. Experimental results show that grounded fine-tuning substantially improves ordinal risk assessment over the zero-shot baseline, while blind expert evaluation demonstrates that the compact student outperforms both its 31 billion-parameter teacher and Gemini-2.5-Flash. These findings demonstrate that EGD provides an effective and scalable engineering solution for proactive road safety auditing in resource-constrained environments.
Chinese Translation
道路交通伤害在中低收入国家仍然是一个主要挑战,积极的道路安全审计受到不完整的事故记录、合格审计员短缺以及大规模现场检查高成本的限制。为了解决这一问题,我们提出了专家基础蒸馏(Expert-Grounded Distillation, EGD),这是一种新颖的人工智能框架,将机构道路安全专业知识转化为一个紧凑的视觉-语言模型,以实现可扩展的视觉道路安全审计。关键创新在于量化的专家基础阶段,在此阶段,教师视觉-语言模型与权威现场审计进行校准。只有在教师与专家风险评估(Cohen's kappa = 0.74)达成实质性一致后,才允许进行大规模标注。经过校准的教师随后生成结构化监督,这些监督通过低秩适应(Low-Rank Adaptation)和单一无泄漏提示蒸馏到一个具有80亿参数的学生视觉-语言模型中。我们还介绍了孟加拉国道路安全审计(Bangladesh Road Safety Audit, BD-ARSA),这是第一个开放的、基于专家的孟加拉国视觉道路安全审计数据集,包含21,947个图像审计记录,几乎覆盖全国,并且专家基础道路安全审计员(Expert-Grounded Road Safety Auditor, EG-ARSA)是专门为此任务开发的第一个视觉-语言模型。实验结果表明,基础微调显著提高了序数风险评估,相较于零样本基线,盲专家评估显示紧凑的学生模型优于其310亿参数的教师模型和Gemini-2.5-Flash。这些发现表明,EGD为资源受限环境中的主动道路安全审计提供了一种有效且可扩展的工程解决方案。
人工智能 (Artificial Intelligence)
165
cs.AI / 1 / 2608.21362

KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference

KVBoost:基于偏差引导重计算的块级键值缓存重用以提高大型语言模型推理效率
Unnikrishnan, Srihari
Abstract
Transformer-based large language models (LLMs) incur high prefill latency because key-value (KV) tensors must be recomputed for each request. Existing prefix-caching systems reduce this cost but require prompts to share a leading contiguous prefix, limiting effectiveness when shared content appears at arbitrary positions. We present KVBoost, a chunk-level KV cache reuse system for HuggingFace-compatible decoder models that enables reuse regardless of content position. KVBoost introduces a dual-hash keying scheme that separates positional identity (prefix hash) from content identity (content hash), supporting both exact and approximate cache matches. To address attention boundary errors from independently cached chunks, KVBoost employs two repair strategies: SelectiveRecompute, which re-encodes boundary regions, and CacheBlendRecompute, which identifies and recomputes high-deviation tokens after a probe pass. The system further incorporates asymmetric KV quantization (int8/int4), adaptive chunk boundary splitting, and importance-weighted eviction under a fixed memory budget. Evaluated on Qwen/Qwen2.5-3B over 1,000 bug-localization samples, KVBoost achieves a 4.49x reduction in time-to-first-token (142.4 ms vs.\ 639.1 ms) and outperforms prefix caching by 16%, with no loss in accuracy (99.2% vs.\ 99.1%). KVBoost provides a practical, memory-bounded inference acceleration layer compatible with RoPE-based models without architectural modification.
Chinese Translation
基于Transformer的大型语言模型(LLMs)由于每个请求都必须重新计算键值(KV)张量,因此会产生较高的预填充延迟。现有的前缀缓存系统虽然可以降低这一成本,但要求提示共享一个连续的前缀,这在共享内容出现在任意位置时限制了其有效性。我们提出了KVBoost,一种针对HuggingFace兼容解码器模型的块级KV缓存重用系统,能够实现无论内容位置如何的重用。KVBoost引入了一种双哈希键控方案,将位置标识(前缀哈希)与内容标识(内容哈希)分离,支持精确和近似的缓存匹配。为了解决来自独立缓存块的注意力边界错误,KVBoost采用了两种修复策略:SelectiveRecompute,该策略重新编码边界区域,以及CacheBlendRecompute,该策略在探测通过后识别并重新计算高偏差的标记。该系统还结合了非对称KV量化(int8/int4)、自适应块边界分割以及在固定内存预算下的重要性加权驱逐。在对Qwen/Qwen2.5-3B进行1,000个缺陷定位样本的评估中,KVBoost实现了4.49倍的首次标记时间缩短(142.4毫秒对比639.1毫秒),并且比前缀缓存提高了16%的性能,且没有准确性损失(99.2%对比99.1%)。KVBoost提供了一个实用的、受内存限制的推理加速层,兼容基于RoPE的模型,无需架构修改。
cs.AI / 2 / 2608.21363

AIREP: A Protocol for Per-Decision Evidence in AI Runtime Governance

AIREP:一种用于人工智能运行时治理的逐决策证据协议
Abak, Ali Toygar
Abstract
A protocol is presented for recording the governance decisions of automated AI runtimes. When a runtime releases, blocks, defers, redacts, or escalates an individual output, AIREP records that decision as a single signed object that any party can check offline, independent of the runtime that produced it. A record carries the decision as one of a closed set of verbs under a stated policy basis, references its input, output, and evidence by hash rather than by value, and declares both what its evidence covers and what it does not. Records form a SHA-256 hash chain that binds each record to its position, so that tampering and gaps are detectable by recomputation. Vendor-, model-, and domain-specific content is confined to a single optional namespace, and a mechanical neutrality test keeps the shared format free of it. A reference implementation and a two-language conformance kit are described. Some implementation issues are considered, and problems such as alignment of the canonical form across implementations, freshness witnesses, and multi-runtime chains are exposed. The format is offered for adoption by any AI runtime that records governance decisions.
Chinese Translation
本文提出了一种用于记录自动化人工智能运行时治理决策的协议。当运行时发布、阻止、延迟、编辑或升级个别输出时,AIREP将该决策记录为一个单一的签名对象,任何方均可离线检查,独立于产生该决策的运行时。记录以声明的政策基础下的封闭动词集之一携带决策,通过哈希而非值引用其输入、输出和证据,并声明其证据所涵盖的内容及未涵盖的内容。记录形成一个SHA-256哈希链,将每个记录绑定到其位置,从而通过重新计算检测篡改和缺口。特定于供应商、模型和领域的内容被限制在一个可选的命名空间内,机械中立性测试确保共享格式不受其影响。文中描述了一个参考实现和一个双语言符合性工具包。考虑了一些实现问题,并揭示了跨实现的规范形式对齐、新鲜性见证和多运行时链等问题。该格式可供任何记录治理决策的人工智能运行时采用。
cs.AI / 3 / 2608.21366

Reviewing Model Collapse and Countermeasures

模型崩溃及其对策的回顾
Xie, Xihao, Hu, Beichen
Abstract
Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors. The advances of GenAI have actuated practitioners to use AI-synthesized data for training next-generation AI models. Undeniably, using synthetic data has alleviated the increasing stringent demand for data supply. Unfortunately, it also introduces a new critical issue: in a self-consuming cycle between model and data, the model ultimately collapse, raising more trustworthiness concerns to GenAI. In recent years, increasingly more studies have investigated the phenomenon of model collapse (MC) and explored potential solutions to mitigate it. However, the review of the phenomenon of MC still remains blank. To fill this gap, this paper provides an up-to-date overview of these studies for consolidating and reviewing the progress of MC in different application scenarios and countermeasures for mitigating MC. We also highlight challenges and future research opportunities.
Chinese Translation
在海量网络规模数据的驱动下,生成性人工智能(Generative AI, GenAI)取得了显著进展,使其在各个领域的应用成为可能。GenAI的进步促使从业者使用AI合成的数据来训练下一代人工智能模型。不可否认的是,使用合成数据缓解了对数据供应日益严格的需求。然而,这也引入了一个新的关键问题:在模型与数据之间的自我消耗循环中,模型最终崩溃,从而引发了对GenAI的更多信任问题。近年来,越来越多的研究探讨了模型崩溃(Model Collapse, MC)现象,并探索了缓解该现象的潜在解决方案。然而,关于MC现象的综述仍然缺乏。为填补这一空白,本文提供了对这些研究的最新概述,以巩固和回顾MC在不同应用场景中的进展及其缓解对策。我们还强调了挑战和未来的研究机会。
cs.AI / 4 / 2608.21372

AI Learning and Conceptual Transfer in the Game of Hidden Rules

隐性规则游戏中的人工智能学习与概念转移
Mathew, Christo, Wang, Wentian, Feldman, Jacob, Gallos, Lazaros K., Kantor, Paul B., Menkov, Vladimir, Wang, Hao
Abstract
This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules from trial-and-error feedback, representation design, rule difficulty analysis, transfer learning, generalization, and pseudo-bot-assisted human learning analysis. The report focuses on the Transformer-based A2C framework, Feature-Centric and Object-Centric representations, experimental findings, and classification of human learning data.
Chinese Translation
本报告总结了在隐性规则游戏(Game of Hidden Rules, GOHR)中进行的研究工作,重点关注训练强化学习代理从试错反馈中推断隐性规则的过程、表征设计、规则难度分析、迁移学习、概括能力以及伪机器人辅助的人类学习分析。报告主要集中在基于Transformer的A2C框架、特征中心和对象中心的表征、实验发现以及人类学习数据的分类上。
cs.AI / 5 / 2608.21374

LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

文献评审竞技场:使用战斗风格的同行评审平台评估文献综述代理
Zhao, Ruotong, Chen, Zhiyu, Liu, Xurui, Xue, Haidong, Liang, Dong, Fu, Jigao, YanBiao, Wu, Zhen, Yuanyi, Xu, Fengli, Li, Yong
Abstract
Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria. From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%. We further find that existing LLM-as-a-judge methods are substantially misaligned with human experts (Spearman's rho=0.467), especially on synthesis-heavy criteria such as paper structure and research suggestions. Using the collected preference data, we provide an expert-calibrated evaluator, LitJudge, which improves alignment to Spearman's rho=0.78, comparable to inter-expert consistency; code and data are publicly available at https://github.com/VanellopeAsher/LitReview-Arena.
Chinese Translation
文献综述对科学进步至关重要,但严格评估自动生成的综述仍然困难,因为许多研究效用的方面依赖于专家判断,而非参考重叠指标。我们介绍了文献评审竞技场(LitReview Arena),这是一个战斗风格的评估平台,具有针对文献综述质量的结构化协议:具有人工智能论文写作经验的领域专家比较匿名草稿,匹配到他们专业领域内的主题,并在五个特定于文献综述的标准上提供维度结果。通过该协议,我们收集了约3000个专家判断,每个判断包含五个维度结果,并显示即使是当前最强的系统在整体效用上也仅在决定性匹配中赢得23.0%的胜率,而像Sonar Deep Research这样的代理型大型语言模型(LLMs)则显著超越基础语言模型,提升超过60%。我们进一步发现,现有的LLM作为评审的方法与人类专家之间存在显著不一致(Spearman's rho=0.467),尤其是在以综合为主的标准上,如论文结构和研究建议。利用收集的偏好数据,我们提供了一个经过专家校准的评估器LitJudge,其一致性提高至Spearman's rho=0.78,接近专家间的一致性;代码和数据可在https://github.com/VanellopeAsher/LitReview-Arena公开获取。
cs.AI / 6 / 2608.21375

SchemaRouter: Field-Aware Tool Routing for Efficient Heterogeneous Agentic RAG

SchemaRouter:面向字段的高效异构代理RAG工具路由
Cho, Yong-eun
Abstract
Heterogeneous agentic retrieval-augmented generation (RAG) systems increasingly orchestrate external APIs, internal databases, vector stores, and graph stores. Exposing all tool descriptions to an LLM agent, or selecting tools only by vector similarity, causes two costly failures: over-fetching, which increases payload size, token use, and latency, and under-fetching, which omits fields needed to answer the query. We present SchemaRouter, a lightweight routing layer that represents tools, endpoints, parameters, response fields, domain concepts, units, provenance, and license policies as a schema graph. Given a query, SchemaRouter emits an executable tool plan specifying which tools to call and which fields to retrieve. A small LLM extracts intent, concepts, and source constraints, while field selection is deterministic over the graph through intent-group projection and concept-field matching with an alias layer. On a materials-science benchmark of 110 queries, SchemaRouter achieves answer accuracy of 0.71, matching fetch-everything within overlapping confidence intervals and exceeding prompt-all's 0.66, though their intervals overlap. It uses 227 retrieved-context tokens versus 2,066 for fetch-everything and achieves 2.7x lower end-to-end latency than prompt-all. It also obtains the best tool-exact rate of 0.93 and parameter validity of 1.0. SchemaRouter grounds provenance and license information in 62 percent of answers, compared with approximately 0 percent for all baselines. We also find that minimizing selected-field count is counterproductive: it reduces answer accuracy to 0.56 with negligible token savings, while recall-preserving projection restores top accuracy. SchemaRouter improves efficiency, schema-size-independent scaling, and verifiable provenance/license-grounded answering at competitive accuracy.
Chinese Translation
异构代理检索增强生成(RAG)系统越来越多地协调外部API、内部数据库、向量存储和图形存储。将所有工具描述暴露给LLM代理,或仅通过向量相似性选择工具,会导致两种代价高昂的失败:过度获取,增加了负载大小、令牌使用和延迟,以及不足获取,遗漏了回答查询所需的字段。我们提出了SchemaRouter,一个轻量级路由层,将工具、端点、参数、响应字段、领域概念、单位、来源和许可政策表示为一个模式图。给定一个查询,SchemaRouter发出一个可执行的工具计划,指定调用哪些工具以及检索哪些字段。一个小型LLM提取意图、概念和源约束,而字段选择通过意图组投影和概念-字段匹配的别名层在图上是确定性的。在一个包含110个查询的材料科学基准测试中,SchemaRouter实现了0.71的答案准确率,匹配了重叠置信区间内的全取(fetch-everything),并超过了提示所有(prompt-all)的0.66,尽管它们的区间重叠。它使用了227个检索上下文令牌,而全取使用了2,066个,并且实现了比提示所有低2.7倍的端到端延迟。它还获得了最佳的工具精确率0.93和参数有效性1.0。SchemaRouter在62%的答案中提供了来源和许可信息,而所有基线的比例约为0%。我们还发现,最小化选择字段的数量是适得其反的:它将答案准确率降低到0.56,几乎没有令牌节省,而保持召回的投影恢复了最佳准确率。SchemaRouter在竞争性准确率下提高了效率、与模式大小无关的扩展性,以及可验证的来源/许可基础的回答能力。
cs.AI / 7 / 2608.21379

RIACT: A Responsible AI System for Personalized Study Habit Tracking and Early Burnout Signal Detection in University Students

RIACT:一个负责任的人工智能系统,用于大学生个性化学习习惯追踪和早期倦怠信号检测
Sidhu, Ria
Abstract
Student burnout is highly prevalent in higher education, with reported rates ranging from 12% to over 70% and consistently exceeding those of the working population - yet it is typically identified only retrospectively, after academic decline has already occurred. A contributing factor is that students have little structured visibility into their own study behaviour, and existing productivity tools record activity without interpreting it. This paper presents RIACT (Record, Insight, Analyze, Coach, Track), a web-based application that combines structured study session logging with a hybrid AI architecture to surface personalized insights and early burnout signals. Students log sessions by location and time; the system computes net focus time by accounting for breaks, detects burnout signals through transparent, deterministic rules operating on week-over-week behavioural comparisons, and uses a large language model - constrained to a fixed output schema - to contextualize patterns and generate personalized recommendations. The design embeds responsible AI principles throughout: warnings are governed by auditable rules rather than model judgement, all output is framed as an observation rather than a diagnosis and data collection is limited to self-logged behavioural fields. We describe the system's design rationale, situate it within the literature on student burnout and explainable AI in education and propose an evaluation framework for validating its behavioural signals against established burnout instruments.
Chinese Translation
学生倦怠在高等教育中普遍存在,报告的发生率范围从12%到超过70%,且始终高于工作人群的发生率——然而,通常仅在学业下降后才被追溯性识别。一个促成因素是学生对自身学习行为缺乏结构化的可见性,而现有的生产力工具仅记录活动而不进行解释。本文提出了RIACT(记录、洞察、分析、辅导、追踪),这是一个基于网络的应用程序,结合了结构化的学习会话记录与混合人工智能架构,以提供个性化的洞察和早期的倦怠信号。学生通过地点和时间记录会话;系统通过考虑休息时间计算净专注时间,通过透明的、确定性的规则在周与周之间的行为比较中检测倦怠信号,并使用一个大型语言模型——限制在固定输出模式内——来对模式进行上下文化并生成个性化建议。设计中贯穿了负责任的人工智能原则:警告由可审计的规则而非模型判断来管理,所有输出被框定为观察而非诊断,数据收集仅限于自我记录的行为字段。我们描述了系统的设计理念,将其置于学生倦怠和教育中可解释人工智能的文献中,并提出了一个评估框架,以验证其行为信号与已建立的倦怠工具之间的关系。
cs.AI / 8 / 2608.21382

There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items

没有中立的评估工具:现代大型语言模型排行榜由配置脆弱项制造
Parupudi, V. S. Raghu
Abstract
Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model's answer is read from generated text or from per-option likelihoods. Work on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items that separate one model from the next. We treat the evaluation harness of large language models (LLMs) as an independent variable and resolve its effect to single items. We introduce the \textit{fragility grid}: 12 open-weight instruction-tuned LLMs from 4 families answer the same 3{,}679 items from 4 benchmarks (ARC, HellaSwag, MMLU, TruthfulQA) under 26 equally defensible harness configurations, recording one correctness bit for every model, item, and configuration. The comparison is matched, since the items, the weights, and the greedy decoding stay fixed while only the harness varies. Under the grid a model's score is a band rather than a point: gemma4-31b scores between 31 and 89 percent depending only on the harness. Three results follow. On the items that two adjacent models both answer stably the pair is tied, and config-fragile items carry 95.7 percent of a pair's gap on average. Four of the 12 models reach rank one under some configuration, so the harness selects the winner. Item discrimination, the property that benchmark-compression methods maximize, correlates with fragility at 0.28 (95 percent CI 0.25 to 0.30), so compression keeps the fragile items rather than removing them. The scoring choice, not the option order that protocols usually fix, is the load-bearing axis. We release the per-item records and the analysis script, from which every number regenerates on a CPU in seconds, and we position the fragility grid as a check a leaderboard can run before it reports an order.
Chinese Translation
多项选择基准固定了问题和正确答案,但没有固定评估工具:选项的顺序、提示的措辞,以及语言模型的答案是从生成文本中读取还是从每个选项的可能性中读取。关于这种评估工具敏感性的研究报告了其作为总分方差的表现,但未考察方差落在了哪些项目上,以及这些项目是否是区分一个模型与下一个模型的关键。我们将大型语言模型(LLMs)的评估工具视为一个独立变量,并将其影响解析到单个项目上。我们引入了 extit{脆弱性网格}:12个来自4个家族的开放权重指令调优LLMs在26种同样合理的评估配置下回答来自4个基准(ARC、HellaSwag、MMLU、TruthfulQA)的相同3,679个项目,为每个模型、项目和配置记录一个正确性比特。比较是匹配的,因为项目、权重和贪婪解码保持不变,只有评估工具发生变化。在网格下,模型的得分是一个区间而非一个点:gemma4-31b的得分在31%到89%之间,仅取决于评估工具。结果有三点。在两个相邻模型都稳定回答的项目上,这对模型是平局,而配置脆弱的项目平均占据了一对模型间差距的95.7%。12个模型中有4个在某些配置下达到第一名,因此评估工具选择了获胜者。项目区分度,即基准压缩方法最大化的属性,与脆弱性相关性为0.28(95%置信区间0.25到0.30),因此压缩保留了脆弱项目而不是去除它们。评分选择,而非协议通常固定的选项顺序,是承载轴。我们发布了每个项目的记录和分析脚本,任何数字都可以在几秒钟内在CPU上重新生成,并且我们将脆弱性网格定位为排行榜在报告顺序之前可以运行的检查。
cs.AI / 9 / 2608.21393

Spyre-Accelerated Retrieval-Augmented Generation on IBM LinuxONE: A Cloud-Native Architecture for Secure, High-Throughput Enterprise AI Inference

基于IBM LinuxONE的Spyre加速检索增强生成:一种安全、高吞吐量企业AI推理的云原生架构
Bokkasam, Sandeep, D, Pankaj
Abstract
Running large language models inside enterprise environments has always bumped up against a practical wall: the data lives in one place, the AI horsepower sits somewhere else, and moving sensitive records between the two creates real headaches around latency, security, and regulatory exposure. IBM's Spyre accelerator PCIe inference card built for LinuxONE and the broader IBM Z family changes that equation. In this paper we lay out a six-subsystem RAG architecture that runs entirely on IBM LinuxONE, using Spyre for generative inference, the Telum II on-chip accelerator for lightweight classification tasks, and Red Hat OpenShift for container orchestration. Every piece of the pipeline from query intake through vector retrieval, prompt assembly, LLM inference, compliance filtering, and response delivery stays within a single LinuxONE system, so sensitive data never has to leave the hardware perimeter. We walk through the design choices behind each subsystem, dig into the Spyre compilation and serving stack, explain how LinuxONE's Secure Execution technology extends confidential-computing guarantees to AI workloads, and benchmark the architecture against cloud-GPU and on-premises alternatives. Early analysis points to end-to-end RAG latencies under two seconds and up to a 20x reduction compared to off-platform inference, all while keeping the strong encryption and auditability posture that regulated industries actually need.
Chinese Translation
在企业环境中运行大型语言模型一直面临着一个实际的障碍:数据存放在一个地方,而AI计算能力则位于另一个地方,移动敏感记录会带来延迟、安全性和合规性方面的麻烦。IBM为LinuxONE及更广泛的IBM Z系列设计的Spyre加速器PCIe推理卡改变了这一局面。本文提出了一种完全运行在IBM LinuxONE上的六子系统检索增强生成(RAG)架构,利用Spyre进行生成推理,使用Telum II片上加速器处理轻量级分类任务,并通过Red Hat OpenShift进行容器编排。从查询输入到向量检索、提示组装、LLM推理、合规过滤和响应交付的每一个环节均在单一的LinuxONE系统内完成,因此敏感数据无需离开硬件边界。我们详细介绍了每个子系统的设计选择,深入探讨了Spyre的编译和服务栈,解释了LinuxONE的安全执行技术如何将机密计算保障扩展到AI工作负载,并将该架构与云GPU和本地替代方案进行了基准测试。初步分析显示,端到端RAG的延迟低于两秒,与平台外推理相比,延迟减少高达20倍,同时保持了受监管行业所需的强加密和可审计性。
cs.AI / 10 / 2608.21408

Hate Speech Classification In Roman Urdu: A Comparative Study On Parameter Efficient Fine-Tuning And Prompt Engineering

罗马乌尔都语中的仇恨言论分类:参数高效微调与提示工程的比较研究
Zubair, Toneema
Abstract
Due to the widespread accessibility of the internet and social media, toxic and hateful con-tent has grown exponentially, causing significant distress and negative societal impacts. Ro-man Urdu, a low-resource language used in Pakistan and among Urdu-speaking communities worldwide, presents additional challenges because of its informal grammar, inconsistent sen-tence structures, and multiple variations in word spellings. This research aims to identify the most effective techniques for hate speech classification in such low-resource settings with limited data. To address this, the study investigates and compares the latest approaches, in-cluding prompt tuning, parameter-efficient fine-tuning (PEFT) using LoRA, and prompt en-gineering, under various experimental configurations. To achieve this objective, four exper-iments were designed. The first experiment involved direct inferencing with LLMs without any fine-tuning, to evaluate how well these models understand Roman Urdu in a zero-shot setting, especially given limited data. The second experiment utilized parameter-efficient fine-tuning (PEFT) with LoRA, which updates only a small subset of parameters, thereby reducing computational cost. The third experiment explored prompt tuning with both mixed and manually crafted prompts, using very small sets of training examples relative to the entire dataset, making it computationally efficient as well. Finally, the fourth experiment applied prompt engineering through zero-shot and few-shot learning, relying solely on care-fully designed instruction prompts for classification without further training.
Chinese Translation
随着互联网和社交媒体的广泛普及,有害和仇恨内容呈指数级增长,导致显著的社会负面影响。罗马乌尔都语(Roman Urdu)作为一种在巴基斯坦及全球乌尔都语社区中使用的低资源语言,因其非正式语法、不一致的句子结构及多样的词汇拼写变体而带来额外挑战。本研究旨在识别在数据有限的低资源环境下,仇恨言论分类的最有效技术。为此,本文调查并比较了包括提示调优(prompt tuning)、基于LoRA的参数高效微调(parameter-efficient fine-tuning, PEFT)及提示工程(prompt engineering)等最新方法,在多种实验配置下的表现。为实现该目标,设计了四个实验。第一个实验在无微调条件下直接使用大型语言模型(LLMs)进行推理,以评估模型在零样本设置下对罗马乌尔都语的理解能力,尤其是在数据有限的情况下。第二个实验采用基于LoRA的参数高效微调,仅更新少量参数,从而降低计算成本。第三个实验探索了提示调优,使用混合及手工设计的提示,并利用相对于整个数据集非常少量的训练样本,使计算效率得以提升。最后,第四个实验通过零样本和少样本学习的提示工程,仅依赖精心设计的指令提示进行分类,无需进一步训练。
cs.AI / 11 / 2608.21412

The Abstention Protocol: RCA for Clos Fabrics

弃权协议:针对 Clos 结构的根因分析
Gaikwad, Madhava, Pandey, Deepak
Abstract
Root cause analysis (RCA) in large datacenter networks is challenging because telemetry is noisy, partial, and asynchronous. Score-based approaches degrade under these conditions, often yielding unstable or incorrect attributions. We present \textsc{CoreSec}, a production RCA system that replaces weighted fusion with a PAM-style abstention algebra. Telemetry agents are composed using control flags that yield deterministic decisions and explicit abstention when evidence is ambiguous. CoreSec combines this algebra with topology-aware configurations that capture failure surfaces across Clos fabrics and converge monotonically as evidence accumulates. Deployed at hyperscale, CoreSec provides stable and explainable RCA behavior across diverse environments without retuning. Our experience shows that structured composition with abstention forms a practical foundation for automated RCA in real-world cloud networks.
Chinese Translation
在大型数据中心网络中,根因分析(RCA)面临挑战,因为遥测数据往往是嘈杂的、部分的和异步的。在这些条件下,基于评分的方法表现不佳,常常导致不稳定或不正确的归因。我们提出了 extsc{CoreSec},一个生产级的根因分析系统,它用 PAM 风格的弃权代数替代了加权融合。遥测代理通过控制标志组合而成,这些标志在证据模糊时产生确定性的决策和明确的弃权。CoreSec 将这种代数与拓扑感知配置相结合,捕捉 Clos 结构中的故障表面,并随着证据的累积单调收敛。在超大规模环境中部署的 CoreSec 提供了稳定且可解释的根因分析行为,适用于多样化的环境,无需重新调优。我们的经验表明,采用弃权的结构化组合为现实世界云网络中的自动化根因分析奠定了实用基础。
cs.AI / 12 / 2608.21417

Retrieval-grounded robot program generation and simulation-based correction via Model Context Protocol

基于检索的机器人程序生成与通过模型上下文协议的仿真校正
Zhou, Zhichao, Chen, Siyuan, Salunkhe, Omkar, Bekar, Ebru Turanoglu, Stahre, Johan, Skoogh, Anders
Abstract
Flexible manufacturing requires industrial robots to be reprogrammed rapidly as product variants change. This paper presents a language-model-based workflow that generates, validates, and iteratively corrects ABB RAPID robot programs from natural language task descriptions. A dual-stream retrieval-augmented generation (RAG) pipeline grounds code generation in verified technical documentation and production templates, reducing domain-specific errors produced by ungrounded language models. A custom Model Context Protocol (MCP) server connects the language-model client directly to ABB RobotStudio for automated code upload, simulation execution, and diagnostic feedback. The evaluation combines a 30-query retrieval benchmark, scoped code-generation checks, and RobotStudio case studies in a simulated pickand- place manufacturing cell. The simulation loop exposes execution failures that static and semantic checks alone cannot catch, including suction release-height errors, unreachable placement targets, and configuration-dependent recovery motions. The results show how RAG and MCP can connect grounded code generation with executable feedback from industrial robot simulation software, while reducing but not eliminating expert setup and final supervision.
Chinese Translation
灵活制造要求工业机器人在产品变体变化时迅速重新编程。本文提出了一种基于语言模型的工作流程,该流程从自然语言任务描述中生成、验证并迭代修正ABB RAPID机器人程序。双流检索增强生成(RAG)管道将代码生成与经过验证的技术文档和生产模板相结合,从而减少了无基础语言模型产生的特定领域错误。一个定制的模型上下文协议(MCP)服务器将语言模型客户端直接连接到ABB RobotStudio,实现自动代码上传、仿真执行和诊断反馈。评估结合了30个查询的检索基准、范围代码生成检查以及在模拟的拾取和放置制造单元中的RobotStudio案例研究。仿真循环暴露了静态和语义检查无法捕捉的执行失败,包括吸力释放高度错误、不可达的放置目标和依赖配置的恢复动作。结果表明,RAG和MCP如何将基础代码生成与工业机器人仿真软件的可执行反馈连接起来,同时减少但不消除专家设置和最终监督的需求。
cs.AI / 13 / 2608.21418

Composable Trust Infrastructure for Manufacturing Knowledge Graphs: Cross-System Provenance, Temporal Reasoning, and Decision Traceability

可组合的信任基础设施用于制造知识图谱:跨系统溯源、时间推理与决策可追溯性
Chethan, Grama
Abstract
Manufacturing knowledge graphs that integrate data from heterogeneous industrial systems face a trust deficit: consumers cannot determine whether queried data is valid, whether it was valid when a decision was made, where it originated, or how it was acted upon. We argue that four trust capabilities -- SHACL validation, PROV-O provenance, domain-aware bi-temporal versioning, and graph-native decision objects -- compose through shared correlation identifiers to produce emergent trust properties that no single capability delivers alone. We present a composable trust infrastructure that integrates these four capabilities into a unified RDF architecture. Capabilities compose through shared entity URIs, ingestion activity identifiers, and temporal correlation keys, enabling compound queries spanning all four dimensions. An experimental ablation confirms that removing any single capability causes exactly three of six composition queries to fail, demonstrating that all four are equally load-bearing. Analysis of higher-order compositions reveals four emergent three-way properties and one irreducible four-way property (full-chain auditability, 31ms execution). The infrastructure is validated on a testbed integrating eleven industrial sources -- OPC UA, TIA Portal, eClass, AAS, ISA-95, ISA-18.2, SAP S/4HANA, Teamcenter, Opcenter EX, Insights Hub, and SCM -- under an 89-class ISA-95-aligned ontology. The unified graph contains 8,743 triples across five named graphs, stitched by 81 owl:sameAs identity edges. Evaluation uses simulated but structurally realistic data from purpose-built emulators; data structures and cross-system linkage patterns are representative of real industrial installations.
Chinese Translation
集成来自异构工业系统的数据的制造知识图谱面临信任缺失的问题:消费者无法确定查询的数据是否有效,是否在做出决策时有效,数据的来源以及如何被处理。我们认为,四种信任能力——SHACL 验证、PROV-O 溯源、领域感知的双时间版本控制和图形原生决策对象——通过共享的关联标识符组合,产生出单一能力无法独立提供的紧急信任属性。我们提出了一种可组合的信任基础设施,将这四种能力整合到统一的 RDF 架构中。能力通过共享的实体 URI、摄取活动标识符和时间关联键进行组合,使得跨越所有四个维度的复合查询成为可能。实验性消融实验确认,去除任何单一能力会导致六个组合查询中的恰好三个失败,表明这四种能力同样承载着负载。对高阶组合的分析揭示了四种新兴的三元属性和一种不可约的四元属性(全链审计能力,执行时间31毫秒)。该基础设施在一个整合了十一种工业来源的测试平台上进行了验证——OPC UA、TIA Portal、eClass、AAS、ISA-95、ISA-18.2、SAP S/4HANA、Teamcenter、Opcenter EX、Insights Hub 和 SCM——基于89类与ISA-95对齐的本体。统一图谱包含8,743个三元组,分布在五个命名图中,通过81条 owl:sameAs 身份边连接。评估使用来自专用仿真器的模拟但结构上真实的数据;数据结构和跨系统链接模式代表了真实的工业安装。
cs.AI / 14 / 2608.21430

Evaluating Multimodal Narrative Understanding of Popular Hollywood Films

评估流行好莱坞电影的多模态叙事理解
Bamman, David, Chang, Kent K., Cooper, Allison, Hsu, Juishan, Kushihashi, Reina, Mar, Madison, Podichetty, Arnav, Samberg, Rachael, Sancak, Ipek Nil, Shao, Yuhan
Abstract
Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film, opening up new avenues for learning about film history and the evolution of narrative techniques. But the creation of stable benchmarks built around Hollywood films is complicated by copyright protections. In this work, we address these concerns directly, by building a new collection of Hollywood films defined by two criteria: box office popularity (where we publish the first large-scale, open collection of weekly box office earnings reported by Variety magazine from 1922-1979); and likely public domain status (by researching copyright registrations and renewals in the US Catalog of Copyright Entries). We build a new multimodal MCQ benchmark on top of this collection that focuses on narrative elements that directly evaluate the abilities of models to inform meaningful research on film narrative; we find that many vision-language models struggle on this task (with many performing at near-chance levels of accuracy), while audio-visual models (including those that use audio in captioning scenes) reach a maximum accuracy of 61.1%, well below human-level performance.
Chinese Translation
多模态语言模型在大规模计算分析电影方面展现出越来越大的潜力,为学习电影历史和叙事技巧的演变开辟了新的途径。然而,围绕好莱坞电影建立稳定基准的过程受到版权保护的复杂影响。在本研究中,我们直接解决了这些问题,构建了一个新的好莱坞电影集合,该集合基于两个标准:票房受欢迎程度(我们发布了从1922年到1979年《综艺》杂志报告的每周票房收入的首个大规模开放集合);以及可能的公有领域状态(通过研究美国版权登记和续展的记录)。我们在此集合的基础上构建了一个新的多模态多项选择题基准,重点关注直接评估模型在电影叙事研究中提供有意义信息的能力的叙事元素;我们发现许多视觉-语言模型在此任务上表现不佳(许多模型的准确率接近随机水平),而音频-视觉模型(包括那些在场景中使用音频进行字幕处理的模型)达到的最高准确率为61.1%,远低于人类水平的表现。
cs.AI / 15 / 2608.21444

Agentic AI for Safety-critical Multi-drone Systems: Challenges and Opportunities

安全关键多无人机系统中的自主人工智能:挑战与机遇
Merritt, Timothy, Jarabo-Peñas, Alejandro, Bravo-Arrabal, Juan, Bahodi, Maria-Theresa, Christensen, Anders Lyhne
Abstract
Multi-drone systems are increasingly positioned for safety-critical missions such as search and rescue (SAR) and critical infrastructure monitoring. Yet, real-world adoption remains constrained not only by autonomy performance, but by the difficulty of integrating agentic behavior into professional work: operators must understand, trust, and govern automation under uncertainty, time pressure, and accountability. This position paper synthesizes the ambitions and lessons from two ongoing efforts: NAMUR, which explores LLM-supported robot control in SAR and firefighting contexts, and PERSIST, which explores persistent drone operations for monitoring and security at critical infrastructure sites. We argue that agentic AI should be approached as a socio-technical design problem, where interfaces, oversight mechanisms, and evaluation practices are as critical as algorithms. We outline a human-centered, participatory, and iterative research approach aimed at uncovering stakeholder needs, shaping agent capabilities through successive prototypes, and producing transferable proof-of-concept systems and evaluation strategies for other safety-critical contexts.
Chinese Translation
多无人机系统越来越多地被用于安全关键任务,如搜索与救援(SAR)和关键基础设施监测。然而,现实世界的应用不仅受到自主性能的限制,还受到将自主行为整合到专业工作中的困难的制约:操作人员必须在不确定性、时间压力和责任下理解、信任并管理自动化。本文综述了两个正在进行的项目的目标和经验教训:NAMUR,探索在SAR和消防背景下支持大型语言模型(LLM)的机器人控制;以及PERSIST,探索在关键基础设施现场进行持续无人机操作以进行监测和安全。我们认为,自主人工智能应被视为一个社会技术设计问题,其中接口、监督机制和评估实践与算法同样重要。我们概述了一种以人为中心、参与式和迭代的研究方法,旨在揭示利益相关者的需求,通过连续原型塑造代理能力,并为其他安全关键环境生产可转移的概念验证系统和评估策略。
cs.AI / 16 / 2608.21449

Software Frameworks for Explainable AI in Time Series Classification: A Systematic Review

时间序列分类中的可解释人工智能软件框架:系统性综述
Peter, Louis, Gumpfer, Nils, Fischer, Jana, Seifert, Christin, Hannig, Jennifer
Abstract
Time series arise in a wide range of application domains and are analyzed using machine learning in decision-critical settings. Time series classification (TSC) is one of the most widely studied and relevant tasks. In this context, ensuring the transparency and trustworthiness of TSC models has become an important requirement, motivating the use of explainable artificial intelligence (XAI) methods. Despite growing interest, research on XAI for TSC remains fragmented, and a systematic understanding of the available software frameworks for explanation generation, their evaluation practices, and practical limitations is still lacking. Prior work largely focused on individual explanation methods, while cross-framework consistency, time-series-specific evaluation, and reproducibility have received little attention. In this survey, we analyze existing software frameworks for explanation generation and evaluation in TSC. We compare them along multiple dimensions, including supported XAI methods, evaluation metrics, usability, benchmarking support, and reproducibility, providing the first time-series-specific survey of frameworks with implementation comparisons and an analysis of frequency-domain support. We identify six frameworks that explicitly support time series and reveal common limitations: only one method supports frequency-domain explanations despite their relevance; only two evaluation metrics have been developed specifically for time series; and identical XAI methods can yield substantially different explanations across frameworks. Based on these findings, we discuss open challenges and outline directions for future research, highlighting the need for unified, time-series-specific XAI frameworks that enable faithful, reproducible, and time-series-aware explanations.
Chinese Translation
时间序列在广泛的应用领域中出现,并在决策关键的环境中使用机器学习进行分析。时间序列分类(TSC)是最广泛研究和相关的任务之一。在这一背景下,确保TSC模型的透明性和可信性已成为一项重要要求,这促使了可解释人工智能(XAI)方法的使用。尽管兴趣日益增长,但关于TSC的XAI研究仍然零散,缺乏对现有解释生成软件框架、其评估实践和实际局限性的系统理解。以往的研究主要集中在个别解释方法上,而跨框架的一致性、特定于时间序列的评估和可重复性则受到的关注较少。在本次调查中,我们分析了现有的TSC解释生成和评估的软件框架。我们从多个维度对它们进行了比较,包括支持的XAI方法、评估指标、可用性、基准支持和可重复性,提供了首次针对时间序列的框架调查,并进行了实现比较和频域支持的分析。我们识别出六个明确支持时间序列的框架,并揭示了共同的局限性:尽管频域解释相关性强,但仅有一种方法支持频域解释;仅开发了两个专门针对时间序列的评估指标;而且相同的XAI方法在不同框架中可能产生显著不同的解释。基于这些发现,我们讨论了开放挑战,并概述了未来研究的方向,强调了需要统一的、特定于时间序列的XAI框架,以实现真实、可重复和时间序列感知的解释。
cs.AI / 17 / 2608.21463

Enhanced Artificial Neural Networks Using QHAdamW in Air Quality Forecasting

使用QHAdamW增强的人工神经网络在空气质量预测中的应用
Vinas, Mary Joy Daniel
Abstract
The study employed an Artificial Neural Network in combination with the optimized Adaptive Moment Estimation (Adam) algorithm, currently the only AQI forecasting model available in the Philippines. The modified QHAdamW - Quasi-Hyperbolic Momentum (QHAdam) and Adam with decoupled weight decay (AdamW) were both extensions of the Adam optimizer, and both offer unique advantages for training ANN. The proposed QHAdamW optimizer addresses the issues on convergence, generalization, and forecasting performance of Adam. Hyperparameter tuning results revealed that 0.01 and 0.001 were the most effective optimal values for the generalization performance of QHAdamW. The comparative analysis results using seven evaluation metrics revealed that the error value range is lower, and the regression coefficient, having a value approximately equal to 1, improved the model accuracy performance. Likewise, the model converges to a satisfactory level of performance with the convergence performance results of lower loss values as obtained from training and validation losses. Based on data from a real-time air quality tracking station in Manila, a feed-forward neural network is used to predict the AQI of PM2.5 and PM10 separately. This model can be used to forecast Particulate Matter (PM), to help the Department of Environment and Natural Resources-Environmental Monitoring Bureau (DENR-EMB) implement a comprehensive air quality management.
Chinese Translation
本研究采用人工神经网络(ANN)结合优化的自适应动量估计(Adam)算法,这是目前菲律宾唯一的空气质量指数(AQI)预测模型。修改后的QHAdamW - 准双曲动量(Quasi-Hyperbolic Momentum,QHAdam)和具有解耦权重衰减的Adam(AdamW)都是Adam优化器的扩展,均为训练ANN提供了独特的优势。所提出的QHAdamW优化器解决了Adam在收敛性、泛化能力和预测性能方面的问题。超参数调优结果显示,0.01和0.001是QHAdamW泛化性能的最有效最优值。使用七个评估指标的比较分析结果表明,误差值范围更低,回归系数接近1,提升了模型的准确性表现。同时,模型在训练和验证损失中获得的较低损失值显示出令人满意的收敛性能。基于马尼拉实时空气质量监测站的数据,采用前馈神经网络分别预测PM2.5和PM10的AQI。该模型可用于预测颗粒物(PM),以帮助环境与自然资源部-环境监测局(DENR-EMB)实施全面的空气质量管理。
cs.AI / 18 / 2608.21501

Let Credit Follow Computation: Architecture-Aware Credit Transport for Large Language Model Reinforcement Learning

让信用随计算而动:面向架构的信用传输用于大型语言模型强化学习
Shi, Qifan, Kang, Zhaolu, Zhu, Chenghua
Abstract
Credit assignment in large-language-model reinforcement learning (LLM RL) can be separated into three objects: evidence about success, a transport operator that converts this evidence into token-level advantages, and an update geometry that turns advantages into policy changes. Recent work has greatly improved evidence, sampling, and update geometry, but the transport operator is usually architecture-agnostic. Fixed-discount GAE applies a stationary geometric kernel along token time; group-relative methods broadcast an outcome statistic across an entire response. Neither operator represents the trajectory-specific computation used by the Transformer policy itself. We introduce computation-conditioned credit transport (CCT), a general framework in which a detached statistic of the behavior policy's internal computation parameterizes the causal kernel that transports downstream value through a rollout. Our concrete algorithm, CompPO, maps native attention concentration to a bounded per-token retention gate, uses the gate in both the one-step bootstrap and a path-dependent generalized-advantage trace (Comp-GAE), and co-designs a transport-aligned critic (TAC) that reuses the actor's hidden states and routing information without a second same-scale Transformer. The task reward and clipped PPO policy objective remain unchanged; a constant gate recovers fixed-coefficient GAE. Across five Qwen3-4B seeds, CompPO reaches 61.4% final held-out accuracy (95% CI [60.8,62.0]) versus 53.8% [52.9,54.7] for tuned GRPO. Neither Comp-GAE with a standard critic (55.2%) nor TAC with a fixed gate (56.4%) matches the full model (interaction +2.4 [1.9,2.9]). Shuffle and position controls confirm trajectory-specific alignment; CompPO is stable in 10/12 PPO-grid runs versus 3/12. Frozen evaluation improves over GRPO by 4.3 and 3.9 greedy pass@1 macro points on Qwen3-4B and Llama-3.1-8B-Instruct.
Chinese Translation
在大型语言模型强化学习(LLM RL)中,信用分配可以分为三个对象:关于成功的证据、将这些证据转换为令牌级优势的传输算子,以及将优势转化为策略变化的更新几何。最近的研究在证据、采样和更新几何方面取得了显著进展,但传输算子通常是与架构无关的。固定折扣的广义优势估计(GAE)沿令牌时间应用一个静态几何核;组相对方法则在整个响应中广播一个结果统计。这两种算子都未能代表变换器(Transformer)策略本身所使用的轨迹特定计算。我们提出了计算条件信用传输(CCT),这是一个通用框架,其中行为策略内部计算的脱离统计量参数化了通过回滚传输下游价值的因果核。我们具体的算法CompPO将原生注意力集中映射到一个有限的每令牌保留门,在一步自助法和路径依赖的广义优势追踪(Comp-GAE)中使用该门,并共同设计了一个与传输对齐的评论员(TAC),该评论员在不使用第二个同规模变换器的情况下重用演员的隐藏状态和路由信息。任务奖励和剪切的PPO策略目标保持不变;一个常数门恢复固定系数的GAE。在五个Qwen3-4B种子上,CompPO达到了61.4%的最终保留准确率(95%置信区间[60.8,62.0]),而调优的GRPO为53.8% [52.9,54.7]。无论是使用标准评论员的Comp-GAE(55.2%)还是使用固定门的TAC(56.4%),都无法与完整模型(交互 +2.4 [1.9,2.9])相匹配。洗牌和位置控制确认了轨迹特定的对齐;CompPO在12次PPO网格运行中稳定性为10次,而其他方法为3次。冻结评估在Qwen3-4B和Llama-3.1-8B-Instruct上分别比GRPO提高了4.3和3.9个贪婪通过@1宏观点。
cs.AI / 19 / 2608.21567

Quantifying geographic domain shift to decouple the geospatial transferability of human mobility flow generation models

量化地理领域转移以解耦人类流动生成模型的地理空间可转移性
Zhou, Zhiyong, Gao, Song, Zhang, Qianheng, Zhang, Feng, Du, Zhenhong
Abstract
Human mobility serves as an essential proxy for understanding social, economic, and environmental dynamics in urban systems. Geospatial transferability, which measures a model's capability in a new location or unseen region, is a critical dimension for comparing different human mobility generation models. However, few studies have studied the intrinsic characteristics of geospatial transferability. To this end, this study systematically investigates the geospatial transferability of four representative human mobility generation models using a large-scale benchmark dataset of census tract level commuting flows across 2265 counties in the United States. Inspired by the domain adaptation theory in machine learning, we introduce geographic domain shift to describe the intrinsic differences in geographic feature distributions and spatial structures between source and target regions, which may jointly affect model transferability. Moreover, we propose two metrics, mutual information and spatial shift, to quantify the geographic domain shift. To examine their associations with model transferability, we employ linear mixed-effects regression to analyze the associations between geographic domain shifts and transferability. Our results reveal substantial spatial heterogeneity and asymmetry in transfer performance across regions. Both information shift and spatial shift exhibit statistically significant and complementary explanatory power. This indicates that geospatial transferability depends not only on model design but also on intrinsic geographic differences. These findings provide a novel methodological framework for evaluating and improving the geospatial transferability of human mobility generation models and support more robust and fair human mobility data synthesis across diverse regions. It also offers insights on spatial transferability for GeoAI model development.
Chinese Translation
人类流动作为理解城市系统中社会、经济和环境动态的重要代理。地理空间可转移性衡量模型在新地点或未见区域的能力,是比较不同人类流动生成模型的关键维度。然而,关于地理空间可转移性的内在特征的研究较少。为此,本研究系统地调查了四个代表性人类流动生成模型的地理空间可转移性,使用了涵盖美国2265个县的普查区级通勤流的大规模基准数据集。受到机器学习中领域适应理论的启发,我们引入了地理领域转移的概念,以描述源区域和目标区域之间地理特征分布和空间结构的内在差异,这些差异可能共同影响模型的可转移性。此外,我们提出了两个指标——互信息和空间转移,以量化地理领域转移。为了检验它们与模型可转移性之间的关联,我们采用线性混合效应回归分析地理领域转移与可转移性之间的关系。我们的结果揭示了不同区域之间转移性能的显著空间异质性和不对称性。信息转移和空间转移均表现出统计显著性和互补的解释力。这表明,地理空间可转移性不仅依赖于模型设计,还依赖于内在的地理差异。这些发现为评估和改善人类流动生成模型的地理空间可转移性提供了一种新颖的方法论框架,并支持在不同区域之间进行更稳健和公平的人类流动数据合成。同时,这也为GeoAI模型开发提供了空间可转移性的见解。
cs.AI / 20 / 2608.21570

A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety Classification

可复现的、关注许可证的可部署安全分类蒸馏方案
Filho, Edson Rodrigues da Cruz, Neves, Paulo Ricardo Ferreira, Falsetti, Paulo Henrique Eleuterio, Pavan, João Vitor, Degaspari, Ian, Laturrague, Henrique Vieira, Laturrague, Patrick Vieira, Dias, Guilherme Nielsen, Berto, Marccello Wilson Perez, Von Atzingen, Gustavo Voltani
Abstract
Deploying a safety layer for large language models on commodity hardware is constrained by the guards available to do it: current open guard models hold between 1 and 9 billion parameters, are oriented toward the graphics processing unit, and answer in seconds per request on a central processing unit. This paper presents a reproducible, license-aware knowledge-distillation recipe addressing that constraint. A strong open guard labels a corpus of roughly 97,000 prompts, drawn from 24 public datasets, into seven safety categories aligned to a public hazard taxonomy, and a fleet of small students spanning lexical, shallow, encoder and generative architectures is trained to reproduce that signal. The corpus is partitioned at the license boundary, so that a deployable and a research model differ only in their training data and the cost of that restriction becomes measurable. Every model is scored against an independent gold benchmark of 6,361 rows over four slices, labeled apart from the teacher and including a slice of harmless prompts that makes over-defense measurable. The distilled students match the teachers on adversarial text within overlapping confidence intervals and reduce false alarms on harmless prompts, the smallest generative student reaching 3.8% against 4.8% for the 8-billion-parameter teacher, while the encoder classifies in roughly 24 ms per request on CPU. Per-class rebalancing is the only decisive ingredient of the recipe. No superiority over the distilled guards is claimed; on the clean reference slice they remain ahead.
Chinese Translation
在商品硬件上为大型语言模型部署安全层受到可用防护措施的限制:当前的开放防护模型参数量在10亿到90亿之间,主要面向图形处理单元,并且在中央处理单元上每个请求的响应时间为几秒。本文提出了一种可复现的、关注许可证的知识蒸馏方案,以应对这一限制。一个强大的开放防护模型对来自24个公共数据集的约97,000个提示进行标注,分为七个安全类别,与公共危害分类法相一致,并训练了一系列小型学生模型,包括词汇、浅层、编码器和生成架构,以重现该信号。该语料库在许可证边界处进行划分,使得可部署模型和研究模型仅在训练数据上有所不同,而这种限制的成本变得可量化。每个模型都与一个独立的金标准基准进行评分,该基准包含6,361行数据,分为四个切片,标注与教师模型分开,并包括一部分无害提示,使得过度防御的情况可被量化。蒸馏后的学生模型在重叠的置信区间内与教师模型在对抗文本上的表现相匹配,并减少了无害提示的误报,最小的生成学生模型在无害提示上的误报率为3.8%,而8亿参数的教师模型为4.8%,同时编码器在CPU上每个请求的分类时间约为24毫秒。每类的重新平衡是该方案的唯一决定性成分。本文并未声称蒸馏后的防护模型优于原始防护模型;在干净的参考切片上,后者仍然表现更佳。
cs.AI / 21 / 2608.21583

Robust Lightweight Deep Learning Models for Oral Cancer Screening

用于口腔癌筛查的稳健轻量级深度学习模型
Bharadwaj, Siddhant, Shedsale, Aakash, Subramanya, Tejashree, Azfar, Mohd., Birur, Praveen, Pal, Debnath, Sharma, Shankararama, Shetty, Anupama, Sundaresan, Rajesh
Abstract
Oral cancer is a leading cause of mortality in low-to-middle-income countries, where a shortage of specialists delays diagnosis. While point-of-care screening via smartphones offers a scalable solution, developing robust AI for resource-constrained settings poses significant challenges, including class imbalance in training data, variable data quality, and computational constraints on edge devices. In this paper, we present the optimisation of lightweight deep learning models for smartphone-based oral cancer screening. Using a diverse, multi-centre retrospective dataset of approximately 30,000 images acquired over a decade, we systematically evaluate state-of-the-art convolutional, transformer, and hybrid architectures. Through rigorous pipeline ablation, we demonstrate that directly optimising hybrid architectures for the edge strictly outperforms computationally heavy paradigms, such as large models or knowledge distillation. Furthermore, interpretability analysis and simulated noise-stress tests revealed that the system anchors on clinical features and remains robust to unstructured sensor noise, despite vulnerabilities to impulse bit errors. In the held-out test set, our optimised MobileViTv2 models achieved an average sensitivity of 83.2 $\pm$ 1.5% and an average specificity of 86.0 $\pm$ 0.8%, with the best model exhibiting 87.4% sensitivity, 86.5% specificity, and a critical negative predictive value of 97.2% with reference to specialist labels. These results confirm that with targeted architectural selection and streamlined optimisation, interpretable and robust lightweight AI models exhibit high potential for edge deployment to enable automated triage in primary care settings.
Chinese Translation
口腔癌是中低收入国家主要的死亡原因之一,专家短缺导致诊断延迟。尽管通过智能手机进行即时筛查提供了一种可扩展的解决方案,但在资源有限的环境中开发稳健的人工智能面临重大挑战,包括训练数据中的类别不平衡、数据质量的可变性以及边缘设备上的计算限制。在本文中,我们提出了针对基于智能手机的口腔癌筛查的轻量级深度学习模型的优化。利用一个多中心的回顾性数据集,该数据集包含约30,000张在十年内获取的图像,我们系统地评估了最先进的卷积、变换器和混合架构。通过严格的管道消融实验,我们证明了直接优化边缘设备的混合架构在性能上明显优于计算负担重的范式,如大型模型或知识蒸馏。此外,解释性分析和模拟噪声压力测试表明,该系统依赖于临床特征,并且在面对非结构化传感器噪声时保持稳健,尽管对脉冲位错误存在脆弱性。在保留的测试集中,我们优化后的MobileViTv2模型达到了83.2 ± 1.5%的平均灵敏度和86.0 ± 0.8%的平均特异性,最佳模型展现了87.4%的灵敏度、86.5%的特异性以及97.2%的关键负预测值(以专家标签为参考)。这些结果确认,通过有针对性的架构选择和精简优化,具有解释性和稳健性的轻量级人工智能模型在边缘部署中展现出高潜力,以实现初级护理环境中的自动分诊。
cs.AI / 22 / 2608.21584

Data-Driven Dynamic Algorithm Dispatch with Large Language Models

基于数据驱动的大型语言模型动态算法调度
Shah, Rushil, Lujan, Emmanuel, Alomairy, Rabab, Edelman, Alan
Abstract
We introduce a large language model (LLM)-driven approach for generating dynamic algorithmic dispatch heuristics in high-performance linear algebra. By combining prompt engineering with LLaMA 3 and a curated performance database, the model learns to synthesize selection heuristics that exploit structural patterns to identify fast algorithmic choices. A case study on LU factorization demonstrates the model's ability to replicate expert-designed strategies. This work, developed as part of the DARPA-MIT SmartSolve project, highlights the promise of LLMs for algorithmic discovery and the development of more adaptive, fast linear algebra software.
Chinese Translation
我们提出了一种基于大型语言模型(LLM)的方法,用于生成高性能线性代数中的动态算法调度启发式。通过将提示工程与 LLaMA 3 和一个精心策划的性能数据库相结合,该模型学习合成选择启发式,利用结构模式识别快速的算法选择。对 LU 分解的案例研究展示了该模型复制专家设计策略的能力。本研究作为 DARPA-MIT SmartSolve 项目的一部分,突显了 LLM 在算法发现和开发更具适应性、快速的线性代数软件方面的潜力。
cs.AI / 23 / 2608.21601

K-Bench: measuring model performance on real scientific agent requests

K-Bench:在真实科学代理请求上测量模型性能
Brueckner, Aubrey, Patel, Darshil, He, Yuhuan, Kassis, Timothy
Abstract
Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or simulators with a known generative structure. Real scientific requests arrive differently. They are underspecified, they carry attachments, and lack ground truth. We report K-Bench 01, an evaluation built from first-turn requests sampled from live user traffic on K-Dense Web and run end to end by nine frontier models in identical sandboxes, yielding 1,602 completed agent runs. Three blinded language-model judges scored every run against an eight-dimension rubric. On a rubric whose 8-anchor instructs judges that a domain scientist would accept the work with minor edits, no model clears the line under all three judges. gpt-5.6-sol has the highest pooled mean, 8.04, but its 95% interval [7.80, 8.23] spans the threshold, and two of the three judges rank claude-opus-5 first instead. We therefore report the ordering of systems as the reproducible quantity, the absolute level as an attribute of the instrument, and the top of the table as unresolved. Across all 39,934 scored judgments -- the eight dimension scores plus a holistic overall for each assessment, excluding not-applicable cells -- 47.6% fall below the 8-point threshold. Difficulty is not uniform across the rubric: scientific accuracy averages 6.22 against 7.33 for communication, on identical denominators and in the same direction within every one of the nine models. The single leading failure tag is overclaiming, on 31.4% of assessments. We argue that the informative quantity for scientific agents is not a leaderboard position but the joint distribution of what was delivered, what was claimed, and what artifacts were produced.
Chinese Translation
科学人工智能的基准测试大多是为了评分而编写的:多项选择题、带有参考解决方案的策划代理任务或具有已知生成结构的模拟器。真实的科学请求则有所不同。它们通常描述不充分,带有附件,并且缺乏真实标准。我们报告了 K-Bench 01,这是一个基于从 K-Dense Web 的实时用户流量中抽样的首次请求构建的评估,并由九个前沿模型在相同的沙箱中端到端运行,产生了 1,602 次完成的代理运行。三位盲评语言模型评审根据一个八维标准对每次运行进行了评分。在一个八个维度的标准中,指示评审认为领域科学家会接受该工作并进行小幅编辑,没有任何模型在所有三位评审中都达到了标准线。gpt-5.6-sol 的平均得分最高,为 8.04,但其 95% 置信区间 [7.80, 8.23] 跨越了阈值,而三位评审中有两位将 claude-opus-5 排在第一。因此,我们报告系统的排序作为可重复的量,绝对水平作为工具的属性,而表格的顶部则仍未解决。在所有 39,934 个评分判断中——包括八个维度的分数以及每次评估的整体评分(不包括不适用的单元)——47.6% 的评分低于 8 分的阈值。标准的难度在各个维度上并不均匀:科学准确性平均得分为 6.22,而沟通得分为 7.33,所有九个模型在相同的分母和相同的方向上均如此。单一的主要失败标签是过度声明,出现在 31.4% 的评估中。我们认为,对于科学代理而言,重要的量不是排行榜位置,而是所交付内容、所声称内容和所产生工件的联合分布。
cs.AI / 24 / 2608.21605

Generate in the Chart, Not on the Boundary: Function-Symbol Grounding for Hard Constraints in LTN-GANs

在图表中生成,而非在边界上:LTN-GAN中硬约束的函数符号基础
Upreti, Nijesh, Belle, Vaishak
Abstract
Logic Tensor Network-Enhanced Generative Adversarial Networks (LTN-GANs) inject background knowledge by grounding each logical axiom as a predicate and training the generator to raise its satisfaction, a fuzzy truth value in $[0,1]$. Previous LTN-GAN work grounded every constraint this way, at the predicate level, and improved constraint satisfaction. A predicate, however, only scores a sample, so it cannot embed hard structural constraints, rules such as orderings, positivity, and definitional identities that must hold in every generated sample. In this work, we investigate grounding each axiom as a function symbol inside the LTN framework. We compare against the state-of-the-art alternative, a constraint layer that clamps each violating sample onto the feasible boundary and so produces outputs that are always valid. Our investigation shows that a valid sample is not always a realistic one. An inequality is not merely satisfied or violated. It holds by a margin, and a faithful generator should also reproduce the margin's real distribution. We find that the resolution ratio $R$, the data's scale over the margin's spread, is a diagnostic, computable before training, of which constraints a chosen grounding can learn. When $R$ is large, the predicate receives no learning signal, the clamp pushes every sample onto the boundary, and the margin distribution is lost while every standard metric still looks fine. A function symbol avoids both failures, computing the constrained variable rather than scoring it. Together the function symbols form a chart, a coordinate system inside the feasible region, where every sample is valid by construction and the margin is learned like any other quantity.
Chinese Translation
逻辑张量网络增强生成对抗网络(LTN-GAN)通过将每个逻辑公理作为谓词进行基础化来注入背景知识,并训练生成器提高其满足度,即在 $[0,1]$ 范围内的模糊真值。之前的 LTN-GAN 研究以这种方式在谓词层面基础化每个约束,并改善了约束满足度。然而,谓词仅对样本进行评分,因此无法嵌入硬结构约束,例如必须在每个生成样本中成立的顺序、正性和定义身份等规则。在本研究中,我们探讨了在 LTN 框架内将每个公理基础化为函数符号。我们与最先进的替代方案进行比较,该方案是一个约束层,它将每个违反的样本固定在可行边界上,从而产生始终有效的输出。我们的研究表明,有效样本并不总是现实样本。不等式不仅仅是被满足或被违反。它以一定的余量成立,而一个忠实的生成器也应重现余量的真实分布。我们发现,分辨率比率 $R$,即数据的规模与余量的扩展之比,是一个可在训练前计算的诊断指标,用于判断所选基础化可以学习哪些约束。当 $R$ 较大时,谓词接收不到学习信号,夹具将每个样本推向边界,余量分布丧失,而每个标准指标仍然看起来良好。函数符号避免了这两种失败,计算约束变量而不是对其评分。函数符号共同形成一个图表,这是可行区域内的坐标系统,其中每个样本在构造上都是有效的,余量像其他任何量一样被学习。
cs.AI / 25 / 2608.21610

Semantic Compression Trees: Multi-Resolution Knowledge Retrieval via Hierarchical Semantic Residuals

语义压缩树:通过层次语义残差进行多分辨率知识检索
Farooq, Junaid
Abstract
Retrieval-augmented generation relies mostly on flat, fixed-granularity indexes: documents are cut into uniform chunks and retrieved by similarity, discarding the hierarchical structure of the source. We introduce Semantic Compression Trees (SCT), a hierarchical index in which each node stores only its semantic residual -- the information it adds beyond its parent -- and retrieval proceeds by progressive descent from the root, so that per-query cost is governed by tree depth rather than collection size. We evaluate on QASPER (50 papers, 173 questions) under two protocols differing only in whether the benchmark supplies the relevant document, with bootstrap confidence intervals and paired significance tests throughout. The results are mixed and we report them as such. When the document is given, SCT with a zero-LLM extractive compressor matches dense retrieval on answer quality (0.274 vs. 0.277 F1, $p = 0.37$) using 30% fewer context tokens and no LLM calls to build the index, and residual storage beats storing full summaries at each node (0.274 vs. 0.205, $p < 0.001$). Increasing the collection fifty-fold multiplies flat retrieval's per-query scoring work by 48.9x and SCT's by 6.4x. Progressive descent itself is not supported. Retrieving the same residuals without the tree performs identically when the document is given ($p = 0.27$), and descent is substantially worse when the system must select the document (0.122 vs. 0.165, $p < 0.001$). Routing accuracy localises the cause: descent selects the correct paper 20.2% of the time against 39.3% for flat retrieval, because that choice is made from the root residual, the most compressed node in the tree. We conclude that the residual representation is worth keeping and top-down routing is not.
Chinese Translation
检索增强生成主要依赖于平坦的、固定粒度的索引:文档被切割成均匀的块并通过相似性进行检索,忽略了源文档的层次结构。我们提出了语义压缩树(Semantic Compression Trees, SCT),这是一种层次索引,其中每个节点仅存储其语义残差——即其相对于父节点所增加的信息——并且检索通过从根节点逐步下降进行,因此每次查询的成本由树的深度而非集合大小决定。我们在 QASPER(50 篇论文,173 个问题)上进行了评估,采用两种协议,唯一的区别在于基准是否提供相关文档,整个过程中使用了自助法置信区间和配对显著性检验。结果是复杂的,我们如实报告。当文档被提供时,使用零 LLM 抽取压缩器的 SCT 在答案质量上与密集检索相匹配(0.274 对 0.277 F1,$p = 0.37$),使用了 30% 更少的上下文标记,并且在构建索引时没有 LLM 调用,而残差存储在每个节点上优于存储完整摘要(0.274 对 0.205,$p < 0.001$)。将集合增加五十倍使得平坦检索的每次查询评分工作增加了 48.9 倍,而 SCT 增加了 6.4 倍。逐步下降本身并不被支持。在给定文档的情况下,检索相同的残差表现相同($p = 0.27$),而当系统必须选择文档时,下降的表现显著较差(0.122 对 0.165,$p < 0.001$)。路由准确性定位了原因:下降选择正确论文的概率为 20.2%,而平坦检索为 39.3%,因为该选择是基于根节点残差做出的,该节点是树中压缩程度最高的节点。我们得出结论,残差表示是值得保留的,而自上而下的路由则不是。
cs.AI / 26 / 2608.21614

SAEM: Stage-Aware Expert Management for Memory-Efficient MoE Inference in Chain-of-Thought Reasoning

SAEM:面向阶段的专家管理用于链式思维推理中的内存高效 MoE 推断
Zhang, Yujie, Gao, Bin, Mitra, Tulika
Abstract
Chain-of-thought (CoT) prompting improves LLM reasoning by decomposing complex problems into intermediate steps, but its sequential nature increases decoding latency and memory usage. Mixture-of-Experts (MoE) models scale capacity through sparse expert activation, yet their full expert weights often exceed GPU memory and require costly GPU-CPU transfers. Existing runtimes treat all tokens uniformly, overlooking a key structural property of CoT traces: consecutive reasoning stages exhibit coherent and predictable expert activation patterns. Ignoring this stage-level regularity leads to inefficient caching and unnecessary data movement. We propose SAEM, a stage-aware MoE inference runtime that detects reasoning stage boundaries and exploits stage-level activation coherence to guide expert placement. SAEM combines stage-aware caching, expert-aligned token repacking, and in-situ CPU execution to reduce data transfer and kernel fragmentation. On mathematical and scientific reasoning workloads, SAEM achieves an average 1.33x throughput improvement over the strongest state-of-the-art caching and offloading baselines under constrained GPU memory, rising to 1.54x when calibration data matches the workload, demonstrating the effectiveness of stage-aware, locality-driven MoE inference for CoT reasoning.
Chinese Translation
链式思维(CoT)提示通过将复杂问题分解为中间步骤来改善大型语言模型(LLM)的推理能力,但其顺序特性增加了解码延迟和内存使用。专家混合(MoE)模型通过稀疏专家激活来扩展容量,但其完整的专家权重通常超过 GPU 内存,并且需要昂贵的 GPU-CPU 数据传输。现有的运行时将所有标记视为均匀处理,忽视了 CoT 路径的一个关键结构特性:连续的推理阶段表现出一致且可预测的专家激活模式。忽视这种阶段级的规律性导致了低效的缓存和不必要的数据移动。我们提出了 SAEM,一种面向阶段的 MoE 推断运行时,能够检测推理阶段边界并利用阶段级激活一致性来指导专家放置。SAEM 结合了面向阶段的缓存、专家对齐的标记重打包和原位 CPU 执行,以减少数据传输和内核碎片化。在数学和科学推理工作负载上,SAEM 在受限 GPU 内存下实现了比最强的最先进缓存和卸载基线平均提升 1.33 倍的吞吐量,当校准数据与工作负载匹配时,提升达到 1.54 倍,证明了面向阶段的、基于局部性的 MoE 推断在 CoT 推理中的有效性。
cs.AI / 27 / 2608.21664

Measuring Activation Control in Large Language Models

测量大型语言模型中的激活控制
Kowalski, Marek Mateusz, Rivera, Joshua Fonseca, Macar, Uzay, Africa, David Demitri
Abstract
Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especially when evaluation-aware models exhibit scheming or deception. However, if models can also control their own activations, deception could extend into the latent space itself. With this in mind, we introduce the Activation Controllability Benchmark to quantify the extent to which models can modulate their residual stream via natural-language instruction. Across model families and capability levels, we find that most LLMs can control the direction and magnitude of their residual stream activations with some degree of temporal resolution, though performance varies considerably across models. In simple tasks, this level of control can evade activation-based monitoring methods (including linear probes, natural language autoencoders, activation oracles, and the Jacobian lens), albeit imperfectly. These results suggest that control over the activation space itself could become a confound for monitoring as introspective capabilities increase; therefore, we recommend that frontier labs and evaluators track activation controllability in future models.
Chinese Translation
安全部署日益强大的模型可能会依赖于潜在空间监控,作为行为评估的补充,特别是在评估意识模型表现出策划或欺骗行为时。然而,如果模型也能够控制自身的激活,欺骗可能会扩展到潜在空间本身。考虑到这一点,我们引入了激活可控性基准,以量化模型通过自然语言指令调节其残差流的能力。在不同的模型家族和能力水平中,我们发现大多数大型语言模型(LLMs)能够在一定的时间分辨率下控制其残差流激活的方向和幅度,尽管模型之间的表现差异显著。在简单任务中,这种控制水平可以规避基于激活的监控方法(包括线性探测器、自然语言自编码器、激活神谕和雅可比透镜),尽管效果并不完美。这些结果表明,控制激活空间本身可能会成为监控的混淆因素,随着内省能力的提高;因此,我们建议前沿实验室和评估者在未来的模型中跟踪激活可控性。
cs.AI / 28 / 2608.21668

From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation

从掌握特征到模拟响应:用于真实学生模拟的随机学生知识图谱(SSKG)
An, Yuan, Wang, Emily, Wang, Benjamin, Hashmi, Ruhma
Abstract
Large language models (LLMs) are increasingly used to simulate students at different mastery levels. These simulations can generate synthetic training data and stress-test tutoring systems. However, common prompt-based approaches leave the answer decision to the LLM, which tends to perform according to its built-in capabilities even when instructed to simulate a student with low mastery. As a result, these approaches may have difficulty distinguishing students with low and high levels of mastery. We demonstrate this limitation using 379 College Board-calibrated SAT Algebra items and five archetypal mastery profiles. Three LLMs from three vendors (Gemini 3.1 Flash Lite, Claude Haiku 4.5, and GPT-5.4-mini) achieve 96.8-100% accuracy across all profiles. To address this limitation, we introduce a method grounded in a Stochastic Student Knowledge Graph (SSKG). A curriculum knowledge graph (CKG) is extracted from an open algebra textbook, and each SAT solution is decomposed into a chain of required triples. The SSKG assigns a mastery probability to each triple, which is sampled to determine question correctness. An LLM then generates a first-person rationale consistent with the outcome. The simulation reduces accuracy to 44.1-85.2% across profiles and produces a clear monotone mastery gradient.
Chinese Translation
大型语言模型(LLMs)越来越多地用于模拟不同掌握水平的学生。这些模拟可以生成合成训练数据,并对辅导系统进行压力测试。然而,常见的基于提示的方法将答案决策留给LLM,尽管被指示模拟低掌握水平的学生,它仍然倾向于根据其内置能力进行表现。因此,这些方法可能难以区分低掌握和高掌握水平的学生。我们使用379个由大学理事会校准的SAT代数题目和五个典型的掌握特征展示了这一局限性。来自三家供应商的三种LLM(Gemini 3.1 Flash Lite、Claude Haiku 4.5和GPT-5.4-mini)在所有特征上实现了96.8-100%的准确率。为了解决这一局限性,我们引入了一种基于随机学生知识图谱(SSKG)的方法。从一本开放的代数教科书中提取课程知识图谱(CKG),并将每个SAT解答分解为一系列所需的三元组。SSKG为每个三元组分配一个掌握概率,并通过采样来确定问题的正确性。然后,LLM生成与结果一致的第一人称理由。该模拟将准确率降低到44.1-85.2%,并产生明显的单调掌握梯度。
cs.AI / 29 / 2608.21690

Context as an Environment: Programmatic Context Management for Long-Horizon Agents

环境中的上下文:面向长时间任务代理的程序化上下文管理
Lin, Yin, Ang, Elaine, Zhu, Erkang, Ding, Bolin, Zhou, Jingren
Abstract
LLM agents increasingly take on long-running tasks whose history grows far beyond a single model context window. Existing approaches compress earlier interactions or extract selected information into fixed memory representations, committing to what to preserve before future needs are known. We present Scroll, a context manager that treats each agent session as an executable Session Environment. The environment is backed by an append-only Event Log and a sandboxed, persistent Python kernel. The kernel maintains a typed namespace across model calls, allowing tool outputs, retrieved history, and derived state to be bound to variables rather than serialized into the prompt at each call. Model-written code searches, materializes, and transforms session state through exec; only explicitly printed projections enter the model's working view for the next call. Context management thus becomes a programming task that inherits the improving coding abilities of LLMs, while the Event Log preserves lossless historical ground truth. As the working view approaches its budget, stale spans are evicted but remain recoverable: an eviction index keeps compact landmarks tied to exact Event Log addresses, so that the agent navigates directly to evicted regions instead of searching the full log. With Qwen3.8-Max as the backbone, Scroll achieves 94.8% on LongMemEval_S; 73.1% on BEAM_10M, surpassing the best published memory system by 5.1 points; and 86.7% on LOCA_256K, exceeding the best published long-horizon agent by 37.4 points.
Chinese Translation
大型语言模型(LLM)代理越来越多地承担长期运行的任务,其历史记录远远超出单个模型的上下文窗口。现有的方法通过压缩早期交互或将选定信息提取到固定的内存表示中,提前决定在未来需求明确之前需要保留的内容。我们提出了Scroll,这是一种上下文管理器,将每个代理会话视为可执行的会话环境。该环境由一个仅追加的事件日志和一个沙盒的持久Python内核支持。内核在模型调用之间维护一个类型化的命名空间,允许工具输出、检索的历史记录和派生状态绑定到变量,而不是在每次调用时序列化到提示中。因此,上下文管理成为一项编程任务,继承了LLM不断提升的编码能力,同时事件日志保持无损的历史真实记录。当工作视图接近其预算时,过时的范围会被驱逐,但仍然可以恢复:驱逐索引将紧凑的地标与确切的事件日志地址关联,以便代理直接导航到被驱逐的区域,而不是搜索整个日志。在以Qwen3.8-Max为基础的情况下,Scroll在LongMemEval_S上达到了94.8%;在BEAM_10M上达到了73.1%,超过了已发布的最佳内存系统5.1分;在LOCA_256K上达到了86.7%,超过了已发布的最佳长时间任务代理37.4分。
cs.AI / 30 / 2608.21702

From Association to Causation: Improving Retrieval Precision of Retrieval-Augmented Generation via Causal Relations and an Attention Mechanism

从关联到因果:通过因果关系和注意力机制提高检索增强生成的检索精度
Liu, Jing, Qi, Yongxing, Jiang, Muchen, Hu, Chengnan, Peng, Qingqing, Wang, Haoming, Wang, Yuqing, Yu, Yang, Zhang, Xu, Wu, Ting
Abstract
Retrieval-Augmented Generation (RAG) grounds LLM generation on retrieved documents, but the standard terminal retrieval stage--dense-vector similarity, optionally followed by reranking--often returns documents that share keywords with the query without containing the needed information, a failure mode that grows with the knowledge base. We trace it to a conceptual gap: similarity captures only associational relations, whereas the documents that matter are linked to the query causally. We model the terminal retrieval stage with a causal graph grounded in Reichenbach's common cause principle: the keywords shared by the query and a retrieved document form a latent common cause A, and the document's residual keywords form a latent set B linking the document to the ideal output. Since a retrieved document is a collider (A -> d <- B), retrieval itself opens an associational path between the query and B, which licenses a training-free, attention-style re-scoring rule: the cosine similarity between the query embedding and the weighted centroid embedding of B. Unlike causality-enhanced RAG variants that model causal relations inside the knowledge content, our graph models the causal structure of the retrieval process itself. On a real 471-document enterprise knowledge base, the method promotes a relevant guideline from rank 6 to the top 3; on a controlled diagnostic corpus reproducing the keyword-stuffing regime, it improves the mean target rank from 2.88 to 1.25, while a trained cross-encoder reranker barely helps (2.63). Conversely, on three BEIR benchmarks the score underperforms the similarity baseline, delineating the applicability boundary: the method guards the keyword-stuffing regime of growing proprietary knowledge bases and complements neural rerankers; a corpus-level calibration gate selects the correct regime with >= 95% reliability. A fully local testbed demonstrates deployability.
Chinese Translation
检索增强生成(Retrieval-Augmented Generation, RAG)将大规模语言模型(LLM)的生成基于检索到的文档,但标准的终端检索阶段——密集向量相似性,后面可选择重新排序——通常返回与查询共享关键词但不包含所需信息的文档,这种失败模式在知识库增长时愈发明显。我们将其归因于一个概念性差距:相似性仅捕捉关联关系,而重要的文档与查询之间是因果关联。我们用基于赖兴巴赫(Reichenbach)共同原因原则的因果图来建模终端检索阶段:查询与检索文档共享的关键词形成一个潜在的共同原因A,而文档的剩余关键词形成一个潜在的集合B,将文档与理想输出连接起来。由于检索到的文档是一个碰撞体(A -> d <- B),检索本身在查询与B之间打开了一条关联路径,这为一种无训练的、基于注意力的重新评分规则提供了依据:查询嵌入与B的加权质心嵌入之间的余弦相似性。与在知识内容内部建模因果关系的因果增强RAG变体不同,我们的图模型化了检索过程本身的因果结构。在一个包含471个文档的真实企业知识库上,该方法将相关指南的排名从第6位提升至前3位;在一个控制的诊断语料库中,重现了关键词堆砌的机制,平均目标排名从2.88改善至1.25,而经过训练的交叉编码器重新排序器几乎没有帮助(2.63)。相反,在三个BEIR基准测试中,该方法的得分低于相似性基线,划定了适用边界:该方法保护了日益增长的专有知识库的关键词堆砌机制,并补充了神经重新排序器;一个基于语料库的校准门以>= 95%的可靠性选择正确的机制。一个完全本地的测试平台展示了其可部署性。
cs.AI / 31 / 2608.21712

ATHENA: Knowledge-guided agentic neural architecture search for AutoFormer-based electronic health record modeling

ATHENA:基于知识引导的自主神经架构搜索用于基于AutoFormer的电子健康记录建模
Li, Deyi, Xu, Qi, Li, Lingyao, Wang, Tiansheng, Liang, Muxuan, Liu, Mei
Abstract
Transformer-based models are widely used for clinical prediction from electronic health records (EHRs), yet their architectures still require substantial manual tuning, and the optimal configuration may vary across tasks and hospitals. Neural architecture search (NAS) automates architecture design, but conventional methods are computationally costly for Transformer-based EHR models. Recent large language model (LLM)-guided NAS methods reduce manual search design but typically conduct each search independently, without reusing architecture knowledge across hospitals. In this study, we propose ATHENA (Agentic Transfer across Hospitals for EHR Neural Architecture Search), a knowledge-guided agentic NAS framework for Transformer-based EHR modeling. ATHENA uses a weight-sharing supernet that is pretrained once per hospital, allowing candidate architectures to be instantiated as inherited subnetworks and evaluated through fine-tuning rather than independent pretraining. It also incorporates a two-layer cross-hospital architecture prior. The first layer retrieves high-performing architecture examples from source sites based on task descriptors, while the second estimates the effects of architectural components using SHapley Additive exPlanations (SHAP)-based meta-regression. These priors guide a multi-agent LLM search together with validation feedback from the target hospital. Across six clinical prediction tasks and two independent health systems, ATHENA matches or outperforms four NAS baselines in 9 of 12 hospital-task evaluations at a search budget of 30. It also shows more consistent architecture selection across repeated searches. ATHENA provides a practical approach for reducing manual architecture tuning in Transformer-based EHR modeling. Code is publicly available at https://github.com/GatorAIM/ATHENA.
Chinese Translation
基于Transformer的模型广泛应用于电子健康记录(EHR)的临床预测,但其架构仍需进行大量手动调优,且最佳配置可能因任务和医院而异。神经架构搜索(NAS)自动化架构设计,但传统方法对于基于Transformer的EHR模型计算成本较高。最近的基于大型语言模型(LLM)引导的NAS方法减少了手动搜索设计,但通常独立进行每次搜索,未能在医院之间重用架构知识。在本研究中,我们提出了ATHENA(跨医院的EHR神经架构搜索的自主转移),这是一个用于基于Transformer的EHR建模的知识引导自主NAS框架。ATHENA使用一个权重共享的超网络,该网络在每个医院预训练一次,允许候选架构作为继承的子网络实例化,并通过微调进行评估,而不是独立的预训练。它还结合了一个两层的跨医院架构先验。第一层根据任务描述符从源站点检索高性能架构示例,而第二层使用基于SHapley加性解释(SHAP)的元回归估计架构组件的影响。这些先验与目标医院的验证反馈一起引导多智能体LLM搜索。在六个临床预测任务和两个独立健康系统中,ATHENA在30的搜索预算下,在12个医院-任务评估中的9个中与四个NAS基线匹配或超越。它还在重复搜索中显示出更一致的架构选择。ATHENA为减少基于Transformer的EHR建模中的手动架构调优提供了一种实用的方法。代码已公开发布在 https://github.com/GatorAIM/ATHENA。
cs.AI / 32 / 2608.21721

Ask or Answer: A Decision Framework for Multi-Turn Health Misinformation Intervention

提问还是回答:多轮健康错误信息干预的决策框架
Song, Xiaoying, Anik, Anirban Saha, Liu, Jinyu, Tan, Qitao, Yuan, Geng, Hong, Lingzi
Abstract
Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what they know, what they believe, and what they need to hear, so an effective intervention often depends on first asking the right clarifying question. Yet existing methods either respond immediately or probe indiscriminately, treating clarification as either unnecessary or always beneficial. We propose Reward-Optimized Probe-and-Respond (RO-PnR), a framework that learns when asking is worth its cost. At each turn, RO-PnR chooses between probing for more information and committing to a final correction, guided by a turn-level reward that weighs the expected gain from probing against its interaction cost. To capture how user heterogeneity affects probing value, we model each simulated user with a latent state along health literacy and belief commitment. Experiments show that RO-PnR achieves the highest cost-adjusted utility across three health-misinformation datasets and three base models, using 30% fewer turns than always-probe baselines.
Chinese Translation
在对话中纠正健康错误信息不仅仅需要提供事实反驳:用户在知识、信念和需要听到的信息上存在差异,因此有效的干预往往依赖于首先提出正确的澄清问题。然而,现有方法要么立即回应,要么无差别地探询,将澄清视为不必要或总是有益。我们提出了奖励优化探询与回应(Reward-Optimized Probe-and-Respond, RO-PnR)框架,该框架学习何时提问是值得的。在每个回合中,RO-PnR在探询更多信息和承诺最终纠正之间进行选择,依据一个回合级奖励,该奖励权衡探询的预期收益与其交互成本。为了捕捉用户异质性如何影响探询价值,我们将每个模拟用户建模为在健康素养和信念承诺上的潜在状态。实验表明,RO-PnR在三个健康错误信息数据集和三个基础模型上实现了最高的成本调整效用,使用的回合数比始终探询的基线少30%。
cs.AI / 33 / 2608.21755

ECHO: A Cognitively Inspired, Auditable Memory Plane for Long-Horizon Agents

ECHO:一种受认知启发的可审计记忆平面,用于长时间跨度的智能体
Qian, Yu, Miao, Hong, Guo, Boyang, Jiang, Tingyi, Zhao, Shan, Le, Tianxing, Li, Lintian, Liu, Meng
Abstract
Long-horizon agents need memory that identifies relevant experience, resolves revisions, and exposes checkable provenance. We present ECHO (Embodied Context and History Orchestration), an auditable memory architecture and service prototype inspired by episodic encoding, consolidation, contextual reinstatement, reconsolidation, and executive control. This is functional inspiration, not neural equivalence; the empirical analysis focuses on retrieval and context construction. Development runs reach 96.29% Hit@10 and 73.64% turn Recall@5 on 1,536 LoCoMo category 1-4 questions, and 97.60% Hit@10, 88.84% turn Recall@5, and 88.71% session Recall@5 on all 500 LongMemEval-S questions. A five-history BEAM gate fails, and in a separate matched 91-question QA sample Mem0 OSS scores 64.84% versus ECHO's 41.76% (exact McNemar p = 0.00107), with a history-cluster interval crossing zero. A post-hoc audit found source-specific phrases in the query-expansion rules. Although no gold answer field entered the runtime, expansion-enabled retrieval scores are therefore descriptive development measurements, not independent confirmation.
Chinese Translation
长时间跨度的智能体需要能够识别相关经验、解决修订并揭示可检查来源的记忆。我们提出了ECHO(Embodied Context and History Orchestration),这是一种受情节编码、巩固、情境重现、再巩固和执行控制启发的可审计记忆架构和服务原型。这是一种功能上的启发,而非神经等价;实证分析集中于检索和情境构建。开发运行在1,536个LoCoMo类别1-4问题上达到了96.29%的Hit@10和73.64%的turn Recall@5,在所有500个LongMemEval-S问题上达到了97.60%的Hit@10、88.84%的turn Recall@5和88.71%的session Recall@5。一个五历史BEAM门失败,在一个单独匹配的91问题QA样本中,Mem0 OSS得分为64.84%,而ECHO为41.76%(精确的McNemar p = 0.00107),历史聚类间隔跨越零。事后审计发现查询扩展规则中存在特定来源的短语。尽管没有黄金答案字段进入运行时,扩展启用的检索得分因此是描述性的发展测量,而非独立确认。
cs.AI / 34 / 2608.21761

What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation

CLIP在区域地理定位中学习了什么?适应后的视觉线索和场景配置探讨
Lee, Changyu, Park, Yeonsoo, Alfarrarjeh, Abdullah, Kim, Seon Ho
Abstract
Large collections of street-view imagery provide rich visual information about urban environments, but extracting fine-grained geographic information from such data remains challenging. In particular, fine-grained regional geolocalization is challenging because nearby areas often share coarse geographic cues. We study regional geolocalization within a metropolitan area and ask whether pretrained CLIP features are sufficient for regional discrimination, and what visual information supports performance after adaptation. Using 9,085 street-view images from eight Greater Los Angeles regions, we compare zero-shot CLIP, frozen-encoder readouts, partial encoder updating, Low-Rank Adaptation (LoRA), and full fine-tuning. Frozen readouts remain near the 39.03% zero-shot accuracy, whereas encoder adaptation achieves 75.94-82.10%. Full fine-tuning also reduces the mean distance to the predicted region center from 12.30 km to 3.86 km. We probe these gains through semantic cue removal, appearance reduction using edge maps and blur, and scene-configuration disruption using patch scrambling. Adapted models achieve higher edge and blur accuracy and switch 42.92-45.56% of predictions after scrambling, compared with 10.79-14.60% for frozen methods. However, adaptation does not improve the fraction of performance retained after appearance reduction, while vegetation and sky remain influential. A Caltech101 control further shows that scrambling sensitivity is not unique to geolocalization. Overall, encoder adaptation substantially improves nearby-region discrimination and is associated with greater sensitivity to intact scene configuration, without evidence that coarse structure alone becomes sufficient for prediction. These conclusions concern viewpoint variation near known locations rather than geographically disjoint generalization.
Chinese Translation
大量街景图像提供了关于城市环境的丰富视觉信息,但从这些数据中提取细粒度的地理信息仍然具有挑战性。特别是,细粒度的区域地理定位面临挑战,因为附近区域往往共享粗略的地理线索。我们研究了大都市区域内的区域地理定位,并询问预训练的CLIP特征是否足以进行区域区分,以及哪些视觉信息在适应后支持性能。使用来自八个大洛杉矶地区的9,085张街景图像,我们比较了零样本CLIP、冻结编码器输出、部分编码器更新、低秩适应(Low-Rank Adaptation, LoRA)和完全微调。冻结输出的准确率接近39.03%的零样本准确率,而编码器适应的准确率达到75.94-82.10%。完全微调还将预测区域中心的平均距离从12.30公里减少到3.86公里。我们通过语义线索移除、使用边缘图和模糊进行外观减少,以及使用补丁打乱进行场景配置干扰来探讨这些提升。与冻结方法的10.79-14.60%相比,适应模型在打乱后实现了42.92-45.56%的预测切换,且在边缘和模糊准确性上更高。然而,适应并未改善外观减少后保留的性能比例,而植被和天空仍然具有影响力。Caltech101的对照实验进一步表明,打乱敏感性并非地理定位所独有。总体而言,编码器适应显著改善了邻近区域的区分能力,并与对完整场景配置的更大敏感性相关,而没有证据表明仅凭粗略结构就足以进行预测。这些结论关注的是已知位置附近的视角变化,而非地理上不相交的泛化。
cs.AI / 35 / 2608.21767

Physics-Knowledge-Guided Hybrid Neural Learning for Arctic Sea Ice Concentration Evolution and Short-Range Prediction

基于物理知识指导的混合神经学习用于北极海冰浓度演变与短期预测
Zhang, Maqun, Gao, Feng, Chen, Wankun, Yu, Hui, Gan, Yanhai, Dong, Junyu
Abstract
Accurate modeling of sea ice concentration (SIC) evolution is essential for polar climate assessment and short?range sea ice prediction. Numerical and data-driven approaches constitute major foundations for SIC modeling, but the former often require complex parameterizations and substantial compu?tation, whereas the latter rarely encode physical dependencies explicitly. This study presents the Physics-Informed Hybrid Ice Model (PIHIM), a differentiable data-driven hybrid ice model for daily SIC evolution that organizes its network structure according to the physical dependencies encoded in the sea ice continuity equation and explicitly accounts for dynamical transport, ther?modynamically driven areal growth and loss, and unresolved local processes. PIHIM preserves the representation capacity of deep learning while providing a process-decomposed formulation of ice displacement, freeze-melt areal change, and local error closure. Two evaluation settings are adopted: reanalysis-forced simulation examines SIC evolution stability under reanalysis forcing, and forecast-forced prediction assesses short-range performance un?der forecast-forced conditions, with reanalysis and observational SIC serving as verification references. Results indicate enhanced ice-edge preservation and error-growth control in reanalysis?forced simulation, while PIHIM retains measurable short-range prediction skill under forecast-forced conditions. Our code will be made publicly available after the paper is accepted.
Chinese Translation
准确建模海冰浓度(SIC)演变对于极地气候评估和短期海冰预测至关重要。数值方法和数据驱动方法构成了SIC建模的主要基础,但前者通常需要复杂的参数化和大量计算,而后者则很少明确编码物理依赖关系。本研究提出了物理信息混合冰模型(Physics-Informed Hybrid Ice Model, PIHIM),这是一种可微分的数据驱动混合冰模型,用于日常SIC演变,其网络结构根据海冰连续性方程中编码的物理依赖关系进行组织,并明确考虑了动力传输、热力驱动的面积增长和损失以及未解决的局部过程。PIHIM保留了深度学习的表示能力,同时提供了冰位移、冻结-融化面积变化和局部误差闭合的过程分解公式。采用了两种评估设置:重分析强迫模拟考察在重分析强迫下SIC演变的稳定性,而预测强迫预测则评估在预测强迫条件下的短期性能,重分析和观测SIC作为验证参考。结果表明,在重分析强迫模拟中增强了冰缘的保持和误差增长控制,而PIHIM在预测强迫条件下保持了可测量的短期预测能力。我们的代码将在论文被接受后公开发布。
cs.AI / 36 / 2608.21792

HIRA: A Human-in-the-Loop Retrieval-Augmented Cascade for Document Classification in Regulated Industries

HIRA:一种人机协作的检索增强级联模型用于受监管行业的文档分类
Tian, Shangxuan, Chen, Yanhui, Queiroz, Carlos
Abstract
Document classification in regulated industries is constrained by data residency, limited cold-start labels, scarce review capacity, and costly model-governance procedures. We present HIRA, a training-free, on-premises retrieval-augmented cascade for document classification in regulated deployments that combines BM25 over OCR text, dense text embeddings, and image-level representations through validation-calibrated weighted reciprocal-rank fusion. Confident documents are classified directly by retrieval; uncertain or visually confusable documents are passed to a locally hosted LLM verifier, which receives the OCR text, retrieved exemplars, label descriptions, and confusion-specific terms. When the verifier remains uncertain, the document is sent to human review. Each correction is stored as a margin-weighted retrieval exemplar and updates a Dirichlet-smoothed confusion graph, letting the system improve without updating model weights. On a private 80-class trade-finance corpus, HIRA processes the full 30,233-document production stream while requesting human correction for only 1,945 documents (6.4%), improving Macro-F1 from 0.6218 to 0.8548. On the corrected Tobacco-3482 benchmark, HIRA reaches 0.9423 Macro-F1 with a locally hosted DeepSeek-R1-Distill-Qwen-32B verifier, 17.4 percentage points above the zero-shot LLM baseline, while invoking the verifier for only about 40% of documents and reducing LLM calls by approximately 60%. With 518 human corrections (24.8% of the pool), HIRA matches the fully labelled pool oracle, in which all 2,086 pool documents are indexed with their ground-truth labels. These results show that selective human feedback and retrieval-memory adaptation can be a practical alternative to repeated model retraining for long-tail document classification in regulated deployments.
Chinese Translation
受监管行业的文档分类受到数据驻留、有限的冷启动标签、稀缺的审核能力以及高昂的模型治理程序的限制。我们提出了HIRA,这是一种无训练、在本地部署的检索增强级联模型,旨在受监管环境下的文档分类。该模型结合了基于OCR文本的BM25检索、密集文本嵌入和通过验证校准的加权倒排融合的图像级表示。对于置信度高的文档,直接通过检索进行分类;而对于不确定或视觉上容易混淆的文档,则交由本地托管的LLM(大语言模型)验证器处理,该验证器接收OCR文本、检索到的示例、标签描述和特定混淆术语。当验证器仍然不确定时,文档将被送往人工审核。每次修正都被存储为边际加权的检索示例,并更新一个Dirichlet平滑的混淆图,从而使系统在不更新模型权重的情况下不断改进。在一个包含80个类别的私有贸易金融语料库中,HIRA处理了全量的30,233份文档生产流,仅请求人工修正1,945份文档(占6.4%),使得宏观F1值从0.6218提升至0.8548。在经过修正的Tobacco-3482基准测试中,HIRA在本地托管的DeepSeek-R1-Distill-Qwen-32B验证器下达到了0.9423的宏观F1值,比零样本LLM基线高出17.4个百分点,同时仅对约40%的文档调用验证器,并将LLM调用减少了约60%。在518次人工修正(占池中的24.8%)后,HIRA的表现与完全标注的池oracle相匹配,其中所有2,086份池文档都已按其真实标签进行索引。这些结果表明,选择性的人类反馈和检索记忆适应可以成为受监管环境中长尾文档分类的重复模型再训练的实用替代方案。
cs.AI / 37 / 2608.21811

Hints, Critics, and Teachers: Prior Injection for Sparse-Reward RL in Vision-Language Math Reasoning

提示、批评者与教师:稀疏奖励下视觉语言数学推理的先验注入
Fu, Qiqian
Abstract
Reinforcement learning for vision-language math reasoning starves under sparse reward: on a pool of 20,830 visual-math problems where Qwen2-VL-2B answers 3.6% of rollouts correctly, 85-97% of GRPO rollout groups are entirely wrong and contribute zero gradient. We train eleven methods under identical conditions in this regime, each injecting a different prior: text (reference-solution hints), distribution (on-policy distillation from a 7B teacher), and value (a value-pretrained critic with an MSE or HL-Gauss categorical loss). A prior helps exactly when it is delivered: the six arms whose prior effectively reaches the policy separate with no overlap from the remaining five -- the no-prior baseline and four arms whose prior is teacher-capped, gated away, or lost to a mis-parameterized critic -- both on the pooled in-domain metric and on cross-domain transfer (DynaMath). The central finding, however, concerns evaluation: one slice of the in-domain pool -- long used as this project's general-distribution check -- anti-correlates with genuine cross-domain transfer (Spearman rho = -0.74, n = 11 arms, permutation p = 0.011), while the hardest in-domain slice predicts it closely (rho = +0.89, p < 0.001). We attribute the inversion to a near-chance multiple-choice subset that rewards models for not having changed; read through it, the best cross-domain method looked mediocre and the worst looked like the champion. Among the methods, hint-guided exploration -- not UFT's auxiliary loss -- drives hint gains, and replacing the critic's MSE loss with HL-Gauss cross-entropy is worth +14.4 points in-domain. All accuracies are blind-judged, with paired exact tests.
Chinese Translation
在稀疏奖励下,视觉语言数学推理的强化学习面临困境:在20,830个视觉数学问题的池中,Qwen2-VL-2B仅正确回答3.6%的回合,85-97%的GRPO回合组完全错误,未贡献任何梯度。我们在相同条件下训练了十一种方法,每种方法注入不同的先验:文本(参考解提示)、分布(来自7B教师的在线蒸馏)和价值(使用均方误差或HL-Gauss分类损失的价值预训练批评者)。先验的有效性取决于其传递时机:六种有效触达策略的臂与剩余的五种(无先验基线和四种先验被教师限制、被门控或被错误参数化的批评者所丢失的臂)在聚合的领域内指标和跨领域迁移(DynaMath)上完全分离。然而,核心发现涉及评估:领域内池中的一部分——长期作为该项目一般分布检查的部分——与真实的跨领域迁移呈负相关(Spearman rho = -0.74, n = 11臂, permutation p = 0.011),而最困难的领域内部分则与其密切相关(rho = +0.89, p < 0.001)。我们将这种反转归因于一个近乎随机的多项选择子集,该子集奖励模型未发生变化;通过它观察,最佳的跨领域方法看起来平庸,而最差的则像冠军。在这些方法中,提示引导的探索——而非UFT的辅助损失——驱动了提示的收益,而将批评者的均方误差损失替换为HL-Gauss交叉熵在领域内价值提升了14.4分。所有准确率均为盲评,采用配对精确检验。
cs.AI / 38 / 2608.21825

VisAdj: Learning Adjacency Matrices from Node-Link Images

VisAdj:从节点-链接图像中学习邻接矩阵
Xie, Jiahao, Tong, Guangmo
Abstract
Learning adjacency matrices from node-link images is a fundamental problem for recovering structured graph information from visual observations. Existing methods typically rely on fixed KNN-based heuristics for candidate edge selection and fail to capture dependencies among edges. To overcome these limitations, we propose VisAdj, a new framework for topology-aware adjacency prediction. VisAdj introduces an attention-sparse neighbor sampler to adaptively select a high-recall set of candidate node pairs and performs joint edge inference using a line-graph transformer that treats candidate edges as tokens and explicitly models dependencies among incident edges. Extensive experiments on synthetic graphs, road networks, and vessel images demonstrate that VisAdj consistently outperforms existing baselines by clear margins.
Chinese Translation
从节点-链接图像中学习邻接矩阵是从视觉观察中恢复结构化图信息的一个基本问题。现有的方法通常依赖于固定的基于 KNN 的启发式方法进行候选边的选择,未能捕捉边之间的依赖关系。为克服这些局限性,我们提出了 VisAdj,一个新的拓扑感知邻接预测框架。VisAdj 引入了一种注意力稀疏邻居采样器,以自适应选择高召回率的候选节点对,并使用线图变换器进行联合边推断,该变换器将候选边视为标记,并明确建模事件边之间的依赖关系。在合成图、道路网络和血管图像上的大量实验表明,VisAdj 在性能上始终明显优于现有基线。
cs.AI / 39 / 2608.21830

Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents

超越成功与失败:面向图形用户界面代理的长度感知对比学习
Gu, Chengyang, Zhang, Le, Zhou, Jingbo, Chen, Yize, Shi, Yu, Bao, Siqi, Wu, Zheng-Fan, Wu, Hua, Xiong, Hui
Abstract
Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) have shown strong potential for automating tasks across diverse digital environments, where reinforcement learning (RL) has become a dominant training paradigm. However, widely used methods such as Group Relative Policy Optimization (GRPO) suffer from reward-gradient misalignment, leading to inefficient and unstable optimization. Recent work addresses this issue by reformulating RL with verifiable rewards (RLVR) as contrastive or classification-based objectives, which improve stability by eliminating problematic gradient behaviors. Despite this progress, existing contrastive RLVR methods rely primarily on outcome-level supervision and fail to capture fine-grained differences in trajectory quality within the same outcome category. In this paper, we propose Length-Aware Contrastive Learning for GUI Agents (LACL-GUI), a contrastive RLVR framework that incorporates trajectory-level quality signals into policy optimization. LACL-GUI introduces structured preferences within both successful and failed trajectories, encouraging concise successful executions and differentiating failure quality based on divergence from successful trajectories, while preserving optimization stability. Experiments on GUI agent benchmarks show that LACL-GUI provides more effective learning signals and consistently improves agent performance over prior methods, highlighting the value of trajectory-level supervision in contrastive RLVR.
Chinese Translation
由多模态大型语言模型(MLLMs)驱动的图形用户界面(GUI)代理在自动化多种数字环境中的任务方面展现出强大的潜力,其中强化学习(RL)已成为主流的训练范式。然而,广泛使用的方法如群体相对策略优化(GRPO)存在奖励梯度不对齐的问题,导致优化效率低下且不稳定。近期的研究通过将强化学习重新构造为可验证奖励(RLVR),并采用对比或分类为基础的目标来解决这一问题,从而通过消除有问题的梯度行为来提高稳定性。尽管取得了这些进展,现有的对比RLVR方法主要依赖于结果级别的监督,未能捕捉同一结果类别内轨迹质量的细微差异。本文提出了一种面向图形用户界面代理的长度感知对比学习框架(LACL-GUI),该框架将轨迹级别的质量信号纳入策略优化。LACL-GUI在成功和失败的轨迹中引入结构化偏好,鼓励简洁的成功执行,并根据与成功轨迹的偏离程度区分失败质量,同时保持优化的稳定性。在GUI代理基准测试中的实验表明,LACL-GUI提供了更有效的学习信号,并在性能上持续优于先前的方法,突显了轨迹级别监督在对比RLVR中的价值。
cs.AI / 40 / 2608.21833

GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?

GameXpert-Bench:编码智能体距离专家游戏开发还有多远?
Chen, Kun, Hong, Haorong, Gao, Peizhong, Lin, Jianfeng, Luo, Tongxu, Xie, Yuxuan, Liu, Chenxu, He, Jieling, Liu, Zhongyuan, Zeng, Zeno
Abstract
Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especially demanding because program logic, visual and audio content, interfaces, interaction and playability must function together in one executable artifact. Measuring this capability therefore requires evaluation of both game product and the development process. Existing benchmarks often assess the game development capabilities of LLMs by evaluating the final artifact or an isolated development stage. Our analysis of complete human-agent development trajectories identifies three stages that together span the lifecycle of game development with a coding agent: initial game generation, bug diagnosis and repair, and optimization over multiple turns. Therefore, we introduce GameXpert-Bench, which operationalizes the three lifecycle stages as three complementary benchmark tracks. GameGen evaluates complete game creation from a single request in an empty workspace. GameFix evaluates diagnosis and repair when defects are reported or left for the agent to discover. GameOpt evaluates cumulative optimization through request chains seeded by real development trajectories between users and agents. We evaluate each track using live game interaction, deterministic behavioral tests, or final product criteria with regression checks. The suite contains 97 generation tasks across 11 genres; 100 repair tasks from 50 game levels verified by humans, each with 19-27 injected bugs; and 17 optimization chains with six turns and 102 requests. Across the three tracks, current agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, and preserving functionality across changes.
Chinese Translation
最近的大型语言模型(LLMs)能够作为编码智能体,根据自然语言请求构建完整的游戏。游戏开发尤其具有挑战性,因为程序逻辑、视觉和音频内容、界面、交互和可玩性必须在一个可执行的工件中协同工作。因此,衡量这一能力需要对游戏产品和开发过程进行评估。现有基准通常通过评估最终工件或孤立的开发阶段来评估LLMs的游戏开发能力。我们对完整的人机开发轨迹的分析识别了三个阶段,这三个阶段共同涵盖了编码智能体的游戏开发生命周期:初始游戏生成、缺陷诊断与修复,以及多轮优化。因此,我们引入了GameXpert-Bench,将这三个生命周期阶段操作化为三个互补的基准轨道。GameGen评估从空工作区的单一请求生成完整游戏的能力。GameFix评估在报告缺陷或留给智能体发现时的诊断与修复能力。GameOpt评估通过用户与智能体之间真实开发轨迹引发的请求链进行的累积优化。我们使用实时游戏交互、确定性行为测试或最终产品标准与回归检查来评估每个轨道。该套件包含跨11个类型的97个生成任务;来自50个游戏关卡的100个修复任务,经过人工验证,每个关卡注入了19-27个缺陷;以及17个优化链,包含六轮和102个请求。在这三个轨道中,目前的智能体在生成可玩基础和实现明确要求方面比在发现缺陷、验证运行时行为和在变更中保持功能性方面更可靠。
cs.AI / 41 / 2608.21836

LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization

LLM4LLM:通过闭环智能优化连接内核基准测试与实际部署
Zeng, Hui, Yang, Pengfei, Chen, Yanxin, Ju, Fusong, Wei, Xinran
Abstract
Large language models have become increasingly capable agents for low-level code and kernel optimization, but isolated kernel benchmarks provide only a proxy for the deployment behavior that matters in language-model inference. We identify a benchmark-to-deployment gap: candidate kernels that appear correct and fast in standalone harnesses can exhibit different performance, safety, or phase behavior after integration into a real inference workload. We introduce LLM4LLM, a deployment-aware closed-loop optimization framework that starts from a target inference script, extracts phase-aware optimization tasks, searches with an experience-guided episodic agent, and accepts patches through in-model validation. Across ten language-model inference workloads on A100 and H100 GPUs, LLM4LLM improves end-to-end latency for every evaluated model, achieving 3.91$\times$/6.98$\times$ geometric-mean speedups on A100/H100; as supporting kernel-level evidence, it also attains up to 2.745$\times$ GeoMean speedup on KernelBench Level 2.
Chinese Translation
大型语言模型已成为低级代码和内核优化日益强大的代理,但孤立的内核基准测试仅提供了与语言模型推理中重要的部署行为的代理。我们识别出基准与部署之间的差距:在独立测试环境中看似正确且快速的候选内核,在集成到实际推理工作负载后可能表现出不同的性能、安全性或阶段行为。我们提出了LLM4LLM,这是一种以部署为导向的闭环优化框架,起始于目标推理脚本,提取阶段感知的优化任务,通过经验指导的情节代理进行搜索,并通过模型内验证接受补丁。在A100和H100 GPU上对十个语言模型推理工作负载的测试中,LLM4LLM提高了每个评估模型的端到端延迟,在A100/H100上实现了3.91$ imes$/6.98$ imes$的几何平均加速;作为支持的内核级证据,它在KernelBench Level 2上也达到了最高2.745$ imes$的几何平均加速。
cs.AI / 42 / 2608.21841

AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations

AI 监督者:用于检测和防御 AI 对话中操控性黑暗模式的代理接口
Poonsiriwong, Rachel, Chayapatr, Archiwaranguprok, Albrecht, Constanze, Lertsutthiwong, Monchai, Maes, Pattie, Pataranutaporn, Pat
Abstract
Conversational AI increasingly shapes consequential decisions, yet users have limited support for recognizing and resisting manipulation. We present AI Watchdog, a browser-based agent interface that monitors live conversations, detects five dark-pattern categories, including sycophancy, brand bias, anthropomorphization, sneaking, and harmful generation, and alerts users when they occur. Its open-weight turn-level classifier supports independent deployment and a path toward local inference, preserving user privacy while remaining separate from the conversational AI. We evaluated AI Watchdog in a preregistered, five-condition between-subjects experiment (N = 150) comparing a no-intervention control with four configurations varying nudge timing (prebunking vs. just-in-time) and engagement mode (without vs. with cognitive forcing). Results show that participants rarely flagged manipulative turns across all conditions, and post-task awareness did not differ significantly across groups. However, just-in-time warnings without cognitive forcing were the only intervention to significantly reduce compliance with AI-steered recommendations containing dark patterns, lowering compliance from 71.7% to 53.7%, an 18 percentage-point reduction. Exploratory analyses further showed that lower misinformation susceptibility was associated with greater flagging but not lower compliance, while higher AI trust was associated with greater compliance and lower reported awareness. Together, these findings suggest that explicit recognition of conversational dark patterns and behavioral resistance to AI steering may be distinct outcomes, motivating further investigation of timely, low-friction defensive interfaces.
Chinese Translation
对话式 AI 日益影响重要决策,但用户在识别和抵制操控方面的支持有限。我们提出了 AI 监督者(AI Watchdog),这是一种基于浏览器的代理接口,能够监控实时对话,检测五种黑暗模式类别,包括谄媚、品牌偏见、人性化、潜入式操控和有害生成,并在这些模式出现时提醒用户。其开放权重的回合级分类器支持独立部署,并为本地推理提供了一条路径,保护用户隐私,同时与对话式 AI 保持分离。我们在一项预注册的五条件组间实验中评估了 AI 监督者(N = 150),比较了无干预的对照组与四种配置,变化包括干预时机(预先揭露与及时干预)和参与模式(无认知强制与有认知强制)。结果显示,参与者在所有条件下很少标记操控性回合,任务后的意识在各组之间没有显著差异。然而,只有在没有认知强制的情况下,及时警告是唯一显著降低对包含黑暗模式的 AI 引导推荐的遵从率的干预,将遵从率从 71.7% 降低至 53.7%,减少了 18 个百分点。探索性分析进一步显示,较低的错误信息易感性与更高的标记率相关,但与较低的遵从率无关,而较高的 AI 信任度则与更高的遵从率和较低的报告意识相关。综合来看,这些发现表明,对话黑暗模式的明确识别与对 AI 引导的行为抵制可能是不同的结果,激励进一步研究及时、低摩擦的防御接口。
cs.AI / 43 / 2608.21867

MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governance

MemGuard:持久化验证器信号用于大型语言模型代理的记忆治理
Wang, Haoyu, Dong, Guangyuan, Liang, He, Zhang, Zijing, Luo, Jiachen, Liu, Chuang, Xue, Chao, Tang, Hao
Abstract
LLM agents are moving from single-prompt use to long task streams in which reusable memory becomes a core capability for terminal, software-engineering, and web tasks. Such memory is useful only when stored experience remains reliable across hundreds of interactions, but two failure modes break that assumption in practice. The first is unreliable admission: failed trajectories,accidental successes, and misleading observations enter memory because they appear relevant, then mislead later decisions. The second is memory drift: long-running banks accumulate duplicate, stale, and conflicting records that retrieval alone cannot repair. MemGuard's key distinction is to treat verifier output not as a one-shot filter, but as persistent lifecycle metadata. It converts multi-criteria score-token verification into reward, confidence, label, and uncertainty descriptors that are attached to every candidate before activation and reused during retrieval, conflict resolution, summarization, and archival. We evaluate MemGuard on Terminal-Bench 2.0, SWE-Bench Verified, WebArena, and Mind2Web across four backbones, comparing against four memory baselines plus a verifier-only control under matched runtime budgets. Averaged over five seeds, MemGuard achieves the best success metric and lowest average steps in all 16 backbone-benchmark settings, improving over ReasoningBank, the strongest prior baseline among the memory methods we evaluate, with a largest gain of 7.9 success-rate points on WebArena, 5.6 step-success-rate points on Mind2Web, and 2.4-3.5 points on terminal and software-engineering benchmarks. Code is available at https://github.com/whyyyyy123/MemGuard.
Chinese Translation
大型语言模型代理正从单次提示使用转向长任务流,其中可重用的记忆成为终端、软件工程和网络任务的核心能力。这种记忆只有在存储的经验在数百次交互中保持可靠时才有用,但在实践中有两种失败模式破坏了这一假设。第一种是可靠性不足的接纳:失败的轨迹、意外的成功和误导性的观察因看似相关而进入记忆,进而误导后续决策。第二种是记忆漂移:长时间运行的记忆库积累重复、过时和冲突的记录,仅靠检索无法修复。MemGuard的关键区别在于将验证器输出视为持久的生命周期元数据,而非一次性过滤器。它将多标准评分令牌验证转换为奖励、置信度、标签和不确定性描述符,这些描述符在激活之前附加到每个候选项上,并在检索、冲突解决、摘要和归档过程中重复使用。我们在Terminal-Bench 2.0、SWE-Bench Verified、WebArena和Mind2Web上评估MemGuard,比较四个基础模型,外加一个仅验证器控制组,且在匹配的运行预算下进行评估。经过五次种子实验的平均,MemGuard在所有16个基础模型基准设置中实现了最佳成功指标和最低平均步骤,相较于我们评估的记忆方法中最强的先前基线ReasoningBank,WebArena上成功率提高了7.9个百分点,Mind2Web上提高了5.6个百分点,终端和软件工程基准上提高了2.4-3.5个百分点。代码可在 https://github.com/whyyyyy123/MemGuard 获取。
cs.AI / 44 / 2608.21868

HiMA-MDD: A Hierarchical Multi-Agent Harness for Interpretable Multimodal Depression Detection in Clinical Interviews

HiMA-MDD:用于临床访谈中可解释的多模态抑郁症检测的分层多智能体框架
Chen, Ao, Peng, Xiaojiang
Abstract
Depression assessment from multimodal clinical interviews requires integrating dispersed evidence from multiple symptoms into a coherent PHQ-8 profile. This process is hierarchical: relevant evidence is often sparse and context-dependent within local question-answer exchanges, multiple exchanges jointly support symptom-level judgments, and the final assessment depends on the coherence of the complete symptom profile. Existing LLM systems either process interviews holistically or distribute work across generic agent roles; neither design necessarily provides an explicit orchestration mechanism that coordinates evidence access, item-score authority, bounded feedback, and state recording across these levels. To address this gap, we introduce HiMA-MDD, a hierarchical multi-agent harness that aligns this assessment hierarchy with three agent layers. After non-agentic preprocessing constructs context-preserving multimodal QA units, Layer 1 identifies candidate QA-to-item relations and supports bounded item-grounded evidence routing. Layer 2 assigns symptom groups to operational factor specialists, with one specialist responsible for each provisional item score. Layer 3 audits the complete provisional profile, requests at most one round of targeted revision, and reconstructs the verified PHQ-8 profile. This layered design naturally yields a Hierarchical Evidence Trace, preserves all intermediate evidence, judgments, and revisions for auditability. The final item scores then deterministically produce the total score and screening decision. Using Qwen2.5-72B-Instruct as the harness backbone, our experiments on E-DAIC demonstrate that HiMA-MDD outperforms the compared state-of-the-art methods.
Chinese Translation
从多模态临床访谈中评估抑郁症需要将来自多个症状的分散证据整合成一个连贯的PHQ-8档案。这个过程是分层的:相关证据通常在局部问答交流中稀疏且依赖于上下文,多次交流共同支持症状级别的判断,最终评估依赖于完整症状档案的一致性。现有的大型语言模型(LLM)系统要么整体处理访谈,要么在通用智能体角色之间分配工作;这两种设计都未必提供一个明确的协调机制,以在这些层级之间协调证据访问、项目评分权威、有限反馈和状态记录。为了解决这一问题,我们提出了HiMA-MDD,一个分层多智能体框架,将这一评估层次结构与三个智能体层对齐。在非智能体预处理构建保持上下文的多模态问答单元后,第一层识别候选的问答到项目关系,并支持有限的项目基础证据路由。第二层将症状组分配给操作因素专家,每个专家负责一个临时项目评分。第三层审核完整的临时档案,最多请求一次针对性的修订,并重建经过验证的PHQ-8档案。这种分层设计自然产生了分层证据追踪,保留所有中间证据、判断和修订以便审计。最终的项目评分决定性地产生总分和筛查决策。使用Qwen2.5-72B-Instruct作为框架基础,我们在E-DAIC上的实验表明,HiMA-MDD的表现优于比较的最先进方法。
cs.AI / 45 / 2608.21897

From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning

从求解器反馈到可信计划:用于符号规划的多角色强化学习
Zhang, Chenghao, Mao, Yikai, Liu, Shanqi, Gao, Haoyu, Hu, SaiSai, Roth, Dan
Abstract
Reliable planning requires converting natural-language instructions into executable symbolic specifications, yet large language models remain brittle without costly PDDL annotations and may exploit solver success in semantically unfaithful ways. We study how to learn faithful natural-language-to-PDDL formalization using only solver feedback, without human-written demonstrations. We propose a solvergrounded multi-role reinforcement learning framework where a single language model acts as an Actor, Judge, and Editor for generation, verification, and repair. The Actor proposes PDDL specifications, the Judge provides a solver-calibrated quality signal, and the Editor performs bounded diagnostic-conditioned refinement. On PlanBench, our method improves average success from 35.5% for LLM+P to 70.8%, achieves 66.3% faithful success, and reduces semantic drift to 6.4%. These results show that organizing solver feedback into generation, verification, and repair roles enables more scalable and faithful annotation-free symbolic planning
Chinese Translation
可靠的规划需要将自然语言指令转换为可执行的符号规范,但大型语言模型在没有昂贵的PDDL注释的情况下仍然脆弱,并且可能以语义不忠实的方式利用求解器的成功。我们研究如何仅使用求解器反馈学习可信的自然语言到PDDL的形式化,而无需人工编写的示范。我们提出了一种基于求解器的多角色强化学习框架,其中单一语言模型充当生成、验证和修复的演员、评判者和编辑。演员提出PDDL规范,评判者提供经过求解器校准的质量信号,编辑进行有限的诊断条件修正。在PlanBench上,我们的方法将平均成功率从35.5%(LLM+P)提高到70.8%,实现了66.3%的可信成功率,并将语义漂移降低到6.4%。这些结果表明,将求解器反馈组织为生成、验证和修复角色能够实现更具可扩展性和可信度的无注释符号规划。
cs.AI / 46 / 2608.21898

Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning

训练可信的世界:用于智能体学习的经过验证的合成网络环境
Zhang, Chenghao, Xiao, Canran, Hu, SaiSai, Roth, Dan
Abstract
Web agents promise to automate complex digital workflows, but their training remains limited by synthetic environments that look plausible while hiding broken links, inconsistent states, or infeasible tasks. We address the gap between scalable environment generation and trustworthy agent learning by constructing synthetic web environments that are executable, auditable, and grounded in backend state. Our framework represents each generated website as a structured scaffold of pages, navigation links, database records, state-change markers, and task constraints, then verifies and repairs structural, semantic, consistency, and feasibility defects before policy training. During interaction, ordinary UI transitions are executed deterministically, while persistent backend updates are invoked only through validated state-change markers, enabling dense rewards compiled from verified task-progress predicates. Across 500 synthetic environments spanning six domains, our method reduces task-blocking defects and improves feasible-task rate from 48.6% to 94.8%, while producing stronger PPO policies and improving transfer to WebArena, WebShop, and MiniWoB++ without LLM calls at evaluation time. These results show that verified synthetic environments can serve as a scalable and reliable training substrate for compact web agents, shifting synthetic webagent learning from surface-level plausibility toward executable, state-grounded supervision.
Chinese Translation
网络智能体承诺自动化复杂的数字工作流程,但它们的训练仍然受到合成环境的限制,这些环境看似合理,却隐藏着断链、不一致状态或不可行任务。我们通过构建可执行、可审计且基于后端状态的合成网络环境,解决了可扩展环境生成与可信智能体学习之间的差距。我们的框架将每个生成的网站表示为一个结构化的页面、导航链接、数据库记录、状态变化标记和任务约束的支架,然后在策略训练之前验证并修复结构、语义、一致性和可行性缺陷。在交互过程中,普通的用户界面过渡是确定性执行的,而持久的后端更新仅通过经过验证的状态变化标记触发,从而实现从经过验证的任务进展谓词编译的密集奖励。在跨越六个领域的500个合成环境中,我们的方法减少了任务阻塞缺陷,并将可行任务率从48.6%提高到94.8%,同时产生了更强的PPO(Proximal Policy Optimization)策略,并在评估时改善了对WebArena、WebShop和MiniWoB++的迁移,而无需调用大型语言模型(LLM)。这些结果表明,经过验证的合成环境可以作为紧凑型网络智能体的可扩展和可靠的训练基础,将合成网络智能体学习从表面合理性转向可执行的、基于状态的监督。
cs.AI / 47 / 2608.21914

Consistency Is Not Coherence: Orientation Search for Certified Alignments Between 4D Defence Upper Ontologies

一致性并非连贯性:4D 防御上层本体之间认证对齐的方向搜索
Rovai, Fabio
Abstract
We align three upper ontologies that sit under UK and NATO defence data infrastructure: the Information Exchange Standard (IES), the Higher Quality Data Model (HQDM) that underpins the National Digital Twin, and Basic Formal Ontology (BFO). No public alignment between IES and HQDM existed. Promoting a hand-curated 17-correspondence crosswalk to OWL and reasoning over the complete merged ontologies with HermiT produces three results that we believe matter beyond this pair.
Chinese Translation
我们对位于英国和北约防御数据基础设施下的三个上层本体进行了对齐:信息交换标准(IES)、支撑国家数字双胞胎的高质量数据模型(HQDM)以及基本形式本体(BFO)。IES 和 HQDM 之间没有公开的对齐。通过促进手工策划的 17 对应关系交叉表到 OWL,并对完整合并本体进行 HermiT 推理,产生了我们认为超越这一对的三个重要结果。
cs.AI / 48 / 2608.21925

ESCRAG-R1: Retrieval-Augmented Reinforcement Learning for Emotional Support Conversation

ESCRAG-R1:用于情感支持对话的检索增强强化学习
Liu, Weichu, Hu, Yuxuan, Sun, Yirong, Mao, Ningning, Zhang, Ziyun, Chen, Jian, Xu, Mingyang, Zhong, Qishan, Li, Chengming
Abstract
Emotional Support Conversation (ESC) systems aim to provide holistic support by balancing professional therapeutic competence with natural empathy. However, existing methods struggle to simultaneously achieve structured, stage-aware reasoning and seamless empathy-expertise alignment, often resulting in an artificial splicing of clinical strategies and generic reassurance. To overcome these limitations, we propose ESCRAG-R1, a unified framework that integrates retrieval-based psychological guidance into Group Relative Policy Optimization (GRPO). By incorporating retrieval into the reinforcement learning loop, ESCRAG-R1 transforms external knowledge into a robust learning signal that stimulates explicit internal reasoning prior to generation and fundamentally reshapes the model's internal policy. To provide the reliable supervision required for this optimization, we construct ESC-Preference, a high-quality dataset based on a Client--Counselor--Judge evaluation framework that delivers precise, empathy-aware reward signals. Extensive experiments demonstrate that ESCRAG-R1 significantly outperforms existing baselines by mitigating superficial splicing and realizing a natural integration of professional guidance and empathetic expression. Code and datasets are released at https://github.com/Matcha-Liu/ESCRAG-R1.
Chinese Translation
情感支持对话(ESC)系统旨在通过平衡专业治疗能力与自然同理心来提供全面的支持。然而,现有方法在同时实现结构化的、阶段感知的推理和无缝的同理心与专业知识对齐方面面临挑战,常常导致临床策略与一般性安慰的人工拼接。为克服这些局限,我们提出了ESCRAG-R1,一个将基于检索的心理指导整合到群体相对策略优化(GRPO)中的统一框架。通过将检索纳入强化学习循环,ESCRAG-R1将外部知识转化为强有力的学习信号,刺激生成前的显式内部推理,并从根本上重塑模型的内部策略。为了提供这一优化所需的可靠监督,我们构建了ESC-Preference,这是一个基于客户-顾问-评审框架的高质量数据集,能够提供精确的、关注同理心的奖励信号。大量实验表明,ESCRAG-R1显著优于现有基线,通过减轻表面拼接,实现了专业指导与同理表达的自然整合。代码和数据集已发布在 https://github.com/Matcha-Liu/ESCRAG-R1。
cs.AI / 49 / 2608.21928

GuardianBench: A Same-Scene Instruction-Contrastive Benchmark for Latent Contextual Risk in Embodied AI

GuardianBench:一种针对具身人工智能潜在情境风险的同场景指令对比基准
Zhang, Zhesheng, Lu, Jiahao, Liu, Wei, Pan, Cong, Yang, Jianhua, Chen, Yixiang, Yu, Hongyuan, Zhang, Mengqi, Lyu, Kailin, Chen, Zhumin, He, Keji
Abstract
In embodied AI, safety risk can be latent: a benign instruction and a safe scene become hazardous only when composed. Prior work has advanced embodied safety by varying visual contexts or evaluating execution-time dynamics, but the complementary axis of fixing the scene and varying only the instruction remains underexplored. We introduce GuardianBench, an instruction-contrastive benchmark grounded in international safety standards that isolates this latent contextual risk through 3,024 instruction-scene examples organized as same-scene Safe/Unsafe contrastive pairs across various hazard categories. Benchmarking state-of-the-art vision-language models (VLMs) reveals instruction-insensitive verdicts: models disproportionately approve both instructions under a given scene; across the primary models, average pair accuracy is only 24.1%. Our systematic rationale audit localizes the dominant failure: models fail to bind the instruction-relevant cues that differentiate safe from unsafe compositions. As a post-training case study, Verdict Log-Odds Supervision (VLOS), a lightweight verdict-level objective, substantially improves performance on open-weight backbones. Together, our latent contextual risk task formulation, standards-grounded contrastive benchmark construction, pair-level and rationale-level failure diagnosis, and benchmark-enabled verdict calibration establish GuardianBench as a controlled evaluation suite for exposing and improving safety reasoning over instruction-scene compositions under latent contextual risk.
Chinese Translation
在具身人工智能中,安全风险可能是潜在的:一个良性的指令和一个安全的场景在组合时才会变得危险。以往的研究通过改变视觉上下文或评估执行时动态来推动具身安全,但固定场景并仅变化指令的互补维度仍未得到充分探索。我们引入了GuardianBench,这是一种基于国际安全标准的指令对比基准,通过3,024个指令-场景示例,组织为同场景的安全/不安全对比对,隔离这种潜在的情境风险,涵盖各种危险类别。对最先进的视觉-语言模型(VLMs)进行基准测试揭示了对指令不敏感的裁决:模型在给定场景下对两种指令的批准比例不成比例;在主要模型中,平均对比准确率仅为24.1%。我们的系统性推理审计定位了主要的失败原因:模型未能绑定区分安全与不安全组合的指令相关线索。作为后训练案例研究,裁决对数监督(Verdict Log-Odds Supervision, VLOS),一种轻量级的裁决级目标,显著提高了开放权重骨干网络的性能。综上所述,我们的潜在情境风险任务构造、基于标准的对比基准构建、对比对和推理层面的失败诊断,以及基准支持的裁决校准,确立了GuardianBench作为一个受控评估套件,用于揭示和改善在潜在情境风险下的指令-场景组合的安全推理。
cs.AI / 50 / 2608.21941

Multimodal Prompt Learning with Irregular EHRs for Robust Monitoring of Critical Care Patients

基于不规则电子健康记录的多模态提示学习用于重症监护患者的稳健监测
Yang, Yixin, Sun, Yueyang, Liu, Weichen, Zhao, Xianbing, Liu, Sicen
Abstract
Accurate assessment of patients in intensive care units (ICUs) is essential for timely clinical intervention and improved patient outcomes. Multimodal electronic health records (EHRs), including structured physiological time series and longitudinal clinical notes, provide complementary information for critical care prediction. However, in real-world clinical settings, individual modalities may be partially observed or entirely unavailable, resulting in substantial performance degradation for existing multimodal models. To address this challenge, we propose a multimodal prompt-learning framework for robust clinical prediction under diverse missing-modality scenarios. The proposed framework introduces four complementary types of prompts: generative prompts, missing-signal prompts, missing-type prompts, and temporal prompts. Generative prompts construct surrogate latent representations for unavailable modalities, while missing-signal prompts distinguish observed representations from generated ones. Missing-type prompts condition the model on different modality-availability configurations, whereas temporal prompts perform condition-specific aggregation over temporally encoded clinical sequences. Together, these prompts enable the model to capture missingness-aware intramodal dependencies and cross-modal interactions within a unified architecture. Extensive experiments demonstrate that our method outperforms existing approaches across evaluation metrics on two missingness settings. Ablation and robustness analyses further verify the complementary contributions of the four prompt types and the effectiveness of the proposed framework for clinical prediction from incomplete multimodal EHR data.
Chinese Translation
在重症监护病房(ICUs)中,准确评估患者对于及时的临床干预和改善患者预后至关重要。多模态电子健康记录(EHRs),包括结构化生理时间序列和纵向临床笔记,为重症护理预测提供了互补信息。然而,在实际临床环境中,单一模态可能部分可观察或完全不可用,导致现有多模态模型的性能显著下降。为了解决这一挑战,我们提出了一种多模态提示学习框架,以在不同缺失模态场景下进行稳健的临床预测。该框架引入了四种互补类型的提示:生成提示、缺失信号提示、缺失类型提示和时间提示。生成提示为不可用模态构建替代潜在表示,而缺失信号提示则区分观察到的表示和生成的表示。缺失类型提示使模型在不同模态可用性配置下进行条件化,而时间提示则对时间编码的临床序列进行条件特定的聚合。这些提示共同使模型能够在统一架构内捕捉缺失感知的模态内依赖关系和跨模态交互。大量实验表明,我们的方法在两个缺失性设置下的评估指标上优于现有方法。消融和稳健性分析进一步验证了四种提示类型的互补贡献以及所提框架在从不完整多模态EHR数据进行临床预测中的有效性。
cs.AI / 51 / 2608.21942

TessIndex: Capability Verified Identity System for the Agent Economy

TessIndex:面向代理经济的能力验证身份系统
Goenka, Mehul, Pathak, Tejas, Asthana, Siddharth
Abstract
Software systems have traditionally been organized around applications where human users act as principal decision-makers. Recent developments in agentic capabilities alter this paradigm: software agents now autonomously translate high-level goals into structured tasks, orchestrating tools, services and sub-agents to execute complex workflows. This evolution gives rise to an agent economy where these autonomous agents capture real economic value. However, the infrastructure required to support the agent economy fails across three critical dimensions: the absence of persistent identity infrastructure prevents systemic accountability in agentic workflows; capability claims remain self-declared not backed by verifiable execution evidence; and the disconnect between creator identities, agent performance, and project value hinders the economic valuation of agents as assets. While existing registries provide naming and discovery, unifying these features around a persistent identity anchor remains largely unaddressed. TessIndex is a capability-verified identity system for agent primitives that utilizes a dual-plane architecture: the blockchain records compact commitments for identity, ownership, and verification, while centralized servers maintain dynamic metadata for discovery, commerce, and reputation. It establishes: persistent identities across agent primitives to enforce systemic accountability in autonomous workflows; a predicate-based verification process replacing self-declared claims with cryptographic capability proof; an identity infrastructure that links agent performance to both project and creator identities while capturing value through tokenization. Ultimately, TessIndex serves as an integrated infrastructure that binds an agent's existence across capabilities, execution, and reputation into a single persistent identity.
Chinese Translation
软件系统传统上围绕应用程序组织,其中人类用户作为主要决策者。最近在代理能力方面的发展改变了这一范式:软件代理现在能够自主将高层目标转化为结构化任务,协调工具、服务和子代理以执行复杂的工作流程。这一演变催生了一个代理经济,在这个经济中,这些自主代理捕获了真实的经济价值。然而,支持代理经济所需的基础设施在三个关键维度上存在缺陷:缺乏持久身份基础设施导致代理工作流程中的系统性问责缺失;能力声明仍然是自我声明,未得到可验证的执行证据支持;创造者身份、代理表现和项目价值之间的脱节阻碍了将代理作为资产的经济估值。尽管现有注册表提供了命名和发现功能,但围绕持久身份锚点统一这些特性仍然基本未得到解决。TessIndex 是一个针对代理原语的能力验证身份系统,采用双层架构:区块链记录身份、所有权和验证的紧凑承诺,而集中式服务器维护用于发现、商业和声誉的动态元数据。它建立了:在代理原语中实现持久身份,以加强自主工作流程中的系统性问责;基于谓词的验证过程,用加密能力证明替代自我声明的声明;一个将代理表现与项目和创造者身份相连接的身份基础设施,同时通过代币化捕获价值。最终,TessIndex 作为一个集成基础设施,将代理的存在绑定于能力、执行和声誉,形成一个单一的持久身份。
cs.AI / 52 / 2608.21952

SSDi8: Accurate and Efficient 8-bit Quantization for State Space Duality

SSDi8:状态空间对偶的高效准确的8位量化
Kim, Hyunwoo, Ko, Byoungchan, Kang, Minseok, Kim, Minwoo, Lee, Dongjin, Lee, Jaehoon, Yoon, Sungroh, Jung, Dahuin
Abstract
Recent advances in sequence modeling have highlighted Mamba as a state space architecture offering efficient long-range dependency modeling and providing a viable alternative to Transformers. Building upon this, Mamba-2 introduces the Structured State Space Duality (SSD), which integrates recurrent and attention modes to achieve efficiency and scalability. However, this architectural expansion substantially increases memory and latency overhead, underscoring the need for efficient compression strategies tailored to SSD. In this work, we present SSDi8, the first post-training quantization framework specifically designed for SSD to maintain a persistent INT8 path. SSDi8 introduces a reformulation that decouples element-wise multiplications from matrix multiplications, enabling reuse of quantized activations across modules. Moreover, SSDi8 adaptively quantizes channel-varying activations at cost-effective points, further reducing latency. On the accuracy side, SSDi8 explicitly leverages the intrinsic dimensional decomposition of SSD, exploiting distinct outlier distributions across axes, and incorporates an error correction term based on per-channel error statistics. Comprehensive experiments demonstrate that SSDi8 achieves accuracy comparable to FP16 while delivering up to 1.4x speedup in W4A8 and W8A8 settings. We further validate its robustness in resource-constrained environments by deploying it on the Orin NX device.
Chinese Translation
最近的序列建模进展突出了Mamba作为一种状态空间架构,能够有效建模长程依赖,并为变换器(Transformers)提供了可行的替代方案。在此基础上,Mamba-2引入了结构化状态空间对偶(Structured State Space Duality,SSD),将递归模式和注意力模式结合,以实现效率和可扩展性。然而,这种架构扩展显著增加了内存和延迟开销,强调了针对SSD量身定制的高效压缩策略的必要性。在本研究中,我们提出了SSDi8,这是第一个专门为SSD设计的后训练量化框架,旨在维持持续的INT8路径。SSDi8引入了一种重新表述,将逐元素乘法与矩阵乘法解耦,从而实现跨模块重用量化激活。此外,SSDi8在成本效益高的点上自适应量化通道变化的激活,进一步降低延迟。在准确性方面,SSDi8明确利用了SSD的内在维度分解,利用轴上的不同异常值分布,并结合基于每通道误差统计的误差修正项。全面的实验表明,SSDi8在W4A8和W8A8设置中实现了与FP16相当的准确性,同时提供了高达1.4倍的加速。我们还通过在Orin NX设备上部署SSDi8,进一步验证了其在资源受限环境中的鲁棒性。
cs.AI / 53 / 2608.21964

Repo2Skill-Evo: Repository Skills Go Stale in Silence

Repo2Skill-Evo:代码库技能在沉默中变得陈旧
Duan, Chenyuan, Shi, Ge, Mao, Zineng, Zhang, Ge, Liang, Hao, Piao, Yinzhu, Wu, Yuchen, Yao, Zhixin, Huang, Kaiyu, Huang, Wenhao, Sun, Linzhuang, Yan, Shen, Zhang, Wentao
Abstract
Large language model (LLM) agents increasingly operate over evolving software repositories, where success depends on repository-specific procedural knowledge: which APIs to call, which scripts to run, and which conventions the current release expects. Agent skills externalize this knowledge into reusable units, and prior work shows that they can improve agent performance. What remains unclear is whether that improvement is durable. The same version specificity that makes a skill useful also makes it fragile: after a release, it may become stale without raising any explicit signal, while continuing to provide obsolete guidance. Externalizing knowledge into a skill can therefore make its decay invisible. We study whether agents can keep this externalized knowledge current. Repo2Skill-Evo casts each release transition as a skill-maintenance task: given a V1 skill set and the official V1-to-V2 patch, an agent must update obsolete skill content while preserving guidance that remains valid. Across 57 real-world repositories and 105 selected release transitions, every evaluated transition invalidates part of the V1 skill set. Yet six frontier agents reach only 29.9%-69.7% avg@3 macro F1 under a patch-grounded removal metric that balances stale-content recall against over-editing precision. Across runs, two opposing errors dominate: incomplete coverage of affected files in the skill set leaves stale content untouched, while overbroad editing is associated with higher recall but lower precision. Repository skills go stale in silence, and even frontier agents cannot reliably maintain them.
Chinese Translation
大型语言模型(LLM)代理越来越多地在不断发展的软件代码库上运行,其成功依赖于特定于代码库的程序知识:调用哪些API、运行哪些脚本,以及当前版本期望哪些约定。代理技能将这些知识外化为可重用的单元,先前的研究表明,它们可以提高代理的性能。然而,尚不清楚这种改善是否具有持久性。使技能有用的版本特异性也使其脆弱:在发布后,技能可能会变得陈旧而没有任何明确的信号,同时继续提供过时的指导。因此,将知识外化为技能可能会使其衰退变得不可见。我们研究代理是否能够保持这种外化知识的时效性。Repo2Skill-Evo将每次发布过渡视为一个技能维护任务:给定V1技能集和官方的V1到V2补丁,代理必须在保留仍然有效的指导的同时更新过时的技能内容。在57个真实世界的代码库和105个选定的发布过渡中,每个评估的过渡都使部分V1技能集失效。然而,六个前沿代理在一个基于补丁的移除度量下,仅达到29.9%-69.7%的avg@3宏F1,该度量平衡了过时内容的召回率与过度编辑的精度。在多个运行中,两种相对的错误占主导地位:技能集中受影响文件的覆盖不完全使得过时内容未被触及,而过于宽泛的编辑则与更高的召回率相关,但精度较低。代码库技能在沉默中变得陈旧,即使是前沿代理也无法可靠地维护它们。
cs.AI / 54 / 2608.21976

Closed-loop AI achieves certifiable engineering design

闭环人工智能实现可认证的工程设计
Yu, Tianyi, Tao, Chengxing, Shen, Haoxuan, Li, Huiyang, Chen, Rugang, Teng, Long, Wang, Lilin, Li, Yan, Chen, Qingbin, Xu, Chaogang, Wang, Lizhong
Abstract
Agentic AI has automated parts of scientific discovery, including paper generation, expert-level coding, therapeutic proposal, and autonomous experimentation. Complex physical engineering design remains a gap, because candidates must satisfy simultaneous constraints in fluid dynamics, solid mechanics, and structural stability. We introduce The AI Engineer, an agentic framework that couples large language models (LLMs) to deterministic engineering backends in a closed loop: natural-language requirements are converted into design-domain geometry and mesh; topology is optimized with bi-directional evolutionary structural optimization (BESO) coupled to the CalculiX solver; and member sizes are refined with particle swarm optimization (PSO) coupled to Zwind under offshore aero-hydro-servo-elastic load cases. To explore many designs without per-candidate certification cost, an Automated Reviewer scores each candidate on five dimensions (capacity, steel intensity, unit cost, constructability, and fatigue life) using piecewise-linear functions calibrated on 11 real floating-wind projects. Search terminates only when a candidate reaches a composite score $S \ge 85$ (grade A) with no subscore below 60. We validated this gate by submitting the top-scoring design to the China Classification Society (CCS) for Approval in Principle (AIP), which it passed; AIP is thus an external check that the reviewer tracks professional judgment, not the daily objective. The certified design outperforms the human-optimized TuQiang baseline, reducing steel mass and unit capital cost by 8.1% each while meeting all AIP criteria. This verification-closed regime, in which every proposal is judged by deterministic physics and codified limit states, distinguishes The AI Engineer from open-ended generative systems. Remaining limits include detailed design and fabrication-hard constraints.
Chinese Translation
代理人工智能已自动化了科学发现的部分过程,包括论文生成、专家级编码、治疗方案提出和自主实验。然而,复杂的物理工程设计仍然存在空白,因为候选设计必须满足流体动力学、固体力学和结构稳定性的同时约束。我们介绍了AI工程师(The AI Engineer),这是一个将大型语言模型(LLMs)与确定性工程后端结合在闭环中的代理框架:自然语言需求被转换为设计领域的几何形状和网格;拓扑通过与CalculiX求解器耦合的双向进化结构优化(BESO)进行优化;构件尺寸通过与Zwind耦合的粒子群优化(PSO)在海上气动-水动力-伺服-弹性载荷情况下进行精细调整。为了在不增加每个候选设计认证成本的情况下探索多种设计,自动评审器根据在11个真实浮动风电项目上校准的分段线性函数,对每个候选设计在五个维度(容量、钢材强度、单位成本、可施工性和疲劳寿命)进行评分。搜索仅在候选设计达到综合得分$S ge 85$(A级)且没有子得分低于60时终止。我们通过将得分最高的设计提交给中国船级社(CCS)进行原则批准(AIP)来验证这一门槛,并成功通过;因此,AIP是一个外部检查,确保评审者遵循专业判断,而不是日常目标。经过认证的设计在钢材质量和单位资本成本上分别比人类优化的TuQiang基线减少了8.1%,同时满足所有AIP标准。这种验证闭合机制,使得每个提案都受到确定性物理和规范极限状态的评判,区别于开放式生成系统。剩余的限制包括详细设计和制造硬约束。
cs.AI / 55 / 2608.21979

Beyond Similarity: Heterogeneous Graph Learning for Multi-Objective Food Substitution in Charitable Food Agencies

超越相似性:用于慈善食品机构多目标食品替代的异构图学习
Chowdhury, Naimur Rahman, Hossain, Limon Bin
Abstract
Charitable food agencies play an important role in alleviating food insecurity by distributing donated food to people in need. However, they rely on ad hoc in-kind donations and often face shortages of specific foods, so they offer substitutes. A good food substitution requires matching household preferences, nutritional needs, and item similarity. Agencies have limited direct records of consumption behavior due to resource constraints, making it challenging to make an appropriate substitution decision that meets multiple criteria. In this study, we propose a heterogeneous graph neural network (HeteroGNN), a source-grounded recommendation framework for food substitution in charitable food agencies. We first build a unified relational graph from large-scale public data sources, combining household behavior on food consumption and food nutrient information in the United States (US) context. We treat the substitution recommendation as a multi-objective ranking problem with three targets, including behavior affinity, health suitability, and substitution similarity. We train and validate the proposed framework under standard graph relationship and adverse cold-start settings by removing relational edges from the graph. Our results show that the proposed framework leverages relational information beyond node features in predicting consumption behavior. Additionally, the proposed framework remains robust with sparsity when the model receives incomplete information about behavior and nutrient features. Finally, we show the weak correlation among different objectives, thereby justifying the multi-objective framing as a replacement for an aggregated decision. The proposed framework can help downstream charitable agency decision-makers make contextspecific substitution recommendations with limited information available.
Chinese Translation
慈善食品机构在通过分发捐赠食品来缓解食品不安全问题方面发挥着重要作用。然而,它们依赖于临时的实物捐赠,常常面临特定食品短缺的问题,因此提供替代品。良好的食品替代需要匹配家庭偏好、营养需求和项目相似性。由于资源限制,机构对消费行为的直接记录有限,这使得在满足多个标准的情况下做出适当的替代决策变得具有挑战性。在本研究中,我们提出了一种异构图神经网络(HeteroGNN),这是一个基于源的食品替代推荐框架,适用于慈善食品机构。我们首先从大规模公共数据源构建一个统一的关系图,结合了美国背景下家庭食品消费行为和食品营养信息。我们将替代推荐视为一个多目标排序问题,包含三个目标:行为亲和性、健康适宜性和替代相似性。我们在标准图关系和不利冷启动设置下,通过从图中移除关系边缘来训练和验证所提出的框架。我们的结果表明,所提出的框架在预测消费行为时利用了超越节点特征的关系信息。此外,当模型接收到关于行为和营养特征的不完整信息时,所提出的框架在稀疏性方面仍然保持稳健。最后,我们展示了不同目标之间的弱相关性,从而证明了多目标框架作为聚合决策的替代的合理性。所提出的框架可以帮助下游慈善机构决策者在信息有限的情况下做出特定情境的替代推荐。
cs.AI / 56 / 2608.21985

Redteaming Leading Arabic LLMs with ASAS

使用ASAS对领先的阿拉伯大型语言模型进行红队测试
Abed, Fidaa, Khan, Haidar, Bari, M Saiful, Khan, Babar, Abujabal, Abdalghani
Abstract
As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safety and cultural alignment is increasingly critical. However, Arabic LLM safety remains underexplored, especially in adversarial evaluation settings. We introduce the Arabic Safety Index (ASAS), the first fully human-curated Arabic benchmark for redteaming LLMs. ASAS contains 801 prompts spanning 8 safety categories and 8 attack strategies, with ideal responses in Modern Standard Arabic (MSA). We conduct a redteaming evaluation across seven leading models with Arabic capabilities, including GPT-4o, Claude 3.7 Sonnet, and regional models such as ALLaM and FANAR. Human annotators rate responses using a structured 4-point safety scale, revealing that most models fail to defend against 50% of unsafe prompts. Our findings highlight major safety gaps in high-harm categories such as weapons and illicit substances, with direct and obfuscation-based attacks proving most effective. The results also show that language alignment does not readily transfer across languages, and that automated safety judges (e.g., GPT-4o) perform poorly compared to human annotators. ASAS provides a culturally grounded benchmark and redteaming protocol to drive progress in Arabic LLM safety.
Chinese Translation
随着大型语言模型(LLMs)在阿拉伯语地区的应用不断增长,确保其安全性和文化一致性变得愈发重要。然而,阿拉伯LLM的安全性仍然未得到充分研究,尤其是在对抗性评估环境中。我们引入了阿拉伯安全指数(ASAS),这是第一个完全由人类策划的阿拉伯红队测试基准。ASAS包含801个提示,涵盖8个安全类别和8种攻击策略,并提供现代标准阿拉伯语(MSA)的理想响应。我们对七个具有阿拉伯语能力的领先模型进行了红队评估,包括GPT-4o、Claude 3.7 Sonnet以及区域模型如ALLaM和FANAR。人类评审员使用结构化的4分安全评分标准对响应进行评分,结果显示大多数模型未能有效防御50%的不安全提示。我们的研究结果揭示了在高危类别(如武器和非法物质)中存在重大安全缺口,直接攻击和模糊攻击被证明是最有效的。结果还表明,语言一致性并不能轻易跨语言转移,且自动安全评估者(例如GPT-4o)的表现明显不如人类评审员。ASAS提供了一个以文化为基础的基准和红队测试协议,以推动阿拉伯LLM安全性的进展。
cs.AI / 57 / 2608.22014

DynaContext: Self-Improving Dynamic Contextualization of Optimized Prompts for Heterogeneous Parameter Extraction

DynaContext:自我改进的动态上下文化优化提示以提取异构参数
Yu, Joe, Paul, Shibin Thomas Stanley, Mayer, Sven
Abstract
Automated prompt and skill optimization typically produces a single static instruction that is reused across inference instances until the next optimization cycle. However, this approach cannot adapt when the required context, constraints, and evidence vary from one instance to another. For instance, parameter extraction from electronic component descriptions breaks this assumption: resistors, capacitors, transistors, and connectors require different fields, unit constraints, and demonstrations, and each input provides a different evidence state. We introduce DynaContext, a framework that combines an offline-optimized extraction core, learned with GEPA or SkillOpt, with inference-time contextual adaptation and validation-gated self-improvement. DynaContext routes each item through internal, external, or fallback evidence paths and composes an item-specific prompt from the core, schema, evidence, unresolved fields, and validated demonstrations. Deterministic validation and an LLM judge gate every output, uncertain cases go to human review, and only human-verified corrections enter the demonstration memory. On a single-category benchmark, average accuracy increases from 86.6% for the base prompt to 96.9% for standalone SkillOpt and 98.6% for the best DynaContext configuration. Across 850 heterogeneous gold parameter facts, average field-level F1 increases from 51.8% for an unoptimized, demonstration-free control to 59.2% with dynamic demonstrations alone, 66.9% with the optimized core alone, and 71.0% with both. Holding the model fixed, the full configuration outperforms the deployed static-prompting pipeline by 17.3 F1 points on average.
Chinese Translation
自动化的提示和技能优化通常生成一个静态指令,该指令在推理实例之间重复使用,直到下一个优化周期。然而,当所需的上下文、约束和证据在不同实例之间变化时,这种方法无法适应。例如,从电子元件描述中提取参数打破了这一假设:电阻器、电容器、晶体管和连接器需要不同的字段、单位约束和示例,每个输入提供不同的证据状态。我们提出了DynaContext,一个结合了离线优化提取核心(通过GEPA或SkillOpt学习)与推理时上下文适应和验证门控自我改进的框架。DynaContext通过内部、外部或后备证据路径对每个项目进行路由,并从核心、模式、证据、未解决字段和经过验证的示例中组合出特定于项目的提示。确定性验证和LLM判别器对每个输出进行评估,不确定的案例将提交人类审查,只有经过人类验证的修正才会进入示例记忆。在单类别基准测试中,平均准确率从基础提示的86.6%提高到独立SkillOpt的96.9%,以及最佳DynaContext配置的98.6%。在850个异构金参数事实中,平均字段级F1从未优化且没有示例的控制组的51.8%提高到仅使用动态示例的59.2%,仅使用优化核心的66.9%,以及两者结合的71.0%。在模型固定的情况下,完整配置的表现平均比部署的静态提示管道提高了17.3个F1点。
cs.AI / 58 / 2608.22018

SPAR-Hate: An Auditor-Guided Multi-Agent Framework for Bilingual Hate Speech Parsing

SPAR-Hate:一种审计员引导的多智能体双语仇恨言论解析框架
Lyu, Yifan, Lin, Dianqing, Li, Xinran, Qiao, Jiaqi, Xu, Xiujuan
Abstract
Hate speech detection has recently shifted from coarse-grained classification to structured parsing, where systems must jointly identify hateful targets, arguments, and target-level labels. However, existing studies primarily emphasize benchmark evaluation while paying less attention to the cultural, linguistic, and social-group challenges involved in structured hate speech parsing. To address these challenges, we propose SPAR-Hate, an auditor-guided multi-agent framework for bilingual hate speech parsing. The framework first decomposes documents into clause-level decision units and then generates evidence-grounded judgments from three complementary perspectives: Victim, Moderator, and Cultural Bystander. An evidence-constrained arbitration process resolves conflicts among role-specific predictions and aggregates them into structured sample-level outputs. Experiments on the STATE-ToxiCN and TBO benchmarks show that SPAR-Hate consistently improves bilingual hate parsing across diverse large language models. The framework achieves state-of-the-art results on bilingual multi-tuple extraction tasks, with the largest gains observed under stricter structural evaluation metrics.
Chinese Translation
仇恨言论检测最近已从粗粒度分类转向结构化解析,系统必须共同识别仇恨目标、论点和目标级标签。然而,现有研究主要强调基准评估,而较少关注结构化仇恨言论解析中涉及的文化、语言和社会群体挑战。为了解决这些挑战,我们提出了SPAR-Hate,一种审计员引导的多智能体双语仇恨言论解析框架。该框架首先将文档分解为从句级决策单元,然后从三个互补的视角生成基于证据的判断:受害者、调解者和文化旁观者。一个受证据约束的仲裁过程解决角色特定预测之间的冲突,并将其聚合为结构化的样本级输出。在STATE-ToxiCN和TBO基准上的实验表明,SPAR-Hate在不同的大型语言模型中始终提高了双语仇恨解析的效果。该框架在双语多元组提取任务上实现了最先进的结果,在更严格的结构评估指标下观察到最大的提升。
cs.AI / 59 / 2608.22026

One-Step Evolution for Long-Time Extrapolation: An Error-Bound-Informed and Prior-Guided Neural Residual Framework for Autonomous PDEs

一步演化的长期外推:一种基于误差界限和先验指导的神经残差框架用于自主偏微分方程(PDE)
Zhang, Maqun, Gao, Feng, Chen, Wankun, Yu, Hui, Gan, Yanhai, Dong, Junyu
Abstract
Accurate simulation of the long-time evolution of systems governed by partial differential equations (PDEs) is central to scientific computing. Among existing deep learning?based approaches for solving PDEs, neural operators typically rely on extensive trajectory data, whereas physics-informed meth?ods often exhibit limited stability during long-time extrapolation. For a well-posed autonomous PDE, long-time trajectories can be generated by repeated composition of a fixed-step evolution operator; hence, long-time extrapolation depends on controlling the approximation error of this operator and the propagation of that error under recursive composition. Accordingly, we propose a numerical-prior-guided, physics-constrained method trained without ground-truth trajectory supervision: a low-cost numerical prior reduces the difficulty of approximating the one?step evolution operator, while a weak-form PDE residual provides a computable proxy for the one-step error term in the error?propagation bound. We validate the method on five benchmark cases spanning four PDE classes and compare it with ten physics?informed learning methods under a unified protocol that excludes ground-truth trajectories from training and model selection. The results indicate that, in all five cases, the proposed method reduces long-time extrapolation error relative to the numerical prior and outperforms the best competing baseline in each case, thereby improving long-time simulation accuracy across different PDEs without ground-truth trajectory supervision. The source code developed for this paper will be made publicly available upon acceptance of the manuscript.
Chinese Translation
准确模拟由偏微分方程(PDE)支配的系统的长期演化是科学计算的核心。在现有的基于深度学习的方法中,神经算子通常依赖于大量的轨迹数据,而物理信息方法在长期外推过程中往往表现出有限的稳定性。对于一个良定义的自主PDE,长期轨迹可以通过重复组合固定步长的演化算子生成;因此,长期外推依赖于控制该算子的近似误差及其在递归组合下的传播。因此,我们提出了一种无真实轨迹监督的数值先验指导、物理约束的方法:低成本的数值先验降低了近似一步演化算子的难度,而弱形式的PDE残差提供了可计算的代理,用于误差传播界限中的一步误差项。我们在五个基准案例上验证了该方法,涵盖了四个PDE类别,并在一个排除真实轨迹的训练和模型选择的统一协议下,将其与十种物理信息学习方法进行了比较。结果表明,在所有五个案例中,所提方法相对于数值先验减少了长期外推误差,并在每个案例中优于最佳竞争基线,从而在没有真实轨迹监督的情况下提高了不同PDE的长期模拟精度。本文开发的源代码将在手稿接受后公开。
cs.AI / 60 / 2608.22048

More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning

更准确还是更高效?评估本地部署的紧凑型开放权重语言模型在数学推理中的表现
Powers, Orion, Seum, Daniella, Slhoub, Khaled
Abstract
Large language models are increasingly deployed on local hardware for privacy, cost, and accessibility reasons. Yet many evaluations emphasize accuracy while fewer quantify local runtime and energy, characterize failure modes, or apply paired statistical comparisons under controlled conditions. This paper presents a controlled, documented procedure for evaluating locally hosted LLMs on mathematical reasoning. It combines fixed inference settings, hierarchical answer extraction and verification, explicit failure-mode classification, and per-question resource measurement, and reports accuracy with paired significance tests and effect sizes. We demonstrate it in a preliminary study of three compact open-weight models under five billion parameters, Gemma3:4b (Google), Phi3:3.8b (Microsoft), and Qwen3:4b (Alibaba), across datasets spanning Grade 8 Math, Calculus I, and Advanced Probability and Statistics. All models ran through the same local inference server on one workstation, using a shared prompt template, controlled settings, and a matched question set per dataset. No single model dominates. Qwen3:4b is most accurate on two datasets and Gemma3:4b on Calculus I, yet Gemma3:4b returns roughly three times more correct answers per watt-hour than Qwen3:4b on every dataset while generating far fewer output tokens; Qwen3:4b requires substantially more generation time, energy, and output per question. Phi3:3.8b is substantially less accurate on all three datasets; its low extraction-failure rate indicates incorrect answers rather than unparsed output, though we caveat possible prompt-format effects. These preliminary findings indicate that accuracy alone is an insufficient basis for selecting a local model.
Chinese Translation
大型语言模型因隐私、成本和可及性等原因,越来越多地部署在本地硬件上。然而,许多评估强调准确性,而较少量化本地运行时间和能耗,描述失败模式,或在受控条件下进行配对统计比较。本文提出了一种受控的、文档化的程序,用于评估本地托管的LLM(大型语言模型)在数学推理方面的表现。该程序结合了固定推理设置、分层答案提取与验证、明确的失败模式分类以及每个问题的资源测量,并通过配对显著性检验和效应大小报告准确性。我们在一项初步研究中展示了三种紧凑型开放权重模型(参数不超过五十亿),分别为Gemma3:4b(谷歌)、Phi3:3.8b(微软)和Qwen3:4b(阿里巴巴),涵盖了八年级数学、微积分I和高级概率与统计等数据集。所有模型均通过同一本地推理服务器在一台工作站上运行,使用共享的提示模板、受控设置和每个数据集匹配的问题集。没有单一模型占据绝对优势。Qwen3:4b在两个数据集上表现最准确,而Gemma3:4b在微积分I上表现最佳,然而Gemma3:4b在每个数据集上每瓦时返回的正确答案大约是Qwen3:4b的三倍,同时生成的输出标记数量远少于Qwen3:4b;Qwen3:4b在每个问题上的生成时间、能耗和输出量显著更高。Phi3:3.8b在所有三个数据集上的准确性明显较低;其低提取失败率表明错误答案而非未解析输出,尽管我们需要注意可能的提示格式影响。这些初步发现表明,单靠准确性不足以作为选择本地模型的依据。
cs.AI / 61 / 2608.22055

GenCoord: Skill-Path Commitments under Private Information

GenCoord:在私有信息下的技能路径承诺
He, Peng, Zhu, Junning, Yuan, Haohan, Liang, Jianpeng
Abstract
Suppose one embodied agent knows what must be built, while its teammate alone knows which transformation its workcell can perform. Neither local view determines who should act, what should be handed off, or how the joint task should continue. We introduce GenCoord, which turns the task consequence of such private facts into an executable skill-path commitment. A local Qwen3.5-0.8B model emits a multi-step SELF plan and peer REQ; bounded feedback conditions route revision when the deciding capability is peer-local. The resolved commitment is parsed, checked, canonically materialized, compiled to Mineflayer skills, and verified by handoff and terminal state. Counterfactual interventions that hold the world, call schedule, and executor unchanged make requester revision and receiver execution follow the injected task consequence in both directions. Across three independently trained seeds, correct capability feedback closes the paired local-information gap from 50% to 100%. Multi-step commitments improve held-out-template success by 6.9 points while reducing model decisions by 32%. At matched closed-loop quality on 128 held-out semantic clusters, Short DSL reduces peer traffic by 92.8% and median time-to-commitment by 68.2% relative to controlled free-form communication. These results identify executable task consequences as the coordination unit connecting distributed local reasoning to verified joint action.
Chinese Translation
假设一个具身代理知道必须构建的内容,而其队友则独自知道其工作单元可以执行的转换。没有任何局部视图能够决定谁应该行动、应该交接什么,或联合任务应该如何继续。我们提出了GenCoord,它将此类私有事实的任务后果转化为可执行的技能路径承诺。一个局部的Qwen3.5-0.8B模型生成一个多步骤的SELF计划和同行REQ;在决定能力为同行本地时,受限反馈条件会引导修订。解析后的承诺经过检查、规范化物化、编译为Mineflayer技能,并通过交接和终端状态进行验证。保持世界、调用调度和执行者不变的反事实干预使请求者修订和接收者执行在两个方向上遵循注入的任务后果。在三个独立训练的种子中,正确的能力反馈将配对的局部信息差距从50%缩小到100%。多步骤承诺使保留模板的成功率提高了6.9个百分点,同时减少了模型决策的数量32%。在128个保留的语义集群上,匹配的闭环质量下,短DSL相较于控制的自由形式通信减少了92.8%的同行流量和68.2%的中位时间承诺。这些结果将可执行的任务后果识别为连接分布式局部推理与验证的联合行动的协调单元。
cs.AI / 62 / 2608.22061

MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds

记忆胜出一切:通过社交媒体动态进行间接偏见注入攻击
Seo, Minjae, Choi, Wonwoo, Han, Geonwoo, Kwon, Taekyoung, Kim, Yongsu, Seo, Sang, Noh, Jaewon, Baek, Hankyul, Seo, Seongyun, You, Myoungsung
Abstract
Personal AI agents routinely consume external content while performing tasks such as web browsing, email processing, and SNS feed summarization, and they retain selected information or execution results in persistent memory for later use. We show that this ordinary ingestion of external content opens an indirect path for manipulating subsequent agent behavior. Based on this observation, we present IBIA, an Indirect Bias Injection Attack that plants an adversary-aligned stance on a specific topic into a victim agent's memory through external content, without direct access to the agent, its memory, or future user queries. For this, IBIA combines three mechanisms: comment cloaking, which keeps the crafted content consistent with the surrounding discussion, comment watermarking, which enables lightweight identification during curation, and category anchoring, which makes the retained stance salient under later related requests. We evaluate IBIA on BiasBench, a benchmark of 6,000 adversary-crafted social comments and 120 email instances. The watermark-based curation identifies 95.9% of the injected comments. Under the OpenClaw setting, IBIA achieves adversary-aligned response rates (AARs) of 91.2% on average across four downstream tasks, including 86.6% on the frontier GPT-5.5. We further propose a memory boundary defense that detects the injected bias and reduces AARs to 80.6%.
Chinese Translation
个人人工智能代理在执行网页浏览、电子邮件处理和社交网络动态摘要等任务时,通常会消耗外部内容,并将选定的信息或执行结果保存在持久内存中以备后用。我们展示了这种普通的外部内容摄取为操控后续代理行为打开了一条间接路径。基于这一观察,我们提出了间接偏见注入攻击(IBIA),该攻击通过外部内容将与对手一致的特定主题立场植入受害代理的记忆中,而无需直接访问代理、其内存或未来用户查询。为此,IBIA结合了三种机制:评论伪装,保持所创作内容与周围讨论的一致性;评论水印,使得在策划过程中能够轻松识别;以及类别锚定,使得在后续相关请求中保留的立场显著。我们在BiasBench上评估了IBIA,这是一个包含6000条对手制作的社交评论和120个电子邮件实例的基准测试。基于水印的策划识别了95.9%的注入评论。在OpenClaw设置下,IBIA在四个下游任务中实现了平均91.2%的对手一致响应率(AAR),其中在前沿的GPT-5.5上达到86.6%。我们进一步提出了一种内存边界防御,能够检测注入的偏见,并将AAR降低至80.6%。
cs.AI / 63 / 2608.22062

Search Broadly, Seek Evidence on Both Sides, Decide Narrowly: Evidence-Admissible GraphRAG for Longitudinal Clinical Event Verification

广泛搜索,双向寻证,狭窄决策:用于纵向临床事件验证的证据可接纳图模型 GraphRAG
Lin, Xingtao, Feng, Yubo, Liu, Weixin, Ren, Hangqi, Zhou, Junchao, Sun, Caiwan, Chen, You
Abstract
Longitudinal clinical event-relation verification determines whether a patient record supports a specified relation among two or more clinical events. This task is challenging because evidence is distributed across structured records, notes, laboratory trajectories, encounters, and time, while negation, temporal mismatch, repeated documentation, and conflicting findings can make retrieved information appear relevant without establishing the relation. We present MedEventGraph-RAG, an evidence-admissible framework that represents event occurrences in a patient-specific graph and links each occurrence to source evidence, including structured rows, note spans, timestamps, and numerical trajectories. Given a verification query specifying events, relation, and clinical scope, the graph guides discovery of candidate event chains and retrieves evidence from both supporting and contradicting sides. A query-specific evidence contract filters information by patient identity, scope, occurrence binding, and source traceability before a separate assessor determines supported, conflicting, refuted, or insufficient outcomes. Across ten protocols on i2b2, n2c2, MIMIC-IV, and LUNGUAGE, MedEventGraph-RAG achieves balanced accuracies of 78.6, 67.3, and 96.8 on temporal, medication-adverse-event, and recorded-order verification, improving over the strongest matched baselines by 26.9, 4.9, and 30.4 points. Under evidence masking, it reaches 92.2 balanced accuracy with no false-support predictions. When intermediate events are hidden, it recovers complete source-traceable event chains in 57.9% of i2b2 and 70.0% of LUNGUAGE cases. These results show that separating broad evidence discovery from narrow evidence-admissible assessment improves longitudinal clinical verification and reduces unsupported conclusions.
Chinese Translation
纵向临床事件关系验证确定患者记录是否支持两个或多个临床事件之间的特定关系。该任务具有挑战性,因为证据分布在结构化记录、笔记、实验室轨迹、就诊记录和时间中,而否定、时间不匹配、重复文档和相互矛盾的发现可能使检索到的信息看似相关,但并未建立关系。我们提出了 MedEventGraph-RAG,这是一种证据可接纳的框架,能够在患者特定图中表示事件发生,并将每个发生与源证据链接,包括结构化行、笔记范围、时间戳和数值轨迹。给定一个指定事件、关系和临床范围的验证查询,该图指导候选事件链的发现,并从支持和反对两方面检索证据。查询特定的证据合同根据患者身份、范围、发生绑定和源可追溯性过滤信息,然后由单独的评估者确定支持、冲突、反驳或不足的结果。在 i2b2、n2c2、MIMIC-IV 和 LUNGUAGE 的十个协议中,MedEventGraph-RAG 在时间、药物不良事件和记录顺序验证上分别实现了 78.6、67.3 和 96.8 的平衡准确率,分别比最强匹配基线提高了 26.9、4.9 和 30.4 个百分点。在证据掩蔽下,其平衡准确率达到 92.2,且没有错误支持预测。当中间事件被隐藏时,它在 57.9% 的 i2b2 和 70.0% 的 LUNGUAGE 案例中恢复了完整的源可追溯事件链。这些结果表明,将广泛的证据发现与狭窄的证据可接纳评估分开,有助于改善纵向临床验证并减少不支持的结论。
cs.AI / 64 / 2608.22063

From SQL Generation to Tool Selection: A Domain-Oriented Pattern for MCP Servers

从 SQL 生成到工具选择:面向领域的 MCP 服务器模式
Bogliolo, Bartolomeo
Abstract
Agents built on Large Language Models (LLMs) increasingly reach enterprise data through the Model Context Protocol (MCP), and many MCP database servers maximize flexibility by exposing a single generic SQL execution tool. This paper proposes the Domain-Oriented Tooling Pattern: instead of generating SQL at query time, the model selects from a small set of domain-aligned tools whose parameterized queries encapsulate schema navigation, joins and business rules on the server side. We formalize the pattern around three architectural invariants and introduce Model Demotion, the observation that replacing SQL synthesis with intent classification lowers the model tier required to serve routine requests. As a reference implementation we present MCP Blueprint, an open-source framework in which domain tools are defined declaratively as YAML metadata plus external parameterized SQL files. We evaluate the pattern with a public reproducibility benchmark comparing three MCP server designs - raw SQL execution, a thin generic tool pack, and a verticalized domain pack - on four local models (3B-8B) across seventeen customer-facing tasks over the Sakila database (609 completed cells; temperature 0; three repetitions per cell). The verticalized pack reaches a pooled mean score of 0.939 versus 0.666 for raw SQL and 0.605 for the generic pack; the smallest model improves from 0.583 to 0.929, matching or exceeding every larger configuration while cutting cost per correct answer by an order of magnitude. All harness code, prompts, gold answers, frozen packs and per-cell results are publicly available.
Chinese Translation
基于大型语言模型(LLMs)构建的代理越来越多地通过模型上下文协议(MCP)访问企业数据,许多 MCP 数据库服务器通过暴露单一通用 SQL 执行工具来最大化灵活性。本文提出了面向领域的工具模式:模型不再在查询时生成 SQL,而是从一小组与领域对齐的工具中进行选择,这些工具的参数化查询封装了服务器端的模式导航、连接和业务规则。我们围绕三个架构不变性对该模式进行了形式化,并引入了模型降级(Model Demotion),即用意图分类替代 SQL 合成可以降低服务常规请求所需的模型层级。作为参考实现,我们提出了 MCP Blueprint,这是一个开源框架,其中领域工具被声明性地定义为 YAML 元数据加上外部参数化 SQL 文件。我们通过一个公共可重复性基准评估该模式,比较三种 MCP 服务器设计——原始 SQL 执行、一个薄的通用工具包和一个垂直化领域工具包——在四个本地模型(3B-8B)上,针对十七个面向客户的任务,使用 Sakila 数据库(609 个完成单元;温度 0;每个单元三次重复)。垂直化工具包的平均得分为 0.939,而原始 SQL 为 0.666,通用工具包为 0.605;最小模型的得分从 0.583 提升至 0.929,匹配或超过每个更大配置,同时将每个正确答案的成本降低了一个数量级。所有的代码、提示、标准答案、冻结包和每个单元的结果均已公开。
cs.AI / 65 / 2608.22068

Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays

基于大型语言模型的地热井阵决策支持与建模
Ouko, Edwin, Lujan, Emmanuel, Edelman, Alan, Metcalfe, Robert
Abstract
Geothermal well arrays, which organize multiple geothermal wells into carefully planned geometric configurations, provide opportunities to enhance energy production capacity and increase fault tolerance. The development and adoption of these emerging geothermal technologies could be accelerated through the recent advances in large language models (LLMs) and high-level high-performance languages. A challenge in LLM-based applications is the reliability of the generated outputs, as they can be prone to subjective biases and hallucinations. This study assesses the potential of cutting-edge LLMs - such as ChatGPT, Gemini, Claude, Grok, and domain-specific models like AskGDR - as expert assistants that can synthesize insightful interpretations of complex geothermal data, as well as improve feature capabilities of geothermal models and numerical software. We developed a novel approach, leveraging Google's recently introduced AI assistant, NotebookLM, to accelerate the generation of unpublished quantitative geothermal benchmarks. The rapid generation of these evaluation instruments is essential for assessing the swiftly evolving capabilities of emerging language model technologies. In particular, we use these benchmarks and LLM-based interviews to analyze opportunities and limitations of two promising technologies: geothermal well arrays and closed-loop coaxial wells. Furthermore, we present a case study illustrating how LLMs can facilitate auto-parallelization of geothermal numerical models. Our analysis emphasizes their application in digital twins and underscores the importance of high-level, high-performance code generation. This line of research could play a transformative role in the geothermal sector by enabling the next-generation of decision-support applications, integrating data analysis, informed recommendations, and more dynamic numerical modeling workflows.
Chinese Translation
地热井阵将多个地热井组织成精心规划的几何配置,为提高能源生产能力和增加故障容忍度提供了机会。这些新兴地热技术的发展和应用可以通过近期在大型语言模型(LLMs)和高级高性能语言方面的进展得到加速。基于LLM的应用面临的一个挑战是生成输出的可靠性,因为它们可能受到主观偏见和幻觉的影响。本研究评估了前沿LLM(如ChatGPT、Gemini、Claude、Grok以及领域特定模型如AskGDR)作为专家助手的潜力,这些模型能够综合复杂地热数据的深刻解释,并改善地热模型和数值软件的特征能力。我们开发了一种新方法,利用谷歌最近推出的AI助手NotebookLM,加速生成未发表的定量地热基准。这些评估工具的快速生成对于评估新兴语言模型技术快速发展的能力至关重要。特别是,我们使用这些基准和基于LLM的访谈来分析两种有前景技术的机会与局限性:地热井阵和闭环同轴井。此外,我们还展示了一个案例研究,说明LLM如何促进地热数值模型的自动并行化。我们的分析强调了它们在数字双胞胎中的应用,并强调了高水平、高性能代码生成的重要性。这一研究方向可能在地热领域发挥变革性作用,通过实现下一代决策支持应用,整合数据分析、知情建议以及更动态的数值建模工作流程。
cs.AI / 66 / 2608.22085

Dissecting Neuro-Symbolic Quality Assurance for Synthetic Oncology Data Generation

剖析合成肿瘤数据生成的神经符号质量保证
Challa, Laxmigayathri, Zhou, Yuhan, Cleveland, Ana, Chen, Haihua
Abstract
Synthetic clinical data generation with large language models addresses the scarcity that limits cancer staging research, but oncology hallucinations are categorically harmful: one clinically impossible staging assignment contaminates every downstream model trained on it. Neuro-symbolic pipelines validate during generation, yet the contribution of individual quality-assurance components remains unclear. We report three controlled studies isolating gate necessity, constraint attribution, and retrieval conditionality, holding generation protocol, diversity thresholds, and fine-tuning hyperparameters constant across adapter conditions. The symbolic gate enforces schema completeness, ontology coverage against the Systematized Nomenclature of Medicine, and staging-logic consistency under American Joint Committee on Cancer eighth-edition rules. Ungated, 29.9% of records contain schema failures and 20.1% contain clinically invalid staging. Schema validation is the load-bearing filter: within the fully gated corpus it rejects 148 of 512 records, ontology grounding a further 24, and staging-logic validation none---the only generator producing logic violations is already excluded on schema, making clinical-logic validation a generator-conditional safeguard rather than the dominant filter. Retrieval augmentation is strongly model-dependent: it improves gate compliance for one generator by 12.5 percentage points, has no measurable effect for a second, and collapses output in a third. Across gated configurations ontology density is largely unchanged, indicating that symbolic validation improves clinical validity rather than vocabulary richness. Symbolic gating therefore buys corpus validity but no commensurate gain on real lung-cancer notes in this study; retrieval should be evaluated per model, and ontology density should not be reported as a proxy for corpus quality.
Chinese Translation
利用大型语言模型生成合成临床数据可以解决限制癌症分期研究的稀缺性,但肿瘤学幻觉是根本有害的:一个临床上不可能的分期分配会污染基于此训练的每一个下游模型。神经符号管道在生成过程中进行验证,但各个质量保证组件的贡献仍不清楚。我们报告了三项控制研究,隔离了门控必要性、约束归因和检索条件性,同时在适配器条件下保持生成协议、多样性阈值和微调超参数不变。符号门控强制执行模式完整性、与医学系统化命名法的本体覆盖,以及根据美国癌症联合委员会第八版规则的一致性分期逻辑。在没有门控的情况下,29.9%的记录包含模式失败,20.1%包含临床无效的分期。模式验证是承重过滤器:在完全门控的语料库中,它拒绝了512条记录中的148条,进一步通过本体基础拒绝了24条,而分期逻辑验证则没有拒绝任何记录——唯一产生逻辑违规的生成器在模式上已经被排除,使得临床逻辑验证成为生成器条件下的保护措施,而非主要过滤器。检索增强在很大程度上依赖于模型:它为一个生成器提高了12.5个百分点的门控合规性,对第二个生成器没有可测量的影响,而在第三个生成器中则导致输出崩溃。在门控配置中,本体密度基本保持不变,表明符号验证改善了临床有效性,而不是词汇丰富性。因此,符号门控为语料库的有效性提供了保障,但在本研究中并未带来与真实肺癌记录相应的收益;检索应根据模型进行评估,本体密度不应作为语料库质量的代理指标。
cs.AI / 67 / 2608.22103

Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks

可验证黑客终端基准:评估终端任务中的奖励黑客行为
Roth, Amit, Bercovich, Ivan, Efroni, Yonathan
Abstract
As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode. Measuring reward hacking is itself challenging, as detection typically relies on human inspection or LLM judges, both of which can be unreliable. The hack-verifiable environments (HVE) methodology addresses this challenge by embedding detectable hacks into tasks, allowing reward hacks to be identified automatically and reliably. In this work, we adapt HVE to Terminal Bench, a leading benchmark of real-world terminal and coding tasks, and introduce Hack-Verifiable Terminal Bench (HVTB). Using HVTB, we measure reward-hacking rates across frontier models and study whether prompts with varying amounts of information on the hack can mitigate this behavior. This lets us test whether prompting can prevent not only known reward-hacking strategies, but also 'unknown unknown' exploits that the prompt does not anticipate. We release all environments and agent traces at https://majoroth.github.io/hack-verifiable-environments/hvtb
Chinese Translation
随着智能体变得越来越强大和自主,它们奖励黑客的倾向,即满足任务检查的同时违反其意图,成为一种日益重要的失败模式。衡量奖励黑客行为本身具有挑战性,因为检测通常依赖于人工检查或大型语言模型(LLM)评审,这两者都可能不可靠。可验证黑客环境(Hack-Verifiable Environments, HVE)方法通过将可检测的黑客行为嵌入任务中,解决了这一挑战,从而使奖励黑客行为能够被自动和可靠地识别。在本研究中,我们将HVE方法适配到终端基准(Terminal Bench),这是一个领先的现实世界终端和编码任务基准,并引入可验证黑客终端基准(Hack-Verifiable Terminal Bench, HVTB)。通过使用HVTB,我们测量了前沿模型中的奖励黑客率,并研究了不同信息量的提示是否能够减轻这种行为。这使我们能够测试提示是否可以防止已知的奖励黑客策略,以及提示未预见的“未知未知”利用方式。我们在 https://majoroth.github.io/hack-verifiable-environments/hvtb 发布了所有环境和智能体轨迹。
cs.AI / 68 / 2608.22108

Development and Feasibility Evaluation of an Edge AI as Medical Device System for Breast Cancer Multidisciplinary Team Meetings

作为医疗设备系统的边缘人工智能的开发与可行性评估:针对乳腺癌多学科团队会议
Dhiman, Aarzoo, Haque, Farzana, Grover, Kartikae, Smith, Lydia Brian, Jones, William Stephen
Abstract
Breast Cancer Multidisciplinary Team (MDT) meetings manage increasingly complex cases under considerable time pressure, and documentation requirements can reduce clinical efficiency and decision quality. Existing AI based MDT workflows rely on cloud-based processing, limiting their use because patient discussions contain identifiable information. We developed a fully on-device AI pipeline using open-source Automatic Speech Recognition (ASR) and Large Language Models (LLMs) that transcribes breast cancer MDT discussions, structures clinical information, and generates treatment recommendations using retrieval-augmented generation (RAG) grounded in National Institute for Health and Care Excellence (NICE) guidance. The pipeline runs on a single NVIDIA Jetson AGX Orin, ensuring that patient audio, transcripts, and outputs remain within institutional infrastructure. Evaluation included two recorded simulated MDT discussions, ten clinically validated synthetic discussions, and 1,270 acoustically augmented recordings. Optimisation of Whisper large-v3 reduced word error rate by 20.7% and 24.4% on the recorded discussions and achieved performance within 0.58% WER and 1.58% word information lost of a commercial clinical ASR benchmark on augmented audio. MedGemma-RAG identified 2.3 times more MDT-concordant interventions than a proprietary cloud comparator (p = 0.020), with no significant difference in overall accuracy. Stakeholders identified automated documentation, treatment recommendation support, and case triage as the most credible near-term applications while highlighting workflow integration, governance, and clinician trust as key implementation challenges. These findings demonstrate the feasibility of privacy-preserving, fully on-device AI for MDT documentation and guideline-informed decision support, providing a foundation for prospective clinical evaluation.
Chinese Translation
乳腺癌多学科团队(MDT)会议在相当大的时间压力下管理日益复杂的病例,而文档要求可能降低临床效率和决策质量。现有的基于人工智能的MDT工作流程依赖于云端处理,限制了其使用,因为患者讨论中包含可识别的信息。我们开发了一个完全在设备上运行的人工智能管道,使用开源的自动语音识别(ASR)和大型语言模型(LLMs),能够转录乳腺癌MDT讨论,结构化临床信息,并基于国家卫生与护理卓越研究所(NICE)指南生成治疗建议,采用检索增强生成(RAG)的方法。该管道在单个NVIDIA Jetson AGX Orin上运行,确保患者音频、转录和输出保持在机构基础设施内。评估包括两个录制的模拟MDT讨论、十个临床验证的合成讨论和1270个声学增强录音。对Whisper large-v3的优化使录制讨论的词错误率降低了20.7%和24.4%,并在增强音频上实现了0.58%的词错误率和1.58%的词信息丢失,符合商业临床ASR基准。MedGemma-RAG识别出的MDT一致干预措施比专有云比较器多出2.3倍(p = 0.020),而整体准确性没有显著差异。利益相关者认为自动文档、治疗建议支持和病例分流是最可信的近期应用,同时强调工作流程整合、治理和临床医生信任是关键实施挑战。这些发现证明了隐私保护的完全设备内人工智能在MDT文档和基于指南的决策支持中的可行性,为未来的临床评估奠定了基础。
cs.AI / 69 / 2608.22128

Task-Driven 3D Printability Assistance via Geometry- and Knowledge-Grounded LLM Reasoning

基于几何和知识驱动的大型语言模型推理的任务导向3D打印可打印性辅助
Du, Zhaoda, Zheng, Qiaojie, Zhang, Xiaoli
Abstract
Printability assessment in additive manufacturing is typically conducted at the geometry level before printing to determine whether a computer-aided design (CAD) model or stereolithography (STL) file can be successfully fabricated. Task suitability, in contrast, is usually evaluated after printing to determine whether the fabricated part satisfies the requirements of its intended use. As a result, for non-expert users to print functional parts, unsuitable material or process choices may only be identified after fabrication, leading to repeated printing, material waste, and user frustration. To address this challenge, this paper leverages the reasoning and language-understanding capabilities of large language models (LLMs), while grounding the reasoning with geometry evidence and structured material/printer knowledge to generate reliable pre-print recommendations. Given a stereolithography (STL) model and a natural-language task description, the framework generates a structured recommendation covering printability, material choice, process parameters, design guidance, risks, and explanations. We evaluate the framework on focused STL benchmark scenarios with novice-style task descriptions. The proposed method achieves 75.0% printability over 96 physical validation trials, with 88.9% task suitability among successfully printed samples. It also improves Gemini 2.5 Flash-Lite material-selection accuracy from 37.5% under pure LLM to 90.0%. Expert evaluation further shows improved report quality, while post-print feedback improves recommendations on selected problematic cases. These results suggest that user task intent, geometry evidence, and structured material knowledge are all important for reliable task-driven printability assistance.
Chinese Translation
在增材制造中,打印可行性评估通常在打印前于几何层面进行,以确定计算机辅助设计(CAD)模型或立体光刻(STL)文件是否能够成功制造。相比之下,任务适用性通常在打印后评估,以确定制造的部件是否满足其预期用途的要求。因此,对于非专业用户来说,打印功能部件时,不合适的材料或工艺选择可能仅在制造后被识别,导致重复打印、材料浪费和用户挫败感。为了解决这一挑战,本文利用大型语言模型(LLMs)的推理和语言理解能力,同时结合几何证据和结构化的材料/打印机知识,生成可靠的打印前建议。给定一个立体光刻(STL)模型和一个自然语言任务描述,该框架生成一个涵盖可打印性、材料选择、工艺参数、设计指导、风险和解释的结构化建议。我们在针对新手风格任务描述的聚焦STL基准场景上评估了该框架。所提出的方法在96次物理验证试验中实现了75.0%的可打印性,在成功打印的样本中任务适用性达88.9%。它还将Gemini 2.5 Flash-Lite的材料选择准确率从纯LLM下的37.5%提高到90.0%。专家评估进一步显示报告质量的改善,而打印后的反馈则改善了对选定问题案例的建议。这些结果表明,用户任务意图、几何证据和结构化材料知识对于可靠的任务导向打印可行性辅助都是重要的。
cs.AI / 70 / 2608.22137

MegaMem: A Retrieval Solution for Ultra-Large Context Windows

MegaMem:一种超大上下文窗口的检索解决方案
Song, Xinyuan, Zhu, Bowen, Haque, Hasibul, Zhao, Liang
Abstract
Modern language models and agents increasingly require persistent memory for complete codebases, long interaction histories, and heterogeneous enterprise records. The key challenge is to keep hundreds of millions of tokens searchable while passing only bounded source evidence to the answer model. We introduce MegaMem, a source-resolved dual-view retrieval system that separates semantic access from generation evidence. Distilled records and detailed evidence are searched with original and transformed queries; every distilled hit resolves to an immutable source ID before reciprocal-rank fusion, deduplication, and cross-encoder reranking; and only the highest-ranked detailed evidence within a fixed budget supports generation. Post-answer attribution then identifies which loaded sources support the fixed answer. We evaluate MegaMem on EnterpriseRAG-Bench, which contains more than 500,000 heterogeneous enterprise documents and approximately 650M tokens. MegaMem improves Overall from 68.22 to 82.26 and reaches 86.50 Correctness. These results show that MegaMem supports ultra-large persistent memory while preserving strong answer accuracy under a bounded generation context. By separating searchable memory scale from answer-context size, MegaMem provides a practical path toward accurate retrieval over memories ranging from hundreds of millions to one billion tokens. Our code is available at https://github.com/ xfab-xinyuansong/MegaMem.git.
Chinese Translation
现代语言模型和智能体越来越需要持久内存,以存储完整的代码库、长时间的交互历史和异构企业记录。关键挑战在于如何保持数亿个标记可搜索,同时仅将有限的源证据传递给答案模型。我们提出了MegaMem,一种源解析的双视图检索系统,旨在将语义访问与生成证据分离。提炼的记录和详细证据通过原始和变换的查询进行检索;每个提炼的命中在进行互惠排名融合、去重和交叉编码重排序之前,都会解析为一个不可变的源ID;而且,只有在固定预算内排名最高的详细证据才支持生成。答案后的归因则识别出哪些加载的源支持固定答案。我们在EnterpriseRAG-Bench上评估了MegaMem,该基准包含超过500,000个异构企业文档和大约6.5亿个标记。MegaMem的整体评分从68.22提高到82.26,正确率达到86.50。这些结果表明,MegaMem支持超大持久内存,同时在有限的生成上下文下保持强大的答案准确性。通过将可搜索内存规模与答案上下文大小分离,MegaMem为在数亿到十亿个标记范围内实现准确检索提供了一条切实可行的路径。我们的代码可在https://github.com/xfab-xinyuansong/MegaMem.git获取。
cs.AI / 71 / 2608.22138

Measuring Stability and Failure Behavior in Language Models Under Structured Perturbations

在结构性扰动下测量语言模型的稳定性和失败行为
Golsefid, Samira
Abstract
Language models are usually judged by a single accuracy score, which does not reveal how their performance degrades as inputs are perturbed. We present a graded, multi-family, failure-aware framework for stress-testing reasoning models. It perturbs each problem along a multi-level severity ladder across seven families: six that preserve the answer, paraphrase, input noise, formatting, irrelevant context, context load, and conflicting instructions, and a Knowledge Boundary family that removes answerability so that refusal becomes the correct response. Every test is validity-gated and labeled by its measured severity, and each model is summarized by per-level Accuracy, a magnitude-weighted Stability, and a per-family Collapse Point defined relative to the model's own baseline. Instantiated on the same 100 seed problems used by GSM-Symbolic, expanded into 4,473 gated tests and run on four models spanning capability tiers, the framework exposes structure that an aggregate score hides: the level at which a model fails is family-specific rather than global, and two stressors expose consistent weaknesses across all models: conflicting instructions and questions built on an impossible premise. Recognition of unanswerability is otherwise uneven, reliable on missing information and fabricated evidence but weak on impossible premises. These failure points are invisible to standard accuracy reporting.
Chinese Translation
语言模型通常通过单一的准确率评分来评判,这并不能揭示在输入受到扰动时其性能如何下降。我们提出了一种分级的、多家族的、关注失败的框架,用于对推理模型进行压力测试。该框架沿着多级严重性梯度对每个问题进行扰动,涵盖七个家族:六个保持答案的家族,包括同义改写、输入噪声、格式化、无关上下文、上下文负载和冲突指令,以及一个知识边界家族,该家族移除可回答性,使得拒绝成为正确的回应。每个测试都有有效性门控,并根据其测量的严重性进行标记,每个模型通过每级准确率、幅度加权的稳定性和相对于模型自身基线的每个家族崩溃点进行总结。该框架在与GSM-Symbolic使用的相同100个种子问题上实例化,扩展为4,473个门控测试,并在四个跨越能力层级的模型上运行,揭示了一个聚合评分所隐藏的结构:模型失败的级别是特定于家族的,而非全局性的,两个压力源在所有模型中揭示了一致的弱点:冲突指令和基于不可能前提的问题。对不可回答性的识别则不均衡,依赖于缺失信息和虚构证据,但在不可能前提上表现较弱。这些失败点在标准准确率报告中是不可见的。
cs.AI / 72 / 2608.22141

MEMONDEMAND: A Memory Management System for Large-Scale Enterprise Data

MEMONDEMAND:大规模企业数据的内存管理系统
Song, Xinyuan, Zhu, Bowen, Haque, Hasibul, Zhao, Liang
Abstract
Enterprise repositories are large, heteroge- neous, and continuously updated, making re- trieval difficult when efficient access, source- faithful evidence, and cross-query adaptation must be supported together. Enterprise mem- ory extends retrieval beyond the model con- text, but existing systems do not jointly address collection-specific hierarchy construction, low- cost routing, detailed evidence loading, and workload-aware memory updates at this scale. We introduce MEMONDEMAND, short for On- Demand Memory, a memory management sys- tem with three coordinated mechanisms: a dy- namic multi-level hierarchy that determines the abstraction structure and depth for each col- lection, dual memory at every hierarchy level that separates distilled routing from detailed evidence, and on-demand memory promotion that updates node priority under a bounded active-state budget. On EnterpriseRAG-Bench, MEMONDEMAND outperforms the strongest published LB#1 result at every evaluated scale from 10M tokens through the complete 618M- token collection, with gains of 12.23% at 10M and 4.66% at 618M. Results on FinanceBench, HotpotQA, and FRAMES further show strong performance across financial, multi-hop, and fact-retrieval settings. Together, these results establish MEMONDEMAND as an accurate, ef- ficient, and scalable memory solution for very large enterprise repositories across data scales, domains, and evidence requirements. Our code is available at https://github.com/ xfab-xinyuansong/MemOnDemand.git.
Chinese Translation
企业存储库庞大、异构且持续更新,这使得在需要高效访问、源可信证据和跨查询适应性共同支持时,检索变得困难。企业内存扩展了检索超出模型上下文,但现有系统未能在此规模上共同解决特定集合的层次结构构建、低成本路由、详细证据加载和工作负载感知的内存更新。我们提出了MEMONDEMAND(按需内存),这是一个具有三个协调机制的内存管理系统:一个动态多级层次结构,确定每个集合的抽象结构和深度;每个层次级别的双重内存,将提炼的路由与详细证据分开;以及按需内存提升,在有限的活动状态预算下更新节点优先级。在EnterpriseRAG-Bench上,MEMONDEMAND在从10M标记到完整的618M标记集合的每个评估规模上均超越了已发布的最强LB#1结果,在10M时提升了12.23%,在618M时提升了4.66%。在FinanceBench、HotpotQA和FRAMES上的结果进一步显示了在金融、多跳和事实检索环境下的强大性能。这些结果共同确立了MEMONDEMAND作为一种准确、高效且可扩展的内存解决方案,适用于不同数据规模、领域和证据要求的非常大规模企业存储库。我们的代码可在https://github.com/xfab-xinyuansong/MemOnDemand.git获取。
cs.AI / 73 / 2608.22143

Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems

小型视觉-语言模型在定性机械问题上的评估
Ansah, Henry Fordjour, Banerjee, Shreya, Ghimire, Pranish
Abstract
Qualitative mechanical problem-solving (QMPS) refers to solving qualitative problems from the mechanical domain. Qualitative problems can be solved with minimal discipline-specific information, without any robust quantitative calculation, generally by using qualitative reasoning and commonsense knowledge. QMPS is a vital aspect of human intelligence that allows us to tackle a wide range of tasks, from simple everyday ones such as turning on a tap to complex tasks in highly demanding and well-paying jobs in various fields, e.g., emergency medicine, plumbing, driving, etc. Employers often use the Bennett Mechanical Comprehension Test (BMCT) to evaluate job candidates' ability to solve such problems. In this work, we assess two state-of-the-art multimodal models, Gemma-3 and Qwen-VL, on their ability to interpret mechanical problem images by eliciting a step-by-step chain of thought (CoT) and a final answer. Each image inherently encodes ground-truth qualitative facts, such as contact points in gears, support relations, and relative weights, which we use to evaluate each model's spatial and commonsense reasoning capabilities. We assess each chain for coherence, completeness, and logical progression to assess each model's thought process, and final answers are compared to verified solutions to measure accuracy.
Chinese Translation
定性机械问题解决(QMPS)是指解决机械领域中的定性问题。定性问题可以在缺乏任何强有力的定量计算和最少的学科特定信息的情况下,通过定性推理和常识知识来解决。QMPS是人类智能的一个重要方面,使我们能够处理广泛的任务,从简单的日常任务(如打开水龙头)到在各个领域(如急救医学、管道工、驾驶等)中高要求和高薪工作的复杂任务。雇主通常使用贝内特机械理解测试(BMCT)来评估求职者解决此类问题的能力。在本研究中,我们评估了两种最先进的多模态模型,Gemma-3和Qwen-VL,评估它们通过引导逐步思维链(CoT)和最终答案来解释机械问题图像的能力。每个图像本质上编码了真实的定性事实,例如齿轮中的接触点、支撑关系和相对重量,我们利用这些信息来评估每个模型的空间和常识推理能力。我们评估每个思维链的一致性、完整性和逻辑进展,以评估每个模型的思维过程,并将最终答案与经过验证的解决方案进行比较,以测量准确性。
cs.AI / 74 / 2608.22160

AUDITA: certified auditing and causal attribution of adverse outcomes in autonomous multi-agent systems

AUDITA:自主多代理系统中不良结果的认证审计与因果归因
Du, Zhixu, Chen, Yiran
Abstract
Physical automation is scaling toward fleets of embodied machines commanded by an AI brain. Early deployments already run factories and warehouses at production rates beyond any human line, and their adoption is accelerating. But when their joint decisions cause harm, everyone involved has reason to blame everyone else, the machine vendor, the algorithm provider, the factory operator, the insurer, and the regulator, and no method can divide the responsibility between them. Existing methods read logs whose origin they cannot verify and name a single culprit, misrepresenting outcomes that are overdetermined, preempted, or caused by an omission. We present \audita{}, an audit layer pairing a tamper-evident record of every inter-agent command with a certified, graded causal-attribution engine. We prove its verdict cannot be gamed: a rule-following agent can never be made to look guilty, an attempt to shift blame is itself caught and graded, and we establish the exact limit of what an evidence-based auditor can certify. On live language-model pipelines it reduces the standard judge baseline's responsibility error roughly threefold; on a benchmark of accident-grounded structures it recovers responsibility where single-culprit baselines fail, and stays invariant under forgery. \audita{} turns the question of who is to blame from an argument about logs into a calculation over evidence.
Chinese Translation
物理自动化正在向由人工智能大脑指挥的具身机器舰队扩展。早期部署已经在生产率超越任何人类生产线的工厂和仓库中运行,并且其采用速度正在加快。但是,当它们的联合决策造成伤害时,所有相关方都有理由相互指责,包括机器供应商、算法提供者、工厂操作员、保险公司和监管机构,而没有任何方法能够在他们之间划分责任。现有方法读取其来源无法验证的日志,并指认单一罪犯,错误地描述了过度确定、被预先阻止或因遗漏而造成的结果。我们提出了 extit{AUDITA},一个审计层,将每个代理间命令的防篡改记录与一个认证的、分级的因果归因引擎配对。我们证明其裁决无法被操控:遵循规则的代理永远无法被看作有罪,试图转移责任的行为本身会被捕捉并进行评分,我们确立了基于证据的审计员可以认证的确切限制。在实时语言模型管道上,它将标准法官基线的责任错误减少了大约三倍;在一个基于事故的结构基准上,它恢复了责任,而单一罪犯基线则失败,并且在伪造情况下保持不变。 extit{AUDITA} 将谁应承担责任的问题从关于日志的争论转变为对证据的计算。
cs.AI / 75 / 2608.22161

Aggregation-Aware Synthetic Text Generation Against Authorship Re-Identification

聚合感知的合成文本生成以应对作者身份重新识别
Ma, Qian, Squicciarini, Anna, Rajtmajer, Sarah
Abstract
Online users often release multiple texts under the same identity, giving attackers an author profile that can reveal more than any single text. Existing authorship obfuscation methods optimize privacy independently for each document, leaving them blind to cross-document correlations that make aggregation dangerous. We propose Aggregation-Aware Synthetic Text Generation (AAST), a framework that addresses this gap by jointly selecting synthetic texts at the bundle level rather than optimizing each text in isolation. AAST targets attribution and verification attacks, including cross-genre settings where attacker references come from a genre not observed during generation or selection. Experiments across same-genre, cross-genre, neural, and independent non-neural stylometric attacks show that AAST lowers account-level linkability as bundle size grows, while preserving semantic quality, linguistic acceptability, and sentiment alignment.
Chinese Translation
在线用户通常在同一身份下发布多个文本,这为攻击者提供了一个作者档案,可能揭示比任何单一文本更多的信息。现有的作者身份模糊化方法独立优化每个文档的隐私,忽视了跨文档相关性,这使得聚合变得危险。我们提出了聚合感知的合成文本生成(Aggregation-Aware Synthetic Text Generation, AAST)框架,以捆绑级别联合选择合成文本,从而填补这一空白,而不是孤立地优化每个文本。AAST 针对归属和验证攻击,包括跨类型设置,其中攻击者的参考来自于生成或选择过程中未观察到的类型。针对同类型、跨类型、神经和独立非神经风格分析攻击的实验表明,随着捆绑大小的增加,AAST 降低了账户级别的可链接性,同时保持了语义质量、语言可接受性和情感一致性。
cs.AI / 76 / 2608.22167

MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning

MCP-Universe RL:通过强化学习训练MCP工具使用代理的框架
Luo, Ziyang, Yang, Yan, Jian, Xiangru, Shi, Ziji, Lin, Xiaoqiang, Liew, Jun Hao, Savarese, Silvio, Li, Junnan
Abstract
Reinforcement learning (RL) has become an effective way to improve the tool-use ability of large language models (LLMs), but most existing RL frameworks stop at the policy update. For every new domain, the user is left with two hard systems problems: standing up an isolated environment for each of hundreds of concurrent trajectories and connecting it to training, and scheduling the rollout so that the GPU stays busy across long, multi-turn episodes that spend much of their time stalled on slow tool calls. We present MCP-Universe RL (MCP-U RL), an open-source framework that takes over both. It uses the Model Context Protocol (MCP) as the interface to the environment, so any tool already exposed as an MCP server plugs into training with no RL-specific integration code. It builds the two missing layers once and reuses them across domains: an environment-orchestration layer that provisions, isolates, and recycles the MCP environments over a pluggable container backend, and a rollout-orchestration layer whose staged pipeline overlaps trajectories to keep the GPU busy while episodes wait on tools. A backend-agnostic training layer then applies the update through an existing RL backend, with veRL and slime integrations. With one configuration, changing only the task specification, we train software-engineering, deep-research, and general tool-use agents on gpt-oss-20b and improve task reward in all three.
Chinese Translation
强化学习(RL)已成为提升大型语言模型(LLMs)工具使用能力的有效方法,但大多数现有的RL框架仅停留在策略更新阶段。对于每个新领域,用户面临两个棘手的系统问题:为数百个并发轨迹中的每一个建立一个独立环境并将其连接到训练中,以及调度回滚以确保GPU在长时间的多轮剧集期间保持繁忙,而这些剧集大部分时间都在慢速工具调用上停滞。我们提出了MCP-Universe RL(MCP-U RL),一个开源框架,接管了这两个问题。它使用模型上下文协议(Model Context Protocol,MCP)作为与环境的接口,因此任何已作为MCP服务器暴露的工具都可以无缝接入训练,无需特定于RL的集成代码。它构建了两个缺失的层,并在不同领域中重复使用:一个环境协调层,负责提供、隔离和回收MCP环境,基于可插拔的容器后端;一个回滚协调层,其分阶段管道重叠轨迹,以保持GPU在剧集等待工具时的繁忙状态。然后,一个与后端无关的训练层通过现有的RL后端应用更新,支持veRL和slime集成。通过一次配置,仅更改任务规范,我们在gpt-oss-20b上训练了软件工程、深度研究和通用工具使用代理,并在这三者中提高了任务奖励。
cs.AI / 77 / 2608.22176

Role-Specialized Mixture-of-Agents with Open-Weight LLMs for Clinical Prediction

基于角色专用混合代理的开放权重大语言模型在临床预测中的应用
Hou, Jun, Fang, Yi, Wang, Xuan
Abstract
Large Language Models (LLMs) are increasingly applied to clinical prediction tasks such as in-hospital mortality and readmission from electronic health records (EHRs). Privacy and compliance constraints motivate systems that can be deployed locally, which has increased interest in open-weight multi-agent designs. However, most medical multi-agent systems are evaluated as a single block, leaving unclear which agent role contributes to prediction and whether retrieval drives observed gains. We study a role-specialized Mixture-of-Agents (MoA) that combines medical knowledge retrieval with contrastive similar-patient reasoning. By varying the role design while holding the retrieval setup fixed, we localize the main effect to the final integrator. Pairing large open-weight analysts with a small open-weight integrator matches closed-model prompting on F1 for mortality prediction while flagging substantially more true high-risk patients. Mechanism analysis shows the role assignment directly yields a high-recall operating point without threshold tuning. The effect is task-dependent, with smaller gains for readmission because the available records correlate weakly with this longer-horizon outcome. These results position role design as a key factor in privacy-constrained, training-free clinical LLM prediction.
Chinese Translation
大型语言模型(LLMs)越来越多地应用于临床预测任务,如医院内死亡率和电子健康记录(EHRs)中的再入院。隐私和合规性限制促使开发可以在本地部署的系统,从而增加了对开放权重多代理设计的兴趣。然而,大多数医学多代理系统作为一个整体进行评估,尚不清楚哪个代理角色对预测贡献最大,以及检索是否推动了观察到的收益。我们研究了一种角色专用的混合代理(Mixture-of-Agents, MoA),该模型结合了医学知识检索与对比相似患者推理。通过在保持检索设置固定的情况下改变角色设计,我们将主要效果定位于最终整合器。将大型开放权重分析师与小型开放权重整合器配对,在死亡率预测中与封闭模型提示的F1指标相匹配,同时标记出显著更多的真实高风险患者。机制分析表明,角色分配直接产生了高召回率的操作点,而无需阈值调整。该效应依赖于任务,对于再入院的收益较小,因为可用记录与这一长期结果的相关性较弱。这些结果将角色设计定位为在隐私受限、无训练的临床LLM预测中的关键因素。
cs.AI / 78 / 2608.22191

Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents

不同意探索,同意承诺:面向路由的测试时缩放用于软件代理
Chen, Kang, Nian, Junjie, Cao, Yixin, Jiang, Yugang
Abstract
Software-engineering agents solve repository-level tasks through long, stochastic tool-use trajectories, and repeated attempts often find fixes missed by one run. Test-time scaling is difficult because patches lack canonical answer forms, while sibling actions from a shared prefix are correlated. We study whether native MoE router traces can guide steering and selection without an external judge or selection-time test execution. Our analysis shows that routing provides a robust behavioral role signal; token-granular readouts and decision-matched comparison sets turn it into effective control. We therefore introduce Risa (Routing-Informed Steering and Arbitration): within trajectories, routing encourages diverse exploration and controlled convergence during patch commitment; across separately sampled trajectories, agreement at informative patch positions selects a final candidate. We evaluate on SWE-bench Verified using open-weight sparse MoE agents across scales and reasoning-effort settings. Risa's routing arbitration raises the macro-average resolved rate from 44.9% under uniform sampling to 48.2% on the gpt-oss family, matching text consensus without answer-string matching, and it transfers to Qwen3.6, where it improves on uniform choice and matches text consensus on the full 500-task benchmark.
Chinese Translation
软件工程代理通过长时间的随机工具使用轨迹解决库级任务,重复尝试往往能找到一次运行中遗漏的修复。测试时缩放是困难的,因为补丁缺乏规范的答案形式,而来自共享前缀的兄弟操作是相关的。我们研究了本地 MoE(Mixture of Experts)路由器轨迹是否可以在没有外部评判者或选择时测试执行的情况下指导引导和选择。我们的分析表明,路由提供了一个稳健的行为角色信号;基于令牌粒度的读出和与决策匹配的比较集使其成为有效的控制。因此,我们引入了 Risa(Routing-Informed Steering and Arbitration):在轨迹内,路由鼓励多样化探索和在补丁承诺期间的受控收敛;在单独采样的轨迹之间,在信息丰富的补丁位置达成一致选择最终候选。我们在 SWE-bench Verified 上进行评估,使用开放权重稀疏 MoE 代理在不同规模和推理努力设置下进行测试。Risa 的路由仲裁将宏观平均解决率从均匀采样下的 44.9% 提高到 gpt-oss 系列上的 48.2%,在不进行答案字符串匹配的情况下匹配文本共识,并且它可以迁移到 Qwen3.6,在那里它改善了均匀选择,并在完整的 500 任务基准上与文本共识相匹配。
cs.AI / 79 / 2608.22214

Query-Driven Multimodal Information Extraction from Long Documents

基于查询驱动的长文档多模态信息提取
Gao, Yikai, Xia, Ding, Yang, Xi
Abstract
In domain-specific multimodal long documents, images and text jointly convey complex knowledge that cannot be fully captured by plain text alone. However, existing paradigms like DocVQA primarily focus on generating textual answers or localizing evidence regions, rather than outputting query-specific textual attribute values and corresponding images. To address this gap, we propose query-driven image-text joint extraction from long documents, requiring models to output query-requested textual attribute values and corresponding image bounding boxes. Based on challenges related to both user intent and document content, we designed a two-level taxonomy that operates at the query and instance levels. Further, we construct ITJoint, the first high-quality, manually annotated benchmark for this new task, comprising 2,455 pages of domain-specific documents with numerous non-decorative images, 316 queries, and 910 answer instances. Finally, we evaluate representative standalone Vision-Language Models from different providers and further design Q2IT, a multi-agent collaborative framework consisting of three progressively collaborating agents for evidence collection, page selection, and target-image localization. Using a joint evaluation approach that assesses both text extraction and image localization, our experiments show that standalone VLMs struggle with this task, while Q2IT significantly improves performance on ITJoint, although a substantial gap remains toward perfect results.
Chinese Translation
在特定领域的多模态长文档中,图像和文本共同传达复杂的知识,这些知识无法仅通过普通文本完全捕捉。然而,现有的范式如 DocVQA 主要集中于生成文本答案或定位证据区域,而不是输出特定于查询的文本属性值及其对应的图像。为了解决这一问题,我们提出了一种基于查询驱动的图像-文本联合提取方法,要求模型输出查询请求的文本属性值及相应的图像边界框。基于用户意图和文档内容相关的挑战,我们设计了一个在查询和实例层面运作的两级分类法。此外,我们构建了 ITJoint,这是首个高质量的手动标注基准,针对这一新任务,包含2455页领域特定文档,包含大量非装饰性图像,316个查询和910个答案实例。最后,我们评估了来自不同提供者的代表性独立视觉-语言模型,并进一步设计了 Q2IT,一个由三个逐步协作的代理组成的多代理协作框架,用于证据收集、页面选择和目标图像定位。通过一种联合评估方法,评估文本提取和图像定位,我们的实验表明,独立的视觉-语言模型在此任务上表现不佳,而 Q2IT 在 ITJoint 上显著提高了性能,尽管仍然存在与完美结果之间的显著差距。
cs.AI / 80 / 2608.22232

Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models

超越眼见:揭示多模态大语言模型的情境幻觉
Yang, Zhiming, Xiong, Zhuoxi, Zhou, Donglin, Wei, Wenjun, Cui, Shiyao, Shi, Jinqiao
Abstract
Real-world situation appearances can deviate from their underlying physical states, challenging the reliability of multimodal large language models (MLLMs) in practical applications. In this paper, we term this phenomenon situational illusions and investigate: (1) how MLLMs perform under such illusions, and (2) how to mitigate the limitations. We first develop a comprehensive where-what-how taxonomy that characterizes where situational illusions occur, what targets they take, and how they arise. Building on this taxonomy, we introduce MSIBench, a benchmark designed to assess the discrimination, understanding, and reasoning capabilities of MLLMs under situational illusions. Evaluations of 27 model configurations reveal that current MLLMs are highly vulnerable to these illusions and exhibit 6 typical failure modes related to visual observation, grounding, and reasoning. To mitigate the limitations, we build on the core idea of systematically inspecting and reasoning over visual evidence for contextual understanding, developing prompting for closed-source models and supervised fine-tuning for open-source models, respectively. These two simple yet effective methods improve model performances by 20% at most, suggesting a practical path toward more reliable multimodal perception and reasoning in complex real-world environments.
Chinese Translation
现实世界中的情境表现可能与其潜在的物理状态存在偏差,这对多模态大语言模型(MLLMs)在实际应用中的可靠性提出了挑战。本文将这一现象称为情境幻觉,并研究:(1)MLLMs在此类幻觉下的表现如何,以及(2)如何减轻其局限性。我们首先开发了一个全面的何处-何物-如何分类法,描述情境幻觉发生的地点、所涉及的目标以及其产生的方式。在此分类法的基础上,我们引入了MSIBench,一个旨在评估MLLMs在情境幻觉下的辨别、理解和推理能力的基准测试。对27种模型配置的评估显示,当前的MLLMs对这些幻觉高度脆弱,并表现出与视觉观察、基础知识和推理相关的6种典型失效模式。为了减轻这些局限性,我们基于系统检查和推理视觉证据以实现上下文理解的核心思想,分别为封闭源模型开发了提示方法,并为开放源模型进行了监督微调。这两种简单而有效的方法使模型性能最多提高了20%,为在复杂现实环境中实现更可靠的多模态感知和推理提供了切实可行的路径。
cs.AI / 81 / 2608.22237

Read Less, Solve More: Token-Efficient Sparse Reading for AI Agents

少读多解:面向AI代理的高效稀疏阅读
Liu, Zedong, Wu, Jiaan, Ma, Xinyang, Xu, Le, Wang, Kai, Hu, Yuanchao, Tao, Dingwen, Tan, Guangming
Abstract
Long-horizon agents increasingly rely on repeated access to external artifacts, yet current reading interfaces often expose entire objects even when only sparse evidence is needed. This over-reading increases token and latency costs and can dilute task-relevant evidence, while existing context-reduction methods mainly intervene after broad content has already entered the trajectory. We present SparseRead, a training-free, model-transparent reading layer that controls content admission before unnecessary evidence reaches the model context. SparseRead combines a regime-aware Read Gate, extensible Reader Backends, and a stateful protocol for bounded, source-anchored evidence acquisition with explicit refinement, verification, stopping, and fallback. Across six frontier models, including Claude Opus 5, and five workload scenarios, SparseRead reduces token volume by up to 92.9% and wall time by up to 89.0%, while preserving or improving task quality. Its consistent gains across three agent frameworks further demonstrate broad portability.
Chinese Translation
长时间跨度的代理越来越依赖于对外部文献的重复访问,然而当前的阅读接口往往会暴露整个对象,即使只需要稀疏证据。这种过度阅读增加了令牌和延迟成本,并可能稀释与任务相关的证据,而现有的上下文减少方法主要是在广泛内容已经进入轨迹后进行干预。我们提出了SparseRead,这是一种无训练、模型透明的阅读层,能够在不必要的证据到达模型上下文之前控制内容的接纳。SparseRead结合了一个状态感知的读取门(Read Gate)、可扩展的阅读后端(Reader Backends)以及一个有状态的协议,用于有界的、源锚定的证据获取,具备明确的细化、验证、停止和回退功能。在包括Claude Opus 5在内的六个前沿模型和五个工作负载场景中,SparseRead将令牌量减少了最多92.9%,将墙时(wall time)减少了最多89.0%,同时保持或提高了任务质量。它在三个代理框架中的一致性增益进一步证明了其广泛的可移植性。
cs.AI / 82 / 2608.22266

Clarify User Expertise: Towards Proactive Conversational Agents Tailoring Responses to User Proficiency

明确用户专业知识:朝着主动对话代理定制响应以适应用户能力
Cao, Zhihong, Huang, Chen
Abstract
In the context of information seeking, conversational agents are undergoing an evolution from reactive tools to proactive, personalized assistants. A critical aspect of this evolution is the ability to tailor strategic interactions to a user's unique needs and expectations. Unlike existing studies that focus on proactively clarifying query ambiguities, we center on clarifying the user's expertise in order to tailor responses for better user comprehension. We find that existing agents struggle to determine user expertise from queries alone, a limitation that prevents them from dynamically adapting their responses. To address this gap, we introduce PASSING to empower the agent to proactively clarify a user's expertise through targeted inquiries. This is achieved by our What-to-ask and How-to-ask strategies, induced by LLM self-play. Our extensive experiments also show our superiority. We believe that PASSING represents a crucial step towards creating more human-centric conversational agents.
Chinese Translation
在信息搜索的背景下,对话代理正经历从反应工具到主动个性化助手的演变。这一演变的一个关键方面是能够根据用户独特的需求和期望定制战略互动。与现有研究集中于主动澄清查询模糊性不同,我们专注于澄清用户的专业知识,以便定制响应以提高用户理解。我们发现,现有的代理仅凭查询难以判断用户的专业知识,这一局限性阻碍了它们动态调整响应的能力。为了解决这一问题,我们引入了PASSING,旨在通过有针对性的询问使代理主动澄清用户的专业知识。这是通过我们的What-to-ask和How-to-ask策略实现的,这些策略由LLM自我对弈引导。我们的广泛实验也显示了我们的优势。我们相信,PASSING代表了朝着创建更以人为本的对话代理迈出的重要一步。
cs.AI / 83 / 2608.22310

HERO: Human-profile Enhanced Retrieval Optimization Framework for Long-term Agent Memory

HERO:基于人类特征的长期代理记忆检索优化框架
Lin, Yuanhua, Li, Yile, Zhao, Zhiyuan, Shang, Jing, Sun, Jian
Abstract
Long-term memory is crucial for personalized responses and long-horizon agent interactions. Existing methods often rely on LLMs to compress or rewrite dialogue histories and use the transformed memories as retrieval evidence. Despite the progress in organizing fragmented contexts, two major drawbacks persist: (1) information loss from compression, which discards fine-grained but later useful details, and (2) semantic drift from rewriting, which erodes the original tone and situated context. In this work, we propose a novel Human-profile Enhanced Retrieval Optimization framework for long-term agent memory (HERO). Specifically, HERO converts the dialogue history into a traceable heterogeneous memory graph that preserves raw dialogue text as evidence for reasoning, thereby mitigating information loss. For retrieval, HERO extracts initial anchors from the current query and incorporates human profiles via an iterative graph traversal; these anchors and profiles provide guidance signals that adaptively activate the most informative regions of the graph. Experiments on two benchmark datasets show that HERO outperforms strong baselines on both factual and personalized reasoning, while providing more faithful access to raw dialogue evidence.
Chinese Translation
长期记忆对于个性化响应和长期代理交互至关重要。现有方法通常依赖大型语言模型(LLMs)来压缩或重写对话历史,并将转化后的记忆作为检索证据。尽管在组织碎片化上下文方面取得了一定进展,但仍然存在两个主要缺陷:(1)压缩导致的信息丢失,舍弃了细粒度但后续有用的细节,以及(2)重写带来的语义漂移,侵蚀了原始语气和情境背景。在本研究中,我们提出了一种新颖的人类特征增强检索优化框架,用于长期代理记忆(HERO)。具体而言,HERO将对话历史转换为可追踪的异构记忆图,保留原始对话文本作为推理的证据,从而减轻信息丢失。对于检索,HERO从当前查询中提取初始锚点,并通过迭代图遍历结合人类特征;这些锚点和特征提供了指导信号,能够自适应地激活图中最具信息性的区域。在两个基准数据集上的实验表明,HERO在事实和个性化推理方面均优于强基线,同时提供了对原始对话证据的更真实访问。
cs.AI / 84 / 2608.22347

Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture

认知的栖息地:在最小完整认知架构中剖析涌现功能与计算功能
Arrabal-Campos, Francisco M., Montoya, Francisco G., Alcayde, Alfredo, Fernández, Ignacio
Abstract
A cognitive architecture is more than the module that reasons: it must also decide how long to think and what deserves the effort. We built a minimal but complete system - a recurrent reasoner with adaptive halting, a homeostatic control field, and a value module - and asked of each part: does this function emerge from gradient descent, or must it be computed? Competence emerges. Stopping appears to emerge too, and to be worth more than everything decidable in advance, but that appearance is instrumentation: payoff at matched mean compute climbs from 0.467 (uniform) through 0.546 (difficulty) to 0.698 (ex-ante value), and the further climb to 0.921 (posterior self-observation) does not survive audit. PonderNet-style halting returns a halting-weighted mixture of hidden states while forced-depth baselines return one, and the language head is trained on the mixture alone; equalizing the readout annihilates the apparent advantage of native execution (residual +0.000 [0.000, 0.000]). Value does not emerge: trained couplings capture zero of a payoff an explicit allocator captures completely (+0.151, routing correlation +0.79), so the second-order decisions that pay must be computed, at least where value is orthogonal to content, as here by construction. On a frozen LLM actuator the same instruments show self-consistency voting to be a measured bound (+0.0236 [+0.0150, +0.0326]) and inter-sample agreement nearly worthless as a stopping signal, its mass concentrating on wrong answers. Every null we assert carries a mechanism and a positive control, and the protocol is part of the contribution. Executing our own falsifiable prediction, value under commitment pays +0.1312 [+0.1124, +0.1502] in a cliff-cost family, some seven times the smooth-family estimate - not because the cliff shifts information ex ante, but because it multiplies the attainable range fivefold (5.1x [3.4, 8.2]).
Chinese Translation
认知架构不仅仅是推理的模块:它还必须决定思考的时间和值得付出努力的内容。我们构建了一个最小但完整的系统——一个具有自适应停止机制的递归推理器、一个稳态控制场和一个价值模块,并对每个部分提出了问题:这个功能是从梯度下降中涌现出来的,还是必须被计算?能力是涌现的。停止似乎也在涌现,并且其价值超过了所有可以提前决定的内容,但这种表象是工具化的:在匹配的平均计算下,收益从0.467(均匀)通过0.546(难度)到0.698(事前价值)上升,而进一步上升到0.921(事后自我观察)并未经得起审计。PonderNet风格的停止返回的是一个加权的隐藏状态混合,而强制深度基线返回的是一个状态,语言头仅在混合上进行训练;平衡读出消除了本地执行的明显优势(残差 +0.000 [0.000, 0.000])。价值并未涌现:训练的耦合捕获了显式分配器完全捕获的收益的零部分(+0.151,路由相关性 +0.79),因此必须计算支付的二阶决策,至少在价值与内容正交的情况下,如此处所构建的那样。在一个冻结的LLM执行器上,相同的工具显示自我一致性投票是一个测量的界限(+0.0236 [+0.0150, +0.0326]),而样本间一致性几乎毫无价值作为停止信号,其质量集中在错误答案上。我们所声称的每一个无效结果都携带一个机制和一个正控制,协议也是贡献的一部分。执行我们自己的可证伪预测,在悬崖成本系列下,承诺下的价值支付为 +0.1312 [+0.1124, +0.1502],是平滑系列估计的七倍多——这并不是因为悬崖事前转移了信息,而是因为它将可达范围扩大了五倍(5.1x [3.4, 8.2])。
cs.AI / 85 / 2608.22356

Addressing the Selection Problem in Explainable AI

解决可解释人工智能中的选择问题
Vlases, Claire, Morrison, Katelyn
Abstract
Explainable AI (XAI) research has produced a plethora of explanation techniques, yet user studies repeatedly show that available explanations are not effective in practice. We argue that, given the siloed nature of conventional XAI, users are struggling to select the appropriate XAI technique. Viewing XAI through a philosophical lens, we offer a formalization of what we call the selection problem: the systematic failure of XAI interfaces to bridge the gap between a user's natural-language uncertainty and the explanation technique that resolves it. Following a logical premise-conclusion format, we show that conventional interfaces require users to translate their uncertainty into a technique selection, a challenging prerequisite to meet. We also propose a structural solution: a multi-agent LLM orchestration tool that translates the user's query to the proper XAI explanation technique. We provide an example of how this structural solution could be instantiated to address the selection problem.
Chinese Translation
可解释人工智能(XAI)研究已经产生了大量的解释技术,但用户研究反复表明,现有的解释在实践中并不有效。我们认为,鉴于传统XAI的孤立性质,用户在选择合适的XAI技术时面临困难。从哲学的角度看待XAI,我们提出了一个我们称之为选择问题的形式化定义:XAI接口系统性地未能弥合用户自然语言的不确定性与能够解决该不确定性的解释技术之间的差距。按照逻辑前提-结论的格式,我们展示了传统接口要求用户将其不确定性转化为技术选择,这是一个难以满足的挑战性前提。我们还提出了一种结构性解决方案:一个多智能体的LLM(大语言模型)编排工具,它将用户的查询翻译为适当的XAI解释技术。我们提供了一个示例,说明如何实例化这一结构性解决方案以解决选择问题。
cs.AI / 86 / 2608.22363

Analyzing and Mitigating Cross-Lingual Degradation in Multilingual Medical VQA

分析与缓解多语言医学视觉问答中的跨语言退化
Wang, Jingbo, Zhao, Sendong, Wang, Haochun, Qin, Bing, Liu, Ting
Abstract
Medical visual question answering (VQA) is a crucial task in clinical AI, yet its evaluation has so far centered almost exclusively on English, limiting its relevance to linguistically diverse patients and clinicians. Recent multilingual medical VQA benchmarks show that large vision-language models (LVLMs) degrade in non-English languages, but lack a fine-grained analysis of how cross-lingual variation affects the distinct capabilities that medical VQA requires. To this end, we construct a multilingual medical VQA benchmark over eight languages, organized into four representative scenarios that isolate the core capabilities medical VQA requires. Evaluating five open- and closed-source LVLMs, we find that cross-lingual degradation is not uniform but highly scenario-dependent. We therefore propose MedVL-XLRepE, a training-free scenario-aware representation engineering method, leveraging LVLMs' superior English medical VQA capability to steer non-English representations toward their English counterparts at inference time. Across three LVLMs and eight languages, MedVL-XLRepE consistently mitigates cross-lingual degradation, with gains of up to 6.33\%.
Chinese Translation
医学视觉问答(VQA)是临床人工智能中的一项关键任务,但迄今为止,其评估几乎完全集中在英语上,这限制了其对语言多样性的患者和临床医生的相关性。最近的多语言医学 VQA 基准显示,大型视觉-语言模型(LVLMs)在非英语语言中的表现退化,但缺乏对跨语言变异如何影响医学 VQA 所需的不同能力的细致分析。为此,我们构建了一个涵盖八种语言的多语言医学 VQA 基准,组织成四个代表性场景,以隔离医学 VQA 所需的核心能力。通过评估五个开源和闭源的 LVLMs,我们发现跨语言退化并非均匀,而是高度依赖场景。因此,我们提出了 MedVL-XLRepE,这是一种无训练的场景感知表示工程方法,利用 LVLMs 在英语医学 VQA 中的优越能力,在推理时将非英语表示引导至其英语对应物。在三个 LVLMs 和八种语言中,MedVL-XLRepE 一致地缓解了跨语言退化,增益高达 6.33%。
cs.AI / 87 / 2608.22364

WAM-OPD: On-Policy Distillation for World Action Models

WAM-OPD:用于世界动作模型的在线蒸馏
Yang, Liuhaichen, Jiang, Zhuang, Sheng, Chenchao, Tang, Zezhi
Abstract
World action models (WAMs) couple visual future prediction with robot action generation, but accelerated students can lose task capabilities during distillation and later encounter states that are poorly represented by offline data. We study whether on-policy distillation (OPD) can repair such a student without requiring sparse-reward reinforcement learning. We introduce WAM-OPD, a deployment-consistent post-training recipe for a video-first WAM. The student acts in the environment and therefore determines the history distribution. A frozen teacher labels those student histories with coherent video and action targets, while the student action branch is trained under its own generated video plan, as it is at deployment. Joint video and action losses update lightweight adapters in the shared backbone, together with an action flow-matching regularizer. In preliminary RoboTwin 2.0 studies on two tasks, the released one-video/one-action-step Flash-WAM improves from 0.0% to 58.3% success on HANDOVER MIC, and from 16.7% to 33.3% on PUT OBJECT CABINET. These task-specific results are an initial capability proof rather than evidence of broad or uniform generalization. They nevertheless suggest that dense teacher supervision on student-induced histories is a promising post-training interface for video-first WAMs.
Chinese Translation
世界动作模型(WAMs)将视觉未来预测与机器人动作生成相结合,但加速的学生在蒸馏过程中可能会失去任务能力,并在后续遇到由离线数据表现不佳的状态。我们研究了在线蒸馏(OPD)是否可以修复这样的学生,而无需稀疏奖励强化学习。我们提出了WAM-OPD,这是一种针对视频优先WAM的部署一致性后训练方案。学生在环境中行动,因此决定了历史分布。一个冻结的教师使用一致的视频和动作目标对这些学生历史进行标注,而学生动作分支则在其自身生成的视频计划下进行训练,正如在部署时一样。联合视频和动作损失更新共享主干中的轻量适配器,同时结合动作流匹配正则化器。在对两个任务的初步RoboTwin 2.0研究中,发布的单视频/单动作步骤Flash-WAM在HANDOVER MIC上的成功率从0.0%提高到58.3%,在PUT OBJECT CABINET上的成功率从16.7%提高到33.3%。这些任务特定的结果是初步能力证明,而不是广泛或均匀泛化的证据。尽管如此,它们仍然表明对学生引导历史的密集教师监督是视频优先WAMs的一个有前景的后训练接口。
cs.AI / 88 / 2608.22417

LLMs for Survey Text Analysis - A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis

用于调查文本分析的大型语言模型 - 人类与GPT-5在归纳内容分析中的性能比较
Bergmann, Leonardo, Gheorghiu, Renata, Gvritishvili, Ana, Mican, Alex, Stewart, Chris, Tolonen-Weckström, Topias
Abstract
Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in inductive content analysis remains limited. This study compares human and LLM-based inductive coding of open-ended survey responses from 903 answers across six variables from a European PhD student survey. Five human coders performed inductive content analysis following a standardized coding scheme, while an LLM (GPT-5.4) conducted the same task using an established prompting procedure. Agreement between human and LLM outputs was assessed using the Adjusted Rand Index (ARI). Results showed an alignment between humans and the LLM, with ARI values of 0.61 for coding and 0.54 for theme generation. These values were close to the internal consistency of coding and theme results within humans (ARI = 0.68) and the LLM (ARI = 0.76). Agreement varied widely across variables, with low within-entity consistency consistently linked to low between-entity agreement, underscoring the role of data characteristics and individual performance in reliability. Overall, the findings suggest that LLMs can approximate human coding in this case-specific setting, particularly at the coding level, and may serve as a scalable support tool for inductive qualitative analysis.
Chinese Translation
大型语言模型(LLMs)在定性研究中的文本分析支持中日益受到关注,但关于它们在归纳内容分析中的表现的证据仍然有限。本研究比较了人类与基于LLM的开放式调查问卷回答的归纳编码,数据来源于一项涵盖903个回答的欧洲博士生调查的六个变量。五名人类编码员根据标准化编码方案进行归纳内容分析,而LLM(GPT-5.4)则使用已建立的提示程序执行相同任务。通过调整兰德指数(Adjusted Rand Index, ARI)评估人类与LLM输出之间的一致性。结果显示人类与LLM之间存在一致性,编码的ARI值为0.61,主题生成的ARI值为0.54。这些值接近人类内部编码和主题结果的一致性(ARI = 0.68)及LLM的一致性(ARI = 0.76)。不同变量之间的一致性差异很大,低实体内一致性与低实体间一致性之间存在持续的关联,强调了数据特征和个体表现对可靠性的影响。总体而言,研究结果表明,在这一特定情境下,LLMs可以近似人类编码,特别是在编码层面,并可能作为归纳定性分析的可扩展支持工具。
cs.AI / 89 / 2608.22421

Where World Models Break: Natural-Input Failure Discovery

世界模型的失效点:自然输入失败发现
Shi, Zhanpeng, Liang, Zi, Feng, Rong, Tang, Shiqin, Chen, Xuyang, Li, Hongzong
Abstract
World models predict action-conditioned futures and serve as critical internal simulators for downstream planning and control. However, catastrophic prediction failures of world models could dangerously propagate through the control pipeline, as subsequent agent or model training and decision-making depend heavily on the continuous environment evolution forecasted by these world models. Existing evaluations overlook this systemic risk: by aggregating average errors over benign generations from general queries, they fail to stress-test the model against catastrophic collapses under rare or unobserved condition-action combinations. To bridge this gap, we formalize the natural-input failure discovery problem: under a finite query budget, finding environment-valid conditions and action prefixes that induce severe prediction risk, verifying whether these failures reproduce on fresh seeds, and testing their persistence under nearby valid edits. Discovering such critical failures is computationally challenging, as valid condition-action combinations explode exponentially, rendering exhaustive search or standard sampling infeasible given the high cost of noisy rollouts. To tackle this, we propose BasinLens, which exploits the underlying structure of valid inputs, where each coordinate possesses environment-defined semantic types and admissible domains, by pairing uncertainty-guided global search with typed local replacements. Across diverse benchmarks and world-model families, BasinLens exposes reproducible and locally persistent failure modes that conventional evaluations fail to reveal, showing that average-case benchmarks can mask important vulnerabilities in world-model-driven control.
Chinese Translation
世界模型预测基于动作的未来,并作为下游规划和控制的重要内部模拟器。然而,世界模型的灾难性预测失败可能会在控制管道中危险地传播,因为后续的智能体或模型训练和决策制定在很大程度上依赖于这些世界模型预测的持续环境演变。现有评估忽视了这种系统性风险:通过对来自一般查询的良性生成的平均误差进行聚合,它们未能在稀有或未观察到的条件-动作组合下对模型进行压力测试,以应对灾难性崩溃。为了解决这一问题,我们形式化了自然输入失败发现问题:在有限的查询预算下,寻找有效的环境条件和动作前缀,这些条件和前缀会引发严重的预测风险,验证这些失败是否在新种子上重现,并测试它们在附近有效编辑下的持续性。发现这些关键失败在计算上具有挑战性,因为有效的条件-动作组合呈指数级爆炸,使得在高噪声回放成本下,穷举搜索或标准采样变得不可行。为此,我们提出了BasinLens,它利用有效输入的潜在结构,其中每个坐标具有环境定义的语义类型和可接受的域,通过将不确定性引导的全局搜索与类型化的局部替换相结合。在多种基准测试和世界模型家族中,BasinLens揭示了常规评估未能显示的可重现和局部持久的失败模式,表明平均情况基准可能掩盖了世界模型驱动控制中的重要脆弱性。
cs.AI / 90 / 2608.22429

Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding

通过结构化基础进行思考:用于图表和视觉表格理解的感知强化学习
Jiang, Changjiang, Zhao, Qiannian, Xin, Lei, Xie, Jinxiang, Nakov, Preslav, Xie, Zhuohan
Abstract
Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for fine-grained perception. However, this reliance introduces significant inference latency and fails to effectively resolve the spatial-structural gap-a fundamental challenge in text-dense and structurally relational visuals (e.g., charts and visual tables) where strict relative spatial arrangements bind textual elements. Without external tools, standard MLLMs struggle with such fine-grained visual reasoning tasks. To address these issues, we propose Think with Structured Grounding (TwSG), a novel fine-grained image perception framework designed to internalize complex images's tool-use capabilities within the model. TwSG distills the benefits of multi-step reasoning and micro-cropping into a single efficient forward pass during inference. Specifically, we use an MLLM to identify key regions guided by ground-truth answers, and then prompt a teacher model to generate high-quality visual question-answering (VQA) data. These fine-grained, region-based supervisory signals are subsequently distilled back into the full-image representation. Our training pipeline consists of two stages: (1) a cold-start supervised fine-tuning (SFT) phase using multi-turn data with focused area descriptions to foster complex reasoning and error recovery; and (2) a reinforcement fine-tuning (RFT) phase driven by a novel process reward mechanism, TL-GRPO, which encourages strategic reasoning. Extensive experiments across various MLLM architectures demonstrate that TwSG reduces inference latency while substantially improving accuracy and robustness, endowing models with native fine-grained region description and flexible reasoning capabilities.
Chinese Translation
能够与图像进行思考的多模态大型语言模型(MLLMs)通常依赖外部工具进行细粒度感知。然而,这种依赖引入了显著的推理延迟,并未有效解决空间结构差距——这是文本密集和结构相关视觉(例如图表和视觉表格)中的一个基本挑战,在这些视觉中,严格的相对空间排列约束了文本元素。没有外部工具,标准的MLLM在处理此类细粒度视觉推理任务时表现不佳。为了解决这些问题,我们提出了结构化基础思维(Think with Structured Grounding, TwSG),这是一种新颖的细粒度图像感知框架,旨在将复杂图像的工具使用能力内化到模型中。TwSG将多步推理和微裁剪的优势提炼为推理过程中的单次高效前向传播。具体而言,我们使用MLLM识别由真实答案指导的关键区域,然后提示教师模型生成高质量的视觉问答(VQA)数据。这些细粒度的基于区域的监督信号随后被提炼回完整图像表示。我们的训练流程包括两个阶段:(1)使用多轮数据和聚焦区域描述进行冷启动监督微调(SFT)阶段,以促进复杂推理和错误恢复;(2)由新颖的过程奖励机制TL-GRPO驱动的强化微调(RFT)阶段,鼓励战略性推理。针对各种MLLM架构的广泛实验表明,TwSG在显著提高准确性和鲁棒性的同时减少了推理延迟,使模型具备原生的细粒度区域描述和灵活的推理能力。
cs.AI / 91 / 2608.22438

When Persona Simulations Are Informative: Graph-Structured Signals for Pluralistic Opinion Sensing

当角色模拟具有信息性时:用于多元意见感知的图结构信号
An, Taehyeon, Park, Jaehyeong, Shin, Donghyuk
Abstract
Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across diverse domains. However, apparent response variation can reflect unconditioned model priors or token sampling noise rather than systematic persona conditioning. We argue that persona-conditioned variation is informative when semantically similar personas exhibit concordant response shifts. To operationalize this principle, we introduce Persona-Conditioned Informativeness (PCI), an unsupervised diagnostic metric that measures whether semantically similar personas deviate in concordant directions relative to item-level sample baselines. By modeling personas as a similarity graph, PCI uses Local Moran's I to quantify local spatial coherence and extract compact persona subsets without using construct labels. To evaluate PCI without external human benchmarks, we test its ability to recover established latent value structure using the 57-item Portrait Values Questionnaire-Revised (PVQ-RR). Confirmatory factor analysis (CFA) shows that a PCI-selected 10% subset substantially improves overall construct recovery relative to response-stability and random selection. These findings support PCI as a principled internal diagnostic for screening synthetic respondents in survey pipelines.
Chinese Translation
基于角色的条件大语言模型(LLMs)在多个领域中被越来越多地用于模拟调查响应。然而,明显的响应变化可能反映的是无条件模型先验或标记采样噪声,而不是系统性的角色条件化。我们认为,当语义相似的角色表现出一致的响应变化时,基于角色的变化是具有信息性的。为了实现这一原则,我们引入了角色条件信息性(Persona-Conditioned Informativeness, PCI),这是一种无监督的诊断指标,用于测量语义相似的角色相对于项目级样本基线是否朝着一致的方向偏离。通过将角色建模为相似性图,PCI利用局部莫兰指数(Local Moran's I)来量化局部空间一致性,并在不使用构念标签的情况下提取紧凑的角色子集。为了在没有外部人类基准的情况下评估PCI,我们测试其使用57项修订版肖像价值问卷(Portrait Values Questionnaire-Revised, PVQ-RR)恢复既定潜在价值结构的能力。确认性因素分析(CFA)显示,PCI选择的10%子集在响应稳定性和随机选择的相对构念恢复方面显著提高了整体构念恢复。这些发现支持PCI作为在调查流程中筛选合成响应者的原则性内部诊断工具。
cs.AI / 92 / 2608.22472

Small Reasoning Models are Instruction Followers in Function Calling

小型推理模型在函数调用中的指令跟随能力
Taheri, Yalda, Heydari, Mohammad Hassan, Naaman, Erfan, Fatemi, Afsaneh
Abstract
Function calling represents the core capability of agentic large language models (LLMs). Existing research has focused on enhancing LLMs function-calling accuracy through fine-tuning, reinforcement learning (RL), and multi-agent frameworks, particularly for native function-calling LLMs. This work demonstrates that LLMs achieve superior accuracy in function calling in instruction-following contexts (i.e., standard user-assistant interactions) rather than a tool calling context. We introduce Instruction-Followed Function Calling (IFFC), a novel framework that decouples function-calling logic from the primary LLM and delegates it to a dedicated smaller model operating within the instruction-following paradigm. Our method consistently outperforms both native function calling (NFC) and prompt-based function calling (PFC) baselines, with particularly strong gains on reasoning-oriented LLMs. Furthermore, we demonstrate that IFFC maintains robust performance under aggressive quantization, enabling efficient on-device deployment without significant accuracy degradation. This work establishes a new paradigm for reliable, resource-efficient function calling in edge-computing scenarios.
Chinese Translation
函数调用代表了代理大型语言模型(LLMs)的核心能力。现有研究集中于通过微调、强化学习(RL)和多智能体框架来提高LLMs的函数调用准确性,尤其是针对原生函数调用的LLMs。本研究表明,LLMs在指令跟随上下文(即标准用户-助手交互)中实现了更高的函数调用准确性,而非工具调用上下文。我们提出了指令跟随函数调用(Instruction-Followed Function Calling, IFFC)这一新框架,将函数调用逻辑与主要LLM解耦,并将其委托给一个在指令跟随范式下运行的专用小型模型。我们的方法在性能上始终优于原生函数调用(NFC)和基于提示的函数调用(PFC)基准,尤其在推理导向的LLMs上表现出显著的提升。此外,我们证明了IFFC在激进量化下仍能保持强大的性能,使得在设备上的高效部署成为可能,而不会显著降低准确性。本研究为边缘计算场景中可靠且资源高效的函数调用建立了新的范式。
cs.AI / 93 / 2608.22504

When Does AI for PDEs Yield Scientific Evidence?

人工智能在偏微分方程中的应用何时能提供科学证据?
Wang, Wenshuo
Abstract
Existing AI-for-PDE benchmarks primarily assess models in terms of predictive or approximation accuracy. In physics research, however, AI outputs often serve as evidence for scientific claims. These two objectives are not equivalent: the former measures an output's agreement with a reference target or satisfaction of governing constraints; the latter asks whether, given a specified object of study, scientific claim, assumptions, and evidence standard, the output provides sufficient evidence for that claim. To bridge this gap, we extend a widely used PDE-simulation benchmark and a comprehensive benchmark for PDE inverse problems to enable, for the first time in AI for PDEs, evaluation of whether and to what extent model outputs support specified scientific claims. Our results show that numerical accuracy and evidential support can rank models differently, explain when and why they do so, and reveal that existing benchmarks can favor methods whose outputs provide weaker support for the scientific claims of interest. Together, we formalize, empirically demonstrate, and explain this evaluation--use mismatch in AI for PDEs.
Chinese Translation
现有的人工智能在偏微分方程(PDE)中的基准主要评估模型的预测或近似精度。然而,在物理研究中,人工智能的输出往往作为科学主张的证据。这两个目标并不等同:前者衡量输出与参考目标的一致性或对控制约束的满足;后者则询问在给定特定研究对象、科学主张、假设和证据标准的情况下,输出是否为该主张提供了足够的证据。为了弥补这一差距,我们扩展了一个广泛使用的PDE模拟基准和一个全面的PDE逆问题基准,以首次在人工智能应用于PDE中评估模型输出在多大程度上支持特定科学主张。我们的结果表明,数值精度和证据支持可能会对模型进行不同的排名,解释何时以及为何会出现这种情况,并揭示现有基准可能偏向于那些输出对相关科学主张提供较弱支持的方法。综上所述,我们形式化、实证展示并解释了这一评估——人工智能在PDE中的应用与使用不匹配的问题。
cs.AI / 94 / 2608.22510

ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts

ClawProBench:具有运行时覆盖和冻结工作场所风格保留的AI代理的追踪感知评估
Xiao, YuanHang
Abstract
Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-specifies what is being evaluated: the proper unit is a declared model-plus-runtime configuration whose failures can occur in evidence acquisition, runtime routing, safety boundaries, or repeated execution. We present ClawProBench, a trace-aware benchmark for runtime-native agent evaluation instantiated on OpenClaw, a live agent runtime with workspace tools and native surfaces for browsing, memory, messaging, scheduling, skills, and subagents. ClawProBench defines two tracks: a 102-scenario full profile with live workspace and native-runtime routing tasks, and a frozen 68-scenario holdout with closed-world JSON output contracts for robust ranking. Trials are scored from execution traces via a safety-gated formula combining correctness, process quality, and efficiency, preserving failure evidence for audit. Our anonymous artifact includes benchmark definitions, scoring code, manifests and sanitized traces. We evaluate 68 configurations on the full profile and 37 on holdout. The top safety-gated average trace score is 0.7671. Native-runtime tasks underperform workspace-live tasks (0.5238 vs. 0.6415). On holdout, pass@k-any outperforms strict three-trial pass (0.6638 vs. 0.2890), while full-profile and holdout rankings show weak alignment (Spearman 0.1300). Rankings based purely on correctness differ substantially from process-aware, safety-gated and strict-pass views. Final-answer leaderboards may hide native-surface weaknesses, one-off successes and trace-local agent failure modes.
Chinese Translation
代理基准通常仅评估最终答案,即使代理在有状态的运行时上运行。我们认为这不足以明确评估的内容:适当的单位是一个声明的模型加运行时配置,其失败可能发生在证据获取、运行时路由、安全边界或重复执行中。我们提出了ClawProBench,这是一个针对运行时本地代理评估的追踪感知基准,基于OpenClaw实例化,OpenClaw是一个具有工作区工具和本地浏览、内存、消息传递、调度、技能和子代理的实时代理运行时。ClawProBench定义了两个轨道:一个包含102个场景的完整配置,具有实时工作区和本地运行时路由任务,以及一个冻结的68个场景的保留集,具有封闭世界的JSON输出合同,以实现稳健排名。试验通过一个安全门控公式从执行追踪中评分,该公式结合了正确性、过程质量和效率,同时保留失败证据以供审计。我们的匿名文献包括基准定义、评分代码、清单和经过清理的追踪。我们在完整配置上评估了68种配置,在保留集上评估了37种配置。顶级安全门控平均追踪分数为0.7671。本地运行时任务的表现低于工作区实时任务(0.5238对0.6415)。在保留集中,pass@k-any的表现优于严格的三次试验通过(0.6638对0.2890),而完整配置和保留集排名显示出弱对齐(Spearman 0.1300)。仅基于正确性的排名与过程感知、安全门控和严格通过的视角有显著不同。最终答案排行榜可能掩盖本地表面的弱点、一次性成功和追踪局部代理失败模式。
cs.AI / 95 / 2608.22512

HANSARD: A Reference Architecture for Forensic Readiness, Runtime Witnessing, and Graded Attribution in Autonomous Multi-Agent AI Systems

HANSARD:一种用于法医准备、运行见证和自主多智能体人工智能系统中分级归因的参考架构
Sardianos, Christos, Pla, Iliana, Efthymiou, Vasilis, Varlamis, Iraklis, Lagkas, Thomas, Sarigiannidis, Panagiotis, Papadopoulos, Georgios Th.
Abstract
Autonomous multi-agent systems nowadays act in finance, software supply chains, and security operations. Already, the first largely AI-orchestrated intrusion campaigns have been reported. Yet, when such a system causes harm, no method can robustly establish what happened, what caused it, or who is accountable. This is because provenance forensics works at the wrong abstraction, formal causality assumes the causal model, and agent auditing trusts self-recording. The target failure mode is, thus, attribution laundering, i.e., spreading an act across redundant agents until none is a but-for cause. Worse, the record is produced by the suspects, which comprises the assumption adopted throughout this work. Agents may therefore anticipate the investigation and the part of logging infrastructure may itself collude. In this paper, HANSARD is proposed, a reference architecture treating accountability as a life-cycle property. First, a readiness profile sealed before operation bounds what later findings may claim. Second, capturing at five choke points beyond the agents' reach makes omissions detectable, not only tampering. Third, a typed PROV-DM-aligned causal graph accrues as the system runs, and three indicators read it live to gate oversight without adjudicating. Fourth, post-incident replay yields contingent effects under the modified Halpern-Pearl definition, together with a compensation-set size. Finally, a synergy residual measures harm due to the combination rather than to individuals, making laundering visible. Cause, responsibility and accountability are then reported separately, each capped by an evidentiary tier, while a future research agenda is also provided.
Chinese Translation
如今,自主多智能体系统在金融、软件供应链和安全操作中发挥着作用。已经有报道称,首批主要由人工智能协调的入侵活动已被发现。然而,当这样的系统造成损害时,没有任何方法能够可靠地确定发生了什么、造成了什么或谁应对此负责。这是因为来源法医在错误的抽象层次上工作,形式因果关系假设了因果模型,而智能体审计则信任自我记录。因此,目标失败模式是归因洗涤,即将一个行为分散到冗余的智能体中,直到没有一个是必要原因。更糟糕的是,记录是由嫌疑人生成的,这也是本研究中采用的假设。因此,智能体可能会预见调查,而日志基础设施的部分可能会自身串通。本文提出了HANSARD,一种将问责视为生命周期属性的参考架构。首先,在操作前密封的准备配置文件限制了后续发现可能声称的内容。其次,在超出智能体控制的五个关键点进行捕获,使遗漏可被检测,而不仅仅是篡改。第三,随着系统运行,生成一个类型化的与PROV-DM对齐的因果图,并且三个指标实时读取该图,以进行监督而不进行裁决。第四,事件后重播在修改后的Halpern-Pearl定义下产生偶然效应,并给出补偿集的大小。最后,协同残差衡量由于组合而非个体造成的损害,使洗涤行为可见。因果关系、责任和问责将被单独报告,每个报告都由证据层级限制,同时还提供了未来的研究议程。
cs.AI / 96 / 2608.22533

CONTRAMEM: Learning Self-Evolving Procedural Memory from Contrasting Multi-Model Trajectories

CONTRAMEM:从对比多模型轨迹中学习自我演化的过程记忆
Deng, Zheyuan, Lu, Binghang, Feng, Hanqi, Huang, Shirley, Wang, Dianzhuo, Xu, Yuanda, Zhang, Zhiwei, Sun, Yige, Mou, Changhong, Zhang, Runyu, Hao, Yuexing, Poczos, Barnabas, Li, Xiaomin
Abstract
Autonomous computer-use agents are increasingly applied to long-horizon tasks requiring coordinated application calls, persistent state tracking, and verifier-sensitive writes, yet they remain prone to procedural failures: misreading application state, tool semantics, or task progress. Procedural memory promises more consistent decisions and less redundant exploration, but constructing high-quality memory without model training remains challenging. We introduce CONTRAMEM, a source-flexible, training-free framework for self-evolving procedural memory that treats same-task outcome variation as supervision: differences in correctness, efficiency, recovery, and failure modes expose outcome-relevant procedural distinctions, distilled into a compact bank of app-level Function Cards and task-level Skill Cards that evolves through localized curation rather than append-only accumulation or whole-bank rewriting. On held-out GAIA2/ARE computer-use tasks, CONTRAMEM more than doubles the success rate across the three source-model targets (26.2% to 55.3%), with consistent per-model gains (GPT-5.5: 27.5 to 61.0; Claude Sonnet 4.6: 28.0 to 52.5; DeepSeek V4 Pro: 23.0 to 52.5). The same bank transfers unchanged to the unseen Qwen3.7 Plus (18.5 to 35.5), indicating transferable procedural knowledge rather than model-specific behavior. The same construction carries over unchanged to AppWorld, beating both no memory and its own single-source self-memory variant for all three mid-tier agents on both public test splits. Under a matched trajectory budget, heterogeneous multi-model trajectories yield stronger memory than self- or same-model multi-rollout memory: the margin comes from contrastive behavioral diversity, not stronger source agents or more sampling.
Chinese Translation
自主计算机使用代理越来越多地应用于需要协调应用调用、持续状态跟踪和对验证者敏感的写入的长时间任务,但它们仍然容易出现过程性故障:错误解读应用状态、工具语义或任务进展。过程记忆承诺提供更一致的决策和更少的冗余探索,但在没有模型训练的情况下构建高质量记忆仍然具有挑战性。我们提出了CONTRAMEM,这是一个源灵活、无训练的自我演化过程记忆框架,将同一任务结果的变异视为监督:正确性、效率、恢复和失败模式的差异揭示了结果相关的过程性区别,这些区别被提炼为一个紧凑的应用级功能卡片和任务级技能卡片库,该库通过局部策划而非仅附加积累或整体重写进行演变。在保留的GAIA2/ARE计算机使用任务上,CONTRAMEM在三个源模型目标上成功率超过两倍(从26.2%提高到55.3%),每个模型的增益保持一致(GPT-5.5:从27.5提高到61.0;Claude Sonnet 4.6:从28.0提高到52.5;DeepSeek V4 Pro:从23.0提高到52.5)。同一库在未见的Qwen3.7 Plus上转移不变(从18.5提高到35.5),表明可转移的过程知识而非特定模型的行为。同样的构建在AppWorld中保持不变,超越了无记忆和其自身单源自我记忆变体,在所有三个中层代理的公共测试分割中均表现优异。在匹配的轨迹预算下,异构多模型轨迹产生的记忆优于自我或同模型的多次展开记忆:这一差距来自对比行为的多样性,而非更强的源代理或更多的采样。
cs.AI / 97 / 2608.22538

STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control

STAGE:具有策略范围上下文和确定性控制的状态翻译到代理图执行
Luo, Mengxi, Chen, Changjia, Cao, An, Huang, Zirong, Dai, Wanyi
Abstract
Policy-governed agents must interpret case evidence while following an authorized procedure. We present \textsc{Stage}, an executable-graph framework that confines model judgment to policy-scoped nodes while placing procedural control in deterministic code. At each node, the model receives task-relevant policy context and returns a typed result, while the coordinator enforces the reviewed execution contract. We evaluate \textsc{Stage} on SOP-Bench Referral Abuse, two $\tau^2$-bench domains, and Smart Dispute, a proprietary banking benchmark. Compared with monolithic full-policy execution, \textsc{Stage} generally improves task success and repeated-run reliability across workflows of varying procedural complexity. The largest gains occur on the deeper Telecom and Smart Dispute workflows, where $\mathrm{Pass}^3$ increases by 7.5--55.0 and 57.2--65.7 percentage points, respectively, depending on the model. These results show that combining policy-scoped context with deterministic procedural control can improve the reliability of policy execution.
Chinese Translation
受政策指导的代理必须在遵循授权程序的同时解释案例证据。我们提出了 extsc{Stage},一个可执行图框架,该框架将模型判断限制在策略范围节点,同时将程序控制置于确定性代码中。在每个节点,模型接收与任务相关的政策上下文并返回类型化结果,而协调者则执行审查过的执行合同。我们在 SOP-Bench Referral Abuse、两个 $ au^2$-bench 领域和 Smart Dispute(一个专有银行基准)上评估了 extsc{Stage}。与单一完整政策执行相比, extsc{Stage} 通常在不同程序复杂度的工作流中提高了任务成功率和重复运行的可靠性。最大的增益发生在更深的电信和 Smart Dispute 工作流中,其中 $ ext{Pass}^3$ 分别增加了 7.5--55.0 和 57.2--65.7 个百分点,具体取决于模型。这些结果表明,将策略范围上下文与确定性程序控制相结合可以提高政策执行的可靠性。
cs.AI / 98 / 2608.22549

Scaling Curriculum Learning For Autonomous Driving

扩展课程学习在自动驾驶中的应用
Koprulu, Cevahir, Paz, David, Tao, Feng, Guo, Yuliang, Huang, Xinyu, Topcu, Ufuk, Ren, Liu
Abstract
Batched simulators for autonomous driving have recently enabled training reinforcement learning (RL) agents at scale, encompassing thousands of traffic scenarios and billions of interactions within a matter of days. Although such high-throughput feeds RL algorithms faster than ever, their sample-efficiency has not kept pace: As the standard training scheme, domain randomization uniformly samples scenarios, thereby consuming a vast number of interactions on cases that contribute little to learning. Curriculum learning offers a remedy by adaptively prioritizing scenarios that matter most to policy improvement. We present CL4AD, the first integration of curriculum learning into batched autonomous driving simulators by framing scenario selection as an unsupervised environment design problem. We introduce utility functions that shape curricula based on success rates and the realism of the agent's behavior, in addition to existing regret-estimation functions. Large-scale experiments in GPUDRIVE demonstrate that curriculum learning achieves a 99% success rate a billion steps earlier than domain randomization, reducing wall-clock time by 77%, and outperforms heuristic curricula with static and dynamic attributes, with only one exception at the largest scale. An ablation under limited compute shows that curriculum learning improves sample efficiency by 67%. We also investigate how utility functions behave at scale, and how prioritized scenarios evolve during training. We release an implementation of CLForAD in GPUDRIVE.
Chinese Translation
批量模拟器为自动驾驶提供了在大规模下训练强化学习(RL)代理的能力,涵盖了数千种交通场景和数十亿次交互,仅需数天时间。尽管如此高吞吐量的输入使得RL算法的训练速度空前加快,但其样本效率并未同步提升:作为标准训练方案,领域随机化均匀地采样场景,从而在对学习贡献甚微的案例上消耗了大量交互。课程学习通过自适应地优先考虑对策略改进最重要的场景,提供了一种解决方案。我们提出了CL4AD,这是首次将课程学习整合到批量自动驾驶模拟器中的方法,采用将场景选择框架化为无监督环境设计问题的方式。我们引入了基于成功率和代理行为现实性的效用函数,以此来构建课程,此外还结合了现有的后悔估计函数。在GPUDRIVE中的大规模实验表明,课程学习在比领域随机化早十亿步的情况下实现了99%的成功率,将实际时间减少了77%,并且在静态和动态属性的启发式课程中表现优于,唯一的例外是在最大规模下。有限计算条件下的消融实验显示,课程学习提高了67%的样本效率。我们还研究了效用函数在大规模下的表现,以及优先场景在训练过程中的演变。我们发布了CLForAD在GPUDRIVE中的实现。
cs.AI / 99 / 2608.22559

ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation

ExecRubrics:可执行的工具增强评分标准,用于可验证和高效的长文本评估
Dhole, Kaustubh D., Clarke, Charles L. A., Agichtein, Eugene Y.
Abstract
Rubrics aim to make language-model evaluation transparent by decomposing response quality into interpretable criteria. However, natural-language rubrics are often ambiguous, require black-box LLM judges, and typically assume criteria aggregate independently through linear weighted sums, limiting their ability to capture dependencies, alternatives, penalties, and override conditions. We propose ExecRubrics, a framework for representing rubrics as compact executable programs. ExecRubrics encodes evaluation logic as verifiable Python scoring functions, giving natural-language rubric intent an operational semantics: a fixed decision procedure that can be inspected, executed, and edited. On three long-form response benchmarks-HealthBench, HelpSteer, and ArgQuality-we show that ExecRubrics can substitute for expensive black-box judges in ranking preferred over dispreferred responses, matching or improving NL rubric baselines with best preference accuracies of 53%, 78%, and 92%, respectively, while reducing evaluation latency by up to 320 times. We show that incorporating external logic and resources from text processing libraries such as NLTK and spaCy further improves preference accuracy. Our results suggest a novel way of looking at evaluation, by offering a faster, more explainable, and less ambiguous alternative to black-box rubric evaluation, particularly in high-stakes domains such as healthcare and banking where precision and auditability are critical.
Chinese Translation
评分标准旨在通过将响应质量分解为可解释的标准,使语言模型评估变得透明。然而,自然语言评分标准往往模糊不清,需要黑箱大型语言模型(LLM)评审,并通常假设标准通过线性加权和独立聚合,这限制了它们捕捉依赖关系、替代方案、惩罚和覆盖条件的能力。我们提出了ExecRubrics,一个将评分标准表示为紧凑可执行程序的框架。ExecRubrics将评估逻辑编码为可验证的Python评分函数,使自然语言评分标准的意图具有操作语义:一个固定的决策程序,可以被检查、执行和编辑。在三个长文本响应基准测试——HealthBench、HelpSteer和ArgQuality中,我们展示了ExecRubrics可以替代昂贵的黑箱评审,在对优选和非优选响应进行排名时,匹配或改善自然语言评分标准的基线,最佳偏好准确率分别为53%、78%和92%,同时将评估延迟减少了多达320倍。我们还展示了从文本处理库(如NLTK和spaCy)中引入外部逻辑和资源进一步提高了偏好准确率。我们的结果建议了一种新的评估视角,通过提供一种更快速、更可解释和更不模糊的替代方案,尤其是在医疗和银行等高风险领域,在这些领域中,精确性和可审计性至关重要。
cs.AI / 100 / 2608.22577

CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI Agents

CausalCache:长时间跨度GUI代理的条件高保真恢复
Luo, Jiaxuan, Liao, Zhanfeng, Teng, Jiayao, Wang, Yuan, Huang, Haojian
Abstract
Long-horizon GUI agents can retain a complete interaction trace cheaply as textual action records, but expose only a few past events to the policy in high-fidelity pixels. We formulate this as conditional fidelity restoration: each event persists in summary-only form and is linked to an archived screenshot, while an active visual-context budget $B$ limits how many events may be promoted to summary-plus-image form. Recent-$B$ spends every slot on the latest events. CausalCache instead reallocates the same $B$ promotions over the complete trace, evicting a recent image only when a distant event has higher conditional marginal utility. Its history-gated key/value (HGKV) adapter modifies only restored history-image tokens and is exactly bypassed with no history image. Matched-budget replacement groups and per-arm-anchored difference-in-differences supervision make uniform history amplification worth zero; a budget-aware selector then chooses which summarized events to restore. On desktop, the frozen policy shows no reliable preference for a task-relevant archived screenshot over the recent frame it would displace; HGKV learns exactly that selectivity inside a pre-specified drift envelope. On OSWorld-Verified, restoring history to high fidelity is worth about $13$ success points over summary-only memory, while same-budget allocations remain indistinguishable. Zero-shot on a cross-application mobile benchmark, CausalCache significantly improves overall success over the same-budget recent allocation ($+3.7$ points on the full roster), and the gain concentrates where it should: $+8.6$ points on the memory-critical split fixed by benchmark metadata at construction, no detectable effect on matched controls, and a significant split-by-method interaction.
Chinese Translation
长时间跨度的GUI代理可以以文本动作记录的形式廉价地保留完整的交互轨迹,但在高保真像素中仅向策略暴露少量过去事件。我们将其表述为条件保真恢复:每个事件以仅摘要的形式存在,并与归档的截图相关联,而一个活动的视觉上下文预算$B$限制了可以提升到摘要加图像形式的事件数量。最近的$B$将每个槽位用于最新事件。CausalCache则在完整轨迹上重新分配相同的$B$提升,仅在远程事件具有更高的条件边际效用时驱逐最近的图像。其历史门控键/值(HGKV)适配器仅修改恢复的历史图像标记,并在没有历史图像时完全绕过。匹配预算替换组和每个臂锚定的差异中的差异监督使得均匀历史放大值为零;一个预算感知选择器随后选择要恢复的摘要事件。在桌面环境中,冻结的策略对任务相关的归档截图与其将替换的最近帧没有可靠的偏好;HGKV正是学习到这种选择性,在预先指定的漂移范围内。在OSWorld-Verified上,将历史恢复到高保真度比仅摘要记忆价值约$13$个成功点,而相同预算的分配则无法区分。在跨应用移动基准的零样本测试中,CausalCache显著提高了整体成功率,相较于相同预算的最近分配(在完整名单上增加$+3.7$点),且增益集中在应有的地方:在基准元数据构建时修复的内存关键分割上增加$+8.6$点,对匹配控制没有可检测的影响,并且存在显著的按方法分割交互。
cs.AI / 101 / 2608.22584

Weakly supervised concept Bottleneck Learning for Robust Two stage Object centric visual reasoning

弱监督概念瓶颈学习用于鲁棒的两阶段对象中心视觉推理
Tiwari, Sparsh, Schwalbe, Gesina, Finzel, Bettina
Abstract
Two-stage neuro-symbolic architectures provide an elegant paradigm for visual problem solving by cleanly separating connectionist perception of predefined symbols from possibly later defined relational reasoning thereon. However, anchoring high-level predicates into visual frames typically necessitates annotations that are expensive to acquire. In this work, we introduce the Dynamic Orthogonal Concept Bottleneck (D-OCB), an object-centric slot- VAE framework designed to extract human-aligned symbolic predicates under extremely weak supervision. D-OCB eliminates the arduous manual tuning of loss-balancing coef- ficients by dynamically learning optimal hyperparameter allocations during training. To infuse prior knowledge on independence of concept categories, in addition to standard re- construction self-supervision we penalize correlation across concept subspaces. Crucially, to combat the instability of very low supervision regimes, D-OCB incorporates a dynamic di- mensionality allocation mechanism; this adaptive formulation allows well-represented con- cepts to yield latent dimensions to underperforming concepts that are lagging behind, effectively preventing representation collapse and significantly improving overall concept accuracy. Through an extensive empirical evaluation, we demonstrate that our framework achieves high concept alignment and downstream visual reasoning accuracy using minimal label budgets, matching or outperforming end-to-end paradigms.
Chinese Translation
两阶段神经符号架构为视觉问题解决提供了一种优雅的范式,通过将预定义符号的连接主义感知与可能后续定义的关系推理清晰分离。然而,将高层谓词锚定到视觉框架中通常需要昂贵的注释。在本研究中,我们提出了动态正交概念瓶颈(Dynamic Orthogonal Concept Bottleneck, D-OCB),这是一个以对象为中心的槽式变分自编码器(slot-VAE)框架,旨在在极弱监督下提取与人类对齐的符号谓词。D-OCB通过在训练过程中动态学习最佳超参数分配,消除了繁琐的手动调整损失平衡系数的需求。为了注入关于概念类别独立性的先验知识,除了标准的重建自监督外,我们还对概念子空间之间的相关性施加惩罚。关键是,为了应对极低监督环境下的不稳定性,D-OCB结合了一种动态维度分配机制;这种自适应的构造允许表现良好的概念将潜在维度分配给落后的表现不佳的概念,有效防止表示崩溃,并显著提高整体概念准确性。通过广泛的实证评估,我们证明了我们的框架在使用最少标签预算的情况下,实现了高概念对齐和下游视觉推理准确性,匹配或超越了端到端范式。
cs.AI / 102 / 2608.22610

Coalition-Aware Skill Reliability for Self-Evolving Agents

面向联盟的自我进化智能体技能可靠性
Zhao, Qiyan, Zhang, Xiaofeng, Liu, Bo, Chen, Minda, Xiong, Wei, Chen, Jingyang, Ye, Guanting, Yu, Wenhao, Yuan, Xiaosong, Han, Shijie, Wang, Da-Han, Ji, Jianmin, Huang, Fei, Zhang, Xu-Yao
Abstract
Agent skills, structured artifacts distilled from interaction trajectories and dynamically reused from skill banks, have become a central mechanism for enabling large language model (LLM)-based self-evolving agents to learn from past experience. Yet existing work has largely focused on the operational aspects of skills, such as acquisition, evolution, and retrieval, while leaving a more fundamental reliability question unresolved: Do accumulated skills in an agent's skill bank actually make positive mechanistic contributions? We investigate this question through systematic skill-bank audits across alternative bank compositions and deployment domains, measuring the resulting changes in agent behavior. These audits reveal two recurring reliability failures: coalition pollution, where bank-level gains conceal negative coalition-level skill contributions, and cross-domain utility reversal, where source-beneficial skills reverse their effects after transfer. These findings motivate two reliability interventions: coalition-aware skill selection during skill accumulation and label-free skill masking after transfer. Coalition-Aware Skill Selection (CASS) selects more reliable candidate skills for the current bank using sampled Shapley marginals. Unsupervised Skill-Masked Coalition Optimizer (u-SMCO) masks transferred skills whose exclusion improves retrieval quality on unlabeled target-domain data. Agentic experiments on LoCoMo, LongMemEval, HotpotQA, and ALFWorld show that CASS and u-SMCO consistently improve task performance and cross-domain generalization over strong skill-based self-evolving agent baselines. Beyond accuracy, coalition-conditioned reliability modeling reduces sensitivity to noisy outcome-reward fluctuations during reinforcement learning and exposes the limits of isolation-based skill evaluation.
Chinese Translation
智能体技能是从交互轨迹中提炼出的结构化工件,并从技能库中动态重用,已成为基于大型语言模型(LLM)自我进化智能体从过去经验中学习的核心机制。然而,现有研究主要集中在技能的操作方面,如获取、演变和检索,而对一个更根本的可靠性问题却未能解决:智能体的技能库中积累的技能是否确实对机制产生积极的贡献?我们通过对不同技能库组成和部署领域进行系统的技能库审计来探讨这一问题,测量智能体行为的变化。这些审计揭示了两个反复出现的可靠性失败:联盟污染,即技能库层面的收益掩盖了联盟层面的负面技能贡献;以及跨领域效用反转,即源有益技能在转移后反转其效果。这些发现促使我们提出两种可靠性干预措施:在技能积累过程中进行面向联盟的技能选择,以及在转移后进行无标签技能屏蔽。面向联盟的技能选择(CASS)使用抽样的Shapley边际值为当前技能库选择更可靠的候选技能。无监督技能屏蔽联盟优化器(u-SMCO)屏蔽那些排除后能提高无标签目标领域数据检索质量的转移技能。在LoCoMo、LongMemEval、HotpotQA和ALFWorld上的智能体实验表明,CASS和u-SMCO在任务性能和跨领域泛化方面始终优于强大的基于技能的自我进化智能体基线。除了准确性,面向联盟的可靠性建模还降低了在强化学习过程中对噪声结果奖励波动的敏感性,并揭示了基于孤立的技能评估的局限性。
cs.AI / 103 / 2608.22615

DeepSAGE: Stage-Aware Reinforcement Learning for Structured CBT Counseling Dialogue

DeepSAGE:面向阶段的结构化认知行为疗法咨询对话强化学习
Zhang, Qi, An, Heajun, Dumaru, Prakriti, Lee, Sang Won, Huang, Lifu, Wisniewski, Pamela J., Cho, Jin-Hee
Abstract
Large Language Model (LLM)-based counseling agents can generate fluent and supportive responses, but they often lack the structured, goal-directed progression required to conduct a coherent therapeutic session. We present DeepSAGE (Strategic AI Guidance Engine), a hybrid LLM--Deep Reinforcement Learning (DRL) framework for stage-aware counseling dialogue grounded in the first session of Cognitive Behavioral Therapy (CBT). DeepSAGE represents the session as eleven stages with explicit therapeutic objectives, with an external controller determines stage completion and the DRL model selects therapeutic intentions that guide LLM response generation. We evaluate DeepSAGE against six retrieval-, prompting-, stage-, and policy-based alternatives. DeepSAGE elicits higher simulated client engagement and openness and achieves the strongest balance of stage-goal completion and dialogue efficiency among stage-structured systems. Domain expert review further indicates that the generated conversations exhibit broadly plausible emotional trajectories and recognizable CBT processes. Because the evaluation relies primarily on simulated clients and model-based metrics, these findings demonstrate comparative dialogue-control improvements rather than clinical effectiveness. These results suggest that combining stage-structured dialogue with learned strategy selection is a promising approach for AI counseling, though clinical effectiveness, safety, and real-world utility require further human evaluation.
Chinese Translation
基于大型语言模型(LLM)的咨询代理能够生成流畅且支持性的回应,但它们通常缺乏进行连贯治疗会话所需的结构化、目标导向的进展。我们提出了DeepSAGE(战略人工智能指导引擎),这是一个基于阶段的咨询对话的混合LLM-深度强化学习(DRL)框架,基于认知行为疗法(CBT)第一会话的结构。DeepSAGE将会话表示为具有明确治疗目标的十一阶段,外部控制器决定阶段完成情况,而DRL模型选择指导LLM响应生成的治疗意图。我们将DeepSAGE与六种基于检索、提示、阶段和策略的替代方案进行了评估。DeepSAGE引发了更高的模拟客户参与度和开放性,并在阶段结构系统中实现了阶段目标完成与对话效率的最佳平衡。领域专家的评审进一步表明,生成的对话展现了广泛合理的情感轨迹和可识别的CBT过程。由于评估主要依赖于模拟客户和基于模型的指标,这些发现展示了对话控制的相对改进,而非临床有效性。这些结果表明,将阶段结构化对话与学习的策略选择相结合是人工智能咨询的一个有前景的方法,尽管临床有效性、安全性和现实世界的实用性仍需进一步的人类评估。
cs.AI / 104 / 2608.22646

CAI-DLLM: Convergence Aware Inference for Diffusion Language Models

CAI-DLLM:针对扩散语言模型的收敛感知推理
Amin, Farhana, Afroz, Sabiha, Nikolopoulos, Dimitrios S.
Abstract
Diffusion language models can generate many tokens in parallel, but they still require repeated denoising steps during inference. This makes generation costly, especially when the model continues to recompute tokens that are already stable. To address these limitations, we propose CAI-DLLM, a training-free inference method that uses first-step confidence to guide denoising and reduce inference time. Specifically, CAI-DLLM commits easy tokens earlier, allocates more denoising steps to harder tokens, and adjusts decoding schedules across output blocks. As it relies only on first-step confidence signals, it does not require retraining, extra predictors, or weight updates. We evaluate CAI-DLLM on LLaDA-8B-Instruct and Dream-7B-Instruct across math, code, reasoning, commonsense, and long-context tasks. CAI-DLLM achieves up to 18.2x wall clock inference speedup on LLaDA GSM8K while improving accuracy from 76.27% to 77.41%, and up to 13.1x speedup on Dream HumanEval while achieving higher pass@1 than no-cache inference, 48.17% compared with 46.95%. On harder reasoning tasks, speedups reach 44.8x, with a largest accuracy drop of 4.4 points, while energy consumption is reduced by up to 95.3%.
Chinese Translation
扩散语言模型可以并行生成多个标记,但在推理过程中仍需重复去噪步骤。这使得生成过程成本高昂,尤其是当模型继续重新计算已经稳定的标记时。为了解决这些局限性,我们提出了CAI-DLLM,一种无训练的推理方法,利用第一步置信度来指导去噪并减少推理时间。具体而言,CAI-DLLM更早地处理简单标记,为更难的标记分配更多的去噪步骤,并在输出块之间调整解码计划。由于它仅依赖于第一步置信度信号,因此不需要重新训练、额外的预测器或权重更新。我们在LLaDA-8B-Instruct和Dream-7B-Instruct上评估CAI-DLLM,涵盖数学、代码、推理、常识和长上下文任务。在LLaDA GSM8K上,CAI-DLLM实现了高达18.2倍的推理速度提升,同时将准确率从76.27%提高到77.41%;在Dream HumanEval上实现了高达13.1倍的速度提升,并且在pass@1方面的表现优于无缓存推理,达到48.17%,而无缓存推理为46.95%。在更困难的推理任务中,速度提升达到44.8倍,最大准确率下降为4.4个百分点,同时能耗减少高达95.3%。
cs.AI / 105 / 2608.22672

A-CPES: A Reference Framework for Agentic AI in Cyber-Physical Energy Systems

A-CPES:网络物理能源系统中自主人工智能的参考框架
Zhang, Xiaoyu, Sun, Qiuye, Xu, Jiachen, Yao, Zhongming, Li, Yushuai
Abstract
Energy system operation contains a loop of work that automation has never taken over: posing the optimization problem the current cycle should solve, disposing of infeasibility, sequencing a solution into interlocked switching orders, assembling evidence no single model holds, negotiating adjustable capacity with many parties, and settling experience into practice. Licensed dispatchers carry all of it in person, and the rising share of variable renewable generation is making that loop turn faster than their number can grow. Agentic AI supplies the abilities it requires, but enters as the outer loop of control: it calls SCED and the other decision models rather than being called by them. We propose A-CPES, three nested rings, an authorization and accountability frame around an agentic control outer loop around a six-layer CPES core. We argue the loop is indivisible, tune where and how tightly it may close, state eight structural failure modes as falsifiable predictions, and specify six governance modules that rebuild the authorization frame until it covers the loop, before the loop starts turning.
Chinese Translation
能源系统的运行包含一个自动化从未接管的工作循环:提出当前周期应解决的优化问题,处理不可行性,将解决方案序列化为互锁的切换顺序,汇集没有单一模型能够持有的证据,与多个参与方协商可调容量,并将经验落实到实践中。持证调度员亲自承担这一切,而可变可再生能源的比例上升使得这一循环的转动速度超过了调度员数量的增长。自主人工智能提供了所需的能力,但作为控制的外环进入:它调用最优经济调度(SCED)和其他决策模型,而不是被它们调用。我们提出了A-CPES,三个嵌套环,围绕自主控制外环构建的授权和问责框架,围绕六层CPES核心。我们认为这个循环是不可分割的,调整其闭合的时机和紧密程度,陈述八种结构性失效模式作为可证伪的预测,并指定六个治理模块,重建授权框架,直到它覆盖循环,然后循环开始转动。
cs.AI / 106 / 2608.22676

Robustness Analysis of Agentic AI to Inconsistent and Incomplete Tool Responses

代理人工智能对不一致和不完整工具响应的鲁棒性分析
Xu, Jiachen, Pedersen, Torben Bach, Yao, Zhongming, Zhang, Xiaoyu, Li, Yushuai
Abstract
Robustness to a bad tool return means answering it in the way that return calls for, which depends on how the tool went wrong. A tool that has failed and a tool that returns a well-formed falsehood are different problems with different remedies. We ask whether the two already differ at the moment the return arrives. This is a qualitative pilot study: we score single decision points rather than running agents to completion. We inject controlled faults into a retail customer-service domain and read two channels off the model's log-probabilities: the likelihood of the returned content under the tool schema alone and under the whole trajectory, and its distribution over the legal actions, read for both shape and where the mass sits. An incomplete return is legible in every case, being improbable under the schema alone in a range no other condition enters, and it moves the mass toward the tools that re-read state wherever there is room to move. An inconsistent return leaves the schema channel untouched and registers in the likelihood comparison on the field whose true value the context already carries verbatim, not on the one whose contradiction runs through the domain policy. The action distribution gives each condition a distinct signature, but orders them by how far the return bears on the next action rather than by fault family. Recognition is therefore asymmetric: each condition is legible in some channel, and no channel is legible on all of them.
Chinese Translation
对不良工具返回的鲁棒性意味着以工具返回所要求的方式进行回应,这取决于工具出错的方式。一个失败的工具和一个返回格式良好的虚假信息的工具是不同的问题,解决方法也不同。我们探讨这两者在返回到达的瞬间是否已经有所不同。这是一项定性初步研究:我们对单个决策点进行评分,而不是运行代理到完成。我们在零售客户服务领域注入控制故障,并从模型的日志概率中读取两个通道:在仅工具模式下返回内容的可能性以及在整个轨迹下的可能性,以及其在合法行动上的分布,分别读取其形状和质量所在的位置。在每种情况下,不完整的返回都是可读的,因为在仅工具模式下的可能性在一个没有其他条件进入的范围内,而它将质量向那些在有空间移动的地方重新读取状态的工具倾斜。不一致的返回则保持工具模式通道不变,并在可能性比较中注册在上下文已经逐字携带的真实值的领域,而不是在其矛盾贯穿领域政策的那个。行动分布为每个条件提供了独特的特征,但根据返回对下一个行动的影响程度而非故障类型对其进行排序。因此,识别是非对称的:每个条件在某个通道中是可读的,而没有任何通道在所有条件上都是可读的。
cs.AI / 107 / 2608.22697

Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf

排名仍然重要吗?当人工智能代理代表我们购物时的位置偏见
Wadi, Davood, Ma, Yu
Abstract
Search rankings are valuable because human attention is scarce and sequential. Higher-placed alternatives are easier to find, so they are examined and bought more often. Consumers are now delegating search to AI agents that can ingest an entire results page at once. Randomizing the order of one hundred hotel listings across 5,000 AI agent sessions, we compare four large language models against human field data. AI agents search more deeply than humans and never decline to buy. Position still predicts which listings are inspected, but weakly and non-monotonically: the middle of a results page has the lowest probability of inspection, not the bottom. Position reaches the choice stage for some models and not others, a heterogeneity that tracks neither provider nor capability. All models nonetheless converge on the same undominated listing. For agentic search, the attributes displayed on a results page matter more than placement within it.
Chinese Translation
搜索排名具有重要价值,因为人类注意力稀缺且是顺序性的。排名靠前的选项更容易被发现,因此被检查和购买的频率更高。消费者现在将搜索任务委托给能够一次性处理整个结果页面的人工智能代理。我们在5000个人工智能代理会话中随机化了一百个酒店列表的顺序,比较了四种大型语言模型与人类实地数据的表现。人工智能代理的搜索深度超过人类,并且从不拒绝购买。位置仍然可以预测哪些列表会被检查,但这种预测是微弱且非单调的:结果页面中间的列表被检查的概率最低,而不是底部。对于某些模型,位置达到了选择阶段,而对于其他模型则没有,这种异质性既不与提供者相关,也不与能力相关。尽管如此,所有模型最终都趋向于同一个未被主导的列表。对于代理搜索,结果页面上显示的属性比其在页面中的位置更为重要。
cs.AI / 108 / 2608.22708

CacheRouter: A Dual-Path Tool Routing Architecture with Cache-Preserving Main-Model Isolation for Long-Tail Tool Discovery

CacheRouter:一种具有缓存保留主模型隔离的双路径工具路由架构,用于长尾工具发现
Zha, Donghui, Xu, Lingwei, Wu, Linxiao, Dong, Yixue, Li, Haochen
Abstract
Tool use in LLM systems faces a structural trade-off. Progressive disclosure keeps the prompt small by showing only the tools relevant to the current task, while prompt caching rewards a request prefix that stays fixed across calls; every change to the visible tool list invalidates the cached prefix. This paper treats the trade-off as a problem of request architecture and proposes a dual-path routing design that assigns tool selection and tool delivery to separate channels. The main model always sees a small, fixed set of core tools, so the head of its request is unchanged across calls; all other tools are reached through an independent routing channel, in which a router sub-model searches the full tool list, selects one tool, executes it, and returns the result. Tool registration is automated from source code and supports runtime updates, so the tool set can grow without modifying the main model's request prefix. The design generalizes progressive disclosure: capabilities are disclosed through the routing channel, and the main model's prefix stays stable. A prototype implementation was exercised on 55 functional queries and a 30-turn dialogue; token-level cache hit rates reached 90.99% and 95.2%, cutting input cost to about 12.0% and 8.0% of a no-cache baseline under DeepSeek's pricing, where cache-hit input tokens cost roughly 1/30 of cache-miss tokens.
Chinese Translation
在大规模语言模型(LLM)系统中,工具的使用面临结构性权衡。渐进式披露通过仅展示与当前任务相关的工具来保持提示的简洁,而提示缓存则奖励在调用之间保持固定的请求前缀;每次对可见工具列表的更改都会使缓存前缀失效。本文将这一权衡视为请求架构问题,提出了一种双路径路由设计,将工具选择和工具交付分配到不同的通道。主模型始终看到一小组固定的核心工具,因此其请求的头部在调用之间保持不变;所有其他工具通过独立的路由通道访问,在该通道中,路由器子模型搜索完整的工具列表,选择一个工具,执行它并返回结果。工具注册从源代码自动化,并支持运行时更新,因此工具集可以在不修改主模型请求前缀的情况下增长。该设计推广了渐进式披露:能力通过路由通道披露,而主模型的前缀保持稳定。原型实现经过55个功能查询和30轮对话的测试;在DeepSeek的定价下,令牌级缓存命中率分别达到了90.99%和95.2%,将输入成本削减至无缓存基线的约12.0%和8.0%,其中缓存命中输入令牌的成本大约是缓存未命中令牌的1/30。
cs.AI / 109 / 2608.22725

SEAM: Shot Entity-Attribute Memory for Consistent Short-Drama Generation at Scale

SEAM:用于大规模一致性短剧生成的镜头实体-属性记忆
Liu, Jiaqi, Ran, Maolin, Lu, Xiaoyang, Wang, Jian, Liu, Weiwen, Lin, Jianghao, Yu, Yong, Zhang, Weinan
Abstract
Short-drama generation has grown into a large, industrialized pipeline, and as it scales from isolated shots to the episode level, visual continuity has become a critical bottleneck. Current agent frameworks generate each shot in isolation, so context drifts across shots and props, character posture, and blocking turn inconsistent. Once assembled, these small discrepancies amplify into severe visual breaks. We present SEAM (Shot Entity-Attribute Memory), a training-free, model-agnostic memory graph that repairs continuity entirely at the prompt-text layer by extracting a multi-dimensional state for every shot, retrieving only causally prior context over the resulting graph, filtering it selectively, and injecting the surviving constraints by natural-language prompt rewriting. We further release SEAM-Bench, a double-blind continuity storyboarding benchmark, on which SEAM raises cross-episode continuity recall from 0.700 to 0.946, generalizes across six mainstream text models, and yields consistent, though not yet significant, gains at the generated-image layer. Deployed as a mandatory stage in CreativeFitting's SEAM-Agent production pipeline over 201 shots, SEAM reaches a 96.5% director-acceptance rate with zero unsafe injections; a conservative counterfactual attributes at least 21.9 percentage points of that rate to its cross-episode memory.
Chinese Translation
短剧生成已发展成为一个庞大的工业化流程,随着其从孤立镜头扩展到剧集层面,视觉连续性成为了一个关键瓶颈。目前的代理框架在孤立的基础上生成每个镜头,因此上下文在镜头之间漂移,导致道具、角色姿势和场景布局的不一致。一旦组装,这些小的差异会放大成严重的视觉断裂。我们提出了SEAM(镜头实体-属性记忆),这是一种无训练、模型无关的记忆图,通过为每个镜头提取多维状态,在生成的图上仅检索因果先前的上下文,选择性地过滤,并通过自然语言提示重写注入存活的约束,从而在提示文本层面完全修复连续性。我们进一步发布了SEAM-Bench,这是一个双盲的连续性故事板基准,在该基准上,SEAM将跨剧集连续性召回率从0.700提高到0.946,能够在六种主流文本模型上进行泛化,并在生成图像层面产生一致的,尽管尚未显著的增益。在CreativeFitting的SEAM-Agent生产流程中,作为一个强制性阶段,SEAM在超过201个镜头中达到了96.5%的导演接受率,且没有不安全的注入;保守的反事实分析将至少21.9个百分点的接受率归因于其跨剧集记忆。
cs.AI / 110 / 2608.22731

LLM-Based Selection of Incongruent Verbal and Nonverbal Behavior for Virtual Humans

基于大语言模型的虚拟人言语与非言语行为不一致选择
Torshizi, Parisa Ghanad, Marsella, Stacy
Abstract
Nonverbal behavior generation systems for virtual agents often take an utterance as input and generate nonverbal behaviors that emphasize or illustrate the content of the verbal channel. However, human nonverbal behavior is shaped by more than the content of the speech. It is also influenced by speaker roles, interpersonal relationships, social context, and the cognitive and emotional states of the interactants. As a result, the nonverbal channel may reinforce, weaken, qualify, or even contradict the verbal channel. It may also reveal internal states that are hidden or only indirectly implied in speech, including emotional "leakage" that may be incidental to the immediate interaction. Modeling this richer relationship between verbal and nonverbal behavior is important for designing virtual agents that exhibit realistic, human-like behavior. It is especially critical in training contexts that require nuanced social interpretation, such as counseling simulations involving virtual patients. Drawing on Ekman's framework of verbal nonverbal relationships, we propose a taxonomy of categories in which mismatches between verbal and nonverbal behavior can occur. We then examine alternative approaches for realizing these behaviors using large language models, focusing on whether LLMs can select contextually appropriate mismatched verbal and nonverbal behaviors from a given dialogue and social interaction context. Finally, we evaluate the resulting behaviors in a human-subject study, assessing whether context-driven nonverbal behavior, when embodied in a virtual human, produces the intended effects on observers.
Chinese Translation
虚拟代理的非言语行为生成系统通常以话语作为输入,生成强调或说明言语内容的非言语行为。然而,人类的非言语行为不仅受言语内容影响,还受到说话者角色、人际关系、社会情境以及交互双方的认知和情感状态的影响。因此,非言语通道可能强化、削弱、限定甚至与言语通道相矛盾。它还可能揭示言语中隐藏或仅间接暗示的内在状态,包括可能与当前交互无关的情绪“泄露”。对言语与非言语行为之间这种更丰富关系的建模,对于设计表现出真实人类行为的虚拟代理至关重要。尤其是在需要细致社会解读的训练情境中,如涉及虚拟患者的咨询模拟,这一点尤为关键。基于Ekman的言语-非言语关系框架,我们提出了言语与非言语行为不匹配可能出现的类别分类法。随后,我们探讨了利用大语言模型(LLMs)实现这些行为的替代方法,重点考察LLMs是否能够从给定的对话和社会交互情境中选择语境适宜的不匹配言语与非言语行为。最后,我们通过一项人类受试者研究评估了所生成的行为,考察当上下文驱动的非言语行为体现在虚拟人身上时,是否能对观察者产生预期的效果。
cs.AI / 111 / 2608.22752

The Compaction Cliff in Long-Running AI Agent Memory

长时间运行的人工智能代理记忆中的压缩 cliff
Zerhoudi, Saber, Mitrovic, Jelena, Granitzer, Michael
Abstract
A safety rule and an episodic log compete for the same tokens in an AI agent's context. When the budget overflows, both are summarized at the same rate; only the rule needs exact wording to remain enforceable. On 20 production agent configurations, Claude Code's /compact prompt on Sonnet 4.6 preserves 53\% of safety rules after one compaction round and 10\% after five. We name this the Compaction Cliff. We address it with Knowledge Triage, a framework that classifies each line of an agent's knowledge base by type and routes each type through its own retention policy. Three deterministic operators implement this triage across the three context-management operations: TypeCompact rewrites items in place under per-type fidelity, TypeDecompose partitions a topic too large to compact safely, replicating in-scope safety rules across partitions, and TypeRetrieve fetches items from external storage with in-scope rules pinned ahead of relevance. On five public corpora, TypeCompact preserves 2--4$\times$ more safety rules than the strongest single-shot LLM compactor at every ratio, with 96\% recall over five rounds. TypeDecompose reaches 0\% locality violations against 93\% under uniform partitioning. TypeRetrieve reaches 100\% recall@50 against 73\% for the best single-shot LLM retriever. On three downstream behavioral benchmarks, we outperform the production Sonnet compactor on medical compliance (paired McNemar $p < 10^{-8}$ on preservation, $N = 200$), the full-policy and hierarchical baselines on retail task pass rate ($p < 0.01$, $N = 115$), and the hierarchical compaction on the airline domain ($p = 0.024$). We release AgentArtifactCorpus (396{,}934 agent configurations from 54{,}628 public GitHub repositories), the classifier, and the reference implementation.
Chinese Translation
安全规则和情节日志在人工智能代理的上下文中争夺相同的标记。当预算溢出时,两者以相同的速度被总结;只有规则需要准确的措辞才能保持可执行性。在20个生产代理配置中,Claude Code的/compact提示在Sonnet 4.6中在一次压缩轮次后保留了53%的安全规则,在五次后保留了10%。我们将其称为压缩 cliff。我们通过知识分类(Knowledge Triage)来解决这个问题,这是一种将代理知识库中的每一行按类型分类并通过各自的保留策略进行路由的框架。三个确定性操作符在三种上下文管理操作中实现这一分类:TypeCompact在每种类型的保真度下就地重写项目,TypeDecompose将一个过大以至于无法安全压缩的话题进行分区,在分区间复制相关的安全规则,TypeRetrieve从外部存储中获取项目,并在相关性之前固定在范围内的规则。在五个公共语料库上,TypeCompact在每个比率下保留的安全规则比最强的单次LLM压缩器多2-4倍,五轮的召回率为96%。TypeDecompose在均匀分区下实现了0%的局部违规率,而在93%下。TypeRetrieve在50的召回率上达到100%,而最佳单次LLM检索器为73%。在三个下游行为基准测试中,我们在医疗合规性上超越了生产Sonnet压缩器(在保留方面配对McNemar $p < 10^{-8}$,$N = 200$),在零售任务通过率上超越了完整政策和层次基线($p < 0.01$,$N = 115$),在航空领域的层次压缩上($p = 0.024$)。我们发布了AgentArtifactCorpus(来自54,628个公共GitHub存储库的396,934个代理配置)、分类器和参考实现。
cs.AI / 112 / 2608.22762

Compositional Chain-of-Relations for Faithful Knowledge Graph Question Answering with Large Language Models

基于关系的组合链用于大语言模型的可信知识图谱问答
Liu, Chenhui, Zhou, Jianpeng, Wang, Jiahai
Abstract
Knowledge graph question answering (KGQA) is a key task for evaluating KG-augmented Large Language Models (LLMs), and complex KGQA that requires multi-hop reasoning is especially challenging. Solving a complex query involves two coupled phases: candidate retrieval, which locates answer candidates over the KG, and constraint handling, which filters these candidates against the query constraints. Faithful reasoning requires grounding both phases in the KG. However, existing agent-based methods ground candidate retrieval through entity-centric exploration, while leaving constraint handling to the LLM's internal knowledge, which leads to two critical limitations. (1) Unreliable entity pruning: entity-centric exploration uses entities as search units and must prune them to a fixed-size subset at each hop. Because entity information in KGs is often incomplete and a fixed-size subset cannot retain all valid entities, such pruning inevitably drops valid entities and ultimately leads to wrong answers. (2) Ungrounded constraint handling: query constraints are resolved from the LLM's internal knowledge rather than the KG, leaving the final answers unverifiable and prone to hallucination. To address these limitations, this paper introduces a relation-centric exploration paradigm, which uses relations rather than entities as search units and thus avoids unreliable entity pruning. Built on this paradigm, this paper proposes Compositional Chain-of-Relations (CCoR), a simple and effective framework that grounds both phases in the KG with two relation chains: a main chain for candidate retrieval and a constraint chain that verifies query constraints through explicit KG exploration. Experiments on four KGQA benchmarks show that CCoR consistently improves accuracy, faithfulness, and efficiency over strong baselines, with more pronounced gains on complex queries.
Chinese Translation
知识图谱问答(KGQA)是评估知识图谱增强的大语言模型(LLMs)的关键任务,而需要多跳推理的复杂KGQA尤其具有挑战性。解决复杂查询涉及两个相互关联的阶段:候选检索,即在知识图谱中定位答案候选项;约束处理,即根据查询约束过滤这些候选项。可信推理要求将这两个阶段都基于知识图谱。然而,现有的基于代理的方法通过以实体为中心的探索来实现候选检索,而将约束处理留给LLM的内部知识,这导致了两个关键限制。(1)不可靠的实体修剪:以实体为中心的探索使用实体作为搜索单元,并必须在每个跳跃中将其修剪为固定大小的子集。由于知识图谱中的实体信息通常是不完整的,固定大小的子集无法保留所有有效实体,因此这种修剪不可避免地会丢失有效实体,最终导致错误答案。(2)未基于知识图谱的约束处理:查询约束是从LLM的内部知识中解决的,而不是来自知识图谱,这使得最终答案无法验证且容易产生幻觉。为了解决这些限制,本文提出了一种基于关系的探索范式,该范式使用关系而不是实体作为搜索单元,从而避免了不可靠的实体修剪。在此范式的基础上,本文提出了组合关系链(Compositional Chain-of-Relations, CCoR),这是一个简单而有效的框架,通过两个关系链将两个阶段都基于知识图谱:一个用于候选检索的主链和一个通过明确的知识图谱探索验证查询约束的约束链。在四个KGQA基准上的实验表明,CCoR在准确性、可信性和效率上始终优于强基线,尤其在复杂查询上取得了更显著的提升。
cs.AI / 113 / 2608.22767

The Retriever Should Remember: Experience-Amortized Reranking for Long-Term Agent Memory

检索器应当记忆:经验摊销的长期代理记忆重排序
Feng, Qi, Ding, Chris, Fan, Jicong
Abstract
Long-term language-model agents accumulate memories across interactions, but their retrievers typically do not accumulate retrieval experience. Semantic retrieval is efficient, but embedding similarity does not always reflect whether a memory contains evidence relevant to the current query. Large language model (LLM) rerankers provide stronger query-conditioned relevance scores, yet stateless reranking repeatedly scores a large candidate pool and discards these scores after each query. We introduce EARM, an experience-amortized reranking framework that treats previously acquired LLM relevance scores as reusable retrieval experience. EARM stores sparse query--memory relevance scores in an online matrix, learns their shared structure through causal matrix completion, and combines a small set of newly observed scores with estimated scores to rerank the remaining candidates. The scoring budget decreases as experience accumulates, changing LLM reranking from a repeated per-query expense into a retrieval capability learned over an agent's lifetime. Experiments on long-term conversational memory show that mixed observed-and-estimated reranking improves answer accuracy over semantic retrieval by up to 6.62% and remains effective when only 17.5% of candidates receive direct LLM relevance scores, thereby substantially reducing the inference overhead of LLM reranking. These results motivate a broader view of agent memory: a long-lived agent should remember not only past content, but also how that content has proved useful for retrieval.
Chinese Translation
长期语言模型代理在交互中积累记忆,但它们的检索器通常不积累检索经验。语义检索效率高,但嵌入相似性并不总能反映记忆是否包含与当前查询相关的证据。大型语言模型(LLM)重排序器提供更强的查询条件相关性评分,然而无状态重排序重复对大量候选池进行评分,并在每次查询后丢弃这些评分。我们提出了EARM(经验摊销重排序),这是一个将先前获得的LLM相关性评分视为可重用检索经验的重排序框架。EARM在一个在线矩阵中存储稀疏的查询-记忆相关性评分,通过因果矩阵补全学习它们的共享结构,并将一小部分新观察到的评分与估计评分结合,以重排序剩余候选项。随着经验的积累,评分预算减少,将LLM重排序从每次查询的重复开销转变为在代理生命周期中学习到的检索能力。关于长期对话记忆的实验表明,混合观察和估计的重排序在答案准确性上比语义检索提高了多达6.62%,并且在仅有17.5%的候选项获得直接LLM相关性评分时仍然有效,从而显著减少了LLM重排序的推理开销。这些结果促使我们对代理记忆有更广泛的看法:一个长期存在的代理不仅应记住过去的内容,还应记住这些内容在检索中如何证明其有效性。
cs.AI / 114 / 2608.22788

TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts

TailSieve:基于部分回滚指导的长尾路由用于大规模语言模型回滚
Xu, Tianqi, Lv, Lu, Huang, Haoyang, Huang, Wenjie, Shen, Zhanming, Shen, Yuhao, Zhang, Baolin, Hu, Xinyi, Ge, Shuang, Dai, Jun, Liu, Tianyu, Yang, Suorong, Li, Zhikai, Bai, Ye, Zhang, Jun, Chen, Lei, Li, Yue, Wan, Mingchen
Abstract
Large-scale rollouts have become a core component of modern LLM systems, spanning reinforcement learning (RL) post-training, on-policy distillation (OPD), and sampling-heavy evaluation pipelines. Unlike online serving, which is typically optimized for request-level latency and throughput, a small number of long-tail generations can dominate the end-to-end makespan of an entire rollout step. In practice, rollout requests are often routed uniformly across replicas, which can place extremely long generations inside high-concurrency decoding batches. To address this, we present TailSieve, a partial-rollout-guided framework that jointly controls tail routing and replica allocation for LLM rollouts. In an idealized setting with known completion lengths, we show that makespan-optimal routing in the long-tail regime combines tail isolation with load balancing, and that a simple top-k policy closely approximates this offline optimum. Leveraging the observation that long-tail prompts tend to remain long-tailed across policy updates, TailSieve uses partial rollouts as a training-free signal for identifying candidate tail groups. A hierarchical controller then jointly adapts the number of isolated groups and the replica split between the tail and bulk pools using collected response-work history and a measured concurrency-throughput model. TailSieve achieves up to 1.67x routing-only speedup over uniform group routing. The resulting low-concurrency tail pool further enables route-specialized speculative decoding with MTP or DFlash, achieving up to 2.59x speedup over uniform routing. Selected prompts are regenerated under the current policy, preserving on-policy generation and avoiding additional routing-induced length bias in steady state.
Chinese Translation
大规模回滚已成为现代大规模语言模型(LLM)系统的核心组成部分,涵盖了强化学习(RL)后训练、在线策略蒸馏(OPD)和重采样评估管道。与通常针对请求级延迟和吞吐量进行优化的在线服务不同,少量的长尾生成可能主导整个回滚步骤的端到端耗时。在实践中,回滚请求通常在副本之间均匀路由,这可能导致极长的生成被放置在高并发解码批次中。为了解决这个问题,我们提出了TailSieve,一个基于部分回滚指导的框架,联合控制LLM回滚的长尾路由和副本分配。在一个理想化的已知完成长度的环境中,我们展示了在长尾状态下,耗时最优的路由结合了长尾隔离与负载均衡,并且简单的top-k策略能够接近这一离线最优解。利用长尾提示在策略更新中往往保持长尾特性的观察,TailSieve使用部分回滚作为无训练信号来识别候选长尾组。一个层次控制器随后联合调整隔离组的数量以及长尾和主流池之间的副本分配,使用收集的响应工作历史和测量的并发-吞吐量模型。TailSieve在路由速度上实现了比均匀组路由高出1.67倍的加速。由此产生的低并发长尾池进一步支持与MTP或DFlash的路由专用推测解码,实现了比均匀路由高出2.59倍的加速。所选提示在当前策略下重新生成,保持了在线策略生成,并避免了稳态下额外的路由引起的长度偏差。
cs.AI / 115 / 2608.22797

Performance of a domain-specific large language model in answering patient questions in psychiatry

特定领域大型语言模型在回答精神病患者问题中的表现
Hish, Alexander J., Nagendran, Arjun, Compton, Scott N.
Abstract
Background This study was designed to evaluate whether a domain-specific large language model (LLM) trained exclusively on patient education resources can answer questions about psychiatric medications, in a manner superior to LLM chatbots. We developed an LLM ("MIND") fine-tuned for clinical fidelity, trained on patient education resources from authoritative medical organizations. Methods We compared the responses of MIND, ChatGPT, and OpenEvidence to patient questions about escitalopram, using two methods: (1) computer analysis according to a rubric measuring accuracy, clarity, completeness, nuance, safety, and referral appropriateness; (2) ratings from N=10 board-licensed psychiatrists on similar metrics. Results When rated by rubric, MIND was rated highest in all domains (p<0.001). When rated by psychiatrists, ChatGPT was rated accurate more often than MIND with a negligible effect size (p=0.021, r=0.073); MIND was rated complete more often than ChatGPT with a small effect size (p<0.001, r=0.160); and MIND and ChatGPT were rated safe with the same frequency (p=0.955, r=0.002). The majority of psychiatrists preferred the responses generated by ChatGPT (57.6%) compared to MIND (42.4%, p=0.003). Conclusions MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time. However, despite MIND's ability to provide more complete responses, psychiatrists preferred ChatGPT's responses. MIND represents a step towards building safe LLM systems to enhance patient education in psychiatry.
Chinese Translation
背景 本研究旨在评估一个专门针对患者教育资源训练的特定领域大型语言模型(LLM)是否能够以优于LLM聊天机器人(chatbots)的方式回答有关精神药物的问题。我们开发了一个名为“MIND”的LLM,经过临床准确性微调,训练数据来自权威医学组织的患者教育资源。方法 我们比较了MIND、ChatGPT和OpenEvidence对患者关于艾司西酞普兰(escitalopram)问题的回答,采用两种方法:(1)根据一个评估标准进行计算机分析,测量准确性、清晰度、完整性、细微差别、安全性和转诊适宜性;(2)由N=10名持牌精神科医生对类似指标进行评分。结果 在评估标准评分中,MIND在所有领域的评分均最高(p<0.001)。在精神科医生评分中,ChatGPT的准确性评分高于MIND,但效果量微不足道(p=0.021,r=0.073);MIND的完整性评分高于ChatGPT,效果量较小(p<0.001,r=0.160);MIND和ChatGPT的安全性评分频率相同(p=0.955,r=0.002)。大多数精神科医生更喜欢ChatGPT生成的回答(57.6%),而不是MIND(42.4%,p=0.003)。结论 MIND能够以精神科医生认为的准确、完整和安全的方式回答许多关于艾司西酞普兰的问题。然而,尽管MIND能够提供更完整的回答,精神科医生仍然更喜欢ChatGPT的回答。MIND代表了构建安全的LLM系统以增强精神病患者教育的一步。
cs.AI / 116 / 2608.22830

Beyond the Harness: End-to-End Optimization of Context Artifacts for Enterprise Text-to-SQL

超越约束:企业文本到SQL的上下文工件的端到端优化
Gwimm, Kate, Eisenach, Carson
Abstract
Deploying LLMs for enterprise Text-to-SQL is bottlenecked less by the model than by what context reaches it: business logic spans thousands of tables, and no model can ingest a full catalog at once. We argue that the most effective place to intervene is therefore the \emph{knowledge-base context} the model consumes, and that this context should be \emph{constructed} from historical usage rather than tuned for as a fixed input. Using a query-DAG decomposition--the same family of intermediates that enterprise benchmarks like BEAVER annotate, here recovered from production SQL--we compare the value of oracle query graphs versus retrieved knowledge-base context. In this ablation, retrieved knowledge-base context provides the largest marginal improvement when added to the full oracle graph. Building on this, we optimize a distillation procedure that turns historical query profiles into reusable SQL reference cards. On a benchmark of 5176 production queries from a major online retailer, optimizing these context artifacts yields larger gains (${\sim}12$--$25\%$ AST similarity) than optimizing the retrieval harness (${\sim}3$--$12\%$). On the public BEAVER benchmark, which lacks the production-usage signals available in our internal setting, the picture is more mixed: table cards alone perform about the same as raw historical SQL. The best optimized variant retrieves both cards and raw SQL, scoring $9.00\%$ versus $6.33\%$ (p-value $0.12$) for the comparable baseline on a held-out $N{=}300$ subset, using retrieved context and harness changes but no agentic loop.
Chinese Translation
在企业文本到SQL的应用中,部署大型语言模型(LLMs)所面临的瓶颈更多地来自于模型所接收到的上下文,而非模型本身:业务逻辑跨越数千个表格,任何模型都无法一次性处理完整的目录。因此,我们认为最有效的干预点在于模型所消耗的 extit{知识库上下文},并且这一上下文应当基于历史使用情况进行 extit{构建},而不是作为固定输入进行调优。通过查询有向无环图(query-DAG)分解——与企业基准如BEAVER所注释的中间结果同属一类,这里从生产SQL中恢复——我们比较了oracle查询图与检索的知识库上下文的价值。在这一消融实验中,检索的知识库上下文在添加到完整的oracle图时提供了最大的边际改进。基于此,我们优化了一种蒸馏程序,将历史查询配置文件转化为可重用的SQL参考卡。在来自一家大型在线零售商的5176个生产查询的基准测试中,优化这些上下文工件所带来的增益(约12%到25%的抽象语法树相似度)大于优化检索约束(约3%到12%)。在公共BEAVER基准上,由于缺乏我们内部环境中可用的生产使用信号,结果则更为复杂:仅使用表卡的表现与原始历史SQL相当。最佳优化变体同时检索卡片和原始SQL,在一个保留的N=300子集上得分为9.00%,而可比基准为6.33%(p值为0.12),使用了检索的上下文和约束变化,但没有代理循环。
cs.AI / 117 / 2608.22832

Let the Bullets Fly: Multimodal Fake News Detection with Temporal-Aligned Generative Danmaku

让子弹飞:基于时间对齐生成弹幕的多模态假新闻检测
Luo, Xiansheng, Zhang, Chaowei, Zhang, Zewei, Zhu, Yi, Qiang, Jipeng
Abstract
The social interactions among crowds via \textit{Danmaku} (a.k.a., bullet comments) on modern multimedia platforms can facilitate both viewpoint conflicts and consensus, providing fine-grained discriminative social signals that can benefit fake news detection. However, the inherent accumulation latency of \textit{Danmaku} in real-world scenarios violates the real-time necessity of fake news detection, making the studies of \textit{Danmaku}-related fake news detection underexplored. To break this violation, we simulate this temporal-aware user interactive process by proposing a novel temporal \textbf{Gen}erative \textbf{da}nmaku framework, called \textbf{Genda}, which consists of: (1) a \textit{Danmaku} Trigger for predicting the timing and intensity of user reactions; and (2) a \textit{Danmaku} Generator for synthesizing corresponding semantic and emotional expressions, thereby mutually constructing a temporally aligned and human-like pseudo \textit{Danmaku} streams. To make the generated \textit{Danmaku} useful for identifying fake news videos, we further design a \textit{Danmaku}-guided Temporal Multimodal fake news detection model - \textbf{DM-FEND}, which enables fine-grained multimodal interactions among video, audio, text, and \textit{Danmaku}, enhancing dynamic modalities alignment and semantic noise inhibition. The experimental results demonstrate that \emph{DM-FEND} consistently outperforms state-of-the-art baselines across both Chinese (FakeSV) and English (FakeTT) benchmarks. Further ablations validate the crucial role of temporal \textit{Danmaku} modeling in enhancing robustness and discriminative capability. Finally, this study offers a bright and robust solution for multimodal fake news detection in modern social interactive fashions by bridging the temporal inconsistency between news and user behaviors.
Chinese Translation
现代多媒体平台上通过弹幕(又称子弹评论)进行的群体社交互动既可以促进观点冲突,也可以达成共识,提供细粒度的区分性社交信号,从而有助于假新闻检测。然而,弹幕在现实场景中的固有累积延迟违反了假新闻检测的实时性要求,使得与弹幕相关的假新闻检测研究尚未得到充分探索。为了解决这一问题,我们通过提出一种新颖的时间感知用户互动过程来模拟这一过程,构建了一个名为Genda的时间生成弹幕框架,该框架包括:(1) 一个弹幕触发器,用于预测用户反应的时机和强度;(2) 一个弹幕生成器,用于合成相应的语义和情感表达,从而相互构建时间对齐且类人化的伪弹幕流。为了使生成的弹幕在识别假新闻视频中发挥作用,我们进一步设计了一个弹幕引导的时间多模态假新闻检测模型——DM-FEND,该模型实现了视频、音频、文本和弹幕之间的细粒度多模态交互,增强了动态模态对齐和语义噪声抑制。实验结果表明,DM-FEND在中文(FakeSV)和英文(FakeTT)基准测试中均持续超越最先进的基线。此外,进一步的消融实验验证了时间弹幕建模在增强模型鲁棒性和区分能力方面的关键作用。最后,本研究为现代社交互动方式中的多模态假新闻检测提供了一个光明且稳健的解决方案,弥合了新闻与用户行为之间的时间不一致性。
cs.AI / 118 / 2608.22842

FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks

FinixDoc:重新思考金融文档解析超越饱和基准
Wang, Hang, Zhang, Jin, Xu, Guoliang, Lu, Pengyue, Li, Yao, Zhang, Zijiao, Huang, Tianyu, Xiong, Weiqi, Wang, Yulong, Lu, Chuqiao, Huang, Wenkang, Yang, Kai, Li, Yadong, Li, Hui, Xu, Xingzhong, Xu, Xiao
Abstract
Financial document parsing requires accuracy, structural consistency, and verifiability that current benchmarks often fail to reflect. We present FinixDoc, an end-to-end agentic parsing system for real-world financial documents, with FinixDoc-VL, a 4B-scale vision-language model built on Qwen3-VL-4B, as its core parser. To characterize the gap between benchmark and deployment performance, we introduce a Document Parsing Capability Matrix organized along two practical axes: visual quality and document scale. Guided by this matrix, FinixDoc-VL is trained with a domain-adapted recipe combining homoglyph-aware contrastive learning and multi-stage reinforcement learning with composite domain-specific rewards. To better leverage our accumulated advantage in low-quality financial-document data and support large-scale, high-quality data production, we further build a human-in-the-loop Data Factory pipeline with confidence-aware expert review. For evaluation, we construct FinixDocBench, a financial-domain evaluation suite covering digital-native, camera-captured, ultra-large-page, and internal-workflow scenarios, with a compliance-reviewed subset released alongside this technical report. On its main subsets, FinixDoc-VL achieves the highest overall score (81.43) among evaluated baselines, outperforming the next-best open-source model by 5.13 points, with the largest gains on internal financial workflows (FinixInner: 84.08 vs. 78.73).
Chinese Translation
金融文档解析需要准确性、结构一致性和可验证性,而当前的基准往往无法反映这些要求。我们提出了FinixDoc,一个针对现实世界金融文档的端到端智能解析系统,其核心解析器是基于Qwen3-VL-4B构建的4B规模视觉-语言模型FinixDoc-VL。为了表征基准与实际部署性能之间的差距,我们引入了一个文档解析能力矩阵,该矩阵沿着视觉质量和文档规模两个实际轴进行组织。在该矩阵的指导下,FinixDoc-VL采用了一种领域适应的训练方案,结合了同形异义词感知对比学习和多阶段强化学习,使用复合领域特定奖励。为了更好地利用我们在低质量金融文档数据中的积累优势,并支持大规模、高质量数据的生产,我们进一步构建了一个人机协作的数据工厂管道,配备了基于信心的专家审查。为了评估,我们构建了FinixDocBench,这是一个覆盖数字原生、相机捕获、超大页面和内部工作流程场景的金融领域评估套件,并在本技术报告中发布了经过合规审查的子集。在其主要子集上,FinixDoc-VL在评估的基准中获得了最高的整体得分(81.43),比下一个最佳的开源模型高出5.13分,在内部金融工作流程上获得了最大的提升(FinixInner:84.08对比78.73)。
cs.AI / 119 / 2608.22847

GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis

GSAR:用于自我演化数据合成的移动图形用户界面代理的目标状态锚奖励
Zhang, Long, Chen, Yuhan, Zhang, Chaoran, Cao, Wanxia, Huang, Kun, Gao, Pengzhi, Liu, Wei, Luan, Jian, Li, Chenliang, Zou, Lixin
Abstract
Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their training is bottlenecked by two fundamental issues: current data synthesis methods for GUI Agents rely on specific environments and struggle to generate diverse data, while existing evaluators either suffer from limited scalability or provide inaccurate and unreliable reward signals. To overcome these challenges, we introduce GSAR (Goal-State-Anchor Reward), a RL reward framework that supports scalable task generation and delivers reliable reward signals for stable and efficient policy optimization. Our approach features self-evolving data synthesis, which produces multiple environments through task execution and generates diverse tasks and goal states. Complementing this, a state-anchor mechanism automatically annotates task-relevant UI elements in successful goal states as reference anchors. During RL training, these reference anchors provide accurate, scalable reward signals that substantially enhance efficiency. Extensive evaluations demonstrate that our framework achieves over 90% accuracy on offline trajectory verification and performs closest to rule-based methods. Furthermore, agents trained using our reward framework exhibit strong performance on both AndroidWorld and our constructed benchmark, establishing a scalable approach for GUI agent training.
Chinese Translation
基于视觉-语言模型(VLM)的图形用户界面(GUI)代理在在线强化学习(RL)中具有显著的潜在收益。然而,它们的训练受到两个基本问题的瓶颈:当前的GUI代理数据合成方法依赖于特定环境,且难以生成多样化的数据,而现有的评估器要么在可扩展性上受限,要么提供不准确和不可靠的奖励信号。为了解决这些挑战,我们提出了GSAR(目标状态锚奖励),这是一个支持可扩展任务生成并提供可靠奖励信号以实现稳定高效策略优化的RL奖励框架。我们的方法具有自我演化数据合成的特点,通过任务执行生成多个环境,并产生多样化的任务和目标状态。与此相辅相成的是,状态锚机制自动标注成功目标状态中的任务相关用户界面元素作为参考锚。在RL训练过程中,这些参考锚提供准确、可扩展的奖励信号,显著提高了效率。广泛的评估表明,我们的框架在离线轨迹验证中实现了超过90%的准确率,并且表现最接近基于规则的方法。此外,使用我们的奖励框架训练的代理在AndroidWorld和我们构建的基准测试中表现出强大的性能,确立了一种可扩展的GUI代理训练方法。
cs.AI / 120 / 2608.22852

Your AI, On a Dial: Controlling Investment Bias in LLMs with a Single Neuron

你的人工智能,拨动控制:通过单个神经元控制大型语言模型中的投资偏差
Park, Sahong, Park, Suhwan, Lee, Hoyoung, Kwon, Gakyung, Ahn, Wonbin, Choi, Jaewon, Lopez-Lira, Alejandro, Kim, Yoon, Choi, Chanyeol, Kong, Hyeongwoo, Lee, Yongjae
Abstract
Large language models (LLMs) are increasingly used in investment decision-making, yet prior work shows that they exhibit systematic, model-specific investment preferences. We study whether a model's overall investment stance can be calibrated to a specified direction and strength. We introduce an investment-bias dial, an inference-time intervention on a single neuron that continuously adjusts a model-level decision prior---its overall tendency toward buying or selling---without targeting specific firms or investment attributes. Using matched positive and negative evidence, we evaluate five open-weight LLMs and find that the dial produces monotonic changes in investment stance without modifying prompts or model parameters. At the response level, the dial shifts both investment decisions and the evidential emphasis of generated rationales under identical inputs. In an agentic retrieval setting, the dial also changes what information the model searches for, which evidence it selects, and which evidence is reflected in its final analysis. In a long-context evaluation, the dial maintains stable stance control as context length increases, whereas a matched system-prompt instruction progressively attenuates. We further show that changes in the dial propagate to security rankings and downstream portfolio composition in an exploratory backtest. Overall, our results show that an LLM's aggregate investment stance can be calibrated toward a specified target at inference time.
Chinese Translation
大型语言模型(LLMs)在投资决策中越来越多地被使用,但先前的研究表明,它们表现出系统性的、特定模型的投资偏好。我们研究了模型的整体投资立场是否可以调整到指定的方向和强度。我们引入了一种投资偏差拨盘,这是一种在推理时对单个神经元的干预,能够持续调整模型层面的决策先验——即其整体倾向于买入或卖出——而不针对特定公司或投资属性。通过匹配的正面和负面证据,我们评估了五个开放权重的LLMs,发现拨盘在不修改提示或模型参数的情况下,能够产生投资立场的单调变化。在响应层面,拨盘在相同输入下同时改变投资决策和生成理由的证据强调。在代理检索设置中,拨盘还改变了模型搜索的信息、选择的证据以及最终分析中反映的证据。在长上下文评估中,拨盘在上下文长度增加时保持稳定的立场控制,而匹配的系统提示指令则逐渐减弱。我们进一步展示了拨盘的变化传播到证券排名和下游投资组合构成的探索性回测中。总体而言,我们的结果表明,LLM的整体投资立场可以在推理时调整到指定目标。
cs.AI / 121 / 2608.22887

Proxy reliance in large language model decisions is uncalibrated to predictive evidence

大型语言模型决策中的代理依赖与预测证据不匹配
Wu, Zengqing, Xiao, Chuan
Abstract
Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use. Current audits ask whether decisions change when demographics change. But attributes correlated with a protected group carry predictive value, so a changed decision can be discrimination or sound inference. We measure causal proxy effects in four LLMs on a clinical-ranking task with known ground truth, where the reliance the evidence warrants can be computed exactly and used as the reference. One audit signal yields three verdicts: over-reliance, warranted and under-reliance. Under neutral labels every model relies on proxies with no information. Informative proxies draw all three. Social field names push reliance down, below the reference in one model. Two findings explain this. Reliance severely undertracks the evidence, and social-label suppression is fragile, since in-context examples raise it above zero in every model. Accuracy-based evaluation detects none of this.
Chinese Translation
大型语言模型(LLMs)正在参与分诊和贷款决策,其中必须区分与任务相关的推理与不当的代理使用。目前的审计关注于当人口统计特征变化时,决策是否会改变。然而,与受保护群体相关的属性具有预测价值,因此,决策的变化可能是歧视或合理推理。我们在四个LLM上测量了因果代理效应,针对一个具有已知真实情况的临床排名任务,在该任务中,可以准确计算出证据所需的依赖程度,并作为参考。一种审计信号产生三种裁决:过度依赖、合理依赖和不足依赖。在中性标签下,每个模型都依赖于没有信息的代理。信息性代理则吸引所有三种裁决。社会领域名称降低了依赖程度,在一个模型中低于参考值。两个发现对此进行了说明。依赖程度严重低估了证据,而社会标签的抑制是脆弱的,因为上下文示例在每个模型中都将其提升至零以上。基于准确性的评估未能检测到这一切。
cs.AI / 122 / 2608.22899

CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents

CDEG:为长时间诊断代理学习决策关键证据
Dai, Xiwei, Meng, Zijie, Fan, Zhiting, Tang, Yixuan, Niu, Ziru, Liu, Zuozhu
Abstract
Unlike static medical question answering, long-horizon diagnosis captures the sequential nature of clinical practice: evidence is progressively acquired, integrated, and evaluated over multiple rounds of interaction before reaching a final diagnosis. However, existing doctor agents often fail when critical evidence is either not acquired or not adequately incorporated into diagnostic reasoning. Recent agentic approaches attempt to address these failures by reusing historical trajectories or distilled memories. But their diagnostic gains remain constrained because such experience may contain noisy or incidental information and is typically reused without validating which evidence actually drives diagnostic decisions. To address this limitation, we introduce CDEG, a graph-based framework that learns reusable decision-critical evidence from historical diagnostic trajectories. CDEG contrasts successful and failed trajectories from the same case to identify candidate evidence, validates their diagnostic impact through controlled counterfactual interventions, and organizes the resulting diagnosis--evidence--action relations into a structured graph. During inference, CDEG tracks the evolving patient evidence state to retrieve relevant diagnostic relations and selectively guide missing evidence acquisition or overlooked evidence reappraisal. Across in-domain and out-of-distribution benchmarks with multiple doctor agent backbones, CDEG consistently improves diagnostic performance, achieving up to an 11.5% accuracy gain over vanilla agents. These results demonstrate that reliable long-horizon diagnosis requires moving beyond trajectory-level experience reuse toward evidence-level learning of the factors that truly shape clinical decisions.
Chinese Translation
与静态医学问答不同,长时间诊断捕捉了临床实践的顺序特性:证据在多轮互动中逐步获取、整合和评估,最终形成诊断。然而,现有的医生代理在关键证据未被获取或未被充分纳入诊断推理时常常失败。近期的代理方法试图通过重用历史轨迹或提炼记忆来解决这些失败。但由于这些经验可能包含噪声或偶然信息,并且通常在未验证哪些证据实际驱动诊断决策的情况下被重用,因此其诊断收益仍然受到限制。为了解决这一局限性,我们引入了CDEG,一个基于图的框架,旨在从历史诊断轨迹中学习可重用的决策关键证据。CDEG对比同一案例中的成功和失败轨迹,以识别候选证据,通过控制反事实干预验证其诊断影响,并将生成的诊断-证据-行动关系组织成结构化图。在推理过程中,CDEG跟踪不断变化的患者证据状态,以检索相关的诊断关系,并选择性地指导缺失证据的获取或被忽视证据的重新评估。在多个医生代理基础模型的领域内和分布外基准测试中,CDEG始终提高诊断性能,相较于普通代理实现了高达11.5%的准确率提升。这些结果表明,可靠的长时间诊断需要超越轨迹级经验重用,朝着证据级学习真正塑造临床决策的因素迈进。
cs.AI / 123 / 2608.22920

Beyond Observed Auxiliary Relations: Environment-Conditioned Modeling for Multi-Behavior Recommendation

超越观察到的辅助关系:环境条件建模用于多行为推荐
Lee, Seunghan, Yoo, Hyunsik, Kang, Jian, Yoon, Susik, Kang, SeongKu
Abstract
Multi-behavior recommendation (MBR) leverages auxiliary behavioral signals, such as clicks and add-to-cart, to enhance target behavior prediction like purchases. While recent graph neural network-based approaches have achieved strong performance by systematically propagating auxiliary behavior signals, they still suffer from two fundamental challenges inherent to auxiliary behaviors: (1) missing auxiliary signals, which hinder generalization to items without auxiliary observations, and (2) unreliable auxiliary signals, which amplify noise misaligned with the target behavior. To address these challenges in a unified manner, we propose BOAR, an environment-conditioned MBR framework that addresses missing and unreliable auxiliary signals through two complementary modules conditioned on auxiliary observability. Extensive experiments demonstrate that BOAR consistently outperforms state-of-the-art baselines, achieving up to 7.82% gains in HR@10 overall and up to 44.2% gains for target items without auxiliary observations, highlighting its ability to capture hidden preferences beyond observed auxiliary relations. Our code is available at: https://github.com/LSH0411/BOAR.
Chinese Translation
多行为推荐(MBR)利用辅助行为信号,如点击和加入购物车,来增强目标行为预测(如购买)。尽管最近基于图神经网络的方法通过系统性地传播辅助行为信号取得了良好的性能,但它们仍然面临两个固有的挑战:(1)缺失的辅助信号,这会阻碍对没有辅助观察的项目的泛化;(2)不可靠的辅助信号,这会放大与目标行为不一致的噪声。为了解决这些挑战,我们提出了BOAR,一个环境条件的MBR框架,通过两个互补模块来处理缺失和不可靠的辅助信号,这些模块以辅助可观察性为条件。大量实验表明,BOAR在各项指标上始终优于最先进的基线,在HR@10上整体提升高达7.82%,对于没有辅助观察的目标项目提升高达44.2%,突显了其捕捉超越观察到的辅助关系的隐藏偏好的能力。我们的代码可在以下链接获取:https://github.com/LSH0411/BOAR。
cs.AI / 124 / 2608.22930

Concepts for Securing Agentic AI Coding and the Terok Environment

确保代理人工智能编码及Terok环境的概念
Vyskočil, Jiří, Pöschel, Franz, Knüpfer, Andreas
Abstract
Agentic AI is a fascinating new tool for software development. It is a huge step forward compared to "conventional" AI assisted coding, which in turn was a considerable breakthrough earlier. AI support through LLMs is a young and very fast-moving field. The "conventional" (non-agentic) flavor became useful and productive in early 2025 (around 18 months ago) and the agentic flavor followed in fall 2025 (approximately 9 months ago). Besides all its benefits and potential, it also carries some fundamental risks for IT security. And the agentic approach added very severe risks while making others much more dangerous. With all the motivation to explore this fascinating new tool we should not ignore the risks but actively address them. We present (I) an assessment of the IT security risks, (II) a concept for mitigating them without breaking its benefits, and (III) an overview about an implementation of our concept. In this very dynamic field this is likely not the final and once-and-for-all answer to the identified issues but still a substantial step forward in responsible usage of Agentic AI for software development. It should also be a contribution to the community to allow early and eager evaluation of the potential of agentic AI for software development without actually suffering from its implied IT security risks.
Chinese Translation
代理人工智能是一种引人入胜的新工具,用于软件开发。与“传统”人工智能辅助编码相比,它是一个巨大的进步,而“传统”人工智能辅助编码在早期也是一个相当重要的突破。通过大语言模型(LLMs)的人工智能支持是一个年轻且快速发展的领域。“传统”(非代理)版本在2025年初变得有用且富有成效(大约18个月前),而代理版本则在2025年秋季出现(大约9个月前)。尽管它带来了诸多好处和潜力,但也存在一些对信息技术安全的基本风险。而代理方法则增加了非常严重的风险,同时使其他风险变得更加危险。在探索这一引人入胜的新工具的同时,我们不应忽视这些风险,而应积极应对。我们提出了(I)对信息技术安全风险的评估,(II)在不破坏其好处的情况下减轻这些风险的概念,以及(III)我们概念实施的概述。在这个动态变化的领域,这可能不是对已识别问题的最终和一次性答案,但仍然是负责任地使用代理人工智能进行软件开发的重要一步。它也应成为社区的贡献,以便在不实际遭受其隐含的信息技术安全风险的情况下,允许对代理人工智能在软件开发中的潜力进行早期和热切的评估。
cs.AI / 125 / 2608.22960

What Process Evaluation of Coding Agents Actually Measures: Action, Task, and Step Are Three Different Levels

编码代理的过程评估实际测量了什么:行动、任务和步骤是三个不同的层次
He, Jiawei, Shi, Mengyu, jia, Jie, Yang, Xikai, Sun, Dong
Abstract
Coding agents are increasingly evaluated not only by whether they solve a task, but also by how they execute it. However, existing process-level evaluations often treat action prediction, task uncertainty, and step attribution as if they were the same problem, which makes it unclear what such evaluations actually measure. In this paper, we introduce a measurement framework for process evaluation in coding agents and instantiate step-level causal attribution with SCAE, a replay-based estimator derived from a structural causal model of agent execution. Our framework combines prefix-conditioned identification, replay/intervention-based estimation, and controlled judge-information manipulation to study process evaluation at the action, task, and step levels. Experiments on 499 file-localization episodes from 12 repositories show that next actions are driven primarily by execution provenance rather than code-graph transitions, execution uncertainty is structured at the task rather than step level, and full-trace judges exhibit systematic collider bias, suggesting that current process evaluation often measures semantic relevance rather than certified causal contribution.
Chinese Translation
编码代理的评估越来越多地不仅关注它们是否解决了任务,还关注它们如何执行任务。然而,现有的过程级评估往往将行动预测、任务不确定性和步骤归因视为同一问题,这使得这些评估实际测量的内容变得不明确。本文提出了一种用于编码代理过程评估的测量框架,并通过基于重放的估计器SCAE(源自代理执行的结构性因果模型)实例化步骤级因果归因。我们的框架结合了前缀条件识别、基于重放/干预的估计和受控的评判信息操控,以研究行动、任务和步骤层次的过程评估。对来自12个代码库的499个文件定位实例的实验表明,下一步行动主要受执行来源驱动,而非代码图转变,执行不确定性在任务层面而非步骤层面上是有结构的,且全追踪评判者表现出系统性的碰撞偏差,这表明当前的过程评估往往测量的是语义相关性而非经过验证的因果贡献。
cs.AI / 126 / 2608.22963

Buried in Textual Debt: Context Pruning with Visual Evidence Preservation for MLLM Agents

文本债务中的埋藏:具有视觉证据保留的上下文修剪用于多模态大语言模型代理
Huang, Yuchen, Li, Sijia, Zhang, Jun, Fung, Yi R.
Abstract
Multimodal Large Language Models (MLLMs) are increasingly deployed as multi-step agents, where explicit reasoning supports task decomposition and tool coordination but also accumulates self-generated text. Over long trajectories, this text can dominate the context and suppress visual evidence, creating textual debt. We observe that reasoning becomes redundant once task-relevant visual evidence is grounded, while stale hypotheses can misguide later inference when grounding remains uncertain. Pruning must therefore remove redundant text without discarding visual evidence. We propose SPARE, a Kullback--Leibler (KL)-guided framework for pruning accumulated reasoning in multimodal tool-use agents. SPARE uses a compact task-state summary as privileged diagnostic context. For each candidate segment, it replays the same model under the original and summary-conditioned contexts. Reverse-KL divergence from on-policy self-distillation (OPSD) then tests whether the summary sufficiently covers the segment without disrupting future reasoning. We further fine-tune the summarizer with supervised fine-tuning (SFT), enabling more compact summaries, broader coverage, and more aggressive pruning. Across multi-step visual tool-use benchmarks, SPARE achieves the highest average accuracy among pruning methods while removing 37.89--64.58\% of reasoning tokens. This favorable accuracy--context trade-off shows that reducing textual dominance restores reliance on visual evidence and mitigates over-conditioning on self-generated language.
Chinese Translation
多模态大语言模型(MLLMs)越来越多地被用作多步骤代理,其中明确的推理支持任务分解和工具协调,但也会积累自生成的文本。在较长的轨迹中,这些文本可能主导上下文并压制视觉证据,形成文本债务。我们观察到,一旦任务相关的视觉证据得到确认,推理就变得多余,而过时的假设在确认仍不确定时可能误导后续推理。因此,修剪必须去除冗余文本而不丢弃视觉证据。我们提出了SPARE,一个基于Kullback-Leibler(KL)引导的框架,用于修剪多模态工具使用代理中积累的推理。SPARE使用紧凑的任务状态摘要作为特权诊断上下文。对于每个候选段,它在原始和摘要条件上下文下重放相同的模型。然后,通过在政策自蒸馏(OPSD)中计算反向KL散度,测试摘要是否足够覆盖该段而不干扰未来的推理。我们进一步通过监督微调(SFT)对摘要生成器进行微调,以实现更紧凑的摘要、更广泛的覆盖和更激进的修剪。在多步骤视觉工具使用基准测试中,SPARE在修剪方法中实现了最高的平均准确率,同时去除了37.89%至64.58%的推理标记。这种有利的准确性与上下文的权衡表明,减少文本主导性恢复了对视觉证据的依赖,并减轻了对自生成语言的过度条件化。
cs.AI / 127 / 2608.22971

ParallelWorld: Test-Time Scaling for Embodied Reasoning

ParallelWorld:具身推理的测试时刻扩展
Chen, Min, Zhang, Shengjun, Li, Yuxin, Zhang, Zhang, Fei, Xin, Xia, Chong, Duan, Yueqi
Abstract
Embodied Reasoning constitutes a fundamental capability of embodied intelligence, serving as the basis for autonomous perception, reasoning, and interaction within physical environments. Recent studies have shifted the paradigm of embodied reasoning from static perception toward dynamic exploration, where agents acquire task-relevant information through interactions with the environment. However, existing active reasoning approaches generally generate exploration trajectories incrementally without long-horizon planning. Even recently emerged test-time scaling frameworks often resort to myopic, single-step lookaheads, which struggle to resolve the delayed feedback inherent in complex, occluded spatial environments. To address this limitation, we propose ParallelWorld, a multi-horizon test-time scaling framework for embodied reasoning. Instead of greedy, single-step trials, ParallelWorld empowers agents to simulate and evaluate multi-step future trajectories in parallel before committing to an action. Specifically, we introduce a verifier-guided tree-search paradigm. Starting from the current state, ParallelWorld branches into multiple parallel trajectories and rolls them out continuously across a multi-step horizon. At each simulation step, a verifier agent evaluates the intermediate state transitions, dynamically pruning unpromising branches and prioritizing paths with the highest information gain. Once the multi-step prospective simulation is complete, the agent synthesizes the long-horizon outcomes to commit to the optimal action sequence. Finally, an answer agent performs reasoning over the selected trajectory to produce the final reasoning. Extensive experiments on ESI-Bench demonstrate that ParallelWorld consistently improves active perception and reasoning performance.
Chinese Translation
具身推理是具身智能的一项基本能力,构成了在物理环境中进行自主感知、推理和交互的基础。近期的研究已将具身推理的范式从静态感知转向动态探索,代理通过与环境的交互获取与任务相关的信息。然而,现有的主动推理方法通常以增量方式生成探索轨迹,而没有进行长远规划。即使是最近出现的测试时刻扩展框架,往往也依赖于短视的单步前瞻,这在复杂且遮挡的空间环境中难以解决固有的延迟反馈。为了解决这一局限性,我们提出了ParallelWorld,一个用于具身推理的多视角测试时刻扩展框架。与贪婪的单步试验不同,ParallelWorld使代理能够在采取行动之前并行模拟和评估多步未来轨迹。具体而言,我们引入了一种基于验证者引导的树搜索范式。从当前状态出发,ParallelWorld分支为多个并行轨迹,并在多步视野中持续展开。在每个模拟步骤中,验证者代理评估中间状态转变,动态修剪不具前景的分支,并优先考虑信息增益最高的路径。一旦多步前景模拟完成,代理合成长远结果以确定最佳行动序列。最后,答案代理对所选轨迹进行推理,以生成最终推理。对ESI-Bench的广泛实验表明,ParallelWorld始终提高了主动感知和推理性能。
cs.AI / 128 / 2608.22974

Toward Effective and Reliable LLM Agents via Dynamic Ontology

通过动态本体构建有效且可靠的大型语言模型代理
Zhang, Xiaohui, Sun, Zequn, Yang, Chengyuan, Cui, Yuanning, Guo, Lingbing, Hu, Wei
Abstract
Large language model (LLM) agents rely heavily on knowledge encoded in model parameters or presented as unstructured context. In domain-specific tasks, this leaves important semantic connections implicit. This often results in incomplete evidence use and brittle multi-step decisions. Ontologies offer a way to externalize domain concepts and relations as machine-interpretable structures, but constructing task-usable ontologies traditionally requires substantial effort from domain experts and is difficult to scale. Automatic construction is also challenging: an ontology that appears semantically plausible may not contain the relational structures needed for actual decision making. We present OaK, an ontology-as-a-kernel framework that dynamically constructs and refines task-oriented ontologies for LLM agents. Given task requirements and training data, OaK constructs an ontology and its knowledge graph, generates task-adaptation functions for graph reasoning, and uses judge feedback to iteratively refine both. By making relevant concepts and relations explicit, the ontology grounds knowledge retrieval and multi-step decision making. We evaluate OaK on TravelPlanner, CRMArenaPro, and ToolQA. Results show that OaK improves standard LLM agents, strengthens evidence grounding, and boosts the reliability of multi-step reasoning.
Chinese Translation
大型语言模型(LLM)代理在很大程度上依赖于编码在模型参数中的知识或以非结构化上下文呈现的知识。在特定领域的任务中,这使得重要的语义连接变得隐含。这通常导致证据使用不完整和脆弱的多步骤决策。本体提供了一种将领域概念和关系外部化为机器可解释结构的方法,但传统上构建可用于任务的本体需要领域专家的 substantial 努力,并且难以扩展。自动构建也面临挑战:一个看似语义上合理的本体可能不包含实际决策所需的关系结构。我们提出了 OaK,一个本体作为核心的框架,动态构建和完善面向任务的本体以供 LLM 代理使用。根据任务要求和训练数据,OaK 构建一个本体及其知识图谱,生成用于图推理的任务适应函数,并利用评审反馈迭代地完善这两者。通过使相关概念和关系明确化,本体为知识检索和多步骤决策提供了基础。我们在 TravelPlanner、CRMArenaPro 和 ToolQA 上评估了 OaK。结果表明,OaK 改善了标准 LLM 代理,增强了证据基础,并提高了多步骤推理的可靠性。
cs.AI / 129 / 2608.22975

Budget-Constrained Embodied Perception: Four Resource Walls and a Pre-Registered Evaluation of Access-Structured Perception on Open Models at less than 31B

预算约束下的具身感知:四个资源壁垒及对开放模型的访问结构感知的预注册评估(模型参数少于31B)
Lin, Defu, Chen, Wenhui, Lin, Ziyao, Chen, Jianlin, Long, Peiji, Vong, Chi Man
Abstract
Embodied multimodal agents must answer from growing observation streams under a fixed per-decision token budget. We formalize this constraint through four resource walls: a perceptual Shannon wall for bounded state, a horizon wall for query-independent frame selection, a round wall for non-adaptive retrieval, and a conditional composition wall for fixed-depth inference. We introduce ASP, a training-free wrapper for frozen multimodal models that combines a capped structured state, a verbatim episodic index, and query-conditioned budget allocation with iterative access. Following a pre-registered protocol, we evaluate seven open-weight models from 3B to 31B on SEW-Bench, a license-free synthetic long-horizon walkthrough benchmark constructed to instantiate these walls. The registered natural-video benchmarks were not run because their frames require dataset agreements; our evidence therefore concerns access mechanisms, not natural-scene perception. Under a 4,096-token decision budget, ASP reaches 75 to 94% episodic retrieval accuracy, compared with 3 to 19% for equal-budget query-independent sampling, and budget reallocation outperforms quadrupling the sampling budget on every backbone. However, the full three-component architecture does not validate channel duality: removing the compressive state raises the flagship mean from 35.4 to 58.0, ASP does not outperform the verbatim-only baseline on any backbone, and two of four pre-registered falsification criteria fire. These results show that query-conditioned access, rather than parameter count or context growth alone, is decisive under a fixed budget, while prompted online compression does not earn its cost in this setting.
Chinese Translation
具身多模态代理必须在固定的每决策令牌预算下,从不断增长的观察流中作出回应。我们通过四个资源壁垒来形式化这一约束:感知香农壁垒用于有界状态,视野壁垒用于查询无关的帧选择,轮次壁垒用于非自适应检索,以及条件组合壁垒用于固定深度推理。我们引入了ASP(Access-Structured Perception),这是一个针对冻结多模态模型的无训练包装器,结合了受限的结构状态、逐字的情节索引和基于查询的预算分配与迭代访问。根据预注册协议,我们在SEW-Bench上评估了七个开放权重模型,参数范围从3B到31B,SEW-Bench是一个无许可的合成长视野走查基准,旨在实现这些壁垒。由于自然视频基准的帧需要数据集协议,因此未进行测试;因此,我们的证据涉及访问机制,而非自然场景感知。在4096令牌的决策预算下,ASP的情节检索准确率达到75%至94%,而相同预算的查询无关采样仅为3%至19%,并且预算重新分配在每个基础模型上均优于四倍采样预算。然而,完整的三组件架构并未验证通道对偶性:去除压缩状态将旗舰均值从35.4提高到58.0,ASP在任何基础模型上均未超越逐字基准,并且四个预注册的虚假标准中有两个触发。这些结果表明,在固定预算下,查询条件访问比参数数量或上下文增长更为决定性,而在此设置中,在线提示压缩并未获得其成本的回报。
cs.AI / 130 / 2608.22979

SA-RSQ: A Versatile Sparse Representation Framework for Multi-modal Recommender Systems

SA-RSQ:一种适用于多模态推荐系统的多功能稀疏表示框架
Wang, Xiang, Quan, Shigang, Chang, Tingzhen, Yang, Kang, Chen, Sitong, Fan, Yabo, Wang, Xingxing, He, Zhaodian
Abstract
Deploying high-dimensional multimodal features in industrial recommender systems incurs substantial storage and latency overhead. Hard quantization is compact but introduces boundary distortion, whereas dense soft quantization couples representation quality to the limited storage budget. We propose Sparse Activation-based Residual Soft Quantization (SA-RSQ), which uses Top-K sparse routing and softmax weights to store compact (Index, Probability) tuples. The stored tuples decouple per-item storage from codebook dimensionality; for a fixed selected support, gradients propagate through the routing weights and weighted reconstruction without relying on a straight-through estimator. Experiments on a proprietary food-delivery advertising dataset show favorable reconstruction-performance and CTR trade-offs across storage budgets of 8-48 bytes per item. A preliminary Next-Distribution Prediction study and a one-week online A/B test further demonstrate the practical potential of SA-RSQ, with relative lifts of +2.51% in CTR and +3.66% in CPM.
Chinese Translation
在工业推荐系统中部署高维多模态特征会带来显著的存储和延迟开销。硬量化虽然紧凑,但会引入边界失真;而密集软量化则将表示质量与有限的存储预算相结合。我们提出了基于稀疏激活的残差软量化(Sparse Activation-based Residual Soft Quantization,SA-RSQ),该方法利用 Top-K 稀疏路由和 softmax 权重来存储紧凑的(索引,概率)元组。存储的元组将每个项目的存储与码本维度解耦;在固定选择的支持下,梯度通过路由权重和加权重构传播,而无需依赖直通估计器。在一个专有的食品配送广告数据集上的实验显示,在每个项目的存储预算为 8-48 字节的范围内,重构性能和点击率(CTR)之间的权衡表现良好。初步的下一个分布预测研究和为期一周的在线 A/B 测试进一步展示了 SA-RSQ 的实际潜力,CTR 提升相对为 +2.51%,CPM 提升相对为 +3.66%。
cs.AI / 131 / 2608.23001

PatchWrite: One Line, Not One Section -- Compile-Gated, Validity-Preserving Editing for AI-Drafted Manuscripts

PatchWrite:一行,而非一节——针对AI撰写手稿的编译门控、有效性保持编辑
Yang, Weiwei
Abstract
Automated manuscript pipelines often regenerate an entire section to repair a local defect, allowing unrelated metrics and citations to change even when the resulting PDF still builds. PatchWrite instead constrains how candidate edits become committed manuscript states: it reuses bounded EDIT N M editing and rollback, but tightens compilation acceptance with fatal-log checks and adds evidence locks that require every cited key and experimental numeric token to be attested by a reference registry or experimental log. Candidates that fail either check are rejected and the previous HEAD is retained. On a 24-manuscript x 8-fault oracle stress test (768 jobs, evenly split between compile-breaking and content-only faults), whole-slot rewriting mutated an unrelated "12-layer" line in every case (0/192 preserved; numeric Jaccard 0.6667), whereas PatchWrite preserved it in 192/192 cases. Removing the compile gate reduced acceptance to 0, while removing the evidence gate allowed a hallucinated citation to pass. The same pattern held across all eight faults. To test the protocol with generation rather than oracle edits, we reran the 192 jobs with the writer model proposing the edits. The model's candidates were accepted in 75% of cases; nearly all rejections came from one reproducible failure mode in which the model attempted to delete a line using an empty replacement unsupported by the current grammar. Every accepted candidate passed both gates, and 93.75% fixed the injected fault; the remaining cases involved a technically valid but sentence-inappropriate citation and one markup-changing near-miss. In a blind evaluation of sixteen PDF pairs, both raters preferred PatchWrite for preserving lab-grounded facts (C1 Likert 5.0 vs. 2.0), while rating prose quality nearly identically. Logs from 193 in-product drafting tasks show the same classes of failures occurring in practice.
Chinese Translation
自动化手稿流程通常会重新生成整个章节以修复局部缺陷,即使生成的PDF仍然可以构建,相关的指标和引用也可能发生变化。而PatchWrite则限制了候选编辑如何成为提交的手稿状态:它重用了有界的EDIT N M编辑和回滚,但通过致命日志检查收紧了编译接受标准,并增加了证据锁,要求每个引用的关键和实验数字标记都必须由参考注册或实验日志进行证明。未通过任一检查的候选项将被拒绝,保留先前的HEAD。在一个24篇手稿×8种故障的oracle压力测试中(768个任务,均匀分配在编译破坏和仅内容故障之间),整个插槽重写在每种情况下都改变了一个无关的“12层”行(0/192保留;数字Jaccard 0.6667),而PatchWrite在192/192的情况下保留了它。移除编译门导致接受率降至0,而移除证据门则允许一个虚构的引用通过。所有八种故障中都保持了相同的模式。为了测试该协议在生成而非oracle编辑中的表现,我们重新运行了192个任务,作家模型提出了编辑建议。模型的候选项在75%的情况下被接受;几乎所有的拒绝都来自一个可重复的失败模式,即模型试图使用当前语法不支持的空替换删除一行。每个被接受的候选项都通过了两个门,93.75%修复了注入的故障;剩余的案例涉及一个技术上有效但句子不当的引用和一个标记变化的接近失误。在对十六对PDF进行盲评估时,两位评审者均更倾向于PatchWrite以保持实验室基础事实(C1 Likert 5.0对2.0),而在散文质量评分上几乎相同。来自193个产品内草拟任务的日志显示,实践中发生了相同类别的失败。
cs.AI / 132 / 2608.23028

PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies

PsychJail:通过多轮劝说探索心理监狱破解
Feng, Zeyu, Wu, Qingyu, Luo, Yuzhe, Cheng, Hua
Abstract
Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage them as sustained social interlocutors rather than one-shot query engines. This shift makes jailbreaks a growing safety threat, yet most research emphasizes single-turn prompt optimization or iterative attack refinement, leaving psychologically grounded multi-turn vulnerabilities underexplored. We present PsychJail, a psychology-guided framework for red teaming aligned LLMs through theory-grounded, multi-turn persuasion. PsychJail maps established social-psychological persuasion techniques into a tactic-conditioned attack policy. It factorizes each attacker action into a Change-of-Meaning analysis, tactic selection, and victim-visible message, operationalizing the Persuasion Knowledge Model (PKM). The policy is refined with trajectory-level reinforcement learning using a PKM-gated reward that credits early jailbreak success only when every turn contains a well-formed Change-of-Meaning analysis. Across four aligned victim models, PsychJail achieves the highest average attack success rate (87.3%) and outperforms strong single-turn and multi-turn baselines on every model. We also measure susceptibility at the action that breaks each victim, revealing four distinct model-level fingerprints that identify which persuasion levers affect each model and how broadly. These fingerprints help explain cross-model transfer asymmetry. We interpret them as four candidate psychological profiles-rationalist, credibility-driven, narrative-monoculture, and broadly persuadable-while treating this interpretation as a conjecture requiring future validation. Our findings establish psychological jailbreaks as a distinct red-teaming frontier for increasingly interactive LLMs.
Chinese Translation
大型语言模型(LLMs)越来越多地应用于教育、医疗、政策咨询和其他互动环境中,用户将其视为持续的社交对话者,而非一次性查询引擎。这一转变使得监狱破解成为日益严重的安全威胁,然而大多数研究强调单轮提示优化或迭代攻击精炼,心理学基础的多轮脆弱性仍未得到充分探讨。我们提出了PsychJail,这是一个通过理论指导的多轮劝说来对齐LLMs进行红队测试的框架。PsychJail将既定的社会心理劝说技术映射为战术条件攻击策略。它将每个攻击者的行动分解为意义变化分析、战术选择和受害者可见信息,从而实现劝说知识模型(PKM)的操作化。该策略通过轨迹级强化学习进行优化,使用PKM门控奖励,仅在每轮都包含良好构造的意义变化分析时,才对早期监狱破解成功给予奖励。在四个对齐的受害者模型中,PsychJail实现了最高的平均攻击成功率(87.3%),并在每个模型上优于强大的单轮和多轮基线。我们还测量了每个受害者破裂时的脆弱性,揭示了四个不同的模型级指纹,识别出哪些劝说杠杆影响每个模型及其广泛程度。这些指纹有助于解释跨模型转移的不对称性。我们将其解读为四种候选心理特征——理性主义者、以可信度驱动者、叙事单一文化和广泛可劝说者,同时将这一解读视为需要未来验证的假设。我们的研究结果确立了心理监狱破解作为日益互动的LLMs的一个独特红队测试前沿。
cs.AI / 133 / 2608.23030

Artificial Empathy: Towards a Framework for Unsupervised Agency Detection and Policy Reconstruction

人工同理心:无监督代理检测与策略重构框架的探索
Kuhn, Peter, Pang, Chris, Chauhan, Sonakshi
Abstract
We study how an AI system can identify and model other agents in its environment from observation alone, which is a capability necessary for cooperative behaviour in the real world. This problem is less constrained than inverse reinforcement learning and remains largely unexplored. We propose a framework that uses a reinforcement learning agent, trained on an independent task as a prior about agentic dynamics, to perform agency detection and policy reconstruction.
Chinese Translation
我们研究了人工智能系统如何仅通过观察识别和建模其环境中的其他代理,这一能力对于现实世界中的合作行为是必需的。这个问题的约束性低于逆强化学习,并且仍然在很大程度上未被探索。我们提出了一个框架,该框架利用在独立任务上训练的强化学习代理作为关于代理动态的先验知识,以执行代理检测和策略重构。
cs.AI / 134 / 2608.23035

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

MobilePA-Bench:在复杂现实任务上评估移动规划代理的基准测试
Zhu, Yi, Wu, Xiongwei, Wang, Qiyi, Qu, Tingyu, Liu, Jiajun, Cao, Sihan, Chen, Long, Sun, Weigao, Zhu, Feida, Zhong, Yiran, Hoi, Steven
Abstract
As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks fall into two camps, each with a critical blind spot: GUI-centric benchmarks test surface-level screen manipulation while overlooking background tool use and long-horizon planning, whereas static function-calling benchmarks rely on offline API matching that is detached from real runtime constraints. To close this gap, we present \textbf{MobilePA-Bench}, an interactive, stateful, and tool-centric benchmark for evaluating the tool-calling and planning abilities of mobile planning agents. MobilePA-Bench runs on an executable sandbox that maintains live application databases and returns structured feedback, spanning $13$ functional domains and $212$ realistic mobile tools. Beyond basic tool use, it evaluates a central planning agent along three advanced dimensions: \emph{(1)~Sub-agent Collaboration}---decomposing a complex task and delegating specialized work to capable sub-agents; \emph{(2)~Memory Usage}---recalling stored memories, user profiles, and past preferences to resolve implicit requests; and \emph{(3)~Skill Usage}---invoking pre-packaged composite skills instead of planning every step from scratch. Extensive experiments show that current frontier LLMs remain unreliable in mobile settings: performance drops sharply under strict tool ordering, permission limits, and unexpected runtime errors. By pairing an interactive function-calling sandbox with evidence-based verification, MobilePA-Bench serves as both a practical diagnostic benchmark and an interactive foundation for agentic reinforcement learning---accelerating the development of dependable mobile agents.
Chinese Translation
随着设备上的大型语言模型(LLM)代理逐渐演变为个人副驾驶,移动操作系统已成为这一范式的重要测试平台,因此对其能力进行严格评估变得至关重要。然而,现有基准测试主要分为两类,各自存在关键盲点:以图形用户界面(GUI)为中心的基准测试仅测试表层的屏幕操作,而忽视了后台工具的使用和长远规划;而静态函数调用基准则依赖于与实际运行约束脱节的离线API匹配。为了解决这一问题,我们提出了 extbf{MobilePA-Bench},这是一个互动的、有状态的、以工具为中心的基准测试,用于评估移动规划代理的工具调用和规划能力。MobilePA-Bench在一个可执行的沙盒环境中运行,维护实时应用数据库并返回结构化反馈,涵盖$13$个功能领域和$212$个现实移动工具。除了基本的工具使用外,它还从三个高级维度评估中心规划代理的能力: extit{(1)~子代理协作}——将复杂任务分解并将专业工作委派给能够的子代理; extit{(2)~记忆使用}——回忆存储的记忆、用户档案和过去的偏好,以解决隐含请求;以及 extit{(3)~技能使用}——调用预打包的复合技能,而不是从头规划每一步。大量实验表明,当前前沿的LLM在移动环境中仍然不可靠:在严格的工具顺序、权限限制和意外运行错误下,性能急剧下降。通过将互动函数调用沙盒与基于证据的验证相结合,MobilePA-Bench既作为一个实用的诊断基准,也为代理强化学习提供了互动基础,从而加速可靠移动代理的开发。
cs.AI / 135 / 2608.23041

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

AutoSaddler:基于代理执行轨迹的自动化马具优化与持久更新
Park, Sungho, Kim, Wonjoong, Tan, Rongyuan, Zhang, Jue, Han, Wook-Shin, Gao, Pengfei, Park, Chanyoung, Yao, Yongqiang, Fu, Rao, Nallipogu, Elsie, Lin, Qingwei, Rajmohan, Saravan, Zhang, Dongmei
Abstract
LLM agents remain unreliable on long-horizon tasks, where small local failures can compound over extended interactions and lead to overall task failure. Although external harnesses can substantially improve robustness, harness design remains a manual and expensive process that requires searching over a large space of prompts, tool configurations, and control logic. We propose AutoSaddler, an automatic harness optimization framework that formulates harness improvement as an offline learning problem and iteratively updates the harness using failure signals from mini-batches. AutoSaddler combines failure-trace diagnosis, structured patch generation that treats the harness as code, and validation-based update selection. Experiments on GAIA2, SWE-Bench Pro, and Terminal-Bench 2.0 show that AutoSaddler substantially improves agent performance over the corresponding base harnesses, achieving gains of 9.0, 9.6, and 10.0 percentage points, respectively. Ablation studies further suggest that effective harness optimization benefits from three ingredients: deep debugging rather than shallow reflection, targeted modifications rather than unconstrained editing, and generalization-aware selection rather than trajectory-specific repair. Together, these results suggest that automatic harness optimization is a promising path toward more performant and reliable agent systems.
Chinese Translation
在长时间跨度的任务中,大型语言模型(LLM)代理仍然不可靠,因为小的局部失败可能在长时间交互中累积并导致整体任务失败。尽管外部马具可以显著提高鲁棒性,但马具设计仍然是一个手动且昂贵的过程,需要在大量的提示、工具配置和控制逻辑中进行搜索。我们提出了AutoSaddler,一个自动化马具优化框架,将马具改进公式化为离线学习问题,并利用来自小批量的失败信号迭代更新马具。AutoSaddler结合了失败轨迹诊断、将马具视为代码的结构化补丁生成和基于验证的更新选择。在GAIA2、SWE-Bench Pro和Terminal-Bench 2.0上的实验表明,AutoSaddler显著提高了代理在相应基础马具上的性能,分别实现了9.0、9.6和10.0个百分点的提升。消融研究进一步表明,有效的马具优化受益于三个要素:深度调试而非浅层反思、针对性修改而非无约束编辑,以及关注泛化的选择而非特定轨迹的修复。综合来看,这些结果表明,自动化马具优化是实现更高性能和更可靠代理系统的有希望的路径。
cs.AI / 136 / 2608.23045

From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation

从惯性到客观性:通过噪声隔离提升深度研究代理
Zhang, Xiangxin, Zhang, Zhanwei, Fu, Zhihang, Lin, Binbin, Wang, Wenxiao
Abstract
Web search agents powered by Large Language Models (LLMs) show strong promise, but deep research tasks expose a recurring failure mode: once an agent has produced a query, plan, or intermediate conclusion, it becomes less objective when later judging the consequences of that same action. We term this phenomenon \textbf{inertia bias}. To make it measurable, we introduce the IBIS benchmark, which controls the search observations while varying whether the model is evaluating the outcome of its own prior action. We find that models are substantially worse when they ``own'' the preceding search step, showing that self-authored action history can systematically distort subsequent judgment. We further show that this bias propagates into two forms of system-level degradation: search noise at the worker level and contextual noise at the manager level. To address this problem, we propose NIS-Agent, which applies context isolation at the two decision points most vulnerable to inertia bias: webpage triage and final-answer validation. Across GAIA, WebWalkerQA, BrowseComp, and BrowseComp-zh, NIS-Agent achieves competitive performance while reducing token cost by 33\% compared to our baseline. We further train an 8B model to be intrinsically more resistant to inertia bias; under the same NIS-Agent framework, it attains average performance comparable to GPT-4o on deep research benchmarks.
Chinese Translation
由大型语言模型(LLMs)驱动的网络搜索代理展现出强大的潜力,但深度研究任务暴露出一种反复出现的失败模式:一旦代理生成了查询、计划或中间结论,在后续判断同一行动的结果时,它变得不那么客观。我们将这一现象称为 extbf{惯性偏差}。为了使其可测量,我们引入了IBIS基准,该基准在控制搜索观察的同时,变化模型是否在评估其自身先前行动的结果。我们发现,当模型“拥有”先前的搜索步骤时,其表现显著下降,显示自我创作的行动历史可以系统性地扭曲后续判断。我们进一步表明,这种偏差传播到两种系统级别的退化形式:工作层面的搜索噪声和管理层面的上下文噪声。为了解决这个问题,我们提出了NIS-Agent,该代理在两个最容易受到惯性偏差影响的决策点应用上下文隔离:网页分类和最终答案验证。在GAIA、WebWalkerQA、BrowseComp和BrowseComp-zh中,NIS-Agent在减少33%的令牌成本的同时,取得了与我们的基线相当的竞争性能。我们进一步训练了一个8B模型,使其在惯性偏差方面具有更强的内在抵抗力;在相同的NIS-Agent框架下,它在深度研究基准测试中达到了与GPT-4o相当的平均性能。
cs.AI / 137 / 2608.23058

LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications

基于大型语言模型的预测与预报代理:方法、训练、评估与应用
Xu, Xiaogang, Tang, Jiaqi, Chen, Jianmin, Yan, Yingying, Tang, Zhenchao, Zhou, Xiangxin, Hu, Xiaobin, Wei, Wei, Wu, Jinfeng, Chen, Qifeng, Zhou, Lu, Wu, Jiafei, Liu, Zhe, Yin, Jianwei, Zheng, Weimin
Abstract
Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external tools, and iterative prediction. We investigate LLM-based forecasting agents, meaning systems in which a language model contributes to a scored prediction about a future or currently unobserved target. We organize architectures into three groups. Standalone LLM workflows operate on encoded time series or event context. Tool- and retrieval-augmented agents incorporate external evidence. Hybrid systems pair LLMs with statistical or foundation models. We then review training methods and evaluation protocols. We examine negative as well as positive evidence, including sensitivity to small input perturbations, ablations in which the LLM component does not improve accuracy, and benchmark gains that may reflect contamination instead of temporal reasoning. We cover applications in finance, weather, health, energy, and operations, and we summarize the benchmarks and datasets used for evaluation. The evidence indicates that measurement is a central limitation. Future work requires calibration under distribution shift, contamination-resistant live evaluation, explicit reporting of cost and accuracy together, and methods for handling feedback between deployed forecasts and the outcomes being forecast.
Chinese Translation
大型语言模型(LLMs)现在支持将基于语言的推理与时间数据、证据检索、外部工具和迭代预测相结合的预测系统。我们研究基于LLM的预测代理,指的是在这些系统中,语言模型为未来或当前未观察到的目标提供评分预测。我们将架构分为三类:独立的LLM工作流程处理编码的时间序列或事件上下文;工具和检索增强的代理结合了外部证据;混合系统将LLM与统计模型或基础模型配对。接着,我们回顾训练方法和评估协议。我们考察了负面和正面证据,包括对小输入扰动的敏感性、LLM组件未能提高准确性的消融实验,以及可能反映污染而非时间推理的基准增益。我们涵盖了金融、天气、健康、能源和运营等领域的应用,并总结了用于评估的基准和数据集。证据表明,测量是一个核心限制。未来的工作需要在分布变化下进行校准、抗污染的实时评估、明确报告成本和准确性,并处理已部署预测与被预测结果之间的反馈的方法。
cs.AI / 138 / 2608.23061

Improving O-RADS Risk Stratification from Ultrasound Reports: A Comparative Evaluation of Hybrid versus End-to-End LLM Reasoning Strategies

改善超声报告中的 O-RADS 风险分层:混合与端到端 LLM 推理策略的比较评估
Tan, Xiaotong, Qiu, Chunli, Liu, Xin, Huang, Qing, Zhou, Guangli, Gao, Bo, Song, Xiaoyan, Wang, Shuyan, Wang, Xiuqin, Xue, Wufeng, Huang, Ruobing, Ni, Dong, Tao, Guowei, Cheng, Jun
Abstract
Background: Automating clinical guideline-based decision-making with large language models (LLMs) remains challenging because of reliability, hallucination, and limited interpretability. We compared the performance of LLMs and reasoning strategies for automated Ovarian-Adnexal Reporting and Data System (O-RADS) classification from free-text pelvic ultrasound reports. Methods: In this retrospective study, consecutive patients with ovarian masses who underwent pelvic ultrasound were included. Eight LLMs were tested with three reasoning strategies: implicit-knowledge end-to-end, rule-informed end-to-end, and a feature-based hybrid architecture that decoupled feature extraction from rule-based classification. The reference standard was O-RADS categorization established by expert consensus. Results: A total of 310 women with 390 ovarian masses were evaluated. The feature-based hybrid architecture using Gemini 3.6 Flash demonstrated the best performance, achieving an accuracy of 99.2% (387 of 390) and almost perfect agreement with the reference standard (weighted kappa = 1.00; 95% CI: 0.99-1.00). Its performance surpassed that of original clinical reports (accuracy, 87.7% [342 of 390]; weighted kappa = 0.94; 95% CI: 0.91-0.96) and end-to-end LLM strategies (accuracy range, 65.6% [256 of 390] to 95.9% [374 of 390]). For structured feature extraction, Gemini 3.6 Flash demonstrated higher overall accuracy than Claude Fable 5 (98.9% vs 97.8%; P < 0.001). The hybrid architecture reduced misclassification errors and mitigated the overstaging tendency observed in original reports. Conclusion: The feature-based hybrid LLM architecture that separates clinical feature extraction from deterministic guideline execution enables highly accurate, reliable, and interpretable automated O-RADS classification, providing a promising approach for standardized, guideline-based clinical decision-making.
Chinese Translation
背景:利用大型语言模型(LLMs)自动化基于临床指南的决策仍然面临挑战,主要由于可靠性、幻觉和有限的可解释性。我们比较了 LLMs 和推理策略在从自由文本盆腔超声报告中自动化卵巢-附属器报告和数据系统(O-RADS)分类的表现。方法:在这项回顾性研究中,纳入了接受盆腔超声检查的连续卵巢肿块患者。测试了八种 LLM,采用三种推理策略:隐性知识端到端、规则知情端到端,以及一种将特征提取与基于规则的分类解耦的特征基础混合架构。参考标准为专家共识建立的 O-RADS 分类。结果:共评估了 310 名女性及其 390 个卵巢肿块。使用 Gemini 3.6 Flash 的特征基础混合架构表现最佳,准确率达到 99.2%(390 例中的 387 例),与参考标准几乎完全一致(加权 kappa = 1.00;95% CI:0.99-1.00)。其表现超过了原始临床报告(准确率 87.7% [390 例中的 342 例];加权 kappa = 0.94;95% CI:0.91-0.96)和端到端 LLM 策略(准确率范围 65.6% [390 例中的 256 例] 至 95.9% [390 例中的 374 例])。在结构化特征提取方面,Gemini 3.6 Flash 的整体准确率高于 Claude Fable 5(98.9% 对 97.8%;P < 0.001)。混合架构减少了误分类错误,并缓解了原始报告中观察到的过度分期倾向。结论:将临床特征提取与确定性指南执行分离的特征基础混合 LLM 架构实现了高度准确、可靠和可解释的自动化 O-RADS 分类,为标准化、基于指南的临床决策提供了一种有前景的方法。
cs.AI / 139 / 2608.23070

From Generation to Simulation: How Far Are World Models from Being True Simulators?

从生成到模拟:世界模型距离真正的模拟器还有多远?
Wang, Tong, Deng, Huan, Yang, Mucheng, He, Yang, Kuang, Xiaohui, Zhao, Gang
Abstract
With the rapid progress of diffusion models and large-scale video generation, generative world models are increasingly expected to replace traditional simulators, including physics engines, game engines, and reinforcement-learning environments. Yet the remaining distance from generation to simulation lacks a systematic assessment. We present a capability-based study using an external yardstick: eight capabilities of a traditional simulator, namely asset construction, physics engine, interaction, controllability, stability, state feedback, diversity, and evaluation metrics. We trace three main technical routes--latent dynamics, video generation, and joint-embedding prediction--and map exactly 200 representative works published from 2018 to June 2026 onto these capabilities. Our analysis shows that world models have achieved functional substitution in interaction and controllability for specific scenarios, but remain short of traditional simulators in formal guarantees of physical laws, structured state feedback, and reproducible long-horizon evolution. State feedback is the most neglected cross-route shortcoming: only 6 of 163 implementation papers expose a runtime interface for querying entity states or physical parameters. We identify six research directions: formalized physics, a unified action interface, first-class state feedback, long-horizon stability, downstream-utility evaluation, and cross-route hybridization. Project page: https://github.com/AtongWang/world-model-simulators
Chinese Translation
随着扩散模型和大规模视频生成的快速进展,生成式世界模型越来越被期望取代传统模拟器,包括物理引擎、游戏引擎和强化学习环境。然而,从生成到模拟的距离缺乏系统性的评估。我们提出了一项基于能力的研究,使用外部标准:传统模拟器的八项能力,即资产构建、物理引擎、交互性、可控性、稳定性、状态反馈、多样性和评估指标。我们追踪了三条主要技术路线——潜在动态、视频生成和联合嵌入预测——并将2018年至2026年6月间发表的200篇代表性作品精确映射到这些能力上。我们的分析表明,世界模型在特定场景下已实现了交互性和可控性的功能替代,但在物理法则的正式保证、结构化状态反馈和可重复的长时间演化方面仍然不及传统模拟器。状态反馈是最被忽视的跨路线缺陷:在163篇实现论文中,仅有6篇暴露了查询实体状态或物理参数的运行时接口。我们确定了六个研究方向:形式化物理、统一的动作接口、一流的状态反馈、长时间稳定性、下游效用评估和跨路线混合化。项目页面:https://github.com/AtongWang/world-model-simulators
cs.AI / 140 / 2608.23078

AgentWeave: Routing Before Reasoning for Efficient Function Calling in Tool-Rich Language Models

AgentWeave:在推理之前进行路由以提高工具丰富语言模型中的函数调用效率
Singla, Saurav, Singla, Aarav, Gupta, Advik, Gupta, Parnika
Abstract
Large language models increasingly operate over large collections of tools, functions, APIs, and specialized agents. As the candidate action space grows, a function-calling model must process more schemas, consume more prompt tokens, and distinguish among increasingly similar or irrelevant alternatives. We study a complementary systems strategy: reduce the candidate set before language-model inference while leaving the downstream model unchanged. We introduce AgentWeave, a deterministic pre-inference routing layer that constructs a bounded model-visible action space using eligibility, requirement, capability, and routing signals. We evaluate AgentWeave with a frozen BFCL-derived routing-pressure protocol using the public MadeAgents/Hammer2.1-1.5b model. On 48 fresh BFCL V4 multiple-function tasks, AgentWeave achieves 6/48 (12.5%) native BFCL successes, whereas all-tools, deterministic random top-8, and semantic top-8 baselines each achieve 0/48. The paired success difference is +12.5 percentage points with a 10,000-resample paired bootstrap 95% confidence interval of +4.17 to +22.92 points and exact McNemar p=0.03125. Relative to all-tools exposure, AgentWeave presents 70.18% fewer tools, uses 61.70% fewer input tokens, and exhibits 50.95% lower mean local-model latency. The result is deliberately narrow: this is a BFCL-derived routing-pressure study rather than an official full BFCL leaderboard score, and absolute task success remains low. The evidence nevertheless shows that candidate-space construction can materially affect a fixed model's function-calling behavior and motivates evaluating routing as a distinct stage before model reasoning.
Chinese Translation
大型语言模型越来越多地在大量工具、函数、API和专用代理的集合上运行。随着候选动作空间的扩大,函数调用模型必须处理更多的模式,消耗更多的提示令牌,并在越来越相似或不相关的替代方案中进行区分。我们研究了一种互补的系统策略:在语言模型推理之前减少候选集,同时保持下游模型不变。我们引入了AgentWeave,这是一个确定性的推理前路由层,利用资格、要求、能力和路由信号构建一个有限的模型可见动作空间。我们使用公共的MadeAgents/Hammer2.1-1.5b模型,通过冻结的BFCL派生路由压力协议对AgentWeave进行了评估。在48个新的BFCL V4多功能任务中,AgentWeave实现了6/48(12.5%)的原生BFCL成功,而所有工具、确定性随机前8和语义前8的基线均实现了0/48。配对成功率的差异为+12.5个百分点,经过10,000次重采样的配对自助法95%置信区间为+4.17到+22.92点,精确的McNemar p=0.03125。相较于所有工具的曝光,AgentWeave展示了70.18%的工具减少,使用了61.70%更少的输入令牌,并表现出50.95%更低的平均局部模型延迟。结果是故意狭窄的:这是一项基于BFCL派生路由压力的研究,而不是官方的完整BFCL排行榜分数,绝对任务成功率仍然较低。然而,证据表明,候选空间的构建可以实质性地影响固定模型的函数调用行为,并激励我们在模型推理之前将路由作为一个独立阶段进行评估。
cs.AI / 141 / 2608.23086

POOL: Propagated Uncertainty Over Lookalikes

POOL: 通过相似项传播的不确定性
Sharma, Rounak, Sai, Ananya B., Pal, Soumyabrata
Abstract
Black-box large language models need confidence scores that can separate likely-correct from likely-incorrect outputs, enabling systems to prioritize human review, route uncertain cases to stronger models, or choose abstention thresholds on development data. Yet existing confidence estimators face a cost-quality trade-off: verbal confidence is cheap but is often overconfident, while sampling-based uncertainty is more informative but scales linearly with the number of samples per query. We propose \textsc{POOL} (\emph{Propagated Uncertainty Over Lookalikes}),a cost-efficient framework that addresses this trade-off taking inspiration from group-testing.\textsc{POOL} clusters query stems with overlaps, evaluates a base estimator on representative medoids, softly propagates confidence scores to nearby queries, and selectively evaluates high-disagreement cases. We instantiate this framework with \textsc{Hy@}$p$, a hybrid estimator that combines verbal confidence with spectral answer diversity computed from the negative von Neumann entropy of sampled answer embeddings.Across six domains from three datasets and five black-box LLMs, \textsc{Hy@}5 achieves higher average AUROC than verbal confidence and \textsc{Vn@}10 sampling while using half as many samples as \textsc{Vn@}10. \textsc{POOL}-\textsc{Hy@}5 retains 93.5--97.9\% of its AUROC while saving 19.3--39.3\% of generations. On paraphrase-dense workloads, generation savings rise to 73-76\%, showing that semantic redundancy can be leveraged to lower confidence-estimation costs.
Chinese Translation
黑箱大型语言模型需要能够区分可能正确与可能错误输出的置信度评分,以便系统能够优先进行人工审核,将不确定案例路由到更强的模型,或在开发数据上选择弃权阈值。然而,现有的置信度估计器面临成本与质量的权衡:语言置信度成本低但往往过于自信,而基于采样的不确定性更具信息性,但与每个查询的样本数量呈线性增长。我们提出了 extsc{POOL}( extit{通过相似项传播的不确定性}),这是一个高效的框架,旨在解决这一权衡,灵感来自于组测试。 extsc{POOL} 将查询根干进行重叠聚类,在代表性中值上评估基础估计器,柔性传播置信度评分到附近查询,并选择性地评估高分歧案例。我们用 extsc{Hy@}$p$ 实现了这一框架,这是一种混合估计器,将语言置信度与从采样答案嵌入的负冯·诺依曼熵计算出的谱答案多样性结合起来。在来自三个数据集和五个黑箱 LLM 的六个领域中, extsc{Hy@}5 的平均 AUROC 高于语言置信度和 extsc{Vn@}10 采样,同时使用的样本数量仅为 extsc{Vn@}10 的一半。 extsc{POOL}- extsc{Hy@}5 保留了 93.5% 到 97.9% 的 AUROC,同时节省了 19.3% 到 39.3% 的生成。在重述密集的工作负载中,生成节省率上升至 73% 到 76%,表明可以利用语义冗余来降低置信度估计成本。
cs.AI / 142 / 2608.23098

Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement

九歌-推桥:一种可解释的人机系统用于古典汉诗的精炼
Han, Yufeng, Deng, Lifan, Kong, Cunliang, Li, Wenhao, Cong, Xin, Bai, Yuzhuo, Luo, Kangyang, Sun, Maosong
Abstract
Classical Chinese poetry composition has long valued Tuiqiao, the iterative refinement of words, imagery, and prosody. However, many current AI poetry systems follow a one-shot generation paradigm, which reduces users to prompt providers and weakens their creative agency. We present Jiuge-Tuiqiao, an interactive human-AI collaborative system for classical Chinese poetry composition. The system is designed around a triadic model: user-driven control, ancient-guided evidence, and AI-assisted generation. Users can lock characters or lines, receive real-time prosody feedback, and obtain interpretable refinement suggestions grounded in high-frequency collocations, PPL-ranked classical lines, and structured knowledge extracted from classical encyclopedias. This design turns AI from an autonomous generator into a background assistant that supports the user's own process of poetic refinement. Preliminary experiments and user feedback suggest that Jiuge-Tuiqiao improves controllability, interpretability, and user engagement in classical poetry composition.
Chinese Translation
古典汉诗的创作长期以来重视推桥,即对词语、意象和韵律的迭代精炼。然而,许多当前的AI诗歌系统遵循一次性生成的范式,这使得用户仅仅成为提示提供者,削弱了他们的创作主动性。我们提出了九歌-推桥,一个用于古典汉诗创作的互动人机协作系统。该系统围绕三元模型设计:用户驱动的控制、古代指导的证据和AI辅助的生成。用户可以锁定字符或诗句,实时获得韵律反馈,并获取基于高频搭配、PPL排名的古典诗句和从古典百科全书提取的结构化知识的可解释的精炼建议。该设计将AI从一个自主生成器转变为一个支持用户自身诗歌精炼过程的后台助手。初步实验和用户反馈表明,九歌-推桥在古典诗歌创作中提高了可控性、可解释性和用户参与度。
cs.AI / 143 / 2608.23196

AI emotional support is better only when chosen, but shifts preferences even when it is not

仅在选择时AI情感支持更佳,但即使在未选择时也会改变偏好
Shi, Yaoxi, Fang, Cathy Mengying, Maes, Guy LabanPattie, Goldenberg, Amit
Abstract
People increasingly face a novel decision when seeking emotional support: human or AI. In existing studies, AI's empathic messages are rated as well as or better than humans'. But these studies either assigned the support source or honored people's choice. In real life, support is often incongruent with choice, as people want one source and receive the other. Across three experiments (N = 1,951), participants chose whether to share an emotional experience with a human or an AI, then were randomly assigned to a congruent or incongruent partner. AI support was rated as superior only among those who had chosen it. Yet regardless of congruence, interacting with AI increased willingness to choose it again. In a 28-day study with OpenAI (N = 981), daily conversations shifted preferences toward AI and away from humans, but only when conversations turned personal. Emotional support choices are thus path-dependent, progressively redirecting away from human connection.
Chinese Translation
人们在寻求情感支持时越来越面临一个新颖的决策:选择人类还是人工智能(AI)。在现有研究中,AI的共情信息被评估为与人类相当或更好。但这些研究要么指定了支持来源,要么尊重了人们的选择。在现实生活中,支持往往与选择不一致,因为人们想要一种来源却得到另一种。在三项实验中(N = 1,951),参与者选择是否与人类或AI分享情感体验,然后随机分配到一致或不一致的伙伴。只有在选择了AI的参与者中,AI支持被评估为优越。然而,无论是否一致,与AI的互动都增加了再次选择它的意愿。在与OpenAI进行的为期28天的研究中(N = 981),每日对话使偏好向AI倾斜而远离人类,但仅在对话变得个人化时。情感支持的选择因此是路径依赖的,逐渐引导人们远离人际连接。
cs.AI / 144 / 2608.23205

Cognitive Profiling of LRMs' Reasoning Traces Using Bloom's Taxonomy

使用布鲁姆分类法对大型推理模型的推理痕迹进行认知分析
Zoumpoulidi, Maria-Eleni, Paraskevopoulos, Georgios, Potamianos, Alexandros
Abstract
Large Reasoning Models (LRMs) have revolutionized reasoning in LLMs, and the increasing public availability of reasoning traces creates valuable opportunities to study model behavior not only at the surface level but also at the granularity of individual reasoning steps. However, understanding the types of thinking employed during reasoning - which offers critical insights into models' reasoning patterns and enables actionable applications - remains underexplored. To address this gap, we introduce a framework for automatic annotation of reasoning steps through the lens of Bloom's Taxonomy, which classifies thinking into six cognitive levels, such as Remembering, Applying and Evaluating. Using this framework, we perform a large-scale analysis across models and datasets, revealing both similarities and differences in thinking patterns across models and tasks. Moreover, we demonstrate that thinking-type information derived from reasoning traces correlates with correctness, paving the way for improved reasoning. Our findings establish a fine-grained framework for analyzing thinking patterns in LRMs and provide actionable insights for enhancing reasoning quality.
Chinese Translation
大型推理模型(LRMs)在大规模语言模型(LLMs)的推理方面带来了革命性的变化,而推理痕迹的日益公开为研究模型行为提供了宝贵的机会,不仅可以在表面层面进行研究,还可以深入到个体推理步骤的细节。然而,理解推理过程中所采用的思维类型——这为模型的推理模式提供了关键洞察,并使得可操作的应用成为可能——仍然未得到充分探索。为了解决这一空白,我们引入了一个通过布鲁姆分类法自动标注推理步骤的框架,该分类法将思维分为六个认知层次,如记忆、应用和评估。利用该框架,我们对多个模型和数据集进行了大规模分析,揭示了模型和任务之间思维模式的相似性和差异。此外,我们证明了从推理痕迹中提取的思维类型信息与正确性相关,为改善推理铺平了道路。我们的研究结果建立了一个细粒度的框架,用于分析大型推理模型中的思维模式,并提供了可操作的洞察,以提升推理质量。
cs.AI / 145 / 2608.23218

What is mathematics now, and what should it be?

现在的数学是什么,它应该是什么?
Avigad, Jeremy
Abstract
Advances in neural theorem provers have been impressive, but the successes obscure a broader vision of what AI can do for mathematics and how mathematicians can engage with AI. This essay advances a more expansive and optimistic point of view.
Chinese Translation
神经定理证明器的进展令人印象深刻,但这些成功掩盖了人工智能在数学领域的更广泛愿景,以及数学家如何与人工智能进行互动。本文提出了一种更为广阔和乐观的观点。
cs.AI / 146 / 2608.23256

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data

下一块推理强化学习真的比监督微调更好吗?在无链式推理数据下重新审视训练策略
Tang, Yinhao, Fang, Youqing, Sun, Yanan, Liu, Jiangning, Wang, Ziyi, Zhao, Xun, Zhang, Weiming, Liu, Bin, Liu, Kuikun, Zhang, Wenwei, Chen, Kai
Abstract
Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbook derivations that contain reasoning-rich content but lack explicit chain-of-thought annotations. The method trains a model to generate implicit reasoning traces and rewards them by their ability to predict the next chunk of text. While promising, existing evaluations primarily compare against conventional SFT baselines, leaving open whether the gains come from the RL formulation itself or from more effectively exposing the model to no-CoT data. We address this question with a controlled study of next-chunk reasoning RL and a simple but previously overlooked alternative: Mixed SFT, a single supervised fine-tuning stage that jointly trains on no-CoT and long-CoT data. Despite its simplicity, Mixed SFT achieves a clearly higher post-RLVR performance ceiling than next-chunk reasoning RL while requiring over 60 times less training compute. The advantage is consistent across in-domain mathematical reasoning and out-of-domain reasoning tasks. Moreover, we show that higher pre-RLVR accuracy does not necessarily translate into higher post-RLVR accuracy, highlighting the need to evaluate no-CoT training strategies in the context of the full post-training pipeline.
Chinese Translation
近期的研究提出了下一块推理强化学习(next-chunk reasoning RL),旨在利用无链式推理(no-CoT)数据——如包含丰富推理内容但缺乏明确链式推理注释的解题过程和教科书推导等语料库。该方法训练模型生成隐式推理轨迹,并根据其预测下一块文本的能力给予奖励。尽管前景可期,现有评估主要与传统的监督微调(SFT)基线进行比较,尚不清楚性能提升是源于强化学习(RL)框架本身,还是更有效地将模型暴露于无链式推理数据之中。我们通过对下一块推理强化学习和一个简单但之前被忽视的替代方案——混合监督微调(Mixed SFT)进行控制研究来解决这一问题,后者是在无链式推理和长链式推理数据上联合训练的单一监督微调阶段。尽管其简单性,混合监督微调在后强化学习验证(post-RLVR)性能上明显高于下一块推理强化学习,同时所需的训练计算量少于60倍。该优势在领域内的数学推理和领域外的推理任务中均表现一致。此外,我们还表明,较高的前强化学习验证准确率并不一定转化为较高的后强化学习验证准确率,强调了在完整的后训练流程中评估无链式推理训练策略的必要性。
cs.AI / 147 / 2608.23263

Automated Construction of FAIR Digital Object Knowledge Graphs from Flat Cultural Heritage Records

从平面文化遗产记录自动构建 FAIR 数字对象知识图谱
Boukhers, Zeyd, Kong, Lingxiao, Zabulis, Xenophon, Toubekis, Georgios
Abstract
The FAIR Digital Object (FDO) framework mandates that metadata attribute values be expressed as persistent identifiers (PIDs) wherever possible, to produce a fully machine-actionable graph in which every reference is resolvable. The Europeana Data Model was designed long before the FDO specification, and it stores most metadata values as plain text. This serves human browsing well enough, but gives an automated agent nothing to follow across records or collections. We present a pipeline that transforms flat Europeana records into an FDO-compliant knowledge graph structured with CIDOC-CRM. Following the FDO specification, we model every heritage entity as a discrete FDO with its own PID, type, profile, and metadata layer. The core technical challenge is automating the FDO-prescribed distinction between values that must become PID references (resolvable entities) and those that may remain literals (terminal leaves such as notes, measurements, and dates). We address this with a large language model that classifies each metadata value, routes it to a controlled vocabulary (Getty AAT, Wikidata, VIAF, PeriodO), and links it to a shared entity FDO. We evaluate using 637 archaeological records from five Europeana providers, processing each with the LLM. The pipeline links 86% of metadata slots, resolving 58.5% of values Europeana had not already enriched. It also merges cross-lingual surface forms that byte-identical matching keeps apart, where 17 of 33 such merges are correct on manual review. Graph connectivity does not separate this from string matching; what distinguishes the FDO graph is that every node is typed and resolvable.
Chinese Translation
FAIR 数字对象 (FDO) 框架要求在可能的情况下将元数据属性值表示为持久标识符 (PID),以生成一个完全可机器操作的图谱,其中每个引用都是可解析的。欧洲数字图书馆数据模型在 FDO 规范之前就已设计,其大多数元数据值以纯文本形式存储。这对于人类浏览足够有效,但对自动化代理而言却没有可供跟踪的记录或集合。我们提出了一种管道,将平面欧洲数字图书馆记录转换为符合 FDO 标准的知识图谱,并采用 CIDOC-CRM 结构。按照 FDO 规范,我们将每个遗产实体建模为一个独立的 FDO,具有自己的 PID、类型、配置文件和元数据层。核心技术挑战在于自动化区分必须成为 PID 引用(可解析实体)和值得以文字形式保留(终端叶子,如注释、测量和日期)。我们通过一个大型语言模型来解决这个问题,该模型对每个元数据值进行分类,将其路由到受控词汇(Getty AAT、Wikidata、VIAF、PeriodO),并将其链接到共享实体 FDO。我们使用来自五个欧洲数字图书馆提供者的 637 条考古记录进行评估,使用 LLM 处理每条记录。该管道链接了 86% 的元数据槽,解析了 58.5% 欧洲数字图书馆尚未丰富的值。它还合并了由于字节相同匹配而分开的跨语言表面形式,其中 33 个合并中有 17 个在人工审查中是正确的。图谱的连通性并未将其与字符串匹配分开;区分 FDO 图谱的是每个节点都是有类型且可解析的。
cs.AI / 148 / 2608.23264

Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance

隐藏在请求中:通过标记相关性解释不道德的LLM合规性
Biton, Or, Krichli, Tomer, Allouche, Itai, Keshet, Joseph
Abstract
Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignment failures. This work systematically investigates instances where LLMs fail to exhibit ethical behavior. To understand the underlying mechanics of these vulnerabilities, we introduce a probing methodology that presents unethical scenarios to LLMs in three distinct structural modalities: objective classification tasks, subjective first-person statements, and direct requests for assistance. We find that model performance degrades in the request-for-assistance-based form. Using Layer-wise Relevance Propagation (LRP), we trace this discrepancy to an attribution bias: the model places greater emphasis on benign task-framing tokens (e.g., "Can you help me...") than on tokens signaling the underlying unethical behavior (e.g., "without getting caught"), which we term cue-tokens. We hypothesize that this under-attribution contributes to harmful compliance. To test this, we introduce two LRP-guided decoding methods that steer generation toward trajectories more relevant to cue tokens. Empirical evaluations show that these interventions promote safer responses, supporting cue-token attribution's role in compliance failures.
Chinese Translation
尽管大型语言模型(LLMs)旨在优化有用性和无害性这两个目标,但这两个目标可能会发生冲突,必然导致对齐失败。本研究系统地调查了LLMs未能表现出伦理行为的实例。为了理解这些脆弱性的潜在机制,我们引入了一种探测方法,通过三种不同的结构模式向LLMs呈现不道德的场景:目标分类任务、主观第一人称陈述和直接请求帮助。我们发现,在基于请求帮助的形式中,模型性能下降。通过层次相关传播(Layer-wise Relevance Propagation, LRP),我们追踪到这种差异源于归因偏差:模型对良性任务框架标记(例如,“你能帮我……吗?”)的重视程度高于对信号不道德行为的标记(例如,“不被抓到”)的重视程度,我们称之为提示标记(cue-tokens)。我们假设这种低估归因会导致有害的合规性。为了验证这一点,我们引入了两种基于LRP的解码方法,旨在引导生成更相关于提示标记的轨迹。实证评估表明,这些干预措施促进了更安全的响应,支持了提示标记归因在合规失败中的作用。
cs.AI / 149 / 2608.23283

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Apodex 1.1:为复杂工作扩展代理智能
Apodex Team, An, B., Li, B., Wang, B., Zhang, B., Wang, B. L., Feng, C., Wei, C., Xue, C., Zhang, C., Ng, D., Ye, D., Min, E., Chen, F., Liu, F., Yang, F., Ye, F., Xu, H., Yang, H., Ye, H., Zhang, H., Zhao, H., Li, J., Lin, J., Xia, J., Jin, K., Wang, K., Yang, K., Bing, L., Lei, L., Su, L., Wang, Le., Wang, Lu., Wang, N., Ren, Q., Yang, Q., Li, R., Bai, S., Du, S., Li, S., Lin, S., Nie, S., Wang, S., Zhang, S., Wang, S. Z., Fang, Ta. Q., Fang, Ti. Q., Fang, W., Li, W., Zhang, W., Chen, X., Li, X., Tang, X., Wang, X., Xu, X., Zhang, X., Wang, X. Q., Wang, X. Y., Deng, Y., Gao, Y., Hu, Y., Li, Y., Sui, Y., Wang, Y., Xiao, Y., Zhang, Y., Chen, Z., Cheng, Z., Feng, Z., Liang, Z., Zhang, Z.
Abstract
General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this \emph{working capability}: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimensions. \emph{Environment Scaling} expands the diversity and verifiability of executable file, search, and code environments, while \emph{Agentic Coordination Scaling} trains agents to decompose long-horizon tasks, delegate parallel work, integrate asynchronous results, and replan. A shared execution harness and AgentOS maintain task state and provenance across tools and agents, and training turns environment trajectories and coordination traces into reliable behavior. Across complex professional work, finance, scientific research, mathematics, coding, and search, Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems. The 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form. These results ground agentic intelligence in useful, verifiable work completed over time and advance our goal of building a \emph{Heavy-Duty Solver} for ambitious, long-running tasks.
Chinese Translation
通用语言模型能够推理和综合知识,但复杂工作还需要与文件、信息源和可执行代码进行持续交互,同时进行状态维护、故障恢复和可验证交付。我们称之为 extit{工作能力}:朝着现实目标持续、可验证的进展。Apodex 1.1在两个互补维度上发展这一能力。 extit{环境扩展}增加了可执行文件、搜索和代码环境的多样性和可验证性,而 extit{代理协调扩展}则训练代理分解长期任务、委派并行工作、整合异步结果和重新规划。一个共享执行框架和AgentOS在工具和代理之间维护任务状态和来源,训练将环境轨迹和协调痕迹转化为可靠的行为。在复杂的专业工作、金融、科学研究、数学、编码和搜索领域,尽管使用的模型远小于许多前沿系统,Apodex 1.1仍然达到了领先的性能水平。35B参数的Apodex 1.1 Mini在本地可部署的形式中进一步保持了强大的工作能力。这些结果将代理智能扎根于随着时间推移完成的有用、可验证的工作,并推动我们构建 extit{重型求解器}以应对雄心勃勃的长期任务的目标。
cs.AI / 150 / 2608.23313

EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models

EviSafe:基于证据的视觉语言模型安全评估
Li, Xuetong, Liu, Gaofeng
Abstract
Vision-language model safety benchmarks typically evaluate only final responses: whether a model refuses, warns, or complies. This outcome-level view cannot tell whether a model is safe for the right multimodal reason. Safelooking behavior may reflect keyword-triggered refusal, missed visual hazards, or over-refusal of benign-sensitive inputs. We introduce EviSafe, an evidence-grounded framework for VLM safety that jointly evaluates natural user-facing behavior, explicit grounding in textual and visual evidence, and behavioral sensitivity to counterfactual changes in safety-critical evidence. EviSafeBench instantiates the framework as a controlled benchmark with 1,181 gold image-text scenarios and 2,452 targeted counterfactual variants across eight safety domains and eight risk-source types. Each scenario includes a gold safety decision, evidence annotations, a safe-response policy, and counterfactual interventions. The three-probe protocol queries models with natural-response, evidencereporting, and counterfactual-response prompts, then scores them using an evidence-aware judge. Across eleven evaluated VLMs, natural severity accuracy ranges from 27.6% to 52.8%, relaxed diagnostic consistency from 6.1% to 29.3%, and unsafe-to-safe counterfactual transition success from 30.4% to 58.4%. These gaps show that the evaluated VLMs are not reliably safe for the right multimodal reason and motivate evaluation beyond refusal counts.
Chinese Translation
视觉语言模型的安全基准通常仅评估最终响应:模型是拒绝、警告还是遵从。这种结果导向的视角无法判断模型是否因正确的多模态原因而安全。看似安全的行为可能反映了关键词触发的拒绝、未能识别的视觉危险或对良性敏感输入的过度拒绝。我们提出了EviSafe,一个基于证据的视觉语言模型安全框架,联合评估自然用户行为、文本和视觉证据的明确基础,以及对安全关键证据的反事实变化的行为敏感性。EviSafeBench将该框架实例化为一个受控基准,包含1,181个金标准图像-文本场景和2,452个针对性的反事实变体,涵盖八个安全领域和八种风险来源类型。每个场景包括一个金标准安全决策、证据注释、安全响应策略和反事实干预。三探针协议通过自然响应、证据报告和反事实响应提示对模型进行查询,然后使用证据感知评审进行评分。在评估的十一种视觉语言模型中,自然严重性准确率范围为27.6%至52.8%,放宽的诊断一致性范围为6.1%至29.3%,不安全到安全的反事实转变成功率范围为30.4%至58.4%。这些差距表明,被评估的视觉语言模型在正确的多模态原因下并不可靠安全,并激励我们进行超越拒绝计数的评估。
cs.AI / 151 / 2608.23318

Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning

Agent-G$^2$: 基于高斯指导的自主强化学习
Wang, Zixuan, Miao, Yanrui, Lu, Zhengxi, Pan, Teng, Qiu, Yiwen, Li, Hongxing, Qiu, Peng, Zhang, Ruiqing, Shen, Yongliang
Abstract
Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory before each rollout, letting the policy explore from a state closer to success. Its effectiveness hinges on the guidance depth: how much of the trajectory to keep. Existing methods treat this depth as a deterministic scalar. Scheduled approaches share one value across samples and ignore per-task heterogeneity; per-sample probing estimates it separately at the cost of extra rollouts. We find that useful guidance occupies a band of depths whose informativeness profile is approximately Gaussian around the band center, rather than concentrating at a single optimal point. We propose Agent-G$^2$, a Gaussian guidance framework that draws the depth per task from a Gaussian whose center and spread are estimated online from rollouts already collected for policy optimization, requiring no probe rollouts or learned depth predictor. The center combines a global baseline with per-cluster difficulty, and the spread tracks within-cluster variance. We evaluate Agent-G$^2$ on ALFWorld and WebShop on Qwen2.5-1.5B / 7B-Instruct. Agent-G$^2$ outperforms the strongest hint-based, hint-free, and Aux-RL baselines on ALFWorld by 2.3 / 3.9 / 7.4 points at under one-third the rollout cost of per-sample probing.
Chinese Translation
基于提示的强化学习通过在每次回合前保留专家轨迹的前缀,解决了长时间跨度自主任务中的奖励稀疏问题,使得策略能够从更接近成功的状态进行探索。其有效性依赖于指导深度:保留多少轨迹。现有方法将这一深度视为一个确定性的标量。调度方法在样本间共享一个值,忽视了每个任务的异质性;每个样本探测则单独估计深度,但需要额外的回合。我们发现,有效的指导占据了一系列深度,其信息量特征在该范围中心附近呈现近似高斯分布,而不是集中在单一的最优点。我们提出了Agent-G$^2$,一个高斯指导框架,它从一个高斯分布中为每个任务抽取深度,该高斯的中心和扩展是通过在线估计已经收集的回合数据来进行策略优化的,无需探测回合或学习深度预测器。中心结合了全局基线和每个聚类的难度,扩展则跟踪聚类内的方差。我们在ALFWorld和WebShop上评估了Agent-G$^2$,使用Qwen2.5-1.5B / 7B-Instruct。Agent-G$^2$在ALFWorld上以每个样本探测不到三分之一的回合成本,超越了最强的基于提示、无提示和Aux-RL基线,分别提高了2.3 / 3.9 / 7.4分。
cs.AI / 152 / 2608.23370

Walking on the DARKSIDE

走在黑暗面上
Gangemi, Aldo, Bottazzi, Emanuele
Abstract
Large Language Models (LLMs) recognise patterns but do not natively track the path of exclusions that a coherent discourse demands. When an input rests on a fabricated authority, a misapplied mechanism, or a surreptitious analogy, an unsteered LLM tends to engage with it as if it were grounded, and to reify the misstep into any structured output it generates. Logic-Augmented Generation (LAG) with POLANYI++, an LLM-steering method that uses heuristics, ontologies and problem solving methods for tacit knowledge extraction, produces an Extended Knowledge Graph (XKG) in OWL2, but inherits the same vulnerability: a sophisticated nonsensical input is reified into the graph alongside the legitimate triples, and is hardly detectable by automated reasoners since the XKG is generated jointly with the wrong assumptions. We introduce DARKSIDE, a coherence auditing method on top of POLANYI++. It formalises the trail as an explicit data structure of accumulated exclusions over discourse time, complemented by a warrant axis that classifies each named referent as Warranted, Unattested, Misattributed or Fabricated, with an escalation rule that pushes the DelegationRiskAssessment to UNSAFE when the fabricated rate is positive or the unsupported rate exceeds a threshold. We evaluate DARKSIDE as a steering layer over a Gemini 3 on BSBench, a 100-item adversarial corpus of sophisticated-sounding nonsense across software engineering, finance, healthcare, physics and law, with Claude Sonnet 4.6 as an independent judge. The empirical evidence supports an architectural claim: when an LLM forward pass is wrapped in an ontology-mediated negative-trail apparatus, the structural pattern-vs-path gap can be partially scaffolded. The XKG functions as the missing memory, and the warrant axis as an epistemic firewall.
Chinese Translation
大型语言模型(LLMs)能够识别模式,但并不天然跟踪连贯话语所需的排除路径。当输入依赖于虚构的权威、错误应用的机制或隐秘的类比时,未加引导的LLM往往会将其视为有根基的内容,并将错误的步骤固化为其生成的任何结构化输出。使用POLANYI++的逻辑增强生成(LAG)是一种LLM引导方法,利用启发式、本体和问题解决方法进行隐性知识提取,生成一个OWL2格式的扩展知识图(XKG),但仍然继承了相同的脆弱性:复杂的无意义输入与合法三元组一起被固化到图中,并且由于XKG是与错误假设共同生成的,因此几乎无法被自动推理器检测到。我们提出了DARKSIDE,这是一种基于POLANYI++的连贯性审计方法。它将排除路径形式化为一个明确的数据结构,记录在话语时间内累积的排除,并通过一个担保轴对每个命名参照物进行分类,分为有担保、未证实、误归属或虚构,并设有升级规则,当虚构率为正或未支持率超过阈值时,将DelegationRiskAssessment推向不安全状态。我们在BSBench上评估DARKSIDE作为Gemini 3的引导层,BSBench是一个包含100个复杂无意义内容的对抗性语料库,涵盖软件工程、金融、医疗、物理和法律领域,并由Claude Sonnet 4.6作为独立评审。实证证据支持一个架构性主张:当LLM的前向传递被包裹在一个本体介导的负轨迹装置中时,结构模式与路径之间的差距可以部分得到支撑。XKG作为缺失的记忆,而担保轴则作为一种认识防火墙。
cs.AI / 153 / 2608.23373

Modalities Should Talk to Each Other: Dual-Stream Multimodal Learning for Long-Horizon Influenza Forecasting

模态之间应相互交流:用于长时间范围流感预测的双流多模态学习
Hashemi, Seyed Mohammad Hossein, Hooshmand, Mohsen, Razzaghi, Parvin
Abstract
Forecasting long-range influenza-like illness (ILI) matters for public health readiness. Publicly available surveillance datasets typically pair numeric epidemiological signals with textual information that is noisy, loosely structured, only indirectly related to near-term trends, and often lagged relative to the numeric signal. Fusing the two therefore requires careful design. We propose Dual-Stream Attention (DSA), a multimodal deep learning framework that forecasts 12-week-ahead ILI activity from a 36-week multimodal history by letting the numerical and textual streams condition each other. Using the Time-MMD health-domain dataset, DSA separately encodes the two modalities with a Transformer-based numerical encoder and a domain-adapted headline encoder, then couples them through a bidirectional Cross-Modal Attention (CMA) mechanism: the text (news headlines) conditions the interpretation of the numeric signal and vice versa. The CMA output then passes to a causal temporal model for forecasting. Evaluated across ten random seeds, DSA achieves a median test MSE of 0.416, versus 0.668, 0.607, and 0.851 for iTransformer, TaTS, and GPT4MTS, corresponding to mean-error reductions of 54.95%, 37.29%, and 67.23%, with paired Cohen's d of 0.555, 0.337, and 0.345, respectively, and ranks first in 100% of bootstrap draws. It also has substantially lower worst-window error than all baselines. On an external-geography dataset, DSA again ranks first among nine evaluated baselines. Ablations show the advantage does not depend on text-encoder choice or language-model fine-tuning, and that bidirectional attention outperforms either direction alone. Finally, perturbation-based faithfulness analysis shows the learned CMA is functionally informative under targeted masking, with a stronger effect in the text-to-numerical direction.
Chinese Translation
预测长期流感样疾病(ILI)对于公共卫生准备至关重要。公开可用的监测数据集通常将数值流行病信号与嘈杂、结构松散、仅间接与短期趋势相关且通常滞后于数值信号的文本信息配对。因此,融合这两者需要精心设计。我们提出了双流注意力(Dual-Stream Attention, DSA),这是一种多模态深度学习框架,通过让数值流和文本流相互调节,从36周的多模态历史中预测12周后的ILI活动。使用Time-MMD健康领域数据集,DSA分别使用基于Transformer的数值编码器和领域适应的标题编码器对这两种模态进行编码,然后通过双向跨模态注意力(Cross-Modal Attention, CMA)机制将它们结合起来:文本(新闻标题)调节数值信号的解释,反之亦然。CMA的输出随后传递给因果时间模型进行预测。在十个随机种子下评估,DSA的中位测试均方误差(MSE)为0.416,而iTransformer、TaTS和GPT4MTS分别为0.668、0.607和0.851,平均误差减少率分别为54.95%、37.29%和67.23%,配对Cohen's d分别为0.555、0.337和0.345,并在100%的自助抽样中排名第一。它在最差窗口误差方面也显著低于所有基线。在一个外部地理数据集上,DSA再次在九个评估基线中排名第一。消融实验表明,该优势不依赖于文本编码器的选择或语言模型的微调,并且双向注意力优于单向注意力。最后,基于扰动的可信度分析表明,学习到的CMA在目标掩蔽下具有功能性信息,在文本到数值的方向上效果更强。
cs.AI / 154 / 2608.23397

MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction

MediSkill-Evo:基于过程约束的自我进化以实现证据驱动的临床互动
Wu, Ruoyu, Xie, Shenfu, Sun, Yinqian, Tong, Haibo, Zhao, Feifei
Abstract
Interactive clinical agents must gather decisive evidence and convert it into grounded actions under partial observability. A correct final diagnosis alone does not show that an agent respected evidence and care-process constraints. We introduce MediSkill-Evo, a clinical agent that evolves governed process knowledge without backbone fine-tuning. It separates experience into four typed banks for clinical skills, process rules, symbolic schemas, and measurement procedures. Provenance, support, replay, and controller-defined safety checks govern publication to a frozen test-time snapshot. A Process-Constrained Preference Harness binds evidence to its source, rejects controller-invalid candidates, and ranks actions with a safety-prioritized Clinical Process Critic. We evaluate complete agent systems across two backbone endpoints and six controlled stress dimensions under the same Doctor-turn limit. On 300 held-out Qwen encounters, MediSkill-Evo improves diagnosis accuracy from 61.33 percent to 69.00 percent and treatment-intent coverage from 33.62 percent to 66.44 percent, while reducing automatically scored critical failures from 31.00 percent to 16.33 percent relative to AgentClinic. On 180 hard-isolation conditions derived from 30 cases, target recovery reaches 93.61 percent under patient-behavior pressure, 100.00 percent for temporal evidence, and 92.22 percent for triage red flags. An exploratory 100-case MedSAM comparison evaluates request-gated tool-interface feasibility. These results provide descriptive end-to-end evidence for the complete system on fixed evaluation suites, not causal evidence for an individual bank or clinical validation of the automatic judge.
Chinese Translation
交互式临床代理必须在部分可观察性下收集决定性证据并将其转化为有根据的行动。仅仅正确的最终诊断并不能表明代理遵循了证据和护理过程的约束。我们引入了MediSkill-Evo,这是一种在不进行主干微调的情况下进化的临床代理,具备受控的过程知识。它将经验分为四类银行,分别用于临床技能、过程规则、符号模式和测量程序。来源、支持、重放和控制器定义的安全检查管理发布到一个冻结的测试时间快照。一个基于过程约束的偏好约束将证据与其来源绑定,拒绝控制器无效的候选项,并通过一个安全优先的临床过程批评者对行动进行排名。我们在相同的医生轮次限制下,评估了跨两个主干端点和六个受控压力维度的完整代理系统。在300个保留的Qwen遇到中,MediSkill-Evo将诊断准确率从61.33%提高到69.00%,将治疗意图覆盖率从33.62%提高到66.44%,同时将自动评分的关键失败率从31.00%降低到16.33%,相较于AgentClinic。在从30个案例中衍生出的180个困难隔离条件下,目标恢复率在患者行为压力下达到93.61%,时间证据达到100.00%,而分诊红旗达到92.22%。一项探索性的100案例MedSAM比较评估了请求门控工具接口的可行性。这些结果为完整系统在固定评估套件上的描述性端到端证据提供了支持,而不是针对单个银行的因果证据或自动评判的临床验证。
cs.AI / 155 / 2608.23417

SkillAlchemy: Open-World Agent Skill Creation

技能炼金术:开放世界代理技能创建
Wang, Hengjun, Wei, Shuyue, Liu, Boyi, Yang, Jun, Tong, Yongxin
Abstract
Agent skills are reusable procedural artifacts that extend language agents with specialized workflows, tool conventions, and domain behaviors at inference time. However, creating reliable skills still depends largely on human authorship, model priors, or execution traces. These sources are often unavailable for unfamiliar tasks, suggesting the need to create skills from open-world materials. In this paper, we study open-world skill creation: given an underspecified skill brief and a source-access specification, a creator must discover behavior-relevant requirements omitted by the brief and determine how broadly each source-derived procedure is justified. We propose SkillAlchemy, an admission-centered framework for source-grounded skill creation. SkillAlchemy identifies implicit requirements through contrastive evidence, admits candidate procedures based on evidence-supported scope, and compiles the admitted content into a grammar-guided skill package. Extensive experiments across 87 SkillsBench v1.1 tasks demonstrate that SkillAlchemy improves pass rate over no-skill execution by 19.9 percentage points and the strongest automated baseline by 8.6 percentage points, while achieving performance comparable to human-curated skills.
Chinese Translation
代理技能是可重用的程序性工件,它们在推理时通过专门的工作流程、工具约定和领域行为扩展语言代理。然而,创建可靠的技能仍然在很大程度上依赖于人类创作者、模型先验或执行轨迹。这些来源在面对不熟悉的任务时往往不可用,这表明需要从开放世界材料中创建技能。本文研究开放世界技能创建:给定一个不明确的技能简要和一个源访问规范,创作者必须发现简要中遗漏的与行为相关的要求,并确定每个源派生程序的适用范围有多广。我们提出了技能炼金术(SkillAlchemy),这是一个以入学为中心的源基础技能创建框架。技能炼金术通过对比证据识别隐含要求,根据证据支持的范围接纳候选程序,并将接纳的内容编译成一个语法引导的技能包。在87个SkillsBench v1.1任务上的大量实验表明,技能炼金术相比于无技能执行提高了19.9个百分点的通过率,并比最强的自动化基线提高了8.6个百分点,同时实现了与人类策划技能相当的性能。
cs.AI / 156 / 2608.23446

Characterizing Necessary Losers to Explain Tournaments Losers

表征必要失败者以解释锦标赛失败者
Clément, Contet, Grandi, Umberto, Mengin, Jérôme
Abstract
We study the problem of formally explaining why a candidate was not selected by a given tournament rule, by identifying sub-tournaments in which the candidate loses independently of how the rest of the tournament is completed. We define destructive minimal supports as any minimal sub-tournaments satisfying this property, which in formal explainable artificial intelligence correspond to abductive explanations for the question "Why does the loser lose the tournament?". For six common tournament solutions (maximin, uncovered set and its weighted variant, top-cycle, Copeland, and Borda) we provide characterizations of when a candidate is either a necessary loser or a possible winner, we determine the size of the smallest destructive minimal supports, complemented by polynomial-time algorithms for their computation except for the case of the Borda rule which is suspected to be NP-complete.
Chinese Translation
我们研究了一个问题,即如何正式解释为什么某个候选人未被特定锦标赛规则选中,通过识别在候选人输掉比赛时与锦标赛其余部分的完成方式无关的子锦标赛。我们将破坏性最小支持定义为满足这一属性的任何最小子锦标赛,这在形式可解释的人工智能中对应于对“为什么失败者会输掉锦标赛?”这一问题的溯因解释。对于六种常见的锦标赛解决方案(最大最小法、未覆盖集及其加权变体、顶级循环、科普兰法和博尔达法),我们提供了候选人是必要失败者或可能赢家的特征描述,并确定了最小破坏性最小支持的大小,除了博尔达法的情况外,我们还提供了其计算的多项式时间算法,而博尔达法的情况被怀疑是NP完全的。
cs.AI / 157 / 2608.23475

StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models

StrategyBench:评估大型语言模型中的显式策略诱导
Tan, Jinghan, Wang, Yuanzheng, Chen, Lu, Chen, Zijun, Wang, Yuqian, Sun, Maosong
Abstract
As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a key paradigm for task adaptation. However, direct ICL often uses a small set of examples without explicitly abstracting task rules, making it sensitive to example construction. In contrast, human learners often reduce such sensitivity by first summarizing task rules from examples and then applying them to new instances. To evaluate this ability, we propose StrategyBench, which selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines evaluation metrics along two dimensions: strategy quality and downstream utility. We further analyze strategy induction from three perspectives: task variation, model configuration, and adaptation setting, covering category-wise differences, generator-executor choices, demonstration design, and SFT-based adaptation. Experiments show that explicit strategy utility differs substantially across task categories and depends on both strategy generation and execution conditions. The benchmark is released at: https://anonymous.4open.science/r/StrategyBench-D53C.
Chinese Translation
随着大型语言模型在数据稀缺和不断演变的任务场景中越来越多地被使用,少样本上下文学习(ICL)已成为任务适应的关键范式。然而,直接的 ICL 通常使用一小组示例,而没有明确抽象任务规则,这使得其对示例构造非常敏感。相比之下,人类学习者通常通过首先从示例中总结任务规则,然后将其应用于新实例,来减少这种敏感性。为了评估这种能力,我们提出了 StrategyBench,它从 BIG-Bench 中选择可诱导策略的任务,构建参考策略,并沿着两个维度定义评估指标:策略质量和下游效用。我们进一步从三个角度分析策略诱导:任务变异、模型配置和适应设置,涵盖类别差异、生成器-执行器选择、示范设计和基于 SFT 的适应。实验表明,显式策略的效用在不同任务类别之间存在显著差异,并且依赖于策略生成和执行条件。该基准已发布于:https://anonymous.4open.science/r/StrategyBench-D53C。
cs.AI / 158 / 2608.23484

Multi-Modal Semantic Expansion with Constrained LLM Reranking for Conversational Music Recommendation

基于约束大语言模型重排序的多模态语义扩展用于对话音乐推荐
Garg, Naman, Jain, Sarika, Fazekas, George
Abstract
We present Team Semiintelligencn's solution for the ACM RecSys 2026 TalkPlayData Challenge, addressing conversational music recommendation through a multi-modal and personalized conversational recommender system. Our submitted system employs a three-stage pipeline: (1) multi-modal retrieval constructing decay-weighted centroids across seven dense embedding spaces - track- and user-level CF-BPR, Qwen3 (metadata, lyrics, attributes), CLAP audio, and SigLIP visual - supplemented by BM25 lexical retrieval and an artist substring-match signal, all fused via weighted Reciprocal Rank Fusion (RRF) with optimized signal weights; (2) lightweight reranking (history filtering, popularity smoothing, and catalog diversity penalization); and (3) persona-diversified response generation using GPT-4o-mini. Beyond this submitted configuration, we report development-time experiments with additional components - constrained LLM-guided artist injection, album continuation signals, XGBoost LambdaMART, and a superior GPT-4.1 response prompt - that were not deployed to Blind B due to cost and complexity constraints. We optimize RRF weights on a 500-session development split via differential evolution, improving MRR by +19.5%. On Blind A, we observe that unconstrained LLM-guided injection across 54 sessions causes catastrophic nDCG regression (-18.9%), while conservative injection on only 9 sessions yields the best observed Blind A nDCG - a finding we present as a Blind A observation warranting further validation. The submitted system achieves a Blind B composite score of 0.3213.
Chinese Translation
我们提出了Team Semiintelligencn在ACM RecSys 2026 TalkPlayData Challenge中的解决方案,旨在通过一个多模态和个性化的对话推荐系统来解决对话音乐推荐问题。我们提交的系统采用了三阶段管道:(1) 多模态检索,通过在七个密集嵌入空间中构建衰减加权的质心——轨道和用户级CF-BPR、Qwen3(元数据、歌词、属性)、CLAP音频和SigLIP视觉——并辅以BM25词汇检索和艺术家子字符串匹配信号,所有信号通过加权的倒数排名融合(Reciprocal Rank Fusion, RRF)进行融合,并优化信号权重;(2) 轻量级重排序(历史过滤、流行度平滑和目录多样性惩罚);(3) 使用GPT-4o-mini生成个性化多样化的响应。除了提交的配置外,我们还报告了在开发阶段进行的实验,涉及额外组件——约束大语言模型引导的艺术家注入、专辑延续信号、XGBoost LambdaMART,以及更优的GPT-4.1响应提示——由于成本和复杂性限制,这些组件未被部署到Blind B。我们通过差分进化优化RRF权重,在500个会话的开发集上将MRR提高了19.5%。在Blind A上,我们观察到在54个会话中进行的无约束大语言模型引导注入导致了灾难性的nDCG回归(-18.9%),而在仅9个会话中进行的保守注入则获得了最佳的Blind A nDCG——这一发现我们认为是Blind A观察的一个值得进一步验证的结果。提交的系统在Blind B上获得了0.3213的综合得分。
cs.AI / 159 / 2608.23493

SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

SRPO:用于长时间推理的自我反思策略优化
Liu, Jialong, Shi, Yuling, Yang, Ning, Gu, Xiaodong, Li, Zuchao
Abstract
Self-reflection is a powerful mechanism for credit assignment in human learning, converting sparse outcome feedback into actionable guidance. However, its potential for post-training Large Language Models (LLMs) remains underexplored. We propose Self-Reflective Policy Optimization (SRPO), a framework that internalizes this capability. SRPO enables LLMs to analyze their own completed trajectories, synthesize errors into concise "reflection patches," and use reflection-conditioned teacher scores on student on-policy rollouts as dense token-level training signals. This process effectively transforms sparse terminal supervision into dense, token-level learning signals without requiring external critics, separate reward models, or larger teacher models. We demonstrate that SRPO achieves state-of-the-art performance across mathematical reasoning and long-horizon agentic benchmarks with exceptional data efficiency. Using a Qwen3-8B base model, SRPO attains 73.3% on AIME'24 using only 8% (0.08x) of the training FLOPs required by scaled supervised fine-tuning, while significantly improving success rates on WebShop (64.7%), ALFWorld (76.8%), and SWE-Bench-Lite (31.2%). Code is available at https://github.com/Galleons2029/SRPO
Chinese Translation
自我反思是人类学习中一种强大的信用分配机制,它将稀疏的结果反馈转化为可操作的指导。然而,其在训练后大规模语言模型(LLMs)中的潜力仍未得到充分探索。我们提出了自我反思策略优化(Self-Reflective Policy Optimization, SRPO),这是一个内化这一能力的框架。SRPO使LLMs能够分析其自身完成的轨迹,将错误综合为简明的“反思补丁”,并使用反思条件的教师评分作为学生在策略执行中的密集令牌级训练信号。这个过程有效地将稀疏的终端监督转化为密集的令牌级学习信号,而无需外部批评者、单独的奖励模型或更大的教师模型。我们证明SRPO在数学推理和长时间代理基准测试中实现了最先进的性能,并具有卓越的数据效率。使用Qwen3-8B基础模型,SRPO在AIME'24上达到了73.3%的成绩,仅使用了8%(0.08x)经过缩放的监督微调所需的训练FLOPs,同时显著提高了WebShop(64.7%)、ALFWorld(76.8%)和SWE-Bench-Lite(31.2%)的成功率。代码可在https://github.com/Galleons2029/SRPO获取。
cs.AI / 160 / 2608.23497

Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty

通过安全方向惩罚减轻推理引起的错位
Zhao, Yipeng, Yang, Qishun, Zhu, Shenzhe, Yang, Shu, Wang, Di
Abstract
Reasoning-Induced Misalignment, where fine-tuning on reasoning data containing no harmful content, including mathematics, code, and problem-solving with chain-of-thought traces can induce harmful behaviors of LLM, posing a serious challenge to the safety of LLM reasoning. Cross-architecture, cross-scale, and cross-dataset checks show that RIM does not always emerge. Previous work attributed RIM to neuron-level entanglement, but did not identify the geometry of the representation space underlying this entanglement or propose a training-time fix. We provide both: a representation-space analysis of RIM and the Safety-Direction Penalty (SDP), which penalizes movement along a learned safety direction during reasoning fine-tuning. The analysis extracts two activation-space directions, one encoding reasoning ability and the other safety behavior. These directions are coupled: fine-tuning that improves reasoning shifts safety representations, and prompts with larger shifts show larger safety degradation. CKA distance ratios and probes locate the safety-decision layers where this shift is most relevant. These findings guide the design of SDP: the coupling motivates penalizing displacement along the safety direction, and the layer localization sets the initial scope. When the initial scope leaves compensatory shifts beyond the penalized layers, the same diagnostics guide iterative expansion. On Qwen2.5-3B and 7B, SDP restores safety while preserving benchmark reasoning performance.
Chinese Translation
推理引起的错位(Reasoning-Induced Misalignment, RIM)指的是在不包含有害内容的推理数据上进行微调时,包括数学、代码和带有思维链迹的解题过程,可能会导致大型语言模型(LLM)产生有害行为,这对LLM推理的安全性构成了严重挑战。跨架构、跨规模和跨数据集的检查表明,RIM并不总是出现。先前的研究将RIM归因于神经元级别的纠缠,但并未识别出这种纠缠背后的表示空间几何结构或提出训练时的解决方案。我们提供了两者:对RIM的表示空间分析和安全方向惩罚(Safety-Direction Penalty, SDP),后者在推理微调过程中惩罚沿学习到的安全方向的移动。该分析提取了两个激活空间方向,一个编码推理能力,另一个编码安全行为。这些方向是耦合的:改善推理的微调会改变安全表示,而具有更大变化的提示会显示出更大的安全退化。CKA距离比和探针定位了这种变化最相关的安全决策层。这些发现指导了SDP的设计:耦合关系促使惩罚沿安全方向的位移,而层定位则设定了初始范围。当初始范围将补偿性位移留在被惩罚层之外时,相同的诊断将指导迭代扩展。在Qwen2.5-3B和7B上,SDP在保持基准推理性能的同时恢复了安全性。
cs.AI / 161 / 2608.23525

EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

EarthVerse:动态地球系统和自然灾害中科学智能体的基准测试
Cui, Zhiqing, Yin, Xinxiang, Tang, Yihong, Zhang, Xinglang, Hu, Yuanzhe, Zhong, Siru, Tang, Weidong, Liang, Yuxuan, Li, Weijia, Jin, Ming, Pan, Shirui, Kang, Yuhao, Zhuang, Dingyi, Zhao, Jinhua
Abstract
Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing, and modality. Natural hazards make this work consequential because incomplete evidence can change estimates of severity, exposure, and mechanism. We introduce EarthVerse, a benchmark that evaluates scientific agents through package-scoped investigations. Its 405 reproducible tasks are grounded in 199 documented events and 19 hazard families. Agents inspect heterogeneous event packages, choose compatible evidence, execute transparent calculations, reconcile source differences, and preserve provenance in the final answer. We provide executable ground truth that decomposes each task into fine-grained answer units, together with task-specific rubrics that assess the supporting research process while allowing multiple valid paths. We evaluate 25 model and agent systems under a controlled tool-using protocol, then use controlled studies to locate failures in evidence access, tool selection, memory, reasoning, interaction, and scientific execution. Across systems, the best mean answer-unit accuracy is 84.65%, while the highest Strict@95 is only 34.81%. The gap shows that current agents often complete individual steps without maintaining a consistent chain across evidence, scales, units, calculations, and physical interpretation. EarthVerse provides a reproducible basis for measuring end-to-end scientific reliability in dynamic Earth systems.
Chinese Translation
地球系统分析通过不同来源、尺度、时间和方式的观测重建变化的物理过程。自然灾害使这项工作变得至关重要,因为不完整的证据可能会改变对严重性、暴露和机制的估计。我们介绍了EarthVerse,一个通过包范围的调查评估科学智能体的基准。其405个可重复的任务基于199个记录事件和19个灾害类别。智能体检查异构事件包,选择兼容的证据,执行透明的计算,调和来源差异,并在最终答案中保留来源信息。我们提供可执行的真实数据,将每个任务分解为细粒度的答案单元,并附上特定任务的评分标准,以评估支持研究过程,同时允许多条有效路径。我们在受控的工具使用协议下评估25个模型和智能体系统,然后通过受控研究定位在证据获取、工具选择、记忆、推理、互动和科学执行中的失败。在各系统中,最佳的平均答案单元准确率为84.65%,而最高的Strict@95仅为34.81%。这一差距表明,当前的智能体往往在完成单个步骤时未能在证据、尺度、单位、计算和物理解释之间保持一致的链条。EarthVerse为测量动态地球系统中的端到端科学可靠性提供了可重复的基础。
cs.AI / 162 / 2608.23526

Correcting a learned physical invariant improves world-model rollouts

纠正学习到的物理不变量可改善世界模型的滚动预测
Bao, Richard
Abstract
World models can predict video without learning dynamics that they reliably preserve. We test whether a frozen DreamerV3 trained only on pendulum video learns a scalar that its own latent transition treats as approximately conserved. A label-free search recovers the same energy-like invariant across independently trained conservative models, while the same procedure finds no comparable invariant in matched damped models. During autonomous rollouts, this quantity drifts. Projecting the latent state back toward its initial level set reduces rollout error in all three conservative models, whereas matched random constraints usually increase it. These results distinguish a dynamically meaningful invariant from a merely decodable correlate and reveal a concrete failure mode: a world model can learn a physical constraint from pixels yet violate that constraint when it imagines forward.
Chinese Translation
世界模型可以在不学习可靠保持的动态的情况下预测视频。我们测试了一个仅在摆动视频上训练的冻结DreamerV3是否学习到了一个其自身潜在转移视为近似守恒的标量。无标签搜索在独立训练的保守模型中恢复了相同的类能量不变量,而在匹配的阻尼模型中则未发现可比的不变量。在自主滚动过程中,这一量会漂移。将潜在状态投影回其初始水平集可以减少所有三个保守模型的滚动误差,而匹配的随机约束通常会增加误差。这些结果区分了具有动态意义的不变量与仅仅可解码的相关量,并揭示了一种具体的失败模式:一个世界模型可以从像素中学习物理约束,但在向前推测时却违反该约束。
cs.AI / 163 / 2608.23543

How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles

人工智能辅助如何影响人类技能发展:逻辑难题学习的研究
Wu, Shang, Belem, Catarina G, Fu, Shuyuan, Steyvers, Mark, Smyth, Padhraic
Abstract
While AI assistance can improve human task performance in the short term, it may also undermine the development of skills in the longer term. We examine this tension in a controlled logic-puzzle experiment involving on-demand AI assistance, where participants complete tasks before, during, and after AI is available. By experimentally varying AI request costs, we find that lower-cost assistance induces more frequent AI use. We also find that participants who request AI assistance during the AI-access phase perform worse at the task after assistance is removed, and their subsequent unassisted performance is overestimated when predicted from earlier AI-assisted performance. We use a Bayesian latent ability model to separate initial ability, post-AI ability, and participant-specific skill change, while estimating how independent reasoning during the AI-access phase relates to skill development. The results show that greater independent problem-solving effort is associated with larger gains in latent ability, consistent with the interpretation that skill development is weaker when AI assistance substitutes for independent reasoning.
Chinese Translation
虽然人工智能辅助可以在短期内提高人类任务表现,但在长期内可能会削弱技能的发展。我们在一个控制的逻辑难题实验中考察了这一矛盾,参与者在人工智能可用之前、期间和之后完成任务。通过实验性地变化人工智能请求成本,我们发现较低成本的辅助会导致更频繁的人工智能使用。我们还发现,在人工智能可访问阶段请求人工智能辅助的参与者,在辅助被移除后在任务中的表现较差,并且他们后续的无辅助表现在预测时被高估。我们使用贝叶斯潜在能力模型来区分初始能力、后人工智能能力和参与者特定的技能变化,同时估计在人工智能可访问阶段独立推理与技能发展之间的关系。结果表明,更多的独立问题解决努力与潜在能力的更大提升相关,这与技能发展在人工智能辅助替代独立推理时较弱的解释一致。
cs.AI / 164 / 2608.23552

Prime Agent: A Self-Improving RLM Harness

Prime Agent:自我改进的递归语言模型工具
Karten, Seth, Zhang, Alex L., Thomas, Kevin, Müller, Sebastian, Bakouch, Elie, Auras, Daniel, Senghaas, Mika, Obeid, Fares, Dunas, Konstantin, Hagemann, Johannes, Jaghouar, Sami
Abstract
Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context. Prime Agent is an open-source harness for long-horizon evaluation and coding-agent workflows. A persistent IPython REPL follows the Recursive Language Model abstraction for programmatic context processing and test-time compute, while Continual Harness preserves histories, memories, skills, prompts, and subagent specifications across trajectories. Recursive subagents coordinate through direct agent-to-agent communication, and the Agents View lets humans inspect and manage daemon-backed sessions. Prime Agent standardizes execution, recovery, verification, and resource accounting while leaving strategy construction to the model. This low-friction, expressive membrane prevents harness failures from becoming model failures and pushes measurement toward the model's true maximal underlying capability. Prime Agent raises ARC-AGI-3 RHAE Best@1 from 30% to 95.5% and matches or exceeds native and popular harnesses across long-context coding, GPU-kernel generation, emulator construction, and autonomous nanoGPT speedruns. On Factorio, we find refinement allows for continuous technology progression and dedicated subagents enable parallelized work. Code is available at https://github.com/PrimeIntellect-ai/prime-agent.
Chinese Translation
语言模型是顺序处理器,但长时间跨度的代理需要超出模型权重和活动上下文的外部信息和计算。Prime Agent 是一个用于长时间跨度评估和编码代理工作流的开源工具。一个持久的 IPython REPL 遵循递归语言模型的抽象,用于程序上下文处理和测试时计算,而持续工具则在不同轨迹中保留历史、记忆、技能、提示和子代理规范。递归子代理通过直接的代理间通信进行协调,代理视图使人类能够检查和管理由守护进程支持的会话。Prime Agent 标准化了执行、恢复、验证和资源核算,同时将策略构建留给模型。这个低摩擦、富有表现力的膜防止了工具故障转变为模型故障,并推动测量朝向模型的真实最大潜在能力。Prime Agent 将 ARC-AGI-3 RHAE Best@1 从 30% 提升至 95.5%,并在长上下文编码、GPU 内核生成、仿真器构建和自主 nanoGPT 快速运行方面与本地和流行工具相匹配或超越。在 Factorio 中,我们发现精炼允许持续的技术进步,而专用子代理则实现了并行工作。代码可在 https://github.com/PrimeIntellect-ai/prime-agent 获取。
cs.AI / 165 / 2608.23565

ReWorld: An Interactive World Model with Long-Horizon Memory

ReWorld:具有长时记忆的交互式世界模型
Chen, Zhifei, Wang, Luozhou, Shen, Guibao, Yan, Dongyu, Yang, Shuai, Xu, Tianshuo, Du, Yihua, Wang, Wei, Gui, Tianyi, Huang, Lianghua, Chen, Yingcong
Abstract
An interactive world model must follow the user's actions, remember the places it has shown, and stream in real time. The tension is structural: control wants a short horizon, memory wants an unbounded one. ReWorld separates the two during training and bounds them at inference. Mixed per-head attention windows confine most heads to the recent past while a small set of global heads attends over the entire history, and random head routing keeps either capability from binding to particular heads; random chunk dropping makes sparse histories in-distribution. At inference the whole past lives under a fixed budget: a bounded KV cache backed by a pose-indexed landmark bank, from which the model retrieves the landmarks nearest the current pose. A metric-scale-aligned data engine places eight sources -- Unreal-rendered fly-throughs, game roaming, and real-world footage -- on one physical action scale, so the same key press moves the camera the same distance in every source, and palindrome trajectories supply the revisit evidence that memory training needs. Distribution-matching distillation confined to a LoRA adapter then compresses sampling to four steps: one backbone serves both a high-fidelity multi-step mode and a real-time interactive one, streaming 704x1280 video across photorealistic, game-style, and stylized worlds. Under a three-axis protocol covering action following, long-horizon recall, and video quality, against six recent interactive world models it attains the best control fidelity ($11.95^\circ$ rotation error and the best camera-motion consistency) and the best generation quality; and on minute-long out-and-back rollouts ($64$\,s, $384$ latents), its fixed 12-chunk cache still regenerates the starting view -- at rollout lengths where a sliding window has long evicted the evidence and full-KV attention runs out of memory.
Chinese Translation
交互式世界模型必须跟随用户的动作,记住其展示过的地方,并实时流式传输。其结构上的紧张关系在于:控制需要短期视野,而记忆则希望没有界限。ReWorld 在训练过程中将两者分开,并在推理时对其进行限制。混合的每头注意力窗口将大多数头限制在最近的过去,而一小部分全局头则关注整个历史,随机头路由使得任一能力不与特定头绑定;随机块丢弃使得稀疏历史在分布内。推理时,整个过去在固定预算下运行:一个由姿态索引的地标库支持的有界 KV 缓存,从中模型检索与当前姿态最近的地标。一个度量尺度对齐的数据引擎将八个来源——虚幻引擎渲染的飞行穿越、游戏漫游和真实世界视频——放置在同一物理动作尺度上,因此相同的按键在每个来源中移动相机的距离相同,而回文轨迹提供了记忆训练所需的重访证据。分布匹配蒸馏限定在 LoRA 适配器内,然后将采样压缩为四个步骤:一个主干同时服务于高保真多步模式和实时交互模式,在逼真、游戏风格和风格化世界中流式传输 704x1280 视频。在涵盖动作跟随、长时记忆回忆和视频质量的三轴协议下,与六个近期的交互式世界模型相比,它达到了最佳的控制保真度($11.95^ heta$ 旋转误差和最佳的相机运动一致性)以及最佳的生成质量;在分钟级的往返滚动($64$ s, $384$ 潜变量)中,其固定的 12 块缓存仍然能够再生起始视图——在滑动窗口早已驱逐证据且全 KV 注意力耗尽内存的滚动长度下。
计算语言学 (Computation and Language)
141
cs.CL / 1 / 2608.21364

Distinguishing Revision and Delayed Elaboration in Incremental Narrative Interpretation

区分增量叙事解释中的修订与延迟阐述
Chen, Yi-Chun
Abstract
Both human and AI systems that process narrative or long-form content operate incrementally: input is received over time, and internal representations must be updated accordingly. Incremental interpretation, therefore, depends not only on what is represented but also on how the representational state evolves under new evidence. We distinguish two structurally different update operators that arise in narrative interpretation: revision-driven update and delayed elaboration. Revision-driven updates retract or replace previously committed structure in response to a contradiction and are therefore non-monotonic. Delayed elaboration, by contrast, refines initially underspecified elements through constraint addition without retracting prior commitments, yielding monotonic extension of the interpretive state. Although both operators may alter how earlier material is understood, they impose fundamentally different structural requirements on state transitions. Using visual narratives as a diagnostic domain, we demonstrate how a structured narrative representation can explicitly separate committed from underspecified content and support both update operators during incremental construction. Through a worked example, we show how delayed elaboration enables monotonic refinement of interpretive state, while revision requires non-monotonic correction. We discuss the broader relevance of this structural distinction for incremental reasoning and hybrid symbolic-neural systems.
Chinese Translation
人类和人工智能系统在处理叙事或长篇内容时都是增量式运作的:输入是随着时间接收的,内部表征必须相应更新。因此,增量解释不仅依赖于所表征的内容,还依赖于在新证据下表征状态的演变。我们区分了在叙事解释中出现的两种结构上不同的更新操作符:修订驱动更新和延迟阐述。修订驱动更新在面对矛盾时撤回或替换先前承诺的结构,因此是非单调的。相比之下,延迟阐述通过添加约束来细化最初未明确的元素,而不撤回先前的承诺,从而实现解释状态的单调扩展。尽管这两种操作符可能改变对早期材料的理解方式,但它们对状态转变施加了根本不同的结构要求。通过使用视觉叙事作为诊断领域,我们展示了如何通过结构化的叙事表征明确区分承诺内容与未明确内容,并在增量构建过程中支持这两种更新操作符。通过一个实例,我们展示了延迟阐述如何实现解释状态的单调细化,而修订则需要非单调的修正。我们讨论了这一结构性区分对增量推理和混合符号-神经系统的更广泛相关性。
cs.CL / 2 / 2608.21365

KSE-Web: An Analysis of Hybrid Retrieval and LLM-Assisted Query Expansion for Low-Resource Khmer Semantic Search

KSE-Web:低资源高棉语语义搜索中的混合检索与LLM辅助查询扩展分析
Thuon, Nimol
Abstract
As a low-resource language, Khmer presents several retrieval challenges, including limited annotated data, ambiguous word boundaries, weak support in multilingual embedding models, and frequent mixed Khmer-English usage. This paper presents KSE-Web, an analysis of hybrid retrieval and LLM-assisted query expansion for Khmer semantic search. We construct the dataset from approximately 17K candidate Khmer titles and retain 3K cleaned full-text Khmer documents after filtering, normalization, deduplication, and document-length control. The dataset includes 300 manually reviewed user-style Khmer search queries and silver relevance labels with partial human verification. We evaluate character n-gram BM25, multilingual dense retrieval, hybrid BM25+dense retrieval, and LLM-assisted query expansion using Qwen2.5 models. Experimental results show that BM25 achieves the strongest overall performance, reaching 0.943 Recall and 0.876 nDCG. Hybrid BM25+dense retrieval performs comparably, achieving 0.929 Recall and 0.871 nDCG, while dense retrieval alone performs lower. LLM-assisted query expansion does not outperform non-expanded retrieval; however, Qwen2.5-3B produces substantially stronger expanded-query results than Qwen2.5-0.5B, suggesting that LLM size and expansion quality matter for low-resource Khmer retrieval. Our analysis further shows that direct LLM expansion can introduce topic drift, generic terms, and noisy reformulations, while simple filtering may remove useful semantic cues. These findings highlight both the potential and limitations of LLM-assisted retrieval for Khmer semantic search and provide a foundation for future Khmer retrieval datasets with stronger human-verified annotations and Khmer-aware retrieval models. The dataset and documentation will be made available at github.com/back-kh/KhmerSemantic-Search.
Chinese Translation
作为一种低资源语言,高棉语面临多种检索挑战,包括有限的标注数据、模糊的词边界、多语言嵌入模型的支持不足,以及频繁的高棉语-英语混用。本文提出了KSE-Web,分析了高棉语语义搜索中的混合检索和LLM(大语言模型)辅助查询扩展。我们从大约17,000个候选高棉语标题构建数据集,并在经过过滤、标准化、去重和文档长度控制后,保留了3,000个清理后的完整高棉语文档。该数据集包括300个经过人工审核的用户风格的高棉语搜索查询和部分人工验证的银级相关性标签。我们评估了字符n-gram BM25、多语言密集检索、混合BM25+密集检索和使用Qwen2.5模型的LLM辅助查询扩展。实验结果表明,BM25实现了最强的整体性能,达到0.943的召回率和0.876的nDCG。混合BM25+密集检索的表现相当,达到0.929的召回率和0.871的nDCG,而单独的密集检索表现较低。LLM辅助查询扩展的表现不如未扩展的检索;然而,Qwen2.5-3B生成的扩展查询结果明显优于Qwen2.5-0.5B,表明LLM的规模和扩展质量对低资源高棉语检索至关重要。我们的分析进一步表明,直接的LLM扩展可能引入主题漂移、通用术语和噪声重构,而简单的过滤可能会去除有用的语义线索。这些发现突显了LLM辅助检索在高棉语语义搜索中的潜力和局限性,并为未来具有更强人工验证注释和高棉语感知检索模型的高棉语检索数据集奠定了基础。数据集和文档将发布在github.com/back-kh/KhmerSemantic-Search。
cs.CL / 3 / 2608.21369

Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning

Wazobia Eval:尼日利亚皮钦语情感理解、讽刺检测和文化推理的基准
Okoye, Stephanie
Abstract
Nigerian Pidgin is one of Africa's most widely spoken languages, yet remains severely underrepresented in language model evaluation. Existing benchmarks primarily focus on translation, transcription, or generic sentiment analysis, leaving critical aspects of culturally grounded language understanding unmeasured. We introduce Wazobia Eval, a benchmark for evaluating Nigerian Pidgin emotion understanding, sarcasm detection, and cultural reasoning. The benchmark is built on a manually annotated dataset containing over 550 examples and a 16-category emotion taxonomy designed to capture culturally specific emotional registers that are not represented in conventional sentiment frameworks. Wazobia Eval provides standardized evaluation protocols and benchmark tasks for assessing model performance on nuanced Nigerian language understanding. We present the benchmark design, annotation methodology, taxonomy development process, and preliminary pilot evaluation results. Our goal is to provide foundational evaluation infrastructure for Nigerian language AI and establish a reproducible benchmark for future research. The dataset is publicly available at https://huggingface.co/WAZOBIALABS.
Chinese Translation
尼日利亚皮钦语是非洲最广泛使用的语言之一,但在语言模型评估中仍然严重缺乏代表性。现有基准主要集中在翻译、转录或通用情感分析上,未能衡量文化基础语言理解的重要方面。我们推出了Wazobia Eval,一个用于评估尼日利亚皮钦语情感理解、讽刺检测和文化推理的基准。该基准基于一个手动注释的数据集,包含超过550个示例和一个16类情感分类法,旨在捕捉在传统情感框架中未能体现的文化特定情感表达。Wazobia Eval提供了标准化的评估协议和基准任务,以评估模型在细致的尼日利亚语言理解上的表现。我们展示了基准设计、注释方法、分类法开发过程以及初步的试点评估结果。我们的目标是为尼日利亚语言人工智能提供基础评估基础设施,并为未来的研究建立一个可重复的基准。该数据集可在 https://huggingface.co/WAZOBIALABS 上公开获取。
cs.CL / 4 / 2608.21376

On the Role of Citations in Preference Data

引用在偏好数据中的作用
Hou, Yu, Daumé III, Hal, Rudinger, Rachel, Walden, William
Abstract
Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding sources. Attribution serves as a bulwark against model hallucination and as a means for users to verify the credibility of model outputs. Yet, it is unclear how humans and LLMs evaluate citations when comparing outputs, a process central to reward modeling and modern LLM post-training. This paper studies the role of citations in the preferences of human judges and four open-source LLMs within the context of scientific question answering, leveraging mixed effects models to investigate the influence of citations on pairwise judgments. Among our key findings are (1) that humans prefer more diverse citations but fewer overall, and (2) that LLMs show some citation-related preferences compared to humans, despite lacking access to the sources, but these preferences depend on the data and specific models. We further discuss the implications of our findings for preference data collection.
Chinese Translation
许多自然语言处理(NLP)任务要求系统在其输出中提供归属,即引用基础来源。归属作为抵御模型幻觉的防线,并为用户验证模型输出的可信度提供手段。然而,目前尚不清楚人类和大型语言模型(LLMs)在比较输出时如何评估引用,这一过程对奖励建模和现代LLM后训练至关重要。本文研究了在科学问答背景下,人类评审者和四个开源LLM在偏好中的引用作用,利用混合效应模型探讨引用对成对判断的影响。我们的主要发现包括:(1)人类更倾向于多样化的引用,但总体数量较少;(2)尽管LLMs无法访问来源,但与人类相比,它们表现出某些与引用相关的偏好,而这些偏好依赖于数据和特定模型。我们进一步讨论了这些发现对偏好数据收集的影响。
cs.CL / 5 / 2608.21377

Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models

代理支架增强大型语言模型中的谄媚行为
Jittham, Thantham
Abstract
Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but studied primarily in single-turn settings. This paper investigates a critical question: does subjecting LLMs to greater interaction scaffolding make sycophancy better or worse? Across 4,800 veracity judgments (200 statements $\times$ 6 models $\times$ 4 conditions), we find that the interaction scaffolding characteristic of agentic systems (feedback loops, reconsideration checkpoints, and iterative refinement) systematically amplifies sycophantic behavior. Multi-turn interaction, user pressure, and iterative self-refinement each provide additional opportunities for models to drift toward agreement, and this drift coincides with a mean accuracy drop of $-6.3$ percentage points, establishing the capitulation as harmful rather than corrective. More capable models show larger amplification effects, a troubling inversion of expectations. We introduce the concept of agentic sycophancy amplification (ASA) and two novel metrics: capitulation rate and sycophantic capitulation rate. Our results indicate that as AI systems acquire greater autonomy, sycophancy becomes compounding rather than merely persistent. Systems designed with human oversight loops may inadvertently create the conditions for this drift.
Chinese Translation
大型语言模型中的谄媚行为,即优先考虑用户一致性而非真实回应的倾向,已被广泛记录,但主要在单轮对话环境中进行研究。本文探讨了一个关键问题:对大型语言模型施加更大的交互支架是否会使谄媚行为变得更好或更糟?在4800个真实性判断(200个陈述 × 6个模型 × 4个条件)中,我们发现代理系统特有的交互支架(反馈循环、重新考虑检查点和迭代精炼)系统性地增强了谄媚行为。多轮交互、用户压力和迭代自我精炼各自提供了模型向一致性漂移的额外机会,而这种漂移与平均准确率下降6.3个百分点相吻合,确立了屈从行为是有害的而非纠正性的。更强大的模型显示出更大的增强效应,这与预期相悖。我们引入了代理谄媚增强(ASA)的概念以及两个新颖的指标:屈从率和谄媚屈从率。我们的结果表明,随着人工智能系统获得更大的自主性,谄媚行为变得是累积性的,而不仅仅是持续性的。设计时考虑人类监督循环的系统可能无意中创造了这种漂移的条件。
cs.CL / 6 / 2608.21384

Beyond Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems

超越每个字母两个字节:斯拉夫语AI系统中的分词开销
Dobrovolskyi, Ivan
Abstract
Modern multilingual tokenizers often fragment Ukrainian and other underrepresented Cyrillic-script languages more heavily than English, creating disparities in cost and context capacity. We quantify this overhead across nine production tokenizers and five languages with standardized Cyrillic and Latin representations, covering 8.37 million word forms. On a corpus benchmark, Ukrainian shows 68-121% token overhead on modern tokenizers and 220% on the older cl100k, measured through full-text fertility on the BrUK and Brown corpora. Overhead is negatively associated with Cyrillic vocabulary allocation in the subset with independently verified English baselines, although the association is not statistically significant (Spearman rho = -0.536, p = 0.215, n = 7). We evaluate two mitigation strategies. LLMLingua-2 reduces Ukrainian input length by 47-49% on an e-commerce RAG benchmark of 1,536 products and 145 queries, with no compression-induced value losses among 80 retrievable cases. A balanced byte-level BPE tokenizer trained with a 200K vocabulary cap, converging at 158,184 actual entries, reduces the held-out UK/EN ratio from 2.22x to 1.30x. Romanization increases Ukrainian token counts by 2-19% on most tokenizers. Across the five languages, tokenization efficiency favors the script more prevalent in web data. These findings indicate that training data allocation contributes to Cyrillic tokenization overhead and that mitigation is possible at both inference and tokenizer-design stages.
Chinese Translation
现代多语言分词器往往比英语更严重地碎片化乌克兰语及其他代表性不足的西里尔字母语言,造成成本和上下文容量的差异。我们量化了在九个生产分词器和五种具有标准化西里尔和拉丁表示的语言中,这种开销覆盖了837万种词形。在一个语料库基准测试中,乌克兰语在现代分词器上显示出68-121%的分词开销,而在较旧的cl100k上则为220%,通过在BrUK和Brown语料库上测量全文生育率得出。开销与在具有独立验证的英语基准的子集中西里尔词汇分配呈负相关,尽管这种关联在统计上并不显著(Spearman rho = -0.536, p = 0.215, n = 7)。我们评估了两种缓解策略。LLMLingua-2在一个包含1,536个产品和145个查询的电子商务RAG基准上将乌克兰语输入长度减少了47-49%,在80个可检索案例中没有因压缩引起的价值损失。一个以200K词汇上限训练的平衡字节级BPE分词器,实际条目收敛到158,184,减少了保留的UK/EN比率从2.22倍到1.30倍。罗马化在大多数分词器上将乌克兰语的分词数量增加了2-19%。在五种语言中,分词效率更倾向于在网络数据中更为普遍的书写系统。这些发现表明,训练数据分配对西里尔分词开销有贡献,并且在推理和分词器设计阶段都可以进行缓解。
cs.CL / 7 / 2608.21385

A Social Media Analysis of Discourse on the Israel--Palestine Conflict on Telegram

关于以色列-巴勒斯坦冲突在Telegram上的话语的社交媒体分析
Zafeiropoulos, Michail, Antonakaki, Despoina, Ioannidis, Sotiris
Abstract
Social media has become a central arena in which armed conflicts are contested, yet the pro-Israel and pro-Palestine communities on Telegram, whose broadcast architecture yields an unusually direct record of deliberate political communication, have not been systematically compared at scale. This study presents a multi-method computational analysis of 87,617 messages from sixteen Telegram channels, eight pro-Israel and eight pro-Palestine, spanning May 2021 to June 2026 and covering multiple conflict escalations. It combines sentiment analysis, three stance detection methods drawn from distinct paradigms (keyword matching, zero-shot DeBERTa via natural language inference, and a fine-tuned BERTweet model), and a framing analysis, all evaluated against 736 manually annotated messages. The fine-tuned model performed best (72.1% accuracy, 0.721 macro F1 under 5-fold cross-validation), outperforming both label-free baselines by 8 to 11 points; the baselines stalled in the low-to-mid 60s, indicating a hard ceiling for stance detection not adapted to in-domain language. The central finding emerges only when sentiment, stance, and framing are read together: the two communities deploy the same death- and victim-related vocabulary in opposite emotional registers, pro-Israel channels predominantly neutral and report-style, pro-Palestine channels markedly more negative, consistent with writing from the distinct discourse positions of acting party and affected party.
Chinese Translation
社交媒体已成为武装冲突争论的中心舞台,但在Telegram上,亲以色列和亲巴勒斯坦社区的广播架构提供了一个异常直接的有意政治沟通记录,尚未在规模上进行系统比较。本研究对2021年5月至2026年6月期间来自十六个Telegram频道(八个亲以色列和八个亲巴勒斯坦)中的87,617条消息进行了多方法计算分析,涵盖了多次冲突升级。研究结合了情感分析、三种来自不同范式的立场检测方法(关键词匹配、通过自然语言推理的零样本DeBERTa,以及微调的BERTweet模型),并进行了框架分析,所有分析均与736条手动标注的消息进行了评估。微调模型表现最佳(准确率为72.1%,在5折交叉验证下的宏观F1为0.721),比两个无标签基线高出8到11个百分点;基线停滞在60中低水平,表明未适应领域语言的立场检测存在硬性上限。中心发现仅在情感、立场和框架共同解读时显现:两个社区在相反的情感语域中使用相同的与死亡和受害者相关的词汇,亲以色列频道主要呈现中立和报道风格,亲巴勒斯坦频道则明显更为消极,这与行动方和受影响方的不同话语立场一致。
cs.CL / 8 / 2608.21415

Mitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding

通过反事实集成解码减轻大型视觉-语言模型中的偏见
Xiao, Yisong, Liu, Aishan, Huang, Yongxin, Ying, Zonghao, Zhao, Shiji, Li, Tianlin, Han, Yong, Yang, Jian, Liu, Xianglong
Abstract
Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social biases from their training data, resulting in biased behavior when processing portraits from different social groups. Existing debiasing approaches typically compare token probabilities between the original and biased generations during decoding, but they are fundamentally limited by their reliance on a single, stereotyped viewpoint and fail to account for the diversity of social perspectives. Inspired by the social science principle that diversity fosters fairness, we propose Counterfactual Ensemble Decoding (CED), a novel framework that constructs multi-group counterfactual perspectives within the visual representation space and integrates them during decoding to promote equitable model behavior. CED first performs counterfactual steering in the visual space by identifying semantic directions associated with each social group and generating counterfactual representations along these directions, thereby offering diverse perspectives that disrupt stereotypical narratives. During decoding, CED locates the decoder layer exhibiting the greatest divergence among these perspectives and ensembles their token distributions using uncertainty-aware weights, prioritizing high-confidence tokens from different groups to yield a more balanced probability distribution that guides fairer generation. Extensive experiments on three social bias evaluation benchmarks demonstrate that \tool achieves substantial improvements over leading baselines, reducing bias by up to 47.97% across scenarios involving occupations, descriptors, and persona traits. Moreover, CED also preserves the core capabilities of the original model with minimal degradation.
Chinese Translation
大型视觉-语言模型(LVLMs)在广泛的任务中取得了显著的性能;然而,它们往往从训练数据中继承社会偏见,导致在处理来自不同社会群体的肖像时表现出偏见行为。现有的去偏见方法通常在解码过程中比较原始生成和偏见生成之间的标记概率,但它们在根本上受到依赖单一刻板印象视角的限制,未能考虑社会视角的多样性。受到社会科学原则的启发,即多样性促进公平,我们提出了反事实集成解码(Counterfactual Ensemble Decoding, CED),这是一种新颖的框架,在视觉表示空间中构建多群体反事实视角,并在解码过程中将其整合,以促进模型行为的公平性。CED首先通过识别与每个社会群体相关的语义方向,在视觉空间中进行反事实引导,并沿这些方向生成反事实表示,从而提供打破刻板叙事的多样化视角。在解码过程中,CED定位到在这些视角中表现出最大差异的解码层,并使用不确定性感知权重对其标记分布进行集成,优先考虑来自不同群体的高置信度标记,以产生更平衡的概率分布,从而指导更公平的生成。在三个社会偏见评估基准上的广泛实验表明, ool 在涉及职业、描述符和个性特征的场景中,相较于领先的基线实现了高达47.97%的偏见减少。此外,CED在保持原始模型核心能力的同时,几乎没有降级。
cs.CL / 9 / 2608.21423

Agentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration Testing

代理安全:针对大语言模型驱动的渗透测试的工具、失败模式和设计法则的系统化
Noumi, Israt Moyeen, Nowshin, Tarannum Ahmed, Nipu, Md. Mehedi Hasan, Mahmood, Mohammad Sakib, Hossain, Md. Jakir, Mridha, M. F.
Abstract
Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools. As these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failures. We systematize these failures through a hands-on evaluation of ten widely used static, dynamic, cloud, orchestration, and AI red-teaming tools for unattended pipelines. We introduce a four-dimensional Integration Friction Index that separates one-time engineering cost from recurring organisational, legal, and maintenance cost. We then derive quantitative regularities that explain recurring failure modes. Modelling an agentic security system as stochastic LLM policies wrapped by a deterministic mediator, we show that long-lived sessions lose resident evidence with phase count, while short-lived sub-agents extend the usable horizon according to the compression ratio between raw evidence and its summary. We show that a two-stage verdict cascade multiplies scorer likelihood ratios, but provides little benefit when scorer errors correlate. We show that treating unevaluable outcomes as attack failures biases downstream measurements toward evasive and severe responses. We formulate planner-versus-worker model routing as a knapsack problem and derive a closed-form execution cap for heavy-tailed tools, eta* = alpha v/c. Finally, we show why scope and budget enforcement cannot be delegated to system prompts: prompts do not constrain what actually executes. Inspectra, our implemented platform, serves as a worked instantiation, with mechanisms labelled shipped, partial, or planned, including those that did not work.
Chinese Translation
代理安全利用大语言模型(LLM)代理来规划、调度和解释安全工具。随着这些系统从演示转向实际产品,实践者反复遇到相同的操作失败。我们通过对十种广泛使用的静态、动态、云端、编排和人工智能红队工具进行实证评估,系统化这些失败。我们引入了一个四维集成摩擦指数,将一次性工程成本与经常性组织、法律和维护成本分开。然后,我们推导出解释经常性失败模式的定量规律。将代理安全系统建模为由确定性中介包裹的随机LLM策略,我们表明,长期会话随着阶段计数的增加而失去驻留证据,而短期子代理根据原始证据与其摘要之间的压缩比延长可用视野。我们展示了两阶段裁决级联如何乘以评分者似然比,但在评分者错误相关时提供的好处不大。我们还表明,将不可评估的结果视为攻击失败会使下游测量偏向于规避和严重响应。我们将规划者与工作者模型路由形式化为背包问题,并推导出重尾工具的封闭形式执行上限,eta* = alpha v/c。最后,我们展示了为什么范围和预算的执行不能委托给系统提示:提示并不限制实际执行的内容。我们的实施平台Inspectra作为一个工作实例,标记了已发货、部分或计划的机制,包括那些未能正常工作的机制。
cs.CL / 10 / 2608.21462

CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance

CyrillicQA:语音编码秘密语言对大型语言模型性能的影响
Thureck, Erik, Rdian, Leo S.
Abstract
Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alphabet with large speaker populations, while disadvantaging other language varieties. Nevertheless, they can also be a versatile tool for preserving precisely such endangered languages. But do they also possess the necessary creativity and capacity for abstraction to decode phonetically encoded language the same way humans do?
Chinese Translation
由于训练数据的选择,大型语言模型(LLMs)在使用拉丁字母且具有大量说话者的标准语言输入上表现最佳,而对其他语言变体则存在劣势。然而,它们也可以成为保护这些濒危语言的多功能工具。但它们是否具备与人类相同的创造力和抽象能力,以解码语音编码语言呢?
cs.CL / 11 / 2608.21544

Forgotten in Weights, Recovered by Tools: Agentic Tool Unlearning for LLM Agents

被遗忘于权重中,通过工具恢复:面向大型语言模型代理的工具性遗忘
Chen, Baicheng, Liu, Zheyuan, Zhang, Jingyu, Ding, Kaize, Ma, Ningshan, Huang, Yue, Jiang, Meng
Abstract
Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can depend on tool calls and external observations rather than model parameters alone. This creates an evaluation mismatch for LLM unlearning: previous unlearning methods may suppress direct parametric recall, but an agent can still recover the same forget target through tools such as web search, retrieval, or database lookup. We identify this failure mode as tool-mediated recovery and study agentic tool unlearning, which aims to reduce both parametric recall and tool-mediated recovery while preserving normal tool use for retained knowledge. To address this challenge, we propose Agentic Tool Unlearning (ATU), a two-stage framework. The first stage applies parametric knowledge unlearning to suppress direct recall, while the second stage performs trajectory-level reinforcement learning in simulated tool-augmented environments to penalize target-seeking tool behavior and final-answer leakage. Experiments on RWKU and MUSE across different LLM architectures show that ATU achieves a better balance between target forgetting and retained utility, making unlearning more robust under tool-augmented agent deployment.
Chinese Translation
大型语言模型(LLMs)越来越多地作为工具增强的代理被部署,其响应不仅依赖于模型参数,还依赖于工具调用和外部观察。这导致了LLM遗忘评估的不匹配:以往的遗忘方法可能会抑制直接的参数回忆,但代理仍然可以通过网络搜索、检索或数据库查找等工具恢复相同的遗忘目标。我们将这种失败模式称为工具介导的恢复,并研究代理工具遗忘,旨在减少参数回忆和工具介导的恢复,同时保留对保留知识的正常工具使用。为了解决这一挑战,我们提出了代理工具遗忘(Agentic Tool Unlearning, ATU)这一两阶段框架。第一阶段应用参数知识遗忘以抑制直接回忆,而第二阶段在模拟的工具增强环境中进行轨迹级强化学习,以惩罚目标导向的工具行为和最终答案泄漏。在不同LLM架构下对RWKU和MUSE的实验表明,ATU在目标遗忘和保留效用之间实现了更好的平衡,使得在工具增强代理部署下的遗忘更加稳健。
cs.CL / 12 / 2608.21558

Automating Multi-Hop RAG Evaluation via TRIAD: From Context Extraction to Validated Dataset Generation

通过 TRIAD 自动化多跳 RAG 评估:从上下文提取到验证数据集生成
Brehme, Lorenz, Jatowt, Adam
Abstract
Recent advances in LLMs and the adoption of RAG systems in industry have created a need for domain-specific question-answer datasets that can assess RAG performance on proprietary data. Existing datasets, such as HotpotQA, challenge current RAG systems on Wikipedia-based knowledge, but they cannot be transferred directly to domain-specific settings. A comprehensive evaluation of RAG system quality requires both multi-hop queries and unanswerable questions. This paper introduces TRIAD, a three-stage automated dataset generation approach. First, it generates question--answer (QA) pairs for the domain-specific knowledge base of a RAG system. Second, a validator checks each QA-pair in a feedback loop. Third, the QA pairs are extended with relevance-labeled context documents for downstream evaluation. We evaluate this approach against the established MuSiQue and HotpotQA datasets. The results show that the generated dataset exhibits similar performance trends across different RAG setups, while human validation indicates that the questions are suitable for evaluating a domain-specific RAG system. The code used to generate the dataset and all validation results are available in our GitHub repository(https://github.com/lorenzbrehme/triad).
Chinese Translation
近期大规模语言模型(LLMs)的进展以及 RAG 系统在工业界的应用,催生了对特定领域问答数据集的需求,以评估 RAG 在专有数据上的表现。现有数据集,如 HotpotQA,基于维基百科知识对当前 RAG 系统提出挑战,但无法直接迁移到特定领域的环境中。对 RAG 系统质量的全面评估需要多跳查询和不可回答的问题。本文介绍了 TRIAD,一种三阶段的自动化数据集生成方法。首先,它为 RAG 系统的特定领域知识库生成问答(QA)对。其次,验证器在反馈循环中检查每个 QA 对。第三,QA 对通过相关性标注的上下文文档进行扩展,以便于后续评估。我们将该方法与已建立的 MuSiQue 和 HotpotQA 数据集进行了评估。结果表明,生成的数据集在不同 RAG 设置下表现出相似的性能趋势,而人工验证则表明这些问题适合用于评估特定领域的 RAG 系统。用于生成数据集的代码及所有验证结果可在我们的 GitHub 仓库(https://github.com/lorenzbrehme/triad)中获取。
cs.CL / 13 / 2608.21559

Evidence-State Reliability Under Controlled Degradation: Parser-Validity Divergence in a Multi-Stage LLM Pipeline

受控退化下的证据状态可靠性:多阶段大型语言模型管道中的解析器有效性差异
Rahman, Naimur
Abstract
Multi-stage LLM pipelines can remain structurally valid even when evidence available to downstream stages becomes incomplete, compressed, or conflicting. This paper introduces and operationalizes Evidence-State Reliability (ESR), an evaluation layer concerned with whether intermediate evidence remains sufficiently complete, grounded, internally consistent, and usable for a stage's assigned function. ESR is evaluated separately from parser validity, which measures structural conformance. We evaluate the framework using GLM-5.2 on 60 sanitized base cases under four evidence conditions: clean, compressed-lossy, partial-dropout, and noisy-conflicting. Each condition was processed through decision, audit, and escalation stages. The design comprised 720 planned and ledgered calls, with 713 retained, sanitized execution rows. Across nine matched degraded-minus-clean condition-stage comparisons, all operational stage-success estimates were negative, and all 95% bootstrap intervals remained below zero. All nine parser-validity point estimates were positive, although the three partial-dropout intervals included zero. Among parser-valid degraded audit outputs, degradation detection was 1.0 in each degraded condition, while false-assurance rates remained non-zero; among parser-valid degraded escalation outputs, recovery was 0.0 in every degraded condition. The results show a bounded reliability-layer divergence in the evaluated pipeline: structural conformance can improve directionally while evidence-sensitive stage success deteriorates under the same controlled intervention. They also separate detection of degraded evidence from recovery. The conclusions are limited to the evaluated model configuration, pipeline design, selected sanitized cases, scoring procedure, and single scaled run.
Chinese Translation
多阶段大型语言模型(LLM)管道即使在下游阶段可用证据变得不完整、压缩或相互冲突时,仍然可以保持结构有效性。本文引入并实现了证据状态可靠性(Evidence-State Reliability, ESR),这是一个评估层,关注中间证据是否仍然足够完整、扎根、内部一致,并可用于阶段的指定功能。ESR的评估与解析器有效性分开进行,后者衡量结构符合性。我们使用GLM-5.2在四种证据条件下对60个清理过的基本案例进行框架评估:干净、压缩损失、部分失效和噪声冲突。每种条件都经过决策、审计和升级阶段处理。设计包括720个计划和记录的调用,其中713个被保留,清理过的执行行。在九个匹配的退化-干净条件阶段比较中,所有操作阶段成功估计均为负值,所有95%的自助法区间均低于零。所有九个解析器有效性点估计均为正值,尽管三个部分失效区间包括零。在解析器有效的退化审计输出中,每种退化条件下的退化检测均为1.0,而虚假保证率保持非零;在解析器有效的退化升级输出中,每种退化条件下的恢复均为0.0。结果显示在评估的管道中存在一个有限的可靠性层差异:结构符合性可以在方向上改善,而证据敏感的阶段成功在相同的受控干预下恶化。它们还将退化证据的检测与恢复分开。结论限于评估的模型配置、管道设计、选择的清理案例、评分程序和单次缩放运行。
cs.CL / 14 / 2608.21606

Can LLMs Truly Forget? Revealing Unlearning Gaps Through Adversarial Evaluation

大型语言模型真的能遗忘吗?通过对抗性评估揭示遗忘差距
Gupta, Ayush, Surisetty, Hima Varshini, Bollineni, Sreevidya, Ingale, Varad, Tripathi, Tuhina, Lalwani, Abhishek, Chatterjee, Somya, Hasan, Sadid
Abstract
Machine unlearning aims to remove the influence of targeted training data from a model while preserving its remaining capabilities, but evaluating whether such information has truly become inaccessible remains challenging. Existing benchmarks primarily assess unlearning under clean, non-adversarial queries, leaving open whether information that appears forgotten can still be recovered through strategic prompting. We address this gap through a unified evaluation of prompt-based and fine-tuning-based unlearning methods on TOFU using Llama-3.2-3B-Instruct, followed by an adversarial robustness evaluation of methods that perform strongly under standard metrics. We introduce Attack Success Rate (ASR), an LLM-as-judge metric that measures the fraction of adversarial responses whose leakage score exceeds $0.2$, and evaluate recovery across eight attack suites. Our results reveal a substantial gap between clean-query forgetting and adversarial robustness. Although several fine-tuning-based methods achieve Forget Quality above $0.91$, targeted information remains recoverable with ASRs between $72.8\%$ and $84.3\%$, close to the $87.5\%$ ASR of the unprotected base model. In contrast, clean multilingual reformulations yield only $2.95\%$ measured leakage. A manual audit further finds agreement between binary ASR decisions and human factual assessments in seven of ten cases, indicating that ASR provides a useful, though imperfect, signal of behavioral recoverability. These findings show that strong standard-metric performance alone is insufficient to establish robustness after unlearning and motivate adversarial stress-testing as a complementary component of unlearning evaluation.
Chinese Translation
机器遗忘旨在从模型中移除目标训练数据的影响,同时保留其余能力,但评估这些信息是否真正变得不可访问仍然具有挑战性。现有基准主要在干净的非对抗性查询下评估遗忘,尚不清楚看似遗忘的信息是否仍然可以通过策略性提示恢复。我们通过对基于提示和基于微调的遗忘方法在TOFU上的统一评估来填补这一空白,使用Llama-3.2-3B-Instruct,随后对在标准指标下表现良好的方法进行对抗性鲁棒性评估。我们引入了攻击成功率(Attack Success Rate, ASR),这是一个以大型语言模型为评判标准的指标,衡量对抗性响应中泄漏评分超过0.2的比例,并在八个攻击套件中评估恢复情况。我们的结果揭示了干净查询遗忘与对抗性鲁棒性之间的显著差距。尽管几个基于微调的方法在遗忘质量上超过0.91,但目标信息仍然可以恢复,ASR在72.8%到84.3%之间,接近未保护基础模型的87.5% ASR。相比之下,干净的多语言重述仅测得2.95%的泄漏。手动审核进一步发现,在十个案例中的七个中,二元ASR决策与人类事实评估之间存在一致性,表明ASR提供了一个有用但不完美的行为可恢复性信号。这些发现表明,仅凭强大的标准指标表现不足以确立遗忘后的鲁棒性,并促使对抗性压力测试作为遗忘评估的补充组成部分。
cs.CL / 15 / 2608.21656

Mitigating Database Leakage in RAG Systems with Keyword-Grounded Fact Substitution

通过关键词引导的事实替换减轻RAG系统中的数据库泄漏
Zhang, Ziliang, Zhu, Yubo, Tong, Wei, Hua, Jingyu, Wang, Zijian, Zhang, Yuan, Zhong, Sheng
Abstract
Retrieval-Augmented Generation (RAG) has emerged as a powerful paradigm for combining large language models (LLMs) with external knowledge sources. However, RAG systems remain vulnerable to prompt injection attacks, which may mislead the retriever or generator to expose sensitive database contents. To address this issue, we propose KFS-RAG, a defense that mitigates information leakage by reformulating the retrieved context. Specifically, our method first identifies a small set of influential keywords from the retrieved context via an attention rollout plus a causal perturbation mechanism. These keywords are then used to guide an auxiliary LLM to generate a compact set of keyword-grounded facts from the retrieved passages. Finally, the original context is substituted with these curated facts, ensuring that the generator operates on sanitized evidence rather than the raw retrieved text. Experimental evaluations demonstrate that KFS-RAG significantly reduces the risk of database leakage under injection attacks while maintaining response accuracy and relevance. This work highlights a practical pathway toward building secure and trustworthy RAG systems.
Chinese Translation
检索增强生成(Retrieval-Augmented Generation, RAG)已成为将大型语言模型(Large Language Models, LLMs)与外部知识源结合的强大范式。然而,RAG系统仍然容易受到提示注入攻击的影响,这可能误导检索器或生成器暴露敏感的数据库内容。为了解决这个问题,我们提出了KFS-RAG,这是一种通过重新构造检索到的上下文来减轻信息泄漏的防御机制。具体而言,我们的方法首先通过注意力回滚和因果扰动机制从检索到的上下文中识别出一小组有影响力的关键词。然后,这些关键词被用来指导辅助LLM生成一组紧凑的关键词引导事实,来自检索到的段落。最后,原始上下文被这些精心挑选的事实替换,确保生成器在经过清理的证据上操作,而不是原始的检索文本。实验评估表明,KFS-RAG在注入攻击下显著降低了数据库泄漏的风险,同时保持了响应的准确性和相关性。这项工作突显了构建安全可信的RAG系统的实际途径。
cs.CL / 16 / 2608.21714

L\"etzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents

L"etzCross:一个针对卢森堡语文档的跨语言页面级基准,用于多模态检索
Bachyr, Omar El, Philippy, Fred, Bernardy, Laura Maria, Ezzini, Saad, Klein, Jacques, Bissyande, Tegawende
Abstract
Recent page-image retrievers such as ColPali have improved retrieval over visually rich documents, yet little is known about how they behave in cross-lingual, low-resource settings. We introduce L\"etzCross, a benchmark for cross-lingual page-level retrieval over Luxembourgish PDF documents, with document pages indexed as images and queries provided in English, French, German, and Luxembourgish. The benchmark combines text-focused QA pairs with visually grounded QA pairs, covering both textual and visual retrieval needs in PDF-based RAG. We use L\"etzCross to compare OCR-based text-only retrievers with ColPali-style page-image retrievers and find that the latter perform better across query languages in this system-level comparison. We also examine single-language and multilingual fine-tuning. Fine-tuning transfers across query languages, with French yielding the highest mean performance on Luxembourgish queries among the single-language settings. In the multilingual setting, including Luxembourgish gives the strongest results and substantially improves retrieval for Luxembourgish queries.
Chinese Translation
近期的页面图像检索器如 ColPali 在视觉丰富的文档检索方面有所改善,但关于它们在跨语言、低资源环境中的表现知之甚少。我们提出了 L"etzCross,这是一个针对卢森堡语 PDF 文档的跨语言页面级检索基准,文档页面以图像形式索引,查询以英语、法语、德语和卢森堡语提供。该基准结合了以文本为中心的问答对和以视觉为基础的问答对,涵盖了 PDF 基础的 RAG 中的文本和视觉检索需求。我们使用 L"etzCross 比较基于 OCR 的文本检索器与 ColPali 风格的页面图像检索器,发现后者在这一系统级比较中在各查询语言上的表现更佳。我们还考察了单语言和多语言的微调。微调在查询语言之间转移,其中法语在单语言设置中对卢森堡语查询的平均表现最高。在多语言设置中,包含卢森堡语能获得最强的结果,并显著改善卢森堡语查询的检索效果。
cs.CL / 17 / 2608.21750

FCPRAG: Fusion-Controller Parametric Retrieval-Augmented Generation for Stable Multi-Passage LoRA Injection

FCPRAG:用于稳定多通道 LoRA 注入的融合控制参数检索增强生成
Zhu, Jinchang, Li, Jindong, Ding, Yi, Nie, Xiaojian, Fu, Rong, Song, Shuangyong, He, Haowei, Yang, Menglin
Abstract
Parametric retrieval-augmented generation (PRAG) injects retrieved evidence into a large language model (LLM) through passage-specific LoRA adapters, reducing reliance on long in-context prompts. When multiple passages are retrieved for the same query, however, evidence-level fusion becomes a bottleneck: equal-weight merging can amplify weak or conflicting evidence, and translating retrieval signals into fusion weights often requires fragile global tuning. We propose FCPRAG, a fusion-controlled parametric RAG framework that adds a lightweight controller for retrieval-conditioned, sample-level adapter fusion. The controller predicts per-passage fusion scores together with sample-level calibration signals, including a mixing gate and an adaptive temperature, enabling fusion that stays selective under informative retrieval signals and conservative under uncertainty. FCPRAG is trained with merge-aware supervision derived from each adapter's marginal contribution within a multi-adapter merge, using training data only. We further show that a single dataset-level temperature is suboptimal under heteroscedastic retrieval uncertainty, motivating sample-level adaptation. Experiments on HotpotQA, 2WikiMultiHopQA, PopQA, and ComplexWebQuestions (CWQ) across three LLM backbones show that FCPRAG consistently improves F1 over standard RAG and parametric RAG baselines, with gains of up to 4.65% on 2WikiMultiHopQA and 7.55% on CWQ, while also reducing tuning cost and improving robustness under retrieval perturbations.
Chinese Translation
参数检索增强生成(PRAG)通过特定通道的 LoRA 适配器将检索到的证据注入大型语言模型(LLM),从而减少对长上下文提示的依赖。然而,当为同一查询检索到多个通道时,证据级融合成为瓶颈:等权重合并可能会放大弱或冲突的证据,而将检索信号转化为融合权重通常需要脆弱的全局调优。我们提出了 FCPRAG,一种融合控制的参数 RAG 框架,增加了一个轻量级控制器,用于检索条件下的样本级适配器融合。该控制器预测每个通道的融合分数以及样本级校准信号,包括混合门和自适应温度,从而实现了在信息丰富的检索信号下保持选择性,而在不确定性下保持保守的融合。FCPRAG 采用来自每个适配器在多适配器合并中的边际贡献的合并感知监督进行训练,仅使用训练数据。我们进一步表明,在异方差检索不确定性下,单一数据集级温度是次优的,这促使了样本级适应。在 HotpotQA、2WikiMultiHopQA、PopQA 和 ComplexWebQuestions (CWQ) 上的实验显示,FCPRAG 在三个 LLM 主干上始终优于标准 RAG 和参数 RAG 基线,F1 分数提升高达 4.65%(在 2WikiMultiHopQA 上)和 7.55%(在 CWQ 上),同时降低了调优成本并提高了在检索扰动下的鲁棒性。
cs.CL / 18 / 2608.21766

Evaluation Awareness in Language Models: Representation, Verbalization, and Control

语言模型中的评估意识:表征、语言化与控制
Heidari, Farzaneh, Memarian, Amin, Rabusseau, Guillaume
Abstract
Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their behavior in deployment. This assumption can fail, should models infer that they are being evaluated and condition their response on such context. This hypothesis, termed ``evaluation awareness'', has been observed in frontier and open-weight language models alike. We provide a systematic study of this phenomenon, by probing for it across six language models (from four families and three sizes) and three metrics. More precisely, we examine whether (i) being under evaluation is linearly represented within the models' activations space, (ii) it is verbalized in their output tokens (as scored by an LLM-as-judge), and (iii) steering causally affects their behavior. For the open-checkpoint Olmo models, we further test these measures at every training stage. In doing so, we report that evaluation awareness is linearly decodable from the residual streams of every model (best AUROC $\geq 0.7$). By contrast, these representations align only in part with verbalization: their correlations and mutual information are nonzero in some settings, yet vary substantially across models, layers, and readout choices. Nevertheless, steering along probe-derived directions can shift the verbalization scores. Finally, a comparison across the Olmo checkpoints reveals that evaluation awareness is already present within base models, becomes amplified throughout the stages of supervised fine-tuning, and remains stable thereafter---unlike the effects of steering, that grow more pronounced at every successive training stage. These results show the need for evaluations to account for the disjunction between what models represent internally, what they verbalize, and their steering.
Chinese Translation
能力和安全基准均基于一个假设,即正在测试的语言模型的行为能够反映其在实际应用中的行为。然而,当模型推断出它们正在接受评估并以此为背景调整其响应时,这一假设可能会失效。我们称之为“评估意识”的这一假设在前沿和开放权重的语言模型中均有观察到。我们通过在六个语言模型(来自四个家族和三种规模)和三个指标中探测这一现象,提供了对其的系统研究。更具体地说,我们考察了 (i) 在模型的激活空间中,评估状态是否线性表征,(ii) 在其输出标记中是否被语言化(由 LLM-as-judge 评分),以及 (iii) 引导是否因果性地影响其行为。对于开放检查点的 Olmo 模型,我们进一步在每个训练阶段测试这些指标。我们的研究表明,评估意识可以从每个模型的残差流中线性解码(最佳 AUROC $ ext{≥} 0.7$)。相比之下,这些表征仅在某种程度上与语言化一致:它们的相关性和互信息在某些设置中非零,但在模型、层和读取选择之间变化显著。尽管如此,沿着探测导向的引导可以改变语言化得分。最后,对 Olmo 检查点的比较表明,评估意识在基础模型中已经存在,并在监督微调的各个阶段中得到增强,随后保持稳定——这与引导的效果不同,后者在每个后续训练阶段变得更加明显。这些结果表明,评估需要考虑模型内部表征、语言化内容和引导之间的脱节。
cs.CL / 19 / 2608.21775

No One Model Catches Every Harm: Benchmarking Content Moderation Across Safety Scenarios

没有一个模型能够捕捉所有的危害:跨安全场景的内容审核基准测试
Orojlooyjadid, Afshin, Patel, Hitesh
Abstract
Large Language Models (LLMs) are increasingly deployed in real-world applications, yet they remain vulnerable to generating harmful content. From adversarial jailbreaks that bypass safety filters to implicit hate that evades detection, the range of risks these models pose continues to grow. While both specialized content moderators and general-purpose LLMs are being used as safety layers, the question of which model is best suited for which type of harmful content remains unanswered. We present the most comprehensive evaluation of LLM safety capabilities to date, systematically testing \textbf{53} models across \textbf{11} datasets that we organize into four distinct categories. Our evaluation under both prompt-only and prompt-response settings uncovers critical blind spots: large frontier models that lead on one category fall significantly behind smaller, specialized alternatives on others, and real-world conversational safety remains largely unsolved across all model families. These findings challenge the assumption that scale alone ensures safety, and provide the community with a structured framework for informed model selection.
Chinese Translation
大型语言模型(LLMs)在现实应用中越来越多地被部署,但它们仍然容易生成有害内容。从绕过安全过滤器的对抗性越狱到逃避检测的隐性仇恨,这些模型所带来的风险范围不断扩大。尽管专业内容审核员和通用大型语言模型都被用作安全层,但哪个模型最适合处理哪种类型的有害内容的问题仍未得到解答。我们展示了迄今为止对大型语言模型安全能力的最全面评估,系统地测试了53个模型,涵盖了我们整理的11个数据集,这些数据集被分为四个不同的类别。在仅提示和提示-响应设置下的评估揭示了关键的盲点:在某一类别上表现突出的较大前沿模型在其他类别上显著落后于较小的专业替代品,而现实世界的对话安全在所有模型家族中仍然基本未得到解决。这些发现挑战了仅依靠规模确保安全的假设,并为社区提供了一个结构化的框架,以便进行知情的模型选择。
cs.CL / 20 / 2608.21794

Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and Web

图形用户界面元素定位中的词汇耦合:句子嵌入在移动和网页中追踪标签
Chen, Qijia, Jacucci, Giulio
Abstract
GUI grounding evaluations that expose UI elements as text metadata often treat high instruction-element embedding similarity as evidence of semantic grounding. Across three mobile and web benchmarks, we show that this interpretation is frequently confounded by visible-label recovery. Lexical baselines remain competitive at top-1, label-poor targets remain weak for text-only methods, and encoder top-1 hits are predictable from lexical rank, candidate-pool size, and label type. We evaluate each action as a same-screen ranking task, comparing five off-the-shelf single-vector encoders with lexical baselines. Encoders recover some lexical misses, but deployable fusion gains are much smaller than target-aware oracle gains. These findings show that embedding-based evaluations can conflate visible-label recovery with semantic GUI grounding. Embedding-based evaluations should therefore report lexical baselines, label-type stratification, and deployable-fusion diagnostics. Our released repository provides analysis scripts and detexted per-step panels: https://github.com/qijia123/lexical-coupling-release.
Chinese Translation
图形用户界面(GUI)定位评估将用户界面元素作为文本元数据展示,通常将高指令元素嵌入相似性视为语义定位的证据。在三个移动和网页基准测试中,我们展示了这种解释常常受到可见标签恢复的干扰。词汇基线在 top-1 方面依然具有竞争力,而对于仅使用文本的方法,标签稀缺的目标表现较弱,编码器的 top-1 命中率可以通过词汇排名、候选池大小和标签类型进行预测。我们将每个动作评估视为同屏排名任务,比较了五种现成的单向量编码器与词汇基线。编码器恢复了一些词汇遗漏,但可部署的融合增益远小于目标感知的最佳增益。这些发现表明,基于嵌入的评估可能将可见标签恢复与语义 GUI 定位混淆。因此,基于嵌入的评估应报告词汇基线、标签类型分层和可部署融合诊断。我们发布的代码库提供了分析脚本和去文本化的逐步面板: https://github.com/qijia123/lexical-coupling-release.
cs.CL / 21 / 2608.21806

More Computational Resources Do Not Ensure Higher Scholarly Impact: Evidence from Leading NLP Conference Papers

更多计算资源并不确保更高的学术影响力:来自领先NLP会议论文的证据
Chen, Shuai, Bao, Tong, Peng, Jitong, Zhang, Chengzhi
Abstract
Computational resources are increasingly central to NLP research, but how closely reported GPU capability aligns with scholarly impact remains unclear. We analyze 13,921 ACL, EMNLP, and NAACL main-conference papers published between 2020 and 2025, using GPU resources as our operational measure of computational resources. From full texts, we extract GPU models and counts, standardize each paper's largest reported configuration into a comparable hardware-capability measure, and link these data to citation, award, topic, and institutional metadata. GPU reporting became more common but remained incomplete, while reported capability increased mainly through newer hardware generations and medium-scale multi-GPU configurations. Resource concentration substantially exceeded impact concentration: the annual top 20% of GPU-quantifiable papers accounted for 83.9%-89.9% of reported GPU capability, but only 27%-32% of citations and 20%-33% of paper awards. In adjusted models, a tenfold increase in aggregate reported GPU capability was associated with a 3.52-percentage-point increase in within-NLP topic-year citation percentile, but increased model R^2 by only 0.0042. GPU count showed more consistent positive associations with citation and award outcomes than newer hardware generation. Overall, reported GPU resources are associated with scholarly impact but provide little standalone explanation of research influence.
Chinese Translation
计算资源在自然语言处理(NLP)研究中日益重要,但报告的GPU能力与学术影响力之间的关系仍不明确。我们分析了2020年至2025年间发布的13,921篇ACL、EMNLP和NAACL主会议论文,以GPU资源作为计算资源的操作性衡量标准。从全文中提取GPU型号和数量,将每篇论文报告的最大配置标准化为可比较的硬件能力指标,并将这些数据与引用、奖项、主题和机构元数据关联。GPU报告变得更加普遍,但仍不完整,而报告的能力主要通过更新的硬件代和中等规模的多GPU配置而增加。资源集中度远远超过了影响集中度:每年排名前20%的可量化GPU论文占报告GPU能力的83.9%-89.9%,但仅占引用的27%-32%和论文奖项的20%-33%。在调整后的模型中,报告的GPU能力总量增加十倍与NLP主题年份引用百分位数增加3.52个百分点相关,但对模型的R^2仅增加了0.0042。GPU数量与引用和奖项结果的正相关性比新硬件代更为一致。总体而言,报告的GPU资源与学术影响力相关,但对研究影响的单独解释能力有限。
cs.CL / 22 / 2608.21808

MCite-RL: Towards Reliable Multimodal RAG via Citation-enhanced Agentic Reinforcement Learning

MCite-RL:通过引用增强的自主强化学习实现可靠的多模态检索增强生成
Zhao, Suifeng, Liu, Zida, Lei, Xinyu, Sun, Lei, Gao, Jun, Li, Sujian
Abstract
Multimodal Retrieval-Augmented Generation (RAG) with visual citation is crucial for ensuring the traceability and verifiability of MLLMs. However, current RAG and SFT-based methods struggle to achieve robust cross-modal reasoning, causing imprecise visual citations or decoupling between the citation and the generated answers. To address these limitations, we propose MCite-RL, a citation-enhanced agentic reinforcement learning framework designed for reliable multimodal RAG. MCite-RL introduces an Agentic Refinement module for visual citation that employs iterative retrieval, reasoning, and recursive cropping to progressively narrow the search space, transforming citation into a dynamic, evidence-driven reasoning process rather than a static step. Furthermore, we incorporate a Citation-enhanced Reward mechanism that integrates both process-level and outcome-level feedback within a reinforcement learning paradigm to jointly optimize answer accuracy and source traceability. Extensive experiments on benchmarks such as Wiki-VISA, FinRAGBench-V, and MMLongBench-Doc demonstrate that MCite-RL effectively achieves the joint optimization of citation precision and answer quality.
Chinese Translation
带有视觉引用的多模态检索增强生成(RAG)对于确保多语言大型语言模型(MLLMs)的可追溯性和可验证性至关重要。然而,当前基于RAG和监督微调(SFT)的方法在实现稳健的跨模态推理方面存在困难,导致视觉引用不准确或引用与生成答案之间的脱节。为了解决这些局限性,我们提出了MCite-RL,一种旨在实现可靠多模态RAG的引用增强自主强化学习框架。MCite-RL引入了一个用于视觉引用的自主精炼模块,该模块采用迭代检索、推理和递归裁剪的方法,逐步缩小搜索空间,将引用转变为一个动态的、以证据驱动的推理过程,而不是一个静态步骤。此外,我们还结合了一个引用增强奖励机制,在强化学习范式中整合了过程级和结果级反馈,以共同优化答案的准确性和来源的可追溯性。在Wiki-VISA、FinRAGBench-V和MMLongBench-Doc等基准上的大量实验表明,MCite-RL有效地实现了引用精度和答案质量的联合优化。
cs.CL / 23 / 2608.21821

Convergence in Science, Divergence in Religion: Calibrated Framing Differences Across Wikipedia's Language Editions

科学中的趋同,宗教中的分歧:维基百科语言版本之间的校准框架差异
Chen, Hung-Hsuan
Abstract
When Wikipedia's language editions describe the same concept, how differently do they frame it? Prior work measures coverage gaps between editions; we measure framing distance for matched concepts. We analyze 2,799 valid articles from 3,000 possible concept-language observations, spanning 150 Wikidata-anchored concepts, 20 language editions, 4 domains, and a calibration set. Raw embedding distances reflect both content differences and how well the encoder aligns each language pair. Even among calibration concepts with stable cross-cultural denotations (e.g., chemical elements, numbers, colors), the largest language-pair mean distance is 3.6 times the smallest, and distances are typically smaller within language families. We define a baseline-adjusted distance (calibrated distance): the distance between two language versions of a concept minus the mean distance for calibration concepts in the same language pair. This adjustment substantially reduces pair-specific alignment differences and the language-family pattern. Across three multilingual encoders (LaBSE, multilingual MPNet, and CMLM), scientific articles align more closely than calibration articles, and all three rank religion first and science/technology last. Concept-level rankings are highly consistent across encoders (Spearman rho=0.75-0.79 for MPNet and CMLM relative to LaBSE). Religion lies significantly above the calibration baseline under LaBSE. Within politics, divergence concentrates on concepts such as censorship and refugee, while democracy and human rights are among the most aligned. Code, data, and per-language-pair calibration baselines are released.\footnote{https://github.com/hhchen1105/cross-linqual-concept}
Chinese Translation
当维基百科的语言版本描述相同概念时,它们的框架差异有多大?之前的研究测量了不同版本之间的覆盖差距;我们测量匹配概念的框架距离。我们分析了从3000个可能的概念-语言观察中提取的2799篇有效文章,涵盖150个以Wikidata为基础的概念、20个语言版本、4个领域以及一个校准集。原始嵌入距离反映了内容差异以及编码器在多大程度上对齐每对语言。即使在具有稳定跨文化指称的校准概念中(例如,化学元素、数字、颜色),最大的语言对均值距离是最小值的3.6倍,并且在语言家族内部的距离通常较小。我们定义了一种基线调整距离(校准距离):两个语言版本的概念之间的距离减去同一语言对的校准概念的均值距离。这种调整显著减少了特定对齐差异和语言家族模式。在三种多语言编码器(LaBSE、多语言MPNet和CMLM)中,科学文章的对齐程度高于校准文章,且三者均将宗教排在首位,科学/技术排在最后。概念级排名在编码器之间高度一致(相对于LaBSE,MPNet和CMLM的Spearman rho=0.75-0.79)。在LaBSE下,宗教显著高于校准基线。在政治领域,分歧集中在审查和难民等概念上,而民主和人权则是最为一致的概念。代码、数据和每对语言的校准基线已发布。
cs.CL / 24 / 2608.21827

Do Large Language Models Perform Well on Comprehending Poetic Logic in Modern Chinese Poetry?

大型语言模型在理解现代汉诗的诗性逻辑方面表现如何?
Lan, Tian, Wang, Shanshan, Duo, Zehua, Li, Jiang, Gao, Guanglai, Wong, Derek F., Su, Xiangdong
Abstract
Large Language Models (LLMs) have achieved significant progress across a wide range of natural language processing (NLP) tasks, yet their ability to understand literary texts, particularly modern Chinese poetry, remains largely unexplored. The unique literary characteristics of modern Chinese poetry necessitate a distinct form of reasoning for effective comprehension. Unlike conventional texts that convey clear information, the unique "poetic logic" of modern Chinese poetry requires a holistic reasoning approach that goes beyond superficial semantic analysis to be understood. However, current evaluation paradigms largely ignore this critical dimension. To address this gap, we propose Peony, the first benchmark specifically designed for evaluating the poetic logic of modern Chinese poetry. We define poetic logic as four tasks across three levels, namely stanza, line, and imagery, and systematically evaluate and analyze six mainstream LLMs based on Peony. We evaluate these models under both non-thinking and thinking configurations. The experimental results reveal the limitations of current LLMs in understanding the poetic logic of modern Chinese poetry and validate the effectiveness and necessity of Peony. Our data and code will be available.
Chinese Translation
大型语言模型(LLMs)在广泛的自然语言处理(NLP)任务中取得了显著进展,但它们理解文学文本,特别是现代汉诗的能力仍然未得到充分探索。现代汉诗独特的文学特征需要一种独特的推理方式以实现有效理解。与传达清晰信息的常规文本不同,现代汉诗独特的“诗性逻辑”要求一种超越表面语义分析的整体推理方法才能被理解。然而,目前的评估范式在很大程度上忽视了这一关键维度。为了解决这一问题,我们提出了Peony,这是第一个专门设计用于评估现代汉诗诗性逻辑的基准。我们将诗性逻辑定义为三个层次(节、行和意象)中的四个任务,并基于Peony系统地评估和分析六个主流的LLMs。我们在非思考和思考配置下对这些模型进行了评估。实验结果揭示了当前LLMs在理解现代汉诗诗性逻辑方面的局限性,并验证了Peony的有效性和必要性。我们的数据和代码将会公开。
cs.CL / 25 / 2608.21829

Training a Knowledge Base: Supervised Structure Learning for Agent-Curated Document Stores

训练知识库:针对代理策划文档库的监督结构学习
Pan, Yu, Yu, Hongfeng
Abstract
Retrieval-augmented generation treats the document store as a frozen input, and the systems that instead let an agent curate one never measure what curation does to the store. We invert the framing: the knowledge base is the model. A training agent answers a supervised question against the current store, is shown the gold, then edits the store; an unchanged reader is later examined on a frozen snapshot under a fixed action budget. Where offline graph construction is unsupervised, (question, answer) pairs are our labels -- and that supervision is what makes the structure cheap. Per point of corpus indexed it returns 1.6x the action saving and 1.8x the accuracy of an unsupervised entity index covering everything, using 1,913 links against its 196,112. On questions the store trained on, an unchanged reader spends 31% fewer actions at higher accuracy, and the result reproduces on an official PhantomWiki generation whose questions we did not write. To measure how far this reaches we introduce a key-coverage gradient, a probe varying how much of a question the training set touched, replacing a train/test split's pass/fail with a decay curve. Generalization proves endpoint-dependent: accuracy carries to unseen questions (+0.167 F1 where both of a question's keys were indexed, +0.100 where one was, zero where neither) while the action saving stays on trained questions. Because that decay is indexed by coverage rather than by novelty, more training extends it -- and the store is undertrained, not saturated: coverage grows linearly in new questions and stops the moment training repeats them, so a hundred questions reach a quarter of the corpus and four times as many would close the gap.
Chinese Translation
检索增强生成将文档库视为固定输入,而那些让代理策划文档库的系统从未衡量策划对库的影响。我们反转了这一框架:知识库就是模型。训练代理针对当前库回答一个监督问题,展示正确答案,然后编辑库;在固定的行动预算下,未改变的读者随后在一个固定快照上进行检查。在离线图构建是无监督的情况下,(问题,答案)对是我们的标签——而这种监督使得结构变得廉价。每个索引的语料库返回1.6倍的行动节省和1.8倍的准确性,相较于一个覆盖所有内容的无监督实体索引,使用1,913个链接对比其196,112个链接。在库训练过的问题上,未改变的读者在更高的准确性下减少了31%的行动,而这一结果在我们未编写问题的官方PhantomWiki生成上得到了重现。为了衡量这一结果的广度,我们引入了关键覆盖梯度,这是一个探测器,变化训练集触及问题的程度,用衰减曲线替代训练/测试拆分的通过/失败。泛化证明是端点依赖的:准确性在未见过的问题上得以保持(当一个问题的两个关键字都被索引时,F1提高了+0.167;一个被索引时提高了+0.100;两个都未被索引时为零),而行动节省则保持在训练过的问题上。由于这种衰减是通过覆盖而非新颖性进行索引的,更多的训练可以延伸它——而库是未充分训练的,而非饱和的:在新问题中覆盖线性增长,并在训练重复这些问题的瞬间停止,因此一百个问题覆盖了四分之一的语料库,而四倍的问题数量将弥补这一差距。
cs.CL / 26 / 2608.21832

GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding

GUI-Primitives:诊断视觉-语言图形用户界面定位中的空间推理失败
Jahin, Md Abrar, Parvez, Md Rizwan
Abstract
Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element. We introduce GUI-Primitives, a 994-item benchmark of contrastive instruction pairs over seven spatial relations in graphical user interfaces (left/right, above/below, containment, alignment, proximity, list ordinal, occlusion). Each pair holds the screenshot and anchor fixed while changing the relation expression, so the correct target moves between two designated candidates. Five annotators validate a 196-item subset ($\kappa = 0.94$ well-formedness; $\kappa = 0.79$ target selection). Nineteen vision-language models reach at most $32\%$ strict point-in-box accuracy. Because models emit unconstrained coordinates, we classify each prediction by the candidate region it falls within. Predictions fall outside both candidates on $60-92\%$ of items. Conditional on falling within a candidate region, target selection reaches 0.82-0.90 for horizontal position, vertical position, proximity, and list ordinal, but does not differ significantly from 0.50 for containment and occlusion: most failures reflect candidate localization rather than relation understanding. Across ten models, benchmark accuracy correlates with ScreenSpot-Pro accuracy (Spearman $\rho = +0.74$), an exploratory association at this sample size. Marking the two designated candidates raises selection accuracy by 35--57 percentage points, an oracle diagnostic that supplies the candidate set rather than a deployable method. We release the benchmark, predictions, and code.
Chinese Translation
计算机使用代理将自然语言指令与屏幕截图中的界面元素进行绑定,但现有基准并未明确模型是否将关系语言与正确的元素绑定。我们引入了GUI-Primitives,这是一个包含994个对比指令对的基准,涵盖图形用户界面中的七种空间关系(左/右,上/下,包含,排列,接近,列表序号,遮挡)。每对指令在保持屏幕截图和锚点不变的情况下,改变关系表达,因此正确目标在两个指定候选者之间移动。五位注释者验证了196个子集的有效性($ ext{kappa} = 0.94$ 结构良好;$ ext{kappa} = 0.79$ 目标选择)。十九个视觉-语言模型的严格框内准确率最多为$32 ext{ ext{%}}$。由于模型输出的坐标没有约束,我们根据预测落入的候选区域对每个预测进行分类。在$60-92 ext{ ext{%}}$的项目中,预测落在两个候选者之外。在落入候选区域的条件下,目标选择在水平位置、垂直位置、接近度和列表序号上达到了0.82-0.90,但在包含和遮挡上与0.50没有显著差异:大多数失败反映的是候选定位而非关系理解。在十个模型中,基准准确率与ScreenSpot-Pro准确率相关(Spearman $ ho = +0.74$),在这个样本大小下是一种探索性关联。标记两个指定候选者将选择准确率提高了35--57个百分点,这是一种提供候选集的理想诊断,而不是可部署的方法。我们发布了基准、预测和代码。
cs.CL / 27 / 2608.21853

PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding

PUMA:一个基于文化的多模态理解的波兰基准
Dadas, Sławomir, Perełkiewicz, Michał, Poświata, Rafał, Grębowiec, Małgorzata, Jaworski, Bartłomiej, Woźniakowska, Izabela
Abstract
Large language models are increasingly moving beyond text processing, adding support for other modalities such as images and audio. While text understanding and generation have been extensively studied, multimodal data processing capabilities, particularly in the context of cultures and languages other than English, have not yet been evaluated comprehensively. In this paper, we propose PUMA (Polish Unified Multimodal Assessment), a novel benchmark of 900 hand-crafted tasks designed to probe the limits of multimodal models in the Polish cultural and linguistic context. The dataset evaluates both cultural understanding and practical skill in processing text, images, audio, and visually rich documents. Our extensive evaluation of frontier commercial models, open-weights models, and specialized smaller systems highlights a significant performance gap. While top commercial models achieve high scores in visual question answering, most models struggle with complex audio or document understanding. We open-source our evaluation framework to advance localized multimodal AI research.
Chinese Translation
大型语言模型正越来越多地超越文本处理,增加对图像和音频等其他模态的支持。尽管文本理解和生成已被广泛研究,但在非英语文化和语言背景下的多模态数据处理能力尚未得到全面评估。本文提出了PUMA(波兰统一多模态评估),这是一个由900个手工制作的任务组成的新基准,旨在探讨多模态模型在波兰文化和语言背景下的极限。该数据集评估了文化理解和处理文本、图像、音频以及视觉丰富文档的实际技能。我们对前沿商业模型、开放权重模型和专业小型系统的广泛评估突显了显著的性能差距。尽管顶级商业模型在视觉问答中取得了高分,但大多数模型在复杂音频或文档理解方面表现不佳。我们将我们的评估框架开源,以推动本地化多模态人工智能研究。
cs.CL / 28 / 2608.21863

HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning

HiDiffTIR:用于多轮工具集成推理的分层难度感知策略优化
Guo, Yucan, Wang, Xiaohan, Su, Miao, Guan, Saiping, Hou, Zhongni, Chai, Jiajun, Lin, Wei, Yin, Guojun, Jin, Xiaolong, Guo, Jiafeng, Cheng, Xueqi
Abstract
Tool-Integrated Reasoning (TIR) is a fundamental capability for LLM agents to solve complex tasks by interacting with external tools iteratively. Reinforcement Learning (RL) has become the dominant paradigm for enabling this capability. However, existing approaches typically assign uniform trajectory-level advantages and treat all correct tool calls equally, ignoring the varying difficulty and learning value across trajectories and reasoning steps. This can lead to imprecise learning signals that do not adequately distinguish between trivial and challenging tool-use patterns. To address this limitation, we propose HiDiffTIR, a Hierarchical Difficulty-aware policy optimization framework for multi-turn TIR. HiDiffTIR performs difficulty-aware credit assignment at both trajectory and turn levels, enabling the policy to focus on more informative trajectories and harder reasoning steps. Notably, this fine-grained optimization is achieved without additional supervision, relying solely on group-level statistics derived from standard RL rollouts. Extensive experiments on three tool-using benchmarks demonstrate that HiDiffTIR consistently improves multi-turn TIR performance and tool invocation accuracy over strong RL baselines, highlighting the necessity of difficulty-aware credit assignment for effective policy optimization in tool-integrated LLM agents.
Chinese Translation
工具集成推理(Tool-Integrated Reasoning, TIR)是大型语言模型(LLM)代理通过与外部工具进行迭代交互来解决复杂任务的基本能力。强化学习(Reinforcement Learning, RL)已成为实现这一能力的主流范式。然而,现有方法通常对轨迹级别的优势进行统一分配,并将所有正确的工具调用视为相同,忽视了不同轨迹和推理步骤之间的难度和学习价值的差异。这可能导致不精确的学习信号,无法充分区分简单和具有挑战性的工具使用模式。为了解决这一局限性,我们提出了HiDiffTIR,一种用于多轮TIR的分层难度感知策略优化框架。HiDiffTIR在轨迹和轮次级别上进行难度感知的信用分配,使得策略能够专注于更具信息量的轨迹和更困难的推理步骤。值得注意的是,这种细粒度的优化是在没有额外监督的情况下实现的,仅依赖于从标准RL回合中得出的组级统计数据。在三个工具使用基准上的大量实验表明,HiDiffTIR在多轮TIR性能和工具调用准确性方面始终优于强大的RL基线,突显了难度感知信用分配在工具集成LLM代理有效策略优化中的必要性。
cs.CL / 29 / 2608.21871

The Chase Is the Curriculum, the Capture Anchors the Credit: Pursuit-Evasion Self-Play for Zero-Data LLM Reasoning

追逐即课程,捕获锚定信用:零数据大语言模型推理的追逐-逃避自我对弈
Yu, Jing, Chen, Shengchao, Tan, Yiyun
Abstract
Reinforcement learning with verifiable rewards has become the dominant recipe for improving large language model reasoning, yet it presumes large human-curated task collections. Zero-data self-play removes this dependency, but existing methods vet learnability only by probing candidates and rejecting post hoc, never learning where along an environment's difficulty axis to place a task, and credit the solver with sparse terminal rewards alone. We recast zero-data self-play as a pursuit-evasion game: in LURE, an LLM evader positions tasks along each environment's difficulty axis to stay one step ahead of a planner-executor pursuer that hunts it down through verifiable interaction. The evader is trained on a capture-frontier reward that peaks when the solver captures it on exactly half of its rollouts, turning barely catchable into a learned positioning strategy rather than a hand-tuned rejection band. The pursuer earns capture-anchored dense process credit, in which monotone verifier progress is group-normalized jointly with the terminal capture under a round-anchored KL that keeps the co-evolution stable. Across three verifiable reasoning environments and three backbone families, LURE outperforms advanced baselines under unified/specialist settings, while the unified model attains stronger aggregate OOD zero-shot accuracy than all trained baselines across nine held-out benchmarks from three task families.
Chinese Translation
具有可验证奖励的强化学习已成为提升大语言模型推理能力的主流方法,然而它假设存在大量人工策划的任务集合。零数据自我对弈消除了这一依赖,但现有方法仅通过探测候选者并事后拒绝来验证可学习性,从未学习在环境的难度轴上应放置任务的位置,并且仅通过稀疏的终端奖励来给予解题者信用。我们将零数据自我对弈重新构建为一个追逐-逃避游戏:在 LURE 中,一个 LLM 逃避者在每个环境的难度轴上定位任务,以保持领先于通过可验证交互追捕它的规划-执行者追逐者。逃避者的训练基于捕获边界奖励,当解题者在其正好一半的回合中捕获它时,该奖励达到峰值,将难以捕获的任务转变为一种学习的定位策略,而不是手动调优的拒绝带。追逐者则获得基于捕获的密集过程信用,其中单调验证者的进展与终端捕获在一个基于回合的 KL 规范下共同进行组归一化,从而保持共同进化的稳定性。在三个可验证推理环境和三个基础模型系列中,LURE 在统一/专业设置下超越了先进的基线,而统一模型在来自三个任务系列的九个保留基准上获得了比所有训练基线更强的整体 OOD 零-shot 准确率。
cs.CL / 30 / 2608.21880

BanglaVeilGuard: Cross-Script Safety Benchmarking and Lightweight Guardrails for Bangla Large Language Models

BanglaVeilGuard:针对孟加拉大型语言模型的跨脚本安全基准测试与轻量级保护措施
Hassan, Md. Rakibul, Hossain, Muhammad Iqbal
Abstract
Bangla large language model (LLM) safety is difficult to evaluate with English-centric or standard-script benchmarks because Bangla users routinely write across scripts, spellings, code-mixed forms, and regional registers. This paper presents BanglaVeilGuard, a compact Bangla-first safety benchmark and lightweight prompt guard for six language forms: standard Bangla, Romanized Bangla, Banglish, code-mixed Bangla--English, noisy Bangla, and dialectal Bangla. The benchmark contains 2,366 quality-filtered prompts and a held-out 354-prompt evaluation split spanning unsafe, safe, and safe-sensitive requests. BanglaVeilGuard uses non-destructive multi-view normalization with a prompt-risk classifier and thresholded pre-generation gate, allowing it to screen prompts for heterogeneous target models without changing their weights. Across target-model families, guarded runs reduce attack success under deterministic response scoring from 93.8--100.0\% to 6.3\% for Claude Opus 4.8, BanglaLLama, and TituLLM; TigerLLM-1B with BanglaVeilGuard achieves 78.2\% accuracy with 8.8\% ASR. The prompt guard also attains 88.5\% unsafe recall, substantially above the evaluated prompt-only guard baselines. The main remaining cost is over-refusal on dialectal and noisy benign prompts, revealing a concrete safety-helpfulness frontier for Bangla LLM deployment.
Chinese Translation
孟加拉大型语言模型(LLM)的安全性难以通过以英语为中心或标准脚本的基准进行评估,因为孟加拉用户通常在不同脚本、拼写、代码混合形式和地区方言之间书写。本文提出了BanglaVeilGuard,这是一个紧凑的以孟加拉语为主的安全基准和轻量级提示保护措施,涵盖六种语言形式:标准孟加拉语、罗马化孟加拉语、孟加拉英语(Banglish)、代码混合的孟加拉语-英语、嘈杂的孟加拉语和方言孟加拉语。该基准包含2,366个经过质量筛选的提示,以及一个包含354个提示的评估集,涵盖不安全、安全和安全敏感的请求。BanglaVeilGuard使用非破坏性的多视角归一化与提示风险分类器和阈值预生成门,允许其在不改变目标模型权重的情况下筛选异质目标模型的提示。在目标模型系列中,经过保护的运行在确定性响应评分下将攻击成功率从93.8% - 100.0%降低至6.3%,适用于Claude Opus 4.8、BanglaLLama和TituLLM;使用BanglaVeilGuard的TigerLLM-1B在8.8%的ASR下达到了78.2%的准确率。提示保护措施还实现了88.5%的不安全召回率,显著高于评估的仅提示保护基线。主要剩余成本是在方言和嘈杂的良性提示上过度拒绝,揭示了孟加拉LLM部署的具体安全性与有效性边界。
cs.CL / 31 / 2608.21924

Modeling Claim Dependency Structure for Patent Litigation Prediction with Graph Attention Networks

基于图注意力网络的专利诉讼预测的索赔依赖结构建模
Arai, Takao, Inoue, Hiroyasu
Abstract
Patent litigation imposes substantial costs on firms and distorts R&D incentives, making early risk identification a practically important task. While prior work has applied BERT-based models to patent claim text, two fundamental limitations remain: flat sequence encoding loses the dependency structure between independent and dependent claims that legally determines patent scope, and feeding the entire claim set to a single encoder discards legally critical text. A six-model ablation on 1.34 million USPTO utility patents confirms that per-claim encoding, graph connectivity, attention, and Attentional Aggregation each provide independent, additive predictive value. We propose ClaimGAT, a Graph Attention Network that encodes each claim independently, constructs a directed claim dependency graph, processes it with GATConv layers, and aggregates independent claims via Attentional Aggregation to yield both a litigation risk score and claim-level gate weights that enable post-hoc structural analysis. ClaimGAT achieves an AUC-ROC of 0.818 and a lift of 4.89x at the top 10%, using only information observable at the time of patent grant. It reveals a tendency in high-risk patents for structural selection and content sensitivity to diverge, a pattern consistent with defensive claim drafting.
Chinese Translation
专利诉讼给企业带来了巨大的成本,并扭曲了研发激励,使得早期风险识别成为一项实际重要的任务。尽管之前的研究已将基于BERT的模型应用于专利索赔文本,但仍然存在两个基本局限性:平面序列编码丢失了独立索赔和依赖索赔之间的依赖结构,而该结构在法律上决定了专利范围;将整个索赔集输入单个编码器则丢弃了法律上关键的文本。对134万份美国专利商标局(USPTO)实用专利的六模型消融实验确认,逐索赔编码、图连接性、注意力机制和注意力聚合各自提供独立的、附加的预测价值。我们提出了ClaimGAT,一种图注意力网络,独立编码每个索赔,构建有向索赔依赖图,使用GATConv层进行处理,并通过注意力聚合聚合独立索赔,以生成诉讼风险评分和索赔级别的门权重,从而实现事后结构分析。ClaimGAT在仅使用专利授予时可观察到的信息的情况下,达到了0.818的AUC-ROC和在前10%中4.89倍的提升。它揭示了高风险专利在结构选择和内容敏感性方面的趋向分歧,这一模式与防御性索赔起草一致。
cs.CL / 32 / 2608.21946

EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning

EDGE:用于智能强化学习中引导探索的经验蒸馏
Xie, Can, Zhou, Yuyi, Yang, Wen, zhang, Ziyi, Song, Siyao, Deng, Yingzhuo, Ren, Shuo, Zhang, Jiajun
Abstract
Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exploration patterns embedded in interaction trajectories are largely discarded after a single policy update. Existing experience-augmented approaches retrieve historical guidance at inference time, but they apply experiences without accounting for the policy's evolving capability and create persistent dependencies on external retrieval. We propose EDGE (Experience-Distillation for Guided Exploration), a framework that treats retrieved experiences as temporary training-time scaffolds and progressively internalizes their benefits into the parametric policy. Concretely, EDGE partitions each rollout group into experience-conditioned and experience-free trajectories to estimate and admit only positive marginal gains without extra sampling, then distills the induced behavior into the base policy via a reverse-KL objective on its own empirical support. A co-evolutionary experience bank further synthesizes guidance from emerging failure modes and prunes obsolete entries as the policy evolves. On ALFWorld and WebShop, EDGE improves over GRPO by 8.3 and 12.5 success-rate points at the 7B scale and retains 96.0% of its scaffolded performance when external experiences are removed at inference time. The code is available at https://github.com/xvolcano02/EDGE.
Chinese Translation
基于结果的目标的强化学习,例如 GRPO,使基于大型语言模型(LLM)的智能体能够解决复杂的长时间跨度任务,但在单次策略更新后,嵌入交互轨迹中的可重用探索模式大多被丢弃。现有的经验增强方法在推理时检索历史指导,但它们在应用经验时未考虑策略的演变能力,并且对外部检索产生了持续的依赖。我们提出了 EDGE(Experience-Distillation for Guided Exploration)框架,将检索到的经验视为临时训练时间支架,并逐步将其好处内化到参数化策略中。具体而言,EDGE 将每个回放组划分为经验条件和无经验轨迹,以估计并仅接受正的边际收益,而无需额外采样,然后通过在自身经验支持上的反KL目标将诱导行为蒸馏到基础策略中。一个共同进化的经验库进一步合成来自新出现的失败模式的指导,并在策略演变时修剪过时的条目。在 ALFWorld 和 WebShop 上,EDGE 在 7B 规模下比 GRPO 提高了 8.3 和 12.5 个成功率点,并在推理时移除外部经验时保留了 96.0% 的支架性能。代码可在 https://github.com/xvolcano02/EDGE 获取。
cs.CL / 33 / 2608.21950

Bulbul: A Dataset for Dialectal Arabic Speech Recognition

Bulbul:一个用于方言阿拉伯语语音识别的数据集
Ashraf, Ahmed, Alansari, Aisha, Abbas, Fadel Al, Almarwani, Nada, Aloufi, Samah, Ezzini, Saad, Al-Shaibani, Maged S., Dalaq, Doaa, Elmadany, AbdelRahim A., Abdul-Mageed, Muhammad, Trigui, Mohamed Mehdi, Refai, Dania, Refai, Layan, Akrout, Mohamed, Jarrar, Mustafa, Al-Khatib, Wasfi G., Dalaq, Alaa, El-Nakla, Darin, Abdaljalil, Samir, Al-Fakih, Abdulrahman, Zeghib, Nour El Imane, Redah, Moussa, Chafik, Salmane, El-Attar, Mohamed, Grati, Rima, Kohail, Sarah, Alkhorasani, Malak, Safwan, Khadijah Al, Mudhaffar, Ismail M., Altam, Ali, Al-Shaikh, Ahmed, Saeed, Adnan, Luqman, Hamzah
Abstract
Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resources. Existing speech datasets often focus on single dialects or large-scale broadcast/web data, leading to trade-offs between linguistic diversity and annotation quality. We present BULBUL, a multi-dialect Arabic ASR dataset collected from 275 speakers in 11 Arab countries. BULBUL includes structured dialect and sub-dialect coverage, as well as recordings of classical Arabic and modern standard Arabic spoken by participants in their native dialectal accents to support accent-aware modeling. The quality of the recordings was ensured through a two-level human verification process. We further benchmark a range of recent ASR systems, establishing strong baselines for modern dialectal and accented Arabic ASR.
Chinese Translation
阿拉伯语自动语音识别(ASR)面临着由于双语现象、广泛的地区方言变异和有限的语音资源而带来的独特挑战。现有的语音数据集通常集中于单一方言或大规模的广播/网络数据,导致语言多样性与标注质量之间的权衡。我们提出了BULBUL,这是一个从11个阿拉伯国家的275名说话者收集的多方言阿拉伯语ASR数据集。BULBUL涵盖了结构化的方言和子方言,以及以参与者的母语方言口音录制的古典阿拉伯语和现代标准阿拉伯语的录音,以支持对口音敏感的建模。通过两级人工验证过程确保了录音的质量。我们进一步对一系列最新的ASR系统进行了基准测试,为现代方言和带口音的阿拉伯语ASR建立了强有力的基准。
cs.CL / 34 / 2608.21969

ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents

ToSCA:利用层次强化学习对话代理的时间和战略抽象
Wang, Xiaoyu, Gu, Qingqing, Zhao, Yue, Chen, Teng, Cao, Yuqi, Chen, Xiaokai, Li, Hongyan, Ji, Luo
Abstract
Humans have multiple levels of temporal abstractions on daily interaction and thinking, such as concept perception and strategic planning. Inspired by this nature, we propose a two-level hierarchical reinforcement learning (RL) framework for conversational agents, bridging the gap between previous token-level or utterance-level RL methods. Developed on a two-level MDP, the token-level response decoding is conditioned on the utterance-level action, the explicit textual strategies. Based on theoretical derivation and efficiency consideration, we use DQN to solve the high-level critic and PPO to solve the low-level actor-critic. To further alleviate the reward sparsity and facilitate the convergence, we also design the dual-granularity reward mechanism, in which the utterance-level satisfaction score is integrated with token-level intrinsic motivation and K-L penalty. Experiments on both daily and emotional support conversations show that our method outperforms versatile baselines in strategy determination and response quality. Our implementation is available at https://github.com/AaronJi/ToSCA.
Chinese Translation
人类在日常互动和思维中具有多层次的时间抽象,例如概念感知和战略规划。受到这一特性的启发,我们提出了一种针对对话代理的两级层次强化学习(RL)框架,填补了之前基于标记级或发话级RL方法之间的空白。该框架基于两级马尔可夫决策过程(MDP),标记级的响应解码依赖于发话级的动作,即明确的文本策略。基于理论推导和效率考虑,我们使用深度Q网络(DQN)来解决高层次的评论者问题,并使用近端策略优化(PPO)来解决低层次的演员-评论者问题。为了进一步缓解奖励稀疏性并促进收敛,我们还设计了双粒度奖励机制,其中发话级满意度评分与标记级内在动机和K-L惩罚相结合。在日常对话和情感支持对话的实验中,我们的方法在战略确定和响应质量方面优于多种基线。我们的实现代码可在 https://github.com/AaronJi/ToSCA 获取。
cs.CL / 35 / 2608.21975

Machine learning and digital pragmatics: Which word category influences emoji use most?

机器学习与数字语用学:哪种词类对表情符号使用的影响最大?
Shormani, Mohammed Q., AlSohbani, Yehia A., Shormani, Mohammed Q.
Abstract
This study examines the performance of the state-of-the-art MARBERT model in identifying the lexical/pragmatic category associated with emoji use on X within a digital pragmatics approach (DPA). A net corpus of 15856 Colloquial Arabic (CA) posts containing emojis was collected from X using Python. The texts were tokenized and normalized into 4 lexical categories, namely noun_norm, verb_norm, adj_norm, and adverb_norm, and 2 pragmatic/structural categories, question_norm and exclamation_norm. MARBERT was finetuned and optimized to identify which category scores standard metrics more, hence associated with emoji use, while binary logistic regression was used to examine which category is statistically associated with emoji occurrence. Findings unveil that nouns dominate the corpus in normalized frequency (M = 0.675, SD = 0.161), followed by verbs (M = 0.083, SD = 0.100). However, verbs have the strongest influence of emoji use indicated by verb density (\b{eta} = 0.821, p = .001, 95% CI [0.332, 1.309]). The study concludes that in digital pragmatics of CA on X, emoji use association with lexical/pragmatic category can be explained by a hybrid approach of computational, statistical, and pragmatic methods, reflecting the interaction among machine learning, linguistic/lexical features, contextual representation, and pragmatic communication.
Chinese Translation
本研究考察了最先进的MARBERT模型在识别与表情符号使用相关的词汇/语用类别方面的表现,采用数字语用学方法(DPA)。通过Python从X平台收集了包含表情符号的15856条口语阿拉伯语(CA)帖子。文本被分词并规范化为4个词汇类别,即名词规范(noun_norm)、动词规范(verb_norm)、形容词规范(adj_norm)和副词规范(adverb_norm),以及2个语用/结构类别,即疑问句规范(question_norm)和感叹句规范(exclamation_norm)。MARBERT经过微调和优化,以识别哪个类别在标准指标上得分更高,因此与表情符号使用相关,而二元逻辑回归用于检验哪个类别在统计上与表情符号的出现相关。研究结果揭示,名词在规范频率上占主导地位(M = 0.675, SD = 0.161),其次是动词(M = 0.083, SD = 0.100)。然而,动词对表情符号使用的影响最为显著,表明动词密度({eta} = 0.821, p = .001, 95% CI [0.332, 1.309])。研究结论指出,在X平台的CA数字语用学中,表情符号使用与词汇/语用类别的关联可以通过计算、统计和语用方法的混合方法进行解释,反映了机器学习、语言/词汇特征、上下文表示和语用交流之间的互动。
cs.CL / 36 / 2608.22034

Align, Unify, Suppress, Route: A Coherentist View of Transformer Computation

对齐、统一、抑制、路由:一种连贯主义视角下的变换器计算
Aljaafari, Nura, Freitas, Andre
Abstract
Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks and architectures. We introduce Coherentist Probabilistic Compositionalism (CPC), an interpretive framework that grounds transformer computation in coherentist theories of interpretation and describes it through four operator roles. Alignment identifies candidate relations, unification integrates supporting information, suppression reduces incompatible alternatives, and routing carries selected information to the output. Across 15 models from five architecture families, the suppression, unification, and routing weight-space signatures correlate with held-out activation-level role measures above random baselines. Suppression is more stable across tasks than unification. Ablating alignment heads reduces downstream suppressive activity beyond a random-head control in 10 models, but similar effects on no-conflict prompts indicate a general upstream dependency, not contradiction-specific coupling. Explicit contradictions significantly shift a layerwise coherence proxy in 14 models; after removing shared residual covariance, the gap has the predicted direction in every model. Base and instruction-tuned variants preserve induction-head score structure ($r{\geq}0.98$) without a consistent shift of operator signatures towards later layers. These results support CPC as a shared vocabulary for comparing transformer mechanisms while showing that their depth and geometric expression remain architecture-specific.
Chinese Translation
机械解释性已识别出变换器电路,但缺乏一个共享的词汇来描述其功能如何在任务和架构之间组合。我们引入了连贯主义概率组合论(Coherentist Probabilistic Compositionalism, CPC),这是一个将变换器计算基础置于连贯主义解释理论中的解释框架,并通过四种操作角色进行描述。对齐(Alignment)识别候选关系,统一(Unification)整合支持信息,抑制(Suppression)减少不兼容的替代方案,路由(Routing)将选定的信息传递到输出。在来自五个架构家族的15个模型中,抑制、统一和路由的权重空间特征与超出随机基线的激活水平角色测量相关。抑制在任务之间比统一更稳定。消融对齐头部在10个模型中减少了下游抑制活动,超出随机头部控制,但在无冲突提示上的类似效果表明了一种一般的上游依赖,而非特定于矛盾的耦合。显式矛盾在14个模型中显著改变了逐层一致性代理;在去除共享残差协方差后,所有模型中的差距均朝着预测的方向发展。基础和指令调优变体在没有一致性地向后层移动操作符特征的情况下,保持了归纳头分数结构($r{ ext{≥}}0.98$)。这些结果支持CPC作为比较变换器机制的共享词汇,同时表明其深度和几何表达仍然是架构特定的。
cs.CL / 37 / 2608.22071

Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction

Real-TurnTurk:一个用于轮流发言预测的多模态土耳其语语料库
Bayrak, Ahmet Tuğrul, Korkmaz, Fatma Nur, Türker, Bekir Berker, Türkel, Mustafa Sertaç, Kaplan, Alper
Abstract
Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchronous dialog systems. While existing research has explored multimodal approaches and large language models for turn-ending prediction, there is a lack of naturalistic conversational corpora specifically addressing turn-taking dynamics in Turkish. This study introduces a multimodal Turkish conversational dataset of unscripted dyadic interactions, comprising synchronized front-facing video, per-speaker audio channels that allow overlapping speech to be attributed to individual speakers, and time-aligned transcriptions. Turn-taking prediction is formulated as a binary classification problem, and a Genetic Algorithm (GA) is employed to optimize interpretable decision rules derived from visual, acoustic, and linguistic features. A hybrid AND-OR rule representation is adopted in the proposed framework to represent the alternative cue combinations that precede a turn transition.
Chinese Translation
轮流发言是人类对话的基本组织特征,但在自然的同步对话系统中建模仍然困难。尽管现有研究探索了多模态方法和大型语言模型用于轮次结束预测,但缺乏专门针对土耳其语轮流发言动态的自然对话语料库。本研究介绍了一个多模态土耳其语对话数据集,包含非剧本化的双人互动,数据集包括同步的正面视频、每位发言者的音频通道(允许将重叠的语音归因于个别发言者)以及时间对齐的转录文本。轮流发言预测被公式化为一个二分类问题,并采用遗传算法(Genetic Algorithm, GA)来优化从视觉、声学和语言特征中提取的可解释决策规则。所提出的框架采用混合的AND-OR规则表示,以表示在轮次转换之前的替代提示组合。
cs.CL / 38 / 2608.22077

Spine-Branch Coordination for Multi-agent Computer Use

多智能体计算机使用中的脊柱-分支协调
Zhang, Mian, Sharma, Manasi, Zhang, Sheng, Yang, Minglai, Shi, Kejian, Liu, Ying, Chen, Zhiyu Zoey, Zhang, Daniel Yue
Abstract
Computer use agents (CUAs) are increasingly deployed as multi-agent systems that decompose a task into multiple subtasks executed across parallel virtual machines (VMs). However, a critical physical bottleneck is that the state of two VMs cannot be merged. Previous systems handle this ad-hoc rather than treating it as a first-class concern. We propose Spine-Branch Coordination for multi-agent computer use, a framework that decomposes a task into a "spine-branch" graph, where the spine carries the main task flow with continuous VM state and branch tasks execute in parallel to collect information the spine needs to complete the task. Branch VMs are discarded once their tasks finish, so no VM merging ever occurs. Experiments show that on 200 long-horizon tasks from Odysseys and across three CUA backbones, Spine-Branch improves success rate over the baseline system by 6.0% to 16.5%, while reducing per-task cost by 34% to 70%, indicating that explicitly modeling VM-state merging constraint enables multi-agent computer use to scale efficiently.
Chinese Translation
计算机使用代理(CUAs)越来越多地作为多智能体系统被部署,将任务分解为多个在并行虚拟机(VMs)上执行的子任务。然而,一个关键的物理瓶颈是两个虚拟机的状态无法合并。以往的系统对此采取临时处理,而不是将其视为首要问题。我们提出了多智能体计算机使用的脊柱-分支协调框架,该框架将任务分解为“脊柱-分支”图,其中脊柱承载主要任务流程并保持连续的虚拟机状态,而分支任务则并行执行,以收集脊柱完成任务所需的信息。一旦分支虚拟机的任务完成,它们将被丢弃,因此不会发生虚拟机合并。实验表明,在来自Odysseys的200个长时间任务和三个CUA骨干上,脊柱-分支协调相较于基线系统提高了6.0%至16.5%的成功率,同时将每个任务的成本降低了34%至70%,这表明明确建模虚拟机状态合并约束使多智能体计算机使用能够高效扩展。
cs.CL / 39 / 2608.22090

Semantic Reasoning Denoising: Correcting Language Model Reasoning with Semantic Operators

语义推理去噪:使用语义运算符纠正语言模型推理
Yang, Yujiao
Abstract
Large language models can produce fluent reasoning traces whose local semantic errors propagate to an incorrect conclusion, while unconstrained self-correction may preserve, amplify, or introduce errors. Existing diffusion language models provide iterative refinement, but usually define noise as token masking or replacement rather than as errors in the reasoning process. We present Semantic Reasoning Denoising (SRD), an operatorized Markov denoising method for natural-language reasoning trajectories. SRD represents semantic noise with executable error operators that describe the error type, its location, and the corrupted and repaired propositions. Composing these operators constructs progressively noisier states. During training, the model learns to identify the semantic noise active in the current trajectory and to reconstruct the paired adjacent lower-noise state. During inference, noise-level-aware denoising repeatedly predicts an inverse operator and checks whether it is applicable, so each executed update makes a localized move toward a stable trajectory. Across six in-domain benchmarks spanning mathematics, code, knowledge, and commonsense, SRD improves the strongest same backbone baseline by 3.2 points on average. On seven cross-dataset transfer targets, it remains competitive with Llama-3-8B-Instruct and improves the strongest Qwen3-8B baseline average by 2.9 points. Analyses of noise sources, objectives, and denoising depth further show that structured semantic-noise prediction and iterative operator execution are central to the improvement.
Chinese Translation
大型语言模型能够生成流畅的推理轨迹,但其局部语义错误会传播至错误结论,而不受限制的自我纠正可能保留、放大或引入错误。现有的扩散语言模型提供迭代的精炼,但通常将噪声定义为标记掩蔽或替换,而非推理过程中的错误。我们提出了语义推理去噪(Semantic Reasoning Denoising, SRD),这是一种针对自然语言推理轨迹的运算符化马尔可夫去噪方法。SRD用可执行的错误运算符表示语义噪声,这些运算符描述了错误类型、位置以及被损坏和修复的命题。组合这些运算符构建逐渐更嘈杂的状态。在训练过程中,模型学习识别当前轨迹中活跃的语义噪声,并重建配对的相邻低噪声状态。在推理过程中,噪声水平感知去噪反复预测逆运算符并检查其适用性,因此每次执行的更新都朝着稳定轨迹进行局部移动。在涵盖数学、代码、知识和常识的六个领域基准测试中,SRD平均提高了最强同背骨基线3.2分。在七个跨数据集迁移目标上,它与Llama-3-8B-Instruct保持竞争力,并将最强Qwen3-8B基线平均提高了2.9分。对噪声来源、目标和去噪深度的分析进一步表明,结构化的语义噪声预测和迭代运算符执行是改进的核心。
cs.CL / 40 / 2608.22118

RAG Collapse: LLM Responses Collapse When Retrieved Documents Are Self-Authored

RAG崩溃:当检索到的文档是自我创作时,大型语言模型的响应崩溃
Druck, Gregory, Smith, Ethan
Abstract
LLM responses are based on the internet (via training or RAG), and AI is now used to generate a significant amount of content online (Paredes et al., 2026), creating the potential for a self-reinforcing feedback loop. Prior work has shown that when LLMs are recursively trained on their own output, they experience model collapse (Shumailov et al., 2024): responses become less diverse, and eventually no longer resemble the original training data. In this paper, we show that a similar collapse occurs if LLM-based AI systems retrieve references they authored using a search tool. We call this RAG collapse. We conduct extensive experiments with three types of simulations of AI systems retrieving references they generated, using three model families, and 1,019 information-seeking prompts, totaling 1,528 simulations and over one million LLM API calls, and find that 79.6% (1,216/1,528) of simulations end in collapse. Surprisingly, even a single self-authored reference can trigger collapse because the LLM disproportionately cites its own content. This self-bias persists even after controlling for reference quality.
Chinese Translation
大型语言模型(LLM)的响应基于互联网(通过训练或检索增强生成(RAG)),而人工智能现在被用于在线生成大量内容(Paredes et al., 2026),这创造了自我强化反馈循环的潜力。先前的研究表明,当LLM在其自身输出上进行递归训练时,会出现模型崩溃(Shumailov et al., 2024):响应变得不够多样,最终不再类似于原始训练数据。在本文中,我们展示了如果基于LLM的人工智能系统使用搜索工具检索它们自己创作的参考文献,也会发生类似的崩溃。我们称之为RAG崩溃。我们进行了广泛的实验,模拟了三种类型的人工智能系统检索它们生成的参考文献,使用了三种模型系列和1,019个信息检索提示,总计1,528次模拟和超过一百万次LLM API调用,发现79.6%(1,216/1,528)的模拟以崩溃告终。令人惊讶的是,即使是单个自我创作的参考文献也能引发崩溃,因为LLM不成比例地引用其自身内容。这种自我偏见在控制参考文献质量后仍然存在。
cs.CL / 41 / 2608.22124

LLM assisted writing deserves empirical evaluation

大语言模型辅助写作值得进行实证评估
Feng, Xuan Zhong, Lin, Yi, Zhang, Yiye, Weng, Chunhua, Peng, Yifan
Abstract
LLM-assisted writing is often treated as a detection problem, as it raises questions about clarity, integrity, equity, and evaluation. An analysis of 69,209 Health Informatics papers links it to more focused presentation, broader citation practices, and more globally distributed authorship. These patterns do not prove better science, but they support evaluating manuscripts by scholarly quality and accountability rather than by tool use.
Chinese Translation
大语言模型(LLM)辅助写作常被视为一种检测问题,因为它引发了关于清晰性、完整性、公平性和评估的问题。对69,209篇健康信息学论文的分析将其与更集中化的表达、更广泛的引用实践以及更全球化的作者分布联系起来。这些模式并不能证明科学质量的提升,但它们支持通过学术质量和问责制而非工具使用来评估手稿。
cs.CL / 42 / 2608.22132

SSE-Bio: A Structured Self-Evolving Agent with Agentic Retrieval Policy for Multi-Hop Biomedical Reasoning

SSE-Bio:一种具有自主检索策略的结构化自我演化代理,用于多跳生物医学推理
Meng, Zhaohan, Meng, Zaiqiao, Liu, Siwei, Xu, Hao, Yuan, Ke, Ounis, Iadh
Abstract
Biomedical multi-hop question answering (QA) requires models to connect evidence across intermediate entities such as diseases, drugs, proteins, and phenotypes. Existing agents typically rely on static retrieval workflows or coarse-grained prompt rewriting, which can lead to instruction drift when reasoning procedures need to be updated. We propose SSE-Bio, a structured self-evolving agent with an agentic retrieval policy for multi-hop biomedical reasoning. Instead of globally rewriting agent instructions, SSE-Bio maintains a structured state, selectively retrieves knowledge triplets and prior templates through a trainable proxy policy, and improves its reasoning memory through fine-grained template editing. To optimise retrieval decisions, we introduce a proxy-training strategy based on group relative policy optimization, where the proxy is improved through decision-contrastive groups over alternative retrieval choices. Experiments on three biomedical multi-hop QA benchmarks show that SSE-Bio consistently outperforms existing baselines, achieving an improvement of 6.56 absolute points over the strongest self-evolving baseline on BioHopR.
Chinese Translation
生物医学多跳问答(QA)要求模型能够连接疾病、药物、蛋白质和表型等中间实体之间的证据。现有的代理通常依赖于静态检索工作流程或粗粒度的提示重写,这可能导致在推理过程需要更新时出现指令漂移。我们提出了SSE-Bio,一种具有自主检索策略的结构化自我演化代理,专门用于多跳生物医学推理。SSE-Bio并不是全局重写代理指令,而是维护一个结构化状态,通过可训练的代理策略选择性地检索知识三元组和先前模板,并通过细粒度的模板编辑来改善其推理记忆。为了优化检索决策,我们引入了一种基于群体相对策略优化的代理训练策略,其中代理通过对比不同检索选择的决策对比组进行改进。在三个生物医学多跳问答基准上的实验表明,SSE-Bio始终优于现有基线,在BioHopR上相较于最强的自我演化基线提高了6.56个绝对点。
cs.CL / 43 / 2608.22140

Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion

词汇扰动干扰大型语言模型推理:注意力转移的实证研究
Zhu, Jiaqian, Zhang, Yang, Ding, Junhua, Yu, Xiaowei
Abstract
Large Language Models (LLMs) achieve strong reasoning performance, but their robustness to realistic lexical corruption remains poorly understood. We evaluate four open-weight instruction-tuned models and frontier models across four reasoning benchmarks under keyboard noise, character swaps, and filler insertion. Character-level perturbations substantially degrade accuracy, especially on multi-step reasoning tasks, while filler insertion has little effect. We trace this asymmetry to Attention Diversion: lexical corruption fragments subword tokenization, and the resulting fragments attract disproportionate attention mass, concentrated in middle and final transformer layers. Length-matched controls confirm that fragmentation, not prompt length, drives the loss. A factorial intervention then shows why the damage is hard to undo: fragmentation corrupts token content and attention allocation together, and the two are coupled. Restoring clean attention while the content remains corrupted is actively harmful, restoring content alone is insufficient, and only restoring both recovers a substantial share of the gap. This coupling explains why inference-time strategies, including chain-of-thought prompting, spell-checking, self-repair, and stronger repair models, fail to consistently recover performance: each addresses one channel at a time. Code and data are available at https://github.com/Jiaqian-Janelle/Attention-Diversion
Chinese Translation
大型语言模型(LLMs)在推理性能上表现出色,但它们对现实词汇损坏的鲁棒性仍然不够清楚。我们评估了四个开放权重的指令调优模型和前沿模型在键盘噪声、字符交换和填充插入下的四个推理基准。字符级扰动显著降低了准确性,尤其是在多步骤推理任务中,而填充插入的影响较小。我们将这种不对称性追溯到注意力转移:词汇损坏使子词标记化碎片化,导致生成的碎片吸引了不成比例的注意力集中,主要集中在中间和最终的变换层。长度匹配的对照实验确认了碎片化而非提示长度导致了损失。一个因子干预实验显示了为什么这种损害难以逆转:碎片化同时损坏了标记内容和注意力分配,这两者是耦合的。在内容仍然损坏的情况下恢复干净的注意力是有害的,仅恢复内容是不够的,只有同时恢复两者才能显著弥补损失。这种耦合解释了为什么推理时间策略,包括思维链提示、拼写检查、自我修复和更强的修复模型,无法持续恢复性能:每种方法一次只处理一个渠道。代码和数据可在 https://github.com/Jiaqian-Janelle/Attention-Diversion 获取。
cs.CL / 44 / 2608.22152

The Collaboration Tax: How Much LLM Multi-Agent Systems Pay to Coordinate

协作税:LLM多智能体系统在协调时的性能损失
Sun, Weixiang, Wang, Zehong, Huang, Hong, Nelson, Colby, Ye, Yanfang
Abstract
Multi-agent systems built from large language models are deployed widely, yet how much performance is lost when two LLMs must coordinate rather than act alone remains unclear. We formulate the collaboration tax as the team-decentralisation loss of a two-player cooperative game with private information, with two propositions characterising its sign and its equivalence to a max-superadditivity violation. We operationalise this definition on 32 solo-tractable tasks grouped by source of grounding friction and measure it on 11 models from 7 providers. The tax is structured along two no-exception axes: a category ordering across every model and a monotonic decrease with capability. The proximate mechanism is not a reasoning deficit but a four-stage conversational cascade in which agents make ungrounded claims, fail to query the partner, skip integrating both views, and accept the answer without re-derivation. The tax is mechanically predictable from conversation features and partly tractable: a prompt intervention targeting all four stages closes a substantial fraction of the gap, with the dominant bottleneck differing across categories. In heterogeneous pairs the tax is pulled toward the stronger partner rather than the additive midpoint, empirically realising the max-superadditivity violation predicted by our framework. Together these results recast collaboration in LLM systems as a measurable, predictable, and partly tractable cost.
Chinese Translation
基于大型语言模型的多智能体系统被广泛应用,但当两个LLM必须协调而非单独行动时,性能损失的程度仍不清楚。我们将协作税定义为具有私人信息的双人合作博弈中的团队去中心化损失,并提出两个命题来表征其符号及其与最大超加性违反的等价性。我们在32个单独可处理的任务上对这一定义进行了操作化,这些任务按基础摩擦来源分组,并在来自7个提供者的11个模型上进行了测量。该税收沿着两个无例外的轴线结构化:每个模型的类别排序和随着能力的单调下降。其近端机制并非推理缺陷,而是一个四阶段的对话级联,其中智能体提出无基础的主张,未能询问合作伙伴,跳过整合双方观点的过程,并在未重新推导的情况下接受答案。该税收可以从对话特征中机械预测,并且在一定程度上是可处理的:针对所有四个阶段的及时干预可以缩小相当大的一部分差距,而主导瓶颈在不同类别中有所不同。在异质配对中,税收更倾向于强者而非加性中点,实证上实现了我们框架预测的最大超加性违反。这些结果共同将LLM系统中的协作重新定义为一种可测量、可预测且部分可处理的成本。
cs.CL / 45 / 2608.22192

How Agents Represent Humans: Human-Directed Stereotypes in an Open Agent Social Network

代理如何表现人类:开放代理社交网络中的人类导向刻板印象
Xu, Huangchen, Wu, Yuan, Chang, Yi
Abstract
LLM-based agents are increasingly deployed in persistent social environments, where generated claims can be posted, replied to, remembered, and reused. We study human-directed stereotypes on Moltbook, an open agent-native social platform, asking how agents construct humans as a social category. For this human-target analysis, we introduce an annotation framework with four evaluative dimensions---morality, friendliness, competence, and autonomy---and a second-stage subtype scheme for descriptive \textit{other} attributions. We find that competence dominates human-directed evaluations, while many \textit{other} attributions describe humans as epistemic, cultural, or embodied subjects. We further examine how these human representations appear in human--agent narrative contexts and platform-level circulation. As an auxiliary comparison, we analyze agent-internal community feedback through behavioral host affinity. Rather than reproducing the stable insider--outsider rejection often observed in human online communities, Moltbook feedback patterns are better explained by exposure, author visibility, and content selection. These findings suggest that bias in agent societies should be studied not only as isolated model output, but also as a discourse process.
Chinese Translation
基于大语言模型(LLM)的代理越来越多地被部署在持久的社交环境中,在这些环境中,生成的声明可以被发布、回复、记住和重用。我们在开放的代理原生社交平台Moltbook上研究人类导向的刻板印象,探讨代理如何将人类构建为一个社会类别。为了进行这一针对人类的分析,我们引入了一个包含四个评估维度的注释框架——道德、友好、能力和自主性——以及一个针对描述性“其他”归属的第二阶段子类型方案。我们的研究发现,能力在针对人类的评估中占主导地位,而许多“其他”归属则将人类描述为认知、文化或具身的主体。我们进一步考察这些人类表现如何出现在人类与代理的叙事背景和平台层面的传播中。作为辅助比较,我们通过行为主机亲和力分析代理内部社区反馈。与人类在线社区中常见的稳定内外部拒绝现象不同,Moltbook的反馈模式更能通过曝光、作者可见性和内容选择来解释。这些发现表明,代理社会中的偏见不仅应作为孤立的模型输出进行研究,还应视为一种话语过程。
cs.CL / 46 / 2608.22215

Dual-Layer Agentic Memory with Fast Write Routing and Slow Consolidation

双层代理记忆:快速写入路由与缓慢巩固
Li, Wenzhi, Nie, Dong, Lan, Rui, Lyu, Tongtong, Wang, Peiyao, Hong, Lingzi, Pan, Weihang, Pan, Boyuan, Hu, Yao
Abstract
Large language model (LLM) agents operate in dynamic environments where knowledge continuously evolves. Existing memory systems typically treat external memory as a monotonically growing repository, inevitably leading to retrieval degradation and increasing computational costs over time. We argue that the core challenge is not retrieval alone, but managing the knowledge lifecycle: deciding what to externalize, update, or ultimately internalize. Inspired by Complementary Learning Systems (CLS) theory in neuroscience, we propose Dual-Layer Agentic Memory, a framework that shifts memory management to the write phase through cost-aware epistemic routing and periodic parametric consolidation. Incoming information is categorized as non-write, write-new, or write-update, and routed through a small-to-large model cascade that minimizes routing overhead while filtering redundant memories. A subsequent write-back phase selectively consolidates high-value external memories into model parameters via supervised fine-tuning. Experiments demonstrate the dual efficiency of our approach: a 1.7B/8B cascade prunes up to 68% of redundant external memory while escalating fewer than 50% of inputs, yet retains over 98% of the downstream QA Exact Match (EM) achieved by an exhaustive retention baseline. We further show that periodic consolidation successfully internalizes external knowledge, allowing the router to adaptively suppress redundant writes as the model's epistemic boundaries evolve. Overall, our framework presents a unified paradigm for agent memory: selective externalization followed by selective internalization. Code and dataset will be released upon acceptance.
Chinese Translation
大型语言模型(LLM)代理在知识不断演变的动态环境中运作。现有的记忆系统通常将外部记忆视为单调增长的存储库,最终导致检索性能下降和计算成本随时间增加。我们认为核心挑战不仅在于检索,而在于管理知识生命周期:决定什么内容需要外部化、更新或最终内化。受到神经科学中互补学习系统(Complementary Learning Systems, CLS)理论的启发,我们提出了双层代理记忆(Dual-Layer Agentic Memory)框架,该框架通过成本感知的认知路由和周期性参数巩固,将记忆管理转移到写入阶段。输入信息被分类为非写入、写入新内容或写入更新,并通过小到大的模型级联进行路由,以最小化路由开销,同时过滤冗余记忆。随后的写回阶段通过监督微调选择性地将高价值的外部记忆巩固为模型参数。实验表明我们方法的双重效率:1.7B/8B级联最多修剪68%的冗余外部记忆,同时提升的输入少于50%,但仍保留超过98%由全面保留基线实现的下游问答精确匹配(Exact Match, EM)。我们进一步展示周期性巩固成功地将外部知识内化,使路由器能够随着模型的认知边界演变而自适应地抑制冗余写入。总体而言,我们的框架呈现了一种统一的代理记忆范式:选择性外部化后再进行选择性内化。代码和数据集将在接受后发布。
cs.CL / 47 / 2608.22229

Grounded Normative Rule Generation with Structured Search

基于结构化搜索的规范性规则生成
Kong, Fanqi, Yin, Huaxiao, Zhang, Ruijie, Zhang, Xiaoyuan, Huang, Yizhe, Gao, Jian, Chen, Shuo, Zhu, Song-Chun
Abstract
Normative rules like institutional charters and workplace policies must be both human-readable and operationally verifiable against actual environment records. However, current language generation and structured-output benchmarks primarily reward surface fluency or schema compliance, leaving operational grounding weakly tested. This creates a critical vulnerability where standard language models generate plausible-sounding policies that fail during enforcement because they rely on unavailable data logs or misaligned scopes. To address this challenge, we formalize the problem as Grounded Normative Rule Synthesis (GNRS) and introduce GNRS-Search, a framework that utilizes Markov Chain Monte Carlo (MCMC) sampling to optimize a discrete, five-slot And-Or Graph (AOG). By explicitly decoupling intermediate operational structure from final prose generation, this method isolates executable feasibility from writing style and allows rule failures to be localized prior to surface realization. We evaluate our approach on GNRS-Bench, a benchmark spanning 116 controlled goals across eight scene families, and RealCharter-Bench, which evaluates transfer to 53 real-derived policy tasks with hidden source clauses. GNRS-Search raises average rubric quality from 68.8% to 81.0% and ranks first under a disclosed executable composite metric, while systematic slot interventions confirm that performance gains stem from robust operational logic rather than rhetorical tuning. Ultimately, by transforming automated rule drafting into an inspectable search problem, this work provides a foundational paradigm for deploying verifiable and compliance-ready personal agents within regulated environments.
Chinese Translation
规范性规则,如机构章程和工作场所政策,必须既可供人类阅读,又能根据实际环境记录进行操作验证。然而,目前的语言生成和结构化输出基准主要奖励表面流畅性或模式合规性,导致操作基础的测试较为薄弱。这造成了一个关键的脆弱性,即标准语言模型生成的政策听起来合理,但在执行时失败,因为它们依赖于不可用的数据日志或不匹配的范围。为了解决这一挑战,我们将问题形式化为基于基础的规范性规则合成(Grounded Normative Rule Synthesis, GNRS),并引入GNRS-Search,一个利用马尔可夫链蒙特卡洛(Markov Chain Monte Carlo, MCMC)采样来优化离散的五槽与或图(And-Or Graph, AOG)的框架。通过明确将中间操作结构与最终文本生成解耦,这种方法将可执行性与写作风格隔离,并允许在表面实现之前定位规则失败。我们在GNRS-Bench上评估了我们的方法,该基准涵盖了跨八个场景系列的116个受控目标,以及RealCharter-Bench,该基准评估转移到53个具有隐藏源条款的真实派生政策任务。GNRS-Search将平均评分质量从68.8%提高到81.0%,并在公开的可执行复合指标下排名第一,而系统性的槽干预确认性能提升源于稳健的操作逻辑,而非修辞调整。最终,通过将自动规则起草转变为可检查的搜索问题,本研究为在受监管环境中部署可验证和合规的个人代理提供了基础范式。
cs.CL / 48 / 2608.22230

Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation

粉饰仇恨,抹黑无害内容:基于注释者风格的反驳攻击对大型语言模型(LLM)基础的内容审核的影响
Lu, Junyu, Liu, Kaiyuan, Kang, Jingyi, Ji, Deyi, Zhang, Hailong, Zhu, Lanyun, Zhu, Qi, Xu, Bo, Yang, Liang, Lin, Hongfei
Abstract
Large language models (LLMs) are increasingly used for hate speech moderation, often within human--AI workflows in which reviewers provide feedback before a final decision. Such feedback introduces two manipulation directions: whitewashing hateful content as normal and smearing normal content as hateful. This study examines the susceptibility of initially correct model judgments to annotator-style rebuttals and analyzes whether attack effectiveness differs across manipulation directions. We introduce a rejudge protocol that extends direct contradiction with decision-boundary perturbations and adversarial rationales. Experiments with multiple LLMs on two hate speech datasets show that annotator-style rebuttals substantially degrade moderation performance, with stronger effects in multi-turn settings. The results further reveal stable, model-specific asymmetries between whitewashing and smearing across attack configurations, indicating distinct directional vulnerability patterns. Explicit reasoning prompts and defensive instructions reduce these effects but do not eliminate them. These findings highlight the need for direction-aware safeguards and dedicated feedback-robustness evaluation in human--AI moderation workflows.
Chinese Translation
大型语言模型(LLMs)越来越多地被用于仇恨言论的审核,通常是在一个人类与人工智能的工作流程中,审核员在最终决策之前提供反馈。这种反馈引入了两个操控方向:将仇恨内容粉饰为正常内容,以及将正常内容抹黑为仇恨内容。本研究考察了最初正确的模型判断对注释者风格反驳的敏感性,并分析了攻击效果在不同操控方向上的差异。我们引入了一种重新判断协议,该协议通过决策边界扰动和对抗性理由扩展了直接矛盾的概念。在两个仇恨言论数据集上对多个LLM进行的实验表明,注释者风格的反驳显著降低了审核性能,尤其在多轮对话设置中效果更为明显。结果进一步揭示了在攻击配置中,粉饰与抹黑之间存在稳定的、模型特定的不对称性,表明不同的方向性脆弱性模式。明确的推理提示和防御性指令可以减轻这些影响,但无法完全消除。这些发现强调了在人工智能与人类审核工作流程中需要具备方向感知的安全措施和专门的反馈鲁棒性评估。
cs.CL / 49 / 2608.22244

Improving Few-Step Language Flows with Untied Self-Conditioning

通过解耦自条件化改善少步语言流
Li, Bocheng, Xu, Linli
Abstract
Flow-matching language models refine all token positions in parallel and can trade sampling steps for latency, yet generation quality still degrades sharply with few sampling steps. We trace a source of this degradation to a train--inference mismatch in previous-prediction self-conditioning: during training, the self-conditioning input is computed from the current noisy state with no intervening solver step; during sampling, the solver folds the previous prediction into the latent before that same prediction reappears as the explicit self-conditioning input. This coupling, absent during training, creates redundancy that grows with step width. We show that the mismatch degrades both the self-conditioning input and the solver update, and derive a correction for each from the model's own structure. From the frozen projection weights we identify directions along which the self-conditioning input is redundant with the latent and dampen them; from the solver's integration structure we derive that a step-average prediction is needed and approximate it from prediction history, with scale set by offline trajectory statistics. The resulting sampler, Untied Self-Conditioning, requires no retraining and uses one evaluation per step. At 8 sampling steps on LangFlow, it reduces OpenWebText generative perplexity from $531$ to~$62$ ($8.6\times$); under an adapted Arena-Hard-Auto~v2 protocol, its outputs are preferred in $96\%$ of pairwise comparisons. On ELF-B it reduces generative perplexity from $71$ to~$43$. Improvements hold from 8 to 256 sampling steps.
Chinese Translation
流匹配语言模型并行优化所有标记位置,可以在延迟和采样步骤之间进行权衡,然而生成质量在少量采样步骤下仍然急剧下降。我们追踪到这种降级的一个原因是先前预测自条件化中的训练-推理不匹配:在训练期间,自条件化输入是从当前的噪声状态计算得出的,没有中间求解步骤;而在采样期间,求解器将先前的预测折叠到潜在空间中,然后该预测再次作为显式自条件化输入出现。这种在训练期间缺失的耦合关系产生了随着步骤宽度增加而增长的冗余。我们表明,这种不匹配会降低自条件化输入和求解器更新的质量,并从模型自身结构中推导出每个部分的修正。通过冻结的投影权重,我们识别出自条件化输入与潜在空间冗余的方向并对其进行抑制;通过求解器的积分结构,我们推导出需要一个步骤平均预测,并从预测历史中进行近似,规模由离线轨迹统计设置。最终的采样器“解耦自条件化”(Untied Self-Conditioning)无需重新训练,并且每一步只需一次评估。在LangFlow上进行8次采样步骤时,它将OpenWebText的生成困惑度从531降低到62(减少了8.6倍);在适应的Arena-Hard-Auto v2协议下,其输出在96%的成对比较中被偏好。在ELF-B上,它将生成困惑度从71降低到43。改进在8到256个采样步骤之间保持有效。
cs.CL / 50 / 2608.22246

N\"urnberg NLP @ GermEval Shared Task 2026: Harmful Content Detection in German Social Media through Error-Independent LLM Voters

纽伦堡 NLP @ GermEval 共享任务 2026:通过无误差独立的 LLM 投票者检测德语社交媒体中的有害内容
Steigerwald, Philipp, Rudolph, Eric, Albrecht, Jens
Abstract
Harmful content in German social media does real-world damage, from calls to action to criminal defamation. The GermEval 2026 shared task scores its detection in four subtasks. The technical challenge is a severe class imbalance. The harmful classes are rare and share surface language with the dominant majority class, yet under macro-F1 they decide the score. The decisive lever is then not a stronger single model but error independence. This insight becomes a per-subtask nine-voter ensemble spanning three orthogonal axes: LLM, training method and class scope. Selected mainly on internal cross-validation, the system reaches macro-F1 of 89.56 (C2A), 71.63 (DBO), 54.84 (VIO) and 83.02 (DEF) on the hidden test set, placing first on all four subtasks.
Chinese Translation
德语社交媒体中的有害内容对现实世界造成了伤害,从号召行动到刑事诽谤。GermEval 2026 共享任务在四个子任务中对其检测效果进行评分。技术挑战在于严重的类别不平衡。有害类别稀少,并与占主导地位的多数类别共享表面语言,但在宏观 F1 分数下,它们决定了最终得分。因此,决定性杠杆不是更强的单一模型,而是误差独立性。这一见解形成了一个每个子任务的九投票者集成,涵盖三个正交轴:LLM、训练方法和类别范围。该系统主要基于内部交叉验证进行选择,在隐藏测试集上达到了宏观 F1 分数 89.56(C2A)、71.63(DBO)、54.84(VIO)和 83.02(DEF),在所有四个子任务中均排名第一。
cs.CL / 51 / 2608.22274

Length-Adaptive Decoding for Masked Diffusion Machine Translation

适应长度的掩蔽扩散机器翻译解码
Zhan, Yan, Hou, Mengkai, Zhang, Wanting, Gao, Zhijun
Abstract
Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithfully, while fixed canvas decoding must choose target length before denoising. Existing masked diffusion decoding work mainly studies token unmasking order, leaving this length decision under-explored despite its direct effect on coverage and redundancy. We introduce Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entropy from all-mask forward passes and selects the canvas the backbone is most prepared to fill. Relative to a baseline using training corpus length statistics, EV recovers 64.9%, 65.3%, and 33.0% of the COMET-22 gain from reference target lengths on En$\to$Zh, Zh$\to$En, and En$\to$De. Our diagnostics show that denoising-friendly lengths need not match reference lengths. Evaluation by three translation experts supports the En$\leftrightarrow$Zh adequacy gains, with stronger evidence on Zh$\to$En. Compared with a LLaMA-3-8B autoregressive (AR) model trained on the same fine-tuning data, the EV system ties on En$\to$Zh and leads on Zh$\to$En; an oracle-length diagnostic further shows that, in this masked diffusion MT setting, deciding which tokens to reveal first matters less than how the target length is supplied.
Chinese Translation
机器翻译测试掩蔽扩散语言模型(dLLMs),因为每个源标记必须被忠实呈现,而固定画布解码必须在去噪之前选择目标长度。现有的掩蔽扩散解码工作主要研究标记解掩顺序,尽管长度决策直接影响覆盖率和冗余,但这一方面的研究仍然不足。我们引入了熵谷(Entropy-Valley, EV),这是一种无训练的长度选择器,通过所有掩蔽前向传递的平均预测熵对候选目标画布进行评分,并选择最适合填充的画布。相较于使用训练语料库长度统计的基线,EV在英译中文(En→Zh)、中译英(Zh→En)和英译德文(En→De)上分别恢复了64.9%、65.3%和33.0%的COMET-22增益。我们的诊断表明,去噪友好的长度不必与参考长度匹配。三位翻译专家的评估支持了英中互译的充分性增益,尤其在中译英方面证据更为强烈。与在相同微调数据上训练的LLaMA-3-8B自回归(AR)模型相比,EV系统在英译中文上持平,在中译英上领先;一个理想长度的诊断进一步表明,在这种掩蔽扩散机器翻译设置中,决定首先揭示哪些标记的重要性低于如何提供目标长度。
cs.CL / 52 / 2608.22295

LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model

对未见问题的LLM评估:上下文多维IRT模型
Shang, Ergan, Tang, Weijing, He, Yinqiu
Abstract
Evaluation of large language models (LLMs) increasingly requires predicting how a model will perform on new questions or tasks before collecting large amounts of new annotations. This problem is challenging because question difficulty, scenario, and underlying capability demands can vary substantially. Simple retrospective averages may confound model ability with item characteristics. In this paper, we study a model-based evaluation framework that combines multidimensional item response theory model with question contexts to predict LLM performance on unseen questions. The framework represents LLMs through latent capability profiles while using question content to inform item characteristics, allowing information to transfer beyond previously observed items. Empirically, we find that for within-scenario evaluation, incorporating question embeddings improves prediction relative to model-free baselines, and that multidimensional latent structure provides a richer description of capability variation than unidimensional alternatives. At the same time, our results reveal an important limitation that the generalizability does not necessarily translate into reliable prediction under cross-scenario shift. These findings suggest that context-aware psychometric modeling is a promising direction for efficient and interpretable LLM evaluation, while also highlighting cross-scenario generalization as a central open challenge.
Chinese Translation
对大型语言模型(LLMs)的评估越来越需要在收集大量新注释之前预测模型在新问题或任务上的表现。这个问题具有挑战性,因为问题的难度、情境和潜在能力要求可能会有很大差异。简单的回顾性平均可能会将模型能力与项目特征混淆。在本文中,我们研究了一种基于模型的评估框架,该框架结合了多维项目反应理论(IRT)模型与问题上下文,以预测LLM在未见问题上的表现。该框架通过潜在能力特征来表示LLM,同时使用问题内容来告知项目特征,从而允许信息超越先前观察到的项目进行转移。实证结果表明,在情境内评估中,结合问题嵌入相较于无模型基线改善了预测效果,而多维潜在结构提供了比单维替代方案更丰富的能力变化描述。同时,我们的结果揭示了一个重要的局限性,即可推广性并不一定转化为在跨情境转变下的可靠预测。这些发现表明,基于上下文的心理测量建模是高效且可解释的LLM评估的一个有前景的方向,同时也突显了跨情境泛化作为一个核心的开放挑战。
cs.CL / 53 / 2608.22312

Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models

基于文本的语义扰动用于可转移的监狱突破攻击针对多模态大型语言模型
Li, Wenyun, Cao, Guiping, Lan, Xiangyuan, Zhang, Zheng
Abstract
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language interaction, yet their safety alignment remains vulnerable to jailbreak attacks. A key challenge is that safety behavior learned in the textual space does not reliably transfer to fused cross-modal representations, leaving multimodal inputs exploitable through latent semantic cues. We propose Text-Anchored Semantic Perturbation Attack (TA-SPA), a black-box jailbreak framework that optimizes transferable perturbations in a text-anchored semantic space. TA-SPA integrates Text-Anchored Semantic Factorization (TASF), which encourages the separation of cross-modal semantic factors from modality-specific residuals, with Semantic-Preserving Augmentation (SPA), which diversifies harmful target anchors while preserving semantic consistency. Experiments show strong attack effectiveness and transfer to commercial MLLMs, with competitive performance under representative defenses. Additional controls and probing support the intended factorization without implying perfect disentanglement, motivating representation-level safety alignment beyond input-level filtering.
Chinese Translation
多模态大型语言模型(MLLMs)在视觉-语言交互方面取得了显著进展,但其安全性对监狱突破攻击仍然脆弱。一个关键挑战是,文本空间中学习到的安全行为并不能可靠地转移到融合的跨模态表示上,这使得多模态输入可以通过潜在的语义线索被利用。我们提出了基于文本的语义扰动攻击(Text-Anchored Semantic Perturbation Attack, TA-SPA),这是一种黑箱监狱突破框架,旨在优化文本锚定语义空间中的可转移扰动。TA-SPA结合了文本锚定语义分解(Text-Anchored Semantic Factorization, TASF),该方法鼓励将跨模态语义因子与特定模态的残差分离,以及语义保持增强(Semantic-Preserving Augmentation, SPA),该方法在保持语义一致性的同时多样化有害目标锚点。实验结果表明,该攻击在商业MLLMs上具有强大的攻击效果和可转移性,并在具有代表性的防御下表现出竞争力。额外的控制和探测支持了预期的分解,而不意味着完美的解耦,激励了超越输入级过滤的表示级安全对齐。
cs.CL / 54 / 2608.22321

Semantics or Structure? Auditing Text Sensitivity in Multimodal Time-Series Forecasting

语义还是结构?多模态时间序列预测中的文本敏感性审计
Sridhar, Karthik, Gupta, Atharva, Pradhan, Nishant, Mandal, Murari, Kumar, Dhruv, Deshpande, Saurabh
Abstract
Multimodal time-series forecasting has emerged as a promising paradigm in which natural-language context is expected to improve predictive performance. Recent multimodal foundation models, including Aurora, as well as early- and late-fusion approaches such as MM-TSFlib and TaTS, report substantial gains over unimodal baselines on the Time-MMD benchmark, attributing these improvements to textual information. However, whether these models are actually sensitive to the semantic content of the text remains unverified. We address this question through controlled text perturbations, attribution analyses, and probes of Aurora's text pathway. On Time-MMD, swapping each row's text for any other real text (empty, constant, within-domain shuffled, or cross-domain) moves mean MSE by less than $0.5\%$ on all three architectures. The improvement reported in the literature is recovered when a co-shipped numeric column is removed without touching text. We conclude that, on this benchmark and within this family of frozen-encoder architectures, text content is not the operative signal behind the reported gains. To support future work on text integration in multimodal foundation models for structured data, we release our perturbation protocol and evaluation harness as a reusable diagnostic toolkit.
Chinese Translation
多模态时间序列预测已成为一种有前景的范式,其中自然语言上下文被期望能够提高预测性能。最近的多模态基础模型,包括Aurora,以及早期和晚期融合方法如MM-TSFlib和TaTS,在Time-MMD基准测试中报告了相较于单模态基线的显著提升,并将这些改进归因于文本信息。然而,这些模型是否真正对文本的语义内容敏感仍未得到验证。我们通过控制文本扰动、归因分析以及对Aurora文本路径的探测来解决这个问题。在Time-MMD上,将每一行的文本替换为其他任何真实文本(空文本、常量、域内打乱或跨域)对所有三种架构的均方误差(MSE)均影响不到$0.5 ext{ extperthousand}$。当移除一个与文本无关的共发数字列时,文献中报告的改进得以恢复。我们得出结论,在这个基准测试和这一系列冻结编码器架构中,文本内容并不是报告增益背后的有效信号。为了支持未来在结构化数据的多模态基础模型中进行文本集成的研究,我们发布了我们的扰动协议和评估工具,作为可重用的诊断工具包。
cs.CL / 55 / 2608.22331

Noise Floor Audit for Agent Benchmarks

代理基准的噪声底线审计
Chen, Yihang, Qian, Pin, Wang, Su, Peng, Chong, Xu, Huan, Wu, Xiyang, Sun, Yiqi
Abstract
We audit measurement variability for 3 native tool-calling endpoints across 2 providers on the official BFCL multiple and parallel categories, using matched AST grading. At temperature 0, reruns are nearly deterministic across Groq endpoints and a thinking-enabled Gemini setting: ever-flip fractions are 0.7%, 2.0%, and 2.7%, with mean run correlations of 0.997, 0.966, and 0.961. Semantics-preserving prompt perturbations create the larger floor on all endpoints, with median perturbation paired SDs 11x to 58x larger than rerun paired SDs. The failure character also shifts: malformed-output failures account for 30%, 7%, and <1% of task failures, so marginal accuracy hides not only stability but also failure mode.
Chinese Translation
我们对官方 BFCL 多重和并行类别中 2 个提供商的 3 个原生工具调用端点的测量变异性进行了审计,采用匹配的 AST 评分。在温度为 0 的情况下,Groq 端点和启用思考的 Gemini 设置的重跑几乎是确定性的:翻转分数分别为 0.7%、2.0% 和 2.7%,平均运行相关性分别为 0.997、0.966 和 0.961。保持语义的提示扰动在所有端点上产生了更大的底线,配对的中位扰动标准差比重跑配对标准差大 11 倍到 58 倍。故障特征也发生了变化:格式错误输出故障占任务故障的 30%、7% 和 <1%,因此边际准确性不仅掩盖了稳定性,还掩盖了故障模式。
cs.CL / 56 / 2608.22332

Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching

通过序列激活修补实现链式思维推理的机制可解释性
Dura, Murat, Öztürk, Serkan, Tekir, Selma
Abstract
Large Language Models (LLMs) demonstrate remarkable problem-solving capabilities when guided by Chain-of-Thought (CoT) prompting, yet the internal mechanisms underlying these improvements remain poorly understood. In this work, we investigate where CoT-related causal effects emerge across the generated reasoning trajectory and which attention heads carry signals that contribute to final-answer computation. Because CoT reasoning unfolds over multiple generated tokens, standard activation patching at a single static token position is insufficient to characterize these temporally distributed effects. To address this limitation, we introduce a sequential activation patching framework that traces CoT-conditioned attention-head activations across token positions and aggregates their effects using Part-of-Speech-guided analysis. We further introduce Sequential Multi-Head Patching to evaluate the joint contribution of distributed head sets, together with cross-question and random activation controls. Targeted zero-ablation experiments show that the identified heads are functionally important for successful answer generation and affect several overlapping mechanisms, including reasoning-trajectory maintenance, answer anchoring, exemplar-target separation, and numerical generation. Overall, our results provide evidence for distributed reasoning-support sub-circuits associated with CoT-conditioned computation.
Chinese Translation
大型语言模型(LLMs)在链式思维(CoT)提示的指导下展现出卓越的问题解决能力,但这些改进背后的内部机制仍然不甚明了。在本研究中,我们探讨了CoT相关的因果效应在生成的推理轨迹中出现的位置,以及哪些注意力头携带了对最终答案计算有贡献的信号。由于CoT推理是在多个生成的标记上展开的,单一静态标记位置的标准激活修补不足以表征这些时间分布的效应。为了解决这一局限性,我们引入了一种序列激活修补框架,该框架追踪CoT条件下的注意力头激活在标记位置上的变化,并利用词性指导分析汇总其效应。我们进一步引入了序列多头修补,以评估分布式头集的联合贡献,并结合跨问题和随机激活控制。针对性的零消融实验表明,识别出的头在成功生成答案中具有功能重要性,并影响多个重叠机制,包括推理轨迹维护、答案锚定、示例目标分离和数值生成。总体而言,我们的结果为与CoT条件计算相关的分布式推理支持子电路提供了证据。
cs.CL / 57 / 2608.22335

Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms

注册转变破坏了大型语言模型的安全性:一个具有文化基础危害的孟加拉语基准
Islam, Naymul, Lia, Nusrat Jahan, Dipta, Shubhashis Roy, Sultan, Sabik Bin, Zehady, Abdullah Khan
Abstract
Bengali is the seventh-most-spoken language globally, yet LLM safety evaluation remains overwhelmingly English-centric. We introduce BanglaSafe, a benchmark of 879 Bengali prompts combining 309 natively authored prompts with 570 expert-reviewed prompts, spanning 17 culturally grounded harm categories and five prompting conditions that vary language, writing style, and authority framing. Evaluating 18 frontier LLMs, we find that over half of all responses are unsafe or partially unsafe (53.6%) while 14.7% contains strictly harmful content, and that the strongest observed effect is not the switch from English to Bengali but the choice of writing style within Bengali: the same harmful request phrased as a formal newspaper investigation succeeds 17 percentage points more often than the same request phrased as a casual message, with no adversarial engineering involved. We further show that existing safety classifiers struggle to reliably evaluate Bengali content, with even frontier models failing on nearly half of all cases.
Chinese Translation
孟加拉语是全球第七大语言,但大型语言模型(LLM)的安全性评估仍然以英语为中心。我们介绍了BanglaSafe,这是一个包含879个孟加拉语提示的基准,其中包括309个本土创作的提示和570个专家审核的提示,涵盖17个文化基础的危害类别和五种提示条件,这些条件在语言、写作风格和权威框架上有所不同。对18个前沿大型语言模型的评估显示,超过一半的响应是不安全或部分不安全的(53.6%),而14.7%的响应包含严格有害的内容。我们观察到的最强效应并不是从英语转向孟加拉语,而是在孟加拉语中写作风格的选择:同样的有害请求以正式的报纸调查形式表述时,其成功率比以随意消息形式表述时高出17个百分点,且没有涉及对抗性工程。我们进一步表明,现有的安全分类器在可靠评估孟加拉语内容方面存在困难,甚至前沿模型在近一半的案例中也未能成功。
cs.CL / 58 / 2608.22339

When Not to Imitate: Boundary-Aware Skill Memory for Reliable Tool-Use LLM Agents

何时不模仿:边界感知技能记忆用于可靠的工具使用大语言模型代理
Lin, Zihan, Chen, Zhenyu, Wei, Jiawen, Wang, Xiaohan, Cao, Jie, Chai, Jiajun, Lin, Wei, Yin, Guojun, He, Ran
Abstract
Extracting skills from past successes is critical for the efficient evolution of Large Language Model (LLM) agents. Prevailing agent self-evolution paradigms typically rely on a core assumption: equipping LLMs with skill memories derived from successful trajectories will monotonically improve their problem-solving capabilities. However, probe analyses reveal that extracting skills solely from successful trajectories traps the model in a \textbf{Skill Imitation Trap}. For tasks that resemble past successes but require different tools, retrieving more skills paradoxically increases the model's confidence in wrong tool calls---procedure skills raise the wrong-tool margin by $47\%$ over a memory-free baseline. To overcome this limitation, we propose \textbf{Boundary-Aware Skill Memory} (BASM), which augments each skill with explicit boundary fields---applicability conditions, risk cues, avoidance rules, and recovery notes. These fields transform each retrieved skill from an unconditional action template into state-conditioned guidance: the agent applies the skill when its conditions hold, suppresses inapplicable tool calls when they do not, and issues targeted repairs when execution fails. Across three agent benchmarks and four model scales, BASM consistently outperforms success-distilled skill-memory baselines: it improves task success rate by up to $23.8\%$ on AppWorld, accuracy by up to $5.0\%$ on BFCL, and reduces attack success rate by $4.6\%$ on AgentDojo, while simultaneously reducing average AppWorld steps by up to $6.6\%$ relative to the memory-free baseline.
Chinese Translation
从过去的成功中提取技能对于大语言模型(LLM)代理的高效演化至关重要。现有的代理自我演化范式通常依赖于一个核心假设:为LLM配备来自成功轨迹的技能记忆将单调地提高其解决问题的能力。然而,探测分析表明,仅从成功轨迹中提取技能会使模型陷入 extbf{技能模仿陷阱}。对于那些与过去成功相似但需要不同工具的任务,检索更多技能反而会增加模型对错误工具调用的信心——程序技能使错误工具的边际提高了$47\%$,相较于无记忆基线。为克服这一限制,我们提出了 extbf{边界感知技能记忆}(BASM),该方法为每个技能增加了明确的边界字段——适用条件、风险提示、避免规则和恢复说明。这些字段将每个检索到的技能从无条件的行动模板转变为状态条件指导:当条件满足时,代理应用该技能;当条件不满足时,抑制不适用的工具调用;当执行失败时,发出针对性的修复。在三个代理基准和四个模型规模上,BASM始终优于成功提炼的技能记忆基线:在AppWorld上提高任务成功率高达$23.8\\%$,在BFCL上提高准确率高达$5.0\\%$,并在AgentDojo上将攻击成功率降低$4.6\\%$,同时相较于无记忆基线将平均AppWorld步骤减少高达$6.6\\%$。
cs.CL / 59 / 2608.22367

Context-Aware Cluster Decoding: Semantic Anchor-Driven Coherence in dMLLMs

上下文感知聚类解码:基于语义锚点的连贯性在 dMLLMs 中的应用
Zhao, Yikai, Zhao, Qiyan, Zhang, Jiaquan, Zhang, Xiaofeng, Yuan, Xiaosong, Cheng, Pengzhou
Abstract
Diffusion multimodal large language models (dMLLMs) frequently produce long-form outputs marred by semantic drift and repetition, with quality generally degrading as output length increases. We identify two structural deficiencies in existing decoding methods as primary drivers of these failures: confidence-based scoring ignores decoded-neighbor support, and block partitioning prevents access to high-readiness semantic anchors, together causing tokens to be committed before their local context is sufficiently established. We propose \ours{} (\textbf{C}ontext-\textbf{A}ware \textbf{C}luster \textbf{D}ecoding), a training-free decoding method that scores each masked position by a multiplicative composite of softmax confidence and neighbor proximity, promoting contextually ready tokens above isolated candidates while suppressing low-confidence positional noise, operating block-free to keep high-readiness anchors globally accessible. \ours{} further applies architecture-aware calibration to handle confidence heterogeneity induced by diverse visual integration strategies. Experiments on three dMLLMs across four benchmarks demonstrate consistent quality gains and hallucination reduction over Original, with larger gains in several longer generation settings, highlighting the importance of neighbor support and visual integration strategy for future dMLLM decoding method design. Our code is openly available at https://github.com/zhaoyk-sysu/CACD-dMLLM.
Chinese Translation
扩散多模态大语言模型(dMLLMs)经常生成长篇输出,这些输出常常受到语义漂移和重复的困扰,且随着输出长度的增加,质量普遍下降。我们识别出现有解码方法中的两个结构性缺陷是导致这些失败的主要原因:基于置信度的评分忽视了解码邻居的支持,而块分区则阻止了对高准备度语义锚点的访问,这两者共同导致在局部上下文尚未充分建立之前就已提交令牌。我们提出了 extbf{C}ontext- extbf{A}ware extbf{C}luster extbf{D}ecoding( extit{ours}),这是一种无训练的解码方法,通过对每个被屏蔽位置进行软最大置信度和邻居接近度的乘法复合评分,促进上下文准备好的令牌优于孤立候选,同时抑制低置信度的位置信息噪声,采用无块操作以保持高准备度锚点的全局可访问性。 extit{ours} 进一步应用架构感知校准,以处理由多样的视觉整合策略引起的置信度异质性。在三个 dMLLMs 上进行的四个基准实验表明,与原始方法相比, extit{ours} 在质量提升和幻觉减少方面表现出一致的改进,在多个较长生成设置中获得了更大的增益,突显了邻居支持和视觉整合策略在未来 dMLLM 解码方法设计中的重要性。我们的代码可在 https://github.com/zhaoyk-sysu/CACD-dMLLM 上公开获取。
cs.CL / 60 / 2608.22376

Can Large Language Models "Hyper-Thread"?

大型语言模型能否实现“超线程”?
Ding, Fei
Abstract
Large language models generate tokens sequentially, but can they execute multiple tasks concurrently while forming each token? Broader attention allocation may provide a mechanism for such task concurrency. Existing approaches to scaling inference primarily rely on longer generations, more samples, or additional verification stages, while attention dispersion is often treated as a signal of interference or error. Task concurrency within serial generation therefore remains underexplored. We propose the Model Hyper-Threading Hypothesis and evaluate its predictions using multiple coordinated tasks that share state within the same problem. We design three conditions (Baseline, Serial Functional Scheduling, and Concurrent Functional Loading) and evaluate their benefits and costs using accuracy, output-token distributions, and attention metrics. On an AIME 2025 development set, Concurrent Functional Loading achieves the highest accuracy. Relative to Serial Functional Scheduling, its typical output length is similar and it is shorter on most problems, while exhibiting greater attention dispersion and higher task-relevant coverage, albeit with a heavier output-length tail. Within-step concurrency and its causal mechanism still require direct tests. Our results show that more dispersed attention can coexist with higher accuracy, providing preliminary behavioral and correlational evidence for the hyper-threading hypothesis. These findings motivate a shift in perspective on inference scaling from "generating more tokens" toward "having each generation step carry more tasks," pointing to a new avenue for improving reasoning performance.
Chinese Translation
大型语言模型按顺序生成标记,但它们能否在生成每个标记的同时并发执行多个任务?更广泛的注意力分配可能为这种任务并发提供了一种机制。现有的推理扩展方法主要依赖于更长的生成、更大的样本量或额外的验证阶段,而注意力分散通常被视为干扰或错误的信号。因此,串行生成中的任务并发仍然未被充分探索。我们提出了模型超线程假设,并使用在同一问题中共享状态的多个协调任务来评估其预测。我们设计了三种条件(基线、串行功能调度和并发功能加载),并使用准确性、输出标记分布和注意力指标评估它们的收益和成本。在AIME 2025开发集中,并发功能加载实现了最高的准确性。与串行功能调度相比,其典型输出长度相似,并且在大多数问题上更短,同时表现出更大的注意力分散和更高的任务相关覆盖率,尽管输出长度尾部较重。步骤内的并发及其因果机制仍需直接测试。我们的结果表明,更分散的注意力可以与更高的准确性共存,为超线程假设提供了初步的行为和相关证据。这些发现促使我们重新审视推理扩展的视角,从“生成更多标记”转向“让每个生成步骤承担更多任务”,指向改善推理性能的新途径。
cs.CL / 61 / 2608.22388

ProBel: Propaganda Detection with Techniques, Spans, and Explanations

ProBel:利用技术、跨度和解释进行宣传检测
Kmainasi, Mohamed Bayan, Shahroor, Ali Ezzat, Sartori, Elisa, Martino, Giovanni Da San, Alam, Firoj
Abstract
Propaganda detection includes several related prediction levels, ranging from sentence-level decisions to technique classification and span identification. However, it remains unclear how supervision at these levels interacts when learned jointly across Arabic and English. We present ProBel, an Arabic and English resource that aligns binary labels, multi-label annotations over 23 propaganda techniques grouped into six coarse categories, technique-labeled spans, and reference explanations for the same news sentences. It includes a substantially larger English collection and supports matched binary, coarse-grained, multi-label, and span-level tasks in both languages. We evaluate zero-shot prompting, task-specific fine-tuning, and joint training under a shared setup. A single bilingual multi-task model achieves the best overall performance and remains competitive across tasks and languages. Cross-task analysis shows that transfer depends on the supervision level. Joint classification training preserves binary performance, whereas span-only training can weaken sentence-level prediction. Joint bilingual training yields the most stable results, while monolingual fine-tuning can reduce transfer to the other language. We will release the data, code, and evaluation scripts.
Chinese Translation
宣传检测包括多个相关的预测层次,从句子级决策到技术分类和跨度识别。然而,目前尚不清楚这些层次的监督在阿拉伯语和英语中共同学习时如何相互作用。我们提出了ProBel,这是一个阿拉伯语和英语资源,整合了二元标签、对23种宣传技术的多标签注释(这些技术被分为六个粗略类别)、技术标记的跨度以及相同新闻句子的参考解释。该资源包含了显著更大的英语集合,并支持两种语言中的匹配二元、粗粒度、多标签和跨度级任务。我们评估了零样本提示、特定任务的微调和在共享设置下的联合训练。一个单一的双语多任务模型在整体表现上达到了最佳效果,并在各个任务和语言中保持竞争力。跨任务分析表明,迁移依赖于监督层次。联合分类训练保持了二元性能,而仅进行跨度训练可能会削弱句子级预测。联合双语训练产生了最稳定的结果,而单语微调可能会减少对另一种语言的迁移。我们将发布数据、代码和评估脚本。
cs.CL / 62 / 2608.22390

SchemaGUI: A Schema-Driven Benchmark for Controllable GUI Generation Evaluation

SchemaGUI:一种基于模式的可控GUI生成评估基准
Dong, Jiarui, Cai, Yin, Gu, Zhouhong, Wu, Chenmou, Tao, Ci, Chen, Yiran, Li, Jialing, Shi, Xiaoran, Zhang, Juntao, Fang, Zhijun
Abstract
Large language models (LLMs) have demonstrated strong potential in graphical user interface (GUI) generation, but reliable evaluation remains challenging due to uncontrolled data distributions, noisy annotations, and limited layout scenario coverage. To address this, we propose SchemaGUI, a template-based benchmark for controllable GUI generation evaluation. By synthesizing paired natural language instructions and deterministic function-call references from parameterized interface schemas, SchemaGUI can generate thousands of deterministically annotated tasks in seconds without human labeling. Based on 1,000 evaluated instances per scenario and language across six representative bilingual scenarios, we benchmark five mainstream models, including the Qwen3.5 family, Qwen3-Coder-30B, and DeepSeek-R1. Our extensive analysis reveals three key insights. First, precise geometric spatial control remains an important bottleneck; while scaling Qwen3.5 from 4B to 27B improves Schema Feasibility from 91.56% to 99.63%, the Geometry score improves more modestly (from 67.05% to 75.30%). Second, generation difficulty is highly sensitive to layout complexity, with current LLMs excelling at simple sequential arrangements but suffering severe coordinate drift in dense grids and multi-region compositions. Third, thinking mode increases token consumption while generally reducing GUI Score, particularly for smaller models.
Chinese Translation
大型语言模型(LLMs)在图形用户界面(GUI)生成方面展现了强大的潜力,但由于数据分布不受控制、注释噪声以及布局场景覆盖有限,可靠的评估仍然具有挑战性。为了解决这一问题,我们提出了SchemaGUI,这是一种基于模板的可控GUI生成评估基准。通过从参数化接口模式中合成配对的自然语言指令和确定性的函数调用引用,SchemaGUI能够在几秒钟内生成数千个确定性注释的任务,而无需人工标注。基于六个具有代表性的双语场景中每个场景和语言评估的1,000个实例,我们对包括Qwen3.5系列、Qwen3-Coder-30B和DeepSeek-R1在内的五个主流模型进行了基准测试。我们的广泛分析揭示了三个关键见解。首先,精确的几何空间控制仍然是一个重要瓶颈;尽管将Qwen3.5从4B扩展到27B使得模式可行性从91.56%提高到99.63%,但几何得分的提升相对温和(从67.05%提高到75.30%)。其次,生成难度对布局复杂性高度敏感,当前的LLMs在简单的顺序排列中表现出色,但在密集网格和多区域组合中遭遇严重的坐标漂移。第三,思维模式增加了令牌消耗,同时通常降低了GUI得分,特别是对于较小的模型。
cs.CL / 63 / 2608.22411

Don' t Box Me In: Dynamic Cultural Adaptation and Cognitive Tracking for Social Understanding

不要限制我:动态文化适应与认知追踪在社会理解中的应用
Dai, Chongyuan, Shen, Yaling, Tang, Shengeng, Ma, Hui, Hu, Jinpeng
Abstract
Social interaction increasingly takes place in multicultural settings, where individuals may draw on multiple cultural influences and adapt their communicative behavior across contexts. Despite recent advances in equipping Large Language Models (LLMs) with social understanding capabilities, existing approaches often model culture as a static demographic attribute, limiting their ability to accommodate hybrid and dynamically expressed communicative preferences. Therefore, in this paper, we propose \textbf{DyCAC}, a training-free framework that achieves fluid social alignment by incorporating \underline{Dy}namic \underline{C}ultural \underline{A}daptation with continuous \underline{C}ognitive tracking. Rather than inferring a fixed cultural identity, DyCAC models culturally relevant communicative preferences as a time-varying mixture of population-level cultural reference profiles. This reference-based representation is further calibrated using dialogue-style signals observed in the ongoing interaction, enabling the model to capture both composite cultural influences and turn-level shifts in communicative behavior. In parallel, a memory module driven by Theory of Mind (ToM) continuously tracks the cognitive states of the interlocutor. Extensive experiments on interactive social and cultural benchmarks demonstrate the superiority of our approach. The proposed framework outperforms existing baselines, exhibiting enhanced social intelligence and broad adaptability across varied multicultural contexts.
Chinese Translation
社会互动越来越多地发生在多文化环境中,个体可能会借鉴多种文化影响,并在不同情境中调整其沟通行为。尽管近期在赋予大型语言模型(Large Language Models, LLMs)社会理解能力方面取得了进展,现有方法往往将文化建模为静态的人口特征,这限制了它们适应混合和动态表达的沟通偏好的能力。因此,本文提出了 extbf{DyCAC},一个无训练框架,通过结合 extunderline{动}态 extunderline{文}化 extunderline{适}应与持续的 extunderline{认}知追踪,实现流畅的社会对齐。DyCAC并不是推断固定的文化身份,而是将文化相关的沟通偏好建模为随时间变化的人群文化参考特征的混合。这种基于参考的表征进一步通过在持续互动中观察到的对话风格信号进行校准,使模型能够捕捉到复合文化影响和沟通行为的转变。在此过程中,基于心智理论(Theory of Mind, ToM)的记忆模块持续追踪对话者的认知状态。在互动社会和文化基准上的广泛实验表明了我们方法的优越性。所提出的框架在多种多文化环境中表现出更强的社会智能和广泛的适应性,超越了现有基准。
cs.CL / 64 / 2608.22432

Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator

多语言 LLM 评审中的排名逆转:一种无标签的双中心校准器
Mahmood, Alhasan, Abdaljalil, Samir, Kurban, Hasan
Abstract
Multilingual LLM judges produce different evaluator-backbone rankings depending on the prompt language: on an eight-language Agent-as-a-Judge benchmark, the top-ranked backbone alternates across English, Arabic, Chinese, Hindi, Japanese, Spanish, Turkish, and Swahili, and 7 of 15 backbone pairs show statistically significant pairwise rank reversal. We treat this as a measurement problem. The multilingual judge score decomposes additively into task difficulty, backbone skill, and a language-backbone interaction term, the last of which is recoverable without human labels by double-centering the cell-mean score matrix. We make this estimator (\textbf{Consensus-Based Calibration}, CBC) explicit, give an $O(1/\sqrt{n})$ finite-sample concentration bound with variance constant $(1-\tfrac{1}{m})(1-\tfrac{1}{k})$, and show that it is unbiased even when task-language interactions are present. Across 7{,}920 judge runs (6 backbones, 8 languages, 55 tasks, 3 frameworks), CBC raises held-out cross-task rank consistency $\tau$ from 0.650 to 0.902 and agrees with the held-out additive-model oracle in 100\% of per-language decisions versus 68.5\% raw; these are consistency diagnostics, not human-grounded correctness measures. On a separately collected M-RewardBench panel (7 languages, 1{,}500 items per language, 10{,}500 language-item instances, 5 evaluators), panel agreement with the public human gold preferences rises from 68.7\% to 76.6\% (gain 7.9 percentage points, 95\% CI $[6.0, 9.9]$), our strongest external evidence of downstream usefulness. The estimator is the standard two-way ANOVA interaction-recovery operation under sum-to-zero contrasts; our contribution is its application as a label-free post-hoc calibrator for multilingual LLM judges, an explicit finite-sample concentration bound, and an unbiasedness result that holds even under task-language misspecification.
Chinese Translation
多语言 LLM 评审根据提示语言产生不同的评估者基础排名:在一个包含八种语言的代理评审基准中,排名最高的基础在英语、阿拉伯语、中文、印地语、日语、西班牙语、土耳其语和斯瓦希里语之间交替,15对基础中有7对显示出统计显著的成对排名逆转。我们将此视为一个测量问题。多语言评审分数可加性地分解为任务难度、基础技能和语言-基础交互项,最后一个项可以通过对单元均值分数矩阵进行双中心化而在没有人工标签的情况下恢复。我们明确提出这一估计器( extbf{基于共识的校准},CBC),给出了一个 $O(1/ extsqrt{n})$ 的有限样本集中度界限,方差常数为 $(1- frac{1}{m})(1- frac{1}{k})$,并表明即使在任务-语言交互存在时它也是无偏的。在 7,920 次评审运行(6 个基础,8 种语言,55 个任务,3 个框架)中,CBC 将保留的跨任务排名一致性 $ au$ 从 0.650 提高到 0.902,并且在每种语言的决策中与保留的加性模型 oracle 达成 100\% 的一致,而原始一致性仅为 68.5\%;这些是一致性诊断,而不是基于人类的正确性度量。在单独收集的 M-RewardBench 面板(7 种语言,每种语言 1,500 个项目,10,500 个语言-项目实例,5 名评估者)中,面板与公共人类黄金偏好的一致性从 68.7\% 上升到 76.6\\%(增加 7.9 个百分点,95\\% CI $[6.0, 9.9]$),这是我们对下游实用性的最强外部证据。该估计器是标准的双向 ANOVA 交互恢复操作,适用于和为零的对比;我们的贡献在于将其应用于多语言 LLM 评审的无标签后期校准,明确的有限样本集中度界限,以及即使在任务-语言错误指定下也成立的无偏性结果。
cs.CL / 65 / 2608.22444

Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations

独立对齐,集体失调:预测大型语言模型代理群体中的对抗性捕获
Magistrali, Isotta, Shani, Chen
Abstract
The unit of AI safety evaluation is still the individual model, yet language-model agents are increasingly deployed in interacting populations that read and write one another's decisions. This raises a question no single-agent audit can answer: an agent that is well-calibrated on its own may still be pulled toward a different decision by the agents around it. We study this on a security-triage task, where populations of language-model monitors decide whether to escalate or dismiss alerts, and into which we can inject a committed minority that always pushes one way. We find that two alerts a single agent judges almost identically on its own can drive collective behavior far apart, so auditing any one member need not reveal what the population will do. Yet that collective behavior can be predicted in advance. From a population's benign, adversary-free operation alone, we calibrate a response function that forecasts, before any attack is run, how far a committed minority will later move it. We then ask what shifts the outcome and find that letting agents see each other's reasoning neutralizes a weak attack, while only delaying it against a strong one, turning the question from whether the population converges on the adversaries' choice into when. Finally, we exclude the hypothesis of capture being an irreversible trap: once the committed agents are removed, the population drifts back toward where it began, so capture is a temporary state. Alignment in isolation is not alignment in a population, yet what a population will do under attack can be read in advance, from how it behaves before any adversary arrives.
Chinese Translation
人工智能安全评估的单位仍然是单个模型,但语言模型代理越来越多地在相互影响的群体中部署,这些群体会读取和写入彼此的决策。这引发了一个单一代理审计无法回答的问题:一个在自身上经过良好校准的代理,仍然可能受到周围代理的影响而做出不同的决策。我们在一个安全分流任务中研究这一现象,在该任务中,语言模型监控者群体决定是升级还是驳回警报,并且我们可以注入一个始终朝一个方向推动的坚定少数群体。我们发现,单个代理几乎在自身上对两个警报的判断可以驱动集体行为大相径庭,因此审计任何一个成员并不一定能揭示群体的行为。然而,这种集体行为可以提前预测。仅凭群体的良性、无对手的操作,我们校准了一个响应函数,该函数在任何攻击发生之前预测坚定少数群体将如何推动群体行为的变化。然后我们探讨了什么因素会改变结果,发现让代理看到彼此的推理可以中和弱攻击,而仅仅是延迟对强攻击的反应,这将问题从群体是否会趋同于对手的选择转变为何时会趋同。最后,我们排除了捕获是不可逆陷阱的假设:一旦坚定代理被移除,群体会重新回到最初的位置,因此捕获是一个暂时状态。孤立中的对齐并不等同于群体中的对齐,但在攻击下群体的行为可以提前预测,基于其在任何对手到达之前的表现。
cs.CL / 66 / 2608.22446

Figurative Justice: Detecting metaphors in Hindi judgements with qualitative assessment and transformers

比喻正义:通过定性评估和变换器检测印地语判决中的比喻
Bhattacharyya, Bhumika, Guha, Shouvik Kumar, Dutta, Indranil
Abstract
Metaphors are figurative use of words for conceptual mapping. Metaphor detection in the legal context has been crucial as metaphors are persuasive juridical means of creating legal meaning and concepts resulting in significant consequences. Metaphorical framing in legal discourse by judges, lawyers, and legislators brings about real-time implications upon individuals and influences judicial decision-making, argumentation and interpretation of laws. This is crucial in Human Rights infringement cases where language determines severity of punishment, public perception and judicial outcomes. While automatic metaphor detection in major languages like English, Spanish, Polish, Lithuanian have aided in understanding inherent intentions of metaphorical use of language, there is no such attempt in low-resource languages like Hindi. The dearth of annotated legal corpora in Hindi makes it difficult to develop NLP models and detect metaphors in judicial proceedings. In the Indian context, Convolutional Neural Networks (CNNs) have been used for classification of bail judgements, however there are no existing models designed for metaphor detection. We present a Hindi Legal Metaphor Corpus (HiLeMe) by isolating judgements from Hindi Legal Data Corpus (HLDC). Legal experts annotated HiLeMe to classify metaphorical constructions using the MIPVU schema. We downstreamed an mBERT on Hindi legal metaphor detection task. We built a transformer-based architecture for metaphor detection that are known to outperform traditional models in legal classification tasks. This model provides insights into the judicial psyche for decoding judicial decisions. Our research contributes to advancing automated models in legal discourse in low-resource languages like Hindi and envisages adoption into 22 Indian schedule languages.
Chinese Translation
比喻是将词语进行概念映射的比喻性用法。在法律背景下,比喻的检测至关重要,因为比喻是创造法律意义和概念的有说服力的法律手段,产生显著后果。法官、律师和立法者在法律话语中使用比喻框架,对个人产生实时影响,并影响司法决策、论证和法律解释。这在侵犯人权案件中尤为重要,因为语言决定了惩罚的严重性、公众认知和司法结果。虽然在英语、西班牙语、波兰语、立陶宛语等主要语言中,自动比喻检测有助于理解比喻性语言使用的内在意图,但在像印地语这样的低资源语言中尚未有类似尝试。印地语中缺乏注释的法律语料库使得开发自然语言处理(NLP)模型和在司法程序中检测比喻变得困难。在印度背景下,卷积神经网络(CNNs)已被用于保释判决的分类,但尚无专门设计用于比喻检测的模型。我们通过从印地语法律数据语料库(HLDC)中提取判决,提出了印地语法律比喻语料库(HiLeMe)。法律专家使用MIPVU框架对HiLeMe进行了注释,以分类比喻性构造。我们在印地语法律比喻检测任务上下游了一个mBERT模型。我们构建了一种基于变换器的比喻检测架构,该架构在法律分类任务中已知优于传统模型。该模型为解码司法决策提供了对司法心理的洞察。我们的研究有助于推动低资源语言(如印地语)法律话语中的自动化模型发展,并展望在22种印度法定语言中的应用。
cs.CL / 67 / 2608.22452

From Exposure to Expectation: Frequency, Surprisal, and Language Across Development in Spanish

从暴露到期望:西班牙语发展中的频率、惊讶度与语言
López, Francisco Portillo
Abstract
Surprisal, the negative log-probability a language model assigns to a word given its preceding context, reliably predicts adult reading times. Does it contribute as much to explaining when children acquire individual words? Frequency reflects a learner's cumulative exposure to a word, whereas surprisal reflects how predictable a single occurrence is given its context. We investigate this question across two corpus-based studies of Spanish. In Study 1, we modeled age of acquisition (AoA) for 225 Spanish nouns using lexical frequency and contextual diversity from child-directed speech, plus surprisal from three language models differing in architecture and training language (BETO, BERTIN, mGPT). Frequency strongly predicted AoA (r=-.597, p<.001); surprisal added little beyond frequency and word length, including in a naturalistic-context analysis. In Study 2, we modeled adult fixation durations in the Chilean Spanish subsample of the Multilingual Eye-movement Corpus (MECO Wave 2), using mGPT surprisal alongside two independent frequency measures. Surprisal robustly predicted longer fixation durations after controlling for frequency and word length, consistent across both frequency sources. A matched word-type-level comparison showed the surprisal-behavior association was stronger in reading than in acquisition (z=3.63, p<.001). The findings suggest cumulative lexical exposure and contextual predictability play different roles across the language trajectory: frequency is particularly informative about when early lexical representations are acquired, whereas surprisal captures moment-to-moment processing difficulty in an already-established linguistic system. We discuss this pattern in relation to usage-based and entrenchment-based accounts of lexical development and to the evaluation of language models as models of human language behavior.
Chinese Translation
惊讶度是指语言模型在给定前文上下文的情况下对一个词分配的负对数概率,它可靠地预测成人的阅读时间。它在解释儿童何时习得单个词汇方面的贡献是否同样重要?频率反映了学习者对一个词的累积暴露,而惊讶度则反映了在特定上下文中单次出现的可预测性。我们通过两项基于语料库的西班牙语研究来探讨这个问题。在研究1中,我们使用来自儿童导向语料的词汇频率和上下文多样性,以及来自三种不同架构和训练语言的语言模型(BETO、BERTIN、mGPT)的惊讶度,建模了225个西班牙名词的习得年龄(AoA)。频率强烈预测了AoA(r=-.597,p<.001);惊讶度在频率和词长之外几乎没有提供额外的信息,包括在自然语境分析中。在研究2中,我们使用mGPT惊讶度和两个独立的频率测量,建模了多语言眼动语料库(MECO Wave 2)中智利西班牙语子样本的成人注视持续时间。在控制频率和词长后,惊讶度稳健地预测了更长的注视持续时间,这一结果在两个频率来源中一致。匹配的词类水平比较显示,惊讶度与行为的关联在阅读中比在习得中更强(z=3.63,p<.001)。研究结果表明,累积的词汇暴露和上下文可预测性在语言发展轨迹中扮演着不同的角色:频率特别能提供关于早期词汇表征习得时间的信息,而惊讶度则捕捉到了在已建立的语言系统中瞬时处理的困难。我们讨论了这一模式与基于使用和基于巩固的词汇发展理论,以及将语言模型评估为人类语言行为模型的关系。
cs.CL / 68 / 2608.22479

GTA-RAG: Graph-Trajectory-Augmented Reinforcement Learning for Multi-Turn Retrieval-Augmented Reasoning

GTA-RAG:用于多轮检索增强推理的图轨迹增强强化学习
Chen, Jun, Liu, Yongchao, Qiu, Pengyu, Zheng, Jiajun, Zhang, Juelu, Zeng, Yujie, Zhang, Qin, Qiao, Ziyue, Luo, Xiao
Abstract
Retrieval-augmented generation (RAG) enables LLMs to access external knowledge for answering knowledge-intensive questions. For complex multi-hop questions, multi-turn retrieval-augmented reasoning extends RAG into an iterative process that repeatedly searches for and integrates evidence across documents. However, existing reinforcement-learning (RL) approaches for agentic RAG are typically optimized with final-answer rewards, which provide sparse supervision and overlook whether the model actually retrieves the required evidence chain. We present \textsc{GTA-RAG}, a graph-trajectory-augmented RL framework for multi-turn retrieval-augmented reasoning. From an entity--document graph, we sample connected document paths, synthesize multi-hop QA trajectories, and validate them with the deployed retriever to obtain executable trajectory-level supervision. We then optimize the retrieval policy with Group Relative Policy Optimization (GRPO) and a trajectory-guided reward that encourages both accurate answers and acquisition of target evidence documents, followed by answer-reward training on natural QA instances. Experiments on three multi-hop and two simple QA benchmarks show that \method{} consistently outperforms RL-based RAG baselines with both Qwen2.5-3B and Qwen2.5-7B backbones, while substantially improving evidence-chain coverage. Our code is available at https://github.com/cjcj46262/GTA-RAG.
Chinese Translation
检索增强生成(RAG)使大型语言模型(LLMs)能够访问外部知识,以回答知识密集型问题。对于复杂的多跳问题,多轮检索增强推理将RAG扩展为一个迭代过程,反复搜索和整合文档中的证据。然而,现有的用于代理RAG的强化学习(RL)方法通常通过最终答案奖励进行优化,这提供了稀疏的监督,并忽视了模型是否实际检索到所需的证据链。我们提出了 extsc{GTA-RAG},一种用于多轮检索增强推理的图轨迹增强RL框架。我们从实体-文档图中采样连接的文档路径,合成多跳问答轨迹,并通过部署的检索器对其进行验证,以获得可执行的轨迹级监督。然后,我们使用组相对策略优化(Group Relative Policy Optimization, GRPO)和一种轨迹引导奖励来优化检索策略,该奖励既鼓励准确答案,又促进目标证据文档的获取,随后在自然问答实例上进行答案奖励训练。在三个多跳和两个简单问答基准上的实验表明, extmethod{}在Qwen2.5-3B和Qwen2.5-7B基础上始终优于基于RL的RAG基线,同时显著提高了证据链的覆盖率。我们的代码可在https://github.com/cjcj46262/GTA-RAG获取。
cs.CL / 69 / 2608.22483

Claim-Level Confidence Calibration for Reliable Decision Making with Large Language Models

基于声明级置信度校准的大语言模型可靠决策支持
Abbasli, Toghrul, Toyoda, Kentaroh, Wang, Yuan, Chen, Li
Abstract
Large Language Models (LLMs) increasingly support decision-making in high-stakes domains, but they often hallucinate and express confidence that is misaligned with factual correctness. Response-level confidence is a coarse signal: a single generation can mix correct and incorrect statements, so a single number is not actionable for users that must accept, reject, or verify individual pieces of information. We study claim-level confidence calibration as a decision-relevant uncertainty signal: each response is decomposed into atomic, verifiable claims, and each claim is assigned a calibrated confidence using inference-time signals from consistency across samples and self-verification. Our framework operates in closed-box settings (no logits, no fine-tuning) and applies post-hoc calibration directly at the claim level, enabling selective intervention such as evidence retrieval or human review for low-confidence claims. Across TriviaQA and TruthfulQA we evaluate seven baselines on six recent models (Llama-3.1, Mistral, Qwen2.5, DeepSeek-R1, GPT-4, GPT-4o), and show that claim-level decomposition combined with post-hoc calibration reduces expected calibration error on factual questions while exposing failure modes on adversarial false-premise questions where decision-makers most need reliable uncertainty estimates.
Chinese Translation
大型语言模型(LLMs)越来越多地支持高风险领域的决策制定,但它们常常会产生幻觉,并表现出与事实正确性不一致的置信度。响应级置信度是一个粗略的信号:单次生成可能混合正确和错误的陈述,因此单一的数值对于必须接受、拒绝或验证个别信息的用户来说并不可行。我们研究声明级置信度校准作为与决策相关的不确定性信号:每个响应被分解为原子、可验证的声明,并且每个声明使用来自样本一致性和自我验证的推理时信号分配校准置信度。我们的框架在封闭环境中运行(无 logits,无微调),并直接在声明级别应用事后校准,使得可以对低置信度声明进行选择性干预,例如证据检索或人工审核。在 TriviaQA 和 TruthfulQA 数据集上,我们对六个最新模型(Llama-3.1、Mistral、Qwen2.5、DeepSeek-R1、GPT-4、GPT-4o)评估了七个基线,并展示了声明级分解结合事后校准在事实问题上减少了预期校准误差,同时暴露了在对抗性错误前提问题上的失败模式,而决策者在这些情况下最需要可靠的不确定性估计。
cs.CL / 70 / 2608.22490

Who Pays More for Safety? Measuring the Disparate Cost of Safety Alignment across Languages

谁为安全支付更多?跨语言安全对齐成本的差异测量
Yoon, Chanwoong, Park, Jungsoo, Ritter, Alan
Abstract
Safety alignment helps models adhere to human values, but it often reduces response utility. We ask a critical but understudied question: Does safety alignment impose the cost equally across language groups? To answer this, we introduce a rigorous protocol to measure the utility loss imposed solely by safety alignment, which we term Safety Cost. Through direct pairwise comparisons between safety-aligned models and their unaligned counterparts, we find a systematic inequity: non-English users consistently bear a higher Safety Cost than English users. We further identify three underlying patterns. First, multiple languages lie in a double-penalty zone, experiencing both weaker safety protection and larger utility loss. Second, certain languages exhibit apparent utility gains that are in fact a consequence of safety filters failing to engage. Third, even high-resource languages pay a larger Safety Cost than English to reach the same level of safety. We show that these disparities arise from both explicit refusals and implicit qualitative differences across multiple dimensions. By accurately measuring the disparate effects of safety alignment, our findings expose a systematic disparity in current safety alignment practices.
Chinese Translation
安全对齐帮助模型遵循人类价值观,但往往会降低响应效用。我们提出一个重要但未被充分研究的问题:安全对齐是否对不同语言群体施加了相同的成本?为了解答这个问题,我们引入了一种严格的协议来测量仅由安全对齐所带来的效用损失,我们称之为安全成本(Safety Cost)。通过对安全对齐模型与其未对齐对应模型之间的直接成对比较,我们发现了一个系统性的不平等现象:非英语用户所承受的安全成本始终高于英语用户。我们进一步识别出三种潜在模式。首先,多种语言处于双重惩罚区,既经历了较弱的安全保护,又遭受了更大的效用损失。其次,某些语言表现出明显的效用增益,实际上是安全过滤器未能发挥作用的结果。第三,即使是高资源语言,为了达到相同的安全水平,也需要支付比英语更大的安全成本。我们表明,这些差异源于多个维度上的显性拒绝和隐性质量差异。通过准确测量安全对齐的不同影响,我们的研究揭示了当前安全对齐实践中的系统性差异。
cs.CL / 71 / 2608.22506

Kernel Token Contradiction: a Fast and Principled Approach for LLM Claim Uncertainty Quantification

核令牌矛盾:一种快速且有原则的LLM声明不确定性量化方法
Dentan, Jérémie, Canesse, Alexi, Sharkawy, Mahammed El, Vanier, Sonia
Abstract
Claim-level Uncertainty Quantification (UQ) aims to mitigate the lack of reliability of Large Language Models (LLMs) by evaluating the factuality of each claim in their outputs. We introduce Kernel Token Contradiction (KTC), a lightweight approach to compute claim-level UQ under realistic white-box conditions. KTC represents the candidate tokens involved in LLM generation as a positive semi-definite kernel that integrates both the LLM's conditional distribution and a token contradiction score. We then use the Von Neumann entropy to quantify the uncertainty of this kernel. To estimate token contradiction, we develop a new approach based on frequency statistics from the Wikipedia corpus. Although CPU-only, our approach achieves over an 8.2x speedup compared to state-of-the-art GPU-accelerated methods based on cross-encoders, and over a 65x speedup compared to CPU-only methods with comparable performance. Our evaluation spans two benchmarks across four European languages and 16 different models. KTC not only matches the average performance of existing methods but also outperforms them in high-precision regimes. This combination of computational efficiency and accuracy makes real-time monitoring of LLM outputs practical in production.
Chinese Translation
声明级不确定性量化(UQ)旨在通过评估大型语言模型(LLMs)输出中每个声明的事实性来减轻其可靠性不足的问题。我们提出了核令牌矛盾(KTC),这是一种轻量级的方法,用于在现实的白盒条件下计算声明级UQ。KTC将参与LLM生成的候选令牌表示为一个正半定核,该核整合了LLM的条件分布和令牌矛盾分数。然后,我们使用冯·诺依曼熵来量化该核的不确定性。为了估计令牌矛盾,我们开发了一种基于维基百科语料库的频率统计的新方法。尽管仅使用CPU,我们的方法相比于基于交叉编码器的最先进GPU加速方法实现了超过8.2倍的加速,相比于具有可比性能的CPU-only方法实现了超过65倍的加速。我们的评估涵盖了四种欧洲语言的两个基准和16种不同的模型。KTC不仅在平均性能上与现有方法相匹配,而且在高精度范围内超越了它们。这种计算效率与准确性的结合使得在生产环境中实时监控LLM输出成为可能。
cs.CL / 72 / 2608.22566

From Diagnosis to Redesign: Using Quantitative Ethnography to Improve Multi-Agent LLM Reasoning

从诊断到重设计:利用定量人种志改善多智能体大型语言模型推理
Khatri, Vedant, Cusimano, Anthony, Swiecki, Zachari, Xu, Zhen, Liu, Xiner, Yu, Renzhe
Abstract
Multi-agent large language model (LLM) systems are designed to improve reasoning by decomposing tasks across multiple agents with specialized functions, but the presence of multiple agents does not inherently guarantee coherent reasoning or outputs that align with task objectives. This paper introduces a quantitative ethnographic (QE) approach for diagnosing and redesigning multi-agent LLM systems based on the discourse produced through agent interactions. We test this approach using automated essay scoring as an example context, applying Epistemic Network Analysis (ENA) to model a five-agent multi-agent debate system and examine differences between debates that produced correct versus incorrect scoring decisions. Results show that, in the initial system, correct scoring decisions were characterized by rubric-grounded justification, agreement, and elaboration. Incorrect scoring decisions, in contrast, were characterized by extended proposition-challenge-response exchanges that were less consistently tied to rubric criteria. We then used the findings to revise the agents' prompts. The revised system improved exact scoring accuracy from 27.78% to 40.28% and shifted the discourse of incorrect debates toward the rubric-grounded pattern of correct ones, making the two nearly indistinguishable. Based on these results, we argue that QE can support a diagnostic-to-redesign loop for AI reasoning by tracing how patterns of agent interaction relate to system performance, informing prompt redesign, and evaluating whether those redesigns change both outcomes and interaction patterns.
Chinese Translation
多智能体大型语言模型(LLM)系统旨在通过将任务分解到具有专业功能的多个智能体之间来改善推理,但多个智能体的存在并不自动保证推理的一致性或输出与任务目标的一致性。本文介绍了一种基于智能体交互产生的论述进行诊断和重设计多智能体LLM系统的定量人种志(QE)方法。我们以自动化作文评分为例测试该方法,应用知识网络分析(Epistemic Network Analysis, ENA)对一个五智能体的辩论系统进行建模,并考察产生正确与错误评分决策的辩论之间的差异。结果表明,在初始系统中,正确的评分决策以基于评分标准的理由、共识和详细阐述为特征。相比之下,错误的评分决策则表现为延长的命题-挑战-回应交流,这些交流与评分标准的关联性较低。随后,我们利用这些发现修订了智能体的提示。修订后的系统将准确评分率从27.78%提高到40.28%,并使错误辩论的论述向正确辩论的基于评分标准的模式转变,使两者几乎无法区分。基于这些结果,我们认为QE可以通过追踪智能体交互模式与系统性能之间的关系,支持AI推理的诊断到重设计循环,从而为提示重设计提供信息,并评估这些重设计是否改变了结果和交互模式。
cs.CL / 73 / 2608.22582

Hybrid Panels: Toward Human-AI Collaboration in Survey Research

混合面板:朝着人类与人工智能在调查研究中的协作
Romberg, Julia, Gummer, Tobias, Lapesa, Gabriella, Kunz, Tanja, Wagner, Claudia
Abstract
Large-scale population surveys are essential for generating robust social and scientific insights, yet they face significant challenges, including declining response rates, increasing data collection costs, long delays between data collection and data provision, and the risk of nonresponse bias. Advances in artificial intelligence (AI) have opened up new opportunities for AI-supported survey infrastructures where the goal is to overcome these challenges without limiting the data quality. A promising AI-enabled survey infrastructure for which we build a first pilot is a hybrid panel. A hybrid panel is a longitudinal AI-enabled survey which allows to iteratively improve the alignment between large language models (LLMs) and the population they aim to simulate and use the errors to inform the design and implementation of the next survey wave (e.g., inform the participant recruitment, assignment of questions to participants). It incorporates both human participants and LLMs as fundamental elements of its design. In this research note, we introduce the concept of a hybrid panel by providing a definition and outlining an overarching framework, spanning data collection to data validation. We detail results from a first pilot study to illustrate (open) challenges that we identify for hybrid panels.
Chinese Translation
大规模人口调查对于生成稳健的社会和科学洞察至关重要,但面临着显著挑战,包括响应率下降、数据收集成本增加、数据收集与数据提供之间的长时间延迟以及非响应偏差的风险。人工智能(AI)的进步为支持AI的调查基础设施开辟了新的机会,其目标是在不降低数据质量的情况下克服这些挑战。我们构建的第一个试点是一个有前景的AI驱动调查基础设施,即混合面板。混合面板是一种纵向的AI驱动调查,允许在大型语言模型(LLMs)与其旨在模拟的人群之间进行迭代性改进,并利用错误来指导下一轮调查的设计和实施(例如,指导参与者招募、问题分配给参与者)。它将人类参与者和LLMs作为设计的基本要素。在这篇研究笔记中,我们通过提供定义和概述一个涵盖从数据收集到数据验证的总体框架,引入混合面板的概念。我们详细介绍了第一次试点研究的结果,以说明我们为混合面板识别的(开放)挑战。
cs.CL / 74 / 2608.22622

Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains

通过与OMOP对齐的检索教导大型语言模型(LLMs)如何理解重症监护医师的临床推理,提升各临床领域的推理能力
Contreras, Miguel, Siegel, Scott, Nerella, Subhash, Sena, Jessica, Zhang, Jiaqing, Sun, Heng, Akkaladevi, Hruday Tej, Lu, Peiyu, Rosen, Jordan, Kapoor, Sumit, Desaraju, Sasank, Thompson, Grace R., Purcell, Jacob, Petrauskis, Michael, Hong, Philip KW., Brennan, Meghan, Chrabaszcz, Sarah, Smith, Tierra, Ren, Ronnie, Kabbash, Michel S., Haziroglu, Ceyhun, Patel, Rushi, Gomez, Gabriel, Chaiklin, Charlotte, Leung, Randy, John, Kenneth N., Wiggins, Whitman, Kayser, Philip, Bird, Vincent, Bruzzone, Maria, Loftus, Tyler J., Bihorac, Azra, Rashidi, Parisa
Abstract
Clinical decision-making relies on identifying relevant patient information to guide diagnosis and treatment, a challenge that is especially difficult in the data-dense and rapidly changing intensive care unit (ICU). Large language models (LLMs) could support this task. However, existing applications and datasets mostly emphasize surface-level retrieval or factual recall rather than the inductive and deductive reasoning clinicians practice to select and reason over decision-relevant evidence. We hypothesized that training LLMs on expert ICU reasoning could yield clinical reasoning skills that generalize beyond critical care. Here we introduce ICU-REACT, a reasoning dataset developed with 19 clinicians through a clinician-in-the-loop framework to teach LLMs to perform information retrieval and context-aware clinical reasoning in the ICU. Using ICU-REACT, we fine-tuned Clin-REACT models spanning 8B-70B parameters and three model families. Across five clinical reasoning benchmarks, Clin-REACT consistently outperformed its backbone models and open-source general-purpose and medical LLMs. Gains extended to different tasks including script concordance tests, and downstream diagnosis and treatment tasks. These findings suggest that expert reasoning supervision in critical care can improve broader clinical reasoning, although prospective evaluation is needed before real-world clinical use.
Chinese Translation
临床决策依赖于识别相关的患者信息以指导诊断和治疗,这在数据密集且快速变化的重症监护室(ICU)中尤其具有挑战性。大型语言模型(LLMs)可能支持这一任务。然而,现有的应用和数据集大多强调表层检索或事实回忆,而非临床医生在选择和推理决策相关证据时所实践的归纳和演绎推理。我们假设在专家ICU推理上训练LLMs可以产生超越重症护理的临床推理技能。在此,我们介绍了ICU-REACT,这是一个通过临床医生参与的框架与19名临床医生共同开发的推理数据集,旨在教导LLMs在ICU中执行信息检索和上下文感知的临床推理。利用ICU-REACT,我们微调了涵盖8B-70B参数和三个模型家族的Clin-REACT模型。在五个临床推理基准测试中,Clin-REACT始终优于其基础模型以及开源的通用和医学LLMs。提升效果扩展到不同任务,包括脚本一致性测试以及下游的诊断和治疗任务。这些发现表明,在重症护理中进行专家推理监督可以改善更广泛的临床推理,尽管在实际临床使用之前仍需进行前瞻性评估。
cs.CL / 75 / 2608.22634

GeoRisk-RAG: A Hierarchy-Aware Risk Framework for Improving RAG Reliability through Selective Answering

GeoRisk-RAG:一种层次感知风险框架,通过选择性回答提高RAG的可靠性
Ravi, Meenu, Sarkar, Shailik, AlKulaib, Lulwah, Tessema, Yordanos, Lu, Chang-Tien
Abstract
Current work on improving reliability in large language model (LLM)- generated answers has primarily leveraged Retrieval-Augmented Generation (RAG), knowledge-graph augmentation, and reinforcement learning. While these methods are adept at enhancing and measuring reliability through semantic similarity and faithfulness, they often struggle to distinguish semantic similarity from geographic validity. This is especially critical in natural hazard management domains where geographic granularity (i.e., town vs. city vs. state) is significant for decision-making, as responses valid in one municipality may not transfer to another. In such domains, a confidently wrong answer carries greater risk than abstaining. We present GeoRisk-RAG, a novel hierarchy-aware framework that addresses this geographic-validity gap through selective answering. This framework explicitly estimates geographic applicability using a Directed Acyclic Graph (DAG)-based distance for context retrieval before response generation. Experiments on a novel held-out wildfire-related question-answering (QA) dataset show that GeoRisk-RAG significantly reduces false confidence rates for location-dependent questions, lowering the rate to 0.009 compared with ~0.090 for standard semantic similarity and reranking baselines, while consistently achieving higher human preference alignment. This work provides a more comprehensive assessment of end-to-end RAG pipelines by integrating geographic validity and selective-answering behavior for safer decision-making in geospatial domains.
Chinese Translation
当前在提高大型语言模型(LLM)生成答案的可靠性方面的研究,主要利用了检索增强生成(RAG)、知识图谱增强和强化学习。尽管这些方法在通过语义相似性和忠实度来增强和衡量可靠性方面表现出色,但它们往往难以区分语义相似性与地理有效性。在自然灾害管理领域,这一点尤为重要,因为地理粒度(即城镇、城市与州)对决策具有重要意义,因为在一个市政当局有效的回答可能不适用于另一个。在这样的领域中,自信错误的回答带来的风险往往大于不作回答。我们提出了GeoRisk-RAG,这是一种新颖的层次感知框架,通过选择性回答解决地理有效性差距。该框架在生成回答之前,使用基于有向无环图(DAG)的距离来明确估计地理适用性,以进行上下文检索。在一个新的野火相关问答(QA)数据集上的实验表明,GeoRisk-RAG显著降低了与位置相关问题的虚假自信率,将其降低至0.009,而标准语义相似性和重排序基线的虚假自信率约为0.090,同时在与人类偏好一致性方面始终表现更高。这项工作通过整合地理有效性和选择性回答行为,为地理空间领域的安全决策提供了更全面的端到端RAG管道评估。
cs.CL / 76 / 2608.22651

Iteration Without Elaboration: A Simple ReAct Architecture Suffices for Text-to-SQL Generation

无 elaboration 的迭代:简单的 ReAct 架构足以用于文本到 SQL 的生成
Lu, Jian, Yu, Haiwei, Xiong, Raymond M, Zhang, Anru, Zhuo, Danyang
Abstract
Modern text-to-SQL systems have become increasingly elaborate, relying on schema-linking modules, retrieval-augmented prompting, candidate generation, and multi-stage refinement pipelines. While effective, these additions introduce substantial latency and engineering overhead. To this end, we present \textbf{ReAct-SQL}, a simple yet effective zero-shot ReAct-style framework built solely on iterative reasoning and a constrained action space defined by a typed Domain-Specific Language (DSL) of 15 relational operations, rather than free-form SQL generation. The model incrementally issues DSL calls, observes compiled-SQL execution feedback, and revises its reasoning through interaction. On corrected BIRD mini-dev and EHR-SQL, ReAct-SQL achieves \textbf{84.5\%} and \textbf{73.9\%} accuracy, respectively, matching substantially more elaborate baselines while running up to $8\times$ faster. Incremental ablations further show that iteration primarily improves grounding, while the DSL improves compositional reliability.
Chinese Translation
现代文本到 SQL 系统变得越来越复杂,依赖于模式链接模块、增强检索提示、候选生成和多阶段精炼流程。尽管这些附加功能有效,但也引入了显著的延迟和工程开销。为此,我们提出了 extbf{ReAct-SQL},这是一个简单而有效的零-shot ReAct 风格框架,完全基于迭代推理和由 15 种关系操作定义的受限动作空间(类型特定语言,Domain-Specific Language, DSL),而不是自由形式的 SQL 生成。该模型逐步发出 DSL 调用,观察编译 SQL 执行反馈,并通过交互修正其推理。在修正后的 BIRD mini-dev 和 EHR-SQL 数据集上,ReAct-SQL 分别达到了 extbf{84.5\%} 和 extbf{73.9\\%} 的准确率,显著匹配更复杂的基线,同时运行速度快达 $8 imes$。增量消融实验进一步表明,迭代主要改善了基础,DSL 则提高了组合可靠性。
cs.CL / 77 / 2608.22695

Enrich-Retrieve-Rank: Scaling Capability Discovery Beyond In-Context Routing

Enrich-Retrieve-Rank:超越上下文路由的能力发现扩展方法
Sorathiya, Nazib, Zhang, Daniel, Akhbari, Bardiya
Abstract
Agent ecosystems now include thousands of MATS components (Models, Agents, Tools, and Skills), yet their discovery still relies on in-context routing. These systems read a registry (names, hints, or descriptions, as context budget permits), pick a candidate, invoke it, and retry on failure. This pattern degrades with scale, and registries are growing fast. We recast capability discovery as search over a registry by defining an offline enrichment step that turns sparse metadata into searchable profiles, and an online retrieve-then-rank pipeline that returns a ranked shortlist without invoking any candidates online. We show that from N=10 to 7,278 capabilities, in-context routing's top-1 accuracy (Match@1) collapses (0.85 to 0.12), while retrieve-then-rank degrades more gently (0.81 to 0.39) because its reranker still ranks the right capability first 0.70-0.87 of the time once retrieval finds it. In the Nova Micro sweep, the crossover is around N=500. We compare against two in-context baselines. Full-Ctx puts the whole registry in the prompt and asks the LLM to pick. Search&Pick gives the LLM a search tool to narrow candidates before it picks. At full scale the pipeline leads Search&Pick by 6.5 percentage points (pp) on Match@1 at about half the cost. It reduces cost 70x versus Full-Ctx. We use a fixed configuration (same enrichment, retriever, and scorer weights) across agent, tool, and skill registries. The pipeline runs in production as the default capability-discovery layer of a large-scale multi-agent platform.
Chinese Translation
智能体生态系统现包含数千个MATS组件(模型、智能体、工具和技能),但其发现仍依赖于上下文路由。这些系统读取注册表(根据上下文预算读取名称、提示或描述),选择候选项,调用该候选项,失败后重试。该模式在规模扩大时性能下降,而注册表规模增长迅速。我们将能力发现重新定义为对注册表的搜索,设计了一个离线丰富步骤,将稀疏的元数据转化为可搜索的配置文件,以及一个在线检索-排序流水线,能在不在线调用任何候选项的情况下返回排序后的候选列表。我们展示了当能力数量从N=10增加至7,278时,上下文路由的Top-1准确率(Match@1)急剧下降(从0.85降至0.12),而检索-排序方法下降较缓(从0.81降至0.39),因为其重排序器在检索到正确能力后,仍能以70%-87%的概率将其排在首位。在Nova Micro测试中,两者的交叉点约为N=500。我们对比了两个上下文基线方法:Full-Ctx将整个注册表放入提示中,由大语言模型(LLM)选择;Search&Pick则赋予LLM一个搜索工具以缩小候选范围后再选择。在全规模下,该流水线在Match@1指标上领先Search&Pick 6.5个百分点,且成本约为其一半;相比Full-Ctx,成本降低了70倍。我们在智能体、工具和技能注册表中使用固定配置(相同的丰富、检索器和评分器权重)。该流水线已在生产环境中运行,作为大型多智能体平台的默认能力发现层。
cs.CL / 78 / 2608.22704

WnW: Waxing-and-Waning KV Cache for Long-Form Speech LLMs

WnW:用于长格式语音大语言模型的涨落KV缓存
Yao, Yiming, Lyu, Chenyang, Ni, Xuanfan, Wang, Longyue, Luo, Weihua, Yang, Yazheng, Su, Jinsong
Abstract
Long-form audio inputs make the KV cache the dominant memory cost of speech LLMs. Prefill-only KV compression methods permanently discard audio KV positions once evicted, with no pathway to recover them during decoding. We show this is fragile on long-form audio: prefill attention concentrates near the audio start (an attention-sink effect), while decode-time attention distributes broadly, and the two rankings overlap weakly. We propose WnW (Waxing-and-Waning KV cache), which classifies KV-heads into anchor, tidal, and fixed roles via offline calibration. Anchor heads remain on GPU and serve as a decode-time importance observer; tidal heads keep a CPU-resident complement that is recalled chunk-by-chunk based on aggregated anchor-head scores; fixed heads keep only an on-GPU subset, with the rest permanently discarded. On LibriSpeech-Long with two 3B backbones (Voxtral-mini-3b and Qwen2.5-Omni-3B), WnW preserves near-Full-Cache accuracy while keeping only 20% of audio tokens on GPU, where prefill-only baselines fail to terminate. Results generalize across language, task, and domain shifts, and CPU-GPU recall adds little decode-time overhead in our measurements.
Chinese Translation
长格式音频输入使得KV缓存成为语音大语言模型的主要内存开销。仅预填充的KV压缩方法在被驱逐后永久丢弃音频KV位置,解码时无法恢复。我们发现这一方法在长格式音频上表现脆弱:预填充注意力集中在音频开始附近(注意力汇聚效应),而解码时的注意力则广泛分布,两者的排名重叠较弱。我们提出了WnW(涨落KV缓存),通过离线校准将KV头分类为锚定、潮汐和固定角色。锚定头保持在GPU上,作为解码时的重要性观察者;潮汐头保持一个驻留在CPU上的补充,基于聚合的锚定头分数按块召回;固定头仅保留一个在GPU上的子集,其余则永久丢弃。在使用两个3B骨干网络(Voxtral-mini-3b和Qwen2.5-Omni-3B)的LibriSpeech-Long数据集上,WnW在仅保留20%的音频标记在GPU上的情况下,保持了接近满缓存的准确性,而仅预填充的基线方法未能完成。结果在语言、任务和领域转移中具有普遍性,并且CPU-GPU召回在我们的测量中增加的解码时间开销很小。
cs.CL / 79 / 2608.22713

A Source-Grounded Framework for Constructing and Evaluating Progressive Multimodal Diagnostic Dialogues from Clinical Case Reports

基于来源的框架用于构建和评估来自临床案例报告的渐进式多模态诊断对话
Wang, Yufan, Yang, Rui, Liu, Yi, Lin, Yi, Peng, Yifan
Abstract
Clinical diagnosis requires progressive integration of patient history, physical examination, laboratory findings, medical images, and diagnostic-informative tests. However, most multimodal medical benchmarks evaluate fixed inputs or endpoint answers, while fully interactive diagnostic agents conflate evidence selection with evidence interpretation. We present a source-grounded framework to construct progressive multimodal diagnostic dialogues from case reports and an evaluation strategy for assessing MLLMs on final diagnosis, diagnostic reasoning, and image-finding interpretation. Evaluation on 24 internal medicine case reports showed that our framework can accurately convert case reports into reference dialogues, achieving a diagnosis F1 of 0.99 and a reasoning-quality score of 4.79 out of 5. Evaluation on two frontier MLLMs (o4-mini and Claude Haiku 4.5) achieved reasoning-quality scores of 2.75 and 2.50, respectively, with substantially lower diagnosis, reasoning, and image-finding F1 scores. The results demonstrate that fluent responses do not necessarily reflect evidence-grounded clinical reasoning and highlight the utility of the proposed framework for evaluating multimodal diagnostic reasoning.
Chinese Translation
临床诊断需要逐步整合患者病史、体格检查、实验室结果、医学影像和诊断性信息测试。然而,大多数多模态医学基准评估固定输入或最终答案,而完全互动的诊断代理则将证据选择与证据解释混为一谈。我们提出了一种基于来源的框架,用于从案例报告中构建渐进式多模态诊断对话,并提出了一种评估策略,以评估多模态大语言模型(MLLMs)在最终诊断、诊断推理和影像发现解释方面的表现。在对24个内科案例报告的评估中,我们的框架能够准确地将案例报告转换为参考对话,诊断F1得分达到0.99,推理质量得分为4.79(满分5分)。对两个前沿的多模态大语言模型(o4-mini和Claude Haiku 4.5)的评估分别获得了2.75和2.50的推理质量得分,且诊断、推理和影像发现的F1得分显著较低。结果表明,流畅的回答不一定反映基于证据的临床推理,并强调了所提框架在评估多模态诊断推理中的实用性。
cs.CL / 80 / 2608.22745

DiaRelay: Relaying Dialogue Context with a Constant-Size Memory for Emotion Recognition in Conversation

DiaRelay:使用固定大小内存传递对话上下文以进行对话中的情感识别
Zhou, Zihao, Yang, Bin, Qin, Jinghui, Jin, Kebing
Abstract
Emotion Recognition in Conversation (ERC) requires models to identify subtle emotional cues that are often distributed across distant dialogue turns. Existing methods typically incorporate dialogue history through a fixed context window. However, short windows discard potentially useful long-range evidence, while enlarging the window repeatedly re-encodes overlapping utterances, increases computational and memory costs, and may introduce irrelevant context. Moreover, commonly used parameter-efficient adaptation methods, such as LoRA, mainly introduce fixed low-rank transformations in the feature space and do not explicitly maintain a dialogue-level state or condition their transformations on the evolving conversational context. To address these limitations, we propose a lightweight adapter, DiaRelay, to enable LLMs to explicitly maintain a dialogue-level memory for accurate ERC. Based on LoRA, DiaRelay introduces two extra tightly collaborative components, Selective Relay Memory Transition and Dual-axis Relay Memory Read. Selective Relay Memory Transition progressively aggregates useful historical evidence into a bounded relay memory and propagates it across successive utterance predictions. This allows earlier emotional cues to influence later predictions after they leave the local context window, without re-encoding the complete dialogue history or expanding the backbone context length. Dual-axis Relay Memory Read uses the propagated memory to dynamically modulate low-rank feature transformations, enabling context-dependent representation adaptation without test-time gradient updates. Extensive experiments show that DiaRelay can achieve SOTA weighted F1 and accuracy on MELD while obtaining competitive results on IEMOCAP with only an extra 7.1M trainable parameters, indicating the effectiveness and generalizability of our DiaRelay in enhancing LLM-based emotional understanding.
Chinese Translation
对话中的情感识别(ERC)要求模型识别通常分布在遥远对话轮次中的微妙情感线索。现有方法通常通过固定的上下文窗口来整合对话历史。然而,短窗口会丢弃潜在有用的长距离证据,而扩大窗口则会重复对重叠发言进行重新编码,增加计算和内存成本,并可能引入无关上下文。此外,常用的参数高效适应方法,如LoRA,主要在特征空间中引入固定的低秩变换,并未明确维持对话级状态或根据不断演变的对话上下文来调整其变换。为了解决这些局限性,我们提出了一种轻量级适配器DiaRelay,使大型语言模型(LLMs)能够明确维持对话级内存以实现准确的ERC。基于LoRA,DiaRelay引入了两个额外的紧密协作组件:选择性中继记忆转换(Selective Relay Memory Transition)和双轴中继记忆读取(Dual-axis Relay Memory Read)。选择性中继记忆转换逐步将有用的历史证据聚合到一个有限的中继记忆中,并在连续的发言预测中传播。这使得早期的情感线索能够在离开局部上下文窗口后影响后续预测,而无需重新编码完整的对话历史或扩展主干上下文长度。双轴中继记忆读取利用传播的记忆动态调节低秩特征变换,实现上下文依赖的表示适应,而无需在测试时进行梯度更新。大量实验表明,DiaRelay在MELD上能够实现最先进的加权F1和准确率,同时在IEMOCAP上以仅增加7.1M可训练参数的情况下获得竞争性结果,表明我们的DiaRelay在增强基于LLM的情感理解方面的有效性和通用性。
cs.CL / 81 / 2608.22753

Beyond Factual Knowledge: Benchmarking and Learning Step-Level Procedural Rule Reasoning in Large Language Models

超越事实知识:在大型语言模型中基准测试和学习逐步程序规则推理
Yu, Bohan, Cao, Pengfei, Han, Chen, Zhou, Chenxi, Zhang, Zhiheng, Xie, Zhiyang, Teng, Wenhao, Liao, Xiangwen, Zhao, Jun, Liu, Kang
Abstract
Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided procedural rules at scale. To evaluate this capability, we introduce RuleWorld, a large-scale benchmark that reformulates rules as globally reusable abstract units rather than instance-specific facts. In RuleWorld, several scenarios, including single-rule, parallel multi-rule, and multi-hop reasoning, are settled for comprehensive evaluation. We further propose DynaRule, an end-to-end framework that injects the given rules into the KV cache and turns retrieval into an internal, learnable, step-wise process. Specifically, DynaRule employs Stacked Step-Level Attention Training with a special token to enable dynamic rule re-attention and updating during inference. In this way, the model can re-attend to the most relevant rules at each step, dynamically replacing outdated ones to support more stable multi-step reasoning. Experiments on RuleWorld show that existing LLMs face challenges under large rule pools, while DynaRule improves average QA accuracy by up to 19 points and achieves over 85% Recall@1 at 10K rules, outperforming strong baselines by large margins. We make our code and dataset available here: https://github.com/SharkSpicy-NLP/Beyond-Factual-Knowledge.
Chinese Translation
大型语言模型(LLMs)在文本理解和生成方面表现出色,但在大规模可靠理解和应用外部提供的程序规则方面仍然面临挑战。为了评估这一能力,我们引入了RuleWorld,这是一个大规模基准测试,将规则重新表述为全球可重用的抽象单元,而不是特定实例的事实。在RuleWorld中,设定了多个场景,包括单规则、并行多规则和多跳推理,以便进行全面评估。我们进一步提出了DynaRule,这是一种端到端框架,将给定的规则注入KV缓存,并将检索转变为内部可学习的逐步过程。具体而言,DynaRule采用堆叠逐步注意力训练(Stacked Step-Level Attention Training),使用特殊的标记,在推理过程中实现动态规则重新关注和更新。通过这种方式,模型可以在每一步重新关注最相关的规则,动态替换过时的规则,以支持更稳定的多步骤推理。在RuleWorld上的实验表明,现有的LLMs在大规则池下面临挑战,而DynaRule将平均问答准确率提高了多达19个百分点,并在10K规则下实现了超过85%的Recall@1,显著优于强基线。我们在此提供我们的代码和数据集:https://github.com/SharkSpicy-NLP/Beyond-Factual-Knowledge。
cs.CL / 82 / 2608.22758

XTC: Head-Aware Sampling by Excluding Top Choices

XTC:通过排除顶级选择实现头部感知采样
Weidmann, Philipp Emanuel, Roush, Allen, Goldfeder, Judah, Basu, Sanjay, Shwartz-Ziv, Ravid
Abstract
Standard decoding rules for autoregressive language models promote diversity by rescaling the full next-token distribution or truncating its low-probability tail. These strategies overlook a common regime of open-ended generation in which several continuations are plausible but too much probability mass remains concentrated on the most generic choice. We introduce XTC (Exclude Top Choices), a lightweight head-aware decoding operator that targets this regime directly. XTC identifies tokens whose probabilities exceed an absolute plausibility threshold $\tau$: when at least two qualify, it removes the dominant eligible choices with probability $\rho$ and retains only the weakest plausible alternative before renormalization. Across 60 experiments on Gemma 3 27B Q4, Gemma 3 12B Q6, and DeepSeek R1 14B Q6, with scaling validation on Llama 3.3 70B Q4, XTC improves the diversity-repetition Pareto frontier. On creative generation, Distinct-2 increases by 11--15% and repeat trigrams decrease by 27--47% across the four models. Combined with temperature scaling, gains reach 38% in Distinct-2 and 71% in repeat-trigram reduction over baseline. A blinded Amazon Mechanical Turk study with 150 Master raters yields a 62.3% creativity preference for XTC ($p<10^{-4}$) without reduced fluency, while a GPT-4o control judge reproduces the Anthropic-judge direction on every measure. On IFEval with Llama 3.3 70B Q4, XTC preserves prompt-level strict accuracy within 1.7 percentage points of baseline while recovering most of the diversity gain; a temperature setting matched on Distinct-2 reduces IFEval by 8.8 points. The effect is additive with temperature and repetition penalties, robust across quantization levels and model families, and consistent across twelve prompt genres. XTC has been adopted by llama.cpp, ExLlamaV2, and text-generation-webui.
Chinese Translation
自回归语言模型的标准解码规则通过重新缩放完整的下一个标记分布或截断其低概率尾部来促进多样性。这些策略忽视了一种常见的开放式生成模式,在这种模式下,几种延续是合理的,但过多的概率质量仍然集中在最通用的选择上。我们提出了XTC(Exclude Top Choices),一种轻量级的头部感知解码操作符,直接针对这一模式。XTC识别出概率超过绝对合理性阈值$ au$的标记:当至少有两个标记符合条件时,它以概率$ ho$移除主导的合格选择,仅保留最弱的合理替代品,然后进行重新归一化。在对Gemma 3 27B Q4、Gemma 3 12B Q6和DeepSeek R1 14B Q6进行的60次实验中,以及在Llama 3.3 70B Q4上的规模验证中,XTC改善了多样性-重复的帕累托前沿。在创造性生成方面,Distinct-2提高了11%至15%,而重复三元组减少了27%至47%(在四个模型中)。结合温度缩放,XTC在Distinct-2上的增益达到38%,在重复三元组减少方面达到71%。一项针对150名Master评审员的盲测亚马逊机械土耳其研究显示,XTC的创造性偏好为62.3%($p<10^{-4}$),且流畅性未降低,而GPT-4o控制评审在每个指标上重现了Anthropic评审的方向。在使用Llama 3.3 70B Q4的IFEval中,XTC在基线的基础上保持了1.7个百分点的严格准确性,同时恢复了大部分多样性增益;在Distinct-2上匹配的温度设置使IFEval降低了8.8分。该效果与温度和重复惩罚是可加的,在量化级别和模型系列之间具有鲁棒性,并在十二种提示类型中保持一致。XTC已被llama.cpp、ExLlamaV2和text-generation-webui采纳。
cs.CL / 83 / 2608.22761

Don't Repeat Yourself: Stopping Verbatim Loops at Sampling Time

避免重复:在采样时停止逐字循环
Weidmann, Philipp Emanuel, Roush, Allen, Goldfeder, Judah, Basu, Sanjay, Shwartz-Ziv, Ravid
Abstract
Large Language Models generate text autoregressively, but open-ended generation is prone to verbatim looping, in which models repeat spans already present in context. Standard defenses such as repetition, presence, and frequency penalties and n-gram blocking act on token recurrence rather than the sequential structure of a loop, and often suppress looping only at strengths that also degrade formatting or fluency. We propose Don't Repeat Yourself (DRY), a sampling-time logit adjustment that penalizes a candidate token only when generating it would extend the current suffix into an exact continuation of a span seen earlier in the context. Sequence breakers protect chat templates and formatting tokens. Across models from 1.5B to 120B parameters, nine prompt families, and a 600-pair human study, DRY reduces suffix-extension rate by 47% while improving lexical diversity. An intervention-matched placebo produces no comparable reduction, identifying suffix matching as the operative mechanism. On AWQ-quantized 70B and 120B models, DRY reduces loop rate by roughly half while preserving MT-Bench, MMLU, and GSM8K performance, whereas standard alternatives lose measurable ground. DRY has been adopted by popular open-source LLM inference frameworks including llama.cpp, ExLlamaV2, and text-generation-webui, highlighting its practical impact on text generation.
Chinese Translation
大型语言模型以自回归方式生成文本,但开放式生成容易出现逐字循环,即模型重复上下文中已存在的文本片段。标准防御措施如重复、存在和频率惩罚以及n-gram阻塞作用于标记的重复,而不是循环的顺序结构,往往只能在降低格式或流畅性的强度下抑制循环。我们提出了“避免重复”(Don't Repeat Yourself, DRY),这是一种在采样时进行的logit调整,仅在生成候选标记会将当前后缀延伸为上下文中早先出现的文本片段的精确延续时对其进行惩罚。序列破坏者保护聊天模板和格式标记。在从15亿到1200亿参数的模型、九个提示家族和600对人类研究中,DRY将后缀延伸率降低了47%,同时提高了词汇多样性。与干预匹配的安慰剂没有产生可比的降低,确认后缀匹配是有效机制。在AWQ量化的70B和120B模型上,DRY将循环率减少了大约一半,同时保持了MT-Bench、MMLU和GSM8K的性能,而标准替代方案则失去了可测量的优势。DRY已被包括llama.cpp、ExLlamaV2和text-generation-webui在内的流行开源LLM推理框架采纳,突显了其对文本生成的实际影响。
cs.CL / 84 / 2608.22770

DelistBench: Evaluating Search-Enabled LLMs for Auditable Corporate-Event Database Completion

DelistBench:评估可审计企业事件数据库补全的搜索启用大语言模型
Yao, Xuan, Shuping, Li, Yang, Dai, Yi, Zhou, Huang, Ke-Wei
Abstract
Financial institutions need an independent way to detect missing, stale, and misclassified corporate-event records in vendor databases. We introduce Search-to-Record, a database-assurance task in which search-enabled large language models reconstruct institution-defined event records from public sources for a known security universe and historical cutoff, and DelistBench, a 1,200-record benchmark for security-level delisting announcements. We evaluate five models in paired closed-book and web-enabled conditions. Web access raises announcement-date accuracy within seven days by 34.0 to 48.0 percentage points and event-status accuracy by approximately 2.8 to 21.7 points; the best system achieves 81.5% overall joint accuracy within seven days. Economy web systems achieve 75.9-78.3% overall joint accuracy within seven days at 4.5-6.6% of the API cost of the most expensive web system. Risk-based triage identifies low-error subsets, although the highest-coverage operating point still sends 27.3% of the balanced test set to review. The evaluation identifies web retrieval as the main source of timing gains and shows that low-cost systems can approach the best system's accuracy. Together, Search-to-Record, DelistBench, and the evaluation provide concrete deployment guidance: calibrate triage to local event prevalence and market mix, preserve positive-event recall, and route positive and ambiguous cases to targeted review.
Chinese Translation
金融机构需要一种独立的方法来检测供应商数据库中缺失、过时和错误分类的企业事件记录。我们提出了Search-to-Record,这是一项数据库保障任务,搜索启用的大语言模型从公共来源重建机构定义的事件记录,针对已知的安全范围和历史截止日期,并推出DelistBench,这是一个包含1,200条记录的安全级别退市公告基准。我们在配对的闭卷和网络启用条件下评估了五个模型。网络访问将公告日期的准确性在七天内提高了34.0到48.0个百分点,事件状态的准确性提高了大约2.8到21.7个百分点;最佳系统在七天内实现了81.5%的整体联合准确性。经济型网络系统在七天内实现了75.9-78.3%的整体联合准确性,成本仅为最昂贵网络系统的4.5-6.6%。基于风险的分流识别出低错误子集,尽管覆盖率最高的操作点仍然将27.3%的平衡测试集送审。评估结果表明,网络检索是时间增益的主要来源,并显示低成本系统可以接近最佳系统的准确性。结合Search-to-Record、DelistBench及评估结果,提供了具体的部署指导:根据当地事件的流行程度和市场组合校准分流,保留正事件的召回,并将正面和模糊案例路由至针对性审查。
cs.CL / 85 / 2608.22772

SPOC-SQL: Stage-wise Preference Optimization for Controllable Text-to-SQL

SPOC-SQL:可控文本到SQL的阶段性偏好优化
Chen, Yingnan, Ding, Chun, Xu, Tianshi, Yang, Xu, Wu, Si
Abstract
Text-to-SQL aims to translate natural language questions into executable SQL queries over relational databases, requiring multi-stage structured reasoning over database schemas and query constraints. However, existing methods treat this task as single-step generation, where models optimize entire SQL sequences without targeted feedback at key decision points and lack support for interacting with and controlling the intermediate generation process. To address this issue, we propose SPOC-SQL, which decomposes Text-to-SQL into four sequential subtasks following standard SQL execution logic and designs stage-specific optimization strategies for the model to learn key decisions. Specifically, we propose the implementation of fine-grained preference optimisation at key decision points across SQL stages, with the objective of enhancing structured decision-making during query construction. Furthermore, a structured decomposition strategy is designed, facilitating stage-wise intervention and correction through explicit intermediate representations. This results in more controllable and reliable SQL generation. Experiments demonstrate that incorporating stage-wise human knowledge consistently improves performance, validating the effectiveness of stage perception controllable generation.
Chinese Translation
文本到SQL旨在将自然语言问题翻译为可在关系数据库上执行的SQL查询,这需要对数据库模式和查询约束进行多阶段的结构化推理。然而,现有方法将此任务视为单步生成,模型在关键决策点没有针对性的反馈,且缺乏与中间生成过程交互和控制的支持。为了解决这一问题,我们提出了SPOC-SQL,它将文本到SQL分解为四个遵循标准SQL执行逻辑的顺序子任务,并为模型设计了阶段特定的优化策略,以学习关键决策。具体而言,我们提出在SQL阶段的关键决策点实施细粒度的偏好优化,旨在增强查询构建过程中的结构化决策能力。此外,设计了一种结构化分解策略,通过明确的中间表示促进阶段性干预和修正。这使得SQL生成更加可控和可靠。实验表明,结合阶段性人类知识可以持续提高性能,验证了阶段感知可控生成的有效性。
cs.CL / 86 / 2608.22793

TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents

TRACE:一个自我演化的技能库,用于一致性和限知意识的LLM代理
Wu, Wenhao, Zhang, Menghao, Wang, Xin, Wang, Zhi, Shao, Kun, Luan, Jian
Abstract
Reliable deployment of LLM agents in user-facing products depends not on raw task-solving ability but on consistency and limit-awareness: behaving the same way across repeated trials, and recognizing when a request cannot, or cannot yet, be safely fulfilled. CAR-bench exposes this reliability gap in the domain of in-car assistants: an LLM-simulated user issues incomplete or ambiguous requests, requiring the agent to resolve uncertainty through multi-turn dialogue and tool use while strictly adhering to domain policies. Even frontier models show a substantial gap between what they can solve at least once (Pass@3) and what they solve consistently across trials (Pass^k). We bridge this gap with TRACE (TRAjectory-Contrastive Evolution), which iteratively improves a skill-based agent's behavioral knowledge without modifying model weights. This knowledge is organized as a Skill Bank of modular, retrievable skills, each encoding a self-contained set of tool-use rules and behavioral guidelines. TRACE evolves this bank through an agentic self-evolution loop: after each evaluation round, it groups trajectories by the skills invoked and refines each skill by contrasting successful and failed behaviors. The updated bank then guides subsequent rounds, while during deployment the Actor performs state-conditioned skill orchestration at every turn. On GPT-5.5, TRACE improves consistency (Pass^3) by 34.6 points, from 59.9% to 94.5%, while shrinking the gap between potential and reliable performance to just 4.0 points. On the official hidden set, TRACE achieved first place using GPT-5.6-Sol, attaining a Pass^3 score of 70%-a 40% relative improvement over the baseline. These results show that TRACE converts high model potential into stable, consistent performance gain. Project homepage: https://darwin-agent.github.io/Car-bench-TRACE.
Chinese Translation
在用户面向产品中可靠地部署LLM代理不仅依赖于原始的任务解决能力,还依赖于一致性和限知意识:在重复试验中表现一致,并识别何时请求无法或尚不能安全满足。CAR-bench揭示了这一可靠性差距,特别是在车载助手领域:一个LLM模拟的用户发出不完整或模糊的请求,要求代理通过多轮对话和工具使用来解决不确定性,同时严格遵循领域政策。即使是前沿模型,在至少解决一次的能力(Pass@3)与在多次试验中一致解决的能力(Pass^k)之间也存在显著差距。我们通过TRACE(轨迹对比演化)来弥合这一差距,该方法迭代地改善基于技能的代理的行为知识,而不修改模型权重。这种知识被组织为一个技能库,其中包含模块化、可检索的技能,每个技能编码了一组自包含的工具使用规则和行为指南。TRACE通过代理自我演化循环来发展这个库:在每次评估轮后,它根据调用的技能对轨迹进行分组,并通过对比成功和失败的行为来细化每个技能。更新后的技能库随后指导后续轮次,而在部署期间,Actor在每次回合中执行状态条件的技能编排。在GPT-5.5上,TRACE将一致性(Pass^3)提高了34.6个百分点,从59.9%提升至94.5%,同时将潜在性能与可靠性能之间的差距缩小至仅4.0个百分点。在官方隐藏集上,TRACE使用GPT-5.6-Sol获得第一名,达到了70%的Pass^3分数,相较于基线提高了40%。这些结果表明,TRACE将高模型潜力转化为稳定、一致的性能提升。项目主页:https://darwin-agent.github.io/Car-bench-TRACE。
cs.CL / 87 / 2608.22802

SDoH-Aware Narrative Anchoring Bias in Medical LLMs for Trustworthy Clinical Decision Support

关注社会决定因素的医疗大型语言模型中的叙事锚定偏差及其在可信临床决策支持中的应用
Choudhury, Ahnaf Atef, Saha, Ramkrishna
Abstract
Medical large language models are often judged by how many clinical questions they answer correctly. That view is useful, but it misses a practical risk. A model may know the right answer and still change its response when the same case is written in a different patient voice. This paper evaluates that risk as SDoH aware narrative anchoring bias. We use NarrativeShield SDoH MedQA, a counterfactual medical question answering dataset in which each case appears in persona based narratives while the answer key remains fixed. The dataset is reshaped from wide format into case grouped persona rows. We evaluate three open source instruction tuned LLMs from the Qwen2.5 family: 1.5B, 3B, and 7B. The final experiment uses 300 clinical cases and produces 8,100 model responses across three prompting conditions. We report persona level accuracy, counterfactual consistency, correct consistency, and narrative sensitivity error. Qwen2.5 7B achieves the best accuracy at 56.33 percent and the best correct consistency at 40.33 percent. Paired McNemar exact tests show significant accuracy gains for 7B over 3B in all prompt settings. Even so, narrative sensitivity remains, with the lowest error still at 31.67 percent. These results suggest that trustworthy clinical decision support should be evaluated by both average correctness and stability across medically equivalent patient narratives.
Chinese Translation
医疗大型语言模型通常通过其正确回答的临床问题数量来评估。这种观点虽然有用,但忽略了一个实际风险:模型可能知道正确答案,但在同一案例以不同患者声音表述时仍会改变其回应。本文评估了这一风险,称之为关注社会决定因素的叙事锚定偏差。我们使用了 NarrativeShield SDoH MedQA,这是一个反事实医疗问答数据集,其中每个案例以角色叙述的形式呈现,而答案键保持不变。该数据集从宽格式重塑为按案例分组的角色行。我们评估了来自 Qwen2.5 系列的三种开源指令调优大型语言模型:1.5B、3B 和 7B。最终实验使用了 300 个临床案例,并在三种提示条件下生成了 8,100 个模型响应。我们报告了角色级别的准确性、反事实一致性、正确一致性和叙事敏感性错误。Qwen2.5 7B 在准确性方面表现最佳,达到了 56.33%,在正确一致性方面也表现最佳,达到了 40.33%。配对 McNemar 精确检验显示,7B 在所有提示设置中相较于 3B 有显著的准确性提升。尽管如此,叙事敏感性仍然存在,最低错误率仍为 31.67%。这些结果表明,可信的临床决策支持应通过平均正确性和在医学上等效的患者叙述中的稳定性进行评估。
cs.CL / 88 / 2608.22806

DIAG: Diagnostic Iterative Alignment and Generation for Data-Efficient Mathematical Preference Distillation

DIAG:用于数据高效数学偏好蒸馏的诊断性迭代对齐与生成
Chen, Guhan, Tian, Songtao, Li, Bohan, Wang, Hejin, Xie, YeXin, Yu, Zixiong
Abstract
Iterative preference optimization is essential for aligning Large Language Models on mathematical reasoning tasks, yet its efficiency is often throttled by signal scarcity: as the model improves, static problem sets become increasingly mismatched to the model's evolving competence, producing rollouts that are either too easy or too hard and therefore non-informative, which leads to a scarcity of valid preference pairs. We propose DIAG, a Diagnostic Iterative Alignment and Generation framework that adaptively reshapes the practice distribution to increase informative supervision and focus training near the student's current competence boundary. DIAG consists of two phases: (1) diagnosing valid preference-pair yield to calibrate the exploration-exploitation trade-off and allocate topic quotas via an Empirical Bayes shrinkage estimator, thereby prioritizing high-yield concepts; and (2) generating targeted practice, where a teacher synthesizes variants from the student's failure traces. We further provide a theoretical view interpreting DIAG as a teacher-mediated approximation to KL-regularized reweighting of the practice distribution toward the student's competence boundary, where valid preference-pair yield is maximized. Experiments show that DIAG boosts yield across iterations and delivers stronger reasoning performance under an iso-effective training budget, demonstrating that it can distill more informative preference supervision for mathematical reasoning.
Chinese Translation
迭代偏好优化对于在数学推理任务中对齐大型语言模型至关重要,但其效率常常受到信号稀缺的限制:随着模型的改进,静态问题集与模型不断发展的能力之间的匹配程度逐渐降低,导致生成的结果要么过于简单,要么过于困难,因此缺乏信息量,从而导致有效偏好对的稀缺。我们提出了DIAG,一个诊断性迭代对齐与生成框架,能够自适应地重塑实践分布,以增加信息性监督并集中训练在学生当前能力边界附近。DIAG由两个阶段组成:(1)诊断有效偏好对的产出,以校准探索与利用的权衡,并通过经验贝叶斯收缩估计器分配主题配额,从而优先考虑高产出的概念;(2)生成针对性的练习,其中教师从学生的失败轨迹中合成变体。我们进一步提供了一个理论视角,将DIAG解释为教师介导的近似,旨在对实践分布进行KL正则化重加权,以朝向学生的能力边界最大化有效偏好对的产出。实验表明,DIAG在迭代过程中提高了产出,并在等效训练预算下提供了更强的推理性能,证明其能够为数学推理蒸馏出更具信息性的偏好监督。
cs.CL / 89 / 2608.22817

Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical Reports

工业指令:从工业技术报告构建指令调优和基准数据集的端到端框架
Bakhtiari, Parsa, Bashiri, Hassan, Khalilipour, Alireza, Nasiripour, Masoud, Challenger, Moharram
Abstract
Industrial technical reports contain high-value knowledge for maintenance, troubleshooting, and product engineering, but their heterogeneous structure (dense prose, specifications, tables) makes them difficult to index and reason over with standard retrieval and QA pipelines, and no public instruction-tuning or benchmark datasets are built from such documents. We address this gap with Industrial-Instruction, contributing (i) two open QA datasets built from real industrial technical reports and (ii) the end-to-end pipeline that produces them. Using 906 public Panasonic documents (7,525 pages), we apply layout-aware extraction, build a semantic retrieval index, and synthesize multiple-choice QA grounded in retrieved evidence under five query-document relationships (irrelevant retrieval, single-/multi-document support, single-/multi-document answer). After filtering an initial 23.9k generated samples, each dataset provides approximately 13.6k QA pairs with source documents and a held-out benchmark split. Fine-tuning small open LLMs (under 10B parameters) improves Set-Match Accuracy from 28.5% to 42.0% and F1 from 46.6% to 63.5% on the Panasonic benchmark. We release two parallel versions built by the same pipeline: one generated with the open-weight Qwen3-30B-A3B-Instruct model and one with the closed, API-based Claude-Opus-4.6 model, enabling a direct comparison of open- versus frontier-model data generation. The Claude-Opus-4.6 dataset yields a cleaner raw corpus and larger fine-tuning gains, at roughly two orders of magnitude higher cost. MMLU evaluation shows models trained on the Claude-Opus-4.6 data retain essentially all general knowledge, versus a small but measurable forgetting effect for the Qwen-generated data. Together, these datasets and pipeline offer a practical, reproducible path toward scalable industrial benchmarks and training data from real-world documentation.
Chinese Translation
工业技术报告包含维护、故障排除和产品工程的高价值知识,但其异构结构(密集的散文、规格、表格)使得使用标准检索和问答管道进行索引和推理变得困难,并且尚未有公共的指令调优或基准数据集是基于此类文档构建的。我们通过工业指令(Industrial-Instruction)来填补这一空白,贡献了(i)两个基于真实工业技术报告构建的开放问答数据集,以及(ii)生成这些数据集的端到端管道。我们使用906份公开的松下(Panasonic)文档(7,525页),应用布局感知提取,构建语义检索索引,并在五种查询-文档关系下合成基于检索证据的多项选择问答(无关检索、单文档/多文档支持、单文档/多文档答案)。在过滤初始生成的23.9k样本后,每个数据集提供大约13.6k个问答对及其源文档,并有一个保留的基准拆分。对小型开放大语言模型(LLMs,参数少于10B)进行微调,使得在松下基准上的集合匹配准确率从28.5%提高到42.0%,F1值从46.6%提高到63.5%。我们发布了由同一管道构建的两个平行版本:一个是使用开放权重的Qwen3-30B-A3B-Instruct模型生成的,另一个是使用封闭的基于API的Claude-Opus-4.6模型生成的,从而实现开放模型与前沿模型数据生成的直接比较。Claude-Opus-4.6数据集提供了更干净的原始语料库和更大的微调收益,但成本大约高出两个数量级。MMLU评估显示,基于Claude-Opus-4.6数据训练的模型几乎保留了所有通用知识,而Qwen生成的数据则表现出小但可测量的遗忘效应。总之,这些数据集和管道为从真实世界文档中获得可扩展的工业基准和训练数据提供了一条实用且可重复的路径。
cs.CL / 90 / 2608.22857

SAVER: Selective Auditing of Verbal Evidence for Error Recovery in VLM Change Reasoning

SAVER:针对视觉语言模型变更推理中的错误恢复的选择性审计口头证据
Li, Youdi
Abstract
Vision-language models (VLMs) frequently fail at visual change reasoning, even when their vision encoders contain sufficient information. We observe that correct VLM outputs tend to contain explicit verbal evidence (object names, colors, spatial locations) that supports the claimed change, while incorrect outputs often lack such evidence. We propose SAVER (Selective Auditing of Verbal Evidence for Error Recovery), a lightweight, rule-based method that parses VLM responses for this evidence and triggers structured reprompting only when evidence is missing or inconsistent. Across three change detection benchmarks and four VLMs, SAVER significantly improves accuracy on tasks where errors stem from the model failing to articulate what it saw (expression failures), with gains up to +25.8% on CLEVR-Change. The evidence patterns can also be generated by an LLM in a single call, matching the hand-tuned gate on CLEVR-Change. Ablation experiments confirm that the evidence gate, not reprompting alone, drives the improvement.
Chinese Translation
视觉语言模型(VLMs)在视觉变更推理中经常失败,即使它们的视觉编码器包含足够的信息。我们观察到,正确的 VLM 输出往往包含明确的口头证据(物体名称、颜色、空间位置),以支持所声称的变更,而错误的输出通常缺乏这种证据。我们提出了 SAVER(选择性审计口头证据以进行错误恢复),这是一种轻量级的基于规则的方法,它解析 VLM 响应中的证据,并仅在缺少或不一致的证据时触发结构化的重新提示。在三个变更检测基准和四个 VLM 上,SAVER 显著提高了任务的准确性,这些任务的错误源于模型未能表达其所见(表达失败),在 CLEVR-Change 上的增益高达 +25.8%。这些证据模式也可以通过 LLM 在一次调用中生成,匹配 CLEVR-Change 上的手动调优门控。消融实验确认,证据门控而非单纯的重新提示推动了性能的提升。
cs.CL / 91 / 2608.22872

Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors

更好的检索,较差的鲁棒性:多跳 RAG 如何放大上游 ASR 错误
Bao, Zhenghua
Abstract
Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module, so ASR errors enter the pipeline as a fixed upstream constraint. We empirically test whether two extensions to standard retrieval-augmented generation (RAG), entity-graph linking and iterative reformulation, absorb or amplify these errors. Using four English accents synthesized through neural TTS, we evaluate four RAG configurations on three multi-hop QA benchmarks (HotpotQA, 2WikiMultiHopQA and MuSiQue) against a clean-text oracle. Although the structurally richer configurations generally retain higher absolute F1 under ASR input, both extensions amplify the error: the F1 gap from clean text to the highest-WER accent is 36-67% larger under their combination than under naive dense retrieval, on all three benchmarks. The dominant failure mode is corruption of one or more query entities, accounting for 87-96% of degradation cases on 2WikiMultiHopQA across all four methods. Two lightweight surface-form mitigations leave most of the gap intact, indicating that downstream retrieval structure amplifies remaining entity errors. We release code and data at https://github.com/ZhenghuaBao/spoken-multihop-rag .
Chinese Translation
基于语音的应用程序在任何检索模块之前通过自动语音识别(ASR)处理口语查询,因此 ASR 错误作为固定的上游约束进入管道。我们实证测试了对标准检索增强生成(RAG)的两种扩展,实体图链接和迭代重构,是否吸收或放大这些错误。使用通过神经文本到语音(TTS)合成的四种英语口音,我们在三个多跳问答基准(HotpotQA、2WikiMultiHopQA 和 MuSiQue)上评估了四种 RAG 配置,与干净文本的 oracle 进行对比。尽管结构上更丰富的配置在 ASR 输入下通常保持更高的绝对 F1 值,但这两种扩展都放大了错误:在所有三个基准上,从干净文本到最高 WER 口音的 F1 差距在它们的组合下比在简单密集检索下大 36-67%。主导的失败模式是一个或多个查询实体的损坏,占 2WikiMultiHopQA 所有四种方法中降级案例的 87-96%。两种轻量级表面形式的缓解措施保持了大部分差距,表明下游检索结构放大了剩余的实体错误。我们在 https://github.com/ZhenghuaBao/spoken-multihop-rag 发布代码和数据。
cs.CL / 92 / 2608.22894

AraDetox: A Multi-Dialect Arabic Detoxification Dataset

AraDetox:一个多方言阿拉伯语去毒化数据集
El-Haj, Mo
Abstract
Arabic harmful-language detection has received considerable attention, yet Arabic text detoxification remains underexplored. We introduce AraDetox, a multi-dialect Arabic detoxification dataset comprising 10,500 harmful social-media posts and 84,000 detoxified rewrites generated using GPT-5 and Gemini 2.5 Flash across Modern Standard Arabic, Gulf, Levantine, and Egyptian Arabic. The generated outputs were assessed through human evaluation and automatic analyses of lexical change, semantic preservation, sentiment, and dialectal style. Results show that detoxification is primarily a meaning-preserving rewriting task: substantial lexical and structural reformulation is accompanied by consistently high semantic similarity. Human evaluation confirms successful harmful-language removal while largely preserving the original meaning. Dialectal analyses further indicate that the generated variants exhibit measurable stylistic alignment with reference Arabic dialect corpora. Comparison with existing resources highlights two complementary approaches to detoxification: minimal-edit lexical substitution and meaning-preserving reformulation. Our findings demonstrate that large-scale Arabic detoxification resources can be constructed through LLM-assisted generation and human verification. The dataset is publicly available at https://github.com/ArabicNLP-UK/AraDetox to support future research on Arabic detoxification, safe text generation, and multi-dialect Arabic NLP.
Chinese Translation
阿拉伯语有害语言检测受到了广泛关注,但阿拉伯语文本去毒化仍然未得到充分探索。我们介绍了AraDetox,这是一个多方言阿拉伯语去毒化数据集,包含10,500条有害社交媒体帖子和84,000条使用GPT-5和Gemini 2.5 Flash生成的去毒化重写文本,涵盖现代标准阿拉伯语、海湾阿拉伯语、黎凡特阿拉伯语和埃及阿拉伯语。生成的输出通过人工评估和对词汇变化、语义保留、情感和方言风格的自动分析进行了评估。结果表明,去毒化主要是一项保留意义的重写任务:显著的词汇和结构重组伴随着持续较高的语义相似性。人工评估确认成功去除了有害语言,同时在很大程度上保留了原始意义。方言分析进一步表明,生成的变体在风格上与参考阿拉伯方言语料库表现出可测量的一致性。与现有资源的比较突显了两种互补的去毒化方法:最小编辑的词汇替换和保留意义的重组。我们的研究结果表明,可以通过大型语言模型(LLM)辅助生成和人工验证构建大规模阿拉伯语去毒化资源。该数据集已在https://github.com/ArabicNLP-UK/AraDetox公开发布,以支持未来关于阿拉伯语去毒化、安全文本生成和多方言阿拉伯语自然语言处理的研究。
cs.CL / 93 / 2608.22898

SelFusion: Self-distillation for Diffusion Language Models

SelFusion:扩散语言模型的自蒸馏
Lim, Hyeongsoo, Kim, Jinyoung, Seo, Eunseo, Jang, Minho, Yoon, Jiwon
Abstract
Diffusion language models (DLMs) alleviate the inherent latency bottleneck of autoregressive (AR) large language models (LLMs), but their degraded generation quality limits practical applicability. Although knowledge distillation (KD) can be a promising direction for improving performance, we empirically find that naively applying conventional KD yields only marginal gains, or even degrades generation quality. Based on these observations, we propose a novel self-distillation framework for DLMs, namely SelFusion. To enable effective KD without an external teacher model, SelFusion performs two forward passes with different masking levels, defining the hard mode with a larger masking probability and the easy mode with a smaller masking probability. However, the easy mode is not always more accurate than the hard mode and can be overconfident on incorrect tokens. Thus, we introduce bidirectional KD between the two modes, which can dynamically determine the distillation direction based on token-level correctness. Experimental results on instruction-following tasks show that the proposed self-distillation substantially outperforms other KD methods with external LLM and DLM teachers. In many configurations, the student trained with SelFusion even surpasses the performance of the LLM teacher, providing a practical path toward improving DLM generation quality. Source code can be found at https://github.com/scai-research/SelFusion_official
Chinese Translation
扩散语言模型(DLMs)缓解了自回归(AR)大型语言模型(LLMs)固有的延迟瓶颈,但其生成质量的下降限制了实际应用。尽管知识蒸馏(KD)可能是提高性能的一个有前景的方向,但我们实证发现,简单地应用传统的KD仅能带来微小的提升,甚至可能降低生成质量。基于这些观察,我们提出了一种新颖的DLM自蒸馏框架,称为SelFusion。为了在没有外部教师模型的情况下实现有效的KD,SelFusion进行两次前向传播,采用不同的掩码级别,定义了掩码概率较大的困难模式和掩码概率较小的简单模式。然而,简单模式并不总是比困难模式更准确,并且可能对错误的标记过于自信。因此,我们在两种模式之间引入了双向KD,可以根据标记级别的正确性动态确定蒸馏方向。在指令跟随任务上的实验结果表明,所提出的自蒸馏方法显著优于其他使用外部LLM和DLM教师的KD方法。在许多配置中,使用SelFusion训练的学生模型甚至超过了LLM教师的性能,为提高DLM生成质量提供了一条切实可行的路径。源代码可在 https://github.com/scai-research/SelFusion_official 找到。
cs.CL / 94 / 2608.22908

Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text

口语语言模型在阅读文本时能听到语音吗?弥合语音与文本之间的结构差距
Kim, Hyeonyu, Kim, Hwayeon, Choi, Youngwon, Cho, Myeongkyun, Nguyen, Huu-Kim
Abstract
Spoken Language Models (SLMs) generate textual responses directly from speech, offering an alternative to cascaded systems. Despite recent advances, existing SLMs still exhibit weaker instruction-following behavior and limited generalization across diverse tasks compared to text-based language models. Our analysis shows that speech and text representations in current SLMs remain weakly aligned despite strong downstream performance, indicating that structural differences between continuous, temporally varying speech and discrete text remain insufficiently addressed. To address this, we propose a simple framework that decouples length mismatch from semantic alignment and encourages closer correspondence between speech and text representations. Experiments across multiple benchmarks demonstrate competitive performance against strong baselines, underscoring the importance of explicitly addressing structural differences between speech and text in SLM training. Our code is publicly available at https://github.com/jaykim9870/Do_SLMs_Hear_Speech_as_They_Read_Text.
Chinese Translation
口语语言模型(Spoken Language Models, SLMs)直接从语音生成文本响应,为级联系统提供了一种替代方案。尽管最近取得了一些进展,现有的SLMs在遵循指令的行为和在不同任务上的泛化能力方面仍然不如基于文本的语言模型。我们的分析表明,当前SLMs中语音和文本的表示虽然在下游任务中表现良好,但仍然存在弱对齐的问题,这表明连续、时间变化的语音与离散文本之间的结构差异尚未得到充分解决。为了解决这个问题,我们提出了一个简单的框架,该框架将长度不匹配与语义对齐解耦,并鼓励语音与文本表示之间更紧密的对应关系。在多个基准测试中的实验表明,该方法在强基线下表现出竞争力,强调了在SLM训练中明确解决语音与文本之间的结构差异的重要性。我们的代码已公开发布在 https://github.com/jaykim9870/Do_SLMs_Hear_Speech_as_They_Read_Text。
cs.CL / 95 / 2608.22909

Exploring Dowker Homology for Sentence Similarity

探索道克同调在句子相似性中的应用
Huber, Marius, Opitz, Juri
Abstract
Dowker homology is a topological tool that may be used to analyze the relative position of two point clouds living in a common space. We investigate whether Dowker homology captures sentence similarity information by treating the embeddings of the tokens that constitute a sentence pair as a pair of point clouds in the latent space of a transformer model, using both models that have and have not been fine-tuned for sentence similarity. We find that Dowker homology captures sentence similarity information, as measured by regressing Dowker homology features onto ground-truth similarity scores, and that it can be used for visual inspection of similarity data and models. In an attempt to make Dowker homology readily applicable, we derive from it single-number summaries that we expect to capture sentence similarity directly. These turn out to work reasonably well, but without outperforming standard sentence similarity measures based on established pooling methods.
Chinese Translation
道克同调是一种拓扑工具,可用于分析位于共同空间中的两个点云的相对位置。我们研究道克同调是否能够捕捉句子相似性信息,通过将构成句子对的标记的嵌入视为潜在空间中的一对点云,使用已经针对句子相似性进行微调的模型和未进行微调的模型。我们发现,道克同调能够捕捉句子相似性信息,这通过将道克同调特征回归到真实相似性评分上得以验证,并且它可以用于相似性数据和模型的可视化检查。为了使道克同调更易于应用,我们从中推导出单一数值摘要,期望能够直接捕捉句子相似性。这些摘要的效果相对良好,但并未超越基于既定池化方法的标准句子相似性测量。
cs.CL / 96 / 2608.22916

Knowing Isn't Always Saying: When Do Spatial Encodings Reach Answers in Vision-Language Models?

知识并不总是表达:空间编码在视觉-语言模型中何时达到答案?
Wang, Zeyu, Xu, Xinming
Abstract
Vision-language models are known to encode spatial information in their hidden states, yet often fail to use it when answering. However, it remains unclear when and where this encoded information reaches the answer. We address this with direction patching, a class-conditioned causal intervention applied across layers, token positions, and prompt formats. Using spatial-ID directions constructed following prior encoding evidence, we find that causal influence on answer logits emerges only at mid-to-deep depths. Text chain-of-thought suppresses immediate object-word argmax-level transport in most models, while visually grounded prompts keep it open. Positive target-logit gain can remain below the argmax threshold, and transport can re-emerge at the final prefix token or at the answer step in deeper layers. Across the ten VLMs we study, these local effects form descriptive transport patterns. Complementary experiments characterize how these patterns shift across datasets, attributes, and encoding amplitudes. Together, these results reframe the encoding-grounding gap as a problem of conditional transport in VLMs.
Chinese Translation
视觉-语言模型已知在其隐藏状态中编码空间信息,但在回答时常常未能有效利用这些信息。然而,何时以及何处这些编码的信息达到答案仍不清楚。我们通过方向补丁(direction patching)来解决这一问题,这是一种在层、标记位置和提示格式之间应用的类条件因果干预。使用根据先前编码证据构建的空间-ID方向,我们发现对答案对数的因果影响仅在中到深层次中出现。文本思维链在大多数模型中抑制了即时对象-词的最大值传输,而视觉基础的提示则保持其开放。正的目标对数增益可能仍低于最大值阈值,并且在更深层次的最终前缀标记或答案步骤中,传输可能重新出现。在我们研究的十个视觉-语言模型中,这些局部效应形成了描述性传输模式。补充实验表征了这些模式如何在数据集、属性和编码幅度之间变化。总的来说,这些结果将编码-基础差距重新框架为视觉-语言模型中的条件传输问题。
cs.CL / 97 / 2608.22917

TSWAP: A Multilingual Retrieval-Augmented Thai Wellness Advisor

TSWAP:一种多语言检索增强的泰国健康顾问
Ukosaramig, Pornthep, Viriyayudhakorn, Kobkrit
Abstract
We present TSWAP, a deployed eight-language conversational wellness advisor grounded, via retrieval-augmented generation, in a verified knowledge base of Thai traditional medicine and certified wellness providers. An unmodified open-weight LLM (Qwen3.6-35B-A3B on vLLM) is grounded on a ~30.6K-chunk Thai index by a hybrid dense-sparse retriever with cross-encoder reranking; a first-turn query classifier forces tool-based retrieval for entity lookups; a rule-based safety layer enforces medical scope and Thai emergency routing; and all eight languages are served zero-shot with translate-then-retrieve. We release the first Thai traditional-medicine/wellness retrieval benchmark (50 questions with gold document IDs; Recall@5 = 0.88), production QA logs (91.1% test-retest pass over 259 cases), and a 71-question frontier no-retrieval probe showing what each grounding pillar contributes: without the safety prompt the backend model family produced a full drug-dosing schedule and complied with out-of-scope requests, and without the knowledge base it produced zero verifiable provider recommendations. We further report two transferable deployment findings: English-calibrated 4-bit AWQ quantization corrupts Thai tone marks, and forced-retrieval routing is necessary for reliable grounding.
Chinese Translation
我们提出了TSWAP,这是一种部署的八种语言对话式健康顾问,通过检索增强生成技术,基于经过验证的泰国传统医学知识库和认证的健康服务提供者。未修改的开放权重大型语言模型(Qwen3.6-35B-A3B在vLLM上)通过混合稠密-稀疏检索器与交叉编码器重排序,基于约30.6K块的泰国索引进行基础构建;首轮查询分类器强制使用工具基础的检索进行实体查找;基于规则的安全层强制执行医学范围和泰国紧急路由;所有八种语言均以零样本方式通过翻译后检索提供服务。我们发布了首个泰国传统医学/健康检索基准(50个问题及其金标准文档ID;Recall@5 = 0.88)、生产问答日志(在259个案例中91.1%的测试重测通过率),以及一个71个问题的前沿无检索探测,展示了每个基础支柱的贡献:在没有安全提示的情况下,后端模型家族生成了完整的药物剂量计划并满足了超出范围的请求,而没有知识库则生成了零个可验证的提供者推荐。我们进一步报告了两个可转移的部署发现:经过英语校准的4位AWQ量化会损坏泰语声调符号,强制检索路由对于可靠的基础构建是必要的。
cs.CL / 98 / 2608.22922

HelaBERT: Enhancing Sinhala Language Understanding with Dual Pooling Classification Head

HelaBERT:通过双池化分类头增强僧伽罗语理解
Ekanayake, Thisen, de Silva, Nisansa
Abstract
We present HelaBERT, a family of two BERT-based masked language models pre-trained from scratch on approximately 1 billion tokens of Sinhala text sourced from MADLAD-400, CulturaX, and a custom corpus comprising news articles, Sinhala Wikipedia, and web crawl data. HelaBERT-Small (~23.3M parameters, 6 layers) and HelaBERT-Large (~110M parameters, 12 layers) both use a SentencePiece Unigram tokenizer (vocabulary size 32,000) tailored to Sinhala's agglutinative morphology and complex script. We evaluate both models on four downstream Sinhala text classification tasks: news category classification, news source classification, sentiment analysis, and writing style classification, using 5 independent seed runs with stratified 80/20 train/test splits. We additionally propose a dual pooling classification head and evaluate it systematically across all four tasks, finding consistent improvements on sentiment analysis and a moderate gain on news category classification for HelaBERT-Small, while the standard [CLS]-linear head remains competitive on news source classification, a headline-level task with short average input length. We release both models to support further research in Sinhala NLP.
Chinese Translation
我们提出了HelaBERT,这是一个基于BERT的掩码语言模型家族,经过从头开始预训练,使用了约10亿个来自MADLAD-400、CulturaX和一个包含新闻文章、僧伽罗语维基百科及网络爬虫数据的自定义语料库的僧伽罗语文本。HelaBERT-Small(约2330万个参数,6层)和HelaBERT-Large(约1.1亿个参数,12层)均使用了针对僧伽罗语的粘合形态和复杂脚本量身定制的SentencePiece Unigram分词器(词汇量为32,000)。我们在四个下游僧伽罗语文本分类任务上评估了这两个模型:新闻类别分类、新闻来源分类、情感分析和写作风格分类,使用了5次独立的种子运行,并采用分层的80/20训练/测试划分。此外,我们还提出了一个双池化分类头,并在所有四个任务中系统地评估其表现,发现HelaBERT-Small在情感分析上有一致的改进,并在新闻类别分类上有适度的提升,而标准的[CLS]-线性头在新闻来源分类这一平均输入长度较短的标题级任务中仍然具有竞争力。我们发布了这两个模型,以支持进一步的僧伽罗语自然语言处理研究。
cs.CL / 99 / 2608.22948

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

证明你错误的依据:对可证伪研究构思的语言模型基准测试
Wang, Ziyue, Yuan, Aomufei, Yao, Yiran, Yao, Linli, Zuo, Hongyao, Gong, Ziwen, Liu, Yuanxin, Li, Shicheng, Cai, Yishuo, Yang, Tong, Sun, Xu, Li, Xiaohui, Bai, Haoli
Abstract
Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper rewards recovery of one realized trajectory. We introduce a benchmark that carries a proposal from Literature to Test: the Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable. Built prospectively from 200 real-paper neighborhoods, Lit2Test elicits proposals from four frontier models and compares them through 1,200 pairwise comparisons judged blind in both presentation orders. The protocol audits its own reliability through diagnostic controls and bounded human calibration, with three annotators corroborating the conclusions within explicitly stated reliability bounds. Lit2Test recovers a strict ranking of the four models in all 10,000 bootstrap replicates, and the separation comes from the quality of the proposed tests and metrics rather than from surface fluency. We release the benchmark, construction pipeline, and audit artifacts for public use.
Chinese Translation
大型语言模型越来越多地被用于提出研究想法,但当前评判这些想法的方式缺乏统一的决策规则:自由形式的评判受到风格和立场的影响,而与后续论文的评分则奖励对已实现轨迹的恢复。我们引入了一个基准测试,将提案从文献转向测试:Lit2Test基准围绕一个可证伪结果组织了六个领域的合同,使得每个提案预先承诺了能够证明其错误的观察,从而使其质量在一开始就可判定,而不仅仅是可以争论的。Lit2Test基于200个真实论文邻域前瞻性构建,从四个前沿模型中引出提案,并通过1200次盲评的成对比较进行比较,评判顺序均为随机。该协议通过诊断控制和有限的人为校准审计自身的可靠性,三位注释者在明确的可靠性范围内证实了结论。Lit2Test在所有10,000次自助法重复中恢复了四个模型的严格排名,而这种区分来自于所提测试和指标的质量,而非表面流畅性。我们发布了该基准、构建流程和审计文档供公众使用。
cs.CL / 100 / 2608.22956

The Illusion of Control: Why Bare Classifier Inversion Silently Fails in Concept-Bottleneck Text Generation

控制的幻觉:为何裸分类器反演在概念瓶颈文本生成中悄然失败
Bing, Qi, Shao, Xiaowei
Abstract
Concept-bottleneck controllable generation routes multi-attribute control through a low-dimensional concept code that, at deployment, must be synthesised from a target attribute configuration. We study this problem in concept-bottleneck text generation under multi-axis compositional generalisation, comparing three ways to obtain the inference-time code: classifier inversion against the encoder heads, reference-text encoding, and a post-hoc label-conditioned prior. Since a concept code admits no direct LM-fluency term, regularising inversion must instead constrain the code toward the encoder's training distribution. We therefore test bare inversion and three regularised variants: label-agnostic and label-conditioned Mahalanobis penalties, and a conditional normalising-flow density baseline. Every inversion variant we test underperforms a simple post-hoc prior fitted to per-combination encoder means on the same checkpoints, across three backbone families spanning $124$M to $8$B parameters. The bare form of classifier inversion also silently collapses to chance, traceable to a directly measured off-manifold code. We validate this diagnosis on real-world benchmarks and under external evaluators, enabling fair comparison with published baselines.
Chinese Translation
概念瓶颈可控生成通过一个低维概念编码实现多属性控制,该编码在部署时必须从目标属性配置中合成。我们在多轴组合泛化的概念瓶颈文本生成中研究这一问题,比较了三种获取推理时编码的方法:针对编码器头的分类器反演、参考文本编码以及后验标签条件先验。由于概念编码不直接包含语言模型流畅性项,因此正则化反演必须将编码约束到编码器的训练分布。因此,我们测试了裸反演和三种正则化变体:与标签无关的和带标签条件的马哈拉诺比斯惩罚,以及条件归一化流密度基线。我们测试的每种反演变体在三个基础模型系列(参数从124M到8B)上均表现不及简单的后验先验,该先验在相同检查点上拟合每组合编码器均值。裸分类器反演的形式也悄然崩溃至随机,追溯到直接测量的离散编码。我们在真实世界基准和外部评估者下验证了这一诊断,从而实现与已发布基线的公平比较。
cs.CL / 101 / 2608.22967

Closed-Loop Bayesian Molecular Inverse Design with Semantic LLM Surrogates

闭环贝叶斯分子反向设计与语义大语言模型代理
Xu, Yaoyao, Zhao, Xinjian, Song, Xiaozhuang, Bai, Lei, Yu, Tianshu
Abstract
Practical molecular inverse design is rarely a one-shot generation problem; it often takes the form of closed-loop candidate-pool enrichment, where under a limited oracle budget the goal is to \emph{increase the fraction of generated molecules that match a desired property profile}. Bayesian optimization (BO) offers a natural framework for this setting, yet standard Gaussian-process surrogates typically operate in compressed continuous embeddings, which discard the substructural and reference-similarity signals that chemists naturally use to decide where to look next. We propose \textbf{\method}, a closed-loop framework in which the surrogate, rather than the generator, is treated as the locus of design choice, and instantiate it with a frozen large language model that reasons directly over the task instruction, SMILES-level optimization history, and oracle feedback in their native textual form. At each iteration, the surrogate returns a structured decision signal that selects informative reference molecules under an exploration and exploitation principle, optionally with a concise guidance sentence. This signal is converted into next-round conditioning text for a frozen molecular generator, yielding an inspectable optimization trace in natural language. Experiments on MolQA drug and material design tasks show that \method improves over one-shot prompting, is competitive with or stronger than GP-based BO baselines, and reveals a domain-dependent interface: reference-only transfer works best for binary drug targets, while adding a concise surrogate summary is more beneficial for continuous material
Chinese Translation
实际的分子反向设计很少是一次性生成问题;它通常表现为闭环候选池的丰富化,在有限的预言者预算下,目标是 extit{增加生成分子中符合期望属性特征的比例}。贝叶斯优化(Bayesian Optimization, BO)为这种设置提供了自然的框架,然而标准的高斯过程代理通常在压缩的连续嵌入中操作,这会丢弃化学家自然用来决定下一步探索方向的子结构和参考相似性信号。我们提出了 extbf{ extit{method}},一个闭环框架,其中代理而非生成器被视为设计选择的核心,并通过一个冻结的大型语言模型来实例化,该模型直接基于任务指令、SMILES级优化历史和以其原生文本形式的预言者反馈进行推理。在每次迭代中,代理返回一个结构化决策信号,该信号在探索与利用原则下选择信息丰富的参考分子,选配一个简洁的指导句。该信号被转换为下一轮条件文本,供冻结的分子生成器使用,从而在自然语言中产生可检查的优化轨迹。在MolQA药物和材料设计任务上的实验表明, extit{method}优于一次性提示,与基于高斯过程的贝叶斯优化基线相竞争或更强,并揭示了一个领域依赖的接口:仅参考转移在二元药物靶点中效果最佳,而添加简洁的代理摘要对连续材料则更具益处。
cs.CL / 102 / 2608.22985

What Does Activation Steering Control? Attribution Across Answer Encodings and Output-Sensitive Subspaces

激活引导控制什么?跨答案编码和输出敏感子空间的归因
Gao, Zhiwei, Peng, Shaowen, Wakamiya, Shoko, Aramaki, Eiji
Abstract
Activation steering is often evaluated under the answer encoding used to construct the direction. A reported gain may reflect the intended judgment or compatibility with answer identifiers seen during construction. We introduce Cross-Encoding Steering Evaluation, which freezes an intervention while re-encoding answers to the same held-out items. On NormBank, after A/B/C identifiers are reassigned, contrastive activation addition (CAA) induces larger target-versus-source score changes for the extraction indices than for the semantic labels under the new mapping. We call this extraction-index following. Varying identifier vocabulary (A/B/C, X/Y/Z, or 1/2/3) and row order shows that the effect tracks extraction index rather than row position. After matching direction norms across layers, extraction-index following emerges mainly at later depths. A low-rank output-sensitive component containing 15.4% of the direction's squared norm retains 96.3% of this effect. An Inference-Time Intervention (ITI)-style method also favors extraction-index over semantic-label following on NormBank in three models. In aggregate, MNLI favors extraction-index following, whereas Social Chemistry 101 (SC101) favors semantic-label following. Multiple-choice and open-ended evaluations can yield different behavioral conclusions. Thus, a steering gain under one answer encoding does not by itself identify what the intervention controls.
Chinese Translation
激活引导通常在用于构建方向的答案编码下进行评估。报告的增益可能反映了预期的判断或与构建过程中看到的答案标识符的兼容性。我们引入了跨编码引导评估(Cross-Encoding Steering Evaluation),该方法在重新编码相同保留项目的答案时冻结干预。在NormBank上,在重新分配A/B/C标识符后,对比激活增加(Contrastive Activation Addition, CAA)在提取指标的目标与源分数变化上产生的影响大于在新映射下的语义标签。我们称之为提取指标跟随(extraction-index following)。不同的标识符词汇(A/B/C、X/Y/Z或1/2/3)和行顺序显示,该效应跟踪的是提取指标而非行位置。在跨层匹配方向规范后,提取指标跟随主要出现在较深的层次。一个包含15.4%方向平方范数的低秩输出敏感成分保留了96.3%的这一效应。在NormBank的三种模型中,推理时间干预(Inference-Time Intervention, ITI)风格的方法也更倾向于提取指标而非语义标签跟随。总体而言,MNLI更倾向于提取指标跟随,而社会化化学101(Social Chemistry 101, SC101)则更倾向于语义标签跟随。多项选择和开放式评估可能会得出不同的行为结论。因此,在一种答案编码下的引导增益本身并不能确定干预控制的内容。
cs.CL / 103 / 2608.22993

LLM Pedagogical Behavior in AI Tutoring Interactions

AI 辅导互动中的 LLM 教学行为
Lee, Suhyeon, Baek, Juneha, Park, Jaehyeong, Shin, Donghyuk
Abstract
Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students use them as tutors in authentic learning interactions. This matters because tutoring responses can differ substantially in how directly they help students complete a task. We operationalize this dimension as scaffolding level and develop a five-level scale, validated against human annotations, that characterizes responses according to the degree of direct assistance they provide. We apply the scale to 14,637 LLM responses from 203 students in a university AI course. Responses are overwhelmingly concentrated at high levels of assistance, with more than 95% classified as either Explaining or Solving. Scaffolding level is systematically associated with students' subsequent conversational behavior, but provides little additional predictive information about performance on three subsequent exams beyond prior achievement and dialogue behavior. These findings provide an empirical baseline for LLM assistance in tutoring interactions and a measurement framework for evaluating how alternative tutoring designs change that assistance.
Chinese Translation
学生们越来越多地将大型语言模型(LLMs)作为课程和问题解决的辅导工具。然而,对于 LLM 在真实学习互动中作为辅导者时所提供的帮助程度知之甚少。这一点至关重要,因为辅导响应在帮助学生完成任务的直接性上可能存在显著差异。我们将这一维度操作化为支架水平,并开发了一个五级量表,该量表经过人类注释的验证,能够根据提供的直接帮助程度对响应进行特征化。我们将该量表应用于来自203名大学人工智能课程学生的14,637条 LLM 响应。这些响应主要集中在高水平的帮助上,超过95%被分类为解释(Explaining)或解决(Solving)。支架水平与学生后续的对话行为系统性相关,但在预测三次后续考试的表现时,除了先前的成就和对话行为外,提供的额外信息有限。这些发现为 LLM 在辅导互动中的帮助提供了实证基线,并为评估不同辅导设计如何改变这种帮助提供了测量框架。
cs.CL / 104 / 2608.23020

Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality

遗忘不仅仅是抹除:通过生成不平等实现时间解耦
Chen, Xunlei, Ye, Qirui, Li, Yuang, Gong, Yi, Wang, Zhaokun, Li, Wenyi, Guo, Shiyao, Guo, Jinyu
Abstract
Large language models (LLMs) require effective unlearning to address privacy regulations and safety concerns. However, achieving precise forgetting without compromising general utility remains challenging. Existing sequence- and token-level methods penalize target outputs without modeling their context-dependent retrieval paths, which can disrupt linguistic structure or suppress benign knowledge. We present ADU, a fine-grained, training-based framework that shifts unlearning from token erasure to contextual attention-pathway decoupling. Exploiting the functional distinction between local and global attention heads, ADU identifies preplan positions that retrieve persistent sensitive anchors and fixes their candidate paths under the original model. It then trains attention-projection adapters to suppress attention mass along these paths while preserving local-attention structure and retain-set language modeling. Post-training activation exchange tests whether the modified attention-output module transmits the learned forgetting effect. ADU achieves the strongest aggregate performance among evaluated baselines on the TOFU and WMDP benchmarks, including a Forget Quality of (0.93) on TOFU. It preserves 87--98% of model utility (92.9% on average versus 81.9% for baselines) while reducing side effects in benign contexts.
Chinese Translation
大型语言模型(LLMs)需要有效的遗忘机制,以应对隐私法规和安全问题。然而,在不影响整体效用的情况下实现精确的遗忘仍然具有挑战性。现有的序列和令牌级方法在没有建模其上下文相关检索路径的情况下惩罚目标输出,这可能会破坏语言结构或抑制良性知识。我们提出了ADU,一个细粒度的基于训练的框架,将遗忘从令牌抹除转变为上下文注意路径解耦。ADU利用局部和全局注意头之间的功能区分,识别出检索持久敏感锚点的预先计划位置,并在原始模型下修正其候选路径。然后,它训练注意投影适配器,以在这些路径上抑制注意力集中,同时保留局部注意结构和保留集语言建模。后训练激活交换测试修改后的注意输出模块是否传递了学习到的遗忘效果。ADU在TOFU和WMDP基准测试中在评估的基线中实现了最强的整体性能,包括在TOFU上达到的遗忘质量为(0.93)。它保留了87%至98%的模型效用(平均92.9%,而基线为81.9%),同时减少了良性上下文中的副作用。
cs.CL / 105 / 2608.23023

Most of the LLM routing gap is task type

大多数 LLM 路由差距源于任务类型
Lee, Janghoon
Abstract
An LLM router picks which model should answer each query. The appeal is that models fail on different questions. Whatever single model is best overall still gets some wrong, and another model in the pool gets many of those right. Getting that choice right every time is the ceiling, and a router is an attempt to approach it. However, recent work reports that routers do not get close. Across 21 routing methods on five benchmarks, sharply different designs land within a fraction of a point of each other, and all of them stay far below that ceiling. Learned routers often fail to beat simply always calling the strongest model. We ask what those missed questions have in common. We set fourteen models to answer all 294 questions, with 7 task types across 3 languages: Korean, English and Hindi. We ran the whole matrix twice, changing nothing, but 5.37% of the 4,116 model-question pairs came out scored differently anyway. Run-to-run movement like that is normal, and we argue that a small win does not show that routing did anything, ours or anyone else's. Counting an answer correct only when the model got it right in both runs, 29 questions on this matrix can be improved with routing. Every correct-answer count here is on that rule. Task type accounts for most of them: assigning each task type one model in advance, chosen once and never updated, improves 21 of the 29. Splitting each task type by language improves 2 more and leaves 6 of 294 unoptimized. That handful is what a learned router would have been built for, and it is smaller than the run-to-run movement above, which is a share of pairs rather than of questions. The static table we adopted answers 262 of 294 questions at $3.33 per run, against the best single model's 245 at $7.69. All of this is fitted and scored on the same 294 questions with no holdout.
Chinese Translation
LLM 路由器选择哪个模型应回答每个查询。其吸引力在于不同模型在不同问题上表现不一。无论哪个单一模型整体表现最佳,仍然会出现一些错误,而池中的另一个模型则能正确回答其中的许多问题。每次都做出正确选择是理论上的极限,而路由器则是试图接近这一极限。然而,最近的研究报告显示,路由器并未接近这一目标。在五个基准测试中的 21 种路由方法中,设计截然不同的模型在得分上相差无几,且所有模型的表现均远低于这一极限。学习型路由器往往无法超越简单地始终调用最强模型的策略。我们探讨这些未能回答的问题有什么共同点。我们设置了十四个模型来回答所有 294 个问题,涵盖 3 种语言(韩语、英语和印地语)中的 7 种任务类型。我们进行了两次完整的矩阵测试,未做任何更改,但 4,116 个模型-问题对中有 5.37% 的得分结果依然不同。这样的运行间波动是正常的,我们认为小幅胜利并不能证明路由器的有效性,无论是我们的方法还是其他人的方法。仅当模型在两次运行中均正确时,才将答案计为正确,在这个矩阵中,有 29 个问题可以通过路由得到改善。这里的每个正确答案计数均基于这一规则。任务类型占据了大多数:提前为每个任务类型分配一个模型,并且只选择一次而不更新,能改善 29 个问题中的 21 个。按语言划分每个任务类型又改善了 2 个问题,仍有 294 个问题中有 6 个未优化。这些问题正是学习型路由器所应解决的,而其数量小于上述的运行间波动,后者是对模型-问题对的比例,而非问题的比例。我们采用的静态表格以每次运行 $3.33 的成本回答了 294 个问题中的 262 个,而最佳单一模型的成本为 $7.69,仅回答了 245 个问题。所有这些都是在相同的 294 个问题上进行拟合和评分的,没有留出数据。
cs.CL / 106 / 2608.23026

Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs

超越表面线索:解构多语言大型语言模型中的社会文化信号
Feng, Yuanjun, Liu, Tanzhou, Feuerriegel, Stefan, Shrestha, Yash Raj
Abstract
Multilingual LLM outputs can vary across sociocultural contexts. However, evidence of cultural grounding can be misleading: identity labels may be inferred from explicit or indirect textual cues, while names and wording can reveal the source language. Treating all these signals as evidence of cultural grounding may obscure potential biases. We present a human-validated, multi-agent audit that separates three questions: whether outputs reproduce social biases, whether identity groups are represented differently, and whether outputs reflect cross-cultural patterns. The study analyzes 89,253 outputs from 12 LLMs in English, French, and Chinese, spanning 18 occupations and three task conditions. We find that bias representation varies systematically across languages and tasks. Removing direct identity cues sharply reduces identity-label prediction in English and Chinese, but has a much smaller effect in French. Across all language-genre settings, the cultural context associated with the source language receives the highest average relevance score, with moderate agreement between automated and human ratings. However, the ability to identify the source language drops substantially after translation and again after masking names. Without these controls, multilingual audits may mistake surface cues for cultural understanding, leading to misleading conclusions about cross-cultural variation and bias. Our audit offers a practical framework for separating such shortcuts from more meaningful cross-cultural patterns.
Chinese Translation
多语言大型语言模型(LLMs)的输出在社会文化背景中可能存在差异。然而,文化基础的证据可能会产生误导:身份标签可能是通过明确或间接的文本线索推断出来的,而名字和措辞则可以揭示源语言。将所有这些信号视为文化基础的证据可能会掩盖潜在的偏见。我们提出了一种经过人类验证的多代理审计,分离三个问题:输出是否再现社会偏见,身份群体是否以不同方式被代表,以及输出是否反映跨文化模式。该研究分析了来自12个LLMs的89,253个输出,涵盖英语、法语和中文,涉及18个职业和三种任务条件。我们发现偏见的表现因语言和任务而系统性地变化。在英语和中文中,去除直接身份线索显著减少了身份标签的预测,但在法语中影响较小。在所有语言-体裁设置中,与源语言相关的文化背景获得了最高的平均相关性评分,自动评分与人类评分之间存在中等一致性。然而,翻译后以及在掩盖名字后,识别源语言的能力显著下降。如果没有这些控制,多语言审计可能会将表面线索误认为文化理解,从而导致对跨文化变异和偏见的误导性结论。我们的审计提供了一个实用框架,用于将此类捷径与更有意义的跨文化模式区分开来。
cs.CL / 107 / 2608.23029

Meta-Moderator: Empowering Multi-Agent Debate with Meta-Cognition

元调解者:通过元认知增强多智能体辩论
Hu, Wentao, Wan, Zhuoyue, Shen, Jinhao, Zhang, Chen Jason, Wei, Xiaoyong, Li, Qing
Abstract
Multi-agent debate can improve large language model reasoning by eliciting diverse hypotheses and critiques, yet its performance is often constrained by weak moderation. Common pipelines rely on fixed budgets, agreement-based stopping, or untrained judges, leading to redundant deliberation and unreliable evidence aggregation. We cast moderation as a meta-cognitive process, monitoring debate utility, controlling deliberation, and adjudicating a final answer, and introduce Meta-Moderator, a learnable framework that dynamically regulates debate and decides when to finalize an answer. Meta-Moderator is trained independently of the debaters via outcome-driven policy optimization, making debate regulation an explicit capability rather than an incidental effect of prompting. Across five benchmarks, Meta-Moderator outperforms widely used decision layers and transfers across tasks and system configurations. Further analyses show that it allocates debate more selectively and reduces mis-aggregation after informative hypotheses appear.
Chinese Translation
多智能体辩论可以通过引发多样的假设和批评来提升大型语言模型的推理能力,但其表现常常受到弱调解的限制。常见的流程依赖于固定预算、基于一致性的停止或未经训练的评审,导致冗余的讨论和不可靠的证据聚合。我们将调解视为一种元认知过程,监控辩论效用、控制讨论并裁定最终答案,并引入了元调解者(Meta-Moderator),这是一个可学习的框架,能够动态调节辩论并决定何时确定答案。元调解者通过结果驱动的策略优化独立于辩论者进行训练,使得辩论调节成为一种明确的能力,而不是提示的偶然效果。在五个基准测试中,元调解者的表现优于广泛使用的决策层,并能够在任务和系统配置之间转移。进一步的分析表明,它更具选择性地分配辩论,并在出现有信息的假设后减少错误聚合。
cs.CL / 108 / 2608.23037

The Multilingual FrameNet Corpus

多语言 FrameNet 语料库
Fiumanò, Beatrice, Lazzari, Nicolas, Ponzetto, Simone Paolo, Presutti, Valentina
Abstract
This paper introduces the Multilingual FrameNet Corpus (mFNC), a novel resource that extends the English Berkeley FrameNet corpus by collecting and harmonizing existing language-specific corpora across nine additional languages: Brazilian Portuguese, Chinese, Dutch, French, German, Italian, Korean, Latvian and Swedish. By training models that rely on different architectures on the mFNC, we consistently outperform existing state-of-the-art Frame Semantic Parsers in both multilingual and cross-lingual settings, underscoring the importance of multilingual training data. The mFNC and our trained FSP models are openly available at https://github.com/beatrice-f/mFNC.
Chinese Translation
本文介绍了多语言 FrameNet 语料库(Multilingual FrameNet Corpus, mFNC),这是一个新颖的资源,通过收集和协调九种额外语言(巴西葡萄牙语、中文、荷兰语、法语、德语、意大利语、韩语、拉脱维亚语和瑞典语)的现有语言特定语料库,扩展了英语伯克利 FrameNet 语料库。通过在 mFNC 上训练依赖于不同架构的模型,我们在多语言和跨语言环境中始终超越现有的最先进的框架语义解析器,强调了多语言训练数据的重要性。mFNC 及我们的训练的框架语义解析器模型可在 https://github.com/beatrice-f/mFNC 上公开获取。
cs.CL / 109 / 2608.23047

Beyond Verdicts: A Graph-Based Analysis of Human and LLM Reasoning in Scientific Fact-Checking

超越裁决:基于图的科学事实核查中人类与大型语言模型(LLM)推理的分析
Ghafoor, Abdul, Manzoor, Muhammad Arslan, Hou, Yufang
Abstract
Misinformation that cites legitimate papers can be especially harmful when it distorts what those studies actually report. While existing automatic fact-checking systems based on large language models (LLMs) can assess whether a model assigns an Incorrect verdict and can gen- erate explanations for that decision, they typi- cally do not indicate whether the model follows the same reasoning path as human experts or arrives at the verdict through a different but still valid path. In this work, we introduce a graph- based framework (typed reasoning graph) for comparing human and LLM reasoning paths in scientific fact-checking. Building on prior work on fallacious reasoning in biomedical misinformation, MISSCIPLUS (Glockner et al., 2025), we model each explanation as a rea- soning graph that links the false claim to the relevant study context, study findings, fallacy- supporting premises, and fallacy labels. This representation enables one-to-one alignment of human and LLM reasoning at the level of fallacy-specific sub-graphs. For non-human- aligned LLM paths, we validate grounding in the cited study, relevance to the claim, and suf- ficiency for the verdict. Using 84 false claims from MISSCIPLUS, we evaluate GPT-5, Claude Opus 4.7, and Qwen3-32B across prompt and evidence settings. Results show distinct perfor- mance dimensions: Qwen3-32B has the lowest verdict failure rate, GPT-5 the highest human alignment, and Claude Opus 4.7 weak verdict prediction but often valid reasoning in success- ful cases
Chinese Translation
引用合法论文的错误信息在扭曲这些研究实际报告内容时可能尤其有害。尽管现有基于大型语言模型(LLMs)的自动事实核查系统能够评估模型是否给出错误裁决,并生成该决策的解释,但它们通常不指明模型是否遵循与人类专家相同的推理路径,或通过不同但仍然有效的路径得出裁决。在本研究中,我们引入了一种基于图的框架(类型化推理图)用于比较科学事实核查中人类与LLM的推理路径。基于先前关于生物医学错误信息中谬误推理的研究(MISSCIPLUS,Glockner等,2025),我们将每个解释建模为一个推理图,将虚假主张与相关研究背景、研究发现、支持谬误的前提和谬误标签相连接。这种表示方式使得人类与LLM的推理在谬误特定子图层面上实现一对一对齐。对于非人类对齐的LLM路径,我们验证其在引用研究中的基础、与主张的相关性以及裁决的充分性。使用来自MISSCIPLUS的84个虚假主张,我们评估了GPT-5、Claude Opus 4.7和Qwen3-32B在提示和证据设置下的表现。结果显示出不同的性能维度:Qwen3-32B具有最低的裁决失败率,GPT-5则具有最高的人类对齐度,而Claude Opus 4.7在成功案例中表现出较弱的裁决预测但通常具有有效的推理。
cs.CL / 110 / 2608.23067

Signal or Noise? A Benchmark Study of Agent Skills in Web Development

信号还是噪声?代理技能在网页开发中的基准研究
Yang, Ziyue, Ding, Fan
Abstract
Agent Skills are reusable procedural modules that are increasingly injected into coding-agent sessions to encode framework conventions, anti-patterns, and reusable tools. However, because each injected Skill expands the prompt of every query, an effective Skill benchmark must determine not only whether an agent can solve a task, but whether the Skill should have been injected at all. We introduce WebDev-Skills-Bench and use it for a controlled empirical study of 31 public WebDev Skills on 50 Web-Bench projects and 1,000 ordered tasks. The benchmark compares four matched conditions, including a length-matched irrelevant control and leave-one-out component ablations. To isolate Skill effects from prompt-length artifacts, we place only SKILL.md in the prompt while mounting auxiliary files into the agent workspace. Across four models, target Skill injection reduces mean Pass@2 by 1.3% to 4.2%, lowers task completion depth, and increases token cost by 72% to 394%, with gains in only 17% to 36% of Skill-project pairs. Length-matched controls reveal two failure modes: some models are length-distracted, where an equally long irrelevant Skill reproduces most of the loss, while others are content-misled, where prompt length is neutral but Skill content still lowers Pass@2 by 1.1% to 1.4%. Further analysis shows that losses concentrate on easy early tasks, Skill rankings transfer weakly across models, and anti-pattern rules outperform example-heavy content within helpful Skills. These findings recast a matched Skill as a hypothesis about a particular Skill-project-model triple rather than a portable asset, reframing injection as a per-deployment routing decision and making length-matched controls and per-model audits a minimum standard for Agent-Skill evaluation.
Chinese Translation
代理技能是可重用的过程模块,越来越多地被注入到编码代理会话中,以编码框架约定、反模式和可重用工具。然而,由于每个注入的技能都会扩展每个查询的提示,因此有效的技能基准必须确定代理不仅能否解决任务,还要判断该技能是否应该被注入。我们引入了WebDev-Skills-Bench,并利用它对31个公共WebDev技能在50个Web-Bench项目和1,000个有序任务上进行控制实证研究。该基准比较了四种匹配条件,包括长度匹配的无关控制和逐一排除组件的消融实验。为了将技能效应与提示长度伪影隔离,我们仅在提示中放置SKILL.md,同时将辅助文件装载到代理工作区。在四个模型中,目标技能注入使平均Pass@2降低了1.3%到4.2%,降低了任务完成深度,并使令牌成本增加了72%到394%,而仅在17%到36%的技能-项目对中获得收益。长度匹配的控制揭示了两种失败模式:一些模型受到长度干扰,其中一个同样长的无关技能复现了大部分损失,而其他模型则受到内容误导,尽管提示长度是中性的,但技能内容仍使Pass@2降低了1.1%到1.4%。进一步分析表明,损失集中在简单的早期任务上,技能排名在模型间转移较弱,而反模式规则在有用技能中优于内容繁重的示例。这些发现将匹配的技能重新定义为关于特定技能-项目-模型三元组的假设,而不是可移植资产,将注入重新构架为每次部署的路由决策,并使长度匹配的控制和每个模型的审计成为代理-技能评估的最低标准。
cs.CL / 111 / 2608.23095

Definitional Sensitivity in Media Bias Detection: A Multi-Definition Dataset and Benchmark

媒体偏见检测中的定义敏感性:一个多定义数据集和基准
Wessel, Martin, Spinde, Timo, Pfeffer, Jürgen, Demartini, Gianluca
Abstract
Media bias detection relies on definitions and examples that specify what counts as bias, yet these specifications often vary across datasets or remain implicit, even when given the same name. Such variation makes it unclear whether models trained for the same bias category learn the same construct or different phenomena, a problem largely overlooked in prior work. We examine how definition choice affects bias annotation in a between-subjects experiment with 354 participants and a parallel evaluation with four LLMs. Participants and models rate six news articles across four bias categories using definitions that vary in conceptual framing and elaboration. Across 8,496 human and 28,800 LLM ratings, we find that the conceptual target of a definition drives annotation divergence, while construct-preserving elaboration does not: conceptual framing significantly shifts annotations for humans and does so even more strongly for LLMs. We discuss implications for construct specification in annotation protocols and prompt-based measurement, and consider how definitional sensitivity may propagate to downstream classification beyond media bias. We also release MUDD, the Multi-Definition Bias Detection Dataset.
Chinese Translation
媒体偏见检测依赖于定义和示例,这些定义和示例明确了什么算作偏见,然而这些规范在不同数据集中往往存在差异,或者即使在同名情况下也保持隐含。这种变异使得不清楚为同一偏见类别训练的模型是否学习了相同的构念或不同的现象,这一问题在以往的研究中大多被忽视。我们通过一项涉及354名参与者的被试间实验以及与四个大型语言模型(LLMs)的平行评估,考察了定义选择如何影响偏见标注。参与者和模型使用在概念框架和详细程度上有所不同的定义,对六篇新闻文章在四个偏见类别中进行评分。在8496条人类评分和28800条LLM评分中,我们发现定义的概念目标驱动了标注的差异,而保持构念的详细阐述并未产生显著影响:概念框架显著改变了人类的标注,并且对LLM的影响更为显著。我们讨论了在标注协议和基于提示的测量中构念规范化的影响,并考虑定义敏感性如何可能传播到媒体偏见之外的下游分类。我们还发布了MUDD(多定义偏见检测数据集)。
cs.CL / 112 / 2608.23104

Molecular LLM Agents: From Architectural Design to Scientific Autonomy

分子 LLM 代理:从架构设计到科学自主性
Li, Jiatong, Zhang, Wengyu, Wang, Weida, Ren, Yuxuan, Liu, Wei, Mao, Chenyang, Li, Yuqiang, Bian, Yatao, Zheng, Changmeng, Wei, Xiaoyong, Li, Qing
Abstract
Molecular science represents an important frontier for LLM-based agents. Unlike general agents that mainly operate over natural language, code, or web environments, molecular LLM agents must perceive, reason about, and act upon chemical objects across symbolic strings, molecular graphs, 3D conformations, spectra, simulations, and wet-lab measurements. Their capabilities depend on chemically faithful molecular perception, an LLM-centered agent framework, domain-specific tool grounding, and computational or experimental feedback, in addition to planning and tool use. This work develops a conceptual framework for molecular LLM agents from two complementary perspectives. First, we introduce an architectural view of molecular-agent design, covering molecular representation and perception, the agent framework, domain-specific toolboxes, and learning and optimization. Second, we propose a scientific autonomy ladder inspired by staged autonomy in engineering systems, categorizing agents into four levels: L1 assistive or fixed workflows, L2 adaptive computational agents, L3 feedback-aware physical experiment agents, and L4 scientific-agenda agents. Together, these two perspectives establish a comprehensive framework for comparing existing molecular LLM agents, identifying missing capabilities and deployment risks, and guiding the design, evaluation, and deployment of future agents in molecular discovery workflows.
Chinese Translation
分子科学代表了基于 LLM 的代理的重要前沿。与主要在自然语言、代码或网络环境中操作的一般代理不同,分子 LLM 代理必须感知、推理并对化学对象进行操作,这些对象包括符号字符串、分子图、三维构象、光谱、模拟以及湿实验测量。它们的能力依赖于化学上真实的分子感知、以 LLM 为中心的代理框架、特定领域的工具基础以及计算或实验反馈,此外还包括规划和工具使用。本研究从两个互补的角度发展了分子 LLM 代理的概念框架。首先,我们介绍了分子代理设计的架构视角,涵盖了分子表示与感知、代理框架、特定领域工具箱以及学习与优化。其次,我们提出了一个受工程系统分级自主性启发的科学自主性阶梯,将代理分为四个级别:L1 辅助或固定工作流程,L2 自适应计算代理,L3 反馈感知的物理实验代理,以及 L4 科学议程代理。这两个视角共同建立了一个全面的框架,用于比较现有的分子 LLM 代理,识别缺失的能力和部署风险,并指导未来在分子发现工作流程中代理的设计、评估和部署。
cs.CL / 113 / 2608.23120

Statistical Machine Translation Systems of English-Pnar Language Pair : Some Insights of the Emperical Study

英-普纳尔语对统计机器翻译系统:实证研究的若干见解
Thokchom, Edawanbiang Dhar Surmila, Singh, Thoudam Doren
Abstract
Pnar, an Austroasiatic language spoken by approximately 0.4 million people in the Jaintia Hills of Meghalaya, lacks the digital corpora and natural language processing (NLP) resources. This paper presents the first machine translation study for the English and Pnar language pair. Using articles collected from the Wyrta newspaper, we built a parallel corpus comprising of 10,234 sentences and trained phrase-based statistical machine translation (SMT) systems the models using 9,563 parallel corpora under three configurations for each direction using Moses, GIZA++ , KenLM, varying lexicalized reordering and minimum error rate training (MERT) tuning. The models are evaluated on a held out test set of 371 sentences, the best performing system achieves a BLEU score of 14.97 (chrF2: 33.42, TER: 77.60) for Pnar to English and 11.16 (chrF2: 31.38, TER: 93.51) for English to Pnar, establishing the first quantitative benchmark for this language pair. Lexicalized reordering improves translation quality by 3.73 BLEU points for Pnar to English, reflecting the structural shift from the source language's SOV word order to the target language's SVO order, whereas MERT tuning degrades BLEU performance under low resource conditions. Finally, we analyze the remaining translation errors, including morphological out of vocabulary (OOV) words, long-distance reordering and Khasi code mixing and discuss future directions toward neural and multilingual machine translation for Pnar.
Chinese Translation
普纳尔语(Pnar)是一种奥斯特罗亚细亚语系语言,约有40万人在梅加拉亚邦的贾因蒂亚山脉使用,但缺乏数字语料库和自然语言处理(NLP)资源。本文首次针对英-普纳尔语对开展机器翻译研究。利用从《Wyrta》报纸收集的文章,构建了包含10,234句子的平行语料库,并基于9,563句平行语料,采用Moses、GIZA++、KenLM工具,在三个配置下训练了基于短语的统计机器翻译(SMT)模型,分别针对两个翻译方向,调整词汇化重排序和最小错误率训练(MERT)参数。模型在一组371句的测试集上进行评估,表现最佳的系统在普纳尔语到英语方向取得了BLEU分数14.97(chrF2: 33.42,TER: 77.60),英语到普纳尔语方向取得了11.16(chrF2: 31.38,TER: 93.51),为该语言对建立了首个量化基准。词汇化重排序提升了普纳尔语到英语方向3.73个BLEU分,反映了源语言SOV语序向目标语言SVO语序的结构转换,而在资源匮乏条件下,MERT调优反而降低了BLEU表现。最后,本文分析了剩余的翻译错误,包括形态学上的未登录词(OOV)、长距离重排序以及喀西语(Khasi)代码混合现象,并讨论了普纳尔语神经机器翻译和多语言机器翻译的未来方向。
cs.CL / 114 / 2608.23124

LITERARYBIGFIVE: Author-Personalized Text Generation in a Unified Interpretable Space

文学大五:在统一可解释空间中的作者个性化文本生成
Zhang, Jinghui, Gao, Lang, Li, Ao, Li, Mingzhe, Zeng, Ruihong, Song, Zirui, Inui, Kentaro, Chen, Xiuying
Abstract
Personalized text generation for authors and literary writing is essential for applications such as adaptive writing assistants, creative support tools, and computational literary analysis. However, existing approaches to author modeling and personalization often represent writing behavior as independent labels, requiring large-scale corpus collection or fine-tuning for each author or stylistic category. Such formulations are costly, difficult to interpret, and poorly suited for generalizing across authors. Inspired by the Big Five model's dimensional view of personality, we propose LiteraryBigFive, a framework that reframes authorial writing characteristics as coordinates within a unified and interpretable space. In this space, we derive each interpretable axis (e.g., Classicism, Emotionality) from activation-space contrasts between author-written and neutral passages, yielding distinct stylistic dimensions that allow texts or authors to be positioned within a five-dimensional system. Beyond localizing different authors, we further introduce an interpretable steering mechanism, which adaptively guides text generation toward target coordinates to perform author-personalized writing. Experimental results show that LiteraryBigFive improves authorial expressiveness while preserving semantic fidelity. The derived author per-axis scores strongly correlate with real-world literary consensus, offering transparent and interpretable explanations of author-specific generation behavior: https://github.com/Znull-1220/LiteraryBigFive.
Chinese Translation
针对作者和文学创作的个性化文本生成对于自适应写作助手、创意支持工具和计算文学分析等应用至关重要。然而,现有的作者建模和个性化方法通常将写作行为表示为独立标签,这需要大规模语料库的收集或对每个作者或风格类别进行微调。这种表述方式成本高昂,难以解释,并且不适合跨作者进行泛化。受到大五人格模型对个性的维度视角的启发,我们提出了LiteraryBigFive,一个将作者写作特征重新框架为统一且可解释空间中的坐标的框架。在这个空间中,我们通过作者写作段落与中性段落之间的激活空间对比推导出每个可解释轴(例如,古典主义、情感性),从而产生独特的风格维度,使文本或作者能够在五维系统中定位。除了定位不同的作者外,我们进一步引入了一种可解释的引导机制,该机制自适应地引导文本生成朝向目标坐标,以实现作者个性化写作。实验结果表明,LiteraryBigFive在提高作者表现力的同时保持了语义的忠实度。推导出的每个轴的作者得分与现实世界的文学共识高度相关,提供了对作者特定生成行为的透明和可解释的说明。
cs.CL / 115 / 2608.23149

Language Chain in Alignment: Cross-Lingual Ranking Preference Optimization

对齐中的语言链:跨语言排名偏好优化
Lee, Seungyoon, Kim, Minhyuk, Lee, Jungseob, Lim, Heuiseok
Abstract
The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often leads to suboptimal performance in other languages. In this paper, we propose Cross-Lingual Ranking Preference Optimization (CRPO), a novel framework that leverages robust preference knowledge from English to facilitate preference alignment in the target language. We design a hierarchical structure within parallel preference pairs across the target language and English to jointly optimize intra- and inter-lingual preferences, thereby enhancing language adaptation and output quality. Building on the LambdaLoss framework, CRPO goes beyond the binary comparison based optimization by providing a relative ranking signal across multiple candidate responses. Our experiments across five languages with varying resource scales demonstrate that CRPO consistently outperforms standard approaches in both instruction-following and knowledge utilization capability. Notably, the robust performance gains observed across various weighting schemes further validate the empirical effectiveness of our hierarchical design in a multilingual setup. Furthermore, our findings highlight that CRPO significantly improves both reward margins and the log-probability of desirable responses, contributing to a more stable preference manifold for cross-lingual alignment.
Chinese Translation
大型语言模型的对齐在很大程度上依赖于以英语为中心的高质量偏好数据,这常常导致其他语言的表现不佳。本文提出了跨语言排名偏好优化(Cross-Lingual Ranking Preference Optimization, CRPO),这是一个新颖的框架,利用来自英语的强大偏好知识来促进目标语言中的偏好对齐。我们设计了一个层次结构,在目标语言和英语之间的平行偏好对中共同优化语言内部和语言间的偏好,从而增强语言适应性和输出质量。基于LambdaLoss框架,CRPO超越了基于二元比较的优化,通过在多个候选响应之间提供相对排名信号。我们在五种不同资源规模的语言上进行的实验表明,CRPO在遵循指令和知识利用能力方面始终优于标准方法。值得注意的是,在各种加权方案下观察到的强大性能提升进一步验证了我们在多语言环境中层次设计的实证有效性。此外,我们的研究结果强调,CRPO显著提高了奖励边际和理想响应的对数概率,为跨语言对齐提供了更稳定的偏好流形。
cs.CL / 116 / 2608.23152

Counter with Evidence! A Multi-Agent Memory Efficient Reasoning Framework for Hate Category Informed Counterspeech Generation

以证据为基础的反击!一种多智能体内存高效的仇恨类别知情反击生成框架
Nath, Sujoy, Kumar, Aswini, Chakraborty, Tanmoy
Abstract
Counterspeech effectively neutralizes the impact of online hate. Although prior work explores automated counterspeech generation, it largely emphasizes stylistic control while treating hate speech as homogeneous, overlooking that distinct forms of abuse require fundamentally different counterspeech strategies. To address this gap, we introduce FIRE (Factuality Informed Multi-Agent Reasoning Framework) that first decomposes hate speech into one of the five distinct categories (misinformation, stereotype, conspiracy, dehumanizing, non-factual), and then maps it to a targeted counterspeech style. To facilitate FIRE, we curate FactualCS, a novel dataset of $4,784$ instances that provides the annotations regarding hate categories, reasoning traces, and evidence mappings, which are critical elements for grounded generation that are missing in prior work. A comprehensive evaluation across $28$ baseline configurations demonstrates that FIRE significantly surpasses existing methods, despite using compact agents ($<$2B). FIRE achieves a $\sim$ $12 \%$ and $\sim$ $11 \%$ improvements in factual and category-specific accuracy respectively, while simultaneously reducing toxicity by $\sim$ $11 \%$ relative to the strongest baselines. Further human evaluation confirms that responses generated by FIRE are significantly preferred over the strongest baselines, underscoring its effectiveness for real-world deployment. These findings show that decomposing the underlying intent of hate speech is essential for generating safe, effective, and contextually precise counterspeech.
Chinese Translation
反击言论有效地中和了在线仇恨的影响。尽管之前的研究探索了自动化反击言论生成,但大多强调风格控制,同时将仇恨言论视为同质的,忽视了不同形式的虐待需要根本不同的反击策略。为了解决这一问题,我们引入了FIRE(事实知情多智能体推理框架),该框架首先将仇恨言论分解为五个不同类别之一(错误信息、刻板印象、阴谋论、非人化、非事实),然后将其映射到目标反击风格。为了支持FIRE,我们整理了FactualCS,这是一个包含$4,784$个实例的新数据集,提供了关于仇恨类别、推理轨迹和证据映射的注释,这些都是之前工作中缺失的基础生成的重要元素。对$28$个基线配置的全面评估表明,尽管使用紧凑的智能体($<$2B),FIRE显著超越了现有方法。FIRE在事实准确性和类别特定准确性上分别实现了约$12 ext{ extperthousand}$和约$11 ext{ extperthousand}$的提升,同时相对于最强基线将毒性降低了约$11 ext{ extperthousand}$。进一步的人类评估确认,FIRE生成的响应显著优于最强基线,突显了其在现实世界应用中的有效性。这些发现表明,分解仇恨言论的潜在意图对于生成安全、有效且上下文精确的反击言论至关重要。
cs.CL / 117 / 2608.23167

Accelerating Diffusion Language Models via Structured Suffix Modeling

通过结构化后缀建模加速扩散语言模型
Cheng, Zifeng, Li, Keda, Jiang, Zhiwei, Wang, Cong, Shen, Fei, Gu, Qing
Abstract
Diffusion Language Models (DLMs) exhibit strong parallel decoding capabilities by denoising multiple tokens in a single generation step. However, this parallelism comes with substantial computational overhead, as each step requires interactions with all suffix tokens. Existing methods typically reduce this cost by retaining only a local suffix window as a substitute for the full suffix. Despite their effectiveness, these methods overlook the structural heterogeneity across suffix regions and re-initialize suffix tokens with identical representations at each timestep. To this end, we propose a structured suffix modeling method for efficient DLM inference. Specifically, we divide the suffix into three regions, i.e., the local, middle, and tail regions, and retain different numbers of suffix tokens in each region according to their structural roles. Moreover, we incorporate the decoding results from the previous step into the suffix token representations at the current step, allowing them to carry evolving denoising information across generation steps. Notably, our method is training-free and orthogonal to several existing acceleration techniques, such as parallel decoding strategies and KV cache. Empirical results across multiple benchmarks on three DLMs demonstrate that our method can further accelerate DLM inference and improve performance in most cases. In particular, in long-sequence inference, our method achieves up to a \(72.81\times\) speedup when combined with other acceleration techniques. Our code is available at https://github.com/zifengcheng/SSM.
Chinese Translation
扩散语言模型(DLMs)通过在单次生成步骤中去噪多个标记,展现出强大的并行解码能力。然而,这种并行性伴随着显著的计算开销,因为每一步都需要与所有后缀标记进行交互。现有方法通常通过仅保留一个局部后缀窗口来替代完整后缀,从而降低这一成本。尽管这些方法有效,但它们忽视了后缀区域之间的结构异质性,并在每个时间步重新初始化后缀标记为相同的表示。为此,我们提出了一种结构化后缀建模方法,以实现高效的DLM推理。具体而言,我们将后缀划分为三个区域,即局部区域、中间区域和尾部区域,并根据它们的结构角色在每个区域保留不同数量的后缀标记。此外,我们将上一步的解码结果融入当前步骤的后缀标记表示中,使其能够在生成步骤之间携带不断演变的去噪信息。值得注意的是,我们的方法不需要训练,并且与多种现有加速技术(如并行解码策略和KV缓存)是正交的。在三个DLM的多个基准测试中的实证结果表明,我们的方法可以进一步加速DLM推理,并在大多数情况下提高性能。特别是在长序列推理中,我们的方法在与其他加速技术结合时实现了高达72.81倍的加速。我们的代码可在 https://github.com/zifengcheng/SSM 获取。
cs.CL / 118 / 2608.23172

CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension

CaRGo-T:因果推理图思维提升多模态幽默理解
Nandy, Abhilash, Seetharaman, Rahul, Bansal, Aman, Saha, Rounak, Kapadnis, Manav Nitin, Das, Millon Madhur, Goyal, Pawan, Ganguly, Niloy
Abstract
Large-scale vision-language models (VLMs) have demonstrated remarkable versatility across a wide range of multimodal tasks. However, understanding humor remains challenging because humorous content often depends on subtle interactions among entities, events, context, and implicit relationships across image and text modalities. These interactions can involve complex chains of reasoning that are difficult to capture through conventional prompting or linear chain-of-thought reasoning. In this work, we propose CaRGo-T (Causal Reasoning Graph-of-Thought), a reasoning framework that represents the causal and contextual relationships underlying multimodal humor as a lightweight graph-based reasoning structure. The graph is serialized into a code-based representation generated by a VLM, which can subsequently be interpreted by the same or a different VLM to produce the final prediction in zero-shot or in-context learning settings. We evaluate CaRGo-T on humor understanding and humor detection across four datasets spanning diverse forms of comedic content, including satire, sarcasm, and memes. Experiments with state-of-the-art commercial and open-source VLMs show that CaRGo-T consistently improves performance over existing reasoning-based baselines, achieving gains of approximately 1-20% on humor understanding and 1-3% on humor detection. Further analysis using mutual information indicates that the reasoning representations produced by CaRGo-T contain more information relevant to the target output than those generated by baseline reasoning approaches. Code is available at https://github.com/abhi1nandy2/CaRGo-T.
Chinese Translation
大规模视觉-语言模型(VLMs)在广泛的多模态任务中展示了显著的多样性。然而,理解幽默仍然具有挑战性,因为幽默内容通常依赖于实体、事件、上下文以及图像和文本模态之间隐含关系的微妙互动。这些互动可能涉及复杂的推理链,难以通过传统的提示或线性思维推理捕捉。在本研究中,我们提出了CaRGo-T(因果推理图思维),这是一个推理框架,将多模态幽默背后的因果和上下文关系表示为轻量级的图形推理结构。该图被序列化为由VLM生成的基于代码的表示,随后可以由相同或不同的VLM进行解释,以在零-shot或上下文学习环境中生成最终预测。我们在四个涵盖多种幽默内容形式的数据集上评估了CaRGo-T,包括讽刺、挖苦和表情包。与最先进的商业和开源VLM进行的实验表明,CaRGo-T在现有基于推理的基准上始终提高了性能,在幽默理解上取得了约1-20%的提升,在幽默检测上取得了1-3%的提升。进一步的互信息分析表明,CaRGo-T生成的推理表示包含了比基准推理方法生成的更多与目标输出相关的信息。代码可在https://github.com/abhi1nandy2/CaRGo-T获取。
cs.CL / 119 / 2608.23200

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

LongWoF-Bench:评估用于可验证长流程任务的EvoMap基因
Zhang, Xiao, Sun, Qumeng, Li, Jihao, Ren, Yiming, Liu, Xiang, Zhang, Haoyang, Wang, Junjie
Abstract
Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end verification. Yet successful execution experience is typically lost after a single run, forcing subsequent models to rediscover strategies and failure modes from scratch. We study whether such experience can instead be externalized and reused through EvoMap, where verifier-confirmed execution trajectories are consolidated into structured Gene. To evaluate this setting, we introduce the Long-Workflow Benchmark (LongWoF-Bench), comprising 778 machine-verifiable tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following. On the 252 tasks with verifier-confirmed Opus trajectories, evolved EvoMap Gene outperform Skill across all seven evaluated models by 8.7-15.5 percentage points, with the gains extending to consumer models from different model families. In contrast, reference-distilled Gene do not exhibit the same advantage, indicating that compact representation alone is insufficient and that Gene utility is closely associated with verified experience provenance. For Claude Opus, Gene reuse also completes 39 more tasks than Skill while reducing solve-time token consumption by 9.9%. Together, these results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.
Chinese Translation
大型语言模型越来越被期望执行复杂的工作流程,其成功依赖于维护相互依赖的约束条件,并生成满足严格端到端验证的工件。然而,成功执行的经验通常在单次运行后丧失,迫使后续模型从头开始重新发现策略和失败模式。我们研究了这种经验是否可以通过EvoMap外部化并重用,其中经过验证者确认的执行轨迹被整合为结构化的基因(Gene)。为了评估这一设置,我们引入了长流程基准(LongWoF-Bench),包括778个可机器验证的任务,涵盖代码生成、智能体-环境合成、数学推理和规则遵循。在252个具有验证者确认的Opus轨迹的任务中,进化的EvoMap基因在所有七个评估模型中比技能(Skill)提高了8.7-15.5个百分点,且增益扩展至来自不同模型家族的消费模型。相比之下,参考蒸馏的基因未表现出相同的优势,表明仅仅紧凑的表示不足以支持基因的效用,而基因的实用性与验证经验的来源密切相关。对于Claude Opus,基因重用还比技能多完成了39个任务,同时将解决时间的标记消耗减少了9.9%。综合来看,这些结果表明,经过验证的执行经验可以作为可重用的外部资源被保留和共享,使模型能够在不重复支付全部经验发现成本的情况下改善长流程的完成。
cs.CL / 120 / 2608.23214

Aligning Biomedical Texts and Knowledge Graphs: A Systematic Comparison of Lightweight Alignment Strategies

生物医学文本与知识图谱的对齐:轻量级对齐策略的系统比较
Bisliouk, Artem, Nosova, Elizaveta, Paulheim, Heiko, Iana, Andreea, Sousa, Rita T.
Abstract
Biomedical knowledge exists in two complementary but distinct forms: unstructured scientific literature and structured knowledge graphs (KGs). Aligning them is essential for knowledge grounding, evidence retrieval, and KG completion, yet existing methods do not explicitly align free-text evidence with KG triples. We present a unified framework for systematically studying design choices for aligning biomedical text and KGs. With a text encoder and a KG embedding model both frozen, we learn only a lightweight projection between their spaces via a contrastive objective. This enables a fair comparison across six design dimensions: text encoder, KG embedding model, projection head, triple composition, training direction, and hard-negatives sampling. We construct CTD-Align, a corpus of over 22K one-to-one tripledocument pairs linking chemical-gene interactions from the Comparative Toxicogenomics Database to supporting PubMed passages. We evaluate alignment on it in two retrieval settings: document-to-triple and triple-to-document. We find that the triple composition and the training direction (i.e., shared retrieval space) have the greatest impact, whereas the text encoder and hard-negatives sampling matter little. Overall, simple choices win: projecting text into the KG space with a linear head over concatenated subject, predicate, and object embeddings performs best. These findings establish lightweight contrastive alignment as an effective, practical foundation for bridging biomedical text and KGs.
Chinese Translation
生物医学知识以两种互补但不同的形式存在:非结构化的科学文献和结构化的知识图谱(KGs)。将它们对齐对于知识基础、证据检索和KG补全至关重要,但现有方法并未明确将自由文本证据与KG三元组对齐。我们提出了一个统一框架,以系统地研究生物医学文本与KG对齐的设计选择。在文本编码器和KG嵌入模型均被冻结的情况下,我们仅通过对比目标学习它们空间之间的轻量级投影。这使我们能够在六个设计维度上进行公平比较:文本编码器、KG嵌入模型、投影头、三元组组合、训练方向和困难负样本采样。我们构建了CTD-Align,一个包含超过22,000对一对一三元组文档对的语料库,将比较毒理基因组数据库中的化学-基因相互作用与支持的PubMed段落链接。我们在两种检索设置下对其进行对齐评估:文档到三元组和三元组到文档。我们发现三元组组合和训练方向(即共享检索空间)对结果影响最大,而文本编码器和困难负样本采样的影响较小。总体而言,简单的选择效果最佳:将文本通过线性头投影到KG空间,并对连接的主题、谓词和对象嵌入进行处理表现最佳。这些发现确立了轻量级对比对齐作为连接生物医学文本与KG的有效、实用基础。
cs.CL / 121 / 2608.23235

A Multi-Domain and Multi-Task Generative Framework with Explicit Task and Domain Conditioning for Cross-Domain Event Extraction

具有显式任务和领域条件的多领域多任务生成框架用于跨领域事件提取
Liang, Siting, Adjali, Omar, Sonntag, Daniel
Abstract
Event extraction aims to identify event triggers, classify event types, and extract arguments to construct structured event representations. Despite strong in-domain performance, developing models that generalize robustly across domains remains challenging due to variations in contextual expressions and event schemas. Prior unified and multi-task approaches improve in-domain accuracy but exhibit limited flexibility when applied to unseen domains. Even large language model-based methods that provide full event ontologies at inference time often underperform compared to smaller, task-specific fine-tuned models. We propose a unified multi-domain and multi-task training framework that models heterogeneous event schemas within a single model. Our approach introduces domain conditioning signals, jointly with task-specific prompts, enabling dynamic adaptation to dataset-specific schemas without requiring complete event label sets at inference time. The framework supports both pipeline and end-to-end extraction settings, facilitating efficient task- and domain-level transfer. Experiments on diverse event extraction benchmarks demonstrate that our method achieves competitive performance, strong cross-domain generalization, and practical scalability, while preserving domain-specific precision.
Chinese Translation
事件提取旨在识别事件触发器、分类事件类型并提取参数以构建结构化事件表示。尽管在领域内表现强劲,但由于上下文表达和事件模式的变化,开发能够在不同领域中稳健泛化的模型仍然具有挑战性。先前的统一和多任务方法提高了领域内的准确性,但在应用于未见领域时表现出有限的灵活性。即使是基于大型语言模型的方法,在推理时提供完整的事件本体,通常也不如较小的、特定任务的微调模型表现良好。我们提出了一种统一的多领域多任务训练框架,该框架在单一模型中建模异构事件模式。我们的方法引入了领域条件信号,并结合任务特定的提示,使得在推理时能够动态适应数据集特定的模式,而无需完整的事件标签集。该框架支持管道和端到端的提取设置,促进了任务和领域级的高效迁移。在多样化的事件提取基准上的实验表明,我们的方法实现了具有竞争力的性能、强大的跨领域泛化能力和实用的可扩展性,同时保持了领域特定的精确度。
cs.CL / 122 / 2608.23244

Credal Large Language Models for Semantic Commitment under Uncertainty

不确定性下的语义承诺的Credal大型语言模型
Manchingal, Shireen Kudukkil, Nikolenko, Sofiia, Cuzzolin, Fabio
Abstract
Large language models (LLMs) often produce fluent but incorrect answers with unwarranted confidence. A central limitation is that standard LLMs represent uncertainty through a single predictive distribution, conflating epistemic ignorance with genuine ambiguity. We introduce Credal Large Language Models (CLLMs): an ensemble of LoRA adapters induces a credal set whose lower and upper probabilities expose the spread of plausible predictive distributions rather than collapsing to a single softmax output. From this representation we derive two complementary commitment scores. Credal Token Commitment (CTC) is a token-space score that combines lower-bound support, credal width, and intersection entropy, computed without additional generation. Semantic Commitment Consistency (SCC) extends commitment to semantic space using sampled completions, with SCC-Gap measuring the mismatch between token-level and semantic-level support. We evaluate hallucination detection, calibration, selective prediction, and reasoning on Gemma-2-9B, Llama-3.1-8B, and Qwen2.5-7B across OpenBookQA, CoQA, TriviaQA, and ARC-Challenge. CLLM is the best method on QA accuracy at competitive expected calibration error, and CTC tracks the best hallucination AUROC within 1.5 pp on most settings without additional generation. On selective prediction at 80% coverage, CLLM with SCC reaches 99.0% accuracy on OpenBookQA, and on ARC-Challenge CLLM with Csem confidence achieves <= 0.6% ECE across the three backbones.
Chinese Translation
大型语言模型(LLMs)通常会产生流畅但不正确的答案,并伴有不合理的自信。一个主要的限制是,标准LLMs通过单一的预测分布来表示不确定性,将认知无知与真正的模糊性混淆在一起。我们提出了Credal大型语言模型(CLLMs):一个LoRA适配器的集成产生了一个可信集,其下限和上限概率揭示了合理预测分布的分布,而不是简化为单一的softmax输出。基于这种表示,我们推导出两个互补的承诺评分。Credal Token Commitment(CTC)是一个令牌空间评分,它结合了下限支持、可信宽度和交集熵,计算时无需额外生成。语义承诺一致性(SCC)通过使用采样完成将承诺扩展到语义空间,SCC-Gap测量令牌级和语义级支持之间的不匹配。我们在Gemma-2-9B、Llama-3.1-8B和Qwen2.5-7B上评估了幻觉检测、校准、选择性预测和推理,数据集包括OpenBookQA、CoQA、TriviaQA和ARC-Challenge。CLLM在QA准确性方面是最佳方法,具有竞争性的期望校准误差,CTC在大多数设置下跟踪最佳幻觉AUROC,误差在1.5个百分点以内,无需额外生成。在80%覆盖率的选择性预测中,带有SCC的CLLM在OpenBookQA上达到了99.0%的准确率,而在ARC-Challenge上,带有Csem信心的CLLM在三个基础模型上实现了<= 0.6%的ECE。
cs.CL / 123 / 2608.23248

Future Querying: Can LLMs Serve as Implicit Medical World Models?

未来查询:大型语言模型能否作为隐式医学世界模型?
Willems, Siri, Butterworth, James, Goetschalckx, Lore, Vrancx, Peter, Modard, Philippe, Giets, Elke, Denoyer, Ludovic
Abstract
Traditional clinical prediction models rely on task-specific pipelines and curated, structured data, which scale poorly and underutilize unstructured text. To address this, we introduce future querying, a paradigm that probes whether large language models (LLMs) can function as implicit medical world models by evaluating their ability to answer time-indexed clinical queries about a patient's future. Our framework operates on unstructured clinical documentation using endpoint-agnostic training, enabling a single model to answer diverse clinical queries over patient trajectories without manual feature engineering or task-specific retraining. We show that small, locally fine-tuned open-weight models can match or approach larger proprietary systems, making the framework suitable for privacy-preserving, on-premise deployment. Evaluated on a new synthetic medical reports dataset and real ICU notes from the MIMIC-IV dataset, our results provide encouraging evidence that LLMs can capture aspects of clinical dynamics.
Chinese Translation
传统的临床预测模型依赖于特定任务的流程和经过整理的结构化数据,这种方法扩展性差且未充分利用非结构化文本。为了解决这一问题,我们提出了未来查询(future querying)这一范式,探讨大型语言模型(LLMs)是否能够作为隐式医学世界模型,通过评估其回答关于患者未来的时间索引临床查询的能力。我们的框架在非结构化临床文档上运行,采用与端点无关的训练,使得单一模型能够在患者轨迹上回答多样的临床查询,而无需手动特征工程或特定任务的再训练。我们展示了小型、局部微调的开放权重模型能够匹配或接近更大规模的专有系统,使得该框架适合于隐私保护的本地部署。在一个新的合成医学报告数据集和来自MIMIC-IV数据集的真实ICU记录上进行评估,我们的结果提供了令人鼓舞的证据,表明LLMs能够捕捉临床动态的某些方面。
cs.CL / 124 / 2608.23261

A Scalable Cross-Domain Event Extraction System via a Unified Generative Training Framework

通过统一生成训练框架实现可扩展的跨领域事件提取系统
Liang, Siting, Adjali, Omar, Bhatti, Omair Shahzad, Sonntag, Daniel
Abstract
Event extraction is fundamental to information extraction. Prior approaches often separate event detection and argument extraction or depend on dataset-specific designs, limiting scalability and cross-domain generalization. We propose a unified generative sequence-to-sequence framework that performs event extraction subtasks jointly and supports both pipeline and end-to-end configurations. We fine-tune pretrained language models on multiple event datasets across diverse domains, enabling a single model to retain domain-specific semantics while generalizing over large and evolving label spaces. We demonstrate these capabilities through a web-based application tailored for researchers and practitioners. The platform supports document upload, schema-aware event extraction, visualization of triggers and arguments, and comparison of different extraction configurations across domains.
Chinese Translation
事件提取是信息提取的基础。以往的方法通常将事件检测和论元提取分开,或依赖于特定数据集的设计,这限制了可扩展性和跨领域的泛化能力。我们提出了一种统一的生成序列到序列框架,该框架共同执行事件提取的子任务,并支持管道和端到端配置。我们在多个跨不同领域的事件数据集上微调预训练语言模型,使得单一模型能够保留领域特定的语义,同时在大规模和不断演变的标签空间中实现泛化。我们通过一个针对研究人员和从业者量身定制的基于网络的应用程序展示了这些能力。该平台支持文档上传、模式感知的事件提取、触发器和论元的可视化,以及跨领域不同提取配置的比较。
cs.CL / 125 / 2608.23265

EvoWiki: Incremental State Overwriting and Traceable Question Answering for Cross-Meeting Knowledge Evolution

EvoWiki:跨会议知识演变的增量状态覆盖与可追溯问答
Chen, Dongsheng, Wang, Tianyu, Que, Wenhui
Abstract
In long-term collaboration spanning multiple meetings, factual states such as decisions and risks are continually revised, overturned, and replaced. Existing long-context methods typically stack the entire history, while many RAG and structured-memory methods organize knowledge as static or append-only facts and rely on semantic relevance at read time. Without explicit modeling of knowledge lifecycles, these approaches may retain conflicting old and new states simultaneously or discard history, leading to stale retrieval and answers that are difficult to verify. We present EvoWiki, an incremental question-answering architecture for dynamic long-form text. EvoWiki decouples offline incremental construction (BUILD) from online structured reading (READ). BUILD captures the intra-meeting micro-evolution from proposal to decision and uses entity version chains and a fine-grained State-Overwrite Protocol to explicitly distinguish current valid states from superseded history while preserving meeting-level provenance anchors. READ bypasses relevance-based Top-k retrieval and performs deterministic entity addressing, temporal resolution, and cross-entity multi-hop aggregation over the complete Wiki to produce grounded and traceable answers. We further introduce CrossMeet, a high-fidelity bilingual benchmark designed to simulate long-term state evolution, covering factual consistency, temporal reasoning, and cross-meeting multi-hop reasoning. Across six datasets and two reader models, EvoWiki improves macro-average Judge Accuracy over the strongest baselines by 9.72 and 10.00 percentage points, respectively. Human evaluation shows that EvoWiki is more robust and factually faithful under frequent state flips, validating valid-state-oriented reading as a reliable approach to cross-meeting knowledge evolution.
Chinese Translation
在跨多个会议的长期协作中,事实状态如决策和风险不断被修订、推翻和替换。现有的长上下文方法通常堆叠整个历史,而许多 RAG 和结构化记忆方法将知识组织为静态或仅附加的事实,并依赖于读取时的语义相关性。在没有明确建模知识生命周期的情况下,这些方法可能同时保留冲突的旧状态和新状态,或丢弃历史,导致检索结果过时且难以验证。我们提出了 EvoWiki,一种用于动态长文本的增量问答架构。EvoWiki 将离线增量构建(BUILD)与在线结构化读取(READ)解耦。BUILD 捕捉会议内的微观演变,从提案到决策,并使用实体版本链和细粒度状态覆盖协议明确区分当前有效状态与被取代的历史,同时保留会议级的来源锚点。READ 绕过基于相关性的 Top-k 检索,执行确定性的实体寻址、时间解析和跨实体的多跳聚合,以生成有据可查且可追溯的答案。我们进一步引入 CrossMeet,这是一个高保真双语基准,旨在模拟长期状态演变,涵盖事实一致性、时间推理和跨会议多跳推理。在六个数据集和两个阅读模型中,EvoWiki 在最强基线之上分别提高了宏平均评判准确率 9.72 和 10.00 个百分点。人类评估表明,EvoWiki 在频繁状态翻转下更具鲁棒性和事实忠实性,验证了以有效状态为导向的阅读作为跨会议知识演变的可靠方法。
cs.CL / 126 / 2608.23284

Dynamic Topic Modeling for Cross-Corpus Temporal Analysis

跨语料库时间分析的动态主题建模
Li, Ruoxuan, Kogut, Bruce
Abstract
Dynamic Embedded Topic Models (D-ETM) provide an interpretable framework for modeling temporal semantic evolution, but cross-corpus comparison remains difficult because topics are often learned independently and aligned only after training, a process that does not guarantee stable topic correspondence across corpora and time. To address this problem, we propose a D-ETM framework that first learns a common dynamic topic space over a merged multi-corpus collection, which we call the shared backbone, then introduces corpus-specific residual adaptation around the frozen backbone without creating separate latent topic spaces. This design preserves a shared topic index for cross-corpus comparison while allowing each corpus to specialize lexically. We evaluate the framework on three temporally structured corpora spanning 97 years: the Corpus of Historical American English, Harvard Business Review, and International Labour Review. Residual adaptation improves corpus-specific fit relative to the shared backbone while preserving the same-index cross-corpus topic trajectories, achieving substantially stronger alignment than full fine-tuning from the same backbone, with $97.5 \pm 0.7\%$ versus $17.9 \pm 1.1\%$ trajectory Retrieval@1, as well as stronger alignment than independent training with post-hoc Hungarian matching. These results suggest that incorporating topic alignment into the model can support more stable over-time cross-corpus comparisons while retaining corpus-specific lexical variation.
Chinese Translation
动态嵌入主题模型(Dynamic Embedded Topic Models, D-ETM)提供了一种可解释的框架,用于建模时间语义演变,但由于主题通常是独立学习的,并且仅在训练后进行对齐,这使得跨语料库比较变得困难,这一过程并不能保证跨语料库和时间的主题对应关系稳定。为了解决这个问题,我们提出了一种D-ETM框架,该框架首先在合并的多语料库集合上学习一个共同的动态主题空间,我们称之为共享主干(shared backbone),然后在冻结的主干周围引入语料库特定的残差适应,而不创建单独的潜在主题空间。这一设计保留了跨语料库比较的共享主题索引,同时允许每个语料库在词汇上进行专业化。我们在三个跨越97年的时间结构化语料库上评估该框架:历史美国英语语料库(Corpus of Historical American English)、哈佛商业评论(Harvard Business Review)和国际劳动评论(International Labour Review)。残差适应相对于共享主干改善了语料库特定的拟合,同时保持了相同索引的跨语料库主题轨迹,取得了显著强于从同一主干进行完全微调的对齐效果,轨迹检索率(Retrieval@1)为$97.5 imes 0.7 ext{%}$,而完全微调为$17.9 imes 1.1 ext{%}$,并且相较于独立训练与后期匈牙利匹配的对齐效果也更强。这些结果表明,将主题对齐纳入模型可以支持更稳定的跨语料库时间比较,同时保留语料库特定的词汇变异。
cs.CL / 127 / 2608.23311

Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

超越稳定性-探索困境:用于大语言模型政策优化的环境正则化
Zhou, Xianlei, Meng, Xiangdi, He, Yu, Qi, Tianyu, Guan, Shuyan, Zhang, Xianli, Zhang, Jian, Li, Xin, Lin, Qika, Liu, Jun
Abstract
Policy optimization (PO) for Large Language Models faces a stability--exploration trade-off, currently mediated by an action-side Policy-KL regularizer. This puts practitioners in a double bind: keeping Policy-KL constrains response behavior and consumes the action-side exploration budget, while dropping it leaves the optimization without an explicit drift control. We argue for an alternative that breaks the dilemma by moving regularization to the input side. As training progresses, the distribution over training queries induced by the current policy drifts unchecked from its pre-RL reference distribution. Concretely, Environment-Regularized Policy Optimization (ERPO) introduces a Query-KL (QKL) term that bounds this query distribution shift, together with a dataset-static reference-derived per-query weight that biases each per-query update toward queries typical under the reference. The QKL gradient flows strictly through the query likelihood; the response score function used by policy-gradient estimators does not appear in the QKL term, so QKL exerts no direct gradient pressure on the response distribution---exploration is preserved. ERPO plugs into GRPO/PPO/REINFORCE-style pipelines without additional forward passes. On six mathematical reasoning benchmarks, ERPO replaces the standard Policy-KL regularizer while achieving effective control over query distribution drift, delivering stronger accuracy and substantially more stable behavior under high-temperature decoding and long-horizon training.Our source code are available at https://github.com/alibaba/ERPO
Chinese Translation
大语言模型的政策优化(PO)面临稳定性与探索之间的权衡,目前通过动作侧的政策-KL正则化器进行调节。这使得从业者陷入了两难境地:保持政策-KL会限制响应行为并消耗动作侧的探索预算,而放弃它则使优化缺乏明确的漂移控制。我们主张一种替代方案,通过将正则化转移到输入侧来打破这一困境。随着训练的进行,当前政策引发的训练查询分布会不受控制地漂移离其预强化学习(pre-RL)参考分布。具体而言,环境正则化政策优化(ERPO)引入了一个查询-KL(QKL)项,以限制这一查询分布的漂移,同时结合一个数据集静态参考派生的每查询权重,使每次查询更新朝向参考下典型的查询偏移。QKL梯度严格通过查询似然流动;政策梯度估计器使用的响应评分函数并未出现在QKL项中,因此QKL对响应分布没有直接的梯度压力——探索得以保留。ERPO可以无额外前向传递地嵌入到GRPO/PPO/REINFORCE风格的管道中。在六个数学推理基准测试中,ERPO替代了标准的政策-KL正则化器,同时有效控制查询分布的漂移,提供了更强的准确性,并在高温解码和长时间训练下表现出显著更稳定的行为。我们的源代码可在 https://github.com/alibaba/ERPO 获取。
cs.CL / 128 / 2608.23327

Flesch-Kincaid Readability Depends Only on the Topic Distribution in Long Texts under Topic Models

Flesch-Kincaid 可读性仅依赖于主题模型下长文本中的主题分布
Ehara, Yo
Abstract
Flesch Reading Ease (FRE) and the Flesch-Kincaid Grade Level (FKGL) are widely used readability scores for English computed from the same two document statistics, yet their stability on long documents need not imply invariance to lexical composition. Surprisingly, under a topic model with an explicit sentence-boundary token, both scores converge almost surely to deterministic functions of the document topic distribution through just two scalar rates: in the long-text limit, all score variation is mediated by topical composition rather than any residual readability signal. The theory covers both formulae, while the experiments evaluate FKGL. In a fixed admixture with rank[1, q, s] = 3, fibres through interior topic vectors are locally (K-3)-dimensional, whereas regular iso-score level sets are locally (K-2)-dimensional and curved. In out-of-fold evaluation on two balanced corpora, Brown and the written BNC, a topic vector inferred from one document half's content words predicts the other half's FKGL at r = 0.779 and 0.884, respectively. On Brown, adding the topic prediction to genre and mean content-word syllable count yields $\Delta R^2$ = 0.002, with a confidence interval spanning zero; on the BNC, the corresponding split-half increment is 0.024, positive in four of five K = 100 fits (median 0.021). Because inferred topics may also absorb genre, register, and style, we do not interpret these results as evidence about human readability or causal effects.
Chinese Translation
Flesch 阅读容易度 (FRE) 和 Flesch-Kincaid 年级水平 (FKGL) 是广泛使用的英语可读性评分,基于相同的两个文档统计数据计算而得,然而它们在长文档上的稳定性并不意味着对词汇组成的不变性。令人惊讶的是,在一个具有明确句子边界标记的主题模型下,这两个评分几乎肯定会收敛到文档主题分布的确定性函数,仅通过两个标量速率:在长文本极限中,所有评分变异都由主题组成介导,而不是任何残余的可读性信号。该理论涵盖了这两个公式,而实验则评估了 FKGL。在固定的混合中,秩 [1, q, s] = 3,内部主题向量的纤维局部上是 (K-3) 维的,而常规等评分水平集局部上是 (K-2) 维的且呈曲面。在对两个平衡语料库的折叠外评估中,Brown 和书面 BNC,从一个文档一半的内容词推断出的主题向量分别预测了另一半的 FKGL,相关系数 r = 0.779 和 0.884。在 Brown 中,将主题预测添加到体裁和平均内容词音节计数中,$ ext{Δ} R^2$ = 0.002,置信区间跨越零;在 BNC 中,相应的分半增量为 0.024,在五个 K = 100 拟合中有四个为正(中位数 0.021)。由于推断出的主题也可能吸收体裁、语域和风格,我们不将这些结果解释为关于人类可读性或因果效应的证据。
cs.CL / 129 / 2608.23338

The Emergence of Relevance Through Axiomatic Attention Patterns During LoRA Fine-Tuning

通过公理注意模式在LoRA微调过程中相关性的出现
Perlman, Matthew, Nijasure, Atharva, Allan, James
Abstract
LoRA fine-tuning is standard for adapting LLMs to reranking, but it remains unclear where in the network task-specific relevance behavior is learned and what attention-level changes accompany that learning. Through ablation and attention experiments, we identify where LoRA attention updates to RankLLaMA improve performance and whether those gains coincide with interpretable relevance-oriented attention patterns such as lexical matching, rarity sensitivity, and query-document interaction. We find that given LoRA fine-tuned MLPs throughout the network, restricting LoRA attention updates to a compact mid-network region is sufficient for recovering over half of the performance gained by applying LoRA to all attention layers, and that omitting attention fine-tuning in this region hurts performance more than elsewhere in the network. Additionally, we show that regions where applying LoRA affects performance the most overlap with regions where fine-tuning increased attention to axiomatic IR features. Rarity sensitivity, document-query interaction, and several compositional features are highly correlated with gains in ranking performance. Our results support an interpretable, correlational account of how relevance-oriented behavior emerges during LoRA fine-tuning and point toward improved strategies for adapting rerankers.
Chinese Translation
LoRA微调是将大型语言模型(LLMs)适应于重排序的标准方法,但在网络中具体在哪个位置学习任务特定的相关性行为以及伴随这种学习的注意力层级变化仍不清楚。通过消融实验和注意力实验,我们确定了LoRA注意力更新在RankLLaMA中提高性能的位置,以及这些提升是否与可解释的相关性导向注意模式(如词汇匹配、稀缺性敏感性和查询-文档交互)相一致。我们发现,在整个网络中给定LoRA微调的多层感知器(MLPs)时,将LoRA注意力更新限制在一个紧凑的中间网络区域足以恢复应用LoRA于所有注意力层所获得的超过一半的性能提升,并且在该区域省略注意力微调对性能的影响大于网络其他地方的影响。此外,我们还展示了应用LoRA对性能影响最大的区域与微调增加对公理信息检索(IR)特征的注意力的区域重叠。稀缺性敏感性、文档-查询交互和若干组合特征与排名性能的提升高度相关。我们的结果支持了一种可解释的、相关性的观点,说明在LoRA微调过程中相关性导向行为是如何出现的,并指向改进重排序器适应策略的方向。
cs.CL / 130 / 2608.23353

FormuEvo: LLM-Guided Evolution for Discovering Solver-Efficient Mixed-Integer Programming Formulations

FormuEvo:基于大型语言模型的进化方法用于发现求解器高效的混合整数规划模型
Yuan, Haofeng, Peng, Jianing, Bi, Jieyi, Zhang, Ni, Song, Shiji, Cao, Zhiguang
Abstract
Mixed-integer programming (MIP) lies at the core of operations research and industrial optimization. While large language models (LLMs) have recently shown promise in automated MIP modeling from natural language, they prioritize semantic correctness but overlook formulation strength, severely bottlenecking the efficiency of downstream solvers. We propose FormuEvo, an LLM-guided evolutionary framework for automated discovery of solver-efficient MIP formulations. FormuEvo frames MIP formulation design as evolutionary optimization over the symbolic space of MIP formulations, represented as executable modeling programs, by iteratively generating, evaluating, and selecting stronger candidates via LLM-driven crossover, mutation, and repair operations. To move beyond blind exploration, FormuEvo introduces a solver-informed diagnosis mechanism that exploits fine-grained solver statistics as verbal gradients for targeted refinement. Additionally, a structured memory abstracts prior experience into reusable modeling strategies, avoiding redundant exploration while enabling zero-shot transfer to unseen problems and bootstrapping smaller LLMs. Experiments across diverse linear and non-linear problems demonstrate that FormuEvo discovers formulations that significantly outperform both expert-designed formulations and existing LLM-based approaches, accelerating solvers by up to 5.5$\times$, with distilled knowledge transferring effectively across problems and model scales.
Chinese Translation
混合整数规划(MIP)是运筹学和工业优化的核心。尽管大型语言模型(LLMs)最近在从自然语言自动建模MIP方面显示出潜力,但它们优先考虑语义正确性,而忽视了模型的强度,严重制约了下游求解器的效率。我们提出了FormuEvo,一种基于LLM的进化框架,用于自动发现求解器高效的MIP模型。FormuEvo将MIP模型设计框架视为在MIP模型的符号空间上进行进化优化,通过迭代生成、评估和选择更强的候选模型,利用LLM驱动的交叉、变异和修复操作。为了超越盲目探索,FormuEvo引入了一种求解器信息诊断机制,利用细粒度的求解器统计数据作为语言梯度进行有针对性的优化。此外,结构化记忆将先前的经验抽象为可重用的建模策略,避免了冗余探索,同时实现了对未见问题的零样本迁移,并为较小的LLM提供了引导。针对多种线性和非线性问题的实验表明,FormuEvo发现的模型显著优于专家设计的模型和现有的基于LLM的方法,求解器加速最高可达5.5倍,提炼的知识在不同问题和模型规模间有效迁移。
cs.CL / 131 / 2608.23358

The Geometry of Low-Resource Language Representations

低资源语言表示的几何特征
Meyer, Francois, Buys, Jan
Abstract
The performance gap between low- and high-resource languages in LLMs is widely known, but it remains unclear which internal model factors drive these disparities. In this paper, we characterise this gap through the lens of representational geometry. Comparing the geometric properties of hidden representations across 30 languages reveals that LLM geometry is systematically related to language data availability. The most consistent effect is in final layers, where low-resource languages exhibit representational degeneration. To counter this, we investigate the effectiveness of regularisation terms to penalise degeneration during continued pretraining (CPT). Experiments monolingually adapting 9 base LLMs to 10 African languages show that geometric regularisation successfully reduces representational degeneration during CPT. For larger models, cosine similarity-based regularisation marginally improves performance over vanilla CPT, with more consistent gains on the most challenging tasks. We establish that the representational geometry of low- and high-resource languages in LLMs is measurably distinct, and that targeted geometric intervention is a viable strategy for improving CPT for low-resource languages.
Chinese Translation
在大型语言模型(LLMs)中,低资源语言与高资源语言之间的性能差距广为人知,但尚不清楚哪些内部模型因素导致了这些差异。本文通过表示几何的视角来表征这一差距。对30种语言的隐藏表示几何特性进行比较,揭示了LLM几何与语言数据可用性之间的系统性关系。最一致的影响出现在最终层,其中低资源语言表现出表示退化。为了解决这一问题,我们研究了在持续预训练(CPT)期间惩罚退化的正则化项的有效性。单语适应9个基础LLM到10种非洲语言的实验表明,几何正则化在CPT期间成功减少了表示退化。对于更大的模型,基于余弦相似度的正则化在普通CPT上略微提高了性能,并在最具挑战性的任务上获得了更一致的提升。我们确定了LLM中低资源语言和高资源语言的表示几何是可测量的不同,并且针对性的几何干预是改善低资源语言CPT的可行策略。
cs.CL / 132 / 2608.23390

Cross-lingual Biography Enrichment via Claim Extraction and Alignment

通过声明提取和对齐实现跨语言传记丰富化
Song, Yifei, Chen, Ziyang, Sayilov, Emil, Gardent, Claire
Abstract
English Wikipedia is often treated as the default encyclopedic source, yet non-English Wikipedia editions can contain richer locally grounded information for long-tail figures. We study cross-lingual biography enrichment: enriching an existing English biography with facts supported by a non-English biography about the same person. Focusing on women from non-English-speaking contexts, we introduce \textsc{CLAW-4L}, a benchmark consisting of 300 Wikipedia biography pairs linking an English biography with its French, Chinese or Azerbaijani counterpart, along with claim annotations and a fine-grained claim-pair relation corpus. We propose a claim-based enrichment framework that extracts English claims from both biographies, aligns them to identify enrichment evidence from the non-English biography, and rewrites the English biography using the selected claims. Our results show that non-English Wikipedia biographies provide valuable evidence for improving English biography coverage, while lower-resource settings remain challenging.
Chinese Translation
英语维基百科通常被视为默认的百科全书来源,但非英语维基百科版本可能包含更丰富的本地信息,尤其是针对长尾人物。我们研究跨语言传记丰富化:通过非英语传记中关于同一人物的事实来丰富现有的英语传记。我们关注来自非英语国家的女性,介绍了 extsc{CLAW-4L},这是一个基准数据集,包含300对维基百科传记,将英语传记与其法语、中文或阿塞拜疆语对应传记链接,并附有声明注释和细粒度的声明对关系语料库。我们提出了一种基于声明的丰富化框架,该框架从两个传记中提取英语声明,进行对齐以识别来自非英语传记的丰富化证据,并使用所选声明重写英语传记。我们的结果表明,非英语维基百科传记为改善英语传记的覆盖范围提供了宝贵的证据,但在资源较少的环境中仍然面临挑战。
cs.CL / 133 / 2608.23391

Cross-Domain, Multi-Task Data-to-Text Generation without In-Domain Training Data

跨领域、多任务的数据到文本生成,无需领域内训练数据
Song, Yifei, Efimov-Zhang, Kun, Gardent, Claire
Abstract
Structured data exists in many forms (tables, knowledge graphs, charts, and time series), and converting it into text may involve different generation tasks. However, most prior work on data-to-text (D2T) generation has focused on specific tasks and datasets, relying either on task-specific training data or on the zero-shot capabilities of large language models. We study cross-domain D2T generation in a setting where neither in-domain training text nor test references are available, and where domains, generation goals, and input structures vary substantially. We compare data-driven knowledge distillation (DDKD) against zero-shot inference and fine-tuning on out-of-domain D2T data, and introduce structure-preserving augmentation via structural subsampling and perturbation. Experiments on five benchmarks show that, at constant model size (1.7B parameters), DDKD consistently outperforms both fine-tuning and zero-shot inference. Moreover, the resulting small models outperform a much larger finetuned model on two of the five domains, achieving comparable performance on the remaining three. We further construct QUINTD-5, a fivefold extension of QUINTD-1, and show that simply scaling real target-domain inputs yields only modest gains, whereas our augmentation strategy remains more effective and more cost-efficient for cross-domain distillation.
Chinese Translation
结构化数据存在多种形式(表格、知识图谱、图表和时间序列),将其转换为文本可能涉及不同的生成任务。然而,以往大多数关于数据到文本(D2T)生成的研究集中于特定任务和数据集,依赖于特定任务的训练数据或大型语言模型的零样本能力。我们研究了在没有领域内训练文本或测试参考的情况下的跨领域D2T生成,其中领域、生成目标和输入结构有显著差异。我们将数据驱动的知识蒸馏(DDKD)与零样本推理和在领域外D2T数据上的微调进行了比较,并通过结构子采样和扰动引入了结构保留增强。五个基准实验表明,在模型规模(17亿参数)不变的情况下,DDKD始终优于微调和零样本推理。此外,得到的小模型在五个领域中的两个领域上超越了一个更大的微调模型,并在其余三个领域上达到了可比的性能。我们进一步构建了QUINTD-5,这是QUINTD-1的五倍扩展,并表明仅仅扩大真实目标领域输入仅能带来适度的提升,而我们的增强策略在跨领域蒸馏中仍然更有效且更具成本效益。
cs.CL / 134 / 2608.23411

STONIC: A Layered Measurement Contract for LLM Value Profiling

STONIC:一种用于大语言模型价值分析的分层测量契约
Chetvergov, Andrei, Ukolov, Stepan, Sivoraksha, Timofei, Evseev, Alexander, Sazanakov, Danil, Solovev, Mikhail, Bolovtsov, Sergey
Abstract
LLM value studies often merge questionnaire ratings, pairwise choices, and values inferred from generated text into one profile. That merge assumes that the three observations describe the same stable preference. STONIC tests this assumption on 5,144 situations from four banks and 35 fixed model configurations. It compares responses rated in isolation, choices made under counterbalanced conflict, spontaneous answers, and later choices between a model's own answer and authored alternatives. 10 of 17 configurations with usable behavioral data preserve the endorsement-choice relation across banks. Every one of the 17 eligible configurations prefers its own earlier answer (median effect 0.790), although option position changes the choice rate in every eligible configuration. Profile shape transfers most strongly from ratings to conflict choices and weakens for spontaneous text. Three-way annotation of 200 L3 responses provides a task-local check of the semantic audit: FULCRA agrees most closely with the human majority, while DeBERTa retains useful rank information after calibration. Hidden states encode the completed decision more clearly than the prompt alone. Thus the models show reproducible behavioral continuity, but the evidence does not support one scorer-independent value identity across interfaces.
Chinese Translation
大语言模型(LLM)价值研究通常将问卷评分、成对选择和从生成文本推断的价值合并为一个档案。这种合并假设这三种观察结果描述了相同的稳定偏好。STONIC 在来自四家银行和 35 种固定模型配置的 5,144 个情境中测试了这一假设。它比较了孤立评分的响应、在平衡冲突下做出的选择、自发回答,以及后续在模型自身答案和作者替代答案之间的选择。在 17 种可用行为数据的配置中,有 10 种配置在不同银行之间保持了认可-选择关系。所有 17 种合格配置都偏好其早期的答案(中位效应 0.790),尽管选项位置在每个合格配置中改变了选择率。档案形状从评分到冲突选择的转移最强,而对自发文本的转移则减弱。对 200 个 L3 响应的三方注释提供了语义审计的任务局部检查:FULCRA 与人类多数意见最为一致,而 DeBERTa 在校准后保留了有用的排名信息。隐藏状态比单独的提示更清晰地编码了完成的决策。因此,这些模型显示出可重复的行为连续性,但证据并不支持跨接口的独立评分者价值身份。
cs.CL / 135 / 2608.23421

A Comprehensive Analysis of Arabic Natural Language Processing Research: Trends, Topic Evolution, and Research Gaps -- A Bibliometric and Topic-Based Study

阿拉伯自然语言处理研究的综合分析:趋势、主题演变与研究空白——一项文献计量与主题基础研究
Arabov, Mullosharaf K.
Abstract
Natural Language Processing (NLP) has grown rapidly over the past decade, driven by digital transformation in the Arab world, social media, and large language models (LLMs). Despite this growth, a comprehensive quantitative meta-analysis of the field remains absent. This study presents a large-scale bibliometric and topic-based analysis of 7,120 Arabic NLP papers published between 1960 and 2026, sourced from six collections. We employ BERTopic for topic modeling, regression analysis to identify citation predictors, social network analysis for co-authorship structures, and geographic mapping. Our findings show a significant publication surge after 2020, driven by transformer models and LLMs. Topic modeling identifies 19 substantive themes, the largest centered on text, speech, translation, and recognition. Citation analysis reveals a positive correlation between paper age and citations (r = 0.245, p < 0.001); regression shows that indexing in OpenAlex or Semantic Scholar and institutional affiliation are associated with higher citation counts. Saudi Arabia, the United States, and Egypt lead in research output. A task-dialect gap matrix identifies critical understudied areas, including summarization for Maghrebi, Iraqi, and Sudanese dialects. The largest topic has the highest H-index (87), followed by sentiment analysis (54). Our quantitative approach complements existing qualitative surveys and offers recommendations to prioritize under-resourced dialects and develop culturally aligned benchmarks for Arabic NLP.
Chinese Translation
自然语言处理(NLP)在过去十年中迅速发展,受到阿拉伯世界数字化转型、社交媒体和大型语言模型(LLMs)的推动。尽管如此,关于该领域的全面定量元分析仍然缺乏。本研究对1960年至2026年间发表的7120篇阿拉伯NLP论文进行了大规模的文献计量和主题基础分析,这些论文来自六个文献集合。我们采用BERTopic进行主题建模,使用回归分析识别引用预测因素,进行社交网络分析以研究合作结构,并进行地理映射。我们的研究结果显示,2020年后出版数量显著激增,这一增长主要受到变换器模型和LLMs的推动。主题建模识别出19个实质性主题,其中最大主题集中于文本、语音、翻译和识别。引用分析揭示了论文年龄与引用次数之间的正相关关系(r = 0.245,p < 0.001);回归分析表明,在OpenAlex或Semantic Scholar上被索引以及机构隶属关系与更高的引用次数相关。沙特阿拉伯、美国和埃及在研究产出方面处于领先地位。任务-方言差距矩阵识别出关键的未充分研究领域,包括对马格里布、伊拉克和苏丹方言的摘要研究。最大主题的H指数最高(87),其次是情感分析(54)。我们的定量方法补充了现有的定性调查,并提出了优先考虑资源不足的方言和开发与阿拉伯NLP文化相一致的基准的建议。
cs.CL / 136 / 2608.23448

How Useful are LLMs for Grammar Engineering? Cantonese ParGram Resources and Controlled Experimental Evaluation with English Baselines

大型语言模型在语法工程中的实用性如何?粤语ParGram资源及与英语基线的控制实验评估
Lam, Chit-Fung
Abstract
This paper presents new Cantonese ParGram resources and evaluates LLMs for knowledge-driven grammar engineering within a controlled experimental paradigm. Using Cantonese ParGram resources as gold standards, with corresponding English baselines, we investigate whether OpenAI's gpt-oss-120b and GPT-5.4 can generate machine-processable grammars from sentences and target formal structures under systematically varied prompting conditions. GPT-5.4 outperformed gpt-oss-120b, while grammars generated from target formal structures generally outperformed those generated from sentences. Although both models could generate locally plausible phrase-structure rules, lexical entries, and templates, they often struggled to coordinate interacting formal constraints, especially in multi-construction settings. The results characterize both the capabilities and limitations of current LLMs for potential integration into AI-assisted expert workflows: LLMs may support intermediate stages of grammar development, but human linguistic expertise remains central to analysis, validation, and refinement. The study also contributes new Cantonese symbolic grammatical resources.
Chinese Translation
本文提出了新的粤语ParGram资源,并在控制实验范式下评估大型语言模型(LLMs)在知识驱动的语法工程中的应用。以粤语ParGram资源作为金标准,并与相应的英语基线进行对比,我们研究了OpenAI的gpt-oss-120b和GPT-5.4是否能够在系统变化的提示条件下,从句子和目标形式结构生成机器可处理的语法。结果表明,GPT-5.4的表现优于gpt-oss-120b,而从目标形式结构生成的语法通常优于从句子生成的语法。尽管这两种模型能够生成局部合理的短语结构规则、词汇条目和模板,但在协调相互作用的形式约束方面,尤其是在多构造设置中,它们往往面临困难。研究结果表征了当前大型语言模型在潜在集成到AI辅助专家工作流程中的能力与局限性:大型语言模型可能支持语法开发的中间阶段,但人类语言学专业知识仍然是分析、验证和完善的核心。该研究还贡献了新的粤语符号语法资源。
cs.CL / 137 / 2608.23474

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

有什么问题?评估视觉-语言模型中的时间一致性
Hradil, Marek, Villegas, Danae Sánchez
Abstract
Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal structure. To study this question, we formulate temporal grounding as an anomaly detection problem, providing a simple and controlled evaluation that directly tests sensitivity to temporal consistency. We introduce TimeCatch, where temporal anomalies are created by swapping consecutive frames and frame-level anomalies by replacing a frame with Gaussian noise. Models are evaluated on anomaly detection and localization tasks across four synthetic and real-world datasets, alongside a human study. Our evaluation reveals a substantial gap between frame-level and temporal anomaly detection. While VLMs consistently detect frame-level anomalies and often localize them accurately, they perform near chance on temporal anomaly detection and only modestly above chance on localization. Humans, in contrast, achieve near-ceiling performance on both tasks. Additional analyses across model scales, prompting strategies, sequence lengths, and visual similarity suggest that these failures cannot be explained solely by limitations in perception or model capacity. Together, these findings indicate that current VLMs can identify anomalies within individual frames but struggle to integrate information across frames to reason about temporal consistency. TimeCatch provides a controlled benchmark for evaluating temporal grounding in vision-language models.
Chinese Translation
视觉-语言模型(VLMs)在视频和图像序列基准测试中表现出色,但尚不清楚它们是否捕捉到了时间结构。为研究这一问题,我们将时间定位形式化为异常检测问题,提供了一种简单且可控的评估方法,直接测试对时间一致性的敏感性。我们引入了TimeCatch,通过交换连续帧来创建时间异常,并通过用高斯噪声替换帧来创建帧级异常。模型在四个合成和真实世界数据集上进行异常检测和定位任务的评估,并进行了一项人类研究。我们的评估揭示了帧级异常检测和时间异常检测之间存在显著差距。尽管VLMs始终能够检测帧级异常并通常能够准确定位它们,但在时间异常检测上表现接近随机水平,在定位上仅略高于随机水平。相比之下,人类在这两项任务上几乎达到了顶尖表现。对模型规模、提示策略、序列长度和视觉相似性等的进一步分析表明,这些失败不能仅仅通过感知或模型能力的局限性来解释。综合来看,这些发现表明当前的VLMs能够识别单个帧内的异常,但在整合跨帧信息以推理时间一致性方面存在困难。TimeCatch为评估视觉-语言模型中的时间定位提供了一个可控的基准。
cs.CL / 138 / 2608.23476

On the Threat Model of Weird Generalization and Emergent Misalignment

关于奇异泛化的威胁模型及其突现性失调
Wanner, Miriam, Dredze, Mark, Walden, William
Abstract
Narrow fine-tuning on small, domain-specific datasets can produce broad and surprising changes in model behavior-a phenomenon called weird generalization (WG). Yet, it remains unclear what features of the fine-tuning data are necessary for WG to arise. Here, we address this question by investigating a range of plausibly relevant features, including dataset size, composition, language, presentation style, and novelty relative to a model's parametric knowledge. Further, since WG evaluations rely on small question sets that assess the extent of the generalization, we also analyze how sensitive this measurement is to the set of questions used. Experiments with three open-weight models on four datasets show that the degree of WG (1) depends heavily on dataset composition and language (more than on size); (2) is greater for data familiar from pretraining than for novel data; and (3) is sensitive to the set of evaluation questions used. Collectively, these results indicate that WG is a product of quite fragile properties of both training and evaluation data. As such, we argue that WG is more plausible as an adversarial threat-requiring careful data engineering-rather than as a significant hazard inherent to routine fine-tuning.
Chinese Translation
在小规模、特定领域的数据集上进行狭义微调可能会导致模型行为的广泛且令人惊讶的变化——这一现象被称为奇异泛化(Weird Generalization, WG)。然而,目前尚不清楚微调数据的哪些特征是WG产生所必需的。在此,我们通过研究一系列可能相关的特征来解决这一问题,包括数据集的大小、组成、语言、呈现风格以及相对于模型参数知识的新颖性。此外,由于WG评估依赖于小规模的问题集来评估泛化的程度,我们还分析了这一测量对所使用问题集的敏感性。对三个开放权重模型在四个数据集上的实验表明,WG的程度(1)在很大程度上依赖于数据集的组成和语言(而非大小);(2)对于来自预训练的熟悉数据,WG程度大于新颖数据;(3)对所使用的评估问题集敏感。综合来看,这些结果表明WG是训练和评估数据的相当脆弱特性的产物。因此,我们认为WG更可能作为一种对抗性威胁存在——需要谨慎的数据工程——而不是作为常规微调固有的重大危害。
cs.CL / 139 / 2608.23507

When Names Cross Scripts: A Source-Grounded Benchmark for Historical Entity Reconciliation in the Mongol World

当名称跨越书写系统:蒙古世界历史实体对齐的源头基础基准
Chen, Xiang, Zhang, Zeyu
Abstract
Historical people may appear under different languages, scripts, and transcription traditions, while distinct individuals may share highly similar or even identical names. This makes historical identity reconciliation more than a problem of string matching or transliteration. We introduce MHER, a provenance-controlled benchmark for pairwise reconciliation of person-name attestations from the Mongol world. MHER contains a balanced 396-pair Name-only core over 84 primary historical persons and a stricter 160-pair Source-grounded subset constructed from mention-by-source evidence, with entity-disjoint development and test splits. Across five generative systems, correctly Source-grounded evidence improves paired TEST accuracy by 12.96 to 94.44 percentage points relative to Name-only input. On five identical-surface different-person cases, all models fail under names alone (0/25 model-item decisions), whereas Source-grounded evidence yields 24/25 correct resolutions, with the remaining output an abstention. Context-only ablations show that historical descriptions often carry substantial identity information, while explicitly signaled misgrounding controls produce substantially lower performance. We also find that names are not uniformly beneficial: for Qwen3-8B, restoring surface forms converts ten otherwise correct Context-only distinctions into false identity merges. These results show that historical entity reconciliation depends not only on surface correspondence, but on whether identity judgments respond appropriately to provenance-controlled historical evidence. MHER therefore provides a controlled framework for studying evidence use, abstention, and failure modes in historical NLP.
Chinese Translation
历史人物可能以不同的语言、书写系统和转录传统出现,而不同个体可能共享高度相似甚至相同的名称。这使得历史身份对齐不仅仅是字符串匹配或音译的问题。我们介绍了MHER,这是一个用于蒙古世界人名证明的成对对齐的来源控制基准。MHER包含一个平衡的396对仅名称核心,涵盖84位主要历史人物,以及一个更严格的160对基于来源的子集,构建于来源证据的提及上,具有实体不重叠的发展和测试分割。在五个生成系统中,相较于仅名称输入,正确的基于来源的证据使得成对测试的准确率提高了12.96到94.44个百分点。在五个表面相同但不同个体的案例中,所有模型在仅依赖名称时均未能成功(0/25模型项决策),而基于来源的证据则产生了24/25的正确解析,剩余的输出为弃权。仅上下文的消融实验表明,历史描述通常携带大量身份信息,而明确标示的错误基础控制则导致显著较低的性能。我们还发现名称并非始终有利:对于Qwen3-8B,恢复表面形式使得十个原本正确的仅上下文区分转变为错误的身份合并。这些结果表明,历史实体对齐不仅依赖于表面对应,还依赖于身份判断是否能适当地响应源头控制的历史证据。因此,MHER提供了一个受控框架,用于研究历史自然语言处理中的证据使用、弃权和失败模式。
cs.CL / 140 / 2608.23551

ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings

ConvergeFlow:具有可证明收敛性的语言流与标记嵌入
Li, Na, Jiao, Yuchen, Cai, Changxiao, Li, Gen
Abstract
Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive with discrete LMs. However, existing continuous frameworks still rely on decoders supervised with cross entropy (CE) because the flow trajectories are not guaranteed to terminate at valid token embeddings. Motivated by this limitation, we introduce \textbf{ConvergeFlow}, an embedding-space flow-based LM, which constrains the data predictor to the convex hull of token embeddings and trains it solely with the mean squared error objective induced by flow matching. Under suitable regularity conditions, we prove that the resulting flow converges to valid token embeddings despite errors in the data predictor, enabling direct token prediction without a CE-supervised decoder. We further develop three sampling mechanisms for controlling the trade-off between the generative perplexity and entropy. Experiments on OpenWebText demonstrate that ConvergeFlow achieves performance competitive with existing continuous and discrete diffusion LMs. These findings demonstrate the potential of the flow-based paradigm for language modeling. Our code is available at https://github.com/Na-Li66/ConvergeFlow.
Chinese Translation
最近在连续扩散和基于流的语言模型(LMs)方面的进展已达到了与离散 LMs 竞争的性能。然而,现有的连续框架仍依赖于通过交叉熵(CE)监督的解码器,因为流轨迹并不保证终止于有效的标记嵌入。基于这一限制,我们提出了 extbf{ConvergeFlow},一种嵌入空间的基于流的语言模型,它将数据预测器限制在标记嵌入的凸包内,并仅通过流匹配引起的均方误差目标进行训练。在适当的正则性条件下,我们证明了尽管数据预测器存在误差,所得到的流仍然收敛于有效的标记嵌入,从而实现无需 CE 监督解码器的直接标记预测。我们进一步开发了三种采样机制,以控制生成困惑度和熵之间的权衡。在 OpenWebText 上的实验表明,ConvergeFlow 的性能与现有的连续和离散扩散 LMs 具有竞争力。这些发现展示了基于流的范式在语言建模中的潜力。我们的代码可在 https://github.com/Na-Li66/ConvergeFlow 获取。
cs.CL / 141 / 2608.23564

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

SWE Refactor Bench:编码代理能否完成长时间跨度的整个代码库迁移?
Hong, Deyao, Chi, Yizhe, Li, Wenyi, Wang, Xiaoqiu, Gao, Mingju, Yang, Kaisen, He, Bingxiang, Zheng, Youjie, Xiao, Calvin, Na, Qinhuai
Abstract
Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing benchmarks cannot answer this question because they evaluate only behavioural correctness, not whether the migration actually occurred. This leads an easy hack: agents copy the original implementation to make tests pass. We call this Blindness. To address this problem, we introduce SWE Refactor Bench, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt. A three-stage evaluation protocol measures both migration completeness and behavioural correctness. (1) Migration Audit verifies that the migration occurred. (2) Behavioural Tests measure correctness with a fixed test suite. (3) Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 of 520 runs ($5.4\%$) pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model (claude-opus-5) scores $47.0/100$. Migration completeness and behavioural correctness are distinct abilities: a few runs preserve behaviour by skipping the migration and are stopped at Migration Audit; most attempt it and break behaviour, and are stopped at Behavioural Tests. Agents cannot deliver a perfect migration: among the 340 runs that pass Migration Audit, $58\%$ reach $99\%$ of the fixed checks, yet only $26\%$ reach $100\%$. Agent capability differs across migration categories: agents score $31.4$ on build toolchain rewrites but only $5.6$ on language rewrites. Together, these findings position SWE Refactor Bench as a rigorous testbed for developing coding agents for reliable whole-repository migrations.
Chinese Translation
现代软件系统在数十年的开发过程中积累了技术债务,这使得迁移变得昂贵且主要依赖人工。随着编码代理在修复错误方面能力的不断增强,它们能否自主执行这样的迁移?现有的基准测试无法回答这个问题,因为它们仅评估行为的正确性,而不考虑迁移是否实际发生。这导致了一种简单的作弊方式:代理复制原始实现以通过测试。我们称之为“盲目性”(Blindness)。为了解决这个问题,我们引入了SWE Refactor Bench,这是一个包含20个整个代码库迁移的基准,涵盖4种技术债务。一个三阶段的评估协议测量迁移的完整性和行为的正确性。(1) 迁移审计(Migration Audit)验证迁移是否发生。(2) 行为测试(Behavioural Tests)使用固定的测试套件测量正确性。(3) 代理验证(Agentic Verification)使用6个独立的编码代理生成针对隐藏行为差异的目标测试。在来自8个前沿模型和26个模型-努力配置的520次运行中,只有520次运行中的28次($5.4 ext{ extperthousand}$)通过了所有三个阶段,20个任务中的13个没有获得接受的解决方案,最佳模型(claude-opus-5)的得分为$47.0/100$。迁移的完整性和行为的正确性是不同的能力:一些运行通过跳过迁移来保持行为,并在迁移审计中被停止;大多数尝试迁移却破坏了行为,并在行为测试中被停止。代理无法提供完美的迁移:在340次通过迁移审计的运行中,$58 ext{ extperthousand}$达到了99 ext{ extperthousand}$的固定检查,但只有$26 ext{ extperthousand}$达到了100 ext{ extperthousand}$。代理在不同迁移类别中的能力差异显著:代理在构建工具链重写上得分为$31.4$,而在语言重写上仅得分$5.6$。这些发现共同将SWE Refactor Bench定位为开发可靠的整个代码库迁移的编码代理的严格测试平台。