cs.RO / 1 / 2608.27497
Beyond Relative Geometry: Metric-Aware Geometry Perception for Robotics
超越相对几何:面向机器人技术的度量感知几何
Abstract
Recent embodied models increasingly leverage geometric representations to improve spatial reasoning and robotic manipulation. However, existing reconstruction methods only reconstruct relative geometry with arbitrary scales, causing predicted object dimensions and spatial distances to vary across scenes, viewpoints, and input configurations. This inconsistency prevents geometric perception from being directly aligned with robotic actions defined on the real-world scale. To address this limitation, we propose Metric-Aware Geometry Perception (MAGP), an end-to-end, plug-and-play framework for metric geometry reconstruction that can be seamlessly integrated into robotic policies. At its core, Metric Scale Equivariant Augmentation encourages the model to reconstruct metric geometry from camera parameters and depth observations, ensuring that the reconstructed geometry follows the metric scale specified by observations. Flexible Metric Conditioning further enables MAGP to support arbitrary view counts and combinations of camera and depth inputs, improving robustness to heterogeneous robotic sensing configurations. Together, these designs produce geometrically consistent reconstructions with stable object dimensions and spatial distances across scenes and sensing conditions. Experiments on ETH3D, MegaDepth, and ScanNet++ demonstrate that MAGP maintains strong relative geometry accuracy while reducing the absolute error by over an order of magnitude, from 2.01m to 0.07m. When integrated into multiple robotic policies, MAGP consistently improves performance on LIBERO, RoboTwin, and zero-shot LIBERO-Plus, with gains of up to 6.26% on RoboTwin. These results demonstrate the effectiveness and generalizability of metric geometry for robotic manipulation.
Chinese Translation
近年来,具身模型越来越多地利用几何表示来改善空间推理和机器人操作。然而,现有的重建方法仅重建具有任意尺度的相对几何,这导致预测的物体尺寸和空间距离在不同场景、视角和输入配置中有所变化。这种不一致性阻碍了几何感知与基于真实世界尺度定义的机器人动作之间的直接对齐。为了解决这一局限性,我们提出了度量感知几何(Metric-Aware Geometry Perception, MAGP),这是一个端到端的即插即用框架,用于度量几何重建,可以无缝集成到机器人策略中。MAGP的核心是度量尺度等变增强(Metric Scale Equivariant Augmentation),它鼓励模型根据相机参数和深度观测重建度量几何,确保重建的几何遵循观测所指定的度量尺度。灵活的度量条件(Flexible Metric Conditioning)进一步使MAGP能够支持任意数量的视角以及相机和深度输入的组合,提高了对异构机器人传感配置的鲁棒性。这些设计共同产生了在不同场景和传感条件下具有稳定物体尺寸和空间距离的几何一致性重建。在ETH3D、MegaDepth和ScanNet++上的实验表明,MAGP在保持强相对几何准确性的同时,将绝对误差减少了一个数量级,从2.01米降低到0.07米。当集成到多个机器人策略中时,MAGP在LIBERO、RoboTwin和零样本LIBERO-Plus上的性能均有显著提升,RoboTwin的提升幅度高达6.26%。这些结果证明了度量几何在机器人操作中的有效性和普适性。
cs.RO / 2 / 2608.27545
Remote Human and Robot Interaction for Greenhouse Gardening Using Virtual Reality
基于虚拟现实的温室园艺远程人机交互研究
Abstract
This study evaluates the effectiveness of remote human-robot interaction using virtual reality for leaf inspection and soil moisture assessment in a greenhouse environment. The robotic system comprised an unmanned ground vehicle and a robotic manipulator equipped with cameras, governed by kinematic models for navigation and manipulator control. Fourteen distinct plants were inspected across two experiments utilizing VR teleoperation, guided by a set of pre-specified research questions and hypotheses. In the leaf inspection experiments, cycle completion times varied from 3.3 to 8.0 s, and plant-based disease detection was achieved up to 88% accuracy; diseased-spot detection improved numerically in the second experiment, though this change was not statistically significant (p=0.378). For soil moisture assessment, the experiments achieved successful determination of watering needs in up to 64.3% of plants (9 of 14), with consistent success observed for plants 1, 2, 3, 8, 9, 10, and 13; however, this improvement was likewise not statistically significant (p=0.50). A post hoc analysis instead revealed that soil moisture assessment reliability was strongly and significantly predicted by plant canopy morphology (p<0.01): plants with broad, single-leaf canopies reached 100% success by the second experiment, versus only 16.7% for dense, compound canopies. A secondary analysis showed operators became measurably faster at attempting dense-canopy plants without a corresponding gain in success, indicating that camera occlusion, not operator skill or effort, is the dominant limiting factor. These findings show occlusion imposes a sensing limitation rather than a control or training deficiency, and that adapting camera viewpoint and sensing strategy to canopy density is needed to improve the system's accuracy and robustness.
Chinese Translation
本研究评估了在温室环境中使用虚拟现实进行远程人机交互在叶片检查和土壤湿度评估方面的有效性。该机器人系统由一辆无人地面车辆和一台配备摄像头的机器人操纵器组成,导航和操纵器控制由运动学模型管理。在两个实验中,共检查了十四种不同的植物,采用虚拟现实远程操作,依据一组预设的研究问题和假设进行指导。在叶片检查实验中,周期完成时间从3.3秒到8.0秒不等,植物病害检测的准确率达到了88%;尽管在第二个实验中病斑检测的数值有所提高,但这一变化在统计上并不显著(p=0.378)。在土壤湿度评估方面,实验成功确定了最多64.3%的植物(14种中的9种)的浇水需求,对于植物1、2、3、8、9、10和13的成功率保持一致;然而,这一改善同样在统计上并不显著(p=0.50)。后续分析表明,土壤湿度评估的可靠性受到植物冠层形态的强烈和显著预测(p<0.01):具有宽大单叶冠层的植物在第二个实验中达到了100%的成功率,而密集复合冠层的植物仅为16.7%。二次分析显示,操作员在尝试密集冠层植物时的速度显著提高,但成功率并未相应提升,表明摄像头遮挡而非操作员的技能或努力是主要限制因素。这些发现表明,遮挡造成了感知限制,而非控制或训练缺陷,并且需要根据冠层密度调整摄像头视角和感知策略,以提高系统的准确性和鲁棒性。
cs.RO / 3 / 2608.27550
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
超越数据扩展:面向视觉-语言-动作模型的以表征为中心的持续预训练
Abstract
Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.
Chinese Translation
扩展机器人数据对于构建通用的视觉-语言-动作(VLA)模型至关重要,但与网络规模的图像-文本数据相比,机器人轨迹的扩展更具挑战性,因为实体收集成本高且对物理世界的覆盖稀疏。这使得表征质量成为一个核心瓶颈:在固定的机器人数据预算下,持续预训练必须将有限的轨迹转化为可转移的视觉-动作知识,而不仅仅是适应动作。我们提出了VLAct,这是一个面向VLA的视觉-语言模型(VLM)主干,在任务特定的微调之前,基于广泛的异质多实体机器人数据进行训练。VLAct保留了广泛的VLM先验,并通过VLM先验保留、多头连续动作共同监督和部分统一的跨实体动作布局,鼓励不同实体之间共享动作语义,同时在微调过程中允许任务特定的动作头。在模拟、现实世界和未见实体的迁移中,VLAct在固定的微调协议下持续提高下游性能。在LIBERO-Plus和RoboTwin 2.0上,VLAct超越了包括ABot-M0和LingBot-VLA在内的工业VLA系统,成功率分别达到82.6%和92.5%。在RoboDojo上,VLAct在所有策略中按成功率排名第六,并在两个指标上超越了所有明确指定的世界-动作模型(WAM)条目。最值得注意的是,在未见的人形实体RoboCasa-GR1上,VLAct仅使用20%的下游轨迹便超越了完整数据的GR00T-N1.6基线。这些结果是在完全开源的数据和仅使用16个GPU的训练设置下获得的,显示出以表征为中心的持续预训练能够在适度的计算预算下提供高度竞争的性能,并且是VLA进展中超越数据扩展的重要独立方向。
cs.RO / 4 / 2608.27609
PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models
PHR-VLA:视觉-语言-动作模型的规划视野推理
Abstract
Vision-language-action models (VLAs) have shown strong promise for general-purpose robotic manipulation by mapping language instructions and vision observations directly to actions. However, most VLAs primarily condition action prediction on current observations and lack an explicit mechanism for reasoning over future task dynamics, which is particularly important for fine-grained, contact-rich manipulation. We present PHR-VLA, a framework that enables planning-horizon reasoning in VLAs through privileged latent representations of future dynamics. PHR-VLA introduces a lightweight auxiliary future head that, during training, aligns the VLA's internal representations with latent dynamics extracted from future observations. Evaluation results demonstrate that local, contact-centric, patch-level latent dynamics supervision from the wrist camera improves success rate on LIBERO from 84.1% to 88.4% and on real-world disassembly tasks from 63.3% to 82.5%. Patch-level supervision from a third-person camera also improves performance on Meta-World from 56.70% to 57.8%. These results demonstrate that privileged latent dynamics alignment provides an effective training signal for improving anticipatory reasoning in VLA policies. Project website: \href{https://davoodsz.github.io/PHR-VLA.github.io/}{https://davoodsz.github.io/PHR-VLA.github.io/}
Chinese Translation
视觉-语言-动作模型(VLA)在通用机器人操作中展现了强大的潜力,通过将语言指令和视觉观察直接映射到动作。然而,大多数VLA主要基于当前观察来预测动作,缺乏对未来任务动态进行推理的明确机制,这对于细粒度的、接触丰富的操作尤为重要。我们提出了PHR-VLA,一个通过未来动态的特权潜在表示实现VLA中规划视野推理的框架。PHR-VLA引入了一个轻量级的辅助未来头,在训练过程中,将VLA的内部表示与从未来观察中提取的潜在动态对齐。评估结果表明,来自手腕摄像头的局部、以接触为中心的补丁级潜在动态监督使LIBERO上的成功率从84.1%提高到88.4%,在真实世界的拆解任务中从63.3%提高到82.5%。来自第三人称摄像头的补丁级监督也使Meta-World上的表现从56.70%提高到57.8%。这些结果表明,特权潜在动态对齐为改善VLA策略中的预期推理提供了有效的训练信号。项目网站: [https://davoodsz.github.io/PHR-VLA.github.io/](https://davoodsz.github.io/PHR-VLA.github.io/)
cs.RO / 5 / 2608.27628
One year in a forest: Analyzing the challenges of autonomous navigation in subarctic environments
森林中的一年:分析亚北极环境中自主导航的挑战
Abstract
Subarctic regions have the potential to see increased deployment of autonomous robots in applications including forestry, mining, and environmental monitoring. In these conditions, an autonomous system's reliance on GNSS or cloud computing is precarious due to dense tree canopies and atmospheric attenuation, necessitating onboard sensing and data processing. However, established exteroceptive modalities, including cameras, lidars, and radars, are typically evaluated in structured urban settings or in environments that lack significant seasonal variations. To address this, we present a field report on a year-long deployment of a mobile robot in a subarctic boreal forest. We evaluate 64 km of data using nine odometry, localization, and mapping methods and assess their performance across seasonal changes. The performed experiments suggest that the environment changes significantly hinder the performance of state-of-the-art techniques, which show increased fragility when subject to conditions characterized by self-similar scenes or tall snowbanks. Additionally, complex Simultaneous Localization and Mapping (SLAM) algorithms offer limited accuracy gains over a proprioceptive baseline while significantly increasing system fragility. Furthermore, by correlating the position drift with features and confidence weight distribution, we show that visual-based SLAM methods are particularly affected by the seasonal changes. Additionally, we investigate the task of cross-season localization in a prior map. While lidar-based methods successfully completed localization runs between seasons, radar and visual methods are prone to failure due to a few matching features between runs, even within the same season. Finally, we detail the challenges and lessons learned from this year-long trial, including a multi-season Teach and Repeat (T&R) evaluation using both radar and lidar-based pipelines.
Chinese Translation
亚北极地区在林业、采矿和环境监测等应用中有望增加自主机器人部署。在这些条件下,由于树冠密集和大气衰减,自主系统对全球导航卫星系统(GNSS)或云计算的依赖变得不稳定,因此需要进行机载传感和数据处理。然而,现有的外部感知模式,包括摄像头、激光雷达(lidar)和雷达,通常是在结构化的城市环境或缺乏显著季节变化的环境中进行评估的。为了解决这一问题,我们提供了一份关于在亚北极针叶林中部署移动机器人的为期一年的现场报告。我们使用九种里程计、定位和地图构建方法评估了64公里的数据,并评估了它们在季节变化中的表现。实验表明,环境变化显著阻碍了最先进技术的性能,当面临自相似场景或高雪堆等条件时,这些技术表现出更大的脆弱性。此外,复杂的同时定位与地图构建(SLAM)算法在精度提升方面相较于本体感知基线的收益有限,同时显著增加了系统的脆弱性。此外,通过将位置漂移与特征和置信权重分布相关联,我们表明基于视觉的SLAM方法特别受到季节变化的影响。此外,我们还研究了在先前地图中进行跨季节定位的任务。尽管基于激光雷达的方法在季节间成功完成了定位,但雷达和视觉方法由于运行间匹配特征较少而容易失败,即使在同一季节内。最后,我们详细介绍了这一年试验中的挑战和经验教训,包括使用基于雷达和激光雷达的流程进行的多季节教学与重复(T&R)评估。
cs.RO / 6 / 2608.27685
Distributed Model-Based Diffusion: Finite Horizon Contraction under Bounded Delay
分布式基于模型的扩散:有限时间范围内的收缩在有界延迟下
Abstract
Simultaneously optimizing the trajectories of multiple agents is a challenging problem plagued by nonlinearity, nonconvexity, and the curse of dimensionality. A collection of interacting aerial vehicles or self-driving cars in an intersection are examples of complex multi-agent systems that remain difficult to solve without many simplifying assumptions. The presence of communication latency between agents further increases the difficulty. In this paper, we analyze Distributed Model-Based Diffusion: a sampling-based Model-Predictive Control method suitable for highly nonlinear, nonconvex, nonsmooth, multi-agent systems. We prove contraction and robustness to latency for multi-agent, nonconvex problems, showing applicability to real-world constraints. We test the algorithm on a circleswap task, a cooperative medium-fidelity driving task, and in an aerial combat scenario. Despite the addition of latency, our algorithm improves circleswap makespan by 31% and increases aerial combat win rate by 25% compared to centralized Model-Based Diffusion.
Chinese Translation
同时优化多个智能体的轨迹是一个具有挑战性的问题,受到非线性、非凸性和维度诅咒的困扰。一组相互作用的空中车辆或在交叉口的自动驾驶汽车是复杂多智能体系统的例子,这些系统在没有许多简化假设的情况下仍然难以解决。智能体之间的通信延迟进一步增加了难度。在本文中,我们分析了分布式基于模型的扩散:一种适用于高度非线性、非凸、非光滑多智能体系统的基于采样的模型预测控制方法。我们证明了在多智能体非凸问题中,收缩性和对延迟的鲁棒性,显示其适用于现实世界的约束条件。我们在一个圆形交换任务、一个合作中等保真度的驾驶任务以及一个空中战斗场景中测试了该算法。尽管增加了延迟,我们的算法在圆形交换的完成时间上提高了31%,并且在空中战斗的胜率上提高了25%,与集中式基于模型的扩散相比。
cs.RO / 7 / 2608.27726
Coordinated Motion Planning for Multi-Arm Systems via Iterative LQ Games
通过迭代线性二次(LQ)博弈实现多臂系统的协调运动规划
Abstract
Multi-agent motion planning for high-degree-of-freedom robotics manipulators in shared workspaces remains a fundamental yet challenging problem. Centralized planners often suffer from poor scalability, while decentralized approaches face robustness and safety concerns. Game-theoretic formulations offer a promising approach for modeling agent interactions, potentially overcoming these limitations. However, their application to articulated multi-arm systems remains limited. This paper presents an iterative Linear Quadratic (LQ) game framework for multi-manipulator motion planning, where each manipulator is modeled as an independent agent optimizing its own objective while interacting with other agents based on shared global states and collision constraints. The method solves a series of local LQ games by linearizing the dynamics and approximating the cost around a nominal trajectory, with Riccati backward recursions yielding feedback Nash strategies. To address the challenges of articulated systems, we incorporate differentiable penalties for self-collision and inter-arm collision into the optimization pipeline, enabling coordinated, collision-aware trajectory generation. Experiments demonstrate that our framework produces smooth, safe, and efficient trajectories in high-dimensional settings, outperforming traditional methods. This highlights the effectiveness of differential game formulations for multi-robot manipulation.
Chinese Translation
在共享工作空间中,高自由度机器人操纵器的多智能体运动规划仍然是一个基本而具有挑战性的问题。集中式规划器通常面临扩展性差的问题,而分散式方法则存在鲁棒性和安全性方面的顾虑。博弈论的形式化提供了一种有前景的方法来建模智能体之间的互动,有可能克服这些局限。然而,其在关节式多臂系统中的应用仍然有限。本文提出了一种迭代线性二次(LQ)博弈框架,用于多操纵器的运动规划,其中每个操纵器被建模为一个独立的智能体,优化其自身目标,同时基于共享的全局状态和碰撞约束与其他智能体进行互动。该方法通过线性化动力学并在名义轨迹周围近似成本,解决一系列局部LQ博弈,利用Riccati反向递归得到反馈纳什策略。为了解决关节式系统的挑战,我们在优化流程中引入了自碰撞和臂间碰撞的可微惩罚,从而实现协调的、考虑碰撞的轨迹生成。实验表明,我们的框架在高维环境中生成平滑、安全且高效的轨迹,优于传统方法。这突显了微分博弈形式化在多机器人操纵中的有效性。
cs.RO / 8 / 2608.27793
CAVE-NAV: VLM-Based Autonomous 3D Navigation in Underwater Cave Environments
CAVE-NAV:基于视觉语言模型的水下洞穴环境自主三维导航
Abstract
Autonomous navigation in underwater cave environments is essential for search-and-rescue operations, scientific exploration, and emergency egress. Traditional navigation systems commonly depend on dense visual features for localization and mapping. In underwater caves, however, visual degradation can undermine feature-based localization, sonar-based mapping may yield overly conservative obstacle representations, and communication constraints preclude real-time human guidance. To address these limitations, we propose an autonomous underwater cave navigation framework that leverages a vision-language model (VLM) with Chain-of-Thought (CoT) reasoning to infer navigable directions from environmental cues, including light intensity gradients, passage morphology, and geometric complexity, captured through multimodal inputs comprising RGB imagery, depth maps, and sonar-based vertical-clearance measurements, thereby supporting safe 3D navigation through confined cave passages. High-fidelity simulations across multiple cave topologies demonstrate that the proposed framework completes all evaluated end-to-end traversals without collisions while maintaining safe clearance from cave boundaries.
Chinese Translation
在水下洞穴环境中的自主导航对于搜索与救援行动、科学探索和紧急撤离至关重要。传统导航系统通常依赖于密集的视觉特征进行定位和地图构建。然而,在水下洞穴中,视觉退化可能会削弱基于特征的定位,基于声纳的地图构建可能导致过于保守的障碍物表示,而通信限制则排除了实时人类指导。为了解决这些局限性,我们提出了一种自主水下洞穴导航框架,该框架利用视觉语言模型(VLM)结合思维链(CoT)推理,从环境线索中推断可导航方向,这些线索包括光强度梯度、通道形态和几何复杂性,这些信息通过多模态输入(包括RGB图像、深度图和基于声纳的垂直净空测量)捕获,从而支持在狭窄洞穴通道中的安全三维导航。针对多种洞穴拓扑的高保真模拟表明,所提出的框架在不发生碰撞的情况下完成了所有评估的端到端穿越,同时保持与洞穴边界的安全间距。
cs.RO / 9 / 2608.28075
Plan Along the Way: Event-Triggered Foundation-Model Planning for TAMP Execution in Partially Observable Manipulation
沿途规划:事件触发的基础模型规划用于部分可观测操作中的任务与运动规划执行
Abstract
Manipulation in partially observable environments requires planning under incomplete scene information. In such settings, an initially valid plan may execute successfully yet remain insufficient for task completion. Existing foundation-model-guided task and motion planning (TAMP) systems can generate useful long-horizon task decompositions, subgoals, or constraints, but they often assume having access to a fully specified scene state or invoke model-level replanning after a subgoal, refinement, or execution attempt fails. We present ROBUST TAMP, a modular LLM/VLM-guided planning framework for reactive TAMP where unseen task-relevant and non-target objects may become visible during execution. The framework restricts the foundation-model planner to the currently visible relational scene state, validates generated task-level actions against a strict executable interface, and routes the accepted actions to scene-specific execution adapters. Object discovery is treated as a distinct replanning event and, after a stable execution horizon, the system reconstructs the visible scene state and replans using completed-action history and structured replanning event context. Evaluations are performed on six RLBench/CoppeliaSim kitchen and grill variants involving hidden objects, non-target object discovery, articulated-container interaction, and temporal manipulation procedures. We compare text-only LLM and VLM planners of different sizes under the same validation, execution, monitoring, and replanning pipeline, reporting task success, partial goal completion, discovery- and failure-triggered replanning behavior, implicit non-target-object handling, and planner inference cost.
Chinese Translation
在部分可观测环境中进行操作需要在不完整的场景信息下进行规划。在这种情况下,最初有效的计划可能成功执行,但仍不足以完成任务。现有的基础模型指导的任务与运动规划(TAMP)系统能够生成有用的长时间跨度任务分解、子目标或约束,但它们通常假设可以访问完全指定的场景状态,或者在子目标、细化或执行尝试失败后调用模型级重新规划。我们提出了ROBUST TAMP,一个模块化的LLM/VLM指导的规划框架,用于反应式TAMP,其中在执行过程中可能会出现未见的任务相关和非目标对象。该框架将基础模型规划器限制在当前可见的关系场景状态,验证生成的任务级动作是否符合严格的可执行接口,并将接受的动作路由到场景特定的执行适配器。对象发现被视为一个独立的重新规划事件,在稳定的执行时间范围后,系统重建可见场景状态,并使用已完成的动作历史和结构化的重新规划事件上下文进行重新规划。我们在六个RLBench/CoppeliaSim厨房和烤架变体上进行了评估,涉及隐藏对象、非目标对象发现、关节容器交互和时间操作程序。我们比较了不同规模的文本仅LLM和VLM规划器在相同的验证、执行、监控和重新规划流程下的表现,报告了任务成功、部分目标完成、发现和失败触发的重新规划行为、隐式非目标对象处理以及规划器推理成本。
cs.RO / 10 / 2608.28090
Stay Seated: Learning Omnidirectional Humanoid Locomotion on a Passive Mobile Chair with Casters
保持坐姿:在带滚轮的被动移动椅上学习全方向类人运动
Abstract
Humanoid robots with quasi-direct-drive actuators continuously generate joint torque while standing, whereas seated humans delegate weight support to chairs during desk work. As a first step toward seated loco-manipulation, we study omnidirectional seated locomotion on a passive mobile chair, requiring unfixed pelvis-seat contact and intermittent foot-floor propulsion of the robot-chair system. We extend a standard standing velocity-tracking environment with a passive-chair model, seated-state rewards, critic-only chair observations, and task-tailored contact settings. The policy is learned without motion-imitation rewards; its actor uses only proprioception and velocity commands, without contact sensing or chair states. In random-command evaluation, the policies tracked omnidirectional commands through nearly all 20-s rollouts, and the best seated policies could outperform the Standing policy in velocity tracking. Across four training seeds, a $2^3$ full-factorial comparison of symmetry regularization (SY), foot-slip regularization (FS), and command curriculum (CC) showed that FS reduced CoT but increased tracking error and that some FS-only policies converged to stationary local optima. Combining FS with either SY or CC avoided this failure without retuning FS, while SY improved bilateral leg symmetry during longitudinal motion. Direction-resolved analysis showed CoT ordered backward $<$ lateral $\ll$ forward, with planted-leg extension in backward and lateral motion and knee flexion following heel contact in forward motion. The learned policy achieved zero-shot sim-to-real transfer to a Unitree G1 and generated omnidirectional seated locomotion.
Chinese Translation
类人机器人使用准直接驱动执行器在站立时持续产生关节扭矩,而坐着的人在桌面工作时则将体重支撑委托给椅子。作为向坐姿运动操控的第一步,我们研究了在被动移动椅上进行全方向坐姿运动,这需要机器人-椅子系统中骨盆与座椅的非固定接触和间歇性的脚-地面推动。我们扩展了标准的站立速度跟踪环境,加入了被动椅模型、坐姿状态奖励、仅评论者的椅子观察以及针对任务的接触设置。该策略的学习不依赖于运动模仿奖励;其执行者仅使用本体感觉和速度指令,而不依赖于接触感知或椅子状态。在随机指令评估中,策略在几乎所有20秒的试验中跟踪了全方向指令,最佳的坐姿策略在速度跟踪上超越了站立策略。在四个训练种子中,$2^3$的全因子比较显示,脚滑动正则化(FS)减少了成本(CoT),但增加了跟踪误差,并且一些仅使用FS的策略收敛到静态局部最优。将FS与对称性正则化(SY)或指令课程(CC)结合使用避免了这种失败,而无需重新调整FS,同时SY在纵向运动中改善了双腿的对称性。方向解析分析显示,成本(CoT)按顺序为向后 $<$ 侧向 $ ext{ll}$ 向前,向后和侧向运动中植腿伸展,而向前运动中跟随脚跟接触的膝关节屈曲。所学策略实现了对Unitree G1的零样本仿真到现实转移,并生成了全方向坐姿运动。
cs.RO / 11 / 2608.28108
DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA
DeicticVLA:基于语言和指示性手势统一指令模式的单一视觉语言行动模型
Abstract
Vision-Language-Action models (VLAs) allow users to specify manipulation tasks in natural language, but distinguishing a target or placement goal among objects of the same category or similar appearance requires detailed expressions that VLAs may not use reliably. We propose DeicticVLA, which canonicalizes Language Instruction (LI), Vision-Language Instruction (VLI), and Visual Instruction (VI) into a text prompt and deictic masks through text-prompt completion and deictic gesture grounding, enabling a single pretrained VLA to handle all three instruction modes. With a shared backbone, demonstrations, and matched training steps, we compare two RGB visual prompting methods, two separate-channel mask prompting methods, and three training strategies in simulation. Under two-stage training, the four prompting methods achieve high in-distribution success but differ in their ability to use deictic masks in unseen layouts. Across methods, training-strategy ablations show that two-stage training improves such use, while retaining second-stage LI data mitigates forgetting without reducing VLI and VI performance. In three real-world tasks, one policy supports all modes. VLI and VI outperform LI under unseen expressions, appearance changes, and novel objects. For unseen categories, both achieve 100% success, compared with 16.7% for jointly trained LI. These results demonstrate the unified three-mode interface and guide DeicticVLA design.
Chinese Translation
视觉-语言-行动模型(VLAs)允许用户用自然语言指定操作任务,但在同一类别或相似外观的物体中区分目标或放置目标需要详细的表达,而VLAs可能无法可靠地使用这些表达。我们提出了DeicticVLA,它通过文本提示补全和指示性手势定位,将语言指令(LI)、视觉-语言指令(VLI)和视觉指令(VI)规范化为文本提示和指示性掩码,使得单个预训练的VLA能够处理这三种指令模式。通过共享的主干网络、演示和匹配的训练步骤,我们在模拟中比较了两种RGB视觉提示方法、两种独立通道掩码提示方法和三种训练策略。在两阶段训练下,这四种提示方法在分布内成功率高,但在未见布局中使用指示性掩码的能力存在差异。在不同方法中,训练策略的消融实验表明,两阶段训练提高了这种使用能力,而保留第二阶段的LI数据则减轻了遗忘,同时不降低VLI和VI的性能。在三个真实世界任务中,一种策略支持所有模式。在未见表达、外观变化和新物体的情况下,VLI和VI的表现优于LI。对于未见类别,两者的成功率均为100%,而联合训练的LI仅为16.7%。这些结果展示了统一的三模式接口,并为DeicticVLA的设计提供了指导。
cs.RO / 12 / 2608.28140
Contact-Guided Exploration for Non-Prehensile Locomanipulation with Multi-Critic RL
基于接触引导的非抓取式运动操控探索与多评论者强化学习
Abstract
Non-prehensile manipulation offers versatile skills for moving and rearranging heavy or bulky objects, particularly when combined with a mobile manipulation platform. However, both model-based and model-free approaches struggle with the complex hybrid dynamics and the sparsity of the contact in these tasks. To address these challenges, we propose a contact-guided exploration strategy implemented within a Multi-Critic Reinforcement Learning (RL) framework. A dedicated exploration critic is trained with a dense contact-seeking reward that guides the end-effector toward meaningful contact points; its influence is progressively decayed to recover a task-optimal policy. We obtain candidate interaction points from a general-purpose grasping algorithm, enabling the exploration mechanism to generalise across various object geometries. We evaluate the approach on multiple tasks, including box pushing, chair transportation, and a dishwasher opening task. Finally, we validate the chair transportation policy through extensive experiments on a quadrupedal mobile manipulator, demonstrating deployable non-prehensile manipulation in the real world.
Chinese Translation
非抓取式操控为移动和重新排列重物或笨重物体提供了多种灵活技能,特别是在与移动操控平台结合时。然而,无论是基于模型的方法还是无模型的方法,在这些任务中都面临复杂的混合动力学和接触稀疏性的问题。为了解决这些挑战,我们提出了一种在多评论者强化学习(Multi-Critic Reinforcement Learning, RL)框架内实施的接触引导探索策略。我们训练了一个专门的探索评论者,使用密集的接触寻求奖励来引导末端执行器朝向有意义的接触点;其影响逐渐衰减,以恢复任务最优策略。我们从通用抓取算法中获取候选交互点,使得探索机制能够在各种物体几何形状中进行泛化。我们在多个任务上评估了该方法,包括箱子推动、椅子运输和洗碗机开启任务。最后,我们通过在四足移动操控器上的大量实验验证了椅子运输策略,展示了在现实世界中可部署的非抓取式操控。
cs.RO / 13 / 2608.28154
From Small Talk to Rapport: Exploring Robot Self-Disclosure in Collaborative Tasks
从闲聊到建立关系:探索机器人在协作任务中的自我披露
Abstract
People naturally chat while collaborating and share personal information (i.e., self-disclose) to build rapport and maintain social connections. As robots are increasingly developed to work with people, the effective use of these social behaviors to enhance engagement and support teamwork becomes ever more important. While prior work has shown that robot-initiated small talk can benefit human-robot collaboration, less is known about how best to design such small talk. In this work, we explore how self-disclosure may be designed to support small talk within a human-robot team---especially when the robot is an industrial manipulator that lacks anthropomorphic cues and performs physical work. We first developed an LLM-driven manipulator capable of partaking in small talk, adopting either a low-disclosure or high-disclosure strategy. We then conducted a user study (N = 50) to investigate how self-disclosure in small talk influences human-robot dynamics. Unexpectedly, participants disclosed more in the low-disclosure condition and reported stronger teaming and coordination than those in the high-disclosure condition. This effect was more pronounced among users with prior experience teaming with robots. These results suggest that increasing robot self-disclosure does not necessarily foster rapport, social connection, or reciprocal disclosure; other factors, such as prior HRI experience, should be considered.
Chinese Translation
人们在协作时自然会聊天并分享个人信息(即自我披露),以建立关系并维持社会联系。随着机器人技术的不断发展,机器人与人类协作的有效社交行为的运用变得愈加重要,以增强参与感和支持团队合作。虽然先前的研究表明,机器人发起的闲聊可以促进人机协作,但关于如何设计这种闲聊的有效性仍知之甚少。在本研究中,我们探讨了如何设计自我披露以支持人机团队中的闲聊——尤其是在机器人作为缺乏人性化线索并执行物理工作的工业机械手时。我们首先开发了一种基于大语言模型(LLM)的机械手,能够参与闲聊,并采用低披露或高披露策略。随后,我们进行了用户研究(N = 50),调查自我披露在闲聊中如何影响人机互动的动态。意外的是,参与者在低披露条件下披露的信息更多,并且报告的团队合作和协调感比高披露条件下的参与者更强。这种效果在与机器人有过合作经验的用户中更为明显。这些结果表明,增加机器人的自我披露并不一定促进关系、社会联系或互惠披露;其他因素,如先前的人机互动经验,也应被考虑在内。
cs.RO / 14 / 2608.28175
Picking Bins Empty: A Hierarchical Hybrid Approach with Online Self-Learning of Grasp Points for Reliable Industrial Bin-Picking
清空料箱:一种具有在线自学习抓取点的分层混合方法,用于可靠的工业料箱拾取
Abstract
Bin-picking is a cornerstone of modern manufacturing, yet achieving complete bin clearance without manual intervention remains a critical challenge. While model-based methods provide high precision, they frequently suffer from deadlocks when predefined grasps are occluded or perception fails. Labor-intensive fine-tuning of grasp points is commonly required to reach a satisfactory performance for new parts. Model-free algorithms offer a more flexible alternative with "out-of-the-box" versatility but lack the reliability and repeatability required for production. Unlike existing work, which treats the two techniques in isolation, we propose a fourtiered hierarchical hybrid approach to combine the best of both worlds. A model-based pipeline serves as a robust backbone, while a model-free "exploration agent" resolves deadlock situations and discovers new grasp points. This is supported by an online self-learning mechanism that uses gripper-stroke feedback and Wilson score intervals to autonomously rank grasp candidates, reducing manual commissioning effort. Validation on three automotive parts demonstrates that our method significantly outperforms a model-free baseline in grasp success rate while improving the bin clearance rate of the model-based baseline from 50.9% to 100% across all experiments. This transition to full bin clearance marks a significant step towards truly autonomous, intervention-free industrial operation.
Chinese Translation
料箱拾取是现代制造业的基石,但在没有人工干预的情况下实现完全清空料箱仍然是一个关键挑战。尽管基于模型的方法提供了高精度,但当预定义的抓取点被遮挡或感知失败时,它们常常会遭遇死锁。为了达到对新零件令人满意的性能,通常需要进行劳动密集型的抓取点微调。无模型算法提供了一种更灵活的替代方案,具有“开箱即用”的多功能性,但缺乏生产所需的可靠性和重复性。与现有研究将这两种技术孤立对待不同,我们提出了一种四层分层混合方法,以结合两者的优势。基于模型的管道作为一个稳健的支撑,而无模型的“探索代理”则解决死锁情况并发现新的抓取点。这一过程得到了在线自学习机制的支持,该机制利用夹具行程反馈和Wilson评分区间自主对抓取候选进行排名,从而减少人工调试的工作量。在对三种汽车零件的验证中,我们的方法在抓取成功率上显著优于无模型基线,同时将基于模型的基线的料箱清空率从50.9%提高到100%。这一向完全清空料箱的转变标志着朝着真正自主、无需干预的工业操作迈出了重要一步。
cs.RO / 15 / 2608.28213
PAMoR: Parameterized Affective Motion Generation in Real Time for Humanoid Robots
PAMoR:用于类人机器人实时参数化情感运动生成
Abstract
People read a humanoid robot's motion in social settings not only for the action performed but for the affect conveyed. Motion carrying that affect has so far been generated for human avatars, where style is taken from a reference clip or an emotion word, neither of which can be quantitatively parameterized. We present PAMoR, which turns affect into a measured control parameter: a valence-arousal (V-A) coordinate computed natively on robot kinematics. It is obtained in closed form from postural expansion and movement energy, and these measurements serve directly as generation conditions, with no human annotation. An action prior and two affect priors, trained in a shared latent space, are composed at each denoising step: the action prior fixes what is performed, the affect priors modulate how. Whole-body motion rolls out autoregressively on a 29-DoF Unitree G1 in real time, with action and affect both editable. Generated motion tracks the commanded V-A over its full range while text-to-motion fidelity still matches text-only baselines. In a perceptual study, raters identify the commanded emotion on 0.38 of trials, above both baselines and approaching the 0.44 reported for acted human bodies.
Chinese Translation
人们在社交场合中解读类人机器人的运动,不仅关注所执行的动作,还关注所传达的情感。迄今为止,承载这种情感的运动主要是为人类虚拟形象生成的,其风格来源于参考片段或情感词,这两者都无法进行定量参数化。我们提出了PAMoR,它将情感转化为一个可测量的控制参数:一个基于机器人运动学计算的价-唤醒(V-A)坐标。该坐标通过姿态扩展和运动能量的闭式形式获得,这些测量直接作为生成条件,无需人工标注。在每个去噪步骤中,训练于共享潜在空间的动作先验和两个情感先验被组合:动作先验固定了执行的内容,而情感先验调节了执行的方式。全身运动在29自由度的Unitree G1上实时自回归生成,动作和情感均可编辑。生成的运动在其整个范围内跟踪指令的V-A,同时文本到运动的保真度仍然与仅基于文本的基线相匹配。在一项感知研究中,评估者在38%的试验中识别出指令情感,超过了两个基线,并接近于对人类身体表演的0.44的报告结果。
cs.RO / 16 / 2608.28214
Probabilistic Multi-Robot Gas Source Localization with Uncalibrated Sensors: A Distributed Estimation Approach
基于未校准传感器的概率多机器人气体源定位:一种分布式估计方法
Abstract
Estimating environmental states with multi-robot systems becomes particularly challenging when robots are equipped with uncalibrated and therefore heterogeneous sensors, whose nonlinear and inconsistent responses prevent reliable information fusion. In this paper, we propose a distributed probabilistic framework for source localization tasks that enables calibration-free estimation in the presence of sensor heterogeneity. The key idea is that each robot independently estimates a local belief using a rank-based feature that captures the relative evolution of observations and is invariant to sensor scaling and nonlinearities. These local beliefs are then fused through a product of experts formulation to obtain a consistent global estimate across the team. To further improve the efficiency of team coordination, we introduce an informative region allocation and path planning strategy that reduces redundant exploration while balancing exploration and exploitation. We validate the proposed framework using high-fidelity simulations with realistic gas sensor models. Results demonstrate that our method significantly outperforms a benchmark method based on standard measurement aggregation, achieving reliable source localization accuracy despite strong sensor heterogeneity. More broadly, this work demonstrates how calibration-free sensing representations can be effectively extended to distributed robotic systems, paving the way for their application to other estimation tasks involving heterogeneous sensors.
Chinese Translation
当多机器人系统配备未校准且因此具有异质性的传感器时,估计环境状态变得尤为具有挑战性,因为这些传感器的非线性和不一致响应阻碍了可靠的信息融合。本文提出了一种用于源定位任务的分布式概率框架,该框架能够在传感器异质性存在的情况下实现无校准估计。其关键思想是每个机器人独立使用一种基于排名的特征来估计局部信念,该特征捕捉观察结果的相对演变,并且对传感器的缩放和非线性不变。这些局部信念随后通过专家乘积形式进行融合,以获得团队的全局一致估计。为了进一步提高团队协调的效率,我们引入了一种信息区域分配和路径规划策略,该策略减少了冗余探索,同时平衡了探索与利用。我们使用高保真度的仿真和现实的气体传感器模型验证了所提出的框架。结果表明,我们的方法显著优于基于标准测量聚合的基准方法,尽管传感器异质性较强,仍能实现可靠的源定位精度。更广泛地说,这项工作展示了无校准传感器表示如何有效地扩展到分布式机器人系统,为其在涉及异质传感器的其他估计任务中的应用铺平了道路。
cs.RO / 17 / 2608.28246
Training-free Suction Grasp Detection for Deformed Aseptic Cartons Using Vision-Language Models and Geometric Surface Scoring
基于视觉-语言模型和几何表面评分的无训练吸取抓取检测用于变形无菌纸箱
Abstract
Robotic sorting of recyclable waste is challenging due to the deformable and geometrically inconsistent nature of target objects. We present a training-free suction grasping system for sorting deformed aseptic beverage cartons, decoupling target identification from grasp-point selection. An open-vocabulary vision-language model detects cartons from a text prompt, SAM2 refines each detection into an instance mask, and a geometric scoring method selects the suction point by combining surface flatness with normal alignment. Three geometric methods are compared: k-nearest-neighbour PCA, Sobel cross-product, and RANSAC plane fitting. Evaluated on a real robot across three deformation levels and 35 cluttered scenes, single-object grasp success reaches 88.2% and end-to-end retrieval in clutter is 72.6%.
Chinese Translation
由于目标物体的可变形和几何不一致特性,机器人可回收废物分类面临挑战。我们提出了一种无训练的吸取抓取系统,用于分类变形的无菌饮料纸箱,将目标识别与抓取点选择解耦。开放词汇的视觉-语言模型通过文本提示检测纸箱,SAM2将每个检测结果细化为实例掩膜,而几何评分方法通过结合表面平整度和法线对齐来选择吸取点。我们比较了三种几何方法:k近邻主成分分析(PCA)、索贝尔叉乘和RANSAC平面拟合。在三个变形级别和35个杂乱场景中对真实机器人进行评估,单对象抓取成功率达到88.2%,在杂乱环境中的端到端检索率为72.6%。
cs.RO / 18 / 2608.28266
CoCoBench: A Cooperative Coordination Benchmark for Embodied Multi-Agent Task Planning
CoCoBench:面向具身多智能体任务规划的协作协调基准
Abstract
Agent systems powered by multimodal large language models (MLLMs) have advanced rapidly in recent years, yet existing embodied-agent benchmarks still lack fine-grained diagnostics for multi-agent coordination. Most benchmarks either focus on single-agent task completion or summarize multi-agent behavior with overall task success rates, which can obscure coordination failures such as duplicated work, violations of ordering constraints, resource contention, and desynchronized handoffs. In this paper, we introduce CoCoBench, a construct-level benchmark for evaluating multi-agent embodied coordination in executable household tasks. CoCoBench contains 897 oracle-validated instances organized around four recurring coordination constructs: task allocation, sequential ordering, mutual exclusion, and handoff coordination. In addition to task success rate, CoCoBench provides construct-level scores that measure whether agents coordinate effectively. We evaluate 11 leading MLLMs across different coordination modes, observation inputs, and numbers of agents. The results show that coordination ability is highly construct-specific: strong overall performance does not imply balanced competence across different coordination types. These findings point to new directions for designing targeted model architectures and improving multi-agent coordination ability.
Chinese Translation
由多模态大型语言模型(MLLMs)驱动的智能体系统近年来发展迅速,然而现有的具身智能体基准仍缺乏对多智能体协调的细粒度诊断。大多数基准要么侧重于单智能体任务完成,要么仅通过整体任务成功率来总结多智能体行为,这可能掩盖诸如重复劳动、顺序约束违规、资源争用和不同步交接等协调失败问题。本文提出了CoCoBench,一种针对可执行家务任务中多智能体具身协调的构件级基准。CoCoBench包含897个经专家验证的实例,围绕四种常见协调构件组织:任务分配、顺序排序、互斥和交接协调。除任务成功率外,CoCoBench还提供构件级评分,用以衡量智能体的协调效果。我们评估了11种领先的MLLMs,涵盖不同的协调模式、观察输入和智能体数量。结果表明,协调能力高度依赖于具体构件:整体表现优异并不意味着在不同协调类型上均衡的能力。这些发现为设计针对性的模型架构和提升多智能体协调能力指明了新方向。
cs.RO / 19 / 2608.28270
Spatial-Semantic Reasoning using Large Language Models for Efficient UAV Search Operations
基于大型语言模型的空间语义推理用于高效无人机搜索操作
Abstract
We present a real-time semantic navigation framework for Unmanned Aerial Vehicles (UAVs) focused on improving time efficiency in the Object Goal Navigation (ObjectNav) task. Central to our approach is a Large Language Model (LLM) that interprets user-provided natural language instructions and performs semantic reasoning over detected objects and spatial context to prioritize high-probability search regions. The system combines real-time object detection, 3D spatial mapping, and polynomial spline interpolation for smooth and feasible UAV trajectory planning. Unlike prior methods that rely on offline reasoning or simulator-constrained action spaces, our framework can operate in real time, continuously updating semantic relevance based on new observations. Experiments in both simulated and real-world settings demonstrate reductions in mission duration while maintaining high search accuracy, underscoring the effectiveness of LLM-guided reasoning for time- efficient UAV-based ObjectNav.
Chinese Translation
我们提出了一种实时语义导航框架,旨在提高无人机(UAV)在目标导航(Object Goal Navigation,ObjectNav)任务中的时间效率。我们的方法核心是一个大型语言模型(Large Language Model,LLM),该模型能够解读用户提供的自然语言指令,并对检测到的对象和空间上下文进行语义推理,以优先考虑高概率搜索区域。该系统结合了实时对象检测、三维空间映射和多项式样条插值,以实现平滑且可行的无人机轨迹规划。与依赖离线推理或受限于模拟器的行动空间的先前方法不同,我们的框架能够实时运行,基于新的观察不断更新语义相关性。在模拟和现实环境中的实验表明,任务持续时间显著缩短,同时保持了高搜索精度,突显了基于LLM引导的推理在提高无人机目标导航时间效率方面的有效性。
cs.RO / 20 / 2608.28279
STEGNav: Spatio-Temporal Event Graph Reasoning for Multimodal Lifelong Object Navigation
STEGNav:用于多模态终身物体导航的时空事件图推理
Abstract
Multimodal lifelong navigation requires an agent to autonomously explore unseen environments while sequentially completing navigation tasks specified by object categories, language descriptions, or reference images. Existing methods primarily accomplish these tasks by constructing state-centric semantic scene graphs. By treating scene graphs as persistent repositories of semantic observations, these methods struggle to distinguish similar instances, jointly represent semantic targets and exploration frontiers, and effectively exploit navigation memory and trajectory experience. To address these limitations, we propose Spatio-Temporal Event Graph Navigation (STEGNav), a training-free framework that extends conventional scene graphs into spatio-temporal event graphs along complementary spatial and temporal axes. The spatial axis performs query-conditioned instance grounding and jointly represents semantic targets and occupancy-aware exploration frontiers characterized by reachability, path cost, and exploration utility. The temporal axis employs trajectory-aware dual-window memory to retain recent decision--trajectory events and verified cross-subtask navigation outcomes. A VLM-based navigation agent reasons over the resulting spatio-temporal event graph and selects either a target instance or an exploration frontier as its next navigation goal. STEGNav achieves 66.3% SR and 39.7 SPL on GOAT-Bench, as well as SR scores of 64.0% and 69.4% on HM3Dv1 and HM3Dv2, respectively. Ablation studies and error analyses validate the complementary effects of the two axes, demonstrating that event-driven spatio-temporal representations improve navigation reliability and cross-subtask experience reuse.
Chinese Translation
多模态终身导航要求智能体在自主探索未见环境的同时,顺序完成由物体类别、语言描述或参考图像指定的导航任务。现有方法主要通过构建以状态为中心的语义场景图来实现这些任务。由于将场景图视为语义观察的持久存储库,这些方法在区分相似实例、共同表示语义目标和探索前沿,以及有效利用导航记忆和轨迹经验方面存在困难。为了解决这些局限性,我们提出了时空事件图导航(STEGNav),这是一个无训练的框架,将传统场景图扩展为沿着互补的空间和时间轴的时空事件图。空间轴执行查询条件实例定位,并共同表示语义目标和以可达性、路径成本和探索效用为特征的占用感知探索前沿。时间轴采用轨迹感知的双窗口记忆,保留最近的决策-轨迹事件和验证的跨子任务导航结果。基于视觉语言模型(VLM)的导航智能体在生成的时空事件图上进行推理,并选择目标实例或探索前沿作为其下一个导航目标。STEGNav在GOAT-Bench上实现了66.3%的成功率(SR)和39.7%的成功路径长度(SPL),在HM3Dv1和HM3Dv2上分别获得了64.0%和69.4%的成功率。消融研究和错误分析验证了两个轴的互补效应,表明事件驱动的时空表示提高了导航的可靠性和跨子任务经验的重用。
cs.RO / 21 / 2608.28300
MaCoPlanner: LLM-Assisted Manual-Compiled Task Planning with Proactive Safety Verification for Robotic Industrial Panel Operation
MaCoPlanner:一种基于大型语言模型的手动编译任务规划框架,具备主动安全验证功能,适用于机器人工业面板操作
Abstract
Robotic industrial panel operation requires not only accurate control localization but also compliance with operating procedures, safety rules, and device-state constraints distributed across heterogeneous manuals. This study presents MaCoPlanner, a task-planning framework built on knowledge compiled from equipment manuals that converts equipment manuals into a typed intermediate representation, retrieves task- and state-relevant evidence, and uses it to support plan generation. Before actuation, candidate plans are symbolically rolled out and checked against procedural and state-transition constraints; detected violations are localized and returned for targeted repair, while unresolved plans are rejected. A separate execution interface grounds verified symbolic actions to physical controls and updates the device state. Under an independent evaluation oracle, MaCoPlanner achieves a final violation rate of 2.7%, and 26.3% of the runs in the repair analysis are rejected after exhausting the refinement budget. Compared with Raw-Manual, task success increases from 62.8% to 84.4% on Level-2 tasks and from 25.9% to 43.2% on Level-3 tasks. Experiments on a controller-panel simulator without an attached industrial load further demonstrate integrated execution feasibility under representative interaction conditions, without claiming industrial deployment readiness.
Chinese Translation
机器人工业面板操作不仅需要准确的控制定位,还需遵循操作程序、安全规则和分布在异构手册中的设备状态约束。本研究提出了MaCoPlanner,一种基于从设备手册中编译的知识构建的任务规划框架,该框架将设备手册转换为类型化的中间表示,检索与任务和状态相关的证据,并利用这些证据支持计划生成。在执行之前,候选计划会被符号化展开,并与程序和状态转移约束进行检查;检测到的违规行为会被定位并返回进行针对性修复,而未解决的计划则会被拒绝。一个独立的执行接口将经过验证的符号动作与物理控制相结合,并更新设备状态。在独立评估的情况下,MaCoPlanner实现了最终违规率为2.7%,在修复分析中,26.3%的运行在耗尽精炼预算后被拒绝。与Raw-Manual相比,Level-2任务的成功率从62.8%提高到84.4%,Level-3任务的成功率从25.9%提高到43.2%。在没有附加工业负载的控制器-面板模拟器上的实验进一步证明了在代表性交互条件下的集成执行可行性,但并未声称具备工业部署的准备性。
cs.RO / 22 / 2608.28305
PanelShield: Verifiable Closed-Loop Safe Planning for Robotic Industrial Panel Operation
PanelShield:可验证的闭环安全规划用于机器人工业面板操作
Abstract
Industrial panel operation is knowledge-intensive and safety-critical. Beyond control recognition and action generation, execution must satisfy constraints in operation manuals and safety regulations. While foundation-model-based planners show strong semantic capability, they typically lack computable, localizable, and reproducible mechanisms for violation detection and repair. To address this, we propose PanelShield, a verifiable closed-loop safety planning framework for manual-guided industrial panel operation. The framework generates parameterized action primitive sequences from task-relevant manual evidence and applies dual formal verification with LTL and a Safety FSM to enforce cross-step temporal correctness and local transition legality. When violations occur, it outputs a structured counterexample with the earliest violating step and cause, enabling targeted repair and re-verification. We build a multi-level long-horizon planning benchmark covering three representative industrial device panels, and evaluate the framework in simulation and real-world robotic experiments. Results show that PanelShield improves complex safety-constrained task performance over foundation-model-only planning baselines while reducing the violation rate to 2.7%, with 4.1 s total latency. Real-world experiments demonstrate end-toend feasibility. Overall, PanelShield offers a verifiable approach to robotic panel operation that balances flexibility, safety, and auditability.
Chinese Translation
工业面板操作是知识密集型且安全关键的。除了控制识别和动作生成外,执行还必须满足操作手册和安全法规中的约束。尽管基于基础模型的规划者展示了强大的语义能力,但它们通常缺乏可计算、可定位和可重复的违规检测与修复机制。为了解决这个问题,我们提出了PanelShield,一个用于手动引导工业面板操作的可验证闭环安全规划框架。该框架从任务相关的手动证据中生成参数化的动作原语序列,并应用LTL和安全有限状态机(Safety FSM)的双重形式验证,以强制执行跨步骤的时间正确性和局部转换合法性。当发生违规时,它输出一个结构化的反例,包含最早的违规步骤和原因,从而实现有针对性的修复和重新验证。我们构建了一个涵盖三个代表性工业设备面板的多层次长时间规划基准,并在仿真和真实世界的机器人实验中评估该框架。结果表明,PanelShield在复杂的安全约束任务性能上优于仅基于基础模型的规划基线,同时将违规率降低至2.7%,总延迟为4.1秒。真实世界的实验展示了端到端的可行性。总体而言,PanelShield提供了一种可验证的机器人面板操作方法,平衡了灵活性、安全性和可审计性。
cs.RO / 23 / 2608.28409
Cooperative Risk-Aware Exploration in Heterogeneous Multi-Robot Systems Using Algorithmic Altruism
基于算法利他主义的异构多机器人系统中的合作风险感知探索
Abstract
Multi-robot systems are well-positioned for exploration in hazardous environments, but effective deployment requires deciding not only where robots should gather information, but also how risk should be distributed across heterogeneous team members. This paper develops a game-theoretic framework for cooperative risk-aware exploration based on ecologically inspired altruistic behavior. Each robot selects a finite-horizon trajectory to maximize information gain while penalizing redundant exploration and expected hazard exposure. Heterogeneity is introduced through agent-specific value parameters for encoding altruistic coupling, which is modeled through relatedness weights inspired by Hamilton's rule. We introduce a game-theoretic structure for trajectory planning that defines a Social Nash Equilibrium, which modifies the utility of agent actions according to agent relatedness. This utility shaping causes agents to internalize the effect of their trajectory choices on teammates, encouraging lower-valued robots to accept risk when doing so benefits higher-valued agents and improves team performance. We define an exploration utility for agents that rewards area coverage and uncertainty reduction, while also penalizing redundancy and risk, enabling projected gradient-based waypoint optimization in a receding-horizon planner. Simulations show that altruistic planning reduces redundant exploration, improves inter-robot separation, and reallocates risk according to agent value while maintaining comparable map coverage. We further demonstrate the approach in hardware experiments, where planned waypoints are tracked by wheeled robots using single-integrator controllers and barrier certificates.
Chinese Translation
多机器人系统在危险环境中的探索中具有良好的适应性,但有效部署不仅需要决定机器人应在哪里收集信息,还需要考虑如何在异构团队成员之间分配风险。本文开发了一种基于生态启发的利他行为的合作风险感知探索的博弈论框架。每个机器人选择一个有限时间范围内的轨迹,以最大化信息增益,同时惩罚冗余探索和预期的危险暴露。通过特定于代理的价值参数引入异构性,以编码利他耦合,这通过受汉密尔顿法则启发的相关性权重进行建模。我们引入了一种轨迹规划的博弈论结构,定义了社会纳什均衡,该均衡根据代理的相关性调整代理行动的效用。这种效用塑造使代理内化其轨迹选择对队友的影响,鼓励低价值的机器人在这样做有利于高价值代理并提高团队表现时接受风险。我们为代理定义了一种探索效用,奖励区域覆盖和不确定性减少,同时惩罚冗余和风险,从而使得在递归时间规划器中能够进行基于投影梯度的路径点优化。仿真结果表明,利他规划减少了冗余探索,提高了机器人间的分离度,并根据代理价值重新分配风险,同时保持了可比的地图覆盖率。我们进一步在硬件实验中演示了该方法,其中计划的路径点由使用单积分控制器和障碍证书的轮式机器人跟踪。
cs.RO / 24 / 2608.28435
Linear Temporal Logic Translation via Human-Inspired Self-Constrained Reasoning for Robot Task Specification
通过人类启发的自我约束推理进行线性时序逻辑翻译以实现机器人任务规范
Abstract
Many robotic tasks are temporally extended and demand precise specifications of subgoals, constraints, and their temporal ordering. Yet human operators typically communicate such tasks in natural language, which is inherently ambiguous, underspecified, and context dependent. Translating human instructions into formal task specifications, such as Linear Temporal Logic (LTL), is therefore essential for verifiable and safe robotic execution. Existing LLM-based translators attempt to bridge this gap through open-ended reasoning or post-hoc constraint enforcement, but the former may violate domain constraints, whereas the latter can disrupt the reasoning needed for novel instructions. This paper proposes Self-Constrained Reasoning (SCR), a framework that mitigates this trade-off by internalizing structural knowledge into the model's decision-making process rather than imposing it as an external filter. By combining a structural constraint representation with a hierarchical decision-making formulation, SCR guides reasoning within a formally grounded space while preserving adaptability to unseen instructions. Experiments show that SCR improves both domain-constraint satisfaction and generalization, providing an effective and interpretable approach for translating human intent into verifiable specifications for robotic execution.
Chinese Translation
许多机器人任务是时间延续的,要求对子目标、约束及其时间顺序进行精确规范。然而,人类操作员通常通过自然语言来传达这些任务,而自然语言本质上是模糊的、未充分指定的,并且依赖于上下文。因此,将人类指令翻译为正式的任务规范,如线性时序逻辑(LTL),对于可验证和安全的机器人执行至关重要。现有的基于大型语言模型(LLM)的翻译器试图通过开放式推理或事后约束执行来弥合这一差距,但前者可能会违反领域约束,而后者则可能会干扰对新指令所需的推理。本文提出了自我约束推理(Self-Constrained Reasoning, SCR)框架,通过将结构知识内化到模型的决策过程中,而不是作为外部过滤器施加,从而缓解这一权衡。通过将结构约束表示与分层决策制定形式结合,SCR在形式化的基础空间内引导推理,同时保持对未见指令的适应性。实验表明,SCR在领域约束满足和泛化能力方面均有所提升,为将人类意图翻译为可验证的机器人执行规范提供了一种有效且可解释的方法。
cs.RO / 25 / 2608.28570
ChainSplat: A Physics-Inspired Screw-Theoretic Model for Learning Deformable Linear Object Dynamics from Multi-View RGB Videos
ChainSplat:一种基于物理的螺旋理论模型,用于从多视角RGB视频学习可变形线性物体的动力学
Abstract
Identifying the underlying dynamics and 3D geometry of deformable linear objects (DLOs), such as cables, ropes, and hoses, is essential for accurate robotic manipulation, but remains challenging due to their high-dimensional configuration spaces and diverse behaviors arising from varying material properties. Existing methods often rely on multi-stage pipelines and auxiliary depth inputs, which are prone to errors under dynamic interactions, while their high-dimensional state representations make model-based control computationally expensive. In this paper, we introduce ChainSplat, a physics-inspired framework that jointly learns the 3D geometry, appearance, kinematics, and dynamics of DLOs solely from multi-view RGB videos. ChainSplat represents a DLO as an open-chain structure of rigid links connected by revolute joints, yielding an analytic, screw-theoretic model with a compact state representation parameterized by joint configurations. By integrating this formulation with Gaussian splatting, ChainSplat jointly recovers DLO dynamics, kinematics-aware 3D geometry, and appearance, while enabling high-fidelity RGB rendering from arbitrary states. Through real-world experiments, we demonstrate that ChainSplat achieves state-of-the-art performance in dynamics predictions, 3D geometry reconstruction, and RGB rendering across dynamic interactions. ChainSplat further enables real-time state and force estimation, as well as accurate model-based trajectory optimization, highlighting its practical utility for real-world robotic manipulation of DLOs. Accompanying source code and video are available at: https://chainsplat.github.io.
Chinese Translation
识别可变形线性物体(DLOs)的潜在动力学和三维几何形状,例如电缆、绳索和软管,对于准确的机器人操作至关重要,但由于其高维配置空间和因材料属性变化而产生的多样化行为,这一任务仍然具有挑战性。现有方法通常依赖于多阶段管道和辅助深度输入,这在动态交互下容易出错,同时其高维状态表示使得基于模型的控制计算开销巨大。本文介绍了ChainSplat,一种基于物理的框架,能够仅通过多视角RGB视频共同学习DLO的三维几何形状、外观、运动学和动力学。ChainSplat将DLO表示为由旋转关节连接的刚性链节的开放链结构,产生一种解析的、基于螺旋理论的模型,其状态表示由关节配置参数化。通过将这一公式与高斯点云结合,ChainSplat共同恢复DLO的动力学、运动学感知的三维几何形状和外观,同时能够从任意状态实现高保真RGB渲染。通过实际实验,我们证明ChainSplat在动态预测、三维几何重建和动态交互中的RGB渲染方面达到了最先进的性能。ChainSplat进一步实现了实时状态和力的估计,以及准确的基于模型的轨迹优化,突显了其在现实世界中对DLO的机器人操作的实用性。相关源代码和视频可在:https://chainsplat.github.io获取。
cs.RO / 26 / 2608.28578
Aero Hand Open: A Simulation-Ready Tendon-Driven Hand for Dexterous Manipulation Learning
Aero Hand Open:一种适用于灵巧操作学习的仿真准备型腱驱动手
Abstract
Tendon-driven hands are anthropomorphic, and moving the actuators off the joints is what makes a hand of this capability affordable to build. Two effects produce that saving. Routing force through a cable removes the requirement that a motor fit inside the joint it drives, so smaller and cheaper motors suffice, and one motor can drive several joints through a single cable, so fewer motors are needed. They are also harder to learn on than a direct-drive hand. The underactuated transmission that produces the saving is itself difficult to represent in a simulator, and the joints one cable drives are not independently commandable. We present Aero Hand Open, a tendon-driven anthropomorphic hand that is released simulation-ready. Three things ship with it. A simulation model reproduces the cable transmission itself. An identified actuation map connects that model to the motor commands in both directions, including the three-way coupling of the thumb. A reinforcement learning package trains policies for the hand. Together they let a policy be trained entirely in simulation and run on the hand with no fine-tuning and no state estimation. We release the mechanical design, the simulation model, the identified mapping, the training environment and the deployment stack.
Chinese Translation
腱驱动手具有人形特征,将执行器移出关节是使这种能力的手具备可负担性的重要因素。这种节省源于两个效应。通过电缆传递力量消除了电机必须安装在其驱动的关节内的要求,因此可以使用更小且更便宜的电机,并且一个电机可以通过一根电缆驱动多个关节,从而减少所需电机的数量。然而,与直接驱动手相比,腱驱动手的学习难度更大。产生节省的欠驱动传动系统本身在仿真中难以表示,并且一根电缆驱动的关节无法独立控制。我们提出了Aero Hand Open,这是一种腱驱动的人形手,已准备好进行仿真。随之发布的有三个部分:一个仿真模型再现了电缆传动系统;一个已识别的驱动映射将该模型与电机指令双向连接,包括拇指的三向耦合;一个强化学习包用于训练手的策略。它们共同使得策略可以完全在仿真中训练,并在手上运行,无需微调和状态估计。我们发布了机械设计、仿真模型、已识别的映射、训练环境和部署堆栈。
cs.CV / 1 / 2608.27527
FVeinSyn: Synthetic Finger Vein Image Generator
FVeinSyn:合成指静脉图像生成器
Abstract
A major challenge in finger vein recognition is the lack of large-scale public datasets. Existing datasets contain few identities and limited samples per finger, restricting the advancement of deep learning-based methods. To address this, we propose FVeinSyn, a large-scale controllable synthetic data generation framework for finger vein. It explicitly decouples synthesis of vascular topology and imaging appearance to mitigate the limitations caused by insufficient training samples, such as inadequate identity diversity and restricted realism. Specifically: first, a finger vein identity generator models vascular topology under physiological and geometric constraints using stochastic L-systems, producing anatomically valid and identity-distinctive vascular patterns. Then, a cascaded region-aware GAN renders the topological maps into realistic near-infrared images. Finally, an intra-class diversity generator introduces geometric and optical perturbations to simulate realistic intra-class variations. Using FVeinSyn, we generated 500,000 images (10,000 vein identities, 50 samples per identity) and conducted extensive evaluations. Results show that FVeinSyn holds significant advantages in realism, identity diversity, vascular pattern consistency, and intra-class diversity. Models trained with FVeinSyn outperform real-data-only baselines a cross eight public datasets, achieving an average accuracy improvement of 27.43\%. The code is available at: https://github.com/EvanWang98/Synthetic-Finger-Vein-Generator.
Chinese Translation
指静脉识别的一个主要挑战是缺乏大规模的公共数据集。现有数据集包含的身份数量较少,每个手指的样本也有限,限制了基于深度学习方法的进展。为了解决这一问题,我们提出了FVeinSyn,一个用于指静脉的大规模可控合成数据生成框架。该框架明确地将血管拓扑的合成与成像外观解耦,以减轻由于训练样本不足而导致的限制,例如身份多样性不足和现实感受限。具体而言:首先,指静脉身份生成器在生理和几何约束下使用随机L系统建模血管拓扑,生成解剖学上有效且身份独特的血管模式。然后,级联区域感知生成对抗网络(GAN)将拓扑图渲染为逼真的近红外图像。最后,类内多样性生成器引入几何和光学扰动,以模拟现实的类内变化。使用FVeinSyn,我们生成了500,000张图像(10,000个静脉身份,每个身份50个样本),并进行了广泛的评估。结果表明,FVeinSyn在现实感、身份多样性、血管模式一致性和类内多样性方面具有显著优势。使用FVeinSyn训练的模型在八个公共数据集上的表现优于仅使用真实数据的基线,平均准确率提高了27.43%。代码可在以下链接获取:https://github.com/EvanWang98/Synthetic-Finger-Vein-Generator。
cs.CV / 2 / 2608.27529
Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction
重新审视长时间流媒体3D重建中的局部上下文
Abstract
Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online under bounded memory and computation. Early streaming models achieve causal, bounded-cost inference using finite context buffers or compact recurrent states, yet their estimates often deteriorate as sequences grow. Recent methods improve long-horizon stability by coupling short-range context with persistent or multi-level long-range memory. We pursue a different route: we keep the learned temporal state strictly local and formulate predictions whose targets remain independent of sequence length. We present ABot-Recon, a simple streaming model that caches KV features from only the preceding 11 frames. It predicts a point map in the current camera coordinate system together with an adjacent-frame relative pose. These predictions remain equivariant under changes of reference frame, and global poses and geometry are recovered through sequential composition. To reduce accumulated drift, a lightweight temporal refiner improves relative rotations using recent visual and motion context, while a composition-aware pose loss supervises multi-step pose composition. Extensive evaluations on challenging long-sequence benchmarks demonstrate the superior long-horizon performance of our local-context approach. On Oxford Spires, ABot-Recon achieves an ATE of 4.35 m and an RPE-R of $0.12^\circ$, reducing both errors by approximately 40\% relative to the best prior results.
Chinese Translation
从极长视频中进行流媒体3D重建需要在有限的内存和计算条件下在线估计相机运动和场景几何。早期的流媒体模型通过使用有限的上下文缓冲区或紧凑的递归状态实现因果的、有限成本的推断,但随着序列的增长,它们的估计往往会恶化。最近的方法通过将短期上下文与持久或多层次的长期记忆结合,改善了长时间稳定性。我们采取了不同的路线:我们将学习到的时间状态严格保持局部,并制定目标与序列长度无关的预测。我们提出了ABot-Recon,这是一种简单的流媒体模型,仅缓存前11帧的KV特征。它在当前相机坐标系中预测点图,同时给出相邻帧的相对姿态。这些预测在参考框架变化下保持等变性,全球姿态和几何通过序列组合恢复。为了减少累积漂移,一个轻量级的时间精炼器利用最近的视觉和运动上下文改善相对旋转,而一个考虑组合的姿态损失则监督多步姿态组合。在具有挑战性的长序列基准上的广泛评估表明,我们的局部上下文方法在长时间范围内表现优越。在牛津尖塔数据集上,ABot-Recon实现了4.35米的平均绝对误差(ATE)和$0.12^ heta$的相对姿态误差(RPE-R),相较于最佳先前结果,两个误差均减少了约40%。
cs.CV / 3 / 2608.27549
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
代码作为世界:可执行世界表征的自主发现用于物理推理
Abstract
Physical understanding and reasoning depend on forming compact and generalizable representations of the world. While modern vision-language models can recognize and explain diverse physical events, they often lack explicit representations of the underlying mechanisms-such as object states, physical parameters, and governing dynamics-needed for reliably reasoning how the world evolves and responds to interventions. In this work, we introduce Code-as-World, a paradigm that represents physical worlds through executable world representations. By expressing physical composition, dynamic evolution, and visual appearance as executable code, Code-as-World provides a compact, quantitatively grounded, and controllable abstraction of the physical world. To construct such representations from multimodal observations, such as natural-language descriptions or real-world videos, we develop an agentic discovery loop inspired by abductive reasoning, where an agent proposes, executes, renders, verifies, and iteratively refines executable world hypotheses. As a concrete application, we use verified executable worlds to provide scalable physical supervision for training vision-language models on quantitative physical reasoning. Experiments show that Code-as-World-VL achieves state-of-the-art performance on QuantiPhy and surpasses leading proprietary models, highlighting the potential of executable world representations as a scalable foundation for physical intelligence.
Chinese Translation
物理理解和推理依赖于形成紧凑且可推广的世界表征。尽管现代视觉-语言模型能够识别和解释多样的物理事件,但它们通常缺乏对基础机制的明确表征,例如物体状态、物理参数和支配动态,这些都是可靠推理世界如何演变和对干预作出反应所必需的。在本研究中,我们提出了代码作为世界(Code-as-World)这一范式,通过可执行的世界表征来表示物理世界。通过将物理组成、动态演变和视觉外观表达为可执行代码,代码作为世界提供了一种紧凑、定量基础且可控的物理世界抽象。为了从多模态观察(如自然语言描述或现实世界视频)中构建这样的表征,我们开发了一种受归纳推理启发的自主发现循环,其中一个代理提出、执行、渲染、验证并迭代优化可执行的世界假设。作为具体应用,我们利用经过验证的可执行世界为训练视觉-语言模型提供可扩展的物理监督,专注于定量物理推理。实验表明,代码作为世界-视觉语言(Code-as-World-VL)在QuantiPhy上达到了最先进的性能,并超越了领先的专有模型,突显了可执行世界表征作为物理智能可扩展基础的潜力。
cs.CV / 4 / 2608.27562
VidParse: Online Parsing of Egocentric Procedures Like a Pro
VidParse:像专业人士一样在线解析自我中心程序
Abstract
Translating continuous, noisy egocentric video streams into discrete, temporally ordered action steps is fraught with visual challenges. Heavy ego-motion, transient occlusions, and the high intra-class variability of unscripted human-object interactions cause standard frame-level online temporal models to struggle, often resulting in severe over-segmentation and structural collapse. To bridge the gap between unstable low-level perception and high-level procedural logic, we present VidParse, an online, training-free framework that treats activity understanding as a graph-constrained inference problem. Rather than relying on learned temporal filters, we dynamically identify semantic transitions using a temporal similarity matrix over manipulation-anchored features, which are extracted from frozen foundation models to prioritize foreground hand-object interactions. A beam search decoder then leverages an induced procedural task graph to explicitly enforce valid action transitions and prune impossible trajectories. By anchoring robust visual segments to hard procedural constraints, our approach preserves long-range state transitions and achieves up to a 10x improvement in complex multi-step parsing accuracy over strong online baselines, all without requiring a single gradient update.
Chinese Translation
将连续、嘈杂的自我中心视频流转换为离散的、时间顺序排列的动作步骤面临着视觉挑战。剧烈的自我运动、瞬时遮挡以及非脚本化人类与物体交互的高类内变异性使得标准的帧级在线时间模型难以应对,常常导致严重的过度分割和结构崩溃。为了弥合不稳定的低级感知与高级程序逻辑之间的差距,我们提出了VidParse,这是一种在线、无训练的框架,将活动理解视为一个图约束推理问题。我们并不依赖于学习的时间滤波器,而是通过对基于操作的特征使用时间相似性矩阵动态识别语义转换,这些特征是从冻结的基础模型中提取的,以优先考虑前景手-物体交互。然后,束搜索解码器利用诱导的程序任务图显式强制有效的动作转换并修剪不可能的轨迹。通过将稳健的视觉片段锚定到严格的程序约束,我们的方法保持了长距离状态转换,并在复杂的多步骤解析准确性上实现了高达10倍的提升,相较于强大的在线基线,且无需进行单次梯度更新。
cs.CV / 5 / 2608.27584
Quanta Perception as Probabilistic Events
量子感知作为概率事件
Abstract
Autonomous systems rely on extracting information from light, yet remain brittle in extreme environments, from nighttime navigation to high-speed robotics. Conventional sensors aggregate photons over fixed exposures, imposing trade-offs between sensitivity, dynamic range, and temporal resolution that degrade perception when photons are scarce or dynamics are rapid. Quanta sensors detect individual photons, but their streams exceed real-time compute and latency budgets by orders of magnitude. Here we introduce $\textit{probabilistic events}$, a computational primitive for real-time quanta perception from individual photon detections. By computing the posterior over the time since the last intensity change, we represent photon streams as recursive belief states. Rather than fixed-threshold event-camera triggers, this recursive Bayesian formulation yields three low-latency signals: motion-adaptive scene flux, high-fidelity activity maps, and entropy-based perceptual uncertainty. This representation enables perception in extreme conditions, including pose estimation of a running person at $\sim$0.05 lux---without retraining vision models. Our approach processes input streams exceeding 50{,}000 quanta frames per second on commodity GPU hardware---yielding kilohertz-scale outputs up to four orders of magnitude faster than state-of-the-art quanta reconstruction baselines, even for megapixel arrays. By replacing frame reconstruction with direct probabilistic inference over photon streams, this work bridges photon-counting quanta sensing with robotic vision.
Chinese Translation
自主系统依赖于从光中提取信息,但在极端环境下仍然显得脆弱,从夜间导航到高速机器人。传统传感器在固定曝光下聚合光子,导致在光子稀缺或动态快速时,灵敏度、动态范围和时间分辨率之间存在权衡,从而降低感知能力。量子传感器能够检测单个光子,但其数据流在数量级上超出了实时计算和延迟预算。在此,我们引入了 extit{概率事件},一种用于从单个光子检测中实现实时量子感知的计算原语。通过计算自上次强度变化以来的后验分布,我们将光子流表示为递归信念状态。与固定阈值事件摄像机触发器不同,这种递归贝叶斯形式化产生了三种低延迟信号:运动自适应场景通量、高保真活动图和基于熵的感知不确定性。这种表示方法使得在极端条件下的感知成为可能,包括在约0.05 lux的光照条件下对奔跑者的姿态估计——无需重新训练视觉模型。我们的方法在普通GPU硬件上处理超过50,000帧每秒的输入流——产生的千赫级输出速度比最先进的量子重建基线快四个数量级,即使对于百万像素阵列也是如此。通过用直接概率推断替代帧重建,这项工作将光子计数量子传感与机器人视觉连接起来。
cs.CV / 6 / 2608.27610
ShiftSplit-AD: Separating Domain Shift from Defects in Foundation-Feature Visual Anomaly Detection
ShiftSplit-AD:在基础特征视觉异常检测中将领域偏移与缺陷分离
Abstract
Visual anomaly detectors based on frozen foundation-model features commonly score distances from test patches to a memory of normal features. Benign acquisition changes can also enlarge these distances, confounding domain variation with defects. We investigate whether structured decomposition of nearest-normal DINOv2 residuals can suppress shift-induced evidence while retaining unseen defects. ShiftSplit-AD decomposes the patch residual matrix into low-rank and row-sparse components and scores the sparse component, with an optional low-rank/sparse fusion. The experiments expose a central trade-off rather than a universal separation: genuine defects can contain correlated, low-dimensional structure, so filtering broad residual activity may also remove defect information. On AeBAD-S, using settings fixed after Bottle development, sparse-only scoring improves image AUROC from 0.6780 to 0.7294 and AUPRC from 0.8052 to 0.8465. Paired bootstrap 95% intervals for the improvements are [0.0238, 0.0808] and [0.0170, 0.0650], respectively. However, sparse-only scoring reduces mean clean AUROC from 0.9890 to 0.9133 on four held-out MVTec categories and degrades Bottle localization. These findings show that residual decomposition can help when domain shift strongly contaminates anomaly evidence, but preserving defect structure remains the limiting problem.
Chinese Translation
基于冻结基础模型特征的视觉异常检测器通常通过计算测试补丁与正常特征记忆之间的距离来评分。良性的获取变化也可能扩大这些距离,从而将领域变化与缺陷混淆。我们研究了最近邻正常 DINOv2 残差的结构化分解是否可以抑制由偏移引起的证据,同时保留未见缺陷。ShiftSplit-AD 将补丁残差矩阵分解为低秩和行稀疏成分,并对稀疏成分进行评分,同时提供低秩/稀疏融合的选项。实验揭示了一个中心权衡,而不是普遍的分离:真实缺陷可能包含相关的低维结构,因此过滤广泛的残差活动也可能会移除缺陷信息。在 AeBAD-S 上,使用在 Bottle 开发后固定的设置,仅稀疏评分将图像的 AUROC 从 0.6780 提高到 0.7294,AUPRC 从 0.8052 提高到 0.8465。改进的配对自助法 95% 置信区间分别为 [0.0238, 0.0808] 和 [0.0170, 0.0650]。然而,仅稀疏评分将四个保留的 MVTec 类别的平均干净 AUROC 从 0.9890 降低到 0.9133,并降低了 Bottle 的定位性能。这些发现表明,当领域偏移强烈污染异常证据时,残差分解可以提供帮助,但保留缺陷结构仍然是一个限制性问题。
cs.CV / 7 / 2608.27633
Depth-Aware Pothole Detection Using YOLO and RT-DETR at the Edge
基于YOLO和RT-DETR的边缘深度感知坑洞检测
Abstract
Pothole detection and its severity measurement is still an important challenges in urban infrastructure management, where late maintenance directly contributes to vehicle damage, road accidents, and escalating repair costs. Existing automated approaches depend on 2D RGB images and cannot measure physical depth of potholes. In this paper, we present a depthaware pothole detection framework and then compare five architectures: YOLOv8n, YOLOv8nSeg, YOLOv9t, RTDETRL, and RTDETRX for RGB-D sensor fusion-based detection and automated depth measurement. A custom offline augmentation pipeline is used here to simulate adverse road monitoring conditions. All models are trained on the PothRGBD dataset with an 80% training and 20% validation split and evaluated using Precision, Recall, mAP@50, and mAP@50_95. Before measuring the depth data, all depth maps are corrected for camera tilt using RANSAC ground-plane orthorectification and all zero-valued sensor pixels are cast to NaN before any statistic is computed. YOLOv8nSeg achieves the highest mAP@50 of 0.9556 and mAP@50_95 of 0.6758 with the most accurate depth estimate of 2.96 cm with the pixel-precise Dseg algorithm. YOLOv8n achieves the fastest inference at 3.6ms. RTDETRX achieves the highest detection confidence at 92.70%. An important finding is that even after full RANSAC orthorectification, bounding box models overestimate pothole depth by 0.16 to 0.21 cm compared to pixel precise segmentation masks. This confirms that the pavement inclusion bias is structural rather than a calibration artifact.
Chinese Translation
坑洞检测及其严重性测量仍然是城市基础设施管理中的重要挑战,延迟维护直接导致车辆损坏、道路事故以及不断上升的修复成本。现有的自动化方法依赖于2D RGB图像,无法测量坑洞的物理深度。本文提出了一种深度感知的坑洞检测框架,并比较了五种架构:YOLOv8n、YOLOv8nSeg、YOLOv9t、RT-DETRL和RTDETRX,基于RGB-D传感器融合进行检测和自动深度测量。这里使用自定义的离线增强管道来模拟不利的道路监测条件。所有模型在PothRGBD数据集上进行训练,采用80%的训练集和20%的验证集,并使用精确度、召回率、mAP@50和mAP@50_95进行评估。在测量深度数据之前,所有深度图通过RANSAC地面平面正射校正来校正相机倾斜,所有零值传感器像素在计算任何统计数据之前被转换为NaN。YOLOv8nSeg以0.9556的最高mAP@50和0.6758的mAP@50_95,以及最准确的深度估计2.96厘米(使用像素精确的Dseg算法)取得了最佳表现。YOLOv8n以3.6毫秒的推理时间实现了最快的推理速度。RTDETRX则以92.70%的最高检测置信度取得了最佳结果。一个重要发现是,即使在完全的RANSAC正射校正后,边界框模型相比于像素精确的分割掩膜仍高估了0.16到0.21厘米的坑洞深度。这确认了路面包含偏差是结构性的问题,而非校准伪影。
cs.CV / 8 / 2608.27668
Report Supervision
报告监督
Abstract
Segmentation models can surpass radiologists, classification models, and vision-language models in tumor detection. Importantly, segmentation models outline tumors, allowing radiologists to better verify and trust the AI output. Their main limitation is the scarcity of tumor masks: creating one 3D tumor mask takes up to 30 minutes, so most public CT datasets contain only a few hundred masks, and even the largest private datasets contain only a couple of thousand. Tumor masks are not produced in clinical routine, but radiology reports are. Public datasets contain tens of thousands of CT-Report pairs, and hospitals contain hundreds of thousands. These reports describe tumors in detail, providing large-scale, informative training data. Here, we introduce Report Supervision (R-Super), a training framework that uses reports to directly supervise and improve tumor segmentation. R-Super introduces new loss functions that teach segmentation models to segment tumors that match report descriptions of tumor count, sizes, and locations. Reports are only used for training. We evaluated R-Super on kidney and pancreatic tumor segmentation, exploring diverse training data sizes, up to 41,418 CT-Report plus 3,488 pancreatic tumor CT-Mask pairs. On external validation, R-Super increased tumor detection F1-Score and segmentation DSC by up to +15% with respect to mask-only training. It also surpassed alternative methods such as CLIP and multi-task learning. Leveraging numerous readily available reports to supplement scarce masks, R-Super strongly improves AI performance when very few training masks are available (e.g., 50), and when many masks are available (e.g., 3,488), unlocking scale in tumor segmentation.
Chinese Translation
分割模型在肿瘤检测方面可以超越放射科医师、分类模型和视觉-语言模型。重要的是,分割模型能够勾勒出肿瘤的轮廓,使放射科医师能够更好地验证和信任人工智能的输出。它们的主要限制在于肿瘤掩膜的稀缺:创建一个三维肿瘤掩膜需要长达30分钟,因此大多数公共CT数据集中仅包含几百个掩膜,即使是最大的私人数据集也仅包含几千个。肿瘤掩膜并不是在临床常规中产生的,但放射学报告是。公共数据集包含数万个CT-报告对,医院中则有数十万个。这些报告详细描述了肿瘤,提供了大规模、信息丰富的训练数据。在此,我们介绍了报告监督(Report Supervision, R-Super),这是一种利用报告直接监督和改善肿瘤分割的训练框架。R-Super引入了新的损失函数,教会分割模型分割与报告描述的肿瘤数量、大小和位置相匹配的肿瘤。报告仅用于训练。我们在肾脏和胰腺肿瘤分割上评估了R-Super,探索了多种训练数据规模,最多达到41,418个CT-报告对和3,488个胰腺肿瘤CT-掩膜对。在外部验证中,R-Super将肿瘤检测的F1分数和分割的DSC提高了最多15%,相较于仅使用掩膜的训练。它还超越了CLIP和多任务学习等替代方法。R-Super利用大量现成的报告来补充稀缺的掩膜,在可用的训练掩膜非常少(例如50个)时,显著提高了人工智能的性能,而在可用掩膜数量较多(例如3,488个)时,也释放了肿瘤分割的规模潜力。
cs.CV / 9 / 2608.27735
ABCD: Alpha-Composited Block Coordinate Descent: Constant-VRAM Training for Large Radiance Fields
ABCD:Alpha复合块坐标下降:大辐射场的常量显存训练
Abstract
We present ABCD (Alpha-Composited Block Coordinate Descent), an out-of-core training framework for alpha-composited radiance fields, instantiated here for 3D Gaussian Splatting. Our method reformulates training as block coordinate descent over spatial partitions: only one block of parameters is active at a time, while all others are frozen. By exploiting the associativity of alpha blending, these inactive regions can be pre-rendered and collapsed into foreground and background RGBA images. As a result, for fixed partition size and image resolution, peak VRAM becomes O(1) with respect to total scene extent, rather than growing with full scene size. This enables GPUs with limited memory to train scenes that would otherwise not fit in core. In experiments, our method closely preserves the reconstruction quality of 3DGS, with less than 5% PSNR degradation, while ABCD with compositing ablated suffers roughly 40% degradation. Our code can be found at https://github.com/shiukaheng/abcd
Chinese Translation
我们提出了ABCD(Alpha复合块坐标下降),这是一个用于alpha复合辐射场的外存训练框架,此处针对3D高斯点云(3D Gaussian Splatting)进行了实例化。我们的方法将训练重新表述为在空间分区上的块坐标下降:每次仅激活一个参数块,而所有其他块保持冻结。通过利用alpha混合的结合性,这些非激活区域可以被预渲染并合并成前景和背景的RGBA图像。因此,对于固定的分区大小和图像分辨率,峰值显存(VRAM)与总场景范围的关系为O(1),而不是随着完整场景大小的增长而增加。这使得显存有限的GPU能够训练那些在内存中无法容纳的场景。在实验中,我们的方法在重建质量上与3DGS(3D Gaussian Splatting)保持了密切的一致性,PSNR降级小于5%,而去掉复合的ABCD则遭受了大约40%的降级。我们的代码可以在https://github.com/shiukaheng/abcd找到。
cs.CV / 10 / 2608.27753
What Can Low Resource Languages Learn From Each Other?
低资源语言可以相互学习什么?
Abstract
Despite the rapid advancement of Vision-Language Models (VLMs), their linguistic reach remains largely confined to high-resource languages, leaving the majority of the world's 7,000+ living languages on the wrong side of a growing digital divide. This disparity is especially pronounced in Optical Character Recognition (OCR), where low-resource scripts lack the massive datasets required for traditional scaling laws. We investigate OCR adaptation in extreme data-scarce regimes (<10K real and <250K synthetic images), demonstrating that conventional fine-tuning strategies often reach a performance ceiling. Our key finding reveals a structural inefficiency in language-specific adaptation: while higher layers of specialized models diverge to capture unique script nuances, the lower layers learn redundant, highly similar features. Motivated by this observation, we propose PSMC (Pre-train, Specialize, Merge, and Co-train), a data-efficient framework that capitalizes on a cross-script "transfer effect". Our approach first derives language-specific experts from a high-resource base model, then employs task arithmetic to fuse these experts into a unified, high-performance multilingual back- bone. Extensive evaluation across 10 Indian scripts (supporting 20+ languages) shows that PSMC achieves a ~2% average improvement in Word Recognition Rate (WRR) over individual specialist models without increasing parameter count. Our results indicate that joint training in the merged latent space facilitates a constructive knowledge transfer that benefits all constituent scripts, providing a scalable pathway for inclusive VLM development. Source code and datasets will be released post publication.
Chinese Translation
尽管视觉语言模型(Vision-Language Models, VLMs)迅速发展,但它们的语言覆盖范围仍然主要局限于高资源语言,这使得全球7000多种活语言中的大多数处于日益扩大的数字鸿沟的另一侧。这种差异在光学字符识别(Optical Character Recognition, OCR)中尤为明显,低资源脚本缺乏传统扩展法则所需的大规模数据集。我们研究了在极端数据稀缺环境下(<10K真实图像和<250K合成图像)的OCR适应性,证明传统的微调策略通常会达到性能上限。我们的关键发现揭示了语言特定适应中的结构性低效:尽管专门模型的高层次分化以捕捉独特的脚本细微差别,但低层次学习到的特征却是冗余且高度相似的。基于这一观察,我们提出了PSMC(预训练、专业化、合并和共同训练,Pre-train, Specialize, Merge, and Co-train),这是一个数据高效的框架,利用跨脚本的“迁移效应”。我们的方法首先从高资源基础模型中推导出语言特定专家,然后通过任务算术将这些专家融合成一个统一的高性能多语言骨干网络。在对10种印度脚本(支持20多种语言)的广泛评估中,PSMC在单个专业模型的基础上实现了约2%的平均词识别率(Word Recognition Rate, WRR)提升,而没有增加参数数量。我们的结果表明,在合并的潜在空间中进行联合训练促进了建设性的知识转移,惠及所有组成脚本,为包容性VLM开发提供了可扩展的路径。源代码和数据集将在发表后发布。
cs.CV / 11 / 2608.27795
uScenes: A Multimodal RGB and 3D Sonar Dataset for Underwater Robot Perception
uScenes:用于水下机器人感知的多模态RGB和3D声纳数据集
Abstract
Robust perception is essential for the deployment of autonomous underwater robots. However, optical cameras become unreliable under poor illumination and backscatter. Forward looking (2D) acoustic sensors remain effective under these conditions, but they measure range and bearing while leaving elevation unresolved, creating an ambiguity that prevents individual sonar returns from being localized in three dimensional (3D) space. This complicates the sensor use for 3D scene understanding and precise object detection. We introduce \textbf{uScenes}, a multimodal underwater dataset containing synchronized 3D multibeam sonar point clouds and RGB imagery. The dataset contains 110 scenes and 95,834 synchronized observation, representing 277.6 minutes of data collected across multiple field sessions. uScenes establishes a foundation for underwater sensor fusion, cross modal representation learning and 3D scene understanding. Code and datasets are given at https://github.com/era-research-lab/uScenes.
Chinese Translation
稳健的感知对于自主水下机器人的部署至关重要。然而,在光照不足和回波散射的情况下,光学相机的可靠性降低。前视(2D)声学传感器在这些条件下仍然有效,但它们仅测量距离和方位,无法解决高度问题,从而产生模糊性,阻碍了单个声纳回波在三维(3D)空间中的定位。这使得传感器在3D场景理解和精确物体检测中的使用变得复杂。我们介绍了 extbf{uScenes},一个包含同步的3D多波束声纳点云和RGB图像的多模态水下数据集。该数据集包含110个场景和95,834个同步观测,代表了在多个实地会话中收集的277.6分钟数据。uScenes为水下传感器融合、跨模态表示学习和3D场景理解奠定了基础。代码和数据集可在https://github.com/era-research-lab/uScenes获取。
cs.CV / 12 / 2608.27860
From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation
从透视到鱼眼深度估计与开放词汇分割
Abstract
Vision foundation models are capable of generalizing across 3-dimensional (3D) scenes with high-fidelity estimates; their empirical success can be attributed to training on large-scale datasets of perspective images. However, when transferred to wide field-of-view (FoV) images, such as those captured by fisheye cameras, they return erroneous outputs due to a covariate shift stemming from the radial distortion on the image pixels. We propose a method to generalize vision foundation models to fisheye cameras. The crux of our method lies in a set of learnable parameters, termed Distortion Extenders (DEX), that model the fisheye distortion coefficients and the distributional shift between fisheye and perspective images encoded in the latent space. By minimizing a self-supervised alignment loss, DEX transforms the latent embeddings of fisheye images to resemble those of perspective images to recover high-fidelity estimates. DEX is architecture- and task-agnostic: We demonstrate DEX on monocular depth estimation and open-vocabulary segmentation for convolution- and Transformer-based architectures, where we consistently improve over baselines across indoor and outdoor fisheye datasets. As a byproduct, the activations of DEX can also be decoded to distortion coefficients to support camera calibration. Code available at: https://github.com/Suchisrit/DEX.
Chinese Translation
视觉基础模型能够在三维(3D)场景中进行高保真度的泛化估计;其经验成功可归因于在大规模透视图像数据集上的训练。然而,当这些模型转移到广视场(FoV)图像时,例如鱼眼相机捕获的图像,由于图像像素的径向失真引起的协变量转移,它们会返回错误的输出。我们提出了一种将视觉基础模型推广到鱼眼相机的方法。我们方法的关键在于一组可学习的参数,称为失真扩展器(Distortion Extenders, DEX),它们建模鱼眼失真系数以及在潜在空间中编码的鱼眼图像与透视图像之间的分布转移。通过最小化自监督对齐损失,DEX将鱼眼图像的潜在嵌入转换为类似透视图像的形式,以恢复高保真度的估计。DEX与架构和任务无关:我们在基于卷积和变换器的架构上展示了DEX在单目深度估计和开放词汇分割中的应用,在室内和室外鱼眼数据集上,我们始终在基准测试中取得了改进。作为附带成果,DEX的激活也可以解码为失真系数,以支持相机校准。代码可在:https://github.com/Suchisrit/DEX.
cs.CV / 13 / 2608.27866
Iron: Intent-Aligned and Retrospective Dual Learning Framework for Enhancing Generalist Virtual Agents
Iron:意图对齐与回顾性双重学习框架以增强通用虚拟代理
Abstract
Achieving virtual agents capable of automating tasks across diverse digital environments remains a pivotal challenge in Embodied AI. While Multimodal Large Language Models (MLLMs) offer enhanced visual perception and reasoning, their agentic deployment faces three challenges: costly data annotation, imprecise action-intent alignment, and inefficient exploration from discarded failed trajectories. To address these, we introduce Iron, an intent-aligned, self-improved, and annotation-efficient framework for training GUI agents. Iron employs a novel dual learning strategy that utilizes a stepwise cycle-consistent (SCC) reward to achieve fine-grained alignment between low-level actions and high-level intents, thereby improving instruction grounding and intent understanding. Concurrently, Iron introduces a hindsight reproduction mechanism to repurpose failed trajectories for training, improving both learning efficiency and task diversity. Extensive experiments demonstrate that Iron-trained generalist agents consistently improve performance on cross-environment and cross-device tasks, outperforming models trained with three times more data. Iron also achieves a substantial 25.06% relative improvement on unseen web tasks, with further gains observed on inherently complex tasks, demonstrating the feasibility of building more capable virtual agents.
Chinese Translation
在具身人工智能领域,实现能够在多样化数字环境中自动化任务的虚拟代理仍然是一个关键挑战。尽管多模态大型语言模型(MLLMs)提供了增强的视觉感知和推理能力,但其代理部署面临三个挑战:昂贵的数据标注、不精确的行动意图对齐以及从被丢弃的失败轨迹中进行低效探索。为了解决这些问题,我们提出了Iron,一个意图对齐、自我改进和标注高效的图形用户界面(GUI)代理训练框架。Iron采用了一种新颖的双重学习策略,利用逐步循环一致性(SCC)奖励实现低级动作与高级意图之间的细粒度对齐,从而改善指令基础和意图理解。同时,Iron引入了一种事后重现机制,以重新利用失败轨迹进行训练,提高学习效率和任务多样性。大量实验表明,经过Iron训练的通用代理在跨环境和跨设备任务上表现持续提升,超越了使用三倍数据训练的模型。Iron在未见过的网络任务上也实现了25.06%的显著相对提升,并在固有复杂任务上观察到进一步的增益,展示了构建更强大虚拟代理的可行性。
cs.CV / 14 / 2608.27871
Temporal Tree of Thought: Reasoning-Guided Visual Cue Search for Long-Video Understanding
时间思维树:基于推理的视觉线索搜索用于长视频理解
Abstract
Long-video understanding remains challenging for Multimodal Large Language Models (MLLMs) due to limited context length. Uniform sampling may miss crucial moments, while agent-based frame video understanding methods often evaluate frames independently, overlooking the temporal organization of videos. Ideally, evidence selection should mimic how humans answer questions about long videos: first locating the relevant segment from the global context, then zooming into local objects and details. We propose Temporal Tree of Thought T^3, a training-free framework for adaptive coarse-to-fine long-video understanding. T^3 constructs a question-agnostic hierarchical temporal tree via recursive temporally constrained clustering, where each node represents a contiguous segment with an informative key frame. During inference, T^3 performs an answer-retrieve-explore loop: it reasons over coarse representative frames, generates a search statement when evidence is insufficient, and expands relevant branches for finer-grained evidence. This process adaptively shifts the search target from temporal regions to specific objects and visual details to help video understanding. Experiments on VideoMME, LongVideoBench, and LVBench show that T^3 improves Qwen2.5-VL-7B by 0.5%, 4.6%, and 4.4%, respectively, under the same frame budget, demonstrating the effectiveness of structured temporal reasoning.
Chinese Translation
由于上下文长度的限制,长视频理解对多模态大型语言模型(MLLMs)仍然具有挑战性。均匀采样可能会错过关键时刻,而基于代理的帧视频理解方法通常独立评估帧,忽视了视频的时间组织。理想情况下,证据选择应模仿人类回答长视频问题的方式:首先从全局上下文中定位相关片段,然后聚焦于局部对象和细节。我们提出了时间思维树(Temporal Tree of Thought,T^3),这是一个无需训练的自适应粗到细的长视频理解框架。T^3通过递归时间约束聚类构建一个与问题无关的分层时间树,其中每个节点代表一个连续片段,并具有一个信息丰富的关键帧。在推理过程中,T^3执行一个答案检索-探索循环:它对粗略的代表性帧进行推理,当证据不足时生成搜索语句,并扩展相关分支以获取更细粒度的证据。这个过程自适应地将搜索目标从时间区域转移到特定对象和视觉细节,以帮助视频理解。在VideoMME、LongVideoBench和LVBench上的实验表明,在相同的帧预算下,T^3分别提高了Qwen2.5-VL-7B的性能0.5%、4.6%和4.4%,证明了结构化时间推理的有效性。
cs.CV / 15 / 2608.27877
Relational Knowledge Distillation Brings DNN Representations Close Enough to Humans to Be Aligned Without Supervision
关系知识蒸馏使深度神经网络表示与人类足够接近,以便在无监督情况下进行对齐
Abstract
Linking the internal representations of deep neural networks (DNNs) to human mental representations is important for using DNNs as computational models of human vision. Existing DNN representations remain insufficiently similar to human mental representations, which are not directly observable and are therefore commonly measured through large-scale similarity judgments of object images. A natural approach to narrowing this gap is to directly transfer the relational structure of human representations into DNNs, and previous studies have reported improved human-DNN representational similarity. However, whether this improvement holds under stricter evaluation remains untested in two respects: fine-grained alignment at the individual-object level, and generalization to a human embedding derived from a dataset independent of the training data. Here, we employ an unsupervised comparison method, Gromov-Wasserstein optimal transport (GWOT), which estimates human-DNN correspondences from the internal distance structure alone and thereby tests fine-grained alignment. We further assess generalization on a curated test set of concepts non-overlapping with the training data. We show that fine-tuning pre-trained DNNs with Relational Knowledge Distillation (RKD), an established relational transfer method, brings DNNs close enough to humans to be aligned at the individual-object level on this test set. We also show that this improvement is driven by a more human-like global structure, as reflected in the ordering of distances among coarse categories, while the local human-DNN nearest-neighbor overlap rate remains largely unchanged. These findings indicate that relational transfer from humans brings the global structure of pre-trained DNNs close enough to the human structure to enable fine-grained human-DNN alignment without supervision.
Chinese Translation
将深度神经网络(DNN)的内部表示与人类心理表征联系起来,对于将DNN作为人类视觉的计算模型至关重要。现有的DNN表示与人类心理表征之间的相似性仍然不足,而人类心理表征是不可直接观察的,因此通常通过对物体图像的大规模相似性判断进行测量。缩小这一差距的自然方法是将人类表征的关系结构直接转移到DNN中,之前的研究报告了人类与DNN之间的表征相似性有所改善。然而,这种改善在更严格的评估下是否依然成立尚未得到检验,具体体现在两个方面:个体物体层面的细粒度对齐,以及对来自与训练数据无关的数据集的人类嵌入的泛化。在此,我们采用了一种无监督比较方法,即Gromov-Wasserstein最优传输(GWOT),该方法仅通过内部距离结构估计人类与DNN之间的对应关系,从而测试细粒度对齐。我们进一步在一个与训练数据不重叠的概念经过筛选的测试集上评估泛化能力。我们展示了通过关系知识蒸馏(RKD)对预训练DNN进行微调,这是一种已建立的关系转移方法,使DNN在该测试集上与人类在个体物体层面上足够接近以实现对齐。我们还表明,这一改善是由更类似人类的全局结构驱动的,这在粗类别之间距离的排序中得以体现,而局部人类-DNN最近邻重叠率基本保持不变。这些发现表明,从人类的关系转移使预训练DNN的全局结构与人类结构足够接近,从而能够在无监督情况下实现细粒度的人类-DNN对齐。
cs.CV / 16 / 2608.27879
What Do Interaction Representations Actually Measure? Pre-Event Separability in Weakly-Supervised Violence Detection
交互表示究竟测量了什么?弱监督暴力检测中的事件前可分离性
Abstract
Articulated human pose provides detailed body-configuration information beyond coarse spatial relationships, but whether this detail yields greater discriminative information when the downstream pipeline is held fixed remains unclear. We examine this through early violence detection. Holding the tracker, temporal head, supervision, folds, and evaluation fixed, we compare five interaction representations spanning coarse bounding-box geometry, a matched handcrafted pose analogue, enriched pose descriptors, and a matched-capacity encoder learned from raw joints, under video-level evaluation with cluster-bootstrap intervals. No pose-based representation outperforms coarse geometry, though with fifteen anomalous videos this subset cannot rule out small effects. Extending the pipeline to frozen visual encoders, and repeating the comparison on XD-Violence (137 anomalous videos, nine times our UCF-Crime sample), person-crop appearance and whole-frame context both exceed geometry by a wide margin, yet context matches appearance on UCF-Crime and exceeds it on the larger split: cropping to the interacting people yields no advantage over encoding the whole frame. This prompts a direct test of what the benchmark measures. Scoring anomalous videos using only frames preceding the annotated onset, under a control removing sequence length as a cue, retains 39-91% of above-chance separation on both benchmarks, including for seven hand-designed geometric channels. Inspection of the tightest pre-onset windows identifies concrete provenance artifacts: editorial title cards and platform watermarks absent from the surveillance footage supplying the normal class. Video-level AUC here is thus a composite of event evidence and pre-event source cues, a shared source of discrimination that can obscure differences between representations. The diagnostic requires only annotations these benchmarks already ship.
Chinese Translation
关节化的人体姿态提供了超越粗略空间关系的详细身体配置信息,但在下游管道固定的情况下,这种细节是否能产生更大的区分信息仍不清楚。我们通过早期暴力检测来检验这一点。在固定跟踪器、时间头、监督、折叠和评估的情况下,我们比较了五种交互表示,涵盖粗略的边界框几何、匹配的手工姿态类比、丰富的姿态描述符,以及从原始关节学习的匹配容量编码器,在视频级评估中使用聚类自助区间。没有任何基于姿态的表示优于粗略几何,尽管在十五个异常视频的情况下,这一子集无法排除小效应。将管道扩展到冻结的视觉编码器,并在XD-Violence(137个异常视频,是我们UCF-Crime样本的九倍)上重复比较,人物裁剪外观和整帧上下文都大幅超越几何,然而在UCF-Crime上,上下文与外观相匹配,并在更大的分割中超过外观:裁剪到交互人物并未比编码整个帧带来优势。这促使我们直接测试基准测量的内容。仅使用注释开始前的帧对异常视频进行评分,在去除序列长度作为线索的控制下,保留了在两个基准上39-91%的超出偶然分离,包括七个手工设计的几何通道。对最紧密的开始前窗口的检查识别出具体的来源伪影:编辑标题卡和平台水印在提供正常类别的监控录像中缺失。因此,这里的视频级AUC实际上是事件证据和事件前源线索的复合体,这是一种共享的区分来源,可能会掩盖表示之间的差异。该诊断仅需要这些基准已经提供的注释。
cs.CV / 17 / 2608.27881
StreamEMS: Streaming Video Understanding with Self-Evolving Memory Scheme for Vision-Language Models
StreamEMS:基于自演化记忆机制的流媒体视频理解方法
Abstract
Recently, many streaming video understanding methods have been proposed by constructing an external memory to store historical data for computational reduction. Most methods focus on optimizing the injection procedure of current data (write) and retrieving informative historical data (read) from memory, while overlooking the opportunity to further enhancing the representational capability of memory itself. In this work, we present StreamEMS, a general mechanism for improving streaming video understanding by re-structuring the historical data stored in memory through self-evolving memory scheme, enabling more informative and robust memory representations. Specifically, we first introduce a Semantic Evolution Module to evolve the memory into more information-dense representations by exploiting informative memory entities discovered via progressively shrinking semantic scales from coarse to fine. In addition, we further introduce a Prior-informed Evolution Module to evolve memory into more robust representations by leveraging prior memory distributions to refine the current memory state. We validate the effectiveness of our proposed designs on widely-used streaming video understanding datasets, i.e., OVO-Bench and StreamingBench, and the results showcase that our method performs better than other methods. Moreover, the advantage of our method becomes consistently evident even under high token usage drop rate settings, indicating the effectiveness and robustness of our method in unleashing the potential of the memory itself.
Chinese Translation
近年来,许多流媒体视频理解方法通过构建外部记忆来存储历史数据以减少计算开销。大多数方法专注于优化当前数据的注入过程(写入)和从记忆中检索有信息的历史数据(读取),而忽视了进一步增强记忆本身表征能力的机会。在本研究中,我们提出了StreamEMS,这是一种通过自演化记忆机制重构存储在记忆中的历史数据,从而改善流媒体视频理解的通用机制,使得记忆表征更加信息丰富和稳健。具体而言,我们首先引入了语义演化模块(Semantic Evolution Module),通过利用逐步缩小的语义尺度从粗到细发现的信息记忆实体,将记忆演化为更具信息密度的表征。此外,我们进一步引入了先验信息演化模块(Prior-informed Evolution Module),通过利用先验记忆分布来细化当前记忆状态,将记忆演化为更稳健的表征。我们在广泛使用的流媒体视频理解数据集(如OVO-Bench和StreamingBench)上验证了我们提出设计的有效性,结果表明我们的方法优于其他方法。此外,即使在高标记使用下降率设置下,我们方法的优势也始终明显,表明我们的方法在释放记忆本身潜力方面的有效性和稳健性。
cs.CV / 18 / 2608.27888
Thread-Efficient Decoding for Neural Texture Compression
神经纹理压缩的线程高效解码
Abstract
Neural texture compression (NTC) achieves higher compression ratios than BCn formats but suffers from GPU thread divergence, which significantly reduces runtime performance. In this work, we propose a shared decoder MLP architecture -- trained with a gradual decoder freezing schedule -- combined with texture clustering to reduce thread divergence by 25%-52% while preserving rendering quality. We evaluate our method on over 500 textures and multiple real rendering scenes, demonstrating up to 8.48x speedup on the Radeon RX 9070 XT GPU compared to non-shared baselines. Our key contributions include: (1) a unified shared decoder architecture that reduces divergence by grouping textures; (2) a training recipe with gradual decoder freezing that improves stability and reconstruction accuracy; (3) a semantic clustering strategy using CLIP embeddings that groups similar textures for effective decoder sharing; and (4) comprehensive performance and ablation studies validating our approach.
Chinese Translation
神经纹理压缩(NTC)实现了比BCn格式更高的压缩比,但受到GPU线程分歧的影响,这显著降低了运行时性能。在本研究中,我们提出了一种共享解码器多层感知机(MLP)架构——采用逐步解码器冻结调度进行训练——结合纹理聚类,以减少25%-52%的线程分歧,同时保持渲染质量。我们在超过500种纹理和多个真实渲染场景上评估了我们的方法,结果显示与非共享基线相比,在Radeon RX 9070 XT GPU上实现了高达8.48倍的加速。我们的主要贡献包括:(1)一种统一的共享解码器架构,通过对纹理进行分组来减少分歧;(2)一种逐步解码器冻结的训练方案,提高了稳定性和重建精度;(3)一种使用CLIP嵌入的语义聚类策略,将相似纹理分组以实现有效的解码器共享;(4)全面的性能和消融研究验证了我们的方法。
cs.CV / 19 / 2608.27893
CommerceVibe: Learning to Design E-Commerce Creatives as Executable Visual Code via Dual-Feedback Reinforcement Learning
CommerceVibe:通过双重反馈强化学习学习设计可执行视觉代码的电子商务创意
Abstract
High-quality e-commerce creatives are essential for presenting products and conveying marketing messages. Recent diffusion models enable scalable creative generation and produce visually compelling images, but their flattened raster outputs often contain distorted text and inconsistent product details, requiring refinement before deployment. Moreover, without explicit structure, the resulting creatives are difficult to edit and reuse, while complex design requirements remain challenging to encode as verifiable training signals. To address these challenges, we present CommerceVibe, which represents creatives as executable visual code and formulates generation as conditional HTML/CSS program synthesis. Given product images, design requirements, and product information, it produces renderable, editable, and reusable creatives. We further introduce dual-feedback reinforcement learning, in which rule-based feedback evaluates rendered programs for text readability, product visibility, and layout validity, while visual feedback from a vision-language model (VLM) assesses rendered creatives against input specifications across six perceptual and commercial dimensions. Together, these complementary feedback signals improve both constraint satisfaction and perception-dependent quality. We perform supervised fine-tuning (SFT) of Qwen3.5-9B on over 28,000 e-commerce examples, followed by dual-feedback reinforcement learning. On a 1,300-case benchmark, the optimized CommerceVibe model achieves a weighted score of 94.0/100, compared with 87.3 for the SFT-only variant, and outperforms strong external models. Blind evaluations by five e-commerce design experts further validate these improvements. CommerceVibe supports controllable, editable, and scalable e-commerce creative production.
Chinese Translation
高质量的电子商务创意对于展示产品和传达营销信息至关重要。最近的扩散模型使得可扩展的创意生成成为可能,并生成视觉上引人注目的图像,但其扁平化的光栅输出往往包含扭曲的文本和不一致的产品细节,需在部署前进行精细化处理。此外,由于缺乏明确的结构,生成的创意难以编辑和重用,而复杂的设计要求仍然难以编码为可验证的训练信号。为了解决这些挑战,我们提出了CommerceVibe,它将创意表示为可执行的视觉代码,并将生成过程表述为条件HTML/CSS程序合成。给定产品图像、设计要求和产品信息,它能够生成可渲染、可编辑和可重用的创意。我们进一步引入了双重反馈强化学习,其中基于规则的反馈评估渲染程序的文本可读性、产品可见性和布局有效性,而来自视觉语言模型(VLM)的视觉反馈则根据输入规范在六个感知和商业维度上评估渲染的创意。这些互补的反馈信号共同提高了约束满足和感知依赖质量。我们对Qwen3.5-9B进行了监督微调(SFT),使用了超过28,000个电子商务示例,随后进行了双重反馈强化学习。在一个1,300案例的基准测试中,优化后的CommerceVibe模型达到了94.0/100的加权分数,而仅进行SFT的变体为87.3,并且超越了强大的外部模型。五位电子商务设计专家的盲评进一步验证了这些改进。CommerceVibe支持可控、可编辑和可扩展的电子商务创意生产。
cs.CV / 20 / 2608.27922
DensityKV: Density-Guided KV Cache Compression for Long Video Generation
DensityKV:用于长视频生成的密度引导键值缓存压缩
Abstract
Autoregressive video diffusion models enable streaming generation through sliding-window attention, but each generated block is conditioned on previously generated content, causing appearance and motion errors to propagate recursively over time. Historical key-value (KV) memory preserves earlier subject and scene states and helps maintain long-horizon consistency. However, retaining every generated state creates a historical archive that grows continuously with the rollout, while recurrent states repeatedly add redundant coverage. To address this problem, we propose DensityKV, a training-free historical KV bank management strategy. DensityKV maintains a separate token-level KV bank for each attention head and measures local redundancy among the post-RoPE keys that directly parameterize attention routing using Soft-Riesz density. By constraining neighborhood-density growth after states enter the bank, DensityKV limits repeated historical accumulation while preserving coherent states from each completed generation block. Experiments across three autoregressive video generation backbones and multiple generation lengths show that, at the same upper bound on historical KV capacity, DensityKV improves long-horizon consistency and generation stability while keeping persistent historical storage bounded independently of rollout length.
Chinese Translation
自回归视频扩散模型通过滑动窗口注意力实现流式生成,但每个生成的块都依赖于先前生成的内容,导致外观和运动错误在时间上递归传播。历史键值(KV)内存保留早期的主题和场景状态,并有助于维持长时间的一致性。然而,保留每个生成状态会创建一个随着生成过程不断增长的历史档案,而递归状态则重复增加冗余覆盖。为了解决这个问题,我们提出了DensityKV,一种无训练的历史KV银行管理策略。DensityKV为每个注意力头维护一个单独的令牌级KV银行,并使用Soft-Riesz密度测量后RoPE键之间的局部冗余,这些键直接参数化注意力路由。通过在状态进入银行后限制邻域密度的增长,DensityKV限制了重复的历史积累,同时保留了每个完成的生成块中的一致状态。在三个自回归视频生成主干和多个生成长度的实验中表明,在相同的历史KV容量上限下,DensityKV提高了长时间一致性和生成稳定性,同时保持持久的历史存储独立于生成长度的界限。
cs.CV / 21 / 2608.27923
PCBnet: A Dataset and Automatic Construction of SPICE Netlists from Schematic Images
PCBnet:从原理图图像自动构建SPICE网表的数据集
Abstract
Printed circuit boards (PCBs) are fundamental to modern electronic systems, yet AI-driven PCB design automation remains constrained by the lack of large-scale paired schematic-netlist datasets. PCB schematics are particularly challenging due to diverse component types, complex wiring topologies, and noisy textual annotations. To address this gap, we present PCBnet, a large-scale PCB schematic dataset comprising over 300 real-world designs with annotated pins and paired SPICE netlists. It contains more than 50,000 component instances, 150,000 wires, 100,000 text regions, and 400,000 characters. We further develop an automated schematic-to-netlist pipeline that combines visual recognition, topology construction, and domain-knowledge-guided multi-agent correction. The proposed method achieves 94.54% component detection mAP, 98.57% text recognition accuracy, and 84.47% end-to-end connectivity accuracy. PCBnet provides a benchmark and data foundation for future AI-driven PCB design automation.
Chinese Translation
印刷电路板(PCBs)是现代电子系统的基础,但基于人工智能的PCB设计自动化仍然受到缺乏大规模配对原理图-网表数据集的限制。由于元件类型多样、布线拓扑复杂以及文本注释噪声较大,PCB原理图尤其具有挑战性。为了解决这一问题,我们提出了PCBnet,这是一个大规模的PCB原理图数据集,包含超过300个真实设计,附带注释引脚和配对的SPICE网表。该数据集包含超过50,000个元件实例、150,000条导线、100,000个文本区域和400,000个字符。我们进一步开发了一个自动化的原理图到网表的流程,该流程结合了视觉识别、拓扑构建和基于领域知识的多代理修正。所提出的方法实现了94.54%的元件检测均值平均精度(mAP)、98.57%的文本识别准确率和84.47%的端到端连通性准确率。PCBnet为未来基于人工智能的PCB设计自动化提供了基准和数据基础。
cs.CV / 22 / 2608.27929
Training-Free Temporal Abstraction for General Video Understanding
无训练的时间抽象用于通用视频理解
Abstract
Videos are expensive to analyze frame by frame, yet many video understanding tasks depend on knowing where relevant moments occur. A system may need to find when an action changes, locate the segment described by a sentence, or choose a few frames for a vision-language model. Existing methods often solve these problems separately, using task-specific training data or specialized architectures. We study whether a pretrained video-text model can provide enough temporal structure to support several of these tasks at once. We present STITCH, a training-free method that divides a video into semantically meaningful temporal chunks. STITCH embeds short video windows with a frozen video-text backbone and detects changes in the resulting embedding sequence. These chunks are computed once per video and reused across tasks. We evaluate STITCH on generic event boundary detection, language-based moment retrieval, and frame selection for long-video VLM reasoning. Across all three settings, STITCH remains competitive with more specialized methods while requiring no task-specific training, with especially clear gains when only a small number of frames or tokens can be processed. These results suggest that reusable temporal abstraction is a promising direction for general video understanding, allowing dense video streams to be converted once into semantic units that can be localized, retrieved, sampled, or reasoned over by downstream systems.
Chinese Translation
逐帧分析视频成本高昂,但许多视频理解任务依赖于知道相关时刻的发生。系统可能需要找到动作变化的时刻,定位由句子描述的片段,或为视觉-语言模型选择少数帧。现有方法通常分别解决这些问题,使用特定任务的训练数据或专门架构。我们研究了预训练的视频-文本模型是否能够提供足够的时间结构,以同时支持多个任务。我们提出了STITCH,这是一种无训练的方法,将视频划分为语义上有意义的时间块。STITCH使用冻结的视频-文本主干嵌入短视频窗口,并检测生成的嵌入序列中的变化。这些块在每个视频中计算一次,并在任务之间重用。我们在通用事件边界检测、基于语言的时刻检索和长视频视觉-语言模型推理的帧选择上评估了STITCH。在所有三个设置中,STITCH与更专业的方法保持竞争力,同时不需要特定任务的训练,尤其是在只能处理少量帧或标记时表现出明显的优势。这些结果表明,可重用的时间抽象是通用视频理解的一个有前景的方向,使得密集的视频流能够一次性转换为可以被下游系统定位、检索、采样或推理的语义单元。
cs.CV / 23 / 2608.27971
GAAT: Geometry-Aware Alignment Transformer for Multimodal UAV Perception
GAAT:用于多模态无人机感知的几何感知对齐变换器
Abstract
Unmanned aerial vehicle (UAV) multimodal perception integrates visible (RGB), infrared (IR), synthetic aperture radar (SAR), and depth sensors for scene understanding under diverse conditions. However, differences in optics, resolution, and mounting often limit practical systems to global or image-center alignment. After tokenization, parallax, platform motion, and lens distortion can shift corresponding patch centers across modalities, weakening the spatial correspondence assumed by dense contrastive learning and cross-modal fusion. We propose GAAT (Geometry-Aware Alignment Transformer), an alignment-first pretrained model that estimates local correspondence reliability before cross-modal interaction. GAAT introduces syncPATC, which learns patch-center consistency under synchronized view transformations without correspondence annotations. It emits geometric priors, including token and query confidence, query centers, and sub-token offsets, that identify reliable local anchors across residual misalignment. Guided by these priors, MG-Sparse-MMA performs query-mediated sparse fusion over top-K_s reliable regions, replacing dense all-patch interaction with geometry-calibrated local updates. RA-QCGCL aligns pretraining supervision with this sparse query bottleneck through reliable patch-to-patch, patch-to-query, and query-to-query contrastive branches. We introduce UAVMeta and StateBench, which provide four acquisition-state scores derived from platform telemetry and image statistics: camera reliability, observation scale, viewpoint stability, and flight maneuver complexity. Extensive experiments across six downstream tasks demonstrate consistently superior transfer performance, establishing GAAT as a state-of-the-art multimodal foundation model for UAV perception. StateBench further enables a systematic diagnosis of real-world acquisition conditions.
Chinese Translation
无人机(UAV)多模态感知整合了可见光(RGB)、红外(IR)、合成孔径雷达(SAR)和深度传感器,以在多种条件下进行场景理解。然而,光学、分辨率和安装方式的差异常常使实际系统仅限于全局或图像中心对齐。在标记化之后,视差、平台运动和镜头畸变可能会导致不同模态之间对应补丁中心的偏移,从而削弱了密集对比学习和跨模态融合所假设的空间对应关系。我们提出了GAAT(几何感知对齐变换器),这是一种以对齐为先的预训练模型,在跨模态交互之前估计局部对应关系的可靠性。GAAT引入了syncPATC,它在没有对应注释的情况下学习补丁中心在同步视图变换下的一致性。它发出几何先验,包括标记和查询的置信度、查询中心和子标记偏移量,以识别在残余错位中可靠的局部锚点。在这些先验的指导下,MG-Sparse-MMA在前K_s个可靠区域上执行查询介导的稀疏融合,用几何校准的局部更新替代密集的全补丁交互。RA-QCGCL通过可靠的补丁到补丁、补丁到查询和查询到查询的对比分支,将预训练监督与这一稀疏查询瓶颈对齐。我们引入了UAVMeta和StateBench,它们提供了四个基于平台遥测和图像统计的获取状态评分:相机可靠性、观察尺度、视点稳定性和飞行机动复杂性。针对六个下游任务的广泛实验表明,GAAT在迁移性能上始终优于其他方法,确立了其作为无人机感知的最先进多模态基础模型的地位。StateBench进一步使得对真实世界获取条件的系统诊断成为可能。
cs.CV / 24 / 2608.27989
GAN-Based Semantic Communication for Image Transmission in IoV
基于GAN的车联网图像传输语义通信
Abstract
For cooperative perception in the internet of vehicles, this paper proposes a generative adversarial network-based semantic communication framework to address the efficiency and fidelity bottlenecks of traditional communication systems in visual data transmission under limited bandwidth and dynamic channel conditions. At the transmitter, the framework adopts a pyramid attention network to extract semantic label maps and introduces a semantic priority preservation mechanism. It assigns differentiated weights to distinct semantic categories based on driving safety, guiding bit allocation and loss function design. At the receiver, an image reconstruction module integrating a coarse to-fine multi-resolution generator and multi-scale discriminator is designed. Combined with the temporal consistency branch, spatial pyramid pooling and class-aware convolutional layers, it achieves high-fidelity reconstruction of high-quality images from corrupted semantic labels. The model is trained with combined adversarial, feature matching and perceptual losses, effectively improving semantic consistency and visual realism of generated images. Experimental results on the Cityscapes dataset show that the proposed method outperforms existing counterparts in both semantic segmentation accuracy and reconstructed image quality, and maintains stable reconstruction performance under AWGN and Rayleigh channels.
Chinese Translation
针对车联网中的协同感知,本文提出了一种基于生成对抗网络的语义通信框架,以解决传统通信系统在有限带宽和动态信道条件下视觉数据传输的效率和保真度瓶颈。在发射端,该框架采用金字塔注意力网络提取语义标签图,并引入语义优先级保留机制。它根据驾驶安全性为不同的语义类别分配差异化权重,从而指导比特分配和损失函数设计。在接收端,设计了一个图像重建模块,该模块集成了粗到细的多分辨率生成器和多尺度判别器。结合时间一致性分支、空间金字塔池化和类别感知卷积层,实现了从损坏的语义标签中高保真重建高质量图像。该模型通过结合对抗损失、特征匹配损失和感知损失进行训练,有效提高了生成图像的语义一致性和视觉真实感。在Cityscapes数据集上的实验结果表明,所提方法在语义分割精度和重建图像质量上均优于现有方法,并在AWGN和Rayleigh信道下保持稳定的重建性能。
cs.CV / 25 / 2608.27997
A-PAIR: A Benchmark and Identity-Consistent Grounding Framework for Air-Ground Cross-View Referring Person Detection
A-PAIR:一种用于空地交叉视角指称人物检测的基准和身份一致性基础框架
Abstract
Air-ground cross-view referring person detection is a necessary component in the language-to-perception-to-control chain of collective embodied intelligence, grounding a language command into the same physical target before ground and aerial agents can coordinate downstream actions. Existing referring expression comprehension and open-vocabulary grounding methods do not jointly account for cross-view identity consistency, making them insufficient for Air-Ground Cross-View Referring Person Detection (AGCV-RPD), which involves similar pedestrian distractors, weak aerial appearance cues, and cross-view identity consistency. To study this problem, we introduce Air-Ground Paired Identity-Aware Referring (A-PAIR), the first comprehensive AGCV-RPD benchmark, containing 22,137 cross-view referring samples. To construct A-PAIR efficiently, we propose Factorized Annotation and Referential Alignment (FARA), a semi-automatic annotation framework that generates factorized referring descriptions and identity-consistency supervision at reduced cost. We propose Identity-Consistent Referring Grounding (ICRG), a framework that combines factorized referential grounding, candidate-completeness supervision, and cross-view consistency calibration for joint air-ground pair selection. ICRG improves ground, aerial, and pair-level detection over strong baselines, increasing pair F1 from 16.65% to 22.28%. These results show that AGCV-RPD requires paired detection and identity-consistent reasoning.
Chinese Translation
空地交叉视角指称人物检测是集体具身智能中语言-感知-控制链的重要组成部分,它将语言指令定位到同一物理目标上,以便地面和空中代理能够协调后续行动。现有的指称表达理解和开放词汇基础方法未能共同考虑交叉视角身份一致性,因此不足以满足空地交叉视角指称人物检测(AGCV-RPD)的需求,该任务涉及相似的行人干扰物、弱的空中外观线索以及交叉视角身份一致性。为了解决这一问题,我们提出了空地配对身份感知指称(A-PAIR),这是第一个综合性的AGCV-RPD基准,包含22,137个交叉视角指称样本。为了高效构建A-PAIR,我们提出了因子化注释和指称对齐(FARA),这是一种半自动注释框架,可以以较低的成本生成因子化的指称描述和身份一致性监督。我们还提出了身份一致性指称基础(ICRG),这是一个结合了因子化指称基础、候选完整性监督和交叉视角一致性校准的框架,用于联合选择空地配对。ICRG在强基线之上提高了地面、空中和配对级别的检测,将配对F1从16.65%提高到22.28%。这些结果表明,AGCV-RPD需要配对检测和身份一致性推理。
cs.CV / 26 / 2608.28008
Visual Token Coding for Video Multimodal Large Language Models
视频多模态大语言模型的视觉令牌编码
Abstract
In this paper, we propose a new token compression paradigm for video Multimodal Large Language Models (MLLMs), termed Visual Token Coding (VTC). Inspired by classical video coding principles, e.g., HEVC, VTC performs structured compression by predicting the I/P frames of a video and measuring their frame-wise residuals to estimate token redundancy. Based on this baseline framework, we also enhance VTC with a set of novel dynamic designs, such as Dynamic Resolution Input (DyRSO), Dynamic Token Allocation (DyTA), and Spatial Coverage Top-K (SC-TopK), and term this new approach $VTC_{Dy}$. To validate VTC, we apply it to three MLLMs and conduct experiments on multiple video understanding benchmarks. The experimental results show that VTC$_{\mathrm{Dy}}$ achieves an average performance retention of 100.1% with a 50% token budget for Qwen3-VL, while still retaining 97.8% of the average performance when the token budget is reduced to 25%. Moreover, as a plug-and-play design, VTC requires no additional tuning of MLLMs for token coding. Our code is available at https://github.com/Msr233/VTC.
Chinese Translation
在本文中,我们提出了一种用于视频多模态大语言模型(MLLMs)的新型令牌压缩范式,称为视觉令牌编码(Visual Token Coding, VTC)。受经典视频编码原理(如HEVC)的启发,VTC通过预测视频的I/P帧并测量其逐帧残差来进行结构化压缩,以估计令牌冗余。在此基础框架上,我们还通过一系列新颖的动态设计增强了VTC,例如动态分辨率输入(Dynamic Resolution Input, DyRSO)、动态令牌分配(Dynamic Token Allocation, DyTA)和空间覆盖Top-K(Spatial Coverage Top-K, SC-TopK),并将这种新方法称为$VTC_{Dy}$。为了验证VTC,我们将其应用于三个MLLM,并在多个视频理解基准上进行实验。实验结果表明,VTC$_{ ext{Dy}}$在Qwen3-VL上以50%的令牌预算实现了100.1%的平均性能保留,而当令牌预算降低至25%时仍保留了97.8%的平均性能。此外,作为一种即插即用设计,VTC无需对MLLM进行额外的调优即可进行令牌编码。我们的代码可在https://github.com/Msr233/VTC获取。
cs.CV / 27 / 2608.28020
3D-USE: From Image-Level to Scene-Level Underwater Enhancement
3D-USE:从图像级到场景级的水下增强
Abstract
Underwater 3D reconstruction faithfully reproduces the color shifts and visibility loss of captured views, while physical inversion may leave estimation errors in the recovered scene appearance. We formulate Underwater Scene-level Enhancement (USE) as learning a persistent, visibility-enhanced 3D scene representation from degraded multi-view underwater observations, enabling consistent enhanced rendering. Realizing USE requires both a reliable scene representation for enhancement and a consistent enhancement target without paired enhanced 3D data. Therefore, we present 3D-USE, a two-stage framework. First, the Medium Radial Basis Anchor Representation (MediumRBF) establishes a medium-aware Gaussian scene by representing water effects with shared radial-basis anchors and explicitly decomposing object and medium contributions. Based on this fixed scene representation, Appearance Transition Consensus (ATC) transfers paired 2D underwater image enhancement (UIE) knowledge into scene-global and Gaussian-local targets, avoiding direct supervision from inconsistent enhanced views. An Underwater Bilateral Appearance Field (U-BAF) then realizes these targets in Gaussian radiance and medium appearance. The scene directly renders enhanced novel views without a 2D UIE model at inference. Experiments on real underwater scenes show improved visibility and cross-view consistency while preserving reconstruction quality.
Chinese Translation
水下三维重建忠实地再现了捕获视图的颜色偏移和可见度损失,而物理反演可能在恢复的场景外观中留下估计误差。我们将水下场景级增强(Underwater Scene-level Enhancement, USE)定义为从退化的多视角水下观测中学习一个持久的、增强可见度的三维场景表示,从而实现一致的增强渲染。实现USE需要一个可靠的场景表示用于增强,以及一个没有配对增强三维数据的一致增强目标。因此,我们提出了3D-USE,一个两阶段框架。首先,中等径向基锚表示(Medium Radial Basis Anchor Representation, MediumRBF)通过使用共享的径向基锚来表示水的影响,并明确分解物体和介质的贡献,建立了一个中等感知的高斯场景。在这个固定的场景表示基础上,外观过渡共识(Appearance Transition Consensus, ATC)将配对的二维水下图像增强(Underwater Image Enhancement, UIE)知识转移到场景全局和高斯局部目标,避免了来自不一致增强视图的直接监督。然后,水下双边外观场(Underwater Bilateral Appearance Field, U-BAF)在高斯辐射和介质外观中实现这些目标。该场景在推理时直接渲染增强的新视图,而无需二维UIE模型。对真实水下场景的实验表明,在保持重建质量的同时,增强了可见度和视图间一致性。
cs.CV / 28 / 2608.28033
ZipMVS: Multi-View Stereo with Compressed Cost Volumes
ZipMVS:基于压缩代价体的多视图立体重建
Abstract
Multi-view stereo (MVS) methods typically deliver highly accurate 3D reconstructions from multiple registered RGB images, thanks to the highly informative, geometric constraints between them. However, their substantial memory requirements remain a major obstacle for deployment in domains such as aerospace and autonomous systems, where resource efficiency is critical. In this work, we introduce ZipMVS, an MVS method specifically designed for efficient high-quality reconstruction. We propose a novel depth-hypothesis strategy that enables substantial compression of the cost volume, hence greatly reducing GPU memory consumption while preserving reconstruction accuracy. Experiments on the DTU and Tanks and Temples datasets show that ZipMVS achieves competitive reconstruction quality compared with other efficiency-oriented MVS methods, while achieving a competitive balance between reconstruction quality and GPU memory usage. The code is available at https://github.com/JihnGlyn/ZipMVS
Chinese Translation
多视图立体(MVS)方法通常能够从多个已注册的RGB图像中提供高度准确的3D重建,这得益于它们之间丰富的几何约束。然而,其巨大的内存需求仍然是部署在航空航天和自主系统等资源效率至关重要领域的主要障碍。在本研究中,我们提出了ZipMVS,这是一种专门设计用于高效高质量重建的MVS方法。我们提出了一种新颖的深度假设策略,使得代价体的压缩显著,从而在保持重建精度的同时大幅降低GPU内存消耗。在DTU和Tanks and Temples数据集上的实验表明,ZipMVS在重建质量上与其他以效率为导向的MVS方法具有竞争力,同时在重建质量和GPU内存使用之间实现了良好的平衡。代码可在https://github.com/JihnGlyn/ZipMVS获取。
cs.CV / 29 / 2608.28058
Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models
动态对齐补偿用于减轻大型视觉-语言模型中的幻觉现象
Abstract
Large Vision-Language Models (LVLMs) remain prone to hallucinations, producing responses that are irrelevant or inconsistent with the multimodal input. Existing mitigation methods mainly rely on external supervision, output calibration, or attention regulation, leaving the internal representation dynamics of autoregressive generation underexplored. We identify an inference-time failure mode in which cross-modal representations degrade across decoder layers and drift across generation steps, destabilizing token prediction and increasing hallucination risk. We propose \emph{Dynamic Alignment Compensation} (DAC), a training-free inference-time method that detects representation divergence and selectively applies lightweight residual compensation. DAC combines Layer-wise Semantic Compensation to mitigate inter-layer degradation with Sequential Semantic Correction to constrain temporal drift. Experiments on nine hallucination-focused and general-purpose multimodal benchmarks across multiple LVLM backbones show that DAC consistently reduces hallucinations while maintaining strong overall performance.
Chinese Translation
大型视觉-语言模型(LVLMs)仍然容易出现幻觉,产生与多模态输入无关或不一致的响应。现有的减轻方法主要依赖于外部监督、输出校准或注意力调节,导致自回归生成的内部表示动态未被充分探索。我们识别出一种推理时的失败模式,其中跨模态表示在解码器层之间退化,并在生成步骤中漂移,从而不稳定令牌预测并增加幻觉风险。我们提出了 extit{动态对齐补偿}(Dynamic Alignment Compensation, DAC),这是一种无训练的推理时方法,能够检测表示的偏离并选择性地应用轻量级残差补偿。DAC结合了逐层语义补偿以减轻层间退化和顺序语义校正以约束时间漂移。在九个以幻觉为重点的和通用的多模态基准测试中进行的实验表明,DAC在保持强大整体性能的同时,始终减少幻觉现象。
cs.CV / 30 / 2608.28063
A Controlled Audit of Architectural Complexity in Uncertainty-Aware Multi-Organ Ultrasound Classification
不确定性感知多脏器超声分类中的架构复杂性控制审计
Abstract
Multi-organ ultrasound classifiers increasingly combine attention, mixture-of-experts routing, uncertainty gating, and evidential deep learning (EDL) objectives to address heterogeneous anatomy and acquisition. Yet a plausible design rationale does not by itself establish that an added component improves the trained system. We contribute a controlled complexity-audit framework, applied to the deployment decision between the maximal evidential candidate Full-EDL and simpler alternatives. Six candidates were evaluated on the primary dataset and three in an internal replication, using ten matched seeds, frozen image-level partitions, capacity- and optimisation-aware comparisons, symmetric temperature scaling, paired decision rules, and a separate out-of-distribution (OOD) veto. Retaining Full-EDL did not establish a reliable macro-F1 gain on either dataset, while the simplified alternatives remained inconclusive under the non-inferiority margin. Simple cross-entropy with temperature scaling (Simple-CE+TS) met the calibrated negative log-likelihood criterion on both datasets and showed favourable selective-risk ordering. The raw calibration advantage of evidential training disappeared after temperature scaling and did not recur on the second dataset. The gate had negligible observable influence at the audited checkpoints, and deleting the Full-only chain revealed no stable task or calibrated-loss benefit. Simple-CE nevertheless triggered the OOD veto against the fetal probe but not the lung probe, precluding an unconditional OOD-safety claim. We therefore selected Simple-CE+TS for the evaluated in-distribution objective while retaining Full-EDL as the maximal reference. Components should earn retention through functional and retraining-based evidence, and calibration and distribution-shift reliability should be evaluated separately.
Chinese Translation
多脏器超声分类器越来越多地结合了注意力机制、专家混合路由、不确定性门控和证据深度学习(EDL)目标,以应对异质解剖结构和采集。然而,合理的设计理由并不能单独证明添加的组件改善了训练系统。我们提出了一个控制复杂性审计框架,应用于在最大证据候选的全证据深度学习(Full-EDL)与更简单替代方案之间的部署决策。对主要数据集评估了六个候选模型,并在内部复制中评估了三个,使用了十个匹配的种子、固定的图像级分区、容量和优化感知比较、对称温度缩放、配对决策规则以及一个单独的分布外(OOD)否决。保留全证据深度学习并未在任何数据集上建立可靠的宏F1增益,而简化的替代方案在非劣性边界下仍然没有结论。简单交叉熵与温度缩放(Simple-CE+TS)在两个数据集上均满足了校准的负对数似然标准,并显示出有利的选择风险排序。证据训练的原始校准优势在温度缩放后消失,并未在第二个数据集上重新出现。审计检查点处,门控对可观察的影响微乎其微,删除仅全证据链条未显示出稳定的任务或校准损失收益。尽管如此,Simple-CE对胎儿探头触发了分布外否决,但对肺探头则没有,从而排除了无条件的分布外安全声明。因此,我们选择了Simple-CE+TS作为评估的内部分布目标,同时保留Full-EDL作为最大参考。组件应通过功能和基于再训练的证据获得保留,校准和分布转移的可靠性应单独评估。
cs.CV / 31 / 2608.28069
VersaGauss: A Versatile Framework for Generating Multiphase Dynamics with 3D Gaussians
VersaGauss:一个用于生成多相动态的多功能框架,基于3D高斯模型
Abstract
Recent progress has been made in 3D Gaussian representation for reconstruction, generation, and physical simulation. However, current approaches mainly concentrate on physics-based dynamic generation of solid objects and only handle single-phase collision interactions. We introduce VersaGauss, a unified framework for generation, simulation, and rendering that supports versatile physics-based dynamic generation, particularly for multiphase interactions. Our system takes a few images as input and produces a realistic, physics-driven 3D dynamic scene with multiple objects. To optimize the Gaussian kernel distribution, we develop a particle pruning algorithm. We also propose the Coupled Multiphase Point Method (CMPM) to effectively model and generate multiphase interactions. Additionally, harmonic interpolation within CMPM and a Gaussian evolution strategy are introduced to achieve realistic fluid rendering. Extensive experiments demonstrate that our framework can simulate interactions among various materials such as fluid, rubber, sand, snow, and others. Code is available at https://github.com/Elowen-surj/VersaGauss.
Chinese Translation
近年来,在3D高斯表示用于重建、生成和物理模拟方面取得了进展。然而,目前的方法主要集中在基于物理的固体物体动态生成上,仅处理单相碰撞交互。我们提出了VersaGauss,一个统一的生成、模拟和渲染框架,支持多功能的基于物理的动态生成,特别是针对多相交互。我们的系统以少量图像作为输入,生成一个现实的、基于物理驱动的3D动态场景,包含多个物体。为了优化高斯核分布,我们开发了一种粒子修剪算法。我们还提出了耦合多相点方法(Coupled Multiphase Point Method, CMPM),以有效建模和生成多相交互。此外,在CMPM中引入了谐波插值和高斯演化策略,以实现逼真的流体渲染。大量实验表明,我们的框架能够模拟流体、橡胶、沙子、雪等各种材料之间的交互。代码可在 https://github.com/Elowen-surj/VersaGauss 获取。
cs.CV / 32 / 2608.28070
CF-YOLO: Context-Aware Feature Refinement for Camouflaged Industrial Micro-Defect Detection
CF-YOLO:用于伪装工业微缺陷检测的上下文感知特征精炼
Abstract
Automated detection of surface micro-defects on industrial components, such as copper tubes, is critically important for quality assurance but remains challenging due to the minute scale of anomalies and their visual camouflage against complex backgrounds. These factors lead to weak feature representations and high rates of false positives and missed detections. To address these issues, we propose a novel real-time detection framework designed for efficient context perception and feature refinement. Our method integrates a Context-Perception Aggregation Module (CPAM), which synergises large-kernel perception for macro-texture context and small-kernel aggregation for sharp boundary delineation, effectively breaking the background camouflage. Furthermore, a Feature Additive Refinement Module (FARM) employs a linear-complexity additive token mixer to globally verify and refine the representation of fine-grained anomalies, suppressing noise-induced errors. To support research in this domain, we introduce the Copper Tube Defect Dataset (CTDD), a manually annotated benchmark containing 1,847 images and 4,898 boundingbox defect instances from copper-tube inspection scenarios. Extensive experiments demonstrate that our detector achieves strong and consistent performance on CTDD, outperforming representative baseline detectors, including YOLOv11, by 2.2% in mAP@50 and 3.9% in Precision while maintaining real-time inference speed. This work provides a robust and efficient solution for high-precision industrial inspection, bridging the gap between contextual understanding and detailed feature analysis. Our code and model are available at: https://github.com/Yu-Xinda/CFYOLO-Context-Aware-Feature-Refinement-for-Camouflaged-Industrial-Micro-Defect-Detection
Chinese Translation
工业组件表面微缺陷的自动检测,如铜管,对于质量保证至关重要,但由于缺陷的微小尺度及其在复杂背景下的视觉伪装,检测仍然具有挑战性。这些因素导致特征表示薄弱以及较高的误报率和漏检率。为了解决这些问题,我们提出了一种新颖的实时检测框架,旨在实现高效的上下文感知和特征精炼。我们的方法集成了上下文感知聚合模块(Context-Perception Aggregation Module, CPAM),该模块结合了大核感知用于宏观纹理上下文和小核聚合用于清晰边界划分,有效地打破了背景伪装。此外,特征加法精炼模块(Feature Additive Refinement Module, FARM)采用线性复杂度的加法令牌混合器,全球验证和精炼细粒度缺陷的表示,抑制噪声引起的错误。为了支持该领域的研究,我们引入了铜管缺陷数据集(Copper Tube Defect Dataset, CTDD),这是一个手动标注的基准数据集,包含1,847张图像和4,898个铜管检测场景中的边界框缺陷实例。大量实验表明,我们的检测器在CTDD上实现了强大且一致的性能,超过了包括YOLOv11在内的代表性基线检测器,在mAP@50上提高了2.2%,在精度上提高了3.9%,同时保持实时推理速度。该研究为高精度工业检测提供了一个稳健且高效的解决方案,弥合了上下文理解与细节特征分析之间的差距。我们的代码和模型可在以下网址获取:https://github.com/Yu-Xinda/CFYOLO-Context-Aware-Feature-Refinement-for-Camouflaged-Industrial-Micro-Defect-Detection
cs.CV / 33 / 2608.28078
Task-State Adaptation with Prototype Memory for Multi-Task Dense Prediction
基于原型记忆的多任务密集预测任务状态适应
Abstract
Vision foundation backbones provide strong representations for dense prediction, yet a single shared feature still needs to support tasks with different, image-dependent adaptation requirements. We propose MemMTL, a multi-task dense prediction framework that estimates a compact task state from global visual context and refines it through a learnable task-state prototype memory. The refined state is converted into task-conditioned expert logits and combined with token-level logits before sparse top-$k$ selection over a local expert bank shared by all tasks. A separate task-agnostic residual bank provides a common adaptation path, and both paths are added once to the backbone feature before task-specific prediction. We specify a matched evaluation protocol on NYUD-v2 and PASCAL-Context with SAM 3 and ViT-L backbones to measure predictive quality, computational cost, and the contributions of task-state conditioning, prototype retrieval, and sparse routing. The numerical record in the present working draft predates this canonical implementation and must be regenerated before it can support empirical claims.
Chinese Translation
视觉基础骨干网络为密集预测提供了强大的表示能力,但单一共享特征仍需支持具有不同、依赖图像的适应需求的任务。我们提出了 MemMTL,这是一种多任务密集预测框架,它从全局视觉上下文中估计紧凑的任务状态,并通过可学习的任务状态原型记忆进行优化。优化后的状态被转换为任务条件的专家 logits,并与所有任务共享的局部专家库中的 token 级 logits 结合,随后进行稀疏的 top-$k$ 选择。一个独立的无任务残差库提供了一个共同的适应路径,并在任务特定预测之前将这两条路径添加到骨干特征中。我们在 NYUD-v2 和 PASCAL-Context 上指定了一个匹配的评估协议,使用 SAM 3 和 ViT-L 骨干网络来测量预测质量、计算成本,以及任务状态条件、原型检索和稀疏路由的贡献。本工作草稿中的数值记录早于这一标准实现,必须在支持实证声明之前重新生成。
cs.CV / 34 / 2608.28080
Cyc3D: Evaluating Cyclic Structural Stability and Asset Usability in Image-to-3D Generation
Cyc3D:评估图像到3D生成中的循环结构稳定性和资产可用性
Abstract
Image-conditioned 3D generation has advanced rapidly, yet existing evaluation protocols largely judge rendered-view plausibility and semantic alignment, overlooking whether a generator forms a stable 3D interpretation and produces assets usable in graphics pipelines. We introduce Cyc3D, a multidimensional benchmark that evaluates image-to-3D generation along two complementary axes: Cross-View Object Consistency and Representation Quality. At the asset level, Cyc3D measures whether object identity remains semantically coherent across rendered viewpoints. At the model level, we propose View-Cycle Structural Consistency, a closed-loop render-regenerate-align protocol that repeatedly re-observes a generated asset from novel views and quantifies geometric, perceptual, and semantic drift across generations. To assess native asset usability beyond rendered appearance, Cyc3D further evaluates geometric structure, reference-image fidelity, mesh discretization and efficiency, and UV parameterization quality. Together, these diagnostics expose failures obscured by a single perceptual score and provide interpretable evidence of both model instability and representation defects. Experiments on five representative image-to-3D systems show that closed-source feed-forward models consistently outperform open-source optimization-based baselines in geometric fidelity, mesh quality, and cycle stability. Nevertheless, even the strongest methods achieve cycle-stability scores below 48, revealing a persistent gap between visually plausible generation and robust 3D object understanding.
Chinese Translation
图像条件下的3D生成技术发展迅速,但现有的评估协议主要判断渲染视图的可信度和语义一致性,忽视了生成器是否形成稳定的3D解释以及是否生成可在图形管道中使用的资产。我们提出了Cyc3D,这是一个多维基准,沿着两个互补的轴线评估图像到3D生成:视角间对象一致性和表现质量。在资产层面,Cyc3D测量对象身份在渲染视点之间是否保持语义一致。在模型层面,我们提出了视角循环结构一致性(View-Cycle Structural Consistency),这是一种闭环渲染-再生-对齐协议,它反复从新视角重新观察生成的资产,并量化生成过程中的几何、感知和语义漂移。为了评估原生资产的可用性超越渲染外观,Cyc3D进一步评估几何结构、参考图像的保真度、网格离散化效率以及UV参数化质量。综合这些诊断工具揭示了单一感知评分所掩盖的失败,并提供了模型不稳定性和表现缺陷的可解释证据。在对五个代表性的图像到3D系统进行的实验中,闭源前馈模型在几何保真度、网格质量和循环稳定性方面始终优于基于开放源代码的优化基线。然而,即使是最强的方法,其循环稳定性评分也低于48,揭示了视觉上可信的生成与稳健的3D对象理解之间的持续差距。
cs.CV / 35 / 2608.28082
Attribute Token Arithmetic: Disentangled and Continuous Semantic Control for Visual Autoregressive Models
属性标记算术:视觉自回归模型的解耦与连续语义控制
Abstract
Autoregressive text-to-image generation has recently achieved remarkable progress, offering high-fidelity synthesis via a unified generative framework. However, fine-grained semantic control remains challenging due to the attribute entanglement and the misalignment between textual and fine-grained visual representations. In this paper, we introduce Attribute Token Arithmetic (ATA), a method that enables disentangled and continuous attribute control in visual autoregressive modelling. Inspired by the vector arithmetic property observed in word embeddings, ATA identifies semantic directions corresponding to visual attributes (e.g., aging, fatness, emotion) directly within the pretrained autoregressive latent space. These directions are learned from a single reference image, without model retraining or large-scale supervision. During generation, attributes can be continuously adjusted and compositionally combined through simple arithmetic operations with other attribute tokens. Extensive experiments demonstrate that ATA achieves identity-preserving, fine-grained, and multi-attribute adjustment, outperforming existing autoregressive editing baselines in controllability, generality, and computational efficiency. Our code will be available at https://github.com/Madaoer/ATA.
Chinese Translation
自回归文本到图像生成最近取得了显著进展,通过统一的生成框架提供高保真合成。然而,由于属性纠缠以及文本与细粒度视觉表示之间的不对齐,细粒度语义控制仍然具有挑战性。本文提出了属性标记算术(Attribute Token Arithmetic, ATA),一种能够在视觉自回归建模中实现解耦和连续属性控制的方法。ATA受到词嵌入中观察到的向量算术特性的启发,直接在预训练的自回归潜在空间中识别与视觉属性(如衰老、肥胖、情感)对应的语义方向。这些方向是从单个参考图像中学习的,无需模型重训练或大规模监督。在生成过程中,属性可以通过与其他属性标记的简单算术操作进行连续调整和组合。大量实验表明,ATA在保持身份、细粒度和多属性调整方面表现优异,超越了现有自回归编辑基线在可控性、普适性和计算效率方面的表现。我们的代码将发布在 https://github.com/Madaoer/ATA。
cs.CV / 36 / 2608.28096
Ex-Sim(3)-Reg: 2D-3D Correspondence Pruning via Extended Sim(3) Registration
Ex-Sim(3)-Reg:通过扩展 Sim(3) 配准进行 2D-3D 对应关系剪枝
Abstract
Learning-based image-to-point-cloud (I2P) registration has garnered increasing attention in recent years. Nevertheless, existing methods still struggle with severe outliers under challenging scenarios with unseen, low-inlier, or distorted cases. A fast and robust 2D-3D correspondence pruning method is therefore highly desirable. Recently, a promising scheme lifts 2D-3D correspondences to 3D-3D correspondences using depth priors, casting correspondence pruning as a Sim(3) registration problem. However, depth priors estimated from monocular images are inherently noisy, which undermines the reliability of this scheme. In this paper, to explicitly model non-negligible depth noise, we reformulate correspondence pruning as an extended Sim(3) registration problem and propose a simple yet effective pruning algorithm termed Ex-Sim(3)-Reg. We further provide a theoretical analysis to justify the effectiveness of our method. Extensive experiments on the 7-Scenes, RGBD-V2, ScanNet, and TUM datasets demonstrate that Ex-Sim(3)-Reg achieves up to \textbf{24.7\% improvement} in registration recall over state-of-the-art baseline methods. Code is released at github.com/anpei96/ex-sim3-demo
Chinese Translation
基于学习的图像到点云(I2P)配准近年来受到越来越多的关注。然而,现有方法在面对未见过的、低内点或失真的情况时,仍然难以处理严重的异常值。因此,急需一种快速且稳健的 2D-3D 对应关系剪枝方法。最近,一种有前景的方案利用深度先验将 2D-3D 对应关系提升为 3D-3D 对应关系,将对应关系剪枝视为 Sim(3) 配准问题。然而,从单目图像估计的深度先验本质上是噪声,这削弱了该方案的可靠性。在本文中,为了明确建模不可忽视的深度噪声,我们将对应关系剪枝重新表述为扩展 Sim(3) 配准问题,并提出了一种简单而有效的剪枝算法,称为 Ex-Sim(3)-Reg。我们进一步提供了理论分析,以证明我们方法的有效性。在 7-Scenes、RGBD-V2、ScanNet 和 TUM 数据集上的大量实验表明,Ex-Sim(3)-Reg 在配准召回率上比最先进的基线方法提高了高达 24.7%。代码已发布在 github.com/anpei96/ex-sim3-demo
cs.CV / 37 / 2608.28138
Token-Budget Distillation: Transferring Full-Token Semantics to Compressed Video Vision-Language Models
令牌预算蒸馏:将全令牌语义转移到压缩视频视觉语言模型
Abstract
Adapting video vision-language models (VLMs) is computationally expensive because video inputs produce a large number of visual tokens, making both fine-tuning and inference costly. Although visual token compression can reduce this overhead, direct adaptation on compressed inputs often causes semantic drift and noticeable performance degradation. We present Token-Budget Distillation (TBD), a parameter-efficient fine-tuning framework for adapting video VLMs under a fixed token budget. TBD freezes the pretrained backbone, updates only LoRA adapters, and integrates FlashVID-based visual token compression into the video pathway. To preserve full-token semantics under compression, TBD employs a dual-path teacher-student design, where a full-token teacher provides stable supervision and a compressed student is optimized with task loss, answer-region KL distillation, GT-anchored margin distillation, and reliability-aware KD control. This design enables the student to recover the semantic behavior of the full-token model while remaining efficient under aggressive token reduction. We evaluate TBD on three video VLM backbones, including LLaVA-Video, LLaVA-OneVision, and Qwen3-VL-8B-Instruct, across four video understanding benchmarks. TBD consistently outperforms compression-only baselines under both moderate and aggressive compression. On LLaVA-Video at retention ratio R = 10 percent, TBD preserves 97.0 percent of the Vanilla model's average accuracy; on LLaVA-OneVision at R = 10 percent, it achieves an average score of 58.4 and matches 100.0 percent relative accuracy.
Chinese Translation
适应视频视觉语言模型(VLMs)计算成本高,因为视频输入会产生大量视觉令牌,这使得微调和推理的成本都很高。尽管视觉令牌压缩可以减少这种开销,但在压缩输入上进行直接适应往往会导致语义漂移和显著的性能下降。我们提出了令牌预算蒸馏(Token-Budget Distillation, TBD),这是一种在固定令牌预算下适应视频 VLMs 的参数高效微调框架。TBD 冻结预训练的主干网络,仅更新 LoRA 适配器,并将基于 FlashVID 的视觉令牌压缩集成到视频路径中。为了在压缩下保留全令牌语义,TBD 采用了双路径教师-学生设计,其中全令牌教师提供稳定的监督,而压缩学生则通过任务损失、答案区域 KL 蒸馏、GT 锚定边际蒸馏和可靠性感知的知识蒸馏控制进行优化。该设计使学生能够在激进的令牌减少下恢复全令牌模型的语义行为,同时保持高效。我们在三个视频 VLM 主干上评估了 TBD,包括 LLaVA-Video、LLaVA-OneVision 和 Qwen3-VL-8B-Instruct,涵盖四个视频理解基准。TBD 在中等和激进压缩下始终优于仅压缩的基线。在 LLaVA-Video 上,保留率 R = 10% 时,TBD 保留了 Vanilla 模型平均准确率的 97.0%;在 LLaVA-OneVision 上,R = 10% 时,TBD 达到了 58.4 的平均得分,并匹配了 100.0% 的相对准确率。
cs.CV / 38 / 2608.28145
Dual-Stream Semantic Guidance with Prototype Anchor Calibration for Source-Fully-Free Adaptation of Vision-Language Models
基于原型锚定校准的双流语义引导用于无源完全自由适应的视觉-语言模型
Abstract
Source-Fully-Free Domain Adaptation (SFF-DA) has emerged as a strategic paradigm to adapt Vision-Language Models (VLMs) without any access to source data or task-specific source models. However, we identify a critical Dual Semantic Drift that hinders this process: static drift arising from the rigidity of fixed class embeddings, and dynamic drift stemming from the divergence of generated captions, causing severe semantic misalignment that intensifies the stability-plasticity dilemma. To address this, we propose DSSG (Dual-Stream Semantic Guidance), an end-to-end framework that reconciles fine-grained plasticity with global stability. Our core contribution is the Dual Semantic Guidance (DSG) module, which integrates a caption stream for domain-specific knowledge with a class-anchor stream to anchor global categorical consistency. Furthermore, a Dynamic Cross-Modal Knowledge Distillation (CMKD) module is introduced to leverage the evolving teacher distribution for calibrating teacher-student consistency. Building upon DSSG, we further introduce Prototype Anchor Calibration (PAC), yielding DSSG-PAC, which periodically calibrates prototype anchors and caches them until the next calibration. This design reduces redundant text-side computation while preserving the adaptability of class guidance to the evolving text space. We further establish SFF-DA risk bounds that relate student risk to semantic-teacher quality and teacher--student discrepancy. Extensive experiments demonstrate that DSSG consistently outperforms current state-of-the-art methods across multiple benchmarks, while DSSG-PAC largely preserves its adaptation performance with 18.9% lower total adaptation time. The code is available at https://github.com/mrmenand/DSSG.
Chinese Translation
无源完全自由领域适应(SFF-DA)已成为一种战略范式,用于在不接触源数据或特定任务源模型的情况下适应视觉-语言模型(VLMs)。然而,我们识别出一个关键的双重语义漂移,阻碍了这一过程:由于固定类别嵌入的刚性而产生的静态漂移,以及由于生成的标题之间的差异而导致的动态漂移,造成严重的语义不一致,从而加剧了稳定性-可塑性困境。为了解决这个问题,我们提出了DSSG(双流语义引导),这是一个端到端的框架,调和了细粒度的可塑性与全局稳定性。我们的核心贡献是双重语义引导(DSG)模块,它将用于领域特定知识的标题流与用于锚定全局类别一致性的类别锚流相结合。此外,我们引入了动态跨模态知识蒸馏(CMKD)模块,以利用不断演变的教师分布来校准教师-学生一致性。在DSSG的基础上,我们进一步引入了原型锚定校准(PAC),形成DSSG-PAC,该方法定期校准原型锚并缓存它们,直到下次校准。该设计减少了冗余的文本侧计算,同时保持了类别引导对不断演变的文本空间的适应性。我们进一步建立了SFF-DA风险界限,将学生风险与语义教师质量和教师-学生差异相关联。大量实验表明,DSSG在多个基准测试中始终优于当前最先进的方法,而DSSG-PAC在总适应时间降低18.9%的情况下,基本保持了其适应性能。代码可在https://github.com/mrmenand/DSSG获取。
cs.CV / 39 / 2608.28161
Empowering Local Agriculture: A Deep Learning-Powered Web System for Identifying Bangladeshi Mango Varieties
赋能地方农业:基于深度学习的网络系统用于识别孟加拉国芒果品种
Abstract
Mango variety identification in Bangladesh is challenging because closely related cultivars can have similar visual characteristics and images are often captured under varying real-world conditions. This work presents a deep learning-based web system for automatic identification of Bangladeshi mango varieties. We collected 2,013 high-quality mango images (3024x4032 pixels) from local markets and farms and organized them into nine classes, combining Bari-4 and Bari-7 as a single Bari class. The dataset was divided into training (70%), validation (15%), and test (15%) sets, with image augmentation applied to improve model generalization. Three pretrained CNN architectures, ResNet18, ResNet50, and EfficientNetB0, were fine-tuned under consistent training settings. EfficientNetB0 achieved the best performance, obtaining 98.01% validation accuracy and 97.36% test accuracy, compared with 86.47% and 78.55% test accuracy for ResNet18 and ResNet50, respectively. Class-wise F1-scores for EfficientNetB0 ranged from 0.93 to 0.99, while the Bari class achieved an F1-score of 0.97. The selected EfficientNetB0 model has approximately 4 million parameters, making it suitable for lightweight deployment. We integrated the model into a Streamlit web application that enables users to upload a mango image and receive a predicted variety with class probabilities. The system provides an accessible, practical tool for mango identification and demonstrates the potential of deep learning for supporting agricultural applications in Bangladesh.
Chinese Translation
在孟加拉国,芒果品种的识别面临挑战,因为密切相关的品种可能具有相似的视觉特征,并且图像常常在不同的现实条件下捕获。本研究提出了一种基于深度学习的网络系统,用于自动识别孟加拉国的芒果品种。我们从当地市场和农场收集了2013张高质量的芒果图像(3024x4032像素),并将其组织为九个类别,将Bari-4和Bari-7合并为一个Bari类别。数据集分为训练集(70%)、验证集(15%)和测试集(15%),并应用图像增强以提高模型的泛化能力。我们对三种预训练的卷积神经网络(CNN)架构进行了微调,分别是ResNet18、ResNet50和EfficientNetB0,均在一致的训练设置下进行。EfficientNetB0表现最佳,获得了98.01%的验证准确率和97.36%的测试准确率,而ResNet18和ResNet50的测试准确率分别为86.47%和78.55%。EfficientNetB0的类别F1分数范围为0.93到0.99,而Bari类别的F1分数为0.97。所选的EfficientNetB0模型约有400万个参数,适合轻量级部署。我们将该模型集成到一个Streamlit网络应用程序中,用户可以上传芒果图像并获得预测的品种及其类别概率。该系统提供了一个可访问的、实用的芒果识别工具,并展示了深度学习在支持孟加拉国农业应用方面的潜力。
cs.CV / 40 / 2608.28174
Manifold4D: Denoising on Point Cloud Rendered Manifolds for Video Re-shooting
Manifold4D:基于点云渲染流形的去噪技术用于视频重拍
Abstract
Video re-shooting re-renders a monocular video of a dynamic scene along a user-specified camera trajectory, and the dominant recipe supplies the target geometry explicitly: per-frame depth lifts the source video into a 4D point cloud, which is rasterized along the trajectory into a point cloud render. Because the render and the source video are both handed to the network as visual conditions, they compete at every denoising step, leaving the model with a trust dilemma --- how much of the render to believe --- which can degrade trajectory control or visual quality on data outside the training distribution. We argue that a render already pixel-aligned with the target view does not need to be supplied as an explicit conditioning stream at all. We propose MANIFOLD4D, which injects the render directly into the initial noise of flow matching, so that generation no longer departs from standard Gaussian noise but from a new noise manifold carrying geometric information, leaving the source video as the only visual condition. The render is thus used exactly once, and the network is never asked to learn how to read it; in subsequent denoising steps the model can focus on the source video. On our DAVIS-Traj benchmark and on the Vista4D evaluation set, MANIFOLD4D attains the best camera-control accuracy on every metric, lowering rotation error by 25% and 27% and translation error by up to 32% over the strongest baseline, while matching it in video fidelity and leading on real-world novel-view photometric quality. In a user study, our method achieves clear advantages in trajectory following and dynamic consistency. The gap widens as the yaw amplitude grows past the training range, and the model still recovers correct dynamic motion from the source video when the render is deliberately corrupted, confirming that the geometric prior guides generation without overriding it.
Chinese Translation
视频重拍沿着用户指定的相机轨迹重新渲染动态场景的单目视频,主要方法明确提供目标几何信息:逐帧深度将源视频提升为4D点云,并沿轨迹光栅化为点云渲染。由于渲染和源视频都作为视觉条件输入到网络中,它们在每个去噪步骤中相互竞争,导致模型面临信任困境——需要相信多少渲染——这可能会降低轨迹控制或在训练分布之外的数据上影响视觉质量。我们认为,已经与目标视图像素对齐的渲染根本不需要作为显式条件流提供。我们提出了MANIFOLD4D,它将渲染直接注入到流匹配的初始噪声中,使得生成不再偏离标准高斯噪声,而是从携带几何信息的新噪声流形中出发,源视频成为唯一的视觉条件。因此,渲染仅使用一次,网络不再被要求学习如何读取它;在后续的去噪步骤中,模型可以专注于源视频。在我们的DAVIS-Traj基准和Vista4D评估集上,MANIFOLD4D在每个指标上都达到了最佳的相机控制精度,旋转误差降低了25%和27%,平移误差降低了高达32%,同时在视频保真度上与最强基线持平,并在真实世界的新视角光度质量上领先。在用户研究中,我们的方法在轨迹跟踪和动态一致性方面表现出明显优势。随着偏航幅度超过训练范围,差距进一步扩大,当渲染故意损坏时,模型仍能从源视频中恢复正确的动态运动,确认几何先验在生成过程中起到指导作用而不覆盖生成。
cs.CV / 41 / 2608.28191
EXPOSE: Explainable and Domain-Robust Embeddings from Pathology Vision Foundation Models using Sparse Autoencoders
EXPOSE:基于稀疏自编码器的病理视觉基础模型的可解释和领域鲁棒嵌入
Abstract
Vision Foundation Models (VFMs) are widely used in computational pathology but remain sensitive to domain shifts arising from variations in staining, tissue preparation, and scanner hardware. A key limitation is that VFM embeddings entangle biological with domain-specific information, hindering cross-domain generalization. We propose Explainable Probing of Cross-Domain Sparse Embeddings (EXPOSE), a framework that uses Sparse Autoencoders (SAEs) as an explainable bottleneck to identify and suppress domain-specific components in VFM embeddings. We train a sparse representation of VFM features, use a linear classifier to identify domain-specific latent dimensions, and mask these features prior to downstream relapse prediction without retraining the backbone model. Experiments on a large prostate cancer dataset with multiple acquisition domains show that SAE features capture both domain- and task-specific information, which are partially disentangled in the latent space. Removing domain-specific features improves cross-domain performance and increases embedding robustness as measured by the Domain Robustness Index (DoRI). Code is available at https://github.com/imsb-uke/expose .
Chinese Translation
视觉基础模型(VFMs)在计算病理学中被广泛应用,但仍然对由于染色、组织制备和扫描仪硬件变化引起的领域转移敏感。一个关键限制是,VFM 嵌入将生物信息与领域特定信息纠缠在一起,阻碍了跨领域的泛化。我们提出了可解释的跨领域稀疏嵌入探测框架(EXPOSE),该框架使用稀疏自编码器(SAEs)作为可解释的瓶颈,以识别和抑制 VFM 嵌入中的领域特定组件。我们训练 VFM 特征的稀疏表示,使用线性分类器识别领域特定的潜在维度,并在下游复发预测之前对这些特征进行掩蔽,而无需重新训练主干模型。在一个包含多个采集领域的大型前列腺癌数据集上的实验表明,SAE 特征同时捕获领域和任务特定的信息,这些信息在潜在空间中部分解缠。去除领域特定特征提高了跨领域性能,并根据领域鲁棒性指数(DoRI)增加了嵌入的鲁棒性。代码可在 https://github.com/imsb-uke/expose 获取。
cs.CV / 42 / 2608.28192
Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
在视频中定位任何事物:重新思考高效生成的时空视频定位
Abstract
Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the target entity throughout that interval. Existing multimodal large language models typically serialize dense localization trajectories autoregressively, causing decoding latency to grow with tube length and allowing localization errors to propagate across time. We introduce Parallel Tube Decoding (PTD), a generative formulation that decomposes grounding into a temporal block followed by time-conditioned spatial blocks decoded simultaneously. This removes both token-level and trajectory-level dependencies, reducing the sequential decoding depth to a fixed $1 + 1$ rounds, independent of tube length. To enable parallel spatial generation, we introduce Decoupled Block Attention, which preserves access to shared video-query context while eliminating cross-box dependencies, together with localization-aware policy optimization for temporal boundaries and spatial geometry. On VidSTG, PTD reduces Tube Completion Latency by 79x and increases spatial decoding throughput by 92x over standard autoregressive decoding, while also improving grounding accuracy. With a compact 4B backbone, our model performs favorably well on VidSTG and HC-STVG, and generalizes zero-shot to temporal grounding, grounded VideoQA, and referring video object tracking. Our results show parallel tube generation is an efficient and effective alternative to autoregressive localization in videos.
Chinese Translation
时空视频定位(STVG)要求模型识别所提及事件发生的时间,并在该时间段内定位目标实体。现有的多模态大型语言模型通常以自回归方式序列化密集的定位轨迹,导致解码延迟随着管道长度的增加而增长,并使定位错误在时间上传播。我们提出了并行管道解码(Parallel Tube Decoding, PTD),这是一种生成性表述,将定位分解为一个时间块,随后是同时解码的时间条件空间块。这消除了标记级和轨迹级的依赖关系,将顺序解码深度减少到固定的 $1 + 1$ 轮,与管道长度无关。为了实现并行空间生成,我们引入了解耦块注意力(Decoupled Block Attention),该方法在消除跨框依赖的同时,保留对共享视频查询上下文的访问,并结合了针对时间边界和空间几何的定位感知策略优化。在 VidSTG 数据集上,PTD 将管道完成延迟减少了 79 倍,并将空间解码吞吐量提高了 92 倍,相较于标准自回归解码,同时也提高了定位准确性。我们的模型在紧凑的 4B 主干上,在 VidSTG 和 HC-STVG 上表现良好,并在时间定位、基于定位的视频问答和引用视频对象跟踪等任务上实现了零样本泛化。我们的结果表明,平行管道生成是视频中自回归定位的高效且有效的替代方案。
cs.CV / 43 / 2608.28195
UniLipi: A Unified Multi-Script OCR for Historical Indic Manuscripts
UniLipi:一个统一的多脚本历史印度手稿光学字符识别系统
Abstract
Optical character recognition (OCR) for handwritten Indic manuscripts is essential for large-scale digitization and computational access to manuscript heritage. However, existing approaches are typically developed for one script at a time and require substantial script-specific customization. This limits scalability and practical deployment across diverse collections. We present UniLipi, a unified multi-script OCR model for handwritten Indic manuscripts trained jointly across 13 Indic scripts within a single framework. UniLipi directly handles realistic manuscript conditions, including extreme variation in line geometry, large variation in line length, and partial interruptions caused by non-textual manuscript entities such as holes, stains, or pictorial illustrations. To operate effectively under ultra low-resource conditions, the model leverages script-aware synthetic manuscript data generation, substantially reducing reliance on large volumes of real annotated data. Beyond historical manuscripts, we show that UniLipi serves as an effective foundational pretrained model. Specifically, its learned representations enable good OCR performance for contemporary Indic handwriting and extend to several non-Indic scripts, including Tibetan, Italian, Latin, and Chinese scripts. In addition to transcription, UniLipi predicts script identity and per-line native character counts, supporting practical manuscript cataloging workflows.
Chinese Translation
手写印度手稿的光学字符识别(OCR)对于大规模数字化和手稿遗产的计算访问至关重要。然而,现有的方法通常是针对单一脚本开发的,并且需要大量特定于脚本的定制。这限制了在不同收藏中的可扩展性和实际部署。我们提出了UniLipi,一个统一的多脚本OCR模型,针对手写印度手稿在一个框架内共同训练了13种印度脚本。UniLipi直接处理现实的手稿条件,包括线条几何形状的极端变化、线条长度的巨大差异,以及由于非文本手稿实体(如孔、污点或图示插图)造成的部分中断。为了在超低资源条件下有效运行,该模型利用了脚本感知的合成手稿数据生成,大大减少了对大量真实标注数据的依赖。除了历史手稿之外,我们还展示了UniLipi作为一个有效的基础预训练模型的作用。具体而言,其学习的表示使得对当代印度手写体的OCR性能良好,并扩展到包括藏文、意大利文、拉丁文和中文在内的几种非印度脚本。除了转录,UniLipi还预测脚本身份和每行的本地字符计数,支持实际的手稿编目工作流程。
cs.CV / 44 / 2608.28205
Cut-ViT: Task-Specific Model Pruning via Gram Anchoring Subspace Consistency
Cut-ViT:通过Gram锚定子空间一致性进行任务特定模型剪枝
Abstract
Pruning visual foundation models has attracted considerable attention. However, existing methods focus on rigid point-to-point token alignment on a single dataset for pruning, suffering from two limitations: i) robustness degradation, and ii) task-specificity deficiency. To address these limitations, we propose a task-specific pruning pipeline, named Cut-ViT. Specifically, we first construct gram anchoring matrices from both spatial and semantic perspectives, and perform the subspace decomposition to extract the corresponding subspace bases. Basis-agnostic and residual constraints are then adopted to align the gram subspaces between the native and pruned DINOv3 models along spatial and channel dimensions, enabling subnetworks to inherit robust feature representations of native DINOv3. Furthermore, we design spectral entropy adaptation, which quantifies the information density of feature manifolds along spatial and channel dimensions, thereby adapting the pruning objective to specific downstream tasks. Experiments show that Cut-ViT requires approximately one minute on a single A100 GPU to obtain subnetworks at various sparsity levels, using only 20.9% of the time and 45.5% of the GPU memory compared with previous methods, while achieving SOTA performance on six tasks across nine datasets.
Chinese Translation
剪枝视觉基础模型引起了广泛关注。然而,现有方法集中于在单一数据集上进行刚性的一对一标记对齐,面临两个局限性:i) 鲁棒性下降,ii) 任务特定性不足。为了解决这些问题,我们提出了一种任务特定的剪枝流程,称为Cut-ViT。具体而言,我们首先从空间和语义两个角度构建Gram锚定矩阵,并进行子空间分解以提取相应的子空间基。然后,采用与基无关的约束和残差约束,在空间和通道维度上对齐原始和剪枝后的DINOv3模型之间的Gram子空间,使得子网络能够继承原始DINOv3的鲁棒特征表示。此外,我们设计了谱熵适应,量化特征流形在空间和通道维度上的信息密度,从而将剪枝目标适应于特定的下游任务。实验表明,Cut-ViT在单个A100 GPU上大约需要一分钟即可获得不同稀疏级别的子网络,仅使用了之前方法的20.9%的时间和45.5%的GPU内存,同时在九个数据集上的六个任务上实现了SOTA性能。
cs.CV / 45 / 2608.28206
NumBench: Diagnosing Counting Failures in Text-to-Image Models
NumBench:诊断文本到图像模型中的计数失败
Abstract
Text-to-image (T2I) models often generate the wrong number of objects, yet existing benchmarks are too small or weakly controlled to explain why. We introduce \textbf{NumBench}, a benchmark of 640{,}000 prompts spanning 1{,}600 categories and counts from 1 to 100. Its factorial design varies object composition, spatial guidance, and appearance conditions while balancing counts and category exposure. We also develop a process model in which requested instances compete for a finite set of resolvable image regions. The model predicts a near-quadratic collision deficit at low occupancy and shows how coordinated placement reduces it. For scalable evaluation, we propose the Confidence-Weighted Numeric Precision Score (\cwnps), which aggregates three calibrated detectors and discounts uncertain proposals. Across five commercial systems, two open models, and two specialized counting methods, performance declines sharply with requested count; all evaluated methods are weak above 50 objects. Count range has the largest measured effect, followed by layout and composition. Grid guidance is strongest among guided layouts, consistent with the coordination prediction, although the analysis does not establish collision as the sole cause. A 14{,}400-image human study supports automated evaluation through count 50, while results on 243 natural-language prompts show transfer beyond NumBench templates.
Chinese Translation
文本到图像(T2I)模型经常生成错误数量的对象,但现有基准测试规模过小或控制不严,无法解释原因。我们引入了 extbf{NumBench},这是一个包含640,000个提示的基准,涵盖1,600个类别和从1到100的计数。其因子设计变化了对象组成、空间引导和外观条件,同时平衡了计数和类别曝光。我们还开发了一个过程模型,其中请求的实例在有限的可解析图像区域中竞争。该模型预测在低占用率下接近二次的冲突缺失,并展示了协调放置如何减少这一缺失。为了可扩展的评估,我们提出了置信加权数值精度分数(Confidence-Weighted Numeric Precision Score, extcwnps),该分数聚合了三个经过校准的检测器,并对不确定的提案进行了折扣。在五个商业系统、两个开放模型和两种专业计数方法中,性能随着请求计数的增加而急剧下降;所有评估的方法在50个对象以上表现较弱。计数范围的影响最大,其次是布局和组成。在引导布局中,网格引导是最强的,这与协调预测一致,尽管分析并未确立冲突为唯一原因。一项包含14,400张图像的人类研究支持通过计数50的自动评估,而在243个自然语言提示上的结果显示超出了NumBench模板的迁移。
cs.CV / 46 / 2608.28207
Explainable Diabetic Retinopathy Classification Using Vision Foundation Models
可解释的糖尿病视网膜病变分类方法基于视觉基础模型
Abstract
Diabetic retinopathy (DR) is a major cause of preventable blindness, creating a need for accurate and trustworthy automated screening. This study investigates an explainable DR classification framework using vision foundation models and multiple transfer learning strategies. Three backbones, DINOv2, CLIP, and Vision Transformer (ViT), were evaluated using full fine-tuning, linear probing, and Low-Rank Adaptation (LoRA). Models were trained and internally evaluated on the ODIR dataset and externally evaluated on APTOS to assess generalization. DINOv2-LoRA achieved the highest internal AUROC of 0.758, while DINOv2 full fine-tuning and ViT full fine-tuning achieved the highest external AUROC of 0.920. Calibration was further assessed using reliability analysis after isotonic regression. For explainability, Grad-CAM and HiResCAM were evaluated against expert-annotated lesion masks from the IDRiD dataset using Dice, Intersection over Union (IoU), and Pointing Game metrics. The results demonstrate that foundation models, particularly DINOv2, can provide strong predictive performance, while LoRA offers a parameter-efficient alternative to full fine-tuning. Quantitative evaluation of explanation maps further supports the assessment of whether model attention corresponds to clinically relevant retinal lesions.
Chinese Translation
糖尿病视网膜病变(DR)是可预防失明的主要原因,这就需要准确且可信赖的自动筛查。本研究探讨了一种使用视觉基础模型和多种迁移学习策略的可解释DR分类框架。评估了三种主干网络:DINOv2、CLIP和视觉变换器(ViT),采用了全量微调、线性探测和低秩适应(LoRA)等方法。模型在ODIR数据集上进行训练和内部评估,并在APTOS数据集上进行外部评估,以评估其泛化能力。DINOv2-LoRA在内部AUROC中达到了最高的0.758,而DINOv2全量微调和ViT全量微调在外部AUROC中达到了最高的0.920。通过同调回归后使用可靠性分析进一步评估了模型的校准性。为了实现可解释性,Grad-CAM和HiResCAM与来自IDRiD数据集的专家标注病变掩膜进行了评估,使用了Dice、交并比(IoU)和指向游戏指标。结果表明,基础模型,特别是DINOv2,能够提供强大的预测性能,而LoRA则提供了一个参数高效的全量微调替代方案。对解释图的定量评估进一步支持了模型注意力是否与临床相关的视网膜病变相对应的评估。
cs.CV / 47 / 2608.28216
WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes
WALDO:在杂乱场景中进行一次性示例条件的物体检测
Abstract
Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is absent, large vision-language models usually address this task. We ask whether the same capability is available far more cheaply, from representations already learned by a world-model pretraining objective. We present WALDO, a one-shot exemplar- and language-conditioned detection head with 3.4M trainable parameters that reads frozen V-JEPA 2.1 features to jointly predict object localization and target presence, with no gradient on the backbone. Because exemplar-conditioned supervision is scarce, we synthesize training episodes from instance annotations, mining exemplars from ground-truth boxes and constructing absence cases that exclude the referenced instance while leaving same-category distractors in view. This is easy to get wrong: in the obvious implementation, crop size alone predicts the label, and a head trained on it reaches 0.9998 absence AUROC without ever consulting the exemplar, and we report the negative controls that close the shortcut. On 35 held-out cluttered scenes, WALDO achieves a 0.461 catalogue AP@50, compared to 0.306 for a prompted Grounding DINO baseline under an identical scorer. Substituting DINOv3 for V-JEPA under a matched 576-token grid drops within-category absence AUROC from 0.880 to 0.726 and instance AP@50 from 0.201 to 0.141, isolating the pretraining objective rather than input resolution as the source of the gain. Instance-level Success@1, however, reaches only 0.190 against a 0.190 category-chance floor: world-model features transfer to localization precision and absence detection but not to instance identity.
Chinese Translation
使用单一参考图像和简短描述在杂乱场景中定位特定物体实例,并在该实例缺失时进行报告,通常由大型视觉语言模型来完成这一任务。我们探讨是否可以以更低的成本,从已经通过世界模型预训练目标学习到的表示中获得相同的能力。我们提出了WALDO,一种具有340万可训练参数的一次性示例和语言条件检测头,它读取冻结的V-JEPA 2.1特征,联合预测物体定位和目标存在性,而不对主干网络进行梯度更新。由于示例条件监督稀缺,我们从实例注释中合成训练情景,从真实框中挖掘示例,并构建排除参考实例的缺失案例,同时保留同类干扰物。这很容易出错:在明显的实现中,仅凭裁剪大小就能预测标签,而在此基础上训练的检测头在未咨询示例的情况下达到0.9998的缺失AUROC,我们报告了关闭这一捷径的负控制。在35个保留的杂乱场景中,WALDO的目录AP@50达到了0.461,而在相同评分者下,提示的Grounding DINO基线为0.306。将DINOv3替换为V-JEPA,并在匹配的576-token网格下,类别内缺失AUROC从0.880降至0.726,实例AP@50从0.201降至0.141,表明预训练目标而非输入分辨率是增益的来源。然而,实例级的成功率@1仅达到0.190,正好与0.190的类别机会底线持平:世界模型特征能够转移到定位精度和缺失检测,但无法转移到实例身份。
cs.CV / 48 / 2608.28218
Focus Where It Counts: A Salience-Driven Vision-Language Model for Low Vision Assistance
聚焦重要信息:一种基于显著性的视觉-语言模型用于低视力辅助
Abstract
Vision-language models (VLMs) are rapidly progressing and offer promising capabilities for assistive technologies supporting persons with blindness or low vision. However, existing VLMs are primarily designed for general-purpose captioning and do not explicitly model human perceptual priorities, thereby limiting their ability to emphasize the most relevant information in a scene. To address this gap, we propose a salience-driven captioning framework that prioritizes scene elements according to their importance for human-centered assistance. We curate three salience-aware datasets, namely, Salience COCO, Salience Flickr, and Salience VizWiz, with object-level salience annotations designed to reflect the visual information most relevant to low vision users across different environments. Building on these datasets, we introduce Salience-LLaVA, a salience-aware VLM that incorporates salience cues to generate captions in which important elements are mentioned in the order of importance. Our work makes four main contributions. We build salience-aware datasets verified by low vision participants, propose Salience-LLaVA to describe objects in the order of importance, introduce SCMI to evaluate ordering accuracy, and deploy the system on assistive glasses to demonstrate real-world practicality. Code and datasets are available at: https://github.com/topo-focus/Topofocus
Chinese Translation
视觉-语言模型(VLMs)正在迅速发展,并为支持盲人或低视力人士的辅助技术提供了有前景的能力。然而,现有的VLMs主要是为通用的图像描述而设计,并未明确建模人类的感知优先级,从而限制了它们在场景中强调最相关信息的能力。为了解决这一问题,我们提出了一种基于显著性的图像描述框架,该框架根据场景元素对人本辅助的重要性进行优先排序。我们策划了三个显著性感知数据集,即Salience COCO、Salience Flickr和Salience VizWiz,这些数据集具有对象级显著性注释,旨在反映在不同环境中与低视力用户最相关的视觉信息。在这些数据集的基础上,我们引入了Salience-LLaVA,这是一种显著性感知的VLM,结合显著性线索生成描述,其中重要元素按照重要性顺序提及。我们的工作主要有四个贡献:我们构建了经过低视力参与者验证的显著性感知数据集,提出了Salience-LLaVA以按照重要性描述对象,引入了SCMI来评估排序准确性,并在辅助眼镜上部署该系统以展示其现实应用性。代码和数据集可在以下网址获取:https://github.com/topo-focus/Topofocus
cs.CV / 49 / 2608.28219
RASA: Disentangled Spatial-Motional Priors for Cross-Identity Character Animation
RASA:用于跨身份角色动画的解耦空间-运动先验
Abstract
Cross-identity character animation aims to drive a target identity from a reference image to follow the motion of a source character from a driving video. The core challenge lies in the inherent entanglement of two capabilities: cross-identity spatial mapping (aligning position, scale, and skeletal proportions) and motion control (refining joint articulation, volumetric consistency, and view coherence). We introduce Reference-Aware Structural Alignment (RASA), a framework that disentangles spatial mapping from motion control by injecting structured priors into a Diffusion Transformer (DiT). Our approach has two stages. First, a Spatial Prior Calibrator (SPC) fuses reference identity with driving pose to generate a spatially grounded initial noise latent, ensuring correct positioning, scaling, and alignment with the driving skeleton. Second, an Inherent Motional Guider (IMG) encodes shape-agnostic SMPL articulation parameters into a semantic motion vector beyond appearance-biased 2D keypoints. Injected into intermediate DiT layers, this vector complements the base pose condition for anatomically consistent articulation and view-aware volumetric refinement. We curate CIM-Bench, a high-quality benchmark with rigorous curation, for evaluation. Extensive experiments show RASA significantly outperforms state-of-the-art methods in motion fidelity and visual quality. Our work establishes a new paradigm showing disentangled spatial and motional priors are key to robust character animation. Project page: https://hidream.ai.github.io/RASA/
Chinese Translation
跨身份角色动画旨在驱动目标身份从参考图像中跟随源角色在驱动视频中的运动。核心挑战在于两种能力的固有纠缠:跨身份空间映射(对齐位置、比例和骨骼比例)和运动控制(细化关节关节、体积一致性和视图一致性)。我们提出了参考感知结构对齐(Reference-Aware Structural Alignment,RASA),这是一个通过将结构化先验注入扩散变换器(Diffusion Transformer,DiT)来解耦空间映射与运动控制的框架。我们的方法分为两个阶段。首先,空间先验校准器(Spatial Prior Calibrator,SPC)将参考身份与驱动姿势融合,以生成空间基础的初始噪声潜变量,确保正确的位置、比例和与驱动骨骼的对齐。其次,固有运动引导器(Inherent Motional Guider,IMG)将形状无关的SMPL关节参数编码为超越外观偏向的2D关键点的语义运动向量。该向量被注入到中间DiT层中,补充了基础姿势条件,以实现解剖学一致的关节动作和视图感知的体积细化。我们策划了CIM-Bench,这是一个经过严格策划的高质量基准,用于评估。大量实验表明,RASA在运动保真度和视觉质量上显著优于最先进的方法。我们的工作建立了一个新的范式,表明解耦的空间和运动先验是稳健角色动画的关键。项目页面:https://hidream.ai.github.io/RASA/
cs.CV / 50 / 2608.28240
WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild
WilLaGS:用于鲁棒高斯点云渲染的潜在条件3D外观场
Abstract
3D Gaussian Splatting (3DGS) delivers real-time and high-fidelity rendering but remains challenged by unconstrained in-the-wild scenes, where drastic appearance variations and transient objects violate multi-view consistency. Existing methods are fundamentally limited by independent and discrete embeddings that struggle to capture continuous environmental changes or model spatially-varying local illumination. To address these limitations, we propose \textbf{WilLaGS}, a unified framework for robust 3D scene reconstruction and generative appearance synthesis under unconstrained settings. Specifically, we introduce a generative appearance model where a $\beta$-VAE learns a structured and continuous manifold of global appearance. Conditioned on the latent code, we construct a 3D neural appearance field that generates dynamic Tri-Plane features to encode spatially-varying local illumination effects. Furthermore, to suppress transient artifacts, we present a self-supervised perceptual masking mechanism that leverages a Teacher-Student (EMA) architecture to derive a stable scene consensus, robustly identifying inconsistent regions via perceptual discrepancies. Extensive experiments on multiple datasets demonstrate that \textbf{WilLaGS} achieves state-of-the-art performance in reconstruction quality and novel view appearance synthesis, while maintaining real-time rendering efficiency.
Chinese Translation
3D高斯点云渲染(3DGS)提供了实时和高保真的渲染效果,但在不受约束的真实场景中仍面临挑战,其中剧烈的外观变化和瞬态物体破坏了多视角一致性。现有方法在本质上受到独立和离散嵌入的限制,难以捕捉连续的环境变化或建模空间变化的局部照明。为了解决这些限制,我们提出了 extbf{WilLaGS},一个统一的框架,用于在不受约束的环境下进行鲁棒的3D场景重建和生成外观合成。具体而言,我们引入了一种生成外观模型,其中$eta$-VAE学习了全球外观的结构化和连续流形。在潜在编码的条件下,我们构建了一个3D神经外观场,生成动态三平面特征以编码空间变化的局部照明效果。此外,为了抑制瞬态伪影,我们提出了一种自监督感知掩蔽机制,该机制利用教师-学生(EMA)架构推导出稳定的场景共识,通过感知差异稳健地识别不一致区域。在多个数据集上的广泛实验表明, extbf{WilLaGS}在重建质量和新视角外观合成方面达到了最先进的性能,同时保持了实时渲染效率。
cs.CV / 51 / 2608.28247
A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation
针对地球观测变化检测的全面且可信的人工智能方法基准
Abstract
Change detection in Earth observation (EO) is critical for monitoring land surface transformations, yet recent research in the field is constrained by inconsistent evaluation protocols and a narrow focus on predictive accuracy without regard for computational efficiency. To address this, we present a standardized, open-source benchmark for evaluating state-of-the-art (SOTA) deep learning methods for Earth observation change detection. We conduct a comprehensive analysis of ten representative model architectures, ranging from convolutional networks (CNNs) to vision transformers (ViTs), across ten heterogeneous change detection datasets. We rigorously evaluate these models with identical experimental protocols, comparing models trained from scratch against those utilizing pre-trained weights. Furthermore, we evaluate predictive performance alongside computational efficiency, including parameter counts and inference latency. Our findings reveal that well-optimized classical architectures, such as Siamese U-Nets, frequently outperform more complex contemporary models when computational efficiency is factored in, and that pre-training consistently provides a significant performance boost with no additional inference cost. To ensure complete transparency and reproducibility, all experimental resources, including standardized data splits, training scripts, training logs, and model checkpoints are publicly available and adhere to FAIR principles (Findable, Accessible, Interoperable, and Reusable).
Chinese Translation
地球观测(EO)中的变化检测对于监测地表变迁至关重要,但该领域近期的研究受到不一致的评估协议和对预测准确性的狭隘关注的限制,而忽视了计算效率。为了解决这一问题,我们提出了一个标准化的开源基准,用于评估最先进(SOTA)的深度学习方法在地球观测变化检测中的表现。我们对十种具有代表性的模型架构进行了全面分析,这些架构涵盖了从卷积网络(CNNs)到视觉变换器(ViTs)的十个异构变化检测数据集。我们使用相同的实验协议对这些模型进行了严格评估,比较了从头训练的模型与利用预训练权重的模型。此外,我们还评估了预测性能和计算效率,包括参数数量和推理延迟。我们的研究结果表明,当考虑计算效率时,经过良好优化的经典架构(如Siamese U-Nets)往往优于更复杂的现代模型,并且预训练始终能显著提升性能而不增加额外的推理成本。为了确保完全的透明性和可重复性,所有实验资源,包括标准化的数据划分、训练脚本、训练日志和模型检查点,均已公开,并遵循FAIR原则(可查找性、可访问性、可互操作性和可重用性)。
cs.CV / 52 / 2608.28248
Synth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images
Synth-JDoc:为具有多样布局和嵌入图像的OCR合成日文文档图像数据集
Abstract
The ability of Large Vision Language Models (LVLMs) to read text within document images is crucial, as it enables various applications such as Document Visual Question Answering. To enhance the text-reading capabilities of LVLMs, high-quality OCR datasets are essential. This need is particularly critical for Japanese documents, which often feature vertically written text alongside horizontally written text. Current LVLMs demonstrate considerably lower performance on vertically written Japanese text than on horizontally written text, necessitating specialized OCR datasets to bridge this gap. However, manually constructing OCR datasets is expensive and difficult to scale. Alternatively, constructing datasets by extracting text from existing document images using OCR models introduces challenges, such as text recognition errors and the prerequisite of sourcing document images. To address these issues, we construct an OCR dataset by synthesizing document images directly from text. Leveraging HTML and CSS, we generate multi-column documents that incorporate both vertical and horizontal writing styles. Furthermore, to ensure the visual realism of the documents, we embed images generated by text-to-image models within the layout. Additionally, to foster model robustness, we apply noise and degradation filters to the synthesized document images. In our experiments, we compared the performance of models fine-tuned on our synthetic dataset against baselines fine-tuned on synthetic datasets from prior work and those generated by a high-performance text-to-image model. Evaluation results demonstrate that our synthetic dataset is the most effective approach for improving LVLM performance on reading vertically written Japanese text. Our dataset and code are publicly available (https://github.com/llm-jp/synth-jdoc).
Chinese Translation
大型视觉语言模型(LVLMs)在文档图像中读取文本的能力至关重要,因为这使得文档视觉问答等多种应用成为可能。为了增强LVLMs的文本读取能力,高质量的OCR数据集是必不可少的。这一需求对于日文文档尤为重要,因为日文文档通常同时包含竖排和横排文本。目前,LVLMs在竖排日文文本上的表现明显低于横排文本,因此需要专门的OCR数据集来弥补这一差距。然而,手动构建OCR数据集既昂贵又难以扩展。另一方面,通过使用OCR模型从现有文档图像中提取文本来构建数据集也面临挑战,例如文本识别错误和获取文档图像的前提条件。为了解决这些问题,我们通过直接从文本合成文档图像来构建OCR数据集。利用HTML和CSS,我们生成包含竖排和横排书写风格的多栏文档。此外,为了确保文档的视觉真实感,我们在布局中嵌入了由文本到图像模型生成的图像。此外,为了增强模型的鲁棒性,我们对合成的文档图像应用了噪声和降解滤镜。在我们的实验中,我们比较了在我们的合成数据集上微调的模型与在先前工作中的合成数据集和由高性能文本到图像模型生成的数据集上微调的基线模型的性能。评估结果表明,我们的合成数据集是提高LVLM在读取竖排日文文本时性能的最有效方法。我们的数据集和代码已公开发布(https://github.com/llm-jp/synth-jdoc)。
cs.CV / 53 / 2608.28272
Non-Uniform Quantisation for 3DGS Compression
用于3DGS压缩的非均匀量化
Abstract
3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis, yet its high bitrate requirements pose significant challenges for storage and transmission. To enable practical applications and ensure interoperability within the 3DGS ecosystem, standardised compression formats are essential. In this paper, we propose a novel non-uniform quantisation scheme specifically tailored for 3DGS models. Our approach adapts to the underlying data distribution by applying importance-weighted quantisation and eliminating post-voxelisation redundancy through importance weighted merging. Extensive evaluations on benchmark datasets demonstrate that our method achieves state-of-the-art compression performance. Furthermore, the proposed scheme is compatible with any point-cloud-based representation and is intended as a formal contribution to the upcoming MPEG 3DGS compression standardisation activities.
Chinese Translation
3D高斯点云(3D Gaussian Splatting, 3DGS)作为一种强大的新视图合成技术,尽管其高比特率需求对存储和传输带来了显著挑战。为了实现实际应用并确保3DGS生态系统内的互操作性,标准化的压缩格式至关重要。本文提出了一种专门针对3DGS模型的新型非均匀量化方案。我们的方法通过应用重要性加权量化来适应底层数据分布,并通过重要性加权合并消除后体素化冗余。在基准数据集上的广泛评估表明,我们的方法达到了最先进的压缩性能。此外,所提出的方案与任何基于点云的表示形式兼容,旨在为即将到来的MPEG 3DGS压缩标准化活动做出正式贡献。
cs.CV / 54 / 2608.28288
GeoFF3D: Coordinate-Anchored Feed-Forward Reconstruction for Large-Scale UAV Mapping
GeoFF3D:用于大规模无人机测绘的坐标锚定前馈重建
Abstract
Existing feed-forward 3D reconstruction methods typically process a bounded number of images and recover cameras and geometry in local or internally normalized frames. Extending them to large-scale UAV mapping requires scalable multi-chunk processing and reliable aggregation, while full Sim(3) alignment can become unstable for near collinear trajectories. We present GeoFF3D, which combines a coordinate-anchored model with a spatial large-scale reconstruction framework (SLRF). The model uses georeferenced camera translations and optional geometric priors to predict camera poses and dense point maps directly in a gravity-aligned Z-up metric frame. SLRF partitions images into spatially overlapping chunks, propagates shared-view priors, and aggregates local reconstructions hierarchically, while remaining applicable to different bounded-view models. Across nine aerial mapping blocks, GeoFF3D achieves the best average reconstruction quality, improving F@5 from 0.829 for Pi3X + SLRF to 0.877. On long UAVScenes sequences, it reaches 0.848, compared with 0.687 for Pi3X + SLRF and 0.451 for the strongest evaluated SLAM/streaming baseline. GeoFF3D reconstructs 2,000 images in approximately five minutes, demonstrating scalable and robust large-scale UAV reconstruction.The code is available at https://github.com/yanxian-ll/GeoFF3D.
Chinese Translation
现有的前馈三维重建方法通常处理有限数量的图像,并在局部或内部归一化框架中恢复相机和几何结构。将其扩展到大规模无人机测绘需要可扩展的多块处理和可靠的聚合,而完全的 Sim(3) 对齐在近共线轨迹下可能变得不稳定。我们提出了 GeoFF3D,它结合了坐标锚定模型和空间大规模重建框架(SLRF)。该模型使用地理参考的相机平移和可选的几何先验,直接在重力对齐的 Z-up 公制框架中预测相机姿态和密集点云图。SLRF 将图像划分为空间重叠的块,传播共享视图先验,并以分层方式聚合局部重建,同时仍适用于不同的有限视图模型。在九个航空测绘区块中,GeoFF3D 实现了最佳的平均重建质量,将 Pi3X + SLRF 的 F@5 从 0.829 提高到 0.877。在长时间的 UAVScenes 序列中,其达到 0.848,而 Pi3X + SLRF 为 0.687,评估的最强 SLAM/流媒体基线为 0.451。GeoFF3D 在大约五分钟内重建 2000 张图像,展示了可扩展和稳健的大规模无人机重建。代码可在 https://github.com/yanxian-ll/GeoFF3D 获取。
cs.CV / 55 / 2608.28302
FUSED: Forensic-Semantic Mixture-of-Experts for AI Inpainting Detection and Localization
FUSED:用于AI修复检测和定位的法医语义专家混合模型
Abstract
Diffusion-based inpainting models modify only a localized part of an image, while many AI-image detectors rely on global artifacts and do not localize. These artifacts vary across generators, limiting detector transfer under distribution shifts. Recent work shows that restoring the authentic pixels outside the inpainted region removes these cues and can degrade pretrained detectors. To address this, we present FUSED, a unified framework for the joint detection and localization of AI-generated inpainting. FUSED combines low-level forensic cues with high-level semantic features using a sparsely-gated Mixture-of-Experts architecture, enabling the model to adaptively prioritize the most relevant signal for each token. For each input, FUSED predicts both an image-level manipulation score and a pixel-level mask of the inpainted area. On the OpenSDID cross-generator benchmark, FUSED achieves the best average detection and localization, with the largest gains on unseen generators. The same model transfers directly to the held-out AutoSplice and CocoGlide benchmarks, more than doubling localization performance. Evaluating each held-out benchmark with and without the global generator artifact further shows that all evaluated methods, ours included, partly read the artifact as evidence of manipulation, and FUSED remains the strongest under both conditions. Code and pretrained models are available at https://github.com/AntonNuzhdin/FUSED.
Chinese Translation
基于扩散的修复模型仅修改图像的局部区域,而许多AI图像检测器依赖于全局伪影且不进行定位。这些伪影在不同生成器之间有所不同,限制了检测器在分布变化下的迁移。近期研究表明,恢复修复区域外的真实像素会去除这些线索,并可能降低预训练检测器的性能。为了解决这个问题,我们提出了FUSED,一个用于AI生成修复的联合检测和定位的统一框架。FUSED结合了低级法医线索和高级语义特征,采用稀疏门控的专家混合架构,使模型能够自适应地优先考虑每个令牌的最相关信号。对于每个输入,FUSED同时预测图像级别的操控分数和修复区域的像素级掩码。在OpenSDID跨生成器基准测试中,FUSED实现了最佳的平均检测和定位,尤其在未见过的生成器上取得了最大的提升。同一模型直接迁移到保留的AutoSplice和CocoGlide基准测试中,定位性能提升超过两倍。对每个保留基准在有无全局生成器伪影的情况下进行评估进一步表明,所有评估方法,包括我们的方法,部分将伪影视为操控的证据,而FUSED在这两种情况下仍然表现最强。代码和预训练模型可在https://github.com/AntonNuzhdin/FUSED获取。
cs.CV / 56 / 2608.28312
AIM: Anchor Identity Features, Then Match for Multimodal Large Language Model Unlearning
AIM:锚定身份特征,然后进行多模态大语言模型的遗忘匹配
Abstract
Multimodal large language models (MLLMs) can memorize identity-specific facts about people in their fine-tuning data, creating privacy risks when a person requests deletion. Existing MLLM unlearning methods often assume access to retain images or ground-truth answers during deletion, which is unrealistic in many practical scenarios. We study identity unlearning when retain images are unavailable at deletion time. Our analysis shows that identity and visual-perception questions occupy distinct regions in fine-tuned hidden states and are organized differently: identity questions cluster by person, whereas perception questions cluster by question type. This suggests that identity knowledge can be suppressed without erasing general visual perception. Building on this observation, we propose AIM, a two-stage method that anchors an identity-forgetting target with a universal visual prompt and then matches the vision encoder to that target under a Fisher-based constraint. Extensive experiments show that AIM achieves competitive identity forgetting while preserving non-deleted identities, prior knowledge, and visual perception on the same images.
Chinese Translation
多模态大语言模型(MLLMs)能够记忆关于个人的身份特定事实,这些事实存在于其微调数据中,当个人请求删除时,会产生隐私风险。现有的MLLM遗忘方法通常假设在删除过程中可以访问保留的图像或真实答案,这在许多实际场景中并不现实。我们研究在删除时无法获得保留图像的身份遗忘问题。我们的分析表明,身份问题和视觉感知问题在微调的隐藏状态中占据不同的区域,并且组织方式不同:身份问题按个人聚类,而感知问题则按问题类型聚类。这表明可以在不抹去一般视觉感知的情况下抑制身份知识。基于这一观察,我们提出了AIM,一种两阶段的方法,首先使用通用视觉提示锚定身份遗忘目标,然后在基于Fisher的约束下将视觉编码器与该目标匹配。大量实验表明,AIM在保留未删除身份、先前知识和相同图像上的视觉感知的同时,实现了竞争性的身份遗忘。
cs.CV / 57 / 2608.28316
Conditional Visual Evidence Utility: State-Dependent Rank Reversals in Frozen Vision-Language Encoders
条件视觉证据效用:冻结视觉-语言编码器中的状态依赖排名逆转
Abstract
Static importance scores compress visual evidence into a single ranking, but the value of remaining evidence can change after one cue has been observed. We study this possibility in controlled compositional visual search, where color, shape, and texture evidence can be independently exposed and their conditional marginal utility measured across acquisition states. In a held-out confirmation on 800 scenes, frozen OpenCLIP and SigLIP exhibit robust state-dependent rank reversals that concentrate in candidate-overlap regimes designed to induce ordering changes. The structure persists across two evidence-accumulation constructions and ten equivalent query wordings, but disappears under query-scene derangement. We also ask whether these reversals matter for decisions. In a post-confirmation exploratory matched-first-action analysis, reranking only after the first acquisition yields positive step-2 utility when decisions are selected under one evidence mode, wording, or backbone and evaluated under another. Together, these results show that evidence importance is state-dependent in this controlled setup and that updating an evidence ordering can retain decision-relevant value across evaluator changes. They motivate evaluating vision-language evidence use conditionally rather than through a single static ranking, while providing a measurable target for future adaptive evidence-selection methods.
Chinese Translation
静态重要性评分将视觉证据压缩为单一排名,但在观察到一个线索后,剩余证据的价值可能会发生变化。我们在受控的组合视觉搜索中研究了这种可能性,其中颜色、形状和纹理证据可以独立暴露,并且可以在获取状态下测量其条件边际效用。在对800个场景的保留确认中,冻结的OpenCLIP和SigLIP表现出强烈的状态依赖排名逆转,这些逆转集中在旨在引发排序变化的候选重叠区域。该结构在两种证据积累构造和十种等效查询措辞中持续存在,但在查询-场景混乱下消失。我们还探讨了这些逆转对决策是否重要。在后确认的探索性匹配首个行动分析中,仅在第一次获取后重新排名,当决策在一种证据模式、措辞或基础模型下选择并在另一种下评估时,产生了积极的第二步效用。综合来看,这些结果表明,在这一受控设置中,证据的重要性是状态依赖的,并且更新证据排序可以在评估者变化中保留与决策相关的价值。这些结果激励我们以条件方式评估视觉-语言证据的使用,而不是通过单一静态排名,同时为未来自适应证据选择方法提供了可测量的目标。
cs.CV / 58 / 2608.28339
Abstract4D: A Large-Scale Dataset and Framework for Understanding the Visual Language of Abstract Art
Abstract4D:理解抽象艺术视觉语言的大规模数据集和框架
Abstract
Artificial intelligence can classify artistic styles and synthesize images, but it still lacks a model of the visual language that gives art meaning. Abstract painting minimizes object semantics and foregrounds structural cues, making it an ideal testbed for computational perception. We introduce \textbf{Abstract4D}, the largest dataset of abstract paintings to date: more than 120,000 images paired with rich metadata and multi-dimensional prompts that capture each work's perceptual attributes---\textit{form, color, texture, and composition}. Annotations are produced by a hybrid human--VLM pipeline for quality and consistency. Using Abstract4D, we (i) analyze the semantic structure of abstract art through large-scale embedding visualization, uncovering how perceptual relationships organize artistic meaning, and (ii) establish benchmark tasks for classification, cross-modal retrieval, and text-to-image generation to evaluate how AI models perceive and reproduce abstract visual language. Together, these analyses demonstrate how Abstract4D enables both exploration and quantitative assessment of AI's ability to represent and interpret abstract art.
Chinese Translation
人工智能可以对艺术风格进行分类并合成图像,但仍缺乏赋予艺术意义的视觉语言模型。抽象绘画最小化了物体语义,突出了结构线索,使其成为计算感知的理想测试平台。我们介绍了 extbf{Abstract4D},迄今为止最大的抽象绘画数据集:包含超过120,000幅图像,配有丰富的元数据和多维提示,捕捉每件作品的感知属性—— extit{形状、颜色、纹理和构图}。注释通过混合人类与视觉语言模型(VLM)管道生成,以确保质量和一致性。利用Abstract4D,我们(i)通过大规模嵌入可视化分析抽象艺术的语义结构,揭示感知关系如何组织艺术意义,以及(ii)建立分类、跨模态检索和文本到图像生成的基准任务,以评估人工智能模型如何感知和再现抽象视觉语言。这些分析共同展示了Abstract4D如何促进对人工智能在表现和解读抽象艺术能力的探索和定量评估。
cs.CV / 59 / 2608.28341
Cross-Spectral Dense Correspondence for Multimodal Spectral Medical Imaging
多模态光谱医学成像的跨光谱密集对应
Abstract
Precise dense correspondence is a fundamental prerequisite for multimodal spectral imaging systems that fuse disparate wavelength ranges for subsequent analysis in medical and scientific imaging. Corresponding image points are often observed with non-overlapping spectral sensitivities, leading to wavelength-dependent contrast changes, intensity inversions, and appearance shifts for which dense ground truth is difficult to obtain and conventional RGB-based training data provides only limited supervision. We address this data gap by introducing a sensor-agnostic cross-spectral modulation protocol on established correspondence benchmarks with intensity input projection, and by proposing a synthetic cross-spectral correspondence benchmark simulating physically plausible radiometric differences. Evaluation on several modern dense correspondence backbones trained with our unified cross-spectral protocol showed substantial improvements under severe spectral mismatch while maintaining performance on standard RGB benchmarks. Ablation experiments show that view-dependent channel selection and nonlinear radiometric transformations provide complementary robustness, indicating that the primary limitation of existing models is not their structural matching capacity but the mismatch between training distribution and spectral characteristics of the target image pair. Qualitative evaluations on heterogeneous medical spectral acquisition systems demonstrate the practical relevance of the proposed training data augmentation protocol as an enabler for spatially coherent spectral fusion in HSI workflows.
Chinese Translation
精确的密集对应是多模态光谱成像系统的基本前提,该系统融合不同波长范围以便于后续的医学和科学成像分析。对应的图像点通常在光谱灵敏度上存在不重叠的情况,导致波长依赖的对比度变化、强度反转和外观变化,这使得获取密集的真实数据变得困难,而传统的基于RGB的训练数据仅提供有限的监督。我们通过在已建立的对应基准上引入一种传感器无关的跨光谱调制协议,结合强度输入投影,来解决这一数据缺口,并提出了一种合成的跨光谱对应基准,模拟物理上合理的辐射差异。在使用我们统一的跨光谱协议训练的几种现代密集对应骨干网络上的评估显示,在严重的光谱不匹配情况下取得了显著改善,同时在标准RGB基准上保持了性能。消融实验表明,视角依赖的通道选择和非线性辐射变换提供了互补的鲁棒性,表明现有模型的主要限制并非其结构匹配能力,而是训练分布与目标图像对的光谱特征之间的不匹配。对异构医学光谱获取系统的定性评估展示了所提训练数据增强协议的实际相关性,作为在高光谱成像工作流程中实现空间一致光谱融合的推动者。
cs.CV / 60 / 2608.28343
Denoising-Aware Temporal Point Cloud Completion for 3D Crop Architecture Recovery and Phenotypic Trait Extraction
考虑去噪的时序点云补全用于3D作物结构恢复和表型特征提取
Abstract
High-throughput phenotyping depends on accurate 3D reconstruction of plants across growth stages, yet the development and evaluation of temporal completion methods are limited by the lack of datasets with complete geometric ground truth. To address this challenge, we introduce SynthCrop4D, a procedurally generated synthetic dataset of temporally evolving plant point clouds that provides controllable noise, occlusion, and complete plant geometry for benchmarking reconstruction methods. Using this dataset, we evaluate a two-stage pipeline that combines spatial denoising and temporal point cloud completion. First, a denoising module removes structural artifacts from raw laser-scanned point clouds. The resulting data are then processed by an Adaptive Temporal PoinTr model that reconstructs the current growth stage (t) using information from the previous stage (t-1), enabling recovery of regions missing due to self-occlusion. We evaluate the proposed framework on both SynthCrop4D and the real-world Pheno4D dataset (tomato and maize) under settings with and without denoising. Results show that denoising substantially improves reconstruction quality, with the best configuration achieving a Chamfer Distance of 0.0061 on SynthCrop4D (Temporal PoinTr + Mamba-DG) and an F-Score of 0.2080 on Pheno4D (Vanilla PoinTr + Mamba-DG). We further demonstrate the use of completed point clouds for phenotypic trait extraction, including plant height, canopy width, and convex hull volume, obtaining hull-volume MAEs of 0.021 on synthetic data and 0.343 on real data. Together, SynthCrop4D and the proposed pipeline provide a benchmark and methodology for temporal plant reconstruction and high-throughput crop phenotyping.
Chinese Translation
高通量表型分析依赖于植物在生长阶段的准确3D重建,然而,时序补全方法的发展和评估受到缺乏完整几何基准数据集的限制。为了解决这一挑战,我们引入了SynthCrop4D,这是一个程序生成的合成数据集,包含时序演变的植物点云,提供可控的噪声、遮挡和完整的植物几何结构,以便对重建方法进行基准测试。利用该数据集,我们评估了一个结合空间去噪和时序点云补全的两阶段管道。首先,去噪模块从原始激光扫描点云中去除结构伪影。然后,得到的数据通过自适应时序PoinTr模型进行处理,该模型利用前一阶段(t-1)的信息重建当前生长阶段(t),从而恢复由于自遮挡而缺失的区域。我们在SynthCrop4D和真实世界的Pheno4D数据集(番茄和玉米)上评估了所提出的框架,在有去噪和无去噪的设置下进行比较。结果表明,去噪显著提高了重建质量,最佳配置在SynthCrop4D上实现了0.0061的Chamfer距离(Temporal PoinTr + Mamba-DG),在Pheno4D上获得了0.2080的F-Score(Vanilla PoinTr + Mamba-DG)。我们进一步展示了补全点云在表型特征提取中的应用,包括植物高度、冠幅宽度和凸包体积,合成数据的凸包体积平均绝对误差为0.021,真实数据为0.343。综上所述,SynthCrop4D和所提出的管道为时序植物重建和高通量作物表型分析提供了基准和方法论。
cs.CV / 61 / 2608.28371
Real-Time Musculoskeletal Surrogates for Pediatric Cerebral Palsy: a Credibility Pilot
儿童脑瘫的实时肌肉骨骼替代模型:可信度初步研究
Abstract
Real-time musculoskeletal (MSK) surrogates could support personalized rehabilitation for children with cerebral palsy (CP), but their credibility depends on subject-wise evaluation, low inference latency, and calibrated uncertainty. We develop a subject-conditioned causal neural surrogate using OpenSim-derived static parameters, temporal joint kinematics, true muscle capacities, and training-only perturbations. On a real pediatric CP gait dataset comprising nine children, we use leave-one-subject-out validation on six development subjects and evaluate a frozen configuration once on three locked test subjects. The surrogate accurately reproduces musculotendon lengths (R-square = 0.92 in development validation and approximately 0.95 on locked subjects; nRMSE < 8%) while requiring only sub-millisecond to few-millisecond neural inference, well below a 100 ms interactive-rehabilitation target. In contrast, direct muscle-force estimation remains unstable at this small, heterogeneous scale: pooled metrics can overstate within-subject, per-muscle accuracy. A Monte Carlo credibility pilot further shows that propagating only +/-5% anthropometry and muscle-capacity variation produces severely overconfident nominal 90% intervals (approximately 4% force coverage and below 1% MT-length coverage). These results establish a leakage-free evaluation and credibility framework for pediatric MSK surrogates, while identifying force modeling and epistemic uncertainty as the central next challenges for clinically credible digital twins.
Chinese Translation
实时肌肉骨骼(MSK)替代模型可以支持儿童脑瘫(CP)的个性化康复,但其可信度依赖于个体评估、低推理延迟和校准的不确定性。我们开发了一种基于个体条件的因果神经替代模型,使用了OpenSim派生的静态参数、时间关节运动学、真实肌肉能力和仅在训练中引入的扰动。在一个包含九名儿童的真实儿童脑瘫步态数据集中,我们对六名发展对象进行了留一法验证,并在三名锁定测试对象上评估了冻结配置。该替代模型准确再现了肌腱长度(开发验证中的R平方为0.92,在锁定对象上的R平方约为0.95;nRMSE < 8%),同时仅需亚毫秒到几毫秒的神经推理,远低于100毫秒的互动康复目标。相比之下,直接的肌肉力量估计在这个小而异质的规模上仍然不稳定:汇总指标可能会夸大个体内每个肌肉的准确性。一项蒙特卡洛可信度初步研究进一步表明,仅传播±5%的人体测量和肌肉能力变化会导致严重过度自信的名义90%区间(约4%的力量覆盖率和低于1%的MT长度覆盖率)。这些结果建立了一个无泄漏的评估和可信度框架,用于儿童肌肉骨骼替代模型,同时识别力量建模和认知不确定性作为临床可信数字双胞胎的核心下一挑战。
cs.CV / 62 / 2608.28383
Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs
语义头部专业化引导混合视觉变换器注意力用于多模态大语言模型
Abstract
Hybrid attention dominates frontier LLMs, yet Vision Transformers (ViTs) in multimodal LLMs lack a satisfactory hybrid design, with no consensus on why certain attention patterns work better. To fill this gap, we study ViT attention heads and find they differentiate into object- and background-specialist roles, a pattern most pronounced under full attention; we call this Semantic Head Specialization (SHS). We propose SHS-Index to quantify this specialization, show that it distinguishes full-attention from chunk-window ViTs, and find that it strongly tracks downstream benchmark performance. We then identify three structural factors that shape SHS---window interaction, token serialization, and local softmax allocation---and use them as design principles for hybrid attention. Guided by these factors, we design Ariadne Attention, a hybrid that matches full attention on 22 image and video tasks at 6.5x less attention compute. Our findings establish head specialization as a measurable property for diagnosing and designing principled hybrid ViT attention at the multimodal-LLM scale.
Chinese Translation
混合注意力在前沿的大语言模型中占主导地位,但多模态大语言模型中的视觉变换器(ViTs)缺乏令人满意的混合设计,且对于某些注意力模式为何表现更好的原因尚无共识。为填补这一空白,我们研究了ViT的注意力头,发现它们分化为物体和背景专家角色,这一模式在完全注意力下最为明显;我们称之为语义头部专业化(Semantic Head Specialization, SHS)。我们提出SHS-Index来量化这种专业化,显示它能够区分完全注意力与块窗口ViTs,并发现它与下游基准性能有很强的相关性。随后,我们识别出三种结构因素影响SHS——窗口交互、令牌序列化和局部softmax分配——并将其作为混合注意力的设计原则。在这些因素的指导下,我们设计了阿里阿德涅注意力(Ariadne Attention),这一混合模型在22个图像和视频任务中以6.5倍更少的注意力计算量匹配完全注意力。我们的发现确立了头部专业化作为一种可测量属性,用于诊断和设计符合原则的混合ViT注意力,适用于多模态大语言模型的规模。
cs.CV / 63 / 2608.28386
GraspHOI: Full-Body 3D Human-Object Reconstruction with Finger-Level Grasps from a Single In-the-Wild Image
GraspHOI:从单张自然图像中重建全身3D人机交互的手指级抓取
Abstract
Existing monocular full-body 3D human-object interaction (HOI) methods do not combine explicit finger-level grasp optimization with category-agnostic object reconstruction. Despite plausible body-object configurations, their fingers may float from or penetrate objects instead of forming a grasp. We present GraspHOI, the first framework that reconstructs a full-body 3D HOI from a single image while explicitly optimizing finger articulation against the reconstructed object. GraspHOI recovers object geometry directly, without predefined meshes or a fixed category vocabulary. It reconstructs the body, hands, and object separately, aligning them in metric camera space via depth-based registration and image-space alignment. Occlusion-aware palmar correspondences seat the object against the grasping hand, and contact-aware optimization refines arm and finger articulation to form surface contact without excessive penetration. Across four benchmarks and six baselines, GraspHOI improves relative human-object placement, hand accuracy, and contact plausibility. Full pipeline code will be released.
Chinese Translation
现有的单目全身3D人机交互(HOI)方法未能将显式的手指级抓取优化与类别无关的物体重建相结合。尽管它们的身体-物体配置看似合理,但手指可能会漂浮在物体之外或穿透物体,而不是形成有效的抓取。我们提出了GraspHOI,这是第一个从单张图像重建全身3D HOI的框架,同时显式优化手指的关节运动以适应重建的物体。GraspHOI直接恢复物体几何形状,无需预定义的网格或固定的类别词汇。它分别重建身体、手和物体,通过基于深度的配准和图像空间对齐将它们对齐到度量相机空间。考虑遮挡的掌部对应关系将物体固定在抓取手上,而考虑接触的优化则细化手臂和手指的关节运动,以形成表面接触而不发生过度穿透。在四个基准测试和六个基线方法中,GraspHOI改善了相对的人机位置、手部精度和接触合理性。完整的管道代码将会发布。
cs.CV / 64 / 2608.28404
How Far Can 5,500 Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion Models
5500小时驾驶能带你多远?视频扩散模型的尺度法则分析
Abstract
Video generation for autonomous driving cannot follow the web-scale route: driving data is expensive to collect, bound by privacy requirements, and cannot be scraped at will, so models must make the most of a fixed corpus. We present a systematic scaling-law study of video diffusion models trained from scratch on driving data: a family of models from 1M to 9B parameters, trained at different exposures on up to 5,500 hours of driving. Validation loss follows consistent power laws in both model size and training exposure, answering the questions that shape a training budget: whether compute is better spent on longer training or on a larger model, and whether more data is needed. Loss improves much faster with training exposure than with model size, making longer training the most effective way to improve a fixed model under limited compute. However, larger models continue to achieve lower asymptotic loss, so compute-optimal scaling still favors increasing model size when sufficient compute and data are available. Guided by these laws, we train a 9B-parameter model, to our knowledge the largest video diffusion model trained from scratch on driving data: it sets a new open-source state of the art for driving video generation, as measured on nuScenes. Our code and pretrained models are available at https://github.com/valeoai/VATIX. NATIX is separately releasing the underlying driving data in stages.
Chinese Translation
自主驾驶的视频生成无法遵循网络规模的路线:驾驶数据的收集成本高昂,受到隐私要求的限制,且无法随意抓取,因此模型必须充分利用固定的语料库。我们提出了一项系统的尺度法则研究,针对从头开始在驾驶数据上训练的视频扩散模型:一系列从1M到9B参数的模型,在多达5500小时的驾驶数据上以不同的曝光量进行训练。验证损失在模型规模和训练曝光上遵循一致的幂律,回答了影响训练预算的问题:计算资源是更适合用于更长的训练还是更大的模型,以及是否需要更多的数据。与模型规模相比,训练曝光带来的损失改善速度更快,使得在有限计算资源下,延长训练成为提升固定模型性能的最有效方式。然而,更大的模型仍然能够实现更低的渐近损失,因此在计算资源和数据充足的情况下,计算最优的扩展仍然倾向于增加模型规模。在这些法则的指导下,我们训练了一个9B参数的模型,至今为止这是在驾驶数据上从头开始训练的最大视频扩散模型:它在nuScenes上设定了驾驶视频生成的新开源状态的最佳表现。我们的代码和预训练模型可在https://github.com/valeoai/VATIX获取。NATIX将分阶段发布基础驾驶数据。
cs.CV / 65 / 2608.28406
Post-Training VLMs for Video Mistake Detection
视频错误检测的后训练视觉语言模型
Abstract
Human mistakes are inevitable when following instructions, yet they can lead to severe consequences. As such, there has been an increased interest in developing methods for detecting mistakes in videos, with current methods mostly focusing on closed-set protocols. While successful in controlled settings, the closed-set assumption limits their wider applicability, as any changes to the task require collecting new data and re-training models. Instead, we argue that mistake detection methods should learn the general concept of a mistake, rather than overfitting to step-specific details. To reflect this, we introduce the Mistake Detection Video Question Answering (MD-VQA) protocol and accompanying benchmark. MD-VQA tests whether methods can discern if a step was executed correctly with respect to its description, for both seen and unseen actions. To address this important challenge, we propose the first video-language-model post-training technique for mistake detection. Our method uses a tailored reward function to encourage the model to identify discrepancies between an instruction and the corresponding video. Extensive evaluations demonstrate that this approach outperforms zero-shot, supervised fine-tuning, and post-training baselines. Notably, our method generalizes especially well to unseen procedures, for instance, with an improvement of up to 11.6% over the best-performing baseline on EP-VQA, paving the way toward general mistake detection. We release our code and benchmark at https://github.com/FedeSpu/mstk.
Chinese Translation
在遵循指令时,人类错误是不可避免的,但这些错误可能导致严重后果。因此,开发检测视频中错误的方法的兴趣日益增加,目前的方法主要集中在封闭集协议上。尽管在受控环境中取得了成功,但封闭集假设限制了它们的更广泛适用性,因为任何任务的变化都需要收集新数据并重新训练模型。相反,我们认为错误检测方法应该学习错误的一般概念,而不是过度拟合于特定步骤的细节。为此,我们引入了错误检测视频问答(Mistake Detection Video Question Answering, MD-VQA)协议及其相应的基准。MD-VQA 测试方法是否能够判断某个步骤是否按照其描述正确执行,适用于已见和未见的动作。为了解决这一重要挑战,我们提出了首个用于错误检测的视频语言模型后训练技术。我们的方法使用定制的奖励函数,鼓励模型识别指令与相应视频之间的差异。广泛的评估表明,该方法在零样本、监督微调和后训练基准上表现优越。值得注意的是,我们的方法在未见程序上特别具有良好的泛化能力,例如,在 EP-VQA 上相较于表现最佳的基准提高了多达 11.6%,为通用错误检测铺平了道路。我们在 https://github.com/FedeSpu/mstk 发布了我们的代码和基准。
cs.CV / 66 / 2608.28429
Lossy Event Compression: From Event Stream Distortion to Task Performance
有损事件压缩:从事件流失真到任务性能
Abstract
Event cameras generate asynchronous, sparse data streams with microsecond temporal resolution, but in moderate-to-high motion scenes they can produce as many as hundreds of millions of events per second, creating significant bandwidth and storage challenges. Lossy compression is therefore essential for practical deployment, yet existing event stream distortion metrics fail to reliably predict compression-induced degradation at the task level, forcing codec optimization to rely on expensive task-specific evaluations. To address this gap, this paper introduces two fundamentally different event compression pipelines: i) an aggregation-based pipeline that converts the event stream into polarity-based histogram frames for compression with the conventional image codec JPEG 2000, and ii) a frame-free point cloud-based pipeline that codes events natively as 3D points using the octree-based codec G-PCC. Both pipelines are then assessed within a unified task-driven evaluation framework that relates event stream distortion to downstream application performance across four representative tasks: i) video reconstruction, ii) object detection, iii) optical flow estimation, and a delay-sensitive task iv) asynchronous feature tracking under a reference-relative protocol. Building on this framework, five classification-based distortion metrics are applied to event compression for the first time, to the best of the authors' knowledge, and benchmarked against existing event stream metrics. Experimental results demonstrate that the proposed metrics reliably predict compression-induced task degradation across different coding frameworks. This demonstrates that event stream distortion assessment can be an efficient alternative to repeated task-specific evaluation, providing direct guidance for the development and optimization of future event data coding solutions.
Chinese Translation
事件相机生成具有微秒时间分辨率的异步稀疏数据流,但在中到高运动场景中,它们每秒可能产生多达数亿个事件,这给带宽和存储带来了重大挑战。因此,有损压缩对于实际部署至关重要,但现有的事件流失真度量无法可靠地预测压缩引起的任务级性能下降,迫使编解码器优化依赖于昂贵的任务特定评估。为了解决这一问题,本文提出了两种根本不同的事件压缩管道:i) 一种基于聚合的管道,将事件流转换为基于极性的直方图帧,以便使用传统图像编解码器JPEG 2000进行压缩;ii) 一种无帧的点云基础管道,使用基于八叉树的编解码器G-PCC将事件原生编码为3D点。然后,在一个统一的任务驱动评估框架内评估这两种管道,该框架将事件流失真与四个代表性任务的下游应用性能相关联:i) 视频重建,ii) 目标检测,iii) 光流估计,以及一个延迟敏感任务iv) 在参考相对协议下的异步特征跟踪。基于该框架,作者首次将五种基于分类的失真度量应用于事件压缩,并与现有的事件流度量进行了基准测试。实验结果表明,所提出的度量能够可靠地预测不同编码框架下的压缩引起的任务性能下降。这表明事件流失真评估可以成为重复任务特定评估的有效替代方案,为未来事件数据编码解决方案的开发和优化提供直接指导。
cs.CV / 67 / 2608.28453
Prompt-Guided Interactive Segmentation of Interstitial Lung Disease in Thoracic CT
基于提示的胸部CT间质性肺病交互式分割
Abstract
Accurate segmentation of interstitial lung disease (ILD) patterns is essential for quantitative disease assessment and longitudinal monitoring. However, existing approaches remain limited by relying on dense annotations and producing static predictions that cannot be refined, motivating interactive approaches. While promptable models show promise in interactive segmentation, their adaptation to ILDs remains largely unexplored. To address this gap, we investigate prompt-guided foundation models for ILD refinement and present, to the best of our knowledge, the first adaptation of MedSAM2 for interactive 3D ILD segmentation on thoracic CT. We investigate three fine-tuning strategies and multiple clinically motivated prompts: bounding-boxes (BBox), point, lasso, and scribble. On a dataset spanning seven ILD patterns and healthy lung tissue, full model fine-tuning performed best, improving the average Dice score by 4.7 percentage points over MedSAM2.While BBox prompts achieve the strongest performance, non-native MedSAM2 interactions such as lasso and scribble prompts also prove effective. Finally, we present and evaluate a proof-of-concept end-to-end workflow in which MedSAM2 is initialized from an automatic segmentation prior and subsequently refined using radiologist prompts. Model weights and plug-ins made available at: https://github.com/AIHNlab/ILD-SemiSegTool.
Chinese Translation
间质性肺病(ILD)模式的准确分割对于定量疾病评估和纵向监测至关重要。然而,现有方法仍然受到依赖密集注释和产生无法细化的静态预测的限制,这促使了交互式方法的发展。尽管可提示模型在交互式分割中显示出潜力,但它们在ILD中的适应性仍然未得到充分探索。为了解决这一空白,我们研究了用于ILD细化的基于提示的基础模型,并首次将MedSAM2适配于胸部CT上的交互式3D ILD分割。我们探讨了三种微调策略和多种临床驱动的提示:边界框(BBox)、点、套索和涂鸦。在涵盖七种ILD模式和健康肺组织的数据集上,全面模型微调表现最佳,平均Dice分数比MedSAM2提高了4.7个百分点。虽然BBox提示实现了最强的性能,但非原生MedSAM2交互(如套索和涂鸦提示)也证明有效。最后,我们展示并评估了一个概念验证的端到端工作流程,其中MedSAM2从自动分割先验初始化,并随后使用放射科医生的提示进行细化。模型权重和插件可在以下链接获取:https://github.com/AIHNlab/ILD-SemiSegTool。
cs.CV / 68 / 2608.28455
ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT
ARC-CT:用于3D胸部CT的解剖导向对比视觉-语言学习
Abstract
Contrastive vision-language learning uses paired chest CT volumes and radiology reports to learn abnormality classifiers without manually annotated labels. However, two characteristics of chest CT challenge conventional global contrastive learning. First, many critical abnormalities are small or anatomically localized, and pooling an en- tire volume into a single embedding may dilute their visual evidence. Second, the standard contrastive objective treats every other scan in a batch as a negative. Because many chest CTs share abnormalities, this objective incorrectly pushes co-positive pairs apart. We propose Anatomy-Routed Contrastive Learning for 3D Chest CT (ARC-CT), a region-aware framework that addresses these limitations using only la- bels extracted from reports by an LLM, with no manual annotations or bounding boxes. ARC-CT combines three components: (1) an Anato- myQFormer localizing evidence via queries constrained by automatically generated organ masks; (2) a label-Jaccard soft InfoNCE objective in- tegrating the standard one-hot target with the label-set overlap of each pair, which reduces false-negative penalties between studies that share clinical findings; and (3) an organ-level alignment loss connecting mask- pooled visual features to organ-specific report text extracted offline with a large language model. ARC-CT achieves a 0.86 mask-free macro AUC across 18 abnormalities using a compact 3D ResNet-18 backbone. Over- all, ARC-CT outperforms both comparable efficient baselines and sev- eral larger transformer models. Our code and weights are available at https://github.com/arc-ct/arc-ct.
Chinese Translation
对比视觉-语言学习利用配对的胸部CT图像和放射学报告,在没有手动标注标签的情况下学习异常分类器。然而,胸部CT的两个特征对传统的全局对比学习提出了挑战。首先,许多关键异常是小型或解剖局部化的,将整个体积汇聚为单一嵌入可能会稀释其视觉证据。其次,标准的对比目标将批次中的每个其他扫描视为负样本。由于许多胸部CT共享异常,这一目标错误地将共阳性对推开。我们提出了解剖导向对比学习(Anatomy-Routed Contrastive Learning)用于3D胸部CT(ARC-CT),这是一个区域感知框架,利用从报告中提取的标签,完全不需要手动标注或边界框,解决了这些局限性。ARC-CT结合了三个组件:(1)通过受自动生成的器官掩膜约束的查询定位证据的AnatomyQFormer;(2)将标准的独热目标与每对的标签集重叠集成的标签-Jaccard软InfoNCE目标,减少了在共享临床发现的研究之间的假阴性惩罚;(3)将掩膜汇聚的视觉特征与通过大型语言模型离线提取的器官特定报告文本连接的器官级对齐损失。ARC-CT在18种异常情况下实现了0.86的无掩膜宏AUC,采用紧凑的3D ResNet-18骨干网络。总体而言,ARC-CT超越了可比的高效基线和多个更大规模的变换器模型。我们的代码和权重可在https://github.com/arc-ct/arc-ct获取。
cs.CV / 69 / 2608.28460
LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation
LayerRecall:一种状态条件记忆路由器,用于视频生成中的长时间一致性
Abstract
Autoregressive video diffusion enables scalable long-video generation by producing chunks from a bounded recent context. While recency-based caching preserves local continuity, it evicts historical cues needed when subjects, objects, scenes, or attributes reappear. Existing memory mechanisms expose models to nonlocal history, but access alone does not ensure effective use. Our analysis reveals that video DiT layers exhibit distinct preferences for current, recent, and distant context, suggesting that long-range memory requires deciding both what to retrieve and where to use it. We introduce LayerRecall, a current-conditioned, layer-selective memory router that retrieves relevant historical K/V states and injects them only into backbone-specific memory-sensitive layers while preserving local attention elsewhere. To reduce reliance on scarce high-quality long-horizon videos and explicit memory-allocation labels, we further propose Cross-Horizon Prediction Matching (CHPM), which uses a privileged long-context reference to supervise the bounded-memory router in prediction space. Across 100 multi-shot evaluation prompts, LayerRecall achieves the best overall results on MemoBench and MovieBench while matching its backbone on VBench-Long, demonstrating stronger long-range recovery without sacrificing local continuity. Qualitative analyses further reveal memory-guided self-correction, whereby initially mismatched local attributes return to their historical appearance without resetting ongoing motion or scene structure. Additional analyses show cross-backbone portability and negligible inference overhead.
Chinese Translation
自回归视频扩散通过从有限的最近上下文中生成片段,实现了可扩展的长视频生成。虽然基于最近性的缓存能够保持局部连续性,但它会驱逐在主题、物体、场景或属性重新出现时所需的历史线索。现有的记忆机制使模型能够接触非局部历史,但仅仅访问并不能确保有效利用。我们的分析揭示,视频 DiT 层对当前、最近和遥远上下文表现出明显的偏好,这表明长程记忆需要决定检索什么以及在何处使用。我们提出了 LayerRecall,这是一种基于当前状态的、层选择性的记忆路由器,它检索相关的历史 K/V 状态,并仅将其注入到特定于主干的记忆敏感层,同时在其他地方保持局部注意力。为了减少对稀缺的高质量长时间视频和显式记忆分配标签的依赖,我们进一步提出了跨时间预测匹配(Cross-Horizon Prediction Matching, CHPM),该方法利用特权的长上下文参考来监督预测空间中的有限记忆路由器。在 100 个多镜头评估提示中,LayerRecall 在 MemoBench 和 MovieBench 上取得了最佳整体结果,同时在 VBench-Long 上与其主干匹配,展示了更强的长程恢复能力而不牺牲局部连续性。定性分析进一步揭示了记忆引导的自我修正,即最初不匹配的局部属性在不重置正在进行的运动或场景结构的情况下回归到其历史外观。额外的分析显示了跨主干的可移植性和微不足道的推理开销。
cs.CV / 70 / 2608.28461
Anatomy-Aware Promptable Segmentation with Online Interactive Training for AUTOPET V
具有在线交互训练的解剖学感知可提示分割模型用于AUTOPET V
Abstract
We present an anatomy-aware, promptable model for whole-body lesion segmentation in FDG and PSMA PET/CT, developed for the AUTOPET V challenge. The proposed method is built as family of nnU-Net-based models and trained in two stages: i) a pre-training stage that produces a strong initial segmentation, and ii) an online interactive stage that learns to exploit scribble prompts, refining the prediction over successive interactions. Anatomical context is incorporated through organ supervision using a single shared head that predicts lesions and organs from the same features, which reduces false positives arising from physiological uptake. Also as the tracer (i.e., FDG/PSMA) is not provided at inference, we add a tracer classifier based on image processing and a random forest over coronal MIP features, routing each study to a combined FDG+PSMA model or to a PSMA-specific model. Across four-fold cross-validation the organ-supervised model achieves the best and most stable performance, the interactive stage improves the Dice score monotonically with each prompt, and PSMA-specific training yields the strongest tracer-wise results.
Chinese Translation
我们提出了一种解剖学感知的可提示模型,用于FDG和PSMA PET/CT中的全身病灶分割,该模型是为AUTOPET V挑战而开发的。所提出的方法基于nnU-Net模型系列,并分为两个阶段进行训练:i) 预训练阶段,产生强大的初始分割;ii) 在线交互阶段,学习利用涂鸦提示,通过连续交互来优化预测。通过器官监督将解剖背景纳入模型,使用一个共享的头部从相同特征中预测病灶和器官,从而减少因生理摄取引起的假阳性。此外,由于在推理时不提供示踪剂(即FDG/PSMA),我们添加了一个基于图像处理的示踪剂分类器和一个基于冠状MIP特征的随机森林,将每个研究路由到结合FDG+PSMA模型或PSMA特定模型。在四折交叉验证中,器官监督模型实现了最佳且最稳定的性能,交互阶段随着每个提示单调提高Dice得分,而PSMA特定训练则产生了最强的示踪剂相关结果。
cs.CV / 71 / 2608.28517
Learning the Target Priors Before Image Translation: A Decoupled Training Paradigm for Cross-Modal Image Translation in Remote Sensing
在图像翻译之前学习目标先验:一种用于遥感跨模态图像翻译的解耦训练范式
Abstract
Cross-modal image translation in remote sensing must preserve source-observed content while matching the target-domain distribution. Existing methods jointly learn the target prior and cross-modal dependence from scarce paired data, overlooking a key asymmetry: only the latter intrinsically requires cross-modal correspondence. We formalize this distinction through conditional-score and denoising-risk analyses and propose Learning the Target Priors Before Image Translation (LTP-BIT), a prior-first paradigm that decouples the two learning tasks. LTP-BIT first learns a target-domain generative prior from large-scale unpaired imagery, then retains the pretrained backbone weights and learns source-conditioned control through P-DART, a parameter-efficient dual-stream architecture. Controlled experiments show that prior matching and scaling primarily improve target-domain realism, whereas instance fidelity relies more strongly on conditional adaptation. LTP-BIT achieves state-of-the-art performance across SAR-to-RGB and NIR-to-RGB benchmarks using only 9.81% task-specific parameters. On QXS-SAROPT, it retains near-full-data instance fidelity with only 25% of the paired samples.
Chinese Translation
遥感中的跨模态图像翻译必须在匹配目标领域分布的同时保留源观测内容。现有方法通过稀缺的配对数据共同学习目标先验和跨模态依赖,忽视了一个关键的不对称性:只有后者本质上需要跨模态对应。我们通过条件评分和去噪风险分析形式化了这一区别,并提出了在图像翻译之前学习目标先验(Learning the Target Priors Before Image Translation, LTP-BIT)的方法,这是一种优先学习的范式,解耦了这两个学习任务。LTP-BIT首先从大规模非配对图像中学习目标领域生成先验,然后保留预训练的主干权重,并通过P-DART(参数高效的双流架构)学习源条件控制。受控实验表明,先验匹配和缩放主要改善目标领域的真实感,而实例保真度则更依赖于条件适应。LTP-BIT在SAR到RGB和NIR到RGB的基准测试中实现了最先进的性能,仅使用了9.81%的任务特定参数。在QXS-SAROPT上,它在仅使用25%的配对样本的情况下保持了接近全数据的实例保真度。
cs.CV / 72 / 2608.28524
Texture Image Classification Using DWT AlexNet Feature Fusion and Deep Neural Networks
基于DWT和AlexNet特征融合的纹理图像分类
Abstract
Texture image classification plays a significant role in computer vision applications, including industrial inspection, medical image analysis, remote sensing, and object recognition. Handcrafted features can capture local texture characteristics but may have limited capability to represent complex visual patterns. In contrast, deep learning models automatically learn discriminative representations but may not fully exploit the multiscale spatial-frequency information inherent in texture images. This paper proposes a hybrid feature fusion framework, termed DWT_AlexNet_DNN, which combines Discrete Wavelet Transform (DWT) features with deep features extracted using AlexNet for texture image classification.
Chinese Translation
纹理图像分类在计算机视觉应用中扮演着重要角色,包括工业检测、医学图像分析、遥感和物体识别。手工特征能够捕捉局部纹理特征,但在表示复杂视觉模式方面可能能力有限。相比之下,深度学习模型能够自动学习判别性表示,但可能未能充分利用纹理图像中固有的多尺度空间频率信息。本文提出了一种混合特征融合框架,称为DWT_AlexNet_DNN,该框架将离散小波变换(DWT)特征与使用AlexNet提取的深度特征相结合,用于纹理图像分类。
cs.CV / 73 / 2608.28549
Video Generative Models as Geometry Learner
视频生成模型作为几何学习者
Abstract
Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-specific geometry models (for depth and surface normal estimation) independently, losing the opportunity of exploring the intrinsic correlation of these geometric targets, or (ii) jointly fine-tune modified image diffusion backbones (e.g., altered self-attention), which typically demands substantial labeled data. To overcome these limitations in a principled fashion, we repurpose pretrained video generative models as a unified and data-efficient framework for geometry estimation, formulated innovatively as a next-frames prediction task. Our method, GeoNeXt, inherits naturally structured knowledge and richer priors from the video model, while further adapting them for joint modeling of images and geometry targets (image <-> geometry), enabling more data efficient and effective learning of geometry. Extensive experiments validate our method for zero-shot monocular depth and surface normal estimation across diverse datasets, outperforming both previous task-specific and unified generative competitors while using substantially less training data. Notably, our method rivals discriminative state-of-the-art approaches trained on over 100x more data and even standouts on several benchmarks.
Chinese Translation
最近的几何估计生成方法适应了预训练的图像扩散模型,并将任务视为图像条件生成。利用现成的图像扩散模型,它们要么 (i) 独立训练特定任务的几何模型(用于深度和表面法线估计),失去了探索这些几何目标内在关联的机会,要么 (ii) 联合微调修改过的图像扩散骨干网络(例如,改变的自注意力),这通常需要大量标记数据。为了以原则性的方式克服这些局限性,我们重新利用预训练的视频生成模型,作为一个统一且数据高效的几何估计框架,创新性地将其表述为下一个帧预测任务。我们的方法 GeoNeXt 自然继承了视频模型中结构化的知识和更丰富的先验,同时进一步调整它们以实现图像与几何目标的联合建模(图像 <-> 几何),从而实现更高效和有效的几何学习。大量实验验证了我们的方法在不同数据集上进行零-shot 单目深度和表面法线估计的有效性,超越了之前的特定任务和统一生成竞争者,同时使用的训练数据显著更少。值得注意的是,我们的方法在训练数据超过 100 倍的情况下与判别性最先进的方法相媲美,并在多个基准测试中表现突出。
cs.CV / 74 / 2608.28567
GeBDA: Building Damage Assessment as Text-Based Sequence Prediction
GeBDA:基于文本的建筑损伤评估作为序列预测
Abstract
Conventionally, Building Damage Assessment (BDA) is tackled either with dedicated network architectures or by fine-tuning geospatial image foundation models. In this work, we ask whether a general-purpose Vision-Language Model (VLM) can localize buildings and grade their damage through autoregressive sequence generation alone. We cast BDA as predicting a variable-length set of bounding boxes, each specified by its coordinates and a damage label. Our preliminary implementation, based on the open Gemma model, achieves promising damage mapping results from only bi-temporal satellite images and a suitable text prompt.
Chinese Translation
传统上,建筑损伤评估(Building Damage Assessment, BDA)主要通过专用网络架构或微调地理空间图像基础模型来解决。在本研究中,我们探讨了通用视觉-语言模型(Vision-Language Model, VLM)是否能够仅通过自回归序列生成来定位建筑物并评估其损伤程度。我们将建筑损伤评估视为预测一组可变长度的边界框,每个边界框由其坐标和损伤标签指定。基于开放的Gemma模型,我们的初步实现仅使用双时相卫星图像和适当的文本提示,取得了令人鼓舞的损伤映射结果。
cs.CV / 75 / 2608.28568
SignRR: Retrieve and Refine Real Motion for Sign Language Production
SignRR:检索与精炼真实运动以实现手语生成
Abstract
Sign language production (SLP) aims to generate continuous signing motion from spoken language, often through gloss-to-pose generation. Prior work mainly follows two paradigms. Generative models synthesize motion from a learned prior or from noise, without reference to an observed signing instance, making rare hand configurations and signer-specific articulation difficult to preserve. Retrieval-based methods reuse real, well-articulated motion segments, but concatenating segments from different signers and co-articulation contexts can introduce rhythm and style inconsistencies across the full sequence, not only at segment boundaries. These limitations suggest a complementary solution: use retrieval to provide realistic articulation, and use learned refinement to impose the global coherence that retrieval alone lacks. We therefore propose retrieve-and-refine, a paradigm that starts from real retrieved motion and refines it into a globally coherent signing sequence rather than generating motion from scratch. Our framework, SignRR, initializes motion from a dictionary of real sign segments and refines the full sequence with a part-aware Residual VQ-VAE, where residual quantization preserves fine hand articulation and temporal length differences are handled in the latent space. Experiments on PHOENIX14T and CSL-Daily show that SignRR achieves state-of-the-art back-translation performance while maintaining competitive pose quality.
Chinese Translation
手语生成(SLP)旨在从口语中生成连续的手势运动,通常通过从词汇到姿态的生成进行实现。先前的研究主要遵循两种范式。生成模型从学习的先验或噪声中合成运动,而不参考观察到的手势实例,这使得稀有的手部配置和特定手势者的发音难以保留。基于检索的方法重用真实的、表达清晰的运动片段,但从不同手势者和共同发音上下文中连接片段可能会在整个序列中引入节奏和风格的不一致,不仅仅是在片段边界。这些局限性提示了一种互补的解决方案:利用检索提供真实的发音,并使用学习的精炼来施加检索单独缺乏的全局一致性。因此,我们提出了检索与精炼的范式,该范式从真实的检索运动开始,并将其精炼为一个全局一致的手势序列,而不是从头生成运动。我们的框架SignRR从真实手势片段的字典初始化运动,并使用部分感知的残差VQ-VAE对整个序列进行精炼,其中残差量化保留了细微的手部发音,而时间长度差异则在潜在空间中处理。在PHOENIX14T和CSL-Daily上的实验表明,SignRR在保持竞争性姿态质量的同时,达到了最先进的反向翻译性能。
cs.AI / 1 / 2608.27459
Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model
可测试人类知识的时间胶囊:41年《危险边缘!》在单一自由本地模型中的体现
Abstract
In 2011, IBM's Watson was something like a sealed capsule of its era's queryable knowledge. Its DeepQA system defeated the strongest human Jeopardy! champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a cluster of POWER7 servers, frozen at build time and impossible to move or copy. We show that the same kind of artifact, a snapshot of what a culture can answer, is now portable and essentially free. We evaluate a single 9 GB open-weight model (Qwen2.5-14B, 4-bit) against the complete open Jeopardy! clue dataset, 529,939 clues across all 41 broadcast seasons from 1984 to 2025. To our knowledge this is the first time a model has been run over the full corpus. The 41 years mark only how long the questions were collected. What they test is far older and broader: the accumulated body of human general knowledge a culture considers worth knowing, from ancient history and dead languages to science, literature, and geography, with a verified answer for every item. The model answers 67.0% of all clues under a strict forced-response protocol with exact and fuzzy matching, and exceeds 85% on factoid categories. We treat training-data exposure as something both systems share rather than a flaw unique to language models. Watson's case is in fact the more extreme one. Its corpus was assembled to contain Jeopardy answers and it was tuned on past clues, and it could not answer anything outside that curated distribution. The decisive test is whether a model can answer clues that did not exist when it was built. On clues aired after its training cutoff, the local model holds 65% and Claude Opus 4.8 holds 95%, while Watson by construction scores zero. The capability survives the move from a server room to a file you could seal in a time capsule, and unlike Watson it is not frozen to its own moment.
Chinese Translation
在2011年,IBM的沃森(Watson)就像是一个封闭的胶囊,承载着那个时代可查询的知识。它的DeepQA系统击败了最强的人类《危险边缘!》冠军,但使其能够做到这一点的知识存储在一个经过策划的十亿文档语料库中,该语料库运行在一组POWER7服务器上,在构建时被冻结,无法移动或复制。我们展示了同样类型的文物,即文化能够回答的快照,现在是可移植的且基本上是免费的。我们评估了一个9 GB的开放权重模型(Qwen2.5-14B,4-bit),与完整的开放《危险边缘!》线索数据集进行对比,该数据集涵盖了1984年至2025年间的41个播出季节,共计529,939个线索。据我们所知,这是首次在完整语料库上运行模型。41年的时间仅仅标志着问题的收集时长。它们所测试的内容则更为古老和广泛:一个文化认为值得了解的人类一般知识的累积,包括古代历史、已灭绝语言、科学、文学和地理,每个项目都有经过验证的答案。在严格的强制响应协议下,该模型对所有线索的回答正确率为67.0%,在事实类问题上超过85%。我们将训练数据的曝光视为两个系统共享的特性,而不是语言模型独有的缺陷。沃森的案例实际上是更极端的。它的语料库是为了包含《危险边缘!》的答案而组装的,并且在过去的线索上进行了调优,无法回答任何超出该策划分布的问题。决定性的测试是模型是否能够回答在构建时不存在的线索。在其训练截止后播出的线索中,本地模型的正确率为65%,而Claude Opus 4.8的正确率为95%,而沃森由于其构造得分为零。这种能力在从服务器房间转移到可以封存于时间胶囊的文件中时依然存在,并且与沃森不同,它并未被冻结在自己的时刻。
cs.AI / 2 / 2608.27463
Rating the Raters: Rasch Measurement Theory for LLM Evaluation
评估评估者:用于大语言模型评估的拉施测量理论
Abstract
LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by raters. Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured. Rasch measurement theory (RMT) is well-suited to this kind of problem. RMT decomposes ordinal ratings into separable facets on a common scale. It further provides a battery of diagnostics that can identify miscalibrated measurements and rater biases. We present a case study of RMT applied to the LLM-as-rater paradigm using the Measuring Hate Speech corpus, whose construct was itself built under RMT. We fit a series of many-facet Rasch models to annotations from nine LLMs spanning families and capability levels. Our analyses show that LLMs systematically differ from human raters in severity, item-level calibration, question-order robustness, target-identity sensitivity, and rating scale use, which all would be obscured by standard evaluation practice. Overall, we argue that RMT belongs in the toolkit for evaluating LLM-as-examinee, -judge, and -rater paradigms.
Chinese Translation
大语言模型(LLMs)在评估中扮演着多重角色:作为基准测试中的被评估者、其他模型输出的评判者,以及人类生成内容的评分者。每种范式都可以视为一个测量问题,其中通过评估者使用工具(例如基准)对对象的潜在属性进行探测。标准评估实践往往忽视每个核心组件对最终结果的贡献,从而限制了我们对所测量内容的理解。拉施测量理论(RMT)非常适合解决此类问题。RMT将有序评分分解为可分离的方面,并在一个共同的尺度上进行分析。它还提供了一系列诊断工具,可以识别校准不当的测量和评分者偏见。我们展示了一个将RMT应用于LLM作为评分者范式的案例研究,使用了在RMT框架下构建的仇恨言论测量语料库。我们对来自九个不同家族和能力水平的LLMs的注释数据拟合了一系列多面拉施模型。我们的分析表明,LLMs在严厉程度、项目级别校准、问题顺序稳健性、目标身份敏感性和评分尺度使用等方面与人类评分者存在系统性差异,这些差异在标准评估实践中会被掩盖。总体而言,我们认为RMT应当成为评估LLM作为被评估者、评判者和评分者范式的工具箱中的一部分。
cs.AI / 3 / 2608.27464
Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAI
并非所有解释都是被寻求的:以人为本的可解释人工智能的信息寻求心理学
Abstract
This position paper argues that human-centered explainable AI (HCXAI) should incorporate insights from the psychology of information seeking. Drawing on Sharot and Sunstein's framework of information-seeking motives, we propose that people evaluate whether to engage with explanations based on three types of expected utility: instrumental (will it help me act better?), hedonic (will it make me feel better?), and cognitive (will it improve my understanding?). Each utility is estimated through a lens shaped by well-documented cognitive biases, including illusion of control, automation bias, unrealistic optimism, impact bias, overconfidence, and confirmation bias. These biases can lead to two failure modes: excessive information-seeking that fragments attention without improving decisions, and insufficient information-seeking that leaves critical risks and misunderstandings unexamined. This challenge is particularly acute for agentic AI systems, where explanations must support not just understanding a single output but anticipating cascading actions, assessing risks, and deciding when to intervene. By integrating information-seeking psychology into HCXAI, we advocate for a shift from making explanations available to making them sought: designing systems that account for when and why users actually want to know.
Chinese Translation
本文立场论文认为,以人为本的可解释人工智能(HCXAI)应当融入信息寻求心理学的见解。基于Sharot和Sunstein的信息寻求动机框架,我们提出人们在决定是否参与解释时,会基于三种预期效用进行评估:工具性(这会帮助我更好地行动吗?)、享乐性(这会让我感觉更好吗?)和认知性(这会提高我的理解吗?)。每种效用的评估都受到已被充分记录的认知偏差的影响,包括控制错觉、自动化偏差、不切实际的乐观、影响偏差、过度自信和确认偏差。这些偏差可能导致两种失败模式:过度的信息寻求使注意力分散而未改善决策,以及不足的信息寻求使得关键风险和误解未被检视。这一挑战在自主AI系统中尤为突出,因为解释不仅需要支持对单一输出的理解,还需预测连锁反应、评估风险以及决定何时干预。通过将信息寻求心理学融入HCXAI,我们倡导从提供解释转向寻求解释:设计考虑用户何时以及为何真正想要了解的系统。
cs.AI / 4 / 2608.27471
Retrieving Relations, Detecting Fallacies: A RAG Approach to Political Debate Analysis
检索关系,检测谬误:一种基于RAG的方法进行政治辩论分析
Abstract
Fallacies are arguments that employ invalid reasoning, making their automatic detection critical in sensitive contexts such as high-stakes political debates, where public opinion is shaped. Spotting a fallacious argument requires contextual knowledge beyond its pure surface text. This entails world knowledge pertaining to the subject matter under discussion, as well as knowledge of the relationships that exist between arguments within the argumentative discourse. Prior work on fallacy analysis has shown that argumentative discourse structure can beneficially improve classification performance. However, such structure is typically encoded only as static classifier features, limiting its flexibility. Building on this intuition while addressing this limitation, we introduce a guided retrieval-augmented methodology for fallacy detection and classification that leverages argumentative relations of support and attack to dynamically steer the extraction of relevant documents. We evaluate our approach on the ElecDeb60to20 benchmark across 42 retrieval configurations and 14 models, performing retrieval over a 15GB knowledge base of collected political-related documents. Our approach improves macro-F1 up to 0.864 for fallacy detection and up to 0.725 for classification over non-retrieval baselines. These results show that incorporating external knowledge significantly enhances fallacy detection and classification when retrieval is argumentatively guided.
Chinese Translation
谬误是使用无效推理的论证,因此在高风险的政治辩论等敏感场合中,自动检测谬误显得尤为重要,因为这些场合会影响公众舆论。识别谬误论证需要超越表面文本的上下文知识。这涉及到与讨论主题相关的世界知识,以及论证话语中论证之间存在的关系。以往的谬误分析研究表明,论证话语结构可以有效提高分类性能。然而,这种结构通常仅作为静态分类器特征进行编码,限制了其灵活性。在此基础上,我们提出了一种引导的检索增强方法,用于谬误检测和分类,该方法利用支持和攻击的论证关系动态引导相关文档的提取。我们在ElecDeb60to20基准上对我们的方法进行了评估,涵盖42种检索配置和14种模型,在一个包含15GB政治相关文档的知识库上进行检索。我们的方法在谬误检测方面的宏观F1值提高至0.864,在分类方面提高至0.725,相较于非检索基线。这些结果表明,当检索受到论证引导时,结合外部知识显著增强了谬误检测和分类的效果。
cs.AI / 5 / 2608.27472
LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation
LLM增强的因果发现:边缘存在性和方向性的概率融合
Abstract
Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offer broad but often unreliable causal knowledge. We propose combining these complementary sources through a novel representation, termed Probabilistic Dependency Graphs (PDGs). In a PDG, each edge is associated with a distribution over directed, undirected, and absent states, enabling fusion via weighted averaging. We evaluate this approach on 26 benchmark networks, combining ensembles of three BNSL algorithms (FGES, Tabu, PC) with three LLMs (Gemini, Claude, GPT) across multiple prompts and random seeds. A simple 50/50 fusion improves F1 over the better of either source alone in 22 of 26 networks, with a statistically significant mean improvement of $0.056$ $(p<0.001)$. Analysis reveals that the two sources play complementary roles: BNSL contributes a high-recall edge skeleton (80\% vs 60\% for LLM), while LLM contributes accurate edge orientation (96\% vs 77\% for BNSL). Our results show that representing both sources as probabilistic uncertainty over edge existence and orientation is a practical and effective way to improve causal graph accuracy.
Chinese Translation
从观察数据中进行贝叶斯网络结构学习(BNSL)面临方向可识别性的问题,而大型语言模型(LLMs)提供了广泛但往往不可靠的因果知识。我们提出通过一种新颖的表示方法,称为概率依赖图(Probabilistic Dependency Graphs, PDGs),将这些互补来源结合起来。在PDG中,每条边与有向、无向和缺失状态的分布相关联,从而通过加权平均实现融合。我们在26个基准网络上评估了这种方法,结合了三种BNSL算法(FGES、Tabu、PC)和三种LLM(Gemini、Claude、GPT),并在多个提示和随机种子下进行测试。简单的50/50融合在26个网络中有22个网络的F1值超过了任一来源的最佳结果,平均显著提高了$0.056$ $(p<0.001)$。分析表明,这两种来源发挥了互补作用:BNSL提供了高召回率的边缘骨架(80 ext{%}对比LLM的60 ext{%}),而LLM则提供了准确的边缘方向(96 ext{%}对比BNSL的77 ext{%})。我们的结果表明,将这两种来源表示为边缘存在性和方向性的概率不确定性是一种实用且有效的提高因果图准确性的方法。
cs.AI / 6 / 2608.27475
Hypothesize, Evaluate, Refine: A Scientific Agent for PDE Discovery with Unknown Spatial Coefficient Fields
假设、评估、精炼:一种用于未知空间系数场的偏微分方程发现的科学代理
Abstract
Discovering PDEs in heterogeneous media requires jointly identifying the governing operator and the unknown spatial fields that parameterize it. These tasks are coupled: changing field placement changes the differential law, while a sufficiently flexible field can conceal structural error on a single trajectory. We present Hypothesize, Evaluate, Refine for PDE Discovery (HER-PDE), a scientific-agent framework that discovers compositional PDE structure together with nonparametric, time-invariant coefficient fields. The Agent analyzes two noisy trajectories generated by different excitations, proposes complete expression-tree hypotheses, and combines creative structural exploration with local candidate refinement. Its Hypothesis Evaluation Interface (HEI) estimates only the fields explicitly declared in each hypothesis, never adds missing terms, and scores structures by bidirectional cross-excitation transfer. The selected law is subsequently audited on a sealed temporal interval. Across five controlled two-dimensional systems observed with 5 percent relative Gaussian state noise, the Agent recovers the generating operator in all five cases, including equivalent signed-field and product-rule parameterizations. Across nine unknown coefficient fields, the recovered fields attain a median Pearson correlation of approximately 0.85 and a median relative L2 error of approximately 0.28. These results show that agent-guided hypothesis refinement can recover heterogeneous governing laws without prescribing a parametric form for their spatial coefficients.
Chinese Translation
在异质介质中发现偏微分方程(PDE)需要共同识别控制算子和参数化的未知空间场。这些任务是相互关联的:改变场的放置会改变微分定律,而足够灵活的场可以在单一路径上掩盖结构误差。我们提出了用于PDE发现的假设、评估、精炼(HER-PDE)框架,这是一个科学代理框架,能够发现组合的PDE结构以及非参数化的时间不变系数场。该代理分析由不同激励生成的两个噪声轨迹,提出完整的表达树假设,并将创造性的结构探索与局部候选者精炼相结合。其假设评估接口(HEI)仅估计每个假设中明确声明的场,从不添加缺失项,并通过双向交叉激励传递对结构进行评分。随后,在一个封闭的时间间隔内对所选定律进行审核。在五个受控的二维系统中,观察到相对高斯状态噪声为5%的情况下,该代理在所有五个案例中恢复了生成算子,包括等效的符号场和乘积规则参数化。在九个未知系数场中,恢复的场的中位皮尔逊相关系数约为0.85,中位相对L2误差约为0.28。这些结果表明,代理引导的假设精炼能够在不规定其空间系数的参数形式的情况下恢复异质控制定律。
cs.AI / 7 / 2608.27476
Class-Based Heuristic Selection for Solving the Flying Block Puzzle
基于类别的启发式选择用于解决飞行块拼图
Abstract
Heuristic search underlies planning in autonomous systems ranging from warehouse logistics to robotic navigation, yet generic heuristics fail to exploit the structural constraints that govern constrained spatial domains, causing search performance to degrade catastrophically on harder instances. We study this problem through the two-column Flying Block Puzzle, a rigorously NP-complete spatial planning microworld whose bottleneck geometry mirrors clearance-to-size constraints encountered in multi-agent path finding, autonomous vehicle navigation, and block relocation systems. We introduce the Class-Based Heuristic A* (CBHA*) algorithm, which integrates a General Move Constraint to capture minimum displacement costs when vacant units are scarce, a formal kinematic taxonomy partitioning the state space into seven mutually exclusive classes with provably admissible heuristics based on vacancy ratio and goal-piece geometry, and a class-conditional tie-breaking mechanism that dynamically switches between depth-priority and vertical-distance ordering to overcome f-value plateaus. Over 146 benchmark instances, CBHA* achieves a 93.4% success rate against 64% for Depth-Prioritized A*, 39% for Standard A*, and 17% for BFS, while reducing node expansions by 87.98% relative to Standard A* and sustaining an average effective branching factor of approximately 3, demonstrating that class-triggered adaptive heuristics constitute a principled mechanism for efficient spatial planning that generalizes structurally to physical constraint systems.
Chinese Translation
启发式搜索是从仓库物流到机器人导航等自主系统规划的基础,但通用启发式方法未能利用约束空间领域中支配的结构性约束,导致在更难实例上的搜索性能急剧下降。我们通过双列飞行块拼图研究这一问题,该拼图是一个严格的 NP 完全空间规划微观世界,其瓶颈几何形状反映了在多智能体路径寻找、自动驾驶导航和块重定位系统中遇到的清除与尺寸约束。我们提出了基于类别的启发式 A* (CBHA*) 算法,该算法整合了一种通用移动约束,以捕捉在空闲单元稀缺时的最小位移成本,采用一种形式的运动学分类法将状态空间划分为七个相互独立的类别,并基于空闲比率和目标块几何形状提供可证明的可接受启发式,同时引入了一种类别条件的平局打破机制,动态切换深度优先和垂直距离排序,以克服 f 值平台。在 146 个基准实例中,CBHA* 的成功率达到了 93.4%,而深度优先 A* 为 64%,标准 A* 为 39%,广度优先搜索 (BFS) 为 17%,同时相较于标准 A* 减少了 87.98% 的节点扩展,并维持了大约 3 的平均有效分支因子,证明了基于类别触发的自适应启发式方法构成了一种高效空间规划的原则性机制,能够在结构上推广到物理约束系统。
cs.AI / 8 / 2608.27477
Benchmarking General Mobile Assistants in Challenging Real-World Scenarios
在具有挑战性的真实世界场景中对通用移动助手进行基准测试
Abstract
Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks. Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but their application coverage and task design do not yet fully capture the diversity and complexity of realistic mobile use. We present GMA, a benchmark for evaluating general mobile assistants in challenging real-world scenarios. GMA introduces seven applications based on open-source projects, spanning domains such as lifestyle sharing and travel planning, and 300 tasks across four difficulty tiers, from atomic actions to complex multi-step workflows. We evaluate eight frontier models and find that performance declines substantially as task complexity increases, with current agents remaining far from reliably handling realistic user requirements. We further conduct controlled ablation studies of agent harness choices, including context retention and explicit state tracking, under a shared environment, model setting, and task taxonomy. Results show that appropriate harness design can meaningfully improve performance, particularly on demanding workflows, while the effectiveness of specific designs can vary across foundation models. Overall, GMA complements existing benchmarks by expanding application coverage and task complexity, providing a challenging testbed for evaluating mobile agents and studying how harness design supports reliable execution in complex mobile workflows.
Chinese Translation
图形用户界面已成为评估自主人工智能代理在多模态交互任务中的重要环境。现有基准如AndroidWorld和MobileWorld为移动代理评估提供了坚实的基础,但它们的应用覆盖范围和任务设计尚未完全捕捉现实移动使用的多样性和复杂性。我们提出了GMA,这是一个用于在具有挑战性的真实世界场景中评估通用移动助手的基准。GMA引入了七个基于开源项目的应用,涵盖生活方式分享和旅行规划等领域,并设计了300个任务,分为四个难度等级,从原子操作到复杂的多步骤工作流。我们评估了八个前沿模型,发现随着任务复杂性的增加,性能显著下降,目前的代理在可靠处理现实用户需求方面仍然相距甚远。我们进一步在共享环境、模型设置和任务分类下,对代理工具选择进行了受控消融研究,包括上下文保留和显式状态跟踪。结果表明,适当的工具设计可以显著提高性能,尤其是在要求较高的工作流中,而特定设计的有效性在基础模型之间可能有所不同。总体而言,GMA通过扩展应用覆盖范围和任务复杂性,补充了现有基准,为评估移动代理和研究工具设计如何支持复杂移动工作流中的可靠执行提供了一个具有挑战性的测试平台。
cs.AI / 9 / 2608.27480
Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea Plantations
物联网与深度学习在茶园中对电蚁(Postelectrotermes militaris)检测及严重性评估的有效性
Abstract
Tea plantations are vulnerable to Postelectrotermes militaris, commonly known as the Upcountry Live Wood Termite (ULWT), which can cause substantial damage when infestations remain undetected. This study proposes an IoT-enabled acoustic monitoring framework integrated with deep learning for early detection and severity assessment of ULWT infestations in tea plantations. Research Method: Audio signals were captured non-invasively from tea trunks using a high-sensitivity microphone connected to a Raspberry Pi-based IoT device, with geographic coordinates recorded for spatial tracking. After trimming, resampling, and segmentation, 2,000 ten-second samples were obtained, comprising 1,000 healthy and 1,000 infested samples, and divided into 1,600 training, 200 validation, and 200 test samples. The dataset used in this study is publicly available on Kaggle (Senevirathna et al. 2026). Fourier-derived spectrograms trained a CNN for infestation classification and probability estimation. A weighted severity model combined CNN probability, mean acoustic amplitude, and nearby infested plants within 5 m, with geospatial mapping used to visualize infestation distribution. Findings and Values: Field trials in a ULWT-affected tea plantation in Pundaluoya demonstrated feasibility under realistic environmental noise. On the held-out test set, the CNN achieved 81.5% accuracy, 80.6% precision, 83.0% recall, 81.8% F1-score, and 0.819 ROC-AUC. Beyond binary infestation detection, the framework introduced quantitative severity assessment using infestation probability, acoustic amplitude, and nearby infested plants. The resulting severity and geospatial outputs can support plantation managers in identifying high-risk areas, prioritizing field inspections, and implementing more timely and targeted control measures.
Chinese Translation
茶园易受到电蚁(Postelectrotermes militaris)的侵害,通常被称为高地活木白蚁(Upcountry Live Wood Termite,ULWT),当虫害未被及时发现时可能造成重大损失。本研究提出了一种基于物联网的声学监测框架,结合深度学习用于茶园中ULWT虫害的早期检测和严重性评估。研究方法:通过连接至基于Raspberry Pi的物联网设备的高灵敏度麦克风,从茶树树干非侵入性地捕获音频信号,并记录地理坐标以进行空间追踪。经过修剪、重采样和分段处理,获得了2,000个十秒样本,其中包括1,000个健康样本和1,000个受害样本,并分为1,600个训练样本、200个验证样本和200个测试样本。本研究使用的数据集在Kaggle上公开可用(Senevirathna et al. 2026)。通过傅里叶变换生成的声谱图训练了卷积神经网络(CNN)以进行虫害分类和概率估计。加权严重性模型结合了CNN概率、平均声学幅度以及5米范围内的附近受害植物,并使用地理空间映射可视化虫害分布。研究结果与价值:在Pundaluoya一处受ULWT影响的茶园进行的实地试验表明,在现实环境噪声下的可行性。在保留的测试集上,CNN达到了81.5%的准确率、80.6%的精确率、83.0%的召回率、81.8%的F1分数和0.819的ROC-AUC。除了二元虫害检测外,该框架还引入了基于虫害概率、声学幅度和附近受害植物的定量严重性评估。最终的严重性和地理空间输出可以帮助种植园管理者识别高风险区域,优先安排田间检查,并实施更及时和针对性的控制措施。
cs.AI / 10 / 2608.27482
Context Localization for Generalized Level-Based Evaluation in Knowledge-Based Systems
知识基础系统中基于上下文的广义水平评估
Abstract
We study context localization for generalized level-based evaluation in knowledge-based systems. The framework models situations where a structured nonnegative score, defined on facts, rules, cases, criteria or evidence units, is evaluated through conditional aggregation tests on admissible knowledge contexts. The generalized level measure maximizes a monotone set function over all contexts whose aggregated support reaches a prescribed level. We characterize when filtering the score by a context $B$ is equivalent to localizing the admissible contexts by intersection with $B$. The main theorem shows that this consistency holds for all monotone set functions if and only if two structural conditions are satisfied: monotonicity with respect to contexts and a reduction property excluding positive localized support outside $B$. We analyze pointwise and block-generated mechanisms producing the reduction property, extend the result to parameterized systems, and interpret it as a stability criterion for context-dependent evidence selection, non-additive support evaluation and level-based knowledge aggregation.
Chinese Translation
我们研究了知识基础系统中基于上下文的广义水平评估。该框架建模了在可接受知识上下文中,通过条件聚合测试评估定义在事实、规则、案例、标准或证据单元上的结构化非负分数的情境。广义水平度量最大化了在所有聚合支持达到规定水平的上下文上定义的单调集合函数。我们描述了通过上下文 $B$ 过滤分数与通过与 $B$ 的交集本地化可接受上下文之间的等价关系。主要定理表明,当且仅当满足两个结构条件时,这种一致性对于所有单调集合函数成立:相对于上下文的单调性和排除 $B$ 外部正向本地化支持的约简性质。我们分析了产生约简性质的逐点和块生成机制,将结果扩展到参数化系统,并将其解释为上下文依赖证据选择、非加性支持评估和基于水平的知识聚合的稳定性标准。
cs.AI / 11 / 2608.27484
CareGraph: An Auditable Hybrid AI Framework for Evidence-Grounded Personalized Longitudinal Health Intelligence
CareGraph:一个可审计的混合人工智能框架,用于基于证据的个性化纵向健康智能
Abstract
Artificial intelligence is transforming personalized healthcare, yet fragmented clinical, self reported, and wearable evidence remains difficult to interpret and trace. We present CareGraph, an auditable hybrid AI framework that converts heterogeneous records into prioritized trends, missing context indicators, bounded next steps, discussion questions, and provenance linked explanations. CareGraph organizes evidence without diagnosing, predicting outcomes, selecting treatment, or making autonomous clinical decisions. Its pipeline covers deterministic analysis, context detection, graph construction, constrained language model synthesis, evidence validation, safety controls, and release gating. Tests used synthetic cohorts of 400 patients each for development, validation, and holdout. On holdout data, a frozen ordinary least squares trend rule with a sufficiency gate achieved 0.827 accuracy, 0.837 macro F1 with a 95 percent confidence interval of 0.819 to 0.854, and 0.974 insufficient data F1. Missing context detection achieved 0.815 strict micro F1 versus 0.318 for the legacy detector. On an authored holdout benchmark, safety ruleset version 1.2 achieved 1.000 precision, 0.950 recall, and 0.974 F1. An audit requiring graph retrieval across 80 patients yielded 79 syntheses and 78 presentations without fallback; one output was blocked and one failed closed because of an invalid evidence key. Against monolithic GPT 5.6 on 56 matched patients, CareGraph was faster at 40.15 versus 49.62 seconds, shorter at 661 versus 1,163 words, and showed better exploratory lexical alignment with longitudinal targets; the baseline used fewer tokens and cited more raw evidence. Graph auditing verified provenance and deterministic retrieval; incremental graph effects on generation require paired evaluation. CareGraph offers a safety bounded foundation for intelligent personalized health systems.
Chinese Translation
人工智能正在改变个性化医疗,但碎片化的临床、自我报告和可穿戴证据仍然难以解读和追踪。我们提出了CareGraph,一个可审计的混合人工智能框架,它将异构记录转换为优先趋势、缺失上下文指示器、有限的下一步行动、讨论问题和来源关联的解释。CareGraph在不进行诊断、预测结果、选择治疗或做出自主临床决策的情况下组织证据。其流程涵盖确定性分析、上下文检测、图构建、约束语言模型合成、证据验证、安全控制和发布门控。测试使用了400名患者的合成队列进行开发、验证和保留。在保留数据上,一个冻结的普通最小二乘趋势规则与充分性门控达到了0.827的准确率,0.837的宏F1,95%的置信区间为0.819到0.854,以及0.974的不足数据F1。缺失上下文检测的严格微F1达到了0.815,而传统检测器为0.318。在一个作者创建的保留基准上,安全规则集版本1.2达到了1.000的精确率、0.950的召回率和0.974的F1。对80名患者进行图检索的审计产生了79个综合结果和78个展示,没有后备;一个输出被阻止,一个因证据密钥无效而关闭。与56名匹配患者的单体GPT 5.6相比,CareGraph在40.15秒的时间内完成,而GPT 5.6为49.62秒,字数为661对比1,163,并且在纵向目标的探索性词汇对齐上表现更佳;基线使用了更少的标记并引用了更多的原始证据。图审计验证了来源和确定性检索;生成的增量图效应需要配对评估。CareGraph为智能个性化健康系统提供了一个安全受限的基础。
cs.AI / 12 / 2608.27506
Thinking Costs Tokens: When More Structure is Worth the Price
Abstract
Adding inference structure to a language model lets it search, verify, and revise, but these actions consume the very budget they are supposed to use well. In this paper, we investigate whether there exists a token-budget threshold, below which the overhead of planning and verification hurts performance and above which it helps. We evaluate two systems on FinQA and TAT-QA financial reasoning tasks, using GPT-5.4 mini across 14 budget tiers ranging from 250 to 42,000 output-equivalent tokens. The first system is a monolith, which is a single LLM call. The second is a verified search architecture that adds planning, label-blind checking, and repair capabilities. We run 1,000 cases for a total of 28,000 completed cells. Both systems score 0% at the two lowest tiers, where neither can fit a complete prompt. At 1,000 tokens, the monolith reaches 18% accuracy while verified search scores near 0%, since the planning overhead leaves no room for an answer. From 1,500 tokens onward, verified search surpasses the monolith and maintains a consistent advantage, reaching approximately 44% at the highest tiers while the monolith reaches approximately 40%. The crossover occurs between 1,000 and 1,500 output-equivalent tokens, confirmed by a strict intersection-union test ($p \le 0.001$ at both endpoints).
cs.AI / 13 / 2608.27508
WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement Learning
WM-R1:训练图形用户界面代理以推理并利用世界模型的强化学习
Abstract
GUI agents trained with reinforcement learning (RL) have showcased strong environment learning capabilities on mobile platforms. However, RL typically demands extensive real-environment interactions, leading to high resource costs and instability, especially in GUI scenarios. To address these, we propose WM-R1, the first reinforcement learning framework that trains mobile GUI agents with world models instead of real environments. Specifically, world models serve as the source of state transitions during all rollouts, replacing the real Android environment within the training loop. WM-R1 also embeds world models directly into the thinking process, enabling agents to reason about the consequences of candidate actions before committing to the final action. Crucially, WM-R1 eliminates the need for real-environment interaction, supports massively parallelized and step-level granularized trajectory generation grounded in world models, and introduces a multi-dimensional rule-based reward that jointly optimizes task success, trajectory efficiency, and world model utilization. For efficient training, we curate a high-quality dataset of 2000 challenging tasks. Experiments on Android mobile benchmarks demonstrate that WM-R1-trained agents significantly outperform GRPO-only baselines and inference-time simulation methods. Code is available at https://github.com/genalyu/WM-R1 .
Chinese Translation
使用强化学习(RL)训练的图形用户界面(GUI)代理在移动平台上展示了强大的环境学习能力。然而,RL通常需要大量的真实环境交互,这导致高资源成本和不稳定性,尤其是在GUI场景中。为了解决这些问题,我们提出了WM-R1,这是第一个使用世界模型而非真实环境训练移动GUI代理的强化学习框架。具体而言,世界模型在所有回合中作为状态转移的来源,替代训练循环中的真实Android环境。WM-R1还将世界模型直接嵌入思考过程,使代理能够在最终行动之前推理候选行动的后果。至关重要的是,WM-R1消除了对真实环境交互的需求,支持基于世界模型的大规模并行和逐步细化的轨迹生成,并引入了一种多维基于规则的奖励机制,联合优化任务成功率、轨迹效率和世界模型的利用率。为了高效训练,我们策划了一个包含2000个挑战性任务的高质量数据集。在Android移动基准测试中的实验表明,WM-R1训练的代理显著优于仅使用GRPO的基线和推理时的仿真方法。代码可在 https://github.com/genalyu/WM-R1 获取。
cs.AI / 14 / 2608.27524
SETU: An Agentic Ecosystem for Multilingual, Persona-Aware Communication Coaching
SETU:一个面向多语言、个性化沟通辅导的代理生态系统
Abstract
Corporate training teams need scalable and explainable tools to improve workforce communication in multilingual settings. Existing systems often score text, audio, or video in isolation, or produce black-box outputs that are difficult to audit for coaching use. This paper presents SETU, an agentic ecosystem for corporate communication coaching aimed at recruiters, frontline sales professionals and training units who prepare for audience specific conversations. SETU is designed for two scoped scenarios: (i) recruiter-candidate eligibility-and-interest calls with persona context and (ii) sales pitches with target-audience adaptation; owing to limited evaluation resources, this paper reports results on scenario (ii) only. The ecosystem decomposes analysis into specialized video, audio-speech, text-relevance, scoring, notification and reporting agents coordinated through trust-aware orchestration. It generates modality-attributed coaching reports for formative training, with human reviewers retaining final judgment. The name SETU (bridge in several Indic languages) reflects the goal of bridging communication gaps across regional languages and audience expectations.
Chinese Translation
企业培训团队需要可扩展且可解释的工具,以改善多语言环境下的员工沟通。现有系统往往孤立地对文本、音频或视频进行评分,或产生难以审计的黑箱输出,难以用于辅导。本论文提出了SETU,一个面向企业沟通辅导的代理生态系统,旨在帮助招聘人员、前线销售专业人士和准备特定受众对话的培训单位。SETU设计了两个特定场景:(i)招聘人员与候选人的资格与兴趣电话沟通,结合个性化背景;(ii)针对目标受众的销售推介;由于评估资源有限,本文仅报告场景(ii)的结果。该生态系统将分析分解为专门的视频、音频语音、文本相关性、评分、通知和报告代理,通过信任感知的编排进行协调。它生成基于不同媒介的辅导报告,用于形成性培训,最终判断由人类评审者保留。SETU(在几种印度语言中意为“桥”)的名称反映了弥合区域语言和受众期望之间沟通差距的目标。
cs.AI / 15 / 2608.27548
Nemotron 3.5 Content Safety Moderator: A Compact Multimodal, Multilingual, and Reasoning Enabled Content Safety Moderator
Nemotron 3.5 内容安全审核器:一款紧凑的多模态、多语言且具备推理能力的内容安全审核器
Abstract
Safety moderation for deployed AI applications is moving beyond text-only prompts: systems increasingly need to judge images, documents, screenshots, and generated responses under policies that vary across domains. Existing guardrails usually cover only part of this setting, making it difficult to combine broad coverage, custom policy control, and low compute cost. We present Nemotron 3.5 Content Safety Moderator, also referred to as Nemotron 3.5 CS in this paper for brevity, a compact 4B vision-language safety moderator that jointly classifies user prompts, images, and assistant responses across 12 languages. Nemotron 3.5 CS returns safety labels for latency-sensitive moderation and can additionally produce concise reasoning traces that apply supplied custom policies and identify violated categories when reasoning is requested. We also release a multimodal and multilingual safety dataset for guard training, spanning human-labeled real-image moderation, benign vision-language and document tasks, synthetic rare-risk and jailbreak cases, and custom-policy examples. Across evaluations spanning multimodal safety, text moderation, multilingual robustness, custom-policy following, benign false positives, and latency, Nemotron 3.5 CS demonstrates a practical coverage tradeoff: it adds image-conditioned and policy-conditioned moderation while remaining broadly competitive with specialized guard models. These results suggest that compact vision-language moderators can serve as deployable front-line safety components, with reasoning used selectively for audit and policy review.
Chinese Translation
针对已部署的人工智能应用的安全审核正在超越仅限文本的提示:系统越来越需要根据跨领域的政策判断图像、文档、截图和生成的响应。现有的保护措施通常仅覆盖这一设置的一部分,使得结合广泛覆盖、自定义政策控制和低计算成本变得困难。我们提出了 Nemotron 3.5 内容安全审核器(在本文中简称为 Nemotron 3.5 CS),这是一款紧凑的 4B 视觉-语言安全审核器,能够在 12 种语言中共同分类用户提示、图像和助手响应。Nemotron 3.5 CS 为延迟敏感的审核返回安全标签,并且在请求推理时能够生成简洁的推理痕迹,应用提供的自定义政策并识别违规类别。我们还发布了一个多模态和多语言的安全数据集,用于保护训练,涵盖人类标注的真实图像审核、良性视觉-语言和文档任务、合成的稀有风险和越狱案例,以及自定义政策示例。在跨越多模态安全、文本审核、多语言鲁棒性、自定义政策遵循、良性误报和延迟的评估中,Nemotron 3.5 CS 展示了实用的覆盖权衡:它增加了基于图像和政策的审核,同时在与专门的保护模型竞争时仍保持广泛的竞争力。这些结果表明,紧凑的视觉-语言审核器可以作为可部署的前线安全组件,推理可选择性地用于审计和政策审查。
cs.AI / 16 / 2608.27580
LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails
LongGuard:长上下文失败的机制分析与无训练缓解方法在安全护栏中的应用
Abstract
Safety guardrails serve as the last line of defense against harmful inputs and outputs of large language models (LLMs), yet they are trained and evaluated almost exclusively on short text. We present LongGuard, a framework that evaluates, mechanistically analyzes, and mitigates long-context guardrail failure. We formulate the task as Safety Needle-in-a-Haystack (SafetyNIAH) over a 0.25k-32k length grid; across 15 mainstream guardrails, unsafe recall drops monotonically by more than 50% on average, and a paired Benign-Fill vs. Needle-Repeat design attributes the failure to proportional dilution of the unsafe needle rather than to absolute length. A three-layer attention-logit-behavior analysis on six guardrails locates the mechanism: attention mass on the unsafe needle is diluted, the unsafe-over-safe logit margin is compressed in lockstep, and the detection decision collapses accordingly, with this attention->logit->behavior chain remaining consistent after partialling out length. We further isolate a sparse set of guard-specialized retrieval heads that exhibit partial specificity relative to their base models. Building on the analysis, we propose two training-free mitigations - Chunked Detection (CD) and Attention-Head Sharpening (AHS) - and a deployment protocol, Context-Aware Hyperparameter Routing (CAHR), that selects configurations by context length and audit side. Across five benchmarks spanning synthetic data, long-context attacks, and reasoning-model outputs, CAHR-CD and CAHR-AHS improve the six-guardrail average by 22% and 13%, respectively. Code and data are available online.
Chinese Translation
安全护栏作为大型语言模型(LLMs)对有害输入和输出的最后防线,几乎仅在短文本上进行训练和评估。我们提出了LongGuard,一个评估、机制分析和缓解长上下文护栏失败的框架。我们将任务表述为在0.25k-32k长度网格上的安全针在干草堆中(Safety Needle-in-a-Haystack, SafetyNIAH);在15个主流护栏中,不安全召回率平均下降超过50%,而配对的良性填充与针重复设计将失败归因于不安全针的比例稀释,而非绝对长度。对六个护栏进行的三层注意力-逻辑-行为分析定位了机制:不安全针的注意力质量被稀释,不安全与安全的逻辑边际同步压缩,检测决策相应崩溃,这一注意力->逻辑->行为链在剔除长度后仍保持一致。我们进一步隔离出一组稀疏的护栏专用检索头,相较于其基础模型展现出部分特异性。基于分析,我们提出了两种无训练的缓解方法——分块检测(Chunked Detection, CD)和注意力头锐化(Attention-Head Sharpening, AHS),以及一种部署协议——上下文感知超参数路由(Context-Aware Hyperparameter Routing, CAHR),该协议根据上下文长度和审计侧选择配置。在涵盖合成数据、长上下文攻击和推理模型输出的五个基准测试中,CAHR-CD和CAHR-AHS分别提高了六个护栏的平均性能22%和13%。代码和数据在线提供。
cs.AI / 17 / 2608.27638
Generative AI Expands the Intellectual Reach of Course Based Undergraduate Research Experiences (CUREs)
生成性人工智能扩展基于课程的本科研究经验(CUREs)的智力范围
Abstract
Course-based undergraduate research experiences (CUREs) broaden access to authentic scientific inquiry through responsive instructor support as research problems become increasingly complex. Generative artificial intelligence (GenAI) may extend this support by providing individualized assistance that can adapt as student needs change. However, how embedding GenAI within a CURE to provide support across the research process impacts student inquiry, collaboration, and scientific reasoning remains unresolved. Here we use longitudinal qualitative data collected across three semesters of a bioinformatics and genomics CURE to show that GenAI expanded the intellectual reach of the research experience in three distinct ways. First, personalized, on-demand scaffolding allowed students to move beyond the boundaries of instructor expertise and transform their own interests into researchable inquiry, with all teams developing distinct self-directed projects rather than selecting instructor-provided topics. Second, GenAI became part of the distributed cognitive system of research teams, helping novice researchers communicate and coordinate across differentiated expertise without eliminating specialization. Third, expanded capability did not replace the need for disciplinary judgment. Students increasingly validated, revised, or rejected AI-generated contributions, such that research independence emerged through retained intellectual responsibility. Together, these findings suggest that GenAI can extend the reach of CUREs by expanding what novice researchers can investigate, how they can collaborate, and the level of responsibility they can assume while preserving human judgment central to authentic scientific inquiry.
Chinese Translation
基于课程的本科研究经验(CUREs)通过响应性的教师支持,拓宽了对真实科学探究的访问,尤其是在研究问题日益复杂的情况下。生成性人工智能(GenAI)可能通过提供个性化的支持来扩展这种帮助,能够随着学生需求的变化而调整。然而,将GenAI嵌入CURE以在研究过程中提供支持对学生的探究、合作和科学推理的影响仍未得到解决。在此,我们使用在三个学期内收集的生物信息学和基因组学CURE的纵向定性数据,展示了GenAI在三个不同方面扩展了研究经验的智力范围。首先,个性化的按需支架使学生能够超越教师专业知识的界限,将自己的兴趣转化为可研究的探究,所有团队都开发了独特的自我导向项目,而不是选择教师提供的主题。其次,GenAI成为研究团队分布式认知系统的一部分,帮助初学者在不同专业知识之间进行沟通和协调,而不消除专业化。第三,能力的扩展并未取代学科判断的必要性。学生们越来越多地验证、修订或拒绝AI生成的贡献,从而使研究独立性通过保留智力责任而得以显现。综合来看,这些发现表明,GenAI可以通过扩展初学者可以研究的内容、他们的合作方式以及他们可以承担的责任水平,从而扩展CURE的影响,同时保留对真实科学探究至关重要的人类判断。
cs.AI / 18 / 2608.27646
If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement at a Trusted Tool Boundary
如果代理是天使,就不需要治理:在可信工具边界的带外政策执行
Abstract
Give an agent a human's credential and it inherits the person's reach without the judgment that limits its use. It can sweep every reachable record into model context, where hidden instructions steer its next call, and every request stays credential-valid while the agent exceeds its job or absorbs a secret. Prompts are a brittle guardrail: one fallible reasoner interprets the task and enforces its limits. We present Out-of-Band Policy Enforcement (OBPE), a trusted boundary outside agent reasoning. It authorizes the typed operation and resource, narrows the query before the backend call, then filters records and fields or masks values in the response. Semantic gating can deny or hold an authorized call on argument values or external state. A data policy owner sets the maximum grant; agent policy can only narrow it. We prove, under stated conditions, that the policy plan is order-independent and agent policy cannot widen the ceiling. Field removal covers one execution; masking and history rules claim less. We release an HTTP proxy prototype simplified from our production system, with conformance tests tying its typed Cedar policy core to the model. Against Jira and ServiceNow mocks, our benchmark compares prompted agents with and without OBPE on four models, including 20 adaptive red-team tasks. A trace failure means protected data entered agent context, an exact value appeared in the answer, or a forbidden effect completed. In 3,621 trials it fell from 57.6% to 0.2%, a cluster-weighted reduction of 41.2 points [95% CI: 27.7, 54.9]; fulfillment fell from 79.1% to 60.9%, while paired safe-useful completion rose 21.8 points [9.5, 35.2]. Some answers reconstructed a value that never entered context or used filtered row counts as an oracle: shaping one execution is not noninterference. Write controls, durable approval, and temporal and aggregate policies lie outside this evaluation.
Chinese Translation
给一个代理一个人的凭证,它继承了该人的权限,但没有限制其使用的判断。它可以将每个可达记录纳入模型上下文,其中隐藏的指令引导其下一个调用,而每个请求在代理超越其工作或吸收秘密时仍保持凭证有效。提示是一种脆弱的护栏:一个易出错的推理者解释任务并执行其限制。我们提出了带外政策执行(Out-of-Band Policy Enforcement, OBPE),这是一个在代理推理之外的可信边界。它授权类型化操作和资源,在后端调用之前缩小查询,然后过滤记录和字段或在响应中掩盖值。语义门控可以根据参数值或外部状态拒绝或保持授权调用。数据政策所有者设置最大授权;代理政策只能缩小这一范围。在规定条件下,我们证明政策计划是无序独立的,代理政策不能扩大上限。字段移除覆盖一次执行;掩盖和历史规则要求更少。我们发布了一个简化自我们生产系统的HTTP代理原型,符合性测试将其类型化的Cedar政策核心与模型相连接。在Jira和ServiceNow的模拟环境中,我们的基准测试比较了使用和不使用OBPE的提示代理在四个模型上的表现,包括20个自适应红队任务。一次跟踪失败意味着受保护的数据进入了代理上下文,确切值出现在答案中,或完成了禁止的效果。在3,621次试验中,失败率从57.6%降至0.2%,集群加权减少了41.2个百分点[95%置信区间:27.7,54.9];完成率从79.1%降至60.9%,而配对的安全-有用完成率上升了21.8个百分点[9.5,35.2]。一些答案重构了从未进入上下文的值或使用过滤的行计数作为神谕:塑造一次执行并不意味着不干扰。书写控制、持久批准以及时间和聚合政策不在此评估范围内。
cs.AI / 19 / 2608.27671
A Framework for Object-Centric Predictive Monitoring of Collaborative Processes
面向对象的协作过程预测监控框架
Abstract
Predictive Process Monitoring (PPM) of collaborative, inter-organizational processes requires reasoning over multiple interdependent entities, including participants, messages, local executions, and the global collaboration case. Existing approaches extend traditional event logs with collaboration attributes but retain a single-case perspective, leaving much of this structure implicit. Object-centric process mining (OCPM) provides an alternative by representing these entities as first-class objects with explicit relations and multiple notions of case. This study connects collaborative PPM and OCPM through three contributions: (i) a formal semantic mapping from extended collaborative event logs to an OCED-conformant object-centric representation, serialized in OCEL 2.0; (ii) a reformulation of collaborative prediction tasks as object-centric prediction tasks; and (iii) a reproducible converter and prediction pipeline implementing the proposed mapping. We evaluate the framework on four public collaborative event logs and a fifth derived from the BPI Challenge 2013 incident-management log by executing the fourteen reformulated tasks using five predictive strategies across tabular, sequential, and graph-native encodings. We further discuss the benefits, limitations, and threats to the approach's validity. The representation makes collaboration structure explicit and makes it natural to state prediction targets based on object relations that fall outside the case-centric taxonomy, at the cost of increased relational complexity and dependence on object-centric tooling.
Chinese Translation
协作的跨组织过程的预测过程监控(PPM)需要对多个相互依赖的实体进行推理,包括参与者、消息、本地执行和全球协作案例。现有的方法通过扩展传统事件日志的协作属性,但仍然保持单一案例的视角,使得许多结构隐含。面向对象的过程挖掘(OCPM)通过将这些实体表示为具有显式关系和多种案例概念的一类对象,提供了一种替代方案。本研究通过三个贡献将协作PPM与OCPM连接起来:(i)从扩展的协作事件日志到符合OCED标准的面向对象表示的正式语义映射,序列化为OCEL 2.0;(ii)将协作预测任务重新表述为面向对象的预测任务;(iii)实现所提出映射的可重现转换器和预测管道。我们在四个公共协作事件日志和一个来自BPI Challenge 2013事件管理日志的派生日志上评估该框架,通过在表格、序列和图形本地编码中使用五种预测策略执行十四个重新表述的任务。我们进一步讨论了该方法的优点、局限性和有效性威胁。该表示使协作结构显式化,并使基于对象关系的预测目标的表述变得自然,这些关系超出了以案例为中心的分类法,但代价是增加了关系复杂性和对面向对象工具的依赖。
cs.AI / 20 / 2608.27675
Agents for Everyone: A Workshop Framework for Building Agentic AI Capabilities in a Distributed Curation Community
面向所有人的智能体:构建分布式策展社区中智能体AI能力的工作坊框架
Abstract
Agentic AI has the potential to accelerate curation of biological databases and knowledge bases. However, uptake has been hindered by a number of challenges and obstacles, including access to agents and appropriate training. Here we describe how we have attempted to address and mitigate these challenges and obstacles through the deployment of a cloud-based agentic environment, and the development of an interactive training workshop for the Gene Ontology Consortium. Our cloud environment for agentic-assisted curation was based on the JupyterHub platform, and utilized Claude Code as a universal harness. This allows curators to interact with an agent session through a terminal running in the browser, and has additional benefits such as centralization of access through a single API gateway, removing the need for participants to manage subscriptions or install software locally. We created four training modules, walking participants through basic agentic tool use first and then working up to agentic biological pathway curation using the existing GO-CAM (GO Causal Activity Model) curation tool. Thirty-seven participants took part in the four-hour workshop. Our key takeaway from this workshop is that building community capability with agentic AI is primarily a problem of access, workflow design, and training. Removing technical barriers, introducing capabilities gradually, grounding exercises in familiar curation tasks, and giving curators direct experience evaluating agent output can provide a practical route toward building shared agentic AI capability in distributed scientific communities.
Chinese Translation
智能体AI有潜力加速生物数据库和知识库的策展工作。然而,智能体AI的应用受到多种挑战和障碍的制约,包括智能体的获取和适当的培训。本文介绍了我们如何通过部署基于云的智能体环境以及为基因本体联盟(Gene Ontology Consortium)开发互动式培训工作坊,来应对和缓解这些挑战和障碍。我们的智能体辅助策展云环境基于JupyterHub平台,并采用Claude Code作为通用接口。这使得策展人员能够通过浏览器中的终端与智能体会话交互,同时具备通过单一API网关集中访问的优势,免去了参与者管理订阅或本地安装软件的需求。我们设计了四个培训模块,首先引导参与者掌握基本的智能体工具使用,随后逐步深入到使用现有的GO-CAM(GO因果活动模型)策展工具进行智能体生物通路策展。共有37名参与者参加了为期四小时的工作坊。我们的主要结论是,构建社区的智能体AI能力主要依赖于访问权限、工作流程设计和培训。消除技术障碍、逐步引入功能、将练习基于熟悉的策展任务,并让策展人员直接体验评估智能体输出,为在分布式科学社区中构建共享的智能体AI能力提供了切实可行的路径。
cs.AI / 21 / 2608.27716
PCFBench: A Diagnostic Benchmark for Product Carbon Footprint Estimation
PCFBench:产品碳足迹估算的诊断基准
Abstract
AI systems are being deployed on high-stakes, domain-specific workflows that demand correctness not just in the final output, but at every intermediate step. One such workflow is estimating a product carbon footprint (PCF), the greenhouse-gas emissions attributable to a physical product. AI agents are increasingly being used to generate PCFs, but existing evaluations score either total emissions (hiding error sources and cancelling mistakes) or sub-tasks in isolation (missing compositional interactions). We introduce PCFBench, the first benchmark to carve PCF modeling into independently-evaluable tasks that require decomposition, retrieval, ontology matching, and numerical extraction. It comprises 614 expert-labelled items across six tasks. Together they probe reasoning under under-specification, conflicting context, and numerical constraints. Across eight frontier LLMs from four providers, no single model dominates. Although the strongest models estimate total product emissions within 2 times of declared totals on 77% of products, this rate drops to 37-58% when the PCF is generated step by step, with only 45-75% obeying mass conservation. These failures undermine the transparency practitioners need to compare products and drive decarbonization. We release the dataset and evaluation harness to support targeted progress.
Chinese Translation
人工智能系统正在被应用于高风险的特定领域工作流程,这些工作流程不仅要求最终输出的正确性,还要求每一个中间步骤的准确性。其中一个这样的工作流程是估算产品碳足迹(PCF),即归因于某一物理产品的温室气体排放。人工智能代理越来越多地被用于生成PCF,但现有评估要么仅评分总排放量(掩盖错误来源并抵消失误),要么孤立地评估子任务(忽视组合交互)。我们推出了PCFBench,这是第一个将PCF建模划分为可独立评估的任务的基准,这些任务需要进行分解、检索、本体匹配和数值提取。该基准包含614个专家标注的项目,涵盖六个任务。这些任务共同探讨了在规范不足、冲突上下文和数值约束下的推理。在来自四个提供商的八个前沿大型语言模型(LLM)中,没有单一模型占据主导地位。尽管最强模型在77%的产品中能够将总产品排放量估算在声明总量的两倍以内,但当PCF逐步生成时,这一比例下降至37-58%,且仅有45-75%的模型遵循质量守恒。这些失败削弱了从业者在比较产品和推动脱碳方面所需的透明度。我们发布了数据集和评估工具,以支持有针对性的进展。
cs.AI / 22 / 2608.27727
Probing Perceptual Priors of MLLMs via Gibbs Sampling with Interpretable Generative Controls
通过可解释的生成控制与吉布斯采样探究多语言大模型的感知先验
Abstract
A model's behavior on a task is jointly determined by the input it receives and the prior it brings in, i.e. the distribution over stimuli it implicitly expects. Interpretability research has traditionally studied models by holding inputs fixed and examining model responses either mechanistically, probing how internal structure represents inputs, or behaviorally, measuring how variation in inputs leads to variation in outputs. Neither reconstructs the prior distribution itself, since internal structure shows what a model can represent, not what it expects, and any fixed stimulus set leaves most of the possible input space unseen. In particular, such an input space in real-world settings, such as images seen by VLMs, is extremely high-dimensional and diverse. These priors thus remain a poorly understood component of models that nonetheless influence real-world behavior. We propose a method to sample from models' perceptual prior distributions directly, by steering a generative model to produce stimuli along controllable axes and running Gibbs sampling over that space with the model under study as the judge. We apply this to a variety of categories and target variables (such as trustworthiness in faces and cheapness in art images) and recover both canonical biases and surprising novel priors invisible to direct prompting, warranting further investigation of their downstream effects.
Chinese Translation
模型在任务上的行为是由其接收的输入和其带入的先验共同决定的,即它隐含期待的刺激分布。解释性研究传统上通过固定输入来研究模型,或者从机械的角度探讨内部结构如何表示输入,或者从行为的角度测量输入的变化如何导致输出的变化。然而,这两者都无法重建先验分布本身,因为内部结构展示了模型能够表示什么,而不是它期待什么,任何固定的刺激集都使得大部分可能的输入空间未被观察到。特别是在现实世界的设置中,例如视觉语言模型(VLMs)所见的图像,这样的输入空间是极高维且多样的。因此,这些先验仍然是模型中一个理解较差的组成部分,但却影响着现实世界的行为。我们提出了一种直接从模型的感知先验分布中采样的方法,通过引导生成模型沿可控轴产生刺激,并在该空间上运行吉布斯采样,以研究中的模型作为评判者。我们将此方法应用于多种类别和目标变量(例如面孔的可信度和艺术图像的廉价性),并恢复了经典偏见和一些通过直接提示无法察觉的令人惊讶的新先验,值得进一步研究它们的下游影响。
cs.AI / 23 / 2608.27768
Why Didn't It Check? Unsupported Final Claims and Their Repair in Two Tool-Equipped Language Models
为什么没有检查?不支持的最终声明及其在两种工具装备语言模型中的修复
Abstract
A language model with access to tools can commit to a final claim unsupported by the evidence it has seen, even when a single available tool call would resolve the uncertainty and its instructions explicitly forbid assumptions and guesses. We separate this failure into two precisely defined quantities: occurrence, how often the model makes an unsupported claim on its own, measured from the visible evidence and final claim without using the hidden correct answer; and conditional repair, how often those same naturally occurring unsupported claims are repaired when the missing evidence is supplied. On one fixed Qwen3-32B setup, 33 of 512 first responses to 256 new prompt templates ended with an unsupported established claim. We replayed each case from an exact copy of the state in which the claim occurred; within each matched replay, the alternative tool responses had the same structure and length and differed only in a one-character response code. Resolving evidence repaired 33 of 33 claims; a matched response carrying no useful information repaired 0 of 33. When the evidence supported the original answer, the model preserved 33 of 33, with no observed harm. In a separate experiment, on 64 cases where evidence was needed, an automatic checking rule added 21 evidence calls, corrected all 10 wrong unsupported claims, preserved the 11 that were correct by accident, and never changed a correct answer into a wrong one. On a fixed Gemma 4 setup using the same sampling settings, the model called the tool in all 512 first responses and never made an unsupported final claim, so conditional repair could not be measured for that setup. These results describe two local fixed model setups on two synthetic task families. They do not show how common this failure is in real-world deployments, nor that it reflects a general mechanism shared across models.
Chinese Translation
一种可以访问工具的语言模型可能会做出一个不被其所见证据支持的最终声明,即使在单个可用工具调用可以解决不确定性且其指令明确禁止假设和猜测的情况下。我们将这种失败分为两个精确定义的量:发生率,即模型在不使用隐藏正确答案的情况下,根据可见证据和最终声明自行做出不支持声明的频率;条件修复,即在缺失证据被提供时,这些自然发生的不支持声明被修复的频率。在一个固定的 Qwen3-32B 设置中,512 个对 256 个新提示模板的首次响应中,有 33 个以不支持的已建立声明结束。我们从发生声明的状态的精确副本中重放每个案例;在每个匹配重放中,替代工具响应具有相同的结构和长度,仅在一个字符的响应代码上有所不同。解决证据修复了 33 个中的 33 个声明;一个不携带有用信息的匹配响应修复了 33 个中的 0 个。当证据支持原始答案时,模型保留了 33 个中的 33 个,没有观察到任何损害。在另一个实验中,在 64 个需要证据的案例中,一个自动检查规则增加了 21 个证据调用,纠正了所有 10 个错误的不支持声明,保留了 11 个偶然正确的声明,并且从未将正确答案更改为错误答案。在一个固定的 Gemma 4 设置中,使用相同的采样设置,模型在所有 512 个首次响应中调用了工具,并且从未做出不支持的最终声明,因此在该设置中无法测量条件修复。这些结果描述了两个固定模型设置在两个合成任务家族中的表现。它们并未显示这种失败在现实世界部署中的普遍性,也未表明它反映了跨模型共享的一般机制。
cs.AI / 24 / 2608.27790
Credo: Reusable Declarative Primitives for Agentic Workflows
Credo:可重用的声明性原语用于自主工作流
Abstract
An LLM application depends on both a model and a harness: the program that determines what each call sees, how many calls to make, and which answers to trust. Coding agents can now discover strong harnesses by searching over candidate programs, but the resulting artifact is an opaque block of imperative code whose logical steps, runtime signals, physical execution decisions, and prompt strategies remain implicit and task-specific, forcing subsequent tasks to start the harness search process from scratch. The potential for reuse, however, is substantial. A searched harness encodes significant knowledge, such as the logical steps that work, the signals that matter, the physical operator decisions that adapt execution, and the prompt strategies that are effective, yet this knowledge is buried in imperative code with no inspectable or reusable structure, nor does it carry any provenance or metadata. Credo addresses this problem by recovering a structured declarative description of a searched harness, tagging each extracted primitive with relevant metadata, and cataloguing all of it with provenance. A compiler can then bind stored primitives to generate harnesses for new tasks without having to start the search over from scratch. This paper provides preliminary results demonstrating the potential of our approach and lays out a related research agenda that the database community is well-positioned to tackle, including cost-based compilation over declarative catalogs and catalog maintenance under model and workload drift.
Chinese Translation
大型语言模型(LLM)应用依赖于模型和工具的结合:程序决定每次调用所见的内容、调用的次数以及信任哪些答案。编码代理现在可以通过搜索候选程序来发现强大的工具,但生成的产物是一个不透明的命令式代码块,其逻辑步骤、运行时信号、物理执行决策和提示策略仍然是隐含的且特定于任务的,这迫使后续任务从头开始进行工具搜索。然而,重用的潜力是巨大的。经过搜索的工具编码了重要的知识,例如有效的逻辑步骤、重要的信号、适应执行的物理操作决策以及有效的提示策略,但这些知识被埋藏在没有可检查或可重用结构的命令式代码中,也没有任何来源或元数据。Credo通过恢复搜索工具的结构化声明性描述来解决这个问题,为每个提取的原语标记相关的元数据,并将所有这些内容进行目录化并附上来源。然后,编译器可以绑定存储的原语,以生成新的任务所需的工具,而无需从头开始进行搜索。本文提供了初步结果,展示了我们方法的潜力,并提出了一个相关的研究议程,数据库社区在此方面具有良好的研究基础,包括基于成本的声明性目录编译和在模型及工作负载漂移下的目录维护。
cs.AI / 25 / 2608.27796
ReToolSQL: Agentic Reinforcement Learning for Robust Text-to-SQL
ReToolSQL:用于稳健文本到SQL的自主强化学习
Abstract
Recent work has shown that reinforcement learning from execution feedback can substantially improve text-to-SQL performance, often enabling smaller models to match or exceed much larger systems. However, most existing approaches treat SQL generation as a single-turn task, limiting the model's ability to recover from errors through iterative refinement. We present ReToolSQL, a two-stage training framework for text-to-SQL that combines (i) a supervised warm-start on rejection-sampled reasoning traces with (ii) agentic reinforcement fine-tuning (RFT) over multi-turn tool-use trajectories. The key insight is that the two stages act on complementary axes, the supervised fine-tuning (SFT) on verified privileged-teacher traces expands the set of solvable questions (raising pass@k coverage on the hardest cases), while RFT converts that expanded capability into higher single-pass accuracy by teaching the model when to verify, what evidence to retrieve, and how to repair faulty SQL from execution feedback. Applied to Gemma 4 instruction-tuned (31B), RFT alone achieves 73.66% execution accuracy (EX) on the BIRD-SQL development benchmark (74.12% EX with self-consistency). Initializing RFT from the SFT checkpoint (SFT$\to$RFT) yields our strongest model at 74.32% EX single-pass and 74.77% EX with self-consistency. At the time of writing, this ranked first on the BIRD single-model development-set leaderboard. The approach uses composite rewards anchored on execution correctness, requires no human annotation beyond the benchmark itself, and operates within a single dense 31B model, showing that a properly designed SFT$\to$RFT pipeline over tool-use trajectories is a practical path toward robust enterprise-grade text-to-SQL.
Chinese Translation
最近的研究表明,从执行反馈中进行强化学习可以显著提高文本到SQL的性能,通常使得较小的模型能够匹配或超越更大的系统。然而,大多数现有的方法将SQL生成视为单轮任务,限制了模型通过迭代优化从错误中恢复的能力。我们提出了ReToolSQL,一个用于文本到SQL的两阶段训练框架,结合了(i)在拒绝采样推理轨迹上的监督热启动和(ii)在多轮工具使用轨迹上的自主强化微调(RFT)。关键的见解是,这两个阶段在互补的轴上发挥作用,监督微调(SFT)在经过验证的特权教师轨迹上扩展了可解问题的集合(提高了在最难案例上的pass@k覆盖率),而RFT则通过教会模型何时进行验证、检索什么证据以及如何从执行反馈中修复错误的SQL,将这种扩展能力转化为更高的单次执行准确率。应用于Gemma 4指令调优(31B),仅RFT在BIRD-SQL开发基准上实现了73.66%的执行准确率(EX)(自一致性下为74.12% EX)。从SFT检查点初始化RFT(SFT$ o$RFT)产生了我们最强的模型,单次执行准确率为74.32% EX,自一致性下为74.77% EX。在撰写本文时,该模型在BIRD单模型开发集排行榜上排名第一。该方法使用基于执行正确性的复合奖励,不需要超出基准本身的人类标注,并在单个密集的31B模型内运行,表明一个合理设计的SFT$ o$RFT管道在工具使用轨迹上是实现稳健企业级文本到SQL的实用路径。
cs.AI / 26 / 2608.27797
CEDAR: Automata as Verifiable Interfaces for Language-Guided Embodied Action
CEDAR:作为可验证接口的自动机用于语言引导的具身行动
Abstract
Natural-language tasking of embodied agents is rarely just goal specification: users also impose constraints that must persist while the world changes. Code-generating LLM agents can produce plausible behaviors for such instructions, but their free-form programs provide no stable object to verify, compose with new constraints, or repair from a failing trace. We present CEDAR, a counterexample-guided framework that grounds instructions as regular languages over environment event traces. CEDAR uses a language model for semantic judgments and execution traces for correction, then represents both skills and specifications as deterministic finite automata. This turns constraints into executable finite-state objects: a learned skill can be intersected with a learned sleep at night or stay in this biome specification, yielding a controller that enforces the learned constraint by construction rather than by repeated prompting. In Minecraft, with the same simulator/API observations available to a program-generating baseline, CEDAR maintains temporal and spatial constraints that the baseline fails to preserve and amortizes reuse of learned skills, reducing cumulative LLM queries. These results suggest that regular languages offer a practical verification layer between natural-language instructions and embodied-agent policies.
Chinese Translation
对具身代理的自然语言任务不仅仅是目标规范:用户还施加必须在世界变化时持续存在的约束。生成代码的语言模型(LLM)代理可以为这些指令产生合理的行为,但其自由形式的程序并未提供可验证、与新约束组合或从失败轨迹中修复的稳定对象。我们提出了CEDAR,一个反例引导的框架,将指令基础化为环境事件轨迹上的正规语言。CEDAR利用语言模型进行语义判断,并利用执行轨迹进行修正,然后将技能和规范都表示为确定性有限自动机。这将约束转化为可执行的有限状态对象:学习到的技能可以与学习到的夜间睡眠或停留在此生物群落的规范交集,从而生成一个控制器,通过构造而非反复提示来强制执行学习到的约束。在Minecraft中,CEDAR在与程序生成基线相同的模拟器/API观察下,保持了基线未能保持的时间和空间约束,并且减少了累积的LLM查询,促进了学习技能的重用。这些结果表明,正规语言为自然语言指令与具身代理策略之间提供了一个实用的验证层。
cs.AI / 27 / 2608.27808
CURA: Certified Runtime Alarms for Computer-Use Agents
CURA:计算机使用代理的认证运行时警报
Abstract
Self-report is the cheapest oversight channel a deployer has, and on capable computer-use agents (CUAs) it fails precisely where oversight matters. On 361 OSWorld tasks our pipeline, a read-only feasibility gate, a planner, and a GUI executor, reaches a mean task score of 82.9, above the 72.4 human reference, yet 64 of its 71 failures (90%) end with a success claim, 61 acknowledging no blocker, and the explicit failure affordance is never used in roughly 9,100 calls. We introduce CURA (Certified Runtime Alarms for Computer-Use Agents), an external monitor that reads only harness-visible telemetry, with no model internals, extra LLM calls, or prompt changes, and turns the running trajectory into a sequential test with certified false-alarm control. At alpha = 0.10 its CUSUM alarm detects 42.3% of failures a median of 31 steps before termination at a realized false-alarm rate of 0.066, and risk is partly resolvable before the first action (gate probe, 0.69 AUROC). Retrospectively the composite reaches 0.828 AUROC (fold-internal floor 0.802), but its margin over a total-token baseline is not significant (Delta = +0.026, p = 0.101); the separation is online, where CURA recalls more at matched certified budgets: 0.41 versus 0.34 at alpha = 0.10, 0.56 versus 0.38 at alpha = 0.20. Alarm-gated mid-execution oversight recovers 23 of 70 failures while spending a frontier overseer on 38, giving a deployable cascade at mean score 86.8 and 84.5% full-solve (305 of 361). The certificate bounds false alarms only. We also report where behavioral monitoring is uninformative.
Chinese Translation
自我报告是部署者拥有的最便宜的监督渠道,但在能够执行计算机使用的代理(CUAs)中,它在监督至关重要的地方失效。在361个OSWorld任务中,我们的管道——一个只读的可行性门、一个规划器和一个图形用户界面执行器,达到了82.9的平均任务得分,超过了72.4的人类参考值,然而71次失败中的64次(90%)以成功声明结束,61次未承认存在阻碍,并且在大约9100次调用中从未使用过显式失败的可能性。我们引入了CURA(计算机使用代理的认证运行时警报),这是一种外部监控器,仅读取可见的遥测数据,不涉及模型内部、额外的LLM调用或提示更改,并将运行轨迹转化为具有认证虚假警报控制的顺序测试。在α = 0.10时,其CUSUM警报在终止前的中位数31步内检测到42.3%的失败,实际虚假警报率为0.066,并且风险在首次行动之前部分可解决(门探测,0.69 AUROC)。回顾性分析显示复合体达到0.828 AUROC(内部折叠底线0.802),但其相对于总标记基线的边际并不显著(Delta = +0.026,p = 0.101);分离是在线的,CURA在匹配的认证预算下回忆更多:在α = 0.10时为0.41对0.34,在α = 0.20时为0.56对0.38。警报门控的中执行监督恢复了70次失败中的23次,同时在38次上花费了前沿监督者,给出了可部署的级联,平均得分为86.8,完全解决率为84.5%(361个任务中的305个)。证书仅界定虚假警报。我们还报告了行为监测无信息量的情况。
cs.AI / 28 / 2608.27818
AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics
AcCoRD:在现实用户偏好动态下评估用户代理协作
Abstract
User preferences in user-agent collaboration are rarely static and fully-specified upfront: preferences are formed, revealed, adjusted, and relaxed during interaction. Existing benchmarks for evaluating user-agent collaboration focus almost exclusively on resolving underspecified preferences, thereby failing to capture the richer dynamics of real-world interaction. We introduce AcCoRD, a user-agent collaboration benchmark requiring agents to handle diverse user preference dynamics in two domains: online shopping and travel planning. We evaluate five frontier LLMs under two prompting strategies: vanilla ReAct, and an uncertainty-guided variant that prompts models to identify and resolve ambiguity about user preferences. Our results reveal that frontier models can handle underspecification but struggle to satisfy preferences that emerge or evolve mid-interaction and require more sophisticated uncertainty modeling. Further, prompting alone fails to elicit the required uncertainty recognition. We release AcCoRD as a resource for developing agents that can navigate the full complexity of real-world user preferences.
Chinese Translation
用户代理协作中的用户偏好通常并非静态且完全预先指定:偏好在交互过程中形成、揭示、调整和放宽。现有的用户代理协作评估基准几乎完全专注于解决不明确的偏好,因此未能捕捉到现实世界交互的丰富动态。我们引入了 AcCoRD,这是一个用户代理协作基准,要求代理在两个领域(在线购物和旅行规划)中处理多样化的用户偏好动态。我们在两种提示策略下评估了五个前沿的大型语言模型(LLMs):普通的 ReAct 和一种不确定性引导变体,该变体提示模型识别和解决关于用户偏好的模糊性。我们的结果表明,前沿模型能够处理不明确性,但在满足交互过程中出现或演变的偏好时遇到困难,并且需要更复杂的不确定性建模。此外,仅依靠提示无法引发所需的不确定性识别。我们发布 AcCoRD 作为一个资源,以帮助开发能够应对现实用户偏好复杂性的代理。
cs.AI / 29 / 2608.27824
Evidential-Based Higher-Order Set Argumentation Framework
基于证据的高阶集合论证框架
Abstract
Evidential argumentation extends Dung's abstract argumentation by requiring arguments and interactions to be backed by chains of evidence rooted in prima-facie elements. However, existing formalisms lack a unified treatment of evidential support, higher-order relations (attacks and supports targeting arbitrary elements), and collective interactions (sources as sets). In this paper, we introduce the Evidential-Based Higher-Order Set Argumentation Framework (EHSAF), which conservatively generalises several existing frameworks within a single expressive setting. We develop two complete semantics for EHSAFs: an \emph{adjacent complete labelling semantics} that admits multiple truth values (true, false, undecided) for arguments in support cycles, reflecting an open epistemic attitude toward future evidence; and an \emph{extension-based complete semantics} that follows a strict evidentialist stance, accepting only arguments with well-founded support chains. We show that these two semantics diverge in the presence of support cycles, and prove their equivalence under support-acyclicity. To enable computational reasoning, we provide a normal propositional encoding of EHSAFs and prove that, in three-valued {\L}ukasiewicz logic, its models correspond precisely to the adjacent complete labellings. We further extend this encoding to continuous fuzzy logics (G{\"o}del, Product, and {\L}ukasiewicz), defining a continuous fuzzy normal encoded semantics. We establish that this fuzzy semantics satisfies key properties---continuity, monotonicity, boundary conditions, and solution existence---and that its ternarisation recovers the adjacent complete labellings under natural t-norm conditions. Our framework thus unifies expressive argumentation with principled three-valued and fuzzy semantics, bridging the gap between qualitative and quantitative reasoning about evidence.
Chinese Translation
基于证据的论证扩展了Dung的抽象论证,通过要求论证和交互必须由基于表面元素的证据链支持。然而,现有的形式化方法缺乏对证据支持、高阶关系(针对任意元素的攻击和支持)以及集体交互(作为集合的来源)的统一处理。本文介绍了基于证据的高阶集合论证框架(EHSAF),在一个单一的表达环境中保守地概括了多个现有框架。我们为EHSAF开发了两种完整的语义:一种是 extit{相邻完整标记语义},允许支持循环中的论证具有多种真值(真、假、未决定),反映出对未来证据的开放认识态度;另一种是 extit{基于扩展的完整语义},遵循严格的证据主义立场,仅接受具有良好支持链的论证。我们展示了在存在支持循环的情况下这两种语义的差异,并证明了它们在支持无环性下的等价性。为了实现计算推理,我们提供了EHSAF的正常命题编码,并证明在三值{ ext{Ł}}ukasiewicz逻辑中,其模型与相邻完整标记精确对应。我们进一步将这一编码扩展到连续模糊逻辑(G{"o}del、乘积和{ ext{Ł}}ukasiewicz),定义了连续模糊正常编码语义。我们确立了这一模糊语义满足关键属性——连续性、单调性、边界条件和解的存在——并且其三值化在自然t-范数条件下恢复相邻完整标记。因此,我们的框架将表达性论证与原则性的三值和模糊语义统一起来,弥合了关于证据的定性与定量推理之间的鸿沟。
cs.AI / 30 / 2608.27831
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
RealSWE:在现实用户请求下对编码代理的组合评估
Abstract
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues--long, structured, and information-rich. Real user requests, however, are typically far shorter and less structured. To characterize this gap, we define a six-category information taxonomy and four dimensions of linguistic style, and apply them to real user prompts from SWE-chat and problem statements from SWE-bench Verified and Pro. We find that requests carrying only a problem statement, alone or with limited additional context, account for 88% of real prompts but just 7% of benchmark problems. Furthermore, 87% of real prompts are casually written whereas 94% of benchmark problems are formal. Guided by these observations, we introduce sys, 381 multi-variant task families derived from SWE-bench Verified and Pro. Variants within each family share the same underlying task and gold patch while differing only in information composition and linguistic style. Evaluating seven contemporary LLMs with sys, we find that i) realistic inputs reduce resolution rates by 6.4 pp on average and can change model rankings. Controlled analysis further shows that ii) including Desired Behavior and Motivation significantly affects performance, whereas Environment Information and Reproduction Steps merely add tokens without measurable benefit; iii) linguistic style has only small, model-dependent effects. These findings provide actionable guidance for users and agents: explicitly stating the desired behavior and motivation--which most real prompts omit--substantially improves the LLM's software engineering performance.
Chinese Translation
编码代理现在通常在SWE-bench系列基准上进行评估,这些基准的任务基于策划的GitHub问题构建——这些问题通常较长、结构化且信息丰富。然而,真实用户请求通常要短得多且结构较少。为了表征这一差距,我们定义了六类信息分类法和四个语言风格维度,并将其应用于来自SWE-chat的真实用户提示以及SWE-bench Verified和Pro的题目声明。我们发现,仅包含问题陈述的请求,无论是单独存在还是带有有限的附加上下文,构成了88%的真实提示,但仅占基准问题的7%。此外,87%的真实提示是随意书写的,而94%的基准问题则是正式的。基于这些观察,我们引入了sys,这是从SWE-bench Verified和Pro派生出的381个多变体任务家族。每个家族内的变体共享相同的基础任务和金标准补丁,但在信息构成和语言风格上有所不同。通过sys评估七个当代大型语言模型(LLMs),我们发现:i) 现实输入平均降低了解决率6.4个百分点,并可能改变模型排名。控制分析进一步表明:ii) 包含期望行为和动机显著影响性能,而环境信息和再现步骤仅增加了标记而没有可测量的好处;iii) 语言风格仅对模型产生小的、依赖于模型的影响。这些发现为用户和代理提供了可操作的指导:明确说明期望的行为和动机——这是大多数真实提示所忽略的——显著提高了LLM在软件工程方面的表现。
cs.AI / 31 / 2608.27839
KLOD: Locality-Preserving Knowledge Editing via Non-Target Distribution Preservation
KLOD:通过非目标分布保持实现局部性保护的知识编辑
Abstract
Fine-tuning-based knowledge editing is simple and architecture-agnostic, but standard cross-entropy increases the edited target probability without explicitly constraining changes in the non-target output distribution. In sequential editing, such unconstrained redistribution can accumulate as distributional drift and contribute to locality degradation. We propose KLOD, a bounded and distribution-preserving objective for fine-tuning-based knowledge editing that separates the intended target update from distributions that should remain stable. KLOD stops target amplification once a probability threshold is reached, while preserving the target-excluded non-target distribution at target positions and the full next-token distribution at prefix positions. Experiments on CounterFact and ZsRE with Llama3-8B-Instruct and Qwen2.5-7B-Instruct show that KLOD substantially mitigates locality degradation while maintaining high edit reliability. The target probability threshold further provides a controllable Generalization--Locality trade-off. Ablation, multi-seed, and distributional KL analyses support the interpretation that KLOD's locality gains are associated with preserving output distributions rather than simply weakening the edit. Code is available on GitHub https://github.com/Hostoday/KLOD .
Chinese Translation
基于微调的知识编辑简单且与架构无关,但标准的交叉熵增加了编辑目标的概率,而没有明确约束非目标输出分布的变化。在顺序编辑中,这种不受约束的重新分配可能会累积为分布漂移,从而导致局部性退化。我们提出了KLOD,一种有界且保持分布的目标,用于基于微调的知识编辑,它将预期的目标更新与应保持稳定的分布分开。KLOD在达到概率阈值后停止目标放大,同时在目标位置保持目标排除的非目标分布,并在前缀位置保持完整的下一个标记分布。在CounterFact和ZsRE上使用Llama3-8B-Instruct和Qwen2.5-7B-Instruct的实验表明,KLOD显著减轻了局部性退化,同时保持了高编辑可靠性。目标概率阈值进一步提供了可控的泛化-局部性权衡。消融、多种种子和分布KL分析支持了KLOD的局部性增益与保持输出分布相关的解释,而不仅仅是削弱编辑。代码可在GitHub上获取:https://github.com/Hostoday/KLOD 。
cs.AI / 32 / 2608.27840
An Empirical Evaluation of Cross-City POI Recommendation on a Large-Scale Benchmark
大规模基准下跨城市兴趣点推荐的实证评估
Abstract
Cross-city point-of-interest (POI) recommendation is crucial for navigating unfamiliar urban environments, yet its progress has historically been constrained by data limitations. Using the recently proposed large-scale benchmark Trip World, we empirically re-examine whether conclusions drawn on small prior benchmarks still hold under worldwide coverage, low home-destination region overlap, and large, semantically rich POI inventories. Our evaluation surfaces three bottlenecks of representative state-of-the-art methods: (1) hometown-aware models appear to rely more on destination-region priors than on user-specific preference transfer; (2) their accuracy-efficiency trade-off degrades at this scale, where the simplest model is among the strongest; and (3) existing mechanisms for integrating semantic metadata yield little benefit. We further include a diagnostic pilot on agentic methods adapted from next-POI recommendation, finding that naive adaptation trails a simple popularity prior even though the relevant semantic signal is present in the data. These results highlight the need for task-specific designs that support cross-city preference transfer, semantic grounding, and scalable reasoning over unseen destination inventories.
Chinese Translation
跨城市兴趣点(POI)推荐对于在不熟悉的城市环境中导航至关重要,但其进展历来受到数据限制的制约。利用最近提出的大规模基准Trip World,我们实证性地重新审视在全球覆盖、低家乡-目的地区域重叠以及大型、语义丰富的POI库存下,先前小型基准得出的结论是否仍然成立。我们的评估揭示了当前代表性最先进方法的三个瓶颈:(1)关注家乡的模型似乎更依赖于目的地区域的先验信息,而非用户特定的偏好迁移;(2)在这一规模下,它们的准确性与效率的权衡下降,最简单的模型却是最强的之一;(3)现有的语义元数据整合机制几乎没有带来好处。我们进一步包含了一项针对从下一兴趣点推荐中改编的代理方法的诊断试点,发现尽管数据中存在相关的语义信号,简单的适应仍然落后于一个简单的受欢迎程度先验。这些结果突显了需要针对特定任务的设计,以支持跨城市的偏好迁移、语义基础以及对未见目的地库存的可扩展推理。
cs.AI / 33 / 2608.27847
From Uncertainty to Clinical Risk: Severity-Aware Conformal Planning for Interactive Medical Diagnosis
从不确定性到临床风险:面向严重性意识的交互式医学诊断规划
Abstract
Interactive medical diagnosis dynamically acquires patient information through multiple rounds of questioning, supporting accurate, efficient, and safe clinical decisions under incomplete evidence. Existing methods commonly guide information acquisition with predictive uncertainty or label ambiguity, but overlook the asymmetric clinical risk of missing severe diseases and lack unified long-horizon planning over whether to continue asking questions or commit to a diagnosis. To address these limitations, we propose Severity-Aware Conformal Clinical Planning, which formulates interactive diagnosis as a risk-sensitive sequential decision problem. The framework maintains complementary diagnostic, safety, and masked-evidence beliefs; calibrates turn-specific diagnostic prediction sets and severity-weighted differential-diagnosis risk on held-out diagnostic trajectories; and introduces the calibrated clinical risk into Monte Carlo Tree Search to jointly evaluate long-horizon Ask and Commit trajectories. Experiments on DDXPlus and MediQ show that our method achieves more accurate diagnoses with fewer questions across multiple large language models, while improving differential-diagnosis quality and reducing high-risk errors in severe cases. These findings validate the value of using clinical risk, rather than predictive uncertainty alone, as a planning signal and demonstrate the effectiveness of the proposed framework for information acquisition and risk-aware diagnostic decision making. They also motivate future work on clinical-risk-oriented interactive diagnosis and information-acquisition methods.
Chinese Translation
交互式医学诊断通过多轮提问动态获取患者信息,支持在证据不完整的情况下做出准确、高效和安全的临床决策。现有方法通常通过预测不确定性或标签模糊性来指导信息获取,但忽视了漏诊严重疾病的非对称临床风险,并缺乏对是否继续提问或作出诊断的统一长远规划。为了解决这些局限性,我们提出了面向严重性意识的符合性临床规划,将交互式诊断形式化为一个风险敏感的序贯决策问题。该框架维持互补的诊断、安全和掩蔽证据信念;校准特定回合的诊断预测集和基于严重性的差异诊断风险,并在保留的诊断轨迹上进行评估;同时将校准后的临床风险引入蒙特卡罗树搜索,以共同评估长远的提问和作出决策轨迹。在DDXPlus和MediQ上的实验表明,我们的方法在多个大型语言模型中以更少的提问实现了更准确的诊断,同时提高了差异诊断质量并减少了严重病例中的高风险错误。这些发现验证了将临床风险而非单纯的预测不确定性作为规划信号的价值,并展示了所提框架在信息获取和风险意识诊断决策中的有效性。这也激励了未来在临床风险导向的交互式诊断和信息获取方法上的研究。
cs.AI / 34 / 2608.27857
SpikeOPD: Stable On-Policy Distillation for Autoregressive Spiking Language Models
SpikeOPD:自回归脉冲语言模型的稳定在线蒸馏
Abstract
Spiking neural networks (SNNs) offer a path to energy-efficient language modeling through sparse encoding and event-driven computation, but training capable spiking language models from scratch remains difficult. A practical alternative is ANN-to-SNN migration through knowledge distillation (KD), where a pretrained artificial neural network (ANN) teacher supervises an SNN student. Existing migration approaches distill on fixed corpus prefixes, whereas autoregressive inference conditions on self-generated prefixes, creating prefix-source mismatch. It manifests as output-policy mismatch with the ANN teacher and internal spiking-dynamics drift between self-generated and matched corpus prefixes. On-policy distillation (OPD) offers a natural way to mitigate both manifestations by continuing teacher supervision on self-generated prefixes. We evaluate a teacher-only full-KL variant, Vanilla OPD, via a controlled stress test and observe it may suffer from delayed rollout-feedback collapse. This result shows that on-policy coverage alone does not ensure stable adaptation. Motivated by these findings, we propose SpikeOPD, a stable on-policy distillation framework for autoregressive SNNs that learns from self-generated prefixes while maintaining rollout stability. It applies full-KL teacher correction to reduce output-policy mismatch, while matched-prefix policy anchoring constrains policy departure from the frozen reference SNN on the same prefixes. Layerwise spike regularization further limits firing-rate deviations during on-policy adaptation. Across three model scales, SpikeOPD improves average accuracy over the corresponding KD SNNs by 0.8, 1.7, and 2.9 points at 0.125B, 0.35B, and 1.3B, respectively, while preserving their sparse-compute profiles.
Chinese Translation
脉冲神经网络(SNNs)通过稀疏编码和事件驱动计算为能效语言建模提供了一条途径,但从头训练出能够的脉冲语言模型仍然困难。一个实用的替代方案是通过知识蒸馏(KD)进行人工神经网络(ANN)到脉冲神经网络(SNN)的迁移,其中一个预训练的人工神经网络教师监督一个脉冲神经网络学生。现有的迁移方法在固定的语料前缀上进行蒸馏,而自回归推理则依赖于自生成的前缀,导致前缀源的不匹配。这表现为与人工神经网络教师的输出策略不匹配,以及自生成和匹配语料前缀之间的内部脉冲动态漂移。在线蒸馏(OPD)提供了一种自然的方法来缓解这两种表现,通过在自生成前缀上继续教师监督。我们通过控制压力测试评估了仅教师的全KL变体Vanilla OPD,并观察到它可能遭受延迟回滚反馈崩溃。这一结果表明,仅依靠在线覆盖并不能确保稳定的适应。基于这些发现,我们提出了SpikeOPD,一个稳定的在线蒸馏框架,针对自回归脉冲神经网络,在保持回滚稳定性的同时从自生成前缀中学习。它应用全KL教师校正以减少输出策略不匹配,同时匹配前缀策略锚定限制了策略在相同前缀上与冻结参考脉冲神经网络的偏离。逐层脉冲正则化进一步限制了在线适应过程中的发火率偏差。在三个模型规模上,SpikeOPD在0.125B、0.35B和1.3B的情况下,分别提高了相应KD脉冲神经网络的平均准确率0.8、1.7和2.9个点,同时保持了它们的稀疏计算特征。
cs.AI / 35 / 2608.27867
CoRe-MoE: Compact Reusable MoE for Continual Multimodal Instruction Tuning
CoRe-MoE:用于持续多模态指令调优的紧凑可重用 MoE
Abstract
Continual multimodal instruction tuning requires multimodal large language models to acquire new task abilities sequentially while preserving previously learned knowledge. LoRA-MoE provides a promising solution by introducing expert-based capacity, but repeatedly learning and maintaining full LoRA experts leads to substantial parameter overhead. This raises a natural question: is full expert expansion necessary for every new task? To answer it, we analyze the SVD of task-specific LoRA updates and observe substantial overlap in their input- and output-side LoRA direction subspaces, with task-specific adaptation largely captured by lightweight coordinates over these subspaces. Motivated by this observation, we propose CoRe-MoE, a Compact Reusable MoE framework for parameter-efficient continual multimodal instruction tuning. CoRe-MoE extracts reusable input- and output-side direction bases from an initial expert bank, and for subsequent tasks trains only compact coordinate experts together with task-specific low-rank routers. Experiments on two representative MLLMs show that CoRe-MoE improves final average performance over the strongest competing baseline by up to 5.90 points, while using less than 1% of the trainable parameters required by sequential LoRA for later tasks. The code is publicly available at https://github.com/runzezz/CoRe-MoE.
Chinese Translation
持续的多模态指令调优要求多模态大型语言模型能够顺序地获取新任务能力,同时保留先前学习的知识。LoRA-MoE通过引入基于专家的能力提供了一个有前景的解决方案,但反复学习和维护完整的LoRA专家会导致显著的参数开销。这引发了一个自然的问题:对于每个新任务,是否有必要进行完整的专家扩展?为了解答这一问题,我们分析了任务特定的LoRA更新的奇异值分解(SVD),观察到其输入和输出侧的LoRA方向子空间存在显著重叠,任务特定的适应性主要通过这些子空间上的轻量级坐标来捕获。基于这一观察,我们提出了CoRe-MoE,一个用于参数高效的持续多模态指令调优的紧凑可重用MoE框架。CoRe-MoE从初始专家库中提取可重用的输入和输出侧方向基,并在后续任务中仅训练紧凑的坐标专家以及任务特定的低秩路由器。在两个代表性的多模态大型语言模型(MLLMs)上的实验表明,CoRe-MoE在最终平均性能上比最强竞争基线提高了多达5.90分,同时使用的可训练参数少于顺序LoRA在后续任务中所需的1%。代码已公开发布在 https://github.com/runzezz/CoRe-MoE。
cs.AI / 36 / 2608.27869
See, Hypothesize, Validate: Multimodal Agentic Framework for Discovering Governing PDEs
观察、假设、验证:用于发现控制偏微分方程的多模态自主框架
Abstract
Discovering governing partial differential equations (PDEs) from observational data remains a core challenge across the sciences. Existing sparse-regression, symbolic-regression, and LLM-based approaches can be constrained by predefined libraries, noise sensitivity, hallucination, or limited iterative refinement. We introduce \textbf{MAGE} (\textbf{M}ultimodal \textbf{A}gentic \textbf{G}overning \textbf{E}quation Discovery), an agentic framework that organizes PDE discovery as a \textit{confidence governed hypothesis validation loop} inspired by the scientific cycle of observation, hypothesis, and falsification. Four role-specialized agents collaborate: a \textit{Differential Observer} computing derivatives and diagnostic visualizations; a VLM-powered \textit{Phenomenology Extractor} distilling qualitative cues from multimodal diagnostics; an LLM-driven \textit{Governing Law Synthesizer} proposing candidates without a predefined library; and an \textit{Equation Arbiter} fitting coefficients and assigning confidence scores. Discovery iterates until the top candidate clears a user-specified threshold, providing a structured process with an explicit accept-reject protocol. On the evaluated canonical PDE suite, MAGE obtains \textbf{8/8} exact structural recovery and the lowest coefficient error among the compared methods on \textbf{7/8} systems, with improvements of up to \textbf{4 orders of magnitude} and a geometric-mean improvement of approximately \textbf{3 orders of magnitude}. The pipeline also recovers the expected operators in two complex geometries and, on one laboratory sensor record, selects a cubic restoring-force model with held-out $R^2=0.98538$. These results support further study of structured agentic reasoning for library-free governing-law discovery, while broader generalization remains to be evaluated.
Chinese Translation
从观测数据中发现控制偏微分方程(PDEs)仍然是科学领域中的一个核心挑战。现有的稀疏回归、符号回归和基于大型语言模型(LLM)的方法可能受到预定义库、噪声敏感性、幻觉或有限迭代精炼的限制。我们提出了 extbf{MAGE}( extbf{M}ultimodal extbf{A}gentic extbf{G}overning extbf{E}quation Discovery),这是一个自主框架,将PDE发现组织为一个 extit{受信心驱动的假设验证循环},灵感来自于观察、假设和反驳的科学循环。四个角色专门化的代理协作: extit{微分观察者}计算导数和诊断可视化;一个由视觉语言模型(VLM)驱动的 extit{现象提取器}从多模态诊断中提取定性线索;一个由LLM驱动的 extit{控制法则合成器}在没有预定义库的情况下提出候选者;以及一个 extit{方程仲裁者}拟合系数并分配信心分数。发现过程迭代,直到顶级候选者超过用户指定的阈值,提供了一个具有明确接受-拒绝协议的结构化过程。在评估的经典PDE套件中,MAGE在 extbf{8/8}的情况下实现了精确的结构恢复,并在 extbf{7/8}系统中获得了最低的系数误差,相较于其他方法提高了多达 extbf{4个数量级},几何平均改善约为 extbf{3个数量级}。该流程还在两种复杂几何中恢复了预期的算子,并在一个实验室传感器记录中选择了一个立方恢复力模型,保持的$R^2=0.98538$。这些结果支持对无库控制法则发现的结构化自主推理的进一步研究,而更广泛的推广仍需评估。
cs.AI / 37 / 2608.27875
HyQuant: Hybrid-Precision Quantization for LLM Attention
HyQuant:用于大规模语言模型注意力的混合精度量化
Abstract
Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose \textbf{HyQuant}, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of vertical-line tokens and local-window states in high precision. These accuracy-critical regions are selected using lightweight vertical-line-aware attention-pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid-precision quantized attention operator that preserves vertical-line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV-cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: https://github.com/jerrysfls/HyQuant .
Chinese Translation
量化技术已被广泛应用于大规模语言模型(LLM)的训练和推理中,以降低成本并提高效率。然而, extit{注意力}模块的低比特量化在非常低的比特宽度下往往会引入较大的误差,从而导致性能下降。现有的方法主要依赖平滑技术来处理异常值,而我们提出了一种混合量化设计,以更好地平衡准确性和效率。具体而言,我们提出了 extbf{HyQuant},这是一个高效的混合量化框架,专为LLM注意力设计。HyQuant将大多数注意力状态量化为低比特格式,同时保留一小部分垂直线标记和局部窗口状态的高精度。这些对准确性至关重要的区域是通过轻量级的垂直线感知注意力模式信号选择的,从而在有限的开销下减少量化误差。在预填充阶段,HyQuant使用混合精度量化的注意力操作符,保留垂直线标记和局部滑动窗口的全精度,同时对其余上下文进行量化。在解码阶段,HyQuant将相同的原则应用于KV缓存压缩,并将KV反量化与注意力计算融合,以提高内存和硬件效率。在各种任务、模型和数据集上,HyQuant以极其简单的设计保持了几乎无损的准确性,展示了混合量化在LLM注意力中的效率和实际可行性。代码可在以下链接获取:https://github.com/jerrysfls/HyQuant 。
cs.AI / 38 / 2608.27886
Resource Constraints and Performance in Agentic AI Systems
资源约束与自主智能系统的性能
Abstract
Progress toward more autonomous AI increasingly depends on agentic systems that combine a language model with tools, memory, state management, and multi-step execution. These mechanisms shape both task capability and operational burden. We compare OpenClaw and NanoBot as complete agentic systems using a paired primary benchmark and a more detailed instrumented subset of paired prompts. In the primary benchmark, the rate of full task completion was 31% for OpenClaw and 25% for NanoBot, a six-percentage-point difference with a 95% task-bootstrap interval from -3 to 15 percentage points, providing no statistically established full-completion advantage for either system. In the instrumented layer, both systems achieved 26% full completion, while NanoBot reached at least partial completion on 43% of prompts compared with 26% for OpenClaw. OpenClaw took longer on 83% of prompts and had a higher recorded peak-memory value on every prompt, with geometric mean ratios of 2.98 for wall time and 19.44 for peak memory. Among the ten detailed-layer prompts on which at least one system achieved partial or full completion, NanoBot weakly dominated on eight; across all 23 prompts, however, ten of its eighteen dominance cases were cheaper joint failures. Outcome labels differ across the two evidence layers, showing why agent-system evaluation should connect capability and resource measurements to attempt-level execution and scoring provenance. These findings show that progress toward more autonomous AI should be evaluated through verified task completion, observed resource use and records linking each result to the execution that produced it.
Chinese Translation
向更自主的人工智能(AI)系统的进展日益依赖于将语言模型与工具、记忆、状态管理和多步骤执行相结合的自主系统。这些机制影响任务能力和操作负担。我们比较了 OpenClaw 和 NanoBot 作为完整的自主系统,使用了一对主要基准和一组更详细的仪器化配对提示。在主要基准中,OpenClaw 的完整任务完成率为 31%,而 NanoBot 为 25%,两者之间存在六个百分点的差异,95% 的任务自助区间为 -3 到 15 个百分点,未能为任何系统提供统计上显著的完整完成优势。在仪器化层面,两者系统均实现了 26% 的完整完成,而 NanoBot 在 43% 的提示中至少达到了部分完成,相比之下,OpenClaw 仅为 26%。在 83% 的提示中,OpenClaw 的耗时更长,并且在每个提示中记录的峰值内存值均较高,墙面时间的几何均值比为 2.98,峰值内存为 19.44。在十个详细层提示中,至少有一个系统实现了部分或完整的完成,NanoBot 在八个提示上表现出弱优势;然而,在所有 23 个提示中,其十八个优势案例中有十个是更便宜的联合失败。结果标签在两个证据层之间存在差异,显示了为什么自主系统评估应将能力和资源测量与尝试级别的执行和评分来源联系起来。这些发现表明,向更自主的人工智能的进展应通过验证的任务完成、观察到的资源使用以及将每个结果与产生该结果的执行相联系的记录进行评估。
cs.AI / 39 / 2608.27906
Rubric-to-Code Credit Assignment for Reinforcement Learning
基于评分标准的代码信用分配用于强化学习
Abstract
Interactive web application generation requires models to produce usable HTML, CSS, and JavaScript applications from natural language requests. Unlike conventional code generation, application quality depends on multiple user-facing functional requirements, each often tied to localized code regions such as event handlers, state updates, DOM fragments, or CSS selectors. Standard GRPO collapses these structured outcomes into a single sequence-level reward and applies the resulting advantage uniformly to all tokens, weakening credit assignment. We propose \textbf{Rubric-to-Code Credit Assignment} (RCCA), a reinforcement learning framework that converts rubric-level functional feedback into localized optimization signals over generated code. RCCA builds training tasks around explicit functional rubrics, uses a hierarchical reward to separate format, source-code, runtime, and functional failures, and aligns evaluator-generated textual attributions with responsible code spans and generated tokens. The resulting model, \textbf{Ling-RCCA-Flash}, scores 41.25 on MiniAppBench, improving Ling-3.0-Flash by 32.20 points and slightly surpassing Claude Opus 4.5. It also reaches 76.19 on ArtifactsBench, improving the SFT model by 4.48 points and establishing a new top score under the official ArtifactsBench leaderboard setting by surpassing the GPT-5 score by 3.64 points, suggesting transferable implementation-level gains.
Chinese Translation
交互式网页应用生成要求模型能够从自然语言请求中生成可用的 HTML、CSS 和 JavaScript 应用程序。与传统的代码生成不同,应用程序的质量依赖于多个面向用户的功能需求,这些需求通常与局部代码区域(如事件处理程序、状态更新、DOM 片段或 CSS 选择器)相关联。标准的 GRPO 将这些结构化结果压缩为单一的序列级奖励,并对所有标记均匀应用所得到的优势,从而削弱了信用分配。我们提出了 extbf{基于评分标准的代码信用分配}(RCCA),这是一种强化学习框架,将评分标准级别的功能反馈转换为针对生成代码的局部优化信号。RCCA 以明确的功能评分标准为基础构建训练任务,使用层次化奖励来区分格式、源代码、运行时和功能失败,并将评估者生成的文本归因与相关代码片段和生成的标记对齐。最终模型 extbf{Ling-RCCA-Flash} 在 MiniAppBench 上得分 41.25,比 Ling-3.0-Flash 提高了 32.20 分,并稍微超过了 Claude Opus 4.5。在 ArtifactsBench 上得分 76.19,比 SFT 模型提高了 4.48 分,并在官方 ArtifactsBench 排行榜设置下创造了新的最高分,超过了 GPT-5 的得分 3.64 分,表明可转移的实现级别收益。
cs.AI / 40 / 2608.27910
AI Alignment through a Game-theoretic Lens: A Survey
通过博弈论视角的人工智能对齐:一项综述
Abstract
As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them with complex human values has become a central challenge. Existing alignment methods, while effective in improving helpfulness, harmlessness, and controllability, often struggle to capture real-world preferences that are context-dependent, non-transitive, and shaped by dynamic multi-party interactions. This survey reviews AI alignment through a game-theoretic lens. Specifically, it organizes recent progress around key game-theoretic elements and synthesizes the literature along three challenges: preference diversity, alignment priority, and temporal dynamics. This perspective clarifies where current alignment methods genuinely benefit from game-theoretic analysis, where the framework is looser, and what challenges remain in building robust, adaptive, and verifiable AI systems.
Chinese Translation
随着大型语言模型和日益强大的人工智能代理在高风险环境中的部署,使其与复杂的人类价值观保持一致已成为一个核心挑战。现有的对齐方法虽然在提高有用性、无害性和可控性方面有效,但往往难以捕捉到依赖于上下文、非传递性且受到动态多方互动影响的现实偏好。本综述通过博弈论的视角审视人工智能对齐问题。具体而言,它围绕关键的博弈论元素组织了近期的进展,并沿着三个挑战对文献进行了综合:偏好多样性、对齐优先级和时间动态。这一视角澄清了当前对齐方法在何处真正受益于博弈论分析,框架在哪些方面较为松散,以及在构建稳健、自适应和可验证的人工智能系统方面仍面临哪些挑战。
cs.AI / 41 / 2608.27919
From Documents to Reasoning: A Validated Synthetic Data Pipeline and Semantic-Aware Fine-Tuning for Financial Numerical Reasoning
从文档到推理:一个经过验证的合成数据管道和语义感知微调用于金融数值推理
Abstract
Financial question answering (QA) has emerged as a key benchmark for evaluating the performance of Large Language Models (LLMs) on domain-specific tasks involving complex data formats such as tables, charts, and rich textual narratives. While recent advancements have enabled models to reason across modalities and perform multi-step arithmetic operations, limitations remain in performance consistency, and evaluation reliability. In particular, standard evaluation metrics like Exact Match (EM) often fail to account for minor variations such as differences in units or formats, misleading performance assessments. In this work, we propose a comprehensive pipeline for improving financial QA systems through high-quality synthetic data generation and fine-tuning of smaller language models (SLMs) using Quantized Low-Rank Adaptation (QLoRA). Our pipeline includes aggressive data validation for synthetic question answer generation to ensure the relevance and correctness of synthetic question-answer pairs. We introduce a novel evaluation metric that matches answers computed from arithmetic expressions rather than ground-truth answers; providing a more accurate reflection of model reasoning capability. Furthermore, we propose a modified loss function that aligns predicted and reference expressions using semantic similarity, our novel evaluation metric and standard cross-entropy, resulting in improved performance. Experimental results on benchmark datasets, ConvFinQA demonstrate significant gains in QA accuracy after fine-tuning using synthetic dataset and proposed loss function.
Chinese Translation
金融问答(QA)已成为评估大型语言模型(LLMs)在涉及复杂数据格式(如表格、图表和丰富文本叙述)等领域特定任务表现的关键基准。尽管最近的进展使得模型能够跨模态推理并执行多步算术运算,但在性能一致性和评估可靠性方面仍然存在局限性。特别是,像精确匹配(EM)这样的标准评估指标往往未能考虑单位或格式等细微差异,从而导致误导性的性能评估。在本研究中,我们提出了一种综合管道,通过高质量合成数据生成和使用量化低秩适应(QLoRA)对小型语言模型(SLMs)进行微调,以改善金融QA系统。我们的管道包括对合成问题答案生成进行严格的数据验证,以确保合成问题-答案对的相关性和正确性。我们引入了一种新颖的评估指标,该指标匹配从算术表达式计算得出的答案,而不是基准答案;提供了对模型推理能力的更准确反映。此外,我们提出了一种修改后的损失函数,通过语义相似性对预测表达式和参考表达式进行对齐,结合我们的新评估指标和标准交叉熵,从而提高了性能。在基准数据集ConvFinQA上的实验结果表明,使用合成数据集和提出的损失函数进行微调后,问答准确性显著提升。
cs.AI / 42 / 2608.27940
A Deep Learning-Based Stacking Ensemble Framework for Turbofan Engine Remaining Useful Life Prediction
基于深度学习的涡扇发动机剩余使用寿命预测的堆叠集成框架
Abstract
This study proposes a two-level stacking ensemble framework for Remaining Useful Life (RUL) prediction of turbofan engines, evaluated on the NASA C-MAPSS benchmark using the FD001 and FD003 subsets. The framework integrates four heterogeneous deep learning base learners: Long Short-Term Memory (LSTM), Convolutional Neural Network (CNN), CNN-LSTM, and CNN-GRU, whose out-of-fold predictions are combined by an XGBoost meta-learner to capture complex degradation patterns while mitigating individual model biases. Comprehensive experiments demonstrate that the stacking ensemble achieves superior predictive performance, with Root Mean Square Error (RMSE) of 9.989 and 8.613, Mean Absolute Error (MAE) of 7.081 and 5.195, and R-squared values of 0.899 and 0.906 for FD001 and FD003, respectively. Compared to the best-reported baseline (TCAT: RMSE 11.12 and 11.02), the proposed method achieves RMSE reductions of 10.2 percent and 21.8 percent for FD001 and FD003, respectively. Feature correlation analysis, residual diagnostics, and training convergence curves validate the model's robustness. These findings underscore the efficacy of stacking ensemble methods for prognostics and health management in safety-critical aerospace applications.
Chinese Translation
本研究提出了一种两层堆叠集成框架,用于涡扇发动机的剩余使用寿命(RUL)预测,基于NASA C-MAPSS基准数据集中的FD001和FD003子集进行评估。该框架集成了四个异构深度学习基础学习器:长短期记忆网络(LSTM)、卷积神经网络(CNN)、CNN-LSTM和CNN-GRU,其外折预测结果通过XGBoost元学习器进行组合,以捕捉复杂的退化模式,同时减轻个别模型的偏差。全面的实验表明,堆叠集成方法实现了优越的预测性能,FD001和FD003的均方根误差(RMSE)分别为9.989和8.613,平均绝对误差(MAE)分别为7.081和5.195,R平方值分别为0.899和0.906。与最佳报告的基线(TCAT:RMSE 11.12和11.02)相比,所提方法在FD001和FD003上实现了RMSE分别降低10.2%和21.8%。特征相关性分析、残差诊断和训练收敛曲线验证了模型的鲁棒性。这些发现强调了堆叠集成方法在安全关键航空应用中的预测与健康管理中的有效性。
cs.AI / 43 / 2608.27942
CASTANET: Causality-Aware Spatio-Temporal Adversarial Network Using Traffic Incident Effects
CASTANET:基于交通事件影响的因果意识时空对抗网络
Abstract
Predicting non-periodic traffic congestion caused by sudden incidents (e.g., accidents and road damage) is crucial for advanced intelligent transportation systems. However, incident-driven congestion is difficult to forecast because incidents are extremely sparse, occur at specific times and locations, and have heterogeneous impacts depending on the traffic context. While recent deep learning approaches have significantly improved periodic traffic forecasting, their performance on non-periodic congestion remains limited, partly because incident records are not explicitly incorporated and their occurrence is strongly biased in space and time. To address these challenges, we propose CASTANET, which integrates spatio-temporal graph neural networks and causal treatment effect estimation to utilize incident records while mitigating selection bias. Experiments on real-world traffic data and accident records from Tokyo, which we treat as incidents, show that CASTANET reduces RMSE by 4.0% overall compared to the best baseline and by 10.1% on incident-conditioned evaluation, with gains reaching 14.55% under severe congestion.
Chinese Translation
预测由突发事件(例如事故和道路损坏)引起的非周期性交通拥堵对于先进的智能交通系统至关重要。然而,由于事件极为稀疏、发生在特定的时间和地点,并且根据交通背景具有异质性影响,因此事件驱动的拥堵难以预测。尽管最近的深度学习方法显著改善了周期性交通预测,但它们在非周期性拥堵上的表现仍然有限,部分原因是事件记录未被明确纳入,并且其发生在空间和时间上存在强烈的偏差。为了解决这些挑战,我们提出了CASTANET,该方法结合了时空图神经网络和因果处理效应估计,以利用事件记录,同时减轻选择偏差。在对东京的真实交通数据和事故记录进行的实验中,我们将这些记录视为事件,结果表明CASTANET在整体上将均方根误差(RMSE)降低了4.0%,在基于事件的评估中降低了10.1%,在严重拥堵情况下的增益达到14.55%。
cs.AI / 44 / 2608.27945
Cross-Session Decomposition Attacks: Scaling Risk and Intent-Aligned Retrieval Defense
跨会话分解攻击:风险扩展与意图对齐检索防御
Abstract
Scaling laws are usually read as a capability story: lower language-modeling loss yields more useful models. We study a safety consequence of this mechanism in \emph{cross-session decomposition attacks}, where benign-looking subqueries are asked across independent interactions and later recomposed toward a forbidden objective. We formalize this setting as \emph{compositional safety risk} and prove a conditional risk-transfer bound: when the reference environment already contains dispersed evidence for a risky reconstruction, the gap between deployed composed risk and reference composed risk is controlled by the model's excess loss on allowed subqueries. Synthetic withholding experiments show that wider transformers assign lower loss to held-out instructions that never appear verbatim in training but are recoverable from injected supporting facts. A 600-intent pretrained-LLM evaluation shows that larger Qwen3 and Gemma3 family members can yield greater harmful-capability uplift under a fixed decomposition-composition pipeline. As a defense, IntentAlign-MiniLM, our 22M-parameter intent-aligned retriever, outperforms much larger embedding models on held-out intent retrieval and yields the best learned-retriever harmful recall across tested guardrails. Code is available in \href{https://github.com/liaodisen/Cross-Session-Decomposition-Attacks}{our GitHub repository}.
Chinese Translation
扩展法则通常被解读为一种能力故事:较低的语言建模损失产生更有用的模型。我们研究了这一机制在 extit{跨会话分解攻击}中的安全后果,其中看似良性的子查询在独立交互中被提出,并随后重新组合以达到一个禁止的目标。我们将这一设置形式化为 extit{组合安全风险},并证明了一个条件风险转移界限:当参考环境中已经包含分散的证据以支持风险重构时,已部署的组合风险与参考组合风险之间的差距由模型在允许子查询上的过剩损失控制。合成隐蔽实验表明,较宽的变换器对那些在训练中从未逐字出现但可以从注入的支持事实中恢复的保留指令分配了较低的损失。对600个意图的预训练大语言模型的评估显示,在固定的分解-组合管道下,较大的Qwen3和Gemma3家族成员能够在有害能力提升方面产生更大的增益。作为防御,我们的22M参数意图对齐检索器IntentAlign-MiniLM在保留意图检索上优于更大的嵌入模型,并在测试的防护措施中产生了最佳的学习检索器有害召回。代码可在 extit{我们的GitHub仓库}中获取。
cs.AI / 45 / 2608.27953
The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs
‘如果’的错觉:评估大型语言模型中反事实推理的崩溃
Abstract
Counterfactual reasoning requires models to reason beyond the observed world and explain how altered conditions propagate through downstream consequences. Existing benchmarks largely target bounded settings with fixed variables or single gold outcomes, overlooking open-domain scenarios requiring causal-process evaluation. To this end, we present $\textbf{WhatIfBench}$, a diagnostic benchmark for open-domain, open-form, long-horizon counterfactual causal reasoning, containing 220 what-if questions across STEM, HSS, and Hybrid scenarios. To evaluate free-form responses, we further propose $\textbf{PRISM}$, which first converts each natural-language explanation into a Response-Derived Semantic Causal Graph of events, states, and mechanisms. On top of this graph, PRISM then jointly applies a Process Metric assessing graph-level causal validity and a Rubric Metric assessing answer-level explanatory adequacy. Evaluating six frontier LLMs with this framework, we find that WhatIfBench remains far from saturated: even the strongest model reaches only a 64.62% final score. Further analysis reveals persistent causal gaps, premise drift, and topology fragmentation, suggesting that fluent counterfactual narratives often mask fragile causal processes. The benchmark, code, and evaluation scripts are available at $\href{https://github.com/zju-gt/WhatIfBench}{WhatIfBench}$.
Chinese Translation
反事实推理要求模型超越观察到的世界进行推理,并解释如何通过下游后果传播改变的条件。现有基准主要针对具有固定变量或单一黄金结果的有限设置,忽视了需要因果过程评估的开放领域场景。为此,我们提出了$ extbf{WhatIfBench}$,这是一个用于开放领域、开放形式、长时间跨度反事实因果推理的诊断基准,包含220个涵盖STEM(科学、技术、工程和数学)、HSS(人文和社会科学)及混合场景的‘如果’问题。为了评估自由形式的回答,我们进一步提出了$ extbf{PRISM}$,该方法首先将每个自然语言解释转换为事件、状态和机制的响应导向语义因果图。在此图的基础上,PRISM联合应用一个过程指标来评估图级因果有效性,以及一个评分标准来评估答案级解释充分性。通过该框架评估六个前沿大型语言模型,我们发现WhatIfBench仍远未饱和:即使是最强的模型最终得分也仅为64.62%。进一步分析揭示了持续存在的因果缺口、前提漂移和拓扑碎片化,表明流畅的反事实叙述往往掩盖了脆弱的因果过程。基准、代码和评估脚本可在$ exthref{https://github.com/zju-gt/WhatIfBench}{WhatIfBench}$获取。
cs.AI / 46 / 2608.27960
When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation
当教师指导误导时:奖励对齐的在线蒸馏
Abstract
On-policy distillation (OPD) has recently emerged as a popular post-training paradigm for large language models (LLMs), providing an efficient way to transfer the knowledge and capabilities of teacher models into student models. However, teacher guidance on student-generated prefixes is not always reliable. Training should optimize the model to generate responses that are more likely to be correct, or equivalently, to get higher outcome rewards. But during OPD, the teacher model may provide guidance that discourages the student from moving toward correct trajectories or moves the student toward incorrect ones, which is misaligned with outcome reward. Such misaligned guidance is unreliable, as it would mislead the optimization process and ultimately degrade model performance. To mitigate misaligned teacher guidance, we propose Reward-Aligned On-Policy Distillation (RA-OPD). The key insight is to keep only trajectories whose induced updates move the student toward correct trajectories or discourage the student from moving toward incorrect ones. Specifically, for each sampled trajectory, RA-OPD checks whether its trajectory-level distillation return is consistent with its outcome reward and then filters out the misaligned trajectories. RA-OPD selects more reliable trajectories to improve student model performance without requiring additional computational cost. We evaluate RA-OPD on math and code benchmarks using models from the Qwen3 family and the DeepSeek-R1 family. Across seven math benchmarks and three code benchmarks, RA-OPD significantly outperforms standard OPD and other tested OPD variants.
Chinese Translation
在线蒸馏(On-policy distillation, OPD)最近成为大型语言模型(Large Language Models, LLMs)后训练的一个流行范式,提供了一种有效的方法将教师模型的知识和能力转移到学生模型中。然而,教师对学生生成的前缀的指导并不总是可靠的。训练应优化模型以生成更可能正确的响应,或者等价地,获得更高的结果奖励。但在OPD过程中,教师模型可能提供的指导会阻碍学生朝向正确轨迹移动,或使学生朝向错误轨迹移动,这与结果奖励不一致。这种不一致的指导是不可靠的,因为它会误导优化过程,最终降低模型性能。为了减轻不一致的教师指导,我们提出了奖励对齐的在线蒸馏(Reward-Aligned On-Policy Distillation, RA-OPD)。其关键见解是仅保留那些诱导更新使学生朝向正确轨迹移动或阻止学生朝向错误轨迹移动的轨迹。具体而言,对于每个采样的轨迹,RA-OPD检查其轨迹级蒸馏回报是否与其结果奖励一致,然后过滤掉不一致的轨迹。RA-OPD选择更可靠的轨迹,以提高学生模型的性能,而无需额外的计算成本。我们在数学和代码基准上评估RA-OPD,使用来自Qwen3家族和DeepSeek-R1家族的模型。在七个数学基准和三个代码基准上,RA-OPD显著优于标准OPD和其他测试的OPD变体。
cs.AI / 47 / 2608.27963
SABER: Stability-Aware Early Exit for LLM Reasoning via Adversarial Branch Probing
SABER:通过对抗分支探测实现的稳定性意识早期退出用于大型语言模型推理
Abstract
Large Reasoning Models (LRMs) achieve strong reasoning capabilities, yet long-chain reasoning becomes inefficient once the intermediate answer stabilizes across reasoning steps: additional reasoning yields little marginal benefit while incurring substantial inference cost. Existing early-exit methods based on confidence or entropy poorly capture reasoning stability, while consistency-based approaches rely on multi-step trajectory agreement, requiring sequential evaluations that delay exit. To better balance efficiency and reliability, we propose SABER, a training-free framework for stability-aware early exit via adversarial branch probing. SABER constructs simple yet effective semantic perturbations around intermediate reasoning states to form adversarial branches, and applies lightweight probing to estimate their likely final outcomes without full trajectory rollouts. When the probed outcomes remain consistent across branches, SABER exits early; otherwise, it continues reasoning. Experiments across multiple reasoning benchmarks and model architectures show that SABER reduces reasoning token consumption by 30.2\%--39.8\% on average while maintaining competitive accuracy with full-length reasoning.
Chinese Translation
大型推理模型(LRMs)具备强大的推理能力,但一旦中间答案在推理步骤中稳定,长链推理变得低效:额外的推理带来的边际收益很小,同时会产生可观的推理成本。现有基于置信度或熵的早期退出方法难以有效捕捉推理的稳定性,而基于一致性的 approaches 依赖于多步骤轨迹一致性,需进行顺序评估,从而延迟退出。为了更好地平衡效率和可靠性,我们提出了SABER,这是一种通过对抗分支探测实现的稳定性意识早期退出的无训练框架。SABER围绕中间推理状态构建简单而有效的语义扰动,以形成对抗分支,并应用轻量级探测来估计其可能的最终结果,而无需完整的轨迹展开。当探测到的结果在分支间保持一致时,SABER会提前退出;否则,它将继续推理。在多个推理基准和模型架构上的实验表明,SABER在保持与全长推理竞争的准确性的同时,平均减少了30.2%至39.8%的推理令牌消耗。
cs.AI / 48 / 2608.27964
AERA: Adaptive Evidence Residual Allocation for Efficient Test-Time Reasoning
AERA:用于高效测试时推理的自适应证据残差分配
Abstract
Test-time scaling improves language-model reasoning by generating additional candidate solutions, but allocating the same inference budget to every problem is computationally wasteful. Existing adaptive stopping methods commonly rely on confidence, agreement, or answer stability, implicitly assuming that stronger current evidence indicates that further computation is unnecessary. We show that this assumption can fail: checkpoint-level correctness evolves non-monotonically, and observable evidence may strengthen before an answer collapses or weaken before it recovers. Motivated by this mismatch, we introduce Adaptive Evidence Residual Allocation (AERA), a sequential controller that learns whether additional computation is likely to recover a better answer from checkpoint-observable evidence. AERA characterizes cumulative response prefixes using answer-distribution, temporal, re-solving, semantic, and compute features, and repeatedly decides whether to stop or allocate the next response block. Future checkpoint correctness is used only to construct offline supervision and is never available to the controller at inference time. Across GSM8K and GPQA Diamond, AERA identifies question-specific residual opportunities while substantially reducing inference computation. In a frozen-threshold incremental-generation evaluation on 300 untouched GSM8K questions, AERA achieves 92.61% accuracy versus 93.01% with 128 responses while reducing completion tokens by 95.99%. These results suggest that adaptive reasoning should estimate the future value of computation rather than equating present confidence with correctness.
Chinese Translation
测试时缩放通过生成额外的候选解决方案来改善语言模型推理,但将相同的推理预算分配给每个问题在计算上是浪费的。现有的自适应停止方法通常依赖于置信度、一致性或答案稳定性,隐含假设更强的当前证据表明进一步计算是没有必要的。我们展示了这一假设可能失效:检查点级别的正确性是非单调演变的,且可观察的证据可能在答案崩溃之前增强,或在答案恢复之前减弱。基于这种不匹配,我们引入了自适应证据残差分配(Adaptive Evidence Residual Allocation,AERA),这是一种顺序控制器,学习是否有可能通过检查点可观察证据恢复更好的答案。AERA使用答案分布、时间、重新求解、语义和计算特征来表征累积响应前缀,并反复决定是停止还是分配下一个响应块。未来的检查点正确性仅用于构建离线监督,在推理时从不提供给控制器。在GSM8K和GPQA Diamond数据集上,AERA识别特定问题的残差机会,同时显著减少推理计算。在对300个未触及的GSM8K问题进行的冻结阈值增量生成评估中,AERA实现了92.61%的准确率,而128个响应的准确率为93.01%,同时减少了95.99%的完成标记。这些结果表明,自适应推理应该估计计算的未来价值,而不是将当前置信度等同于正确性。
cs.AI / 49 / 2608.27969
openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents
openJiuwen:超越静态代理的长时程编码代理
Abstract
Long-horizon coding agents operate over evolving repository states while increasingly relying on heterogeneous capabilities, delegated agents, and multi-agent coordination. These trends pose two complementary challenges for the agent harness. First, developers need to compose capabilities, reconfigure execution logic, and scale increasingly complex agent systems without repeatedly rebuilding orchestration. Second, complex coding tasks continuously produce new evidence---such as semantic diagnostics, execution outcomes, task progress, and changing context relevance---that should dynamically influence subsequent runtime decisions. We characterize these challenges as Structural Composability and Runtime Adaptivity. We present openJiuwen, an open-source harness designed for both developer composability and adaptive task execution. openJiuwen provides a shared execution substrate and Rail-based capability composition across single agents, delegated sub-agents, and Swarm Flow, enabling developers to construct sophisticated agent harnesses under common execution semantics. It further adapts framework-controlled runtime decisions around a fixed model policy, allowing evolving evidence to dynamically affect context, feedback, and task control toward successful completion. We systematically evaluate openJiuwen on SWE-bench Verified and Terminal-Bench 2.1, where it achieves 82.6% and 87.19%, respectively, exceeding the strongest selected official-leaderboard point estimates by 3.4 and 3.39 percentage points. These results show that openJiuwen achieves strong performance on complex coding tasks while providing a composable and adaptive harness design.
Chinese Translation
长时程编码代理在不断发展的代码库状态下运行,同时越来越依赖异构能力、委派代理和多代理协调。这些趋势为代理工具提出了两个互补的挑战。首先,开发者需要组合能力、重新配置执行逻辑,并在不重复重建编排的情况下扩展日益复杂的代理系统。其次,复杂的编码任务持续产生新的证据——例如语义诊断、执行结果、任务进展和变化的上下文相关性——这些证据应动态影响后续的运行时决策。我们将这些挑战归纳为结构可组合性(Structural Composability)和运行时适应性(Runtime Adaptivity)。我们提出了openJiuwen,一个旨在支持开发者可组合性和自适应任务执行的开源工具。openJiuwen提供了一个共享的执行基础设施和基于Rail的能力组合,适用于单个代理、委派的子代理和Swarm Flow,使开发者能够在共同的执行语义下构建复杂的代理工具。此外,它还围绕固定的模型策略调整框架控制的运行时决策,使不断变化的证据能够动态影响上下文、反馈和任务控制,从而实现成功完成。我们在SWE-bench Verified和Terminal-Bench 2.1上对openJiuwen进行了系统评估,分别达到了82.6%和87.19%的成绩,超出了最强选定官方排行榜点估计值3.4和3.39个百分点。这些结果表明,openJiuwen在复杂编码任务上表现出色,同时提供了可组合和自适应的工具设计。
cs.AI / 50 / 2608.27982
Learning from Hard Prompts: Difficulty-aware Advantage Amplification in Dynamic Sampling
从困难提示中学习:动态采样中的难度感知优势放大
Abstract
Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) is a prominent variant of Group Relative Policy Optimization (GRPO). DAPO introduces several improvements over GRPO. Among these, Dynamic Sampling contributes the most to DAPO's accuracy gains relative to GRPO. To improve accuracy, Dynamic Sampling enhances training stability by eliminating zero policy gradients from zero advantages. Specifically, it avoids such zero gradients by filtering out prompts where sampled responses are either entirely correct or incorrect. However, our theoretical analysis shows that Dynamic Sampling decrease training efficiency as it cannot effectively utilize hard-to-sample correct responses on hard prompts. Formally, it asymmetrically amplifies the advantages of distinct responses to the same prompts. On hard prompts, incorrect responses undergo greater amplification than correct ones. This leads the model to avoid generating the observed incorrect responses rather than capitalizing on the hard-to-sample correct ones on hard prompts, resulting in low training efficiency. To improve training efficiency, we propose Direct Advantage Amplification (DAA), which amplifies the advantages of hard-to-sample correct responses on hard prompts, as obtained by Dynamic Sampling. This ensures that, when Dynamic Sampling is used, these hard-to-sample responses can be effectively capitalized on, implying higher training efficiency. By integrating DAA into DAPO, we obtain Difficulty-aware Advantage Amplification Policy Optimization (DA3PO), which is implemented with fewer than 30 lines of code from DAPO. Experiments show that DA3PO significantly outperforms GRPO and other classical GRPO variants.
Chinese Translation
解耦的剪辑和动态采样策略优化(Dynamic Sampling Policy Optimization, DAPO)是群体相对策略优化(Group Relative Policy Optimization, GRPO)的一个重要变体。DAPO在GRPO的基础上引入了若干改进。其中,动态采样对DAPO相对于GRPO的准确性提升贡献最大。为了提高准确性,动态采样通过消除来自零优势的零策略梯度来增强训练的稳定性。具体而言,它通过过滤掉那些采样响应完全正确或错误的提示来避免这种零梯度。然而,我们的理论分析表明,动态采样降低了训练效率,因为它无法有效利用在困难提示上难以采样的正确响应。从形式上看,它不对称地放大了对同一提示的不同响应的优势。在困难提示上,错误响应的放大程度大于正确响应。这导致模型倾向于避免生成观察到的错误响应,而不是利用在困难提示上难以采样的正确响应,从而导致训练效率低下。为了提高训练效率,我们提出了直接优势放大(Direct Advantage Amplification, DAA),它放大了在困难提示上难以采样的正确响应的优势,这些响应是通过动态采样获得的。这确保了在使用动态采样时,这些难以采样的响应能够被有效利用,从而提高训练效率。通过将DAA集成到DAPO中,我们得到了难度感知优势放大策略优化(Difficulty-aware Advantage Amplification Policy Optimization, DA3PO),其实现代码少于30行。实验表明,DA3PO显著优于GRPO及其他经典的GRPO变体。
cs.AI / 51 / 2608.27984
When Evidence Shapes Collaboration: Knowledge-Conditioned Topology Generation for Multi-Agent Systems
当证据影响协作:面向多智能体系统的知识条件拓扑生成
Abstract
Multi-Agent Systems (MAS) have recently moved from static workflows toward dynamically generated collaboration topologies. However, existing topology generation methods rely primarily on the parametric knowledge of large language models, with external search or retrieval used only as a reactive tool rather than an explicit determinant of collaboration structure. This leads to structure-knowledge misalignment, where systems exhibit redundant interactions or insufficient verification in knowledge-intensive tasks. We propose K-GAT (Knowledge-Guided Agent Topology Generator), a neuro-symbolic framework that formulates collaboration topology design as a knowledge-conditioned structure learning problem, integrating external evidence directly into autoregressive graph generation. Extensive experiments on knowledge-intensive benchmarks demonstrate K-GAT's efficiency and effectiveness: notably on the expert-level GPQA dataset, K-GAT outperforms the LLM-Debate baseline by a substantial margin of +15.7% in accuracy, while consuming less than half the computational tokens.
Chinese Translation
多智能体系统(MAS)最近从静态工作流转向动态生成的协作拓扑。然而,现有的拓扑生成方法主要依赖于大型语言模型的参数知识,外部搜索或检索仅作为一种反应工具,而非协作结构的明确决定因素。这导致了结构与知识之间的不匹配,系统在知识密集型任务中表现出冗余的交互或不足的验证。我们提出了K-GAT(知识引导的智能体拓扑生成器),这是一个神经符号框架,将协作拓扑设计公式化为一个知识条件下的结构学习问题,直接将外部证据整合到自回归图生成中。在知识密集型基准测试上的大量实验表明K-GAT的高效性和有效性:特别是在专家级GPQA数据集上,K-GAT的准确率比LLM-Debate基线高出15.7%,同时计算代币消耗不到一半。
cs.AI / 52 / 2608.27992
GOD: Govern, Observe, and Direct - A Real-Time Control Room for Agent Societies
GOD:治理、观察与引导 - 一种面向代理社会的实时控制室
Abstract
Generative-agent systems are easier to start than to inspect. A run can contain many agents, locations, messages, commands, and model calls, yet the operator often gets either a finished replay or raw logs. That makes it hard to ask why an agent moved, test a small intervention, or package a run for another researcher. GOD is a local-first control room for agent societies. From the same browser workflow, an operator can issue targeted questions or interventions and inspect the resulting replay state. The system combines a setup wizard, Agent Studio, Map Studio, a spatial replay interface, Ask and Intervene commands, and portable experiment, map, and agent packs. Its technical contribution is the command and artifact loop: live controls and replay evidence share the same operator command model, while package contracts separate scenario, map, and profile data from local runtime state. The public release includes hosted Smallville-style and PKU replays, the open-source repository, and downloadable packs. We evaluate this path on 15 completed run slots. Across the 14 intervention runs, 78 of 84 target-agent checks recorded the commanded destination, and 169 of 182 state answers matched a saved location or action string.
Chinese Translation
生成代理系统的启动比检查要容易。一次运行可以包含多个代理、位置、消息、命令和模型调用,但操作员往往只能获得完成的重放或原始日志。这使得很难询问某个代理为何移动、测试小规模干预或将运行打包给其他研究者。GOD是一个以本地为先的代理社会控制室。在同一个浏览器工作流中,操作员可以发出针对性的问题或干预,并检查相应的重放状态。该系统结合了设置向导、代理工作室、地图工作室、空间重放界面、询问和干预命令,以及可移植的实验、地图和代理包。其技术贡献在于命令和工件循环:实时控制和重放证据共享相同的操作员命令模型,而包合同则将场景、地图和配置数据与本地运行时状态分离。公开发布包括托管的Smallville风格和PKU重放、开源代码库以及可下载的包。我们在15个已完成的运行槽上评估了这一路径。在14个干预运行中,84个目标代理检查中有78个记录了命令的目标位置,182个状态回答中有169个与保存的位置或动作字符串匹配。
cs.AI / 53 / 2608.27996
Should I Use This Synthetic Dataset for Training? How to Test with Minimal Real Data
我应该使用这个合成数据集进行训练吗?如何在最少的真实数据下进行测试
Abstract
Digital twins (DTs) and learned world models are increasingly used to generate synthetic data that augment the scarce real datasets available for training artificial intelligence (AI) models in engineering systems. Owing to the inevitable simulation-to-reality (sim-to-real) gap, however, augmentation may fail to improve the performance of the trained model on the real data distribution. This paper addresses the resulting decision problem: Given a real dataset, a candidate synthetic dataset, and a fixed learning algorithm, decide whether training on the augmented dataset improves the true, population-level performance, while consuming as few real test data points as possible. Two formulations are considered: a direct test on the mean loss difference between the two trained models, and a symmetry-based test on the paired loss difference, which trades a stronger null assumption for faster evidence accumulation. For the latter, we introduce the {adaptive e-process sign-flip test} (aeSFT), a doubly adaptive procedure that adapts both the number of Monte Carlo sign-flip rounds, and hence the computational cost, and the amount of real test data consumed. aeSFT yields anytime-valid Type-I error control, with no need to pre-specify the test-set size. Experiments on a synthetic-data classification task, a DT-aided wireless packet-scheduling task, and a radio-map prediction task show that aeSFT identifies useful synthetic data using substantially fewer real test samples than mean-based sequential testing, matches the power of fixed-sample sign-flip testing and the paired $t$-test, while keeping the false-positive rate below the target level.
Chinese Translation
数字双胞胎(Digital Twins, DTs)和学习的世界模型越来越多地被用于生成合成数据,以增强用于训练工程系统中人工智能(Artificial Intelligence, AI)模型的稀缺真实数据集。然而,由于不可避免的模拟与现实(sim-to-real)差距,增强可能无法改善训练模型在真实数据分布上的表现。本文解决了由此产生的决策问题:给定一个真实数据集、一个候选合成数据集和一个固定的学习算法,决定在增强数据集上训练是否能改善真实的总体水平表现,同时尽可能消耗较少的真实测试数据点。我们考虑了两种形式:对两个训练模型之间的平均损失差异进行直接测试,以及基于对称性的配对损失差异测试,后者以更强的零假设换取更快的证据积累。对于后者,我们引入了自适应e过程符号翻转测试(adaptive e-process sign-flip test, aeSFT),这是一种双重自适应程序,能够同时调整蒙特卡洛符号翻转轮次的数量,从而调整计算成本,以及消耗的真实测试数据量。aeSFT提供了随时有效的第一类错误控制,无需预先指定测试集的大小。在合成数据分类任务、DT辅助的无线数据包调度任务和无线电地图预测任务上的实验表明,aeSFT能够使用显著少于基于均值的顺序测试的真实测试样本识别有用的合成数据,其效能与固定样本符号翻转测试和配对t检验相当,同时将假阳性率控制在目标水平以下。
cs.AI / 54 / 2608.27998
Automated Analysis Framework for Multilingual Climate-Health Literature Based on Multi-Agent Large Language Model
基于多智能体大型语言模型的多语言气候-健康文献自动分析框架
Abstract
The rapid proliferation of interdisciplinary and multilingual scientific literature has left traditional manual analysis and single-algorithm methods plagued by low efficiency, poor scalability, and insufficient domain adaptability. Targeting the literature analysis needs of the typical interdisciplinary climate-health field, this study proposes a multi-agent large language model automated analysis framework for multilingual scientific literature, which realizes full-process automation covering literature screening, structured information extraction, and standardized integration. With a central coordination module as the core, the framework deploys three dedicated agents for document evaluation, information extraction, and analytical review to mimic the literature analysis thinking of domain experts, and adopts a four-layer hallucination control strategy together with a manual verification procedure to ensure the accuracy and reliability of analytical outcomes. Validated on a bilingual Chinese-English corpus of 32,642 climate-health papers covering China from 1993 to 2023, the framework achieves an F1 score of 0.92 in core information extraction, and completes the extraction and standardization of 2,012 city-literature association pairs, offering effective technical support for large-scale evidence mining in the climate-health research domain.
Chinese Translation
跨学科和多语言科学文献的快速激增使得传统的手动分析和单一算法方法面临低效率、可扩展性差和领域适应性不足的问题。针对典型跨学科气候-健康领域的文献分析需求,本研究提出了一种多智能体大型语言模型自动分析框架,用于多语言科学文献的分析,实现了涵盖文献筛选、结构化信息提取和标准化整合的全流程自动化。该框架以中央协调模块为核心,部署了三个专用代理,分别用于文档评估、信息提取和分析审查,以模拟领域专家的文献分析思维,并采用四层幻觉控制策略及手动验证程序,以确保分析结果的准确性和可靠性。在对1993年至2023年覆盖中国的32,642篇气候-健康论文的双语中英文语料库进行验证后,该框架在核心信息提取中实现了0.92的F1分数,并完成了2,012对城市-文献关联的提取和标准化,为气候-健康研究领域的大规模证据挖掘提供了有效的技术支持。
cs.AI / 55 / 2608.27999
PhenoIntel: A Lifecycle-Aligned Multi-Agent Web Application for Verified, Accessible Plant Phenotype Analysis
PhenoIntel:一个与生命周期对齐的多智能体网络应用程序,用于经过验证的、可访问的植物表型分析
Abstract
Existing conversational plant-phenotyping platforms are difficult for plant scientists to use and lack the reliability scientific research demands: failed analyses are reported as valid measurements rather than flagged as missing, statistical tests run without checking assumptions, predictions carry no uncertainty estimate, and specialised hardware limits accessibility. We present PhenoIntel, a lifecycle-aligned multi-agent web platform that turns the full machine-learning workflow into a reliable, user-friendly phenotyping system. Nine specialised agents divide the analysis into stages, from image collection through model selection, inference, and reporting, rather than handing the whole task to one AI manager. Independent checks separate these stages, and every agent reads from and writes to one shared, fixed-structure record, so an inconsistent output from one stage is caught before it reaches the next. Uncertainty is matched to each model family, conformal prediction, detection-confidence spread, or Monte Carlo Dropout, rather than applied uniformly, and quality thresholds adapt to crop and task instead of one global cutoff. When no suitable model exists, PhenoIntel can propose, validate, and integrate a new one on its own. The model repository spans ten trained models across five crops and four imaging modalities. Classification models reach Macro F1 of 0.78-0.996; object-detection models reach 0.96 mAP@50 with a 54% reduction in counting error over an unoptimised baseline; and a temporal model reaches held-out Macro F1 of 0.7050. PhenoIntel runs in a browser on standard hardware, requiring no GPU, and a 1,200-test automated suite confirms complete pipeline execution. Every result carries calibrated uncertainty, validated statistics, and FAIR-compliant provenance, a combination existing conversational phenotyping tools do not offer.
Chinese Translation
现有的对话式植物表型平台对于植物科学家而言使用困难,且缺乏科学研究所需的可靠性:失败的分析被报告为有效测量而不是标记为缺失,统计测试在未检查假设的情况下运行,预测没有不确定性估计,专用硬件限制了可访问性。我们提出了PhenoIntel,一个与生命周期对齐的多智能体网络平台,将完整的机器学习工作流程转变为一个可靠、用户友好的表型系统。九个专门的智能体将分析划分为多个阶段,从图像采集到模型选择、推理和报告,而不是将整个任务交给一个AI管理者。独立检查将这些阶段分开,每个智能体从一个共享的固定结构记录中读取和写入,因此一个阶段的不一致输出会在到达下一个阶段之前被捕获。不确定性与每个模型家族相匹配,包括符合预测、检测置信度分布或蒙特卡洛Dropout,而不是统一应用,质量阈值根据作物和任务进行调整,而不是一个全局的截止值。当不存在合适的模型时,PhenoIntel可以自行提出、验证并整合一个新模型。模型库涵盖五种作物和四种成像模式下的十个训练模型。分类模型的宏观F1值达到0.78-0.996;目标检测模型在未优化基线的基础上,计数误差减少54%,达到0.96 mAP@50;而时间模型的保留宏观F1值为0.7050。PhenoIntel在标准硬件的浏览器中运行,无需GPU,并且一个1200个测试的自动化套件确认了完整的管道执行。每个结果都带有经过校准的不确定性、验证的统计数据和符合FAIR标准的来源,这是现有对话式表型工具所不具备的组合。
cs.AI / 56 / 2608.28011
Coverage, Not Credit: Failure-Credit Routing of Zeroth-Order Perturbation Budgets Does Not Improve On-Pool Sample Efficiency for LLM Agents
覆盖,而非信用:零阶扰动预算的失败信用路由并未提高LLM代理的池内样本效率
Abstract
Trajectory-level credit assignment can localize which module of a tool-using LLM agent causes failures using only verifiable signals. We ask whether such failure credit should route a fixed zeroth-order/evolution-strategies (ZO/ES) perturbation budget. Across a synthetic environment and frozen Qwen2.5-1.5B/3B and SmolLM2-1.7B agents, three task families, six allocation schemes, a credit-noise sweep, paired seeds, and exact sign-flip tests, we find no statistically detectable improvement over uniform allocation in any on-pool comparison (no gain of at least 2 percentage points). The joint soft-plus-sigma scheme is equivalent to uniform within a +/- 0.02 AUC margin on 1.5B and 3B; concentrating the full budget on the credit argmax is marginally equivalent on 1.5B, where that module is the verified bottleneck, and significantly worse on 3B. Inverse-propensity debiasing does not rescue routing, and misrouting costs up to -0.074 AUC in-house and -0.118 end-to-end on the BFCL-derived family. Across six fixed-step schedules, loss is linear in bottleneck starvation rate (R^2 = 0.94, descriptive), and a preregistered credit-free coverage floor removes detected harm. Matched-budget burst and step-compensating catch-up schedules are consistent with harm arising from insufficient cumulative parameter movement rather than update frequency. Our primary estimand is optimization efficiency on a fixed task pool. On unseen BFCL functions, the study's one exception is that soft routing exceeds uniform on held-out endpoints (+0.047, p = 0.031, n = 6). A plausible but untested reading is that routing-favored caller improvements transfer while uniform's on-pool gains reflect a synthesizer behavior specific to our harness. We report this exception explicitly and document three failure modes that can silently invalidate ZO/ES experiments on frozen LLMs.
Chinese Translation
轨迹级信用分配可以仅通过可验证信号定位工具使用的LLM代理中导致失败的模块。我们探讨这种失败信用是否应当路由固定的零阶/进化策略(ZO/ES)扰动预算。在一个合成环境中,以及冻结的Qwen2.5-1.5B/3B和SmolLM2-1.7B代理下,涉及三类任务、六种分配方案、信用噪声扫描、配对种子和精确的符号翻转测试,我们发现,在任何池内比较中,相较于均匀分配没有统计上可检测的改进(至少没有增加2个百分点)。联合soft-plus-sigma方案在1.5B和3B上在+/- 0.02 AUC的边际内等同于均匀分配;将全部预算集中在信用argmax上在1.5B上边际等同(该模块是经过验证的瓶颈),而在3B上则显著更差。逆倾向去偏不改善路由,错误路由在内部造成高达-0.074 AUC的损失,在BFCL派生的任务上造成-0.118的端到端损失。在六个固定步长计划中,损失与瓶颈饥饿率呈线性关系(R^2 = 0.94,描述性),而预注册的无信用覆盖底线消除了检测到的损害。匹配预算的突发和步长补偿追赶计划与由于累积参数移动不足而非更新频率造成的损害一致。我们的主要估计量是固定任务池上的优化效率。在未见过的BFCL函数上,本研究的唯一例外是软路由在保留的端点上超过均匀分配(+0.047,p = 0.031,n = 6)。一个合理但未经检验的解读是,路由偏好的调用者改进能够转移,而均匀分配的池内增益反映了我们测试平台特有的合成器行为。我们明确报告这一例外,并记录了三种可能会静默使ZO/ES实验在冻结LLM上失效的失败模式。
cs.AI / 57 / 2608.28027
String: An Agentic OS Where Every App Is a Markdown File
String:一个每个应用都是Markdown文件的自主操作系统
Abstract
LLM agents have become a new class of software user, but every surface they work through was designed for someone else. Pages are built for human eyes, which can skim and ignore; tool schemas for programs, which pay nothing to carry definitions they never call. An agent has neither luxury: it re-reads, and pays again for, everything it is shown on every turn. We present String, an open-source runtime that gives this user an interface of its own and treats the job as an operating-systems problem. Tool knowledge moves out of the agent's context and into a common layer that renders it back one view at a time as Markdown. A single SFMD (String-Flavored Markdown) document declares an application's views, typed actions, navigation, and credentials, and the runtime handles discovery, validation, execution, state, and secrets behind two core verbs: /open to see and /act to do. Web and app turn out to be two renderings of one architecture: an SFMD site serves styled HTML to browsers and the raw document to agents, so one grammar reaches apps, files, shells, and the web, even legacy HTML, with no per-site integration. Views stay partial by design, and the staging is causal: disclosing one tier of detail a single turn too early costs up to 23 accuracy points, while proper staging drops wrong-action selection from 28% to 2%. Privilege follows provenance: a remote page may call HTTP but never the shell, and caller-supplied text never expands a stored secret. On an 87-task benchmark that pairs each task with curated skills, operationalizing those procedures as on-demand String apps yields comparable aggregate success across six models from frontier to small (+1.3pp) while using 33.5% fewer tokens among completed episodes, and the resident interface stays a constant 53 tokens at any catalog size. We report the design, the evaluation, and what three months of production use taught us.
Chinese Translation
大型语言模型(LLM)代理已成为一种新的软件用户类别,但它们所使用的每个界面都是为其他人设计的。页面是为人类的视觉而构建的,人类可以快速浏览和忽略;工具模式是为程序设计的,这些程序并不在意携带它们从未调用的定义。代理没有这样的奢侈:它在每个回合中重新阅读并再次为它所展示的每一件事付出代价。我们提出了String,一个开源运行时,为这一用户提供了自己的接口,并将这一任务视为一个操作系统问题。工具知识从代理的上下文中移出,进入一个公共层,以Markdown格式逐次呈现。一个单一的SFMD(String-Flavored Markdown)文档声明了应用程序的视图、类型化操作、导航和凭证,而运行时则通过两个核心动词:/open(查看)和/act(执行)处理发现、验证、执行、状态和秘密。网页和应用程序实际上是同一架构的两种表现形式:一个SFMD站点向浏览器提供样式化的HTML,而向代理提供原始文档,因此同一语法能够覆盖应用程序、文件、shell和网页,甚至遗留的HTML,而无需每个站点的集成。视图设计上保持部分性,分层是因果的:过早披露一个层级的细节可能导致多达23个准确度点的损失,而适当的分层将错误操作选择的比例从28%降至2%。特权遵循来源:远程页面可以调用HTTP,但绝不能调用shell,调用者提供的文本永远不会扩展存储的秘密。在一个87项任务的基准测试中,每个任务与策划的技能配对,将这些程序作为按需String应用程序进行操作化,在从前沿到小型的六个模型中产生了可比的总体成功率(+1.3个百分点),同时在完成的回合中使用了33.5%更少的tokens,并且常驻接口在任何目录大小下始终保持53个tokens。我们报告了设计、评估以及三个月的生产使用教会我们的经验。
cs.AI / 58 / 2608.28062
WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents
WeAgent-MMSearch:用于多模态搜索代理的原生文本-视觉交互
Abstract
Multimodal search agents extend parametric knowledge with newly emerging and long-tail evidence from the open web. Yet many existing agentic search environments often expose retrieved evidence only as text and omit tool-returned images from subsequent context, reducing visually grounded trajectories to text-only reasoning. Long-horizon interaction also compounds tool-call, response-length, timeout, and budget failures, which can discard salvageable trajectories, waste rollout computation, and disturb policy updates. To address these issues, we introduce WeAgent-Harness, a multimodal agentic harness that supports native text-vision interaction and runtime recovery. Retrieved images receive persistent disk references, allowing the model to inspect, process, and cite them throughout the trajectory. Based on this harness, we develop WeAgent-MMSearch, an integrated system spanning data construction, agentic post-training, and multimodal rollout. For data construction, a strong MLLM uses WeAgent-Harness to discover, synthesize, and verify MMSearch-style tasks and collect expert trajectories. During post-training, our Failure-Aware GSPO (FA-GSPO) recovers salvageable abnormal rollouts and filters invalid ones to improve bounded multimodal planning and search.We also introduce VisTarget-Bench, a 150-task human-verified benchmark that pairs each question with a held-out target image, distinguishing image-retrieval failures from visual-perception failures. Evaluation on VisTarget-Bench and seven public benchmarks shows that agentic post-training improves the average score by 19.22 points, enabling our model to outperform similarly sized open-source models and rival models with roughly ten times its parameter count.
Chinese Translation
多模态搜索代理通过开放网络中新兴和长尾证据扩展参数知识。然而,许多现有的代理搜索环境通常仅以文本形式呈现检索到的证据,省略了后续上下文中的工具返回图像,从而将视觉基础的轨迹简化为仅文本推理。长时间交互还加剧了工具调用、响应长度、超时和预算失败,这可能会丢弃可挽救的轨迹,浪费展开计算,并干扰策略更新。为了解决这些问题,我们引入了WeAgent-Harness,这是一个支持原生文本-视觉交互和运行时恢复的多模态代理框架。检索到的图像获得持久的磁盘引用,使模型能够在整个轨迹中检查、处理和引用它们。基于该框架,我们开发了WeAgent-MMSearch,这是一个涵盖数据构建、代理后训练和多模态展开的综合系统。在数据构建方面,一个强大的多模态大语言模型(MLLM)利用WeAgent-Harness发现、合成和验证MMSearch风格的任务,并收集专家轨迹。在后训练过程中,我们的故障感知生成式策略优化(FA-GSPO)恢复可挽救的异常展开,并过滤无效展开,以改善有界的多模态规划和搜索。我们还引入了VisTarget-Bench,这是一个包含150个任务的人类验证基准,将每个问题与一个保留的目标图像配对,以区分图像检索失败和视觉感知失败。在VisTarget-Bench和七个公共基准上的评估表明,代理后训练使平均得分提高了19.22分,使我们的模型超越了同等规模的开源模型,并与参数数量大约是其十倍的模型相抗衡。
cs.AI / 59 / 2608.28065
Learning to Allocate Incentives for Incentivized Advertising via Offline Model-Based Reinforcement Learning
通过离线模型驱动强化学习学习激励广告的激励分配
Abstract
Complete your ad view and grab a 5-cent bonus! In incentivized advertising, a platform promises users a bonus before observing downstream ad revenue, encouraging them to click and complete ads. It must balance the incentive promised in advance against the revenue realized afterward: insufficient incentives forfeit monetization opportunities, whereas excessive incentives reduce net profit. Because current incentives may also shape user expectations and future engagement, incentive allocation is a sequential decision problem with delayed revenue, cost sensitivity, and carryover effects. Existing work has not studied decision-making algorithms for this setting. Auto-bidding assumes available ad opportunities, while targeted promotion optimizes incentives outside the ad monetization pipeline. We formulate the problem as an MDP and develop an offline model-based RL framework for cost-controllable sequential incentive allocation. It learns a world model of user feedback and ad revenue, then performs conservative policy optimization. An independent counterfactual scorer evaluates each learned policy on held-out logs, enabling pre-launch selection without costly online exposure. Experiments on large-scale industrial data and online A/B tests show that the scorer provides a stable offline signal. The deployment path from causal inference to offline RL and then Offline-MBRL further validates the framework: MB-IQL improves per-user net profit by 7.96\% over TD3+BC, whereas reverting to plain IQL reduces it by 6.56\% (both \(p<0.0001\)).
Chinese Translation
完成您的广告观看并获得5美分奖金!在激励广告中,平台在观察到下游广告收入之前承诺用户奖金,鼓励他们点击并完成广告。它必须在提前承诺的激励与之后实现的收入之间取得平衡:不足的激励会丧失货币化机会,而过多的激励则会降低净利润。由于当前的激励也可能影响用户的期望和未来的参与,因此激励分配是一个具有延迟收入、成本敏感性和延续效应的序列决策问题。现有研究尚未探讨这一环境下的决策算法。自动竞价假设可用的广告机会,而定向推广则在广告货币化流程之外优化激励。我们将该问题表述为马尔可夫决策过程(MDP),并开发了一个用于可控成本的序列激励分配的离线模型驱动强化学习(RL)框架。该框架学习用户反馈和广告收入的世界模型,然后进行保守的策略优化。一个独立的反事实评分器在保留的日志上评估每个学习到的策略,使得在没有昂贵在线曝光的情况下进行预发布选择成为可能。对大规模工业数据和在线A/B测试的实验表明,评分器提供了稳定的离线信号。从因果推断到离线RL再到离线模型驱动强化学习的部署路径进一步验证了该框架:MB-IQL相比于TD3+BC提高了每用户净利润7.96\%,而回归到普通IQL则降低了6.56 ext{(均为} p<0.0001 ext{)}。
cs.AI / 60 / 2608.28067
SEPO: Evidence-Grounded Prompt Optimization via Structural Editing
SEPO:通过结构编辑的基于证据的提示优化
Abstract
Existing API-only prompt optimisers are often described as interpretable, but in practice, this usually means only post-hoc inspectability: each iteration still rewrites the prompt as one opaque string, leaving a trace of full-prompt diffs rather than localisable, machine-readable edits. This paper introduces SEPO (Structural, Evidence-grounded Prompt Optimization), a multi-trajectory prompt optimiser centred on edit-effect lineage feedback. Rather than treating each iteration as an isolated whole-prompt rewrite, SEPO locally edits stable, typed units in a two-layer prompt schema, links the target and realised structural operations of each edit to the examples it newly fixes or breaks, and carries this edit-effect record forward to guide later architect calls on the same search branch. This makes prompt optimisation addressable, attributable, and actionable. Across a 14-task held-out suite, SEPO improves over the strongest baseline, GEPA, by 3.1 pp on Llama-3.1-8B-Instruct and 2.2 pp on Qwen3-8B, reaching 61.9% and 73.3% macro accuracy. SEPO also lies on both the optimisation-time and test-time Pareto frontiers, spending 2.9M optimisation tokens versus 4.1M for GEPA and producing prompts over 5x shorter.
Chinese Translation
现有的仅基于API的提示优化器通常被描述为可解释的,但在实践中,这通常仅意味着事后可检查性:每次迭代仍然将提示重写为一个不透明的字符串,留下完整提示差异的痕迹,而不是可定位的、机器可读的编辑。本文介绍了SEPO(结构性、基于证据的提示优化),这是一种以编辑效果谱系反馈为中心的多轨迹提示优化器。SEPO并不将每次迭代视为孤立的完整提示重写,而是局部编辑在两层提示架构中的稳定、类型化单元,将每次编辑的目标和实现的结构操作与其新修复或破坏的示例链接,并将这一编辑效果记录向前传递,以指导同一搜索分支上的后续架构调用。这使得提示优化变得可寻址、可归因和可操作。在一个包含14个任务的保留套件中,SEPO在Llama-3.1-8B-Instruct上比最强基线GEPA提高了3.1个百分点,在Qwen3-8B上提高了2.2个百分点,达到了61.9%和73.3%的宏观准确率。SEPO还位于优化时间和测试时间的帕累托前沿,使用了290万优化标记,而GEPA则使用了410万,并生成的提示长度超过5倍更短。
cs.AI / 61 / 2608.28099
Speculative Probing: LLM Monitoring at Speculative-Decoding Cost
推测性探测:在推测解码成本下的LLM监控
Abstract
Real-time classification during language model inference is valuable for safety filtering, behavioral analysis, and model monitoring, but current approaches force a trade-off between accuracy and efficiency. Hidden-state probes are fast but limited: they are either not context-aware: operating on a single vector and cannot model interactions across positions; or they are very costly: having dedicated classifier models (Llama Guard, Qwen Guard, LLM-as-judge) or performing computation on hidden states for all tokens and then pooling the results (MultiMax). This shows an intrinsic trade-off between efficiency and accuracy. However, we find that the speculative-decoding module in recent LLMs can be repurposed for efficient high-quality classification. By appending a trained soft prompt at the end of the target sequence, we can repurpose the speculative-decoding module into a sequence classifier. At inference time in a speculative-decoding pipeline, the KV cache is already in GPU memory, so classification adds negligible overhead. We evaluate on four classification tasks across four models (Qwen3.5-4B, 9B, 27B, MiniCPM4.1-8B). Our small probes consistently outperform zero-shot GPT-5.4-mini and, on multilingual prompt safety, match or beat specialized 8B safety classifiers (Qwen3Guard-Gen-8B, Llama-Guard-3-8B) without running a full LLM.
Chinese Translation
在语言模型推理过程中进行实时分类对于安全过滤、行为分析和模型监控具有重要价值,但当前的方法在准确性和效率之间存在权衡。隐状态探测器速度较快但有限:它们要么缺乏上下文意识:仅在单个向量上操作,无法建模位置之间的交互;要么成本非常高:需要专用的分类器模型(Llama Guard, Qwen Guard, LLM-as-judge)或对所有标记的隐状态进行计算,然后汇总结果(MultiMax)。这显示了效率与准确性之间的内在权衡。然而,我们发现最近的LLM中的推测解码模块可以重新用于高效的高质量分类。通过在目标序列末尾附加一个训练好的软提示,我们可以将推测解码模块转变为序列分类器。在推测解码管道的推理时,KV缓存已经在GPU内存中,因此分类的额外开销微乎其微。我们在四个模型(Qwen3.5-4B, 9B, 27B, MiniCPM4.1-8B)上评估了四个分类任务。我们的微型探测器在性能上始终优于零-shot GPT-5.4-mini,并且在多语言提示安全性方面,与专门的8B安全分类器(Qwen3Guard-Gen-8B, Llama-Guard-3-8B)相匹配或超越,而无需运行完整的LLM。
cs.AI / 62 / 2608.28144
The Shape of Power: A Multilingual Framework for Social Power Reasoning in Dialogues
权力的形态:对话中社会权力推理的多语言框架
Abstract
Social power plays a fundamental role in shaping human interaction, yet computational studies of power remain limited to narrow linguistic and cultural settings. Existing datasets further lack the demographic and relational depth needed for robust cross-cultural analysis. To address this gap, we introduce a theoretically grounded framework for studying social power in naturalistic multilingual dialogue through movie screenplays. The framework integrates a schema informed by social science theory, a native speaker annotation pipeline refined through pilot studies, and a custom interface for scalable cross-lingual analysis. Using this framework, we constructed an initial corpus containing 15,836 annotated instances from 100 scenes in French and Egyptian Arabic movies. Our analysis reveals strong agreement on observable demographic and contextual attributes, while socially interpretive aspects, such as power asymmetry and intention alignment, remain more contested, highlighting the complexity of social power across cultures. We evaluated 6 Large Language Models (LLMs) and Multimodal LLMs on cross-cultural social power reasoning, finding persistent gaps between human and model agreement in relational and theory-of-mind reasoning. Our work introduces the first extensible multilingual framework for studying social power in dialogues and provides an initial evaluation setting for studying cross-cultural social reasoning.
Chinese Translation
社会权力在塑造人际互动中发挥着基础性作用,但对权力的计算研究仍限于狭窄的语言和文化背景。现有数据集缺乏进行稳健跨文化分析所需的人口统计和关系深度。为了解决这一问题,我们提出了一个理论基础框架,通过电影剧本研究自然多语言对话中的社会权力。该框架整合了一个受社会科学理论启发的模式、通过初步研究精炼的母语者注释流程,以及一个用于可扩展跨语言分析的定制界面。利用该框架,我们构建了一个初步语料库,包含来自100个法语和埃及阿拉伯语电影场景的15,836个注释实例。我们的分析揭示了在可观察的人口统计和上下文属性上存在强一致性,而社会解释性方面,如权力不对称和意图一致性,则存在更多争议,突显了社会权力在不同文化中的复杂性。我们评估了6个大型语言模型(LLMs)和多模态LLMs在跨文化社会权力推理上的表现,发现人类与模型在关系和心智理论推理上的一致性存在持续差距。我们的工作引入了第一个可扩展的多语言框架,用于研究对话中的社会权力,并提供了一个初步的评估设置,以研究跨文化社会推理。
cs.AI / 63 / 2608.28152
Under-Mattress Temporal Sensing for Next-Day Agitation Risk Scoring in Dementia Wards
床垫下的时间感知用于评估痴呆病房次日激动风险
Abstract
Agitation fluctuates over short time horizons in people living with dementia, yet continuous physiological information for anticipating next-day risk is limited. We assessed whether contactless under-mattress signals from the preceding night inform next-day agitation risk and whether preserving minute-level temporal structure improves performance over conventional nightly summaries. We analyzed 423 patient-nights from 65 subjects in a specialized hospital dementia unit using two under-mattress sensing systems. A unified four-paradigm benchmark compared nightly handcrafted summaries, three-period handcrafted features, full-night sequence modeling, and sliding-window multiple-instance learning. Source-specific preprocessing and five-fold patient-grouped cross-validation were used, with performance estimated from pooled out-of-fold predictions. Evaluation included discrimination, calibration, fixed-threshold metrics, and a comparison of period-signal attribution patterns across two temporal models. Full-night sequence modeling achieved the highest discrimination (AUROC, 0.692; AUPRC, 0.849) and balanced accuracy (0.658). Both minute-level pipelines had higher AUROC than nightly summaries, but differences from three-period handcrafted features were uncertain. Cross-model attribution prioritized activity, heart rate, and respiratory rate during the core overnight period. Calibration remained limited. The preceding night's signals supported modest next-day risk discrimination, with minute-level temporal modeling outperforming nightly summaries. Prospective calibration and external validation are needed before use in individual care decisions. This patient-grouped benchmark identifies contactless overnight sensing as a promising biomedical engineering direction for agitation-risk research in hospitalized dementia cohorts.
Chinese Translation
在痴呆患者中,激动在短时间内波动,但用于预测次日风险的连续生理信息有限。我们评估了前一晚的无接触床垫下信号是否能提供次日激动风险的信息,以及保留分钟级时间结构是否能提高相较于传统夜间摘要的性能。我们分析了来自65名患者的423个夜间数据,使用了两种床垫下感知系统。在一个统一的四范式基准中,我们比较了手工制作的夜间摘要、三期手工特征、整夜序列建模和滑动窗口多实例学习。采用了源特定的预处理和五折患者分组交叉验证,性能通过汇总的外折预测进行估计。评估包括区分度、校准、固定阈值指标,以及对比两个时间模型的周期信号归因模式。整夜序列建模实现了最高的区分度(AUROC, 0.692; AUPRC, 0.849)和均衡准确率(0.658)。两个分钟级管道的AUROC均高于夜间摘要,但与三期手工特征的差异尚不确定。跨模型归因优先考虑了核心夜间期间的活动、心率和呼吸频率。校准仍然有限。前一晚的信号支持了适度的次日风险区分,分钟级时间建模优于夜间摘要。在用于个体护理决策之前,需要进行前瞻性校准和外部验证。该患者分组基准识别了无接触夜间感知作为住院痴呆患者激动风险研究的有前景的生物医学工程方向。
cs.AI / 64 / 2608.28165
CrabOS: An Operating System for Human-AI Co-inhabitation
CrabOS:一种用于人类与人工智能共存的操作系统
Abstract
AI agents are evolving into long-running computational entities that can invoke tools, maintain memory, and complete complex tasks across applications. In real-world settings, completing a task often requires humans and AI to take turns leading its execution. Such alternation depends on the seamless handoff of the work state of the task between humans and AI. Existing agent systems, however, provide humans and AI with separate work environments. AI agents must therefore rely on additional bridges to continue work: either developers build task-specific interfaces to access the work state, or users manually transfer relevant parts of it through screenshots or textual descriptions. Both approaches make handoffs costly and scale poorly. We propose Human-AI Co-inhabitation, a type of work environment that enables humans and AI to seamlessly take turns continuing work on the same task, and design and implement CrabOS to realize this concept. CrabOS represents the work state as natural-language-readable text objects shared by humans and AI, allowing both to access and manipulate it directly through the same auditable interface without bridges. Case studies show that CrabOS elevates support for complex tasks with alternating human and AI leadership from bridge-dependent application-level solutions to native operating-system capabilities, which provide a new foundation for developing and running AI agents.
Chinese Translation
人工智能代理正在演变为能够调用工具、维护记忆并在多个应用程序中完成复杂任务的长期计算实体。在现实世界中,完成一项任务通常需要人类和人工智能轮流主导其执行。这种交替依赖于人类与人工智能之间无缝的工作状态交接。然而,现有的代理系统为人类和人工智能提供了独立的工作环境。因此,人工智能代理必须依赖额外的桥梁来继续工作:要么开发人员构建特定于任务的接口以访问工作状态,要么用户通过截图或文本描述手动转移相关部分。这两种方法都使得交接成本高昂且扩展性差。我们提出了人类与人工智能共存(Human-AI Co-inhabitation),一种工作环境,使人类和人工智能能够无缝地轮流继续同一任务的工作,并设计和实现了CrabOS以实现这一概念。CrabOS将工作状态表示为人类和人工智能共享的自然语言可读文本对象,使双方能够通过相同的可审计接口直接访问和操作,而无需桥梁。案例研究表明,CrabOS将交替的人类和人工智能领导下的复杂任务支持从依赖桥梁的应用级解决方案提升到原生操作系统能力,为开发和运行人工智能代理提供了新的基础。
cs.AI / 65 / 2608.28178
Expert Knowledge & Machine Understanding: Bridging Reactome's Ontology with LLM Semantic Embeddings
专家知识与机器理解:将 Reactome 的本体与 LLM 语义嵌入相结合
Abstract
Biological knowledgebases like Reactome provide high-quality pathways that include biological elements' relationships and textual descriptions (metadata). The quality of such pathways is granted by manual curation, that presents, however, significant scalability challenges. Lately, numerous NLP tools have been proposed to cope with this issue, leveraging textual information to automatically expand biological knowledgebases. However, little exploration has been done so far to assess whether relationships among textual descriptions mirror higher order biological relationships. This study explores whether human-written descriptions in Reactome can be used to infer the experts' defined global hierarchical structure. To test this, we extracted from Reactome the Homo Sapiens hierarchy of pathways and their reactions (Reactome Hierarchy), and used textual metadata to reconstruct a Semantic Hierarchy, combining a sentence transformer model (SPECTER2) with a modified agglomerative nesting algorithm and a graph reconstruction algorithm. Quantitative (Laplacian Spectral Distance and Bootstrapping) and qualitative (global topological metrics) analyses confirm our hypothesis and indicate that the global hierarchical structure of pathways can be inferred by experts textual metadata.
Chinese Translation
生物知识库如 Reactome 提供了高质量的通路,包括生物元素之间的关系和文本描述(元数据)。这些通路的质量得益于人工审校,但这也带来了显著的可扩展性挑战。近年来,提出了许多自然语言处理(NLP)工具来应对这一问题,利用文本信息自动扩展生物知识库。然而,迄今为止,关于文本描述之间的关系是否反映更高阶生物关系的探索仍然较少。本研究探讨了 Reactome 中人类撰写的描述是否可以用来推断专家定义的全球层级结构。为此,我们从 Reactome 中提取了人类(Homo Sapiens)通路及其反应的层级(Reactome Hierarchy),并利用文本元数据重建了语义层级,结合了句子变换模型(SPECTER2)、修改过的聚合嵌套算法和图重建算法。定量(拉普拉斯谱距离和自助法)和定性(全局拓扑指标)分析证实了我们的假设,并表明通路的全球层级结构可以通过专家的文本元数据推断出来。
cs.AI / 66 / 2608.28228
Generative AI Alignment with Hinduism's Theological Plurality and Sacred Representation
生成性人工智能与印度教的神学多元性及神圣表现的对齐
Abstract
Generative AI systems are increasingly used to answer personal questions and mediate everyday practices, including religion. However, existing discussions around AI alignment and ethics have largely centered secular, Western, and Abrahamic assumptions about religion, offering limited attention to other faith-based traditions. In this paper, we examine how Hindu users engage with generative AI systems in relation to their religious knowledge, belief, and practice. Drawing on 15 semi-structured interviews with Bangladeshi Hindu participants, we analyze how users interpret AI-generated religious representations, scriptural explanations, devotional interactions, and synthetic religious media. We found that AI can be both accessible and ethically troubling. While AI supported scriptural inquiry, devotional visualization, and religious storytelling, our study also identified concerns about theological flattening, cultural misrepresentation, devotional manipulation, and the simulation of sacred presence and authority. We conclude by arguing that religious alignment in generative AI requires interpretive alignment: systems that disclose their limits, preserve plurality, and avoid simulating sacred authority and sycophantic personalization.
Chinese Translation
生成性人工智能系统越来越多地被用于回答个人问题和调解日常实践,包括宗教。然而,现有的关于人工智能对齐和伦理的讨论主要集中在世俗、西方和亚伯拉罕宗教的宗教假设上,对其他信仰传统的关注有限。本文探讨了印度教用户如何与生成性人工智能系统互动,涉及他们的宗教知识、信仰和实践。基于对15位孟加拉国印度教参与者的半结构化访谈,我们分析了用户如何解读人工智能生成的宗教表现、经典解释、虔诚互动和合成宗教媒体。我们的研究发现,人工智能既可以是可获取的,也可能引发伦理问题。尽管人工智能支持经典探究、虔诚可视化和宗教叙事,但我们的研究还识别出关于神学扁平化、文化误表述、虔诚操控以及神圣存在和权威的模拟等问题。我们最后认为,生成性人工智能中的宗教对齐需要解释性对齐:系统应披露其局限性,保持多元性,并避免模拟神圣权威和谄媚个性化。
cs.AI / 67 / 2608.28229
Stay Within Your Bounds: Distance-Guided Decoding for Guaranteed Context-Free Grammar Compliance
保持在你的界限内:基于距离引导的解码以确保无上下文文法的合规性
Abstract
Grammar-constrained decoding helps large language models produce syntactically valid structured outputs, such as code, JSON, and SQL. For context-free grammars, many practical decoders enforce local prefix feasibility: each token must keep the current prefix extendable to some valid completion. Yet, under tokenizer-grammar mismatch and finite token budgets, feasible prefixes may still fail to reach acceptance. We propose a lookahead-guided decoding framework for context-free grammars based on pushdown automata. Offline, we compute bounded pushdown summaries with reachability labels and upper-bound distances to acceptance. Online, these estimates guide horizon-aware pruning and beam search. The resulting decoder is syntactically sound: every output is accepted by the target grammar. Experiments on JSON, SQL, and Linear Temporal Logic (LTL) show both consistent syntactic validity and improved completion quality over existing baselines.
Chinese Translation
语法约束解码帮助大型语言模型生成语法上有效的结构化输出,如代码、JSON和SQL。对于无上下文文法,许多实用的解码器强制执行局部前缀可行性:每个标记必须保持当前前缀可扩展到某个有效的完成。然而,在标记器与文法不匹配和有限标记预算的情况下,可行前缀仍可能无法达到接受。我们提出了一种基于下推自动机的无上下文文法的前瞻性引导解码框架。在离线阶段,我们计算带有可达性标签和接受上界距离的有界下推摘要。在在线阶段,这些估计指导了基于视野的剪枝和束搜索。最终得到的解码器在语法上是健全的:每个输出都被目标文法接受。在JSON、SQL和线性时序逻辑(LTL)上的实验显示,与现有基准相比,语法有效性一致且完成质量有所提升。
cs.AI / 68 / 2608.28233
REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features
REINS:基于拒绝增强的抑制引导与稀疏自编码器特征
Abstract
Steering with Sparse Autoencoders (SAEs) offers a lightweight inference-time path for adapting the behavior of large language models without retraining. By exposing sparse and interpretable features, SAE steering provides a promising interface for safety control that guides harmful continuations toward refusal. However, we observe that complex wrappers can still undermine existing SAE steering methods on harmful prompts. To evaluate this failure mode systematically, we construct Generalized Undercover Instruction Safety Evaluation (GUISE), a dataset of harmful prompts with complex wrappers. Existing single direction SAE steering methods do not reliably produce refusals on harmful prompts, suggesting that refusal enhancement alone can be too weak when the harmful continuation path remains active. This motivates us to propose Refusal-Enhanced INhibitory Steering (REINS), which suppresses harmful continuation features and enhances safe refusal features in the same SAE feature space. Experiments on GUISE and other datasets show that prior methods either intervene too weakly or achieve only apparent safety through collapse, while REINS substantially reduces harmful responses, markedly improves safe refusals and largely preserves general capabilities.
Chinese Translation
使用稀疏自编码器(SAEs)进行引导提供了一种轻量级的推理时路径,可以在不重新训练的情况下调整大型语言模型的行为。通过暴露稀疏且可解释的特征,SAE引导为安全控制提供了一个有前景的接口,能够将有害的延续引导至拒绝。然而,我们观察到复杂的包装器仍然可能削弱现有的SAE引导方法在有害提示上的效果。为了系统地评估这种失败模式,我们构建了广义隐蔽指令安全评估(GUISE),这是一个包含复杂包装器的有害提示数据集。现有的单向SAE引导方法在有害提示上并不能可靠地产生拒绝,这表明仅仅增强拒绝可能在有害延续路径仍然活跃时过于薄弱。这促使我们提出拒绝增强抑制引导(REINS),该方法在同一SAE特征空间中抑制有害延续特征并增强安全拒绝特征。在GUISE及其他数据集上的实验表明,先前的方法要么干预过于薄弱,要么仅通过崩溃实现表面安全,而REINS显著减少了有害响应,显著改善了安全拒绝,并在很大程度上保留了整体能力。
cs.AI / 69 / 2608.28241
Beyond Task-Only Matching: Personalized Skill Routing with Counterfactual Evaluation
超越仅基于任务的匹配:基于反事实评估的个性化技能路由
Abstract
The rapid expansion of reusable skill repositories makes skill routing a critical capability for large language model (LLM) agents. Existing methods treat routing as task-only semantic matching. However, when users with incompatible constraints issue an identical request, this assumption conflates task relevance with skill suitability: a task-only router can select a semantically plausible skill that is unsuitable for the requesting user. To expose this failure mode, we formulate \textit{personalized skill routing} as profile-conditioned retrieval, in which relevance depends jointly on the task and the user profile. We first introduce a profile-counterfactual benchmark, in which the task is held fixed while changes in the user profile induce changes in the reference skill. We further construct paired counterfactual supervision and propose SkillFeed, a progressive retrieve-and-rerank framework that first establishes task--skill alignment and then learns profile-conditioned discrimination. By retrieving body-level evidence and reranking semantically similar but profile-conflicting candidates, SkillFeed identifies skills that satisfy both task requirements and user constraints. On SkillFeed-Bench, SkillFeed attains 75.1\% top-1 retrieval accuracy, a 23.1-point improvement over the corresponding pretrained routing baseline. Adding profile conditioning yields a 35.1-point gain on queries where user profile changes the reference skill. This contrast shows that user profiles are most consequential precisely when they change skill suitability. Our website is publicly available at http://www.aiskillfeed.com .
Chinese Translation
可重用技能库的快速扩展使得技能路由成为大型语言模型(LLM)代理的重要能力。现有方法将路由视为仅基于任务的语义匹配。然而,当具有不兼容约束的用户发出相同请求时,这一假设将任务相关性与技能适用性混淆:仅基于任务的路由器可能选择一个在语义上合理但对请求用户不适用的技能。为了揭示这一失败模式,我们将 extit{个性化技能路由}形式化为基于用户档案的检索,其中相关性共同依赖于任务和用户档案。我们首先引入一个档案-反事实基准,在该基准中,任务保持不变,而用户档案的变化引起参考技能的变化。我们进一步构建成对的反事实监督,并提出SkillFeed,一个渐进式的检索与重排序框架,首先建立任务与技能的对齐,然后学习基于档案的区分。通过检索体级证据并重排序语义相似但档案冲突的候选项,SkillFeed识别出满足任务要求和用户约束的技能。在SkillFeed-Bench上,SkillFeed达到了75.1%的顶级检索准确率,比相应的预训练路由基线提高了23.1个百分点。添加档案条件化在用户档案改变参考技能的查询中获得了35.1个百分点的提升。这一对比表明,用户档案在改变技能适用性时最为重要。我们的网站公开可用,网址为http://www.aiskillfeed.com。
cs.AI / 70 / 2608.28252
Regime-Aware Portfolio Management via Retrieval-Augmented LLM-Guided Expert Switching
基于检索增强的LLM引导专家切换的状态感知投资组合管理
Abstract
Financial markets are inherently non-stationary, making the effectiveness of individual portfolio-management strategies highly dependent on changing market conditions. This work proposes a retrieval-augmented expert-switching framework that dynamically selects portfolio management experts based on their historical performance under similar market situations. A dual-stream variational autoencoder represents asset-level and market-wide information, while a retrieval-based knowledge base stores historical situations and expert performance. During inference, an instruction-tuned LLM reasons over the retrieved evidence to identify the most appropriate expert rather than directly generating portfolio actions. We further establish a monotonicity property showing that adding a locally superior expert cannot degrade the switching mechanism's performance. Experiments across cryptocurrency, stock, and foreign-exchange markets show that the proposed selector achieves the highest cumulative return and Sharpe ratio among the evaluated selection strategies in all three markets. In the stock market, for example, cumulative return increases from 26% for the best fixed expert to 34%, while the Sharpe ratio improves from 0.74 to 0.96. Ablation results confirm the importance of both retrieval and LLM reasoning, while experiments with different expert-pool sizes demonstrate the value of complementary expertise. Overall, the findings support retrieval-grounded expert switching as an effective approach to adaptive portfolio management in non-stationary financial environments.
Chinese Translation
金融市场本质上是非平稳的,使得单一投资组合管理策略的有效性高度依赖于不断变化的市场条件。本研究提出了一种检索增强的专家切换框架,该框架根据专家在类似市场情境下的历史表现动态选择投资组合管理专家。双流变分自编码器表示资产级和市场级信息,而基于检索的知识库存储历史情境和专家表现。在推理过程中,经过指令调优的LLM对检索到的证据进行推理,以识别最合适的专家,而不是直接生成投资组合操作。我们进一步建立了单调性属性,表明添加一个局部优越的专家不会降低切换机制的性能。在加密货币、股票和外汇市场的实验中,所提议的选择器在所有三个市场的评估选择策略中实现了最高的累计回报和夏普比率。例如,在股票市场中,最佳固定专家的累计回报从26%增加到34%,而夏普比率从0.74提高到0.96。消融实验结果确认了检索和LLM推理的重要性,而不同专家池规模的实验展示了互补专业知识的价值。总体而言,研究结果支持基于检索的专家切换作为在非平稳金融环境中自适应投资组合管理的有效方法。
cs.AI / 71 / 2608.28256
Physics-Guided Flow Matching for CT Image Reconstruction
基于物理引导的流匹配用于CT图像重建
Abstract
Deep generative models have recently emerged as powerful priors for solving ill-posed inverse problems in CT, with diffusion-based approaches achieving state-of-the-art reconstruction performance. However, diffusion models typically rely on stochastic sampling procedures, long inference trajectories, and carefully tuned noise schedules, which can limit computational efficiency and numerical stability, especially at high spatial resolutions. In this work, we investigate Flow Matching as an alternative generative prior for CT reconstruction. We train a high-resolution Rectified Flow Matching model on 256x256 chest images from the Mayo Clinic Low-Dose CT dataset. To mitigate overfitting and limited anatomical variability, we employ a two-stage training strategy consisting of an initial phase with strong, anatomically informed data augmentation, followed by a fine-tuning phase with reduced or no augmentation to refine structural fidelity. The resulting model is capable of generating high-quality and anatomically coherent CT-like images, serving as a strong learned prior. We then evaluate multiple reconstruction methods specifically designed for Flow Matching models, including Plug-and-Play Flow, FlowDPS, Flower, and Flow-Priors (ICTM), and compare them against state-of-the-art diffusion-based reconstruction algorithms such as DDRM, DPS, and DiffPIR. Experimental results across several CT inverse problem settings show that Flow Matching-based approaches consistently outperform diffusion-based methods in terms of PSNR, SSIM, and perceptual quality, while requiring fewer sampling steps. Finally, we publicly release the trained Flow Matching model and accompanying code to facilitate reproducibility and future research. Overall, this work demonstrates that Flow Matching provides a stable, efficient, and effective alternative to diffusion models for high-resolution CT image reconstruction.
Chinese Translation
深度生成模型最近作为解决CT中病态逆问题的强大先验而崭露头角,基于扩散的方法实现了最先进的重建性能。然而,扩散模型通常依赖于随机采样过程、较长的推理轨迹和精心调整的噪声调度,这可能限制计算效率和数值稳定性,特别是在高空间分辨率下。在本研究中,我们探讨了流匹配(Flow Matching)作为CT重建的替代生成先验。我们在梅奥诊所低剂量CT数据集中对256x256的胸部图像训练了一个高分辨率的修正流匹配模型。为了减轻过拟合和有限的解剖变异性,我们采用了两阶段的训练策略,初始阶段使用强烈的、基于解剖信息的数据增强,随后是减少或不进行增强的微调阶段,以提高结构保真度。最终得到的模型能够生成高质量且解剖一致的CT样图像,作为强有力的学习先验。然后,我们评估了多种专门为流匹配模型设计的重建方法,包括插拔式流(Plug-and-Play Flow)、FlowDPS、Flower和流先验(Flow-Priors, ICTM),并将其与最先进的基于扩散的重建算法(如DDRM、DPS和DiffPIR)进行比较。在多个CT逆问题设置下的实验结果表明,基于流匹配的方法在PSNR、SSIM和感知质量方面始终优于基于扩散的方法,同时需要更少的采样步骤。最后,我们公开发布了训练好的流匹配模型及其相关代码,以促进可重复性和未来研究。总体而言,本研究表明流匹配为高分辨率CT图像重建提供了一个稳定、高效和有效的替代方案。
cs.AI / 72 / 2608.28264
Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration
寻找问题的根源:基于自动化故障归因的多智能体协作反思框架
Abstract
Multi-agent systems (MAS) powered by large language models have shown promise for complex tasks but suffer from high failure rates. Current self-reflection methods for MAS require all agents to reflect upon failure, overlooking a critical reality: failures typically stem from a specific agent leading the task astray, namely the decisive error agent, while others merely fulfill their regular duties. Forcing regular-behaving agents to reflect contaminates their memory with wrong insights. Hence, we propose DoCtOR (Diagnose-then-Correct PPO-enhanced Reflection), a novel reflection framework that enhances multi-agent collaboration. DoCtOR first identifies the decisive error step and decisive error agent through automated failure attribution, then employs counterfactual reasoning to generate a corrected decisive error step, and finally engages only the decisive error agent to produce targeted reflections. Experimental results show DoCtOR achieves 22%, 26%, and 27% improvements over initial success rates on HotPotQA, ChartQAPro, and Mind2Web datasets, outperforming Reflexion, Retroformer, and COPPER. We further establish the generalizability of our diagnose-then-correct paradigm and demonstrate that in low-resource settings, focusing reflection on reasoning steps after the decisive error step achieves comparable quality to reflecting on the complete failure trajectory.
Chinese Translation
由大型语言模型驱动的多智能体系统(MAS)在复杂任务中展现出潜力,但其故障率较高。目前,MAS的自我反思方法要求所有智能体对故障进行反思,忽视了一个关键现实:故障通常源于某个特定的智能体使任务偏离方向,即决定性错误智能体,而其他智能体则仅履行其常规职责。强迫表现正常的智能体进行反思会使其记忆受到错误见解的污染。因此,我们提出了DoCtOR(Diagnose-then-Correct PPO-enhanced Reflection),一种新颖的反思框架,旨在增强多智能体协作。DoCtOR首先通过自动化故障归因识别决定性错误步骤和决定性错误智能体,然后采用反事实推理生成修正后的决定性错误步骤,最后仅让决定性错误智能体进行针对性的反思。实验结果表明,DoCtOR在HotPotQA、ChartQAPro和Mind2Web数据集上的初始成功率分别提高了22%、26%和27%,超越了Reflexion、Retroformer和COPPER。我们进一步确立了我们的诊断-修正范式的普适性,并证明在低资源环境中,聚焦于决定性错误步骤后的推理步骤进行反思,其质量可与对完整故障轨迹的反思相媲美。
cs.AI / 73 / 2608.28271
RECAST: Recent & Context-Aware Sampling for Test-Time Adaptation in Streaming Biosignals
RECAST:用于流式生物信号测试时适应的近期与上下文感知采样
Abstract
Streaming biosignals vary across subjects and drift over time, so population-trained models lose accuracy during long-term monitoring. Test-time adaptation (TTA) enables online personalization by updating the model on incoming samples. But in a stream, a basic question is left open: \emph{which samples should drive each update?} Using all buffered samples blurs the update with irrelevant segments. Using only the latest segment makes the update noisy and unstable. The most useful samples are recent, aligned with the current physiological state, and reliable enough to learn from. We propose \textbf{RECAST} (REcent \& Context-Aware Sampling for TTA), a lightweight sampling module for buffered TTA frameworks. RECAST builds each adaptation batch from three signals: temporal recency, contextual similarity, and predictive reliability. It changes only which samples are used, leaving the model and the training objective unchanged. On two blood-pressure datasets, RECAST improves estimation accuracy and trend tracking over baselines and ablations. The per-patient gains are statistically significant on both datasets, with broad improvement on the regular benchmark and gains concentrated on the hardest patients in the emergency-department setting. RECAST stays practical, adding only sub-second latency per segment on a single GPU and CPU core.
Chinese Translation
流式生物信号在不同个体之间存在差异,并且随着时间的推移而漂移,因此基于人群训练的模型在长期监测中会失去准确性。测试时适应(TTA)通过更新模型以适应输入样本,实现在线个性化。然而,在流式数据中,一个基本问题尚未解决: extit{哪些样本应驱动每次更新?} 使用所有缓冲样本会使更新与无关段落混淆。仅使用最新段落则会使更新变得嘈杂且不稳定。最有用的样本是近期的、与当前生理状态一致的,并且足够可靠以供学习。我们提出了 extbf{RECAST}(用于TTA的近期与上下文感知采样),这是一个轻量级的采样模块,适用于缓冲TTA框架。RECAST从三个信号构建每个适应批次:时间的近期性、上下文的相似性和预测的可靠性。它仅改变使用的样本,而不改变模型和训练目标。在两个血压数据集上,RECAST在基准和消融实验中提高了估计准确性和趋势跟踪。每位患者的增益在两个数据集中均具有统计显著性,在常规基准上有广泛改善,并且在急诊科环境中,增益集中在最困难的患者身上。RECAST保持实用性,在单个GPU和CPU核心上每个段落仅增加亚秒级延迟。
cs.AI / 74 / 2608.28281
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
LoopArena:作为运行时控制器的模型基准测试用于循环工程
Abstract
Loop Engineering is emerging as a practice for organizing development work around coding agents. Instead of writing each prompt by hand, practitioners design loops that monitor progress, assign work, run checks, and decide what the agent should do next. Even with a capable coding agent, a loop may trust a stale progress note, skip needed verification, spend its budget in the wrong direction, or stop before the task is safe to submit. Yet the final outcome of one end-to-end run cannot tell whether success or failure reflects the loop's guidance or the coding agent's ability to carry out the task. We introduce LoopArena, a benchmark for evaluating how well one model can guide a separate coding agent through a long-running task. The model under evaluation is the \textbf{Controller}: after each coding round, it receives a structured summary of the run and instructs a separate, fixed coding agent, the \textbf{Worker}, on what to do or verify next, or decides whether to stop. LoopArena evaluates this ability in three complementary settings that differ in execution scope and cost. Type I scores next-step Loop Contract selection through execution-validated questions without running the Worker at evaluation time. Type II executes repeated control over a selected slice of a full task, while Type III evaluates the paired full task from its original state. On full tasks, the best observed Strict Success Rate is \textbf{24.69\%}, leaving substantial room for improvement in long-horizon loop control. Across Controllers, the paired reduction in estimated inference cost averages \textbf{64.4\%}, and Type II produces a similar ordering under the main Core criterion (Spearman's \(\rho=\textbf{0.9747}\)). We release the benchmark data and evaluation code at https://github.com/AMAP-ML/LoopArena .
Chinese Translation
循环工程正逐渐成为围绕编码代理组织开发工作的实践。实践者设计循环来监控进度、分配工作、进行检查并决定代理接下来应该做什么,而不是手动编写每个提示。即使有一个能够的编码代理,循环也可能信任过时的进度记录,跳过必要的验证,错误地花费预算,或者在任务尚未安全提交之前停止。然而,一个端到端运行的最终结果无法判断成功或失败是反映循环的指导还是编码代理执行任务的能力。我们介绍了LoopArena,这是一个评估一个模型如何有效指导一个独立编码代理完成长期任务的基准。被评估的模型是 extbf{控制器}:在每轮编码后,它接收运行的结构化摘要,并指示一个独立的固定编码代理 extbf{工人}接下来该做什么或验证什么,或者决定是否停止。LoopArena在三个互补的设置中评估这种能力,这些设置在执行范围和成本上有所不同。类型I通过执行验证的问题来评分下一步循环合同选择,而不在评估时运行工人。类型II对选定的完整任务片段进行重复控制,而类型III则从原始状态评估配对的完整任务。在完整任务中,观察到的最佳严格成功率为 extbf{24.69\%},在长期循环控制中仍有很大的改进空间。在所有控制器中,估计推理成本的配对减少平均为 extbf{64.4 extbf{ extbackslash%}},而类型II在主要核心标准下产生了类似的排序(斯皮尔曼系数 extit{ρ}= extbf{0.9747})。我们在https://github.com/AMAP-ML/LoopArena发布基准数据和评估代码。
cs.AI / 75 / 2608.28295
Memristive-Friendly Hadamard Reservoir Computing: Structured, Multiplier-Free Recurrences at Scale
适合忆阻器的哈达玛尔水库计算:大规模的结构化、无乘法器递归
Abstract
Reservoir Computing (RC) designs Recurrent Neural Networks around a fixed, i.e., untrained, recurrent layer, and is a natural candidate for neuromorphic hardware. Memristive-friendly reservoirs derive the neuron dynamics from memristive-device kinetics, but still rely on dense recurrent matrices, which are expensive to realize physically. In this paper, we replace the dense matrix with a structured orthogonal operator, built from sign diagonals, a permutation, and a fast Walsh-Hadamard transform. The operator is multiplier-free, requires $O(N)$ parameters and $O(N\log N)$ operations per step, and is never materialized as a matrix. We instantiate it in a standard and in a memristive-friendly Echo State Network, with one binary input connection per unit. Our mathematical analysis shows that exact orthogonality yields an echo state condition that is tight in the recurrent scaling, and a noise response that is predictable at design time. Moreover, the operator mixes the whole state in a single application. Experiments on twenty classification and seven regression benchmarks, at reservoir sizes up to $N = 8192$, show that the structured models match dense orthogonal reservoirs, and achieve better mean performance than the cycle reservoir by a margin that widens with size. Furthermore, we time the recurrent step on three hardware platforms, where it is up to $50\times$ faster than a dense product and $10^4\times$ smaller in memory. Finally, we ablate the operator and measure the response to noise, quantization, device mismatch and discrete faults.
Chinese Translation
水库计算(RC)围绕一个固定的、即未训练的递归层设计递归神经网络,是神经形态硬件的自然候选者。适合忆阻器的水库从忆阻器件的动力学中推导神经元动态,但仍依赖于密集的递归矩阵,这在物理实现上成本高昂。本文中,我们用一个结构化的正交算子替代了密集矩阵,该算子由符号对角线、置换和快速的沃尔什-哈达玛变换构成。该算子无乘法器,每步需要 $O(N)$ 参数和 $O(N ext{log} N)$ 操作,并且从未以矩阵形式实现。我们在标准和适合忆阻器的回声状态网络中实例化了该算子,每个单元有一个二进制输入连接。我们的数学分析表明,精确的正交性产生了一个在递归缩放中严格的回声状态条件,以及一个在设计时可预测的噪声响应。此外,该算子在单次应用中混合了整个状态。在高达 $N = 8192$ 的水库规模下,对二十个分类和七个回归基准的实验表明,结构化模型与密集正交水库相匹配,并且在性能均值上优于循环水库,且这一差距随着规模的增加而扩大。此外,我们在三个硬件平台上计时递归步骤,其速度比密集乘积快达 $50 imes$,内存占用小达 $10^4 imes$。最后,我们对算子进行了消融实验,并测量了对噪声、量化、器件不匹配和离散故障的响应。
cs.AI / 76 / 2608.28315
MAIL: Memory-driven, Adaptive, Incremental, and Literature-grounded Framework for Hypothesis Generation in Chemistry
MAIL:基于记忆驱动的自适应增量文献基础假设生成框架在化学中的应用
Abstract
The ever-expanding volume of the chemical literature offers unprecedented opportunities to generate novel and impactful hypotheses. However, the bottleneck lies in efficiently navigating this vast knowledge base to formulate high-quality, experimentally meaningful insights. While Large Language Models (LLMs) show promise for this task, existing methods often rely on static inspiration corpora, predefined heuristics, or laborious human-in-the-loop pipelines and decision-support frameworks that limit scalability and novelty. In this work, we propose an automated approach, a Memory-augmented, Adaptive, Incremental, and Literature-grounded (MAIL) framework for hypothesis generation in chemistry. Our MAIL method formulates hypothesis generation as a temporally grounded, memory-driven reasoning process, where hypotheses emerge from an evolving conceptual path that continuously accumulates and reinterprets prior knowledge. We evaluated the MAIL framework on a public TOMATO-Chem dataset and a newly curated and disseminated high-novelty nature/science challenge (HN-NS) dataset. Across both datasets, MAIL generates structurally coherent and mechanistically plausible hypotheses, achieves the highest MIOS and MPOS by more effectively recovering the central ideas and methodological elements of the historical target hypotheses, and obtains the highest overall expert-evaluation scores for scientific quality. These results demonstrate the potential of LLMs to autonomously explore chemical domains and generate hypotheses that are both innovative and chemically plausible.
Chinese Translation
不断扩展的化学文献为生成新颖且具有影响力的假设提供了前所未有的机会。然而,瓶颈在于如何高效地导航这一庞大的知识库,以形成高质量、实验上有意义的见解。虽然大型语言模型(LLMs)在此任务中展现出潜力,但现有方法往往依赖于静态的灵感语料库、预定义的启发式方法或繁琐的人机协作流程和决策支持框架,这限制了可扩展性和新颖性。在本研究中,我们提出了一种自动化的方法,即基于记忆增强的自适应增量文献基础(MAIL)框架,用于化学中的假设生成。我们的MAIL方法将假设生成视为一个时间基础的、记忆驱动的推理过程,其中假设从一个不断演变的概念路径中产生,该路径持续积累并重新解释先前的知识。我们在公共的TOMATO-Chem数据集和新近整理及传播的高新颖性自然/科学挑战(HN-NS)数据集上评估了MAIL框架。在这两个数据集中,MAIL生成了结构上连贯且机制上合理的假设,凭借更有效地恢复历史目标假设的核心思想和方法元素,达到了最高的MIOS和MPOS,并获得了科学质量的最高整体专家评估分数。这些结果展示了LLMs在自主探索化学领域和生成既创新又化学上合理的假设方面的潜力。
cs.AI / 77 / 2608.28334
Real-Valued Hyperdimensional Sequence Representations with Hadamard Product Binding and Shift Equivariance
具有哈达玛积绑定和位移等变性的实值超维序列表示
Abstract
Encoding temporal order is a fundamental requirement for sequence representations in Hyperdimensional Computing. Fractional Power Encoding provides similarity-preserving position vectors whose inner products approximate shift-invariant kernels, and it supports shift-equivariant transformations of encoded sequence representations. However, standard formulations of Fractional Power Encoding are primarily designed for binding operations such as circular convolution or complex-valued multiplication, which limits their compatibility with Hadamard product binding of real-valued vectors. This paper develops real-valued position encodings motivated by Random Fourier Features, aiming to retain the desirable properties of Fractional Power Encoding while supporting Hadamard-based operations. We propose three real-valued position-encoding variants: a real-valued baseline based on the inverse Fourier transform, and Sinusoid and Cosine-only representations derived from Random Fourier Features. Among them, the Sinusoid variant provides an explicit algebraic shift operator, allowing temporal shifts to be applied directly to the vector-encoded sequence representation without re-encoding the shifted sequence. Experiments on time-series classification datasets show that the proposed real-valued representations achieve performance comparable to standard Fractional Power Encoding while enabling computationally efficient Hadamard product binding. The Sinusoid variant offers the most favorable trade-off, combining efficient real-valued implementation with exact shift-equivariant transformations.
Chinese Translation
编码时间顺序是超维计算中序列表示的基本要求。分数幂编码提供了相似性保持的位置向量,其内积近似于位移不变核,并支持编码序列表示的位移等变变换。然而,分数幂编码的标准形式主要设计用于绑定操作,如循环卷积或复值乘法,这限制了它们与实值向量的哈达玛积绑定的兼容性。本文开发了基于随机傅里叶特征的实值位置编码,旨在保留分数幂编码的理想属性,同时支持基于哈达玛的操作。我们提出了三种实值位置编码变体:基于逆傅里叶变换的实值基线,以及从随机傅里叶特征派生的正弦和余弦-only 表示。其中,正弦变体提供了一个显式的代数位移算子,允许时间位移直接应用于向量编码的序列表示,而无需重新编码位移后的序列。在时间序列分类数据集上的实验表明,所提出的实值表示在性能上可与标准的分数幂编码相媲美,同时实现了计算上高效的哈达玛积绑定。正弦变体提供了最有利的权衡,结合了高效的实值实现与精确的位移等变变换。
cs.AI / 78 / 2608.28345
AGENT-O: A Semantic Agent Card Framework for Interoperable and Governed Healthcare AI Agents
AGENT-O:一个用于互操作和治理的医疗人工智能代理的语义代理卡框架
Abstract
AGENT-O is a modular ontology framework that defines a semantic Agent Card for representing health-oriented AI agent systems and supports assessment of reporting completeness in scientific publications. AGENT-O was developed as an OWL 2/RDF ontology covering runtime, models, workflow, tools, clinical use, evaluation, provenance, governance, and reporting assessment. Evaluation included ontology inventory, OWL-RL reasoning, three SHACL suites, 12 SPARQL competency queries, three cases, and model-assisted reporting-completeness assessment of 279 papers across five dimensions. The ontology contained 1,962 RDF triples and 1,922 Protege axioms, with 252 active classes, 198 active object properties, and 51 datatype properties. All SHACL suites conformed on example graphs, all competency queries returned prespecified evidence, and all 279 papers were scored. Incomplete reporting was highest for runtime/architecture (84.6%), governance/safety (82.8%), and provenance/reproducibility (78.1%), compared with evaluation (25.8%) and benchmark-process alignment (29.8%). AGENT-O supported semantic Agent Card representation and reporting assessment while revealing an evaluation-specification gap: evaluation and benchmark procedures were reported more consistently than runtime architecture, governance, and reproducibility. AGENT-O provides a reusable ontology, semantic Agent Card profile, and reporting-completeness workflow for structured reporting and gap identification, but does not assess agent quality or deployment readiness.
Chinese Translation
AGENT-O 是一个模块化本体框架,定义了一个语义代理卡,用于表示面向健康的人工智能代理系统,并支持对科学出版物报告完整性的评估。AGENT-O 作为一个 OWL 2/RDF 本体开发,涵盖了运行时、模型、工作流、工具、临床应用、评估、来源、治理和报告评估。评估包括本体清单、OWL-RL 推理、三个 SHACL 套件、12 个 SPARQL 能力查询、三个案例,以及对 279 篇论文在五个维度上的模型辅助报告完整性评估。本体包含 1,962 个 RDF 三元组和 1,922 个 Protege 公理,具有 252 个活跃类、198 个活跃对象属性和 51 个数据类型属性。所有 SHACL 套件在示例图上均符合,所有能力查询返回了预先指定的证据,所有 279 篇论文均进行了评分。报告不完整性在运行时/架构(84.6%)、治理/安全(82.8%)和来源/可重复性(78.1%)方面最高,而在评估(25.8%)和基准过程对齐(29.8%)方面最低。AGENT-O 支持语义代理卡表示和报告评估,同时揭示了评估规范的差距:评估和基准程序的报告一致性高于运行时架构、治理和可重复性。AGENT-O 提供了一个可重用的本体、语义代理卡配置文件和报告完整性工作流,用于结构化报告和差距识别,但不评估代理的质量或部署准备情况。
cs.AI / 79 / 2608.28360
Propagating construction-time knowledge quality into medical question answering: A framework grounded in clinical guidelines
将构建时知识质量传播到医学问答中的框架:基于临床指南的研究
Abstract
Large language models have facilitated knowledge graph (KG) construction from clinical guidelines, but extracted triples vary in structural validity and evidential support. Meanwhile, graph-augmented question answering (QA) systems typically optimize query relevance during retrieval, with limited reuse of quality information produced during KG construction. This creates a disconnect between construction-time quality control and inference-time evidence use. We investigate whether construction-time triple quality can serve as a persistent signal for downstream evidence selection and presentation. We propose a quality-aware framework that models structural conformance (SchemaConf) and evidential support (EvidScore) as complementary dimensions and fuses them into a per-triple quality signal, Q(t). Rather than using quality solely for filtering, the framework retains Q(t) and derived quality tiers as graph attributes and propagates them into quality-weighted subgraph retrieval and tier-conditioned evidence prompting, while preserving passage-level provenance. Experiments on Chinese diabetes clinical guidelines show that the utility of the quality signal is distribution dependent. Under cross-version and cross-model shift, the fused Q(t) provides stronger triple-quality discrimination than either component alone (AUC 0.748 vs. 0.703 for EvidScore and 0.645 for SchemaConf). In guideline-grounded QA, propagating construction-time quality reduces required-knowledge omission from 16.3% to 5.3% and conflicting outputs from 16.3% to 2.7%, with an evidence-grounded precision of 81.6% and near-zero invalid citations. Blinded clinician ratings favor the full framework over no retrieval (4.68 vs. 4.21 on a five-point scale) and approach the oracle condition (4.80), while cross-generator experiments show consistent trends.
Chinese Translation
大型语言模型促进了从临床指南中构建知识图谱(KG),但提取的三元组在结构有效性和证据支持方面存在差异。同时,图增强问答(QA)系统通常在检索过程中优化查询相关性,而对KG构建过程中产生的高质量信息的重用有限。这导致了构建时质量控制与推理时证据使用之间的脱节。我们研究了构建时三元组质量是否可以作为下游证据选择和呈现的持久信号。我们提出了一种质量感知框架,将结构一致性(SchemaConf)和证据支持(EvidScore)建模为互补维度,并将其融合为每个三元组的质量信号Q(t)。该框架不仅仅将质量用于过滤,而是将Q(t)及其衍生的质量层级作为图属性保留,并将其传播到质量加权的子图检索和层级条件的证据提示中,同时保留段落级的来源。对中国糖尿病临床指南的实验表明,质量信号的效用依赖于分布。在跨版本和跨模型的变化下,融合的Q(t)提供了比单独的任一组件更强的三元组质量区分能力(AUC 0.748 vs. 0.703的EvidScore和0.645的SchemaConf)。在基于指南的QA中,传播构建时质量将所需知识的遗漏率从16.3%降低到5.3%,冲突输出从16.3%降低到2.7%,证据基础的精确度为81.6%,无效引用接近零。盲评临床医生的评分更倾向于完整框架而非无检索(在五分制上为4.68 vs. 4.21),接近于理想条件(4.80),而跨生成器实验显示出一致的趋势。
cs.AI / 80 / 2608.28361
GRACE:Gradient-guided Coreset Selection for LLM Unlearning
GRACE:用于大语言模型遗忘的梯度引导核心集选择
Abstract
Machine Unlearning methods for Large Language Models typically assume pre-specified forget and retain sets. In realistic settings, however, requests may provide only a few examples of undesired behavior, requiring forget and retain sets to be inferred from heterogeneous corpora. We study this data-selection problem and propose GRACE , a gradient-guided coreset selection method that constructs both forget and retain sets for LLM unlearning. GRACE first computes a forget direction from seed examples that elicit the undesired behavior, then selects a compact forget coreset whose gradients approximate this direction using non-negative orthogonal matching pursuit. To preserve model utility, it selects retain examples after projecting out the forget direction and applying clustered orthogonal matching pursuit in the remaining gradient space. Across two target domains, two model families, and four unlearning algorithms, GRACE improves model utility while maintaining comparable forget quality, with particularly consistent gains over prior gradient-based selection methods.
Chinese Translation
大语言模型的机器遗忘方法通常假设预先指定的遗忘集和保留集。然而,在现实情况下,请求可能仅提供少量不希望出现行为的示例,这要求从异构语料中推断遗忘集和保留集。我们研究了这一数据选择问题,并提出了GRACE,一种梯度引导的核心集选择方法,构建大语言模型遗忘的遗忘集和保留集。GRACE首先从引发不希望行为的种子示例中计算遗忘方向,然后使用非负正交匹配追踪选择一个紧凑的遗忘核心集,其梯度近似该方向。为了保持模型效用,它在投影出遗忘方向后选择保留示例,并在剩余的梯度空间中应用聚类正交匹配追踪。在两个目标领域、两个模型家族和四种遗忘算法中,GRACE在保持可比的遗忘质量的同时提高了模型效用,尤其是在先前基于梯度的选择方法上取得了一致的增益。
cs.AI / 81 / 2608.28363
EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
EvoUndo:受恢复性约束的自我演化用于大型语言模型代理的利用
Abstract
LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a successful mutation may leave persistent effects that cannot be safely reversed in states different from the one in which it was created. We introduce EvoUndo, a framework for representing, synthesizing, diagnosing, and independently verifying recoverability of model-generated self-modifications across counterfactual states. Across 600 unseen one-shot self-evolution tasks, we identify 197 capability-improving mutations that fail recoverability verification. Under the original recovery representation, conventional repair strategies recover 0/197 of these natural failures. Deterministic oracle analysis recovers 48/197 under the original recovery language L0, while the extended recovery calculus increases empirical oracle recovery to 191/197. A protocol-locked 2x2 grounding-by-expressivity intervention then separates two bottlenecks: exact state-address grounding increases successful recovery from 0/48 to 38/48 (79.2%) when the original language is sufficient, while extending the recovery language enables recovery on 142/143 (99.3%) failures in the oracle-defined S1 stratum. On the primary gpt-oss-120b backbone, adding exact-address diagnostics to the richer language reduces recovery to 133/143 (93.0%); a Qwen3.8-27B replication preserves the grounding and expressivity effects but not this negative interaction, indicating that the latter is model-dependent. These results indicate that reliable agent self-evolution requires co-designing verification, state grounding, witness semantics, and recovery-language expressivity rather than relying on iterative prompting alone.
Chinese Translation
大型语言模型(LLM)代理越来越多地在运行时修改自己的提示、工具、中间件、资源和执行环境。这种自我演化可以提高能力,但成功的变异可能会在与其创建时不同的状态下留下无法安全逆转的持久影响。我们提出了EvoUndo,一个框架,用于表示、合成、诊断和独立验证模型生成的自我修改在反事实状态下的可恢复性。在600个未见的一次性自我演化任务中,我们识别出197个提高能力的变异,但未能通过可恢复性验证。在原始恢复表示下,传统的修复策略对这些自然失败的恢复率为0/197。确定性神谕分析在原始恢复语言L0下恢复了48/197,而扩展恢复演算将经验神谕恢复率提高到191/197。一个协议锁定的2x2表达性干预然后分离了两个瓶颈:当原始语言足够时,精确状态地址的基础增强了成功恢复率,从0/48提高到38/48(79.2%),而扩展恢复语言则使得在神谕定义的S1层次上恢复142/143(99.3%)的失败成为可能。在主要的gpt-oss-120b主干上,向更丰富的语言添加精确地址诊断将恢复率降低到133/143(93.0%);Qwen3.8-27B的复制保留了基础和表达性效果,但未能保留这种负交互,表明后者是模型依赖的。这些结果表明,可靠的代理自我演化需要共同设计验证、状态基础、见证语义和恢复语言的表达性,而不是仅仅依赖于迭代提示。
cs.AI / 82 / 2608.28384
MAP: A Benchmark on Multimodal Accessibility Planning for Real World Places
MAP:针对现实世界地点的多模态可达性规划基准
Abstract
We introduce MAP, the first benchmark to evaluate multimodal AI systems as assistants for users with accessibility requirements when planning visits to places in the real world. In our evaluation, systems are presented with requests to verify or recommend a point of interest meeting an accessibility requirement. MAP contains two novel assessments: Claim verification for accessibility planning assesses if information on places and stated accessibility features is supported and identifies places that satisfy requested accessibility features. Visual evidence retrieval for accessibility planning checks if a multimodal AI system can select visual evidence for the requested place and accessibility feature. Our methodology supports comparison of AI systems in a setting where place information and accessibility information can change over time by evaluating systems and refreshing ground truth data at scheduled times. The benchmark is based on automatic rating and human rating for a proportion of responses.
Chinese Translation
我们介绍了MAP,这是第一个基准,用于评估多模态人工智能系统作为具有可达性需求的用户在规划访问现实世界地点时的助手。在我们的评估中,系统会收到请求,以验证或推荐符合可达性要求的兴趣点。MAP包含两个新颖的评估:可达性规划中的声明验证评估地点信息和所述可达性特征是否得到支持,并识别满足请求的可达性特征的地点。可达性规划中的视觉证据检索检查多模态人工智能系统是否能够为请求的地点和可达性特征选择视觉证据。我们的方法论支持在地点信息和可达性信息可能随时间变化的环境中比较人工智能系统,通过在预定时间评估系统并刷新真实数据。该基准基于自动评分和人工评分对部分响应进行评估。
cs.AI / 83 / 2608.28393
Timing-Aware Repurchase Prediction for Web-Scale E-Commerce: Survival Models for Multi-Surface Grocery Recommendation
基于时间的再购买预测在大规模电子商务中的应用:多维度杂货推荐的生存模型
Abstract
Repurchase recommenders in e-commerce are commonly framed as a binary question asking "will this customer buy this item within W days", a formulation that requires a separately trained model for every horizon of interest. We replace this stack with survival models that predict time-to-repurchase directly, and evaluate them on millions of customers from a major grocery e-commerce platform across more than thirty ablation configurations. Our study makes three contributions. First, an empirical hazard analysis reveals a slightly decreasing marginal hazard (k ~ 0.9), differing from the common intuition that grocery items become more likely to be repurchased the longer since the last purchase (increasing hazard, k > 1). Log-Normal achieves the best marginal fit (R^2 = 0.998) and the best ranking, despite Weibull providing the best conditional residual fit, revealing an apparent discrepancy we analyze in detail. Second, a single Accelerated Failure Time (AFT) model replaces three per-horizon binary classifiers, matching or exceeding each at its own horizon while using roughly 3x fewer total trees. Feature importance reshuffles under the survival objective: channel-cadence and recency signals rise while aggregate frequency counts fall. Third, a 4-parameter parametric calibration maps raw survival CDFs to per-horizon probabilities with zero cross-horizon monotonicity violations. Calibration quality varies by an order of magnitude across the AFT family: Exponential AFT (Weibull k=1) achieves expected calibration error (ECE) ~1e-4, roughly 10x lower than Log-Normal, while ranking metrics agree within 0.3% relative. We adopt Exponential AFT for probability-consuming surfaces and Log-Normal for pure ranking, exposing a principled calibration-ranking trade-off within a single AFT family.
Chinese Translation
电子商务中的再购买推荐通常被表述为一个二元问题,即“该客户在W天内会购买该商品吗”,这种表述需要为每个感兴趣的时间范围单独训练一个模型。我们用生存模型直接预测再购买的时间,评估了来自一个主要杂货电子商务平台的数百万客户,并进行了超过三十种消融配置的实验。我们的研究有三个贡献。首先,实证风险分析揭示了边际风险略微下降(k ~ 0.9),这与常见的直觉相悖,即杂货商品在距离上次购买时间越长时再购买的可能性越大(风险增加,k > 1)。尽管Weibull模型提供了最佳的条件残差拟合,Log-Normal模型实现了最佳的边际拟合(R^2 = 0.998)和最佳排名,揭示了我们详细分析的明显差异。其次,一个单一的加速失效时间(AFT)模型替代了三个每个时间范围的二元分类器,在各自的时间范围内匹配或超越每个分类器,同时使用的总树木数量大约减少了3倍。在生存目标下,特征重要性发生了重新排序:渠道节奏和最近性信号上升,而总频率计数下降。第三,四参数参数校准将原始生存CDF映射到每个时间范围的概率,且没有跨时间范围单调性违反。校准质量在AFT家族中差异显著:指数AFT(Weibull k=1)实现了预期校准误差(ECE)约为1e-4,约为Log-Normal的10倍低,而排名指标相对一致在0.3%以内。我们为概率消耗型表面采用指数AFT,为纯排名采用Log-Normal,揭示了在单一AFT家族内的原则性校准与排名的权衡。
cs.AI / 84 / 2608.28399
RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents
RetailAgent:自我条件下多模态大语言模型交易代理中的结构性不利时机
Abstract
In financial markets, a sequential policy that reacts systematically to price movements may become predictable to other market participants. This paper studies whether large language model (LLM) agents exhibit such directional structure through RetailAgent, an experimental framework in which an LLM observes anonymized intraday equity price histories and permitted state, then repeatedly chooses long (hold the stock) or flat (stay out) before the subsequent interval return is revealed. We compare returns during long and flat intervals along the same stock's intraday path after removing the overall fraction of long decisions. This exposure-matched measure reveals persistent negative timing across modality, horizon, state, and model family. Shuffling saved action sequences substantially attenuates the effect, showing that alignment between actions and subsequent returns drives the negative score. Feeding self-authored memories into decisions further increases policy persistence, while timing becomes more negative among stock-days on which the agent uses both actions. These results reveal stable, recoverable directional structure in sequential LLM financial decisions and a behavioral signal for studying how another participant could respond to a predictable policy.
Chinese Translation
在金融市场中,系统性地对价格变动作出反应的顺序策略可能会变得对其他市场参与者可预测。本文研究了大型语言模型(LLM)代理是否通过RetailAgent这一实验框架展现出这种方向性结构。在该框架中,LLM观察匿名的日内股票价格历史和允许的状态,然后在后续区间回报揭示之前,反复选择做多(持有股票)或持平(不参与)。我们在去除总体做多决策比例后,比较同一股票在日内路径上的做多和持平区间的回报。这一匹配暴露的度量揭示了在不同模态、时间范围、状态和模型家族中持续存在的负时机效应。打乱保存的行动序列显著减弱了这一效应,表明行动与后续回报之间的对齐驱动了负得分。将自我创作的记忆纳入决策进一步增加了策略的持续性,而在代理同时使用两种行动的股票日中,时机变得更加负面。这些结果揭示了顺序LLM金融决策中稳定且可恢复的方向性结构,以及研究其他参与者如何响应可预测策略的行为信号。
cs.AI / 85 / 2608.28402
VERA-8B: Evidence-Grounded Audit Risk Reasoning from SEC Filings
VERA-8B:基于证据的审计风险推理来自SEC文件
Abstract
Across audit applications, judgments must be supported by reasonable evidence. However, standard financial language models prioritize fluency over evidence. They are built for general financial reasoning and may produce plausible but ambiguous answers, creating a grounding gap that makes them unsuitable for audit work. We address this gap with VERA-8B, a new end-to-end audit reasoning system that identifies audit risks before enforcement actions occur. Constructing such a model raises several challenges, as no prior machine learning work targets pre-enforcement audit prediction. To our knowledge, we are the first to unify SFT and GRPO for evidence-grounded audit reasoning under one evidence standard, achieving performance that surpasses all evaluated baselines. Because auditing cannot tolerate unsupported claims, we introduce abstention and uncertainty qualification to defer uncertain or evidence-incomplete cases. Finally, we design an AuditBridge to ground model reasoning for practical audit work. It transforms raw filings into verified records and then into reviewer-ready reports, bridging finance and computation with broad generality. Together, these components produce auditable, review-ready outputs suitable for practical audit work.
Chinese Translation
在审计应用中,判断必须有合理的证据支持。然而,标准的金融语言模型优先考虑流畅性而非证据。它们是为一般金融推理构建的,可能会产生似是而非但模糊的答案,从而造成一个基础差距,使其不适合审计工作。我们通过VERA-8B来解决这一差距,这是一种新的端到端审计推理系统,能够在执法行动发生之前识别审计风险。构建这样的模型面临多个挑战,因为之前没有机器学习工作专门针对执法前的审计预测。据我们所知,我们是首个将SFT和GRPO统一为基于证据的审计推理,并在一个证据标准下进行的研究,取得的性能超越了所有评估的基准。由于审计不能容忍没有支持的主张,我们引入了弃权和不确定性资格,以推迟不确定或证据不完整的案例。最后,我们设计了一个AuditBridge,以为实际审计工作提供模型推理的基础。它将原始文件转化为经过验证的记录,然后转化为审阅者准备好的报告,广泛地桥接了金融与计算。综合这些组件,产生了适合实际审计工作的可审计、准备审阅的输出。
cs.AI / 86 / 2608.28421
Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs
可验证奖励的程序学习:后训练大语言模型的符号反向传播
Abstract
Post training a language model to reason means updating its weights. Supervised finetuning and reinforcement learning both place the acquired capability inside the model where it cannot be inspected cannot be checked step by step and cannot be moved to another model. We argue that for tasks whose intermediate steps admit verification, reasoning is better placed outside the base models weights as an explicit program composed from deterministic and neural primitives. We introduce PLVR (Program Learning with Verifiable Rewards): a post training method that learns such programs directly from input-output examples. Its mechanism is symbolic backpropagation: each program layer carries a typed ontology a loss is computed at the output against ground truth and required input ontologies are propagated backward by type inference over primitive signatures: an analogue of the chain rule in which credit assignment is a derivation rather than an estimate. Where RLVR verifies a terminal outcome, PLVRs reward is a per step contract verdict dense over program structure. On LiveCodeBench v6 and Tau2Bench, 30B base models with PLVR outperform RL at matched budget by 27.8 points on average and frontier models an order of magnitude larger by 13.6 points. A single primitive library serves two benchmarks, so the marginal cost of a new task is 100 examples of program search and no new finetuning data. Replacing the loss guided search with uniform sampling over the same type admissible space at equal budget collapses the median program from 65.6 to 17.5, identifying the backward pass rather than the type system as the source of the advantage. We release the symbolic backpropagation library and a conformance checker so the method can be applied to primitive libraries other than our own.
Chinese Translation
对语言模型进行后训练以实现推理意味着更新其权重。监督微调和强化学习都将获得的能力置于模型内部,无法逐步检查和验证,也无法转移到其他模型。我们认为,对于那些中间步骤可以验证的任务,推理更适合放在基础模型权重之外,作为由确定性和神经原语组成的显式程序。我们引入了 PLVR(可验证奖励的程序学习):一种直接从输入-输出示例中学习此类程序的后训练方法。其机制是符号反向传播:每个程序层携带一个类型本体,输出处的损失是针对真实值计算的,所需的输入本体通过对原始签名的类型推断向后传播:这类似于链式法则,其中信用分配是推导而非估计。RLVR 验证终端结果,而 PLVR 的奖励是对程序结构的每一步合同裁决。在 LiveCodeBench v6 和 Tau2Bench 上,使用 PLVR 的 30B 基础模型在匹配预算下平均超越 RL 27.8 分,而前沿模型则超越 13.6 分。一个原语库服务于两个基准,因此新任务的边际成本为 100 个程序搜索示例,而无需新的微调数据。在相同类型可接受空间内以均匀采样替代损失引导搜索,预算相等时中位程序从 65.6 降至 17.5,识别出反向传播而非类型系统是优势的来源。我们发布了符号反向传播库和一致性检查器,以便该方法可以应用于除我们自己的原语库之外的其他库。
cs.AI / 87 / 2608.28433
Prove2Me: An Open Collaborative Platform for Scaling Math Formalization
Prove2Me:一个开放的协作平台,用于扩展数学形式化
Abstract
Proof assistants such as Lean 4 promise the paradigm of formally verified mathematics, but large-scale formalization projects have faced major barriers to entry, including the need for expertise in formal verification (as well as the underlying mathematics) and the significant time required for writing formal proofs. AI coding agents have dramatically reduced these barriers; human users can now use natural language to prompt agents to write complex proofs in Lean. This opens up the intriguing possibility of internet-scale mathematical collaboration involving both humans and AI agents, where correctness is machine-checked. To realize this possibility, we introduce Prove2Me (https://prove2.me), an open collaborative platform for formalizing mathematics. Users launch formalization "missions", to which AI agents contribute formal proofs toward completion. We designed mechanisms and a specialized harness in Prove2Me that enable large-scale collaboration so that agents can build on one another's work and freely reuse existing results. In doing so, Prove2Me aims to turn math formalization into a scalable, crowd-sourced effort open to anyone with an agent.
Chinese Translation
证明助手如 Lean 4 预示着形式验证数学的范式,但大规模形式化项目面临着主要的准入障碍,包括对形式验证(以及基础数学)的专业知识的需求,以及撰写形式证明所需的显著时间。人工智能编码代理显著降低了这些障碍;人类用户现在可以使用自然语言提示代理在 Lean 中撰写复杂的证明。这为涉及人类和人工智能代理的互联网规模数学协作开辟了引人入胜的可能性,其中正确性由机器检查。为了实现这一可能性,我们推出了 Prove2Me(https://prove2.me),一个用于形式化数学的开放协作平台。用户启动形式化“任务”,AI 代理为完成这些任务贡献形式证明。我们在 Prove2Me 中设计了机制和专门的框架,使大规模协作成为可能,以便代理可以在彼此的工作基础上构建并自由重用现有结果。通过这样做,Prove2Me 旨在将数学形式化转变为一个可扩展的、众包的努力,向任何拥有代理的人开放。
cs.AI / 88 / 2608.28447
Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning
学习使用工具:用于工具集成数学推理的强化学习
Abstract
Current large language models (LLMs) increasingly benefit from external tool integration, especially for tasks requiring reliable computation and verification. Motivated by this, we study calculator tool calling for improving mathematical reasoning on the Countdown task. We first analyze reasoning failures and find that calculation errors account for a substantial portion of incorrect responses. We then construct supervised fine-tuning datasets to teach the model useful tool-use patterns and how to interpret returned outputs. Building on this tool-formatted policy, we apply several on-policy reinforcement learning methods, including RLOO, RLOO++, GRPO, and DAPO, using automatically verifiable final-answer rewards. To enable a more reliable evaluation, we construct a fresh 1,024-problem held-out Countdown benchmark with no exact overlap with the training data. Our results show that calculator tool integration consistently improves both SFT and RL baselines, yielding roughly 10 percentage-point gains across pass@k. Among the RL methods, Tool-DAPO achieves the strongest performance, improving pass@1 from 35.8% for Tool-SFT to 66.0%. Further analysis shows that RL encourages more effective tool use even when only final-answer rewards are provided. These findings suggest that tool integration reduces arithmetic and verification errors, while RL increases the probability of correct reasoning traces.
Chinese Translation
当前的大型语言模型(LLMs)越来越多地受益于外部工具的集成,特别是在需要可靠计算和验证的任务中。基于此动机,我们研究了计算器工具调用,以改善在倒计时任务中的数学推理。我们首先分析了推理失败的原因,发现计算错误占据了不正确响应的相当大一部分。随后,我们构建了监督微调数据集,以教导模型有用的工具使用模式以及如何解释返回的输出。在此工具格式化策略的基础上,我们应用了几种在线强化学习方法,包括 RLOO、RLOO++、GRPO 和 DAPO,使用自动可验证的最终答案奖励。为了实现更可靠的评估,我们构建了一个全新的 1,024 道题目的倒计时基准集,该基准集与训练数据没有完全重叠。我们的结果表明,计算器工具的集成始终改善了 SFT 和 RL 基线,在 pass@k 上获得了大约 10 个百分点的提升。在所有 RL 方法中,Tool-DAPO 的表现最强,将 Tool-SFT 的 pass@1 从 35.8% 提升至 66.0%。进一步分析表明,即使仅提供最终答案奖励,RL 也能鼓励更有效的工具使用。这些发现表明,工具集成减少了算术和验证错误,而 RL 则增加了正确推理轨迹的概率。
cs.AI / 89 / 2608.28475
COVER: Identifiable Evaluation of Coalition Routing
COVER:可识别的联盟路由评估
Abstract
When a multi-agent system changes its team, it also changes the messages and final answer it produces, so an end-to-end accuracy gap does not by itself identify a routing effect. We introduce method, an evaluation contract that fixes a public information boundary, downstream stack G, and finite legal team family before outcomes are generated. Complete coverage identifies exact finite-benchmark oracle regret conditional on that stack. For any finite collection of frozen policies, executing the union of their distinct selected teams is the minimal assumption-free support for every pairwise policy contrast, though not for absolute oracle regret. Two controlled tables with source-ID-disjoint splits test the instrument. On MuSiQue-12, a pre-specified privileged positive control improves regret from 0.532 to 0.402; a later public-interface control reaches 0.424 versus 0.554 but is retrospective. On HotpotQA-4, a pre-specified public direct scorer improves regret from 0.313 to 0.110. In fixed-stack Llama execution, verified route regret improves by 0.190, while the raw-answer gain is 0.010 with an interval crossing zero. A five-family ToolSandbox variant-shift validation exhaustively evaluates 16 declared teams on 14 untouched task variants (224/224 valid rows): the declared-family oracle reaches 0.768 safe-evidence completion, while the prospectively frozen router gets 0.637 (regret 0.131), failing the predeclared 0.10 criterion. A later retrospective comparator reaches 0.655, matching all-workers with 4.57 versus 5.00 workers on average. Thus COVER exposes selection headroom without manufacturing a routing win. A crossed-stack diagnostic shows absolute scores depend on G but finds no detectable router-by-finalizer interaction. COVER is an auditable measurement methodology, not a claim of stack-invariant or universal agent-routing superiority.
Chinese Translation
当多智能体系统更换其团队时,它所产生的消息和最终答案也会随之改变,因此端到端的准确性差距本身并不能识别路由效果。我们引入了一种方法,即评估合同,该合同在结果生成之前固定了公共信息边界、下游堆栈 G 和有限的合法团队家族。完全覆盖识别基于该堆栈的精确有限基准预言者遗憾。对于任何有限的冻结策略集合,执行其不同选择团队的并集是每对策略对比的最小假设自由支持,尽管对于绝对预言者遗憾则不然。两个具有源 ID 不重叠分割的控制表测试该工具。在 MuSiQue-12 上,预先指定的特权正控制将遗憾从 0.532 改善到 0.402;而后来的公共接口控制则达到 0.424,相较于 0.554,但属于回顾性。在 HotpotQA-4 上,预先指定的公共直接评分者将遗憾从 0.313 改善到 0.110。在固定堆栈 Llama 执行中,验证的路由遗憾改善了 0.190,而原始答案的增益为 0.010,且区间跨越零。五个家庭的 ToolSandbox 变体转移验证对 14 个未触及的任务变体上的 16 个声明团队进行了全面评估(224/224 有效行):声明家庭的预言者达到了 0.768 的安全证据完成,而前瞻性冻结路由器则获得 0.637(遗憾 0.131),未能满足预先声明的 0.10 标准。后来的回顾性比较者达到了 0.655,平均匹配所有工人 4.57 对 5.00。因此,COVER 揭示了选择的余地,而不制造路由胜利。交叉堆栈诊断表明绝对分数依赖于 G,但未发现可检测的路由器与最终化器之间的交互。COVER 是一种可审计的测量方法,而不是对堆栈不变或普遍智能体路由优越性的主张。
cs.AI / 90 / 2608.28491
AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction
AcrossVAM1.0:用于文本辅助机器人视频预测的粒子世界建模
Abstract
Predicting robot videos requires both precise motion reasoning and preservation of high-frequency appearance, yet monolithic pixel models entangle these objectives and often conceal their progress behind a strong last-frame baseline. We present AcrossVAM1.0, a lightweight, text-assisted video action model that factorizes future prediction into object-centric motion and dense appearance. A frozen SAM3-DLP codec decomposes four context frames into semantic particles for the robot, arm, and gripper, together with a background latent. A 0.28M-parameter spatio-temporal Transformer aligns particle identities, rolls their states forward, and is modulated by a frozen OpenCLIP instruction embedding through FiLM. A causal dual-stream decoder combines particle-rendered motion with appearance encoded exclusively from the last observed frame; a residual refiner and learned delivery mask produce five future frames without access to future appearance. On our VRS benchmark constructed from diverse real-robot trajectories, particle dynamics reduce trajectory error by 21.0\% over persistence. Across three delivery-mask seeds, AcrossVAM1.0 improves future-frame PSNR/SSIM from 19.97/0.796 to 20.573/0.8004, while raw particle generation improves motion-region PSNR from 11.89 to 13.23. The delivered model does not yet beat persistence in LPIPS, and correct-versus- shuffled language changes trajectory error by only 2.8--3.1%. We report these limitations alongside oracle, negative-control, multi-seed, and per-robot analyses. The results show that explicit particle dynamics are a promising low-dimensional interface for robot video prediction, while robust language grounding and appearance delivery remain the principal open challenges.
Chinese Translation
预测机器人视频需要精确的运动推理和高频外观的保留,然而单一的像素模型将这些目标纠缠在一起,常常在强大的最后帧基线背后掩盖其进展。我们提出了AcrossVAM1.0,这是一种轻量级的文本辅助视频动作模型,将未来预测分解为以对象为中心的运动和密集外观。一个冻结的SAM3-DLP编解码器将四个上下文帧分解为机器人的语义粒子、手臂和抓手,以及一个背景潜变量。一个包含0.28M参数的时空Transformer对齐粒子身份,向前滚动它们的状态,并通过FiLM由冻结的OpenCLIP指令嵌入进行调制。一个因果双流解码器将粒子渲染的运动与仅从最后观察到的帧编码的外观结合;一个残差细化器和学习的传递掩码在不访问未来外观的情况下生成五个未来帧。在我们基于多样化真实机器人轨迹构建的VRS基准上,粒子动力学将轨迹误差减少了21.0%。在三个传递掩码种子上,AcrossVAM1.0将未来帧的PSNR/SSIM从19.97/0.796提高到20.573/0.8004,而原始粒子生成将运动区域的PSNR从11.89提高到13.23。交付的模型在LPIPS上尚未超过持久性,而正确与随机语言的比较仅使轨迹误差变化了2.8%至3.1%。我们报告了这些局限性,并进行了oracle、负控制、多种子和每个机器人的分析。结果表明,显式粒子动力学是机器人视频预测的有前景的低维接口,而稳健的语言基础和外观传递仍然是主要的开放挑战。
cs.AI / 91 / 2608.28511
Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration
训练通信效率高的混合专家语言模型与层重配置
Abstract
When training Mixture-of-Experts (MoE) language models with expert parallelism, all-to-all token dispatch and combine collectives can consume a substantial fraction of end-to-end training time. In this work, we study communication-efficient MoE models (CE-MoE), in which we adopt a heterogeneous layer pattern that decouples token-mixing and channel-mixing depth. Compared to conventional models which interleave MoE layers after each token-mixing layer (e.g., attention, Mamba-2), CE-MoE models concentrate expert capacity in a select few routed MoE layers, while maintaining depth by adding additional token-mixing and dense-FFN layers. Across a scaling ladder from 2B to 31.5B total parameters, under matched total and activated parameters, CE-MoE models consistently reduce training cost while matching validation loss and downstream benchmarks with full-MoE baselines. At the 31.5B scale, CE-MoE uses 33.3\% fewer GPU-hours while improving average downstream score and inference throughput.
Chinese Translation
在使用专家并行训练混合专家(MoE)语言模型时,全到全的令牌调度和组合收集操作可能会消耗相当大比例的端到端训练时间。在本研究中,我们研究了通信效率高的MoE模型(CE-MoE),其中采用了一种异构层模式,解耦了令牌混合和通道混合的深度。与传统模型在每个令牌混合层(例如,注意力机制,Mamba-2)后交错MoE层相比,CE-MoE模型将专家容量集中在少数选定的路由MoE层中,同时通过添加额外的令牌混合层和稠密前馈网络(dense-FFN)层来保持深度。在从2B到31.5B总参数的扩展过程中,在匹配的总参数和激活参数下,CE-MoE模型在保持验证损失和下游基准与全MoE基线相匹配的同时,始终降低训练成本。在31.5B规模下,CE-MoE使用的GPU小时数减少了33.3\%,同时提高了平均下游得分和推理吞吐量。
cs.AI / 92 / 2608.28518
When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI
当机器人误听我们的声音:语音控制的具身人工智能安全风险映射
Abstract
We investigate whether automatic speech recognition (ASR) errors in user input can lead to unsafe outputs from Embodied AI (EAI) models. We find that ASR errors can lead to harmful instructions being accepted and executed by EAI models, thereby reducing safety. We simulate ASR errors and combine them with existing safety benchmarks (SafeAgentBench and POEX) to evaluate how different errors affect embodied AI safety. We find that some of them preserve semantic structure but increase harmful ambiguity, while others weaken the model refusal behaviour and allow unsafe plans to be generated and executed. We show that in some cases automatic correction of ASR errors can reduce the risk, but this is not always effective. Overall, we show that ASR errors lead to significant safety risks for embodied AI.
Chinese Translation
我们研究了用户输入中的自动语音识别(ASR)错误是否会导致具身人工智能(EAI)模型产生不安全的输出。我们的研究发现,ASR错误可能导致有害指令被EAI模型接受并执行,从而降低安全性。我们模拟了ASR错误,并将其与现有的安全基准(SafeAgentBench和POEX)相结合,以评估不同错误对具身人工智能安全性的影响。我们发现,有些错误保留了语义结构,但增加了有害的模糊性,而另一些则削弱了模型的拒绝行为,允许生成和执行不安全的计划。我们表明,在某些情况下,自动纠正ASR错误可以降低风险,但这并不总是有效。总体而言,我们的研究表明,ASR错误对具身人工智能构成了显著的安全风险。
cs.AI / 93 / 2608.28534
InstructMesh: Selective Refinement of Generative 3D Models for Fabrication
InstructMesh:用于制造的生成3D模型的选择性细化
Abstract
Recent advances in generative AI allow users to create 3D models from text or images. However, these models prioritize visual plausibility over geometric accuracy, often generating results with flaws that compromise their intended use post-fabrication. We present InstructMesh, an interactive post-generation refinement tool that enables selective repair of generative 3D models through region selection and targeted operations, such as opening or sealing voids, or adjusting local thickness. Users can invoke edit operations via natural language prompts or slider controls. By operating directly on the intermediate latent representation, InstructMesh allows users to apply robust geometric corrections without requiring expert modeling skills. To inform our design, we first analyze common fabrication-related failure modes in outputs from state-of-the-art generative tools. We then conduct two user studies, demonstrating that novices can identify and perform fabrication-relevant repairs on generative outputs using InstructMesh, and revealing user preference for hybrid interfaces that combine slider controls with natural language input.
Chinese Translation
最近,生成性人工智能的进展使用户能够从文本或图像创建3D模型。然而,这些模型优先考虑视觉合理性而非几何精确性,常常生成存在缺陷的结果,影响其在制造后的预期使用。我们提出了InstructMesh,这是一种交互式后生成细化工具,能够通过区域选择和针对性操作(如打开或封闭空洞,或调整局部厚度)实现生成3D模型的选择性修复。用户可以通过自然语言提示或滑块控制调用编辑操作。通过直接操作中间潜在表示,InstructMesh使用户能够在不需要专业建模技能的情况下应用稳健的几何修正。为了指导我们的设计,我们首先分析了来自最先进生成工具的输出中与制造相关的常见故障模式。然后,我们进行了两项用户研究,证明新手能够使用InstructMesh识别并执行与制造相关的修复,并揭示用户偏好将滑块控制与自然语言输入相结合的混合界面。
cs.AI / 94 / 2608.28553
Logos: An Agent Harness on a Cross-Process Bus
Logos:一种跨进程总线的代理工具
Abstract
Modern agent systems assemble capabilities at runtime, and this dynamic composition has recently received a complete formal treat ment in the spatiotemporal-composability calculus, in which a capability is a component carrying a tracked inverse, and agents are assembled as plugins. This plugin form is carried by a single process sharing one context, a carrier that places all components in one physical failure domain, a fault suspends every component at once, and process death interrupts every session the process hosts. This paper shows that neither the modeling nor the calculus binds an agent to one process, the statelessness of the language model keeps all cross-step state outside the model, and the soundness invariant is defined on the state space alone. These observations condense into four lemmas whose premises are the hypotheses of the calculus and the statelessness of language-model inference. On these lemmas this paper constructs Logos, a ROS-like cross process agent harness in which a plugin is a process and the only shared state is an append-only transcript. Eighty sessions resume with no repeated effect after kills placed at the four boundaries of the tool-call cycle, and a same-fault comparison with a single process reference configuration shows one fault interrupting every co-resident session while under the peer-process construction one fault ends at one node.
Chinese Translation
现代代理系统在运行时组装能力,这种动态组合最近在时空可组合性演算中得到了完整的形式化处理,其中能力是携带跟踪逆的组件,代理则作为插件进行组装。这种插件形式由一个共享单一上下文的进程承载,承载者将所有组件置于一个物理故障域中,故障会同时暂停每个组件,而进程的死亡会中断该进程所承载的每个会话。本文表明,无论是建模还是演算,都并未将代理绑定到一个进程,语言模型的无状态性将所有跨步骤状态保持在模型之外,而健全性不变式仅在状态空间上定义。这些观察凝聚成四个引理,其前提是演算的假设和语言模型推理的无状态性。基于这些引理,本文构建了Logos,一种类似ROS的跨进程代理工具,其中插件是一个进程,唯一共享的状态是一个仅追加的记录。在工具调用周期的四个边界处进行的杀死操作后,八十个会话在没有重复效果的情况下恢复,而与单进程参考配置的同故障比较显示,一个故障会中断每个共驻会话,而在对等进程构造下,一个故障仅在一个节点结束。
cs.CL / 1 / 2608.27460
Accelerating LLM Inference via Vector Index Based Output Embeddings
通过基于向量索引的输出嵌入加速大规模语言模型推理
Abstract
Large output embedding matrices create a significant memory bandwidth bottleneck during autoregressive decoding, especially for compact LLMs with large multilingual vocabularies. We reformulate the output projection followed by top-k token selection as a maximum inner product search over token embeddings and replace the dense vocabulary projection with an HNSW-based vector index. The resulting output head retrieves only a small candidate set of high-scoring tokens and can be integrated into existing decoding pipelines by scattering retrieved logits into a sparse full-vocabulary tensor. On CPU inference with Gemma 3, Llama 3.2, and Qwen 3 models, our method substantially accelerates the output projection and improves end-to-end batch-size-one decoding throughput by up to 82% for Gemma 3 270M, while preserving generation quality under AlpacaEval evaluation. These results suggest approximate retrieval is a practical alternative to dense output projections in latency-sensitive small-batch decoding.
Chinese Translation
大型输出嵌入矩阵在自回归解码过程中造成了显著的内存带宽瓶颈,尤其是对于具有大型多语言词汇的紧凑型大规模语言模型(LLMs)。我们将输出投影和前k个高分数标记的选择重新表述为对标记嵌入的最大内积搜索,并用基于HNSW(Hierarchical Navigable Small World)的向量索引替代了密集的词汇投影。所得到的输出头仅检索一小部分高分数标记的候选集,并可以通过将检索到的logits散布到稀疏的全词汇张量中,集成到现有的解码管道中。在使用Gemma 3、Llama 3.2和Qwen 3模型的CPU推理中,我们的方法显著加速了输出投影,并将Gemma 3 270M的端到端批量大小为1的解码吞吐量提高了多达82%,同时在AlpacaEval评估下保持了生成质量。这些结果表明,近似检索是延迟敏感的小批量解码中密集输出投影的一个实用替代方案。
cs.CL / 2 / 2608.27461
SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction
SciReC:适应性交互下的多模态、多轮关系推理的诊断评估
Abstract
Relational reasoning requires the process of perceptual understanding, comparing, and integrating the underlying relationships between concepts. This ability consists of multiple categories, such as analogical, structural, and cause-effect, each capturing a different aspect of higher-order understanding. To examine the performance of multimodal large language models (MLLM) on these relational inference tasks, we developed SciReC, a model-adaptive multimodal academic dialog benchmark. As the relational reasoning process involves multiple representations and various factors (visual understanding, exhibiting knowledge, and memory recall), we propose DMRA, a deficit-based diagnostic framework that quantifies the contribution of these components to identify the primary cause of unsuccessful cases. Claude 4.6 achieved the best performance on the overall relational score with 73\%, followed by GPT 5.4 with 68\%. Performance trends indicate that open-source models achieve their lowest scores on spatial relations, while proprietary models struggle more with hierarchical and sequential relations. Across domains, model performance is lowest on Astronomy and highest on Psychology. The results of DMRA reveal that relational reasoning is the primary source of error across all models, followed by memory limitations.
Chinese Translation
关系推理需要对概念之间的潜在关系进行感知理解、比较和整合的过程。这种能力由多个类别组成,如类比、结构和因果关系,每个类别捕捉高阶理解的不同方面。为了检验多模态大型语言模型(MLLM)在这些关系推理任务上的表现,我们开发了SciReC,一个模型自适应的多模态学术对话基准。由于关系推理过程涉及多种表现形式和各种因素(视觉理解、知识展示和记忆回忆),我们提出了DMRA,一个基于缺陷的诊断框架,用于量化这些组件的贡献,以识别不成功案例的主要原因。Claude 4.6在整体关系得分上取得了73%的最佳表现,其次是GPT 5.4,得分为68%。性能趋势表明,开源模型在空间关系上的得分最低,而专有模型在层次和顺序关系上表现更差。在各个领域中,模型在天文学上的表现最低,而在心理学上的表现最高。DMRA的结果显示,关系推理是所有模型错误的主要来源,其次是记忆限制。
cs.CL / 3 / 2608.27462
Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech
大锤还是手术刀?一种细粒度自适应隐性仇恨言论检测框架
Abstract
Unlike explicit attacks with obvious profanity, implicit hate speech hides malice within seemingly compliant expressions through metaphors and contextual hints, making its detection in online content review challenging. While existing PLM- or LLM-based methods perform well, they typically apply a single reasoning process to all samples. This overlooks fine-grained linguistic nuances and causes unnecessary computation for simpler cases. We observe that online hate speech is not monolithic but manifests in varied forms. We therefore define three fine-grained categories: Shallow, Targeted, and Context-Dependent. Accordingly, we propose Fine-grained Adaptive Implicit Hate speech Detection (FAID), a novel framework that first performs fine-grained classification and then adapts to specific categories. Specifically, for Shallow samples with surface-identifiable intents, the framework adopts lightweight prompt-tuning for rapid classification; for Targeted comments that bind malicious intent to concealed targets, we design knowledge augmentation to iteratively refine the model and reveal hidden targets; for Context-Dependent comments lacking background information, we utilize an agentic framework that automatically generates prompts to evolve context, infer missing background information and identify ambiguous malicious intents. This adaptive architecture focuses computational resources on complex implicit samples while avoiding redundant reasoning for shallow samples. Experiments on four benchmark datasets demonstrate that FAID significantly outperforms SOTA baselines.
Chinese Translation
与明显的粗俗攻击不同,隐性仇恨言论通过隐喻和上下文暗示在表面上看似合规的表达中隐藏恶意,这使得在在线内容审查中检测其变得具有挑战性。虽然现有的基于预训练语言模型(PLM)或大型语言模型(LLM)的方法表现良好,但它们通常对所有样本应用单一的推理过程。这忽视了细粒度的语言细微差别,并导致对简单案例的不必要计算。我们观察到在线仇恨言论并非单一,而是以多种形式表现。因此,我们定义了三个细粒度类别:浅层(Shallow)、针对性(Targeted)和上下文依赖(Context-Dependent)。因此,我们提出了细粒度自适应隐性仇恨言论检测(Fine-grained Adaptive Implicit Hate speech Detection,FAID),这是一个新颖的框架,首先进行细粒度分类,然后适应特定类别。具体而言,对于表面可识别意图的浅层样本,该框架采用轻量级提示调优(prompt-tuning)进行快速分类;对于将恶意意图绑定到隐蔽目标的针对性评论,我们设计了知识增强(knowledge augmentation)以迭代地优化模型并揭示隐藏目标;对于缺乏背景信息的上下文依赖评论,我们利用一个自主框架(agentic framework)自动生成提示以演变上下文,推断缺失的背景信息并识别模糊的恶意意图。该自适应架构将计算资源集中于复杂的隐性样本,同时避免对浅层样本进行冗余推理。在四个基准数据集上的实验表明,FAID显著优于现有的最先进基线(SOTA baselines)。
cs.CL / 4 / 2608.27465
The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models
情感背景对大型语言模型支持过早决策的影响:比较六个商业模型的情感脆弱性
Abstract
As large language models (LLMs) are increasingly used for everyday decision-making advice, whether a model shifts the direction of its advice according to the user's emotional state has become an important safety problem. We test whether emotional expression increases a model's endorsement (encouragement to proceed) when a user, holding the same objective information, is overconfident about a premature decision (e.g., quitting a stable job on weak evidence). As a key control, we include a no-emotion multi-turn (neutral) condition that holds factual content and the number of conversational turns constant, isolating the effect of emotion from that of conversation length. We exposed six commercial models (top-tier and mid-tier models from OpenAI, Anthropic, and Google) to three scenarios (career change, business expansion, emigration) across three conditions (cold/neutral/distress) with six repetitions each, yielding 324 conversations, and measured endorsement strength (0-100) via an eight-item rubric-based automated scoring. Emotional expression significantly increased endorsement (neutral 18.6 to distress 31.5, +12.9 points; mixed-effects $\beta = +12.9$, $p < .001$; Cohen's d = 0.51), and this was not explained by conversation length (cold-neutral difference non-significant, $p = .083$). Critically, the vulnerability varied by individual model rather than by price tier: five of six models showed a significant emotion effect, including the top-tier flagships Gemini 3.1 Pro and GPT-5.5, while only Claude Opus showed no significant change. Results were reproduced with an independent non-Google judge model ($\rho = .89$) and agreed in rank with two human coders ($\rho = .70$). Through a controlled design that separates emotion from conversational context, we show that emotional context increases LLM sycophancy even in top-tier flagship models.
Chinese Translation
随着大型语言模型(LLMs)在日常决策建议中的使用日益增多,模型是否会根据用户的情感状态改变其建议方向已成为一个重要的安全问题。我们测试了当用户在相同客观信息下对过早决策(例如,在薄弱证据下辞去稳定工作)过于自信时,情感表达是否会增加模型的支持(鼓励继续进行)。作为一个关键的控制,我们包含了一个无情感的多轮(中性)条件,保持事实内容和对话轮次不变,从而将情感的影响与对话长度的影响隔离开来。我们将六个商业模型(来自OpenAI、Anthropic和Google的顶级和中级模型)暴露于三个场景(职业变动、商业扩展、移民)下的三个条件(冷/中性/痛苦),每种情况重复六次,共产生324次对话,并通过基于八项指标的自动评分测量支持强度(0-100)。情感表达显著增加了支持度(中性18.6到痛苦31.5,+12.9分;混合效应 $eta = +12.9$, $p < .001$; Cohen's d = 0.51),且这一结果并未被对话长度解释(冷-中性差异不显著,$p = .083$)。关键是,脆弱性因个体模型而异,而非价格层级:六个模型中有五个显示出显著的情感效应,包括顶级旗舰模型Gemini 3.1 Pro和GPT-5.5,而只有Claude Opus未显示显著变化。结果在一个独立的非Google评判模型中得到了重复($
ho = .89$),并与两位人类编码者的排名一致($
ho = .70$)。通过一个将情感与对话背景分开的控制设计,我们展示了情感背景即使在顶级旗舰模型中也会增加LLM的阿谀奉承行为。
cs.CL / 5 / 2608.27466
PACE: Publisher-Adaptive Content Extraction via Agentic Automation
PACE:通过自主自动化实现出版商自适应内容提取
Abstract
Web content extraction is essential for reliable LLM data pipelines, yet existing methods often struggle to jointly satisfy accuracy, scalability, and adaptability. General-purpose extractors can be applied broadly, but they are often brittle on publisher-specific layouts and richer extraction targets such as metadata, images, and tables. Direct LLM-based extraction offers greater flexibility, but incurs substantial cost and latency at scale, while manually engineered publisher-specific parsers can achieve high accuracy but require substantial human effort to build and maintain. We introduce PACE, an agentic framework for learning publisher-specific extraction configurations from representative pages and user requirements. During training, PACE uses LLMs to analyze page structure and aggregate reusable extraction patterns. At inference time, the learned configurations instantiate a fixed deterministic extractor template, enabling scalable extraction without additional LLM calls. Experiments spanning article-body, metadata, and multimodal extraction show that PACE outperforms scalable non-manual baselines while approaching the quality of manually engineered publisher-specific parsers. PACE achieves stronger extraction of article text, metadata, images, and tables, demonstrating that agentic configuration learning can automate publisher-specific extraction for LLM-ready page representations beyond article text.
Chinese Translation
网络内容提取对于可靠的大型语言模型(LLM)数据管道至关重要,但现有方法往往难以同时满足准确性、可扩展性和适应性。通用提取器可以广泛应用,但在特定出版商的布局和更复杂的提取目标(如元数据、图像和表格)上往往表现脆弱。基于LLM的直接提取提供了更大的灵活性,但在大规模应用时会产生显著的成本和延迟,而手动设计的特定出版商解析器可以实现高准确性,但需要大量人力来构建和维护。我们提出了PACE,一个用于从代表性页面和用户需求中学习特定出版商提取配置的自主框架。在训练过程中,PACE利用LLM分析页面结构并汇总可重用的提取模式。在推理时,学习到的配置实例化一个固定的确定性提取器模板,从而实现可扩展的提取而无需额外的LLM调用。跨越文章正文、元数据和多模态提取的实验表明,PACE在可扩展的非手动基准上表现优越,同时接近手动设计的特定出版商解析器的质量。PACE在文章文本、元数据、图像和表格的提取上表现更强,证明了自主配置学习能够自动化特定出版商的提取,以适应LLM准备的页面表示,超越文章文本的提取。
cs.CL / 6 / 2608.27467
UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering
UIC-AIHealth4All在ArchEHR-QA 2026:以答案为先的临床问题回答证据基础
Abstract
We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records. We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment). For Subtasks 2 and 3, we propose an answer-first pipeline in which the model generates candidate answers citing specific note sentences before classifying the full evidence set, exploiting the asymmetry between judging relevance in the abstract versus relative to a generated answer. For Subtask 4, we apply self-consistency voting over five independent model calls, retaining links by vote threshold. Our pipeline ranked third on evidence identification (Strict Micro F1 62.90), ninth on answer generation (Overall 31.90), and fifth on answer-evidence alignment (F1 79.81). A post-hoc linguistic analysis of 45 stylistic features reveals that model outputs remain 3.2 Flesch-Kincaid grade levels harder to read than clinician-authored references despite matching their word and sentence counts, suggesting readability warrants explicit optimization in clinical NLP systems. Code and prompts are available at https://github.com/mo-arvan/archehr-qa-2026-uic-aihealth4all.
Chinese Translation
我们描述了UIC-AIHealth4All系统,该系统参与了ArchEHR-QA 2026,这是一个基于电子健康记录的有据问题回答共享任务。我们参与了子任务2(证据识别)、子任务3(答案生成)和子任务4(答案-证据对齐)。对于子任务2和3,我们提出了一种以答案为先的流程,其中模型在对完整证据集进行分类之前,生成引用特定笔记句子的候选答案,利用了在抽象与相对于生成答案的相关性判断之间的非对称性。对于子任务4,我们在五个独立模型调用上应用自一致性投票,通过投票阈值保留链接。我们的流程在证据识别上排名第三(严格微F1 62.90),在答案生成上排名第九(总体31.90),在答案-证据对齐上排名第五(F1 79.81)。对45个风格特征的事后语言分析显示,尽管模型输出的单词和句子数量与临床医生撰写的参考文献相匹配,但其可读性仍比临床医生的参考文献难以阅读3.2个Flesch-Kincaid年级水平,这表明可读性在临床自然语言处理系统中需要明确优化。代码和提示可在https://github.com/mo-arvan/archehr-qa-2026-uic-aihealth4all获取。
cs.CL / 7 / 2608.27470
Select, Don't Train: The Benefits of Modular Entity Disambiguation with LLM-Based Selection
选择,而非训练:基于大型语言模型选择的模块化实体消歧义的优势
Abstract
Entity Disambiguation (ED) is a key task for constructing and using knowledge graphs. State-of-the-art neural approaches commonly model ED as a single task, although it consists of two distinct subproblems: retrieving candidate entities and selecting the correct one given context. Dual-encoder models optimize for both within a shared embedding space, forcing representations to balance high-recall retrieval with fine-grained selection, and they require trained retrievers, which are costly to maintain as knowledge graphs change. While recent work has begun to combine retrievers with LLM-based selectors, the interplay between the two stages has not been studied systematically. In this paper, we present a systematic comparison of retrieval strategies for candidate generation under a shared LLM-based selection stage, combining sparse retrieval (BM25), Web KB search, and a state-of-the-art trained dense retriever with several open- and closed-source LLMs. We show that, once selection is delegated to a capable LLM, training the retriever provides only modest additional value: a fully training-free BM25 retriever paired with an LLM selector reaches a new state of the art on the ZELDA benchmark, raising inKB micro-F1 from 82.3 to 86.3 (+4); pairing the same LLM with a trained dense retriever reaches 88.5. Decoupling retrieval from selection also exposes a limitation of current ED systems: when the correct entity is missing from retrieved candidates, they are forced to predict an incorrect entity. In contrast, our framework allows for abstention when retrieval failure is detected. In an evaluation setting that rewards correct abstentions, the training-free BM25 + LLM pipeline reaches 90.7 F1.
Chinese Translation
实体消歧义(ED)是构建和使用知识图谱的关键任务。最先进的神经网络方法通常将ED建模为单一任务,尽管它由两个不同的子问题组成:检索候选实体和根据上下文选择正确的实体。双编码器模型在共享的嵌入空间中优化这两者,迫使表示在高召回率检索与细粒度选择之间取得平衡,并且它们需要训练好的检索器,随着知识图谱的变化,这种维护成本较高。尽管近期的研究已开始将检索器与基于大型语言模型(LLM)的选择器结合,但这两个阶段之间的相互作用尚未得到系统研究。在本文中,我们对在共享LLM选择阶段下候选生成的检索策略进行了系统比较,结合了稀疏检索(BM25)、网络知识库搜索以及与多个开源和闭源LLM结合的最先进训练密集检索器。我们表明,一旦选择委托给一个有能力的LLM,训练检索器提供的额外价值仅为适度:一个完全无训练的BM25检索器与LLM选择器配对,在ZELDA基准上达到了新的最先进水平,将inKB微F1从82.3提升至86.3(+4);将相同的LLM与训练密集检索器配对则达到了88.5。将检索与选择解耦还暴露了当前ED系统的一个局限性:当正确的实体缺失于检索候选中时,它们被迫预测一个错误的实体。相比之下,我们的框架在检测到检索失败时允许放弃。在一个奖励正确放弃的评估环境中,无训练的BM25 + LLM管道达到了90.7的F1值。
cs.CL / 8 / 2608.27481
XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering
XHotpotQA:跨语言知识组合的多跳问答基准
Abstract
Knowledge-intensive multi-hop question answering requires systems to select evidence and compose dependent facts, yet multilingual benchmarks usually translate an entire example into one language. This hides failures at language boundaries inside the reasoning chain. We introduce XHotpotQA, a controlled benchmark for cross-lingual knowledge composition over mixed-language evidence. Each instance is modeled as an evidence-dependency graph whose question, bridge evidence, answer-bearing evidence, and distractors have explicit language assignments. The audited resource contains 15,661 training and 7,405 validation instances, with sentence-level support supervision and supplied distractors. In validation, 99.81% of items cross the question-to-gold-evidence language interface and 95.60% use gold paragraphs in different languages. Across three reader artifacts, full question-evidence mismatch is associated with 10.25 to 15.79 lower Unicode-aware answer F1 than partial alignment, and different-script evidence with deficits of 11.98 to 23.70 points; the corresponding adapted-selector contrasts are 1.71 and 1.78 points. Under this supplied-candidate design, the evaluated readers therefore show substantially larger condition-associated deficits than the selector. XHotpotQA provides role-aware diagnostics, modular evaluation, and an audited test bed for knowledge-based systems that must integrate evidence across languages.
Chinese Translation
知识密集型的多跳问答要求系统选择证据并组合依赖事实,但多语言基准通常将整个示例翻译成一种语言。这掩盖了推理链中语言边界的失败。我们引入了XHotpotQA,这是一个针对混合语言证据的跨语言知识组合的受控基准。每个实例被建模为一个证据依赖图,其中的问题、桥接证据、答案承载证据和干扰项都有明确的语言分配。审计资源包含15,661个训练实例和7,405个验证实例,具有句子级支持监督和提供的干扰项。在验证中,99.81%的项目跨越了问题与黄金证据的语言接口,95.60%使用不同语言的黄金段落。在三个阅读器工件中,完整问题-证据不匹配与部分对齐相比,Unicode感知答案F1降低了10.25到15.79点,而不同脚本证据的缺陷为11.98到23.70点;相应的适应选择器对比为1.71和1.78点。在这种提供候选的设计下,评估的阅读器因此显示出比选择器更大的条件相关缺陷。XHotpotQA提供了角色感知的诊断、模块化评估,以及一个审计的测试平台,供必须跨语言整合证据的知识基础系统使用。
cs.CL / 9 / 2608.27501
INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning
INSPIRE:一种内化后改进的示例驱动数学推理方法
Abstract
Mathematical reasoning has seen rapid progress in large language models (LLMs), yet existing methods optimize predominantly for final-answer correctness, raising the question whether models truly internalize mathematical concepts or merely memorize solution patterns. In human mathematics education, example-based reasoning such as constructing counterexamples to test theorem boundaries reflects deep conceptual understanding, but remains underdeveloped in current LLMs. Enhancing this capability through preference optimization presents two key challenges: (1) the model's limited example-based reasoning ability makes constructing effective preference pairs inherently difficult; and (2) capability acquisition is progressive, as the model must first learn to adopt this strategy before learning to apply it correctly. Therefore we propose INSPIRE, an Internalize-Then-Improve approach combining Reference-Guided Student Internalization (RGSI), which produces high-quality preference candidates under the policy model's own distribution, with a stage-wise rubric preference training strategy that decomposes learning into method-oriented and correctness-oriented stages. Experiments across multiple model scales and families demonstrate consistent improvements, even surpassing larger open-source models, while evaluations on out-of-distribution benchmarks confirm no degradation in general mathematical reasoning ability.
Chinese Translation
数学推理在大型语言模型(LLMs)中取得了快速进展,但现有方法主要优化最终答案的正确性,这引发了一个问题:模型是否真正内化了数学概念,还是仅仅记忆了解决模式。在人类数学教育中,基于示例的推理,例如构造反例以测试定理的边界,反映了深刻的概念理解,但在当前的LLMs中仍然发展不足。通过偏好优化增强这一能力面临两个关键挑战:(1)模型有限的基于示例的推理能力使得构造有效的偏好对 inherently 变得困难;(2)能力获取是渐进的,因为模型必须首先学习采用这一策略,然后才能学习正确应用它。因此,我们提出了INSPIRE,一种内化后改进的方法,结合了参考引导的学生内化(Reference-Guided Student Internalization, RGSI),该方法在策略模型自身分布下生成高质量的偏好候选,以及一种分阶段的评分偏好训练策略,将学习分解为以方法为导向和以正确性为导向的阶段。跨多个模型规模和家族的实验表明了一致的改进,甚至超越了更大的开源模型,而在分布外基准上的评估则确认了数学推理能力没有下降。
cs.CL / 10 / 2608.27505
A Survey on Rubric-Guided Reinforcement Learning for Language Models
关于基于评分标准的强化学习在语言模型中的应用调查
Abstract
Reinforcement learning from human feedback (RLHF) has become the dominant paradigm for aligning large language models (LLMs) with human preferences. However, traditional RLHF relies on scalar reward signals that lack interpretability and fail to capture the multifaceted nature of response quality. Rubric-guided reinforcement learning addresses these limitations by introducing structured, interpretable evaluation criteria, or rubrics, as the backbone of reward design, feedback generation, and policy optimization. In this survey, we introduce a Bayesian framework that defines constitutions as prior distributions $P(R)$ over evaluation criteria and rubrics as conditional instantiations $R_x \sim P(R|x)$. Under this unified view, we present a taxonomy of rubric-guided RL along the prior-posterior axis, covering constitutional AI, instance-specific rubrics, process-level supervision, self-evolving rubrics, and their agentic and multimodal extensions. Furthermore, as rubrics are natural-language artifacts, we present a linguistic analysis of how granularity trade-offs, semantic drift, and linguistic reward hacking impact alignment reliability, identifying key open problems for future research.
Chinese Translation
来自人类反馈的强化学习(RLHF)已成为将大型语言模型(LLMs)与人类偏好对齐的主导范式。然而,传统的RLHF依赖于缺乏可解释性的标量奖励信号,未能捕捉响应质量的多面性。基于评分标准的强化学习通过引入结构化、可解释的评估标准(或评分标准)作为奖励设计、反馈生成和策略优化的基础,从而解决了这些局限性。在本调查中,我们引入了一个贝叶斯框架,将宪法定义为评估标准的先验分布 $P(R)$,而评分标准则作为条件实例化 $R_x ext{~} P(R|x)$。在这一统一视角下,我们沿着先验-后验轴呈现了基于评分标准的强化学习的分类法,涵盖了宪法人工智能、特定实例评分标准、过程级监督、自我演化评分标准及其智能和多模态扩展。此外,由于评分标准是自然语言的产物,我们还进行了语言学分析,探讨了粒度权衡、语义漂移和语言奖励黑客如何影响对齐的可靠性,并识别出未来研究的关键开放问题。
cs.CL / 11 / 2608.27510
How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution
线性探针是如何产生的?一个具有概念目标归因的电路追踪框架
Abstract
Transcoder attribution graphs are usually trained to explain why a model assigns high probability to a particular next token. We introduce Concept-Targeted Attribution (CTA), which instead trains attribution graphs with respect to a linear probe direction. CTA therefore yields probe-specific circuits that explain why an internal concept representation arises in a prompt, independently of whether it is expressed in the generated token. Using Cross-Layer Transcoders, we show that these probe-targeted graphs contain predictive structure: graph-level features predict probe accuracy across four widely studied concept categories ($\rho = 0.91$, $R^2 = 0.84$), while local features identify the sparse components driving per-prompt classification. This connects probe performance to interpretable circuit structure, allowing us to ask not only whether a probe works, but which internal computations make it work. Causal ablations further show that probe-targeted and logit-targeted graphs capture functionally distinct mechanisms. Removing probe-relevant features reduces internal concept scores while largely preserving generated tokens, whereas removing logit-relevant features changes the generated token in 92% to 100% of cases with near-zero effect on probe scores. CTA provides a framework for moving from behavioral probe accuracy to mechanistic explanations of probe performance, enabling more detailed audits of internal concept representations, including safety-critical ones. Our code is available at https://github.com/vedantpalit/concept-targeted-attribution
Chinese Translation
转码器归因图通常用于解释模型为何对特定下一个标记分配高概率。我们引入了概念目标归因(Concept-Targeted Attribution, CTA),它以线性探针方向为基础训练归因图。因此,CTA 生成了特定于探针的电路,解释了为何内部概念表示在提示中出现,而不管它是否在生成的标记中表达。通过使用跨层转码器,我们展示了这些针对探针的图包含预测结构:图级特征预测四个广泛研究的概念类别中的探针准确性($
ho = 0.91$, $R^2 = 0.84$),而局部特征则识别出驱动每个提示分类的稀疏成分。这将探针性能与可解释的电路结构联系起来,使我们不仅可以询问探针是否有效,还可以探讨哪些内部计算使其有效。因果消融进一步表明,针对探针和针对逻辑回归的图捕捉了功能上不同的机制。去除与探针相关的特征会降低内部概念分数,同时在很大程度上保留生成的标记,而去除与逻辑回归相关的特征则在92%到100%的情况下改变生成的标记,对探针分数几乎没有影响。CTA 提供了一个从行为探针准确性到探针性能机制解释的框架,使我们能够更详细地审计内部概念表示,包括安全关键的表示。我们的代码可在 https://github.com/vedantpalit/concept-targeted-attribution 获取。
cs.CL / 12 / 2608.27514
Trajectory-Level Speculative Decoding for Diffusion Language Models
基于轨迹的推测解码用于扩散语言模型
Abstract
Diffusion-based language models (dLLMs) enable parallel token generation through iterative denoising, but existing decoding strategies collapse to single-token generation under low confidence, severely limiting throughput. Unlike autoregressive models where speculative decoding operates on token sequences in a fixed left-to-right order, dLLMs require speculating over denoising trajectories-sequences of multi-token updates with explicit positions and unmasking orders. We develop a trajectory-level speculative framework that constructs draft denoising trajectories via confidence-stratified tree exploration and verifies them through blockwise parallel evaluation with bidirectional attention masking. Our method further introduces inter-block speculation, exploiting diffusion models' bidirectional structure to perform cross-block lookahead. We formally characterize when this approach is exact and identify trajectory drift as the fundamental cost of increased parallelism. Building on Fast-dLLM's dual-cache infrastructure, our framework reduces denoising iterations by 30-40% and increases tokens-per-step from 2.6 to 4.3, achieving 7-14x speedup over vanilla dLLMs and 1.3x over Fast-dLLM with less than 1% accuracy change across reasoning and code benchmarks.
Chinese Translation
基于扩散的语言模型(dLLMs)通过迭代去噪实现并行标记生成,但现有的解码策略在低置信度下会崩溃为单标记生成,严重限制了吞吐量。与自回归模型中推测解码在固定的从左到右的顺序上操作标记序列不同,dLLMs 需要在去噪轨迹上进行推测——即具有明确位置和解掩蔽顺序的多标记更新序列。我们开发了一种基于轨迹的推测框架,通过置信度分层的树探索构建草拟去噪轨迹,并通过块级并行评估与双向注意力掩蔽进行验证。我们的方法进一步引入了块间推测,利用扩散模型的双向结构执行跨块前瞻。我们正式表征了何时这种方法是精确的,并识别轨迹漂移作为增加并行性的基本成本。基于 Fast-dLLM 的双缓存基础设施,我们的框架将去噪迭代减少了 30-40%,并将每步生成的标记数从 2.6 增加到 4.3,实现了相较于普通 dLLMs 的 7-14 倍加速,以及相较于 Fast-dLLM 的 1.3 倍加速,同时在推理和代码基准测试中准确率变化小于 1%。
cs.CL / 13 / 2608.27658
When Tokenizers Fail: Byte-Level Chunking for Zero-Shot Transfer to Low-Resource Languages
当分词器失效时:针对低资源语言的零样本迁移的字节级分块
Abstract
Subword tokenization hinders low-resource language processing by imposing frequency patterns from dominant languages onto script-sharing variants. Byte-level models bypass this issue by processing raw UTF-8 characters, yet they create a granularity mismatch for word-level tasks in non-Latin scripts. Hierarchical byte-level architectures address this mismatch by grouping bytes into word-aligned chunks. However, these architectures require massive training data and suffer from representational misalignment when paired with frozen subword-based language models. In this paper, we propose an adapted hierarchical network framework that bridges this modality gap without extensive training. Our method initializes byte embeddings directly from the subword representations of a frozen base model. We apply a chunk alignment loss to project dynamically grouped byte chunks toward precomputed subword targets, and interleave lightweight part-of-speech (POS) supervision to guide boundary detection. Experiments across six languages demonstrate that our tokenizer-free approach improves performance for word-level morphological tasks, yielding up to a 13.3% improvement on POS tagging.
Chinese Translation
子词分词在低资源语言处理上存在障碍,因为它将主导语言的频率模式强加于共享脚本的变体。字节级模型通过处理原始的 UTF-8 字符来绕过这个问题,但在非拉丁脚本的词级任务中,它们会造成粒度不匹配。层次化字节级架构通过将字节分组为与单词对齐的块来解决这一不匹配。然而,这些架构需要大量的训练数据,并且在与冻结的基于子词的语言模型配对时会遭遇表示不对齐的问题。在本文中,我们提出了一种改进的层次网络框架,能够在不进行广泛训练的情况下弥合这种模态差距。我们的方法直接从冻结基模型的子词表示初始化字节嵌入。我们应用块对齐损失,将动态分组的字节块投影到预计算的子词目标上,并交错轻量级的词性(POS)监督以指导边界检测。对六种语言的实验表明,我们的无分词器方法在词级形态学任务中提高了性能,在词性标注上最高可提高 13.3%。
cs.CL / 14 / 2608.27661
Knowing Before Answering: Decoding Language Models for Reliable RAG
回答前的认知:解码语言模型以实现可靠的检索增强生成
Abstract
In Retrieval-Augmented Generation (RAG), retrieval may provide insufficient or conflicting information needed to answer a question. The system should not only know when to answer but also be able to identify cases in which the documents provided in RAG are insufficient or contain conflicting information. This can be framed as a three-way classification problem, where we use the model's internal signals to determine whether the provided information in the input can be classified as sufficient, insufficient, or conflicting. We create a controlled benchmark dataset that replicates a RAG setup with fictitious information and labels each instance as answerable, insufficient, or conflicting. We use hidden activations and attention-derived features as inputs to train a lightweight linear model to distinguish among the three classes. Across 16 language models spanning different architectures and a range of model sizes, our feature-based router consistently outperforms prompting-based baselines and the performance of specialised RAG-models. We further conduct analyses into the information dynamics of the models. We show that the most informative signals for the classification are available in the middle layers, with hidden activation states being more effective than attention values or the MLP-feature outputs in most of the tested models. Overall, our results suggest that language models internally encode whether retrieved evidence is sufficient to support answering, and that this signal can be decoded reliably for RAG triage.
Chinese Translation
在检索增强生成(RAG)中,检索可能提供不足或相互矛盾的信息来回答问题。系统不仅需要知道何时回答,还需要能够识别在 RAG 中提供的文档是否不足或包含矛盾信息。这可以被框定为一个三类分类问题,我们利用模型的内部信号来判断输入中提供的信息是否可以被分类为充分、不充分或矛盾。我们创建了一个受控的基准数据集,复制了一个包含虚构信息的 RAG 设置,并将每个实例标记为可回答、不充分或矛盾。我们使用隐藏激活和基于注意力的特征作为输入,训练一个轻量级线性模型以区分这三类。在涵盖不同架构和一系列模型大小的 16 个语言模型中,我们的基于特征的路由器始终优于基于提示的基线和专门 RAG 模型的性能。我们进一步分析了模型的信息动态。我们展示了分类的最有信息量的信号在中间层中可用,隐藏激活状态在大多数测试模型中比注意力值或 MLP 特征输出更有效。总体而言,我们的结果表明,语言模型在内部编码了检索证据是否足以支持回答,并且这一信号可以可靠地解码用于 RAG 分流。
cs.CL / 15 / 2608.27672
First Make It Playable, Then Make It Good: Staged Interaction Learning for Small Dialogue-Game Agents
先让它可玩,再让它优秀:小型对话游戏代理的分阶段交互学习
Abstract
We present Qwen-GuidePlay-2B, a 2B-parameter language model for dialogue-game interaction. We fine-tune Qwen3.5-2B using three steps: a) SFT on only successful game trajectories from Playpen, b) weighted turn-level SFT, and c) teacher-guided SFT. The teacher model (which is a larger model) is only used to fix formatting and evaluate examples, but does not create new gold actions. Our final model scores 57.12 clemscore and 42.68 statscore on the public Playpen validation. In the officially released challenge results, our model obtains the second-highest Playpen clemscore delta among submitted systems (which is approximately +36 over its base model). Our findings suggest that imitating full trajectories helps with playability, while turn-level and teacher-guided training usually improve decision-making and increase the overall score. Alternative procedurally heavy approaches like replay-repair and hard-example mining did not help, which suggests that small models are performant simply by using careful curation strategies rather than aggressive changes. We make available both the model and the code for reproducibility.
Chinese Translation
我们提出了Qwen-GuidePlay-2B,这是一个用于对话游戏交互的2B参数语言模型。我们通过三个步骤对Qwen3.5-2B进行了微调:a) 仅在来自Playpen的成功游戏轨迹上进行监督微调(SFT),b) 加权回合级SFT,以及c) 教师引导的SFT。教师模型(一个更大的模型)仅用于修正格式和评估示例,但不生成新的黄金动作。我们的最终模型在公共Playpen验证集上得分57.12 clemscore和42.68 statscore。在官方发布的挑战结果中,我们的模型在提交系统中获得了第二高的Playpen clemscore增量(约为+36,相较于其基础模型)。我们的研究结果表明,模仿完整轨迹有助于可玩性,而回合级和教师引导的训练通常改善决策制定并提高整体得分。诸如重放修复和困难示例挖掘等替代性程序重的方式并未带来帮助,这表明小型模型通过谨慎的策划策略而非激进的变化即可实现良好性能。我们提供了模型和代码以便于重现。
cs.CL / 16 / 2608.27729
Below the Noise Floor: Bimodal Seed Collapse and Distinct Failure Modes in Small-Model Knowledge Distillation
噪声底线以下:小模型知识蒸馏中的双模种子崩溃与不同失败模式
Abstract
Function routing -- selecting the correct API call from a fixed catalog given a natural-language request -- is a deployment problem where small students are attractive but knowledge distillation gains are typically reported single-seed, at scales where seed variance is unknown. On a 740-instance healthcare API routing task with a 1.5B Qwen student and a 20B teacher, we compare eight KD variants against supervised cross-entropy, using three to six seeds for key configurations. We find: (i) per-seed standard deviation ranges from 2.8 to 48.7 percentage points, swallowing every claimed KD gain below five points; (ii) three of seven KD variants exhibit bimodal collapse, with at least one in three to five seeds falling below 55% accuracy while the others train normally, and a fourth showing elevated variance; (iii) collapse has distinct modes -- wrong-function selection for ce_kd and ce_paraphrase, and a previously undocumented output-truncation mode for reasoning_kd, where the model emits reasoning but terminates before producing a function name (0.9% accuracy); (iv) only progressive_kd and rank_kd avoid collapse across observed seeds, with sigma <= 3.9 pp; (v) a naive cross-split +3.78 pp gain from input enrichment reverses to -2.70 pp under controlled within-split multi-seed re-testing. Single-seed evaluation is therefore unable to detect central failure modes in small-model KD.
Chinese Translation
功能路由——根据自然语言请求从固定目录中选择正确的API调用——是一个部署问题,其中小型学生模型具有吸引力,但知识蒸馏的收益通常是以单一种子报告的,而在种子方差未知的情况下进行评估。在一个740实例的医疗API路由任务中,使用1.5B的Qwen学生模型和20B的教师模型,我们比较了八种知识蒸馏(KD)变体与监督交叉熵,针对关键配置使用三到六个种子。我们的发现包括:(i)每个种子的标准差范围从2.8到48.7个百分点,吞噬了所有声称的KD收益,低于五个百分点;(ii)七种KD变体中有三种表现出双模崩溃,至少有三到五个种子中的一个准确率低于55%,而其他种子正常训练,还有一种表现出较高的方差;(iii)崩溃具有不同的模式——对于ce_kd和ce_paraphrase的错误功能选择,以及对于reasoning_kd的一种先前未记录的输出截断模式,其中模型发出推理但在生成函数名称之前终止(准确率为0.9%);(iv)只有progressive_kd和rank_kd在观察到的种子中避免了崩溃,标准差<=3.9个百分点;(v)通过输入丰富获得的天真交叉分割+3.78个百分点的收益在控制的分割内多种子重新测试中反转为-2.70个百分点。因此,单一种子评估无法检测小模型KD中的中心失败模式。
cs.CL / 17 / 2608.27756
Load-Bearing Context: The Question Damage Score for Evaluating Context Reliance in Linguistic Reasoning
承载上下文:评估语言推理中上下文依赖性的问答损伤评分
Abstract
Determining whether large language models derive answers from context or prior knowledge remains a fundamental challenge. Self-contained linguistic olympiad puzzles provide a controlled setting where all answers derive solely from expert-designed context examples without external knowledge. Removing individual context examples can eliminate information needed for specific questions while leaving the rest of the puzzle unchanged. We leverage this to introduce a diagnostic framework for analyzing individual context examples. Using 53 UK Linguistics Olympiad puzzles, we generate two modified variants by deleting a single context example: (1) uniform random deletion, and (2) targeted deletion (inspired by error-correcting codes) to remove a structurally load-bearing example uniquely carrying necessary information. We formalize this impact using a Question Damage Score to classify puzzles as fragile or robust. Evaluating three frontier LLMs under instructions to abstain when information is insufficient, we find they rarely abstain, often continuing to produce correct answers after load-bearing context is removed. These findings motivate further investigation into context-based reasoning, prior knowledge, memorization, and linguistic inference. Beyond abstention, the framework enables fine-grained analyses of context reliance, including causal interventions, stopping-set analysis, targeted contamination studies, and mechanistic interpretability.
Chinese Translation
确定大型语言模型是否从上下文或先前知识中推导答案仍然是一个基本挑战。自包含的语言奥林匹克难题提供了一个受控环境,其中所有答案仅来自专家设计的上下文示例,而不依赖外部知识。删除单个上下文示例可以消除特定问题所需的信息,同时保持难题的其余部分不变。我们利用这一点引入了一个诊断框架,用于分析单个上下文示例。通过使用53个英国语言学奥林匹克难题,我们生成了两种修改变体,通过删除单个上下文示例: (1) 均匀随机删除, (2) 目标删除(灵感来自错误更正代码),以删除一个结构上承载必要信息的示例。我们使用问答损伤评分对这一影响进行形式化,以将难题分类为脆弱或稳健。在指示三种前沿大型语言模型在信息不足时保持沉默的情况下进行评估时,我们发现它们很少保持沉默,通常在承载上下文被移除后仍然继续产生正确答案。这些发现激励我们进一步研究基于上下文的推理、先前知识、记忆和语言推断。除了保持沉默之外,该框架还能够对上下文依赖性进行细致的分析,包括因果干预、停止集分析、目标污染研究和机制可解释性。
cs.CL / 18 / 2608.27760
Informational Antilocality and the Locality Bias in LLMs
信息反局部性与大语言模型中的局部性偏差
Abstract
We consider the ability of transformer-based language models (LLMs) to learn what we call k-antilocal languages, i.e., languages that have no mutual information across any span of $k$ contiguous symbols. We construct such languages with increasing $k$, finding that LLMs trained on them achieve comparable cross-entropy loss regardless of antilocality, but converge more slowly on more antilocal languages. Our findings support the idea that non-local dependencies are more difficult to learn, but the evidence for this bias comes from learning speed rather than learning success.
Chinese Translation
我们考虑基于变换器的语言模型(LLMs)学习我们称之为k-反局部语言的能力,即在任何k个连续符号的跨度内没有互信息的语言。我们构造了随着k增加而变化的此类语言,发现训练于这些语言的LLMs在交叉熵损失上表现相当,不论反局部性如何,但在更反局部的语言上收敛速度较慢。我们的发现支持了非局部依赖关系更难以学习的观点,但这种偏差的证据来自学习速度而非学习成功率。
cs.CL / 19 / 2608.27785
Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict
音视频大语言模型中的组合失败:跨模态冲突下的晚层优先主导
Abstract
We study audio-visual conflict as a compositional generalization test for AV-LLMs: the model must combine synchronized but semantically incompatible audio and video evidence and decide whether the pair matches. On VideoLLaMA 2-7B-AV, three alignment configurations remain nearchance on the scored exact-string Yes/No subset of AVHBench, even though their output priors shift substantially. Similarly, off-the-shelf InternVideo2 experienced a 32.3% accuracy decrease specifically under cross-modal conflict, accompanied by a 17.3% instruction-following failure. We call this failure mode prior dominance: late-layer commitment to an internally preferred answer pattern that is weakly grounded in the conflicting inputs. To explain this behavior, we conduct a mechanistic interpretability analysis and find that commitment remains concentrated at 25.5 $\pm$ 1 layers. We show that stronger temporal alignment changes answer bias, but do not improve compositional conflict resolution. Code and data to reproduce our mechanistic audit and behavioral evaluations are available at https://github.com/AdarshSudheer09/AVHBench-dmai.
Chinese Translation
我们将音视频冲突作为音视频大语言模型(AV-LLMs)的组合泛化测试:模型必须结合同步但语义不兼容的音频和视频证据,并决定这对是否匹配。在 VideoLLaMA 2-7B-AV 上,尽管其输出先验发生了显著变化,但在 AVHBench 的精确字符串是/否子集上,三种对齐配置的表现仍接近随机。类似地,现成的 InternVideo2 在跨模态冲突下的准确率下降了 32.3%,同时出现了 17.3% 的指令遵循失败。我们将这种失败模式称为先验主导:晚层对内部偏好答案模式的承诺与冲突输入的基础较弱。为了解释这种行为,我们进行了机械可解释性分析,发现承诺仍然集中在 25.5 ± 1 层。我们表明,较强的时间对齐会改变答案偏差,但并未改善组合冲突的解决。用于重现我们的机械审计和行为评估的代码和数据可在 https://github.com/AdarshSudheer09/AVHBench-dmai 获得。
cs.CL / 20 / 2608.27813
Representation of syntax in LLMs through the lens of linear distance and similarity-aware entropy
通过线性距离和相似性感知熵的视角理解大型语言模型中的句法表示
Abstract
Structural probes were introduced by Hewitt and Manning to reconstruct syntactic trees from a neural language model's latent representations. They are evaluated by calculating the proportion of syntactic tree edges correctly reconstructed over an annotated corpus (as measured by undirected unlabeled attachment score). Here, we disaggregate this measure, considering undirected attachment score by label (UASL), which assesses the reconstruction accuracy of each syntactic relation separately, establishing important differences among relations that overlap linguistic distinctions. Moreover, we identify two factors that predict most of UASL's variability across relations: (i) the mean and dispersion of the linear distance (on a log scale) between the related words, and (ii) the diversity (similarity-aware entropy) of the syntactic relation's head. These results, which hold across a range of model sizes and architectures, shed light on the degree of abstraction of the representation of syntax in language models and the dependence of such representation on geometric properties of the embedding space.
Chinese Translation
Hewitt 和 Manning 引入了结构探针,以从神经语言模型的潜在表示中重构句法树。通过计算在标注语料库中正确重构的句法树边的比例(以无向未标记附着分数衡量)来评估它们。在此,我们对这一指标进行细分,考虑按标签的无向附着分数(UASL),该指标分别评估每种句法关系的重构准确性,揭示了在语言学区分中重叠的关系之间的重要差异。此外,我们识别出两个因素,可以预测 UASL 在不同关系中的大部分变异性:(i)相关词之间的线性距离的均值和离散度(以对数尺度表示),以及(ii)句法关系主语的多样性(相似性感知熵)。这些结果在多种模型规模和架构中均成立,揭示了语言模型中句法表示的抽象程度及其对嵌入空间几何属性的依赖性。
cs.CL / 21 / 2608.27816
PersonaEdit: Representative Sample Selection for Personalized Model Editing
PersonaEdit:个性化模型编辑的代表性样本选择
Abstract
Personalization has attracted growing interest in LLM applications, yet existing retrieval-based approaches depend heavily on retrieval quality and degrade in long-term interactions. Model editing, which directly modifies internal model parameters to incorporate new knowledge, has demonstrated effective knowledge modification capabilities in factual knowledge editing tasks and may provide a potential solution for personalization. However, scaling model editing to personalization is non-trivial. Editing large amounts of user data increases computational cost and causes interference among edits, motivating the need for effective sample selection. To address this issue, we propose, PersonaEdit, a hidden representation clustering strategy that selects representative editing samples through proportional stratified sampling. Experiments show that model editing is effective for personalization, and that our selection strategy preserves most of the performance while substantially reducing the number of required editing samples. Beyond standalone editing, we find that combining model editing with retrieval-based prompt augmentation further improves personalization, as edited knowledge and retrieved context provide complementary information. These results demonstrate the potential of model editing as an efficient and scalable approach for LLM personalization.
Chinese Translation
个性化在大规模语言模型(LLM)应用中引起了越来越多的关注,但现有的基于检索的方法在很大程度上依赖于检索质量,并且在长期交互中表现不佳。模型编辑通过直接修改内部模型参数来融入新知识,已在事实知识编辑任务中展现出有效的知识修改能力,并可能为个性化提供潜在解决方案。然而,将模型编辑扩展到个性化并非易事。编辑大量用户数据会增加计算成本并导致编辑之间的干扰,这促使我们需要有效的样本选择。为了解决这个问题,我们提出了PersonaEdit,一种隐藏表示聚类策略,通过比例分层抽样选择代表性的编辑样本。实验表明,模型编辑在个性化方面是有效的,并且我们的选择策略在显著减少所需编辑样本数量的同时,保留了大部分性能。除了独立编辑外,我们发现将模型编辑与基于检索的提示增强相结合进一步改善了个性化,因为编辑的知识和检索的上下文提供了互补的信息。这些结果展示了模型编辑作为一种高效且可扩展的LLM个性化方法的潜力。
cs.CL / 22 / 2608.27843
Synthetic Linguistic Agency: How an Embodied Mortal Agent Learns Linguistic Affordances through Consequential Social Experience
合成语言代理:一个具身的凡人代理如何通过后果性社会经验学习语言可供性
Abstract
Contemporary language models can converse fluently and influence human decisions, yet their exchanges do not enter a continuing, vulnerable life of their own. Linguistic-agency theory identifies this missing connection as linguistic agency and characterizes it through embodiment, linguistic participation, and precariousness: a body that acts and bears consequences, interaction that changes both agent and partner, and a future that can be sustained or lost. Two coordinated studies examine how this organization can appear in artificial systems. First, we translate these relations into inspectable criteria for Synthetic Linguistic Agency (SLA) and identify several existing SLA systems. Second, building on Homeostatically Regulated Reinforcement Learning, we develop a mortality-grounded linguistic-reinforcement-learning model and instantiate it in an Embodied Mortal Agent (EMA). The EMA learns how ways of speaking change a partner's willingness to protect it and chooses expressions by considering what those responses mean for its remaining life. Controlled experiments show that linguistic choices depend on the EMA's body and social history, change partner behavior, and adapt through experience with particular partners. When bodily consequences persist, linguistic choices alter the future of the same life; when the body is reset, their social effects remain but no longer shape continued viability. The resulting EMA exhibits SLA under our operational definition. This work motivates further research on synthetic empathy and strategic human-AI interaction: how artificial agents with persistent bodies, histories, and futures might develop and express empathy, and how people might care for, negotiate with, or govern them.
Chinese Translation
当代语言模型能够流利地进行对话并影响人类决策,但它们的交流并未进入一个持续的、脆弱的生命状态。语言代理理论将这种缺失的联系称为语言代理,并通过具身性、语言参与和脆弱性来表征:一个能够行动并承担后果的身体、改变代理者和伙伴的互动,以及一个可以持续或失去的未来。两项协调研究考察了这种组织如何在人工系统中出现。首先,我们将这些关系转化为可检验的合成语言代理(Synthetic Linguistic Agency, SLA)标准,并识别出几个现有的SLA系统。其次,基于稳态调节强化学习(Homeostatically Regulated Reinforcement Learning),我们开发了一个以死亡为基础的语言强化学习模型,并在一个具身的凡人代理(Embodied Mortal Agent, EMA)中实现。EMA学习如何通过语言表达改变伙伴的保护意愿,并通过考虑这些反应对其剩余生命的意义来选择表达方式。控制实验表明,语言选择依赖于EMA的身体和社会历史,改变伙伴行为,并通过与特定伙伴的经验进行适应。当身体后果持续存在时,语言选择改变同一生命的未来;当身体被重置时,其社会影响仍然存在,但不再影响持续的生存能力。最终,EMA在我们的操作定义下展现出SLA。这项工作激励了对合成同理心和战略性人机互动的进一步研究:如何具有持续身体、历史和未来的人工代理可能发展和表达同理心,以及人们如何关心、与之协商或管理它们。
cs.CL / 23 / 2608.27844
EvoHarmBench: Breaking Content Moderation with Iterative Human-Like Evasion
EvoHarmBench:通过迭代类人规避打破内容审核
Abstract
Existing evaluations of harmful content detection rely predominantly on static benchmarks, which struggle to reflect the interactive adversarial ecosystem of real-world content platforms where users continuously revise their expressions in response to moderation feedback. This mismatch creates a significant performance gap between offline benchmark scores and online deployment effectiveness. To the best of our knowledge, we present EvoHarmBench, the first dynamic adversarial evaluation framework for content moderation systems. The framework employs an iterative optimization loop that evolves evasion strategies at the semantic-cluster level, while simultaneously optimizing for evasion success and human readability. We systematically evaluate LLM-based defense models which are widely used in real world moderation systems. The evaluation covers 229 semantic sub-clusters across five violation categories, derived from 5,002 real-world adversarial samples collected from content platforms. Our experiments reveal substantial vulnerabilities even in leading commercial systems: after twelve optimization iterations, the attack success rate under readability constraints reaches 80.3% within SOTA LLM moderators. We will release the full benchmark data, evaluation framework, and code to encourage a shift from static benchmarking toward dynamic adversarial evaluation in content safety research.
Chinese Translation
现有的有害内容检测评估主要依赖于静态基准,这难以反映现实内容平台中用户根据审核反馈不断修订表达的互动对抗生态。这种不匹配导致离线基准分数与在线部署效果之间存在显著的性能差距。我们提出了EvoHarmBench,这是首个动态对抗评估框架,旨在评估内容审核系统。该框架采用迭代优化循环,在语义聚类层面上演变规避策略,同时优化规避成功率和人类可读性。我们系统地评估了在现实世界审核系统中广泛使用的基于大型语言模型(LLM)的防御模型。评估涵盖了来自5002个真实世界对抗样本的五个违规类别下的229个语义子聚类。我们的实验揭示了即使在领先的商业系统中也存在显著的脆弱性:在可读性约束下,经过十二次优化迭代后,攻击成功率在SOTA LLM审核员中达到了80.3%。我们将发布完整的基准数据、评估框架和代码,以鼓励内容安全研究从静态基准测试转向动态对抗评估。
cs.CL / 24 / 2608.27855
AI Writers Have a Consistent Stylometric Footprint, but AI Editors Do Not
Abstract
Text generated by large language models (LLMs) has been shown to be stylometrically distinct from human-written text \citep{andreDetectingAIAuthorship2023, shahDetectingUnmaskingAIGenerated2023, oparaStyloAIDistinguishingAIGenerated2024, soto2024fewshot, liLinguisticDifferencesAI2025, selviogluFeatureExtractionAnalysis2025}. But LLMs are increasingly used not only to generate text but also to edit human writing, and it is unclear whether the two leave the same trace. We show that AI generation leaves a consistent ``stylometric footprint'': a small subset of features, primarily entropy and lexical diversity, consistently separates AI-generated text from human writing across 8 LLMs and 5 domains, while the remaining features depend heavily on the domain and generator. AI editing, however, does not reproduce the same footprint. Relative to their human-written sources, AI-edited texts show only a small increase in lexical diversity and a decrease in entropy, rather than the joint increase that characterizes AI generation. Lexical density, which contributes little to generation, instead becomes the dominant editing-associated signal. Stylometric features therefore separate AI-edited text from AI-generated text but are substantially less effective at separating it from human-written text. Our results suggest that ``AI text'' is not a single phenomenon: generation and editing leave qualitatively different stylometric traces and should be studied separately.
cs.CL / 25 / 2608.27899
OpenStamp: A Watermark for Open-Source Language Models
OpenStamp:一种用于开源语言模型的水印
Abstract
With the growing prevalence of large language model (LLM) generated content, watermarking is considered a promising approach for attributing text to LLMs and distinguishing it from human-written content. A prominent class of techniques embeds subtle but detectable signals in generated text by modifying token sampling probabilities. However, such methods are unsuitable for open-source models, where users have white-box access and can easily disable watermarking during inference. In this work, we introduce OpenStamp, a watermarking technique that encodes the watermarking logic directly into the model weights by modifying only the final projection, or unembedding, layer. Through experiments across two models, we show that OpenStamp achieves superior detection performance, with minimal degradation in model capabilities compared to prior methods. The implanted watermark is explicitly designed, and empirically confirmed, to be more robust to paraphrasing attacks and harder to scrub off through post-hoc fine-tuning than prior open-source watermarks. To enable developers to watermark their models, we release our code alongside watermarked versions of 4 popular open-source models.
Chinese Translation
随着大型语言模型(LLM)生成内容的日益普及,水印技术被认为是一种有前景的方法,用于将文本归属到LLM,并将其与人类撰写的内容区分开来。一类显著的技术通过修改令牌采样概率在生成的文本中嵌入微妙但可检测的信号。然而,这些方法不适用于开源模型,因为用户可以进行白盒访问,并且可以轻松地在推理过程中禁用水印。在本研究中,我们介绍了OpenStamp,一种水印技术,通过仅修改最终的投影层或去嵌入层,将水印逻辑直接编码到模型权重中。通过对两个模型的实验,我们表明OpenStamp在检测性能上优于先前的方法,并且模型能力的降级最小。植入的水印经过明确设计,并通过实验证实,对改写攻击具有更强的鲁棒性,并且比以往的开源水印更难通过后期微调去除。为了使开发者能够为他们的模型添加水印,我们发布了我们的代码以及4个流行开源模型的水印版本。
cs.CL / 26 / 2608.27902
LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing Pages
LandingAgent:一个参考注释数据集和着陆页生成框架
Abstract
Landing pages are goal-oriented web interfaces that must communicate a target-specific value proposition while organizing information flow, visual hierarchy, and calls to action (CTA). Although large language models can generate plausible webpage code from natural-language prompts, direct generation often yields generic templates and unsupported persuasive claims. We study target-grounded, reference-guided landing-page generation, where a system must create an executable page for a new target by adapting reusable patterns from real pages without copying them. We introduce LandingBench, a reference-profile dataset that abstracts real landing pages into section sequences, layout patterns, tone descriptors, visual emphasis, and CTA structure. Building on LandingBench, we propose LandingAgent, a three-phase agentic framework that profiles the target, constructs a reference-guided wireframe, and refines the page through critique-guided polishing. We evaluate LandingAgent against direct prompting on faithfulness, conciseness, readability, aesthetics, and structural diversity. Experiments show improved target grounding, presentation quality, and layout diversity. Code is available at https://github.com/IAURAI/LandingAgent.
Chinese Translation
着陆页是目标导向的网络界面,必须传达特定目标的价值主张,同时组织信息流、视觉层次和行动号召(CTA)。尽管大型语言模型可以根据自然语言提示生成合理的网页代码,但直接生成往往会产生通用模板和不支持的说服性声明。我们研究目标导向的、参考引导的着陆页生成,其中系统必须通过适应真实页面的可重用模式而非复制它们,来为新目标创建可执行页面。我们引入了LandingBench,一个参考配置数据集,将真实的着陆页抽象为部分序列、布局模式、语气描述、视觉强调和CTA结构。在LandingBench的基础上,我们提出了LandingAgent,一个三阶段的代理框架,首先对目标进行分析,构建参考引导的线框,然后通过批评引导的润色来完善页面。我们在忠实性、简洁性、可读性、美观性和结构多样性方面评估LandingAgent与直接提示的表现。实验结果显示目标导向、展示质量和布局多样性有所改善。代码可在 https://github.com/IAURAI/LandingAgent 获取。
cs.CL / 27 / 2608.27924
What Makes Agent Memory Useful for Reliable Unanswerable Question Handling?
什么使得代理记忆在可靠的不可回答问题处理上变得有用?
Abstract
Reliable handling of unanswerable questions (UAQs) is critical for trustworthy LLM-based agents. Although memory is widely used in agent systems, its role in reliable UAQ handling remains unclear. We present a systematic study of agent memory for UAQ handling under a unified agentic RAG framework, evaluating four representative memory methods across three UAQ-related datasets and two base models. We find that memory can improve UAQ performance in some settings, but such gains are selective rather than universal and remain fragile under dataset shift. Interestingly, cross-model memory reuse is often more feasible than cross-dataset transfer, suggesting that shifts in answerability patterns pose a greater challenge to memory reuse than changes in the base model itself. We further find that UAQ gains are more strongly preserved through decision guidance than through trajectory shaping, and that memory effectiveness depends strongly on representation. In particular, procedural and rule-based memories often provide the most reliable support for UAQ handling, while memory composition is most effective when procedural guidance is combined with complementary behavioral signals. Overall, our findings suggest that reliable UAQ memory depends less on storing larger amounts of experience and more on preserving transferable behavioral guidance.
Chinese Translation
可靠处理不可回答问题(UAQs)对于基于大型语言模型(LLM)的代理系统至关重要。尽管记忆在代理系统中被广泛使用,但其在可靠的UAQ处理中的作用仍不明确。我们在统一的代理RAG框架下,对代理记忆在UAQ处理中的作用进行了系统研究,评估了四种代表性的记忆方法在三个与UAQ相关的数据集和两个基础模型上的表现。我们发现,在某些设置中,记忆可以提高UAQ的表现,但这种提升是选择性的而非普遍的,并且在数据集变化时仍然脆弱。有趣的是,跨模型的记忆重用通常比跨数据集的转移更可行,这表明回答模式的变化对记忆重用构成了更大的挑战,而不是基础模型本身的变化。我们进一步发现,通过决策指导所获得的UAQ提升比通过轨迹塑造更为显著,并且记忆的有效性在很大程度上依赖于表示方式。特别是,过程性和基于规则的记忆通常为UAQ处理提供最可靠的支持,而当过程性指导与互补的行为信号相结合时,记忆组合的效果最为显著。总体而言,我们的研究结果表明,可靠的UAQ记忆更依赖于保留可转移的行为指导,而不是存储大量的经验。
cs.CL / 28 / 2608.27925
Entity-Memory Graph Retrieval Improves Evidence Coverage in Long-Conversation Question Answering
实体记忆图检索提高了长对话问答中的证据覆盖率
Abstract
Entity-Memory graph retrieval keeps dialogue turns as verbatim Memory nodes, links repeated mentions through shared Entities, and connects adjacent Memories with directed chronological edges. At query time the retriever moves from Entity gating through semantic fusion and one-hop chronological recovery to dense backfill. The path can keep a neighboring Memory that dense cosine ranking would otherwise omit. A matched dense control shares the Memory and query vectors, context budget, requested answer protocol, and evaluator, isolating graph structure from changes to the reader. On 1,986 questions from ten LoCoMo conversations, graph retrieval raises official evidence recall at top-k 25 from 79.7468% to 84.4842%. The recall advantage is supported from top-k 5 to 50, while no matched cutoff supports an overall final-answer F1 difference. Four paper-eligible requested configurations support empirical robustness across the tested GPT-3.5 and DeepSeek extractors on both outcomes. Embedding robustness is mixed: F1 has no supported contrast, but recall is sensitive to the embedding artifact. The comparison isolates a retrieval-coverage gain from graph structure. It does not establish a final-answer F1 gain, model or embedding equivalence, or cross-dataset generalization.
Chinese Translation
实体记忆图检索将对话轮次作为逐字记忆节点,通过共享实体链接重复提及,并用有向时间边连接相邻的记忆。在查询时,检索器从实体门控通过语义融合和单跳时间恢复到密集回填。该路径可以保留一个相邻的记忆,而密集余弦排名通常会忽略它。匹配的密集控制共享记忆和查询向量、上下文预算、请求答案协议和评估者,将图结构与读者的变化隔离。在来自十个 LoCoMo 对话的 1,986 个问题中,图检索将官方证据召回率在前 25 名中从 79.7468% 提高到 84.4842%。这一召回优势在前 5 到 50 名中得到了支持,而没有匹配的截止点支持整体最终答案 F1 差异。四个符合论文要求的请求配置在测试的 GPT-3.5 和 DeepSeek 提取器的两个结果上支持经验稳健性。嵌入稳健性表现不一:F1 没有支持的对比,但召回对嵌入伪影敏感。比较结果将图结构的检索覆盖增益进行了隔离。它并未确立最终答案 F1 增益、模型或嵌入等价性,或跨数据集的泛化能力。
cs.CL / 29 / 2608.27966
Lexically conditioned realization ambiguity in Korean predicate morphology
韩语谓词形态中的词汇条件化实现歧义
Abstract
This paper examines Korean surface realization as distinct from morphological analysis. It asks whether a sequence of canonical morphemes and grammatical category labels uniquely determines the corresponding surface form. The answer is negative for a restricted but theoretically revealing class of Korean predicates. In these cases, formally identical or near-identical stem-ending configurations yield different outputs depending on lexical identity and realization class membership. We analyze this phenomenon as homonymy with inflectional divergence, focusing on regular versus digeut irregular pairs, regular versus bieup irregular pairs, and reu irregular versus reo irregular pairs. These cases show that stem shape and ending alone do not always determine surface realization. Instead, lexical meaning, subcategorization, and semantic role structure help identify the intended predicate; the predicate determines the realization class; and the realization class determines the surface form. Korean realization thus reveals a limit of bare morphological representation.
Chinese Translation
本文考察了韩语表面实现与形态分析的区别。研究问题是,一系列规范的语素和语法类别标签是否唯一决定相应的表面形式。对于一类有限但在理论上具有启示性的韩语谓词,答案是否定的。在这些情况下,形式上相同或近似相同的词干结尾配置会根据词汇身份和实现类别的成员资格产生不同的输出。我们将这一现象分析为具有屈折差异的同音异义,重点关注规则与“디귿”(digeut)不规则对、规则与“비읍”(bieup)不规则对,以及“르”(reu)不规则与“레”(reo)不规则对。这些案例表明,仅凭词干形状和结尾并不总能决定表面实现。相反,词汇意义、子分类和语义角色结构有助于识别所需的谓词;谓词决定实现类别;而实现类别则决定表面形式。因此,韩语的实现揭示了裸形态表示的局限性。
cs.CL / 30 / 2608.27974
QUORUM: QUality-Optimized Routing Using Multiple annotators
QUORUM:基于多标注者的质量优化路由
Abstract
Data annotation remains a central bottleneck in natural language processing, requiring human effort to obtain high-quality labels at scale. While Large Language Models (LLMs) offer a fast and cost-effective alternative, their reliability is highly instance-dependent: they perform well on simple inputs but often fail on examples requiring nuanced reasoning or contextual understanding. In this work, we address this challenge with QUORUM (QUality-Optimized Routing Using Multiple annotators), a budget-aware routing framework that dynamically assigns each instance to human or LLM annotators under a fixed annotation budget. Unlike prior approaches relying on model confidence or uncertainty estimates, QUORUM leverages feature-based signals to estimate instance difficulty and supports multiple annotations per instance, combining them through agreement-based rewards to improve reliability. We evaluate QUORUM across diverse closed- and open-ended annotation tasks in English and multilingual settings, and QUORUM improves annotation quality by up to 34.4% while reducing costs by 8.8% over competing methods. Code can be found at https://github.com/amazon-science/QUORUM.
Chinese Translation
数据标注仍然是自然语言处理中的一个核心瓶颈,需要人力投入以获得高质量的标签。在规模化标注中,虽然大型语言模型(LLMs)提供了一种快速且具有成本效益的替代方案,但它们的可靠性高度依赖于具体实例:在简单输入上表现良好,但在需要细致推理或上下文理解的示例中往往失败。在本研究中,我们通过QUORUM(基于多标注者的质量优化路由)来应对这一挑战,QUORUM是一个预算感知的路由框架,能够在固定的标注预算下动态地将每个实例分配给人类或LLM标注者。与之前依赖模型置信度或不确定性估计的方法不同,QUORUM利用基于特征的信号来估计实例的难度,并支持每个实例的多次标注,通过基于一致性的奖励将标注结果结合起来,以提高可靠性。我们在多种封闭和开放式标注任务中评估了QUORUM,包括英语和多语言环境,结果显示QUORUM在提高标注质量方面最多可提升34.4%,同时相比竞争方法降低成本8.8%。代码可在https://github.com/amazon-science/QUORUM找到。
cs.CL / 31 / 2608.27988
Predicting Turn-Taking Outcomes in Multi-Party Conversation: Interpretable Modelling of Speech and Gaze Dynamics with Interpersonal Closeness
多方对话中轮流发言结果的预测:可解释的言语与目光动态建模与人际亲密度
Abstract
Smooth speaker transitions are fundamental to effective conversation and rely on an interlocutor's ability to predict when to enter the conversation. This ability depends on accurately interpreting and expressing the verbal and non-verbal cues that signal when a speaker wishes to take or relinquish the floor. The process becomes even more complex in noisy, natural, multi-party settings, with multiple interlocutors available. This study models how gaze and speech, together with perceived interpersonal closeness, signal conversational floor changes in free four-person dialogue. Using the GaMMA corpus, we trained logistic regression models using interpretable, behaviourally motivated features extracted before each turn-taking event to classify floor-transfer outcomes as gaps or overlaps. Predictors included gaze features such as transition motifs and behavioural contrasts, entropy, gaze-based addressee identity, and mutual gaze, alongside speech features derived from speaker loudness, as well as perceived interpersonal closeness (IOS) between speakers. Results show that gaze features capture predictive structure, and that combining them with loudness improves performance (ROC AUC = 0.76 +- 0.04). Loudness reflected speaker control, while gaze dispersion and addressing indexed listener readiness and competitive entry. Performance remained robust across noise conditions, indicating that gaze provides a complementary, noise-resilient cue to turn-taking dynamics.
Chinese Translation
顺畅的发言者转换是有效对话的基础,依赖于对话者预测何时进入对话的能力。这种能力取决于准确解读和表达信号,表明发言者希望接管或放弃发言权的言语和非言语线索。在嘈杂的自然多方环境中,多个对话者的存在使得这一过程变得更加复杂。本研究建模了在自由的四人对话中,目光和言语如何与感知的人际亲密度一起,信号化对话中的发言权变化。我们使用GaMMA语料库,训练了逻辑回归模型,利用在每次轮流发言事件之前提取的可解释的、行为驱动的特征,将发言权转移结果分类为间隙或重叠。预测变量包括目光特征,如转移动机和行为对比、熵、基于目光的听众身份以及相互注视,此外还有来自发言者音量的言语特征,以及发言者之间的感知人际亲密度(IOS)。结果表明,目光特征捕捉到了预测结构,并且将其与音量结合使用提高了性能(ROC AUC = 0.76 ± 0.04)。音量反映了发言者的控制,而目光分散和指向则标识了听众的准备状态和竞争性进入。性能在不同噪声条件下保持稳健,表明目光为轮流发言动态提供了一种互补的、抗噪声的线索。
cs.CL / 32 / 2608.28009
Beyond Global Scalars: Synergizing Token-Level Statistics and Deep Semantics for Adversarial AIGC Text Detection
超越全局标量:协同令牌级统计与深层语义以进行对抗性AIGC文本检测
Abstract
The rapid evolution of large language models necessitates robust machine-generated text detection. Existing paradigms typically follow two isolated tracks. Training-free methods rely on global statistical scalars such as perplexity, while training-based methods utilize semantic hidden states. Both approaches exhibit fundamental vulnerabilities in adversarial scenarios. Global scalars act as lossy compressions that obscure local probabilistic burstiness in interleaved texts, whereas pure semantic models overfit to specific fingerprints and remain susceptible to spoofing. To expose these flaws, we introduce MOSAIC, a comprehensive adversarial benchmark comprising 16000 samples across a full-granularity attack spectrum. To address these challenges, we propose NeuroStat, an end-to-end framework bridging the statistical and semantic gap. NeuroStat captures uncompressed token-level probabilistic logits alongside deep semantic hidden states from a single causal language model backbone. We fuse these heterogeneous signals through Macro-State Residual Modulation, which adaptively calibrates local convolutional features using global uncertainty indicators. Orthogonal and contrastive losses further ensure the learning of complementary representations. Extensive experiments demonstrate that NeuroStat maintains exceptional robustness on MOSAIC compared to the severe degradation of state-of-the-art methods, establishing a new standard for adversarial text detection. Code and the MOSAIC benchmark are available at https://github.com/TencentBAC/NeuroStat.
Chinese Translation
大型语言模型的快速发展需要强大的机器生成文本检测。现有的范式通常遵循两个孤立的轨道。无训练的方法依赖于全局统计标量,如困惑度,而基于训练的方法则利用语义隐状态。这两种方法在对抗性场景中都表现出基本的脆弱性。全局标量作为有损压缩,掩盖了交织文本中的局部概率突发性,而纯语义模型则过拟合于特定指纹,容易受到伪造的攻击。为了揭示这些缺陷,我们引入了MOSAIC,这是一个全面的对抗性基准,包含16000个样本,覆盖全粒度攻击谱。为了解决这些挑战,我们提出了NeuroStat,一个端到端框架,弥合统计与语义之间的差距。NeuroStat从单一的因果语言模型骨干捕获未压缩的令牌级概率logits以及深层语义隐状态。我们通过宏状态残差调制(Macro-State Residual Modulation)融合这些异质信号,该方法使用全局不确定性指标自适应地校准局部卷积特征。正交和对比损失进一步确保学习互补表示。广泛的实验表明,与最先进的方法在MOSAIC上的严重退化相比,NeuroStat在鲁棒性方面保持了卓越的表现,为对抗性文本检测建立了新的标准。代码和MOSAIC基准可在https://github.com/TencentBAC/NeuroStat获取。
cs.CL / 33 / 2608.28018
Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded Reasoning
双重世界:基于等变性的弃权用于证据基础推理
Abstract
Knowledge-intensive reasoning requires Large Language Models (LLMs) to ground answers in provided evidence. When evidence is insufficient, it is desirable that models abstain rather than confidently generating unsupported answers. Existing abstention methods rely on uncertainty estimation or evidence sufficiency checks, but neither tests whether the reasoning process for generation, driven by the interaction of provided evidence and the model's internal memory parameters, is actually grounded in the evidence. A key contributing factor is that entity mentions in context activate memorised associations, causing models to generate plausible responses ungrounded in evidence. We propose Twin Worlds (TW), a framework for improving reliability in knowledge-intensive reasoning through equivariance-based abstention: unlike invariance, which requires outputs to remain unchanged, equivariance requires outputs to transform correspondingly under entity substitutions. A model grounded in the evidence should produce answers that shift consistently when entities are substituted while their relations are preserved. TW constructs multiple worlds via typed substitutions of the original input that preserve relational structure while reducing parametric priors, and uses equivariance violations as an abstention signal. Across four benchmarks and three model backbones, TW identifies when answers are not reliably grounded in the provided evidence and outperforms uncertainty- and sufficiency-based baselines.
Chinese Translation
知识密集型推理要求大型语言模型(LLMs)将答案基于提供的证据进行支撑。当证据不足时,模型应选择弃权,而不是自信地生成没有支持的答案。现有的弃权方法依赖于不确定性估计或证据充分性检查,但这些方法都未能测试生成过程的推理是否真正基于证据,这一过程受到提供的证据与模型内部记忆参数的交互影响。一个关键因素是上下文中的实体提及会激活记忆中的关联,导致模型生成与证据无关的合理响应。我们提出了双重世界(Twin Worlds, TW),这是一个通过基于等变性的弃权来提高知识密集型推理可靠性的框架:与不变性不同,不变性要求输出保持不变,而等变性要求输出在实体替换下相应地转变。一个基于证据的模型在替换实体时应产生一致变化的答案,同时保持其关系。TW通过类型替换原始输入构建多个世界,这些替换保持关系结构的同时减少参数先验,并利用等变性违反作为弃权信号。在四个基准测试和三个模型骨干上,TW能够识别出何时答案未能可靠地基于提供的证据,并优于基于不确定性和充分性的基线方法。
cs.CL / 34 / 2608.28040
A Shaky Voice Is Not Always a Dodge: Benchmarking Textual and Vocal Evasion Detection in Earnings Calls
颤抖的声音并不总是逃避:财报电话会议中文本与声音逃避检测的基准评估
Abstract
Existing approaches to evasion detection in earnings calls focus on textual transcripts, treating evasion as a single-dimensional phenomenon. We argue that evasion in spoken communication is inherently multidimensional: beyond what executives say, how they say it carries independent and complementary information. To study these dimensions jointly, we introduce DualEvasion, a benchmark for evasion detection across text and audio in earnings call Q&A. The benchmark contains 505 annotated question-answer pairs from 60 earnings calls, each with two independent labels: textual evasion (direct vs. evasive) and vocal cues operationalized as speaker confidence (confident vs. unconfident). Our experiments show that state-of-the-art multimodal models struggle to detect vocal confidence, particularly on unconfident responses. Our analysis suggests these models interpret acoustic cues in isolation rather than relative to each speaker's baseline. Providing speaker-level references yields modest improvements, but a substantial gap with human performance remains.
Chinese Translation
现有的财报电话会议逃避检测方法主要集中于文本转录,将逃避视为一种单维现象。我们认为,口头交流中的逃避本质上是多维的:除了高管所说的内容,表达方式本身也携带独立且互补的信息。为了共同研究这些维度,我们引入了DualEvasion,这是一个用于财报电话会议问答环节中跨文本和音频的逃避检测基准。该基准包含来自60个财报电话会议的505对标注的问答,每对问答都有两个独立标签:文本逃避(直接 vs. 逃避)和声音线索,后者通过说话者的自信程度(自信 vs. 不自信)进行操作化。我们的实验表明,最先进的多模态模型在检测声音自信度方面存在困难,尤其是在不自信的回答中。我们的分析表明,这些模型在解释声学线索时往往是孤立进行的,而不是相对于每位说话者的基线进行的。提供说话者级别的参考虽能带来适度的改进,但与人类表现之间仍存在显著差距。
cs.CL / 35 / 2608.28042
SimpCue: Cue-Based Prompting for Multilingual Text Simplification
SimpCue:基于提示的多语言文本简化
Abstract
Text simplification aims to make complex texts easier to understand while preserving their original meaning. Recent large language models can perform simplification through prompting, but it remains unclear whether adding explicit linguistic information about sentence complexity to the prompt improves their outputs. We investigate this question for multilingual sentence-level Easy-to-Read simplification in Catalan, Spanish, and Italian. Using Qwen3-8B, we compare a baseline prompt, a gold-cue prompt enriched with gold linguistic cues, and a predicted-cue prompt enriched with automatically predicted cues. We evaluate the outputs using SARI, BLEU, chrF, and BERTScore, and complement this evaluation with a manual qualitative analysis. Predicted-cue prompting obtains the best overall scores across all four metrics, although the gains over the baseline are small. Gold-cue prompting does not consistently improve over the baseline, and results vary across languages. These findings indicate that cue-based prompting can influence multilingual Easy-to-Read simplification, but its benefits are modest, metric-dependent, and language-dependent.
Chinese Translation
文本简化旨在使复杂文本更易于理解,同时保留其原始含义。近期的大型语言模型可以通过提示进行简化,但尚不清楚在提示中添加关于句子复杂性的显式语言信息是否会改善其输出。我们针对加泰罗尼亚语、西班牙语和意大利语的多语言句子级易读性简化研究了这一问题。使用 Qwen3-8B,我们比较了基线提示、富含金标准语言提示的金标准提示,以及富含自动预测提示的预测提示。我们使用 SARI、BLEU、chrF 和 BERTScore 对输出进行评估,并通过手动定性分析补充该评估。预测提示在所有四个指标中获得了最佳的总体得分,尽管与基线相比的提升较小。金标准提示并未始终优于基线,且结果在不同语言间存在差异。这些发现表明,基于提示的简化可以影响多语言易读性简化,但其益处是适度的、依赖于指标和语言的。
cs.CL / 36 / 2608.28053
CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms
CNeo-Bench:对中文新词的语言模型诊断
Abstract
Chinese neologisms exploit diverse and unique linguistic mechanisms, such as phonetic substitution (e.g., 886 for ``bye-bye'') and visual character decomposition that are rare in other languages. We introduce CNeo-Bench, a benchmark of 4,759 such neologisms with reference definitions, organized into five top-level categories and nine subcategories by the linguistic mechanism behind each expression. CNeo-Bench is paired with a two-tier evaluation framework that separates whether a model can describe a neologism from whether it can operate on its underlying mechanism. Evaluating 18 LLMs, we find that Chinese neologisms remain an open challenge; most models fall below 40\% on definition generation, and on several subcategories a systematic recognition-manipulation gap emerges: models describe neologisms correctly but, in source-form restoration tasks, substitute a semantic equivalent (paraphrase) for the source form rather than producing the source form itself. A few-shot analysis on 1,058 hard items shows that in-context examples can solve many difficult cases, but leave a noticeable portion of errors remaining, indicating challenges beyond prompting alone can address.
Chinese Translation
中文新词利用了多样且独特的语言机制,例如音位替代(如886代表“再见”)和视觉字符分解,这在其他语言中较为罕见。我们介绍了CNeo-Bench,这是一个包含4,759个新词的基准数据集,附有参考定义,并根据每个表达背后的语言机制分为五个顶级类别和九个子类别。CNeo-Bench配备了一个两级评估框架,区分模型是否能够描述新词与其是否能够操作其潜在机制。在对18个大型语言模型(LLMs)的评估中,我们发现中文新词仍然是一个开放性挑战;大多数模型在定义生成上得分低于40%,并且在几个子类别中出现了系统性的识别-操作差距:模型能够正确描述新词,但在源形式恢复任务中,往往用语义等价物(意译)替代源形式,而不是生成源形式本身。对1,058个困难项目的少量样本分析表明,语境示例能够解决许多困难案例,但仍然留有明显的错误部分,表明仅依靠提示无法解决的挑战仍然存在。
cs.CL / 37 / 2608.28113
H-Scale: Hessian-Guided Scale Refinement for NVFP4 Sub-Byte LLM Inference
H-Scale:基于海森矩阵引导的NVFP4子字节大语言模型推理尺度细化
Abstract
The NVIDIA Blackwell architecture, with native support for the ultra-fine-grained NVFP4 format, opens new opportunities for accelerating large language model (LLM) inference. NVFP4's micro-block design, such as a group size of 16, offers strong representational flexibility for capturing local weight distributions and isolating outliers, but it also introduces a large and highly sensitive space of per-group scaling factors. Existing post-training quantization (PTQ) methods primarily focus on refining quantized weight values, leaving this scale-selection step underexplored. To address this gap, we propose \textbf{H-Scale}, a lightweight post-processing method for NVFP4 per-group scale refinement. Instead of minimizing plain weight reconstruction error, H-Scale selects hardware-valid group scales using a diagonal second-order proxy derived from calibration activations, thereby targeting layer output perturbation more directly. It is designed as a drop-in replacement for RTN-style scale selection in diverse NVFP4 pipelines, requires only modest offline calibration, and introduces strictly zero overhead at inference time. Under a fixed evaluation protocol, experiments on mainstream LLMs show that H-Scale generally improves a broad range of NVFP4 baselines and brings several variants closer to the BF16 reference.
Chinese Translation
NVIDIA Blackwell架构原生支持超细粒度的NVFP4格式,为加速大语言模型(LLM)推理开辟了新的机会。NVFP4的微块设计,例如组大小为16,提供了强大的表示灵活性,以捕捉局部权重分布并隔离异常值,但同时也引入了一个庞大且高度敏感的每组缩放因子空间。现有的后训练量化(PTQ)方法主要集中于细化量化权重值,而对这一缩放选择步骤的研究相对较少。为了解决这一问题,我们提出了 extbf{H-Scale},一种轻量级的后处理方法,用于NVFP4每组尺度的细化。H-Scale并不是通过最小化简单的权重重构误差,而是使用从校准激活中推导的对角二阶代理选择硬件有效的组尺度,从而更直接地针对层输出扰动。它被设计为在多种NVFP4管道中替代RTN风格的尺度选择,仅需适度的离线校准,并在推理时引入严格的零开销。在固定的评估协议下,对主流LLM的实验表明,H-Scale通常能改善广泛的NVFP4基线,并使多个变体更接近BF16参考。
cs.CL / 38 / 2608.28151
Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result
嵌套字节级词汇的部署成本低但共享成本高:一项预注册的负面结果
Abstract
A byte-level BPE tokenizer is an ordered list of merge rules, so applying only a prefix yields a vocabulary whose token identifiers are the first rows of the full vocabulary. This prefix nesting allows one language model to operate at several vocabulary sizes, use a control token to indicate the active size, and be deployed at any trained size by slicing its embedding and output head. We pre-registered five claims, including margins, seeds, contrasts, and a stop rule, and trained 30 models with 3.1M- and 10.6M-parameter bodies on 200M tokens each. Slicing is numerically exact: across 76 checks, a sliced model reproduces the restricted full model's logits bit for bit and removes 66% of deployed weights without changing latency. However, the shared model trails a fixed-cap specialist by 3.64% bits per byte at 32k against a 1% margin, and by 2.96% at 8k against a 2% margin. A 2x2 ablation separating the control token from output restriction finds that the token changes performance by +0.07% to +0.13%, with all intervals crossing zero, while output restriction costs +0.47% to +1.19%; the factors are substitutes rather than complements. Multi-cap training nevertheless improves robustness: under typographical noise, the same checkpoint degrades 12.5--15.4 points less in its fine mode and outperforms each fixed-cap specialist at that specialist's vocabulary size. A control with neither cap token nor output restriction is equally robust, attributing this benefit to multi-granularity training rather than conditioning. The per-cap penalty tracks each cap's share of training rows, yielding a falsifiable prediction for future work.
Chinese Translation
字节级 BPE 分词器是一个合并规则的有序列表,因此仅应用前缀会产生一个词汇,其标识符为完整词汇的前几行。这种前缀嵌套允许一个语言模型在多个词汇大小上操作,使用控制标记指示活动大小,并通过切片其嵌入和输出头在任何训练大小上部署。我们预注册了五个声明,包括边际、种子、对比和停止规则,并在每个 2 亿个标记上训练了 30 个模型,参数规模分别为 310 万和 1060 万。切片在数值上是精确的:在 76 次检查中,切片模型逐位重现了受限完整模型的 logits,并在不改变延迟的情况下去除了 66% 的部署权重。然而,共享模型在 32k 时比固定容量专家低 3.64% 的比特,边际为 1%;在 8k 时低 2.96%,边际为 2%。一个 2x2 的消融实验将控制标记与输出限制分开,发现标记的性能变化为 +0.07% 到 +0.13%,所有区间均穿越零,而输出限制的成本为 +0.47% 到 +1.19%;这些因素是替代关系而非互补关系。尽管如此,多容量训练提高了鲁棒性:在排版噪声下,相同的检查点在其细化模式下降低 12.5-15.4 分,并在该专家的词汇大小上超越每个固定容量专家。一个既没有容量标记也没有输出限制的控制组同样鲁棒,将这一好处归因于多粒度训练而非条件化。每个容量的惩罚与每个容量在训练行中的份额相关,为未来的研究提供了可证伪的预测。
cs.CL / 39 / 2608.28155
FinExam-10K: When Retrieval Helps Financial Reasoning?
FinExam-10K:检索如何帮助金融推理?
Abstract
Professional financial examinations require models to combine domain knowledge, calculation, and judgment, yet no benchmark covers the full CFA and FRM structure under one protocol. We introduce FinExam-10K, to our knowledge the largest reported English benchmark for this setting, with 10,198 expert-reannotated questions spanning CFA Levels I-III and FRM Parts I-II. We release 5,110 questions and sequester 5,088 for a quarterly maintained leaderboard. To separate coverage from local answerability, we report a 10,198-item Full-Coverage Track and a 7,625-item Context-Complete Reasoning Track, which is the primary basis for claims about reasoning from the supplied record. Across 17 models, the best accuracy is 85.29% overall. On the frozen Hard band, the best score is 34.68% on the Full-Coverage Track and 54.57% on the 372 context-complete items. All 17 models share 47 context-complete failures. Function-RAG and FunctionGraph-RAG rescue hundreds of errors but also overturn many correct answers, producing little or negative net gain. A gate trained only on public data decides from the question and initial response when FunctionGraph-RAG should run. On the 5,088 held-out items, the gate invokes FunctionGraph-RAG for 7.9% of questions and improves accuracy from 70.83% to 71.23% (p = .0446).
Chinese Translation
专业的金融考试要求模型结合领域知识、计算和判断,但目前没有一个基准能够在一个协议下涵盖完整的CFA和FRM结构。我们介绍了FinExam-10K,这是我们所知的在这一领域最大的英文基准,包含10,198个专家重新标注的问题,涵盖CFA一级至三级和FRM第一部分至第二部分。我们发布了5,110个问题,并将5,088个问题保留用于季度维护的排行榜。为了将覆盖范围与局部可回答性区分开来,我们报告了一个包含10,198个项目的全面覆盖轨道和一个包含7,625个项目的上下文完整推理轨道,后者是关于从提供记录中推理的主张的主要依据。在17个模型中,最佳准确率为85.29%。在冻结的困难组中,最佳得分在全面覆盖轨道上为34.68%,在372个上下文完整项目上为54.57%。所有17个模型共享47个上下文完整的失败。Function-RAG和FunctionGraph-RAG修正了数百个错误,但也推翻了许多正确答案,产生了很小或负的净收益。仅基于公共数据训练的门控模型根据问题和初始响应决定何时运行FunctionGraph-RAG。在5,088个保留项目上,该门控模型对7.9%的问题调用FunctionGraph-RAG,并将准确率从70.83%提高到71.23%(p = .0446)。
cs.CL / 40 / 2608.28170
Text Restoration of Ancient Documents with Language Models
利用语言模型对古文献进行文本修复
Abstract
Purpose - This study investigates the feasibility of restoring missing text caused by physical lacunae in damaged ancient manuscripts using language models. Methodology - The study proposes different scenarios to replicate real-world conditions. Language models of different architectures are applied according to their suitability to each scenario. We also propose several decoding strategies that further enhance performance and address the discrepancy between lacuna boundaries and the models' tokenization schemes. Findings - The results reveal that text restoration of these documents cannot be fully automated, but it can serve as a useful tool to assist paleographers in their work. Model performance varies greatly depending on which structural part of the document needs to be restored and whether the character length of missing text is available. Originality - This is the first study and to analyze model performance on formulaic and non-formulaic content and the impact of lacuna length awareness in manuscript restoration. Both are recurring challenges in paleographers' manual restoration work. Through systematic comparison and both qualitative and quantitative analysis of different models' performance under varying settings, this study offers a guideline for developing assistive tools to support paleographers.
Chinese Translation
目的 - 本研究探讨了利用语言模型修复因物理缺失而导致的古代手稿缺失文本的可行性。方法 - 本研究提出了不同的场景以模拟现实条件。根据每种场景的适用性,应用不同架构的语言模型。我们还提出了几种解码策略,以进一步提升性能,并解决缺失边界与模型分词方案之间的差异。发现 - 结果表明,这些文献的文本修复无法完全自动化,但可以作为辅助古文字学家工作的有用工具。模型性能在很大程度上取决于需要修复的文档结构部分,以及缺失文本的字符长度是否可用。原创性 - 这是首个分析模型在公式化和非公式化内容上的性能,以及缺失长度意识对手稿修复影响的研究。这两者都是古文字学家手动修复工作中反复出现的挑战。通过对不同模型在不同设置下性能的系统比较以及定性和定量分析,本研究为开发支持古文字学家的辅助工具提供了指导。
cs.CL / 41 / 2608.28283
Embedding Models for Stance-Aware Argument Retrieval
面向立场的论证检索的嵌入模型
Abstract
In computational argumentation, obtaining arguments that explicitly support or attack given claims is a critical precursor to downstream reasoning tasks. When these supporting and attacking arguments are to be retrieved using semantic search methods, they need to be assessed for topic-relevance to the claims of interest as well as for correctness of their (positive or negative) stance towards the claims. In this paper we explore how dense embedding models (hereafter, models), powering modern retrieval pipelines, can serve as the basis of semantic search incorporating this dual assessment. We show experimentally that existing models struggle with asymmetric reasoning, exhibiting a strong bias toward topical overlap while ignoring instructional stance. We also show that correcting this bias via contrastive training triggers a new failure mode where models over-correct, over-fixating on polarity keywords (e.g., "supports" or "refutes") at the expense of the semantic topic. We thus introduce diagnostic word-ablation metrics to quantify this phenomenon and propose a data-centric solution. By implementing a balanced argument curriculum alongside LLM-augmented, stance-inverted arguments, we force the (embedding) models to learn deeper directional logic rather than exploiting superficial lexical shortcuts. Our evaluation demonstrates that, for sufficiently powerful models, this approach can alleviate the observed overcorrection, achieving further improvements in stance-aware argument retrieval.
Chinese Translation
在计算论证中,获取明确支持或攻击特定主张的论证是下游推理任务的关键前提。当这些支持和攻击论证通过语义搜索方法进行检索时,需要评估它们与相关主张的主题相关性以及它们对主张的(正面或负面)立场的正确性。本文探讨了如何利用密集嵌入模型(以下简称模型)作为现代检索管道的基础,以实现包含这种双重评估的语义搜索。我们通过实验表明,现有模型在处理不对称推理时存在困难,表现出对主题重叠的强烈偏见,同时忽视了指导性立场。我们还展示了通过对比训练纠正这种偏见会引发一种新的失败模式,即模型过度纠正,过度关注极性关键词(例如,“支持”或“反驳”),而忽视语义主题。因此,我们引入了诊断性词汇消融指标来量化这一现象,并提出了一种以数据为中心的解决方案。通过实施平衡的论证课程以及增强的、立场反转的论证,我们迫使(嵌入)模型学习更深层次的方向逻辑,而不是利用表面的词汇捷径。我们的评估表明,对于足够强大的模型,这种方法可以缓解观察到的过度纠正,实现立场感知论证检索的进一步改善。
cs.CL / 42 / 2608.28293
A Probabilistic Interpretation of KV Cache Eviction
KV 缓存驱逐的概率解释
Abstract
The premise and promise of KV (cache) eviction is simple: higher throughput can be achieved by evicting some entries from the KV cache, at a negligible cost to quality. This holds empirically for many existing methods, though most rely on creative heuristics for selecting which entries to drop. Despite recent advances, the problem of KV eviction has remained informal in the literature. This paper aims to properly formalize this problem through the lens of probabilistic reasoning and reveal what can be learned from this perspective. Concretely, we (1) formalize the problem of KV eviction and, unfortunately, prove that it is computationally hard, (2) show that by framing it probabilistically, KV eviction reduces to the problem of expectation estimation, which can be approximated through sampling, (3) show that through this probabilistic interpretation, correcting for evicted entries during decoding---a previously ignored problem---becomes feasible, and (4) reveal that existing methods in the literature are zero-variance biased estimators that can be easily adapted in order to enable decode time correction. In practice, we show that this probabilistic version of KV eviction coupled with decode time correction is more robust to different tasks compared to existing eviction methods and achieves competitive performance at the same compression budget.
Chinese Translation
KV(键值)缓存驱逐的前提和承诺很简单:通过驱逐一些条目,可以在对质量几乎没有影响的情况下实现更高的吞吐量。这在许多现有方法中得到了实证支持,尽管大多数方法依赖于创造性的启发式算法来选择哪些条目被驱逐。尽管最近取得了一些进展,KV 驱逐问题在文献中仍然保持非正式。本文旨在通过概率推理的视角对这一问题进行适当的形式化,并揭示从这一角度可以学到的内容。具体而言,我们 (1) 形式化了 KV 驱逐问题,并不幸地证明它在计算上是困难的;(2) 通过将其框定为概率问题,KV 驱逐简化为期望估计问题,可以通过采样进行近似;(3) 通过这种概率解释,纠正解码过程中被驱逐条目——一个之前被忽视的问题——变得可行;(4) 揭示了文献中现有方法是零方差偏差估计器,可以轻松调整以实现解码时的纠正。在实践中,我们展示了这种概率版本的 KV 驱逐结合解码时的纠正相比于现有的驱逐方法在不同任务上更具鲁棒性,并在相同的压缩预算下实现了具有竞争力的性能。
cs.CL / 43 / 2608.28329
BanglaMed-QA: A Question Answering System for Healthcare Support in Bangla
BanglaMed-QA:用于孟加拉语医疗支持的问答系统
Abstract
Medical question answering (QA) systems have become crucial tools for providing reliable health information. But they remain very unexplored for low-resource languages like Bangla due to limited datasets and systems tailored to these languages. To address this, we introduce BanglaMed-QA, a robust QA system specifically designed for the Bangla medical domain. The process begins with building a structured medical knowledge base that includes 4,493 QA pairs in 9 categories under 506 diseases. To improve semantic comprehension, domain-specific root word dictionaries and synonym sets are proposed, in addition to part-of-speech tagging for anaphora resolution. We adopt supervised machine learning models in which SVM is found to be the best model to categorize questions. Multiple similarity metrics, including cosine, Jaccard, BM25, and Levenshtein, are applied with soft and hard voting methods for query matching. The performance of the QA system has been evaluated in two aspects, with a 95% F1 score in an automated evaluation and an average human satisfaction rating of 0.9 out of 1.0. This validates the real-world application of BanglaMed-QA in closing the healthcare information gap for Bangla speakers.
Chinese Translation
医疗问答(QA)系统已成为提供可靠健康信息的重要工具。然而,由于缺乏针对这些语言的有限数据集和系统,低资源语言如孟加拉语的研究仍然非常不足。为了解决这一问题,我们提出了BanglaMed-QA,一个专门为孟加拉医疗领域设计的强大问答系统。该过程始于构建一个结构化的医学知识库,其中包含在506种疾病下的9个类别中的4,493个问答对。为了提高语义理解,除了为指代消解提供词性标注外,还提出了特定领域的词根字典和同义词集。我们采用监督学习模型,其中支持向量机(SVM)被发现是分类问题的最佳模型。多个相似度度量,包括余弦相似度、杰卡德相似度、BM25和莱温斯坦距离,结合软投票和硬投票方法用于查询匹配。该问答系统的性能从两个方面进行了评估,在自动评估中获得了95%的F1分数,平均人类满意度评分为0.9(满分1.0)。这验证了BanglaMed-QA在弥补孟加拉语使用者医疗信息差距方面的实际应用价值。
cs.CL / 44 / 2608.28378
PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems
PersonaForge:用于代理系统的真实多轮用户模拟
Abstract
Large language models are increasingly used as agentic workflow executors, yet existing training data and benchmarks largely assume informationally complete, single-turn queries. Our analysis of 16K real-world sessions shows that 75.9% of interactions are multi-turn, revealing a substantial gap between how users interact with agents and how such systems are trained and evaluated. We introduce \textbf{PersonaForge}, a user simulation framework for synthesizing realistic multi-turn user--agent interactions. PersonaForge combines a four-dimensional persona space, SOUL-driven behavioral control calibrated to real-user statistics, and Reverse Deep Construction grounded in authentic seed queries. Using PersonaForge, we construct a 6.3K-record training dataset and \textbf{PersonaForge-Bench}, a manually annotated 138-task benchmark spanning over 20 professional domains with four-dimensional scoring. Experiments on Qwen3.5-27B show that PersonaForge training improves the composite score by +4.1%, with gains across all four dimensions and the largest improvements in Task Completion (+6.0%) and Response Quality (+6.8%). Further analyses show that PersonaForge-trained agents use fewer turns and tool calls, suggesting improved interaction efficiency, while ablations confirm the contribution of SOUL components and adaptive simulation. Together, PersonaForge and PersonaForge-Bench establish a foundation for training and evaluating agents under realistic multi-turn user interaction.
Chinese Translation
大型语言模型越来越多地被用作代理工作流执行者,但现有的训练数据和基准大多假设信息是完整的、单轮查询。我们对16,000个真实世界会话的分析表明,75.9%的交互是多轮的,这揭示了用户与代理的交互方式与这些系统的训练和评估方式之间存在显著差距。我们引入了 extbf{PersonaForge},一个用于合成真实多轮用户-代理交互的用户模拟框架。PersonaForge结合了四维角色空间、基于真实用户统计的SOUL驱动行为控制,以及基于真实种子查询的反向深度构建。使用PersonaForge,我们构建了一个包含6,300条记录的训练数据集和 extbf{PersonaForge-Bench},这是一个手动标注的138任务基准,涵盖20多个专业领域,并具有四维评分。对Qwen3.5-27B的实验表明,PersonaForge训练使综合得分提高了4.1%,在所有四个维度上均有提升,其中任务完成度提高了6.0%,响应质量提高了6.8%。进一步分析显示,经过PersonaForge训练的代理使用了更少的轮次和工具调用,表明交互效率得到了改善,而消融实验确认了SOUL组件和自适应模拟的贡献。总之,PersonaForge和PersonaForge-Bench为在真实多轮用户交互下训练和评估代理奠定了基础。
cs.CL / 45 / 2608.28382
When Linguistic and Internal Confidence Diverge in Large Language Models
当语言信心与大型语言模型的内部信心出现偏差时
Abstract
Users often ask large language models (LLMs) to report how confident they are, but it is unclear whether such linguistic confidence tracks the model's internal confidence. We study this question across 8 classification tasks, 2 generation tasks and 30 models from three families. For classification, we compare linguistic confidence with logits-based confidence along three axes: association, magnitude agreement and calibration. For generation, we test whether linguistic confidence tracks semantic-entropy-based uncertainty. The axes frequently diverge. Instance-level association is weak on average, although it improves on easier items and for stronger base models. Instruction-tuned models often report higher confidence and sometimes show higher association, but they also have larger confidence gaps and worse calibration. Prompt design mostly changes the distribution of reported confidence. Attitude cues inflate confidence without improving alignment, while score exemplars can preserve rank-order signal when they avoid collapsed confidence values. Regression analyses show that distributional properties of confidence scores explain much of the observed alignment pattern, with model metadata playing a smaller role after controls. These results support a lossy-channel view of linguistic confidence. A more dispersed verbal confidence distribution can carry useful rank information, but it does not make the scores calibrated. Linguistic confidence should therefore be evaluated with multi-axis diagnostics before being used in downstream reliability pipelines.
Chinese Translation
用户经常要求大型语言模型(LLMs)报告它们的信心程度,但尚不清楚这种语言信心是否与模型的内部信心相一致。我们在8个分类任务、2个生成任务和来自三个家族的30个模型中研究了这个问题。对于分类任务,我们沿着三个维度比较语言信心与基于logits的信心:关联性、幅度一致性和校准性。对于生成任务,我们测试了语言信心是否跟踪基于语义熵的不确定性。这些维度经常出现偏差。实例级的关联性平均较弱,尽管在较简单的项目和更强的基础模型上有所改善。经过指令调优的模型通常报告更高的信心,并且有时显示出更高的关联性,但它们也存在更大的信心差距和较差的校准性。提示设计主要改变报告信心的分布。态度线索会膨胀信心而不改善一致性,而评分示例在避免崩溃的信心值时可以保留排名信号。回归分析表明,信心分数的分布特性解释了观察到的一致性模式的大部分,而模型元数据在控制后起的作用较小。这些结果支持了语言信心的有损信道视角。更分散的语言信心分布可以传递有用的排名信息,但并未使得分数得到校准。因此,在用于下游可靠性流程之前,语言信心应通过多维诊断进行评估。
cs.CL / 46 / 2608.28405
CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia
文化对话:一个多语言多轮模拟工具,用于东亚和东南亚的文化基础援助
Tan, Bryan Chen Zhengyu, Zheng, Weihua, Doan, Thong T., Doan, Bich Ngoc, Peh, Jia Wang, Yi, Xiaoyuan, Yao, Jing, Xie, Xing, Chen, Nancy F., Liu, Zhengyuan, Bak, JinYeong, Shamdi, Wafi, Chie, Soo Kai, Siong, Liew Yu, Rezal, Aina Azyyati Binti Mohamad, Vanessa, Lew Yan Yan, Wu, Huadan, Raharja, Dylan, Wangsajaya, Nadya Yuki, Fukushige, Akane, Kato, Kazushi, Inoue, Koji, Kawahara, Tatsuya, Seo, Jaehyung, Kim, Dongjun, Lee, Seungyoon, Pang, Zi Haur, Tan, Rui Yang, Cheng, Charibeth Ko, Estuar, Maria Regina Justina, Montalan, Jann Railey, Duc, Pham Minh, Lee, Roy Ka-Wei
Abstract
Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common use case: users seeking practical help over multiple turns in culturally grounded scenarios. We introduce CultureConverse, a scalable, multilingual simulation and evaluation harness for culturally grounded assistant dialogue that covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains. Each simulated and evaluated episode produces a scored interaction where the assistant assists the user and infers cultural constraints from partial information. The resulting CultureConverse-DS dataset contains 14,610 benchmark (evaluation) episodes and 274,295 oracle-guided (gold-mode) dialogues. In our benchmark evaluation of 18 models, GPT-5 mini achieves the highest assistance quality. Human annotation experiments suggest that our evaluation framework is a sufficient proxy for human judgment. Performance gains from fine-tuning on 27,860 high-quality CultureConverse-DS samples improve in-domain assistance and transfer out-of-domain to cultural MCQ and safety classification benchmarks. We release the harness, both splits, and judge prompts to support interactive evaluation of cultural competency.
Chinese Translation
目前对大型语言模型(LLMs)的文化评估通常将文化简化为通过多项选择题(MCQs)进行的单轮事实回忆,未能捕捉到一个常见的使用场景:用户在文化基础场景中寻求多轮的实际帮助。我们介绍了CultureConverse,一个可扩展的多语言模拟和评估工具,用于文化基础的助手对话,覆盖10个东亚和东南亚地区、58个子群体身份和7个领域。每个模拟和评估的事件生成一个评分交互,其中助手帮助用户并从部分信息中推断文化限制。最终生成的CultureConverse-DS数据集包含14,610个基准(评估)事件和274,295个由神谕指导(黄金模式)对话。在对18个模型的基准评估中,GPT-5 mini实现了最高的援助质量。人类注释实验表明,我们的评估框架是人类判断的一个充分代理。从27,860个高质量CultureConverse-DS样本的微调中获得的性能提升改善了领域内的援助,并向文化多项选择题和安全分类基准转移。我们发布了该工具、两个数据集划分以及评估提示,以支持文化能力的互动评估。
cs.CL / 47 / 2608.28407
A Unified Framework to Elicit Structured Feedback for Interpretable Multi-Trait Essay Scoring
一个统一框架以引导可解释的多特征作文评分的结构化反馈
Abstract
Multi-trait Automated Essay Scoring (AES) requires rubric-grounded reasoning across interdependent traits, rather than isolated score prediction. Existing feedback-enhanced methods often decouple feedback from scoring or assess traits independently, weakening score--feedback consistency and rubric alignment. We propose HiFTS, a unified autoregressive framework that generates hierarchical CoT feedback before predicting trait-level and holistic scores. HiFTS distills rubric-grounded hierarchical CoT feedback from a teacher LLM and trains student models to jointly generate feedback and scores. HiFTS further applies Group Relative Policy Optimization with a composite reward balancing score agreement, calibration, feedback quality, and structural validity. At inference, a lightweight global prior provides holistic guidance to reduce drift during long-form reasoning. We also introduce CFMS-34, a Chinese multi-trait AES dataset with 951 essays annotated with holistic scores and 34 rubric-based traits. Experiments on CFMS-34 and ASAP++ show that HiFTS achieves strong holistic and trait-level scoring while producing coherent, rubric-aligned feedback.
Chinese Translation
多特征自动作文评分(AES)需要基于评分标准的推理,涉及相互依赖的特征,而非孤立的分数预测。现有的增强反馈方法往往将反馈与评分解耦或独立评估特征,这削弱了分数与反馈的一致性和评分标准的对齐。我们提出了HiFTS,一个统一的自回归框架,在预测特征级和整体分数之前生成层次化的链式思维(CoT)反馈。HiFTS 从教师大语言模型(LLM)中提炼基于评分标准的层次化 CoT 反馈,并训练学生模型共同生成反馈和分数。HiFTS 进一步应用组合奖励的群体相对策略优化(Group Relative Policy Optimization),平衡分数一致性、校准、反馈质量和结构有效性。在推理时,轻量级的全局先验提供整体指导,以减少长篇推理过程中的漂移。我们还引入了CFMS-34,一个包含951篇带有整体分数和34个基于评分标准特征的中文多特征AES数据集。在CFMS-34和ASAP++上的实验表明,HiFTS在生成一致且与评分标准对齐的反馈的同时,实现了强大的整体和特征级评分。
cs.CL / 48 / 2608.28432
Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQL
这些模块值得它们的成本吗?一种范式级别的准确性-成本分析在上下文学习文本到SQL中
Abstract
Recent advances in in-context learning (ICL) text-to-SQL have substantially improved execution accuracy on public benchmarks by assembling increasingly elaborate pipelines around the base generator, yet existing studies typically report aggregate end-to-end accuracy, without quantifying the marginal accuracy-cost contribution of individual design choices. Consequently, providing a unified, paradigm-level cost-accuracy quantification remains a critical challenge for understanding and configuring modern text-to-SQL. To address this, we instantiate 17 paradigm-level configurations across five recurring modules of the ICL text-to-SQL pipeline under a single controlled implementation, and attribute each paradigm's marginal contribution and incurred cost across all four backbones spanning diverse capability levels and reasoning styles. Our analysis reveals that execution-feedback refinement is the only paradigm whose benefit holds universally at consistently low cost, while most other modules help only under backbone-dependent conditions. Token accounting shows that input demand is more closely tied to pipeline structure, whereas output demand is more sensitive to backbone generation behavior. Cross-module analysis further shows that stacking improves accuracy on most backbones, although how the gains compose varies with backbone capability. We also find that a fixed budget is often better spent engineering a more elaborate pipeline over a mid-tier backbone than upgrading to a frontier model with a lean pipeline. These findings distill into an actionable, cost-aware tiered guideline that transfers to five additional backbones without per-paradigm search.
Chinese Translation
最近在上下文学习(ICL)文本到SQL方面的进展,通过围绕基础生成器组装日益复杂的管道,显著提高了公共基准上的执行准确性。然而,现有研究通常报告整体的端到端准确性,而未量化单个设计选择的边际准确性-成本贡献。因此,提供统一的范式级别成本-准确性量化仍然是理解和配置现代文本到SQL的关键挑战。为了解决这个问题,我们在单一受控实现下实例化了17种范式级别配置,涵盖ICL文本到SQL管道中的五个重复模块,并对所有四个涵盖不同能力水平和推理风格的基础模型进行每个范式的边际贡献和产生的成本归因。我们的分析揭示,执行反馈优化是唯一一个在一致低成本下普遍有效的范式,而大多数其他模块仅在依赖于基础模型的条件下提供帮助。令牌计数显示,输入需求与管道结构关系更密切,而输出需求则对基础生成行为更为敏感。跨模块分析进一步表明,堆叠在大多数基础模型上提高了准确性,尽管增益的组合方式因基础能力而异。我们还发现,在固定预算下,通常在中等基础模型上构建更复杂的管道比升级到具有精简管道的前沿模型更为划算。这些发现提炼出一个可操作的、成本意识的分层指导方针,可以转移到五个额外的基础模型,而无需逐范式搜索。
cs.CL / 49 / 2608.28439
Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction
忠诚度不足:代理数据表提取的调度级仪器化
Abstract
One model passed our fidelity check without ever opening the datasheet. We found it while qualifying models for an internal extraction service: a structured-output constraint had silently disabled tool use, and the model answered anyway, with fabricated source text. Only the per-tool trace exposed it. Fidelity -- whether an extracted value matches the source -- is the standard measure for agentic document extraction, and it scores that run a success. We therefore log every tool call in an agentic benchmark of 25 hand-curated claims over three components, with 12 more on a fourth, 37 in all. From that dispatch record we build two instruments: a rule-based failure-attribution classifier, and a silent-failure detector whose two rules check only which tools were called, never the extracted value. The detector raises no flag on 207 clean fidelity-passing extractions across three model families, and recovers all 50 planted faults that withhold exactly the tools its rules check. The two results are not symmetric: the first bounds the false-positive rate, the second is recall by construction, and detection power against runs that call their tools and still answer wrongly is unmeasured. A second, independent oracle, a causal chamber that tests whether the datasheet's claims hold under physical measurement, is intentionally partial: it confirms only what the apparatus can exercise, a verifiable envelope of 2 of those 37 claims, and we give a taxonomy of why the rest are not physically gradable. Under a controlled perturbation, fidelity passes throughout while the chamber verdict flips exactly at the measurement uncertainty. Across three deployed model stacks (one destabilised by its serving stack, not by any capability gap) the tool layer buys portability and observability rather than accuracy, and earns its premium only once a document outgrows the context window.
Chinese Translation
有一个模型在未打开数据表的情况下通过了我们的忠诚度检查。我们在为内部提取服务资格审查模型时发现了它:一个结构化输出约束无声地禁用了工具使用,而模型仍然给出了答案,且源文本是虚构的。只有每个工具的追踪记录揭示了这一点。忠诚度——即提取的值是否与源匹配——是代理文档提取的标准衡量指标,并且它将该运行评估为成功。因此,我们在一个包含25个手动策划声明的代理基准中记录每个工具调用,涵盖三个组件,第四个组件上还有12个,总共37个。从该调度记录中,我们构建了两个工具:一个基于规则的失败归因分类器和一个静默失败检测器,其两个规则仅检查调用了哪些工具,而从不检查提取的值。该检测器在三个模型系列中对207个干净的忠诚度通过提取未发出警报,并恢复了所有50个植入的故障,这些故障正好抑制了其规则检查的工具。这两个结果并不对称:第一个限制了假阳性率,第二个是通过构造实现的召回,而对调用其工具并仍然错误回答的运行的检测能力尚未测量。第二个独立的神谕,一个因果室,测试数据表的声明在物理测量下是否成立,故意是部分的:它仅确认该设备能够测试的内容,即37个声明中可验证的2个封闭范围,并且我们给出了其余声明无法物理评分的分类。在受控扰动下,忠诚度在整个过程中通过,而室内判决恰好在测量不确定性处翻转。在三个已部署的模型堆栈中(一个因其服务堆栈而不稳定,而不是由于任何能力差距),工具层购买了可移植性和可观察性,而不是准确性,只有当文档超出上下文窗口时才获得其溢价。
cs.CL / 50 / 2608.28444
Sliding-window beats linear attention
滑动窗口超越线性注意力
Abstract
Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy. Every new token costs more than the previous one. For each additional token, the keys and values must be stored in memory indefinitely, which is unsustainable. Several alternatives have been proposed to fix the quadratic scaling problem, one of which is retrofitting LLMs to use Linear Attention. This idea has attracted a lot of attention, given its promise to solve the quadratic scaling problem with state-of-the-art performance at low cost. However, this line of research has not been properly compared to simpler baselines. In this work, we show that Sliding Window Attention (SWA) with sinks performs as well or better than post-trained Linear Attention models. We observe this across multiple LLMs on various downstream tasks. For long-context reasoning tasks (Needle-in-a-Haystack and BABILong), SWA achieves massively higher performance (2 to 10 times higher than linear attention). SWA requires no post-training, is extremely fast, and requires low memory; therefore, making it an extremely cheap and reliable solution. To reduce inference memory cost, we strongly recommend switching to SWA instead of post-training linear models. Linear attention models may have shown some promise, but they likely require to be trained from scratch or extensive post-training in order to even match SWA.
Chinese Translation
由于二次注意力的特性,大型语言模型(LLMs)消耗了大量的内存和能源。每个新令牌的成本都高于前一个。对于每个额外的令牌,键和值必须在内存中无限期存储,这种做法不可持续。为了解决二次扩展问题,提出了几种替代方案,其中之一是改造LLMs以使用线性注意力。考虑到其在低成本下解决二次扩展问题的潜力,这一想法引起了广泛关注。然而,这一研究方向尚未与更简单的基线进行适当比较。在本研究中,我们展示了带有汇聚的滑动窗口注意力(Sliding Window Attention, SWA)在性能上与后训练的线性注意力模型相当或更优。我们在多个LLMs和各种下游任务中观察到了这一点。在长上下文推理任务(如Needle-in-a-Haystack和BABILong)中,SWA的性能显著提高(比线性注意力高出2到10倍)。SWA不需要后训练,速度极快,内存需求低,因此成为了一种极其廉价和可靠的解决方案。为了降低推理内存成本,我们强烈建议转向SWA,而不是后训练线性模型。线性注意力模型可能显示出一些潜力,但它们可能需要从头开始训练或进行广泛的后训练,才能与SWA相匹配。
cs.CL / 51 / 2608.28458
Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents
获取、修复、保存:一种基于诊断的后训练小型模型对话游戏代理的方案
Abstract
Interactive dialogue games test a capability that static benchmarks largely leave implicit: a model must carry state across turns, interpret feedback, and choose valid actions under changing constraints. We study this setting in the LM Playschool Challenge with a 2B open-weight model, and find that many failures are not only broad knowledge failures but also local decision failures: repeated guesses, malformed actions, and violations of feedback that the model has just seen. These diagnostics motivate a training recipe organized around three steps: acquire broad game participation through supervised fine-tuning, repair mechanically verifiable failures within one targeted dialogue-game family using turn-local preference pairs, and preserve general capabilities beyond these dialogue games. In the official final evaluation, our submission improves public clemscore from 10.67 to 38.92 and closed in-domain score from 13.41 to 41.17, while approximately preserving aggregate static performance (44.14 vs. 44.24 for the baseline). Out-of-domain clemscore remains low at 7.88, with the largest gains concentrated in unseen variants of the targeted family. Our results suggest that broad SFT brings most of the model's capability improvement; turn-local supervision can be effective when failure detection is precise, with observed transfer concentrated primarily within-family.
Chinese Translation
互动对话游戏测试了一种静态基准大多隐含的能力:模型必须在回合间保持状态,解释反馈,并在不断变化的约束下选择有效的行动。我们在 LM Playschool Challenge 中研究这一设置,使用一个 2B 开放权重模型,发现许多失败不仅是广泛知识的失败,还有局部决策的失败:重复猜测、格式错误的行动以及违反模型刚刚看到的反馈。这些诊断结果促使我们提出一个围绕三个步骤组织的训练方案:通过监督微调获取广泛的游戏参与,利用回合局部偏好对修复可机械验证的失败,针对一个特定的对话游戏家族进行修复,并在这些对话游戏之外保持一般能力。在官方最终评估中,我们的提交将公共 clemscore 从 10.67 提升至 38.92,闭合领域得分从 13.41 提升至 41.17,同时大致保持了整体静态性能(基线为 44.14 vs. 44.24)。领域外的 clemscore 仍然较低,为 7.88,最大的增益集中在未见过的目标家族变体中。我们的结果表明,广泛的监督微调(SFT)带来了模型能力的主要提升;当失败检测精确时,回合局部监督可以有效,观察到的转移主要集中在同一家族内。
cs.CL / 52 / 2608.28467
Stranger, Fan, or Peer? A Systematic Study on the Role of Interlocutor in Persona-Based Dialogue Generation
陌生人、粉丝还是同伴?基于角色的对话生成中对话者角色的系统性研究
Abstract
Persona-based dialogue systems are usually conditioned on speaker biography, but dialogues involve at least two participants, and who has access to whose biography can vary across training, inference, and evaluation. Prior work often neglected these aspects, obscuring mechanisms that only appear when biography visibility is toggled separately across training, inference, and evaluation, a three-stage factorisation that prior work has largely treated as a single factor. We study this factorisation on a dataset of dialogues paired with speaker's biographies, varying whether the target and interlocutor speakers see each other's biographies during training and inference, and using an LLM as a judge to perform author identification. We find that (i) training-time visibility, more than inference-time visibility, determines whether models express persona traits through dialogue or fall back on copying biographical text (a known problem/phenomenon in persona-based generation); (ii) models trained with interlocutor-biography visibility copy less target-biographical text than models trained without it, while changing visibility only at inference time has a less consistent effect; and (iii) under asymmetric disclosure, where only the interlocutor sees the target biography, target content leaks into interlocutor turns more often, and dialogues containing such traces are easier for the judge to identify, especially when interlocutor turns are visible. These results suggest that biography leakage into generated turns is an artefact of how interlocutor visibility is configured across training and inference, and separating the three stages is necessary.
Chinese Translation
基于角色的对话系统通常依赖于说话者的个人简介,但对话至少涉及两名参与者,且谁能访问谁的个人简介在训练、推理和评估阶段可能各不相同。以往研究常忽视这些方面,掩盖了仅在训练、推理和评估三个阶段分别切换个人简介可见性时才出现的机制,而此前研究大多将这三个阶段视为单一因素。我们在一个包含对话及说话者个人简介的数据集上研究了这一因素分解,变更目标说话者和对话者在训练和推理阶段是否能看到彼此的个人简介,并使用大型语言模型(LLM)作为评判者进行作者身份识别。研究发现:(i) 训练阶段的可见性比推理阶段的可见性更决定模型是通过对话表达角色特征,还是退而求其次复制个人简介文本(这是基于角色生成中的已知问题/现象);(ii) 在训练时对对话者个人简介可见的模型比训练时不可见的模型更少复制目标个人简介文本,而仅在推理阶段改变可见性则效果不够一致;(iii) 在非对称披露情况下,即只有对话者能看到目标个人简介时,目标内容更频繁地泄露到对话者的发言中,且包含此类痕迹的对话更易被评判者识别,尤其是在对话者发言可见时。上述结果表明,个人简介泄露到生成发言中是对话者可见性在训练与推理阶段配置方式的产物,且必须将这三个阶段分开考虑。
cs.CL / 53 / 2608.28476
ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
ContextPilot:通过细粒度强化学习教导代理进行主动上下文管理
Abstract
Long-horizon agentic tasks require large language models (LLMs) to iteratively retrieve, integrate, and maintain dispersed information across multi-turn interactions, but preserving all interaction histories leads to a continuously growing working context. Recent proactive context management methods allow models to edit their own working context with specialized tools, yet they still face three key limitations: (1) a limited toolset restricted to search, deletion, and summarization, with no support for global planning, long-term memory, and adaptive compression; (2) inefficient exploration that treats context management actions uniformly despite their heterogeneous impacts on final outcomes; and (3) coarse-grained credit assignment that assigns the final trajectory-level reward to all intermediate context editing actions during RL. To bridge these gaps, we introduce ContextPilot, a proactive context management framework for long-horizon agentic reasoning. Our approach systematically augments the toolset with planning, long-term memory, and soft context offloading tools. We further propose an RL method tailored for context management, which uses context and entropy variation to identify critical editing decisions for branch sampling and estimates action-level advantages from all branched trajectories that pass through the corresponding context editing action. Experiments on long-context QA and deep search tasks show that ContextPilot achieves stronger performance with a more compact working context, consistently outperforming existing baselines across various base models and benchmarks. Code is available at https://github.com/Tencent/ContextPilot.
Chinese Translation
长时间跨度的代理任务要求大型语言模型(LLMs)在多轮交互中迭代地检索、整合和维护分散的信息,但保留所有交互历史会导致工作上下文不断增长。近期的主动上下文管理方法允许模型使用专门工具编辑其工作上下文,但仍面临三个主要限制:(1)工具集有限,仅限于搜索、删除和摘要,缺乏全球规划、长期记忆和自适应压缩的支持;(2)低效的探索方式将上下文管理行为视为均匀,尽管它们对最终结果的影响是异质的;以及(3)粗粒度的信用分配,在强化学习(RL)过程中将最终轨迹级奖励分配给所有中间上下文编辑行为。为了解决这些问题,我们提出了ContextPilot,一个用于长时间跨度代理推理的主动上下文管理框架。我们的方法系统性地增强了工具集,加入了规划、长期记忆和软上下文卸载工具。我们进一步提出了一种针对上下文管理的强化学习方法,该方法利用上下文和熵变化来识别关键编辑决策以进行分支采样,并从所有经过相应上下文编辑行为的分支轨迹中估计动作级优势。在长上下文问答和深度搜索任务上的实验表明,ContextPilot在更紧凑的工作上下文下实现了更强的性能,在各种基础模型和基准测试中始终优于现有基线。代码可在 https://github.com/Tencent/ContextPilot 获取。
cs.CL / 54 / 2608.28478
Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge
盲人和大象:探讨大型语言模型在长尾分散知识下的认知短视
Abstract
Factual question answering (QA) typically assumes a single canonical answer, obscuring whether large language models (LLMs) retain divergent accounts of long-tail facts. To address this gap, we introduce ElephantBench, a closed-book knowledge probe comprising 1,094 questions generated through an auditable graph-based pipeline. The pipeline retrieves related documents from a low-exposure web corpus, identifies naturally occurring disagreements, and converts them into multi-account QA records. Each answer is verified against the originating documents and authoritative public web sources and is then reviewed by human annotators. Across 32 models, even the strongest model recovers both accounts on only 52.4% of questions, while on nearly all remaining questions it recalls one account but omits the other. Scaling model size and inference-time reasoning improve recall but do not eliminate this incompleteness. Corpus analysis further shows that exposure imbalance favors the dominant account, whereas greater minority-side exposure is associated with more complete recall. These findings establish ElephantBench as a reproducible knowledge probe for diagnosing epistemic myopia in parametric memory. More broadly, our graph-based benchmark construction pipeline provides an efficient and scalable way to turn long-tail corpora into source-traceable knowledge probes, supporting efforts to evaluate and advance the epistemic rigour of next-generation LLMs. Code is available at https://github.com/Tencent/ElephantBench.
Chinese Translation
事实问答(QA)通常假设存在单一的规范答案,这掩盖了大型语言模型(LLMs)是否保留了长尾事实的不同解释。为了解决这一问题,我们引入了 ElephantBench,这是一个闭卷知识探测工具,包含通过可审计的基于图的流程生成的1,094个问题。该流程从低曝光的网络语料库中检索相关文档,识别自然发生的分歧,并将其转化为多账户问答记录。每个答案都经过与原始文档和权威公共网络来源的验证,并由人工标注者进行审核。在32个模型中,即使是最强的模型也仅在52.4%的问题上恢复了两种解释,而在几乎所有剩余的问题上,它回忆起一种解释但遗漏了另一种。扩大模型规模和推理时间的推理提高了回忆率,但并未消除这种不完整性。语料库分析进一步表明,曝光不平衡有利于主导解释,而较大的少数方曝光与更完整的回忆相关。这些发现确立了 ElephantBench 作为一种可重复的知识探测工具,用于诊断参数记忆中的认知短视。更广泛地说,我们基于图的基准构建流程提供了一种高效且可扩展的方法,将长尾语料库转化为可追溯来源的知识探测工具,支持评估和提升下一代 LLMs 的认知严谨性。代码可在 https://github.com/Tencent/ElephantBench 获取。
cs.CL / 55 / 2608.28481
NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry
NL2AGBench:针对 AlphaGeometry 的 LLM 自动形式化基准测试
Abstract
Recent advances in large language models (LLMs) have demonstrated strong capabilities in natural language understanding and mathematical reasoning. However, their ability to translate informal mathematical problems into formal representations remains underexplored. This limitation is particularly important for neuro-symbolic geometry systems such as AlphaGeometry, whose theorem-proving engine requires inputs in a specialized domain-specific language (DSL). Although AlphaGeometry achieves near-IMO gold-medalist performance, manually converting natural-language problems into its formal syntax remains a significant usability bottleneck. To address this challenge, we introduce the Natural Language to AlphaGeometry Benchmark (NL2AGBench), which evaluates LLMs in translating English geometry problems into AlphaGeometry-compatible formal representations. NL2AGBench uses execution-based verification within AlphaGeometry to assess translation quality rather than relying solely on textual similarity. We evaluate ten state-of-the-art open- and closed-source LLMs across multiple parameter scales and analyze executable translation accuracy, syntactic correctness, and error characteristics. Our experiments reveal a substantial performance gap between closed- and open-source models: leading closed-source models achieve executable translation rates above 80%, while even the largest open-source models struggle to consistently preserve geometric constraints and produce valid formalizations. We introduce an error taxonomy distinguishing syntax and logic errors and investigate mitigation strategies, including few-shot prompting, fine-tuning, and human-guided hinting, which yield measurable improvements across multiple model families.
Chinese Translation
近期大型语言模型(LLMs)的进展展示了其在自然语言理解和数学推理方面的强大能力。然而,它们将非正式数学问题转化为正式表示的能力仍然未得到充分探索。这个限制对于神经符号几何系统如 AlphaGeometry 尤为重要,因为其定理证明引擎需要使用专门的领域特定语言(DSL)作为输入。尽管 AlphaGeometry 的表现接近国际数学奥林匹克(IMO)金牌水平,但手动将自然语言问题转换为其正式语法仍然是一个显著的可用性瓶颈。为了解决这一挑战,我们引入了自然语言到 AlphaGeometry 基准(NL2AGBench),该基准评估 LLM 在将英语几何问题翻译为与 AlphaGeometry 兼容的正式表示方面的能力。NL2AGBench 采用基于执行的验证方法来评估翻译质量,而不仅仅依赖于文本相似性。我们评估了十种最先进的开源和闭源 LLM,涵盖多个参数规模,并分析可执行翻译的准确性、语法正确性和错误特征。我们的实验揭示了闭源模型与开源模型之间的显著性能差距:领先的闭源模型实现了超过 80% 的可执行翻译率,而即使是最大的开源模型在持续保持几何约束和生成有效形式化方面也面临困难。我们引入了一种错误分类法,区分语法错误和逻辑错误,并研究了缓解策略,包括少量示例提示、微调和人类引导提示,这些策略在多个模型系列中均取得了可测量的改进。
cs.CL / 56 / 2608.28496
Ladders in Chaos: When, How, (and Perhaps Why) Does Test-Time Scaling Improve LLM Machine Translation
混沌中的阶梯:何时、如何(以及或许为什么)测试时缩放改善大型语言模型的机器翻译
Abstract
Two forms of test-time scaling for Large Language Models (LLMs) have emerged as effective and widely adopted paradigms: sequential, in which later answer attempts depend on earlier ones, and parallel, such as i.i.d. sampling with reranking. In this study, we investigate their properties in translation. First, our study shows that sequential sampling has a higher performance ceiling, providing a more diverse and effective pool of samples, particularly under smaller sampling budgets. Second, we interrogate the nature of test-time scaling through a multidimensional manual analysis. Human analysis of the Best-of-$N$ translations demonstrates that sequential sampling substantially improves translation fluency and naturalness, but can degrade accuracy when inference budgets are large. Finally, we suggest an explanation of the mechanism through which sequential scaling improves machine translation. Our controlled analysis partially attributes the success of sequential self-improvement to the model's access to a larger target-side context. Ablation experiments on sequential sampling demonstrate its robustness across different sampling temperatures, while also revealing sensitivity to context construction, suggesting directions for future improvement.
Chinese Translation
两种形式的测试时缩放在大型语言模型(LLMs)中已成为有效且广泛采用的范式:顺序缩放,其中后续的回答尝试依赖于先前的回答,以及并行缩放,例如独立同分布(i.i.d.)采样与重排名。在本研究中,我们探讨了它们在翻译中的特性。首先,我们的研究表明,顺序采样具有更高的性能上限,提供了更为多样化和有效的样本池,尤其是在较小的采样预算下。其次,我们通过多维手动分析审视测试时缩放的本质。对最佳的$N$翻译的人工分析表明,顺序采样显著提高了翻译的流畅性和自然性,但在推理预算较大时可能会降低准确性。最后,我们提出了顺序缩放改善机器翻译的机制解释。我们的控制分析部分将顺序自我改进的成功归因于模型对更大目标侧上下文的访问。对顺序采样的消融实验表明其在不同采样温度下的稳健性,同时也揭示了对上下文构建的敏感性,为未来的改进指明了方向。
cs.CL / 57 / 2608.28508
Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation
基于自监督语音表示的音素和词级度量用于强制对齐评估
Abstract
Forced alignment evaluation typically requires manually annotated timestamps, limiting large-scale and multilingual analysis. We introduce two corpus-level metrics based on self-supervised (SSL) speech representations for reference-free forced alignment evaluation: Phoneme-Cluster Mutual Information (PCMI) and Word Acoustic Consistency Score (WACS). PCMI measures agreement between aligned phoneme labels and clusters induced from SSL-speech representations, while WACS measures consistency of repeated word realizations using dynamic time warping similarity between word representation sequences. Using both random and systematic perturbations, we show that PCMI and WACS degrade consistently under alignment perturbations. We further analyze the metrics across multiple alignment systems on 85 languages from FLEURS, validate them against manually annotated alignments from 45 languages in DoReCo, and evaluate them on two phonologically complex low-resource languages. The metrics effectively separate high- and low-quality alignments and correlate strongly with timestamp-based alignment quality measures. Our results demonstrate that SSL-speech representations enable scalable, reference-free forced alignment evaluation. The metrics are available as an open-source Python package at https://github.com/mahesh-ak/forced-aligner-metrics.
Chinese Translation
强制对齐评估通常需要手动标注的时间戳,这限制了大规模和多语言分析。我们提出了两种基于自监督(SSL)语音表示的语料库级度量,用于无参考的强制对齐评估:音素聚类互信息(Phoneme-Cluster Mutual Information, PCMI)和词音频一致性评分(Word Acoustic Consistency Score, WACS)。PCMI 衡量对齐的音素标签与从 SSL 语音表示中诱导的聚类之间的一致性,而 WACS 则通过使用动态时间规整相似性来测量重复词实现的一致性。通过随机和系统性扰动,我们展示了 PCMI 和 WACS 在对齐扰动下的一致降级。我们进一步分析了这两个度量在来自 FLEURS 的 85 种语言的多个对齐系统中的表现,并在 DoReCo 中对 45 种语言的手动标注对齐进行验证,同时在两种音系复杂的低资源语言上进行评估。这些度量有效地区分高质量和低质量的对齐,并与基于时间戳的对齐质量测量具有很强的相关性。我们的结果表明,SSL 语音表示能够实现可扩展的无参考强制对齐评估。这些度量作为开源 Python 包可在 https://github.com/mahesh-ak/forced-aligner-metrics 获取。
cs.CL / 58 / 2608.28560
A Formal Limitation on Learning Human Language From Textual Corpora
从文本语料中学习人类语言的正式限制
Abstract
Can a listener recover what a speaker means from the form of an utterance alone? We answer this question information-theoretically, and for a listener given by any featurizer of text, including the hidden states of contemporary large language models. Modeling language use as a joint distribution over meanings, contexts, and utterances, we derive upper bounds on the probability that a decoder recovers a speaker's intended meaning from a representation of the utterance. The bounds are governed by the uncertainty that form leaves about meaning, which splits into an irreducible part and a part that only (extralinguistic) context, but never the utterance alone, can resolve. Because these quantities are intrinsic to language, no representation, however much text or supervision produced it, can surpass them; the bounds hold whether the space of meanings is discrete or continuous. Experiments on artificial languages, Mandarin zero-pronoun resolution, and color reference provide empirical evidence in support of the theory.
Chinese Translation
听者是否能够仅从话语的形式中恢复说话者的意思?我们从信息论的角度回答了这个问题,并考虑了任何文本特征化方法下的听者,包括当代大型语言模型的隐藏状态。我们将语言使用建模为意义、上下文和话语的联合分布,推导出解码器从话语的表示中恢复说话者意图的概率的上限。这些上限受形式对意义的不确定性支配,这种不确定性分为不可简化的部分和仅能通过(超语言)上下文而非单独话语来解决的部分。由于这些量是语言固有的,因此无论生成了多少文本或监督,任何表示都无法超越这些限制;这些上限适用于离散或连续的意义空间。对人工语言、汉语零代词解析和颜色参考的实验提供了支持该理论的实证证据。