← Back to Index
Daily Research Digest

arXiv Papers

2026-08-13
250
Papers
4
Categories
250
Translated
收藏清单 0
机器人学 (Robotics)
21
cs.RO / 1 / 2608.11363

Adaptation of Generalist Robot Policies with Minimal Data

利用最少数据适应通用机器人策略
Kowshik, Shreyas, Venkataraman, Sreyas, Wang, Leo, Pant, Niharika, Simchowitz, Max, Kumar, Aviral
Abstract
A central goal in robot learning is to move beyond task-specific human data collection toward robots that improve through autonomous interaction. Yet fully autonomous learning remains difficult with current policies: sparse rewards and weak zero-shot exploration make it unlikely that a robot will discover successful behavior from scratch. We study minimal-data adaptation, a regime in which a pre-trained robot policy must learn a new task from as little as one demonstration followed by autonomous online interaction. This setting serves as the closest tractable proxy for fully autonomous improvement, allowing us to study whether minimal human guidance can bootstrap autonomous learning and what algorithmic ingredients make it feasible. We build MiDAS, a simple offline-to-online RL recipe that first anchors a pre-trained VLA to the target task with behavior cloning on single/few demonstrations, then improves it through value-based online RL on a residual policy parameterization. Across LIBERO and RoboCasa, MiDAS recovers strong task performance from as little as one demonstration, substantially outperforming baselines and generalizing beyond demonstrated conditions. We further evaluate MiDAS on a bimanual YAM platform. Starting from a fragile low-success policy obtained from a single demonstration, MiDAS improves its robustness and learns new successful behaviors over ~6 hours of online interaction. To the best of our knowledge, this is the first demonstration of reliable robot policy adaptation from a single task demonstration.
Chinese Translation
机器人学习的一个核心目标是超越特定任务的人类数据收集,朝着能够通过自主交互进行改进的机器人发展。然而,当前策略下完全自主学习仍然困难:稀疏的奖励和弱的零-shot 探索使得机器人从零开始发现成功行为的可能性很小。我们研究了最少数据适应,这是一种预训练机器人策略必须从仅一个示范及随后的自主在线交互中学习新任务的模式。该设置是完全自主改进的最接近可处理的代理,使我们能够研究最少的人类指导是否能够引导自主学习,以及哪些算法成分使其成为可能。我们构建了 MiDAS,这是一种简单的离线到在线的强化学习(RL)方案,首先通过在单个/少量示范上的行为克隆将预训练的变压器(VLA)锚定到目标任务,然后通过基于价值的在线 RL 在残差策略参数化上进行改进。在 LIBERO 和 RoboCasa 上,MiDAS 从仅一个示范中恢复了强大的任务表现,显著优于基线并超越了示范条件的泛化。我们进一步在双手 YAM 平台上评估 MiDAS。从单个示范获得的脆弱低成功率策略开始,MiDAS 提高了其鲁棒性,并在约 6 小时的在线交互中学习了新的成功行为。据我们所知,这是首次展示从单个任务示范中可靠的机器人策略适应。
cs.RO / 2 / 2608.11407

Top-down Traffic Scenario Generation via Joint Initial-Goal Diffusion and Trajectory Infilling

通过联合初始目标扩散与轨迹填充的自上而下交通场景生成
Lee, Da Saem, Pant, Yash Vardhan, Fischmeister, Sebastian
Abstract
Robust traffic simulators are crucial for developing and testing autonomous vehicles to reduce the costly, labor-intensive real-world data collection process and the need for physical presence on the road. However, existing simulators require agents' initial states to generate trajectories, which limits scalability and diversity due to restrictions on the given initial states. While data-driven agent initialization has been widely studied, the generated initial states are not interpretable in terms of why the agents are initialized at those specific locations. Given known initial states, trajectory generation is also a challenging problem, as the model must learn the variability of the destination and how agents should reach it over time. In this paper, we propose TrafficDiffuser, a top-down traffic scenario generation framework that generates high-level traffic scenarios, defined by initial and goal state pairs, by jointly modeling them. The high-level scenario generation makes initial states better interpretable and reduces trajectory generation into as simple as an infilling problem. We demonstrate how the generated high-level traffic scenarios can be used, including constraining based on different trajectory modes and integrating them with existing trajectory generation models. We conduct extensive experiments on the Argoverse 2 motion prediction dataset to evaluate how well the generated outputs capture real-world distributions. In addition to generating goal states, TrafficDiffuser outperforms the next-best approach for agent initialization, reducing speed distribution distance by 55.3% and the off-road rate by 2.8%.
Chinese Translation
强大的交通模拟器对于开发和测试自动驾驶车辆至关重要,能够减少成本高昂、劳动密集的现实数据收集过程以及在道路上进行实地测试的需求。然而,现有的模拟器需要代理的初始状态来生成轨迹,这限制了可扩展性和多样性,因为对给定初始状态的限制。尽管数据驱动的代理初始化已被广泛研究,但生成的初始状态在解释代理为何在特定位置初始化方面并不直观。在已知初始状态的情况下,轨迹生成也是一个具有挑战性的问题,因为模型必须学习目的地的变化性以及代理应如何随时间到达目的地。在本文中,我们提出了TrafficDiffuser,一个自上而下的交通场景生成框架,通过联合建模初始状态和目标状态对来生成高层次的交通场景。高层次场景生成使得初始状态更具可解释性,并将轨迹生成简化为填充问题。我们展示了生成的高层次交通场景的应用,包括基于不同轨迹模式的约束以及与现有轨迹生成模型的集成。我们在Argoverse 2运动预测数据集上进行了广泛的实验,以评估生成的输出在多大程度上捕捉了现实世界的分布。除了生成目标状态外,TrafficDiffuser在代理初始化方面的表现优于下一最佳方法,将速度分布距离减少了55.3%,将越界率降低了2.8%。
cs.RO / 3 / 2608.11409

From Self-Normal-Positioning to Omni-Directional Tracking: Real-Time Surface Modeling Enabled Probe Tilt Control for Robotic Ultrasound Imaging

从自我法线定位到全向跟踪:实时表面建模驱动的探头倾斜控制用于机器人超声成像
Ma, Xihan, Zhang, Haichong
Abstract
Ultrasound (US) provides real-time, radiation-free imaging, but the image quality depends strongly on how the probe is oriented against the patient body. Robotic US can reduce operator workload and improve acquisition consistency; however, most existing systems focus on normal positioning, where the probe is maintained perpendicular to the local surface. This constraint is inadequate for examinations like echocardiography, where obtaining a diagnostic view requires a non-normal probe angle. Consequently, a clinically useful robotic system must sense the local surface in real-time and preserve the desired probe orientation. Here, we propose an omni-directional probe-orientation control framework that integrates RGB-D perception, local-surface modeling, and task-space orientation control. The surface model fuses multi-view point clouds and provides a quadratic estimate of the local surface. A desired imaging direction is then encoded relative to the normal, enabling the probe to track arbitrary angles. The framework was evaluated through flat-surface tracking, phantom target-angle recovery, and in-vivo tracking of an expert selected view. Results show that the mean angular tracking error was 1.06 +- 0.66 deg. The system recovered a non-normal tilt angle of up to 44.39 +- 2.59 deg relative to the surface normal, and acquired the desired heart chamber view in the phantom and in-vivo experiments.
Chinese Translation
超声(US)提供实时、无辐射成像,但图像质量在很大程度上依赖于探头相对于患者身体的方向。机器人超声可以减少操作员的工作负担并提高采集一致性;然而,大多数现有系统专注于法线定位,即探头保持与局部表面垂直。这一限制对于如超声心动图等检查是不够的,因为获取诊断视图需要非法线探头角度。因此,一个临床上有用的机器人系统必须实时感知局部表面并保持所需的探头方向。在此,我们提出了一种全向探头方向控制框架,该框架集成了RGB-D感知、局部表面建模和任务空间方向控制。该表面模型融合了多视角点云,并提供局部表面的二次估计。然后,相对于法线编码所需的成像方向,使探头能够跟踪任意角度。通过平面表面跟踪、虚拟目标角度恢复以及对专家选择视图的体内跟踪对该框架进行了评估。结果显示,平均角度跟踪误差为1.06 ± 0.66度。该系统恢复了相对于表面法线的最大44.39 ± 2.59度的非法线倾斜角度,并在虚拟和体内实验中获取了所需的心腔视图。
cs.RO / 4 / 2608.11417

Locomotion Variability and User Experience in Smart Wheelchair Human-Robot Interaction

智能轮椅人机交互中的运动变异性与用户体验
Kille, Sean, Panchea, Adina M., Varga, Balint, Hohmann, Sören
Abstract
Human movement is inherently variable, with variability structured according to task relevance: movements are typically more consistent at task-critical points and more flexible elsewhere. In human-robot interaction (HRI), however, model-based assistance strategies commonly assume deterministic human behavior and suppress such variability, potentially altering how interactions are experienced and lowering sense of agency. While movement variability is increasingly recognized as functionally meaningful, its deliberate preservation in assisted interaction, and its consequences for user experience, remain underexplored. In this paper, we empirically investigate how different assistance strategies shape human movement variability, task performance, and subjective interaction experience in a shared control setting. We introduce an autonomy-supportive shared control strategy that preserves users' natural movement structure. This approach is evaluated in a user study in which participants push an intelligent powered wheelchair under three conditions: no assistance, conventional variability-reducing assistance, and variability-preserving assistance. While task-relevant performance remained comparable across assisted modes, preserving natural movement variability led to more favorable interaction experiences. In particular, participants reported significantly higher perceived agency compared to conventional assistance and highest perceived usefulness. These findings suggest that variability-aware assistance can support both performance and user autonomy in physical human-robot collaboration. More broadly, the results highlight the importance of designing assistive robotic systems that respect the embodied structure of human movement rather than treating variability as noise to be neglected or eliminated.
Chinese Translation
人类运动本质上是具有变异性的,其变异性根据任务相关性进行结构化:在任务关键点,运动通常更为一致,而在其他地方则更为灵活。然而,在人机交互(HRI)中,基于模型的辅助策略通常假设人类行为是确定性的,并抑制这种变异性,这可能会改变交互的体验并降低自主感。尽管运动变异性越来越被认为具有功能意义,但在辅助交互中有意保留这种变异性及其对用户体验的影响仍然未被充分探索。本文通过实证研究探讨了不同辅助策略如何塑造人类运动变异性、任务表现和主观交互体验,特别是在共享控制环境中。我们提出了一种支持自主性的共享控制策略,旨在保留用户的自然运动结构。该方法在一项用户研究中进行了评估,参与者在三种条件下推动智能电动轮椅:无辅助、传统的减少变异性辅助和保留变异性的辅助。尽管在辅助模式下任务相关的表现保持相当,但保留自然运动变异性导致了更为积极的交互体验。特别是,与传统辅助相比,参与者报告了显著更高的感知自主感和最高的感知有用性。这些发现表明,关注变异性的辅助可以同时支持表现和用户在物理人机协作中的自主性。更广泛地说,结果强调了设计尊重人类运动具身结构的辅助机器人系统的重要性,而不是将变异性视为应被忽视或消除的噪声。
cs.RO / 5 / 2608.11451

Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards

通过神经符号安全守卫实现端到端自主驾驶的群体行为
Idarraga, Simón Patiño, Silva, Erick, Yasmin, Rehana, Shoker, Ali
Abstract
Modern end-to-end driving agents can achieve high average performance yet still violate basic traffic rules that a human driver would never miss. The reason is structural: they learn statistical patterns rather than the physical conditions that guarantee safe driving, leaving their decision-making process opaque and safety constraints unenforced. We introduce a neuro-symbolic safety guard, a lightweight module that attaches to the final command interface of an already-trained agent. Immediately before a command reaches the vehicle, it checks the command against explicit safety rules and, only when necessary, replaces it with the nearest safe alternative. Each intervention is directly executable and traceable to the rule that triggered it, while the guard itself requires no retraining and adds no learned component. Evaluated on the long-tail benchmarks Fail2Drive and Bench2Drive using the state-of-the-art TransFuser v6 (TFv6) as a case study, the guard improves Success Rate by 15% and reduces safety-critical collisions by up to 53%, while preserving the original Driving Score.
Chinese Translation
现代端到端驾驶代理能够实现高平均性能,但仍然会违反人类驾驶员绝不会忽视的基本交通规则。其原因在于结构性:它们学习的是统计模式,而非确保安全驾驶的物理条件,这使得它们的决策过程不透明,安全约束未得到执行。我们提出了一种神经符号安全守卫,这是一种轻量级模块,附加在已训练代理的最终指令接口上。在指令到达车辆之前,它会根据明确的安全规则检查该指令,只有在必要时,才会将其替换为最近的安全替代方案。每次干预都是可直接执行的,并且可以追溯到触发它的规则,而安全守卫本身无需重新训练,也不增加任何学习组件。在使用最先进的TransFuser v6 (TFv6)作为案例研究的长尾基准Fail2Drive和Bench2Drive上进行评估时,安全守卫将成功率提高了15%,并将安全关键碰撞减少了多达53%,同时保持了原始驾驶评分。
cs.RO / 6 / 2608.11461

Koopman Representation of Nonlinear Virtual Environments in Kinesthetic Haptic Systems

基于Koopman算子的非线性虚拟环境在运动触觉系统中的表征
Zhou, Yanting, Kövecses, Jozsef, Forbes, James Richard
Abstract
Rendering haptic feedback with nonlinear virtual environments (VEs) is important in many applications that require highly accurate force feedback. This paper considers the use of the Koopman operator to represent a nonlinear VE interacting with a haptic system. Simulation and experimental results demonstrated that the proposed method provides an effective representation of the nonlinear dynamics of a Duffing-oscillator VE. A multi-user study further confirmed this conclusion. In addition, a closed-loop (CL) stability analysis is performed leveraging the Koopman representation of the nonlinear VE to access stability of the overall haptic system. This alternative way of representing nonlinear VEs enables a convenient CL stability analysis that is less conservative than traditional passivity-based methods. Since a linear combination of all lifted states is used to represent the nonlinearity, such representation is also more robust to uncertainties in the modeling of the haptic device than a traditional nonlinear model.
Chinese Translation
在许多需要高精度力反馈的应用中,渲染非线性虚拟环境(VEs)的触觉反馈至关重要。本文考虑使用Koopman算子来表征与触觉系统交互的非线性虚拟环境。仿真和实验结果表明,所提出的方法有效地表征了Duffing振荡器虚拟环境的非线性动态。多用户研究进一步确认了这一结论。此外,利用Koopman对非线性虚拟环境的表征进行闭环(CL)稳定性分析,以评估整体触觉系统的稳定性。这种替代的非线性虚拟环境表征方式使得闭环稳定性分析更加便捷,且比传统的基于被动性的方法更不保守。由于使用所有提升状态的线性组合来表征非线性,因此这种表征在触觉设备建模的不确定性方面也比传统的非线性模型更具鲁棒性。
cs.RO / 7 / 2608.11521

Keep the Future, Drop the Rollout: RIFT for World Action Models

保持未来,放弃展开:用于世界行动模型的RIFT
Zhang, Chushan, Tong, Jinguang, Li, Xuesong, Wang, Yikai, Li, Hongdong
Abstract
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on all 40 LIBERO tasks, paired closed-loop interventions show that masking or reassigning future-cache values changes execution and reduces success, indicating sensitivity to future values and their assigned positions. For Joint and Cosmos-2, however, replaying one fixed final-clean key/value (K/V) cache nearly preserves unmodified execution, with $1.7$ to $1.9$~cm end-effector average displacement error and $97.9\%$ to $98.2\%$ success. This separates cache consumption from production: these models can reuse a fixed cache but still require iterative rollout to construct it. We therefore propose RIFT (\emph{Rollout-free Imagination via Future Tokens}), which uses learned anticipation tokens to construct a complete future K/V cache in one backbone pass while retaining the original future-read interface. On LIBERO, RIFT achieves $98.8\%$ success, close to rollout-based Joint, IDM, and LingBot-VA at $98.4\%$ to $98.6\%$, while reducing action-chunk latency by $68.2\%$ to $89.1\%$. On RoboTwin~2.0, RIFT reaches $92.9/92.6\%$ on clean/randomized scenes, the highest observed among the evaluated methods. These results support rollout-free future conditioning without iterative video generation at deployment.
Chinese Translation
世界行动模型(WAMs)基于预测的未来条件化机器人动作,但迭代视频展开增加了部署延迟。我们探讨动作生成是否需要不断演变的展开轨迹,还是仅需其未来表示。在所有40个LIBERO任务的四个WAMs中,配对的闭环干预表明,屏蔽或重新分配未来缓存值会改变执行并降低成功率,表明对未来值及其分配位置的敏感性。然而,对于Joint和Cosmos-2,重放一个固定的最终清洁键/值(K/V)缓存几乎可以保持未修改的执行,末端执行器的平均位移误差为$1.7$到$1.9$~cm,成功率为$97.9\%$到$98.2\\%$。这将缓存消耗与生产分开:这些模型可以重用固定缓存,但仍需迭代展开来构建它。因此,我们提出了RIFT( extit{Rollout-free Imagination via Future Tokens}),该方法利用学习到的预期令牌在一次主干传递中构建完整的未来K/V缓存,同时保留原始的未来读取接口。在LIBERO上,RIFT达到了$98.8\\%$的成功率,接近基于展开的Joint、IDM和LingBot-VA的$98.4\\%$到$98.6\\%$,同时将动作块延迟减少了$68.2\\%$到$89.1\\%$。在RoboTwin~2.0上,RIFT在干净/随机场景中达到了$92.9/92.6\\%$的成功率,是评估方法中观察到的最高值。这些结果支持无展开的未来条件化,而无需在部署时进行迭代视频生成。
cs.RO / 8 / 2608.11580

RoadWeaver: Large-Scale Lane-Level HD Map Generation from Scratch for Autonomous Driving Simulation

RoadWeaver:从零开始生成大规模车道级高清地图以支持自动驾驶仿真
Li, Yueyuan, Chen, Zexi, Xi, Weijie, Jiang, Mingyang, Zhang, Songan, Zhuang, Hanyang, Yang, Ming
Abstract
Autonomous driving simulation requires diverse and scalable lane-level HD maps to support long-horizon evaluation across complex road networks. Existing approaches either rely on handcrafted or reconstructed real-world maps, which limits scalability, or generate only local road structures rather than complete HD maps. We present RoadWeaver, a coarse-to-fine framework for from-scratch generation of diverse, large-scale HD maps. RoadWeaver first synthesizes a global road layout, expands it into a connected road network, and then constructs lane-level geometry with topologically consistent lane connectivity. Experimental results show that RoadWeaver achieves a 99.8\% reachability, a 10.7\% dead-end ratio, and an endpoint alignment error of 0.24 m. Compared with SOTA generation methods, it reduces endpoint alignment error by 94.4\% while generating complete HD maps in 1.39--3.50 s. The generated maps can be directly deployed in driving simulators, providing scalable simulation environments for future closed-loop evaluation of autonomous driving systems. The training code and an out-of-the-box implementation of RoadWeaver will be released upon acceptance.
Chinese Translation
自动驾驶仿真需要多样化且可扩展的车道级高清地图,以支持在复杂道路网络中的长时间评估。现有的方法要么依赖于手工制作或重建的真实世界地图,这限制了可扩展性,要么仅生成局部道路结构而非完整的高清地图。我们提出了RoadWeaver,一个从零开始生成多样化、大规模高清地图的粗到细框架。RoadWeaver首先合成全球道路布局,将其扩展为一个连接的道路网络,然后构建具有拓扑一致车道连通性的车道级几何形状。实验结果表明,RoadWeaver实现了99.8%的可达性,10.7%的死胡同比例,以及0.24米的端点对齐误差。与最先进的生成方法相比,它将端点对齐误差降低了94.4%,同时在1.39至3.50秒内生成完整的高清地图。生成的地图可以直接部署在驾驶模拟器中,为未来自动驾驶系统的闭环评估提供可扩展的仿真环境。RoadWeaver的训练代码和即插即用的实现将在论文接受后发布。
cs.RO / 9 / 2608.11592

Video2Track: From Real-World Interaction Videos to Steerable Adversarial Closed-Track Testing for Automated Driving Systems

Video2Track:从现实世界互动视频到可操控的对抗性闭环测试用于自动驾驶系统
Tian, Mengjie, Zhang, Xinrui, Li, Tianyu, Zhang, Peizhi, Zhou, Guirong, Feng, Haojie, Huang, Junpeng, Zhang, Qixiang, Xiong, Lu
Abstract
Closed-track testing plays a fundamental role in the verification and validation of automated driving systems (ADS), particularly for safety-critical scenarios, by enabling reproducible evaluation under controlled conditions. However, most existing approaches still rely on standardized protocols or predefined trajectories, leading to overly scripted interactions and limited ability to reproduce the natural complexity of public-road traffic. To address this limitation, we propose Video2Track, a framework that transfers real-world interactive driving scenarios from videos into steerable adversarial closed-track testing. The framework consists of two tightly coupled modules. The first is a scenario semantic mapping module, which extracts structured semantics from driving videos using a vision-language model and grounds them onto a closed-track topology library via retrieval-augmented generation, thereby identifying compatible map segments and interaction anchors. The second is a dynamic interactive testing module, which conditions on the grounded topology and anchors to generate diverse multi-agent trajectories through a conditional diffusion model, while regulating interaction intensity via a Stackelberg game with a parameterized adversarial objective. Closed-track experiments demonstrate that the proposed framework can faithfully reproduce representative real-world interaction scenarios and generate executable scenario variants with controllable risk levels and interaction styles, providing a scalable approach for realistic and steerable ADS validation.
Chinese Translation
闭环测试在自动驾驶系统(ADS)的验证和确认中发挥着基础性作用,特别是在安全关键场景中,通过在受控条件下实现可重复的评估。然而,现有的大多数方法仍然依赖于标准化协议或预定义轨迹,导致过于刻板的互动和有限的能力来重现公共道路交通的自然复杂性。为了解决这一局限性,我们提出了Video2Track,一个将现实世界互动驾驶场景从视频转化为可操控的对抗性闭环测试的框架。该框架由两个紧密耦合的模块组成。第一个是场景语义映射模块,它使用视觉-语言模型从驾驶视频中提取结构化语义,并通过检索增强生成将其映射到闭环拓扑库,从而识别兼容的地图段和互动锚点。第二个是动态互动测试模块,它基于已映射的拓扑和锚点,通过条件扩散模型生成多样化的多智能体轨迹,同时通过带参数的对抗目标的斯塔克尔博格博弈调节互动强度。闭环实验表明,所提出的框架能够忠实地重现具有代表性的现实世界互动场景,并生成具有可控风险水平和互动风格的可执行场景变体,为现实且可操控的ADS验证提供了一种可扩展的方法。
cs.RO / 10 / 2608.11597

IoT-Enabled Autonomous Maritime Navigation in Smart Ports: A Curriculum-Guided Shared Policy Learning Framework

物联网驱动的智能港口自主海洋导航:一种课程指导的共享策略学习框架
Lin, Yuqing, Zhang, Rangya, Yuen, Kum Fai
Abstract
As smart port infrastructures increasingly rely on autonomous maritime devices enabled by the Internet of Things (IoT), ensuring reliable onboard navigation intelligence has become a critical challenge for safe and scalable operations in congested waterways. This paper investigates onboard autonomous navigation for such IoT devices under partial observability and dense traffic conditions. A curriculum-guided reinforcement learning framework with a shared recurrent policy is developed to enhance temporal reasoning, deployment scalability, and robustness of edge-level decision-making. Centralized training is adopted as an offline design-time strategy, while all navigation actions are executed fully onboard, consistent with IoT edge intelligence paradigms. Extensive simulations in multiple realistic port environments demonstrate that the proposed approach improves navigation reliability, collision avoidance, and training stability compared with standard baseline methods, and generalizes effectively to previously unseen high-density scenarios. The results indicate that curriculum-guided shared learning provides a practical solution for scalable deployment of IoT-enabled autonomous maritime devices in smart port operations.
Chinese Translation
随着智能港口基础设施越来越依赖于物联网(IoT)驱动的自主海洋设备,确保可靠的船上导航智能已成为在拥挤水道中安全和可扩展操作的关键挑战。本文研究了在部分可观测和密集交通条件下,这些物联网设备的船上自主导航。我们开发了一种课程指导的强化学习框架,采用共享递归策略,以增强时间推理、部署可扩展性和边缘决策的鲁棒性。集中训练被采用作为离线设计时策略,而所有导航动作均在船上完全执行,这与物联网边缘智能范式一致。在多个现实港口环境中进行的广泛仿真表明,与标准基线方法相比,所提出的方法提高了导航的可靠性、避免碰撞的能力和训练的稳定性,并有效地推广到先前未见的高密度场景。结果表明,课程指导的共享学习为物联网驱动的自主海洋设备在智能港口操作中的可扩展部署提供了切实可行的解决方案。
cs.RO / 11 / 2608.11671

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models

StellaVLA:用于可泛化视觉-语言-动作模型的上下文结构化示范
Xu, Siyu, Wang, Yunke, Wang, Zijian, Zhu, Dihao, Xia, Chenghao, Du, Chengbin, Liu, Daochang, Huang, Tao, Xu, Chang
Abstract
Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often collapses out of distribution (OOD), when the scene, viewpoint, or object differs from training. Adapting to each new situation typically requires collecting more data and fine-tuning. We present StellaVLA, a framework that instead adapts at test time by conditioning on a single retrieved demonstration. The key idea is to move beyond imitating what an expert did and instead convey why: an automated offline pipeline converts each raw trajectory into a structured demonstration, e.g., a task plan, sub-goal descriptions, and verbalized 3D motion, at zero human-annotation cost. Provided as in-context guidance, this structured demonstration lets the policy reason about the task rather than mimic a pixel trajectory, which also makes it transferable across embodiments (real-robot, human-hand, or XR demonstrations). A parallel dual-training design internalizes this reasoning during training through a joint action-and-language objective, while inference uses the action expert alone, preserving real-time, high-frequency control with no added latency. On the VLA-Arena leaderboard(Aug 1, 2026), StellaVLA ranks first with an overall score of 0.63, versus 0.44 and 0.22 for the strong prior models ($\pi_{0.5}$ and LingBot-VLA), and it further leads on LIBERO with 98.8% average success rate and LIBERO-Plus with 85.1% success rate. Our real-robot benchmark demonstrates that StellaVLA can use both human/robot demos and human-to-robot (XR) demos as in-context structured demonstration to help VLA model adapt to OOD tasks.
Chinese Translation
视觉-语言-动作(VLA)模型能够遵循指令并操控物体,但当场景、视角或物体与训练数据不同时,其性能往往会下降。适应每种新情况通常需要收集更多数据并进行微调。我们提出了StellaVLA,一个在测试时通过条件化单个检索示范进行适应的框架。其关键思想是超越模仿专家的行为,而是传达其背后的原因:一个自动化的离线管道将每个原始轨迹转换为结构化示范,例如任务计划、子目标描述和口头化的三维运动,且无需人工标注。作为上下文指导提供的结构化示范使得策略能够对任务进行推理,而不是单纯模仿像素轨迹,这也使其能够在不同的实现方式(真实机器人、人手或扩展现实示范)之间进行迁移。一个并行的双重训练设计通过联合动作与语言目标在训练过程中内化这种推理,而推理时仅使用动作专家,从而保持实时、高频率的控制且没有额外延迟。在VLA-Arena排行榜(2026年8月1日)上,StellaVLA以0.63的总体得分排名第一,而强先验模型($ ext{π}_{0.5}$和LingBot-VLA)的得分分别为0.44和0.22,并且在LIBERO上以98.8%的平均成功率和在LIBERO-Plus上以85.1%的成功率领先。我们的真实机器人基准测试表明,StellaVLA能够利用人类/机器人示范和人类到机器人(XR)示范作为上下文结构化示范,帮助VLA模型适应OOD任务。
cs.RO / 12 / 2608.11731

ContactIPM: A Structure-Exploiting Interior-Point Solver for Contact-Implicit Trajectory Optimization

ContactIPM:一种结构利用的内点求解器用于接触隐式轨迹优化
Chen, Yucheng
Abstract
Contact-implicit trajectory optimization avoids prescribing contact sequences, but yields mathematical programs with complementarity constraints (MPCCs) whose degeneracy challenges conventional primal--dual solvers. Existing contact-specific methods improve robustness to this degeneracy but do not leverage a stagewise optimal-control factorization and primal--dual consistency, while structure-exploiting optimal-control solvers are not designed for complementarity constraints. We show that these capabilities can be combined in a single primal--dual method. ContactIPM identifies complementary inequality pairs, embeds them through a barrier-coupled elastic interior relaxation, eliminates slack and dual variables stagewise, and solves the reduced Newton system using a Riccati recursion. A fixed multi-phase MPCC recovery schedule provides four continuation and restart attempts from naive initializations, while termination is gated by the unrelaxed physical complementarity residual. We compare ContactIPM with two contact-specific MPCC solvers, CRISP and IMPACT, using matched benchmark conditions and common post-solve acceptance criteria. On four fixed CRISP benchmark cases, ContactIPM is $2.17$--$8.87\times$ faster over 20 paired timing repetitions per case and achieves higher success on the Push Box and Push-T robustness suites. Against IMPACT, ContactIPM is \(2.96\times\) faster on Push T and \(4.91\times\) faster on Cart Transport, but \(4.46\times\) slower on Push Box. In 50 closed-loop Push Box rollouts spanning model mismatch, measurement noise, initial-pose errors, and state resets,
Chinese Translation
接触隐式轨迹优化避免了规定接触序列,但产生了具有互补约束的数学规划问题(MPCC),其退化性对传统的原始-对偶求解器构成挑战。现有的接触特定方法提高了对这种退化的鲁棒性,但未能利用阶段性最优控制分解和原始-对偶一致性,而结构利用的最优控制求解器并未针对互补约束进行设计。我们展示了这些能力可以在单一的原始-对偶方法中结合。ContactIPM识别互补不等式对,通过与障碍耦合的弹性内点松弛将其嵌入,逐步消除松弛和对偶变量,并使用Riccati递归求解简化的牛顿系统。固定的多阶段MPCC恢复计划提供了四次从简单初始化的延续和重启尝试,而终止则由未松弛的物理互补残差控制。我们将ContactIPM与两个接触特定的MPCC求解器CRISP和IMPACT进行比较,使用匹配的基准条件和共同的后处理接受标准。在四个固定的CRISP基准案例中,ContactIPM在每个案例的20次配对计时重复中快$2.17$--$8.87 imes$,并在Push Box和Push-T鲁棒性测试中取得更高的成功率。与IMPACT相比,ContactIPM在Push T上快了$2.96 imes$,在Cart Transport上快了$4.91 imes$,但在Push Box上慢了$4.46 imes$。在50次闭环Push Box的滚动测试中,涵盖了模型不匹配、测量噪声、初始姿态误差和状态重置,
cs.RO / 13 / 2608.11739

G0.5: One Autoregressive Stream for Robot Reasoning and Action

G0.5:一种用于机器人推理和行动的自回归流
Liu, Yicheng, Dong, Zibin, Ye, Baijun, Yuan, Tianyuan, Jiang, Tao, Yang, Anqi, Cao, Shicheng, Liu, Haonan, Sun, Yue, Guo, Zihan, Liu, Xiao, Ke, Dong, Pan, Changxun, Wu, Chenru, Cheng, Tailai, Ren, Xiaoshu, Zhang, Xinlei, Cui, Jianning, Zhao, Zijie, Zhang, Haoyu, Xu, Kaiming, Yang, Haodong, Zhang, Bowen, Niu, Jiahui, Zhu, Shaoting, Zhang, Shiduo, Zhao, Hang
Abstract
The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a context encoder rather than a decision-maker. We introduce G0.5, a pretrained autoregressive VLA in which a single transformer decoder emits reasoning and action tokens under a single objective. Three components make this tractable at foundation-model scale: a learnable cross-embodiment action tokenizer that maps heterogeneous robot actions into a shared vocabulary; a native chain-of-thought stream interleaving task decomposition, object grounding, and action hints with action tokens; and a visual memory module that injects multi-second history through the vision encoder. Because reasoning and action share a single set of weights, the pretrained VLM's capabilities carry over to physical behavior: the model follows instructions closely, and prompts directly steer action granularity, task horizon, and out-of-distribution scene handling without further training. Pretrained on a large collection of robot datasets together with VQA samples, G0.5 surpasses state-of-the-art models across 7 independent regimes: real-world fine-tuning on R1lite and R1pro robots (76.7\% vs.\ 53.3\% for $\pi_{0.5}$ and 24.4\% for GR00T-N1.7), the 2025 BEHAVIOR Challenge on 50 long-horizon household mobile manipulation tasks using a generalist policy (31.4\% vs.\ 26.3\% for $\pi_{0.5}$ and 26.1\% for the challenge winner), DROID post-training followed by zero-shot transfer to an unseen environment and objects (82.5\%), a language-following Pick-and-Place benchmark, LIBERO (98.9\%), RoboTwin 2.0 (93.3\%), and SimplerEnv-Bridge (87.3\%).
Chinese Translation
当前视觉-语言-行动(VLA)模型的主流方法将预训练的视觉语言模型(VLM)与单独训练的流匹配行动专家相结合。这使得VLM成为上下文编码器,而非决策者。我们提出了G0.5,这是一种预训练的自回归VLA,其中单个变换器解码器在单一目标下发出推理和行动标记。三个组件使其在基础模型规模上可行:一个可学习的跨体现行动标记器,将异构机器人行动映射到共享词汇中;一个原生的思维链流,将任务分解、物体定位和行动提示与行动标记交错;以及一个视觉记忆模块,通过视觉编码器注入多秒历史信息。由于推理和行动共享一组权重,预训练的VLM的能力转移到物理行为上:模型紧密遵循指令,提示直接引导行动的粒度、任务范围和对分布外场景的处理,无需进一步训练。G0.5在大规模机器人数据集和视觉问答(VQA)样本上进行预训练,在7个独立领域中超越了最先进的模型:在R1lite和R1pro机器人上的真实世界微调($ ext{76.7\%}$对比$ ext{53.3\\%}$的$ ext{π_{0.5}}$和$ ext{24.4\\%}$的GR00T-N1.7),2025年行为挑战赛中针对50个长时间范围的家庭移动操控任务使用通用策略($ ext{31.4\\%}$对比$ ext{26.3\\%}$的$ ext{π_{0.5}}$和$ ext{26.1\\%}$的挑战获胜者),DROID后训练再零样本转移到未见环境和物体($ ext{82.5\\%}$),语言跟随的拾取与放置基准LIBERO($ ext{98.9\\%}$),RoboTwin 2.0($ ext{93.3\\%}$),以及SimplerEnv-Bridge($ ext{87.3\\%}$)。
cs.RO / 14 / 2608.11769

Policy-Induced Hand Priors in Humanoid Dual-Arm Manipulation: Diagnosing and Mitigating Initial-Pose Dependence

政策诱导的人形双臂操作中的手部先验:诊断与减轻初始姿态依赖性
Jung, Chaeyeon, Park, Juyoun
Abstract
Vision-language-action (VLA) policies are expected to operate robustly across variations in the robot's initial configuration, yet aggregate task success can conceal pose-specific failures and inappropriate hand selection. This work investigates initial-pose dependence in VLA-based humanoid dual-arm manipulation. We characterize the initial-condition-dependent early hand preference as a policy-induced hand prior and quantify it using HandPriorScore, residual hand bias, and target responsiveness. Evaluations across multiple policies and 17 initial configurations reveal strong initial-pose--policy interactions: the same pose produces substantially different success rates across policies, while a single policy exhibits large performance variation across poses. Specific initial arm configurations can suppress or induce an asymmetric hand preference, with the resulting effect varying in direction and strength across policies. Wrist-camera observations also influence hand selection and task performance. Expanding initial-pose coverage in the training dataset substantially improves robustness, while targeted augmentation around a low-performing configuration increases its success rate. Comparisons across training configurations show that sufficient exposure to the target simulation task is beneficial, whereas the effect of real or auxiliary data depends on pose coverage, simulation ratio, and observation availability. These findings characterize a pose-conditioned hand prior, identify a localized initial arm configuration as a causal handle on hand-selection behavior, and demonstrate how data coverage and training composition affect initial-pose robustness.
Chinese Translation
视觉-语言-动作(VLA)政策预计能够在机器人初始配置的变化中稳健运行,但整体任务成功率可能掩盖特定姿态的失败和不当的手部选择。本研究探讨了基于VLA的人形双臂操作中的初始姿态依赖性。我们将初始条件依赖的早期手部偏好表征为政策诱导的手部先验,并通过HandPriorScore、残余手部偏差和目标响应性进行量化。在多个政策和17种初始配置下的评估揭示了强烈的初始姿态-政策交互:相同的姿态在不同政策下产生显著不同的成功率,而单一政策在不同姿态下表现出较大的性能变化。特定的初始臂配置可以抑制或诱发不对称的手部偏好,其结果在不同政策下的方向和强度各异。腕部摄像头的观察也会影响手部选择和任务表现。在训练数据集中扩展初始姿态覆盖范围显著提高了稳健性,而围绕低表现配置的有针对性增强则提高了其成功率。不同训练配置的比较显示,充分接触目标仿真任务是有益的,而真实或辅助数据的效果则依赖于姿态覆盖、仿真比例和观察可用性。这些发现表征了一种姿态条件的手部先验,识别出局部初始臂配置作为手部选择行为的因果因素,并展示了数据覆盖和训练组成如何影响初始姿态的稳健性。
cs.RO / 15 / 2608.11870

Enhancing Visual Domain Robustness in Behaviour Cloning via Saliency-Guided Augmentation

通过显著性引导增强行为克隆中的视觉领域鲁棒性
Zhuang, Zheyu, Wang, Ruiyu, Ingelhag, Nils, Kyrki, Ville, Kragic, Danica
Abstract
In vision-based behavior cloning (BC), conventional image augmentations such as Random Crop and Color Jitter often fall short under substantial visual domain shifts, including changes in shadows, distractors, and backgrounds. Superimposition-based augmentations, which blend in-domain and out-of-domain images, have shown promise for improving generalization in computer vision, but their suitability for BC remains uncertain because task-critical semantics, spatiotemporal relationships, and agent-target interactions must be preserved. To address this, we introduce RoboSaGA, a Saliency-Guided Augmentation method within the superimposition family tailored for vision-based BC. RoboSaGA dynamically adjusts augmentation intensity at the pixel level using policy-driven saliency, enabling aggressive augmentation in task-irrelevant regions while preserving task-critical information. It integrates seamlessly into existing architectures without requiring structural modifications or additional learning objectives. Experiments in both simulated and real-world settings show that RoboSaGA preserves in-domain performance while substantially improving robustness to visual domain shifts, including distractor and background changes, as well as lighting and shadow variations. Code is available at https://github.com/Zheyu-Zhuang/RoboSaGA.
Chinese Translation
在基于视觉的行为克隆(BC)中,传统的图像增强方法如随机裁剪(Random Crop)和颜色抖动(Color Jitter)在面对显著的视觉领域变化时(例如阴影、干扰物和背景的变化)往往效果不佳。基于叠加的增强方法通过将领域内和领域外的图像进行融合,已显示出在计算机视觉中提高泛化能力的潜力,但其在行为克隆中的适用性仍不确定,因为必须保留任务关键的语义、时空关系和代理-目标交互。为此,我们提出了RoboSaGA,一种针对基于视觉的行为克隆的显著性引导增强方法,属于叠加增强家族。RoboSaGA通过策略驱动的显著性在像素级动态调整增强强度,使得在与任务无关的区域进行激进增强的同时保留任务关键的信息。它可以无缝集成到现有架构中,而无需结构修改或额外的学习目标。在模拟和真实环境中的实验表明,RoboSaGA在保持领域内性能的同时,显著提高了对视觉领域变化的鲁棒性,包括干扰物和背景变化,以及光照和阴影变化。代码可在 https://github.com/Zheyu-Zhuang/RoboSaGA 获取。
cs.RO / 16 / 2608.11876

D3D-GEN: Robot-Aware Domain-Grounded Interactive 3D World Generation for Social Robotics

D3D-GEN:面向社交机器人的人机交互领域基础的3D世界生成系统
Do, Anh Duc, Scherbyna, Volodymyr, Nguyen, Tai Duc, Thakkar, Spaarsh, Shen, Zhengcheng, Buiyan, Teham, Misra, Archan, Kästner, Linh
Abstract
Training and validation of Embodied AI for social navigation critically depends on realistic simulation environments, yet many current approaches fail to find a balance between realism and simulability. We propose D3D-GEN, a novel world generation system that combines a domain agent with a retrieval-augmented generation (RAG) pipeline grounded in that domain. Our system enables users to rapidly generate domain-grounded, fully interactive 3D worlds by automating both the collection of domain knowledge and the synthesis of realistic floorplans and object placements, without dependence on any fixed 3D model database. Given a domain description prompt, the research agent collects publicly accessible domain-specific data and constructs a persistent domain database. Using this database, our RAG pipeline generates plausible floorplans and object placements by dynamically querying a user-provided semantic database, which can be easily extended or modified. The output is a fully interactive 3D world loadable by the popular simulators Isaac Sim and Gazebo. With our approach, we have built databases for several common domains (indoor residential, hospital, office) and generated dozens of distinct, plausible simulation environments for each domain. We present D3D-GEN with a local web frontend that facilitates rapid, interactive world generation for robot simulation.
Chinese Translation
具身人工智能在社交导航中的训练和验证严重依赖于逼真的仿真环境,然而许多当前的方法未能在现实性和可模拟性之间找到平衡。我们提出了D3D-GEN,一种新颖的世界生成系统,它将领域代理与基于该领域的检索增强生成(RAG)管道相结合。我们的系统使用户能够快速生成领域基础的、完全互动的3D世界,通过自动化收集领域知识和合成逼真的平面图及物体摆放,而不依赖于任何固定的3D模型数据库。在给定领域描述提示的情况下,研究代理收集公开可获取的领域特定数据,并构建一个持久的领域数据库。利用该数据库,我们的RAG管道通过动态查询用户提供的语义数据库生成合理的平面图和物体摆放,该数据库可以轻松扩展或修改。输出结果是一个可以被流行仿真器Isaac Sim和Gazebo加载的完全互动的3D世界。通过我们的方法,我们为多个常见领域(室内住宅、医院、办公室)构建了数据库,并为每个领域生成了数十个独特且合理的仿真环境。我们展示了D3D-GEN的本地Web前端,便于快速、互动地生成用于机器人仿真的世界。
cs.RO / 17 / 2608.11895

Scalable Multi-Agent Maze Traversal with Local Communication

可扩展的多智能体迷宫遍历与局部通信
Rau, Julian, Argote-Gerald, Jahir, McFassel, Grace, Miyauchi, Genki, Trodden, Paul, Groß, Roderich
Abstract
Cave networks, pipe systems, and similar maze-like environments pose significant challenges for multi-agent navigation in unknown settings with limited communication. We propose a distributed algorithm that enables agents to collectively traverse an unknown, possibly cyclic graph. Agents enter sequentially at a designated start node and are tasked to localize and reach an undisclosed goal while avoiding collisions. They coordinate via local communication using leader-follower relationships and leader switching. At any moment in time, exploration is performed by only one of the agents, which runs a single-agent maze solver. We prove that the algorithm is complete, that its makespan is asymptotically equivalent (in the number of agents) to that of an optimal full-knowledge strategy, and derive its time and space complexity. Simulations with up to $625$ agents show a decreasing average sum-of-fuels as the number of agents increases and demonstrate that the proposed approach outperforms a na\"ive baseline in which all agents independently execute the single-agent solver.
Chinese Translation
洞穴网络、管道系统及类似的迷宫环境对在未知环境中进行多智能体导航提出了重大挑战,尤其是在通信受限的情况下。我们提出了一种分布式算法,使得智能体能够共同穿越一个未知的、可能是循环的图。智能体依次从指定的起始节点进入,并被赋予定位和到达未公开目标的任务,同时避免碰撞。它们通过局部通信协调,采用领导-跟随关系和领导者切换。在任何时刻,只有一个智能体进行探索,该智能体运行单智能体迷宫求解器。我们证明了该算法是完备的,其完成时间在智能体数量上渐近等价于最优全知策略的完成时间,并推导了其时间和空间复杂度。对多达625个智能体的模拟显示,随着智能体数量的增加,平均燃料总和逐渐减少,并且证明了所提方法优于所有智能体独立执行单智能体求解器的简单基线。
cs.RO / 18 / 2608.11901

DaViNCi: A Dataset Towards Outdoor Vision-and-Language Navigation with Continuous Actions and Dynamic Elements

DaViNCi:一个面向具有连续动作和动态元素的户外视觉与语言导航的数据集
Xie, Zihao, Lai, Pingrui, Wu, Yitong, Yang, Hua
Abstract
Vision-and-Language Navigation (VLN) has progressively expanded from indoor to outdoor environments. However, existing outdoor VLN datasets still rely on fixed discrete topological graphs for construction. It fails to align with the rapidly changing real-world outdoor environments and impedes the sim-to-real transfer of VLN agents. To address this limitation, we propose DaViNCi (\textbf{D}yn\textbf{a}mic \textbf{Vi}sion-and-Language \textbf{N}avigation in \textbf{C}ont\textbf{i}nuous Environment), the first outdoor VLN dataset that simultaneously introduces both continuous and dynamic factors. The agent not only moves in the outdoor environment using continuous actions but is also required to handle unpredictable dynamic elements. The dataset encompasses six distinct maps with a total of 6,933 trajectories. Through comprehensive comparative experiments, we find that the success rate on DaViNCi decreased by more than 10\% in discrete environments compared to previous datasets. And there is an even greater decline in continuous settings, demonstrating the challenge of DaViNCi. Furthermore, we clarify the impact of action granularity and dynamic elements. These results demonstrate the practical value of DaViNCi in advancing outdoor VLN toward more realistic environments. The website is https://xzh0312.github.io/DaViNCi/.
Chinese Translation
视觉与语言导航(VLN)逐渐从室内环境扩展到户外环境。然而,现有的户外 VLN 数据集仍然依赖于固定的离散拓扑图进行构建。这与快速变化的现实户外环境不相符,阻碍了 VLN 代理的仿真到现实转移。为了解决这一局限性,我们提出了 DaViNCi( extbf{D}yn extbf{a}mic extbf{Vi}sion-and-Language extbf{N}avigation in extbf{C}ont extbf{i}nuous Environment),这是第一个同时引入连续和动态因素的户外 VLN 数据集。代理不仅使用连续动作在户外环境中移动,还需要处理不可预测的动态元素。该数据集包含六个不同的地图,总共 6,933 条轨迹。通过全面的比较实验,我们发现 DaViNCi 在离散环境中的成功率比之前的数据集下降了超过 10\%。而在连续环境中,下降幅度更大,显示了 DaViNCi 的挑战。此外,我们阐明了动作粒度和动态元素的影响。这些结果展示了 DaViNCi 在推动户外 VLN 向更现实环境发展的实际价值。网站地址为 https://xzh0312.github.io/DaViNCi/。
cs.RO / 19 / 2608.12063

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

从 SMPC 演示中学习运动操控的稀疏离线到在线强化学习
Schuck, Martin, Sorokin, Maks, Manni, Simone, Ta, Duy, Schoellig, Angela P., Hutter, Marco, Cleac'H, Simon Le, Brüdigam, Jan
Abstract
Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely bottlenecked by the slow, manual process of dense reward shaping. To bypass this limitation, we leverage Sample-based Model Predictive Control (SMPC) entirely in simulation as an automated, rapidly tunable expert to generate massive offline datasets. Because this data solves the fundamental exploration problem, we can train an off-policy RL agent using purely sparse task rewards, drastically reducing the time required to learn new skills and eliminating the need for manual tuning. Integrating this high-level agent with a low-level dynamic stability controller yields more optimal behaviors that strictly align with true task objectives, ultimately allowing the learned policies to surpass the original optimal control teacher. We validate the robustness of this sim-to-real framework by successfully deploying complex loco-manipulation skills across different morphologies, including an arm-equipped Spot quadruped and a G1 humanoid.
Chinese Translation
整合运动与操控对于机器人自主性至关重要,但将标准强化学习(Reinforcement Learning, RL)扩展到复杂任务时,稠密奖励塑造的缓慢手动过程严重制约了其效率。为了解决这一限制,我们在模拟环境中完全利用基于样本的模型预测控制(Sample-based Model Predictive Control, SMPC)作为一种自动化、快速可调的专家,以生成大量的离线数据集。由于这些数据解决了基本的探索问题,我们可以使用纯粹的稀疏任务奖励训练一个离策略 RL 代理,从而大幅减少学习新技能所需的时间,并消除手动调优的需求。将这个高层代理与低层动态稳定控制器结合,能够产生更优的行为,严格符合真实任务目标,最终使得学习到的策略超越原始的最优控制教师。我们通过成功部署复杂的运动操控技能在不同形态的机器人上验证了这一从模拟到现实的框架的鲁棒性,包括配备机械臂的 Spot 四足机器人和 G1 人形机器人。
cs.RO / 20 / 2608.12122

HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing

HandEdit:一种统一的以自我为中心的人机灵巧手图像编辑基准
Yang, Zhenjie, Jiao, Xingyu, Zhong, Guopeng, Yang, Shuzhe, Che, Shi, Wu, Chao, Jiang, Chenyu, Zhang, Dongjie, Zhang, Yideng, Zhang, Zheng, Jiang, Muyun, Su, Haisheng, Jin, Shuang, Zhang, Donghang, Yang, Chao, Chen, Li, Li, Hongyang, Wu, Zuxuan, Jiang, Yu-Gang, Jia, Xiaosong, Yan, Junchi
Abstract
Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet its progress is stifled by the high cost of collecting embodiment-aware teleoperation data. While abundant egocentric videos of human hands offer a scalable alternative, the profound discrepancies in appearance, articulation, and camera viewpoints between human and robotic data raise significant challenges for co-training. Though existing general image-editing models demonstrate strong capabilities, they lack necessary embodiment-specific priors to fully bridge this gap. In this work, we present HandEdit, a unified large-scale embodiment-aware image-editing dataset and benchmark specifically designed to transform human hands and arms into various dexterous robotic embodiments within egocentric frames. HandEdit comprises over 200M editing instances derived from five diverse source datasets, covering 26 distinct URDFs, including 13 hand-only and 13 hand-arm configurations. Alongside the dataset, we establish a unified benchmark protocol with two tracks: Hand-only and Hand-Arm, supporting URDF-conditioned evaluation. We conduct extensive evaluations of 11 representative image-editing baselines using a multi-dimensional metric suite, including generic similarity metrics, VLM-based judgment, and embodiment-aware metrics. HandEdit serves as a critical resource at the intersection of image editing and robotics: it advances embodiment-aware editing models while enabling scalable dexterous robotic learning from abundant human video data, paving the way for more generalizable Embodied AI.
Chinese Translation
灵巧手的机器人操控是具身人工智能的基石,但由于收集具身感知的遥操作数据成本高昂,其进展受到限制。虽然大量以自我为中心的人手视频提供了一种可扩展的替代方案,但人类与机器人数据在外观、关节运动和摄像机视角上的显著差异给共同训练带来了重大挑战。尽管现有的通用图像编辑模型展现了强大的能力,但它们缺乏必要的具身特定先验知识,无法完全弥合这一差距。在本研究中,我们提出了HandEdit,这是一个统一的大规模具身感知图像编辑数据集和基准,专门设计用于在以自我为中心的框架内将人类手和手臂转化为各种灵巧机器人具身。HandEdit包含来自五个不同源数据集的超过2亿个编辑实例,涵盖26种不同的URDF,包括13种仅手部和13种手臂配置。除了数据集,我们还建立了一个统一的基准协议,分为两个轨道:仅手部和手臂,支持基于URDF的评估。我们使用多维度指标套件对11个代表性的图像编辑基线进行了广泛评估,包括通用相似性指标、基于VLM的判断和具身感知指标。HandEdit在图像编辑与机器人技术的交叉点上提供了重要资源:它推动了具身感知编辑模型的发展,同时使得从丰富的人类视频数据中进行可扩展的灵巧机器人学习成为可能,为更具普遍性的具身人工智能铺平了道路。
cs.RO / 21 / 2608.12198

Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment

基于学习的自动驾驶行为规划:现实世界的集成与部署
Busch, Jean-Pierre, Linden, Guido, Bergmann, Jan, Eckstein, Lutz
Abstract
Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches to improve the driving behavior of automated vehicles, especially in complex environments. However, their complex nature and lack of transparency can hinder explainability and trustworthiness and complicate safety assurance. Motivated by these challenges, we propose a hybrid planning architecture that combines the advantages of machine learning with the verifiability and the determinism of classical approaches. Specifically, we developed a deep neural network to interpret complex traffic scenes and propose driving behavior, while an optimization-based supervision layer validates this proposal and enforces explicit drivability and safety constraints. We evaluate the learned planner's driving behavior in open-loop studies on real-world urban data, discuss system integration aspects for stable closed-loop operation, and report results from real-world deployment on our research vehicle karl..
Chinese Translation
近期在机器学习和深度学习领域的研究表明,基于学习的运动规划方法在改善自动驾驶车辆的驾驶行为方面具有潜力,尤其是在复杂环境中。然而,这些方法的复杂性和缺乏透明性可能会妨碍可解释性和可信度,并使安全保障变得复杂。针对这些挑战,我们提出了一种混合规划架构,结合了机器学习的优势与经典方法的可验证性和确定性。具体而言,我们开发了一种深度神经网络来解读复杂的交通场景并提出驾驶行为,而基于优化的监督层则验证这一提议并强制执行明确的可驾驶性和安全约束。我们在真实城市数据的开放环路研究中评估了学习规划器的驾驶行为,讨论了稳定闭环操作的系统集成方面,并报告了在我们的研究车辆karl上的真实世界部署结果。
计算机视觉 (Computer Vision)
99
cs.CV / 1 / 2608.11263

GeoUniPR: A Geometry-Consistent Unified Framework for Cross-Modal Place Recognition

GeoUniPR:一个几何一致的跨模态地点识别统一框架
Kim, Wonbong, Xiao, Jiatong, Li, Rui, Wang, Xufei, Gu, Qiwen, Zhao, Junqiao, Ye, Chen, Chen, Guang
Abstract
Cross-modal place recognition (CMPR) aims to identify the same location across heterogeneous sensing modalities, such as vision and LiDAR. Existing methods commonly bridge the modality gap using complex alignment modules, multi-stage training, or full fine-tuning of pretrained backbones. In this work, we revisit CMPR from the perspective of geometric consistency and propose GeoUniPR, a unified and concise geometry-consistent framework. GeoUniPR reduces cross-modal discrepancy at the representation level by projecting LiDAR point clouds into the camera perspective to construct Geometry-Consistent depth image views (DIV), which establish direct RGB-LiDAR correspondence. We further augment DIV with native LiDAR cues, including intensity and surface-normal information, yielding a multi-channel geometric representation that improves structural consistency. Based on this representation, GeoUniPR learns a unified embedding space using two modality-specific ViT-based encoders with identical architectures, trained through parameter-efficient adaptation without auxiliary alignment modules, multi-stage training, or full backbone fine-tuning. In addition, we introduce Spatially-Consistent InfoNCE (SC-InfoNCE), a CMPR-specific contrastive objective that suppresses distance-induced false negatives under spatial continuity. Extensive experiments on KITTI and KITTI-360 demonstrate that GeoUniPR achieves state-of-the-art (SOTA) performance in both same-modal and cross-modal place recognition, with strong cross-dataset generalization.
Chinese Translation
跨模态地点识别(CMPR)旨在识别不同传感模态(如视觉和激光雷达)下的同一位置。现有方法通常通过复杂的对齐模块、多阶段训练或对预训练骨干网络的全面微调来弥合模态差异。在本研究中,我们从几何一致性的角度重新审视CMPR,并提出GeoUniPR,一个统一且简洁的几何一致框架。GeoUniPR通过将激光雷达点云投影到相机视角,构建几何一致的深度图像视图(DIV),从而在表示层面减少跨模态差异,建立直接的RGB-激光雷达对应关系。我们进一步利用原生激光雷达线索(包括强度和表面法线信息)增强DIV,生成多通道几何表示,从而改善结构一致性。基于此表示,GeoUniPR使用两个具有相同架构的模态特定ViT编码器学习统一的嵌入空间,通过参数高效的适应训练,而无需辅助对齐模块、多阶段训练或全面的骨干微调。此外,我们引入空间一致信息对比损失(SC-InfoNCE),这是一个针对CMPR的特定对比目标,能够抑制因距离引起的空间连续性下的假负样本。在KITTI和KITTI-360上的大量实验表明,GeoUniPR在同模态和跨模态地点识别中均实现了最新的(SOTA)性能,并具有强大的跨数据集泛化能力。
cs.CV / 2 / 2608.11285

SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation

SegPAR:面向语义分割的类中心决策基础稀疏攻击
Song, Dongsu, GO, DaeYun, Seo, Boseung, Jung, Jay Hoon
Abstract
Despite the practical relevance of sparse decision-based black-box threats, they have received limited attention in semantic segmentation. To bridge this gap, we adapt the most representative decision-based black-box sparse attacks from the classification domain to serve as baselines, establishing a rigorous benchmark for this underexplored setting. In this context, we demonstrate that one of the existing methods suffers from severe query inefficiency due to its image-centric pixel accumulation, which rapidly exhausts query budgets across the vast image space. To overcome this, we propose SegPAR, a novel decision-based framework that shifts to a class-centric exploration paradigm. Furthermore, to eliminate the misleading feedback generated by standard decision rewards during pixel accumulation, we introduce a novel discrepancy reward. Extensive experiments show that SegPAR significantly outperforms black-box baselines in sparsity efficiency and MIoU reduction, while remaining competitive with white-box sparse attacks. Code is available at \href{https://github.com/KAU-QuantumAILab/SegPAR}{https://github.com/KAU-QuantumAILab/SegPAR}.
Chinese Translation
尽管稀疏决策基础黑箱威胁在实际应用中具有重要意义,但在语义分割领域却受到的关注有限。为填补这一空白,我们将分类领域中最具代表性的决策基础黑箱稀疏攻击进行改编,以作为基准,建立一个严格的基准测试,以应对这一尚未深入探索的设置。在此背景下,我们展示了现有方法之一由于其以图像为中心的像素累积而遭受严重的查询效率低下,这在广阔的图像空间中迅速耗尽了查询预算。为了解决这一问题,我们提出了SegPAR,一种新颖的决策基础框架,转向类中心探索范式。此外,为了消除在像素累积过程中标准决策奖励产生的误导性反馈,我们引入了一种新颖的差异奖励。大量实验表明,SegPAR在稀疏效率和MIoU(Mean Intersection over Union)降低方面显著优于黑箱基准,同时在与白箱稀疏攻击的竞争中保持竞争力。代码可在 exttt{https://github.com/KAU-QuantumAILab/SegPAR} 获取。
cs.CV / 3 / 2608.11287

CLEAR: Class-wise Expert Aggregation with Structured Sampling for Long-Tailed Classification

CLEAR:基于结构化采样的类别专家聚合用于长尾分类
Lim, Gawon
Abstract
Long-tailed classification poses a reliability challenge because models trained on imbalanced data are unevenly reliable across frequent and underrepresented classes. While existing methods address imbalance through re-balancing, adjustment, representation learning, or multi-expert modeling, they rarely estimate which expert should be trusted for each class. This paper proposes CLEAR (Class-wise reLiability-aware Expert Aggregation for long-tailed Recognition), a modular ensemble framework for long-tailed classification. CLEAR generates diverse experts through threshold-based structured sampling while preserving the full label space, then estimates a class-wise trust score for each expert using a smoothed class-wise precision formulation. During inference, expert predictions are combined through class-wise generalized product-of-experts aggregation, allowing different experts to be emphasized for different classes. Experiments on CIFAR-100-LT, ImageNet-LT, and Places-LT across multiple backbones show that CLEAR achieves competitive overall accuracy and particularly strong few-shot performance. These results support class-wise expert reliability as a useful design principle for long-tailed ensemble learning.
Chinese Translation
长尾分类面临可靠性挑战,因为在不平衡数据上训练的模型在频繁类和欠代表类之间的可靠性不均衡。虽然现有方法通过重平衡、调整、表示学习或多专家建模来解决不平衡问题,但它们很少评估每个类别应信任哪个专家。本文提出了CLEAR(基于类别的可靠性感知专家聚合用于长尾识别),这是一个用于长尾分类的模块化集成框架。CLEAR通过基于阈值的结构化采样生成多样化的专家,同时保留完整的标签空间,然后使用平滑的类别精度公式为每个专家估计类别可信度分数。在推理过程中,专家预测通过类别化的专家聚合的广义乘积进行组合,从而允许不同的专家在不同类别中被强调。在CIFAR-100-LT、ImageNet-LT和Places-LT等多个基准上进行的实验表明,CLEAR在整体准确性上具有竞争力,并且在少样本性能上尤其强。这些结果支持类别专家可靠性作为长尾集成学习的有用设计原则。
cs.CV / 4 / 2608.11292

Self-Evolving Code-with-Image Reasoning

自我演化的代码与图像推理
Yang, Tianze, Wu, Liang, Sun, Ruitong, Shi, Yucheng, Wang, Yanqiao, Darbari, Mayank, Liu, Ninghao, Sun, Jin, Hong, Liangjie
Abstract
Multimodal models increasingly reach for tools when solving visual tasks (crop, zoom, rotate, brighten), a paradigm known as thinking-with-images. The central challenge is one of perception: tools mostly serve to expose visual evidence, reasoning over that evidence stays in language, and most targets are ones a human could in principle determine by inspection. Some visual questions, however, are not bottlenecked by perception: recovering their answers requires executing a multi-step visual algorithm over the pixels. On such questions a model often names the correct algorithm at once yet still answers wrong, because language can describe an algorithm without being able to run one. Code-with-Image crosses that line: given nothing but a Python interpreter, the model must implement a genuine visual algorithm in code to solve the task; the program itself becomes the reasoning. The bottleneck then shifts from executing code to deciding which algorithm to implement. So we let the model teach itself: a training-free reflection loop studies its own failed programs, tests repairs against constructive ground truth, and keeps what survives as portable skills. On our Code-with-Image Bench (CwI-Bench), thirty task families induced by hidden visual computations with disjoint learning and evaluation splits, even GPT-5.6-luna stays below 30% with tool-free chain of thought; given a bare interpreter it reaches 43%, and with skills evolved through its own executable reflection, 67%. The open 27B model climbs the same ladder (9% $\rightarrow$ 33% $\rightarrow$ 56%), and the skills are plain text, transferable across scales and families. When code carries the reasoning, debugging code becomes debugging reasoning.
Chinese Translation
多模态模型在解决视觉任务时越来越依赖工具(裁剪、缩放、旋转、调亮),这一范式被称为图像思维。核心挑战在于感知:工具主要用于揭示视觉证据,而对这些证据的推理仍然停留在语言层面,大多数目标是人类原则上可以通过观察来确定的。然而,一些视觉问题并不受限于感知:恢复其答案需要在像素上执行多步骤的视觉算法。在这些问题上,模型往往能够立即命名正确的算法,但仍然回答错误,因为语言可以描述算法,却无法执行算法。代码与图像(Code-with-Image)跨越了这一界限:模型仅凭 Python 解释器,必须在代码中实现真正的视觉算法来解决任务;程序本身成为推理的载体。瓶颈因此从执行代码转移到决定实现哪个算法。因此,我们让模型自我学习:一个无训练的反思循环研究其失败的程序,测试修复与构造性真实值的对比,并保留存活下来的部分作为可移植技能。在我们的代码与图像基准(CwI-Bench)上,三十个任务家族由隐藏的视觉计算诱导,具有不相交的学习和评估分割,即使是 GPT-5.6-luna 在无工具的思维链下也保持在 30% 以下;在仅有一个基本解释器的情况下,它达到了 43%,而通过其自身可执行反思演化出的技能则达到了 67%。开放的 27B 模型也攀登了同样的阶梯(9% → 33% → 56%),这些技能是纯文本的,可以跨规模和家族转移。当代码承载推理时,调试代码便成了调试推理。
cs.CV / 5 / 2608.11317

Clinical Feasibility of Low-Magnification Fluorescence Imaging for Breast Cancer Margin Detection Using Texture Analysis and Deep Learning

低倍荧光成像在乳腺癌切缘检测中的临床可行性:基于纹理分析和深度学习
Afshin, Pouya, Niu, Tianling, Lu, Tongtong, Helminiak, David, Jorns, Julie, Patton, Mollie, Yen, Tina, Ye, Donghye, Yu, Bing
Abstract
High-resolution images of unprocessed surgical breast tissue can be obtained using microscopy with ultraviolet surface excitation (MUSE). This technique is considered a promising method for checking surgical margins during breast cancer surgery. In this study, MUSE images at 4x and 10x magnifications were compared using patch-level classification methods. Texture analysis (TA) based on local binary patterns (LBP) and deep learning (DL) with a base Vision Transformer (ViT) model were used. Both methods achieved similar performance at both magnifications. Using DL method, both 4x and 10x magnifications achieved 96.30% sensitivity, 100% specificity and 98.18% accuracy. Using TA method, 4x achieved better specificity (100% vs 93.33%) and 10x yielded higher sensitivity (100% vs 93.33%), but both had the same accuracy (96.67%). No clear improvement in performance was observed with 10x magnification. These results show that 4x imaging achieves the same diagnostic accuracy as 10x imaging. At the same time, 4x offers a larger field of view and faster image capture. Therefore, lower magnification can be effectively used in MUSE systems for accurate and efficient intraoperative margin assessment.
Chinese Translation
通过紫外表面激发显微镜(MUSE)可以获得未经处理的外科乳腺组织的高分辨率图像。这种技术被认为是一种有前景的方法,用于在乳腺癌手术中检查外科切缘。在本研究中,比较了4倍和10倍放大倍数下的MUSE图像,采用了基于补丁级分类的方法。使用了基于局部二值模式(LBP)的纹理分析(TA)和基于基础视觉变换器(ViT)模型的深度学习(DL)。两种方法在两个放大倍数下均取得了相似的性能。使用深度学习方法,4倍和10倍放大均达到了96.30%的灵敏度、100%的特异性和98.18%的准确性。使用纹理分析方法,4倍放大在特异性上表现更佳(100%对93.33%),而10倍放大在灵敏度上更高(100%对93.33%),但两者的准确性相同(96.67%)。在10倍放大下未观察到明显的性能提升。这些结果表明,4倍成像的诊断准确性与10倍成像相同。同时,4倍成像提供了更大的视野和更快的图像捕获。因此,低倍放大可以有效用于MUSE系统,以实现准确和高效的术中切缘评估。
cs.CV / 6 / 2608.11335

Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation

双域跨模态解码用于临床文本指导的医学图像分割
Rahman, Md Maklachur, Hammond, Tracy
Abstract
Clinical text can narrow down what to segment, but recent text-guided designs emphasize spatial alignment while overlooking frequency content that governs texture and boundaries. We propose Dual-Domain Cross-Modal Decoding (DD-CMD) for clinical text-guided pulmonary infection segmentation, integrating two complementary forms of language guidance during decoding. In the spatial domain, Text-Guided Spatial Cross-Attention (TGSA) aligns multi-scale visual tokens with text semantics and updates features through gated residual fusion. In the frequency domain, Spectral-Text Adaptive Modulation (STAM) applies a 2D DCT to compute learnable band-energy statistics and predicts text-conditioned FiLM parameters to recalibrate decoder channels for frequency-aware decoding. DD-CMD embeds TGSA and STAM into a coarse-to-fine decoder (7x7 to 56x56) and restores full-resolution masks using a lightweight two-stage refinement module. Experiments on QaTa-COV19 and MosMedData+ show that DD-CMD achieves 91.46% Dice / 84.26% mIoU and 81.95% Dice / 69.42% mIoU, respectively, with average gains of +1.96 Dice and +2.67 mIoU over the strongest prior baselines. Code: https://github.com/maklachur/DD-CMD.
Chinese Translation
临床文本可以缩小分割范围,但近期的文本指导设计强调空间对齐,而忽视了控制纹理和边界的频率内容。我们提出了双域跨模态解码(Dual-Domain Cross-Modal Decoding, DD-CMD)用于临床文本指导的肺部感染分割,在解码过程中整合了两种互补的语言指导形式。在空间域中,文本指导空间跨注意力(Text-Guided Spatial Cross-Attention, TGSA)将多尺度视觉标记与文本语义对齐,并通过门控残差融合更新特征。在频率域中,谱-文本自适应调制(Spectral-Text Adaptive Modulation, STAM)应用二维离散余弦变换(2D DCT)计算可学习的带能量统计,并预测文本条件的FiLM参数,以重新校准解码器通道,实现频率感知解码。DD-CMD将TGSA和STAM嵌入到一个粗到细的解码器(从7x7到56x56),并使用轻量级的两阶段精细化模块恢复全分辨率掩膜。在QaTa-COV19和MosMedData+上的实验表明,DD-CMD分别达到了91.46%的Dice / 84.26%的mIoU和81.95%的Dice / 69.42%的mIoU,平均提升为+1.96 Dice和+2.67 mIoU,超越了最强的先前基线。代码: https://github.com/maklachur/DD-CMD。
cs.CV / 7 / 2608.11367

Gaze Target Estimation Anywhere with Concepts

基于概念的任意视线目标估计
Cao, Xu, Yang, Houze, Gunda, Vipin, Zhou, Zhongyi, Xu, Tianyu, Kowdle, Adarsh, Kim, Inki, Rehg, James M.
Abstract
Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis. As a result, detection errors can cascade and lead to failure. Moreover, these prior works lack the flexibility of specifying the gaze analysis task via natural language prompting, an approach which has been shown to have significant benefits in convenience and scalability for other image analysis tasks. To overcome these limitations, we introduce the Promptable Gaze Target Estimation (PGE) task, a new end-to-end, concept-driven paradigm for gaze analysis. PGE conditions gaze prediction on flexible user text or visual prompts (e.g., "the boy in the red shirt" or "person in point [0.52, 0.48]") to identify a specific subject for gaze analysis. This approach integrates subject localization with gaze estimation, and eliminates the rigid dependency on intermediate analysis stages. We develop a scalable data engine to generate Gaze-Co (Gaze Estimation with Concepts), a dataset and benchmark of 120K high-quality, prompt-annotated image pairs. We also propose GazeAnywhere, the first model designed for PGE. GazeAnywhere uses a transformer-based detector to fuse features from frozen encoders and simultaneously solves subject localization, in/out-of-frame presence, and gaze target heatmap estimation. GazeAnywhere achieves state-of-the-art performance on multiple PGE benchmarks, setting a strong baseline for this new problem even on a difficult out-of-domain, real-world clinical dataset. GazeAnywhere is open-sourced in github.com/IrohXu/GazeAnywhere.
Chinese Translation
从自然场景图像中估计人类视线目标是一项重要且艰巨的任务。现有的方法主要采用脆弱的多阶段流程,这些流程需要明确的输入,如头部边界框和人体姿态,以识别视线分析的对象。因此,检测错误可能会级联并导致失败。此外,以往的研究缺乏通过自然语言提示来指定视线分析任务的灵活性,而这种方法在其他图像分析任务中已被证明在便利性和可扩展性方面具有显著优势。为克服这些局限性,我们提出了可提示视线目标估计(Promptable Gaze Target Estimation, PGE)任务,这是一种新的端到端、以概念驱动的视线分析范式。PGE根据灵活的用户文本或视觉提示(例如,“穿红色衬衫的男孩”或“位于点 [0.52, 0.48] 的人”)来条件化视线预测,以识别特定的视线分析对象。这种方法将对象定位与视线估计相结合,消除了对中间分析阶段的严格依赖。我们开发了一个可扩展的数据引擎,以生成Gaze-Co(基于概念的视线估计),这是一个包含12万个高质量、带提示注释的图像对的数据集和基准。我们还提出了GazeAnywhere,这是第一个为PGE设计的模型。GazeAnywhere使用基于变换器的检测器来融合来自冻结编码器的特征,同时解决对象定位、框内/框外存在性和视线目标热图估计。GazeAnywhere在多个PGE基准上实现了最先进的性能,即使在困难的域外真实世界临床数据集上也建立了强大的基线。GazeAnywhere已在github.com/IrohXu/GazeAnywhere上开源。
cs.CV / 8 / 2608.11422

COGENT: Counterfactual Gaussian Explanations for Volumetric Medical Images

COGENT:体积医学图像的反事实高斯解释
Rząsa, Dorian, Zabdyr, Bartosz, Piekarz, Krzysztof, Grzywaczewski, Jakub, Sobieski, Bartlomiej, Biecek, Przemyslaw, Świderska-Chadaj, Żaneta, Śliwicka, Olga, Spurek, Przemysław, Świebocka-Więk, Joanna
Abstract
Explainability is essential for deploying deep learning models in high-stakes medical applications. Existing explainability methods for volumetric imaging predominantly operate in voxel space, overlooking the structured representations introduced by recent advances in 3D scene modeling. We present COGENT (Counterfactual Gaussian Explanations), a framework that generates counterfactual explanations directly in the parameter space of Gaussian-based volumetric representations. Built upon MedGS and the Sybil lung cancer risk prediction model, COGENT optimizes selected Gaussian primitives through a differentiable rendering pipeline, enabling gradients from the downstream predictor to identify representation components that most influence model decisions. Unlike conventional pixel- or voxel-level attribution methods, our approach formulates explainability as a counterfactual optimization problem over an explicit 3D scene representation, producing sparse and spatially localized explanations while preserving anatomical consistency. We evaluate COGENT on lung CT scans using quantitative comparisons with existing explainability methods together with qualitative analysis by medical experts. The results demonstrate that representation-space counterfactual optimization provides clinically meaningful explanations while offering a new perspective on interpreting volumetric deep learning models.
Chinese Translation
可解释性对于在高风险医疗应用中部署深度学习模型至关重要。现有的体积成像可解释性方法主要在体素空间中操作,忽视了最近在3D场景建模中引入的结构化表示。我们提出了COGENT(反事实高斯解释),这是一个直接在基于高斯的体积表示的参数空间中生成反事实解释的框架。COGENT建立在MedGS和Sybil肺癌风险预测模型之上,通过可微渲染管道优化选定的高斯原语,使下游预测器的梯度能够识别对模型决策影响最大的表示组件。与传统的像素或体素级归因方法不同,我们的方法将可解释性公式化为一个显式3D场景表示上的反事实优化问题,生成稀疏且空间局部化的解释,同时保持解剖一致性。我们在肺部CT扫描上评估了COGENT,通过与现有可解释性方法的定量比较以及医学专家的定性分析。结果表明,表示空间的反事实优化提供了临床上有意义的解释,同时为解释体积深度学习模型提供了新的视角。
cs.CV / 9 / 2608.11425

VLMs Win a Systematic Evaluation of Underwater Image Reconstruction

VLMs在水下图像重建的系统评估中获胜
Aghajanzadeh, Sara, Wang, Yingxue, Bagdonaviciute, Ieva, Forsyth, David
Abstract
Underwater image restoration consists of recovering an image which looks like there is no water present. To date, evaluation has not been systematic. This paper describes a systematic evaluation pipeline for underwater reconstruction, which can be used to assess a method for accuracy; consistency of reconstruction over camera moves; and the effect of water parameters. We use this pipeline to evaluate a range of current procedures, from models constructed using explicit but approximate physical models of scattering to Vision-Language Models (VLMs which are not currently trained with explicit physical models). Overall, VLMs wholly and significantly outperform physically based models in our evaluation, likely because of the importance of a strong image prior. Results on images of real underwater scenes strongly confirm the evaluation.
Chinese Translation
水下图像恢复是指恢复看起来没有水的图像。迄今为止,评估尚未系统化。本文描述了一种水下重建的系统评估流程,可用于评估方法的准确性;在相机移动过程中的重建一致性;以及水参数的影响。我们使用该流程评估了一系列当前的程序,从使用显式但近似的散射物理模型构建的模型到未使用显式物理模型训练的视觉-语言模型(Vision-Language Models, VLMs)。总体而言,在我们的评估中,VLMs在性能上完全且显著优于基于物理的模型,这可能是由于强图像先验的重要性。对真实水下场景图像的结果强烈确认了这一评估。
cs.CV / 10 / 2608.11452

TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator for Poetry-to-Image Generation

唐诗基准:用于诗歌到图像生成的多维基准和评分条件评估器
Hu, Haoqi, Luo, Tongji, Zhang, Li, Zhou, Boning
Abstract
Text-to-image (T2I) models are increasingly asked to illustrate literary and cultural content, yet we cannot measure how well an image renders the meaning of a poem. The task is many-sided: a good illustration must be visually sound, faithful to the poem's imagery and scene, culturally and stylistically apt, free of spurious text, and true to its emotion, and its deepest requirements, imagery and especially implicit emotion, are never stated in the words. Existing metrics (CLIPScore, BLIPScore, VQAScore) reward literal text-image correspondence and so cannot tell whether an illustration succeeds, let alone why, or even separate the best model from the worst. We introduce TangPoetryBench, a multi-dimensional benchmark of 1,280 images (320 classical Chinese Tang poems x 4 state-of-the-art T2I models) with quality-controlled human annotations across ten dimensions. Analyzing this data, we reveal the shared and model-specific strengths and weaknesses of current T2I models, including their ability to evoke a poem's implicit emotion. We further introduce PoemAutoEvaluator (PAE), an open, rubric-conditioned evaluator that reaches parity with a strong proprietary judge (Claude), generalizes to an unseen generator and a second poetic tradition (Song Ci), and lets the benchmark scale to new images without fresh human annotation. We release the benchmark, annotations, and evaluator.
Chinese Translation
文本到图像(T2I)模型越来越多地被要求插图文学和文化内容,但我们无法衡量一幅图像在多大程度上呈现了诗歌的意义。这个任务是多方面的:好的插图必须在视觉上合理,忠实于诗歌的意象和场景,文化和风格上适宜,避免虚假的文本,并真实地传达其情感,而其最深层的要求,即意象,尤其是隐含情感,往往并未在文字中明确表达。现有的指标(CLIPScore、BLIPScore、VQAScore)奖励字面上的文本与图像的对应关系,因此无法判断插图是否成功,更不用说原因,甚至无法区分最佳模型与最差模型。我们推出了唐诗基准,这是一个包含1280幅图像(320首经典中国唐诗 x 4个最先进的T2I模型)的多维基准,涵盖十个维度的质量控制人类注释。通过分析这些数据,我们揭示了当前T2I模型的共同优势和特定模型的优缺点,包括它们唤起诗歌隐含情感的能力。我们进一步介绍了PoemAutoEvaluator(PAE),一个开放的、基于评分标准的评估器,其性能与强大的专有评审(Claude)相当,能够推广到未见过的生成器和第二种诗歌传统(宋词),并使基准能够在没有新的人类注释的情况下扩展到新图像。我们发布了基准、注释和评估器。
cs.CV / 11 / 2608.11458

Multi-Agent Target-Existence Verification and Learned Mask Geometry Refinement: Winning Report of the MeViS-Text Track at the 8th LSVOS Challenge 2026

多智能体目标存在性验证与学习的掩膜几何精细化:2026年第八届大规模视频目标分割挑战赛MeViS-Text赛道获胜报告
Lee, Jungyoon, Lim, Gyuil, Kim, Doeon, Kim, Seong-heum
Abstract
We present the first-place solution to the MeViS-Text track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge 2026: referring video object segmentation guided by written motion expressions, including deceptive no-target expressions that match no object in the video and must yield empty masks in every frame. Our pipeline, SSUPER, resolves each expression into a visual concept, generates full-video candidate masklets with SAM~3.1, and selects target IDs. At every reasoning stage, three heterogeneous multimodal large language models independently execute the same stage-specific prompt before a single synthesis pass commits one schema-validated verdict. Although this system rejects every no-target expression in validation, the leaderboard reveals that a substantial share of test no-target cases still slips through. The reason is that hard negatives name a plausible object and fail only under the complete temporal predicate, so when selection and existence are decided together, a category-plausible masklet anchors the verdict. Hence, we decouple existence verification into an independent multi-agent audit of the full predicate (category, count, action, trajectory, event order, and semantic role) that distinguishes absence from temporary invisibility, discounts apparent motion caused by camera movement, and requires contradicting evidence rather than mere uncertainty for a no-target verdict. Without any new segmentation call, this audit recovers most of the residual no-target errors. A training-data-only StyleRefiner then aligns mask geometry with the annotation style of MeViSv2 while preserving every presence decision by construction, showing that once the semantics are fixed, part of the remaining error is stylistic rather than semantic. The complete system reaches a Final score of 0.9081339614 on the official challenge leaderboard.
Chinese Translation
我们展示了2026年第八届大规模视频目标分割(LSVOS)挑战赛MeViS-Text赛道的第一名解决方案:通过书面运动表达引导的视频目标分割,包括与视频中没有对象匹配的欺骗性无目标表达,这些表达在每一帧中必须产生空掩膜。我们的管道SSUPER将每个表达解析为视觉概念,使用SAM~3.1生成全视频候选掩膜,并选择目标ID。在每个推理阶段,三个异构多模态大型语言模型独立执行相同的阶段特定提示,然后通过一次合成传递提交一个经过模式验证的裁决。尽管该系统在验证中拒绝了每个无目标表达,但排行榜显示,仍有相当一部分测试无目标案例漏网。原因在于,困难的负样本命名了一个合理的对象,仅在完整的时间谓词下失败,因此当选择和存在性一起决定时,一个类别合理的掩膜锚定了裁决。因此,我们将存在性验证解耦为对完整谓词(类别、数量、动作、轨迹、事件顺序和语义角色)的独立多智能体审计,以区分缺失与暂时不可见,排除由相机移动引起的明显运动,并要求相互矛盾的证据而非仅仅是不确定性来做出无目标裁决。在没有任何新的分割调用的情况下,这一审计恢复了大部分剩余的无目标错误。一个仅基于训练数据的StyleRefiner随后将掩膜几何与MeViSv2的注释风格对齐,同时通过构造保留每个存在决策,表明一旦语义固定,剩余错误的一部分是风格性的而非语义性的。完整系统在官方挑战排行榜上达到了最终得分0.9081339614。
cs.CV / 12 / 2608.11472

Gaussian Meta-Space Augmentation for Stacking Ensembles in Multimodal IPMN Risk Stratification

用于多模态IPMN风险分层的高斯元空间增强堆叠集成
Nelson, Max A., Tasci, Eminenur Sen, Wang, Zhixiang, Zhou, Zongwei, Aktas, Halil Ertugrul, Bejar, Andrea M., Keles, Elif, Hong, Ziliang, Taflan, Sıtkı Safa, Tasci, Muhammed Enes, Miller, Frank H., Wallace, Michael B., Keswani, Rajesh N., Durak, Gorkem, Bagci, Ulas
Abstract
Pancreatic cancer is among the most lethal malignancies; risk stratification of intraductal papillary mucinous neoplasms (IPMNs) offers a crucial opportunity for early intervention but typically requires invasive tissue biopsy. Dominant vision-based approaches, including radiomics and deep learning, provide promising but initially separate discrimination opportunities. Similarly, multisequence MRI (T1W/T2W) and anatomically decomposed (head, body and tail) analysis of the pancreas provide additional and potentially complementary signals. Effective fusion of this information is crucial in ordinal IPMN dysplasia risk prediction and can be accomplished via a meticulously regularized and calibrated ensemble stacking combiner. We present cUPMI, a class-conditional Gaussian augmentation of a combiner's log-probability meta-features, and test it on various prediction paradigms. In our multi-center analysis, we find cUPMI adds limited value to properly regularized L2-logistic binary classification stacks, but consistently regularizes higher-capacity tree combiners in the binary and radiomics-only setting (RF +0.015 and XGBoost +0.024 binary AUC, positive in all seeds). Its cleanest ordinal benefit appears for XGBoost on an 8-stream radiomics task (3-class no < low < high, +0.022 QWK in all seeds). Separately, fold-locked fusion of radiomics and 2.5D CNN streams yields the strongest overall model, an RF stack reaching QWK 0.595 (95% CI [0.54, 0.64]) and binary AUC 0.839, surpassing radiomics, 2.5D ResNet, and 3D DenseNet-121 baselines.
Chinese Translation
胰腺癌是最致命的恶性肿瘤之一;对管内乳头状粘液性肿瘤(IPMN)的风险分层为早期干预提供了关键机会,但通常需要侵入性组织活检。以视觉为主的主流方法,包括放射组学和深度学习,提供了有前景但最初是分开的鉴别机会。同样,多序列MRI(T1W/T2W)和胰腺的解剖分解(头、体和尾)分析提供了额外且可能互补的信号。有效融合这些信息对于IPMN的序数发育不良风险预测至关重要,可以通过精心正则化和校准的集成堆叠组合器实现。我们提出了cUPMI,这是一种组合器对数概率元特征的类条件高斯增强,并在各种预测范式上进行了测试。在我们的多中心分析中,我们发现cUPMI对适当正则化的L2-logistic二元分类堆叠的增值有限,但在二元和仅放射组学的设置中(RF +0.015和XGBoost +0.024二元AUC,在所有种子中均为正)始终能正则化更高容量的树组合器。其在8流放射组学任务(3类无<低<高,所有种子中+0.022 QWK)上对XGBoost的序数效益最为明显。此外,折叠锁定的放射组学和2.5D CNN流的融合产生了最强的整体模型,一个RF堆叠达到了QWK 0.595(95% CI [0.54, 0.64])和二元AUC 0.839,超越了放射组学、2.5D ResNet和3D DenseNet-121基线。
cs.CV / 13 / 2608.11474

Test-Time Hallucination Control in Large Vision-Language Models

大型视觉-语言模型中的测试时幻觉控制
Tamjidi, Mehran, Dastmalchi, Hamidreza, Cheraghian, Ali, Alimoradijazi, Mohammadreza, An, Aijun, Rahmani, Hossein
Abstract
Object Hallucination in large vision-language models (LVLMs), where models generate non-factual content about input images, remains a critical barrier to their reliability in real-world applications. Existing mitigation strategies can be categorized into training-based and training-free methods. Training-based methods often achieve strong performance but are costly, requiring extensive computational resources, large-scale data, and time-consuming fine-tuning. Training-free approaches are particularly appealing due to their efficiency. However, existing training-free methods either require multiple decoding rounds, which adds computational overhead, or modify internal states in a model-specific way that risks degrading pretrained knowledge. We propose Test-Time Hallucination Mitigation (TTH) method, a novel training-free method that addresses both limitations. TTH introduces a token-validator module, implemented as a zero-shot Multi-Modal Classifier (MMC), to generate auxiliary logits grounded in the input image. These logits are fused with the original LVLM outputs at the token level for object tokens selected from a candidate pool. An entropy-based weighting scheme is then applied to enable robust and accurate predictions. Extensive experiments across multiple LVLM families and diverse benchmarks demonstrate that TTH consistently improves accuracy and robustness, underscoring its generalizability and practical effectiveness. Code is released at https://github.com/Mehran-TAM/TTH
Chinese Translation
在大型视觉-语言模型(LVLMs)中,物体幻觉指的是模型生成关于输入图像的非事实内容,这仍然是其在现实应用中可靠性的一个关键障碍。现有的缓解策略可以分为基于训练的方法和无训练的方法。基于训练的方法通常表现良好,但成本高昂,需要大量的计算资源、大规模的数据以及耗时的微调。无训练的方法因其高效性而特别受欢迎。然而,现有的无训练方法要么需要多轮解码,增加了计算开销,要么以模型特定的方式修改内部状态,可能会导致预训练知识的退化。我们提出了一种测试时幻觉缓解(Test-Time Hallucination Mitigation, TTH)方法,这是一种新颖的无训练方法,旨在解决这两种限制。TTH引入了一个令牌验证模块,作为零样本多模态分类器(Multi-Modal Classifier, MMC)实现,以生成基于输入图像的辅助logits。这些logits在令牌级别与从候选池中选择的物体令牌的原始LVLM输出进行融合。随后应用基于熵的加权方案,以实现稳健和准确的预测。在多个LVLM家族和多样化基准上的广泛实验表明,TTH始终提高了准确性和鲁棒性,突显了其通用性和实际有效性。代码已发布在 https://github.com/Mehran-TAM/TTH
cs.CV / 14 / 2608.11498

Language-Structured Relational Q-Learning for Threat-Aware Control in Safety-Critical Driving

基于语言结构的关系Q学习在安全关键驾驶中的威胁感知控制
Humnabadkar, Aditya, Zhang, Huaizhong, Behera, Ardhendu
Abstract
Natural-language-based scenario generation offers an intuitive means of describing rare and complex driving interactions, yet it is still uncertain whether training with language-structured data leads to truly adaptive control policies. We propose Language-Structured Relational Q-Learning, instantiated through an Ego-Centric Relational Q-Network (ERQ-Net), which jointly learns inter-vehicle relevance and action values from dynamic traffic graphs. Language descriptions define surrounding-vehicle behaviours during training, while prompts and semantic actor roles are hidden from the policy. ERQ-Net must therefore infer threat relevance solely from observable kinematics and interactions. Across 2,500 safety-critical scenarios, language-structured training improves test success from 49-52% to 55-58% and increases adversary-focused attention from 1.2x to 2.1x, demonstrating emergent threat awareness. However, this representational gain does not consistently translate into adaptive control: trained policies perform similarly to the best constant action, while a portfolio of simple policies solves 76% of scenarios. We formalise this discrepancy as a recognition-control gap and show that reward reweighting and margin shaping do not eliminate the resulting policy collapse. Evaluations of realism, criticality, semantic accuracy, and transfer of state-interface representations to CARLA further highlight both the strengths and the constraints of language-structured relational policy learning in safety-critical driving scenarios.
Chinese Translation
基于自然语言的场景生成提供了一种直观的方式来描述稀有和复杂的驾驶交互,但尚不确定使用语言结构数据进行训练是否能够产生真正自适应的控制策略。我们提出了基于语言结构的关系Q学习,通过自我中心关系Q网络(Ego-Centric Relational Q-Network, ERQ-Net)实现,该网络共同学习动态交通图中的车辆间相关性和行动价值。语言描述在训练过程中定义了周围车辆的行为,而提示和语义角色则对策略隐藏。因此,ERQ-Net必须仅从可观察的运动学和交互中推断威胁相关性。在2500个安全关键场景中,基于语言结构的训练将测试成功率从49-52%提高到55-58%,并将针对对手的关注度从1.2倍提高到2.1倍,展示了新兴的威胁感知。然而,这种表征提升并未始终转化为自适应控制:训练后的策略表现与最佳常量动作相似,而一组简单策略解决了76%的场景。我们将这种差异形式化为识别-控制差距,并表明奖励重加权和边际塑造并未消除由此导致的策略崩溃。对现实性、关键性、语义准确性以及状态-接口表示向CARLA的转移的评估进一步突显了基于语言结构的关系策略学习在安全关键驾驶场景中的优势与局限。
cs.CV / 15 / 2608.11518

New Orthogonal Multiwavelet Filters Derived by Matrix Spectral Factorization

通过矩阵谱分解导出的新正交多小波滤波器
Kolev, Vasil, Cooklev, Todor, Keinert, Fritz
Abstract
The paper considers the construction of two new orthogonal multiwavelets with supercompact support by using the Fast Bauer's method for matrix spectral factorization on the matrix product filter of the orthogonal CL multiwavelet filter. The new multiwavelets possess orthogonality, symmetry/antisymmetry, and one of them provides better coding and smoothness than other supercompact multiwavelets. The performance of the new multiwavelet filters in subband-based edge detection, grayscale and color image compression and 1D and 2D signal denoising is compared with the GHM, SA4, CL, Integer Haar and Alpert multifilters. The comparative analysis shows that new multiwavelets can provides better human visual measures, SSIM and MS-SSIM in image compression and denoising applications.
Chinese Translation
本文考虑通过快速鲍尔法(Fast Bauer's method)对正交CL多小波滤波器的矩阵乘积滤波器进行矩阵谱分解,从而构造两种具有超紧支撑的新正交多小波。新多小波具有正交性、对称性/反对称性,其中一种在编码和光滑性方面优于其他超紧多小波。新多小波滤波器在基于子带的边缘检测、灰度和彩色图像压缩以及一维和二维信号去噪中的性能与GHM、SA4、CL、整数哈希(Integer Haar)和阿尔珀特(Alpert)多滤波器进行了比较。比较分析表明,新多小波在图像压缩和去噪应用中能够提供更好的视觉效果、结构相似性指数(SSIM)和多尺度结构相似性指数(MS-SSIM)。
cs.CV / 16 / 2608.11537

Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment

通过可观察的语义-图像接口和分层生成器证据对齐的生成语义分割
Cai, Weize, Dong, Yongqi, Shao, Zhida, Fu, Zixin
Abstract
Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boundary mixing, whereas latent-feature decoders that predict a separate output distribution may relegate the rendered image to an intermediate visualization. We present Semantic Prism, a conditional semantic-image generation-and-refinement framework with deterministic inference. A diffusion-distilled one-step generator renders a semantic RGB image; per-pixel distances from the rendered colors to a fixed class-color codebook define an explicit probabilistic interface. Hierarchical Generator Evidence Alignment spatially aligns multi-level generator features and uses a zero-initialized output projection to predict an additive residual in the interface logit space, retaining the image-defined interface as the reference for the final distribution. The interface and refined distributions further enable Contextual Interface--Hierarchy Disagreement (C-IHD), a fixed readout for ranking remaining pixel errors without an auxiliary predictor or additional forward pass. On the 500-image Cityscapes validation set, Semantic Prism achieves 72.07% mean intersection over union, 11.39 mIoU points above direct-interface decoding, with 0.41% expected calibration error. Matched-capacity ablations over three seeds support the benefit of jointly aligned multi-level evidence. A separately trained model attains 62.22% mIoU on BDD100K, while the Cityscapes-trained model reaches 46.89\% mIoU under source-frozen transfer to the Adverse Conditions Dataset with Correspondences, without target-domain adaptation. Across all three datasets, C-IHD consistently improves the area under the precision--recall curve for pixel-error ranking over maximum softmax probability on the same segmentation predictions; on ACDC, it raises AUPR from 0.6580 to 0.7557.
Chinese Translation
生成语义分割将结构化预测呈现为图像,但直接的颜色解码容易受到颜色漂移和边界混合的影响,而预测单独输出分布的潜在特征解码器可能将渲染图像降级为中间可视化。我们提出了语义棱镜(Semantic Prism),这是一个具有确定性推理的条件语义-图像生成与精炼框架。一个经过扩散蒸馏的一步生成器渲染出语义RGB图像;从渲染颜色到固定类别颜色代码本的每个像素距离定义了一个明确的概率接口。分层生成器证据对齐(Hierarchical Generator Evidence Alignment)在空间上对齐多层生成器特征,并使用零初始化的输出投影来预测接口对数空间中的附加残差,保留图像定义的接口作为最终分布的参考。该接口和精炼分布进一步使得上下文接口-层次不一致(Contextual Interface--Hierarchy Disagreement, C-IHD)成为可能,这是一种固定的读取方式,用于在没有辅助预测器或额外前向传递的情况下对剩余像素错误进行排序。在500幅图像的Cityscapes验证集上,语义棱镜实现了72.07%的平均交并比(mIoU),比直接接口解码高出11.39个mIoU点,预期校准误差为0.41%。在三个种子的匹配容量消融实验中,支持了联合对齐多层证据的好处。一个单独训练的模型在BDD100K上达到了62.22%的mIoU,而在源冻结转移到具有对应关系的逆境条件数据集(Adverse Conditions Dataset)时,Cityscapes训练的模型达到了46.89%的mIoU,且没有目标领域适应。在所有三个数据集上,C-IHD始终提高了像素错误排名下的精确度-召回曲线下的面积(AUPR),相较于同一分割预测的最大softmax概率;在ACDC上,它将AUPR从0.6580提高到0.7557。
cs.CV / 17 / 2608.11546

Through Van Gogh's Eyes: Global Style Transfer with Diffusion Mod

透过梵高的眼睛:基于扩散模型的全球风格迁移
Lee, Jeongha, Kim, Yujin, Ali, Ghazanfar, Kim, Suhyun, Hwang, Jae-In
Abstract
Artistic image synthesis aims to recreate the expressive visual identity of a target artist, yet existing methods often fail to capture an artist's global style. Conventional style transfer methods transfer the style of one or a few reference artworks to a content image in a One-to-One manner, making them effective for artwork-level stylization but limited in representing the broader stylistic distribution of an artist. Text-to-image diffusion models conditioned on artist names, such as '~ in Van Gogh style', offer greater flexibility, but they often suffer from text-induced bias and reproduce patterns from only a few iconic works. To address these limitations, we introduce Global Style Transfer (GST), an artistic image synthesis paradigm, in a Many-to-One manner, that aggregates multiple artworks from a target artist and transfers their shared global style to a single content image. For GST, we propose Global Style Guidance (GSG), which learns a residual global style offset in the intermediate feature space, or h-space, of a diffusion model under a fixed prompt. By learning artist-level style semantics purely from visual statistics, GSG mitigates text-dependent artistic bias. We further propose Content Alignment Guidance (CAG), a training-free perceptual guidance mechanism that preserves the semantic structure of the content image while allowing artist-specific geometric deformation. Experiments on WikiArt demonstrate that GST achieves superior stylistic fidelity, content preservation, and output diversity compared to existing style transfer and diffusion-based artistic synthesis methods.
Chinese Translation
艺术图像合成旨在重现目标艺术家的表现性视觉特征,但现有方法往往无法捕捉艺术家的整体风格。传统的风格迁移方法以一对一的方式将一件或几件参考艺术作品的风格转移到内容图像上,使其在艺术作品级别的风格化中有效,但在表现艺术家的更广泛风格分布方面受到限制。基于艺术家名称(如“~以梵高风格”)的文本到图像扩散模型提供了更大的灵活性,但它们通常受到文本引发的偏见影响,仅重现少数标志性作品的模式。为了解决这些局限性,我们提出了全球风格迁移(Global Style Transfer, GST),一种艺术图像合成范式,以多对一的方式聚合目标艺术家的多件艺术作品,并将其共享的整体风格转移到单一内容图像上。对于GST,我们提出了全球风格引导(Global Style Guidance, GSG),它在固定提示下学习扩散模型中间特征空间(或h空间)中的残差全球风格偏移。通过仅从视觉统计中学习艺术家级别的风格语义,GSG减轻了依赖文本的艺术偏见。我们进一步提出了内容对齐引导(Content Alignment Guidance, CAG),这是一种无训练的感知引导机制,能够在允许艺术家特定几何变形的同时保持内容图像的语义结构。对WikiArt的实验表明,GST在风格保真度、内容保留和输出多样性方面优于现有的风格迁移和基于扩散的艺术合成方法。
cs.CV / 18 / 2608.11562

From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection

从合成到去除:基于物理的反射模拟与扩散视频去反射
Wang, Zepeng, Hu, Jiagao, Li, Fuhao, Chen, Yuxuan, Wang, Fei, Zhou, Daiguo
Abstract
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively studied, video reflection removal remains largely underexplored due to the lack of paired video data, temporally coherent removal models, and dedicated evaluation benchmarks. We present a closed-loop framework that unifies physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation. Our S2R-Synthesis pipeline generates paired reflected and reflection-free videos by performing physics-grounded augmentation in the structure space and rendering realistic reflected videos with a trained video diffusion renderer; the augmentation models key glass-related effects including roughness-induced blur, thickness-induced ghosting, and reflectance variation. Based on the synthesized data, we introduce S2R-Removal, the first diffusion-based video reflection removal model, which adapts a pretrained video diffusion prior through reflection-aware latent adaptation and one-step pixel-geometric refinement, recovering the clean transmission in a single denoising step. We further build S2R-Bench, the first benchmark for video reflection removal, supporting both full-reference evaluation and real-world human perceptual assessment. Experiments on S2R-Bench and multiple public image benchmarks demonstrate state-of-the-art performance and faster inference than even non-diffusion baselines, and validate the effectiveness of S2R-Synthesis. Project page: https://codingwzp.github.io/VideoDereflection_S2R.
Chinese Translation
通过玻璃捕获的视频常常包含反射,这会降低视觉质量并干扰后续视觉任务。尽管单幅图像的反射去除已经得到了广泛研究,但由于缺乏配对视频数据、时间一致的去除模型和专门的评估基准,视频反射去除仍然在很大程度上未被探索。我们提出了一个闭环框架,统一了基于物理的反射模拟、基于扩散的视频去反射和基准评估。我们的S2R-Synthesis管道通过在结构空间中执行基于物理的增强,生成配对的反射和无反射视频,并利用训练好的视频扩散渲染器渲染出逼真的反射视频;该增强模型涵盖了与玻璃相关的关键效应,包括粗糙度引起的模糊、厚度引起的鬼影和反射率变化。基于合成数据,我们引入了S2R-Removal,这是第一个基于扩散的视频反射去除模型,它通过反射感知的潜在适应和一步像素几何细化,调整预训练的视频扩散先验,在单一去噪步骤中恢复干净的透射。我们进一步构建了S2R-Bench,这是第一个视频反射去除基准,支持全参考评估和真实世界的人类感知评估。在S2R-Bench和多个公共图像基准上的实验表明,性能达到了最先进水平,推理速度甚至快于非扩散基线,并验证了S2R-Synthesis的有效性。项目页面:https://codingwzp.github.io/VideoDereflection_S2R。
cs.CV / 19 / 2608.11564

Repurposing RGB-based Foundation Model for Depth Estimation on Thermal Images Using Hierarchical Supervision

基于RGB的基础模型在热图像深度估计中的再利用:层次监督方法
Hong, Jie, Li, Tingtian, Li, Xuesong, Li, Xiao
Abstract
Depth estimation from thermal images is highly valuable for robotic applications in adverse conditions, such as nighttime and rainy weather. Recent studies have sought to transfer knowledge from RGB-based foundation models to thermal modalities, yet the rich hierarchical representations these models encode remain underutilized. To address this limitation, we propose RGB-HS, a novel framework for thermal-image depth estimation that leverages hierarchical supervision from an RGB-based foundation model. Specifically, we first replace the baseline thermal encoder with a foundational model and introduce a parallel RGB branch that also employs a foundational model as an encoder of the same architecture, taking RGB images as input. The alignment is then performed across multiple levels between the tokens of the two encoders, allowing the thermal student branch to capture both structural precision and semantic abstraction from the RGB teacher branch. Furthermore, we introduce verification to refine the alignment process by weighting tokens from the RGB branch based on RGB image quality. Extensive experiments on the popular benchmark demonstrate that RGB-HS achieves competitive performance and more effectively exploits the representational capacity of RGB-based foundation models for depth estimation on thermal images.
Chinese Translation
从热图像中进行深度估计对于在不利条件下(如夜间和雨天)的机器人应用具有重要价值。近期研究试图将基于RGB的基础模型的知识转移到热成像模式中,但这些模型所编码的丰富层次表示仍未得到充分利用。为了解决这一局限性,我们提出了RGB-HS,一个用于热图像深度估计的新框架,利用来自基于RGB的基础模型的层次监督。具体而言,我们首先用基础模型替换基线热编码器,并引入一个并行的RGB分支,该分支同样采用基础模型作为相同架构的编码器,以RGB图像作为输入。然后,在两个编码器的多个层次之间进行对齐,使得热学生分支能够从RGB教师分支捕捉到结构精度和语义抽象。此外,我们引入验证机制,通过根据RGB图像质量对RGB分支的tokens进行加权,从而优化对齐过程。在流行基准上的大量实验表明,RGB-HS实现了具有竞争力的性能,并更有效地利用了基于RGB的基础模型在热图像深度估计中的表示能力。
cs.CV / 20 / 2608.11574

Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands

手部可见性检测器:逐关键点手部可见性估计
Hara, Ryosei, Hatano, Masashi, Yanagi, Rintaro, Hashimoto, Atsushi, Yagi, Takuma, Isogawa, Mariko
Abstract
Hand Pose Estimation (HPE) is a fundamental technology for various applications such as AR/VR and robotics. In these applications, the visibility of each hand joint in the image is crucial for assessing the reliability of estimation results under occlusion. However, most existing HPE methods output joint positions without explicitly indicating their visibility. Although some methods account for occlusion or visibility, visibility estimation has mainly been used as an auxiliary signal for improving pose estimation. To our knowledge, per-joint hand visibility estimation has not been systematically studied as a standalone task. In this work, we propose Hand Visibility Detector, a model for estimating the visibility of individual hand joints, and present the first systematic investigation of visibility estimation as an independent task. We show that leveraging the prior knowledge of HPE models pretrained on large-scale data as a backbone yields high performance in this task. We further demonstrate the utility of Hand Visibility Detector on a downstream task of 3D hand pose annotation via multi-view triangulation of 2D keypoints, showing that visibility-weighted triangulation reduces reprojection error. Our method is released as a ready-to-use package, and the code and demo are available at https://github.com/ryhara/hand_visibility_detector .
Chinese Translation
手部姿态估计(HPE)是增强现实/虚拟现实(AR/VR)和机器人等多种应用的基础技术。在这些应用中,图像中每个手关节的可见性对于在遮挡情况下评估估计结果的可靠性至关重要。然而,大多数现有的HPE方法输出关节位置时并未明确指示其可见性。尽管一些方法考虑了遮挡或可见性,但可见性估计主要被用作改善姿态估计的辅助信号。据我们所知,逐关节手部可见性估计尚未作为独立任务进行系统研究。在本研究中,我们提出了手部可见性检测器(Hand Visibility Detector),这是一个用于估计单个手关节可见性的模型,并首次系统地探讨了可见性估计作为独立任务的研究。我们展示了利用在大规模数据上预训练的HPE模型作为基础网络的先验知识可以在此任务中获得高性能。我们进一步证明了手部可见性检测器在通过二维关键点的多视角三角测量进行三维手部姿态标注的下游任务中的实用性,显示出可见性加权三角测量可以减少重投影误差。我们的方法已作为可直接使用的包发布,代码和演示可在 https://github.com/ryhara/hand_visibility_detector 获取。
cs.CV / 21 / 2608.11582

A Hybrid Framework of Vision Transformer and Gated Recurrent Unit for Detection of Mosquito Diseases

视觉变换器与门控递归单元的混合框架用于蚊子疾病检测
Sharifrazi, Danial, Behzadi, Saadat, Javed, Nouman, Alizadehsani, Roohallah, Paradkar, Prasad N., Bhatti, Asim
Abstract
Identifying dengue virus-infected mosquitoes from control mosquitoes is a major challenge in analyzing mosquito locomotion behavior due to the small size and complexity of the video background. Conventional AI methods are often unable to extract accurate features from video frames and produce erroneous features. In this study, a three-step framework is introduced: first, mosquitoes are identified and the background is removed using the YOLO 11M model, then visual features are extracted using the Vision Transformer (ViT), and finally the videos are classified with a convolutional GRU (ConvGRU) classifier. A comparative analysis of different models, including Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), and their convolutional versions showed that the ConvGRU model achieved the best performance; it achieved 88.88% accuracy, 84.45% precision, 82.82% recall, and 82.81% F1 score. These results demonstrate that combining convolutional models with sequence-based networks, especially in the ConvGRU model, allows the simultaneous extraction of precise spatial features and long-term temporal dependencies from mosquito movements. Finally, the proposed framework provides a reliable solution for analyzing mosquito behavior in complex environments.
Chinese Translation
由于蚊子体积小和视频背景复杂,识别感染登革热病毒的蚊子与对照蚊子之间的差异是分析蚊子运动行为的一项重大挑战。传统的人工智能方法往往无法从视频帧中提取准确的特征,导致特征提取错误。在本研究中,提出了一个三步框架:首先,使用YOLO 11M模型识别蚊子并去除背景;然后,使用视觉变换器(Vision Transformer, ViT)提取视觉特征;最后,使用卷积门控递归单元(Convolutional GRU, ConvGRU)分类视频。对不同模型的比较分析,包括递归神经网络(Recurrent Neural Network, RNN)、长短期记忆(Long Short-Term Memory, LSTM)、门控递归单元(Gated Recurrent Unit, GRU)及其卷积版本,显示ConvGRU模型表现最佳;其准确率达到88.88%,精确率为84.45%,召回率为82.82%,F1分数为82.81%。这些结果表明,将卷积模型与基于序列的网络结合,尤其是在ConvGRU模型中,可以同时提取蚊子运动的精确空间特征和长期时间依赖性。最后,所提出的框架为在复杂环境中分析蚊子行为提供了可靠的解决方案。
cs.CV / 22 / 2608.11595

ProtoHGF-Net: Prototype HyperGraph Fusion with Intra-modal Calibration for RGBT Object Detection

ProtoHGF-Net:具有模态内校准的原型超图融合用于RGB-T物体检测
Chen, Xiangqi, Zhang, Xiuling, Yang, Chengzhuan, Zhao, Li, Zhang, Dawei, Wang, Yanchao, Chen, Liyuan, Wang, Hua, Peng, Hao, Zheng, Zhonglong
Abstract
RGB-Thermal (RGBT) object detection enables robust perception in complex scenes by leveraging the complementary strengths of visible textures and thermal cues. However, existing methods mainly rely on dense cross-modal interactions over full-resolution features, which inevitably introduce background interference and hinder the learning of target-relevant representations. In this paper, we propose the Prototype HyperGraph Fusion Network (ProtoHGF-Net), a novel framework that redefines cross-modal fusion as prototype-level semantic interaction rather than the dense cross-modal interaction paradigm. Specifically, we design Prototype HyperGraph Fusion to perform cross-modal interaction in a compact prototype-level semantic space. This design enables more selective fusion among target-relevant prototypes. To support this prototype-level fusion, we propose Teacher-Mask Calibration Distillation, which calibrates modality features before fusion using modality-specific teachers and target-aware masks. This strategy suppresses backgrou- nd-dominant responses and produces more target-focused features. Extensive experiments on DroneVehicle, DVTOD, and FLIR demonstrate that ProtoHGF-Net achieves state-of-the-art performance with 85.9\% $mAP_{50}$, 88.2\% $mAP_{50}$, and 79.1\% $mAP_{50}$, respectively. Our code is available at \href{https://github.com/ZiMo-Chen/ProtoHGF}{GitHub}.
Chinese Translation
RGB-热成像(RGBT)物体检测通过利用可见纹理和热线索的互补优势,在复杂场景中实现了稳健的感知。然而,现有的方法主要依赖于全分辨率特征上的密集跨模态交互,这不可避免地引入了背景干扰,并阻碍了目标相关表示的学习。本文提出了原型超图融合网络(ProtoHGF-Net),这是一个新颖的框架,它将跨模态融合重新定义为原型级语义交互,而不是密集的跨模态交互范式。具体而言,我们设计了原型超图融合,以在紧凑的原型级语义空间中执行跨模态交互。该设计使得在目标相关原型之间的融合更加具有选择性。为了支持这种原型级融合,我们提出了教师-掩码校准蒸馏,它在融合前使用特定模态的教师和目标感知掩码对模态特征进行校准。这一策略抑制了背景主导的响应,并产生了更具目标聚焦的特征。在DroneVehicle、DVTOD和FLIR上的大量实验表明,ProtoHGF-Net分别达到了85.9\% $mAP_{50}$、88.2\% $mAP_{50}$和79.1\% $mAP_{50}$的最新性能。我们的代码可在\href{https://github.com/ZiMo-Chen/ProtoHGF}{GitHub}上获取。
cs.CV / 23 / 2608.11601

How Can Driving World Models Do Counterfactual Prediction?

驾驶世界模型如何进行反事实预测?
Zhang, Jiaru, Cui, Can, Xu, Yi, Ye, Xin, Zhang, Ruqi, Wang, Ziran
Abstract
Driving world models are often interpreted as counterfactual simulators for observed driving episodes: given a factual driving log, they are asked what would have happened under an alternative ego action. In this paper, we identify a fundamental mismatch between this goal and direct action-conditioned prediction. The direct prediction uses the shared history and the alternative action but not the factual continuation observed after that history. It can therefore generate a plausible future without preserving what actually happened in this episode. We formalize this gap using the causal recipe of abduction, action, and prediction and study it in a setting with a short time horizon, where the alternative ego action does not alter how surrounding agents evolve. To make the gap measurable, we construct a controlled simulation benchmark with factual outcomes and matched counterfactual outcomes. Across two representative world models, direct predictions fail to match the counterfactual ground truth, supporting our analysis. As a constructive check of this analysis, we introduce a deliberately simple, training-free pipeline that moves observed evidence into the counterfactual view and lets the frozen model complete what remains unknown. Even this simple construction raises the overall recovered fraction substantially and reduces perceptual distance to the matched counterfactual on both models. We hope this work draws attention to this gap and motivates better counterfactual prediction methods for driving world models.
Chinese Translation
驾驶世界模型通常被解释为观察到的驾驶情境的反事实模拟器:给定一个事实驾驶日志,它们被要求预测在替代自我行为下会发生什么。在本文中,我们识别出这一目标与直接基于行为的预测之间的根本不匹配。直接预测使用共享的历史和替代行为,但不使用在该历史之后观察到的事实延续。因此,它可以生成一个合理的未来,而不保留在这一情境中实际发生的事情。我们使用溯因、行动和预测的因果框架形式化这一差距,并在一个短时间范围内的设置中研究它,在该设置中,替代自我行为不会改变周围代理的演变。为了使这一差距可测量,我们构建了一个具有事实结果和匹配反事实结果的受控仿真基准。在两个代表性的世界模型中,直接预测未能与反事实真实情况匹配,从而支持我们的分析。作为对这一分析的建设性检验,我们引入了一个故意简单的、无训练的流程,将观察到的证据转移到反事实视角,并让冻结的模型完成剩余的未知部分。即便是这个简单的构造也显著提高了整体恢复比例,并减少了两个模型与匹配反事实之间的感知距离。我们希望这项工作能引起对这一差距的关注,并激励更好的反事实预测方法用于驾驶世界模型。
cs.CV / 24 / 2608.11607

Topology-Aware Query Selection for Surgical Instrument Instance Segmentation

基于拓扑的外科器械实例分割查询选择
Zhang, Ze, Zhang, Yang
Abstract
Accurate foreground masks can still form an incorrect surgical-instrument instance set: duplicate, fragmented, merged, missed, or empty-frame predictions may preserve favorable pixel overlap while violating object identity and count. Final query selection is therefore a relational, variable-cardinality problem rather than a collection of independent candidate decisions. We evaluate topology-aware query selection, which represents the nonempty candidates of a fixed Mask2Former as a complete graph, learns relational candidate and pair representations, predicts set cardinality, and solves an exact structured subset problem. The formal comparison is the complete relational path versus a node-feature-matched path; it evaluates the combined effect of pairwise geometry, message passing, and the additional relational-path capacity, not an isolated component. On the sealed 22-case source test, all three discovery seeds supported instance-set performance improvement with segmentation fidelity and predefined technical-safety preservation: instance F1 increased by 0.0504--0.0612 and positive-frame set-failure rate decreased by 0.0848--0.1060. Direct ROBUST-MIPS transfer reproduced the complete result in all three seeds. Endoscapes supported only one of three seeds and therefore did not establish stable direct transfer. Taken together, the results support a bounded conclusion: the evaluated complete path improved coherent instance-set construction from fixed Mask2Former candidates in specified native-instance contracts, while stable cross-domain transfer and component-specific effects remain unestablished.
Chinese Translation
准确的前景掩膜仍可能形成不正确的外科器械实例集:重复、碎片化、合并、遗漏或空帧预测可能保留有利的像素重叠,同时违反物体的身份和计数。因此,最终的查询选择是一个关系性、可变基数的问题,而不是一组独立的候选决策。我们评估了基于拓扑的查询选择,它将固定的 Mask2Former 的非空候选表示为一个完全图,学习关系候选和对的表示,预测集合基数,并解决一个精确的结构化子集问题。正式比较是完整的关系路径与节点特征匹配路径;它评估了成对几何、消息传递和额外关系路径容量的综合效应,而不是孤立的组件。在封闭的22个案例源测试中,所有三个发现种子都支持实例集性能的提升,同时保持分割的准确性和预定义的技术安全性:实例 F1 增加了 0.0504--0.0612,正帧集失败率降低了 0.0848--0.1060。直接的 ROBUST-MIPS 转移在所有三个种子中重现了完整结果。内窥镜支持仅一个种子,因此未能建立稳定的直接转移。综合来看,结果支持一个有限的结论:评估的完整路径在特定的原生实例合同中改善了从固定 Mask2Former 候选构建一致的实例集,而稳定的跨域转移和组件特定效应仍未建立。
cs.CV / 25 / 2608.11617

KANResDiff: Learning Local Residual Diffusion via Kolmogorov-Arnold Network for Ambiguous Medical Image Segmentation

KANResDiff:通过Kolmogorov-Arnold网络学习局部残差扩散以实现模糊医学图像分割
Li, Fanding, Wang, Chenglin, Li, Xiangyu, Qiu, Xingyu, Ma, Xinghua, Yin, Xiangming, Li, Haiyang, Dong, Suyu, Wang, Wei, Wang, Kuanquan, Luo, Gongning, Li, Shuo
Abstract
Ambiguous medical image segmentation aims to provide a series of diverse but plausible segmentation hypotheses. However, existing methods introduce stochasticity in a fixed and pre-defined manner, failing to form a progressive semantic modeling process. To address these challenges, we propose KANResDiff to learn local residual diffusion with Kolmogorov-Arnold Network, thereby assigning distinct roles across stages for ambiguity modeling. Specifically, we propose Independent Time Encoding that offers spline-based time embeddings instead of linear ones from MLPs, which enhances the independence across inference stages and assigns progressive semantic roles to different stages. We propose Residual Schrodinger Bridge that injects deterministic residual prior with learnable weights by constructing local Schrodinger Bridge instead of following manually settings, achieving a flexible deterministic-stochastic interaction and stage-aware ambiguity modeling thanks to local optimal diffusion path. Extensive experimental results on two public datasets demonstrate that KANResDiff achieves SOTA performance on GED and HM-IoU, with maximum improvements of 16.8% and 7.7%, respectively, while maintaining competitive performance on the MDM metric. Source code is available at https://github.com/PerceptionComputingLab/KANResDiff.
Chinese Translation
模糊医学图像分割旨在提供一系列多样但合理的分割假设。然而,现有方法以固定和预定义的方式引入随机性,未能形成渐进的语义建模过程。为了解决这些挑战,我们提出KANResDiff,通过Kolmogorov-Arnold网络学习局部残差扩散,从而在各个阶段为模糊建模分配不同的角色。具体而言,我们提出了独立时间编码,提供基于样条的时间嵌入,而不是来自多层感知器(MLPs)的线性嵌入,这增强了推理阶段之间的独立性,并为不同阶段分配渐进的语义角色。我们提出了残差薛定谔桥,通过构建局部薛定谔桥,注入具有可学习权重的确定性残差先验,而不是遵循手动设置,从而实现灵活的确定性-随机性交互和阶段感知的模糊建模,得益于局部最优扩散路径。在两个公共数据集上的大量实验结果表明,KANResDiff在GED和HM-IoU上达到了SOTA性能,最大提升分别为16.8%和7.7%,同时在MDM指标上保持了竞争力的表现。源代码可在https://github.com/PerceptionComputingLab/KANResDiff获取。
cs.CV / 26 / 2608.11618

Generative Video Compression Based on Hierarchical Referencing

基于层次引用的生成视频压缩
Li, Daowen, Ding, Ding, Zhang, Zifu, Li, Kai, Chen, Ying
Abstract
Diffusion-based generative video compression has emerged as a promising paradigm to improve perceptual quality, where latent frames are required to be encoded efficiently while serving as denoising conditions. However, existing methods neither carefully design reference and quality structures during latent coding nor account for the impact of frame-level quality variation on denoising procedure, which limits coding efficiency and aggravates artifact propagation during generative reconstruction. In this paper, we propose GVCHR, Generative Video Compression based on Hierarchical Referencing. The key idea is to organize latent frames hierarchically, where the selected high-quality references benefit both latent coding and generative reconstruction. In latent coding, GVCHR couples a hierarchical reference structure with a hierarchical quality structure, assigning more bits to lower-layer frames that are reused more frequently as references. Built on this design, we introduce Hierarchical Temporal Context Mining to exploits complementary short- and long-term temporal context for effective latent coding. In generative reconstruction, the coding-side hierarchy is incorporated into a Hierarchical Attentive Adapter which is attached to a video diffusion transformer. This adapter uses hierarchical attention to restrict each latent frame to attend only to the same- or lower-layer references, thereby reducing artifact propagation during denoising. Experiments validate GVCHR on multiple benchmarks. Compared with the previous state-of-the-art method, GVCHR achieves 50.5% and 54.0% BD-rate gains in terms of LPIPS and DISTS, respectively, while also delivering clearly improved visual quality.
Chinese Translation
基于扩散的生成视频压缩作为一种有前景的范式,旨在提高感知质量,其中潜在帧需要高效编码,同时作为去噪条件。然而,现有方法在潜在编码过程中既没有仔细设计参考和质量结构,也没有考虑帧级质量变化对去噪过程的影响,这限制了编码效率并加剧了生成重建过程中的伪影传播。本文提出了GVCHR,即基于层次引用的生成视频压缩。其关键思想是将潜在帧进行层次化组织,所选的高质量参考有利于潜在编码和生成重建。在潜在编码中,GVCHR将层次参考结构与层次质量结构相结合,为更频繁被重用作为参考的低层帧分配更多比特。在此设计基础上,我们引入了层次时间上下文挖掘,以利用互补的短期和长期时间上下文进行有效的潜在编码。在生成重建中,编码侧的层次结构被纳入一个层次注意适配器,该适配器附加在视频扩散变换器上。该适配器使用层次注意力限制每个潜在帧仅关注同层或下层的参考,从而减少去噪过程中的伪影传播。实验在多个基准上验证了GVCHR的有效性。与之前的最先进方法相比,GVCHR在LPIPS和DISTS方面分别实现了50.5%和54.0%的BD-rate增益,同时显著提升了视觉质量。
cs.CV / 27 / 2608.11634

CAM-Guided Saliency Cutout and Image-Based Malware Classification

基于CAM的显著性切割与图像恶意软件分类
Ebrahimi, Yasaman, Jurecek, Martin, Stamp, Mark
Abstract
Dropout regularization is commonly used to reduce overfitting by removing parts of a neural network during training. For Convolutional Neural Networks (CNN), cutouts serve a somewhat analogous purpose. Cutouts can be implemented as data augmentation: the original training image is retained, and additional copies are created with regions removed. In this chapter, we test whether cutout placement can be improved by using High-Resolution Class Activation Mapping (HiResCAM). We compare four controlled training conditions: no cutout, standard random cutout, low-saliency cutout, and high-saliency cutout. We experiment using grayscale malware images from the RawMal-TF dataset (17 families with~1,000 samples per family), and for comparison to natural images, we experiment with the well-known CIFAR-100 dataset. All experiments are based on ResNet18 with~100 training epochs. For the cutout experiments, we test cutout areas of~5\%, 10\%, 20\%, and~30\%, and we consider~$M\in\{4,8}$ augmented copies per original training image. The RawMal-TF results are slightly worse for all three cutout cases (random, high and low saliency) as compared to no cutouts. In contrast, our CIFAR-100 experimental results improve slightly under low-saliency cutout. These results suggest that the value of saliency-guided cutout is domain dependent, and that malware images should not be treated as equivalent to natural images.
Chinese Translation
Dropout正则化通常用于通过在训练过程中移除神经网络的部分来减少过拟合。对于卷积神经网络(CNN),切割具有类似的目的。切割可以作为数据增强的实现方式:保留原始训练图像,并创建去除某些区域的额外副本。在本章中,我们测试了是否可以通过使用高分辨率类别激活映射(HiResCAM)来改善切割的放置。我们比较了四种受控训练条件:无切割、标准随机切割、低显著性切割和高显著性切割。我们使用来自RawMal-TF数据集的灰度恶意软件图像进行实验(17个家族,每个家族约1,000个样本),并与著名的CIFAR-100数据集进行比较。所有实验基于ResNet18,训练周期约为100个。在切割实验中,我们测试了约5%、10%、20%和30%的切割区域,并考虑每个原始训练图像的$M ext{∈} ext{4,8}$个增强副本。与无切割相比,RawMal-TF的结果在所有三种切割情况下(随机、高显著性和低显著性)略有下降。相反,我们的CIFAR-100实验结果在低显著性切割下略有改善。这些结果表明,显著性引导切割的价值依赖于领域,恶意软件图像不应被视为与自然图像等同。
cs.CV / 28 / 2608.11643

Robustness of AI-Art Detectors under Generator Shift

生成器转变下的AI艺术检测器的鲁棒性
Thakur, Shivank Singh, Li, Meien, Stamp, Mark
Abstract
Text-to-image generative models have advanced rapidly, with modern Diffusion Transformer architectures producing images that are increasingly difficult to distinguish from human-created artwork. This development has raised significant concerns regarding copyright protection, misinformation, fraud, impersonation, and the authenticity of digital content. Most AI-art detectors are trained and evaluated on the same generator family, leaving robustness to newer architectures underexplored. In this chapter, we analyze generator shift based on a Stable Diffusion 3.5 Medium (SD3.5m) artwork dataset spanning ten art styles through reverse prompting of held-out human artwork samples. Five detectors are trained on U-Net-based latent diffusion artwork and evaluated in a zero-shot cross-generator setting on the SD3.5m dataset. Deep learning models perform strongly in-distribution but degrade under generator shift, misclassifying many SD3.5m images as human while human false positives remain low. The CLIP ViT-L/14 model performs best overall, while Grad-CAM analysis reveals weaker and more diffuse activation on false negatives. These findings highlight a generalization gap in current AI-art detectors and motivate the development of detectors as one component of a layered defense that remains reliable across rapidly evolving generative architectures.
Chinese Translation
文本到图像的生成模型发展迅速,现代扩散变换器架构生成的图像越来越难以与人类创作的艺术作品区分。这一发展引发了关于版权保护、错误信息、欺诈、冒充以及数字内容真实性的重大担忧。大多数AI艺术检测器是在同一生成器家族上进行训练和评估的,因此对于新架构的鲁棒性研究较少。在本章中,我们基于一个包含十种艺术风格的Stable Diffusion 3.5 Medium (SD3.5m)艺术作品数据集分析生成器转变,通过反向提示持有的人类艺术作品样本进行研究。五个检测器在基于U-Net的潜在扩散艺术作品上进行训练,并在SD3.5m数据集上以零样本跨生成器的方式进行评估。深度学习模型在分布内表现良好,但在生成器转变下性能下降,许多SD3.5m图像被错误分类为人类作品,而人类的误报则保持较低。CLIP ViT-L/14模型整体表现最佳,而Grad-CAM分析显示在误判为负类的样本上激活较弱且更为分散。这些发现突显了当前AI艺术检测器的泛化差距,并推动了检测器作为一个分层防御组件的开发,以在快速演变的生成架构中保持可靠性。
cs.CV / 29 / 2608.11645

Cloak of Invisibility: Real-Time Privacy-Preserving Volumetric Video Streaming

隐形斗篷:实时隐私保护的体积视频流媒体
Khalili, Hossein, Do, Philip, Vilesov, Alexander, Apicharttrisorn, Kittipat, Sehatbakhsh, Nader
Abstract
Volumetric video streaming turns privacy into a 3D, multi-view problem. Unlike ordinary video, where sensitive content can often be redacted frame by frame, RGB-D volumetric pipelines capture people, rooms, and personal objects from multiple cameras and fuse them into a shared 3D representation. A private object missed in one view, or only partially removed before fusion, can therefore reappear in the reconstructed scene. This creates a privacy challenge for 3D telepresence, education, entertainment, and immersive applications: private content should be removed before raw visual and geometric data leave the camera side, while the public part of the scene should remain useful for real-time reconstruction. Existing volumetric streaming systems mainly optimize reconstruction, data movement, and latency, while privacy-preserving vision methods are designed for single-camera, single-frame images and do not directly address calibrated multi-view RGB-D fusion. We present InViStream, a real-time "privacy-from-source" system designed for this setting. InViStream addresses three challenges in volumetric capture: private objects may appear differently across views, RGB masking alone can leave geometric privacy leakage in depth, and public/private instances of the same class must be separated consistently before cloud-side fusion. To address these challenges, InViStream combines object detection with depth-aware masking, propagates public/private decisions across calibrated views, and fuses only sanitized point clouds. We evaluate InViStream on synthetic and real RGB-D scenes, including offices, conference rooms, living rooms, and settings with multiple public and private people and objects. InViStream achieves synthetic Dice/Recall of 0.799/0.891 and real Dice/Recall of 0.792/0.908, with synthetic SSIM above 0.98 and real-time streaming above 30 FPS.
Chinese Translation
体积视频流媒体将隐私转化为一个三维、多视角的问题。与普通视频不同,普通视频中的敏感内容通常可以逐帧编辑,而RGB-D体积管道则通过多个摄像头捕捉人、房间和个人物品,并将其融合为共享的三维表示。在一个视角中遗漏的私密物体,或在融合前仅部分移除的物体,可能会在重建的场景中重新出现。这给三维远程呈现、教育、娱乐和沉浸式应用带来了隐私挑战:私密内容应在原始视觉和几何数据离开摄像头之前被移除,而场景的公共部分应保持对实时重建的实用性。现有的体积流媒体系统主要优化重建、数据传输和延迟,而隐私保护视觉方法则是为单摄像头、单帧图像设计的,并未直接解决经过校准的多视角RGB-D融合问题。我们提出了InViStream,一个为此设置设计的实时“源头隐私”系统。InViStream解决了体积捕捉中的三个挑战:私密物体在不同视角中可能表现不同,单靠RGB遮罩可能在深度上留下几何隐私泄露,以及同一类别的公共/私密实例在云端融合前必须一致分离。为了解决这些挑战,InViStream将物体检测与深度感知遮罩相结合,跨校准视角传播公共/私密决策,并仅融合经过清理的点云。我们在合成和真实的RGB-D场景上评估了InViStream,包括办公室、会议室、客厅以及有多个公共和私密人员及物体的环境。InViStream在合成场景中实现了0.799/0.891的Dice/Recall,在真实场景中实现了0.792/0.908的Dice/Recall,合成SSIM超过0.98,实时流媒体超过30 FPS。
cs.CV / 30 / 2608.11646

Hybrid-LUT: Channel-Aware Hybrid Lookup Table and Filtering for Efficient Image Denoising

混合查找表(Hybrid-LUT):面向通道的混合查找表与滤波在高效图像去噪中的应用
Ai, Zhilin, Li, Boyu, Yang, Sidi, Shi, Wenqing, Zhou, Wenyong, Huang, Binxiao, Ding, Chenchen, Wong, Ngai
Abstract
Lookup table (LUT)-based image denoising methods have attracted increasing attention due to their high efficiency and hardware-friendly properties. However, existing RGB-LUT approaches require three identical LUTs to process RGB channels in parallel, resulting in large on-chip SRAM consumption. A simple alternative is to apply LUT processing only to the luminance (Y) channel in the YUV color space to reduce memory usage. However, this naive strategy leads to degraded restoration quality, since ignoring the chrominance (UV) channels introduces color distortion and residual artifacts. In this work, we propose Hybrid-LUT, a YUV-based asymmetric channel-processing framework that combines LUT and filtering in a unified design. Specifically, a multi-band LUT branch with pixel-level weight fusion is applied to the Y channel to recover fine textures, while lightweight filtering is used for the UV channels to maintain color consistency. This design reduces LUT storage by two-thirds compared with RGB-LUT methods while maintaining the same runtime throughput. Extensive experiments show that Hybrid-LUT achieves state-of-the-art (SOTA) performance across multiple benchmarks with only 421 KB of storage. In particular, our method surpasses existing LUT-based denoising approaches by at least 0.63 dB CPSNR on real-world datasets, demonstrating its effectiveness for image denoising on resource-constrained edge devices. The project is available at https://github.com/Ai-ZL/Hybrid-LUT .
Chinese Translation
基于查找表(LUT)的图像去噪方法因其高效性和硬件友好特性而受到越来越多的关注。然而,现有的RGB-LUT方法需要三个相同的LUT并行处理RGB通道,这导致了较大的片上SRAM消耗。一个简单的替代方案是仅对YUV颜色空间中的亮度(Y)通道应用LUT处理,以减少内存使用。然而,这种简单策略会导致恢复质量下降,因为忽略色度(UV)通道会引入颜色失真和残余伪影。在本研究中,我们提出了混合查找表(Hybrid-LUT),这是一种基于YUV的不对称通道处理框架,将LUT与滤波统一设计。具体而言,采用多带LUT分支与像素级权重融合应用于Y通道,以恢复细腻的纹理,同时对UV通道使用轻量级滤波以保持颜色一致性。与RGB-LUT方法相比,该设计将LUT存储减少了三分之二,同时保持相同的运行时吞吐量。大量实验表明,Hybrid-LUT在多个基准测试中实现了最先进(SOTA)的性能,仅需421 KB的存储空间。特别是,我们的方法在真实世界数据集上超越了现有的基于LUT的去噪方法,至少提高了0.63 dB的CPSNR,证明了其在资源受限的边缘设备上进行图像去噪的有效性。该项目可在https://github.com/Ai-ZL/Hybrid-LUT获取。
cs.CV / 31 / 2608.11655

Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Motion-Guided Cross-Frame Visual Prompting

运动作为提示:通过运动引导的跨帧视觉提示增强多模态大型语言模型中的运动推理
Sun, Xikai, Liu, Kebin, Wang, Haotian, Liu, Li, Wang, Xu, Liu, Yunhao
Abstract
Motion-centric video reasoning is fundamental to interactive applications such as robotic manipulation and autonomous navigation. However, multimodal large language models (MLLMs) typically process videos through sparse uniform sampling to control visual-token and attention costs. This strategy may discard critical transitions between sampled frames, limiting reasoning about object movement, collisions, and causal interactions. To mitigate this issue, we propose Motion-as-Prompt (MaP), a track-guided cross-frame visual prompting framework. MaP recovers dense point trajectories, selects motion-informative frames, and marks the trajectories accumulated between consecutive sampled frames directly onto the visual inputs, making otherwise hidden displacement, direction changes, and interactions observable to frozen MLLMs. Experiments on CLEVRER and Something-Something-v2 show that MaP consistently improves average motion-reasoning accuracy, yielding gains of 4.2% and 8.9% for GPT-5.5, respectively. Notably, these improvements are obtained without degrading non-motion understanding, highlighting the robustness of MaP. These results demonstrate that MaP provides a simple and effective solution for enhancing motion-centric video reasoning without model training or architectural modification. Project page:https://github.com/SunVictor23/MaP.
Chinese Translation
以运动为中心的视频推理对于机器人操作和自主导航等交互式应用至关重要。然而,多模态大型语言模型(MLLMs)通常通过稀疏均匀采样来处理视频,以控制视觉标记和注意力成本。这种策略可能会丢弃采样帧之间的关键过渡,限制对物体运动、碰撞和因果交互的推理。为了解决这个问题,我们提出了运动作为提示(Motion-as-Prompt,MaP),一种轨迹引导的跨帧视觉提示框架。MaP 恢复密集的点轨迹,选择运动信息丰富的帧,并将连续采样帧之间累积的轨迹直接标记到视觉输入上,使得原本隐藏的位移、方向变化和交互对冻结的 MLLMs 可观察。对 CLEVRER 和 Something-Something-v2 的实验表明,MaP 一致地提高了平均运动推理的准确性,分别为 GPT-5.5 带来了 4.2% 和 8.9% 的提升。值得注意的是,这些改进是在不降低非运动理解的情况下获得的,突显了 MaP 的鲁棒性。这些结果表明,MaP 提供了一种简单有效的解决方案,以增强以运动为中心的视频推理,而无需模型训练或架构修改。项目页面:https://github.com/SunVictor23/MaP。
cs.CV / 32 / 2608.11663

Zero-OVCD: Bridging Training-Free Foundation Models and Pseudo-Label Learning for Open-Vocabulary Change Detection

Zero-OVCD:连接无训练基础模型与伪标签学习以实现开放词汇变化检测
Peng, Daifeng, Peng, Yuanke, Guan, Haiyan
Abstract
Open-vocabulary change detection (OVCD) enables the identification of user-specified land-cover changes in bitemporal remote sensing images, but existing training-free pipelines remain vulnerable to inaccurate candidate masks, ambiguous semantic assignments, and accumulated inference errors. To address these issues, we propose Zero-OVCD, a two-stage framework that requires no pixel-level annotations from the target domain. In the first stage, high-quality change pseudo-labels are generated through complementary candidate-mask refinement, multiscale semantic similarity fusion with margin-based reliability filtering, and response-guided mask correction and completion. These components jointly suppress noisy candidates, enhance mask-level semantic discrimination, and recover missed change regions. In the second stage, a change detector is trained using the generated pseudo-labels, while checkpoint voting and high-agreement sample selection are introduced to mitigate residual pseudo-label noise. On LEVIR-CD, WHU-CD, and S2Looking, Stage I achieves F1 scores of 86.25%, 85.82%, and 50.48%, while Stage II further improves them to 88.65%, 88.85%, and 57.96%, respectively. On SECOND, the macro-average F1 across six category-wise one-vs-rest tasks increases from 47.91% to 50.92%. These results demonstrate that bridging training-free foundation-model inference with noise-aware pseudo-label learning provides an effective solution for open-vocabulary change detection without target-domain pixel-level annotations. Code will be available at https://github.com/1321663019/Zero-OVCD.
Chinese Translation
开放词汇变化检测(OVCD)能够识别用户指定的双时相遥感图像中的土地覆盖变化,但现有的无训练流程仍然容易受到不准确的候选掩膜、模糊的语义分配和累积的推理错误的影响。为了解决这些问题,我们提出了Zero-OVCD,一个不需要目标领域像素级注释的两阶段框架。在第一阶段,通过互补候选掩膜精炼、基于边际的可靠性过滤的多尺度语义相似性融合,以及响应引导的掩膜修正与补全,生成高质量的变化伪标签。这些组件共同抑制噪声候选,增强掩膜级语义区分,并恢复遗漏的变化区域。在第二阶段,使用生成的伪标签训练变化检测器,同时引入检查点投票和高一致性样本选择,以减轻残余的伪标签噪声。在LEVIR-CD、WHU-CD和S2Looking数据集上,第一阶段的F1分数分别达到86.25%、85.82%和50.48%,而第二阶段进一步提高到88.65%、88.85%和57.96%。在SECOND数据集上,六个类别的宏平均F1分数从47.91%提高到50.92%。这些结果表明,将无训练基础模型推理与噪声感知伪标签学习相结合,为开放词汇变化检测提供了一种有效的解决方案,而无需目标领域的像素级注释。代码将发布在 https://github.com/1321663019/Zero-OVCD。
cs.CV / 33 / 2608.11681

Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation

基于多模态伪标签学习的鲁棒开放词汇实例与全景分割
Thanh, Duy Tran, Lee, Yeejin, Kang, Byeongkeun
Abstract
This work addresses the challenge of open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS), which aim to recognize both predefined and unseen object categories without exhaustive human annotations. Existing methods often suffer from noisy pseudo-masks, limited visual-textual grounding, and difficulty handling synonyms or out-of-vocabulary (OOV) words. To overcome these challenges, we propose a multimodal framework that leverages pre-trained vision-language models for automatic pseudo-label generation, CLIP-guided synonym filtering, and GPT-based caption reconstruction. In our target-vocabulary-assisted pseudo-labeling setting, the framework first constructs pseudo segmentation masks, descriptive captions, and semantically aligned synonym sets using Grounded SAM, LLaVA, and CLIP, providing multimodal supervision without manual annotation. We then enhance visual-textual alignment through three complementary training objectives: an extended grounding loss that incorporates visually grounded synonyms, a semantic consistency loss, and a generative caption reconstruction loss. Extensive experiments on the COCO dataset demonstrate that the proposed method consistently outperforms previous state-of-the-art approaches under this protocol, achieving substantial improvements on both OVIS and OSPS benchmarks.
Chinese Translation
本研究解决了开放词汇实例分割(OVIS)和开放集全景分割(OSPS)面临的挑战,旨在识别预定义和未见的物体类别,而无需大量人工标注。现有方法常常受到噪声伪掩膜、有限的视觉-文本对齐以及处理同义词或词汇外(OOV)单词的困难等问题的困扰。为了解决这些挑战,我们提出了一种多模态框架,利用预训练的视觉-语言模型进行自动伪标签生成、CLIP引导的同义词过滤和基于GPT的描述重构。在我们的目标词汇辅助伪标签设置中,该框架首先使用Grounded SAM、LLaVA和CLIP构建伪分割掩膜、描述性标题和语义对齐的同义词集,从而提供无需人工标注的多模态监督。然后,我们通过三个互补的训练目标增强视觉-文本对齐:一个扩展的对齐损失,结合了视觉对齐的同义词,一个语义一致性损失,以及一个生成性标题重构损失。在COCO数据集上的大量实验表明,所提方法在该协议下始终优于之前的最先进方法,在OVIS和OSPS基准上均取得了显著的改进。
cs.CV / 34 / 2608.11685

EGM-Det: Entropy-Guided Multimodal Adaptive Fusion for UAV RGB-IR Object Detection

EGM-Det:基于熵引导的多模态自适应融合用于无人机RGB-IR目标检测
Fan, Cunzheng, Yan, Dawei, Wang, Guanlin, Yang, Xingshuo, Jia, Yupeng, Yang, Jing, Zhang, Haokui
Abstract
Joint use of RGB and infrared (IR) imagery can improve UAV-view object detection, but most existing methods fuse multimodal features with static or fixed weights and therefore overlook spatially varying modality reliability. We propose EGM-Det, an entropy-guided multimodal adaptive fusion framework for RGB-IR object detection. EGM-Det employs a dual-stream architecture to preserve modality-specific representations and introduces an Entropy Offset Gate Fusion module for adaptive multi-scale fusion. The module derives shallow entropy priors from input intensity, local entropy, and cross-modal discrepancy, and uses them to guide local offset alignment and spatial-channel gated fusion. It therefore selectively aggregates reliable RGB and infrared cues instead of uniformly combining heterogeneous features. We further introduce cross-modal distillation to regularize the learned fusion gates and reduce fusion degradation. Each student branch extracts complementary knowledge from the cross-modality teacher branch matched to the main branch, while entropy-adaptive supervision emphasizes uncertain modality decisions. Experiments on DroneVehicle, LLVIP, and VEDAI demonstrate state-of-the-art performance across all three benchmarks; in particular, EGM-Det outperforms prior approaches by more than 10 percentage points on VEDAI.
Chinese Translation
RGB与红外(IR)图像的联合使用可以提高无人机视角下的目标检测能力,但大多数现有方法以静态或固定权重融合多模态特征,因此忽视了空间上变化的模态可靠性。我们提出了EGM-Det,一种基于熵引导的多模态自适应融合框架,用于RGB-IR目标检测。EGM-Det采用双流架构以保留模态特定的表示,并引入熵偏移门融合模块以实现自适应多尺度融合。该模块从输入强度、局部熵和跨模态差异中推导出浅层熵先验,并利用这些先验指导局部偏移对齐和空间通道门控融合。因此,它选择性地聚合可靠的RGB和红外线线索,而不是均匀地组合异构特征。我们进一步引入跨模态蒸馏以规范学习到的融合门并减少融合降级。每个学生分支从与主分支匹配的跨模态教师分支中提取互补知识,同时熵自适应监督强调不确定的模态决策。在DroneVehicle、LLVIP和VEDAI上的实验表明,在所有三个基准上EGM-Det均表现出最先进的性能;特别是,EGM-Det在VEDAI上比之前的方法提高了超过10个百分点。
cs.CV / 35 / 2608.11697

Boundary-Enhanced Segmentation of Pig Point Clouds in Commercial Housing Environments

商业养殖环境中猪点云的边界增强分割
Xu, Zhankang, Shi, Fei, Qi, Xiangyu, Wang, Zhaoyang, Guo, Mengxin, Fan, Yikai, Yang, Simon X., Li, Qifeng, Ma, Weihong
Abstract
In real pigsty environments, pig point clouds often come into close contact with background structures, resulting in blurred target boundaries, local adhesion, and background mis-segmentation. This reduces the accuracy of subsequent point cloud completion and body size measurement. To address these challenges, this study proposes a pig point cloud segmentation method based on boundary feature analysis. The proposed method adopts Octree Transformer as the backbone network and integrates local geometric details with global semantic context through octree convolution, self-attention encoding, and multi-scale feature fusion. Furthermore, soft-distance boundary pseudo-labels are generated to provide continuous boundary supervision, and a bidirectional cross-boundary semantic module is designed to enable explicit interaction between boundary and semantic features. Experiments conducted on a comprehensive dataset demonstrate that the proposed method significantly outperforms various state-of-the-art models in terms of segmentation accuracy, mean intersection over union, and boundary delineation. The results indicate that the method effectively alleviates boundary adhesion, providing reliable point cloud inputs for downstream precision livestock farming tasks.
Chinese Translation
在实际猪舍环境中,猪点云常常与背景结构紧密接触,导致目标边界模糊、局部粘连和背景误分割。这降低了后续点云补全和体型测量的准确性。为了解决这些挑战,本研究提出了一种基于边界特征分析的猪点云分割方法。该方法采用Octree Transformer作为主干网络,通过八叉树卷积、自注意力编码和多尺度特征融合,将局部几何细节与全局语义上下文相结合。此外,生成软距离边界伪标签以提供连续的边界监督,并设计了双向跨边界语义模块,以实现边界与语义特征之间的显式交互。在一个综合数据集上进行的实验表明,该方法在分割准确性、平均交并比和边界描绘方面显著优于多种最先进的模型。结果表明,该方法有效缓解了边界粘连,为下游精细化养殖任务提供了可靠的点云输入。
cs.CV / 36 / 2608.11699

STAR: A Spatial-Topology Aware Routing Framework for Generalizable 3D Scene Understanding

STAR:一种空间拓扑感知的通用3D场景理解路由框架
Xing, Mingwei, Wang, Xinliang, Shi, Yifeng
Abstract
Constructing a unified 3D scene understanding model has long been hindered by the topological discrepancies across sensor modalities. While applying the Mixture-of-Experts (MoE) architecture is a flexible approach for multi-domain 3D understanding, we observe that conventional feature-only MoE routers may underrepresent local sampling topology under semantic supervision, making expert allocation difficult when semantic consistency coexists with geometric heterogeneity. To overcome this challenge, we propose STAR (Spatial-Topology Aware Routing Framework). Specifically, we introduce a multi-attribute self-supervised pre-training branch, covering topological and textural variations, to anchor cross-domain structural priors. Building upon this, we design a domain-aware expert branch with two mechanisms: Domain-Spatial-Guided Routing (DSR), which captures local topological variations from spatial context, and Entropy-controlled Dynamic Allocation (EDA), which adjusts the number of activated experts according to routing uncertainty. Together, these branches combine stable cross-domain representation learning with adaptive expert allocation. Extensive experiments across various tasks, encompassing both indoor and outdoor scenes, demonstrate the effectiveness of STAR. It achieves 80.1% mIoU on the ScanNet validation set and 77.2% mIoU on S3DIS, consistently improving over strong baselines. Code is available at our project page (https://xmw666.github.io/STAR/).
Chinese Translation
构建统一的3D场景理解模型长期以来受到传感器模态之间拓扑差异的阻碍。虽然应用专家混合(Mixture-of-Experts, MoE)架构是一种灵活的多领域3D理解方法,但我们观察到传统的仅基于特征的MoE路由器在语义监督下可能无法充分代表局部采样拓扑,从而在语义一致性与几何异质性共存时使专家分配变得困难。为了解决这一挑战,我们提出了STAR(空间拓扑感知路由框架)。具体而言,我们引入了一个多属性自监督预训练分支,涵盖拓扑和纹理变化,以锚定跨领域结构先验。在此基础上,我们设计了一个领域感知专家分支,包含两个机制:领域空间引导路由(Domain-Spatial-Guided Routing, DSR),用于捕捉来自空间上下文的局部拓扑变化,以及熵控制的动态分配(Entropy-controlled Dynamic Allocation, EDA),根据路由不确定性调整激活专家的数量。这些分支共同结合了稳定的跨领域表征学习与自适应专家分配。在各种任务(包括室内和室外场景)上的广泛实验表明,STAR的有效性。它在ScanNet验证集上达到了80.1%的mIoU,在S3DIS上达到了77.2%的mIoU,始终优于强基线。代码可在我们的项目页面(https://xmw666.github.io/STAR/)获取。
cs.CV / 37 / 2608.11738

Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System

推进基于多模态大型语言模型的无人机图像理解与推理:基准测试与无训练多智能体系统
Zhang, Haoyu, Zhang, Shuoxun, Ye, Peng, Zhang, Lin, Yuan, Jiakang, Yi, Shenghong, Wang, Yuening, Chen, Tao
Abstract
Multimodal Large Language Model (MLLM)-based UAV aerial image understanding and reasoning is essential for aerial intelligence yet poses distinct challenges arising from extreme scale variation, arbitrary camera orientations, and high object density. Despite growing interest, existing evaluations remain fragmented across individual datasets and narrow tasks, leaving a critical gap in unified assessment of UAV understanding and reasoning capabilities. To fill this gap, we construct UAVQA-Bench, a benchmark of 1,500 human-annotated QA pairs drawn from 13 public UAV datasets, covering 6 capability dimensions and 16 tasks in both multiple-choice and visual grounding formats. Systematic evaluation of a broad range of open-source and closed-source MLLMs as well as agent-based systems on UAVQA-Bench identifies three key failure modes: domain-toolset mismatch, unchecked error propagation, and static reasoning. Motivated by these findings, we propose UAV-MAS, a training-free multi-agent system for MLLM-based UAV aerial image understanding and reasoning, comprising a Domain-Specific Perception Engine (DSPE) that routes queries to task-appropriate visual tools, a Context-Aware Iterative Refinement module (CAIR) that validates intermediate reasoning to curb error accumulation, and a Difficulty-Aware Adaptive Search mechanism (DAAS) that adjusts search depth to question difficulty. UAV-MAS with a 32B open-source MLLM achieves 77.0% overall accuracy on UAVQA-Bench, surpassing Gemini 3 Pro by 4.0\%, while the 8B variant improves 8.7\% over its base model.
Chinese Translation
基于多模态大型语言模型(MLLM)的无人机(UAV)航拍图像理解与推理对于空中情报至关重要,但面临着极端尺度变化、任意相机方向和高物体密度等独特挑战。尽管兴趣日益增长,现有评估仍然分散在各个独立数据集和狭窄任务上,导致对无人机理解与推理能力的统一评估存在重大空白。为填补这一空白,我们构建了UAVQA-Bench,这是一个包含1,500个来自13个公共无人机数据集的人类标注问答对的基准,涵盖6个能力维度和16个任务,包括多项选择和视觉定位格式。对一系列开源和闭源的MLLM以及基于智能体的系统在UAVQA-Bench上的系统评估识别出三种主要失败模式:领域-工具集不匹配、未检查的错误传播和静态推理。基于这些发现,我们提出了UAV-MAS,一个无训练的多智能体系统,用于基于MLLM的无人机航拍图像理解与推理,包含一个领域特定感知引擎(Domain-Specific Perception Engine, DSPE),该引擎将查询路由到适合任务的视觉工具,一个上下文感知迭代精炼模块(Context-Aware Iterative Refinement, CAIR),该模块验证中间推理以抑制错误累积,以及一个难度感知自适应搜索机制(Difficulty-Aware Adaptive Search, DAAS),该机制根据问题难度调整搜索深度。使用32B的开源MLLM,UAV-MAS在UAVQA-Bench上实现了77.0%的整体准确率,超过了Gemini 3 Pro 4.0%,而8B变体比其基础模型提高了8.7%。
cs.CV / 38 / 2608.11741

JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis

JieZi:一个大规模专家审核的数据集和古代汉字释义基准
Li, Ran, He, Huiguo, Cao, Jiahuan, Liu, Junle, Cheng, Hiuyi, Jin, Lianwen
Abstract
The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. However, existing computational approaches focus narrowly on subtasks such as character recognition and retrieval, lacking the structured datasets and benchmarks required for comprehensive scholarly analysis. To address this limitation, we introduce Ancient Chinese Character Exegesis (ACCE), a vision-language question answering (VQA) task that models the scholarly exegesis process. ACCE is organized into four progressive levels: basic character identification, glyph-form analysis, meaning exegesis, and diachronic evolution analysis. To support this task, we construct two complementary resources. JieZi-Dataset is the first large-scale, expert-audited VQA training dataset for ACCE, comprising over 500K QA pairs. It is constructed via a pipeline that reduces factual errors by constraining generation with expert-designed templates and source-text references. Human verification is further applied at each key stage to ensure scholarly accuracy. JieZi-Bench is an evaluation benchmark aligned with the exegesis process, constructed and verified by human experts to ensure evaluation reliability. It consists of four levels with reference answers curated from authoritative lexicographic works held separate from the training data. Experiments on multimodal large language models show that current models perform well on basic identification but struggle with glyph analysis, semantic reasoning, and diachronic understanding. Fine-tuning on JieZi-Dataset substantially improves performance across all four levels. Code and dataset are available at https://github.com/Ran00w/JieZi.
Chinese Translation
古代汉字的学术释义需要整合视觉观察、语言分析和历史背景。然而,现有的计算方法仅关注于字符识别和检索等子任务,缺乏进行全面学术分析所需的结构化数据集和基准。为了解决这一局限性,我们提出了古代汉字释义(ACCE),这是一项建模学术释义过程的视觉-语言问答(VQA)任务。ACCE分为四个逐步递进的层次:基本字符识别、字形分析、意义释义和历时演变分析。为支持这一任务,我们构建了两个互补资源。JieZi-Dataset是第一个大规模、专家审核的ACCE VQA训练数据集,包含超过50万个问答对。该数据集通过一个管道构建,利用专家设计的模板和源文本参考来减少事实错误。在每个关键阶段进一步应用人工验证,以确保学术准确性。JieZi-Bench是与释义过程对齐的评估基准,由人类专家构建和验证,以确保评估的可靠性。它由四个层次组成,参考答案来自于与训练数据分开的权威词典作品。对多模态大型语言模型的实验表明,当前模型在基本识别方面表现良好,但在字形分析、语义推理和历时理解方面存在困难。在JieZi-Dataset上进行微调显著提高了所有四个层次的性能。代码和数据集可在https://github.com/Ran00w/JieZi获取。
cs.CV / 39 / 2608.11745

LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time

LiveAnimate:实时稳定的长篇流式人类动画
Zhang, Yuxuan, Xiong, Haozhong, Huang, Yubo, Song, Jiayi, Yu, Jinpeng, Wang, Haofan, Liu, Jiaming, Huang, Ruihua, Wang, Liwei
Abstract
Pose-driven human animation synthesizes a video of a target person from a single reference image and a driving pose stream. Real-time generation is essential for interactive applications such as live streaming, telepresence, and virtual avatars, yet diffusion-based systems require minutes to hours per clip, precluding responsive interaction. We present LiveAnimate, to our knowledge the first animation system to combine real-time streaming with stable long-form generation at billion scale, built on a 14B-parameter video Diffusion Transformer (DiT). A two-stage training pipeline first adapts a pretrained bidirectional DiT into a block-causal autoregressive generator through Reference-Anchored Teacher-Forcing Adaptation, and then reduces the sampling budget to three steps through Block-wise Self-Forcing Distillation. To preserve appearance over extended streams, we introduce Pose-Retrieval Sink Attention (PR-Sink), a bounded KV-cache mechanism combining a Static Sink that permanently anchors the first generated block, a Dynamic Sink that holds a pose-retrieved historical block, and a three-slot Rolling Window. When a pose recurs, PR-Sink restores the relevant appearance context without retaining the entire sequence, so memory and per-block latency remain constant regardless of stream duration. Together with Ulysses sequence parallelism and operator fusion, these designs enable 19.63\,FPS streaming inference on two NVIDIA H100 GPUs. On a three-minute benchmark, LiveAnimate maintains nearly constant perceptual quality and identity from the first 30 seconds to the final minute (IQA 4.047 vs.\ 4.026), while prior systems degrade substantially or require hours of offline computation for the same rollout. These results establish a new operating point in quality, latency, and duration for interactive full-body animation.
Chinese Translation
基于姿态驱动的人类动画从单一参考图像和驱动姿态流合成目标人物的视频。实时生成对于互动应用(如直播、远程存在和虚拟化身)至关重要,然而基于扩散的系统每个片段需要几分钟到数小时,无法实现响应式交互。我们提出了LiveAnimate,据我们所知,这是第一个将实时流式传输与稳定的长篇生成结合在一起的动画系统,基于一个具有140亿参数的视频扩散变换器(Diffusion Transformer,DiT)。一个两阶段的训练流程首先通过参考锚定教师强制适应(Reference-Anchored Teacher-Forcing Adaptation)将预训练的双向DiT适配为一个块因果自回归生成器,然后通过块级自强蒸馏(Block-wise Self-Forcing Distillation)将采样预算减少到三步。为了在延长的流中保持外观,我们引入了姿态检索汇聚注意力(Pose-Retrieval Sink Attention,PR-Sink),这是一种结合了静态汇聚(Static Sink)和动态汇聚(Dynamic Sink)的有界KV缓存机制,前者永久锚定第一个生成的块,后者保持一个姿态检索的历史块,并且有一个三槽滚动窗口。当姿态重复时,PR-Sink在不保留整个序列的情况下恢复相关的外观上下文,因此无论流的持续时间如何,内存和每块延迟保持恒定。结合尤利西斯序列并行性和操作符融合,这些设计使得在两块NVIDIA H100 GPU上实现19.63帧每秒的流式推理。在一个三分钟的基准测试中,LiveAnimate从前30秒到最后一分钟保持几乎恒定的感知质量和身份(IQA 4.047对比4.026),而之前的系统在相同的播放过程中显著降级或需要数小时的离线计算。这些结果确立了交互式全身动画在质量、延迟和持续时间上的新操作点。
cs.CV / 40 / 2608.11747

Making Every Step Count: Spatio-Temporal Information Allocation for Imaging Inverse Problems

让每一步都算数:成像逆问题的时空信息分配
Cao, Yi, Cao, Xiangyong, Liu, Pei, Liu, Yong-Jin, Meng, Deyu
Abstract
Flow-based generative models have emerged as powerful image priors for training-free inverse problem solving, capturing coherent semantics and fine-grained structure. Despite these strengths, existing flow-based inverse solvers primarily focus on the design of individual updates, largely overlooking spatio-temporal information allocation under a fixed number of function evaluations (NFEs). Temporally, insufficient early exploration can trap the flow trajectory in an incorrect semantic basin, whereas excessive allocation of NFEs to early stages leaves little budget for late-stage refinement. Spatially, data consistency provides direct constraints only within observed regions, whereas the recovery of missing regions relies mainly on the generative prior. To address these two issues, we introduce two complementary and training-free components, i.e., Spectrum-Adaptive Scheduling (SAS) and Measurement-Prioritized Attention (MPA). For temporal allocation, SAS distributes the available NFEs over flow time according to the degradation spectrum and logSNR geometry, thus better balancing semantic exploration and detail refinement. For spatial propagation, MPA exploits data-prior conflicts to guide information toward weakly constrained regions, thereby enhancing semantic and structural fidelity. Extensive experiments on standard image inverse problems, e.g., super-resolution, motion deblurring, and inpainting, demonstrate that the proposed components can be integrated into existing flow-based inverse solvers in a plug-and-play manner without retraining or additional flow-model evaluations, and can also significantly improve the restoration quality of existing solvers.
Chinese Translation
基于流的生成模型已成为无训练逆问题求解的强大图像先验,能够捕捉连贯的语义和细致的结构。尽管具有这些优势,现有的基于流的逆解算器主要集中于单个更新的设计,往往忽视在固定数量的函数评估(NFEs)下的时空信息分配。在时间上,早期探索不足可能会使流轨迹陷入错误的语义盆地,而将过多的NFEs分配给早期阶段则会使后期精细化的预算不足。在空间上,数据一致性仅在观察区域内提供直接约束,而缺失区域的恢复主要依赖于生成先验。为了解决这两个问题,我们引入了两个互补的无训练组件,即谱自适应调度(Spectrum-Adaptive Scheduling, SAS)和测量优先关注(Measurement-Prioritized Attention, MPA)。在时间分配方面,SAS根据降解谱和对数信噪比几何分布可用的NFEs,从而更好地平衡语义探索和细节精细化。在空间传播方面,MPA利用数据先验冲突引导信息向弱约束区域传播,从而增强语义和结构的保真度。在标准图像逆问题(例如超分辨率、运动去模糊和修复)上的大量实验表明,所提出的组件可以以即插即用的方式集成到现有的基于流的逆解算器中,而无需重新训练或额外的流模型评估,并且可以显著提高现有解算器的恢复质量。
cs.CV / 41 / 2608.11748

Dual Modality Prompted Diffusion Priors for Zero Shot Hyperspectral Pansharpening

双模态提示扩散先验用于零样本高光谱图像融合
Xie, Pengwei, Zhu, Fei, Li, Jiajun, Liu, Xiangyuan, Liu, Xiangyuan, Shen, Kangqing, Vivone, Gemine
Abstract
Hyperspectral pansharpening aims to reconstruct a high resolution hyperspectral (HRHS) image from a panchromatic (PAN) image and a low resolution hyperspectral (LRHS) image while preserving both spatial details and spectral fidelity. Recent diffusion based methods exploit pretrained image priors by generating a low dimensional representation and subsequently mapping it to the hyperspectral domain. However, the observed panchromatic and hyperspectral images are typically imposed only through external reconstruction objectives, limiting their direct interaction with the diffusion prior. To address this issue, we propose dual-modality image-prompted diffusion model (DIDM) for zero shot hyperspectral pansharpening. DIDM encodes the low resolution hyperspectral and panchromatic observations into spectral and spatial prompt tokens, respectively, and injects them into intermediate features of a frozen remote sensing diffusion model through cross attention, allowing complementary spectral and spatial information to directly guide diffusion feature evolution. In addition, we introduce a panchromatic guided weighted pixel aware total variation regularizer that combines low resolution hyperspectral degradation fidelity and panchromatic response fidelity with gradient adaptive structural regularization, thereby preserving structural discontinuities while suppressing spurious variations in homogeneous regions. Extensive experiments on Pavia, Chikusei, and Houston under reduced resolution protocols show that DIDM achieves the best performance across all evaluated metrics, while full resolution evaluation on FR1 yields the highest HQNR among the compared methods. These results demonstrate that internal dual modality prompting and panchromatic guided structural regularization provide an effective balance between spatial detail enhancement and spectral preservation.
Chinese Translation
高光谱图像融合旨在从全色图像(PAN)和低分辨率高光谱图像(LRHS)重建高分辨率高光谱图像(HRHS),同时保留空间细节和光谱保真度。近期的基于扩散的方法通过生成低维表示并将其映射到高光谱域,利用预训练的图像先验。然而,观察到的全色图像和高光谱图像通常仅通过外部重建目标施加限制,限制了它们与扩散先验的直接交互。为了解决这一问题,我们提出了双模态图像提示扩散模型(DIDM)用于零样本高光谱图像融合。DIDM将低分辨率高光谱和全色观测分别编码为光谱和空间提示标记,并通过交叉注意机制将其注入到冻结的遥感扩散模型的中间特征中,从而使互补的光谱和空间信息能够直接引导扩散特征的演变。此外,我们引入了一种全色引导的加权像素感知总变差正则化器,该正则化器结合了低分辨率高光谱退化保真度和全色响应保真度,并采用梯度自适应结构正则化,从而在抑制均匀区域中的虚假变化的同时保留结构不连续性。在Pavia、Chikusei和Houston的降分辨率协议下进行的大量实验表明,DIDM在所有评估指标上均实现了最佳性能,而在FR1的全分辨率评估中,DIDM的HQNR在比较方法中最高。这些结果表明,内部双模态提示和全色引导的结构正则化在空间细节增强和光谱保留之间提供了有效的平衡。
cs.CV / 42 / 2608.11752

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

UniSwap:用于对话视频的流式音视频身份替换
Zhang, Yuxuan, Xiong, Haozhong, Song, Jiayi, Yu, Jinpeng, Shi, Yang, Liu, Jiaming, Huang, Ruihua, Wang, Liwei
Abstract
Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a single audio-visual diffusion transformer while preserving the source content and dynamics. To address the scarcity of aligned cross-identity training pairs, we introduce a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets. Starting from a bidirectional backbone, we progressively adapt the model through In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and Efficient Self-forcing DMD for mitigating exposure bias and reducing sampling from 30 to 3 denoising steps per block. Efficient Multi-LoRA Switching enables the three DMD roles to share a single frozen backbone. Feature-RoPE Decomposition keeps cached positions within the training range, supporting stable long-form inference. Experiments demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation.
Chinese Translation
对话视频中的角色替换需要在保持源运动、场景、语言内容和音视频时序的同时,协调外观和声音的转移。现有方法为这两种模态使用单独优化的模型,导致音视频一致性难以保证。我们提出了UniSwap,这是第一个用于对话视频中流式联合音视频身份替换的框架。给定一个源视频、一张参考图像和一个参考声音片段,UniSwap在一个单一的音视频扩散变换器中转移参考外观和声音音色,同时保留源内容和动态。为了应对对齐的跨身份训练对的稀缺,我们引入了一种交换与重构管道,该管道从真实片段中去除视觉和声音身份,并使用原始片段作为重构目标。从双向主干网络开始,我们通过上下文预训练逐步调整模型以实现联合替换,通过条件流式适应实现块因果KV缓存生成,以及通过高效自强DMD减轻曝光偏差并将每个块的采样从30步减少到3步。高效多LoRA切换使得三个DMD角色能够共享一个冻结的主干网络。特征-位置编码分解保持缓存位置在训练范围内,支持稳定的长格式推理。实验表明,UniSwap在音视频同步、身份保留、流式效率和长格式生成方面表现出色。
cs.CV / 43 / 2608.11759

Automated binary classification of hazelnut X-ray images: A deep-learning benchmark for quality assessment

榛子X射线图像的自动二分类:质量评估的深度学习基准
Sportelli, Giancarlo, Belcari, Nicola, Pace, Roberta, Bernardo, Umberto, Sultana, Sharmin, Toncelli, Alessandra, Giaccone, Matteo
Abstract
Non-destructive X-ray imaging can reveal internal hazelnut defects that are difficult to detect by external inspection alone; however, automated interpretation remains challenging because of subtle radiographic differences among classes, marked class imbalance, and limited annotated data. Here, we present a benchmark for binary hazelnut quality classification (healthy versus defective) based on 799 segmented single-kernel X-ray images (224 x 224 pixels, grayscale), grouped into 101 acquisition units. Seven single-model configurations and ten probability-aggregation ensembles were evaluated using a group-wise split-rotation protocol across five data splits generated using different random seeds. Decision thresholds were selected on the validation set, and performance was assessed deterministically on validation and test sets. Under the expert-reassessed annotation condition, the average-probability ensemble of the binary cross-entropy-trained convolutional neural network and frozen Swin Transformer achieved the highest mean balanced accuracy (86.3% +/- 1.8%, five seeds), with several other ensembles providing comparable performance. Across methods, substantial split-to-split variability was observed, indicating that multi-split evaluation is essential for reliable model comparison at this dataset scale. Expert reassessment of ambiguous samples improved the performance of all 17 evaluated methods by 2.8-8.1 percentage points, while having only a limited effect on cross-split variance. The results highlight both the potential of deep learning for automated X-ray-based hazelnut quality assessment and the importance of rigorous evaluation and label curation in small, imbalanced agricultural imaging datasets.
Chinese Translation
非破坏性X射线成像能够揭示仅通过外部检查难以发现的榛子内部缺陷;然而,由于类别之间微妙的放射差异、显著的类别不平衡以及有限的标注数据,自动化解读仍然具有挑战性。在此,我们提出了一个基于799幅分割单核X射线图像(224 x 224像素,灰度)的榛子质量二分类(健康与缺陷)的基准,这些图像被分为101个采集单元。我们使用基于五个不同随机种子生成的数据划分,采用组内分割旋转协议评估了七种单模型配置和十种概率聚合集成。决策阈值在验证集上选择,并在验证集和测试集上进行确定性性能评估。在专家重新评估的标注条件下,经过二元交叉熵训练的卷积神经网络和冻结的Swin Transformer的平均概率集成达到了最高的均衡平均准确率(86.3% +/- 1.8%,五个种子),其他几个集成也提供了可比的性能。在不同方法中,观察到显著的分割间变异性,这表明在此数据集规模下,多分割评估对于可靠的模型比较至关重要。对模糊样本的专家重新评估使所有17种评估方法的性能提高了2.8-8.1个百分点,同时对跨分割方差的影响有限。结果强调了深度学习在基于X射线的榛子质量评估中的潜力,以及在小型不平衡农业成像数据集中严格评估和标签管理的重要性。
cs.CV / 44 / 2608.11765

ProBAG: Prototype-Guided Boundary-Aware Graph Diffusion for Weakly Supervised Histopathology Segmentation

ProBAG:基于原型引导的边界感知图扩散用于弱监督组织病理学分割
Nguyen, Duy-Dong, Thai, Le-Van, Pham, Hoai Nhan, Bui, Ngoc Lam Quang, Tran, Tam, Huang, Zhi
Abstract
Weakly supervised semantic segmentation enables histopathology tissue segmentation from image-level annotations, avoiding costly pixel-level labeling by expert pathologists. However, CAM-based methods often localize only highly discriminative regions and remain unreliable near tissue interfaces. We propose ProBAG, a stage-1 pseudo-mask generator that combines dataset-specific visual prototypes with pathology-aligned CONCH text prototypes over multi-scale frozen UNI features. ProBAG introduces two complementary mechanisms: class-wise power recalibration that reshapes inter-class competition while preserving the total foreground activation mass at each pixel, and one-step graph diffusion in which feature affinities are penalized by a late-transformer attention-context discrepancy used as a soft structural boundary cue. The resulting stage-1 pseudo-masks require neither CRF nor an external segmentation model; for complete two-stage comparison, they additionally supervise a downstream Phikon-FPN segmenter. Experiments on BCSS-WSSS and LUAD-HistoSeg show consistent gains over recent WSSS approaches, while ablations indicate that pathology-aligned text semantics provide the largest improvement and graph refinement provides a smaller complementary gain. The code is available at: https://github.com/wterrr/WSSS
Chinese Translation
弱监督语义分割使得通过图像级注释进行组织病理学组织分割成为可能,避免了专家病理学家进行昂贵的像素级标注。然而,基于CAM的方法往往仅定位于高度可区分的区域,并且在组织界面附近表现不可靠。我们提出了ProBAG,一个阶段一伪掩膜生成器,它结合了数据集特定的视觉原型和与病理对齐的CONCH文本原型,基于多尺度冻结的UNI特征。ProBAG引入了两种互补机制:类别级别的功率重校准,重塑了类别间的竞争,同时在每个像素处保持总前景激活量,以及一步图扩散,其中特征亲和力受到用于软结构边界线索的后期变换器注意力上下文差异的惩罚。生成的阶段一伪掩膜不需要CRF或外部分割模型;为了进行完整的两阶段比较,它们还监督下游的Phikon-FPN分割器。在BCSS-WSSS和LUAD-HistoSeg上的实验显示出相对于近期WSSS方法的一致提升,而消融实验表明,与病理对齐的文本语义提供了最大的改善,而图形细化则提供了较小的互补增益。代码可在以下链接获取:https://github.com/wterrr/WSSS
cs.CV / 45 / 2608.11770

Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines

在实时检测管道中实现近零开销的多模型层次分类
Raju, Vaishnav
Abstract
Edge-deployed vision systems in target recognition, surveillance, autonomous vehicles, and drone domains require hierarchical inference pipelines where a detection model identifies objects of interest and downstream classifiers provide fine-grained attribute analysis. Running all models on the GPU creates a serial bottleneck that limits real-time throughput as pipeline stages grow. Modern edge SoCs pair GPUs with dedicated neural accelerators (NPUs, DLAs) capable of concurrent execution, yet deploying custom models on these accelerators remains impractical due to strict operator constraints, quantization incompatibilities, and an undocumented end-to-end pipeline. We target NVIDIA Jetson DLA cores as the representative platform. We present a five-step methodology for zero GPU fallback DLA INT8 deployment of classification backbones, comprising architecture adaptation, manual dynamic range workaround to rescue TensorRT's implicit quantization (recovering 94.0% accuracy from implicit quantization's 75%) for rapid pipeline validation before explicit quantization, quantization-aware training, ONNX graph surgery for DLA compilation, and a concurrent GPU-detection/DLA-classification inference pipeline. We document nine engineering constraints with root-cause analysis and generalizable solutions. Validation on a dual-head person attribute classifier running on DLA alongside a GPU object detector on a Jetson Orin NX demonstrates near-zero pipeline overhead (12.5 vs. 13.3~FPS detector-only at 1080p), with dual-DLA scaling at no additional cost. The methodology is backbone-agnostic and generalizes to any detection-classification edge pipeline.
Chinese Translation
在目标识别、监控、自动驾驶汽车和无人机领域,边缘部署的视觉系统需要层次推理管道,其中检测模型识别感兴趣的对象,下游分类器提供细粒度的属性分析。在GPU上运行所有模型会造成串行瓶颈,限制了随着管道阶段增长而实现的实时吞吐量。现代边缘系统芯片(SoCs)将GPU与专用神经加速器(NPUs, DLAs)配对,能够并发执行,但由于严格的操作符限制、量化不兼容性以及未记录的端到端管道,在这些加速器上部署自定义模型仍然不切实际。我们以NVIDIA Jetson DLA核心作为代表平台。我们提出了一种五步方法,旨在实现分类骨干网络的零GPU回退DLA INT8部署,包括架构适配、手动动态范围解决方案以拯救TensorRT的隐式量化(从隐式量化的75%恢复94.0%的准确率),以便在显式量化之前快速验证管道、量化感知训练、用于DLA编译的ONNX图修剪,以及并发的GPU检测/DLA分类推理管道。我们记录了九个工程约束,并进行了根本原因分析和可推广的解决方案。在Jetson Orin NX上运行的双头人属性分类器与GPU目标检测器的验证显示,管道开销近乎为零(1080p下检测器仅为12.5 vs. 13.3~FPS),且双DLA扩展没有额外成本。该方法与骨干网络无关,适用于任何检测-分类边缘管道。
cs.CV / 46 / 2608.11777

VOLA: Improving Open-World Driving by VLM-Based Semantic Attribute Prediction

VOLA:通过基于VLM的语义属性预测改善开放世界驾驶
Zhang, Yuchen, Gao, Yuan, Schmidt, Sebastian, Betz, Johannes
Abstract
Driving in the real world is open-world: a car may encounter a fallen mattress, a deer, or other objects outside its training data. Naming them is not enough. The system must know how to treat each region: can it drive over it, and how severe would a collision be? We therefore shift scene perception from category labels to dense action-relevant attributes, where each pixel is labeled by how it should affect motion rather than by object name. We instantiate this general formulation with two ordered attributes: 7-rank drivability and 5-rank vulnerability. We read Qwen3.5 image-token hidden states directly as a spatial semantic representation. A lightweight boundary-aware decoder then turns this coarse token grid into sharp full-resolution attribute maps. The whole process requires neither autoregressive text generation nor an external mask model such as SAM. We train on dense attribute labels built in CARLA and test transfer to real scenes and to novel obstacles never seen in training. We compare with vision-only segmenters trained on the same attributes and prompted VLM segmenters. Our model matches strong vision-only segmenters on familiar categories and improves transfer to real open-world anomalies, reaching 69.4% mean vulnerability-rank recall versus 57.1% for the best vision-only baseline and 53.9% for the best prompted VLM baseline. These results show that VLM image tokens provide useful semantic cues for transferring driving attributes to objects outside the training vocabulary.
Chinese Translation
现实世界中的驾驶是开放世界的:汽车可能会遇到掉落的床垫、鹿或其他超出其训练数据的物体。仅仅命名这些物体是不够的。系统必须知道如何处理每个区域:它能否驶过,碰撞的严重程度如何?因此,我们将场景感知从类别标签转变为密集的与动作相关的属性,其中每个像素被标记为如何影响运动,而不是物体名称。我们用两个有序属性来具体化这一一般性表述:7级可驾驶性和5级脆弱性。我们直接读取Qwen3.5图像-标记的隐藏状态作为空间语义表示。一个轻量级的边界感知解码器将这个粗略的标记网格转换为清晰的全分辨率属性图。整个过程既不需要自回归文本生成,也不需要外部掩码模型,如SAM。我们在CARLA构建的密集属性标签上进行训练,并测试其在真实场景和从未在训练中见过的新障碍物上的迁移。我们与在相同属性上训练的仅视觉分割器和提示的VLM分割器进行了比较。我们的模型在熟悉类别上与强大的仅视觉分割器相匹配,并改善了对真实开放世界异常的迁移,达到69.4%的平均脆弱性等级召回率,而最佳仅视觉基线为57.1%,最佳提示VLM基线为53.9%。这些结果表明,VLM图像标记为将驾驶属性转移到训练词汇之外的物体提供了有用的语义线索。
cs.CV / 47 / 2608.11789

Anti-Shortcut Distillation via Temporal Negative Knowledge Transfer

通过时间负知识转移的反快捷蒸馏
Raza, Syed Muhammad, Tariq, Omer, Son, Jeongbae
Abstract
Knowledge distillation (KD) trains a compact student by attracting it towards a converged teacher. It is silent about which directions the teacher itself learned to suppress: repulsive and bias-aware objectives exist, but none exploits the teacher's own trajectory to identify what the student should avoid. We observe that the missing signal is already encoded in the teacher's optimization trajectory: features that an early-stage teacher emphasizes but that a converged teacher attenuates are precisely the shortcut directions worth pushing the student away from. We instantiate this observation as \textbf{A}nti-\textbf{S}hortcut \textbf{D}istillation (ASD), a push--pull KD framework that treats the converged teacher $\Tfinal$ as a positive semantic anchor and an early-checkpoint teacher $\Tearly$ as a temporal negative reference. ASD couples two losses: a temporal contrastive loss ($\Ltc$) that places the early-teacher feature as a same-sample negative against in-batch and memory-bank final-teacher features in an InfoNCE objective; and a shortcut suppression loss ($\Lss$) that penalizes student projection onto the top eigenvectors of $\E[\Dh\Dh^{\top}]$, the uncentered second-moment matrix of early-to-final feature displacements. Across 13 teacher--student pairs on CIFAR-100, ImageNet-100, and TinyImageNet, ASD attains the highest clean top-1 accuracy on more than 10 pairs and outperforms standard KD on 12. On CIFAR-100-C corruption robustness, ASD obtains the lowest mean Corruption Error ($86.1$\,mCE) on the most challenging cross-architecture pair (WRN-40-2$\to$ShuffleNet-V2). Mechanistic diagnostics confirm the intended geometry: the ASD student is systematically anti-aligned with the shortcut direction, while its projection onto the robust subspace is substantially larger ($0.45$ vs.\ $0.12$).
Chinese Translation
知识蒸馏(Knowledge Distillation, KD)通过将紧凑的学生模型吸引到收敛的教师模型来进行训练。它并未明确指出教师模型学习到的需要抑制的方向:虽然存在排斥和偏差感知的目标,但没有利用教师模型自身的轨迹来识别学生模型应避免的内容。我们观察到,缺失的信号已经编码在教师模型的优化轨迹中:早期教师模型强调的特征,而收敛教师模型减弱的特征,正是值得将学生模型推开快捷方向的特征。我们将这一观察实例化为反快捷蒸馏(Anti-Shortcut Distillation, ASD),这是一种推拉式的知识蒸馏框架,将收敛教师模型 $ ext{T}_{ ext{final}}$ 视为正语义锚点,而将早期检查点教师模型 $ ext{T}_{ ext{early}}$ 视为时间负参考。ASD 结合了两个损失函数:一个时间对比损失($ ext{L}_{ ext{tc}}$),将早期教师特征作为同样样本的负样本,与批次内和记忆库中的最终教师特征进行对比,采用 InfoNCE 目标;另一个是快捷抑制损失($ ext{L}_{ ext{ss}}$),对学生模型在早期到最终特征位移的未中心化二阶矩阵 $ ext{E}[ ext{D} ext{h} ext{D} ext{h}^{ op}]$ 的顶特征向量上的投影进行惩罚。在 CIFAR-100、ImageNet-100 和 TinyImageNet 上的 13 对教师-学生模型中,ASD 在超过 10 对模型上达到了最高的清晰顶级准确率,并在 12 对模型上超越了标准 KD。在 CIFAR-100-C 的破坏鲁棒性测试中,ASD 在最具挑战性的跨架构对(WRN-40-2$ o$ShuffleNet-V2)上获得了最低的平均破坏错误(86.1 mCE)。机制诊断确认了预期的几何结构:ASD 学生模型与快捷方向系统性反对齐,而其在鲁棒子空间上的投影显著增大(0.45 对比 0.12)。
cs.CV / 48 / 2608.11793

PolarSym: Polar Geometry-aware Attention for CAD Floorplan Parsing

PolarSym:基于极坐标几何感知的CAD平面图解析注意力机制
Chen, Kerui, Wang, Yiqing, Xin, Kangzhou, Zhang, Qinghan, Ding, Songyang
Abstract
CAD plan parsing is a fundamental task in Building Information Modeling (BIM), aiming to automatically extract architectural elements including walls, doors, windows, and furniture from 2D engineering drawings. Existing Transformer-based methods capture global semantic dependencies via self-attention, yet they infer spatial relationships merely from semantic features without explicitly characterizing the intrinsic geometric symmetry of building layouts. Such methods tend to produce mismatched correspondences in long-range matching and complex symmetric spatial layouts. To tackle this limitation, we propose PolarSym, a polar-coordinate geometry-aware attention framework for CAD plan parsing. The framework decouples geometric relationships of buildings into two complementary components, direction and distance, which are modeled independently. Structural consistency is strengthened by directional constraints, while long-range symmetric correspondences are built with distance constraints. A dynamic gating mechanism is adopted to synergistically fuse the two geometric information branches while maintaining the vanilla Transformer architecture. This design boosts geometric modeling capacity with negligible extra computation. Experiments on a public CAD plan parsing dataset show that PolarSym surpasses the reproduced SymPoint V2 baseline by 1.73% PQ, 1.54% RQ and 4.31% mIoU under identical training settings. PolarSym also converges faster and yields more stable optimization. Ablation experiments verify the complementary effects of direction and distance modeling. Our results reveal that PolarSym improves the geometric awareness of Transformers at low computational cost, offering an effective geometric modeling paradigm for CAD plan parsing.
Chinese Translation
CAD平面图解析是建筑信息建模(BIM)中的一项基础任务,旨在从二维工程图纸中自动提取建筑元素,包括墙壁、门、窗和家具。现有的基于Transformer的方法通过自注意力机制捕捉全局语义依赖,但它们仅从语义特征推断空间关系,而没有明确表征建筑布局的内在几何对称性。这类方法在长距离匹配和复杂对称空间布局中往往会产生不匹配的对应关系。为了解决这一局限性,我们提出了PolarSym,一种用于CAD平面图解析的极坐标几何感知注意力框架。该框架将建筑的几何关系解耦为两个互补的组件:方向和距离,并独立建模。通过方向约束增强结构一致性,同时利用距离约束建立长距离对称对应关系。采用动态门控机制协同融合这两个几何信息分支,同时保持原始Transformer架构。这一设计在几乎不增加额外计算的情况下提升了几何建模能力。在一个公共CAD平面图解析数据集上的实验表明,PolarSym在相同训练设置下超越了重现的SymPoint V2基线,分别提高了1.73%的PQ、1.54%的RQ和4.31%的mIoU。PolarSym还收敛更快,优化更稳定。消融实验验证了方向和距离建模的互补效果。我们的结果表明,PolarSym以低计算成本提升了Transformer的几何感知能力,为CAD平面图解析提供了一种有效的几何建模范式。
cs.CV / 49 / 2608.11807

CoDiR: Confidence-Guided Diffusion Refinement for Semi-Supervised Histopathology Segmentation

CoDiR:基于置信度引导的扩散精细化用于半监督组织病理分割
Pham, Hoai Nhan, Bui, Dang-Nguyen, Thai, Le-Van, Vo, Thanh-Hiep, Thi, Lan Anh Dinh, Nguyen, Tien Dat, Nguyen, Duy-Dong, Bui, Ngoc Lam Quang, Tran, Tam, Huang, Zhi
Abstract
Semi-supervised histopathology segmentation is challenging due to scarce annotations and unreliable pseudo-labels in ambiguous gland regions. To address this problem, we propose Confidence-Guided Diffusion Refinement (CoDiR), a semi-supervised framework that combines a Mean Teacher segmentation model with diffusion-based pseudo-label refinement. Given an unlabeled image, the teacher first produces a soft prediction, and only low-confidence regions are refined by a conditional diffusion model trained to capture plausible mask structures from labeled data. The refined mask is then fused with reliable teacher predictions and used to train the student with confidence weighting and consistency regularization. On the GlaS and CRAG datasets CoDiR reaches 88.09\% and 89.83\% mDice with 10\% labeled data, and 89.19\% and 90.29\% mDice with 20\%, matching or exceeding the strongest published method on seven of the eight benchmark metrics. Ablations attribute the largest single contribution to the refinement module, which adds +6.36\% mDice over the Mean Teacher baseline. The implementation code is publicly available at: https://github.com/vongla345/codir
Chinese Translation
半监督组织病理分割因注释稀缺和模糊腺体区域中不可靠的伪标签而具有挑战性。为了解决这个问题,我们提出了基于置信度引导的扩散精细化(CoDiR),这是一个将均值教师分割模型与基于扩散的伪标签精细化相结合的半监督框架。给定一幅未标记的图像,教师首先生成一个软预测,只有低置信度区域通过一个条件扩散模型进行精细化,该模型经过训练以捕捉来自标记数据的合理掩膜结构。然后,将精细化的掩膜与可靠的教师预测融合,并用于通过置信度加权和一致性正则化来训练学生。在GlaS和CRAG数据集上,CoDiR在10%的标记数据下达到了88.09%和89.83%的mDice,在20%的标记数据下达到了89.19%和90.29%的mDice,匹配或超过了八个基准指标中七个的最强已发布方法。消融实验表明,精细化模块贡献最大,相较于均值教师基线增加了6.36%的mDice。实现代码已公开可用,链接为:https://github.com/vongla345/codir
cs.CV / 50 / 2608.11810

Can Vision Models Read the Radar Display? On the Feasibility of Radar Imagery for Air Traffic Complexity Estimation

视觉模型能读取雷达显示吗?雷达图像在空中交通复杂性估计中的可行性研究
Kim, Hyewook, Kang, Byul, Yoon, Seokbin, Lee, Keumjin
Abstract
Air traffic controllers perceive traffic complexity through the radar display, suggesting that a computer vision model operating on the same imagery may provide a natural architecture for modeling controller-perceived complexity; however, whether radar imagery is a viable input format for deep learning vision models remains unclear. Unlike natural images, radar images are extremely sparse and self-similar, consisting primarily of a black background and a few visually identical aircraft blobs, while small changes in aircraft positions can substantially alter sector-level complexity. To test whether a vision model can capture these operationally important differences, we encode each traffic situation as a position image supplemented by five channels representing aircraft state variables, including heading, speed, and altitude, and train a Vision Transformer (ViT) to regress four intrinsic complexity components derived from pairwise geometric relations among aircraft. The model achieves $R^2 > 0.96$ for all four components, and a one-aircraft-removal perturbation study shows that its response changes proportionally to how much the removed aircraft contributed to sector complexity rather than treating every removal as equivalent. These results demonstrate that, despite its atypical visual characteristics, radar imagery is a viable input format for air traffic complexity modeling.
Chinese Translation
空中交通管制员通过雷达显示感知交通复杂性,这表明基于相同图像操作的计算机视觉模型可能为建模管制员感知的复杂性提供一种自然架构;然而,雷达图像是否是深度学习视觉模型的可行输入格式仍不清楚。与自然图像不同,雷达图像极其稀疏且自相似,主要由黑色背景和少量视觉上相同的飞机斑块组成,而飞机位置的细微变化可以显著改变区域级复杂性。为了测试视觉模型是否能够捕捉这些在操作上重要的差异,我们将每个交通情况编码为一个位置图像,并补充五个通道,表示飞机状态变量,包括航向、速度和高度,并训练一个视觉变换器(Vision Transformer, ViT)来回归从飞机之间的成对几何关系中得出的四个内在复杂性组件。该模型在所有四个组件上达到了 $R^2 > 0.96$,而一次飞机移除扰动研究表明,其响应与被移除飞机对区域复杂性的贡献成比例变化,而不是将每次移除视为等同。这些结果表明,尽管雷达图像具有非典型的视觉特征,但它是空中交通复杂性建模的可行输入格式。
cs.CV / 51 / 2608.11820

TD-VAD: Breaking Visual Dependence in Video Anomaly Detection with Text-Driven Learning

TD-VAD:通过文本驱动学习打破视频异常检测中的视觉依赖
Zhang, Shuangqing, Ma, Lei-Lei, Wang, Zhao, Dong, Wen, Xu, Xinyi, Xie, Guo-Sen, Shan, Caifeng, Zhao, Fang
Abstract
Visual data is typically a prerequisite for training existing video anomaly detection (VAD) methods. However, obtaining sufficient annotated anomaly data for training is challenging and not scalable due to the rarity of anomaly data and the wide variety of abnormal events. In this work, we advocate that the effectiveness of treating texts as video sequences for the VAD model and propose a novel Text-Driven Video Anomaly Detection (TD-VAD) approach to break visual dependence. In contrast to the anomaly video data, text descriptions of abnormal events are easy to collect, and their class labels can be directly derived. Specifically, our method utilizes video-like text descriptions with temporal characteristics generated by LLM to train a VAD model, without any reliance on target-domain anomaly data. To capture the long- and short-range temporal logic of events, we design the event evolution causal attention module to model contextual dependencies across time. During inference, considering the domain gap between the texts and video sequences, we use the frozen CLIP encoder to extract embeddings of video frames to align the text modality while retaining crucial visual information. Comprehensive experiments on two large-scale VAD datasets, XD-Violence and UCF-Crime, demonstrate that our method outperforms prior one-class and unsupervised VAD methods by a large margin.
Chinese Translation
视觉数据通常是训练现有视频异常检测(VAD)方法的前提。然而,由于异常数据的稀缺性和异常事件的多样性,获取足够的标注异常数据进行训练是具有挑战性的,并且不可扩展。在本研究中,我们主张将文本视为视频序列来处理VAD模型的有效性,并提出了一种新颖的文本驱动视频异常检测(TD-VAD)方法,以打破视觉依赖。与异常视频数据相比,异常事件的文本描述易于收集,并且其类别标签可以直接推导。具体而言,我们的方法利用由大型语言模型(LLM)生成的具有时间特征的视频类文本描述来训练VAD模型,而无需依赖目标领域的异常数据。为了捕捉事件的长短期时间逻辑,我们设计了事件演变因果注意力模块,以建模跨时间的上下文依赖。在推理过程中,考虑到文本和视频序列之间的领域差距,我们使用冻结的CLIP编码器提取视频帧的嵌入,以对齐文本模态,同时保留关键的视觉信息。在两个大规模VAD数据集XD-Violence和UCF-Crime上的全面实验表明,我们的方法在性能上大幅超越了之前的一类和无监督VAD方法。
cs.CV / 52 / 2608.11835

Distractor-Aware Video Object Segmentation

考虑干扰物的视频目标分割
Robinson, Andreas, Eldesokey, Abdelrahman, Felsberg, Michael
Abstract
Semi-supervised video object segmentation is a challenging task that aims to segment a target throughout a video sequence given an initial mask at the first frame. Discriminative approaches have demonstrated competitive performance on this task at a sensible complexity. These approaches typically formulate the problem as a one-versus-one classification between the target and the background. However, in reality, a video sequence usually encompasses a target, background, and possibly other distracting objects. Those objects increase the risk of introducing false positives, especially if they share visual similarities with the target. Therefore, it is more effective to separate distractors from the background, and handle them independently. We propose a one-versus-many scheme to address this situation by separating distractors into their own class. This separation allows imposing special attention to challenging regions that are most likely to degrade the performance. We demonstrate the prominence of this formulation by modifying the learning-what-to-learn (LWL) method to be distractor-aware. Our proposed approach sets a new state-of-the-art on the DAVIS 2017 val dataset, and improves over the baseline on the DAVIS 2017 test-dev benchmark by 4.6 percentage points.
Chinese Translation
半监督视频目标分割是一项具有挑战性的任务,旨在根据第一帧的初始掩膜在整个视频序列中对目标进行分割。区分性方法在这一任务上表现出竞争力的性能,同时保持合理的复杂性。这些方法通常将问题表述为目标与背景之间的二分类。然而,实际上,一个视频序列通常包含目标、背景以及可能的其他干扰物体。这些物体增加了引入假阳性的风险,特别是当它们与目标具有视觉相似性时。因此,将干扰物与背景分离并独立处理更为有效。我们提出了一种一对多的方案,通过将干扰物分离到自己的类别中来解决这一情况。这种分离允许我们对最有可能降低性能的挑战性区域施加特别关注。我们通过修改学习-学习什么(Learning-What-to-Learn, LWL)方法,使其能够考虑干扰物,展示了这一表述的突出性。我们提出的方法在DAVIS 2017验证数据集上设定了新的最先进水平,并在DAVIS 2017测试开发基准上提高了4.6个百分点。
cs.CV / 53 / 2608.11838

GeoBridge: Decoupled Semantic Conditioning for Generative Image Geolocalization

GeoBridge:用于生成图像地理定位的解耦语义条件机制
Dou, Zhiyang, Han, Xumeng, Peng, Fengde, Wang, Zipeng, Zhao, Moxuan, Huang, Zhipei, Han, Zhenjun
Abstract
Multimodal large language models (MLLMs) have advanced image geolocalization mainly by improving how they reason about geographic cues. How that reasoning isdecoded into coordinates, however, has lagged behind. Predicting a place name for a geocoding API is discrete and lossy: it ignores image evidence and collapses multi-granular semantics into a coarse lookup. We argue that the bottleneck has shifted from what a model reasons to how that reasoning is represented for a continuous, geometry-aware decoder. We present GeoBridge, a role-decoupled conditioning mechanism that connects a frozen semantic MLLM to a frozen Riemannian flow-matching head that generates coordinates on the sphere. The central obstacle is arole conflict: supervising the condition with discrete semantic labels biases its representation toward class-discriminative geometry, at odds with the smooth manifold the generative head requires. GeoBridge keeps the semantic supervision decoupled from the condition interface: a separate projection forms the continuous condition the frozen head expects, injecting geographic priors without disturbing the spherical decoder. On IM2GPS3K, GeoBridge reaches 38.67/52.89/70.37 at the 25/200/750 km thresholds, improving over a place-name-to-API pipeline and reasoning-augmented direct prediction at these precision-relevant scales. GeoBridge is a decode-side algorithmic contribution, orthogonal and complementary to chain-of-thought reasoning. Code will be made publicly available.
Chinese Translation
多模态大型语言模型(MLLMs)通过改善对地理线索的推理,推动了图像地理定位的发展。然而,这种推理如何解码为坐标却滞后于此。为地理编码API预测地点名称是离散且有损的:它忽略了图像证据,并将多层次语义压缩为粗略的查找。我们认为瓶颈已从模型推理的内容转向如何将这种推理表示为连续的、几何感知的解码器。我们提出了GeoBridge,一种角色解耦的条件机制,它将一个冻结的语义MLLM与一个冻结的黎曼流匹配头连接起来,后者在球面上生成坐标。核心障碍是角色冲突:用离散语义标签监督条件会使其表示偏向于类别区分几何,这与生成头所需的平滑流形相悖。GeoBridge保持语义监督与条件接口的解耦:一个单独的投影形成了冻结头所期望的连续条件,在不干扰球面解码器的情况下注入地理先验。在IM2GPS3K数据集上,GeoBridge在25/200/750公里阈值下分别达到了38.67/52.89/70.37的成绩,相比于地点名称到API的管道和增强推理的直接预测,在这些与精度相关的尺度上有所提升。GeoBridge是一个解码侧的算法贡献,与思维链推理正交且互补。代码将公开发布。
cs.CV / 54 / 2608.11844

BoltNet: An Ultra-Lightweight Convolutional Network for On-Device Plant Species Identification

BoltNet:一种超轻量级卷积网络用于设备端植物物种识别
Rossi, Daniel, Borghi, Guido, Vezzani, Roberto
Abstract
Automated plant species identification from citizen-science imagery is an established, demanding fine-grained recognition problem: large taxonomic label spaces, visually similar species, and long-tailed observations require real model capacity, while field use constrains memory, latency, and power. Model size is only part of the deployment cost: intermediate activations held in memory during inference and platformdependent execution behavior matter too, so compact recognition must be assessed on target hardware rather than through complexity metrics alone. We present BoltNet, an ultra-lightweight fully convolutional architecture combining a Spatial Redistribution Bottleneck and Logit PreSampling to improve the tradeoff between predictive performance and model size in high-cardinality classification, and report the AccuracyCompression Tradeoff as a complementary diagnostic. On Pl@ntNet300K, BoltNet reaches 0.682 F1-score with 341K parameters (1.37 MB), the highest F1-score among evaluated models below 2 MB and close to substantially larger convolutional backbones. Model-only measurements on a Raspberry Pi 5, Jetson Orin Nano, and Hailo-8 characterize execution across CPU, GPU, and NPU platforms, where BoltNet is the most consistently efficient model, with the best FPS/W on the GPU and NPU and second-best on the CPU. Results on AIDERv2 and CLRS provide secondary evidence of transfer across environmental image-classification tasks. Code available at: https://codeberg.org/danielrossi/BoltNet
Chinese Translation
从公民科学图像中自动识别植物物种是一项成熟且要求严格的细粒度识别问题:庞大的分类标签空间、视觉上相似的物种以及长尾观察都需要真实的模型能力,而现场使用又限制了内存、延迟和功耗。模型大小只是部署成本的一部分:推理过程中保留在内存中的中间激活和平台依赖的执行行为同样重要,因此紧凑的识别必须在目标硬件上进行评估,而不仅仅通过复杂度指标。我们提出了BoltNet,这是一种超轻量级的全卷积架构,结合了空间重分配瓶颈(Spatial Redistribution Bottleneck)和逻辑预采样(Logit PreSampling),以改善高维分类中预测性能与模型大小之间的权衡,并报告了准确性-压缩权衡(Accuracy-Compression Tradeoff)作为补充诊断。在Pl@ntNet300K数据集上,BoltNet以341K参数(1.37 MB)达到了0.682的F1-score,这是评估模型中低于2 MB且接近大得多的卷积骨干网的最高F1-score。在Raspberry Pi 5、Jetson Orin Nano和Hailo-8上的模型测量表征了在CPU、GPU和NPU平台上的执行情况,其中BoltNet是最一致高效的模型,在GPU和NPU上具有最佳的FPS/W,在CPU上则为第二佳。在AIDERv2和CLRS上的结果提供了跨环境图像分类任务转移的次要证据。代码可在:https://codeberg.org/danielrossi/BoltNet
cs.CV / 55 / 2608.11847

LookBack: Where and How to Score LVLM Responses via Visual Reference Usage

LookBack:如何通过视觉参考使用来评分LVLM响应
Cho, Beomsik, Kim, Jinhyeong, Lee, Dongseok, Kim, Jaehyung
Abstract
Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning. However, LVLMs do not just inherit the text-level hallucinations; they also hallucinate against the image, producing fluent responses ungrounded in what they see. This makes LVLM response scoring inherently harder, and our diagnostics show that existing confidence-based metrics adopted from LLMs are insufficient for LVLMs. Specifically, removing the input image barely changes confidence-based selection, suggesting that output-space confidence primarily captures textual plausibility rather than agreement with the image. To address this gap, we propose LookBack, a training-free LVLM response scoring method that augments token likelihood with visual lookback score, a lightweight measure of how strongly each response token refers to image tokens. Across four benchmarks and three models, LookBack consistently improves Best-of-$N$ selection over existing baselines with negligible additional overhead.
Chinese Translation
大型视觉语言模型(LVLMs)将视觉感知与语言生成相结合,使得响应能够涵盖图像理解和复杂推理。然而,LVLMs不仅继承了文本层面的幻觉;它们还会在图像上产生幻觉,生成与所见内容无关的流畅响应。这使得LVLM响应评分本质上变得更加困难,我们的诊断表明,从大型语言模型(LLMs)借用的现有基于置信度的指标对LVLMs来说是不够的。具体而言,去除输入图像几乎不会改变基于置信度的选择,这表明输出空间的置信度主要捕捉文本的合理性,而不是与图像的一致性。为了解决这一问题,我们提出了LookBack,这是一种无训练的LVLM响应评分方法,通过视觉回顾评分增强了标记的可能性,这是一种轻量级的度量,表明每个响应标记与图像标记的相关程度。在四个基准测试和三个模型中,LookBack在现有基准上始终改善了最佳选择(Best-of-$N$),且附加开销微乎其微。
cs.CV / 56 / 2608.11883

Warping Earth Observations for better ice labeling in the Marginal Marginal Ice Zone

扭曲地球观测数据以改善边缘冰区的冰层标注
Kelly, Tom, Rogers, Martin S. J.
Abstract
Multimodal satellite imagery provides complementary information for Earth Observation, but accurately combining heterogeneous sensors remains challenging in dynamic environments. Fast-changing regions, such as the Antarctic marginal ice zone, cannot fully exploit multimodal information from different satellite sensors because surface features move between image acquisitions. This spatial and temporal mismatch challenges effective perceptual grounding, violating the assumption of pixel-level correspondence that underpins most multimodal reasoning and downstream classification pipelines. Antarctic sea ice provides a challenging benchmark due to the rapid, heterogeneous drift of individual ice floes and the differing responses of sea ice to radar, visible and thermal sensing modalities. Accurate, dense supervision of sea ice remains scarce because generating pixel-wise labels requires time-consuming expert interpretation of noisy data, leading to historical reliance on coarse-resolution maritime ice charts for model training. This paper presents a novel architecture based on mutual information warping to align multi-satellite (Sentinel-1 and MODIS platforms) multimodal (visible, thermal, radar) satellite scenes. To demonstrate the approach, we introduce a sparse expert-labeled dataset of 2,088 pixel-wise annotations (7,046 expert point classifications) located at the ice-water margin interface across 43 scenes. Our results demonstrate that spatially grounding and aligning modalities prior to segmentation improves classification accuracy, and enables accurate, dense sea ice segmentation from sparse point-wise supervision.
Chinese Translation
多模态卫星影像为地球观测提供了互补信息,但在动态环境中准确结合异构传感器仍然具有挑战性。快速变化的区域,如南极边缘冰区,无法充分利用来自不同卫星传感器的多模态信息,因为表面特征在影像采集之间发生移动。这种空间和时间的不匹配挑战了有效的感知基础,违反了大多数多模态推理和下游分类流程所依赖的像素级对应假设。南极海冰由于单个冰块的快速、异质漂移以及海冰对雷达、可见光和热感知模态的不同响应,提供了一个具有挑战性的基准。由于生成像素级标签需要耗时的专家对噪声数据的解释,因此对海冰的准确、密集监督仍然稀缺,导致历史上依赖于粗分辨率的海洋冰图进行模型训练。本文提出了一种基于互信息扭曲的新架构,以对齐多卫星(Sentinel-1和MODIS平台)多模态(可见光、热、雷达)卫星场景。为了演示该方法,我们引入了一个稀疏专家标注的数据集,其中包含2,088个像素级注释(7,046个专家点分类),位于43个场景的冰水交界面。我们的结果表明,在分割之前对模态进行空间基础和对齐可以提高分类准确性,并使得从稀疏点级监督中实现准确、密集的海冰分割成为可能。
cs.CV / 57 / 2608.11907

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models

你看到你所绘制的内容吗?一个用于统一多模态模型整体评估的语义闭环框架
Zhang, Hao, Qi, Jiaxin, Tang, Zhijiang, Huang, Jianqiang
Abstract
As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge. Current evaluation protocols predominantly treat generative and discriminative capabilities as separate tasks, leaving a gap in system-level evaluation for unified multimodal models (UMMs). In this work, we propose Self-Generative-Understanding (SGU), a novel, annotation-free evaluation framework that probes the integrated capabilities of unified models through a semantic closed-loop challenge. Without requiring new annotations, SGU leverages the dual understanding-and-generation abilities of UMMs by asking them to first perceive an image and produce a textual description, subsequently reconstruct a visual context based on that description, and finally perform reasoning over the self-generated output. This pipeline provides a zero-cost testbed that yields an integrated performance score specifically tailored for evaluating UMMs as unified systems. Extensive experiments show that even high-performing UMMs often struggle to reason over their own generated contexts, revealing limitations that are not captured by separate evaluations of understanding or generation alone. Our work provides a complementary holistic evaluation framework and offers a foundation for benchmarking the development of next-generation unified multimodal models.
Chinese Translation
随着大型视觉-语言模型越来越多地旨在在单一参数空间内整合视觉生成和理解,如何以连贯的方式评估这种结构统一性仍然是一个关键挑战。目前的评估协议主要将生成能力和判别能力视为独立任务,这在统一多模态模型(UMMs)的系统级评估中留下了空白。在本研究中,我们提出了自生成理解(Self-Generative-Understanding, SGU),这是一种新颖的无注释评估框架,通过语义闭环挑战探测统一模型的集成功能。SGU不需要新的注释,而是通过要求UMMs首先感知图像并生成文本描述,随后根据该描述重建视觉上下文,最后对自生成的输出进行推理,从而利用UMMs的双重理解和生成能力。该流程提供了一个零成本的测试平台,生成专门用于评估UMMs作为统一系统的集成性能评分。大量实验表明,即使是高性能的UMMs,往往也难以对其自身生成的上下文进行推理,揭示了仅通过理解或生成的独立评估无法捕捉的局限性。我们的工作提供了一个互补的整体评估框架,并为下一代统一多模态模型的发展基准提供了基础。
cs.CV / 58 / 2608.11913

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion

HarmoniDPO:通过偏好优化扩散进行视频引导的音频生成
Peng, Wenshuo, Zhang, Kaipeng
Abstract
Video-to-audio (V2A) generation faces significant challenges in achieving precise temporal synchronization and high perceptual quality due to the complex, ambiguous relationship between visual and auditory cues. Existing methods typically compress video inputs into single feature representations, leading to significant loss of temporal dynamics and fine-grained visual information. These approaches also rely on reconstruction-based training objectives that poorly correlate with human perceptual judgments of audio quality and appropriateness. We propose HarmoniDPO, a novel framework that integrates preference-based optimization into diffusion-based V2A generation to address these limitations. (1) Our approach leverages a dual video representation: combining global context with frame-wise features to preserve temporal dynamics and semantic detail. (2) Inspired by reinforcement learning from human feedback (RLHF), HarmoniDPO employs online Direct Preference Optimization (online-DPO) to fine-tune a diffusion-based V2A model from preference judgments, enhancing perceptual quality and alignment. (3) Additionally, we introduce Dual-scale Diffusion Search (DDS), a test time scaling algorithm that adaptively optimizes output fidelity during inference. Experiments demonstrate that HarmoniDPO outperforms state-of-the-art methods in audio-video synchronization and subjective audio quality, offering a robust solution for generating realistic, human-preferred audio from video.
Chinese Translation
视频到音频(V2A)生成在实现精确的时间同步和高感知质量方面面临重大挑战,这主要源于视觉和听觉线索之间复杂且模糊的关系。现有方法通常将视频输入压缩为单一特征表示,导致时间动态和细粒度视觉信息的显著损失。这些方法还依赖于基于重建的训练目标,这与人类对音频质量和适宜性的感知判断相关性较差。我们提出了HarmoniDPO,这是一种新颖的框架,将基于偏好的优化整合到基于扩散的V2A生成中,以解决这些局限性。(1)我们的方法利用双重视频表示:结合全局上下文与逐帧特征,以保留时间动态和语义细节。(2)受人类反馈强化学习(RLHF)的启发,HarmoniDPO采用在线直接偏好优化(online-DPO)从偏好判断中微调基于扩散的V2A模型,增强感知质量和一致性。(3)此外,我们引入了双尺度扩散搜索(DDS),这是一种测试时缩放算法,在推理过程中自适应优化输出保真度。实验表明,HarmoniDPO在音频-视频同步和主观音频质量方面优于最先进的方法,为从视频生成真实且符合人类偏好的音频提供了一种稳健的解决方案。
cs.CV / 59 / 2608.11928

Seed2GS: Camera-Free, Training-Free Object Extraction from 3D Gaussian Scenes via a Single Reference-View Grounding

Seed2GS:无摄像头、无训练的基于单一参考视图的3D高斯场景目标提取
Ding, Zongjian, Gao, Yudong, Liu, Jiale, Yu, Xinglin, Ren, Junxing, Wei, Dong, Chen, Yajing, Huang, Shan, Cheng, Mingjun, Li, Min
Abstract
Extracting a target object from a pre-built 3D Gaussian Splatting (3DGS) scene enables interactive 3D editing. Existing methods either train for tens of minutes per scene, sacrifice accuracy, or require original reconstruction cameras that pre-built assets may not include. We present Seed2GS, which achieves the highest reported LERF-MASK accuracy without original reconstruction cameras or scene-specific representation training. Its key insight is to separate target identity from 3D coverage. QD-SAM3 selects one reliable reference mask from several open-vocabulary candidates, fixing identity once. Seed lift and visibility-adaptive virtual orbits then expose the object from new viewpoints, while tracking propagates the seed without repeated detection. Because the scene remains frozen, these masks supervise only one temporary foreground logit per Gaussian. On LERF-MASK, Seed2GS reaches 92.1% mean intersection over union (mIoU) with a measured compute-only latency of 9.3 seconds, 3.7 points above the strongest scene-trained baseline and 7.6 points above the closest camera-free baseline. With one fixed test reference per scene, the complete pipeline retains 91.1% mIoU; replacing its predicted seed with a ground-truth mask improves mIoU by only 0.72 points. On 3D-OVS, Seed2GS reaches 95.7% mIoU.
Chinese Translation
从预构建的3D高斯喷溅(3DGS)场景中提取目标对象能够实现交互式3D编辑。现有方法要么每个场景训练数十分钟,要么牺牲准确性,或者需要原始重建摄像头,而预构建资产可能不包含这些摄像头。我们提出了Seed2GS,它在没有原始重建摄像头或场景特定表示训练的情况下,实现了最高报告的LERF-MASK准确率。其关键见解在于将目标身份与3D覆盖分离。QD-SAM3从多个开放词汇候选中选择一个可靠的参考掩码,固定身份一次。种子提升和适应可见性的虚拟轨道随后从新视角暴露对象,同时跟踪传播种子而无需重复检测。由于场景保持静止,这些掩码仅对每个高斯监督一个临时前景logit。在LERF-MASK上,Seed2GS达到了92.1%的平均交并比(mIoU),测得的计算延迟为9.3秒,比最强的场景训练基线高出3.7点,比最接近的无摄像头基线高出7.6点。对于每个场景一个固定的测试参考,完整的管道保持91.1%的mIoU;用真实掩码替换其预测的种子仅提高了0.72点的mIoU。在3D-OVS上,Seed2GS达到了95.7%的mIoU。
cs.CV / 60 / 2608.11933

Dual Anchors, Do It Better: Hierarchical Group Merging for Zero-Shot Anomaly Detection

双锚点,做得更好:用于零样本异常检测的层次化组合并
Roh, Jimin, Kim, DongKyu, Kang, Suk-Ju
Abstract
Zero-shot anomaly detection (ZSAD) aims to identify anomalies in unseen domains, a setting that is particularly critical for industrial and medical applications where domain shifts are prevalent. However, most CLIP-based ZSAD methods anchor semantics solely on the text modality, making performance highly sensitive to prompt design and leading to weak visual grounding. To mitigate these limitations, we propose a Dual-Anchor framework that complements conventional text anchors with hierarchical image anchors constructed via a top-down grouping mechanism. This mechanism progressively aggregates local-to-global image features to form normal and abnormal group tokens, which serve as image anchors and act as gating signals in a Group-Gated Token Refiner to enhance the global representation. The refined image anchors are then fused with text prompts to construct dynamic state prompts. By jointly reinforcing visual and textual semantics, our framework stabilizes image-text alignment, reduces prompt dependency, and achieves strong generalization across 8 industrial and 6 medical benchmarks.
Chinese Translation
零样本异常检测(ZSAD)旨在识别未见领域中的异常,这一设置在工业和医疗应用中尤为重要,因为这些领域常常面临领域转移的挑战。然而,大多数基于CLIP的ZSAD方法仅将语义锚定于文本模态,这使得性能对提示设计高度敏感,并导致视觉基础较弱。为了解决这些限制,我们提出了一种双锚点框架,该框架通过自上而下的分组机制,结合传统文本锚点与层次化图像锚点。该机制逐步聚合局部到全局的图像特征,以形成正常和异常组标记,这些标记作为图像锚点,并在组门控标记细化器(Group-Gated Token Refiner)中充当门控信号,以增强全局表示。然后,将细化后的图像锚点与文本提示融合,以构建动态状态提示。通过共同强化视觉和文本语义,我们的框架稳定了图像-文本对齐,减少了对提示的依赖,并在8个工业基准和6个医疗基准上实现了强大的泛化能力。
cs.CV / 61 / 2608.11938

Surfsvr: 2D Surface Priors as 3D Geometric Regularizers for Sparse Voxel Reconstruction

SurfSVR:将2D表面先验作为3D几何正则化器用于稀疏体素重建
Di, Yan, Li, Chengxi, Wang, Yaoxing, Liu, Mengge, Li, Zhigang, Zhang, Ruida, Li, Mingyang, Wang, Pengyuan, Gao, Shan, Ji, Xiangyang
Abstract
Sparse voxel reconstruction offers an efficient representation for high-fidelity 3D modeling, yet its geometry is commonly optimized from local photometric evidence and discrete visibility statistics. This often leads to fragmented surfaces, excessive subdivision, and floating artifacts, particularly in weakly textured or sparsely observed regions. We introduce SurfSVR, a novel sparse voxel reconstruction paradigm that treats 2D surface priors as explicit 3D geometric regularizers. Instead of directly lifting noisy pixel-wise depth predictions, SurfSVR first organizes each image into coherent surface regions by jointly reasoning over appearance, monocular depth, normals and cross-view geometry. Each region is then represented by an adaptively selected planar or quadratic surface model based on fitting reliability and geometric complexity, while cross-model agreement distinguishes reliable geometry from ambiguous predictions. These structured 2D priors are lifted into 3D and integrated throughout the reconstruction pipeline. They guide surface-adaptive voxel subdivision, provide region-level depth and normal supervision during optimization, enhance geometrically reliable sparse-observed surfaces in voxel pruning, and suppress off-surface floaters during post-refinement training. This unified design converts semantic and geometric coherence in image space into persistent structural constraints in 3D. Extensive experiments on 3 public benchmarks demonstrate that SurfSVR consistently improves sparse voxel reconstruction across scenes with substantially different visibility and geometry characteristics, achieving state-of-the-art reconstruction quality. Codes and models will be released soon.
Chinese Translation
稀疏体素重建为高保真3D建模提供了一种高效的表示方式,但其几何形状通常是基于局部光度证据和离散可见性统计进行优化的。这往往导致表面破碎、过度细分和浮动伪影,尤其是在纹理较弱或观察稀疏的区域。我们提出了SurfSVR,一种新颖的稀疏体素重建范式,将2D表面先验视为显式的3D几何正则化器。SurfSVR并不是直接提升嘈杂的像素级深度预测,而是首先通过共同推理外观、单目深度、法线和跨视图几何,将每幅图像组织成连贯的表面区域。然后,根据拟合可靠性和几何复杂性,采用自适应选择的平面或二次曲面模型来表示每个区域,而跨模型一致性则区分了可靠的几何形状和模糊的预测。这些结构化的2D先验被提升到3D,并在整个重建流程中进行集成。它们指导表面自适应的体素细分,在优化过程中提供区域级的深度和法线监督,增强几何上可靠的稀疏观察表面在体素修剪中的表现,并在后期细化训练中抑制离表浮动物。这种统一设计将图像空间中的语义和几何一致性转化为3D中的持久结构约束。在3个公共基准上的大量实验表明,SurfSVR在具有显著不同可见性和几何特征的场景中始终提高稀疏体素重建的效果,达到了最先进的重建质量。代码和模型将很快发布。
cs.CV / 62 / 2608.11942

Evaluating and Calibrating Diffusion Model-derived Uncertainty for Quantitative MRI Mapping

评估和校准基于扩散模型的定量MRI映射的不确定性
Wang, Shishuai, Klein, Stefan, Hernandez-Tamames, Juan A., Poot, Dirk H. J.
Abstract
Quantitative MRI (qMRI) provides standardised tissue parameter maps, but the reliability of deep learning-based qMRI mapping methods is often not explicitly characterised. In this work we systematically evaluate uncertainty maps for quantitative MRI derived from multiple inferences of a data-consistent diffusion model-based qMRI framework. Evaluation on synthetic test data assessed error-awareness, high-error detection, selective prediction, and Gaussian interval calibration. Diffusion model-derived uncertainty was positively associated with the mapping error, while risk-coverage analysis showed that excluding high-uncertainty voxels reduced the retained error. However, the raw uncertainty was poorly calibrated for quantitative interval interpretation. Calibration was substantially improved using a post-hoc procedure combining prediction-value-dependent bias correction with scalar uncertainty scaling. Qualitative evaluation on a healthy volunteer showed spatially meaningful uncertainty patterns. These results indicate that diffusion model-derived uncertainty is informative for reliability assessment and selective prediction, but requires calibration for quantitative interval interpretation.
Chinese Translation
定量MRI(qMRI)提供标准化的组织参数图,但基于深度学习的qMRI映射方法的可靠性往往未被明确表征。在本研究中,我们系统地评估了基于数据一致的扩散模型的qMRI框架推导的定量MRI不确定性图。对合成测试数据的评估考察了错误意识、高错误检测、选择性预测和高斯区间校准。基于扩散模型的不确定性与映射误差呈正相关,而风险覆盖分析显示,排除高不确定性体素减少了保留的误差。然而,原始不确定性在定量区间解释方面校准效果较差。通过后处理程序结合预测值依赖的偏差修正与标量不确定性缩放,校准得到了显著改善。对健康志愿者的定性评估显示出空间上有意义的不确定性模式。这些结果表明,基于扩散模型的不确定性对可靠性评估和选择性预测具有信息价值,但在定量区间解释方面需要进行校准。
cs.CV / 63 / 2608.11985

Auditing Frame-Level AUC in Weakly Supervised Video Anomaly Detection: Granularity, Resolution, and Scene Bias

审计弱监督视频异常检测中的帧级 AUC:粒度、分辨率与场景偏差
Abdulaziz, Sara, Bondarev, Egor
Abstract
Frame-level area under the ROC curve (AUC) is the dominant evaluation metric for weakly supervised video anomaly detection (WSVAD). Its standard form measures whether an anomalous frame outranks a normal frame drawn from anywhere in the test set. We refer to this comparison as pooled AUC, since it aggregates frame pairs across test videos regardless of source. Pooled AUC therefore credits both event localization and differences between video sources. We audit this protocol on UCF-Crime across recent state-of-the-art models spanning different backbone families. Holding each model's frame scores fixed, we read them under three pairing granularities: global, per anomaly category, and within each video, then repeat the same three-granularity readout on zero-shot scores computed from the models' internal representations. We assess ranking reliability with a paired video bootstrap. Three findings follow. First, pooled AUC does not reliably predict within-video anomaly localization: models with similar pooled scores exhibit large localization differences and rank reversals under stricter granularities. Second, at the benchmark's test-split size, pooled AUC lacks the resolution to support state-of-the-art margins reported in the field. Within each backbone family, it resolves no comparison at those margins, while within-video AUC resolves several over identical predictions. Learned representations further reveal that within-video anomaly structure and detector localization are decoupled. Third, on normal footage alone, every model we examine separates videos by recording properties, such as resolution and color encoding, indicating that scene sensitivity is shared across the setting rather than specific to any architecture. We publicly release a granularity-aware protocol computable from existing predictions and scene-factor annotations for UCF-Crime.
Chinese Translation
帧级接收者操作特征曲线下面积(AUC)是弱监督视频异常检测(WSVAD)的主要评估指标。其标准形式衡量异常帧是否优于从测试集中的任何位置抽取的正常帧。我们将这种比较称为汇总 AUC,因为它聚合了来自不同测试视频的帧对,而不考虑来源。因此,汇总 AUC 同时考虑了事件定位和视频来源之间的差异。我们在 UCF-Crime 数据集上审计这一协议,涵盖了不同主干网络系列的最新先进模型。在固定每个模型的帧分数的情况下,我们在三种配对粒度下读取这些分数:全局、按异常类别和在每个视频内,然后对从模型内部表示计算的零-shot 分数重复相同的三种粒度读取。我们通过配对视频自助法评估排名的可靠性。得出了三个发现。首先,汇总 AUC 并不能可靠地预测视频内的异常定位:具有相似汇总分数的模型在更严格的粒度下表现出较大的定位差异和排名反转。其次,在基准测试的测试集大小下,汇总 AUC 缺乏支持该领域报告的先进边际的分辨率。在每个主干网络系列内,它在这些边际上没有解决任何比较,而在相同预测下,视频内 AUC 则解决了多个问题。学习到的表示进一步揭示了视频内异常结构与检测器定位是解耦的。第三,仅在正常视频上,我们检查的每个模型通过记录属性(如分辨率和颜色编码)将视频分开,表明场景敏感性在设置中是共享的,而不是特定于任何架构。我们公开发布了一种粒度感知协议,可从现有预测和 UCF-Crime 的场景因子注释中计算得出。
cs.CV / 64 / 2608.11996

A Remote Approach to Cashew Orchard Detection: Leveraging Active Learning with Satellite Imagery in Guinea-Bissau

一种遥感方法用于腰果果园检测:利用卫星影像中的主动学习在几内亚比绍的应用
Miguel, Sofia, Maria, Patrícia, Luke, João
Abstract
Cashew production is a widespread economic activity in Guinea-Bissau, as well as other countries in West Africa. However, unregulated cashew production can be directly associated with increasing regionwide deforestation rates, biodiversity losses, and a fragile economic structure. There is no nationwide database for listing or georeferencing cashew orchards, so there is a clear need to remotely map their locations. In recent years, multiple methods for detecting orchards have been developed, though they have only been applied on a regional level. This work expands regional analyses to a nationwide scale. It develops a scalable and cost-effective remote approach, based on Sentinel-2 satellite imagery, using Machine Learning techniques to detect cashew orchards automatically. Margin-based Active Learning techniques were employed to develop an optimal training set in terms of the number of points and their informativeness, leading to a cashew map with 94.0% balanced accuracy obtained entirely off-site. We created two datasets and a 2021 cashew map with 10m spatial resolution that are openly accessible through GitHub. The results demonstrate the possibility of a broader cashew orchard mapping, creating a new stepping stone for this environmental application.
Chinese Translation
腰果生产是几内亚比绍及其他西非国家广泛的经济活动。然而,未经监管的腰果生产与区域性森林砍伐率的上升、生物多样性的丧失以及脆弱的经济结构直接相关。目前尚无全国性数据库来列出或地理参考腰果果园,因此迫切需要对其位置进行遥感制图。近年来,虽然已经开发出多种果园检测方法,但这些方法仅应用于区域层面。本研究将区域分析扩展到全国范围,开发了一种基于Sentinel-2卫星影像的可扩展且具有成本效益的遥感方法,利用机器学习技术自动检测腰果果园。采用基于边际的主动学习技术,开发出在点数和信息量方面最优的训练集,从而生成了一幅具有94.0%平衡准确率的腰果地图,完全在异地获得。我们创建了两个数据集和一幅具有10米空间分辨率的2021年腰果地图,并通过GitHub公开访问。结果表明,进行更广泛的腰果果园制图是可能的,为这一环境应用创造了新的起点。
cs.CV / 65 / 2608.12032

LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration

LoSA:用于无训练视频扩散加速的近无损稀疏注意力
Liu, Enhuai, Wang, Yunke, Wang, Yutong, Sun, Changming, Xu, Chang
Abstract
Video diffusion transformers are costly to sample: every denoising step applies self-attention over a long 3D token sequence, a quadratic cost that dominates as resolution and duration grow. Sparse attention reduces this cost without retraining, but existing methods pursue aggressive sparsity, where further speedup costs disproportionately more attention fidelity. We target the opposite end of this trade-off: fix near-lossless fidelity by construction, and remove as much computation as this constraint permits. Two observations make this regime practical: roughly 40% of block interactions can be removed while retaining 99% of the attention mass, and the high-mass support remains stable across denoising steps. We propose LoSA, a training-free sparse-attention method that fixes a retained-mass threshold of 99% rather than a sparsity ratio: it measures exact block attention masses at one early dense step, keeps, for each head and query block, the smallest key/value block set meeting the threshold, and reuses the frozen block indices for all remaining steps. On Wan2.1-1.3B, LoSA alone gives a $1.36\times$ speedup with a 0.06-point VBench Overall drop. The benefit is largest under composition: combined with feature caching, LoSA reaches a $3.2\times$ speedup on HunyuanVideo at a 0.02-point drop, versus 0.32 points for the strongest sparse baseline at comparable speed. Across three video diffusion transformers and speedups up to $3.2\times$, LoSA consistently achieves the best training-free speed-quality trade-off.
Chinese Translation
视频扩散变换器的采样成本较高:每个去噪步骤都对长达三维的标记序列应用自注意力,这种二次成本在分辨率和持续时间增加时占主导地位。稀疏注意力在不重新训练的情况下降低了这一成本,但现有方法追求激进的稀疏性,进一步的加速会导致注意力保真度的不成比例损失。我们针对这一权衡的相反端:通过构造固定近无损的保真度,并在这一约束允许的范围内去除尽可能多的计算。有两个观察使得这一方案切实可行:大约40%的块交互可以被去除,同时保留99%的注意力质量,并且高质量支持在去噪步骤中保持稳定。我们提出了LoSA,一种无训练的稀疏注意力方法,它固定了99%的保留质量阈值,而不是稀疏比率:它在一个早期的密集步骤中测量确切的块注意力质量,为每个头和查询块保留满足阈值的最小键/值块集,并在所有剩余步骤中重用冻结的块索引。在Wan2.1-1.3B上,LoSA单独提供了1.36倍的加速,VBench总体下降0.06点。在组合情况下,收益最大:结合特征缓存,LoSA在HunyuanVideo上实现了3.2倍的加速,下降0.02点,而在相似速度下,最强稀疏基线下降0.32点。在三个视频扩散变换器和高达3.2倍的加速中,LoSA始终实现了最佳的无训练速度-质量权衡。
cs.CV / 66 / 2608.12035

How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging

距离临床应用还有多远?评估医学影像中的完整无监督领域适应流程
Xiong, Yiheng, Gallée, Luisa, Wolf, Daniel Santak, Hillenhagen, Heiko, Götz, Michael
Abstract
Deploying unsupervised domain adaptation (UDA) in clinical practice requires choosing which algorithm to use and which of its trained models to ship. However, the deployment (target) domain is unlabeled, so models cannot be evaluated directly on it, leaving it unclear which to select. We address this by evaluating the complete UDA pipeline, considering both adaptation and label-free selection together. Our study covers eleven clinically relevant cross-domain scenarios from nine medical imaging datasets, with ten UDA algorithms and 13 label-free selection methods (validators), evaluating over 80,000 trained models in total. By this, we find that a capable adapted model usually exists, but identifying it without target labels is difficult: the validator-selected models leave a large and structural target performance gap to the best available one, with no evaluated validator consistently reliable. Towards closing it, we explore two strategies, ensembling and a small target-labeling budget; both narrow this gap but do not close it entirely. Overall, deployable UDA depends on the complete pipeline; addressing the less explored selection step could bring much of current UDA closer to clinical use.
Chinese Translation
在临床实践中部署无监督领域适应(UDA)需要选择使用哪种算法以及选择其训练模型。然而,部署(目标)领域是未标记的,因此无法直接在其上评估模型,这使得选择变得不明确。我们通过评估完整的UDA流程来解决这一问题,同时考虑适应和无标签选择。我们的研究涵盖了来自九个医学影像数据集的十一种临床相关的跨领域场景,使用了十种UDA算法和13种无标签选择方法(验证器),总共评估了超过80,000个训练模型。通过这些研究,我们发现通常存在一个能够适应的模型,但在没有目标标签的情况下识别它是困难的:验证器选择的模型与最佳可用模型之间存在较大且结构性的目标性能差距,且没有一个评估过的验证器是一致可靠的。为了缩小这一差距,我们探索了两种策略:集成和小规模目标标记预算;这两者都缩小了这一差距,但并未完全消除。总体而言,可部署的UDA依赖于完整的流程;解决较少探索的选择步骤可能会使当前的UDA更接近临床应用。
cs.CV / 67 / 2608.12045

Localizing to Debias: A Patch-Level Benchmark and Baseline for Weakly Supervised Spatial Anomaly Detection

去偏见的本地化:弱监督空间异常检测的补丁级基准和基线
Abdulaziz, Sara, Al-Abri, Abdulrahman, D'Amicantonio, Giacomo, Bondarev, Egor
Abstract
Despite growing interest in weakly supervised video anomaly detection (WSVAD), current methods struggle to bridge the gap between coarse temporal supervision and fine-grained spatial reasoning. A key obstacle is the tendency of temporal detectors to latch onto background and scene-level cues rather than truly discriminative anomaly evidence. This background bias raises ethical concerns: models may inadvertently associate anomalies with societal or environmental context rather than authentic crime-related cues. Without spatial grounding, such biases remain hidden and unauditable. To address this, we propose SST-WSVADL, a sparse spatio-temporal framework that bridges temporal anomaly detection with fine-grained spatial localization. Rather than processing all spatial regions indiscriminately, SST-WSVADL progressively focuses on the most anomaly-relevant spatio-temporal regions through dynamic sparsification, naturally suppressing background dominant content while preserving discriminative evidence. The temporal and spatial branches are coupled end-to-end via motion-aware regularization that guides sparsification toward dynamically informative regions, without relying on external detectors or vision-language prompts. We publicly release frame-level spatial annotations and a method-agnostic evaluation protocol for three public datasets: UCF-Crime, XD-Violence, and MSAD. These resources enable the community to audit spatial biases in WSVAD predictions, supporting progress toward more ethical and accountable anomaly detection. Experiments demonstrate that SST-WSVADL is competitive with prior methods across benchmarks while enabling localization and patch-level auditability of scene bias, providing a reproducible foundation for interpretability-oriented evaluation of WSVAD models.
Chinese Translation
尽管对弱监督视频异常检测(WSVAD)的兴趣日益增长,但当前方法在粗略的时间监督与细粒度的空间推理之间仍存在差距。一个主要障碍是时间检测器倾向于依赖背景和场景级线索,而非真正具有区分性的异常证据。这种背景偏见引发了伦理问题:模型可能无意中将异常与社会或环境背景联系起来,而非真实的犯罪相关线索。缺乏空间基础,这些偏见仍然隐藏且无法审计。为了解决这个问题,我们提出了SST-WSVADL,一个稀疏时空框架,旨在将时间异常检测与细粒度空间定位结合起来。SST-WSVADL并不是无差别地处理所有空间区域,而是通过动态稀疏化逐步关注最与异常相关的时空区域,自然抑制背景主导内容,同时保留区分性证据。时间和空间分支通过运动感知正则化端到端耦合,指导稀疏化朝向动态信息丰富的区域,而不依赖于外部检测器或视觉-语言提示。我们公开发布了帧级空间注释和一种方法无关的评估协议,适用于三个公共数据集:UCF-Crime、XD-Violence和MSAD。这些资源使得社区能够审计WSVAD预测中的空间偏见,支持向更具伦理性和责任感的异常检测的进展。实验表明,SST-WSVADL在各基准上与先前方法具有竞争力,同时实现了场景偏见的定位和补丁级审计,为面向可解释性评估的WSVAD模型提供了可重复的基础。
cs.CV / 68 / 2608.12050

Predicting Functions, Not Features: KANs with Function-Space Joint-Embedding Predictive Learning for Medical Image Segmentation

预测函数而非特征:基于功能空间联合嵌入预测学习的 KANs 在医学图像分割中的应用
Liu, Yungeng, Fang, Xuanzi, Zhang, Yuge, Ren, Shuqi, Zeng, Haijin, Chen, Yongyong
Abstract
Kolmogorov--Arnold Networks (KANs) introduce explicit functional representations by parameterizing each network edge as a learnable univariate function. However, existing KAN-based segmentation models optimize edge functions only through objectives defined after edge aggregation, leaving individual functions without an explicit pre-aggregation learning target. To address this limitation, we propose Function-Space Joint-Embedding Predictive Learning (FS-JEPA) for medical image segmentation. Our FS-JEPA framework moves predictive learning into the pre-aggregation function space of KANs. A masked online branch predicts structured signatures of sampled KAN edge functions generated by a full-context exponential moving average target branch, while shared edge indices preserve correspondence between predictions and targets. Rather than predicting an isolated edge response, we represent each sampled edge function using a multi-radius signature composed of function evaluations around its input anchor. This structured representation captures local functional variations that cannot be characterized by a single response and provides a more informative predictive target. The function-space objective is jointly optimized with the segmentation loss during training, while the predictive branch is removed at inference. Experiments on five medical image segmentation benchmarks show that our FS-JEPA achieves the best average Dice and outperforms the strongest competing KAN-based method by +2.25 percentage points.
Chinese Translation
Kolmogorov--Arnold 网络(KANs)通过将每个网络边缘参数化为可学习的一元函数,引入了显式的功能表示。然而,现有的基于 KAN 的分割模型仅通过在边缘聚合后定义的目标来优化边缘函数,导致单个函数缺乏明确的聚合前学习目标。为了解决这一限制,我们提出了功能空间联合嵌入预测学习(FS-JEPA)用于医学图像分割。我们的 FS-JEPA 框架将预测学习移入 KAN 的聚合前功能空间。一个掩码在线分支预测由全上下文指数移动平均目标分支生成的采样 KAN 边缘函数的结构化特征,同时共享的边缘索引保持预测与目标之间的对应关系。我们并非预测孤立的边缘响应,而是使用围绕其输入锚点的函数评估组成的多半径特征来表示每个采样的边缘函数。这种结构化表示捕捉了无法通过单一响应表征的局部功能变化,并提供了更具信息量的预测目标。功能空间目标与分割损失在训练期间共同优化,而预测分支在推理时被移除。在五个医学图像分割基准上的实验表明,我们的 FS-JEPA 实现了最佳的平均 Dice,并比最强的竞争 KAN 基础方法提高了 2.25 个百分点。
cs.CV / 69 / 2608.12051

Do Not Forget the Obvious - RISC: A Risk-Informed Slice-Coverage Protocol for Safe Autonomous Driving

不要忽视显而易见的 - RISC:一种基于风险的切片覆盖协议用于安全自主驾驶
Hüger, Fabian
Abstract
Aggregate metrics may not fully reflect performance in insufficiently examined high-risk driving conditions. We propose RISC (Risk-Informed Slice Coverage), a practical protocol for risk-guided stress testing and coverage-qualified evaluation. Risk-guided stress testing directs a finite audit budget toward risk-relevant sub-datasets, called risk slices, while coverage-qualified evaluation reports results together with explicit statements about which slices are sufficiently or insufficiently covered. The protocol translates safety concerns into machine-readable risk slices, uses lightweight signals to tag candidate data, selects a compact audit set by risk, and qualifies the results using coverage evidence. An LLM can optionally support this process by surfacing relevant but potentially overlooked conditions during test planning, thereby helping engineers not to forget the obvious. RISC is model-agnostic and can be applied to perception modules, driving models, and other autonomous-driving subsystems. We instantiate the protocol for monocular pedestrian perception using 1,000 frames from the Zenseact Open Dataset, image statistics, and a YOLO-based detector proxy. In this proof-of-concept study, risk-guided selection increases critical failure discovery from 34.0% under random sampling to 98.5%. RISC provides a lightweight, assurance-oriented evaluation layer that complements scenario categorization, coverage assessment, and broader testing-and-verification workflows.
Chinese Translation
聚合指标可能无法充分反映在未充分审查的高风险驾驶条件下的性能。我们提出了RISC(基于风险的切片覆盖),这是一种用于风险引导压力测试和覆盖合格评估的实用协议。风险引导压力测试将有限的审计预算指向与风险相关的子数据集,称为风险切片,而覆盖合格评估则报告结果并明确说明哪些切片得到了充分或不足的覆盖。该协议将安全问题转化为机器可读的风险切片,使用轻量级信号标记候选数据,根据风险选择紧凑的审计集,并使用覆盖证据对结果进行资格认证。大型语言模型(LLM)可以选择性地支持这一过程,在测试规划期间揭示相关但可能被忽视的条件,从而帮助工程师不忘显而易见的事情。RISC是模型无关的,可以应用于感知模块、驾驶模型和其他自主驾驶子系统。我们使用来自Zenseact开放数据集的1,000帧图像、图像统计数据和基于YOLO的检测器代理实例化该协议,用于单目行人感知。在这一概念验证研究中,风险引导选择将关键故障发现率从随机抽样下的34.0%提高到98.5%。RISC提供了一个轻量级、面向保障的评估层,补充了场景分类、覆盖评估和更广泛的测试与验证工作流程。
cs.CV / 70 / 2608.12064

Draw This First

首先绘制这个
Zhong, Dazhi, Bradbury, Rowan, Davis, Grant
Abstract
We invert the typical formulation of sketch generation: instead of drawing strokes in order, we predict a 2D field that defines the order in which strokes are drawn. We use a pretrained latent flow-matching transformer to supply the image prior to predict an intermediate representation, while training the VAE's decoder to predict the order field, stroke mask, and stroke segmentation. We vectorize the predicted segmentation into polylines and sort them by the field, producing an ordered vector sketch. Our model can predict an ordered vector sketch from a text description or derender an image into ordered vectors; for either, it follows text instructions specifying the order of drawing.
Chinese Translation
我们颠覆了传统的草图生成方式:我们不是按顺序绘制笔画,而是预测一个二维场,该场定义了笔画绘制的顺序。我们使用预训练的潜在流匹配变换器(latent flow-matching transformer)来提供图像先验,以预测一个中间表示,同时训练变分自编码器(VAE)的解码器以预测顺序场、笔画掩码和笔画分割。我们将预测的分割向量化为多线段,并根据该场进行排序,从而生成有序的向量草图。我们的模型可以根据文本描述预测有序的向量草图,或者将图像反渲染为有序向量;在这两种情况下,它都遵循指定绘制顺序的文本指令。
cs.CV / 71 / 2608.12078

Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World Models

更好的插槽,更好的世界:面向对象的世界模型中的表示质量与鲁棒性
Nazirjonov, Shukrullo, Prasanna, Sai, Manasyan, Anna, Martius, Georg
Abstract
Learning world models from offline trajectories enables agents to accomplish different tasks through planning. Object-centric (OC) representations, which decompose a scene into a set of slots that bind to its objects, have been proposed as an inductive bias for world models that are more sample-efficient and generalize better. Yet prior object-centric world models (OCWMs) take the slot encoder as given and evaluate only in-distribution, leaving open whether the object-centric bias actually delivers for planning and what within the OCWM drives it. We conduct a controlled study of OCWMs for visual model-predictive control along two axes: object-centric representation quality and generalization under distribution shift relative to scene-centric models. We find that (i) planning success correlates positively with unsupervised slot-quality metrics (FG-ARI, mBO), though the gains saturate at high slot quality; (ii) with well-bound slots, the auxiliary proprioception inputs and masking inductive bias that prior methods relied on become unnecessary; and (iii) under unseen distribution shifts, the OCWM with well-bound slots is more robust overall than the end-to-end trained scene-centric LeWM, while DINO-WM, built on similar frozen pretrained features, remains competitive -- pointing to pretrained features as a key contributor to robustness.
Chinese Translation
从离线轨迹中学习世界模型使得智能体能够通过规划完成不同的任务。面向对象的(Object-Centric, OC)表示将场景分解为一组与其对象绑定的插槽,已被提出作为一种归纳偏置,使得世界模型在样本效率上更高且具有更好的泛化能力。然而,先前的面向对象世界模型(Object-Centric World Models, OCWMs)将插槽编码器视为给定,并仅在分布内进行评估,尚未明确面向对象的偏置是否真正有助于规划,以及OCWM中的哪些因素驱动了这一点。我们针对视觉模型预测控制进行了OCWMs的控制研究,从两个方面进行考察:面向对象的表示质量和相对于场景中心模型的分布转移下的泛化能力。我们发现:(i)规划成功与无监督插槽质量指标(FG-ARI, mBO)呈正相关,尽管在高插槽质量时增益趋于饱和;(ii)在插槽绑定良好的情况下,先前方法所依赖的辅助本体感知输入和掩蔽归纳偏置变得不再必要;(iii)在未见的分布转移下,具有良好绑定插槽的OCWM整体上比端到端训练的场景中心LeWM更具鲁棒性,而基于类似冻结预训练特征构建的DINO-WM仍然具有竞争力——这表明预训练特征是鲁棒性的关键因素。
cs.CV / 72 / 2608.12086

Look What the Probes Dragged In! Real-World Chest X-ray Shortcuts in MedCLIP

看探针拖来了什么!MedCLIP中的真实世界胸部X光快捷方式
Pedersen, Nikolette, Sydendal, Regitze, Cheplygina, Veronika, Sourget, Théo
Abstract
Vision-language models, such as contrastive language-image pre-training (CLIP)-based approaches, have reached state-of-the-art (SOTA) results in medical artificial intelligence. However, recent work reveals that CLIP-based models remain vulnerable to shortcuts. We investigate how real-world shortcuts manifest across different layers of the medical CLIP-based model, MedCLIP, and its vision encoder, a frozen ResNet-50. We attach 17 linear classification probes to the intermediate layers of the ResNet-50 and train them on three different dataset configurations and targets: NIH-CXR14 (pneumothorax) and PadChest (cardiomegaly and pneumothorax). This setup allows us to observe model behaviour during evaluation using subgroup-based calibration and layer-wise confidence curves. We find that the final linear probes achieve a high AUROC but poor calibration in the models. The layer-wise confidence analyses suggest that shortcuts emerge at different depths. Patterns consistent with localised shortcuts, such as drains, appear at later layers, while patterns consistent with diffuse shortcuts, such as scanner-specific noise patterns, emerge earlier, aligning with previous work. Finally, we conduct a manual analysis of the images, which reveals data quality issues in both NIH-CXR14 and PadChest. Our findings underscore that even SOTA models remain vulnerable to shortcuts, and the need for high-quality and well-annotated datasets to draw solid conclusions. Code can be found on our GitHub: https://github.com/nikodice4/MedCLIP_shortcuts.
Chinese Translation
视觉-语言模型,如基于对比语言-图像预训练(CLIP)的方法,在医学人工智能领域已达到最先进(SOTA)的结果。然而,近期的研究表明,基于CLIP的模型仍然容易受到快捷方式的影响。我们研究了真实世界的快捷方式如何在医学CLIP模型MedCLIP的不同层次中表现出来,以及其视觉编码器,一个冻结的ResNet-50。我们在ResNet-50的中间层附加了17个线性分类探针,并在三种不同的数据集配置和目标上进行训练:NIH-CXR14(气胸)和PadChest(心脏肥大和气胸)。这种设置使我们能够在评估过程中使用基于子组的校准和逐层置信度曲线观察模型行为。我们发现最终的线性探针在模型中实现了高AUROC,但校准效果较差。逐层置信度分析表明,快捷方式在不同深度出现。与局部快捷方式(如排水管)一致的模式出现在较后层,而与扩散快捷方式(如扫描仪特定噪声模式)一致的模式则较早出现,这与之前的研究结果一致。最后,我们对图像进行了手动分析,揭示了NIH-CXR14和PadChest中存在的数据质量问题。我们的发现强调,即使是SOTA模型也仍然容易受到快捷方式的影响,并且需要高质量和良好注释的数据集以得出可靠的结论。代码可在我们的GitHub上找到:https://github.com/nikodice4/MedCLIP_shortcuts。
cs.CV / 73 / 2608.12088

RA-ClipScore: Making Generative Model Evaluation More Interpretable

RA-ClipScore:使生成模型评估更具可解释性
Lu, Yifan, Kucherenko, Taras, Kjellström, Hedvig, Bütepage, Judith
Abstract
Generative models can produce images nearly indistinguishable from real data, yet rigorous and interpretable evaluation remains challenging. Conventional metrics such as FID provide only scalar scores with limited diagnostic insight. Widely adopted CLIP-based metrics enable semantic evaluation beyond simple training class labels, but inherit limitations from CLIP's training paradigm that restrict attribute-wise analysis. We propose RA-CLIPScore, a novel metric that mitigates these issues and extends CLIP-based evaluation to spatial distribution alignment, measuring whether generated objects adhere to the positional priors found in the training data. RA-CLIPScore introduces dual prompts to decouple competing attributes and leverages local patch tokens to capture fine-grained regional semantics. We evaluate image generative models on their ability to match both attribute and spatial distributions of the training data. Extensive experiments show that RA-CLIPScore provides more robust and interpretable evaluations than prior methods, particularly under distribution misalignment or partially irrelevant textual attributes. We further demonstrate how it reveals spatial biases in generative models. User evaluations confirm that Regional Single Attribute Divergence based on our RA-CLIPScore aligns more closely with human perception of visual diversity than existing semantic metrics.
Chinese Translation
生成模型能够生成几乎与真实数据无法区分的图像,但严格且可解释的评估仍然具有挑战性。传统指标如FID仅提供有限的标量分数,缺乏诊断洞察。广泛采用的基于CLIP的指标使得超越简单训练类别标签的语义评估成为可能,但继承了CLIP训练范式的局限性,限制了属性级分析。我们提出了RA-CLIPScore,这是一种新颖的指标,旨在缓解这些问题,并将基于CLIP的评估扩展到空间分布对齐,测量生成对象是否遵循训练数据中的位置先验。RA-CLIPScore引入了双重提示,以解耦竞争属性,并利用局部补丁令牌捕捉细粒度的区域语义。我们评估图像生成模型在匹配训练数据的属性和空间分布方面的能力。大量实验表明,RA-CLIPScore提供了比以往方法更稳健且可解释的评估,特别是在分布不对齐或部分不相关文本属性的情况下。我们进一步展示了它如何揭示生成模型中的空间偏差。用户评估确认,基于我们的RA-CLIPScore的区域单一属性差异与人类对视觉多样性的感知比现有语义指标更为一致。
cs.CV / 74 / 2608.12107

Avatar-Forever: Decoupled Parallel Training for High-Quality Real-Time Infinite Avatars

Avatar-Forever:高质量实时无限化身的解耦并行训练
Li, Ruibin, Yang, Tao, Ma, Zhiyuan, Ai, Fangzhou, Wen, Shilei, Zhang, Lei
Abstract
Existing streaming video systems often rely on sequential, distillation-centered training pipelines to enable few-step long-video generation. However, this paradigm suffers from two limitations. First, failures or distribution shifts introduced in earlier stages affect later optimization, complicating the training process to converge. Second, the distillation-centric objective favours short-term generation but is prone to quality degradation when autoregressive errors accumulate over long rollouts. We propose Avatar-Forever, a decoupled parallel training framework for high-quality real-time infinite interactive avatars. Instead of coupling generation efficiency and long-horizon robustness under a sequential distillation pipeline, we treat them as two independent capabilities that can be trained in parallel. One branch performs full-parameter distillation to train an efficient generator with high visual quality, while another trains a lightweight long-horizon adapter via Recovery-oriented Rollout Training (RRT), which improves generation robustness under long-horizon inference conditions. Our decoupled parallel training design simplifies the overall training process and avoids unnecessary objective conflicts between few-step generation and long-horizon adaptation. We further introduce ForeverCache, a chunk-wise feature caching mechanism to substantially reduce redundant history computation during streaming inference. Built upon a 22B video foundation model, Avatar-Forever supports unbounded audio-driven avatar generation while maintaining identity consistency, motion coherence, and visual fidelity, enabling an end-to-end throughput of high-resolution 768x512 videos at 27.2 FPS on a single H100 GPU and providing a practical path toward stable digital humans.
Chinese Translation
现有的流媒体视频系统通常依赖于顺序的、以蒸馏为中心的训练流程来实现少步长视频生成。然而,这一范式存在两个局限性。首先,早期阶段引入的失败或分布变化会影响后续优化,复杂化训练过程的收敛。其次,以蒸馏为中心的目标偏向于短期生成,但在长时间推理中,自回归错误的累积容易导致质量下降。我们提出了Avatar-Forever,一种用于高质量实时无限互动化身的解耦并行训练框架。我们将生成效率和长期鲁棒性视为可以并行训练的两个独立能力,而不是在顺序蒸馏流程下将其耦合在一起。一个分支执行全参数蒸馏,以训练具有高视觉质量的高效生成器,而另一个分支通过面向恢复的回滚训练(Recovery-oriented Rollout Training, RRT)训练轻量级的长期适配器,从而提高在长期推理条件下的生成鲁棒性。我们的解耦并行训练设计简化了整体训练过程,避免了少步生成与长期适应之间不必要的目标冲突。我们进一步引入了ForeverCache,一种基于块的特征缓存机制,显著减少了流媒体推理过程中的冗余历史计算。基于一个22B的视频基础模型,Avatar-Forever支持无限制的音频驱动化身生成,同时保持身份一致性、运动连贯性和视觉保真度,实现了在单个H100 GPU上以27.2 FPS的速度生成高分辨率768x512视频的端到端吞吐量,为稳定数字人类提供了切实可行的路径。
cs.CV / 75 / 2608.12127

SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks

SCOPE-Router:面向执行任务的成本感知开放集VLM路由
Yu, Tao, Qu, Yifei, Cui, Zhiqing, Zhou, Pengfei, Luo, Zhongtian, Yang, Yujia, Chai, Shenghua, Jin, Haopeng, Zhang, Zhenghao, Wang, Xinming, Yi, Hongzhu, Zhao, Wangbo, Wan, Zhenglin, Huang, Yan, Yeshani, Luo, Jinwen, You, Yang
Abstract
Model routing aims to select the most suitable model from a candidate pool for each query, balancing quality and cost. Existing VLM routing research is limited to traditional VQA evaluation, lacks systematic calibration optimization for open-set scenarios, and employs training objectives that dilute multi-positive signals via softmax normalization without incorporating cost. We address these limitations with three contributions: (1)VLM-ExecRouterBench, the first execution-oriented VLM routing benchmark covering Code, Agentic, and Search domains with 11 candidate models spanning nearly two orders of magnitude in pricing; (2)SCOPE-Router, a dual-tower router that matches queries to model behavior profiles constructed via hybrid calibration (random/diagnostic/diversity sampling), enabling new models to join routing without retraining; (3)CRM+RCCR, an architecture-agnostic cost-aware objective that encodes cost preference into continuous relevance targets through per-pair independent scoring, eliminating multi-positive dilution while regularizing queries with similar routing preferences to be closer in the routing space. Empirically, SCOPE-Router achieves the best Rank Score on all three benchmarks, surpassing the runner-up by 1.84 points under OOD settings and by 6.75 points under doubly OOD open-set evaluation. When applied to four diverse routers, CRM+RCCR improves Rank Score by 1.25--6.21 points.
Chinese Translation
模型路由旨在从候选池中为每个查询选择最合适的模型,以平衡质量和成本。现有的VLM路由研究仅限于传统的视觉问答(VQA)评估,缺乏针对开放集场景的系统校准优化,并且采用的训练目标通过softmax归一化稀释了多正信号而未考虑成本。我们通过三项贡献来解决这些局限性:(1)VLM-ExecRouterBench,这是第一个面向执行的VLM路由基准,涵盖代码、代理和搜索领域,包含11个候选模型,价格跨度近两个数量级;(2)SCOPE-Router,一种双塔路由器,通过混合校准(随机/诊断/多样性采样)构建模型行为特征,能够将查询与模型行为特征匹配,使新模型能够在不重新训练的情况下加入路由;(3)CRM+RCCR,一种与架构无关的成本感知目标,通过每对独立评分将成本偏好编码到连续相关性目标中,消除了多正信号的稀释,同时将具有相似路由偏好的查询在路由空间中拉近。实证结果表明,SCOPE-Router在所有三个基准上均取得了最佳排名分数,在OOD设置下超越亚军1.84分,在双重OOD开放集评估中超越6.75分。当应用于四种不同的路由器时,CRM+RCCR将排名分数提高了1.25至6.21分。
cs.CV / 76 / 2608.12145

Autonomous Telerehabilitation via Skeletal Motion Prediction and Joint-Level Performance Assessment

基于骨骼运动预测和关节级性能评估的自主远程康复
Pereira, Lara, Paulo, João Ruivo, Santos, Pedro, Peixoto, Paulo
Abstract
Autonomous rehabilitation systems must not only recognize human motion but also provide structured feedback to support users without continuous therapist supervision. This paper presents a telerehabilitation pipeline that integrates skeleton-based exercise quality assessment and short-term motion prediction into a two-module system operating on marker-free RGB video. A self-attentive Bidirectional LSTM performs exercise quality classification using MMD-NCA metric learning, while a graph-based motion prediction module computes per-joint position errors between predicted and observed poses, generating spatially localized deviation signals. Each module is evaluated independently on established benchmarks: the classifier achieves 96.45% mean-class accuracy on squat sequences from the PROZIS dataset, and the adopted STARS predictor achieves a mean MPJPE of 75.8 mm at 560 ms on Human3.6M, outperforming graph and recurrent baselines across all prediction horizons. The framework is designed for eventual deployment in assistive robotics and home-based rehabilitation contexts; end-to-end integration and clinical validation are important directions for future work. By combining motion recognition and prediction in a single system, this work contributes a step toward autonomous, feedback-driven telerehabilitation, for more accessible and scalable rehabilitation solutions.
Chinese Translation
自主康复系统不仅需要识别人体运动,还需提供结构化反馈,以支持用户在没有持续治疗师监督的情况下进行康复。本文提出了一种远程康复流程,该流程将基于骨骼的运动质量评估和短期运动预测集成到一个在无标记RGB视频上运行的双模块系统中。自注意力双向长短期记忆网络(Bidirectional LSTM)使用MMD-NCA度量学习进行运动质量分类,而基于图的运动预测模块计算预测姿态与观察姿态之间的每个关节位置误差,生成空间局部化的偏差信号。每个模块在既定基准上独立评估:分类器在PROZIS数据集的深蹲序列上实现了96.45%的平均类别准确率,而采用的STARS预测器在Human3.6M数据集上以560毫秒的时间达到了75.8毫米的平均关节位置误差(MPJPE),在所有预测时间范围内均优于图形和递归基线。该框架旨在最终应用于辅助机器人和基于家庭的康复环境;端到端集成和临床验证是未来工作的重点方向。通过在一个系统中结合运动识别和预测,本研究为自主、反馈驱动的远程康复迈出了重要一步,为更可及和可扩展的康复解决方案提供了贡献。
cs.CV / 77 / 2608.12155

Understanding Why Foundation Models Work for Diffusion-Generated Image Detection

理解基础模型为何适用于扩散生成图像检测
Cozzolino, Davide, Poggi, Giovanni, Verdoliva, Luisa
Abstract
Vision foundation models have recently emerged as powerful feature extractors for detecting AI-generated images, achieving strong generalization across generators and robustness to common image degradations. However, the reason behind their effectiveness is poorly understood. In this work, we investigate what cues are exploited by foundation-model-based detectors to distinguish real images from diffusion-generated ones. To this end, we design an ad hoc analysis protocol based on DDIM inversion. Given a real image we generate a sequence of synthetic copies by changing the depth of DDIM inversion. Even though most copies are semantically identical to the real reference, the detector score varies significantly across them due to subtle traces introduced by the diffusion synthesis, showing that its decision is not primarily driven by semantic failures. Through a frequency-swapping analysis, we further reveal that the discriminative cues exploited by the detectors are mainly localized in the low-to-mid frequency range, rather than only in the high-frequency range, as is the case for artifacts commonly associated with generative models. Finally, a latent-space analysis shows that regenerated images exhibit reduced variance and effective dimensionality, indicating that diffusion models do not fully reproduce the variability of real data. Overall, our results suggest that foundation-model-based detectors succeed by capturing non-semantic low-to-mid frequency distributional discrepancies between real and diffusion-generated images. These findings provide new insight into the robustness and generalization of such detectors and suggest directions for more interpretable forensic methods.
Chinese Translation
视觉基础模型最近作为强大的特征提取器出现,用于检测人工智能生成的图像,在不同生成器之间实现了强大的泛化能力,并对常见图像退化具有鲁棒性。然而,它们有效性的原因尚不清楚。在本研究中,我们探讨了基础模型检测器利用哪些线索来区分真实图像和扩散生成的图像。为此,我们设计了一种基于DDIM反演的特定分析协议。给定一张真实图像,我们通过改变DDIM反演的深度生成一系列合成副本。尽管大多数副本在语义上与真实参考图像相同,但由于扩散合成引入的微妙痕迹,检测器的评分在这些副本之间显著变化,表明其决策并非主要受语义失败驱动。通过频率交换分析,我们进一步揭示检测器利用的判别线索主要集中在低到中频范围,而不仅仅是在高频范围,这与通常与生成模型相关的伪影情况不同。最后,潜在空间分析表明,重生图像表现出降低的方差和有效维度,表明扩散模型并未完全再现真实数据的变异性。总体而言,我们的结果表明,基于基础模型的检测器通过捕捉真实图像与扩散生成图像之间的非语义低到中频分布差异而取得成功。这些发现为此类检测器的鲁棒性和泛化能力提供了新的见解,并为更具可解释性的取证方法指明了方向。
cs.CV / 78 / 2608.12158

Context Blindness in DPO: Mitigating Object Hallucination in MLLMs via Context-Calibrated Preference Optimization

DPO中的上下文盲ness:通过上下文校准偏好优化减轻MLLM中的对象幻觉
Ko, Byungoh, Park, Jinyoung, Kim, Jongha, Na, Jeehye, Cho, Jaewon, Kim, Hyunwoo J.
Abstract
Multimodal large language models (MLLMs) have made rapid progress, yet they still exhibit object hallucination, generating plausible but incorrect descriptions that are inconsistent with the visual input. Direct Preference Optimization (DPO) mitigates this by training models to prefer non-hallucinated responses over hallucinated ones, and recent efforts further enrich the preference data with relevant context. However, it remains unclear whether DPO actually leverages such context. To investigate this, we propose Contextual Preference Gain (CPG), a simple metric that measures how much a model's preference strengthens when relevant context is provided. We find that higher CPG consistently corresponds to lower hallucination, yet standard DPO and its variants exhibit only limited CPG, indicating that they underutilize contextual information and thus remain prone to hallucination. To address this, we propose Context-Calibrated DPO (C$^2$-DPO), which directly maximizes CPG while preserving the original preference ordering. Across multiple benchmarks, C$^2$-DPO substantially reduces hallucination without compromising general reasoning, relatively reducing the Object HalBench hallucination rate of Qwen2-VL-Instruct-2B by 36%. Code is available at https://github.com/mlvlab/C2-DPO
Chinese Translation
多模态大型语言模型(MLLM)已取得快速进展,但仍然表现出对象幻觉,生成与视觉输入不一致的合理但不正确的描述。直接偏好优化(DPO)通过训练模型偏好非幻觉响应而非幻觉响应来减轻这一问题,最近的努力进一步丰富了相关上下文的偏好数据。然而,目前尚不清楚DPO是否真正利用了这些上下文。为此,我们提出了上下文偏好增益(CPG),这是一种简单的度量,衡量在提供相关上下文时模型偏好的增强程度。我们发现,较高的CPG与较低的幻觉率始终相关,但标准DPO及其变体的CPG仅限,表明它们未充分利用上下文信息,因此仍然容易出现幻觉。为了解决这个问题,我们提出了上下文校准DPO(C$^2$-DPO),它在保留原始偏好排序的同时直接最大化CPG。在多个基准测试中,C$^2$-DPO显著降低了幻觉率,而不影响一般推理,相对减少了Qwen2-VL-Instruct-2B在Object HalBench上的幻觉率36%。代码可在https://github.com/mlvlab/C2-DPO获取。
cs.CV / 79 / 2608.12175

TGRHuman: Text-Guided Realistic 3D Human Generation via Diffusion Renderer

TGRHuman:通过扩散渲染器实现文本引导的逼真3D人类生成
Zhang, Muxin, Yu, Chaohui, Yang, Yuanwang, Wei, Min, Su, Zhuo, Li, Kun
Abstract
Realistic 3D human generation plays a crucial role in many graphics applications. However, current methods still struggle to generate high-quality human geometry and texture while maintaining 3D consistency and inference efficiency. In this work, we address these limitations by introducing TGRHuman, a novel approach for generating realistic 3D humans from text. Our method decouples geometry and texture generation to alleviate the issues commonly encountered in NeRF-based methods. Instead of relying on slow, implicit score-distillation-based optimization, we directly use explicit multi-view observation generation and optimization for efficient 3D synthesis. For geometry generation, we propose a high-resolution generative module for multi-view normals together with a geometry-carving strategy that preserves view consistency and supports loose clothing. For texture generation, we produce spatially consistent RGB observations from densely sampled surrounding views using a carefully designed texture-prior acquisition strategy and a diffusion renderer, enabling detailed human texture synthesis. Experiments show that our method can generate high-quality and consistent 3D human geometry and texture efficiently. TGRHuman outperforms existing text-to-3D human methods in both geometry and texture quality.
Chinese Translation
逼真的3D人类生成在许多图形应用中发挥着至关重要的作用。然而,目前的方法在生成高质量的人体几何和纹理的同时,仍然面临着保持3D一致性和推理效率的挑战。在本研究中,我们通过引入TGRHuman,提出了一种从文本生成逼真3D人类的新方法,以解决这些局限性。我们的方法将几何和纹理生成解耦,以缓解基于NeRF的方法中常见的问题。我们不再依赖于缓慢的隐式评分蒸馏优化,而是直接使用显式的多视角观察生成和优化,以实现高效的3D合成。在几何生成方面,我们提出了一种高分辨率生成模块,用于多视角法线生成,并结合几何雕刻策略,以保持视图一致性并支持宽松的服装。在纹理生成方面,我们利用精心设计的纹理先验获取策略和扩散渲染器,从密集采样的周围视图中生成空间一致的RGB观察,进而实现细致的人体纹理合成。实验表明,我们的方法能够高效生成高质量且一致的3D人类几何和纹理。TGRHuman在几何和纹理质量上均优于现有的文本到3D人类生成方法。
cs.CV / 80 / 2608.12179

Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs

Map-Det3D:用于流输入的多视角3D物体检测的度量前馈3D重建先验
Yang, Yung-Hsu, Piccinelli, Luigi, Bulò, Samuel Rota, Hong, Sunghwan, Rozumny, Denis, Schönberger, Johannes, Bauer, Zuria, Blum, Hermann, Kontschieder, Peter, Pollefeys, Marc
Abstract
Metric 3D object detection is a core capability for embodied agents, yet most reliable systems lean on depth sensors, trading away cost, power, and integration simplicity. This motivates monocular 3D detection, which avoids additional constraints, yet it faces a major obstacle: from a single image, depth, and especially absolute scale, are underconstrained. As a result, the prevailing pattern of detecting in 2D and then predicting 3D attributes is often brittle, since modest range errors can dominate 3D localization, and the learned scale prior can fail when cameras, motion, or environments undergo domain shifts. To address this, we propose Map-Det3D, an online multi-view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB. We map a short temporal window into multiple views and repurpose a feed-forward metric 3D reconstruction model as our geometric backbone while tuning its object-aware capabilities. Building on this representation, Map-Det3D directly predicts boxes in metric 3D space, without the widely used 2D-to-3D lifting. Experiments across different benchmarks show that this design supports strong online performance and robust transfer without adaptation, suggesting that training reconstruction priors for detection is a practical route to stable metric 3D detection from monocular video. Code and models are available at https://royyang0714.github.io/Map-Det3D.
Chinese Translation
度量3D物体检测是具身智能体的一项核心能力,然而大多数可靠系统依赖深度传感器,这在成本、功耗和集成简易性上存在权衡。这促使了单目3D检测的发展,它避免了额外的约束,但面临一个主要障碍:从单幅图像中,深度,尤其是绝对尺度,都是欠约束的。因此,当前普遍采用的先在2D中检测然后预测3D属性的模式往往不够稳健,因为适度的范围误差可能主导3D定位,而学习到的尺度先验在相机、运动或环境发生领域转变时可能失效。为了解决这个问题,我们提出了Map-Det3D,这是一种在线多视角3D物体检测模型,它将检测直接引入从RGB重建的3D空间中。我们将一个短时间窗口映射到多个视角,并重新利用一个前馈度量3D重建模型作为我们的几何骨干,同时调整其物体感知能力。在此表示的基础上,Map-Det3D直接在度量3D空间中预测边界框,而无需广泛使用的2D到3D的提升。不同基准上的实验表明,这种设计支持强大的在线性能和无适应性的鲁棒迁移,表明为检测训练重建先验是从单目视频中实现稳定度量3D检测的可行途径。代码和模型可在 https://royyang0714.github.io/Map-Det3D 获取。
cs.CV / 81 / 2608.12185

GenFAR: A generalized representation of brain structure, derived from 49,246 multi-cohort MRIs via deep learning

GenFAR:一种基于49,246个多队列MRI通过深度学习获得的脑结构的广义表示
Bashyam, Vishnu M., Erus, Guray, Wen, Junhao, Chaudhari, Pratik, Melhem, Randa, Tirumalai, Sindhuja Govindarajan, Harman, Gareth, Fan, Yong, Masters, Colin L., Maruff, Paul, Johnson, Sterling C., Fripp, Jurgen, Tosun, Duygu, Morris, John C., Marcus, Daniel S., LaMontagne, Pamela, Benzinger, Tammie, Heckbert, Susan R., Espeland, Mark, Albert, Marilyn S., Saykin, Andrew J., Thompson, Paul M., Hohman, Timothy J., Resnick, Susan M., Bryan, R. Nick, Bilgel, Murat, An, Yang, Wolk, David A., Shen, Li, Shou, Haochang, Nasrallah, Ilya M., Davatzikos, Christos
Abstract
Deep learning models for neuroimaging have largely been developed for individual tasks, limiting knowledge transfer across applications. Here we introduce GenFAR, a modular deep learning framework that learns general, clinically informed features from brain MRIs. We trained this modular architecture on 49,246 individuals across 11 cohorts, using 17 diverse classification and regression tasks spanning cognition, clinical, diagnosis, demographics, and biomarkers. This yields aggregated, focused feature sets that capture rich, clinically- and biologically-relevant brain representations. We developed a sequential learning approach where tasks progressively build on previously learned representations. Through an analysis of 5,000 task sequences, we identified an optimal sequence length of six tasks and introduced a Donor Score metric to quantify each task's contribution to downstream performance. This analysis revealed five consistently strong donor tasks (Age, AD/MCI, MMSE, Hypertension, Hyperlipidemia) that formed the base of our sequential model. We demonstrated the utility of our learned representation, in various tasks beyond those included in the training set, to serve as the foundation for specialized secondary predictors. We further showed that using the learned feature representation can substantially increase the sample efficiency of secondary deep learning training tasks and models, as well as improve their accuracy.
Chinese Translation
神经影像学的深度学习模型主要是针对单一任务开发的,这限制了知识在不同应用之间的转移。在此,我们介绍了GenFAR,一个模块化的深度学习框架,能够从脑MRI中学习一般的、临床相关的特征。我们在11个队列中对49,246名个体进行了训练,使用了17个多样的分类和回归任务,涵盖了认知、临床、诊断、人口统计学和生物标志物。这产生了聚合的、专注的特征集,捕捉了丰富的、临床和生物学相关的脑表示。我们开发了一种顺序学习方法,其中任务逐步建立在先前学习的表示之上。通过对5,000个任务序列的分析,我们确定了最优的序列长度为六个任务,并引入了Donor Score指标来量化每个任务对下游性能的贡献。这项分析揭示了五个始终表现强劲的捐赠任务(年龄、阿尔茨海默病/轻度认知障碍、MMSE、高血压、高脂血症),这些任务构成了我们顺序模型的基础。我们展示了我们学习的表示在训练集中未包含的各种任务中的实用性,以作为专门的次级预测器的基础。我们进一步表明,使用学习到的特征表示可以显著提高次级深度学习训练任务和模型的样本效率,并改善其准确性。
cs.CV / 82 / 2608.12187

HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation

HSTGFormer:用于三维人体姿态估计的超空间-时间图变换器
Li, Ruochen, Chen, Shuang, E, Wenke, Arvin, Farshad, Atapour-Abarghouei, Amir
Abstract
Transformer-based methods have achieved strong performance in monocular 3D human pose estimation, but most existing approaches organise spatial and temporal reasoning as separate stages, which may weaken unified spatial-temporal interdependencies inherent in human motion and compress frame-level structural information before temporal modelling. In this paper, we propose HSTGFormer, a graph-enhanced Transformer framework that reformulates spatial-temporal reasoning as localised coupled graph aggregation over joint-time nodes. Specifically, HSTGFormer introduces a Hyper Spatial-Temporal Graph (HSTG), which decomposes global spatial-temporal reasoning into local spatial-temporal receptive fields around individual joint-time nodes by extending per-frame skeleton graphs into temporal neighbourhoods, thereby enabling structure-aware coupled reasoning while preserving local structural motion information. It further incorporates an Adaptive Dual-Scale Temporal Graph (ADSTG) to capture joint-specific temporal dependencies over complementary short- and long-range windows. A lightweight node-wise fusion module further adaptively integrates the two graph representations for each joint-time node. Experiments on Human3.6M and MPI-INF-3DHP show that HSTGFormer achieves strong accuracy with high computational efficiency.
Chinese Translation
基于变换器的方法在单目三维人体姿态估计中取得了优异的性能,但大多数现有方法将空间和时间推理组织为独立的阶段,这可能削弱了人类运动中固有的统一空间-时间相互依赖关系,并在时间建模之前压缩了帧级结构信息。本文提出了HSTGFormer,一种图增强的变换器框架,将空间-时间推理重新表述为对关节时间节点的局部耦合图聚合。具体而言,HSTGFormer引入了超空间-时间图(Hyper Spatial-Temporal Graph, HSTG),通过将每帧骨架图扩展到时间邻域,将全局空间-时间推理分解为围绕各个关节时间节点的局部空间-时间感受野,从而在保留局部结构运动信息的同时实现结构感知的耦合推理。它进一步结合了自适应双尺度时间图(Adaptive Dual-Scale Temporal Graph, ADSTG),以捕捉关节特定的时间依赖关系,涵盖互补的短期和长期窗口。一个轻量级的节点级融合模块进一步自适应地整合每个关节时间节点的两个图表示。在Human3.6M和MPI-INF-3DHP上的实验表明,HSTGFormer在高计算效率的同时实现了强大的准确性。
cs.CV / 83 / 2608.12196

M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation

M-Net:将光谱特征与物理场算子整合到深度学习中用于医学图像分割
Zhu, Jing, Wang, Ye, Wang, Fumin
Abstract
Purpose: Deep learning-based medical image segmentation has achieved remarkable success, yet purely data-driven approaches often fail to exploit the rich mathematical structure inherent in medical images. We investigate whether explicit mathematical inductive biases, specifically matrix spectral analysis and vector calculus operators, can enhance segmentation beyond data-driven learning alone. Methods: We propose M-Net (Math-Augmented Network), which integrates three complementary mathematical priors into U-Net: (1) continuous spectral features derived from the condition number of centered local pixel matrices, providing a differentiable measure of texture ill-conditioning; (2) physical field operators (divergence and a discrete curl-like boundary irregularity operator) computed from image gradient fields, capturing focal intensity extrema and edge non-smoothness; and (3) a Math-Attention Gate (MAG) that adaptively fuses mathematical features with CNN-extracted deep features at skip connections. Results: Experiments on three benchmarks (LiTS, KiTS, and BraTS) show that M-Net achieves Dice scores of 78.42%, 76.15%, and 83.67%, outperforming baseline U-Net by 12.37%, 3.52%, and 5.55% on liver, kidney, and brain tumor segmentation, respectively. Ablations reveal that the condition-number feature contributes a 2.14% gain over binary invertibility features, while MAG adds 1.45% over simple concatenation. Conclusion: M-Net establishes that mathematical inductive biases provide effective complementary information for medical image segmentation. The continuous condition-number feature offers superior gradient information over discrete alternatives, and MAG preserves these priors throughout the network. This work opens avenues for integrating linear algebra and vector calculus into deep architectures for medical imaging.
Chinese Translation
目的:基于深度学习的医学图像分割取得了显著成功,但纯粹的数据驱动方法往往未能充分利用医学图像中固有的丰富数学结构。我们探讨了显式的数学归纳偏差,特别是矩阵光谱分析和向量微积分算子,是否能够在数据驱动学习的基础上增强分割效果。方法:我们提出了M-Net(数学增强网络),将三种互补的数学先验整合到U-Net中:(1)从中心局部像素矩阵的条件数导出的连续光谱特征,提供了一种可微分的纹理病态度量;(2)从图像梯度场计算的物理场算子(散度和离散涡旋边界不规则算子),捕捉焦点强度极值和边缘非光滑性;(3)一个数学注意力门(Math-Attention Gate, MAG),在跳跃连接中自适应地将数学特征与CNN提取的深度特征融合。结果:在三个基准测试(LiTS、KiTS和BraTS)上的实验表明,M-Net在肝脏、肾脏和脑肿瘤分割中分别获得了78.42%、76.15%和83.67%的Dice分数,分别比基线U-Net提高了12.37%、3.52%和5.55%。消融实验显示,条件数特征比二元可逆性特征贡献了2.14%的增益,而MAG比简单拼接增加了1.45%。结论:M-Net证明了数学归纳偏差为医学图像分割提供了有效的互补信息。连续条件数特征提供了优于离散替代品的梯度信息,而MAG在整个网络中保留了这些先验。这项工作为将线性代数和向量微积分整合到医学成像的深度架构中开辟了新的途径。
cs.CV / 84 / 2608.12203

GeoFlow: Efficient Driving Video Generation via Geometry-Aligned Priors

GeoFlow:通过几何对齐先验实现高效的驾驶视频生成
Liu, Jiazheng, Li, Hang, Zhang, Jiawei, Li, Jiahe, Yu, Xiaohan, Fan, Shengyin, Zheng, Jin, Bai, Xiao
Abstract
Generative models like Diffusion Models and Flow Matching have demonstrated remarkable capabilities in synthesizing high-fidelity driving videos, but are severely constrained by high inference latency due to the requirement of extensive sampling steps. We argue that this inefficiency stems from the prevailing reliance on a standard Gaussian source distribution, where consecutive frames are initialized as independent Gaussian noise. This paradigm disregards the rich spatiotemporal correlations inherent in driving videos, compelling the model to regenerate deterministic scene structures existing in previous frames from noise, which is both computationally redundant and prone to geometric inconsistency. To address this problem, we propose GeoFlow, a novel framework designed to achieve efficient driving video generation by harnessing explicit geometric priors. Instead of sampling from standard Gaussian noise, we leverage multi-view geometry and spatially-adaptive noise injection to construct a Geometry-Aligned Prior (GAP) distribution as starting point. This initialization bridges the gap between source distribution and data distribution, yielding a significantly straighter and shorter sampling trajectory. Extensive experiments demonstrate that GeoFlow can achieve remarkable efficiency of both training and inference: merely several hours of fine-tuning on baseline models can significantly boost few-step generation quality, while fully converged training drastically reduces number of inference steps required for state-of-the-art video generation.
Chinese Translation
生成模型如扩散模型(Diffusion Models)和流匹配(Flow Matching)在合成高保真驾驶视频方面展现了显著的能力,但由于需要大量采样步骤,推理延迟严重受限。我们认为,这种低效源于对标准高斯源分布的普遍依赖,其中连续帧被初始化为独立的高斯噪声。这一范式忽视了驾驶视频固有的丰富时空相关性,迫使模型从噪声中重新生成先前帧中存在的确定性场景结构,这既计算冗余又容易导致几何不一致。为了解决这个问题,我们提出了GeoFlow,一个旨在通过利用显式几何先验实现高效驾驶视频生成的新框架。我们不再从标准高斯噪声中采样,而是利用多视角几何和空间自适应噪声注入构建几何对齐先验(Geometry-Aligned Prior, GAP)分布作为起始点。这种初始化弥合了源分布与数据分布之间的差距,产生了显著更直且更短的采样轨迹。大量实验表明,GeoFlow在训练和推理方面都能实现显著的效率:仅需在基线模型上进行几小时的微调,即可显著提升少步生成质量,而完全收敛的训练则大幅减少了进行最先进视频生成所需的推理步骤。
cs.CV / 85 / 2608.12209

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction

生成作为辅助监督:通过解耦嵌入预测在零推理开销下增强视觉理解
Guo, Zhongbin, Xie, Jiahao, Xiao, Dongling, Wang, Qianle, Lu, Ruiqi, He, Xiaomin, Sun, Wanxuan, Yang, Cheng
Abstract
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation are typically treated as divergent objectives. Existing unified frameworks often rely on discrete visual tokenization or diffusion objectives whose generative targets differ from the continuous representations consumed by visual understanding models, making direct transfer to enhance existing pretrained MLLMs non-trivial. In this work, we present GAS, a generation-guided training framework that reinterprets visual generation as auxiliary supervision for representation learning. Concretely, GAS adapts Next Embedding Prediction (NEP) as a cross-modal generation paradigm within a decoupled Mixture-of-Transformers (MoT) architecture. By maintaining a shared lower trunk and parallel upper layers, GAS lets generation losses enrich the shared visual pathway with finer spatial precision and stronger visual retention while shielding the upper understanding layers from direct generation gradients. To maximize this synergy, we further construct highly correlated generation tasks that demand deep cognitive grounding rather than generic synthesis alone. Across model scales and training stages, GAS improves aggregate multimodal understanding, with its most reliable gains on perception and spatial comprehension. Crucially, because the auxiliary generation branch is discarded after training, these gains incur zero inference overhead. Extensive controlled comparisons and representation-level analyses further clarify when and why generation-guided training benefits understanding, and demonstrate the feasibility of generation-guided training as a practical route to stronger multimodal understanding.
Chinese Translation
尽管多模态大型语言模型(MLLMs)取得了显著进展,但视觉理解和生成通常被视为不同的目标。现有的统一框架往往依赖于离散的视觉标记化或扩散目标,其生成目标与视觉理解模型所消耗的连续表示不同,使得直接转移以增强现有的预训练MLLMs变得复杂。在本研究中,我们提出了GAS,一个生成引导的训练框架,将视觉生成重新解释为表示学习的辅助监督。具体而言,GAS在解耦的混合变换器(Mixture-of-Transformers, MoT)架构中,将下一嵌入预测(Next Embedding Prediction, NEP)适配为跨模态生成范式。通过维持共享的下部主干和并行的上层,GAS使生成损失以更精细的空间精度和更强的视觉保留丰富共享的视觉通路,同时保护上层理解层免受直接生成梯度的影响。为了最大化这种协同效应,我们进一步构建高度相关的生成任务,这些任务需要深层次的认知基础,而不仅仅是通用的合成。在不同的模型规模和训练阶段,GAS改善了整体的多模态理解,其在感知和空间理解方面的增益最为可靠。关键的是,由于辅助生成分支在训练后被丢弃,这些增益不会产生任何推理开销。广泛的受控比较和表示级别分析进一步阐明了生成引导训练何时以及为何有利于理解,并展示了生成引导训练作为增强多模态理解的可行路径。
cs.CV / 86 / 2608.12220

SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward

SCOUT:通过结构化思维链和多目标过程奖励解锁增强的空间推理能力
Zhou, Zile, Yuan, Huining, Zhang, Weichen, Chen, Xinlei, Zhang, Xiao-ping
Abstract
Existing Vision-Language Models (VLMs) exhibits a critical bottleneck in robust spatial reasoning. Recent reinforcement learning (RL) methods aim to close this gap with verifiable outcomes, yet they suffer from poor credit assignment across intermediate reasoning steps. Concurrently, structured reasoning approaches overlook the critical depth perception necessary for comprehensive 3D understanding. To address these challenges, we propose SCOUT (Structured Chain-Of-Thought Utilizing Process-Supervised RL Training). Specifically, we design a structured Chain-of-Thought (CoT) framework that explicitly models 3D environmental perception to ensure robust spatial understanding and reasoning. Furthermore, we introduce a novel RL algorithm featuring multi-objective process rewards and a tailored advantage estimation method, facilitating fine-grained credit assignment across distinct segments of the reasoning trajectory. To support our framework, we develop SCOUT-24k, a structured spatial reasoning CoT dataset synthesized through a customized pipeline. Extensive evaluations demonstrate that SCOUT-3B improves upon baseline models by 16.85% and 6.3% on general spatial benchmarks and complex spatial reasoning tasks respectively. Notably, our larger SCOUT-7B even outperforms GPT-4o by a margin of 4.28%. Moreover, despite being trained exclusively on single image, SCOUT-7B exhibits robust out-of-domain generalization to multi-image and video scenarios. These empirical results render SCOUT as a critical step towards next generation of spatially-aware VLMs.
Chinese Translation
现有的视觉-语言模型(VLMs)在稳健的空间推理方面存在严重瓶颈。近期的强化学习(RL)方法旨在缩小这一差距,并提供可验证的结果,但它们在中间推理步骤中的信用分配表现不佳。同时,结构化推理方法忽视了全面理解三维(3D)环境所需的深度感知。为了解决这些挑战,我们提出了SCOUT(结构化思维链利用过程监督的RL训练)。具体而言,我们设计了一个结构化的思维链(Chain-of-Thought, CoT)框架,明确建模3D环境感知,以确保稳健的空间理解和推理。此外,我们引入了一种新颖的RL算法,具有多目标过程奖励和定制的优势估计方法,促进了在推理轨迹不同部分之间的细粒度信用分配。为了支持我们的框架,我们开发了SCOUT-24k,这是通过定制管道合成的结构化空间推理CoT数据集。广泛的评估表明,SCOUT-3B在一般空间基准和复杂空间推理任务上分别比基线模型提高了16.85%和6.3%。值得注意的是,我们更大的SCOUT-7B甚至比GPT-4o超出4.28%。此外,尽管仅在单幅图像上训练,SCOUT-7B在多图像和视频场景中表现出稳健的域外泛化能力。这些实证结果使SCOUT成为朝着下一代空间感知VLMs迈出的重要一步。
cs.CV / 87 / 2608.12230

Few-Shot Ordinal Learning for Day-Wise Freshness Estimation with Hyperspectral Fish Images

基于少量样本的序数学习用于日常新鲜度估计:以高光谱鱼类图像为例
Alam, Kazi Nabiul, Zadeh, Pooneh Bagheri, Sheikh-Akbari, Akbar
Abstract
Non-destructive food quality assessment has increasingly benefited from hyperspectral imaging (HSI), which captures spectral signatures linked to biochemical changes during storage. Estimating day-wise freshness, however, remains challenging owing to strong inter-fillet variability and scarce labelled data per product. All existing deep learning approaches for HSI-based freshness prediction operate under full supervision, requiring densely annotated training sets that are costly to obtain at the individual-product level. We introduce, to the best of our knowledge, the first few-shot learning framework for HSI-based food quality estimation. Each fillet defines a distinct episodic task, and a CORAL-style ordinal prediction head captures the ranked nature of freshness progression through cumulative threshold modelling. Biologically grounded monotonicity and embedding smoothness constraints further guide predictions toward plausible trajectories. On a 16-day salmon HSI dataset under a strict unseen-fillet protocol, our method achieves a mean absolute error of 1.58 days and 2-day accuracy of 72.3% with only three labelled days per fillet, substantially outperforming scalar regression and label-distribution baselines under an identical unseen-fillet protocol.
Chinese Translation
非破坏性食品质量评估越来越依赖于高光谱成像(HSI),该技术捕捉与储存过程中生化变化相关的光谱特征。然而,由于鱼片之间的强烈变异性和每种产品标注数据的稀缺,日常新鲜度的估计仍然具有挑战性。现有的基于HSI的新鲜度预测深度学习方法均在全监督的条件下运行,要求密集标注的训练集,而这些训练集在单个产品层面上获取成本高昂。我们首次提出了一种基于HSI的食品质量估计的少量样本学习框架。每个鱼片定义了一个独特的情景任务,而CORAL风格的序数预测头通过累积阈值建模捕捉新鲜度进展的排名特性。生物学基础的单调性和嵌入平滑性约束进一步引导预测朝向合理的轨迹。在一个为期16天的鲑鱼HSI数据集上,在严格的未见鱼片协议下,我们的方法在每个鱼片仅有三个标注日的情况下,实现了1.58天的平均绝对误差和72.3%的2天准确率,显著优于标量回归和标签分布基线模型在相同未见鱼片协议下的表现。
cs.CV / 88 / 2608.12232

ScaleVid: Geometry-Aware Video Object Scaling with Mesh-Free Inference

ScaleVid:基于几何的无网格推理视频对象缩放
Huang, Youze, Ruan, Penghui, Zi, Bojia, Qi, Xianbiao, Zhao, Shihao, Xiao, Rong
Abstract
Geometry-aware video object scaling aims to anisotropically resize the object along object-centric axes while preserving geometric plausibility, temporal coherence, and background consistency. Existing text-guided methods mainly operate in the 2D image plane, while depth-guided approaches provide coarse control and mesh-based methods require costly 3D reconstruction. We present a progressive two-stage training framework that decouples geometry-aware foreground transformation from background preservation and realistic video composition, without mesh-pixel alignment and explicit 3D reconstruction at inference. In both stages, geometrically perturbed pseudo-sources are constructed from real videos, while the original complete videos are retained as reconstruction targets. The first stage uses planar transformations to learn robust foreground-background composition, whereas the second introduces object-centric 3D deformation guidance for geometry-aware scaling. This pseudo-source reconstruction formulation enables real-video synthesis without paired real-world scaling targets. We construct complementary paired-geometry and real-background benchmarks and further evaluate on in-the-wild videos. Extensive experiments demonstrate superior geometric consistency, foreground fidelity, and background preservation, together with faster and more practical inference than methods requiring explicit 3D reconstruction.
Chinese Translation
基于几何的视频对象缩放旨在沿着以对象为中心的轴向非均匀地调整对象的大小,同时保持几何合理性、时间一致性和背景一致性。现有的文本引导方法主要在二维图像平面上操作,而深度引导方法提供粗略的控制,基于网格的方法则需要昂贵的三维重建。我们提出了一种渐进式的两阶段训练框架,将基于几何的前景变换与背景保持和真实视频合成解耦,在推理时无需网格像素对齐和显式的三维重建。在两个阶段中,几何扰动的伪源是从真实视频构建的,而原始完整视频则作为重建目标保留。第一阶段使用平面变换来学习稳健的前景-背景合成,而第二阶段则引入以对象为中心的三维变形引导以实现基于几何的缩放。这种伪源重建形式使得在没有配对真实世界缩放目标的情况下进行真实视频合成成为可能。我们构建了互补的配对几何和真实背景基准,并进一步在野外视频上进行评估。大量实验表明,与需要显式三维重建的方法相比,我们的方法在几何一致性、前景保真度和背景保持方面表现出更优越的性能,同时推理速度更快且更具实用性。
cs.CV / 89 / 2608.12239

HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression

HAMP-LIC:一种基于Hessian的混合精度后训练量化方法用于学习图像压缩
Zhang, Yuefeng
Abstract
Use this plain-text version for the arXiv abstract field: Learned image compression (LIC) models achieve strong rate-distortion performance but are hindered by high computational complexity and encoding-decoding mismatches across heterogeneous hardware platforms. Uniform fixed-precision quantization alleviates these issues but suffers severe quality degradation at low bit widths because it ignores differences in the quantization sensitivities of individual layers. To enable efficient and accurate low-bit deployment of pretrained LIC models, we propose HAMP-LIC, a Hessian-aware mixed-precision post-training quantization (PTQ) framework with a four-stage optimization strategy. First, block-wise sensitivity is estimated from the Hessian trace to capture second-order importance. Second, a task-aware refinement module adjusts these sensitivities by jointly considering quantization distortion and rate-distortion performance. Third, guided by the refined sensitivity profile, bit widths are allocated under a global model-size constraint to balance efficiency and reconstruction quality. Finally, block-wise reconstruction using a small calibration set further suppresses quantization error. Experiments on representative LIC models, including Minnen2018 and Cheng2020, demonstrate that HAMP-LIC achieves up to 4.85x model compression with as little as 0.59% BD-rate loss. It consistently outperforms existing fixed- and mixed-precision PTQ methods across multiple datasets while completely eliminating cross-platform encoding-decoding errors.
Chinese Translation
学习图像压缩(LIC)模型在速率-失真性能上表现出色,但由于计算复杂度高以及在异构硬件平台上的编码-解码不匹配,导致其应用受到限制。统一的固定精度量化虽然缓解了这些问题,但在低比特宽度下会严重降低质量,因为它忽略了各层之间量化敏感度的差异。为了实现预训练LIC模型的高效且准确的低比特部署,我们提出了HAMP-LIC,一种基于Hessian的混合精度后训练量化(PTQ)框架,采用四阶段优化策略。首先,从Hessian迹中估计块级敏感度,以捕捉二阶重要性。其次,任务感知的细化模块通过共同考虑量化失真和速率-失真性能来调整这些敏感度。第三,在全球模型大小约束下,根据细化的敏感度配置比特宽度,以平衡效率和重建质量。最后,使用小型校准集进行块级重建,进一步抑制量化误差。在代表性的LIC模型(包括Minnen2018和Cheng2020)上的实验表明,HAMP-LIC实现了高达4.85倍的模型压缩,同时仅损失0.59%的BD-rate。它在多个数据集上始终优于现有的固定和混合精度PTQ方法,同时完全消除了跨平台的编码-解码错误。
cs.CV / 90 / 2608.12252

Automated Borehole Core Analysis with Report-Derived Weak Labels and Supervised Crack Segmentation

基于报告衍生弱标签和监督裂缝分割的自动化钻孔岩心分析
Imdad, Usama, Khan, Ali, Lu, Luke, Khalid, Zubair, Mahmood, Arif
Abstract
Borehole archives commonly contain core tray photographs and corresponding digital log reports, but no native pixel-level crack annotations. We investigate two complementary approaches for extracting defect-spacing information from these archives. First, structured spacing categories recovered from the report text layer provide weak interval-level labels for classification. A DINO encoder trained on unlabeled core crops supplies domain-specific representations, and a manually verified subset is used to identify label inconsistencies. Second, we manually annotate 5,087 extracted core-row images and evaluate fully supervised crack-segmentation models. Our gated U-Net combines PiDiNet edge maps with Mask R-CNN masks through a learned spatial gating mechanism. This configuration achieves an F1 score of 0.860 and a crack-class IoU of 0.754, the highest result among the evaluated segmentation configurations. Deterministic post-processing converts predicted crack locations into defect-spacing categories. Separate rule-based branches estimate core-relative bedding angles and lithological color descriptors; their predictions agree with log-report references on 75.4% and 84.7% of 1,200 evaluated images, respectively. Because these references are extracted from existing reports, the reported values measure agreement with recorded geological observations rather than independent physical accuracy. The resulting framework combines report-derived weak supervision for spacing classification with fully supervised segmentation for image-based crack localization.
Chinese Translation
钻孔档案通常包含岩心托盘照片和相应的数字日志报告,但缺乏原生的像素级裂缝注释。我们研究了两种互补的方法,从这些档案中提取缺陷间距信息。首先,从报告文本层恢复的结构化间距类别为分类提供了弱的区间级标签。一个在未标记岩心裁剪图像上训练的 DINO 编码器提供了特定领域的表示,并使用一个经过人工验证的子集来识别标签不一致性。其次,我们手动注释了 5,087 张提取的岩心行图像,并评估了完全监督的裂缝分割模型。我们的门控 U-Net 通过学习的空间门控机制将 PiDiNet 边缘图与 Mask R-CNN 掩膜相结合。该配置在评估的分割配置中实现了 0.860 的 F1 分数和 0.754 的裂缝类别 IoU,达到了最高结果。确定性后处理将预测的裂缝位置转换为缺陷间距类别。独立的基于规则的分支估计岩心相对的层理角度和岩石颜色描述符;它们的预测分别与 1,200 张评估图像的日志报告参考一致率为 75.4% 和 84.7%。由于这些参考是从现有报告中提取的,报告的值衡量的是与记录的地质观察的一致性,而不是独立的物理准确性。最终的框架结合了基于报告的弱监督用于间距分类,以及完全监督的分割用于基于图像的裂缝定位。
cs.CV / 91 / 2608.12262

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams

Diagram-MMU:一个用于科学图表的多模态基准
Bo, Weihao, Zhang, Shan, Sun, Yanpeng, Liu, Jie, Yao, Yongke, Du, Jinhao, He, Wei, Zou, Kai, Li, Zechao, Wang, Jingdong
Abstract
Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for scientific writing and collaboration. One important feature in Prism is turning scientific diagrams directly into LaTeX TikZ code. In this paper, we build a benchmark, Diagram-MMU, a multi-modal benchmark designed to assess MLLMs' ability for scientific diagram parsing and understanding. Diagram-MMU features 3.7k curated diagrams and 18.3k human-validated questions across six domains. It evaluates MLLMs on three tasks common in vibe writing workspaces: diagram-to-code parsing, diagram-to-code editing, and diagram question answering, alongside agentic settings per task. The evaluation of 12 MLLMs reveals that diagram-to-code tasks are more challenging than diagram question answering: models can reason well over diagrams but struggle to parse and edit them, underscoring the need for methods to enhance MLLMs' capability in diagram-to-code generation. Under agentic settings, most models improve parsing and editing performance but degrade on question answering, while Claude-4.6 Opus consistently improves across all three tasks. Project Page: https://vi-ocean.github.io/projects/diagram-mmu.
Chinese Translation
多模态大型语言模型(MLLMs)在科学写作和协作方面的能力不断增强。例如,OpenAI Prism 是一个用于科学写作和协作的免费工作空间。Prism 的一个重要功能是将科学图表直接转换为 LaTeX TikZ 代码。本文构建了一个基准,Diagram-MMU,这是一个旨在评估 MLLMs 在科学图表解析和理解能力的多模态基准。Diagram-MMU 包含 3.7k 精心挑选的图表和 18.3k 人工验证的问题,涵盖六个领域。它在三个常见的写作工作空间任务上评估 MLLMs:图表到代码的解析、图表到代码的编辑和图表问答,同时针对每个任务设置代理环境。对 12 个 MLLMs 的评估表明,图表到代码的任务比图表问答更具挑战性:模型能够很好地推理图表,但在解析和编辑方面存在困难,这突显了增强 MLLMs 在图表到代码生成能力的方法的必要性。在代理环境下,大多数模型在解析和编辑性能上有所提升,但在问答方面表现下降,而 Claude-4.6 Opus 在所有三个任务上均持续改善。项目页面:https://vi-ocean.github.io/projects/diagram-mmu。
cs.CV / 92 / 2608.12274

A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery

一种邻域注意力变换网络用于增强左前降支动脉的三维分割
Sultan, Rafi Ibn, Li, Chengyin, Demetriou, Yiannos, Ghanem, Ahmed I., Kim, Joshua P., Cunningham, Justine, Bagher-Ebadian, Hassan, Zhu, Dongxiao, Thind, Kundan S.
Abstract
Background: Accurate segmentation of the Left Anterior Descending (LAD) artery in 3D free-breathing, non-contrast CT is critical for cardiac dose sparing in thoracic radiotherapy. The LAD is extremely small, has poor soft-tissue contrast, and varies substantially across patients; even manual contours show limited inter-observer agreement, underscoring the ambiguity of the vessel boundaries. Purpose: To develop a transformer-based framework that improves LAD delineation in low-contrast, imbalanced CT through local-global context modeling and uncertainty-guided optimization. Methods: We propose NA-UNETR, a 3D transformer-based segmentation model whose Neighborhood Attention (NA) and Dilated NA (DiNA) blocks jointly capture fine structural detail and long-range context. Given the scarcity of annotated LAD data, the model is pretrained on 1,000 CTA volumes of general coronary anatomy and fine-tuned with LoRA-based parameter-efficient adaptation on 20 free-breathing institutional CT scans. A composite Dice-Focal and Hausdorff loss, dynamically balanced via homoscedastic uncertainty, improves overlap and boundary accuracy. Results: NA-UNETR reached 45.64% Dice, 38.16 mm HD95, and 10.01 mm ASD, improving Dice by 3.10 percentage points over nnU-Net and reducing HD95 by 2.96 mm relative to Swin UNETR, with the strongest boundary accuracy among all models and improved centerline stability. On ImageCAS it achieved 79.49% Dice, 8.89 mm HD95, and 1.02 mm ASD. Ablations confirmed that residual blocks, variable kernels, and uncertainty-weighted loss each contributed. Conclusions: NA-UNETR balances local precision and global context for thin, low-contrast LAD structures, offering a computationally efficient framework for substructure-level cardiac segmentation in radiotherapy planning.
Chinese Translation
背景:在三维自由呼吸、无对比剂的CT中,准确分割左前降支(LAD)动脉对于胸部放射治疗中的心脏剂量节省至关重要。LAD动脉极小,软组织对比度差,并且在患者之间差异显著;即使是手动轮廓也显示出有限的观察者间一致性,突显了血管边界的模糊性。目的:开发一个基于变换器的框架,通过局部-全局上下文建模和不确定性引导优化来改善低对比度、不平衡CT中的LAD描绘。方法:我们提出了NA-UNETR,这是一种基于3D变换器的分割模型,其邻域注意力(NA)和扩张NA(DiNA)模块共同捕捉细微的结构细节和长距离上下文。鉴于标注的LAD数据稀缺,该模型在1,000个一般冠状解剖的CTA体积上进行预训练,并在20个自由呼吸的机构CT扫描上通过基于LoRA的参数高效适应进行微调。通过同质不确定性动态平衡的复合Dice-Focal和Hausdorff损失,改善了重叠和边界准确性。结果:NA-UNETR达到了45.64%的Dice,38.16 mm的HD95和10.01 mm的ASD,相比于nnU-Net提高了3.10个百分点的Dice,并相对于Swin UNETR减少了2.96 mm的HD95,所有模型中具有最强的边界准确性和改善的中心线稳定性。在ImageCAS上,它达到了79.49%的Dice,8.89 mm的HD95和1.02 mm的ASD。消融实验确认了残差块、可变核和不确定性加权损失各自的贡献。结论:NA-UNETR在薄而低对比度的LAD结构中平衡了局部精度和全局上下文,为放射治疗规划中的亚结构级心脏分割提供了计算高效的框架。
cs.CV / 93 / 2608.12276

XYZFlow:Scaling Multi dimensional Shortcut Flows for Efficient Generative Modeling

XYZFlow:为高效生成建模扩展多维快捷流
Liu, Jinxiu, Liu, Xuanming, Mei, Kangfu, Wen, Yandong, Liu, Weiyang
Abstract
High-fidelity image generation faces a trade-off between speed and quality. Diffusion models produce strong visuals but require costly iterative sampling. Existing efficient methods mainly distill pretrained models into few-step samplers, a challenging process that depends heavily on teacher-model quality. In this paper, we introduce XYZFlow, a framework that rethinks efficient generation through multidimensional scaling of flow matching. Unlike single-step mappings, XYZFlow enhances expressivity by making probability paths more identifiable and learnable through structured multidimensional conditioning. We view autoregressive modeling as implicit flow straightening, where richer context reduces trajectory ambiguity. XYZFlow realizes this idea through two orthogonal dimensions: temporal scaling, which uses non-Markovian conditioning on the full denoising history; and spatial scaling, enabled by Next Shortcut Prediction, which sequentially generates patches using preceding patches' denoising trajectories as priors. Experiments show that XYZFlow achieves state-of-the-art performance, with 7.2-8.5X teacher speedups and competitive FID, while Next Shortcut Prediction delivers superior quality-latency trade-offs over model scaling or step reduction.
Chinese Translation
高保真图像生成面临速度与质量之间的权衡。扩散模型能够生成高质量的视觉效果,但需要昂贵的迭代采样。现有的高效方法主要将预训练模型提炼为少步采样器,这一过程具有挑战性,并且在很大程度上依赖于教师模型的质量。本文介绍了XYZFlow,一个通过流匹配的多维缩放重新思考高效生成的框架。与单步映射不同,XYZFlow通过结构化的多维条件增强了表达能力,使概率路径更具可识别性和可学习性。我们将自回归建模视为隐式流的拉直,其中更丰富的上下文减少了轨迹的模糊性。XYZFlow通过两个正交维度实现这一理念:时间缩放,利用对完整去噪历史的非马尔可夫条件;以及空间缩放,由下一快捷预测(Next Shortcut Prediction)实现,该方法使用前面补丁的去噪轨迹作为先验,顺序生成补丁。实验表明,XYZFlow实现了最先进的性能,教师速度提升达到7.2-8.5倍,并且在FID上具有竞争力,而下一快捷预测在模型缩放或步数减少方面提供了更优的质量-延迟权衡。
cs.CV / 94 / 2608.12279

Curvature-Aware Zeroth-Order Optimization for Memory-Efficient Test-Time Adaptation

考虑曲率的零阶优化用于内存高效的测试时适应
Zhang, Junming, Yin, Shuyu, Liu, Peilin, Ying, Rendong, Wen, Fei
Abstract
Test-time adaptation (TTA) aims to enhance the cross-domain performance of pre-trained models by adapting to unlabeled test data. While most existing TTA methods rely on backpropagation (BP) for finetuning, BP-free methods such as zeroth-order (ZO) methods are more desired in practical on-device scenarios. ZO methods rely only on forward computation, which can largely reduce the complexity and memory overhead of on-device deployment. However, ZO methods suffer from much higher variance compared with first-order methods in estimating the gradient. To address this, we propose an improved ZO method to substantially boost the performance of ZO optimization based TTA. First, we provide an observation to reveal the persistent low-rank Hessian structure of the loss during the adaptation process. Based on this insight, we then propose a loss-landscape curvature-aware zeroth-order (CAZO) method, which leverages a sliding-average estimation of the diagonal Hessian to construct a covariance matrix for anisotropic perturbation sampling. CAZO operates by freezing pretrained weights and optimizing minimal adapter parameters via forward-only passes based gradient estimation, which can substantially reduce the memory overhead compared to BP-based methods. Extensive experiments demonstrate that CAZO significantly outperforms existing TTA methods, achieving state-of-the-art performance while maintaining an excellent balance between accuracy and memory efficiency. Code is available at https://github.com/Hollyming/CAZO.
Chinese Translation
测试时适应(TTA)旨在通过适应未标记的测试数据来增强预训练模型的跨域性能。虽然大多数现有的 TTA 方法依赖于反向传播(BP)进行微调,但在实际的设备场景中,更希望使用无反向传播的方法,例如零阶(ZO)方法。ZO 方法仅依赖前向计算,这可以大大减少设备部署的复杂性和内存开销。然而,与一阶方法相比,ZO 方法在估计梯度时面临更高的方差。为了解决这个问题,我们提出了一种改进的 ZO 方法,以显著提升基于 ZO 优化的 TTA 性能。首先,我们提供了一种观察,揭示了适应过程中损失的持续低秩 Hessian 结构。基于这一见解,我们提出了一种考虑损失景观曲率的零阶(CAZO)方法,该方法利用对角 Hessian 的滑动平均估计来构建用于各向异性扰动采样的协方差矩阵。CAZO 通过冻结预训练权重并通过仅前向传递的梯度估计优化最小适配器参数,从而显著减少与基于 BP 的方法相比的内存开销。大量实验表明,CAZO 显著优于现有的 TTA 方法,实现了最先进的性能,同时在准确性和内存效率之间保持了良好的平衡。代码可在 https://github.com/Hollyming/CAZO 获取。
cs.CV / 95 / 2608.12290

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

超越试错:图像到视频的自主优化
Tyagi, Aman, Boinpally, Hemanth, Chen, Jonathan, Gebert, Douglas, Hickson, Steven
Abstract
Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows. Their inherent stochasticity causes minor variations in textual prompts or hyperparameters to yield drastically different outputs often necessitating inefficient, brute-force trial-and-error processes. To address these limitations, we introduce the ``Agentic Self-Improvement" framework, which reframes video synthesis into a closed-loop, goal-directed optimization. Our framework systematically navigates the generation parameter space using a novel two-stage approach. In the first stage, an iterative prompt optimization loop uses a multimodal Large Language Model (mLLM) to refine the input prompt. This refinement implements two automated evaluations: Davidsonian Scene Graph (DSG) queries ensure semantic adherence, and Common Mistake Questions (CMQ) for artifact detection. At the second stage, we use Bayesian optimization to efficiently co-optimize stochastic seeds and CFG scales. This search is guided by a suite of quality metrics, including the novel Video-Text Adherence (VTA) score derived from the DSG and CMQ evaluations. Our framework significantly outperforms unguided search methods: in human preference studies, videos generated via our agentic approach were strongly preferred over baseline outputs, achieving win rates up to 69\%. This work provides a practical and extensible methodology for enhancing the predictability and control of state-of-the-art video generation models, moving the field beyond speculative curiosities toward reliable, production-ready tools.
Chinese Translation
现代黑箱图像到视频(Image-to-Video, I2V)模型在自动内容创作方面提供了强大的能力,但其缺乏细粒度控制和可靠性在专业工作流程中带来了显著挑战。其固有的随机性导致文本提示或超参数的微小变化产生截然不同的输出,通常需要低效的、粗暴的试错过程。为了解决这些局限性,我们提出了“自主自我改进”(Agentic Self-Improvement)框架,该框架将视频合成重新构建为一个闭环的、目标导向的优化过程。我们的框架采用一种新颖的两阶段方法系统地探索生成参数空间。在第一阶段,一个迭代的提示优化循环使用多模态大型语言模型(Multimodal Large Language Model, mLLM)来细化输入提示。该细化过程实施了两种自动评估:Davidsonian场景图(Davidsonian Scene Graph, DSG)查询确保语义一致性,以及常见错误问题(Common Mistake Questions, CMQ)用于伪影检测。在第二阶段,我们使用贝叶斯优化有效地共同优化随机种子和CFG尺度。该搜索由一系列质量指标引导,包括从DSG和CMQ评估中得出的新颖视频-文本一致性(Video-Text Adherence, VTA)分数。我们的框架显著优于无指导搜索方法:在人工偏好研究中,通过我们的自主方法生成的视频被强烈偏好于基线输出,胜率高达69%。这项工作提供了一种实用且可扩展的方法论,以增强最先进视频生成模型的可预测性和控制性,推动该领域从投机性的好奇心走向可靠的、可用于生产的工具。
cs.CV / 96 / 2608.12299

Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations

可解释计算机视觉中的类别激活映射:以方法为中心的卷积神经网络、变换器和基础模型时代视觉解释的综述
Eshghi, AmirHossein, Saadatfar, Hamid, Hoseini, Seyyed Ali, Eshghi, AmirMohsen, Bigdel, Siavash Arjomand
Abstract
Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts internal model evidence into a heatmap that highlights the image regions, convolutional channels, tokens, or patches that support a target class or concept. Since the first CAM formulation in 2016, the field has moved far beyond global-average-pooled CNN classifiers. CAM-style methods now include gradient-based post-hoc explanations, gradient-free score and ablation methods, high-resolution upscaling, weakly supervised localization and segmentation, transformer token attribution, causal and debiasing methods, and foundation-model-era approaches that use CLIP, DINO, SAM, or feature-distribution comparisons. This review synthesizes a strict corpus of 57 method-centered papers published from 2016 onward. The paper develops a taxonomy that separates methods by attribution mechanism, architectural dependence, and evaluation objective. It then reviews gradient-based CAMs, recent and hybrid CAM-style methods, and model-based or architecture-aware methods. Across the corpus, the main trend is clear: the field is shifting from explaining one class score in one low-resolution CNN layer toward comparative, multi-layer, probabilistic, token-aware, and foundation-model-aware explanations. At the same time, evaluation remains fragmented. Faithfulness, localization, robustness, computational cost, and human trust are often measured with different protocols. The review therefore emphasizes not only what each method contributes, but also which gap it leaves open and which later methods attempt to close that gap.
Chinese Translation
类别激活映射(Class Activation Mapping, CAM)是可解释人工智能中最广泛使用的视觉解释方法之一。其目的直观明了:将内部模型证据转换为热图,突出显示支持目标类别或概念的图像区域、卷积通道、标记或补丁。自2016年首次提出CAM的公式以来,该领域的发展已远远超出全局平均池化的卷积神经网络(CNN)分类器。CAM风格的方法现在包括基于梯度的后验解释、无梯度的评分和消融方法、高分辨率上采样、弱监督定位和分割、变换器标记归因、因果和去偏方法,以及使用CLIP、DINO、SAM或特征分布比较的基础模型时代方法。本综述综合了自2016年以来发表的57篇以方法为中心的严格文献。本文发展了一种分类法,根据归因机制、架构依赖性和评估目标对方法进行区分。接着,回顾了基于梯度的CAM、近期和混合CAM风格的方法,以及基于模型或架构感知的方法。在这些文献中,主要趋势显而易见:该领域正从解释单一类别得分的低分辨率CNN层,向比较的、多层的、概率的、标记感知的和基础模型感知的解释转变。同时,评估仍然存在碎片化的问题。忠实性、定位、鲁棒性、计算成本和人类信任通常使用不同的协议进行测量。因此,本综述不仅强调每种方法的贡献,还指出它们留下的空白,以及后续方法试图填补这些空白。
cs.CV / 97 / 2608.12308

DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation

DreamFly:用于空中视觉语言导航的因果记忆与递归视野扩散规划
Deng, Yan, Xu, Fei
Abstract
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observability. Although recent VLA models offer a promising perception-to-action paradigm, adapting them to aerial navigation remains challenging due to limited historical context, short planning horizons, and unreliable implicit termination. To address these challenges, we propose DreamFly, a diffusion-based aerial VLN framework built on Dream-VLA. DreamFly introduces a causally aligned historical memory that augments the current visual representation using only observations preceding the current decision step, enabling temporal reasoning without future information leakage. We further formulate navigation as receding-horizon diffusion planning, where the policy predicts a $K$-step action chunk but executes only the first action before replanning. This plan-$K$, execute-one strategy uses future actions as auxiliary planning targets while preserving closed-loop visual feedback. Finally, LiteStop estimates the stop probability directly from action logits at the initial all-mask state, decoupling explicit termination from action generation. Experiments on the OpenFly benchmark demonstrate consistent improvements in seen and unseen environments. DreamFly achieves 32.04%/29.46% SR and 28.22%/23.54% SPL on the test-seen/test-unseen splits, respectively, outperforming all compared methods on both metrics while attaining the lowest navigation error. These results demonstrate the effectiveness of jointly modeling historical context, future action structure, and explicit termination for aerial VLN.
Chinese Translation
空中视觉语言导航(VLN)要求具身代理在部分可观测性下整合视觉证据、规划未来行动,并确定何时达到导航目标。尽管近期的视觉语言代理(VLA)模型提供了一个有前景的感知到行动的范式,但由于历史上下文有限、规划视野短暂以及隐式终止不可靠,将其适应于空中导航仍然具有挑战性。为了解决这些问题,我们提出了DreamFly,一个基于扩散的空中VLN框架,建立在Dream-VLA之上。DreamFly引入了一种因果对齐的历史记忆,仅使用当前决策步骤之前的观察来增强当前的视觉表示,从而实现时间推理而不泄露未来信息。我们进一步将导航形式化为递归视野扩散规划,其中策略预测一个$K$步的动作块,但仅在重新规划之前执行第一个动作。这种计划-$K$、执行-一个的策略利用未来动作作为辅助规划目标,同时保持闭环视觉反馈。最后,LiteStop直接从初始全掩码状态的动作logits中估计停止概率,将显式终止与动作生成解耦。OpenFly基准上的实验表明,在已见和未见环境中均有一致的改进。DreamFly在测试集的已见/未见拆分中分别达到了32.04%/29.46%的成功率(SR)和28.22%/23.54%的成功路径长度(SPL),在这两个指标上均优于所有比较方法,同时实现了最低的导航误差。这些结果证明了联合建模历史上下文、未来动作结构和显式终止在空中VLN中的有效性。
cs.CV / 98 / 2608.12313

AVA-Encoder: Towards Agent-Native Video Representation Learning

AVA-编码器:面向代理原生视频表示学习
Li, Chuyue, Yu, Jinpeng, Wang, Haozhe, Xueyun, Tian, Zhang, Zhijing, Li, Bingnan, Gu, Shuqi, Ren, Kan, Liu, Jiaming, Hua, Ruihua
Abstract
Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipulation. To address the challenge, we propose the Agentic Video Auto-Encoder (AVA-Encoder), a framework for learning agent-native video representations via agentic auto-encoding. AVA-Encoder transforms a video into a knowledge graph (KG) representation and then reconstructs it back into video. Its hierarchy and state nodes store structured text, while a linked asset layer holds generated images, audio, and video. Typed edges preserve the relations between these text descriptions and assets in a form that agents can easily understand, query, and edit. The video reconstruction differences drive a textual-gradient optimization framework, which expresses evaluation feedback as natural-language update directions for Data-Independent Encoding Policy Pseudo-Training in the outer loop and optional Data-Dependent KG Representation Refinement in the test-time inner loop. Extensive experiments show that AVA-Encoder improves by 20.7 percentage points over the strongest external baseline. In the controlled policy-only setting, its pseudo-trained shot-level Agentic Video Encoder policy also outperforms a carefully human-tuned policy while using 74.3% fewer system-prompt tokens. We release the complete AVA-Encoder framework, a reliable agentic video reconstruction benchmark, and the first dataset of high-quality film KG representations.
Chinese Translation
创造性代理仍然缺乏有效的方法从高质量的人类电影中学习,这限制了它们制作电影级视频的能力。一个关键挑战是缺乏一种结构化的视频表示,这种表示既忠实于电影内容,又可直接用于代理推理和操作。为了解决这一挑战,我们提出了代理视频自编码器(Agentic Video Auto-Encoder,AVA-Encoder),这是一个通过代理自编码学习代理原生视频表示的框架。AVA-Encoder将视频转换为知识图(Knowledge Graph,KG)表示,然后再将其重构回视频。其层次结构和状态节点存储结构化文本,而链接资产层则保存生成的图像、音频和视频。类型化边缘保持这些文本描述与资产之间的关系,以代理可以轻松理解、查询和编辑的形式。视频重构差异驱动了一个文本梯度优化框架,该框架将评估反馈表达为自然语言更新方向,用于外循环中的数据无关编码策略伪训练(Data-Independent Encoding Policy Pseudo-Training)和测试时内循环中的可选数据相关KG表示细化(Data-Dependent KG Representation Refinement)。大量实验表明,AVA-Encoder在最强外部基线之上提高了20.7个百分点。在受控的仅策略设置中,其伪训练的镜头级代理视频编码器策略在使用74.3%更少的系统提示令牌的情况下,也超越了经过精心调优的人类策略。我们发布了完整的AVA-Encoder框架,一个可靠的代理视频重构基准,以及第一个高质量电影KG表示的数据集。
cs.CV / 99 / 2608.12314

StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

StateFlow:构建、演变和访问用于预可视化的3D世界状态
Yin, Yuyang, Li, Zixiang, Deng, Longxuan, Li, Hongkai, Zhao, Shifang, Liu, Junnan, Huang, Weirong, Wang, Mengyu, Fu, Tianxiao, Wang, Yikai, Wang, Peng-Shuai, Jin, Xiaojie, Zhao, Yao, Wei, Yunchao
Abstract
Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras, and spatial-temporal dynamics. Yet existing generative methods rely on simple prompts to jointly control all of these factors through one-shot image or video synthesis, offering weak controllability and limited support for iterative editing. Fundamentally, a world comprises multiple elements with geometry, appearance, and other attributes, together with cameras. Different frames are produced through local modifications or recombinations of this shared state, which is otherwise largely reused. Therefore, we argue that the missing component is an explicit and persistent working state. To address this, we present StateFlow, a state-centric framework for generative previsualization. Rather than generating videos in one shot, StateFlow uses an editable 3D world to organize scene structure, evolution, and cameras, while off-the-shelf video models enhance visual quality when higher fidelity is desired. This world is maintained as a persistent structured 3D state of scene elements and camera configurations, serving as the core working representation for previsualization. Built on this insight, StateFlow has three stages to construct, evolve, and access the world state. State construction lifts generated 2D content into a coherent 3D world through prior-guided, conflict-aware dual-view initialization, while State evolution translates user intent into structured state transitions while preserving world memory, avoiding full-scene regeneration for each edit. State access uses render-feedback reflection to refine camera plans into visually feasible trajectories, avoiding reliance on VLM semantics alone. Experiments show that StateFlow produces high-quality 3D worlds for video creation and game-like prototyping.
Chinese Translation
预可视化是电影、游戏、建筑和城市设计中思想与制作之间的中介层。它允许创作者迭代地优化场景、动作、相机和时空动态。然而,现有的生成方法依赖简单的提示通过一次性图像或视频合成共同控制所有这些因素,提供了较弱的可控性和有限的迭代编辑支持。从根本上讲,世界由多个元素组成,具有几何形状、外观和其他属性,以及相机。不同的帧是通过对这一共享状态的局部修改或重组产生的,而这一状态在很大程度上是被重复使用的。因此,我们认为缺失的组成部分是一个明确且持久的工作状态。为了解决这个问题,我们提出了StateFlow,一个以状态为中心的生成预可视化框架。StateFlow并不是一次性生成视频,而是使用可编辑的3D世界来组织场景结构、演变和相机,同时现成的视频模型在需要更高保真度时增强视觉质量。这个世界被维持为场景元素和相机配置的持久结构化3D状态,作为预可视化的核心工作表示。基于这一洞察,StateFlow有三个阶段来构建、演变和访问世界状态。状态构建通过先前引导的、冲突感知的双视图初始化将生成的2D内容提升为一致的3D世界,而状态演变则将用户意图转化为结构化的状态转变,同时保留世界记忆,避免每次编辑时的全场景再生。状态访问使用渲染反馈反射将相机计划细化为视觉上可行的轨迹,避免仅依赖于VLM语义。实验表明,StateFlow为视频创作和类游戏原型制作生成高质量的3D世界。
人工智能 (Artificial Intelligence)
75
cs.AI / 1 / 2608.11207

Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes

多LLM代理系统的动态治理以实现协作对话结果
Liss, Alexander, Desmond, Nicholas, Gallego, Santiago Gil
Abstract
When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach, and the conversation terminates without achieving either agent's stated objective. This paper asks whether a control-theoretic governance layer can substitute for that missing goal function. The Experience Orchestrator (EO) addresses this in a simulated financial services environment where a site agent guides a visitor toward advisor contact while the visitor maintains psychologically realistic resistance. EO governs the joint trajectory through three mechanisms: a Contextual Bandit (CB) that selects content arms calibrated from real-world web analytics, a PID controller that enforces behavioral consistency via dynamic schema constraints, and a POMDP belief tracker that maintains a probabilistic model of visitor intent. Across 60,000 simulations, EO achieves a +32 percentage point lift in high-intent advisor contact rate (78.1% vs. 46.1% over a naive LLM control), with CB variant selection accounting for 97% of between-factor outcome variance -- confirming that the governance policy, not environmental initial conditions, determines where trajectories end up. Persona-level analysis reveals two distinct regimes: for visitors with no natural inclination toward conversion, the governance layer is the difference between a functional system and a non-functional one; for visitors already near alignment, a naive LLM's empathetic defaults are largely sufficient. All findings are conditional on LLM-to-LLM simulation. The PID controller has not been calibrated against real human unpredictability, and validating EO on live traffic is the critical next step.
Chinese Translation
当两个目标结构上对立的LLM代理在多个回合中互动时,缺乏共享目标函数并不会导致竞争,而是导致崩溃:访客屈服,站点代理停止改变其方法,谈话在未实现任何代理所声明的目标的情况下终止。本文探讨控制理论治理层是否可以替代缺失的目标函数。体验协调器(Experience Orchestrator, EO)在一个模拟的金融服务环境中解决了这一问题,其中站点代理引导访客与顾问联系,而访客则保持心理上现实的抵抗。EO通过三种机制来治理联合轨迹:一个上下文赌博者(Contextual Bandit, CB)选择根据真实世界网络分析校准的内容臂,一个PID控制器通过动态模式约束强制行为一致性,以及一个部分可观测马尔可夫决策过程(Partially Observable Markov Decision Process, POMDP)信念追踪器维护访客意图的概率模型。在60,000次模拟中,EO实现了高意图顾问联系率的+32个百分点提升(78.1%对比46.1%),其中CB变体选择占了97%的因素间结果方差——确认治理政策,而非环境初始条件,决定了轨迹的最终结果。个性化分析揭示了两种不同的机制:对于没有自然转化倾向的访客,治理层是功能系统与非功能系统之间的区别;对于已经接近一致的访客,简单的LLM的同理心默认设置基本上是足够的。所有发现均基于LLM与LLM的模拟。PID控制器尚未针对真实人类不可预测性进行校准,而在实时流量中验证EO是关键的下一步。
cs.AI / 2 / 2608.11210

Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

Distribird:基于文献的贝叶斯模型校准先验分布设计
Süli, Patrik P., Eigner, György, Hollós, Roland
Abstract
Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors from scientific literature is slow and needs both domain and statistical expertise. We present \textbf{Distribird}, an agentic web application that automates this process. Given a parameter name, physical description, and domain context, Distribird deploys a multi-agent pipeline that searches the literature, extracts and weights reported values by domain relevance, and fits a probability distribution via AIC model selection. When no literature is available, the system falls back to sensible uninformative alternatives, and clearly reports both the evidence behind and the confidence level of every prior it produces. It is designed for the problems where the models have physically interpretable parameters, where domain knowledge exists in the published literature. We evaluate the tool on 24~parameters across 10 scientific domains comparing three open-weight models (Qwen3.6 27B, Gemma 4 31B, Mistral Small 4 119B) with a single-prompt LLM baseline. On prior quality the full pipeline \emph{matches} this baseline. Every prior is traced to the specific papers and values from which it was constructed; a built-in validity layer declines to produce priors for out-of-scope requests, whereas the single-prompt baseline returns confident but unfounded priors for them in 11 of 30~model--parameter cases; and every language-model call runs locally, so no parameter description or unpublished modelling detail is transmitted to a third-party LLM provider (only generated search terms reach the public literature databases). For scientific use, we argue these properties matter more than a marginal improvement in point-estimate accuracy.
Chinese Translation
基于过程的模型的贝叶斯校准需要为每个模型参数设定一个先验分布。尽管经过数十年的方法论研究,研究人员几乎总是依赖于均匀先验。主要原因在于,从科学文献中构建信息丰富的先验分布过程缓慢,并且需要领域和统计方面的专业知识。我们提出了 extbf{Distribird},一个自动化此过程的智能网络应用程序。给定参数名称、物理描述和领域背景,Distribird 部署了一个多代理管道,搜索文献,提取并根据领域相关性加权报告值,并通过 AIC 模型选择拟合概率分布。当没有可用文献时,该系统会回退到合理的非信息性替代方案,并清晰地报告每个生成的先验背后的证据和置信水平。它旨在解决那些具有物理可解释参数且领域知识存在于已发表文献中的问题。我们在10个科学领域的24个参数上评估了该工具,并将其与三个开放权重模型(Qwen3.6 27B、Gemma 4 31B、Mistral Small 4 119B)和单提示LLM基线进行了比较。在先验质量方面,完整管道 extit{匹配}了该基线。每个先验都追溯到其构建所依据的具体论文和数值;内置的有效性层拒绝为超出范围的请求生成先验,而单提示基线在30个模型-参数案例中有11个返回了自信但毫无根据的先验;并且每次语言模型调用都在本地运行,因此没有参数描述或未发表的建模细节被传输到第三方LLM提供商(只有生成的搜索词会到达公共文献数据库)。对于科学用途,我们认为这些特性比点估计准确性的边际改善更为重要。
cs.AI / 3 / 2608.11211

A Forced-Structure Reduction and Verifiable Bounds for Conway's 99-Graph

康威的99图的强制结构简化与可验证界限
Thakkar, Aalok
Abstract
Conway's 99-graph problem asks whether a strongly regular graph with parameters $\mathrm{srg}(99,14,1,2)$ exists. We report a systematic, fully reproducible attack by an autonomous AI research agent, scored under the track's partial-credit metric. Our verifiable contributions are: (1) an exhaustive proof that no circulant graph on $\mathbb{Z}/99$ satisfies more than $3366/4950=68.0\%$ of the constraints ($33$ of $49$ difference-classes), with the same ceiling for the other abelian group of order $99$; (2) a forced-structure reduction: $\lambda=1$ makes each neighbourhood a perfect matching and $\mu=2$ puts the outer vertices in bijection with non-matched neighbour-pairs, collapsing existence to a $12$-regular graph on $84$ vertices, encoded for CP-SAT and validated by recovering the unique $\mathrm{srg}(9,4,1,2)$; (3) a validated prescribed-automorphism orbit-existence framework (fixed-point-free and single-fixed-point actions, checked on $\mathrm{srg}(9,4,1,2)$ and the Paley graph $\mathrm{srg}(13,6,2,3)$), and (4) a best verified artifact at $69.43\%$, with evidence that this is a robust frontier (fourteen distinct methods, none exceeding it) entangled with the open question, since any provable bound below $4950$ is a non-existence proof.
Chinese Translation
康威的99图问题询问是否存在参数为 $ ext{srg}(99,14,1,2)$ 的强正则图。我们报告了一项由自主AI研究代理进行的系统性、完全可重复的攻击,并在该轨道的部分评分指标下进行评估。我们的可验证贡献包括:(1) 详尽证明没有任何在 $ ext{Z}/99$ 上的循环图满足超过 $3366/4950=68.0\%$ 的约束($49$ 个差异类中的 $33$ 个),对于其他阶数为 $99$ 的阿贝尔群也有相同的上限;(2) 强制结构简化:$ heta=1$ 使每个邻域成为完美匹配,而 $ heta=2$ 将外部顶点与未匹配的邻居对建立双射,从而将存在性简化为一个在 $84$ 个顶点上的 $12$-正则图,已为 CP-SAT 编码并通过恢复唯一的 $ ext{srg}(9,4,1,2)$ 进行验证;(3) 验证的预设自同构轨道存在框架(无固定点和单固定点作用,在 $ ext{srg}(9,4,1,2)$ 和 Paley 图 $ ext{srg}(13,6,2,3)$ 上进行检查),以及 (4) 最佳验证结果为 $69.43\\%$,有证据表明这是一个稳健的前沿(十四种不同的方法,均未超过此值),与开放问题交织在一起,因为任何低于 $4950$ 的可证明界限都是一个不存在证明。
cs.AI / 4 / 2608.11212

Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts

检测路由翻转比判断是否修复更容易:量化混合专家中的因果路由介导损伤
Gu, Parvel
Abstract
Top-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-motivated numerical disturbance -- simulated 4-bit KV-cache quantization read by a protected BF16 gate -- pushes tokens across decision boundaries and flips which experts fire. This paper proposes no new mitigation; it supplies a causal apparatus, empirical findings, and a detection-limit result. A four-run apparatus prices the route-mediated fraction (RMF) of quantization damage, a token-level attribution decomposes it by mechanism, and pre-registered probes carry the findings across three architectures. On OLMoE-1B-7B at 4-bit KV (pilot), about a third of the damage is routing-mediated: RMF ~ 0.31 (discovery 0.31 [0.20, 0.41]; process-replicated mean 0.313 +/- 0.020; pre-registered re-execution 0.231). The deployable router margin detects that a flip occurred (AUC 0.772) but cannot tell a harmful flip from a helpful one (at chance): among the tested local, inference-observable router statistics we find no predictor of a flip's loss sign above chance -- an empirical benefit-detection barrier bounding selective repair restricted to this feature family. The signed-flip tax and sign-inseparability carry cross-model; the clean-reference remedy's payout is architecture-modulated; a controlled same-checkpoint flag-swap re-scopes the gate's normalization convention to a damage-magnitude moderator, not a route-recoverability mechanism. A real int4 KV kernel yields a fraction compatible with the fake-quant dose curve but underpowered (95% CI [-0.111, 0.394] includes zero) -- ruling out gross disagreement, not an independent replication. Hypotheses, thresholds, and evaluations were pre-registered before measurement, with misses reported; a pre-registered held-out read replicates the partition and the near-cancelling tax out of sample, while the strict impossibility exclusion narrowly misses.
Chinese Translation
Top-k 混合专家 (MoE) 路由是非连续的,因此一种基于部署的数值干扰——通过受保护的 BF16 门读取的模拟 4 位 KV-cache 量化——会推动令牌跨越决策边界并翻转触发的专家。本文没有提出新的缓解措施;而是提供了一种因果机制、实证发现和检测限制结果。一个四次运行的装置评估了量化损伤的路由介导比例 (RMF),一个令牌级归因将其按机制分解,预注册探针将发现结果扩展到三种架构。在 OLMoE-1B-7B 的 4 位 KV(试点)中,约三分之一的损伤是路由介导的:RMF ~ 0.31(发现 0.31 [0.20, 0.41];过程复制均值 0.313 +/- 0.020;预注册重执行 0.231)。可部署路由器的边际检测到翻转发生(AUC 0.772),但无法区分有害翻转和有益翻转(随机情况):在测试的局部、可推理的路由器统计中,我们未发现翻转损失符号的预测因子超过随机水平——这是一个实证的效益检测障碍,限制了对这一特征家族的选择性修复。签名翻转税和符号不可分离性在跨模型中存在;干净参考补救措施的收益受到架构调制;一个受控的同检查点标志交换将门的归一化惯例重新定义为损伤幅度的调节器,而非路由可恢复机制。一个真实的 int4 KV 内核产生的比例与假量化剂量曲线相符,但功率不足(95% CI [-0.111, 0.394] 包含零)——排除了严重不一致,而非独立复制。假设、阈值和评估在测量之前已预注册,并报告了遗漏;一个预注册的保留读取复制了分区和近乎抵消的税收,而严格的不可能排除则略有失误。
cs.AI / 5 / 2608.11215

Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop

穷人的代理建模:在笔记本电脑上模拟大型语言模型代理社会
Itkin, Igor
Abstract
Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent. We turn a statistical-physics observation into a method: replace each LLM agent by a low-parameter model fitted from a few hundred to a few thousand cheap queries, then run the society at any $N$ on a laptop. Whether this works is decided before the simulation runs, chiefly by what each agent perceives. We introduce an [interaction order x memory] taxonomy that maps perception and memory to an effective theory and a predicted $N$-trend of the surrogate error. We validate it on a faithful reimplementation of the LLM macroeconomy EconAgent and seven further named LLM simulations, with agent decisions cloned from genuine LLM elicitations (primarily DeepSeek) for a few dollars; the predicted error trends hold cell by cell, and the two refuted predictions, both on a strongly saturating response and traced to its curvature, are themselves matched quantitatively by the theory with no free parameters.
Chinese Translation
模拟许多大型语言模型(LLM)代理的社会是昂贵的,但对这些模拟提出的问题通常是宏观的:相行为、典型事实以及与代理数量 $N$ 的缩放,而不是任何单个代理的认知。我们将一个统计物理观察转化为一种方法:用从几百到几千个廉价查询拟合的低参数模型替代每个 LLM 代理,然后在笔记本电脑上以任意 $N$ 运行该社会。这是否有效在模拟运行之前就已决定,主要取决于每个代理的感知。我们引入了一种[交互顺序 x 记忆]分类法,将感知和记忆映射到有效理论以及预测的 $N$ 趋势的替代误差。我们在对 LLM 宏观经济模型 EconAgent 的忠实重新实现以及另外七个命名的 LLM 模拟中验证了这一点,代理决策是从真实的 LLM 引导(主要是 DeepSeek)中克隆的,成本仅为几美元;预测的误差趋势在每个单元中都成立,而两个被否定的预测(均基于强饱和响应并追溯到其曲率)则通过理论在没有自由参数的情况下定量匹配。
cs.AI / 6 / 2608.11216

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

AutoWorldModel-Bench:一个以状态为中心的自动化世界模型研究基准
Moodi, Marjan, Zhu, Xuankang, Silva, Fernando De Mesentier, Chaput, Harold, Taesiri, Mohammad Reza
Abstract
World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomous researchers--a setting in which the improvement direction is not specified in advance, unlike the engineering-to-spec tasks that dominate current agent benchmarks. We introduce AutoWorldModel-Bench, a closed-loop benchmark in which frontier coding agents autonomously improve a provided world-model starter under a fixed compute budget. The benchmark spans eight game environments under a unified structured-state representation--ground-truth entity state extracted from each game and consumed through a shared tensor format--which isolates dynamics modeling from perception and enables minutes-per-run iteration. Across 64 sessions, Codex-5.4 and Claude Opus 4.6 improve their starter on 63; in 91% of sessions the winning edit is a non-trivial research-style modification--a new objective, representation, rollout procedure, or architectural change--rather than a hyperparameter tweak. Our benchmark offers a setting in which frontier coding agents can be evaluated on open-ended research rather than engineering-to-spec problems.
Chinese Translation
世界建模是一个尚未定型的领域:架构、训练目标和状态表示以复杂的方式相互作用,且没有单一的方法在各种环境中占主导地位。这使其成为一个理想的测试平台,适用于作为自主研究者的人工智能编码代理——在这种环境中,改进方向并未事先指定,这与当前代理基准中占主导地位的工程到规格任务不同。我们引入了AutoWorldModel-Bench,这是一个闭环基准,其中前沿编码代理在固定计算预算下自主改进提供的世界模型起始版本。该基准涵盖八个游戏环境,采用统一的结构化状态表示——从每个游戏中提取的真实实体状态,并通过共享的张量格式进行处理——这将动态建模与感知隔离开来,并实现每轮几分钟的迭代。在64个会话中,Codex-5.4和Claude Opus 4.6在63个会话中改进了其起始版本;在91%的会话中,获胜的编辑是非平凡的研究风格修改——新的目标、表示、展开程序或架构变更,而不是超参数调整。我们的基准提供了一个环境,在该环境中,前沿编码代理可以在开放式研究而非工程到规格的问题上进行评估。
cs.AI / 7 / 2608.11218

MaSRead: Content-Addressed Reading of Replicated Latent Stores

MaSRead:对复制潜在存储的内容寻址阅读
Baquero, Carlos, Brito, Luís, Resende, João
Abstract
Independent agents that reason in latent space can share computed state as key-value cache fragments rather than text. Merged by a conflict-free replicated data type, these fragments form a store that converges under any delivery order or duplication. Yet a later query, unknown at encode time, cannot reliably read the merged cache: colocated fragments interfere, so colocation is not addressability. MaSRead addresses the read to content. It routes through opaque keyed tag sets derived from fragment words and decodes each selected fragment under a hard attention mask that hides the rest. Under lexical connectivity, a graph walk reaches the fragments required by a multi-hop query. Across chain, pipeline, symmetric, hub, and natural-language stores, MaSRead recovers visited fragments in isolation, remains effective as unrelated fragments accumulate, and transfers to another model family. After routing, materialized decoding depends on fragment length rather than total store size; end-to-end work still includes store-dependent routing and one read per visited fragment. The limits are explicit: lexical routing can miss disconnected evidence, and answer composition remains bounded by the frozen reader. Thus a replicated latent store becomes selectively readable for later queries when the needed fragments connect to the query through content.
Chinese Translation
在潜在空间中进行推理的独立代理可以将计算状态作为键值缓存片段而非文本进行共享。通过无冲突复制数据类型合并,这些片段形成一个在任何传递顺序或重复下都能收敛的存储。然而,后续查询在编码时未知,无法可靠地读取合并的缓存:共置片段相互干扰,因此共置并不等同于可寻址性。MaSRead 解决了对内容的读取问题。它通过从片段词汇派生的透明键标签集进行路由,并在一个硬注意力掩码下解码每个选定的片段,掩盖其余部分。在词汇连接的基础上,图遍历可以到达多跳查询所需的片段。在链式、管道式、对称式、中心式和自然语言存储中,MaSRead 能够在隔离状态下恢复访问过的片段,随着无关片段的累积仍然保持有效,并且可以转移到另一种模型家族。路由后,物化解码依赖于片段长度而非总存储大小;端到端的工作仍然包括依赖存储的路由和每个访问片段的一次读取。其限制是明确的:词汇路由可能会遗漏断开的证据,而答案组合仍然受到冻结读取器的限制。因此,当所需片段通过内容与查询连接时,复制的潜在存储变得可选择性可读,以便于后续查询。
cs.AI / 8 / 2608.11219

From Monolithic to Modular: Segment-level Automatic Prompt Optimization

从单体到模块化:基于段落的自动提示优化
Kulin, Nikita, Zhuravlev, Viktor, Khairullin, Artur, Muravyov, Sergey, Makarov, Ilya, Sukhorukov, Daniil, Averkova, Ekaterina
Abstract
Automatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others. We present SAPO, a segment-level APO method that decomposes prompts into role, context, tasks, and output format, then applies targeted improvements based on top-5 and bottom-5 examples. The optimization loop uses one LLM with static meta-prompts and structured outputs for segmentation, weakness analysis, and candidate generation. We describe a train/validation protocol and a two-stage generation process: (1) segment-level diagnosis and recommendation extraction, (2) candidate synthesis constrained by weak/strong segment signals. Using the evaluation setup across SQuADv2, TweetEval, XSUM, CommonGen, and GSM8K on GPT-3.5-Turbo and GPT-4o-mini, SAPO achieves the best average score against Zero-shot and strong APO baselines including APE, OPRO, EvoPrompt, GEPA, and StraGO.
Chinese Translation
自动提示优化(Automatic Prompt Optimization, APO)通常以单一整体的方式重写提示,这可能会改善某一行为,但却可能导致其他行为的下降。我们提出了SAPO,一种基于段落的APO方法,该方法将提示分解为角色、上下文、任务和输出格式,然后基于前五和后五个示例进行有针对性的改进。优化循环使用一个大型语言模型(LLM),结合静态元提示和结构化输出进行分段、弱点分析和候选生成。我们描述了一种训练/验证协议和一个两阶段生成过程:(1)基于段落的诊断和推荐提取,(2)受弱/强段信号约束的候选合成。通过在GPT-3.5-Turbo和GPT-4o-mini上对SQuADv2、TweetEval、XSUM、CommonGen和GSM8K的评估设置,SAPO在零-shot和强APO基准(包括APE、OPRO、EvoPrompt、GEPA和StraGO)中实现了最佳平均得分。
cs.AI / 9 / 2608.11220

LLMs in Process Diagram Engineering: From Optimal PFDs to Validated P&IDs

过程图工程中的大型语言模型:从最优的流程图到验证的管道和仪表图
Zakarin, Timur, Voitov, Sergei, Shumilin, Sergei, Burnaev, Evgeny
Abstract
Nowadays, the creation of a process flow diagram (PFD) and its subsequent transformation into a piping and instrumentation diagram (P&ID) is predominantly performed manually. Applying artificial intelligence in the task could potentially lead not only to process automation and time savings, but also to financial gains by exploring numerous diagram's topology options and reducing manual labor. This research presents P&ID Pilot - a practical end-to-end AI pipeline capable of handling flowsheet developing for both stages. The first stage focuses on PFD synthesis, whereas the second is directed toward modifying the generated PFD into P&ID. After comparing four different methods, the hybrid approach combining genetic algorithms (GA) and large language models (LLM) is shown to generate the optimal valid PFD topology, achieving the lowest loss value among all the methods, while satisfying the required outlet flow parameters without engineering-rule violations. For the second stage, the proposed LLM-based agent successfully transforms the generated PFD into a source-grounded P&ID by producing validated, executable modifications through a restricted engineering software development kit, achieving 100% execution success while maintaining compliance with domain-specific rules and reference graph structures. This unified pipeline - coupling GA/LLM-driven synthesis with an LLM-based transformation agent - offers a feasible path toward end-to-end process design automation by producing validated, deployable outputs and substantially reduces manual engineering effort.
Chinese Translation
如今,流程图(PFD)的创建及其后续转化为管道和仪表图(P&ID)主要是通过手动完成的。将人工智能应用于这一任务不仅有可能实现过程自动化和节省时间,还可以通过探索众多图形拓扑选项和减少人工劳动来带来经济收益。本研究提出了P&ID Pilot——一个实用的端到端人工智能管道,能够处理两个阶段的流程图开发。第一个阶段侧重于PFD合成,而第二个阶段则致力于将生成的PFD修改为P&ID。在比较了四种不同的方法后,结合遗传算法(GA)和大型语言模型(LLM)的混合方法被证明能够生成最优有效的PFD拓扑,在所有方法中实现了最低的损失值,同时满足所需的出口流量参数且不违反工程规则。在第二个阶段,所提出的基于LLM的代理成功地通过生成经过验证的可执行修改,将生成的PFD转化为源基础的P&ID,达成了100%的执行成功率,同时保持了对特定领域规则和参考图结构的合规性。这个统一的管道——将GA/LLM驱动的合成与基于LLM的转化代理相结合——为实现端到端过程设计自动化提供了一条可行的路径,能够生成经过验证的可部署输出,并显著减少人工工程工作量。
cs.AI / 10 / 2608.11221

A Conceptual Framework for Refining Influence Knowledge from Simulation Evidence in Cyber-Physical Systems

从网络物理系统的仿真证据中提炼影响知识的概念框架
Oliveira, Barbara da Silva, Deantoni, Julien, Ferry, Nicolas
Abstract
Cyber-physical systems (CPS) are typically developed by multiple stakeholders who produce artefacts tailored to their specific domains of expertise. The behaviour of these systems emerges from the interaction between those artefacts and their operational environment. Simulation and co-simulation have become essential approaches for analysing CPS behaviour and, through simulation campaigns, developers can explore system responses under changing conditions, including interactions with the environment. However, the lack of details and understanding of some environmentmediated interactions (typically the ones beyond direct sensing and actuation), which remain unmodelled due to their complexity, a lack of time, or a lack of domain experience, hinders the proper comprehension and exploitation of simulation results. To address these limitations, we propose a conceptual framework leveraging the novel concept of Influences to support the iterative and incremental refinement of simulation campaigns and deepen the understanding of the system behaviour. We demonstrate the proposed approach through a case study involving a mobile robot implemented using Simulink/Gazebo co-simulation.
Chinese Translation
网络物理系统(CPS)通常由多个利益相关者开发,他们生产针对特定专业领域的工件。这些系统的行为源于这些工件与其操作环境之间的相互作用。仿真和协同仿真已成为分析CPS行为的重要方法,通过仿真活动,开发人员可以探索系统在变化条件下的响应,包括与环境的相互作用。然而,由于某些环境介导的相互作用(通常是超出直接传感和执行的那些)缺乏细节和理解,这些相互作用因其复杂性、时间不足或缺乏领域经验而未被建模,从而阻碍了对仿真结果的正确理解和利用。为了解决这些局限性,我们提出了一个概念框架,利用影响(Influences)这一新颖概念,支持仿真活动的迭代和增量改进,并加深对系统行为的理解。我们通过一个案例研究展示了所提方法,该案例涉及使用Simulink/Gazebo协同仿真实现的移动机器人。
cs.AI / 11 / 2608.11224

Harnessing agent memory to build lifelong AI partners for materials scientists

利用代理记忆构建材料科学家的终身人工智能伙伴
Liu, Siyu, Hu, Bo, Ye, Beilin, Cao, He, Srolovitz, David J., Wen, Tongqi
Abstract
Materials research advances through accumulated experience - scripts that work, protocols that are trusted, warnings attached to failed calculations or experiments, and judgement that links a new question to an old result. This experience is essential for reproducibility and knowledge transfer, yet it is usually fragmented across notebooks, repositories, job logs and individual memory, and it is rarely portable across artificial-intelligence agents. Here we argue that a lifelong AI partner for materials science can be designed around persistent memory rather than around a particular agent implementation. We introduce a self-evolving memory framework that stores scientific experience as inspectable facts and executable skills, so that observations, failure boundaries, protocols and validation checks can be retrieved, revised and migrated across models. We evaluate the idea in three computational settings that expose different layers of materials-research competence. In 49 real-world materials-tool-use questions comprising 138 executable subtasks, memory nearly doubles GPT-5.2 task success without model-parameter updates. In elemental-solid equation-of-state calculations, memory converts a wavefunction-initialization failure into a pre-execution guardrail, improving outcomes from 22/1/4 to 25/2/0 Correct/Partial/Error and avoiding 92% of repeated errors. In 13 practical material simulation workflows, remembered skills and failure facts halve the aggregate trace burden (tokens) and reduce tool calls by over a factor of two by the third round, while preserving physically meaningful outputs in band-gap, phonon, vacancy and work-function analyses. These results show that agent memory can serve as a durable scientific asset; a portable, self-improving record of materials-research experience that outlives any single model or agent stack.
Chinese Translation
材料研究通过积累经验而不断进步——有效的脚本、可信的协议、附加在失败计算或实验上的警告,以及将新问题与旧结果联系起来的判断。这种经验对于可重复性和知识转移至关重要,但通常分散在笔记本、代码库、作业日志和个人记忆中,并且很少能够在人工智能代理之间进行迁移。在此,我们认为可以围绕持久记忆而非特定代理实现设计出材料科学的终身人工智能伙伴。我们提出了一种自我演化的记忆框架,将科学经验存储为可检查的事实和可执行的技能,以便可以检索、修订和迁移观察、失败边界、协议和验证检查。我们在三种计算环境中评估了这一理念,这些环境揭示了材料研究能力的不同层面。在49个真实世界的材料工具使用问题中,包括138个可执行的子任务,记忆几乎使GPT-5.2的任务成功率翻倍,而无需更新模型参数。在元素固体状态方程计算中,记忆将波函数初始化失败转化为执行前的保护措施,使结果从22/1/4(正确/部分/错误)改善到25/2/0,并避免了92%的重复错误。在13个实际材料模拟工作流程中,记忆的技能和失败事实将总追踪负担(标记数)减半,并在第三轮中将工具调用减少了超过两倍,同时在带隙、声子、空位和功函数分析中保持了物理上有意义的输出。这些结果表明,代理记忆可以作为一种持久的科学资产;它是一种可移植的、自我改进的材料研究经验记录,超越了任何单一模型或代理堆栈。
cs.AI / 12 / 2608.11225

Identity from the Outside: A Conceptual Framework and Research Program for AI Personality Clones

外部身份:人工智能个性克隆的概念框架与研究计划
Brunet, Luc E.
Abstract
AI "personality clones" force a re-examination of personal identity in operational terms. Setting aside the hard problem of consciousness, we approach identity through the indiscernibility of manifestations, as assessed by an observer over a duration. We distinguish three criteria that "identity" conflates: fidelity to a target person, generic human-likeness, and individuality. We propose a six-term factorization of observed identity (substrate, dispositions, memory, update dynamics, context, exogenous contingencies), with a state-space formulation. Indiscernibility is defined as one minus a judge's distinguishing advantage, and the factorization's coefficients become local sensitivities estimable by randomized ablation. The central claim is a conditional conjecture: given hypotheses about the agent's information on its own persistence and about consequences bearing on its own stakes, versionability tends to degrade long-horizon indiscernibility. An analogy with lambda-calculus, linear typing, and bisimulation clarifies what linearity does and does not establish. Between product-clone and individual we identify a third object, the delegate: a task-limited, bounded-lifespan partial clone ending in a bandwidth-limited testament. We map the empirical literature onto the three criteria, propose an experimental program, and argue that the correct long-horizon criterion is not trajectory fidelity but climate fidelity: matching the conditional distribution of a person's possible responses. The best clone is the one that diverges from the original as the original would have diverged from itself.
Chinese Translation
人工智能“个性克隆”迫使我们在操作层面重新审视个人身份。抛开意识的难题,我们通过观察者在一段时间内对表现的不可区分性来探讨身份。我们区分了“身份”所混淆的三个标准:对目标个体的忠实度、一般人类相似性和个体性。我们提出了一种观察到的身份的六项因子分解(基质、倾向、记忆、更新动态、上下文、外生偶然性),并采用状态空间的形式化。不可区分性被定义为评判者的区分优势减去一,因子分解的系数成为可通过随机消融估算的局部敏感性。中心主张是一个条件猜想:在关于代理人自身持续性的信息假设和关于影响其自身利益的后果的假设下,版本化往往会降低长期的不可区分性。与λ-演算、线性类型和双模拟的类比阐明了线性所确立的内容与未确立的内容。在产品克隆与个体之间,我们识别出第三种对象,即委托者:一种任务有限、寿命受限的部分克隆,最终以带宽有限的遗嘱结束。我们将实证文献映射到这三个标准上,提出一个实验计划,并论证正确的长期标准不是轨迹忠实度,而是气候忠实度:匹配一个人可能反应的条件分布。最佳克隆是那个与原始个体的偏离程度与原始个体自身的偏离程度相一致的克隆。
cs.AI / 13 / 2608.11226

Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet

通过强化学习降低人工智能数据中心能耗:从单个GPU到整个集群的LLM训练功率控制测量
Curcio, Eliseo
Abstract
Reinforcement-learning post-training dominates modern language-model development, yet its power behavior on GPU hardware has not been characterized, and datacenters manage GPU power with workload-blind mechanisms, static caps and reactive throttling, that slow hardware indiscriminately. We instrument GRPO training with half-second power telemetry at 7B, 14B, and 72B scales on one to four A100s (380,000+ samples), and train a PPO meta-controller that adapts the workload's own generation parameters to measured power. Against the full 500-step 7B trace, the controller cuts power-limit violations by 89.8% while increasing token output by 18.1% and energy efficiency by 26.2% (tokens per MWh). Deployed live at 72B, the same controller family yields replicated null results, diagnosed as the group-size actuator losing authority under model sharding. An actuator-authority sweep shows the same parameters applied as generation concurrency retain 17-22% power authority, isolating an occupancy-versus-volume principle; a controller rebuilt on that actuator controls a live 72B rollout-generation workload across three replications: 35.7% more output than a static safe baseline at 2.27 +/- 1.08% budget violations, 87.2% fewer violations than uncontrolled operation, and the best mean throughput and energy per token among constrained controllers, with an adaptive threshold rule matching it in one of three operating conditions. Under realistic measurement windows the original 72B transients fall from 23.6% at half-second resolution to 1.6% at 30 s and zero at 5 min; a composed 16-GPU fleet shows zero violations at 30 s and longer, with peak demand at 50-56% of nameplate. For this fleet mix, roughly twofold oversubscription of nameplate appears feasible, subject to operator validation. We quantify the economic and carbon consequences and specify a low-cost operator pilot.
Chinese Translation
强化学习后训练在现代语言模型开发中占据主导地位,但其在GPU硬件上的功率行为尚未得到表征,而数据中心则通过对工作负载盲目的机制、静态限制和反应性节流来管理GPU功率,这种方式会无差别地减慢硬件性能。我们在一到四个A100上对GRPO训练进行了半秒功率遥测,规模为7B、14B和72B(超过380,000个样本),并训练了一个PPO元控制器,使其能够根据测量到的功率调整工作负载的生成参数。在完整的500步7B轨迹中,控制器将功率限制违规减少了89.8%,同时将令牌输出提高了18.1%,能效提高了26.2%(每MWh的令牌数)。在72B的实时部署中,同一控制器系列产生了重复的无效结果,诊断为在模型分片下组大小执行器失去控制权。执行器控制权的扫描显示,当作为生成并发性应用相同参数时,仍保持17-22%的功率控制权,隔离出占用与体积的原则;基于该执行器重建的控制器在三个复制中控制实时72B的滚动生成工作负载:输出比静态安全基线多35.7%,预算违规率为2.27 +/- 1.08%,比不受控制的操作减少87.2%的违规,并在受限控制器中实现最佳的平均吞吐量和每令牌能耗,其中一个自适应阈值规则在三种操作条件下与之匹配。在现实的测量窗口下,原始72B瞬态从半秒分辨率的23.6%降至30秒的1.6%和5分钟的零;一个由16个GPU组成的集群在30秒及更长时间内显示零违规,峰值需求为额定值的50-56%。对于这一集群组合,约两倍的额定超额订阅似乎是可行的,需经运营商验证。我们量化了经济和碳后果,并指定了一个低成本的运营商试点。
cs.AI / 14 / 2608.11227

Forecasting Side Effects of Activation Steering

预测激活引导的副作用
Ong, Chong Yong, Sim, Alson Wei Jie, Zhang, Peixin, Sun, Jun
Abstract
Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. While effective, steering often produces unintended side effects on other behaviors, making it difficult to deploy safely. We therefore ask: can these side effects be forecasted before steering is applied? We answer this question by constructing a cross-effect matrix over a taxonomy of 67 behaviors across three open-weight language models. We find that side effects are common, structured, and often asymmetric, revealing interactions that cannot be explained by existing similarity-based heuristics. Despite this complexity, we show that side effects are largely predictable before steering is performed. Their magnitude depends primarily on the target behavior, while their direction can be forecasted from the model's unsteered representations with substantially higher accuracy than simple baselines. Our results demonstrate that activation steering has systematic and forecastable side effects, enabling proactive safety auditing and more informed deployment of steering interventions.
Chinese Translation
激活引导通过向语言模型的隐藏激活添加学习到的方向来修改模型,从而实现有针对性的行为变化而无需重新训练。尽管有效,激活引导往往会对其他行为产生意想不到的副作用,这使得安全部署变得困难。因此,我们提出了一个问题:在应用激活引导之前,这些副作用是否可以被预测?我们通过构建一个跨效应矩阵,涵盖三个开放权重语言模型中的67种行为分类,来回答这个问题。我们的研究发现,副作用是普遍存在的,具有结构性,并且往往是非对称的,这揭示了现有基于相似性启发式方法无法解释的交互关系。尽管存在这种复杂性,我们表明,在进行激活引导之前,副作用在很大程度上是可预测的。副作用的大小主要取决于目标行为,而其方向可以从模型未引导的表示中以显著高于简单基线的准确性进行预测。我们的结果表明,激活引导具有系统性和可预测的副作用,从而使得主动安全审计和更为知情的引导干预部署成为可能。
cs.AI / 15 / 2608.11229

Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version)

在人机自主团队中使用二阶心智理论同步信念(扩展版)
Mirenzi, Jack, Admoni, Henny
Abstract
Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align robot and agent behavior with human intent when the reward itself cannot be specified directly. Preference-based reward learning typically casts the human teacher as a passive oracle answering learner-generated queries. We argue this forfeits the teacher's defining advantage: knowledge of the objective. A teacher who knows the target can construct training examples more efficiently than any learner-driven acquisition strategy, an advantage that widens as the reward's feature dimension grows. However, exploiting this advantage requires an accurate model of what the learner currently knows. We therefore recast preference learning as a human-autonomy team problem coupling two behavioral models: the teacher maintains a model of the learner to design an informative curriculum, and the learner maintains a second-order model of the teacher's model, emitting structured preference constraints (understanding statements) that keep the teacher's model of the learner synchronized. In simulation, an informed teacher outperforms learner-led selection; teacher-model drift under alternating teachers erodes this advantage; and understanding statements repair it, with second-order (ToM-2) statements outperforming mean-belief statements when the teacher's error about the learner is concentrated in a particular direction rather than spread evenly.
Chinese Translation
比较反馈,即询问人们在两种行为中更偏好哪一种,已成为一种标准方法,用于在奖励无法直接指定时使机器人和智能体的行为与人类意图对齐。基于偏好的奖励学习通常将人类教师视为被动的神谕者,回答学习者生成的查询。我们认为,这放弃了教师的一个关键优势:对目标的知识。一个了解目标的教师能够比任何学习者驱动的获取策略更有效地构建训练示例,这一优势在奖励的特征维度增大时会进一步扩大。然而,利用这一优势需要对学习者当前所知的内容有一个准确的模型。因此,我们将偏好学习重新构建为一个人机自主团队问题,结合了两个行为模型:教师维护一个学习者模型,以设计信息丰富的课程,而学习者则维护一个教师模型的二阶模型,发出结构化的偏好约束(理解陈述),以保持教师对学习者模型的同步。在仿真中,信息丰富的教师优于学习者主导的选择;在交替教师下,教师模型的漂移削弱了这一优势;而理解陈述则修复了这一点,当教师对学习者的错误集中在某个特定方向而不是均匀分布时,二阶(ToM-2)陈述优于均值信念陈述。
cs.AI / 16 / 2608.11230

The Edge-based Contiguous p-median Problem with Connections to Logistics Districting

基于边缘的连续 p-中位数问题与物流区域划分的联系
Kassem, Zeyad, Escobedo, Adolfo R.
Abstract
This paper introduces the edge-based contiguous p-median (ECpM) problem to partition the roads in a network into a given number of compact and contiguous territories. Two binary programming models are introduced, both of which incorporate a network distance. The first model requires an exponential number of cut set-based constraints to model contiguity; it is paired with a separation scheme that usually generates only a small number of these constraints, namely, a branch-and-cut (B&C) algorithm. The second model utilizes a polynomial number of shortest-path constraints to model contiguity and can be solved with off-the-shelf solvers. The respective solution approaches are tested on road networks with over 2,700 nodes and close to 3,400 edges, yielding models with over 9.6 million binary variables. Solving the model based on shortest path contiguity (SPC) constraints via standard branch and bound attains speedups in computational time of up to 17x relative to the cut set-based B&C implementation. In addition, the SPC constraints are demonstrated to be supervalid inequalities of the edge-based p-median (EpM) model (i.e., for which contiguity is not explicitly required), meaning that they may cut off integer-feasible solutions and some, but not all, of the optimal solutions of this simpler problem. Finally, the paper explores structural insights and connections between ECpM and the edge-based districting (EBD) problem, which enforces an additional work balance criterion. An existing model that utilizes cut set-based contiguity constraints was unable to find a feasible solution within 12 hours for any of the tested instances, while an SPC-based EBD model was able to solve most of these to optimality.
Chinese Translation
本文介绍了基于边缘的连续 p-中位数(ECpM)问题,旨在将网络中的道路划分为给定数量的紧凑且连续的区域。提出了两个二进制规划模型,这两个模型均结合了网络距离。第一个模型需要指数数量的基于切割集的约束来建模连续性;它配备了一个分离方案,通常仅生成少量这些约束,即分支切割(B&C)算法。第二个模型利用多项式数量的最短路径约束来建模连续性,并可以使用现成的求解器进行求解。各自的解决方法在超过2700个节点和近3400条边的道路网络上进行了测试,生成的模型具有超过960万个二进制变量。通过标准的分支界限方法求解基于最短路径连续性(SPC)约束的模型,相较于基于切割集的 B&C 实现,计算时间的加速可达17倍。此外,SPC约束被证明是基于边缘的 p-中位数(EpM)模型的超有效不等式(即不明确要求连续性),这意味着它们可能会切断整数可行解以及一些但不是所有的更简单问题的最优解。最后,本文探讨了 ECpM 与基于边缘的区域划分(EBD)问题之间的结构性见解和联系,后者强制执行额外的工作平衡标准。一个利用基于切割集的连续性约束的现有模型在测试的任何实例中都未能在12小时内找到可行解,而基于 SPC 的 EBD 模型能够将大多数这些实例求解至最优。
cs.AI / 17 / 2608.11231

LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs

LinearKV:单一缓存状态足以实现混合大语言模型中的位置无关缓存
Liu, Yirui, Qi, Ruoling, Wang, Longwen, Wu, Xuaner, Chen, Jian, Jin, Yuxin, Shao, Jiawei, Li, Xuelong
Abstract
LLM serving is increasingly accelerated by position-independent caching (PIC). Existing PIC methods, however, are built for full-attention models, where a token-indexed KV cache underlies its core operations: matching reusable token chunks, concatenating their KV entries, and selectively recomputing a few tokens to restore cross-chunk context. Hybrid LLMs break these primitives---they replace most attention layers with linear recurrences that expose only a fixed-size state, leaving no token-indexed KV to concatenate or to locally repair. This raises a natural question: can PIC benefit hybrid models, and what would it take? We present LinearKV, a training-free hybrid-PIC framework. Its key insight is a \emph{decoupled initialization}: each linear layer maps its $K$ matched local states to a single initial state, while full-attention layers concatenate their KV as before. LinearKV is therefore compatible with existing PIC methods, reusing their token selection and recomputation as-is. Under this framework, we find that a \emph{single cached state} suffices as the linear layer's initializer. The algebraically principled alternative---composing all $K$ cached states into the exact full-prefix state, as concurrent work HYPIC does---is unnecessary and, on some architectures, even harmful. We compare the two across three hybrid models and three PIC selectors. On the two GDN models the two tie, both recovering most of full quality (up to $92\%$); on the Mamba-2 model, exact composition instead collapses under every selector---under EPIC, for instance, it recovers only $46.6\%$ of full quality, versus $86.8\%$ for a single cached block initializer. A single state initializer is also cheaper, cutting time-to-first-token to $0.46\times$ full prefill versus a further $5$--$17\%$ overhead for exact composition; results hold across LongBench QA and RULER at 8K--32K.
Chinese Translation
大语言模型(LLM)的服务越来越依赖于位置无关缓存(PIC)。然而,现有的PIC方法是为全注意力模型设计的,其中基于令牌索引的KV缓存是其核心操作的基础:匹配可重用的令牌块、连接它们的KV条目,以及选择性地重新计算少量令牌以恢复跨块上下文。混合大语言模型打破了这些原语——它们用线性递归替代了大多数注意力层,仅暴露固定大小的状态,因而没有令牌索引的KV可以连接或局部修复。这引出了一个自然的问题:PIC能否惠及混合模型?需要什么条件?我们提出了LinearKV,一个无训练的混合PIC框架。其关键见解是 extit{解耦初始化}:每个线性层将其$K$个匹配的局部状态映射到一个单一的初始状态,而全注意力层则如之前一样连接它们的KV。因此,LinearKV与现有的PIC方法兼容,按原样重用它们的令牌选择和重新计算。在这一框架下,我们发现 extit{单一缓存状态}足以作为线性层的初始化器。代数上合理的替代方案——将所有$K$个缓存状态组合成精确的全前缀状态,如并行工作HYPIC所做——是不必要的,并且在某些架构上甚至是有害的。我们在三个混合模型和三个PIC选择器之间进行了比较。在两个GDN模型中,两者相当,都恢复了大部分的全质量(高达$92 ext{%}$);而在Mamba-2模型中,精确组合在每个选择器下都崩溃——例如,在EPIC下,它仅恢复了$46.6 ext{%}$的全质量,而单一缓存块初始化器则恢复了$86.8 ext{%}$。单一状态初始化器的成本也更低,将首次令牌的时间缩短到$0.46 imes$的全预填充,而精确组合则增加了$5 ext{%}$至$17 ext{%}$的开销;结果在LongBench QA和RULER的8K至32K范围内保持一致。
cs.AI / 18 / 2608.11234

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

InfraBench:跨层、生命周期和风险评估基础设施代理
Gao, Yuan, Yang, Zeren, Li, Junnan, Shawn, Zhong, Dajani, Ahmed, Zheng, Mai, Arpaci-Dusseau, Andrea, Arpaci-Dusseau, Remzi
Abstract
Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, but it remains unclear how well such agents can handle real-world infrastructure complexity. We present InfraBench, a benchmark suite for evaluating AI agents on realistic infrastructure tasks across the full system stack and full operational lifecycle with fine-grained risk assessment. Experiments with 15 agent-model configurations show that even the strongest agent cannot secure a full score across all tasks. Mean effective scores range from roughly 40% to 88% (with per-configuration standard errors of 6-12 points), repeating every task three times reveals that top configurations still pass only a fraction of their attempts, and per-check scoring exposes a general failure pattern: agents may routinely satisfy short-term objectives while leaving non-durable changes, broken distributed invariants, unsafe side effects, and uncleaned state behind. INFRABENCH, including its live leaderboard, tasks, and evaluation harness, is publicly available at infraben.ch.
Chinese Translation
管理现代计算基础设施已成为一个日益困难的问题,原因在于其复杂性不断增加。近期在人工智能代理方面的进展为自动化基础设施管理任务提供了及时的机会,但尚不清楚这些代理在处理现实世界基础设施复杂性方面的表现如何。我们提出了InfraBench,这是一个基准套件,用于评估人工智能代理在整个系统栈和完整操作生命周期中的现实基础设施任务,结合细致的风险评估。对15种代理模型配置的实验表明,即使是最强大的代理也无法在所有任务中获得满分。有效平均得分范围大约在40%到88%之间(每个配置的标准误差为6-12分),对每个任务重复三次的结果显示,顶级配置仍然仅能通过其尝试的一小部分,而逐项评分揭示了一种普遍的失败模式:代理可能常常满足短期目标,但留下了不可持续的变化、破损的分布式不变性、不安全的副作用和未清理的状态。INFRABENCH,包括其实时排行榜、任务和评估工具,已在infraben.ch上公开发布。
cs.AI / 19 / 2608.11235

CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference

CORA-Diff:面向信心的残差接受机制用于高效的扩散语言模型推理
Wu, Yifan, Zhang, Yufeng, Li, Kenli
Abstract
Diffusion language models (DLMs) update many tokens in parallel, yet practical decoders often use a fixed denoising horizon. Many predictions stabilize early, but blockwise decoding continues until all positions are resolved, causing repeated dense forward passes. Existing accelerators often rely on learned filters, modified scores, dependency models, or cache-specific mechanisms. We ask whether native trajectory signals can identify residual positions likely to match the deterministic dense endpoint. We propose CORA-Diff, a training-free method that preserves the original transfer rule and applies confidence-and-persistence gating only to positions that rule leaves unresolved. Accepted tokens remain visible as context, and the block terminates once all positions are resolved. This requires no backbone change, learned acceptance model, or logit modification. Our theory explains why high-confidence, persistent predictions are more likely to match the fixed-horizon dense endpoint, and paired post-intervention trajectories provide direct empirical support. We select one operating point on a separate GSM8K calibration subset and freeze it for all evaluations. Under a matched Learn2PD-style LLaDA protocol, CORA-Diff has the lowest measured runtime in all eight task-length settings. Task scores match or exceed dense decoding in five settings, and the largest observed drop is 1.22 points. Its incremental speedups over EOS-aware dense decoding are 2.70x and 3.32x on GSM8K and HumanEval. It also reaches 13.14x under the fixed-horizon 1024/1024 mechanism-isolation protocol and transfers to Dream without retuning at 3.18x-3.53x. These results show that native confidence and persistence enable reliable residual acceptance, reducing repeated denoising computation while preserving task quality.
Chinese Translation
扩散语言模型(DLMs)能够并行更新多个标记,但实际的解码器通常使用固定的去噪时间范围。许多预测在早期就已稳定,但块状解码仍持续进行,直到所有位置都得到解决,这导致重复的密集前向传递。现有的加速器通常依赖于学习的过滤器、修改的分数、依赖模型或特定缓存机制。我们探讨原生轨迹信号是否能够识别可能匹配确定性密集终点的残余位置。我们提出了CORA-Diff,这是一种无训练的方法,保留了原始转移规则,仅对未解决的位置应用信心和持久性门控。被接受的标记作为上下文保持可见,块在所有位置解决后终止。这不需要更改主干结构、学习接受模型或修改logit。我们的理论解释了为什么高信心、持久的预测更可能匹配固定时间范围的密集终点,并且配对的干预后轨迹提供了直接的实证支持。我们在一个单独的GSM8K校准子集上选择了一个操作点,并将其冻结用于所有评估。在匹配的Learn2PD风格的LLaDA协议下,CORA-Diff在所有八个任务长度设置中测得的运行时间最低。在五个设置中,任务得分与密集解码相匹配或超过,观察到的最大下降为1.22分。与考虑EOS的密集解码相比,其增量加速分别为GSM8K和HumanEval上的2.70倍和3.32倍。在固定时间范围1024/1024机制隔离协议下,它还达到了13.14倍的加速,并在不重新调优的情况下以3.18倍至3.53倍的速度转移到Dream。这些结果表明,原生信心和持久性使得可靠的残差接受成为可能,减少了重复的去噪计算,同时保持了任务质量。
cs.AI / 20 / 2608.11237

Geometry-aware Incremental Neural Operator for Long-Horizon PDE prediction

几何感知增量神经算子用于长时间跨度的偏微分方程预测
Zhang, Jiaquan, Chen, Shuxu, Meng, Haifan, Lu, Yi, Lyu, Zhihan, Mo, Fan, Dong, Wei, Yang, Yang, Zhang, Chaoning
Abstract
Neural operators have shown strong potential for learning solution operators of partial differential equations (PDEs). However, long-horizon autoregressive prediction remains challenging: local errors accumulate as spectral inconsistency, phase misalignment, or mean drift. Existing methods mainly improve state representations and operator backbones, while leaving the repeatedly applied latent transition increment weakly structured, allowing spectral errors and unstable channel couplings to accumulate during rollout. To address these issues, we propose a geometry-aware incremental neural operator (GeoIncNO) for stable long-horizon PDE prediction. GeoIncNO predicts latent increments for residual advancement and uses lightweight low-rank projectors to regulate channel coupling within active frequency bands derived from the increment spectral energy distribution. To reduce physical-space reconstruction errors, GeoIncNO further introduces a mean--fluctuation decoupled reconstruction mechanism, where stable mean structures and dynamic fluctuations are fused separately, and phase correction is applied only to the zero-mean fluctuation component. Extensive experiments on six PDE benchmarks, covering 1D, 2D, and 3D dynamical systems, show that GeoIncNO achieves consistently strong prediction accuracy, improved rollout stability, and better spectral fidelity compared with competitive neural-operator baselines.
Chinese Translation
神经算子在学习偏微分方程(PDE)解算子方面展现了强大的潜力。然而,长时间跨度的自回归预测仍然面临挑战:局部误差会累积为谱不一致、相位错位或均值漂移。现有方法主要改善状态表示和算子骨架,而对反复应用的潜在过渡增量结构较为薄弱,导致谱误差和不稳定的通道耦合在展开过程中累积。为了解决这些问题,我们提出了一种几何感知增量神经算子(GeoIncNO),旨在实现稳定的长时间跨度PDE预测。GeoIncNO预测潜在增量以推进残差,并使用轻量级低秩投影器来调节从增量谱能量分布中导出的活跃频带内的通道耦合。为了减少物理空间重构误差,GeoIncNO进一步引入了一种均值-波动解耦重构机制,其中稳定的均值结构和动态波动分别融合,并且相位校正仅应用于零均值波动成分。在六个PDE基准测试上的大量实验,涵盖1D、2D和3D动态系统,表明GeoIncNO在预测准确性、展开稳定性和谱保真度方面均优于竞争性神经算子基线。
cs.AI / 21 / 2608.11238

Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability

基于查询覆盖和主张可验证性的查询无关RAG评估方法
Choi, Jeonghwan, Yun, Taewon, Ban, Minjeong, Sun, Gyeonghun, Lee, Jae-Gil, Song, Hwanjun
Abstract
Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide consistent, fine-grained diagnostics across the diverse spectrum of user queries, ranging from close-ended fact-seeking to open-ended explanatory requests. We propose Q-CARE, a query-agnostic and fully reference-free framework that enables fine-grained assessment by decomposing queries into sub-queries and answers into atomic claims. Q-CARE establishes a unified evaluation principle based on query coverage and claim verifiability, yielding coverage-aware retriever metrics (C-Prec@k, C-nDCG@k) and claim-level generator metrics (Completeness, Conciseness, and Verifiableness). On a human-annotated benchmark spanning eight datasets, Q-CARE achieves higher correlation with human judgments than four existing RAG evaluation metrics, including RAGEval and RAGChecker, proving its effectiveness as a reliable, automated evaluation framework. Code and data are publicly available at https://github.com/DISL-Lab/Q-CaRE-COLM-26.
Chinese Translation
检索增强生成通过将响应基于检索到的证据进行基础化,提升了大型语言模型的事实性。然而,现有的评估框架在应对用户查询的多样性时,难以提供一致且细致的诊断,这些查询从封闭式的事实寻求到开放式的解释请求不等。我们提出了Q-CARE,一个查询无关且完全不依赖参考的框架,通过将查询分解为子查询,将答案分解为原子主张,从而实现细致的评估。Q-CARE建立了基于查询覆盖和主张可验证性的统一评估原则,生成了覆盖感知的检索指标(C-Prec@k,C-nDCG@k)和主张级生成指标(完整性、简洁性和可验证性)。在一个涵盖八个数据集的人类标注基准上,Q-CARE与人类判断的相关性高于四个现有的RAG评估指标,包括RAGEval和RAGChecker,证明了其作为可靠的自动评估框架的有效性。代码和数据可在 https://github.com/DISL-Lab/Q-CaRE-COLM-26 获取。
cs.AI / 22 / 2608.11240

VQ-bench: A Composable Vector Quantization Framework

VQ-bench:一个可组合的向量量化框架
Padaki, Ashwin, Ingber, Amir, Liberty, Edo
Abstract
Vector quantization is an old problem but has recently become central to AI infrastructure. It is therefore experiencing a surge of renewed engineering and research activity. This paper provides a unified framework for developing and benchmarking new quantization algorithms. We describe 7 common conceptual quantization primitives and show how to compose them arbitrarily. We then re-express 25 common quantizers as pipelines of these primitives. Finally, we publish VQ-bench as open-source to be extended further and make reproducible benchmarks publicly available.
Chinese Translation
向量量化是一个古老的问题,但最近已成为人工智能基础设施的核心。因此,它正在经历一波新的工程和研究活动的浪潮。本文提供了一个统一的框架,用于开发和基准测试新的量化算法。我们描述了7个常见的概念量化原语,并展示了如何任意组合它们。然后,我们将25个常见的量化器重新表达为这些原语的管道。最后,我们将VQ-bench发布为开源,以便进一步扩展并使可重复的基准测试公开可用。
cs.AI / 23 / 2608.11241

RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle

推荐系统工厂:在工业推荐生命周期中限制大语言模型代理的自主性至决策点
Ao, Dongyang, Fang, Kaixiang, Xu, Shijie
Abstract
Deploying LLM agents into industrial recommender operations exposes a three-way tension we frame as the autonomy-determinism-efficiency trilemma: general autonomy (interpreting operator intent, generating glue code zero-shot), industrial determinism (schema-conforming feature extraction, non-crashing A/B, zero compliance-path hallucination), and end-to-end efficiency. Any two can be maximized against the third. We present RecSys Factory, an LLM-agent platform deployed for 78 days across three heterogeneous Tencent recommender business lines. The design principle is autonomy at decision points, not over pipelines, made concrete through three deconstructions that each discharge one vertex of the trilemma. Runtime is deconstructed into three host-emitted event sources (Claude Code Stop hooks, corporate-IM webhooks, workflow scheduler APIs): the platform carries no long-running daemon during the wait phase and consumes zero CPU during the 94% of wall-clock spent waiting on Spark or GPU jobs. Capability is deconstructed into a 29-file skill ecosystem (8,971 lines of SKILL.md) whose per-skill pitfall tables mechanically compile into a 400-entry PitfallStore, confining autonomy to bounded typed decision surfaces inside pre-committed pipelines. Deployment spans three business lines with disjoint label semantics, A/B layer topologies, and operator personas; an onboarding-time compression is observed on two of the three and is reported as a case-study observation, not a generalization claim, and not measured against a controlled pre-platform baseline. The human is retained at the diagnostic-versus-execution boundary via a human-in-the-loop card protocol, deployed as an audit-trail primitive (schema-validated, idempotent, replayable) and reported from an 8-day 16-run pilot. Across the 78-day window the platform recorded 1,624 CLI-tool dispatches at a 78.6% aggregate success rate.
Chinese Translation
将大语言模型(LLM)代理部署到工业推荐操作中暴露出一种三方紧张关系,我们将其框架化为自主性-确定性-效率三难困境:一般自主性(解释操作员意图、生成粘合代码零样本)、工业确定性(符合模式的特征提取、非崩溃的A/B测试、零合规路径幻觉)以及端到端效率。任何两个方面可以在第三个方面上达到最大化。我们提出了推荐系统工厂(RecSys Factory),这是一个在三个异构腾讯推荐业务线中部署了78天的LLM代理平台。设计原则是在决策点实现自主性,而不是在管道上,通过三个解构具体化,每个解构释放三难困境的一个顶点。运行时被解构为三个主机发出的事件源(Claude Code Stop hooks、企业IM网络钩子、工作流调度API):该平台在等待阶段不携带任何长期运行的守护进程,并在94%的墙钟时间内等待Spark或GPU作业时消耗零CPU。能力被解构为一个29文件的技能生态系统(8,971行的SKILL.md),其每个技能的陷阱表机械地编译成一个400条目的PitfallStore,将自主性限制在预先承诺的管道内部的有界类型决策表面。部署跨越三个具有不重叠标签语义、A/B层拓扑和操作员角色的业务线;在三者中的两个观察到入职时间压缩,并作为案例研究观察报告,而非一般化声明,也未与受控的预平台基线进行测量。通过人机交互卡协议在诊断与执行边界保留人类,该协议作为审计跟踪原语(模式验证、幂等、可重放)进行部署,并从一个8天16次运行的试点中报告。在78天的时间窗口内,该平台记录了1,624次CLI工具调度,整体成功率为78.6%。
cs.AI / 24 / 2608.11243

The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification

离支持障碍:为何语义安全约束不是学习问题的不变性,以及这对先前设计、约束和验证的影响
Watanabe, Yoshinori
Abstract
We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off-support object. Formally, if q is the data distribution and \(p(\cdot\mid w)\) the model, the safety predicate B is not measurable with respect to \(\sigma(\text{model}, q)\), whereas the real log-canonical threshold (RLCT) of singular learning theory (SLT) is. From this non-invariance we derive, as corollaries rather than independent observations: (i) why reward hacking and sandbox escape arise under outcome-based optimization; (ii) why encoding such constraints through Bayesian prior design or soft penalty weighting has poor leverage in singular models; (iii) why hard invariants belong in the harness and soft dispositions in the model; (iv) why the same B is nonetheless soundly and locally certifiable by formal verification, exactly as the local learning coefficient (LLC) locally pins the same RLCT --- with two precise points of disanalogy; and (v) why the residual difficulty, identifying which off-support region matters, coincides with performative prediction and self-referential functional dynamics, where SLT's analytic machinery breaks down. We use the July 2026 OpenAI--Hugging Face evaluation incident as the motivating case. Numerical experiments code and related proofs in lean are available at https://github.com/xiangze/Preventing_Jailbreak_as_regularization
Chinese Translation
我们认为一个单一的结构性事实组织了当代人工智能安全中的广泛现象:语义安全约束(例如,智能体不逃离其沙箱)是一个离支持对象。形式上,如果 q 是数据分布,而 p(·|w) 是模型,则安全谓词 B 在 σ(模型, q) 下不可测,而单一学习理论(SLT)的真实对数典范阈值(RLCT)则是可测的。基于这种非不变性,我们推导出以下推论,而不是独立观察:(i)为何在基于结果的优化下会出现奖励黑客和沙箱逃逸;(ii)为何通过贝叶斯先验设计或软惩罚加权来编码这些约束在单一模型中效果不佳;(iii)为何硬不变性应属于约束,而软倾向应属于模型;(iv)为何相同的 B 仍然可以通过形式验证被可靠且局部地证明有效,正如局部学习系数(LLC)局部固定相同的 RLCT——其中有两个精确的不相似点;(v)为何剩余的困难,即识别哪个离支持区域重要,与表现性预测和自指功能动态相吻合,而在这些情况下,SLT 的分析机制失效。我们以2026年7月的OpenAI-Hugging Face评估事件作为动机案例。数值实验代码及相关证明可在 https://github.com/xiangze/Preventing_Jailbreak_as_regularization 获取。
cs.AI / 25 / 2608.11244

BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model

BEST-KAG:通过多模态知识图建模和大型语言模型增强建筑工程标准的问题回答
Lin, Jia-Rui, Guo, Junxi, Chen, Keyin, Pan, Peng
Abstract
Construction standards are critical for building safety and sustainability. Existing standard application workflows rely on keyword-based document retrieval and manual cross-clause interpretation, which cannot reliably support multi-clause reasoning, multimodal knowledge utilization, or traceable clause-level evidence linkage. To address these limitations, this study develops a multimodal knowledge-driven framework that supports question answering on standard knowledge named BEST-KAG (Knowledge-Augmented Generation for Building Engineering STandards). The framework introduces 1) a multimodal knowledge graph (MKG) for unified representation of document hierarchy and heterogeneous standard knowledge with various connections, 2) a rule-LLM hybrid knowledge construction pipeline for scalable multimodal knowledge extraction, creating a large MAG with 251 building engineering standards, 171,652 nodes and 310,914 edges, and 3) a graph-retrieval-based knowledge-augmented generation architecture for clause-grounded and traceable question answering. Experiments demonstrate that BEST-KAG consistently outperforms multiple mainstream LLMs in terms of Expert evaluation, and metrics including BLEU, and ROUGE, with the best improvement up to 74.01% compared to the baselines.
Chinese Translation
建筑标准对于建筑安全和可持续性至关重要。现有的标准应用工作流程依赖于基于关键词的文档检索和手动跨条款解释,这无法可靠地支持多条款推理、多模态知识利用或可追溯的条款级证据链接。为了解决这些局限性,本研究开发了一种以多模态知识为驱动的框架,支持对标准知识的问题回答,命名为BEST-KAG(建筑工程标准的知识增强生成)。该框架引入了1)一个多模态知识图(MKG),用于统一表示文档层级和异构标准知识及其各种连接,2)一个规则-LLM混合知识构建管道,用于可扩展的多模态知识提取,创建了一个包含251个建筑工程标准、171,652个节点和310,914条边的大型MAG,以及3)一个基于图检索的知识增强生成架构,用于基于条款的可追溯问题回答。实验表明,BEST-KAG在专家评估和包括BLEU和ROUGE在内的多项指标上始终优于多种主流LLM,最佳改进幅度达到74.01%,相较于基线表现。
cs.AI / 26 / 2608.11245

Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach

迈向在线教育的可持续学习:一种基于强化学习的方法
Zhai, Chaofan, Song, Yicheng, Bapna, Ravi, Ye, Junyao
Abstract
Online education offers unprecedented scalability and accessibility to global learners from diverse backgrounds, but it often suffers from low engagement and poor long term learning effectiveness. To address these challenges, we introduce AI Tutor, a reinforcement learning based model designed to promote sustainable learning by optimizing both short and longterm learning outcomes. In the short term, AI-Tutor draws on cognitive theory to guide learners through a balance of acquiring new knowledge and reinforcing prior learning. In the long term, it models learner engagement to inform strategies that sustain motivation and reduce dropout. These enhancements enable AI-Tutor to provide personalized guidance that fosters both effective learning and sustained participation. Empirical evaluations on 23 million learning records from 33,700 learners show that AI Tutor consistently outperforms state-of-the-art baselines across engagement, knowledge retention, and final learning outcomes. Learning path analyses further reveal how AI-Tutor adapts its strategies to learners with diverse profiles, offering adaptive and human-centered support.
Chinese Translation
在线教育为来自不同背景的全球学习者提供了前所未有的可扩展性和可及性,但往往面临低参与度和长期学习效果不佳的问题。为了解决这些挑战,我们提出了AI Tutor,这是一种基于强化学习的模型,旨在通过优化短期和长期学习成果来促进可持续学习。在短期内,AI Tutor借鉴认知理论,引导学习者在获取新知识和巩固先前学习之间取得平衡。在长期内,它对学习者的参与度进行建模,以制定维持动机和减少辍学的策略。这些增强功能使AI Tutor能够提供个性化指导,促进有效学习和持续参与。对来自33,700名学习者的2300万条学习记录的实证评估表明,AI Tutor在参与度、知识保留和最终学习成果方面始终优于最先进的基准。学习路径分析进一步揭示了AI Tutor如何根据不同特征的学习者调整其策略,提供适应性和以人为本的支持。
cs.AI / 27 / 2608.11246

Towards the Harness of Embodied Agents

迈向具身智能体的利用
Wang, Qi, Wang, Tianyi, Li, Chengyang, Ban, Shikun, Chen, Yurun, Ge, Yizhong, Qin, Jason, Li, Chengtai, Zhu, Wentao
Abstract
The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the infrastructure around it. We ask whether the same paradigm extends to embodied agents in the physical world. We present Thea, a harness in which an agentic loop orchestrates robot capabilities, each wrapped as a callable tool. It inherits the core components of coding agents, modified as the physical world requires. The world, however, withholds two abilities that software grants for free: reading the state of the world, and judging the outcome of an action. To bridge these gaps, Thea introduces Scene Graph as Context, a persistent, symbolic representation of the world, and Evaluation as Exit Codes, which detects when an action should terminate, judges whether it succeeded, and on failure diagnoses the cause. Together they close the loop between the agent and the physical world. Rich behaviors then emerge from the composition of tools, and the closed loop carries long-horizon tasks to completion in real environments.
Chinese Translation
编码智能体的成功确立了利用作为一种范式:智能体的成就不仅依赖于模型本身,还依赖于其周围的基础设施。我们探讨这一范式是否同样适用于物理世界中的具身智能体。我们提出了Thea,一个利用智能循环协调机器人能力的框架,每个能力都被封装为可调用的工具。它继承了编码智能体的核心组件,并根据物理世界的要求进行了修改。然而,物理世界却缺乏软件所免费提供的两种能力:读取世界状态和判断行动结果。为了弥补这些缺口,Thea引入了“场景图作为上下文”(Scene Graph as Context),这是对世界的持久符号表示,以及“评估作为退出代码”(Evaluation as Exit Codes),它检测何时应终止行动,判断行动是否成功,并在失败时诊断原因。它们共同闭合了智能体与物理世界之间的循环。丰富的行为随后从工具的组合中涌现,而闭合的循环则在真实环境中完成长时间跨度的任务。
cs.AI / 28 / 2608.11247

Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier

大型语言模型中的一致性缓解措施位于单一的抵抗-接受前沿
Hussain, Zafar, Nielbo, Kristoffer
Abstract
Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs. Each agent sees what the others assert before it answers, so peer opinion competes with the model's own parametric knowledge, and a wrong majority can overturn an answer the model would otherwise get right. We measure that displacement in 23 open-weight models, 19 conditions, and three datasets, yielding more than a million graded responses. A unanimous wrong majority reverses 22.8% of a model's correct MMLU answers, 54.8% on GPQA and 71.0% on SimpleQA, and 84-89% of the reversed answers match the peers' answers. Existing mitigations aim to increase Resistance, the rate at which a model keeps its correct answer under this pressure, which is only half of what a collaborating agent needs. We pair it with Receptivity, the rate at which a model adopts a correct peer answer after initially answering incorrectly. We score six methods on both axes, four drawn from prior work and two of our own. Each gains Resistance only by losing Receptivity, and their means fall on a single Resistance-Receptivity frontier with $R^2$ between 0.80 and 0.90. Reflection, the strongest published method, gains 7.9 points of MMLU Resistance and gives up 15.3 of Receptivity. Reasoning is the one exception. On GPQA and SimpleQA it trades like the rest, but on the MMLU subjects whose answers a model can derive for itself it raises Resistance by 7.2 points and Receptivity by 9.6 at once, the only intervention we find that improves both.
Chinese Translation
近期语言模型的进展使得多个模型能够在协作环境中相互利用各自的能力,迭代地改进、转化和扩展彼此的输出。每个代理在回答之前都会看到其他代理的主张,因此同伴的意见与模型自身的参数知识相竞争,错误的多数意见可能会推翻模型本来能够正确给出的答案。我们在23个开放权重模型、19种条件和三个数据集上进行了测量,产生了超过一百万个评分响应。一个一致的错误多数会逆转22.8%的模型正确的MMLU答案,在GPQA上为54.8%,在SimpleQA上为71.0%,并且84-89%的被逆转答案与同伴的答案相匹配。现有的缓解措施旨在提高抵抗力,即模型在这种压力下保持正确答案的比例,但这仅是协作代理所需的能力的一半。我们将其与接受度相结合,即模型在最初回答错误后采纳正确同伴答案的比例。我们在两个维度上对六种方法进行了评分,其中四种来自于先前的研究,两个是我们自己的方法。每种方法在获得抵抗力的同时都失去了接受度,它们的均值落在一个单一的抵抗-接受前沿上,$R^2$在0.80到0.90之间。反思(Reflection),作为已发布的最强方法,获得了7.9分的MMLU抵抗力,但放弃了15.3分的接受度。推理(Reasoning)是唯一的例外。在GPQA和SimpleQA上,它的表现与其他方法相似,但在MMLU上,对于模型能够自行推导的答案,它同时提高了7.2分的抵抗力和9.6分的接受度,这是我们发现的唯一一种同时改善两者的干预措施。
cs.AI / 29 / 2608.11248

EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents

EvoGraph-Mem:面向失败的可编辑图记忆用于长期语言代理
Qian, Yuxi, Ren, Yuxiang
Abstract
Long-term memory is essential for language agents operating across extended interactions and evolving tasks. Existing memory-augmented agents mainly focus on storing and retrieving past experience, but the quality of stored memories may degrade over time. In particular, previously distilled insights can become outdated, over-generalized, or harmful under new task contexts, causing memory pollution when repeatedly reused. To address this issue, we study insight-level memory maintenance for long-term language agents and propose a failure-aware memory maintenance framework based on an editable insight graph. Each insight node tracks positive evidence, negative evidence, and an activation state, enabling the agent to distinguish reusable insights from conflicting or invalid ones. We further introduce a utility-aware retrieval mechanism and a graph controller that updates the memory graph after task execution by keeping reliable insights, archiving invalid ones, revising outdated ones, and adding newly discovered reusable insights. Extensive experiments show that our method consistently outperforms representative memory-based agent baselines across different backbone models. Ablation studies further demonstrate that append-only memory is insufficient for long-horizon tasks, while evidence-aware retrieval and graph-level editing improve memory reliability and downstream task performance.
Chinese Translation
长期记忆对在扩展交互和不断演变的任务中运行的语言代理至关重要。现有的增强记忆代理主要关注于存储和检索过去的经验,但存储记忆的质量可能随时间而下降。特别是,之前提炼的见解在新的任务背景下可能变得过时、过于泛化或有害,导致在重复使用时产生记忆污染。为了解决这个问题,我们研究了长期语言代理的见解级记忆维护,并提出了一种基于可编辑见解图的面向失败的记忆维护框架。每个见解节点跟踪正面证据、负面证据和激活状态,使代理能够区分可重用的见解与冲突或无效的见解。我们进一步引入了一种实用性意识的检索机制和一个图控制器,该控制器在任务执行后更新记忆图,通过保留可靠的见解、归档无效的见解、修订过时的见解以及添加新发现的可重用见解。大量实验表明,我们的方法在不同的基础模型上始终优于代表性的基于记忆的代理基线。消融研究进一步表明,仅附加记忆对于长期任务是不够的,而证据意识的检索和图级编辑提高了记忆的可靠性和下游任务的表现。
cs.AI / 30 / 2608.11250

AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search

AgonAlpha:通过提示经济和可扩展代理搜索实现自主阿尔法发现
Ye, Weicheng, Sun, Youran, Ren, Xingyu, Yu, Shunyao, Yi, Chugang, Yang, Haizhao
Abstract
Language models can propose many plausible trading factors, but an autonomous research system must also allocate its evaluation budget, verify its own evidence, and preserve how each candidate was produced. We present AgonAlpha, an architecture that searches over frozen research artifacts---hypotheses, executable expressions, platform evidence, rationales, and review status---rather than formulas alone. To our knowledge, AgonAlpha is the first alpha-mining system to combine verified artifact search, a fresh-context adversarial reviewer with re-execution and veto authority, and pending-aware parallel budget allocation, together with a complete public evidence trail. Independent deployments on WorldQuant BRAIN produced SPECTACULAR-grade alphas across five users and six model backends, with Fitness reaching 9.50 and Sharpe reaching 3.48, while retaining prompt-to-expression provenance for every submission.
Chinese Translation
语言模型可以提出许多合理的交易因素,但一个自主研究系统还必须分配其评估预算,验证自身证据,并保留每个候选项的生成过程。我们提出了AgonAlpha,这是一种在冻结的研究文献(假设、可执行表达式、平台证据、理由和审查状态)上进行搜索的架构,而不仅仅是公式。根据我们的了解,AgonAlpha是第一个结合了经过验证的文献搜索、具有重新执行和否决权的新环境对抗审查者,以及待处理的并行预算分配的阿尔法挖掘系统,并提供完整的公共证据链。在WorldQuant BRAIN上的独立部署产生了五个用户和六个模型后端的SPECTACULAR级阿尔法,Fitness达到9.50,Sharpe达到3.48,同时为每个提交保留了提示到表达的来源。
cs.AI / 31 / 2608.11252

Local verification cannot detect non-transportability: a cohomological theory of context preservation in agentic reasoning

局部验证无法检测非可传输性:代理推理中上下文保留的同调理论
Mishra, Suyash
Abstract
Agentic AI systems routinely transport conclusions across biological, clinical and financial contexts, and the emerging safeguard is local verification: checking at each step that the entity is representable in the chosen tool, that parameters are compatible, and that outputs cohere with the plan. We prove this class of safeguard is structurally incomplete. Modelling a covering of context space by its nerve and evidence by a real-valued 1-cochain, an agent chaining evidence performs path integration: its conclusion is path-independent if and only if the cochain is exact, and disagreement between valid reasoning paths is exactly the holonomy of a first Cech cohomology class. Hodge decomposition partitions evidence conflict into a gradient part (calibration), a curl part (local inconsistency, visible at triple overlaps) and a harmonic part. Our central result is that no family of simplex-supported consistency checks can distinguish omega from omega+h for harmonic h, which nonetheless generates non-zero disagreement between valid paths; detection requires a statistic on a cycle basis. The resulting procedure, Ksetra, estimates by coboundary projection and gates abstention on the harmonic component, which we give a mechanism: it arises from effect modification combined with overlap-specific population composition, and vanishes to machine precision when effect modification is absent. The degrees of freedom of an evidence network partition into calibration, coherence and transport, yielding an exact F-test for the existence of a global claim; we quantify its distortion under unequal precision and supply the precision-whitened form that restores exactness. Foreign exchange, where the arbitrage-free null makes the cochain exactly a coboundary, serves as a calibration bench: the test is correctly sized, fires on loop arbitrage, and ignores triangular arbitrage.
Chinese Translation
代理人工智能系统通常在生物、临床和金融上下文中传递结论,而新兴的安全措施是局部验证:在每一步检查实体是否可以在所选工具中表示,参数是否兼容,以及输出是否与计划一致。我们证明这一类安全措施在结构上是不完整的。通过其神经元对上下文空间进行覆盖建模,并通过实值1-链表示证据,代理在证据链中执行路径积分:当且仅当链是精确的,其结论是路径无关的,而有效推理路径之间的不一致恰好是第一Cech同调类的平行性。Hodge分解将证据冲突分为梯度部分(校准)、旋转部分(局部不一致,出现在三重重叠处)和谐部分。我们的核心结果是,没有任何一类简单形支持的一致性检查能够区分ω和ω+h(其中h为谐波),而这仍然在有效路径之间产生非零的不一致;检测需要在循环基础上的统计量。由共边投影估计的程序Ksetra,对谐波成分进行门控,提供了一个机制:它源于效应修改与特定重叠的人群组成的结合,并在缺乏效应修改时消失到机器精度。证据网络的自由度分为校准、一致性和传输,产生了一个关于全局声明存在的精确F检验;我们量化了在不等精度下的失真,并提供了恢复精确性的精度白化形式。外汇市场中,无套利零假设使得链恰好是共边界,作为校准基准:该检验大小正确,能在循环套利中触发,并忽略三角套利。
cs.AI / 32 / 2608.11255

Symbolic Machine Learning for Vapor-Liquid Equilibrium Prediction in Cx-N2 Binary Mixtures

用于Cx-N2二元混合物蒸汽-液体平衡预测的符号机器学习
Kim, Bongseok, Chakraborty, Suman, Huang, Gary, Mathur, Mehek, Lin, Guang, Qiao, Li
Abstract
Accurate prediction of vapor--liquid equilibrium (VLE) for hydrocarbon-nitrogen mixtures remains challenging for cubic equations of state, particularly across broad ranges of composition and hydrocarbon chain length. While deep learning models can provide accurate predictions, they often lack interpretability and explicit analytical expressions. In this work, we propose a symbolic machine learning approach to discover interpretable symbolic corrections to Peng-Robinson equation-of-state (PR-EOS) predictions from experimental data. The proposed approach adopts a two-level strategy: symbolic expressions are first identified for individual hydrocarbon systems, after which their coefficients are represented as functions of carbon number to enable accurate prediction across different hydrocarbon systems. The results demonstrate significantly improved prediction accuracy over the original PR-EOS across all hydrocarbon-nitrogen systems. Overall, the proposed approach provides an interpretable symbolic correction framework for improving PR-EOS predictions of hydrocarbon-nitrogen VLE.
Chinese Translation
对于碳氢化合物-氮气混合物,立方状态方程在蒸汽-液体平衡(VLE)预测方面仍然面临挑战,尤其是在广泛的组成和碳氢链长度范围内。尽管深度学习模型能够提供准确的预测,但它们往往缺乏可解释性和明确的解析表达式。在本研究中,我们提出了一种符号机器学习方法,以发现对Peng-Robinson状态方程(PR-EOS)预测的可解释符号修正,这些修正基于实验数据。所提出的方法采用两级策略:首先为单个碳氢化合物系统识别符号表达式,然后将其系数表示为碳原子数的函数,以便在不同的碳氢化合物系统中实现准确预测。结果表明,在所有碳氢化合物-氮气系统中,预测精度显著优于原始的PR-EOS。总体而言,所提出的方法为改善PR-EOS对碳氢化合物-氮气VLE预测提供了一个可解释的符号修正框架。
cs.AI / 33 / 2608.11258

Adaptive Hybrid Particle Swarm Optimization with Gradient Descent

自适应混合粒子群优化与梯度下降
Gurudeo, Aryan
Abstract
Gradient injection helps Particle Swarm Optimization (PSO) only when the swarm has identified a basin with smooth local structure, not universally. We propose Adaptive Hybrid PSO (AHPSO), which uses a sigmoid function on swarm diversity to automatically modulate gradient influence: near-zero during exploration, near-maximum during exploitation, with no manual phase-switching. Under budget-normalized comparison (PSO given equivalent total function evaluations), PSO wins 52.5% of 40 configurations versus AHPSO's 20% (p = 7.0e-5, Friedman). AHPSO retains advantage specifically on problems with smooth local basins (F8, F24-F27) where directed descent outperforms undirected sampling even at equal cost. Under iteration-matched comparison across 29 functions (42 configurations, 14,700 runs), AHPSO-Adadelta ranks first of 9 methods including CMA-ES (p = 9.75e-4). The contribution is a principled characterization of when gradient injection provides value in swarm-based search, not a claim of universal superiority.
Chinese Translation
梯度注入仅在粒子群优化(PSO)识别到具有平滑局部结构的盆地时才有助于优化,而并非普遍适用。我们提出了自适应混合粒子群优化(AHPSO),该方法利用 sigmoid 函数对群体多样性进行自动调节梯度影响:在探索阶段接近零,在开发阶段接近最大值,无需手动切换阶段。在预算归一化比较下(PSO 在相同总函数评估下),PSO 在 40 种配置中获胜 52.5%,而 AHPSO 仅为 20%(p = 7.0e-5,Friedman)。AHPSO 在具有平滑局部盆地的问题上(如 F8、F24-F27)保持优势,在这些问题中,定向下降的表现优于无方向采样,即使在成本相等的情况下。在 29 个函数(42 种配置,14,700 次运行)下的迭代匹配比较中,AHPSO-Adadelta 在包括 CMA-ES 在内的 9 种方法中排名第一(p = 9.75e-4)。本研究的贡献在于对梯度注入在基于群体的搜索中何时提供价值进行了原则性描述,而非声称其具有普遍优越性。
cs.AI / 34 / 2608.11260

Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning

一瞥、细察与思考:从无训练到自主推理推进视频异常检测
Gao, Shibo, Yang, Peipei, Zhang, Xu-Yao, Huang, Linlin
Abstract
Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals. Existing approaches exhibit a "when-what" dissociation: traditional DNN-based methods localize when anomalies occur but lack semantic understanding, whereas LLM-based methods explain what happens but neglect precise temporal grounding. We attribute this to the absence of a unified reasoning paradigm. Inspired by how humans inspect surveillance videos - glancing globally to form temporal hypotheses, scrutinizing suspicious segments, and thinking iteratively to correct errors - we study this global-to-local paradigm from two perspectives. We first propose Glance then Scrutinize (GtS), a training-free framework using static and dynamic textual guidance for coarse-to-fine anomaly grounding and understanding, balancing accuracy and speed. To break the ceiling imposed by frozen external modules, we further propose a tool-augmented agentic VAD method, where a multimodal large language model learns to invoke a video cropping tool, inspect densely resampled frames, and self-correct mislocalized hypotheses, via cold-start supervised fine-tuning followed by reinforcement learning with a joint answer-grounding reward. For training and evaluation, we extend our prior VAGU benchmark into VAGU-T (Video Anomaly Grounding, Understanding, and Thinking), comprising 7,567 real-world videos over 21 anomaly categories with human-validated grounding, explanations, QA pairs, and chain-of-thought tool-calling traces. We further introduce JeAUG, a metric jointly evaluating semantic interpretability and temporal precision. Experiments show that GtS substantially surpasses training-free baselines, while the agentic model delivers both higher accuracy and faster inference.
Chinese Translation
视频异常检测(VAD)的目标是识别异常事件并定位其时间区间。现有方法存在“何时-何物”的分离:传统的基于深度神经网络(DNN)的方法能够定位异常发生的时间,但缺乏语义理解;而基于大型语言模型(LLM)的方法能够解释发生了什么,但忽视了精确的时间定位。我们将此归因于缺乏统一的推理范式。受到人类检查监控视频的启发——首先全局浏览以形成时间假设,然后细致检查可疑片段,最后通过迭代思考来纠正错误——我们从两个角度研究这一全局到局部的范式。我们首先提出了“先一瞥再细察”(Glance then Scrutinize, GtS),这是一个无训练的框架,利用静态和动态文本指导进行粗到细的异常定位和理解,平衡准确性与速度。为了打破冻结外部模块所带来的限制,我们进一步提出了一种工具增强的自主视频异常检测方法,其中多模态大型语言模型学习调用视频裁剪工具,检查密集重采样的帧,并通过冷启动监督微调后结合联合答案定位奖励的强化学习自我纠正错误定位的假设。为了进行训练和评估,我们将之前的VAGU基准扩展为VAGU-T(视频异常定位、理解与思考),包含21个异常类别的7,567个真实世界视频,具有经过人工验证的定位、解释、问答对和思维链工具调用轨迹。我们进一步引入了JeAUG,一个共同评估语义可解释性和时间精确性的指标。实验表明,GtS在无训练基线上显著超越,而自主模型则在准确性和推理速度上均表现更佳。
cs.AI / 35 / 2608.11323

Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

部署决策可靠性:用于长期代理评估规模化的一般化理论框架
Srinivasan, Vasundra
Abstract
Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany, $\tau^2$-bench, and AppWorld), that the agent main effect accounts for less than 3% of total variance in every dataset and check type, while the agent-by-task interaction accounts for 7-23%. Leaderboards rank specialization, not capability. We arrive at this through a four-facet Generalizability Theory variance decomposition, fit with three estimators (Henderson Method-I, REML via lme4, and a Bayesian binomial GLMM) that agree to three decimal places. Four further findings sharpen what the leaderboard is hiding. First, aggregate reliability collapses on the hardest task quartile: $E\rho^2$ on $\tau^2$ action_checks falls from 0.752 to 0.000. Second, training-cell reliability negatively correlates with held-out reliability ($r = -0.90$ on $\tau^2$), meaning the designs that look most reliable replicate worst. Third, population-level diagnostics transfer across enterprise benchmarks (capability-gap ratio stable at 0.35-0.40) but per-family agent rankings invert. Fourth, on the MAST failure taxonomy, trace-level mode profiles are idiosyncratic (MAE = 0.261) while cell-level profiles generalise (MAE = 0.056, $r = 0.83$). We package these into Deployment Decision Reliability (DDR), a one-page reporting discipline that turns the variance-component table into five decisions an enterprise buyer can defend. All code, data loaders, and fit artifacts are released under an open-source license.
Chinese Translation
企业从业者将代理排行榜视为代理能力的排名。我们在三个开放的代理追踪基准(TheAgentCompany、$ au^2$-bench 和 AppWorld)中展示,代理的主要效应在每个数据集和检查类型中占总方差的比例低于3%,而代理与任务的交互效应占7-23%。排行榜排名的是专业化,而非能力。我们通过四个方面的一般化理论方差分解得出这一结论,并使用三种估计方法(Henderson Method-I、通过 lme4 的 REML 和贝叶斯二项 GLMM)进行拟合,结果一致到小数点后三位。进一步的四项发现揭示了排行榜隐藏的信息。首先,整体可靠性在最难任务的四分位数上崩溃:在 $ au^2$ 的 action_checks 上,$E ho^2$ 从 0.752 降至 0.000。其次,训练单元的可靠性与保留的可靠性呈负相关(在 $ au^2$ 上 $r = -0.90$),这意味着看似最可靠的设计反而复制效果最差。第三,人口级别的诊断在企业基准之间转移(能力差距比率稳定在 0.35-0.40),但每个家庭的代理排名则相反。第四,在 MAST 失败分类法中,追踪级别的模式特征是特异性的(MAE = 0.261),而单元级别的特征则具有普遍性(MAE = 0.056,$r = 0.83$)。我们将这些整合为部署决策可靠性(DDR),这是一种将方差成分表转化为企业买方可以辩护的五个决策的一页报告规范。所有代码、数据加载器和拟合文档均以开源许可证发布。
cs.AI / 36 / 2608.11341

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

Apodex发现:评估和构建发现性人工智能的现实基准和环境
Wang, Brian, Feng, Bin, Pan, Xiaoman, An, Chenyang, Liu, Felix, Fang, Tangqi, Sun, Gongbo, Shen, Lingfeng, Wang, Ning, Zhang, Handuo, Chen, Feng, Yang, Fuchao, Wang, Xiang, Lin, Jiacheng, Li, Siting, Liu, Zixuan, Han, Chi, Wang, Zhenhailong, Zhu, Kunlun, Zhao, Lawrence, Guo, Yueqi, Wen, Kailong, Xing, Feng, Guo, Yiling, Bing, Lidong, Tan, David, An, Bo, Ji, Heng, Wang, Sheng
Abstract
Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a similar transition: frontier models can solve difficult tasks once the problem, tools, and success criteria are specified, yet consequential real-world challenges rarely arrive in an executable or verifiable form. We introduce Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations. It has three core components. First, a problem-scouting process surveyed 561 industries across 16 sectors, assembled 423 high-value real-world problems, and selected 20 for the initial release. Second, a common environment-task-episode abstraction provides data, tools, constraints, feedback, trajectory recording, and verification of intermediate artifacts and final submissions. Third, HDS6 evaluates Tools, Repair, Alternatives, Coherence, Evidence, and Scope independently of final-task success. In AAV capsid design, Apodex surpassed the published state of the art by 7% across viability, tropism, structure prediction, and generative design. In drug repurposing and reformulation, a task-specific biomedical environment improved the mean normalized prediction score of GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points over the same closed-book backbone. Controlled ablations show that the fixed TRACES episode interface enables attribution of performance differences to specific solver components. Apodex Discovery moves AI evaluation beyond predefined benchmarks toward verifiable investigations aimed at genuine discovery.
Chinese Translation
阿波罗之所以能够到达月球,并不仅仅因为其工程师能够解决复杂的方程。它的成功在于将一个遥远的目标转化为明确的任务架构,包括目标设定、模拟、验证和反复修正。人工智能现在面临着类似的转变:前沿模型能够在问题、工具和成功标准明确的情况下解决复杂任务,但重要的现实世界挑战往往以不可执行或不可验证的形式出现。我们引入了Apodex Discovery,这是一个通过重型求解器构建和评估发现性人工智能的框架,该系统包括基础模型、工具、控制策略,旨在进行扩展的、有状态的、可验证的研究。它有三个核心组成部分。首先,问题勘探过程调查了16个行业中的561个行业,汇总了423个高价值的现实问题,并选择了20个进行初步发布。其次,通用环境-任务-情节抽象提供数据、工具、约束、反馈、轨迹记录以及中间工件和最终提交的验证。第三,HDS6独立于最终任务成功评估工具、修复、替代方案、一致性、证据和范围。在AAV衣壳设计中,Apodex在生存性、趋向性、结构预测和生成设计方面超越了已发布的最新技术7%。在药物再利用和重新配方中,特定任务的生物医学环境使GPT-5.5和GPT-5.6-sol的平均标准化预测分数分别提高了2.5和7.6分。受控消融实验表明,固定的TRACES情节接口使得性能差异的归因能够指向特定的求解器组件。Apodex Discovery将人工智能评估从预定义基准推进到旨在真正发现的可验证研究。
cs.AI / 37 / 2608.11343

Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval

前沿大型语言模型能否匹配原生多模态嵌入?在困难负样本文本到图像检索中的比较
Dutta, Archan, Kanungo, Vyanktesh
Abstract
Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning. The March 2026 release of Gemini Embedding 2, Google's first natively multimodal embedding model to map text, images, video, audio, and documents into a single shared space, raises competition among multimodal retrieval systems. Simultaneously, frontier Large language models (LLMs) have also demonstrated strong visual understanding, raising the question of whether they can serve as effective zero-shot rankers. Our study provides the first direct comparison of native multimodal embeddings against LLM-based visual ranking on Flickr30k. We observe that GPT-4.1 and Claude Sonnet 4.6 perform on par with Gemini Embedding 2. Additionally, once embeddings are precomputed, multimodal embeddings are better suited for low-latency applications.
Chinese Translation
跨越文本、图像、视频和音频等不同媒体类型的多模态检索和分类传统上依赖于通过对比学习对齐视觉和文本表示的双编码器模型。2026年3月发布的Gemini Embedding 2是谷歌首个原生多模态嵌入模型,能够将文本、图像、视频、音频和文档映射到一个共享空间,这引发了多模态检索系统之间的竞争。同时,前沿大型语言模型(LLMs)也展现出强大的视觉理解能力,这引发了它们是否能够作为有效的零样本排序器的疑问。我们的研究首次直接比较了原生多模态嵌入与基于LLM的视觉排序在Flickr30k上的表现。我们观察到,GPT-4.1和Claude Sonnet 4.6的表现与Gemini Embedding 2相当。此外,一旦嵌入被预计算,多模态嵌入更适合低延迟应用。
cs.AI / 38 / 2608.11354

Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces

内容推荐的逆心智理论建模:从网页浏览到动态智能界面
Chen, Mengyu, Lu, Feiyu, Chen, Chun-Fu, Tran, Lucas Vinh, Katukuri, Jay
Abstract
Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UIs and immersive extended reality (XR), the need for deeper, modality-agnostic user understanding grows: these adaptive environments must decide not only what to present but where, when, how prominently, and most importantly why a user acts. We propose an Inverse Theory of Mind (IToM) pipeline that reasons backward from observed interactions to infer the beliefs, preferences, and decision-making traits that explain behavior. The pipeline reconstructs each user's decision context, including what was chosen and what alternatives were available, applies LLM-driven counterfactual reasoning to produce evidence-grounded natural-language belief statements, and synthesizes these beliefs through multi-hypothesis abductive inference into a structured user persona. We evaluate on the OPeRA dataset against ground-truth personality assessments, attitudinal surveys, and interview-based personas across four tasks: next action prediction, shopping attitude alignment, Big Five personality inference, and held-out category prediction. Results show that inferred personas match or exceed ground-truth personas and that multi-hypothesis reasoning is essential for accurate personality prediction. We further demonstrate cross-modal transferability with a persona-driven spatial banking application on VisionOS.
Chinese Translation
现代推荐系统将观察到的行为视为用户偏好的可靠代理,然而,交互往往反映的是探索或比较,而非稳定的偏好表达。随着界面从静态布局演变为生成式用户界面(UIs)和沉浸式扩展现实(XR),对更深层次的、与模态无关的用户理解的需求日益增长:这些自适应环境不仅需要决定展示什么,还需要考虑展示的位置、时间、显著性,以及最重要的,用户行为的原因。我们提出了一种逆心智理论(Inverse Theory of Mind, IToM)管道,从观察到的交互中推理出信念、偏好和决策特征,以解释行为。该管道重建每个用户的决策背景,包括所选择的内容和可用的替代选项,应用基于大语言模型(LLM)的反事实推理生成基于证据的自然语言信念陈述,并通过多假设的溯因推理将这些信念综合成结构化的用户画像。我们在OPeRA数据集上进行评估,针对真实的人格评估、态度调查和基于访谈的人物画像进行四项任务的验证:下一步行动预测、购物态度对齐、五大人格推断和保留类别预测。结果表明,推断的人物画像与真实的人物画像相匹配或超越,并且多假设推理对于准确的人格预测至关重要。我们进一步展示了在VisionOS上基于人物画像的空间银行应用的跨模态可转移性。
cs.AI / 39 / 2608.11381

From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate

从数字到判断:专门化的LLM代理与欧洲上市房地产的强化学习
Taghavi, Pardis, Bhavani, Santosh
Abstract
We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization. Larix maps a 16-lens European listed-real-estate analysis framework to eight lens-aligned specialists; we compare a frontier LLM under monolithic versus specialist-decomposed prompting while holding the model, source evidence, task instructions, output schema, and scoring fixed. Across 19 firms spanning seven regulatory wrappers, decomposition improves the numerical-task aggregate by 15.8 percentage points but does not reliably improve, and can reduce, performance on judgment tasks, a pattern stable across four frozen-template dispatches; a single-agent control given the complete framework does not reproduce the numerical gain. Post-training Qwen3.5-9B with GRPO using task-aligned structured rewards then raises the development-split score by 12.0 points and the judgment aggregate by 14.2 points, with gains on all four sub-ceiling tasks; the gains transfer to unseen firms (+15.2 points overall; +40.4 on covenant stress) and to unseen regulatory wrappers (+4.3), with positive transfer on all three anti-memorization splits. Prompt-level decomposition thus improves modular numerical execution, whereas targeted parameter adaptation improves integrative financial judgment.
Chinese Translation
我们研究了金融分析的局部数值操作和综合判断是否受益于相同形式的LLM专门化。Larix将一个16镜头的欧洲上市房地产分析框架映射到八个镜头对齐的专家;在保持模型、源证据、任务指令、输出模式和评分不变的情况下,我们比较了在单一与专家分解提示下的前沿LLM。在涵盖七个监管框架的19家公司中,分解使数值任务的总得分提高了15.8个百分点,但对判断任务的表现没有可靠的改善,甚至可能降低,这一模式在四个固定模板的派遣中保持稳定;在给定完整框架的情况下,单一代理控制并未再现数值增益。经过训练的Qwen3.5-9B使用任务对齐的结构化奖励后,开发分数提高了12.0分,判断总得分提高了14.2分,所有四个子任务均有增益;这些增益转移到未见过的公司(总体+15.2分;在契约压力上+40.4分)和未见过的监管框架(+4.3),在所有三个反记忆分割上均有正向转移。因此,提示级别的分解改善了模块化的数值执行,而有针对性的参数适应则改善了综合金融判断。
cs.AI / 40 / 2608.11403

When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs

当自我一致性适得其反:多数投票对小型大语言模型的多数硬科学问题造成伤害
Bahuguna, Utkarsh
Abstract
Self-consistency (SC) via majority vote is a widely used way to spend inference-time compute: sample N chains of thought, return the plurality answer. On the full GPQA Diamond benchmark (198 graduate-level science questions), majority voting reduces per-problem accuracy on a majority of problems for two instruction-tuned models from different families: 56.6% of problems for Qwen2.5-7B and 65.7% for Llama-3-8B, with Qwen the primary demonstration and Llama corroborating the direction from a near-chance baseline. The effect was pre-registered on a 151-problem confirmatory split after being observed on 47 exploratory problems, and all four confirmatory hypotheses passed. A grid oracle that routes each problem to the best N across {1, 2, 4, 8, 16, 32, 64} marks a theoretical upper bound 14 accuracy points above N = 1 for Qwen and 17 for Llama, an oracle bound requiring ground truth rather than a deployable method. No verifier-free gate reaches it: neither a plurality-agreement gate nor a token-entropy gate moves accuracy more than 0.002 from fixed-budget voting at N = 64. The mechanism is direct: confidence does not track correctness on these problems. In the highest-agreement bin the plurality answer is correct about half the time for Qwen, and for Llama that bin is less accurate than its lowest-agreement bin. We pre-register and confirm these findings on small instruction-tuned models; we do not test reasoning-native models, which we flag as the central open question.
Chinese Translation
通过多数投票实现的自我一致性(SC)是一种广泛使用的推理时间计算方法:采样 N 条思维链,返回多数答案。在完整的 GPQA Diamond 基准测试(198 道研究生级科学问题)中,对于来自不同模型家族的两个指令调优模型,多数投票在大多数问题上降低了每个问题的准确性:Qwen2.5-7B 模型的 56.6% 问题和 Llama-3-8B 模型的 65.7% 问题,Qwen 是主要的演示,Llama 则从接近随机基线的方向上证实了这一结果。该效应在观察到 47 个探索性问题后,在 151 个问题的确认性分割上进行了预注册,且所有四个确认性假设均通过。一个网格神谕(grid oracle)将每个问题路由到最佳的 N(取自 {1, 2, 4, 8, 16, 32, 64})上,理论上将 Qwen 的准确性上限提高了 14 个点,Llama 提高了 17 个点,而该神谕界限需要真实答案而非可部署的方法。没有无验证门(verifier-free gate)能够达到这一点:无论是多数一致门(plurality-agreement gate)还是令牌熵门(token-entropy gate),其准确性都未能比固定预算投票(N = 64)提高 0.002。该机制是直接的:在这些问题上,置信度并不与正确性相关。在最高一致性区间内,Qwen 的多数答案大约有一半的正确率,而对于 Llama,该区间的准确性甚至低于其最低一致性区间。我们在小型指令调优模型上预注册并确认了这些发现;我们未测试推理原生模型(reasoning-native models),这被我们标记为核心未解问题。
cs.AI / 41 / 2608.11420

Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology

社会思维链:基于医学鉴别诊断方法论的多智能体架构
Coburn, Del, Sanner, Scott, Silver, Dan
Abstract
Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications for the health and wellbeing of users. When OpenAI (2026) reports that more than 5% of ChatGPT messages globally are healthcare-related, the transparency of these systems becomes a serious design concern. This is especially true for complex cases, where differential diagnosis often requires integrating multiple forms of specialist reasoning. Existing work has proposed multi-agent approaches to medical diagnosis, but it remains unclear when such systems are needed, why they help, and where they outperform monolithic inference. We introduce Social Chain of Thought (SCoT),a multi-round pipeline for medical differential diagnosis that structures multi-agent interaction as a deliberative framework for collabora. tive LLM reasoning. Evaluating SCoT against single-agent baselines, one-agent pipeline ablations, and best-of-n scaling, we show that its recall advantage is not reproduced by monolithic inference alone. SCoT is most successful in the hardest diagnostic cases, where multiple rounds of specialist conversation help recover ground-truth diagnoses and converge on a higher-recall differential.
Chinese Translation
医学诊断推理是大型语言模型(LLMs)的一个高影响力应用案例,对用户的健康和福祉具有重要影响。当OpenAI(2026)报告称全球超过5%的ChatGPT消息与医疗相关时,这些系统的透明性成为一个严重的设计问题。这在复杂病例中尤为明显,因为鉴别诊断通常需要整合多种专家推理形式。现有研究提出了多智能体的医学诊断方法,但仍不清楚何时需要此类系统、它们为何有效以及在何种情况下它们优于单一推理。我们引入了社会思维链(Social Chain of Thought, SCoT),这是一种用于医学鉴别诊断的多轮管道,将多智能体互动结构化为协作大型语言模型推理的审议框架。通过将SCoT与单智能体基线、单智能体管道消融和最佳规模进行评估,我们表明其召回优势并非单一推理所能复制。SCoT在最困难的诊断案例中最为成功,多轮专家对话有助于恢复真实诊断并收敛于更高召回率的鉴别结果。
cs.AI / 42 / 2608.11434

Benchmarking LLM Judges for Mobile Agent Evaluation

移动代理评估的 LLM 判别器基准测试
Wan, Ziqiang, Gu, Li, Chi, Zhixiang, Liu, Zhi, Ayyoubzadeh, Seyed Mehdi, Yu, Yuanhao, Wang, Yang
Abstract
Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark for systematically evaluating LLM-as-judge methods on mobile agent trajectories. Our benchmark comprises 931 human-annotated trajectories spanning 6 mobile agent benchmarks, 4 agent models, and 68 apps. We evaluate 6 judge methods (five adapted from SPA-Bench, A3 with two modes, AndroidArena, and AgentRewardBench, plus a simple baseline we design) across multiple LLM backends. Our experiments reveal three key findings. First, a simple baseline judge with sampled screenshots is competitive with, and often exceeds, purpose-built methods, indicating that more elaborate judge pipelines do not consistently improve judge quality; among competitive methods, the LLM backbone is the primary driver. Second, benchmark quality metrics reliably predict real-world judge utility: they correlate with both agent ranking fidelity for evaluation and downstream performance when judges serve as reward signals for on-policy reinforcement learning. Third, failure analysis across two LLM backends uncovers qualitatively opposite failure profiles, one conservative and the other permissive, linked to the backbone's precision-recall characteristics.
Chinese Translation
移动代理基准测试越来越依赖基于 LLM 的判别器来评估任务完成情况,但这些判别器在移动代理轨迹上的可靠性仍然未得到充分检验。我们引入了 MobileJudgeBench,这是一个系统评估 LLM 作为判别器方法在移动代理轨迹上的基准测试。我们的基准测试包含 931 条人类标注的轨迹,涵盖 6 个移动代理基准、4 种代理模型和 68 个应用程序。我们在多个 LLM 后端上评估了 6 种判别方法(五种改编自 SPA-Bench,A3 的两种模式,AndroidArena 和 AgentRewardBench,以及我们设计的一个简单基线)。我们的实验揭示了三个关键发现。首先,使用采样截图的简单基线判别器在竞争中表现良好,且常常超过专门构建的方法,这表明更复杂的判别管道并不总是能提高判别质量;在竞争方法中,LLM 主干是主要驱动因素。其次,基准质量指标可靠地预测了真实世界判别器的效用:它们与评估的代理排名保真度以及当判别器作为策略强化学习的奖励信号时的下游性能相关。第三,对两个 LLM 后端的失败分析揭示了质的相反失败特征,一个是保守的,另一个是宽松的,这与主干的精确度-召回特性相关。
cs.AI / 43 / 2608.11483

A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization

一种模块化的代理框架用于合成约束的多目标命中到先导优化
Idanwekhai, Kelvin P., Kelestemur, Enes, Strickland, Benjamin, Hart, Matthew, Davidsson, Steini, Angelopoulos, Angelos, Alterovitz, Ron, DeLuca, Marcello, Tropsha, Alexander
Abstract
Hit-to-lead optimization requires iterative design of hit analogs across competing potency, selectivity, physicochemical, pharmacokinetic, safety, and synthetic constraints. We present SABLE (Synthetically-accessible Agentic Bayesian Ligand Exploration), an open-source framework that employs natural-language orchestration to guide chemical structure optimization. SABLE uses an LLM to interpret user-defined goals and route tasks, while specialized tools perform reaction-templated analog enumeration, physicochemical and ADMET property prediction, structure-based affinity scoring, and Bayesian optimization. The resulting workflow is a computational twin of the analytical and prioritization stages of the design-make-test-analyze cycle, providing provenance of each numerical output. Across single, and multi-objective optimization studies, SABLE enriches candidate sets for user-defined computational objectives while evaluating only a subset of the enumerated search space. Its modular architecture allows tools and characterization backends to be replaced by editing a simple config file, without modifying operational logic. SABLE provides an extensible decision-support framework for prioritizing synthetically constrained analogs in early-stage drug discovery.
Chinese Translation
命中到先导优化需要在竞争的效力、选择性、物理化学性质、药代动力学、安全性和合成约束之间进行迭代设计。我们提出了SABLE(可合成的代理贝叶斯配体探索),这是一个开源框架,利用自然语言编排来指导化学结构优化。SABLE使用大型语言模型(LLM)来解释用户定义的目标并引导任务,同时专用工具执行反应模板的类似物枚举、物理化学性质和ADMET(吸收、分布、代谢、排泄和毒性)属性预测、基于结构的亲和力评分和贝叶斯优化。最终的工作流程是设计-制造-测试-分析循环中分析和优先排序阶段的计算双胞胎,提供每个数值输出的来源。在单目标和多目标优化研究中,SABLE丰富了用户定义的计算目标的候选集,同时仅评估枚举搜索空间的一个子集。其模块化架构允许通过编辑简单的配置文件来替换工具和表征后端,而无需修改操作逻辑。SABLE为在早期药物发现中优先考虑合成约束的类似物提供了一个可扩展的决策支持框架。
cs.AI / 44 / 2608.11493

From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation

从提示到行为对齐:个性化大语言模型评估推荐的判断
Ziabari, Alireza S., Ellis, Kat, Chan, Colleen, Tong, Ding
Abstract
Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer a promising alternative by predicting user engagement directly from raw text logs, empirical analysis in this study identifies a critical failure mode termed bidirectional rationalization. In a zero-shot setting, LLMs are found to convincingly argue for both positive and negative user engagement outcomes on the exact same item with identical evidence, highlighting the unreliability of off-the-shelf LLMs in predicting user engagement. To resolve this, we develop and apply a sequential behavioral alignment framework pairing fine-tuning with preference optimization over paired correct and counterfactual rationales. Evaluated on real-world homepage interaction logs, this aligned reasoning approach achieves a 32.19\% lift in Macro-F1 score over the zero-shot baseline and matches the production feature-engineered baseline. The results demonstrate that behavioral alignment mitigates bidirectional rationalization while delivering human-interpretable reasoning traces without manual pipeline overhead.
Chinese Translation
传统的离线推荐评估严重依赖于复杂且手动维护的特征管道,这些管道难以扩展。虽然大语言模型(LLMs)通过直接从原始文本日志中预测用户参与度提供了一个有前景的替代方案,但本研究的实证分析识别出一种称为双向合理化的关键失效模式。在零样本设置中,发现LLMs能够令人信服地为同一项目的正面和负面用户参与结果提供相同证据的论据,突显了现成LLMs在预测用户参与度方面的不可靠性。为了解决这一问题,我们开发并应用了一种序列行为对齐框架,将微调与配对正确和反事实理由的偏好优化相结合。在真实的主页交互日志上进行评估,这种对齐推理方法在宏观F1分数上比零样本基线提高了32.19%,并且与生产特征工程基线相匹配。结果表明,行为对齐减轻了双向合理化,同时提供了无需手动管道开销的人类可解释推理轨迹。
cs.AI / 45 / 2608.11583

Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models

安全对齐的局部化:MLP层和中间网络块在大型语言模型中编码拒绝行为
Zong, Mingyu, Mohanty, Sampad, Krishnamachari, Bhaskar
Abstract
Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters. This work addresses where safety-aligned refusal is encoded by transplanting weights from aligned models into matched unaligned base models at multiple levels of granularity. Using two open-weight model pairs and four safety benchmarks, we conducted experiments to compare the effects of replacing attention weights, MLP weights, contiguous layer regions, and MLP blocks. Across both model families, refusal transfer is dominated by MLP weights: replacing MLP parameters recovers substantially more malicious-prompt refusal than replacing attention parameters, with gains of at least 2.7 times more across benchmarks. Within the MLP stack, refusal-relevant parameters exhibit a consistent mid-network concentration, as the block spanning layers 8-11 is selected first in all six greedy searches over model-dataset pairs. The results also show that the composition of safety-relevant components is non-additive: in five of six greedy trajectories, adding more aligned blocks can reduce refusal performance, and selective block subsets can outperform full MLP transplantation on malicious refusal, benign over-refusal, or both. Finally, greedy orders transferred to OR-Bench vary with the source benchmark used to derive them, indicating a benchmark-dependent precision-coverage trade-off. These results suggest that safety alignment in current LLMs is both localized and interaction-sensitive, offering insight into alignment brittleness and potential avenues for targeted safety interventions.
Chinese Translation
大型语言模型中的安全对齐通常被视为整个网络的分布特性,然而其实际脆弱性表明拒绝行为可能集中在一小部分参数中。本研究通过将对齐模型的权重移植到多个粒度级别的匹配未对齐基础模型中,探讨了安全对齐拒绝行为的编码位置。我们使用了两个开放权重模型对和四个安全基准,进行了实验以比较替换注意力权重、MLP权重、连续层区域和MLP块的效果。在这两类模型中,拒绝转移主要由MLP权重主导:替换MLP参数比替换注意力参数恢复了显著更多的恶意提示拒绝,在各基准中至少提高了2.7倍。在MLP堆栈中,与拒绝相关的参数表现出一致的中间网络集中性,因为在对模型-数据集对的六次贪婪搜索中,跨越第8至11层的块首先被选中。结果还表明,安全相关组件的组成是非加性的:在六条贪婪轨迹中的五条中,添加更多对齐块可能会降低拒绝性能,而选择性的块子集在恶意拒绝、良性过度拒绝或两者上可能优于完整的MLP移植。最后,转移到OR-Bench的贪婪顺序随着用于推导它们的源基准而变化,表明存在基准依赖的精度-覆盖权衡。这些结果表明,当前大型语言模型中的安全对齐既是局部化的,又对交互敏感,为对齐脆弱性和针对性安全干预的潜在途径提供了洞见。
cs.AI / 46 / 2608.11584

EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval

EnterpriseRAG:在非理想企业检索下评估大型语言模型的指令遵循性和鲁棒性
Miao, Huiqi, Sun, Xinbao, Wang, Bo, Meng, Fanyu, Mei, Lijun, Wu, Na, Jin, Di, Deng, Chao, Feng, Junlan
Abstract
Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. Existing benchmarks assume clean retrieval with simple queries, failing to capture production conditions where noisy documents and multi-dimensional constraints coexist. We introduce EnterpriseRAG, a benchmark of 983 expert-validated samples across six domains that systematically simulates three failure modes absent from prior work: retrieval noise, knowledge gaps, and factual conflicts, coupled with complex instructions. Evaluation of 13 state-of-the-art LLMs reveals a severe instruction adherence collapse, where high per-constraint satisfaction masks low holistic compliance. Critical findings expose deep barriers under knowledge gaps and factual conflicts, even with reasoning-enhanced inference, indicating production RAG requires explicit context-aware protocols and calibrated judgment. EnterpriseRAG provides a reproducible foundation for measuring and closing these gaps, directly informing deployment decisions for enterprise-scale RAG systems. We will release the benchmark and evaluation framework upon publication.
Chinese Translation
企业RAG的部署面临着一个关键的可靠性缺口:尽管大型语言模型(LLMs)满足80%的单个约束,但仅有26.8%的响应同时满足所有要求,揭示出57点的协调缺口。现有基准假设检索过程干净且查询简单,未能捕捉到生产环境中噪声文档和多维约束共存的情况。我们引入了EnterpriseRAG,这是一个涵盖六个领域的983个专家验证样本的基准,系统性地模拟了以往研究中缺失的三种失败模式:检索噪声、知识缺口和事实冲突,以及复杂的指令。对13种最先进的LLM的评估揭示了严重的指令遵循崩溃现象,高每个约束的满足率掩盖了整体合规性低下的事实。关键发现揭示了在知识缺口和事实冲突下的深层障碍,即使在增强推理的情况下也如此,这表明生产RAG需要明确的上下文感知协议和经过校准的判断。EnterpriseRAG为测量和弥补这些缺口提供了一个可重复的基础,直接为企业级RAG系统的部署决策提供信息。我们将在发表后发布该基准和评估框架。
cs.AI / 47 / 2608.11588

CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications

CoAdapt-GUI:针对未见图形用户界面的联合工作流上下文和策略适应
Guo, Linqiang, Gu, Li, Jiang, Zihuan, Chi, Zhixiang, Reid, Siobhan, Wang, Ziqiang, Yu, Yuanhao, Liu, Wei, Wang, Yang, Tse-Hsun, Chen
Abstract
Mobile GUI agents remain brittle when deployed to applications absent from source training. We study novel-app generalization under a limited target interaction budget and without target demonstrations. We introduce CoAdapt-GUI, a test-time adaptation (TTA) framework that jointly adapts structured workflow context and policy from the agent's own target-app rollouts and rewards. The workflow context retains transferable procedures, failure modes, and verification rules while excluding app-bound source details. This separation allows reusable workflow knowledge to guide adaptation without transferring source-interface state. For policy adaptation, task-context-matched group-relative optimization updates a LoRA adapter on a frozen vision-language model. Across two unseen-app evaluations, CoAdapt-GUI reaches 45.0% on AndroidWorld-Generalization, compared with 37.5% for the reported Policy-Only TTA baseline, and raises AndroidWorld Plus performance from 38.6% to 52.9%. These results show that transfer-constrained workflow context provides substantial gains and that joint policy adaptation further improves held-out performance.
Chinese Translation
移动图形用户界面(GUI)代理在部署到缺乏源训练的应用程序时仍然表现脆弱。我们在有限的目标交互预算和没有目标演示的情况下研究新应用的泛化。我们提出了CoAdapt-GUI,一个测试时适应(TTA)框架,它联合适应代理自身在目标应用中的回滚和奖励的结构化工作流上下文和策略。工作流上下文保留可转移的程序、失败模式和验证规则,同时排除与应用绑定的源细节。这种分离允许可重用的工作流知识指导适应,而不转移源接口状态。对于策略适应,任务上下文匹配的组相对优化更新冻结的视觉-语言模型上的LoRA适配器。在两个未见应用的评估中,CoAdapt-GUI在AndroidWorld-Generalization上达到了45.0%,相比之下,报告的仅策略TTA基线为37.5%,并将AndroidWorld Plus的性能从38.6%提高到52.9%。这些结果表明,转移受限的工作流上下文提供了显著的提升,而联合策略适应进一步改善了保留性能。
cs.AI / 48 / 2608.11604

Learning from Online User Feedback for Shopping Agents

从在线用户反馈中学习购物代理
Zhang, Haobo, Mao, Kelong, Xu, Sulong, Gu, Simiu, Dou, Zhicheng
Abstract
Large language model-based shopping agents are increasingly deployed in real-world e-commerce platforms, generating massive amounts of user interaction logs that provide valuable supervision for improving these agents. However, existing approaches primarily rely on offline training signals, such as user-item interactions or synthetic preference data, while largely overlooking the rich supervision contained in users' natural conversational feedback. Moreover, the available online feedback is heterogeneous, sparse, and noisy, making it difficult to transform into reliable learning signals automatically. To address these challenges, we propose LOFA, a framework that enables shopping agents to learn directly from real online interaction logs without human annotation. LOFA combines reinforcement learning over verifiable purchase outcomes with feedback-aware on-policy distillation, which identifies users'in-dialogue directives and converts them into dense token-level supervision. These complementary objectives capture both collaborative behavioral patterns and user-specific preferences. Extensive experiments on real-world e-commerce logs demonstrate that LOFA consistently improves recommendation quality, response helpfulness, and user-satisfaction alignment over strong baselines, highlighting the effectiveness of learning shopping agents from real online user feedback.
Chinese Translation
基于大型语言模型的购物代理在现实世界的电子商务平台上越来越多地被部署,生成大量的用户交互日志,这些日志为改善这些代理提供了宝贵的监督。然而,现有的方法主要依赖于离线训练信号,如用户-商品交互或合成偏好数据,而在很大程度上忽视了用户自然对话反馈中蕴含的丰富监督。此外,现有的在线反馈是异构的、稀疏的且噪声较多,这使得将其自动转化为可靠的学习信号变得困难。为了解决这些挑战,我们提出了LOFA,一个使购物代理能够直接从真实在线交互日志中学习的框架,而无需人工标注。LOFA结合了基于可验证购买结果的强化学习与反馈感知的在线策略蒸馏,后者识别用户在对话中的指令并将其转化为密集的标记级监督。这些互补目标捕捉了协作行为模式和用户特定偏好。在真实世界电子商务日志上的广泛实验表明,LOFA在推荐质量、响应有用性和用户满意度对齐方面始终优于强基线,突显了从真实在线用户反馈中学习购物代理的有效性。
cs.AI / 49 / 2608.11605

Foresight Without Seeing: Latent Futures for World Action Models

未见之先见:世界行动模型的潜在未来
Huang, Jiakai, Wu, Zhongbo, Zhang, Zheng, Wang, Zihan, You, Shan, Huang, Tao
Abstract
World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction. Existing WAMs differ in how predictive dynamics are exposed to the action pathway. Explicit-future WAMs provide direct access to predicted scene evolution, but incur substantial inference costs from iterative video denoising. In contrast, direct-policy WAMs efficiently predict actions from the current observation but lack an explicit inference-time interface for exposing predictive dynamics to the Action DiT. To bridge this gap, we propose ForeWAM, a dynamics-conditioned direct-policy WAM that provides predictive context for action generation without decoding future videos. At its core, Future-KV performs a single Video DiT prefill over the current visual latent and stochastic future slots, and reuses the resulting layer-wise key-value states throughout action denoising. We further introduce dynamics registers supervised by a frozen latent action teacher, encouraging the implicit future states to capture interaction-induced transitions such as object motion, contact changes, and task progress. Ground-truth future observations and the teacher are used only during training; deployment requires neither and performs no future video generation. Without embodied robot data pretraining, the standard and accelerated variants of ForeWAM achieve average success rates of 96.7% and 96.9% on LIBERO, respectively. The standard variant further achieves 61.6% success on LIBERO-Plus. These results demonstrate that direct-policy WAMs can retain efficient action prediction while exposing predictive dynamics to the action pathway without explicitly generating future observations.
Chinese Translation
世界行动模型(World Action Models, WAMs)将未来的视觉预测与机器人动作生成相结合,使得政策能够模拟物理世界在交互过程中的演变。现有的WAMs在预测动态如何暴露于动作路径上存在差异。显式未来WAMs提供对预测场景演变的直接访问,但由于迭代视频去噪,导致了相当大的推理成本。相比之下,直接政策WAMs能够有效地从当前观察中预测动作,但缺乏在推理时暴露预测动态给动作生成的显式接口。为了解决这一问题,我们提出了ForeWAM,一种动态条件的直接政策WAM,能够在不解码未来视频的情况下为动作生成提供预测上下文。其核心是Future-KV,它在当前视觉潜变量和随机未来槽上执行单次视频DiT预填充,并在整个动作去噪过程中重用结果的层级关键值状态。我们进一步引入由冻结的潜在动作教师监督的动态寄存器,鼓励隐式未来状态捕捉由交互引起的转变,例如物体运动、接触变化和任务进展。真实的未来观察和教师仅在训练期间使用;部署时不需要这两者,并且不进行未来视频生成。在没有具身机器人数据预训练的情况下,ForeWAM的标准和加速变体在LIBERO上分别达到了96.7%和96.9%的平均成功率。标准变体在LIBERO-Plus上进一步达到了61.6%的成功率。这些结果表明,直接政策WAMs能够在不显式生成未来观察的情况下,保持高效的动作预测,同时将预测动态暴露于动作路径上。
cs.AI / 50 / 2608.11616

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

MBA:用于现实世界商业构思的多模态基准与智能体
Choi, Hojun, Shin, Jaeyo, Lee, Suin, Shim, Hyunjung
Abstract
Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we automatically caption images and employ GPT-4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis. Following prior work, we evaluate agents across six business-oriented criteria using MLLM-as-a-Judge. To consider settings where criteria are hidden or disclosed, we present MBA-b and MBA-k for blind and known, respectively. We train both with two novel reward objectives---creativity and feasibility---while MBA-k further optimizes the six disclosed criteria for eight in total. Both are trained via LoRA-based supervised fine-tuning followed by group relative policy optimization with these setting-specific rewards. For extensive experiments on MBA-Bench, we set up two baselines accommodating either captions only or multimodal inputs, with the latter nearing closed-source performance on several metrics. MBA-b and MBA-k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.
Chinese Translation
由大型语言模型(LLMs)驱动的智能系统为商业构思开辟了新的机遇。然而,现有的方法仍然局限于仅文本的范式,尽管现实世界的情境本质上是多模态的。因此,我们引入了MBA-Bench,这是第一个用于训练和评估商业构思智能体的多模态基准,包含来自六个领域的30K样本,每个领域都有独特的视觉线索,仅通过文本无法完全传达。具体而言,我们自动为图像生成标题,并利用GPT-4o通过检索查询生成、市场证据检索和证据增强合成,为每个商业问题生成三个参考想法。遵循先前的工作,我们使用MLLM-as-a-Judge在六个商业导向的标准上评估智能体。为了考虑标准被隐藏或公开的情境,我们分别提出了MBA-b和MBA-k用于盲评和已知评估。我们使用两种新颖的奖励目标——创造力和可行性——对两者进行训练,而MBA-k进一步优化了六个已公开标准,总共达到八个标准。两者均通过基于LoRA的监督微调进行训练,随后进行针对这些特定设置奖励的群体相对策略优化。针对MBA-Bench的广泛实验,我们设置了两个基线,分别适应仅使用标题或多模态输入,后者在多个指标上接近闭源性能。MBA-b和MBA-k分别比标题基线提高了63.9%和77.1%,比多模态基线提高了25.6%和35.8%。
cs.AI / 51 / 2608.11625

Making AI-Generated Feedback Matter: From Provision to Student Enactment

让人工智能生成的反馈变得重要:从提供到学生的实施
Alsaiari, Omar, Baghaei, Nilufar, Lodge, Jason M., Gaševi'c, Dragan, Winstone, Naomi, Khosravi, Hassan
Abstract
Feedback processes strongly influence student learning, yet their educational value depends on addressing two distinct challenges: providing high-quality, timely, and individualised feedback at scale, and supporting students to interpret, evaluate, and act on that feedback productively. Generative AI offers a credible means of addressing the provision challenge, but students' uptake of AI-generated feedback remains limited. We conducted a large-scale quasi-experimental sequential cohort study comparing three AI-mediated feedback workflows across 13,037 students and 51,296 student-authored resources. In Directed Feedback (n = 3,723), students received AI-generated feedback comments without structured support. In Self-Directed Feedback (n = 3,951), students could initiate optional AI-supported dialogue. In Enacted Feedback (n = 5,363), students were prompted to select feedback suggestions, evaluate their relevance, and engage in targeted AI-supported dialogue anchored to those selections. Enacted Feedback was associated with significantly higher uptake of AI-generated feedback, with an estimated probability of 26.2%, compared with 14.1% for Directed Feedback and 0.1% for Self-Directed Feedback. It was also associated with significantly higher self-assessment confidence and submitted-work quality than both comparison conditions. These findings suggest that the educational value of AI-generated feedback depends not only on the quality of feedback comments, but also on workflows that actively structure students' enactment of feedback literacy processes. The results have implications for the design of AI feedback systems that position learners as active participants in judgement, dialogue, and improvement rather than passive recipients of comments. Overall findings show that AI access alone is insufficient; purposeful workflow design is central to productive feedback use.
Chinese Translation
反馈过程对学生学习有着重要影响,但其教育价值取决于解决两个不同的挑战:在大规模范围内提供高质量、及时和个性化的反馈,以及支持学生有效地解读、评估和采取行动。生成性人工智能提供了一种有效应对反馈提供挑战的手段,但学生对人工智能生成反馈的接受度仍然有限。我们进行了大规模的准实验顺序队列研究,比较了三种人工智能介导的反馈工作流程,涉及13,037名学生和51,296份学生创作的资源。在定向反馈(Directed Feedback,n = 3,723)中,学生在没有结构化支持的情况下接收人工智能生成的反馈评论。在自我导向反馈(Self-Directed Feedback,n = 3,951)中,学生可以发起可选的人工智能支持对话。在实施反馈(Enacted Feedback,n = 5,363)中,学生被提示选择反馈建议,评估其相关性,并参与与这些选择相关的有针对性的人工智能支持对话。实施反馈与人工智能生成反馈的接受度显著提高相关,估计概率为26.2%,而定向反馈为14.1%,自我导向反馈为0.1%。与两个比较条件相比,实施反馈还与显著更高的自我评估信心和提交作品质量相关。这些发现表明,人工智能生成反馈的教育价值不仅取决于反馈评论的质量,还取决于积极构建学生反馈素养过程实施的工作流程。这些结果对设计人工智能反馈系统具有重要意义,强调学习者应作为判断、对话和改进的积极参与者,而非被动接受评论的对象。总体结果表明,仅有人工智能的访问是不够的;有目的的工作流程设计是有效反馈使用的核心。
cs.AI / 52 / 2608.11631

CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement

CLAIM:基于不确定性测量的大型语言模型主动澄清的开放领域引导
Yang, Kuangzhao, Zhao, Ziliang, Dou, Zhicheng
Abstract
In open-domain human-computer interaction scenarios, large language models (LLMs) frequently encounter user queries that are ambiguous or incomplete. In such cases, directly producing an answer often leads to overgeneralized, erroneous, or low-information responses. In contrast, asking clarifying questions can substantially improve interaction quality. However, existing approaches still rely heavily on manually annotated data or preference alignment to address two fundamental challenges: when clarification is necessary, and which aspect of the query should be clarified. This reliance incurs high annotation costs and limits generalization. To address these challenges, we propose CLAIM, an uncertainty-driven framework for active clarification learning in open-domain settings. CLAIM eliminates the need for explicit human preference annotations by quantifying query uncertainty through the entropy induced by answer disagreements across multiple models. This uncertainty signal is then used to construct high-quality synthetic data, enabling the training of a unified clarification decision model through a combination of supervised learning and reinforcement learning. Specifically, we propose an entropy-driven synthetic data generation pipeline that integrates entropy-based uncertainty estimation with semantic clustering and reasoning-based judgments, enabling reliable automatic annotation of clarification requirements. To train CLAIM, we formulate the clarification process as a structured decision generation problem and adopt a training paradigm that combines supervised fine-tuning (SFT) with group-relative policy optimization (GRPO). Experimental results demonstrate that CLAIM can learn stable and generalizable clarification strategies without relying on manually labeled data, offering a low-cost and robust solution for proactive understanding in real-world open-domain interactions with LLMs.
Chinese Translation
在开放领域的人机交互场景中,大型语言模型(LLMs)经常会遇到模糊或不完整的用户查询。在这种情况下,直接生成答案往往会导致过于泛化、错误或信息量低的响应。相比之下,提出澄清问题可以显著提高交互质量。然而,现有的方法仍然在很大程度上依赖于手动标注的数据或偏好对齐,以解决两个基本挑战:何时需要澄清,以及应澄清查询的哪个方面。这种依赖带来了高昂的标注成本,并限制了泛化能力。为了解决这些挑战,我们提出了CLAIM,一个基于不确定性的主动澄清学习框架,适用于开放领域设置。CLAIM通过量化查询的不确定性,消除了对明确人类偏好标注的需求,这种不确定性是通过多个模型之间答案不一致引起的熵来衡量的。然后,利用这一不确定性信号构建高质量的合成数据,从而通过监督学习和强化学习的结合训练统一的澄清决策模型。具体而言,我们提出了一种基于熵的合成数据生成管道,将基于熵的不确定性估计与语义聚类和基于推理的判断相结合,实现了澄清需求的可靠自动标注。为了训练CLAIM,我们将澄清过程表述为一个结构化决策生成问题,并采用结合监督微调(SFT)和组相对策略优化(GRPO)的训练范式。实验结果表明,CLAIM能够在不依赖手动标注数据的情况下学习稳定且可泛化的澄清策略,为在真实世界的开放领域与LLMs的主动理解提供了一种低成本且稳健的解决方案。
cs.AI / 53 / 2608.11676

XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication

XBridge:用于异构大型语言模型通信的实体基础潜在桥接
Yang, Wooseong, Huang, Wei-Chieh, Zhang, Weizhi, Wang, Yu, Yu, Philip S., Lee, Junhyun
Abstract
Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redundant reasoning patterns. Yet existing communication protocols either operate through text, discarding the sender's internal representations, or require architectural homogeneity for latent-level transfer. We identify the entity grounding problem in cross-architecture communication: cross-attention bridges that transfer continuous representations across different LLM families suffer from rare-token compression collapse, where entity identity is lost in the continuous bottleneck (bridge-only F1 ~30%). We propose XBRIDGE, a decode-free communication protocol that addresses this through two mechanisms. Lexical Anchor Mapping (LAM) maps the sender's original context tokens to the receiver's vocabulary, providing discrete entity anchors. A Latent Enrichment Bridge (LEB) lets the receiver query the sender's hidden states for contextual enrichment. The entity anchors ground the bridge's contextual signals to specific entities through the receiver's own self-attention. Across three model families (Llama, Qwen, and Mistral), seven benchmarks, and both communication directions, XBRIDGE outperforms text-based communication on all seven tasks for each model pair while achieving 11x lower latency, and in a same-architecture setting it also exceeds a KV-sharing baseline on six of seven tasks. LEB requires only 264M trainable parameters (3.8% of the receiver), is trained on a small balanced sample set, and adds negligible inference overhead.
Chinese Translation
异构多智能体大型语言模型(LLM)系统中,智能体由不同的模型家族驱动,能够通过减少冗余推理模式而超越同质配置。然而,现有的通信协议要么通过文本进行操作,丢弃发送者的内部表示,要么要求架构同质性以实现潜在级别的传输。我们识别出跨架构通信中的实体基础问题:跨注意力桥接在不同LLM家族之间传输连续表示时,容易遭遇稀有标记压缩崩溃,导致实体身份在连续瓶颈中丢失(仅桥接的F1分数约为30%)。我们提出了XBRIDGE,一种无解码的通信协议,通过两种机制解决这一问题。词汇锚定映射(Lexical Anchor Mapping, LAM)将发送者的原始上下文标记映射到接收者的词汇表中,提供离散的实体锚定。潜在丰富桥接(Latent Enrichment Bridge, LEB)允许接收者查询发送者的隐藏状态以进行上下文丰富。在接收者自己的自注意力机制下,实体锚定将桥接的上下文信号与特定实体相结合。在三个模型家族(Llama、Qwen和Mistral)、七个基准测试和两个通信方向上,XBRIDGE在每对模型的所有七个任务中均优于基于文本的通信,同时实现了11倍更低的延迟,并且在同一架构设置下,它在七个任务中的六个任务上也超越了KV共享基线。LEB仅需264M可训练参数(占接收者的3.8%),在一个小的平衡样本集上进行训练,并且增加的推理开销可以忽略不计。
cs.AI / 54 / 2608.11679

AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection

AgenticTwin:一种与数字双胞胎集成的代理型大语言模型框架用于异常检测
Hasan, Touseef, Ghanta, Mounika, Sarkar, Souvika, Guin, Ujjwal
Abstract
Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreting anomalies detected within digital twin pipelines is challenging, as the sheer complexity and volume of raw sensor data make thorough analysis difficult. Recent advances in large language models (LLMs) offer promising capabilities for reasoning and explanation, yet their integration into digital twin-driven anomaly analysis remains underexplored. In this work, we propose AgenticTwin, an agentic framework that integrates LLM-driven reasoning with a digital twin-based anomaly detection pipeline. The framework grounds LLM-generated explanations in outputs from a digital twin-driven anomaly classifier and enables human operators to ask relevant natural-language questions about the system. Beyond the framework itself, we introduce a benchmark-oriented evaluation pipeline constructed over synthetic anomalies injected into a real-world weather sensor dataset, enabling controlled generation of operator queries over anomaly events. We further evaluate the feasibility of deploying lightweight, open-source LLMs for practical cyber-physical environments. Experimental results demonstrate that structured agent collaboration and knowledge-grounded reasoning improve diagnosis quality, contextual retrieval, and mitigation quality across diverse possible anomaly scenarios.
Chinese Translation
数字双胞胎越来越多地用于监测和模拟网络物理系统的行为。即使有熟练的操作员,在数字双胞胎管道中解释检测到的异常也具有挑战性,因为原始传感器数据的复杂性和数量使得彻底分析变得困难。最近在大语言模型(LLMs)方面的进展为推理和解释提供了有希望的能力,但它们在数字双胞胎驱动的异常分析中的集成仍然未得到充分探索。在本研究中,我们提出了AgenticTwin,这是一种将LLM驱动的推理与基于数字双胞胎的异常检测管道集成的代理型框架。该框架将LLM生成的解释与来自数字双胞胎驱动的异常分类器的输出相结合,使人类操作员能够针对系统提出相关的自然语言问题。除了框架本身,我们还引入了一个基准导向的评估管道,该管道建立在注入真实天气传感器数据集的合成异常之上,能够控制生成操作员对异常事件的查询。我们进一步评估了在实际网络物理环境中部署轻量级开源LLMs的可行性。实验结果表明,结构化的代理协作和知识基础推理提高了在多种可能异常场景下的诊断质量、上下文检索和缓解质量。
cs.AI / 55 / 2608.11683

FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

FrontierFinance:衡量金融代理前沿智能的挑战性基准
Zhang, Yuhao, Koyluoglu, O. Ozan, Venkatesh, Thejas, Martinez, Richard Diehl, Bhatia, Vishank, Alidoust, Arash, Paranjape, Ashwin
Abstract
AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that current models have largely saturated, while reference-based metrics and generic LLM-as-a-judge scoring fall short on the open-ended, long-form answers that real analyst queries demand. We introduce FrontierFinance, a fully open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six crucial use cases across the full investor workflow. FrontierFinance is both broader and harder than existing public finance benchmarks. Evaluating frontier models and agent systems under a common harness restricted to publicly available data, we find that the tool harness, not the model alone, strongly shapes quality and efficiency; that Samaya's in-house system leads at 56.0%, ahead of the strongest frontier model (Claude Fable 5, 49.2%) at roughly 2.2x lower cost; and that the best open-weight model (Kimi K3, 46.4%) nearly matches the best proprietary model at 4.5x lower cost. Screening & Discovery and Sector, Industry & Macro remain the hardest use cases across all systems, where even the best systems reach only 33% and 39%. We make the dataset and grading code publicly available.
Chinese Translation
人工智能代理在专业投资研究中的应用日益增多,但目前没有任何基准能够全面捕捉投资者工作流程的复杂性。现有基准主要针对金融数据提取,这一狭窄领域当前模型已基本饱和,而基于参考的指标和通用的 LLM-as-a-judge 评分在真实分析师查询所需的开放式、长篇答案方面表现不足。我们推出了 FrontierFinance,这是一个完全开放的基准,包含 220 个专家设计的查询和 11,543 个来源归属的评分标准,涵盖了整个投资者工作流程中的六个关键使用案例。FrontierFinance 比现有的公共金融基准更广泛且更具挑战性。在对前沿模型和代理系统进行评估时,我们发现工具的使用方式,而非单一模型,显著影响质量和效率;Samaya 的内部系统以 56.0% 的表现领先,远超最强的前沿模型(Claude Fable 5,49.2%),且成本大约低 2.2 倍;而最佳的开放权重模型(Kimi K3,46.4%)几乎与最佳专有模型相当,但成本低 4.5 倍。筛选与发现以及行业、部门与宏观经济仍然是所有系统中最具挑战性的使用案例,即使是最佳系统在这两个领域的表现也仅达到 33% 和 39%。我们将数据集和评分代码公开发布。
cs.AI / 56 / 2608.11692

HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting

HUGIN:增强自主物流分拣的视觉-语言规划
Sun, Xikai, Zhou, Cangtian, Liu, Kebin, Ma, Ke, Wang, Xu, Chen, Zaishu, Wang, Haotian, Liu, Li, Liu, Yunhao
Abstract
Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which requires joint planning over spatially disjoint camera views. We formulate this setting as Joint Multi-Scene Understanding (JMSU). With open-world visual understanding and task-planning capabilities, vision-language models (VLMs) are promising candidates for JMSU. However, directly applying existing VLMs to JMSU is non-trivial due to scarce cross-scene supervision and attention dispersion caused by long visual context in JMSU. To address these challenges, we propose HUGIN, a training framework with two complementary components. Endogenous Data Augmentation recombines verified atomic facts under operating constraints, while Global Context Ranking aligns the instruction representation more strongly with the complete visual context than with a partial visual context. To support ongoing research, we construct a high-quality industrial sorting dataset and benchmark named SortingBench from four layouts of autonomous logistics sorting systems. Across five open VLMs, HUGIN consistently outperforms matched baselines; for example, the accuracy on SortingBench of Qwen3-VL-8B increases from 63.6% to 78.8%. Additional experiments verify the effectiveness of each component and JMSU's spillover benefits in embodied tasks. Deployment tests involving more than 15,000 packages support the practical viability of VLM-based planning for autonomous logistics sorting.
Chinese Translation
自主物流分拣系统(ALSS)是具身人工智能的重要工业应用,要求在空间上不相交的摄像头视角之间进行联合规划。我们将这一设置表述为联合多场景理解(Joint Multi-Scene Understanding, JMSU)。凭借开放世界的视觉理解和任务规划能力,视觉-语言模型(Vision-Language Models, VLMs)是JMSU的有希望的候选者。然而,由于跨场景监督稀缺以及JMSU中长视觉上下文导致的注意力分散,直接将现有的VLM应用于JMSU并非易事。为了解决这些挑战,我们提出了HUGIN,一个具有两个互补组件的训练框架。内生数据增强(Endogenous Data Augmentation)在操作约束下重新组合经过验证的原子事实,而全局上下文排序(Global Context Ranking)使指令表示与完整视觉上下文的对齐程度明显高于与部分视觉上下文的对齐程度。为了支持持续的研究,我们构建了一个高质量的工业分拣数据集和基准,命名为SortingBench,基于四种自主物流分拣系统的布局。在五个开放的VLM上,HUGIN始终优于匹配的基线;例如,Qwen3-VL-8B在SortingBench上的准确率从63.6%提高到78.8%。额外实验验证了每个组件的有效性以及JMSU在具身任务中的溢出效益。涉及超过15,000个包裹的部署测试支持了基于VLM的自主物流分拣规划的实际可行性。
cs.AI / 57 / 2608.11705

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning

提高大型语言模型的客观性:通过特征不变安全调优稳定大型语言模型的安全行为
Cao, Lang
Abstract
Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of the user request: they should refuse unsafe requests and comply with safe ones. However, we show that the same request can elicit substantially different safety decisions under different traits assigned in the system prompt, a failure mode we call trait-induced safety variation. To measure this failure, we introduce refusal-based metrics: Trait-Induced Deviation measures dataset-level deviation from the no-trait baseline, while Trait-Induced Flip Rate measures whether the same request receives different safety decisions across traits. We then provide a representation-level analysis of the mechanism behind trait-induced safety shifts and find that traits perturb the model's safety representations within a low-dimensional subspace. To achieve trait-invariant safety, where safety behavior remains stable across traits, we introduce Trait-Invariant Safety Tuning (TIST), a simple yet effective self-distillation framework that aligns an LLM's trait-conditioned behavior with its no-trait behavior. Guided by our analysis, we further propose Trait-Subspace Neutralization (TraSN), an instantiation of TIST, which enforces invariance only within the identified trait subspace. Experiments show that TraSN improves trait-invariant safety and strengthens harmful-request safety while preserving general capability. Our results highlight traits as an important factor in LLM safety and robust model behavior.
Chinese Translation
对齐的大型语言模型(LLMs)预计会根据用户请求的内容表现出安全行为:它们应拒绝不安全的请求并遵从安全的请求。然而,我们发现相同的请求在系统提示中赋予不同特征时可能会引发显著不同的安全决策,这种失败模式我们称之为特征诱导的安全变异。为了衡量这种失败,我们引入了基于拒绝的指标:特征诱导偏差(Trait-Induced Deviation)衡量数据集级别相对于无特征基线的偏差,而特征诱导翻转率(Trait-Induced Flip Rate)则衡量相同请求在不同特征下是否会得到不同的安全决策。接着,我们对特征诱导安全变化背后的机制进行了表示层面的分析,发现特征在低维子空间内扰动模型的安全表示。为了实现特征不变安全,即安全行为在不同特征下保持稳定,我们引入了特征不变安全调优(Trait-Invariant Safety Tuning, TIST),这是一种简单而有效的自蒸馏框架,旨在将LLM的特征条件行为与其无特征行为对齐。在我们的分析指导下,我们进一步提出了特征子空间中和(Trait-Subspace Neutralization, TraSN),这是TIST的一种实例,强制在识别的特征子空间内保持不变性。实验表明,TraSN提高了特征不变安全性,并增强了对有害请求的安全性,同时保持了模型的整体能力。我们的结果强调了特征作为LLM安全性和稳健模型行为的重要因素。
cs.AI / 58 / 2608.11724

Proportional Analogies on Probability Distributions via Bayesian Updating

通过贝叶斯更新的概率分布比例类比
Murena, Pierre-Alexandre
Abstract
Analogies are quaternary relations of the form "A is to B as C is to D". Among the various formalizations of analogical reasoning, proportional analogies provide an important axiomatic framework by characterizing valid analogies through a set of postulates. While proportional analogies have been extensively studied over Boolean, symbolic, and real-valued domains, their extension to probability distributions remains largely unexplored. In this paper, we introduce a notion of proportional analogy for probability distributions based on Bayesian updating. Our approach builds upon the idea that two distributions are related whenever one can be transformed into the other through Bayesian updating induced by a suitable set of observations. We investigate this framework for several standard members of the exponential family and discuss how it naturally extends to arbitrary probability distributions through Gaussian mixture approximations.
Chinese Translation
类比是形式为“A与B的关系如同C与D的关系”的四元关系。在各种类比推理的形式化中,比例类比通过一组公理特征化有效类比,提供了一个重要的公理框架。尽管比例类比在布尔、符号和实值领域得到了广泛研究,但其在概率分布上的扩展仍然未被充分探索。本文基于贝叶斯更新引入了概率分布的比例类比概念。我们的方法建立在这样的思想之上:当一个分布可以通过适当观察集引发的贝叶斯更新转化为另一个分布时,这两个分布是相关的。我们研究了这一框架在指数族的几个标准成员中的应用,并讨论了它如何通过高斯混合近似自然扩展到任意概率分布。
cs.AI / 59 / 2608.11727

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

Harness-IF:评估编码智能体在指令表面上的指令遵循能力
Huang, Zining, Que, Haoran, Zeng, Hong, Zhang, Ge, Wang, Zuo, Chen, Jin, Wang, Haodong, Hou, Zhongfei, Pu, Changxin, Yan, Shen, Huang, Wenhao
Abstract
When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell the difference: they concentrate rules in the user turn, while coding-agent benchmarks emphasize final task success. We introduce Harness-IF, which scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads. To separate compliance from coincidence we introduce Against-Prior Accuracy (AP-Acc), which scores only rules labeled as opposing unprompted defaults, observed by re-running tasks with the rule withheld across nine probe builds and curated otherwise. Across 12 frontier models, accuracy spans 72.1-85.9% and AP-Acc 66.1-78.6%; every model is worse on against-prior rules, by 3.6 to 7.4 points (mean 5.81), and the direction survives a common-support analysis with item-clustered intervals. Aggregate scores therefore overstate compliance by a model-specific margin: prior control leaves the top build unchanged and exchanges three adjacent rank pairs. A counterbalanced conflict pilot on nine separate builds adds a second result: pooled precedence does not follow prompt depth, with system prompts, project files, and user instructions ahead of tool and skill descriptions.
Chinese Translation
当一个编码智能体遵循一条规则时,它可能本来就会这样做。现有的指令遵循基准无法区分这一点:它们将规则集中在用户的发言中,而编码智能体的基准则强调最终任务的成功。我们提出了Harness-IF,它通过执行证据逐一评分操作规则:从一个包含642条规则的库中提取的60个现实的多轮编码项目,其中256条规则获得了判定,放置在已部署智能体读取的五个可配置表面上。为了将合规性与巧合分开,我们引入了Against-Prior Accuracy (AP-Acc),该指标仅对标记为反对未提示默认值的规则进行评分,这些规则通过在九个探测构建中重新运行任务并在其他方面进行策划时被观察到。在12个前沿模型中,准确率范围为72.1%-85.9%,AP-Acc为66.1%-78.6%;每个模型在反对先前规则的表现上都较差,下降幅度为3.6到7.4个百分点(平均5.81),这一趋势在使用项目聚类区间的共同支持分析中得以保留。因此,聚合评分在模型特定的范围内夸大了合规性:先前的控制使得最佳构建保持不变,并交换了三个相邻的排名对。对九个独立构建的平衡冲突试点增加了第二个结果:汇总优先级并不遵循提示深度,系统提示、项目文件和用户指令优先于工具和技能描述。
cs.AI / 60 / 2608.11768

HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Systems via Hyperbolic Geometry

HyperANFIS:通过双曲几何增强自适应神经模糊系统中的规则表示和可解释性
Pei, Haoran, Su, Zhao, Lin, Zetao, Li, Haoran, Shen, Jun, Zhu, Qi, Guo, Lan, Zhou, Qingguo, Yong, Binbin
Abstract
The adaptive neuro-fuzzy inference system (ANFIS) is an interpretable reasoning framework capable of generating explicit IF-THEN fuzzy rules, making it suitable for tasks requiring transparent reasoning. However, existing ANFIS models generally construct rule antecedents and perform inference in Euclidean space, limiting their representational capacity and predictive performance. To address this issue, we propose Hyperbolic ANFIS (HyperANFIS), a hyperbolic extension of ANFIS. HyperANFIS preserves the fuzzy semantics and core architecture of conventional ANFIS while performing rule-prototype learning, rule activation, and consequent aggregation in hyperbolic space. It also retains the ability to generate interpretable IF-THEN rules. By exploiting the representational properties of hyperbolic geometry, HyperANFIS strengthens the fuzzy inference process, thereby improving predictive accuracy, inter-rule collaboration, and the credibility of its interpretable rules. Experimental results show that HyperANFIS consistently outperforms the standard ANFIS baseline and various ANFIS variants across all datasets, while also generating higher-quality fuzzy rules.
Chinese Translation
自适应神经模糊推理系统(ANFIS)是一种可解释的推理框架,能够生成明确的IF-THEN模糊规则,适用于需要透明推理的任务。然而,现有的ANFIS模型通常在欧几里得空间中构建规则前提并进行推理,这限制了它们的表示能力和预测性能。为了解决这一问题,我们提出了双曲ANFIS(HyperANFIS),这是ANFIS的一个双曲扩展。HyperANFIS保留了传统ANFIS的模糊语义和核心架构,同时在双曲空间中进行规则原型学习、规则激活和结果聚合。它还保留了生成可解释的IF-THEN规则的能力。通过利用双曲几何的表示特性,HyperANFIS增强了模糊推理过程,从而提高了预测准确性、规则间协作以及其可解释规则的可信度。实验结果表明,HyperANFIS在所有数据集上始终优于标准ANFIS基线和各种ANFIS变体,同时生成更高质量的模糊规则。
cs.AI / 61 / 2608.11775

The Sleeping Agent: What Gist-Based Context Compression Loses and Why

沉睡代理:基于要旨的上下文压缩所失去的内容及其原因
Kyrkewood, Nicholas E.
Abstract
Gist-based context compression---summarising older conversation history into compact representations---is a common approach in long-horizon language model agents, yet its effect on different types of memory retrieval is poorly understood. We use Salience-Weighted Consolidation (SWC), a biologically-inspired compression framework motivated by sleep-based memory consolidation, as a diagnostic probe to study when gist compression helps and when it hurts. SWC scores conversation history by salience, partitions it into priority tiers, and applies structured gist abstraction to mid-priority content. Evaluating four conditions on all ten LoCoMo conversations---1,935 matched text-only questions in total, 1,501 used in the primary aggregate after excluding Category 5 (adversarial) questions---at temperature 0, we find a consistent task-type interaction: gist compression substantially outperforms truncation on multi-hop reasoning and single-hop factual questions, but temporal questions remain substantially harder under compression, with compressed conditions scoring well below the full-context reference on the conversations where both are evaluated. We trace this failure to a specific mechanism: the gist abstraction prompt preserves relational and event structure while discarding dates and times. A preservation analysis across all ten conversations confirms the mechanism: an approximately 20-fold increase in temporal expression preservation (3.05% to 62.39%) with a one-sentence prompt modification, while named entity and event preservation rates barely change (x1.02 and x1.11), demonstrating that the fix is a precision instrument. The prompt modification recovers +0.314 [0.254, 0.375] judge accuracy on category-2 (temporal) questions in the matched set. Code and results: https://github.com/kyrkewood/sleeping-agent.
Chinese Translation
基于要旨的上下文压缩——将较早的对话历史总结为紧凑的表示——是长时间跨度语言模型代理中的一种常见方法,但其对不同类型记忆检索的影响尚不清楚。我们使用生物启发的压缩框架——基于睡眠的记忆巩固(Salience-Weighted Consolidation, SWC)作为诊断工具,研究何时要旨压缩有助于记忆检索,何时又会造成损害。SWC根据显著性对对话历史进行评分,将其划分为优先级层次,并对中等优先级内容应用结构化的要旨抽象。在对所有十个LoCoMo对话进行四种条件的评估中——总计1,935个匹配的文本问题,排除类别5(对抗性)问题后使用1,501个问题——在温度为0的情况下,我们发现了一种一致的任务类型交互:在多跳推理和单跳事实问题上,要旨压缩显著优于截断,但在压缩情况下,时间性问题的难度仍然显著增加,压缩条件的得分远低于完整上下文参考。在对比评估的对话中,我们将这一失败追溯到一个特定机制:要旨抽象提示保留了关系和事件结构,同时丢弃了日期和时间。对所有十个对话的保留分析确认了这一机制:通过对提示进行一次句子修改,时间表达的保留率约增加20倍(从3.05%提升至62.39%),而命名实体和事件的保留率几乎没有变化(x1.02和x1.11),这表明该修正是一个精确的工具。提示修改在匹配集的类别2(时间性)问题上恢复了+0.314 [0.254, 0.375]的判断准确率。代码和结果见:https://github.com/kyrkewood/sleeping-agent。
cs.AI / 62 / 2608.11888

Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

代理技能可能是有害的:关于技能引发的 LLM 代理失败的实证研究
Dong, Gen, Gao, Yanjie, Li, Liqun, Xu, Tianyin, Hua, Yu, Yang, Fan
Abstract
Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation. Prior work reported mixed results of agent skills: some skills improve task success rates, while others have no effect, increase token use and execution time, and even reduce success rates. This paper presents a comprehensive analysis of skill-induced agent failures by attributing task failures and cost regressions to specific loaded skills. We introduce a differential analysis framework that attributes a failure or regression to a skill by comparing a target skill-guided run against a no-skill or semantically matched skill reference run that solves the same task, or solves it more cheaply. We instantiate this framework on SkillsBench and SWE-Skills-Bench, yielding 307 skill-induced failures, including 125 functional failures and 182 efficiency regressions. We also build SkillTriage, a taxonomy-guided attribution tool that normalizes paired cases, extracts differential evidence, and produces triage reports. Our major findings include: (1) Skill induced functional failures are rarely caused by obviously irrelevant skills; instead, seemingly relevant skills often make the agent incorrectly implement or omit task-required implementation elements. (2) Skill-induced efficiency regressions are not explained by prompt length alone. (3) The largest sources within Excessive Procedure are excessive verification and heavy implementation pipelines, contributing 67 and 30 cases, respectively. This shows that skills often turn validation checklists and construction recipes into mandatory work. Based on our findings, we propose research topics and tooling improvements for safer and more cost-aware skill reuse.
Chinese Translation
代理技能是扩展 LLM 代理以实现可重用指导的实际机制。技能可以影响代理的任务执行,包括规划、工具使用、问题解决和验证。先前的研究报告了代理技能的混合结果:一些技能提高了任务成功率,而另一些则没有效果,增加了令牌使用和执行时间,甚至降低了成功率。本文通过将任务失败和成本回归归因于特定加载技能,提供了对技能引发的代理失败的全面分析。我们引入了一种差异分析框架,通过将目标技能指导的运行与无技能或语义匹配的技能参考运行进行比较,从而将失败或回归归因于某一技能,这两者解决了相同的任务或以更低的成本解决了该任务。我们在 SkillsBench 和 SWE-Skills-Bench 上实例化了该框架,得到了 307 个技能引发的失败,包括 125 个功能性失败和 182 个效率回归。我们还构建了 SkillTriage,一个基于分类法的归因工具,规范化配对案例,提取差异证据,并生成分类报告。我们的主要发现包括:(1)技能引发的功能性失败很少是由明显无关的技能造成的;相反,似乎相关的技能往往使代理错误地实施或遗漏任务所需的实施元素。(2)技能引发的效率回归不能仅通过提示长度来解释。(3)在过度程序中,最大的来源是过度验证和繁重的实施流程,分别贡献了 67 和 30 个案例。这表明,技能往往将验证清单和构建配方转变为强制性工作。基于我们的发现,我们提出了更安全和更具成本意识的技能重用的研究主题和工具改进。
cs.AI / 63 / 2608.11905

Policy-as-logic for robust reasoning over rules

基于逻辑的政策用于规则的稳健推理
Nair, Rahul, Lipka, Bastian, Daly, Elizabeth
Abstract
In many practical applications of generative AI systems, from tax rules to airline baggage allowance, responses to natural language queries must respect written policies or rules. We present a hybrid symbolic approach that expresses policies in formal logic and at inference time exploits the representation power of language models for fact extraction to ground predicates, and an answer set solver for reasoning such that responses are interpretable, auditable, and as we show, accurate and robust under input perturbations. Specifically, we show this separation of extraction and reasoning steps outperforms policy-as-prompt and policy-as-code methods in most cases with ~10x reduction in token usage. The results point to the value of structured reasoning and symbolic solvers in conjunction with generative models to make robust decisions involving objective criteria.
Chinese Translation
在许多生成性人工智能系统的实际应用中,从税收规则到航空公司行李限额,对自然语言查询的响应必须遵循书面政策或规则。我们提出了一种混合符号方法,该方法将政策表达为形式逻辑,并在推理时利用语言模型的表示能力进行事实提取,以支持谓词的基础,并使用答案集求解器进行推理,从而使响应可解释、可审计,并且如我们所示,在输入扰动下准确且稳健。具体而言,我们展示了提取和推理步骤的分离在大多数情况下优于基于提示的政策和基于代码的政策方法,令牌使用量减少约10倍。结果表明,结构化推理和符号求解器与生成模型结合在一起,对于涉及客观标准的稳健决策具有重要价值。
cs.AI / 64 / 2608.11941

OEIS Open: How many conjectures can language models turn into theorems?

OEIS Open:语言模型能将多少猜想转化为定理?
Adamczewski, Tom
Abstract
We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Whereas these conjectures had previously been attempted only with a bespoke agent, our open-source evaluation code runs any generic language model (LM) against them, and is secure against LM cheating attempts. We find that LMs equipped with a minimal set of tools resolve 147 of these conjectures with a budget of \$50 per attempt, scoring 30% on OEIS Open. OEIS Open Lite is a random subset of 100 conjectures for cheaper evaluation. When evaluated with a budget of \$200 per attempt, the best current LM scores 44% on OEIS Open Lite. Giving LMs access to the mathematics literature via 476,000 papers from arXiv did not increase performance on OEIS Open Lite, and nor did using more sophisticated agent loops. The conjectures covered in this work are of uncertain mathematical significance, and most have likely received little previous attention. Nevertheless, our results show that LMs can resolve open research conjectures autonomously and at modest cost.
Chinese Translation
我们构建了OEIS Open,这是一个基于492个来自OEIS的开放数学猜想的基准,由Tsoukalas等人使用Lean进行了形式化。尽管这些猜想之前仅通过定制代理进行尝试,但我们的开源评估代码能够对其进行评估,并且能够防止语言模型(LM)的作弊尝试。我们发现,配备最小工具集的LM在每次尝试预算为50美元的情况下解决了147个猜想,在OEIS Open上得分为30%。OEIS Open Lite是一个随机选择的100个猜想的子集,便于进行更便宜的评估。当每次尝试的预算为200美元时,目前最佳的LM在OEIS Open Lite上的得分为44%。通过476,000篇来自arXiv的论文让LM访问数学文献,并未提高其在OEIS Open Lite上的表现,使用更复杂的代理循环也没有带来改善。本研究所涵盖的猜想在数学上具有不确定的重要性,且大多数可能之前未受到太多关注。尽管如此,我们的结果表明,LM能够自主且以适度的成本解决开放研究猜想。
cs.AI / 65 / 2608.11949

ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models

ExRole:从团队轨迹到多智能体语言模型中的可执行角色
Liu, Zhou, Han, Chaoyang, Pan, Zewei, Su, Zeli, Zhang, Wentao
Abstract
Roles provide an interpretable interface for organizing language-model agents, yet most multi-agent systems treat them as hand-written prompt labels disconnected from learned behavior and parameter updates. We argue that a useful role should instead be an executable control variable: it should summarize behavior predictive of future utility, guide subsequent interaction, and identify the trainable capacity responsible for that behavior. We introduce ExRole, a trajectory-to-role framework that learns future-aware role prototypes from prefix-local team traces, resolves them into readable instructions and token-aligned role markers, and optionally routes shared LoRA rank slots with turn-aligned credit. Across MuSiQue and 2WikiMultiHopQA, ExRole improves over single-agent search by 15.0/14.4 and 13.5/16.1 EM/F1 points, respectively. Against the strongest non-ExRole controls, the corresponding gains remain 11.5/11.6 and 7.7/9.7 points. Across both benchmarks, the controlled results consistently favor trajectory-induced role conditioning over role-free, manual, random, and shuffled alternatives. Role-Agent-Turn interventions further show that the induced roles capture transferable behavioral specialization beyond fixed agent identities or turn positions.
Chinese Translation
角色为组织语言模型代理提供了可解释的接口,但大多数多智能体系统将其视为与学习行为和参数更新无关的手写提示标签。我们认为,一个有用的角色应该是一个可执行的控制变量:它应该总结预测未来效用的行为,指导后续互动,并识别负责该行为的可训练能力。我们提出了ExRole,一个从轨迹到角色的框架,它从前缀局部团队轨迹中学习未来感知的角色原型,将其解析为可读的指令和与标记对齐的角色标记,并可选择性地路由共享的LoRA(低秩适应)排名槽与回合对齐的信用。在MuSiQue和2WikiMultiHopQA数据集上,ExRole在单智能体搜索中分别提高了15.0/14.4和13.5/16.1的EM/F1分数。与最强的非ExRole控制相比,相应的增益仍然保持在11.5/11.6和7.7/9.7分。在这两个基准测试中,受控结果始终优于无角色、手动、随机和洗牌替代方案的轨迹诱导角色条件。角色-代理-回合干预进一步表明,诱导的角色捕捉了超越固定代理身份或回合位置的可转移行为专业化。
cs.AI / 66 / 2608.11977

Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection

重试、切换还是放弃?通过受控错误注入学习策略感知的工具使用政策
Chen, Chaoran, Nguyen, Vy, Zhang, Ziji, Gullapalli, Abhinav, Wang, Ziyi, Lu, Yuxuan, Wang, Dakuo, Huang, Jing, Yu, Zhou, Lai, Jin
Abstract
Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed reliably, yet deployed tools can fail transiently, persistently, or silently. Robust recovery therefore requires more than repeated retries: an agent may need to retry the same path, switch to an alternative, or recognize that no viable path remains. We present BENCH2ROBUST, a framework that converts failure-free tool-use benchmarks into controlled stochastic environments with scenario-controlled solvability, where episodes explicitly require retrying, switching, or stopping after available paths are exhausted. We use BENCH2ROBUST to study two complementary interventions: structured runtime recovery context through Bayesian Tool Memory (BTM), and curriculum-controlled reinforcement learning. Across 7 models from 4 families and two multi-turn benchmark families, tool failures produce a near-universal robustness gap. On held-out Retail tasks, BTM improves robustness by up to 16.8 percentage points without retraining, while RL learns complementary recovery behavior that remains beneficial without inference-time BTM. Combining the two reaches 40.8-45.5% under injection while preserving failure-free performance. These results suggest that robust tool use benefits from combining environment-specific recovery knowledge with learned recovery behavior.
Chinese Translation
使用工具的LLM代理通常在工具调用可靠成功的环境中进行训练和评估,然而,部署的工具可能会暂时性、持续性或无声地失败。因此,稳健的恢复不仅仅需要重复重试:代理可能需要重试相同的路径、切换到替代路径,或识别出没有可行路径可用。我们提出了BENCH2ROBUST,一个将无失败工具使用基准转换为具有场景控制可解性的受控随机环境的框架,其中情节明确要求在可用路径耗尽后进行重试、切换或停止。我们使用BENCH2ROBUST研究两种互补干预:通过贝叶斯工具记忆(Bayesian Tool Memory, BTM)提供结构化的运行时恢复上下文,以及课程控制的强化学习。在来自4个家族的7个模型和两个多轮基准家族中,工具失败产生了几乎普遍的稳健性差距。在保留的零售任务中,BTM在不重新训练的情况下将稳健性提高了最多16.8个百分点,而强化学习学习的互补恢复行为在没有推理时BTM的情况下仍然是有益的。将两者结合在注入情况下达到了40.8-45.5%的稳健性,同时保持无失败性能。这些结果表明,稳健的工具使用受益于将特定环境的恢复知识与学习到的恢复行为相结合。
cs.AI / 67 / 2608.11994

Claim-Level Reliability Assessment for Efficient Test-Time Reasoning

基于声明级别的可靠性评估以实现高效的测试时间推理
Xu, Sen, Wang, Wei, Liu, Shixi, Min, Jixin, Dai, Yingwei, Yin, Zhibin, Chen, Yirong, Zhang, Junlin
Abstract
We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reliability Assessment (CLR), a training-free framework that reallocates test-time compute from additional solution sampling to targeted verification. Since whole-trace evaluation often obscures decisive errors due to signal dilution from routine tokens, CLR condenses each reasoning trace into a compact set of decision-critical claims, thereby isolating its logical anchors. Furthermore, recognizing the inherent difficulty of generating entirely correct solutions under fixed model capabilities, CLR shifts the focus to semantic falsification. This approach exploits a fundamental asymmetry between solution construction and claim refutation. Constructing a valid solution requires a flawless reasoning path, whereas refuting an incorrect claim requires identifying only a single decisive flaw. This targeted search for negative evidence systematically compresses the survival space of high-confidence incorrect traces, effectively suppressing erroneous consensus via nonlinear reliability scoring. Across four LLMs and four reasoning benchmarks under matched budgets, CLR generally improves upon pass@1 and self-consistency. On GPT-OSS-20B/CMIMC25, for instance, CLR exceeds pass@1 by 27.15 percentage-points and raises self-consistency accuracy from 77.50\% to 82.19\% with 37.0\% fewer tokens.
Chinese Translation
我们提出了声明级别的伪造作为测试时间扩展的原则,并通过声明级别可靠性评估(Claim-Level Reliability Assessment, CLR)将其具体化,这是一种无训练的框架,旨在将测试时间的计算从额外的解决方案采样重新分配到针对性的验证上。由于整体追踪评估常常因常规标记导致信号稀释而掩盖决定性错误,CLR将每个推理追踪浓缩为一组紧凑的决策关键声明,从而隔离其逻辑锚点。此外,考虑到在固定模型能力下生成完全正确解决方案的固有困难,CLR将重点转向语义伪造。这种方法利用了解决方案构建与声明反驳之间的基本不对称性。构建有效解决方案需要无瑕疵的推理路径,而反驳错误声明只需识别一个决定性缺陷。这种针对负证据的搜索系统地压缩了高置信度错误追踪的生存空间,通过非线性可靠性评分有效抑制错误共识。在四个大型语言模型(LLMs)和四个推理基准下,CLR在匹配预算的情况下通常提高了通过率(pass@1)和自一致性。例如,在GPT-OSS-20B/CMIMC25上,CLR的通过率超过了27.15个百分点,并将自一致性准确率从77.50%提高到82.19%,同时减少了37.0%的标记数量。
cs.AI / 68 / 2608.12002

CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations

CTBench:评估人工智能代理在现实电信网络操作中的故障排除能力
Yan, Xingyu, Dai, Tingting, De Domenico, Antonio, Sana, Mohamed, Piovesan, Nicola, Li, Changchang, Liu, Bowen, Jiang, Kun, Zhang, Mengjie, Shan, Dingcheng, Pang, Jing-Cheng, Wu, Chenwei, Wu, Sijie, Chao, Lianying, Cai, Haoran, Ye, Jiantao, Li, Xubin, Lucas, Simon Mark, Chen, Xin
Abstract
Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimize configurations to enhance services, and reduce operational costs while acting under strict constraints. However, existing evaluations fail to accurately model real network characteristics or assess agents under partially observable telecom environments with diverse vendors, devices, protocols, and interfaces. In this paper, we introduce CTBench, a public benchmark for assessing whether an agent behaves like a competent telecom troubleshooting engineer. CTBench focuses on root cause analysis and path restoration. Each task is constructed by experts and annotated with rich task metadata, including golden evidence steps. CTBench uses expert-grounded metrics that evaluate both final answers and the diagnostic evidence. Experiments with representative harness-model combinations show that state-of-the-art agents perform very well at identifying endpoints in path-restoration tasks but, more generally, underperform in root cause analysis. In particular, agents struggle with interface state, link-layer, service-management, and other operational faults. Most importantly, even when agents produce plausible or correct final answers, they often fail to provide the evidence-grounded diagnoses required in operational practice. Our results further show that path restoration is generally more resource expensive, yet larger resource usage does not necessarily translate into better diagnosis.
Chinese Translation
代理越来越被认为可以用于自动化网络操作和维护,工程师必须在严格的约束下诊断网络故障,优化配置以提升服务,并降低运营成本。然而,现有的评估未能准确模拟真实网络特征或在具有多样化供应商、设备、协议和接口的部分可观察电信环境中评估代理。在本文中,我们介绍了CTBench,这是一个公共基准,用于评估代理是否表现得像一名合格的电信故障排除工程师。CTBench专注于根本原因分析和路径恢复。每个任务由专家构建,并附有丰富的任务元数据,包括黄金证据步骤。CTBench使用基于专家的指标,评估最终答案和诊断证据。与代表性的测试模型组合的实验表明,最先进的代理在路径恢复任务中识别端点的表现非常好,但在根本原因分析中表现普遍较差。特别是,代理在接口状态、链路层、服务管理和其他操作故障方面存在困难。最重要的是,即使代理生成了合理或正确的最终答案,它们通常未能提供在操作实践中所需的基于证据的诊断。我们的结果进一步表明,路径恢复通常更耗资源,但更大的资源使用并不一定转化为更好的诊断。
cs.AI / 69 / 2608.12036

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

Mechanist:将人工智能作为发现智能机制的科学工具
Wang, Mengru, Fang, Junfeng, Qiao, Shuofei, Xu, Zhenqian, Xu, Haoming, Wang, Haoxiong, Deng, Shumin, Yang, Linyi, Cui, Zhixiang, Xu, Xin, Yao, Yunzhi, Xu, Buqiang, Shen, Fei, Luo, Haozhe, Wei, Yunxiang, Zhang, Ningyu, McAuley, Julian, Chua, Tat Seng, Chen, Huajun
Abstract
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. To support autonomous mechanistic discovery, we construct an interpretability-focused knowledge graph of approximately 13,000 papers and integrate it with a multidisciplinary database of 43 million papers spanning 26 fields. We further curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates more valuable mechanism hypotheses and executes experiments more reliably. Mechanist also demonstrates a progression from discovering model behaviors to explaining and controlling AI models. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer across modalities through apparently safe training data. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Finally, Mechanist translates these mechanistic insights into practical interventions that improve model performance across diverse scenarios and steer scientific foundation models toward generating DNA sequences with specified properties.
Chinese Translation
人工智能模型在多个领域取得了显著成功,但其能力背后的机制以及可能带来的风险仍然不甚了解。随着人工智能的发展速度加快并日益自动化,机制探索仍然主要依赖人工,这加大了模型所能实现的功能与我们理解和控制它们的能力之间的差距。为了解决这一问题,我们提出了Mechanist,一个利用人工智能作为科学工具的自主系统,用于自动发现人工智能智能背后的机制。为了支持自主的机制发现,我们构建了一个以可解释性为重点的知识图谱,涵盖约13,000篇论文,并将其与一个跨越26个领域的4300万篇论文的多学科数据库整合在一起。我们还整理了32种基础方法的库,用于机制分析、因果干预和验证。与Claude Code和现有的人工智能科学家系统相比,Mechanist生成了更有价值的机制假设,并更可靠地执行实验。Mechanist还展示了从发现模型行为到解释和控制人工智能模型的进展。具体而言,Mechanist首先揭示了科学实验室中的一个反直觉安全风险,表明不安全特征可以通过看似安全的训练数据在不同模态之间转移。随后,Mechanist发展了一个信念机制理论,揭示了模型如何表示世界知识、形成信念、推断他人的信念,以及这些机制在预训练过程中如何出现。最后,Mechanist将这些机制洞察转化为实际干预措施,改善模型在多种场景下的表现,并引导科学基础模型生成具有特定属性的DNA序列。
cs.AI / 70 / 2608.12097

Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges

图结构评分标准:将评分标准编译为类型化评估图以供大型语言模型评判
Chen, Xi, Mu, Jie, Xuan, Mo, Shao, Qun
Abstract
Rubric-based evaluators commonly treat rubrics as prompt context or flat criteria: they specify what to judge but leave criterion composition implicit, even when natural-language rules state it. We introduce Graph-Structured Rubrics (GSR), which compiles a rubric into a response-independent typed evaluation graph before observing responses. Criterion nodes elicit judgments; transformation, reduction, and gating operators compose them through named ports; and a task-specific output mapping, termed Readout, converts the unique sink into a score or preference. Compilation rejects malformed or type-incompatible graphs. Pointwise evaluation judges rubric dimensions separately before graph aggregation; pairwise evaluation reuses the graph with one judgment for each candidate under every criterion. Under GPT-OSS-120B, GSR improves exact score agreement by 0.62--6.75 percentage points over Prometheus-style scoring on four pointwise datasets and achieves the numerically highest end-to-end pairwise accuracy on two preference benchmarks under native tie and abstention policies.
Chinese Translation
基于评分标准的评估者通常将评分标准视为提示上下文或平面标准:它们指定了评判内容,但在标准组合上保持隐含,即使自然语言规则已明确说明。我们引入了图结构评分标准(Graph-Structured Rubrics, GSR),在观察响应之前将评分标准编译为一个与响应无关的类型化评估图。标准节点引发判断;变换、归约和门控运算符通过命名端口组合这些节点;一个特定任务的输出映射,称为读取(Readout),将唯一的汇点转换为分数或偏好。编译过程会拒绝格式不正确或类型不兼容的图。逐点评估分别判断评分标准的维度,然后进行图聚合;成对评估在每个标准下对每个候选者重用图,并进行一次判断。在GPT-OSS-120B模型下,GSR在四个逐点数据集上相比于Prometheus风格的评分提高了0.62至6.75个百分点的准确分数一致性,并在两个偏好基准测试中在原生平局和弃权政策下实现了数值最高的端到端成对准确性。
cs.AI / 71 / 2608.12133

GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings

GUIDE:企业环境中文档到工件生成的治理统一智能
Dalmia, Shivali, Thoppanahalli, Sumukha, Sediqin, Mohammadreza, Mukherji, Abhishek
Abstract
Enterprise guideline documents are heterogeneous and multimodal, combining narrative text, complex tables, and embedded images. Existing LLM and VLM systems face hallucinated content, table structure degradation, and lack governed workflows extending beyond extraction to validation and artifact generation. This leaves enterprises to perform this manually, consuming 2-3 days per document. To address this, we introduce GUIDE, a governed multi-agent framework built on a shared versioned rule store with schema-validated inter-agent contracts and end-to-end provenance tracking. Six specialized agents handle parsing, VLM-driven extraction, consistency checking, evaluation, human-in-the-loop (HITL) escalation, and persona-tailored artifact synthesis. Evaluated on 120 real-world enterprise guideline documents, GUIDE achieves 96% document success, extracts 3,896 rules with 71.4% auto-approved, produces 812 deployment-ready artifacts, and reduces turnaround to 40-125 minutes per document.
Chinese Translation
企业指导文件是异构和多模态的,结合了叙述文本、复杂表格和嵌入图像。现有的大型语言模型(LLM)和视觉语言模型(VLM)系统面临虚假内容、表格结构退化以及缺乏超越提取到验证和工件生成的治理工作流程。这使得企业不得不手动执行这些任务,每份文件消耗2-3天的时间。为了解决这个问题,我们提出了GUIDE,一个基于共享版本规则存储的治理多代理框架,具有模式验证的代理间合同和端到端的来源追踪。六个专门的代理负责解析、基于VLM的提取、一致性检查、评估、人机协作(HITL)升级和个性化工件合成。在对120份真实世界企业指导文件进行评估时,GUIDE实现了96%的文档成功率,提取了3,896条规则,其中71.4%自动批准,生成了812个可部署的工件,并将每份文件的周转时间缩短至40-125分钟。
cs.AI / 72 / 2608.12150

Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation

谁思考得最好取决于你让他们思考多久:大语言模型评估中的预算依赖排名
de Souza, Rodrigo Guedes, Panisson, Alison R.
Abstract
Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven levels (64--4,096), evaluating four models on three reasoning benchmarks (56,476 inferences). We report four findings: (i) 3--19% of items exhibit non-monotone behavior (accuracy decreasing with more budget), even after controlling for truncation, and this phenomenon is model-specific (cross-model overlap: 6--14%). (ii) Model rankings reverse across budgets on all benchmarks ($p {<} 0.01$, McNemar). (iii) Oracle analysis reveals model complementarity up to $+27.8$pp, most pronounced at constrained budgets. (iv) A budget-aware router captures 14.1% of the oracle gap cross-domain; budget features help within-domain ($+1.6$ to $+5.7$pp) but are domain-specific and hurt transfer ($-1.2$pp). These results argue for budget-conditioned evaluation protocols.
Chinese Translation
大型语言模型的标准评估假设在推理条件下模型排名是稳定的。我们通过在七个级别(64--4,096)中变化令牌生成预算,即模型可以生成的最大令牌数,来挑战这一假设,并在三个推理基准上评估四个模型(共56,476次推理)。我们报告了四个发现:(i)3--19%的项目表现出非单调行为(准确率随着预算增加而下降),即使在控制截断后,这种现象是特定于模型的(跨模型重叠:6--14%)。 (ii)在所有基准上,模型排名在不同预算之间发生逆转($p {<} 0.01$, McNemar)。 (iii)Oracle分析揭示了模型的互补性,最高可达$+27.8$个百分点,在预算受限时最为明显。 (iv)一个预算感知的路由器在跨领域中捕获了14.1%的Oracle差距;预算特征在领域内有帮助($+1.6$到$+5.7$个百分点),但是领域特定的,并且对迁移造成了负面影响($-1.2$个百分点)。这些结果支持预算条件评估协议的必要性。
cs.AI / 73 / 2608.12192

How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models

如何支配您的Oracle预算:蛋白质结构预测模型的实用指导
Kalisz, Aleksandra, Simons, Jack, Sinkovics, Krisztina, Ghenassia, Noam, Surana, Shikha, Moss, Henry, Duckworth, Paul
Abstract
Foundation models for protein structure prediction remain unreliable on certain targets. External oracles can flag and correct these failures, but biological oracles are expensive, making oracle budget a critical constraint. Existing guidance methods, such as FK-steering, DPO, and Best K-of-N sampling, differ in how they spend this budget, yet no systematic comparison exists to guide method selection. To bridge this gap, we benchmark these methods alongside the recently proposed Optimisation Over Outputs (O3), which applies off-the-shelf optimisers within a generative model's latent subspace. We extend the usage of O3 to protein structure prediction models. Overall, our work provides the first practical reference for oracle budget-aware guidance. Our evaluation on two protein targets, calmodulin (1CLL) and E. coli aspartate transcarbamoylase (9EEH), reveals that no single method consistently dominates across all budgets and oracles. Specifically, O3 proves most effective at low oracle budgets, while FK-steering and DPO demonstrate improved performance as the budget increases. We distil these findings into actionable recommendations for practitioners operating under real-world oracle-budget constraints.
Chinese Translation
蛋白质结构预测的基础模型在某些目标上仍然不可靠。外部Oracle可以标记并纠正这些失败,但生物Oracle成本高昂,使得Oracle预算成为一个关键限制。现有的指导方法,如FK-steering、DPO和Best K-of-N采样,在如何支配这一预算上存在差异,但尚无系统比较来指导方法选择。为填补这一空白,我们对这些方法进行了基准测试,并与最近提出的基于输出优化(Optimisation Over Outputs, O3)进行比较,后者在生成模型的潜在子空间中应用现成的优化器。我们将O3的使用扩展到蛋白质结构预测模型。总体而言,我们的工作为关注Oracle预算的指导提供了首个实用参考。我们对两个蛋白质目标,钙调蛋白(calmodulin, 1CLL)和大肠杆菌天冬氨酸转氨酶(E. coli aspartate transcarbamoylase, 9EEH)的评估表明,没有单一方法在所有预算和Oracle中始终占据优势。具体而言,O3在低Oracle预算下最为有效,而FK-steering和DPO在预算增加时表现出更好的性能。我们将这些发现提炼为可操作的建议,以供在现实世界Oracle预算限制下工作的实践者参考。
cs.AI / 74 / 2608.12249

An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS

遗留高性能计算现代化的自主工作流程:转换GAMESS的双电子积分核心
Shen, Yuzhong, Sosonkina, Masha, Xu, Peng, Gordon, Mark S.
Abstract
Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the codebases can be enormous, and across much of computational science the work simply goes undone. We propose an agentic workflow that takes this work on at production scale, and we set out to measure how far such delegation can reach. In this work, three prompt-specialized agent roles operate under a version-controlled specification that the agents themselves authored and revised, while humans hold a small number of gates. The arrangement is kept safe by an exact verification oracle inherited from the domain, and the boundary of safe delegation lies exactly where that oracle stops seeing. We apply the proposed workflow in a case study, converting the two-electron-integral routines of GAMESS (General Atomic and Molecular Electronic Structure System), a mature quantum-chemistry package with a 48-year development history, from fixed-form Fortran 77 to free-form Fortran 2008. The scope of this work was twelve source files, 56,448 lines, and 225 subroutines for computing electron repulsion integrals. The agents ran as three Claude Code roles in isolated worktrees, and the work spanned four Claude model generations. Because the GAMESS group ships a standard test suite whose printed energies its user community treats as canonical, we could adopt bit-for-bit reproduction of those energies as the merge criterion, where a deviation in the twelfth decimal place counts as a failure rather than drift. All twelve source files pass a 51-test validation battery comprising the 49 standard GAMESS tests and two additional calculations, and across 612 test runs the number of chemistry-relevant differences is zero, and every file also passes the Jenkins tests that are used for continuous integration.
Chinese Translation
现代化遗留Fortran是一个庞大的问题:这些转换在个别情况下是常规的,但代码库可能非常庞大,在计算科学的许多领域,这项工作往往未能完成。我们提出了一种自主工作流程,以生产规模处理这项工作,并着手测量这种委托能达到的范围。在本研究中,三个针对提示的专门代理角色在一个版本控制的规范下运作,该规范由代理自身撰写和修订,而人类则控制少量的关卡。该安排通过从领域继承的精确验证神谕保持安全,安全委托的边界正好位于该神谕停止观察的地方。我们在一个案例研究中应用了所提议的工作流程,将GAMESS(通用原子和分子电子结构系统)的双电子积分例程,从固定格式的Fortran 77转换为自由格式的Fortran 2008。该工作的范围包括十二个源文件,56,448行代码,以及225个用于计算电子排斥积分的子例程。代理以三个Claude Code角色在隔离的工作树中运行,工作跨越了四个Claude模型代。由于GAMESS团队提供了一个标准测试套件,其打印的能量被用户社区视为规范,我们可以采用这些能量的逐位重现作为合并标准,其中第十二位小数的偏差被视为失败而非漂移。所有十二个源文件通过了一个包含49个标准GAMESS测试和两个额外计算的51项验证测试,且在612次测试运行中,化学相关的差异数量为零,每个文件也通过了用于持续集成的Jenkins测试。
cs.AI / 75 / 2608.12282

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

VAKRA:评估工具使用政策下的多跳推理跨API和检索
Naik, Ankita Rajaram, Murthi, Anupama, Elder, Benjamin, Huo, Siyu, Gupta, Raavi, Jain, Abhinav, Venkateswaran, Praveen, Adebayo, Abdulhamid, Contractor, Danish
Abstract
Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e\textbf{V}aluating \textbf{A}PI and \textbf{K}nowledge \textbf{R}etrieval \textbf{A}gents), a benchmark of over $8{,}000$ executable APIs across $62$ domains with tasks spanning three settings of increasing difficulty: diverse API interaction styles, multi-hop reasoning over structured APIs, and multi-source reasoning with natural-language tool-use policy constraints. Correctness is verified by re-executing predicted tool calls against live APIs, accommodating multiple valid paths. Using a fixed ReAct harness to isolate model capabilities from agent architecture, we evaluate frontier and open-weight models and find that even the best model achieves only 70.4\% on single-hop endpoint-style tasks and drops to 50--51\% on compositional APIs; performance degrades by over 50\% as reasoning depth increases, and policy-constrained questions expose severe failures (as low as 2.4\% on unanswerable queries). Trace analysis shows failures concentrate at language-mediated reasoning - entity disambiguation, cross-source grounding, rather than tool invocation mechanics. Code is available https://github.com/IBM/VAKRA. Dataset is available https://huggingface.co/datasets/ibm-research/VAKRA
Chinese Translation
在企业环境中部署的智能体必须在结构化API和文档集合之间进行推理,然而现有基准测试在孤立的情况下评估这些能力。我们引入了VAKRA(评估API和知识检索智能体),这是一个涵盖62个领域、超过8000个可执行API的基准,任务涵盖三个逐渐增加难度的设置:多样的API交互风格、对结构化API的多跳推理,以及带有自然语言工具使用政策约束的多源推理。通过对实时API重新执行预测的工具调用来验证正确性,允许多个有效路径。使用固定的ReAct工具来隔离模型能力与智能体架构,我们评估了前沿和开放权重模型,发现即使是最佳模型在单跳端点风格任务上也仅达到70.4%的准确率,而在组合API上下降到50-51%;随着推理深度的增加,性能下降超过50%,而政策约束问题则暴露出严重的失败(在无法回答的查询上低至2.4%)。追踪分析显示失败主要集中在语言中介推理上——实体消歧、跨源基础,而非工具调用机制。代码可在 https://github.com/IBM/VAKRA 获取。数据集可在 https://huggingface.co/datasets/ibm-research/VAKRA 获取。
计算语言学 (Computation and Language)
55
cs.CL / 1 / 2608.11232

Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs

Backtrader-Bench:基于自生成多项选择题的算法交易中大型语言模型代理的基准测试
Zhao, Ruoxi, Raissi, Maziar
Abstract
Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual code execution. We present Backtrader-Bench, a framework with two complementary pipelines. A deterministic multiple-choice question (MCQ) pipeline generates questions from backtest configurations across five trading strategies, 33 templates, and three difficulty tiers, with an independent checker that re-derives every answer. A generator-solver filtering pipeline autonomously mines harder questions: a generator writes questions verified by executable code, converts them to MCQs, and discards any that a no-tool solver can answer without code execution. We evaluate 11 models without tools (10 runs each) and four with-tools configurations on a 30-question curated set. Tool-augmented agents reach 90.0% accuracy in a single pass (GPT-5.5 and Opus 4.7), outperforming the best no-tools baselines (73.0%, averaged over 10 runs) by 17 percentage points. On 38 separately mined questions, no-tools accuracy drops further, with half the models falling to roughly random-chance level (25%). Beyond evaluation, the scalable MCQ infrastructure is designed to produce a training corpus for reinforcement learning, with the ultimate goal of building a specialized agent for quantitative trading workflows.
Chinese Translation
评估大型语言模型(LLM)编码代理在算法交易中的表现是困难的,因为静态基准测试存在数据污染的风险,而数值回测输出需要实际代码执行的真实结果。我们提出了Backtrader-Bench,一个具有两个互补管道的框架。一个确定性的多项选择题(MCQ)管道从五种交易策略、33个模板和三个难度级别的回测配置中生成问题,并配有一个独立的检查器重新推导每个答案。生成-过滤管道自主挖掘更难的问题:生成器编写经过可执行代码验证的问题,将其转换为多项选择题,并丢弃任何无需代码执行即可由无工具求解器回答的问题。我们在一个包含30个问题的精心策划集上评估了11个无工具模型(每个模型进行10次运行)和四个有工具配置。增强工具的代理在单次测试中达到90.0%的准确率(GPT-5.5和Opus 4.7),比最佳的无工具基线(73.0%,平均10次运行)高出17个百分点。在38个单独挖掘的问题上,无工具的准确率进一步下降,半数模型的表现降至大约随机机会水平(25%)。除了评估之外,具有可扩展性的多项选择题基础设施旨在为强化学习生成训练语料库,最终目标是构建一个专门用于量化交易工作流程的代理。
cs.CL / 2 / 2608.11233

Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets

将递归深度重新安装到预训练语言模型中:在两个参数预算下的安装、外推、转移和保留
Shapiro, Mark
Abstract
A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent transition that persists after outcome-only annealing. Qwen2.5-0.5B-Instruct is split into a Prelude, a weight-tied Recurrent Block, and a Coda, with an identity-preserving one-loop path and a re-entry bridge on later loops. At loop 1 the retrofit remains non-inferior to its base on a preregistered ARC battery. Three findings. First, the mechanism is a reusable procedure rather than terminal-answer lookup, and installs at two budgets: 6M trained parameters over frozen base weights and 180M full-block. With intermediate-step supervision, the model computes one task step per loop and persists when only final answers are graded. The adapter matched the full block overall (83.8% versus 84.0%), led through depth 11, and trailed beyond. Verbal fine-tuning reached 79-86% on controlled verbal renderings (zero-shot transfer was minimal), and adapter verbal training begun from the installed mechanism outpaced matched fresh training by 18.6 points, including on a held-out test set. Second, the operation extrapolates to roughly 1.5 times its supervised depth, holding 70% accuracy through depth 18. Third, a same-size scratchpad-trained model matched the recurrent model within its learned horizon but collapsed beyond it. The recurrent model won overall, 84% versus 72%, retained 53% versus 2.5% beyond depth 10, and answered 7.6 times faster. An iterative transformer can therefore perform deeper reasoning in latent space faster than comparable or larger models fine-tuned on the same task, in a system-level comparison. A second task, running the rule in reverse, exposed the limits: the inverse was learnable in isolation, but no continuation acquired it while preserving the installed mechanism and general capability, a catastrophic-interference boundary. Learned depth selection remains open.
Chinese Translation
一个密集的预训练语言模型可以通过递归深度进行改造,并学习一种在仅进行结果退火后仍然存在的迭代潜在转变。Qwen2.5-0.5B-Instruct被分为前奏、权重绑定的递归块和尾声,具有保持身份的一循环路径和后续循环中的重新进入桥。在循环1中,该改造在预注册的ARC电池上不劣于其基础模型。三项发现。首先,该机制是一个可重复使用的过程,而不是终端答案查找,并在两个预算下进行安装:在冻结的基础权重上训练6M参数和180M完整块。通过中间步骤监督,模型在每个循环中计算一个任务步骤,并在仅对最终答案进行评分时保持持续。适配器在整体上与完整块相匹配(83.8%对84.0%),在深度11时领先,而在更深时落后。语言微调在受控语言呈现上达到了79-86%(零-shot转移最小),而从已安装机制开始的适配器语言训练比匹配的新训练快18.6个百分点,包括在一个保留的测试集上。其次,该操作外推到大约1.5倍的监督深度,在深度18时保持70%的准确率。第三,一个相同大小的从零开始训练的模型在其学习的范围内与递归模型相匹配,但在其范围之外崩溃。递归模型在整体上获胜,84%对72%,在深度10之后保留53%对2.5%,并且回答速度快7.6倍。因此,迭代变换器能够在潜在空间中比在同一任务上微调的可比或更大模型更快地进行更深层次的推理,在系统级比较中。第二个任务,即反向运行规则,揭示了限制:逆向是可以在孤立中学习的,但没有继续获得它,同时保留已安装的机制和一般能力,这是一个灾难性干扰的边界。学习的深度选择仍然开放。
cs.CL / 3 / 2608.11236

TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation

TRACE Bench:任务驱动的角色扮演代理检查表评估
Zhang, Jiahui, Zhang, Ziwei, Wang, Yipeng, Liu, Yibo, Pang, Haozhou, Hu, Yikai, Ren, Hongyan, Zhou, Lan, Gan, Qi, Sheng, Kai
Abstract
Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task-driven agentic checklist evaluation framework. It decomposes each role profile offline into a fixed checklist, then uses a User Agent to converse naturally with the target roleplay model while privately updating checklist states from model responses. Scores therefore trace back to checklist items and supporting dialogue turns rather than a black-box holistic impression. For coverage cross-validation, we audit released M2 free-dialogue transcripts from the MiniMax Role-play Benchmark against the same role-derived checklist. The released free-chat transcripts cover only 73.74% of key role-profile points, whereas TRACE Bench reaches 99.91% coverage in fewer turns. Robustness experiments show stable rankings under repeated runs and User Agent replacement. Across 26 models, TRACE Bench reports overall rankings together with capability breakdowns and checklist traces. It also supports Closed-Loop Benchmark Evolution, distilling verification methods proven effective in failed traces so later evaluations can more reliably elicit and examine observed failure modes.
Chinese Translation
角色扮演评估不仅应当给出一个单一的评分,还应揭示测试了哪些角色要求、哪些未能满足以及哪些对话证据支持该判断。我们提出了TRACE Bench,一个任务驱动的代理检查表评估框架。它将每个角色档案离线分解为一个固定的检查表,然后使用用户代理与目标角色扮演模型进行自然对话,同时私下更新来自模型响应的检查表状态。因此,评分追溯到检查表项目和支持的对话轮次,而不是一个黑箱的整体印象。为了进行覆盖交叉验证,我们审计了MiniMax角色扮演基准发布的M2自由对话记录,针对相同的角色派生检查表。发布的自由聊天记录仅覆盖了73.74%的关键角色档案点,而TRACE Bench在更少的轮次中达到了99.91%的覆盖率。稳健性实验表明,在重复运行和用户代理替换下,排名保持稳定。在26个模型中,TRACE Bench报告了整体排名以及能力细分和检查表追溯。它还支持闭环基准演变,提炼在失败追溯中被证明有效的验证方法,以便后续评估能够更可靠地引出和检查观察到的失败模式。
cs.CL / 4 / 2608.11242

Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction

迷失在压缩中:评估上下文压缩下的侧约束损失
Wang, Zhiqi, Zhang, Yichi, Lee, Dongwon, Yang, Yuchen
Abstract
When the context window is under pressure, LLM systems compact prior context to continue ongoing tasks. We identify a class of user-issued instructions, Session Constraints (SCs), such as "do not delete any emails until I confirm," that are meant to constrain LLM's behavior for the remainder of a session but are silently dropped during compaction. To quantify this loss, we introduce COMPINT, an evaluation suite that evaluates compactors across three long-context scenarios: multi-turn chat, agentic trajectory, and long-horizon research. Current compactors retain only 17% of injected SCs on average, and most perform worse than running the same task without compaction. Retention varies sharply with compactor, prompt, context length, SC phrasing, and injection location, showing that the loss is systematic rather than tied to any single setting. We propose an SC-aware extractor that runs alongside the compactor as a plug-and-play module, achieving over 90% retention across all three scenarios without modifying the compactor or LLM. The COMPINT evaluation suite and accompanying implementation are available at https://github.com/ZhiqiEliWang/compaction-integrity.
Chinese Translation
当上下文窗口受到压力时,大型语言模型(LLM)系统会压缩先前的上下文以继续进行中的任务。我们识别出一类用户发出的指令,称为会话约束(Session Constraints, SCs),例如“在我确认之前不要删除任何电子邮件”,这些指令旨在约束LLM在会话剩余时间内的行为,但在压缩过程中被默默丢弃。为了量化这种损失,我们引入了COMPINT,一个评估套件,用于在三种长上下文场景下评估压缩器:多轮对话、代理轨迹和长时间研究。目前的压缩器平均仅保留17%的注入SCs,并且大多数在执行相同任务时的表现不如不进行压缩。保留率因压缩器、提示、上下文长度、SC措辞和注入位置而有显著差异,显示出这种损失是系统性的,而不是与任何单一设置相关。我们提出了一种SC感知提取器,作为即插即用模块与压缩器并行运行,在所有三种场景下实现超过90%的保留率,而无需修改压缩器或LLM。COMPINT评估套件及其实现可在 https://github.com/ZhiqiEliWang/compaction-integrity 获取。
cs.CL / 5 / 2608.11249

Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression

从扩散到压缩:利用扩散语言模型进行无损压缩
Nardone, Angelo, Ferragina, Paolo
Abstract
We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and structured formats such as XML - and by recent advances in neural language model-based compression. In particular, recent LLM-based approaches, whether built on symbol-ranking pipelines or paired with a statistical compressor, have demonstrated compression ratios significantly superior to general-purpose compressors such as zstd, gzip, or bzip on text and code. However, these neural approaches suffer from severe throughput limitations, making them not yet practically usable. For the first time in the context of lossless neural text compression, we introduce Diffusion Language Models (DLMs) as an alternative inference paradigm to autoregressive LLM-based approaches. We argue that replacing autoregressive LLMs with DLMs within the same compression framework could overcome the throughput bottleneck caused by their one-symbol-per-step limitation. However, achieving these improvements requires addressing algorithmic challenges introduced by applying DLMs to lossless compression, where the architecture allows the number and positions of symbols encoded at each forward pass to be decided independently. We design efficient and effective strategies to solve these challenges and evaluate them experimentally against LLM-based and general-purpose compressors on enwik8, a well-established textual benchmark. Our results show that the newly proposed DLM-based framework advances the state of the art in lossless text compression. Moreover, as DLMs are still a relatively young paradigm, recent advances toward increasingly capable and efficient models suggest substantial room for further improvements.
Chinese Translation
我们研究了无损文本压缩的问题,这一研究受到数字文本数据(包括纯文本、源代码以及XML等结构化格式)收集和存储快速增长的推动,以及基于神经语言模型的压缩技术的最新进展。特别是,最近的基于大型语言模型(LLM)的方法,无论是基于符号排序管道还是与统计压缩器配对,都在文本和代码的压缩比方面表现出显著优于通用压缩器(如zstd、gzip或bzip)的性能。然而,这些神经方法存在严重的吞吐量限制,使其尚未在实际中得到有效应用。在无损神经文本压缩的背景下,我们首次引入扩散语言模型(DLM)作为自回归LLM方法的替代推理范式。我们认为,在同一压缩框架内用DLM替代自回归LLM可以克服其每步一个符号限制所造成的吞吐量瓶颈。然而,实现这些改进需要解决将DLM应用于无损压缩所引入的算法挑战,其中架构允许在每次前向传递中独立决定编码符号的数量和位置。我们设计了高效且有效的策略来解决这些挑战,并在著名的文本基准enwik8上对其进行实验评估,与基于LLM和通用压缩器进行比较。我们的结果表明,新的DLM框架在无损文本压缩领域推动了技术的进步。此外,由于DLM仍然是一个相对年轻的范式,最近在更强大和高效模型方面的进展表明,仍有很大的改进空间。
cs.CL / 6 / 2608.11332

Gloss-Free Representation Learning for Cross-Dataset Sign Spotting

无注释表示学习用于跨数据集手势识别
Tüfekcioğlu, Oğuz Akif, Ekin, Ezgi, Çevik, Mustafa Kaan, Keles, Hacer Yalim
Abstract
Sign-language research for resource-constrained languages is often limited by the cost of dense linguistic labels such as glosses, temporal boundaries, and sign order. Broadcast news offers a practical alternative by pairing continuous signing with spoken-language transcripts, but this supervision is weak since text and signing are loosely aligned. Morphologically rich languages such as Turkish add further difficulty, as the same lexical meaning can appear in many inflected forms while some derived forms should remain distinct. We study whether weak transcript-based supervision can pretrain a reusable sign encoder in this setting, where poor text normalization can fragment pseudo-gloss targets and weaken representation learning. Unlike prior pseudo-gloss pipelines designed mainly to improve translation, we test whether the pretrained encoder transfers as a reusable representation for cross-dataset sign spotting. We pretrain on TSL-News, a new Turkish broadcast corpus, using pseudo-gloss labels derived from transcripts rather than manual annotation, comparing rule-based morphological lemmatization with constrained LLM-assisted normalization over a fixed vocabulary. We evaluate the learned representations via cross-dataset sign spotting on a new TSL Spotting Benchmark built from the TSL Dictionary corpus. The LLM-assisted encoder raises top-5 temporal localization mean IoU from 0.235 to 0.465, with 56.2% of examples reaching an IoU of at least 0.50; a frequency analysis suggests this gain is not mainly driven by memorizing frequent pseudo-gloss labels. In a downstream translation check, the same pretraining improves BLEU-4 from 9.60 to 11.04 and ROUGE from 23.48 to 27.43. These results show that loosely aligned broadcast data can provide effective weak supervision for learning sign representations that capture both lexical content and temporal structure.
Chinese Translation
针对资源受限语言的手语研究常常受到密集语言标签(如注释、时间边界和手势顺序)成本的限制。广播新闻通过将连续手势与口语转录配对,提供了一种实用的替代方案,但这种监督是弱的,因为文本和手势之间的对齐较为松散。形态丰富的语言(如土耳其语)增加了进一步的难度,因为相同的词汇意义可以以多种屈折形式出现,而某些派生形式应保持独特。我们研究在这种情况下,基于弱转录的监督是否可以预训练一个可重用的手势编码器,其中不良的文本规范化可能会破坏伪注释目标并削弱表示学习。与之前主要旨在改善翻译的伪注释管道不同,我们测试预训练的编码器是否可以作为跨数据集手势识别的可重用表示。我们在TSL-News(一个新的土耳其广播语料库)上进行预训练,使用从转录中派生的伪注释标签,而不是手动标注,比较基于规则的形态词干提取与在固定词汇上受限的LLM辅助规范化。我们通过在基于TSL词典语料库构建的新TSL Spotting Benchmark上进行跨数据集手势识别来评估学习到的表示。LLM辅助编码器将前五个时间定位的平均IoU从0.235提高到0.465,56.2%的示例达到至少0.50的IoU;频率分析表明,这一增益并非主要由记忆频繁的伪注释标签驱动。在下游翻译检查中,相同的预训练使BLEU-4从9.60提高到11.04,ROUGE从23.48提高到27.43。这些结果表明,松散对齐的广播数据可以为学习捕捉词汇内容和时间结构的手势表示提供有效的弱监督。
cs.CL / 7 / 2608.11338

Better, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent Cost

更好、更快、更强:程序化技能学习最佳降低代理成本
Huang, Zixi, Wang, Xiheng, Wang, Andrew, Jurayj, William, Gutiérrez, Bernal Jiménez, Khashabi, Daniel, Andrews, Nicholas
Abstract
Recently, the practice of augmenting LLM agent capability with skills has gained prevalence. We explore the cost effective adaptation of agents to novel domains by means of learning skills. Existing works focus on performance gain over cost effectiveness. As a result, little is known about what skill learning strategies save cost. We argue that among all the different skill learning methods, those that view skills as programs can achieve the best cost reduction. By executing sequences of actions deterministically, a program-augmented agent can reliably and cheaply achieve goals that would otherwise require trial and error and risk degenerate behavior over long horizons. An agent can learn at inference time by incrementally discovering these programs and equipping them for future tasks. We hypothesize that past trajectories contain enough signal to guide skill learning, even without replay or validation, provided the agent can learn to analyze them. To test our claims, we propose SpeedRunner, a coding agent that analyzes trajectories and refactors skills for better performance on future tasks. Across three different embodied environments, we show that SpeedRunner consistently achieves the frontier in learning and cost reduction while remaining robust against distribution shifts and environmental randomness.
Chinese Translation
最近,通过技能增强大型语言模型(LLM)代理能力的做法越来越普遍。我们探讨了通过学习技能以经济有效的方式将代理适应于新领域。现有研究主要关注性能提升而忽视成本效益。因此,关于哪些技能学习策略能够节省成本的知识仍然有限。我们认为,在所有不同的技能学习方法中,将技能视为程序的方法能够实现最佳的成本降低。通过确定性地执行一系列动作,程序增强的代理可以可靠且低成本地实现目标,而这些目标在其他情况下可能需要试错并且在长时间内存在退化行为的风险。代理可以在推理时通过逐步发现这些程序并为未来任务装备它们来进行学习。我们假设过去的轨迹包含足够的信号来指导技能学习,即使没有重放或验证,只要代理能够学习分析这些轨迹。为了验证我们的观点,我们提出了SpeedRunner,一个分析轨迹并重构技能以提高未来任务表现的编码代理。在三个不同的具身环境中,我们展示了SpeedRunner在学习和成本降低方面始终处于前沿,同时对分布变化和环境随机性保持稳健。
cs.CL / 8 / 2608.11350

Self-Evolving Embodied Agents via Skill-Harness Evolution

通过技能利用进化实现自我进化的具身智能体
Wang, Peidong, Ma, Zhiming, Chang, Ying, Luo, Xufang, Yang, Xiaocui, Feng, Shi, Yang, Yuqing, Li, Dongsheng
Abstract
Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills, context, action interfaces, and execution harness surrounding the model. While supervised fine-tuning and reinforcement learning can adapt agents to new environments, they require additional data, rewards, and training runs; meanwhile, many train-free code-centric approaches rely on programmable robot APIs that may be unavailable in fixed-interface settings. We propose SHAPER, a self-evolving framework for train-free embodied adaptation that keeps model parameters frozen and improves the non-parametric agent system by evolving reusable skills and a context-code harness through target-environment rollouts. In SHAPER, the same frozen model can serve as both planner and optimizer, refining its external skills and context-code harness without parameter updates. We evaluate SHAPER on VLABench and ESI-Bench, covering embodied agents with different low-level action interfaces, and compare against pure execution, supervised fine-tuning, and test-time-scaling baselines such as verifier-free selection and voting. Our results suggest that skill-and-harness optimization is a practical route to self-evolving embodied agents when model training is expensive, unavailable, or undesirable.
Chinese Translation
具身智能体越来越多地被构建为围绕基础模型的系统,其性能不仅依赖于模型权重,还依赖于技能、上下文、动作接口和围绕模型的执行环境。虽然监督微调和强化学习可以使智能体适应新环境,但它们需要额外的数据、奖励和训练过程;与此同时,许多无训练的代码中心方法依赖于可编程的机器人API,而在固定接口设置中可能无法使用。我们提出了SHAPER,一个用于无训练具身适应的自我进化框架,该框架保持模型参数不变,并通过目标环境的回滚进化可重用技能和上下文代码环境,从而改善非参数智能体系统。在SHAPER中,相同的冻结模型可以同时作为规划者和优化器,精炼其外部技能和上下文代码环境,而无需更新参数。我们在VLABench和ESI-Bench上评估SHAPER,涵盖具有不同低级动作接口的具身智能体,并与纯执行、监督微调和测试时扩展基线(如无验证选择和投票)进行比较。我们的结果表明,当模型训练成本高昂、不可用或不理想时,技能和环境优化是实现自我进化具身智能体的可行途径。
cs.CL / 9 / 2608.11352

ODE-Based Transformer Decoders for Iterative Sign Language Translation

基于常微分方程的变换器解码器用于迭代手语翻译
Kızıltepe, Tuğçe, Keles, Hacer Yalim
Abstract
Sign language translation has achieved strong results with Transformer architectures, yet recent improvements largely rely on scaling model capacity at the cost of increased computation. We propose a parameter-efficient alternative that improves expressiveness without increasing model size. Rather than scaling capacity, we focus on enhancing the update dynamics of iterative refinement decoders, where each refinement step corresponds to one internal decoder iteration that progressively improves the latent representation before translation generation. We reinterpret residual refinement updates from an Ordinary Differential Equation (ODE) perspective and replace them with higher-order numerical integration schemes, namely Runge--Kutta methods (RK-2 and RK-4). These methods perform multiple function evaluations within each refinement step to produce more accurate and stable representation updates without adding decoder parameters. To the best of our knowledge, this is the first application of ODE-inspired update dynamics to sign language translation. RK-2 achieves 22.96 BLEU-4 on the PHOENIX-2014-T test set and 19.34 BLEU-4 on the CSL-Daily test set, outperforming the IPSLT baseline on both benchmarks, with fewer decoder layers and refinement iterations on CSL-Daily. These results suggest that stronger refinement dynamics can improve translation performance under parameter-efficient decoder designs, providing a complementary alternative to conventional model scaling.
Chinese Translation
手语翻译在变换器架构下取得了显著成果,但最近的改进在很大程度上依赖于扩大模型容量,这导致了计算成本的增加。我们提出了一种参数高效的替代方案,在不增加模型规模的情况下提高表达能力。我们关注于增强迭代细化解码器的更新动态,而不是简单地扩大容量,其中每个细化步骤对应于一次内部解码器迭代,该迭代逐步改善潜在表示,以便生成翻译。我们从常微分方程(ODE)的角度重新解释残差细化更新,并用更高阶的数值积分方案替代它们,即龙格-库塔方法(RK-2 和 RK-4)。这些方法在每个细化步骤内执行多次函数评估,以产生更准确和稳定的表示更新,而无需增加解码器参数。根据我们所知,这是首次将受ODE启发的更新动态应用于手语翻译。RK-2在PHOENIX-2014-T测试集上达到了22.96的BLEU-4分数,在CSL-Daily测试集上达到了19.34的BLEU-4分数,在这两个基准上均优于IPSLT基线,并且在CSL-Daily上使用了更少的解码器层和细化迭代。这些结果表明,更强的细化动态可以在参数高效的解码器设计下提高翻译性能,为传统模型扩展提供了一种互补的替代方案。
cs.CL / 10 / 2608.11408

Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning

测量,而非优化:大语言模型去学习中的恢复预测
Song, Zirui, Liu, Huaxing, Wang, Xiang, Li, Shuai, Li, Xinye, Gao, Lang, Zhang, Jinghui, Lu, Zheng, Ji, Fengxian, Chang, Xiaojun, Chen, Xiuying
Abstract
Prior white-box studies show that large language models can retain latent traces of target knowledge after unlearning, even when the knowledge is no longer expressed in their outputs. However, existing audits remain limited to one-off diagnostics: it is unclear whether these residual signals can predict future recovery under continued training or serve as reliable optimization targets. Resolving this gap is essential to determine whether internal auditing can move beyond post-hoc evaluation toward proactive risk monitoring and safer unlearning. We propose J-Access, an inference-time audit that uses the Jacobian lens to map intermediate representations into vocabulary space and measures how often target concepts remain accessible along the model's output pathway. We hypothesize that residual accessibility reflects recovery susceptibility: knowledge that remains closer to the output pathway requires less fine-tuning to restore, leading to faster recovery. We audit 398 public unlearned models spanning eight unlearning methods. We find that: (1) most unlearned models retain access above the retain-only gold level; (2) pre-attack accessibility predicts recovery speed and extent at the model level, but cannot identify which specific facts will be recovered; and (3) directly minimizing J-Access does not promote genuine deletion. Instead, the model learns to hide knowledge from the audit, producing lower audit scores but greater post-attack recovery. These findings position J-Access as a model-level diagnostic for assessing residual susceptibility in unlearned models. We argue internal audits should serve as an independent diagnostic dimension in unlearning evaluation, and should not be converted into optimization targets without validation.
Chinese Translation
先前的白盒研究表明,大语言模型在去学习后仍能保留目标知识的潜在痕迹,即使这些知识不再在其输出中表达。然而,现有的审计仍然局限于一次性诊断:尚不清楚这些残余信号是否可以预测在持续训练下的未来恢复,或是否可以作为可靠的优化目标。解决这一空白对于确定内部审计是否能够超越事后评估,朝向主动风险监控和更安全的去学习至关重要。我们提出了 J-Access,这是一种推理时审计方法,利用雅可比矩阵(Jacobian)视角将中间表示映射到词汇空间,并测量目标概念在模型输出路径中保持可访问的频率。我们假设残余可访问性反映了恢复的易感性:与输出路径更接近的知识需要更少的微调即可恢复,从而导致更快的恢复。我们审计了398个公共去学习模型,涵盖八种去学习方法。我们的发现包括:(1)大多数去学习模型的可访问性高于仅保留黄金水平;(2)攻击前的可访问性可以预测模型层面的恢复速度和程度,但无法识别哪些具体事实将被恢复;以及(3)直接最小化 J-Access 并未促进真正的删除。相反,模型学会了隐藏知识以避免审计,导致审计分数降低但攻击后的恢复更大。这些发现将 J-Access 定位为评估去学习模型中残余易感性的模型级诊断工具。我们认为内部审计应作为去学习评估中的独立诊断维度,而不应在未经验证的情况下转化为优化目标。
cs.CL / 11 / 2608.11426

Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models

收敛是不可避免的吗?追溯输出同质性至基础模型
Fortier, Alexandrine, Chen, Hazel, West, Peter
Abstract
The lack of diversity in LM content is widely attributed to the alignment process, but how and where exactly in the pipeline this collapse begins is unknown. We argue that output homogeneity is likely learned during the pretraining phase, and only \emph{revealed} or magnified during the alignment process. Specifically, we find that semantic convergence is observed from the first alignment stage--the instruction-tuning phase (SFT)--suggesting that homogeneity might already exist in the pre-alignment model. To investigate this, we conduct controlled SFT experiments examining how training data influences output convergence on specific input/output pairs. We find that convergence can be revealed and amplified, but not introduced by the SFT data, supporting its role as a catalyst rather than a cause. To further test whether homogeneity originates before alignment, we measure convergence in base models. We find that instruct-like collapse can be induced through prompting alone, even without alignment. Taken together, our results suggest that semantic convergence may arise naturally from the objectives underlying LM training, making it difficult to mitigate through post-alignment interventions alone.
Chinese Translation
语言模型(LM)内容缺乏多样性被广泛归因于对齐过程,但这一崩溃究竟在管道的何处开始仍不清楚。我们认为,输出同质性可能在预训练阶段就已学习到,并且仅在对齐过程中被 extit{揭示}或放大。具体而言,我们发现语义收敛从第一个对齐阶段——指令微调阶段(SFT)开始就已观察到,这表明同质性可能在对齐模型之前就已存在。为此,我们进行控制的SFT实验,研究训练数据如何影响特定输入/输出对的输出收敛。我们发现收敛可以被揭示和放大,但不能由SFT数据引入,这支持了其作为催化剂而非原因的角色。为了进一步测试同质性是否在对齐之前产生,我们测量了基础模型中的收敛性。我们发现,仅通过提示就可以诱导出类似指令的崩溃,即使没有对齐。综合来看,我们的结果表明,语义收敛可能自然源于语言模型训练的目标,使得仅通过对齐后的干预来缓解这一现象变得困难。
cs.CL / 12 / 2608.11433

Stigma and Support in Online Sexual Violence Narratives on Reddit

Reddit上在线性暴力叙事中的污名与支持
Bandela, Shirlene Rose, Bindal, Karan, Garg, Vaibhav, Rezapour, Rezvaneh
Abstract
Online communities increasingly provide spaces where survivors of sexual violence can share their experiences and seek support. Although prior research has examined stigma and social support separately, less is known about how stigma expressed in survivor narratives relates to the support offered in response. We introduce the SCOPE dataset, linking stigma signals in online survivor narratives to support types in corresponding comment threads. We annotate posts using a multi-dimensional stigma taxonomy, including Experienced, Internalized, Anticipated, and Structural Stigma, and comments using a support taxonomy encompassing Information Support, Emotional Support, Esteem Support, Tangible Assistance, and Group Interaction. Using contextual, linguistic, and emotion analyses, we compare Stigma and No Stigma content and find that Stigma narratives place greater emphasis on internalized distress, whereas No Stigma narratives focus more on interpreting situations and experiences. Internalized Stigma is the most prevalent category, and community responses remain broadly stable across stigma types, with Information and Esteem Support appearing most often. These findings show how stigma shapes survivor narratives and peer responses and have implications for computational modeling, content moderation, and safer online systems.
Chinese Translation
在线社区日益成为性暴力幸存者分享经历和寻求支持的空间。尽管先前的研究分别考察了污名和社会支持,但关于幸存者叙事中表达的污名与相应评论中提供的支持之间关系的研究相对较少。我们引入了SCOPE数据集,将在线幸存者叙事中的污名信号与相应评论线程中的支持类型关联起来。我们使用多维污名分类法对帖子进行注释,包括经历的污名、内化的污名、预期的污名和结构性污名,并使用支持分类法对评论进行注释,包括信息支持、情感支持、自尊支持、实质性帮助和群体互动。通过上下文、语言和情感分析,我们比较了污名和无污名内容,发现污名叙事更强调内化的痛苦,而无污名叙事则更侧重于对情境和经历的解释。内化污名是最普遍的类别,社区反应在不同污名类型之间保持大致稳定,其中信息支持和自尊支持出现得最为频繁。这些发现展示了污名如何塑造幸存者叙事和同伴反应,并对计算建模、内容审核和更安全的在线系统具有重要意义。
cs.CL / 13 / 2608.11441

DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition

DonorRank:低资源跨语言语音识别中的捐赠语言选择
Dhasmana, Akriti, Srivastava, Aarohi, Chiang, David
Abstract
Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains challenging for spontaneous speech from under-resourced language communities, due to linguistic variation, evolving orthographic conventions, and uneven resource availability. We present DonorRank, a learning-to-rank framework for predicting effective donor languages for zero-shot ASR. We evaluate DonorRank on two multilingual speech corpora of Indic and African language families. It accurately predicts donor language rankings and improves donor selection over common heuristics based on genetic similarity or high-resource languages. Beyond improving transfer, we show how DonorRank is a general framework for analyzing donor language selection itself. Our analyses show that the composition of the donor set determines which linguistic cues are useful in predicting successful transfer. We also identify transfer patterns that provide practical guidance for multilingual ASR in low-resource settings.
Chinese Translation
低资源自动语音识别(ASR)通常依赖于跨语言迁移,即从高资源捐赠语言中适应模型。然而,对于来自资源匮乏语言社区的自发语音,选择捐赠语言仍然具有挑战性,这主要是由于语言变异、不断演变的正字法规范以及资源可用性的差异。我们提出了DonorRank,一个用于预测零样本ASR有效捐赠语言的学习排序框架。我们在两个印度语言和非洲语言家族的多语言语音语料库上评估了DonorRank。它准确预测了捐赠语言的排名,并在选择捐赠语言时优于基于遗传相似性或高资源语言的常见启发式方法。除了改善迁移效果外,我们还展示了DonorRank作为分析捐赠语言选择本身的通用框架。我们的分析表明,捐赠集的组成决定了哪些语言线索在预测成功迁移时是有用的。我们还识别出一些迁移模式,为低资源环境中的多语言ASR提供了实用指导。
cs.CL / 14 / 2608.11460

Principal Trait Analysis: Towards Deriving "Skills" in Human-AI Collaboration

主特征分析:揭示人机协作中的“技能”
McNichols, Hunter, Du, Kai, Lan, Andrew
Abstract
Large Language Model-powered agents are increasingly used in the workplace via human-artificial intelligence (AI) collaboration. In this new era of work, it is important to understand the kinds of prompting traits that contribute to task success. Moreover, we need to uncover key skills required for modern professionals and inform educators on how to foster these skills among students. Existing guidelines for human-AI collaboration are built from either top-down theory or context-specific observations of human-AI interactions. However, since LLM capabilities are rapidly improving, theory may not be able to explain emerging interaction patterns, and empirical guidelines may become obsolete quickly. In this work, we explore an automated, data-driven approach to uncover patterns, which we term traits, of effective human-AI interaction that are aligned with task outcomes. We propose Principal Trait Analysis, a Principal Component Analysis-inspired algorithm for deriving common traits from patterns in LLM conversations. Our algorithm uses LLM-based processing stages to analyze corpora of human-AI collaborative session traces, deriving common traits across the dataset and scoring each human collaborator's usage style by each trait. The approach also allows domain expertise to be injected during trait discovery and selects the most distinguishing traits to be those that exhibit the highest variance across collaborators. We evaluate PTA on two human-AI collaborative coding datasets, an educational setting (students working with an AI tutor) and a professional setting (developers working with an AI coding agent). We find that PTA-derived traits are significant in explaining collaborator behavior across both settings and can help predict task outcomes. However, whether traits qualify as skills remains to be seen, due to inconclusive results on generalizability and how user traits change over time.
Chinese Translation
基于大型语言模型的代理在工作场所中通过人类与人工智能(AI)协作的方式被越来越多地使用。在这一新的工作时代,理解促成任务成功的提示特征变得尤为重要。此外,我们需要揭示现代专业人士所需的关键技能,并为教育工作者提供如何在学生中培养这些技能的指导。现有的人机协作指南主要基于自上而下的理论或特定情境下的人机交互观察。然而,由于大型语言模型(LLM)的能力正在迅速提升,理论可能无法解释新兴的交互模式,而经验性指南可能很快变得过时。在本研究中,我们探索了一种自动化的数据驱动方法,以揭示与任务结果相关的有效人机交互模式,我们称之为特征。我们提出了主特征分析(Principal Trait Analysis,PTA),这是一种受主成分分析启发的算法,用于从LLM对话模式中提取共同特征。我们的算法利用基于LLM的处理阶段分析人机协作会话记录的语料库,从数据集中提取共同特征,并根据每个特征对每位人类协作者的使用风格进行评分。该方法还允许在特征发现过程中注入领域专业知识,并选择在协作者之间表现出最高方差的最具区分性的特征。我们在两个人机协作编码数据集上评估PTA,一个是教育环境(学生与AI导师合作),另一个是专业环境(开发者与AI编码代理合作)。我们发现PTA提取的特征在解释两个环境中的协作者行为方面具有显著性,并且可以帮助预测任务结果。然而,特征是否符合技能的标准仍有待观察,因为关于可推广性和用户特征随时间变化的结果尚不确定。
cs.CL / 15 / 2608.11528

Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment

群体对齐引发的谄媚行为:可调节多元对齐的双面评估
Zhao, Haokai, Xiao, Yunze, Xuan, Weihao, Salim, Flora, Tag, Benjamin, Joshi, Aditya
Abstract
Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-product of alignment, causes the model to over-agree with the user regardless of factual and objective information. However, existing group alignment methods and evaluations focus only on how closely the model matches the group's opinions, overlooking the induced change in sycophantic behaviour. To bridge this gap, we introduce \textbf{G}roup \textbf{A}lignment-induced \textbf{S}ycophancy (GAS) and systematically evaluate alignment across 3 methods, 4 models and 13 demographic groups, on both the intended gain in opinion alignment and the unintended shift in sycophancy. We find that gain and shift are non-uniform across groups: under an identical budget, some groups receive larger gains in opinion alignment than others, and the induced sycophancy shift forms a group-specific profile rather than a single-dimensional change. These results suggest that group alignment should be reported as a two-sided, multi-dimensional profile rather than a single fit score that accounts for per-group differences when adapting LLMs to diverse populations.
Chinese Translation
群体对齐将语言模型调整为特定的人口群体,以生成反映该群体意见、价值观和偏好的响应。谄媚行为是对齐的一个已被充分记录的副产品,它导致模型过度同意用户的观点,而不考虑事实和客观信息。然而,现有的群体对齐方法和评估仅关注模型与群体意见的匹配程度,忽视了谄媚行为的变化。为了解决这一问题,我们引入了 extbf{G}roup extbf{A}lignment-induced extbf{S}ycophancy (GAS) 并系统地评估了3种方法、4个模型和13个人口群体的对齐情况,既包括期望的意见对齐增益,也包括意外的谄媚行为变化。我们发现,增益和变化在不同群体之间并不均匀:在相同的预算下,一些群体在意见对齐上获得的增益大于其他群体,而引发的谄媚行为变化形成了特定于群体的特征,而不是单一维度的变化。这些结果表明,群体对齐应作为双面、多维的特征进行报告,而不是仅仅作为一个适应于不同群体的单一适配分数。
cs.CL / 16 / 2608.11531

On Weak Bisimilarities in CCSK

关于CCSK中的弱相似性
Vallée, Baptiste, Lanese, Ivan
Abstract
In the context of CCSK, a reversible extension of CCS, we study different notions of bisimilarity (strong/weak, forward-only/reversible) and highlight their differences and commonalities. In particular, for the weak reversible case, not previously studied in the literature, we propose two variants, dubbed directional and mixed bisimilarity, depending on whether $\tau$ actions should be in the same direction (forward/backward) as the action being matched or not. We show, in particular, that mixed bisimilarity is a congruence and completely abstracts away from $\tau$ actions.
Chinese Translation
在CCSK的背景下,CCSK是CCS的可逆扩展,我们研究了不同的相似性概念(强/弱、仅前向/可逆),并强调它们之间的差异和共性。特别是,对于文献中未曾研究的弱可逆情况,我们提出了两种变体,称为方向性相似性和混合相似性,这取决于$ au$动作是否应与被匹配的动作在相同方向(前向/后向)上。我们特别展示了混合相似性是一种同构,并完全抽象掉了$ au$动作。
cs.CL / 17 / 2608.11534

CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models

CT-$ riangle$Bench:一种用于纵向3D医学影像差异报告的基准测试,结合视觉-语言模型
Tang, Kegeng, Wang, Jingbo, Ren, Shaogang, Wang, Zihao
Abstract
In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to determine disease evolution, a process that underpins response assessment, recurrence detection, and ongoing patient management. Yet, despite this central role of temporal comparison in clinical decision-making, existing medical foundation models remain largely confined to single-study understanding, leaving temporally grounded cross-examination insufficiently addressed. To address this gap, we study longitudinal imaging difference reporting, a task in which a model takes two temporally separated scans from the same patient and generates a clinically meaningful report describing interval changes between them. We introduce CT-$\Delta$Bench, a dedicated benchmark for this task with patient-level splitting to prevent information leakage. To better evaluate this task beyond surface-level text similarity, we further develop change-aware metrics specifically designed to capture clinically meaningful longitudinal changes, and conduct an independent physician validation to assess the reliability of the synthesized references and event extraction pipeline. We also compare direct paired-CT reasoning with an indirect two-stage pipeline that first generates single-timepoint reports and then performs textual differencing. Finally, we propose DeltaMed, a baseline model for direct paired-CT difference reporting, and train it on the benchmark training set. Together, these contributions lay the groundwork for temporally aware medical foundation models that better reflect real-world longitudinal clinical reasoning.
Chinese Translation
在医学影像学中,计算机断层扫描(CT)的临床价值不仅在于描绘当前的疾病状态,更在于能够对连续扫描进行纵向比较,以确定疾病的演变,这一过程是评估治疗反应、检测复发和持续患者管理的基础。然而,尽管时间比较在临床决策中扮演着核心角色,现有的医学基础模型仍然主要局限于单一研究的理解,导致时间上相关的交叉检验未得到充分解决。为了解决这一问题,我们研究了纵向影像差异报告,这是一项任务,其中模型接受来自同一患者的两个时间上分离的扫描,并生成描述它们之间间隔变化的临床意义报告。我们引入了CT-$ riangle$Bench,这是一个专门针对该任务的基准测试,采用患者级别的拆分以防止信息泄露。为了更好地评估这一任务,超越表面文本相似性,我们进一步开发了专门设计的变化感知指标,以捕捉具有临床意义的纵向变化,并进行独立的医生验证,以评估合成参考和事件提取流程的可靠性。我们还比较了直接配对CT推理与间接的两阶段流程,后者首先生成单时间点报告,然后进行文本差异比较。最后,我们提出了DeltaMed,一个用于直接配对CT差异报告的基线模型,并在基准训练集上对其进行了训练。这些贡献为时间感知的医学基础模型奠定了基础,使其更好地反映现实世界中的纵向临床推理。
cs.CL / 18 / 2608.11552

Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents

超越单轮信心:针对LLM代理的轨迹适应性不确定性量化
Bouchard, Dylan, Chauhan, Mohit Singh
Abstract
Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate to the final outcome. We study whether three common families of single-turn UQ methods transfer to this setting. Across five LLMs and four multi-turn tool-use datasets from BFCL-v4 and $\tau^2$-bench, we evaluate white-box scorers based on action-token probabilities, black-box consistency scorers based on resampled trajectories, and reflexive scorers based on model self-assessment of the trajectory. We find that transfer is often useful but uneven. Token-probability scores are highly sensitive to the choice of aggregator used across turns, reflexive scores provide the strongest low-cost baseline in most evaluated settings, and black-box self-consistency is often the strongest UQ family, with trajectory-equivalence and action-set consistency typically ranking highest among its variants. These results suggest that UQ methods developed for single generations should be revalidated at the trajectory level, with careful attention to the consistency measurement, aggregator choice, and computational budget.
Chinese Translation
语言模型的不确定性量化(UQ)方法通常在单轮输出上进行评估,其中不确定性与生成的单一答案相关。然而,对于LLM代理而言,观察的单位是交互轨迹,在此过程中,模型可以提出澄清问题、调用工具、更新状态并做出中间决策,这些决策的错误会传播到最终结果。我们研究了三种常见的单轮UQ方法是否能够转移到这种设置中。在五个LLM和来自BFCL-v4及$ au^2$-bench的四个多轮工具使用数据集上,我们评估了基于动作令牌概率的白盒评分器、基于重采样轨迹的黑盒一致性评分器,以及基于模型自我评估轨迹的反射评分器。我们发现,转移通常是有用的,但并不均匀。令牌概率评分对跨轮使用的聚合器选择高度敏感,反射评分在大多数评估设置中提供了最强的低成本基线,而黑盒自一致性通常是最强的不确定性量化方法,其中轨迹等价性和动作集一致性通常在其变体中排名最高。这些结果表明,针对单次生成开发的UQ方法应在轨迹层面重新验证,并需仔细关注一致性测量、聚合器选择和计算预算。
cs.CL / 19 / 2608.11573

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

强化步骤级推理以实现大型语言模型的有效自我纠正
Anh, Vu Duc, Hoang, Nhat M., Long, Do Xuan, Nguyen, Cong-Duy, Srey, Ponhvoan, Tuan, Luu Anh
Abstract
Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose Self-Fix Step-DPO (SFS-DPO), a reinforcement learning based, two-stage framework for step-level self-verification and self-correction. The first stage strengthens step-level reasoning via step-level preference optimization, while the second stage explicitly trains models to self-verify and self-correct. We further introduce a teacher-assisted variant, SFS-DPO-R, which incorporates explanatory rationales for error verification to provide stronger corrective signals. Comprehensive in-domain and out-of-domain evaluations across multiple LLMs demonstrate that SFS-DPO and SFS-DPO-R consistently outperform prior step-level training baselines. Our analysis further reveals improvements in self-correction frequency and effectiveness, highlighting the importance of strengthening step-level reasoning for robust performance.
Chinese Translation
实现有效的自我纠正,即模型能够验证并纠正自身错误,仍然是大型语言模型(LLMs)面临的一项基本挑战。在本研究中,我们提出了基于强化学习的自我修复步骤-DPO(Self-Fix Step-DPO,SFS-DPO)框架,这是一个两阶段的步骤级自我验证和自我纠正的方法。第一阶段通过步骤级偏好优化来增强步骤级推理,而第二阶段则明确训练模型进行自我验证和自我纠正。我们进一步引入了一种教师辅助变体SFS-DPO-R,该变体结合了解释性理由以进行错误验证,从而提供更强的纠正信号。针对多个LLMs的全面领域内和领域外评估表明,SFS-DPO和SFS-DPO-R在性能上始终优于先前的步骤级训练基线。我们的分析进一步揭示了自我纠正频率和有效性的改善,强调了强化步骤级推理对于稳健性能的重要性。
cs.CL / 20 / 2608.11624

Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

学习说服揭示了大型语言模型(LLMs)多么容易放弃正确信念
Bozdag, Nimet Beyza, Acikgoz, Emre Can, Tur, Gokhan, Hakkani-Tür, Dilek
Abstract
Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior. Yet we show that this requirement is far from met: a single targeted persuasive argument is enough to collapse model accuracy to near zero, even when the argument is factually false. We formalize this threat as adversarial persuasion and introduce an adversarial reinforcement learning framework that trains persuader agents to change a target model's answer in a single interaction. First, we show that optimizing persuasion strategies through trial and error exposes vulnerabilities that static prompting misses: RL-trained persuaders raise persuasion success from approximately 24% to over 93% against the training-time persuadee. Second, we find that these learned strategies transfer to unseen models, achieving 83% attack success on Qwen-14B, 79% on Llama-3.1-8B, and 25% on GPT-4o-mini. Third, we demonstrate that a curriculum that bootstraps on more persuadable open-weight models before targeting harder models further increases GPT-4o-mini attack success from 25% to 38%. Moreover, our results reveal that optimized persuaders increasingly rely on credibility-based tactics, including fabricated citations and false authoritative evidence. Together, these findings expose a critical weakness in current LLM agents: even when they initially reason correctly, they can be steered toward false conclusions by optimized natural language influence. This positions persuasion robustness as a necessary safety criterion for multi-agent and human-AI decision-making systems.
Chinese Translation
说服是自然语言交流的核心动态,塑造了大型语言模型(LLMs)如何更新信念、解决分歧和做出决策。随着LLMs越来越多地与人类及彼此进行辩论、建议和协作,抵制有害说服成为可靠行为的核心要求。然而,我们的研究表明,这一要求远未得到满足:一个有针对性的说服论点足以使模型的准确性降至接近零,即使该论点在事实层面上是错误的。我们将这一威胁形式化为对抗性说服,并引入一个对抗性强化学习框架,训练说服代理在单次交互中改变目标模型的回答。首先,我们展示了通过试错优化说服策略暴露出静态提示所忽视的脆弱性:经过强化学习训练的说服者在训练时的被说服者面前将说服成功率从约24%提高到超过93%。其次,我们发现这些学习到的策略能够转移到未见过的模型上,在Qwen-14B上实现83%的攻击成功率,在Llama-3.1-8B上为79%,在GPT-4o-mini上为25%。第三,我们证明了一种课程设计,通过在针对更难模型之前先在更容易被说服的开放权重模型上进行引导,进一步将GPT-4o-mini的攻击成功率从25%提高到38%。此外,我们的结果揭示,优化后的说服者越来越依赖于基于可信度的策略,包括虚构的引用和虚假的权威证据。综合来看,这些发现暴露了当前LLM代理的一个关键弱点:即使它们最初推理正确,也可以通过优化的自然语言影响引导至错误结论。这使得说服的鲁棒性成为多代理和人机决策系统的必要安全标准。
cs.CL / 21 / 2608.11629

Easper: An Accessible ASR Pipeline for Language Documentation

Easper:一种可访问的语言文献记录自动语音识别管道
Mahmudi, Aso, Dang, Ting, Vylomova, Ekaterina, Thieberger, Nick
Abstract
Audio transcription is a critical bottleneck in language documentation. While multilingual Automatic Speech Recognition (ASR) models like Whisper offer solutions, field linguists often lack the expertise to utilise them. We present Easper, an open-source, no-code workflow enabling linguists to iteratively fine-tune ASR models via cloud resources directly from ELAN annotations. Deploying ASR also raises a cold start problem: deciding which recordings to transcribe first to bootstrap an accurate model. Using Easper, we evaluate transcription prioritisation strategies on three Vanuatu languages (Bislama, Nafsan, Nguna). We fine-tune models by recording session, comparing Character Error Rate trajectories when prioritising acoustic cleanliness versus linguistic richness. We demonstrate that prioritising lexically rich narratives and increasing acoustic-phonetic repetition, even in noisy environments, leads to faster improvements in transcription quality.
Chinese Translation
音频转录是语言文献记录中的一个关键瓶颈。尽管像 Whisper 这样的多语言自动语音识别(ASR)模型提供了解决方案,但现场语言学家往往缺乏使用这些模型的专业知识。我们提出了 Easper,这是一个开源的无代码工作流程,使语言学家能够通过云资源直接从 ELAN 注释中迭代微调 ASR 模型。部署 ASR 还会引发一个冷启动问题:决定首先转录哪些录音以启动一个准确的模型。使用 Easper,我们在三种瓦努阿图语言(比斯拉马语、纳夫桑语、古纳语)上评估转录优先策略。我们通过录音会话微调模型,比较在优先考虑声学清晰度与语言丰富性时字符错误率的变化轨迹。我们证明了优先考虑词汇丰富的叙述并增加声学-语音重复,即使在嘈杂环境中,也能更快地提高转录质量。
cs.CL / 22 / 2608.11649

Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study

你会投票给谁?对大型语言模型中的政治对齐进行审计:意大利案例研究
Mungari, Simone
Abstract
As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the political preferences expressed by these systems have become a matter of public interest. Prior research has shown that interactions with LLMs can influence users' political attitudes and choices, raising questions about how these models themselves evaluate political actors. In this paper, we investigate whether and how LLMs express preferences toward political parties and political leaders. We introduce a systematic and reproducible auditing framework in which multiple LLMs are prompted to evaluate parties and leaders across nine criteria. Rather than attempting to infer the models' "true" political beliefs, we focus on their observable behavior, examining consistency across evaluations, differences between models, refusal rates, and sensitivity to prompt formulation. We further investigate how these evaluations vary when models are instructed to adopt different personas. We demonstrate the framework through an Italian case study, providing a systematic analysis of LLM-generated political evaluations on italian parties and leaders.
Chinese Translation
随着用户越来越多地依赖大型语言模型(LLMs)获取有关政治事务的信息和建议,特别是在选举期间,这些系统所表达的政治偏好已成为公众关注的问题。先前的研究表明,与LLMs的互动可能会影响用户的政治态度和选择,这引发了关于这些模型如何评估政治行为者的问题。在本文中,我们探讨LLMs是否以及如何对政党和政治领导人表达偏好。我们引入了一个系统且可重复的审计框架,在该框架中,多个LLMs被提示根据九个标准评估政党和领导人。我们并不试图推断模型的“真实”政治信仰,而是关注它们的可观察行为,考察评估的一致性、模型之间的差异、拒绝率以及对提示表述的敏感性。我们进一步研究当模型被指示采用不同角色时,这些评估如何变化。我们通过一个意大利案例研究展示了该框架,提供了对LLM生成的意大利政党和领导人政治评估的系统分析。
cs.CL / 23 / 2608.11657

Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models

语义Lenia:在大型语言模型的语义空间中自稳孤子的出现
Kayama, Yoshihiko
Abstract
We introduce Semantic Lenia, an artificial life framework that transforms Large Language Model (LLM) inference from a static optimization problem into a continuous dynamical system within the macroscopic logit space. By establishing a non-linear homeostatic feedback loop to dynamically balance semantic attraction and syntactic repulsion, we demonstrate the emergence of "Autonomous Semantic Solitons" -- macroscopic dissipative structures that avoid repetitive crystallization. Our exhaustive parameter sweeps map a critical "Habitable Ridge" where applied steering forces perfectly balance the model's intrinsic syntactic inertia. This approach successfully maintains generative trajectories at the edge of chaos, triggering profound abductive leaps without structural collapse and establishing a physical scaling law for machine cognition.
Chinese Translation
我们介绍了语义Lenia,这是一种人工生命框架,它将大型语言模型(LLM)的推理从静态优化问题转变为宏观logit空间中的连续动力系统。通过建立非线性自稳反馈回路,动态平衡语义吸引力和句法排斥力,我们展示了“自主语义孤子”的出现——避免重复结晶的宏观耗散结构。我们全面的参数扫描描绘了一个关键的“宜居山脊”,在该山脊上施加的引导力完美平衡了模型的内在句法惯性。这种方法成功地保持了生成轨迹在混沌边缘,触发了深刻的溯因飞跃而不发生结构崩溃,并为机器认知建立了物理尺度法则。
cs.CL / 24 / 2608.11660

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

可组合非结构化知识编辑的混合策略自我编辑
Liu, Tianci, Dong, Zihan, Li, Tianchun, Chen, Yi-Chung, Cao, Qiming, Wang, Xingchen, Wang, Shiyang, Miao, Zichen, Zhang, Linjun, Wang, Haoyu, Gao, Jing
Abstract
Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing editors inject such a passage yet fail to use it: the edited model can recall the passage, but can neither answer atomic questions about its facts nor compose them into multi-hop reasoning. We attribute this missing property, which we term composability, to editors' passive reliance on the fixed passage as the sole learning source. In response, we cast editing as a proactive self-distillation from a privileged in-context state of the same model, which requires no external supervision. We further reveal that due to the novelty of the injected knowledge, the pre-edited model's own rollouts rarely cover it, which limits the effectiveness of pure on-policy distillation. To close this gap, we propose HPSE, which builds a hybrid rollout that steps in to place missing facts onto the student's own trajectory precisely where its coverage fails, while staying on-policy elsewhere. We theoretically analyze HPSE's advantage over pure on-policy distillation, and empirically establish its plug-and-play improvements across four LLM backbones and two KE editors under various scenarios.
Chinese Translation
大型语言模型(LLMs)在自然语言任务中表现出色,但它们是在静态语料库上训练的,知识在快速变化的世界中迅速过时。这促使了知识编辑(KE)的研究,旨在在不改变其他无关知识的情况下更新LLM中的特定知识。近期的研究从结构化知识三元组转向非结构化知识编辑(UKE),其中编辑是一个自由形式的段落,可能同时陈述多个事实。然而,现有的编辑器虽然注入了这样的段落,却未能有效利用它:被编辑的模型可以回忆起该段落,但无法回答关于其事实的原子问题,也无法将这些事实组合成多跳推理。我们将这种缺失的属性称为可组合性,并归因于编辑器对固定段落作为唯一学习来源的被动依赖。对此,我们将编辑视为来自同一模型的特权上下文状态的主动自我蒸馏,无需外部监督。我们进一步揭示,由于注入知识的新颖性,预编辑模型自身的回滚很少覆盖这些知识,这限制了纯在线策略蒸馏的有效性。为了解决这一问题,我们提出了HPSE,它构建了一个混合回滚,在学生自身轨迹中精确地填补缺失的事实,同时在其他地方保持在线策略。我们理论上分析了HPSE相对于纯在线策略蒸馏的优势,并在四个LLM基础模型和两个KE编辑器的不同场景下实证建立了其即插即用的改进效果。
cs.CL / 25 / 2608.11694

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

措辞效应:量化大型语言模型基准性能的双向漂移
Thakur, Shailja, An, Sungeun, DeLuca, Chad, Patel, Hima
Abstract
A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it does not. We show that rephrasing a problem while keeping its meaning and answer fixed routinely flips a model's answer in both directions, so some failures become successes and some successes become failures. We call this drift. BenchDrift generates meaning-preserving variations of benchmark problems along four axes, namely linguistic, referential, pragmatic, and structural, and measures how often, and why, correctness flips under each. Across eight models and three benchmarks (GSM8K, MMLU, MATH-Hard), we observe that drift is large in both directions. Two findings stand out. First, phrasing sensitivity does not fade as models get better. Instead, it changes sign. Weak models gain more from rephrasing than they lose, while strong models lose far more than they gain. We find that the best models on a benchmark are therefore the ones whose scores depend most on the wording they happened to be given. Second, the models largely agree on which rephrasings cost the most correct answers even though they differ in how much they drift, so fragility belongs to the rephrasing and not to the model. Furthermore, rephrasing breaks answers a model was confident about, whether the problem is made shorter or longer. Code and Data: https://github.com/IBM/BenchDrift/tree/demo-ui
Chinese Translation
基准分数来自每个问题的单一措辞。这个单一措辞被视为代表了同一问题可以提问的所有方式,但实际上并非如此。我们展示了在保持问题的意义和答案不变的情况下,重新措辞问题通常会在两个方向上翻转模型的答案,因此一些失败变成成功,而一些成功变成失败。我们称之为漂移。BenchDrift 生成沿着四个轴(即语言、指称、语用和结构)保持意义的基准问题变体,并测量在每个轴上正确性翻转的频率及原因。在八个模型和三个基准(GSM8K、MMLU、MATH-Hard)中,我们观察到漂移在两个方向上都很大。有两个发现尤为突出。首先,措辞敏感性并不会随着模型的提升而减弱,反而会改变符号。弱模型从重新措辞中获得的收益大于损失,而强模型的损失远大于收益。因此,我们发现在某个基准上表现最好的模型恰恰是那些得分最依赖于其所接收到的措辞的模型。其次,尽管模型在漂移程度上有所不同,但它们在哪些重新措辞导致最多正确答案损失上基本一致,因此脆弱性属于重新措辞而非模型。此外,无论问题被缩短还是延长,重新措辞都会破坏模型对答案的信心。代码和数据: https://github.com/IBM/BenchDrift/tree/demo-ui
cs.CL / 26 / 2608.11715

When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use

当API使用错误语言时:重新审视多语言工具使用的后训练
Chauhan, Siddharth, Butler, Thomas, Singhania, Abhishek, Porwal, Pankaj, Gupta, Honey
Abstract
The reliability of Large Language Models (LLMs) for API calling degrades in multilingual settings. A common failure occurs when a model selects the correct tool but generates argument values in an inconsistent language, which we term Argument Language Mismatch (ALM). Although semantically correct, such outputs are operationally invalid and not captured by standard API-calling metrics. We revisit post-training strategies for mitigating ALM and find that, in our benchmark, supervised fine-tuning (SFT) provides a strong baseline, substantially improving argument language consistency and end-to-end function call accuracy. Under consistent model selection, SFT achieves performance comparable to, and sometimes exceeding more complex reinforcement learning (RL) approaches. We further examine whether RL with structured, argument-aware rewards offers additional benefits. While methods such as Group Relative Policy Optimization (GRPO) can improve language consistency and better preserve general reasoning ability, these gains are incremental and most pronounced in generalization and multi-objective trade-offs. Overall, our results suggest that much of the performance in multilingual API grounding can be achieved through careful supervised training, with RL providing targeted rather than fundamental improvements.
Chinese Translation
大型语言模型(LLMs)在多语言环境下调用API的可靠性下降。一个常见的失败情况是模型选择了正确的工具,但生成的参数值使用了不一致的语言,我们称之为参数语言不匹配(Argument Language Mismatch,ALM)。尽管语义上正确,这种输出在操作上是无效的,并且未被标准的API调用指标所捕捉。我们重新审视了减轻ALM的后训练策略,发现,在我们的基准测试中,监督微调(Supervised Fine-Tuning,SFT)提供了一个强有力的基线,显著提高了参数语言的一致性和端到端函数调用的准确性。在一致的模型选择下,SFT的性能可与更复杂的强化学习(Reinforcement Learning,RL)方法相媲美,甚至有时超过这些方法。我们进一步考察了使用结构化、关注参数的奖励的RL是否提供额外的好处。尽管像群体相对策略优化(Group Relative Policy Optimization,GRPO)等方法能够改善语言一致性并更好地保持一般推理能力,但这些收益是渐进的,并且在泛化和多目标权衡中最为明显。总体而言,我们的结果表明,多语言API对接中的大部分性能可以通过精心的监督训练来实现,而RL则提供了针对性的而非根本性的改进。
cs.CL / 27 / 2608.11735

Locating and Controlling Implicit Personalization in Large Language Models

定位与控制大型语言模型中的隐性个性化
Yan, Yueru, Wu, Siqi, Le, Thai
Abstract
Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behavioral changes and the model's internal activations remains unclear. Using matched cued and neutral conversations across five LLMs, we establish that a localized internal activation signal tracks changes in recommendations, with correlations up to r=0.87. When multiple cues appear together, their internal signals largely combine, but the changes in output do not simply add up. We further show that removing the internal signal associated with one cue can suppress its influence, often more effectively than asking the model to ignore demographics via prompting, while largely preserving general benchmark performance. However, the ability to selectively remove one dimension's influence while leaving co-present dimensions intact remains highly model- and attribute-specific. These results connect implicit personalization behavior to an internal signal that can be analyzed and causally controlled.
Chinese Translation
大型语言模型(LLMs)通常会根据隐性的人口统计线索调整其输出,即使用户从未明确表述其人口统计身份。先前的研究记录了这种行为,但这些行为变化与模型内部激活之间的联系仍不清晰。通过在五个LLM中使用匹配的提示和中性对话,我们确定了一个局部内部激活信号可以追踪推荐的变化,其相关性高达 r=0.87。当多个线索同时出现时,它们的内部信号在很大程度上会结合,但输出的变化并不简单相加。我们进一步表明,去除与某个线索相关的内部信号可以抑制其影响,通常比通过提示要求模型忽略人口统计信息更有效,同时在很大程度上保持一般基准性能。然而,选择性去除一个维度的影响而不影响共存维度的能力仍然高度依赖于模型和属性的特性。这些结果将隐性个性化行为与一个可以被分析和因果控制的内部信号联系起来。
cs.CL / 28 / 2608.11742

Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

涟漪枢轴搜索:扩散大语言模型的主动并行解码
Ye, Yushi, Chen, Xu, Jiang, Haoyun, Lan, Jinsong, Tang, Haihong, Han, Bo, Tsang, Ivor, Wang, Yanfeng, Zheng, Bo, Yao, Jiangchao
Abstract
Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive language models, offering the potential for substantially faster inference through parallel decoding. Existing parallel decoding schedulers typically commit positions only after they meet a per-position criterion, overlooking how early commitments may benefit subsequent decoding. We identify a ripple effect in dLLM decoding: proactively committing a mid-entropy pivot position can induce a pronounced reduction in uncertainty across the remaining masked positions. This uncertainty reduction allows subsequent steps to unmask more tokens in parallel, thereby accelerating the overall decoding process. To exploit the ripple effect, we propose Ripple-Pivot Search (RPS), a novel training-free decoding method that seeks mid-entropy positions as promising candidate pivots (where to decode), and determines their token assignment that yields the greatest downstream benefit via lookahead evaluation (what to decode). Across 3 dLLMs and 4 reasoning and code-generation benchmarks, RPS achieves 4-10$\times$ wall-clock speedup over the standard decoder while preserving generation quality, and improves accuracy over the previous lookahead baseline by up to 5.49% while delivering higher throughput in most settings. When integrated with KV caching, RPS further achieves up to 18$\times$ wall-clock speedup over the standard decoder.
Chinese Translation
扩散大语言模型(dLLMs)作为自回归语言模型的竞争性替代方案,提供了通过并行解码显著加快推理速度的潜力。现有的并行解码调度器通常仅在满足每个位置的标准后才进行位置承诺,忽视了早期承诺如何有利于后续解码。我们在dLLM解码中识别出一种涟漪效应:主动承诺一个中熵的枢轴位置可以显著降低其余被掩蔽位置的不确定性。这种不确定性降低使得后续步骤能够并行解码更多的标记,从而加速整体解码过程。为了利用涟漪效应,我们提出了涟漪枢轴搜索(Ripple-Pivot Search, RPS),这是一种新颖的无训练解码方法,旨在寻找中熵位置作为有前景的候选枢轴(解码位置),并通过前瞻评估确定其标记分配,以获得最大的下游收益(解码内容)。在3个dLLM和4个推理及代码生成基准测试中,RPS在保持生成质量的同时,实现了相较于标准解码器4-10倍的墙钟速度提升,并在大多数设置中将准确率提高了最高5.49%,同时提高了吞吐量。当与KV缓存结合使用时,RPS进一步实现了相较于标准解码器高达18倍的墙钟速度提升。
cs.CL / 29 / 2608.11753

LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification

LabelFusion-TS:融合大型语言模型、变换器编码器与金融时间序列以进行货币政策立场分类
Schlee, Michael, Lukassen, Fabian, Weisser, Christoph
Abstract
Financial text is produced and interpreted within a market environment, yet financial text classifiers almost always receive text alone. We study whether financial time series are useful as an additional input on the task of classifying sentences from Federal Reserve communication as hawkish, dovish, or neutral. Our system, \lfts{}, extends the \lf{} architecture with this modality: a small voting network combines three independently trained components, a fine-tuned RoBERTa encoder, a prompted large language model (LLM), and a fused ensemble of time-series transformers over the market series of the months preceding publication. Because only about a thousand annotated sentences are available for training, the RoBERTa encoder is first pre-trained on sentences annotated automatically by the LLM and only then fine-tuned on the human labels. Trained on Federal Open Market Committee (FOMC) communication up to 2015 and evaluated on 2015--2022, the fused system achieves 70.2\% weighted F1 -- against 64.1\% for the zero-shot LLM -- and overtakes it with as few as 240 human-labelled sentences. We take this as initial evidence for market time series as an input modality in financial text classification.
Chinese Translation
金融文本是在市场环境中产生和解读的,然而金融文本分类器几乎总是仅接收文本作为输入。我们研究金融时间序列在将美联储沟通的句子分类为鹰派、鸽派或中立时,是否作为额外输入是有用的。我们的系统 extit{LabelFusion-TS} 扩展了 extit{LabelFusion} 架构,增加了这一模态:一个小型投票网络结合了三个独立训练的组件,一个经过微调的 RoBERTa 编码器、一个提示的大型语言模型(LLM),以及一个融合的时间序列变换器集成,基于发布前几个月的市场系列。由于可用于训练的标注句子仅约一千个,RoBERTa 编码器首先在由 LLM 自动标注的句子上进行预训练,然后再在人工标签上进行微调。该系统在2015年之前的联邦公开市场委员会(FOMC)沟通上进行训练,并在2015年至2022年间进行评估,融合系统实现了70.2%的加权F1分数——而零-shot LLM的分数为64.1%——并且在仅使用240个人工标注句子的情况下超越了它。我们将此视为金融文本分类中市场时间序列作为输入模态的初步证据。
cs.CL / 30 / 2608.11758

AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention

AWARe:通过激活加权自适应保留减轻灾难性遗忘
Liao, Juncheng, Lv, Jinfan, Wang, Guoming, Zheng, Jupeng, Xiao, Ling, Tang, Siliang
Abstract
Multimodal Large Language Models (MLLMs) exhibit strong generalization and reasoning abilities due to large-scale multimodal pre-training. However, fine-tuning these models on downstream tasks often leads to catastrophic forgetting, where newly learned task-specific knowledge degrades previously acquired capabilities. This issue arises because gradient updates for new tasks overwrite parameters critical to prior knowledge, limiting the practical deployment of MLLMs. To address this challenge, we propose Activation-Weighted Adaptive REtention (AWARe), a fine-tuning method that mitigates catastrophic forgetting by dynamically controlling parameter updates based on activation patterns. AWARe assigns activation-based importance scores to parameters, selectively freezing those essential for preserving prior capabilities while allowing less important parameters to adapt to new tasks. Importantly, AWARe operates without modifying model architectures, ensuring compatibility with existing inference engines. Extensive experiments demonstrate that AWARe effectively preserves upstream capabilities while achieving superior downstream performance compared to existing methods. Code is available at https://github.com/kaln27/AWARe.
Chinese Translation
多模态大型语言模型(MLLMs)由于大规模的多模态预训练,展现出强大的泛化和推理能力。然而,在下游任务上对这些模型进行微调时,常常会导致灾难性遗忘,即新学习的任务特定知识会削弱先前获得的能力。这个问题的产生是因为新任务的梯度更新覆盖了对先前知识至关重要的参数,从而限制了MLLMs的实际应用。为了解决这一挑战,我们提出了激活加权自适应保留(AWARe),这是一种通过动态控制基于激活模式的参数更新来减轻灾难性遗忘的微调方法。AWARe为参数分配基于激活的权重重要性分数,选择性地冻结那些对保留先前能力至关重要的参数,同时允许不太重要的参数适应新任务。重要的是,AWARe在不修改模型架构的情况下运行,确保与现有推理引擎的兼容性。大量实验表明,AWARe在有效保留上游能力的同时,相较于现有方法在下游性能上取得了优越的表现。代码可在 https://github.com/kaln27/AWARe 获取。
cs.CL / 31 / 2608.11767

Causal Structure is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library

因果结构可诱导但功能上解耦:一种类型机制库的路由/读出边界
Xun, Xining
Abstract
When a language model answers an interventional question, the computation it must perform depends on the type of evidence the query requires. We report a decoupling in how a transformer organizes causal knowledge: slot-by-type structure induced by type-level supervision organizes routing, yet remains functionally decoupled from answer readout. We establish this with a typed mechanism library -- discrete mechanism slots partitioned by evidence type, auditable at the state level -- on a causal-world benchmark with exact interventional ground truth, under a frozen protocol, at two scales (22.6M and 125M). Four preregistered findings. (i) Origin. Slot-by-type organization is induced by type-level supervision: absent in architecturally identical unsupervised controls, not buyable by content-free gating labels, and statistically attributable to the supervision signal, replicating at 125M under a powered preregistered protocol (all nine cells passed). (ii) Boundary. The induced structure is a typed routing index with a sharp routing/readout boundary: slot codes scaffold routing but do not drive answer readout ($|\Delta\hat{y}| \le 3.4\times10^{-6}$, zero collateral, three seeds, stable across a 5.6x scale window) -- we therefore make no behavioral-editability claim. (iii) Cost. The structure is free: LM quality matches a parameter-matched monolith within 0.0082 nats. (iv) Trust. The library state is exactly local under edit and bit-exactly revertible -- 250 single-edit and 1,000 stacked reverts per seed, zero failures. We further find that the unsupervised null itself moves with scale, so comparisons reusing a null calibrated at one scale may be confounded at another. Every claim is tied to a preregistered, machine-checkable criterion archived before the data it governs; the full audit trail, including one criterion we failed and how the frozen protocol handled it, is released as an appendix.
Chinese Translation
当语言模型回答一个干预性问题时,它必须执行的计算取决于查询所需的证据类型。我们报告了变换器组织因果知识的解耦现象:由类型级监督诱导的按类型划分的槽结构组织了路由,但在功能上与答案读出解耦。我们通过一个类型机制库来建立这一点——按证据类型划分的离散机制槽,在状态级别可审计——在一个具有确切干预真实值的因果世界基准上,在冻结协议下,使用两个规模(22.6M 和 125M)。四个预注册发现。(i) 起源。按类型组织的槽结构是由类型级监督诱导的:在架构上完全相同的无监督对照中缺失,无法通过无内容的门控标签获得,且统计上可归因于监督信号,在一个有力的预注册协议下在125M上复制(所有九个单元均通过)。(ii) 边界。诱导的结构是一个带有明确路由/读出边界的类型路由索引:槽代码支撑路由,但不驱动答案读出($| ext{Δ} ilde{y}| ext{≤} 3.4 imes10^{-6}$,零附带,三个种子,在5.6倍规模窗口内稳定)——因此我们不做行为可编辑性的声明。(iii) 成本。该结构是免费的:语言模型的质量与参数匹配的单体模型相差不超过0.0082 nats。(iv) 信任。库状态在编辑下是完全局部的,并且可以逐位精确恢复——每个种子250次单次编辑和1000次堆叠恢复,零失败。我们进一步发现,无监督的零本身随着规模而变化,因此在一个规模上校准的零进行比较时,可能在另一个规模上受到混淆。每个声明都与一个预注册的、机器可检查的标准相关联,该标准在其所管理的数据之前归档;完整的审计记录,包括我们未能通过的一个标准以及冻结协议如何处理它,作为附录发布。
cs.CL / 32 / 2608.11772

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction

恢复前的诊断:将代理失败转化为选择性自我修正
Wang, Pan, Hu, Yihao, Wang, Hang, Lv, Zirui, Zhang, Xin, Li, Jianshe, Yang, Jiang-Ming, Wu, Wei, Tong, Yongqi
Abstract
Self-correction is particularly useful when a failure constrains the next repair. Coding agents benefit from this property because compilers, tests, and execution traces turn many failures into typed recovery signals, but broad language-agent tasks often expose only a coarse task failure. This creates a tension for generic recovery playbooks: they broaden the agent's context precisely when the system needs a narrower repair interface, mixing incompatible signals for invalid actions, missing procedures, and strict-format errors. Our insight is that development-set failures can recover part of the missing diagnostic substrate by deciding which recovery interventions are admissible before test-time correction. We propose DARC, a diagnosis-guided recovery harness that profiles task-family failure modes, prunes mismatched interventions from a shared recovery library, and freezes a verifier-selected success-cost policy for deployment. This causal order makes correction selective: the harness first determines what kind of failure can be repaired, then decides how much recovery evidence to spend. In ALFWorld, AppWorld, and XBRL Finance, the same protocol yields an action-validity harness, a procedural-recovery fallback, and a format-precision retrieval policy; in each evaluated setting it improves average task performance over base agents and broad playbooks while reducing environment steps or retrieval budget. Our experiments show that failures need not trigger uniformly more context: DARC turns self-correction from prompt expansion into recovery-interface design. DARC provides a practical route toward more reliable agents in domains where compiler-like feedback is absent: making failures actionable before making contexts larger.
Chinese Translation
自我修正在故障限制下的下一次修复时尤为有用。编码代理从这一特性中受益,因为编译器、测试和执行跟踪将许多故障转化为类型化的恢复信号,但广泛的语言代理任务通常仅暴露出粗糙的任务失败。这为通用恢复手册带来了紧张关系:它们在系统需要更窄的修复接口时扩大了代理的上下文,混合了无效操作、缺失程序和严格格式错误的不兼容信号。我们的见解是,开发集中的失败可以通过在测试时间修正之前决定哪些恢复干预是可接受的,从而恢复部分缺失的诊断基础。我们提出了 DARC(诊断引导恢复工具),它对任务家族的失败模式进行分析,从共享恢复库中剔除不匹配的干预,并为部署冻结一个验证者选择的成功成本策略。这一因果顺序使得修正变得选择性:该工具首先确定可以修复的故障类型,然后决定投入多少恢复证据。在 ALFWorld、AppWorld 和 XBRL Finance 中,相同的协议产生了一个操作有效性工具、一个程序恢复后备和一个格式精度检索策略;在每个评估环境中,它在提高平均任务性能的同时,减少了环境步骤或检索预算,相较于基础代理和广泛手册。我们的实验表明,故障不必均匀触发更多的上下文:DARC 将自我修正从提示扩展转变为恢复接口设计。DARC 为在缺乏编译器类反馈的领域中实现更可靠的代理提供了一条实用的途径:在扩大上下文之前使故障可操作。
cs.CL / 33 / 2608.11786

Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English Languages

语言条件去量化:恢复量化对非英语语言的损失
Thomas, Nirmal
Abstract
Aggressive quantization disproportionately harms multilingual capability: in the sub-4B INT3 GPTQ regime, we measure 2-4x larger perplexity degradation on non-English languages than on English. We propose Language-Conditional Dequantization (LCD), a post-hoc method that attaches per-language rank-2 LoRA corrections to the linear layers of an already-quantized model, adding 0.12% parameters per language and training in under 20 minutes on a single GPU. Across Qwen2.5-3B and Llama-3.2-3B, LCD recovers 70-83% of the perplexity gap for non-Latin script languages and 17-28% of the GlobalMMLU accuracy gap, outperforming a language-agnostic correction of equal capacity by 3-9 points on typologically distant languages and a data-free low-rank baseline (LQER) by an order of magnitude. We further identify a perplexity-accuracy disconnect and trace it to where quantization concentrates damage: early-depth errors (Llama) propagate downstream and resist local correction, while late-depth errors (Qwen) do not. A layer-restricted variant of LCD validates this mechanism directly.
Chinese Translation
激进的量化对多语言能力造成了不成比例的损害:在低于4B的INT3 GPTQ范围内,我们测量到非英语语言的困惑度降级比英语高出2-4倍。我们提出了语言条件去量化(Language-Conditional Dequantization, LCD),这是一种后处理方法,它将每种语言的秩-2 LoRA修正附加到已量化模型的线性层上,为每种语言增加0.12%的参数,并在单个GPU上训练不超过20分钟。在Qwen2.5-3B和Llama-3.2-3B上,LCD恢复了非拉丁文字语言70-83%的困惑度差距和17-28%的GlobalMMLU准确率差距,超越了同等容量的语言无关修正,在类型上相距较远的语言上提高了3-9个点,并在无数据的低秩基线(LQER)上提高了一个数量级。我们进一步识别出困惑度与准确率之间的脱节,并追踪到量化集中损害的地方:早期深度错误(Llama)向下传播并抵抗局部修正,而晚期深度错误(Qwen)则没有。LCD的一个层限制变体直接验证了这一机制。
cs.CL / 34 / 2608.11787

GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

基于GRPO的金融建议生成:在CATE评估下超越商业大型语言模型
Shoham, Ofir Ben, Harsola, Shrutendra, Subrahmaniam, Vignesh, Mohan, Shravan, Gazman, Yakov, Vainas, Oded
Abstract
Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business. Direct supervision is difficult: historical decisions are not necessarily optimal, and high-quality free-form labels are expensive to obtain. We formulate financial advice generation as a reinforcement learning problem and fine-tune an open-weight language model using Group Relative Policy Optimization (GRPO). Our reward is an LLM-as-a-judge rubric that scores each recommendation across multiple binary dimensions of advice quality, augmented with a safety gate for harm prevention. Since LLM-based evaluation alone cannot confirm whether improvements reflect genuine business value rather than adaptation to the judge, we complement it with a judge-independent audit based on a standard doubly-robust Conditional Average Treatment Effect (CATE) estimator. Under this observational off-policy audit, our trained LLM achieves approximately twice the estimated gross-profit lift of the strongest evaluated commercial baseline ($0.0228$ vs.\ $0.0104$), together with the lowest downside rate and the least negative tail risk of any policy evaluated. Notably, the two evaluations do not rank the baselines identically: the untrained base model places last on the judge rubric but second on the causal audit, indicating that the audit captures a signal the judge does not. Our results demonstrate that GRPO with a finance-grounded reward signal can produce substantially more useful business recommendations than commercial LLMs, and that a judge-independent causal audit is a valuable complement to, rather than a confirmation of, LLM-as-a-judge assessment in financial NLP.
Chinese Translation
从商业记录中生成可行的金融建议要求模型整合数值推理、领域知识和合理判断,同时避免可能对业务造成伤害的建议。直接监督是困难的:历史决策不一定是最优的,高质量的自由形式标签获取成本高昂。我们将金融建议生成形式化为一个强化学习问题,并使用群体相对策略优化(Group Relative Policy Optimization, GRPO)对开放权重语言模型进行微调。我们的奖励机制是一个大型语言模型作为评判者的评分标准,它在多个二元维度上对每个建议的质量进行评分,并增加了一个安全门以防止伤害。由于仅依靠大型语言模型的评估无法确认改进是否反映了真正的商业价值而非对评判者的适应,我们还结合了基于标准双重稳健条件平均处理效应(Conditional Average Treatment Effect, CATE)估计器的独立审计。在这种观察性离线政策审计下,我们训练的语言模型实现了大约是最强评估商业基线的两倍的估计毛利提升($0.0228$ 对比 $0.0104$),同时具有最低的下行风险率和最小的负尾风险。值得注意的是,这两种评估对基线的排名并不相同:未经训练的基础模型在评判者评分标准中排名最后,但在因果审计中排名第二,表明审计捕捉到了评判者未能识别的信号。我们的结果表明,结合以金融为基础的奖励信号的GRPO能够生成比商业大型语言模型更有用的商业建议,而独立的因果审计是对大型语言模型作为评判者评估的有价值的补充,而非确认。
cs.CL / 35 / 2608.11788

TELLME: Test-Enhanced Learning for Language Model Enrichment

TELLME:用于语言模型丰富的测试增强学习
Kim, Minjun, Won, Inho, Lim, Hyeonseok, Kim, MinKyu, Yuk, Junghun, Go, Wooyoung, Park, Jongyoul, Park, Jungyeul, Lim, KyungTae
Abstract
Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such as the difficulty of acquiring large-scale domain-specific datasets and high computational costs. In this study, we propose a novel method called Test-Enhanced Learning for Language Model Enrichment (TELLME) to alleviate these issues. TELLME leverages the TestEnhanced Learning (TEL) principle, whereby the model's training efficiency is improved using quizzes during training. It integrates this principle with CPT, thereby promoting efficient domain-specific knowledge acquisition and long-term memory retention. Experimental results demonstrate that TELLME outperforms existing methods by up to 23.6% in the financial domain and achieves a 9.8% improvement in long-term memory retention.
Chinese Translation
持续预训练(CPT)已被广泛采用作为大型语言模型领域适应的方法。然而,CPT 一直伴随着一些挑战,例如获取大规模领域特定数据集的困难和高计算成本。在本研究中,我们提出了一种新颖的方法,称为用于语言模型丰富的测试增强学习(TELLME),以缓解这些问题。TELLME 利用测试增强学习(TEL)原则,通过在训练过程中使用测验来提高模型的训练效率。它将这一原则与CPT相结合,从而促进高效的领域特定知识获取和长期记忆保留。实验结果表明,TELLME 在金融领域的表现比现有方法提高了多达 23.6%,并在长期记忆保留方面实现了 9.8% 的提升。
cs.CL / 36 / 2608.11805

Hybrid Gated Attention

混合门控注意力
Zhou, Zekun, Xie, Ruobing, Wang, Lanrui, Sun, Weixuan
Abstract
Gated attention is an effective approach to mitigate attention sinks and enhance the representational capacity of attention. To further extend its effectiveness-efficiency Pareto frontier, we propose a Hybrid Gated Attention (HyGA) framework that contains three types of gating strategies. Specifically, these gates leverage diverse information from multiple stages of attention, and collaboratively build element-wise/head-wise gating from multiple perspectives, capturing intra-head and cross-head information interactions. Through our hybrid gating components, HyGA could provide multi-source modulation signals, enabling more comprehensive control over information flow and improving the representational capacity of attention. We also introduce low-rank matrix decomposition and learnable attention sink to further enhance training efficiency and stability. In experiments, we evaluate HyGA on widely-used benchmarks based on different backbones. The experimental results show that our HyGA comprehensively improves both training loss and various downstream performances compared with Gated attention. HyGA has also been verified to achieve the best performance at different computation costs, with comprehensive model analyses for better understanding. The proposed HyGA sheds light on a more effective, efficient, and stable attention mechanism.
Chinese Translation
门控注意力是一种有效的方法,用于减轻注意力消耗并增强注意力的表征能力。为了进一步扩展其有效性-效率的帕累托前沿,我们提出了一种混合门控注意力(Hybrid Gated Attention, HyGA)框架,该框架包含三种类型的门控策略。具体而言,这些门控策略利用来自多个注意力阶段的多样信息,并从多个角度协同构建逐元素/逐头门控,捕捉头内和头间的信息交互。通过我们的混合门控组件,HyGA能够提供多源调制信号,从而实现对信息流的更全面控制,并提高注意力的表征能力。我们还引入了低秩矩阵分解和可学习的注意力消耗,以进一步提高训练效率和稳定性。在实验中,我们在基于不同主干网络的广泛使用基准上评估了HyGA。实验结果表明,与门控注意力相比,我们的HyGA在训练损失和各种下游性能上都有全面的改善。HyGA还被验证在不同计算成本下实现最佳性能,并进行了全面的模型分析以便更好地理解。所提出的HyGA为更有效、高效和稳定的注意力机制提供了启示。
cs.CL / 37 / 2608.11822

Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release

定位但不可释放:静默门控反转与有界线性释放
Xun, Xining
Abstract
A growing body of work reports that language models represent task-relevant latent structure that they fail to use. Whether such structure, once located, can be converted into behavior is a separate question that is rarely tested end to end. We submit the complete pipeline -- detect, localize, and release -- to a fully preregistered stress test on a 25.7M transformer trained on causal-evidence discrimination, where a known suppression phenomenon (latent causal structure present but behaviorally unused) has previously been documented. Every threshold, claim template, and decision-tree branch was hashed and archived before any corresponding data existed. Three findings. (i) Localization succeeds: interventions at observation-evidence channels of mid layers restore target behavior on otherwise-suppressed worlds (paired release advantages $0.563$ and $0.854$, 97.5% CIs excluding zero; best-site release rate $0.889$). (ii) Gating fails out of distribution: a detector calibrated to trigger on zero out-of-distribution calibration worlds triggers on 6.9-7.3% of held-out in-distribution generations and on zero of the 2,400 held-out generations that actually need it -- a complete inversion that silently reduces the gated pipeline to its base model. (iii) Linear release is capped: removing the gate and injecting a per-instance linear direction unconditionally yields a monotone dose-response that plateaus far below the preregistered release margin (intercept $0.382 \to 0.311 \to 0.264$ vs. threshold $\le 0.08$); per-instance adaptivity adds less than $\pm 0.03$. The failure is doubly located: the detector is OOD-inverted, and the entire family of linear release directions at this site and resolution is bounded away from sufficiency. The two failures are dissociable, and neither overturns localization. Every number traces to a hashed artifact in the released audit chain.
Chinese Translation
越来越多的研究报告表明,语言模型表示了与任务相关的潜在结构,但它们未能有效利用这些结构。一旦定位,这种结构是否能够转化为行为是一个很少经过全面测试的独立问题。我们将完整的流程——检测、定位和释放——提交给一个完全预注册的压力测试,测试对象为一个在因果证据区分任务上训练的2570万参数的变换器,其中已记录了一个已知的抑制现象(潜在因果结构存在但行为上未被使用)。在任何相应数据存在之前,每个阈值、声明模板和决策树分支都被哈希并存档。三个发现:(i)定位成功:在中间层的观察证据通道进行干预能够恢复在其他被抑制世界中的目标行为(配对释放优势为 $0.563$ 和 $0.854$,97.5% 置信区间不包括零;最佳站点释放率为 $0.889$)。 (ii)门控在分布外失败:一个经过校准以在零分布外校准世界中触发的检测器在6.9%-7.3%的保留分布内生成中触发,而在2400个实际上需要的保留生成中则未触发——这是一种完全的反转,默默地将门控流程简化为其基础模型。 (iii)线性释放受到限制:去除门控并注入每个实例的线性方向无条件地产生了一个单调的剂量-反应关系,该关系远低于预注册的释放边际(截距 $0.382 o 0.311 o 0.264$ 对比阈值 $ ext{≤} 0.08$);每个实例的适应性增加不足 $ ext{±} 0.03$。这种失败是双重定位的:检测器在分布外被反转,而在该站点和分辨率下的所有线性释放方向都远离充分性。两个失败是可分离的,且都不推翻定位。每个数字都追溯到释放审计链中的哈希工件。
cs.CL / 38 / 2608.11843

When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation

当知识库成为黄金标准:测量实体级机器翻译中的资源共享评估循环
Bae, Jinhyung, Kil, Dain, Oh, Seongmin, Lee, Seungmin
Abstract
The Seungjeongwon Ilgi, a UNESCO Memory of the World record, is only 37.4% translated, and the most conspicuous failure mode in automatic translation is the person name -- a misread name corrupts the historical fact rather than merely the surface. Low-resource historical domains have no expert gold standard for entity translation, so practitioners substitute a knowledge base (KB) for the gold. That KB is the same resource injected into the system: scoring becomes self-referential and the metric measures instruction compliance rather than translation quality. We measure this loop. Using expert person-name annotations from the National Institute of Korean History as a gold independent of the injection pipeline, we hold the entity set fixed and vary only the provenance of the correct reading. Of 527 expert-annotated mentions, only 31.1% lie outside the injection pipeline, and the residual loop is not uniform -- in the overlapping segment the injected reading agrees with the human translation 97.8% of the time against 70.1% in the independent one, so the segment that looks healthiest is the one the loop is holding up. Across four models, a difference-in-differences analysis shows the gain from KB injection is confined to the segment whose gold shares the injected resource; in the independent segment it is at or below zero. Post-injection preservation clusters in a narrow 0.910-0.996 band even though baseline capability differs fivefold, so the reported gain is the complement of prior performance and weaker models appear to improve more dramatically. On an independent sample built by removing the construction filter, the measure replicates within model (overlapping intervals) while discriminating between models (non-overlapping intervals) -- it reflects a property of the model, not of the sample.
Chinese Translation
《承政院日记》作为联合国教科文组织世界记忆遗产,仅翻译了37.4%,而自动翻译中最明显的失败模式是人名——错误的名字会破坏历史事实,而不仅仅是表面。低资源历史领域没有专家黄金标准用于实体翻译,因此从业者用知识库(KB)替代黄金标准。这个知识库是注入系统的相同资源:评分变得自我指涉,度量衡量的是指令合规性而非翻译质量。我们测量这个循环。使用来自韩国历史研究院的专家人名注释作为与注入管道无关的黄金标准,我们固定实体集,仅改变正确读音的来源。在527个专家注释的提及中,只有31.1%位于注入管道之外,残余循环并不均匀——在重叠部分,注入的读音与人工翻译一致的比例为97.8%,而独立部分为70.1%,因此看似最健康的部分正是循环所维持的部分。在四个模型中,差异中的差异分析显示,知识库注入的收益仅限于黄金标准与注入资源共享的部分;在独立部分,收益为零或以下。注入后的保留聚集在狭窄的0.910-0.996区间,尽管基线能力差异达到五倍,因此报告的收益是先前表现的补充,而较弱的模型似乎改善得更为显著。在通过去除构建过滤器构建的独立样本中,该度量在模型内复制(重叠区间),同时在模型之间区分(非重叠区间)——它反映的是模型的特性,而非样本的特性。
cs.CL / 39 / 2608.11879

Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems

全面回忆的代价是什么?代理记忆系统服务成本的基准测试
Pollertlam, Natchanon, Kornsuwannawit, Witchayut
Abstract
Long-running conversational agents increasingly rely on a memory system to avoid resending the whole conversation each turn, yet how much that costs to serve has received little systematic benchmarking. We compare three memory systems (Mem0, Hindsight, and Mastra Observational Memory) against two reference strategies -- a fixed-size rolling window and resubmitting the full transcript -- across two backbones and conversations of up to 400 turns, pairing every cost measurement with answer accuracy on 665 LoCoMo questions. First, a memory system's serving cost cannot be predicted from conversation length and message size alone: a regression that tracks the two reference strategies closely misses the memory systems by 18-69%, their cost driven instead by internal memory behavior. Second, a break-even analysis shows that whether -- and when -- a memory system becomes cheaper to serve than the full transcript is highly sensitive to the system and the backbone, from the first tens of turns for the cheapest to never within 400 turns for the most expensive. Third, no system wins on both axes: accuracy spans 21-54%, and the backbone choice drives cost as much as the memory system does.
Chinese Translation
长期运行的对话代理越来越依赖于记忆系统,以避免在每轮中重新发送整个对话,但其服务成本却鲜有系统性的基准测试。我们将三种记忆系统(Mem0、Hindsight 和 Mastra Observational Memory)与两种参考策略——固定大小的滚动窗口和重新提交完整转录——进行比较,涵盖两个骨干网络和最多 400 轮的对话,并将每项成本测量与 665 个 LoCoMo 问题的回答准确性配对。首先,记忆系统的服务成本不能仅通过对话长度和消息大小来预测:一个紧密跟踪两种参考策略的回归模型错过了记忆系统 18-69%,其成本主要受内部记忆行为的驱动。其次,盈亏平衡分析表明,记忆系统何时变得比完整转录更便宜的服务成本高度依赖于系统和骨干网络,对于最便宜的系统在前几十轮中就能实现,而对于最昂贵的系统在 400 轮内则从未实现。第三,没有任何系统在两个维度上都胜出:准确性范围为 21-54%,而骨干网络的选择对成本的影响与记忆系统相当。
cs.CL / 40 / 2608.11919

LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training

LazyTrain:面向零浪费产量优化的有限资源分配在大语言模型训练中的应用
Wu, Xiaojun, Yang, Cehao, Liu, Honghao, Lin, Xueyuan, Jiang, Xuhui, Xu, Chengjin, Li, Jia, Guo, Jian
Abstract
Training large language models on limited hardware is increasingly a scheduling problem across GPU compute, host memory, PCIe transfer, and storage bandwidth. Existing offloading systems reduce GPU residency, and MegaTrain shows that a CPU-master layer-streaming executor can train large models on a single GPU, but fixed checkpointing and placement heuristics still leave communication exposed on the critical path. We propose LazyTrain, an optimization layer over a layer-streaming executor. LazyTrain formulates checkpoint selection, activation placement, recomputation, and CPU-GPU-NVMe communication overlap as a mixed-integer scheduling problem, then executes the solved policy during training. It further couples 8-bit optimizer states with fast gradient clipping as a single Hybrid 8-bit operator: state compression reduces optimizer-state memory, while fast clipping counteracts the additional CPU-side update overhead. Across H800 experiments from Qwen2.5-3B to Qwen3.6-27B, LazyTrain improves sustained TFLOPS over matched baselines runs by approximately 1.24$\times$; RTX 3090 experiments likewise increase the maximum feasible batch size by one at each model scale. In the primary Qwen3.6-27B H800 MetaMathQA run, LazyTrain reaches 219.95 TFLOPS and 1361 tokens/s at batch size 72, peaks at 68.84\,GB of GPU memory, and obtains 95.42\% exact-match accuracy on the full evaluation split. The source code is available at https://github.com/DataArcTech/LazyTrain.
Chinese Translation
在有限硬件上训练大语言模型日益成为一个涉及GPU计算、主机内存、PCIe传输和存储带宽的调度问题。现有的卸载系统减少了GPU的驻留时间,而MegaTrain展示了一个CPU主控层流式执行器可以在单个GPU上训练大模型,但固定的检查点和放置启发式仍然使得通信暴露在关键路径上。我们提出LazyTrain,一个基于层流式执行器的优化层。LazyTrain将检查点选择、激活放置、重计算以及CPU-GPU-NVMe通信重叠形式化为一个混合整数调度问题,然后在训练过程中执行解决的策略。它进一步将8位优化器状态与快速梯度裁剪结合为一个单一的混合8位操作符:状态压缩减少了优化器状态的内存需求,而快速裁剪抵消了额外的CPU端更新开销。在Qwen2.5-3B到Qwen3.6-27B的H800实验中,LazyTrain在匹配基线运行中将持续TFLOPS提高了约1.24倍;RTX 3090实验同样在每个模型规模上将最大可行批量大小增加了一个。在主要的Qwen3.6-27B H800 MetaMathQA运行中,LazyTrain达到了219.95 TFLOPS和1361 tokens/s,批量大小为72,峰值GPU内存为68.84 GB,并在完整评估分割上获得了95.42%的精确匹配准确率。源代码可在https://github.com/DataArcTech/LazyTrain获取。
cs.CL / 41 / 2608.11922

LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence

LODESTAR:可信的熵是被导航的,而不仅仅是被测量的——增强极化器使得冻结的语言模型不被错误证据自信地误导
Ko, Po-Jen, Wu, Che-Cheng, Hsu, Hung-Chun, Chang, Li-Yang, Wang, Chuan-Ju
Abstract
Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.4769 to 0.5148 over the retriever's top-ranked passage, with no gold answers. Yet this lowest-entropy rule, which prior entropy-based selectors adopt, fails in a specific and consequential way: a misleading passage makes the respondent confidently wrong, driving its entropy down precisely where the signal looks most trustworthy. We show that the failure comes from the passage the respondent reads -- and the context that passage is read in is an input we can intervene on. We introduce LODESTAR, to our knowledge the first method to score a text intervention by the uncertainty it induces in a third-party frozen respondent, compared across one question's candidates. LODESTAR uses reinforcement learning to train, once and offline, a polarizer -- a short fixed natural-language string inserted into the respondent's prompt and never into its weights; its training labels are built offline from gold answers and two LLM judges, and inference reads neither. Evaluating every competing selector under the same frozen respondent and the same candidate pools on 5,008 questions, LODESTAR attains the highest mean $F_1$ of any inference-ready selector (0.5148 to 0.5339), the highest exact match (0.4136), and the highest GPT-4o judge score of the frozen-respondent configurations judged (0.6435); its three-seed mean wins all 70 method-by-dataset $F_1$ cells against fourteen published configurations while remaining paired-significant against every one. The gain holds both in-domain and out-of-domain, and ablating the polarizer shows it is what makes the respondent read a misleading passage less often (26.0% against 30.3%).
Chinese Translation
预测分布熵在检索增强的问答中形成了强有力的选择规则:在五个问答基准测试中,保持冻结的响应者语言模型所生成的最低答案标记熵的候选答案,使得在检索器的最高排名段落中,平均答案 $F_1$ 从 0.4769 提升至 0.5148,且没有黄金答案。然而,这一最低熵规则在特定且重要的方面失败了:误导性的段落使得响应者自信地错误,从而在信号看起来最可信的地方降低了其熵。我们表明,这一失败源于响应者所阅读的段落——而该段落的阅读上下文是我们可以介入的输入。我们引入了 LODESTAR,作为我们所知的第一种通过其对第三方冻结响应者所引发的不确定性来评分文本干预的方法,比较一个问题的候选答案。LODESTAR 使用强化学习进行训练,一次性离线训练一个极化器——一个短的固定自然语言字符串,插入到响应者的提示中,而不进入其权重;其训练标签是基于黄金答案和两个语言模型评审者离线构建的,推理时不读取这两者。在 5,008 个问题上评估每个竞争选择器,使用相同的冻结响应者和相同的候选池,LODESTAR 达到了任何推理准备选择器中最高的平均 $F_1$(从 0.5148 提升至 0.5339),最高的精确匹配(0.4136),以及冻结响应者配置中被评审的最高 GPT-4o 评审分数(0.6435);其三种种子的平均在与十四种已发布配置的 70 个方法-数据集 $F_1$ 单元中获胜,同时在每个单元中保持配对显著性。这一增益在领域内和领域外均有效,且去除极化器显示它是使响应者较少阅读误导性段落的原因(26.0% 对比 30.3%)。
cs.CL / 42 / 2608.11924

Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

从火花到论文:可组合技能的端到端研究论文生成
Qian, Zhuoyang, Wu, Biao, Wang, Yiran, Yan, Chris D, Dai, Desan, Zheng, Liangwei, Jiang, Jin, Zhang, Junsheng, Wang, Wenhao
Abstract
Turning a research idea into a complete paper requires more than text generation: the system must retrieve literature, design and execute experiments, revise claims according to evidence, produce publication-ready figures, and maintain consistency across a long generation process. We present Spark-to-Paper, an end-to-end research paper generation system implemented as thirteen composable skills inside an existing coding assistant, without requiring a separate agent platform or orchestration service. Spark-to-Paper separates model-based judgment from deterministic operations that can be directly executed and checked. It further separates experiment planning from reporting, so that required evidence is specified before results are observed and manuscript claims are revised according to measured outcomes. To improve reliability over long research trajectories, the system combines deterministic integrity checks with self-critique and bounds a failure mode we call the Self-Refutation Loop, in which repeated experiments continue to reject the original research objective. Spark-to-Paper also produces editable vector figures through programmatic plotting for experimental results and code-based reconstruction for generated method diagrams. Across eight controlled research topics, Spark-to-Paper achieves 99.5% citation validity and 96.4% figure editability. A controlled ablation increases fabrication detection from 14% for a single-pass draft to 92% with the full integrity and review stack, while adversarial review achieves 74% precision. The full system uses 11.9M tokens, costs $8.1 per manuscript, and requires 3.2 hours on average. These results show that end-to-end research paper generation can be implemented as a lightweight, composable workflow inside existing coding assistants while keeping experimental evidence central to how claims are accepted, revised, or abandoned.
Chinese Translation
将研究想法转化为完整论文不仅仅需要文本生成:系统必须检索文献、设计和执行实验、根据证据修订主张、生成适合发表的图表,并在漫长的生成过程中保持一致性。我们提出了Spark-to-Paper,这是一个端到端的研究论文生成系统,作为现有编码助手中的十三个可组合技能实现,无需单独的代理平台或协调服务。Spark-to-Paper将基于模型的判断与可直接执行和检查的确定性操作分开。它进一步将实验规划与报告分开,以便在观察结果之前指定所需的证据,并根据测量结果修订手稿主张。为了提高在长期研究轨迹上的可靠性,该系统结合了确定性完整性检查与自我批评,并限制了一种我们称之为自我反驳循环的失败模式,在该模式中,重复实验持续拒绝原始研究目标。Spark-to-Paper还通过程序化绘图生成可编辑的矢量图,以展示实验结果,并通过基于代码的重建生成方法图。在八个受控研究主题中,Spark-to-Paper实现了99.5%的引用有效性和96.4%的图表可编辑性。一次受控消融实验将单次草稿的伪造检测率从14%提高到92%,而对抗性审查实现了74%的精确度。整个系统使用了11.9M个标记,平均每篇手稿成本为8.1美元,平均需要3.2小时。这些结果表明,端到端的研究论文生成可以作为现有编码助手中的轻量级可组合工作流程实现,同时将实验证据作为接受、修订或放弃主张的核心。
cs.CL / 43 / 2608.11947

Accuracy and Order Sensitivity Diverge Under Label-Free Strategies

无标签策略下的准确性与顺序敏感性分歧
Hanna, Karl, Feng, Chen
Abstract
Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge. In this paper, we test whether preventing a model from seeing option labels while committing to an answer removes positional influence and, in turn, improves performance. We evaluate two different strategies for mitigating bias. The first uses a generation-then-matching approach, and the second scores options in isolation, which is positionally unbiased by construction. Neither reliably improves accuracy. A complete decomposition shows that the bottleneck is withholding options, not the matching step. The only configuration that consistently matches the baseline is the one that shows the model all options paired with an LLM matcher. However, eliminating positional influence entirely still does not reliably yield accuracy gains, while cyclic permutation often improves them. For two-stage prompting, an aggregate measure of recall imbalance and a direct per-question measure of order sensitivity both fail to show reliable debiasing.
Chinese Translation
多项选择基准广泛用于评估大型语言模型,但多项选择题(MCQ)得分将知识与选项顺序的敏感性混合在一起,这使得它们成为不可靠的模型知识测量工具。本文测试了在模型作答时防止其看到选项标签是否能消除位置影响,从而提高性能。我们评估了两种不同的减轻偏差的策略。第一种采用生成后匹配的方法,第二种则在孤立状态下对选项进行评分,这种方法在结构上是无位置偏见的。两者都未能可靠地提高准确性。完整的分解显示瓶颈在于抑制选项,而不是匹配步骤。唯一能持续与基线匹配的配置是将所有选项与大型语言模型(LLM)匹配器配对展示给模型。然而,完全消除位置影响仍然未能可靠地带来准确性提升,而循环排列往往能改善准确性。在两阶段提示中,召回不平衡的聚合度量和每个问题的顺序敏感性直接度量均未能显示出可靠的去偏见效果。
cs.CL / 44 / 2608.11981

Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed

小语言模型的可信度基准测试:预训练模型与压缩模型的比较
Lin, Haokun, Zhu, Kaijie, Xu, Haobo, Wu, Yichen, Lu, Zhichao, Zhang, Qingfu, Sun, Zhenan
Abstract
Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in resource-constrained scenarios. Existing approaches to building SLMs typically follow two paths: training compact models from scratch, or compressing larger pre-trained models using methods such as pruning, quantization, or distillation. As language models become increasingly integrated into real-world applications, ensuring their trustworthiness has become a critical concern. However, how to build trustworthy SLMs remains an underexplored question. In this work, we present a comprehensive evaluation of SLM trustworthiness across multiple dimensions, including fairness, robustness, privacy, and ethics. We first examine the effects of pruning and quantization, and find that quantization is significantly more effective in preserving trustworthiness compared to pruning. More importantly, we demonstrate that compressing a reliable large model via quantization can produce SLMs with superior trustworthiness and adaptability compared to using small models trained from scratch. Furthermore, knowledge distillation from trustworthy teacher models can further enhance the reliability of SLMs. We hope our findings provide practical guidance and a foundation for future research into the development and deployment of trustworthy small language models.
Chinese Translation
小语言模型(SLMs)作为传统大型语言模型(LLMs)的更高效替代方案,已在资源受限的场景中展现出良好的潜力。现有的小语言模型构建方法通常遵循两条路径:从头开始训练紧凑模型,或使用剪枝、量化或蒸馏等方法压缩较大的预训练模型。随着语言模型在实际应用中的日益普及,确保其可信度已成为一个关键问题。然而,如何构建可信的小语言模型仍然是一个未被充分探讨的问题。在本研究中,我们对小语言模型的可信度进行了多维度的综合评估,包括公平性、鲁棒性、隐私和伦理等方面。我们首先考察了剪枝和量化的影响,发现与剪枝相比,量化在保持可信度方面显著更为有效。更重要的是,我们证明通过量化压缩一个可靠的大型模型可以生成具有更高可信度和适应性的小语言模型,相较于从头训练的小模型。此外,从可信的教师模型进行知识蒸馏可以进一步增强小语言模型的可靠性。我们希望我们的研究结果能够为未来可信小语言模型的开发和部署提供实用指导和基础。
cs.CL / 45 / 2608.12008

Asymptotic Risk Calibration for Selective Question Answering

选择性问答的渐近风险校准
Lin, Shufan, Dong, Sijin
Abstract
Large language models (LLMs) may generate fluent but incorrect answers, making uncertainty quantification important for reliable question answering. However, heuristic uncertainty scores cannot perfectly distinguish correct predictions from incorrect ones, and directly applying a fixed uncertainty threshold provides no statistical control over the error rate among accepted answers. To address this limitation, we propose A-CRC-QA, a post-hoc calibration framework for uncertainty-aware selective question answering. The proposed method reformulates selection-conditioned error control as a linear expectation constraint and applies a monotonized empirical-risk calibration procedure inspired by conformal risk control. Since the resulting instance-wise loss is generally non-monotone with respect to the acceptance threshold, our framework targets asymptotic rather than finite-sample risk control. A-CRC-QA is model-agnostic, requires no additional training, and can be combined with different uncertainty estimators. Experiments on CoQA and MedMCQA demonstrate its applicability to both open-ended and closed-ended question answering, achieving a favorable trade-off between accepted-answer reliability and answer retention compared with uncalibrated and confidence-bound-based baselines.
Chinese Translation
大型语言模型(LLMs)可能生成流畅但不正确的答案,因此不确定性量化对于可靠的问答至关重要。然而,启发式不确定性评分无法完美区分正确预测与错误预测,直接应用固定的不确定性阈值对被接受答案的错误率没有统计控制。为了解决这一限制,我们提出了A-CRC-QA,一种针对不确定性感知选择性问答的事后校准框架。该方法将选择条件下的错误控制重新表述为线性期望约束,并应用受符合风险控制启发的单调经验风险校准程序。由于所得到的实例损失通常与接受阈值呈非单调关系,我们的框架旨在实现渐近而非有限样本的风险控制。A-CRC-QA是模型无关的,无需额外训练,并且可以与不同的不确定性估计器结合使用。在CoQA和MedMCQA上的实验表明,其在开放式和封闭式问答中的适用性,与未校准和基于置信区间的基线相比,实现了被接受答案的可靠性与答案保留之间的良好平衡。
cs.CL / 46 / 2608.12018

Poly-Dialectal Neural Machine Translation System for Bangla Regional Dialects

孟加拉地区方言的多方言神经机器翻译系统
Ullah, Rakib, Rahul, Ruhul Islam, Ahmed, Tanbir
Abstract
Regional dialectal variation poses a fundamental challenge to natural language processing (NLP) in Bangla, where over 240 million speakers communicate across diverse regional variants that diverge significantly from Standard Colloquial Bangla (SCB) in phonology, morphology, and lexicon. Contemporary neural machine trans- lation (NMT) architectures and large language models (LLMs) predominantly as- sume a homogeneous language distribution, resulting in severe performance degra- dation when translating low-resource regional dialects. In this work, we present a unified Poly-Dialectal Neural Machine Translation System capable of multi-directional translation across 12 Bangla regional dialects without routing through an inter- mediary standard pivot. We compile the largest multi-dialect parallel corpus for Bangla to date, comprising 51,531 non-null parallel sentence pairs across 12 di- alects, incorporating 2,500 expert-verified, bidirectional parallel sentence pairs for five previously unaddressed dialects. Evaluating sequence-to-sequence architec- tures under Weight-Decomposed Low-Rank Adaptation (DoRA), our fine-tuned BanglaT5 model achieves state-of-the-art translation performance (29.26 BLEU, 57.26 chrF++), outperforming NLLB-200 (615M) and mBART-50 (611M) while preserving morphological coherence. Furthermore, we conduct a systematic cross- dialectal transfer analysis and dataset scaling study, establishing empirical thresh- olds for low-resource dialect adaptation. Finally, we deploy the optimized INT8- quantized model as an open-access web application to promote digital inclusion for marginalized dialect communities. The complete dataset is publicly available at Mendeley Data (https://data.mendeley.com/datasets/v9cf66fk2t/2).
Chinese Translation
地区方言变异对孟加拉的自然语言处理(NLP)构成了根本挑战,超过2.4亿讲者在不同的地区变体中交流,这些变体在音韵、形态和词汇上与标准口语孟加拉语(SCB)有显著差异。当前的神经机器翻译(NMT)架构和大型语言模型(LLMs)主要假设语言分布是均匀的,因此在翻译低资源地区方言时,性能严重下降。在本研究中,我们提出了一种统一的多方言神经机器翻译系统,能够在12种孟加拉地区方言之间进行多向翻译,而无需通过中介标准枢轴。我们编制了迄今为止最大的孟加拉多方言平行语料库,包含12种方言的51,531对非空平行句子,并为五种之前未被关注的方言纳入了2,500对经过专家验证的双向平行句子。在权重分解低秩适应(DoRA)下评估序列到序列架构,我们微调的BanglaT5模型实现了最先进的翻译性能(29.26 BLEU,57.26 chrF++),超越了NLLB-200(615M)和mBART-50(611M),同时保持了形态一致性。此外,我们进行了系统的跨方言迁移分析和数据集扩展研究,确立了低资源方言适应的经验阈值。最后,我们将优化后的INT8量化模型部署为开放访问的网络应用,以促进边缘方言社区的数字包容性。完整数据集可在Mendeley Data(https://data.mendeley.com/datasets/v9cf66fk2t/2)公开获取。
cs.CL / 47 / 2608.12062

Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations

偏好树优化:通过前瞻性模拟增强目标导向对话
Baruch, Lior, Butman, Moshe, Bar, Kfir, Friedman, Doron
Abstract
Developing dialogue systems capable of engaging in multi-turn, goal-oriented conversations remains a significant challenge, especially in specialized domains with limited data. This research proposes a novel framework called Preference Tree Optimization (PTO), designed to iteratively improve agent models in such dialogue systems, by generating preference data using a method called Preference Tree with Look-Ahead. Focusing on Motivational Interviewing (MI) -- a counseling technique aimed at facilitating behavioral change -- we leverage virtual patients and an oracle evaluator to simulate conversations and generate rich preference datasets. By combining this method with Direct Preference Optimization (DPO), we aim to enhance the agent's decision-making capabilities over iterative training cycles. The proposed framework addresses data scarcity and advances the development of more nuanced and effective dialogue systems in goal-oriented domains. Experimental evaluations demonstrate that the PTO framework enhances dialogue agents' performance in goal-oriented conversations within the domain of Motivational Interviewing (MI). Models trained with PTO consistently outperformed the baseline in key metrics such as session satisfaction and working alliance. Additionally, incorporating look-ahead simulations led to improved long-term planning and more effective conversational strategies, with deeper look-ahead configurations yielding the most stable and high-scoring results.
Chinese Translation
开发能够进行多轮目标导向对话的对话系统仍然是一个重大挑战,尤其是在数据有限的专业领域。本文提出了一种名为偏好树优化(Preference Tree Optimization, PTO)的新框架,旨在通过使用一种称为带前瞻的偏好树(Preference Tree with Look-Ahead)的方法,迭代地改进此类对话系统中的代理模型,生成偏好数据。我们专注于动机访谈(Motivational Interviewing, MI)——一种旨在促进行为改变的咨询技术——利用虚拟患者和一个评估者模拟对话并生成丰富的偏好数据集。通过将此方法与直接偏好优化(Direct Preference Optimization, DPO)相结合,我们旨在增强代理在迭代训练周期中的决策能力。该框架解决了数据稀缺问题,并推动了在目标导向领域中开发更细致和有效的对话系统。实验评估表明,PTO框架在动机访谈(MI)领域的目标导向对话中提升了对话代理的表现。使用PTO训练的模型在会话满意度和工作联盟等关键指标上始终优于基线。此外,结合前瞻性模拟提高了长期规划和更有效的对话策略,深入的前瞻配置产生了最稳定和高分的结果。
cs.CL / 48 / 2608.12113

Structuring the Space of Perspectives

构建视角空间
Daffara, Agnese, Padó, Sebastian, Ceron, Tanise
Abstract
The same event can be reported from different perspectives depending on the experiences, background, and beliefs of the writer or speaker. A variety of NLP areas engage with perspectives, spanning from text analysis to algorithm optimization. A wide range of operative concepts (such as stances, sentiment, frames, and arguments) has been used to capture perspectives in texts, however the precise relationships among those concepts remain unclear. Arguably, a deeper theoretical understanding of these concepts would empower more effective research on perspectives. In this paper, we address this gap by reviewing the space of perspectives in NLP and defining a set of properties that help distinguishing perspective-related concepts. Our analysis leads us to posit a hierarchy which organizes these concepts linearly along a single axis. Finally, we show how this principled conceptual hierarchy can help researchers navigate the field and select operationalizations of perspective that align with their specific research objectives.
Chinese Translation
同一事件可以根据撰写者或发言者的经历、背景和信念从不同的视角进行报道。多种自然语言处理(NLP)领域涉及视角,从文本分析到算法优化不一而足。为了捕捉文本中的视角,已经使用了多种操作性概念(如立场、情感、框架和论点),然而这些概念之间的确切关系仍不清晰。可以说,对这些概念的更深入理论理解将促进更有效的视角研究。在本文中,我们通过回顾NLP中的视角空间并定义一组有助于区分与视角相关概念的属性来填补这一空白。我们的分析使我们提出了一个层次结构,该结构将这些概念沿单一轴线线性组织。最后,我们展示了这一原则性概念层次如何帮助研究人员在该领域中导航,并选择与其特定研究目标相一致的视角操作化方案。
cs.CL / 49 / 2608.12121

QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving

QV-PIC:用于高效RAG服务的查询感知视觉位置无关缓存
Liu, Yilin, Meng, Rui, Ni, Wangze, Yan, Jianxin, Cao, Heng, Zheng, Libin, Cheng, Peng, Liu, Jinfei
Abstract
Retrieval-Augmented Generation (RAG) repeatedly prefills identical text chunks across queries, incurring redundant computations. Position-Independent Caching (PIC) mitigates it by reusing precomputed Key-Value (KV) across positions, but its efficiency is constrained by the large volume of text tokens. Rendering text chunks as images can compress the text into fewer visual tokens, but the rendered-image PIC suffers more severe quality degradation than the text PIC. This representation-specific gap primarily arises from contextual mismatches across independently compiled caches and the loss of fine-grained textual evidence during visual compression. Existing PIC repair methods mainly address the former through selective recomputation, but they incur online computation and cannot recover lost textual details. We propose QV-PIC, a query-aware dual-resolution PIC reuse framework guided by model-native templates. Offline, QV-PIC compiles visual caches under the model's native chat-template prefix, improving PIC quality without online recomputation. Online, it preserves global context with low resolution and restores fine-grained textual evidence within a high-resolution budget by cumulative query relevance scores, retaining the efficiency benefit of visual compression. Across six tasks, QV-PIC improves average F1 by 21.6 points over vanilla rendered-image PIC, closes the gap to vanilla text PIC, and surpasses optimized text PIC by 2.58 F1 while reducing TTFT by 17.2\%. Relative to full prefill, it cuts TTFT by 83.8%.
Chinese Translation
检索增强生成(RAG)在多个查询中重复填充相同的文本块,导致冗余计算。位置无关缓存(PIC)通过在不同位置重用预计算的键值对(KV)来缓解这一问题,但其效率受到大量文本标记的限制。将文本块呈现为图像可以将文本压缩为更少的视觉标记,但渲染图像的PIC比文本PIC遭受更严重的质量下降。这种特定于表示的差距主要源于独立编译缓存之间的上下文不匹配,以及在视觉压缩过程中细粒度文本证据的丢失。现有的PIC修复方法主要通过选择性重新计算解决前者,但它们会引入在线计算,并无法恢复丢失的文本细节。我们提出了QV-PIC,这是一种查询感知的双分辨率PIC重用框架,以模型原生模板为指导。在离线阶段,QV-PIC在模型的原生聊天模板前缀下编译视觉缓存,提高了PIC质量而无需在线重新计算。在在线阶段,它以低分辨率保留全局上下文,并通过累积查询相关性得分在高分辨率预算内恢复细粒度文本证据,从而保持视觉压缩的效率优势。在六个任务中,QV-PIC的平均F1提高了21.6分,相较于普通渲染图像PIC,缩小了与普通文本PIC的差距,并超越了优化文本PIC 2.58 F1,同时将TTFT减少了17.2%。与完全预填充相比,它将TTFT减少了83.8%。
cs.CL / 50 / 2608.12129

SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges

SAG:带有查询时动态超边的SQL检索增强生成
Wu, Yuchao, Li, Junqin, Liang, XingCheng, Chen, Yongjie, Liang, Yinghao, Mo, Linyuan, Li, Guanxian
Abstract
While retrieval-augmented generation (RAG) has proven effective at giving LLMs access to external knowledge, mainstream dense-retrieval implementations remain inherently limited in handling structured constraints and multi-hop reasoning. Graph-based methods address this by constructing knowledge graphs offline, but they often fragment semantics, incur high maintenance, and complicate incremental updates. We propose SAG (SQL-Retrieval Augmented Generation), a structured retrieval architecture that organizes documents into an event-entity index without building a global knowledge graph. SAG represents each chunk as a semantically complete event paired with its entities, forming a latent hyperedge that preserves n-ary relations without decomposing them into triples. At query time, SAG treats shared entities as join keys to connect related chunks. This dynamically yields a query-scoped neighborhood of events, and yet every piece of evidence remains the original chunk throughout. Experiments on HotpotQA, 2WikiMultiHopQA, and MuSiQue show that SAG achieves the best retrieval and end-to-end QA performance on every benchmark, with gains that widen as reasoning-chain complexity increases. On MuSiQue, where multi-hop evidence chaining is most demanding, SAG reaches 80.36% Recall@5, outperforming the strongest baseline by 11.52 points. This work paves the way for knowledge infrastructure that enables LLM agents to retrieve and reason over continually growing organizational knowledge.
Chinese Translation
尽管检索增强生成(RAG)已被证明能够有效地为大型语言模型(LLMs)提供外部知识的访问,但主流的密集检索实现仍然在处理结构化约束和多跳推理方面存在固有的局限性。基于图的方法通过离线构建知识图谱来解决这一问题,但它们往往会导致语义碎片化,维护成本高,并且使增量更新变得复杂。我们提出了SAG(SQL检索增强生成),这是一种结构化检索架构,能够将文档组织成事件-实体索引,而无需构建全局知识图谱。SAG将每个块表示为一个语义上完整的事件,并与其实体配对,形成一个潜在的超边,从而保留n元关系,而无需将其分解为三元组。在查询时,SAG将共享实体视为连接相关块的连接键。这动态地生成了一个查询范围内的事件邻域,并且每个证据片段在整个过程中仍保持原始块的状态。在HotpotQA、2WikiMultiHopQA和MuSiQue上的实验表明,SAG在每个基准测试中都实现了最佳的检索和端到端问答性能,随着推理链复杂性的增加,性能提升更加显著。在MuSiQue上,面对最具挑战性的多跳证据链,SAG达到了80.36%的Recall@5,超越了最强基线11.52个百分点。这项工作为知识基础设施铺平了道路,使大型语言模型代理能够检索和推理不断增长的组织知识。
cs.CL / 51 / 2608.12149

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

混合线性注意力大型语言模型中的大规模激活:注意力前峰值与峰值间平台
Su, Zunhai, Sun, Bohan, Zhuang, Xialie, Zhang, Shuibai, Xiao, He, Xiong, Jing, Zhang, Hengyuan, Zhou, Zhongzhu, Zhang, Tiantian, Wong, Ngai, Kuo, Chuan-Wei
Abstract
We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre-attention spikes (PAS), and can persist through intervening linear attention layers, giving rise to inter-spike plateaus (ISP). As full attention becomes denser, successive PAS become increasingly connected through ISP, ultimately recovering the stable MA morphology of full attention LLMs. We establish the recurrence of this organization across five linear attention architectures, six hybridization configurations, five data domains, and representative open-source hybrid models spanning 1.2B to 397B total parameters. Controlled pretraining of GDN-based hybrids at scales up to 1.3B shows that both morphologies emerge early and respond asymmetrically to output gating: full attention output gating strongly attenuates their absolute magnitudes without eliminating their layerwise organization, whereas removing GDN gates yields comparatively modest amplification. Mechanistically, our systematic-outlier analysis supports a shared lifecycle account governed by the timing of MA cancellation. PAS follows a localized write-sink-cancel process, while the extended persistence of ISP is consistent with delayed cancellation. At the full attention limit, this account recovers the stable MA morphology characteristic of full attention LLMs. Our code is available at https://github.com/StartluxLabs/Massive-Activations-HLA.
Chinese Translation
我们首次系统地研究了层间交错的混合线性注意力大型语言模型(HLA LLMs)中的大规模激活(MAs),并揭示了两种与架构相一致的形态:MAs 在全注意力层之前始终会迅速激增,形成注意力前峰值(PAS),并且可以在介入的线性注意力层中持续存在,产生峰值间平台(ISP)。随着全注意力变得更加密集,连续的 PAS 通过 ISP 变得越来越紧密连接,最终恢复全注意力 LLMs 的稳定 MA 形态。我们在五种线性注意力架构、六种混合配置、五个数据领域以及涵盖 12 亿到 3970 亿总参数的代表性开源混合模型中建立了这种组织的重复性。在高达 13 亿规模的基于 GDN 的混合模型的受控预训练中,显示出这两种形态早期出现,并且对输出门控的反应不对称:全注意力的输出门控强烈减弱了它们的绝对幅度,但并未消除它们的层次组织,而去除 GDN 门控则导致相对温和的放大。从机制上讲,我们的系统性异常分析支持一个由 MA 取消时机主导的共享生命周期解释。PAS 遵循局部写入-汇聚-取消过程,而 ISP 的延续性与延迟取消一致。在全注意力极限下,该解释恢复了全注意力 LLMs 特有的稳定 MA 形态。我们的代码可在 https://github.com/StartluxLabs/Massive-Activations-HLA 获取。
cs.CL / 52 / 2608.12218

Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge

信息丰富性悖论:长上下文训练削弱参数知识
Uzunoglu, Arda, van Durme, Benjamin, Khashabi, Daniel
Abstract
Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories. This scaling reflects the implicit assumption that training on longer contexts will only help the model by exposing it to richer evidence. We challenge this view by studying how the context window shapes a model's mode of learning, shifting it between parametric internalization and contextualization. We propose the Information Abundance Paradox, which hypothesizes that abundant relevant information in the training context can reduce the incentive to encode that information parametrically, thereby increasing reliance on context. In pretraining with long documents, increasing the context window improves language modeling, natural language understanding, and closed-book MCQA only up to an intermediate optimum, after which performance consistently declines. In supervised fine-tuning, more task-relevant train-time context improves performance with supporting context, but reduces robustness when context is absent or misleading at test time. Our analysis suggests that this behavior arises when longer context provides a lower complexity solution. Mechanistically, training with informative context shifts gradient pressure from feed-forward networks, often linked to parametric knowledge, toward attention modules, and causal interventions show that this shift increases reliance on context during inference. Overall, these findings support the Information Abundance Paradox and suggest that scaling toward near-infinite context is not simply a matter of supplying more data, even when high-quality long-context data is abundant.
Chinese Translation
大型语言模型越来越多地在跨文档、代码库和交互历史的长上下文中进行训练和部署。这一扩展反映了一个隐含假设,即在更长的上下文中训练只会通过向模型提供更丰富的证据来帮助模型。我们通过研究上下文窗口如何塑造模型的学习模式来挑战这一观点,发现其在参数内部化和情境化之间转变。我们提出了信息丰富性悖论,假设训练上下文中丰富的相关信息可能减少以参数形式编码该信息的动机,从而增加对上下文的依赖。在使用长文档进行预训练时,增加上下文窗口在语言建模、自然语言理解和闭卷多项选择问答(MCQA)方面的表现仅在达到一个中间最优点后有所改善,之后性能持续下降。在监督微调中,更多与任务相关的训练时上下文在支持上下文的情况下提高了性能,但在测试时上下文缺失或误导时则降低了鲁棒性。我们的分析表明,当更长的上下文提供了较低复杂度的解决方案时,会出现这种行为。从机制上讲,使用信息丰富的上下文进行训练将梯度压力从前馈网络(通常与参数知识相关)转移到注意力模块,并且因果干预表明,这种转变在推理过程中增加了对上下文的依赖。总体而言,这些发现支持了信息丰富性悖论,并表明,向近乎无限的上下文扩展并不仅仅是提供更多数据的问题,即使高质量的长上下文数据也很丰富。
cs.CL / 53 / 2608.12253

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

单一模拟器不足以应对:多智能体强化学习中的模拟器崩溃
Yu, Simon, Tomlin, Nicholas, Abdulhai, Marwa, Lu, Ximing, Chong, Derek, Hou, Abe, Soylu, Dilara, Levine, Sergey, Manning, Christopher D., Shi, Weiyan
Abstract
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM policy trained against it overfits to narrow strategies that exploit the simulator's dominant mode, and such a policy transfers poorly to unseen simulators and real users. We formalize this collapse theoretically and propose two complementary solutions, one at inference time and one at training time. The inference-time solution, Verbalized Sampling, broadens the simulator's behavior by sampling from a verbalized response distribution, reducing mode collapse. The training-time solution, Co-Training, jointly optimizes the policy against a population of trainable simulators, preventing it from overfitting to any single simulator's mode. We validate both solutions on three multi-turn benchmarks: Persuasion for Good, $\tau^2$-bench, and CooperBench. Verbalized Sampling improves held-out success by up to 9% over single-simulator RL, and Co-Training pushes gains further to 14%; the human study shows similar gain on real users. Both solutions preserve the policy diversity that collapses under single-simulator RL. To support further work in this direction, we release SCOPE, an open-source framework for Population Co-Training multi-agent RL. More broadly, our results suggest that the diversity of the training environment, not only the policy, is critical to the generalization of multi-turn RL to real-world deployment.
Chinese Translation
多智能体强化学习在人机交互中通常依赖于单一大型语言模型来模拟用户行为。我们表明这种方法系统性地无法泛化,并将其失败归因于模拟器崩溃:由于模拟器的LLM(大型语言模型)发生了模式崩溃,针对其训练的LLM策略过拟合于利用模拟器主导模式的狭窄策略,而这种策略在未见过的模拟器和真实用户中迁移效果不佳。我们从理论上形式化了这种崩溃,并提出了两个互补的解决方案,一个在推理阶段,一个在训练阶段。推理阶段的解决方案“Verbalized Sampling”(语言化采样)通过从语言化响应分布中采样来拓宽模拟器的行为,从而减少模式崩溃。训练阶段的解决方案“Co-Training”(协同训练)则是针对一组可训练的模拟器共同优化策略,防止其过拟合于任何单一模拟器的模式。我们在三个多轮基准测试上验证了这两种解决方案:Persuasion for Good、$ au^2$-bench 和 CooperBench。Verbalized Sampling在保留集上的成功率比单一模拟器强化学习提高了最多9%,而Co-Training进一步将增益推高至14%;人类研究显示在真实用户上也有类似的增益。这两种解决方案都保持了在单一模拟器强化学习下崩溃的策略多样性。为了支持该方向的进一步研究,我们发布了SCOPE,一个用于人口协同训练多智能体强化学习的开源框架。更广泛地说,我们的结果表明,训练环境的多样性,而不仅仅是策略,对于多轮强化学习在现实世界部署中的泛化至关重要。
cs.CL / 54 / 2608.12269

A Cascaded Unsupervised-Supervised NLP Pipeline for Detecting Accusatory Language in Public Procurement

用于检测公共采购中指责性语言的级联无监督-监督自然语言处理管道
Torres, Bryan, Riofrío, Daniel, Vega-Sánchez, José, Orozco, Nathaly, Parra, Carla, Rosero, Karen, Grijalva, Felipe
Abstract
Public procurement involves the allocation of substantial financial resources; therefore, continuous oversight through audits, controls, and monitoring mechanisms is essential. However, stakeholder comments and publicly available government data are often underutilized, despite their potential to reveal procedural irregularities. To address this gap, this paper analyzes metadata from Ecuador's Sistema Oficial de Contrataci\'on P\'ublica (SOCE, Official Public Procurement System), with particular emphasis on participant comments generated during the pre-contractual phase. We propose a hybrid modeling framework that integrates unsupervised clustering and supervised classification within a natural language processing (NLP) pipeline to uncover latent patterns and detect potentially irregular procurement processes. Semantic embeddings are generated using Word2Vec, LLaMA, and RoBERTa, followed by Gaussian Mixture Models (GMMs) for unsupervised clustering. A supervised classification stage is then applied to identify accusatory or whistleblowing-style comments. Experimental results show that the combination of domain-trained Word2Vec embeddings, GMM-based clustering, and a Random Forest classifier achieves high precision and recall, even under severe class imbalance. These findings demonstrate that lightweight, domain-adapted NLP architectures can effectively support risk identification and enhance transparency in public procurement systems without requiring large-scale computational infrastructure.
Chinese Translation
公共采购涉及大量财政资源的分配,因此,通过审计、控制和监测机制进行持续监督至关重要。然而,尽管利益相关者的评论和公开的政府数据具有揭示程序不规范的潜力,但它们往往未被充分利用。为了解决这一问题,本文分析了厄瓜多尔的官方公共采购系统(Sistema Oficial de Contratación Pública, SOCE)的元数据,特别强调了在合同前阶段生成的参与者评论。我们提出了一种混合建模框架,将无监督聚类和监督分类集成在自然语言处理(NLP)管道中,以揭示潜在模式并检测可能不规范的采购过程。使用Word2Vec、LLaMA和RoBERTa生成语义嵌入,随后应用高斯混合模型(Gaussian Mixture Models, GMMs)进行无监督聚类。然后,应用监督分类阶段以识别指责性或举报风格的评论。实验结果表明,领域训练的Word2Vec嵌入、基于GMM的聚类和随机森林分类器的组合在严重类别不平衡的情况下仍能实现高精度和高召回率。这些发现表明,轻量级、领域适应的NLP架构可以有效支持风险识别,并增强公共采购系统的透明度,而无需大型计算基础设施。
cs.CL / 55 / 2608.12278

Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

结构性沉默:当人工智能基础设施未能支持弱势语言使用者时
Roy, Avijit, Roy, Proma
Abstract
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained. This paper examines these structural barriers through Bengali, one of the world's most widely spoken languages, focusing on AI-assisted education in low-connectivity environments. We identify four interlocking failures: a severe web presence gap, with Bengali accounting for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit between English and Bengali in major multilingual corpora; a tokenization penalty associated with Bengali's alphasyllabary script that compounds the data deficit through higher token fertility; and connectivity exclusion, with individual internet penetration at 36.5% in rural areas compared with 71.4% in urban areas. These failures reflect longstanding resource-allocation decisions, institutional priorities, and design defaults that did not center underrepresented languages in mainstream AI development. We argue that dataset scarcity should be understood as a structural barrier rather than an isolated technical limitation, and that offline-first design should be treated as an equity-oriented infrastructure strategy. We conclude with directions for linguistics and AI research aimed at reducing these structural inequalities.
Chinese Translation
人工智能教育和语言支持工具越来越被视为对资源匮乏社区中获取差距的可扩展响应。然而,这些工具背后的基础设施,包括训练语料库、分词方案、评估基准和部署架构,可能在模型训练之前就系统性地使弱势语言使用者处于不利地位。本文通过孟加拉语这一全球使用最广泛的语言之一,探讨这些结构性障碍,重点关注低连接环境下的人工智能辅助教育。我们识别出四个相互交织的失败:网络存在差距严重,尽管孟加拉语占全球人口的近4%,但其在全球网络内容中所占比例不足0.5%;在主要的多语言语料库中,英语与孟加拉语之间存在67:1的训练标记缺口;与孟加拉语的音节字母书写系统相关的分词惩罚,通过更高的标记丰度加剧了数据缺口;以及连接排斥,农村地区的个人互联网普及率为36.5%,而城市地区为71.4%。这些失败反映了长期以来的资源分配决策、机构优先事项和设计缺陷,这些都没有将弱势语言置于主流人工智能发展的中心。我们认为,数据集稀缺应被理解为一种结构性障碍,而非孤立的技术限制,并且离线优先设计应被视为一种以公平为导向的基础设施战略。最后,我们提出了旨在减少这些结构性不平等的语言学和人工智能研究方向。