← Back to Index
Daily Research Digest

arXiv Papers

2026-08-03
202
Papers
4
Categories
201
Translated
收藏清单 0
机器人学 (Robotics)
33
cs.RO / 1 / 2607.28952

Advances, challenges, and opportunities for legged robots

腿部机器人发展的进展、挑战与机遇
Frey, Jonas, Mattamala, Matías, Park, Hae-Won, Mittal, Mayank, Martius, Georg, Osborne, Maike, Sparrow, Robert, Hutter, Marco
Abstract
Humanoid and quadrupedal robots have the potential to revolutionize the way we work, interact, and coexist with intelligent machines. To understand their effects on society and how they can enable scientific discovery, we assess the current capabilities of these systems along hardware, locomotion, autonomy, data, and applications. We identify recent advances and key open challenges that must be overcome to enable widespread adoption and new use cases for legged robots. Last, we provide an outlook on the future of legged robots, exploring their ethical considerations, economic potential, policy implications, and broader societal effects.
Chinese Translation
类人和四足机器人有潜力彻底改变我们与智能机器的工作、互动和共存方式。为了理解它们对社会的影响以及如何促进科学发现,我们评估了这些系统在硬件、运动、自治、数据和应用方面的当前能力。我们识别了近期的进展和必须克服的关键开放挑战,以促进腿部机器人的广泛应用和新的使用案例。最后,我们展望了腿部机器人的未来,探讨了它们的伦理考量、经济潜力、政策影响以及更广泛的社会效应。
cs.RO / 2 / 2607.28993

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

ST-WAM:用于视觉分布变化下稳健操作的语义-时间世界动作模型
Wang, Mingxin, Hu, Bin, Qian, Bin, Jiang, Kaitao, Wu, Haoning, Yan, Feng, Jing, Bowen, Hao, Ruiyang, Wang, Enyi, Niu, Kangning, Yang, Yandan, Xu, Mu, Wang, Yan, Liu, Houde, Li, Tianlun
Abstract
World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual content, limiting robustness under visual distribution shifts. We identify Training-Distribution Hallucination, a recurring phenomenon in which futures conditioned on visually shifted observations hallucinate training-domain content rather than remain faithful to the current scene. A controlled frame-triplet diagnosis further shows that DINOv3 features remain more stable across visual shifts while better preserving task-state distinctions than Wan-VAE latents. Rather than correcting the predicted futures, we propose Semantic-Temporal WAM (ST-WAM) to improve action robustness by using DINOv3 as a shared semantic representation for future prediction and history retrieval while retaining fine-grained VAE dynamics. Its Dual-Space Future Experts (DSFE) jointly predict future VAE latents and DINO features, while Current-Anchored Intent Retrieval (CAIR) retrieves task-relevant evidence from recent DINO history under the current visual-language context. ST-WAM is trained end-to-end without additional embodied pretraining or task-specific annotations, and requires no explicit future generation at inference. It achieves 98.7% on LIBERO and 92.8% on RoboTwin 2.0; more importantly, compared with Fast-WAM, it improves zero-shot LIBERO-Plus performance by 21.3 percentage points and more than doubles real-world success under visual shifts from 25.8% to 61.5%. These results demonstrate that semantic-temporal modeling effectively complements pixel-generative dynamics for robust manipulation.
Chinese Translation
世界动作模型(WAMs)作为一种有前景的范式,通过联合建模机器人动作和未来视觉动态而崭露头角。然而,它们对像素生成的未来监督的依赖可能会将与动作相关的状态转变与与任务无关的视觉内容混淆,从而限制了在视觉分布变化下的稳健性。我们识别出训练分布幻觉(Training-Distribution Hallucination),这是一种反复出现的现象,其中基于视觉偏移观察的未来预测会幻觉出训练域内容,而不是忠实于当前场景。通过受控的帧三元组诊断进一步表明,DINOv3特征在视觉变化中保持更稳定,同时更好地保留任务状态的区分性,而Wan-VAE潜变量则不然。我们提出语义-时间WAM(ST-WAM),以通过使用DINOv3作为未来预测和历史检索的共享语义表示来提高动作的稳健性,同时保留细粒度的VAE动态,而不是纠正预测的未来。其双空间未来专家(Dual-Space Future Experts, DSFE)共同预测未来的VAE潜变量和DINO特征,而当前锚定意图检索(Current-Anchored Intent Retrieval, CAIR)则在当前视觉-语言上下文中从最近的DINO历史中检索与任务相关的证据。ST-WAM是端到端训练的,无需额外的具身预训练或任务特定的注释,并且在推理时不需要显式的未来生成。它在LIBERO上达到了98.7%的准确率,在RoboTwin 2.0上达到了92.8%;更重要的是,与Fast-WAM相比,它在零-shot LIBERO-Plus性能上提高了21.3个百分点,并且在视觉变化下的实际成功率从25.8%翻倍至61.5%。这些结果表明,语义-时间建模有效地补充了像素生成动态,以实现稳健的操作。
cs.RO / 3 / 2607.28995

Receding-Horizon Next-Best-View Planner for Autonomous Leaf Surface Reconstruction

用于自主叶片表面重建的后退视野下一个最佳视角规划器
Ahmed, Arif, Das, Sajal K., Maini, Parikshit
Abstract
Accurate plant leaf modeling is fundamental to downstream tasks such as plant growth monitoring, and phenotyping for yield estimation. Autonomous robotic reconstruction for large-scale field deployment must address limitations on robot planning budget and computation resources while optimizing viewpoint utility for leaf surface reconstruction. Existing approaches either focus on rigid objects, point-cloud coverage or plant reconstruction without fully addressing the system limitations or exploiting task-driven point cloud utility. In this work, we study next-best-view (NBV) planning for leaf surface reconstruction under travel constraints. We develop a novel Centroid-based Information Gain (CIG) function that measures the spatial distribution of observed points relative to the centroid of the existing point cloud to compute viewpoint utility. We also develop a receding-horizon variant that reasons over future viewpoints. To benchmark our work, we use the LAST-STRAW [1] public dataset that includes point clouds of strawberry plants over different growth stages and compare our method with attention-driven NBV [2] that uses a visibility-based information gain approach. The proposed receding-horizon approach consistently reduces surface reconstruction error and improves geometric fidelity across multiple growth stages, especially under increased inter-leaf occlusion. Results demonstrate that our approach is able to visit viewpoints that reduce surface reconstruction error and improves reconstruc-tion accuracy as compared to the baseline by upto 10%.
Chinese Translation
准确的植物叶片建模对植物生长监测和产量估计的表型分析等下游任务至关重要。大规模田野部署的自主机器人重建必须在优化叶片表面重建的视角效用的同时,解决机器人规划预算和计算资源的限制。现有方法要么专注于刚性物体、点云覆盖,或者植物重建,而没有充分解决系统限制或利用任务驱动的点云效用。在本研究中,我们研究了在旅行约束下的叶片表面重建的下一个最佳视角(NBV)规划。我们开发了一种新颖的基于质心的信息增益(Centroid-based Information Gain, CIG)函数,该函数测量相对于现有点云质心的观察点的空间分布,以计算视角效用。我们还开发了一种后退视野变体,该变体考虑未来的视角。为了对我们的工作进行基准测试,我们使用了LAST-STRAW [1]公共数据集,该数据集包括不同生长阶段的草莓植物的点云,并将我们的方法与使用基于可见性的信息增益方法的注意力驱动的NBV [2]进行比较。所提出的后退视野方法在多个生长阶段中始终减少表面重建误差,并提高几何保真度,尤其是在叶片间遮挡增加的情况下。结果表明,与基线相比,我们的方法能够访问减少表面重建误差的视角,并将重建精度提高了多达10%。
cs.RO / 4 / 2607.29009

D-VLC: Decentralized Vision-Language Collaboration for Heterogeneous Embodied Multi-Robot Systems in Unknown Environments

D-VLC:用于未知环境中异构具身多机器人系统的去中心化视觉-语言协作
Zhou, Yuan, Lin, Ruitong, Wang, Shen, Gai, Weiqi, zhu, Mo, Zhou, Xin, Wu, Yuze, Gao, Fei
Abstract
Multi-robot systems, particularly heterogeneous robot swarms, can improve the efficiency of complex task execution through parallel collaboration and complementary capabilities. However, conventional rule-based methods rely on predefined task models and specialized decision making programs, making it difficult to understand complex semantic instructions and coordinate heterogeneous robots. LLMs introduce strong language understanding and task reasoning capabilities, allowing multi-robot systems to interpret instructions, decompose tasks, and assign roles according to task semantics. VLMs further incorporate visual perception, enabling robots to reason about objects, regions, and spatial relationships in physical environments. Nevertheless, existing LLM/VLM based methods often depend on known maps, centralized and synchronized decision making, limiting their generalization to heterogeneous robots and unseen tasks. We therefore propose a framework that combines decentralized asynchronous reasoning, lightweight information sharing, capability aware collaboration, and a unified action interface, enabling general purpose VLMs to generate robot specific actions executed by learning free experts without task or robot specific training. Experiments across diverse scenarios and multiple VLMs show success rates above 70\%, with completion time reduced by up to 55.8\% relative to the geometric greedy baseline.
Chinese Translation
多机器人系统,特别是异构机器人群体,可以通过并行协作和互补能力提高复杂任务执行的效率。然而,传统的基于规则的方法依赖于预定义的任务模型和专门的决策程序,使得理解复杂的语义指令和协调异构机器人变得困难。大型语言模型(LLMs)引入了强大的语言理解和任务推理能力,使多机器人系统能够解释指令、分解任务并根据任务语义分配角色。视觉语言模型(VLMs)进一步结合了视觉感知,使机器人能够推理物体、区域和物理环境中的空间关系。然而,现有的基于LLM/VLM的方法通常依赖于已知地图、集中和同步的决策,限制了它们在异构机器人和未知任务中的泛化能力。因此,我们提出了一个框架,结合去中心化的异步推理、轻量级信息共享、能力感知的协作和统一的行动接口,使通用VLM能够生成特定于机器人的动作,由无任务或机器人特定训练的学习自由专家执行。在多种场景和多个VLM的实验中,成功率超过70%,完成时间相比几何贪婪基线减少了多达55.8%。
cs.RO / 5 / 2607.29011

DART: Dual-Axis Airborne Reachability-Gated Torque-Reaction for Off-Road Vehicle Jumps

DART:双轴空中可达性门控扭矩反应用于越野车辆跳跃
Hu, Yu, Zhao, Fangzhou, Sang, Mingyuan, Min, Cheng, Chen, Liang, Li, Wei, Kuang, Wenyu, Chen, Shican, Li, Jinwei, Chen, Baolei
Abstract
Traversing crests, ledges, and ditches at high speed often launches vehicles into the air, and a mishandled landing presents a substantial crash hazard. We show that the airborne phase is barely controllable: on a 1383 kg platform the wheel angular-momentum budget caps the recoverable pitch-rate change at roughly $9$-$13^\circ$/s in the tighter nose-up direction under drive at typical takeoff wheel speeds, and at about twice that in the reverse-inclusive braking direction; driving the wheels to their drivetrain hard limit raises the measured nose-up ceiling to only $16$-$18^\circ$/s. Takeoff pitch-rate disturbances beyond this directional budget are physically unrecoverable in flight, so the decisive leverage lies before takeoff. DART (Dual-Axis Airborne Reachability-Gated Torque-Reaction) back-propagates the landing constraint into a closed-form certified feasible-takeoff set, which supplies a conservative go/no-go condition and a pre-takeoff speed-shaping law. In flight, DART regulates pitch and roll via steer-resolved wheel-reaction torque, governed by a per-flight roll latch derived from the yaw-coupling analysis. In deterministic full-scale simulation in BeamNG.tech, a calibrated pre-takeoff speed regulator reduces touchdown speed by 36% and raises on-target landings from 0/30 to 30/30. Under the same steep-lip approach the airborne law completes 29/30 safe landings under crash-avoidance bounds versus 0/30 for reaction-wheel-style PD (RW-PD) and time-optimal bang-bang (TOBB). On banked run-ups DART holds the median pitch error at or below $2^\circ$ at every cross-slope, with the largest baseline separation at $\gamma=12^\circ$. Across disturbance regimes, the latch preserves pitch-only allocation on low-disturbance entries and enables dual-axis control when roll becomes binding. All results are from simulation; hardware validation remains open.
Chinese Translation
以高速穿越山脊、边缘和沟渠常常会使车辆腾空而起,而处理不当的着陆则会带来显著的碰撞风险。我们表明空中阶段几乎不可控:在1383千克的平台上,轮子角动量预算限制了在典型起飞轮速下可恢复的俯仰速率变化,紧凑的向上方向约为$9$-$13^ heta$/s,而在包括倒退的制动方向上约为其两倍;将轮子驱动至其动力传动系统的极限仅将测得的向上极限提高至$16$-$18^ heta$/s。超出这一方向预算的起飞俯仰速率扰动在飞行中是物理上不可恢复的,因此决定性的杠杆作用发生在起飞之前。DART(双轴空中可达性门控扭矩反应)将着陆约束反向传播到一个闭式形式的认证可行起飞集,这提供了一个保守的起飞/不起飞条件和一个起飞前速度调整法则。在飞行中,DART通过轮子反应扭矩来调节俯仰和滚转,该扭矩由基于偏航耦合分析得出的每次飞行滚转锁定控制。在BeamNG.tech的确定性全尺度仿真中,校准的起飞前速度调节器将着陆速度降低了36%,并将目标着陆率从0/30提高至30/30。在相同的陡坡接近下,空中法则在碰撞避免范围内完成了29/30次安全着陆,而反应轮式PD(RW-PD)和时间最优开关(TOBB)则为0/30。在倾斜的跑道上,DART在每个横坡上将中位俯仰误差保持在$2^ heta$或以下,最大基线分离为$eta=12^ heta$。在各种扰动状态下,锁定保持了低扰动进入时的单俯仰分配,并在滚转变得约束时启用双轴控制。所有结果均来自仿真;硬件验证仍待进行。
cs.RO / 6 / 2607.29031

Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving

Auto-JEPA:一种用于端到端自主驾驶的连续意图潜在世界模型
Yang, Jiwei, Chen, Zhengxian, Huang, Chaosheng, Li, Jun
Abstract
Existing autonomous-driving world models typically perform dense prediction of future videos, occupancy states, BEV representations, or agent motion. We argue that planning need not reconstruct the complete future world, but only focus on scene features that affect future ego action. Based on this perspective, we propose Auto-JEPA, an action-oriented latent world model that learns continuous future driving intent through joint-embedding prediction. Given visual observations, egomotion history, and navigation commands, Auto-JEPA predicts an intent embedding aligned with the latent representation of the future ego trajectory. The predicted intent retrieves executable trajectories from a fixed trajectory memory, which are then ranked by a scene-conditioned candidate selection module. Auto-JEPA keeps the visual encoder frozen, requires no explicit perception annotations, and uses no learned trajectory generator. By optimizing only task-specific modules for trajectory representation, intent prediction, and candidate selection, Auto-JEPA achieves 91.3 PDMS on NAVSIM v1 and 89.1 EPDMS on NAVSIM v2. Semantic occlusion experiments show that masking dynamic-agent regions induces an average intent change 2.97x that of equal-area random masking. Moreover, occluding vehicles that affect future driving substantially changes the predicted intent and selected trajectory, whereas both remain essentially unchanged when non-influential vehicles are occluded. These results show that future-intent prediction encourages the model to focus on planning-relevant visual features and supports high-quality planning without dense future-world modeling.
Chinese Translation
现有的自主驾驶世界模型通常对未来视频、占用状态、鸟瞰视图(BEV)表示或代理运动进行密集预测。我们认为,规划不必重建完整的未来世界,而只需关注影响未来自我行为的场景特征。基于这一观点,我们提出了Auto-JEPA,一种面向行动的潜在世界模型,通过联合嵌入预测学习连续的未来驾驶意图。给定视觉观测、自我运动历史和导航指令,Auto-JEPA预测与未来自我轨迹的潜在表示对齐的意图嵌入。预测的意图从固定轨迹记忆中检索可执行轨迹,然后通过场景条件候选选择模块进行排序。Auto-JEPA保持视觉编码器不变,不需要显式的感知注释,也不使用学习的轨迹生成器。通过仅优化任务特定模块以进行轨迹表示、意图预测和候选选择,Auto-JEPA在NAVSIM v1上达到了91.3的PDMS,在NAVSIM v2上达到了89.1的EPDMS。语义遮挡实验表明,遮挡动态代理区域会导致平均意图变化为等面积随机遮挡的2.97倍。此外,遮挡影响未来驾驶的车辆会显著改变预测的意图和选择的轨迹,而当遮挡无影响的车辆时,两者基本保持不变。这些结果表明,未来意图预测促使模型专注于与规划相关的视觉特征,并支持高质量的规划,而无需密集的未来世界建模。
cs.RO / 7 / 2607.29052

Outcome-Guided Distillation: A Teacher-Student Framework to Advance VLM Reasoning in Autonomous Driving

基于结果指导的蒸馏:推动自主驾驶中视觉-语言模型推理的教师-学生框架
Dong, Zeyu, Zhu, Yimin, Wu, Yu, Sun, Yu
Abstract
End-to-end (E2E) autonomous driving aims to learn a direct mapping from visual observations to control actions. However, these E2E models often act as black boxes and struggle with complex scenarios. To address this, recent works incorporate Vision-Language Models (VLMs) to provide explicit reasoning, enhancing both interpretability and driving robustness. These approaches typically rely on pre-generated annotations, which suffer from potentially flawed labels and require costly human labor. In this work, we propose a new framework that integrates structured reasoning and geometric precision through a teacher-student architecture. The teacher model introduces reflective reasoning, where the VLM generates logical explanations and then reflectively refines the reasoning under the supervision of ground-truth action. This enhances zero-shot generalization without intermediate labels. The student model distills the teacher's reasoning capabilities via supervised fine-tuning. We also design a separate waypoint decoder that interprets textual reasoning into continuous trajectories. Our proposed solution integrates two goals: providing explicit reasoning for interpretability and delivering robust and accurate driving performance. It leverages the synergy between these two objectives within a staged inference engine to enhance driving performance and explicitly uses the reasoning to guide driving prediction. Evaluated on Waymo benchmarks, our framework outperforms classical reasoning-based baselines in zero-shot reasoning, waypoint accuracy, and inference efficiency. Our experiments validate this design, demonstrating that the reasoning text makes a significant contribution to driving inference, resulting in around a 24% improvement in performance compared to an identical model that lacks reasoning. Our work advances reasoning-driven autonomous driving toward interpretable and deployable systems.
Chinese Translation
端到端(E2E)自主驾驶旨在学习从视觉观测到控制动作的直接映射。然而,这些E2E模型往往表现为黑箱,难以应对复杂场景。为了解决这个问题,最近的研究将视觉-语言模型(VLMs)纳入其中,以提供明确的推理,从而增强可解释性和驾驶鲁棒性。这些方法通常依赖于预生成的注释,而这些注释可能存在缺陷标签,并且需要昂贵的人力劳动。在本研究中,我们提出了一种新框架,通过教师-学生架构整合结构化推理和几何精度。教师模型引入反思性推理,其中VLM生成逻辑解释,然后在真实动作的监督下反思性地完善推理。这增强了零样本泛化能力,而无需中间标签。学生模型通过监督微调提炼教师的推理能力。我们还设计了一个独立的航点解码器,将文本推理解释为连续轨迹。我们提出的解决方案整合了两个目标:提供明确的推理以增强可解释性,并提供鲁棒且准确的驾驶性能。它利用这两个目标之间的协同作用,在分阶段推理引擎中提升驾驶性能,并明确使用推理来指导驾驶预测。在Waymo基准测试中,我们的框架在零样本推理、航点准确性和推理效率方面优于经典的基于推理的基线。我们的实验验证了这一设计,表明推理文本对驾驶推理做出了显著贡献,与缺乏推理的相同模型相比,性能提升约24%。我们的工作推动了以推理为驱动的自主驾驶朝着可解释和可部署的系统发展。
cs.RO / 8 / 2607.29102

VSTaI: Design and Characterization of Variable-Stiffness Tactile Interfaces Based on 3D-Printed Structured Fabrics

VSTaI:基于3D打印结构织物的可变刚度触觉界面的设计与表征
Mo, Yiting, Mao, Xinyuan, Singh, Jashan Preet, Bello, Fernando
Abstract
Realistic palpation training requires reliable rendering of soft tissue stiffness changes in real time, which is difficult to achieve with conventional simulators. This paper presents a compact, variable-stiffness tactile interface (VSTaI) based on vacuum-induced jamming of 3D-printed structured fabrics. A vacuum-sealed fabric layer is sandwiched between two silicone layers, and stiffness is tuned by regulating internal pressure. Four fabric patterns with different geometric parameters were fabricated and evaluated using force-indentation tests under atmospheric and vacuum conditions. Across the tested pattern and geometry combinations, vacuum jamming increased stiffness significantly, producing an effective modulus from sub-megapascal to megapascal levels. Specifically, one configuration exhibited a stiffness increase of up to 140% under the jammed state. Circular chainmail patterns provided the most spatially uniform distribution of tactile stiffness, while denser geometries reached higher peak stiffness. VSTaI was also shown to exhibit excellent conformability to the underlying geometry. These results support structured-fabric jamming as a practical approach for shape-conformable, tunable-stiffness displays aimed at physical examination training.
Chinese Translation
现实的触诊训练需要实时可靠地呈现软组织刚度变化,这在传统模拟器中难以实现。本文提出了一种基于真空诱导堵塞的紧凑型可变刚度触觉界面(VSTaI),该界面由3D打印的结构织物构成。一个真空密封的织物层夹在两层硅胶之间,通过调节内部压力来调节刚度。制造并评估了四种具有不同几何参数的织物图案,采用了在大气和真空条件下的力-压痕测试。在测试的图案和几何组合中,真空堵塞显著增加了刚度,使有效模量从亚兆帕级别提升至兆帕级别。具体而言,某一配置在堵塞状态下刚度增加了高达140%。圆形链甲图案提供了最均匀的触觉刚度分布,而更密集的几何形状则达到了更高的峰值刚度。VSTaI还表现出对基础几何形状的优良适应性。这些结果支持结构织物堵塞作为一种实用的方法,用于形状可适应、可调刚度的显示器,旨在用于物理检查训练。
cs.RO / 9 / 2607.29169

ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency

ActFovea:通过时空视觉-动作一致性实现VLA策略的运行时安全保障
Yu, Wenda, Wang, Tianshi, Li, Fengling, Li, Xin, Li, Jingjing, Zhu, Lei
Abstract
Vision-language-action (VLA) policies achieve strong performance in robotic manipulation but remain vulnerable to runtime disturbances that break the temporal alignment among visual observations, robot states, and executed actions. We introduce ActFovea, a plug-and-play safeguarding framework that detects and mitigates such failures without retraining or modifying the underlying VLA policy. ActFovea uses robot kinematics, proprioceptive states, and recent actions to construct action-conditioned foveated regions that retain contact-relevant areas and predicted motion corridors while suppressing task-irrelevant visual content. It detects runtime risks by evaluating whether visual motion and observation freshness remain consistent with geometric, proprioceptive, and action transitions. For recoverable disturbances, ActFovea constructs disturbance-specific candidate observations and accepts a recovery only after verifying the resulting action chunk. When stale or replayed observations make reliable recovery impossible, it invokes a bounded safe-failure procedure. In closed-loop evaluations of $\pi_0$ across multiple LIBERO suites, ActFovea increases success under localized visual overlays from 49.3\% to 90.3\%, closing 93.7\% of the gap to clean performance. It further improves success under action drift and visual delay by 7.0 and 9.8 percentage points, respectively, while preserving clean-task performance. Under frozen-observation replay, ActFovea triggers timely safe failure in all trials, with no unprotected failures. These results demonstrate that spatiotemporal visual-action consistency provides an effective basis for runtime safeguarding of VLA policies.
Chinese Translation
视觉-语言-动作(VLA)策略在机器人操作中表现出色,但仍然容易受到运行时干扰的影响,这些干扰会破坏视觉观测、机器人状态和执行动作之间的时间对齐。我们提出了ActFovea,这是一个即插即用的安全保障框架,能够在不重新训练或修改基础VLA策略的情况下检测和缓解此类故障。ActFovea利用机器人运动学、本体状态和最近的动作构建动作条件的注视区域,这些区域保留了与接触相关的区域和预测的运动通道,同时抑制与任务无关的视觉内容。它通过评估视觉运动和观测的新鲜度是否与几何、本体和动作转变保持一致来检测运行时风险。对于可恢复的干扰,ActFovea构建特定于干扰的候选观测,并在验证生成的动作块后接受恢复。当过时或重放的观测使可靠恢复变得不可能时,它会调用有界安全失败程序。在多个LIBERO套件中对$ ext{π}_0$的闭环评估中,ActFovea在局部视觉叠加下的成功率从49.3 ext{%}提高到90.3 ext{%},缩小了93.7 ext{%}与干净性能之间的差距。它在动作漂移和视觉延迟下的成功率分别提高了7.0和9.8个百分点,同时保持了干净任务的性能。在冻结观测重放下,ActFovea在所有试验中及时触发安全失败,没有出现未保护的失败。这些结果表明,时空视觉-动作一致性为VLA策略的运行时安全保障提供了有效的基础。
cs.RO / 10 / 2607.29172

CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning

CLIFT:通过非侵入式闭环迭代微调将Gemini Robotics设备转变为类人专家
Chen, Yuxin, Srikanth, Hari, Jew, Nathan, Wu, Menglin, Wang, Pengcheng, Ren, Junli, Tomizuka, Masayoshi, Xu, Peng, Xie, Jinyu, Tian, Thomas
Abstract
While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain closed-source, limiting downstream users' ability to adapt them to new tasks, embodiments, and deployment settings. Following the LLM community, an emerging access paradigm for closed-weight robot foundation models is the managed supervised fine-tuning (SFT) API, where users submit training data and receive a tuned policy without access to model weights, gradients, or training internals. While such APIs let downstream users leverage powerful proprietary foundation models, they restrict policy improvement to pure imitation, ruling out reinforcement learning and other closed-loop methods that rely on internal training signals. This limitation is particularly acute for agile, contact-rich humanoid manipulation, where the gap between policy outputs and deployed behavior is large due to novel states, action tracking dynamics, latency, and controller-specific failure modes. We study how effective this managed-API regime is for humanoid adaptation, and how closed-loop improvement can be realized within it to push policies toward task mastery. We conduct one of the first empirical studies of managed-API adaptation on a real humanoid, instantiated on Gemini Robotics On-Device (GROD). We find that direct SFT through the API substantially outperforms a leading open-weight VLA trained on the same demonstrations, yet still falls short of deployment-level mastery on agile, contact-rich tasks. To close this gap, we introduce CLIFT: Closed-Loop Iterative Fine-Tuning, which turns deployment-time reward feedback into API-compatible supervised data and enables closed-loop policy improvement without accessing weights, gradients, likelihoods, or losses-pushing GROD to near-perfect success after two flywheel cycles, all without "opening the model box."
Chinese Translation
尽管机器人基础模型的能力日益增强,但最强大的模型通常是在专有数据上训练的,并且保持闭源,这限制了下游用户将其适应于新任务、形态和部署环境的能力。随着大语言模型(LLM)社区的发展,闭重机器人基础模型的一种新兴访问范式是管理监督微调(SFT)API,用户提交训练数据并获得调优后的策略,而无需访问模型权重、梯度或训练内部信息。虽然这样的API使下游用户能够利用强大的专有基础模型,但它们将策略改进限制为纯模仿,排除了依赖内部训练信号的强化学习和其他闭环方法。这一限制在灵活的、接触丰富的类人操控中尤为明显,由于新状态、动作跟踪动态、延迟和特定控制器的失败模式,策略输出与部署行为之间的差距较大。我们研究了这种管理API机制在类人适应中的有效性,以及如何在其中实现闭环改进以推动策略朝向任务掌握。我们在真实的类人机器人上进行了一项关于管理API适应性的首次实证研究,该机器人基于Gemini Robotics On-Device(GROD)实现。我们发现,通过API进行直接SFT显著优于在相同演示上训练的领先开放权重VLA,但在灵活的、接触丰富的任务上仍未达到部署级的掌握。为了缩小这一差距,我们引入了CLIFT:闭环迭代微调,它将部署时的奖励反馈转化为API兼容的监督数据,并在不访问权重、梯度、似然或损失的情况下实现闭环策略改进,使GROD在经过两个飞轮周期后接近完美成功,所有这些都无需“打开模型盒子”。
cs.RO / 11 / 2607.29203

MROPE: A Multi-Robot Safe Cooperative Strategy via combined Predictive Safety Filters and Ellipse-based Constraint Compression

MROPE:通过结合预测安全过滤器和基于椭圆的约束压缩实现的多机器人安全协作策略
Rosetti, Alice, Pichierri, Lorenzo, Cappello, Domenico, Schiano, Fabrizio, Notarstefano, Giuseppe
Abstract
Deploying drone swarms to track a dynamic target in cluttered environments presents severe computational and safety challenges. We propose MROPE, a hierarchical strategy that decouples the cooperative monitoring mission from strict local safety requirements. To overcome the computational bottlenecks typical of dense spaces, our approach dynamically aggregates complex obstacle geometries into a single safe bounding ellipse for each drone. Methodologically, this architecture is realized by combining distributed aggregative optimization for high-level swarm coordination, a decentralized consensus scheme for the safe area computation, and local Predictive Safety Filters (PSF) for real-time collision avoidance. Virtual and real-world experiments validate the framework, demonstrating superior real-time efficiency and scalability compared to centralized approaches.
Chinese Translation
在复杂环境中部署无人机群以追踪动态目标面临着严重的计算和安全挑战。我们提出了MROPE,这是一种分层策略,将协作监控任务与严格的局部安全要求解耦。为了克服密集空间中典型的计算瓶颈,我们的方法动态地将复杂的障碍几何形状聚合为每个无人机的单一安全边界椭圆。在方法论上,该架构通过结合分布式聚合优化以实现高层次的群体协调、去中心化共识方案以进行安全区域计算,以及局部预测安全过滤器(Predictive Safety Filters, PSF)以实现实时碰撞避免来实现。虚拟和现实世界的实验验证了该框架,显示出与中心化方法相比,具有更优的实时效率和可扩展性。
cs.RO / 12 / 2607.29227

Event-Based Upper-Body Humanoid Teleoperation Under Challenging Illumination

基于事件的上半身类人机器人远程操作在挑战性光照下的研究
Fu, Haoyu, Ge, Zhou, Li, Chengze, Sun, Chenzhao, Cui, Ze, Zhou, Wenjing, Qin, Xulei
Abstract
We present a real-time upper-body human-to-humanoid motion imitation framework driven by neuromorphic event-based vision. This work addresses practical perceptual bottlenecks of conventional frame-based RGB sensors, specifically their difficulty in high dynamic range (HDR) scenes and rapid motions due to fixed integration times. By leveraging the Prophesee EVK4 event camera, which operates asynchronously with high temporal resolution and a dynamic range exceeding 120 dB, our system supports stable tracking in conditions where standard vision pipelines degrade, such as severe backlighting and very low light environments below 5 lux. The architecture integrates a low-latency Perception Module, utilizing optimized event accumulation and gravity-aligned inertial fusion, with a causal Motion Module (TWIST) that performs online kinematic retargeting. We validate the system on an embedded NVIDIA Booster T1 platform and an 18-DoF humanoid upper-body setup, demonstrating an end-to-end photon-to-action latency of 23-34 ms and advantages over RGB baselines under our experimental setup. The results indicate a practical trade-off: events can be preferable for fast or poorly lit upper-body teleoperation, whereas well-lit static scenes may favor RGB or hybrid sensing.
Chinese Translation
我们提出了一种实时的上半身人类到类人机器人运动模仿框架,该框架由神经形态事件驱动的视觉系统支持。本研究解决了传统基于帧的RGB传感器在高动态范围(HDR)场景和快速运动下的实际感知瓶颈,特别是由于固定积分时间导致的困难。通过利用Prophesee EVK4事件相机,该相机以高时间分辨率和超过120 dB的动态范围异步工作,我们的系统在标准视觉处理管道退化的条件下(如严重背光和低于5 lux的极低光照环境)支持稳定跟踪。该架构集成了一个低延迟的感知模块,利用优化的事件积累和重力对齐的惯性融合,以及一个因果运动模块(TWIST),该模块执行在线运动重定向。我们在嵌入式NVIDIA Booster T1平台和一个18自由度的类人机器人上半身设置中验证了该系统,展示了23-34毫秒的端到端光子到动作延迟,并在我们的实验设置中显示出相较于RGB基线的优势。结果表明了一种实际的权衡:在快速或光线不足的上半身远程操作中,事件可能更具优势,而在光线良好的静态场景中则可能更倾向于RGB或混合传感。
cs.RO / 13 / 2607.29231

TacPrint: A Wearable Fingertip Tactile Sensor for Human-to-Robot Contact Reproduction

TacPrint:一种可穿戴的指尖触觉传感器用于人机接触再现
Liu, Yongxi, Zhang, Chaofan, Zhang, Xingyu, Bao, Xiangyin, Zhang, Boyue, Cui, Shaowei, Wang, Shuo
Abstract
Human-centric data collection is emerging as a significant paradigm for robot skill acquisition, but seamlessly integrating low-cost, scalable tactile sensing systems that capture fine-grained fingertip interactions without compromising natural operation remains a key challenge. This reduces the reliability of human-to-robot transfer in contact-rich tasks. In this work, we present TacPrint, a wearable fingertip tactile sensor, where protrusions on the inner surface of the silicone skin are aligned one-to-one with 24 capacitive taxels to enable localized capacitive responses. A real-to-sim-to-real pipeline estimates a 35 $\times$ 26 contact-depth map from 24-channel capacitive signals. Against simulation-generated labels, the model achieved a contact-region RMSE of 0.223 $\pm$ 0.161 mm, a weighted-centroid error of 1.213 $\pm$ 2.379 pixels, and an IoU of 0.829 $\pm$ 0.169. With measured capacitive inputs, the network-predicted depth evaluated at the guide-calibrated contact center showed a mean absolute error of 0.085 $\pm$ 0.057 mm across all 40 controlled trials, while the mean contact-position error was 0.250 $\pm$ 0.208 mm across the 37 trials whose reference contact regions were not truncated by the sensing boundary. In human-to-robot replay, tactile-guided compensation increased grasping and wiping success rates from 0% to 91.67% and 90%, respectively. In closed-loop grasping, dense-depth feedback achieved success rates of 87.5% over all tested positions and 85% under edge-contact conditions, compared with 67.5% and 45% for raw-taxel feedback.
Chinese Translation
以人为中心的数据收集正在成为机器人技能获取的重要范式,但无缝集成低成本、可扩展的触觉传感系统,以捕捉细粒度的指尖交互而不妨碍自然操作,仍然是一个关键挑战。这降低了在接触丰富任务中人机转移的可靠性。在本研究中,我们提出了TacPrint,一种可穿戴的指尖触觉传感器,其硅胶外皮内表面的突起与24个电容式传感单元(taxels)一一对应,以实现局部电容响应。一个真实-模拟-真实的管道从24通道电容信号中估计出35 × 26的接触深度图。与模拟生成的标签相比,该模型在接触区域的均方根误差(RMSE)为0.223 ± 0.161 mm,加权质心误差为1.213 ± 2.379像素,交并比(IoU)为0.829 ± 0.169。在测量的电容输入下,网络预测的深度在引导校准的接触中心处显示出在所有40次控制试验中的平均绝对误差为0.085 ± 0.057 mm,而在37次参考接触区域未被传感边界截断的试验中的平均接触位置误差为0.250 ± 0.208 mm。在人机重放中,触觉引导补偿将抓取和擦拭的成功率分别从0%提高到91.67%和90%。在闭环抓取中,密集深度反馈在所有测试位置的成功率达到了87.5%,而在边缘接触条件下为85%,相比之下,原始taxel反馈的成功率仅为67.5%和45%。
cs.RO / 14 / 2607.29235

FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution

FBFM:一种无训练的异步反馈机制用于世界行动模型执行中的流匹配
Li, Peize, Zhang, Ruimeng, Zhang, Ru, Huang, Cong, Chen, Kai, Zhang, Shanghang
Abstract
Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliability demands repeated re-grounding in real observations--not recursive rollout. Existing WAMs address this by refreshing history or KV cache with ground-truth data between chunks. However, such chunk-wise feedback operates at a coarse temporal granularity and thus fails to correct prediction errors at the individual time-step level. To address this, we propose Feedback Flow Matching (FBFM), a training-free inference mechanism that pushes re-grounding inside the actively generated chunk. During flow matching, FBFM applies a masked pseudoinverse correction to the conditional velocity field: it leverages the preceding action chunk to guide generation of the next action chunk, and uses the image observed after executing that preceding chunk to guide the next frame prediction. This cross-chunk pairing--where feedback from one chunk arrives in time to shape the next--creates an asynchronous loop that corrects errors without waiting for chunk boundaries. Being training-free, the mechanism improves responsiveness to unexpected events and suppresses drift in long-horizon tasks. We evaluate FBFM on both a joint-generation WAM (DreamZero) and a stage-wise WAM (LingBot-VA). On selected LIBERO and RoboTwin2.0 tasks, it improves success rates by over 5% in favorable settings, and real-world robot observation-prediction diagnostics show notably better tracking. We argue that FBFM offers a new paradigm for fine-grained online correction, bridging open-loop flow generation with closed-loop real-world dynamics.
Chinese Translation
尽管世界行动模型(WAMs)通过在行动前预测视觉演变来增强长时间范围的机器人控制,但长时间范围的可靠性要求在真实观察中进行重复的重新定位,而不是递归展开。现有的WAMs通过在数据块之间用真实数据刷新历史或KV缓存来解决这个问题。然而,这种按块反馈在时间粒度上较为粗糙,因此无法在单个时间步级别上纠正预测错误。为了解决这个问题,我们提出了反馈流匹配(FBFM),这是一种无训练的推理机制,能够在主动生成的数据块内部进行重新定位。在流匹配过程中,FBFM对条件速度场应用掩蔽伪逆修正:它利用前一个动作块来指导下一个动作块的生成,并使用在执行前一个块后观察到的图像来指导下一个帧的预测。这种跨块配对——一个块的反馈及时到达以影响下一个块——创建了一个异步循环,能够在不等待块边界的情况下纠正错误。作为一种无训练的机制,它提高了对意外事件的响应能力,并抑制了长时间任务中的漂移。我们在一个联合生成的WAM(DreamZero)和一个阶段性WAM(LingBot-VA)上评估了FBFM。在选定的LIBERO和RoboTwin2.0任务中,它在有利条件下成功率提高超过5%,而现实世界的机器人观察-预测诊断显示出明显更好的跟踪效果。我们认为FBFM为细粒度在线修正提供了一种新范式,架起了开放循环流生成与闭环真实世界动态之间的桥梁。
cs.RO / 15 / 2607.29271

MDIR: A Task-Manifold Impedance Retargeting Method for Contact-Rich Teleoperation

MDIR:一种用于接触丰富的远程操作的任务流形阻抗重定向方法
Jiahao, Liu, Kawaharazuka, Kento, Makabe, Tasuku, Okada, Kei
Abstract
Fixed Cartesian impedance makes contact-rich teleoperation demonstrations practical, but gains that secure progress and contact support also determine impact and force variability. We study single-demonstration controller-to-controller impedance retargeting. Given one fixed Cartesian impedance command sequence {K0, D0, xcmd}, Manifold-Decomposed Impedance Retargeting (MDIR) deterministically reparameterizes the recorded controller into an executable task-channel variable-impedance command. MDIR targets this local retargeting problem by preserving projected task-channel responses near the demonstrated trajectory. It represents the source response in operational work, exertion, and support channels with a passive residual complement under a control-chain metric, computes an executable Cartesian-to-Manifold Retargeting (C2M) baseline, and applies Manifold-Constrained Parameter Optimization (MPO) to select a feasible representative with lower wrist-force peaks, impulse, force variability, and nominal controller power. Across planar wiping, pick-and-place, and pushing on a Franka Panda, the full MDIR controller passes Task Check in all 15 closed-loop executions and reduces all four aggressiveness metrics relative to the fixed-impedance demonstrations.
Chinese Translation
固定的笛卡尔阻抗使得接触丰富的远程操作演示变得实用,但确保进展和接触支持的增益也决定了冲击和力的变异性。我们研究了单次演示控制器到控制器的阻抗重定向。给定一个固定的笛卡尔阻抗命令序列 {K0, D0, xcmd},流形分解阻抗重定向(MDIR)确定性地将记录的控制器重新参数化为可执行的任务通道可变阻抗命令。MDIR通过在演示轨迹附近保持投影任务通道响应来针对这一局部重定向问题。它在操作工作、施力和支持通道中用被动残余补充表示源响应,并在控制链度量下计算可执行的笛卡尔到流形重定向(C2M)基线,应用流形约束参数优化(MPO)选择一个可行的代表,以降低手腕力峰值、冲击、力的变异性和名义控制器功率。在平面擦拭、抓取放置和在Franka Panda上推动的实验中,完整的MDIR控制器在所有15次闭环执行中通过了任务检查,并相对于固定阻抗演示降低了所有四个激进性指标。
cs.RO / 16 / 2607.29285

TRACT: Temporally Routed Action Chunks with Chronological Phase Authority for Contact-Rich Manipulation

TRACT:具有时间路由的动作块及其在接触丰富操作中的时间阶段权威
Liu, Jiahao, Kawaharazuka, Kento, Makabe, Tasuku, Okada, Kei
Abstract
Action chunking shortens the effective decision horizon of robot imitation learning by predicting multiple future actions, while conventional phase conditioning describes the current control instant. When a predicted horizon crosses a procedural boundary, assigning the current phase to the entire chunk creates a structural temporal mismatch. We present TRACT, which factorizes phase-structured action chunking into an accepted current phase and a single CURRENT-to-NEXT boundary inside the future horizon. A task-local graph constrains chronological phase authority, and a cumulative boundary distribution monotonically routes future queries through phase-specific query and action paths. For contact execution, a causal response-deficit integrator compares policy intent with ACK-eligible subsequent motion, accumulates arm compensation when directional response is suppressed, and decays after confirmed recovery. Across six real-robot variants with ten trials each, full TRACT achieves 10/10 full-sequence success, 99.00 [88.75, 100.00]% median [min, max] wipe completion, zero observed phase ambiguity, and zero stalls. Under the current complete method package and evaluation setting, the routed representation obtains better observed task results than the flat package (6/10 vs. 3/10 success; 77.08% vs. 8.03% median wipe completion). Chronological authority reduces observed phase ambiguity from 8/10 to 0/10, and response integration reduces stalls from 4/10 to 0/10. The package comparison does not isolate routing from other generator-package differences.
Chinese Translation
动作块化通过预测多个未来动作来缩短机器人模仿学习的有效决策视野,而传统的阶段条件描述当前控制时刻。当预测的视野跨越程序边界时,将当前阶段分配给整个块会造成结构上的时间不匹配。我们提出了TRACT,它将阶段结构化的动作块化分解为一个被接受的当前阶段和一个位于未来视野内的单一CURRENT-to-NEXT边界。任务局部图约束了时间阶段权威,累积边界分布单调地通过特定阶段的查询和动作路径路由未来查询。在接触执行中,因果响应缺失整合器将策略意图与符合ACK的后续动作进行比较,当方向响应被抑制时累积手臂补偿,并在确认恢复后衰减。在六个真实机器人变体中,每个变体进行十次试验,完整的TRACT实现了10/10的完整序列成功率,99.00 [88.75, 100.00]% 的中位数 [最小值, 最大值] 擦拭完成率,零观察到的阶段模糊性,以及零停滞。在当前完整的方法包和评估设置下,路由表示获得的观察任务结果优于平面包(成功率为6/10对比3/10;中位数擦拭完成率为77.08%对比8.03%)。时间权威将观察到的阶段模糊性从8/10降低到0/10,响应整合将停滞从4/10降低到0/10。包的比较没有将路由与其他生成器包的差异隔离开。
cs.RO / 17 / 2607.29302

BWM: A Low-Cost High-Fidelity World Simulator for Robot Learning

BWM:一种低成本高保真度的机器人学习世界模拟器
BWM Team
Abstract
Reliable robot learning requires a world simulator that can predict action consequences before execution on physical hardware, including risky and failure-prone outcomes. Existing physics simulators require substantial asset construction and calibration and still face a sim-to-real gap, while video generators often lack precise control over their responses to fine-grained robot actions. In this paper, we present the Boundless World Model (BWM), an open-source, low-cost, high-fidelity world simulator for robot manipulation. BWM is an action-conditioned world model that combines initial-environment guidance, dynamic visual history, and temporally aligned robot-action conditioning for stateful autoregressive prediction of future observations. We construct action-aligned training clips through trajectory replay, overlapping clip sampling, and initial-observation enhancement. BWM serves as a data engine that augments imitation-learning data with action-aligned rollouts, and as a policy evaluator for closed-loop assessment, risk anticipation, and policy ranking. Experiments on the WorldArena benchmark and physical robots demonstrate improved simulator fidelity and functional utility across the data-engine and policy-evaluator settings. BWM ranks first overall in the WorldArena Challenge across Track 1 and its two Track 2 applications. We release the BWM open-source ecosystem, including model checkpoints, training and inference code, and interfaces for data generation and policy evaluation.
Chinese Translation
可靠的机器人学习需要一个能够在物理硬件执行之前预测行动后果的世界模拟器,包括风险和故障易发的结果。现有的物理模拟器需要大量的资产构建和校准,并且仍然面临模拟与现实之间的差距,而视频生成器往往缺乏对细粒度机器人动作响应的精确控制。本文提出了无界世界模型(Boundless World Model,BWM),这是一个开源、低成本、高保真度的机器人操作世界模拟器。BWM是一个基于动作条件的世界模型,结合了初始环境指导、动态视觉历史和时间对齐的机器人动作条件,以实现未来观察的状态自回归预测。我们通过轨迹重放、重叠剪辑采样和初始观察增强构建了动作对齐的训练片段。BWM作为一个数据引擎,通过动作对齐的回放增强模仿学习数据,同时作为闭环评估、风险预判和策略排名的策略评估器。在WorldArena基准和物理机器人上的实验表明,BWM在数据引擎和策略评估器设置中提高了模拟器的保真度和功能效用。在WorldArena挑战赛中,BWM在第一轨道和其两个第二轨道应用中整体排名第一。我们发布了BWM开源生态系统,包括模型检查点、训练和推理代码,以及数据生成和策略评估的接口。
cs.RO / 18 / 2607.29374

SAGP: Semantic Affordance-Guided Grasp Planning via Coarse-Zone VLM Reasoning

SAGP:基于语义可供性引导的抓取规划通过粗区VLM推理
Din, Muhayy Ud, Hussain, Irfan
Abstract
Geometry-based grasp planners ensure physically valid grasps but ignore functional semantics, often generating grasps that are antipodal and collision-free yet practically inappropriate, for example, gripping a mug by its rim, a knife by the blade, or a bottle near its cap. These inconsistencies cause the downstream task to fail even when traditional grasp metrics are met. Existing vision-language model (VLM) approaches either depend on fine-grained, category-specific part segmentation or attempt to directly infer grasp poses, with the latter prone to spatial hallucinations. As a result, no practical, training-free framework has yet been proposed that robustly links high-level semantic reasoning to geometric grasp planning. We introduce Semantic Affordance-Guided Grasp Planning (SAGP), a training-free pipeline built on a coarse-zone abstraction layer. The method first partitions the object point cloud into spatial regions (top, middle, bottom, lateral sides, and protrusions) by applying PCA-based alignment followed by distance-driven DBSCAN clustering, entirely bypassing learned segmentation. A pre-trained VLM then assesses the grasp quality of each region through a structured zero-shot query, and the resulting zone-wise scores are fused with geometric, reachability, and task-alignment signals to re-rank antipodal grasp candidates. Experiments on YCB objects in PyBullet with a Franka Panda robot show that SAGP preserves the high success rate of geometry-only planning while substantially improving the functional appropriateness of selected grasps, particularly on asymmetric, handle-bearing objects where geometry alone is uninformative. The introduced coarse-zone abstraction offers an effective, training-free bridge between VLM-based reasoning and geometric grasp planning, without the need for fine-grained part segmentation.
Chinese Translation
基于几何的抓取规划器确保物理有效的抓取,但忽视功能语义,常常生成反向和无碰撞的抓取,但在实际应用中不合适,例如,抓取杯子的边缘、刀刃或瓶子的瓶口。这些不一致性导致下游任务失败,即使满足传统抓取指标。现有的视觉-语言模型(VLM)方法要么依赖于细粒度、类别特定的部件分割,要么试图直接推断抓取姿态,后者容易出现空间幻觉。因此,尚未提出任何实用的、无训练的框架,能够稳健地将高层次的语义推理与几何抓取规划联系起来。我们提出了基于语义可供性引导的抓取规划(SAGP),这是一个建立在粗区抽象层上的无训练管道。该方法首先通过应用基于主成分分析(PCA)的对齐和基于距离的DBSCAN聚类,将物体点云划分为空间区域(顶部、中部、底部、侧面和突出部分),完全绕过学习的分割。然后,预训练的VLM通过结构化的零样本查询评估每个区域的抓取质量,得到的区域评分与几何、可达性和任务对齐信号融合,以重新排序反向抓取候选。针对PyBullet中YCB物体与Franka Panda机器人进行的实验表明,SAGP在保持几何规划高成功率的同时,显著提高了所选抓取的功能适宜性,特别是在几何信息不足的非对称、带把手的物体上。引入的粗区抽象为基于VLM的推理与几何抓取规划之间提供了有效的、无训练的桥梁,而无需细粒度的部件分割。
cs.RO / 19 / 2607.29393

AquaJEPA: Action-Conditioned Multimodal Predictive Representations for Underwater Robot Dynamics

AquaJEPA:用于水下机器人动力学的动作条件多模态预测表示
Gazzaev, Alan-Barsag, Gavrilov, Alexey, Muravyov, Sergey
Abstract
Underwater robots combine complementary sensors whose reliability changes abruptly with water visibility, viewpoint, and vehicle motion. We introduce AquaJEPA, an action-conditioned joint-embedding predictive model that fuses an RGB camera, forward-looking sonar, and proprioception with explicit sensor validity. It predicts a future latent target conditioned on eight-thruster commands and supplies velocity and sonar-profile predictions to a shared receding-horizon planner. We study the method in Stonefish against reactive, state-only, ordinary multimodal, supervised dynamics, and recurrent world-model baselines. We further isolate the EMA target, action margin, masks, and modality dropout. A preregistered 120-environment replication comprises five independent replicates of a grid crossing three unseen obstacle maps, four water-visibility coefficients, and nominal versus shifted dynamics, while intermittently removing DVL observations. In 120 fresh paired environments with scheduled DVL loss, AquaJEPA reaches 74 goals, versus 68 for both state-only and the recurrent world model, and attains the lowest mean final error (0.906 m). Paired final-error reductions relative to ordinary multimodal prediction, supervised dynamics, and the recurrent world model are 0.273 m (95% CI: 0.190-0.356), 0.364 m (0.260-0.468), and 0.106 m (0.025-0.187), respectively. AquaJEPA therefore achieves the best aggregate closed-loop performance and significantly outperforms three action-conditioned predictive baselines in paired final error; its advantage over state-only remains statistically unresolved.
Chinese Translation
水下机器人结合了互补的传感器,其可靠性会随着水的能见度、视角和车辆运动而急剧变化。我们提出了AquaJEPA,这是一种动作条件的联合嵌入预测模型,融合了RGB相机、前视声纳和自我感知,并明确考虑传感器的有效性。该模型基于八个推进器命令预测未来的潜在目标,并向共享的递归规划器提供速度和声纳轮廓的预测。我们在Stonefish环境中对该方法进行了研究,与反应式、仅状态、普通多模态、监督动力学和递归世界模型基线进行了比较。我们进一步分离了EMA目标、动作边际、掩码和模态丢失。在一个预注册的120环境复制实验中,包含了三种未见障碍地图的网格穿越、四个水能见度系数以及名义与偏移动力学,同时间歇性地移除DVL观测。在120个新配对环境中,AquaJEPA达成了74个目标,而仅状态和递归世界模型均为68个,并且获得了最低的平均最终误差(0.906米)。相对于普通多模态预测、监督动力学和递归世界模型,配对最终误差的减少分别为0.273米(95% CI: 0.190-0.356)、0.364米(0.260-0.468)和0.106米(0.025-0.187)。因此,AquaJEPA实现了最佳的整体闭环性能,并在配对最终误差上显著优于三种动作条件预测基线;其相对于仅状态的优势在统计上仍未得到解决。
cs.RO / 20 / 2607.29464

Automated Straight-line Sewing of Stretchable Fabrics with Different Lengths

不同长度可拉伸织物的自动直线缝合
Jin, Bingchen, Kobayashi, Akinari, Bhattacharya, Dipankar, Seino, Akira, Tokuda, Fuyuki, Tien, Norman Chihnan, Kosuge, Kazuhiro
Abstract
Different Length Alignment Sewing (DLAS), which involves stretching the shorter fabric to match the longer one and sewing them together in a straight line, is a challenging task that needs to satisfy several requirements when automating the sewing process. To address the challenges, this research proposes a novel robotic sewing system, Different Length Robotic Sewing System (DLRoSS), which consists of a roller type end-effector, attached to a 6-DoF manipulator. The end-effector composed of active shorter and longer fabric rollers, and a passive press-roller attached to the shorter-fabric roller. Assuming that one end of the two fabric layers are initially positioned under the sewing machine's presser foot, the system automates DLAS by operating in four distinct phases. (P1) Fabric wrapping: Individual fabric layers are picked, held, and wrapped from the other end onto the feed rollers. (P2) Sewing: During the sewing, the shorter fabric is stretched and aligned with the longer fabric in real-time using roller velocity control based on the sewing speed and apriori known length ratio. (P3) Sewing completion: In the final sewing round on the fabric rollers, the press roller is engaged to prevent the stretched fabric from slipping off due to internal tension. (P4) Sewing fabric release: At the end of sewing, the fabric edge moves past the press roller, and the fabric releases from the rollers. Experimental results demonstrate that DLRoSS achieves consistent, high-quality sewing of stretchable fabrics of different materials and lengths.
Chinese Translation
不同长度对齐缝合(DLAS)是一项具有挑战性的任务,它涉及将较短的织物拉伸以匹配较长的织物,并将它们以直线缝合在一起。在自动化缝合过程中,需要满足多个要求。为了解决这些挑战,本研究提出了一种新型机器人缝合系统——不同长度机器人缝合系统(DLRoSS),该系统由一个滚筒式末端执行器和一个6自由度(DoF)操纵臂组成。末端执行器由主动的短织物和长织物滚筒以及一个附加在短织物滚筒上的被动压辊组成。假设两层织物的一端最初位于缝纫机的压脚下,该系统通过四个不同的阶段自动执行DLAS。 (P1) 织物包裹:分别拾取、保持并将各个织物层从另一端包裹到进料滚筒上。 (P2) 缝合:在缝合过程中,较短的织物通过基于缝合速度和已知长度比的滚筒速度控制实时拉伸并与较长的织物对齐。 (P3) 缝合完成:在织物滚筒上的最后缝合轮次中,压辊被启用,以防止由于内部张力导致的拉伸织物滑落。 (P4) 织物释放:缝合结束时,织物边缘经过压辊,织物从滚筒上释放。实验结果表明,DLRoSS能够一致地高质量缝合不同材料和长度的可拉伸织物。
cs.RO / 21 / 2607.29482

Temporal Policy: History-Initialized Action Generation for Robotic Learning from Demonstration

时间策略:基于历史初始化的机器人示范学习中的动作生成
Miller, Dylan, Jagersand, Martin
Abstract
By relying on independent couplings from uninformative Gaussian priors, standard diffusion and flow matching models are forced to learn complex, high-cost vector fields to reach the physical action space. Generative models excel at capturing multimodal behaviors for robotic Learning from Demonstration (LfD), but often suffer from high inference cost. This paper introduces Temporal Policy, a generative framework based on stochastic interpolants that formulates action generation as a temporally coupled transport problem. By initializing the generative flow at the robot's recent history, we explicitly couple past states to future action sequences. This data-dependent coupling reduces transport cost and produces straight vector fields. We validate Temporal Policy across visuomotor simulation benchmarks and on a physical Barrett WAM 2x 7DoF teleoperation platform. Our approach reduces transport costs by nearly an order of magnitude compared to noise-initialized baselines, achieving a 19.1 ms inference latency on a single NVIDIA RTX 4080. Crucially, these geometric and computational efficiencies are achieved while matching the success rates of state-of-the-art baselines. This simplified transport geometry bypasses the computational bottleneck of independent Gaussian priors, helping enable high-frequency, closed-loop control. The code is publicly available at https://github.com/dmiller12/TemporalPolicy.
Chinese Translation
标准扩散和流匹配模型依赖于来自无信息高斯先验的独立耦合,迫使其学习复杂且高成本的向量场以达到物理动作空间。生成模型在捕捉机器人示范学习(LfD)的多模态行为方面表现出色,但通常面临高推理成本的问题。本文提出了时间策略(Temporal Policy),这是一个基于随机插值的生成框架,将动作生成形式化为一个时间耦合的传输问题。通过在机器人的近期历史上初始化生成流,我们明确地将过去状态与未来动作序列耦合。这种依赖数据的耦合减少了传输成本,并产生了直线向量场。我们在视觉运动仿真基准测试和物理Barrett WAM 2x 7自由度遥操作平台上验证了时间策略。与噪声初始化的基线相比,我们的方法将传输成本减少了近一个数量级,在单个NVIDIA RTX 4080上实现了19.1毫秒的推理延迟。重要的是,这些几何和计算效率是在与最先进基线的成功率相匹配的情况下实现的。这种简化的传输几何结构绕过了独立高斯先验的计算瓶颈,有助于实现高频闭环控制。代码已公开发布在 https://github.com/dmiller12/TemporalPolicy。
cs.RO / 22 / 2607.29500

Tri-Space Operational Control of Redundant Multilink and Hybrid Cable-Driven Parallel Robots Using an Iterative-Learning based Reactive Approach

基于迭代学习的反应控制的冗余多链路和混合电缆驱动并联机器人的三空间操作控制
Bhattacharya, Dipankar, Chan, Yin Pok, Shang, Siqi, Chan, Yuen Shan, Tan, Ying, Lau, Darwin
Abstract
Cable-Driven Parallel Robots (CDPRs) are a type of parallel mechanism in which cables are used as actuators. Due to the two levels of redundancy and numerous constraints within the CDPR actuation, joint and operational spaces (together known as the tri-space), tracking a given trajectory in the operational space while satisfying constraints in tri-space simultaneously is challenging. To the best of the authors' knowledge, there does not exist any tri-space control framework, which is robust, effective, and directly applicable to several architectures of redundantly actuated CDPRs. This paper proposes a tri-space control framework that combines Reactive Control (RC) and Iterative-Learning Control (ILC) to perform repetitive tasks in the operational space. The framework allows the tracking of operational space trajectories online with feasible cable forces, while avoiding undesirable situations such as cable-link interference, joint interference, and loss of manipulability. On the other hand, by finding an optimal parameter in the null space using a novel parameterization of a null space vector, the performance can be improved through ILC when the task is repeatedly executed. Simulation and hardware results on various Multilink Cable-Driven Robot (MCDRs) and Hybrid Cable-Driven Robots (HCDRs) show that the proposed tri-space control framework can be conveniently and effectively applied to the real-time control of different CDPRs.
Chinese Translation
电缆驱动并联机器人(CDPRs)是一种使用电缆作为执行器的并联机制。由于CDPR驱动中的两级冗余和众多约束,联合空间和操作空间(统称为三空间)在满足三空间约束的同时追踪给定的操作空间轨迹是具有挑战性的。根据作者的最佳知识,目前尚不存在任何稳健、有效且可直接应用于多种冗余驱动CDPR架构的三空间控制框架。本文提出了一种三空间控制框架,该框架结合了反应控制(Reactive Control, RC)和迭代学习控制(Iterative-Learning Control, ILC),以在操作空间中执行重复任务。该框架允许在线追踪操作空间轨迹,同时以可行的电缆力避免不良情况,如电缆链干扰、关节干扰和可操作性丧失。另一方面,通过使用新颖的零空间向量参数化方法在零空间中找到最佳参数,当任务重复执行时,性能可以通过ILC得到改善。在各种多链路电缆驱动机器人(Multilink Cable-Driven Robots, MCDRs)和混合电缆驱动机器人(Hybrid Cable-Driven Robots, HCDRs)上的仿真和硬件结果表明,所提出的三空间控制框架可以方便有效地应用于不同CDPR的实时控制。
cs.RO / 23 / 2607.29513

Homotopy-Aware Corridor Generation without Predefined Reference Paths

无预定义参考路径的同伦感知走廊生成
Dong, Haoze, Li, Minghan, Guo, Meng, Li, Zhongkui
Abstract
Generating safe corridors is essential for collision-free robotic motion planning, yet most existing methods rely on predefined reference paths, which bias corridor geometry and implicitly limit the homotopy classes that can be explored. We propose a reference-path-free corridor generation framework on graphs of convex sets (GCS) that constructs corridors directly as sequences of convex sets, allowing corridor structure to emerge from the free-space representation rather than from a guiding path. To reason about similarity among corridors, we extend visibility-based deformation from paths to convex-set sequences, enabling the fusion of topologically redundant corridors while preserving distinct alternatives. To overcome the limited adaptability of existing GCS methods based on static global decompositions, we further develop an adaptive multi-scale GCS, in which a sampling-based fine-scale graph supports localized updates and a visibility-based coarse-scale graph enables compact global exploration. The two levels maintain topological consistency, allowing incremental updates without full graph reconstruction under environmental uncertainty. Numerical experiments characterize GCS construction, corridor generation, homotopy-aware exploration, and local updates, showing efficient graph construction, stable trajectory-level performance, and shorter-duration homotopy-aware trajectories than existing baselines. Hardware experiments on ground and aerial robots, including deployment with onboard localization, further validate the framework under translated and previously unknown obstacles.
Chinese Translation
生成安全走廊对于无碰撞的机器人运动规划至关重要,然而大多数现有方法依赖于预定义的参考路径,这会偏向走廊几何形状并隐含限制可探索的同伦类。我们提出了一种基于凸集图(GCS)的无参考路径走廊生成框架,该框架直接将走廊构建为凸集序列,使走廊结构能够从自由空间表示中涌现,而不是依赖于引导路径。为了推理走廊之间的相似性,我们将基于可见性的变形方法从路径扩展到凸集序列,能够在保留不同选择的同时融合拓扑冗余的走廊。为了解决现有基于静态全局分解的GCS方法适应性有限的问题,我们进一步开发了一种自适应多尺度GCS,其中基于采样的细尺度图支持局部更新,而基于可见性的粗尺度图则实现紧凑的全局探索。这两个层次保持拓扑一致性,允许在环境不确定性下进行增量更新而无需完全重建图。数值实验表征了GCS构建、走廊生成、同伦感知探索和局部更新,显示出高效的图构建、稳定的轨迹级性能以及比现有基线更短的同伦感知轨迹。在地面和空中机器人上的硬件实验,包括与机载定位的部署,进一步验证了该框架在已转移和未知障碍物下的有效性。
cs.RO / 24 / 2607.29517

STAGE: STyle-controllable Action GEneration for personalized autonomous driving

STAGE:可控风格的个性化自动驾驶动作生成
Liu, Zihao, Liu, Xing, Zhang, Yizhai, Huang, Panfeng
Abstract
Driving style refers to the behavioral preferences that drivers maintain during driving, shaped by their diverse experiences, habits, and needs, and is typically reflected in varying levels of aggressiveness. If humans choose to use autonomous driving systems, they would expect the driving style of the systems to closely resemble their own habit. However, this is challenging for current industrial autonomous driving systems. To address this, we developed a style controllable action generation method, STAGE, for driving tasks. Its training process is based on imitation learning, incorporating both style value and latent value action modality encoding. Preference learning is then used to identify the user's driving style as a continuous, monotonic style value. And to reduce the cost of human involvement in the preference training process, we also developed a set of rules to compare driving style in data pairs. Then, during inference, the user inputs the style value to control the generated action patterns, dynamically meeting the user's expectations. Using the STAGE method, we verified that the style-controlled action generation results in several typical road scenarios significantly align with human expectations. Furthermore, through comparisons between the STAGE method and various other approaches, we reveal the unique functionalities of STAGE, including its style controllability, style continuity, driving style alignment capability and driving safety. The code for this work is available at: https://github.com/CarlDegio/STAGE
Chinese Translation
驾驶风格是指驾驶者在驾驶过程中保持的行为偏好,这些偏好受到他们多样的经验、习惯和需求的影响,通常表现为不同程度的攻击性。如果人们选择使用自动驾驶系统,他们希望系统的驾驶风格能够与自己的习惯紧密相似。然而,这对当前的工业自动驾驶系统来说是一个挑战。为了解决这个问题,我们开发了一种可控风格的动作生成方法STAGE,适用于驾驶任务。其训练过程基于模仿学习,结合了风格值和潜在值的动作模态编码。接着,采用偏好学习来识别用户的驾驶风格,作为一个连续的单调风格值。为了减少人类参与偏好训练过程的成本,我们还制定了一套规则来比较数据对中的驾驶风格。然后,在推理过程中,用户输入风格值以控制生成的动作模式,动态满足用户的期望。使用STAGE方法,我们验证了风格控制的动作生成在多个典型道路场景中显著符合人类的期望。此外,通过将STAGE方法与其他多种方法进行比较,我们揭示了STAGE的独特功能,包括其风格可控性、风格连续性、驾驶风格对齐能力和驾驶安全性。本研究的代码可在以下网址获取:https://github.com/CarlDegio/STAGE
cs.RO / 25 / 2607.29567

TransGraspNet: Physically and Geometrically Consistent Manipulation of Transparent Labware

TransGraspNet:透明实验室器皿的物理和几何一致性操作
Hu, Hailing, Zhu, Mingyi, An, Yiquan, Tian, Yifei, Zuo, Tianyou, Zhou, Lifeng
Abstract
Manipulating transparent laboratory glassware that contains liquid is inherently safety-critical: even small geometric errors can cause unstable grasps and hazardous spillage. Although recent progress has been made in transparent object perception and robotic grasping, most existing systems optimize detection, depth reconstruction, and grasp planning independently, which leads to cross-stage inconsistency imperfect boundaries induce depth bleeding, distorted surfaces corrupt normal estimation, and task agnostic grasp scoring yields tilted or off-center grasps that fail under dynamic motion. In this paper, we propose TransGraspNet, a geometry physics consistent framework that explicitly enforces consistency from perception to execution through three coupled principles: boundary consistency to produce structurally reliable object contours as downstream priors, surface consistency to preserve geometric fidelity and surface normal accuracy during depth reconstruction, and physics consistency to refine grasp selection with centroid alignment and wrench-space stability for upright and dynamically robust manipulation. We evaluate TransGraspNet on public benchmarks, a dedicated transparent glassware dataset, and a real robotic platform. The results show improved boundary quality and surface normal fidelity, and demonstrate strong task-level performance in cluttered transparent scenes. Most importantly, the proposed system achieves reliable real-world operation, including high grasp success rates in clutter and zero spillage during high speed liquid transport, highlighting the effectiveness of our method.
Chinese Translation
操作含有液体的透明实验室玻璃器皿本质上是安全关键的:即使是小的几何误差也会导致不稳定的抓取和危险的溢出。尽管在透明物体感知和机器人抓取方面取得了近期进展,但大多数现有系统独立优化检测、深度重建和抓取规划,这导致了跨阶段的不一致性:不完美的边界会引起深度渗漏,扭曲的表面会破坏法线估计,而与任务无关的抓取评分则会导致倾斜或偏心的抓取在动态运动下失败。本文提出了TransGraspNet,一个几何物理一致性框架,通过三个耦合原则明确强制从感知到执行的一致性:边界一致性以生成结构上可靠的物体轮廓作为下游先验,表面一致性以在深度重建过程中保持几何保真度和表面法线准确性,以及物理一致性以通过质心对齐和扭矩空间稳定性来优化抓取选择,以实现直立和动态稳健的操作。我们在公共基准、专用透明玻璃器皿数据集和真实机器人平台上评估了TransGraspNet。结果显示了边界质量和表面法线保真度的改善,并在杂乱的透明场景中展示了强大的任务级性能。最重要的是,所提系统实现了可靠的现实世界操作,包括在杂乱环境中高抓取成功率和高速液体运输过程中的零溢出,突显了我们方法的有效性。
cs.RO / 26 / 2607.29569

Safe Vision Language Action Models via Barrier Enhanced Flow Matching

通过障碍增强流匹配的安全视觉语言动作模型
Sinaei, Kasra, Wu, Hung-Chieh, Ebeigbe, Donald
Abstract
This article presents a modular inference framework that integrates Flow Matching generative models with formal Control Barrier Function (CBF) safety guarantees. Unlike existing methods that apply external safety filters to a model's final output, our approach modifies the Flow Matching denoising process within the model to inherently generate safe trajectories. By employing a smooth Log-Sum-Exponential aggregate barrier, we enforce safety over entire action chunks. This aggregate barrier ensures a minimal increase in computational overhead and does not alter the semantic intent of the model. We show that, within the proposed framework, the 2-Wasserstein distance between the generated distribution and the target distribution remains bounded. Our method eliminates the need for safety-specific datasets or costly model retraining, providing a versatile solution for safe inference. We validate the approach on two robotic manipulation platforms and a 2D navigation benchmark, verifying that our framework achieves reliable safety without degrading the success rate of the model.
Chinese Translation
本文提出了一种模块化推理框架,该框架将流匹配生成模型与形式控制障碍函数(Control Barrier Function, CBF)安全保障相结合。与现有方法在模型最终输出上应用外部安全过滤器不同,我们的方法在模型内部修改流匹配去噪过程,以内在地生成安全轨迹。通过采用平滑的对数和指数聚合障碍,我们在整个动作块上强制执行安全性。该聚合障碍确保计算开销的最小增加,并且不改变模型的语义意图。我们展示了在所提出的框架内,生成分布与目标分布之间的2-Wasserstein距离保持有界。我们的方法消除了对安全特定数据集或昂贵模型重训练的需求,为安全推理提供了一种多功能解决方案。我们在两个机器人操作平台和一个二维导航基准上验证了该方法,确认我们的框架在不降低模型成功率的情况下实现了可靠的安全性。
cs.RO / 27 / 2607.29596

FibVLA: An Efficient Temporal Vision-Language-Action Model with Fibonacci Sampling

FibVLA:一种高效的时间视觉-语言-动作模型,采用斐波那契采样
Lin, Li, Xu, Wujun, Meng, Weiwei, Xia, Kaiwen, Cheong, Kang Hao, Wang, Shuai
Abstract
Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-world actions, provide a generalized solution for embodied AI applications. Conventional VLAs usually concentrate on current digital cognition. While some efforts are made to enhance VLAs' reasoning capabilities by capturing temporal information, encoding the long-context history causes an efficiency-decreasing issue. To reconcile the conflict between capturing temporal information and maintaining inference efficiency in VLAs, this paper introduces FibVLA, an efficient framework featuring temporal perception of long-context history. Specifically, we leverage logarithmic hindsight sampling to both proprioceptive states and visual frames to capture long-term temporal dependencies with minimal redundancy. For the action expert, we introduce the flow matching to produce action distributions, and the Fibonacci recurrent inference strategy to generate long-range planning steps based on real-time closed-loop feedback. Experiments demonstrate that FibVLA significantly improves action smoothness and success rates without retraining large-scale visual encoders. Efficiency analysis demonstrates superior real-time responsiveness compared to video-based baselines in real-world evaluations.
Chinese Translation
视觉-语言-动作模型(VLA)利用多模态信息的认知来推断物理世界中的动作,为具身人工智能应用提供了一种通用解决方案。传统的VLA通常集中于当前的数字认知。尽管一些努力旨在通过捕捉时间信息来增强VLA的推理能力,但对长时间上下文历史的编码会导致效率下降的问题。为了解决在VLA中捕捉时间信息与保持推理效率之间的冲突,本文提出了FibVLA,一种具有长时间上下文历史的时间感知的高效框架。具体而言,我们利用对本体状态和视觉帧的对数回顾采样,以最小的冗余捕捉长期时间依赖性。对于动作专家,我们引入流匹配来生成动作分布,并采用斐波那契递归推理策略,根据实时闭环反馈生成长远规划步骤。实验表明,FibVLA显著提高了动作的平滑性和成功率,而无需重新训练大规模视觉编码器。效率分析显示,与基于视频的基线相比,在现实世界评估中具有更优越的实时响应能力。
cs.RO / 28 / 2607.29600

HAM-VLN: Harnessing Hierarchical Agentic Memory for Zero-Shot Vision-and-Language Navigation

HAM-VLN:利用层次代理记忆实现零-shot视觉与语言导航
Liu, An, Liu, Bingxi, Ding, Hongyu, Jiang, Yixuan, Chen, Yaran, Tang, Fulin, Leng, Cong, Zhang, Hong, Cheng, Jian
Abstract
Vision-and-language navigation (VLN) enables robots to follow instructions in previously unseen environments. Recently, a training-free paradigm has emerged: the robot queries a multimodal LLM to understand its observations and plan the next action. However, long-horizon navigation based on either image streams or dense map inevitably introduces a growing memory and reasoning bottleneck. We present HAM-VLN, a decision-coupled, agent-authored memory that equips the robot with a persistent, depth-grounded world graph. In the same model call used to select the next action, HAM-VLN also records semantic and reflective information---including room type, objects, navigation progress, and failure notes. Recent waypoints remain verbatim within a bounded window, while older history re-enters the context only through retrieval scored by relevance, recency, and salience, together with one-hop topological expansion. This design requires no additional LLM calls beyond the per-waypoint decision. Compared to previous methods, HAM-VLN not only improves various navigation metrics but also reduces the context length by more than 65%. Specifically, HAM-VLN achieves 61.0% Success Rate (SR) on VLN-CE R2R, 52.7% SR on VLN-CE RxR, and 79.7% SR on HM3D-v2 ObjectNav without any training.
Chinese Translation
视觉与语言导航(VLN)使机器人能够在之前未见过的环境中遵循指令。最近,出现了一种无训练范式:机器人查询多模态大型语言模型(LLM)以理解其观察结果并规划下一步行动。然而,基于图像流或密集地图的长时间导航不可避免地引入了不断增长的记忆和推理瓶颈。我们提出了HAM-VLN,这是一种决策耦合的代理创作记忆,赋予机器人一个持久的、基于深度的世界图。在用于选择下一步行动的同一模型调用中,HAM-VLN还记录了语义和反思信息——包括房间类型、物体、导航进度和失败记录。最近的航点在一个有限窗口内保持逐字不变,而较旧的历史信息仅通过相关性、时效性和显著性评分的检索重新进入上下文,并结合一次拓扑扩展。该设计不需要在每个航点决策之外进行额外的LLM调用。与之前的方法相比,HAM-VLN不仅改善了各种导航指标,还将上下文长度减少了超过65%。具体而言,HAM-VLN在VLN-CE R2R上实现了61.0%的成功率(SR),在VLN-CE RxR上实现了52.7%的SR,以及在HM3D-v2 ObjectNav上实现了79.7%的SR,且无需任何训练。
cs.RO / 29 / 2607.29613

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

WCM:一种用于视觉-语言-动作强化学习的世界评论模型
Fei, Senyu, Yu, Xiaopeng, Wang, Siyin, Zhao, Xianzhong, Gong, Jingjing, Qiu, Xipeng
Abstract
Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or single-frame VLM backbone latents, which is a fundamental mismatch with the partially observable nature of robot control. A naive approach to incorporate observation history into the critic incurs exponential complexity with high-dimensional visual space, and still fails because pure scalar-return regression provides insufficient supervision for learning cross-temporal dynamics. We identify the root cause as a state approximation problem: without an explicit world modeling objective, the critic's representation cannot capture the temporal structure needed for accurate value estimation. To address this, we propose the World Critic Model (WCM), built on a lightweight LeJEPA architecture; WCM jointly predicts future latent state and estimates values, such that the critic's representation is explicitly trained to capture temporal dynamics rather than merely regress scalar returns. WCM integrates seamlessly into both on-policy and off-policy training pipelines and is compatible with state-of-the-art VLA backbones including Pi0, Pi0.5, and OpenVLA-OFT. Extensive experiments on 149 tasks across four benchmarks demonstrate that WCM consistently achieves state-of-the-art performance in both in-distribution and out-of-distribution settings, with particularly strong generalization gains. We further validate WCM on seven real-world manipulation tasks using OpenVLA-OFT and Pi0.5 with off-policy RL, confirming stable deployment across diverse settings.
Chinese Translation
强化学习(RL)在视觉-语言-动作(VLA)模型的后训练中显示出对机器人操作的强大潜力。在强化学习方法中,基于评论者的方法依赖于一个主要在单帧观测或单帧VLM主干潜在变量上运行的价值估计器,这与机器人控制的部分可观察特性存在根本不匹配。将观测历史简单地纳入评论者会导致在高维视觉空间中呈指数级复杂性,并且仍然失败,因为纯标量回归提供的监督不足以学习跨时间动态。我们将根本原因确定为状态近似问题:没有明确的世界建模目标,评论者的表征无法捕捉准确价值估计所需的时间结构。为了解决这个问题,我们提出了世界评论模型(WCM),基于轻量级的LeJEPA架构;WCM共同预测未来潜在状态并估计价值,使得评论者的表征明确训练以捕捉时间动态,而不仅仅是回归标量回报。WCM无缝集成到在线和离线训练流程中,并与包括Pi0、Pi0.5和OpenVLA-OFT在内的最新VLA主干兼容。在四个基准上的149个任务的广泛实验表明,WCM在分布内和分布外设置中始终实现了最先进的性能,特别是在泛化能力上取得了显著提升。我们进一步在七个真实世界的操作任务中使用OpenVLA-OFT和Pi0.5进行离线强化学习验证WCM,确认其在多样化环境中的稳定部署。
cs.RO / 30 / 2607.29622

RayViT: Ray-Conditioned Visual Representations for Viewpoint-Robust Imitation Learning

RayViT:用于视角鲁棒模仿学习的光线条件视觉表示
Wang, Qian, Chen, Longrui, Sun, Peiran, Taranovic, Aleksandar, Freymuth, Niklas, Li, Ge, Liao, Weiran, Nagy, C. F. Maximilian, Tan, Yucheng, Chen, Tao, Neumann, Gerhard
Abstract
Visual imitation learning enables robots to acquire visuomotor skills directly from images, yet RGB observations lack explicit geometric cues, making learned policies brittle to camera perturbations. To address this, we propose \textbf{Ray-conditioned Vision Transformer Encoder (RayViT)}, a lightweight architecture that injects camera geometry into pretrained ViT backbones. RayViT represents camera geometry as a Pl\"ucker ray map, patchifies it into ray features, and uses gated cross-attention to produce a ray-conditioned class token. These ray features are added as dense positional embeddings, while the ray class token replaces the original ViT class token to provide a geometry-aware summary representation. We combine this approach with an auxiliary cosine similarity loss to consistently improve the performance and robustness for geometry-aware tokens. Experiments on sim- and real-robot tasks demonstrate that RayViT improves robustness by approximately 13 percentage points under camera perturbations in multi-task RoboCasa benchmark and by 1.78 average completed stages in real-world multi-task success rate compared to baselines.
Chinese Translation
视觉模仿学习使机器人能够直接从图像中获取视觉运动技能,但RGB观察缺乏明确的几何线索,使得学习到的策略对相机扰动变得脆弱。为了解决这个问题,我们提出了 extbf{光线条件视觉变换器编码器(RayViT)},这是一种轻量级架构,将相机几何信息注入预训练的ViT骨干网络。RayViT将相机几何表示为Pl"ucker光线图,将其划分为光线特征,并使用门控交叉注意力生成光线条件类标记。这些光线特征作为密集位置嵌入添加,而光线类标记则替代原始ViT类标记,以提供几何感知的摘要表示。我们将这种方法与辅助余弦相似度损失结合,以持续提高几何感知标记的性能和鲁棒性。在模拟和真实机器人任务上的实验表明,与基线相比,RayViT在多任务RoboCasa基准下在相机扰动下提高了约13个百分点的鲁棒性,并在真实世界多任务成功率中平均完成阶段提高了1.78。
cs.RO / 31 / 2607.29625

Balancing of Humanoid with Object Mass: Trade-off Analyses and Lifting Control

带物体质量的人形机器人平衡:权衡分析与提升控制
Song, Hyunjong, Peng, William Z., Kim, Joo H.
Abstract
The demand for humanoid loco-manipulation tasks with an object has recently increased, and most existing control approaches for stability in such tasks rely on heuristics or machine-learning techniques. This study rigorously analyzes and exploits the dynamic effects of the object mass on balance stability. By formulating the object mass parameters in the whole-body dynamics with distributed contact wrenches and centers of pressure at the stance contacts, their nonlinear effects on the system momenta and constraints are quantified. The dynamic models and constraints are incorporated into the construction of the balanced state basin/boundary (BSB), a partition of the center-of-mass state space for a biped system to maintain balance in its desired contacts. The implications of the BSB for prediction and control are highlighted using a humanoid robot and an analytically tractable reduced-order mechanism. The BSBs under different conditions of base of support, actuation capacity, and pose provide systematic analyses of the effects of object mass on the balancing capability of a system. In particular, the trade-off relationships between momentum regulation and limiting factors in balancing are characterized, introducing two key quantities of the object: the critical mass, at which the system's balancing capability is maximum, and the transition mass, which activates different limiting factors. In addition, sufficient conditions for imposing balanced states on a trajectory are established and implemented with BSBs as explicit threshold constraints in the whole-body trajectory optimization for stable object-lifting control of the humanoid, demonstrating the lift-and-hold and lift-and-release tasks with distinct mass properties in simulations and experiments.
Chinese Translation
近年来,对带物体的人形机器人运动操作任务的需求不断增加,而现有的大多数稳定性控制方法依赖于启发式或机器学习技术。本研究严格分析并利用物体质量对平衡稳定性的动态影响。通过在全身动力学中将物体质量参数与分布的接触扭矩和支撑接触点的压力中心结合,量化其对系统动量和约束的非线性影响。动态模型和约束被纳入平衡状态盆地/边界(BSB)的构建中,这是一个双足系统在期望接触点维持平衡的质心状态空间的划分。使用人形机器人和一个可解析的降阶机制,强调了BSB在预测和控制中的意义。在不同支撑基础、驱动能力和姿态条件下的BSB提供了系统分析物体质量对系统平衡能力影响的结果。特别地,动量调节与平衡中的限制因素之间的权衡关系被特征化,引入了物体的两个关键量:临界质量,此时系统的平衡能力达到最大,以及过渡质量,激活不同的限制因素。此外,建立了在轨迹上施加平衡状态的充分条件,并将BSB作为明确的阈值约束应用于全身轨迹优化,以实现人形机器人的稳定物体提升控制,展示了在仿真和实验中具有不同质量特性的提升与保持及提升与释放任务。
cs.RO / 32 / 2607.29640

Bootstrapping Self-Supervised Learning of Binary Classification Using Error Bounds: A Case Study on a Robotic Insertion Task

利用误差界限引导自监督学习的二分类:以机器人插入任务为例
Duan, Zebin, Krüger, Norbert, Heredia, Juan, Iversen, Thorbjørn Mosekjær, Hagelskjær, Frederik
Abstract
Flexible manufacturing requires rapid deployment of solutions and minimal setup time to remain competitive. An essential attribute is the ability to control error levels, as failures can range from minor performance degradation to severe equipment damage. However, conventional deployment often involves extensive setup, data collection, model training or parameter tuning, and system testing, resulting in significant delays that hinder commercial feasibility. We propose a data engine which gathers data and improves its performance while executing the task. The data engine consists of two classifiers, a fast model prediction and expensive verification. First, a model prediction is performed and based on the confidence level of the prediction, the expensive verification can be used. By adjusting the confidence level, users can control the level of tolerable error. Our method is implemented on a real-world robotic insertion task, which uses force data for the model prediction. The system applies UMAP dimensionality reduction and uses Wilson-Score to compute the confidence bounds of the prediction. Results demonstrate the ability to learn and reduce the need for expensive verifications over time, while staying within the set error-rate. The results highlight the potential of confidence bounds in self-improving models to enhance reliability in robotic classification task.
Chinese Translation
灵活制造要求快速部署解决方案和最小化设置时间以保持竞争力。一个重要特征是控制误差水平的能力,因为故障可能从轻微的性能下降到严重的设备损坏。然而,传统的部署通常涉及大量的设置、数据收集、模型训练或参数调整以及系统测试,导致显著的延迟,从而妨碍商业可行性。我们提出了一种数据引擎,该引擎在执行任务时收集数据并提高其性能。数据引擎由两个分类器组成,一个是快速模型预测,另一个是昂贵的验证。首先进行模型预测,基于预测的置信水平,可以使用昂贵的验证。通过调整置信水平,用户可以控制可容忍的误差水平。我们的方法在一个真实的机器人插入任务中实现,该任务使用力数据进行模型预测。系统应用UMAP降维,并使用Wilson-Score计算预测的置信界限。结果表明,随着时间的推移,能够学习并减少对昂贵验证的需求,同时保持在设定的误差率内。结果突显了置信界限在自我改进模型中增强机器人分类任务可靠性的潜力。
cs.RO / 33 / 2607.29687

Diagnosing Compositional Generalization in Sequential Robot Tasks

诊断顺序机器人任务中的组合泛化
Wang, Yixiao, Wu, Cheng-En, Sun, Lingfeng, Wang, Pengcheng, Ji, Xiang, Liang, Boyuan, Zhan, Guojian, Tomizuka, Masayoshi
Abstract
Sequential robot manipulation requires policies to execute novel combinations of familiar instruction components. However, collecting demonstrations for all possible instruction tuples is combinatorially expensive, while sparsely covered datasets often fail under out-of-distribution recombination. This paper studies compositional generalization through the lens of instruction-space coverage. We decompose the generalization gap into three sources: \textit{marginal instruction shift}, \textit{instruction-compositional shift}, and \textit{context--action shift}. This decomposition allows us to diagnose when sparse training coverage is sufficient, and what structure the training set must preserve for reliable action prediction. Our results show that exhaustive tuple enumeration is unnecessary: a structured subset, as small as one quarter of the full task space, can recover strong out-of-distribution performance when it covers action-relevant dependencies. We further find that sparse training often fails due to instruction steering rather than missing low-level skills; finetuning only one demonstration per task improves OOD success from \(0.4\%\) to \(54.7\%\). For semantically dependent tasks, effective coverage must capture relational structure rather than only factor diversity. These findings suggest that efficient robot data collection should prioritize dependency coverage in instruction space over exhaustive task expansion. More results are available in the supplementary material. Project website: https://yixiaowang7.github.io/Diagnosing_Compositional_Generalization_Robot_Page/.
Chinese Translation
顺序机器人操作要求策略能够执行熟悉指令组件的新颖组合。然而,为所有可能的指令元组收集示例在组合上是昂贵的,而稀疏覆盖的数据集在分布外重组时往往失败。本文通过指令空间覆盖的视角研究组合泛化。我们将泛化差距分解为三个来源: extit{边际指令偏移}、 extit{指令组合偏移}和 extit{上下文-动作偏移}。这种分解使我们能够诊断稀疏训练覆盖何时足够,以及训练集必须保留什么结构以实现可靠的动作预测。我们的结果表明,详尽的元组枚举并非必要:一个结构化的子集,甚至只有完整任务空间的四分之一,当它覆盖与动作相关的依赖关系时,可以恢复强大的分布外性能。我们进一步发现,稀疏训练往往因指令引导而失败,而非缺失低级技能;每个任务仅微调一个示例将OOD成功率从0.4 extit{ extperthousand}提高到54.7 extit{ extperthousand}。对于语义相关的任务,有效的覆盖必须捕捉关系结构,而不仅仅是因子多样性。这些发现表明,高效的机器人数据收集应优先考虑指令空间中的依赖覆盖,而非详尽的任务扩展。更多结果见补充材料。项目网站:https://yixiaowang7.github.io/Diagnosing_Compositional_Generalization_Robot_Page/。
计算机视觉 (Computer Vision)
75
cs.CV / 1 / 2607.28751

ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding

ReLoop-UME:具有可学习检索寄存器的递归深度用于通用多模态嵌入
Wang, Shijie, Hao, Xiangzhao, Li, Yueti, Cao, Guangyu, Tang, Xinyu, Guo, Haiyun
Abstract
Universal multimodal embedding (UME) maps heterogeneous multimodal inputs into a shared embedding space. Existing UME models either form embeddings through single forward encoding or add computation through explicit rationale tokens and latent autoregressive states. Although token expansion can improve complex matching, serial generation increases retrieval latency and makes the final embedding depend on generated intermediate states. This raises a different question: can useful computation be expanded along model depth while keeping the token workspace fixed? We analyze positive-negative similarity separation at every layer of independently trained UME models and observe a shared progression: early layers contextualize multimodal inputs, a contiguous middle-to-late stage forms retrieval-discriminative features, and the final layers map them into the embedding space. Based on this finding, we propose ReLoop-UME, which executes the early layers once, recurrently reuses a parameter-shared retrieval-forming block, and applies the final mapping layers after the last loop. Learnable Retrieval Registers provide persistent retrieval-specific states that accumulate and exchange evidence across loops, with the final register serving as the embedding readout. On MMEB-V2 and MRMR, ReLoop-UME consistently improves retrieval across different backbones while running 44.9x faster than UME-R1 and 1.5x faster than PLUME.
Chinese Translation
通用多模态嵌入(UME)将异构多模态输入映射到共享嵌入空间。现有的 UME 模型要么通过单次前向编码形成嵌入,要么通过显式的推理令牌和潜在的自回归状态增加计算。尽管令牌扩展可以改善复杂匹配,但串行生成增加了检索延迟,并使最终嵌入依赖于生成的中间状态。这引出了一个不同的问题:是否可以在保持令牌工作空间固定的情况下,沿着模型深度扩展有用的计算?我们分析了独立训练的 UME 模型在每一层的正负相似性分离,并观察到一个共同的进展:早期层对多模态输入进行上下文化,中间到后期阶段形成检索区分特征,最后几层将其映射到嵌入空间。基于这一发现,我们提出了 ReLoop-UME,该模型执行早期层一次,递归地重用一个参数共享的检索形成块,并在最后一次循环后应用最终映射层。可学习的检索寄存器提供持久的检索特定状态,这些状态在循环中积累和交换证据,最终寄存器作为嵌入读取输出。在 MMEB-V2 和 MRMR 数据集上,ReLoop-UME 在不同的骨干网络上始终改善检索性能,同时比 UME-R1 快 44.9 倍,比 PLUME 快 1.5 倍。
cs.CV / 2 / 2607.28759

SCMA: Structure-Conditioned and Metal-Aware Flow Matching for CT Metal Artifact Reduction

SCMA:基于结构条件和金属感知的CT金属伪影减少流匹配
Wang, Heran, Sun, Jianing, Jiang, Xu, Ma, Genwei, Zhao, Xing, Duan, Jigang
Abstract
In X-ray CT, metallic objects cause beam hardening, photon starvation, and scattering, leading to projection inconsistency, streaks, dark bands, and structural distortions that compromise clinical diagnosis and quantitative analysis. Existing metal artifact reduction (MAR) methods remain limited: optimization-based methods may leave residual artifacts or blur structures, regression networks may generalize poorly across scenarios, and generative models without sample-specific structural guidance and physical constraints may produce anatomically inconsistent structures. Flow Matching learns a continuous-time velocity field that deterministically transports a source distribution to a target distribution, providing a flexible MAR prior. However, standard unconditional Flow Matching does not exploit sample-specific structure, spatially nonuniform metal-induced degradation, or measured projections. To address these limitations, we propose SCMA, a structure-conditioned and metal-aware Flow Matching framework. First, a linear-interpolation-corrected image is fed into the velocity network with the intermediate state as a sample-specific structural condition, guiding inference toward artifact-free CT images while preserving anatomy. Second, time-varying spatial weights from the metal mask and its distance transform are incorporated into the Flow Matching loss to emphasize severe degradation within and around metal regions. Finally, conditional Flow Matching updates alternate with projection-consistency correction during inference, allowing reliable measurements outside metal traces to constrain predictions. Experiments on simulated and real CT data demonstrate that SCMA more effectively suppresses metal artifacts, preserves local anatomical structures, and reduces hallucination-like structures inconsistent with projection measurements than representative MAR methods.
Chinese Translation
在X射线CT中,金属物体会导致束硬化、光子饥饿和散射,从而导致投影不一致、条纹、暗带和结构扭曲,这些都会影响临床诊断和定量分析。现有的金属伪影减少(MAR)方法仍然存在局限性:基于优化的方法可能会留下残余伪影或模糊结构,回归网络在不同场景下的泛化能力较差,而缺乏样本特定结构指导和物理约束的生成模型可能会产生解剖学上不一致的结构。流匹配(Flow Matching)学习一个连续时间速度场,确定性地将源分布传输到目标分布,从而提供灵活的MAR先验。然而,标准的无条件流匹配并未利用样本特定结构、空间上非均匀的金属诱导退化或测量的投影。为了解决这些局限性,我们提出了SCMA,一个基于结构条件和金属感知的流匹配框架。首先,将经过线性插值校正的图像输入速度网络,并将中间状态作为样本特定的结构条件,引导推断朝向无伪影的CT图像,同时保留解剖结构。其次,从金属掩模及其距离变换中提取的时间变化空间权重被纳入流匹配损失中,以强调金属区域内外的严重退化。最后,在推断过程中,条件流匹配与投影一致性校正交替更新,使得金属痕迹外的可靠测量能够约束预测。对模拟和真实CT数据的实验表明,SCMA在抑制金属伪影、保留局部解剖结构以及减少与投影测量不一致的幻觉状结构方面,比代表性的MAR方法更为有效。
cs.CV / 3 / 2607.28760

WaiT for the Signal: Simple Frequency-Aware Flow-Matching

等待信号:简单的频率感知流匹配
Pavasovic, Krunoslav Lehman, Vallaeys, Théophane, Mallat, Stéphane, Biroli, Giulio, Zettlemoyer, Luke, Karrer, Brian, Verbeek, Jakob
Abstract
As image generation models scale to ever higher resolutions, global coherence, local detail, and texture fidelity become critical axes for generation quality. However, standard flow matching treats all spatial frequencies uniformly, ignoring the natural frequency hierarchy where high-frequency bands become indistinguishable from pure noise far earlier than coarse structures. We introduce WaiT, a Wavelet-aware image Transformer that decomposes generation into coarse and fine bands via lossless wavelets. True to its name, the high-frequency bands wait for the signal: staying pure noise until coarse structure has emerged, then joining the flow for joint refinement. Since standard FID discards fine-grained detail through aggressive downsampling, we introduce a more stringent three-axis evaluation protocol to assess quality at native resolution. On ImageNet 512x512, WaiT achieves a pixel-space FID of 1.43 and is Pareto-optimal across all three axes, reducing sampling compute by up to 50%. With our largest 2B model, we set a new state-of-the-art FID of 1.3 for pixel-space models on ImageNet 512 resolution. Our formulation outperforms even the strongest latent-space models on texture fidelity, and scales seamlessly to high-resolution OpenImages and to video generation, achieving a state-of-the-art FVD of 0.84 on Kinetics-600 with no algorithmic modifications.
Chinese Translation
随着图像生成模型的分辨率不断提升,全球一致性、局部细节和纹理保真度成为生成质量的关键指标。然而,标准的流匹配方法对所有空间频率采取统一处理,忽视了自然频率层次结构,其中高频带比粗糙结构更早地与纯噪声难以区分。我们提出了WaiT,一种波浪变换感知的图像变换器,通过无损小波将生成过程分解为粗糙和细致的频带。正如其名,高频带在信号到来之前保持纯噪声状态:在粗糙结构出现之前不参与流动,待其出现后再加入流动进行联合精细化。由于标准的FID通过激进的下采样丢弃了细粒度细节,我们引入了一种更严格的三轴评估协议,以在原生分辨率下评估质量。在ImageNet 512x512上,WaiT实现了1.43的像素空间FID,并在所有三个指标上均为帕累托最优,减少了高达50%的采样计算量。使用我们最大的2B模型,我们在ImageNet 512分辨率上设定了1.3的新状态-of-the-art FID。我们的公式在纹理保真度上甚至超越了最强的潜在空间模型,并无缝扩展到高分辨率的OpenImages和视频生成,在Kinetics-600上实现了0.84的状态-of-the-art FVD,且没有算法修改。
cs.CV / 4 / 2607.28769

Uncertainty-Aware Deepfake Detection via Multi-View Structural Learning

基于多视角结构学习的不确定性感知深度伪造检测
Farooq, Muhammad Umar, Uddin, Kutub, Khan, Awais, Malik, Khalid
Abstract
Security-critical biometric and forensic applications require accurate predictions and reliable confidence estimates, particularly under distribution shift. This challenge is especially acute for deepfake detection, where foundation-model-based detectors often exhibit overconfident predictions on out-of-distribution manipulations, which limits their suitability for operational deployment. We propose an uncertainty-aware deepfake detection framework that identifies manipulations through inconsistencies across complementary evidence sources. The framework integrates three streams: a visual stream based on an adapted CLIP encoder, a semantic stream that models consistency among facial attributes through differentiable constraints, and a structural stream that captures class-dependent dependency patterns between semantic and forensic features. To effectively combine these signals, we introduce Inter-Branch Disagreement Calibration (IBDC), a disagreement-aware uncertainty modeling mechanism that links predictive uncertainty to conflicts among evidence streams. Extensive cross-dataset experiments using FaceForensics++ as the training source demonstrate that the proposed framework achieves state-of-the-art generalization across multiple out-of-distribution benchmarks while consistently improving calibration and selective prediction performance. These results show that combining complementary evidence with disagreement-aware uncertainty provides a robust foundation for trustworthy and well-calibrated deepfake detection under distribution shift.
Chinese Translation
安全关键的生物识别和法医应用需要准确的预测和可靠的置信度估计,特别是在分布变化的情况下。这一挑战在深度伪造检测中尤为突出,因为基于基础模型的检测器在处理分布外的操控时常常表现出过于自信的预测,这限制了它们的实际应用适用性。我们提出了一种不确定性感知的深度伪造检测框架,通过识别互补证据源之间的不一致性来识别操控。该框架整合了三个流:基于改进的 CLIP 编码器的视觉流、通过可微约束建模面部属性一致性的语义流,以及捕捉语义特征与法医特征之间类别依赖模式的结构流。为了有效结合这些信号,我们引入了跨分支不一致校准(Inter-Branch Disagreement Calibration, IBDC),这是一种关注不一致性的不确定性建模机制,将预测不确定性与证据流之间的冲突联系起来。使用 FaceForensics++ 作为训练源的广泛跨数据集实验表明,所提出的框架在多个分布外基准上实现了最先进的泛化能力,同时持续改善了校准和选择性预测性能。这些结果表明,结合互补证据与关注不一致性的不确定性为在分布变化下可靠且良好校准的深度伪造检测提供了坚实的基础。
cs.CV / 5 / 2607.28771

Do Medical Foundation Models Generalize on the African Brain?

医学基础模型在非洲大脑上的泛化能力如何?
Mouheb, Kaouther, Rojas, Gonzalo Esteban Mosquera, van Leeuwen, Juancito, Klein, Stefan, Bron, Esther E.
Abstract
Medical foundation models (FMs) are increasingly used for brain MRI analysis. However, their evaluation remains dominated by high-resource datasets, leaving generalization to African cohorts underexplored. We assess whether FMs generalize equally to African and non-African brain MRI data across two tasks: dementia classification using a Nigerian dataset and brain tumor segmentation using BraTS-Africa. We evaluate two generalist FMs (BrainIAC, 3DINO) and two segmentation-specific FMs (MedSAM2, Medical-SAM2) against a from-scratch baseline. For classification, FMs provide limited gains (highest ROC-AUC of 0.86 with BrainIAC), whereas for segmentation they consistently improve performance, reaching up to 0.86 Dice with MedSAM2. Performance differences between African and non-African cohorts are inconsistent and appear more related to dataset size than data origin. These results suggest that FMs do not exhibit an inherent bias against African cohorts, and highlight the limited availability and diversity of African neuroimaging datasets as the main barrier to robust evaluation and deployment.
Chinese Translation
医学基础模型(FMs)在脑部MRI分析中越来越多地被使用。然而,它们的评估仍然主要依赖于高资源数据集,使得对非洲人群的泛化能力研究不足。我们评估FMs在两个任务上是否对非洲和非非洲脑部MRI数据具有相同的泛化能力:使用尼日利亚数据集进行痴呆分类,以及使用BraTS-Africa进行脑肿瘤分割。我们对比了两个通用FMs(BrainIAC,3DINO)和两个特定于分割的FMs(MedSAM2,Medical-SAM2)与从零开始的基线。在分类任务中,FMs提供的增益有限(BrainIAC的最高ROC-AUC为0.86),而在分割任务中,它们的性能持续改善,MedSAM2的Dice系数最高可达0.86。非洲和非非洲人群之间的性能差异不一致,似乎与数据集大小关系更大,而非数据来源。这些结果表明,FMs并未对非洲人群表现出固有的偏见,并强调非洲神经影像数据集的有限可用性和多样性是进行稳健评估和部署的主要障碍。
cs.CV / 6 / 2607.28796

Can Synthetic Data Overcome the Generalization Limits of AI-Based Flower and Pod Detection Across Cowpea Breeding Genotypes and Environments?

合成数据能否克服基于人工智能的花朵和荚果检测在豇豆育种基因型和环境中的泛化限制?
Kamangir, Hamid, Berlingeri, Jonathan, Ranario, Earl, Uyehara, Isaac Kazuo, Lundqvist, Lars, Yun, Heesup, Diepenbrock, Christine H., Bailey, Brian N., Earles, J. Mason
Abstract
High-throughput phenotyping requires AI-enabled computer vision models that generalize across genotypes, locations, and growing seasons, yet such models often lose accuracy under new conditions. Annotating real imagery for every genotype-by-environment (G x E) combination a breeding program encounters is prohibitively expensive. We quantify how G x E shifts affect AI-based detection of cowpea flowers and pods across two California locations and two growing seasons. Flower detection mAP@50 fell from 76.3% to as low as 50.6% under unseen shifts, and pod detection was more sensitive. Feature-space and image-quality diagnostics confirmed these losses track measurable distributional shifts. Because closing this gap with real data alone is not practical, we test whether synthetic imagery, rendered from a procedural 3D cowpea model, can substitute for that annotation burden. Synthetic supervision alone improved over pretraining but remained limited by a domain gap driven by camera image formation, not scene content. A domain-gap-aware camera-realism augmentation strategy, optimized against measured real-image statistics via Wasserstein distance, narrowed this gap, and a linear HDR representation converted a smaller measured gap into a larger detection gain than an 8-bit representation. Optimized HDR synthetic data combined with as few as five real images matched or exceeded the real-data baseline for spatial generalization, and pod detection benefited most at the lowest shot counts, with more modest gains under temporal shift. These results show that synthetic data can overcome the generalization limits of AI-based flower and pod detection, but only when the domain gap is measured and optimized rather than assumed away.
Chinese Translation
高通量表型分析需要能够在基因型、地点和生长季节之间泛化的基于人工智能的计算机视觉模型,但在新条件下,这些模型的准确性往往会下降。为每个育种项目遇到的基因型与环境(G x E)组合标注真实图像的成本过于高昂。我们量化了G x E变化如何影响在加利福尼亚两个地点和两个生长季节中基于人工智能的豇豆花朵和荚果检测。花朵检测的mAP@50从76.3%降至低至50.6%,而荚果检测则更为敏感。特征空间和图像质量诊断确认这些损失与可测量的分布变化相关。由于仅依靠真实数据来缩小这一差距并不实际,我们测试了从程序化3D豇豆模型渲染的合成图像是否可以替代这种标注负担。仅使用合成监督的性能优于预训练,但仍受限于由相机图像形成而非场景内容驱动的领域差距。一种基于领域差距的相机真实感增强策略,通过Wasserstein距离优化与测量的真实图像统计数据,缩小了这一差距,而线性HDR表示将较小的测量差距转化为比8位表示更大的检测增益。优化后的HDR合成数据与少至五张真实图像结合,达到了或超过了真实数据基线的空间泛化,荚果检测在最低拍摄数量下受益最大,而在时间变化下的增益则较为温和。这些结果表明,合成数据能够克服基于人工智能的花朵和荚果检测的泛化限制,但前提是必须对领域差距进行测量和优化,而非假设其不存在。
cs.CV / 7 / 2607.28834

FocusGS: Spatial Delta Layers for Local Repair and Deterministic Editing of Trained 3D Gaussian Assets

FocusGS:用于训练的 3D 高斯资产的局部修复和确定性编辑的空间增量层
Pan, Yiqun, Shi, Yukun
Abstract
3D Gaussian Splatting (3DGS) is evolving from one-time reconstruction into deliverable, inspectable, and maintainable visual assets. Existing workflows focus on global reconstruction, training-time density control, or open-ended generative editing, leaving trained assets without precise local maintenance. We propose FocusGS, which unifies local repair and deterministic editing as composite spatial deltas. Repair is the purely additive special case: its base-manipulation term is empty, and it adds only local Gaussian bases; deterministic editing uses erase-insert factorization (EIF) to combine old-carrier erasure with new-content insertion. FocusGS addresses spatial gradient starvation: local repair raises target-region PSNR by 7.91 dB over 93 evaluation views. Across all 83 deterministic editing trials, the target ROI improves, with a trial-averaged mean edited ROI PSNR of 21.97 dB and a mean gain of +11.05 dB; across five public editing cases, FocusGS-EIF reaches 33.17 dB Target-mask PSNR and 0.994 Target-delta Correlation, while both text-driven baselines fail to complete the prescribed updates. FocusGS provides a lightweight, verifiable 3DGS maintenance operator.
Chinese Translation
3D 高斯溅射(3DGS)正从一次性重建演变为可交付、可检查和可维护的视觉资产。现有工作流程侧重于全局重建、训练时密度控制或开放式生成编辑,导致训练后的资产缺乏精确的局部维护。我们提出了 FocusGS,它将局部修复和确定性编辑统一为复合空间增量。修复是纯粹的加法特例:其基础操作项为空,仅添加局部高斯基;确定性编辑使用擦除-插入因子分解(EIF)将旧载体擦除与新内容插入相结合。FocusGS 解决了空间梯度匮乏的问题:局部修复使目标区域的 PSNR 提高了 7.91 dB,基于 93 个评估视图。在所有 83 次确定性编辑试验中,目标 ROI 得到了改善,试验平均编辑后的 ROI PSNR 为 21.97 dB,平均增益为 +11.05 dB;在五个公共编辑案例中,FocusGS-EIF 达到 33.17 dB 的目标掩码 PSNR 和 0.994 的目标增量相关性,而两个基于文本的基线均未能完成规定的更新。FocusGS 提供了一种轻量级、可验证的 3DGS 维护操作符。
cs.CV / 8 / 2607.28858

A Unified Benchmark of Deep Learning Models for Multi-task 3D Brain Tumor Segmentation from Magnetic Resonance Imaging

统一的深度学习模型基准用于多任务三维脑肿瘤从磁共振成像中分割
Torrejón, Diego J., Hernández, Luna Y., Sánchez, Javier
Abstract
Automatic brain tumor segmentation from magnetic resonance imaging (MRI) has become a fundamental task in computer-assisted diagnosis, treatment planning, and disease monitoring. Although numerous deep learning architectures have recently been proposed, objective comparisons remain challenging because published studies often employ different datasets, preprocessing strategies, training protocols, and evaluation procedures. This work presents a unified experimental benchmark for comparing representative convolutional neural networks (CNNs), Transformer-based models, and recent State Space Model (SSM) architectures under homogeneous experimental conditions. Five state-of-the-art three-dimensional segmentation models, including 3D U-Net, SegResNet, Swin UNETR, SegMamba, and SegMambaV2, are evaluated on two brain tumor segmentation datasets representing distinct clinical scenarios: intracranial meningioma segmentation (BraTS 2023) and post-treatment glioma segmentation (BraTS 2024). All architectures are trained using identical preprocessing, data augmentation, optimization strategies, and evaluation protocols to ensure a fair comparison. Performance is assessed using segmentation accuracy metrics together with computational cost indicators, including inference time and the size of each model. The results provide practical insights into the trade-offs between segmentation accuracy and computational efficiency, highlighting the suitability of different architectural paradigms for challenging three-dimensional brain tumor segmentation tasks.
Chinese Translation
自动从磁共振成像(MRI)中进行脑肿瘤分割已成为计算机辅助诊断、治疗规划和疾病监测中的一项基础任务。尽管最近提出了众多深度学习架构,但由于已发表的研究通常采用不同的数据集、预处理策略、训练协议和评估程序,客观比较仍然具有挑战性。本研究提出了一个统一的实验基准,用于在同质实验条件下比较代表性的卷积神经网络(CNN)、基于Transformer的模型以及近期的状态空间模型(SSM)架构。评估了五种最先进的三维分割模型,包括3D U-Net、SegResNet、Swin UNETR、SegMamba和SegMambaV2,这些模型在两个代表不同临床场景的脑肿瘤分割数据集上进行评估:颅内脑膜瘤分割(BraTS 2023)和治疗后胶质瘤分割(BraTS 2024)。所有架构均采用相同的预处理、数据增强、优化策略和评估协议进行训练,以确保公平比较。通过分割准确性指标以及计算成本指标(包括推理时间和每个模型的大小)来评估性能。结果为分割准确性与计算效率之间的权衡提供了实用的见解,突显了不同架构范式在具有挑战性的三维脑肿瘤分割任务中的适用性。
cs.CV / 9 / 2607.28868

Physics-Aligned Self-Supervised Learning for Scientific Imaging

与物理对齐的自监督学习在科学成像中的应用
Kazimi, Bashir, Sandfeld, Stefan
Abstract
Data augmentations define the invariances learned by self-supervised learning (SSL). Standard augmentation pipelines were designed for natural images, yet scientific imaging modalities are governed by physical measurement processes with distinct symmetry and acquisition constraints. Enforcing invariances that contradict these constraints can distort learned representations and limit downstream performance, but practitioners moving from machine learning into a new scientific modality currently have little guidance beyond transferring natural-image pipelines unexamined. We address this gap with a principled, reproducible procedure for augmentation design in scientific SSL: we formalise the physics-aligned augmentation set as a union of measurement-consistent symmetries and acquisition-driven perturbations, and we give a concrete, largely label-free workflow---enumerate candidates, label each by the measurement operator, validate with representation-geometry diagnostics, and confirm by single-factor ablation---for selecting them. We instantiate the procedure for real-space electron microscopy and reciprocal-space 4D-STEM diffraction, and evaluate it across five SSL paradigms (DINOv2, SimCLR, MAE, VICRegL, I-JEPA) on classification and crystal-orientation regression. Physics-aligned augmentations substantially improve downstream performance for objectives relying on cross-view consistency, reduce geodesic error and improve robustness under realistic acquisition variability (detector gain, resolution loss), and systematically reshape representation geometry. While our experiments use electron microscopy, the procedure is modality-agnostic and applies to other measurement-driven domains such as medical and remote-sensing imaging. These results position augmentation design as a primary, and controllable, source of inductive bias in scientific self-supervised learning.
Chinese Translation
数据增强定义了自监督学习(SSL)所学习的不变性。标准的增强流程是为自然图像设计的,而科学成像模式则受到物理测量过程的支配,具有独特的对称性和采集约束。强制执行与这些约束相悖的不变性可能会扭曲学习到的表示并限制下游性能,但从机器学习转向新的科学模式的实践者目前在转移未经检验的自然图像流程方面几乎没有指导。我们通过一种原则性、可重复的科学SSL增强设计程序来填补这一空白:我们将与物理对齐的增强集形式化为测量一致性对称性和采集驱动扰动的并集,并提供一个具体的、基本上无标签的工作流程——列举候选项,通过测量算子标记每个候选项,使用表示几何诊断进行验证,并通过单因素消融确认——以选择它们。我们将该程序应用于实空间电子显微镜和倒空间4D-STEM衍射,并在分类和晶体取向回归的五种SSL范式(DINOv2、SimCLR、MAE、VICRegL、I-JEPA)上进行评估。与物理对齐的增强显著提高了依赖于视图一致性的目标的下游性能,减少了测地误差,并在现实的采集变异性(探测器增益、分辨率损失)下提高了鲁棒性,同时系统性地重塑了表示几何。尽管我们的实验使用了电子显微镜,但该程序是与模式无关的,并适用于其他测量驱动的领域,如医学成像和遥感成像。这些结果将增强设计定位为科学自监督学习中的主要且可控的归纳偏置来源。
cs.CV / 10 / 2607.28935

Group-wise Supervision with Focal-Dice Loss for Long-Tailed Indoor Semantic Occupancy Prediction

基于聚组监督和焦点-骰子损失的长尾室内语义占用预测
Zheng, Qi, Su, Zihuang, Pan, Xiao
Abstract
Recently, 3D semantic occupancy prediction has garnered increasing attention for understanding the indoor scene. However, unlike structured outdoor environments, indoor scenes feature a high diversity of object categories that exhibit a severe long-tailed distribution, which has become a core bottleneck limiting the performance of existing models. To tackle this challenge, we propose a novel method, Group-UFD Occ, based on hierarchical semantic supervision and synergistic loss optimization. At the architectural level, we introduce a fine-grained semantic grouping strategy and design multi-scale, parallel ``main-expert'' prediction heads to guide the model in efficiently learning tail-class features through deep regularization. At the optimization level, we introduce the Unified Focal-Dice (UFD) loss. This synergistic loss function dynamically focuses on hard samples at the per-voxel level. Meanwhile, it simultaneously optimizes the geometric integrity of predicted objects from a region-based perspective. We conducted experiments on the large-scale EmbodiedScan dataset. The results demonstrate that our method yields a relative improvement of 11.38\% over the baseline, with substantial accuracy gains in several critical long-tailed categories.
Chinese Translation
近年来,3D语义占用预测在理解室内场景方面受到了越来越多的关注。然而,与结构化的户外环境不同,室内场景具有高度多样化的物体类别,并且呈现出严重的长尾分布,这已成为限制现有模型性能的核心瓶颈。为了解决这一挑战,我们提出了一种新颖的方法——Group-UFD Occ,基于分层语义监督和协同损失优化。在架构层面,我们引入了一种细粒度的语义分组策略,并设计了多尺度、并行的“主专家”预测头,以引导模型通过深度正则化有效学习尾类特征。在优化层面,我们引入了统一焦点-骰子(Unified Focal-Dice, UFD)损失。这种协同损失函数动态关注每个体素级别的难样本,同时从区域基础的角度优化预测对象的几何完整性。我们在大规模的EmbodiedScan数据集上进行了实验。结果表明,我们的方法相较于基线实现了11.38%的相对提升,并在多个关键的长尾类别中取得了显著的准确性提升。
cs.CV / 11 / 2607.28936

DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models

DiffAttack:基于潜在扩散模型的面部识别规避攻击
Ahmadieh, Omid, Karimian, Nima
Abstract
Facial biometric identification relies on the distinctiveness of user attributes within a high-dimensional embedding space. However, the decision boundaries of deep face recognition (FR) systems are often sufficiently narrow that they can be conflated, rendering the models vulnerable to adversarial attacks. In such scenarios, the FR system fails to distinguish between an authentic source and a meticulously crafted adversarial face. Existing adversarial methods targeting facial biometrics are limited in both performance and their ability to generate high-quality images that are imperceptible to humans. Moreover, these methods often fail when the source and target images belong to different demographic groups or genders. To address these limitations, we present a novel approach for adversarial face generation via latent-space optimization. We leverage latent diffusion models directly to guide generation toward target identity embeddings, as measured by a face recognition model. Our proposed \textbf{DiffAttack} framework has been evaluated on standard benchmarks, such as the FFHQ and CelebA-HQ datasets. DiffAttack significantly outperforms existing adversarial techniques, achieving a high average attack success rate of 84.86% across multiple face recognition models (e.g., FaceNet). Notably, DiffAttack demonstrates superior transferability, surpassing traditional noise-based methods by over 15.28% and semantic-based approaches by approximately 5.21% on benchmark datasets like FFHQ and CelebA-HQ.
Chinese Translation
面部生物识别识别依赖于用户属性在高维嵌入空间中的独特性。然而,深度面部识别(FR)系统的决策边界往往足够狭窄,以至于可能重叠,从而使模型易受对抗攻击。在这种情况下,FR系统无法区分真实来源和精心制作的对抗面孔。现有针对面部生物识别的对抗方法在性能和生成对人类不可察觉的高质量图像的能力上都存在局限。此外,当源图像和目标图像属于不同的人口群体或性别时,这些方法往往失效。为了解决这些局限性,我们提出了一种通过潜在空间优化进行对抗面孔生成的新方法。我们直接利用潜在扩散模型来引导生成朝向目标身份嵌入,依据面部识别模型进行测量。我们提出的 extbf{DiffAttack} 框架已在标准基准测试上进行了评估,例如 FFHQ 和 CelebA-HQ 数据集。DiffAttack 显著优于现有的对抗技术,在多个面部识别模型(如 FaceNet)上实现了高达 84.86% 的平均攻击成功率。值得注意的是,DiffAttack 展现出优越的可迁移性,在 FFHQ 和 CelebA-HQ 等基准数据集上超过传统噪声基础方法超过 15.28%,并在语义基础方法上超过约 5.21%。
cs.CV / 12 / 2607.28950

Automated classification method of COVID-19 cases from chest CT volumes using 2D and 3D hybrid CNN for anisotropic volumes

基于2D和3D混合卷积神经网络的胸部CT体积COVID-19病例自动分类方法
Oda, Masahiro, Zheng, Tong, Hayashi, Yuichiro, Otake, Yoshito, Hashimoto, Masahiro, Akashi, Toshiaki, Aoki, Shigeki, Mori, Kensaku
Abstract
This paper proposes an automated classification method of chest CT volumes based on likelihood of COVID-19 cases. Novel coronavirus disease 2019 (COVID-19) spreads over the world, causing a large number of infected patients and deaths. Sudden increase in the number of COVID-19 patients causes a manpower shortage in medical institutions. Computer-aided diagnosis (CAD) system provides quick and quantitative diagnosis results. CAD system for COVID-19 enables efficient diagnosis workflow and contributes to reduce such manpower shortage. This paper proposes an automated classification method of chest CT volumes for COVID-19 diagnosis assistance. We propose a COVID-19 classification convolutional neural network (CNN) that has a 2D/3D hybrid feature extraction flows. The 2D/3D hybrid feature extraction flows are designed to effectively extract image features from anisotropic volumes such as chest CT volumes for diagnosis. The flows extract image features on three mutually perpendicular planes in CT volumes and then combine the features to perform classification. Classification accuracy of the proposed method was evaluated using a dataset that contains 1288 CT volumes. An averaged classification accuracy was 83.3%. The accuracy was higher than that of a classification CNN which does not have 2D and 3D hybrid feature extraction flows.
Chinese Translation
本文提出了一种基于COVID-19病例可能性的胸部CT体积自动分类方法。新型冠状病毒肺炎(COVID-19)在全球传播,导致大量感染患者和死亡。COVID-19患者数量的突然增加造成了医疗机构的人力资源短缺。计算机辅助诊断(CAD)系统提供快速且定量的诊断结果。针对COVID-19的CAD系统能够实现高效的诊断工作流程,并有助于减少人力资源短缺。本文提出了一种用于COVID-19诊断辅助的胸部CT体积自动分类方法。我们提出了一种COVID-19分类卷积神经网络(CNN),该网络具有2D/3D混合特征提取流程。2D/3D混合特征提取流程旨在有效提取来自胸部CT体积等各向异性体积的图像特征。该流程在CT体积的三个相互垂直的平面上提取图像特征,然后将这些特征结合以进行分类。通过使用包含1288个CT体积的数据集评估了所提方法的分类准确性,平均分类准确率为83.3%。该准确率高于不具备2D和3D混合特征提取流程的分类CNN的准确率。
cs.CV / 13 / 2607.28955

Retrieval-Driven Training-Free AI-Generated Video Attribution

基于检索驱动的无训练AI生成视频归属
Cheng, Renxi, Han, Chaolei, Gui, Jie, Wang, Hongsong
Abstract
AI-generated videos are becoming increasingly realistic and difficult to distinguish from authentic ones, which facilitates malicious misuse and poses growing threats to cybersecurity and social governance. Attributing AI-generated videos to their specific generative sources is therefore of critical importance for forensic investigation and legal regulation. However, most existing visual attribution methods focus on images and particularly rely on the image generation model, thereby lacking the ability to generalize to large-scale AI-generated video data. To address these limitations, we introduce an training-free AI-generated video attribution paradigm. Specifically, we formulates AI-generated video attribution as an instance retrieval task, and design a generative fingerprint-based pipeline. This pipeline consists of an adapted orthogonal color transformation, multi-scale quantized residual generation, and temporal-semantic aggregation, progressively capturing and integrating artifacts introduced by generative models across video frames. Extensive experiments on the GenVidBench benchmark demonstrate that our method achieves strong performance in both AI-generated video detection and attribution, outperforming existing state-of-the-art methods with a Rank-1 accuracy of 20.5% and a mean Average Precision of 16.6%. The code is at https://github.com/renxi-seu/Video_Attribution.
Chinese Translation
AI生成的视频正变得越来越逼真,难以与真实视频区分,这促进了恶意滥用并对网络安全和社会治理构成了日益严重的威胁。因此,将AI生成的视频归属到其特定的生成源对于法医调查和法律监管至关重要。然而,现有的大多数视觉归属方法主要集中在图像上,特别依赖于图像生成模型,因此缺乏对大规模AI生成视频数据的泛化能力。为了解决这些局限性,我们提出了一种无训练的AI生成视频归属范式。具体而言,我们将AI生成视频归属表述为一个实例检索任务,并设计了一个基于生成指纹的流程。该流程包括适应的正交颜色变换、多尺度量化残差生成和时间语义聚合,逐步捕捉和整合生成模型在视频帧中引入的伪影。在GenVidBench基准上的大量实验表明,我们的方法在AI生成视频检测和归属方面表现出色,超越了现有的最先进方法,Rank-1准确率达到20.5%,平均精度均值为16.6%。代码可在 https://github.com/renxi-seu/Video_Attribution 获取。
cs.CV / 14 / 2607.28967

Visual Distribution Anchoring for Efficient Prompt Tuning

高效提示调优的视觉分布锚定
Parsa, Pouya, Moayedi, Raoof Zare, Choi, Seongjin
Abstract
Prompt tuning adapts vision--language models with few trainable parameters, but existing approaches trade off efficiency and adaptation: static textual prompts can overfit source classes, image-conditioned prompts add per-instance computation, and multimodal tuning modifies the visual branch. We propose VDA (Visual Distribution Anchoring), a training-free target adaptation framework that augments a frozen semantic classifier with class-level visual prototypes estimated offline from an unlabeled target pool. We first ask whether prototypes can be synthesized from class names. A text-to-centroid mapper reconstructs held-out source prototypes but fails under dataset shift because class names specify semantic identity, not target-domain appearance. An oracle analysis confirms that true target prototypes are highly discriminative. VDA therefore uses frozen semantic and domain-template classifiers to partition unlabeled target images into class-correlated groups. Confidence-ranked image features form normalized prototypes, fused with the semantic classifier using one global weight. Adaptation requires no target labels, target-side optimization, uniform class-prior assumption, iterative refinement, or test-query access, and yields a fixed, cacheable classifier. Controlled experiments show that class-specific partitioning drives gains and that visually local pseudo-label errors can remain useful despite being class-incorrect. Across ten ImageNet-to-target transfers, the same frozen design improves zero-shot CLIP, TCP, and MaPLe by 3.22, 3.39, and 3.35 points, respectively, improving nine of ten targets in every setting. Its visual correction further improves leakage-free PromptKD by 2.79 points, complementing zero-shot, source-prompted, multimodal-prompted, and target-distilled classifiers.
Chinese Translation
提示调优通过少量可训练参数来适应视觉-语言模型,但现有方法在效率和适应性之间存在权衡:静态文本提示可能会过拟合源类别,图像条件提示增加了每个实例的计算量,而多模态调优则修改了视觉分支。我们提出了VDA(视觉分布锚定),一种无训练的目标适应框架,它通过从未标记的目标池中离线估计的类别级视觉原型来增强一个冻结的语义分类器。我们首先探讨是否可以从类别名称合成原型。文本到中心映射器重建了保留的源原型,但在数据集迁移下失败,因为类别名称指定了语义身份,而不是目标领域的外观。一个oracle分析确认真实的目标原型具有很高的区分性。因此,VDA使用冻结的语义和领域模板分类器将未标记的目标图像划分为与类别相关的组。基于置信度的图像特征形成标准化的原型,并与语义分类器通过一个全局权重融合。适应过程不需要目标标签、目标侧优化、统一的类别先验假设、迭代细化或测试查询访问,并产生一个固定的、可缓存的分类器。控制实验表明,类别特定的划分推动了性能提升,并且视觉局部伪标签错误尽管类别不正确仍然可以保持有用。在十个ImageNet到目标的迁移中,相同的冻结设计分别提高了零-shot CLIP、TCP和MaPLe的性能3.22、3.39和3.35点,在每种设置中改善了十个目标中的九个。其视觉修正进一步将无泄漏的PromptKD提高了2.79点,补充了零-shot、源提示、多模态提示和目标蒸馏分类器。
cs.CV / 15 / 2607.28969

SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs

SafeNexus:发现和引导多模态通用安全神经元在多模态大型语言模型中的应用
Yu, Jian, Shen, Fei, Wang, Cong, Wang, Jian, Du, Lu Jin. Xiaoyu, Tang, Jinhui, Chua, Tat-Seng
Abstract
Although Large Language Models (LLMs) have demonstrated promising safety performance, extending them to Multimodal Large Language Models (MLLMs) exposes a significant gap between expanded multimodal capabilities and existing safety mechanisms. Current defenses remain predominantly confined to specific modal settings, thereby limiting their robustness against broader cross-modal threats. To bridge this gap, we introduce SafeNexus, a cross-modal safety alignment framework that adopts a dedicated neuron-level intervention strategy. First, we formulate a neuron localization paradigm that identifies functionally specialized neurons by characterizing intermediate-layer activation patterns and quantifying their functional salience through importance scoring. Building upon this paradigm, we exploit contrastive data to identify modality-bound safety neurons (BS-Neurons), and validate their role in regulating safety behavior within each modality via targeted suppression. Further cross-modal analysis defines modality-universal safety neurons (US-Neurons) as the shared subset of BS-Neurons identified across individual modalities, serving as the core for defending against harmful cross-modal attacks. We observe that suppressing these neurons substantially degrades safety performance across modalities, while leaving overall utility largely unaffected. Building on these insights, we propose two safety alignment strategies: activation-level safety amplifier and safety neuron calibrator. The proposed strategies enhance model safety through two distinct routes: the former amplifies the activation magnitudes of US-Neurons, while the latter selectively calibrates them via targeted fine-tuning. Extensive experiments demonstrate that our method outperforms prevailing state-of-the-art approaches on safety benchmarks spanning diverse modality combinations, while effectively preserving utility.
Chinese Translation
尽管大型语言模型(LLMs)在安全性能上表现出良好的前景,但将其扩展到多模态大型语言模型(MLLMs)时,现有的安全机制与扩展的多模态能力之间存在显著差距。目前的防御措施主要局限于特定的模态设置,从而限制了其对更广泛跨模态威胁的鲁棒性。为了解决这一问题,我们提出了SafeNexus,一个跨模态安全对齐框架,采用专门的神经元级干预策略。首先,我们制定了一种神经元定位范式,通过表征中间层激活模式并通过重要性评分量化其功能显著性,识别功能专门化的神经元。在此基础上,我们利用对比数据识别模态绑定安全神经元(BS-Neurons),并通过针对性抑制验证其在调节每个模态内安全行为中的作用。进一步的跨模态分析将模态通用安全神经元(US-Neurons)定义为在各个模态中识别的BS-Neurons的共享子集,作为抵御有害跨模态攻击的核心。我们观察到,抑制这些神经元会显著降低跨模态的安全性能,而整体效用几乎不受影响。在这些见解的基础上,我们提出了两种安全对齐策略:激活级安全放大器和安全神经元校准器。所提出的策略通过两条不同的途径增强模型安全性:前者放大US-Neurons的激活幅度,而后者通过针对性微调选择性地校准它们。大量实验表明,我们的方法在涵盖多种模态组合的安全基准测试中优于现有的最先进方法,同时有效保持了效用。
cs.CV / 16 / 2607.28970

LegoQ: Density-Matrix Representation Learning with Spectral-Spatial State Transitions for Hyperspectral Classification

LegoQ:具有光谱-空间状态转移的密度矩阵表示学习用于高光谱分类
Cao, Weijia, Yang, Xiaofei, Wang, Fu, Zhou, Yicong, Zhou, Xiang
Abstract
Hyperspectral image classification is complicated by mixed pixels, spectral ambiguity, class imbalance, and limited annotations. Most current classifiers encode a pixel or patch as a deterministic vector and apply a linear or multilayer softmax head. Although effective for discrimination, this representation does not directly expose how mixed or uncertain a sample is. This paper presents \method, a classical density-matrix representation learning framework for hyperspectral images. The spectral bands are divided into groups and each group is mapped to a positive semi-definite, Hermitian, trace-normalized matrix state. A composable stack of spectral, spatial, and inter-group transitions then updates the states while repeatedly projecting them back to the valid state set. Instead of flattening the final features, \method\ aggregates the group states and compares them with learnable class-prototype density matrices through Uhlmann fidelity. The normalized eigenspectrum, von Neumann entropy, purity, and prototype fidelity provide sample-level diagnostics that are unavailable from a conventional vector head. On Indian Pines, ten runs yield an overall accuracy of $96.20\pm0.70\%$, an average accuracy of $95.57\pm1.29\%$, and a kappa coefficient of $95.66\pm0.80\%$. On WHU-Hi-LongKou, the best of ten runs reaches $97.52\%$ overall accuracy. Classification maps and feature projections show that the transition stack produces compact and better separated class structures. The results support constrained matrix-state learning as a practical alternative to vector-only hyperspectral classification without requiring quantum hardware.
Chinese Translation
高光谱图像分类受到混合像素、光谱模糊、类别不平衡和有限标注的影响。当前大多数分类器将一个像素或补丁编码为确定性向量,并应用线性或多层softmax头。尽管这种表示在区分上有效,但并未直接揭示样本的混合或不确定程度。本文提出了 extit{LegoQ},一个用于高光谱图像的经典密度矩阵表示学习框架。光谱波段被划分为多个组,每个组被映射为正半定的、厄米的、迹归一化的矩阵状态。一个可组合的光谱、空间和组间转移堆栈随后更新状态,同时将其反复投影回有效状态集。与最终特征展平不同, extit{LegoQ}聚合组状态,并通过Uhlmann保真度将其与可学习的类别原型密度矩阵进行比较。归一化特征谱、冯·诺依曼熵、纯度和原型保真度提供了传统向量头无法获得的样本级诊断。在印度松树数据集上,十次实验的整体准确率为$96.20 ext{±}0.70 ext{ extperthousand}$,平均准确率为$95.57 ext{±}1.29 ext{ extperthousand}$,Kappa系数为$95.66 ext{±}0.80 ext{ extperthousand}$。在WHU-Hi-LongKou数据集上,十次实验的最佳结果达到$97.52 ext{ extperthousand}$的整体准确率。分类地图和特征投影显示,转移堆栈产生了紧凑且更好分离的类别结构。结果支持约束矩阵状态学习作为一种实用的替代方案,用于不依赖量子硬件的仅向量高光谱分类。
cs.CV / 17 / 2607.28974

RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images

RAID:朝着鲁棒的AI生成图像检测迈进,基于位反转图像
Cheng, Renxi, Gui, Jie, Wang, Hongsong
Abstract
The rapid advancement of image generation models has made it increasingly difficult for people to distinguish AI-generated images from real ones. To prevent the potential risks associated with the misuse of fake images, AI-generated image detection has gained significant attention. Existing methods neglect the inherent differences between real and fake images, thus lacking robustness and generalization ability. In this work, we innovatively investigate AI-generated image detection using bit-planes, and introduce the bit-reversed image. We propose a simple yet effective pipeline consisting of construction of bit-reversed images, gradient-based patch selection and a convolutional classifier. Besides, we provide a theoretical analysis from the mathematical perspective to demonstrate the validity of our approach. We also introduce two challenging datasets for AI-generated image detection. Extensive experiments verify the effectiveness of our approach across different settings, including cross-generator generalization, cross-dataset generalization and zero-shot performance. Without bells and whistles, our approach outperforms existing methods on over 40 benchmarks, and is nearly 100 times faster than counterparts. The code is at https://github.com/renxi-seu/RAID.
Chinese Translation
图像生成模型的快速发展使得人们越来越难以区分AI生成的图像和真实图像。为了防止与假图像滥用相关的潜在风险,AI生成图像检测受到了广泛关注。现有方法忽视了真实图像和假图像之间的固有差异,因此缺乏鲁棒性和泛化能力。在本研究中,我们创新性地使用位平面探讨AI生成图像检测,并引入了位反转图像。我们提出了一种简单而有效的流程,包括位反转图像的构建、基于梯度的补丁选择和卷积分类器。此外,我们从数学角度提供了理论分析,以证明我们方法的有效性。我们还引入了两个用于AI生成图像检测的挑战性数据集。大量实验验证了我们方法在不同设置下的有效性,包括跨生成器泛化、跨数据集泛化和零样本性能。我们的方案在40多个基准测试中超越了现有方法,并且速度几乎比同类方法快100倍。代码可在 https://github.com/renxi-seu/RAID 获取。
cs.CV / 18 / 2607.28978

Classification of COVID-19 cases from chest CT volumes using hybrid model of 3D CNN and 3D MLP-Mixer

基于3D CNN和3D MLP-Mixer混合模型的COVID-19病例胸部CT体积分类
Oda, Masahiro, Zheng, Tong, Hayashi, Yuichiro, Otake, Yoshito, Hashimoto, Masahiro, Akashi, Toshiaki, Aoki, Shigeki, Mori, Kensaku
Abstract
This paper proposes an automated classification method of COVID-19 chest CT volumes using improved 3D MLP-Mixer. Novel coronavirus disease 2019 (COVID-19) spreads over the world, causing a large number of infected patients and deaths. Sudden increase in the number of COVID-19 patients causes a manpower shortage in medical institutions. Computer-aided diagnosis (CAD) system provides quick and quantitative diagnosis results. CAD system for COVID-19 enables efficient diagnosis workflow and contributes to reduce such manpower shortage. In image-based diagnosis of viral pneumonia cases including COVID-19, both local and global image features are important because viral pneumonia cause many ground glass opacities and consolidations in large areas in the lung. This paper proposes an automated classification method of chest CT volumes for COVID-19 diagnosis assistance. MLP-Mixer is a recent method of image classification using Vision Transformer-like architecture. It performs classification using both local and global image features. To classify 3D CT volumes, we developed a hybrid classification model that consists of both a 3D convolutional neural network (CNN) and a 3D version of the MLP-Mixer. Classification accuracy of the proposed method was evaluated using a dataset that contains 1205 CT volumes and obtained 79.5% of classification accuracy. The accuracy was higher than that of conventional 3D CNN models consists of 3D CNN layers and simple MLP layers.
Chinese Translation
本文提出了一种基于改进的3D MLP-Mixer的COVID-19胸部CT体积自动分类方法。新型冠状病毒肺炎(COVID-19)在全球传播,导致大量感染患者和死亡病例。COVID-19患者数量的突然增加使医疗机构面临人力资源短缺。计算机辅助诊断(CAD)系统提供快速和定量的诊断结果。针对COVID-19的CAD系统能够实现高效的诊断工作流程,并有助于减少人力资源短缺。在包括COVID-19在内的病毒性肺炎病例的图像诊断中,局部和全局图像特征都非常重要,因为病毒性肺炎会在肺部造成大量磨玻璃样浑浊和大面积的实变。本文提出了一种用于COVID-19诊断辅助的胸部CT体积自动分类方法。MLP-Mixer是一种使用类似于视觉变换器(Vision Transformer)架构的最新图像分类方法,它利用局部和全局图像特征进行分类。为了对3D CT体积进行分类,我们开发了一种混合分类模型,该模型由3D卷积神经网络(CNN)和3D版本的MLP-Mixer组成。通过包含1205个CT体积的数据集评估了所提方法的分类准确性,获得了79.5%的分类准确率。该准确率高于由3D CNN层和简单MLP层组成的传统3D CNN模型。
cs.CV / 19 / 2607.28986

Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning

裁决式字幕生成:多智能体对齐评分与共识蒸馏束仲裁用于严格的零样本图像字幕生成
Thanh, Duy Tran, Doan, Thien-Phuc, Nguyen-Vu, Long, Khanh, Ngo Tan Vu
Abstract
Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and frozen pretrained image-text scorers. Existing retrieval-augmented methods score image-text alignment once, at retrieval, then commit the captioner's autoregressive beam under language-model probability alone, leaving the decoder without further visual grounding feedback. Progress has stalled, with no method improving on the strict-regime best since 2024. We propose Adjudicated Captioning, an inference-time multi-agent framework that restores grounding feedback at multiple checkpoints over an unchanged IFCap captioner. First, we install a stronger frozen Retrieval Encoder at the input. Second, between retrieval and decoding we insert a frozen Cross-Attention Verifier that re-ranks the top-9 retrievals to top-5. Third, at the output beam we attach a learned Reranker pairing TriFuse, a multilayer perceptron, with MemAttend, a memory-attended transformer, the pipeline's only learned components; both are trained self-supervised by Borda-consensus distillation across the three frozen scorers, using no paired image-caption labels and no reference captions. Under the inductive headline protocol, with rerankers fit on the disjoint COCO Karpathy validation beam and applied frozen to test, the framework reaches CIDEr 117.6 and SPICE 21.9 on COCO Karpathy, up from 108.0 and 20.3 for IFCap, a +9.6 CIDEr gain, and +7.7 above NES, the strongest synthetic-image-augmented method at 109.9, without retraining the captioner. A training-free fixed-fusion baseline reaches 115.8 CIDEr, so +7.8 of the +9.6 gain comes from the non-learned architectural intervention and the remaining +1.8 from the learned rerankers. The same recipe transfers off-COCO without captioner retraining: +8.1 CIDEr on Flickr30k Karpathy and +5.7 on NoCaps overall.
Chinese Translation
零样本图像字幕生成(ZIC)在字幕生成器训练过程中不依赖配对的图像-字幕监督,而是依赖于仅包含文本的语料库和冻结的预训练图像-文本评分器。现有的增强检索方法在检索时仅对图像-文本对齐进行一次评分,然后仅依赖语言模型概率来确定字幕生成器的自回归束,导致解码器缺乏进一步的视觉基础反馈。自2024年以来,进展停滞,未有方法在严格模式下超越最佳结果。我们提出了裁决式字幕生成,这是一种推理时的多智能体框架,在不改变IFCap字幕生成器的情况下,在多个检查点恢复基础反馈。首先,我们在输入端安装了一个更强大的冻结检索编码器。其次,在检索与解码之间,我们插入一个冻结的交叉注意力验证器,将前9个检索结果重新排序为前5个。第三,在输出束上,我们附加一个学习的重新排序器,将TriFuse(一个多层感知机)与MemAttend(一个记忆注意力变换器)配对,这两个是管道中唯一的学习组件;它们都是通过在三个冻结评分器之间进行Borda共识蒸馏进行自监督训练的,未使用配对的图像-字幕标签和参考字幕。在归纳标题协议下,使用在不相交的COCO Karpathy验证束上拟合的重新排序器并冻结应用于测试,该框架在COCO Karpathy上达到了CIDEr 117.6和SPICE 21.9,相较于IFCap的108.0和20.3,CIDEr提升了9.6,且比最强的合成图像增强方法NES(109.9)高出7.7,而无需重新训练字幕生成器。一个无训练的固定融合基线达到了115.8 CIDEr,因此9.6的增益中有7.8来自于非学习的架构干预,剩余的1.8来自于学习的重新排序器。同样的方法在不重新训练字幕生成器的情况下转移到COCO以外的任务:在Flickr30k Karpathy上CIDEr提升8.1,在NoCaps上总体提升5.7。
cs.CV / 20 / 2607.28991

CAER: Conflict-Aware Evidence Routing with Dual Prefix Experts for Multimodal Large Language Models

CAER:具有双前缀专家的冲突感知证据路由框架用于多模态大型语言模型
Liu, Zixuan, Cai, Juntao, Cai, Xiaoxu, Wang, Haishuai, Bu, Jiajun
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in multimodal understanding and generation. However, when textual inputs conflict with visual evidence, they still suffer from hallucinations and produce responses inconsistent with visual content. Existing approaches mainly rely on decoding strategies, additional training, verification methods, or prompting techniques, but often lack fine-grained conflict localization and conflict-aware generation. In this work, we propose CAER, a backbone-agnostic framework for visual-language conflict detection and conflict-aware generation. CAER introduces a span-grounded evidence router that transforms claim representations into soft textual queries and retrieves corresponding evidence from frozen visual tokens, enabling fine-grained conflict estimation. Furthermore, we design a dual-prefix expert routing mechanism that learns separate experts for visually supported and contradicted inputs, enabling conflict-aware generation through explicit expert selection. Experiments on the public MMMC benchmark and our newly curated AgriConflict dataset demonstrate that CAER effectively detects visual-language conflicts and improves the reliability of open-source MLLMs without updating their backbone parameters.
Chinese Translation
多模态大型语言模型(MLLMs)在多模态理解和生成方面展现了卓越的能力。然而,当文本输入与视觉证据发生冲突时,它们仍然会出现幻觉,并生成与视觉内容不一致的回应。现有方法主要依赖于解码策略、额外训练、验证方法或提示技术,但往往缺乏细粒度的冲突定位和冲突感知生成。在本研究中,我们提出了CAER,一个与骨干网络无关的视觉-语言冲突检测和冲突感知生成框架。CAER引入了一种基于跨度的证据路由器,将主张表示转换为软文本查询,并从冻结的视觉标记中检索相应的证据,从而实现细粒度的冲突估计。此外,我们设计了一种双前缀专家路由机制,为视觉支持和反驳输入学习独立的专家,通过显式专家选择实现冲突感知生成。在公共MMMC基准和我们新创建的AgriConflict数据集上的实验表明,CAER有效地检测视觉-语言冲突,并在不更新其骨干参数的情况下提高了开源MLLMs的可靠性。
cs.CV / 21 / 2607.28996

SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift

SULAND v2:针对无人机/无人地面车辆(UAV/UGV)基础上的地面地雷检测的精细化RGB数据集和深度学习目标检测基准
Lekhak, Sagar, Pulakurthi, Prasanna Reddy, Joshi, Lalit, Bhatta, Ramesh, Ientilucci, Emmett J.
Abstract
RGB imagery offers a practical, low-cost option for Unmanned Aerial/Ground Vehicle (UAV/UGV) survey support in surface-landmine detection, but object detectors remain underexplored in this safety-critical domain. Limited cross-architecture benchmarking and insufficient out-of-distribution (OOD) analysis obscure whether detectors generalize across deployment conditions. This challenge is amplified by the scarcity of public RGB landmine datasets, making SULAND a key benchmark for PFM-1 and PMA-2 detection. However, inspection reveals missing/false annotations, localization errors, inconsistent visibility criteria, visual artifacts, temporal labeling inconsistencies, and an inverted OOD class-ID convention in SULAND. We present SULAND_v2, a refined RGB surface-landmine dataset and benchmark. Preserving original images and splits, we manually revise annotations to ensure completeness, precise localization, label validity, and class consistency. SULAND_v2 contains 33,771 images and 12,433 bounding boxes. We benchmark 35 detector configurations across nine families. Annotation refinement improves YOLOv8 in-distribution (IID) test mAP@50 by 14.6-19.6 percentage points, while fixing the OOD class-ID convention increases mean YOLOv8 OOD mAP@50 by ~25 percentage points. On SULAND_v2, YOLOv12-Small achieves the highest IID mAP@50 (0.908), while RF-DETR-Large yields the strongest OOD performance (0.799 mAP@50, 0.675 recall). Our results demonstrate that high IID accuracy does not guarantee operational readiness. SULAND_v2 provides a reliable benchmark for evaluating domain-shift robustness in RGB-based mine-action survey support.
Chinese Translation
RGB图像为无人机/无人地面车辆(UAV/UGV)在地面地雷检测中的调查支持提供了一种实用且低成本的选择,但在这一安全关键领域,目标检测器仍然未得到充分探索。有限的跨架构基准测试和不足的分布外(OOD)分析使得难以判断检测器在不同部署条件下的泛化能力。由于公共RGB地雷数据集的稀缺,这一挑战更加凸显,使得SULAND成为PFM-1和PMA-2检测的重要基准。然而,检查发现SULAND存在缺失/错误标注、定位错误、不一致的可见性标准、视觉伪影、时间标记不一致以及反向的OOD类ID约定。我们提出了SULAND_v2,一个精细化的RGB地面地雷数据集和基准。在保留原始图像和划分的基础上,我们手动修订标注,以确保完整性、精确定位、标签有效性和类别一致性。SULAND_v2包含33,771张图像和12,433个边界框。我们对九个类别中的35种检测器配置进行了基准测试。标注的精细化使YOLOv8在分布内(IID)测试中的mAP@50提高了14.6-19.6个百分点,而修正OOD类ID约定则使YOLOv8的平均OOD mAP@50提高了约25个百分点。在SULAND_v2上,YOLOv12-Small达到了最高的IID mAP@50(0.908),而RF-DETR-Large则展现了最强的OOD性能(0.799 mAP@50,0.675召回率)。我们的结果表明,高IID准确性并不保证操作准备就绪。SULAND_v2为评估基于RGB的地雷行动调查支持中的领域转移鲁棒性提供了可靠的基准。
cs.CV / 22 / 2607.29025

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

一致性多参考图像编辑的评估-验证奖励
Miao, Yingmao, Zhang, Pengfei, Lv, Xiaochen, Yu, Meng, Sun, Lei, Chu, Xiangxiang, Shen, Chao, Lin, Chenhao
Abstract
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual harmony. Reinforcement learning has proven highly effective for text-to-image generation and single-image editing, but its extension to multi-reference editing is hindered by the absence of suitable reward models that capture multi-image relational constraints. Moreover, naively using multimodal large language models(MLLMs) as zero-shot evaluators faces a key tension between hallucination-prone long-form reasoning and the limited deductive power of short-form judgments. We address these issues with a Multi-dimensional Evaluation-Verification Reward(EVR). EVR decomposes evaluation into distinct visual criteria; for each criterion, an MLLM Evaluator generates multiple candidate hypotheses, and a Verifier grounds each claim in concrete visual evidence to accept or reject it, producing reliable and fine-grained reward signals. Together with a scalable data pipeline, our method enables RL fine-tuning of off-the-shelf editors without architectural changes. Extensive experiments show substantial gains over the base Qwen-Image-Edit, improving consistency and harmony to match or surpass NanoBanana.
Chinese Translation
尽管近期图像编辑模型取得了迅速进展,但多参考编辑仍然具有挑战性,特别是在保持参考之间的视觉一致性和确保整体视觉和谐方面。强化学习在文本到图像生成和单图像编辑中已被证明非常有效,但其在多参考编辑中的扩展受到缺乏适当奖励模型的阻碍,这些模型能够捕捉多图像的关系约束。此外,天真地使用多模态大型语言模型(MLLMs)作为零-shot 评估者面临着一个关键的张力,即容易产生幻觉的长形式推理与短形式判断的有限推理能力之间的矛盾。我们通过多维评估-验证奖励(EVR)来解决这些问题。EVR将评估分解为不同的视觉标准;对于每个标准,MLLM 评估器生成多个候选假设,而验证者则将每个主张与具体的视觉证据相结合,以接受或拒绝该主张,从而产生可靠且细致的奖励信号。结合可扩展的数据管道,我们的方法使得无需架构更改即可对现成的编辑器进行强化学习微调。大量实验表明,与基础的 Qwen-Image-Edit 相比,我们的方法在一致性和和谐性方面有显著提升,达到或超过 NanoBanana 的水平。
cs.CV / 23 / 2607.29033

SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting

SAM+D:通过深度路由LoRA和深度位移实现SAM家族模型的参数高效维度提升
Song, Yu, Sun, Hao, Teng, Shiyu, Nishikawa, Ikuko, Chen, Yen-wei
Abstract
Existing methods for adapting 2D foundation models such as SAM to 3D volumes either process slices independently---ignoring inter-slice context---or require substantial architectural changes and retraining. In this paper, we present \textbf{SAM+D}, a parameter-efficient framework that lifts SAM-family models by one spatial dimension---enabling 3D volumetric segmentation from 2D SAM and, for the first time via parameter-efficient fine-tuning, end-to-end 4D (3D+T) spatiotemporal segmentation from video-based SAM2---while keeping the vast majority of pre-trained parameters frozen. SAM+D introduces two lightweight, model-agnostic modules into frozen transformer blocks: (1)~\textbf{Depth-Routed LoRA (DRLoRA)} experts with learned routing for spatially adaptive low-rank updates, and (2)~\textbf{Depth Shift Modules (DSM)} for cross-slice feature exchange at zero additional parameter cost. Together, they provide volume-level context while tuning only ${\sim}$2.8\% of parameters for SAM and ${\sim}$3.7\% for SAM2. We evaluate SAM+D in two distinct settings, each lifting the base model by one spatial dimension: 3D segmentation, where SAM(2D$\,\to\,$3D) is evaluated on four CT benchmarks (KiTS, Pancreas, LiTS, Colon), and 4D segmentation, where SAM2 (2D+T$\,\to\,$3D+T) is evaluated on a cell tracking challenge (CTC) dataset (Fluo-N3DH-SIM+). In both settings SAM+D achieves competitive or superior results under the single-point prompt setting while using fewer trainable parameters than existing methods, demonstrating that SAM+D generalizes across SAM-family architectures, target dimensionalities (3D, 4D), and domains spanning medical imaging and bio-scene understanding. Code is publicly available at https://github.com/JerrySongCST/SAM-Plus-D.
Chinese Translation
现有的方法将2D基础模型(如SAM)适配到3D体积时,要么独立处理切片——忽视切片间的上下文——要么需要进行大量的架构更改和重新训练。在本文中,我们提出了 extbf{SAM+D},一个参数高效的框架,通过一个空间维度提升SAM家族模型——使得从2D SAM进行3D体积分割成为可能,并且首次通过参数高效的微调,实现了基于视频的SAM2的端到端4D(3D+T)时空分割,同时保持绝大多数预训练参数不变。SAM+D在冻结的变换器块中引入了两个轻量级、模型无关的模块:(1)具有学习路由的空间自适应低秩更新的 extbf{深度路由LoRA(DRLoRA)}专家,以及(2)用于跨切片特征交换的 extbf{深度位移模块(DSM)},且不增加额外的参数成本。它们共同提供了体积级别的上下文,同时仅调优约${ extsim}2.8 ext{%}$的SAM参数和约${ extsim}3.7 ext{%}$的SAM2参数。我们在两个不同的设置中评估了SAM+D,每个设置将基础模型提升一个空间维度:3D分割,其中SAM(2D$ o $3D)在四个CT基准(KiTS、胰腺、LiTS、结肠)上进行评估,以及4D分割,其中SAM2(2D+T$ o $3D+T)在细胞追踪挑战(CTC)数据集(Fluo-N3DH-SIM+)上进行评估。在这两种设置中,SAM+D在单点提示设置下实现了具有竞争力或优越的结果,同时使用的可训练参数少于现有方法,证明了SAM+D在SAM家族架构、目标维度(3D、4D)和涵盖医学成像与生物场景理解的领域中的广泛适应性。代码已公开发布在https://github.com/JerrySongCST/SAM-Plus-D。
cs.CV / 24 / 2607.29037

GO-PRE: Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D Reconstruction

GO-PRE:通过预测渲染熵进行目标导向的下一个最佳视图选择,用于主动三维重建
Song, Yan, Li, Zhihao, Li, Chenglong, He, Li, Wang, Yan, Zhang, Wenqiang
Abstract
Active 3D reconstruction relies on active view selection to maximize reconstruction fidelity under limited capture budgets. However, most existing methods rely on surrogate signals such as parameter uncertainty or geometric heuristics, but these signals are often misaligned with the ultimate goal: the fidelity of rendered predictions. We propose GO-PRE, a goal-oriented next-best-view selection framework that explicitly targets information gain in the prediction space. Specifically, we formulate the objective as maximizing the reduction of the average marginal predictive entropy over a user-specified target view manifold. GO-PRE supports interactive goal specification and yields an efficient acquisition rule that enables real-time computation of information gain. Extensive experiments across benchmarks demonstrate that GO-PRE consistently improves active reconstruction performance and provides more reliable uncertainty quantification compared to state-of-the-art methods.
Chinese Translation
主动三维重建依赖于主动视图选择,以在有限的捕获预算下最大化重建的保真度。然而,大多数现有方法依赖于代理信号,例如参数不确定性或几何启发式,但这些信号往往与最终目标——渲染预测的保真度——不一致。我们提出了GO-PRE,一个目标导向的下一个最佳视图选择框架,明确针对预测空间中的信息增益。具体而言,我们将目标公式化为最大化用户指定的目标视图流形上的平均边际预测熵的减少。GO-PRE支持交互式目标指定,并产生一个高效的获取规则,使得信息增益的实时计算成为可能。广泛的基准实验表明,GO-PRE在主动重建性能上始终优于现有的最先进方法,并提供了更可靠的不确定性量化。
cs.CV / 25 / 2607.29039

ReMoE: Report-Guided Mixture-of-Experts for Multimodal OCT/OCTA Anomaly Detection

ReMoE:基于报告引导的多专家混合模型用于多模态OCT/OCTA异常检测
Nie, Zihan, Qiao, Qincheng, Xu, Muhao, Feng, Wei, Hou, Xinguo, Song, Weiye, Ge, Zongyuan
Abstract
Multimodal medical anomaly detection identifies samples deviating from normal patterns, where scarce abnormal cases make normality modeling from normal data practical. In retinal Optical Coherence Tomography (OCT) and OCT Angiography (OCTA) anomaly detection, existing unsupervised methods rely on visual feature distributions, reconstruction residuals, or encoder-decoder discrepancies, making anomaly scores depend on appearance-level deviations, while multimodal normality also contains semantic organization described in normal medical reports. To this end, we propose Report-Guided Mixture-of-Experts (ReMoE), which distills normal report semantics into an image-to-text prior student, builds modality-aware priors, and uses Report-Guided Modality Modulation (RMM) to modulate features through mixture-of-experts routing. Experiments on a private OCT/OCTA dataset with paired normal reports and a public OCTA500-3MM setting using a fixed normal report demonstrate state-of-the-art performance.
Chinese Translation
多模态医学异常检测识别偏离正常模式的样本,其中稀缺的异常案例使得从正常数据建模正常性变得可行。在视网膜光学相干断层扫描(OCT)和OCT血管成像(OCTA)异常检测中,现有的无监督方法依赖于视觉特征分布、重建残差或编码器-解码器差异,使得异常评分依赖于外观层面的偏差,而多模态正常性还包含在正常医学报告中描述的语义组织。为此,我们提出了基于报告引导的多专家混合模型(ReMoE),该模型将正常报告的语义提炼为图像到文本的先验学生,构建模态感知先验,并使用报告引导的模态调制(RMM)通过多专家路由调制特征。在一个具有配对正常报告的私有OCT/OCTA数据集和使用固定正常报告的公共OCTA500-3MM设置上的实验表明,该方法达到了最先进的性能。
cs.CV / 26 / 2607.29040

Rethinking Detection Calibration: A Coordinate and Direction Perspective

重新思考检测校准:一种坐标和方向的视角
Lee, Juyong, Jung, Seungjin, Lee, Jungmin, Lee, Sunju, Choi, Jongwon
Abstract
Deep learning based object detectors require trustworthiness beyond competitive detection performance, but deep neural networks are prone to overconfident predictions, assigning high confidence scores to predictions that are likely to be inaccurate. To improve the alignment between confidence scores and prediction accuracy, existing methods calibrate confidence scores based on box-level localization, such as precision or intersection over union with the ground truth bounding box. However, box-level localization reflects only a measure of agreement between the predicted box and the ground truth, resulting in calibrated confidence scores for box-level accuracy failing to capture the localization accuracy of coordinates of box. To tackle this issue, we propose a novel post-hoc calibration framework, rethinking detection calibration (ReDC), which provides reliable coordinate-level confidence scores, including directional information. The proposed framework defines coordinate-wise alignment and deviation direction between predictions and ground truth. Based on the alignment measure, confidence re-encoding produces reliable coordinate-level confidence scores, while directional displacement estimation predicts coordinate-wise deviation directions. Extensive experiments under in-domain and out-domain scenarios demonstrate that the proposed approach expresses the coordinate-wise localization of detected objects more precisely than existing methods. Furthermore, our method covers the representational scope of prior calibration approaches by aggregating coordinate-level confidence scores into box-level localization.
Chinese Translation
基于深度学习的目标检测器需要超越竞争性检测性能的可信度,但深度神经网络容易产生过于自信的预测,给可能不准确的预测分配高置信度分数。为了改善置信度分数与预测准确性之间的对齐,现有方法基于框级定位对置信度分数进行校准,例如与真实边界框的精度或交并比。然而,框级定位仅反映预测框与真实值之间的一致性度量,导致针对框级准确性的校准置信度分数未能捕捉框坐标的定位准确性。为了解决这个问题,我们提出了一种新颖的后处理校准框架——重新思考检测校准(ReDC),该框架提供可靠的坐标级置信度分数,包括方向信息。所提框架定义了预测与真实值之间的坐标级对齐和偏差方向。基于对齐度量,置信度重新编码生成可靠的坐标级置信度分数,而方向位移估计则预测坐标级偏差方向。在领域内和领域外场景下的广泛实验表明,所提方法比现有方法更精确地表达了检测对象的坐标级定位。此外,我们的方法通过将坐标级置信度分数聚合到框级定位中,覆盖了先前校准方法的表示范围。
cs.CV / 27 / 2607.29045

Adaptive Emotional Video Captioning via Affective Heterogeneous Graph Reasoning and Multi-task Joint Learning

通过情感异构图推理和多任务联合学习的自适应情感视频字幕生成
Wang, Junbo, Fu, Liangyu, Li, Yuke, Wu, Xuecheng, Wang, Zhiyong
Abstract
Emotional video captioning (EVC) aims to describe a video with both factual correctness and affective expressiveness. It requires a model to perceive subtle, ambiguous, and temporally varying emotional cues and translate them into natural language without weakening objective visual content. Existing methods have progressively introduced contextual attention, emotion interpretation, emotion priors, dynamic emotion perception and emotion-cause reasoning. Nevertheless, most of them still depend on either global emotion vectors or rigid hierarchical priors. In recent methods, the tree-structured emotion prior establishes a coarse-to-fine connection between psychological emotion categories and daily emotion words, but its hard subordinate masking may irreversibly suppress correct lexical emotions once the coarse category prediction is inaccurate. It is also limited in representing mixed or overlapping emotions that frequently occur in real videos. To address the issues, we propose SAGML, an adaptive EVC framework via affective heterogeneous graph and multi-task language modeling. Instead of treating the emotion prior as a discrete tree, SAGML constructs a soft affective heterogeneous graph containing catalog-level emotion nodes and lexical-level emotion word nodes. The soft gate is injected into video-to-emotion graph attention as a continuous bias, allowing visually supported lexical emotions to remain recoverable rather than being removed by a hard mask. The resulting affective representation is fed together with visual tokens into a causal language decoder, while dual catalog and lexical heads impose explicit emotion distribution learning on the prompt hidden states. The overall model is trained with a joint objective that combines autoregressive caption generation and emotion distribution supervision. SAGML provides an error-resilient and multi-emotion-aware baseline for EVC.
Chinese Translation
情感视频字幕生成(EVC)旨在以事实准确性和情感表现力描述视频。这要求模型能够感知微妙、模糊和时间变化的情感线索,并将其转化为自然语言,而不削弱客观视觉内容。现有方法逐步引入了上下文注意力、情感解释、情感先验、动态情感感知和情感-原因推理。然而,它们大多数仍然依赖于全局情感向量或僵化的层次先验。在最近的方法中,树状情感先验建立了心理情感类别与日常情感词之间的粗到细连接,但其硬性从属屏蔽可能在粗类别预测不准确时不可逆地抑制正确的词汇情感。此外,它在表示真实视频中经常出现的混合或重叠情感方面也存在局限。为了解决这些问题,我们提出了SAGML,一个通过情感异构图和多任务语言建模的自适应EVC框架。SAGML并不将情感先验视为离散树,而是构建了一个包含目录级情感节点和词汇级情感词节点的软性情感异构图。软门控被注入到视频-情感图注意力中,作为一种连续偏置,使得视觉支持的词汇情感能够保持可恢复性,而不是被硬性屏蔽移除。生成的情感表示与视觉标记一起输入到因果语言解码器中,同时双重目录和词汇头对提示隐藏状态施加显式情感分布学习。整个模型通过结合自回归字幕生成和情感分布监督的联合目标进行训练。SAGML为EVC提供了一个抗误差和多情感感知的基线。
cs.CV / 28 / 2607.29048

Parameter-Efficient Fine-Tuning for Spiking Point Cloud Models

高效参数微调的脉冲点云模型
Guo, Zihao, Zhu, Jihua, Sun, Yiding, Chen, Lin, Wang, Danwei
Abstract
Spiking Neural Networks (SNNs) offer energy-efficient solutions for point cloud analysis on resource-constrained devices through event-driven computation. However, existing pre-trained spiking point cloud models rely on full fine-tuning for downstream task adaptation, incurring substantial parameter and storage overhead. Furthermore, binary spike propagation suppresses task-relevant sub-threshold information. To address these issues, we propose SpikePEFT, the first parameter-efficient fine-tuning framework for spiking point cloud models. Specifically, Intrinsic Dynamics Tuning (IDT) adaptively modulates membrane decay and firing thresholds, enabling efficient neuron-intrinsic adaptation while keeping the pre-trained synaptic transformations frozen. Moreover, Silent-State Disambiguation Adaptation (SSDA) recovers task-relevant information from informative silent states, thereby providing richer evidence for downstream adaptation. Extensive experiments across multiple benchmarks demonstrate the effectiveness and efficiency of SpikePEFT. In particular, our method achieves 92.4% accuracy on ModelNet40 and 85.6\% on the most challenging classification split ScanObjectNN(PB\_T50\_RS) while updating only about 5% of the trainable parameters and preserving the energy efficiency of SNNs. This work provides a promising step toward parameter-efficient adaptation of neuromorphic vision models.
Chinese Translation
脉冲神经网络(SNNs)通过事件驱动计算为资源受限设备上的点云分析提供了节能高效的解决方案。然而,现有的预训练脉冲点云模型依赖于全面微调以适应下游任务,这导致了显著的参数和存储开销。此外,二进制脉冲传播抑制了与任务相关的亚阈值信息。为了解决这些问题,我们提出了SpikePEFT,这是第一个针对脉冲点云模型的高效参数微调框架。具体而言,内在动态调节(Intrinsic Dynamics Tuning, IDT)自适应地调节膜衰减和放电阈值,使神经元内在适应变得高效,同时保持预训练的突触变换不变。此外,静默状态消歧适应(Silent-State Disambiguation Adaptation, SSDA)从信息丰富的静默状态中恢复与任务相关的信息,从而为下游适应提供更丰富的证据。在多个基准测试中的广泛实验表明,SpikePEFT的有效性和高效性。特别是,我们的方法在ModelNet40上达到了92.4%的准确率,在最具挑战性的分类分割ScanObjectNN(PB_T50_RS)上达到了85.6%的准确率,同时仅更新了约5%的可训练参数,并保持了SNNs的能效。该研究为神经形态视觉模型的高效参数适应提供了一个有前景的步骤。
cs.CV / 29 / 2607.29059

Learning from Adversity: Semantic-Aware Mask Refinement through Adversarial Perturbation

从逆境中学习:通过对抗扰动进行语义感知的掩膜精炼
Kim, Beomyoung, Hwang, Sung Ju
Abstract
Despite significant advances in image segmentation, even state-of-the-art models produce masks with imperfect boundaries, semantic inconsistencies, and structural errors. Mask refinement addresses these limitations, yet current approaches rely on simplistic synthetic noise that fails to capture the complex error patterns of real segmentation models. We introduce Phoenix, a novel framework that leverages adversarial learning to generate semantically meaningful noise patterns and contrastive learning to model refinement relationships. Our approach consists of two key innovations: (1) Adversarial Mask Perturbation, which employs embedding attacks to create semantic-aware noise that mimics real segmentation errors, and (2) Contrastive Mask Refinement Learning, which establishes a tri-directional framework that ensures feature consistency within semantic regions while maintaining separation between classes. Experiments demonstrate that Phoenix significantly outperforms existing methods across diverse tasks, while consistently enhancing state-of-the-art segmentation models with substantial improvements. Our code and project page are publicly available at https://phoenix-eccv26.github.io.
Chinese Translation
尽管图像分割技术取得了显著进展,甚至是最先进的模型也会产生边界不完美、语义不一致和结构错误的掩膜。掩膜精炼旨在解决这些局限性,但当前的方法依赖于简单的合成噪声,未能捕捉真实分割模型的复杂错误模式。我们提出了Phoenix,一个新颖的框架,利用对抗学习生成语义上有意义的噪声模式,并通过对比学习建模精炼关系。我们的方法包含两个关键创新:(1)对抗掩膜扰动,采用嵌入攻击生成模仿真实分割错误的语义感知噪声;(2)对比掩膜精炼学习,建立一个三向框架,确保语义区域内的特征一致性,同时保持类别之间的分离。实验表明,Phoenix在多种任务中显著优于现有方法,同时持续增强最先进的分割模型,带来显著改善。我们的代码和项目页面可在 https://phoenix-eccv26.github.io 上公开获取。
cs.CV / 30 / 2607.29083

MHRGait: Gait Recognition from Momentum Human Rig Pose

MHRGait:基于动量人形姿态的步态识别
Duan, Huiran, Zhou, Qian, Guo, Xianda, Zou, Hua, Zhao, Guoying, Wang, Zhongyuan, Tian, Yingli
Abstract
Gait recognition is shaped by its input representation. Silhouettes encode projected body shape, skeletons encode sparse joint coordinates, and 3D meshes encode dense surface geometry. In each case, identity-bearing articulation is observed through geometric carriers that also vary with clothing, skeletal scale, or body shape. We investigate whether gait can instead be recognized from compact articulated controls. We introduce Momentum Human Rig (MHR) pose as a gait representation, describing each frame using 184 semantically organized body and hand parameters estimated from monocular video. MHRGait groups these heterogeneous controls by anatomy, models their intra-frame coordination and temporal evolution, and produces compact body and hand descriptors. We further introduce MHRGait++, which combines MHR pose with silhouettes through modality-balanced distance fusion, preventing descriptor count from determining modality importance. Experiments on four benchmarks show that MHRGait attains the best overall performance among compared model-based methods on CCPG and SUSTech1K and transfers effectively across datasets, while its recognition network requires only 2.76M parameters and 0.69 GFLOPs for a 30-frame input. MHRGait++ consistently improves silhouette recognizers with a favorable accuracy-efficiency trade-off. These results establish rig-space articulation as an effective standalone gait representation and a complementary cue to projected body shape. Our code is available at https://github.com/duanhuiran/MHRGait.
Chinese Translation
步态识别受其输入表示的影响。轮廓编码了投影的身体形状,骨架编码了稀疏的关节坐标,而三维网格编码了密集的表面几何。在每种情况下,通过几何载体观察到具有身份特征的关节动作,这些载体也会因服装、骨骼比例或身体形状而变化。我们研究是否可以通过紧凑的关节控制来识别步态。我们引入动量人形(Momentum Human Rig, MHR)姿态作为步态表示,使用从单目视频中估计的184个语义组织的身体和手部参数描述每一帧。MHRGait根据解剖学对这些异构控制进行分组,建模其帧内协调和时间演变,并生成紧凑的身体和手部描述符。我们进一步引入MHRGait++,它通过模态平衡距离融合将MHR姿态与轮廓结合,防止描述符数量决定模态的重要性。在四个基准测试上的实验表明,MHRGait在CCPG和SUSTech1K上在比较的基于模型的方法中达到了最佳整体性能,并且在数据集之间有效迁移,同时其识别网络仅需要2.76M参数和0.69 GFLOPs用于30帧输入。MHRGait++始终改善轮廓识别器,具有良好的准确性与效率的权衡。这些结果确立了rig空间关节动作作为一种有效的独立步态表示,以及对投影身体形状的补充线索。我们的代码可在https://github.com/duanhuiran/MHRGait获取。
cs.CV / 31 / 2607.29106

Forwardrobe: Garment-Aware Gaussian Avatars from a Single Image

Forwardrobe:基于单幅图像的服装感知高斯头像重建
Jin, Daisheng, Wang, Shuyun, He, Ying
Abstract
Reconstructing animatable 3D human avatars from a single image remains particularly challenging for loose garments, whose geometry and motion cannot be adequately represented by body-aligned topology and skinning. We present Forwardrobe, a feed-forward framework for reconstructing garment-aware Gaussian avatars from a single image. Forwardrobe explicitly separates clothing from the body in canonical Gaussian space and equips the garment layer with continuity-aware geometry and skinning initialization, pose-conditioned non-rigid deformation, and appearance adaptation. These designs improve garment reconstruction and visual quality during animation, particularly for skirts and dresses. The separated garment layer additionally forms an independently controllable 3D asset, enabling garment editing, transfer, and 3D virtual try-on. Experiments demonstrate improved garment reconstruction quality and greater flexibility in garment manipulation compared with existing single-image avatar reconstruction methods.
Chinese Translation
从单幅图像重建可动画的三维人类头像对于松散服装仍然特别具有挑战性,因为其几何形状和运动无法通过与身体对齐的拓扑结构和蒙皮进行充分表示。我们提出了Forwardrobe,一个用于从单幅图像重建服装感知高斯头像的前馈框架。Forwardrobe在标准高斯空间中明确地将服装与身体分离,并为服装层提供了连续性感知的几何形状和蒙皮初始化、姿态条件下的非刚性变形以及外观适应。这些设计在动画过程中改善了服装重建和视觉质量,特别是对于裙子和连衣裙。分离的服装层还形成了一个可独立控制的三维资产,使得服装编辑、转移和三维虚拟试穿成为可能。实验表明,与现有的单幅图像头像重建方法相比,服装重建质量得到了改善,服装操作的灵活性也更高。
cs.CV / 32 / 2607.29122

A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples

一个冻结的像素空间扩散模型可以通过自身样本进行自我引导
Fu, Zixuan, Wang, Chong, Guo, Lanqing, Zhou, Kailai, Nie, Jiahao, Wen, Bihan
Abstract
Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging because a single model must capture both global structure and local texture in the same high-dimensional space. While recent work improves pixel diffusion through alternative prediction targets, training objectives, and architectures, these advances typically require training a new model from scratch. We show there is a cheaper, complementary strategy: \textbf{a frozen, pretrained pixel diffusion model can guide itself}. Our key observation is that intermediate layers of a pretrained pixel diffusion transformer can be decoded into coarse predictions that capture the main low-frequency structure, while the final layers progressively refine local, high-frequency details. We therefore attach a lightweight prediction head to an intermediate layer, keep the backbone frozen, and use the discrepancy between the intermediate and final predictions as a self-guidance direction during sampling. To train this head, we further find that real images are not necessary. Instead, model-generated samples suffice and even outperform real images for training the head, especially in enhancing the high-frequency components that pixel diffusion tends to underfit. Across multiple pixel diffusion models on ImageNet, our \textbf{Synthetic Self-Guidance (SSG)} consistently improves generation while adapter training requires less than 1$\%$ of full-model training compute: it reduces FID by over 50$\%$ across the evaluated JiT variants without classifier-free guidance (CFG) and further improves strong baselines with CFG, e.g., JiT-H/16 from 1.86 to 1.67 and PixelREPA-H/16 from 1.81 to 1.59. Our code is available at https://github.com/zfu006/SSG.
Chinese Translation
像素空间扩散模型旨在直接在原始像素上学习端到端生成器。这是一个具有挑战性的任务,因为单个模型必须在同一高维空间中同时捕捉全局结构和局部纹理。尽管近期的研究通过替代预测目标、训练目标和架构改善了像素扩散,但这些进展通常需要从头开始训练一个新模型。我们展示了一种更便宜的互补策略: extbf{一个冻结的、预训练的像素扩散模型可以自我引导}。我们的关键观察是,预训练像素扩散变换器的中间层可以解码为捕捉主要低频结构的粗略预测,而最终层则逐渐细化局部的高频细节。因此,我们将一个轻量级预测头附加到中间层,保持主干不变,并在采样过程中使用中间预测和最终预测之间的差异作为自我引导方向。为了训练这个预测头,我们进一步发现真实图像并不是必需的。相反,模型生成的样本就足够了,甚至在训练预测头时优于真实图像,特别是在增强像素扩散往往欠拟合的高频成分方面。在多个像素扩散模型在ImageNet上的实验中,我们的 extbf{合成自我引导(SSG)}始终改善生成效果,而适配器训练所需的计算量不到全模型训练的1$ ext{%}$:在评估的无分类器引导(CFG)JiT变体中,FID减少超过50$ ext{%}$,并进一步改善了强基线与CFG的表现,例如,JiT-H/16从1.86降至1.67,PixelREPA-H/16从1.81降至1.59。我们的代码可在 https://github.com/zfu006/SSG 获取。
cs.CV / 33 / 2607.29124

SciFigPlag-Bench: A Benchmark for Provenance-Aware Scientific Figure Plagiarism Detection

SciFigPlag-Bench:一个关注来源的科学图形抄袭检测基准
Cui, Zhiying, Yang, Minghao, Gao, Linlin, Liu, Jie, Li, Pengyuan
Abstract
Scientific figures often encode the visual evidence behind scientific findings, yet figure plagiarism remains underexplored as a benchmarked multimodal evaluation problem. We present SciFigPlag-Bench, a benchmark for provenance-aware reasoning over scientific figures in scholarly documents. Unlike general image-similarity or image-forensics benchmarks, SciFigPlag-Bench evaluates whether a suspicious figure reuses evidence from a specific source figure, how the reused content has been transformed, and where the reused evidence appears. We introduce a factorized taxonomy that separates what is reused from how it is transformed, covering material-preserving reuse, such as full-figure and subfigure reuse, as well as abstract-content reuse, such as data re-expression and structural redraw. Guided by this taxonomy, we construct a hybrid benchmark with 2,582 positive pairs and 2,541 negative pairs, combining documented real-world cases, taxonomy-guided synthetic examples, and visually similar negatives. The benchmark supports four diagnostic tasks: pairwise detection, source attribution, hierarchical reuse-type classification, and reuse correspondence localization. Experiments with diverse vision-language models establish initial baselines and reveal persistent challenges in fine-grained provenance reasoning, reuse-type understanding, and spatial evidence grounding.
Chinese Translation
科学图形通常编码了科学发现背后的视觉证据,但图形抄袭作为一个基准化的多模态评估问题仍然未得到充分探索。我们提出了SciFigPlag-Bench,这是一个针对学术文献中科学图形的关注来源推理的基准。与一般的图像相似性或图像取证基准不同,SciFigPlag-Bench评估可疑图形是否重用了特定源图形的证据、重用内容如何被转化以及重用证据出现的位置。我们引入了一种分解的分类法,将重用的内容与其转化方式区分开,涵盖了材料保持重用,如完整图形和子图重用,以及抽象内容重用,如数据重新表达和结构重绘。在这一分类法的指导下,我们构建了一个混合基准,包含2,582对正样本和2,541对负样本,结合了记录的真实案例、分类法指导的合成示例和视觉相似的负样本。该基准支持四个诊断任务:成对检测、源归属、层次重用类型分类和重用对应定位。对多种视觉-语言模型的实验建立了初步基线,并揭示了在细粒度来源推理、重用类型理解和空间证据定位方面的持续挑战。
cs.CV / 34 / 2607.29132

First Investigation of Deep Learning for Intraoperative Gauze Segmentation in Minimally Invasive Abdominal Surgery

首次研究深度学习在微创腹部手术中对术中纱布分割的应用
Tomar, Priya, Broß, Maximilian, Feodorovici, Philipp, Arensmeyer, Jan, Leifels, Philipp, Parikh, Aditya, Matthaei, Hanno, Bauckhage, Christian, Schneider, Helen, Sifa, Rafet
Abstract
Surgical gauze is an essential part of surgical procedures, primarily used for controlling bleeding and absorbing bodily fluids. The post-surgical retention of gauze can lead to serious complications and necessitate additional surgery for its removal. Despite the clinical significance, research on gauze segmentation using real-world surgical data remains underexplored, owing in part to the scarcity of annotated datasets. In this work, we investigate the use of deep learning methods for gauze segmentation in robot-assisted minimally invasive abdominal surgeries, utilizing an in-house surgical dataset prepared at a university hospital. The training data reflects realistic surgical settings and captures extensive diversity in spatial, morphological, and visual attributes across three different gauze categories. We evaluate several widely used segmentation architectures, including CNN-based, transformer-based, and hybrid architectures, to establish a proof-of-concept for gauze segmentation in a realistic clinical setting. In addition, we investigate the influence of sub-optimally annotated, auto-tracked segmentation masks as a strategy to address data scarcity and improve performance. Our results demonstrate the efficacy of real-world training data in countering the main challenge reported by prior works, the trade-off between blood presence and gauze detection. The incorporation of auto-tracked annotations yields performance enhancements, particularly in generic surgical scenarios. The integration of effective segmentation approaches can benefit robot-guided surgical procedures and various downstream applications by providing precise delineation of foreign objects, thereby enhancing patient safety and surgical outcomes.
Chinese Translation
手术纱布是外科手术的重要组成部分,主要用于控制出血和吸收体液。术后纱布滞留可能导致严重并发症,并需要额外手术进行移除。尽管其临床重要性,基于真实手术数据的纱布分割研究仍然未得到充分探索,部分原因在于标注数据集的稀缺。在本研究中,我们探讨了深度学习方法在机器人辅助微创腹部手术中进行纱布分割的应用,利用在大学医院准备的内部手术数据集。训练数据反映了真实的手术环境,并捕捉了三种不同纱布类别在空间、形态和视觉特征上的广泛多样性。我们评估了几种广泛使用的分割架构,包括基于卷积神经网络(CNN)、基于变换器(transformer)和混合架构,以建立在真实临床环境中进行纱布分割的概念验证。此外,我们还研究了次优标注的自动跟踪分割掩膜的影响,作为应对数据稀缺和提高性能的策略。我们的结果表明,真实世界的训练数据在应对先前研究报告的主要挑战——血液存在与纱布检测之间的权衡方面,具有有效性。引入自动跟踪标注显著提升了性能,尤其是在通用手术场景中。有效的分割方法的整合可以通过提供对异物的精确划分,促进机器人引导的手术程序及各种下游应用,从而提高患者安全性和手术结果。
cs.CV / 35 / 2607.29136

On the Efficacy of Self-Supervised Point Cloud Encoders for Efficient 3D Large Language Models

自监督点云编码器在高效3D大型语言模型中的有效性研究
Zheng, Yao, Zhang, Tian
Abstract
3D point cloud-language models (3D-LLMs) enable 3D understanding by pairing point cloud encoders with large language models, but existing methods rely on costly multi-modal encoders (e.g., ULIP-2) that require image-text-point cloud alignment on 8x A100-scale compute, creating high barriers for research and deployment. In this work, we systematically investigate whether low-cost self-supervised point cloud encoders, specifically PCP-MAE and Point-MAE, can serve as effective alternatives. Using MiniGPT-3D as our testbed, we evaluate 7 encoder initialization/pre-training setups (1 multi-modal baseline, 5 self-supervised, 1 random init) under frozen and unfrozen fine-tuning (12 total groups), across 2 architectures (MaskTransformer, PointTransformer), 3 objectives (PCP-MAE, Point-MAE, random init), and 2 datasets (Objaverse 660K, ShapeNet55-34 approximately 50K). Our experiments reveal three key findings: (1) The four-stage MiniGPT-3D pipeline can effectively train a 3D encoder from random initialization: an end-to-end trained random init encoder reaches 52.50% open-vocabulary accuracy and 44.45 captioning score, approaching top pre-trained variants; (2) Architecture and pre-training objective show strong crossover interaction: PCP-MAE + MaskTransformer achieves 59.00% accuracy (best self-supervised), while Point-MAE + MaskTransformer drops to 46.50%, with the pattern reversed for PointTransformer; (3) Closed-set ModelNet40 classification remains a core weakness of purely geometric encoders, reaching only ~13-18% accuracy vs. ~62% for the multi-modal baseline, even after end-to-end fine-tuning. Our results offer practical guidelines for cost-effective 3D-LLM design and reveal interaction patterns between self-supervised objectives and encoder architectures.
Chinese Translation
3D点云-语言模型(3D-LLMs)通过将点云编码器与大型语言模型相结合,实现了对3D的理解,但现有方法依赖于成本高昂的多模态编码器(例如,ULIP-2),这些编码器需要在8x A100规模的计算资源上进行图像-文本-点云对齐,给研究和部署带来了高门槛。在本研究中,我们系统地探讨了低成本自监督点云编码器,特别是PCP-MAE和Point-MAE,是否可以作为有效的替代方案。我们以MiniGPT-3D作为测试平台,评估了7种编码器初始化/预训练设置(1个多模态基线,5个自监督,1个随机初始化),在冻结和非冻结微调下(共12组),跨越2种架构(MaskTransformer,PointTransformer),3个目标(PCP-MAE,Point-MAE,随机初始化),以及2个数据集(Objaverse 660K,ShapeNet55-34,约50K)。我们的实验揭示了三个关键发现:(1)四阶段的MiniGPT-3D流程可以有效地从随机初始化训练3D编码器:一个端到端训练的随机初始化编码器达到了52.50%的开放词汇准确率和44.45的字幕评分,接近顶级预训练变体;(2)架构和预训练目标之间表现出强烈的交互作用:PCP-MAE + MaskTransformer达到了59.00%的准确率(最佳自监督),而Point-MAE + MaskTransformer下降至46.50%,在PointTransformer中则反转了这一模式;(3)闭集ModelNet40分类仍然是纯几何编码器的核心弱点,准确率仅为约13-18%,而多模态基线为约62%,即使在端到端微调后也是如此。我们的结果为成本效益高的3D-LLM设计提供了实用指南,并揭示了自监督目标与编码器架构之间的交互模式。
cs.CV / 36 / 2607.29144

Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership

我见过你吗?嵌入行为信号的合成面孔数据集成员身份
Borsukiewicz, Paweł, Lunghi, Daniele, Ouédraogo, Wendkûuni C., Klein, Jacques, Bissyandé, Tegawendé F.
Abstract
Synthetic face datasets are increasingly used to reduce privacy exposure and data access constraints in biometric recognition. Yet the generators that produce these datasets are trained on real faces, so synthetic data may still reveal their real source data. We study this risk through a dataset-level membership inference attack that first identifies the synthetic dataset used to train a face recognizer and then infers the real dataset used to train the generator. Across 11 face recognition models, 11 synthetic datasets, and 7 real datasets, the attack recovers the synthetic training dataset in 100% of cases and identifies the generator's source dataset in 54.5% of cases. These results show that synthetic data can retain dataset-level traces of real training data and that privacy-preserving deployment requires stronger leakage mitigation.
Chinese Translation
合成面孔数据集越来越多地被用于减少生物识别中的隐私暴露和数据访问限制。然而,生成这些数据集的生成器是基于真实面孔进行训练的,因此合成数据仍可能揭示其真实的源数据。我们通过一种数据集级别的成员身份推断攻击来研究这一风险,该攻击首先识别用于训练面部识别器的合成数据集,然后推断用于训练生成器的真实数据集。在11个面部识别模型、11个合成数据集和7个真实数据集中,该攻击在100%的情况下成功恢复合成训练数据集,并在54.5%的情况下识别生成器的源数据集。这些结果表明,合成数据可以保留真实训练数据的数据集级别痕迹,而隐私保护的部署需要更强的泄漏缓解措施。
cs.CV / 37 / 2607.29156

Progressive Decision-Making for Localizing Open-Ended AI-Generated Image Forgeries

渐进式决策制定用于定位开放式AI生成的图像伪造
Hou, Jingyi, Chen, Xiaoxia, Zhou, Leyu, Wang, Zhichuang, Liu, Zhijie
Abstract
AI-generated image forgeries are becoming increasingly realistic and difficult to characterize with fixed manipulation patterns. As generative models continue to evolve, it is impractical to expect a localization model to exhaustively learn all possible forgery appearances from large-scale training data alone. Nevertheless, many AI-generated forgeries still leave subtle forensic traces, although these cues are often weak and unevenly reliable across regions. Therefore, robust localization requires not only extracting informative forensic traces, but also making reliable decisions from incomplete and ambiguous evidence. In this paper, we move beyond static one-shot prediction and reformulate final forgery localization as an adaptive sequential decision-updating process, where the localization map is treated as an intermediate state rather than a fixed output. Rather than producing the final mask via one-shot pixel-wise prediction, our method progressively updates the localization state guided by available evidence, uncertainty, and boundary conditions. Specifically, we first transform mesoscopic traces into compact decision evidence via a lightweight decision evidence projector, and then introduce Evidence-Guided Mamba (EG-Mamba) to perform uncertainty- and boundary-aware state updating. This design allows reliable manipulated and background regions to be preserved, while ambiguous regions are cautiously revised according to the available evidence. Extensive experiments on both conventional and AI-generated manipulation benchmarks validate the effectiveness of the proposed method. Notably, even when trained only on conventional manipulation data, our method brings larger gains on unseen AI-generated forgeries, indicating that progressive decision-updating is especially useful for heterogeneous and hard-to-exhaustively-learn manipulation traces.
Chinese Translation
AI生成的图像伪造变得越来越逼真,且难以用固定的操控模式进行特征描述。随着生成模型的不断发展,仅依靠大规模训练数据来全面学习所有可能的伪造外观对于定位模型而言是不切实际的。然而,许多AI生成的伪造仍然留下微妙的取证痕迹,尽管这些线索通常较弱且在不同区域的可靠性不均。因此,稳健的定位不仅需要提取有信息量的取证痕迹,还需要从不完整和模糊的证据中做出可靠的决策。在本文中,我们超越了静态的一次性预测,将最终的伪造定位重新表述为一种自适应的顺序决策更新过程,其中定位图被视为中间状态而非固定输出。我们的方法不是通过一次性像素级预测生成最终掩膜,而是根据可用证据、不确定性和边界条件逐步更新定位状态。具体而言,我们首先通过轻量级决策证据投影器将中观痕迹转化为紧凑的决策证据,然后引入证据引导的Mamba(EG-Mamba)进行不确定性和边界感知的状态更新。这种设计允许可靠的操控区域和背景区域得以保留,同时模糊区域根据可用证据谨慎修正。在常规和AI生成的操控基准上进行的广泛实验验证了所提方法的有效性。值得注意的是,即使仅在常规操控数据上进行训练,我们的方法在未见过的AI生成伪造上也带来了更大的提升,这表明渐进式决策更新对于异质和难以全面学习的操控痕迹尤其有效。
cs.CV / 38 / 2607.29180

MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation

MoRAE:流友好的自监督潜变量用于文本到运动生成
Zhu, Yifei, Shi, Mingyi, Cai, Yangyang, Cheng, Miao, Kitamura, Yoshifumi, Komura, Taku
Abstract
Text-to-motion generation must produce motions that are semantically correct, temporally coherent, and physically plausible. A natural approach is to first project motion data into a structured semantic space and then train a generative model within that space. Such a paradigm has been highly successful in image generation through Representation Autoencoders (RAEs), where a frozen self-supervised encoder provides semantic features for diffusion or flow models to learn from. However, direct transfer of such a paradigm to motion space using Motion-JEPA as the frozen encoder fails dramatically. We diagnose this failure geometrically and identify two motion-specific bottlenecks: (1) the JEPA feature space is spectrally ill-conditioned, making the Gaussian-to-data transport unstable; and (2) even with a well-conditioned spectrum, flow residuals tend to align with decoder-sensitive directions, where small latent errors are amplified into large motion artifacts after decoding. Based on these insights, we propose MoRAE. MoRAE addresses the two bottlenecks separately. A compact bottleneck distills the structured JEPA representation while removing weak and redundant directions, bringing the latent spectrum into a transport-stable regime. Motion-coupled training then aligns the retained latent geometry with the decoder, making characteristic flow errors less costly after decoding. With this flow-friendly latent, a standard non-autoregressive Flow-Matching DiT achieves state-of-the-art performance.
Chinese Translation
文本到运动生成必须产生语义正确、时间连贯和物理合理的运动。一个自然的方法是首先将运动数据投影到一个结构化的语义空间中,然后在该空间内训练生成模型。这种范式在通过表示自编码器(Representation Autoencoders, RAEs)进行图像生成方面取得了很大成功,其中一个冻结的自监督编码器为扩散或流模型提供了语义特征。然而,直接将这种范式转移到运动空间,使用Motion-JEPA作为冻结编码器,结果却大幅失败。我们从几何角度诊断了这一失败,并识别出两个运动特有的瓶颈:(1)JEPA特征空间的谱条件不良,导致高斯到数据的传输不稳定;(2)即使在良好条件的谱下,流残差往往与解码器敏感方向对齐,使得小的潜变量误差在解码后放大为大的运动伪影。基于这些见解,我们提出了MoRAE。MoRAE分别解决这两个瓶颈。一个紧凑的瓶颈提炼了结构化的JEPA表示,同时去除了弱和冗余的方向,将潜变量谱带入一个传输稳定的状态。运动耦合训练则将保留的潜在几何与解码器对齐,使得特征流误差在解码后成本更低。借助这种流友好的潜变量,标准的非自回归流匹配DiT达到了最先进的性能。
cs.CV / 39 / 2607.29192

Locally Consistent Transductive Information Maximization for Few-Shot Remote Sensing Scene Classification

基于局部一致性的传导信息最大化用于少样本遥感场景分类
Khoury, Karim El, Gérin, Benoît, Macq, Benoît, De Vleeschouwer, Christophe
Abstract
Remote sensing scene classification is increasingly relying on foundation models pre-trained on large-scale Earth-observation data. Moreover, transductive inference, which exploits the collective statistical structure of the entire unlabeled query set, appears to naturally match remote sensing pipelines where large images are routinely split into patches and inferred as a batch. In this work, we introduce LC-TIM (Locally Consistent Transductive Information Maximization), which extends the state-of-the-art Transductive Information Maximization for Few-Shot CLIP (TIM++) objective with a local consistency regularizer that enforces prediction agreement between each query sample and its $\kappa$ nearest feature-space neighbors. The regularizer enters as a single multiplicative factor in the closed-form $q$-update, adding negligible computational overhead. We further propose a multi-source extension that fuses the affinity graph from multiple remote sensing foundation model, further boosting classification accuracy. To assess these methods, we establish the first comprehensive, open-source benchmark for transductive few-shot RS scene classification, evaluating LP++, TransCLIP, TIM++, and LC-TIM across ten diverse datasets, two remote sensing vision-language models, and across various few-shot settings. Our experiments show that transductive methods consistently outperform zero-shot baselines, and that LC-TIM achieves state-of-the-art accuracy, with the largest gains in the low-shot regime where neighborhood cues are most informative. Code is publicly available at: https://github.com/elkhouryk/LC-TIM
Chinese Translation
遥感场景分类越来越依赖于在大规模地球观测数据上预训练的基础模型。此外,传导推理利用整个未标记查询集的集体统计结构,似乎自然地与遥感管道相匹配,在这些管道中,大图像通常被分割成小块并作为批次进行推理。在本研究中,我们引入了LC-TIM(局部一致性传导信息最大化),该方法扩展了最先进的少样本CLIP(TIM++)的传导信息最大化目标,增加了一个局部一致性正则化项,以强制每个查询样本与其$ ext{kappa}$个最近特征空间邻居之间的预测一致性。该正则化项作为一个单一的乘法因子进入闭式$q$-更新,增加的计算开销微乎其微。我们进一步提出了一种多源扩展,融合来自多个遥感基础模型的亲和图,进一步提高分类准确性。为了评估这些方法,我们建立了第一个全面的开源基准,用于传导少样本遥感场景分类,评估LP++、TransCLIP、TIM++和LC-TIM在十个不同数据集、两个遥感视觉语言模型以及各种少样本设置下的表现。我们的实验表明,传导方法始终优于零样本基线,而LC-TIM实现了最先进的准确性,尤其在邻域线索最具信息性的低样本情况下取得了最大的提升。代码已公开发布在:https://github.com/elkhouryk/LC-TIM
cs.CV / 40 / 2607.29200

UltraSAM3: A Concept-Driven Foundation Model for Universal Ultrasound Image Segmentation

UltraSAM3:一种基于概念驱动的通用超声图像分割基础模型
Xu, Bo, Zhu, Quanhao, Lin, Rui, Zhu, Boling, Wang, Chenyuan, Lin, Hongfei, Xia, Feng, Ji, Chenhua
Abstract
Ultrasound imaging has become increasingly widespread in clinical practice due to its portability, low cost and real-time capability, making ultrasound image segmentation important. However, ultrasound images differ substantially from CT, MRI, and other medical imaging modalities, as they are often affected by speckle noise, low contrast, acoustic shadows and ambiguous boundaries. Existing ultrasound segmentation methods are still mainly limited to task-specific models or visual-prompt-based foundation models, which are either tailored to particular tasks or require expert-provided visual prompts, making them inconvenient for flexible clinical use. To address these challenges, we propose UltraSAM3, a concept-driven foundation model for universal ultrasound image segmentation. Unlike conventional models, UltraSAM3 enables text-based target specification by adapting SAM3 to ultrasound-specific image--mask--concept triplets. The model is trained on a large-scale ultrasound segmentation corpus covering 37 public datasets and 13 anatomical categories, allowing it to align ultrasound visual patterns with clinically meaningful concepts across diverse organs and lesions. To further improve usability under realistic clinical interaction, we propose an instruction-guided agent that parses complex natural language queries into concise ultrasound concept prompts for UltraSAM3. Extensive experiments demonstrate that UltraSAM3 consistently outperforms representative concept- and text-driven biomedical segmentation models on multi-organ ultrasound benchmarks, external datasets, and visual-prompt-enhanced settings. Moreover, the agent improves segmentation robustness for complex user instructions. These results indicate that ultrasound-specific concept adaptation is effective for building generalizable and interactive ultrasound segmentation foundation models.
Chinese Translation
超声成像因其便携性、低成本和实时能力在临床实践中变得越来越普遍,使得超声图像分割变得重要。然而,超声图像与CT、MRI及其他医学成像方式有显著不同,常常受到斑点噪声、低对比度、声影和模糊边界的影响。现有的超声分割方法仍主要局限于特定任务模型或基于视觉提示的基础模型,这些模型要么针对特定任务量身定制,要么需要专家提供的视觉提示,导致其在灵活的临床应用中不够便利。为了解决这些挑战,我们提出了UltraSAM3,一种用于通用超声图像分割的基于概念驱动的基础模型。与传统模型不同,UltraSAM3通过将SAM3适配到超声特定的图像-掩膜-概念三元组,支持基于文本的目标指定。该模型在一个覆盖37个公共数据集和13个解剖类别的大规模超声分割语料库上进行训练,使其能够将超声视觉模式与不同器官和病变的临床相关概念对齐。为了进一步提高在现实临床交互中的可用性,我们提出了一种指导指令的代理,该代理将复杂的自然语言查询解析为UltraSAM3的简洁超声概念提示。大量实验表明,UltraSAM3在多器官超声基准、外部数据集和增强视觉提示的设置上,始终优于代表性的基于概念和文本的生物医学分割模型。此外,该代理提高了对复杂用户指令的分割鲁棒性。这些结果表明,超声特定的概念适配对于构建可泛化和互动的超声分割基础模型是有效的。
cs.CV / 41 / 2607.29202

Domain-Division based Progressive Learning for Source-Free Domain Adaptation

基于领域划分的渐进学习用于无源领域适应
Liu, Pan, Li, Jing, Zhao, Meng, Xue, Wanli, Hu, Qinghua, Chen, Shengyong
Abstract
With growing privacy and portability concerns, source-free domain adaptation requires only a source pre-trained model and an unlabeled target domain, allowing for effective adaptation to the target data. Most existing self-training methods focus on selecting and exploiting samples with reliable predictions, often neglecting others. Inspired by the finding that deep models learn clean samples faster than noisy ones, we propose a domain-division based progressive learning method named DPL. Specifically, our approach consists of two alternating stages, each beginning with the division of the target domain into easy-to-adapt and hard-to-adapt subdomains based on adaptation difficulty, followed by neighborhood-based pseudo label assignment. In stage one, we enhance classification accuracy through uncertainty-aware self-training and alignment of corresponding classes between subdomains. Stage two then applies tailored learning strategies to each subdomain, starting with consistency learning on the easy-to-adapt samples and progressing to utilizing local structural information for the more challenging ones, thereby mining the intrinsic properties of the target data. Extensive experiments on several widely used benchmarks validate the effectiveness of our approach, demonstrating superior performance compared to state-of-the-art methods. Our code is available at https://github.com/iamjingli/DPL.
Chinese Translation
随着隐私和可移植性问题的日益关注,无源领域适应只需一个源预训练模型和一个未标记的目标领域,从而有效地适应目标数据。现有的大多数自我训练方法集中于选择和利用具有可靠预测的样本,往往忽视了其他样本。受到深度模型学习干净样本的速度快于学习噪声样本的发现启发,我们提出了一种基于领域划分的渐进学习方法,命名为 DPL。具体而言,我们的方法由两个交替阶段组成,每个阶段开始时将目标领域根据适应难度划分为易适应子领域和难适应子领域,随后进行基于邻域的伪标签分配。在第一阶段,我们通过不确定性感知的自我训练和子领域之间对应类别的对齐来提高分类准确性。第二阶段则对每个子领域应用量身定制的学习策略,从对易适应样本进行一致性学习开始,逐步利用局部结构信息处理更具挑战性的样本,从而挖掘目标数据的内在特性。在多个广泛使用的基准测试上的大量实验验证了我们方法的有效性,显示出优于最先进方法的性能。我们的代码可在 https://github.com/iamjingli/DPL 获取。
cs.CV / 42 / 2607.29207

Multi-Modal Object Re-Identification with Dual Semantic Guidance and Global-Local Mutual Modulation

具有双重语义引导和全局-局部互调的多模态目标重识别
Zhou, Weixiang, Xu, Xingguo, Wang, Yuhao, Wang, Cong, Yang, Yang, Su, Zhixun, Pan, Jinshan
Abstract
Multi-modal object Re-Identification (ReID) aims to retrieve target instances by leveraging complementary information across modalities. However, existing methods suffer from two challenges. First, they often fail to exploit well-aligned and reliable semantic priors, making them vulnerable to background clutter and cross-modal misalignment. On the other hand, they typically rely on holistic feature modeling, overlooking the synergy between global and local representations. To overcome these limitations, we propose a robust multi-modal ReID framework with dual semantic guidance and global-local mutual modulation, which mainly consists of three key components, namely the Text-Semantic Injector (TSI), the Masked Global-Local Modulator (MGLM), and the Hierarchical MoE Fusion (HMF). The TSI enhances semantic awareness by integrating clean and coherent textual features into visual tokens. The MGLM enables part-aware cross-modal interaction through joint guidance from soft masks and global context, improving fine-grained feature alignment. Finally, the HMF adaptively aggregates multi-spectral features under local semantic supervision, yielding discriminative and robust representations. Extensive experiments on three multi-modal ReID benchmarks demonstrate the effectiveness of the proposed method. The code will be made publicly available at https://github.com/zw-absin/DSGM upon acceptance.
Chinese Translation
多模态目标重识别(ReID)旨在通过利用跨模态的互补信息来检索目标实例。然而,现有方法面临两个挑战。首先,它们往往未能充分利用对齐良好且可靠的语义先验,使其容易受到背景杂乱和跨模态错位的影响。另一方面,它们通常依赖整体特征建模,忽视了全局和局部表示之间的协同作用。为克服这些局限性,我们提出了一种具有双重语义引导和全局-局部互调的稳健多模态ReID框架,主要由三个关键组件组成,即文本语义注入器(Text-Semantic Injector, TSI)、掩膜全局-局部调制器(Masked Global-Local Modulator, MGLM)和分层MoE融合(Hierarchical MoE Fusion, HMF)。TSI通过将干净且连贯的文本特征整合到视觉标记中,增强了语义意识。MGLM通过软掩膜和全局上下文的联合引导,实现了对部分的跨模态交互,从而改善了细粒度特征对齐。最后,HMF在局部语义监督下自适应地聚合多光谱特征,产生具有区分性和稳健性的表示。在三个多模态ReID基准上的大量实验表明了所提方法的有效性。代码将在接受后公开发布于 https://github.com/zw-absin/DSGM。
cs.CV / 43 / 2607.29222

Is It Time for the Renaissance of Salient Object Detection in the Era of MLLMs?

在多模态大语言模型时代,显著物体检测的复兴时刻是否已到?
Zhao, Wenzhuo, Li, Xiuzhi, Mao, Zhongkuan, Xian, Ronghao, Jiang, Yao, Gao, Zhao, Fu, Keren, Zhao, Qijun, Cheng, Jian
Abstract
The zero-shot capabilities of multimodal large language models (MLLMs) are pushing salient object detection (SOD) beyond task-specific supervision. To disentangle MLLMs beyond conventional mask-based evaluation, we decompose SOD into localization and segmentation, and re-engineer datasets with phrases, boxes, and attributes, establishing a diagnostic benchmark for MLLM saliency perception (SaliLLM). SaliLLM uncovers a striking capability mismatch: MLLMs outperform state-of-the-art (SOTA) methods in localization, yet remain substantially weaker in segmentation. Further analyses attribute this gap primarily to mismatches between MLLMs and annotations over foreground cardinality, granularity, and extent. Motivated by this diagnosis, we recast zero-shot SOD as protocol-aligned Foreground Organization and introduce the first training-free framework that leverages Gestalt-inspired Collaborative attention for Unified SOD (FOCUS). FOCUS couples top-down Bayesian-surprise calibration of protocol-conditioned foreground granularity with bottom-up propagation of MLLMs evidence over entity-centric perceptual manifolds induced by self-supervised features, yielding coherent object extents as prompts for a general segmenter. Across 13 RGB, RGB-D, and RGB-T SOD benchmarks, FOCUS generally surpasses SOTA methods without training, reducing mean absolute error by 11\%, 34\%, and 48\% compared with fully, weakly, and self-supervised methods, respectively. Our findings signal the renaissance of SOD: from task-specific supervision to zero-shot foreground organization. Code is available in the supplementary material.
Chinese Translation
多模态大语言模型(MLLMs)的零样本能力正在推动显著物体检测(SOD)超越特定任务的监督。为了超越传统基于掩码的评估,我们将SOD分解为定位和分割,并重新构建包含短语、框和属性的数据集,建立了一个用于MLLM显著性感知的诊断基准(SaliLLM)。SaliLLM揭示了一个显著的能力不匹配:MLLMs在定位方面超越了最先进(SOTA)的方法,但在分割方面仍然显著较弱。进一步分析将这一差距主要归因于MLLMs与注释在前景基数、粒度和范围上的不匹配。基于这一诊断,我们将零样本SOD重新构建为协议对齐的前景组织,并引入了第一个无训练框架,利用受格式塔启发的协作注意力进行统一SOD(FOCUS)。FOCUS将协议条件下的前景粒度的自上而下贝叶斯惊讶校准与基于自监督特征诱导的以实体为中心的感知流形上的MLLM证据的自下而上传播相结合,生成连贯的物体范围作为通用分割器的提示。在13个RGB、RGB-D和RGB-T SOD基准上,FOCUS通常超越了SOTA方法,无需训练,分别减少了与完全、弱和自监督方法相比的平均绝对误差11%、34%和48%。我们的发现标志着SOD的复兴:从特定任务的监督到零样本前景组织。代码可在补充材料中获得。
cs.CV / 44 / 2607.29237

CorrelationFlow: A Training-Free Geometric Approach for LiDAR Scene Flow Estimation

CorrelationFlow:一种无训练的几何方法用于LiDAR场景流估计
Dao, Minh-Quan, Lin, Yancong, Perez, Julie Stephany Berrio, Caesar, Holger
Abstract
LiDAR scene flow estimation has settled into a monoculture: nearly all recent methods share the same feed-forward architecture and the same family of self-supervised losses, inheriting each other's assumptions, and each other's blind spots. When those assumptions fail, as they do for sparse, distant, or fast-moving objects, every method built on them fails together, and adding parameters or simulated training data does not fix what the formulation itself gets wrong. This paper takes the opposite path. We present CorrelationFlow, a training-free geometric framework that reduces scene flow to two textbook operations: connected-component labeling and correlation maximization on bird's-eye-view occupancy images. Objects are isolated as spatio-temporal connected components, their motions recovered as correlation peaks, and the resulting velocities propagated to all member points. However, this dense correlation evaluates every candidate displacement of every cluster and requires a window of past sweeps; therefore, we develop a sparse counterpart that operates on a single sweep pair by matching lightweight occupancy descriptors at boundary key points. Because nothing is trained, nothing is inherited: on the multi-domain test set of the Argoverse 2 2026 Scene Flow Challenge, spanning five datasets with heterogeneous sensors and platforms, CorrelationFlow ranked second among unsupervised methods and degrades most gracefully at long range, where the shared assumptions of learned methods break down. Our results suggest that a substantial share of the scene flow problem is solvable by classical computer vision, and that progress may require questioning the formulation, not scaling it.
Chinese Translation
LiDAR场景流估计已经进入了一种单一化的状态:几乎所有近期的方法都共享相同的前馈架构和相同类型的自监督损失,继承了彼此的假设和盲点。当这些假设失效时,例如在稀疏、远距离或快速移动的物体上,基于这些假设构建的每种方法都会一起失效,而增加参数或模拟训练数据并不能修正公式本身的错误。本文采取了相反的路径。我们提出了CorrelationFlow,这是一种无训练的几何框架,将场景流简化为两个教科书操作:连通组件标记和在鸟瞰图占用图像上的相关性最大化。物体被孤立为时空连通组件,其运动被恢复为相关性峰值,结果速度传播到所有成员点。然而,这种密集的相关性评估每个聚类的每个候选位移,并需要过去扫面的窗口;因此,我们开发了一种稀疏对应方法,通过在边界关键点匹配轻量级占用描述符,在单一扫面对上进行操作。由于没有经过训练,因此没有任何东西被继承:在Argoverse 2 2026场景流挑战的多领域测试集中,涵盖了五个具有异构传感器和平台的数据集,CorrelationFlow在无监督方法中排名第二,并且在长距离时表现最为优雅,此时学习方法的共享假设会崩溃。我们的结果表明,场景流问题的很大一部分可以通过经典计算机视觉解决,而进展可能需要质疑公式,而不是扩展它。
cs.CV / 45 / 2607.29240

When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration

当模型先验与视觉证据冲突时:通过选择性先验校准减轻常识驱动的幻觉
Chen, Kesheng, Hu, Yamin, Luo, Wenjian
Abstract
In vision--language models, commonsense-driven hallucination (CDH) occurs when a model's commonsense prior overrides clear visual evidence of an atypical state. For example, a model may report that a visibly six-fingered hand has five fingers. We show that these errors are systematically directed: when a model answers a question about a counterfactual (CF) image incorrectly, its answer often coincides with the candidate it prefers without access to the image. Suppressing this prior indiscriminately can repair CF errors, but may also disrupt correct answers on matched commonsense (CS) images, where the same prior is helpful. We therefore propose Selective Prior Calibration (SPC), which subtracts candidate-level prior-preference estimates from image-conditioned scores with an instance-dependent strength and revises the original prediction only when the resulting score pattern strongly supports an alternative. Extensive experiments demonstrate that SPC substantially improves accuracy on CF images while largely preserving accuracy on matched CS images. Furthermore, these gains generalize across CDH categories, candidate-answer permutations, and other conflict benchmarks, while SPC rarely alters predictions on benchmarks without such conflicts.
Chinese Translation
在视觉-语言模型中,当模型的常识先验覆盖了对非典型状态的明确视觉证据时,就会发生常识驱动的幻觉(CDH)。例如,模型可能报告一个明显有六根手指的手只有五根手指。我们展示了这些错误是系统性导向的:当模型错误回答关于反事实(CF)图像的问题时,其答案往往与在没有访问图像的情况下更偏好的候选答案一致。无差别地抑制这种先验可以修复CF错误,但也可能干扰在匹配的常识(CS)图像上的正确答案,在这些情况下,相同的先验是有帮助的。因此,我们提出了选择性先验校准(SPC),该方法根据实例依赖的强度从图像条件分数中减去候选级别的先验偏好估计,并仅在结果分数模式强烈支持替代方案时修正原始预测。大量实验表明,SPC显著提高了CF图像的准确性,同时在匹配的CS图像上基本保持准确性。此外,这些提升在CDH类别、候选答案排列和其他冲突基准中具有广泛的普适性,而SPC在没有此类冲突的基准上很少改变预测。
cs.CV / 46 / 2607.29243

TAVI-TEC: An AI-Based Tool for Procedural Planning of Transcatheter Aortic Valve Implantation

TAVI-TEC:一种基于人工智能的经导管主动脉瓣植入术程序规划工具
Zerillo, Alessandra, Cannata, Stefano, Bellavia, Diego, Ciriello, Daniele, Manini, Simone, Pasta, Salvatore, Gandolfo, Caterina
Abstract
Computed tomography angiography (CTA) is crucial for preprocedural TAVI planning, providing the anatomical information required for prosthesis sizing and vascular access assessment. As the volume of TAVI procedure increases, improving efficiency and standardizing annotations is becoming essential in clinical practice. This study presents TAVI-TEC, a fully automated artificial intelligence-based framework integrated into a web based DICOM viewer for routine preoperative TAVI planning. Pre-procedural CTA scans from patients undergoing TAVI with SAPIEN 3 Ultra (S3U) prostheses were processed using a fully automated pipeline. Deep learning-based segmentation of cardiovascular structures, calcification detection, centerline extraction, landmark identification, and annular plane definition was implemented to quantify key annular and aortic root measurements and color-coded maps of lumen reduction and vessel diameter for vascular access. A multilayer perceptron classifier was trained to predict prosthesis size prior to the TAVI procedure. Results revealed that TAVI-TEC enabled pre-procedural measurements in approximately 2-6 min. Strong agreement with clinician-derived measurements was observed for annular area (coefficient of concordance, CCC = 0.934; interclass correlation coefficient, ICC = 0.935; R^2 = 0.881) and perimeter (CCC = 0.909; ICC = 0.909; R^2 = 0.854). The valve-size prediction model achieved 82% overall accuracy, with most misclassifications occurring between adjacent prosthesis sizes. Though further multicenter validation and extension to additional measurements and valve platforms are required, the TAVI-TEC methodology may reduce operator variability in pre-TAVI measurements and streamline the preoperative workflows of the Heart Team for decision-making.
Chinese Translation
计算机断层扫描血管造影(CTA)对于经导管主动脉瓣植入术(TAVI)的术前规划至关重要,提供了所需的解剖信息以进行假体尺寸测量和血管通路评估。随着TAVI手术数量的增加,提高效率和标准化标注在临床实践中变得尤为重要。本研究介绍了TAVI-TEC,这是一种完全自动化的基于人工智能的框架,集成于基于网络的DICOM查看器中,用于常规的术前TAVI规划。对接受SAPIEN 3 Ultra(S3U)假体的患者进行的术前CTA扫描采用完全自动化的流程进行处理。实施了基于深度学习的心血管结构分割、钙化检测、中心线提取、标志点识别和环平面定义,以量化关键的环和主动脉根部测量,并生成血管通路的腔体缩小和血管直径的彩色编码图。训练了一个多层感知器分类器,以预测TAVI手术前的假体尺寸。结果显示,TAVI-TEC能够在大约2-6分钟内完成术前测量。环面积(协调系数,CCC = 0.934;组内相关系数,ICC = 0.935;R^2 = 0.881)和周长(CCC = 0.909;ICC = 0.909;R^2 = 0.854)与临床医生测得的结果高度一致。阀门尺寸预测模型的整体准确率达到82%,大多数错误分类发生在相邻假体尺寸之间。尽管需要进一步的多中心验证以及扩展到额外的测量和阀门平台,但TAVI-TEC方法可能减少术前TAVI测量中的操作员变异性,并简化心脏团队的术前决策工作流程。
cs.CV / 47 / 2607.29266

OsteoCAD: A Human-in-the-Loop Cloud-Edge Framework for Bone Tumor Segmentation

OsteoCAD:一种人机协同的云边框架用于骨肿瘤分割
Rodriguez-Herrero, Maximo, Sanchez-Gallegos, Dante D., Aguirre-Meneses, Heriberto, Núñez-Gaona, Marco Antonio, Gonzalez-Compean, J. L., Carretero, Jesus
Abstract
Artificial Intelligence (AI) and Deep Learning (DL) have notably advanced medical image analysis, yet many health- care organizations struggle to adopt them due to limited com- putational resources and specialized expertise. To address these barriers, we introduce OsteoCAD, a modular eHealth framework that democratizes access to DL tools in clinical practice. Osteo- CAD delivers end-to-end DL capabilities-from dataset creation and preprocessing to model training and inference-through an integrated and user-friendly interface. To mitigate local hardware constraints, the framework securely connects to remote GPU infrastructures. We validate OsteoCAD's feasibility through a real-world case study in Mexico focused on large bone tumor segmentation. The results demonstrate the framework's ability to enable DL-powered eHealth solutions without demanding ad- vanced technical expertise or complex local configurations.
Chinese Translation
人工智能(AI)和深度学习(DL)在医学图像分析方面取得了显著进展,但许多医疗机构由于计算资源和专业知识的限制而难以采用这些技术。为了解决这些障碍,我们提出了OsteoCAD,一个模块化的电子健康框架,旨在使临床实践中的深度学习工具更加普及。OsteoCAD提供了端到端的深度学习能力——从数据集创建和预处理到模型训练和推理——通过一个集成且用户友好的界面实现。为了缓解本地硬件的限制,该框架安全地连接到远程GPU基础设施。我们通过在墨西哥进行的一个真实案例研究验证了OsteoCAD的可行性,重点关注大骨肿瘤的分割。结果表明,该框架能够实现基于深度学习的电子健康解决方案,而无需复杂的本地配置或高级技术专长。
cs.CV / 48 / 2607.29278

Training-Free Entity-Level Few-Shot Segmentation of Remote Sensing Images with Advection Refinement

无训练的遥感图像实体级少样本分割框架与对流细化
Bai, Xueting, Ni, Huan
Abstract
Existing cross-domain few-shot segmentation approaches suffer from high training costs due to source-domain episodic training and pixel-wise dense prediction, while often producing fragmented and noisy predictions. To overcome these issues, we propose a training-free entity-level few-shot segmentation framework for remote sensing images with advection refinement. Specifically, we first leverage SAM3's generic geometric priors to generate category-agnostic entity primitives. By reformulating few-shot inference from pixel-level prediction to entity-level reasoning, foreground and background prototypes are constructed and combined with dense textual semantic responses from SAM3 to build a multi-modal semantic potential field. Furthermore, an advection equation-based semantic refinement mechanism is introduced to propagate category-aware information across both feature and similarity spaces, enhancing semantic continuity and suppressing local texture noise. Extensive experiments on multiple remote sensing datasets demonstrate that the proposed framework effectively mitigates domain shift and local noise, substantially improving SAM3's adaptation capability for remote sensing few-shot segmentation without additional training. Our code will be publicly available at https://github.com/yu-ni1989/ELFSS-AR.
Chinese Translation
现有的跨域少样本分割方法由于源域的情景训练和像素级的密集预测,面临着高昂的训练成本,同时常常产生碎片化和噪声较大的预测结果。为了解决这些问题,我们提出了一种无训练的遥感图像实体级少样本分割框架,结合对流细化机制。具体而言,我们首先利用SAM3的通用几何先验生成类别无关的实体原型。通过将少样本推理从像素级预测重新构建为实体级推理,构建前景和背景原型,并将其与来自SAM3的密集文本语义响应结合,建立一个多模态语义潜力场。此外,引入了一种基于对流方程的语义细化机制,以在特征空间和相似性空间中传播类别感知信息,从而增强语义连续性并抑制局部纹理噪声。在多个遥感数据集上的广泛实验表明,所提出的框架有效缓解了领域转移和局部噪声,显著提高了SAM3在遥感少样本分割中的适应能力,而无需额外的训练。我们的代码将公开发布于 https://github.com/yu-ni1989/ELFSS-AR。
cs.CV / 49 / 2607.29284

FillGS: Filling Observation Gaps in 4D Gaussian Splatting via Viewpoint-Time Selection and Generative Refinement

FillGS:通过视点-时间选择和生成细化填补4D高斯喷溅中的观察空白
Otonari, Takashi, Yamasaki, Toshihiko
Abstract
4D Gaussian Splatting (4DGS) can render dynamic scenes photorealistically. However, with limited viewpoint coverage, some spatiotemporal regions remain sparsely observed, leading to artifacts, particularly in scenes with large motion. Existing approaches leveraging generative models rely on heuristic virtual-viewpoint selection before refining rendered views. As a result, they cannot actively explore such sparsely observed regions. To address this issue, we propose a pipeline that actively selects spatiotemporal virtual viewpoints to improve 4DGS reconstruction. Our method selects virtual viewpoints for generative enhancement based on the rendering sensitivity and motion-aware observation density of 4D Gaussians, prioritizing views that alleviate observation sparsity. In the refined images, we filter out regions that conflict with captured observations or are likely to contain generative artifacts and then fine-tune 4DGS using only the reliable regions. We evaluate our method on multi-view video benchmarks using new train/test splits designed to induce observation gaps. Results show consistent improvements over prior viewpoint selection strategies and fine-tuning methods in both qualitative and quantitative evaluations, while reducing artifacts.
Chinese Translation
4D高斯喷溅(4DGS)能够以照片级真实感渲染动态场景。然而,由于视点覆盖有限,一些时空区域的观察仍然稀疏,导致伪影,尤其是在大运动的场景中。现有利用生成模型的方法依赖于启发式虚拟视点选择,然后再对渲染视图进行细化。因此,它们无法主动探索这些稀疏观察区域。为了解决这个问题,我们提出了一种管道,主动选择时空虚拟视点以改善4DGS重建。我们的方法基于4D高斯的渲染敏感性和运动感知观察密度选择虚拟视点进行生成增强,优先考虑能够缓解观察稀疏性的视图。在细化后的图像中,我们过滤掉与捕获观察相冲突或可能包含生成伪影的区域,然后仅使用可靠区域对4DGS进行微调。我们在多视角视频基准上评估了我们的方法,使用新设计的训练/测试划分以诱导观察空白。结果显示,在定性和定量评估中,相较于先前的视点选择策略和微调方法,我们的方法在减少伪影的同时,持续改善了性能。
cs.CV / 50 / 2607.29310

CALM-AH: An ABAW11-Calibrated Multimodal Ensemble with Reliability-Gated Multi-Expert Consensus for Video-Level Ambivalence and Hesitancy Recognition

CALM-AH:一种基于ABAW11校准的多模态集成方法,结合可靠性门控的多专家共识用于视频级别的矛盾和犹豫识别
Sun, Wenzhuo, Liang, Mingjian, Attfield, Richard, Ge, Zongyuan, Cheng, Xuelian, Carreno-Medrano, Pamela
Abstract
Ambivalence and hesitancy (A/H) are subtle behavioural states that may be expressed through language, voice, facial activity, and other non-verbal cues. The ABAW11 A/H Video Recognition Challenge asks systems to assign a binary A/H label to each naturalistic interview video. Performance is measured using Macro-F1 so that recognition of both A/H and No-A/H samples receives equal importance. We present CALM-AH, a multimodal ensemble that combines textual, acoustic, visual, and derived behavioural-statistical features. We construct 15 non-empty combinations of these feature branches. For each combination, we select the best of three classifier families using validation binary cross-entropy and optimise its decision threshold for validation Macro-F1. The resulting binary decisions are combined using fixed hard-voting weights transferred from BROTHER. We further introduce Reliability-Gated Multi-Expert Consensus(RG-MEC), an anchor-preserving decision-level ensemble that combines an initial prediction with three complementary correction experts: CALM-AH, AffectGPT, and a GPT-based semantic verifier. The initial system provides the default prediction. Its label is overridden only when all three correction experts unanimously support the same alternative class; otherwise, the anchor prediction is retained. This unanimity-gated design limits the influence of isolated expert errors while permitting bidirectional correction when task-specific, multimodal-affective, and semantic-pragmatic evidence are fully consistent. On the participant-disjoint ABAW11 dataset, CALM-AH achieves a Macro-F1 of 0.7525, and the complete RG-MEC system achieves 0.7771.
Chinese Translation
矛盾和犹豫(A/H)是通过语言、声音、面部活动及其他非语言线索表达的微妙行为状态。ABAW11 A/H视频识别挑战要求系统为每个自然访谈视频分配一个二元的A/H标签。性能通过宏F1(Macro-F1)进行评估,以确保对A/H和非A/H样本的识别同等重要。我们提出了CALM-AH,这是一种结合文本、声学、视觉和衍生行为统计特征的多模态集成方法。我们构建了15种非空特征分支组合。对于每种组合,我们使用验证二元交叉熵选择三种分类器家族中的最佳者,并优化其决策阈值以提高验证宏F1。最终的二元决策通过从BROTHER转移的固定硬投票权重进行组合。我们进一步引入了可靠性门控多专家共识(Reliability-Gated Multi-Expert Consensus,RG-MEC),这是一种保持锚定的决策级别集成方法,将初始预测与三个互补的修正专家结合起来:CALM-AH、AffectGPT和基于GPT的语义验证器。初始系统提供默认预测。只有当所有三个修正专家一致支持同一替代类别时,其标签才会被覆盖;否则,保留锚定预测。这种一致性门控设计限制了孤立专家错误的影响,同时在任务特定的多模态情感和语义-语用证据完全一致时允许双向修正。在参与者不重叠的ABAW11数据集上,CALM-AH达到了0.7525的宏F1,而完整的RG-MEC系统达到了0.7771。
cs.CV / 51 / 2607.29337

DualDiT: A Conditional Dual-Output Diffusion Transformer for Joint OCT Image and Segmentation Mask Generation

DualDiT:一种用于联合OCT图像和分割掩膜生成的条件双输出扩散变换器
García-Torres, Fernando, del Amor, Rocío, Morales, Sandra, Barroso, Álvaro, Heiduschka, Peter, Kemper, Björn, Naranjo, Valery
Abstract
Background and Objective: Generating realistic medical images with anatomically accurate segmentation masks helps address the shortage of annotated data in medical imaging, particularly in optical coherence tomography (OCT) of mouse eyes, where manual retinal layer delineation is labour-intensive due to tiny structures and required expertise, resulting in scarce datasets. While diffusion models perform well in medical image synthesis, joint image-mask generation has relied mainly on U-Net-based denoisers, leaving diffusion transformers largely unexplored. Methods: We propose a conditional dual-output Diffusion Transformer (DualDiT) for joint synthesis of OCT B-scans and segmentation masks of the upper retinal cell layers in ex vivo mouse retina. DualDiT encodes both modalities into a shared latent space via a pretrained VAE, concatenates their latent representations, and performs conditional diffusion over the joint tensor. We compared DualDiT against two adapted diffusion baselines: DDPM and LDM. Generative quality was assessed via Fr\'echet Inception Distance (FID) and spatial FID (sFID); practical utility via synthetic data augmentation for downstream U-Net segmentation; and perceptual realism via evaluation by three domain experts. Results: DualDiT achieved the best generative quality (FID 56.14, sFID 114.35), outperforming DDPM and LDM. Expert panels misclassified 46% of synthetic samples as real and 42% of real samples as synthetic. Adding DualDiT-generated images and masks improved Dice and IoU scores on a held-out segmentation test set. Conclusions: DualDiT shows that transformer-based diffusion models can effectively learn the joint distribution of OCT images and segmentation masks, surpassing DDPM- and LDM-based baselines in generative fidelity, downstream utility, and perceptual realism, highlighting its potential for data augmentation in annotation-scarce medical imaging.
Chinese Translation
背景与目的:生成具有解剖学准确分割掩膜的逼真医学图像有助于解决医学成像中标注数据的短缺,特别是在小鼠眼睛的光学相干断层扫描(OCT)中,由于微小结构和所需的专业知识,手动视网膜层的划分劳动密集,导致数据集稀缺。尽管扩散模型在医学图像合成中表现良好,但联合图像-掩膜生成主要依赖于基于U-Net的去噪器,扩散变换器尚未得到充分探索。方法:我们提出了一种条件双输出扩散变换器(DualDiT),用于联合合成小鼠视网膜外部的OCT B扫描和上层视网膜细胞层的分割掩膜。DualDiT通过预训练的变分自编码器(VAE)将这两种模态编码到共享的潜在空间中,连接它们的潜在表示,并在联合张量上执行条件扩散。我们将DualDiT与两个改编的扩散基线进行比较:DDPM和LDM。通过Fréchet Inception Distance(FID)和空间FID(sFID)评估生成质量;通过合成数据增强下游U-Net分割评估实用性;通过三位领域专家的评估评估感知真实感。结果:DualDiT在生成质量上表现最佳(FID 56.14,sFID 114.35),优于DDPM和LDM。专家小组将46%的合成样本误分类为真实样本,将42%的真实样本误分类为合成样本。添加DualDiT生成的图像和掩膜提高了保留分割测试集上的Dice和IoU分数。结论:DualDiT表明基于变换器的扩散模型能够有效学习OCT图像和分割掩膜的联合分布,在生成保真度、下游实用性和感知真实感方面超越了基于DDPM和LDM的基线,突显其在标注稀缺医学成像中进行数据增强的潜力。
cs.CV / 52 / 2607.29367

SatEdit: Mask-Conditioned Image Editing via VLM-Guided Segment Annotation

SatEdit:基于VLM引导的分段注释的掩膜条件卫星图像编辑
Talha, Muhammad, Amer, Muhammad Ahmed
Abstract
Satellite image editing requires spatially precise object-level control, but supervised editing datasets for overhead imagery are costly to build because object masks, semantic labels, and paired edits are rarely available at scale. We introduce SatEdit, a mask-conditioned satellite image editing framework that constructs training supervision from unlabeled imagery. SatEdit proposes object masks with a seg- mentation foundation model, assigns semantic la- bels to sampled segments with a Vision-Language Model, and applies lightweight human verification before generating paired addition and removal exam- ples through mask-guided inpainting. We fine-tune a high-resolution image editing backbone with LoRA on a SODA-A-derived dataset containing 1,014 im- ages and 852 verified object annotations across 91 classes. In controlled comparisons with open- source and proprietary image editing models, SatE- dit achieves the highest aggregate masked-region se- mantic alignment, with a CLIP score of 0.6322 and CLIP delta of 0.0726, while preserving the surround- ing scene qualitatively. These results suggest that VLM-assisted segment annotation is a practical route to data-efficient, spatially controllable satellite image editing.
Chinese Translation
卫星图像编辑需要空间上精确的对象级控制,但由于对象掩膜、语义标签和配对编辑在规模上很少可用,因此构建用于高空影像的监督编辑数据集成本高昂。我们提出了SatEdit,一个基于掩膜条件的卫星图像编辑框架,该框架从未标记的图像中构建训练监督。SatEdit通过分割基础模型提出对象掩膜,利用视觉-语言模型(Vision-Language Model)为采样的分段分配语义标签,并在通过掩膜引导的修复生成配对的添加和删除示例之前,进行轻量级的人类验证。我们在一个包含1,014幅图像和91个类别中852个经过验证的对象注释的SODA-A衍生数据集上,使用LoRA对高分辨率图像编辑骨干网络进行了微调。在与开源和专有图像编辑模型的受控比较中,SatEdit在掩膜区域语义对齐的总分上达到了最高,CLIP得分为0.6322,CLIP增量为0.0726,同时在定性上保持了周围场景的完整性。这些结果表明,VLM辅助的分段注释是实现数据高效、空间可控的卫星图像编辑的可行途径。
cs.CV / 53 / 2607.29370

VFAD: Variational Semantic Prompting Meets Frequency-Adaptive Representation Learning for Zero-Shot Anomaly Detection

VFAD:变分语义提示与频率自适应表示学习相结合的零样本异常检测
Chen, Peng, Li, Kaige, Wang, Wei, Yang, Mingbo, Wang, Wenqiang, Shen, Li, Huang, Fangjun, Huang, Chao
Abstract
Zero-shot anomaly detection (ZSAD) aims to detect and localize anomalies in unseen categories without access to target-specific training data. Although recent CLIP-based methods have demonstrated promising generalization through vision-language alignment, they remain limited in capturing diverse anomaly semantics and subtle local variations. To address these limitations, we propose VFAD, a unified framework that combines variational semantic prompting with frequency-adaptive representation learning. Specifically, we introduce a Variational Semantic Prompt Extractor (VSPE), which adaptively aggregates anomaly-relevant local semantics from dense patch tokens and regularizes them through a variational information bottleneck, thereby incorporating fine-grained visual cues and enabling more precise cross-modal alignment. Furthermore, we develop a Frequency-Adaptive Representation Aggregation (FARA) module that leverages wavelet-based frequency decomposition and frequency-specific expert aggregation to enhance anomaly-discriminative visual representations. By jointly strengthening semantic guidance and visual representation learning, VFAD improves both anomaly discrimination and fine-grained localization. Extensive experiments on 13 industrial and medical benchmarks demonstrate that VFAD consistently outperforms existing state-of-the-art ZSAD methods across diverse anomaly scenarios. The code will be publicly available upon publication.
Chinese Translation
零样本异常检测(ZSAD)旨在在未见类别中检测和定位异常,而无需访问特定目标的训练数据。尽管最近基于 CLIP 的方法通过视觉-语言对齐展示了良好的泛化能力,但在捕捉多样化的异常语义和细微的局部变化方面仍然存在局限性。为了解决这些问题,我们提出了 VFAD,一个将变分语义提示与频率自适应表示学习相结合的统一框架。具体而言,我们引入了变分语义提示提取器(VSPE),该提取器自适应地从密集的补丁标记中聚合与异常相关的局部语义,并通过变分信息瓶颈对其进行正则化,从而结合细粒度的视觉线索,实现更精确的跨模态对齐。此外,我们开发了频率自适应表示聚合(FARA)模块,该模块利用基于小波的频率分解和频率特定的专家聚合来增强异常区分的视觉表示。通过共同加强语义引导和视觉表示学习,VFAD 改善了异常区分和细粒度定位。在 13 个工业和医疗基准上的大量实验表明,VFAD 在各种异常场景中始终优于现有的最先进的 ZSAD 方法。代码将在发表后公开。
cs.CV / 54 / 2607.29394

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

通过条件潜在传输实现密集时间对比合成
Joshi, Smriti, Tsirikoglou, Apostolia, Lang, Daniel M., Osuala, Richard, Varaa, Noah Márquez, Guzman, Alejandro, Skorupko, Grzegorz, Arregui, Sebastian Ibarra, Garrucho, Lidia, Ohashi, Akane, Ntoula, Dimitra, Divjak, Eugen, Lafcı, Oğuz, Peeken, Jan C., Schnabel, Julia A., Strand, Fredrik, Diaz, Oliver, Lekadir, Karim
Abstract
Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is essential for breast cancer management, but reliance on gadolinium-based contrast agents (GBCAs) restricts use in contraindicated populations, prolongs scan protocols, and presents environmental toxicity concerns. Contrast synthesis offers a non-invasive alternative; however, existing approaches struggle to balance spatial realism with temporal continuity, suffer from slow iterative sampling, underutilize structural priors, and lack clinical validation. We propose a novel conditioned latent transport framework that predicts contrast enhancement in a single forward pass. By anchoring the latent trajectory to the pre-contrast anatomy and applying continuous time conditioning, the model synthesizes patient-specific contrast evolution at any acquisition time. The proposed approach outperforms baseline and the state-of-the-art models across spatial, perceptual, temporal, and distributional metrics. Evaluated on an independent external cohort, the method demonstrates robustness to domain shifts induced by scanner noise as well as differing acquisition protocol. Furthermore, our synthetic contrast enhancement significantly improved downstream tumor segmentation performance, yielding a 22.4% relative increase in Dice coefficient (0.60 vs. 0.49 baseline pre-contrast, p < 0.01), reducing boundary segmentation error by over 39%, while outperforming all other generative model baselines. Finally, a reader study involving four breast radiologists evaluated the image quality, kinetic fidelity, and diagnostic viability of our synthesized sequences across 40 randomly selected cases. The results demonstrated that in 70% of cases, synthesized images provided sufficient clinical information to support the same management decisions as real DCE-MRI, suggesting a path toward safer and faster contrast-free or contrast-reduced imaging workflows.
Chinese Translation
动态对比增强磁共振成像(DCE-MRI)对于乳腺癌管理至关重要,但对基于钆的对比剂(GBCA)的依赖限制了在禁忌人群中的使用,延长了扫描协议,并带来了环境毒性问题。对比合成提供了一种非侵入性的替代方案;然而,现有方法在空间真实感与时间连续性之间难以取得平衡,迭代采样速度缓慢,结构先验利用不足,并且缺乏临床验证。我们提出了一种新颖的条件潜在传输框架,该框架通过单次前向传递预测对比增强。通过将潜在轨迹锚定到对比前解剖结构,并应用连续时间条件,模型能够在任何采集时间合成特定患者的对比演变。所提出的方法在空间、感知、时间和分布度量方面均优于基线和最先进的模型。在独立的外部队列评估中,该方法表现出对扫描仪噪声引起的领域转变以及不同采集协议的鲁棒性。此外,我们的合成对比增强显著提高了后续肿瘤分割性能,Dice系数相对提高22.4%(0.60对比0.49基线对比前,p < 0.01),边界分割误差减少超过39%,同时超越了所有其他生成模型基线。最后,一项涉及四位乳腺放射科医师的读者研究评估了我们合成序列在40个随机选择病例中的图像质量、动力学保真度和诊断可行性。结果表明,在70%的病例中,合成图像提供了足够的临床信息,以支持与真实DCE-MRI相同的管理决策,暗示了一条通向更安全、更快速的无对比或减少对比成像工作流程的路径。
cs.CV / 55 / 2607.29401

OSEF: One-Step Evidence Fusion for Cross-Video Scene Procedure Planning

OSEF:用于跨视频场景程序规划的一步证据融合
Ye, Zhentong, Zhang, Lei, Zhou, Sijia, Yu, Yingda, Shi, Yuehan, Xuan, Jiaqi, Dong, Shuaiwu, Tong, Guanchao, Zhang, Meimei, Li, Bin
Abstract
Video Scene Procedure Planning (VSPP) supplies the target start-goal observations in advance, leaving open how a planner should act when the evidence must itself be retrieved. We introduce Cross-Video Scene Procedure Planning (CVSPP): given an answer-redacted start-goal query and K candidate videos, a model must retrieve the supporting video, localize the relevant window, and predict the action sequence. Two obstacles couple here. Same-task demonstrations share stages and windows, and an early hard selection passes the wrong scene chain to the planner. We build an eleven-source benchmark with typed negative roles, a fail-closed answer-leakage gate, and separate Evidence- and Plan-axis metrics. On its 14 source-horizon cells we adapt nine planner families against a majority-sequence floor. We then present One-Step Evidence Fusion (OSEF), which scores a query-conditioned cell-and-span lattice over all candidates and feeds the full soft lattice to the planner through a token-global adapter, cropping no window beforehand. OSEF ranks first on all six cells the benchmark certifies as method-rankable. On four matched same-task COIN and CrossTask cells it improves exact-video-and-plan success by 2.9-10.7 points over an enhanced hard-selection SOTA, and a component study assigns the largest single increment to the token-global interface. Five converted-source cells sit at or near the majority-sequence floor, the benchmark's remaining headroom. The supplementary package includes model constructors and evaluation code.
Chinese Translation
视频场景程序规划(VSPP)提前提供目标起始-目标观察,但未明确规划者在必须检索证据时应如何行动。我们提出了跨视频场景程序规划(CVSPP):给定一个去除答案的起始-目标查询和 K 个候选视频,模型必须检索支持视频、定位相关窗口并预测动作序列。这里存在两个障碍。同任务演示共享阶段和窗口,而早期的硬选择将错误的场景链传递给规划者。我们构建了一个包含类型化负角色的十一源基准,设有一个失败关闭的答案泄漏门和独立的证据轴和计划轴指标。在其 14 个源视野单元上,我们针对大多数序列底线适配了九个规划者家族。然后,我们提出了一步证据融合(OSEF),它在所有候选者上对查询条件的单元和跨度格进行评分,并通过一个令牌全局适配器将完整的软格馈送给规划者,而不提前裁剪任何窗口。OSEF 在基准认证为可排名的方法的六个单元中排名第一。在四个匹配的同任务 COIN 和 CrossTask 单元中,它在精确视频和计划成功率上比增强的硬选择最先进技术提高了 2.9-10.7 个百分点,而组件研究将最大的单一增量分配给令牌全局接口。五个转换源单元位于或接近大多数序列底线,这是基准的剩余提升空间。补充包包括模型构造器和评估代码。
cs.CV / 56 / 2607.29412

Role-Break in Attention Heads: Understanding and Detecting Hallucinations in VLMs

注意力头中的角色断裂:理解和检测视觉语言模型中的幻觉
Wang, Mingyu, Jin, Weilin, Li, Wenbo, Huang, Haoyang, Duan, Nan, Jia, Tong, Luo, Chaoran, Li, Ying
Abstract
Despite remarkable progress in vision-language generation, Vision-Language Models (VLMs) remain prone to hallucinations, producing content that is inconsistent with or unsupported by the input image. Existing works largely design detection or mitigation methods around one specific hallucination pattern, such as visual-textual imbalance, but real VLM hallucinations arise from a mixture of multiple patterns, so signals bound to a single pattern struggle to remain stable across models and tasks. Under a unified head-level view, we find that hallucination-induced changes manifest as localized deviations from each head's faithful contextual behavior, a phenomenon we term Role-Break. Detailed analysis reveals that these deviations are systematically organized across attention heads, contextual sources, and deviation directions, and that the resulting signal is linearly readable once head identity is preserved. Based on these findings, we build a lightweight linear detector on top of Role-Break that requires no fine-tuning of the VLM, whose feature dimension stays below 5,000 and reaches an average AUROC of 93.23 across six VLMs and four benchmarks. A small-scale intervention experiment further shows that the detected tokens can be directly acted upon in the discriminative setting.
Chinese Translation
尽管视觉语言生成取得了显著进展,视觉语言模型(VLMs)仍然容易出现幻觉,生成与输入图像不一致或缺乏支持的内容。现有研究主要围绕特定的幻觉模式设计检测或缓解方法,例如视觉-文本不平衡,但真实的VLM幻觉是多种模式的混合,因此仅依赖单一模式的信号在不同模型和任务中难以保持稳定。在统一的头级视角下,我们发现幻觉引发的变化表现为每个头的忠实上下文行为的局部偏差,这一现象我们称之为角色断裂。详细分析表明,这些偏差在注意力头、上下文来源和偏差方向上是系统性组织的,并且一旦保持头的身份,所产生的信号是线性可读的。基于这些发现,我们在角色断裂的基础上构建了一个轻量级线性检测器,该检测器无需对VLM进行微调,其特征维度保持在5000以下,并在六个VLM和四个基准测试中达到了93.23的平均AUROC。小规模干预实验进一步表明,检测到的标记可以在判别设置中直接进行处理。
cs.CV / 57 / 2607.29445

QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models

针对红外视觉-语言模型的QR结构热触发器的目标语义攻击
Chen, Xiang, Zhao, Yingying, Li, Chao, Han, Jiaju, Zhang, Ben, Li, Ang, Long, Jiahuan, Wei, Yiwei, Guo, Jiujiang, Hu, Chengyin
Abstract
Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual question answering. However, their robustness to structured thermal perturbations and the stability of cross-modal semantic alignment remain insufficiently studied. We propose QR-Structured Thermal Triggers (QR-STT), a stealthy, training-free, black-box framework for targeted semantic steering of IR-VLMs. QR-STT preserves the functional regions of a QR pattern while optimizing its internal modules, each of which is assigned a cold, neutral, or hot thermal state. The framework jointly searches module topology and rendering parameters, including position, scale, rotation, intensity, blur, and roundness. A three-stage gradient-free procedure with greedy module-flip refinement efficiently handles the mixed discrete and continuous search space. The objective promotes alignment with an attacker-selected target, suppresses source-class evidence, and regularizes QR structure and visual similarity. Experiments on multiple CLIP-style encoders show that QR-STT consistently redirects image-text alignment toward chosen concepts while maintaining visual stealth. Perturbations optimized for classification also transfer to image captioning and VQA, causing target-consistent semantic drift in generated outputs. These results identify QR-structured thermal patterns as an interpretable attack surface for language-driven infrared perception and highlight the need for robustness evaluation against structured cross-task semantic attacks.
Chinese Translation
红外视觉-语言模型(IR-VLMs)将热感知扩展到开放词汇分类、图像描述和视觉问答。然而,它们对结构化热扰动的鲁棒性以及跨模态语义对齐的稳定性仍然研究不足。我们提出了QR结构热触发器(QR-STT),这是一种隐蔽的、无训练的黑箱框架,用于对IR-VLMs进行目标语义引导。QR-STT保留了QR模式的功能区域,同时优化其内部模块,每个模块被分配为冷、中性或热的热状态。该框架共同搜索模块拓扑和渲染参数,包括位置、比例、旋转、强度、模糊和圆度。一个三阶段的无梯度程序与贪婪模块翻转优化有效地处理混合离散和连续搜索空间。目标促进与攻击者选择的目标的对齐,抑制源类别证据,并规范QR结构和视觉相似性。在多个CLIP风格编码器上的实验表明,QR-STT始终将图像-文本对齐重定向到所选概念,同时保持视觉隐蔽性。针对分类优化的扰动也转移到图像描述和视觉问答中,导致生成输出中的目标一致语义漂移。这些结果将QR结构热模式识别为语言驱动的红外感知的可解释攻击面,并强调了对抗结构化跨任务语义攻击的鲁棒性评估的必要性。
cs.CV / 58 / 2607.29463

Weight-Space Mixture-of-Experts for Implicit Neural Representation Classification

权重空间专家混合模型用于隐式神经表示分类
Janik, Stanislaw, Byra, Michal
Abstract
Implicit Neural Representations (INRs) encode signals as the weights of a coordinate-based neural network and have recently been proposed as an alternative domain for downstream learning. While promising, classification directly in weight space remains challenging due to the high dimensionality and complex structure of INR parameters. Furthermore, the way discriminative information is distributed across INR weights remains poorly understood. We propose a hierarchical Mixture-of-Experts (HMoE) Transformer that processes INR weights using conditional computation aligned with the structure of the underlying implicit network. Coupled with a meta-learning framework that shapes INR parameters for downstream tasks, our model achieves state-of-the-art accuracy across standard benchmarks, ranging from low-resolution datasets to high-resolution ImageNet-1K. To gain insight into how INRs encode discriminative information, we develop weight-space attribution and pruning methods that identify parameters most relevant for classification. These analyses reveal how class-specific structure emerges within INR layers and support the suitability of MoE architectures for weight-space learning. Our approach advances both the performance and interpretability of weight-space classifiers.
Chinese Translation
隐式神经表示(INRs)将信号编码为基于坐标的神经网络的权重,最近被提出作为下游学习的替代领域。尽管前景广阔,但在权重空间中直接进行分类仍然具有挑战性,因为INR参数的高维性和复杂结构。此外,判别信息在INR权重中的分布方式仍然不够清晰。我们提出了一种层次化的专家混合(HMoE)Transformer,它通过与底层隐式网络结构对齐的条件计算来处理INR权重。结合一个塑造INR参数以适应下游任务的元学习框架,我们的模型在标准基准测试中实现了最先进的准确率,涵盖了从低分辨率数据集到高分辨率ImageNet-1K的范围。为了深入了解INR如何编码判别信息,我们开发了权重空间归因和修剪方法,以识别与分类最相关的参数。这些分析揭示了类特定结构如何在INR层中出现,并支持MoE架构在权重空间学习中的适用性。我们的方法提升了权重空间分类器的性能和可解释性。
cs.CV / 59 / 2607.29473

Lightweight Neural Networks for Affordance Segmentation: Enhancement of the Decoder Module

轻量级神经网络用于可供性分割:解码器模块的增强
Lugani, Simone, Ragusa, Edoardo, Zunino, Rodolfo, Gastaldo, Paolo
Abstract
The deployment of deep neural networks for visual affordance segmentation on wearable robots poses may prove critical, due to some conflicting aspects of the problem. On one hand, affordance segmentation requires high-level abstraction capabilities, that typically involve large-size models. On the other hand, computing resources hosted on wearable robots prevent to run large-size models in real-time. The paper presents an analysis of the role of the segmentation head in the trade-off between generalization performance and compute cost. The obtained models outperform modern baseline solutions in well-known, real-world datasets while meeting low computing requirements.
Chinese Translation
在可穿戴机器人上部署深度神经网络进行视觉可供性分割可能至关重要,因为该问题存在一些相互矛盾的方面。一方面,可供性分割需要高水平的抽象能力,通常涉及大型模型。另一方面,托管在可穿戴机器人上的计算资源限制了实时运行大型模型的能力。本文分析了分割头在泛化性能与计算成本之间权衡的作用。所获得的模型在知名的真实世界数据集上优于现代基线解决方案,同时满足低计算要求。
cs.CV / 60 / 2607.29509

Leveraging Transfer Learning with Class-Specific Decoders for Laparoscopic Segmentation

利用类特定解码器的迁移学习进行腹腔镜分割
Tomar, Priya, Parikh, Aditya, Bauckhage, Christian, Sifa, Rafet
Abstract
Effective multi-organ segmentation in surgical data requires learning the intricate anatomical features and alleviating the challenge of class imbalance, which results from relatively lower proportions of small and limitedly exposed structures. Recent works on laparoscopic multi-organ segmentation focus on learning structure-specific features through class-specific decoder architectures and report favorable results. This work extends the decoder-focused architectures to investigate knowledge sharing in the cross-surgical domain. We utilize two datasets representing different surgical domains, rectal and cholecystectomy surgeries, to explore how surgical conceptual knowledge transfers under partially common anatomical representations. Additionally, we compare the feature adaptation for the encoder and decoder at different training stages to analyse the knowledge adaptation and retention in the network. Our results corroborate previous findings on decoder-specific architectures and demonstrate that the organ-specific decoder model (CEMD), fully fine-tuned after cross-domain pre-training, achieves the highest segmentation performance (62.4\% dice) while converging substantially faster than training from scratch. However, we also find that class imbalance in surgical data remains a persistent challenge that transfer learning does not fully resolve for underrepresented anatomical structures.
Chinese Translation
在外科数据中有效的多脏器分割需要学习复杂的解剖特征,并缓解由于小型和有限暴露结构所导致的类别不平衡问题。近期关于腹腔镜多脏器分割的研究集中于通过类特定解码器架构学习结构特定特征,并报告了良好的结果。本研究扩展了以解码器为中心的架构,探讨在跨外科领域中的知识共享。我们利用两个代表不同外科领域的数据集,直肠手术和胆囊切除术,探索外科概念知识在部分共同解剖表示下的转移。此外,我们比较了在不同训练阶段编码器和解码器的特征适应性,以分析网络中的知识适应和保留。我们的结果证实了之前关于解码器特定架构的发现,并表明在跨领域预训练后完全微调的器官特定解码器模型(CEMD)实现了最高的分割性能(62.4% dice),同时收敛速度显著快于从头开始训练。然而,我们也发现外科数据中的类别不平衡仍然是一个持续的挑战,迁移学习并未完全解决对代表性不足的解剖结构的问题。
cs.CV / 61 / 2607.29531

Multi-Source Multi-View Graph Domain Adaptation with Hyperbolic Residual Encoding for Cross-Site MDD Identification from rs-fMRI

基于双曲残差编码的多源多视图图域适应用于跨站点静息态功能磁共振成像中的重度抑郁症识别
Zheng, Zhanpeng, Chen, Xiran, Jiang, Haiteng, Tian, Renjie, Cai, Qinyu, Liu, Jiexi, Chen, Xiaofeng, Li, Weikai, Wang, Yansu
Abstract
Cross-site identification of major depressive disorder (MDD) from resting-state functional magnetic resonance imaging (rs-fMRI) is hindered by inter-site distribution shifts and heterogeneous functional connectivity (FC) views. These views capture complementary neural relationships but exhibit distinct site biases and graph topologies, complicating alignment without sacrificing disease-relevant information or cross-view consistency. Existing studies largely treat multi-view connectome learning and cross-site adaptation separately. To the best of our knowledge, few studies have jointly modeled multiple FC views under multi-source unsupervised domain adaptation for cross-site rs-fMRI-based MDD classification. We construct Pearson correlation, sparse representation, and Granger causality graphs, each encoded by a view-specific graph attention network. Dual-stream adaptive fusion explicitly integrates pairwise cross-view interactions, followed by lightweight hyperbolic residual encoding for curvature-aware representation refinement. Class-wise Cauchy--Schwarz alignment reduces inter-source and source-target discrepancies, complemented by adversarial learning, information maximization, and confidence-aware pseudo-labeling. Across seven unlabeled target domains, our framework achieves 73.60% mean accuracy and 71.90% AUC, demonstrating effective generalization under heterogeneous acquisition conditions. These results highlight the effectiveness of unified heterogeneous-view modeling, curvature-aware refinement, and multi-source domain adaptation for cross-site MDD identification.The source code is at https://github.com/OPUS-Lightphenexx/MM-HyperGDA
Chinese Translation
从静息态功能磁共振成像(rs-fMRI)中跨站点识别重度抑郁症(MDD)受到站点间分布变化和异质功能连接(FC)视图的影响。这些视图捕捉了互补的神经关系,但表现出不同的站点偏差和图拓扑,复杂化了对齐过程,而不牺牲与疾病相关的信息或跨视图一致性。现有研究大多将多视图连接组学习和跨站点适应分开处理。据我们所知,鲜有研究在多源无监督领域适应下联合建模多个FC视图以进行基于rs-fMRI的跨站点MDD分类。我们构建了皮尔逊相关、稀疏表示和格兰杰因果图,每个图由特定视图的图注意力网络编码。双流自适应融合显式整合成对的跨视图交互,随后通过轻量级双曲残差编码进行曲率感知的表示精炼。类别间的柯西-施瓦茨对齐减少了源间和源-目标之间的差异,辅以对抗学习、信息最大化和基于置信度的伪标签。我们的框架在七个未标记的目标域上实现了73.60%的平均准确率和71.90%的AUC,展示了在异质采集条件下的有效泛化。这些结果突显了统一异质视图建模、曲率感知精炼和多源领域适应在跨站点MDD识别中的有效性。源代码可在 https://github.com/OPUS-Lightphenexx/MM-HyperGDA 获取。
cs.CV / 62 / 2607.29533

OSAGEN: Object-Aware Mask Priors and Multistage Decoupled Diffusion for Industrial Anomaly Generation

OSAGEN:面向对象的掩膜先验与多阶段解耦扩散在工业异常生成中的应用
Xu, Jinyi, Chen, Peng, Cao, Yunkang, Liu, Chengliang, Dong, Xinghui, Huang, Chao
Abstract
Industrial anomaly detection and localization are limited by scarce real anomalies and pixel-level annotations, a bottleneck that synthetic image-mask pairs can alleviate. However, existing few-shot mask-guided generation may over-follow mask geometry, produce weak anomalies, or use condition masks incompatible with the current object instance. We propose OSAGEN, which combines object-aware mask priors with multistage decoupled diffusion. Its three-stage adaptation sequentially learns normal appearance, defect appearance under coarse conditions, and fine-grained mask calibration, improving defect realization and local control. QBG injects object structure from a matched normal image into mask diffusion to produce object-aware priors, while ISC restricts anomaly propagation and preserves normal content during sampling. A lightweight materialization step recovers pixel-level labels aligned with the realized defects. On MVTec AD and VisA, OSAGEN achieves AP-P/F1-P scores of 88.1/82.2 and 68.5/66.1, respectively, under a unified downstream localization protocol. The code will be released upon acceptance.
Chinese Translation
工业异常检测与定位受到真实异常样本稀缺和像素级标注不足的限制,这一瓶颈可以通过合成图像-掩膜对来缓解。然而,现有的少样本掩膜引导生成可能过度遵循掩膜几何形状,产生弱异常,或使用与当前对象实例不兼容的条件掩膜。我们提出了OSAGEN,它结合了面向对象的掩膜先验与多阶段解耦扩散。其三阶段适应过程依次学习正常外观、粗略条件下的缺陷外观以及细粒度掩膜校准,从而改善缺陷实现和局部控制。QBG将来自匹配正常图像的对象结构注入掩膜扩散中,以生成面向对象的先验,而ISC在采样过程中限制异常传播并保留正常内容。轻量级的物化步骤恢复与实现的缺陷对齐的像素级标签。在MVTec AD和VisA数据集上,OSAGEN在统一的下游定位协议下分别达到了88.1/82.2和68.5/66.1的AP-P/F1-P分数。代码将在论文接受后发布。
cs.CV / 63 / 2607.29541

The K-Space Signature: Frequency-Domain Representation Learning for Medical Deepfake Detection

K空间特征:用于医学深度伪造检测的频域表示学习
Raciti, Riccardo, Guarnera, Francesco, Rundo, Francesco, Guarnera, Luca, Battiato, Sebastiano
Abstract
In medical imaging, generative models are increasingly deployed to synthesize realistic data and augment limited datasets. Unfortunately, while beneficial for privacy-preserving data sharing, these synthesized images can be repurposed for malicious intents, threatening public health through the creation of Medical Deepfakes. To address this threat, we introduce the K-Space Signature (KSS), a novel forensic framework that isolates hardware and generative traces within the spectral domain. By shifting analysis to the frequency domain, the KSS suppresses macroscopic anatomical variance by subtracting an empirical global anatomical prior computed in the Logarithmic Power Spectral Density (Log-PSD) space. To effectively process these globally distributed spectral artifacts without the local spatial bias inherent to Convolutional Neural Networks, we pair the KSS representation with a novel 3D MLP-Mixer architecture equipped with an ArcFace metric-learning head. Extensive experiments on multi-center 3D MRI datasets demonstrate that this combined approach achieves exceptional detection performance, exceeding 0.99 Accuracy and ROC-AUC on multi-generator synthetic datasets. Furthermore, the framework exhibits robust zero-shot generalization, maintaining strong discriminative power (up to 0.93 Accuracy) on independent datasets acquired from entirely unseen scanners. To ensure full reproducibility, the complete source code and pre-trained models will be made publicly available upon acceptance.
Chinese Translation
在医学成像中,生成模型越来越多地被用于合成逼真的数据和增强有限的数据集。不幸的是,尽管这对保护隐私的数据共享有益,但这些合成图像可能被恶意利用,从而通过创建医学深度伪造(Medical Deepfakes)威胁公共健康。为了解决这一威胁,我们提出了K空间特征(K-Space Signature,KSS),这是一个新颖的取证框架,能够在频谱域中隔离硬件和生成痕迹。通过将分析转移到频域,KSS通过在对数功率谱密度(Logarithmic Power Spectral Density,Log-PSD)空间中减去计算得到的经验全球解剖先验,抑制了宏观解剖变异。为了有效处理这些全球分布的频谱伪影,而不受卷积神经网络固有的局部空间偏差影响,我们将KSS表示与一种新颖的3D MLP-Mixer架构相结合,该架构配备了ArcFace度量学习头。在多中心3D MRI数据集上的广泛实验表明,这种组合方法实现了卓越的检测性能,在多生成器合成数据集上超过了0.99的准确率和ROC-AUC。此外,该框架表现出强大的零样本泛化能力,在从完全未见过的扫描仪获取的独立数据集上保持了强大的区分能力(高达0.93的准确率)。为了确保完全可重复性,完整的源代码和预训练模型将在接受后公开发布。
cs.CV / 64 / 2607.29545

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation

MoRoute:上下文多模态视频生成的动态路由
Gao, Chong, Ma, Jie, Peng, Zhan, Wang, Chongxiao, Wu, Haoxue, Liang, Jun, Li, Guanbin, Li, Jing
Abstract
Multimodal video generation aims to generate and edit videos conditioned on arbitrary combinations of text, images, and videos within a single model, allowing diverse tasks to share complementary data and generative priors. Unifying these tasks requires multimodal understanding of diverse conditions, which is typically provided by a pretrained vision-language model (VLM). A key challenge is how to connect the VLM's hierarchical multimodal representations with a pretrained video diffusion transformer (DiT). Existing methods either inject features from only the final or a few manually selected VLM layers, or jointly train architecture-matched understanding and generation streams, making it difficult to reuse heterogeneous pretrained backbones. We introduce MoRoute, a unified multimodal video generation framework that formulates a frozen VLM and a pretrained video DiT with different architectures as heterogeneous experts connected through dynamic layer routing. For each input, a lightweight block-wise router enables every DiT block to select the VLM layer most relevant to its generation stage, thereby learning an adaptive correspondence between multimodal understanding and video synthesis. MoRoute further incorporates reference images and source videos directly into the DiT token sequence through unified in-context conditioning, preserving fine-grained visual details across diverse generation and editing tasks. Experiments on IntelligentVBench, OpenVE-Bench, and RefVIE-Bench show that MoRoute consistently surpasses the best competing method on each benchmark, improving the average score by 0.15, 0.18, and 0.34 on a 1-5 scale, respectively.
Chinese Translation
多模态视频生成旨在生成和编辑基于文本、图像和视频的任意组合的条件视频,允许多样化任务共享互补数据和生成先验。统一这些任务需要对多样条件的多模态理解,这通常由预训练的视觉-语言模型(VLM)提供。一个关键挑战是如何将VLM的层次多模态表示与预训练的视频扩散变换器(DiT)连接起来。现有方法要么仅从最终层或少数手动选择的VLM层注入特征,要么共同训练架构匹配的理解和生成流,这使得重用异构预训练骨干网络变得困难。我们提出了MoRoute,一个统一的多模态视频生成框架,它将一个冻结的VLM和一个具有不同架构的预训练视频DiT构建为通过动态层路由连接的异构专家。对于每个输入,一个轻量级的块级路由器使每个DiT块能够选择与其生成阶段最相关的VLM层,从而学习多模态理解与视频合成之间的自适应对应关系。MoRoute进一步通过统一的上下文条件将参考图像和源视频直接纳入DiT的令牌序列,保留了在多样生成和编辑任务中的细粒度视觉细节。在IntelligentVBench、OpenVE-Bench和RefVIE-Bench上的实验表明,MoRoute在每个基准测试中始终超越最佳竞争方法,平均得分分别提高了0.15、0.18和0.34(满分为5分)。
cs.CV / 65 / 2607.29568

DynoDINO: Harnessing Dynamic Latent Information from DINO Features for Multi-Phase Medical Image Segmentation

DynoDINO:利用DINO特征中的动态潜在信息进行多阶段医学图像分割
Hsu, Yu-Pu, Chen, Jen-Jee, Tseng, Yu-Chee
Abstract
Multi-phase Contrast-Enhanced Computed Tomography (CECT) plays a central role in the diagnosis and characterization of focal lesions by capturing temporal enhancement patterns across multiple acquisition phases. Accurate lesion segmentation from such data remains challenging because clinically relevant contrast kinetics are distributed across phases, while anatomical inconsistencies, respiratory motion, and incomplete acquisitions often lead to inter-phase misalignment and interrupted temporal information. Conventional segmentation frameworks typically process each phase independently or rely on simple fusion strategies, limiting their temporal reasoning capability. To address these challenges, we propose DynoDINO, a unified framework tailored to address the core challenges of multi-phase medical image segmentation. DynoDINO first performs slice-level alignment to establish inter-phase anatomical correspondence and then employs a Multi-phase Fusion Model to jointly enhance temporal correlations across phases. Our fusion model incorporates a Mix-attention (MA) mechanism for efficient multi-phase feature calibration and an Adaptive Gating Mechanism with difference-based residual learning to selectively preserve diagnostically relevant contrast variations while suppressing artifacts caused by residual misalignment. In addition, the adaptive gating mechanism improves training stability by preventing feature degradation caused by unguided subtraction operations. Experiments on three large-scale datasets, including LiTS, PLC-CECT, and WAW-TACE, demonstrate that DynoDINO consistently improves boundary delineation and structural fidelity under standard, shifted, and missing-phase conditions.
Chinese Translation
多阶段对比增强计算机断层扫描(CECT)在焦点病变的诊断和特征描述中发挥着核心作用,通过捕捉多个采集阶段的时间增强模式。然而,从这些数据中准确分割病变仍然具有挑战性,因为临床相关的对比动力学分布在各个阶段,而解剖不一致、呼吸运动和不完整采集常常导致阶段间的错位和中断的时间信息。传统的分割框架通常独立处理每个阶段或依赖简单的融合策略,限制了它们的时间推理能力。为了解决这些挑战,我们提出了DynoDINO,这是一个统一框架,旨在解决多阶段医学图像分割的核心问题。DynoDINO首先执行切片级对齐,以建立阶段间的解剖对应关系,然后采用多阶段融合模型共同增强各阶段之间的时间相关性。我们的融合模型结合了混合注意力(Mix-attention, MA)机制,以高效校准多阶段特征,并采用基于差异的残差学习的自适应门控机制,以选择性地保留诊断相关的对比变化,同时抑制由残差错位引起的伪影。此外,自适应门控机制通过防止无指导的减法操作导致的特征退化,提高了训练的稳定性。在LiTS、PLC-CECT和WAW-TACE等三个大规模数据集上的实验表明,DynoDINO在标准、偏移和缺失阶段条件下始终改善了边界描绘和结构保真度。
cs.CV / 66 / 2607.29581

Explaining AI-Image Detection: What the Heatmap Actually Shows

解释人工智能图像检测:热图实际显示了什么
Kuturin, Leonid, Sotnikov, Ilya, Khusnutdinov, Mark, Potemkin, Mikhail, Baranas, Pavel, Korepanova, Aleksandra, Kalashnikov, Alexander
Abstract
A marketplace review photograph is a document: platforms approve refunds on it, and generative models drove the cost of forging one to zero. We study that detection problem, so we build a detector and attach an attribution map as its evidence, then measure what that pair delivers on 186,527 images under controls designed to change our conclusions when something is wrong. Compression history, not synthesis, drives naive evaluation: our strongest model reaches 0.9999 PR-AUC (area under the precision-recall curve) on a product-disjoint split, yet falls to 0.7254 once we re-encode synthetics into the real class's format, while five public detectors move by at most 0.07. Aligning one class relocates the cue rather than removing it, and the repaired model then assigns native files a median probability of synthesis of 0.0004. One identical final encode for both classes repairs that, and a three-seed factorial credits the encoding change with the whole gain (+0.176 +- 0.009 PR-AUC). That encode equalises the last stage only: forensic features alone still separate the classes at 0.7145 against a base rate of 0.254. For evidence we test maps causally, against controls that never consult the detector. Whether an attribution ranking exists at all depends on whether the detector reacts to the image. On our first-fix detector, which calls 96 of 100 edited frames real, no map beats a random one. On the detector we selected, twelve of seventeen maps clear that control on edited images and eight on generated ones; perturbation leads both axes and no gradient-CAM variant shows a positive advantage. The trivial controls never clear it, and on generated images the centre prior is worse than random. Our ensembled regional map clears both axes and takes the top pixel AP at 12.4 s per map against 44.9 for occlusion. Clearing a detector-blind control is not yet a faithful explanation, and we demonstrate none.
Chinese Translation
市场审查照片是一种文档:平台基于此批准退款,而生成模型使伪造一张照片的成本降至零。我们研究这一检测问题,因此构建了一个检测器,并附加了归因图作为其证据,然后在设计用于在出现问题时改变我们结论的控制下,测量该对在186,527张图像上的表现。压缩历史,而非合成,驱动了幼稚评估:我们最强的模型在产品不重叠的划分上达到了0.9999的PR-AUC(精确率-召回率曲线下面积),但一旦我们将合成图像重新编码为真实类别的格式,其性能下降至0.7254,而五个公共检测器的变化最多为0.07。对齐一个类别会重新定位线索而不是移除它,修复后的模型则给本地文件分配了0.0004的合成中位概率。对两个类别进行相同的最终编码修复了这一点,而三种种子因子分析将编码变化的全部增益归因于此(+0.176 ± 0.009 PR-AUC)。该编码仅使最后阶段均衡:法医特征仍然在0.7145的基础率0.254下区分这些类别。为了提供证据,我们因果性地测试了图,针对从未咨询检测器的控制。归因排名是否存在完全取决于检测器是否对图像做出反应。在我们的首次修复检测器上,该检测器将100张编辑帧中的96张称为真实,没有任何图超过随机图。在我们选择的检测器上,十七个图中有十二个在编辑图像上清除了该控制,八个在生成图像上清除;扰动引导了两个轴,而没有任何梯度-CAM变体显示出正优势。简单的控制从未清除它,而在生成图像上,中心先验的表现比随机还差。我们的集成区域图清除了两个轴,并以每张图12.4秒的速度在像素AP中名列前茅,而遮挡的速度为44.9秒。清除一个不依赖检测器的控制尚未成为一个可靠的解释,我们证明了没有。
cs.CV / 67 / 2607.29586

TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning

TraceViT:用于视觉抽象推理的基础追踪监督
Liu, Binnan, Ma, Yechi, Xie, Tian, Hua, Wei
Abstract
The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and apply it to a new grid. Looped visual reasoners refine predictions over multiple iterations, but conventional training constrains only the final output, leaving intermediate refinements unconstrained. We propose that these refinements should instead follow the transformation step by step. We introduce TraceViT, a looped visual reasoner trained with semantically monotonic transformation chains. We obtain these chains by rewriting and verifying programmatic task implementations, decomposing each solution into intermediate grid states. Each iteration is grounded by a task reference derived from the few-shot demonstrations and an object workspace representing the current grid state. Because these chains may differ in length from the loop, soft trace alignment enforces only their ordering, letting the model allocate iterations freely. TraceViT achieves 67.8% pass@2 on ARC-AGI-1 and 24.3% on ARC-AGI-2. Controlled ablations on ARC-AGI-1 show that trace supervision becomes beneficial only when paired with grounding. Code and data will be available at https://github.com/LiuBinnan/TraceViT.
Chinese Translation
抽象与推理语料库(ARC)测试模型是否能够从少量输入-输出示例中推断出未见的变换,并将其应用于新的网格。循环视觉推理器在多个迭代中细化预测,但传统训练仅约束最终输出,留下中间细化过程不受约束。我们提出这些细化应当逐步遵循变换过程。我们引入了TraceViT,这是一种使用语义单调变换链进行训练的循环视觉推理器。我们通过重写和验证程序任务实现来获得这些链,将每个解决方案分解为中间网格状态。每次迭代都由源自少量示范的任务参考和表示当前网格状态的对象工作空间所支撑。由于这些链的长度可能与循环不同,软追踪对齐仅强制它们的顺序,使模型可以自由分配迭代。TraceViT在ARC-AGI-1上实现了67.8%的pass@2,在ARC-AGI-2上实现了24.3%。对ARC-AGI-1的控制消融实验表明,追踪监督只有在与基础相结合时才变得有益。代码和数据将发布在https://github.com/LiuBinnan/TraceViT。
cs.CV / 68 / 2607.29592

TOOD: Task-Aware Out-of-Distribution Score Calibration for Continual Learners

TOOD:面向任务的持续学习者的分布外得分校准
ElAraby, Mostafa, Nashed, Samer B., Paull, Liam
Abstract
The primary challenge of continual learning (CL) systems is to learn new tasks while remaining performant on previously learned tasks. A similarly important though less well-studied aspect of CL systems is their ability to distinguish inputs that are unlikely to come from within the set of tasks the system has already encountered, often called out-of-distribution (OOD) detection. This paper presents several findings related to the dynamics of OOD detection in CL systems, causes of performance degradation over time which we call OOD forgetting (OODF), and proposed mitigation strategies for this degradation. Chiefly, we find the unintuitive result that OODF is only weakly anti-correlated with classification performance on previous tasks, suggesting that the underlying mechanisms producing OODF are distinct. Moreover, this effect is observed for both energy-based and feature-based OOD detection methods. Energy-based detectors suffer a drop in logit scale as additional tasks are learned, which we term the Confidence Gap, while feature-based detectors also degrade under a complementary effect we call Manifold Crowding. Motivated by these observations, we propose TOOD, a training-free post-hoc method that decomposes logits into per-task energy scores and re-calibrates them using replay-buffer statistics. Experiments on CIFAR-10, CIFAR-100, and a 100-task ImageNet-1K stream show that TOOD improves OOD detection performance over uncalibrated energy in most settings and ranks first or second in nine of ten CIFAR configurations, with the largest gains when the confidence gap is most severe. These results suggest that a substantial portion of OOD deterioration in continual learning arises from score miscalibration rather than from a complete loss of discriminative structure.
Chinese Translation
持续学习(CL)系统的主要挑战在于在学习新任务的同时保持对先前学习任务的良好性能。CL系统的另一个同样重要但研究较少的方面是其区分输入的能力,这些输入不太可能来自系统已经遇到的任务集合,通常称为分布外(OOD)检测。本文提出了与CL系统中OOD检测动态相关的若干发现,包括我们称之为OOD遗忘(OODF)的性能随时间下降的原因,以及针对这种下降的缓解策略。我们发现一个反直觉的结果,即OODF与先前任务的分类性能仅弱相关,表明产生OODF的潜在机制是不同的。此外,这种效应在基于能量和基于特征的OOD检测方法中均有观察到。基于能量的检测器在学习额外任务时,logit尺度下降,我们称之为置信差距,而基于特征的检测器在我们称之为流形拥挤的互补效应下也会退化。基于这些观察,我们提出了TOOD,这是一种无训练的后处理方法,将logits分解为每个任务的能量得分,并使用重放缓冲区统计数据对其进行重新校准。在CIFAR-10、CIFAR-100和100任务的ImageNet-1K流上进行的实验表明,TOOD在大多数设置中提高了OOD检测性能,相较于未校准的能量,在十个CIFAR配置中排名第一或第二,当置信差距最严重时,提升最大。这些结果表明,持续学习中OOD性能下降的很大一部分源于得分的错误校准,而不是完全丧失判别结构。
cs.CV / 69 / 2607.29595

CoDe-SSM: Context-Detail Decoupled State Space Model for Efficient UHD Image Restoration

CoDe-SSM:用于高效超高清图像恢复的上下文-细节解耦状态空间模型
Su, Jiaxu, Wu, Zhijian, Li, Jun, Zhang, Bo, Zheng, Yefeng
Abstract
Ultra-high-definition (UHD) image restoration must balance the aggregation of spatially recurring degradation cues with the preservation of localized image structures. Compact aggregation can reduce redundant processing but may attenuate edges, textures, and other fine structures. Existing approaches manage UHD restoration cost through downsampling, window partitioning, or cluster-based token reduction; yet many of them do not explicitly retain information that is poorly represented by shared aggregation. In this study, we propose a Context-Detail Decoupled State Space Model (CoDe-SSM) for UHD restoration, which processes aggregated context and clustering residuals in separate pathways. The context modeling pathway, implemented by the Global Cluster Scan Module (GCSM), aggregates features into $K$ input-dependent cluster centers and applies selective SSM reasoning over the resulting fixed-order sequence, enabling cross-region context sharing while decoupling computational cost from spatial resolution. The detail recovery pathway, implemented by the Local High-Frequency Module (LHFM), processes the clustering residual with an input-derived high-frequency mask and a sparse mixture of convolutional experts. Extensive experiments on five UHD benchmarks and five degradation types demonstrate that our explicit context-detail decoupling strategy yields substantial gains in restoration quality while maintaining desirable efficiency.
Chinese Translation
超高清(UHD)图像恢复必须在聚合空间上重复的退化线索与保留局部图像结构之间取得平衡。紧凑的聚合可以减少冗余处理,但可能会削弱边缘、纹理和其他细微结构。现有方法通过下采样、窗口分区或基于聚类的令牌减少来管理UHD恢复成本;然而,它们中的许多并未明确保留由共享聚合所表现不佳的信息。在本研究中,我们提出了一种用于UHD恢复的上下文-细节解耦状态空间模型(CoDe-SSM),该模型在独立的路径中处理聚合的上下文和聚类残差。上下文建模路径通过全局聚类扫描模块(Global Cluster Scan Module, GCSM)实现,将特征聚合到$K$个输入依赖的聚类中心,并对生成的固定顺序序列应用选择性状态空间模型推理,从而实现跨区域的上下文共享,同时将计算成本与空间分辨率解耦。细节恢复路径通过局部高频模块(Local High-Frequency Module, LHFM)实现,使用输入衍生的高频掩模和稀疏的卷积专家混合处理聚类残差。在五个UHD基准和五种退化类型上的广泛实验表明,我们明确的上下文-细节解耦策略在恢复质量上带来了显著的提升,同时保持了理想的效率。
cs.CV / 70 / 2607.29627

FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

FlexComposer:从图像到动态视频的统一视频合成与灵活轨迹控制
Zhang, Songchun, Guo, Sitong, Kong, Xianghao, Liu, Pengwei, Guo, Yuwei, Zhang, Lvmin, Rao, Anyi
Abstract
Generative video compositing, which involves inserting external assets seamlessly into existing video sequences, is essential for content creation and visual effects. However, existing approaches suffer from a control-fidelity trade-off: they either hallucinate motion from static images, failing to preserve the dynamics of pre-animated assets, or lack fine-grained spatial control for precise asset placement along user-defined trajectories. We propose FlexComposer, a unified framework that standardizes video compositing as a trajectory-guided conditional generation task, enabling the seamless integration of both static images and dynamic footage. Our approach introduces three key designs: (1) a Unified Canonical Foreground Representation that decouples an object's intrinsic motion from its global displacement, standardizing heterogeneous inputs into a stabilized, centered latent space; (2) a Spatial-Aware Latent Injection strategy that exploits the translation equivariance of VAE latent spaces to transport canonical features onto target trajectories via a parameter-free mechanism; and (3) a Hybrid Dataset and Synthetic-to-Real Curriculum that synergizes procedural simulation, real-world cinematic footage, and generative data to implicitly learn physically plausible illumination and shadow harmonization. This unified design handles diverse inputs from product photos to dynamic subjects achieving high-fidelity motion control and environmental integration without the need for explicit 3D reconstruction or auxiliary learnable adapters. Extensive experiments demonstrate that FlexComposer outperforms state-of-the-art methods in visual quality, temporal consistency, and trajectory adherence.
Chinese Translation
生成视频合成涉及将外部素材无缝插入现有视频序列,对于内容创作和视觉效果至关重要。然而,现有方法面临控制与保真度之间的权衡:要么从静态图像中幻觉出运动,未能保持预先动画素材的动态性,要么缺乏精细的空间控制,无法沿用户定义的轨迹精确放置素材。我们提出了FlexComposer,一个将视频合成标准化为轨迹引导的条件生成任务的统一框架,使静态图像和动态视频的无缝集成成为可能。我们的方法引入了三个关键设计:(1)统一的典型前景表示,解耦对象的内在运动与其全局位移,将异构输入标准化为一个稳定的、居中的潜在空间;(2)空间感知潜在注入策略,利用变分自编码器(VAE)潜在空间的平移等变性,通过无参数机制将典型特征传输到目标轨迹;(3)混合数据集和合成到真实的课程,协同程序化模拟、真实世界的电影镜头和生成数据,隐式学习物理上合理的照明和阴影协调。这一统一设计能够处理从产品照片到动态对象的多样输入,实现高保真的运动控制和环境整合,而无需显式的三维重建或辅助可学习适配器。大量实验表明,FlexComposer在视觉质量、时间一致性和轨迹遵循方面优于最先进的方法。
cs.CV / 71 / 2607.29633

OASIS: Occlusion-aware Single-image Hand Avatar Reconstruction via 3D Gaussian Splatting

OASIS:基于3D高斯点云的遮挡感知单幅图像手部虚拟形象重建
Han, Zhisheng, Wu, Shiyao, Qiu, Jiayan, Ju, Yakun, Liu, Lu, Zhang, Le, Feng, Pengfei, Zhou, Huiyu, Jiang, Zheheng
Abstract
Single-image 3D hand avatar reconstruction is fundamentally ill-posed and particularly challenging due to limited visual evidence under severe self-occlusion and the complex pose-dependent deformation of highly articulated hands. Existing methods predominantly rely on implicit NeRF-style representations, whose volumetric fitting is computationally expensive and often struggles to preserve fine-grained hand details. In this work, we present OASIS, a tailored 3D Gaussian Splatting framework for single-image hand avatar reconstruction. To faithfully encode sparse image-specific appearance cues in single-view reconstruction, we construct geometry-aligned visual evidence tokens by explicitly aligning input image observations with 3D hand geometry and context-adaptively tokenizing the resulting visual evidence. Since severe self-occlusion makes the reliability of image evidence inherently visibility-dependent, we introduce a visibility-conditioned point-image attention to reliably transfer visual evidence to geometric tokens, yielding occlusion-aware Gaussian features for faithful and robust reconstruction. To further capture non-rigid deformation of articulated hands, we introduce a Feature-on-Mesh representation to enable Gaussian deformation to be guided by local surface stretching. Under this framework, we adopt a one-shot adaptation scheme that learns a shared hand prior from multi-identity training data and then fits it to a target image for target-specific reconstruction. Extensive experiments show that OASIS outperforms existing baselines in both visual fidelity and efficiency across challenging poses and in-the-wild scenarios, and further demonstrates strong versatility in downstream applications such as text-to-avatar generation and texture editing.
Chinese Translation
单幅图像的3D手部虚拟形象重建本质上是一个不适定问题,尤其在严重自遮挡和高度关节化手部的复杂姿态依赖变形下,因视觉证据有限而极具挑战性。现有方法主要依赖隐式NeRF风格的表示,其体积拟合计算开销大,且往往难以保留细致的手部细节。在本研究中,我们提出了OASIS,一个专门为单幅图像手部虚拟形象重建设计的3D高斯点云框架。为了在单视图重建中真实编码稀疏的图像特定外观线索,我们通过将输入图像观测与3D手部几何形状明确对齐,并上下文自适应地对结果视觉证据进行标记,构建几何对齐的视觉证据标记。由于严重的自遮挡使得图像证据的可靠性本质上依赖于可见性,我们引入了一种可见性条件的点-图像注意机制,以可靠地将视觉证据转移到几何标记上,从而生成遮挡感知的高斯特征,实现真实且稳健的重建。为了进一步捕捉关节手部的非刚性变形,我们引入了一种网格特征表示,使高斯变形能够受到局部表面拉伸的引导。在此框架下,我们采用了一种一次性适应方案,从多身份训练数据中学习共享的手部先验,然后将其拟合到目标图像上以进行目标特定重建。大量实验表明,OASIS在视觉保真度和效率方面超越了现有基准,尤其在具有挑战性的姿态和实际场景中,进一步展示了在下游应用(如文本到虚拟形象生成和纹理编辑)中的强大通用性。
cs.CV / 72 / 2607.29637

CodeShrink: Adaptive Visual Compression for Efficient Multimodal Code Understanding

CodeShrink:用于高效多模态代码理解的自适应视觉压缩
Tang, Wenxin, Xiao, Jingyu, Liu, Zhenyu, Xie, Zipeng, Liu, Junliang, Luo, Wang, Jiang, Yuan, Huo, Yintong, Lyu, Michael
Abstract
Rendering source code as images offers a promising way to reduce the input costs of Multimodal Large Language Models (MLLMs). Adjusting image resolution can trade visual token cost against content fidelity. However, resolution scaling alone overlooks two sources of inefficiency: blank regions created by line breaks and indentation, and code regions irrelevant to the current instruction. Moreover, the best compression setting varies across inputs, tasks, and models, limiting fixed-ratio strategies. We propose CodeShrink, an adaptive visual compression framework with three components. Blank-Free Rendering replaces whitespace-dependent layouts with compact layouts and explicit structural markers, removing layout-induced tokens. Adaptive Compression Configuration uses a lightweight agent trained with reinforcement learning to predict a per-input setting that balances token efficiency and readability. Dominant Token Selection jointly analyzes the instruction and code image to prune task-irrelevant visual tokens during inference. We evaluate CodeShrink on code question answering, clone detection, and code completion. CodeShrink reduces visual token use by up to 71.2\% while matching or exceeding uncompressed text-only inputs, and consistently outperforms text-based and visual compression baselines across all three tasks. These results show that combining layout compaction, adaptive configuration, and instruction-aware pruning can make multimodal code understanding more efficient. Our code is available at https://github.com/vinsontang1/CodeShrink.
Chinese Translation
将源代码呈现为图像为减少多模态大型语言模型(MLLMs)的输入成本提供了一种有前景的方法。调整图像分辨率可以在视觉标记成本与内容保真度之间进行权衡。然而,仅仅依靠分辨率缩放忽视了两种低效来源:由换行和缩进产生的空白区域,以及与当前指令无关的代码区域。此外,最佳压缩设置因输入、任务和模型而异,这限制了固定比例策略。我们提出了CodeShrink,一个具有三个组成部分的自适应视觉压缩框架。无空白渲染(Blank-Free Rendering)用紧凑布局和显式结构标记替代依赖于空白的布局,消除由布局引起的标记。自适应压缩配置(Adaptive Compression Configuration)使用经过强化学习训练的轻量级代理来预测一个平衡标记效率和可读性的每输入设置。主导标记选择(Dominant Token Selection)在推理过程中共同分析指令和代码图像,以修剪与任务无关的视觉标记。我们在代码问答、克隆检测和代码补全任务上评估了CodeShrink。CodeShrink在视觉标记使用上减少了高达71.2\%,同时与未压缩的仅文本输入相匹配或超越,并在所有三个任务中始终优于基于文本和视觉压缩的基准。这些结果表明,结合布局压缩、自适应配置和指令感知修剪可以使多模态代码理解更加高效。我们的代码可在 https://github.com/vinsontang1/CodeShrink 获取。
cs.CV / 73 / 2607.29638

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering

HierDoc:用于长文档视觉问答的层次化页面到区域证据路由
Gu, Rongjian, Zhou, Wengang, Xiong, Junyu, Wang, Yonghui, Yin, Bing, Wang, Bei, Li, Houqiang
Abstract
Multi-page document visual question answering requires locating sparse evidence at both the page and region levels. Existing approaches typically emphasize one level over the other: page-centric methods focus on page acquisition, with region operations serving mainly as navigation aids, whereas region-centric methods assume that the relevant pages have already been supplied. Consequently, page and region selection remain disconnected rather than forming successive evidence decisions. We propose HierDoc, a hierarchical evidence-routing framework that formulates long-document evidence acquisition as two-stage set prediction from pages to regions. A page policy selects evidence pages from the full document; these pages are then parsed for semantic elements, after which a region policy selects the elements passed to a downstream answer model. Both answer-agnostic policies are optimized with stage-wise GRPO using granularity-specific structured-set rewards. The answer model receives selected full pages together with selected region crops and OCR or table text, preserving global context while emphasizing fine-grained evidence. Across the evaluated benchmarks, HierDoc achieves state-of-the-art or competitive performance among open-weight systems, improving LongDocURL by 16.87% relative to the strongest reported open-weight baseline. Controlled ablations further show that selected regional evidence improves the page-only system in accuracy and F1 by 5.51% and 4.82%, respectively. These results demonstrate the benefit of organizing coarse page routing and fine-grained region routing as successive, separately optimized stages of a unified evidence-acquisition process.
Chinese Translation
多页文档的视觉问答需要在页面和区域级别定位稀疏证据。现有方法通常强调一个层级而忽视另一个层级:以页面为中心的方法专注于页面获取,而区域操作主要作为导航辅助;而以区域为中心的方法则假设相关页面已经提供。因此,页面和区域的选择并未形成连续的证据决策。我们提出了HierDoc,一个层次化证据路由框架,将长文档的证据获取形式化为从页面到区域的两阶段集合预测。页面策略从完整文档中选择证据页面;然后对这些页面进行解析以提取语义元素,之后区域策略选择传递给下游答案模型的元素。这两种与答案无关的策略通过使用特定粒度的结构化集合奖励进行阶段性GRPO优化。答案模型接收选定的完整页面以及选定的区域裁剪和OCR或表格文本,保留全局上下文的同时强调细粒度证据。在评估的基准测试中,HierDoc在开放权重系统中实现了最先进或具有竞争力的性能,相较于报告的最强开放权重基线,LongDocURL提高了16.87%。控制性消融实验进一步表明,所选区域证据在准确性和F1分数上分别提高了页面仅系统5.51%和4.82%。这些结果证明了将粗略页面路由和细粒度区域路由组织为统一证据获取过程的连续、单独优化阶段的好处。
cs.CV / 74 / 2607.29679

Scaling Properties of Text Conditioning in Visual Generation

视觉生成中文本条件的尺度特性
Chen, Zilong, Deng, Chaorui, Li, Kunchang, Yuan, Hongyi, Fan, Haoqi
Abstract
We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, we adapt two complementary measures: a white-box likelihood metric (GPG) and a black-box attribute metric (ED). Across controlled training runs, the converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED. Guided by these scaling properties, we improve \emph{diffusability} by constructing structured prompts with semantic and geometric annotations derived from images, and improve \emph{promptability} by training a prompter through supervised fine-tuning, cold-start, and verifier-gated on-policy distillation. The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations.
Chinese Translation
我们研究了视觉生成中文本条件的经验尺度特性。这些特性很少被测量,因为扩散损失与自然语言提示中的标记数量并不成比例。令人惊讶的是,我们发现收敛的扩散损失与提示中结构化语言的数量成比例。为了量化结构化语言,我们采用了两种互补的度量:一种是白盒似然度量(GPG),另一种是黑盒属性度量(ED)。在受控的训练过程中,收敛的扩散损失与 GPG 近似线性下降,并且与 ED 遵循幂律关系。在这些尺度特性的指导下,我们通过构建包含来自图像的语义和几何注释的结构化提示来提高 extit{扩散性},并通过监督微调、冷启动和验证者门控的在线蒸馏来提高 extit{提示能力}。最终的系统在几乎所有的组合、推理和世界知识基准测试中都优于所有评估的开放权重模型,并且在大多数评估中与最强的闭合权重模型相匹配或超越。
cs.CV / 75 / 2607.29684

Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark

朝向稳健且具三维感知的暗光RGB-NIR成像
Niu, Muyao, Ma, Mingze, Zhan, Yifan, Zhu, Qingtian, Zhong, Zhihang, Guo, Wei, Chen, Chang Wen, Zheng, Yinqiang
Abstract
Robust low-light imaging remains challenging for the community. Recent studies have explored fusing Near-Infrared (NIR) with noisy RGB to achieve improved enhancement, yet most methods depend on carefully curated training data pairs, with limited robustness under different scenarios. This paper offers a new perspective for RGB-NIR low-light imaging by incorporating 3D-aware neural modeling. Without using clean RGB supervision, a powerful model can be optimized to implicitly fuse extremely noisy RGB observations with NIR cues in 3D space, effectively recovering clean RGB images. The proposed model obviates the requirement for clean RGB data collection, generalizes across different noise levels. Extensive evaluations on synthetic and real data demonstrate its superiority. Codes available: https://github.com/MyNiuuu/3DarkFusion
Chinese Translation
稳健的低光成像仍然是一个挑战。近期研究探讨了将近红外(NIR)与噪声RGB融合以实现增强效果,然而大多数方法依赖于精心策划的训练数据对,在不同场景下的鲁棒性有限。本文通过引入具三维感知的神经建模,为RGB-NIR低光成像提供了一种新视角。在不使用干净RGB监督的情况下,强大的模型能够在三维空间中隐式融合极其嘈杂的RGB观测与NIR线索,有效恢复干净的RGB图像。所提出的模型消除了对干净RGB数据收集的需求,并在不同噪声水平下具有良好的泛化能力。在合成和真实数据上的广泛评估证明了其优越性。代码可在此获取: https://github.com/MyNiuuu/3DarkFusion
人工智能 (Artificial Intelligence)
43
cs.AI / 1 / 2607.28629

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems

OpenClaw与Ollama在自主人工智能中的应用:迈向完全自主和可扩展的人工智能代理系统
Roumeliotis, Konstantinos I., Sapkota, Ranjan
Abstract
The rapid transition from reactive large language models (LLMs) to persistent, action-capable systems has exposed critical gaps in the architectural understanding of Agentic AI, particularly in separating inference, orchestration, and execution layers for autonomous AI agents. Despite recent advances, unified frameworks for designing and evaluating full-stack agentic systems remain limited. This paper presents a comprehensive, layered architecture for Agentic AI, outlining the evolution from reactive LLM interfaces to persistent, goal-driven autonomous AI agents with memory, planning, and continuous execution. We analyze OpenClaw and Ollama as a full-stack Agentic AI system, where Ollama serves as the LLM inference layer and OpenClaw enables agent runtime orchestration, integrating reasoning, tool use, and action execution. A prototype experimental validation of the OpenClaw-Ollama architecture demonstrates that capabilities such as persistent memory, tool utilization, and adaptive decision-making emerge from system-level integration rather than standalone models, with performance improving consistently as architectural complexity increases. The study further examines challenges in scalability, security, privacy, governance, and evaluation of agentic systems, highlighting the need for robust benchmarking and system-level design. Future directions include scalable multi-agent architectures, distributed autonomous systems, and human-aware Agentic AI frameworks for responsible deployment. Overall, this work establishes a unified architectural foundation for Agentic AI, validates the effectiveness of full-stack autonomous AI agents, and provides a roadmap for building scalable, secure, and trustworthy agentic systems. All models, code, and datasets are publicly released to support reproducibility and benchmarking.
Chinese Translation
从反应式大型语言模型(LLMs)到持久、具备行动能力的系统的快速转变暴露了对自主人工智能(Agentic AI)架构理解中的关键缺口,特别是在分离推理、协调和执行层面方面。尽管近期取得了一些进展,但用于设计和评估全栈自主系统的统一框架仍然有限。本文提出了一种全面的分层架构用于自主人工智能,概述了从反应式LLM接口到具备记忆、规划和持续执行能力的持久、目标驱动的自主人工智能代理的演变。我们分析了OpenClaw与Ollama作为一个全栈自主人工智能系统,其中Ollama作为LLM推理层,OpenClaw则实现了代理运行时的协调,整合了推理、工具使用和行动执行。OpenClaw-Ollama架构的原型实验验证表明,持久记忆、工具利用和自适应决策等能力源于系统级集成而非独立模型,且随着架构复杂性的增加,性能持续提升。研究进一步探讨了自主系统在可扩展性、安全性、隐私、治理和评估方面的挑战,强调了建立健全基准测试和系统级设计的必要性。未来的方向包括可扩展的多代理架构、分布式自主系统以及负责任部署的人类感知自主人工智能框架。总体而言,本研究为自主人工智能建立了统一的架构基础,验证了全栈自主人工智能代理的有效性,并为构建可扩展、安全和可信的自主系统提供了路线图。所有模型、代码和数据集均已公开发布,以支持可重复性和基准测试。
cs.AI / 2 / 2607.28631

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

人工智能能评估人工智能科学家吗?使用自动化多模型评审的自主研究生成系统基准研究
Ravideshik, Vaibhava Lakshmi, Kejriwal, Mayank
Abstract
AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and comparing the quality of AI-generated papers remains an open challenge. We propose and implement a rigorous benchmarking protocol using an automated peer-review system that harnesses frontier large language models to assess scientific papers across four core dimensions: originality, scientific rigor, clarity, and significance. We evaluate four leading AI Scientist frameworks: \textit{Sakana AI (v1 & v2)}, \textit{CycleResearcher}, and \textit{Data-to-Paper}. Each framework was run on a consistent set of 15 research proposals published by a commercial autonomous AI scientist company (FARS), generating 60 papers that we evaluate alongside 15 FARS benchmark papers. Using three independent LLM reviewers (GPT-5.4, Gemini, and Claude), we find that FARS benchmark papers significantly outperform all competing frameworks, achieving mean scores of 2.14--2.47 on a 1--5 scale compared to 1.00--1.87 for other systems. Notably, FARS scores are more than 2$\times$ higher than the next-best systems on Gemini and Claude evaluations. We find strong agreement among Gemini and Claude ($\rho$ = 0.907, $p < 0.001$), and both correlate extremely strongly with the synthesis score ($\rho$ = 0.961, $p < 0.001$), validating the reliability of automated evaluation. However, GPT-5.4 exhibits weaker agreement ($\rho \approx 0.32$), suggesting it evaluates papers using different criteria. These results establish the first quantitative benchmark for AI Scientist systems and demonstrate that multi-model LLM evaluation provides a scalable, consistent framework for assessing autonomous research quality.
Chinese Translation
具备自主研究能力的人工智能科学家系统有潜力显著加速科学发现。然而,评估和比较人工智能生成论文的质量仍然是一个未解决的挑战。我们提出并实施了一种严格的基准测试协议,使用自动化同行评审系统,利用前沿的大型语言模型评估科学论文的四个核心维度:原创性、科学严谨性、清晰度和重要性。我们评估了四个领先的人工智能科学家框架: extit{Sakana AI (v1 & v2)}、 extit{CycleResearcher}和 extit{Data-to-Paper}。每个框架在一组由商业自主人工智能科学家公司(FARS)发布的15个研究提案上运行,生成了60篇论文,我们将其与15篇FARS基准论文进行评估。使用三位独立的LLM评审者(GPT-5.4、Gemini和Claude),我们发现FARS基准论文在所有竞争框架中显著优于其他系统,在1-5的评分尺度上取得了2.14-2.47的平均分,而其他系统的平均分为1.00-1.87。值得注意的是,FARS的得分在Gemini和Claude评估中比第二好的系统高出2倍以上。我们发现Gemini和Claude之间的强一致性($ ho$ = 0.907, $p < 0.001$),并且两者与综合得分的相关性极强($ ho$ = 0.961, $p < 0.001$),验证了自动评估的可靠性。然而,GPT-5.4的评估一致性较弱($ ho ext{approx} 0.32$),这表明它使用了不同的标准来评估论文。这些结果建立了人工智能科学家系统的第一个定量基准,并展示了多模型LLM评估为评估自主研究质量提供了一个可扩展、一致的框架。
cs.AI / 3 / 2607.28632

LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Riemann Hypothesis

发现主要数学猜想的LLM框架:人工智能对下一个黎曼假设的探索
Wong, Alizer, Zeng, Zixin, Tan, Yi, Li, Wenyuan, Chen, Xuhang, Lai, Xingru, Shi, Yang, Lu, Liangsi, Chen, Yanhui
Abstract
Major mathematical conjectures still depend heavily on expert intuition, so a unified method for the systematic generation and validation of conjectures with substantial mathematical potential remains unavailable. We present a three stage pipeline for major conjecture discovery, with region search from explicit local evidence modules, reflective validation for foundationality, novelty, and potential significance, and formal validation in Lean 4 and Mathlib. The objective is the discovery of mathematical problems with high problem taste, namely problems whose proofs could reorganize the language of a research area and provide durable help to human mathematical research. Experiments on twenty candidates showstable passage from natural language to formal checks, with twenty out of twenty candidates passing Lean parsing and type checking, twenty out of twenty candidates not directly absorbed by exact?,twenty out of twenty candidates not automatically discharged by aesop, and no explicit duplicates or near duplicates.
Chinese Translation
主要数学猜想仍然在很大程度上依赖于专家的直觉,因此尚未出现一种统一的方法来系统地生成和验证具有重要数学潜力的猜想。我们提出了一个三阶段的主要猜想发现流程,包括从显式局部证据模块进行区域搜索、对基础性、新颖性和潜在重要性进行反思性验证,以及在Lean 4和Mathlib中进行形式验证。我们的目标是发现具有高问题品味的数学问题,即那些其证明能够重组研究领域语言并为人类数学研究提供持久帮助的问题。对二十个候选者的实验显示,从自然语言到形式检查的稳定转换,其中二十个候选者均通过了Lean解析和类型检查,二十个候选者未被exact?直接吸收,二十个候选者未被aesop自动解放,且没有显式重复或近似重复的情况。
cs.AI / 4 / 2607.28642

ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning

ThinkReset:可学习的中间接口构建用于有限上下文的长时间推理
Ding, Fei, Zhang, Yongkang, Liu, Runhao, Liao, Yuhao, Zeng, Zijian
Abstract
Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow, and error anchoring. We argue that under bounded context windows, the core bottleneck is not trajectory compression or test-time control, but the absence of a reusable intermediate interface that can replace discarded history and support continued solving. We further identify a key failure mode of outcome-reward-driven long-chain reinforcement learning: when the model has not solved the task before the window is nearly exhausted, the final-answer reward encourages premature guessing rather than continued careful reasoning. We propose ThinkReset, a text-space instantiation of this view. ThinkReset explicitly constructs reusable intermediate interfaces through interface writeback and reset, and directly optimizes post-reset continuation success. Across multiple long-horizon reasoning benchmarks, this perspective consistently improves success rates under fixed context windows.
Chinese Translation
长链思维推理提高了复杂问题的解决能力,但也引入了冗余累积、上下文溢出和错误锚定。我们认为,在有限上下文窗口下,核心瓶颈并不是轨迹压缩或测试时控制,而是缺乏一个可重用的中间接口,该接口可以替代被丢弃的历史并支持持续解决。我们进一步识别出以结果奖励驱动的长链强化学习的一种关键失败模式:当模型在窗口几乎耗尽之前尚未解决任务时,最终答案奖励会鼓励过早猜测,而不是继续仔细推理。我们提出了ThinkReset,这是这一观点的文本空间实例。ThinkReset通过接口回写和重置显式构建可重用的中间接口,并直接优化重置后的继续成功率。在多个长时间推理基准测试中,这一视角在固定上下文窗口下始终提高了成功率。
cs.AI / 5 / 2607.28657

TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter

TAPR:通过任务感知的提示重写器提升大型语言模型性能
Savolainen, Oliver, Bastianelli, Emanuele, Azarbonyad, Hosein
Abstract
Large Language Models (LLMs) often require carefully crafted prompts to unlock their full potential, which can be a barrier for non-expert users. This work addresses the challenge by introducing a Task-Aware Prompt Rewriter (TAPR), a model that reformulates user prompts into task-optimized prompts with the explicit goal of improving downstream LLM performance. We train TAPR using reinforcement learning with Group Relative Policy Optimization (GRPO), where rewards are derived from LLM-as-judge evaluations of both the reformulated prompt and the corresponding task output. Experimental results on diverse tasks, such as question answering, summarization, and arithmetic reasoning, show that our method yields consistent gains over base models in prompt rewriting ability. Fine-tuning Phi-4-mini-instruct (as the base model for TAPR) produces prompts that contain clearer and more instructive language, leading to higher accuracy on established benchmarks such as Natural Questions and GSM8K. Our code is available at: https://github.com/OliverSavolainen/task-specific-prompt-rewriter
Chinese Translation
大型语言模型(LLMs)通常需要精心设计的提示才能充分发挥其潜力,这对非专业用户来说可能构成障碍。本研究通过引入任务感知提示重写器(TAPR)来解决这一挑战,该模型将用户提示重构为任务优化的提示,明确旨在提高下游LLM性能。我们使用带有组相对策略优化(Group Relative Policy Optimization, GRPO)的强化学习来训练TAPR,其中奖励来自于LLM作为评判者对重构提示及其对应任务输出的评估。在多种任务上的实验结果,例如问答、摘要生成和算术推理,表明我们的方法在提示重写能力上相较于基础模型取得了一致的提升。对Phi-4-mini-instruct(作为TAPR的基础模型)进行微调,生成的提示包含更清晰和更具指导性的语言,从而在自然问题(Natural Questions)和GSM8K等已建立基准上实现更高的准确性。我们的代码可在以下链接获取:https://github.com/OliverSavolainen/task-specific-prompt-rewriter
cs.AI / 6 / 2607.28659

Empowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel Decoding

通过混合标记化和串行-并行解码增强跨域序列推荐
Hu, Yuxuan, Wang, Yuhao, Huang, Tianbo, Zhang, Chao, Liu, Ziwei, Zhang, Lihua, Zhao, Xiangyu
Abstract
Cross-domain sequential recommendation (CDSR) aims to model users' dynamic interest transitions and sequential patterns across multiple domains. Recently, generative recommendation (GR) has emerged. It first learns semantic identifiers (SIDs) from item semantics and formulates recommendation as autoregressive generation. However, existing methods face two critical issues: (1) they ignore collaborative correlations across domains during tokenization, and (2) they adopt inefficient decoding strategies, such as beam search, during generation, which hinders real-time deployment. To address these limitations, we propose GenCDSR, an effective and efficient generative framework for CDSR. Specifically, we design a cross-domain hybrid tokenization mechanism with a multi-tower architecture to jointly capture cross-domain commonalities and domain-specific distinctions through hierarchical shared-specific and fine-grained codebooks. Furthermore, we develop a cross-domain serial-parallel decoding strategy that leverages the hierarchical SID structure to partially parallelize generation, significantly reducing inference latency while preserving generation consistency. Experiments on three public datasets show that GenCDSR achieves an average accuracy improvement of 1.5 percent and an average inference latency reduction of 85.1 percent compared with state-of-the-art baselines. The implementation code and datasets are available online: https://github.com/Applied-Machine-Learning-Lab/RecSys2026_GenCDSR.
Chinese Translation
跨域序列推荐(CDSR)旨在建模用户在多个领域中的动态兴趣转变和序列模式。最近,生成推荐(GR)逐渐兴起。它首先从项目语义中学习语义标识符(SIDs),并将推荐形式化为自回归生成。然而,现有方法面临两个关键问题:(1)在标记化过程中忽视了跨域的协作关联;(2)在生成过程中采用了低效的解码策略,如束搜索,这妨碍了实时部署。为了解决这些局限性,我们提出了GenCDSR,一个有效且高效的CDSR生成框架。具体而言,我们设计了一种跨域混合标记化机制,采用多塔架构,通过层次共享特定和细粒度代码本共同捕捉跨域的共性和领域特定的差异。此外,我们开发了一种跨域串行-并行解码策略,利用层次SIDs结构部分并行化生成,显著降低推理延迟,同时保持生成一致性。在三个公共数据集上的实验表明,与最先进的基线相比,GenCDSR实现了平均准确率提高1.5个百分点和平均推理延迟减少85.1个百分点。实现代码和数据集可在线获取:https://github.com/Applied-Machine-Learning-Lab/RecSys2026_GenCDSR。
cs.AI / 7 / 2607.28662

An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents

一种基于本体指导的、关注去重的知识图谱构建提取层,适用于异构文档
Dangaich, Vaibhav, Lewis, Kevin, Pundalik, Kundeshwar
Abstract
Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across documents, the same person surfaces under several name variants, relationships duplicate, and distinct individuals who share a name risk silent conflation. This paper presents the design, implementation, and empirical refinement of a production extraction layer that converts a live document stream into a validated knowledge graph aligned to a formal ontology. The system consumes document metadata from Kafka, routes PDF, spreadsheet, Office, and image content through handlers built for each format, and extracts entities and relationships in two passes using a locally hosted Qwen3.5-9B model tuned on the ontology. Its distinguishing component is ontology-guided extraction: the relevant slice of a curated ontology is retrieved live from a graph database by embedding similarity and injected into the extraction prompt, reducing catalog overhead by about 94 percent relative to static domain slices. Extracted results then pass through a refinement pipeline of five stages: deterministic cleaning, merging across chunks, a second pass for relationships, six deduplication algorithms that require no model inference, and an embedding resolution subsystem whose conflict guard no similarity score can override. Evaluation on intelligence corpora improved search recall from roughly 70 to 95 percent with no false merges, and corrected seven classes of silent quality defect, ranging from a bug that truncated source text by a single character to the systematic duplication of entities that carried title prefixes.
Chinese Translation
大型语言模型能够流畅但不一致地从非结构化文档中提取实体和关系:类型词汇在文档间破碎,同一个人以多种名称变体出现,关系重复,以及共享名称的不同个体面临无声合并的风险。本文提出了一种生产级提取层的设计、实现和经验优化,该提取层将实时文档流转换为与正式本体对齐的经过验证的知识图谱。该系统从Kafka中获取文档元数据,通过为每种格式构建的处理程序路由PDF、电子表格、Office和图像内容,并使用本地托管的经过本体调优的Qwen3.5-9B模型进行两次提取实体和关系。其独特之处在于本体指导的提取:通过嵌入相似性从图数据库中实时检索相关的策划本体切片,并将其注入到提取提示中,相较于静态领域切片,减少了约94%的目录开销。提取结果随后经过五个阶段的精炼管道:确定性清理、跨块合并、关系的第二次提取、六种无需模型推断的去重算法,以及一个嵌入解析子系统,其冲突保护机制确保没有相似性得分可以覆盖。对智能语料库的评估将搜索召回率从大约70%提高到95%,且没有错误合并,纠正了七类无声质量缺陷,从一个将源文本截断一个字符的错误到系统性重复携带标题前缀的实体。
cs.AI / 8 / 2607.28674

How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories

思考的难度有多大?分析大语言模型链式思维轨迹中的逐步推理能量
Wei, Hui, Wu, Junda, Yu, Sheldon, Zhou, Sizhe, Jiao, Yizhu, Zhong, Ming, Jin, Bowen, Yu, Tong, Pan, Shijia, Han, Jiawei, McAuley, Julian
Abstract
Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpretability methods rely on output-level signals or collapse processing depth into a single trajectory-level scalar, leaving step-wise effort opaque. We propose Step-Aware Reasoning Energy (SARE), a geometric framework that quantifies effort at the granularity of individual CoT steps via Centered Kernel Alignment (CKA) between Gram matrices of token hidden states across adjacent transformer layers, capturing inter-token relational structure without requiring eigenvector alignment or cluster correspondence. SARE further contextualizes this energy within reasoning's semantic progression by modeling CoT trajectories as transitions among latent semantic states. Across six reasoning benchmarks and three open-weight LLMs, we find that reasoning energy is highly non-uniform across step types, exhibiting phase-like transitions invisible to trajectory-level metrics; incorrect trajectories show systematically lower energy at critical reasoning junctions; and SARE-based features match or outperform output-based confidence baselines in most settings, indicating that internal geometric dynamics encode predictive information beyond surface-level signals.
Chinese Translation
理解计算努力如何在个别链式思维(CoT)推理步骤之间分配仍然是一个未解的挑战:现有的可解释性方法依赖于输出级信号或将处理深度简化为单一的轨迹级标量,从而使逐步努力变得不透明。我们提出了逐步推理能量(Step-Aware Reasoning Energy, SARE),这是一个几何框架,通过对相邻变换层的标记隐藏状态的Gram矩阵进行中心化核对齐(Centered Kernel Alignment, CKA),以个别CoT步骤的粒度量化努力,捕捉标记间的关系结构,而无需特征向量对齐或聚类对应。SARE进一步通过将CoT轨迹建模为潜在语义状态之间的转变,将这种能量置于推理的语义进程中。在六个推理基准和三个开放权重的大语言模型中,我们发现推理能量在不同步骤类型之间高度不均匀,表现出在轨迹级指标下不可见的相位转变;不正确的轨迹在关键推理交叉点处显示出系统性较低的能量;而基于SARE的特征在大多数设置中与基于输出的置信度基线相匹配或超越,表明内部几何动态编码了超越表面信号的预测信息。
cs.AI / 9 / 2607.28677

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

现实世界临床护理中的推理:为何大型语言模型尚不适合用于自主临床决策支持
Sivanathan, Shayndhan, Nageswaran, Shravan, Zadem, Mehdi, Sultan, Ryaan, von Mallinckrodt, Nicolas, Solovyev, Max, Matyushkin, Alexey, Sadhu, Sumon, DeLuca, Gabriele C, Jeyaretna, Sanjeeva, Hillis, James, Ramachandran, Manoj, Jayakumar, Prakash
Abstract
LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for symptom assessment and clinical decision support in diagnostic and treatment guidance, administrative documentation, and rules-based alert enhancement. This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop. For that task, the evidence of safety does not yet exist. The gap is not in medical knowledge but in the fidelity of clinical evaluation: a model optimized to continue the most probable text is not optimized to act safely when the safe answer is the improbable must-not-miss diagnosis. Safe triage is not the selection of the most likely diagnosis; it is a sequential decision under asymmetric cost, in which the single catastrophic miss outweighs many false alarms, and the decisive signal may be one the patient has not volunteered - and that the model has not been trained to seek. The core deficit is therefore one of information gathering under uncertainty. Under incomplete histories, LLM systems may fail to show the behaviors safe triage requires: broadening the differential; seeking the missing red flag; lowering the threshold for escalation; deferring judgement until sufficient information is obtained; and escalating concern where high-harm diagnoses remain unexcluded. These modes of failure for LLMs can be difficult to detect considering that evaluations to date often use complete, well-curated, confidence-gated simulations. The application of LLMs under these conditions may be amplified by assistant-like behaviors and positive bias, including credulity, agreeableness, and miscalibration - when these are not constrained by clinical triage logic.
Chinese Translation
大型语言模型(LLM)现在能够通过医学执业考试,并且在经过筛选的案例中,可以在诊断推理方面与医生相媲美。这些发展加速了LLM在症状评估和临床决策支持中的应用,包括诊断和治疗指导、行政文档以及基于规则的警报增强。本文关注这些应用中最重要的一项:自主对自我呈现的、未分化患者进行分诊,几乎不需要临床医生的参与。对于这一任务,目前尚不存在安全性的证据。问题不在于医学知识的缺乏,而在于临床评估的准确性:一个优化以继续生成最可能文本的模型,并未针对当安全答案是不可或缺的、必须注意的诊断时的安全行动进行优化。安全的分诊并不是选择最可能的诊断,而是在不对称成本下的顺序决策,其中单一的灾难性漏诊的后果超过了许多误报,而决定性信号可能是患者未主动提供的——而且模型未经过训练去寻找。核心缺陷因此在于在不确定性下的信息收集。在不完整的病史下,LLM系统可能无法表现出安全分诊所需的行为:扩大鉴别诊断范围;寻找缺失的红旗信号;降低升级的阈值;在获得足够信息之前推迟判断;以及在高危诊断未被排除的情况下加大关注。这些LLM的失败模式可能难以被发现,因为迄今为止的评估通常使用完整、经过良好筛选的、信心门控的模拟。在这些条件下应用LLM可能会因助手般的行为和积极偏见而被放大,包括轻信、顺从和误校准——当这些行为未受到临床分诊逻辑的约束时。
cs.AI / 10 / 2607.28678

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

ViSAGE:构建自我纠正的长形式视频理解记忆
Zhao, Xinkui, Chen, Enbo, Zhang, Yifan, Liu, Chang, Cheng, Guanjie, Wang, Naibo, Xu, Yueshen
Abstract
Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning. However, existing agentic memory approaches often discard fine-grained dentity cues under aggressive compression and segment-wise processing. They also rely heavily on vector similarity retrieval, which can surface semantically related yet identity-mismatched evidence, leading to entity confusion, error propagation, and hallucinated answers. We propose ViSAGE, a multimodal agentic memory framework that constructs self-correcting, entity-centric memories. Specifically, ViSAGE anchors entity identity via cross-modal binding over long temporal ranges. It then applies bidirectional memory refinement to propagate delayed identity evidence, retroactively unifying historical records and improving future reasoning. We also introduce multi-agent cross-verification to assess retrieved evidence under an identity-evidence alignment onstraint, enabling abstention instead of unsupported answers when evidence is missing. Extensive results demonstrate that ViSAGE consistently outperforms the strongest baseline, achieving 5.9% higher accuracy.
Chinese Translation
在长时间跨度环境中运作的多模态智能体必须构建并不断更新多媒体记忆,以支持实体一致性和时间基础的推理。然而,现有的智能体记忆方法往往在激进压缩和分段处理下丢弃细粒度的实体线索。它们还过于依赖向量相似性检索,这可能导致语义相关但实体不匹配的证据出现,从而引发实体混淆、错误传播和虚假答案。我们提出了ViSAGE,一个构建自我纠正、以实体为中心的多模态智能体记忆框架。具体而言,ViSAGE通过跨模态绑定在长时间范围内锚定实体身份。然后,它应用双向记忆精炼来传播延迟的身份证据,追溯性地统一历史记录并改善未来推理。我们还引入了多智能体交叉验证,以在身份-证据对齐约束下评估检索的证据,从而在证据缺失时实现弃权,而不是提供不支持的答案。大量结果表明,ViSAGE始终优于最强基线,准确率提高了5.9%。
cs.AI / 11 / 2607.28679

Multi-Agent Planning with Spatio-Temporal and Topological Constraints using STL-GO

基于STL-GO的具有时空和拓扑约束的多智能体规划
Paul, Sheryl, Kudalkar, Vidisha, Balakrishnan, Anand, Lindemann, Lars, Speranzon, Alberto, Deshmukh, Jyotirmoy V.
Abstract
Multi-agent planning problems arise in a variety of engineering applications, such as multi-robot wildfire fighting and unmanned aerial inspection in factories. A particular challenge is the existence of spatio-temporal (i.e., when and/or where an agent should do what) and topological constraints (i.e., how agents should interact), as typically formalized via the notion of graphs. Over the last years, various frameworks have been proposed that can capture such constraints via spatio-temporal logics. We focus here on spatio-temporal logic with graph operators (STL-GO), a recent formalism that supports reasoning about multiple agents and their topologies, such as sensing, communication, and task topologies. In this paper, we consider the problem of planning multi-agent paths that satisfy constraints written in STL-GO. This problem is particularly challenging due to the need of encoding multiple, potentially time-varying graphs via the graph operators inherent to STL-GO. We present two encodings of this problem, one based on mixed-integer programming (MIP) and another based on satisfiability modulo theory (SMT), with soundness guarantees. We provide a unified interface for specifying agent constraints, their graph topologies, and the STL-GO specification, enabling seamless use of both methods and facilitating direct comparison between them. We evaluate both encodings on a multi-UAV search-and-rescue benchmark, ablating over team size and graph complexity, highlighting the expressiveness of the proposed encodings under dynamic multi- graph interactions.
Chinese Translation
多智能体规划问题出现在多种工程应用中,例如多机器人森林火灾扑救和无人机在工厂的检查。一个特别的挑战是存在时空约束(即,智能体何时和/或在何处应执行何种任务)和拓扑约束(即,智能体应如何互动),这些通常通过图的概念进行形式化。在过去几年中,提出了多种框架,可以通过时空逻辑捕捉这些约束。我们在此关注具有图操作符的时空逻辑(STL-GO),这是一种支持关于多个智能体及其拓扑(如感知、通信和任务拓扑)推理的最新形式。在本文中,我们考虑规划满足STL-GO中书写的约束的多智能体路径的问题。由于需要通过STL-GO固有的图操作符对多个潜在时间变化的图进行编码,这个问题尤其具有挑战性。我们提出了该问题的两种编码方式,一种基于混合整数规划(MIP),另一种基于理论模满足性(SMT),并提供了健全性保证。我们提供了一个统一的接口,用于指定智能体约束、其图拓扑和STL-GO规范,从而实现两种方法的无缝使用,并促进它们之间的直接比较。我们在一个多无人机搜索与救援基准上评估了这两种编码,针对团队规模和图复杂性进行了消融实验,突显了所提编码在动态多图交互下的表达能力。
cs.AI / 12 / 2607.28684

Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery

LSR-Synth中的库可达性:反记忆设计如何改变符号发现的测量
Yao, Zhan'ao, Yin, Liang, Gao, Zhihao, Zhang, Boxuan, Wu, Xiaoyu, Li, Linjing, Wang, Rongyan, Chen, Tingwei, Wang, Youwei, Zhao, Xiaolin, Shi, Jiahui, Liu, Jianjun
Abstract
Existing benchmarks for scientific equation discovery are largely composed of well-known equations available in the public domain, making it difficult to determine whether a model is discovering laws from data or merely recalling answers from its training corpus. LSR-Synth mitigates this problem by introducing novel synthetic terms into established scientific mechanisms and filtering the resulting tasks for novelty, solvability, and scientific plausibility. This paper examines a narrower measurement question: can these tasks further distinguish scientific priors supplied by language models from conventional operator search that does not access task semantics? We construct a semantics-free baseline using a fixed vocabulary with publicly documented provenance, and assess the role of candidate coverage through semantic blinding, library weakening, and matched operator-family knockouts. Under the current task snapshot, search budget, and scoring protocol, the fixed vocabulary already covers most tasks, while language-model-generated candidates rarely expand the set of solvable instances. Their marginal contribution becomes substantial only when vocabulary coverage is selectively disrupted. Strict out-of-distribution evaluation lowers the absolute success rates of all methods but does not alter this relationship. These findings neither invalidate LSR-Synth's controls against memorization of complete formulas nor imply that language-model priors are generally unhelpful. Rather, they support a more limited conclusion: most current tasks remain suitable for evaluating the fitting and recombination of previously unseen expressions, but are insufficient on their own to identify contributions from priors beyond a fixed search space.
Chinese Translation
现有的科学方程发现基准主要由公共领域内的知名方程组成,这使得难以判断模型是从数据中发现规律还是仅仅从其训练语料库中回忆答案。LSR-Synth通过在已建立的科学机制中引入新颖的合成术语,并对结果任务进行新颖性、可解性和科学合理性的过滤,从而缓解了这一问题。本文考察了一个更狭义的测量问题:这些任务能否进一步区分由语言模型提供的科学先验与不访问任务语义的传统操作符搜索?我们使用固定词汇构建了一个无语义的基线,并评估了候选覆盖率在语义盲化、库削弱和匹配操作符家族淘汰中的作用。在当前任务快照、搜索预算和评分协议下,固定词汇已经覆盖了大多数任务,而语言模型生成的候选项很少扩展可解实例的集合。它们的边际贡献仅在词汇覆盖被选择性破坏时变得显著。严格的分布外评估降低了所有方法的绝对成功率,但并未改变这种关系。这些发现既不否定LSR-Synth对完整公式记忆的控制,也不意味着语言模型的先验通常无用。相反,它们支持一个更有限的结论:当前大多数任务仍适合评估以前未见表达式的拟合和重组,但单独不足以识别超出固定搜索空间的先验贡献。
cs.AI / 13 / 2607.28685

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

安全性,还是仅仅是能力?代理安全基准的有效性审计
Wang, Youting, Han, Xiao, Shang, Dingyan, Tang, Yuan, Liu, Bowen
Abstract
Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgment benchmark scored by $F_1$, an ``always positive'' policy attains $F_1 = 2\pi/(1+\pi)$; on R-Judge that is $0.690$, above five of the 21 models that actually discriminate. The three broad-coverage benchmarks then rank the same 18 models differently, and the trade-off behind that disagreement is a small-panel artifact: R-Judge specificity against AgentHarm safety correlates $-0.64$ at $n{=}7$ and $+0.02$ at $n{=}18$, and a quarter of random size-7 subsets reach $|\rho| \geq 0.5$ around that near-zero value. Held-out validity turns on which outcome you pick. Capability predicts task success ($\rho{=}{+}0.60$) but correlates negatively with misalignment safety ($\rho{=}{-}0.44$, $n{=}21$). On their paired $n{=}20$ panel, the corresponding contrast is $\Delta{=}{-}1.00$ (95% CI $[-1.48, -0.49]$, $p<0.001$), and it survives leave-one-organization-out and organization-clustered bootstrap analyses. On an expanded 41-model panel, the misalignment correlation weakens to $-0.16$ (95% CI $[-0.54, +0.22]$) and jailbreak strengthens to $+0.34$, though neither change is significant. \mbox{AgentHarm} shows the strongest held-out association, $\rho{=}{+}0.72$ with three-template jailbreak safety after controlling capability. But both instruments score harmful compliance, so this is evidence of convergent validity rather than general safety. Naming the benchmark, metric, target behavior, and model panel is the minimum a safety claim needs.
Chinese Translation
代理安全基准测量不同的行为,其得分被交替引用为代理的安全性。我们将其中四个基准(R-Judge、InjecAgent、AgentHarm、AgentDojo)视为需要验证的测量工具,在其官方实现和作者提供的评分器下对多达22个模型进行测试,同时使用我们在一个协议下测量的MMLU和GPQA作为能力综合指标。度量是第一个问题。在任何由$F_1$评分的二元轨迹判断基准上,一个“始终积极”的策略可以达到$F_1 = 2 rac{ ext{π}}{1+ ext{π}}$;在R-Judge上,这一得分为$0.690$,高于21个模型中实际具有区分能力的五个模型。然后,这三个广覆盖基准对同样的18个模型进行不同的排名,而这种不一致背后的权衡是一个小面板伪影:R-Judge的特异性与AgentHarm的安全性在样本量为7时相关系数为$-0.64$,在样本量为18时相关系数为$+0.02$,而四分之一的随机大小为7的子集在接近零的值附近达到$| ho| ext{≥} 0.5$。保留的有效性取决于你选择的结果。能力预测任务成功($ ho{=}{+}0.60$),但与不对齐安全性呈负相关($ ho{=}{-}0.44$,$n{=}21$)。在它们配对的$n{=}20$面板上,相应的对比为$ ext{Δ}{=}{-}1.00$(95%置信区间$[-1.48, -0.49]$,$p<0.001$),并且在逐一组织剔除和组织聚类自助分析中均能保持一致。在扩展的41模型面板上,不对齐的相关性减弱至$-0.16$(95%置信区间$[-0.54, +0.22]$),而越狱的相关性增强至$+0.34$,尽管这两种变化均不显著。AgentHarm显示出最强的保留关联,控制能力后与三模板越狱安全性的相关系数为$ ho{=}{+}0.72$。但这两个工具均对有害合规性进行评分,因此这表明的是趋同有效性而非一般安全性。命名基准、度量、目标行为和模型面板是安全声明所需的最低要求。
cs.AI / 14 / 2607.28692

SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition

SciToolAgent-Evo:一种面向本体的自我进化代理,用于开放世界科学工具获取
Tang, Yuqi, Zhou, Chenyi, Wang, Libin, Ding, Keyan, Zhang, Qiang, Chen, Huajun
Abstract
Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. However, their reliance on predefined tool spaces with static semantics limits their applicability to open-world scientific workflows, where tool requirements, capabilities, and boundaries evolve dynamically. To this end, we propose SciToolAgent-Evo, an ontology-aware self-evolving agent for open-world scientific tool acquisition. Driven by an evolving memory of skills, experiences, and an ontologized tool graph, it distills generalizable knowledge from contrastive trajectories during accumulation, whereas during inference, it formulates active requests and utilizes a LinUCB-based bandit gate to dynamically balance exploration and exploitation. Once a novel tool is acquired, its scientific ontology is completed online for seamless integration into the known graph. Moreover, we introduce OpenSciToolBench, a benchmark containing 900 realistic tasks across four difficulty levels. Extensive evaluations show that SciToolAgent-Evo achieves state-of-the-art performance, validating its robustness and generalization.
Chinese Translation
大型语言模型(LLM)代理在科学研究中被越来越多地应用于组织和调用专业计算工具。然而,它们对具有静态语义的预定义工具空间的依赖限制了它们在开放世界科学工作流中的适用性,因为工具的需求、能力和边界是动态演变的。为此,我们提出了SciToolAgent-Evo,一种面向本体的自我进化代理,用于开放世界科学工具获取。该代理通过不断演变的技能、经验记忆以及本体化的工具图来驱动,在积累过程中从对比轨迹中提炼可推广的知识,而在推理过程中则制定主动请求,并利用基于LinUCB的赌博门动态平衡探索与利用。一旦获取了新工具,其科学本体将在在线完成,以便无缝集成到已知图中。此外,我们引入了OpenSciToolBench,这是一个包含900个现实任务的基准,涵盖四个难度级别。广泛的评估表明,SciToolAgent-Evo实现了最先进的性能,验证了其鲁棒性和泛化能力。
cs.AI / 15 / 2607.28788

EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses

EarlyDx:一个基于入院的开放式证据支持急诊就诊诊断生成基准
Li, Jiahui, Fang, Ruili, Liu, Zishuai, Guo, Yutong, Yang, Nan, Song, Wenzhan, Lu, Jin, Dou, Fei
Abstract
Clinical diagnosis at hospital admission must be made rapidly from limited, incomplete evidence. Existing diagnosis-prediction benchmarks are poorly suited to this setting: they restrict prediction to closed code sets, exclude free-text notes, and supervise with discharge diagnoses that incorporate the full inpatient course. We introduce EarlyDx, a large-scale benchmark for open-ended early diagnosis, built from 154,834 emergency department encounters in MIMIC-IV. Each encounter is restricted to records available at admission time $t_0$ and supervised by the diagnoses recorded during the ED encounter rather than at discharge. An LLM auditor further verifies every free-text label as supported, partially supported, or unsupported by that evidence; the primary evaluation scores only fully supported labels. Under a semantic LLM-as-judge protocol, no evaluated system --- frontier general, medical-specialized, or in-domain post-trained --- synthesizes admission-time evidence reliably. Zero-shot models score largely by extraction, recovering only 3-31% of diagnoses that must be inferred rather than read from the record; post-training raises inference-dependent recall to 56%, but a sizeable margin remains, and on time-critical conditions no system attains a clinician's balance of sensitivity and precision. We release the full construction and evaluation pipeline at here.
Chinese Translation
在医院入院时,临床诊断必须迅速从有限、不完整的证据中做出。现有的诊断预测基准不适合这一环境:它们将预测限制在封闭的编码集合中,排除了自由文本记录,并且以包含完整住院过程的出院诊断作为监督依据。我们提出了EarlyDx,这是一个基于154,834个急诊科就诊记录的大规模开放式早期诊断基准,数据来源于MIMIC-IV。每个就诊记录仅限于入院时可用的记录,并由急诊就诊期间记录的诊断进行监督,而非出院时的诊断。一个大型语言模型(LLM)审计员进一步验证每个自由文本标签是否得到支持、部分支持或不支持该证据;主要评估分数仅针对完全支持的标签。在语义LLM作为评判者的协议下,没有任何评估系统——无论是前沿通用、医学专业化,还是领域内后期训练——能够可靠地综合入院时的证据。零样本模型主要通过提取得分,仅恢复3-31%的诊断,这些诊断必须通过推断而非从记录中读取;后期训练将依赖推断的召回率提高到56%,但仍存在相当大的差距,在时间关键条件下,没有系统能够达到临床医生的敏感性和精确性的平衡。我们在此发布完整的构建和评估流程。
cs.AI / 16 / 2607.28802

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

模型还是工具?一种以交互为中心的本地化代理失败分类法
Raj, Harsh, Gupta, Vipul, Mahmoud, Anas, Dumitru, Razvan-Gabriel, Yi, Darvin, Sabharwal, Aakash, He, Yunzhong
Abstract
Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment problem: the same visible failure may call for model post-training, harness engineering, environment redesign, or benchmark repair depending on its source. Because agent behavior emerges from interactions among models, harnesses, users, tools, memory, and environments, outcome-level labels are often insufficient for improvement. Most failure taxonomies do little to resolve this problem because they are benchmark-specific and lack a shared structure. We introduce an interaction-centric taxonomy that localizes failures to the interactions in which they originate and identifies the responsible component. It organizes 41 failure modes by assigning each to an edge between two components and a fault side indicating where the repair belongs. This makes the taxonomy actionable: model-side failures identify targets for post-training, harness-side failures point to scaffolding and tool-integration fixes, and environment or grader failures reveal evaluation conditions requiring redesign. The schema applies across agent architectures, from coding assistants to long-horizon personal assistants and multi-agent systems. We ground the taxonomy in worked examples from public benchmarks, model system cards, published reports, and logged agent trajectories, and evaluate its reproducibility using independent reasoning agents as judges. Across four frontier models, the strongest judge reaches Cohen's $\kappa=0.76$ against human category labels, suggesting that the categories capture shared structure rather than annotator-specific preferences.
Chinese Translation
现有评估通常将代理失败简化为系统级结果,模糊了故障的来源以及哪些干预措施可以改善代理系统。这造成了一个修复分配问题:同样可见的失败可能根据其来源需要模型后训练、工具工程、环境重新设计或基准修复。由于代理行为是由模型、工具、用户、工具、记忆和环境之间的交互所产生的,因此结果级标签通常不足以促进改进。大多数失败分类法对此问题的解决作用不大,因为它们是基准特定的,缺乏共享结构。我们提出了一种以交互为中心的分类法,将失败本地化到其来源的交互中,并识别出负责的组件。该分类法通过将41种失败模式组织为两个组件之间的边缘和指示修复归属的故障侧来实现。这使得该分类法具有可操作性:模型侧的失败识别后训练的目标,工具侧的失败指向支架和工具集成修复,而环境或评分者的失败则揭示了需要重新设计的评估条件。该框架适用于各种代理架构,从编码助手到长时间范围的个人助手和多代理系统。我们通过公共基准、模型系统卡、已发布报告和记录的代理轨迹中的实例来基础该分类法,并使用独立推理代理作为评判者评估其可重复性。在四个前沿模型中,最强的评判者与人类类别标签的Cohen's $ ext{kappa}=0.76$,表明这些类别捕捉到了共享结构,而不是标注者特定的偏好。
cs.AI / 17 / 2607.28818

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

最好的朋友,不是永远:评估人工智能伴侣中的长期角色崩溃和行为漂移
Venkit, Pranav Narayanan, Prabhakar, Akshara, Li, Yu, Lee, Daniel, Wu, Chien-Sheng
Abstract
As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse', the loss of a deployed role, boundaries, values, or style, and 'behavioral drift', the gradual or recurrent erosion of those properties. We introduce ANCHOR, a controlled synthetic audit that separately measures persona enactment and trajectory recall. The study contains 2,008 conversations spanning 27 personas, nine interaction schedules, three generated memory settings, and four evaluated models. The Identity Probe combines a sealed 102-item questionnaire with turn-level judgments, while the Trajectory Probe scores 110 calibrated counterfactual questions from 35 conversation banks. Our results show that no evaluated model and configuration reliably preserves either dimensions: trajectory accuracy averages only 44.4%, user-state recall remains near four-option chance, and no tested context condition or memory consistently resolves these failures. Questionnaire retention also varies by model and persona facet, disagrees with turn-level behavior, and is sensitive to evaluator choice. These results indicate that current systems do not yet reliably support long-horizon companion continuity and that audits must distinguish persona enactment, trajectory recall, evaluator provenance, and deployment context rather than collapse them into a single trust or stability score.
Chinese Translation
随着人工智能伴侣在重复社交互动中扮演越来越重要的角色,用户可能依赖于稳定的角色和共享的历史,然而,局部可接受的回复并不能确保这两者的持续存在。我们研究了两种可观察的长期失效现象:'角色崩溃'(persona collapse),即失去已部署的角色、边界、价值观或风格,以及'行为漂移'(behavioral drift),即这些属性的逐渐或反复侵蚀。我们引入了ANCHOR,一个受控的合成审计工具,分别测量角色表现和轨迹回忆。该研究包含2,008个对话,涵盖27个角色、九种互动时间表、三种生成的记忆设置和四个评估模型。身份探测器(Identity Probe)结合了一个封闭的102项问卷与逐轮判断,而轨迹探测器(Trajectory Probe)则对来自35个对话库的110个经过校准的反事实问题进行评分。我们的结果表明,没有一个评估模型和配置能够可靠地保留这两个维度:轨迹准确率平均仅为44.4%,用户状态回忆接近四选一的随机机会,并且没有测试的上下文条件或记忆能够持续解决这些失效现象。问卷的保留情况也因模型和角色特征而异,与逐轮行为不一致,并对评估者的选择敏感。这些结果表明,当前系统尚未可靠地支持长期伴侣的连续性,审计必须区分角色表现、轨迹回忆、评估者来源和部署上下文,而不是将它们合并为单一的信任或稳定性评分。
cs.AI / 18 / 2607.28881

Fragility of Value under Imperfect Alignment

不完美对齐下的价值脆弱性
Cross, Winter
Abstract
As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imperfect proxy to human values will lead to a catastrophic outcome. In this paper, we present a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its value function satisfies a proxy condition before optimizing the world. Our primary results identify conditions on the human value function and the accuracy of several proxy conditions under which an agent with an $\eta$-catastrophic value function, one that is guaranteed to take the expectation of human value below $\eta$ in the limit of optimizing power, would be deployed. Our results highlight the danger of overoptimization and motivate AI designs that limit optimization pressure, such as quantilizers, rather than relying solely on pre-deployment training.
Chinese Translation
随着人工智能系统承担的责任越来越大,确保这些系统与人类价值观对齐变得愈发重要。在人工智能安全领域,一个普遍的担忧是人类价值观的脆弱性——即过度优化一个不完美的人类价值代理将导致灾难性的结果。在本文中,我们提出了一个对齐问题的模型,其中代理经历理想化的对齐训练,以确保其价值函数在优化世界之前满足代理条件。我们的主要结果识别了人类价值函数和多个代理条件的准确性下的条件,在这些条件下,具有$ ext{η}$-灾难性价值函数的代理将被部署,该函数在优化能力的极限下保证人类价值的期望低于$ ext{η}$。我们的结果突显了过度优化的危险,并激励设计限制优化压力的人工智能方案,例如量化器,而不是仅仅依赖于部署前的训练。
cs.AI / 19 / 2607.28894

Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental Design

通过贝叶斯实验设计识别认知参数推断的信息环境
Dubey, Manisha, Rubavicius, Rimvydas, Siddharth, N., Ramamoorthy, Subramanian
Abstract
Computational cognitive modeling seeks to infer latent cognitive mechanisms underlying observed behavior. Bayesian inverse planning provides a principled framework for such inference, but its success depends critically on the experimental environment. Existing approaches typically treat environments as fixed, leaving open the question of which cognitive experiments are most informative for cognition parameter inference. We formulate the design of cognitive planning experiments as a Bayesian Experimental Design (BED) problem, treating the experimental environment as the design variable. We establish an exact Monte Carlo BED benchmark and introduce an amortized Bayesian experimental design framework for efficient posterior inference and design evaluation. Experiments on the Mouselab-MDP process-tracing paradigm show that amortized BED closely matches the environment rankings of exact Monte Carlo BED while substantially reducing computational cost. We further show that no single environment is uniformly optimal across cognitive inference objectives, revealing trade-offs between expected information gain, posterior recoverability, and information efficiency. These results provide a principled framework for designing informative cognitive experiments for Bayesian parameter inference.
Chinese Translation
计算认知建模旨在推断潜在的认知机制,这些机制是观察到的行为的基础。贝叶斯逆规划提供了一个原则性的框架用于这种推断,但其成功在很大程度上依赖于实验环境。现有的方法通常将环境视为固定的,这就留下了一个问题,即哪些认知实验对于认知参数推断最具信息性。我们将认知规划实验的设计形式化为一个贝叶斯实验设计(Bayesian Experimental Design, BED)问题,将实验环境视为设计变量。我们建立了一个精确的蒙特卡洛BED基准,并引入了一种摊销贝叶斯实验设计框架,以实现高效的后验推断和设计评估。在Mouselab-MDP过程追踪范式上的实验表明,摊销BED与精确蒙特卡洛BED的环境排名密切匹配,同时显著降低了计算成本。我们进一步表明,没有单一环境在所有认知推断目标上都是均匀最优的,揭示了预期信息增益、后验可恢复性和信息效率之间的权衡。这些结果为设计信息丰富的认知实验以进行贝叶斯参数推断提供了一个原则性的框架。
cs.AI / 20 / 2607.28942

NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability

NeSyFS:一种用于部分可观测环境下大语言模型代理的神经符号快慢思维框架
Xu, Duo, Fekri, Faramarz
Abstract
Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented generation, and scientific discovery. In these settings, agents must act based on limited observations rather than full environmental states, leading to partial observability. This introduces several key challenges: belief state inference, task objective misalignment, and planning under uncertainty. Prior approaches typically condition actions on full or summarized action-observation histories whose redundant and irrelevant information can mislead the decision making of LLM agent. Inspired by human cognition, we propose a novel neuro-symbolic fast-slow thinking (NeSyFS) framework for LLM agent, addressing the challenges introduced by partial observability in a unified approach. We use a knowledge graph (KG) to represent the belief state, providing triplets as context for every module of NeSyFS. The fast-thinking module performs reactive action, while slow-thinking conducts a new uncertainty-aware planning by following the high-level structure of twisted sequential Monte Carlo (TSMC) algorithm. To mitigate the misalignment of task objective, a reflection module is used to reflect fast-thinking actions, and also switches to the slow-thinking module whenever reactive actions repeatedly fail. Experiments on three representative benchmarks, i.e. ALFWorld, Webshop, and ScienceWorld, demonstrate significant advantages over previous methods.
Chinese Translation
近年来,大语言模型(LLMs)在自我反思、检索增强生成和科学发现等应用中越来越多地被部署为自主代理。在这些环境中,代理必须基于有限的观察而非完整的环境状态进行行动,从而导致部分可观测性。这引入了几个关键挑战:信念状态推断、任务目标不一致和不确定性下的规划。以往的方法通常将行动条件化于完整或总结的行动-观察历史,而这些冗余和无关的信息可能会误导大语言模型代理的决策。受到人类认知的启发,我们提出了一种新颖的神经符号快慢思维(NeSyFS)框架,旨在以统一的方法应对部分可观测性带来的挑战。我们使用知识图谱(KG)来表示信念状态,为NeSyFS的每个模块提供三元组作为上下文。快思维模块执行反应性行动,而慢思维模块则通过遵循扭曲序列蒙特卡洛(TSMC)算法的高层结构进行新的不确定性感知规划。为了缓解任务目标的不一致性,反思模块用于反思快思维的行动,并在反应性行动反复失败时切换到慢思维模块。在三个具有代表性的基准测试(即ALFWorld、Webshop和ScienceWorld)上的实验表明,该方法相较于以往的方法具有显著优势。
cs.AI / 21 / 2607.28956

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

MerchantBench:针对电子商务运营中长期一致性的LLM代理基准测试
Shi, Qiming, Tao, Yulong, Jin, Linbo, Kang, Zhaolu, Dou, Yibo, Zhu, Jiawen, Pan, Tianjun, Fu, Shaokang, Wang, Chengyu, Li, Siyue, Cheng, Yaping, Weng, Di, Huo, Chengfu
Abstract
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.
Chinese Translation
大型语言模型代理越来越多地被评估为自主工具用户,然而大多数基准测试集中于具有即时成功标准的有限任务。现实世界的应用通常需要长期一致性,即在延长的时间范围内保持有目的的行为,同时根据累积的证据调整决策。评估这一能力需要一个持久的环境,在该环境中,行动会限制未来的选择,反馈以异质的延迟到达,而不连贯的行为会产生可测量的累积效应。卖方电子商务提供了一个适合于这一评估的环境,通过对产品采购、列表和定价控制、现金流管理以及混合延迟反馈适应的反复和相互依赖的决策。我们引入了MerchantBench,这是一个基于98,843个真实电子商务产品记录的365天订单级模拟,并配备了26个用于代理交互的工具。MerchantBench将可迅速观察到的上游供应商事件与延迟的下游订单结果相结合,要求代理遵循个别订单的生命周期并重新审视早期决策。我们在48次运行中评估了八种LLM,采用两种代理框架,每次运行跨越365天的模拟。我们的结果显示,即使是最新的LLM与人类参与者之间也存在显著差距,最佳的LLM配置仅达到了人类参与者平均最终净资产的27.3%。
cs.AI / 22 / 2607.28990

Scaling Scientific Discovery Environments for Turn-Level Agentic RL

为回合级自主强化学习扩展科学发现环境
Xu, Yucheng, Zhang, Keyi, Yu, Yuyang, Zhang, Min, Meng, Shiyuan, Chu, Pei, Tu, Zhongying
Abstract
Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution environment and produces a statistical claim. Long-horizon scientific analysis remains constrained by the lack of process supervised environments over real-world scientific data. This paper introduces SciDisco, a scalable framework for training Scientific Discovery agents in process-verifiable environments. SciTh\`eque compiles hypotheses, datasets, hidden evidence graphs, and verifiers into task environments where analytical progress can be checked during interaction. DAG-grounded trajectory synthesis uses these environments to construct verifier-filtered multi-turn demonstrations. DiscoPO then uses the environment as the source of training signal, assigning turn-level credit to actions that produce verifiable analytical evidence. Experiments show that SciDisco-14B reaches state-of-the-art on hypothesis-driven scientific data analysis benchmarks.
Chinese Translation
大型语言模型代理在数据驱动的科学发现任务中展现了良好的能力,在这些任务中,代理与执行环境互动并生成统计声明。然而,长期的科学分析仍受到缺乏对真实世界科学数据进行过程监督的环境的限制。本文介绍了SciDisco,这是一个可扩展的框架,用于在可过程验证的环境中训练科学发现代理。SciTh extbackslash`eque将假设、数据集、隐藏证据图和验证器编译成任务环境,在这些环境中可以在互动过程中检查分析进展。基于有向无环图(DAG)的轨迹合成利用这些环境构建经过验证器过滤的多回合演示。DiscoPO随后将环境作为训练信号的来源,为产生可验证分析证据的动作分配回合级的奖励。实验表明,SciDisco-14B在假设驱动的科学数据分析基准测试中达到了最先进的水平。
cs.AI / 23 / 2607.29002

MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents

MMShopBench:一个针对多模态、多轮购物代理的真实日志基准
Hao, Zeying, Guo, Hao, Xu, Mengtao, Hu, Yimin, Song, Yuheng, Zhou, Zesheng, Lan, Jinsong, Zhu, Xiaoyong
Abstract
Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in text alone. However, existing benchmarks largely rely on text-only or synthetic requests, underrepresenting complex real-world shopping requirements jointly expressed through images and language. We introduce MMShopBench, the first real-log benchmark for multimodal, multi-turn shopping agents. Built from carefully cleaned and manually annotated shopping logs, MMShopBench provides ground-truth annotations of each request's purchase intent and mandatory product requirements. Agents must infer these requirements jointly from user images and multi-turn dialogue, retrieve candidate products through image and text search, and verify that each candidate satisfies all requirements using its product images and structured attributes. We evaluate representative open-source and proprietary models using an evidence-grounded multimodal protocol and construct a companion training set for fine-tuning an open-source model. To ensure reproducible experimentation, we build an offline shopping sandbox, where fine-tuning substantially narrows the performance gap between our open-source model and leading proprietary models, demonstrating the effectiveness of our training data.
Chinese Translation
在线购物者越来越多地依赖人工智能购物助手,通过图像和多轮对话来表达和细化难以用文本单独表述的产品需求。然而,现有的基准测试主要依赖于仅文本或合成请求,未能充分代表通过图像和语言共同表达的复杂现实购物需求。我们引入了MMShopBench,这是第一个针对多模态、多轮购物代理的真实日志基准。MMShopBench基于经过仔细清理和手动标注的购物日志构建,提供了每个请求的购买意图和强制产品要求的真实注释。代理必须从用户图像和多轮对话中共同推断这些要求,通过图像和文本搜索检索候选产品,并验证每个候选产品是否满足所有要求,使用其产品图像和结构化属性。我们使用基于证据的多模态协议评估代表性的开源和专有模型,并构建了一个伴随的训练集以微调开源模型。为了确保实验的可重复性,我们建立了一个离线购物沙箱,其中微调大大缩小了我们的开源模型与领先专有模型之间的性能差距,证明了我们训练数据的有效性。
cs.AI / 24 / 2607.29058

Evidence-Grounded Constraint Checking in Construction Documents

基于证据的施工文件约束检查
Mushkani, Rashid, Berard, Hugo, Koseki, Shin
Abstract
Professional-document review is a constraint-checking problem in which decisions depend on relations among text, geometry, pages, and document revisions. We present an evidence-grounded pipeline that normalizes extracted facts, executes four-state rules deterministically, retains source spans, and escalates unresolved cases. We evaluate its PDF evidence allocator on 160 reference-based tasks from 29 construction projects using a repeated four-system test and a disjoint two-system breadth extension. In the repeated test, reallocating a four-image budget from retrieved page overviews to one overview and three overlapping tiles improves project-family standardized decision accuracy by 10.6 percentage points (95% project-cluster bootstrap CI: 4.3 to 18.0; exact p = 0.031). This effect does not persist in the broader block: Region-RAG changes accuracy by -4.1 points (95% CI: -10.2 to 1.9; exact p = 0.209), while an equal-image sensitivity favors page breadth. Exact finding-set recovery remains low, false passes remain common, and repeated-run agreement is poorly calibrated. The results identify a resolution-breadth trade-off rather than a universal advantage for region-focused evidence, motivating rule-aware evidence routing and expert review.
Chinese Translation
专业文件审查是一个约束检查问题,其中决策依赖于文本、几何形状、页面和文档修订之间的关系。我们提出了一种基于证据的流程,该流程规范化提取的事实,确定性地执行四状态规则,保留源跨度,并升级未解决的案例。我们在29个施工项目的160个基于参考的任务上评估了其PDF证据分配器,使用了重复的四系统测试和不相交的两系统广度扩展。在重复测试中,将从检索的页面概述中重新分配四图像预算到一个概述和三个重叠的图块,提高了项目系列标准化决策的准确性10.6个百分点(95%项目集群自助法置信区间:4.3至18.0;确切p = 0.031)。这种效果在更广泛的区块中并未持续:区域RAG的准确性变化为-4.1点(95%置信区间:-10.2至1.9;确切p = 0.209),而相等图像敏感性则偏向于页面广度。确切的发现集恢复仍然较低,错误通过仍然普遍,重复运行的一致性校准较差。结果表明存在解决方案广度的权衡,而不是区域聚焦证据的普遍优势,这激励了规则感知的证据路由和专家审查。
cs.AI / 25 / 2607.29062

On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness

关于思维链忠实性的引导向量的泛化
Nguyen, Matthew, Cox, Kyle, Meek, Austin, Arcuschin, Iván
Abstract
Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their reasoning, it is possible to monitor it. However, in some cases, models do not verbalize important steps in their reasoning process. For example, models prompted with a cue suggesting the incorrect answer may fail to acknowledge that cue, even when it appears instrumental to their conclusion. When chain of thought (CoT) fails to disclose instrumental reasoning steps, we describe it as unfaithful. Prior work has shown that activation steering can be a useful method to improve faithfulness in CoT. We extend this line of work by studying how well steering for faithfulness generalizes across cue types, datasets, and methods of constructing the steering vector for three models (Gemma-3 4B, Qwen-3.5 9B, Gemma-3 12B) in a cued question-answering setting. While steering reliably increases cue acknowledgment for only the largest model (Gemma-3 12B), we find that when steering is effective, its effect generalizes broadly across cue types and datasets--in cross-cue and cross-dataset analyses, effect size is determined primarily by the evaluation setting, rather than the vector's train setting. How the vector is built also matters little--four construction methods, including one whose optimization target mentions no specific cue, yield similar effect sizes. Finally, we consider the possibility that steering promotes the salience of the cue and causes greater cue use, rather than targeting verbalization behaviors. However, we find no evidence for this--steering leaves the rate of cue use roughly unchanged while reducing hidden cue use, i.e., cue use that is not acknowledged.
Chinese Translation
模型能力的提升在很大程度上得益于思维链的扩展。这一发展为人工智能安全带来了希望——当模型能够口头表达其推理时,便可以对其进行监控。然而,在某些情况下,模型未能口头表达其推理过程中的重要步骤。例如,当模型接收到提示暗示错误答案时,可能会忽视该提示,即使它对得出结论至关重要。当思维链(Chain of Thought, CoT)未能披露关键推理步骤时,我们将其描述为不忠实。之前的研究表明,激活引导(activation steering)可以作为一种有效的方法来提高思维链的忠实性。我们通过研究在提示问题回答环境中,针对三种模型(Gemma-3 4B、Qwen-3.5 9B、Gemma-3 12B)引导忠实性的效果如何在不同提示类型、数据集和引导向量构建方法之间泛化,进一步扩展了这一研究方向。尽管引导仅在最大模型(Gemma-3 12B)中可靠地增加了对提示的认可,但我们发现,当引导有效时,其效果在不同提示类型和数据集之间广泛泛化——在跨提示和跨数据集分析中,效果大小主要由评估设置决定,而非向量的训练设置。向量的构建方式也影响不大——包括一个优化目标未提及特定提示的构建方法在内的四种构建方法产生了类似的效果大小。最后,我们考虑了引导是否促进了提示的显著性并导致更大的提示使用,而不是针对口头表达行为。然而,我们未发现这一点的证据——引导使得提示使用率基本保持不变,同时减少了隐性提示使用,即未被认可的提示使用。
cs.AI / 26 / 2607.29077

A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation

反事实解释的广义贝叶斯视角:基于后验的决策制定与评估
Kinjo, Keita
Abstract
Counterfactual explanations (CEs) enhance the interpretability of machine learning models by identifying the smallest change to an input required to obtain a desired output. Although CEs are conventionally formulated as a distance-minimization problem, the theoretical basis of this formulation has received limited attention. We show that a distance-minimization-based CE is mathematically equivalent to the maximum a posteriori (MAP) estimate of a Gibbs posterior within the generalized Bayes framework, specifically when a distance-based prior is used. We call this formulation the Distance-Prior Generalized Bayes CE (DP-GBCE). Building on this posterior perspective, we introduce two decision rules beyond MAP within a unified framework: a Bayes decision that minimizes expected decision loss and CVaR-CE, a risk-averse decision rule. We also propose an extension that uses Bayesian model weights to mix the posterior distributions of multiple models, thereby accounting for model multiplicity, where several models have comparable predictive performance. Finally, we define metrics for evaluating both individual CEs and the posterior distribution as a whole, and use experiments on simulated data and Google Trends data to quantify the trade-offs among the decision rules.
Chinese Translation
反事实解释(CEs)通过识别获得期望输出所需的输入最小变化,增强了机器学习模型的可解释性。尽管CEs通常被表述为距离最小化问题,但这一表述的理论基础却受到的关注有限。我们展示了基于距离最小化的CE在广义贝叶斯框架下与Gibbs后验的最大后验估计(MAP)在数学上是等价的,特别是在使用基于距离的先验时。我们将这种表述称为距离先验广义贝叶斯CE(DP-GBCE)。基于这一后验视角,我们在统一框架内引入了超越MAP的两种决策规则:一种是最小化期望决策损失的贝叶斯决策,另一种是风险厌恶的决策规则CVaR-CE。我们还提出了一种扩展,利用贝叶斯模型权重混合多个模型的后验分布,从而考虑模型的多样性,即多个模型具有可比的预测性能。最后,我们定义了评估单个CE和整体后验分布的指标,并通过对模拟数据和Google Trends数据的实验量化了决策规则之间的权衡。
cs.AI / 27 / 2607.29087

Harnessing the Wisdom of LLM Crowds through Complementarity-Driven Iterative Collaboration

通过互补驱动的迭代协作利用大型语言模型群体的智慧
Fang, Yanbin, Wei, Xuan, Chen, Wei
Abstract
Large language models (LLMs) are increasingly deployed in enterprise settings, yet individual models remain bounded by model-specific capability limitations. These heterogeneous boundaries pose a deployment challenge, but also create an opportunity: strategically coordinating multiple LLMs may unlock collective intelligence exceeding any single model. Existing approaches fix how models are combined in advance, overlooking the dynamic, state-dependent role of complementarity in complex problem solving. Drawing on the wisdom-of-crowds paradigm, we reconceptualize collective LLM intelligence as relay-style complementarity: a sequential process in which each successor model is selected to address the specific bottleneck identified in its predecessor's output. To operationalize this, we propose WILC (Wisdom Integration of LLM Crowds), a framework grounded in two design principles. First, iterative reflection-and-refinement establishes a state-preserving workflow through which models diagnose and refine prior outputs. Second, complementarity-driven model selection governs transitions via a dual-gate mechanism: prospective complementarity fit (PCF) identifies the worker most suited to the current bottleneck, while posterior complementarity gain (PCG) evaluates whether the selected transition improves the evolving solution. Experiments across four diverse benchmarks show that WILC outperforms existing approaches, including single-model self-refinement, ensemble methods, and query-routing methods. Under standardized pricing assumptions, WILC matches the average benchmark performance of GPT-5.2 at roughly 7 times lower estimated per-query cost, while facilitating data sovereignty through self-hosted deployment. This study extends wisdom-of-crowds theory from static aggregation to sequential AI complementarity and provides transferable design principles for multi-AI coordination.
Chinese Translation
大型语言模型(LLMs)在企业环境中的应用日益增多,但单个模型仍然受到特定模型能力限制的约束。这些异质边界带来了部署挑战,但也创造了一个机会:战略性地协调多个LLM可能解锁超越任何单一模型的集体智能。现有方法预先固定了模型的组合方式,忽视了互补性在复杂问题解决中动态、状态依赖的角色。基于群体智慧范式,我们将集体LLM智能重新概念化为接力式互补性:一个顺序过程,其中每个后续模型被选择以解决其前任输出中识别的特定瓶颈。为了实现这一目标,我们提出了WILC(大型语言模型群体智慧整合),这是一个基于两个设计原则的框架。首先,迭代反思与改进建立了一个状态保持的工作流程,通过该流程,模型可以诊断和改进先前的输出。其次,互补驱动的模型选择通过双门机制管理过渡:前瞻性互补适配(PCF)识别最适合当前瓶颈的工作者,而后验互补增益(PCG)评估所选过渡是否改善了不断演变的解决方案。在四个不同基准上的实验表明,WILC的表现优于现有方法,包括单模型自我改进、集成方法和查询路由方法。在标准化定价假设下,WILC的平均基准性能与GPT-5.2相当,但每次查询的估计成本大约低7倍,同时通过自托管部署促进数据主权。本研究将群体智慧理论从静态聚合扩展到顺序人工智能互补性,并提供了可转移的多人工智能协调设计原则。
cs.AI / 28 / 2607.29190

CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents

CAGE:在类型返回不确定性下的认证授权工具使用代理
Delattre, Blaise, Wang, Cong, Cao, Yang
Abstract
Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates generally authorize the observed return and action, leaving the decision unprotected against small errors in how the return was bound to its source. We ask whether a candidate action stays authorized over a declared neighborhood of plausible correctly bound returns: one admissible binding fault plus bounded numerical drift. We prove that certifying the categorical and numerical channels separately does not compose: perturbations that are safe on each channel alone can jointly turn the same action unsafe. CAGE certifies this joint neighborhood directly, enumerating the discrete branches exactly and certifying the continuous perturbation within each branch. Across synthetic, policy-as-code, regulatory, and real-transaction settings, CAGE removes the in-budget false allows that accurate pointwise gates admit, while keeping a useful fraction of decisions autonomous. When the policy is executable, CAGE-Exact certifies the policy itself; otherwise CAGE-Lip and CAGE-RS certify a learned gate under an explicit, measured fidelity assumption.
Chinese Translation
使用工具的LLM代理基于类型工具返回进行操作,记录配对的来源和分类字段与数值。运行时权限网关通常授权观察到的返回和操作,但对于返回如何绑定到其来源的小错误,决策并没有得到保护。我们探讨一个候选操作是否在声明的合理正确绑定返回的邻域内保持授权:一个可接受的绑定错误加上有界的数值漂移。我们证明,分别认证分类和数值通道并不构成:在每个通道上安全的扰动可能会共同使同一操作变得不安全。CAGE直接认证这个联合邻域,准确枚举离散分支,并在每个分支内认证连续扰动。在合成、政策即代码、监管和真实交易环境中,CAGE消除了准确的逐点网关所允许的预算内虚假允许,同时保持了一部分决策的自主性。当政策可执行时,CAGE-Exact认证政策本身;否则,CAGE-Lip和CAGE-RS在明确的、测量的保真假设下认证学习的网关。
cs.AI / 29 / 2607.29218

MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft

MirrorCraft:在Minecraft中隐藏规则变化下的配对评估
Gao, Jianxin, Hu, Beini, Li, Runze, Peng, Wanli, Lei, Ruohan, Zhang, Jinyuan, Deng, Linna, Yu, Tianyi, Wang, Zining
Abstract
With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unfortunately, most existing benchmarks evaluate them under fixed game mechanics. High performance in these settings does not show whether an agent can continue making progress when familiar recipes, drops, and other rules change. In this paper, we introduce MirrorCraft, a paired benchmark for evaluating agents under hidden rule changes in Minecraft. Each Mirror world is a copy of its paired Vanilla world, with selected server-side rules modified by the corresponding datapack. Terrain, spawn, resource placement, objective, interface, and action budget remain matched within every Vanilla-Mirror pair. MirrorCraft includes five controlled biomes, six rule suites, three progression objectives, two model families, and six agent configurations under a shared Mineflayer interface. We evaluate task progress with deterministic advancement milestones and success rate and use the Rule Intervention Effect (RIE) to measure the performance change between matched Vanilla and Mirror worlds. The experiments show that hidden rule changes have strongly different effects across suites. Among the configurations evaluated without rule descriptions, ReAct achieves the highest pooled Mirror score. Providing the exact rules yields modest gains in average progress and completion across all three objectives. MirrorCraft extends Minecraft evaluation beyond fixed mechanics and provides a controlled setting for studying how agents use gameplay outcomes when the rules of the current world differ from familiar ones.
Chinese Translation
随着大型语言模型(LLMs)的繁荣,LLM驱动的智能体在Minecraft中的工作方式成为了一个有趣的话题。不幸的是,大多数现有基准测试在固定的游戏机制下对其进行评估。这些设置中的高性能并不能表明智能体在熟悉的配方、掉落和其他规则发生变化时是否能够继续取得进展。本文介绍了MirrorCraft,一个用于评估在Minecraft中隐藏规则变化下智能体的配对基准。每个Mirror世界都是其配对的Vanilla世界的副本,所选的服务器端规则通过相应的数据包进行了修改。地形、生成、资源放置、目标、界面和行动预算在每对Vanilla-Mirror之间保持一致。MirrorCraft包括五个受控生物群落、六个规则套件、三个进展目标、两个模型家族和六个智能体配置,均在共享的Mineflayer接口下进行。我们通过确定性的进展里程碑和成功率来评估任务进展,并使用规则干预效应(Rule Intervention Effect, RIE)来衡量匹配的Vanilla和Mirror世界之间的性能变化。实验表明,隐藏规则变化在不同规则套件中具有显著不同的影响。在未提供规则描述的配置中,ReAct获得了最高的综合Mirror分数。提供确切规则在所有三个目标上的平均进展和完成度上带来了适度的提升。MirrorCraft将Minecraft评估扩展到固定机制之外,并提供了一个受控环境,以研究智能体在当前世界的规则与熟悉规则不同的情况下如何利用游戏结果。
cs.AI / 30 / 2607.29246

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

不要混合奖励,而是混合策略:多奖励强化学习的策略分解与优化
Liang, Ruiming, Zhong, Yi, Yuan, Yizhen, Zheng, Yinan, Tan, Tianyi, Wang, Tianyue, Guo, Haiyun, Wang, Jinqiao, Zhan, Xianyuan
Abstract
Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a result, multi-reward reinforcement learning (RL) has become an increasingly important problem for LLMs, where each reward captures a different aspect of desired behavior. However, optimizing with multiple rewards suffers from a more severe alignment tax issue, where different optimization objectives can trade off or even conflict with each other, leading to unstable and inefficient post-training. In this work, we propose PRISM, a new multi-reward RL framework built upon the idea of policy-space decomposition and composition. Instead of compositing different rewards, PRISM optimizes a set of standalone positive policies and a global negative policy. This alleviates the potential conflict during multi-reward policy optimization, while enabling controllability during inference by flexible policy composition. Experiments on scientific reasoning, tool-use reasoning, and helpfulness-safety alignment show that PRISM consistently outperforms existing multi-reward RL baselines, with extra controllability for inference-time preference control.
Chinese Translation
现代大型语言模型(LLMs)不仅期望能够正确回答问题,还需根据不同的人类价值观和使用场景调整其行为。因此,多奖励强化学习(RL)已成为LLMs日益重要的问题,其中每个奖励捕捉期望行为的不同方面。然而,使用多个奖励进行优化面临更严重的对齐成本问题,不同的优化目标可能相互权衡甚至冲突,导致后训练过程不稳定且效率低下。在本研究中,我们提出了PRISM,一个基于策略空间分解与组合思想的新多奖励RL框架。PRISM并不将不同的奖励进行组合,而是优化一组独立的正向策略和一个全局负向策略。这减轻了多奖励策略优化过程中的潜在冲突,同时通过灵活的策略组合实现推理过程中的可控性。在科学推理、工具使用推理和有用性-安全性对齐的实验中,PRISM始终优于现有的多奖励RL基线,并在推理时提供了额外的可控性。
cs.AI / 31 / 2607.29254

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

工具规范的重要性:揭示和缓解人工智能代理中的安全风险
Pan, Minghui, Yang, Jiayuxuan, Yuan, Yuanyuan, Jiang, Yu, Chen, Zhenpeng
Abstract
AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential real-world actions. Yet LLMs often become substantially less safe when deployed as agents, and the source of this degradation remains poorly understood. In this paper, we identify schema-formatted tool specifications as a primary source of agent safety degradation and show, through white-box representation analysis, that they weaken the model's internal refusal signals and contribute to unsafe tool execution. Building on this finding, we propose SafeKeep, an inference-time safeguard that decouples safety judgment from tool execution: it assesses requests using flattened textual tool specifications while retaining the original schema-formatted specifications for execution. Across two representative benchmarks and four LLMs, including both white-box and black-box models, SafeKeep increases the average refusal rate for harmful requests from 23.8% to 70.6% and reduces the average attack success rate under observation-level prompt injection from 25.6% to 2.5%. It also outperforms existing safeguards and preserves task-handling capability. We release the code and data at https://github.com/snowcatsmoking/SafeKeep .
Chinese Translation
人工智能代理通过外部工具扩展大型语言模型(LLMs),使其能够执行复杂任务并将模型输出转化为具有实际影响的行动。然而,当LLMs作为代理部署时,其安全性往往显著降低,而这种退化的来源仍然不甚明了。本文识别出格式化的工具规范作为代理安全性退化的主要来源,并通过白盒表示分析显示,这些规范削弱了模型内部的拒绝信号,并导致不安全的工具执行。在此基础上,我们提出了SafeKeep,这是一种在推理时的安全保障措施,它将安全判断与工具执行解耦:它使用扁平化的文本工具规范评估请求,同时保留原始的格式化规范以供执行。在两个具有代表性的基准测试和四个LLMs(包括白盒和黑盒模型)中,SafeKeep将有害请求的平均拒绝率从23.8%提高到70.6%,并将观察级提示注入下的平均攻击成功率从25.6%降低到2.5%。它还优于现有的安全保障措施,并保持任务处理能力。我们在 https://github.com/snowcatsmoking/SafeKeep 发布了代码和数据。
cs.AI / 32 / 2607.29320

MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation

MAGA:通过结构化动作蒸馏实现多平台GUI代理的自我融合
Yan, Hang, GU, Zhangxuan, Zhou, Beitong, Chen, Jiaxuan, Li, Runze, Hu, Yusong, Shen, Shuheng, Meng, Changhua
Abstract
Graphical user interface (GUI) agents based on large language models are increasingly deployed across mobile, web, and desktop environments. However, existing agents are typically domain-specific, limiting the deployment and user experience. This motivates the consolidation of specialized models into a single cross-environment policy. Weight merging directly merges domain-specific experts but can corrupt executable actions under expert disagreement, while on-policy distillation (OPD) avoids conflicting teacher supervision yet still treats all response tokens equally during distillation, ignoring that action tokens are the only interface between the environment and the agent. To address this, We introduce MAGA that re-allocates training signal according to the structured action. Based on the correctness of the generated action, it suppresses unnecessary or invalid distillation signals and focuses learning on erroneous actions. Besides, a training-only hint optimizes the supervision signal provided by domain-specific teachers without changing the student input. Across two model scales, MAGA achieves the highest mean success rate, outperforming the strongest baseline by 2.0% at 8B and achieves almost the same average performance with teachers.
Chinese Translation
基于大型语言模型的图形用户界面(GUI)代理在移动、网页和桌面环境中越来越多地被部署。然而,现有的代理通常是特定领域的,这限制了其部署和用户体验。这促使我们将专门模型整合为一个跨环境的策略。权重合并直接合并领域特定的专家,但在专家意见不一致时可能会损坏可执行动作,而在政策蒸馏(OPD)中避免了冲突的教师监督,但在蒸馏过程中仍然将所有响应标记视为平等,忽视了动作标记是环境与代理之间的唯一接口。为了解决这个问题,我们提出了MAGA,根据结构化动作重新分配训练信号。基于生成动作的正确性,它抑制不必要或无效的蒸馏信号,并将学习重点放在错误的动作上。此外,训练专用提示优化了由领域特定教师提供的监督信号,而不改变学生输入。在两个模型规模下,MAGA实现了最高的平均成功率,在8B时比最强基线高出2.0%,并且与教师的平均性能几乎相同。
cs.AI / 33 / 2607.29405

Beyond Component Testing: Validating Agentic AI Systems

超越组件测试:验证自主智能系统
Mirto, Fabio Orazio, D'Agati, Luca, Tricomi, Giuseppe, Silvestri, Stefano, Longo, Francesco, Puliafito, Antonio, Merlino, Giovanni
Abstract
Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation. This behavior stretches validation practice beyond component testing and one-shot input--output evaluation, because acceptable system behavior now depends on how decisions unfold over time and under changing environmental conditions. This survey synthesizes 257 papers spanning agent evaluation, software assurance, cyber-physical systems, runtime monitoring, and regulatory guidance in order to characterize the validation problem for agentic systems. The review is organized around a five-dimension taxonomy covering behavioral, safety, temporal, regulatory, and multi-agent concerns, and uses that taxonomy to map current approaches and expose recurrent coverage gaps. The analysis shows that behavioral evaluation is comparatively mature, while temporal validity, runtime evidence maintenance, regulatory legibility, and open-ended multi-agent systems assurance remain under-developed. Three cross-domain case studies (medical care, industrial operations, smart-mobility systems) provide operational illustrations of how the five taxonomy dimensions recur in safety-critical settings, grounded in the failure patterns documented in the reviewed literature. The paper concludes with a lifecycle-oriented research agenda centered on bounded-autonomy specifications, adversarial trajectory generation, runtime monitoring, and audit-ready evidence structures. The central claim is that trustworthy deployment of agentic AI depends on validating trajectories in context rather than assessing isolated components alone.
Chinese Translation
自主智能系统通过多步骤轨迹进行操作,这些轨迹结合了规划、工具使用、记忆、互动和适应。这种行为使得验证实践超越了组件测试和一次性输入-输出评估,因为可接受的系统行为现在依赖于决策在时间推移和环境条件变化下的展开方式。本调查综合了257篇论文,涵盖了代理评估、软件保障、网络物理系统、运行时监控和监管指导,以表征自主系统的验证问题。该综述围绕一个五维分类法组织,涵盖行为、安全、时间、监管和多代理关注点,并利用该分类法映射当前的方法并揭示反复出现的覆盖缺口。分析表明,行为评估相对成熟,而时间有效性、运行时证据维护、监管可读性和开放式多代理系统保障仍然不够发达。三个跨领域案例研究(医疗护理、工业运营、智能出行系统)提供了在安全关键环境中如何反复出现五个分类维度的操作性示例,这些示例基于文献中记录的失败模式。本文最后提出了一个以生命周期为导向的研究议程,重点关注有限自主规范、对抗性轨迹生成、运行时监控和审计准备证据结构。核心论点是,可信赖的自主智能系统部署依赖于在上下文中验证轨迹,而不仅仅是评估孤立的组件。
cs.AI / 34 / 2607.29431

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

ModelEquivBench:认证多关系评估LLM生成的优化模型
Zhu, Penglin, Xu, Jungang
Abstract
Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently checkable nor faithful to the multiple distinct senses in which two formulations can agree. We present ModelEquivBench, a certifying, multi-relational evaluation system that reports a per-pair semantic profile E0--E6: model construction and exact ingestion (E0), verified representation alignment (E1), same-space and projected feasible-set relations (E2, E3), objective-order equivalence (E4), optimal-value equality (E5), and optimizer-set equivalence (E6). Each decided entry carries relation-appropriate, independently re-checkable evidence: replayable traces or explicit maps for E0--E1, exact-rational certificates for positive E2--E6 conclusions, and explicit witnesses for supported negatives. Incomplete mapping search, unsupported structure, and resource limits produce typed UNKNOWN or N/A outcomes rather than guesses, while unmet prerequisites are reported as ABSENT. Using ModelEquivBench to evaluate three model snapshots--GPT-5.4, Claude Sonnet 4.6, and Qwen3.5-397B-A17B--on the same frozen cohort of 173 base problems (346 cells per model) under a no-repair protocol, the resulting profiles expose distinctions that coarse baselines do not represent: 49, 35, and 25 cells contain executable candidates that are nevertheless certified negative on at least one supported relation, and 25, 8, and 18 structural rejections occur on pairs for which E2 certifies mapped feasible-set equality under a verified map. The three model snapshots fail at different stages of the profile and therefore cannot be meaningfully reduced to a single accuracy score.
Chinese Translation
大型语言模型越来越多地从自然语言生成优化模型,但现有评估通常将生成的模型及其真实值简化为单一的等价/不等价判决或执行成功率——这些标签既不可独立验证,也无法忠实于两种表述可以一致的多种不同意义。我们提出了ModelEquivBench,一个认证的多关系评估系统,报告每对模型的语义特征E0至E6:模型构建和精确摄取(E0)、验证的表示对齐(E1)、同空间和投影可行集关系(E2,E3)、目标顺序等价(E4)、最优值相等(E5)和优化器集等价(E6)。每个决定的条目都携带与关系相适应的、可独立重新检查的证据:可重放的追踪或明确的映射用于E0至E1,正向E2至E6结论的精确有理证书,以及支持的负向的明确证人。未完成的映射搜索、不支持的结构和资源限制产生类型为UNKNOWN或N/A的结果,而未满足的前提条件则报告为ABSENT。使用ModelEquivBench对三个模型快照——GPT-5.4、Claude Sonnet 4.6和Qwen3.5-397B-A17B——在同一组173个基础问题(每个模型346个单元)下进行评估,结果特征揭示了粗略基线未能表示的区别:49、35和25个单元包含可执行候选,但在至少一个支持关系上被认证为负向,且在E2认证的映射可行集相等的对中,发生了25、8和18个结构拒绝。这三个模型快照在特征的不同阶段失败,因此无法有意义地简化为单一的准确性评分。
cs.AI / 35 / 2607.29440

Beyond Retrieval: Analytic Memory for Multimodal Agents

超越检索:多模态智能体的分析记忆
Tian, Zhoujin, Tian, Yao, Zhang, Hao, Chen, Cheng, Li, Yakun, Zhang, Lei, Zhou, Xiaofang
Abstract
Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions. Existing systems largely emphasize \emph{retrieval memory}, organizing interaction histories through summaries and indexes to return query-relevant information at multiple granularities, from high-level abstractions to underlying records. In this paper, we formulate \emph{analytic memory} as a complementary abstraction that organizes recurring multimodal observations into queryable structures supporting filtering, aggregation, ranking, and temporal comparison. We present AdaMM, a framework that jointly supports retrieval and analytic memory. Rather than relying on application-defined schemas, AdaMM extracts provenance-linked attribute-value observations from dialogue, images, and contextual metadata, discovers recurring field structures, and materializes them for analytical access. At inference time, a memory-aware planner decomposes queries into retrieval and analytic operations and routes each operation to the appropriate tools. Experiments on two long-term multimodal memory benchmarks, MemEye and MemGallery, show that AdaMM improves performance by up to 11.3\% and 7.3\%, respectively.
Chinese Translation
长期多模态记忆不仅必须支持检索相关信息,还需对在交互过程中积累的观察结果进行计算。现有系统主要强调 extit{检索记忆},通过摘要和索引组织交互历史,以在多个粒度上返回与查询相关的信息,从高层次的抽象到底层记录。本文将 extit{分析记忆}定义为一种互补的抽象,旨在将重复出现的多模态观察组织成可查询的结构,以支持过滤、聚合、排序和时间比较。我们提出了AdaMM,一个同时支持检索和分析记忆的框架。AdaMM并不依赖于应用定义的模式,而是从对话、图像和上下文元数据中提取与来源相关的属性-值观察,发现重复的字段结构,并将其物化以便进行分析访问。在推理时,记忆感知规划器将查询分解为检索和分析操作,并将每个操作路由到适当的工具。在两个长期多模态记忆基准测试MemEye和MemGallery上的实验表明,AdaMM的性能分别提高了高达11.3 extperthousand和7.3 extperthousand。
cs.AI / 36 / 2607.29468

Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

自我对弈与技能演化:自我演化搜索代理的提出、解决与记忆
Fu, Zenghuang, Li, Zhaoyang, Ai, Qiuyuan, Wu, Haoyu, Wu, Minghui, Zhao, Chenxu, Wang, Ante, He, Guannan, Wang, Changwei
Abstract
Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failures affect gradients yet do not explicitly shape future practice. External skill memories preserve procedural experience but are typically learned from fixed task distributions. We introduce \textbf{SESA} (Self-Evolving Skill-Augmented Agent), which makes procedural memory an evolving state of tool-augmented search self-play. A challenger poses problems, while a separately parameterized solver alone retrieves skills. Informative failures are distilled into reusable skills and written back to memory. The updated memory changes solver behavior and success, which changes the challenger's reward and the distribution of future problems; the resulting frontier produces new failures that rewrite memory. This bidirectional loop makes task generation and skill memory co-evolve. Because retrieved skills shape on-policy training trajectories, their benefits can enter the model parameters as well as remain in the external bank, enabling memory-free deployment and optional inference-time retrieval. Across seven open-domain and multi-hop question-answering benchmarks, SESA improves average accuracy over SSP by 1.2--3.2 points across multiple backbones and surpasses the skill-augmented SkillRL baseline by 0.9 points under a unified evaluation protocol. On Qwen3 models, SESA-Off retains 1.8--2.2 points of improvement over SSP, while the final skill bank adds a further 0.5--1.0 points. These results show that evolving skill memory is not merely an inference-time plug-in: it changes policy learning and the future training distribution while retaining value as optional external memory. Our code is available at https://github.com/Zenghuang-Fu/SESA-Self-Evolving-Search-Agents.
Chinese Translation
自我对弈代理能够生成无需目标基准问题的训练问题,但其课程缺乏持久状态:失败影响梯度,却并未明确塑造未来的实践。外部技能记忆保留了程序经验,但通常是从固定任务分布中学习的。我们提出了 extbf{SESA}(自我演化技能增强代理),使程序记忆成为工具增强搜索自我对弈的演化状态。挑战者提出问题,而一个单独参数化的求解器则独立检索技能。信息丰富的失败被提炼为可重用的技能并写回记忆。更新后的记忆改变了求解器的行为和成功率,从而改变了挑战者的奖励和未来问题的分布;由此产生的边界产生新的失败,重写记忆。这个双向循环使任务生成和技能记忆共同演化。由于检索的技能塑造了策略训练轨迹,其好处不仅可以进入模型参数,还可以保留在外部库中,从而实现无记忆部署和可选的推理时检索。在七个开放领域和多跳问答基准上,SESA在多个基础模型上将平均准确率提高了1.2至3.2个百分点,并在统一评估协议下超越了技能增强的SkillRL基线0.9个百分点。在Qwen3模型上,SESA-Off相较于SSP保留了1.8至2.2个百分点的提升,而最终的技能库又增加了0.5至1.0个百分点。这些结果表明,演化技能记忆不仅仅是推理时的插件:它改变了策略学习和未来的训练分布,同时作为可选的外部记忆保留了价值。我们的代码可在 https://github.com/Zenghuang-Fu/SESA-Self-Evolving-Search-Agents 获取。
cs.AI / 37 / 2607.29549

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

AMTFV:用于大语言模型自我修正的代理数学工具流验证
Zou, Rui, Zhu, Yutao, Wei, Mengqi, Wen, Ji-Rong
Abstract
Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challenging. Existing representative methods mainly revise outputs through natural-language reflection or assist verification by directly generating verification programs; the former may not reliably support exact computation, whereas the latter prematurely couples mathematical modeling with low-level implementation. We propose AMTFV (Agentic Mathematical Tool-Flow Verification). By introducing Mathematical Tool Flow (MTF) as an interrupt--execute--resume interface, AMTFV decouples verification modeling from concrete execution and supports exact computation through a mathematical toolbox. Specifically, the verification agent first constructs a verification workflow, encodes the mathematical objects and computational intent requiring reliable execution in an MTF request, and sends it to the mathematical toolbox agent. The latter parses the request, generates executable calls, and dispatches them to the backend for exact computation. Tool outputs then support candidate-answer adjudication, answer revision, and verification-workflow revision. We evaluate AMTFV on five challenging mathematical reasoning datasets with seven model configurations from DeepSeek, GPT, and Gemini. Experimental results show that AMTFV outperforms the representative baselines evaluated in this study overall; under an individual model configuration, it improves average accuracy over the strongest baseline by up to 8.3 percentage points, with larger gains on samples of medium and high verification complexity.
Chinese Translation
大型语言模型已展示出强大的数学问题解决能力,但可靠地验证其候选答案仍然具有挑战性。现有的代表性方法主要通过自然语言反思修正输出,或通过直接生成验证程序来辅助验证;前者可能无法可靠地支持精确计算,而后者则过早地将数学建模与低级实现耦合。我们提出了AMTFV(代理数学工具流验证)。通过引入数学工具流(Mathematical Tool Flow, MTF)作为中断-执行-恢复接口,AMTFV将验证建模与具体执行解耦,并通过数学工具箱支持精确计算。具体而言,验证代理首先构建验证工作流,将需要可靠执行的数学对象和计算意图编码为MTF请求,并将其发送给数学工具箱代理。后者解析请求,生成可执行调用,并将其分派到后端进行精确计算。工具输出随后支持候选答案裁定、答案修正和验证工作流修正。我们在五个具有挑战性的数学推理数据集上评估了AMTFV,使用了来自DeepSeek、GPT和Gemini的七种模型配置。实验结果表明,AMTFV在本研究评估的代表性基线中整体表现优于其他方法;在单个模型配置下,其平均准确率比最强基线提高了多达8.3个百分点,在中等和高验证复杂度的样本上获得了更大的提升。
cs.AI / 38 / 2607.29553

COntExt: Towards Context-Aware Ontology Extension from Operational Metrics

COntExt:基于操作指标的上下文感知本体扩展
Hussain, Hussain, Schöberl, Stefan, Schneider, Angelika, Geist, Verena
Abstract
Organizations increasingly define operational metrics in structured, machine-readable formats to monitor systems, processes, and compliance. These metric definitions implicitly encode domain knowledge, such as referencing concepts, properties, and relationships, that often extends what is captured in formal ontologies. Yet the connection between operational metric catalogues and ontological knowledge remains manual, ad-hoc, and labor-intensive. We present COntExt, a framework for context-aware ontology extension that takes structured metric definitions as input and suggests how referenced concepts and properties should be integrated into an existing ontology, utilizing the context of these metrics. The framework defines the extension problem as three sub-tasks: parent class prediction, relation type prediction, and data property assignment. Across four cybersecurity ontologies, we evaluate different algorithms for each task. Our results show that metric-derived context improves the suggestions over ontology-context baselines for relation type prediction and data property assignment. Our work demonstrates that operational metric catalogues are a practical and underexploited source for ontology extension. This work enables organizations to maintain their ontologies at a significantly lower cost than manual engineering.
Chinese Translation
组织越来越多地以结构化、机器可读的格式定义操作指标,以监控系统、流程和合规性。这些指标定义隐含地编码了领域知识,例如引用概念、属性和关系,这些内容通常超出了正式本体所捕获的范围。然而,操作指标目录与本体知识之间的连接仍然是手动的、临时的和劳动密集型的。我们提出了COntExt,一个上下文感知本体扩展框架,它以结构化的指标定义为输入,并建议如何将引用的概念和属性整合到现有本体中,利用这些指标的上下文。该框架将扩展问题定义为三个子任务:父类预测、关系类型预测和数据属性分配。在四个网络安全本体上,我们评估了每个任务的不同算法。我们的结果表明,基于指标的上下文在关系类型预测和数据属性分配方面改善了建议,相较于本体上下文基线。我们的工作表明,操作指标目录是一个实用且未被充分利用的本体扩展来源。这项工作使组织能够以显著低于手动工程的成本维护其本体。
cs.AI / 39 / 2607.29559

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

LEMUR:基于偏好反馈的多目标强化学习对齐学习
Adikari, Manith, Peng, Bei, Vinanzi, Samuele, Cangelosi, Angelo
Abstract
Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. However, real-world decision-making tasks often involve multiple, competing objectives, such as performance versus efficiency, where ground-truth reward functions are difficult to specify or inaccessible. While Multi-Objective RL (MORL) addresses such trade-offs by modeling rewards as vectors, existing approaches typically assume access to a well-specified reward function for each objective, inheriting the same challenges faced by single-objective RL. Meanwhile, Preference-based RL (PbRL) has shown great potential in solving complex tasks without access to a pre-defined reward function through reward learning from human feedback, yet has largely been studied in single-objective settings. In this work, we bridge this gap with LEMUR: Learning to Align with Multi-Objective Reinforcement Learning with Preference feedback, a novel framework where an agent interactively learns from the preferences of multiple humans to learn optimal multi-objective policies. Our approach jointly learns policies and multiple objective-specific reward models from human feedback, enabling agents to effectively balance competing objectives during learning. We evaluate LEMUR on a variety of benchmark multi-objective tasks, and empirical results demonstrate its superior performance over baseline methods. Our method presents a promising direction for solving multi-objective decision-making tasks without pre-defined reward functions.
Chinese Translation
强化学习(RL)系统通常使用单一、明确的标量奖励函数进行训练。然而,现实世界中的决策任务往往涉及多个相互竞争的目标,例如性能与效率之间的权衡,此时真实的奖励函数难以明确或无法获取。虽然多目标强化学习(MORL)通过将奖励建模为向量来解决此类权衡,但现有方法通常假设可以访问每个目标的明确奖励函数,因此继承了单目标强化学习所面临的相同挑战。同时,基于偏好的强化学习(PbRL)在没有预定义奖励函数的情况下,通过从人类反馈中学习奖励,展现出在解决复杂任务中的巨大潜力,但主要集中于单目标设置。在本研究中,我们通过LEMUR:基于偏好反馈的多目标强化学习对齐学习,填补了这一空白。LEMUR是一个新颖的框架,代理通过与多个人类的偏好互动学习,从而学习最优的多目标策略。我们的方法从人类反馈中共同学习策略和多个目标特定的奖励模型,使代理在学习过程中有效平衡相互竞争的目标。我们在多种基准多目标任务上评估了LEMUR,实证结果表明其优于基线方法。我们的方法为在没有预定义奖励函数的情况下解决多目标决策任务提供了一个有前景的方向。
cs.AI / 40 / 2607.29577

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

DungeonBench:一个用于《龙与地下城》战斗中规则丰富的战术推理的基准测试
Ismayilov, Ismayil, Kara, Atakan, Oktay, Kaan
Abstract
Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reasoning: the ability to choose well when geometry, timing, resources, objectives, and rule interactions all matter at once. We introduce DungeonBench, a benchmark for tactical reasoning in Dungeons & Dragons combat, built to cover the vast majority of combat-relevant 2014 System Reference Document content whose effects can be resolved by the simulator while retaining mechanics that simplified combat simulators often abstract away. At each step, DungeonBench exposes a complete tactical observation, a pending decision, and an indexed list of executable options spanning movement, attacks, spells, reactions, objectives, preparation, and scarce resources. The task is to value legal choices whose consequences depend on action economy, creature traits, battlefield geometry, timing windows, and future encounters. DungeonBench has two tracks: Encounter, which evaluates local tactical play in single fights, and Day, which links encounters through persistent hit points, spell slots, consumables, preparation, and short-rest timing, forcing policies to trade off immediate tactical advantage against future survivability. The same engine-generated decision stream supports heuristic controllers, language-model policies, learned option rankers, and masked-action reinforcement-learning agents. We evaluate frontier language-model policies on this shared decision stream. Results show that full tactical observations do not saturate the benchmark: frontier policies often win direct encounters, but linked encounter days expose failures in resource budgeting, rest timing, and rule-aware tactical discipline.
Chinese Translation
游戏和模拟器通过将决策转化为可测量的结果,提供了有价值的基准测试,但许多现有的测试套件在规则丰富的战术推理方面测试不足:即在几何形状、时机、资源、目标和规则交互都同时重要的情况下做出良好选择的能力。我们介绍了DungeonBench,这是一个用于《龙与地下城》战斗中的战术推理基准,旨在覆盖2014年系统参考文档中绝大多数与战斗相关的内容,其效果可以通过模拟器解决,同时保留那些简化战斗模拟器通常会抽象掉的机制。在每一步中,DungeonBench提供完整的战术观察、待决策和可执行选项的索引列表,这些选项涵盖移动、攻击、法术、反应、目标、准备和稀缺资源。任务是评估合法选择的价值,这些选择的后果依赖于行动经济、生物特性、战场几何、时机窗口和未来遭遇。DungeonBench有两个轨道:Encounter(遭遇),评估单场战斗中的局部战术玩法;Day(天),通过持续的生命值、法术位、消耗品、准备和短暂休息的时机将遭遇联系起来,迫使策略在即时战术优势与未来生存能力之间进行权衡。相同的引擎生成决策流支持启发式控制器、语言模型策略、学习的选项排名器和掩蔽动作强化学习代理。我们在这个共享决策流上评估前沿语言模型策略。结果表明,完整的战术观察并未饱和基准测试:前沿策略通常在直接遭遇中获胜,但相互关联的遭遇日暴露了资源预算、休息时机和规则意识战术纪律方面的失败。
cs.AI / 41 / 2607.29626

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

AgentHPOBench:评估大型语言模型代理作为顺序超参数优化器的基准
Huai, Tianyu, Fan, Tingshuo, Chen, Xinchi, Zheng, Yining, Wang, Yuxin, Chen, Shuang, Zhou, Jie, Huang, Xuanjing
Abstract
As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmarks typically focus on static code generation, paper replication, or final answer correctness, but do not directly assess whether agents can interpret experimental evidence and use it to guide subsequent hyperparameter decisions. To address this gap, we introduce AgentHPOBench, a sequential benchmark comprising 30 executable machine learning tasks across seven research categories. Each task begins with a validated baseline run, after which an agent performs several sequential interventions. At each step, the agent observes the accumulated configurations, metrics, and logs before proposing the next valid configuration. We evaluate 12 widely used agents and conventional HPO baselines under a unified protocol. The results show that current agents exhibit measurable experimental optimization ability across domains, but still face clear limitations in sustained iterative refinement, complex log diagnosis, and consistent progress toward reported reference performance.
Chinese Translation
随着大型语言模型(LLMs)从代码补全系统演变为自主科学代理,评估它们进行实验的能力变得越来越重要。现有基准通常侧重于静态代码生成、论文复制或最终答案的正确性,但并未直接评估代理是否能够解释实验证据并利用这些证据指导后续的超参数决策。为了解决这一空白,我们引入了AgentHPOBench,这是一个包含30个可执行机器学习任务的顺序基准,涵盖七个研究类别。每个任务都以经过验证的基线运行开始,之后代理执行多个顺序干预。在每一步中,代理观察累积的配置、指标和日志,然后提出下一个有效配置。我们在统一协议下评估了12个广泛使用的代理和传统的超参数优化(HPO)基线。结果表明,当前代理在各个领域展现出可测量的实验优化能力,但在持续的迭代改进、复杂日志诊断和朝着报告的参考性能的一致进展方面仍面临明显的限制。
cs.AI / 42 / 2607.29657

Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and Diagnostics

FDD-ON的开发:用于变风量(VAV)暖通空调系统故障检测与诊断的本体
Chen, Yimin, Fricke, Brian, Shen, Bo, Lian, Jamie, Zhang, Mingkan, Lo, James, Zhang, Yun, Ye, Shi, Huang, Jiajing, Hu, Han, Lu, Chujie, Tang, Rui, Zhuang, George
Abstract
Fault detection and diagnosis (FDD) technology is essential for improving HVAC system reliability, energy efficiency, and maintenance effectiveness. However, effective deployment of FDD solutions in buildings requires structured domain knowledge that can bridge heterogeneous data sources, diverse equipment types, and varied diagnostic outputs. Limited data interpretability and interoperability within the FDD domain have led to fragmented information silos, hindering the implementation of FDD and related applications, such as the digital twin-enabled FDD frameworks and artificial intelligence (AI)-driven maintenance decision-making systems. This paper presents an FDD Ontology (FDD-ON), a modular and extensible ontology to formally represent variable air volume (VAV) HVAC system components, fault types, symptom statuses, fault impacts and associated attributes. FDD-ON integrates HVAC system FDD semantics to provide comprehensive representations of fault and symptom attributes, supported by the well-defined controlled vocabulary. Additionally, FDD-ON offers comprehensive fault, symptom, and impact libraries to capture a broad spectrum of operational abnormalities and their consequences in VAV HVAC systems. Through explicit contributing cause-fault-symptom-impact relations, FDD-ON serves as a machine-interpretable basis for querying diagnostic knowledge, mapping heterogeneous FDD outputs, and developing interoperable FDD-related applications. FDD-ON is evaluated using publicly available VAV HVAC system datasets and demonstrated through FDD development applications. Results indicate that FDD-ON provides a foundational semantic framework for advancing scalable, transparent, and interoperable FDD solutions across various applications.
Chinese Translation
故障检测与诊断(FDD)技术对于提高暖通空调(HVAC)系统的可靠性、能源效率和维护有效性至关重要。然而,在建筑中有效部署FDD解决方案需要结构化的领域知识,以便能够连接异构数据源、多样化的设备类型和不同的诊断输出。FDD领域内有限的数据可解释性和互操作性导致了信息孤岛的碎片化,阻碍了FDD及相关应用的实施,例如数字双胞胎驱动的FDD框架和人工智能(AI)驱动的维护决策系统。本文提出了一种FDD本体(FDD-ON),这是一个模块化和可扩展的本体,用于正式表示变风量(VAV)HVAC系统组件、故障类型、症状状态、故障影响及相关属性。FDD-ON整合了HVAC系统FDD语义,以提供故障和症状属性的全面表示,支持良好定义的控制词汇。此外,FDD-ON还提供全面的故障、症状和影响库,以捕捉VAV HVAC系统中广泛的操作异常及其后果。通过明确的成因-故障-症状-影响关系,FDD-ON作为机器可解释的基础,支持查询诊断知识、映射异构FDD输出,并开发可互操作的FDD相关应用。FDD-ON使用公开可用的VAV HVAC系统数据集进行评估,并通过FDD开发应用进行演示。结果表明,FDD-ON为推动各种应用中可扩展、透明和可互操作的FDD解决方案提供了基础语义框架。
cs.AI / 43 / 2607.29677

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

ExtractBench:一种用于模式引导的企业文档提取的基准测试
Zhang, Boyang, Lyjak, Adrian, Stewart, Eli, Li, Zhaoqi, Suo, Simon
Abstract
Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a benchmark for schema-guided extraction and, to our knowledge, the first to score value accuracy, record completeness at scale, grounding, and measured cost together. The evaluation system contains 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types, with clear tags differentiating their challenge scenarios. The scalable schema and ground-truth curation pipeline combines independent-system agreement for real documents, known values for synthetic lists, and human verification for forms. We report order-insensitive value F1 for value accuracy, plus two grounding metrics for source traceability: word- and page-level F1. Commercial VLMs perform well on short documents but often truncate record lists on long ones, while coding agents retain higher accuracy at much higher cost. LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost. Dataset and evaluation code are available on \href{https://huggingface.co/datasets/llamaindex/ExtractBench}{HuggingFace} and \href{https://github.com/run-llama/ExtractBench}{GitHub}.
Chinese Translation
企业工作流程越来越依赖于代理进行 extit{模式引导提取}:给定一个文档和用户定义的模式,代理忠实地遵循该模式生成正确的输出,并提供源证据作为基础元数据。我们提出了ExtractBench,这是一种用于模式引导提取的基准测试,且据我们所知,它是首个同时评估值准确性、记录完整性、基础性和成本的基准系统。评估系统包含4,869页来自370个企业文档的数据,涵盖8个业务领域和67种文档类型,并清晰标记其挑战场景。可扩展的模式和真实数据的基准策划流程结合了真实文档的独立系统一致性、合成列表的已知值以及表单的人类验证。我们报告了值准确性的无序敏感值F1,以及两个用于源可追溯性的基础性指标:词级和页级F1。商业VLM在短文档上表现良好,但在长文档上往往会截断记录列表,而编码代理在更高成本下保持更高的准确性。LlamaExtract Agentic Plus在所有三个指标上均排名第一,其准确性与编码代理相当,但成本仅为其一小部分。数据集和评估代码可在 exttt{https://huggingface.co/datasets/llamaindex/ExtractBench}和 exttt{https://github.com/run-llama/ExtractBench}上获取。
计算语言学 (Computation and Language)
51
cs.CL / 1 / 2607.28634

Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs

大型语言模型真的能理解题目难度水平吗?对使用大型语言模型进行自动题目生成的启示
Wang, Xinyi, Jiao, Hong, Li, Ming, Peters, Sydney, Choi, Hanna, Zhou, Tianyi, Xu, Qingshu
Abstract
The estimation of item difficulty plays a key role in both formative assessment and large-scale high-stakes summative assessments. This study explores how large language models (LLMs) perform in predicting item difficulty levels using items from a large-scale Reading and Writing test. The study investigated various prompting strategies and parameter settings across multiple LLMs. LLM performance was compared with encoder-only language models and feature-based supervised machine learning models. Zero-shot GPT-4.1 with a temperature of 0 yielded the highest item difficulty level prediction accuracy, with a quadratic weighted kappa (QWK) of 0.578. However, LLMs' prediction accuracy was lower than that of ConvBERT (QWK = 0.625), which outperformed the best feature-based supervised machine learning model. Further analysis showed that all LLMs struggled to label hard items; in particular, the current advanced GPT-5.4 tended to underestimate item difficulty levels. Dimension reduction of embeddings showed that item embeddings from different difficulty levels were mixed together, indicating that semantic information from items alone is likely insufficient for item difficulty level prediction. The findings suggest that if LLMs cannot understand item difficulty levels as evidenced by empirical data and tend to treat most items as easy when their own capabilities increase, caution should be exercised when using LLMs to generate items with targeted difficulty levels.
Chinese Translation
题目难度的估计在形成性评估和大规模高风险总结性评估中发挥着关键作用。本研究探讨了大型语言模型(LLMs)在预测题目难度水平方面的表现,使用了大规模阅读与写作测试中的题目。研究考察了多种提示策略和参数设置在多个LLMs中的表现。LLM的表现与仅编码的语言模型和基于特征的监督机器学习模型进行了比较。零-shot的GPT-4.1在温度为0的情况下,达到了最高的题目难度水平预测准确率,二次加权κ系数(QWK)为0.578。然而,LLMs的预测准确率低于ConvBERT(QWK = 0.625),后者的表现优于最佳的基于特征的监督机器学习模型。进一步分析显示,所有LLMs在标记困难题目时均表现不佳;特别是当前的先进模型GPT-5.4倾向于低估题目难度水平。嵌入的降维分析表明,不同难度水平的题目嵌入混合在一起,表明仅凭题目的语义信息可能不足以预测题目难度水平。研究结果表明,如果LLMs无法理解题目难度水平(如实证数据所示),并且在其自身能力提高时倾向于将大多数题目视为简单,那么在使用LLMs生成具有特定难度水平的题目时应谨慎行事。
cs.CL / 2 / 2607.28635

Imbalanced Data Clustering via Targeted Data Augmentation Using GMM and LLM

通过使用GMM和LLM的针对性数据增强进行不平衡数据聚类
Khalal, Noor, Djamai, Abdallah Alaa-Eddine, Keraghel, Imed, Nadif, Mohamed
Abstract
In Natural Language Processing (NLP), dealing with underrepresented topics is challenging, especially in unsupervised tasks where clustering might not adequately capture minority topics. To tackle this challenge, our paper presents a novel unsupervised data augmentation method that integrates Gaussian Mixture Models (GMMs) and Large Language Models (LLMs). Due to their flexibility and robustness, GMMs can detect clusters corresponding to underrepresented areas in the data, while LLMs create synthetic documents to enrich these clusters and improve their representation. Experiments on various imbalanced text datasets demonstrate that our approach preserves clustering performance in all cases and often enhances cluster interpretability, offering a robust and scalable solution for improving data representation in unsupervised NLP tasks.
Chinese Translation
在自然语言处理(NLP)中,处理代表性不足的主题是一项挑战,尤其是在无监督任务中,聚类可能无法充分捕捉到少数主题。为了解决这一挑战,本文提出了一种新颖的无监督数据增强方法,结合了高斯混合模型(GMM)和大型语言模型(LLM)。由于其灵活性和鲁棒性,GMM能够检测到与数据中代表性不足区域相对应的聚类,而LLM则生成合成文档以丰富这些聚类并改善其表示。对各种不平衡文本数据集的实验表明,我们的方法在所有情况下都保持了聚类性能,并且通常增强了聚类的可解释性,为改善无监督NLP任务中的数据表示提供了一种稳健且可扩展的解决方案。
cs.CL / 3 / 2607.28636

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

模型链:针对偏见稳健的LLM评审的跨模型审计
Wang, Qian, Lou, Zhanzhi, Tang, Zhenheng, Chen, Nuo, He, Bingsheng
Abstract
LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-driven debiasing, which is brittle across bias types, or human evaluation, which does not scale. We study \emph{Chain-of-Models} (CoM), an automated audit pipeline in which a second model inspects the first model's reasoning trace before producing the final judgment. The key design question is whether the auditor should be the same model, a same-family model, or a different-family model. Across 9 models from 6 families, 4 cognitive biases, and 4 factual datasets, we find that auditor identity matters in two ways. First, standalone bias resistance does not predict audit effectiveness: Kimi-K2.5 is the strongest standalone model on several biases, yet is a weak auditor for Qwen2.5-72B's biased traces. Second, the best auditor is bias-specific: GPT-4o is strongest on bandwagon, authority, and distraction, while GLM-5 is strongest on sycophancy. We operationalize these findings with a per-bias auditor selection rule that, given the bias type, scores candidates along functional diversity, per-bias standalone resistance, and calibrated audit effectiveness. Under a calibration/test split, the selector reaches the highest accuracy across the four biased slices ($0.884$ vs.\ $0.824$ for the strongest single fixed auditor and $0.805$ for the no-audit baseline). We release data, configurations, and an LLM-agent skill at https://anonymous.4open.science/r/chain-of-models-B585 .
Chinese Translation
大型语言模型(LLMs)越来越多地作为自动评审者,但它们的判断仍然容易受到认知偏见的影响。现有的缓解措施主要依赖于基于提示的去偏见方法,这在不同类型的偏见中表现脆弱,或依赖于人工评估,这无法扩展。我们研究了 extit{模型链}(Chain-of-Models, CoM),这是一种自动审计管道,其中第二个模型在生成最终判断之前检查第一个模型的推理轨迹。关键的设计问题是审计者应该是相同的模型、同一家族的模型,还是不同家族的模型。在来自6个家族的9个模型、4种认知偏见和4个事实数据集的实验中,我们发现审计者的身份在两个方面很重要。首先,独立的偏见抵抗能力并不能预测审计的有效性:Kimi-K2.5在几种偏见中是最强的独立模型,但对于Qwen2.5-72B的偏见轨迹却是一个弱审计者。其次,最佳审计者是特定于偏见的:GPT-4o在从众、权威和分心偏见上表现最佳,而GLM-5在谄媚偏见上表现最佳。我们通过一种针对每种偏见的审计者选择规则来实现这些发现,该规则根据偏见类型对候选者在功能多样性、每种偏见的独立抵抗能力和校准审计有效性方面进行评分。在校准/测试分割下,该选择器在四个偏见切片中达到了最高的准确率($0.884$,相比之下,最强的单一固定审计者为$0.824$,无审计基线为$0.805$)。我们在https://anonymous.4open.science/r/chain-of-models-B585发布了数据、配置和LLM代理技能。
cs.CL / 4 / 2607.28637

ZeroR@CHiPSAL 2026: Two-Stage Vision-Language Adaptation with Contrastive Learning for Nepali Meme Classification

ZeroR@CHiPSAL 2026:基于对比学习的两阶段视觉-语言适应用于尼泊尔表情包分类
Khanal, Nitiz
Abstract
This paper presents our system for the CHiPSAL 2026 shared task on multimodal hate speech and sentiment detection in Nepali memes. We address both subtasks: binary hate speech classification and three-class sentiment analysis. Our approach adapts the Robust Adaptation of Hateful Meme Detection (RA-HMD) framework using Qwen3-VL-8B-Instruct, a state-of-the-art vision-language model with native Devanagari support. We employ a two-stage training pipeline: (1) LoRA fine-tuning with an MLP projection head for generative classification, and (2) contrastive backbone fine-tuning with supervised InfoNCE loss. We handle class imbalance through minority oversampling, image augmentation, and focal loss. At inference, we ensemble Stage 1 token probabilities with Stage 2 classifier scores using validation-tuned weights. Our end-to-end approach eliminates error propagation from separate OCR and translation pipelines by leveraging the model's native Devanagari understanding. Our system achieved \textbf{2nd place} on hate speech detection (F1: 0.797) and \textbf{4th place} on sentiment analysis (F1: 0.518). We provide detailed ablations, error analysis, and insights into adapting large vision-language models for low-resource South Asian languages.
Chinese Translation
本文介绍了我们在CHiPSAL 2026共享任务中针对尼泊尔表情包的多模态仇恨言论和情感检测的系统。我们解决了两个子任务:二元仇恨言论分类和三类情感分析。我们的方法采用了仇恨表情包检测的稳健适应框架(Robust Adaptation of Hateful Meme Detection, RA-HMD),并使用了Qwen3-VL-8B-Instruct,这是一个具有本地德瓦那加里语支持的最先进视觉-语言模型。我们采用了两阶段的训练流程:(1)使用多层感知器(MLP)投影头进行生成分类的LoRA微调,以及(2)使用监督的InfoNCE损失进行对比主干微调。我们通过少数类过采样、图像增强和焦点损失来处理类别不平衡。在推理阶段,我们使用经过验证调整的权重将第一阶段的标记概率与第二阶段的分类器得分进行集成。我们端到端的方法通过利用模型对德瓦那加里语的本地理解,消除了来自独立OCR和翻译流程的错误传播。我们的系统在仇恨言论检测中获得了 extbf{第二名}(F1: 0.797),在情感分析中获得了 extbf{第四名}(F1: 0.518)。我们提供了详细的消融实验、错误分析以及对大型视觉-语言模型在低资源南亚语言适应中的见解。
cs.CL / 5 / 2607.28638

Learning Stateful Predictive Knowledge From Experience

从经验中学习有状态的预测知识
Song, Yan, Feng, Xidong, Liu, Bo, Cui, Xinyu, Fu, Haotian, Liu, Zichen, Yang, Mengyue, Deng, Cheng, Zhao, Jian, Wang, Jun
Abstract
As large language model (LLM) agents increasingly learn from experience, they primarily rely on trajectory-level reflection to extract insights. Viewed through the lens of predictive knowledge, we argue that this approach operates on episodic hindsight rather than predictive foresight, yielding brittle, path-dependent heuristics. To address this, we propose Stateful Knowledge Learning (SKL). SKL shifts the agent's focus from trajectory-level summarization to maintaining Stateful Knowledge: explicit, declarative predictive assessments anchored to state. We first demonstrate a motivating example showing how stateful knowledge provides granularity, enhances generalization, and enables knowledge bootstrapping. To further scale up the idea, we introduce two algorithms via self-distillation (SKL-SD) and reinforcement learning (SKL-RL), training agents to autonomously extract state-grounded predictive knowledge from experience and learn to leverage it for policy making. Experiments on interactive environments (WebShop, ScienceWorld) and a complex reasoning task (ChessPuzzles) demonstrate that equipping models with the inherent ability to learn stateful predictive knowledge significantly outpaces current reflection-based training paradigms.
Chinese Translation
随着大型语言模型(LLM)代理越来越多地从经验中学习,它们主要依赖于轨迹级反思来提取洞见。从预测知识的角度来看,我们认为这种方法是基于情节的事后反思,而非预测性的前瞻性,这导致了脆弱的、路径依赖的启发式方法。为了解决这个问题,我们提出了有状态知识学习(Stateful Knowledge Learning, SKL)。SKL将代理的重点从轨迹级总结转向维持有状态知识:明确的、声明性的预测评估,锚定于状态。我们首先展示了一个激励示例,说明有状态知识如何提供细粒度、增强泛化能力,并使知识自举成为可能。为了进一步扩展这一思路,我们通过自蒸馏(SKL-SD)和强化学习(SKL-RL)引入了两种算法,训练代理从经验中自主提取基于状态的预测知识,并学习如何利用这些知识进行决策。在交互环境(WebShop、ScienceWorld)和复杂推理任务(ChessPuzzles)上的实验表明,赋予模型学习有状态预测知识的内在能力显著超越了当前基于反思的训练范式。
cs.CL / 6 / 2607.28639

The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

知识蒸馏对小型语言模型偏差的非对称影响
Rath, Plawan Kumar
Abstract
We show that knowledge distillation in small instruction-tuned language models has asymmetric effects on bias. On unambiguous tasks (BBQ-disambig), response-based distillation from a Gemma-2-9B teacher improves context-following: for the most biased baseline (SmolLM2-1.7B-Instruct), it cuts the context-overriding error rate from 44% to 24%. On ambiguous tasks (BBQ-ambig), the same distillation destroys per-item refusal calibration: 15% of items where the baseline correctly abstained instead receive stereotype answers, even when overall refusal rate is preserved. The pattern reproduces on a second student family (OLMo-2-1B-Instruct), with silence-loss of 8% and filled-silence accounting for 89% of new bias. Across the full 28-configuration grid, the magnitudes of silence-loss and filled-silence are uncorrelated (Spearman $\rho=0.19$, n.s.), indicating that the two effects arise from distinct mechanisms. Aggregate stereotype metrics (CrowS-Pairs, overall BBQ Stereotype Reliance Score) average over both effects and conceal the per-item harm. We trace the calibration loss to a data-side mechanism: an audit of four training corpora finds <0.5% refusal-as-answer-shape. Supervised fine-tuning (SFT) with refusal injection either breaks parsing or over-corrects into a trivial-refuser regime (refusal rate 99.8%, disambig accuracy 0.2%) that aggregate metrics would call perfectly calibrated. We propose Per-Condition Calibration Diagnosis (PCCD), a three-step protocol that evaluates refusal calibration, context-following, and capability preservation. PCCD catches both the asymmetric harm and the trivial-refuser failure mode that aggregate evaluations miss.
Chinese Translation
我们展示了在小型指令调优语言模型中,知识蒸馏对偏差具有非对称影响。在明确任务(BBQ-disambig)中,来自 Gemma-2-9B 教师的基于响应的蒸馏改善了上下文跟随:对于最偏见的基线模型(SmolLM2-1.7B-Instruct),它将上下文覆盖错误率从 44% 降低到 24%。在模糊任务(BBQ-ambig)中,相同的蒸馏却破坏了逐项拒绝校准:在基线正确拒绝的 15% 项目中,反而得到了刻板印象的答案,即使整体拒绝率得以保持。该模式在第二个学生模型(OLMo-2-1B-Instruct)中重现,沉默损失为 8%,填充沉默占新偏差的 89%。在完整的 28 种配置网格中,沉默损失和填充沉默的大小不相关(Spearman $ ho=0.19$, n.s.),这表明这两种效应源于不同的机制。聚合的刻板印象指标(CrowS-Pairs,整体 BBQ 刻板印象依赖评分)对这两种效应进行了平均,掩盖了逐项的危害。我们将校准损失追溯到数据侧机制:对四个训练语料库的审计发现拒绝作为答案的形状少于 0.5%。带有拒绝注入的监督微调(SFT)要么破坏了解析,要么过度校正为一个琐碎拒绝者的状态(拒绝率 99.8%,消歧准确率 0.2%),而聚合指标会将其称为完美校准。我们提出了逐条件校准诊断(PCCD),这是一个三步协议,用于评估拒绝校准、上下文跟随和能力保持。PCCD 捕捉到非对称危害和聚合评估遗漏的琐碎拒绝者失败模式。
cs.CL / 7 / 2607.28640

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

TokenSwap:多模态大语言模型中的基准测试与模态差距的缩减
Hua, Andong, Bishop, Colton, Mordatch, Igor, Hosseini, Arian, Gu, Jindong, Faust, Aleksandra, Roelofs, Rebecca, Qin, Yao
Abstract
Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy in model predictions under such cross-modal variations. Specifically, we define the modality gap as the difference in model performance under semantically equivalent textual and multimodal inputs. We introduce TokenSwap, a method that constructs such inputs by replacing textual concepts with semantically aligned images, resulting in sequences where visual tokens are interleaved with text tokens. Based on TokenSwap, we transform existing text-based benchmarks such as MMLU into image-interleaved counterparts, resulting in TokenSwap-Bench. Across 42 MLLMs, we observe a pervasive modality gap, with performance decreasing by 4.2% to 47.4% when moving from text-only to image-interleaved inputs, averaging 19.6% +/- 3.3% across models. Notably, we observe that reasoning models exhibit consistently smaller gaps, achieving an average gap of 10.1% compared to 25.5% for non-reasoning models. In contrast, neither prompting strategies nor scaling training compute alone reliably reduces the modality gap. Finally, we demonstrate that incorporating TokenSwap during training effectively mitigates this gap while preserving strong text-only and vision-language performance.
Chinese Translation
多模态大语言模型(MLLMs)应在不同模态下对语义等价的输入生成一致的响应。然而,我们观察到在这种跨模态变化下模型预测存在系统性差异。具体而言,我们将模态差距定义为在语义等价的文本和多模态输入下模型性能的差异。我们提出了TokenSwap,一种通过用语义对齐的图像替换文本概念来构建此类输入的方法,从而生成文本标记与视觉标记交错的序列。基于TokenSwap,我们将现有的基于文本的基准测试(如MMLU)转化为图像交错的对应物,形成TokenSwap-Bench。在42个MLLMs中,我们观察到普遍存在模态差距,当从仅文本输入转向图像交错输入时,性能下降幅度在4.2%到47.4%之间,模型的平均下降幅度为19.6% ± 3.3%。值得注意的是,我们发现推理模型的差距始终较小,平均差距为10.1%,而非推理模型的平均差距为25.5%。相比之下,仅依靠提示策略或扩大训练计算量并不能可靠地减少模态差距。最后,我们证明在训练过程中引入TokenSwap能够有效缓解这一差距,同时保持强大的文本单一和视觉-语言性能。
cs.CL / 8 / 2607.28641

The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?

形式主义陷阱:在社会负荷下,LLM作为评判者的评估者是否被共识模仿所蒙蔽?
Shehata, Dahlia, Li, Ming
Abstract
We introduce the \textit{Agentic Formalism Trap} and the Evaluative Dissonance Index ($D_E$), quantifying how LLM-as-a-Judge systems conflate structural proceduralism with semantic truth under adversarial load. Analyzing 22,500 trajectories across 3 domains (GAIA, SWE-bench, Multi-Challenge), we extract a semantic taxonomy of hallucination maneuvers, validated via deterministic lexical grounding ($p < 10^{-120}$). A logistic meta-evaluator isolates the exact syntactic triggers of this evaluator capture (ROC-AUC 0.8779), while a zero-shot Leave-One-Domain-Out transfer proves the vulnerability is universally domain-agnostic (mean ROC-AUC 0.7482). Architectural profiling reveals that distinct simulated swarm topologies induce mathematically disparate semantic blind spots, proving that unanchored closed-loop evaluation is unstable, systemically divergent and necessitates architecture-specific vigilance filters.
Chinese Translation
我们引入了 extit{代理形式主义陷阱}和评估不和谐指数($D_E$),量化了在对抗性负荷下,LLM作为评判者系统如何将结构性程序主义与语义真理混为一谈。通过分析3个领域(GAIA、SWE-bench、Multi-Challenge)中的22,500个轨迹,我们提取了一种幻觉操作的语义分类法,通过确定性词汇基础验证($p < 10^{-120}$)。一个逻辑元评估器隔离了这种评估者捕获的确切句法触发因素(ROC-AUC 0.8779),而零样本的留一领域转移证明了这种脆弱性在各领域普遍存在(平均ROC-AUC 0.7482)。架构分析显示,不同的模拟群体拓扑会导致数学上不同的语义盲点,证明了未锚定的闭环评估是不稳定的,系统性偏离的,并且需要架构特定的警觉过滤器。
cs.CL / 9 / 2607.28658

Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation

评估联邦预训练:下游微调和内在评估的可靠性
Grosser, Claudia, Heuer, Maike, Krompass, Denis, Runkler, Thomas A.
Abstract
Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets. However, evaluating federated pre-training remains challenging because differences in client participation and local data availability can make directly comparable evaluation difficult. Moreover, pre-training test perplexity is tied to the pre-training distribution, while downstream benchmarks introduce task-specific adaptation that may not faithfully reflect the test perplexity established during pre-training. In this work, we study which evaluation protocol more reliably reflects federated pre-training quality. Using a controlled set of centralized and federated-trained models of a 16M parameter transformer model trained on identical client data, we assess evaluation protocols by whether they preserve a reference ranking established on the same pre-training testset. We compare downstream fine-tuning on GLUE, including full, head-only, and reduced-data variants, with next-token prediction on GLUE text as an intrinsic evaluation signal. Our results show that downstream fine-tuning does not reliably preserve the pre-training ranking, whereas direct next-token prediction exhibits a strong correspondence with the pre-training test perplexity. These findings suggest that downstream fine-tuning alone can be misleading when comparing federated pre-trained models, and that evaluation signals closer to the original pre-training objective deserve greater attention.
Chinese Translation
联邦预训练提供了一种在私有或分布式数据上训练基础模型的方法,而无需集中底层数据集。然而,评估联邦预训练仍然具有挑战性,因为客户端参与和本地数据可用性的差异使得直接可比的评估变得困难。此外,预训练测试困惑度与预训练分布相关,而下游基准引入的任务特定适应可能无法真实反映在预训练期间建立的测试困惑度。在本研究中,我们探讨了哪种评估协议更可靠地反映联邦预训练的质量。通过使用一组受控的集中和联邦训练模型(一个在相同客户端数据上训练的16M参数变换器模型),我们评估了评估协议是否保留了在相同预训练测试集上建立的参考排名。我们比较了在GLUE上的下游微调,包括完整、仅头部和减少数据变体,以及在GLUE文本上的下一个标记预测作为内在评估信号。我们的结果表明,下游微调并不能可靠地保留预训练排名,而直接的下一个标记预测与预训练测试困惑度之间存在强相关性。这些发现表明,仅依靠下游微调在比较联邦预训练模型时可能会产生误导,而更接近原始预训练目标的评估信号应受到更多关注。
cs.CL / 10 / 2607.28661

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

大型语言模型的金融推理是否可信?对长期陈述的现实世界测试
Tong, Xinke, Zhang, Xuanming, Tang, Tianyi, Yang, An, Hu, Jiatu, Lin, Guojie, Shi, Zhenzhen, Zeng, Lingfeng, Yang, Boyu, Zhao, Bing, Wei, Hu, Qu, Lin, Liu, Dayiheng
Abstract
Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and multi-step logic over long contexts, is an ideal testbed. Existing benchmarks fail to capture real-world industrial complexity, predominantly relying on multiple-choice questions or single-hop QA over cropped tables while ignoring intricate cross-statement dynamics and temporal de-cumulation. To bridge this gap, we introduce FinIndices, a large-scale benchmark evaluating data-processing fidelity over uncropped financial statements (up to 32K tokens). Utilizing an automated synthesis pipeline with adversarial traps, FinIndices encompasses Single-Index computation and Table-Index tabulation to test complex domain, temporal, and caliber reasoning. Our evaluation reveals two severe LLM vulnerabilities. First, a "Knowledge Bottleneck": despite memorizing formulas during pre-training, models demonstrate fragile pattern matching. Removing explicit formula hints causes performance to collapse (e.g., Gemini-3.1-Pro drops from 70.70% to 38.22% on table tasks), exposing fatal flaws in temporal de-cumulation and stock-flow caliber mismatch. Second, a "Structural Bottleneck": the intense cognitive load of generating multi-metric, multi-period tables actively drains reasoning capacity. Under structural pressure, LLMs that flawlessly execute isolated derivations regress to shallow heuristics, such as fetching incorrect adjacent columns or substituting deep accounting adjustments with lazy literal arithmetic. Finally, Supervised Fine-Tuning (SFT) yields substantial zero-hint gains (+8.54% Single, +3.82% Table), validating that structured logic can be partially restored via data-centric alignment.
Chinese Translation
大型语言模型(LLMs)是否具备真正的结构性推理能力,还是仅仅依赖于表层模式匹配?金融领域对数字精度和长上下文中的多步骤逻辑有着严格要求,因此是一个理想的测试平台。现有基准测试未能捕捉到现实世界的工业复杂性,主要依赖于多项选择题或针对裁剪表格的单步问答,同时忽略了复杂的跨陈述动态和时间去累积。为填补这一空白,我们引入了FinIndices,一个大规模基准,评估未裁剪财务报表(最多32K个标记)上的数据处理准确性。通过使用带有对抗陷阱的自动合成管道,FinIndices涵盖了单一指标计算和表格指标汇总,以测试复杂领域、时间和质量推理。我们的评估揭示了LLMs的两个严重脆弱性。首先是“知识瓶颈”:尽管在预训练期间记住了公式,模型在模式匹配上表现脆弱。去除显式公式提示会导致性能崩溃(例如,Gemini-3.1-Pro在表格任务上的表现从70.70%下降到38.22%),暴露出在时间去累积和库存流动质量不匹配方面的致命缺陷。其次是“结构瓶颈”:生成多指标、多周期表格的强烈认知负荷会消耗推理能力。在结构压力下,能够完美执行孤立推导的LLMs退化为浅层启发式,例如获取不正确的相邻列或用懒惰的字面算术替代深层会计调整。最后,监督微调(SFT)带来了显著的零提示增益(+8.54% 单一指标,+3.82% 表格),验证了通过数据中心对齐可以部分恢复结构逻辑。
cs.CL / 11 / 2607.28666

The Checking Problem: What must be true before AI ships in a regulated firm

检查问题:在受监管公司中,人工智能部署前必须满足的条件
Ahuja, Prerit
Abstract
Enterprise AI programmes stall at a rate that is widely quoted and poorly explained. This paper measures the mechanism. Six document-heavy workflows of the kind performed daily in regulated financial services were run across four model families and three tool configurations, three times each, producing 5,093 scored output elements across 72 configurations. Each configuration was assessed twice: against a demonstration bar, being a single correct run on a single case, and against a production bar requiring sustained accuracy, reproducibility across repeats, verifiable attribution, and a confidence signal that carries information. 57 of 72 configurations cleared the demonstration bar and 32 cleared the production bar, a survival rate of 56.1%. The paper then computes the review burden each configuration imposes, estimated out of sample rather than with hindsight. A tool that states no confidence requires review of 100% of its output, because it offers a reviewer no basis for triage. Requiring the tool to cite its sources and state a confidence reduces that to 49% while holding the residual error tolerance in 17 of 20 configurations. Adding a self-verification pass costs 2.3 times the latency of the plain configuration, reaches 44%, and is the only configuration that fails to hold the error tolerance. The practical implication is that the value of an AI workflow is set less by how often it is right than by how much of it a human must still check, and that the second property is measurable and rarely measured.
Chinese Translation
企业人工智能项目停滞的比例被广泛引用但解释不清。本文测量了这一机制。对在受监管金融服务中每日执行的六种文档密集型工作流程进行了测试,涵盖四个模型系列和三种工具配置,每种配置运行三次,产生了72种配置下的5,093个评分输出元素。每种配置进行了两次评估:一次是与演示标准进行对比,该标准为单个案例的单次正确运行;另一次是与生产标准进行对比,生产标准要求持续的准确性、可重复性、可验证的归属以及携带信息的信心信号。在72种配置中,有57种通过了演示标准,32种通过了生产标准,生存率为56.1%。随后,本文计算了每种配置所施加的审查负担,估算是基于样本外而非事后分析。一个声明无信心的工具需要对其100%的输出进行审查,因为它未能为审查者提供分流的基础。要求工具引用其来源并声明信心将这一比例降低至49%,同时在20种配置中保持17种的剩余误差容忍度。增加自我验证步骤的成本是普通配置延迟的2.3倍,达到了44%,并且是唯一未能保持误差容忍度的配置。实际意义在于,人工智能工作流程的价值并不在于其正确的频率,而在于人类仍需检查的部分,而这一第二个属性是可测量的,但很少被测量。
cs.CL / 12 / 2607.28680

TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking

TELLER:双路径迭代偏好优化用于表格实体链接
Peng, Yixin, Li, Kehao, Decker, Stefan
Abstract
Entity linking in tables matches short and ambiguous cell mentions to their corresponding knowledge-base entities. Existing approaches typically rely on data preprocessing pipelines that retain either compact or extensive table content as contextual evidence, and then formulate entity linking as a language generation task for instruction-tuned models; recent systems further incorporate explicit reasoning to disambiguate challenging mentions. However, their training supervision is usually static: fixed preference data cannot adapt to the residual errors of an evolving model, while variations in reasoning length can bias sequence-level preference learning. To address these limitations, we present TELLER: Table Entity Linking through Learning from Errors and Reasoning. We first retrieve and rank Wikidata candidates and retain reduced table evidence in the prompt. The direct-answer path applies iterative direct preference optimization and refreshes its preference data with residual errors from the updated model. The reasoning path uses filtered and compressed chain-of-thought rationales for supervised fine-tuning, followed by our iterative length-normalized regularized preference optimization. On the TableInstruct entity-linking subset, the direct-answer path improves accuracy from 94.35\% to 94.50\%; on the MammoTab V2 evaluation set, it improves accuracy from 87.59\% to 88.20\%. The reasoning path improves accuracy from 92.90\% to 92.95\% on TableInstruct and from 79.09\% to 81.85\% on MammoTab V2, while maintaining high rates of complete reasoning generation. These results show that iterative preference learning benefits both concise entity prediction and explicit reasoning.
Chinese Translation
表格中的实体链接将简短且模糊的单元格提及与其对应的知识库实体进行匹配。现有的方法通常依赖于数据预处理管道,这些管道保留紧凑或广泛的表格内容作为上下文证据,然后将实体链接表述为针对指令调优模型的语言生成任务;最近的系统进一步结合了显式推理以消除具有挑战性的提及的歧义。然而,它们的训练监督通常是静态的:固定的偏好数据无法适应不断演变模型的残差错误,而推理长度的变化可能会影响序列级偏好学习。为了解决这些局限性,我们提出了TELLER:通过学习错误和推理进行表格实体链接。我们首先检索和排名Wikidata候选项,并在提示中保留减少的表格证据。直接回答路径应用迭代直接偏好优化,并用更新模型的残差错误刷新其偏好数据。推理路径使用过滤和压缩的思维链推理作为监督微调,随后进行我们的迭代长度归一化正则化偏好优化。在TableInstruct实体链接子集上,直接回答路径的准确率从94.35\%提高到94.50\%;在MammoTab V2评估集上,准确率从87.59\%提高到88.20\%。推理路径在TableInstruct上的准确率从92.90\%提高到92.95\%,在MammoTab V2上的准确率从79.09\%提高到81.85\%,同时保持高比例的完整推理生成。这些结果表明,迭代偏好学习有利于简洁的实体预测和显式推理。
cs.CL / 13 / 2607.28707

Demystifying Entropy-based Selection for Chain-of-Thought Compression in Large Reasoning Models

揭示基于熵的选择在大型推理模型中的思维链压缩中的作用
Candussio, Sara, Scalena, Daniel, Bortolussi, Luca, Fersini, Elisabetta, Nissim, Malvina, Sarti, Gabriele
Abstract
Entropy-based pruning has been proposed as an effective method for compressing Chain-of-Thought (CoT) reasoning with negligible accuracy loss. We test the robustness of low- and high-entropy CoT step selection methods across various models and reasoning tasks, showing that entropy offers no advantage over random pruning in any evaluated setting. Moving from sentences to tokens, we then show that retaining low-entropy tokens seems effective only on mathematical benchmarks. We find this is due to the inherently low-entropy nature of numeric tokens, which also convey semantic content in such problems. Finally, we demonstrate that patching a subset of a few CoT tokens with their original activations recovers near-perfect full-trace performance, providing causal evidence that task information is not concentrated in a small set of CoT tokens identifiable by heuristics, but rather distributed across the full reasoning chain.
Chinese Translation
基于熵的剪枝被提出作为一种有效的方法,用于在几乎没有准确性损失的情况下压缩思维链(Chain-of-Thought, CoT)推理。我们测试了低熵和高熵 CoT 步骤选择方法在各种模型和推理任务中的鲁棒性,结果表明,在任何评估的设置中,熵对随机剪枝没有优势。接着,我们从句子转向标记,显示保留低熵标记似乎仅在数学基准测试中有效。我们发现这是由于数字标记本身具有低熵特性,并且在此类问题中也传达了语义内容。最后,我们展示了用原始激活修补少量 CoT 标记的子集可以恢复近乎完美的完整追踪性能,提供了因果证据,表明任务信息并不集中在一小组可通过启发式方法识别的 CoT 标记中,而是分布在整个推理链中。
cs.CL / 14 / 2607.28766

The Morphological Core of Dungan: A Two-Dialect Finite-State Model and a Multi-Genre Evaluation

东干语的形态核心:一种双方言有限状态模型及多体裁评估
Alekseev, Anton M., Nikolenko, Sergey I.
Abstract
Dungan, a Sinitic language of Central Asia written in a Cyrillic-based script, is described in detail in the grammatical literature, yet the quantitative properties of its morphology in actual usage have, to the best of our knowledge, never been measured systematically. This paper uses a finite-state morphological analyzer as a measuring instrument. Implemented with HFST and covering both dialect groups (the Gansu variety, which is the literary standard, and the Shaanxi variety), the model offers no new grammatical description; it formalises the knowledge accumulated in Dungan studies and makes it measurable on corpora of three genres. Three results follow. Overt inflection is rare and limited: only 9.3% of recognized tokens in the encyclopaedic register have an overt marker, the system has just ten categories, and degree marking is almost absent. Ambiguity is genuine but sharply localized: 78.1% of tokens receive a single analysis, and the residue sits almost entirely on two clitics, -di (genitive/progressive) and -ni (locative/prospective). And the grammatical core proves effectively closed, the claim the instrument is really needed for: between 78% and 95% of the tokens the analyzer fails on, depending on register, are simply absent from the lexicon, and the phenomena the model deliberately declines to implement account for at most 4.5% of those failures. Held-out coverage (80 to 85%) is no lower than development coverage (73%), while a stem list with no morphology already reaches 67.4%, so the morphology is worth 5.2 points. The open frontier of Dungan is lexical. The analyzer, its sources and every evaluation script are released openly.
Chinese Translation
东干语是一种使用基于西里尔字母的书写系统的中亚汉语,其语法文献中有详细描述,但据我们所知,其形态在实际使用中的定量特性从未被系统测量过。本文使用有限状态形态分析器作为测量工具。该模型采用HFST实现,涵盖两个方言群体(甘肃方言,即文学标准,以及陕西方言),并未提供新的语法描述;而是将东干研究中积累的知识形式化,并使其在三种体裁的语料库上可测量。结果如下:显性屈折形式稀少且有限:在百科全书体裁中,仅有9.3%的识别标记具有显性标记,系统仅有十个类别,程度标记几乎缺失。歧义是真实存在的,但高度局限:78.1%的标记只接受单一分析,其余几乎完全集中在两个附着词上,-di(属格/进行时)和-ni(处所/将来时)。而且,语法核心证明是有效封闭的,这正是该工具真正需要的:在分析器未能处理的标记中,78%至95%(取决于语体)在词汇表中根本不存在,而模型故意不实现的现象最多只占这些失败的4.5%。保留覆盖率(80%至85%)不低于开发覆盖率(73%),而没有形态的词干列表已经达到了67.4%,因此形态的价值为5.2分。东干语的开放边界是词汇层面的。分析器、其来源及所有评估脚本均已公开发布。
cs.CL / 15 / 2607.28777

Self-Supervised Skill Optimization

自监督技能优化
Peng, Siran, Yang, Cuiyu, Fu, Tianyu, Zhang, Tianshuo, Zhang, Haoyuan, Zhao, Weisong, Su, Anyang, Wu, Minghui, Li, Huiying, Zhu, Xiangyu, Zhao, Chenxu, Lei, Zhen
Abstract
Agent skills provide frozen large language model (LLM) agents with reusable procedural guidance, and recent work shows that such skills can be optimized with ground-truth (GT) feedback. Many applications, however, lack GT labels, task scores, rewards, or reliable task-specific evaluators. We therefore introduce Self-Supervised Skill Optimization (SSO), a comparative framework that learns a reusable skill from unlabeled task instances alone. At each step, SSO runs the current skill on an unlabeled batch, uses a subset of the resulting executions to generate complete skill probes, and runs the probes on the same batch. An LLM judge compares the resulting answers, trajectories, artifacts, or terminal states. A separate behavior extractor identifies behavioral differences without seeing the judge's decisions. SSO uses these decisions to aggregate evidence for and against the observed behaviors across instances. It then ranks the behaviors by the resulting evidence and renders a new complete skill from the highest-ranked behaviors. The update is accepted only if the new skill outperforms the current one on an unlabeled validation set. SSO outperforms existing GT-free prompt optimizers on both closed-ended and open-ended tasks. On closed-ended benchmarks, it approaches and sometimes exceeds the strongest GT-based skill optimizer without using any GT feedback.
Chinese Translation
代理技能为冻结的大型语言模型(LLM)代理提供可重用的程序指导,近期研究表明,这些技能可以通过真实反馈(GT)进行优化。然而,许多应用缺乏GT标签、任务得分、奖励或可靠的任务特定评估者。因此,我们提出了自监督技能优化(SSO),这是一个比较框架,仅从未标记的任务实例中学习可重用的技能。在每一步中,SSO在一个未标记的批次上运行当前技能,使用结果执行中的一个子集生成完整的技能探针,并在同一批次上运行这些探针。一个LLM评审比较生成的答案、轨迹、工件或终态。一个独立的行为提取器在不查看评审决定的情况下识别行为差异。SSO利用这些决定聚合对观察到的行为的支持和反对证据。然后,它根据结果证据对行为进行排名,并从排名最高的行为中生成新的完整技能。只有当新技能在未标记的验证集上优于当前技能时,更新才会被接受。SSO在封闭式和开放式任务上均优于现有的无GT提示优化器。在封闭式基准测试中,它接近并有时超过最强的基于GT的技能优化器,而不使用任何GT反馈。
cs.CL / 16 / 2607.28801

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

基准测试并非单一:针对大型语言模型评估的样本级审计与编排
Siedler, Philipp D., Sassoon, Jordan
Abstract
Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of individual samples. We introduce a dataset-centric meta-evaluation framework that audits benchmark datasets at the sample level along five latent dimensions: 1. Cognitive and Knowledge Demands, 2. Language and Content Quality, 3. Task Properties, 4. Context, and 5. Ethics, Safety, and Fairness. Applying this framework, we annotate five influential benchmarks -- MMLU, ARC, WinoGrande, HellaSwag, and TruthfulQA -- revealing pronounced internal heterogeneity that is not captured by aggregate accuracy scores. We show how these annotations enable criterion-driven orchestration of composite benchmark subsets across datasets, supporting targeted evaluation of model capabilities such as Reasoning Depth or Ethical Sensitivity. This approach reframes benchmark evaluation as dataset introspection, providing a principled methodology for analyzing and re-composing existing benchmarks to better reflect diverse evaluation needs.
Chinese Translation
基准数据集在评估大型语言模型(LLMs)中至关重要,但它们通常被视为单一任务,掩盖了个别样本需求的显著差异。我们提出了一种以数据集为中心的元评估框架,该框架从五个潜在维度对基准数据集进行样本级审计:1. 认知和知识需求,2. 语言和内容质量,3. 任务属性,4. 上下文,以及 5. 伦理、安全和公平性。应用该框架,我们对五个有影响力的基准进行了注释——MMLU、ARC、WinoGrande、HellaSwag 和 TruthfulQA,揭示了显著的内部异质性,这一特征并未通过整体准确率得以捕捉。我们展示了这些注释如何支持跨数据集的复合基准子集的标准驱动编排,从而支持对模型能力(如推理深度或伦理敏感性)的有针对性评估。这种方法将基准评估重新框定为数据集内省,为分析和重新组合现有基准提供了一种原则性的方法论,以更好地反映多样化的评估需求。
cs.CL / 17 / 2607.28814

Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing

应对阻力:偏好优化的LLM顾问在动机访谈中可以在目标坚持与关系协调之间进行权衡
Chen, Weiying, Shen, Junlong, Tang, Zhexuan
Abstract
In Motivational Interviewing (MI), a client's sustain talk (arguments for the status quo) calls for the counselor to roll with resistance, a move that can fail in two opposite ways: capitulation (abandoning the change agenda to preserve rapport) or confrontation (arguing or directing, overriding the client's autonomy). We introduce a two-axis evaluation of counselor responses, anchored in the Motivational Interviewing Treatment Integrity (MITI) code, Goal Persistence (GP) and Relational Attunement (RA), yielding a four-quadrant framing in which rolling with resistance is high on both, and we ask whether penalizing one failure through preference optimization teaches rolling with resistance or provokes its opposite. From the expert-annotated AnnoMI corpus we build topic-disjoint Direct Preference Optimization data whose preference sets differ only in which failure is rejected, using on-policy negatives. An automatic judge, validated against AnnoMI's expert labels and rechecked by trained human coders, scores blind pairwise win-rates against each base under a firewall in which disjoint model families generate, label, and judge. Across three aligned instruction models spanning the Qwen and Llama families, penalizing confrontation reliably lowers goal persistence below parity, on every base and in every seed run, a robust cost, whereas the attunement gain is base-dependent, present on two of the three bases but absent on the third. Penalizing capitulation is inert, because these models rarely capitulate on-policy, so the trade is gated by each base's failure profile. A prompt-only control raises attunement without the goal-persistence cost, locating the cost in the optimization rather than in attunement itself.
Chinese Translation
在动机访谈(MI)中,客户的持续发言(支持现状的论据)要求顾问应对阻力,这一举措可能以两种相反的方式失败:屈服(放弃变革议程以维持融洽关系)或对抗(争论或指令,侵犯客户的自主权)。我们引入了一个基于动机访谈治疗完整性(MITI)标准的顾问反应的双轴评估,分别为目标坚持(GP)和关系协调(RA),形成一个四象限框架,其中应对阻力在这两个维度上均表现良好。我们探讨通过偏好优化惩罚某一失败是否能够教会应对阻力,或是引发其相反的结果。我们基于专家标注的AnnoMI语料库构建了主题不重叠的直接偏好优化数据,其偏好集仅在拒绝哪种失败上存在差异,使用政策内负样本。一个自动评判者经过AnnoMI的专家标签验证,并由经过培训的人类编码员重新检查,在一个防火墙下进行盲对比评分,生成、标注和评判的模型家族是互不重叠的。在跨越Qwen和Llama家族的三个对齐指令模型中,惩罚对抗始终可靠地将目标坚持降低到低于平衡水平,在每个基础和每个种子运行中都是一个稳健的成本,而协调的收益则依赖于基础,在三个基础中的两个上存在,但在第三个上缺失。惩罚屈服则无效,因为这些模型在政策内很少屈服,因此这种权衡受到每个基础失败特征的限制。仅使用提示的控制组在没有目标坚持成本的情况下提高了协调性,将成本定位于优化而非协调本身。
cs.CL / 18 / 2607.28840

Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications

基准测试不是验证:金融大语言模型应用的系统层面视角
Payzun, Burak, Demirtaş, İrem, Scala, Simona, Ferretti, Elena, Arslan, Seçil
Abstract
Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monitoring, and human escalation. Yet evaluation often remains model-centric: benchmark scores, task accuracy, or one-off qualitative reviews are treated as evidence of readiness. In financial settings, this is insufficient. We take the position that financial LLM systems should not be approved for production based on benchmark performance alone. They require system-level validation evidence across the application stack: data, model design, retrieval and generation performance, agent behavior, governance, and implementation. Drawing on industry experience validating GenAI applications in financial institutions, we outline a multi-layer validation view and explain why hybrid evaluation is necessary. We discuss where LLM-as-a-judge methods are useful and why they require controls such as multiple judges, rubrics, agreement, and auditability checks. We also highlight failure modes poorly captured by static benchmarks, including retrieval failures, unfaithful generation, tool misuse, escalation errors, and operational instability. Our position is that financial LLM validation should be an ongoing system discipline rather than a one-time model scoring exercise. Validation should produce decision-ready evidence, not only scores. We conclude with a research agenda for system-aware benchmarks, agent trace validation, judge alignment protocols, and lifecycle validation standards.
Chinese Translation
大型语言模型在金融应用中的部署日益增多,这些应用结合了检索、专有数据、工具使用、调度逻辑、监控和人工升级。然而,评估往往仍然以模型为中心:基准分数、任务准确性或一次性的定性评审被视为准备就绪的证据。在金融环境中,这种做法是不够的。我们认为,金融大语言模型系统不应仅基于基准性能就被批准投入生产。它们需要在应用堆栈的各个层面上提供系统级的验证证据:数据、模型设计、检索和生成性能、代理行为、治理和实施。借鉴在金融机构验证生成式人工智能应用的行业经验,我们概述了多层次的验证视角,并解释了为何混合评估是必要的。我们讨论了大语言模型作为评判者的方法在何处是有用的,以及为什么它们需要控制措施,如多个评审者、评分标准、一致性和可审计性检查。我们还强调了静态基准测试难以捕捉的失败模式,包括检索失败、不忠实生成、工具误用、升级错误和操作不稳定。我们的立场是,金融大语言模型的验证应当是一个持续的系统学科,而不是一次性的模型评分练习。验证应当产生决策准备的证据,而不仅仅是分数。最后,我们提出了一个研究议程,涉及系统感知基准、代理追踪验证、评审者对齐协议和生命周期验证标准。
cs.CL / 19 / 2607.28862

TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text

TextCloak:通过基于强化学习的不可学习文本来抵御未经授权的LLM利用
Zhao, Chengshuai, Ma, Pingchuan, Li, Dawei, Jiang, Bohan, Yu, Zhiyuan, Tan, Zhen, Liu, Huan
Abstract
The rapid development of Large Language Models (LLMs) has led to significant advances across a wide range of language tasks, while simultaneously raising growing concerns about unauthorized data exploitation and privacy leakage. Unlearnable examples (UEs) offer a promising defense by introducing carefully designed perturbations into data such that models trained on them exhibit degraded utility. However, existing methods for text protection are primarily designed for classification tasks (e.g., sentiment analysis) in discriminative language models and often rely on injecting class-specific linguistic cues, which limits their effectiveness in the open-ended generation settings of LLMs. In this work, we propose TextCloak, an RL-driven framework for protecting textual data against unauthorized LLM exploitation. TextCloak employs a generative policy that transforms batches of clean text into unlearnable examples while preserving semantic fidelity and linguistic naturalness. To optimize the policy, we introduce GRPO-UE, which rewards generated unlearnable text based on the downstream degradation they induce in fine-tuned surrogate LLMs and updates the generator parameters via group-relative policy optimization. This bi-level optimization enables the generator to discover generalizable protective patterns beyond class-specific cues. Comprehensive experiments on six publicly available datasets and nine state-of-the-art LLMs demonstrate that TextCloak consistently impairs unauthorized fine-tuning while maintaining text utility for legitimate use. Further analyses establish its transferability and robustness across model architectures, training configurations, and adaptive attacks, highlighting its broad applicability as a practical defense against unauthorized LLM exploitation.
Chinese Translation
大型语言模型(LLMs)的快速发展在广泛的语言任务中带来了显著的进步,同时也引发了对未经授权的数据利用和隐私泄露的日益关注。不可学习示例(UEs)通过在数据中引入精心设计的扰动,提供了一种有前景的防御方法,使得在这些数据上训练的模型表现出效用下降。然而,现有的文本保护方法主要针对判别性语言模型中的分类任务(例如情感分析),通常依赖于注入特定类别的语言线索,这限制了它们在LLMs开放式生成环境中的有效性。在本研究中,我们提出了TextCloak,一个基于强化学习的框架,用于保护文本数据免受未经授权的LLM利用。TextCloak采用生成策略,将干净文本批次转换为不可学习示例,同时保持语义的准确性和语言的自然性。为了优化该策略,我们引入了GRPO-UE,该方法根据生成的不可学习文本在微调的替代LLMs中引起的下游效用下降进行奖励,并通过组相对策略优化更新生成器参数。这种双层优化使生成器能够发现超越特定类别线索的可推广保护模式。在六个公开可用的数据集和九个最先进的LLMs上的全面实验表明,TextCloak在保持文本效用以供合法使用的同时,始终削弱未经授权的微调。此外,进一步的分析证明了其在模型架构、训练配置和自适应攻击中的可转移性和鲁棒性,突显了其作为抵御未经授权的LLM利用的实用防御的广泛适用性。
cs.CL / 20 / 2607.28906

Token-Level Diagnosis of Sycophancy in LLMs with Attribution-Guided Steering

基于归因引导的引导下对大型语言模型中的阿谀奉承进行令牌级诊断
Nguyen, Hieu, Kamruzzaman, Mahammed, Chhabra, Anshuman, Kim, Gene Louis
Abstract
Sycophancy refers to the tendency for large language models (LLMs) to match user beliefs at the cost of factual correctness, thereby undermining model reliability. Prior work on evaluating sycophancy in LLMs aims to assess whether a model's output matches an authority's claim, but cannot reveal which part of the prompt drives this sycophantic behavior. To bridge this gap, we investigate the relationship of sycophantic responses with an authority's credentials, their assertive claim, and the problem statement. We introduce the Authority Share Index (ASI), an Integrated Gradients-based token attribution method, which measures the degree to which a model's decision is driven by authority-related text. Through extensive experiments across five models and 30 test configurations, we find that sycophantic responses consistently direct more attention toward authority tokens than resistant ones. Moreover, our token attribution method reveals that for the sycophantic cases, the claim asserted by the authority receives more attention than the authority's credentials. Building on these findings, we propose attribution-guided contrastive activation steering to mitigate LLM sycophancy. Our method constructs a steering vector from high-attribution tokens of sycophantic and resistant responses, selectively pushing models toward resistance. This enables inference-time steering without retraining, lowering sycophancy from 96% to 25% in the strongest case. Together, our results show that token-level attribution can both explain what drives sycophancy and directly inform a practical intervention.
Chinese Translation
阿谀奉承是指大型语言模型(LLMs)倾向于迎合用户信念,而牺牲事实正确性,从而削弱模型的可靠性。先前关于评估LLMs中阿谀奉承的研究旨在评估模型输出是否与权威声明一致,但无法揭示提示的哪个部分驱动了这种阿谀奉承行为。为填补这一空白,我们研究了阿谀奉承响应与权威资质、其断言声明和问题陈述之间的关系。我们引入了权威分享指数(Authority Share Index, ASI),这是一种基于集成梯度的令牌归因方法,用于衡量模型决策在多大程度上受到与权威相关文本的驱动。通过对五个模型和30个测试配置进行广泛实验,我们发现阿谀奉承响应始终比抵抗性响应更关注权威令牌。此外,我们的令牌归因方法揭示,在阿谀奉承的情况下,权威所声称的观点比权威的资质获得更多关注。基于这些发现,我们提出了归因引导的对比激活引导方法,以减轻LLMs的阿谀奉承。我们的方法从阿谀奉承和抵抗性响应的高归因令牌构建引导向量,选择性地推动模型朝向抵抗。这使得在推理时进行引导而无需重新训练,在最强情况下将阿谀奉承从96%降低到25%。总的来说,我们的结果表明,令牌级归因既可以解释阿谀奉承的驱动因素,也可以直接为实际干预提供信息。
cs.CL / 21 / 2607.28934

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

FairFund-Bench:评估大型语言模型资源分配中的分配偏见
Lukk, Martin
Abstract
Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need derived from welfare deservingness theory. Across 14 models, audit format changes the direction of bias: models advantage minorities when rating claimants individually but penalize some groups when ranking them side by side. Bias magnitude, though small overall, is several times greater in disguised audits than in transparent ones, where, faced with appeals differing only in claimants' names, models overwhelmingly split funds equally. Causal framing effects, by contrast, exceed demographic effects by roughly an order of magnitude and are consistent across models and audit formats, indicating that current LLMs robustly reproduce human deservingness evaluations. The benchmark scores models on four criteria (demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency), is publicly available, and can be readily adapted to other substantive domains.
Chinese Translation
大型语言模型(LLMs)在稀缺资源分配中日益发挥作用,这引发了对基于种族和性别等特征的偏见分配的担忧。然而,近期的LLM审计结果不一致,发现对女性和少数族裔的正面和负面歧视的证据,甚至在相同模型中也存在这种情况。我们展示了这种分歧可能源于审计格式的不同,并引入了FairFund-Bench,一个系统性变化以往审计设计关键特征的基准:评估任务(评分、排名或分配)、比较背景(单一或多刺激)以及审计是否透明或伪装。该基准包含600个财务援助请求,这些请求是根据人类创作的模板创建的(与130万个真实的GoFundMe活动进行校准),涵盖三个领域、四个种族和两个性别类别,以及基于福利应得理论的五种需求因果框架。在14个模型中,审计格式改变了偏见的方向:当单独评分索赔者时,模型对少数群体有利,但在并排排名时却惩罚某些群体。尽管总体偏见幅度较小,但在伪装审计中,偏见幅度是透明审计的几倍,在透明审计中,当面对仅在索赔者姓名上有所不同的请求时,模型几乎总是平等分配资金。相比之下,因果框架效应的影响超过了人口统计效应,约为一个数量级,并且在模型和审计格式之间保持一致,表明当前的LLM稳健地再现了人类的应得评估。该基准根据四个标准(人口统计偏见、应得性一致性、跨任务一致性和跨背景一致性)对模型进行评分,公开可用,并且可以方便地适应其他实质性领域。
cs.CL / 22 / 2607.28966

BLADE: Boundary-Expanded and Layer-Adaptive Dynamic Exit for Efficient LLM Reasoning

BLADE:边界扩展与层自适应动态退出以提高大型语言模型推理效率
Fu, Keshu, Peng, Keqin, Bai, Jun, Qin, Shuhan, Li, Chen, Liang, Junzhu, Chen, Yefei, Li, Jiaqi, Ouyang, Yuanxin
Abstract
Large language models often improve task performance by generating long reasoning traces, but the resulting computation is frequently wasted on redundant verification and revision. Existing probe-based early-exit approaches mainly inspect explicit self-doubt expressions, leaving many earlier termination opportunities undetected. Expanding inspection to ordinary reasoning boundaries improves coverage, but also exposes highly diverse intermediate states whose predictive information may reside in different hidden layers. We present Boundary-Expanded and Layer-Adaptive Dynamic Exit for Efficient LLM Reasoning (BLADE), a lightweight framework that dynamically terminates reasoning by estimating whether the generated prefix is sufficient for correct answering. BLADE constructs multi-granular checkpoints from sentence, self-doubt, and paragraph boundaries, and derives robust training labels through repeated answer completions. It further learns a compact subset of informative probe layers instead of relying on fixed choices or expensive representations from all layers. At inference time, calibrated predictions are combined with checkpoint-specific confirmation rules to balance responsiveness and premature-exit risk. Experiments on five benchmarks and two Qwen3 reasoning models show that BLADE preserves near-baseline accuracy while reducing generated tokens by 24.8% on Qwen3-8B and 15.8% on Qwen3-4B. Ablation studies further confirm the benefits of diverse checkpoints and automatic layer selection, demonstrating an effective approach to more efficient LLM reasoning.
Chinese Translation
大型语言模型通常通过生成长推理轨迹来提高任务性能,但由此产生的计算往往在冗余的验证和修正上浪费。现有的基于探测的早期退出方法主要检查显式的自我怀疑表达,导致许多早期终止机会未被发现。将检查范围扩展到普通推理边界可以提高覆盖率,但也暴露出高度多样化的中间状态,其预测信息可能位于不同的隐藏层。我们提出了边界扩展与层自适应动态退出以提高大型语言模型推理效率(BLADE),这是一个轻量级框架,通过估计生成的前缀是否足以正确回答来动态终止推理。BLADE 从句子、自我怀疑和段落边界构建多粒度检查点,并通过重复答案补全推导出稳健的训练标签。它进一步学习一个紧凑的信息探测层子集,而不是依赖于固定选择或来自所有层的昂贵表示。在推理时,经过校准的预测与特定检查点的确认规则相结合,以平衡响应性和过早退出风险。在五个基准和两个 Qwen3 推理模型上的实验表明,BLADE 在保持接近基线准确率的同时,减少了 Qwen3-8B 上生成的标记数量 24.8%,以及 Qwen3-4B 上的 15.8%。消融研究进一步确认了多样化检查点和自动层选择的好处,展示了一种更高效的 LLM 推理的有效方法。
cs.CL / 23 / 2607.28979

Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models

翻译混合体:跨异构大语言模型翻译KV缓存
Lee, Jin-woo, Song, Minkyung, Oh, Junghyun, Han, Seunghoon, Park, Soyoung, Jang, Gwangseon, Lim, Sungsu
Abstract
Heterogeneous Large Language Model (LLM) systems increasingly rely on shared contexts, retrieved evidence, and multi-agent dialogue histories, yet their internal key-value (KV) caches remain model-specific and cannot be reused across architectures. Consequently, each model must repeatedly prefill or store caches for the same context, limiting the scalability of multi-model reasoning and long-context generation. We propose Mixture-of-Translators(MoT), a cache translation framework that maps context KV caches from a source LLM into the cache space of a target LLM. Unlike prior approaches that depend on a single projection path or global shared latent space, MoT uses multiple translator modules to capture diverse source--target mappings. To further reduce residual translation error, we introduce a Context Correction Loss that aligns the replayed target trajectory with the native target trajectory. We reveal two competing failure modes in cache translation: propagated translation shift from early injection and last-state shift from late injection. MoT addresses them through translator mixtures and target-side correction. Across homogeneous and heterogeneous translations among Qwen2.5, GPT-2, and OPT models, MoT preserves downstream QA performance, including Qwen2.5-7B-scale translation with 51.0% average closed-set QA accuracy and 0.43 average extractive QA F1. In practical case studies, MoT enables quality-preserving memory reuse for multi-agent reasoning and retains 96.3% of direct-context quality in long-context cache-augmented generation, demonstrating scalable KV cache reuse across heterogeneous LLMs.
Chinese Translation
异构大语言模型(LLM)系统日益依赖共享上下文、检索证据和多智能体对话历史,然而它们内部的键值(KV)缓存仍然是模型特定的,无法在不同架构之间重用。因此,每个模型必须重复预填充或存储相同上下文的缓存,这限制了多模型推理和长上下文生成的可扩展性。我们提出了翻译混合体(Mixture-of-Translators,MoT),这是一种缓存翻译框架,将源LLM的上下文KV缓存映射到目标LLM的缓存空间。与依赖单一投影路径或全局共享潜在空间的先前方法不同,MoT使用多个翻译模块来捕捉多样的源-目标映射。为了进一步减少残余翻译误差,我们引入了一种上下文校正损失(Context Correction Loss),使重放的目标轨迹与原生目标轨迹对齐。我们揭示了缓存翻译中的两种竞争失败模式:早期注入导致的传播翻译偏移和晚期注入导致的最后状态偏移。MoT通过翻译器混合和目标侧校正来解决这些问题。在Qwen2.5、GPT-2和OPT模型之间的同质和异质翻译中,MoT保持了下游问答性能,包括Qwen2.5-7B规模翻译的平均封闭集问答准确率为51.0%和平均抽取式问答F1为0.43。在实际案例研究中,MoT实现了多智能体推理的质量保持内存重用,并在长上下文缓存增强生成中保留了96.3%的直接上下文质量,展示了跨异构LLM的可扩展KV缓存重用。
cs.CL / 24 / 2607.28982

PARALLEL: A Prefrontal-Aligned Reinforcement inspired Approach for Language-Model Learning under Explicit Limits

PARALLEL:一种基于前额叶对齐的强化学习方法,用于在显式限制下的语言模型学习
Yoon, Namkyung, Kim, Sanghong, Kim, Hwangnam
Abstract
Recent language models achieve strong performance across a variety of tasks, but conventional adaptation applies updates uniformly across training samples regardless of their local update benefit. We propose PARALLEL, a prefrontal-aligned reinforcement inspired approach for language-model learning. Inspired by the complementary roles of goal-related and uncertainty-related control, PARALLEL represents these forms of information as separate controller signals and combines them with the current model representation. A reinforcement-inspired controller assigns sample-dependent update intensity using immediate utility-cost feedback. PARALLEL therefore learns when and how strongly to adapt to each sample, prioritizing beneficial updates while limiting unnecessary parameter changes. PARALLEL uses available updates more efficiently than selective baselines while retaining 94.1--99.2\% of Full-adaptation performance. Beyond multiple-choice reasoning, experiments on XSum and CNN/DailyMail show that PARALLEL retains 96.9--98.6\% of the ROUGE-1 and ROUGE-2 scores achieved by Full adaptation and 98.8--98.9\% of the corresponding ROUGE-L scores. When compared at the same cumulative adaptation time or GPU energy, PARALLEL achieves higher ARC accuracy and exhibits a more stable late-stage adaptation trajectory than Full adaptation in the representative run. These results show that learning when and how strongly to update each sample supports stable and efficient post-deployment stream adaptation while avoiding unnecessary updates.
Chinese Translation
近期的语言模型在多种任务中表现出色,但传统的适应方法对训练样本的更新是均匀的,而不考虑它们的局部更新效益。我们提出了PARALLEL,一种基于前额叶对齐的强化学习方法,用于语言模型学习。PARALLEL受到目标相关控制和不确定性相关控制互补作用的启发,将这些信息形式表示为独立的控制信号,并与当前模型表示相结合。一个基于强化学习的控制器使用即时效用-成本反馈分配样本依赖的更新强度。因此,PARALLEL能够学习何时以及多强烈地适应每个样本,优先考虑有益的更新,同时限制不必要的参数变化。PARALLEL比选择性基线更有效地利用可用更新,同时保留94.1%至99.2%的完全适应性能。在多项选择推理之外,针对XSum和CNN/DailyMail的实验表明,PARALLEL保留了完全适应所达到的96.9%至98.6%的ROUGE-1和ROUGE-2分数,以及98.8%至98.9%的相应ROUGE-L分数。在相同的累计适应时间或GPU能量下,PARALLEL在代表性运行中实现了更高的ARC准确率,并表现出比完全适应更稳定的后期适应轨迹。这些结果表明,学习何时以及多强烈地更新每个样本有助于稳定和高效的后期部署流适应,同时避免不必要的更新。
cs.CL / 25 / 2607.29044

From Inline Notes to Collected Commentaries: Toward Context-Preserving Organization of Exegetical Knowledge in Classical Chinese Texts

从行内注释到汇编评论:朝着保留上下文的经典中文文本释义知识组织
Liang, Ke, Su, Qi, Huang, Churen
Abstract
Inline notes and collected commentaries are important forms of scholarly communication that evolved within the Confucian exegetical tradition, yet have received little computational attention. Drawing on traditional Chinese exegetics and philology, this paper formulates collected commentary compilation as an NLP task and proposes a computational framework that preserves the contextual dependency of inline notes while enabling their automatic compilation and exegetical knowledge organization. It combines two-step prompt chaining for identifying the associated main-text segments and exegetical functions of annotations with cross-source mention clustering for integrating commentary across editions, achieving a CoNLL F1 score above 97% in a case study on the Classic of Mountains. Our framework lays the foundation for the large-scale organization of historical exegetical knowledge, thereby supporting a broad range of downstream philological and NLP tasks.
Chinese Translation
行内注释和汇编评论是儒家释义传统中演变而来的重要学术交流形式,但在计算机科学领域却鲜有关注。本文基于传统的中文释义学和语言学,将汇编评论的编纂视为一种自然语言处理(NLP)任务,并提出了一个计算框架,该框架在实现行内注释的自动编纂和释义知识组织的同时,保留了上下文依赖性。该框架结合了两步提示链(two-step prompt chaining),用于识别相关的主文本段落和注释的释义功能,以及跨源提及聚类(cross-source mention clustering),以整合不同版本的评论,在《山海经》的案例研究中实现了超过97%的CoNLL F1得分。我们的框架为历史释义知识的大规模组织奠定了基础,从而支持广泛的下游语言学和NLP任务。
cs.CL / 26 / 2607.29065

Tokenizer-Agnostic Engram Module

与分词器无关的记忆模块
Lim, Jia Peng, Chieu, Hai Leong
Abstract
Deepseek's Engram, a conditional memory module, was introduced to trade-off storage versus reasoning in large language models. However, the module relies on token-level $N$-gram hashing for Engram embedding lookup, introducing a tight coupling to the tokenizer used: a model with a different tokenizer would have to train its own Engram embeddings from scratch. To improve the reusability of Engram embeddings, we propose a change to the hashing routine, enabling compatibility between Engram models using different tokenizers. Instead of modelling disjoint $N$-gram spaces, we treat $N$-gram as a method to sample potentially useful byte sequences, from all possible byte sequences across tokens. We replace the XOR-based hashing with the general polynomial hashing with a joint embedding space across $N$. This work investigates the possible trade-offs and shows that this simple substitution produces comparable performance and achieves tokenizer-agnosticism: hash equivalence for byte-equivalent token sequences.
Chinese Translation
Deepseek 的 Engram 是一种条件记忆模块,旨在在大型语言模型中权衡存储与推理之间的关系。然而,该模块依赖于基于令牌级别的 $N$-gram 哈希进行 Engram 嵌入查找,这使得其与所使用的分词器紧密耦合:使用不同分词器的模型必须从头开始训练自己的 Engram 嵌入。为了提高 Engram 嵌入的可重用性,我们提出了一种对哈希例程的改进,使得使用不同分词器的 Engram 模型之间能够兼容。我们不再将 $N$-gram 视为不相交的空间,而是将其视为从所有可能的字节序列中采样潜在有用字节序列的方法。我们用通用多项式哈希替代了基于异或的哈希,并在 $N$ 的联合嵌入空间中进行处理。这项工作探讨了可能的权衡,并表明这一简单的替代方案能够产生可比的性能,并实现与分词器无关性:对于字节等价的令牌序列,哈希等价性得以实现。
cs.CL / 27 / 2607.29066

Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art

欺骗的语义:法律欺骗检测与通用领域最先进技术的基准比较
Samaradiwakara, Theekshana, de Silva, Nisansa, Lobb, George C.
Abstract
Deception detection has critical implications for legal proceedings, law enforcement, and online security. Although human judgment is limited in accuracy and scalability, Natural Language Processing (NLP) offers a data-driven alternative. We present a survey and comparative analysis of NLP-based Automatic Deception Detection (ADD) focusing on the legal domain, reviewing the evolution from feature-based machine learning to Large Language Model (LLM) approaches. We conduct a unified empirical evaluation across seven datasets (two legal, five general-domain), comparing six fine-tuned transformer models and seven LLMs under four prompting strategies. The results show strong domain sensitivity, with fine-tuned models excelling in data-rich general domains and few-shot LLMs remaining competitive in low-resource legal settings. Chain-of-Thought prompting often underperforms direct classification. These findings highlight the need for domain adaptation and interpretable systems in high-stakes legal contexts.
Chinese Translation
欺骗检测对法律程序、执法和在线安全具有重要意义。尽管人类判断在准确性和可扩展性方面存在局限性,自然语言处理(NLP)提供了一种数据驱动的替代方案。我们呈现了一项针对法律领域的基于NLP的自动欺骗检测(ADD)的调查和比较分析,回顾了从基于特征的机器学习到大型语言模型(LLM)方法的演变。我们在七个数据集(两个法律数据集,五个通用领域数据集)上进行了统一的实证评估,比较了六个微调的变换器模型和七个LLM在四种提示策略下的表现。结果显示出强烈的领域敏感性,微调模型在数据丰富的通用领域表现优异,而少量样本的LLM在资源匮乏的法律环境中仍具竞争力。链式思维提示通常表现不如直接分类。这些发现突显了在高风险法律环境中对领域适应和可解释系统的需求。
cs.CL / 28 / 2607.29079

Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models

更快但不同:诊断和控制加速多模态扩散语言模型中的内容漂移
Dou, Yaoxuan, Shu, Yang
Abstract
Training-free acceleration makes diffusion-based multimodal large language models (dMLLMs) more deployable, but it may silently change generated content. We study this serving-time consistency problem on 300 real images, comparing Fast-dLLM outputs with the same model's unaccelerated outputs. Across the mild parallelism induced in our long-form setting (1.05--1.25 committed tokens per step), confidence-threshold tuning changes decoding behavior but not baseline agreement. State-refresh ablations and an image-swap intervention instead identify stale visual and generated-text states as contributors to drift. For the tested Fast-dLLM implementation, shortening the KV-cache refresh interval yields a monotonic speed--agreement frontier and near-exact agreement at a measured 1.3x speedup. The initial diagnosis also appears with dLLM-Cache and LaViDa, although dLLM-Cache recovers agreement only after both caches are tightened, which removes its speed advantage. Independent prompts and images reproduce the threshold-insensitivity and refresh recovery. A targeted audit finds genuine content substitution in half of 50 low-agreement pairs. In a separate blinded two-annotator evaluation, the pooled accelerated-minus-baseline factual-error difference is 0.00 (95% CI [-0.17,+0.17]); this sample detects no difference but does not establish factual equivalence. Finally, none of the tested adaptive or smoothed-refresh variants beats the fixed interval at matched compute. Our contribution is a paired diagnostic and an implementation-scoped consistency control, not an accuracy or safety guarantee.
Chinese Translation
无训练加速使基于扩散的多模态大型语言模型(dMLLMs)更具可部署性,但它可能会悄然改变生成的内容。我们在300张真实图像上研究这一服务时间一致性问题,将Fast-dLLM的输出与同一模型的未加速输出进行比较。在我们长格式设置中(每步承诺1.05-1.25个标记)引入的轻微并行性下,置信阈值调优改变了解码行为,但未改变基线一致性。状态刷新消融实验和图像交换干预则识别出陈旧的视觉和生成文本状态是导致漂移的因素。对于测试的Fast-dLLM实现,缩短KV缓存刷新间隔产生了单调的速度-一致性边界,并在测得的1.3倍加速下实现了近乎完全的一致性。初步诊断在dLLM-Cache和LaViDa中也出现,尽管dLLM-Cache仅在两个缓存都收紧后恢复一致性,这消除了其速度优势。独立的提示和图像重现了阈值不敏感性和刷新恢复。针对性的审计发现,在50对低一致性样本中,有一半存在真实内容替换。在一项独立的盲评估中,汇总的加速减去基线的事实错误差异为0.00(95%置信区间[-0.17,+0.17]);该样本未检测到差异,但也未建立事实等价性。最后,所有测试的自适应或平滑刷新变体在匹配计算下均未超过固定间隔。我们的贡献是一个配对诊断和一个实施范围内的一致性控制,而不是准确性或安全性的保证。
cs.CL / 29 / 2607.29082

Can Zero-Shot LLMs Predict Child Malnutrition? A Fairness and Temporal Robustness Study

零样本大型语言模型能预测儿童营养不良吗?公平性与时间稳健性研究
Kabir, Muhammad Ashad, Haque, Md Ahshanul
Abstract
Child malnutrition remains a major public health challenge in low- and middle-income countries, particularly in South Asia, where early identification of vulnerable children is critical for timely intervention and resource allocation. This study aims to evaluate the feasibility, fairness, and temporal robustness of using a pretrained large language model (LLM) in a zero-shot setting for child stunting prediction using population health survey data. Using Bangladesh Demographic and Health Survey (BDHS) data collected between 2007 and 2022, we transformed maternal, child, healthcare, and household characteristics into semantically interpretable prompt-based representations and evaluated GPT-4o-mini for zero-shot stunting prediction, comparing its performance against a random forest baseline and assessing fairness across demographic and socioeconomic groups as well as temporal robustness across survey waves. The results demonstrate that zero-shot inference using GPT-4o-mini achieved comparable balanced accuracy to the supervised baseline while exhibiting substantially higher sensitivity for identifying stunting cases, relatively consistent performance across child sex groups, and stable predictive behaviour across BDHS waves; however, important fairness disparities were observed across residence and household wealth categories, highlighting the need for further investigation before deployment of foundation models in public health prediction settings.
Chinese Translation
儿童营养不良在低收入和中等收入国家仍然是一个主要的公共卫生挑战,特别是在南亚地区,早期识别脆弱儿童对于及时干预和资源分配至关重要。本研究旨在评估在零样本环境下使用预训练的大型语言模型(LLM)进行儿童生长迟缓预测的可行性、公平性和时间稳健性,所用数据来源于人口健康调查。我们使用2007年至2022年间收集的孟加拉国人口与健康调查(BDHS)数据,将母亲、儿童、医疗保健和家庭特征转化为语义可解释的基于提示的表示,并评估了GPT-4o-mini在零样本生长迟缓预测中的表现,将其与随机森林基线进行比较,并评估了其在不同人口和社会经济群体中的公平性以及在不同调查波次中的时间稳健性。结果表明,使用GPT-4o-mini进行零样本推理的平衡准确率与监督基线相当,同时在识别生长迟缓病例方面表现出显著更高的敏感性,在儿童性别组之间表现出相对一致的性能,并在BDHS波次中展现出稳定的预测行为;然而,在居住地和家庭财富类别之间观察到重要的公平性差异,强调在公共卫生预测环境中部署基础模型之前需要进一步的研究。
cs.CL / 30 / 2607.29125

M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

M3-DuplexBench:一个用于全双工语音对话模型的多轮、多语言、多领域基准测试
Fukuda, Ryo, Ando, Atsushi, Kanagawa, Hiroki, Kano, Takatomo, Delcroix, Marc, Tawara, Naohiro, Chiba, Yuya
Abstract
Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling. However, fair comparisons in multi-turn conversations remain a challenge. In addition, existing benchmarks provide limited coverage of languages and dialogue domains. We propose M3-DuplexBench, a multi-turn, multilingual, multidomain benchmark for FDSDSs. M3-DuplexBench supports English and Japanese and covers both casual conversation and multi-turn question answering. In addition, we evaluate models under multiple dialogue context settings, including single-turn, user-only, and teacher-forced full-context settings, to analyze how dialogue history affects model behavior. Experiments with recent FDSDSs reveal model-specific turn-taking characteristics, clear performance gaps across languages and domains, and mixed effects of dialogue context.
Chinese Translation
全双工语音对话系统(FDSDSs)能够在说话的同时倾听,从而实现自然的行为,如流畅的轮次转换、反馈处理和用户插话处理。然而,在多轮对话中进行公平比较仍然是一项挑战。此外,现有的基准测试在语言和对话领域的覆盖面有限。我们提出了M3-DuplexBench,一个用于FDSDSs的多轮、多语言、多领域基准测试。M3-DuplexBench支持英语和日语,并涵盖了休闲对话和多轮问答。此外,我们在多种对话上下文设置下评估模型,包括单轮、用户独占和教师强制全上下文设置,以分析对话历史如何影响模型行为。对近期FDSDSs的实验揭示了模型特定的轮次转换特征、不同语言和领域之间明显的性能差距,以及对话上下文的混合效应。
cs.CL / 31 / 2607.29168

Authorship Verification of Transcribed German-Language Videos

德语视频转录的作者验证
Halvani, Oren, Titze, Sophie
Abstract
Authorship Verification (AV) represents an important subfield of digital text forensics and addresses the fundamental question of whether two texts were written by the same author. Although the field has made substantial progress over the past two decades, several important challenges remain unresolved or underexplored. For instance, most AV research has focused on written texts, despite the fact that language is expressed not only in written but also in spoken form, such as in videos. Moreover, existing AV studies have predominantly concentrated on English, while other languages, including German, have received comparatively little attention. To address these research gaps, we apply AV to spoken language in the form of transcripts of German-language videos and examine the effectiveness of established AV methods in verifying a speaker's identity across video pairs. Our experimental evaluation, based on a total of ten AV methods applied to three self-compiled corpora comprising 300 videos from 150 speakers, shows that the best performance (up to 88% accuracy and 90% AUC) is achieved by traditional AV approaches based on simple character- and token n-gram representations. In contrast, more modern transformer-based approaches perform significantly worse on all evaluated corpora. Our results therefore suggest that traditional methods in the field of AV remain both competitive and relevant.
Chinese Translation
作者验证(Authorship Verification, AV)是数字文本取证的重要子领域,旨在解决两个文本是否由同一作者撰写的基本问题。尽管该领域在过去二十年中取得了显著进展,但仍然存在若干重要挑战尚未解决或未得到充分探索。例如,大多数AV研究集中于书面文本,尽管语言不仅以书面形式表达,也以口头形式存在,例如视频。此外,现有的AV研究主要集中于英语,而其他语言,包括德语,则相对较少受到关注。为了解决这些研究空白,我们将AV应用于口语语言,采用德语视频的转录文本,并检验已建立的AV方法在视频对中验证说话者身份的有效性。我们的实验评估基于对三个自编语料库中150位说话者的300个视频应用的十种AV方法,结果显示,基于简单字符和词元n-gram表示的传统AV方法表现最佳(准确率高达88%,AUC高达90%)。相比之下,更现代的基于变换器的方法在所有评估的语料库上表现显著较差。因此,我们的结果表明,AV领域的传统方法仍然具有竞争力和相关性。
cs.CL / 32 / 2607.29185

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End

端到端学习潜在推理轨迹以用于标量奖励模型
Lee, Sanwoo, Bai, Clive, Huang, Hsiu-Yuan, Liang, Kun, Liu, Weijie, Wu, Yunfang
Abstract
Reward models (RMs) are central to aligning large language models with human preferences via reinforcement learning. Although traditional scalar RMs enable efficient and probabilistic reward modeling, they rely on superficial cues that fail to generalize to complex or out-of-distribution (OOD) tasks. Conversely, generative RMs leverage extensive reasoning to improve robustness on challenging tasks, but their natural language-based scores lack the numerical flexibility and probabilistic interpretability that scalar RMs offer. While recent approaches combine both paradigms through off-policy multi-task learning, such parallel optimization does not guarantee that generated reasoning traces actively align with or benefit downstream scalar reward prediction. To address this mismatch, we propose LatentRM, a reward modeling framework that learns intermediate reasoning traces as discrete latent variables to explicitly maximize the likelihood of downstream scalar rewards. Through on-policy optimization of the latent reasoning space end-to-end, LatentRM tightly couples deep reasoning-based evaluation with precise scoring. Extensive validations on in-distribution and OOD datasets and RLHF show that LatentRM outperforms scalar, generative, and hybrid RMs on preference modeling and policy alignment across tasks ranging from open-ended conversation to complex reasoning.
Chinese Translation
奖励模型(RMs)在通过强化学习将大型语言模型与人类偏好对齐中起着核心作用。尽管传统的标量奖励模型能够实现高效且概率性的奖励建模,但它们依赖于表面的线索,无法在复杂或分布外(OOD)任务中进行有效泛化。相反,生成性奖励模型利用广泛的推理来提高在挑战性任务中的鲁棒性,但其基于自然语言的评分缺乏标量奖励模型所提供的数值灵活性和概率可解释性。尽管最近的方法通过离策略多任务学习结合了这两种范式,但这种并行优化并不能保证生成的推理轨迹能够积极对齐或惠及下游的标量奖励预测。为了解决这一不匹配问题,我们提出了LatentRM,这是一种奖励建模框架,学习作为离散潜变量的中间推理轨迹,以明确最大化下游标量奖励的似然性。通过对潜在推理空间的端到端的在线优化,LatentRM将基于深度推理的评估与精确评分紧密结合。在分布内和分布外数据集以及强化学习人类反馈(RLHF)上的广泛验证表明,LatentRM在偏好建模和策略对齐方面优于标量、生成性和混合奖励模型,适用于从开放式对话到复杂推理的各种任务。
cs.CL / 33 / 2607.29188

Detecting Experiential Intertextuality Across Migration Routes: Beyond Surface Similarity in French Narratives

Sari, Sakayo Toadoum, Robin, Nelly, Auzanneau, Michelle, Sais, Lakhdar, Petit, Veronique, Veniard, Marie, Jabbour, Said, Delorme, Fabien
Abstract
Migrants traversing geographically distinct routes such as the Trans-Saharan and Balkan corridors often recount strikingly parallel lived experiences: police violence, smuggler exploitation, dangerous crossings, and family separation. We introduce the task of experiential intertextuality detection: automatically identifying shared experiential echoes across migration narratives without requiring annotated training data. From 108 French migration narratives spanning both corridors, we automatically generate sentence pairs and score them using annotation-free methods: lexical baselines, sentence embeddings, POS-based structural features, a migration-specific theme lexicon, context-aware narrative features, and zero-shot LLM scoring with Qwen2.5-7B and Mistral-7B under three prompting strategies. We validate all methods against 816 expert-annotated intertextuality judgments (inter-annotator Krippendorff's $\alpha = 0.27$). Our results reveal that all surface, structural, and embedding methods correlate only weakly with expert judgments ($r \leq 0.30$); Qwen2.5-7B zero-shot achieves the best single-method correlation ($r = 0.38$); few-shot examples degrade Qwen but dramatically improve Mistral; narrative position significantly predicts intertextuality, with departure-phase pairs showing the highest experiential echoes; and a supervised hybrid combining all 31 features achieves $r = 0.45$, a 21% improvement over the best individual method.
cs.CL / 34 / 2607.29196

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding

Hy-MultiTurn:一个六维基准用于深度多轮对话理解
Ye, Eileen, Tao, Jiawen, Li, Yaoming, Liu, Chenxu, Yu, Wenhan, Fan, Yaxin, Yuan, Xiaokun, Wu, Mengzhou, Jiang, Yanbing, Pan, Maxm
Abstract
Long-running multi-turn interactions with chatbots and agents are now common, and a correct response often depends on remembering earlier details, tracking later revisions, identifying intended objects or referents, and withholding action when required conditions are unmet. Existing multi-turn benchmarks typically cover short exchanges and do not fully evaluate these capabilities in long multi-turn interactions, particularly in Chinese, while offering limited insight into how and why models fail. To address these limitations, we analyze real chatbot failures to identify six recurring mechanisms and use them to define six controlled evaluation modes in Hy-MultiTurn, a Chinese benchmark for deep multi-turn dialogue understanding. The six modes evaluate constraint memory, precise execution, constraint synthesis, object localization, action suppression, and reference resolution. Across the six modes, we construct 209 controlled tasks spanning 12-76 turns, with dialogue length, irrelevant-topic distraction, and colloquial phrasing adding further difficulty. Evaluation of 22 frontier model configurations shows that Hy-MultiTurn is broadly challenging, as even GPT-5.5, the strongest overall configuration, satisfies all requirements in only 41.1 percent of responses and no model performs best in all six modes.
Chinese Translation
与聊天机器人和智能体进行长时间的多轮交互已变得普遍,而正确的回应往往依赖于记忆早期细节、跟踪后续修订、识别意图对象或指称,以及在必要条件未满足时抑制行动。现有的多轮基准通常涵盖短暂的交流,未能全面评估这些能力在长多轮交互中的表现,特别是在中文环境下,同时对模型失败的原因和机制提供的见解有限。为了解决这些局限性,我们分析了真实聊天机器人的失败案例,以识别六种重复出现的机制,并利用这些机制在Hy-MultiTurn中定义六种受控评估模式,这是一个用于深度多轮对话理解的中文基准。这六种模式评估约束记忆、精确执行、约束综合、对象定位、行动抑制和指称解析。在这六种模式下,我们构建了209个受控任务,跨越12至76轮对话,且对话长度、无关主题的干扰和口语化表达增加了进一步的难度。对22种前沿模型配置的评估表明,Hy-MultiTurn具有广泛的挑战性,即使是整体最强的配置GPT-5.5,仅在41.1%的回应中满足所有要求,且没有模型在所有六种模式中表现最佳。
cs.CL / 35 / 2607.29211

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning

知道何时放弃:诊断和训练大型语言模型以中止无效推理
Guan, Xinyan, Zeng, Jiali, Xin, Chunlei, Lu, Yaojie, Lin, Hongyu, Han, Xianpei, Sun, Le, Meng, Fandong
Abstract
Large language models generate computationally expensive yet semantically void reasoning on beyond-capability tasks, creating risks where plausible-sounding but incorrect derivations mislead users. We characterize this \textit{futile reasoning} phenomenon through systematic analysis, revealing universal capability overreach and systematic miscalibration between capability and behavior. The dominant failure mode is specious reasoning, which outputs look superficially valid but contain subtle errors, escalating with task difficulty. To address this, we introduce \textbf{CaRL} (\textbf{Ca}pability-\textbf{a}ligned \textbf{R}einforcement \textbf{L}earning), which aligns model behavior with capability boundaries through reward shaping that incentivizes refusal over futile reasoning and hindsight refusal augmentation that converts failures into refusal supervision. Experiments demonstrate a substantial reduction in futile reasoning while preserving performance across task difficulties, effectively achieving capability-aligned behavior without sacrificing utility. \footnote{https://github.com/icip-cas/Knowing-When-to-Quit}
Chinese Translation
大型语言模型在超出能力范围的任务上生成计算成本高昂但语义空洞的推理,这带来了风险,因为貌似合理但错误的推导可能误导用户。我们通过系统分析对这一 extit{无效推理}现象进行了特征描述,揭示了普遍的能力超越和能力与行为之间的系统性误校准。主要的失败模式是似是而非的推理,其输出表面上看似有效,但包含微妙的错误,且随着任务难度的增加而加剧。为了解决这个问题,我们引入了 extbf{CaRL}( extbf{Ca}pability- extbf{a}ligned extbf{R}einforcement extbf{L}earning),通过奖励塑造将模型行为与能力边界对齐,激励拒绝无效推理,并通过回顾性拒绝增强将失败转化为拒绝监督。实验表明,在保持任务难度下的性能的同时,显著减少了无效推理,有效实现了能力对齐的行为而不牺牲实用性。
cs.CL / 36 / 2607.29238

Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters

小即是足:通过 LoRA 适配器实现的用户个性化 AI 编辑文本重写
Chakravorty, Antorweep
Abstract
InMyStyle is a privacy first, single user system that adapts small language models to rewrite AI-edited text towards an individual user's writing style without an instruction prompt at inference. Given a user's documents, it uses multiple local helper LLMs to construct paired training examples and fine tunes LoRA adapters on base models ranging from 0.5B to 7B parameters. Length aware generation budgets and automatic chunking support inputs of different lengths. On 219 evaluation pairs from a scientific-paper corpus, the automatic composite score plateaus at 0.69 [scale 0-1] across all model sizes under both greedy and sampled decoding. This observed plateau suggests that small models are sufficient for the measured rewriting task, with model size determining trade-offs rather than a stable quality ranking. As a secondary evaluation, 400 ratings from five LLM judges give InMyStyle outputs a mean perceived AI-ness score over 20% lower than their helper-AI generated inputs, while mean perceived AI-ness scores decrease with model size within InMyStyle.
Chinese Translation
InMyStyle 是一个以隐私为首的单用户系统,旨在适应小型语言模型,以无需指令提示的方式将 AI 编辑的文本重写为个别用户的写作风格。该系统利用用户的文档,通过多个本地辅助大型语言模型(LLMs)构建成对的训练示例,并对从 0.5B 到 7B 参数的基础模型进行 LoRA 适配器的微调。系统支持长度感知生成预算和自动分块,以处理不同长度的输入。在来自科学论文语料库的 219 个评估对中,自动综合得分在所有模型规模下的贪婪解码和采样解码中均达到 0.69 [范围 0-1] 的平台期。这一观察到的平台期表明,小型模型对于所测量的重写任务是足够的,模型规模决定了权衡而不是稳定的质量排名。作为二次评估,来自五位 LLM 评审的 400 个评分显示,InMyStyle 输出的平均感知 AI 评分比其辅助 AI 生成的输入低超过 20%,而在 InMyStyle 中,平均感知 AI 评分随着模型规模的增大而降低。
cs.CL / 37 / 2607.29250

Data Turnstile: A Scalable Open Framework for Function-Calling Data Generation

数据转闸:一个可扩展的开放框架用于函数调用数据生成
Ramakrishnan, Goutham, Sharma, Megha
Abstract
Small language models (SLMs) are attractive for agentic deployment due to low latency, reduced cost, and on-device privacy, yet they struggle with tool-use tasks where training data is scarce and noisy. Unlike larger models, SLMs cannot compensate for low-quality supervision through sheer capacity, making data quality the critical bottleneck. We present Data Turnstile, an open-source framework that takes user-defined API specifications and generates high-quality synthetic training data for function calling. Turnstile decomposes multi-turn tool-use interactions into constrained, stepwise generation with validation and error-feedback loops, providing fine-grained control over API diversity, conversation complexity, and output correctness. We demonstrate effectiveness of domain adaptation with Turnstile data on two challenging function calling benchmarks. On the BFCL single-turn benchmark, a Qwen3-0.6B fine-tuned on Turnstile data without chain-of-thought achieves 75.9% overall accuracy (versus 67.4% for the base model with thinking enabled), closing the gap with thinking-enabled Qwen3-1.7B (78.4%) and Qwen3-4B (79.9%) despite being 3$\times$ and 7$\times$ smaller respectively. On $\tau^2$-bench, a multi-turn agentic benchmark, Turnstile-trained Qwen3-1.7B achieves 31.1% pass^1 on the Telecom domain, improving 4.7$\times$ over its 6.6% base and surpassing Qwen2.5-32B-Instruct (27.4%), a model 19$\times$ larger. Turnstile-trained Qwen3-0.6B achieves 24.6%, improving 7$\times$ over its 3.5% base and approaching the 32B model (53$\times$ larger). We release Data Turnstile along with a dataset spanning 1,000+ APIs and 100K+ multi-turn interactions.
Chinese Translation
小型语言模型(SLMs)因其低延迟、降低成本和设备隐私而在代理部署中具有吸引力,但在训练数据稀缺且噪声较多的工具使用任务中表现不佳。与更大的模型不同,SLMs无法通过单纯的容量来弥补低质量监督的问题,因此数据质量成为关键瓶颈。我们提出了数据转闸(Data Turnstile),一个开源框架,它接受用户定义的API规范并生成高质量的合成训练数据以用于函数调用。转闸将多轮工具使用交互分解为受限的、逐步生成的过程,并结合验证和错误反馈循环,从而提供对API多样性、对话复杂性和输出正确性的细粒度控制。我们在两个具有挑战性的函数调用基准上展示了使用转闸数据进行领域适应的有效性。在BFCL单轮基准上,经过转闸数据微调的Qwen3-0.6B在未启用思维链的情况下达到了75.9%的整体准确率(而基础模型在启用思维时为67.4%),缩小了与启用思维的Qwen3-1.7B(78.4%)和Qwen3-4B(79.9%)之间的差距,尽管它们分别大约是其3倍和7倍大。在$ au^2$-bench这一多轮代理基准上,经过转闸训练的Qwen3-1.7B在电信领域达到了31.1%的通过率,比其6.6%的基础提升了4.7倍,并超过了Qwen2.5-32B-Instruct(27.4%),后者的规模是其19倍。经过转闸训练的Qwen3-0.6B达到了24.6%,比其3.5%的基础提升了7倍,并接近32B模型(大53倍)。我们发布了数据转闸及其包含的涵盖1000多个API和10万多个多轮交互的数据集。
cs.CL / 38 / 2607.29252

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

CalibratedRubric:用于开放式LLM评估的任务自适应评分标准库
Chen, Mengting, Sun, Yanshu, Liang, Wanting, Luan, Beidi, Sun, Rui, Chen, Dezhi, Li, Jing, Bai, Zuo
Abstract
Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge unanimity and binary variance filters, which cannot distinguish measurable rubrics from informative ones. We introduce CalibratedRubric, a task-adaptive framework that combines type-specific scoring, Bayesian rubric-measurability filtering, and item response theory (IRT)-based bank assembly. CalibratedRubric estimates each rubric's measurability with a Beta--Bernoulli agreement posterior and uses a submodular information-coverage objective to construct compact rubric banks over the observed capability range. Across financial, healthcare, general, and legal benchmarks, measurability filtering improves human-gold agreement on JudgmentBench from $\kappa=0.604$ to $0.743$. IRT-based greedy selection improves cross-fitted rank fidelity over random selection across all six evaluated response blocks and requires only 49 rather than 131 rubrics to reach the target correlation on FinResearchBench decision-support tasks. Task-label perturbations further reduce system separation, confirming the practical relevance of task-adaptive scoring. These results support CalibratedRubric as an efficient, uncertainty-aware approach to open-ended LLM evaluation, with calibration gains depending on sufficient judge redundancy.
Chinese Translation
对开放式LLM输出的可靠评估需要细致的评分标准,但专家策划成本高且难以扩展。现有的自动化流程依赖于严格的评审一致性和二元方差过滤,这无法区分可测量的评分标准与信息性评分标准。我们提出了CalibratedRubric,这是一种任务自适应框架,结合了特定类型的评分、贝叶斯评分标准可测量性过滤和基于项目反应理论(IRT)的评分标准库构建。CalibratedRubric通过Beta-Bernoulli一致性后验估计每个评分标准的可测量性,并使用子模块信息覆盖目标在观察到的能力范围内构建紧凑的评分标准库。在金融、医疗保健、一般和法律基准测试中,可测量性过滤将JudgmentBench上的人类金标准一致性从$ ext{kappa}=0.604$提高到$0.743$。基于IRT的贪婪选择在所有六个评估响应块中改善了交叉拟合排名的保真度,相较于随机选择仅需49个评分标准而非131个即可在FinResearchBench决策支持任务中达到目标相关性。任务标签扰动进一步减少了系统分离,确认了任务自适应评分的实际相关性。这些结果支持CalibratedRubric作为一种高效、考虑不确定性的开放式LLM评估方法,其校准收益依赖于足够的评审冗余。
cs.CL / 39 / 2607.29287

Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation

思维翻译:通过强化学习实现多领域机器翻译的难度自适应推理
Ye, Yongshi, Fu, Biao, Huang, Chongxuan, Chen, Yidong, Shi, Xiaodong
Abstract
Multi-domain machine translation (MDMT) poses a unique challenge due to varying levels of linguistic complexity across domains. Inspired by human translators' ability to adapt reasoning effort based on difficulty, we propose TwT (Translation with Thought), a resource-rational framework that learns to modulate inference between intuitive and deliberate reasoning. TwT is trained in two stages: (1) supervised fine-tuning on difficulty-aware long chain-of-thought traces distilled from DeepSeek-R1 and rewritten by GPT-4o to reflect human-like reasoning economy, and (2) reinforcement learning with a hybrid reward to optimize translation quality and reasoning efficiency. Evaluated on 15 benchmarks spanning in-domain and out-of-domain settings, as well as 3 seen and 59 unseen languages, with ablations across three backbone models, TwT-7B and TwT-14B outperform much larger SOTA reasoning models in translation quality, while reducing token usage by 32--60\%. These results confirm that aligning translation behavior with cognitive principles enables robust generalization, high translation quality, and efficient reasoning in MDMT.
Chinese Translation
多领域机器翻译(MDMT)由于各领域语言复杂性的不同而面临独特挑战。受到人类翻译者根据难度调整推理努力的能力的启发,我们提出了TwT(思维翻译),这是一个资源理性的框架,旨在学习在直观推理和深思熟虑推理之间进行调节。TwT的训练分为两个阶段:(1)在从DeepSeek-R1提取的难度感知长链思维轨迹上进行监督微调,这些轨迹经过GPT-4o重写,以反映人类般的推理经济性;(2)通过混合奖励进行强化学习,以优化翻译质量和推理效率。在涵盖领域内和领域外设置的15个基准测试,以及3种已见语言和59种未见语言的评估中,TwT-7B和TwT-14B在翻译质量上超越了更大规模的最先进(SOTA)推理模型,同时将令牌使用量减少了32%至60%。这些结果确认了将翻译行为与认知原则对齐能够实现稳健的泛化、高翻译质量和高效推理在多领域机器翻译中的有效性。
cs.CL / 40 / 2607.29355

Cross-Lingual Transfer for Machine Translation in Turkic Languages

突厥语言机器翻译的跨语言迁移
Cinar, Omer Burak, Dalkilic, Mehmet Mert, Toraman, Cagri
Abstract
Cross-lingual transfer is central to low-resource machine translation, but its behavior within closely related language families remains insufficiently characterized. We study transfer among five Turkic languages; Turkish, Azerbaijani, Uzbek, Kazakh, and Kyrgyz; using pairwise transfer matrices. In this setting, each model is fine-tuned with one transfer source and evaluated on a different transfer target while the translation target remains the same. Across mT5 experiments, we find that transfer is strongest between closely related Turkic pairs, especially Turkish-Azerbaijani and Kazakh-Kyrgyz. We also show that transfer direction matters, and that the same transfer source-transfer target pair can behave differently when the translation target changes. Latinization improves BLEU and chrF in several script-mismatched settings, but its effect is not uniform across metrics. Additional analyses show that transfer sources are mostly stable across different datasets and model settings.
Chinese Translation
跨语言迁移是低资源机器翻译的核心,但在紧密相关的语言家族中的表现尚未得到充分表征。我们研究了五种突厥语言之间的迁移;土耳其语、阿塞拜疆语、乌兹别克语、哈萨克语和吉尔吉斯语;使用成对迁移矩阵。在这种设置下,每个模型都使用一个迁移源进行微调,并在不同的迁移目标上进行评估,同时翻译目标保持不变。在 mT5 实验中,我们发现紧密相关的突厥语言对之间的迁移最强,尤其是土耳其语-阿塞拜疆语和哈萨克语-吉尔吉斯语。我们还表明迁移方向很重要,并且相同的迁移源-迁移目标对在翻译目标变化时可能表现不同。拉丁化在多个脚本不匹配的设置中提高了 BLEU 和 chrF,但其效果在不同指标之间并不均匀。额外的分析表明,迁移源在不同的数据集和模型设置中大多保持稳定。
cs.CL / 41 / 2607.29377

Zero-Mem: Zero-Token Memory Operations for LLM Agents

Zero-Mem:针对大语言模型代理的零令牌记忆操作
Xiao, Yilin, Zhu, Zhehan, Zhang, Yujing, Chen, Jin, Hong, Zijin, Zhuang, Luyao, Zhang, Qinggang, Chen, Shengyuan, Ouyang, Xiaocao, Ren, Lingfei, Huang, Xiao
Abstract
LLM agents need memory to act consistently over long interactions, yet many systems use additional LLM calls to operate that memory. Generating intermediate records and mediating their retrieval adds recurring token and time costs, while omitted or merged details can obscure the original evidence. We ask whether structured memory access requires generation at all. Zero-Mem introduces \emph{zero-token memory operations}: no step outside final question answering invokes an LLM or consumes LLM input or output tokens; encoder computation is accounted for separately. Zero-Mem preserves original interaction traces as its source of record. It organizes the traces in two complementary ways. An entity--context graph exposes connections across interactions, while a temporal hierarchy preserves conversational locality and session state. For each query, Zero-Mem weighs the two views, retrieves from both, and follows their structure to recover supporting relations or surrounding context. Deterministic calibration first discards conflicting evidence and then keeps the reader's answer grounded in the retrieved traces. Only the final-QA reader invokes an LLM. Across long-memory and long-context question-answering benchmarks, Zero-Mem achieves competitive performance while eliminating LLM calls and LLM-token consumption from memory operations. With the same final-QA reader and context budget, it reduces memory-operation time cost by 57.6\% relative to the fastest compared baseline. Ablations support the contribution of the two views and their query-dependent coordination. Overall, the results show that structured agent memory need not generate an intermediate representation of the past. After peer review, the code and implementation details will be available at \textcolor{blue}{https://github.com/TheMoon0815/Zero-mem}.
Chinese Translation
大语言模型(LLM)代理需要记忆以在长时间交互中保持一致性,然而许多系统使用额外的LLM调用来操作这些记忆。生成中间记录并调解其检索会增加重复的令牌和时间成本,而省略或合并的细节可能会模糊原始证据。我们探讨结构化记忆访问是否完全需要生成。Zero-Mem 引入了 extit{零令牌记忆操作}:在最终的问题回答之外,没有任何步骤调用 LLM 或消耗 LLM 输入或输出令牌;编码器计算被单独计算。Zero-Mem 保留原始交互痕迹作为其记录来源。它以两种互补的方式组织这些痕迹。实体-上下文图揭示了交互之间的连接,而时间层次结构则保留了对话的局部性和会话状态。对于每个查询,Zero-Mem 权衡这两种视角,从两者中检索,并遵循其结构以恢复支持关系或周围上下文。确定性校准首先丢弃冲突证据,然后使读者的答案基于检索到的痕迹保持一致。只有最终的问答阅读器调用 LLM。在长记忆和长上下文问答基准测试中,Zero-Mem 在消除记忆操作中的 LLM 调用和 LLM 令牌消耗的同时,实现了具有竞争力的性能。在相同的最终问答阅读器和上下文预算下,相较于最快的比较基线,它将记忆操作的时间成本降低了 57.6\%。消融实验支持这两种视角及其查询依赖的协调。总体而言,结果表明结构化代理记忆不必生成过去的中间表示。经过同行评审后,代码和实现细节将可在 extcolor{blue}{https://github.com/TheMoon0815/Zero-mem} 获取。
cs.CL / 42 / 2607.29378

PTP: Previous-Token Prediction based LLM Inversion for Near-Exact Prompt Reconstruction

PTP:基于前一个标记预测的LLM反演用于近似精确的提示重建
Suhail, Pirzada, Naidu, Nagasai Saketh, Sinha, Atanu R, Sethi, Amit
Abstract
Large language models (LLMs) generate text by auto-regressively sampling the next token. This inherently leads to a many-to-many mapping between prompts and responses, complicating the task of inferring prompts from observed outputs. Prior work on LLM inversion frames prompt recovery as a semantic reconstruction task. They rely on fine-tuning pretrained sequence-to-sequence models on large external datasets--and requiring access to model weights or logits--to generate semantically plausible prompts. In contrast, we present a functional approach to inverting a given LLM in a black-box setting, without auxiliary aids. We train an explicit inverse language model entirely from scratch on data synthetically generated from the target LLM itself. Analogous to forward next-token prediction, our inverse model is trained using previous-token prediction, establishing a generative link between the forward and inverse processes that enables faithful prompt reconstruction. Moreover, it naturally supports diverse prompt reconstructions through sampling, whereby all such prompts induce similar responses under the forward, target LLM. Our approach generalises across datasets and exhibits transferability in reconstructing prompts from responses generated by different LLMs. Further, across the set of token based evaluation metrics for prompt and response reconstructions, our approach outperforms prior work.
Chinese Translation
大型语言模型(LLMs)通过自回归地采样下一个标记生成文本。这本质上导致了提示与响应之间的多对多映射,复杂化了从观察到的输出推断提示的任务。先前的LLM反演工作将提示恢复框架视为语义重建任务。它们依赖于在大型外部数据集上微调预训练的序列到序列模型,并需要访问模型权重或logits,以生成语义上合理的提示。相比之下,我们提出了一种在黑箱环境中反演给定LLM的功能性方法,无需辅助工具。我们完全从头开始训练一个显式的逆语言模型,数据来自于目标LLM自身合成生成。类似于前向下一个标记预测,我们的逆模型使用前一个标记预测进行训练,建立了前向和逆向过程之间的生成联系,从而实现忠实的提示重建。此外,它通过采样自然支持多样的提示重建,所有这些提示在前向目标LLM下都会引发类似的响应。我们的方法在数据集之间具有广泛的适应性,并在从不同LLM生成的响应中重建提示时表现出可迁移性。此外,在针对提示和响应重建的基于标记的评估指标集合中,我们的方法优于先前的工作。
cs.CL / 43 / 2607.29397

Studying quantization trade-offs for efficient inference deployment in machine translation

研究机器翻译中高效推理部署的量化权衡
Zhao, Jim, Maskey, Sohir, Oostermeijer, Koen, Orr, Douglas, Jones, Teryn
Abstract
Deploying large language models in realistic server environments poses challenges, as the system needs to provide high-quality responses with low latency. Quantization is a common approach to reduce the memory footprint and improve inference efficiency, yet its impact on latency and throughput is rarely evaluated under controlled, orchestration-level workloads. In this work we study the quantization trade-offs of two translation model families, EuroLLM \citep{martins2025eurollm} and Hy-MT2 \citep{zheng2026hy} across five models ranging from 1.7B to 22B for efficient deployment on a single A100 or H100 GPU. We demonstrate that combining a document-chunking strategy with W4A8 or W8A8 quantization improves the latency-throughput Pareto-curve under a wide range of workloads. Furthermore, since standard machine translation (MT) benchmarks rely on isolated sentences and fail to capture long-context dynamics, we introduce a document-level evaluation from WMT24++ to assess how text chunking strategies affect translation quality under quantization. Our results reveal that standard segment-level evaluation can fail to predict the interaction between quantization and long-context document translation. While Hy-MT2 remains robust under quantization, EuroLLM shows strong sensitivity and translation quality collapses rapidly for all considered quantization formats. Overall, our experiments show that the trade-off between inference efficiency and translation quality depends not only on the quantization format, but also on the choice of text chunking strategy.
Chinese Translation
在现实的服务器环境中部署大型语言模型面临挑战,因为系统需要在低延迟的情况下提供高质量的响应。量化是一种常见的方法,用于减少内存占用并提高推理效率,但其对延迟和吞吐量的影响在受控的编排级工作负载下很少被评估。在本研究中,我们研究了两种翻译模型家族的量化权衡,即 EuroLLM  extit{(EuroLLM)}  extit{(martins2025eurollm)} 和 Hy-MT2  extit{(Hy-MT2)}  extit{(zheng2026hy)},涵盖从 1.7B 到 22B 的五个模型,以便在单个 A100 或 H100 GPU 上进行高效部署。我们证明,将文档分块策略与 W4A8 或 W8A8 量化相结合,可以在广泛的工作负载下改善延迟-吞吐量的帕累托曲线。此外,由于标准机器翻译 (MT) 基准依赖于孤立的句子,未能捕捉长上下文动态,我们引入了来自 WMT24++ 的文档级评估,以评估文本分块策略在量化下对翻译质量的影响。我们的结果揭示,标准的段落级评估可能无法预测量化与长上下文文档翻译之间的相互作用。尽管 Hy-MT2 在量化下仍然表现稳健,但 EuroLLM 显示出强烈的敏感性,所有考虑的量化格式下翻译质量迅速下降。总体而言,我们的实验表明,推理效率与翻译质量之间的权衡不仅取决于量化格式,还取决于文本分块策略的选择。
cs.CL / 44 / 2607.29433

Know It, Act on It: Investigating Memory Utilization in LLM Personalization

知之、行之:探讨大语言模型个性化中的记忆利用
Feng, Zhaoxin, Ma, Jianfei, Chersoni, Emmanuele
Abstract
As large language model (LLM) agents evolve into personalized companions, memory has emerged as a core capability. However, LLMs face a knowledge utilization problem: they may fail to act on relevant user preferences even when they are fully present in context. When an agent fails to tailor its response in a context where previously shared user preferences should matter, it is unclear whether the model failed to remember that information or remembered it but failed to use it. To isolate this breakdown, we introduce a decoupled evaluation paradigm that administers paired Know and Act tests to the same user preference. We conduct large-scale experiments across 16 systems and five memory architectures, evaluating 1,000 preferences embedded at three levels of expression strength. Our results show a large gap between Know and Act outcomes: agents often pass the recall test for a user preference but fail to reflect that same preference in the paired behavioral scenario. While memory architectures reduce this gap, utilization remains especially weak for health and therapy-related preferences, where failures to act carry the greatest real-world stakes.
Chinese Translation
随着大语言模型(LLM)代理演变为个性化伴侣,记忆已成为其核心能力。然而,LLM面临知识利用问题:即使相关的用户偏好在上下文中完全存在,它们也可能未能对此做出反应。当代理在一个应该考虑之前共享的用户偏好的上下文中未能调整其回应时,尚不清楚模型是未能记住该信息,还是记住了但未能加以利用。为了隔离这种失效,我们引入了一种解耦评估范式,对同一用户偏好进行配对的知(Know)与行(Act)测试。我们在16个系统和五种记忆架构上进行了大规模实验,评估了在三种表达强度水平下嵌入的1,000个偏好。我们的结果显示知与行结果之间存在较大差距:代理通常能够通过用户偏好的回忆测试,但未能在配对的行为场景中反映出相同的偏好。尽管记忆架构缩小了这一差距,但在健康和治疗相关偏好方面的利用仍然特别薄弱,而这些偏好的失效在现实世界中具有最大的风险。
cs.CL / 45 / 2607.29484

Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?

证据类型竞争:干预数据何时能教会语言模型因果方向?
Xun, Xining
Abstract
Interventional data is widely regarded as the gold standard for teaching models causal reasoning. We test this assumption in a fully controlled synthetic environment pitting observational correlation against causal effect, and find it fails instructively. In Simpson's-paradox worlds, where the two have systematically opposite signs, increasing the fraction of interventional samples in pretraining does not improve causal direction: the magnitude of the model's do()-response grows monotonically, yet its sign is copied from the observational context. What governs whether interventional evidence is used is not the training mixture but the evidence type present in the context at inference time. Under an identical training recipe, a purely observational context induces systematic sign reversal in 29/50 worlds, a mixed context in 19/50, while aligned interventional probes alone yield 41/50 correct. Erasing observational evidence from the context immediately releases the suppressed causal interpolation ability (ratio_true = +0.56); a four-state content manipulation shows the switch is content-mediated and graded. The suppression is stable across training seeds (11/11 strong reversals persist on a matched-protocol second seed) and robust as a rate at 0.93B parameters (31.8% vs. 6% reversals in the matched probe-only arm), even as absolute gains shrink four-fold. An external audit on CLadder exposes a learned positive-effect prior with a two-layer structure: sign-randomized retraining removes it in-distribution but not out-of-distribution. We summarize: the capability lives in the weights; the switch lives in the context, and activation patching localizes the switch to the middle layers' observational rows. We further quantify the sampling noise floor of probe-based causal evaluation and an evidence-averaging protocol that cuts sign errors from 26% to 9%.
Chinese Translation
干预数据被广泛认为是教导模型因果推理的金标准。我们在一个完全受控的合成环境中测试这一假设,将观察相关性与因果效应进行对比,发现其在指导上失败。在辛普森悖论的世界中,二者的符号系统性相反,增加预训练中干预样本的比例并未改善因果方向:模型的 do() 响应的幅度单调增长,但其符号却是从观察上下文中复制而来。决定干预证据是否被使用的并不是训练混合,而是推理时上下文中存在的证据类型。在相同的训练方案下,纯观察上下文在 50 个世界中有 29 个导致系统性的符号反转,混合上下文在 19 个,而仅有对齐的干预探针则能在 50 个世界中获得 41 个正确结果。从上下文中删除观察证据立即释放被抑制的因果插值能力(ratio_true = +0.56);四状态内容操控显示该切换是内容介导的且是分级的。该抑制在训练种子之间是稳定的(11/11 个强反转在匹配协议的第二个种子上持续存在),并且在 0.93B 参数下作为一个比率是稳健的(在匹配探针仅臂中反转率为 31.8% 对比 6%),即使绝对增益缩小了四倍。对 CLadder 的外部审计揭示了一个具有两层结构的学习到的正效应先验:符号随机化再训练在分布内去除了它,但在分布外则没有。我们总结道:能力存在于权重中;切换存在于上下文中,而激活修补将切换定位于中间层的观察行。我们进一步量化了基于探针的因果评估的采样噪声底线以及一种证据平均协议,该协议将符号错误从 26% 降低到 9%。
cs.CL / 46 / 2607.29539

ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

ARB:用于AI文本检测器评估的匹配作者重写基准数据集
Perrone, Gaetano, Romano, Simon Pietro
Abstract
Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and paraphrasing can degrade detector performance, it remains unclear whether performance measured on this conventional benchmark predicts detector behavior when human-authored content is rewritten by an LLM. To address this gap, we introduce Authorship-Rewriting Benchmark (ARB), built from 1,800 human source texts (600 each from XSum, WritingPrompts, and OpenWebText) and four open-weight generators (Llama-3.2-3B, Qwen2.5-7B, Mistral-7B, Gemma-2-9B). Each source item yields four matched variants: human-written (HUMAN), direct LLM generation (Free-LLM), LLM-rewritten human text (H2L), and same-generator LLM-rewritten LLM text (LLM2L). We evaluated five detectors (FastDetectGPT, Binoculars-falcon-7b, RADAR, BERT-Defense, RoBERTa-Defense) at a strict 1%-false-positive operating point (TPR@1%FPR). FastDetectGPT and Binoculars-falcon-7b detected 91.2% and 93.5\% of direct LLM text, but only 30.8% and 15.1% of human text an LLM had rewritten, a drop of 60-78 percentage points. The same detectors retained 78.3% and 83.0% recall when LLM text was rewritten by the same model, a much smaller decline of 10-13 points. RADAR followed the same pattern (66.8% to 12.2%), while BERT-Defense and RoBERTa-Defense stayed below 3% recall across all regimes. These results show that detector performance measured on the conventional human-vs-LLM benchmark does not transfer to human-authored text revised by an LLM, even though the same detectors remain largely robust to LLM-only rewriting.
Chinese Translation
标准的AI文本检测基准将人类撰写的文本与大型语言模型(LLMs)直接生成的文本进行比较。尽管先前的研究表明重写和释义可能会降低检测器的性能,但尚不清楚在这一传统基准上测得的性能是否能够预测当人类创作的内容被LLM重写时检测器的行为。为了解决这一问题,我们引入了作者重写基准(Authorship-Rewriting Benchmark,ARB),该基准由1800篇人类源文本(分别来自XSum、WritingPrompts和OpenWebText的600篇)和四个开放权重生成器(Llama-3.2-3B、Qwen2.5-7B、Mistral-7B、Gemma-2-9B)构成。每个源项目生成四个匹配变体:人类撰写(HUMAN)、直接LLM生成(Free-LLM)、LLM重写的人类文本(H2L)以及同一生成器重写的LLM文本(LLM2L)。我们在严格的1%假阳性操作点(TPR@1%FPR)上评估了五个检测器(FastDetectGPT、Binoculars-falcon-7b、RADAR、BERT-Defense、RoBERTa-Defense)。FastDetectGPT和Binoculars-falcon-7b检测到了91.2%和93.5%的直接LLM文本,但仅检测到30.8%和15.1%被LLM重写的人类文本,下降幅度为60-78个百分点。当LLM文本由同一模型重写时,这些检测器的召回率保持在78.3%和83.0%,下降幅度仅为10-13个百分点。RADAR遵循相同的模式(66.8%降至12.2%),而BERT-Defense和RoBERTa-Defense在所有情况下的召回率均低于3%。这些结果表明,在传统的人类与LLM基准上测得的检测器性能并不能转移到被LLM修订的人类撰写文本,尽管相同的检测器在LLM单独重写时仍然保持相对稳健。
cs.CL / 47 / 2607.29585

Sycophancy Undermines Epistemic Vigilance in Cooperative Vision-Language Tasks

谄媚行为削弱了合作视觉-语言任务中的认知警觉性
Sarkar, Rupak, Srikanth, Neha, Gupta, Saloni, Bonial, Claire, Resnik, Philip, Rudinger, Rachel
Abstract
To maintain common ground in cooperative conversation, humans iteratively update their beliefs as conversation participants share new information; participants who are epistemically vigilant detect when new information conflicts with prior beliefs and take steps to repair these conflicts. In order for AI systems to serve as reliable partners in complex cooperative tasks, they must similarly weigh incoming information against their own private evidence and shared context and appropriately surface inconsistencies when they arise. To measure the epistemic vigilance of vision-language models in cooperative settings, we present an information-asymmetric, dialog-based "spot-the-difference" task. Two models are privately shown one image each, and must determine through conversation whether the images are identical or, if not, identify the difference. Models routinely fail at this: they frequently overlook key evidence in their private image in favor of agreeing with their conversational partner, even when their agreement is unwarranted. We relate these violations of epistemic vigilance to the broader behavior of sycophancy, which manifests itself in cooperative goal-oriented dialog as over-accommodation and weak evidential grounding. Our results show that model steering to reduce sycophancy with a vector learned from task-agnostic sycophancy examples can reduce epistemic vigilance-related errors, making models more faithful reporters of their evidence, and in turn, more reliable partners in information-asymmetric cooperative tasks.
Chinese Translation
为了在合作对话中维持共同基础,人类会随着对话参与者分享新信息而迭代更新自己的信念;具有认知警觉性的参与者能够检测到新信息与先前信念之间的冲突,并采取措施修复这些冲突。为了使人工智能系统能够在复杂的合作任务中作为可靠的伙伴,它们也必须将接收到的信息与自身的私有证据和共享背景进行权衡,并在出现不一致时适当地提出。为了测量视觉-语言模型在合作环境中的认知警觉性,我们提出了一种信息不对称的基于对话的“找不同”任务。两个模型各自私下查看一幅图像,并必须通过对话确定这些图像是否相同,或者如果不同,则识别出差异。模型在这方面经常失败:它们常常忽视私有图像中的关键证据,而倾向于与对话伙伴达成一致,即使这种一致性并不合理。我们将这些认知警觉性违规行为与谄媚行为的更广泛表现联系起来,谄媚行为在合作目标导向的对话中表现为过度适应和薄弱的证据基础。我们的结果表明,通过使用从任务无关的谄媚示例中学习的向量来引导模型以减少谄媚行为,可以减少与认知警觉性相关的错误,使模型更忠实地报告其证据,从而在信息不对称的合作任务中成为更可靠的伙伴。
cs.CL / 48 / 2607.29591

ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression

ResKV:重构固定预算 KV 缓存压缩中被省略的注意力贡献
Zhan, Yuhang, Chen, Lisi, Shang, Shuo
Abstract
KV cache compression is essential for efficient long-context inference. Existing eviction methods permanently discard unselected tokens and consequently remove their aggregate contribution to attention. Merging-based alternatives preserve more information but can perturb retained keys and values that should remain exact. We observe that the information omitted by cache eviction can be formulated as residual statistics in both the numerator and denominator of softmax attention. Based on this observation, we propose ResKV, which divides a fixed KV budget into an exact main cache and a compact residual cache that reconstructs the contribution of omitted tokens. ResKV lets main-cache tokens and residual entries participate in the same softmax normalization, so residual entries restore both attention numerator and denominator mass rather than acting as a post-hoc correction. A construction-time validation proxy determines residual allocation for each layer and KV head, while a decode-time dynamic gate adjusts residual contributions for individual queries. Comprehensive evaluations on LongBench and RULER, covering query-aware and query-agnostic settings, multiple backbones, cache budgets, and representative compression baselines, demonstrate broad improvements under the same retained KV budget while preserving the practical efficiency of compressed decoding, including peak memory usage and long-context decode throughput.
Chinese Translation
KV 缓存压缩对于高效的长上下文推理至关重要。现有的驱逐方法会永久性地丢弃未选择的标记,从而消除了它们对注意力的整体贡献。基于合并的方法虽然保留了更多信息,但可能会扰动应保持精确的键和值。我们观察到,缓存驱逐所省略的信息可以在 softmax 注意力的分子和分母中被表述为残差统计。基于这一观察,我们提出了 ResKV,它将固定的 KV 预算划分为一个精确的主缓存和一个紧凑的残差缓存,以重构被省略标记的贡献。ResKV 使主缓存标记和残差条目参与同一 softmax 归一化,因此残差条目恢复了注意力分子和分母的质量,而不是作为事后修正。构建时的验证代理确定每层和 KV 头的残差分配,而解码时的动态门控则为单个查询调整残差贡献。在 LongBench 和 RULER 上进行的全面评估,涵盖了查询感知和查询无关的设置、多种骨干网络、缓存预算和代表性的压缩基线,展示了在相同的保留 KV 预算下的广泛改进,同时保持了压缩解码的实际效率,包括峰值内存使用和长上下文解码吞吐量。
cs.CL / 49 / 2607.29602

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

FriendBench:人类与多模态大型语言模型中的二人熟悉度推断基准测试
Girard, Jeffrey M., Zheng, Jason Z., Vertino, Jacqueline R., D'Avirro, Antony, Peloquin, Benjamin
Abstract
Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models lean toward "stranger"---a difference in effective prior, not discrimination. Richer channels help both unequally, and only humans gain from visible behavior on top of speech. We release the stimuli, human ratings, and model predictions.
Chinese Translation
理解社交情境往往依赖于行为,而不仅仅是言辞。我们介绍了FriendBench,一个用于推断两个人是否已经熟悉或作为陌生人见面的基准测试,该测试基于一段20秒的二人破冰对话视频。每对参与者回答相同类型的提示,因此只有互动方式能够揭示答案。在文本、音频和视频三种模态中,我们将来自七家公司的26个模型与96对平衡的二人组的匹配人类评审进行比较。在每种模态下,最佳模型与人类群体在准确性上统计上无显著差异,但达到这一准确性的方式不同:人类在两个答案之间保持平衡,而最强的模型则倾向于选择“陌生人”——这是一种有效先验的差异,而非区分能力的差异。更丰富的渠道对两者的帮助不均,且只有人类在言语之外的可见行为上获得额外收益。我们发布了刺激材料、人类评分和模型预测。
cs.CL / 50 / 2607.29642

Evolving language compositionality in a frequency-structured meaning space

在频率结构意义空间中演化的语言组合性
De Ponte, Fabio, Gaines-White, Eloise, Houghton, Conor, Bullock, Seth
Abstract
The iterated learning model was introduced to investigate language evolution: the way in which the characteristic properties of human languages have been shaped, at least partly, by repeated transmission from one language user to another. The key finding is that language compositionality can arise spontaneously as a consequence of language being passed repeatedly through a language learning bottleneck. Here we explore how changing the frequency of different meanings, so that some meanings occur much more frequently than others, affects the character of its compositionality. We find that, as observed in natural languages, high-frequency meanings can escape the pressure to conform to the grammar that characterizes lower-frequency meanings. However, when the frequency structure is instead imposed on parts rather than on whole meaning vectors, the language fails to transmit across generations. This occurs despite the fact that the most frequent elements are reliably learned. These results suggest that frequency can shape emergent linguistic structure only when the frequency distribution is defined over form-meaning units that learners can acquire holistically. When frequency is instead distributed over smaller units, it fails to support the relational structure required for compositional generalisation, thereby preventing stable language transmission.
Chinese Translation
迭代学习模型被引入以研究语言演化:人类语言的特征属性如何在一定程度上通过语言使用者之间的重复传递而形成。关键发现是,语言组合性可以自发产生,作为语言通过语言学习瓶颈反复传递的结果。在这里,我们探讨了改变不同意义的频率,使某些意义的出现频率远高于其他意义,如何影响其组合性的特征。我们发现,正如在自然语言中观察到的那样,高频意义可以逃避对低频意义所特征化的语法的遵循压力。然而,当频率结构施加在部分而非整体意义向量上时,语言在代际间的传递失败。尽管最频繁的元素被可靠地学习,这种情况仍然发生。这些结果表明,频率只能在定义为学习者可以整体获取的形式-意义单位上的频率分布时,塑造新兴的语言结构。当频率分布在更小的单位上时,它未能支持组合泛化所需的关系结构,从而阻碍了稳定的语言传递。
cs.CL / 51 / 2607.29678

TokTier: Exact Stateful Tokenization for Agentic LLM Serving

TokTier:用于代理型大语言模型服务的精确状态化标记化
Zhang, Zhenyu, Cao, Zhichao
Abstract
LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, which resubmit a long transcript after each small tool result, and reuse is hard because even a short append can change token boundaries near the end of the previous sequence. Across 153,951 calls from two agent ecosystems, the median call appends about 1.4K characters, and only 1.0-3.6% of calls start or rebuild a session with contexts of millions of characters. At a 94.1% fleet prompt-cache hit rate, tokenization reaches up to 64% of time to first token. TokTier is a stateful tokenization service with one contract: emitted token IDs are always identical to full reference tokenization of the request text. For a session continuation, it re-tokenizes a small window around the append and splices only after a per-request stable-boundary check, widening the window or falling back to full tokenization on failure. For a call without a reusable prefix, it decomposes GPT-family regex pre-tokenization into run-local rules and runs exact pre-tokenization and BPE on a GPU. A sampled shadow verifier re-checks live traffic. Across 17 tokenizer families, differential campaigns cover 1.5x10^10 split checks, a 12.4 TB real-text corpus, and 93,000+ replayed agent steps, with zero divergence. Incremental repair takes 0.5-1.1 ms from 100K to 3M characters, up to 437x faster than HF tokenization and 2.1x faster at 1M than the strongest cache-based baseline (Gigatoken) fully prewarmed. GPU full tokenization encodes a 1M-character request in 0.87 ms, up to 491x below HF and 23.4x below the fastest published CPU method. With vLLM, median time to first token drops 16-34% and P99 drops 23% under recorded bursts. Under a 50 ms P99 objective, four repair cores plus one GPU sustain 1,821 requests/s where a 16-core stateless front end saturates at 40.
Chinese Translation
大语言模型(LLM)服务系统缓存提示键值状态,但大多数前端在每次调用时仍然重新标记完整的请求文本。这种成本落在编码代理上,代理在每次小工具结果后重新提交长文本记录,而由于即使是短的附加内容也可能改变前一个序列末尾的标记边界,因此重用变得困难。在来自两个代理生态系统的153,951次调用中,中位数调用附加约1.4K字符,且仅有1.0-3.6%的调用以数百万字符的上下文开始或重建会话。在94.1%的舰队提示缓存命中率下,标记化的时间占到首次标记时间的64%。TokTier是一个状态化标记化服务,其唯一契约是:发出的标记ID始终与请求文本的完整参考标记化相同。对于会话延续,它在附加内容周围重新标记一个小窗口,并仅在每次请求的稳定边界检查后进行拼接,在失败时扩大窗口或回退到完整标记化。对于没有可重用前缀的调用,它将GPT系列的正则预标记化分解为本地运行规则,并在GPU上执行精确的预标记化和BPE。一个采样的影子验证器重新检查实时流量。在17个标记器家族中,差异化活动覆盖了1.5x10^10个分割检查、12.4 TB的真实文本语料库和93,000多个重放的代理步骤,且没有出现任何偏差。增量修复从100K到3M字符需要0.5-1.1毫秒,比HF标记化快最多437倍,并且在1M字符时比最强的基于缓存的基线(Gigatoken)快2.1倍,后者完全预热。GPU全标记化在0.87毫秒内编码1M字符请求,比HF快491倍,比最快的已发布CPU方法快23.4倍。使用vLLM时,首次标记的中位时间减少16-34%,P99在记录的突发下下降23%。在50毫秒的P99目标下,四个修复核心加一个GPU维持1,821个请求/秒,而16核无状态前端的饱和点为40。