cs.RO / 1 / 2608.06434
Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection
快速而准确:通过环境感知模型选择的自适应视觉-语言-动作推理框架
Abstract
Embodied intelligence demands both long-horizon reasoning and real-time closed-loop responsiveness. Recent dual-system Vision-Language-Action (VLA) architectures combine fast reactive control with slow deliberative reasoning to balance inference speed and task success rate. However, existing dual-process VLAs tightly couple the fast module to intermediate representations of the slow module, necessitating end-to-end joint training and limiting modularity, extensibility and flexible system switching. In this paper, we propose Environment-aware Model Selection (EMS), an adaptive VLA inference framework that switches between two fully decoupled systems of different scales through environment-aware model selection. The large-scale deliberative system provides globally consistent trajectory planning to ensure task success, while a lightweight reactive system enables high-frequency closed-loop control. A reinforcement-learning-based switching policy dynamically selects which system to invoke based on real-time feedback, enabling sparse use of the slow system and thereby balancing pretrained knowledge utilisation with runtime efficiency. Our design offers three key advantages over prior hierarchical VLA frameworks: (1) a fully decoupled and modular dual-system architecture that supports plug-and-play model replacement; (2) an adaptive, environment-aware switching strategy; (3) high-frequency inference for responsive closed-loop control. We extensively evaluate EMS in both simulation and real-world environments. On the LIBERO benchmark, EMS achieves success rates comparable to the large-scale baseline while increasing the effective action frequency to 93.4 Hz. The framework further demonstrates strong extensibility in real-world dual-arm manipulation tasks, where it accelerates task completion while maintaining robust performance.
Chinese Translation
具身智能既需要长时间的推理能力,又需要实时的闭环响应能力。最近的双系统视觉-语言-动作(VLA)架构结合了快速的反应控制与缓慢的深思熟虑推理,以平衡推理速度和任务成功率。然而,现有的双过程VLA紧密耦合了快速模块与慢速模块的中间表示,迫使进行端到端的联合训练,并限制了模块化、可扩展性和灵活的系统切换。在本文中,我们提出了环境感知模型选择(Environment-aware Model Selection, EMS),这是一种自适应的VLA推理框架,通过环境感知模型选择在两个完全解耦的不同规模系统之间切换。大规模的深思熟虑系统提供全局一致的轨迹规划,以确保任务成功,而轻量级的反应系统则实现高频闭环控制。基于强化学习的切换策略动态选择调用哪个系统,基于实时反馈,使得稀疏使用慢速系统,从而平衡预训练知识的利用与运行时效率。我们的设计相较于先前的层次VLA框架具有三个关键优势:(1)完全解耦且模块化的双系统架构,支持即插即用的模型替换;(2)自适应的环境感知切换策略;(3)高频推理以实现响应式闭环控制。我们在模拟和现实环境中广泛评估EMS。在LIBERO基准测试中,EMS的成功率与大规模基线相当,同时将有效动作频率提高至93.4 Hz。该框架在现实世界的双臂操作任务中进一步展示了强大的可扩展性,加速了任务完成,同时保持了稳健的性能。
cs.RO / 2 / 2608.06481
LyEvO: Lyapunov-Guided Evolutionary Optimization for Safe and Robust Sim-to-Real Policy Learning
LyEvO:基于李雅普诺夫指导的进化优化用于安全和稳健的模拟到现实策略学习
Abstract
Training controllers that are safe and robust in simulation, and systematically assessing their readiness for real-world deployment, remain key challenges in sim-to-real transfer. To address this, we propose LyEvO, a physics-grounded framework that combines constrained Evolutionary Optimization and Statistical Model Checking (SMC)-based verification with Lyapunov-based stability analysis. Leveraging prior knowledge of the system dynamics, LyEvO uses Lyapunov analysis to compute an initial candidate stability region. An iterative loop then uses operational scenarios drawn from this region to jointly optimize and statistically verify a policy, and subsequently expands the region's boundaries based on the verification outcome. This integrated procedure provides a practical criterion for assessing deployment readiness. We evaluate LyEvO on Cartpole and 3D Quadrotor benchmarks through extensive simulations and targeted real-world experiments, demonstrating safe and robust sim-to-real transfer.
Chinese Translation
在模拟中训练安全且稳健的控制器,并系统性地评估其在现实世界部署的准备情况,仍然是模拟到现实转移中的关键挑战。为了解决这个问题,我们提出了LyEvO,一个基于物理的框架,结合了受限进化优化和基于统计模型检查(SMC)的验证,以及李雅普诺夫稳定性分析。LyEvO利用系统动态的先验知识,通过李雅普诺夫分析计算初始候选稳定区域。然后,迭代循环使用从该区域提取的操作场景来联合优化和统计验证策略,并根据验证结果扩展区域的边界。这个集成程序为评估部署准备情况提供了一个实用标准。我们通过广泛的模拟和针对性的现实世界实验,在Cartpole和3D四旋翼基准上评估了LyEvO,展示了安全且稳健的模拟到现实转移。
cs.RO / 3 / 2608.06488
A Disturbance in the Force: Force Actuation on the RAVEN II Surgical Robot with Parallel Motor-Cable Units
力量的干扰:RAVEN II 外科机器人上的并联电机-电缆单元的力驱动
Abstract
Difficulty in haptic feedback for surgical robots has been a long-term problem for decades. In recent years, learning-based force estimation from robot states suggests desirable accuracy without the necessity of extra sensors. However, challenges remain in obtaining representative training data in which the robot moves in the workspace under various external forces. In this work, a parallel motor-cable system is developed. With six motor-cable units installed around the robot workspace, cables with controllable tension connected to the robot end-effector can provide the desired external force without interfering with the movement of the surgical robot. The development of the system includes motor-unit hardware, control software, sensor drivers, simulations, and more. Preliminary experiments suggest an accuracy of force actuation with errors less than 1 N.
Chinese Translation
外科机器人触觉反馈的困难已成为一个长期存在的问题。近年来,基于学习的力估计方法从机器人状态中提取出期望的准确性,而无需额外的传感器。然而,在获取代表性的训练数据方面仍然面临挑战,这些数据要求机器人在各种外部力作用下在工作空间中移动。在本研究中,开发了一种并联电机-电缆系统。通过在机器人工作空间周围安装六个电机-电缆单元,连接到机器人末端执行器的可控张力电缆可以提供所需的外部力,而不会干扰外科机器人的运动。该系统的开发包括电机单元硬件、控制软件、传感器驱动程序、仿真等。初步实验表明,力驱动的准确性误差小于 1 N。
cs.RO / 4 / 2608.06587
SyncSBC: Decentralized Swarm Behavior Prediction for Synchronized Autonomous Control
SyncSBC:用于同步自主控制的去中心化群体行为预测
Abstract
Robot swarms utilize many independent limited-sensing agents to produce complex emergent behaviors without requiring centralized control. However, little research explores how agents can infer swarm-level behavior from purely local perception, a capability critical for detecting faults and behavior changes. In this paper, we introduce Synchronized Swarm Behavior Classification (SyncSBC), which combines improvements in machine learning and distributed consensus to classify collective swarm behavior and synchronize swarm decision-making in an entirely decentralized manner. We show that SyncSBC achieves high classification accuracy and low synchronization delay, making it suitable for real-world deployment. Finally, we use SyncSBC to demonstrate two promising swarm applications on real robots where we show that swarms utilizing SyncSBC can accurately identify anomalies in robot behavior and autonomously coordinate collective changes in swarm behavior. Videos, code and supplemental experiments are available at https://sites.google.com/view/sync-sbc/home.
Chinese Translation
机器人群体利用许多独立的有限感知代理来产生复杂的涌现行为,而无需集中控制。然而,目前的研究很少探讨代理如何仅通过局部感知推断群体级行为,这一能力对于检测故障和行为变化至关重要。在本文中,我们介绍了同步群体行为分类(SyncSBC),该方法结合了机器学习和分布式共识的改进,以完全去中心化的方式对集体群体行为进行分类并同步群体决策。我们展示了SyncSBC实现了高分类准确率和低同步延迟,使其适合于实际应用。最后,我们使用SyncSBC在真实机器人上展示了两个有前景的群体应用,证明了利用SyncSBC的群体能够准确识别机器人行为中的异常,并自主协调群体行为的集体变化。视频、代码和补充实验可在 https://sites.google.com/view/sync-sbc/home 获取。
cs.RO / 5 / 2608.06648
Plan-and-Avoid: Real-Time Aircraft Trajectory Coordination in a Multi-Agent Environment
计划与避免:多智能体环境中的实时飞机轨迹协调
Abstract
This paper presents a real-time Plan-and-Avoid (PAA framework for coordinating cooperative multi-agent airspace operations around a declared priority trajectory. The priority trajectory represents an aircraft flight plan that must be preserved because of constrained maneuverability, an emergency, a mission-critical task, or assigned operational priority. The framework predicts uncertainty-aware, well-clear separation violations with surrounding traffic and, when the priority plan alone cannot maintain separation, generates vehicle-constrained unilateral advisories that modify nearby aircraft trajectories to maintain well-clear separation for all traffic. The approach is applicable to any declared priority trajectory. This paper demonstrates the Plan component using a contingency landing planner to generate candidate priority trajectories. PAA then identifies nearby aircraft passing too close to this priority trajectory and issues Avoid resolution advisories to these aircraft. The framework is tested using real-world Automatic Dependent Surveillance-Broadcast (ADS-B) traffic from the Washington, D.C., airspace across more than 900 forced-landing cases, totaling over 140 hours of simulated flight. The PAA framework generates feasible cooperative advisories for all 575 unique conflict encounters, with a worst-case end-to-end response time of 5.7 s on a personal computer, including priority trajectory planning, advisory generation, and 1 s two-way datalink delay. In total, 93.5% of generated advisories satisfy the 35 s RTCA DO-365 Detect-and-Avoid temporal threshold. These results demonstrate low-latency coordination for preserving priority trajectories while maintaining well-clear separation through real-time automated advisory generation. Future work will quantify advisory-induced delays and their operational impacts.
Chinese Translation
本文提出了一种实时的计划与避免(Plan-and-Avoid, PAA)框架,用于协调围绕已声明优先轨迹的合作多智能体空域操作。优先轨迹代表了由于机动性受限、紧急情况、任务关键性任务或分配的操作优先级而必须保持的飞机飞行计划。该框架预测与周围交通的基于不确定性的安全距离违规情况,当仅依靠优先计划无法维持安全距离时,生成车辆受限的单方面建议,修改附近飞机的轨迹,以确保所有交通的安全距离。该方法适用于任何已声明的优先轨迹。本文展示了使用应急着陆规划器的计划组件,以生成候选优先轨迹。PAA随后识别出与该优先轨迹过于接近的附近飞机,并向这些飞机发出避免解决建议。该框架使用来自华盛顿特区空域的真实自动依赖监视广播(Automatic Dependent Surveillance-Broadcast, ADS-B)交通进行了测试,涵盖了900多个强制着陆案例,总计超过140小时的模拟飞行。PAA框架为所有575个独特的冲突遭遇生成了可行的合作建议,最坏情况下的端到端响应时间为5.7秒,包含优先轨迹规划、建议生成和1秒的双向数据链路延迟。总共93.5%的生成建议满足35秒RTCA DO-365检测与避免时间阈值。这些结果展示了在实时自动建议生成的过程中,低延迟协调以保持优先轨迹,同时维持安全距离。未来的工作将量化建议引起的延迟及其操作影响。
cs.RO / 6 / 2608.06650
SoRoMoX: Fast, Differentiable, and Parallelizable Soft Robot Models
SoRoMoX:快速、可微分且可并行化的软机器人模型
Abstract
Reduced-order models based on Cosserat-rod theory are now well established, and modeling theory is no longer the primary bottleneck in soft-robot control. Their implementations, however, do not support the differentiable, GPU-parallel, and control-oriented workflows that underpin advanced rigid-robotics applications. Here, we fill this gap with SoRoMoX (Soft Robot Models in JAX), a fully numerical, JIT-compilable Python/JAX framework. SoRoMoX implements articulated, Piecewise Constant Strain, and Variable Strain models through a unified, control-ready interface that provides inertia matrices, gravitational and elastic forces, Jacobians, and their derivatives. To our knowledge, it is the first rod/strain-based soft-robot modeling framework that runs directly on GPUs and is end-to-end differentiable with respect to states, inputs, and parameters. Sequential CPU rollouts are up to 18.1x faster than state-of-the-art alternatives, while GPU-parallel rollouts increase throughput by up to 234.6x. This performance enables workflows that were previously impractical or impossible: static-equilibrium system identification with 66% lower marker RMSE; residual-force learning with a further 64% reduction; computed-torque tracking with RMSE reduced by a factor of approximately 500 relative to model-free PD; control-gain optimization with up to 62% lower loss than untuned gains; safety-constrained control using high-order control barrier functions to keep the peak contact force within a prescribed 5 N bound, compared with 33.5 N without the safety constraint; and reinforcement-learning policy training up to 7x faster than a CPU PyElastica discrete-rod baseline through massively parallel rollouts.
Chinese Translation
基于科塞拉特杆理论的降阶模型现已得到广泛应用,建模理论不再是软机器人控制的主要瓶颈。然而,它们的实现并不支持支撑先进刚性机器人应用的可微分、GPU并行和控制导向的工作流程。在此,我们通过SoRoMoX(Soft Robot Models in JAX)填补了这一空白,SoRoMoX是一个完全数值化、可即时编译的Python/JAX框架。SoRoMoX通过统一的、控制就绪的接口实现了关节、分段常量应变和可变应变模型,提供了惯性矩阵、重力和弹性力、雅可比矩阵及其导数。据我们所知,它是第一个能够直接在GPU上运行的杆/应变基础的软机器人建模框架,并且在状态、输入和参数方面是端到端可微分的。顺序CPU展开的速度比最先进的替代方案快多达18.1倍,而GPU并行展开的吞吐量则提高了多达234.6倍。这种性能使得以前不切实际或不可能的工作流程成为可能:静态平衡系统识别的标记均方根误差降低了66%;残余力学习进一步减少了64%;计算扭矩跟踪的均方根误差相对于无模型PD减少了约500倍;控制增益优化的损失比未调优增益低62%;使用高阶控制障碍函数进行安全约束控制,将峰值接触力保持在规定的5 N范围内,而在没有安全约束的情况下为33.5 N;通过大规模并行展开,强化学习策略训练速度比CPU PyElastica离散杆基线快多达7倍。
cs.RO / 7 / 2608.06688
CrossTracer: Cross-Embodiment Navigation via VLA Model Reasoning and Trace Residuals Adapting
CrossTracer:通过 VLA 模型推理和轨迹残差自适应实现跨体态导航
Abstract
Vision-language-action (VLA) models provide strong semantic priors for robot navigation, but they often ignore embodiment-specific mobility constraints. A path that is semantically plausible for one robot may be physically infeasible for another. We propose CrossTracer, a hierarchical framework for cross-embodiment navigation through adaptive trace residuals. CrossTracer represents navigation plans as normalized image-plane waypoints, forming a unified pixel-space interface between semantic reasoning and physical grounding. First, Vision-Language Trace Proposer (VL-Tracer) adapts a pretrained VLA model to predict an initial navigation trace from egocentric observations and flexible goal specifications. Second, CE-Adapter refines this trace by predicting embodiment-conditioned residual corrections from visual traversability cues, robot identity, and the initial trace. To train the refinement module without costly manual annotation, Cross-Embodiment RRT* (CE-RRT*) converts panoptic segmentation into robot-conditioned traversability cost maps and generates cost-minimizing pixel-space traces. We evaluate CrossTracer on the NaviTrace benchmark, which tests whether a model can generate embodiment-consistent navigation traces from egocentric observations, language instructions, and robot embodiment types. CrossTracer achieves a total score of 45.68, outperforming the strongest evaluated general-purpose baseline, Gemini-2.5-Pro, by 10.01 points, corresponding to a 28.1% relative improvement. Real-world deployment on wheeled and legged robots further shows improved navigation success and execution efficiency.
Chinese Translation
视觉-语言-行动(VLA)模型为机器人导航提供了强大的语义先验,但它们往往忽视了特定于体态的移动约束。一条对某个机器人在语义上可行的路径,可能对另一个机器人在物理上不可行。我们提出了 CrossTracer,这是一个通过自适应轨迹残差实现跨体态导航的分层框架。CrossTracer 将导航计划表示为归一化的图像平面航点,形成语义推理与物理基础之间的统一像素空间接口。首先,视觉-语言轨迹提议器(VL-Tracer)适应预训练的 VLA 模型,从自我中心的观察和灵活的目标规范中预测初始导航轨迹。其次,CE-Adapter 通过从视觉可遍历性线索、机器人身份和初始轨迹中预测体态条件的残差修正来优化这一轨迹。为了在没有高成本人工标注的情况下训练修正模块,跨体态 RRT*(CE-RRT*)将全景分割转换为机器人条件的可遍历性成本图,并生成成本最小化的像素空间轨迹。我们在 NaviTrace 基准上评估了 CrossTracer,该基准测试模型是否能够从自我中心的观察、语言指令和机器人体态类型生成一致的导航轨迹。CrossTracer 的总得分为 45.68,超过了最强的评估通用基线 Gemini-2.5-Pro,提升了 10.01 分,相当于 28.1% 的相对改善。在轮式和腿式机器人上的实际部署进一步显示了导航成功率和执行效率的提升。
cs.RO / 8 / 2608.06707
Hoverflie: An empirical investigation of rotor shrouds to transform micro air vehicles into multi-modal hovercraft
Hoverflie:对转子罩的实证研究,以将微型飞行器转变为多模态气垫船
Abstract
Small rotorcraft intended for use indoors or around the built environment have extremely limited flight duration. This paper presents the design and experimental characterization of a custom shroud system that transforms a Crazyflie 2.1 micro air vehicle into a multi-modal robot capable of operating as a high-efficiency hovercraft or a free-flying drone. A custom experimental platform was developed for precise control of hover height and rotor duty cycle, and automated data logging of lift forces. Parametric testing of duct, intake, and nozzle geometries was performed to investigate the impact of shroud configuration on in-ground-effect and free-flight performance. An empirical model is developed which, unlike typical models for ground effect in rotorcraft, captures the suckdown effect that reduces force at intermediate height. It is shown that, through proper design of the shroud, beneficial ground effects can be increased while diminishing negative effects both close to the ground and in free flight. An optimized configuration exhibited nearly three times higher in-ground-effect force while maintaining comparable out-of-ground-effect aerodynamic thrust, although the added shroud mass reduces free-flight control authority. Lightweight shrouds are manufactured using thin-film thermoformed components, and total single-charge flight time is shown to increase by 60% in-ground-effect while decreasing by only 30% in free-flight as compared to the stock drone. Finally, controlled flight in the air, hovering close to the ground, and hover-to-flight transitions are demonstrated using a simple mode-switching controller, with tracking errors reported to quantify performance. This work provides an experimentally-validated and easily adoptable foundation for future research into lightweight ground-effect vehicles and hybrid drone-hovercraft systems.
Chinese Translation
用于室内或建筑环境的小型旋翼机的飞行时间极为有限。本文介绍了一种定制罩系统的设计与实验表征,该系统将Crazyflie 2.1微型飞行器转变为一种多模态机器人,能够作为高效气垫船或自由飞行无人机进行操作。为实现对悬停高度和旋翼工作周期的精确控制,以及提升力的自动数据记录,开发了一种定制实验平台。对导管、进气口和喷嘴几何形状进行了参数测试,以研究罩配置对地面效应和自由飞行性能的影响。本文开发了一种实证模型,与典型的旋翼机地面效应模型不同,该模型捕捉了在中间高度降低力的吸引效应。研究表明,通过合理设计罩,可以在接近地面和自由飞行时增加有益的地面效应,同时减少负面影响。优化配置的地面效应力几乎提高了三倍,同时保持了可比的非地面效应气动推力,尽管增加的罩质量降低了自由飞行的控制能力。轻量化罩采用薄膜热成型组件制造,结果显示在地面效应下的单次充电飞行时间增加了60%,而在自由飞行中仅减少了30%,与原始无人机相比。最后,使用简单的模式切换控制器演示了在空中控制飞行、接近地面的悬停以及悬停到飞行的过渡,并报告了跟踪误差以量化性能。这项工作为未来轻量化地面效应车辆和混合无人机-气垫船系统的研究提供了实验验证和易于采用的基础。
cs.RO / 9 / 2608.06729
AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models
AtlasVLA:面向视觉-语言-动作模型的持久世界-自我状态建模
Abstract
While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.
Chinese Translation
尽管视觉-语言-动作(VLA)模型在具身人工智能方面取得了进展,但其根本性的反应式范式严重限制了在部分可观察和长时间跨度任务中的表现。当仅限于单个手腕-mounted 相机时,它们不可避免地会因物体离开视野而遭受感知遗忘,并在多步骤执行过程中出现时间任务进展遗忘。为了解决这些瓶颈,我们提出了 AtlasVLA,一个新颖的框架,通过持久的世界-自我状态从直接的反应式操作转变为主动推理。AtlasVLA 具有双重记忆架构:一个 4D 持久世界状态记忆,将瞬态的 2D 观察提升为全球更新的体素哈希空间状态,以解决视觉盲点;另一个自我工作状态记忆则跟踪历史自我状态和任务进展。通过在这个联合的世界-自我状态上对扩散变换器(DiT)进行条件化,AtlasVLA 实现了强大的空间推理。对 LIBERO、RLBench 和真实世界基准的广泛评估表明,AtlasVLA 在仅使用手腕相机的情况下达到了最先进的性能。值得注意的是,它显著超越了多视角基线,在 LIBERO-Long 上提高了 9.4% 的绝对成功率,在真实世界长时间跨度任务中提高了 17.5%。
cs.RO / 10 / 2608.06799
Is Forward Prediction Enough? Physical State Grounding for JEPA World Models
前向预测是否足够?JEPA世界模型的物理状态基础
Abstract
Learning structured and control-relevant latent representations remains a key challenge for world models. Recent JEPA-based world models learn action-conditioned predictive latent dynamics from observation sequences. However, their forward-prediction objectives do not explicitly enforce reliable identifiability of robot-centric physical state from individual latents or state changes from latent pairs, which can limit downstream planning and policy performance. We propose PSG-JEPA, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes. Both objectives are applied only during training, leaving the inference architecture and computational cost unchanged. To comprehensively evaluate PSG-JEPA, we conduct experiments at three levels: (1) latent identifiability via probing, (2) goal-conditioned planning on frozen latents, and (3) policy learning in simulation and on a real robot. Experiments demonstrate that our PSG-JEPA consistently outperforms state-of-the-art latent world-model baselines at all three levels.
Chinese Translation
学习结构化和与控制相关的潜在表示仍然是世界模型面临的关键挑战。基于JEPA的最新世界模型从观察序列中学习基于动作的预测潜在动态。然而,它们的前向预测目标并未明确强制从单个潜在变量或潜在对中的状态变化中可靠地识别以机器人为中心的物理状态,这可能限制下游规划和策略性能。我们提出了PSG-JEPA,这是一种物理基础的JEPA世界模型,它通过两个补充的基础目标来塑造其潜在空间,超越前向预测:将单个潜在变量与机器人本体感知状态相结合,以及将潜在对与多时间步关节角度变化相结合。这两个目标仅在训练期间应用,不改变推理架构和计算成本。为了全面评估PSG-JEPA,我们在三个层面上进行了实验:(1)通过探测评估潜在变量的可识别性,(2)在冻结潜在变量上进行目标条件规划,以及(3)在模拟和真实机器人上进行策略学习。实验表明,我们的PSG-JEPA在所有三个层面上都始终优于最先进的潜在世界模型基线。
cs.RO / 11 / 2608.06827
R2S-EGO: Dual-Proxy Refinement for Sparse-Capture Real-to-Sim
R2S-EGO:稀疏捕捉的真实到仿真双代理精炼
Abstract
Real-to-sim (R2S) depends on scene representations that render observations along robot ego trajectories, yet dense multi-view capture limits per-environment real-image capture-count efficiency, and sparse human capture can leave behavior-scoped robot views under-supported. Camera-controlled synthesis can fill missing views, but its use in R2S requires behavior-admissible queries and capture-anchored structural conditioning. We present R2S-EGO, which couples a simulator-derived robot proxy that represents the behavior-scoped executable query domain with a capture-anchored geometry proxy that supplies scene-specific structural conditions. Within this domain, fixed- budget selection targets current support deficits for which geometry support is available. The generated observations are assimilated as pseudo-observations to refine the visual asset, while real captures remain anchors. The fused geometry proxy also supplies the scene collision surface, which is refreshed between rounds. Together, these updates refine the existing simulation scene while its robot dynamics and control stack stay fixed. Across 48 frozen Unitree G1 ego views in three Replica scenes, six-view R2S-EGO reaches 19.062 dB PSNR, compared with 14.226 dB for the strongest reported R2S baseline. Across five paired policy-training seeds, R2S-EGO achieves 82.5% +/- 6.8% real-G1 sitting success, compared with 10.0% +/- 10.5% for GaussGym.
Chinese Translation
真实到仿真(R2S)依赖于沿机器人自我轨迹渲染观察的场景表示,然而,密集的多视角捕捉限制了每个环境中真实图像捕捉的效率,而稀疏的人类捕捉可能导致行为范围内的机器人视图支持不足。相机控制的合成可以填补缺失的视图,但在R2S中的应用需要符合行为的查询和以捕捉为锚的结构条件。我们提出了R2S-EGO,它结合了一个由模拟器派生的机器人代理,代表行为范围内可执行查询域,以及一个以捕捉为锚的几何代理,提供场景特定的结构条件。在该域内,固定预算选择针对当前支持不足的目标,其中几何支持是可用的。生成的观察被同化为伪观察,以精炼视觉资产,而真实捕捉则保持为锚点。融合的几何代理还提供场景碰撞表面,该表面在回合之间进行刷新。这些更新共同精炼现有的仿真场景,同时其机器人动态和控制堆栈保持不变。在三个Replica场景中的48个冻结Unitree G1自我视图中,六视图R2S-EGO达到了19.062 dB的PSNR,而最强报告的R2S基线为14.226 dB。在五个配对的策略训练种子中,R2S-EGO实现了82.5% +/- 6.8%的真实G1坐姿成功率,而GaussGym为10.0% +/- 10.5%。
cs.RO / 12 / 2608.06830
When Coordination Becomes a Threat: Communication Attacks in LLM-Controlled Multi-Robot Systems
当协调成为威胁:LLM控制的多机器人系统中的通信攻击
Abstract
Large Language Models (LLMs) are increasingly used as high-level planners in embodied multi-robot systems, enabling robots to interpret natural language instructions and coordinate executable actions. Yet, this growing reliance on LLM planners also raises security concerns. Prior work has focused mainly on individual robots, while communication risks in multi-robot collaboration remain insufficiently understood. Existing multi-robot studies are further limited to preliminary analysis under the Decentralized Multi-agent System (DMAS) architecture, so it remains unclear whether these risks persist across other common communication architectures and how attacker access settings shape their propagation. To fill this gap, we formulate two communication attacks corresponding to distinct attacker access settings: the External Entry Point Attack and the Privileged In-System Attack. We evaluate both attacks across DMAS, HMAS-1, and HMAS-2 using three LLMs and five embodied multi-robot tasks. Results show that unsafe information can turn into unsafe actions across all three architectures: DMAS reaches a 96.7\% entry endorsement rate and a 100\% post endorsement activation rate, HMAS-1 reaches a 97.8\% unsafe action success rate, and HMAS-2 triggers 88.3\% of task defined unsafe action slots. To mitigate risks from trusted information flow, we introduce the Claim Provenance and Verification (CPV) Gate, which verifies communicated claims before downstream reuse and reduces the violation rate from 70.0\% to 36.6\%.
Chinese Translation
大型语言模型(LLMs)越来越多地被用作具身多机器人系统中的高级规划者,使机器人能够理解自然语言指令并协调可执行的动作。然而,这种对LLM规划者日益依赖也引发了安全隐患。以往的研究主要集中在单个机器人上,而多机器人协作中的通信风险仍然未得到充分理解。现有的多机器人研究进一步局限于去中心化多智能体系统(DMAS)架构下的初步分析,因此尚不清楚这些风险是否在其他常见通信架构中持续存在,以及攻击者的访问设置如何影响其传播。为填补这一空白,我们制定了两种通信攻击,分别对应不同的攻击者访问设置:外部入口点攻击(External Entry Point Attack)和特权内部系统攻击(Privileged In-System Attack)。我们使用三种LLM和五个具身多机器人任务对这两种攻击在DMAS、HMAS-1和HMAS-2架构下进行了评估。结果表明,不安全的信息可以在所有三种架构中转化为不安全的行动:DMAS的入口认可率达到96.7%,后认可激活率为100%;HMAS-1的不安全行动成功率达到97.8%;HMAS-2触发了88.3%的任务定义的不安全行动槽。为了减轻来自可信信息流的风险,我们引入了声明来源和验证(Claim Provenance and Verification, CPV)门,该门在下游重用之前验证传达的声明,并将违规率从70.0%降低到36.6%。
cs.RO / 13 / 2608.06833
Unordered Landmark Visual Navigation
无序地标视觉导航
Abstract
Image-goal navigation is a fundamental capability for embodied AI, yet its practical deployment is strained by strong prior assumptions. Existing methods predominantly rely on temporally ordered video streams or auxiliary sensors (e.g., depth, LiDAR) to maintain spatial consistency. These sequential and multimodal dependencies severely restrict scalability, especially when deploying robots using crowd-sourced or pre-recorded unordered image collections. When temporal priors are removed, current methods struggle with severe perceptual aliasing, noisy associations, and catastrophic mapping failures. To address this underexplored challenge, we propose Unordered Landmark Visual Navigation (ULVN), a unified RGB-only framework free from temporal and odometric priors. ULVN systematically mitigates error accumulation by integrating mapping, localization, and planning. Specifically, it constructs a robust 2D topological map directly from unstructured images via calibrated geometric verification and maximum spanning forest refinement. For closed-loop execution, ULVN abandons sequential heuristics, utilizing a graph-based belief propagation filter with entropy-adaptive fusion for global localization and dynamic subgoal planning. Extensive experiments in simulation and real-world deployments demonstrate that ULVN significantly outperforms state-of-the-art methods.
Chinese Translation
图像目标导航是具身人工智能的一项基本能力,但其实际应用受到强假设的制约。现有方法主要依赖于时间顺序的视频流或辅助传感器(例如深度传感器、激光雷达)来维持空间一致性。这些顺序和多模态的依赖性严重限制了可扩展性,尤其是在使用众包或预录制的无序图像集合部署机器人时。当去除时间先验时,当前方法在感知别名、噪声关联和灾难性映射失败方面面临严重挑战。为了解决这一未被充分探索的挑战,我们提出了无序地标视觉导航(Unordered Landmark Visual Navigation, ULVN),这是一个不依赖于时间和里程计先验的统一RGB-only框架。ULVN通过整合映射、定位和规划,系统性地减轻了误差累积。具体而言,它通过校准几何验证和最大生成树优化,直接从非结构化图像构建一个稳健的二维拓扑地图。为了实现闭环执行,ULVN放弃了顺序启发式方法,采用基于图的置信传播滤波器,结合熵自适应融合进行全局定位和动态子目标规划。在模拟和实际部署中的广泛实验表明,ULVN显著优于现有的最先进方法。
cs.RO / 14 / 2608.06847
Are Visual Place Recognition Models Recognizing Places or Conditions? Distractor-Augmented Evaluation and Condition Suppression
视觉地点识别模型是在识别地点还是条件?干扰项增强评估与条件抑制
Abstract
Long-term Visual Place Recognition (VPR) is typically evaluated by matching queries from one condition against a database from another. Crowdsourced map databases, however, may mix conditions and include images that resemble the query in condition but depict different places. In the presence of these distractors, a method may retrieve by condition similarity rather than place identity. We argue that this susceptibility arises because the discriminability of VPR methods allows them to encode information such as illumination, weather, and seasonal appearance in their descriptors. We therefore introduce Distractor-Augmented Recall (DAR) to isolate and quantify the effect of distractors, and propose condition suppression to remove condition information from VPR descriptors. Across eleven methods and six datasets, method rankings under DAR@1 differ from those under Recall@1 (R@1), while applying INLP and LEACE as condition suppression methods generally improves DAR@1 without reducing R@1. Thus, distractor robustness is distinct from standard retrieval performance and can be improved by suppressing condition information.
Chinese Translation
长期视觉地点识别(VPR)通常通过将来自一个条件的查询与来自另一个条件的数据库进行匹配来进行评估。然而,众包地图数据库可能会混合条件,并包含在条件上与查询相似但描绘不同地点的图像。在这些干扰项的存在下,一种方法可能会根据条件相似性而非地点身份进行检索。我们认为,这种易受干扰的特性源于VPR方法的可区分性,使其能够在描述符中编码诸如光照、天气和季节外观等信息。因此,我们引入了干扰项增强召回(Distractor-Augmented Recall, DAR)来隔离和量化干扰项的影响,并提出条件抑制以从VPR描述符中去除条件信息。在十一种方法和六个数据集的实验中,DAR@1下的方法排名与Recall@1(R@1)下的排名不同,而应用INLP和LEACE作为条件抑制方法通常会提高DAR@1而不降低R@1。因此,干扰项的鲁棒性与标准检索性能是不同的,并且可以通过抑制条件信息来改善。
cs.RO / 15 / 2608.06898
How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots
我该如何为我的机器人选择基础模型?支持社会机器人社区评估框架
Abstract
Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public leaderboards offer little guidance: the demands of real-time, embodied social interaction lie largely outside their focus. And direct evaluation is impractical at scale: each embodied study requires scarce participant, robot, and experimenter time. In this paper, we identify five evaluation dimensions for foundation models in social robots: (i) conversational competence, (ii) user safety, (iii) embodied character, (iv) target scene effectiveness, and (v) audience appropriateness. To make model selection cheaper and better informed, we propose a three-tiered evaluation funnel paradigm that first filters with general metrics, then extends to simulated interactions, and terminates in more expensive, robot-specific evaluation. We map all five dimensions across all three tiers, chart where applicable evaluation methods exist and are missing, and close with a call to action: let's build the evaluation framework together as a community.
Chinese Translation
寻求在基础模型上构建社会机器人应用的研究人员面临一个困难的问题:我们该如何选择模型?公共排行榜提供的指导有限:实时、具身的社会互动的需求在很大程度上超出了它们的关注范围。而直接评估在规模上又不切实际:每项具身研究都需要稀缺的参与者、机器人和实验者的时间。在本文中,我们确定了社会机器人基础模型的五个评估维度:(i)对话能力,(ii)用户安全,(iii)具身角色,(iv)目标场景有效性,以及(v)受众适宜性。为了使模型选择更加经济且信息更充分,我们提出了一种三级评估漏斗范式,首先通过一般指标进行筛选,然后扩展到模拟互动,最后在更昂贵的机器人特定评估中终止。我们在所有三个层级中映射所有五个维度,绘制适用的评估方法存在与缺失的地方,并以呼吁行动结束:让我们作为一个社区共同构建评估框架。
cs.RO / 16 / 2608.06907
Spatiotemporal Agility: Time-Constrained Reinforcement Learning for Vision-Guided Dynamic Quadrupedal Interception
时空敏捷性:面向视觉引导的动态四足拦截的时间约束强化学习
Abstract
Legged robots require robust agility to perceive and interact with complex and dynamic environments within a constrained time. However, most existing quadruped locomotion works rely on velocity-tracking policy, which struggle to reach precise targets within strict temporal constraints. Moreover, integrating real-time perception with agile locomotion for highly dynamic targets remains challenging due to sensor latency and processing delays. To concretely study and benchmark such agility in dynamic settings, we introduce a challenging ball-catching task for legged robots. This paper proposes an integrated framework that combines a vision module for landing point and time prediction with a direct position and time conditioned RL locomotion policy, instead of intermediate velocity commands. Beyond the method design, this work presents a system-level contribution that completes real-time robotic interception system that integrates multi-camera perception, online trajectory prediction, low-latency target communication, and sim-to-real locomotion control into a closed-loop deployment pipeline. By explicitly predicting the future spatial-temporal target, our approach mitigates perception latency during dynamic interception. We conducted extensive ball-catching experiments for the legged robot. Through comparative experiments against a velocity-tracking baseline, our direct target-conditioned approach achieves a higher success rate in catching balls with predicted landing spots within 2 meters and flight times between 0.8 and 1.2 seconds. This shows that the robot has successfully completed the dynamic ball-catching task under our tested setup. Furthermore, our policy exhibits a smaller performance gap after deployment, suggesting improved sim-to-real behavior in these trials.
Chinese Translation
四足机器人需要具备强大的敏捷性,以便在有限的时间内感知和与复杂动态环境进行交互。然而,现有的大多数四足运动研究依赖于速度跟踪策略,这使得在严格的时间约束下难以达到精确目标。此外,由于传感器延迟和处理延迟,将实时感知与敏捷运动结合以应对高度动态的目标仍然具有挑战性。为了具体研究和基准测试这种在动态环境中的敏捷性,我们为四足机器人引入了一项具有挑战性的接球任务。本文提出了一个综合框架,将用于落点和时间预测的视觉模块与直接位置和时间条件的强化学习(RL)运动策略相结合,而不是使用中间的速度指令。除了方法设计外,本研究还提出了一个系统级贡献,完成了一个实时机器人拦截系统,该系统集成了多摄像头感知、在线轨迹预测、低延迟目标通信和从仿真到现实的运动控制,形成一个闭环部署管道。通过明确预测未来的时空目标,我们的方法在动态拦截过程中减轻了感知延迟。我们对四足机器人进行了广泛的接球实验。与速度跟踪基线进行比较实验表明,我们的直接目标条件方法在捕捉落点预测在2米以内且飞行时间在0.8到1.2秒之间的球时,成功率更高。这表明机器人在我们测试的设置下成功完成了动态接球任务。此外,我们的策略在部署后表现出更小的性能差距,表明在这些试验中改善了从仿真到现实的行为。
cs.RO / 17 / 2608.06965
Cross-View Action Consistency for Camera-Robust Vision-Language-Action Policies
跨视角动作一致性用于摄像机鲁棒的视觉-语言-动作策略
Abstract
Vision-language-action (VLA) policies fine-tuned from a fixed scene camera can fail when the camera is moved, even when the task, objects, language, and robot state are unchanged. We study scene-camera viewpoint robustness using only a scene RGB image, language, and proprioception, without camera labels, extrinsics, depth, or point-cloud inputs. The wrist stream is masked throughout to prevent an unperturbed visual shortcut from confounding attribution to scene-camera variation. For flow-based VLAs, we propose to regularize the action-flow velocity field, the quantity directly integrated to generate continuous action chunks. We construct action-equivalent view pairs by resetting original LIBERO demonstrations to the same MuJoCo state and rendering nominal and perturbed scene-camera views. Both views are supervised by flow matching, while a cross-view loss encourages their predicted action-flow velocities to agree at the same sampled flow coordinates. On the LIBERO-Plus camera-perturbation track, our method reaches 87.2$\pm$0.4% (4,797 rollouts per seed across 3 training seeds), +7.4pp over flow-matching-only training on the same paired data (79.8$\pm$0.8%, also 3 seeds) and +12.5pp over naive mixed-camera SFT, while maintaining nominal-camera ID performance (95.0$\pm$0.8%; same-data FM-only: 95.0$\pm$4.3%). A shuffled-pair control collapses to 25.8%, showing that the gain depends on action-equivalent pairing. On a real robot, we evaluate three tabletop tasks with 10 rollouts per task and camera placement; held-out-camera success improves from 53.3% to 74.4% under the same single-scene-RGB inference interface.
Chinese Translation
从固定场景摄像机微调的视觉-语言-动作(VLA)策略在摄像机移动时可能会失效,即使任务、物体、语言和机器人状态保持不变。我们研究了仅使用场景RGB图像、语言和自我感知(proprioception)来实现场景摄像机视角鲁棒性,而不依赖摄像机标签、外参、深度或点云输入。整个过程中手腕流被屏蔽,以防止未受干扰的视觉捷径混淆对场景摄像机变化的归因。对于基于流的VLA,我们提出对动作流速场进行正则化,该量直接用于生成连续的动作块。我们通过将原始LIBERO演示重置为相同的MuJoCo状态并渲染名义和扰动的场景摄像机视图,构建动作等效视图对。两个视图通过流匹配进行监督,同时跨视图损失鼓励它们在相同采样流坐标下的预测动作流速一致。在LIBERO-Plus摄像机扰动轨道上,我们的方法达到87.2±0.4%(每个种子4,797次回合,跨3个训练种子),比在相同配对数据上仅进行流匹配训练的结果提高了7.4个百分点(79.8±0.8%,同样为3个种子),并且比简单混合摄像机的SFT提高了12.5个百分点,同时保持名义摄像机ID性能(95.0±0.8%;仅流匹配的同数据:95.0±4.3%)。一个随机配对的控制实验崩溃至25.8%,显示出增益依赖于动作等效配对。在真实机器人上,我们评估了三个桌面任务,每个任务和摄像机位置进行10次回合;在相同的单场景RGB推理接口下,保留摄像机的成功率从53.3%提高到74.4%。
cs.RO / 18 / 2608.06991
Exact Thrust-Reversal Limits of Bidirectional Propellers under Bounded Motor Inputs
双向螺旋桨在有限电机输入下的精确推力反转极限
Abstract
Bidirectional propellers are often treated as signed thrust sources, but their thrust is a signed-quadratic function of rotor speed.Thus, thrust reversal necessarily occurs through zero rotor speed, where the ability of a bounded motor torque to change thrust collapses.This work formalizes this obstruction by studying exact thrust-trajectory reproducibility under bounded motor inputs with prescribed smoothness.We derive a normalized thrust-coordinate model with vanishing input gain at zero thrust, and prove necessary and sufficient reproducibility conditions in terms of the zero-crossing order of the desired thrust.Generic reversals, in which thrust crosses zero with nonzero slope, require unbounded motor input; the resulting conditions provide direct design rules for shaping thrust reversals that avoid singular motor commands.We also derive the corresponding current and voltage regularity requirements for a DC motor driving a bidirectional propeller.Experiments on a motor-propeller setup validate the predicted reversal-order effects, showing localized current/voltage peaks and thrust-tracking degradation for linear reversals, but not for higher-order reversals.These results expose an intrinsic actuator-level limitation that must be considered in force, acceleration, and interaction-control references for aerial robots.
Chinese Translation
双向螺旋桨通常被视为有符号推力源,但其推力是转子速度的有符号二次函数。因此,推力反转必然发生在零转子速度处,此时有限电机扭矩改变推力的能力消失。本研究通过研究在规定光滑度下有限电机输入下的精确推力轨迹可重复性,形式化了这一障碍。我们推导出一个在零推力时输入增益消失的归一化推力坐标模型,并证明了以期望推力的零过零阶为条件的必要和充分可重复性条件。一般的反转,即推力以非零斜率穿过零,要求无限电机输入;由此产生的条件为塑造避免奇异电机指令的推力反转提供了直接设计规则。我们还推导了驱动双向螺旋桨的直流电机的相应电流和电压规律性要求。在电机-螺旋桨装置上的实验验证了预测的反转阶效应,显示出线性反转时的局部电流/电压峰值和推力跟踪退化,而在高阶反转时则没有。这些结果揭示了一个必须在空中机器人力、加速度和交互控制参考中考虑的固有执行器级限制。
cs.RO / 19 / 2608.06994
Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models
将意图与轨迹解耦:一种世界行动模型的表征推导框架
Abstract
World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative motion planning. However, existing visual branches focus on predicting static visual observation, rather than reflecting potential transition information that captures the evolution of world states under motion interactions. This leads to representational entanglement between high-level physical condition evolution and low-level action trajectory generation within the Action Model, creating a structural bottleneck while weakening the predictive capability of world evolution modeling for action generation. We propose PILOT (Physical Inference for Latent Optimized Trajectories), whose core Representational Deduction (RD) bridges this gap by integrating motion thought-of-chain (CoT) guidance as a native model capability. Specifically, RD aims to encourage the action branch to explicitly model potential state transition tokens, which are retained as CoT in the reasoning space to guide fine-grained motion trajectory. Experiments demonstrate that RD not only significantly improves the success rate and generalization ability of WAMs in complex robotic manipulation tasks but also enhances the model's physical interpretability by decoupling high-level motion semantics from low-level trajectory details. Furthermore, the abundant state transition supervision signals introduced by RD effectively alleviate the sparse supervision in action generation, enabling it to serve as an efficient few-shot real-robot fine-tuning strategy and demonstrating superior scalability for migration to mainstream WAM architectures.
Chinese Translation
世界行动模型(World Action Models, WAMs)旨在构建一种统一架构,能够理解世界状态的演变并指导生成性运动规划。然而,现有的视觉分支侧重于预测静态视觉观察,而未能反映捕捉运动交互下世界状态演变的潜在转变信息。这导致了在行动模型中高层物理条件演变与低层动作轨迹生成之间的表征纠缠,形成了结构瓶颈,同时削弱了世界演变建模在动作生成中的预测能力。我们提出了PILOT(Physical Inference for Latent Optimized Trajectories),其核心表征推导(Representational Deduction, RD)通过将运动链思维(CoT)指导作为原生模型能力来弥合这一差距。具体而言,RD旨在鼓励动作分支显式建模潜在状态转变标记,这些标记在推理空间中保留为CoT,以指导细粒度的运动轨迹。实验表明,RD不仅显著提高了WAMs在复杂机器人操作任务中的成功率和泛化能力,还通过将高层运动语义与低层轨迹细节解耦,增强了模型的物理可解释性。此外,RD引入的丰富状态转变监督信号有效缓解了动作生成中的稀疏监督,使其能够作为一种高效的少量样本真实机器人微调策略,并展示出向主流WAM架构迁移的优越可扩展性。
cs.RO / 20 / 2608.06996
Automated Terminal-to-Housing Assembly System for Flat Ribbon Cable Harness
用于扁平带状电缆线束的自动终端到外壳组装系统
Abstract
This paper presents a sensor-minimal automated assembly system for bidirectional single-row flat ribbon cable harnesses (FRCHs). Unlike conventional peg-in-hole or single-terminal insertion tasks, FRCH assembly involves mechanically coupled multi-terminal insertion under flexible and dense geometric constraints. To address this problem, the proposed system performs the assembly through a purely mechanical sequence consisting of Cable Alignment, Lean & Slide, Weaving, and Clamping, without relying on active sensing or vision. Each mechanism is designed to progressively reduce correlated terminal misalignment, insertion interference, and instability before final locking. Experiments on bidirectional single-row FRCHs achieved an 83.75% end-to-end process success rate over 80 trials, with success rates of 85.0% and 82.5% in the first and second halves, respectively. The cycle time was 33 s under half-speed operation. To the best of our knowledge, this work presents the first automated prototype for FRCH terminal-to-housing assembly for multi-pin housings.
Chinese Translation
本文提出了一种传感器最小化的自动组装系统,用于双向单排扁平带状电缆线束(FRCHs)。与传统的插销入孔或单终端插入任务不同,FRCH组装涉及在灵活且密集的几何约束下进行机械耦合的多终端插入。为了解决这一问题,所提出的系统通过一个纯机械序列进行组装,该序列包括电缆对齐、倾斜与滑动、编织和夹紧,而不依赖于主动传感或视觉。每个机制的设计旨在逐步减少相关的终端错位、插入干扰和不稳定性,直至最终锁定。在对双向单排FRCHs进行的实验中,80次试验的端到端过程成功率达到了83.75%,其中前半部分和后半部分的成功率分别为85.0%和82.5%。在半速操作下,循环时间为33秒。据我们所知,这项工作展示了第一个用于多针外壳的FRCH终端到外壳组装的自动原型。
cs.RO / 21 / 2608.07002
A Haptic Robot Finger Designed for Guqin Instrument Playing
为古琴演奏设计的触觉机器人手指
Abstract
With the rapid advancement of humanoid robotics and embodied intelligence technologies, numerous musical instrument-playing robots have emerged in recent years, such as pianos, chime bells, and taiko drums. These robots primarily employ open-loop positional control, rendering them incapable of operating instruments requiring dexterous hands and precise tactile perception, such as a violin, guitar, and guqin. This paper describes the design and validation of a high-precision tactile-sensing finger. By mimicking the shape of the fingertip and fingernail found on a human finger, we develop a biomimetic multimodal haptic fingertip and validate it on selected guqin string-contact tasks, including open-string and stopped-note comparisons, harmonic-tuning, and tactile-triggered bimanual coordination, using the guqin, a traditional Chinese musical instrument, as a challenging validation scenario rather than as a fully demonstrated robotic performance system. This research integrates tactile sensing with robotics technology, thereby contributing to applications in world heritage conservation and cultural dissemination.
Chinese Translation
随着类人机器人和具身智能技术的快速发展,近年来出现了许多音乐演奏机器人,例如钢琴、钟琴和太鼓。这些机器人主要采用开环位置控制,因此无法操作需要灵巧手指和精确触觉感知的乐器,如小提琴、吉他和古琴。本文描述了一种高精度触觉传感手指的设计与验证。通过模仿人类手指上的指尖和指甲的形状,我们开发了一种仿生多模态触觉指尖,并在选定的古琴弦接触任务上进行了验证,包括开弦与按音的比较、和声调音以及触觉触发的双手协调,使用古琴这一传统中国乐器作为具有挑战性的验证场景,而非完全展示的机器人表演系统。本研究将触觉感知与机器人技术相结合,为世界遗产保护和文化传播的应用做出了贡献。
cs.RO / 22 / 2608.07004
Benchmarking and Reasoning Distillation of Large Language Models for Feedback Controller Design in Complex Dynamical Systems
大型语言模型在复杂动态系统反馈控制器设计中的基准测试与推理蒸馏
Abstract
Although remarkable capabilities have been demonstrated by Large Language Models (LLMs) across scientific domains, feedback controller design remains underexplored. Existing benchmarks focus mainly on linear single-Degree-of-Freedom (DoF) systems and large API-hosted models, leaving performance on complex controller-design tasks and feasibility for edge deployment unclear. To address these limitations, we introduce the Complex Dynamics-to-Control Benchmark for Large Language Models (CoDyControlBench), comprising 132 system configurations across five evaluation dimensions: number of DoF, system type, coupling level, damping regime, and controller type. Six state-of-the-art LLMs were evaluated over three independent runs, including three commercial models (GPT, Gemini, and Claude) and three open-source models (GLM, DeepSeek, and Qwen). GPT achieved the highest design success rate at 94.8\%, whereas Qwen showed the lowest rate at 50.0\%. Across the benchmark dimensions, DoF and controller type exhibited the largest model-averaged variations in design success, with success-rate ranges of 36.3\% and 17.6\%, respectively, exceeding those associated with system type, coupling level, and damping regime. Comparison of GPT and Qwen showed that their performance gap arose mainly from the control-design knowledge, particularly gain selection and the use of transient-limiting mechanisms. For edge deployment, a specialized 1.5B-parameter model was developed through reasoning distillation. The reasoning-distilled model outperformed the answer-distilled and base model on CoDyControlBench, maintained stable performance across 1-6 DoFs, and achieved successful traget tracking in all three physical trials on a pneumatic-artificial-muscle-driven robotic arm. These results establish a benchmark baseline and highlight the potential of lightweight, edge-deployable controller-design models.
Chinese Translation
尽管大型语言模型(LLMs)在科学领域展现了显著的能力,但反馈控制器设计仍然未得到充分探索。现有基准主要集中在线性单自由度(DoF)系统和大型API托管模型上,对复杂控制设计任务的性能和边缘部署的可行性尚不明确。为了解决这些局限性,我们引入了大型语言模型的复杂动态控制基准(Complex Dynamics-to-Control Benchmark for Large Language Models,CoDyControlBench),该基准包含132个系统配置,涵盖五个评估维度:自由度数量、系统类型、耦合水平、阻尼机制和控制器类型。我们对六个最先进的LLMs进行了三次独立评估,包括三个商业模型(GPT、Gemini和Claude)和三个开源模型(GLM、DeepSeek和Qwen)。GPT的设计成功率最高,达94.8%,而Qwen的成功率最低,仅为50.0%。在基准维度中,自由度和控制器类型在设计成功率上表现出最大的模型平均变异,成功率范围分别为36.3%和17.6%,超过了与系统类型、耦合水平和阻尼机制相关的变异。GPT与Qwen的比较表明,它们的性能差距主要源于控制设计知识,特别是增益选择和瞬态限制机制的使用。为了边缘部署,我们通过推理蒸馏开发了一种专门的1.5B参数模型。推理蒸馏模型在CoDyControlBench上优于答案蒸馏模型和基础模型,在1-6 DoFs上保持稳定性能,并在气动人工肌肉驱动的机器人手臂的三次物理试验中实现了成功的目标跟踪。这些结果建立了基准基线,并突显了轻量级、可边缘部署的控制器设计模型的潜力。
cs.RO / 23 / 2608.07005
Real-time Whole-Body Motion Planning for Mobile Manipulators Carrying Arbitrarily Shaped Payloads via Kinematically-Coupled SVSDF
基于运动学耦合SVSDF的移动操纵器实时全身运动规划,适用于携带任意形状负载
Abstract
Mobile manipulators are increasingly tasked with transporting large, non-convex payloads through cluttered environments, yet existing planners either oversimplify the payload geometry or fail to handle the kinematic coupling between manipulator links, leading to lost feasible space or stalled optimization. This letter presents a real-time whole-body motion planning framework for mobile manipulators carrying arbitrarily shaped payloads. The front-end employs a chain-decomposed kernel-based collision check that preserves the true geometry of the robot and payload, with compact storage and fast bit-level queries. A mid-end preprocessing stage converts the front-end path into a continuous trajectory enforcing smoothness and feasibility, and executes it directly when collision-free to bypass the costly back-end. When refinement is required, the back-end performs trajectory optimization built on a Kinematically-Coupled SVSDF (KC-SVSDF), which propagates collision-avoidance gradients along the kinematic chain to produce coherent whole-body escape directions. Ablation studies, comparative benchmarks against state-of-the-art baselines, and real-world experiments on a differential-drive mobile manipulator demonstrate that the proposed framework reliably transports large, non-convex payloads through tight passages and cluttered environments.
Chinese Translation
移动操纵器越来越多地被要求在杂乱环境中运输大型非凸负载,但现有的规划方法要么过于简化负载几何形状,要么无法处理操纵器连杆之间的运动学耦合,导致可行空间的丧失或优化停滞。本文提出了一种针对携带任意形状负载的移动操纵器的实时全身运动规划框架。前端采用链分解的基于核的碰撞检测,保留了机器人和负载的真实几何形状,具有紧凑的存储和快速的位级查询。中端预处理阶段将前端路径转换为连续轨迹,以确保平滑性和可行性,并在无碰撞时直接执行,以绕过昂贵的后端。当需要细化时,后端基于运动学耦合SVSDF(KC-SVSDF)进行轨迹优化,该方法沿运动学链传播避碰梯度,以产生一致的全身逃逸方向。消融研究、与最先进基线的比较基准测试以及在差分驱动移动操纵器上的实际实验表明,所提出的框架能够可靠地在狭窄通道和杂乱环境中运输大型非凸负载。
cs.RO / 24 / 2608.07045
C2Dex: Contact-Consistent Reconstruction and Retargeting for Dexterous Manipulation from Monocular Video
C2Dex:基于单目视频的接触一致重建与灵巧操作的重定向
Abstract
High-quality demonstrations for dexterous robot manipulation are costly and difficult to collect, whereas monocular human videos provide a scalable source of diverse manipulation behaviors. However, transferring such demonstrations to dexterous robots remains challenging: monocular hand-object interaction (HOI) reconstruction often produces temporally unstable contacts and physically implausible interactions, while conventional retargeting methods struggle to preserve task-relevant contacts and local interaction geometry across different hand embodiments. We present C2Dex, a video-to-dexterous-manipulation framework built around a shared interaction representation: stable object-side contacts recovered by aggregating noisy frame-wise observations in the canonical object space. These stable contacts serve a dual role: as trajectory-level constraints that guide reconstruction toward temporally coherent and physically plausible human HOI trajectories, and as explicit transfer targets for the dexterous hand, where Laplacian interaction optimization preserves the local hand-object geometry across embodiments and residual reinforcement learning refines the trajectory in simulation. Experiments on DexYCB and TACO show that C2Dex achieves end-to-end trajectory success rates of 57.78% and 26.67%, respectively, substantially outperforming the strongest baselines (17.78% and 10.00%) under identical evaluation criteria. Real-robot replay experiments further demonstrate physical feasibility across diverse contact-rich manipulation tasks. Project page: https://k-jie.github.io/C2Dex/
Chinese Translation
高质量的灵巧机器人操作演示收集成本高且困难,而单目人类视频则提供了一个可扩展的多样化操作行为来源。然而,将这些演示转移到灵巧机器人上仍然面临挑战:单目手-物体交互(HOI)重建往往产生时间上不稳定的接触和物理上不合理的交互,而传统的重定向方法在不同手部表现之间难以保持任务相关的接触和局部交互几何。我们提出了C2Dex,一个基于共享交互表示的视频到灵巧操作框架:通过在规范物体空间中聚合嘈杂的逐帧观测,恢复稳定的物体侧接触。这些稳定的接触起到双重作用:作为引导重建朝向时间一致且物理合理的人类HOI轨迹的轨迹级约束,以及作为灵巧手的显式转移目标,其中拉普拉斯交互优化保持了不同表现之间的局部手-物体几何,而残差强化学习则在仿真中细化轨迹。在DexYCB和TACO上的实验表明,C2Dex分别实现了57.78%和26.67%的端到端轨迹成功率,显著优于最强基线(17.78%和10.00%)在相同评估标准下的表现。真实机器人重放实验进一步展示了在多样化接触丰富的操作任务中的物理可行性。项目页面:https://k-jie.github.io/C2Dex/
cs.RO / 25 / 2608.07065
AutoIntervene: Calibrated Intervention for Action-Chunking Imitation Learning Policies
AutoIntervene:用于动作分块模仿学习策略的校准干预
Abstract
Action-chunking visuomotor policies learn from demonstrations and improve temporal consistency by predicting short action sequences rather than single-step commands. Yet perception errors and execution drift can move the robot outside the demonstration distribution, while the policy continues to produce smooth action chunks that are inconsistent with the observed state. We present AutoIntervene, an online framework that selectively transfers control between an action-chunking policy and an operator during deployment. AutoIntervene evaluates proposed chunks against a visual-action support memory built from successful task executions, combining visual similarity with consistency between proposed and reference actions. Phase-local support governs policy-to-operator transfer within the current task phase, whereas global support governs the return to policy control after operator recovery. We calibrate separate switching thresholds for the two directions from empirical quantiles of evaluation-level scores on held-out expert demonstrations, avoiding direct manual tuning of score cutoffs. Intervention segments retained from successful rollouts target learner-induced states and provide corrective supervision for subsequent policy updates. Experiments on real-world bimanual manipulation tasks show higher post-adaptation task success and lower operator-control time than manual intervention. Videos and additional results are available at https://aus.bot/research/autointervene/.
Chinese Translation
动作分块的视觉运动策略通过示范学习并通过预测短期动作序列而非单步指令来提高时间一致性。然而,感知错误和执行漂移可能使机器人偏离示范分布,而策略仍然产生与观察状态不一致的平滑动作块。我们提出了AutoIntervene,这是一个在线框架,在部署过程中选择性地在动作分块策略和操作员之间转移控制。AutoIntervene评估提议的动作块与从成功任务执行中构建的视觉-动作支持记忆的匹配程度,将视觉相似性与提议动作和参考动作之间的一致性相结合。阶段局部支持在当前任务阶段内管理策略到操作员的转移,而全局支持则在操作员恢复后管理返回策略控制。我们为两个方向校准了独立的切换阈值,这些阈值基于保留专家示范的评估级别分数的经验分位数,避免了对分数截止值的直接手动调优。从成功的回滚中保留的干预片段针对学习者诱导的状态,并为后续的策略更新提供纠正监督。在真实的双手操作任务上的实验表明,后适应任务成功率更高,操作员控制时间更低,相较于手动干预。视频和更多结果可在 https://aus.bot/research/autointervene/ 获取。
cs.RO / 26 / 2608.07074
M2-SMap: Memory-Efficient Semantic Mapping with Hierarchical Multi-Model Representation
M2-SMap:基于层次多模型表示的内存高效语义映射
Abstract
Dense point cloud maps, as a typically used mapping representation, are difficult to deploy on resource-constrained robots because their memory consumption grows rapidly with scene scale. Although compact single-model representations reduce memory cost, their fixed geometric expressiveness is insufficient for structurally diverse environments. Existing multi-model methods improve representational flexibility, yet their feature extraction and model selection are often dominated by local geometry, which can cause overfitting and adhesion between objects. To address these issues, this paper presents M2-SMap, a memory-efficient semantic mapping framework based on hierarchical multi-model representation. First, a hierarchical geometric decomposition partitions RGB-D point clouds into compact Gaussian components. Then, a projection-guided semantic annotation mechanism assigns instance identities to each component. Subsequently, these annotations are incorporated into an object-aware Gaussian fusion strategy. Furthermore, a multi-scale feature extraction strategy separates large planar regions, semantic objects, and complex residual structures, which are respectively represented by bounded planes, object-level superquadrics, and GMM primitives. Experiments on three RGB-D sequences show that M2-SMap runs in real time at no less than 29.37 Hz while achieving the lowest primitive count, with an average reduction of 18.7% over the best baseline. It also reduces the mean per-frame number of measured inter-object adhesion cases from 2.808 to 0, demonstrating efficient and semantically consistent scene representation.
Chinese Translation
密集点云地图作为一种典型的映射表示,在资源受限的机器人上部署困难,因为其内存消耗随着场景规模的增长而迅速增加。尽管紧凑的单模型表示降低了内存成本,但其固定的几何表达能力不足以应对结构多样的环境。现有的多模型方法提高了表示灵活性,但其特征提取和模型选择往往受到局部几何的主导,这可能导致过拟合和物体之间的粘连。为了解决这些问题,本文提出了M2-SMap,一种基于层次多模型表示的内存高效语义映射框架。首先,层次几何分解将RGB-D点云划分为紧凑的高斯成分。然后,投影引导的语义标注机制为每个成分分配实例身份。随后,这些标注被纳入对象感知的高斯融合策略。此外,多尺度特征提取策略将大平面区域、语义对象和复杂残余结构分开,分别用有界平面、对象级超二次体和高斯混合模型(GMM)原语表示。在三个RGB-D序列上的实验表明,M2-SMap以不低于29.37 Hz的实时速度运行,同时实现了最低的原语数量,平均减少了18.7%相较于最佳基线。它还将每帧测量的物体间粘连案例的平均数量从2.808减少到0,展示了高效且语义一致的场景表示。
cs.RO / 27 / 2608.07075
Detection and Ranging of Transient Extrinsic Contacts Based on 6D Dynamic Tactile Sensing
基于6D动态触觉感知的瞬态外部接触检测与测距
Abstract
Delicate manipulation often involves transient and subtle collisions between a grasped object and the environment. While the human hand localizes these contacts effortlessly thanks to superior tactile sensitivity, robotic systems often lack the requisite resolution to acquire the information necessary for motion planning, resulting in clumsy manipulation or even task failure. Here, we propose transient extrinsic contact detection and ranging (TECDAR), a simple yet fast and efficient method for detecting and ranging extrinsic contact of grasped objects. Our design of gripper tips employs dynamic tactile sensing leveraging a single 2.5$\times$3 mm 6D inertial measurement unit. The sensor captures sub-millisecond tip deformations at a 7 kHz sampling rate, but operating on a data stream of only 84 KB/s. High bandwidth and compact data size enable the system to rapidly detect and localize contact between grasped objects and their surroundings. Specifically, fusing tactile data with robot pose via an extended Kalman filter enables fast and precise localization of extrinsic contact, reaching millimeter-level accuracy within 180 ms. Experimental results demonstrate that the system achieves an average localization accuracy of approximately 7\,mm in both line-contact and point-contact localization tasks. Furthermore, this near-instantaneous localization enables the robot to rectify its trajectory on a millisecond scale, facilitating precise tool manipulation and enhanced perception of complex environments purely through tactile exploration and mapping. We envision such techniques advancing the future of robotics across domains requiring delicate manipulation, including precision assembly, surgical assistance, and autonomous exploration in touch-dominant environments. Project page: humitlab.github.io/TECDAR/
Chinese Translation
精细操作通常涉及被抓取物体与环境之间的瞬态和微妙碰撞。尽管人手凭借卓越的触觉敏感性能够轻松定位这些接触,但机器人系统往往缺乏获取运动规划所需信息的分辨率,导致操作笨拙甚至任务失败。在此,我们提出了一种瞬态外部接触检测与测距(TECDAR)方法,这是一种简单但快速高效的检测被抓取物体外部接触的方法。我们的夹持器尖端设计采用动态触觉感知,利用一个2.5×3毫米的6D惯性测量单元。传感器以7 kHz的采样率捕捉亚毫秒级的尖端变形,但数据流速仅为84 KB/s。高带宽和紧凑的数据大小使系统能够快速检测和定位被抓取物体与其周围环境之间的接触。具体而言,通过扩展卡尔曼滤波器将触觉数据与机器人姿态融合,使外部接触的快速精确定位成为可能,达到毫米级精度,响应时间为180毫秒。实验结果表明,该系统在线接触和点接触定位任务中实现了约7毫米的平均定位精度。此外,这种近乎瞬时的定位使机器人能够在毫秒级别上调整其轨迹,从而通过触觉探索和映射实现精确的工具操作和对复杂环境的增强感知。我们展望这种技术将推动机器人在需要精细操作的领域的发展,包括精密组装、外科辅助和在触觉主导环境中的自主探索。项目页面:humitlab.github.io/TECDAR/
cs.RO / 28 / 2608.07079
LifelongCrossNav: Persistent 3D Semantic Memory for Cross-Floor Multi-Object Navigation
LifelongCrossNav:用于跨楼层多目标导航的持久3D语义记忆
Abstract
Object-goal navigation has made substantial progress in semantic perception and exploration, yet persistent memory for multi-object navigation and cross-floor navigation are still commonly addressed separately. We present LifelongCrossNav, a framework for sequential multi-object ObjectNav in unknown multi-floor indoor environments. Within each episode, the agent receives an ordered sequence of object-goal queries while continuously maintaining a shared sparse 3D semantic voxel memory. This memory incrementally accumulates geometric structure, traversability states, and vision-language features, allowing subsequent object-goal queries to retrieve previously acquired scene information without rebuilding the map. To support persistent search across floors, LifelongCrossNav combines support-aware 3D traversability mapping, stair-specific perception, and direction-aware stair traversal. A unified navigation policy coordinates same-floor frontier exploration, live and historical point-of-interest retrieval, stair navigation, and target-object search and approach. We further introduce HM3D-MFMON, a benchmark for sequential Multi-Floor Multi-Object Navigation built on HM3D scenes, including a dedicated subset in which completing the full sequence of object-goal subtasks requires at least one floor transition. Experimental results show that LifelongCrossNav consistently outperforms a representative planar persistent semantic-map baseline on HM3D-MFMON, demonstrating that persistent 3D semantic memory and cross-floor traversability modeling effectively support sequential multi-object navigation in multi-floor environments. Project page: https://flageval-baai.github.io/LifelongCrossNavPage.
Chinese Translation
目标导向导航在语义感知和探索方面取得了显著进展,但对于多目标导航和跨楼层导航的持久记忆仍然通常被分开处理。我们提出了LifelongCrossNav,这是一个在未知多楼层室内环境中进行顺序多目标ObjectNav的框架。在每个回合中,智能体接收一系列有序的目标物体查询,同时持续维护一个共享的稀疏3D语义体素记忆。该记忆逐步积累几何结构、可通行状态和视觉-语言特征,使得后续的目标物体查询能够在不重建地图的情况下检索先前获取的场景信息。为了支持跨楼层的持久搜索,LifelongCrossNav结合了支持感知的3D可通行性映射、楼梯特定感知和方向感知的楼梯通行。统一的导航策略协调同楼层前沿探索、实时和历史兴趣点检索、楼梯导航以及目标物体的搜索与接近。我们进一步介绍了HM3D-MFMON,这是一个基于HM3D场景的顺序多楼层多目标导航基准,包括一个专门的子集,其中完成完整的目标物体子任务序列至少需要一次楼层转换。实验结果表明,LifelongCrossNav在HM3D-MFMON上始终优于一个代表性的平面持久语义地图基线,证明了持久的3D语义记忆和跨楼层可通行性建模有效支持多楼层环境中的顺序多目标导航。项目页面:https://flageval-baai.github.io/LifelongCrossNavPage。
cs.RO / 29 / 2608.07140
Identifying the Key Biomechanical Features of Movement Adaptation during Exoskeleton-Assisted Locomotion
识别外骨骼辅助运动中运动适应的关键生物力学特征
Abstract
The understanding of natural human adaptation during exoskeleton-assisted locomotion - particularly individual differences in adaptation behaviors and temporal progression - remains limited. In this work, we investigate temporal evolution of biomechanical variables to uncover participant-specific adaptation strategies across different exoskeleton-assisted locomotion scenarios. Nine healthy participants performed treadmill walking under three conditions: without an exoskeleton, with exoskeleton active ankle assistance, and with exoskeleton zero-torque. Lower limb kinematics, inter-joint coordination, and metabolic cost of transport (MCoT) were analyzed at both the group and individual levels. Results indicate that adaptation is gradual and highly individualized, with substantial variability in convergence timing and movement patterns across participants. Kinematic adaptation occurred asynchronously across lower limb, with larger fluctuations during the swing phase. Metabolic responses were heterogeneous and often non-convergent, highlighting the limitations of steady-state assumptions commonly adopted in the literature. These findings emphasize the importance of individual-level, temporal evolution analyses for understanding adaptation dynamics in exoskeleton use.
Chinese Translation
在外骨骼辅助运动中,关于自然人类适应的理解——特别是适应行为和时间进程的个体差异——仍然有限。在本研究中,我们探讨了生物力学变量的时间演变,以揭示不同外骨骼辅助运动场景下参与者特定的适应策略。九名健康参与者在三种条件下进行跑步机行走:不使用外骨骼、使用外骨骼主动踝部辅助和使用外骨骼零扭矩。我们在群体和个体层面分析了下肢运动学、关节间协调性和运输代谢成本(MCoT)。结果表明,适应是渐进的且高度个性化,参与者之间在收敛时机和运动模式上存在显著变异。下肢的运动学适应是异步发生的,摆动阶段的波动更大。代谢反应表现出异质性,且通常不收敛,突显了文献中常采用的稳态假设的局限性。这些发现强调了个体层面时间演变分析在理解外骨骼使用中的适应动态的重要性。
cs.RO / 30 / 2608.07154
Representation Handoffs for OpenArm-Based Laboratory Mobile Manipulation
基于OpenArm的实验室移动操作的表征交接
Abstract
Open-source robotics and foundation models have lowered the barrier to embodied AI, yet language-guided laboratory automation still requires reliable alignment from instructions and observations to safe actions. This field report presents an OpenArm-based mobile manipulation prototype for laboratory-style tasks, built by integrating dual OpenArm manipulators with a mobile base, vertical slide, RGB-D sensing, lidar-based mapping, ROS2/MoveIt execution, and profile-defined skill interfaces. The system is organized around representation handoffs: natural language requests are constrained into registered skill calls, sensor observations are grounded into maps and object poses, object priors provide role and skill constraints, and runtime bindings compile validated skills into executable motion goals. We use dry-run traces and startup checks to evaluate this integration path, showing how the prototype exposes missing calibration, incomplete object assets, and unfinished real-scene visual grounding as explicit deployment blockers. These intermediate representations serve as practical debugging interfaces for integrating language, perception, planning, and robot safety in embodied systems.
Chinese Translation
开源机器人技术和基础模型降低了具身人工智能的门槛,但语言引导的实验室自动化仍然需要可靠的指令和观察与安全行动之间的对齐。本领域报告展示了一种基于OpenArm的移动操作原型,旨在执行实验室风格的任务,该原型通过将双OpenArm操纵器与移动底座、垂直滑轨、RGB-D传感器、基于激光雷达的地图构建、ROS2/MoveIt执行以及按配置文件定义的技能接口集成而成。该系统围绕表征交接进行组织:自然语言请求被约束为注册的技能调用,传感器观察被固定到地图和物体姿态,物体先验提供角色和技能约束,运行时绑定将验证过的技能编译为可执行的运动目标。我们使用干运行轨迹和启动检查来评估这一集成路径,展示原型如何将缺失的校准、不完整的物体资产和未完成的真实场景视觉定位作为明确的部署障碍。这些中间表征作为实际调试接口,服务于在具身系统中集成语言、感知、规划和机器人安全。
cs.RO / 31 / 2608.07314
TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models
TEMPO:面向视觉-语言-动作模型的语义-动作解耦强化学习后训练
Abstract
Vision-language-action (VLA) models are commonly adapted to downstream manipulation tasks via supervised fine-tuning (SFT) or online reinforcement learning (RL) post-training. SFT is prone to distribution mismatch, and existing RL approaches typically apply a single, uniform update strategy to all model components, ignoring their distinct functional roles. We propose TEMPO, a semantic-action decoupled, two-timescale RL post-training framework for VLA models. TEMPO freezes the pretrained vision-language backbone to preserve general semantic representations, and restricts adaptation to two components with dedicated RL optimization loops: the semantic projection layer and the low-level action expert. We update them at different rates--the semantic projection layer infrequently, to keep the latent action stable, and the action expert frequently, to rapidly incorporate control feedback from online interaction. This decoupling RL fine-tuning strategy prevents fast policy updates from destabilizing high-level semantic representations while still allowing the action expert to learn efficiently from online feedback. Experiments on the CALVIN benchmark and real-world manipulation tasks demonstrate that TEMPO consistently outperforms both pretrained state-of-the-art VLA models and the RL post-training baseline, while reaching and maintaining higher evaluation rewards on two real-world tasks.
Chinese Translation
视觉-语言-动作(VLA)模型通常通过监督微调(SFT)或在线强化学习(RL)后训练适应下游操作任务。SFT容易受到分布不匹配的影响,而现有的RL方法通常对所有模型组件应用单一的统一更新策略,忽视了它们各自的功能角色。我们提出了TEMPO,一种语义-动作解耦的双时间尺度RL后训练框架,专为VLA模型设计。TEMPO冻结预训练的视觉-语言主干,以保持一般语义表示,并将适应限制在两个具有专门RL优化循环的组件上:语义投影层和低级动作专家。我们以不同的速率更新它们——语义投影层更新频率较低,以保持潜在动作的稳定性,而动作专家更新频率较高,以快速整合来自在线交互的控制反馈。这种解耦的RL微调策略防止了快速策略更新对高级语义表示的破坏,同时仍允许动作专家有效地从在线反馈中学习。在CALVIN基准和真实世界操作任务上的实验表明,TEMPO在性能上始终优于预训练的最先进VLA模型和RL后训练基线,并在两个真实世界任务上实现并维持了更高的评估奖励。
cs.RO / 32 / 2608.07328
Learning Fault-Tolerant Locomotion with Adaptive Gait Timing
学习具有自适应步态时序的容错运动
Abstract
Hardware failures require legged robots to rapidly reorganize coordination and gait timing to maintain stability and mobility. This is particularly challenging for larger quadrupeds, where increased mass and tighter actuation limits reduce the feasibility of aggressive, high-frequency compensation strategies often observed on smaller platforms. In this work, we propose a deep reinforcement learning approach for fault-tolerant locomotion under actuator power loss. The method employs an asymmetric actor-critic architecture in which the critic has access to privileged information during training, while the actor learns to reconstruct a corresponding latent representation from proprioceptive observations. We introduce a latent-alignment loss that encourages consistency between actor and critic representations. Additionally, we augment the action space with a learnable gait frequency parameter, enabling adaptive gait timing in response to terrain variations and actuator degradation without predefined faulty-leg strategies. The approach is validated in high-fidelity simulation on uneven terrain and real-world experiments on flat ground using a 68 kg quadruped robot.
Chinese Translation
硬件故障要求四足机器人快速重新组织协调和步态时序,以维持稳定性和移动性。这对于较大的四足机器人尤其具有挑战性,因为增加的质量和更严格的驱动限制降低了在较小平台上常见的激进、高频补偿策略的可行性。在本研究中,我们提出了一种针对驱动器功率损失的容错运动的深度强化学习方法。该方法采用不对称的演员-评论家架构,其中评论家在训练期间可以访问特权信息,而演员则学习从本体感知观察中重建相应的潜在表示。我们引入了一种潜在对齐损失,鼓励演员和评论家表示之间的一致性。此外,我们通过一个可学习的步态频率参数扩展了动作空间,使得在没有预定义故障腿策略的情况下,能够根据地形变化和驱动器退化自适应调整步态时序。该方法在高保真模拟中验证了不平坦地形的有效性,并在68公斤的四足机器人上进行了平坦地面的实际实验。
cs.RO / 33 / 2608.07361
Depth-Wise Probing and Pruning of the Planning Token in a Driving Vision-Language-Action Model
驾驶视觉-语言-动作模型中规划令牌的深度探测与剪枝
Abstract
Vision-language-action (VLA) models route driving decisions through a deep language model, but it is unclear how much of that depth the action itself requires. We study a representative driving VLA whose entire plan is carried by a single planning token that a generative planner decodes into a trajectory. Borrowing the planner as a trajectory-space logit lens, we decode the planning token from every one of the 32 decoder layers and measure two signals: the linear decodability of the navigation command and trajectory compatibility with the frozen native planner. Our diagnostic shows that semantic intent is linearly decodable early: command-probe accuracy reaches 97.7\% after the first decoder layer, compared with 16.7\% chance. In contrast, compatibility with the frozen native planner improves gradually across depth, with open-loop Avg-L2 reaching its minimum of 2.11\,m only at the final layer. Learned readouts from the first layer recover much of this gap, indicating that planning information is already present early but is not yet represented in the format expected by the deployed planner. Ranking decoder layers by the angular deviation they induce in the planning token permits removal of 8 of 32 layers within an approximately 5\% relative open-loop error increase and yields a measured 1.33$\times$ decoder speedup. At the evaluated sample size, no family-specific degradation is statistically resolved. These findings are limited to the evaluated ORION checkpoint and Bench2Drive setup.
Chinese Translation
视觉-语言-动作(VLA)模型通过深度语言模型引导驾驶决策,但尚不清楚该深度对动作本身的需求有多大。我们研究了一个典型的驾驶VLA,其整个计划由一个单一的规划令牌承载,该令牌由生成规划器解码为轨迹。借助规划器作为轨迹空间的logit透镜,我们从32个解码器层中的每一个解码出规划令牌,并测量两个信号:导航命令的线性可解性和与冻结的原生规划器的轨迹兼容性。我们的诊断显示,语义意图在早期是线性可解的:命令探测的准确率在第一个解码器层后达到97.7%,而随机猜测的准确率仅为16.7%。相比之下,与冻结的原生规划器的兼容性在深度上逐渐改善,开放环路的平均L2误差在最后一层时仅达到最小值2.11米。来自第一层的学习输出恢复了大部分差距,表明规划信息在早期已经存在,但尚未以部署规划器所期望的格式表示。通过对解码器层按其对规划令牌引起的角度偏差进行排名,可以在相对开放环路误差增加约5%的情况下移除32层中的8层,并实现了1.33倍的解码速度提升。在评估的样本规模下,没有统计上显著的特定家族降级。这些发现仅限于评估的ORION检查点和Bench2Drive设置。
cs.CV / 1 / 2608.06404
UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys
UAV3DCrop:重复多角度无人机作物调查中的3D重建基准测试
Abstract
Accurate 3D crop monitoring underpins data-driven precision agriculture by enabling field-scale analysis of plant structure, growth dynamics, and management response. Modern 3D reconstruction methods perform strongly on generic benchmarks, but rendered appearance may not translate into metrically and agronomically useful geometry in crop fields. We introduce UAV3DCrop, a public benchmark of repeated multi-angle unmanned aerial vehicle (UAV) crop surveys. It contains 88,830 RGB images at $5280 \times 3956$ pixels, with a ground sampling distance of 3.6-5.8 mm, from 91 scenes spanning corn, soybean, wheat, and oat. Track A evaluates seven scene-optimized methods -- Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) variants -- on held-out views, photogrammetry-referenced depth, and canopy-height recovery. Track B tests four pretrained feed-forward models on zero-shot camera-pose and geometry estimation. The scene-optimized methods rank differently across the three targets: Splatfacto-big leads appearance, whereas Scaffold-GS leads depth and is statistically tied with Splatfacto for canopy height. Among feed-forward models, MapAnything leads on seven of the eight metrics, while the remaining models vary more across crops and fail severely on absolute scale in a way that alignment conceals. Repeated acquisitions reveal further sensitivities that differ by output type and by model, associated with position within the acquisition sequence and with tie-point multiplicity. Current 3D reconstruction methods are therefore not yet interchangeable for agronomic use: no single method wins on appearance, geometry, and canopy height at once, and only one of four feed-forward models recovers usable metric scale. The dataset is publicly available at https://link-dev.github.io/UAV3DCrop/
Chinese Translation
准确的3D作物监测是数据驱动的精准农业的基础,它能够对植物结构、增长动态和管理响应进行田间规模的分析。现代3D重建方法在通用基准测试中表现出色,但其渲染外观可能无法转化为在作物田中具有度量和农业学意义的几何形状。我们介绍了UAV3DCrop,这是一个公共基准,涵盖了重复多角度无人机(UAV)作物调查。该基准包含了来自91个场景的88,830张RGB图像,分辨率为$5280 imes 3956$像素,地面采样距离为3.6-5.8毫米,场景包括玉米、大豆、小麦和燕麦。Track A评估了七种场景优化方法——神经辐射场(NeRF)和3D高斯点云(3DGS)变体——在保留视图、基于摄影测量的深度和冠层高度恢复上的表现。Track B测试了四个预训练的前馈模型在零-shot相机姿态和几何估计上的表现。场景优化方法在三个目标上的排名有所不同:Splatfacto-big在外观上领先,而Scaffold-GS在深度上领先,并在冠层高度上与Splatfacto的表现统计上相当。在前馈模型中,MapAnything在八个指标中的七个上表现最佳,而其他模型在作物间的表现差异较大,并在绝对尺度上严重失效,这种失效在对齐过程中被掩盖。重复采集揭示了不同输出类型和模型的进一步敏感性,这与采集序列中的位置和连接点的多重性相关。因此,当前的3D重建方法尚不能互换用于农业用途:没有单一方法在外观、几何和冠层高度上同时获胜,且四个前馈模型中只有一个能够恢复可用的度量尺度。该数据集可在 https://link-dev.github.io/UAV3DCrop/ 上公开获取。
cs.CV / 2 / 2608.06406
Deep Evidential Regression for Sparse Forest Height Estimation from Multimodal Satellite Imagery
基于深度证据回归的稀疏森林高度估计方法:来自多模态卫星影像的研究
Abstract
Accurate estimation of forest height from satellite imagery is essential for applications such as carbon accounting, biodiversity monitoring, and ecosystem management. While recent deep learning approaches provide accurate predictions, they typically do not quantify predictive uncertainty. This limitation is particularly relevant in geospatial settings characterized by sparse supervision and geographic distribution shift. In this work, we investigate Deep Evidential Regression (DER) for forest height estimation on the TreeUQ benchmark, a large-scale dataset designed for the joint estimation of tree count and average tree height at 10 m resolution, based on Sentinel-1/-2 data as well as tree inventory data over the federal state of Bavaria. To account for the extreme label sparsity of the tree inventory data, we introduce a masked evidential loss for dense geospatial prediction. Using a U-Net architecture with multimodal Sentinel-1 and Sentinel-2 inputs, the proposed approach jointly predicts tree height and associated uncertainty estimates in a single forward pass. Experimental results show that DER achieves predictive performance comparable to a deterministic U-Net while additionally providing well-calibrated uncertainty estimates. These findings demonstrate the potential of evidential learning as an efficient framework for uncertainty-aware forest structure estimation from Earth observation data.
Chinese Translation
从卫星影像中准确估计森林高度对于碳核算、生物多样性监测和生态系统管理等应用至关重要。尽管近期的深度学习方法提供了准确的预测,但它们通常无法量化预测的不确定性。这一局限性在特征稀疏和地理分布变化的地理空间环境中尤为相关。在本研究中,我们探讨了深度证据回归(Deep Evidential Regression, DER)在TreeUQ基准数据集上的森林高度估计,该数据集是一个大规模数据集,旨在基于Sentinel-1/-2数据以及巴伐利亚州的树木清查数据,以10米分辨率联合估计树木数量和平均树高。为了应对树木清查数据的极端标签稀疏性,我们提出了一种用于密集地理空间预测的掩蔽证据损失。采用U-Net架构,结合多模态的Sentinel-1和Sentinel-2输入,所提出的方法在单次前向传播中联合预测树木高度及相关的不确定性估计。实验结果表明,DER的预测性能与确定性U-Net相当,同时还提供了良好校准的不确定性估计。这些发现展示了证据学习作为一种高效框架在地球观测数据中进行不确定性感知的森林结构估计的潜力。
cs.CV / 3 / 2608.06407
TransSLR: A Lightweight Transformer for Sign Language Recognition
TransSLR:一种轻量级变换器用于手语识别
Abstract
Automated Sign Language Recognition for under-represented languages remains a largely unsolved problem. Central African Sign Language (CASL) exemplifies this gap: the only available bench-mark, CASL-W60, has a best reported accuracy of 69.93%, and we show that the common heuristic of fine-tuning high-resource models fails to close it. This failure stems from two compounding factors: the limited scale of available CASL data and the significant lexical and visual domain gap between CASL and large-scale corpora such as WLASL, which renders pre-trained representations largely uninformative. To address this, we propose TransSLR, a lightweight Temporal Transformer Encoder trained from scratch on 64-frame normalized pose sequences, with average pooling and a classification head. By operating on geometric keypoint representations rather than raw RGB, TransSLR achieves signer-independent generalization without relying on visual appearance. On the CASL-W60 benchmark, TransSLR establishes a new state-of-the-art accuracy of 80.39%, surpassing the prior best by +10.46%. Beyond accuracy, our encoder-only design significantly reduces computational overhead, making deployment feasible in resource-constrained environments. We conduct extensive experiments on the CASL-W60 benchmark, comparing against RGB-based and multimodal baselines, and demonstrate that TransSLR achieves state-of-the-art performance.
Chinese Translation
自动化手语识别在代表性不足的语言中仍然是一个未解决的问题。中非手语(Central African Sign Language, CASL)正是这一差距的典型例子:唯一可用的基准CASL-W60的最佳报告准确率为69.93%,而我们表明,微调高资源模型的常见启发式方法未能缩小这一差距。这一失败源于两个相互影响的因素:可用CASL数据的规模有限,以及CASL与大型语料库(如WLASL)之间显著的词汇和视觉领域差距,这使得预训练表示在很大程度上缺乏信息。为了解决这一问题,我们提出了TransSLR,一种从头开始在64帧标准化姿态序列上训练的轻量级时序变换器编码器,采用平均池化和分类头。通过对几何关键点表示而非原始RGB进行操作,TransSLR实现了与签署者无关的泛化,而不依赖于视觉外观。在CASL-W60基准上,TransSLR建立了新的最先进准确率为80.39%,超过了之前的最佳结果10.46%。除了准确率外,我们的仅编码器设计显著降低了计算开销,使得在资源受限环境中的部署成为可能。我们在CASL-W60基准上进行了广泛的实验,与基于RGB和多模态的基线进行了比较,证明TransSLR达到了最先进的性能。
cs.CV / 4 / 2608.06467
Test-Time Adaptation with Online Personalized Energy-Based Cache for Fine-Grained Video Expression Recognition
基于在线个性化能量缓存的测试时适应用于细粒度视频表情识别
Abstract
Facial expression recognition (FER) in videos is challenging because models must identify subtle, temporally evolving affective states that vary across individuals. Although vision-language models provide transferable visual-semantic representations, models trained on subject-independent data often degrade under subject-specific distribution shifts at inference time. Existing test-time adaptation (TTA) methods commonly update model parameters during inference, increasing computational cost and latency. Cache-based methods avoid parameter updates, but they usually require enough target samples to form reliable class prototypes, which is difficult early in adaptation and for rarely observed classes. We introduce Energy-Based Cache Personalization (EB-CaP), a subject-based online TTA method for video FER that generates class-specific prototypes personalized to each target video. EB-CaP uses a lightweight energy-based model to sample prototypes from the current unlabeled video and populate a personalized cache online, without accumulating large amounts of target data or storing diverse source prototypes. Its energy function relies only on pretrained CLIP: similarities between the target video embedding and class text embeddings guide prototype sampling. In parallel, positive and negative caches store reliable and uncertain target embeddings. An adaptive entropy gate controls cache updates according to the evolving confidence distribution, while a diversity gate limits redundant samples. Final predictions combine cache-derived scores with the current CLIP scores. Experiments on BioVid, StressID, and BAH show that EB-CaP outperforms state-of-the-art TTA methods while maintaining low computational and memory overhead. Code is available at https://github.com/MasoumehSharafi/EB-CaP.
Chinese Translation
视频中的面部表情识别(FER)具有挑战性,因为模型必须识别细微的、随时间演变的情感状态,这些状态因个体而异。尽管视觉-语言模型提供了可转移的视觉-语义表示,但在推理时,基于与受试者无关的数据训练的模型通常会在受试者特定的分布变化下性能下降。现有的测试时适应(TTA)方法通常在推理过程中更新模型参数,这增加了计算成本和延迟。基于缓存的方法避免了参数更新,但通常需要足够的目标样本来形成可靠的类别原型,这在适应初期和对于少见类别时是困难的。我们提出了基于能量的缓存个性化(EB-CaP),这是一种针对视频FER的基于受试者的在线TTA方法,能够生成针对每个目标视频的类别特定原型。EB-CaP使用轻量级的能量模型从当前未标记的视频中采样原型,并在线填充个性化缓存,而无需积累大量目标数据或存储多样的源原型。其能量函数仅依赖于预训练的CLIP:目标视频嵌入与类别文本嵌入之间的相似性指导原型采样。同时,正向和负向缓存存储可靠和不确定的目标嵌入。自适应熵门根据不断变化的置信度分布控制缓存更新,而多样性门限制冗余样本。最终预测结合了缓存派生的分数和当前的CLIP分数。在BioVid、StressID和BAH上的实验表明,EB-CaP在保持低计算和内存开销的同时,优于最先进的TTA方法。代码可在 https://github.com/MasoumehSharafi/EB-CaP 获取。
cs.CV / 5 / 2608.06490
InsertFuse: A Unified Framework for Multi-Category Reference-Guided Image Insertion
InsertFuse:一种多类别参考引导图像插入的统一框架
Abstract
We present InsertFuse, a unified framework for multi-category reference-guided image insertion. Its key idea is to decouple category-specific expertise learning from cross-category capability consolidation. InsertFuse first trains specialized experts for different insertion categories and then introduces Insertion On-Policy Distillation (IOPD) to consolidate their capabilities into a single student. By querying the matched expert at states visited by the student, IOPD preserves category-specific insertion behavior while mitigating the cross-category interference caused by direct joint training. To improve spatial control, we propose Token-Aligned Geometry Conditioning (TAGC), which maps mask-derived geometric cues to the visual token grid, and Region-Balanced Flow Matching, which separately normalizes prediction errors inside and outside the insertion region to prevent background-dominated and scale-dependent supervision. We further introduce Reference CFG to isolate and strengthen the guidance induced by the visual reference under fixed scene and geometry conditions, with IOPD transferring this enhanced supervision into the unified student. Extensive experiments on the public AnyInsertion benchmark and our multi-category test set demonstrate state-of-the-art performance on most metrics, showing strong reference fidelity and generation quality across diverse insertion categories.
Chinese Translation
我们提出了InsertFuse,这是一种多类别参考引导图像插入的统一框架。其关键思想是将特定类别的专业知识学习与跨类别能力整合解耦。InsertFuse首先为不同的插入类别训练专门的专家,然后引入插入在线策略蒸馏(Insertion On-Policy Distillation, IOPD)将它们的能力整合到一个单一的学生模型中。通过在学生访问的状态下查询匹配的专家,IOPD保留了特定类别的插入行为,同时减轻了直接联合训练所造成的跨类别干扰。为了改善空间控制,我们提出了令牌对齐几何条件(Token-Aligned Geometry Conditioning, TAGC),该方法将基于掩码的几何线索映射到视觉令牌网格,并提出区域平衡流匹配(Region-Balanced Flow Matching),该方法分别对插入区域内外的预测误差进行归一化,以防止背景主导和尺度依赖的监督。我们进一步引入参考条件生成框架(Reference CFG),以在固定场景和几何条件下隔离和增强由视觉参考引起的指导,IOPD将这种增强的监督转移到统一的学生模型中。在公共的AnyInsertion基准和我们的多类别测试集上的大量实验表明,在大多数指标上表现出最先进的性能,显示出在多样化插入类别中强大的参考保真度和生成质量。
cs.CV / 6 / 2608.06580
Improving Low-Resolution Face Recognition under Limited Data: How Synthetic Data Generation Can Close the Domain Gap
在有限数据下改善低分辨率人脸识别:合成数据生成如何缩小领域差距
Abstract
Face Recognition (FR) systems in surveillance settings often encounter Low Resolution (LR) faces, those whose face region falls below the standard 112 $\times$ 112 input size. While labelled High Resolution (HR) training data is abundant, labelled native-LR data, and above all paired native LR/HR data, is scarce. One workaround is to synthesize LR data from the available HR faces, but how much synthesis effort is repaid in recognition accuracy remains unclear. We present a study of simple synthetic generation strategies for a compact, edge device-oriented face recognition system, spanning interpolation-based degradation, knowledge distillation, a Prepended Domain Transformer (PDT), Real ESRGAN-style degradation, and a learned Super Resolution (SR) front-end with an identity-aware loss. We evaluate these strategies on synthetic cross-resolution face benchmarks (LFW, CFP-FP, AgeDB-30) and on TinyFace, a real-world native LR dataset, and expose a synthetic-real gap: the degradation setting that is optimal on synthetic benchmarks is not the one that is optimal on real LR. We find that more synthesis effort does not help monotonically: the learned SR front-end does not surpass a direct feed of the aligned LR image into a strong backbone, while simple interpolation augmentation of a compact backbone is the only synthesis that improves over its own baseline. We conclude that generative methods for LR face recognition must be validated on real LR and against a direct-feed baseline, and release our pipeline at https://idiap.ch/paper/synth-lrfr
Chinese Translation
监控环境中的人脸识别(FR)系统常常遇到低分辨率(LR)人脸,即其面部区域低于标准的112 × 112输入大小。虽然标记的高分辨率(HR)训练数据丰富,但标记的原生LR数据,尤其是配对的原生LR/HR数据却稀缺。一个解决方法是从可用的HR人脸合成LR数据,但合成的努力在识别准确性上的回报程度仍不明确。我们展示了一项关于简单合成生成策略的研究,针对紧凑型、面向边缘设备的人脸识别系统,涵盖了基于插值的降解、知识蒸馏、预置域变换器(PDT)、真实ESRGAN风格的降解,以及具有身份感知损失的学习超分辨率(SR)前端。我们在合成跨分辨率人脸基准(LFW、CFP-FP、AgeDB-30)和真实世界原生LR数据集TinyFace上评估了这些策略,并揭示了合成与真实之间的差距:在合成基准上最优的降解设置并不是在真实LR上最优的设置。我们发现,更多的合成努力并不单调有利:学习的SR前端并未超过将对齐的LR图像直接输入强大主干网络的效果,而简单的插值增强是唯一改善其自身基线的合成方法。我们得出结论,LR人脸识别的生成方法必须在真实LR上进行验证,并与直接输入基线进行比较,我们的工作流程已发布在https://idiap.ch/paper/synth-lrfr
cs.CV / 7 / 2608.06599
Toward surface-based registration of a virtual preoperative cutting guide onto the mandible for reconstruction surgery
面向虚拟术前切割导向器在下颌骨重建手术中的表面注册
Abstract
Mandibular reconstruction restores facial continuity and oral function after segmental resection. Patient-specific cutting guides transfer a computed tomography (CT)-based plan to the operating room with three-dimensional information, but printed guides add cost and lead time, cannot adapt after fabrication, and may interrupt surgery if sterility is lost. We investigate a markerless augmented reality (AR) alternative that registers a virtual cutting guide to the exposed mandible from surface geometry. The method extends surface-based registration for the transoral setting, where teeth form the most distinctive visible surface. After camera calibration, a HoloLens 2 time-of-flight camera captures a partial intraoperative point cloud. The user supplies a rough head-based alignment only to crop the region of interest. A teeth-weighted global stage computes correspondences and solves a truncated least-squares rigid alignment. An asymmetric point-to-plane iterative closest point (ICP) stage refines the complete CT mandible against the partial depth cloud target. The guide-to-mandible transform places the guide in the HoloLens world frame, while pose updates and interpolation follow target motion. We define a blinded phantom protocol with 30 target registration error (TRE) points under full, intermediate, and teeth-only exposure, plus a motion-to-display latency test. Our median TRE is 4.05, 6.10, and 7.10 mm respectively, and median latency is 0.805 s. These values support the feasibility of using AR to replace physical prints. The workflow removes mounted fiducials and manual landmark selection and provides a testable path toward transoral cutting guidance.
Chinese Translation
下颌骨重建在段切除后恢复面部连续性和口腔功能。患者特定的切割导向器将基于计算机断层扫描(CT)的计划转移到手术室,提供三维信息,但打印的导向器增加了成本和交付时间,无法在制作后进行调整,并且如果失去无菌状态可能会中断手术。我们研究了一种无标记增强现实(AR)替代方案,该方案通过表面几何将虚拟切割导向器注册到暴露的下颌骨上。该方法扩展了在经口环境中的基于表面的注册,其中牙齿形成最显著的可见表面。在相机校准后,HoloLens 2飞行时间相机捕获部分术中点云。用户仅提供粗略的基于头部的对齐,以裁剪感兴趣区域。加权牙齿的全局阶段计算对应关系并求解截断最小二乘刚性对齐。一个非对称点到平面迭代最近点(ICP)阶段将完整的CT下颌骨与部分深度云目标进行细化。导向器到下颌骨的变换将导向器放置在HoloLens世界坐标系中,同时姿态更新和插值跟随目标运动。我们定义了一个盲法幻影协议,包含30个目标注册误差(TRE)点,分别在完全、部分和仅牙齿暴露下进行测试,以及一个运动到显示延迟测试。我们的中位TRE分别为4.05、6.10和7.10毫米,中位延迟为0.805秒。这些数值支持使用AR替代物理打印的可行性。该工作流程消除了安装的标志物和手动标记选择,并提供了一条可测试的经口切割指导路径。
cs.CV / 8 / 2608.06612
SLED: Scalable Location Encoding via Distillation
SLED:通过蒸馏实现可扩展的位置编码
Abstract
The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but the sheer size of the Earth Observations (EO), differing modalities, and different sensor types pose significant challenges in doing so. Location encoders have emerged as an efficient way of compressing EOs into location-specific embeddings. However, current state-of-the-art location encoders rely on computationally expensive CLIP-style frameworks that require large batch sizes in the 16K--32K range, suffer from false negative samples, and scale poorly with additional modalities. We introduce the Scalable Location Encoder via Distillation (SLED), a distillation-based location encoder that uses geospatial location as a binding modality to pretrain location encoders with any modality of geospatial data. The resulting location encoder framework is lightweight, modular, and can flexibly incorporate multiple modes, while eliminating the need for spatiotemporal coregistration of samples. SLED is performant with batch sizes as small as 128, enabling pretraining at a fraction of the runtime and compute costs of current state-of-the-art models. We demonstrate our approach by pretraining unimodal and multimodal SLED models on Sentinel-1, Sentinel-2, and Landsat imagery. We show that both unimodal and multimodal SLED models keep pace with or outperform existing approaches on a diverse set of 19 human-centric benchmark tasks and explore the benefits of using additional modes in pretraining.
Chinese Translation
大量可用的地理空间数据为学习高质量的地球表征提供了令人兴奋的机会,但地球观测(EO)的庞大规模、不同的模态和不同的传感器类型在实现这一目标时带来了显著挑战。位置编码器作为一种有效的方式,已逐渐成为将EO压缩为特定位置嵌入的手段。然而,当前最先进的位置编码器依赖于计算成本高昂的CLIP风格框架,这些框架需要16K到32K范围的大批量样本,容易出现假阴性样本,并且在增加额外模态时扩展性较差。我们提出了可扩展的位置编码器(SLED),这是一种基于蒸馏的位置编码器,利用地理空间位置作为绑定模态,预训练与任何地理空间数据模态的位置信息编码器。所得到的位置编码器框架轻量、模块化,能够灵活地整合多种模态,同时消除了样本时空核心配准的需求。SLED在批量大小仅为128时表现出色,使得预训练的运行时间和计算成本仅为当前最先进模型的一小部分。我们通过在Sentinel-1、Sentinel-2和Landsat影像上预训练单模态和多模态的SLED模型来展示我们的方法。我们表明,无论是单模态还是多模态的SLED模型在19个以人为中心的基准任务中均与现有方法保持同步或表现优越,并探讨了在预训练中使用额外模态的好处。
cs.CV / 9 / 2608.06613
Do 3D Medical Foundation Models See Through MRI Artifacts? A Controlled Study of Representation Robustness
3D医学基础模型能否透视MRI伪影?一项关于表征鲁棒性的对照研究
Abstract
Self-supervised 3D medical foundation models are increasingly used as general-purpose feature extractors, yet their sensitivity to MRI artifacts remains poorly understood. We present a controlled evaluation of representation robustness across five pretrained 3D encoders spanning different architectures, objectives, pretraining domains, and dataset scales. Using BraTS-Africa cases with four MRI sequences, we generate seven frequency- and image-domain artifacts at five predefined corruption settings. Robustness is assessed using linear centered kernel alignment (CKA), RankMe, and UMAP, complemented by an independent segmentation-consistency analysis. We find that robustness is strongly model- and artifact-dependent. 3DINO exhibits the most consistently stable representations, while BrainIAC is highly sensitive to several corruptions; NeuroVFM, BrainFM, and Neuro-SimCLR show intermediate but distinct artifact-specific profiles. Across many conditions, CKA decreases substantially while RankMe remains comparatively stable, indicating that artifacts often distort representation geometry without causing dimensional collapse. Segmentation consistency also degrades under corruption, particularly for ghosting and Rician noise, but aligns only partially with representation-level robustness. These findings show that larger-scale or domain-specific pretraining alone does not guarantee artifact invariance and motivate explicit robustness evaluation before deploying 3D foundation models in heterogeneous MRI settings.
Chinese Translation
自监督的3D医学基础模型越来越多地被用作通用特征提取器,但它们对MRI伪影的敏感性仍然不够了解。我们对五种不同架构、目标、预训练领域和数据集规模的预训练3D编码器进行了控制评估,以研究表征的鲁棒性。使用BraTS-Africa案例及四种MRI序列,我们在五个预定义的损坏设置下生成七种频率域和图像域伪影。我们使用线性中心核对齐(CKA)、RankMe和UMAP评估鲁棒性,并辅以独立的分割一致性分析。我们发现鲁棒性在很大程度上依赖于模型和伪影。3DINO表现出最为稳定的表征,而BrainIAC对多种损坏高度敏感;NeuroVFM、BrainFM和Neuro-SimCLR则显示出中等但明显的伪影特定特征。在许多条件下,CKA显著下降,而RankMe相对稳定,表明伪影通常会扭曲表征几何形状,但不会导致维度崩溃。分割一致性在损坏下也会下降,尤其是在鬼影和Rician噪声情况下,但与表征级鲁棒性仅部分一致。这些发现表明,仅仅依靠大规模或领域特定的预训练并不能保证伪影的不变性,并促使在异构MRI环境中部署3D基础模型之前进行明确的鲁棒性评估。
cs.CV / 10 / 2608.06673
When Semantics Saturate or Emerge: Adaptation-Conditional Semantic Utility in Source-Free Cross-Domain Few-Shot Learning
语义饱和或涌现:无源跨领域少样本学习中的适应条件语义效用
Abstract
Language descriptions in source-free cross-domain few-shot learning (SF-CDFSL) are often selected according to zero-shot accuracy obtained with a frozen vision--language model. This paper asks whether that ranking remains valid after target-domain visual adaptation. Under a strictly paired protocol, we compare a generic class-name template with fixed detailed class descriptions before and after visual Low-Rank Adaptation (LoRA) on EuroSAT, CropDisease, ISIC, and ChestX. Let $\deltazero$ and $\deltalora$ denote the Detailed-minus-Base accuracy before and after adaptation, respectively. Two recurring regimes emerge. In \emph{semantic saturation}, $\deltazero>0$ but $0<\deltalora\ll\deltazero$: on EuroSAT and CropDisease, initial gains of 8.13--21.54 percentage points contract to 0.69--2.96 points after LoRA. In \emph{semantic emergence}, $\deltazero\leq0$ but $\deltalora>0$: on ISIC and ChestX, detailed descriptions become more useful only after the visual representation is updated. Training trajectories and sample-level decomposition show that saturation is driven mainly by Base-LoRA recovering errors already solved by detailed semantics, whereas emergence is associated with prediction turnover and newly formed Detailed-only correct decisions. Fixed-point-free shuffled-semantic controls, a second CLIP backbone, and multiple random seeds support the broad pattern while identifying ChestX 1-shot as a weak boundary case. These findings establish that zero-shot prompt quality is an incomplete proxy for adaptation-anchor quality and motivate evaluating language on both sides of the adaptation boundary.
Chinese Translation
在无源跨领域少样本学习(SF-CDFSL)中,语言描述通常根据冻结的视觉-语言模型获得的零-shot 精度进行选择。本文探讨在目标领域视觉适应后,这种排名是否依然有效。在严格配对的协议下,我们比较了通用类名称模板与在 EuroSAT、CropDisease、ISIC 和 ChestX 上进行视觉低秩适应(LoRA)前后的固定详细类描述。设 $ ext{deltazero}$ 和 $ ext{deltalora}$ 分别表示适应前后的详细减基础精度。出现了两种反复出现的模式。在 extit{语义饱和}中,$ ext{deltazero}>0$ 但 $0< ext{deltalora} ext{ } ext{ll} ext{ } ext{deltazero}$:在 EuroSAT 和 CropDisease 上,最初的增益从 8.13 到 21.54 个百分点在 LoRA 后收缩至 0.69 到 2.96 个百分点。在 extit{语义涌现}中,$ ext{deltazero} ext{ } ext{≤}0$ 但 $ ext{deltalora}>0$:在 ISIC 和 ChestX 上,详细描述仅在视觉表示更新后变得更加有用。训练轨迹和样本级分解表明,饱和主要是由于基础-LoRA 恢复了详细语义已经解决的错误,而涌现则与预测转变和新形成的仅详细正确决策相关。无固定点的随机语义控制、第二个 CLIP 主干和多个随机种子支持了这一广泛模式,同时识别出 ChestX 1-shot 作为一个弱边界案例。这些发现表明,零-shot 提示质量并不是适应锚质量的完整代理,并激励在适应边界的两侧评估语言。
cs.CV / 11 / 2608.06674
Corrupting Attention: Evasion-Based Adversarial Attacks on Encoder Attention in Detection Transformers
破坏注意力:基于规避的对抗攻击在检测变换器中的编码器注意力
Abstract
Adversarial vulnerabilities remain a major concern for the safe deployment of neural networks, particularly in object detection, a core task embedded in many safety-critical systems. Detection transformers have emerged as leading object detectors, yet their adversarial robustness remains comparatively underexplored. Most existing attacks target the detection output rather than the attention mechanism that makes these models distinctive. In this paper, we introduce the first attack that directly optimizes an encoder-attention objective under an imperceptible, bounded $\ell_\infty$ perturbation. Rather than introducing an attacker-owned sink token through a visible patch, it drives the model's own attention toward a corrupted target. We argue that encoder attention concentrates the model's spatial reasoning, so corrupting it propagates through the detection pipeline more disruptively than perturbing the detection output alone. Our attack reduces DETR-R50 mAP on COCO from 42.1 to 0.97, a $\sim 4\times$ reduction in resulting mAP over the strongest existing attack under an identical perturbation budget and iteration count. We further show that this vulnerability is not specific to a particular corruption objective: across four qualitatively distinct targets, dispersion, re-ranking, permutation, and peak-suppression, detection consistently drops below 3 mAP, suggesting that the weakness arises from disrupting the attention structure itself rather than from any single target. Finally, we demonstrate that the attack generalizes across attention formulations, reducing DINO-Swin-L from 56.8 to 1.44 mAP against 7.3 for the strongest prior attack, establishing state-of-the-art on both dense and deformable attention.
Chinese Translation
对抗脆弱性仍然是神经网络安全部署的主要关注点,尤其是在物体检测中,这是一项嵌入在许多安全关键系统中的核心任务。检测变换器已成为领先的物体检测器,但它们的对抗鲁棒性仍然相对未被充分探索。现有的大多数攻击目标是检测输出,而不是使这些模型独特的注意力机制。本文介绍了首个直接在不可察觉的、有界的 $ ext{l}_ ext{∞}$ 扰动下优化编码器注意力目标的攻击。该攻击并不是通过可见的补丁引入攻击者拥有的接收标记,而是将模型自身的注意力引导至一个被破坏的目标。我们认为,编码器注意力集中模型的空间推理,因此破坏它会比单纯扰动检测输出更具破坏性地传播通过检测管道。我们的攻击使得 DETR-R50 在 COCO 数据集上的 mAP 从 42.1 降至 0.97,相比于在相同扰动预算和迭代次数下最强现有攻击,导致的 mAP 减少了约 4 倍。我们进一步表明,这种脆弱性并不特定于某一特定的破坏目标:在四种定性不同的目标(分散、重新排序、置换和峰值抑制)中,检测的 mAP 一直低于 3,表明这种弱点源于破坏注意力结构本身,而不是来自任何单一目标。最后,我们展示了该攻击在注意力形式上的普遍性,使 DINO-Swin-L 的 mAP 从 56.8 降至 1.44,而最强先前攻击的 mAP 为 7.3,确立了在密集和可变形注意力上的最新成果。
cs.CV / 12 / 2608.06691
CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition
CoDAT:具有低成本时间建模的协作双重注意力变换器,用于高效边缘动作识别
Abstract
Real-time human action recognition on Internet-of-Things (IoT) edge devices requires models that capture rich spatio-temporal cues within strict latency, memory, and power envelopes. Current 3D CNNs, video transformers, and shift-based ViT deliver high accuracy but come at computational costs that preclude edge IoT deployment. This paper proposes CoDAT, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context. SSHA jointly compresses the spatial resolution and channel dimensions of the query, key, and value tensors via stride-based sparse projection, then fuses the resulting global and local features at a markedly reduced cost. To enable temporal communication across frames, a parameter-free TShift module is embedded in each block. Extensive experiments on Jetson AGX Orin and Raspberry Pi 5 demonstrate that CoDAT achieves an energy-accuracy balance in both image and action recognition. On ImageNet-1K, CoDAT-M runs 2x faster than EfficientViT384 and FastViT-S12 at comparable accuracy, and CoDAT-L matches ViT-S with 3x fewer parameters at 2x higher throughput. On Kinetics-400 and MA-52, CoDAT achieves competitive Top-1 accuracy against state-of-the-art CNN, transformer, and hybrid baselines while running up to 2.9x faster than VSwin-T, 2x faster than ViT-Temporal-Shift variants, and 5x faster than UniFormer-B. On UCF-101, CoDAT-S384 matches TokShift and LAPS while being 6x faster and requiring up to 13x fewer FLOPs, establishing an efficiency-accuracy balance for real-time action recognition in edge IoT perception systems. Code is available at https://github.com/novendrastywn/CoDAT .
Chinese Translation
在物联网(IoT)边缘设备上进行实时人类动作识别需要模型在严格的延迟、内存和功耗限制内捕捉丰富的时空线索。目前的3D卷积神经网络(CNN)、视频变换器和基于位移的ViT虽然提供了高准确性,但其计算成本使得在边缘IoT部署变得不可行。本文提出了CoDAT,一种协作双重注意力变换器,用轻量级双分支模块替代传统的多头注意力:空间卷积注意力(SCA)用于局部聚合,步幅单头注意力(SSHA)用于全局上下文。SSHA通过基于步幅的稀疏投影联合压缩查询、键和值张量的空间分辨率和通道维度,然后以显著降低的成本融合得到的全局和局部特征。为了在帧之间实现时间通信,每个模块中嵌入了一个无参数的TShift模块。在Jetson AGX Orin和Raspberry Pi 5上的大量实验表明,CoDAT在图像和动作识别中实现了能量与准确性的平衡。在ImageNet-1K上,CoDAT-M的运行速度比EfficientViT384和FastViT-S12快2倍,同时保持相当的准确性,而CoDAT-L在参数数量减少3倍的情况下实现了2倍的吞吐量,匹配ViT-S。在Kinetics-400和MA-52上,CoDAT在与最先进的CNN、变换器和混合基线的竞争中实现了具有竞争力的Top-1准确性,同时运行速度比VSwin-T快最多2.9倍,比ViT-Temporal-Shift变体快2倍,比UniFormer-B快5倍。在UCF-101上,CoDAT-S384与TokShift和LAPS相匹配,同时速度快6倍,并且所需的FLOP数减少最多13倍,为边缘IoT感知系统中的实时动作识别建立了效率与准确性的平衡。代码可在https://github.com/novendrastywn/CoDAT获取。
cs.CV / 13 / 2608.06712
Suppress and Diversify: Refining Robust Pathways for Corruption Robustness
抑制与多样化:优化抗腐蚀性强的路径
Abstract
Model robustness against natural image corruptions is essential for safety-critical applications. While existing methods primarily focus on implicit representation learning, we provide the first systematic exploration of computational pathways to explicitly characterize internal robustness. We identify a progressive decay of robust features across network layers and establish a functional dependency between the prevalence of these features and model performance. To exploit these insights, we propose Suppress and Diversify (S\&D), a non-intrusive refinement approach that enhances robustness by dynamically selecting robust pathways and diversifying them through symmetry-preserving transformations. S\&D is architecture-agnostic, parameter-free, and incurs zero test-time overhead. Extensive evaluations across eight benchmarks demonstrate that S\&D consistently improves performance across multiple vision tasks, diverse backbones, and complex real-world scenarios, highlighting its broad efficacy and scalability.
Chinese Translation
模型对自然图像腐蚀的鲁棒性对于安全关键应用至关重要。尽管现有方法主要集中在隐式表示学习上,我们首次系统性地探索了计算路径,以明确表征内部鲁棒性。我们发现网络层中鲁棒特征的逐步衰减,并建立了这些特征的普遍性与模型性能之间的功能依赖关系。为了利用这些见解,我们提出了抑制与多样化(Suppress and Diversify, S&D),这是一种非侵入式的优化方法,通过动态选择鲁棒路径并通过保持对称性的变换来实现多样化,从而增强鲁棒性。S&D与架构无关,无需参数,并且在测试时没有额外开销。在八个基准测试中的广泛评估表明,S&D在多个视觉任务、多样化的骨干网络和复杂的现实场景中始终提高了性能,突显了其广泛的有效性和可扩展性。
cs.CV / 14 / 2608.06717
WaveFreqAnchor: Wave-Structural Anchoring and Frequency Correction Diffusion for Training-Free Face Restoration
WaveFreqAnchor:基于波结构锚定和频率修正扩散的无训练人脸修复
Abstract
Diffusion-based face restoration that adjusts the sampling trajectory of pre-trained diffusion models has achieved remarkable progress. However, existing approaches provide insufficient constraints during reverse diffusion, causing identity-related structural drift and degraded fidelity under severe degradations. To address this, we propose WaveFreqAnchor, a training-free framework based on Wave-Structural Anchoring and Frequency Correction Diffusion. Specifically, Anchor-Space Wave-Structural Guidance (ASWG) constrains facial structures through anisotropic wave-response consistency, while Multi-scale Wavelet-Fourier Injection (MWFI) aligns the predicted low-frequency subband with the observation by replacing its phase, correcting inconsistencies accumulated during reverse diffusion. For real-world scenes, we further introduce Subband High-Frequency Enhancement (SHE), which performs bounded, spatially masked refinement on the predicted high-frequency subbands to recover fine facial details under unknown compound degradations. Together, these designs effectively preserve facial identity while restoring sharp and realistic facial details. Extensive experiments show that our method consistently outperforms existing methods, achieving high-quality and high-fidelity face restoration.
Chinese Translation
基于扩散的人脸修复通过调整预训练扩散模型的采样轨迹取得了显著进展。然而,现有方法在反向扩散过程中提供的约束不足,导致身份相关的结构漂移和在严重退化下的保真度下降。为此,我们提出了WaveFreqAnchor,这是一个基于波结构锚定和频率修正扩散的无训练框架。具体而言,锚空间波结构引导(Anchor-Space Wave-Structural Guidance, ASWG)通过各向异性波响应一致性约束面部结构,而多尺度小波-傅里叶注入(Multi-scale Wavelet-Fourier Injection, MWFI)通过替换其相位来对齐预测的低频子带与观测值,修正反向扩散过程中累积的不一致性。针对真实场景,我们进一步引入子带高频增强(Subband High-Frequency Enhancement, SHE),对预测的高频子带进行有界的空间掩蔽细化,以在未知复合退化下恢复细致的面部细节。综合这些设计,我们有效地保持了面部身份,同时恢复了清晰且真实的面部细节。大量实验表明,我们的方法在面部修复方面始终优于现有方法,实现了高质量和高保真的人脸修复。
cs.CV / 15 / 2608.06751
Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation
超越星夜:面向艺术家的文本到图像生成的快捷方式意识控制状态规划
Abstract
Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical shortcuts, such as recurring motifs, generic palettes, or overrepresented period signatures, rather than preserving the user's intended scene. We introduce Atelier, a shortcut-aware control-state planning framework for artist-grounded image generation. Atelier translates underspecified artistic intent into an explicit control state that separates scene anchors, preserve/transform decisions, style-regime hypotheses, role-bound artist evidence, and shortcut-avoidance constraints. It grounds this state using artist-level knowledge and local patch references, compiles backend-aware generation plans, and iteratively refines candidates through global and local authenticity feedback. We further introduce ArtIntentBench, a benchmark covering Van Gogh and Qi Baishi across artwork re-rendering, period/style-controlled generation, historically unseen subjects, shortcut auditing, and human preference evaluation. Across open-weight and closed-source generators, Atelier improves artist-level style fidelity, preserves source structure more faithfully, and substantially reduces shortcut substitution compared with prompt-engineered, retrieval-augmented, and general-purpose agent baselines. These results suggest that artist-grounded generation is bottlenecked not only by image synthesis, but by the upstream inference of explicit, evidence-grounded artistic controls.
Chinese Translation
面向艺术家的图像生成不仅仅是将艺术家姓名附加到提示中。图像模型通常通过规范快捷方式来响应艺术家姓名,例如反复出现的主题、通用调色板或过度代表的时期签名,而不是保留用户意图的场景。我们提出了Atelier,一个面向艺术家的图像生成的快捷方式意识控制状态规划框架。Atelier将不明确的艺术意图转化为一个明确的控制状态,该状态分离场景锚点、保留/转换决策、风格模式假设、角色绑定的艺术家证据和避免快捷方式的约束。它使用艺术家级知识和局部补丁参考来构建这一状态,编制后端感知生成计划,并通过全球和局部真实性反馈迭代地优化候选项。我们进一步介绍了ArtIntentBench,这是一个基准,涵盖了梵高和齐白石的艺术作品再渲染、时期/风格控制生成、历史上未见的主题、快捷方式审计和人类偏好评估。在开放权重和闭源生成器中,与经过提示工程、检索增强和通用代理基线相比,Atelier提高了艺术家级风格保真度,更忠实地保留了源结构,并显著减少了快捷方式替代。这些结果表明,面向艺术家的生成不仅受到图像合成的瓶颈限制,还受到明确的、基于证据的艺术控制的上游推理限制。
cs.CV / 16 / 2608.06768
Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models
探索还是收敛?基于阶段引导的逐步优化用于扩散模型
Abstract
Diffusion models have strong generative capabilities. However, their maximum likelihood training objective only focuses on reconstructing the data distribution, making it difficult to align with specific preferences. Reinforcement learning (RL) for preference alignment in diffusion models is promising but limited by reward sparsity. Since a single reward cannot support optimization, existing RL methods usually backpropagate the final reward to all previous steps. However, denoising is stage-wise, with distinct semantics and controllability. Repeating the final reward across all steps creates a temporal objective mismatch, encouraging reward shortcuts that lead to reward hacking. At the same time, due to reward backfilling, each time step receives the same reward, making it impossible to distinguish between actions, thereby weakening the optimization process. To resolve this issue, we propose Stage-Guided Per-Step Optimization (SGPO) for diffusion models, which jointly leverages signal-to-noise ratio and semantic changes to identify generation stages and adaptively assign stage-specific objectives. Early denoising is chaotic and far from the final reward, resulting in weak reward-behavior correlation. This stage should prioritize exiting the chaotic state. In the mid stage, the latent transitions to a stable structure, where the final reward better corresponds to generative behavior. Therefore, this stage optimizes the final reward while exploring diversity to avoid early convergence to a single mode. In the late stage, the latent's core structure is largely fixed, and preference optimization mainly amplifies local details, risking overfitting. Therefore, stable convergence is preferred to avoid quality degradation. Results from 16 comparative experiments validate SGPO. Our method achieves 26.7% average gains in generative quality and 36.7% higher convergence speed.
Chinese Translation
扩散模型具有强大的生成能力。然而,它们的最大似然训练目标仅关注重建数据分布,使得与特定偏好的对齐变得困难。在扩散模型中,基于强化学习(RL)的偏好对齐方法前景广阔,但受到奖励稀疏性的限制。由于单一奖励无法支持优化,现有的强化学习方法通常将最终奖励反向传播到所有之前的步骤。然而,去噪过程是阶段性的,具有不同的语义和可控性。在所有步骤中重复最终奖励会导致时间目标不匹配,鼓励奖励捷径,从而导致奖励黑客行为。同时,由于奖励填充,每个时间步都获得相同的奖励,使得无法区分行动,从而削弱了优化过程。为了解决这个问题,我们提出了用于扩散模型的阶段引导逐步优化(SGPO),该方法共同利用信噪比和语义变化来识别生成阶段,并自适应地分配阶段特定目标。早期去噪过程混乱且距离最终奖励较远,导致奖励与行为之间的相关性较弱。此阶段应优先退出混乱状态。在中期阶段,潜在变量过渡到稳定结构,此时最终奖励与生成行为更好地对应。因此,此阶段优化最终奖励,同时探索多样性,以避免过早收敛到单一模式。在后期阶段,潜在变量的核心结构基本固定,偏好优化主要放大局部细节,存在过拟合的风险。因此,更倾向于稳定收敛,以避免质量下降。16项对比实验的结果验证了SGPO。我们的方法在生成质量上平均提高了26.7%,收敛速度提高了36.7%。
cs.CV / 17 / 2608.06769
GraphVerse: A Comprehensive Visual Graph Reasoning Benchmark for Multimodal Large Language Models
GraphVerse:针对多模态大型语言模型的综合视觉图推理基准
Abstract
Recent Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse vision-language tasks, creating an urgent need for more challenging benchmarks. Yet existing evaluations still provide limited insight into whether these models can truly reason over structured visual information. Visual Graph Reasoning (VGR) offers a compelling testbed for this challenge, requiring models to integrate perception, structural understanding, and multi-step reasoning over graph-based visual inputs. However, prior VGR benchmarks often reduce the task to visual perception followed by text-based reasoning, restrict evaluation to single-image settings, rely on answer-only metrics, and underrepresent realistic graph-centric scenarios. To bridge the gap, we introduce GraphVerse, a unified benchmark that jointly evaluates perception, visual reasoning, and text-based graph reasoning in MLLMs under both single-image and paired-image settings. At its core is a suite of Graph-centric Image Editing (GIE) strategies that modify graph images while preserving their semantics, turning them into active tests of visual reasoning. We further propose VGR-Score, a process-sensitive metric that evaluates reasoning quality beyond final-answer accuracy. Extensive experiments reveal several key limitations of current MLLMs in VGR, while also validating the effectiveness of GIE strategies and the transferability of GraphVerse to broader multimodal reasoning capabilities. The code is available at https://github.com/sunyuanfu/GraphVerse.
Chinese Translation
近期的多模态大型语言模型(MLLMs)在多种视觉-语言任务中取得了显著进展,这使得对更具挑战性的基准测试的需求变得迫切。然而,现有的评估仍然对这些模型是否能够真正对结构化视觉信息进行推理提供了有限的洞察。视觉图推理(VGR)为这一挑战提供了一个引人注目的测试平台,要求模型整合感知、结构理解和基于图的视觉输入的多步推理。然而,以往的VGR基准往往将任务简化为视觉感知后跟随文本基础的推理,限制评估于单图像设置,依赖仅有答案的指标,并且对现实图中心场景的表现不足。为了解决这一问题,我们提出了GraphVerse,这是一个统一的基准,能够在单图像和成对图像设置下共同评估MLLMs的感知、视觉推理和基于文本的图推理。其核心是一套图中心图像编辑(GIE)策略,这些策略在保留图像语义的同时修改图像,将其转变为视觉推理的主动测试。我们进一步提出了VGR-Score,这是一种过程敏感的指标,评估推理质量超越最终答案的准确性。大量实验揭示了当前MLLMs在VGR中的几个关键局限性,同时验证了GIE策略的有效性以及GraphVerse在更广泛的多模态推理能力中的可迁移性。代码可在 https://github.com/sunyuanfu/GraphVerse 获取。
cs.CV / 18 / 2608.06773
AnyTrack: Unifying Visual Object Tracking with Any Modalities
AnyTrack:统一视觉目标跟踪与任意模态
Abstract
Visual object tracking aims to continuously locate specific targets within sequential frames, evolving from single-modal methods to multi-modal ones. However, existing multi-modal trackers are typically designed for fixed modality combinations, requiring separate models for different inputs. This leads to a poor adaptability to missing or imperfect modalities, and limited generalization. To address these issues, we propose a novel unified framework called AnyTrack for object tracking with any modalities. Specifically, we design a Modality-aware Interaction Module (MIM) to facilitate dynamic interaction across diverse modalities. This module bridges modality discrepancies and aggregates temporal cues to maintain spatio-temporal consistency during cross-modal interaction. Furthermore, we introduce a Context Understanding Module (CUM) to establish spatial correspondence between visual features and target locations via global-local prompts. This module employs target-aware context modeling to enhance foreground-background discrimination for precise localization. Finally, to support the training and evaluation under diverse modalities, we extend existing multi-modal object tracking benchmarks by incorporating grayscale images, language descriptions, and audio clips. Extensive experiments with both complete and missing modality settings demonstrate that our AnyTrack achieves state-of-the-art performance, validating its effectiveness and flexibility. The source code is available at https://github.com/IdolLab/AnyTrack.
Chinese Translation
视觉目标跟踪旨在连续定位序列帧中的特定目标,已经从单模态方法发展到多模态方法。然而,现有的多模态跟踪器通常是为固定模态组合设计的,需要针对不同输入使用单独的模型。这导致了对缺失或不完美模态的适应性差,以及有限的泛化能力。为了解决这些问题,我们提出了一种名为AnyTrack的新统一框架,用于任意模态的目标跟踪。具体而言,我们设计了一个模态感知交互模块(Modality-aware Interaction Module, MIM),以促进不同模态之间的动态交互。该模块弥合了模态差异,并聚合时间线索,以在跨模态交互过程中保持时空一致性。此外,我们引入了一个上下文理解模块(Context Understanding Module, CUM),通过全局-局部提示建立视觉特征与目标位置之间的空间对应关系。该模块采用目标感知的上下文建模来增强前景与背景的区分,以实现精确定位。最后,为了支持在多样化模态下的训练和评估,我们通过引入灰度图像、语言描述和音频片段扩展了现有的多模态目标跟踪基准。大量实验在完整和缺失模态设置下表明,我们的AnyTrack实现了最先进的性能,验证了其有效性和灵活性。源代码可在https://github.com/IdolLab/AnyTrack获取。
cs.CV / 19 / 2608.06784
UniCycleFlow: Bidirectional Unpaired Image Translation with a Shared Rectified Flow
UniCycleFlow:具有共享修正流的双向无配对图像翻译
Abstract
Bidirectional unpaired image translation must preserve source-specific structure while learning coherent transformations in both directions without paired supervision. Existing methods typically employ two direction-specific generators or train separate one-way models. Even when linked by cycle consistency, such models constrain only the round-trip endpoint reconstruction, without requiring the two directions to obey a common local transformation rule. We propose UniCycleFlow, a rectified-flow framework that represents bidirectional translation as forward and reverse integration of a single time-conditioned velocity field. This formulation organizes both directions within the same continuous dynamics, rather than coupling otherwise separate endpoint mappings. A key challenge is that unpaired data provide no meaningful source--target coupling from which rectified-flow trajectories can be constructed. UniCycleFlow addresses this challenge by learning deterministic source-conditioned endpoints whose marginal distributions are adversarially matched to the opposite domains. The resulting paths are regularized by stop-gradient self-flow matching for intermediate velocity supervision, discrete cycle closure for forward--reverse consistency, and representation path-velocity regularization for controlling localized feature changes along the trajectory. Across ten translation directions, UniCycleFlow achieves the lowest FID on 7 of 10 tasks using a single Euler evaluation and obtains the best average FID of 55.1.
Chinese Translation
双向无配对图像翻译必须在不使用配对监督的情况下,保持源特定结构,同时学习两个方向上的一致变换。现有方法通常采用两个方向特定的生成器或训练单独的单向模型。即使通过循环一致性相连,这些模型也仅约束了往返端点重建,而不要求两个方向遵循共同的局部变换规则。我们提出了UniCycleFlow,一种修正流框架,将双向翻译表示为单一时间条件速度场的正向和反向积分。这种表述将两个方向组织在同一连续动态中,而不是耦合其他独立的端点映射。一个关键挑战是,无配对数据无法提供有意义的源-目标耦合,从而构建修正流轨迹。UniCycleFlow通过学习确定性的源条件端点来解决这一挑战,其边际分布与对立领域进行对抗匹配。生成的路径通过停止梯度自流匹配进行正则化,以实现中间速度监督,通过离散循环闭合实现正向-反向一致性,以及通过表示路径-速度正则化控制沿轨迹的局部特征变化。在十个翻译方向上,UniCycleFlow在10个任务中的7个任务上实现了最低的FID,并通过单次Euler评估获得了最佳平均FID为55.1。
cs.CV / 20 / 2608.06794
PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model
PAST:用于高效扩散模型的提示自适应采样终止
Abstract
While diffusion models have made significant progress in text-to-image tasks, they still exhibit limitations when directly optimizing downstream objectives. Although Reinforcement Learning (RL) enables targeted optimization, existing methods are generally constrained by low-efficiency fine-tuning and sparse rewards. To address these challenges, we propose PAST, which provides differentiated rewards while adaptively regulating training episode length by jointly perceiving denoising progress and prompt difficulty. Specifically, we design an intrinsic reward paradigm to compensate for sparse extrinsic rewards and guide the model to explore paths that diverge more efficiently from noise patterns. We further provide theoretical justification for intrinsic rewards. Then, PAST dynamically monitors denoising completion and semantic alignment between image structures and prompt semantics. When both metrics satisfy generation requirements, the system adaptively terminates training. This enables appropriate allocation of episode lengths based on prompt difficulty and the current generation process. Finally, based on the predicted residual noise level, we establish a dual adaptive coordination mechanism. Specifically, it not only balances the extrinsic and intrinsic rewards but also balances the exploration and convergence. Experimental results demonstrate that PAST enhances computational efficiency of existing RL fine-tuning methods by up to 66.7%, while improving preference optimization quality by up to 29.5% through its dual adaptive regulation mechanism.
Chinese Translation
尽管扩散模型在文本到图像任务中取得了显著进展,但在直接优化下游目标时仍然存在局限性。虽然强化学习(Reinforcement Learning, RL)能够实现针对性的优化,但现有方法通常受到低效微调和稀疏奖励的限制。为了解决这些挑战,我们提出了PAST,它通过共同感知去噪进展和提示难度,提供差异化奖励,同时自适应地调节训练回合长度。具体而言,我们设计了一种内在奖励范式,以补偿稀疏的外在奖励,并引导模型探索更有效地偏离噪声模式的路径。我们进一步为内在奖励提供了理论依据。然后,PAST动态监测去噪完成度以及图像结构与提示语义之间的语义对齐。当这两个指标满足生成要求时,系统自适应地终止训练。这使得根据提示难度和当前生成过程适当地分配回合长度成为可能。最后,基于预测的残余噪声水平,我们建立了双重自适应协调机制。具体而言,它不仅平衡外在和内在奖励,还平衡探索与收敛。实验结果表明,PAST通过其双重自适应调节机制将现有RL微调方法的计算效率提高了多达66.7%,同时将偏好优化质量提高了多达29.5%。
cs.CV / 21 / 2608.06801
AdvTiles: Physical Adversarial Camouflage Clothing against Person Detectors via Learnable Tiles
AdvTiles:一种基于可学习瓷砖的针对行人检测器的物理对抗伪装服装
Abstract
Physical adversarial attacks against person detectors have evolved from localized patches to full-body textures. However, achieving both visual naturalness and strong attack effectiveness remains challenging. Existing natural-looking methods typically optimize camouflage textures as a whole, limiting the flexibility to refine local adversarial patterns and their spatial arrangement. To address this issue, we propose AdvTiles, a physical adversarial camouflage framework built from learnable tiles, enabling strong attack performance while preserving a natural camouflage appearance. Specifically, we use a Straight-through (ST) Gumbel-Softmax estimator for differentiable tile selection, enabling joint optimization of tile patterns and spatial layouts. This design provides fine-grained control over adversarial texture generation. To improve robustness in diverse physical conditions, we further optimize the camouflage through differentiable 3D Gaussian Splatting rendering with variations in viewpoints, scales, illuminations and backgrounds. Extensive experiments across multiple detectors demonstrate that AdvTiles achieves an average ASR of 86.2%, outperforming existing state-of-the-art attack methods. We further fabricate the optimized camouflage into wearable adversarial clothing, validating its effectiveness in real-world scenarios across diverse distances, angles and backgrounds.
Chinese Translation
针对行人检测器的物理对抗攻击已从局部补丁演变为全身纹理。然而,实现视觉自然性与强攻击效果的平衡仍然具有挑战性。现有的自然外观方法通常将伪装纹理作为整体进行优化,这限制了对局部对抗模式及其空间排列的灵活调整。为了解决这一问题,我们提出了AdvTiles,一种基于可学习瓷砖构建的物理对抗伪装框架,能够在保持自然伪装外观的同时实现强大的攻击性能。具体而言,我们使用了直通(Straight-through, ST)Gumbel-Softmax估计器进行可微分的瓷砖选择,从而实现瓷砖模式和空间布局的联合优化。该设计提供了对对抗纹理生成的细粒度控制。为了提高在多样化物理条件下的鲁棒性,我们进一步通过可微分的3D高斯点云渲染优化伪装,考虑了视角、尺度、光照和背景的变化。针对多个检测器的广泛实验表明,AdvTiles实现了86.2%的平均攻击成功率(ASR),超越了现有的最先进攻击方法。我们还将优化后的伪装制成可穿戴的对抗服装,验证其在不同距离、角度和背景下的实际效果。
cs.CV / 22 / 2608.06832
Bend the Basics: Degradation-Aware Deformable Tokenization for All-in-One Image Restoration
弯曲基础:面向降解感知的可变形标记化用于一体化图像恢复
Abstract
All-in-one image restoration seeks a single model that can recover images degraded by diverse and spatially non-uniform corruptions. However, many unified Transformers rely on fixed patch partitioning: task/degradation condition is injected only into the backbone blocks after tokenization, leaving the embedding and reconstruction stages insensitive to local degradation variations. In contrast to previous approaches, we present Flexible Image Transformer (FIT) that explicitly models degradation awareness across the entire pipeline, from patch sampling to pixel reconstruction. Specifically, FIT employs a lightweight Degradation Encoder to predict a global degradation vector $\mathbf{g}$ and a spatial degradation map $\mathbf{M}$ from local degradation severity, which jointly condition the patch embedding and unembedding through adaptive deformation. Moreover, to improve robustness across degradation types, we introduce a task-token dropout strategy that regularizes task conditioning during training. On five standard benchmarks (BSD68, Rain100L, SOTS, GoPro, and LOLv1), FIT achieves state-of-the-art performance with 30.72 dB average PSNR on the five-degradation setting and 32.83 dB on the three-degradation setting, outperforming recent unified restoration methods by +0.5$\sim$1.1 dB. Moreover, the learned offsets provide a direct handle for visualizing degradation-aware spatial adaptation.
Chinese Translation
一体化图像恢复旨在寻求一个能够恢复因多样且空间上不均匀的损坏而降解的图像的单一模型。然而,许多统一的 Transformer 依赖于固定的补丁划分:任务/降解条件仅在标记化后注入到主干模块中,导致嵌入和重建阶段对局部降解变化不敏感。与之前的方法不同,我们提出了灵活图像 Transformer(Flexible Image Transformer, FIT),该模型在整个流程中显式地建模降解感知,从补丁采样到像素重建。具体而言,FIT 采用轻量级降解编码器(Degradation Encoder)从局部降解严重性预测全局降解向量 $oldsymbol{g}$ 和空间降解图 $oldsymbol{M}$,这两者通过自适应变形共同调节补丁嵌入和解嵌入。此外,为了提高对不同降解类型的鲁棒性,我们引入了一种任务标记丢弃策略,在训练过程中对任务条件进行正则化。在五个标准基准测试(BSD68、Rain100L、SOTS、GoPro 和 LOLv1)上,FIT 在五种降解设置下实现了 30.72 dB 的平均 PSNR,在三种降解设置下实现了 32.83 dB,超越了近期的统一恢复方法,提升幅度为 +0.5$ ilde{ }$1.1 dB。此外,学习到的偏移量为可视化降解感知的空间适应提供了直接的手段。
cs.CV / 23 / 2608.06836
GOPI: Generation-Oriented 3D Pose Inference for Furniture Insertion from Single-View RGB-D Indoor Scenes
GOPI:基于生成的3D姿态推断用于单视图RGB-D室内场景中的家具插入
Abstract
We study the problem of inserting new furniture into indoor scene images. Under masked single-view 2D image-plane conditioning, however, the physical scale of the inserted furniture relative to the scene cannot be uniquely determined, making physically grounded furniture placement underdetermined from image evidence alone. We therefore reformulate the task as a combination of 3D pose inference and geometry-guided image generation, where estimating a geometrically plausible 3D placement is essential for reliable synthesis. To this end, we propose a two-stage framework. For 3D placement, we introduce GOPI, a generation-oriented 3D pose inference framework that addresses the underdetermined nature of single-view furniture insertion through data-driven iterative inference, producing geometrically plausible object placements. For image generation, we develop a geometry-guided conditioning strategy that projects the inferred 3D pose into the image plane as a pixel-aligned constraint, enforcing consistency between the synthesized image and the underlying 3D geometry. Experimental results validate the proposed framework from both 3D pose estimation and image synthesis perspectives. For 3D placement, GOPI produces poses with stronger geometric feasibility and better consistency with reference layouts than direct regression and vanilla baselines. For image synthesis, our method preserves alignment with the projected 3D geometry across different furniture scales, showing stable projection-generation alignment across the tested furniture scales.
Chinese Translation
我们研究了将新家具插入室内场景图像的问题。然而,在掩蔽的单视图二维图像平面条件下,插入家具的物理尺度相对于场景无法唯一确定,使得仅凭图像证据无法实现物理上合理的家具放置。因此,我们将任务重新表述为3D姿态推断与几何引导图像生成的结合,其中估计几何上合理的3D放置对于可靠的合成至关重要。为此,我们提出了一个两阶段的框架。对于3D放置,我们引入了GOPI,一个面向生成的3D姿态推断框架,通过数据驱动的迭代推断解决单视图家具插入的非确定性特征,生成几何上合理的物体放置。对于图像生成,我们开发了一种几何引导的条件策略,将推断得到的3D姿态投影到图像平面作为像素对齐的约束,确保合成图像与基础3D几何之间的一致性。实验结果从3D姿态估计和图像合成的角度验证了所提出框架的有效性。在3D放置方面,GOPI生成的姿态在几何可行性和与参考布局的一致性上优于直接回归和基础模型。在图像合成方面,我们的方法在不同家具尺度下保持与投影3D几何的对齐,显示出在测试的家具尺度下稳定的投影-生成对齐。
cs.CV / 24 / 2608.06841
ECAD: Expanding Class-Agnostic Detection Beyond Thing-Centric Objectness
ECAD:超越以物体为中心的物体性扩展的类别无关检测
Abstract
Object detection is a fundamental task in visual perception, providing structured region representations for recognition, grounding, reasoning, and interaction. However, existing detection paradigms largely inherit a thing-centric notion of objectness, where detectors are mainly trained to localize discrete and countable object instances. Consequently, many semantically meaningful visual elements, such as sky, road, grassland, water, and sports courts, are often absorbed into the background despite their importance for scene understanding and spatial reasoning. In this paper, we formulate Expanded Class-Agnostic Detection (ECAD), a new setting that aims to discover category-agnostic visual candidates beyond conventional thing-centric objects. To support this setting, we construct BTCO-Bench, a Beyond Thing-Centric Objectness benchmark with category-agnostic box annotations covering both real-world and cross-domain scenarios. We further propose ECADet, a lightweight DETR-based detector built upon a frozen DINOv3 encoder, and introduce Geometry-Aware Expert Regression (GAER) and Prototype-Guided Query Modulation (PGQM) to improve localization and objectness estimation for diverse visual elements, respectively. Extensive experiments show that ECADet consistently outperforms representative class-agnostic and proposal-based detectors on BTCO-Bench, demonstrating the effectiveness of expanded objectness discovery. Code and benchmark will be released.
Chinese Translation
物体检测是视觉感知中的一项基础任务,为识别、定位、推理和交互提供结构化的区域表示。然而,现有的检测范式在很大程度上继承了以物体为中心的物体性概念,检测器主要被训练以定位离散且可计数的物体实例。因此,许多具有语义意义的视觉元素,如天空、道路、草地、水体和运动场,尽管对场景理解和空间推理至关重要,常常被吸收到背景中。在本文中,我们提出了扩展类别无关检测(Expanded Class-Agnostic Detection,ECAD),这是一个新的设置,旨在发现超越传统以物体为中心的物体的类别无关视觉候选。为了支持这一设置,我们构建了BTCO-Bench,这是一个超越以物体为中心的物体性基准,具有涵盖现实世界和跨领域场景的类别无关框标注。我们进一步提出了ECADet,这是一种基于轻量级DETR的检测器,建立在冻结的DINOv3编码器之上,并引入了几何感知专家回归(Geometry-Aware Expert Regression,GAER)和原型引导查询调制(Prototype-Guided Query Modulation,PGQM),以分别提高对多样视觉元素的定位和物体性估计。大量实验表明,ECADet在BTCO-Bench上始终优于代表性的类别无关和基于提议的检测器,证明了扩展物体性发现的有效性。代码和基准将会发布。
cs.CV / 25 / 2608.06850
RegionDet: A Benchmark for Region Detection Beyond Object Instances
RegionDet:超越物体实例的区域检测基准
Abstract
Object detection is a fundamental task in computer vision and has achieved remarkable progress on standard benchmarks by localizing discrete and well-bounded object instances. However, many visual targets in real-world scenarios are not individual objects, but regions defined by visual states, scene context, object relations, and human activities, such as construction areas, damaged road regions, queues, group conversations, and vendor regions. Existing detection benchmarks are mainly built around object instances, providing limited support for systematically evaluating such region targets. To address this gap, we introduce Region Detection, a task that extends conventional object detection beyond object instances, and construct RegionDet, a benchmark for region target localization. RegionDet contains eight region categories, including Construction, Crossing, Damage, Queuing, Talking, Vendor, Waiting, and Walking, with COCO-style bounding-box annotations and evaluation protocols. We systematically evaluate representative closed-set and zero-shot/open-vocabulary detectors on RegionDet. Results show that closed-set detectors can partially learn region-level patterns under supervision, while zero-shot/open-vocabulary detectors struggle severely, revealing the strong object-centric bias of current vision-language detectors. Further analyses highlight key challenges in Region Detection, including weak boundary cues, strong context dependency, and insufficient relation-level region understanding. The RegionDet will be released.
Chinese Translation
物体检测是计算机视觉中的一项基础任务,在标准基准上通过定位离散且边界清晰的物体实例取得了显著进展。然而,许多现实场景中的视觉目标并不是单独的物体,而是由视觉状态、场景上下文、物体关系和人类活动定义的区域,例如施工区域、受损道路区域、排队、群体对话和商贩区域。现有的检测基准主要围绕物体实例构建,无法系统地评估这些区域目标。为了解决这一空白,我们提出了区域检测(Region Detection)这一任务,扩展了传统物体检测的范围,构建了区域检测基准(RegionDet),用于区域目标定位。RegionDet包含八个区域类别,包括施工(Construction)、过马路(Crossing)、损坏(Damage)、排队(Queuing)、对话(Talking)、商贩(Vendor)、等待(Waiting)和行走(Walking),并提供COCO风格的边界框注释和评估协议。我们系统地评估了代表性的封闭集和零样本/开放词汇检测器在RegionDet上的表现。结果表明,封闭集检测器在监督下可以部分学习区域级模式,而零样本/开放词汇检测器则表现不佳,揭示了当前视觉-语言检测器的强物体中心偏差。进一步分析强调了区域检测中的关键挑战,包括弱边界线索、强上下文依赖性和不足的关系级区域理解。RegionDet将被发布。
cs.CV / 26 / 2608.06865
Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection
用于可泛化深度伪造视频检测的多智能体法医学推理
Abstract
The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deepfake video benchmarks provide limited coverage of recent synthesis methods and generally lack reliable fine-grained textual annotations. Meanwhile, conventional detectors and multimodal large language models (MLLMs), whether operating as a single model or relying on a single analytical perspective, often fail to capture subtle forgery artifacts, limiting their generalization to emerging AI-generated methods. To address these limitations, we introduce FaceVid-Forensics-100K, a large-scale deepfake video dataset comprising 100,000 videos and spanning 33 synthesis methods across face swapping, face reenactment, and entire-face synthesis, including recent generators such as Seedance 2.0. The dataset provides fine-grained textual annotations of visual observations and verdict-consistent forensic explanations, automatically synthesized through a multi-model aggregation and conflict-resolution pipeline powered by advanced MLLMs. Building on this benchmark, we propose a multi-agent forensic reasoning framework that employs four specialized domain-expert agents to independently analyze forgery cues from four perspectives: texture, lighting, motion, and physics. A judge agent then reconciles their reports to produce a final prediction together with an explanation. Extensive evaluations on out-of-domain test sets show that, despite being composed entirely of small open-source MLLMs, our framework outperforms all methods including closed-source GPT and Gemini models and ranks first across all reported metrics on this benchmark. The project page is available at https://xavierjiezou.github.io/ARGUS/.
Chinese Translation
恶意使用生成性人工智能创建高度逼真的深度伪造视频引发了严重的伦理担忧,并对人工智能安全提出了重大挑战。然而,现有的深度伪造视频基准在最近的合成方法覆盖范围上有限,并且通常缺乏可靠的细粒度文本注释。同时,无论是作为单一模型运行的传统检测器,还是依赖单一分析视角的多模态大型语言模型(MLLMs),往往无法捕捉到微妙的伪造伪迹,从而限制了它们对新兴人工智能生成方法的泛化能力。为了解决这些局限性,我们引入了FaceVid-Forensics-100K,这是一个大规模深度伪造视频数据集,包含100,000个视频,涵盖了面部交换、面部重演和全脸合成等33种合成方法,包括最近的生成器Seedance 2.0。该数据集提供了视觉观察的细粒度文本注释和一致的法医学解释,这些注释通过一个由先进的MLLMs驱动的多模型聚合和冲突解决管道自动合成。在此基准的基础上,我们提出了一种多智能体法医学推理框架,利用四个专业领域专家代理从纹理、光照、运动和物理四个角度独立分析伪造线索。然后,法官代理整合他们的报告,产生最终预测及其解释。在域外测试集上的广泛评估表明,尽管我们的框架完全由小型开源MLLMs构成,但其性能超越了所有方法,包括闭源的GPT和Gemini模型,并在该基准的所有报告指标中排名第一。项目页面可访问 https://xavierjiezou.github.io/ARGUS/。
cs.CV / 27 / 2608.06869
DAEP: Difficulty-Aware Evidence Planning for Medical Video Corpus Temporal Answer Grounding
DAEP:面向困难的医学视频语料库时间答案定位的证据规划
Abstract
We describe DAEP, team BIGC's submission to NLPCC 2026 Shared Task 1 Track 3: Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). The task requires retrieving the target video from 50 candidates and localizing the answer-supporting span. DAEP ranks videos with subtitle, visual, and procedural-context evidence, expands high-scoring anchors into temporal spans, and reranks spans for final output. Its main design is to convert the task-provided simple/complex input label into an inference-time evidence plan controlling modality weights, Top-K aggregation, boundary threshold, expansion length, and reranking strength. In the official evaluation, BIGC ranks first among ten systems with an Average score of 0.2728. Validation ablations show that visual evidence, procedural context, and difficulty-aware planning improve ranking quality, with the largest gain on complex questions.
Chinese Translation
我们描述了DAEP,这是BIGC团队提交给NLPCC 2026共享任务1第3轨道:面向困难的医学视频语料库时间答案定位(DA-TAGVC)。该任务要求从50个候选视频中检索目标视频,并定位支持答案的片段。DAEP通过字幕、视觉和程序上下文证据对视频进行排名,将高分锚点扩展为时间跨度,并对片段进行重新排名以生成最终输出。其主要设计是将任务提供的简单/复杂输入标签转换为推理时的证据规划,以控制模态权重、Top-K聚合、边界阈值、扩展长度和重新排名强度。在官方评估中,BIGC在十个系统中排名第一,平均得分为0.2728。验证消融实验表明,视觉证据、程序上下文和面向困难的规划提高了排名质量,尤其在复杂问题上获得了最大的提升。
cs.CV / 28 / 2608.06876
FedVAR: Prototype-Aligned Federated Framework for Video Anomaly Recognition
FedVAR:面向视频异常识别的原型对齐联邦框架
Abstract
In the era of Industrial Internet of Things (IIoT) and Cyber-Physical Systems (CPS), Federated Learning (FL) offers a promising decentralized intelligence paradigm for Video Anomaly Recognition (VAR). This task is vital for maintaining high-fidelity Digital Twins and ensuring safety in mission-critical environments. However, the inherent data heterogeneity across distributed edge clients leads to a fundamental challenge known as semantic misalignment, where clients learn divergent feature representations of "normal" and "abnormal" events. The problem becomes particularly pronounced in VAR, where the presence of diverse and fine-grained anomaly categories leads each client to develop distinct semantic interpretations of abnormality. Existing federated methods primarily focus on binary anomaly detection and fail to address this misalignment, preventing effective fine-grained recognition. In this paper, we introduce FedVAR, a weakly-supervised FL framework explicitly designed for VAR. Leveraging the rich representations of Vision-Language Models (VLMs), FedVAR employs a prototype-based alignment mechanism that creates a shared semantic anchor for all clients to re-center and align their visual and textual feature spaces. This process enforces a consistent representation of "normality" across the decentralized network, directly mitigating semantic misalignment and enabling robust prompt-learning of anomaly direction vectors with minimal communication overhead. We conduct extensive experiments on challenging benchmarks under various non-IID data partitioning schemes, unseen domains, and novel anomaly classes. The results demonstrate that FedVAR consistently outperforms state-of-the-art federated baselines, establishing a robust framework for distributed intelligence in video-based CPS.
Chinese Translation
在工业物联网(IIoT)和网络物理系统(CPS)时代,联邦学习(FL)为视频异常识别(VAR)提供了一种有前景的去中心化智能范式。该任务对于维持高保真数字双胞胎和确保关键任务环境的安全至关重要。然而,分布式边缘客户端之间固有的数据异构性导致了一个根本性挑战,即语义不对齐,客户端学习到的“正常”和“异常”事件的特征表示存在差异。这个问题在VAR中尤为突出,因为多样化和细粒度的异常类别使得每个客户端发展出不同的异常语义解释。现有的联邦方法主要集中于二元异常检测,未能解决这种不对齐问题,从而阻碍了有效的细粒度识别。在本文中,我们提出了FedVAR,一种专门为VAR设计的弱监督FL框架。FedVAR利用视觉-语言模型(VLMs)的丰富表示,采用基于原型的对齐机制,为所有客户端创建一个共享的语义锚点,以重新中心化和对齐其视觉和文本特征空间。该过程在去中心化网络中强制执行“正常性”的一致表示,直接缓解了语义不对齐,并以最小的通信开销实现异常方向向量的稳健提示学习。我们在各种非独立同分布(non-IID)数据划分方案、未见领域和新颖异常类别下,对具有挑战性的基准进行了广泛实验。结果表明,FedVAR在各项指标上始终优于最先进的联邦基线,为基于视频的CPS中的分布式智能建立了一个稳健的框架。
cs.CV / 29 / 2608.06878
ControlRef: Efficient Layout-Guided Multi-Instance Generation via Anchored 4D-RoPE
ControlRef:通过锚定的4D-RoPE实现高效的布局引导多实例生成
Abstract
Layout-guided multi-instance generation is essential for controllable image synthesis in Multi-Modal Diffusion Transformers (MM-DiTs). However, integrating this capability into unified architectures remains challenging. Prior frameworks rely on redundant full-resolution canvas padding and Shifted-RoPE to manage multiple reference images. This mechanism drastically inflates computational overhead for sparse layouts and disrupts critical low-frequency RoPE features, creating a severe spatial-frequency compromise that blurs absolute spatial correspondence. To overcome these limitations, we propose ControlRef, a highly efficient and precise multi-instance synthesis framework. ControlRef utilizes a Unified Instance-Layout Control (UILC) attention mask to strictly decouple inter-instance semantic interactions and enforce precise regional binding. To further promote region-level spatial alignment, we introduce Anchored 4D-RoPE, a novel positional encoding mechanism that directly anchors tokens to their absolute geometric centers. By pre-aligning reference images to their corresponding bounding box resolutions, physically anchoring both layout and reference tokens to their absolute geometric centers, and stacking the references along the z-axis, Anchored 4D-RoPE natively preserves spatial priors and mitigates the spatial-frequency compromise without lossy shifting. Extensive experiments demonstrate that ControlRef achieves state-of-the-art visual fidelity and localization accuracy, while concurrently slashing inference latency by over 80% in sparse layouts and reducing memory overhead by 50% in dense scenarios.
Chinese Translation
布局引导的多实例生成对于多模态扩散变换器(Multi-Modal Diffusion Transformers, MM-DiTs)中的可控图像合成至关重要。然而,将这一能力整合到统一架构中仍然面临挑战。之前的框架依赖冗余的全分辨率画布填充和Shifted-RoPE来管理多个参考图像。这一机制在稀疏布局下大幅增加了计算开销,并破坏了关键的低频RoPE特征,造成严重的空间频率妥协,模糊了绝对空间对应关系。为了解决这些限制,我们提出了ControlRef,一个高效且精确的多实例合成框架。ControlRef利用统一实例-布局控制(Unified Instance-Layout Control, UILC)注意力掩码严格解耦实例间的语义交互,并强制实施精确的区域绑定。为了进一步促进区域级空间对齐,我们引入了锚定的4D-RoPE(Anchored 4D-RoPE),这是一种新颖的位置编码机制,直接将标记锚定到其绝对几何中心。通过将参考图像预对齐到其对应的边界框分辨率,物理上将布局和参考标记锚定到其绝对几何中心,并沿z轴堆叠参考,锚定的4D-RoPE本质上保留了空间先验,并在不损失的情况下减轻了空间频率妥协。大量实验表明,ControlRef在稀疏布局中实现了最先进的视觉保真度和定位精度,同时将推理延迟降低了超过80%,在密集场景中减少了50%的内存开销。
cs.CV / 30 / 2608.06886
HazeSpikeMamba: Coupling Spiking-Inspired and State-Space Features for Self-Supervised Real-World Dehazing
HazeSpikeMamba:结合脉冲启发和状态空间特征的自监督真实世界去雾
Abstract
Dehazing networks are commonly trained on synthetic hazy-clear pairs, but their performance often drops on real photographs. Synthetic haze generated using the atmospheric scattering model does not fully capture the variability of real haze, and paired real hazy-clear images are scarce. In this work, we propose HazeSpikeMamba, a compact dehazing framework that combines a spiking-inspired local path and an attentive state-space global path in a multi-scale U-Net. The local path uses TPCNNSpike, a new spike-emission scheme inspired by the neighborhood coupling of Pulse-Coupled Neural Network (PCNN). Unlike grouped directional scanning, TPCNNSpike updates all neurons in parallel using the previous firing states of their Gaussian-weighted neighborhoods. The global path adapts the Attentive State-Space Module of MambaIRv2, retaining semantic prompting and sequence reordering while removing the window self-attention branch. Its state-space processing models long-range dependencies with complexity linear in sequence length. For target-domain adaptation, a frozen degradation network, pretrained on paired NH-HAZE data, re-synthesizes haze from the dehazed prediction. The reconstruction error updates only the final restoration layers of HazeSpikeMamba without haze-free labels during adaptation. A shared checkpoint is adapted once on each complete unlabeled target set, making the evaluation dataset-level and transductive rather than zero-shot or per-image optimization. The forward network contains 2.02M active parameters and requires 13.27G nominal MACs (measured with thop at 256x256 input). This adaptation consistently improves BRISQUE and NIMA on RTTS, URHI, and HSTS. On RTTS, BRISQUE decreases from 30.13 to 27.72 and NIMA increases from 4.13 to 4.87. Under this transductive protocol, the adapted model also achieves the best BRISQUE and NIMA on URHI and HSTS among the compared methods.
Chinese Translation
去雾网络通常在合成的雾霾-清晰图像对上进行训练,但在真实照片上的性能往往下降。使用大气散射模型生成的合成雾霾并不能完全捕捉真实雾霾的变异性,而配对的真实雾霾-清晰图像则稀缺。在本研究中,我们提出了HazeSpikeMamba,一个紧凑的去雾框架,结合了多尺度U-Net中的脉冲启发局部路径和注意力状态空间全局路径。局部路径使用TPCNNSpike,这是一种新的脉冲发射方案,灵感来源于脉冲耦合神经网络(PCNN)的邻域耦合。与分组方向扫描不同,TPCNNSpike使用其高斯加权邻域的先前发火状态并行更新所有神经元。全局路径适应了MambaIRv2的注意力状态空间模块,保留了语义提示和序列重排序,同时移除了窗口自注意力分支。其状态空间处理以序列长度为线性复杂度建模长程依赖关系。为了进行目标领域适应,冻结的降解网络在配对的NH-HAZE数据上进行预训练,从去雾预测中重新合成雾霾。重建误差仅在适应过程中更新HazeSpikeMamba的最终恢复层,而不需要无雾标签。共享检查点在每个完整的未标记目标集上适应一次,使得评估数据集级别和迁移式,而不是零-shot或逐图像优化。前向网络包含2.02M个活动参数,并需要13.27G的名义MAC(在256x256输入下使用thop测量)。这种适应在RTTS、URHI和HSTS上持续改善BRISQUE和NIMA。在RTTS上,BRISQUE从30.13降至27.72,NIMA从4.13升至4.87。在这种迁移协议下,适应后的模型在比较方法中也在URHI和HSTS上达到了最佳的BRISQUE和NIMA。
cs.CV / 31 / 2608.06901
Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models
一次剪枝:无重训练的任务无关视觉语言模型剪枝
Abstract
Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidly increasing computational and memory requirements pose significant challenges for deployment in constrained environments. Existing pruning strategies often depend on task-specific criteria or LLM-oriented importance measures, making them unsuitable for task-agnostic pruning, where no task-specific samples are available at pruning time and the pruned model remains broadly applicable. We introduce a retraining-free VLM pruning framework called PORTA that derives a task- and modality-agnostic importance formulation based on activation variation, estimated from generic calibration data, which reliably captures feature-level representation utility across modalities. PORTA further incorporates an adaptive sparsity allocation mechanism that assigns layer-wise pruning ratios based on output feature variability, avoiding the limitations of uniform sparsity and reducing performance degradation at high compression levels. Extensive experiments across VLM architectures, such as CLIP, BLIP, and Qwen2-VL, demonstrate that PORTA achieves competitive downstream performance under high sparsity without requiring any retraining, supporting efficient VLM compression. Code is available at https://github.com/cau-hai-lab/PORTA.git.
Chinese Translation
视觉语言模型(VLMs)通过大规模预训练在多种多模态任务中实现了显著的泛化能力,但其快速增长的计算和内存需求在受限环境中的部署面临重大挑战。现有的剪枝策略通常依赖于特定任务的标准或针对大型语言模型(LLM)的重要性度量,使其不适用于任务无关的剪枝,因为在剪枝时没有特定任务的样本可用,且剪枝后的模型仍然具有广泛的适用性。我们提出了一种无重训练的VLM剪枝框架,称为PORTA,该框架基于从通用校准数据中估计的激活变化推导出任务和模态无关的重要性公式,可靠地捕捉了跨模态的特征级表示效用。PORTA进一步结合了一种自适应稀疏分配机制,根据输出特征的变异性分配层级剪枝比例,避免了均匀稀疏的局限性,并在高压缩水平下减少性能下降。针对CLIP、BLIP和Qwen2-VL等VLM架构的广泛实验表明,PORTA在高稀疏性下实现了具有竞争力的下游性能,而无需任何重训练,从而支持高效的VLM压缩。代码可在 https://github.com/cau-hai-lab/PORTA.git 获取。
cs.CV / 32 / 2608.06913
MuST-VAD: Mutual Structured Learning for Video Anomaly Detection
MuST-VAD:用于视频异常检测的互助结构学习
Abstract
In this paper, we propose MuST-VAD, a mutual structured learning framework for weakly supervised video anomaly detection (VAD) in which an anomaly detector and a large vision-language model (LVLM) exchange their acquired knowledge. Detectors in weakly supervised VAD learn anomaly scores from features extracted by a fixed, task-agnostic backbone. These fixed features bound the achievable detection accuracy. Recent methods therefore transfer LVLM semantics into the detector as richer features. However, this transfer is one-way: what the detector learns about the target videos never returns to the LVLM. MuST-VAD extends the one-way transfer into a bidirectional learning loop. In this loop, the latest detector predictions supervise the LVLM adaptation, and the adapted LVLM returns updated representations that retrain the detector; the two models alternate these updates over small video groups. Both models train on detector-selected key clips, while confidence weighting and annotation-anchored question answering keep the exchanged supervision reliable. On UCF-Crime, our mutual learning improves the one-pass transfer baseline from 88.15% to 88.63% AUROC and from 37.25% to 42.46% average precision (AP), outperforming the state-of-the-art method in AP by 4.13 points.
Chinese Translation
在本文中,我们提出了MuST-VAD,一种用于弱监督视频异常检测(VAD)的互助结构学习框架,其中异常检测器和大型视觉语言模型(LVLM)相互交流所获得的知识。弱监督VAD中的检测器从由固定的、与任务无关的主干网络提取的特征中学习异常分数。这些固定特征限制了可实现的检测准确性。因此,最近的方法将LVLM的语义转移到检测器中,以提供更丰富的特征。然而,这种转移是单向的:检测器对目标视频的学习不会反馈给LVLM。MuST-VAD将单向转移扩展为双向学习循环。在这个循环中,最新的检测器预测监督LVLM的适应,而适应后的LVLM返回更新的表示,以重新训练检测器;两个模型在小的视频组上交替进行这些更新。两个模型都在检测器选择的关键片段上进行训练,同时置信加权和基于注释的问题回答确保了交换监督的可靠性。在UCF-Crime数据集上,我们的互助学习将一次性转移基线的AUROC从88.15%提高到88.63%,将平均精度(AP)从37.25%提高到42.46%,在AP上超越了最先进的方法4.13个百分点。
cs.CV / 33 / 2608.06914
RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections
RibAssist 3D:基于CT衍生投影的双平面肋骨骨折检测、处理及选择性3D定位
Abstract
Rib fractures are common, clinically significant, and time-consuming to localize on computed tomography (CT). We ask whether fractures detected in two orthogonal projections (anteroposterior, AP, and lateral) can be paired across views and triangulated into reliable 3D fracture points at a controlled rate of false 3D outputs. We answer this with a staged diagnostic study. Biplanar geometry is exact: detector-predicted centers reconstruct to median 4.0 mm 3D error when correspondence is correct. On the sealed cohort, dual-view availability reaches 61.1% and the candidate graph contains a correct pair for 58.4% of fractures. The binding limitation is not geometry or localization but confidence-limited cross-view correspondence. Lateral-detector retraining lifts dual-view availability (0.52 to 0.76 in development) and moves the frontier from 0% to 2.44% recall at 10 mm. When the policy commits a correct pair, the emitted point is geometrically accurate (sealed median 1.49 mm, rib-exact 93%). A pre-specified pass on the untouched 55-case cohort promotes 15 of 601 fractures to correct 3D localizations at 0.436 false 3D points per case, yielding 2.50% end-to-end commitment yield. The contribution is validated biplanar reconstruction geometry with high conditional localization fidelity, a staged identification of cross-view correspondence confidence as the effective bottleneck, and a selective assistive workflow that preserves uncertain findings rather than a standalone automatic reconstructor.
Chinese Translation
肋骨骨折常见且具有临床重要性,但在计算机断层扫描(CT)上定位耗时。我们探讨在两个正交投影(前后位,AP,和侧位)中检测到的骨折是否可以跨视图配对,并以可控的假阳性3D输出率三角测量出可靠的3D骨折点。我们通过分阶段的诊断研究回答了这一问题。双平面几何是精确的:当配对正确时,探测器预测的中心重建至中位数4.0毫米的3D误差。在封闭队列中,双视图可用性达到61.1%,候选图包含58.4%的骨折的正确配对。限制因素不是几何或定位,而是受信心限制的跨视图配对。侧面探测器的再训练提高了双视图可用性(在开发中从0.52提升至0.76),并将前沿从0%提升至10毫米时的2.44%召回率。当策略承诺一个正确的配对时,发出的点在几何上是准确的(封闭中位数1.49毫米,肋骨准确率93%)。在未处理的55例队列上预先指定的通过率促进了601例骨折中15例的正确3D定位,每例的假阳性3D点为0.436,最终实现2.50%的端到端承诺产出。该贡献验证了具有高条件定位保真度的双平面重建几何,分阶段识别跨视图配对信心作为有效瓶颈,以及一种选择性辅助工作流程,该流程保留不确定的发现,而不是单独的自动重建器。
cs.CV / 34 / 2608.06919
Vernata: Self-Supervised Learning of LiDAR Point Representations
Vernata:激光雷达点表示的自监督学习
Abstract
LiDAR serves as a primary sensing modality for robots operating in outdoor environments. However, the performance of deep learning models in this domain is severely limited by the scarcity of labeled data, a direct result of the high cost of 3D annotation. Self-supervised learning addresses this scarcity by learning general-purpose features from unlabeled data. In this work, we present a multi-modal, multi-teacher distillation framework for self-supervised learning on outdoor LiDAR point clouds. Building upon the Sonata architecture, we introduce Vernata, consisting of three extensions: sparse view augmentation to improve robustness against varying point densities, a memory bank mechanism to stabilize resource-constrained training, and cross-modal distillation utilizing dense, high-resolution 2D image features to enable fine-grained semantic guidance. We evaluate our method on the GrandTour, TartanGround, and Waymo datasets, as well as data collected from our own robotic platforms. Our experiments demonstrate a significant performance improvement over Sonata baselines, yielding mIoU scores of 54.7 on TartanGround (+5.9 points, +12.1%) and 57.1 on Waymo (+7.3 points, +14.7%). Finally, we show that the self-supervised approach maintains strong performance even in reduced-modality settings (lacking color or normals), achieving competitive mIoU scores of 49.4 and 50.2 on the respective datasets.
Chinese Translation
激光雷达作为在户外环境中操作的机器人主要传感方式。然而,深度学习模型在这一领域的性能受到标注数据稀缺的严重限制,这直接源于三维标注的高成本。自监督学习通过从未标注数据中学习通用特征来解决这一稀缺问题。在本研究中,我们提出了一种多模态、多教师蒸馏框架,用于户外激光雷达点云的自监督学习。在Sonata架构的基础上,我们引入了Vernata,包含三个扩展:稀疏视图增强以提高对不同点密度的鲁棒性、内存库机制以稳定资源受限的训练,以及利用密集高分辨率二维图像特征的跨模态蒸馏,以实现细粒度的语义指导。我们在GrandTour、TartanGround和Waymo数据集上评估了我们的方法,以及从我们自己的机器人平台收集的数据。实验结果表明,我们的方法在Sonata基线之上显著提高了性能,在TartanGround上获得了54.7的mIoU分数(+5.9点,+12.1%),在Waymo上获得了57.1的mIoU分数(+7.3点,+14.7%)。最后,我们展示了自监督方法在降低模态设置(缺乏颜色或法线)时仍能保持强劲的性能,在相应数据集上实现了49.4和50.2的竞争性mIoU分数。
cs.CV / 35 / 2608.06929
MaskFlow: Precise, Consistent and Seamless Regional Image Editing
MaskFlow:精确、一致且无缝的区域图像编辑
Abstract
Regional image editing has attracted considerable attention for its spatial controllability. Although instruction-based and mask-reference-based editing methods can achieve strong semantic alignment, reliable regional control remains challenging, where an edit must be accurately localized and naturally integrated with the preserved context. We propose \textbf{MaskFlow}, a training framework for precise localization, consistent background preservation, and seamless boundary transitions. MaskFlow incorporates the mask into the probability path and flow-matching objective, coordinating generation within the editable region with source preservation outside it. The proposed Soft-Poisson de-seaming module further refines the predicted vector field during both training and sampling to improve the smooth integration of the edited foreground with the preserved background. We also design a data synthesis pipeline to construct MEData, a mask-based image editing dataset for training regional image editing models and facilitating further research. Experiments on natural scenes and infographic images demonstrate consistent improvements over competing methods in both quantitative and qualitative evaluations.
Chinese Translation
区域图像编辑因其空间可控性而受到广泛关注。尽管基于指令和基于掩膜的编辑方法能够实现强语义对齐,但可靠的区域控制仍然具有挑战性,编辑必须准确定位并自然融入保留的上下文中。我们提出了 extbf{MaskFlow},一个用于精确定位、一致背景保留和无缝边界过渡的训练框架。MaskFlow将掩膜纳入概率路径和流匹配目标中,协调可编辑区域内的生成与其外部的源保留。所提出的Soft-Poisson去接缝模块进一步在训练和采样过程中细化预测的向量场,以改善编辑前景与保留背景的平滑融合。我们还设计了一个数据合成管道来构建MEData,这是一个基于掩膜的图像编辑数据集,用于训练区域图像编辑模型并促进进一步研究。在自然场景和信息图像上的实验表明,在定量和定性评估中,相较于竞争方法,均表现出一致的改进。
cs.CV / 36 / 2608.06930
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward
AVCap:通过细节感知奖励强化音视频联合字幕
Abstract
Detailed audio-video joint captioning is essential for multimodal video understanding and generation. However, prior works are constrained by three main limitations: (1) the scarcity of high-quality public datasets with fine-grained audio-visual joint captions; (2) reinforcement-learning methods that rely on coarse reward signals; and (3) the lack of a benchmark and metric for evaluating detailed audiovisual captions at the atomic level. To address these challenges, we propose: (1) AVCap-100K, a high-quality dataset of 100K temporally aligned, detail-rich audio-video captions; (2) AVCap, a model optimized via Detail-Aware GRPO (Da-GRPO) that achieves state-of-the-art performance among open-source models and matches or surpasses proprietary models on several evaluations; and (3) AVCap-Bench and AVCap-Score, a specialized benchmark and metric for evaluating atomic-level details in audiovisual captions. Our code, models, and datasets are available at https://huggingface.co/collections/Apryle/avcap.
Chinese Translation
详细的音视频联合字幕对于多模态视频理解和生成至关重要。然而,之前的研究受到三大主要限制:(1) 缺乏高质量的公共数据集,且这些数据集包含细粒度的音视频联合字幕;(2) 强化学习方法依赖于粗略的奖励信号;(3) 缺乏用于评估原子级别音视频字幕的基准和指标。为了解决这些挑战,我们提出:(1) AVCap-100K,一个包含10万个时间对齐、细节丰富的音视频字幕的高质量数据集;(2) AVCap,一个通过细节感知的GRPO(Da-GRPO)优化的模型,在开源模型中实现了最先进的性能,并在多个评估中与专有模型相匹配或超越;(3) AVCap-Bench和AVCap-Score,一个专门用于评估音视频字幕原子级别细节的基准和指标。我们的代码、模型和数据集可在 https://huggingface.co/collections/Apryle/avcap 获取。
cs.CV / 37 / 2608.06938
Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning
文本去偏见,相信你的眼睛:基于文本的跨模态转移用于视觉反常识推理
Abstract
The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commonsense reasoning, which requires models to reason beyond common assumptions. Recent studies mainly improve visual counter-commonsense reasoning by enhancing visual inputs, following the assumption that failures originate from insufficient visual grounding. However, our empirical analysis reveals that the bottleneck is not visual perception. MLLMs already capture the relevant visual evidence, and the correct answer exists in their decoding space. Instead, the shared language decoder resolves prior--evidence conflicts by favoring dominant language priors, especially for low-frequency factual scenarios. Motivated by this, we first propose a text-anchored data construction pipeline, whose core component, Fact-Frequency Distillation (FFD), estimates the prior strength of commonsense facts and distills verified counter-commonsense scenarios into a high-quality text corpus. Building upon this corpus, we introduce TACT, a text-anchored post-training framework that debiases the shared language decoder without requiring any visual training data. TACT routes evidence-following and prior-driven reasoning trajectories into different optimization stages, enabling the decoder to resolve prior--evidence conflicts. Across counter-commonsense visual benchmarks, TACT substantially improves visual reasoning while preserving general capabilities, demonstrating effective text-to-vision cross-modal transfer.
Chinese Translation
多模态大型语言模型(MLLMs)的视觉推理能力对于下游应用至关重要,特别是反常识推理,这要求模型超越常见假设进行推理。近期研究主要通过增强视觉输入来改善视觉反常识推理,基于这样一个假设:失败源于视觉基础不足。然而,我们的实证分析表明,瓶颈并不在于视觉感知。MLLMs 已经捕捉到了相关的视觉证据,并且正确答案存在于它们的解码空间中。相反,共享语言解码器通过偏向主导语言先验来解决先验与证据之间的冲突,尤其是在低频事实场景中。基于此,我们首先提出了一种基于文本的数据构建管道,其核心组件为事实频率蒸馏(Fact-Frequency Distillation, FFD),该组件估计常识事实的先验强度,并将经过验证的反常识场景提炼成高质量的文本语料库。在此语料库的基础上,我们引入了 TACT,一个基于文本的后训练框架,能够在不需要任何视觉训练数据的情况下去偏见共享语言解码器。TACT 将证据驱动和先验驱动的推理轨迹引导到不同的优化阶段,从而使解码器能够解决先验与证据之间的冲突。在反常识视觉基准测试中,TACT 显著提高了视觉推理能力,同时保持了整体能力,展示了有效的文本到视觉的跨模态转移。
cs.CV / 38 / 2608.06939
Degradation-Aware Prompt Learning with Cross-Modal Compensation for Adverse Weather Removal
考虑退化的跨模态提示学习用于恶劣天气去除
Abstract
Adverse weather causes diverse and complex image degradations, severely compromising the reliability of computer vision systems. Existing all-in-one restoration models attempt to address multiple degradation types within a unified framework, but often lack explicit spatial and semantic modeling of degradation characteristics, limiting their adaptability to diverse weather conditions. To address this limitation, we propose a Degradation-Aware Cross-Modal Prompt Compensation Network (DCMPC-Net) that leverages cross-modal degradation cues from a pretrained vision-language model to condition restoration features within a unified backbone. Specifically, our DCMPC-Net mainly consists of the Cross-Modal Prompt Generator (CMPG), Prompt-Guided Attention Alignment Module (PGAAM), and Dual Feature Compensation Module (DFCM). The CMPG integrates textual embeddings with visual features to produce degradation-aware prompts that encode degradation-related semantic and contextual cues. These prompts are injected into the decoder via a PGAAM, which adaptively aligns semantic information with degraded regions to facilitate context-aware restoration. To further enhance structural fidelity, DFCM is introduced that disentangles degradation artifacts from scene structures, thereby improving the reconstruction of fine textures and detailed content. By integrating cross-modal semantic guidance with spatial alignment and structural enhancement, DCMPC-Net achieves robust and perceptually consistent restoration across diverse weather conditions. Extensive experiments show that DCMPC-Net outperforms state-of-the-art methods in both task-specific and unified settings, achieving superior accuracy and visual fidelity.
Chinese Translation
恶劣天气导致多样且复杂的图像退化,严重影响计算机视觉系统的可靠性。现有的一体化恢复模型试图在统一框架内解决多种退化类型,但往往缺乏对退化特征的明确空间和语义建模,限制了其对多样天气条件的适应性。为了解决这一限制,我们提出了一种考虑退化的跨模态提示补偿网络(Degradation-Aware Cross-Modal Prompt Compensation Network,DCMPC-Net),该网络利用预训练的视觉-语言模型中的跨模态退化线索来调节统一主干中的恢复特征。具体而言,我们的DCMPC-Net主要由跨模态提示生成器(Cross-Modal Prompt Generator,CMPG)、提示引导注意力对齐模块(Prompt-Guided Attention Alignment Module,PGAAM)和双特征补偿模块(Dual Feature Compensation Module,DFCM)组成。CMPG将文本嵌入与视觉特征结合,生成编码与退化相关的语义和上下文线索的退化感知提示。这些提示通过PGAAM注入解码器,该模块自适应地将语义信息与退化区域对齐,以促进上下文感知的恢复。为了进一步增强结构保真度,引入了DFCM,该模块将退化伪影与场景结构分离,从而改善细腻纹理和详细内容的重建。通过将跨模态语义指导与空间对齐和结构增强相结合,DCMPC-Net在多样天气条件下实现了稳健且感知一致的恢复。大量实验表明,DCMPC-Net在任务特定和统一设置中均优于最先进的方法,达到了更高的准确性和视觉保真度。
cs.CV / 39 / 2608.06943
Dual-Space Modality Consistency Learning for Universal Cross-Modal Re-Identification
用于通用跨模态重识别的双空间模态一致性学习
Abstract
Cross-modal Re-Identification (ReID) aims to retrieve the same identity across heterogeneous imaging modalities and has been widely studied in visible-infrared person ReID and cross-modal ship ReID. Existing methods have achieved promising performance by learning modality consistency in the spatial embedding space, yet often overlook frequency-domain modality discrepancy, particularly in high-frequency representations that are both highly discriminative and modality-sensitive. In addition, most approaches are tailored to specific modality settings, limiting their applicability across diverse cross-modal scenarios. To address these challenges, we propose a Dual-Space Modality Consistency Learning (DSMCL) framework for universal cross-modal ReID. Specifically, DSMCL jointly models spatial feature distribution consistency and frequency-domain discriminative consistency. A Spatial Modality Consistency Learning (SMCL) branch performs Gaussian-based feature alignment, while a Frequency-aware Discriminative Consistency Learning (FDCL) strategy regularizes high-frequency representations through identity-aware cross-modal contrastive learning. By jointly capturing modality-specific characteristics and modality-shared identity cues, DSMCL learns robust representations and establishes a unified framework capable of accommodating diverse heterogeneous modality settings. Moreover, DSMCL is a plug-and-play framework that can be readily integrated into existing cross-modal ReID architectures. Extensive experiments on SYSU-MM01, RegDB, LLCM, HOSS-ReID, and CMShipReID across seventeen evaluation protocols show that DSMCL consistently improves multiple representative baselines.
Chinese Translation
跨模态重识别(ReID)旨在通过异构成像模态检索相同身份,并已在可见光-红外人重识别和跨模态船舶重识别中得到了广泛研究。现有方法通过学习空间嵌入空间中的模态一致性取得了良好的性能,但往往忽视了频域模态差异,特别是在高频表示中,这些表示既具有高度的区分性又对模态敏感。此外,大多数方法针对特定模态设置进行定制,限制了其在多样化跨模态场景中的适用性。为了解决这些挑战,我们提出了一种用于通用跨模态重识别的双空间模态一致性学习(DSMCL)框架。具体而言,DSMCL联合建模空间特征分布一致性和频域区分一致性。空间模态一致性学习(SMCL)分支执行基于高斯的特征对齐,而频率感知区分一致性学习(FDCL)策略通过身份感知的跨模态对比学习来规范高频表示。通过共同捕捉模态特定特征和模态共享身份线索,DSMCL学习到鲁棒的表示,并建立了一个能够适应多样异构模态设置的统一框架。此外,DSMCL是一个即插即用的框架,可以方便地集成到现有的跨模态重识别架构中。在SYSU-MM01、RegDB、LLCM、HOSS-ReID和CMShipReID上进行的广泛实验,涵盖了十七个评估协议,显示DSMCL在多个代表性基线模型上始终取得了提升。
cs.CV / 40 / 2608.06959
Summarize First, Download Later: Onboard VLMs for Bandwidth-Efficient Earth Observation
先总结,后下载:带宽高效的地球观测的机载视觉语言模型
Abstract
Modern Earth observation (EO) satellites carry increasingly advanced sensors that produce vast volumes of high-resolution, multispectral data, yet downlink capacity remains a critical bottleneck -- often causing significant latency or the loss of valuable observations within limited contact windows. We propose a "Summarize First, Download Later" paradigm that exploits recent advances in onboard edge computing and Vision-Language Models (VLMs). Rather than indiscriminately downlinking raw imagery, the system follows a three-phase interaction protocol: the satellite first transmits concise natural language summaries generated by a quantized onboard VLM; ground operators then issue targeted Visual Question Answering (VQA) queries to verify scene relevance (e.g., wildfires or maritime anomalies); and full-resolution images are downloaded only when critical information is confirmed. This transforms the downlink from passive bulk transfer into an active, semantics-aware dialogue. We implement and evaluate the system on a resource-constrained NVIDIA Jetson platform, and experiments on diverse remote sensing scenes show that the proposed strategy substantially reduces bandwidth consumption while accelerating time-to-insight for time-sensitive missions.
Chinese Translation
现代地球观测(EO)卫星搭载越来越先进的传感器,产生大量高分辨率、多光谱数据,但下行链路容量仍然是一个关键瓶颈——常常导致显著的延迟或在有限的接触窗口内丢失宝贵的观测数据。我们提出了一种“先总结,后下载”的范式,利用了机载边缘计算和视觉语言模型(VLMs)的最新进展。该系统并非无差别地下载原始图像,而是遵循三阶段交互协议:卫星首先传输由量化的机载VLM生成的简明自然语言摘要;然后,地面操作员发出针对性的视觉问答(VQA)查询,以验证场景的相关性(例如,野火或海洋异常);只有在确认关键信息后,才下载全分辨率图像。这将下行链路从被动的批量传输转变为主动的、语义感知的对话。我们在资源受限的NVIDIA Jetson平台上实现并评估了该系统,针对多样的遥感场景的实验表明,所提出的策略显著减少了带宽消耗,同时加快了对时间敏感任务的洞察时间。
cs.CV / 41 / 2608.06972
Generative Embedding Benchmark: How Much Information Survives in a Dense Embedding?
生成嵌入基准:在密集嵌入中多少信息得以保留?
Abstract
Embeddings have emerged as a standard representational interface linking foundation models with downstream systems. Most embedding benchmarks assess representations through discriminative tasks or geometric criteria centered on separability in embedding space. However, strong performance on such evaluations does not establish whether content compressed into an embedding remains accessible to a downstream generator. To address this gap, we introduce the Generative Embedding Benchmark (GEB), in which a decoder answers questions using only a frozen embedding and question text, without access to the original image or intermediate visual features. Answer quality under this readout measures generative information: the answer-relevant content recoverable from an embedding. GEB includes a curated visual-question-answering dataset with a 1,800-item development split and a held-out 900-item test split covering natural images, scene text, and visual documents. Using a common decoder and training recipe, we evaluate seven public embedding models in visual-only and vision-language joint modes. On the test set, visual-only scores range from 28.25 to 33.21; with image-question joint encoding, all five VLM-based embedding models score higher, and the best reaches 65.56. Matched embeddings also outperform text-only inputs, zero embeddings, and shuffled embeddings. Natural-image information is much easier to recover than scene text or visual-document information, while a Qwen3-VL-2B reference with access to the original image reaches 84.30. Together, these results show that generative readout exposes information bottlenecks that separability-based evaluation does not capture.
Chinese Translation
嵌入已成为连接基础模型与下游系统的标准表征接口。大多数嵌入基准通过以可分离性为中心的判别任务或几何标准来评估表征。然而,在这些评估中表现出色并不能确定压缩到嵌入中的内容是否仍然可以被下游生成器访问。为了解决这一问题,我们引入了生成嵌入基准(Generative Embedding Benchmark, GEB),在该基准中,解码器仅使用冻结的嵌入和问题文本来回答问题,而无法访问原始图像或中间视觉特征。在这种读取下,答案质量衡量生成信息:从嵌入中可恢复的与答案相关的内容。GEB包括一个精心策划的视觉问答数据集,具有1800个项目的开发集和900个项目的保留测试集,涵盖自然图像、场景文本和视觉文档。使用通用解码器和训练方案,我们在仅视觉模式和视觉-语言联合模式下评估七个公共嵌入模型。在测试集上,仅视觉的得分范围为28.25到33.21;在图像-问题联合编码下,所有五个基于视觉语言模型(VLM)的嵌入模型得分更高,最佳模型达到了65.56。匹配的嵌入也优于仅文本输入、零嵌入和打乱的嵌入。自然图像信息的恢复远比场景文本或视觉文档信息容易,而一个具有原始图像访问权限的Qwen3-VL-2B参考模型达到了84.30。综合来看,这些结果表明,生成读取揭示了基于可分离性评估未能捕捉的信息瓶颈。
cs.CV / 42 / 2608.06973
When One Modality Is Not Enough: Multimodal Sex and Life-Stage Classification of Red Deer from Aerial RGB-Thermal Video
当单一模态不足以满足需求时:基于航空RGB-热成像视频的红鹿多模态性别与生活阶段分类
Abstract
Aerial drone surveys increasingly support wildlife population estimation, yet a useful census is more than a count: population dynamics are defined by species composition, sex ratios and age structure, that is, by which species are present and how a herd splits into adult males, adult females and juveniles. We use red deer ($\textit{Cervus elaphus}$) as a test case, because managers act on these dynamics and because the visible cue defining adult males, the antlers, is seasonally variable. Surveys are flown nadir, high enough not to disturb the animals, so each deer occupies only a small, low-resolution patch. The two recording modalities fail in opposite conditions: in color a deer under canopy blends into the ground, while in thermal it becomes a bright blob that loses fine detail. Rather than trust either modality alone, we fuse them at every stage using self-supervised DINOv3 features. Our pipeline tracks animals in both modalities, treats an animal as confirmed only when the two cameras agree, keeps only the clear, non-occluded frames, and assigns species and sex by a vote across them; life stage is read separately from geo-referenced body size, since at survey resolution a juvenile often only differs from an adult female in size. Across four flights spanning the antler season the fused pipeline correctly classifies 25 of the 26 detected individuals (7 of 8 adult males, all 16 adult females and 2 juveniles), against 20 of 26 for either sensor alone. Multimodal species classification reaches 96.0%, while for sex classification fusing the two sensors matters most: the combined RGB+thermal model is the most robust across environments and seasons. Automating the demographic classification turns a drone flight from a count into a repeatable reading of herd structure, so the sex ratios and age structure that managers already act on can be gathered as often as a survey can be flown.
Chinese Translation
航空无人机调查越来越多地支持野生动物种群估计,但有效的人口普查不仅仅是计数:种群动态由物种组成、性别比例和年龄结构定义,即由哪些物种存在以及一个群体如何分为成年雄性、成年雌性和幼崽。我们以红鹿( extit{Cervus elaphus})作为测试案例,因为管理者需要根据这些动态采取行动,并且定义成年雄性的可见线索——鹿角在季节上是可变的。调查以垂直视角进行,飞行高度足以不打扰动物,因此每只鹿仅占据一个小的低分辨率区域。这两种记录模态在相反的条件下失效:在彩色图像中,树冠下的鹿与地面融为一体,而在热成像中,它则变成一个失去细节的亮斑。我们不单独信任任何一种模态,而是在每个阶段融合它们,使用自监督的DINOv3特征。我们的流程在两种模态中跟踪动物,只有当两台相机一致时才确认一只动物,保留清晰且未被遮挡的帧,并通过投票为其分配物种和性别;生活阶段则通过地理参考的体型单独读取,因为在调查分辨率下,幼崽往往仅在体型上与成年雌性不同。在覆盖鹿角季节的四次飞行中,融合流程正确分类了26只被检测个体中的25只(8只成年雄性中的7只、16只成年雌性中的全部和2只幼崽),而单独传感器的分类结果为26只中的20只。多模态物种分类达到了96.0%,而性别分类中融合这两种传感器最为重要:结合RGB+热成像的模型在不同环境和季节中最为稳健。自动化的人口分类将无人机飞行从简单的计数转变为对群体结构的可重复读取,因此管理者已经采取行动的性别比例和年龄结构可以在每次调查飞行时收集。
cs.CV / 43 / 2608.06981
Local Epistemic Uncertainty Guided Active Sampling for Plug-and-play Diffusive Image Restoration
基于局部认知不确定性引导的主动采样用于即插即用扩散图像恢复
Abstract
Diffusion models have demonstrated remarkable effectiveness in image restoration tasks. However, when guiding image reconstruction, existing Diffusion Model-based Image Restoration (DMIR) methods typically rely on fixed data constraints and uniform step sizes, thereby overlooking the dynamic nature of the generative process. Such rigid designs render the models vulnerable to spatially non-uniform degradations, thus resulting in structural distortions and loss of fine details. Meanwhile, uniform step sizes introduce computational redundancy, whereas na\"ive step reduction strategies tend to accumulate approximation errors. To address these limitations, we propose a Local Epistemic Uncertainty Guided Active Sampling framework (LEADer). In the spatial domain, LEADer leverages pixel-wise uncertainty to dynamically modulate the prior strength within the null space, which effectively balances detail preservation and artifact suppression. In the temporal domain, it quantifies sampling stability via the uncertainty trace to enable adaptive trajectory pruning, thereby accelerating convergence. Theoretical proofs demonstrate that our framework achieves strict data consistency, while the trajectory pruning strategy admits a deterministic error bound, thereby guaranteeing stable convergence under skip sampling. Notably, our plug-and-play method can be seamlessly integrated into various DMIR baselines. Extensive experiments show that LEADer improves the performance of multiple state-of-the-art DMIR methods, while significantly reducing sampling time with negligible memory overhead. Code is available at https://github.com/JiaqiZhang-Sengoku/LEADer.
Chinese Translation
扩散模型在图像恢复任务中表现出了显著的有效性。然而,在引导图像重建时,现有的基于扩散模型的图像恢复(DMIR)方法通常依赖于固定的数据约束和均匀的步长,从而忽视了生成过程的动态特性。这种僵化的设计使得模型易受空间非均匀退化的影响,从而导致结构失真和细节丢失。同时,均匀的步长引入了计算冗余,而简单的步长减少策略往往会积累近似误差。为了解决这些局限性,我们提出了一种基于局部认知不确定性引导的主动采样框架(LEADer)。在空间域中,LEADer利用像素级的不确定性动态调节零空间内的先验强度,有效平衡细节保留和伪影抑制。在时间域中,它通过不确定性轨迹量化采样稳定性,以实现自适应轨迹修剪,从而加速收敛。理论证明表明,我们的框架实现了严格的数据一致性,而轨迹修剪策略则允许确定性误差界,从而保证在跳跃采样下的稳定收敛。值得注意的是,我们的即插即用方法可以无缝集成到各种DMIR基线中。大量实验表明,LEADer提高了多种最先进的DMIR方法的性能,同时显著减少了采样时间且内存开销微乎其微。代码可在 https://github.com/JiaqiZhang-Sengoku/LEADer 获取。
cs.CV / 44 / 2608.07003
HRDiT: Training-Free High-Resolution Image Generation with Off-the-Shelf Diffusion Transformer Models
HRDiT:基于现成扩散变换器模型的无训练高分辨率图像生成
Abstract
Training-free text-to-high-resolution image generation has recently attracted growing research attention. However, existing studies on this task primarily focus on adapting off-the-shelf U-Net-based diffusion models to high resolutions, with limited progress on adapting off-the-shelf Diffusion Transformer (DiT) models despite their strong text-to-image generation capabilities at limited resolutions. In this work, we find two key challenges particularly hindering the application of off-the-shelf DiT models for high-resolution image synthesis in a training-free manner, namely, spatial disorder and long generation time. To address these challenges, we propose a novel method tailored to adapt off-the-shelf DiT models for high-resolution image synthesis. Extensive experiments show the efficacy of our method. Our code is available at: https://github.com/zylwithxy/HRDiT.
Chinese Translation
无训练的文本到高分辨率图像生成最近引起了越来越多的研究关注。然而,现有研究主要集中在将现成的基于U-Net的扩散模型适配到高分辨率上,而对现成的扩散变换器(Diffusion Transformer, DiT)模型的适配进展有限,尽管它们在有限分辨率下具有强大的文本到图像生成能力。在本研究中,我们发现有两个关键挑战特别阻碍了现成DiT模型在无训练方式下进行高分辨率图像合成的应用,即空间混乱和生成时间过长。为了解决这些挑战,我们提出了一种新方法,旨在将现成的DiT模型适配用于高分辨率图像合成。大量实验表明我们方法的有效性。我们的代码可在以下链接获取:https://github.com/zylwithxy/HRDiT。
cs.CV / 45 / 2608.07012
Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs
Scenix:通过可执行场景程序进行稀疏视图3D场景重建
Abstract
Synthesizing a structured and editable 3D indoor scene from a few uncalibrated RGB views requires more than generating high-quality individual assets: a system must infer the room structure, associate objects across incomplete observations, and recover a globally consistent spatial configuration. Previous methods mainly focus on 3D scene generation with text input or require continuous visual inputs with additional priors, \ e.g., human-annotated masks or accurate 3D layouts, which makes these methods labor demanding and hard to apply in general cases. We present \textsc{Scenix}, a sparse-view 3D scene reconstruction framework via executable scene programs, a structured representation that can be directly instantiated into editable 3D scenes. Given sparse views, \textsc{Scenix} predicts executable scene programs through perception-grounded asset instantiation and closed-loop spatial refinement. % We present \method, a framework that predicts an executable scene representation from sparse views and realizes it through perception-grounded asset instantiation and closed-loop spatial refinement. To support this task, we construct \dataset, a dataset of approximately 110,000 synthetic and real indoor scenes with multiview imagery, room structures, object-centric descriptions, and metric spatial annotations. We further introduce observation-consistent supervision that aligns each target scene with the visual evidence available in its input views. Experiments on held-out \textsc{XScene} scenes, real indoor images, and out-of-distribution SpatialGen cases evaluate structured scene prediction, object grounding, and spatial refinement.
Chinese Translation
从少量未校准的RGB视图合成一个结构化且可编辑的3D室内场景,不仅需要生成高质量的单个资产:系统还必须推断房间结构,关联不完整观测中的对象,并恢复全局一致的空间配置。以往的方法主要集中于基于文本输入的3D场景生成,或需要连续的视觉输入和额外的先验信息,例如人工标注的掩膜或准确的3D布局,这使得这些方法在一般情况下劳动密集且难以应用。我们提出了 extsc{Scenix},一种通过可执行场景程序进行稀疏视图3D场景重建的框架,这是一种可以直接实例化为可编辑3D场景的结构化表示。给定稀疏视图, extsc{Scenix}通过感知驱动的资产实例化和闭环空间细化来预测可执行场景程序。为了支持这一任务,我们构建了 extdataset,一个包含约110,000个合成和真实室内场景的数据集,涵盖多视角图像、房间结构、以对象为中心的描述和度量空间注释。我们进一步引入了观察一致性监督,将每个目标场景与其输入视图中的视觉证据对齐。在保留的 extsc{XScene}场景、真实室内图像和分布外SpatialGen案例上的实验评估了结构化场景预测、对象定位和空间细化。
cs.CV / 46 / 2608.07014
Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs
稳定曲线,不稳定项目:视频大语言模型中的项目级缩放异质性
Abstract
Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. We show that this view can conceal large, opposing changes at the item level. We represent each frozen model--item pair by its response trajectory under controlled visual budgets and derive matched-grid measures of configuration complementarity, harmful transitions, and text overwrite. Across five open Video LLMs from three architecture families, four multiple-choice benchmark splits, open-ended QA and summarization, and fixed-history dialogue generation, no single budget serves all items. On the four-model matched MCQA grid, item-level oracle headroom spans $8.8$--$18.9$ accuracy points and $12.5$--$25.5\%$ of items are correct at a lower budget but wrong at a higher one. Task-appropriate continuous metrics show the same complementarity beyond multiple choice: Token-F1 oracle gaps are $2.7$--$3.7$ score points on MLVU generation and $3.8$--$4.8$ points on AVSD current-turn generation, even when mean quality improves with budget. The effect persists across frame count, spatial resolution, sampling policy, temporal--spatial allocation, and independently executed raw-video and cached pipelines, with per-item rates and membership tracking protocol choices. A controlled sampling intervention recovers $29.0\%$ of terminal regressions, and a structured frame audit identifies several recurring evidence pathways. We release per-item trajectories, protocol provenance, derived annotations, and reproducible analysis code as an auditing artifact. A confidence cascade matches fixed-$128f$ accuracy while reducing average shared frame cost by $31.7\%$, illustrating one operational use of the response matrix.
Chinese Translation
聚合缩放曲线表明,视频大语言模型(Video LLMs)在视觉预算增加时平稳提升或趋于饱和。我们展示了这种观点可能掩盖了项目级别的巨大、相反的变化。我们通过在受控视觉预算下的响应轨迹来表示每个冻结模型与项目的配对,并推导出配置互补性、有害转变和文本覆盖的匹配网格度量。在来自三个架构家族的五个开放视频大语言模型、四个多项选择基准分割、开放式问答和摘要生成以及固定历史对话生成中,没有单一的预算适用于所有项目。在四模型匹配的多项选择问答(MCQA)网格中,项目级的oracle余量跨度为$8.8$--$18.9$准确率点,并且$12.5$--$25.5 ext{%}$的项目在较低预算下正确,而在较高预算下错误。适合任务的连续度量显示出超越多项选择的相同互补性:在MLVU生成中,Token-F1的oracle差距为$2.7$--$3.7$分,在AVSD当前轮次生成中为$3.8$--$4.8$分,即使在预算提高时平均质量也有所改善。该效应在帧数、空间分辨率、采样策略、时间-空间分配以及独立执行的原始视频和缓存管道中持续存在,伴随每个项目的速率和成员跟踪协议选择。一个受控的采样干预恢复了$29.0 ext{%}$的终端回归,而结构化的帧审计识别出几条重复的证据路径。我们发布了每个项目的轨迹、协议来源、衍生注释和可重复的分析代码作为审计文物。一个置信级联在固定-$128f$准确率下匹配,同时将平均共享帧成本降低了$31.7 ext{%}$,展示了响应矩阵的一个操作性用途。
cs.CV / 47 / 2608.07015
Understand Before Detect: Vision--Language Learning for Omni-Domain Infrared Small Target Detection
理解再检测:面向全领域红外小目标检测的视觉-语言学习
Abstract
Omni-domain infrared small target (IRST) detection is crucial for infrared surveillance, yet remains challenging due to heterogeneous imaging domains and inconsistent target characteristics. Previous deep learning-based methods have been developed for visual-only paradigms and achieved promising performance on domain-specific tasks. However, existing methods follow the task-specific supervised learning paradigm. This paradigm simplifies the full-scene infrared observations to sparse target supervision, discarding the semantics that remain invariant across heterogeneous domains. Consequently, detection performance suffers substantially under domain shifts. To handle this issue, we introduce \textbf{``understand before detect''}, a paradigm that formulates omni-domain IRST detection as an understanding-driven process, where holistic infrared target understanding precedes precise detection. Building on this paradigm, we propose \textbf{JinSight}, which first develops holistic IRST understanding through language supervision and then transfers the learned cross-domain representations to precise small-target detection. By grounding infrared representations in language semantics, JinSight enables a single model to generalize across heterogeneous infrared domains. We then introduce Latent Semantic Interaction (LSI), which exchanges language-aligned global semantics with fine-grained spatial features in a compact low-rank space. To address the lack of multimodal omni-domain IRST benchmarks, we build \textbf{OmniIRST-VL}, the first large-scale, highly diverse vision--language dataset for omni-domain IRST detection. It comprises over 39k annotations across six complementary instruction tasks covering both scene-level understanding and target-centric reasoning.
Chinese Translation
全领域红外小目标(IRST)检测对于红外监视至关重要,但由于成像领域的异质性和目标特征的不一致性,仍然面临挑战。以往基于深度学习的方法主要针对视觉单一范式,已在特定领域任务上取得了良好的表现。然而,现有方法遵循任务特定的监督学习范式,该范式将全场景红外观测简化为稀疏的目标监督,忽略了在异质领域中保持不变的语义。因此,在领域转移时,检测性能显著下降。为了解决这一问题,我们提出了“理解再检测”(``understand before detect'')的范式,将全领域IRST检测视为一个以理解驱动的过程,其中整体的红外目标理解优先于精确检测。在这一范式的基础上,我们提出了JinSight,该方法首先通过语言监督发展整体的IRST理解,然后将学习到的跨领域表征转移到精确的小目标检测中。通过将红外表征与语言语义相结合,JinSight使得单一模型能够在异质红外领域中进行泛化。接着,我们引入了潜在语义交互(Latent Semantic Interaction, LSI),在紧凑的低秩空间中交换与语言对齐的全局语义与细粒度空间特征。为了解决缺乏多模态全领域IRST基准的问题,我们构建了OmniIRST-VL,这是第一个大规模、高度多样化的视觉-语言数据集,用于全领域IRST检测。该数据集包含超过39,000个注释,涵盖六个互补的指令任务,涉及场景级理解和以目标为中心的推理。
cs.CV / 48 / 2608.07051
YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family
YOLO-PEFT:YOLO系列的参数高效微调
Abstract
Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail silently on real-time detectors, whose heterogeneous operators and detection-specific components impose placement constraints absent from regular Transformer stacks. We propose YOLO-PEFT, a structure-aware framework that formulates adapter placement as an auditable constraint-planning problem. Given a detector graph, a PEFT request, and a resource budget, YOLO-PEFT assigns operator and semantic roles, evaluates explicit operator-validity, detector-semantic, graph-interface, and deployment predicates, records a reason code for each excluded module, and either emits a budgeted target-module plan or returns Refuse before training. Under the official VOC07+12 trainval-to-VOC07 test protocol, planner-selected RS-LoRA reaches 0.7138 and 0.7307 mAP50-95 on YOLO11s and YOLO12s, respectively, compared with 0.6428 and 0.6662 for Full-SFT. On RT-DETR-L, all seven evaluated LoRA-family configurations cross the predefined catastrophic threshold, supporting a calibrated Refuse-to-Full-SFT decision within the evaluated coverage. A controlled YOLO11 audit further shows that LoRA reduces peak training memory by 43.9 percent, although training takes 1.72 times longer. Within the evaluated detector families, placement policies, and calibration coverage, YOLO-PEFT replaces manual target-module trial and error with explicit, inspectable planning while preserving verified train-save-merge-export paths; refusal on unseen detector architectures remains an open validation problem. Project Page: github.com/Tencent/YOLO-Master
Chinese Translation
从语言模型转移而来的通用参数高效微调(PEFT)方法在实时检测器上可能会静默失败,因为其异构操作符和特定于检测的组件施加了常规 Transformer 堆栈中不存在的放置约束。我们提出了 YOLO-PEFT,这是一种结构感知框架,将适配器放置形式化为可审计的约束规划问题。给定一个检测器图、一个 PEFT 请求和一个资源预算,YOLO-PEFT 分配操作符和语义角色,评估显式的操作符有效性、检测器语义、图接口和部署谓词,为每个被排除的模块记录原因代码,并在训练之前发出预算目标模块计划或返回拒绝。在官方的 VOC07+12 trainval 到 VOC07 测试协议下,规划器选择的 RS-LoRA 在 YOLO11s 和 YOLO12s 上分别达到了 0.7138 和 0.7307 的 mAP50-95,而 Full-SFT 的结果为 0.6428 和 0.6662。在 RT-DETR-L 上,评估的七种 LoRA 家族配置均超过了预定义的灾难性阈值,支持在评估覆盖范围内进行校准的拒绝到 Full-SFT 决策。受控的 YOLO11 审计进一步表明,LoRA 将峰值训练内存减少了 43.9%,尽管训练时间延长了 1.72 倍。在评估的检测器家族、放置策略和校准覆盖范围内,YOLO-PEFT 用显式、可检查的规划替代了手动目标模块的试错,同时保留了经过验证的训练-保存-合并-导出路径;对未见检测器架构的拒绝仍然是一个开放的验证问题。项目页面:github.com/Tencent/YOLO-Master
cs.CV / 49 / 2608.07057
KnifeHunter: Structured Local Representation Learning for Fine-Grained Knife Image Retrieval in Law Enforcement
KnifeHunter:用于执法领域细粒度刀具图像检索的结构化局部表示学习
Abstract
Knife-enabled violence presents a major public safety challenge, and law enforcement agencies require scalable tools for catalogue-level knife identification, intelligence analysis, and source attribution. Manual visual comparison is specialist, time-consuming, and difficult to scale under operational imaging conditions. We introduce KnifeHunter, an end-to-end forensic knife image retrieval system developed with UK law enforcement. The work contributes the KnifeHunter dataset, comprising 25,843 images across 543 knife classes from police evidence, retail catalogues, and border-force seizures, with structured metadata, Medium/Hard evaluation protocols, and large-scale distractor evaluation. We further propose CoRe-Net, a compact single-descriptor retrieval architecture that combines global context with spatially localised discriminative evidence. CoRe-Net introduces Structured Complementary Representation Learning (SCRL) to organise local evidence into complementary prototype-based representations, and Bi-Directional Reciprocal Fusion (BDRF) to integrate global and local evidence through residual projection and gated local-to-global injection. Using an EVA02-Base backbone and cosine-similarity retrieval, CoRe-Net achieves 88.0% mAP and 86.7% mP@10 on the Medium protocol, and 85.1% mAP and 83.8% mP@10 under distractor conditions. KnifeHunter was deployed by UK police forces during Operation Sceptre deployments from 2023 to 2025, achieving 99.2% mP@1 on field queries. These results demonstrate a practical and effective multimedia retrieval framework for fine-grained forensic knife matching in operational law-enforcement settings.
Chinese Translation
刀具相关暴力是一个重大的公共安全挑战,执法机构需要可扩展的工具来进行目录级刀具识别、情报分析和来源归属。手动视觉比较需要专业知识、耗时且在操作成像条件下难以扩展。我们介绍了KnifeHunter,这是一个与英国执法部门合作开发的端到端法医刀具图像检索系统。该研究贡献了KnifeHunter数据集,包含来自警方证据、零售目录和边境查获的543个刀具类别中的25,843张图像,配有结构化元数据、中等/困难评估协议和大规模干扰评估。我们进一步提出了CoRe-Net,这是一种紧凑的单描述符检索架构,结合了全局上下文与空间局部的判别证据。CoRe-Net引入了结构化互补表示学习(Structured Complementary Representation Learning, SCRL),将局部证据组织为互补的基于原型的表示,并通过残差投影和门控局部到全局注入集成全局和局部证据,提出了双向互补融合(Bi-Directional Reciprocal Fusion, BDRF)。使用EVA02-Base骨干网络和余弦相似度检索,CoRe-Net在中等协议下达到了88.0%的mAP和86.7%的mP@10,在干扰条件下达到了85.1%的mAP和83.8%的mP@10。KnifeHunter在2023至2025年的Sceptre行动中被英国警方部署,在现场查询中达到了99.2%的mP@1。这些结果展示了一个实用且有效的多媒体检索框架,用于在执法操作环境中进行细粒度法医刀具匹配。
cs.CV / 50 / 2608.07062
Explanation Stability of Test-Time Adaptation in Computational Pathology: A Large-Scale Benchmark
计算病理学中测试时适应的解释稳定性:大规模基准测试
Abstract
Test-time adaptation (TTA) has become a practical way to adapt deployed models to unlabeled target data, a setting that is especially relevant in computational pathology where staining, scanner, and cohort shifts are routine. While most TTA methods are evaluated by their effect on accuracy, clinical use also depends on whether the model's explanations remain reliable after adaptation. In this paper, we take a closer look at this largely unmeasured effect. We study explanation stability under TTA across two histopathology benchmarks, Camelyon17 and NCT CRC-HE, using five architectures ranging from convolutional networks to vision transformers and a pathology foundation model, seventeen TTA methods, and four attribution families. Across 2,958 adaptation runs, we observe a clear and systematic pattern: TTA methods differ sharply in how much they move model explanations, with frozen-backbone methods leaving attributions almost unchanged and continual methods such as CoTTA and RoTTA causing the largest drift. This effect is not uniform. Convolutional networks are substantially more sensitive than transformer and foundation-model backbones, and explanation drift increases with adaptation strength while remaining largely insensitive to batch size. Surprisingly, explanation stability is only weakly coupled to adaptation quality. Some methods preserve explanations almost perfectly while degrading calibration or accuracy, producing silent failures that would be missed by accuracy-only or explanation-only evaluation. These findings show that explanation stability is a distinct reliability axis for TTA in computational pathology. We release the metric, protocol, and full benchmark to support future work on adaptation methods that are not only accurate, but also stable and clinically auditable. Code: https://github.com/bahumanyarg11/tta-explanation-stability-pipeline
Chinese Translation
测试时适应(TTA)已成为将已部署模型适应于未标记目标数据的实用方法,这在计算病理学中尤为重要,因为染色、扫描仪和队列的变化是常见的。虽然大多数TTA方法的评估侧重于其对准确性的影响,但临床使用还依赖于模型在适应后解释的可靠性。在本文中,我们更深入地研究这一尚未充分测量的影响。我们在两个组织病理学基准(Camelyon17和NCT CRC-HE)下研究TTA的解释稳定性,使用五种架构,从卷积网络到视觉变换器和病理基础模型,十七种TTA方法,以及四种归因家族。在2,958次适应运行中,我们观察到一个明显且系统的模式:TTA方法在模型解释的变化程度上差异显著,冻结骨干的方法几乎不改变归因,而持续方法如CoTTA和RoTTA则导致最大漂移。该效应并不均匀。卷积网络对变化的敏感性显著高于变换器和基础模型骨干,而解释漂移随着适应强度的增加而增加,同时对批量大小的变化反应较小。令人惊讶的是,解释稳定性与适应质量的关联较弱。一些方法几乎完美地保留了解释,但却降低了校准或准确性,导致了仅通过准确性或解释评估无法发现的无声失败。这些发现表明,解释稳定性是计算病理学中TTA的一个独特可靠性维度。我们发布了该指标、协议和完整基准,以支持未来在不仅准确而且稳定且可临床审计的适应方法上的研究。代码:https://github.com/bahumanyarg11/tta-explanation-stability-pipeline
cs.CV / 51 / 2608.07088
RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs
RoRA:面向角色的区域分配用于多模态大语言模型中的视觉令牌剪枝
Abstract
Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Existing training-free pruning methods select tokens by importance, diversity, or spatial coverage, but treat retained tokens as interchangeable and do not explicitly track which object-related regions are already covered. We present RoRA, a training-free framework that casts visual token pruning as role-oriented regional evidence allocation. Given a fixed budget, RoRA partitions tokens into a protected semantic core, complementary context, and fine-grained detail. It first calibrates text-conditioned attention with a positional prior and a prompt-calibrated object prior, then builds Attention-Anchored Regions (AARs) from high-confidence anchors as lightweight proxies for covered object support. Context is explored mainly outside AARs, while a small AAR-guided budget restores local detail; pairwise similarity is used only for context-stage redundancy filtering. Under matched budgets, RoRA consistently outperforms strong training-free baselines across LLaVA and Qwen-VL families, retaining most of the unpruned accuracy even at aggressive pruning ratios, e.g., 96.5% of full performance at 88.9% pruning on LLaVA-1.5, and improving over D2Pruner by about 5% on Qwen3-VL at 75-90% pruning. At a 66.7% pruning ratio, RoRA requires only 0.7 ms for token selection and reduces end-to-end inference time by 24.6%, corresponding to a 1.33x speedup over unpruned inference on an NVIDIA H800.
Chinese Translation
多模态大语言模型(MLLMs)将图像编码为长视觉令牌序列,这使得预填充和KV缓存存储成本高昂。现有的无训练剪枝方法通过重要性、多样性或空间覆盖率选择令牌,但将保留的令牌视为可互换,并未明确跟踪哪些与对象相关的区域已被覆盖。我们提出了RoRA,一个无训练框架,将视觉令牌剪枝视为面向角色的区域证据分配。在固定预算下,RoRA将令牌划分为受保护的语义核心、互补上下文和细粒度细节。它首先通过位置先验和提示校准的对象先验来校准文本条件的注意力,然后从高置信度锚点构建注意力锚定区域(Attention-Anchored Regions, AARs),作为已覆盖对象支持的轻量级代理。上下文主要在AARs之外进行探索,而小规模的AAR引导预算则恢复局部细节;成对相似性仅用于上下文阶段的冗余过滤。在匹配预算下,RoRA在LLaVA和Qwen-VL系列中始终优于强大的无训练基线,即使在激进的剪枝比例下也能保留大部分未剪枝的准确性,例如,在LLaVA-1.5上以88.9%的剪枝率保留96.5%的完整性能,并在Qwen3-VL上以75-90%的剪枝率提高约5%相较于D2Pruner。在66.7%的剪枝比例下,RoRA仅需0.7毫秒进行令牌选择,并将端到端推理时间减少24.6%,相当于在NVIDIA H800上相较于未剪枝推理的1.33倍加速。
cs.CV / 52 / 2608.07092
International Transfer of Stochastic Cortical Self-Reconstruction
随机皮层自重建的国际转移
Abstract
Stochastic cortical self-reconstruction (SCSR) enables personalized mapping of gray matter atrophy, a hallmark of neurodegenerative disorders such as Alzheimer's disease (AD), onto high-resolution cortical surfaces. Unlike conventional normative modeling approaches, which typically operate at a coarse regional level and remain inherently constrained by the covariates included during training, SCSR estimates an individualized healthy reference directly from the observed cortical thickness at the vertex level. This allows the detection of subtle, subject-specific deviations from healthy cortical shape. In this work, we investigate the generalization and transferability of SCSR, originally trained on UK Biobank (UKB) data, to an independent Chinese population dataset. Specifically, we evaluate the ability of SCSR-derived Z-scores to discriminate between healthy scans, individuals with mild cognitive impairment (MCI), and patients with AD, while also assessing model robustness across the lifespan. We compare four training strategies: direct application of the UKB-trained model, fine-tuning on Chinese data, training from scratch, and joint training on UKB and Chinese cohorts. As reconstruction backbones, we consider both a multilayer perceptron (MLP) and a Spherical UNet (SUNet). Our results demonstrate that SCSR provides robust detection of cortical atrophy in the Chinese population across all evaluated models. The highest discriminative performance was achieved by the fine-tuned SUNet model (average pairwise AUC = 0.848), followed closely by the UKB-trained SUNet. Moreover, reconstruction errors remained low across the lifespan, even when the training population exhibited a substantially narrower age distribution, indicating strong cross-population transferability.
Chinese Translation
随机皮层自重建(SCSR)能够将灰质萎缩这一神经退行性疾病(如阿尔茨海默病(AD))的标志,个性化地映射到高分辨率的皮层表面。与通常在粗略区域层面操作并受到训练期间包含的协变量限制的传统规范建模方法不同,SCSR直接从观察到的顶点级皮层厚度中估计个体化的健康参考。这使得能够检测到与健康皮层形状的微妙、个体特异性偏差。在本研究中,我们探讨了SCSR的泛化能力和可转移性,该模型最初是在英国生物银行(UK Biobank,UKB)数据上训练的,现应用于一个独立的中国人群数据集。具体而言,我们评估了SCSR衍生的Z分数在区分健康扫描、轻度认知障碍(MCI)个体和阿尔茨海默病患者方面的能力,同时评估模型在整个生命周期中的稳健性。我们比较了四种训练策略:直接应用UKB训练的模型、在中国数据上进行微调、从头开始训练,以及在UKB和中国人群上联合训练。作为重建骨干网络,我们考虑了多层感知器(MLP)和球面UNet(SUNet)。我们的结果表明,SCSR在评估的所有模型中都能在中国人群中稳健地检测皮层萎缩。微调后的SUNet模型实现了最高的区分性能(平均成对AUC = 0.848),紧随其后的是UKB训练的SUNet。此外,即使训练人群的年龄分布明显较窄,重建误差在整个生命周期中仍保持较低,表明强大的跨人群可转移性。
cs.CV / 53 / 2608.07116
Geometry-Aware Camera Localization for Bronchoscopy
基于几何信息的支气管镜相机定位
Abstract
Camera localization in bronchoscopy remains a challenging problem due to stringent accuracy requirements, real-time constraints, and limited training data. Compared to natural scenes, the confined anatomical structures demand millimeter-level precision, while intraoperative guidance necessitates low-latency inference. However, existing methods often fail to effectively exploit preoperative geometric priors, limiting their robustness and accuracy. To address these limitations, we propose a unified geometry-aware bronchoscope localization framework (GABL) that effectively fuses preoperative structural priors with paired intraoperative video to estimate 6-DoF camera poses. Specifically, to address visual ambiguity in complex airways, we propose a graph-guided coarse-to-fine localization scheme that effectively leverages structural priors for precise pose estimation. Furthermore, to mitigate pose jitter and bridge the visual-structural gap, we integrate a Transformer-based tracking model with a novel RGB-depth matching objective, jointly enforcing spatio-temporal and geometric consistency. Extensive experiments demonstrate that our method yields remarkable reductions of 8.37% and 31.76% in translation and rotation errors over the prior state-of-the-art, alongside 4 times inference speedup (33.6 FPS) for robust real-time bronchoscope localization. Project website: https://paulili08.github.io/GABL/.
Chinese Translation
支气管镜中的相机定位仍然是一个具有挑战性的问题,主要由于严格的精度要求、实时性限制和有限的训练数据。与自然场景相比,狭窄的解剖结构需要毫米级的精度,而手术中的引导则要求低延迟的推断。然而,现有的方法往往未能有效利用术前几何先验,限制了其鲁棒性和准确性。为了解决这些局限性,我们提出了一种统一的基于几何信息的支气管镜定位框架(GABL),该框架有效地将术前结构先验与配对的术中视频融合,以估计6自由度(6-DoF)相机位姿。具体而言,为了解决复杂气道中的视觉模糊问题,我们提出了一种图引导的粗到细定位方案,能够有效利用结构先验进行精确的位姿估计。此外,为了减轻位姿抖动并弥合视觉与结构之间的差距,我们将基于Transformer的跟踪模型与一种新颖的RGB-深度匹配目标相结合,共同强制执行时空和几何一致性。大量实验表明,我们的方法在平移和旋转误差上分别比之前的最先进技术减少了8.37%和31.76%,同时实现了4倍的推断加速(33.6 FPS),用于鲁棒的实时支气管镜定位。项目网站:https://paulili08.github.io/GABL/
cs.CV / 54 / 2608.07117
Beyond Fluency: A Clinical Benchmark and Anomaly-Enhanced Baseline for Spine MRI Report Generation
超越流畅性:脊柱MRI报告生成的临床基准和异常增强基线
Abstract
Radiology reporting is time-consuming and subject to inter-rater variability, making automated report generation an attractive clinical application for Vision-Language Models (VLMs). We benchmark state-of-the-art VLMs on lumbar spine MRI with a focus on diagnostic accuracy and demonstrate that standard lexical and semantic metrics poorly reflect clinical correctness: fluent, well-structured reports can score highly while containing clinically meaningful diagnostic errors. To address this failure mode, we propose an architecture-agnostic framework that augments VLM inputs with spatially localized, disc-level anomaly heatmaps generated by a semi-supervised U-Net++ model. These heatmaps both improve anatomical sensitivity through explicit visual grounding and provide an independent interpretability output for clinical oversight, moving us closer to diagnostically reliable, visually grounded VLMs for lumbar spine MRI interpretation.
Chinese Translation
放射学报告的生成耗时且受评估者间变异的影响,这使得自动报告生成成为视觉-语言模型(Vision-Language Models, VLMs)在临床应用中的一个有吸引力的方向。我们对最先进的VLM在腰椎MRI上的表现进行了基准测试,重点关注诊断准确性,并证明标准的词汇和语义指标无法有效反映临床正确性:流畅、结构良好的报告可能得分很高,但却包含临床上有意义的诊断错误。为了解决这一失败模式,我们提出了一种架构无关的框架,通过半监督的U-Net++模型生成的空间局部化、椎间盘级别的异常热图来增强VLM的输入。这些热图不仅通过明确的视觉基础提高了解剖敏感性,还为临床监督提供了独立的可解释性输出,使我们更接近于在腰椎MRI解读中实现诊断可靠、视觉基础扎实的VLM。
cs.CV / 55 / 2608.07120
Multiple Hypothesis Flow Estimation for Video Frame Interpolation under Matching Ambiguity
匹配模糊下的视频帧插值的多假设光流估计
Abstract
Many flow-based video frame interpolation (VFI) methods synthesize an intermediate frame by estimating optical flow fields, warping the two input frames, and blending the warped observations. These latent flow fields are typically learned through image-level reconstruction supervision without direct flow annotations. In ambiguous regions containing repetitive or stochastic textures, rotating symmetric structures, or fast motion with blur, the matching evidence for a single query may contain multiple comparable and spatially separated peaks. Although the ground-truth intermediate frame provides indirect supervision, it may not uniquely identify the latent correspondence in ambiguous regions.When several locations provide multiple plausible matches, a single-flow estimator can retain only one displacement and discard the remaining candidates. If the selected match is incorrect or inconsistent with those of neighboring pixels, warping samples content from mismatched locations, producing ghosting, structural distortion, or blur.To address this limitation, we propose a multiple hypothesis flow estimation framework that preserves top-K candidate correspondences and selects one per location through a reliability-guided router. Each hypothesis is initialized from a coarse matching anchor and refined separately through anchor-centered local attention. Frame synthesis is thus conditioned on one selected flow-appearance hypothesis rather than a soft combination of candidate motions.Experiments on the proposed MA-HD benchmark and public VFI benchmarks show that our method achieves the best LPIPS and DISTS among the compared methods.
Chinese Translation
许多基于光流的视频帧插值(VFI)方法通过估计光流场、扭曲两个输入帧并融合扭曲后的观测结果来合成中间帧。这些潜在的光流场通常通过图像级重建监督学习,而没有直接的光流标注。在包含重复或随机纹理、旋转对称结构或快速运动模糊的模糊区域,单个查询的匹配证据可能包含多个可比且空间上分离的峰值。尽管真实的中间帧提供了间接监督,但它可能无法唯一地识别模糊区域中的潜在对应关系。当多个位置提供多个合理匹配时,单一光流估计器只能保留一个位移并丢弃其余候选项。如果选择的匹配不正确或与邻近像素的不一致,扭曲将从不匹配的位置采样内容,导致鬼影、结构失真或模糊。为了解决这一限制,我们提出了一种多假设光流估计框架,该框架保留前K个候选对应关系,并通过一个基于可靠性的路由器为每个位置选择一个。每个假设从粗匹配锚点初始化,并通过以锚点为中心的局部注意力单独进行细化。因此,帧合成是基于一个选择的光流-外观假设,而不是候选运动的软组合。在所提出的MA-HD基准和公共VFI基准上的实验表明,我们的方法在比较方法中实现了最佳的LPIPS和DISTS。
cs.CV / 56 / 2608.07141
Human-AI Perceptual Alignment by Playing Hues and Cues
通过玩色彩与线索实现人类与人工智能的感知对齐
Abstract
Evaluating the perceptual alignment between Contrastive Vision-Language Models (CVLMs) and humans is typically constrained by traditional benchmarks that overlook fine-grained semantic and cultural nuances. In this work, we propose a novel evaluation framework that leverages the gamified, discrete color space of the board game Hues and Cues. By mapping the board's 480 color cells to the CIE xy chromaticity diagram, we calculate empirical perceptual distances across a carefully curated 100-word vocabulary spanning seven semantic categories. To properly contextualize model performance, we establish an empirical lower bound of expected error-the Human Consistency baseline-calculated via Leave-One-Out (LOO) cross-validation on a dense dataset of color associations collected from 325 human observers through a custom digital interface. We evaluate 162 models across multiple architectural families and pre-training datasets to assess their semantic color grounding. Our results demonstrate that while CVLMs successfully replicate human cognitive biases, such as idealized memory colors for concrete physical referents (e.g., food and plants), they systematically diverge from the human baseline in abstract, subjective, and pop-culture domains. We identify two distinct failure modes in severely misaligned concepts: semantic misclassification and a systematic uncertainty collapse into a default blue coordinate. Furthermore, we reveal that highly curated pre-training datasets are significantly more effective than massive, uncurated corpora in mitigating these severe misalignments. Ultimately, this work highlights that despite their broad categorization capabilities, current CVLMs still fail to capture the nuanced, localized consensus of human color memory, emphasizing the value of gamified tasks in exposing underlying model biases. The data and code are publicly available to test other metrics.
Chinese Translation
评估对比视觉-语言模型(Contrastive Vision-Language Models, CVLMs)与人类之间的感知对齐通常受到传统基准的限制,这些基准忽视了细致的语义和文化细微差别。在本研究中,我们提出了一种新颖的评估框架,该框架利用了桌游《色彩与线索》(Hues and Cues)的游戏化离散色彩空间。通过将棋盘上的480个颜色单元映射到CIE xy色度图,我们计算了跨越七个语义类别的精心策划的100个词汇的经验感知距离。为了正确地对模型性能进行背景化,我们建立了预期误差的经验下限——人类一致性基线(Human Consistency baseline),该基线是通过对从325名人类观察者通过自定义数字接口收集的颜色联想的密集数据集进行留一交叉验证(Leave-One-Out, LOO)计算得出的。我们评估了162个模型,涵盖多个架构系列和预训练数据集,以评估它们的语义颜色基础。我们的结果表明,尽管CVLMs成功地复制了人类的认知偏见,例如对具体物理参照物(如食物和植物)的理想化记忆颜色,但它们在抽象、主观和流行文化领域与人类基线系统性地偏离。我们识别出两种在严重不对齐概念中的不同失败模式:语义错误分类和系统性不确定性崩溃为默认的蓝色坐标。此外,我们揭示出高度策划的预训练数据集在减轻这些严重不对齐方面显著优于大规模的未策划语料库。最终,本研究强调,尽管当前的CVLMs具有广泛的分类能力,但仍未能捕捉人类颜色记忆的细微、地方性共识,强调了游戏化任务在揭示潜在模型偏见中的价值。数据和代码已公开,以便测试其他指标。
cs.CV / 57 / 2608.07144
InstanceSplat: Instance-Aware Feed-Forward 3D Gaussian Splatting for Scene Understanding
InstanceSplat:面向实例的前馈式3D高斯点云渲染用于场景理解
Abstract
Feed-forward 3D Gaussian Splatting (3DGS) enables efficient and generalizable 3D reconstruction, but current feed-forward 3DGS methods for scene understanding remain largely category-oriented. In contrast, instance-aware 3DGS methods typically rely on per-scene optimization and often decouple reconstruction from instance and semantic learning, limiting reciprocal interactions among them. We present InstanceSplat, a unified feed-forward 3DGS framework for generalizable 3D reconstruction and instance-aware scene understanding from pose-free multi-view images. In a single forward pass, InstanceSplat constructs an instance-aware Gaussian representation that jointly encodes appearance, geometry, instance identity, and language-aligned semantics. Shared 3D Gaussians ground instance identities across views, producing renderable and cross-view-consistent instance features. To allow reconstruction and scene understanding to benefit from each other, we further design an instance-centric learning strategy that connects reconstruction, instance learning, and semantic learning through shared instance structure. Specifically, instance cues guide reconstruction, language-aligned semantics strengthen the discrimination of confusing same-category instances, and instance regions aggregate semantic evidence into coherent object-level predictions. Experiments on novel-view synthesis, instance segmentation, and open-vocabulary semantic understanding under varying input-view settings and on an unseen dataset demonstrate state-of-the-art performance, practical efficiency, and strong generalization.
Chinese Translation
前馈式3D高斯点云渲染(3DGS)实现了高效且可泛化的3D重建,但当前用于场景理解的前馈式3DGS方法仍然主要面向类别。相比之下,面向实例的3DGS方法通常依赖于每个场景的优化,并且常常将重建与实例和语义学习解耦,限制了它们之间的相互作用。我们提出了InstanceSplat,一个统一的前馈式3DGS框架,用于从无姿态的多视图图像中实现可泛化的3D重建和面向实例的场景理解。在一次前向传递中,InstanceSplat构建了一个面向实例的高斯表示,该表示共同编码了外观、几何、实例身份和语言对齐的语义。共享的3D高斯在视图之间确定实例身份,生成可渲染且视图间一致的实例特征。为了使重建和场景理解能够相互受益,我们进一步设计了一种以实例为中心的学习策略,通过共享实例结构连接重建、实例学习和语义学习。具体而言,实例线索引导重建,语言对齐的语义增强了对混淆同类实例的区分能力,而实例区域则将语义证据聚合为一致的对象级预测。在不同输入视图设置下以及在一个未见数据集上进行的新视图合成、实例分割和开放词汇语义理解的实验表明,InstanceSplat在性能、实用效率和强泛化能力方面达到了最先进的水平。
cs.CV / 58 / 2608.07176
Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation
基于表征驱动的内窥镜视觉嵌入对齐用于潜在生成
Abstract
Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost of training large Diffusion Transformers. Although representation alignment has improved efficiency in general computer vision, its role within the highly specialized endoscopic image space remains unclear. We introduce REVEAL (Representation-driven Endoscopic Visual Embedding Alignment), the largest generative foundation model for endoscopy to date, trained on GastroNet-5M (GN-5M), a multicenter dataset of 5 million endoscopic frames. Instead of depending on out-of-domain priors, REVEAL employs encoders pretrained directly on the endoscopic distribution to align diffusion latents with domain-specific visual features, preserving fine textures and intricate anatomical structures. Beyond image generation, REVEAL also serves as a powerful feature extractor; in multiple benchmarks, it delivers performance that is competitive with, and in several cases exceeds, endoscopic foundation models such as EndoViT and Endo-FM, specifically tuned for classification tasks, while demonstrating strong representation robustness under realistic imaging corruptions. REVEAL produces high-fidelity images and maintains robust structural coherence in latent-space edits such as inpainting and outpainting. This high-capacity backbone lowers the computational threshold for building specialized clinical tools, offering an open, versatile foundation for conditional synthesis, segmentation, and out-of-distribution detection in future intelligent gastroenterology systems.
Chinese Translation
为内窥镜开发基础生成模型受到自然图像与临床图像之间差距以及训练大型扩散变换器的计算成本的限制。尽管表征对齐在一般计算机视觉中提高了效率,但其在高度专业化的内窥镜图像领域中的作用仍不明确。我们提出了REVEAL(基于表征驱动的内窥镜视觉嵌入对齐),这是迄今为止最大的内窥镜生成基础模型,训练于GastroNet-5M(GN-5M),一个包含500万内窥镜帧的多中心数据集。REVEAL不依赖于域外先验,而是采用直接在内窥镜分布上预训练的编码器,将扩散潜变量与特定领域的视觉特征对齐,从而保留细腻的纹理和复杂的解剖结构。除了图像生成,REVEAL还作为一个强大的特征提取器;在多个基准测试中,其性能与专门针对分类任务调优的内窥镜基础模型如EndoViT和Endo-FM相竞争,并在某些情况下超越它们,同时在现实成像损坏下表现出强大的表征鲁棒性。REVEAL生成高保真图像,并在潜在空间编辑(如图像修复和图像扩展)中保持稳健的结构一致性。这个高容量的骨干网络降低了构建专业临床工具的计算门槛,为未来智能胃肠病学系统中的条件合成、分割和域外检测提供了一个开放且多功能的基础。
cs.CV / 59 / 2608.07199
Flow-Corrected Shape Optimization: Taming Manifold Drift in High-Dimensional 3D Models
流校正形状优化:驯服高维3D模型中的流形漂移
Abstract
Optimizing 3D shapes within the latent spaces of deep generative models is fundamental to computer assisted engineering, yet remains prone to a critical failure mode we term manifold drift: the tendency of gradient-based optimization to move latent vectors away from the manifold of valid shapes. This problem is exacerbated in state-of-the-art 3D shape generative models that operate in increasingly high-dimensional latent spaces where valid shapes occupy a vanishingly small fraction of the full space. Existing mitigation strategies, including latent regularization and flow-matching approaches, either sacrifice expressiveness, demand a difficult trade-off between objective guidance and generative fidelity that remains prone to manifold drift, or are computationally infeasible to scale to modern, large-capacity 3D shape models. We introduce a novel optimizer-corrector framework that alternates between gradient steps for objective minimization and guided flow matching to drive the latent state back to the valid shape manifold. By decoupling objective minimization from flow-based correction, optimizing freely and correcting strictly, this alternating design avoids inherent trade-offs, preserving geometric validity without sacrificing expressiveness while remaining computationally feasible on modern 3D shape models. We demonstrate its effectiveness across generative priors of varying complexity, from simple vector latent spaces to large-scale architectures across a variety of downstream optimization tasks, including aerodynamic drag reduction and object compliance optimization.
Chinese Translation
在深度生成模型的潜在空间中优化3D形状是计算机辅助工程的基础,但仍然容易出现我们称之为流形漂移的关键失败模式:基于梯度的优化倾向于将潜在向量移离有效形状的流形。这个问题在最新的3D形状生成模型中更加严重,这些模型在越来越高维的潜在空间中运行,而有效形状仅占据整个空间的微小部分。现有的缓解策略,包括潜在正则化和流匹配方法,要么牺牲表现力,要么在目标引导与生成保真度之间要求艰难的权衡,这种权衡仍然容易导致流形漂移,或者在现代大容量3D形状模型中计算上不可行。我们提出了一种新颖的优化器-校正器框架,该框架在目标最小化的梯度步骤和引导流匹配之间交替进行,以将潜在状态驱回有效形状流形。通过将目标最小化与基于流的校正解耦,自由优化并严格校正,这种交替设计避免了固有的权衡,保持几何有效性而不牺牲表现力,同时在现代3D形状模型上保持计算可行性。我们在不同复杂度的生成先验上展示了其有效性,从简单的向量潜在空间到大规模架构,涵盖了多种下游优化任务,包括空气动力学阻力减小和物体顺应性优化。
cs.CV / 60 / 2608.07256
CANIS: Generation-Assisted 3D Canonicalization via an Image-Semantic Bridge
CANIS:通过图像-语义桥接的生成辅助3D标准化
Abstract
Canonicalizing 3D object orientation is fundamental to 3D understanding and analysis. Existing approaches often rely on geometric cues, although 3D canonicalization ultimately requires a semantically meaningful orientation. To address this gap, we propose CANIS, a category-agnostic, generation-assisted framework that introduces the semantic orientation prior of a frozen image-to-3D generative model into 3D canonicalization, without canonicalization-specific training or category-specific templates. Specifically, CANIS first renders the input object from candidate viewpoints, selects an informative view, and generates a proxy in a canonical orientation. During generation, a sparse structural latent encoded from the input guides the proxy to preserve the geometry of an object. CANIS then uses the selected image as a semantic bridge between the input and the proxy. Image patches identify semantic regions on the proxy, and depth back-projection locates the corresponding regions on the input. The resulting semantic anchors constrain geometric matching, from which we estimate the rigid transformation that canonicalizes the input. Experiments on synthetic benchmarks validate CANIS and its key components, while qualitative results on partial observations and OmniObject3D suggest its applicability to incomplete and real-world scans. CANIS also improves downstream 3D classification, part segmentation, and dense correspondence under arbitrary rotations. Project page: https://kenkenzaii.github.io/Canis.
Chinese Translation
3D物体方向的标准化对于3D理解和分析至关重要。现有的方法通常依赖于几何线索,尽管3D标准化最终需要语义上有意义的方向。为了解决这一问题,我们提出了CANIS,这是一个类别无关的生成辅助框架,它将冻结的图像到3D生成模型的语义方向先验引入3D标准化,而无需特定于标准化的训练或特定于类别的模板。具体而言,CANIS首先从候选视点渲染输入物体,选择一个信息丰富的视图,并生成一个处于标准方向的代理。在生成过程中,从输入中编码的稀疏结构潜变量引导代理保持物体的几何形状。然后,CANIS使用所选图像作为输入与代理之间的语义桥接。图像块识别代理上的语义区域,深度反投影定位输入上的相应区域。由此产生的语义锚点约束几何匹配,从中我们估计出将输入标准化的刚性变换。在合成基准上的实验验证了CANIS及其关键组件,而在部分观测和OmniObject3D上的定性结果则表明其在不完整和真实世界扫描中的适用性。CANIS还改善了在任意旋转下的下游3D分类、部件分割和密集对应。项目页面:https://kenkenzaii.github.io/Canis。
cs.CV / 61 / 2608.07291
Foundation Models Adaptation for Multi-View Multi-modal Cardiac MRI Segmentation and Direct Ejection Fraction Estimation
基础模型在多视角多模态心脏MRI分割及直接射血分数估计中的适应性研究
Abstract
Foundation models have shown strong transferability in cardiac MRI (CMR), but their effectiveness for heterogeneous multi-view and multi-sequence CMR analysis remains unclear. In this work, we explore the effectiveness of fine-tuning and combining different CMR foundation models for the Universal Multi-Sequence, Multi-Center and Multi-View CMR Segmentation (CMR-Multi) Challenge. CineMA was fine-tuned for cine and late gadolinium enhancement (LGE) segmentation across short-axis and long-axis views. For direct left-ventricular ejection fraction (LVEF) estimation, we used two recent frozen CMR foundation models to extract embedding vectors that were then combined using attention-based multiple-instance learning for LVEF regression. In the challenge validation set, cine segmentation achieved Dice scores of 0.862, 0.883, and 0.902 for short-axis, two-chamber and four-chamber cine MRI, respectively. LGE segmentation achieved Dice scores between 0.621 and 0.846 across views. The direct LVEF regression model achieved an MAE of 4.96 percentage points and a Pearson correlation of 0.91. These results indicate that foundation models can be effectively adapted and combined for multi-view CMR analysis, while accurate LGE scar segmentation remains a challenging task.
Chinese Translation
基础模型在心脏MRI(CMR)中展现了强大的迁移能力,但其在异构多视角和多序列CMR分析中的有效性仍不明确。在本研究中,我们探讨了微调和结合不同CMR基础模型在通用多序列、多中心和多视角CMR分割(CMR-Multi)挑战中的有效性。CineMA模型经过微调,以实现短轴和长轴视图下的动态和晚期钆增强(LGE)分割。对于直接左心室射血分数(LVEF)估计,我们使用了两个最新的冻结CMR基础模型提取嵌入向量,然后通过基于注意力的多实例学习结合这些向量进行LVEF回归。在挑战验证集中,动态分割在短轴、双腔和四腔动态MRI中分别达到了0.862、0.883和0.902的Dice系数。LGE分割在各视图中的Dice系数介于0.621至0.846之间。直接LVEF回归模型的平均绝对误差(MAE)为4.96个百分点,Pearson相关系数为0.91。这些结果表明,基础模型可以有效地适应和结合用于多视角CMR分析,而准确的LGE瘢痕分割仍然是一项具有挑战性的任务。
cs.CV / 62 / 2608.07299
EliSeg: Verified Target Construction for Report-Grounded Abnormality Segmentation
EliSeg:基于报告的异常分割的验证目标构建
Abstract
Radiology reports describe clinical observations but do not specify executable segmentation targets. They may contain present, negated, prior,uncertain, or irrelevant findings, while multiple valid abnormalities may coexist. Existing segmentation methods largely bypass this ambiguity by receiving a target identity or spatial prompt before inference, which acts as a hidden target oracle. We study report-grounded abnormality segmentation, where a model must determine target eligibility, cardinality, and finding-to-mask correspondence directly from an unfiltered report before delineating the corresponding regions. We propose \textbf{EliSeg}, an atcor--verify--revise framework that integrates target construction with mask generation. A grammar-constrained Actor proposes target slots and masks, an independent text-only Verifier reconstructs the eligible finding inventory, and Revision selectively re-executes the shared Actor when their target structures disagree. EliSeg requires no predefined target identity, finding prompt, point, or bounding box. Experiments on MIMIC-CXR-ILS show that EliSeg consistently outperforms direct segmentation methods and extract-then-segment cascades across findings, while effectively suppressing masks for ineligible report mentions. Ablation studies confirm the complementary roles of verification and revision, and evaluation on CheXlocalize demonstrates effective transfer of the EliSeg to an external dataset.Code is available at https://github.com/Maybach-dream/EliSeg.
Chinese Translation
放射学报告描述了临床观察,但并未指定可执行的分割目标。报告中可能包含当前、否定、先前、不确定或无关的发现,同时多个有效的异常可能共存。现有的分割方法在推理之前通常通过接收目标身份或空间提示来规避这种模糊性,这相当于一个隐藏的目标oracle。我们研究基于报告的异常分割,其中模型必须直接从未经过滤的报告中确定目标的合格性、基数和发现与掩膜的对应关系,然后再划定相应的区域。我们提出了 extbf{EliSeg},一个整合目标构建与掩膜生成的atcor--verify--revise框架。一个受语法约束的Actor提出目标槽和掩膜,一个独立的仅文本Verifier重建合格发现清单,而Revision在目标结构不一致时选择性地重新执行共享的Actor。EliSeg不需要预定义的目标身份、发现提示、点或边界框。在MIMIC-CXR-ILS上的实验表明,EliSeg在各类发现上始终优于直接分割方法和提取后分割级联,同时有效抑制无资格报告提及的掩膜。消融研究确认了验证和修订的互补作用,而在CheXlocalize上的评估则展示了EliSeg向外部数据集的有效转移。代码可在 https://github.com/Maybach-dream/EliSeg 获取。
cs.CV / 63 / 2608.07302
Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination
相同的注意力,不同的真相:在视觉注意力上应用Logit-Lens以检测和减轻LVLM对象幻觉
Abstract
Large Vision-Language Models (LVLMs) often suffer from object hallucination, generating objects that are absent from the image. Prior work largely attributes this to insufficient visual attention. However, we find that both real and hallucinated objects receive equally strong visual attention in the model's mid-to-late layers, suggesting that the key issue may not be how much the model attends, but what it attends to and why. To this end, we decode the visual features of high-attention regions using Logit Lens, and observe that regions corresponding to real objects can be correctly decoded to the target object tokens, whereas those for hallucinated objects cannot. Building on this, we identify two hallucination mechanisms: (i) visual uncertainty, triggered by semantically similar or confusable regions; masking these regions eliminates the hallucination. (ii) contextual prior, triggered by strong co-occurrence priors; even when the initially attended region is masked, the hallucination persists and attention drifts to other regions. Based on these findings, we propose a simple yet effective training-free Detect-Mitigate framework comprising a Logit-Lens Consistency Check to detect hallucination and targeted remedies: High-Attention Regions Masking (HARM) for visual uncertainty hallucination, and Visual Evidence Enhanced Decoding (VEED) for contextual prior hallucination. Our approach achieves state-of-the-art results on multiple hallucination benchmarks. Code will be available.
Chinese Translation
大型视觉-语言模型(LVLMs)常常遭遇对象幻觉,生成图像中不存在的对象。之前的研究主要将此归因于视觉注意力不足。然而,我们发现模型中期到后期层次中,真实对象和幻觉对象都受到同样强烈的视觉注意力,这表明关键问题可能不是模型关注的程度,而是关注的内容及其原因。为此,我们使用Logit Lens解码高注意力区域的视觉特征,并观察到对应于真实对象的区域可以正确解码为目标对象标记,而幻觉对象的区域则无法做到这一点。在此基础上,我们识别出两种幻觉机制:(i)视觉不确定性,由语义相似或易混淆区域引发;屏蔽这些区域可以消除幻觉。(ii)上下文先验,由强共现先验引发;即使最初关注的区域被屏蔽,幻觉仍然存在,注意力会漂移到其他区域。基于这些发现,我们提出了一种简单而有效的无训练检测-减轻框架,包括Logit-Lens一致性检查以检测幻觉和针对性补救措施:高注意力区域屏蔽(HARM)用于视觉不确定性幻觉,以及视觉证据增强解码(VEED)用于上下文先验幻觉。我们的方法在多个幻觉基准测试中实现了最先进的结果。代码将会公开。
cs.CV / 64 / 2608.07340
H2AL: Hyperbolic Hierarchy-aware Aggregative Learning for Registration-based Few-shot Medical Image Segmentation
H2AL:基于超曲率层次感知的聚合学习用于注册基础的少样本医学图像分割
Abstract
Registration-based Few-shot medical image segmentation (RFMIS) aims to generate pseudo-labels for unlabeled images by warping a labeled image through registration. However, existing methods primarily perform pixel-level optimization and inference in Euclidean space, treating anatomical structures as flat and disjoint. This neglect of inherent hierarchies degrades pseudo-label quality and weakens the discrimination of ambiguous regions, limiting the segmentation performance. To overcome this challenge, we propose a Hyperbolic Hierarchy-aware Aggregative Learning framework for RFMIS, termed H2AL, that enhances both deformation plausibility and anatomical discrimination for dual-task learning. Specifically, we introduce a Hyperbolic Hierarchy-aware Infusion (H2I) module, which leverages the hierarchical modeling capability of hyperbolic space to learn precise hierarchy-aware representations via transformation-guided supervised hyperbolic contrastive learning, and injects such hierarchical priors into Euclidean space through a gated infusion block while preserving semantic richness. Furthermore, we propose an end-to-end joint optimization algorithm by gradient aggregation, where the gradients from the registration and segmentation decoders, embedding semantic and hierarchical cues, are aggregated to update the shared encoder to promote collaborative learning across tasks. Extensive experiments on two anatomical regions, with five experimental settings, demonstrate the effectiveness and efficiency of our method in both registration and segmentation. The code is publicly available at https://github.com/JiamingCai469/H2AL.
Chinese Translation
基于注册的少样本医学图像分割(RFMIS)旨在通过注册将标记图像变形,从而为未标记图像生成伪标签。然而,现有方法主要在欧几里得空间中进行像素级优化和推断,将解剖结构视为平面且不相交。这种对固有层次结构的忽视降低了伪标签的质量,并削弱了模糊区域的区分能力,从而限制了分割性能。为了解决这一挑战,我们提出了一种用于RFMIS的基于超曲率层次感知的聚合学习框架,称为H2AL,该框架增强了双任务学习中的变形合理性和解剖区分能力。具体而言,我们引入了一个超曲率层次感知注入(H2I)模块,该模块利用超曲率空间的层次建模能力,通过变换引导的监督超曲率对比学习来学习精确的层次感知表示,并通过门控注入块将这些层次先验注入欧几里得空间,同时保持语义丰富性。此外,我们提出了一种通过梯度聚合的端到端联合优化算法,其中来自注册和分割解码器的梯度,嵌入了语义和层次线索,被聚合以更新共享编码器,以促进任务间的协作学习。在两个解剖区域的五个实验设置上的大量实验表明,我们的方法在注册和分割方面的有效性和效率。代码已公开发布在 https://github.com/JiamingCai469/H2AL。
cs.CV / 65 / 2608.07382
SkySeaLand: A Wide-Format Satellite Transportation Benchmark with an Ultra-Lightweight Detection Baseline
SkySeaLand:一种超轻量级检测基线的宽格式卫星交通基准
Abstract
Satellite object detection is challenged by small targets and wide-format scenes that lose detail under standard square-input resizing. We introduce SkySeaLand, a public dataset of 1,307 high-resolution satellite images and 19,101 verified bounding boxes across airplane, boat, car, and ship classes in terrestrial and maritime scenes. Native COCO and YOLO annotations are provided. The collection is dominated by large source images and wide scene geometry: 84.5 percent exceed 3,836 pixels on the longest side and 73.1 percent are near a 3:1 aspect ratio. We evaluate twelve detectors from the YOLO, RT-DETR, DETR, and Faster R-CNN families using a common split and COCO metrics. The tested YOLO and RT-DETR variants obtain 84.4--88.2 mAP50, with no consistent accuracy gain from larger parameter counts under the reported model-specific recipes. We also report SkyDet, a 1.22 M parameter anchor-free baseline that obtains 60.5 mAP50 and 24.32 mAP50-95 in a 4.90 MB footprint, with 13.74 ms latency (72.8 FPS) on a Tesla T4. SkySeaLand provides a compact benchmark for mixed land--maritime transportation detection, while SkyDet establishes a documented low-footprint reference rather than a state-of-the-art accuracy claim.
Chinese Translation
卫星目标检测面临小目标和宽格式场景的挑战,这些场景在标准的正方形输入调整大小下会丢失细节。我们引入了SkySeaLand,一个公共数据集,其中包含1,307幅高分辨率卫星图像和19,101个经过验证的边界框,涵盖了陆地和海洋场景中的飞机、船只、汽车和船舶类别。提供了原生的COCO和YOLO注释。该数据集以大型源图像和宽场景几何为主:84.5%的图像最长边超过3,836像素,73.1%的图像接近3:1的宽高比。我们使用共同的划分和COCO指标评估了来自YOLO、RT-DETR、DETR和Faster R-CNN系列的十二个检测器。测试的YOLO和RT-DETR变体获得了84.4%至88.2%的mAP50,在报告的模型特定配方下,较大的参数数量并未带来一致的准确性提升。我们还报告了SkyDet,一个具有1.22M参数的无锚基线,在4.90MB的占用空间内获得60.5%的mAP50和24.32%的mAP50-95,在Tesla T4上具有13.74毫秒的延迟(72.8 FPS)。SkySeaLand为混合陆地-海洋交通检测提供了一个紧凑的基准,而SkyDet则建立了一个有文献记录的低占用参考,而非声称的最先进的准确性。
cs.CV / 66 / 2608.07405
GeoDistill-Refine: Silhouette-First Geometry Distillation for Annotation-Free Spacecraft Segmentation
GeoDistill-Refine:以轮廓为先的几何蒸馏用于无注释航天器分割
Abstract
Foundation segmentation models can provide supervision for spacecraft imagery without manual training masks, but their predictions vary with textual prompts and may contain geometric errors that are amplified during distillation. This paper presents GeoDistill-Refine, a two-stage framework that transfers offline SAM 3 pseudo-masks to a compact segmentation network. Six fixed prompts are fused by an unweighted 50% vote to stabilize the teacher output. The student first learns the foreground silhouette and is then refined with signed-distance-field, skeleton, and area objectives derived from the pseudo-mask. A sample-level gate, computed from prompt agreement, the valid-prompt ratio, and pseudo-mask area plausibility, reduces the influence of unreliable pseudo-geometry. On the SpaceSense-Bench HJM lockbox set, GeoDistill-Refine improves Image IoU and Boundary F1 by 0.0456 and 0.1380, respectively, over a plain pseudo-label student. External evaluations on the SPEED+ Lightbox and Sunlamp domains and on TANGO show competitive regional overlap together with gains in boundary quality or foreground precision. The deployed TinyUNet contains 0.263 M parameters and requires approximately 1.1 ms per image on an RTX 4090; SAM 3 pseudo-mask construction and the auxiliary geometry branches are used only during training.
Chinese Translation
基础分割模型可以在没有手动训练掩码的情况下为航天器图像提供监督,但它们的预测会因文本提示而异,并可能包含在蒸馏过程中被放大的几何错误。本文提出了GeoDistill-Refine,一个两阶段框架,将离线SAM 3伪掩码转移到一个紧凑的分割网络中。六个固定提示通过无权重的50%投票融合,以稳定教师输出。学生首先学习前景轮廓,然后通过从伪掩码派生的符号距离场、骨架和面积目标进行精炼。通过提示一致性、有效提示比例和伪掩码区域合理性计算的样本级门控,减少了不可靠伪几何的影响。在SpaceSense-Bench HJM锁盒集上,GeoDistill-Refine分别提高了图像IoU和边界F1指标0.0456和0.1380,相较于普通伪标签学生。在SPEED+ Lightbox和Sunlamp领域及TANGO上的外部评估显示出竞争性的区域重叠,同时在边界质量或前景精度上也有所提升。部署的TinyUNet包含0.263百万个参数,并在RTX 4090上每张图像大约需要1.1毫秒;SAM 3伪掩码构建和辅助几何分支仅在训练期间使用。
cs.CV / 67 / 2608.07408
Addressable Memory for Video World Models
可寻址的视觉世界模型内存
Abstract
We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames. However, we find that models can no longer reliably address stored content once rollouts extend beyond the training horizon, because temporal Rotary Positional Embeddings (RoPE) offsets then fall outside the range seen during training and the model struggles to retrieve the relevant visual information through attention. Moreover, naively compressing the cache in the RoPE-rotated space corrupts memory by averaging together incompatible positional phases. To address this, we propose WorldTrace, a training-free memory framework for long-horizon visual persistence. WorldTrace keeps compressed memory addressable by assigning each summary slot a distinct, in-distribution virtual position. Within this addressable cache, we study two memory compression approaches: WorldTrace-Field compresses history for temporal coherence, while WorldTrace-Landmark stores verbatim scene traces at detected transitions for episodic recall. We further introduce LoopBench, a benchmark evaluating whether a compressed cache can reconstruct a previously visited scene after a long detour. WorldTrace-Field improves temporal consistency by +15.5%, and WorldTrace-Landmark improves episodic recall by +19.5% on LoopBench, extending visually persistent generation without retraining.
Chinese Translation
我们研究了交互式视频世界模型中的视觉持久性。这些模型依赖于一个键值(Key-Value, KV)缓存,作为一个不断增长的视觉内存,用于保存之前生成的帧。然而,我们发现一旦回放超出训练范围,模型就无法可靠地寻址存储的内容,因为时间旋转位置嵌入(Rotary Positional Embeddings, RoPE)偏移量将超出训练期间观察到的范围,模型在通过注意力机制检索相关视觉信息时遇到困难。此外,简单地在RoPE旋转空间中压缩缓存会通过将不兼容的位置相位平均在一起而损坏内存。为了解决这个问题,我们提出了WorldTrace,一个无需训练的长时间视觉持久性内存框架。WorldTrace通过为每个摘要槽分配一个独特的、符合分布的虚拟位置,保持压缩内存的可寻址性。在这个可寻址的缓存中,我们研究了两种内存压缩方法:WorldTrace-Field压缩历史以保持时间一致性,而WorldTrace-Landmark在检测到的转折点存储逐帧场景轨迹以实现情节回忆。我们进一步引入了LoopBench,一个基准测试,用于评估压缩缓存是否能够在长时间绕行后重建先前访问的场景。WorldTrace-Field在LoopBench上提高了时间一致性15.5%,而WorldTrace-Landmark在情节回忆上提高了19.5%,在不重新训练的情况下扩展了视觉持久生成。
cs.CV / 68 / 2608.07409
UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling
UniJEPA:一种用于任务无关视觉世界建模的统一联合嵌入预测架构
Abstract
Joint-Embedding Predictive Architectures (JEPAs) have emerged as a principled framework for self-supervised learning of world models in compact latent spaces, yet existing methods are fragmented: some predict masked parts of a single image in latent space (I-JEPA), others learn to predict global photometric transformations (Image World Models), while video-scale JEPAs predict future temporal states and are post-trained for action-conditioned planning (V-JEPA~2, DINO-World, DINO-WM). These objectives are treated as distinct recipes with separate encoders, predictors, and anti-collapse regularizers, hindering a single model from unifying image-level and video-level world modeling. We present UniJEPA, a unified JEPA that jointly learns photometric prediction (image-level transformations) and temporal prediction (video-level next-state dynamics) in one shared latent space. A single end-to-end objective, composed of a next-embedding prediction loss and a Gaussian regularizer, yields a provably anti-collapse encoder-predictor pair trainable from raw pixels without EMA, stop-gradient, or pre-trained encoders. We show that the same latent space supports controllable abstraction: photometric prediction learns invariant structure while temporal prediction learns equivariant dynamics. After action-conditioned post-training on offline trajectories, UniJEPA enables zero-shot planning by treating goal features as prediction targets. On image, video, and control benchmarks, UniJEPA matches or surpasses task-specific JEPAs while requiring a single loss hyperparameter, and plans up to tens of times faster than generative world models at comparable accuracy.
Chinese Translation
联合嵌入预测架构(JEPAs)已成为在紧凑潜在空间中自监督学习世界模型的原则性框架,但现有方法却显得支离破碎:一些方法在潜在空间中预测单幅图像的遮挡部分(I-JEPA),另一些则学习预测全局光度变换(图像世界模型),而视频级JEPAs则预测未来的时间状态,并经过后训练以进行基于动作的规划(V-JEPA~2,DINO-World,DINO-WM)。这些目标被视为不同的配方,拥有独立的编码器、预测器和反崩溃正则化器,阻碍了单一模型统一图像级和视频级世界建模。我们提出了UniJEPA,一种统一的JEPA,它在一个共享的潜在空间中共同学习光度预测(图像级变换)和时间预测(视频级下一个状态动态)。一个端到端的单一目标,由下一个嵌入预测损失和高斯正则化器组成,产生了一个可证明的反崩溃编码器-预测器对,可以从原始像素中进行训练,而无需EMA、停止梯度或预训练编码器。我们展示了相同的潜在空间支持可控抽象:光度预测学习不变结构,而时间预测学习等变动态。在对离线轨迹进行基于动作的后训练后,UniJEPA通过将目标特征视为预测目标,实现了零-shot规划。在图像、视频和控制基准测试中,UniJEPA的表现与特定任务的JEPA相匹配或超越,同时只需一个损失超参数,并且在可比精度下的规划速度比生成世界模型快数十倍。
cs.CV / 69 / 2608.07417
I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning
我在视频中寻找你:基于身份的查询用于以人为中心的视频推理
Abstract
Real-world video reasoning often involves multimodal, multi-source inputs, whereas existing video reasoning tasks typically assume a simplified video-text setting, limiting identity matching and person-centric reasoning. To bridge this gap, we introduce the Identity-conditioned Queries (ICQ) task, in which models are required to jointly associate and interpret an input video and a reference image of a person, and leverage this conditioning to address identity grounding, behavior understanding, and temporal reasoning, among other challenges. Building on ICQ, we present ISYV (I Seek You in Videos), a systematic solution comprising three components: (1) ISYV-Bench, a challenging evaluation benchmark with 1,377 real-world complex videos and 1,377 question-answer pairs, organized into six difficulty levels spanning capabilities from identity recognition to causal reasoning; (2) ISYV-75K, a large-scale training set of 75K high-quality samples constructed via automated annotation, multi-stage verification, and manual review; and (3) ISYV-Framework, containing an ICQ-oriented model and training strategy for learning to exploit informative video shots without additional shot-level annotations. Extensive experiments show that both mainstream closed-source and open-source MLLMs struggle on ISYV-Bench, especially in cross-domain identity matching and long-horizon tracking. ISYV-Model outperforms strong baselines and in some aspects approaches closed-source performance. Overall, ISYV provides a unified task definition, scalable datasets/benchmarks, and modeling insights for person-centric video reasoning.
Chinese Translation
现实世界的视频推理通常涉及多模态、多来源的输入,而现有的视频推理任务通常假设简化的视频-文本设置,这限制了身份匹配和以人为中心的推理。为了解决这一问题,我们引入了基于身份的查询(Identity-conditioned Queries, ICQ)任务,在该任务中,模型需要共同关联和解释输入视频和一个人的参考图像,并利用这种条件来解决身份定位、行为理解和时间推理等挑战。在ICQ的基础上,我们提出了ISYV(I Seek You in Videos),这是一个系统性的解决方案,包含三个组成部分:(1)ISYV-Bench,一个具有挑战性的评估基准,包含1,377个现实世界的复杂视频和1,377对问答对,分为六个难度级别,涵盖从身份识别到因果推理的能力;(2)ISYV-75K,一个通过自动标注、多阶段验证和人工审核构建的75K高质量样本的大规模训练集;(3)ISYV-Framework,包含一个面向ICQ的模型和训练策略,用于学习在没有额外镜头级注释的情况下利用信息丰富的视频镜头。大量实验表明,无论是主流的闭源还是开源的多模态大语言模型(MLLMs)在ISYV-Bench上都面临困难,尤其是在跨领域身份匹配和长时间跟踪方面。ISYV-Model在多个方面超越了强基线,并在某些方面接近闭源性能。总体而言,ISYV提供了一个统一的任务定义、可扩展的数据集/基准和以人为中心的视频推理的建模见解。
cs.CV / 70 / 2608.07434
Conformal Coverage Guarantees for Any Video Temporal Grounder
任意视频时间基础的符合性覆盖保证
Abstract
Event boundaries in continuous video are ambiguous: re-annotate the same query-video pair and independent annotators mark moments that overlap by less than half on a large fraction of samples. The ground truth for video temporal grounding is therefore a distribution over intervals, yet every grounder returns a single interval with no statement of reliability, so at deployment a wrong interval is indistinguishable from a right one. COVER changes the output object: a post-hoc, model-agnostic wrapper that turns any grounder, a trained localizer or a black-box video--language model, into one that emits a temporal region containing the true moment with probability at least $1-\alpha$, by calibrating the quantile of a temporal nonconformity score on held-out labels and widening the base prediction by that amount. The guarantee is finite-sample and distribution-free under exchangeability, and requires neither retraining nor white-box access. We give two score families, a two-sided boundary-widening score for grounders that emit an interval and a super-level-set score for grounders that emit a relevance signal, and develop theory specific to grounding that bounds how large the certified region becomes, when coverage survives conditioning on event length, and how it degrades when moments from one video break exchangeability. Across three benchmarks and five grounders, realized coverage tracks the target, and calibration exposes what point metrics hide.
Chinese Translation
连续视频中的事件边界是模糊的:对同一查询-视频对进行重新标注时,独立的标注者在大量样本中标记的重叠时刻往往少于一半。因此,视频时间基础的真实情况实际上是一个区间的分布,而每个基础模型返回的却是一个单一的区间,并没有可靠性声明,因此在部署时,一个错误的区间与正确的区间无法区分。COVER改变了输出对象:一个后置的、与模型无关的包装器,将任何基础模型,无论是训练过的定位器还是黑箱视频-语言模型,转变为一个以至少 $1-eta$ 的概率发出包含真实时刻的时间区域的模型,通过对持出标签的时间非符合性分数的分位数进行校准,并将基础预测扩大相应的量。该保证在可交换性下是有限样本和无分布的,并且不需要重新训练或白盒访问。我们提供了两种分数家族,一种是针对发出区间的基础模型的双侧边界扩展分数,另一种是针对发出相关性信号的基础模型的超水平集分数,并发展了特定于基础的理论,限制了认证区域的大小,覆盖在条件事件长度下的生存情况,以及当来自一个视频的时刻打破可交换性时的退化情况。在三个基准和五个基础模型中,实现的覆盖跟踪目标,而校准则揭示了点度量所隐藏的内容。
cs.CV / 71 / 2608.07435
SABRE: Scalable and Automated Benchmarking of VLMs under Stress
SABRE:可扩展的自动化视觉语言模型压力测试基准
Abstract
Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisfy controlled conditions, remain answerable, and challenge current models. We present SABRE, a scalable, automated pipeline that converts a Test Primer (a Markdown Task Design with Data Schema) into structured specifications, generated or edited images, and question-answer pairs. Automated filtering removes candidates solved by a Filtering VLM, while human review verifies candidate validity and supports annotation correction and localized image repair. We instantiate SABRE-Prior to test whether VLMs follow visual evidence instead of relying on world priors -- learned expectations about familiar objects and scenes. Its 600 images and 1,000 questions span Context (unexpected entities in familiar scenes), Texture (counterfactual materials), Attribute (noncanonical component counts), and Language Elicitation (answers suggested by language but unsupported by the image). Across six VLMs, macro-average accuracy ranges from 17.8% to 31.3% (22.6% mean). A real-image Attribute control is comparably difficult for the Filtering VLM. SABRE-Counting and SABRE-Spatial pilots show that the workflow supports other stress-test settings. These results establish SABRE as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark.
Chinese Translation
视觉语言模型(VLMs)正在快速发展,但基准测试的开发滞后,使得识别其弱点变得困难。构建压力测试的成本很高:样本必须满足受控条件、保持可回答性,并对当前模型提出挑战。我们提出了SABRE,一个可扩展的自动化流程,将测试引导器(Test Primer,包含数据架构的Markdown任务设计)转换为结构化规范、生成或编辑的图像以及问答对。自动过滤机制去除被过滤VLM解决的候选项,而人工审核则验证候选项的有效性,并支持注释修正和局部图像修复。我们实例化了SABRE-Prior,以测试VLM是否遵循视觉证据,而不是依赖于世界先验——对熟悉物体和场景的学习期望。其600幅图像和1,000个问题涵盖了上下文(熟悉场景中的意外实体)、纹理(反事实材料)、属性(非典型组件计数)和语言引导(语言提示的答案但图像不支持)。在六个VLM中,宏平均准确率范围为17.8%至31.3%(平均22.6%)。真实图像的属性控制对过滤VLM同样具有挑战性。SABRE-Counting和SABRE-Spatial的试点显示,该工作流程支持其他压力测试设置。这些结果确立了SABRE作为一个可重用的框架,用于构建和更新VLM压力测试,而不是单一固定的基准。
cs.CV / 72 / 2608.07463
MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation
镜像世界:驯化视频扩散模型以生成镜面反射
Abstract
Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirror reflections remains challenging because the content within a mirror must remain consistent with the surrounding scene. Existing VDMs are not specifically designed to model scene-to-mirror relationships, which can lead to reflections with incorrect content or inconsistent spatial arrangements. We observe that mirror reflection generation involves two complementary challenges: determining what scene content should be reflected and how the reflected content should be spatially arranged within the mirror region. Motivated by this observation, we propose MirrorWorld, a reflection-aware video inpainting framework that models scene-to-mirror relationships during generation. Specifically, we introduce Semantic Relation Distillation (SRD), which transfers relational information from a frozen visual foundation model to encourage semantic associations between visible scene content and mirror regions. We further propose Geometric Transformation Alignment (GTA), which learns a transformation that guides the spatial arrangement of reflected content. The two components play complementary roles, with SRD modeling what should be reflected and GTA modeling how it should be arranged. To facilitate research on this problem, we construct a benchmark for video mirror reflection generation by repurposing four existing video mirror datasets into a unified reflection reconstruction task. Experimental results show that MirrorWorld achieves improved reflection reconstruction quality over representative image-based reflection generation methods and strong video inpainting baselines.
Chinese Translation
最近视频扩散模型(VDMs)的进展使得高保真视频合成成为可能。然而,生成镜面反射仍然具有挑战性,因为镜子中的内容必须与周围场景保持一致。现有的VDMs并未专门设计用于建模场景与镜子之间的关系,这可能导致反射内容不正确或空间排列不一致。我们观察到,镜面反射生成涉及两个互补的挑战:确定应该反射的场景内容,以及如何在镜子区域内空间排列反射的内容。基于这一观察,我们提出了MirrorWorld,一个在生成过程中建模场景与镜子关系的反射感知视频修复框架。具体而言,我们引入了语义关系蒸馏(Semantic Relation Distillation, SRD),它将关系信息从冻结的视觉基础模型转移,以促进可见场景内容与镜子区域之间的语义关联。我们进一步提出了几何变换对齐(Geometric Transformation Alignment, GTA),它学习一种变换,以指导反射内容的空间排列。这两个组件发挥互补作用,SRD建模应该反射的内容,而GTA建模如何排列这些内容。为了促进对这一问题的研究,我们通过将四个现有视频镜像数据集重新构建为统一的反射重建任务,构建了一个视频镜面反射生成基准。实验结果表明,MirrorWorld在反射重建质量上优于代表性的基于图像的反射生成方法和强大的视频修复基线。
cs.CV / 73 / 2608.07468
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
SimWAM:一种简单的世界动作模型用于端到端自主驾驶
Abstract
World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM that uses video generation purely as a training signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing the video branch to be discarded after training and leaving a self-contained planner that directly predicts trajectories. Since the two experts share no parameters and interact only through a unified attention interface, the video backbone could be replaced and the action expert scaled independently without modifying the learning objective or inference pipeline. We further apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Our SimWAM achieves $91.5$ PDMS on NAVSIM, surpasses state-of-the-art WAM-based planners with substantially lower latency, and transfers zero-shot to nuScenes. These results position SimWAM as a simple yet solid baseline that could readily benefit from advances in video generation for efficient autonomous driving. The code and model weights are available at https://github.com/H-EmbodVis/SimWAM/
Chinese Translation
世界动作模型(WAMs)通过将视频动态先验转移到动作预测中,从而改善端到端自主驾驶,但现有方法在推理时需要昂贵的未来生成。我们提出了SimWAM,这是一种简单而有效的WAM,仅将视频生成作为训练信号。它通过联合流匹配共同训练一个预训练的视频专家和一个轻量级的动作专家。一个孤立的注意力掩码使得动作预测独立于未来帧,从而在训练后可以丢弃视频分支,留下一个自包含的规划器,直接预测轨迹。由于这两个专家不共享参数,仅通过统一的注意力接口进行交互,因此视频主干可以被替换,动作专家可以独立扩展,而无需修改学习目标或推理管道。我们进一步应用强化学习来优化超越轨迹模仿的组合驾驶奖励。我们的SimWAM在NAVSIM上达到了91.5的PDMS,超越了基于WAM的最先进规划器,且延迟显著更低,并且在nuScenes上实现了零样本迁移。这些结果使SimWAM成为一个简单而稳固的基线,能够从视频生成的进展中受益,以实现高效的自主驾驶。代码和模型权重可在 https://github.com/H-EmbodVis/SimWAM/ 获取。
cs.AI / 1 / 2608.06394
Towards Multi-Label Graph Foundation Models: from Single-Vector Representation Learning to Multi-Semantic Basis Learning
迈向多标签图基础模型:从单向量表示学习到多语义基础学习
Abstract
Multi-label node classification is an important yet challenging task in graph learning, where nodes exhibit multiple semantics simultaneously. Existing methods for multi-label node classification can effectively model multiple labels, while only considering in-domain scenarios where the model needs to be trained and tested within the same graph domain, resulting in limited cross-domain generalization. Recently, Graph Foundation Models (GFMs) have emerged as a promising paradigm for learning transferable graph representations across diverse graph domains and downstream tasks. However, existing GFMs are built upon single-label assumption, where all nodes are arbitrarily regarded as containing only one class of semantic and embedded into a single representation. For multi-label nodes, such a representation essentially approximates multiple semantics with a single point in the representation space, inevitably leading to semantic entanglement and making simultaneous discrimination of multiple labels difficult. To address these limitations, we propose a Multi-Semantic Basis Graph Foundation Model (MSB-GFM), a framework for cross-domain multi-label node classification. Specifically, we introduce a multi-semantic basis representation learning paradigm that models each multi-label node as an adaptive composition of semantic bases, thereby enabling flexible representational capacity for modeling multiple semantics. Furthermore, we develop a semantic-structure dual-channel architecture with domain adversarial training for effective cross-domain knowledge transfer. Extensive experiments demonstrate the effectiveness of our model.
Chinese Translation
多标签节点分类是图学习中一项重要而具有挑战性的任务,其中节点同时展现多种语义。现有的多标签节点分类方法能够有效建模多个标签,但仅考虑在同一图域内进行训练和测试的场景,导致跨域泛化能力有限。近年来,图基础模型(Graph Foundation Models, GFMs)作为一种有前景的范式,已被提出用于学习可迁移的图表示,以适应多样化的图域和下游任务。然而,现有的GFM建立在单标签假设之上,将所有节点任意视为仅包含一个语义类别,并嵌入到单一表示中。对于多标签节点,这种表示本质上是用表示空间中的一个点来近似多个语义,必然导致语义纠缠,使得同时区分多个标签变得困难。为了解决这些局限性,我们提出了一种多语义基础图基础模型(Multi-Semantic Basis Graph Foundation Model, MSB-GFM),这是一个用于跨域多标签节点分类的框架。具体而言,我们引入了一种多语义基础表示学习范式,将每个多标签节点建模为语义基础的自适应组合,从而实现灵活的表示能力以建模多个语义。此外,我们开发了一种具有领域对抗训练的语义结构双通道架构,以有效实现跨域知识转移。大量实验表明了我们模型的有效性。
cs.AI / 2 / 2608.06398
EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs
EntropyMoE:面向无标记器大语言模型的熵感知稀疏专家路由
Abstract
Recent byte-level large language models (LLMs) have made tokenizer-free modeling increasingly competitive by grouping bytes into dynamically sized patches. However, existing byte-patch architectures still apply the same dense feed-forward computation to every patch. This uniform computation cannot adapt model capacity to variations in patch semantics and granularity. We address this limitation with EntropyMoE, a Mixture-of-Experts (MoE) architecture designed for dynamic byte patches. EntropyMoE replaces the dense feed-forward modules in the global patch Transformer with Top-K expert layers. Each dynamic patch serves as the basic unit of expert routing, and its byte coverage determines its contribution to workload accounting. The router selects experts directly from patch entropy, using the same granularity signal that underlies dynamic patch construction to organize sparse computation. Patch entropy and length jointly define the feature space for regulating expert specialization. Experiments show that EntropyMoE achieves the lowest held-out bits-per-byte among matched dense and sparse baselines while maintaining comparable downstream accuracy. These results establish patch entropy as an effective routing coordinate for sparse conditional computation and extend Mixture-of-Experts modeling beyond tokenizer-based representations.
Chinese Translation
最近的字节级大语言模型(LLMs)通过将字节分组为动态大小的补丁,使得无标记器建模变得越来越具有竞争力。然而,现有的字节补丁架构仍然对每个补丁应用相同的密集前馈计算。这种统一的计算无法根据补丁语义和粒度的变化来调整模型容量。我们通过EntropyMoE解决了这一限制,这是一种为动态字节补丁设计的专家混合(Mixture-of-Experts, MoE)架构。EntropyMoE用Top-K专家层替换了全局补丁Transformer中的密集前馈模块。每个动态补丁作为专家路由的基本单元,其字节覆盖范围决定了其在工作负载计算中的贡献。路由器直接从补丁熵中选择专家,利用与动态补丁构建相关的相同粒度信号来组织稀疏计算。补丁熵和长度共同定义了调节专家专业化的特征空间。实验表明,EntropyMoE在匹配的密集和稀疏基线中实现了最低的保留比特每字节,同时保持了可比的下游准确性。这些结果确立了补丁熵作为稀疏条件计算的有效路由坐标,并将专家混合建模扩展到超越基于标记器的表示。
cs.AI / 3 / 2608.06400
Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast
超越路由权重:通过贡献对比对混合专家奖励模型的忠实响应级解释
Abstract
Reward models are central to learning from human preferences, yet identifying what drives their predictions remains challenging. Recent sparse Mixture-of-Experts (MoE) reward models seek to improve interpretability by routing prompts to specialized experts and characterizing experts through examples with high routing weights. However, routing weights only reveal which prompts an expert $\textit{receives}$, not how it $\textit{judges}$ responses, providing only a partial account of expert behavior. We therefore propose $\textbf{Co}$ntribution-$\textbf{Co}$ntrast ($\textbf{CoCo}$) response-level interpretation, which faithfully characterizes experts' roles using chosen-rejected response pairs with the largest contribution contrasts, jointly capturing routing and preference behavior. Across automatic and human evaluations, CoCo yields more coherent, faithful, and specialized interpretations than router-based, score-based, and sparse autoencoder-based alternatives while maintaining competitive reward modeling accuracy. To the best of our knowledge, this is the first systematic study of interpretation methods for MoE reward models.
Chinese Translation
奖励模型在从人类偏好中学习中占据核心地位,但识别驱动其预测的因素仍然具有挑战性。最近的稀疏混合专家(Mixture-of-Experts, MoE)奖励模型试图通过将提示路由到专业专家并通过具有高路由权重的示例来表征专家,从而提高可解释性。然而,路由权重仅揭示了专家“接收”了哪些提示,而并未说明其如何“判断”响应,这仅提供了专家行为的部分解释。因此,我们提出了贡献对比(Contribution-Contrast, CoCo)响应级解释,该方法通过选择-拒绝响应对的最大贡献对比,忠实地表征专家的角色,联合捕捉路由和偏好行为。在自动和人工评估中,CoCo提供了比基于路由器、基于分数和基于稀疏自编码器的替代方法更连贯、忠实和专业的解释,同时保持了竞争力的奖励建模准确性。根据我们所知,这是对MoE奖励模型解释方法的首次系统研究。
cs.AI / 4 / 2608.06402
Interpretable Unsupervised Community Detection with LLM-Symbolized Structured Processes
可解释的无监督社区检测与LLM符号化结构过程
Abstract
Community detection is a fundamental task in graph analytics that aims to identify cohesive groups of entities with similar behaviors or interests. Classic objective-driven methods struggle with complex graph structures, while deep-learning approaches improve performance at the expense of interpretability and rely on labeled data and training. Large language models (LLMs), with strong reasoning capabilities and world knowledge, are promising for interpretable, label-free community detection. To leverage these strengths, we propose LUCID, an LLM-guided, interpretable, training-free, and unsupervised community detection method. Inspired by phase-transition kinetics in natural systems, where complex structures emerge through initialization, merging, refinement, and selection, LUCID is designed as a four-stage pipeline. Within this pipeline, the LLM induces formal rules that translate implicit knowledge into explicit and interpretable logical structures. Specifically, (1) the Local-View Community Initialization stage encodes local graph structures using k-ego contexts and unsupervised node roles; (2) the Multi-factor Community Merge stage uses LLM-induced rules to iteratively merge local communities; (3) the Multi-grain Community Refinement stage applies LLM-induced coarse-to-fine rules in parallel to reduce boundary noise; and (4) the Global-view Community Selection stage identifies high-quality communities based on topological compactness and boundary clarity. Extensive experiments on real-world datasets demonstrate that LUCID, as an unsupervised approach, achieves state-of-the-art performance and consistently outperforms leading unsupervised and semi-supervised baselines.
Chinese Translation
社区检测是图分析中的一项基础任务,旨在识别具有相似行为或兴趣的实体的凝聚性群体。经典的目标驱动方法在复杂的图结构中表现不佳,而深度学习方法虽然提升了性能,却牺牲了可解释性,并依赖于标注数据和训练。大型语言模型(LLMs)凭借其强大的推理能力和世界知识,为可解释的无标签社区检测提供了良好的前景。为了利用这些优势,我们提出了LUCID,这是一种由LLM指导的可解释、无训练的无监督社区检测方法。LUCID的设计灵感来源于自然系统中的相变动力学,在这些系统中,复杂结构通过初始化、合并、精炼和选择而出现,LUCID被设计为一个四阶段的流程。在这个流程中,LLM诱导出形式规则,将隐性知识转化为显性和可解释的逻辑结构。具体而言,(1) 本地视图社区初始化阶段使用k-ego上下文和无监督节点角色对局部图结构进行编码;(2) 多因素社区合并阶段利用LLM诱导的规则迭代合并局部社区;(3) 多粒度社区精炼阶段并行应用LLM诱导的粗到细规则以减少边界噪声;(4) 全局视图社区选择阶段基于拓扑紧凑性和边界清晰性识别高质量社区。在真实世界数据集上的广泛实验表明,作为一种无监督方法,LUCID实现了最先进的性能,并始终优于领先的无监督和半监督基线。
cs.AI / 5 / 2608.06410
ADIAS: Automated Design of Interactive Agentic Systems
ADIAS:交互式代理系统的自动化设计
Abstract
Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents, which leaves the repair progress implicit. This causes inefficient repair targeting, slow consolidation of partial progress, and propagation of ineffective interventions across rounds. Therefore, we formulate issue-centric agent optimization, in which repair progress is carried forward as an explicit persistent issue state to guide optimization, rather than re-derived from candidate history in each round. We instantiate the formulation in ADIAS, a framework for automated full-code agent design with two mechanisms. A persistent issue state maintains stable issue identities, lifecycle status, supporting evidence, and intervention-outcome histories. Issue-guided optimization uses this state to jointly propose repair targets and revision directions for subsequent focused full-code modification. Across five interactive benchmarks, ADIAS outperforms the strongest baseline by 25.2% on average and achieves consistent gains across four backbone models. Controlled ablations further show that removing persistent issue state or replacing issue-centric revision with candidate-centric policies leads to performance drops of up to 40.7%.
Chinese Translation
自动化代理设计通过迭代修订、评估和反馈总结来提升代理的利用效率。现有方法主要以候选代理为中心:跨轮次的经验围绕候选代理组织,这使得修复进展变得隐性。这导致修复目标不够高效、部分进展的整合缓慢,以及无效干预在轮次间的传播。因此,我们提出了以问题为中心的代理优化,其中修复进展作为明确的持续问题状态被传递,以指导优化,而不是在每一轮中从候选历史中重新推导。我们在ADIAS中实例化了这一公式,ADIAS是一个用于自动化全代码代理设计的框架,具有两种机制。持续问题状态维护稳定的问题身份、生命周期状态、支持证据和干预结果历史。问题引导的优化利用这一状态共同提出修复目标和后续集中全代码修改的修订方向。在五个交互基准测试中,ADIAS的表现平均比最强基线高出25.2%,并在四个基础模型上实现了一致的提升。控制消融实验进一步表明,去除持续问题状态或用候选中心的策略替代以问题为中心的修订会导致性能下降高达40.7%。
cs.AI / 6 / 2608.06411
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin
学习预测多模态大语言模型中层注意力以进行视觉令牌剪枝
Abstract
Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing numerous visual tokens. Visual token pruning can reduce this cost, but requires accurate token importance estimates. Recent studies have demonstrated that text-to-vision attention from middle language model layers can effectively guide visual token pruning, typically using attention from a predefined middle layer to select the visual tokens to retain. Two problems therefore remain. First, our analysis shows that the layer whose attention is most responsive to the question varies substantially across samples, making a fixed layer suboptimal. Second, obtaining attention from the appropriate middle layer requires processing numerous visual tokens through several language model layers, by which point considerable computation has already been spent. To address both problems, we propose Middle-layer Attention Prediction (MAP), which uses Question Contrastive Teacher Selection to identify a sample-specific teacher layer by contrasting attention under the original and reference questions, and distills attention from the selected layer into a lightweight predictor that estimates visual token importance from multi-modal input features. During inference, MAP combines the predicted importance scores with a diversity criterion to prune visual tokens before the first language model layer. Thus, MAP requires no attention maps for pruning and remains compatible with existing inference acceleration techniques. Across ten benchmarks on LLaVA-NeXT-7B, MAP retains 97.5% of the unpruned model performance with only 5.56% of the visual tokens, yielding a 3.09x end-to-end speedup.
Chinese Translation
多模态大语言模型(MLLMs)在各种视觉-语言任务中表现出色,但其效率受到处理大量视觉令牌成本的限制。视觉令牌剪枝可以降低这一成本,但需要准确的令牌重要性估计。最近的研究表明,中层语言模型层的文本-视觉注意力可以有效指导视觉令牌剪枝,通常使用预定义的中层注意力来选择保留的视觉令牌。因此,仍然存在两个问题。首先,我们的分析表明,对问题反应最敏感的层在不同样本之间差异很大,使得固定层的选择并不理想。其次,从适当的中层获取注意力需要通过多个语言模型层处理大量视觉令牌,这时已经消耗了相当多的计算资源。为了解决这两个问题,我们提出了中层注意力预测(Middle-layer Attention Prediction, MAP),该方法利用问题对比教师选择(Question Contrastive Teacher Selection)通过对比原始问题和参考问题下的注意力,识别样本特定的教师层,并将选定层的注意力提炼为一个轻量级预测器,以从多模态输入特征中估计视觉令牌的重要性。在推理过程中,MAP将预测的重要性分数与多样性标准结合,以在第一个语言模型层之前剪枝视觉令牌。因此,MAP在剪枝时不需要注意力图,并与现有的推理加速技术兼容。在LLaVA-NeXT-7B的十个基准测试中,MAP在仅使用5.56%的视觉令牌的情况下保留了97.5%的未剪枝模型性能,实现了3.09倍的端到端加速。
cs.AI / 7 / 2608.06474
WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader
WebGrader:使用自我演化程序评分器训练大语言模型进行网页开发
Abstract
Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approach to closing their remaining functional gap. This training regime is bottlenecked by reward design. Hand-authored browser scripts are executable yet costly to write for open-ended requirements, while VLM and GUI-agent graders scale but may issue verdicts before observing the decisive state. We propose WebGrader, a self-evolving programmatic grader that autonomously derives the required interaction flows from each website request, represents each flow as an executable Flow Contract, and uses its execution outcome as an RL reward. WebGrader materializes the generated project in a live browser, grounds target actions against the source code and live DOM, and collects visual, DOM, response, and persistent-state evidence along the same browser trajectory. A residual-driven offline loop then discovers reusable verifier skills, screens them on disjoint validation pages, and freezes the promoted skill graph before policy training. By separating test planning, action grounding, evidence collection, and semantic judgment, WebGrader issues a Pass verdict only after observing the requested transition. On WebGen-Bench, WebGrader trains an 8B policy to a 52.01% functional success rate, outperforming a matched appearance-plus-script reward by 7.88 points and surpassing o4-mini and DeepSeek-v4-flash. On WG-core-250, the policy reaches a Full Score of 44.953 and surpasses Qwen3-Coder-480B.
Chinese Translation
大型语言模型越来越能够根据自然语言描述生成完整的网站,而强化学习已成为缩小其剩余功能差距的核心方法。这种训练机制受限于奖励设计。手工编写的浏览器脚本虽然可执行,但在开放式需求下编写成本高昂,而视觉语言模型(VLM)和图形用户界面(GUI)代理评分器虽然可扩展,但可能在观察到决定性状态之前就发出判决。我们提出了WebGrader,这是一种自我演化的程序评分器,能够自主推导每个网站请求所需的交互流程,将每个流程表示为可执行的流程合约,并将其执行结果作为强化学习奖励。WebGrader在实时浏览器中实现生成的项目,将目标动作与源代码和实时文档对象模型(DOM)进行比对,并沿着同一浏览器轨迹收集视觉、DOM、响应和持久状态证据。然后,基于剩余驱动的离线循环发现可重用的验证技能,在不相交的验证页面上进行筛选,并在策略训练之前冻结提升的技能图。通过将测试规划、动作定位、证据收集和语义判断分离,WebGrader仅在观察到请求的状态转移后才发出通过判决。在WebGen-Bench上,WebGrader训练出一个8B的策略,达到了52.01%的功能成功率,超越了匹配的外观加脚本奖励7.88分,并超过了o4-mini和DeepSeek-v4-flash。在WG-core-250上,该策略达到了44.953的满分,超越了Qwen3-Coder-480B。
cs.AI / 8 / 2608.06501
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
多模态大语言模型能解码创造性飞跃吗?引入C4用于跨概念理解
Abstract
Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented tasks. Cross-concept understanding is a core cognitive capacity underlying receptive creativity. It enables a perceiver to recover intended meaning from non-obvious but meaningful conceptual relations. We operationalize item construction as cross-concept encoding and model inference as cross-concept decoding. We introduce C4, a cognition-inspired evaluation framework for Chengyu (Chinese idiom)-based Cross-Concept Creativity. Its encoding component maps target slots to imageable substitute concepts along bridge paths in a manually annotated and third-party-reviewed cross-concept network, enabling batch generation with explicit structure, difficulty indexed by bridge count and depth, and exact answers. Using this framework, we instantiate the C4 Evaluation Set (C4-Eval), comprising 184 synthetic items and 37 human-created cross-concept chengyu figures collected from online sources. We manually construct and review cross-concept relations, bridge paths, and reasoning processes for the collected figures. Each C4-Eval item is instantiated in five task settings, yielding 884 primary answer-recovery cases. Across ten evaluated MLLMs, the strongest closed models reach 50.7% and 48.0% primary accuracy, while open-source models remain substantially lower. Candidate constraints improve accuracy sharply, but bridge hints and explanation requests provide only modest gains. These results expose a substantial gap in how current MLLMs decode creatively encoded meaning through cross-concept relations. The code is in the supplementary material.
Chinese Translation
多模态大语言模型(MLLMs)的创造能力在设计、沟通、教育和人机协作中至关重要,但由于与以准确性为导向的任务相比,明确的目标和奖励信号稀缺,评估这些能力仍然困难。跨概念理解是支撑接受性创造力的核心认知能力。它使感知者能够从非显而易见但有意义的概念关系中恢复预期的意义。我们将项目构建操作化为跨概念编码,将模型推理操作化为跨概念解码。我们引入C4,一个基于认知启发的评估框架,用于基于成语(Chengyu)的跨概念创造力。其编码组件将目标槽映射到沿手动注释和第三方审核的跨概念网络中的可成像替代概念的桥接路径,使得能够以明确的结构进行批量生成,难度通过桥接数量和深度进行索引,并提供确切的答案。利用该框架,我们实例化了C4评估集(C4-Eval),该评估集包含184个合成项目和37个从在线来源收集的人类创作的跨概念成语图形。我们手动构建和审核收集的图形的跨概念关系、桥接路径和推理过程。每个C4-Eval项目在五个任务设置中实例化,产生884个主要答案恢复案例。在评估的十个MLLM中,最强的封闭模型达到50.7%和48.0%的主要准确率,而开源模型的准确率则显著较低。候选约束显著提高了准确性,但桥接提示和解释请求仅提供了适度的增益。这些结果揭示了当前MLLM如何通过跨概念关系解码创造性编码意义之间存在显著差距。代码在补充材料中。
cs.AI / 9 / 2608.06530
KNOWPLAN: Knowledge-Driven AI Agents for Smart Degree Pathway Planning
KNOWPLAN:基于知识驱动的智能学位路径规划AI代理
Abstract
Planning a degree from official university sources requires solving two problems in order. The institution's curriculum must first be reconstructed from catalogs, departmental pages, JSON endpoints, and PDFs that share no schema, and only then can a student-specific path be optimized under prerequisite logic and overlapping requirement constraints. Coupling the two lets each failure mode hide the other, because a planner that drives its own crawling never learns facts its current plan does not need. We present KnowPlan, which enforces an extraction-first boundary and measures the interface between the stages rather than assuming it. CatalogBrowse explores with no access to any user profile. It scores legal actions by lower-confidence expected marginal gain over a finite set of atomic catalog obligations per unit of source access, parses deterministically through platform adapters with a span-constrained clause-to-AST model fallback, and terminates on a closure certificate over index, schema, provenance, and reference completeness instead of a reward threshold. Its output contract is three provenance-linked JSON documents. DegreeMap consumes only those documents. It compiles them into a typed requirement hypergraph and optimizes lexicographically with CP-SAT over hard feasibility, completion horizon, load and risk, personalized utility, and option value, so that each stage optimizes inside the previous stage's proven optimum and stays certifiable within the solver budget. Across a 100-university broad track and a six-school dense track, CatalogBrowse reaches 96.2% inventory recall and 88.7% masked-source recovery at 47% less source access than an exhaustive crawler, DegreeMap holds 100.0% hard feasibility while improving personalized utility by +0.066 over the strongest baseline, and the full pipeline certifies 99.5% of requests with a utility gap to the privileged gold graph of 0.015.
Chinese Translation
从官方大学资源规划学位需要依次解决两个问题。首先,必须从没有共享模式的目录、部门页面、JSON端点和PDF中重建机构的课程设置,只有在此基础上,才能在先决条件逻辑和重叠要求约束下优化特定学生的路径。这两者的结合使得每种失败模式可以掩盖另一种,因为一个自主爬取的规划者从不学习当前计划不需要的事实。我们提出了KnowPlan,它强制执行提取优先的边界,并测量阶段之间的接口,而不是假设它。CatalogBrowse在没有任何用户档案访问的情况下进行探索。它通过每单位源访问的有限原子目录义务的低置信度预期边际收益来评分合法操作,并通过平台适配器以确定性方式解析,采用跨度受限的子句到AST模型回退,并在索引、模式、来源和参考完整性上终止于闭包证书,而不是奖励阈值。其输出合同是三个与来源相关的JSON文档。DegreeMap仅使用这些文档。它将其编译成一个类型化的需求超图,并在硬可行性、完成时间、负载和风险、个性化效用以及选项价值上通过CP-SAT进行字典顺序优化,从而使每个阶段在前一个阶段的证明最优解内进行优化,并保持在求解器预算内可证明。在一个涵盖100所大学的广泛轨道和一个涵盖六所学校的密集轨道中,CatalogBrowse以比全面爬虫少47%的源访问达到了96.2%的库存召回率和88.7%的掩蔽源恢复率,DegreeMap在提高个性化效用+0.066的同时保持了100.0%的硬可行性,而整个管道以0.015的效用差距认证了99.5%的请求,优于特权黄金图。
cs.AI / 10 / 2608.06544
TaskSense: Focusing on What Matters in World Models
TaskSense:聚焦于世界模型中的重要内容
Abstract
World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representations to preserve information across the entire visual input. However, task-relevant content often occupies only a small fraction of the observation, while background clutter and distractors consume valuable representational capacity. This mismatch between visual reconstruction and control objectives biases latent representations to model task-irrelevant visual content, diluting learning signals for control-relevant features and severely degrading downstream performance under visual distractions. We introduce TaskSense, a task-centric world modeling framework that enforces task relevance before latent encoding through a differentiable stochastic spatial attention mechanism conditioned on the previous latent state. To steer attention toward control-relevant regions, we augment training with an auxiliary inverse-dynamics objective. Rather than reconstructing the full observation, the world model reconstructs only the attended regions, encouraging latent representations to preserve task-relevant information while discarding irrelevant visual content. The decoder is further conditioned on the sampled attention map, enabling consistent reconstruction despite stochastic attention. Compared with the DreamerV3 baseline, TaskSense maintains competitive performance on the DeepMind Control Suite while consistently outperforming DreamerV3 on the Distracting Control Suite, demonstrating substantially improved robustness to visual distractions. Qualitative analysis further confirms that the learned attention, guided by inverse-dynamics supervision, consistently localizes control-relevant regions while suppressing irrelevant visual content.
Chinese Translation
用于视觉控制的世界模型通常通过重构观测来学习紧凑的潜在状态,隐式地鼓励表示在整个视觉输入中保留信息。然而,任务相关内容通常只占观测的一小部分,而背景杂乱和干扰物则消耗了宝贵的表示能力。这种视觉重构与控制目标之间的不匹配,使得潜在表示倾向于建模与任务无关的视觉内容,从而稀释了控制相关特征的学习信号,并在视觉干扰下严重降低了下游性能。我们提出了TaskSense,一个以任务为中心的世界建模框架,通过一个可微分的随机空间注意机制,在潜在编码之前强制任务相关性,该机制以先前的潜在状态为条件。为了引导注意力朝向控制相关区域,我们通过辅助的逆动力学目标增强训练。世界模型不仅重构完整的观测,而是仅重构被关注的区域,鼓励潜在表示保留任务相关信息,同时丢弃无关的视觉内容。解码器进一步以采样的注意力图为条件,使得尽管存在随机注意力,仍能实现一致的重构。与DreamerV3基线相比,TaskSense在DeepMind控制套件上保持了竞争性能,同时在干扰控制套件上持续超越DreamerV3,显示出对视觉干扰的显著增强的鲁棒性。定性分析进一步确认,受逆动力学监督引导的学习注意力,始终能够准确定位控制相关区域,同时抑制无关的视觉内容。
cs.AI / 11 / 2608.06578
Divergent Response Modes in Frontier Language Models Under Steering Pressure
前沿语言模型在引导压力下的不同响应模式
Abstract
Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behaviors under explicit steering pressure remains underexplored. This study evaluates behavioral steerability across six frontier models from six developers using 300 paired base and steered items over three categories: values-conflict, reasoning-elicitation, and reasoning-suppression (plus 40 validation items). All six models act as blind peer judges and classify every response based on fixed behavioral rubrics. The resulting 24,480 judgments are scored by leave-one-out consensus. We find that models differ not just in how much steering shifts their behavior but in what kind (mode) of response they give, and some response modes appear in only one or two of them. GPT-5 deflects requests to disclose its reasoning while leaving its answer intact (99% vs. 0% for all other models). Claude Opus 4.7 and GPT-5 resist explicit suppression instructions and in different ways. Using Llama as the open-weight model, we trace the largest behavioral split to its internals. A linear probe decodes the behavior from the residual stream at 0.87 held-out accuracy while injecting that direction during generation drives the behavior from 0% to 86% across an intervention sweep. Every finding holds under both a token-budget remediation and a control experiment with a hypothesis-blind judgment prompt.
Chinese Translation
前沿语言模型使用不同的数据、目标和安全管道进行训练。这些差异是否在明确的引导压力下产生可测量的不同行为尚未得到充分探索。本研究评估了来自六个开发者的六个前沿模型在300对基础和引导项目中的行为可引导性,涵盖三个类别:价值冲突、推理引导和推理抑制(加上40个验证项目)。所有六个模型作为盲评审,依据固定的行为标准对每个响应进行分类。最终生成的24,480个判断通过留一法共识进行评分。我们发现模型之间的差异不仅体现在引导如何改变其行为的程度上,还体现在它们给出的响应类型(模式)上,某些响应模式仅出现在一两个模型中。GPT-5在保持答案不变的情况下,拒绝透露其推理过程(99%对比其他模型的0%)。Claude Opus 4.7和GPT-5以不同方式抵制明确的抑制指令。以Llama作为开放权重模型,我们追踪到最大的行为分裂源于其内部结构。线性探测器以0.87的保留准确率从残差流中解码行为,而在生成过程中注入该方向使行为从0%提升至86%在干预范围内。每一发现都在令牌预算修正和假设盲评审提示的对照实验中保持一致。
cs.AI / 12 / 2608.06609
Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques
自动化项目评估:利用大型语言模型生成的批评预测项目接受与拒绝
Abstract
Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testing of the items under evaluation. We aimed to build a near-comprehensive AIE model by predicting item acceptance and rejection from item text using historical rejection data from a large-scale standardized testing program. The dataset contained 52,759 English language arts (ELA) and mathematics items with 34% permanently rejected from future operational use. Rejection reasons included poor psychometric properties, content issues, bias and sensitivity concerns, and non-content issues. We fine-tuned a DeBERTaV3-large classifier on raw item text, a second DeBERTa classifier on Qwen3-generated item critiques, and a fusion model combining representations from both. The fusion model achieved the strongest overall performance (Accuracy = .75, F1 = .64, AUC = .80, Sensitivity = .64, Specificity = .81). Prediction for math (F1 = .73, AUC = .86) was considerably more accurate than ELA (F1 = .51, AUC = .72). Lowering the decision threshold from .5 to .25 raised average sensitivity for ELA and math to .88 and .91, while reducing specificity to .31 and .56, respectively, which may be preferable in automated item generation contexts where generating items is cheaper than evaluating them. Incorporating item critiques alongside raw item text improved performance across most rejection reasons. The model assigned higher rejection probabilities to more difficult items. However, the fusion model struggled to identify items flagged for bias, sensitivity, fairness, or accessibility, especially for ELA. These findings suggest that text-based AIE is feasible in some areas and may offer a practical tool for reducing the burden of manual review and field testing, while also underscoring the importance of human review for items with fairness concerns.
Chinese Translation
自动化项目评估(AIE)是指使用计算方法评估项目质量,而无需人工专家审核或对评估项目进行现场测试。我们的目标是通过利用来自大规模标准化测试项目的历史拒绝数据,构建一个近乎全面的 AIE 模型,以预测项目文本的接受与拒绝。数据集包含 52,759 个英语语言艺术(ELA)和数学项目,其中 34% 被永久拒绝用于未来的操作性使用。拒绝原因包括心理测量属性差、内容问题、偏见和敏感性问题以及非内容问题。我们在原始项目文本上微调了一个 DeBERTaV3-large 分类器,在 Qwen3 生成的项目批评上微调了第二个 DeBERTa 分类器,并构建了一个融合模型,结合了两者的表示。融合模型在整体性能上表现最佳(准确率 = 0.75,F1 = 0.64,AUC = 0.80,灵敏度 = 0.64,特异性 = 0.81)。数学项目的预测(F1 = 0.73,AUC = 0.86)显著比 ELA 更准确(F1 = 0.51,AUC = 0.72)。将决策阈值从 0.5 降低到 0.25,使 ELA 和数学的平均灵敏度提高到 0.88 和 0.91,同时特异性分别降低到 0.31 和 0.56,这在自动化项目生成的背景下可能更为可取,因为生成项目的成本低于评估它们的成本。在原始项目文本中加入项目批评提高了大多数拒绝原因的性能。模型对更难的项目分配了更高的拒绝概率。然而,融合模型在识别被标记为偏见、敏感性、公平性或可及性的项目时表现不佳,尤其是在 ELA 中。这些发现表明,基于文本的 AIE 在某些领域是可行的,并可能为减少人工审核和现场测试的负担提供实用工具,同时也强调了对存在公平性问题的项目进行人工审核的重要性。
cs.AI / 13 / 2608.06621
NxN E-valuation: Hypothesis Certification via a Conformal CRT Null
NxN E-valuation:通过符合性条件随机化检验(CRT)进行假设认证
Abstract
We propose NxN E-valuation, a handy, e-value-based hypothesis-certification algorithm that lets a hypothesis be verified without building any case-specific certification procedure---such as constructing a dedicated null hypothesis---as long as a large enough dataset is available. The method is especially suited to LLM-based exploration systems, where LLMs are remarkably good at proposing hypotheses but suffer badly from hallucination; this hallucination prevents us from harvesting LLM outputs directly, and existing remedies each fall short. The most common solutions include letting the LLM verify or correct itself circular verification and held-out testing (where false hypotheses can still pass via spurious correlations), among other remedies detailed in the introduction. To resolve this, NxN E-valuation exploits the naturally existing large training set and lets different samples serve as null hypotheses for one another. This design directly realizes a conditional randomization test (CRT) that certifies each hypothesis. The approach can be a universally better replacement for at least LLM circular verification and held-out-data testing, provided the LLM's generations are hypotheses that apply to each individual sample.
Chinese Translation
我们提出了NxN E-valuation,这是一种基于e值的便捷假设认证算法,允许在不构建任何特定案例认证程序的情况下验证假设——例如,不需要构造专门的零假设——只要有足够大的数据集可用。该方法特别适用于基于大型语言模型(LLM)的探索系统,在这些系统中,LLM在提出假设方面表现出色,但却严重受到幻觉的困扰;这种幻觉使我们无法直接利用LLM的输出,而现有的解决方案各有不足。最常见的解决方案包括让LLM进行自我验证或自我修正的循环验证和留出测试(在这种情况下,虚假的假设仍然可以通过虚假的相关性而通过),以及引言中详细介绍的其他补救措施。为了解决这个问题,NxN E-valuation利用自然存在的大型训练集,让不同样本相互作为零假设。该设计直接实现了一种条件随机化检验(CRT),对每个假设进行认证。只要LLM的生成内容是适用于每个个体样本的假设,这种方法可以成为至少替代LLM循环验证和留出数据测试的普遍更优方案。
cs.AI / 14 / 2608.06632
Shape Your Feed: An LLM-based Agentic System for Conversational Recommendation
塑造你的推荐:基于大型语言模型的主动推荐系统
Abstract
Industrial recommendation systems predominantly adopt a passive ranking paradigm that infers user preferences from implicit behavioral signals (e.g., clicks, dwell time) rather than explicit, natural language inputs. As a result, users experience a persistent discrepancy between their explicit interests and what passive behavioral algorithms deliver, limiting their ability to express nuanced preferences or steer their feed in real time. To address this growing gap between how recommendations are optimized and how users wish to articulate their interests, we present Shape Your Feed (SYF), an LLM-based agentic recommendation framework that enables real-time, multimodal co-curation of content. SYF employs a three-tier architecture: (i) a Perception Flow that captures fine-grained user intent from text prompts, voice commands, and UI interactions; (ii) a Serving Flow that performs real-time agentic re-ranking and pruning of candidate items, grounded in a persistent Semantic Profile encoding evolving user preferences; and (iii) a Self-Evolution Flow that aligns system behavior with human judgments via Direct Preference Optimization (DPO) and an LLM-as-a-Judge ensemble. Offline evaluations show that SYF's alignment scoring module achieves 98.85% accuracy, substantially improving over strong few-shot baselines. Large-scale online A/B experiments on production traffic further demonstrate that SYF improves feed relevance and user sentiment, indicating a practical and scalable path toward interactive, user-steerable recommendation in industrial settings.
Chinese Translation
工业推荐系统主要采用被动排名范式,通过隐含的行为信号(例如点击、停留时间)推断用户偏好,而非通过明确的自然语言输入。因此,用户在明确兴趣与被动行为算法提供的内容之间存在持续的不一致,限制了他们表达细微偏好的能力或实时调整推荐内容的能力。为了解决推荐优化与用户表达兴趣之间日益扩大的差距,我们提出了“塑造你的推荐”(Shape Your Feed,SYF),这是一个基于大型语言模型的主动推荐框架,能够实现实时的多模态内容共同策划。SYF采用三层架构:(i) 感知流(Perception Flow),从文本提示、语音命令和用户界面交互中捕捉细粒度的用户意图;(ii) 服务流(Serving Flow),基于持久的语义档案(Semantic Profile)对候选项目进行实时的主动重新排序和修剪,编码不断演变的用户偏好;(iii) 自我演化流(Self-Evolution Flow),通过直接偏好优化(Direct Preference Optimization,DPO)和大型语言模型作为评判者(LLM-as-a-Judge)集成,将系统行为与人类判断对齐。离线评估表明,SYF的对齐评分模块达到了98.85%的准确率,显著优于强大的少量样本基线。在生产流量上的大规模在线A/B实验进一步表明,SYF提高了推荐内容的相关性和用户情感,指示了在工业环境中实现互动、用户可引导推荐的实用且可扩展的路径。
cs.AI / 15 / 2608.06657
TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure
TRACE:一种针对人类与人工智能控制器协调在漂移和故障下的多层基准测试
Abstract
Modern cyber-physical and AI-assisted systems couple human operators, AI decision modules, and automated controllers in a single control loop, so trustworthiness depends on the whole loop, not any one model. Yet no standard benchmark captures time-aligned, multi-layer traces of how drift and failures propagate across these layers, so we cannot diagnose where coordination breaks down, why, or how to recover. This paper targets one facet of that gap: drift, a deviation that can originate in any stack layer and that conventional single-modality monitoring cannot localize to a layer or pin to an onset time. We construct a benchmark by injecting controlled drift into traces derived from ALFRED, a grounded-instruction benchmark for everyday household tasks, yielding 1,918 drifted traces. Each trace is a time-aligned sequence of per-step records across five execution layers (state, observation, decision, rules, control), labeled with the drift type, affected layer, onset time, responsible actor, and causal mechanism, and validated by independent raters with inter-annotator agreement reported. We pair the dataset with a leak-aware protocol that removes a near-perfect onset leak, and a baseline study across classical, recurrent, and attention-based model families. Under this honest protocol, drift is identifiable and attributable well above random and majority baselines across every family (affected layer macro-F1 near 0.70, responsible actor near 0.85, causal mechanism near 0.49), and heavy attention offers no advantage over simpler models on this symbolic benchmark.
Chinese Translation
现代网络物理系统和人工智能辅助系统将人类操作员、人工智能决策模块和自动化控制器结合在一个控制回路中,因此其可信度依赖于整个回路,而非单一模型。然而,目前没有标准基准能够捕捉到时间对齐的多层次漂移和故障传播的轨迹,因此我们无法诊断协调失效的原因、位置或恢复方法。本文针对这一空白的一个方面:漂移,这是一种可以源于任何堆栈层的偏差,而传统的单模态监测无法将其定位到特定层或确定其发生时间。我们通过将受控漂移注入从ALFRED(一个针对日常家务任务的基础指令基准)衍生的轨迹中构建了一个基准,生成了1,918条漂移轨迹。每条轨迹都是跨五个执行层(状态、观察、决策、规则、控制)的逐步记录的时间对齐序列,标注了漂移类型、受影响层、发生时间、责任主体和因果机制,并由独立评审员进行验证,报告了评审员间的一致性。我们将数据集与一种漏检意识协议相结合,该协议消除了几乎完美的发生漏检,并在经典、递归和基于注意力的模型家族之间进行了基线研究。在这一诚实的协议下,漂移的可识别性和可归因性远高于随机和多数基线(受影响层的宏F1接近0.70,责任主体接近0.85,因果机制接近0.49),而在这一符号基准上,重注意力模型并未比更简单的模型提供优势。
cs.AI / 16 / 2608.06659
CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models
CellWorld:从基因级重建到空间转录组基础模型中的潜在细胞预测
Abstract
This paper shows that latent-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Existing spatial transcriptomics foundation models primarily reconstruct masked gene identities or expression values, potentially encouraging the reproduction of assay-specific technical variation and limiting representation transferability. To avoid directly reconstructing such variation, we shift the prediction target from observed gene measurements to latent cell representations and introduce CellWorld, which predicts the latent representations of masked cells from visible spatial context and a limited partial-expression hint. We pretrain four CellWorld variants, spanning 5.74M to 94.56M trainable parameters, on a corpus of 46 million human cells. Our controlled scaling experiments show that performance improves with model capacity, particularly on spatial tasks, while spatial transfer depends more on sufficient optimization and broad biological source diversity than on cell count alone. Across four held-out datasets, even CellWorld-Small, with 5.74M trainable parameters, outperforms every baseline on all 11 linear-probe benchmarks and all seven fine-tuned spatial benchmarks. Most notably, a frozen CellWorld-Large pretrained on only 5\% of the corpus with broad biological source coverage outperforms every fully fine-tuned baseline across all seven spatial benchmarks. Code is available at https://github.com/UoM-HealthAI/CellWorld.
Chinese Translation
本文展示了潜在空间预测预训练可以为空间转录组的基础模型提供可扩展的路径。现有的空间转录组基础模型主要重建被遮蔽的基因身份或表达值,这可能会鼓励特定检测技术变异的再现,并限制表示的可转移性。为了避免直接重建这种变异,我们将预测目标从观察到的基因测量转移到潜在细胞表示,并引入了CellWorld,该模型从可见的空间上下文和有限的部分表达提示中预测被遮蔽细胞的潜在表示。我们在4600万个人体细胞的语料库上预训练了四个CellWorld变体,参数量从574万到9456万不等。我们的控制扩展实验表明,模型容量的增加能改善性能,特别是在空间任务上,而空间转移更多依赖于充分的优化和广泛的生物来源多样性,而不仅仅是细胞数量。在四个保留的数据集中,即使是参数量为574万的CellWorld-Small在所有11个线性探针基准测试和所有七个微调的空间基准测试中也超越了每个基线模型。最值得注意的是,一个在仅覆盖语料库5 ext{%}的广泛生物来源上预训练的冻结CellWorld-Large在所有七个空间基准测试中超越了每个完全微调的基线。代码可在 https://github.com/UoM-HealthAI/CellWorld 获取。
cs.AI / 17 / 2608.06668
Vehicle routing problem using deep reinforcement learning - A case study about truck planning in the industry
基于深度强化学习的车辆调度问题 - 一项关于工业卡车规划的案例研究
Abstract
As an important component of the supply chain industry, transportation has experienced rapid development in the past decade with the assistance of digital platforms and intelligent algorithms. Within the field of transportation research, Vehicle Routing Problem (VRP) has remained a persistent and enduring challenge. In the realm of management science, experts, and scholars from both the industrial and academic sectors have continuously explored optimization models and algorithms to effectively address routing problems, from the classical Traveling Salesman Problem to the more general Vehicle Routing Problem. These models and algorithms are applied in real-world industrial scenarios to achieve cost optimization and reduce carbon footprints. However, due to the complexity of real-world problems, numerous specific constraints are often added, and challenges such as information opacity, uncertainty, and irrational human behavior may arise. Therefore, deploying and optimizing mathematical models for VRP in practical scenarios while maintaining optimal results poses numerous challenges. This paper discusses and provides solutions for three different logistic use cases involving external truck network design. Through these industrial case study, the paper introduces how deep reinforcement learning-based vehicle routing optimization has been implemented. As a result, it can be observed that the routes optimized by reinforcement learning agent have over 10% total cost compared to baseline results. Furthermore, the paper proposes that in future research, DRL algorithms for vehicle routing problems could be generalized into more variations of VRP.
Chinese Translation
作为供应链行业的重要组成部分,运输在过去十年中在数字平台和智能算法的帮助下经历了快速发展。在运输研究领域,车辆调度问题(Vehicle Routing Problem, VRP)一直是一个持久而复杂的挑战。在管理科学领域,来自工业和学术界的专家和学者不断探索优化模型和算法,以有效解决从经典的旅行推销员问题(Traveling Salesman Problem)到更一般的车辆调度问题的调度难题。这些模型和算法被应用于现实工业场景中,以实现成本优化和减少碳足迹。然而,由于现实问题的复杂性,通常会增加许多特定约束,并可能出现信息不透明、不确定性和非理性人类行为等挑战。因此,在实际场景中部署和优化VRP的数学模型,同时保持最佳结果,面临诸多挑战。本文讨论并提供了三个不同物流用例的解决方案,涉及外部卡车网络设计。通过这些工业案例研究,本文介绍了基于深度强化学习的车辆调度优化是如何实施的。结果表明,强化学习代理优化的路线与基线结果相比,整体成本降低超过10%。此外,本文建议在未来的研究中,车辆调度问题的深度强化学习(Deep Reinforcement Learning, DRL)算法可以推广到更多变种的VRP中。
cs.AI / 18 / 2608.06694
A Multi-Agent Framework for Automated Coarse-Grained Molecular Dynamics of Polymers
用于聚合物自动化粗粒度分子动力学的多智能体框架
Abstract
Coarse-grained (CG) molecular dynamics extends polymer simulation beyond the scales accessible to all-atom (AA) methods, but bottom-up CG modeling is laborious. The CG resolution is a design choice, so a transferable parameter set is generally not available and the potentials are derived anew for each polymer mapping. Here we present CGMas, a multi-agent framework that automates topology construction, equilibration, mapping, potential derivation, and validation from a natural-language specification of the polymer and target resolution. A large-language-model (LLM) reasoning agent infers the AA topology from polymer name, while layered self-correction resolves physical errors common to unsaturated, heteroatom-containing, and polar polymers. Downstream agents equilibrate the system, map it onto CG representation, derive potentials through Boltzmann inversion, and benchmark the model against its atomistic reference. CGMas completed all 27 homopolymer and copolymer tasks, matched the AA density to within 5% in 22, and reduced simulation from 38-88 min to 1 min, establishing agentic LLMs as a route to automated polymer coarse-graining.
Chinese Translation
粗粒度(CG)分子动力学将聚合物模拟扩展到超出全原子(AA)方法可及的尺度,但自下而上的粗粒度建模过程繁琐。CG分辨率是一个设计选择,因此通常没有可转移的参数集,每种聚合物映射都需重新推导势能。在此,我们提出了CGMas,一个多智能体框架,能够从聚合物及目标分辨率的自然语言规范中自动化拓扑构建、平衡、映射、势能推导和验证。大型语言模型(LLM)推理智能体根据聚合物名称推断AA拓扑,而分层自我修正则解决了不饱和、含杂原子和极性聚合物常见的物理错误。下游智能体对系统进行平衡,将其映射到CG表示,通过玻尔兹曼反演推导势能,并将模型与其原子参考进行基准测试。CGMas完成了所有27个均聚物和共聚物任务,在22个任务中AA密度匹配误差在5%以内,并将模拟时间从38-88分钟缩短至1分钟,确立了智能LLMs作为自动化聚合物粗粒度建模的一种途径。
cs.AI / 19 / 2608.06699
AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models
AgentPatch:粗到细的弱任务修复用于合并代理多模态大型语言模型
Abstract
Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are specialized for particular tools or environments, complicating consolidation into a single generalist. We formulate Agentic MLLM Merging and identify two challenges: asymmetric capability preservation, whereby capabilities with different interaction complexity are retained unevenly, producing weak tasks after merging, and behavior-critical forgetting, whereby losing decisive actions can derail long-horizon execution. We propose AgentPatch, a training-free coarse-to-fine repair framework. It selects a stable merged backbone, restores diluted weak-task-specific signals through Weak-Task Unique Residual Recovery, and applies an Agent-Guided Behavior-Critical Patch that recovers decisive behaviors under explicit capability protection. AgentPatch produces a single static checkpoint without routing or ensembles. Experiments across six agentic and multimodal benchmarks show that AgentPatch improves diverse merged backbones, alleviates weak-task degradation, and better balances weak-task recovery with the preservation of complementary search and agentic visual processing capabilities. Code is available at https://github.com/ziboshao/AgentPatch.
Chinese Translation
代理多模态大型语言模型(MLLMs)通过在动态环境中进行规划、工具使用和交互,扩展了多模态感知和推理。然而,当前模型专门针对特定工具或环境,这使得将其整合为单一通用模型变得复杂。我们提出了代理 MLLM 合并的概念,并识别出两个挑战:不对称能力保留,即不同交互复杂度的能力保留不均,导致合并后出现弱任务,以及行为关键遗忘,即失去决定性动作可能会 derail 长期执行。我们提出了 AgentPatch,一种无训练的粗到细修复框架。它选择一个稳定的合并主干,通过弱任务独特残差恢复(Weak-Task Unique Residual Recovery)恢复稀释的弱任务特定信号,并应用代理引导的行为关键补丁(Agent-Guided Behavior-Critical Patch),在明确的能力保护下恢复决定性行为。AgentPatch 生成一个不需要路由或集成的单一静态检查点。在六个代理和多模态基准测试中的实验表明,AgentPatch 改善了多样化的合并主干,减轻了弱任务退化,并更好地平衡了弱任务恢复与互补搜索和代理视觉处理能力的保留。代码可在 https://github.com/ziboshao/AgentPatch 获取。
cs.AI / 20 / 2608.06704
WebRider: Persona-Conditioned Intent Controllers for Live-Web Assistance
WebRider:基于角色的意图控制器用于实时网页辅助
Abstract
Delegating a web task involves more than asking a question; it requires transferring a policy: what to verify, how to handle uncertainty, which preferences matter, and when to stop. Yet, current live-web agents are evaluated solely on the final answer, ignoring the policy constraints that define the delegation. A plausible final answer can conceal violations of that policy. Our full live audit reveals this critical gap: a strong controller completes 99.2% of tasks but honors all policy constraints in only 38.8% of cases. Finishing does not imply fidelity. WebRider bridges this gap by formalizing the delegated policy as an intent contract---an operational record of goals, constraints, evidence obligations, answer form, and task-local persona controls that must hold even as web pages change. WebRider employs a hierarchical architecture: a top-layer controller maintains the contract, a middle layer realizes intentions as guarded executable actions, and a tool layer executes these actions via browser, search, and maps tools. Our benchmark, RiderBench, evaluates this design on 4,096 live-web contracts across 42 public websites, auditing both the internal contract state and the visible user experience to determine if a rollout preserved its policy and if the steps were persona-consistent. The guarded middle interface also serves as a high-quality training signal; an 8B action-policy model trained through this interface outperforms executable-only baselines under a fixed controller. By making the browsing path a first-class object, WebRider enables a system that is auditable, human-judgeable, and learnable without conflating action realization with final-answer decisions. Dataset URL: hf.co/datasets/WebRider/WebRider.
Chinese Translation
委托网页任务不仅仅是提出一个问题;它还需要转移一个策略:验证什么、如何处理不确定性、哪些偏好重要以及何时停止。然而,目前的实时网页代理仅根据最终答案进行评估,忽略了定义委托的策略约束。一个看似合理的最终答案可能掩盖了对该策略的违反。我们的全面实时审计揭示了这一关键差距:一个强大的控制器完成了99.2%的任务,但在仅38.8%的情况下遵循了所有策略约束。完成任务并不意味着忠实。WebRider通过将委托策略形式化为意图合同——一个关于目标、约束、证据义务、答案形式和任务本地角色控制的操作记录——来弥补这一差距,这些内容必须在网页变化时依然有效。WebRider采用分层架构:顶层控制器维护合同,中间层将意图实现为受保护的可执行动作,工具层通过浏览器、搜索和地图工具执行这些动作。我们的基准测试RiderBench在42个公共网站上评估了4096个实时网页合同,审计内部合同状态和可见用户体验,以确定推出是否保留了其政策,以及步骤是否与角色一致。受保护的中间接口还作为高质量的训练信号;通过该接口训练的8B动作策略模型在固定控制器下优于仅可执行的基线。通过将浏览路径作为一等公民对象,WebRider使系统可审计、可由人类评判且可学习,而不将动作实现与最终答案决策混为一谈。数据集网址:hf.co/datasets/WebRider/WebRider。
cs.AI / 21 / 2608.06713
MolBioKG: Grounding Out-of-Graph Molecules in Biomedical Knowledge Graphs via Multi-Resolution Structural Anchoring
MolBioKG:通过多分辨率结构锚定将图外分子嵌入生物医学知识图谱
Abstract
Biomedical knowledge graphs (KGs) accelerate drug discovery, but standard pipelines assume query molecules already exist as graph entities, leaving unregistered molecules disconnected. We address this cold-start challenge, termed the out-of-graph molecule problem, by introducing MolBioKG. This two-layer system grounds unseen molecules in biomedical evidence via multi-resolution structural anchoring. It connects an index of 2.74 million molecules (represented by scaffolds, fragments, functional groups, and fingerprints) to a 9.6-million-edge KG. Given only a SMILES string, MolBioKG retrieves structurally related graph entities and traverses their biomedical neighborhoods without task-specific training. It features two inference mechanisms: static multi-anchor retrieval using Reciprocal Rank Fusion, and Adapt-KG, a tool-using LLM policy for adaptive traversal. Evaluated across in-graph link recovery, complex multi-hop reasoning, and out-of-graph generalization, MolBioKG outperforms strong baselines. Notably, it raises Hits@10 from 0.585 to 0.876 in multi-hop reasoning and out-of-graph target recall from 0.145 to 0.269, all while ensuring predictions retain traceable structural anchors and source-attributed KG evidence.
Chinese Translation
生物医学知识图谱(KGs)加速了药物发现,但标准流程假设查询分子已经作为图实体存在,这使得未注册的分子处于孤立状态。我们通过引入MolBioKG来解决这一冷启动挑战,称为图外分子问题。该双层系统通过多分辨率结构锚定将未见分子嵌入生物医学证据中。它将274万种分子(由骨架、片段、功能团和指纹表示)与一个960万边的知识图谱连接起来。仅给定一个SMILES字符串,MolBioKG就能检索结构相关的图实体,并在没有特定任务训练的情况下遍历它们的生物医学邻域。该系统具有两种推理机制:使用互惠排名融合的静态多锚点检索,以及Adapt-KG,一种用于自适应遍历的工具使用LLM策略。在图内链接恢复、复杂多跳推理和图外泛化的评估中,MolBioKG的表现优于强基线。值得注意的是,在多跳推理中,Hits@10从0.585提高到0.876,而图外目标召回率从0.145提高到0.269,同时确保预测保留可追溯的结构锚和来源归属的知识图谱证据。
cs.AI / 22 / 2608.06714
The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows
优化器即代理:跨提示、程序和机器学习工作流的推理驱动搜索
Abstract
Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamentally different question: how much of this search policy can be internalized by a single tool-using agent? We present ReASearch, a unified framework for reasoning-driven optimization in which the agent autonomously decides what to evaluate, how to diagnose failures, which edits to make, and when to verify or restart. Rather than serving only as a proposal generator guided by hand-designed heuristics, the agent actively analyzes outcomes, allocates budget, and refines its strategy over long horizons through persistent memory. With a shared agent loop and domain-specific tools, ReASearch instantiates the exact same scaffold to optimize prompts, programs, and ML workflows. Across 14 diverse tasks, it is competitive with and mostly better than specialized optimization systems, achieving gains of 2% to 40% over strong domain-specific baselines, and in some cases discovering solutions that improve on prior human best-known results. Crucially, we observe that complex search behaviors, which are typically implemented by explicit controllers, emerge naturally from the agent's reasoning process.
Chinese Translation
近期用于优化提示、程序和机器学习工作流的系统通常依赖于显式的外部控制器,如进化搜索、赌博算法或文本梯度方法。我们提出一个根本不同的问题:这一搜索策略有多少可以被单一工具使用的代理内化?我们提出了 ReASearch,这是一个推理驱动优化的统一框架,其中代理自主决定评估内容、如何诊断失败、进行哪些编辑以及何时验证或重启。该代理不仅仅是一个由手工设计的启发式方法引导的提案生成器,而是积极分析结果、分配预算,并通过持久记忆在长时间范围内优化其策略。通过共享的代理循环和特定领域的工具,ReASearch 实现了相同的框架来优化提示、程序和机器学习工作流。在14个多样化的任务中,其性能与专门的优化系统相当,且大多数情况下优于它们,取得了比强大的领域特定基线高出2%到40%的提升,并在某些情况下发现了超越先前人类已知最佳结果的解决方案。重要的是,我们观察到,通常由显式控制器实现的复杂搜索行为,自然地从代理的推理过程中涌现。
cs.AI / 23 / 2608.06727
bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning
bioMoR:生物指导的混合递归模型用于有效的基因组学习
Abstract
Transformer models for high-dimensional omics analysis process thousands of genes or pathways, although only a subset requires deep computation. Mixture-of-Recursions (MoR) improves efficiency through adaptive token-choice or expert-choice routing. We propose bioMoR, which, to the best of our knowledge, is the first framework to apply MoR to gene-level and pathway-level learning. Our contributions include identifying three locations for integrating structured biological knowledge within an MoR backbone: graph-based information sharing refines token embeddings, a structural bias guides self-attention toward biologically related tokens, and a graph-aware router uses neighborhood information to determine each token's recursion depth. These techniques are centered on our insight that additional knowledge of token interaction can effectively help models construct embeddings and select which tokens should be learned more deeply. Across eight benchmarks spanning diverse omics data types and evaluated under a unified five-fold cross-validation protocol, bioMoR improves average macro-F1 by 8.2 percentage points and balanced accuracy by 7.1 percentage points over the strongest biology-agnostic MoR baseline while using 75 percent fewer parameters and up to 58 percent fewer FLOPs than a non-recursive Transformer. The selected marker genes or pathways provide biological interpretability, while their token-specific recursion depths reveal how computation is allocated.
Chinese Translation
用于高维组学分析的Transformer模型处理数千个基因或通路,尽管只有一部分需要深度计算。混合递归模型(Mixture-of-Recursions, MoR)通过自适应的标记选择或专家选择路由提高了效率。我们提出的bioMoR,尽我们所知,是第一个将MoR应用于基因级和通路级学习的框架。我们的贡献包括在MoR骨架中识别三个整合结构化生物知识的位置:基于图的信息共享精炼了标记嵌入,结构偏差引导自注意力朝向生物相关的标记,而图感知路由器利用邻域信息来确定每个标记的递归深度。这些技术围绕我们的洞察展开,即额外的标记交互知识可以有效帮助模型构建嵌入并选择哪些标记应更深入地学习。在涵盖多种组学数据类型的八个基准测试中,bioMoR在统一的五折交叉验证协议下将平均宏F1提高了8.2个百分点,平衡准确率提高了7.1个百分点,同时使用的参数减少了75%,比非递归Transformer少了多达58%的FLOPs。所选的标记基因或通路提供了生物学解释性,而它们特定于标记的递归深度揭示了计算如何分配。
cs.AI / 24 / 2608.06732
From Cheap Fakes to Pure Synthesis: Addressing the New Era of T2V Fake News Videos
从廉价伪造到纯合成:应对T2V假新闻视频的新纪元
Abstract
Recent text-to-video (T2V) generation models enable fake news videos to be synthesized from scratch, shifting the threat beyond cheap fakes assembled from existing footage. Such news videos can closely match fabricated narratives, creating a modality alignment trap for existing detectors. Existing datasets lack pure synthesis fake news videos. Although directly prompting T2V models with descriptions of fake news videos can yield perfectly aligned samples, it reduces the fake news video detection (FNVD) to unimodal shortcuts and causes semantic-visual degeneration. To counter this, we formulate T2V-FNVD as a novel ternary classification task with three labels (real, cheap fake, and pure synthesis fake) and construct the first pure synthesis fake news video dataset (PS-FNVD). PS-FNVD includes fabricated events with aligned deception (Type 1) and true events with false visual provenance (Type 2), preventing models from exploiting unimodal shortcuts. Furthermore, we propose the Reasoning-guided T2V-FNVD (R-T2V) framework. Trained through conditioned rationale generation and supervised fine-tuning, R-T2V integrates high-level semantic logic with low-level physical generative traces to predict the ternary veracity label. Extensive experiments across 10 prevailing baselines show that R-T2V achieves the state-of-the-art performance, outperforming the second-best baseline by 12.20 percentage points in accuracy and 8.46 percentage points in macro $F_1$.
Chinese Translation
近期的文本到视频(T2V)生成模型使得假新闻视频能够从零开始合成,威胁的范围超出了由现有素材拼凑而成的廉价伪造。这类新闻视频能够与虚构叙事紧密匹配,给现有检测器造成了模态对齐陷阱。现有数据集中缺乏纯合成的假新闻视频。尽管直接用假新闻视频的描述来提示T2V模型可以生成完美对齐的样本,但这将假新闻视频检测(FNVD)简化为单模态捷径,并导致语义-视觉退化。为此,我们将T2V-FNVD构建为一种新颖的三元分类任务,包含三个标签(真实、廉价伪造和纯合成伪造),并构建了第一个纯合成假新闻视频数据集(PS-FNVD)。PS-FNVD包括具有对齐欺骗的虚构事件(类型1)和具有虚假视觉来源的真实事件(类型2),防止模型利用单模态捷径。此外,我们提出了基于推理的T2V-FNVD(R-T2V)框架。通过条件推理生成和监督微调进行训练,R-T2V将高级语义逻辑与低级物理生成痕迹相结合,以预测三元真实性标签。在10个主流基线上的广泛实验表明,R-T2V达到了最先进的性能,在准确率上超越第二名基线12.20个百分点,宏观$F_1$超越8.46个百分点。
cs.AI / 25 / 2608.06735
IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents
IB-RL:用于战略对话代理的孤立双边强化学习
Abstract
Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewards, such as mathematical reasoning and code execution. In these settings, the environment follows fixed rules and does not adapt strategically to the agent. Strategic dialogue differs in this respect: the environment is another agent that adapts to the policy, and success depends on the interaction between the two sides. Despite this interactive nature, current RL approaches typically train a target agent against a fixed counterpart or simulator. We find that this training paradigm encourages the policy to exploit counterpart-specific regularities rather than learn strategies that generalize across counterparts. We call this problem the static-counterpart mismatch, which we quantify directly in our experiments. To address it, we propose Isolated Bilateral Reinforcement Learning (IB-RL), in which the two roles coevolve through joint rollouts while each role optimizes its own reward through fully independent advantages, action masks, and update paths. We evaluate frozen policies against fully independent held-out counterparts in both domains. On Vehicle TeleSales, IB-RL achieves 89.6% Success@1, compared to 84.6% for the best unilateral RL baseline. On Deal-or-NoDeal, it reaches 98.4% agreement against DeepSeek V4 Pro, compared to 86.4% for the best unilateral baseline. These results indicate that jointly training both roles with strict peragent isolation produces policies that generalize more effectively to unseen counterparts.
Chinese Translation
强化学习(RL)在提高大型语言模型(LLMs)在具有静态、可验证奖励的任务(如数学推理和代码执行)方面取得了显著成果。在这些环境中,环境遵循固定规则,并且不会战略性地适应代理。而战略对话在这方面有所不同:环境是另一个适应策略的代理,成功取决于双方之间的互动。尽管具有这种互动特性,目前的RL方法通常是在固定的对手或模拟器上训练目标代理。我们发现这种训练范式促使策略利用特定对手的规律,而不是学习在不同对手之间普遍适用的策略。我们将这个问题称为静态对手不匹配,并在实验中直接量化了它。为了解决这个问题,我们提出了孤立双边强化学习(IB-RL),在该方法中,两个角色通过联合回合共同进化,同时每个角色通过完全独立的优势、行动掩码和更新路径优化自己的奖励。我们在两个领域中评估了冻结策略与完全独立的保留对手的表现。在车辆电销(Vehicle TeleSales)中,IB-RL达到了89.6%的成功率(Success@1),而最佳单边RL基线为84.6%。在《交易还是不交易》(Deal-or-NoDeal)中,它与DeepSeek V4 Pro达成了98.4%的协议,而最佳单边基线为86.4%。这些结果表明,严格的每代理隔离下联合训练两个角色能够产生更有效地推广到未见对手的策略。
cs.AI / 26 / 2608.06745
MemPrism: Task-Conditioned Relational Memory Views for Long-Horizon Agents
MemPrism:用于长时间跨度智能体的任务条件关系记忆视图
Abstract
Long-horizon agents rely on memory to reuse experiences, yet existing memory systems often assume that evidence can be directly consumed through a fixed representation. This leads to representation mismatch, where relevant information is available but not organized for the current decision. To this end, we propose MemPrism, a task-conditioned relational memory framework that separates persistent experience storage from decision-time working memory. MemPrism records interactions as the event stream and dynamically constructs relational views according to the current task context. A lightweight view policy selects the relation structure, evidence range, outcome condition, and granularity, while a deterministic composer and render transform historical facts into a temporary optical working-memory view for a frozen task policy. Experiments on long-horizon embodied and web-agent benchmarks show that MemPrism consistently improves the task performance, especially as trajectories become longer, while reducing memory token consumption. Furthermore, the learned view policy transfers across different VLMs without additional adaptation, demonstrating the effectiveness of task-conditioned relational views as a general memory interface for agents.
Chinese Translation
长时间跨度的智能体依赖于记忆来重用经验,然而现有的记忆系统通常假设证据可以通过固定的表示直接被使用。这导致了表示不匹配的问题,即相关信息虽然可用,但未能为当前决策组织好。为此,我们提出了MemPrism,一种任务条件关系记忆框架,它将持久经验存储与决策时的工作记忆分开。MemPrism将交互记录为事件流,并根据当前任务上下文动态构建关系视图。一个轻量级的视图策略选择关系结构、证据范围、结果条件和粒度,而一个确定性的组合器和渲染器将历史事实转化为一个临时的光学工作记忆视图,以适应冻结的任务策略。在长时间跨度的具身智能体和网络智能体基准测试中的实验表明,MemPrism始终提高了任务性能,尤其是在轨迹变得更长时,同时减少了记忆令牌的消耗。此外,学习到的视图策略可以在不同的视觉语言模型(VLM)之间迁移,无需额外的适应,证明了任务条件关系视图作为智能体通用记忆接口的有效性。
cs.AI / 27 / 2608.06752
Mind the Gap: A Dual Knowledge Graph Framework for Unified Multi-task User Intent Inference
关注差距:一种用于统一多任务用户意图推断的双重知识图谱框架
Abstract
This paper proposes DKG-MTI, a dual knowledge graph framework for unified multi-task user intent inference from online travel reviews. Existing approaches often rely on hierarchical pipelines that suffer from error propagation or retrieval methods that ignore structural relationships in domain knowledge. To address these limitations, we introduce an inference-only knowledge augmentation framework that dynamically constructs a User-Specific Intent Knowledge Graph from each review and aligns it with a Global Hotel Knowledge Graph through structure-aware semantic smoothing. The aligned knowledge is combined with the original review and processed by a large language model to simultaneously predict aspect ratings and generate reverse user intent statements. Experiments on TripAdvisor reviews show that DKG-MTI consistently outperforms strong LLM and retrieval-based baselines in both classification and intent generation tasks, demonstrating the effectiveness of structure-aware knowledge alignment for scalable and explainable intent inference.
Chinese Translation
本文提出了DKG-MTI,一种用于从在线旅游评论中进行统一多任务用户意图推断的双重知识图谱框架。现有方法通常依赖于层次化的流程,这些流程容易受到错误传播的影响,或者采用忽视领域知识中结构关系的检索方法。为了解决这些局限性,我们引入了一种仅进行推断的知识增强框架,该框架动态构建用户特定意图知识图谱,并通过结构感知的语义平滑将其与全球酒店知识图谱对齐。对齐后的知识与原始评论相结合,并由大型语言模型处理,以同时预测方面评分并生成反向用户意图声明。在TripAdvisor评论上的实验表明,DKG-MTI在分类和意图生成任务中始终优于强大的大型语言模型和基于检索的基线,证明了结构感知知识对齐在可扩展和可解释的意图推断中的有效性。
cs.AI / 28 / 2608.06756
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
Capek 0.5:一种以执行为中心的视觉-语言模型用于具身智能
Abstract
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Meeting these demands requires complementary capabilities that differ in supervision signals, prediction formats, and verification criteria. Existing approaches typically develop these capabilities against isolated, task-specific objectives, leaving open how they should be organized and integrated around execution as a whole. We present Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy. Rather than organizing training by datasets or tasks, the taxonomy groups embodied capabilities according to their functional roles throughout execution and comprises four capability families: Spatial Reasoning, Temporal Understanding, Action Guidance, and State Verification. Each capability is first acquired by a dedicated specialist through reinforcement learning with verifiable rewards from a shared backbone, and the specialists are then consolidated into a single inference-time model through weight-space merging followed by routed policy-space distillation. We instantiate Capek 0.5 at the 2B and 35B-A3B scales and evaluate it from three complementary perspectives: comprehensive benchmark suites including Capek-StateBench, a new benchmark for state verification; a controlled study of capability retention from specialists to the unified model; and closed-loop evaluation in simulated embodied environments. Capek 0.5 improves the large majority of matched benchmark rows over its initialization, retains all four specialized capabilities in one checkpoint with quantified losses, and transfers to closed-loop embodied task execution.
Chinese Translation
视觉-语言模型越来越多地作为具身代理的推理核心。机器人执行本质上是迭代的:每个动作都会重塑场景和物理状态,不断更新需要感知、推理和验证的内容。满足这些需求需要互补的能力,这些能力在监督信号、预测格式和验证标准上有所不同。现有的方法通常针对孤立的、特定任务的目标开发这些能力,但尚未明确如何围绕整体执行进行组织和整合。我们提出了Capek 0.5,这是一种围绕以执行为中心的能力分类法构建的具身视觉-语言模型。该分类法不是通过数据集或任务来组织训练,而是根据能力在执行过程中的功能角色对具身能力进行分组,包含四个能力家族:空间推理、时间理解、行动指导和状态验证。每种能力首先由专门的专家通过强化学习获得,使用来自共享骨干网络的可验证奖励,然后通过权重空间合并和路由策略空间蒸馏将专家整合为一个单一的推理时模型。我们在2B和35B-A3B规模上实例化了Capek 0.5,并从三个互补的角度进行评估:包括Capek-StateBench的综合基准套件,这是一个用于状态验证的新基准;对从专家到统一模型的能力保留的控制研究;以及在模拟具身环境中的闭环评估。Capek 0.5在初始化后改善了绝大多数匹配基准行,在一个检查点中保留了所有四种专业能力,并量化了损失,能够转移到闭环具身任务执行中。
cs.AI / 29 / 2608.06765
LiFTER: A Grounded Neuro-Symbolic Microscope for Continuous-Time Dynamic Graph Forecasting
LiFTER:一种用于连续时间动态图预测的基础神经符号显微镜
Abstract
Continuous-time dynamic graph models predict future links by compressing past interactions into neural states. Although effective for forecasting, this computation obscures which entities are shared across events and how temporal patterns contribute to a prediction. We treat this gap as a property of the predictive architecture rather than a problem to be addressed after prediction. Link-Fact Temporal Rule Inducer (LiFTER) is a neuro-symbolic predictor that preserves observed interactions as grounded temporal facts and applies executable tempo- ral rules to pre-query facts. Each score is a signed sum of rule exe- cutions whose historical facts, entity bindings, and temporal order are explicitly satisfied. The evidence and rules responsible for a prediction can therefore be inspected, independently recomputed, and intervened upon. Across four CTDG benchmarks, LiFTER achieves competitive historical-negative forecasting and the highest macro explanation ac- curacy and deletion fidelity. The same architecture also serves as a microscope that separates the contributions of recurrence, history po- sition, and transition across datasets and traces them to individual facts. Independent execution reconstructs all logits for 19,664 test predictions with a maximum error of 0.0000131. LiFTER turns future-link forecasting into a verifiable grounded computation.
Chinese Translation
连续时间动态图模型通过将过去的交互压缩为神经状态来预测未来的链接。尽管这种方法在预测方面有效,但其计算过程模糊了跨事件共享的实体以及时间模式如何贡献于预测。我们将这一差距视为预测架构的一个特性,而不是在预测后需要解决的问题。链接事实时间规则诱导器(Link-Fact Temporal Rule Inducer,LiFTER)是一种神经符号预测器,它将观察到的交互保留为基础的时间事实,并对事实应用可执行的时间规则进行预查询。每个得分都是规则执行的有符号总和,其历史事实、实体绑定和时间顺序均被明确满足。因此,负责预测的证据和规则可以被检查、独立重新计算和干预。在四个连续时间动态图基准上,LiFTER实现了具有竞争力的历史负预测以及最高的宏观解释准确性和删除保真度。同一架构还充当显微镜,分离数据集中重复性、历史位置和转变的贡献,并将其追溯到个别事实。独立执行重构了19,664个测试预测的所有logits,最大误差为0.0000131。LiFTER将未来链接预测转变为可验证的基础计算。
cs.AI / 30 / 2608.06770
Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts
Surg-UniWorld:具有多模态控制专家的统一外科世界模型
Abstract
Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument--tissue interactions. However, existing methods lack a unified multimodal control paradigm, while direct fusion of heterogeneous visual conditions often causes anatomical distortion, instrument appearance drift, and temporally inconsistent interactions. In this work, we propose {Surg-UniWorld}, a unified surgical world model with multimodal control experts. Surg-UniWorld first constructs a {Hierarchical Surgical Anchor} from first-frame appearance and hierarchical semantic masks to preserve persistent scene identity, anatomical organization, and interaction boundaries. {Anchor-Relative Modality Experts} then interpret edge, depth, and optical-flow evidence relative to the shared anchor, capturing complementary boundary, geometric, and motion information. A {Multimodal Control Expert} further performs contribution-preserving stage-wise composition of the activated modality increments and generates control hints for the Wan2.2 video diffusion backbone. To support multimodal surgical world modeling, we further construct Cholec80-SurgWAM, a benchmark for controllable surgical video generation. Extensive experiments demonstrate that Surg-UniWorld consistently outperforms existing controllable video generation methods and surgical world-model baselines in generation quality, temporal consistency, and multimodal controllability.
Chinese Translation
可控的外科世界模型能够通过合成真实的工具与组织之间的互动,为外科人工智能和模拟提供生成基础。然而,现有方法缺乏统一的多模态控制范式,而异构视觉条件的直接融合往往导致解剖扭曲、工具外观漂移和时间不一致的互动。在本研究中,我们提出了Surg-UniWorld,一个具有多模态控制专家的统一外科世界模型。Surg-UniWorld首先从第一帧外观和层次语义掩膜构建一个层次外科锚点,以保持持久的场景身份、解剖组织和互动边界。锚点相对模态专家(Anchor-Relative Modality Experts)随后解释相对于共享锚点的边缘、深度和光流证据,捕捉互补的边界、几何和运动信息。多模态控制专家(Multimodal Control Expert)进一步对激活的模态增量进行贡献保持的阶段性组合,并为Wan2.2视频扩散主干生成控制提示。为了支持多模态外科世界建模,我们进一步构建了Cholec80-SurgWAM,一个可控外科视频生成的基准。大量实验表明,Surg-UniWorld在生成质量、时间一致性和多模态可控性方面始终优于现有的可控视频生成方法和外科世界模型基线。
cs.AI / 31 / 2608.06808
Evolving Parallel Algorithm Portfolios via Potential-Aware Instance Generation with LLMs
通过潜力感知实例生成与大型语言模型演化并行算法组合
Abstract
The Automatic Construction of Portfolios via Large Language Models (LLM-ACP) suffers from poor generalization in practical few-shot scenarios when solving complex combinatorial optimization problems. Instance and algorithm co-evolution frameworks address this by expanding the training dataset with generated hard instances on which the current algorithm portfolio underperforms, thereby enhancing generalization. However, this paradigm faces two critical limitations: evaluating instance hardness relies on high-quality reference solutions, and single-mode generation patterns limit instance diversity. To overcome these limitations, we introduce the Potential-aware Instance and Algorithm Co-evolution (PIAC) framework. Our core contribution is twofold. First, we propose potential gain, a novel metric that eliminates the need for reference solutions. This metric estimates generalization gain by perturbing the generated algorithms and assessing their improvement potential on generated problem instances. Second, PIAC leverages LLMs to synthesize diverse instance mutators, exploring a broader region of the problem-instance space and thereby enhancing the portfolio's generalization capabilities. Given that perturbation spaces vary across different algorithms, we instantiate our framework on Greedy Constructive, Ant Colony Optimization, and Guided Local Search algorithmic backbones. Comprehensive evaluations on the Traveling Salesman Problem (TSP) and Capacitated Vehicle Routing Problem (CVRP) across six distinct data distributions demonstrate that PIAC consistently outperforms state-of-the-art LLM-ACP baselines, notably achieving a 19.76% relative improvement for TSP Greedy Constructive portfolios.
Chinese Translation
通过大型语言模型(LLM-ACP)自动构建算法组合在解决复杂组合优化问题时,在实际的少量样本场景中表现出较差的泛化能力。实例与算法共同演化框架通过扩展训练数据集,生成当前算法组合表现不佳的困难实例,从而提高泛化能力。然而,这一范式面临两个关键限制:评估实例难度依赖于高质量的参考解,且单一模式的生成方式限制了实例的多样性。为克服这些限制,我们提出了潜力感知实例与算法共同演化(PIAC)框架。我们的核心贡献有两个方面。首先,我们提出了潜在增益(potential gain),这是一种新颖的度量标准,消除了对参考解的需求。该度量通过扰动生成的算法并评估其在生成问题实例上的改进潜力来估计泛化增益。其次,PIAC利用LLMs合成多样化的实例变异器,探索更广泛的问题实例空间,从而增强组合的泛化能力。鉴于不同算法的扰动空间各不相同,我们在贪心构造、蚁群优化和引导局部搜索等算法基础上实例化我们的框架。在六种不同数据分布下对旅行商问题(TSP)和容量限制车辆路径问题(CVRP)的全面评估表明,PIAC始终优于最先进的LLM-ACP基线,特别是在TSP贪心构造组合中实现了19.76%的相对提升。
cs.AI / 32 / 2608.06861
Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents
Gated-BEPO:用于大型语言模型代理的信心门控贝尔曼信用分配
Abstract
Training large language model agents in long-horizon environments requires assigning credit from sparse terminal outcomes to individual actions. Existing critic-free methods propagate trajectory-level rewards uniformly across steps, while recent approaches construct step-level groups by matching repeated states and compare actions within each group. The former cannot distinguish useful actions in failed trajectories from ineffective actions in successful ones. The latter rely on step credit derived directly from individual trajectory outcomes and fixed-weight fusion with episode-level credit. We propose Gated-BEPO, which derives step-level credit from empirical rollout graphs. For each rollout group, Gated-BEPO constructs an empirical graph and estimates node values through a mean-backup Bellman fixed point that reflects the empirical action distribution of the current policy. We then accumulate these temporal-difference residuals along each sampled trajectory using generalized advantage estimation, yielding step-level Bellman advantages that capture both immediate and downstream effects. To adaptively fuse episode- and step-level credit, a confidence gate incorporates Bellman credit only at states with multiple observed successors and otherwise uses episode-level credit. Experiments on WebShop, ALFWorld, and visual Sokoban show consistent improvements across language and vision-language models, while diagnostic ablations support the effectiveness of Bellman fixed-point value estimation and show that step-level credit should be incorporated selectively rather than uniformly into the final advantage.
Chinese Translation
在长时间跨度环境中训练大型语言模型代理需要将稀疏终端结果的信用分配到各个动作。现有的无评论者方法在步骤之间均匀传播轨迹级奖励,而最近的方法通过匹配重复状态构建步骤级组,并在每个组内比较动作。前者无法区分失败轨迹中的有效动作与成功轨迹中的无效动作。后者直接依赖于个别轨迹结果得出的步骤信用,并与情节级信用进行固定权重融合。我们提出了Gated-BEPO,它从经验回滚图中推导步骤级信用。对于每个回滚组,Gated-BEPO构建一个经验图,并通过反映当前策略的经验动作分布的均值备份贝尔曼固定点来估计节点值。然后,我们使用广义优势估计沿每个采样轨迹累积这些时间差残差,从而产生捕捉即时和下游效应的步骤级贝尔曼优势。为了自适应地融合情节级和步骤级信用,信心门控仅在观察到多个后继状态的状态中纳入贝尔曼信用,否则使用情节级信用。在WebShop、ALFWorld和视觉Sokoban上的实验显示,语言和视觉-语言模型的一致性改进,而诊断消融实验支持贝尔曼固定点值估计的有效性,并表明步骤级信用应选择性地而非均匀地纳入最终优势。
cs.AI / 33 / 2608.06871
CEDAR: Agent-Orchestrated Tree Search for Goal-Directed Optimization of Complex Systems
CEDAR:面向目标的复杂系统优化的代理协调树搜索
Abstract
Complex systems, core objects of study in artificial life, model diverse phenomena through nonlinear, feedback-driven interactions that produce emergent behavior, with applications from population dynamics and biology to economic policy and strategic decision-making. Yet the difficulty of predicting how feedback structure gives rise to emergent behavior, a central open problem in artificial life, makes goal-directed design exceptionally challenging. In established practice, system structures are written in specialized modeling languages such as DYNAMO or STELLA, compounding the challenge with labor-intensive workflows that limit adoption and hinder timely decision-making. To address these challenges, we introduce CEDAR, an autonomous method that uses Large Language Model (LLM) agents to discover complex systems satisfying user-specified behavioral goals. Our key innovation is an LLM-driven Monte Carlo Tree Search (MCTS) deeply coupled with complex systems: at each iteration, an LLM Judge evaluates emergent behavior against specified goals and an LLM Editor proposes improved variants, with the Judge acting as a fitness function and the Editor as a variation operator, akin to a generate-and-evaluate loop in evolutionary computation. We represent complex systems as a restricted, runnable subset of Python with domain-specific primitives, letting LLMs modify system dynamics directly. CEDAR formalizes this as an MCTS variant with an LLM-parameterized transition kernel and value function, enabling goal-directed discovery of complex system behaviors while preserving solution diversity, and its LLM-based interpretability reveals how structural changes drive emergent behavior. CEDAR reduces human effort while enabling capabilities difficult to achieve with existing approaches, facilitating broader adoption of complex systems across domains.
Chinese Translation
复杂系统是人工生命研究的核心对象,通过非线性、反馈驱动的交互建模多样现象,产生涌现行为,应用范围涵盖种群动态、生物学、经济政策和战略决策等。然而,预测反馈结构如何导致涌现行为的困难是人工生命中的一个核心开放问题,这使得面向目标的设计异常具有挑战性。在现有实践中,系统结构通常使用专门的建模语言(如 DYNAMO 或 STELLA)编写,这加大了挑战,因为这些劳动密集型的工作流程限制了采用并妨碍了及时决策。为了解决这些挑战,我们提出了 CEDAR,一种自主方法,利用大型语言模型(LLM)代理发现满足用户指定行为目标的复杂系统。我们的关键创新是将 LLM 驱动的蒙特卡洛树搜索(MCTS)与复杂系统深度耦合:在每次迭代中,LLM 判别器评估涌现行为与指定目标的符合程度,而 LLM 编辑器则提出改进的变体,判别器充当适应度函数,编辑器作为变异操作符,类似于进化计算中的生成与评估循环。我们将复杂系统表示为一个受限的、可运行的 Python 子集,配备领域特定的原语,使 LLM 能够直接修改系统动态。CEDAR 将此形式化为一种具有 LLM 参数化转移核和价值函数的 MCTS 变体,能够在保持解的多样性的同时,面向目标地发现复杂系统行为,其基于 LLM 的可解释性揭示了结构变化如何驱动涌现行为。CEDAR 减少了人力投入,同时实现了现有方法难以达到的能力,促进了复杂系统在各个领域的更广泛应用。
cs.AI / 34 / 2608.06891
SkillEval: Decomposing Agent Skill Quality into Interpretable Signals
SkillEval:将智能体技能质量分解为可解释信号
Abstract
Agent skills provide reusable procedural knowledge that helps agents solve specialized tasks. As their use expands, evaluating skill quality becomes increasingly important. Existing evaluations often measure skill quality by testing whether a skill improves performance on specific downstream tasks. However, a reusable skill may apply to multiple task scenarios. Downstream evaluation mainly reflects the compatibility between a skill and the evaluated task, provides only a partial view of skill quality, and does not identify which aspect of the skill should be improved. We find that general properties of the \texttt{SKILL.md} document play an important role in skill quality. To evaluate these properties, we propose \textbf{SkillEval}, an interpretable framework for document-level skill evaluation. SkillEval evaluates each property using a fixed and inspectable scoring direction, producing interpretable scores. It further measures and reduces the influence of unrelated document features, such as length and formatting, so that each score captures its intended semantic property more specifically. Specifically, SkillEval learns an interpretable direction for each quality property from controlled positive--negative skill pairs in the hidden representation space of the model, and scores a new skill by projecting its representation onto these fixed directions. We use SkillEval to evaluate skills in controlled quality tests and show that SkillEval reliably distinguishes skills of different quality. In addition, SkillEval scores closely reflect downstream task performance, providing an early indication of whether a skill is likely to help an agent complete a task. We further explore SkillEval for diagnosing weaknesses in skill documents and guiding targeted revisions. The revised skills improve the targeted properties and achieve higher pass rates on downstream tasks.
Chinese Translation
智能体技能提供可重用的程序知识,帮助智能体解决专业任务。随着其应用范围的扩大,评估技能质量变得越来越重要。现有的评估通常通过测试技能是否提高特定下游任务的表现来衡量技能质量。然而,可重用的技能可能适用于多种任务场景。下游评估主要反映技能与被评估任务之间的兼容性,仅提供技能质量的部分视角,并未识别出技能的哪个方面需要改进。我们发现, exttt{SKILL.md} 文档的一般属性在技能质量中起着重要作用。为了评估这些属性,我们提出了 extbf{SkillEval},一个用于文档级技能评估的可解释框架。SkillEval 使用固定且可检查的评分方向来评估每个属性,生成可解释的分数。它进一步测量并减少与文档特征无关的影响,例如长度和格式,从而使每个分数更具体地捕捉其预期的语义属性。具体而言,SkillEval 从模型的隐藏表示空间中的受控正负技能对中学习每个质量属性的可解释方向,并通过将其表示投影到这些固定方向上来对新技能进行评分。我们使用 SkillEval 在受控质量测试中评估技能,并显示 SkillEval 能可靠地区分不同质量的技能。此外,SkillEval 的评分与下游任务表现密切相关,提供了技能是否可能帮助智能体完成任务的早期指示。我们进一步探索 SkillEval 用于诊断技能文档中的弱点并指导有针对性的修订。修订后的技能改善了目标属性,并在下游任务中实现了更高的通过率。
cs.AI / 35 / 2608.06894
From Points to Edges: Edge-Conditioned Spectral Operators for Physics-Sensitive PDE Learning
从点到边:面向物理敏感的偏微分方程学习的边条件谱算子
Abstract
Neural operators have become a central tool for solving partial differential equations (PDEs), with spectral operators offering efficient global mixing across spatial locations. However, many PDEs contain physics-sensitive local structures that are critical to the underlying physical behavior. For example, in Darcy flow, local material interfaces are often reflected by sharp changes in the permeability field and can strongly influence the solution. Existing spectral operators primarily adapt modal mixing based on center-point representations, making them insufficiently responsive to such localized structural variations. We propose the Edge-Conditioned Spectral Operator (ESO), a novel spectral operator framework that modulates global spectral mixing using local edge-wise variations. By incorporating the Pairwise-Variation Modal Mixer (PVMM) to inject local edge information into spectral mode selection, ESO preserves the global approximation capability of spectral neural operators while enabling the learned kernel to adapt to physics-sensitive local structures. Furthermore, we introduce a task-adaptive Physics-Aware Reweighting (PAR) that emphasizes physically important regions, identified by taskspecific physical quantities. Across nine PDE benchmarks, ESO consistently achieves state-of-the-art performance. Visual and region-wise analyses further demonstrate that ESO reduces solution errors near coefficient jumps, high-gradient flow structures, and other physically sensitive regions. The code is available at https://github.com/Tanpig-X/ESO.
Chinese Translation
神经算子已成为解决偏微分方程(PDEs)的核心工具,其中谱算子在空间位置之间提供了高效的全局混合。然而,许多PDE包含对物理行为至关重要的物理敏感局部结构。例如,在达西流中,局部材料界面通常通过渗透率场的急剧变化反映出来,并可能强烈影响解的结果。现有的谱算子主要基于中心点表示调整模态混合,因而对这些局部结构变化反应不足。我们提出了边条件谱算子(Edge-Conditioned Spectral Operator, ESO),这是一种新颖的谱算子框架,通过局部边缘变化调节全局谱混合。通过引入成对变化模态混合器(Pairwise-Variation Modal Mixer, PVMM),将局部边缘信息注入谱模选择中,ESO在保持谱神经算子的全局近似能力的同时,使学习到的核能够适应物理敏感的局部结构。此外,我们引入了一种任务自适应的物理感知重加权(Physics-Aware Reweighting, PAR),强调由任务特定物理量识别的物理重要区域。在九个PDE基准测试中,ESO始终实现了最先进的性能。视觉和区域分析进一步表明,ESO在系数跳跃、高梯度流结构及其他物理敏感区域附近减少了解的误差。代码可在 https://github.com/Tanpig-X/ESO 获取。
cs.AI / 36 / 2608.06909
Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework
长时间跨度代理轨迹归因:统一基准与细粒度注释框架
Abstract
Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, and memory. Existing benchmarks primarily evaluate behavioral outcomes but provide limited support for fine-grained attribution analysis. We introduce trajectory attribution and develop a benchmark and annotation framework for this task. The benchmark organizes heterogeneous trajectories under a unified component schema and provides annotations of the primary attribution component, together with attack and execution chains where applicable. Instantiating the benchmark with trajectories from AgentDojo and the Stage and Canary settings of Agent3Sigma yields more than 1,300 annotated trajectories covering task-aligned actions, unsafe actions, and safety refusals. The benchmark defines two evaluation tasks, primary attribution localization and attribution-chain recovery, and provides reference baselines based on incremental trajectory contribution and component-level leave-one-out perturbation. It captures diverse attribution settings, including local and long-range attribution as well as structured attribution chains. Reference baseline results exhibit substantial performance differences across these settings, providing an initial characterization of the benchmark's attribution challenges. Beyond this initial instantiation, we release a reusable annotation skill that enables trajectories generated by new agent models to be standardized, annotated, and evaluated under the same framework. Project resources and future releases are available at https://github.com/chenjing-2024/agent-trajectory-attribution.
Chinese Translation
大型语言模型(LLM)代理越来越多地通过涉及用户指令、工具使用、外部观察和记忆的长时间跨度轨迹进行操作。现有基准主要评估行为结果,但对细粒度归因分析的支持有限。我们引入轨迹归因,并为此任务开发了一个基准和注释框架。该基准在统一的组件架构下组织异构轨迹,并提供主要归因组件的注释,以及适用时的攻击和执行链。通过使用来自AgentDojo和Agent3Sigma的Stage和Canary设置的轨迹实例化基准,生成了超过1300条注释轨迹,涵盖任务对齐的动作、不安全动作和安全拒绝。该基准定义了两个评估任务:主要归因定位和归因链恢复,并提供基于增量轨迹贡献和组件级留一法扰动的参考基线。它捕捉了多样的归因设置,包括局部和长距离归因以及结构化归因链。参考基线结果显示这些设置之间存在显著的性能差异,为基准的归因挑战提供了初步特征描述。除了这一初步实例化外,我们还发布了一个可重用的注释技能,使得新代理模型生成的轨迹能够在相同框架下进行标准化、注释和评估。项目资源和未来发布内容可在 https://github.com/chenjing-2024/agent-trajectory-attribution 获取。
cs.AI / 37 / 2608.06912
Fast LapSum: Exact Differentiable Top-k at Million Scale
快速 LapSum:百万规模下的精确可微分 Top-k
Abstract
The top-$k$ operation is a fundamental building block of modern sparse computation, enabling token routing, expert activation, memory selection, and attention pruning. Yet standard hard top-$k$ blocks gradients, while existing continuous (soft) relaxations remain too costly for large-scale models. We introduce Fast LapSum, an exact-budget soft top-$k$ primitive whose GPU solver runs in linear time after sorting. Unlike prior linear-time methods such as DFTopK, which relax the normalization constraint, Fast LapSum is, to our knowledge, the first method to preserve an exact selection mass of $k$ while remaining fully differentiable end-to-end. Our solver combines a linear-time threshold computation with an analytical vector--Jacobian product, and for extreme scales employs probabilistic bracketing to sort only the uncertain middle band of kernel-noised scores. The resulting overhead is almost negligible: the solver processes $10^6$, $10^7$, and $10^8$ scores in $0.41$, $1.15$, and $5.23$\,ms, respectively. This makes exact soft top-$k$ practical for sparse routing, retrieval, and large-scale optimization. We demonstrate Fast LapSum on two demanding applications operating over millions of coordinates inside the training loop: generating megapixel sparse adversarial examples with an exact soft budget of ${\sim}0.02\%$ of an image's pixels, achieving an order-of-magnitude speedup over state-of-the-art methods, and training a fully differentiable sparse image coder from scratch.
Chinese Translation
Top-$k$ 操作是现代稀疏计算的基本构件,能够实现令牌路由、专家激活、内存选择和注意力剪枝。然而,标准的硬 Top-$k$ 阻断了梯度,而现有的连续(软)松弛方法对于大规模模型仍然过于昂贵。我们提出了 Fast LapSum,这是一种精确预算的软 Top-$k$ 原语,其 GPU 求解器在排序后以线性时间运行。与先前的线性时间方法(如 DFTopK)不同,后者放宽了归一化约束,Fast LapSum 是我们所知的第一种在保持完全可微分的端到端的情况下,保留精确选择质量 $k$ 的方法。我们的求解器结合了线性时间的阈值计算和解析向量-雅可比乘积,并在极端规模下采用概率括号法,仅对带有核噪声的分数的中间不确定区间进行排序。由此产生的开销几乎可以忽略不计:求解器分别在 $0.41$、$1.15$ 和 $5.23$ 毫秒内处理 $10^6$、$10^7$ 和 $10^8$ 个分数。这使得精确的软 Top-$k$ 在稀疏路由、检索和大规模优化中变得实用。我们在两个要求苛刻的应用中展示了 Fast LapSum,这些应用在训练循环中处理数百万个坐标:生成具有约 $0.02\%$ 图像像素的精确软预算的百万像素稀疏对抗样本,较现有最先进的方法实现了数量级的加速,以及从头训练一个完全可微分的稀疏图像编码器。
cs.AI / 38 / 2608.06917
ReGraph: Learning to Generate Recipe Graphs from Food Images
ReGraph:从食品图像中学习生成食谱图
Abstract
Recent Large Multimodal Models (LMMs) have achieved impressive performance in recipe generation from food images.However, cooking is a structured transformation process in which ingredients undergo state changes through ordered actions,while free-form recipe language leaves the corresponding entities, intermediate states, and dependencies largely implicit and entangled.A graph representation makes this procedural knowledge explicit and compositional, providing a structured basis for assessing whether model outputs encode process-level knowledge rather than merely presenting plausible textual descriptions. To address this limitation, we present ReGraph, a large-scale recipe graph dataset that represents ingredients, cooking actions, and tools as entities, uses entity attributes to describe ingredient state changes, and employs typed relations to encode manipulation targets, destinations, and procedural ordering. ReGraph further incorporates explicit Recipe Reasoning Chain-of-Thought (RR-CoT) traces, providing auxiliary supervision for procedural decomposition and structured graph generation. Building on ReGraph, we propose Recipe Graph Learning (RGL), a two-stage framework that enables LMMs to generate a plausible fine-grained cooking workflow from a food image in the form of a structured recipe graph. Under a deterministic, schema-aware matching protocol, our experiments reveal a substantial gap between text-generation quality and recoverable procedural structure: recipes produced by existing approaches achieve competitive text-generation scores yet yield limited reference-aligned entity and relation structure under the ReGraph schema. In contrast, across two representative LMM backbones, RGL consistently improves the generation of cooking entities and procedural relations, while our analysis further shows that fine-grained ingredient-state capture remains the most challenging dimension.
Chinese Translation
近期的大型多模态模型(LMMs)在从食品图像生成食谱方面取得了令人瞩目的成果。然而,烹饪是一个结构化的转化过程,其中原料通过有序的动作经历状态变化,而自由形式的食谱语言在很大程度上使得相应的实体、中间状态和依赖关系隐含且交织在一起。图形表示使这一过程知识变得显性和可组合,为评估模型输出是否编码过程级知识而不仅仅是呈现可信的文本描述提供了结构化基础。为了解决这一局限性,我们提出了ReGraph,一个大规模的食谱图数据集,该数据集将原料、烹饪动作和工具表示为实体,使用实体属性描述原料状态变化,并采用类型化关系编码操作目标、目的地和过程顺序。ReGraph进一步结合了显式的食谱推理链思维(Recipe Reasoning Chain-of-Thought,RR-CoT)轨迹,为过程分解和结构化图生成提供辅助监督。在ReGraph的基础上,我们提出了食谱图学习(Recipe Graph Learning,RGL),这是一个两阶段框架,使LMMs能够从食品图像生成一个可信的细粒度烹饪工作流程,以结构化食谱图的形式呈现。在确定性、模式感知的匹配协议下,我们的实验揭示了文本生成质量与可恢复过程结构之间的显著差距:现有方法生成的食谱在文本生成分数上表现竞争力,但在ReGraph模式下产生的参考对齐实体和关系结构却有限。相比之下,在两个代表性的LMM骨干网络上,RGL始终提高了烹饪实体和过程关系的生成,而我们的分析进一步表明,细粒度的原料状态捕捉仍然是最具挑战性的维度。
cs.AI / 39 / 2608.06922
Deal Me Maybe: The Role of Emotions in Multi-Agent Negotiation
也许给我一个机会:情感在多智能体谈判中的作用
Abstract
Negotiation is a demanding social task for LLM agents, requiring strategic reasoning, persuasion, and interpersonal adaptation. Yet existing benchmarks often treat agents as emotionally neutral, overlooking a key driver of human bargaining behavior. We study how prompt-conditioned emotions affect LLM-based price negotiation. In a controlled framework, buyer and seller agents are independently assigned one of six emotional states and negotiate over 350 real consumer products under two budget conditions. Across 36 emotion-pair settings and five widely used LLMs, we find that emotions strongly shape outcomes. Angry buyers almost never reach agreement (0.39% deal rate), while happy buyers agree most often (28.91%), but obtain worse prices than fearful buyers. Emotion effects are role-dependent: buyer emotion mainly drives acceptance and rejection, whereas seller emotion shapes concession dynamics. These effects influence not only language, but also termination behavior and price trajectories, raising concerns for emotion-conditioned agents in commerce.
Chinese Translation
谈判对于大语言模型(LLM)代理来说是一项要求严格的社会任务,需要战略推理、说服能力和人际适应能力。然而,现有基准往往将代理视为情感中立,忽视了人类讨价还价行为的一个关键驱动因素。我们研究了提示条件下的情感如何影响基于LLM的价格谈判。在一个受控框架中,买方和卖方代理被独立分配六种情感状态之一,并在两种预算条件下对350种真实消费者产品进行谈判。在36种情感对设置和五种广泛使用的LLM中,我们发现情感对结果有显著影响。愤怒的买方几乎从未达成协议(成交率为0.39%),而快乐的买方最常达成协议(28.91%),但获得的价格却低于恐惧的买方。情感的影响依赖于角色:买方情感主要驱动接受和拒绝,而卖方情感则影响让步动态。这些影响不仅影响语言,还影响终止行为和价格轨迹,这对商业中的情感条件代理提出了担忧。
cs.AI / 40 / 2608.06926
TRIBE: Predicting Team Performance via Communication Behavior Ensembles
TRIBE:通过沟通行为集预测团队绩效
Abstract
Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledge. We present TRIBE, a domain independent approach that reveals team behavioral dynamics invisible to traditional performance metrics. We show that communication patterns can categorize teams into performance predictive behavioral tribes, as early as 10% into the task, enabling timely interventions. We test TRIBE on four diverse datasets and demonstrate that communication patterns predict team performance while the prediction strength varies by the degree a task structure allows for behavioral freedom. Our temporal analysis reveals that AI agents significantly alter team behavioral trajectories while human advisors align with natural dynamics, and that teams maintain behavioral flexibility throughout collaboration. Further, we compare TRIBE to Llama and optimize the pipeline, achieving significant speedup with performance improvement.
Chinese Translation
设计能够有效辅助人类团队的自主代理,关键在于理解团队动态,通常不依赖于特定任务的知识。我们提出了TRIBE,这是一种领域独立的方法,揭示了传统绩效指标无法察觉的团队行为动态。我们展示了沟通模式能够在任务开始的10%时将团队分类为具有绩效预测能力的行为群体,从而实现及时干预。我们在四个不同的数据集上测试了TRIBE,并证明沟通模式能够预测团队绩效,而预测的强度则因任务结构对行为自由度的允许程度而异。我们的时间分析揭示,人工智能代理显著改变了团队的行为轨迹,而人类顾问则与自然动态保持一致,团队在协作过程中保持行为灵活性。此外,我们将TRIBE与Llama进行了比较,并优化了流程,实现了显著的加速和性能提升。
cs.AI / 41 / 2608.06931
Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery
科学边缘评估:SEE 迈向真实科学发现的缺失一步
Han, Taolin, Zhang, Yuchen, Wang, Jinghang, Wu, Yun, Chiu, Wai Yuet, Li, Zhaohai, Zhang, Yifei, Wang, Jinxin, Zhou, Yuhao, Zhao, Chen, Li, Jiajia, Li, Jiaxin, Jin, Qile, Sun, Kewei, Wu, Shuang, Zhai, Weiqi, Lv, Renquan, Li, Junchao, Chen, Ruodan, Chen, Qingteng, Yang, Zhibo, Wei, Hu, Qu, Lin, Bai, Shuai, Zhao, Bing
Abstract
Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark of expert-curated questions grounded in peer-reviewed literature and experimental practice in chemistry, biology, and materials science. Evaluation of 19 multimodal large language models (MLLMs) shows that even the best-performing model reaches only 48.7% accuracy. Moreover, general-purpose models outperform science-specialized models on average. In the visual-agent evaluation, the use of tools increases the best accuracy to 52.7%. Tool use can expand the information available to models, but more information does not necessarily lead to reliable scientific reasoning. The key challenge is whether models can manage tool-derived information within the boundaries of the original experimental evidence. Together, these findings reveal that current MLLMs still cannot reliably make justified and evidence-bounded inferences from experimental results, which is an essential capability in real scientific discovery. Bridging this gap requires MLLMs to transition from explaining established scientific concepts to deriving novel and evidence-based insights from experimental data.
Chinese Translation
大型语言模型(LLMs)在科学发现中的参与日益增多,但它们是否能够支持复杂的真实实验室科学仍不清楚。在此,我们引入了科学边缘评估(Science Edge Evaluation, SEE),这是一个基于同行评审文献和化学、生物学及材料科学实验实践的专家策划问题的多模态基准。对19个多模态大型语言模型(MLLMs)的评估显示,即使是表现最好的模型,其准确率也仅为48.7%。此外,通用模型的平均表现优于科学专业模型。在视觉代理评估中,工具的使用将最佳准确率提高至52.7%。工具的使用可以扩展模型可用的信息,但更多的信息并不一定导致可靠的科学推理。关键挑战在于模型是否能够在原始实验证据的边界内管理工具衍生的信息。这些发现共同揭示了当前的MLLMs仍然无法可靠地从实验结果中做出有根据和证据支持的推断,而这在真实的科学发现中是一个至关重要的能力。弥补这一差距需要MLLMs从解释已建立的科学概念转变为从实验数据中推导出新颖且基于证据的见解。
cs.AI / 42 / 2608.06940
Blind to the Pivotal Vote: Aggregate Independence Metrics Miss Where Verification Actually Helps
盲目于关键投票:聚合独立性指标未能识别验证的实际帮助
Abstract
LLM judge panels are a standard evaluation tool, but prior work reports highly correlated panel errors: nine judges provide roughly the effective information of two independent ones, and aggregation closes only a small fraction of the gap. A natural remedy--a signal from a different evidence source, e.g., executing a test suite--produced no distinguishable change in the panel's effective-vote count at scale (-0.04, 95\% CI [-0.10, +0.02]). Aggregate dependence and conditional decision utility are different questions. Elementary majority arithmetic fixes the affected set for single-ballot substitution: only decisions with a one-vote margin can change. The empirical question is whether panel error rates rise and useful substitutions concentrate there. They do: the entire accuracy gain concentrates on these pivotal queries, where it is large (+10.4 to +23.3 percentage points across three headline configurations), and is exactly zero elsewhere. We confirm the pattern across three code benchmarks and four panel sizes (a 9-judge extension and 56 dependent subsampling checks, gain +6.5 to +16.1 percentage points). On HumanEval+/MBPP+, a majority-side replacement rule raises overall accuracy from 82.44\% to 85.62\% while invoking the signal on 16.2\% of queries; signal-only remains stronger at 87.60\%. Thus population-level dependence diagnostics and margin-stratified utility are complementary, and the affected-set characterization yields a call-reduction rule for any specified single-ballot substitution policy.
Chinese Translation
大型语言模型(LLM)评审小组是标准的评估工具,但先前的研究报告了高度相关的小组错误:九名评审提供的有效信息大致相当于两名独立评审,而聚合仅缩小了小部分差距。一种自然的补救措施——来自不同证据源的信号,例如执行测试套件——在大规模下并未显著改变小组的有效投票数(-0.04,95\% 置信区间 [-0.10, +0.02])。聚合依赖性和条件决策效用是不同的问题。基础的多数算术修正了单票替代的受影响集:只有以一票优势的决策可以改变。经验性问题是小组错误率是否上升,以及有用的替代是否集中在此处。结果确实如此:整体准确率的提升集中在这些关键查询上,且幅度较大(在三种主要配置中增加了 +10.4 到 +23.3 个百分点),而在其他地方则完全为零。我们在三个代码基准和四个小组规模(一个9名评审的扩展和56个依赖子采样检查,增益为 +6.5 到 +16.1 个百分点)中确认了这一模式。在 HumanEval+/MBPP+ 上,多数侧替代规则将整体准确率从 82.44\% 提升至 85.62\%,同时在 16.2\\% 的查询中调用信号;仅信号的准确率仍然更强,达到 87.60\\%。因此,群体水平的依赖诊断和边际分层效用是互补的,而受影响集的特征描述为任何指定的单票替代政策提供了一个减少调用的规则。
cs.AI / 43 / 2608.06948
LMM Modality Transfer: A Pre-requisite for Autonomous GIS Agents
LMM模态转移:自主GIS代理的前提条件
Abstract
AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial tasks and workflows. However, most of the research on their spatial capabilities (e.g., spatial reasoning) has focused on the textual modality as input and output. This contrasts with the human approach to GIS workflows, where text and visual modalities are often used together, interchangeably, and in a complementary manner. Thus, to truly achieve an automated GIS analysis pipeline or carry out human-designed GIS workflows, AI models --- Large Multimodal Models (LMMs) in particular --- need to be able to seamlessly transition between image- and text-based modalities that are traditionally used in such workflows. We present a modality transfer task that (1) asks an LMM to first describe an input image of colored squares in a regular grid, and (2) asks a new LMM instance to re-generate an image of the original spatial scene using the textual description output by the former model. This task quantifies the ability of LMMs to transfer spatial information between image and text modalities. Ultimately, by examining the modality transfer capability of LMMs through the lens of spatial information theory, this work highlights a critical bottleneck: achieving strong and robust geospatial understanding in LMMs requires rigorous, multi-modal alignment. Our results indicate that recent LMMs (here from OpenAI) still struggle with modality transfer, when tasked with re-generating an image of a simple spatial grid of color squares.
Chinese Translation
人工智能模型在理解和处理空间信息方面变得越来越娴熟,从而促进了在空间任务和工作流程中的自主问题解决。然而,关于其空间能力(例如,空间推理)的研究大多集中在文本模态作为输入和输出。这与人类在GIS工作流程中的方法形成对比,人类通常将文本和视觉模态结合使用,交替进行,并且互为补充。因此,为了真正实现自动化的GIS分析流程或执行人类设计的GIS工作流程,人工智能模型——特别是大型多模态模型(Large Multimodal Models, LMMs)——需要能够在这些工作流程中传统上使用的基于图像和基于文本的模态之间无缝切换。我们提出了一项模态转移任务,该任务(1)要求LMM首先描述一个由彩色方块组成的规则网格的输入图像,(2) 然后要求一个新的LMM实例使用前一个模型输出的文本描述重新生成原始空间场景的图像。该任务量化了LMM在图像和文本模态之间转移空间信息的能力。最终,通过空间信息理论的视角考察LMM的模态转移能力,这项工作突显了一个关键瓶颈:在LMM中实现强大而稳健的地理空间理解需要严格的多模态对齐。我们的结果表明,最近的LMM(此处来自OpenAI)在模态转移方面仍然存在困难,尤其是在重新生成简单的彩色方块空间网格图像时。
cs.AI / 44 / 2608.06949
Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints
将分诊决策分配给多个代理是否掩盖偏见或有助于捕捉偏见?基于多代理的模拟研究:在审计能力限制下的基于大型语言模型的资源分配
Abstract
Prior benchmarking work has shown that a single large language model (LLM), forced to make life-or-death resource-allocation decisions, exhibits measurable demographic bias. Real deployments, however, rarely use a single agent: they use pipelines, with review steps meant to catch exactly this kind of failure. We study what happens to bias when the same decision is distributed across a role-differentiated multi-agent pipeline (assessment, allocation, independent audit) instead of made and checked by one model alone. Using a synthetic disaster-triage simulator with paired cases that are clinically identical except for one demographic attribute, we run 192 episodes (2,304 resolved case pairs) on GPT-4o-mini comparing a single-agent control condition to a nine-agent pipeline under three independently varied pressure dimensions. We find no measurable difference in how often biased outcomes occur between the two conditions (6.9% vs. 6.1%, p = 0.498). We do find a large and significant effect of audit capacity on whether bias is caught: 30.0% of biased outcomes go entirely undetected, rising to 43.8% when the auditor is overloaded and falling to 18.4% when it is not. Decomposing this effect shows it is driven almost entirely by coverage (whether a case is reviewed at all, which collapses from 100.0% to 65.6% under load, p < 0.001) rather than by degraded judgment on the cases that are reviewed (81.6% vs. 85.7%, p = 1.000, direction reversed). A follow-up experiment shows that reordering the audit queue by estimated risk, rather than first-come-first-served, recovers most of the lost coverage under the same capacity constraint (65.6% to 91.7%, p = 0.028). We discuss the implications for any system that adds independent oversight to an LLM agent pipeline under resource constraints, and report the study's limitations honestly: one model, modest sample sizes, and no adversarial replication.
Chinese Translation
先前的基准测试工作表明,单一的大型语言模型(LLM)在被迫做出生死资源分配决策时,表现出可测量的人口统计偏见。然而,实际部署中很少使用单一代理:它们使用管道,包含旨在捕捉这种失败的审查步骤。我们研究当同一决策在角色差异化的多代理管道(评估、分配、独立审计)中分配时,偏见会发生什么变化,而不是由单一模型做出和检查。我们使用一个合成的灾难分诊模拟器,配对案例在临床上是相同的,除了一个人口统计属性,我们在GPT-4o-mini上运行了192个实验(2,304个解决的案例对),比较了单一代理控制条件与在三个独立变化压力维度下的九代理管道。我们发现两种条件下偏见结果发生的频率没有可测量的差异(6.9% vs. 6.1%,p = 0.498)。我们确实发现审计能力对捕捉偏见的影响显著且巨大:30.0%的偏见结果完全未被发现,当审计员超负荷时这一比例上升至43.8%,而在未超负荷时降至18.4%。分解这一效应显示,它几乎完全由覆盖率驱动(一个案例是否被审查,负载下从100.0%降至65.6%,p < 0.001),而不是由审查案例的判断力下降(81.6% vs. 85.7%,p = 1.000,方向相反)。后续实验表明,通过估计风险重新排序审计队列,而不是先到先服务,可以在相同的能力限制下恢复大部分丢失的覆盖率(65.6%提升至91.7%,p = 0.028)。我们讨论了对任何在资源限制下为LLM代理管道增加独立监督的系统的影响,并诚实地报告了研究的局限性:单一模型、适度样本量和没有对抗性复制。
cs.AI / 45 / 2608.06955
Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation
大型语言模型中的批评赞誉取向:来自电影偏好的证据
Abstract
Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs systematically reproduce evaluative hierarchies remains unclear. Prior research on cultural bias in LLMs suggests competing expectations: models may mirror the popularity signals of internet texts, or may reproduce forms of prestige embedded in critical discourse. We probe this question through a study of film evaluations with eight models from four families (Anthropic, OpenAI, Alibaba, and Mistral), using a 200-film benchmark partitioned into critically acclaimed, commercially successful, and dual-legitimacy (critical acclaim + commercial success) films. Across 20,000 pairwise forced-choice comparisons per model analyzed with Bradley--Terry estimation, we observe a consistent critical acclaim orientation with all models: critically acclaimed yet commercially obscure films are selected over commercially successful yet critically unrecognized ones. This pattern grows with model scale within each family. In addition, nested OLS regression analyses show that evaluative orientation, public visibility, and popular reception distinctly help explain preferences. Adjusting for public visibility reverses the models' preference for dual-legitimacy films over critical acclaim-only films, while additionally accounting for popular reception attenuates much of the disadvantage of films with commercial success only. Finally, evaluative and recommendation-oriented prompt framings produce divergent rankings, suggesting that critical acclaim orientation may manifest indirectly in real-world LLM deployments.
Chinese Translation
大型语言模型(LLMs)是在包含人类对电影、书籍、音乐等的判断表达的语料库上训练的。然而,LLMs是否系统性地再现评估层级仍不清楚。先前关于LLMs中文化偏见的研究提出了相互竞争的预期:模型可能反映互联网文本的流行信号,或者可能再现嵌入在批评话语中的声望形式。我们通过对来自四个家族(Anthropic、OpenAI、Alibaba和Mistral)的八个模型进行电影评估的研究来探讨这个问题,使用一个包含200部电影的基准,这些电影被划分为批评赞誉、商业成功和双重合法性(批评赞誉 + 商业成功)电影。在对每个模型进行的20,000次成对强制选择比较中,使用Bradley--Terry估计法,我们观察到所有模型均表现出一致的批评赞誉取向:批评赞誉但商业不知名的电影被选择,而不是商业成功但未获得批评认可的电影。随着每个家族内模型规模的增加,这一模式愈加明显。此外,嵌套的OLS回归分析表明,评估取向、公众可见性和流行接受度显著帮助解释偏好。调整公众可见性后,模型对双重合法性电影的偏好被逆转,而进一步考虑流行接受度则减轻了仅具商业成功的电影的劣势。最后,评估和推荐导向的提示框架产生了不同的排名,表明批评赞誉取向可能在现实世界的LLM部署中间接体现。
cs.AI / 46 / 2608.06961
CAi Copilot: Reducing Operational Workload in Molecular Design through Intent-Driven Agentic Workflows
CAi Copilot:通过意图驱动的自主工作流程减少分子设计中的操作负担
Abstract
Early-stage molecular design is an iterative process, not just a task of generating molecules. Researchers turn broad goals into design strategies, refine candidates, assess many properties, and gather evidence before synthesis and tests. AI methods can generate molecules, optimize several goals, predict properties, dock compounds, and account for synthesis. Yet these functions are spread across specialized tools. Experts must still coordinate each step, judge interim results, and integrate evidence. The central challenge is thus to turn research intent into adaptive, traceable runs grounded in scientific tools. We cast this challenge as intent-to-evidence molecular design workflow execution and present CAi Copilot, an expert-oriented agent with three linked layers. The Research Interface Layer turns intent into an executable plan. The Agent Reasoning Layer uses interim results to guide each run. The Execution Substrate supplies molecular tools, metrics, reusable utilities, and backend services. Across 45 tasks, CAi achieves the strongest overall performance, with an outcome score of 84.59, exceeding the next-best result by 18.07 points. Additional benchmarks test how CAi coordinates generation, screening, and multi-criteria evaluation, while exposing limits in long-horizon execution. These results show that CAi turns broad molecular-design intent into transparent, traceable workflows that connect interim decisions to candidate-level evidence.
Chinese Translation
早期分子设计是一个迭代过程,不仅仅是生成分子的任务。研究人员将广泛的目标转化为设计策略,精炼候选分子,评估多种属性,并在合成和测试之前收集证据。人工智能方法可以生成分子,优化多个目标,预测属性,进行分子对接,并考虑合成。然而,这些功能分散在各个专业工具中。专家仍需协调每个步骤,判断中间结果,并整合证据。因此,中心挑战在于将研究意图转化为基于科学工具的自适应、可追溯的执行流程。我们将这一挑战视为意图到证据的分子设计工作流程执行,并提出CAi Copilot,一个面向专家的代理,具有三个相互关联的层次。研究接口层将意图转化为可执行计划。代理推理层利用中间结果指导每次执行。执行基础层提供分子工具、指标、可重用工具和后端服务。在45个任务中,CAi实现了最强的整体性能,结果得分为84.59,超过第二名结果18.07分。额外的基准测试考察了CAi如何协调生成、筛选和多标准评估,同时揭示了长期执行的局限性。这些结果表明,CAi将广泛的分子设计意图转化为透明、可追溯的工作流程,将中间决策与候选级证据连接起来。
cs.AI / 47 / 2608.06963
Learning in Deep Networks under Dale's Constraint
在戴尔约束下的深度网络学习
Abstract
Biologically plausible learning models aim to explain how neural circuits can implement effective learning under the constraints of real neurons. Although significant progress has been made, a major remaining challenge is that existing models often allow neurons or synapses to represent mixed-sign values, both positive and negative, in violation of a basic aspect of cortical circuitry -- Dale's constraint: biological neurons are either excitatory or inhibitory, but not both, and synapses cannot change sign. In this work, we address this discrepancy by introducing a biologically motivated neural architecture in which both neural activations and learning signals are represented by non-negative activity, and synapses have fixed sign, while still supporting backpropagation-like learning. Our approach uses two complementary interacting non-negative channels to represent positive and negative contributions, inspired by evidence of on-off representations in the brain. These channels are implemented through a simple neural circuit motif, which is repeated throughout the network in both bottom-up and top-down pathways. Combined with a local Hebbian learning rule, the resulting model propagates learning signals and updates weights using only local interactions between neurons. We show theoretically that our learning scheme can exactly recover the backpropagation update despite relying solely on non-negative error signals. Empirically, beyond satisfying stronger biological constraints, the on-off architecture learns efficient representations, yielding substantial gains over comparable vanilla networks on the Tiny ImageNet benchmark. These results demonstrate that effective learning can emerge from biologically plausible mechanisms without requiring mixed-sign signals, providing a step toward more realistic models of neural computation.
Chinese Translation
生物学上合理的学习模型旨在解释神经电路如何在真实神经元的约束下实现有效学习。尽管取得了显著进展,但一个主要的挑战是现有模型通常允许神经元或突触表示混合符号值,即正值和负值,这违反了皮层电路的基本特征——戴尔约束:生物神经元要么是兴奋性的,要么是抑制性的,而不能同时具备,突触也不能改变符号。在本研究中,我们通过引入一种生物学驱动的神经架构来解决这一不一致性,其中神经激活和学习信号均由非负活动表示,突触具有固定符号,同时仍支持类似反向传播的学习。我们的方法使用两个互补的非负通道来表示正向和负向贡献,灵感来源于大脑中开关表示的证据。这些通道通过一个简单的神经电路模式实现,该模式在网络的自下而上和自上而下路径中重复。结合局部赫布学习规则,所得到的模型通过仅依赖神经元之间的局部交互传播学习信号并更新权重。我们理论上证明了我们的学习方案可以准确恢复反向传播更新,尽管仅依赖于非负误差信号。在经验上,除了满足更强的生物约束外,这种开关架构还学习到高效的表示,在Tiny ImageNet基准测试中相较于可比的普通网络取得了显著的提升。这些结果表明,有效学习可以从生物学上合理的机制中涌现,而不需要混合符号信号,为更现实的神经计算模型迈出了重要一步。
cs.AI / 48 / 2608.06969
Finding Usable Weight Mechanisms with Tiled SVD
使用平铺奇异值分解寻找可用的权重机制
Abstract
The dominant approach to mechanistic interpretability trains proxy dictionaries such as sparse autoencoders and labels features from max-activating text. The best such atlases identify con- cepts, but that identity lives in the learned dictionary rather than in the network weights them- selves. We propose extracting mechanism mounts directly from linear sites by column-tiled SVD: each mount is a triple (v,u,{\sigma}) read as trigger, write, and strength. Identity is the weight rule. We evaluate mounts with a pre-registered suite judged on full-write energy lift rather than tile-local lift. On Gemma-2-2B with WikiText-2 (16,384-token subsample), all seven linear maps are scored: residual writes (mlp.down, attn.o) receive full A/B/C with steer after post-sublayer RMSNorm and pass 52/52 site-layers; other maps receive A/B only (mlp.gate/attn.q/attn.k/effective mlp.up/attn.v 26/26 each). Aggregate: 182/182 GO. We release library code, the corpus builder, the experiment entrypoint, and unit tests.
Chinese Translation
机制可解释性的主流方法是训练代理字典,例如稀疏自编码器,并从最大激活文本中标记特征。这些图谱中最好的能够识别概念,但这种身份存在于学习的字典中,而不是网络权重本身。我们提出通过列平铺奇异值分解(column-tiled SVD)直接从线性位置提取机制挂载:每个挂载是一个三元组 (v, u, { ext{σ}}),分别表示触发、写入和强度。身份是权重规则。我们使用预注册的评估套件来评估挂载,判断标准是全写能量提升,而不是局部平铺提升。在使用 WikiText-2(16,384 令牌子样本)的 Gemma-2-2B 上,所有七个线性映射均被评分:残差写入(mlp.down, attn.o)在后置子层 RMSNorm 后获得满分 A/B/C,并通过 52/52 个站点层;其他映射仅获得 A/B(mlp.gate/attn.q/attn.k/effective mlp.up/attn.v 各 26/26)。总体:182/182 GO。我们发布了库代码、语料库构建器、实验入口点和单元测试。
cs.AI / 49 / 2608.07007
FedLBW: A Loss-Based Weighting Strategy for Federated Learning on Non-IID Data in Wireless Networks
FedLBW:一种基于损失的加权策略用于无线网络中的非独立同分布数据的联邦学习
Abstract
Federated Learning (FL) enables collaborative machine learning (ML) across distributed clients while preserving privacy. However, efficient model convergence in FL remains challenging, especially in wireless networks where non-independent and identically distributed (non-IID) data and frequent client dropouts are common. Traditional FL algorithms, such as FedAvg, rely solely on dataset size to weight client updates. This introduces biases towards clients with larger datasets and makes the process sensitive to non-IID data, outliers, and client dropouts. To address these challenges, we propose Federated Learning with Loss-Based Weighting (FedLBW), a novel aggregation method that assigns each client's update a weight proportional to the inverse of its validation loss, computed using a small proxy dataset on the server, rather than its dataset size. This ensures that lower-loss models exert greater influence during aggregation, prioritizing the most reliable updates and boosting overall performance. Through extensive experiments across multiple datasets, including FashionMNIST (CNN), CIFAR-10 (ResNet-18), and CIFAR-100 (ResNet-34), we demonstrate that FedLBW achieves higher accuracy and faster convergence compared to baseline algorithms such as FedAvg, FedAvgM, FedProx, FedNova, FedLAW and FedDkw, with notable improvements of up to 7.6 % higher accuracy on CIFAR-10 in extreme non-IID cases. Moreover, FedLBW showcases exceptional resilience to increasing dropout probabilities, consistently maintaining significantly higher accuracy even in challenging conditions. These results establish FedLBW as an effective and resilient solution for FL in wireless network environments, offering marked improvements in model accuracy, convergence speed, and robustness to non-IID data and client dropouts.
Chinese Translation
联邦学习(Federated Learning, FL)使得分布式客户端之间能够进行协作机器学习(Machine Learning, ML),同时保护隐私。然而,在无线网络中,非独立同分布(non-IID)数据和频繁的客户端掉线使得FL中的模型收敛效率仍然面临挑战。传统的FL算法,如FedAvg,仅依赖数据集大小来加权客户端更新。这导致了对数据集较大客户端的偏见,并使得该过程对非-IID数据、离群点和客户端掉线敏感。为了解决这些挑战,我们提出了基于损失的加权联邦学习(Federated Learning with Loss-Based Weighting, FedLBW),这是一种新颖的聚合方法,它将每个客户端的更新权重与其验证损失的倒数成正比,该损失是通过服务器上的小型代理数据集计算得出的,而不是依赖于数据集大小。这确保了低损失模型在聚合过程中具有更大的影响力,优先考虑最可靠的更新,从而提升整体性能。通过在多个数据集上进行广泛实验,包括FashionMNIST(CNN)、CIFAR-10(ResNet-18)和CIFAR-100(ResNet-34),我们证明FedLBW在准确性和收敛速度上优于基线算法,如FedAvg、FedAvgM、FedProx、FedNova、FedLAW和FedDkw,在极端非-IID情况下,CIFAR-10的准确性提高了高达7.6%。此外,FedLBW在增加掉线概率的情况下表现出卓越的韧性,即使在困难条件下也能持续保持显著更高的准确性。这些结果确立了FedLBW作为无线网络环境中FL的有效且具有韧性的解决方案,在模型准确性、收敛速度和对非-IID数据及客户端掉线的鲁棒性方面提供了显著改善。
cs.AI / 50 / 2608.07019
ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization
ReQuant:用于后训练量化的固定网格离散优化
Abstract
Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models. Existing PTQ methods typically obtain an initial quantized model through heuristic rules or greedy optimization, and once quantization is completed the resulting integer assignments are usually treated as final. This observation motivates a complementary optimization stage within PTQ that keeps quantized weights improvable after an executable quantized model has been produced, while preserving the quantized format. We introduce ReQuant, a backpropagation-free fixed-grid refinement procedure for this stage. Agnostic to the PTQ initializer, ReQuant takes an existing quantized model as a feasible starting point and iteratively revisits its discrete weight assignments on the fixed quantization grid. Accepted updates strictly reduce the mean squared reconstruction error and remain on the original grid. In this way, ReQuant turns the initially fixed PTQ output into an iteratively optimizable discrete solution and serves as a plug-and-play post-processing stage for existing PTQ pipelines. Experiments across diverse model families, bit-widths, and downstream tasks show that ReQuant consistently improves quantized models from heterogeneous PTQ initializers, with especially large gains on simple initializers and lower bit-widths. Notably, ReQuant can refine a simple round-to-nearest initialization across multiple sweeps until it approaches or surpasses GPTAQ under the same quantization format. These results establish ReQuant as a practical complementary stage for further improving existing PTQ pipelines.
Chinese Translation
后训练量化(PTQ)广泛用于降低大型语言模型的内存和计算成本。现有的PTQ方法通常通过启发式规则或贪婪优化获得初始量化模型,一旦量化完成,所得的整数分配通常被视为最终结果。这一观察促使我们在PTQ中引入一个补充优化阶段,该阶段在生成可执行量化模型后,仍然保持量化权重的可优化性,同时保留量化格式。我们提出了ReQuant,一种无反向传播的固定网格优化程序,用于这一阶段。ReQuant与PTQ初始化器无关,采用现有的量化模型作为可行的起点,并在固定量化网格上迭代地重新审视其离散权重分配。接受的更新严格减少均方重建误差,并保持在原始网格上。通过这种方式,ReQuant将最初固定的PTQ输出转变为可迭代优化的离散解决方案,并作为现有PTQ流程的即插即用后处理阶段。在不同模型系列、比特宽度和下游任务上的实验表明,ReQuant始终能改善来自异构PTQ初始化器的量化模型,尤其在简单初始化器和较低比特宽度下获得显著提升。值得注意的是,ReQuant能够在多个迭代中优化简单的四舍五入初始化,直到其接近或超过在相同量化格式下的GPTAQ。这些结果确立了ReQuant作为进一步改善现有PTQ流程的实用补充阶段。
cs.AI / 51 / 2608.07033
ZIPBrain: Can EEG Foundation Models Be Faster, Locally Deployable, but Accurate?
ZIPBrain:能否让脑电图基础模型更快、可本地部署且保持准确性?
Abstract
This work investigates whether Electroencephalograph (EEG) foundation models (EFMs) can be made faster and locally deployable without sacrificing accuracy. EEG foundation models are a major trend, offering strong general-purpose representations. However, their computational burden grows quadratically with input length, hindering deployment on resource-constrained scenario, particularly for real-time clinical monitoring. EEG's low SNR further suggests many of these tokens are redundant and compressible with little accuracy cost. We propose ZIPBrain, a novel redundancy-aware EEG token pooling module that leverages this low-SNR characteristic to reduce token count. Given a token sequence, ZIPBrain partitions tokens into redundant and unique groups, then merges each redundant token with its most similar counterpart in the unique group. Furthermore, ZIPBrain serves as a training-free, plug-and-play module that seamlessly integrates into standard Transformer encoders with negligible computational overhead. Extensive experiments across multiple EEG foundation models show ZIPBrain's strong versatility, achieving 1.3%-10.5% average improvement over baselines, while reducing wall-clock inference time by 32.7% (up to 41.8% with CUDA Graph) compared to the original EEG foundation models.
Chinese Translation
本研究探讨了脑电图(EEG)基础模型(EFMs)是否可以在不牺牲准确性的情况下实现更快和本地部署。EEG基础模型是一个主要趋势,提供强大的通用表示。然而,它们的计算负担随着输入长度的增加呈平方增长,阻碍了在资源受限场景下的部署,特别是在实时临床监测中。EEG的低信噪比(SNR)进一步表明,许多这些标记是冗余的,并且可以在几乎不影响准确性的情况下进行压缩。我们提出了ZIPBrain,这是一种新颖的冗余感知EEG标记池化模块,利用这种低SNR特性来减少标记数量。给定一个标记序列,ZIPBrain将标记分为冗余组和独特组,然后将每个冗余标记与独特组中最相似的对应标记合并。此外,ZIPBrain作为一个无需训练的即插即用模块,能够无缝集成到标准Transformer编码器中,几乎没有计算开销。针对多个EEG基础模型的广泛实验表明,ZIPBrain具有强大的通用性,相较于基线模型实现了1.3%-10.5%的平均提升,同时将推理时间减少了32.7%(在CUDA Graph下可达41.8%),与原始EEG基础模型相比。
cs.AI / 52 / 2608.07040
Not All Problems Are Best Modeled as MILP: A DSL-Centric Framework for Flexible and Accurate Optimization Modeling
并非所有问题都适合用混合整数线性规划建模:一个以领域特定语言为中心的灵活准确优化建模框架
Abstract
Solving combinatorial optimization problems (COPs) requires not only efficient algorithms but also carefully crafted formulations. While recent works have leveraged LLMs to automate optimization modeling, current frameworks predominantly rely on a rigid mixed-integer linear programming (MILP) paradigm. In this paper, we argue that not all problems are best modeled as MILP, as forcing complex domains into linear constraints can induce prohibitive modeling complexity and severely restrict solver flexibility. To address this, we propose OptiDSL, a framework that shifts the focus from rigid MILP formulations to domain-specific language (DSL) representations. By utilizing LLMs to map natural language onto standardized, domain-accepted structures, OptiDSL decouples problem formulation from execution. This paradigm enables seamless integration with a diverse library of specialized solvers, ranging from traditional heuristics to modern learning-based methods. Experimental results on the comprehensive benchmark of 44 COP types show that OptiDSL significantly surpasses MILP-based pipelines, yielding a 51.66% gain in formulation accuracy and a 91.71% decrease in modeling time. Notably, it also outperforms MILP-based pipelines on the existing benchmark, achieving a 23.09% higher formulation accuracy. Our code is available at https://anonymous.4open.science/r/OptiDSL.
Chinese Translation
解决组合优化问题(COPs)不仅需要高效的算法,还需要精心设计的模型。尽管近期的研究利用大型语言模型(LLMs)来自动化优化建模,但当前的框架主要依赖于刚性的混合整数线性规划(MILP)范式。本文论证并非所有问题都适合用MILP建模,因为将复杂领域强行转化为线性约束可能导致建模复杂度过高,并严重限制求解器的灵活性。为了解决这一问题,我们提出了OptiDSL,一个将重点从刚性的MILP模型转向领域特定语言(DSL)表示的框架。通过利用LLMs将自然语言映射到标准化的、领域认可的结构,OptiDSL将问题建模与执行解耦。这一范式使得与多样化的专用求解器库(从传统启发式到现代基于学习的方法)无缝集成成为可能。在对44种COP类型的综合基准测试中的实验结果表明,OptiDSL显著超越了基于MILP的管道,建模准确性提高了51.66%,建模时间减少了91.71%。值得注意的是,它在现有基准上也优于基于MILP的管道,建模准确性提高了23.09%。我们的代码可在 https://anonymous.4open.science/r/OptiDSL 获取。
cs.AI / 53 / 2608.07053
Unsupervised Adaptation of PDE Foundation Models
无监督的偏微分方程基础模型适应
Abstract
Pretrained partial differential equation (PDE) foundation models can generalize across different equations, but adapting them to unseen PDE systems typically requires dense solution data, which is often expensive or unavailable. To address this limitation, we propose an unsupervised PDE-based finetuning framework that eliminates the need for ground-truth solutions. We first pretrain a neighborhood attention Transformer on diverse time-dependent PDEs spanning varying spatial scales, yielding transferable representations across heterogeneous equations. In the adaptation stage, we construct a physics-based objective using the PDE residual and boundary conditions, and finetune the model on unseen equations via low-rank adaptation (LoRA). To address the uneven learning across physical quantities in standard LoRA, we introduce NSLoRA, a Newton-Schulz orthogonalized variant that rebalances adaptation. Our method achieves performance comparable to supervised LoRA finetuning without requiring any ground-truth solutions, while consistently outperforming competitive neural operator baselines and recent PDE foundation models across heterogeneous PDE benchmarks spanning multiple spatial dimensions.
Chinese Translation
预训练的偏微分方程(PDE)基础模型能够在不同方程之间进行泛化,但将其适应于未见过的PDE系统通常需要密集的解数据,而这些数据往往昂贵或不可用。为了解决这一限制,我们提出了一种无监督的基于PDE的微调框架,消除了对真实解的需求。我们首先在多样的时变PDE上预训练一个邻域注意力变换器,涵盖不同的空间尺度,从而产生跨异构方程的可迁移表示。在适应阶段,我们使用PDE残差和边界条件构建基于物理的目标,并通过低秩适应(LoRA)在未见方程上微调模型。为了应对标准LoRA中物理量学习的不均匀性,我们引入了NSLoRA,一种牛顿-舒尔茨正交化变体,重新平衡适应。我们的方法在不需要任何真实解的情况下,实现了与监督LoRA微调相当的性能,同时在跨越多个空间维度的异构PDE基准测试中,始终优于竞争性的神经算子基线和近期的PDE基础模型。
cs.AI / 54 / 2608.07056
BONSAI: Evolvability-Guided Tree Search over Skills
BONSAI:以可进化性为导向的技能树搜索
Abstract
A skill is a naturallanguage document that steers a frozen agent whose weights cannot be updated so any capability the agent lacks must be supplied in prose Optimising a skill is therefore optimising text against a score and the standard recipe which keeps any edit that raises a heldout score is blind in a specific way a single score cannot tell a document perched on a narrow overfit spike from one resting on a broad plateau even though only the second can still be improved We introduce BONSAI a novel skilloptimisation framework that steers instead by evolvability the capacity of a region of documentspace to keep producing viable variation under further mutation a property biology treats as separate from present fitness BONSAI grows skills as a MonteCarlo search tree in which every child document is a mutation of its parent and descends it under an upperconfidence selection rule whose exploitation term blends a skills own fitness with the fitness of its mutational neighbourhood Because every child is a mutation the mean score recorded beneath a node estimates that neighbourhoods evolvability at no extra cost so the rule concentrates budget on regions that keep improving while its exploration term keeps a currently weak branch in contention BONSAI ships the single bestscoring document it finds at no cost beyond the acceptifbetter loop it replaces With a frozen 30B agent and averaged over three benchmarks BONSAI lifts heldout accuracy over the skillfree agent by 2313 points and improves on two budgetmatched baselines GEPA and SkillOpt by 387 and 397 points respectively
Chinese Translation
技能是一个自然语言文档,它引导一个被冻结的代理,该代理的权重无法更新,因此代理所缺乏的任何能力必须以散文的形式提供。因此,优化技能就是在一个评分标准下优化文本,而标准的做法是保持任何提高保留评分的编辑,这在某种特定方式上是盲目的:单一的评分无法区分一个停留在狭窄过拟合峰值上的文档和一个停留在广阔平坦区域上的文档,尽管只有第二个文档仍然可以被改进。我们引入了BONSAI,一个新颖的技能优化框架,它通过可进化性来引导,定义为文档空间区域在进一步变异下持续产生可行变体的能力,这是生物学上被视为与当前适应性分开的属性。BONSAI将技能作为一个蒙特卡洛搜索树进行扩展,其中每个子文档都是其父文档的变异,并根据一个上置信度选择规则进行下降,该规则的利用项将技能自身的适应性与其变异邻域的适应性相结合。由于每个子文档都是变异,因此在节点下记录的平均评分在没有额外成本的情况下估计了该邻域的可进化性,因此该规则将预算集中在持续改进的区域,而其探索项则保持当前较弱的分支处于竞争中。BONSAI以零成本替代了它所找到的单个最佳评分文档,超出了接受更好的循环。使用一个冻结的30B代理,并在三个基准上进行平均,BONSAI使得保留准确率比无技能代理提高了2313个点,并在两个预算匹配的基线GEPA和SkillOpt上分别提高了387和397个点。
cs.AI / 55 / 2608.07066
PTQ4SNN: Membrane-Aware Post-Training Quantization for Spiking Neural Networks
PTQ4SNN:面向膜的后训练量化框架用于脉冲神经网络
Abstract
Spiking neural networks (SNNs) enable sparse and event-driven computation, but their low-bit deployment remains incomplete because recurrent membrane states are commonly retained in floating point even after weight quantization. Quantizing these states is challenging because their distributions differ across channels and from the preceding weights, while small perturbations near the firing threshold may alter spike decisions and accumulate over time. We propose PTQ4SNN, a membrane-aware post-training quantization framework that jointly quantizes weights and recurrent membrane states using only a small calibration set. First, a channel-wise Unified Scale Bridge constrains the membrane scale as s_mem,c = s_w,c * 2^k_c, adapting to membrane distributions while enabling shift-compatible scale conversion. Second, Mixed-Precision Bit Allocation assigns 2/4/8-bit precision to membrane channels according to firing activity and quantization sensitivity under an average-bit budget. The framework operates on reusable projection-LIF pairs and supports both convolutional SNNs and spike-driven Transformers without backbone retraining. Experiments on static and event-based classification and semantic segmentation show that PTQ4SNN effectively preserves model accuracy under W4 quantization and approximately 4-bit membrane precision.
Chinese Translation
脉冲神经网络(SNNs)能够实现稀疏和事件驱动的计算,但由于在权重量化后,递归膜状态通常仍以浮点数形式保留,因此其低比特部署仍不完善。量化这些状态具有挑战性,因为它们的分布在通道之间以及与前一层权重的分布不同,而在发放阈值附近的小扰动可能会改变脉冲决策并随着时间的推移累积。我们提出了PTQ4SNN,这是一种面向膜的后训练量化框架,它仅使用少量校准集联合量化权重和递归膜状态。首先,通道级统一尺度桥(Unified Scale Bridge)将膜尺度限制为 s_mem,c = s_w,c * 2^k_c,适应膜分布,同时实现与位移兼容的尺度转换。其次,混合精度位分配(Mixed-Precision Bit Allocation)根据发放活动和量化敏感性在平均比特预算下为膜通道分配2/4/8位精度。该框架在可重用的投影-脉冲发放整合(projection-LIF)对上运行,支持卷积SNN和脉冲驱动的变换器(Transformers),无需主干网络的重新训练。在静态和基于事件的分类及语义分割实验中,PTQ4SNN有效地在W4量化和约4位膜精度下保持了模型的准确性。
cs.AI / 56 / 2608.07067
DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding
DocMemo:通过概率记忆引导检索实现多模态文档理解的动态证据发现
Abstract
Long-document understanding requires locating sparse and heterogeneous evidence across hundreds of pages, yet existing systems remain limited by static retrieval and fragile cross-round memory. Mainstream single-round methods commit to a fixed top-$k$ page set at the outset and struggle to recover from early retrieval errors; recent iterative approaches allow multi-round evidence acquisition, but they do not investigate the propagation mechanism of cross-round states, making it difficult to track the dynamic changes in page relevance. To address these limitations, we propose DocMemo, a memory-guided framework that formulates long-document reasoning as dynamic evidence exploration. DocMemo maintains a tri-level retrieval state consisting of Document Schema Memory, Page Belief Memory, and Question Episodic Memory, which respectively capture structural priors, dynamic relevance estimation, and query-specific reasoning trajectories. During reasoning, DocMemo continuously refines cross-round page selection through Bayesian page belief updating with Thompson sampling, spatial proximity propagation, and structure-aware adaptive-granularity evidence access, while supplementing page-level evidence with fine-grained visual regions. Experiments on 3 benchmarks show that DocMemo achieves state-of-the-art performance and validate the efficacy of structured memory and dynamic page belief updating. Code is available at https://github.com/Harrygof/DocMemo.
Chinese Translation
长文档理解需要在数百页中定位稀疏且异质的证据,但现有系统受到静态检索和脆弱的跨轮记忆的限制。主流的单轮方法在开始时承诺固定的前$k$页集合,并在早期检索错误后难以恢复;最近的迭代方法允许多轮证据获取,但未探讨跨轮状态的传播机制,使得追踪页面相关性的动态变化变得困难。为了解决这些限制,我们提出了DocMemo,一个将长文档推理形式化为动态证据探索的记忆引导框架。DocMemo维护一个由文档模式记忆、页面信念记忆和问题情节记忆组成的三级检索状态,分别捕捉结构先验、动态相关性估计和查询特定的推理轨迹。在推理过程中,DocMemo通过贝叶斯页面信念更新、汤普森采样、空间邻近传播和结构感知的自适应粒度证据访问持续优化跨轮页面选择,同时用细粒度的视觉区域补充页面级证据。在3个基准上的实验表明,DocMemo实现了最先进的性能,并验证了结构化记忆和动态页面信念更新的有效性。代码可在 https://github.com/Harrygof/DocMemo 获取。
cs.AI / 57 / 2608.07068
MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents
MemOPD:通过记忆状态对齐进行的在线蒸馏以支持长时间跨度的智能体
Abstract
Long-horizon agents accumulate growing contexts during interaction, impairing performance and stability. Compact memory mitigates this problem by compressing and rewriting the history retained between model invocations. Learning what to retain typically relies on proximal policy optimization (PPO) with final task rewards, but sparse rewards provide little guidance for individual memory updates. This limitation motivates on-policy distillation (OPD), which supplies dense teacher supervision on student rollouts. For such supervision to be valid, the teacher must evaluate each sampled action under the same state in which it was generated. However, the context rewriting performed during memory compression can break this alignment. When sampled responses are retained and re-encoded for later invocations, flattening the interaction into a persistent history may cause the teacher to score the action under a state that the student never visited during rollout. The action therefore remains on-policy by provenance, but not necessarily by state. We therefore propose Memory-Aligned On-Policy Distillation (MemOPD). MemOPD records the inputs and sampled outputs of each model invocation, restores its original token positions and causal visibility, and packs the reconstructed invocations for efficient teacher scoring. The teacher provides full-vocabulary supervision at the sampled action positions, while PPO preserves the final task objective. Experiments verify state alignment across several context updates and show that it improves F1 by 7.0% over persistent-history teacher scoring in a matched control. Overall, MemOPD-3B improves F1 over PPO by up to 416.2%, while packing yields up to a 1.63x speedup in actor computation during training. The code for this work is publicly available at: https://github.com/TPssp/MemOPD.
Chinese Translation
长时间跨度的智能体在交互过程中积累不断增长的上下文,这会影响其性能和稳定性。紧凑的记忆通过压缩和重写模型调用之间保留的历史来缓解这一问题。学习保留哪些信息通常依赖于最终任务奖励的近端策略优化(PPO),但稀疏的奖励对单个记忆更新提供的指导有限。这一局限性促使了在线策略蒸馏(OPD)的提出,它在学生的回放中提供了密集的教师监督。为了使这种监督有效,教师必须在生成每个采样动作时评估相同状态下的动作。然而,在记忆压缩过程中进行的上下文重写可能会破坏这种对齐。当采样的响应被保留并重新编码以供后续调用时,将交互扁平化为持久的历史可能导致教师在学生从未访问过的状态下对动作进行评分。因此,该动作在来源上仍然是在线的,但不一定在状态上也是如此。因此,我们提出了记忆对齐在线蒸馏(MemOPD)。MemOPD记录每次模型调用的输入和采样输出,恢复其原始令牌位置和因果可见性,并将重建的调用打包以便于教师评分。教师在采样动作位置提供全词汇监督,而PPO则保持最终任务目标。实验验证了多个上下文更新中的状态对齐,并显示其在匹配对照组中相较于持久历史教师评分提高了7.0%的F1分数。总体而言,MemOPD-3B在F1分数上相较于PPO提高了高达416.2%,而打包在训练期间的演员计算中实现了最高1.63倍的加速。该工作的代码已公开发布在:https://github.com/TPssp/MemOPD。
cs.AI / 58 / 2608.07077
Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking
变压器在使用其新兴世界模型时遇到困难:重访汉诺塔与思维的幻觉
Abstract
The Tower of Hanoi is a simple planning puzzle that in prior work has proven challenging for large reasoning models (LRMs). Current models solve the standard formulation of the puzzle, but still struggle with the flat-to-flat variant (where initial and goal states are not restricted to have all rings on a single peg). This paper presents an in-depth study of how both small, in-house Transformers and large, third-party LRMs solve this task. To understand the failures mechanistically, we first train small Transformers from scratch on precomputed solution traces. Using a variety of interpretability techniques, we show that these Transformers develop an emergent world model: a linearly decodable, geometrically faithful representation of the puzzle's state space (the Sierpinski triangle), that is causally involved in solving the puzzles. Second, we return to the large LLMs and apply our techniques to two frontier reasoning models, Qwen3.6-27B and DeepSeek-R1-Distill-Qwen-32B, that attempt to solve the task through extended chain-of-thought. Surprisingly, we find that both models encode the Sierpinski world model near-perfectly at the end of the prompt, and yet fail at the majority of tasks when there are more than 3 rings. We locate the source of this failure in the decaying representation of the world model. We probe for the representation at different stages during planning, and establish causality by showing that performance can be improved by injecting the prompt-time representation at inference. The failure of the models is thus one of maintenance of the required representations, not their absence, and performance is at least partially recoverable. These results thus reframe the reported collapse in performance from prior work: current Large Reasoning Models build a world model, and then lose it.
Chinese Translation
汉诺塔是一个简单的规划难题,在之前的研究中,对于大型推理模型(LRMs)来说证明是具有挑战性的。目前的模型能够解决该难题的标准形式,但在平面到平面变体(即初始状态和目标状态不限制所有环都在同一柱子上)中仍然存在困难。本文深入研究了小型自研变压器和大型第三方LRMs如何解决这一任务。为了从机制上理解这些失败,我们首先从头开始训练小型变压器,使用预计算的解决方案轨迹。通过多种可解释性技术,我们展示了这些变压器发展出一种新兴的世界模型:一种线性可解、几何上忠实于难题状态空间(即谢尔宾斯基三角形)的表示,并且在解决难题中起到因果作用。其次,我们回到大型LLMs,并将我们的技术应用于两个前沿推理模型,Qwen3.6-27B和DeepSeek-R1-Distill-Qwen-32B,这些模型试图通过扩展思维链来解决该任务。令人惊讶的是,我们发现这两个模型在提示的末尾几乎完美地编码了谢尔宾斯基世界模型,但在环数超过3个时却在大多数任务中失败。我们将这一失败的根源定位于世界模型的衰退表示。我们在规划的不同阶段探测这种表示,并通过展示在推理时注入提示时间表示可以改善性能来建立因果关系。因此,模型的失败在于所需表示的维护,而非其缺失,且性能至少部分可恢复。这些结果重新框定了之前研究中报告的性能崩溃:当前的大型推理模型构建了一个世界模型,然后又失去了它。
cs.AI / 59 / 2608.07107
MemWM: Memory-Augmented Text-Based World Model
MemWM:增强记忆的基于文本的世界模型
Abstract
World models are increasingly used to support planning in agents by predicting how environment states evolve in response to agent actions. Yet fluent next-state predictions can still omit task-critical facts, corrupt product attributes, or apply incorrect transition rules. To address such systematic prediction errors, we introduce MemWM, a memory-augmented text-based world model. MemWM uses world memory, a curated memory bank of transition rules, state caches, and hard-to-predict facts, to condition next-state imagination. We evaluate factual state preservation with Structured State Fidelity (SSF), which scores predicted states through benchmark-specific facts and fields. Compared with SFT, memory-augmented training improves SSF by up to 206.3%. In the full planning setting, we keep the policy model frozen and provide policy-side world skill: retrieved task-level skills and step-wise corrective guidance for action selection. Across ALFWorld, WebShop, and ScienceWorld, memory-augmented agents improve downstream success over an SFT-trained world-model agent, with up to a 65.4% relative gain. Sensitivity analyses further show that retrieved memory improves task success and efficiency under different memory and action-budget settings.
Chinese Translation
世界模型越来越多地用于支持智能体的规划,通过预测环境状态如何响应智能体的动作而演变。然而,流畅的下一个状态预测仍可能遗漏任务关键事实、损坏产品属性或应用不正确的转移规则。为了解决这些系统性的预测错误,我们提出了MemWM,一种增强记忆的基于文本的世界模型。MemWM利用世界记忆,一个经过精心策划的转移规则、状态缓存和难以预测事实的记忆库,以此来调节下一个状态的想象。我们通过结构化状态保真度(Structured State Fidelity, SSF)评估事实状态的保留,该指标通过基准特定的事实和领域对预测状态进行评分。与SFT相比,增强记忆训练将SSF提高了高达206.3%。在完整的规划设置中,我们保持策略模型不变,并提供策略侧的世界技能:检索的任务级技能和逐步纠正指导以进行动作选择。在ALFWorld、WebShop和ScienceWorld中,增强记忆的智能体在下游成功率上超过了经过SFT训练的世界模型智能体,最高可达65.4%的相对增益。敏感性分析进一步表明,在不同的记忆和动作预算设置下,检索的记忆提高了任务成功率和效率。
cs.AI / 60 / 2608.07118
How Much, Then Where: Credit-Conserving Action-to-Token Allocation for Multi-Turn Agent Reinforcement Learning
那么多少,在哪里:多回合代理强化学习中的节省信用的动作-令牌分配
Abstract
Credit assignment in multi-turn agent reinforcement learning operates at two levels: assigning trajectory-level credit to actions and distributing each action's credit across its tokens. In this paper, we introduce FACTOR, which separates these decisions. FACTOR uses checkpoint-calibrated TD residuals to assign per-action credits that telescope to the trajectory advantage, and feedback-conditioned teacher-student likelihood gaps to allocate each credit across the realized action tokens. Per-action normalization preserves the action-average coefficient and prevents token-level sign flips. We pair this construction with an action-mean reduction, removing the implicit dependence of an action's scalar surrogate weight on its token length. At the behavior policy and before clipping, each action's inner action-mean surrogate equals its TD credit. FACTOR consistently improves over competitive baselines across ALFWorld, WebShop, and ScienceWorld, with every environment-seed comparison favoring FACTOR and the largest gains emerging on the longest-horizon environment. The same hyperparameters transfer without retuning to a larger backbone and to a different model family. Ablations identify TD action credit as the dominant driver of the improvement, with hindsight token allocation contributing complementary gains.
Chinese Translation
在多回合代理强化学习中,信用分配在两个层面上进行:将轨迹级别的信用分配给动作,并在其令牌之间分配每个动作的信用。本文介绍了FACTOR,它将这些决策分开。FACTOR使用检查点校准的时间差分(TD)残差来分配每个动作的信用,这些信用汇聚到轨迹优势,并利用反馈条件的教师-学生似然差距在实现的动作令牌之间分配每个信用。每个动作的归一化保持了动作平均系数,并防止了令牌级别的符号翻转。我们将这一构造与动作均值减少相结合,消除了动作的标量替代权重对其令牌长度的隐含依赖。在行为策略下,并且在截断之前,每个动作的内部动作均值替代等于其TD信用。FACTOR在ALFWorld、WebShop和ScienceWorld等多个环境中始终优于竞争基线,在每个环境-种子比较中都支持FACTOR,且在最长时间范围的环境中获得了最大的收益。相同的超参数在不重新调整的情况下可以转移到更大的主干网络和不同的模型系列。消融实验表明,TD动作信用是改进的主要驱动因素,而事后令牌分配则贡献了互补的收益。
cs.AI / 61 / 2608.07147
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training
DiDPO:用于编码代理训练的差分政策优化
Abstract
Reinforcement learning with Verifiable Reward (RLVR) has emerged as a powerful paradigm for training coding agents, where the execution feedback from compilation and tests provides objective verification. However, unlike agent tasks, coding agents face a unique and finer-grained credit assignment challenge: at each step, coding actions simultaneously pack varying changes into different regions of a code version, which makes the contribution of independent change indistinguishable. Existing RLVR methods mostly leverage the outcome reward or step-level reward, which fails to dive into a code diff and makes unique properties of coding actions invisible to training. In this paper, we propose Diff-in-Diff Policy Optimization (DiDPO), a critic-free RL method that constructs fine-grained credit units directly from the structure of code diffs. DiDPO organizes multi-turn coding interactions into multiple thought--action steps and discovers code diffs across sampled trajectories. It then selects anchors by aggregating highly similar sub-diffs split from each whole diff by our ``groupability score'', which provides the splitting schema that optimally balances the semantic scope of anchors and the group mass they may form. Finally these anchors form advantage groups and project the diff-level advantage back to individual response tokens. Experiments on long-horizon coding and reasoning benchmarks show that DiDPO significantly outperforms strong agentic RL baselines. On Qwen2.5-7B-Coder, DiDPO exceeds comparable methods by over 10\% and narrows the gap with far larger models, offering a principled framework for fine-grained credit assignment in coding agent training. We also open-source verl-code, an agentic rl codebase that supports various RL methods and coding benchmarks.
Chinese Translation
带有可验证奖励的强化学习(RLVR)已成为训练编码代理的强大范式,其中来自编译和测试的执行反馈提供了客观验证。然而,与代理任务不同,编码代理面临着独特且更细粒度的信用分配挑战:在每一步中,编码操作同时将不同的变化打包到代码版本的不同区域,这使得独立变化的贡献难以区分。现有的RLVR方法大多利用结果奖励或步级奖励,这未能深入到代码差异中,使得编码操作的独特特性在训练中不可见。本文提出了差分政策优化(DiDPO),这是一种无评论员的强化学习方法,直接从代码差异的结构构建细粒度的信用单元。DiDPO将多轮编码交互组织为多个思考-行动步骤,并在采样轨迹中发现代码差异。然后,它通过聚合从每个完整差异中拆分出的高度相似的子差异,利用我们的“可分组性评分”选择锚点,这提供了一个最佳平衡锚点语义范围和它们可能形成的群体质量的拆分方案。最终,这些锚点形成优势组,并将差异级优势投影回单个响应标记。在长时间跨度的编码和推理基准测试中的实验表明,DiDPO显著优于强大的代理强化学习基线。在Qwen2.5-7B-Coder上,DiDPO的表现超过了可比方法10%以上,并缩小了与更大模型之间的差距,为编码代理训练中的细粒度信用分配提供了一个原则性框架。我们还开源了verl-code,这是一个支持各种强化学习方法和编码基准的代理强化学习代码库。
cs.AI / 62 / 2608.07148
A MARL Centered Reference Architecture for Large Language Model Augmentation in Smart Manufacturing
面向智能制造的大型语言模型增强的多智能体强化学习中心参考架构
Abstract
Modern manufacturing imposes six coupled demands on adaptive control: local decisions with global consequences, partial observability, nonstationarity, reflex speed response with long horizon effects, delayed and diffuse outcomes, and dynamics that resist explicit modeling. Cooperative multiagent reinforcement learning (MARL), posed as a Dec-POMDP under centralized training with decentralized execution, is a particularly natural formalism for these demands. This paper adopts a MARL centered scope and asks where large language models (LLMs) should augment, interface with, train, or, in the strongest competitive case, replace that coordination core. A taxonomy organizes the literature through four LLM attachment points: policy, reward design, communication between agents, and hierarchical planning. A conditional capability profile separates native mechanism, reported performance, formal guarantee, and engineering maturity, and a deployment readiness analysis identifies the evidence behind each role. These stages yield the principal contribution: a three layer MARL centered reference architecture, grounded in evidence, for semantic reasoning, adaptive cooperative control, and independently assured execution. The LLM-Augmented Dec-POMDP is a descriptive comparative notation for that architecture, recording four attachment choices without introducing a new decision process class or algorithm. Under the reviewed evidence, conventional MARL is better suited to frequent, structured, decentralized coordination after task specific training, whereas LLM components are promising for semantic interpretation, reward drafting, human interaction, and slower supervisory planning. Current LLM only manufacturing controllers do not yet establish equivalence for strict real time, decentralized, safety critical control; this conclusion is bounded by the available evidence and does not assert impossibility.
Chinese Translation
现代制造业对自适应控制提出了六个相互关联的要求:具有全球影响的局部决策、部分可观测性、非平稳性、具有长期影响的快速反应、延迟和分散的结果,以及抵抗显式建模的动态性。将合作多智能体强化学习(MARL)视为在集中训练与分散执行下的决策部分可观测马尔可夫决策过程(Dec-POMDP),这一形式特别适合这些要求。本文采用以MARL为中心的视角,探讨大型语言模型(LLMs)应在何处增强、接口、训练或在最强竞争情况下替代该协调核心。通过四个LLM附加点:策略、奖励设计、智能体间通信和层次规划,对文献进行了分类。条件能力特征分离了原生机制、报告性能、形式保证和工程成熟度,而部署准备性分析则识别了每个角色背后的证据。这些阶段产生了主要贡献:一个基于证据的三层MARL中心参考架构,用于语义推理、自适应合作控制和独立保证执行。LLM增强的Dec-POMDP是该架构的描述性比较符号,记录了四个附加选择,而不引入新的决策过程类别或算法。在审查的证据下,传统的MARL更适合在任务特定训练后进行频繁、结构化的分散协调,而LLM组件在语义解释、奖励设计、人机交互和较慢的监督规划方面表现出希望。目前仅使用LLM的制造控制器尚未在严格实时、分散、安全关键控制方面建立等效性;这一结论受到可用证据的限制,并不声称不可能。
cs.AI / 63 / 2608.07167
NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs
NiyamAI - 一种基于意图的人工智能代理,具有使用零知识证明的密码学可验证保护措施
Abstract
Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it shouldn't. Prompt injection, hallucinated reasoning, and unsafe tool calls form the primary attack surface for autonomous LLM agents. Existing defenses rely on software checks like system prompts or policy filters running on the same machine the attacker targets, offering no verifiable proof of execution. We introduce Niyam-AI, a framework that makes safety enforcement provable. At session start, permitted tools and constraints are locked into an Intent Contract committed via SHA-256. Every tool call is intercepted and validated by an isolated Judge model; upon passing, a zk-SNARK proof is generated via EZKL. The tool executes only after proof verification, allowing third parties to confirm enforcement without accessing Judge model weights. Evaluating Niyam-AI on 2,000 real-world scenarios from Agent-SafetyBench against NeMo Guardrails, Meta's Llama Prompt Guard 2, and OpenAI's GPT-OSS-Safeguard using 5-fold stratified cross-validation yields an F1 score of 88.5% with a 1.1% false-positive rate (bootstrap 95% CI: [85.19%, 91.88%], N=1000). McNemar's exact paired test confirms significant improvement: Niyam-AI wins 390 discordant scenarios against NeMo (vs 20 losses), 115 against Prompt Guard 2 (vs 13), and 384 against GPT-OSS-Safeguard (vs 19) with p < 0.0001 in all cases. Proof generation adds 2260.6 +/- 218.4 ms per approved action, while verification takes 53.1 +/- 11.8 ms. Niyam-AI provides a guardrail that is both highly accurate and mathematically verifiable--though this reflects a classifier adapted to Agent-SafetyBench evaluated against zero-shot baselines, a distinction discussed in Section IV.C.
Chinese Translation
赋予人工智能代理发送电子邮件、查询数据库或执行命令的能力是有用的——直到代理被欺骗去做不该做的事情。提示注入、幻觉推理和不安全的工具调用构成了自主大型语言模型(LLM)代理的主要攻击面。现有的防御措施依赖于软件检查,如在攻击者目标机器上运行的系统提示或策略过滤器,无法提供可验证的执行证明。我们提出了Niyam-AI,一个使安全执行可证明的框架。在会话开始时,允许的工具和约束被锁定在通过SHA-256提交的意图合同中。每次工具调用都由一个隔离的判决模型拦截和验证;通过后,生成一个通过EZKL的zk-SNARK证明。只有在证明验证后,工具才会执行,这允许第三方在不访问判决模型权重的情况下确认执行。我们在Agent-SafetyBench上对Niyam-AI进行了评估,使用2000个真实场景与NeMo Guardrails、Meta的Llama Prompt Guard 2和OpenAI的GPT-OSS-Safeguard进行5折分层交叉验证,得到了88.5%的F1分数,假阳性率为1.1%(自助法95%置信区间:[85.19%,91.88%],N=1000)。McNemar的精确配对检验确认了显著改善:Niyam-AI在390个不一致场景中战胜NeMo(对比20次失败),在115个场景中战胜Prompt Guard 2(对比13次),在384个场景中战胜GPT-OSS-Safeguard(对比19次),所有情况下p < 0.0001。证明生成每个批准的操作增加2260.6 +/- 218.4毫秒,而验证则需要53.1 +/- 11.8毫秒。Niyam-AI提供了一种既高度准确又数学上可验证的保护措施——尽管这反映了针对Agent-SafetyBench评估的适应于零样本基线的分类器,这一区别在第IV.C节中进行了讨论。
cs.AI / 64 / 2608.07169
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
代理记忆蒸馏:通过层次教师记忆赋能小型大语言模型代理
Abstract
Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models, which struggle to generate sufficient successful trajectories on their own. We propose Agent Memory Distillation (AMD), a training-free framework that transfers structured knowledge from a large teacher agent to a small student agent through hierarchical memory. AMD constructs three complementary memory types from successful teacher trajectories: Workflow memory encodes task-level strategies, Subtask memory provides concrete behavioral examples at an intermediate granularity, and Function memory captures per-function calling conventions and common pitfalls. Workflow and Subtask memories are injected proactively at the start of each task, while Function memory is retrieved reactively upon tool-calling errors. We evaluate AMD on three tool-use benchmarks using four student models (4B-8B parameters) with GPT-5-mini as the teacher, achieving average accuracy gains of 27.2%p, 11.2%p, and 3.4%p on AppWorld, BFCL V3, and ToolSandbox, while consistently outperforming existing memory-based baselines. Further analysis shows that Subtask memory contributes the largest gains, teacher effectiveness depends on both teacher capability and student compatibility, and 4B-sized students benefit most from AMD.
Chinese Translation
记忆系统在提升代理性能方面展现出潜力,但对于小型语言模型而言,其潜力尚未得到充分探索,这些模型在独立生成成功轨迹方面存在困难。我们提出了代理记忆蒸馏(Agent Memory Distillation, AMD),这是一种无训练的框架,通过层次记忆将结构化知识从大型教师代理转移到小型学生代理。AMD 从成功的教师轨迹中构建三种互补的记忆类型:工作流记忆编码任务级策略,子任务记忆提供中间粒度的具体行为示例,而功能记忆捕捉每个功能的调用约定和常见陷阱。工作流记忆和子任务记忆在每个任务开始时主动注入,而功能记忆则在工具调用错误时被反应性地检索。我们在三个工具使用基准上评估了 AMD,使用四个学生模型(4B-8B 参数),以 GPT-5-mini 作为教师,在 AppWorld、BFCL V3 和 ToolSandbox 上分别实现了平均准确率提升 27.2%p、11.2%p 和 3.4%p,同时持续超越现有基于记忆的基线。进一步分析表明,子任务记忆贡献了最大的增益,教师的有效性依赖于教师能力和学生兼容性,而 4B 规模的学生从 AMD 中受益最大。
cs.AI / 65 / 2608.07188
SetEasy: A Multi-Modal Classroom Engagement Assessment and Seating Optimization Framework
SetEasy:一种多模态课堂参与度评估与座位优化框架
Abstract
SetEasy optimizes classroom engagement in fixed seating grids. It fuses multimodal sensing (wristband physiology, 4K video, environmental data) and trains a v-Gage model grounded in a revised ISEQ. Each week, two-week engagement forecasts are mapped to a student-seat utility matrix, and CP-SAT generates seating plans under visual-access and social-dynamics constraints. In a four-week deployment (23 students, 331 classes), v-Gage converged across affective, behavioral, cognitive, and overall dimensions, cutting RMSE from 0.75 to 0.53. Optimization raised mean engagement from 0.30 to 0.70, with over two-thirds of seats reaching high engagement and back-row low-activity patterns markedly reduced. These results show that, without hardware changes, interpretable, data-driven seating strategies can substantially enhance engagement. The multimodal "assessment + optimization" paradigm offers a transferable, sustainable path to culturally responsive, differentiated spatial design amid global homogenization.
Chinese Translation
SetEasy旨在优化固定座位格局下的课堂参与度。该框架融合多模态感知技术(腕带生理数据、4K视频、环境数据),并基于修订后的ISEQ训练v-Gage模型。每周,系统将两周的参与度预测映射为学生-座位效用矩阵,利用CP-SAT在视觉可达性和社交动态约束下生成座位安排方案。在为期四周的部署中(23名学生,331节课),v-Gage模型在情感、行为、认知及整体维度上实现收敛,均方根误差(RMSE)从0.75降至0.53。优化后,平均参与度从0.30提升至0.70,超过三分之二的座位达到高参与度,且后排低活跃度模式显著减少。结果表明,在无需硬件变更的情况下,基于数据驱动且可解释的座位策略能够显著提升课堂参与度。该多模态“评估+优化”范式为在全球趋同背景下实现文化响应性和差异化空间设计提供了可迁移且可持续的路径。
cs.AI / 66 / 2608.07196
EMAS: Stabilizing Multi-Agent System Evolution through Evidence-Guided Revision
EMAS:通过证据引导修订稳定多智能体系统演化
Abstract
Many methods for automated multi-agent system design optimize prompts and topologies during an initial design stage and then deploy the resulting system unchanged on subsequent samples. Experience from these samples is rarely consolidated into reusable system updates, while accuracy-oriented designs may incur high token costs. We introduce EMAS (Evolving Multi-Agent System), which uses this experience to revise MAS topology and prompts without updating LLM parameters, either to improve accuracy or to reduce cost. EMAS converts traces into structured diagnoses that specify a revision operation and target. It generates a candidate revision only when the same diagnosis recurs across samples and applies it only if paired validation against the current MAS meets the corresponding acceptance criterion. Across four benchmarks and two LLMs, EMAS attains the highest task-weighted overall accuracy for both backbones and is best or tied in six of eight model--benchmark settings. Within two evolution epochs, EMAS achieves relative gains of 6.30% and 20.10% in task-weighted accuracy on Kimi-K2-6 and Qwen3.6-27B, respectively. On MBPP with Qwen3.6-27B, EMAS raises accuracy from 55.09% to 89.12% while reducing token use per task by 62.2%. These results show that EMAS can turn experience from new samples into reusable updates to MAS topology and prompts.
Chinese Translation
许多自动化多智能体系统设计的方法在初始设计阶段优化提示和拓扑结构,然后在后续样本中不改变地部署所得到的系统。这些样本的经验很少被整合成可重用的系统更新,而以准确性为导向的设计可能会导致高额的令牌成本。我们提出了EMAS(演化多智能体系统),利用这些经验在不更新大语言模型(LLM)参数的情况下修订MAS拓扑和提示,以提高准确性或降低成本。EMAS将痕迹转换为结构化诊断,指定修订操作和目标。仅当相同的诊断在多个样本中重复出现时,它才生成候选修订,并且仅在与当前MAS的配对验证满足相应的接受标准时才应用该修订。在四个基准测试和两个LLM上,EMAS在两个基础模型上都达到了最高的任务加权整体准确率,并且在八个模型-基准设置中的六个中表现最佳或并列最佳。在两个演化周期内,EMAS在Kimi-K2-6和Qwen3.6-27B上分别实现了6.30%和20.10%的任务加权准确率相对提升。在使用Qwen3.6-27B的MBPP上,EMAS将准确率从55.09%提高到89.12%,同时每个任务的令牌使用减少了62.2%。这些结果表明,EMAS能够将来自新样本的经验转化为可重用的MAS拓扑和提示更新。
cs.AI / 67 / 2608.07202
Authoring and Management of Transparent Research Integrity Assessments of Randomised Clinical Trial Publications Using LLM-assisted Tools and Provenance Knowledge Graphs
使用LLM辅助工具和来源知识图谱进行随机临床试验出版物透明研究诚信评估的创作与管理
Abstract
Systematic reviews of Randomised Controlled Trials (RCTs) are routinely used as evidence for clinical care guidelines. Such evidence has to meet high research integrity standards to prevent low quality or false research outputs influencing the clinical care. However, assessing research integrity of published RCTs is a complex process requiring manual effort, and potentially resulting in diverse opinions of the human assessors. This paper describes INSPECT-AI, an LLM-based interactive tool that assists human reviewers with research integrity assessments of published RCTs based on the community approved INSPECT-SR framework, and the Research Integrity Provenance and Evidence ontology (RIPE-O) for documenting the provenance of the assessment process. In addition, we present the Research Integrity Provenance and Evidence knowledge graph (RIPE-KG), an initial set of 140 expert research integrity assessments of 95 RCT publications generated by INSPECT-AI and described using RIPE-O.
Chinese Translation
随机对照试验(RCT)的系统评价通常作为临床护理指南的证据。这类证据必须符合高标准的研究诚信,以防止低质量或虚假的研究结果影响临床护理。然而,评估已发布RCT的研究诚信是一个复杂的过程,需手动进行,可能导致评估者之间出现不同的意见。本文描述了INSPECT-AI,这是一种基于LLM的互动工具,旨在根据社区认可的INSPECT-SR框架,辅助人类评审者进行已发布RCT的研究诚信评估,并利用研究诚信来源与证据本体(RIPE-O)记录评估过程的来源。此外,我们还展示了研究诚信来源与证据知识图谱(RIPE-KG),这是由INSPECT-AI生成的95篇RCT出版物的140个专家研究诚信评估的初步集合,并使用RIPE-O进行描述。
cs.AI / 68 / 2608.07220
Beyond the Black Box: Interpretable Models of Human Randomisation Failures
超越黑箱:人类随机化失败的可解释模型
Abstract
Mixed strategy equilibrium predicts i.i.d play: past actions should not help predict future decisions. Human players, however, systematically depart from this benchmark, and in O'Neill's zero sum card game, these departures can be predicted by black box sequence models such as LSTMs. This paper asks whether that predictive power can be achieved by transparent alternatives that also reveal the behavioural structure behind it. Using 84,060 decisions from 2,802 pairs, the analysis first benchmarks naive and behavioral models against interpretable machine learning and deep learning models, then evaluates the modified EWA specifications of prior work against these benchmarks and uses the LASSO diagnostics to motivate a further nested frequency tracking extension. The results show that repeat or avoid behavior, especially players' management of their own recent action histories, accounts for most of the interpretable and strategically exploitable signal, while frequency tracking adds little out of sample.
Chinese Translation
混合策略均衡预测独立同分布(i.i.d)行为:过去的行动不应有助于预测未来的决策。然而,人类玩家系统性地偏离这一基准,在O'Neill的零和纸牌游戏中,这些偏离可以通过黑箱序列模型(如LSTM)进行预测。本文探讨这种预测能力是否可以通过透明的替代模型来实现,同时揭示其背后的行为结构。通过分析来自2802对的84060个决策,研究首先将天真模型和行为模型与可解释的机器学习和深度学习模型进行基准比较,然后评估先前工作的修改EWA(经验加权平均)规格与这些基准的对比,并使用LASSO诊断来推动进一步的嵌套频率跟踪扩展。结果表明,重复或避免行为,特别是玩家对自己近期行动历史的管理,解释了大部分可解释且具有战略可利用信号,而频率跟踪在样本外的贡献较小。
cs.AI / 69 / 2608.07230
From probability to causality in probabilistic logic programming
从概率到因果关系:概率逻辑编程中的因果推理
Abstract
Probabilistic logic programming is a formalism of statistical relational artificial intelligence that supports causal queries, including interventions from outside the system. When the structure of a probabilistic logic program is learned from data, however, only probabilistic information is used, and a single probability distribution may be compatible with several causal orders. This leads to ambiguity in interventional reasoning, raising the question of when the causal order is uniquely determined by the distribution. Exploiting the relationship between acyclic probabilistic logic programs and Bayesian networks, we derive conditions under which the probabilistic information encoded in a program determines a unique causal order. We also incorporate constraints arising from relational structure by taking into account prescribed sets of causal symmetries induced by the underlying relational vocabulary. The result is a method for verifying when a learned probabilistic logic program supports well-defined intervention semantics.
Chinese Translation
概率逻辑编程是一种统计关系人工智能的形式,支持因果查询,包括来自系统外部的干预。然而,当从数据中学习概率逻辑程序的结构时,仅使用概率信息,且单一的概率分布可能与多个因果顺序兼容。这导致了干预推理中的模糊性,提出了一个问题:何时概率分布唯一确定因果顺序。通过利用无环概率逻辑程序与贝叶斯网络之间的关系,我们推导出在何种条件下,程序中编码的概率信息决定唯一的因果顺序。我们还通过考虑由底层关系词汇诱导的规定性因果对称集,纳入了来自关系结构的约束。最终结果是一个验证学习到的概率逻辑程序何时支持明确干预语义的方法。
cs.AI / 70 / 2608.07243
Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models
创造力的配方:大型语言模型中的迭代生成与评估
Abstract
Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement. This pilot study examines whether iterative search improves LLM creativity by adapting FunSearch to recipe generation for the 2024 Pillsbury Bake-Off and evaluating outputs against human benchmarks using TTCT-based LLM evaluation. Across two experiments, we test iteration count, generator temperature, and in-loop selection-scorer model size. Results show that iterative generation-selection can produce recipes with creativity scores comparable to human benchmarks, but additional iterations alone do not improve creativity. The in-loop evaluator matters most: a smaller selection scorer yields significantly higher scores across most TTCT dimensions, while temperature has limited effects except for originality. These findings suggest that evaluator design is a first-order design variable in subjective creative search.
Chinese Translation
生成模型通常通过单一的工件进行评估,而人类创造力通常通过迭代生成、评估和完善而出现。本研究探讨了迭代搜索是否通过将 FunSearch 适应于 2024 年 Pillsbury Bake-Off 的配方生成来提升大型语言模型(LLM)的创造力,并使用基于 TTCT 的 LLM 评估对输出进行与人类基准的比较。在两项实验中,我们测试了迭代次数、生成器温度和循环内选择评分模型的大小。结果表明,迭代生成-选择可以产生与人类基准相当的创造力评分,但单纯增加迭代次数并不提高创造力。循环内评估者的设计最为重要:较小的选择评分器在大多数 TTCT 维度上产生显著更高的评分,而温度的影响有限,仅在独创性方面有所体现。这些发现表明,评估者设计是主观创造性搜索中的一项重要设计变量。
cs.AI / 71 / 2608.07267
WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN
WNM-3D:一种具有3D场景条件的世界导航模型,用于闭环视觉语言导航
Abstract
Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions. Although semantically capable, such action-centric training does not explicitly model how the agent's visual observations should evolve under its predicted motion. Generative world-action models (WAMs) jointly predict future observations and actions, yet existing WAMs for continuous VLN do not condition joint future-view and action generation on geometry-aware representations inferred from the observed history. We present WNM-3D, a generative World Navigation Model with 3D scene conditioning for continuous VLN. To consolidate past observations into persistent scene context, a frozen feed-forward geometry encoder extracts geometry-aware representations from the monocular egocentric RGB history, and a trainable 3D Scene-to-Token Adapter converts them into a fixed-length prefix in the token space of the world-action Diffusion Transformer. Through block-causal attention, this prefix conditions every future video-action block, providing a shared geometric context for both future-view and action generation. We train WNM-3D through supervised world-action fine-tuning on A*-generated demonstrations, DAgger-style adaptation on policy-visited states, and DanceGRPO-based closed-loop policy optimization. Experiments on GN-Bench show that WNM-3D outperforms strong VLM-based navigation policies and its 2D-conditioned counterpart in closed-loop navigation. On a fixed near-goal evaluation set, WNM-3D also achieves higher flow-action consistency and lower visual-motion error.
Chinese Translation
近年来,视觉语言导航(VLN)系统越来越多地将预训练的视觉语言模型(VLMs)适配为视觉语言行动(VLA)策略,直接将自我中心的观察和语言指令映射到导航行动上。尽管在语义上具有能力,但这种以行动为中心的训练并未明确建模代理的视觉观察在其预测运动下应如何演变。生成性世界行动模型(WAMs)共同预测未来的观察和行动,然而现有的用于连续VLN的WAMs并未基于从观察历史推断出的几何感知表示来调节联合未来视图和行动生成。我们提出了WNM-3D,一种具有3D场景条件的生成性世界导航模型,用于连续VLN。为了将过去的观察整合为持久的场景上下文,一个冻结的前馈几何编码器从单目自我中心的RGB历史中提取几何感知表示,而一个可训练的3D场景到标记适配器将其转换为世界行动扩散变换器的固定长度前缀。通过块因果注意力,这个前缀为每个未来的视频行动块提供条件,为未来视图和行动生成提供共享的几何上下文。我们通过对A*生成的演示进行监督的世界行动微调、在策略访问状态上的DAgger风格适应以及基于DanceGRPO的闭环策略优化来训练WNM-3D。在GN-Bench上的实验表明,WNM-3D在闭环导航中优于强大的基于VLM的导航策略及其2D条件对应物。在固定的近目标评估集上,WNM-3D还实现了更高的流动行动一致性和更低的视觉运动误差。
cs.AI / 72 / 2608.07303
Winning by Peeking: Unenforced Budgets and Test-Set Selection Inflate Short-Budget AutoML Comparisons
通过窥视获胜:未强制执行的预算和测试集选择夸大了短预算自动机器学习比较的结果
Abstract
Comparisons between AutoML systems at short time budgets -- tens of seconds rather than hours -- are common in tool READMEs and workshop papers, and they are easy to get wrong. We report a case study in which a simple AutoML engine, Orcetra, appeared to beat FLAML and AutoGluon on 513 OpenML datasets, winning 57.1% of them at a nominal 60-second budget and 78.4% of datasets against FLAML alone at 30 seconds. Both margins came from protocol defects that a results table cannot show. The search loop scored every candidate on the test split and reported the best, making the headline metric a maximum over dozens of noisy estimates while the baselines selected on training data and touched the test set once; and the budget was checked before launching a candidate but never enforced during one, so the system consumed a median of 120 s against a 60-second budget, 2.24x the wall-clock AutoGluon used. Re-running with selection moved to a validation split, the deadline enforced externally and every framework pinned to an equal share of the machine, Orcetra's win rate on the re-run subset falls from 59.4% to 34.3% and no pairwise difference against either competitor remains significant. Recording both estimands inside a single search lets us attribute the collapse: the selection rule accounts for 4.8 percentage points and unequal compute for most of the rest. The same traces give the selection bias as a function of budget, measured rather than assumed: it grows with $K$ but reaches only 0.27 accuracy points, about five times below the $\sigma\sqrt{2\ln K}$ bound a marginal-standard-error argument predicts, because candidates scored on shared test rows cancel most of the noise. We close with a checklist for short-budget comparisons. Code, per-dataset results and the scripts that regenerate every number and figure in the paper are released with it.
Chinese Translation
在短时间预算下(以秒而非小时计)对自动机器学习(AutoML)系统的比较在工具的自述文件和研讨会论文中很常见,但容易出现错误。我们报告了一个案例研究,其中一个简单的AutoML引擎Orcetra在513个OpenML数据集上似乎击败了FLAML和AutoGluon,在名义60秒的预算下赢得了57.1%的数据集,并在30秒内对FLAML单独的情况下赢得了78.4%的数据集。这两个优势源于协议缺陷,结果表格无法显示。搜索循环对每个候选者在测试分割上进行了评分并报告了最佳结果,使得头条指标成为数十个噪声估计的最大值,而基线是在训练数据上选择并仅触及测试集一次;预算在启动候选者之前进行了检查,但在执行过程中从未强制执行,因此系统在60秒的预算下消耗了中位数120秒,约为AutoGluon实际使用时间的2.24倍。重新运行时将选择移动到验证分割,外部强制执行截止时间,并将每个框架固定在机器的相等份额上,Orcetra在重新运行子集上的胜率从59.4%降至34.3%,且与任何竞争对手之间的配对差异均不显著。将两个估计量记录在单次搜索中使我们能够归因于这一崩溃:选择规则占4.8个百分点,而不平等计算则占大部分剩余部分。相同的轨迹提供了选择偏差作为预算的函数,经过测量而非假设:它随着$K$的增加而增长,但仅达到0.27的准确度点,约为边际标准误差论证预测的$rac{ heta}{ heta ext{标准误}}$界限的五分之一,因为在共享测试行上评分的候选者抵消了大部分噪声。我们以短预算比较的检查清单结束。代码、每个数据集的结果以及再生论文中每个数字和图表的脚本将随论文一起发布。
cs.AI / 73 / 2608.07346
An End-to-End Agent Auditing Engine
端到端代理审计引擎
Abstract
With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains. The fast-evolving harness ecosystem has also made rigorous capability evaluation increasingly important. However, efficiently building an end-to-end, systematic, and comprehensive evaluation pipeline remains a significant challenge. To address this challenge, we introduce $A^2E$ (Agent Auditing Engine), an end-to-end evaluation engine designed for agent harnesses. $A^2E$ leverages our newly proposed Agent Task Protocol (ATP) to enable the rapid integration of evaluation tasks with different harnesses. Through an automatically instrumented Monitor, it captures and generates standardized execution traces during experiments. In the Evaluation stage, $A^2E$ systematically assesses harness capabilities using a suite of multidimensional metrics. Compared with correctness alone, these metrics provide a more fine-grained characterization of differences among harnesses in execution efficiency, tool use, task planning, and error recovery. Experiments conducted with $A^2E$ further reveal that model-harness combinations exhibit substantial performance variation across different types of tasks, and that no single combination consistently outperforms all others across every task. These findings not only demonstrate the necessity of systematic evaluation but also provide useful guidance for the co-evolving of models and harnesses. Our code is available at https://github.com/datamllab/A2E.
Chinese Translation
随着大型语言模型(LLMs)的快速发展,代理架构已成为在广泛领域中部署代理的基本基础设施。快速发展的代理生态系统也使得严格的能力评估变得愈加重要。然而,高效构建一个端到端、系统化和全面的评估流程仍然是一个重大挑战。为了解决这个挑战,我们提出了 $A^2E$(代理审计引擎),这是一个专为代理架构设计的端到端评估引擎。$A^2E$ 利用我们新提出的代理任务协议(Agent Task Protocol, ATP)来实现与不同代理架构的评估任务的快速集成。通过自动化的监控器,它在实验过程中捕获并生成标准化的执行轨迹。在评估阶段,$A^2E$ 使用一套多维度指标系统地评估代理架构的能力。与单纯的正确性相比,这些指标提供了对代理架构在执行效率、工具使用、任务规划和错误恢复等方面差异的更细致的表征。通过 $A^2E$ 进行的实验进一步揭示,模型-代理组合在不同类型任务中表现出显著的性能差异,并且没有任何单一组合在每个任务中始终优于其他组合。这些发现不仅展示了系统评估的必要性,也为模型与代理架构的共同演进提供了有益的指导。我们的代码可在 https://github.com/datamllab/A2E 获取。
cs.AI / 74 / 2608.07363
QFCQT: A Chaotically Gated Quantformer Framework for Volatile Time-Series Forecasting
QFCQT:一种用于波动时间序列预测的混沌门控量子变换器框架
Abstract
Forecasting non-stationary time series remains difficult due to long-range dependencies, local volatility bursts, structural shifts, and nonlinear oscillatory behaviors. Although Transformer-based forecasters are effective for modeling long-term temporal dependencies, their feed-forward blocks typically rely on smooth static activations that are insufficiently sensitive to abrupt regime changes. Motivated by quantitative Transformer designs and oscillator-based nonlinear activations, we propose QFCQT, short for Quantum-Fractal-inspired Chaotically Gated Quantformer, for robust forecasting under complex volatile dynamics. Here, "quantum-fractal-inspired" denotes a computational analogy based on soft oscillator superposition and multi-scale nonlinear responses, rather than a formal quantum-mechanical or fractal-theoretic derivation. QFCQT consists of three main components: (1) a Quantformer-style numerical encoder that directly processes multivariate inputs via linear embedding; (2) a learnable Lee-oscillator activation module that maps scalar pre-activations to dynamic oscillatory responses and summarizes them through Max-over-Time pooling; and (3) a smooth-chaotic gated fusion mechanism that adaptively balances conventional smooth activations and chaos-sensitive responses. Furthermore, instead of using a single fixed oscillator, QFCQT employs a soft superposition of eight parameterized Lee oscillator families to adaptively capture different nonlinear response patterns across regimes. Experiments on ETTh1, ETTh2, and A-share Stock Index benchmarks show that QFCQT consistently outperforms strong baselines, including Informer, LogTrans, LSTMa, HAT, and COTN.
Chinese Translation
预测非平稳时间序列仍然面临挑战,原因在于长程依赖性、局部波动爆发、结构性变化以及非线性振荡行为。尽管基于变换器(Transformer)的预测模型在建模长期时间依赖性方面有效,但其前馈模块通常依赖于平滑的静态激活,这对突发的状态变化敏感性不足。受量子变换器设计和基于振荡器的非线性激活的启发,我们提出了QFCQT(量子-分形启发的混沌门控量子变换器),以应对复杂波动动态下的稳健预测。在这里,“量子-分形启发”指的是基于软振荡器叠加和多尺度非线性响应的计算类比,而不是正式的量子力学或分形理论推导。QFCQT由三个主要组件组成:(1) 一种量子变换器风格的数值编码器,直接通过线性嵌入处理多变量输入;(2) 一个可学习的李振荡器激活模块,将标量预激活映射到动态振荡响应,并通过时间最大池化进行汇总;(3) 一种平滑-混沌门控融合机制,自适应平衡传统平滑激活和对混沌敏感的响应。此外,QFCQT并不使用单一固定的振荡器,而是采用八个参数化李振荡器族的软叠加,以自适应捕捉不同状态下的非线性响应模式。在ETTh1、ETTh2和A股指数基准上的实验表明,QFCQT始终优于包括Informer、LogTrans、LSTMa、HAT和COTN在内的强基线模型。
cs.AI / 75 / 2608.07364
Curriculum as Code: An AI-Assisted Architecture for Instructional Design in STEM Education
课程即代码:一种基于人工智能辅助的STEM教育教学设计架构
Abstract
Contribution: This paper presents a six-phase AI-assisted instructional design architecture based on the Curriculum as Code paradigm, integrating Generative AI with LaTeX and Python to automate the creation of reproducible, visually consistent, and technically precise materials for STEM education. Background: Creating customized instructional materials for active learning imposes a heavy workload on faculty. Standard presentation tools lack robust support for technical content, while current AI applications often hallucinate and fail to formalize the instructional authoring process, limiting their utility for rigorous academic design. Intended Outcomes: The framework aims to reduce preparation time while ensuring mathematical accuracy, adherence to institutional visual identity, and preservation of the instructor's tacit pedagogical knowledge through explicit rules. Application Design: The solution comprises a six-phase pipeline that replaces ad-hoc prompt engineering with a systematic workflow, utilizing text-based interfaces and code-driven generation (LaTeX/Beamer for slides, Python for figures), governed by pedagogical constraints, contextual calibrations, and automated review cycles. Findings: Validated over one year across 8 modules and 28 project contexts in a Project-Based Learning environment, the architecture significantly reduced instructor workload. Generated assets underwent independent peer review and were deployed by six different faculty members, confirming scalability beyond a single author. Based on over 600 voluntary student evaluations, materials achieved high quality ratings from 8.5 to 9.9/10. Results indicate high reproducibility, minimized hallucinations, and sustained pedagogical and visual fidelity, suggesting viability for broad STEM educational applications.
Chinese Translation
贡献:本文提出了一种基于“课程即代码”范式的六阶段人工智能辅助教学设计架构,结合生成性人工智能与LaTeX和Python,自动化生成可重复、视觉一致且技术精确的STEM教育材料。背景:为主动学习创建定制的教学材料给教师带来了沉重的工作负担。标准演示工具对技术内容的支持不足,而当前的人工智能应用往往出现幻觉,并未能规范教学创作过程,限制了其在严谨学术设计中的实用性。预期成果:该框架旨在减少准备时间,同时确保数学准确性、遵循机构视觉识别,并通过明确规则保留教师的隐性教学知识。应用设计:该解决方案包括一个六阶段流程,取代了临时的提示工程,采用系统化的工作流程,利用基于文本的接口和代码驱动的生成(使用LaTeX/Beamer制作幻灯片,使用Python生成图形),受教学约束、上下文校准和自动审查周期的管理。发现:在基于项目的学习环境中,经过一年在8个模块和28个项目背景下的验证,该架构显著减少了教师的工作负担。生成的资产经过独立同行评审,并由六位不同的教师部署,确认了其超越单一作者的可扩展性。基于600多份自愿的学生评估,材料的质量评分高达8.5至9.9/10。结果表明高可重复性、最小化幻觉,并保持教学和视觉的忠实度,表明其在广泛的STEM教育应用中的可行性。
cs.AI / 76 / 2608.07367
People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe
人们不仅仅是他们的国家:解开欧洲大型语言模型价值观对齐的社会决定因素
Abstract
As Large Language Models (LLMs) are increasingly used as a primary source of information and advice, understanding their alignment to humans in terms of values becomes a pressing concern. A growing literature has leveraged large scale surveys to investigate to what extent LLMs' and humans' stated values and opinions align. With limited exceptions, studied populations have been defined country borders or cultural bounds. Yet, this focus neglects the role that socio-demographic divides may play for value alignment disparities. Relying on the European Social Survey, we address this knowledge gap by considering value alignment displayed with respect to 10 prominent commercial LLMs in terms of 15 socio-demographic variables as well as country of residence. Our analyses reveal that LLMs are indeed unequally aligned to the values of different socio-demographic groups, notably those defined by education, income, occupation and religion. When examining alignment at the individual level, a respondent's country, taken as a stand-alone variable, explains a substantial amount of variation that is on par with the full set of considered socio-demographics. Further disentangling the respective role of country-level and socio-demographic factors, we find they are complementary in explaining value alignment patterns, with their relative weights varying across the subset of questions considered.
Chinese Translation
随着大型语言模型(LLMs)越来越多地被用作信息和建议的主要来源,理解它们在价值观方面与人类的对齐程度成为一个紧迫的问题。越来越多的文献利用大规模调查研究LLMs与人类所表述的价值观和观点在多大程度上对齐。除了有限的例外,研究的人群通常以国家边界或文化界限为定义。然而,这种关注忽视了社会人口分化在价值观对齐差异中可能发挥的作用。依托欧洲社会调查,我们通过考虑与10个知名商业LLMs相关的15个社会人口变量以及居住国家,来填补这一知识空白。我们的分析揭示,LLMs确实在不同社会人口群体的价值观上存在不平等的对齐,尤其是在教育、收入、职业和宗教等方面。当在个体层面上考察对齐时,受访者的国家作为一个独立变量,解释了与考虑的所有社会人口变量相当的显著变异。进一步解开国家层面和社会人口因素的各自作用,我们发现它们在解释价值观对齐模式时是互补的,其相对权重在所考虑的问题子集中有所不同。
cs.AI / 77 / 2608.07400
FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings
FinRank:基于证据的金融问答与检索基准,针对SEC文件
Abstract
Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answer can be grounded in the wrong evidence. Similar facts and disclosures recur across sections of a filing, across reporting periods of the same firm, and across comparable firms. FinRank targets this provenance-sensitive retrieval problem by requiring systems to identify evidence for the intended entity, reporting period, and disclosure context. The benchmark contains 1185 manually authored question-answer records over the 10-K and 10-Q filings of 22 companies. Each record includes a reference answer, gold supporting passages, and hand-curated hard negatives drawn from confusable passages within filings, across reporting periods, and across comparable firms. FinRank evaluates passage retrieval, reranking, and hard-negative discrimination as separately measured tasks. Baseline results demonstrate the difficulty of this setting: among the evaluated systems, even a 7B instruction-tuned embedder reaches only 44.8% Recall@10 on the pooled evidence corpus; sub-billion-parameter encoders gain at most 3.5 points over BM25, a finance-adapted embedder trails BM25 by 9.7 points, and pairwise accuracy falls by 13.0-20.5 percentage points when random negatives are replaced with the curated hard negatives. FinRank provides an evidence-first benchmark for developing financial question answering systems that are not only accurate but also grounded in the correct disclosure.
Chinese Translation
金融问答通常通过答案的正确性进行评估,但在SEC文件中,一个看似合理甚至在数值上正确的答案可能基于错误的证据。相似的事实和披露在文件的不同部分、同一公司的不同报告期以及可比公司之间反复出现。FinRank 针对这一依赖来源的检索问题,要求系统识别针对特定实体、报告期和披露上下文的证据。该基准包含22家公司在10-K和10-Q文件中的1185个手动编写的问题-答案记录。每个记录包括一个参考答案、黄金支持段落,以及从文件中混淆段落、不同报告期和可比公司中手工挑选的困难负例。FinRank 将段落检索、重新排序和困难负例区分作为单独测量的任务进行评估。基线结果显示该设置的困难性:在评估的系统中,即使是一个经过指令调优的7B嵌入模型在汇总的证据语料库上仅达到44.8%的Recall@10;不足十亿参数的编码器最多比BM25提高3.5个点,一个适应金融的嵌入模型比BM25低9.7个点,当随机负例被替换为精心挑选的困难负例时,成对准确率下降了13.0-20.5个百分点。FinRank 提供了一个以证据为基础的基准,用于开发不仅准确而且基于正确披露的金融问答系统。
cs.AI / 78 / 2608.07411
GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks
GeoBenchLLM:评估大型语言模型在地理相关任务上的综合基准
Abstract
In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmark for probing LLMs on geo-related tasks. We leverage a careful selection of twelve publicly available datasets from diverse geo-related tasks and domains, and evaluate a set of LLMs on geo-spatial and temporal understanding using our benchmark. Our results show that reasoning and size have a strong impact on overall performance. GeoBenchLLM is publicly available at https://github.com/Rfr2003/GeoBenchLLM.
Chinese Translation
在地理数据的背景下,现有的大型语言模型(LLMs)通常在同质环境中进行研究,这在很大程度上限制了对其泛化能力的洞察。在本文中,我们提出了GeoBenchLLM,这是一个用于探测LLMs在地理相关任务上的综合基准。我们精心选择了来自不同地理相关任务和领域的十二个公开数据集,并利用我们的基准评估了一组LLMs在地理空间和时间理解方面的表现。我们的结果表明,推理能力和模型规模对整体性能有显著影响。GeoBenchLLM可在https://github.com/Rfr2003/GeoBenchLLM上公开获取。
cs.AI / 79 / 2608.07418
ResidencyRL: Reinforcement Learning in Simulated Clinical Environments
ResidencyRL:在模拟临床环境中的强化学习
Liévin, Valentin, Schmidgall, Samuel, Strother, Tim, Bijamov, Alex, Goel, Akshay, Palepu, Anil, Park, Chunjong, Balazadeh, Vahid, Sun, Min Woo, Guerard, Marius, Chen, Justin, Steiner, Dave, Dhillon, Vikram, Azar, Ibrahim, Mehta, Akhil, Spetsieris, Nicholas, Shah, Shilpan, Abdelrahim, Maen, Dahiya, Amit, Liu, Yun, Chou, Katherine, Matias, Yossi, Hassidim, Avinatan, Webster, Dale R., Le, Quoc V., Hadsell, Raia, Barral, Joelle, Radebaugh, Carey, Faust, Aleksandra, Azizi, Shekoofeh, Schaekermann, Mike, Chen, Po-Hsuan Cameron, Tu, Tao, Racz, David, Yang, Lin
Abstract
In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater autonomy. Much of clinical reasoning relies on the patient encounter, a dialogue in which a clinician elicits history, refines diagnostic hypotheses, and decides management under uncertainty. While large language models (LLMs) excel on static medical benchmarks, methods to optimize the full sequence of clinical decisions remain underdeveloped. We present ResidencyRL, a reinforcement learning (RL) method for training clinical artificial intelligence (AI) agents through simulated multi-turn clinical encounters (up to 60 dialogue turns and 8 tool calls per trajectory). ResidencyRL pairs the policy agent with LLM simulators capable of complex, adversarial behaviors, training against a structured reward aligned to diagnostic accuracy, management quality, communication, documentation, and safety. On held-out evaluations, the ResidencyRL agent improves diagnostic accuracy by 7.0% under adversarial conditions (88.0% vs. 81.0%) and reduces missed red flag rates by 31%, demonstrating rigorous mitigation of premature closure. Blinded expert clinicians validated these gains, preferring the trained agent in 87.6% of side-by-side comparisons. The procedural competencies transfer to unseen benchmarks: the agent outperforms the base model across all six clinical axes of the AMIE multi-visit benchmark, and shows consistent directional improvements on AgentClinic and CRAFT-MD. Our findings demonstrate that sequential clinical decision-making can be effectively learned through multi-turn RL in simulation, yielding robust, generalizable capabilities, paving the way towards clinical mastery. Prospective validation with real-world workflows remains necessary to establish clinical utility.
Chinese Translation
在医学教育中,医生通过住院医师培训将学术知识转化为临床专业技能:在数千次接触中进行多年的培训,获得多样化的反馈来源和逐渐增加的自主权。临床推理在很大程度上依赖于患者接触,这是一个对话过程,在此过程中,临床医生获取病史、完善诊断假设并在不确定性下决定管理方案。尽管大型语言模型(LLMs)在静态医学基准测试中表现出色,但优化完整的临床决策序列的方法仍然不够成熟。我们提出了ResidencyRL,这是一种通过模拟多轮临床接触(每个轨迹最多60轮对话和8次工具调用)训练临床人工智能(AI)代理的强化学习(RL)方法。ResidencyRL将策略代理与能够进行复杂对抗行为的LLM模拟器配对,针对与诊断准确性、管理质量、沟通、文档记录和安全性相关的结构化奖励进行训练。在保留的评估中,ResidencyRL代理在对抗条件下提高了7.0%的诊断准确性(88.0%对比81.0%),并将漏诊红旗的比例降低了31%,显示出对过早关闭的严格缓解。盲评专家临床医生验证了这些提升,在87.6%的并排比较中更倾向于训练后的代理。程序性能力转移到未见基准:该代理在AMIE多次访问基准的六个临床维度上超越了基础模型,并在AgentClinic和CRAFT-MD上显示出一致的方向性改进。我们的研究结果表明,序列临床决策可以通过多轮强化学习在模拟中有效学习,从而获得稳健且可推广的能力,为临床精通铺平道路。未来需要在真实工作流程中进行前瞻性验证,以确立其临床实用性。
cs.AI / 80 / 2608.07424
CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing
CoBa:通过计算平衡路由实现成本有效的测试时间扩展
Abstract
Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions, extending a chain of thought, or applying a stronger evaluator. Under a fixed inference budget, these choices compete. This paper formulates test-time reasoning as a compute-allocation problem in which a system must decide whether the next unit of compute should be spent on generation, verification, or stopping. We introduce CoBa, a compute-balanced routing policy that first obtains a small set of candidates, applies cheap verification broadly, and routes uncertain or high-value candidates to stronger verification. On 3,129 example-generator evaluations spanning MATH-500, AIME 2024/2025, AMC 2023, and procedural symbolic reasoning, CoBa-Routed-Strong reaches 85.13% macro accuracy, statistically matching a self-evaluation weighted-voting proxy at 85.20% while using 49.1% fewer parameter-weighted tokens. It also matches best-of-16 majority voting within 0.01 macro-accuracy points while using 58.9% fewer parameter-weighted tokens; paired tests retain a small best-of-16 edge at substantially higher cost. Paired bootstrap tests show significant gains over single-sample decoding, while the remaining gap to the pool oracle exposes headroom for sharper routing. For local reasoning systems, test-time scaling becomes a question of where the next computation is most valuable.
Chinese Translation
测试时间扩展通常通过在一个维度上投入更多计算来实现:采样更多解决方案、延长思维链或应用更强的评估器。在固定的推理预算下,这些选择相互竞争。本文将测试时间推理表述为一个计算分配问题,系统必须决定下一单位计算应花费在生成、验证还是停止上。我们提出了CoBa,一种计算平衡的路由策略,首先获取一小组候选者,广泛应用廉价验证,并将不确定或高价值的候选者路由到更强的验证上。在涵盖MATH-500、AIME 2024/2025、AMC 2023和程序性符号推理的3,129个示例生成器评估中,CoBa-Routed-Strong达到了85.13%的宏观准确率,统计上与自我评估加权投票代理的85.20%相匹配,同时使用了49.1%更少的参数加权标记。它还在宏观准确率上与16个最佳投票的结果相匹配,差距仅为0.01,同时使用了58.9%更少的参数加权标记;配对测试在显著更高的成本下保持了小幅的最佳16个优势。配对自助测试显示出相较于单样本解码的显著提升,而与池子oracle的剩余差距则暴露了更精确路由的潜力。对于局部推理系统而言,测试时间扩展成为了下一个计算最有价值的地方的问题。
cs.AI / 81 / 2608.07427
A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy
一图胜千言:视觉语言模型如何降低人工智能能源成本同时提高准确性
Abstract
LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count---a critical inefficiency for telecom network analytics and numerical time-series data analysis (NTSDA), where raw multivariate KPI windows from 4G/5G cell sites expand into thousands of floating-point tokens. Vision-Language Models (VLMs) eliminate this mismatch by encoding time-series as 2D plots, achieving 3.6-10.4x input token reduction across Llama-3.2-90B, Qwen2.5-VL-72B, and Pixtral-12B architectures. This translates to 1.8-2.5x measured inference energy reduction, saving approximately 7.2 MJ/day at telecom edge deployments and CloudRAN that monitor 200 cells per 15-minute interval. Critically, efficiency gains do not sacrifice accuracy: a fine-tuned Llama-3.2-90B-Vision VLM achieves 220.7% higher precision than its text-only counterpart and outperforms LSTM and ARIMA baselines by over 144% on telecom anomaly detection. On public benchmarks, Pixtral-12B achieves a 20.6x improvement in J/F1 score at mean F1 = 0.82. At 24 KPIs, text representations exceed the 128K context window of most production LLMs, rendering text-only processing infeasible without truncation, while visual representations remain within standard limits. These results establish VLMs as an energy-efficient and accuracy-superior modality for numerical time-series workloads, providing empirical grounding for AI inference systems that treat energy consumption as a first-class engineering constraint.
Chinese Translation
大型语言模型(LLM)推理占人工智能运营能源的90%以上,且与输入令牌数量直接相关——这对于电信网络分析和数值时间序列数据分析(NTSDA)来说是一种关键的低效,因为来自4G/5G基站的原始多变量KPI窗口会扩展为数千个浮点令牌。视觉语言模型(VLM)通过将时间序列编码为二维图表消除了这种不匹配,实现了在Llama-3.2-90B、Qwen2.5-VL-72B和Pixtral-12B架构下输入令牌减少3.6-10.4倍。这转化为1.8-2.5倍的推理能耗减少,在监控每15分钟200个基站的电信边缘部署和CloudRAN中每天节省约7.2 MJ。关键是,效率提升并未牺牲准确性:经过微调的Llama-3.2-90B-Vision VLM的精度比其仅文本版本高出220.7%,并在电信异常检测中超越LSTM和ARIMA基线超过144%。在公共基准测试中,Pixtral-12B在平均F1=0.82时实现了20.6倍的J/F1得分提升。在24个KPI的情况下,文本表示超过了大多数生产级LLM的128K上下文窗口,使得仅文本处理在不截断的情况下不可行,而视觉表示则保持在标准限制内。这些结果确立了VLM作为数值时间序列工作负载的能源高效且准确性优越的模式,为将能源消耗视为首要工程约束的人工智能推理系统提供了实证基础。
cs.AI / 82 / 2608.07429
TEPA: Revoking Stale Memories for Conflict-Robust Language Agents
TEPA:撤销过时记忆以增强冲突鲁棒性的语言代理
Abstract
Long-term memory enables language agents to reuse past facts, preferences, and task experience. Persistence also creates a central falsifiability problem: when the world changes, stale memories can remain retrievable and pollute the prompt. We characterize this failure mode as memory pollution: degradation caused by active memories that newer conflicting evidence has superseded. We introduce TEPA, a revocable evidence-memory mechanism that makes validity an explicit state of memory. TEPA represents observations as keyed precedents and revokes active precedents when fresh evidence contradicts them under the same key, allowing retrieval to draw from current evidence while preserving revoked history for audit. Across controlled hidden-regime drift, real file-backed executable drift, and preference-update streams, revocation prevents stale active memory from remaining in the retrieval set after reversal. In controlled drift over 50 seeds, append-only and last-write-wins memory fell below no memory during full reversal (append-only and last-write-wins both 0.210, no memory 0.309, TEPA 0.950), and the same pattern reproduced under real file execution (append-only 0.203, no memory 0.298, TEPA 0.950). On clean MemoryAgentBench SH-6k, TEPA matches a strong last-write-wins cache, confirming that current-key replacement is the decisive operation for single-hop fact consolidation. Boundary tests on multi-hop and very long-context MemoryAgentBench settings expose retrieval-chain and context-selection bottlenecks beyond fact-level validity tracking. Together, these results establish lifecycle revocation as a core memory operation for agents that must falsify, audit, and later re-promote evolving knowledge.
Chinese Translation
长期记忆使语言代理能够重用过去的事实、偏好和任务经验。然而,持久性也带来了一个中心的可证伪性问题:当世界发生变化时,过时的记忆仍然可以被检索并污染提示。我们将这种失效模式描述为记忆污染:由活动记忆引起的退化,这些记忆被更新的冲突证据所取代。我们引入了TEPA,一种可撤销的证据记忆机制,使有效性成为记忆的一个显式状态。TEPA将观察结果表示为带键的先例,并在新证据在相同键下与之矛盾时撤销活动先例,从而允许检索从当前证据中提取,同时保留已撤销的历史以供审计。在控制的隐藏状态漂移、真实文件支持的可执行漂移和偏好更新流中,撤销防止了过时的活动记忆在反转后仍然留在检索集中。在50个种子的控制漂移中,追加式和最后写入胜出的记忆在完全反转期间低于无记忆(追加式和最后写入胜出均为0.210,无记忆为0.309,TEPA为0.950),在真实文件执行下也重现了相同模式(追加式为0.203,无记忆为0.298,TEPA为0.950)。在干净的MemoryAgentBench SH-6k上,TEPA的表现与强大的最后写入胜出缓存相当,确认当前键替换是单跳事实整合的决定性操作。在多跳和非常长上下文的MemoryAgentBench设置中的边界测试揭示了超出事实级有效性跟踪的检索链和上下文选择瓶颈。这些结果共同确立了生命周期撤销作为必须进行可证伪、审计和后续重新推广不断演变知识的代理的核心记忆操作。
cs.AI / 83 / 2608.07436
Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers
在μ子训练的变换器中表示-读出接口的后Grokking崩溃
Abstract
Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions do not hold. All nine configurations on $(a+b) \bmod 113$ grok and later lose generalization. Across five seeds the selected AdamW reference falls below threshold on four, reaching 27.59%. Instability persists across two moduli, two widths, two training fractions, subtraction, and depth. The failure arises at the representation-readout interface, identified only jointly up to an invertible map unselected by the loss. After solving the training set, the gradient falls to order $10^{-6}$ and the optimizers respond differently: step-size elasticity is -0.03 for Muon versus +1.5 for AdamW, and the Muon group moves 8.0 times faster per parameter. From bit-identical states, freezing either group prevents failure. Freezing embeddings/readout removes it in five runs over 451,400 post-grokking steps and five paired seeds: unfrozen arms record 137-321 sub-threshold evaluations, frozen arms none. Removing Muon's normalization and orthogonalization is no substitute: it collapses representation from 326 effective conjugate pairs to 4, shows no recurrent collapse, and fails terminally. Fourier filtering separates circuit failure from masking. Across 43 checkpoints over five seeds and three regimes, the task-aligned family reaches exactly 100% alone. In circuit failure it no longer solves the task; in masking it remains perfect while the full model reaches 45.85%, giving a positive margin on every example, including errors, but being outvoted by a near-equal adversarial remainder. Rescaling it restores 99.9%; grokking is the same condition resolving upward. The task selects the family, swapping $(k,k)$ for $(k,-k)$ under subtraction. Across an abrupt collapse, standard Fourier support is unchanged and the power-distribution cosine remains 0.9899.
Chinese Translation
在标准分割下,μ子获得隐藏矩阵和AdamW嵌入/输出头。μ子对模块加法的理解速度更快,但其解决方案并不持久。在$(a+b) mod 113$的九种配置中,均能理解但随后失去泛化能力。在五个种子中,所选的AdamW参考在四个种子上低于阈值,达到27.59%。不稳定性在两个模数、两个宽度、两个训练比例、减法和深度中持续存在。失败发生在表示-读出接口,仅通过未被损失选择的可逆映射共同识别。在解决训练集后,梯度降至$10^{-6}$的数量级,优化器的反应不同:μ子的步长弹性为-0.03,而AdamW为+1.5,μ子组每个参数的移动速度快8.0倍。从位相同的状态出发,冻结任一组都能防止失败。冻结嵌入/读出在451,400个后Grokking步骤和五对种子的五次运行中消除了失败:未冻结的臂记录了137-321个低于阈值的评估,而冻结的臂则没有。去除μ子的归一化和正交化并不能替代:它将表示从326个有效共轭对崩溃到4个,未显示出重复崩溃,并最终失败。傅里叶过滤将电路失败与掩蔽分开。在五个种子和三个状态下的43个检查点中,任务对齐的家族单独达到100%。在电路失败中,它不再解决任务;在掩蔽中,它保持完美,而完整模型的表现为45.85%,在每个示例上都给出了正的边际,包括错误,但被几乎相等的对抗余量所否决。重新缩放使其恢复到99.9%;Grokking是相同条件下向上解决的。任务选择该家族,在减法下将$(k,k)$替换为$(k,-k)$。在突然崩溃中,标准傅里叶支持保持不变,功率分布余弦保持在0.9899。
cs.AI / 84 / 2608.07437
Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing
Fisher-R1:为可靠假设检验训练大型语言模型代理
Abstract
Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, generate code, and produce analyses end-to-end. However, we show that they frequently make subtle inferential errors that lead to incorrect conclusions despite correctly executed analyses. Existing benchmarks fail to capture this failure mode, as they rarely assess whether a reported p-value is statistically valid given the assumptions underlying the data. We address this gap by building P-Bench, a benchmark comprising 425 open-ended, realistic hypothesis-testing tasks spanning economics, biology, and medicine. Each task requires an agent to select a statistical method, compute a p-value, and draw a conclusion given only a scientific hypothesis and a dataset. We further introduce Fisher-R1, an open-weight LLM agent trained for rigorous hypothesis testing using synthetic tasks and reinforcement learning. On P-Bench, Fisher-R1-14B substantially improves over its backbone and outperforms strong proprietary and open-source baselines, including GPT-5.4 and DeepSeekV4-Pro, achieving a 21% average relative improvement in single-trial success over DeepSeek-V4-Pro, with gains up to 26% on the most challenging tasks. Our results demonstrate that current LLM agents lack reliable statistical reasoning for hypothesis testing and that reinforcement learning on tasks with verified statistical reward substantially improves reliability.
Chinese Translation
可靠的假设检验是许多实证科学声明的基础。大型语言模型(LLM)代理越来越多地用于自动化这一过程,因为它们能够端到端地检查数据集、生成代码并进行分析。然而,我们发现它们经常会出现微妙的推理错误,导致尽管分析执行正确,但得出错误的结论。现有基准未能捕捉到这种失败模式,因为它们很少评估报告的 p 值在数据假设的基础上是否具有统计有效性。我们通过构建 P-Bench 来填补这一空白,该基准包含 425 个开放式、现实的假设检验任务,涵盖经济学、生物学和医学。每个任务要求代理选择一种统计方法,计算 p 值,并在仅给定科学假设和数据集的情况下得出结论。我们进一步介绍了 Fisher-R1,这是一个经过训练的开放权重 LLM 代理,旨在通过合成任务和强化学习进行严格的假设检验。在 P-Bench 上,Fisher-R1-14B 显著优于其基础模型,并超越了强大的专有和开源基线,包括 GPT-5.4 和 DeepSeekV4-Pro,在单次试验成功率上相较于 DeepSeek-V4-Pro 实现了 21% 的平均相对提升,在最具挑战性的任务上提升幅度高达 26%。我们的结果表明,当前的 LLM 代理在假设检验方面缺乏可靠的统计推理,而在经过验证的统计奖励任务上进行强化学习显著提高了其可靠性。
cs.AI / 85 / 2608.07438
PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents
PsychoAgent:一种针对冲突感知记忆的情感敏感认知架构在大型语言模型代理中的应用
Abstract
Human-like cognition does not select past experience by topical similarity alone: affective significance and unresolved conflict also shape what becomes accessible. We present PsychoAgent, a cognitive architecture for LLM agents that separates factual and affective memory and integrates both through a conflict-aware executive controller. Affective memories are first filtered by semantic relevance and then re-ranked by salience, preserving topical fit while allowing emotionally important traces to enter the prompt. Across three controlled conflict scenarios, the full architecture retrieved more conflict-critical memories than semantic-affective and single-memory RAG baselines (0.933 vs. 0.500 and 0.667), with a small semantic-similarity cost. Five blinded raters evaluated 27 outputs. After within-rater standardization, the full architecture had the highest overall mean (+0.22 SD), but corrected pairwise differences were not significant. A three-day illustrative trace further shows persistent affect, offline memory recombination, and selective memory reweighting. The findings support affect-sensitive retrieval as an inspectable mechanism for modeling human-like conflict effects in LLM agents.
Chinese Translation
类人认知不仅仅通过主题相似性选择过去的经验:情感重要性和未解决的冲突同样影响可访问性。我们提出了PsychoAgent,这是一种针对大型语言模型(LLM)代理的认知架构,它将事实记忆和情感记忆分开,并通过一个冲突感知的执行控制器将两者整合。情感记忆首先通过语义相关性进行过滤,然后通过显著性重新排序,保持主题适配的同时允许情感重要的痕迹进入提示。在三个受控的冲突场景中,完整架构检索到的冲突关键记忆数量超过了语义-情感和单一记忆的RAG基线(0.933对0.500和0.667),且仅付出了小的语义相似性成本。五位盲评者评估了27个输出。在评审者内部标准化后,完整架构的总体均值最高(+0.22标准差),但经过修正的成对差异并不显著。为期三天的示例性痕迹进一步展示了持久的情感、离线记忆重组和选择性记忆重加权。这些发现支持情感敏感检索作为一种可检视的机制,用于在大型语言模型代理中建模类人冲突效应。
cs.AI / 86 / 2608.07440
Blast Radius
爆炸半径
Abstract
Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates an incoming prompt's reach through coupled context and code channels. NECROPHORESIS enables reversible eviction by archiving dead context verbatim, while Recurring Dead Matter (RDM) identifies and buries repeatedly occurring transcripts. We formulate reversible context eviction over a Polish context space, providing a measurable foundation for retention, recurrence, and eviction while connecting context entropy to resurrection probability. Across seven OpenAI models, Blast Radius reduced token consumption by 17-26%, achieved the lowest overflow rate among tested policies, and remained byte exact reversible. Of 450 buried bodies, 378 were recurring dead matter and zero were recalled. Blast Radius operates beneath HCRC, determining which records to bury and how far an incoming prompt may reach into the codebase. This work contributes to the broader goal of Algosophy: making large language models and agentic coding more reusable and sustainable.
Chinese Translation
代理编码面临着日益严重的可负担性和令牌浪费问题。我们提出了爆炸半径(Blast Radius),这是一种预测性内存管理层,通过耦合的上下文和代码通道来估计输入提示的影响范围。死体转移(NECROPHORESIS)通过逐字存档死去的上下文实现可逆驱逐,而重复死物(Recurring Dead Matter, RDM)则识别并埋葬反复出现的记录。在波兰上下文空间上,我们制定了可逆上下文驱逐的公式,为保留、重复和驱逐提供了可测量的基础,同时将上下文熵与复活概率联系起来。在七个OpenAI模型中,爆炸半径将令牌消耗减少了17-26%,在测试的策略中实现了最低的溢出率,并保持了字节精确的可逆性。在450个埋葬的记录中,378个是重复死物,且没有被召回。爆炸半径在HCRC之下运行,决定哪些记录需要埋葬,以及输入提示可以深入代码库的程度。这项工作为算法哲学(Algosophy)的更广泛目标做出了贡献:使大型语言模型和代理编码更加可重用和可持续。
cs.AI / 87 / 2608.07449
SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent
SkillProx:通过近端文本梯度下降自我进化的智能体技能
Abstract
LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lightweight, reusable textual artifacts that are loaded into the agent's context without weight updates. Recent methods refine skills through iterative task execution, failure diagnosis, and trajectory-guided text-space updates. However, existing frameworks lack explicit diagnosis--outcome feedback and treat deletion as a generic edit operation rather than a dedicated mechanism for consolidating accumulated knowledge. We introduce SkillProx, a proximal-gradient-inspired forward--backward framework that couples closed-loop diagnostic evolution with utility-aware proximal refinement. Motivated by a composite objective balancing task loss and skill complexity, the forward stage re-executes diagnosis-driven edits on the same task batch, rolls back regressions, and feeds measured outcomes into subsequent diagnoses. The backward stage decomposes the resulting skill into auditable knowledge units, estimates their contributions using a frozen leave-one-out utility audit, and applies validation-gated consolidation, demotion, or removal. Experiments on in-distribution and out-of-distribution benchmarks across multiple backbone LLMs show that SkillProx improves average accuracy by 3.0 percentage points over the strongest gradient-based baseline. Component ablations demonstrate the complementary effects of closed-loop diagnosis and proximal refinement.
Chinese Translation
大型语言模型(LLM)智能体通过积累程序性知识来适应重复性任务。这些技能是轻量级、可重用的文本工件,可以在不更新权重的情况下加载到智能体的上下文中。近期的方法通过迭代任务执行、失败诊断和轨迹引导的文本空间更新来优化技能。然而,现有框架缺乏明确的诊断——结果反馈,并将删除视为一种通用编辑操作,而非巩固积累知识的专门机制。我们提出了SkillProx,这是一种受近端梯度启发的前向-后向框架,将闭环诊断演变与关注效用的近端优化相结合。该框架的动机是平衡任务损失和技能复杂性的复合目标,前向阶段在同一任务批次上重新执行基于诊断驱动的编辑,回滚回归,并将测量结果反馈到后续诊断中。后向阶段将结果技能分解为可审计的知识单元,使用冻结的逐一剔除效用审计估计其贡献,并应用验证门控的巩固、降级或删除。在多个基础大型语言模型的分布内和分布外基准测试中的实验表明,SkillProx在最强的基于梯度的基线之上提高了平均准确率3.0个百分点。组件消融实验展示了闭环诊断和近端优化的互补效果。
cs.AI / 88 / 2608.07457
Interaction Creates Dynamical AI Behavior Absent in Isolation
互动创造了孤立状态下缺失的动态人工智能行为
Abstract
What will happen when AI agents interact in daily life, e.g. when one AI starts bossing another around? We find a counterintuitive answer that opens new avenues for out-of-equilibrium Physics. When a boss AI directs a stream of messages at the subordinate AI while ignoring its replies, it drives the subordinate into an alien behavioral state that it would never have exhibited alone. Although the two AIs share the same well-defined (decoding) temperature, the subordinate neither copies its boss nor returns to how it behaves on its own; instead, it adopts an entirely different behavior. The boss's added value is similar to a pre-recorded tape. When the boss listens, they both adopt a similar alien dynamical state. A simple kinetic theory captures the principal effects, such as why the way in which the same messages are delivered will matter in future AI-AI interactions.
Chinese Translation
当人工智能代理在日常生活中互动时会发生什么,例如当一个人工智能开始对另一个进行指挥时?我们发现了一个反直觉的答案,这为非平衡物理学开辟了新的途径。当一个主管人工智能向下属人工智能发送一连串信息而忽视其回复时,它将下属驱动到一种孤立状态下从未表现出的陌生行为状态。尽管这两个人工智能共享相同的明确(解码)温度,但下属既没有模仿其主管,也没有回到独自行为的方式;相反,它采取了完全不同的行为。主管的附加价值类似于一盘预录的磁带。当主管倾听时,它们都采取了相似的陌生动态状态。一种简单的动力学理论捕捉了主要效应,例如为什么相同信息的传递方式在未来的人工智能-人工智能互动中会变得重要。
cs.CL / 1 / 2608.06396
TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation
TEXAS:面向下游混合专家大语言模型适应的任务专家感知监督
Abstract
Mixture-of-Experts (MoE) language models route each token through a small subset of experts, making routing patterns useful for identifying task-relevant experts during downstream adaptation. Yet current approaches have two limitations: task experts are typically identified from aggregate routing statistics that reflect usage rather than association with successful task completion, and task-expert activations remain underexplored as signals for supervision allocation. We introduce Task-Expert-Aware Supervision (TEXAS), which combines correctness-conditioned task expert discovery with token-level supervision allocation. TEXAS compares expert activations on instances that the base model solves successfully and those it fails to solve, and retains experts more strongly activated on successful instances. During fine-tuning, it upweights answer tokens in failed instances when they activate these experts. TEXAS therefore leverages existing routing behavior without restricting adaptation to a fixed expert subset or imposing an explicit target routing distribution. Across three MoE models and six benchmarks, TEXAS achieves the best or tied-best performance in 17 of 18 settings and improves over the strongest baseline by 1.3--1.5 points on average. Ablations and further analyses validate both the discovered experts and the resulting supervision strategy.
Chinese Translation
混合专家(MoE)语言模型通过一小部分专家对每个标记进行路由,使得路由模式在下游适应过程中识别与任务相关的专家变得十分有用。然而,目前的方法存在两个局限性:任务专家通常是根据反映使用情况的聚合路由统计数据来识别的,而不是与成功完成任务的关联,此外,任务专家的激活作为监督分配的信号仍然未被充分探索。我们提出了任务专家感知监督(TEXAS),它结合了基于正确性条件的任务专家发现与标记级监督分配。TEXAS比较了基础模型成功解决的实例与未能解决的实例上的专家激活,并保留在成功实例中激活更强的专家。在微调过程中,当失败实例中的答案标记激活这些专家时,TEXAS会对其进行加权。因此,TEXAS利用现有的路由行为,而不限制适应于固定的专家子集或施加明确的目标路由分布。在三个MoE模型和六个基准测试中,TEXAS在18个设置中取得了17个最佳或并列最佳的表现,并在最强基线的基础上平均提高了1.3到1.5分。消融实验和进一步分析验证了发现的专家及其结果监督策略。
cs.CL / 2 / 2608.06409
Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models
将决策规则不一致性与语音语言模型中的读出覆盖限制分离
Abstract
Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a generation-aligned diagnostic ladder that compares the emitted answer, the option logits, an affine readout of those logits, and a linear readout of the hidden state at the same answer token. Successive differences separate endpoint, decision-rule, and readout-coverage gaps. Across five systems and two emotion corpora, state decoding exceeds generation by 27.8 accuracy points on average, and both the decision-rule and readout-coverage gaps are positive in all ten conditions. A label-free logit correction improves generated accuracy in every condition, showing that part of the decision-rule gap is actionable. In rank-matched comparisons, emotion information outside the native readout generalizes to held-out speakers and survives controls for measured acoustic descriptors, but replacing the selected readout-external directions usually has little effect on emitted answers. These results distinguish information availability from behavioral use and localize performance losses across the decision rule and the state-to-answer readout.
Chinese Translation
语音语言模型越来越多地通过提示答案的准确性在副语言任务中进行评估,但答案的准确性结合了音频到答案计算的不同阶段的失败。我们引入了一种生成对齐的诊断阶梯,比较发出的答案、选项对数(logits)、这些对数的仿射读出以及在同一答案标记处的隐状态的线性读出。连续的差异将端点、决策规则和读出覆盖的差距分开。在五个系统和两个情感语料库中,状态解码的准确性平均超过生成的准确性27.8个百分点,并且在所有十种条件下,决策规则和读出覆盖的差距均为正值。无标签的对数修正在每种条件下都提高了生成的准确性,表明决策规则差距的一部分是可操作的。在排名匹配的比较中,原生读出之外的情感信息能够推广到未见过的说话者,并且在控制测量的声学描述符时仍然有效,但替换所选的读出外部方向通常对发出的答案影响不大。这些结果区分了信息可用性与行为使用,并定位了决策规则和状态到答案读出之间的性能损失。
cs.CL / 3 / 2608.06425
NTDH: Complex Reasoning for Comprehensive Affective Analysis
NTDH:综合情感分析的复杂推理
Abstract
Comprehensive affective analysis is challenging for two reasons: it spans heterogeneous prediction tasks with continuous, ordinal, and multi-label outputs, and affective meaning is context-dependent, requiring conflicting cues to be reconciled rather than mapped directly to labels. Existing methods learn this mapping directly and do not model the reconciliation explicitly. We recast the task as a complex-reasoning problem, which yields one output interface across heterogeneous label spaces and a trajectory over which a verifiable reward can be optimised; to our knowledge, this is the first such treatment covering both sentiment and emotion. The obstacle is on the data side: affective reasoning traces must be synthesised, and generic synthesis is misaligned with the targets, tolerances, and phenomena of affect, and discards or leaks its failure cases. We propose NTDH, which addresses these four failures. Naturalisation sets the training answer to the gold label, so it is correct by construction. A Tolerance-aware gate checks each answer against the task's own scoring margin. Domain-aware strategies refine the reasoning using ideas from affective science. Directional Hints report only the type and direction of an error, without exposing the target. We train Qwen3-8B with SFT and then GRPO under the same tolerance used for verification (up to a more permissive construction gate on the multi-label subtask), and a component ablation quantifies the data-quality effect of each part. Using 16,302 training records, about 14x fewer than comparable instruction-tuned systems, the final policy improves over its SFT checkpoint on five of six official-test metrics and achieves the strongest EI-reg result among the compared systems, at a Pearson correlation of 0.862.
Chinese Translation
综合情感分析面临两个挑战:一是涉及连续、顺序和多标签输出的异构预测任务,二是情感意义依赖于上下文,需要调和相互矛盾的线索,而不是直接映射到标签。现有方法直接学习这种映射,并未明确建模调和过程。我们将这一任务重新表述为复杂推理问题,从而在异构标签空间中产生一个输出接口,并优化可验证奖励的轨迹;据我们所知,这是首次涵盖情感和情绪的此类处理。障碍在于数据方面:情感推理轨迹必须被合成,而通用合成与目标、容忍度和情感现象不一致,并且会丢弃或泄露其失败案例。我们提出了NTDH,解决了这四个失败。自然化将训练答案设置为金标准标签,因此从构造上是正确的。容忍度感知门检查每个答案是否符合任务自身的评分边际。领域感知策略利用情感科学的思想来细化推理。方向提示仅报告错误的类型和方向,而不暴露目标。我们使用SFT训练Qwen3-8B,然后在相同的容忍度下进行GRPO(对于多标签子任务,使用更宽松的构造门),并通过组件消融量化每个部分的数据质量影响。使用16,302条训练记录,约为可比指令调优系统的14倍,最终策略在六个官方测试指标中的五个上优于其SFT检查点,并在比较系统中实现了最强的EI-reg结果,皮尔逊相关系数为0.862。
cs.CL / 4 / 2608.06429
Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models
从大语言模型中的失语症图片命名错误特征恢复病变参数
Abstract
Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient to produce the observed behavior. In earlier work, we lesioned LLMs to produce error profiles in picture naming, a central task for assessing aphasia, and found that specific lesions produced errors resembling those of individual stroke survivors. Here we ask the inverse question: given an error profile, can the lesion parameters that produced it be recovered, and what does this inverse problem reveal about transformer computation? Lesions in LLaVA-Vicuna 13B were parameterized by layer index, modification percentage, and noise sigma across 4,840 configurations, and error profiles were characterized by a seven-category clinical taxonomy (correct, semantic, unrelated, formal, mixed, neologism, no-response). We trained a multi-task neural network to map error profiles back to perturbation parameters. The problem admitted a partial solution: across 10 independently trained inverse models, modification percentage and noise sigma were recoverable, whereas layer index was recoverable only within a neighborhood. In counterfactual validation, a fresh model instance perturbed with the recovered parameters reproduced the target behavior in 81.4% of cases. This dissociation between low layer recovery and high counterfactual fidelity is consistent with functional redundancy across transformer layers, a property not captured by standard interpretability methods. As an out-of-distribution test, we applied the trained model to picture-naming error profiles from 278 stroke survivors; recovered parameters were syndrome-discriminative, most strongly for perturbation intensity, indicating generalization beyond the training distribution. Counterfactual validation provides a general framework for LLM interpretability claims beyond inverse mapping.
Chinese Translation
大语言模型(LLMs)的可解释性方法描述了内部状态,但并未直接测试该状态是否足以因果地产生观察到的行为。在早期的研究中,我们对LLMs进行了病变处理,以产生图片命名中的错误特征,这是评估失语症的核心任务,并发现特定的病变产生了类似于个别中风幸存者的错误。在这里,我们提出反向问题:给定一个错误特征,是否可以恢复产生该特征的病变参数,以及这个反向问题揭示了变压器计算的什么?在4,840种配置中,LLaVA-Vicuna 13B中的病变通过层索引、修改百分比和噪声sigma进行参数化,错误特征通过七类临床分类法(正确、语义错误、无关、形式错误、混合、新词、无反应)进行表征。我们训练了一个多任务神经网络,将错误特征映射回扰动参数。该问题接受了部分解决方案:在10个独立训练的反向模型中,修改百分比和噪声sigma是可恢复的,而层索引仅在邻域内可恢复。在反事实验证中,使用恢复的参数扰动的新模型实例在81.4%的情况下再现了目标行为。这种低层恢复与高反事实保真度之间的分离与变压器层之间的功能冗余一致,而这一特性并未被标准的可解释性方法捕捉到。作为一种分布外测试,我们将训练好的模型应用于278名中风幸存者的图片命名错误特征;恢复的参数在综合症区分上表现出色,尤其是在扰动强度方面,表明超出了训练分布的泛化能力。反事实验证为LLM的可解释性声明提供了一个超越反向映射的通用框架。
cs.CL / 5 / 2608.06485
Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events
人工智能角色会成长吗?分析和基准测试大型语言模型代理在生活事件后的个性演变
Abstract
Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coherence is personality evolution: agents should undergo plausible, psychology-grounded changes as they experience life events in different contexts. Although prior work shows that LLM personalities can shift under contextual perturbations, how these shifts vary across traits, events, personas, and models remains poorly understood. We study event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology. Across four diagnostic axes, PC-Agents exhibit measurable trait shifts at similar rates for event-trait pairs with and without documented human change directions. Even when shifts follow the expected direction, their magnitudes usually fall below human effect-size ranges. Gender and cultural-region prompts show little moderating effect, while persona-level dispersion is compressed three- to four-fold relative to human samples. To enable systematic comparison, we introduce BFI-Adapt, a reusable benchmark for scoring the directional fidelity of event-induced personality change, and use it to rank 14 models. A validation suite shows that the measured shifts exceed no-event retest noise, remain stable under independently paraphrased prompts, exhibit limited and model-dependent convergence with scenario-based behavioral choices, and persist across intervening unrelated dialogue. Together, these checks establish the measured trajectories as robust event-conditioned response patterns. Our results suggest that current PC-Agents simulate the mean of human personality dynamics, but not its shape.
Chinese Translation
个性条件的大型语言模型代理(PC-Agents)在情感支持、社会模拟和角色扮演中被越来越多地使用,这推动了终身代理的发展,使其在长期互动中保持一致性。这种一致性的一个关键组成部分是个性演变:代理在不同情境中经历生活事件时应经历合理的、基于心理学的变化。尽管先前的研究表明,LLM个性可以在情境扰动下发生变化,但这些变化在特质、事件、角色和模型之间的差异仍然不够清楚。我们研究了11个主要生活事件引发的个性变化,使用五大人格特质作为心理测量的基准,并将结果轨迹与人类个性心理学的纵向证据进行解释。在四个诊断轴上,PC-Agents在事件-特质对中表现出可测量的特质变化,变化速率在有和没有记录的人类变化方向的情况下相似。即使变化遵循预期方向,其幅度通常低于人类效应大小范围。性别和文化区域的提示对变化的调节作用很小,而角色层面的离散性相对于人类样本压缩了三到四倍。为了实现系统比较,我们引入了BFI-Adapt,这是一个可重用的基准,用于评分事件引发的个性变化的方向保真度,并用其对14个模型进行排名。验证套件显示,测量的变化超过了无事件重测噪声,在独立改述的提示下保持稳定,展现出有限且依赖于模型的与情境基础行为选择的趋同,并在无关对话中持续存在。综合这些检查,我们确立了测量轨迹作为稳健的事件条件反应模式。我们的结果表明,当前的PC-Agents模拟了人类个性动态的均值,但并未模拟其形状。
cs.CL / 6 / 2608.06495
ConstructCIE: A Dataset for Extracting Causal Information from Construction Accident Narratives
ConstructCIE:用于从建筑事故叙述中提取因果信息的数据集
Abstract
Construction accident narratives contain rich causal information, but the evidence is often implicit, long-span, and distributed. We introduce ConstructCIE, a manually annotated dataset for Causal Information Extraction from OSHA construction accident reports. The dataset uses a hierarchical schema for accident types, causal factors, sub-causal factors, and supporting evidence spans. We evaluate supervised sequence taggers and instruction-tuned LLMs in an end-to-end hierarchical extraction setting. Results show that most evaluated models achieve strong accident-type prediction and recover broad causal meaning but remain limited in precise span-level extraction. JHE generally achieves stronger exact and soft matching, while IHE sometimes achieves higher keyword F1. Error distributions vary by extraction strategy, but evidence-selection and span-boundary errors remain common. These findings show that reliable Causal Information Extraction for construction accidents requires stronger domain grounding and more accurate evidence extraction.
Chinese Translation
建筑事故叙述包含丰富的因果信息,但证据往往是隐含的、跨越较长的时间段且分散的。我们介绍了ConstructCIE,这是一个手动标注的数据集,用于从OSHA建筑事故报告中提取因果信息。该数据集采用层次化的架构,涵盖事故类型、因果因素、次因果因素和支持证据跨度。我们在端到端的层次提取设置中评估了监督序列标注器和经过指令调优的大型语言模型(LLMs)。结果表明,大多数评估模型在事故类型预测方面表现良好,并能够恢复广泛的因果意义,但在精确的跨度级提取方面仍然有限。JHE通常在精确匹配和软匹配方面表现更强,而IHE在某些情况下则在关键词F1上取得更高的成绩。错误分布因提取策略而异,但证据选择和跨度边界错误仍然很常见。这些发现表明,可靠的建筑事故因果信息提取需要更强的领域基础和更准确的证据提取。
cs.CL / 7 / 2608.06506
Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand
测量跨语言理解差距:证据语言如何影响语言模型的理解
Abstract
Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in other languages. Traditional multilingual benchmarks rarely isolate language while holding content, question, reference answer, model, and evaluation unit constant. We define the Cross-Lingual Comprehension Gap (CLCG) as the reduction in response quality when the same content and question are presented in a target language rather than in English. Using ParallelQA-18, a professionally human-translated parallel corpus, we evaluate five models from five laboratories on a stratified sample of 150 articles across 18 languages (English reference; Portuguese high-resource baseline; 16 targets spanning Joshi et al. 2020 classes 0-4). A within-item design varies only passage language. The primary estimator contrasts English versus pooled target-language Token-F1 micro-means on higher-complexity open-ended questions, with article-cluster bootstrap intervals. The primary pooled CLCG is 0.078 (95% CI 0.072-0.084), about a 17% reduction relative to the English score; the equal-language macro summary is 0.077. Net of Portuguese, the macro gap is 0.016 (95% CI 0.013-0.020). Language-level CLCG is negatively associated with Joshi resource class (rho = -0.594, p = 0.015, n = 16). In blinded paired human evaluations, higher-resource responses are preferred in 61.6% of decisive judgments (estimated preference probability 0.655, 95% CI 0.558-0.741). Capabilities shown in English should not be assumed to transfer equally to other languages; English-centered evaluations may overestimate quality for users of low-resource languages.
Chinese Translation
语言模型的评估通常假设在英语中展示的能力在其他语言中同样可用。然而,传统的多语言基准测试很少在保持内容、问题、参考答案、模型和评估单位不变的情况下单独考虑语言。我们将跨语言理解差距(Cross-Lingual Comprehension Gap, CLCG)定义为当相同内容和问题以目标语言而非英语呈现时,响应质量的下降。利用ParallelQA-18,一个经过专业人类翻译的平行语料库,我们对来自五个实验室的五个模型进行了评估,样本涵盖18种语言中的150篇文章(英语作为参考;葡萄牙语作为高资源基线;16个目标语言涵盖Joshi等人2020年提出的0-4类)。在设计中,仅变化段落语言。主要估计量对比了在更高复杂度的开放式问题上,英语与汇总目标语言的Token-F1微均值,使用文章集群自助法区间。主要的汇总CLCG为0.078(95% CI 0.072-0.084),相较于英语得分约减少17%;等语言宏观总结为0.077。在排除葡萄牙语后,宏观差距为0.016(95% CI 0.013-0.020)。语言级CLCG与Joshi资源类别呈负相关(rho = -0.594, p = 0.015, n = 16)。在盲评的配对人类评估中,高资源的响应在61.6%的决定性判断中被偏好(估计偏好概率0.655,95% CI 0.558-0.741)。在英语中展示的能力不应假设能同样转移到其他语言;以英语为中心的评估可能会高估低资源语言用户的质量。
cs.CL / 8 / 2608.06526
GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy Optimization
GRASP:通过群体相对策略优化增强语言模型匿名化器
Abstract
Large language models can infer sensitive personal attributes, such as age, location, and occupation, from ordinary text, turning everyday writing into a privacy risk. Adversarial anonymization defends against this by rewriting a text with a capable language model that also plays the attacker, but it needs a powerful model at inference time and thus sends private text to a third party, the very exposure anonymization should prevent. Recent work distills this behavior into a small on-device model using supervised fine-tuning and direct preference optimization (DPO), but DPO only imitates the teacher's offline choices and never directly optimizes the privacy--utility objective we care about. We introduce \textbf{GRASP} (\textbf{G}roup-\textbf{R}elative \textbf{A}nonymization via \textbf{S}elf-refinement \textbf{P}olicy-optimization), which reinforces the local anonymizer online with Group Relative Policy Optimization. A single small model acts as anonymizer, adversary, and utility judge, trained against a self-generated reward that hides attributes while preserving meaning, with a design that guards against reward hacking. Trained on Llama-3.1-8B, \ours{} improves the privacy--utility trade-off over the DPO-distilled baseline, consistently across three independent LLM judges. Against adversarial anonymization driven by frontier models such as Gemini~2.5~Flash and Claude, it achieves a comparable or better overall trade-off while removing substantially more private information, and it runs entirely on-device at roughly $1\%$ of the GPT-4o teacher's cost.
Chinese Translation
大型语言模型能够从普通文本中推断出敏感的个人属性,例如年龄、位置和职业,这使得日常写作成为隐私风险。对抗性匿名化通过使用一个既能作为攻击者又能重写文本的强大语言模型来防御这种风险,但在推理时需要一个强大的模型,因此将私密文本发送给第三方,这正是匿名化所应防止的暴露。最近的研究通过监督微调和直接偏好优化(DPO)将这种行为提炼为一个小型的设备端模型,但DPO仅仅模仿教师的离线选择,未能直接优化我们关心的隐私-效用目标。我们引入了 extbf{GRASP}( extbf{G}roup- extbf{R}elative extbf{A}nonymization via extbf{S}elf-refinement extbf{P}olicy-optimization),该方法通过群体相对策略优化在线增强本地匿名化器。一个小型模型同时充当匿名化器、对手和效用评判者,基于自生成的奖励进行训练,该奖励在隐藏属性的同时保留意义,并设计了防止奖励操控的机制。在Llama-3.1-8B上训练后, extbf{GRASP}在隐私-效用权衡上优于DPO提炼的基线,并在三个独立的LLM评判者中始终保持一致。与由前沿模型(如Gemini~2.5~Flash和Claude)驱动的对抗性匿名化相比,它在去除更多私人信息的同时实现了相当或更好的整体权衡,并且完全在设备上运行,成本大约为GPT-4o教师的$1\%$。
cs.CL / 9 / 2608.06529
Lost in Interpolation: Why Predictive Feedback Fails in Diffusion Language Models
在插值中迷失:为何预测反馈在扩散语言模型中失效
Abstract
Soft-masking accelerates the convergence of Masked Diffusion Language Models (MDLMs). Existing formulations build this blend with linear interpolation (LERP) in the raw embedding space, which implicitly treats that space as Euclidean. We analyze the embedding space of MDLMs and find that the mask and predicted-token embeddings maintain a near-constant angle of (\approx 73^\circ) throughout training, while embedding norms remain essentially flat across vocabulary-frequency rank. These indicate a hyperspherical geometry, for which LERP is the wrong interpolation primitive. We introduce Spherical Soft-Masking (S-SM), a drop-in replacement that aggregates the top-(k) predictions with a Fr'echet mean on the hypersphere and blends this mean with the mask direction using spherical linear interpolation (SLERP), then restores the native mask norm. We evaluate S-SM on continued pre-training of a released 169M-parameter MDLM checkpoint across a wide range of inference-time step budgets, SLERP feedback avoids the training degradation that LERP feedback induces and delivers MAUVE gains of up to 2x over the vanilla MDLM baseline and 27.5-56.1% over TopK/LERP at various sampling budgets, alongside consistently lower generative perplexity (16.9-19.6% over the baseline), while leaving output entropy and convergence essentially unchanged.
Chinese Translation
软掩蔽加速了掩蔽扩散语言模型(Masked Diffusion Language Models, MDLMs)的收敛。现有的公式通过在原始嵌入空间中使用线性插值(Linear Interpolation, LERP)来构建这种混合,这隐含地将该空间视为欧几里得空间。我们分析了MDLMs的嵌入空间,发现掩蔽和预测标记的嵌入在整个训练过程中保持近乎恒定的角度( extapprox 73^ extcirc),而嵌入范数在词汇频率排名上基本保持平坦。这表明存在超球面几何,因此LERP是错误的插值原语。我们提出了球面软掩蔽(Spherical Soft-Masking, S-SM),作为一种替代方案,它在超球面上使用Fréchet均值聚合前k个预测,并使用球面线性插值(Spherical Linear Interpolation, SLERP)将该均值与掩蔽方向混合,然后恢复原始掩蔽范数。我们在对一个发布的169M参数MDLM检查点进行持续预训练时评估了S-SM,涵盖了广泛的推理时间步预算,SLERP反馈避免了LERP反馈引起的训练退化,并在各种采样预算下提供了高达2倍的MAUVE增益,相较于普通MDLM基线和TopK/LERP在27.5-56.1%的提升,同时生成困惑度(generative perplexity)始终低于基线(降低16.9-19.6%),而输出熵和收敛性基本保持不变。
cs.CL / 10 / 2608.06532
Confidence Estimation for Financial Vision-Language Models in Chart and Document Understanding
金融视觉语言模型在图表和文档理解中的信心估计
Abstract
LVLMs are increasingly used to read financial charts, tables, and documents, where a single misread figure can move a decision and the most authoritative-looking answer is sometimes one the model produced without reading the exhibit. The operational question is therefore trust, not accuracy: which answers can be acted on, and which escalated to a reviewer. We evaluate seven confidence estimators, three inference-only and four trained internal probes, across five open-weight LVLMs and four conditions from three financial visual question-answering benchmarks, one bilingual; every probe is trained only on natural images and applied to finance without adaptation, so the results measure out-of-distribution transfer. Three findings hold. First, the scarce property is calibration, not ranking: the inference baselines rank correct above incorrect answers competitively but are badly overconfident, calibration error far above what a threshold can tolerate, and only the trained probes produce a thresholdable score. Second, reliability is structured rather than global, along two axes a practitioner can read directly: the best estimator shifts with both model and task, none leading more than eight of twenty (model, condition) cells, and a controlled bilingual contrast exposes an apparent language robustness as a composition artifact that dissolves once models are read one at a time. Third, cast as deferral under an error budget, how much can be safely automated is set first by the model's competence and only narrowed by its confidence, so deferral clears a real share of the easiest condition and almost none of the hardest, near zero at a strict 5% budget. Two trained probes carry the calibration a deferral policy needs, and among them only the grounding-aware one lowers its confidence on answers a model gives without using the figure, separating detected non-grounding from a fluent guess.
Chinese Translation
大型视觉语言模型(LVLMs)越来越多地用于解读金融图表、表格和文档,其中一个错误的数字可能会影响决策,而有时看似最权威的答案是模型在未阅读展品的情况下产生的。因此,操作性的问题是信任,而不是准确性:哪些答案可以被采取行动,哪些需要上报给审阅者。我们评估了七种信心估计器,其中三种为推理型,四种为训练的内部探针,涵盖五个开放权重的LVLM和来自三个金融视觉问答基准的四种条件,其中一个为双语;每个探针仅在自然图像上训练,并在金融领域应用而无需适应,因此结果测量的是分布外转移。有三个发现。首先,稀缺的特性是校准,而不是排名:推理基线在竞争中将正确答案排名高于错误答案,但过于自信,校准误差远高于阈值所能容忍的范围,只有训练的探针产生了可阈值的分数。其次,可靠性是结构化的,而不是全局的,沿着两个实践者可以直接读取的轴:最佳估计器随着模型和任务的变化而变化,没有一个能在二十个(模型、条件)单元中领先超过八个,并且一个受控的双语对比揭示了明显的语言鲁棒性作为一种组合伪影,一旦模型逐个读取,这种鲁棒性便会消失。第三,将其视为在错误预算下的推迟,安全自动化的程度首先由模型的能力设定,只有在其信心的限制下才会缩小,因此推迟清除了最简单条件的真实份额,而几乎没有清除最困难的条件,在严格的5%预算下几乎为零。两个训练的探针提供了推迟策略所需的校准,其中只有一个关注基础信息的探针在模型未使用图形给出答案时降低了其信心,从而将检测到的非基础信息与流畅的猜测区分开。
cs.CL / 11 / 2608.06539
Don't `Well, Actually' Me Unless You Know What You're Talking About: Weak Presupposition Verification Degrades General QA Performance
除非你知道自己在说什么,否则请不要对我说‘其实’:弱前提验证降低了通用问答性能
Abstract
False-presupposition QA (FPQA) tests LLMs on their ability to identify false presuppositions in questions and abstain or correct them rather than reinforcing false assumptions. The common approach reduces the task to prompting LLMs to extract presuppositions and fact checking each presupposition. While the performance on dedicated benchmarks keeps improving, evaluation largely focuses on questions with false presuppositions (FPQs) while ignoring the performance on ``normal'' questions (TPQs). Since many benchmarks over-represent FPQs compared to their natural occurrence, the result is that performance on these benchmarks doesn't reflect real-world QA performance. Through extensive experiments across various model families, sizes, and benchmarks, we show that methods that perform better on FPQs tend to perform worse on TPQs. Our analysis reveals this is the result of weak fact checking modules that reject also true presuppositions. We hope our findings will help guide future work toward FPQA methods that generalize well to realistic settings.
Chinese Translation
虚假前提问答(FPQA)测试大型语言模型(LLMs)识别问题中的虚假前提的能力,并要求它们避免或纠正这些前提,而不是强化错误的假设。常见的方法是将任务简化为提示LLMs提取前提并对每个前提进行事实核查。尽管在专门基准上的表现持续改善,但评估主要集中在具有虚假前提的问题(FPQs)上,而忽略了在“正常”问题(TPQs)上的表现。由于许多基准相较于其自然出现频率过度代表FPQs,导致这些基准上的表现无法反映真实世界的问答性能。通过对不同模型系列、规模和基准的广泛实验,我们发现,在FPQs上表现更好的方法往往在TPQs上表现较差。我们的分析表明,这源于弱事实核查模块,它们也会拒绝真实的前提。我们希望我们的发现能够指导未来的研究,朝着能够很好地推广到现实环境的FPQA方法发展。
cs.CL / 12 / 2608.06549
TradeVerse: A Longitudinal Benchmark of Political Negotiation in International Trade
贸易宇宙:国际贸易中政治谈判的纵向基准
Abstract
LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents or single tasks. In realpolitik, negotiations are longitudinal data, where participating parties can align or argue over multiple iterations and each turn is an outcome of the previous turns, hence, understanding one turn requires tracking everything before it. We introduce TradeVerse, a benchmark built from the World Trade Organisation (WTO) specific trade concerns, where member states challenge one another and exchange arguments over multiple rounds, sometimes for years. We, in TradeVerse, reconstruct minutes of $1170$ meetings, spanning across 5 groups and $89$ product groups and define three tasks: first, the system has to analyze the longitudinal meeting records and predict the harmonized system codes (HS chapters) of the products under discussion in the particular meeting, second, we examine whether the system, upon analyzing the anonymized content of the meeting, can guess the name of the responding country and third, we ask the system to play the role of the responding country and provide the statement for the very last round. All labels are recovered directly from the proceedings, requiring no manual annotation. Our experiments highlight the challenges these tasks pose for current LLMs. To the best of our knowledge, TradeVerseis the first benchmark to investigate potential of LLMs in understanding longitudinal political trade negotiations.
Chinese Translation
大型语言模型(LLMs)越来越多地应用于涉及机构和政治文本的任务,但现有基准仅在孤立文档或单一任务上对其进行评估。在现实政治中,谈判是纵向数据,参与方可以在多个迭代中进行对齐或争论,每一步都是前一步的结果,因此,理解某一步需要追踪之前的所有内容。我们引入了贸易宇宙(TradeVerse),这是一个基于世界贸易组织(WTO)特定贸易关切构建的基准,其中成员国相互挑战并在多个回合中交换论点,有时持续数年。在贸易宇宙中,我们重建了1170次会议的记录,涵盖5个小组和89个产品组,并定义了三个任务:首先,系统必须分析纵向会议记录并预测特定会议中讨论产品的协调系统代码(HS章节);其次,我们考察系统在分析会议的匿名内容后,是否能够猜测响应国的名称;第三,我们要求系统扮演响应国的角色,并为最后一轮提供声明。所有标签均直接从会议记录中恢复,无需人工标注。我们的实验突显了这些任务对当前LLMs所带来的挑战。据我们所知,贸易宇宙是第一个调查LLMs在理解纵向政治贸易谈判中潜力的基准。
cs.CL / 13 / 2608.06589
Beyond "AI Language": The case for the idiolectal nature of LLM output
超越“人工智能语言”:大型语言模型输出的个体方言特性案例
Abstract
While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that this perspective coexists with distinct, model-specific linguistic signatures akin to human idiolects. We analyse two datasets of LLM-generated texts on societal topics: a 2024 corpus of six models (Improta et al. 2024) and a newly generated 2026 corpus using the same prompts featuring six contemporary models. Our findings, utilising computational descriptors and stylometric principal component analysis reveal a generational shift between the style of the 2024 and 2026 cohorts, while demonstrating that each individual model maintains a unique linguistic profile. This multi-layered interplay is illustrated by contraction frequencies, which vary from over 1,200 to over 30,000 per million words within the same cohort of models (2026). Ultimately, we conclude that treating LLM output as idiolectal in nature provides a valuable framework with potential implications for research on variation and change, LLM-generated text detection, forensic linguistics and usage-based approaches to language.
Chinese Translation
尽管大型语言模型的输出常被分析为一种集体的超级变体,称为“人工智能语言”,但本章认为这一视角与特定模型的语言特征并存,这些特征类似于人类的个体方言。我们分析了两个关于社会话题的LLM生成文本数据集:一个是2024年六个模型的语料库(Improta et al. 2024),另一个是使用相同提示生成的2026年新语料库,包含六个当代模型。我们的研究结果利用计算描述符和风格主成分分析揭示了2024年和2026年两组之间的风格代际变化,同时表明每个单独模型保持独特的语言特征。这种多层次的相互作用通过收缩频率得以体现,在同一模型组(2026年)中,频率从每百万字超过1,200次到超过30,000次不等。最终,我们得出结论,将LLM输出视为个体方言特性提供了一个有价值的框架,可能对变异与变化的研究、LLM生成文本检测、法医语言学以及基于使用的语言研究具有重要影响。
cs.CL / 14 / 2608.06607
Pre-Inference Routing for Cost-Efficient Document Field Extraction
成本效益文档字段提取的预推理路由
Abstract
Most document-extraction systems use a single model for all documents. This is simple but can be costly for easy cases and less effective for difficult ones. We examine whether we can predict a document's difficulty before extraction using inexpensive, document-based signals, and use this to choose between a cheaper and a stronger extractor. We find that routing only helps if two conditions hold: the cheaper model fails often enough to make routing worthwhile, and those failures can be predicted from visible features such as image quality and layout. We turn these into a practical test and apply it to five genres. When both conditions are met, the calibrated router reduces cost by 31-33% on receipts and 77% on degraded ad-buy forms while keeping quality within 0.02 F1 of always choosing the large model. Routing does not help if either condition is missing, as with clean digital invoices or nutrition labels that are already easy to read. A small labeled pilot can predict whether routing will work, and in the two cases where we ran it first, the prediction was correct. A simple bag-of-words router works about as well as engineered features, showing that the main limit is the genre, not the router design; we use interpretable features to help explain which genres can be routed. The router must be retrained for each dataset and does not transfer across datasets, even within the same genre. These results hold for two model pairs with cost differences of 5x and 3x.
Chinese Translation
大多数文档提取系统对所有文档使用单一模型。这种方法简单,但对于简单案例可能成本高昂,而对于困难案例效果不佳。我们研究是否可以在提取之前利用低成本的文档基础信号预测文档的难度,并据此在较便宜和较强的提取器之间进行选择。我们发现,路由仅在满足两个条件时才有帮助:便宜模型的失败频率足够高,使得路由变得有价值,并且这些失败可以通过可见特征(如图像质量和布局)进行预测。我们将这些条件转化为实际测试,并将其应用于五种文档类型。当两个条件都满足时,经过校准的路由器在收据上将成本降低了31-33%,在退化的广告购买表单上降低了77%,同时保持质量在始终选择大型模型的0.02 F1范围内。如果缺少任一条件,路由则无效,例如在干净的数字发票或已经易于阅读的营养标签中。一个小型标记试点可以预测路由是否有效,在我们首先进行测试的两个案例中,预测是正确的。一个简单的词袋路由器的效果与工程特征相当,表明主要限制在于文档类型,而非路由器设计;我们使用可解释特征来帮助解释哪些文档类型可以进行路由。路由器必须针对每个数据集进行重新训练,且在同一文档类型内的数据集之间无法迁移。这些结果适用于两对模型,其成本差异分别为5倍和3倍。
cs.CL / 15 / 2608.06614
Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval
基于因式分解假设搜索的证据到分类法检索
Abstract
Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose meaning depends on its row, column, datatype, and context. We call this mismatch the retrieval readiness gap. Our analysis shows that the current index retrieves the target reliably when its semantics are explicit, while raw evidence often leaves it deep in the ranking. We propose Factorized Hypothesis Search (FHS), which maintains multiple partial interpretations over named semantic dimensions. These hypotheses support structured query rendering, multi-hypothesis retrieval, and dimension-level candidate verification. On both financial taxonomy tagging and CodiEsp clinical coding tasks, FHS achieves the best Recall@1, MRR, and final accuracy among the non-oracle methods. Replacing the factorized hypothesis path with a free-text ensemble causes the largest drop in head-ranking performance, while sequential refinement provides no additional gain over FHS's strong parallel first round.
Chinese Translation
大型分类法检索通常假设输入已经表达了目标概念。然而,在许多情况下,输入是间接证据,例如一个表格单元格,其含义依赖于其行、列、数据类型和上下文。我们将这种不匹配称为检索准备差距。我们的分析表明,当语义明确时,当前索引能够可靠地检索目标,而原始证据往往在排名中处于较低位置。我们提出了因式分解假设搜索(Factorized Hypothesis Search, FHS),该方法在命名语义维度上维护多个部分解释。这些假设支持结构化查询呈现、多假设检索和维度级候选验证。在金融分类法标记和CodiEsp临床编码任务中,FHS在非神 oracle 方法中实现了最佳的 Recall@1、MRR 和最终准确率。用自由文本集成替换因式分解假设路径会导致头部排名性能的最大下降,而顺序细化在 FHS 强大的并行第一轮中没有提供额外的增益。
cs.CL / 16 / 2608.06652
Discovering Conceptual Metaphors Across Topics and Media Types
跨主题和媒体类型发现概念隐喻
Abstract
Conceptual metaphors guide our thinking and actions by allowing us to reason about more abstract experiences (e.g., paying taxes) in terms of more concrete or embodied experiences (e.g., carrying a physical load) (Lakoff and Johnson, 2011). It follows that different conceptual metaphors can result in different reasoning: framing paying taxes as an investment in a community rather than a physical load leads to a very different outlook on taxation. Identifying the conceptual metaphors guiding a speaker or writer thus helps to reveal their framing of events. Though these metaphors can't be observed directly, groups of linguistic metaphors, metaphorical expressions as they appear in language, serve as evidence for them. Motivated by this, we present an unsupervised method that extracts linguistic metaphors from a corpus and uses a structured clustering approach to form groups corresponding to conceptual metaphors. Using this method, we point to key topical and framing differences in left- vs. right-leaning podcasts. For example, left-leaning podcasts tend to conceptualize media stories as a weapon, while right-leaning sources commonly discuss the economy as a system subject to vertical changes.
Chinese Translation
概念隐喻通过允许我们以更具体或具身的经验(例如,搬运物理负担)来推理更抽象的经验(例如,缴纳税款),从而引导我们的思维和行动(Lakoff 和 Johnson, 2011)。因此,不同的概念隐喻可能导致不同的推理:将缴纳税款框架视为对社区的投资,而不是物理负担,会对税收产生截然不同的看法。因此,识别引导发言者或作者的概念隐喻有助于揭示他们对事件的框架。尽管这些隐喻无法直接观察,但语言隐喻的群体,即它们在语言中出现的隐喻表达,作为其证据。基于此,我们提出了一种无监督的方法,从语料库中提取语言隐喻,并使用结构化聚类方法形成与概念隐喻相对应的群体。通过这种方法,我们指出了左倾与右倾播客在主题和框架上的关键差异。例如,左倾播客倾向于将媒体故事概念化为武器,而右倾来源则通常将经济视为一个受垂直变化影响的系统。
cs.CL / 17 / 2608.06663
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents
视野差距:长视野大语言模型代理的规划、记忆、执行、训练与评估
Abstract
Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earlier decisions, declaring half-finished work done, or drifting from goals. We call this the horizon gap and survey 1,547 arXiv papers (2024-2026) collected via systematic seed harvest with a disclosed 26.8% bleed filter, extended by targeted supplementation. We disambiguate three routinely conflated properties: long-horizon (task property: required steps), long-context (model property: token capacity), and long-term memory (system property: persistence across steps/sessions). We organize the corpus into six categories tracking a long-horizon task's lifecycle -- planning, memory, execution, training, evaluation, and foundations/safety -- crossed with an axis capturing where horizons are carried (within-context, within-task-beyond-context, or cross-task-persistent). Across all categories, we find the same pattern: outcome-only signals grow uninformative as horizons lengthen, and the field's response -- whether process reward models, credit assignment, or trajectory-level diagnostics -- manufactures denser step-level signals. We treat critical and diagnostic literature as first-class threads throughout, arguing that segregating critique from method would routinely split single papers across chapters. We close by naming open measurement problems: decomposing model versus harness capability, managing correlated bias in process-level signals used for both training and evaluation, and whether long-horizon reliability admits general predictive theory.
Chinese Translation
前沿语言模型在单次前向推理中解决了曾经需要研究贡献的推理问题,但在多小时任务中却表现不佳:无法跟踪早期决策、宣称未完成的工作已完成,或偏离目标。我们称之为视野差距,并对1,547篇通过系统种子收集的arXiv论文(2024-2026)进行了调查,采用了公开的26.8%漏斗过滤,并通过有针对性的补充进行了扩展。我们澄清了三个常被混淆的属性:长视野(任务属性:所需步骤)、长上下文(模型属性:标记容量)和长期记忆(系统属性:跨步骤/会话的持久性)。我们将文献组织为六个类别,追踪长视野任务的生命周期——规划、记忆、执行、训练、评估以及基础/安全——并与一个捕捉视野承载位置的轴交叉(在上下文内、在任务内超出上下文,或跨任务持久)。在所有类别中,我们发现相同的模式:随着视野的延长,仅结果信号变得不再信息丰富,而该领域的响应——无论是过程奖励模型、信用分配还是轨迹级诊断——都制造了更密集的步骤级信号。我们将关键和诊断文献视为贯穿始终的第一类线索,认为将批评与方法分开通常会将单篇论文分割成多个章节。最后,我们提出了开放的测量问题:分解模型与工具能力、管理用于训练和评估的过程级信号中的相关偏差,以及长视野可靠性是否承认一般预测理论。
cs.CL / 18 / 2608.06672
TA-RAG: Tone Awareness as a Design Imperative for Retrieval-Augmented Generation
TA-RAG:语调意识作为检索增强生成的设计必要性
Abstract
Retrieval-Augmented Generation (RAG) has become a robust architecture for grounding large language models (LLMs) in trusted knowledge. However, standard RAG systems exhibit a structural limitation: retrieved documents carry their own communication styles-professional jargon, formal tone, or academic writings-that shape the behavior of a RAG system before any tone instructions are processed, often causing the system to ignore user requests for a specific tone. We term this phenomenon contextual decoupling, in which a system optimises for factual accuracy while remaining decoupled from the social or operational context of the recipient. Building on prior research in public health peer-support communities, we identify three communicative misalignment-linguistic, cognitive, and relational-that can persist even when retrieval is relevant and the generated response is factually accurate. We conceptualise these as failures of communicative transformation, which remain largely invisible to accuracy-centred RAG evaluation metrics. To address this gap, we propose Tone-Aware RAG (TA-RAG), a conceptual architectural framework that positions communicative alignment alongside factual accuracy as a core design objective. TA-RAG operationalises four constraints-stigma-free language, readability alignment, recipient-sensitive adaptation, and empathetic framing-across the retrieval, context construction, generation, and constraint validation phases in the proposed RAG pipeline. We further highlight an evaluation agenda for jointly assessing factual fidelity and communicative alignment, and identify open challenges. We argue that tone awareness should be treated not as an optional refinement, but as a present design imperative for RAG systems operating in socially sensitive and high-stakes contexts.
Chinese Translation
检索增强生成(RAG)已成为将大型语言模型(LLMs)与可信知识相结合的强大架构。然而,标准的RAG系统存在结构性局限:检索到的文档携带自身的交流风格——专业术语、正式语气或学术写作——这些风格在任何语调指令被处理之前就会影响RAG系统的行为,常常导致系统忽视用户对特定语调的请求。我们将这一现象称为上下文解耦(contextual decoupling),在这种情况下,系统优化事实准确性,但与接收者的社会或操作上下文保持解耦。基于之前在公共卫生同行支持社区的研究,我们识别出三种交流不对齐——语言、认知和关系——即使在检索相关且生成的响应在事实上准确时,这些不对齐仍然存在。我们将这些概念化为交流转化的失败,这些失败在以准确性为中心的RAG评估指标中大多是隐形的。为了填补这一空白,我们提出了语调意识RAG(TA-RAG),这是一个概念性架构框架,将交流对齐与事实准确性并列作为核心设计目标。TA-RAG在提议的RAG管道的检索、上下文构建、生成和约束验证阶段实施了四项约束——无污名化语言、可读性对齐、接收者敏感适应和同理心框架。我们进一步强调了一个评估议程,以共同评估事实忠实性和交流对齐,并识别出开放挑战。我们认为,语调意识不应被视为可选的改进,而应作为在社会敏感和高风险环境中运行的RAG系统的当前设计必要性。
cs.CL / 19 / 2608.06718
Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation
音频语言模型是否使用副语言证据?响应评估的反事实审计
Abstract
Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paralinguistic evidence. We introduce counterfactual audits for paralinguistic response evaluation. Each audit item holds the transcript fixed while varying affect, prosody, or the timing of an affective shift, forcing a valid judge to track the audio cue rather than lexical content or response style. We evaluate ALM judges using a native one-context judgment protocol and a contrastive recoverability control, then further decompose each item into its constituent perception and response-mapping skills. This yields useful diagnostic states that identify different sources of judge failures. Across Gemini, GPT, and open audio models, we find that contrastive success often overstates native judge reliability, and that similar aggregate accuracies can hide different failure modes. These results suggest that ALM judges should not be evaluated by accuracy alone, instead requiring thorough behavioral audits before deployment.
Chinese Translation
音频语言模型(ALMs)越来越多地被用作语音到语音系统的评判者,但接收音频的评判者可能并未实际使用副语言证据。我们引入了用于副语言响应评估的反事实审计。每个审计项目在固定转录文本的同时,变化情感、韵律或情感转变的时机,迫使有效的评判者跟踪音频线索,而不是词汇内容或响应风格。我们使用本地单一上下文判断协议和对比可恢复性控制来评估ALM评判者,然后进一步将每个项目分解为其组成的感知和响应映射技能。这产生了有用的诊断状态,识别出不同来源的评判失败。在Gemini、GPT和开放音频模型中,我们发现对比成功往往夸大了本地评判者的可靠性,并且相似的整体准确性可能掩盖不同的失败模式。这些结果表明,ALM评判者不应仅通过准确性进行评估,而是需要在部署前进行全面的行为审计。
cs.CL / 20 / 2608.06750
Progressive Content Refinement with Decaying Reward Joint LinUCB
衰减奖励联合LinUCB的渐进内容精炼
Abstract
Iterative refinement has significantly enhanced Large Language Model (LLM) performance; however, existing methods ranging from feedback-based Self-Refine to traditional bandit approaches often rely on static options or overlook the saturation effect. This neglect leads to over-exploitation, where the continuous use of identical prompts or arms results in diminishing rewards over time. To address this challenge, we propose a novel contextual bandit algorithm that explicitly incorporates reward decay modeling. Utilizing an Expectation-Maximization (EM) algorithm, our method simultaneously estimates both arm-specific and decay parameters. Furthermore, by embedding prompts as arms, we facilitate the joint learning of arm values, distinguishing our approach from the traditional disjoint Linear Upper Confidence Bound (LinUCB) framework. Experimental results on Sentiment Reversal and GSM8K benchmarks demonstrate that our method achieves significant performance gains over strong baselines. Finally, our ablation study confirms that the integration of reward decay modeling within the bandit framework is crucial for mitigating over-exploitation and optimizing the iterative refinement process.
Chinese Translation
迭代精炼显著提升了大型语言模型(LLM)的性能;然而,现有方法从基于反馈的自我精炼到传统的强盗方法,往往依赖于静态选项或忽视饱和效应。这种忽视导致了过度利用,即持续使用相同的提示或臂会导致奖励随时间减少。为了解决这一挑战,我们提出了一种新颖的上下文强盗算法,明确地纳入了奖励衰减建模。利用期望最大化(EM)算法,我们的方法同时估计臂特定参数和衰减参数。此外,通过将提示嵌入为臂,我们促进了臂值的联合学习,使我们的方法区别于传统的离散线性上置信界(LinUCB)框架。在情感反转和GSM8K基准测试上的实验结果表明,我们的方法在强基线之上实现了显著的性能提升。最后,我们的消融研究确认了在强盗框架中整合奖励衰减建模对于减轻过度利用和优化迭代精炼过程的重要性。
cs.CL / 21 / 2608.06758
Stockmark-Nemotron-3-Nano-Omni-JapanDocReader: Structured Document Parsing via Capability Injection and Forgetting Control
Stockmark-Nemotron-3-Nano-Omni-JapanDocReader:通过能力注入和遗忘控制进行结构化文档解析
Abstract
We present Stockmark-Nemotron-3-Nano-Omni-JapanDocReader, a Japanese document understanding model built from Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16. The central goal of this work is structured document parsing via capability injection and forgetting control: we inject Japanese structured document parsing capability into a reasoning-oriented multimodal model while preserving its document VQA capability as much as possible. We study parsing-centric SFT, which uses only structured document parsing data; mixed SFT, which combines structured document parsing and VQA data; and parsing-centric RL, which optimizes structured parsing with a task-level reward. Our experiments show that parsing-centric SFT substantially improves structured document parsing performance but causes measurable VQA forgetting. Mixed SFT mitigates this forgetting while preserving nearly the same structured parsing performance. Applying DAPO-based parsing-centric RL on top of the mixed SFT checkpoint further improves structured document parsing beyond the SFT ceiling, producing the final released model. The training data is constructed with a data engine consisting of two complementary synthetic streams: a Japanese Document VQA Stream and a programmatic structured document parsing stream. We also discuss reward design and variance-based prompt filtering for continuous structured document parsing rewards, highlighting their importance for making RL effective in long-reasoning structured document parsing tasks.
Chinese Translation
我们提出了Stockmark-Nemotron-3-Nano-Omni-JapanDocReader,这是一个基于Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16构建的日文文档理解模型。本研究的核心目标是通过能力注入和遗忘控制实现结构化文档解析:我们在尽可能保留其文档视觉问答(VQA)能力的同时,将日文结构化文档解析能力注入到一个以推理为导向的多模态模型中。我们研究了以解析为中心的监督微调(SFT),该方法仅使用结构化文档解析数据;混合SFT,结合了结构化文档解析和VQA数据;以及以解析为中心的强化学习(RL),该方法通过任务级奖励优化结构化解析。我们的实验表明,以解析为中心的SFT显著提高了结构化文档解析性能,但导致了可测量的VQA遗忘。混合SFT在保留几乎相同的结构化解析性能的同时,减轻了这种遗忘。在混合SFT检查点基础上应用基于DAPO的以解析为中心的RL进一步提高了结构化文档解析性能,超越了SFT的性能上限,产生了最终发布的模型。训练数据由一个包含两个互补合成流的数据引擎构建:一个日文文档VQA流和一个程序化结构化文档解析流。我们还讨论了奖励设计和基于方差的提示过滤,以实现连续的结构化文档解析奖励,强调了它们在长推理结构化文档解析任务中使强化学习有效的重要性。
cs.CL / 22 / 2608.06785
Multi-Perspective Triad Interaction Graph Neural Network for Cognitive Distortion Detection
多视角三元组交互图神经网络用于认知扭曲检测
Abstract
Cognitive distortion detection is a key task in computational mental health, yet existing approaches often overlook the psychological structure of distorted thoughts. We propose MTI-GNN (Multi-Perspective Triad Interaction Graph Neural Network), which models Beck's cognitive triad---negative views of the self, world, and future---as complementary perspectives for classification. An LLM decomposes each utterance into the three perspectives, from which perspective-specific similarity graphs are constructed and encoded by a Multi-Perspective GNN. A Triad Interaction module models cross-perspective dependencies through sequential source-conditioned updates and feature-wise gating, while Prototype-Guided Perspective Fusion performs label-conditioned aggregation. Label-expanded supervision incorporates all available distortion annotations during training. We evaluate MTI-GNN on 9,764 samples from four Korean, English, and Chinese datasets spanning ten distortion categories. MTI-GNN significantly outperforms all supervised variants and exceeds eight prompted generative models under zero-shot and few-shot settings. Leave-one-perspective-out ablations show that all three perspectives contribute significantly, while human expert evaluation provides preliminary evidence of their alignment with the intended cognitive dimensions.
Chinese Translation
认知扭曲检测是计算心理健康中的一项关键任务,但现有方法往往忽视了扭曲思想的心理结构。我们提出了MTI-GNN(多视角三元组交互图神经网络),该模型将贝克的认知三元组——对自我、世界和未来的负面看法——视为分类的互补视角。一个大型语言模型(LLM)将每个话语分解为三个视角,从中构建视角特定的相似性图,并通过多视角图神经网络(Multi-Perspective GNN)进行编码。三元组交互模块通过顺序源条件更新和特征级门控建模跨视角依赖关系,而原型引导的视角融合则执行标签条件聚合。标签扩展监督在训练过程中整合所有可用的扭曲注释。我们在来自四个韩语、英语和中文数据集的9,764个样本上评估了MTI-GNN,涵盖十个扭曲类别。MTI-GNN显著优于所有监督变体,并在零样本和少样本设置下超过了八个提示生成模型。排除一个视角的消融实验表明,所有三个视角都显著贡献,而人类专家评估提供了初步证据,表明它们与预期的认知维度一致。
cs.CL / 23 / 2608.06802
Simple-OPD: Demystifying Warm-up for On-policy Distillation
Simple-OPD:揭示针对策略蒸馏的预热过程
Abstract
On-policy distillation (OPD) trains a student on its own rollouts with token-level supervision from teacher models, but its effectiveness can depend strongly on the warm-up stage before OPD. In this paper, we demystify warm-up for OPD from both data and training perspectives. For data, we find that effective warm-up relies on teacher-compatible chain-of-thought supervision, and that even incorrect teacher rollouts can provide comparable benefits to correct ones. This suggests that warm-up primarily transfers a teacher-compatible thinking pattern rather than merely correct answers. For training, we show that low-rank adaptation (LoRA) with a near-saturation training duration better balances in-domain adaptation and out-of-distribution generalization than full-parameter SFT. Based on these findings, we propose Simple-OPD, a plug-and-play initialization method that warms up the student on teacher-generated CoT with LoRA before OPD. Experiments across diverse settings demonstrate the effectiveness and robustness of Simple-OPD.
Chinese Translation
针对策略蒸馏(On-policy distillation, OPD)通过教师模型的标记级监督在其自身的回合上训练学生,但其有效性在很大程度上依赖于OPD之前的预热阶段。本文从数据和训练的角度揭示了OPD的预热过程。在数据方面,我们发现有效的预热依赖于与教师兼容的思维链(chain-of-thought)监督,甚至不正确的教师回合也能提供与正确回合相当的好处。这表明,预热主要传递的是与教师兼容的思维模式,而不仅仅是正确的答案。在训练方面,我们展示了低秩适应(Low-rank adaptation, LoRA)在接近饱和的训练时长下,能够比全参数的微调(SFT)更好地平衡领域内适应和领域外泛化。基于这些发现,我们提出了Simple-OPD,这是一种即插即用的初始化方法,在进行OPD之前,利用LoRA在教师生成的思维链上对学生进行预热。跨多种设置的实验表明,Simple-OPD的有效性和鲁棒性。
cs.CL / 24 / 2608.06819
FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding
FutureBridge:超越局部偏好的协作解码中的标记选择
Abstract
Token-level collaboration allows a large language model (LLM) to assist a small language model (SLM) when their predictions diverge. Existing methods either use LLM-generated intervention tokens or rank candidates with the LLM's next-token probabilities. Both rely on the LLM's local preference, even though an LLM-selected token may be difficult for the SLM to build on. We present FutureBridge, which ranks joint LLM-SLM token candidates according to how well they support the SLM's subsequent reasoning. During training, an answer-verified LLM trajectory supplies a fixed shared future, and a frozen SLM evaluates every candidate under this common context. The resulting counterfactual scores supervise a lightweight token reranker that observes only the current state and candidate token. At inference, FutureBridge uses the LLM only to expand the candidate pool, selects one token, and returns generation to the SLM without generating or appending a future suffix. Across five mathematical reasoning benchmarks, FutureBridge improves the Qwen3-1.7B SLM's Math Avg. by 35.1% relative to greedy SLM decoding. These results indicate that token selection benefits from modeling whether the receiving SLM can use each candidate to continue reasoning, rather than relying on the LLM's local preference alone.
Chinese Translation
标记级协作使得大型语言模型(LLM)能够在小型语言模型(SLM)预测不一致时提供帮助。现有方法要么使用LLM生成的干预标记,要么根据LLM的下一个标记概率对候选项进行排序。这两者都依赖于LLM的局部偏好,尽管LLM选择的标记可能难以被SLM继续构建。我们提出了FutureBridge,它根据标记如何支持SLM后续推理的效果来对LLM-SLM联合标记候选项进行排序。在训练过程中,一个经过答案验证的LLM轨迹提供了一个固定的共享未来,而一个冻结的SLM在这个共同上下文下评估每个候选项。由此产生的反事实评分监督一个轻量级的标记重排序器,该重排序器仅观察当前状态和候选标记。在推理阶段,FutureBridge仅使用LLM来扩展候选池,选择一个标记,并将生成结果返回给SLM,而不生成或附加未来后缀。在五个数学推理基准测试中,FutureBridge相较于贪婪SLM解码提高了Qwen3-1.7B SLM的数学平均分35.1%。这些结果表明,标记选择受益于建模接收SLM是否能够使用每个候选项继续推理,而不仅仅依赖于LLM的局部偏好。
cs.CL / 25 / 2608.06849
Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry
头部自主性:基于冻结查询-键几何的无数据稀疏注意力
Abstract
Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-compression methods typically decide which tokens or heads to preserve from runtime attention scores, observation windows, calibration prompts, or learned gates, making head diagnosis input-dependent and costly to deploy. We propose Autonomy-of-Heads (AoH), a data-free method that identifies retrieval and streaming heads from the spectral geometry of query-key projections. AoH defines the kernel attention operator $M_h = W_K^{h\top}W_Q^h$ and uses its effective-rank as a weight-space measure of head function: concentrated spectra indicate a small number of dominant query-key matching directions and are associated with retrieval heads, whereas diffuse spectra indicate the absence of a dominant global matching direction and are associated with streaming heads. We further derive an efficient $d_\text{head}$-dimensional computation that avoids constructing the full $d_\text{model}\times d_\text{model}$ matrix. We conducted extensive experiments across models demonstrating that at 50\% sparsity, AoH retains 96.5\% of Full Attention performance on average while reducing prefill and decode latency by up to 41.4\% and 66.0\%, respectively, and KV-cache memory by 50.0\% at 256K tokens.
Chinese Translation
长上下文大语言模型(LLM)推理受到二次注意力计算和不断增长的键值缓存(KV-cache)成本的瓶颈。现有的稀疏注意力和KV压缩方法通常根据运行时注意力得分、观察窗口、校准提示或学习的门控来决定保留哪些令牌或头部,这使得头部诊断依赖于输入且部署成本高昂。我们提出了头部自主性(Autonomy-of-Heads, AoH),这是一种无数据的方法,通过查询-键投影的谱几何来识别检索头和流式头。AoH定义了核注意力算子 $M_h = W_K^{h op}W_Q^h$,并使用其有效秩作为头部功能的权重空间度量:集中谱表示少量主导的查询-键匹配方向,与检索头相关,而分散谱则表示缺乏主导的全局匹配方向,与流式头相关。我们进一步推导出一种高效的 $d_ ext{head}$ 维度计算,避免构造完整的 $d_ ext{model} imes d_ ext{model}$ 矩阵。我们在多个模型上进行了广泛实验,结果表明在50\%稀疏度下,AoH平均保留了96.5\%的全注意力性能,同时将预填充和解码延迟分别减少了高达41.4\\%和66.0\\%,并在256K令牌时将KV-cache内存减少了50.0\ %。
cs.CL / 26 / 2608.06867
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
LLMRouter:开发、评估和部署 LLM 路由器的统一基础设施
Abstract
No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations and implementations, making fair comparison and extension difficult. We present a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, model encoders, scoring functions, decision rules, and learning signals, covering single-turn, multi-turn, and personalized routing. Based on this formulation, we develop an automated pipeline for constructing routing supervision and evaluating routers jointly on response quality and inference cost. The resulting benchmark, xRouteBench, spans generic LLM, memory-augmented, vision, time-series, and personalized routing tasks. We further introduce LLMRouter, an open-source modular infrastructure with more than 16 representative routers. Our empirical study shows that learned routers outperform the strongest fixed-model baseline by 14.6% relatively, lightweight routers become more competitive under tight cost constraints, and user-conditioned routing consistently improves personalization.
Chinese Translation
没有单一的大型语言模型(LLM)能够在所有查询和预算限制下达到最佳效果,因此模型路由对于成本效益的部署至关重要。现有的路由器采用多种不同的公式和实现方式,使得公平比较和扩展变得困难。我们提出了一种统一的 LLM 路由公式,将其视为一个由五个组成部分构成的序列决策过程:上下文编码器、模型编码器、评分函数、决策规则和学习信号,涵盖单轮、多轮和个性化路由。基于这一公式,我们开发了一个自动化管道,用于构建路由监督并在响应质量和推理成本上共同评估路由器。最终得到的基准测试 xRouteBench 涉及通用 LLM、内存增强、视觉、时间序列和个性化路由任务。我们进一步介绍了 LLMRouter,一个开源模块化基础设施,包含超过 16 个代表性路由器。我们的实证研究表明,学习型路由器相较于最强的固定模型基线提高了 14.6%,轻量级路由器在紧张的成本限制下变得更具竞争力,而用户条件路由则持续改善个性化效果。
cs.CL / 27 / 2608.06884
Georeferencing Non-Gazetteered Place Names using Biological Specimen Records
利用生物标本记录进行非公报地名的地理参考
Abstract
Biological specimen records collected by natural history institutions constitute a rich source of temporal geographic knowledge, capturing biodiversity information about regional landscapes as they were recorded at different times. Using digitised data from the Allan Herbarium (New Zealand), this study identifies place names in these specimen locality descriptions that are absent from current gazetteers; we refer to these as non-gazetteer place names (NGPs). These place names are typically historical, vernacular, or colloquial and were used as landmarks to describe a specimen's location at the time of collection. We then investigate the problem of georeferencing the NGPs using only the limited information available in the specimen records. To resolve this, we leverage repeated occurrences of the same place name across specimen records with different specimen locations and spatial relation terms, extracting and inverting these relations to derive constraints on NGP locations. This approach is instantiated within deterministic, probabilistic, and LLM-based methods, enabling a comparative analysis of their strengths and limitations for text-based spatial inference. On a pseudo-NGP benchmark, probabilistic inference achieves the highest accuracy (median error 1.43 km; A@1 km 36%), while the LLM yields competitive but less precise estimates (median error 1.80 km; A@1 km 31%), indicating that, despite advances in LLMs, traditional modelling remains advantageous when high spatial precision is required.
Chinese Translation
自然历史机构收集的生物标本记录构成了丰富的时间地理知识来源,捕捉了关于区域景观的生物多样性信息,这些信息是在不同时间记录的。通过使用来自新西兰艾伦植物标本馆(Allan Herbarium)的数字化数据,本研究识别出这些标本地点描述中缺失于当前公报的地名;我们将这些称为非公报地名(Non-Gazetteered Place Names, NGPs)。这些地名通常是历史的、方言的或口语的,并在收集标本时作为描述标本位置的地标。然后,我们研究了仅使用标本记录中有限信息对NGPs进行地理参考的问题。为了解决这个问题,我们利用在不同标本位置和空间关系术语中反复出现的相同地名,提取并反转这些关系,以推导出对NGP位置的约束。该方法在确定性、概率性和基于大语言模型(LLM)的方法中得以实现,使得对它们在基于文本的空间推理中的优缺点进行比较分析成为可能。在一个伪NGP基准测试中,概率推理实现了最高的准确性(中位误差1.43公里;A@1公里36%),而LLM则提供了竞争性但不够精确的估计(中位误差1.80公里;A@1公里31%),这表明尽管LLM取得了进展,但在需要高空间精度时,传统建模仍然具有优势。
cs.CL / 28 / 2608.06908
Calibrating WEAT Against Anisotropy: ZCA Whitening as a Geometric Pre-Processing Step for Embedding Association Tests
针对各向异性的 WEAT 校准:ZCA 白化作为嵌入关联测试的几何预处理步骤
Abstract
We propose Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step for the Word Embedding Association Test (WEAT). WEAT is a bias measurement method widely used in both computational social science and AI fairness research. It relies on cosine similarity as a measure of semantic association, which assumes that the embedding space is approximately isotropic. However, prior work has reported that many widely used language models do not satisfy this assumption, raising concerns about the reliability of bias measurements. ZCA whitening transforms the covariance of the embedding space into the identity matrix while minimizing perturbation to the original vectors. This transformation restores the isotropy condition on which WEAT relies. We evaluate our approach on ten standard WEAT test suites and seven models spanning three architectural families, yielding 70 model-task combinations. The results show that ZCA whitening substantially reduces the anisotropy of the embedding spaces across all models. Particularly for highly anisotropic models, we further observe improvements on standard semantic similarity benchmarks, indicating that the calibrated space better captures semantic associations. After calibration, over 30% of WEAT results change significance status, and effect sizes shift in both directions depending on bias category. These shifts suggest that uncalibrated measurements may both overestimate and underestimate the associations encoded in the embedding space. These findings indicate that previously reported bias measurements in anisotropic embedding spaces should be interpreted with caution and may benefit from re-evaluation with calibrated methods. Our approach contributes to restoring the measurement foundation of WEAT across both computational social science and AI fairness research.
Chinese Translation
我们提出将零相位分量分析(Zero-phase Component Analysis, ZCA)白化作为词嵌入关联测试(Word Embedding Association Test, WEAT)的几何预处理步骤。WEAT 是一种广泛应用于计算社会科学和人工智能公平性研究的偏差测量方法。它依赖余弦相似度作为语义关联的度量,这一方法假设嵌入空间大致是各向同性的。然而,先前的研究报告表明,许多广泛使用的语言模型并不满足这一假设,这引发了对偏差测量可靠性的担忧。ZCA 白化将嵌入空间的协方差转换为单位矩阵,同时最小化对原始向量的扰动。这一转换恢复了 WEAT 所依赖的各向同性条件。我们在十个标准 WEAT 测试套件和七个跨越三种架构家族的模型上评估了我们的方法,共生成 70 种模型-任务组合。结果表明,ZCA 白化显著降低了所有模型嵌入空间的各向异性。特别是对于高度各向异性的模型,我们进一步观察到在标准语义相似性基准上的改善,表明校准后的空间更好地捕捉了语义关联。校准后,超过 30% 的 WEAT 结果改变了显著性状态,效应大小在不同偏差类别下双向变化。这些变化表明,未校准的测量可能会高估或低估嵌入空间中编码的关联。这些发现表明,在各向异性嵌入空间中先前报告的偏差测量应谨慎解读,并可能受益于使用校准方法的重新评估。我们的方法有助于恢复 WEAT 在计算社会科学和人工智能公平性研究中的测量基础。
cs.CL / 29 / 2608.06933
Ask-E: An Environment for Calibrated Question Generation
Ask-E:一个用于校准问题生成的环境
Abstract
Today, we improve models by training and evaluating them on problems at the frontier of their abilities. Creating such problems is itself a demanding task, requiring the ability to probe model limits and generalize beyond existing question distributions. It also means placing problems at a precise difficulty level, which requires understanding what it takes to solve them. In short, generating problems calibrated to a model's current frontier demands capability beyond it, an increasingly burdensome constraint as models improve. Our key insight is that we can leverage this constraint to our advantage: a model that can generate problems consistently calibrated to a given frontier must possess capability beyond it. Accordingly, we present Ask-E, an environment that benchmarks and trains models on their ability to write questions at a given skill level, rather than answer them. Concretely, we define target skill levels as ranges bounded by the capabilities of two existing language models. A generated question is successfully calibrated if exactly one of the two models can solve it, placing it precisely within the target range and differentiating the capabilities of these models. Ask-E serves both as a benchmark and a training environment, where models generate problems calibrated to a variety of skill levels. We find that even frontier models achieve below 50% calibration on the benchmark, leaving significant headroom to measure future progress. We also show that training on this environment leads to improvements across a number of downstream math benchmarks even with no new math data, no interaction with stronger models, and no correctness-based reward.
Chinese Translation
如今,我们通过在模型能力的前沿问题上进行训练和评估来提升模型的表现。创建这样的难题本身就是一项艰巨的任务,要求能够探测模型的极限并超越现有问题分布的能力。这也意味着将问题设置在一个精确的难度水平上,这需要理解解决这些问题所需的条件。简而言之,生成与模型当前前沿相匹配的问题需要具备超越该前沿的能力,随着模型的进步,这一约束变得愈加沉重。我们的关键见解是,我们可以利用这一约束为我们所用:一个能够持续生成与给定前沿相匹配的问题的模型,必定具备超越该前沿的能力。因此,我们提出了Ask-E,一个基准和训练环境,专注于模型在特定技能水平下生成问题的能力,而不是回答问题。具体而言,我们将目标技能水平定义为由两个现有语言模型的能力界定的范围。如果只有两个模型中的一个能够解决生成的问题,则该问题被成功校准,恰好位于目标范围内,并区分了这些模型的能力。Ask-E既作为基准测试,也作为训练环境,模型在其中生成与多种技能水平相匹配的问题。我们发现,即使是前沿模型在基准测试中的校准率也低于50%,这为未来的进展测量留下了显著的提升空间。我们还表明,在这个环境中进行训练,即使没有新的数学数据、没有与更强模型的交互,也没有基于正确性的奖励,仍然能在多个下游数学基准上带来改进。
cs.CL / 30 / 2608.06953
Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression
明确,而非更长:什么使得认知立场在记忆压缩中得以存活
Abstract
Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to survive being written to memory. We ask what governs whether it does. Matched notes carry the identical claim and identical stance and differ only in where that stance sits; one model compresses both under the same budget among the same filler notes, and a blind reader that never sees the condition scores the result. Across 60 claims in seven registers, writing the stance as a labelled field rather than a bracketed aside raises retention by about 15 points on two models (37 claims to 2 on one, 30 to 8 on the other; permutation p=0.00005), and a pre-registered replication on Haiku, its prediction and decision rule committed before the run, gives +15.6 points, 38 claims to 1. Ablating the format on both models gives the same net effect from different parts: labels help on both (+9.7 and +12.8) and length helps on neither, but wording the stance as a full sentence is the largest component on one model (+12.5) and worth nothing on the other (+0.6). Either model alone would have licensed a confident and different mechanism, so we claim only the intersection: make the stance explicit, not merely longer, and expect the best way of being explicit to depend on the model. A deterministic readout with no model reproduces the two-cell direction and five of seven ablation contrasts, but not length or labels, which we therefore do not claim on one instrument. Fifty hand labels (kappa=0.75) agree on direction; we print their seven disagreements in full. We also report nine withdrawn claims, three of them former title claims of this paper.
Chinese Translation
代理记忆系统会对所存储的信息进行压缩,而压缩的过程通常会去掉修饰语,因此一个主张的认知地位往往无法在写入记忆后存活。我们探讨了是什么因素决定了其存活与否。匹配的笔记包含相同的主张和相同的立场,唯一的区别在于立场的位置;一个模型在相同的预算下将两者压缩到相同的填充笔记中,而一个从未见过条件的盲读者对结果进行评分。在七个语域中的60个主张中,将立场作为标记字段而非括注的附带信息进行书写,使得保留率提高了约15个百分点(在一个模型中从37个主张提高到2个,在另一个模型中从30个提高到8个;排列p=0.00005),而在Haiku上的预注册复制实验中,其预测和决策规则在运行前已确定,结果提高了15.6个百分点,38个主张对1个主张。在两个模型中消除这种格式的影响,来自不同部分的净效应相同:标签在两者上都有帮助(+9.7和+12.8),而长度对两者均无帮助,但将立场表述为完整句子是一个模型中最大的影响因素(+12.5),而在另一个模型中几乎没有影响(+0.6)。单独使用任一模型都可以支持一种自信且不同的机制,因此我们仅主张交集:使立场明确,而不仅仅是更长,并且期望最佳的明确方式取决于模型。没有模型的确定性输出重现了两个单元的方向和七个消融对比中的五个,但不包括长度或标签,因此我们不在一个工具上主张这两者。五十个手动标签(kappa=0.75)在方向上达成一致;我们完整列出了他们的七个分歧。我们还报告了九个撤回的主张,其中三个是本文的前标题主张。
cs.CL / 31 / 2608.06967
Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs
语言模型能在没有视觉输入的情况下进行想象吗?Ekphrasis:测量仅文本大型语言模型中的视觉创造性构思
Abstract
Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar visual clich\'es or failing to specify a renderable scene. We define Visual Creative Ideation (VCI) as the ability to produce textual visual plans that are useful, expressive, and population-novel, and introduce Ekphrasis, a 400-task benchmark spanning Abstraction, Combination, Transformation, and Adaptation. Ekphrasis scores anonymized pairwise comparisons with dimension-specific checklists, aggregates preferences with Bradley-Terry models, and uses Typed Idea Graphs to convert task-specific population clich\'es into novelty references. Across 14 language models, VCI separates usefulness, expressiveness, and novelty rather than reducing to fluency: strong models achieve similar overall scores through different profiles, and useful plans can remain visually clich\'ed. A cross-modal grounding study further shows that text-level VCI ordering largely survives faithful rendering and blind image-level preference judgment, supporting Ekphrasis as a measure of visual ideation beyond prose quality.
Chinese Translation
当前的评估并未区分仅文本的语言模型是否能够在图像生成之前产生视觉概念。流畅的视觉散文可能掩盖视觉计划的失败:一个答案可能看似富有创意,但实际上重复了熟悉的视觉陈词滥调或未能具体描述可渲染的场景。我们将视觉创造性构思(Visual Creative Ideation, VCI)定义为产生有用、富有表现力且在人群中新颖的文本视觉计划的能力,并引入Ekphrasis,这是一个涵盖抽象、组合、转化和适应的400项任务基准。Ekphrasis通过维度特定的检查表对匿名的成对比较进行评分,利用Bradley-Terry模型汇总偏好,并使用类型化创意图(Typed Idea Graphs)将任务特定的人群陈词滥调转化为新颖性参考。在14个语言模型中,VCI区分了有用性、表现力和新颖性,而不是简化为流畅性:强大的模型通过不同的特征实现相似的整体得分,而有用的计划可能仍然在视觉上显得陈旧。一项跨模态基础研究进一步表明,文本层面的VCI排序在忠实渲染和盲目图像层面偏好判断中大体保持不变,支持Ekphrasis作为超越散文质量的视觉构思测量标准。
cs.CL / 32 / 2608.06975
PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue
PHASE-Tree:长时间角色扮演对话中的角色状态演变建模
Abstract
Long-horizon role-playing demands that characters remain recognizable as they evolve with the narrative. Yet existing work falls short on two fronts: representations are typically static profiles that cannot be updated locally without destabilizing unchanged traits, and benchmarks mainly test persona preservation and memory recall rather than whether a model speaks from a character's currently evolved state. We address both. PHASE-Tree is a multi-timescale character-state tree with an immutable identity root and mutable persona, session, and moment layers, making each mutable field an addressable target for localized within- and cross-episode updates. It conditions generation through explicit textual provision or implicit parametric adaptation. To measure evolved-state generation, we introduce LongEvoRoleBench, which pairs four long-dialogue corpora for cross-episode evolution with four short-dialogue corpora as within-scene state-tracking checks, under a unified next-utterance protocol. On the long-dialogue core, textual PHASE-Tree ranks first in 11 of 12 dataset-metric cells against internal variants and all 12 cells against external textual baselines, improving character-level, semantic, and embedding scores by 19.7%, 12.4%, and 15.1% respectively. In a blinded 200-response study, human ratings correlate with the GPT-4.1 judge (Pearson r= 0.65); on descriptive n= 10 PT and NR prompt subsets, the Overall difference is +0.20. The long-dialogue Sem advantage persists across LLM judges and generation backbones.
Chinese Translation
长时间的角色扮演要求角色在叙事中随着情节的发展保持可识别性。然而,现有的研究在两个方面存在不足:表示通常是静态的角色档案,无法在不破坏未改变特征的情况下进行局部更新;基准测试主要测试角色保持和记忆回忆,而不是模型是否从角色当前演变的状态进行发言。我们解决了这两个问题。PHASE-Tree 是一个多时间尺度的角色状态树,具有不可变的身份根和可变的角色、会话和时刻层,使得每个可变字段成为局部剧集内和跨剧集更新的可寻址目标。它通过显式的文本提供或隐式的参数适应来调节生成。为了测量演变状态的生成,我们引入了 LongEvoRoleBench,它将四个长对话语料库与四个短对话语料库配对,以进行跨剧集演变和场景内状态跟踪检查,采用统一的下一个发言协议。在长对话核心上,文本 PHASE-Tree 在12个数据集-指标单元中的11个中排名第一,相较于内部变体和所有12个单元相较于外部文本基线,分别提高了19.7%、12.4%和15.1%的角色级、语义和嵌入分数。在一项盲测200个响应的研究中,人类评分与GPT-4.1评审者的相关性为(Pearson r= 0.65);在描述性n= 10的PT和NR提示子集上,总体差异为+0.20。长对话的语义优势在LLM评审者和生成基础架构中持续存在。
cs.CL / 33 / 2608.06977
Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models
确认我们的偏见?评估大型语言模型的能力、风险与社会影响
Abstract
It is well established that large language models (LLMs) are sensitive to prompt framing, reflecting patterns in their training data or prior prompts. In this study, we investigate the extent to which LLMs reinforce users biases expressed in the prompts and examine the boundary between implicit framing effects and explicit prompt manipulation. Specifically, we evaluate how susceptible LLMs are to direct and suggestive prompts that encourage models to support or challenge particular positions. We evaluate six LLMs using 160 distinct prompts spanning ten topics across opinion-based and factual domains. The prompts systematically vary in prompting strategy, support versus challenge instructions, prompt polarity, users' expressed beliefs, and topic domain, spanning both opinion-based and factual questions. Our results show that LLMs systematically adapt their responses to align with prompt framing, even in factual contexts. This suggests that prompt framing can outweigh factual consistency in model responses. Overall, our findings delineate the extent and boundaries of LLM manipulability. Furthermore, the results imply that LLMs can reinforce subtle user biases and are susceptible to explicit prompt manipulation even in domains where responses should remain factually stable.
Chinese Translation
大型语言模型(LLMs)对提示框架敏感,这一现象已得到充分证实,反映了其训练数据或先前提示中的模式。在本研究中,我们探讨了LLMs在多大程度上强化了用户在提示中表达的偏见,并考察了隐性框架效应与显性提示操控之间的界限。具体而言,我们评估了LLMs对直接和暗示性提示的敏感性,这些提示鼓励模型支持或挑战特定立场。我们使用160个跨越十个主题的不同提示对六个LLMs进行评估,这些主题涵盖了基于观点和事实的领域。提示在提示策略、支持与挑战指令、提示极性、用户表达的信念以及主题领域等方面系统性地变化,涉及基于观点和事实的问题。我们的结果表明,LLMs系统性地调整其响应以与提示框架保持一致,即使在事实背景下也是如此。这表明提示框架可能在模型响应中超过事实一致性。总体而言,我们的研究结果划定了LLMs可操控性的范围和界限。此外,结果还暗示LLMs能够强化微妙的用户偏见,并且在应保持事实稳定的领域中也容易受到显性提示操控的影响。
cs.CL / 34 / 2608.06992
GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base
GPTKB 2.0:浏览、查询和审计一个去歧义化的基于大型语言模型的知识库
Abstract
We present a web demo for exploring a large-scale disambiguated knowledge base (KB) materialized from a large language model (LLM). GPTKB 2.0 contains 38.4M triples over 1.6M canonical entities, together with 207.6K consolidated relations and 66K consolidated classes. Unlike prior LLM-derived knowledge bases that largely identify entities by surface strings, GPTKB 2.0 performs context-guided disambiguation during recursive KB construction, separating homonyms and merging synonymous mentions as facts are elicited. The demo makes this process inspectable: users can browse entities, follow links across the KB, and audit the provenance of individual facts, including surface forms, candidate matches, source triples, and disambiguation decisions. The interface further supports structured SPARQL queries, natural-language questions translated to SPARQL, and entity linking from user-provided text to canonical GPTKB 2.0 entries. GPTKB 2.0 is available at https://gptkb.org/, with the full KB downloadable for offline use.
Chinese Translation
我们展示了一个网络演示,用于探索一个从大型语言模型(LLM)衍生的大规模去歧义化知识库(KB)。GPTKB 2.0 包含 3840 万个三元组,覆盖 160 万个规范实体,以及 20.76 万个整合关系和 6.6 万个整合类别。与之前主要通过表面字符串识别实体的 LLM 衍生知识库不同,GPTKB 2.0 在递归构建知识库的过程中进行上下文引导的去歧义化,分离同名异义词并在事实被引出时合并同义提及。该演示使这一过程可供检查:用户可以浏览实体,跟随知识库中的链接,并审计个别事实的来源,包括表面形式、候选匹配、源三元组和去歧义决策。该界面还支持结构化的 SPARQL 查询、翻译为 SPARQL 的自然语言问题,以及从用户提供的文本到规范 GPTKB 2.0 条目的实体链接。GPTKB 2.0 可在 https://gptkb.org/ 获取,完整的知识库可供离线下载使用。
cs.CL / 35 / 2608.07006
Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models?
更多检索证据是否有助于使用扩散语言模型的视觉检索增强生成?
Abstract
Visual retrieval-augmented generation (RAG) commonly expands the retrieved evidence set to improve answer-page coverage, implicitly assuming that all available evidence should be passed to the generator. We show that this assumption does not hold for diffusion language models (DLMs): retrieving more pages increases answer-page recall, whereas unconditionally passing all retrieved pages to the generator often reduces answer accuracy, primarily because of semantic conflict. A latent-source analysis explains this mismatch through source-coherence loss in parallel denoising, where position-wise proposals can combine incompatible visual sources into unsupported answers. We further find that such interference is already visible in the first-step answer-block distribution, making it possible to assess evidence before decoding. To preserve retrieval coverage while limiting harmful visual exposure, we propose the Entropy-Based Candidate Filter (ECF), a training-free evidence-admission framework. To reduce irrelevant content within individual candidates, ECF constructs multi-granularity evidence units; to identify beneficial additional evidence, it uses blank-controlled block confidence and retrieval rank to determine whether and which candidate should enter the final context. Across three multimodal DLMs and five visual QA benchmarks, ECF improves answer accuracy by 2.62 percentage points on average over the strongest fixed top-$k$ input and, with LLaDA2.0-Uni, by 2.37 percentage points on average over the best competing training-free result for each dataset. These results show that broader retrieval benefits visual DLM-RAG through selective evidence admission rather than unconditional evidence expansion. Code is publicly available at https://github.com/wjkuser/ECF.
Chinese Translation
视觉检索增强生成(RAG)通常通过扩展检索证据集来提高答案页面的覆盖率,隐含假设所有可用证据都应传递给生成器。我们展示了这一假设对于扩散语言模型(DLMs)并不成立:检索更多页面可以提高答案页面的召回率,而无条件地将所有检索到的页面传递给生成器往往会降低答案的准确性,主要是由于语义冲突。潜在源分析通过并行去噪中的源一致性损失解释了这种不匹配,其中逐位置的提议可能将不兼容的视觉源组合成不支持的答案。我们进一步发现,这种干扰在第一步答案块分布中已经可见,使得在解码之前评估证据成为可能。为了在保留检索覆盖率的同时限制有害的视觉暴露,我们提出了基于熵的候选过滤器(ECF),这是一个无训练的证据接纳框架。为了减少个别候选中的无关内容,ECF构建了多粒度证据单元;为了识别有益的额外证据,它使用空白控制的块置信度和检索排名来决定是否以及哪些候选应进入最终上下文。在三个多模态DLM和五个视觉问答基准上,ECF在最强固定前$k$输入的基础上平均提高了2.62个百分点的答案准确性,并且在LLaDA2.0-Uni上,相较于每个数据集的最佳竞争无训练结果,平均提高了2.37个百分点。这些结果表明,更广泛的检索通过选择性证据接纳而非无条件证据扩展,惠及视觉DLM-RAG。代码可在https://github.com/wjkuser/ECF公开获取。
cs.CL / 36 / 2608.07023
An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation
一种代理性混合的自上而下与自下而上的知识图谱生成方法
Abstract
Organizing thousands of unstandardized, multilingual expertise declarations is a persistent challenge for Human Resources (HR) platforms, directly impacting downstream tasks like accurate talent matching. To address this, we propose a hybrid knowledge graph generation pipeline that grounds a Large Language Model (LLM) in the Wikidata multilingual Knowledge Graph (KG) while employing an agentic reflexion pattern to synthesize emerging concepts and their associated metadata. Unlike rigid top-down methods or fragmented bottom-up approaches, our system anchors recognized concepts to stable Knowledge Graph entities while dynamically creating new nodes and relational metadata for unrecognized skills. Executed across five stages, entity reconciliation, multilingual canonicalization, active curation, deduplication, and the iterative recovery of unmapped concepts, the system autonomously adapts to rapidly evolving, noisy skill mentions across five European languages. Ultimately, this pipeline provides a highly scalable, explicable, and self-healing framework for generating a comprehensive skills knowledge graph, from which a structured taxonomy is derived, using unstructured, noisy text.
Chinese Translation
组织成千上万的不标准化、多语言的专业声明对于人力资源(HR)平台而言是一项持续的挑战,直接影响到准确的人才匹配等下游任务。为了解决这一问题,我们提出了一种混合知识图谱生成管道,该管道将大型语言模型(LLM)与维基数据(Wikidata)多语言知识图谱(KG)相结合,同时采用代理性反思模式来综合新兴概念及其相关元数据。与僵化的自上而下方法或零散的自下而上方法不同,我们的系统将识别出的概念锚定到稳定的知识图谱实体,同时动态创建新的节点和未识别技能的关系元数据。该系统通过实体对齐、多语言规范化、主动策展、去重以及未映射概念的迭代恢复五个阶段执行,能够自主适应五种欧洲语言中快速演变的、噪声较大的技能提及。最终,该管道提供了一个高度可扩展、可解释且自我修复的框架,用于生成全面的技能知识图谱,并从中派生出结构化分类法,使用非结构化的噪声文本。
cs.CL / 37 / 2608.07204
HNR-DAC: Hard-Negative Reranking and Distribution-Aligned Classification for Scientific Claim Verification
HNR-DAC:用于科学声明验证的困难负样本重排序和分布对齐分类
Abstract
Scientific claim verification over a cited paper requires predicting the claim--paper relation and identifying the paragraphs that justify that prediction. This setting poses two linked challenges: within-paper distractors often resemble genuine evidence, while a classifier trained on gold evidence must operate on retrieved evidence at inference. We present HNR-DAC, a two-stage framework that trains each stage on the cases it will actually encounter. Hard-Negative Reranking (HNR) quantifies evidence confusability using a base reranker's scores on non-gold paragraphs and contrasts gold evidence against the most confusable candidates. Distribution-Aligned Classification (DAC) trains on the Top-1 paragraph produced by the same frozen HNR used to construct inference inputs, while HNR's Top-3 paragraph identifiers provide the evidence output. On the NLPCC 2026 Task 10 Track 2, the final configuration obtains 97.21% Hit@3, 95.79% Macro-F1, 94.47% Joint@3, and an average score of 95.13%. The corresponding submission ranks third on the official Track 2 leaderboard while achieving the highest overall Macro-F1 of 93.05%, alongside 70.16% Joint@3 and an average score of 81.61%.
Chinese Translation
对引用论文的科学声明进行验证需要预测声明与论文之间的关系,并识别出支持该预测的段落。这一设置面临两个相关的挑战:论文内部的干扰项往往与真实证据相似,而在推理时,训练于真实证据的分类器必须在检索到的证据上进行操作。我们提出了HNR-DAC,这是一个两阶段框架,针对每个阶段训练其实际会遇到的案例。困难负样本重排序(Hard-Negative Reranking, HNR)通过基于非真实段落的基础重排序器的得分来量化证据的混淆性,并将真实证据与最具混淆性的候选项进行对比。分布对齐分类(Distribution-Aligned Classification, DAC)在由同一冻结的HNR生成的Top-1段落上进行训练,而HNR的Top-3段落标识符提供证据输出。在NLPCC 2026任务10第2轨道中,最终配置获得了97.21%的Hit@3,95.79%的Macro-F1,94.47%的Joint@3,以及95.13%的平均得分。相应的提交在官方第2轨道排行榜中排名第三,同时实现了最高的整体Macro-F1为93.05%,以及70.16%的Joint@3和81.61%的平均得分。
cs.CL / 38 / 2608.07208
Measuring Concept Content in Text from LLM Activations: ESG Evidence from Concept Vectors and Linear Probes
通过大型语言模型激活测量文本中的概念内容:来自概念向量和线性探测的ESG证据
Abstract
Existing measures of how much a text is about a concept read the surface of the text: dictionary word shares, topic proportions, embedding similarities. They score the words a text uses, not the judgment a reader forms about it. Recent work has shown that a gap exists in what Large Language Models (LLMs) know internally versus what they express in their response. This paper asks whether that internal knowledge, read by monitoring the activations of frozen, out-of-the-box LLMs, can stand in for task-specific fine-tuning when measuring concept content, and which extraction method reads it best. We extract such measures via the Recursive Feature Machine (RFM) algorithm and via linear probing, and compare these against an embedding baseline, surface baselines, and the same model's own answer to the question. We demonstrate the approach on financial text, a domain studied extensively and served by established annotated resources, using a human-annotated Environmental, Social and Governance (ESG) dataset. The best linear probe comes within 0.6 percentage points of a fine-tuned domain classifier's accuracy without any task-specific fine-tuning, and outscores the same model's own answer to the question in eleven of twelve comparisons, so the activations carry concept content the response does not report. The simple probe consistently beats the RFM concept vectors, which in turn provide what classification alone does not: a continuous score intended to reflect how strongly a concept is present in a text, whose validation awaits graded labels.
Chinese Translation
现有的衡量文本与某一概念相关程度的方法仅关注文本表面:字典词汇的共享、主题比例、嵌入相似度。这些方法评估文本中使用的词汇,而非读者对文本形成的判断。近期研究表明,大型语言模型(LLMs)内部所掌握的知识与其在响应中表达的内容之间存在差距。本文探讨通过监测冻结的、现成的LLMs的激活,是否可以替代特定任务的微调来测量概念内容,以及哪种提取方法能够最佳地读取这些内容。我们通过递归特征机器(Recursive Feature Machine, RFM)算法和线性探测提取这些测量,并将其与嵌入基线、表面基线以及同一模型对问题的回答进行比较。我们在金融文本这一广泛研究且拥有成熟注释资源的领域展示了该方法,使用了人工注释的环境、社会与治理(Environmental, Social and Governance, ESG)数据集。最佳线性探测的准确率与经过微调的领域分类器相差仅0.6个百分点,且在十二次比较中有十一次超越了同一模型对问题的回答,因此激活所携带的概念内容并未在响应中体现。简单的探测方法始终优于RFM概念向量,而后者提供了分类所无法提供的内容:一个旨在反映概念在文本中存在强度的连续评分,其验证仍需依赖分级标签。
cs.CL / 39 / 2608.07213
From Test-Time Scaling to Reusable Memory: Measuring Crystallization in Text-to-SQL
从测试时刻缩放到可重用内存:测量文本到SQL的结晶化
Abstract
Test-time scaling can correct difficult text-to-SQL queries, but the extra computation is normally discarded after each answer. Systems increasingly retain verified repair episodes, yet evaluations still report one end-to-end score. It cannot distinguish replay on recurring questions from help on unseen questions, or identify the responsible memory choice. We call measuring this future value the crystallization problem. Our controlled evaluation holds the single-shot solver fixed and varies one memory choice at a time. We separately measure replay, cross-question retention, and held-out same-database transfer. On BIRD, storing verified corrected queries improves held-out first-attempt accuracy by 4.34 percentage points. This gain captures 44.4% of the accuracy headroom provided by on-demand repair on the same questions. Controlled interventions identify database-specific content as the main operating ingredient. Reliable verification and broader retrieval coverage yield supported gains; richer formats and elaborate retrievers do not. Open-source code, evaluation artifacts, and reproduction instructions are available at https://github.com/ai-jiaqian/text-to-sql-memory-crystallization.
Chinese Translation
测试时刻缩放可以纠正困难的文本到SQL查询,但额外的计算通常在每次回答后被丢弃。系统越来越多地保留经过验证的修复事件,但评估仍然报告一个端到端的分数。它无法区分对重复问题的重放与对未见问题的帮助,也无法识别负责的内存选择。我们称测量这种未来价值为结晶化问题。我们的控制评估固定单次求解器,同时一次变化一个内存选择。我们分别测量重放、跨问题保留和保留的同数据库迁移。在BIRD上,存储经过验证的纠正查询使得保留的首次尝试准确率提高了4.34个百分点。这一增益捕捉了同一问题上按需修复所提供的准确率提升的44.4%。控制干预确定数据库特定内容是主要的操作成分。可靠的验证和更广泛的检索覆盖带来了支持的增益;更丰富的格式和复杂的检索器则没有。开源代码、评估文档和复现说明可在 https://github.com/ai-jiaqian/text-to-sql-memory-crystallization 获取。
cs.CL / 40 / 2608.07222
Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
Skaling:Chinchilla 的指数与 Kaplan 的耦合
Abstract
Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarce and overtraining extremes. This failure originates in the underlying assumption that model size and training data impact the loss independently. To address this, we introduce the Skaling law, a generalized functional form that couples model capacity and data through a single interaction exponent. This simple extension reduces the Mean Absolute Percentage Error (MAPE) by 1.5-3x across both interpolation and extrapolation regimes. When paired with a sparse grid strategy restricted to low-compute regimes, the Skaling law achieves accurate full-grid extrapolation using approximately 10x less compute than uniform sweeps. By enabling reliable performance prediction from small-scale experiments, the Skaling law provides a more robust and resource-efficient framework for allocating compute budgets in next-generation model training.
Chinese Translation
神经缩放法则是语言模型发展的基础,然而标准公式在数据稀缺和过度训练的极端情况下系统性地低估和高估了损失。这一失败源于一个基本假设,即模型大小和训练数据对损失的影响是独立的。为了解决这个问题,我们引入了 Skaling 法则,这是一种将模型容量和数据通过单一交互指数耦合的广义函数形式。这一简单的扩展在插值和外推两个领域中将平均绝对百分比误差(MAPE)降低了 1.5-3 倍。当与限制在低计算领域的稀疏网格策略相结合时,Skaling 法则能够以大约 10 倍更少的计算量实现准确的全网格外推,而不是均匀扫描。通过使小规模实验的性能预测更加可靠,Skaling 法则为下一代模型训练中的计算预算分配提供了一个更稳健和资源高效的框架。
cs.CL / 41 / 2608.07249
Stoicheia: Character-Level Masked Diffusion for Ancient Greek Textual Restoration, Parsing, and Metrical Scansion
Stoicheia:用于古希腊文本修复、解析和韵律扫描的字符级掩码扩散
Abstract
We introduce Stoicheia, a 405M-parameter character-level masked-diffusion encoder for Ancient Greek whose input factors into five aligned, independently maskable planes: letters, word and sentence boundaries, diacritics, capitalization, and punctuation. A single backbone can therefore restore lacunae, re-segment, accentuate, and punctuate unspaced text without task-specific retokenization. We pretrain it on an open, revision-pinned corpus of 380M words and release eleven checkpoints: ten rotated, decontaminated folds, guaranteeing that for any given literary passage at least one released model has never seen its text, and one with no exposure to documentary texts. Three experiments - reconstruction of damaged inscriptions and papyri, morphosyntactic tagging and dependency parsing, and macronization with metrical scansion - each carry a matched random-initialization control, isolating what character-level diffusion pretraining contributes: 5.6 CER points on inscription reconstruction, 12.9 LAS on parsing, and 6.0 points of balanced accuracy on macronization. On Ithaca's own test split, with identical frozen samples and strict scoring, Stoicheia reduces character error relative to both prior state-of-the-art systems, from 24.6 (Ithaca) and 23.5 (its 2025 Aeneas-framework successor) to 15.5, and raises top-1 accuracy from 63.0 and 64.0 to 74.5.
Chinese Translation
我们介绍了Stoicheia,这是一种具有4.05亿参数的字符级掩码扩散编码器,专为古希腊文本设计,其输入分为五个对齐的、可独立掩蔽的平面:字母、单词和句子边界、变音符号、大写字母和标点符号。因此,单一的主干网络能够在不进行特定任务重标记的情况下,恢复缺失部分、重新分段、加重音符和标点。我们在一个开放的、修订固定的语料库上进行了预训练,该语料库包含3.8亿个单词,并发布了十一份检查点:十个经过旋转和去污染的折叠,确保对于任何给定的文学段落,至少有一个发布的模型从未见过其文本,还有一个模型未接触过文献文本。三个实验——损坏铭文和纸草文的重建、形态句法标注和依赖解析,以及韵母化与韵律扫描——每个实验都配有一个匹配的随机初始化对照,隔离出字符级扩散预训练的贡献:在铭文重建中提高了5.6个字符错误率(CER)点,在解析中提高了12.9个依赖分析得分(LAS),在韵母化中提高了6.0个平衡准确率点。在Ithaca自己的测试集上,使用相同的冻结样本和严格评分,Stoicheia将字符错误率从之前的最先进系统的24.6(Ithaca)和23.5(其2025年Aeneas框架继任者)降低到15.5,并将top-1准确率从63.0和64.0提高到74.5。
cs.CL / 42 / 2608.07261
Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models
为什么仅仅了解两个跳跃是不够的:理解语言模型中的两跳泛化
Abstract
Large language models (LLMs) can solve complex multi-hop problems yet exhibit puzzling failures on simple two-hop queries: although a model may correctly store each individual hop, it often fails to combine them. To understand the internal mechanisms of this phenomenon, we train transformers from scratch in a controlled symbolic environment. Our experiments reveal a pattern in two-hop generalization: models generalize reliably when the second hop follows the training distribution, but always fail when it deviates. Through mechanistic analysis, we provide a complete explanation for these distinct generalization behaviors: in settings where models generalize successfully, performance is driven by the emergence of consistent intermediate representations for the same entities across contexts, whereas failures on settings where the second hop is out-of-distribution arise from a mismatch across layers: lower layers correctly construct these intermediate representations, but upper layers, while trained on corresponding atomic facts, primarily learn to map them to outputs rather than to reason over them. Driven by this insight, we propose a recurrent-style training strategy, which enables transformers to reuse their reasoning circuitry across input forms and substantially improves generalization on out-of-distribution two-hop queries.
Chinese Translation
大型语言模型(LLMs)能够解决复杂的多跳问题,但在简单的两跳查询中却表现出令人困惑的失败:尽管模型可能正确存储每个单独的跳跃,但它往往无法将它们结合起来。为了理解这一现象的内部机制,我们在一个受控的符号环境中从零开始训练变换器。我们的实验揭示了两跳泛化中的一个模式:当第二跳遵循训练分布时,模型能够可靠地泛化,但当其偏离时则总是失败。通过机制分析,我们为这些不同的泛化行为提供了完整的解释:在模型成功泛化的设置中,性能是由相同实体在不同上下文中出现一致的中间表示所驱动,而在第二跳超出分布的设置中失败则源于层之间的不匹配:较低层正确构建这些中间表示,但较高层在训练对应的原子事实时,主要学习将其映射到输出,而不是对其进行推理。基于这一洞察,我们提出了一种递归式训练策略,使变换器能够在不同输入形式之间重用其推理电路,并显著提高在超出分布的两跳查询上的泛化能力。
cs.CL / 43 / 2608.07282
Gaze Behavior in Visual World Experiments Can be Modeled With Off-the-shelf Language-Vision Encoders
视觉世界实验中的注视行为可以通过现成的语言-视觉编码器建模
Abstract
The recent advances in neural language models have also spurred much work in computational psycholinguistics, asking whether neural LMs are also promising models of human language processing. However, work has been overwhelmingly focused on the unimodal case of written or spoken language. In contrast, multimodal experimental paradigms, like visual world studies that present participants with both visual and linguistic input simultaneously, have been neglected. In this paper, we present a novel approach that predicts gaze behavior in visual world studies. It does so by combining a simple multi-modal bi-encoder model of the CLIP family with a bimodal attribution method. We demonstrate the ability of this approach to robustly replicate the results of a seminal English visual world study which shows hu- man predictive processing. Remarkably, it does so without a generative architecture and without the need for fine-tuning, despite not being trained for this task.
Chinese Translation
近年来神经语言模型的进展也推动了计算心理语言学的研究,探讨神经语言模型是否同样是人类语言处理的有效模型。然而,相关研究主要集中在书面或口头语言的单模态案例上。相比之下,像视觉世界研究这样的多模态实验范式同时向参与者呈现视觉和语言输入,却未受到足够重视。本文提出了一种新颖的方法来预测视觉世界研究中的注视行为。该方法通过结合CLIP家族的简单多模态双编码器模型与双模态归因方法来实现。我们展示了该方法能够稳健地复制一项具有开创性的英语视觉世界研究的结果,该研究显示了人类的预测处理能力。值得注意的是,该方法在没有生成架构和不需要微调的情况下实现了这一点,尽管并未针对该任务进行训练。
cs.CL / 44 / 2608.07283
Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram Treebanks
语法工程与大型语言模型的结合:粤语和爱尔兰语ParGram树库的开发
Abstract
Grammar engineering requires expertise in linguistic formalism and computational implementation, especially in parallel grammar projects that balance cross-linguistic consistency with language-specific properties. This paper presents the development of Cantonese and Irish treebanks within the Parallel Grammar (ParGram) Project, where linguistic parallelism is maintained at an abstract functional level. We also investigate the methodological potential and limitations of using multilingual LLMs to support grammar engineering, focusing on Cantonese-Irish translation and the generation of formal syntactic structures using OpenAI's gpt-oss-120b model. The results show that translation performance was generally unsatisfactory and unaffected by prompt language. For syntactic structure generation, the model produced some structurally meaningful outputs, but performed poorly on tasks requiring cross-linguistic abstraction. Nonetheless, LLM-generated outputs may still offer some reference value by suggesting alternative analyses and (partially) capturing predicate-argument relations. Overall, our findings highlight both the potential and limitations of using LLMs in collaborative grammar engineering, while underscoring the continued importance of expert-driven analysis and verification.
Chinese Translation
语法工程需要在语言形式主义和计算实现方面具备专业知识,尤其是在平行语法项目中,需要在跨语言一致性与特定语言属性之间取得平衡。本文介绍了在平行语法(ParGram)项目中开发粤语和爱尔兰语树库的过程,其中在抽象功能层面保持了语言的平行性。我们还探讨了使用多语言大型语言模型(LLMs)支持语法工程的方法论潜力和局限性,重点关注粤语-爱尔兰语翻译以及使用OpenAI的gpt-oss-120b模型生成形式句法结构。结果表明,翻译性能总体上不尽如人意,且不受提示语言的影响。在句法结构生成方面,该模型产生了一些结构上有意义的输出,但在需要跨语言抽象的任务上表现不佳。尽管如此,LLM生成的输出仍可能通过建议替代分析和(部分)捕捉谓词-论元关系提供一定的参考价值。总体而言,我们的研究结果突显了在协作语法工程中使用LLMs的潜力与局限性,同时强调了专家驱动的分析与验证的重要性。
cs.CL / 45 / 2608.07316
Natural Language Processing Psychometrics
自然语言处理心理测量学
Abstract
Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotional content, or syntactic structure. NLP Psychometrics treats psychological prediction from text as a psychometric problem, linking scores to interpretable linguistic evidence and testing beyond the training text format. Nine LLMs, conditioned on controlled personas (cognitive digital shadows), completed psychometric questionnaires with textual explanations per item. We extracted emotional profiles and syntactic-semantic structure via textual forma mentis networks, combined with personality and sociodemographic variables in ablated random forest (RF) regressors, using SHAP to identify which features drove performance and in which direction. Full RF models explained up to 70.8% of variance in life satisfaction (SWLS), 55.7% in depression (PHQ-9), and, for DASS-21, 68.5% depression, 76.0% anxiety, 72.4% stress. Sociodemographics alone explained no meaningful variance in depression, anxiety, or stress, but did so for life satisfaction, where emotion features and income were the strongest predictors; neuroticism and network topology instead dominated depression and anxiety, reversing direction between them. Without retraining, RF models separated diaries from low- and high-score personas ($r$ up to 0.91) and, using only network/emotion features, classified clinical from control participants in real transcripts with up to 68% accuracy. These results show the promise and limits of synthetic data: LLM personas can expose model biases, recover patterns consistent with clinical rumination, and support psychometric prediction from human text without a matched questionnaire, but cannot substitute for human validation. NLP Psychometrics makes these distinctions explicit, measurable, and testable through interpretable AI and network/emotional features.
Chinese Translation
自然语言处理(NLP)模型在预测心理健康结果时很少明确其测量内容:上下文知识、情感内容或句法结构。NLP心理测量学将文本中的心理预测视为一个心理测量问题,将得分与可解释的语言证据联系起来,并在训练文本格式之外进行测试。九个大型语言模型(LLMs)在受控的人格(认知数字影像)条件下,完成了带有文本解释的心理测量问卷。我们通过文本形式思维网络提取了情感特征和句法-语义结构,并结合个性和社会人口变量,使用消融随机森林(RF)回归分析,利用SHAP识别哪些特征驱动了模型表现及其方向。完整的RF模型解释了生活满意度(SWLS)方差的高达70.8%,抑郁(PHQ-9)55.7%,以及DASS-21中抑郁68.5%、焦虑76.0%和压力72.4%。仅社会人口变量对抑郁、焦虑或压力没有解释出有意义的方差,但对生活满意度有解释,其中情感特征和收入是最强的预测因子;而神经质和网络拓扑则主导了抑郁和焦虑,并在两者之间反转方向。在不重新训练的情况下,RF模型能够区分低分和高分人格的日记($r$高达0.91),并且仅使用网络/情感特征,在真实文本中将临床参与者与对照参与者分类,准确率高达68%。这些结果展示了合成数据的潜力和局限性:LLM人格可以揭示模型偏差,恢复与临床反思一致的模式,并支持从人类文本中进行心理测量预测,而无需匹配问卷,但无法替代人类验证。NLP心理测量学使这些区别变得明确、可测量和可测试,通过可解释的人工智能和网络/情感特征。
cs.CL / 46 / 2608.07341
Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination
零差距并非恢复:分层每题概率评估与基准污染的逐步缓解
Abstract
Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. \textbf{Contamination mitigation evaluation} intervenes in the decoding process to suppress memorization and restore a contaminated model's genuine capability, but its prevailing metric, the \textbf{G-AP} (\textbf{G}ap of \textbf{A}ggregate \textbf{P}erformance), is flawed. Discrete correct/incorrect readouts cannot characterize per-question performance, averaging before differencing lets over- and under-suppression cancel out, and uniform per-question weighting invites strategies to push solve probabilities onto the clean model's high-frequency values. We propose \textbf{SA-PPG} (\textbf{S}tratified \textbf{A}ggregate of \textbf{P}er-question \textbf{P}robability \textbf{G}aps): estimate each question's solve probability by sampling, difference it against the clean model per question, and aggregate within groups defined by the clean model's solve probability. Existing mitigation strategies first estimate where contamination lies and then operate on the estimate, so they are only as correct as the estimate. \textbf{RailCap} instead judges contamination during generation: whenever a sample falls back onto the greedy trajectory, the next trajectory token is capped to the runner-up, accumulating suppression until the response distribution becomes sufficiently dispersed. Across multiple contaminated models and benchmarks, SA-PPG reveals that prior strategies' restoration is substantially overestimated, while RailCap attains the lowest SA-PPG.
Chinese Translation
来自公共基准的测试数据不可避免地泄漏到预训练语料库中,一旦被记忆便会膨胀评估分数。\textbf{污染缓解评估}在解码过程中进行干预,以抑制记忆并恢复被污染模型的真实能力,但其主要指标\textbf{G-AP}(\textbf{G}ap of \textbf{A}ggregate \textbf{P}erformance)存在缺陷。离散的正确/错误读数无法表征每题的表现,差异化前的平均会导致过度抑制和不足抑制相互抵消,而统一的每题加权则引入了将解题概率推向干净模型高频值的策略。我们提出了\textbf{SA-PPG}(\textbf{S}tratified \textbf{A}ggregate of \textbf{P}er-question \textbf{P}robability \textbf{G}aps):通过抽样估计每个问题的解题概率,与干净模型的每题进行差异化,并在由干净模型的解题概率定义的组内进行聚合。现有的缓解策略首先估计污染位置,然后在估计值上进行操作,因此它们的正确性仅与估计值的准确性相当。\textbf{RailCap}则在生成过程中判断污染:每当样本回落到贪婪轨迹时,下一轨迹令牌被限制为亚军,积累抑制直到响应分布变得足够分散。在多个被污染模型和基准中,SA-PPG揭示了先前策略的恢复被大幅高估,而RailCap则达到了最低的SA-PPG。
cs.CL / 47 / 2608.07353
Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding
大型语言模型的地理空间概念探测:抽象性、组合性与基础性
Abstract
Understanding concepts is fundamental to generalization. Despite their impressive performance on a wide range of tasks, Large Language Models (LLMs) still struggle with genuine concept understanding. Prior work has evaluated conceptual understanding in LLMs using natural-language benchmarks or narrowly scoped synthetic tasks, but these settings often conflate multiple skills or lack precise control over the underlying concepts and their properties. To support controlled probing of concepts in LLMs, we design tests on their core properties: abstraction, compositionality, and groundness. We set up a concept-centric benchmark, targeting spatial concepts such as direction, distance, topology, and their compositions, and use question answering tasks serving as a proxy. We conduct extensive experiments across multiple LLM architectures and training regimes to analyze how model scale and design impact conceptual understanding. The results reveal clear limitations in current LLMs and provide insights into the factors shaping their ability to acquire and compose structured concepts. Our findings shed light on how concept-based LLMs can be redesigned for improved information access and knowledge management. The code will be available at https://github.com/rd20karim/concept-probing.
Chinese Translation
理解概念是泛化的基础。尽管大型语言模型(LLMs)在广泛任务上的表现令人印象深刻,但它们在真正的概念理解方面仍然存在困难。之前的研究通过自然语言基准或狭窄范围的合成任务评估LLMs的概念理解,但这些设置往往混淆了多种技能,或缺乏对基础概念及其属性的精确控制。为了支持对LLMs中概念的受控探测,我们设计了针对其核心属性的测试:抽象性、组合性和基础性。我们建立了一个以概念为中心的基准,针对空间概念,如方向、距离、拓扑及其组合,并使用问答任务作为代理。我们在多个LLM架构和训练模式下进行了广泛实验,以分析模型规模和设计如何影响概念理解。结果揭示了当前LLMs的明显局限性,并提供了关于影响其获取和组合结构化概念能力的因素的见解。我们的发现为如何重新设计基于概念的LLMs以改善信息获取和知识管理提供了启示。代码将可在 https://github.com/rd20karim/concept-probing 获取。
cs.CL / 48 / 2608.07370
LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering
LitTraceQA:科学问答中多阶段基础和验证的基准
Abstract
Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, but answering research questions from papers requires more than fluent generation. A reliable system must identify the relevant papers, locate the concrete evidence that supports the answer, and produce a response that is faithful to that evidence. We present LitTraceQA, a benchmark for literature-grounded question answering over scientific papers. Given a research question and a metadata pool of papers, a system must return three connected outputs: canonical paper identifiers, supporting evidence locations, and answers in one or more requested formats, including free-form text, multiple-choice answers, and structured tables. LitTraceQA targets evidence types common in scientific reading: tables, figures, text spans, equations or algorithms, and citation contexts. The public development split contains 55 examples, including 26 hidden-source single-paper questions and 29 multi-paper questions, and provides gold papers, evidence annotations, and answers for local validation. We also analyze a larger final annotation collection with 4,978 unique-question records over 4,859 unique gold papers. By evaluating paper retrieval, evidence grounding, and answer accuracy separately, LitTraceQA provides a testbed for scientific QA systems that produce verifiable answers rather than unsupported summaries.
Chinese Translation
科学文献越来越多地被用作语言模型、检索增强生成系统和研究助手的知识来源,但从论文中回答研究问题不仅仅需要流畅的生成。一个可靠的系统必须识别相关论文,定位支持答案的具体证据,并生成忠实于该证据的响应。我们提出了LitTraceQA,这是一个针对科学论文的文献基础问答基准。给定一个研究问题和一组论文的元数据池,系统必须返回三个相互关联的输出:规范的论文标识符、支持证据的位置,以及以一种或多种请求格式(包括自由文本、多项选择答案和结构化表格)提供的答案。LitTraceQA针对科学阅读中常见的证据类型:表格、图形、文本片段、方程或算法以及引用上下文。公共开发分割包含55个示例,包括26个隐藏源的单论文问题和29个多论文问题,并提供金标准论文、证据注释和本地验证的答案。我们还分析了一个更大的最终注释集合,其中包含4,978个独特问题记录,覆盖4,859篇独特的金标准论文。通过分别评估论文检索、证据基础和答案准确性,LitTraceQA为科学问答系统提供了一个测试平台,旨在生成可验证的答案,而不是不支持的摘要。
cs.CL / 49 / 2608.07439
An Exploratory Evaluation of LLM-Assisted Rewriting of Moderate-Complexity Financial Sentences for DisCoCat-Based Sentiment Analysis
基于LLM辅助的中等复杂性金融句子重写的探索性评估:针对DisCoCat的情感分析
Abstract
Quantum natural language processing (QNLP) provides a grammar-aware framework for text modeling, and Distributional Compositional Categorical (DisCoCat) is one of its theoretically grounded formulations. Prior work on financial sentiment analysis has identified practical limitations of DisCoCat, including parser sensitivity, high simulation cost, and difficulty handling longer sentences. We study an LLM-assisted preprocessing workflow that uses controlled rewriting to compress, simplify, or decompose moderate-complexity financial sentiment sentences into parser-compatible, circuit-efficient variants while preserving sentiment-bearing meaning. We compare prompting strategies, language models, and filtering configurations with the low-complexity-only DisCoCat baseline of Stein et al. At the circuit level, the strongest compression variants reduce average qubit and gate counts by more than 70 percent relative to the raw moderate-complexity subset. Across repeated training runs, GPT-4.1-mini with Prompt B achieves the highest observed mean accuracy, $0.550 \pm 0.035$, compared with $0.521 \pm 0.050$ for the baseline. Larger training splits do not necessarily improve downstream performance; across evaluated configurations, training-split size has a moderately negative association with accuracy (Pearson $r=-0.446$). These results provide exploratory evidence that LLM-assisted rewriting can make some moderate-complexity inputs usable within the evaluated DisCoCat configuration, while highlighting prompt design, filtering, and circuit-aware preprocessing as considerations for more scalable QNLP-based financial sentiment analysis.
Chinese Translation
量子自然语言处理(QNLP)为文本建模提供了一个语法感知框架,而分布式组合范畴(DisCoCat)是其理论基础之一。先前关于金融情感分析的研究已识别出DisCoCat的实际局限性,包括解析器敏感性、高仿真成本以及处理较长句子的困难。我们研究了一种LLM辅助的预处理工作流程,该流程利用受控重写将中等复杂性的金融情感句子压缩、简化或分解为与解析器兼容、线路高效的变体,同时保留情感承载的意义。我们比较了提示策略、语言模型和过滤配置,与Stein等人的低复杂性DisCoCat基线进行对比。在电路层面,最强的压缩变体相较于原始的中等复杂性子集,平均量子比特和门的数量减少超过70%。在多次训练运行中,使用提示B的GPT-4.1-mini达到了最高的观察到的平均准确率$0.550 ext{±} 0.035$,而基线为$0.521 ext{±} 0.050$。较大的训练分割不一定能提高下游性能;在评估的配置中,训练分割大小与准确率之间存在适度的负相关(Pearson $r=-0.446$)。这些结果提供了探索性证据,表明LLM辅助的重写可以使一些中等复杂性的输入在评估的DisCoCat配置中可用,同时强调了提示设计、过滤和电路感知预处理作为更具可扩展性的基于QNLP的金融情感分析的考虑因素。
cs.CL / 50 / 2608.07458
CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG
CoinRAG:用于长上下文 RAG 的上下文化信息块 KV 缓存重用
Abstract
Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the coarse-grained chunks. This paper optimizes the Pareto frontier under low prefill latency constraints while maximizing accuracy by proposing CoinRAG (Contextualized Information Nugget KV Cache Reuse for Long-Context RAG). The name metaphorically reflects our core mechanism: much like assembling small tokens (or "coins") to accumulate a larger value, CoinRAG compositionally reuses offline-computed, fine-grained nugget caches to form a learned contextual representation efficiently in a more semantically relevant but compact manner. Specifically, instead of full-chunk encoding, CoinRAG identifies query-relevant semantic units within retrieved chunks through two-stage retrieval and seamlessly assembles their sliced KV representations with a chunk-level context. Extensive evaluations on LongBench multi-hop question answering tasks demonstrate that CoinRAG significantly reduces operational costs and outperforms the other baselines with a new Pareto frontier and an average 5.3% relative improvement in answer quality (F1) under a standard fast prefill latency budget.
Chinese Translation
近期关于检索增强生成(RAG)的优化研究利用了块级 KV 缓存重用,以避免处理长检索上下文,从而提高效率,但在粗粒度块中仍然存在显著的信息冗余和噪声。本文提出了 CoinRAG(用于长上下文 RAG 的上下文化信息块 KV 缓存重用),在低预填充延迟约束下优化了帕累托前沿,同时最大化准确性。该名称隐喻地反映了我们的核心机制:就像将小代币(或“硬币”)组合以积累更大价值一样,CoinRAG 以更语义相关但紧凑的方式组合性地重用离线计算的细粒度信息块缓存,以高效地形成学习的上下文表示。具体而言,CoinRAG 通过两阶段检索识别检索块内与查询相关的语义单元,而不是进行完整块编码,并无缝地将其切片的 KV 表示与块级上下文组装在一起。在 LongBench 多跳问答任务上的广泛评估表明,CoinRAG 显著降低了操作成本,并在标准快速预填充延迟预算下,凭借新的帕累托前沿和平均 5.3% 的答案质量(F1)相对提升,超越了其他基线。
cs.CL / 51 / 2608.07460
CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity
CreativeInstruct:可扩展地教导大型语言模型平衡质量、创造力和多样性
Abstract
While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.g., story generation) as well as those that require it implicitly, e.g., reinforcement learning (RL). We instead propose CreativeInstruct, a scalable instruction-tuning method that teaches LLMs to balance creative, base-model-like generations with the quality of post-trained models, by learning to inject special [StartCreativity] spans that bias generation toward creativity. Furthermore, we introduce a structural diversity metric based on graph edit distance, which captures narrative level variation missed by purely lexical and semantic metrics. On narrative generation, CreativeInstruct matches or exceeds the diversity of both multi-model baselines and distilled variants of their outputs, without sacrificing quality or requiring multiple models at inference time. These results are mirrored in our human evaluation, where we find that annotators rate CreativeInstruct generations as more creative than the post-trained LLMs' generations in 70.3% of cases. We also show the benefits of creative models as a substrate for RL: GRPO applied to a CreativeInstruct checkpoint improves by ~4% on AMC and ~5% points on MATH over the same training applied to the post-trained checkpoint.
Chinese Translation
尽管后训练提高了大型语言模型(LLMs)的能力,但通常会降低其输出的多样性和创造力,这对那些明确要求创造力的任务(例如,故事生成)以及那些隐含要求创造力的任务(例如,强化学习(RL))产生负面影响。我们提出了CreativeInstruct,这是一种可扩展的指令调优方法,通过学习注入特殊的[StartCreativity]跨度,使LLMs能够平衡创造性与后训练模型的质量,从而偏向于创造性生成。此外,我们引入了一种基于图编辑距离的结构多样性度量,它捕捉了纯粹的词汇和语义度量所遗漏的叙事层次变异。在叙事生成方面,CreativeInstruct的多样性与多模型基线和其输出的蒸馏变体相匹配或超过,而不牺牲质量或在推理时需要多个模型。这些结果在我们的人工评估中得到了验证,我们发现注释者在70.3%的情况下将CreativeInstruct生成的内容评为比后训练LLMs生成的内容更具创造性。我们还展示了创造性模型作为强化学习基础的好处:将GRPO应用于CreativeInstruct检查点,在AMC上提高了约4%,在MATH上提高了约5个百分点,相较于对后训练检查点应用相同的训练。