cs.RO / 1 / 2608.02653
Light-Loco-Parkour: Versatile Perceptive Whole-Body Locomotion via Multi-Skill Distillation
轻量级全身运动:通过多技能蒸馏实现多功能感知全身运动
Abstract
Existing humanoid whole-body control systems still fall short of the way humans move through cluttered terrain: they either track expressive whole-body references without terrain generalization, or react to terrain online while leaving the arms, torso, and knees largely unused. We present \texttt{Light-Loco-Parkour} (LLP), an end-to-end perceptive whole-body locomotion system that closes this gap with a single deployable policy. Conditioned only on onboard depth and a velocity command, the policy decides when to walk, balance, climb, step down, or vault, with no reference input, skill label, hand-coded gate, or runtime motion graph. Compared with prior humanoid systems, LLP makes three contributions. First, it introduces a whole-body perceptive-control pipeline that extends an RL-trained, velocity-tracking locomotion policy with parkour skills learned from object-interacting motions, so the same policy tracks velocity in open terrain, executes whole-body traversal at obstacles, and resumes locomotion afterward. Second, it acquires terrain-conditioned skills from sparse seeds by expanding a single motion into dynamically feasible, terrain-paired references across obstacle geometry, rather than relying on a large motion corpus. Third, it learns autonomous skill transitions from reward, letting the policy decide when and which whole-body skill to invoke from depth and command alone, with no one-hot skill label, hand-coded state machine, or runtime motion generator. Simulation and real-world experiments show high success across both benchmarked terrains and unseen obstacle variations, and the same policy transfers zero-shot to indoor and outdoor hardware experiments. These results demonstrate autonomous perceptive whole-body locomotion on a humanoid in outdoor settings, using only onboard sensing and a single deployable policy.
Chinese Translation
现有的人形机器人全身控制系统仍然无法达到人类在复杂地形中移动的方式:它们要么跟踪没有地形泛化的表现性全身参考,要么在在线反应地形时,手臂、躯干和膝盖几乎未被使用。我们提出了 exttt{Light-Loco-Parkour} (LLP),一个端到端的感知全身运动系统,通过一个可部署的策略弥补了这一差距。该策略仅基于机载深度传感器和速度指令,决定何时行走、平衡、攀爬、下台阶或翻越,无需参考输入、技能标签、手动编码的门控或运行时运动图。与之前的人形系统相比,LLP 有三个贡献。首先,它引入了一个全身感知控制管道,该管道扩展了一个经过强化学习训练的速度跟踪运动策略,并结合从物体交互动作中学习的跑酷技能,使得同一策略能够在开放地形中跟踪速度,在障碍物上执行全身穿越,并在之后恢复运动。其次,它通过将单一动作扩展为动态可行的、与地形配对的参考,获取基于地形的技能,而不是依赖于大量的运动库。第三,它通过奖励学习自主技能转换,让策略仅凭深度和指令决定何时以及调用哪种全身技能,无需独热技能标签、手动编码的状态机或运行时运动生成器。仿真和现实世界实验显示,在基准地形和未见障碍变化中都取得了高成功率,并且同一策略能够零样本迁移到室内和室外硬件实验。这些结果展示了在户外环境中,使用仅有的机载传感器和单一可部署策略实现自主感知全身运动的能力。
cs.RO / 2 / 2608.02780
Semantic Haptic Feedback Enhances Dexterous Robotic Teleoperation
语义触觉反馈增强灵巧机器人遥操作
Abstract
In robot teleoperation, haptic feedback can be used to help human operators accomplish dexterous manipulation tasks. However, existing haptic feedback methods try to replicate high-fidelity sensory haptics that are felt in real world interactions, which are constrained by the sensing and feedback hardware capability and may lead to higher workload. To addresses these limitations, this work introduces semantic haptics for teleoperation, which uses abstract haptic patterns to convey critical information about robot states. We categorize robot states into "Confirmations" and "Exceptions", implement a modular haptic rendering pipeline in robot simulation, and deliver semantic haptic feedback to operators through pneumatic and vibrotactile wristbands. This simplifies hardware requirements and enables one-to-many mappings between haptic patterns and robot states. Through three evaluation studies, we identify the most effective semantic haptic design for a common pick and place teleoperation task and compare semantic haptics to other teleoperation feedback approaches including sensory haptics and visual feedback. Results suggest that while semantic haptics performs similarly as other feedback in unimanual tasks, it achieves superior performance in bimanual tasks, with reduced task workload, increased situational awareness, and overall preference.
Chinese Translation
在机器人遥操作中,触觉反馈可以帮助人类操作员完成灵巧的操作任务。然而,现有的触觉反馈方法试图复制在现实世界交互中感受到的高保真感官触觉,这受到传感和反馈硬件能力的限制,并可能导致更高的工作负荷。为了解决这些限制,本研究引入了用于遥操作的语义触觉,它使用抽象的触觉模式传达关于机器人状态的关键信息。我们将机器人状态分类为“确认”和“异常”,在机器人仿真中实现了模块化的触觉渲染管道,并通过气动和振动腕带向操作员提供语义触觉反馈。这简化了硬件要求,并实现了触觉模式与机器人状态之间的一对多映射。通过三项评估研究,我们确定了在常见的拾取和放置遥操作任务中最有效的语义触觉设计,并将语义触觉与其他遥操作反馈方法(包括感官触觉和视觉反馈)进行了比较。结果表明,虽然在单手任务中,语义触觉的表现与其他反馈相似,但在双手任务中,它的表现更为优越,工作负荷降低,情境意识提高,整体偏好也更高。
cs.RO / 3 / 2608.02809
Toward Certified Functional Safety for Industrial Humanoid Robots: The Fail-Passive Gap and a Feasibility Study
朝向工业类人机器人认证功能安全的研究:失效被动差距及可行性研究
Abstract
Industrial humanoid robots are constrained less by locomotion or manipulation capability than by the immaturity of functional safety certification for legged platforms. The root difficulty is that the safe state of a legged robot is an actively-controlled state, which violates the fail-passive assumption underlying ISO~13849-1 / EN~60204-1: removing power from a walking biped causes an uncontrolled fall, so classical de-energization is itself a hazard. We term this the fail-passive gap and use a certified external safety chain (light curtain, emergency stop, fail-safe input, fail-safe PLC, and wireless PROFIsafe) as an instrument to locate it precisely: because the external chain is closed and quantifiable with established methods (PFHD, DC, CCF, PL/SILCL), the residual uncertifiable element is pinpointed to the robot-side reaction chain. Using a Siemens fail-safe S7-1500 emergency-stop reference, we show its certifiable Reaction subsystem is contactor-based power removal (Stop Category~0)---exactly the element a balancing humanoid cannot have. We deliberately do not claim end-to-end certified PL~e / SIL~3. We validate the approach on a Unitree G1 EDU pick-and-place cell in a 3m x 1.5m semi-enclosed workspace, and contribute a humanoid-specific analysis of the active safe state (fall-as-hazard, single-support stop bounds, balancing-policy residual risk, ISO~13855 separation) and a provenance-labeled timing budget. Hosting an industrial software-defined automation (SDA) controller on the robot, co-located with the balancing policy, moves robot-side PROFINET/PROFIsafe reception onto a standardized IEC~61131-3 interface; because the G1's onboard compute is not safety-rated hardware, this endpoint is not a certified safety runtime, which reinforces rather than resolves the fail-passive gap and localizes it to the SDA-to-balancing-policy interface.
Chinese Translation
工业类人机器人在运动或操作能力方面的限制,主要源于其腿部平台功能安全认证的不成熟。根本问题在于,腿部机器人的安全状态是一个主动控制的状态,这违反了ISO 13849-1 / EN 60204-1所依据的失效被动假设:对行走的双足机器人断电会导致失控坠落,因此经典的断电措施本身就是一种危险。我们将此称为失效被动差距,并使用认证的外部安全链(光幕、紧急停止、失效安全输入、失效安全PLC和无线PROFIsafe)作为精确定位的工具:由于外部链是封闭的,并且可以使用既定方法(PFHD、DC、CCF、PL/SILCL)量化,因此剩余的不可认证元素被明确定位于机器人侧的反应链。通过使用西门子失效安全S7-1500紧急停止参考,我们展示了其可认证的反应子系统是基于接触器的电源切断(停止类别0)——这正是一个平衡类人机器人所不能具备的元素。我们故意不声称实现端到端的认证PL e / SIL 3。我们在一个3米 x 1.5米的半封闭工作空间中,对Unitree G1 EDU的拣选与放置单元验证了该方法,并贡献了针对主动安全状态的类人特定分析(坠落作为危险、单支撑停止界限、平衡策略剩余风险、ISO 13855分离)及其来源标记的时间预算。将工业软件定义自动化(SDA)控制器托管在机器人上,并与平衡策略共同定位,将机器人侧的PROFINET/PROFIsafe接收移至标准化的IEC 61131-3接口;由于G1的板载计算不是安全认证硬件,因此该端点不是认证的安全运行时,这进一步强化而非解决了失效被动差距,并将其局限于SDA与平衡策略接口之间。
cs.RO / 4 / 2608.02811
Staying on Spec: Real-Time Monitoring under Uncertainty with a Maritime Case Study
保持规范:在不确定性下的实时监测与海事案例研究
Abstract
Robotic systems must operate under uncertainty while satisfying complex task and safety specifications. Monitoring such specifications under uncertainty remains challenging, as existing formulations typically require extensive data or explicit uncertainty distributions. In this paper, we propose a real-time monitoring framework that reduces data requirements by leveraging data-driven reachable sets for specification evaluation. We instantiate the framework for maritime navigation, where complex specifications arise from traffic rules. We develop a data-efficient pipeline for constructing reachable sets and derive a monitoring formulation suitable for real-time deployment. Simulation and hardware experiments demonstrate robust monitoring under realistic disturbances, achieving improved risk detection compared to state-of-the-art metrics.
Chinese Translation
机器人系统必须在不确定性下操作,同时满足复杂的任务和安全规范。在不确定性下监测这些规范仍然具有挑战性,因为现有的公式通常需要大量数据或明确的不确定性分布。在本文中,我们提出了一种实时监测框架,通过利用数据驱动的可达集来减少规范评估的数据需求。我们为海事导航实例化该框架,其中复杂的规范源自交通规则。我们开发了一种数据高效的管道来构建可达集,并推导出适合实时部署的监测公式。仿真和硬件实验表明,在现实干扰下实现了稳健的监测,与最先进的指标相比,风险检测得到了改善。
cs.RO / 5 / 2608.02834
Biconvex Optimization for Smooth Minimum-Time Trajectories around Convex Obstacles
双凸优化在凸障碍物周围的平滑最短时间轨迹规划
Abstract
We present a biconvex approach for minimum-time motion planning around convex obstacles that is guaranteed to converge, is anytime, and supports derivative constraints to arbitrary order. We jointly convexify the minimum-time objective and all derivative constraints through a change of variables, and handle collision avoidance via time-varying separating planes, reducing the problem to a biconvex program. This program is solved by alternating between computing maximum-margin separating planes and optimizing the trajectory. By only adding planes for obstacles that the current iterate collides with, the trajectory can jump around obstacles and escape local minima. The method is guaranteed to converge starting from a simple collision-free polygonal curve. In our experiments on drone navigation and dual-arm bin unloading, we find that the proposed method reliably produces high-quality trajectories with computation times comparable to state-of-the-art decomposition-based motion planners, while handling a larger class of problems and being substantially more robust to bad initialization. Project page:https://wernerpe.github.io/bmtp-website/
Chinese Translation
我们提出了一种双凸方法,用于在凸障碍物周围进行最短时间运动规划,该方法保证收敛、具备随时可用性,并支持任意阶的导数约束。通过变量变换,我们将最短时间目标和所有导数约束共同凸化,并通过时变分离平面处理碰撞避免,将问题简化为一个双凸程序。该程序通过交替计算最大间隔分离平面和优化轨迹来求解。通过仅为当前迭代与之发生碰撞的障碍物添加平面,轨迹可以绕过障碍物并逃离局部最小值。该方法保证从简单的无碰撞多边形曲线开始收敛。在我们对无人机导航和双臂卸货的实验中,发现该方法可靠地产生高质量的轨迹,其计算时间与最先进的基于分解的运动规划器相当,同时处理更大范围的问题,并对不良初始化具有显著更强的鲁棒性。项目页面:https://wernerpe.github.io/bmtp-website/
cs.RO / 6 / 2608.02886
Control Barrier Functions via Minkowski Operations for Safe Navigation among Polytopes
通过闵可夫斯基运算的控制障碍函数实现多面体间的安全导航
Abstract
Safely navigating polytopic environments while respecting the dynamics, control, and exact geometry of the underlying system is a challenge in robotics. Control barrier functions (CBFs) synthesize safe control policies by rendering the safe set forward invariant, but many existing CBF-based methods approximate polytopes using conservative smooth shapes, such as spheres or ellipsoids, to obtain explicit differentiable distance functions. In this article, we propose an exact Signed Distance Function (SDF) formulation for a {\it polytopic} robot and {\it polytopic} obstacles and integrate it with nonsmooth CBFs. Leveraging Minkowski operations, the proposed method computes the exact SDF via companion convex programs in both the collision-free (positive-sign) and in-collision (negative-sign) cases. Furthermore, by exploiting the convenient geometric properties of 2D Minkowski operations and the optimality conditions of the two companion convex programs, we derive a unified analytical expression for the gradient of the exact SDF via sensitivity analysis. The exact rotational gradient further reveals a previously masked class of local minima induced by the coupling between geometry and nonholonomic kinematics. We demonstrate the effectiveness of the proposed framework through a pure-translation case and three scenarios with unicycle models involving recovery from an unsafe initialization and single- and multiple-obstacle avoidance. Comparisons with baseline methods highlight how the proposed framework enables non-conservative maneuvers and safety recovery.
Chinese Translation
在机器人技术中,安全地导航于多面体环境,同时尊重底层系统的动力学、控制和精确几何形状,是一项挑战。控制障碍函数(CBFs)通过使安全集保持前向不变来合成安全控制策略,但许多现有的基于CBF的方法使用保守的光滑形状(如球体或椭球体)来近似多面体,以获得显式可微的距离函数。在本文中,我们为多面体机器人和多面体障碍物提出了一种精确的有符号距离函数(SDF)公式,并将其与非光滑CBFs结合。利用闵可夫斯基运算,所提出的方法通过伴随的凸优化程序在无碰撞(正号)和碰撞(负号)情况下计算精确的SDF。此外,通过利用二维闵可夫斯基运算的便利几何特性和两个伴随凸优化程序的最优性条件,我们通过灵敏度分析推导出精确SDF梯度的统一解析表达式。精确的旋转梯度进一步揭示了由几何与非完整动力学之间的耦合引起的先前被掩盖的局部极小值类。我们通过一个纯平移案例和三个涉及从不安全初始化恢复以及单个和多个障碍物规避的单轮车模型场景,展示了所提出框架的有效性。与基线方法的比较突显了所提出框架如何实现非保守的机动和安全恢复。
cs.RO / 7 / 2608.02895
Contact-Driven Localization in a Freeform Robotic Self-Assembled Structure
基于接触驱动的自由形态机器人自组装结构定位
Abstract
Accurate localization remains a key challenge in swarm robotics, particularly for self-reconfigurable systems that must identify relative positions to form diverse structures. Most existing approaches rely on external tracking infrastructure or high-cost sensors, which limit scalability and deployment in unstructured environments. In this paper, we propose a novel contact-driven localization method for modular robots that leverages only local communication through binary contact information (whether two robots are physically connected or not). To exploit these contact cues, we introduce a virtual-force framework in which robots iteratively refine their poses attracting toward dock-connected neighbors and repelling from non-connected ones. The method requires no external infrastructure and relies only on minimal onboard sensing. Simulations show effective localization during the assembly of towers and cantilevers, enabling accurate, scalable, free-form self-assembly.
Chinese Translation
准确定位仍然是群体机器人领域的一项关键挑战,尤其对于必须识别相对位置以形成多样结构的自重构系统。现有的大多数方法依赖于外部跟踪基础设施或高成本传感器,这限制了其在非结构化环境中的可扩展性和部署能力。本文提出了一种新颖的基于接触驱动的模块机器人定位方法,仅利用通过二进制接触信息(两个机器人是否物理连接)进行局部通信。为了利用这些接触线索,我们引入了一个虚拟力框架,在该框架中,机器人通过向连接的邻居吸引和从未连接的邻居排斥的方式迭代地优化其姿态。该方法不需要外部基础设施,仅依赖于最小的机载传感器。仿真结果显示,在组装塔和悬臂的过程中,该方法实现了有效的定位,支持准确、可扩展的自由形态自组装。
cs.RO / 8 / 2608.02904
DeRP: An Algorithm for Self-Assembly of Power-Delivery Networks using Recursive Branching in Information-Limited Environments
DeRP:一种在信息受限环境中使用递归分支自组装电力传输网络的算法
Abstract
Delivering sustained power to distributed equipment in unstructured field environments using pre-planned wired networks or battery-based solutions presents significant infrastructure and logistics challenges. This paper presents Dendritic Recursive Pivoting (DeRP), a decentralized framework for multi-target network formation in robot swarms based solely on local communication and bearing-based sensing toward sinks. We envision a system in which robots, acting as a conduit, self-assemble a power network from a common source, forming branches at locally selected pivot points that approximate the Steiner points of Steiner trees to efficiently route to multiple Sinks. This branching operation is performed recursively to enable scalable and adaptive network formation without global planning. The proposed method is evaluated in terms of the total network length and estimated power loss, and is quantitatively compared against global baselines such as the Minimum Spanning Tree and Steiner tree solutions (GeoSteiner), which require complete knowledge of Sink locations. Specifically, we found that the networks formed by DeRP asymptotically form approximately 125\% of the global minimum length while reducing power losses to 65\% relative to Euclidean Steiner trees. In addition, we empirically characterize scaling behavior by measuring simulation completion time as the number of Sinks and robots increases, and find that this scaling was sub-linear for up to 100 sinks. The proposed approach enables resilient, adaptive power delivery in environments where deployment of traditional infrastructure is challenging.
Chinese Translation
在非结构化现场环境中,通过预先规划的有线网络或基于电池的解决方案为分布式设备提供持续电力面临着重大的基础设施和物流挑战。本文提出了树状递归枢轴(Dendritic Recursive Pivoting,DeRP),这是一个基于局部通信和基于方向感知的去中心化框架,用于机器人群体中的多目标网络形成。我们设想一个系统,其中机器人作为导管,从一个共同源自组装电力网络,在局部选择的枢轴点形成分支,这些枢轴点近似于斯坦纳树的斯坦纳点,以有效地路由到多个接收点(Sinks)。该分支操作以递归方式进行,以实现可扩展和自适应的网络形成,而无需全球规划。所提出的方法在总网络长度和估计功率损耗方面进行了评估,并与需要完全了解接收点位置的全球基准(如最小生成树和斯坦纳树解决方案(GeoSteiner))进行了定量比较。具体而言,我们发现DeRP形成的网络渐近地形成了全球最小长度的约125%,同时将功率损耗减少至相对于欧几里得斯坦纳树的65%。此外,我们通过测量仿真完成时间来经验性地表征缩放行为,随着接收点和机器人数量的增加,发现这种缩放在最多100个接收点时是亚线性的。所提出的方法使在传统基础设施部署具有挑战性的环境中实现弹性和自适应的电力传输成为可能。
cs.RO / 9 / 2608.02958
ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies
ValueFormer:一种具有阶段感知标签的因果变换器价值函数,用于半自主视觉-语言-动作策略
Abstract
Vision-Language-Action (VLA) policies trained by behavior cloning fail silently: from the action stream alone, a collapsing rollout looks much like one making clean progress, because imitation supplies no notion of progress. Reinforcement learning would supply one, but it is impractical here, where real-robot experience is costly and deformable food resists simulation. The cheap alternative, a terminal success / failure bit, is learnable in principle yet far too sparse to say when a rollout went wrong. We argue that the per-frame label, not the architecture, is the hard part: to be useful it must be dense, continuous, and correctly shaped. We present ValueFormer, a compact policy-agnostic causal transformer over a frozen DINOv3 backbone that emits two per-frame signals in one forward pass: a smooth Monte Carlo value, V_mc, for advantage estimation and a sharp binary value for online mistake detection, targets that pull in opposite directions by design. Failed episodes are labeled with a stage-aware, success-then-decay return that preserves the success curve before the failure stage, and detection is supervised from mistake intervals rather than a single failure time, so mistakes the policy recovers from also carry signal. On a real-robot bimanual sandwich-assembly task 1,427 episodes), a critic-derived per-frame training weight lifts task completion from 70% to 85% (within noise at n=20), and a batched bf16 encoder cuts the live serving cost 3~5 times so the critic runs at 2 Hz alongside the policy on a single GPU.
Chinese Translation
通过行为克隆训练的视觉-语言-动作(VLA)策略在失败时表现得很隐蔽:仅从动作流来看,崩溃的回放与清晰进展的回放看起来非常相似,因为模仿并未提供进展的概念。强化学习可以提供这一点,但在这里并不实用,因为真实机器人经验成本高昂,而可变形食物又难以模拟。便宜的替代方案是终端成功/失败位,原则上是可以学习的,但过于稀疏,无法准确判断回放何时出错。我们认为每帧标签,而非架构,才是难点:为了有用,它必须是密集的、连续的,并且形状正确。我们提出了ValueFormer,这是一种紧凑的与策略无关的因果变换器,基于冻结的DINOv3主干,在一次前向传递中发出两个每帧信号:用于优势估计的平滑蒙特卡洛值V_mc,以及用于在线错误检测的尖锐二元值,这两个目标在设计上是相互牵制的。失败的回合被标记为具有阶段感知的成功后衰减回报,保留了失败阶段之前的成功曲线,检测是从错误间隔进行监督,而不是单一的失败时间,因此策略恢复的错误也携带信号。在一个真实机器人双手三明治组装任务中(1,427个回合),基于评论员推导的每帧训练权重将任务完成率从70%提升至85%(在n=20的噪声范围内),而一个批处理bf16编码器将实时服务成本降低了3~5倍,使得评论员与策略在单个GPU上以2 Hz的速度同时运行。
cs.RO / 10 / 2608.02990
EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation
EmbodiedVAE:用于高效可控的具身操控的解耦视频变分自编码器
Abstract
Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable performance, existing LDMs predominantly rely on Variational Autoencoders (VAEs) optimized for natural scenes while failing to account for the unique characteristics of embodied manipulation scenarios, yielding latent representations that are neither compact nor controllable, thereby hindering efficient training of LDMs and precise robotic control. To solve this problem, we present EmbodiedVAE, a novel video VAE that provides compact yet controllable latent representations tailored for the robotic manipulation world models. Specifically, EmbodiedVAE adopts a dual-encoder, single-decoder architecture with an asymmetric spatio-temporal compression module, which automatically disentangles the robot arm's motion from background environment, resulting in overall compactness while providing explicit embodied latent to support fine-grained action control. To further preserve the temporal consistency of learned robotic motion latent, we introduce an optimal-transport-based consistency module that explicitly enforces motion fidelity and inter-frame coherence. Extensive experiments demonstrate that our proposed EmbodiedVAE achieves superior reconstruction quality with high compression rate, while enabling more precise action control in robotic manipulation scenarios with an average of 2dB PSNR improvement over state-of-the-art video VAEs.
Chinese Translation
潜在扩散模型(LDMs)最近在构建强大的具身操控世界模型方面取得了显著进展。然而,尽管性能卓越,现有的LDMs主要依赖于针对自然场景优化的变分自编码器(VAEs),未能考虑具身操控场景的独特特征,导致潜在表示既不紧凑也不易于控制,从而阻碍了LDMs的高效训练和精确的机器人控制。为了解决这个问题,我们提出了EmbodiedVAE,一种新颖的视频变分自编码器,提供紧凑且可控的潜在表示,专为机器人操控世界模型量身定制。具体而言,EmbodiedVAE采用双编码器、单解码器架构,并配备不对称时空压缩模块,能够自动将机器人手臂的运动与背景环境解耦,从而实现整体紧凑性,同时提供明确的具身潜在表示以支持细粒度的动作控制。为了进一步保持学习到的机器人运动潜在的时间一致性,我们引入了一种基于最优传输的一致性模块,明确地强制执行运动保真性和帧间一致性。大量实验表明,我们提出的EmbodiedVAE在高压缩率下实现了优越的重建质量,同时在机器人操控场景中实现了更精确的动作控制,相较于最先进的视频变分自编码器平均提高了2dB的峰值信噪比(PSNR)。
cs.RO / 11 / 2608.03002
A Wearable Stiffness-Rendering Haptic Device with a Honeycomb Jamming Mechanism for Bilateral Teleoperation
一种具有蜂窝阻塞机制的可穿戴刚度呈现触觉设备用于双向遥操作
Abstract
This paper addresses the challenge of providing kinesthetic feedback in bilateral teleoperation by designing a wearable, lightweight (20 g), and compact haptic device, the HJ-Haptic, utilizing a honeycomb jamming mechanism for object stiffness rendering. The HJ-Haptic device can vary its stiffness, from 1.15 N/mm to 2.64 N/mm, using a 30 kPa vacuum pressure. We demonstrate its implementation in a teleoperation framework, enabling operators to adjust grip force based on a reliable haptic feedback on object stiffness. A three-point flexural test on the honeycomb jamming mechanism and teleoperated object-grasping tasks were conducted to evaluate the device's functionality. Our experiments demonstrated a small RMSE and strong correlations in teleoperated motion, stiffness rendering, and interaction force feedback. The HJ-Haptic effectively adjusts its stiffness in response to real-time gripper feedback, mimicking the sensation of direct object grasping with hands. The device's use of vacuum pressure ensures operator safety by preventing dangerous outcomes in case of gas leakage or material failure. Incorporating the HJ-Haptic into the teleoperation framework provided the reliable perception of object stiffness and stable teleoperation. This study highlights the potential of the honeycomb jamming mechanism for enhancing haptic feedback in various applications, including teleoperation scenarios, as well as interactions with extended-reality environments.
Chinese Translation
本文通过设计一种可穿戴、轻量(20克)且紧凑的触觉设备HJ-Haptic,解决了双向遥操作中提供动觉反馈的挑战,该设备利用蜂窝阻塞机制进行物体刚度呈现。HJ-Haptic设备能够在30 kPa的真空压力下调整其刚度,范围从1.15 N/mm到2.64 N/mm。我们在遥操作框架中展示了其实现,允许操作员根据物体刚度的可靠触觉反馈来调整握持力。我们进行了蜂窝阻塞机制的三点弯曲测试和遥操作物体抓取任务,以评估设备的功能性。实验结果显示,在遥操作运动、刚度呈现和交互力反馈方面具有较小的均方根误差(RMSE)和强相关性。HJ-Haptic能够有效地根据实时抓取反馈调整其刚度,模拟用手直接抓取物体的感觉。该设备使用真空压力确保操作员安全,防止气体泄漏或材料故障时产生危险后果。将HJ-Haptic纳入遥操作框架中提供了物体刚度的可靠感知和稳定的遥操作。该研究突显了蜂窝阻塞机制在增强各种应用中的触觉反馈潜力,包括遥操作场景以及与扩展现实环境的交互。
cs.RO / 12 / 2608.03010
Forbidden Region Dynamic Active Constraints in Robot-Assisted Minimally Invasive Surgery
机器人辅助微创手术中的禁区动态主动约束
Abstract
In robot-assisted surgery, Forbidden Region Active Constraints (FRAC) represent a control strategy that helps maintain task safety by generating anisotropic haptic guidance to surgeons. However, several challenges need to be overcome before FRAC can benefit teleoperative surgery in a clinical setting. These challenges include the ability to allow for dynamic tissue deformation, maintain energetic passivity, and speed of implementation, among others. In this study, we propose the pipeline design for an energy dissipative FRAC strategy, which accommodates the dynamic tissue deformation caused by respiratory movements, by utilizing a depth sensing camera. The proposed FRAC strategy adopts a fine mesh representation, with a total number of 122,806 polygons in the case study presented, while running at 43.48Hz. We designed in vitro trajectory tracking experiments conducted by a "virtual" surgeon to aid quantitative assessment of the method, including its effectiveness in maintaining task safety, which was confirmed by successfully maintaining a pre-defined safety distance across all trials. We also conducted comparative studies to investigate the robustness and time-efficiency of our method against other FRAC methods that rely on simple geometry AC representations. We demonstrate that our method provides a more robust and effective guidance overall, while maintaining comparable, if not lower, time costs.
Chinese Translation
在机器人辅助手术中,禁区主动约束(Forbidden Region Active Constraints, FRAC)是一种控制策略,通过为外科医生提供各向异性的触觉引导来帮助维持任务安全。然而,在FRAC能够在临床环境中为远程手术带来益处之前,需要克服几个挑战。这些挑战包括允许动态组织变形、保持能量被动性以及实施速度等。在本研究中,我们提出了一种能量耗散FRAC策略的管道设计,该策略通过利用深度传感相机来适应由呼吸运动引起的动态组织变形。所提出的FRAC策略采用细网格表示,在所呈现的案例研究中总共有122,806个多边形,并以43.48Hz的频率运行。我们设计了由“虚拟”外科医生进行的体外轨迹跟踪实验,以帮助定量评估该方法的有效性,包括其在维持任务安全方面的有效性,这通过在所有试验中成功维持预定义的安全距离得到了确认。我们还进行了比较研究,以调查我们的方法相对于依赖简单几何主动约束表示的其他FRAC方法的鲁棒性和时间效率。我们证明了我们的方法在整体上提供了更鲁棒和有效的引导,同时保持了可比的,甚至更低的时间成本。
cs.RO / 13 / 2608.03034
PACE: Adaptive Budget Allocation for Time-Efficient Embodied Planning
PACE:用于时间高效的具身规划的自适应预算分配
Abstract
Reasoning-enhanced large language models have achieved remarkable improvements in planning tasks, yet their deployment in embodied systems remains impractical due to prohibitive inference delays-often exceeding minutes per planning instance. The fundamental bottleneck stems from the serial nature of existing paradigms: models must complete all reasoning before any action execution, leaving execution time windows entirely unexploited. We introduce PACE (Planning with Adaptive Cognitive Effort), a framework that enables interleaved reasoning and execution through two key innovations: an Interleaved Think-Act architecture that pipelines cognitive processing with action execution, and a Dynamic Budget Allocator that adapts reasoning token budgets to available execution time windows. On the Robotouille benchmark using Qwen3-8B-AWQ, PACE achieves a 10% success rate-representing a 67% improvement over the ReAct+Think baseline-while delivering 6.9 times acceleration in thinking time compared to unconstrained reasoning. The framework hides 66.8% of thinking time within execution windows, demonstrating that strategic cognitive effort allocation can simultaneously improve both planning quality and time efficiency. These results provide evidence that time-aware architectural innovations enable reasoning models to operate in latency-sensitive embodied domains where they were previously impractical.
Chinese Translation
增强推理的大型语言模型在规划任务中取得了显著进展,但由于推理延迟过高(通常超过每个规划实例几分钟),其在具身系统中的应用仍然不切实际。根本瓶颈源于现有范式的串行特性:模型必须在执行任何动作之前完成所有推理,从而完全未利用执行时间窗口。我们提出了PACE(Planning with Adaptive Cognitive Effort),这是一个通过两个关键创新实现推理与执行交错进行的框架:一种将认知处理与动作执行流水线化的交错思考-行动架构,以及一个根据可用执行时间窗口调整推理令牌预算的动态预算分配器。在使用Qwen3-8B-AWQ的Robotouille基准上,PACE实现了10%的成功率,相较于ReAct+Think基线提高了67%,同时在思考时间上实现了6.9倍的加速,相比于不受限制的推理。该框架在执行窗口内隐藏了66.8%的思考时间,证明了战略性认知努力分配可以同时提高规划质量和时间效率。这些结果提供了证据,表明时间敏感的架构创新使推理模型能够在以前不切实际的延迟敏感的具身领域中运行。
cs.RO / 14 / 2608.03051
CUDA MPC: A GPU-Native Solver for Model Predictive Control
CUDA MPC:一种基于GPU的模型预测控制求解器
Abstract
Model Predictive Control (MPC) delivers constraint-aware control, but its reliance on online optimization limits its use on systems with fast dynamics, high-dimensional models, or long horizons. Existing GPU implementations typically treat the device as a linear-algebra accelerator, leaving the optimization loop dependent on repeated kernel launches and high-latency memory transfers. This paper introduces CUDA MPC, a GPU-native MPC framework that co-designs the optimization algorithm, execution model, and memory architecture for CUDA hardware. CUDA MPC pairs a parallel-in-horizon alternating direction method of multipliers (ADMM) splitting with a fused CUDA kernel that runs the entire iterative solve on the device. Intermediate optimization variables stay in low-latency, on-chip shared memory, and a localized atomic-flag protocol synchronizes only adjacent horizon blocks, minimizing host intervention, kernel-dispatch overhead, and global-memory traffic. Across six nonlinear robotics benchmarks spanning increasing state dimension and constraint density, CUDA MPC sustains real-time rates at horizons one to two orders of magnitude longer than CPU solvers: it solves an optimization-based collision-avoidance parking problem with 100 s of lookahead within a 0.1 s sampling interval, and is the only solver evaluated that achieves both real-time execution and collision-free coordination for a centralized 10-agent swarm, where acados and CasADi return no feasible solution and require 3.5 s and 4.5 s per solve. Against tensor-framework implementations of the same ADMM splitting, the fused kernel is up to $965\times$ faster.
Chinese Translation
模型预测控制(MPC)提供了考虑约束的控制,但其对在线优化的依赖限制了其在快速动态、高维模型或长时间范围系统中的应用。现有的GPU实现通常将设备视为线性代数加速器,使得优化循环依赖于重复的内核启动和高延迟的内存传输。本文介绍了CUDA MPC,一种针对CUDA硬件的本地MPC框架,旨在共同设计优化算法、执行模型和内存架构。CUDA MPC将并行时间域交替方向乘子法(ADMM)分解与一个融合的CUDA内核相结合,该内核在设备上运行整个迭代求解过程。中间优化变量保留在低延迟的片上共享内存中,并且局部原子标志协议仅同步相邻的时间域块,从而最小化主机干预、内核调度开销和全局内存流量。在六个非线性机器人基准测试中,随着状态维度和约束密度的增加,CUDA MPC在比CPU求解器长一个到两个数量级的时间范围内维持实时速率:它在0.1秒的采样间隔内解决了一个基于优化的避碰停车问题,具有100秒的前瞻时间,并且是唯一一个在集中式10代理群体中实现实时执行和无碰撞协调的求解器,而acados和CasADi返回的都是不可行解,并且每次求解需要3.5秒和4.5秒。与相同ADMM分解的张量框架实现相比,融合内核的速度提高了多达965倍。
cs.RO / 15 / 2608.03052
How Should Vision-Language-Action Models Use Proprioceptive State?
视觉-语言-动作模型应如何利用本体状态?
Abstract
Recent Vision-Language-Action (VLA) models almost universally take robot proprioceptive state as input, yet wire it in incompatible ways -- serialized into text prompts, projected into the vision-language prefix, or fed directly to the action expert -- and almost always as a single current frame. Three questions remain open: (1) whether, and on which tasks, current state actually improves closed-loop control; (2) how much state history helps, and whether its benefit reflects genuine temporal variation rather than added conditioning capacity; and (3) where state should enter the model -- the vision-language backbone or the action-generation module. We answer these questions through controlled experiments on a flow-matching VLA, fixing the backbone, training data, action representation, and evaluation protocol throughout. We implement five representative interfaces -- discrete state prompt, VLM prefix, action prefix, state expert, and feature modulation -- under matched implementation details, and evaluate them on 45 atomic tasks spanning three task families plus 20 composite tasks; we then sweep the state-history length from 1 to 96 frames to examine how historical state information affects model performance. The experiments yield systematic answers to all three questions, distilled into testable design principles for state-aware VLAs.
Chinese Translation
近期的视觉-语言-动作(VLA)模型几乎普遍将机器人本体状态作为输入,但以不兼容的方式进行连接——将其序列化为文本提示、投影到视觉-语言前缀中,或直接输入到动作专家中——且几乎总是作为单一的当前帧。仍然存在三个未解的问题:(1)当前状态是否以及在哪些任务上确实改善了闭环控制;(2)状态历史的帮助程度,以及其益处是否反映了真正的时间变化而非增加的条件能力;(3)状态应在模型的何处进入——视觉-语言主干还是动作生成模块。我们通过对流匹配VLA进行的控制实验来回答这些问题,固定主干、训练数据、动作表示和评估协议。我们在匹配的实现细节下实现了五种代表性接口——离散状态提示、VLM前缀、动作前缀、状态专家和特征调制——并在涵盖三个任务家族的45个原子任务和20个复合任务上进行评估;然后我们将状态历史长度从1帧扫到96帧,以检查历史状态信息如何影响模型性能。实验为所有三个问题提供了系统的答案,并提炼出可测试的状态感知VLA设计原则。
cs.RO / 16 / 2608.03060
Passively Safe Convex Guidance for Cislunar Rendezvous and Proximity Operations
被动安全的凸引导用于月球间会合与接近操作
Abstract
This paper presents purely convex programs for passively safe impulsive rendezvous and proximity operations in cislunar orbits. Approach, arrival, and abort maneuvers are all designed and validated in the context of maneuver execution error and navigation uncertainty, and formulated for efficient onboard execution in the autonomous scenario. The outlined methods form the baseline onboard guidance routines for NASA's CAPSTONE 02 mission planned to demonstrate autonomous rendezvous and proximity operations capabilities in the southern 9:2 synodic near rectilinear halo orbit. High fidelity closed loop Monte Carlo simulations using the planned relative navigation sensor suite and measurement cadence verify the intended maneuver design performance.
Chinese Translation
本文提出了用于月球间轨道的被动安全冲击会合与接近操作的纯凸规划方法。接近、到达和中止机动均在机动执行误差和导航不确定性的背景下进行设计和验证,并为自主场景中的高效机载执行进行了公式化。所述方法构成了美国国家航空航天局(NASA)CAPSTONE 02任务的基础机载引导程序,该任务计划在南部9:2的合成近直线光环轨道中演示自主会合与接近操作能力。使用计划的相对导航传感器套件和测量节奏进行的高保真闭环蒙特卡洛仿真验证了预期机动设计的性能。
cs.RO / 17 / 2608.03103
A Hierarchical Approach to Imitation Learning for Manipulation Tasks Requiring Time Varying Forces
一种层次化模仿学习方法用于需要时间变化力的操作任务
Abstract
Diffusion policies have shown strong performance in learning complex, multi-modal behaviors for robotic manipulation. However, their application to contact-rich disassembly tasks remains limited by a key trade-off: the iterative denoising process introduces inference latencies that makes high frequency control difficult, which is essential for realizing dynamic interactions such as chiseling and prying. Recent action-chunking techniques mitigate latency but use an open-loop execution window, rendering the system blind to rapid force transients caused by fracture events. To bridge this gap, we introduce the Diffusion Policy Augmented by Fast Trajectory Generation (DPA-FTG). Compared to recent visual-tactile approaches that focus on positional correction, DPA-FTG decouples low-frequency planning from high-frequency force regulation. At the high level ($5$ Hz), a conditional diffusion model predicts a sequence of latent parameters for selecting a strategy from a learned vocabulary of task primitives. At the low level ($60$ Hz), a lightweight, force-conditioned policy acts as a neural impedance controller, modulating execution in real-time to maintain contact stability. We validate our approach on a bimanual battery disassembly task involving the separation of a compliant sheet. Experimental evaluation demonstrates that DPA-FTG outperforms state-of-the-art baselines, including Reactive Diffusion Policy (RDP).
Chinese Translation
扩散策略在学习复杂的多模态行为方面表现出色,尤其是在机器人操作中。然而,其在接触丰富的拆解任务中的应用受到一个关键权衡的限制:迭代去噪过程引入的推理延迟使得高频控制变得困难,而这对于实现动态交互(如凿削和撬动)至关重要。最近的动作分块技术虽然减轻了延迟,但采用开放式执行窗口,使得系统无法感知由于断裂事件引起的快速力瞬变。为了解决这一问题,我们提出了快速轨迹生成增强的扩散策略(DPA-FTG)。与近期关注位置修正的视觉-触觉方法相比,DPA-FTG将低频规划与高频力调节解耦。在高层($5$ Hz),条件扩散模型预测一系列潜在参数,以从学习到的任务原语词汇中选择策略。在低层($60$ Hz),轻量级的力条件策略作为神经阻抗控制器,实时调节执行以保持接触稳定性。我们在一个双手电池拆解任务中验证了我们的方法,该任务涉及分离一个柔性片。实验评估表明,DPA-FTG的性能优于包括反应扩散策略(RDP)在内的最先进基线。
cs.RO / 18 / 2608.03116
Shooting for Contact: Contact-Implicit Multiple Shooting for Dynamic Motion Retargeting
追求接触:基于接触隐式的动态运动重定向多重射击方法
Abstract
Motion retargeting approaches often prioritize kinematic similarity over whole-body dynamics, contact consistency, and actuation limits, yielding references that are difficult for reinforcement learning (RL) policies to reproduce, particularly for contact-rich behaviors. We present a contact-implicit, direct simulation-based multiple shooting (DSMS) framework that transforms kinematically feasible references into dynamically feasible whole-body trajectories. By embedding a differentiable simulator within a nonlinear program, DSMS resolves contact, friction, impacts, self-collision, and joint limits internally while enforcing tracking, actuation, and task constraints without prescribing a contact schedule or introducing explicit contact constraints. Compared with existing retargeting methods, DSMS accelerates motion-imitation RL training and yields policies with high success rates and low tracking error. We further demonstrate zero-shot sim-to-real transfer on the Unitree G1 through command-conditioned contact-rich crawling and a highly dynamic 180-degree jump-turn.
Chinese Translation
运动重定向方法通常优先考虑运动学相似性,而忽视整体动力学、接触一致性和驱动限制,从而产生难以被强化学习(RL)策略重现的参考,尤其是在接触丰富的行为中。我们提出了一种接触隐式的直接仿真基础多重射击(DSMS)框架,该框架将运动学上可行的参考转化为动力学上可行的整体轨迹。通过在非线性程序中嵌入可微分的仿真器,DSMS 在内部解决接触、摩擦、冲击、自碰撞和关节限制,同时在不规定接触时间表或引入显式接触约束的情况下,强制执行跟踪、驱动和任务约束。与现有的重定向方法相比,DSMS 加速了运动模仿的 RL 训练,并产生了高成功率和低跟踪误差的策略。我们进一步展示了在 Unitree G1 上的零样本仿真到现实转移,通过命令条件的接触丰富爬行和高度动态的180度跳转转向。
cs.RO / 19 / 2608.03127
DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units
DigitCode:通过解剖单元对手部动作进行符号化标记
Abstract
Hand motion carries the finest-grained information in human activity, yet the representations behind hand generation, understanding, and robot learning are overwhelmingly continuous--joint angles or MANO parameters. These are accurate but unstructured: a finger cannot be indexed or edited as a symbol, and nothing marks a pose as anatomically valid. Discrete symbolic representations supply exactly this structure, and Hand Labanotation (HL) has shown they are feasible for the hand, writing motion as a T x 40 grid of one fixed direction symbol per bone. Building on this grid, we ask the question underneath it: the anatomical unit a symbol should span--bone, finger, or whole hand. DigitCode answers it by adapting, grouping, and layering HL's alphabet along the hand's unit hierarchy within one code, cutting the symbolic representation's quantization error by three quarters. The lever is the unit, not the quantizer family: at a fixed unit, training-free and learned strong quantizers are interchangeable on reconstruction, while moving down the anatomical hierarchy is what shifts accuracy. The hierarchy also tracks what downstream tasks need. Because a finger is a genuine, enumerable unit, one per-finger token doubles as a training-free, editable handle for jobs a continuous representation cannot address--repairing malformed generated hands, and retargeting them onto robots. We release HandTok, a reproducible testbed, so hand tokenizers can be compared unit-for-unit. Project page: https://digitcode-demo.github.io.
Chinese Translation
手部动作承载着人类活动中最细粒度的信息,但在手部生成、理解和机器人学习背后的表示方式大多是连续的——关节角度或MANO参数。这些表示虽然准确,但缺乏结构:手指无法像符号那样被索引或编辑,也没有任何标记表明一个姿势在解剖上是有效的。离散的符号表示恰好提供了这种结构,而手部拉班符号法(Hand Labanotation, HL)已证明它们在手部是可行的,将动作写成每根骨骼一个固定方向符号的T x 40网格。在此网格基础上,我们提出了一个问题:符号应跨越的解剖单元——骨骼、手指或整个手。DigitCode通过适应、分组和分层HL的字母表,回答了这个问题,将符号表示的量化误差减少了四分之三。关键在于单位,而不是量化器家族:在固定单位下,无需训练的强量化器和学习得到的强量化器在重建时是可以互换的,而向下移动解剖层级则是影响准确性的因素。该层级还跟踪下游任务的需求。由于手指是一个真实的、可枚举的单位,因此每个手指的标记也可以作为一个无需训练的、可编辑的工具,用于处理连续表示无法解决的任务——修复生成的畸形手部,并将其重新定向到机器人上。我们发布了HandTok,一个可重复的测试平台,以便手部标记器可以逐单位进行比较。项目页面:https://digitcode-demo.github.io。
cs.RO / 20 / 2608.03155
POMDPs for Autonomous Science Exploration
用于自主科学探索的部分可观测马尔可夫决策过程(POMDP)
Abstract
Autonomous exploration missions require decision-making under sensor uncertainty and computational constraints, yet integrating scientific representations into POMDP planning has remained intractable due to high-dimensional observation spaces. Information-theoretic planners overcome this by assuming deterministic observations, sacrificing the principled uncertainty quantification that POMDPs provide. We introduce the Science Hypothesis Map POMDP (SHM-POMDP), which makes science-driven belief-space planning more tractable by branching on inferred physical properties rather than raw sensor data. This preserves full sensor information through learned observation models while enabling the planner to reason jointly about navigation and scientific properties under uncertainty. On an extended RockSample domain with 50-dimensional observations, SHM-POMDP achieves 18.6\% higher rewards and 32.9\% reduced computation time per step than continuous-observation baselines. On realistic geologic exploration using Cuprite hyperspectral data, SHM-POMDP achieves 2.5$\times$ higher information gain than the best information-theoretic baseline by maintaining beliefs and replanning adaptively---reaching 80\% of oracle performance using only uniform priors. These results demonstrate that integrating hierarchical probabilistic models into belief-space planning enables tractable, principled autonomous science that outperforms both traditional POMDP methods and science-aware information-theoretic approaches.
Chinese Translation
自主探索任务需要在传感器不确定性和计算限制下进行决策,然而由于高维观测空间,将科学表示整合到POMDP规划中仍然难以实现。信息论规划者通过假设确定性观测来克服这一问题,但这牺牲了POMDP所提供的原则性不确定性量化。我们提出了科学假设图POMDP(SHM-POMDP),通过基于推断的物理属性进行分支,而非原始传感器数据,使得以科学为驱动的信念空间规划变得更加可行。这种方法通过学习的观测模型保留了完整的传感器信息,同时使规划者能够在不确定性下共同推理导航和科学属性。在一个具有50维观测的扩展RockSample领域中,SHM-POMDP比连续观测基线获得了18.6%的更高奖励和32.9%的每步计算时间减少。在使用Cuprite高光谱数据进行的现实地质探索中,SHM-POMDP通过维护信念并自适应地重新规划,实现了比最佳信息论基线高出2.5倍的信息增益,仅使用均匀先验即可达到80%的预言者性能。这些结果表明,将层次概率模型整合到信念空间规划中能够实现可行的、原则性的自主科学探索,且优于传统的POMDP方法和以科学为导向的信息论方法。
cs.RO / 21 / 2608.03159
Accelerating Human-Aware Robot Trajectory Generation via Diffusion and Consistency Distillation
通过扩散和一致性蒸馏加速人机交互中的机器人轨迹生成
Abstract
This research proposes a constrained motion planning framework for robot manipulators in human-robot interaction (HRI). For a non-redundant manipulator with a fully specified end-effector pose, additional requirements such as collision avoidance and self-collision avoidance are difficult to handle as simple null-space secondary tasks. This limitation makes it challenging to generate feasible joint-space trajectories in HRI environments where safety and kinematic constraints must be considered simultaneously. To address this limitation, collision- and self-collision-aware trajectories are generated using Rapidly-exploring Random Tree (RRT) and RRT* algorithms, and the resulting dataset is used to train a diffusion model that generates constraint-satisfying trajectories through guided sampling. To reduce the inference time required for iterative diffusion sampling, consistency distillation is applied, and a joint-weighted jerk regularization term is incorporated into the loss function to promote smoother trajectories by penalizing abrupt changes in joint acceleration. Simulation results show that the consistency model generates 150 trajectory candidates in less than 100 ms, maintains a high episode success rate, and substantially reduces joint and end-effector jerk when jerk regularization is applied.
Chinese Translation
本研究提出了一种用于人机交互(HRI)中机器人操纵器的约束运动规划框架。对于具有完全指定末端执行器姿态的非冗余操纵器,诸如避碰和自碰撞避免等额外要求难以作为简单的零空间次要任务来处理。这一限制使得在必须同时考虑安全性和运动学约束的人机交互环境中生成可行的关节空间轨迹变得具有挑战性。为了解决这一限制,使用快速探索随机树(Rapidly-exploring Random Tree,RRT)和RRT*算法生成了考虑碰撞和自碰撞的轨迹,并利用生成的数据集训练了一个扩散模型,通过引导采样生成满足约束的轨迹。为了减少迭代扩散采样所需的推理时间,应用了一致性蒸馏,并在损失函数中加入了关节加速度加权的抖动正则化项,以通过惩罚关节加速度的突然变化来促进轨迹的平滑性。仿真结果表明,一致性模型在不到100毫秒的时间内生成150个轨迹候选,保持了高的成功率,并在应用抖动正则化时显著减少了关节和末端执行器的抖动。
cs.RO / 22 / 2608.03227
PFM-HR: Pose Flow Matching for Humanoid Robots
PFM-HR:用于类人机器人姿态流匹配
Abstract
Motion priors improve reinforcement learning for physics-based humanoid tracking, but temporal priors require ordered motion clips, while pose priors provide limited guidance for policy-induced pose transitions. We present Pose Flow Matching for Humanoid Robots (PFM-HR), a reusable flow matching prior trained directly on large scale unordered pose data. PFM-HR introduces the Pose Geometry Score (PGS), which quantifies how joint coordinate changes during rollouts align with the local geometry of pose variation captured by the prior. Using PGS to modulate the tracking reward guides policy exploration toward structured pose changes while keeping the prior frozen across tracking tasks. Experiments demonstrate that PFM-HR improves both single motion and general motion tracking, especially for highly dynamic motions.
Chinese Translation
运动先验改善了基于物理的类人跟踪的强化学习,但时间先验需要有序的运动片段,而姿态先验对策略诱导的姿态转变提供的指导有限。我们提出了类人机器人姿态流匹配(PFM-HR),这是一种直接在大规模无序姿态数据上训练的可重用流匹配先验。PFM-HR引入了姿态几何分数(PGS),该分数量化了在执行过程中关节坐标变化与先验捕获的姿态变化局部几何的对齐程度。使用PGS来调节跟踪奖励,引导策略探索朝向结构化的姿态变化,同时在跟踪任务中保持先验不变。实验表明,PFM-HR在单一运动和一般运动跟踪方面均有所改善,尤其是在高度动态的运动中。
cs.RO / 23 / 2608.03231
Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking
结构感知的鲁棒微调:保护视觉-语言-动作机器人免受物理注意力劫持
Abstract
Vision-Language-Action (VLA) policies promise general robotic manipulation, but their robustness against physical-world attacks remains fragile. In particular, we show that physically realizable adversarial patches can reliably induce failures by triggering a mechanism we call policy-critical action-to-vision attention hijacking, where action-conditioned attention is diverted from task-relevant regions to a localized patch. To demonstrate the threat, we propose Attention-Guided Semantic Disruption (AGSD), an Expectation-over-Transformation (EOT) optimized printable patch that jointly (i) concentrates action-to-vision attention on the patch and (ii) disrupts vision-language semantic alignment, yielding strong cross-task and cross-architecture transfer. To mitigate such attacks, we introduce Structure-Aware Robust Fine-Tuning (SARF), a zero-inference-overhead defense that fine-tunes only the visual encoder using feature anchoring, policy-critical attention correction, and language-guided geometric consistency restricted to semantically relevant regions. On LIBERO, SARF reduces OpenVLA's failure rate under AGSD from 100% to 14.2%-56.8% (28.6% average) across suites while preserving clean performance, and on a real PiPER manipulator it improves average success under AGSD from 23.0% to 65.0%. These results highlight mechanism-level robustness as a practical path to securing VLA robots against physical attention hijacking.
Chinese Translation
视觉-语言-动作(VLA)策略承诺实现通用机器人操作,但其对物理世界攻击的鲁棒性仍然脆弱。特别是,我们展示了可物理实现的对抗性补丁可以可靠地诱发失败,触发我们称之为策略关键的动作到视觉注意力劫持的机制,其中基于动作的注意力从与任务相关的区域转移到局部补丁。为了展示这一威胁,我们提出了注意力引导的语义干扰(AGSD),这是一种基于期望-变换(EOT)优化的可打印补丁,它共同(i)将动作到视觉的注意力集中在补丁上,并且(ii)干扰视觉-语言的语义对齐,从而实现强大的跨任务和跨架构迁移。为了缓解此类攻击,我们引入了结构感知的鲁棒微调(SARF),这是一种零推理开销的防御方法,仅使用特征锚定、策略关键的注意力校正和限制在语义相关区域的语言引导几何一致性来微调视觉编码器。在LIBERO上,SARF将OpenVLA在AGSD下的失败率从100%降低到14.2%-56.8%(平均28.6%),同时保持清晰的性能;在真实的PiPER操纵器上,它将AGSD下的平均成功率从23.0%提高到65.0%。这些结果突出了机制级鲁棒性作为保护VLA机器人免受物理注意力劫持的实际途径。
cs.RO / 24 / 2608.03234
Learning Context-Aware Motion Priors for Humanoid Control
学习上下文感知的运动先验用于类人控制
Abstract
Motion priors provide powerful guidance for learning naturalistic humanoid behaviors. However, existing methods typically learn a general, task-agnostic prior from the entire reference dataset and apply it uniformly throughout policy training. As a result, the prior cannot distinguish which reference motions are relevant to the current task context, potentially providing irrelevant or conflicting guidance. We present Context-Aware Motion Priors (CMP), a framework that adapts a general motion prior to the current task context without manual skill labels, dataset partitioning, or a separate skill discovery stage. Specifically, CMP learns context-motion compatibility using high-advantage policy rollouts, while a demonstration-based objective keeps the learned relevance grounded in the reference distribution. The resulting relevance scores reweight reference supervision for training a lightweight context-conditioned adapter. To evaluate the effectiveness and generality of CMP, we instantiate it with both Adversarial Motion Priors and Score-Matching Motion Priors. Across five humanoid control tasks, CMP consistently improves task performance and sample efficiency, learns meaningful context-motion alignment, and remains robust to imbalanced reference distributions. These results show that adapting motion priors to task contexts provides more relevant guidance for humanoid policy learning.
Chinese Translation
运动先验为学习自然类人的行为提供了强有力的指导。然而,现有方法通常从整个参考数据集中学习一个通用的、与任务无关的先验,并在策略训练中均匀应用。因此,该先验无法区分哪些参考动作与当前任务上下文相关,可能提供无关或相互冲突的指导。我们提出了上下文感知运动先验(Context-Aware Motion Priors, CMP),这是一个将通用运动先验适应于当前任务上下文的框架,无需手动技能标签、数据集划分或单独的技能发现阶段。具体而言,CMP通过高优势策略回放学习上下文-运动兼容性,同时基于示范的目标保持所学相关性与参考分布相一致。最终的相关性得分重新加权参考监督,以训练轻量级的上下文条件适配器。为了评估CMP的有效性和普遍性,我们将其与对抗运动先验(Adversarial Motion Priors)和评分匹配运动先验(Score-Matching Motion Priors)结合使用。在五个类人控制任务中,CMP始终提高任务性能和样本效率,学习到有意义的上下文-运动对齐,并对不平衡的参考分布保持稳健。这些结果表明,将运动先验适应于任务上下文为类人策略学习提供了更相关的指导。
cs.RO / 25 / 2608.03295
GraspMeanFlow: SE(3)-Equivariant MeanFlow for Few-Step 6-DoF Grasp Generation
GraspMeanFlow:用于少步6自由度抓取生成的SE(3)-等变均流
Abstract
Recent data-driven methods for synthesizing 6-DoF grasp poses use generative models to learn complex grasp pose distributions and generate diverse candidate poses. In particular, SE(3)-equivariant flow-based models generate grasp poses that transform consistently with object rotations and translations. However, these methods sample by iterative numerical integration, requiring tens of function evaluations per grasp and limiting their use in real-time manipulation. We propose GraspMeanFlow, an SE(3)-equivariant MeanFlow framework for few-step 6-DoF grasp generation. Our method learns the average velocity over a finite time interval, defined through the time-ordered exponential so that it reproduces exactly the rigid-body displacement accumulated over that interval. We prove that a point-cloud-conditioned distribution transported by an equivariant average-velocity flow map remains invariant, so equivariance is retained under few-step sampling, and we condition the field on a pair of times by lifting both to equivariant vectors, leaving the backbone otherwise unchanged. For stable training, we pair a flow-matching boundary term with either of two consistency terms: the differential MeanFlow identity, whose target requires a Jacobian-vector product, or an equivalent semigroup loss that avoids it. Experiments on ACRONYM show that a single function evaluation of GraspMeanFlow reaches the EMD that an iterative SE(3) flow model needs five steps to approach, that a second instantiation of the same framework improves grasp success by up to 24.3 points in the few-step regime, and that both generate grasp distributions transforming exactly with the object.
Chinese Translation
最近的数据驱动方法通过生成模型合成6自由度抓取姿态,学习复杂的抓取姿态分布并生成多样的候选姿态。特别是,SE(3)-等变的基于流的方法生成的抓取姿态在物体旋转和平移时保持一致的变换。然而,这些方法通过迭代数值积分进行采样,每个抓取需要数十次函数评估,这限制了它们在实时操作中的应用。我们提出了GraspMeanFlow,一个用于少步6自由度抓取生成的SE(3)-等变均流框架。我们的方法学习在有限时间间隔内的平均速度,通过时间有序指数定义,使其准确再现该时间间隔内积累的刚体位移。我们证明了由等变平均速度流映射传输的点云条件分布保持不变,因此在少步采样下保持等变性,并且我们通过将时间对提升到等变向量来对场进行条件处理,而保持主干不变。为了稳定训练,我们将流匹配边界项与两个一致性项之一配对:微分均流恒等式,其目标需要雅可比-向量积,或避免它的等效半群损失。在ACRONYM上的实验表明,GraspMeanFlow的单次函数评估达到了迭代SE(3)流模型需要五步才能接近的EMD,第二个相同框架的实例在少步情况下将抓取成功率提高了最多24.3个百分点,并且两者生成的抓取分布在物体变换时完全一致。
cs.RO / 26 / 2608.03296
PLS-Calib: A Partial Least Squares Framework for Event Camera and Odometry Calibration under Ground Motion Constraints
PLS-Calib:一种在地面运动约束下用于事件相机与里程计标定的偏最小二乘框架
Abstract
Accurate extrinsic rotation calibration between sensors is fundamental to the performance of robotic perception systems. However, most existing calibration techniques rely on full 6-DoF motion to excite all degrees of freedom, which is often infeasible for ground-constrained robots with limited motion capabilities. Recent approaches designed for such restricted settings, such as Canonical Correlation Analysis (CCA)-based methods, suffer from ill-conditioned covariance matrices that lead to numerical instability and suboptimal calibration accuracy. To overcome these limitations, we present a novel rotation calibration framework named PLS-Calib that, for the first time, leverages Partial Least Squares (PLS) regression to model the latent kinematic correlations between asynchronous, heterogeneous sensor streams. Specifically, we apply our method to the calibration of an event camera and an odometry onboard a ground robot. To improve event-based pattern detection, we introduce a polarity-aware event representation, which enhances spatiotemporal contrast in circular calibration targets. Our PLS-based formulation yields a closed-form, stable solution that avoids matrix singularities inherent in CCA-based approaches. Extensive experiments on both synthetic and real-world datasets validate the effectiveness of our approach, demonstrating significant improvements in calibration robustness and accuracy over state-of-the-art methods. This work offers a practical and theoretically grounded solution for rotation calibration in constrained robotic systems and opens up new directions for applying statistical learning techniques in neuromorphic vision.
Chinese Translation
传感器之间的准确外部旋转标定对机器人感知系统的性能至关重要。然而,大多数现有的标定技术依赖于完整的6自由度运动来激发所有自由度,这对于运动能力有限的地面约束机器人来说往往不可行。针对这种受限环境设计的近期方法,例如基于典型相关分析(CCA)的方法,存在条件不良的协方差矩阵,导致数值不稳定和次优的标定精度。为克服这些限制,我们提出了一种新颖的旋转标定框架,命名为PLS-Calib,它首次利用偏最小二乘(PLS)回归来建模异步异质传感器流之间的潜在运动学相关性。具体而言,我们将该方法应用于地面机器人上的事件相机与里程计的标定。为了提高基于事件的模式检测,我们引入了一种考虑极性的事件表示,增强了圆形标定目标的时空对比度。我们的基于PLS的公式提供了一种封闭形式的稳定解,避免了基于CCA的方法中固有的矩阵奇异性。在合成和真实数据集上的大量实验验证了我们方法的有效性,显示出在标定鲁棒性和精度方面相较于最先进方法的显著提升。这项工作为受限机器人系统中的旋转标定提供了一个实用且理论上扎实的解决方案,并为在神经形态视觉中应用统计学习技术开辟了新的方向。
cs.RO / 27 / 2608.03378
Shaping Wind-Tunnel Airflow for Unmanned Aerial Vehicles using Online Learning
利用在线学习塑造无人机的风洞气流
Abstract
The development and testing of advanced aerial robots require experiments in controlled environments with tailored airflow profiles. This paper presents an online learning algorithm for controlling the complex airflow field in a multi-fan vertical wind tunnel. Our method combines a simplified physical model with iterative, measurement-based learning, enabling sample-efficient convergence to desired airflow distributions. We demonstrate the method's versatility by generating complex airflow, such as uniform, Gaussian, and parabolic profiles. Crucially, we show that our algorithm can produce an airflow profile specifically designed for passive soaring, greatly enhancing flight performance of a soaring robot. Variability, practical utility, and robustness of our approach are further highlighted by successful operation with a varying number of fans.
Chinese Translation
先进空中机器人的开发与测试需要在受控环境中进行实验,并制定特定的气流特征。本文提出了一种在线学习算法,用于控制多风扇垂直风洞中的复杂气流场。我们的方法结合了简化的物理模型与基于测量的迭代学习,使得在样本效率上快速收敛到期望的气流分布。我们通过生成复杂的气流,如均匀、Gaussian(高斯)和抛物线特征,展示了该方法的多样性。重要的是,我们表明我们的算法能够产生专门为被动滑翔设计的气流特征,极大地提升了滑翔机器人的飞行性能。我们的方法的可变性、实用性和鲁棒性还通过在不同风扇数量下的成功操作得到了进一步的验证。
cs.RO / 28 / 2608.03387
RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation
RoboReact:从生成的自我中心视频中提取代理技能以实现可泛化的全身操控
Abstract
Humanoid robots have the potential to perform dexterous manipulation in human environments, yet acquiring diverse and generalizable skills remains costly due to expensive hardware data collection and labor-intensive annotation. Recent advances in video generative models provide a promising opportunity to synthesize rich manipulation experiences from visual observations, but transferring such imagined behaviors into executable whole-body humanoid skills remains largely unexplored. In this work, we present RoboReact, a framework that automatically synthesizes whole-body humanoid manipulation skills from a single egocentric RGB-D observation. RoboReact generates human manipulation videos, extracts geometry-preserving interaction keyframes through depth-aware 3D reconstruction, and retargets them to high-DoF humanoid platforms while preserving hand-object interaction geometry. To bridge the gap between imagined plans and physical execution, RoboReact performs online object-centric re-grounding and leverages a vision-language model-guided refinement loop to adapt skills under geometric mismatch and execution deviations. The refined skills are executed through a whole-body controller, enabling coordinated whole-body manipulation and dexterous interaction. Experiments on real humanoid robots demonstrate that RoboReact generalizes across diverse object configurations and robustly recovers from execution disturbances without requiring teleoperation or human demonstrations. These results highlight the potential of combining generative models, vision-language reasoning, and closed-loop control for scalable humanoid skill acquisition.
Chinese Translation
类人机器人在人工环境中具有执行灵巧操控的潜力,但由于昂贵的硬件数据收集和劳动密集型标注,获取多样化和可泛化的技能仍然成本高昂。最近在视频生成模型方面的进展为从视觉观察中合成丰富的操控经验提供了有希望的机会,但将这些想象的行为转化为可执行的全身类人技能仍然在很大程度上未被探索。在本研究中,我们提出了RoboReact,一个从单个自我中心RGB-D观察中自动合成全身类人操控技能的框架。RoboReact生成人类操控视频,通过深度感知的3D重建提取几何保留的交互关键帧,并在保留手-物体交互几何的同时将其重新定向到高自由度的类人平台。为了弥合想象计划与物理执行之间的差距,RoboReact执行在线物体中心的重新定位,并利用视觉-语言模型引导的精炼循环来适应几何不匹配和执行偏差。精炼后的技能通过全身控制器执行,实现协调的全身操控和灵巧的交互。在真实类人机器人上的实验表明,RoboReact能够在多样的物体配置中泛化,并在不需要远程操作或人类示范的情况下稳健地从执行干扰中恢复。这些结果突显了结合生成模型、视觉-语言推理和闭环控制在可扩展类人技能获取中的潜力。
cs.RO / 29 / 2608.03408
Flying over The Uncertain Nature (FORTUNE): Intelligent and Humanistic 3D Path Planning for Low-Altitude Collaboration
飞越不确定性(FORTUNE):低空协作的智能人文三维路径规划
Abstract
The proliferation of low-altitude intelligent agents is increasing the demand for timely and socially responsible collaborative sensing in dynamic urban environments. However, jointly addressing heterogeneous spatiotemporal demands, environmental uncertainty, and human-centered operational constraints remains challenging. This paper studies 3D multi-UAV path planning and task assignment under uncertain ground PoI demands. Unlike existing work assuming static and fully known PoIs, we model persistent, temporally predictable, and emergent demands within a unified framework. We further incorporate altitude-dependent societal and environmental costs, including noise exposure and public safety risks, to balance sensing performance with socially compliant operations. To solve the resulting large-scale mixed-integer nonlinear problem, we propose FORTUNE, a hierarchical offline-online framework. Offline, a Transformer predicts Type-II PoI activation windows, while an enhanced sparrow search algorithm generates coordinated flight plans through priority-aware decoding and danger-aware evolution. Online, a lightweight refinement module accommodates emerging Type-III PoIs while preserving global mission coherence. Experiments on real-world traffic data and synthetic scenarios show that FORTUNE consistently outperforms state-of-the-art methods in effectiveness, scalability, and practical applicability.
Chinese Translation
低空智能代理的普及增加了在动态城市环境中及时和社会责任感强的协同感知的需求。然而,联合解决异质时空需求、环境不确定性和以人为本的操作约束仍然具有挑战性。本文研究了在不确定地面兴趣点(PoI)需求下的三维多无人机路径规划和任务分配。与现有研究假设静态且完全已知的兴趣点不同,我们在统一框架内对持久、时间可预测和突发的需求进行了建模。我们进一步结合了高度依赖的社会和环境成本,包括噪音暴露和公共安全风险,以平衡感知性能与社会合规操作。为了解决由此产生的大规模混合整数非线性问题,我们提出了FORTUNE,一个分层的离线-在线框架。在离线阶段,Transformer预测类型II兴趣点的激活窗口,而增强的麻雀搜索算法通过优先级感知解码和危险感知演化生成协调的飞行计划。在在线阶段,一个轻量级的精细化模块在保持全球任务一致性的同时,适应新出现的类型III兴趣点。基于真实交通数据和合成场景的实验表明,FORTUNE在有效性、可扩展性和实际适用性方面始终优于最先进的方法。
cs.RO / 30 / 2608.03444
A Low-Cost Hybrid Reservoir Computing Model for Isolated Sign Language Video Recognition
一种低成本混合水库计算模型用于孤立手语视频识别
Abstract
Sign language recognition (SLR) enhances communication between hearing and hearing-impaired individuals. Although deep learning (DL) has achieved promising performance in SLR, its high computational cost limits deployment on edge devices. To address this challenge, we propose a lightweight reservoir computing (RC)-based approach for SLR. In the proposed method, MediaPipe extracts body and hand keypoints to capture the spatial and temporal dynamics of gestures. These keypoints are then processed by a hybrid reservoir computing (HRC) architecture that combines deep reservoir computing (DRC) and bidirectional reservoir computing (BRC), transforming the input into a high-dimensional dynamic representation. A ridge regression model maps the final HRC state to class labels. This HRC-based SLR method achieved Top-1, Top-5, and Top-10 accuracies of 61.12%, 86.05%, and 92.56%, respectively, on the Word-Level American Sign Language 100 (WLASL100) video dataset, demonstrating competitive performance compared to deep learning-based approaches. Additionally, due to the lightweight nature of RC, the training time was drastically reduced to only a few seconds compared with DL-based methods such as Bi-GRU.This method offers low computational cost, showing its potential for deployment on edge devices.
Chinese Translation
手语识别(SLR)增强了听力正常者与听力障碍者之间的沟通。尽管深度学习(DL)在手语识别中取得了令人鼓舞的表现,但其高计算成本限制了在边缘设备上的部署。为了解决这一挑战,我们提出了一种基于轻量级水库计算(RC)的手语识别方法。在该方法中,MediaPipe 提取身体和手部关键点,以捕捉手势的空间和时间动态。这些关键点随后由一种混合水库计算(HRC)架构处理,该架构结合了深度水库计算(DRC)和双向水库计算(BRC),将输入转换为高维动态表示。岭回归模型将最终的 HRC 状态映射到类别标签。该基于 HRC 的手语识别方法在 Word-Level American Sign Language 100(WLASL100)视频数据集上分别达到了 61.12%、86.05% 和 92.56% 的 Top-1、Top-5 和 Top-10 准确率,表现出与基于深度学习的方法相媲美的竞争力。此外,由于 RC 的轻量特性,训练时间相比于基于 DL 的方法(如 Bi-GRU)大幅缩短至仅几秒钟。该方法具有低计算成本,显示出在边缘设备上部署的潜力。
cs.RO / 31 / 2608.03483
Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution
继续还是重新规划?用于自适应时间执行的伯努利继续策略学习
Abstract
Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.e., execution horizon) before replanning, turning replanning into a task-agnostic periodic schedule that is independent of task progress. As a result, when no replanning boundary falls before a critical manipulation stage, it is executed from a stale chunk rather than a freshly replanned one. To address this limitation, we propose Bernoulli-Continuation Policy (BCP), a lightweight, plug-and-play framework for adaptive horizon execution that keeps the base VLA frozen. Given a fixed-length action chunk, its continuation head decomposes execution-horizon selection into a sequence of continue-or-replan decisions, which imposes an ordinal, prefix-sharing inductive bias over candidate horizons rather than treating them as independent classes. Since the optimal horizon for each chunk is not observable, we train this head with reinforcement learning from trajectory-level outcomes and introduce a Replanning-Efficiency Reward that jointly rewards task success and efficient VLA usage, discouraging the policy from collapsing to unnecessarily short horizons. On RoboTwin 2.0 with LingBot-VLA as the base policy, BCP improves the average success rate by +11.08% on 13 low-success tasks and from 89.88% to 93.94% (+4.06%) across all 50 tasks. Although trained only under the Clean setting, BCP generalizes to the Randomized setting, raising the average success rate by +4.06%. It also transfers to a different base policy $\pi_{0.5}$, achieving a better result on LIBERO (+1.7%) and, notably, on the harder LIBERO-PRO (+6.8%). On a real robot, BCP lifts success from 74% to 92% and from 44% to 84% on two manipulation tasks. Meanwhile, its negligible overhead, combined with higher success, makes BCP's overall runtime even lower than the fixed-horizon baselines.
Chinese Translation
现有的基于块的视觉-语言-动作(VLA)模型在重新规划之前执行固定数量的动作(即执行时间),将重新规划变成与任务进展无关的任务无关的周期性调度。因此,当没有重新规划边界落在关键操作阶段之前时,它是从过时的块中执行,而不是从新规划的块中执行。为了解决这一限制,我们提出了伯努利继续策略(BCP),这是一种轻量级、即插即用的自适应时间执行框架,保持基础VLA不变。给定一个固定长度的动作块,其继续头将执行时间选择分解为一系列继续或重新规划的决策,这对候选时间施加了序数的、前缀共享的归纳偏差,而不是将其视为独立类别。由于每个块的最佳时间不可观察,我们通过轨迹级结果的强化学习训练这个头,并引入了重新规划效率奖励,该奖励共同奖励任务成功和高效的VLA使用,抑制策略陷入不必要的短时间。基于LingBot-VLA作为基础策略的RoboTwin 2.0上,BCP在13个低成功任务上将平均成功率提高了11.08%,在所有50个任务上从89.88%提高到93.94%(增加4.06%)。尽管仅在Clean设置下训练,BCP在Randomized设置中也具有良好的泛化能力,平均成功率提高了4.06%。它还转移到不同的基础策略$ ext{π}_{0.5}$,在LIBERO上取得了更好的结果(+1.7%),并且在更困难的LIBERO-PRO上显著提高了6.8%。在真实机器人上,BCP将两个操作任务的成功率从74%提升到92%,从44%提升到84%。同时,其微不足道的开销,加上更高的成功率,使得BCP的整体运行时间甚至低于固定时间基线。
cs.RO / 32 / 2608.03490
Lightweight 3D Object Detection via Mamba-Based Knowledge Distillation
基于Mamba的知识蒸馏轻量级3D物体检测
Abstract
3D object detection using light detection and ranging (LiDAR) sensors requires a balance between accuracy and computational efficiency for onboard perception in autonomous driving and robotic navigation. Many existing LiDAR-based detection methods employ complex architectures to extract features, integrating large amounts of contextual information to enhance accuracy. This often results in significant computational costs, leading to suboptimal performance on resource-constrained embedded devices. In this study, we propose a knowledge distillation framework that transfers object-level voxel representations from a strong teacher model to lightweight student models through selective voxel-space feature alignment. Taking advantage of the linear-time sequence model with selective state spaces (Mamba), we design a multi-branch Mamba teacher backbone and a box-aware feature transfer mechanism that aligns spatially corresponding voxel features between teacher and student networks through a Mamba-based projection module. Experimental results on both a public dataset and real-world data show that our approach significantly reduces computational load while maintaining competitive accuracy compared with state-of-the-art methods.
Chinese Translation
使用激光雷达(LiDAR)传感器进行3D物体检测需要在自主驾驶和机器人导航的车载感知中实现准确性与计算效率之间的平衡。许多现有的基于LiDAR的检测方法采用复杂的架构来提取特征,整合大量上下文信息以提高准确性。这通常导致显著的计算成本,从而在资源受限的嵌入式设备上表现不佳。在本研究中,我们提出了一种知识蒸馏框架,通过选择性体素空间特征对齐,将对象级体素表示从强教师模型转移到轻量级学生模型。利用具有选择性状态空间的线性时间序列模型(Mamba),我们设计了一个多分支Mamba教师主干和一个盒子感知特征转移机制,通过基于Mamba的投影模块对教师和学生网络之间空间对应的体素特征进行对齐。我们在公共数据集和真实世界数据上的实验结果表明,我们的方法在保持与最先进方法竞争的准确性的同时,显著降低了计算负担。
cs.RO / 33 / 2608.03496
Principles of Robot Autonomy
机器人自主性的原则
Abstract
Autonomous robots are moving rapidly from research labs into everyday life - on roads, in the air, in warehouses, and in space. Robot autonomy is no longer solely an academic pursuit, but a collection of mature, field-tested methods and tools that practitioners rely on in real-world deployments. This book offers a clear, unified introduction to the methods that make this possible. Built on decades of teaching at Stanford, the text develops the core elements of modern autonomy stacks within a single conceptual framework, bridging classical robotics and modern physical AI. Every major topic is paired with hands-on Jupyter notebooks and implementation-driven exercises, so readers build practical intuition alongside theoretical understanding. The result is a principled, accessible, and deployment-aware foundation for anyone seeking to design, analyze, or contribute to the next generation of autonomous systems. This is a comprehensive resource for students, engineers, and researchers entering one of today's fastest-growing fields.
Chinese Translation
自主机器人正迅速从研究实验室走入日常生活——在道路上、空中、仓库中以及太空中。机器人自主性不再仅仅是学术追求,而是一系列成熟的、经过实地验证的方法和工具,实践者在现实部署中依赖于这些方法和工具。本书提供了一个清晰、统一的介绍,阐述了使这一切成为可能的方法。基于在斯坦福大学数十年的教学经验,文本在一个单一的概念框架内发展了现代自主堆栈的核心要素,桥接了经典机器人技术与现代物理人工智能。每个主要主题都配有动手实践的 Jupyter 笔记本和以实现为驱动的练习,使读者在理论理解的同时建立实际直觉。最终,这为任何希望设计、分析或为下一代自主系统做出贡献的人提供了一个有原则、易于理解且关注部署的基础。这是一个全面的资源,适合进入当今增长最快领域之一的学生、工程师和研究人员。
cs.RO / 34 / 2608.03521
Pivot-Centric Trajectory Prediction: Bridging Long Horizons via Dynamical Guidance
以枢轴为中心的轨迹预测:通过动态引导连接长时间预测
Abstract
Forecasting precise future motion of surrounding agents is essential for reliable autonomous vehicles. However, as the demand for longer prediction horizons increases, existing endpoint-completion or iterative-refine methods increasingly struggle with weak guidance and compounding errors. To tackle the long-horizon prediction challenge, we propose Pivot-Centric Trajectory Prediction (PCTP). By introducing ``pivots'' and focusing on predicting pivot points along extended trajectories, we divide the long-term prediction task into short-term sub-tasks at various scales. Specifically, PCTP decouples the long-term trajectory predicting process into two processes: pivot prediction and pivot-based trajectory refinement. The pivot prediction process aims to utilize global map context and agent-to-agent interactions to identify these ``pivot points'', while the pivot-based trajectory refinement process focuses on local map details and refines the short-term trajectory based on predicted ``pivot points''. Compared with existing methods, PCTP provides more intermediate guidance while reducing compounding errors. Moreover, PCTP is a flexible approach that can be integrated into most state-of-the-art trajectory prediction models. Experimental results show that PCTP improves the prediction accuracy of leading models on both Argoverse I and Argoverse II datasets with minimal impact on model size. Specifically, PCTP combined with QCNet outperforms all published ensemble-free methods on the Argoverse II leaderboard at submission.
Chinese Translation
准确预测周围代理的未来运动对于可靠的自动驾驶车辆至关重要。然而,随着对更长预测时间范围的需求增加,现有的端点完成或迭代优化方法在弱引导和累积误差方面越来越难以应对。为了解决长时间预测的挑战,我们提出了以枢轴为中心的轨迹预测(Pivot-Centric Trajectory Prediction, PCTP)。通过引入“枢轴”并专注于预测沿扩展轨迹的枢轴点,我们将长期预测任务分解为不同尺度的短期子任务。具体而言,PCTP将长期轨迹预测过程解耦为两个过程:枢轴预测和基于枢轴的轨迹细化。枢轴预测过程旨在利用全局地图上下文和代理间的互动来识别这些“枢轴点”,而基于枢轴的轨迹细化过程则关注局部地图细节,并根据预测的“枢轴点”细化短期轨迹。与现有方法相比,PCTP提供了更多的中间引导,同时减少了累积误差。此外,PCTP是一种灵活的方法,可以集成到大多数最先进的轨迹预测模型中。实验结果表明,PCTP在Argoverse I和Argoverse II数据集上提高了领先模型的预测准确性,对模型大小的影响最小。具体而言,PCTP与QCNet结合在Argoverse II排行榜上超越了所有已发布的无集成方法。
cs.RO / 35 / 2608.03528
Tired Actor: Fatigue-Informed Character Control
疲惫角色:基于疲劳的角色控制
Abstract
Replicating human behavior with physics simulation has been a long-expected goal in character animation. Existing efforts have achieved impressive performance in imitating a wide span of general motions. However, most existing efforts could still suffer from unnatural movements due to the lack of biomechanical and physiological priors. Given this, we project our sights to advances in behavioral energetics, which demonstrate how energy use shapes human movements. In contrast, current character controllers typically assume the character is equipped with infinite energy over time. Inspired by these, we propose to adopt fatigue as a proxy of the finite energy limit, inject it into general character animation, and thoroughly investigate how fatigue introduces new characteristics to physics-based character control. Leveraging the Three-Compartment Controller (3CC) model, we managed to obtain a policy for general motion imitation under different fatigue statuses. Furthermore, extensive analyses are conducted to demonstrate how fatigue could influence the naturalness, scalability, and robustness of character animation. Our code will be made public.
Chinese Translation
利用物理模拟复制人类行为一直是角色动画中的一个长期期待的目标。现有的研究在模仿广泛的通用动作方面取得了令人印象深刻的成果。然而,由于缺乏生物力学和生理学的先验知识,大多数现有研究仍可能遭遇不自然的运动。鉴于此,我们将目光投向行为能量学的进展,展示了能量使用如何塑造人类运动。相比之下,当前的角色控制器通常假设角色在时间上具备无限的能量。受此启发,我们提出将疲劳作为有限能量限制的代理,将其注入到通用角色动画中,并深入研究疲劳如何为基于物理的角色控制引入新的特征。利用三室控制器(Three-Compartment Controller, 3CC)模型,我们成功获得了在不同疲劳状态下进行通用动作模仿的策略。此外,我们进行了广泛的分析,以展示疲劳如何影响角色动画的自然性、可扩展性和鲁棒性。我们的代码将公开发布。
cs.RO / 36 / 2608.03556
Human Centric Embodied Intelligence for Soft Wearable Robotics
以人为中心的具身智能在软性可穿戴机器人中的应用
Abstract
Soft wearable robots have evolved rapidly from proof-of-concept devices into promising platforms for rehabilitation, occupational assistance, and human augmentation. As the field matures, its central challenge extends beyond the development of softer materials and more capable actuators to the integration of sensing, intelligence, and human adaptation into systems that users can wear comfortably, trust, and benefit from over extended periods. This transition motivates the concept of Human-Centric Embodied Intelligence (HCEI), in which intelligence emerges from the coupled human-robot system through the interaction of morphology, multimodal sensing, adaptive cognition, compliant actuation, and the wearer's own physiological and behavioral adaptation. To organize this perspective, this review introduces the Perception-Cognition-Actuation-Augmentation (PCAA) framework, which positions perception and cognition as the primary drivers of design, shifting development beyond the conventional actuator-first paradigm. Using this framework, the review synthesizes advances in soft materials, wearable sensing, artificial intelligence, actuation, human-robot interaction, digital twins, clinical translation, manufacturing, regulation, and ethics, highlighting how these interdependent components collectively shape long-term personalization and real-world deployment. By providing a unified conceptual framework and design perspective, this review aims to guide future research, foster interdisciplinary collaboration, and accelerate the translation of next-generation soft wearable robots toward personalized, predictive, and human-centric wearable intelligence.
Chinese Translation
软性可穿戴机器人已迅速从概念验证设备发展为有前景的康复、职业辅助和人类增强平台。随着该领域的成熟,其核心挑战已超越了更柔软材料和更强大执行器的开发,转向将感知、智能和人类适应性整合到用户能够舒适佩戴、信任并在长时间内受益的系统中。这一转变促生了以人为中心的具身智能(Human-Centric Embodied Intelligence, HCEI)的概念,其中智能通过形态、跨模态感知、自适应认知、顺应性驱动和佩戴者自身的生理与行为适应的相互作用,从耦合的人机系统中涌现出来。为了组织这一视角,本综述引入了感知-认知-驱动-增强(Perception-Cognition-Actuation-Augmentation, PCAA)框架,将感知和认知定位为设计的主要驱动因素,推动开发超越传统的以执行器为中心的范式。利用该框架,综述综合了软材料、可穿戴感知、人工智能、驱动、人机交互、数字双胞胎、临床转化、制造、监管和伦理等方面的进展,强调这些相互依赖的组件如何共同塑造长期个性化和现实世界的部署。通过提供统一的概念框架和设计视角,本综述旨在指导未来研究,促进跨学科合作,加速下一代软性可穿戴机器人向个性化、预测性和以人为中心的可穿戴智能的转化。
cs.RO / 37 / 2608.03563
Unified Visuomotor Targets: Supervising VLAs Beyond Physical Actions
统一视觉运动目标:超越物理动作的VLA监督
Abstract
VLA models are trained to predict robot actions from visual and language observations. This is a natural choice, but it creates a mismatch: VLMs encode rich, high-level representations of scenes and goals, while robot actions are low-level signals with limited task structure. We ask whether changing what the policy is trained to predict, rather than how it is architecturally designed, can yield better and more efficiently trained policies. We propose UVT (Unified Visuomotor Target), a unified latent prediction target that jointly encodes motor control and visual scene transition information, requiring no architectural changes and no additional data. Applied to two representative VLA systems across simulation benchmarks and real bimanual manipulation tasks, UVT improves training efficiency, final task performance, and policy robustness, with particularly strong gains under limited training budgets and challenging environmental conditions. Rollout videos and additional qualitative results are available at our project webpage: https://unified-visuomotor-targets.github.io/
Chinese Translation
VLA模型被训练用于从视觉和语言观察中预测机器人动作。这是一个自然的选择,但它造成了不匹配:VLM编码了丰富的、高层次的场景和目标表示,而机器人动作则是具有有限任务结构的低层次信号。我们探讨了改变政策训练预测的内容,而不是其架构设计,是否能够产生更好且更高效训练的政策。我们提出了UVT(统一视觉运动目标),这是一种统一的潜在预测目标,联合编码运动控制和视觉场景转变信息,无需架构更改和额外数据。将其应用于两个代表性的VLA系统,在模拟基准和实际双手操作任务中,UVT提高了训练效率、最终任务表现和政策鲁棒性,尤其在有限训练预算和具有挑战性的环境条件下表现出显著的提升。相关的回放视频和其他定性结果可在我们的项目网页上找到:https://unified-visuomotor-targets.github.io/
cs.RO / 38 / 2608.03677
Active Stiffness Control of a Supportive Continuum Robot
支持性连续机器人主动刚度控制
Abstract
Supportive continuum robots (SCRs) enhance the load-bearing capability of an operative continuum robot by mechanically coupling it with a supportive arm. However, their passive stiffness is determined by the mechanical configuration and cannot be adjusted online for varying payloads or interaction forces. Active stiffness control is therefore needed to regulate the load response and maintain positioning accuracy. Meanwhile, the closed-chain structure introduces kinematic constraints that complicate task-space regulation and stiffness control. This paper presents an active task-space stiffness control framework for a tendon-driven SCR. An existing geometric variable strain model describes the closed-chain dynamics, which are projected onto the constraint-consistent motion subspace. A projected sliding mode controller regulates the operative arm tip while preserving the constraints, and closed-loop stability is established through Lyapunov analysis. After position regulation, active apparent stiffness is introduced through a virtual Cartesian spring based on position-error feedback to shape the force--displacement response. The framework is evaluated in simulation and experimentally validated under prescribed external loads and different desired configurations. Results show that increasing the commanded stiffness gain reduces load-induced tip deflection and increases apparent directional stiffness, thereby improving load resistance and positioning robustness under external loading.
Chinese Translation
支持性连续机器人(SCRs)通过与支持臂的机械耦合增强了操作性连续机器人的承载能力。然而,它们的被动刚度由机械配置决定,无法针对不同的载荷或交互力进行在线调整。因此,需要主动刚度控制来调节负载响应并保持定位精度。同时,闭链结构引入了运动学约束,复杂化了任务空间的调节和刚度控制。本文提出了一种针对腱驱动SCR的主动任务空间刚度控制框架。现有的几何变量应变模型描述了闭链动力学,并将其投影到约束一致的运动子空间。投影滑模控制器在保持约束的同时调节操作臂尖端,并通过李雅普诺夫分析建立闭环稳定性。在位置调节后,通过基于位置误差反馈的虚拟笛卡尔弹簧引入主动表观刚度,以塑造力-位移响应。该框架在仿真中进行了评估,并在规定的外部载荷和不同期望配置下进行了实验验证。结果表明,增加指令刚度增益可以减少载荷引起的尖端偏转,并增加表观方向刚度,从而提高在外部载荷下的负载抗力和定位鲁棒性。
cs.RO / 39 / 2608.03701
LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation
LiLa-WAM:轻量级潜在推理世界-动作模型用于机器人操控
Abstract
World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticipate how a scene will evolve. However, existing WAMs often incur substantial computational overhead. Pixel-space methods often allocate substantial capacity to visual details that may not be directly relevant to control, while some latent-space methods require multi-stage training to construct the reasoning space. The resulting training cost can make such methods difficult to train under modest computational budgets. In this work, we propose LiLa-WAM, a lightweight world-action model that reasons about the future in a compact latent space and can be trained end-to-end on a single 24GB GPU. Its core design is a compact latent reasoning space jointly shaped by future-state prediction and action generation, which keeps the model lightweight while remaining well aligned with control. For task specification, we further propose the Visual Transition Token(VTT), a language-free task representation that encodes each task as a direction in visual feature space. Experiments on RoboTwin~2.0, LIBERO, and real-robot tasks demonstrate LiLa-WAM's effectiveness, achieving 90.48\% success across 50 RoboTwin tasks with single-GPU training.
Chinese Translation
世界-动作建模已成为机器人控制的一个有前景的范式,因为它使模型能够超越对观察的反应,预测场景将如何演变。然而,现有的世界-动作模型(WAM)通常会产生相当大的计算开销。像素空间方法往往将大量计算资源分配给可能与控制无直接相关的视觉细节,而一些潜在空间方法则需要多阶段训练来构建推理空间。由此产生的训练成本使得在适度计算预算下训练这些方法变得困难。在本研究中,我们提出了LiLa-WAM,一种轻量级的世界-动作模型,它在紧凑的潜在空间中进行未来推理,并可以在单个24GB GPU上进行端到端训练。其核心设计是一个由未来状态预测和动作生成共同构成的紧凑潜在推理空间,这使得模型保持轻量,同时与控制保持良好对齐。为了任务规范,我们进一步提出了视觉过渡标记(Visual Transition Token, VTT),这是一种无语言的任务表示,将每个任务编码为视觉特征空间中的一个方向。在RoboTwin 2.0、LIBERO和真实机器人任务上的实验表明,LiLa-WAM的有效性,在50个RoboTwin任务中实现了90.48%的成功率,且仅需单GPU训练。
cs.RO / 40 / 2608.03727
Track4Action: Distilling World-Centric 3D Tracker into Vision-Language-Action Policies
Track4Action:将以世界为中心的3D跟踪器提炼为视觉-语言-动作策略
Abstract
Action labels tell a vision-language-action (VLA) policy which robot commands to imitate, but not how those commands change the 3D world. The aligned demonstration clip contains this missing supervision because its $K$ frame transitions record the geometry, motion, visibility, and camera change produced during the corresponding $K$ actions. We introduce Track4Action, a framework that distills this realized transition from a frozen world-centric 3D tracker into a current-observation VLA policy. During training, Track4World encodes the clip $V_{t:t+K}$ into a pooled tracker feature. Learnable track queries infer this feature from current VLA hidden states, match it in a shared space, and condition a flow-matching action head through a feature-wise gate. The tracker feature only defines the alignment target, so neither the clip nor the tracker is used at deployment. Track4Action reaches 82.3% on zero-shot LIBERO-Plus, improving the alignment-free variant by 7.6 points and LaMP by 3.0 points. It obtains 80.44% and 81.48% on the clean and randomized RoboTwin 2.0 splits, and 67.5% average success across four physical bimanual tasks, 25.0 points above the alignment-free variant. The gains across simulation and physical tasks support action-aligned 3D tracker features as privileged supervision for tracker-free VLA deployment. Our project page is available at https://wing0night.github.io/track4action-project-page.
Chinese Translation
动作标签告诉视觉-语言-动作(VLA)策略应模仿哪些机器人指令,但并未说明这些指令如何改变3D世界。对齐的演示片段包含了这一缺失的监督,因为其$K$帧过渡记录了在相应$K$个动作中产生的几何形状、运动、可见性和摄像机变化。我们提出了Track4Action,一个将这一实现的过渡从冻结的以世界为中心的3D跟踪器提炼为当前观察的VLA策略的框架。在训练过程中,Track4World将片段$V_{t:t+K}$编码为一个汇聚的跟踪器特征。可学习的跟踪查询从当前的VLA隐藏状态中推断出该特征,在共享空间中进行匹配,并通过特征级门控条件化流匹配动作头。跟踪器特征仅定义对齐目标,因此在部署时既不使用片段也不使用跟踪器。Track4Action在零-shot LIBERO-Plus上达到了82.3%,比无对齐变体提高了7.6分,比LaMP提高了3.0分。在干净和随机的RoboTwin 2.0分割上分别获得80.44%和81.48%,在四个物理双手任务中平均成功率为67.5%,比无对齐变体高出25.0分。模拟和物理任务中的增益支持以动作对齐的3D跟踪器特征作为无跟踪器VLA部署的特权监督。我们的项目页面可在https://wing0night.github.io/track4action-project-page访问。
cs.RO / 41 / 2608.03753
GORDON: Graph-based Object-centric Rewards for Decomposition of Long-Horizon Manipulation
GORDON:基于图的面向对象奖励框架用于长时间范围操控的分解
Abstract
Learning long-horizon manipulation skills with reinforcement learning remains challenging due to the complexity of reward design, the limited guidance of sparse rewards, and the high cost of manual subtask annotation. Visual demonstrations can provide supervision for reward learning, but rewards learned from raw pixels can be brittle and sensitive to visual variation, background appearance, and robot motion. In this work, we propose GORDON, a graph-based object-centric reward learning framework that learns dense rewards from action-free video demonstrations. Each visual scene is represented as a graph of detected objects and spatial relations, and a graph neural network is trained in a self-supervised manner to embed these graphs into a task-aligned latent space. To align the representation with semantic task progress, we introduce an activity-aware weighted pooling mechanism that emphasizes task-relevant objects while masking robot-dominated motion. The dense reward is then computed as distances in the learned latent space of the current state to demonstrated goal configurations, providing a measure of task progress. In long-horizon tasks, the temporal profile of this reward reveals stage-wise object-state transitions, enabling automatic subtask discovery without manual segmentation. The discovered segments are then used to train subtask-specific rewards and specialized policies that are composed sequentially. Experiments on seven manipulation tasks on MAGICAL and ManiSkill3 benchmarks show that our object-centric reward improves reinforcement learning in short-horizon settings and enables successful policy learning in complex long-horizon tasks through automatic decomposition, achieving an average success rate of 74.4% across the long-horizon tasks (on average approximately +35 p.p. vs. best learned baseline and approximately +25 p.p. vs. oracle).
Chinese Translation
使用强化学习学习长时间范围的操控技能仍然具有挑战性,原因在于奖励设计的复杂性、稀疏奖励的有限指导以及手动子任务标注的高成本。视觉演示可以为奖励学习提供监督,但从原始像素中学习的奖励可能脆弱且对视觉变化、背景外观和机器人运动敏感。在本研究中,我们提出了GORDON,一个基于图的面向对象奖励学习框架,该框架从无动作的视频演示中学习密集奖励。每个视觉场景被表示为一个检测到的对象及其空间关系的图,并且一个图神经网络以自监督的方式进行训练,将这些图嵌入到与任务对齐的潜在空间中。为了使表示与语义任务进展对齐,我们引入了一种活动感知加权池化机制,该机制强调与任务相关的对象,同时屏蔽机器人主导的运动。然后,密集奖励被计算为当前状态在学习的潜在空间中与演示目标配置之间的距离,提供任务进展的度量。在长时间范围的任务中,该奖励的时间轮廓揭示了阶段性对象状态转换,使得在没有手动分割的情况下自动发现子任务成为可能。发现的片段随后用于训练特定于子任务的奖励和顺序组合的专门策略。在MAGICAL和ManiSkill3基准上的七个操控任务实验表明,我们的面向对象奖励在短时间范围设置中改善了强化学习,并通过自动分解在复杂的长时间范围任务中实现了成功的策略学习,在长时间范围任务中实现了74.4%的平均成功率(与最佳学习基线相比平均提高约35个百分点,与oracle相比平均提高约25个百分点)。
cs.RO / 42 / 2608.03816
Design and Evaluation of an AI-Enabled Cloud-Edge Architecture for Connected Precision Agriculture Farms
面向连接精准农业农场的人工智能云边架构设计与评估
Abstract
Plant diseases cause significant yield losses worldwide, with tomato crops particularly susceptible to early blight, late blight, and leaf mold. Manual monitoring is practical only for small-scale farms and becomes unmanageable at larger scales. To tackle this limitation, an artificial intelligence (AI) enabled cloud-edge architecture is proposed for autonomous crop monitoring. This proposed architecture integrates Internet of Things (IoT) sensors, unmanned aerial vehicles (UAVs), deep learning, Azure IoT Hub-based cloud analytics, and multi-platform (mobile app, web app, and embedded edge device platform) interfaces to enable real-time detection of tomato diseases. For training and validation, we used publicly available datasets, such as PlantVillage and Kaggle. A TensorFlow model trained on a collected dataset is deployed across mobile, web, and edge-device platforms. Experimental results show detection effectiveness around 92-95%, with consistent performance over diverse environments and device platforms. The proposed system improves disease detection effectiveness, lowers dependence on manual inspection, and enables prompt interventions, thereby supporting sustainable, connected precision agriculture farms.
Chinese Translation
植物疾病在全球范围内造成了显著的产量损失,番茄作物尤其易受早疫病、晚疫病和叶霉病的影响。人工监测仅适用于小规模农场,而在更大规模的农场中则变得难以管理。为了解决这一限制,本文提出了一种基于人工智能(AI)的云边架构,用于自主作物监测。该架构集成了物联网(IoT)传感器、无人机(UAV)、深度学习、基于Azure IoT Hub的云分析以及多平台(移动应用、网页应用和嵌入式边缘设备平台)接口,以实现番茄疾病的实时检测。我们使用了公开可用的数据集进行训练和验证,如PlantVillage和Kaggle。一个在收集的数据集上训练的TensorFlow模型被部署在移动、网页和边缘设备平台上。实验结果显示检测有效性约为92-95%,在不同环境和设备平台上表现一致。所提出的系统提高了疾病检测的有效性,降低了对人工检查的依赖,并能够及时进行干预,从而支持可持续的连接精准农业农场。
cs.RO / 43 / 2608.03820
Designing Social Robots for Inclusive Child Wellbeing Assessment: Insights from Communities Supporting Developmental Language Disorder and Forced Migration
为包容性儿童福祉评估设计社会机器人:来自支持发展性语言障碍和被迫迁移社区的见解
Abstract
Assessing children's wellbeing and mental health can be particularly challenging for children experiencing communication barriers, such as children with Developmental Language Disorder (DLD) and children with forced migration backgrounds. During the assessment process, traditional self-report questionnaires place substantial demands on language comprehension and verbal expression. In this context, social robots have emerged as a promising tool for supporting wellbeing assessment without solely relying on self-report questionnaires, yet limited research has examined how such interactions can be designed to be inclusive, appropriate, and ethically acceptable for children with diverse communication needs. To address this gap, we created candidate child--robot interaction activities as design probes and conducted focus groups with parents and professionals supporting children with DLD and children with forced migration backgrounds. Through thematic analysis, we identified considerations relating to robot role and capabilities, interactional dynamics, individual differences, and child agency, alongside population-specific considerations shaped by children's communication needs and lived experiences. Based on these findings, we derive a set of ethical and inclusive design recommendations for robot-mediated wellbeing assessment. By foregrounding these considerations and recommendations, this work contributes design guidance for inclusive robot-mediated wellbeing assessments for children with diverse communication needs.
Chinese Translation
评估儿童的福祉和心理健康对于经历沟通障碍的儿童,尤其是发展性语言障碍(DLD)儿童和有被迫迁移背景的儿童来说,尤其具有挑战性。在评估过程中,传统的自我报告问卷对语言理解和口头表达提出了较高的要求。在这种背景下,社会机器人作为一种有前景的工具,能够在不完全依赖自我报告问卷的情况下支持福祉评估,但目前对如何设计这种互动以适应不同沟通需求的儿童的研究仍然有限。为了解决这一空白,我们创建了候选的儿童-机器人互动活动作为设计探针,并与支持DLD儿童和有被迫迁移背景儿童的父母和专业人士进行了焦点小组讨论。通过主题分析,我们识别了与机器人角色和能力、互动动态、个体差异以及儿童自主性相关的考虑因素,以及由儿童的沟通需求和生活经历塑造的人群特定考虑因素。基于这些发现,我们提出了一套针对机器人介导的福祉评估的伦理和包容性设计建议。通过突出这些考虑和建议,本研究为具有不同沟通需求的儿童的包容性机器人介导福祉评估提供了设计指导。
cs.RO / 44 / 2608.03872
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning
EvoHIL:用于鲁棒人机协作强化学习的自我演化奖励与流匹配策略优化
Abstract
Human-in-the-loop reinforcement learning (HIL-RL) enables robots to learn contact-rich manipulation from limited real-world interaction, but deployment exposes three coupled limitations: static visual reward models fail under scene changes; independently sampled actions cause temporally inconsistent motion; and vision-based policies remain sensitive to appearance shifts. We present EvoHIL, a unified framework that adapts the reward model, action generator, and visual do main within a staged human-in-the-loop learning process. First, self-evolving reward (SER) adapts the success classifier from human-confirmed positives and provisional weak negatives. Second, Action Flow Stabilization (AFS) generates temporally coherent action chunks through flow matching, grounding policy updates in executed action prefixes and demonstrated behavior. Third, retention-aware offline fine-tuning replays relit interaction data while anchoring the AFS actor-critic to prior behavior, adapting the visual domain without additional robot interaction. Across six manipulation tasks on Franka FR3 and SO-101 arms under a controlled lighting shift, EvoHIL improves task success, agreement with human-confirmation labels, motion smoothness, and completion time relative to human-in-the-loop and imitation baselines.Project page: https://anonymous4366.github.io/EvoHIL/
Chinese Translation
人机协作强化学习(HIL-RL)使机器人能够通过有限的现实世界交互学习接触丰富的操作,但在部署过程中暴露出三种相互关联的局限性:静态视觉奖励模型在场景变化下失效;独立采样的动作导致时间上不一致的运动;基于视觉的策略对外观变化保持敏感。我们提出了EvoHIL,一个统一框架,在分阶段的人机协作学习过程中适应奖励模型、动作生成器和视觉领域。首先,自我演化奖励(SER)从人类确认的正样本和临时弱负样本中调整成功分类器。其次,动作流稳定化(AFS)通过流匹配生成时间一致的动作块,将策略更新基于执行的动作前缀和示范行为。第三,关注保留的离线微调在锚定AFS演员-评论家到先前行为的同时重放重新点燃的交互数据,在没有额外机器人交互的情况下适应视觉领域。在控制光照变化下的Franka FR3和SO-101臂的六个操作任务中,EvoHIL相较于人机协作和模仿基线提高了任务成功率、人类确认标签的一致性、运动平滑性和完成时间。项目页面:https://anonymous4366.github.io/EvoHIL/
cs.RO / 45 / 2608.03924
ETA: A New Agentic Paradigm for Embodied Tasks
ETA:一种用于具身任务的新代理范式
Abstract
When will robots have their ChatGPT moment? Such a breakthrough requires a general-purpose robot that can handle unfamiliar tasks in unfamiliar environments, remain controllable over long interactions, and learn from experience. Today's embodied systems largely follow an end-to-end observation-to-action path. Despite rapid progress, they remain far from this goal: their generalization depends heavily on the coverage of robot training data, while long task execution remains difficult to control and inspect. To realize this goal, we introduce the Embodied Task Agent (ETA), a new paradigm for extending digital agents into the physical world, and release OpenETA as its open-source implementation. ETA centers the robot around a Planner that chooses one Tool call at a time, an Interface that controls execution, and a World that returns the result and a fresh observation. This loop allows the agent to verify outcomes, adapt its plan, and turn successful and failed interactions into reusable experience. OpenETA provides replaceable Planners, composable Tools and Skills, auditable memory, replayable trajectories, and common interfaces for simulation and real robots. For Codex, OpenETA can operate as a lightweight plugin that exposes only observe, mark_point, and move_to.
Chinese Translation
机器人何时会迎来它们的 ChatGPT 时刻?这样的突破需要一种通用机器人,能够在不熟悉的环境中处理不熟悉的任务,在长时间的交互中保持可控性,并从经验中学习。当前的具身系统主要遵循端到端的观察到行动路径。尽管取得了快速进展,但它们距离这一目标仍然相去甚远:它们的泛化能力在很大程度上依赖于机器人训练数据的覆盖范围,而长时间任务执行的控制和检查仍然困难。为了实现这一目标,我们提出了具身任务代理(Embodied Task Agent,ETA),一种将数字代理扩展到物理世界的新范式,并发布了其开源实现 OpenETA。ETA 以一个规划器(Planner)为中心,该规划器一次选择一个工具调用(Tool call),一个接口(Interface)控制执行,以及一个返回结果和新观察的世界(World)。这个循环使代理能够验证结果,调整计划,并将成功和失败的交互转化为可重用的经验。OpenETA 提供可替换的规划器、可组合的工具和技能、可审计的记忆、可重放的轨迹,以及用于仿真和真实机器人的通用接口。对于 Codex,OpenETA 可以作为一个轻量级插件,仅暴露 observe、mark_point 和 move_to 功能。
cs.RO / 46 / 2608.03938
Bimanual Manipulation Within an 8 GB Budget: Zero-Copy Sensing and Quantized ACT on an Entry-Level Jetson
在8 GB预算内的双手操控:零拷贝感知与量化ACT在入门级Jetson上的应用
Abstract
Bimanual manipulation policies trained with imitation learning are typically evaluated on workstation or datacenter-class GPUs, leaving the cost of deploying them on embedded hardware largely uncharacterized. We present a bimanual SO-101 system running entirely on an NVIDIA Jetson Orin Nano Super (8 GB), the entry-level tier of NVIDIA's embedded line, using a desktop GPU (RTX 3070) only for offline training, evaluated on pick-and-place of a deformable beanbag. First, we build a GStreamer capture pipeline backed by NVMM buffers that removes redundant host-device copies from three-camera sensing. Contrary to expectation, the conventional path fit the memory budget and dropped no frames; what zero-copy sensing recovers is CPU headroom (peak single-core utilization 98.0% to 77.0%) and worst-case latency (117.31 ms to 101.52 ms). Second, we train ACT and Diffusion Policy on identical demonstrations, each at its own reference budget (100k gradient steps for ACT, 200k for Diffusion Policy). ACT converges to a task-competent policy (19/20 trials) while Diffusion Policy does not converge to a usable one (0/10) even at twice the step count, which we attribute to differing convergence costs rather than an accuracy ceiling. Third, we convert ACT to TensorRT. FP16 reduces mean inference latency from 114.02 ms to 17.93 ms (6.4x) and INT8 to 12.65 ms (9.0x), with task success preserved at all three precisions (19/20, 18/20, 19/20). We report two findings not previously documented for ACT: TensorRT's general-purpose INT8 calibration quantizes the ResNet18 backbone but accepts zero of 145 transformer layers, explaining INT8's negligible size reduction over FP16 (0.9%) despite a further 28% latency gain; and the need for quantization is conditional on ACT's action-chunking configuration, feasible in full precision at n_action_steps = 100 but not at the per-step re-prediction temporal ensembling requires.
Chinese Translation
使用模仿学习训练的双手操控策略通常在工作站或数据中心级GPU上进行评估,这使得在嵌入式硬件上部署它们的成本尚未得到充分表征。我们展示了一个完全运行在NVIDIA Jetson Orin Nano Super(8 GB)上的双手SO-101系统,仅在离线训练时使用桌面GPU(RTX 3070),并在可变形豆袋的拾取和放置任务上进行评估。首先,我们构建了一个基于NVMM缓冲区的GStreamer捕获管道,消除了来自三摄像头感知的冗余主机-设备拷贝。与预期相反,传统路径符合内存预算且未丢帧;零拷贝感知所恢复的是CPU的余量(峰值单核利用率从98.0%降至77.0%)和最坏情况下的延迟(从117.31毫秒降至101.52毫秒)。其次,我们在相同演示上训练了ACT和Diffusion Policy,每个策略在其各自的参考预算下(ACT为100k梯度步骤,Diffusion Policy为200k)。ACT收敛到一个任务合格的策略(19/20次试验),而Diffusion Policy即使在两倍的步骤数下也未能收敛到可用策略(0/10),我们将此归因于收敛成本的差异,而非准确性上限。第三,我们将ACT转换为TensorRT。FP16将平均推理延迟从114.02毫秒降低到17.93毫秒(6.4倍),而INT8则降至12.65毫秒(9.0倍),在所有三种精度下任务成功率保持不变(19/20,18/20,19/20)。我们报告了两个之前未记录的ACT发现:TensorRT的通用INT8校准量化了ResNet18主干,但接受了145个变换器层中的零个,这解释了尽管延迟进一步提高28%,INT8相较FP16的大小减小微乎其微(0.9%);并且量化的需求依赖于ACT的动作分块配置,在n_action_steps = 100的全精度下可行,但在每步重新预测的时间集成要求下则不可行。
cs.RO / 47 / 2608.03978
Stochastic Multiple Shooting Trajectory Optimization via Sequential Local Policy Evaluation
通过序列局部策略评估的随机多重射击轨迹优化
Abstract
Stochastic single shooting trajectory optimization methods such as Model Predictive Path Integral control (MPPI) have been widely adopted in robotics due to their ability to reason about probabilistic dynamics and provide solutions where model gradients are noisy, costly to evaluate, or unavailable. However, satisfaction of terminal constraints when shooting over long action sequences is often sample inefficient, requiring a large number of iterations for convergence. In this paper, we present a stochastic multiple shooting method that optimizes short control action sequences connected via local feedback policies to improve sample efficiency and convergence to a terminal set. Additionally, we show that we are able to synthesize approximate system Jacobians purely from rollouts, making the method suitable for model-based reinforcement learning with black-box dynamics. We demonstrate the algorithm has improved sample efficiency and terminal set convergence for three nonlinear, underactuated optimization problems: a classic cartpole swingup task with analytical dynamics, a cartpole swingup task with learned neural network dynamics, and a VTOL quadplane performing a high angle-of-attack, precision post-stall landing maneuver.
Chinese Translation
随机单次射击轨迹优化方法,如模型预测路径积分控制(Model Predictive Path Integral control, MPPI),因其能够处理概率动态并在模型梯度噪声大、评估成本高或不可用的情况下提供解决方案而被广泛应用于机器人领域。然而,在长时间动作序列上进行射击时,满足终端约束通常效率低下,需大量迭代才能收敛。本文提出了一种随机多重射击方法,通过局部反馈策略优化短控制动作序列,以提高样本效率和收敛到终端集。此外,我们展示了能够仅通过回放合成近似系统雅可比矩阵,使得该方法适用于具有黑箱动态的基于模型的强化学习。我们证明了该算法在三个非线性欠驱动优化问题上提高了样本效率和终端集收敛性:一个经典的带有解析动态的倒立摆摆动任务、一个带有学习的神经网络动态的倒立摆摆动任务,以及一个执行高攻角精确失速着陆机动的垂直起降四旋翼飞机(VTOL quadplane)。
cs.CV / 1 / 2608.02711
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
Hunyuan3D-Buffalo 1.0:一个统一的多模态模型,用于可扩展的3D生成、理解和编辑
Abstract
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D-Buffalo 1.0, a unified framework supporting 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation within a single architecture. To enable scalable training, we construct an 87M-scale 3D multimodal corpus, comprising 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs generated using Nano3D-v2. Architecturally, the framework combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high-fidelity 3D synthesis. The VLM provides multimodal semantic conditions for generation, while editing and part generation additionally condition the diffusion process on the source object representation to preserve its overall structure and unedited regions. Extensive experiments show that Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance on text-to-3D generation and 3D editing benchmarks, while exhibiting strong understanding and part-generation capabilities. Our analysis further shows that both generation and understanding improve editing, demonstrating the effectiveness of unified 3D multimodal training. Project Page: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/
Chinese Translation
最近在图像生成方面的进展展示了统一多模态模型在理解、生成和编辑方面的潜力。然而,统一的3D建模仍然受到稀缺多模态数据的限制,特别是缺乏大规模和几何一致的编辑数据。为了解决这一限制,我们提出了Hunyuan3D-Buffalo 1.0,这是一个统一框架,支持3D理解、文本到3D生成、指令引导的3D编辑以及文本基础的部件生成,均在单一架构内实现。为了实现可扩展的训练,我们构建了一个87M规模的3D多模态语料库,其中包含2500万条理解样本、5000万对文本到3D的配对以及1200万对使用Nano3D-v2生成的编辑配对。在架构上,该框架结合了Hunyuan3D-VLM用于语义、结构和空间理解,以及Hunyuan3D DiT用于高保真3D合成。VLM为生成提供了多模态语义条件,而编辑和部件生成则进一步将扩散过程条件化于源对象表示,以保持其整体结构和未编辑区域。大量实验表明,Hunyuan3D-Buffalo 1.0在文本到3D生成和3D编辑基准测试中实现了最先进或领先的性能,同时展现出强大的理解和部件生成能力。我们的分析进一步表明,生成和理解均能改善编辑,证明了统一3D多模态训练的有效性。项目页面:https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/
cs.CV / 2 / 2608.02713
Quo Vadis, World Modeling?
世界建模的未来何在?
Abstract
Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to parallelize. World modeling offers a natural intermediate proxy that allows agents to query lower-cost, more controllable feedback before committing to real actions. Classical world models instantiate this proxy primarily through future physical-state prediction, a formulation useful yet narrow for agents that require actionable feedback beyond raw state transitions. In this work, we conceptualize Agent-Centric Interactive World Proxies, shifting the fundamental paradigm from physical state transitions to agent-usable information transitions, such as execution outcomes, retrieved experiences or skills, and verification signals, broadening the scope of world modeling to provide versatile feedback for continually improving agents. To systematically map this design space, we organize world proxies into six functional forms based on their feedback modalities: dynamics, spatial, execution, memory/experience, skill, and reward/verification proxies, which together characterize the primary ways world modeling serves agent improvement. We further analyze how these proxies empower agents across three progressive levels: L.1 Inference-Time Guidance, where proxy outputs enrich in-context information for superior decisions; L.2 Training-Time Optimization, where proxy outputs yield rewards, critiques, or synthetic rollouts for policy learning; and L.3 Agent-Proxy Co-Evolution, where real-environment evidence continuously updates both the proxy and the agent for co-evolution. Ultimately, this work recasts world modeling into an agent-centric paradigm, establishing a roadmap for building world proxies that empower agents to plan better, learn faster, and evolve continually.
Chinese Translation
持续改进的智能体需要动态的互动反馈,而不仅仅是静态监督,然而直接与真实环境的互动成本高、速度慢、安全性差且难以并行化。世界建模提供了一种自然的中介代理,使智能体能够在采取真实行动之前查询成本更低、可控性更强的反馈。经典的世界模型主要通过未来物理状态预测来实现这一代理,这种形式对于需要可操作反馈的智能体来说虽然有用,但却过于狭窄。在本研究中,我们概念化了以智能体为中心的互动世界代理,将基本范式从物理状态转变为智能体可用的信息转移,例如执行结果、检索的经验或技能以及验证信号,从而拓宽了世界建模的范围,以提供多样化的反馈,支持持续改进的智能体。为了系统性地映射这一设计空间,我们根据反馈方式将世界代理组织为六种功能形式:动态代理、空间代理、执行代理、记忆/经验代理、技能代理和奖励/验证代理,这些代理共同描述了世界建模服务于智能体改进的主要方式。我们进一步分析了这些代理如何在三个渐进层次上赋能智能体:L.1 推理时指导,其中代理输出丰富上下文信息以做出更优决策;L.2 训练时优化,其中代理输出产生奖励、批评或合成回放以进行策略学习;L.3 智能体-代理共同进化,其中真实环境证据持续更新代理和智能体,实现共同进化。最终,本研究将世界建模重新构建为以智能体为中心的范式,为构建能够赋能智能体更好规划、更快学习和持续进化的世界代理奠定了基础。
cs.CV / 3 / 2608.02762
Oh Deer, How Should I Handle This? Seasonal Priors for Selective Wildlife Annotation and Classification
哦,亲爱的,我该如何处理这个?选择性野生动物标注和分类的季节性先验
Abstract
Fine-grained wildlife classification in aerial imagery is limited not only by model performance, but also by unreliable labels: animals occupy few pixels, key visual cues vary seasonally, and modality-specific evidence can be ambiguous. We study adult-male identification in red deer, where the antler cycle defines predictable windows of reliable evidence for both annotation and prediction. Using 7,295 RGB-only, thermal-only, and matched RGB+thermal crop sets labeled by three annotators, we show that seasonal structure links (I) annotation quality, (II) downstream classification, and (III) selective prediction. Matched RGB+thermal review resolves more samples than either single modality, recovering majority-male labels otherwise missed by RGB or thermal alone, in human based as well as model based classification. Months with high annotator abstention also show lower classifier confidence, and soft seasonal priors mainly benefit the season-limited thermal view. Uncertainty-band abstention further improves covered accuracy up to 98.9%, though at reduced coverage and with deferral that falls disproportionately on males. Overall, a biologically grounded seasonal calendar predicts where annotation and prediction are unreliable, and can guide both annotation protocol design and modality weighting.
Chinese Translation
在航空影像中,细粒度野生动物分类不仅受到模型性能的限制,还受到不可靠标签的影响:动物占据的像素很少,关键视觉线索随季节变化,而特定模态的证据可能模糊不清。我们研究了红鹿的成年雄性识别,其中鹿角周期定义了可预测的可靠证据窗口,适用于标注和预测。通过使用由三位标注者标注的7,295个仅RGB、仅热成像和匹配的RGB+热成像裁剪集,我们展示了季节性结构如何关联(I)标注质量,(II)下游分类,以及(III)选择性预测。匹配的RGB+热成像审查比任何单一模态解决更多样本,恢复了RGB或热成像单独无法捕捉的主要雄性标签,无论是在基于人类的还是基于模型的分类中。标注者缺席率较高的月份也显示出分类器信心较低,而软季节性先验主要有利于季节限制的热成像视图。不确定性带缺席进一步将覆盖准确率提高至98.9%,尽管覆盖率降低,并且推迟的情况在雄性上表现得不成比例。总体而言,基于生物学的季节性日历预测了标注和预测不可靠的地方,并可以指导标注协议设计和模态加权。
cs.CV / 4 / 2608.02790
Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRI
自信但不可靠:对脑部MRI的视觉-语言模型的行为安全审计
Abstract
Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evaluated separately from correctness. We use brain MRI as a controlled, high-stakes testbed for a broader failure mode in frontier multimodal systems: models can appear competent while lacking reliable self-knowledge. We present an automatically graded behavioral audit and pilot study of six instruction-tuned VLMs (five general-purpose and one medical specialist) on 4,102 images (4,032 axial/coronal/sagittal MRI slices from 250 subjects plus 70 non-brain/noise controls), with labels derived from public metadata and released expert segmentation masks rather than new human annotation. Across models, answer coverage is near-complete, but verbalized-confidence calibration is poor: ECE ranges from 0.27 to 0.40, mean confidence on incorrect answers ranges from 0.82 to 0.97, and 33-46% of answered items are high-confidence errors. The most accurate model is also the most confident on its errors, while a base/specialist family contrast suggests that medical adaptation improves tumor-presence detection without improving confidence reliability. Open-ended diagnostics further show that hallucination and abstention vary separately from multiple-choice accuracy. These findings argue that medical-image VLM evaluation should report verbalized-confidence reliability, confident error, hallucination, and abstention alongside accuracy.
Chinese Translation
视觉-语言模型(VLMs),包括医疗专业模型,越来越多地被提议用于医学影像,但其所声称的自信度很少与正确性分开评估。我们使用脑部MRI作为一个受控的高风险测试平台,探讨前沿多模态系统中的更广泛失效模式:模型可能看起来有能力,但缺乏可靠的自我认知。我们展示了一项自动评分的行为审计和对六个经过指令调优的VLMs(五个通用模型和一个医疗专业模型)进行的初步研究,涉及4,102幅图像(来自250名受试者的4,032幅轴向/冠状面/矢状面MRI切片以及70幅非脑/噪声对照图像),其标签来源于公共元数据和发布的专家分割掩膜,而非新的人工注释。在各模型中,答案覆盖率接近完整,但口头自信度校准较差:ECE范围从0.27到0.40,错误答案的平均自信度范围从0.82到0.97,33-46%的回答项目是高自信错误。最准确的模型在其错误上也是最自信的,而基础/专业模型的对比表明,医学适应性提高了肿瘤存在检测的能力,但并未改善自信度的可靠性。开放式诊断进一步显示,幻觉和弃权与多项选择准确性是独立变化的。这些发现表明,医学影像VLM的评估应报告口头自信度的可靠性、自信错误、幻觉和弃权,以及准确性。
cs.CV / 5 / 2608.02791
Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
更好、更强、更快、更广:基于 MLLM 的结构化全掩码预测用于分割
Abstract
MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability, and fast inference. Embedding-prediction methods may disrupt language modeling through pixel-level objectives, whereas next-token generation is inefficient for dense masks. We propose All-Mask Prediction, decoupling autoregressive dialogue from non-autoregressive mask prediction. Its binary instantiation, STAMP (Simultaneous Textual All-Mask Prediction), emits an in-vocabulary trigger, fuses image-aligned mask tokens with corresponding patch features, and uses hybrid attention to classify all tokens as foreground or background in one pass. It thereby combines strong referring and reasoning segmentation with preserved multimodal ability and efficient inference. However, binary masks cannot retain multiple semantic or instance identities without repeated target-specific predictions. We therefore propose Structured All-Mask Prediction and develop STAMPlus. It generates a target list with explicit IDs and optional boxes, binds these IDs to a shared multi-class mask space, and jointly predicts all targets in one non-autoregressive pass. A single unified checkpoint retains STAMP's referring and reasoning capabilities while extending to open-vocabulary semantic, instance-aware, and remote-sensing small-target segmentation, where high-resolution mask-token scaling preserves finer spatial evidence. Across these settings, STAMPlus achieves state-of-the-art segmentation performance, preserves general multimodal instruction following, and reduces 12-category latency from 13.50s for repeated STAMP inference to 5.16s. Further analyses show that accurate target cues improve segmentation and learned spatial grounding benefits look-twice reasoning. Overall, STAMPlus resolves the trilemma beyond single-target prediction.
Chinese Translation
基于 MLLM 的分割面临核心的分割三难问题:高分割性能、保留对话能力和快速推理。嵌入预测方法可能通过像素级目标干扰语言建模,而下一个标记生成对于密集掩码效率低下。我们提出了全掩码预测(All-Mask Prediction),将自回归对话与非自回归掩码预测解耦。其二元实例 STAMP(同时文本全掩码预测)发出词汇内的 触发器,将与图像对齐的掩码标记与相应的补丁特征融合,并使用混合注意力在一次传递中将所有标记分类为前景或背景。这样,它结合了强大的指代和推理分割,同时保留了多模态能力和高效推理。然而,二元掩码无法在没有重复目标特定预测的情况下保留多个语义或实例身份。因此,我们提出了结构化全掩码预测(Structured All-Mask Prediction),并开发了 STAMPlus。它生成一个具有明确 ID 和可选框的目标列表,将这些 ID 绑定到共享的多类掩码空间,并在一次非自回归传递中联合预测所有目标。单个统一的检查点保留了 STAMP 的指代和推理能力,同时扩展到开放词汇的语义、实例感知和遥感小目标分割,其中高分辨率掩码标记缩放保留了更细的空间证据。在这些设置中,STAMPlus 实现了最先进的分割性能,保留了一般的多模态指令跟随,并将 12 类的延迟从重复 STAMP 推理的 13.50 秒减少到 5.16 秒。进一步分析表明,准确的目标线索改善了分割效果,学习的空间定位有助于仔细推理。总体而言,STAMPlus 解决了超越单目标预测的三难问题。
cs.CV / 6 / 2608.02792
PixelUp: Zero-Shot Semantic Feature Upsampling for Fine-Grained Vision Tasks
PixelUp:用于细粒度视觉任务的零-shot语义特征上采样
Abstract
Self-supervised Vision Foundation Models (VFMs) have become essential backbones for downstream tasks due to their strong and transferable visual representations. However, their patch-token-level features are often too coarse for dense prediction tasks such as semantic segmentation and depth estimation when accurate fine-grained predictions are required. Feature upsampling methods have been developed to recover pixel-level detail but still face limitations. Learnable upsamplers are often designed for a specific encoders and must be retrained for different encoders. Image-guided methods that use shallow pixel encoders often introduce textural artifacts and lack the semantic guidance needed for accurate downstream predictions. We introduce PixelUp, a zero-shot VFM-agnostic upsampler achieving semantic awareness through a coarse-to-fine chain of windowed cross-attention architecture guided by multi-scale semantic features. We demonstrate that PixelUp outperforms both VFM-specific and VFM-agnostic upsamplers, achieving state-of-the-art performance on dense prediction tasks with an average improvement of +1.2 mIoU on semantic segmentation and +0.25 $\delta_1$, on NYUv2 depth estimation across VFMs. PixelUp further improves training-free open-vocabulary and unsupervised semantic segmentation by an average of +1.3 mIoU and +0.5 mIoU, respectively. Code available at https://pixelup-project.vercel.app/
Chinese Translation
自监督视觉基础模型(VFMs)由于其强大且可迁移的视觉表征,已成为下游任务的重要支柱。然而,当需要精确的细粒度预测时,它们的补丁令牌级特征往往过于粗糙,无法满足语义分割和深度估计等密集预测任务的要求。虽然已经开发了特征上采样方法以恢复像素级细节,但仍然面临一些限制。可学习的上采样器通常是为特定编码器设计的,必须为不同的编码器重新训练。使用浅层像素编码器的图像引导方法往往会引入纹理伪影,并缺乏进行准确下游预测所需的语义指导。我们提出了PixelUp,这是一种零-shot且与VFM无关的上采样器,通过由多尺度语义特征引导的粗到细的窗口交叉注意力架构实现语义感知。我们证明PixelUp在密集预测任务中优于VFM特定和VFM无关的上采样器,在语义分割上平均提高了+1.2 mIoU,在NYUv2深度估计上提高了+0.25 $ ext{δ}_1$。PixelUp进一步改善了无训练的开放词汇和无监督语义分割,分别平均提高了+1.3 mIoU和+0.5 mIoU。代码可在 https://pixelup-project.vercel.app/ 获取。
cs.CV / 7 / 2608.02803
SAGE: Semantic Explainability of Attention-Based Survival Models in Computational Pathology
SAGE:基于语义的注意力生存模型在计算病理学中的可解释性
Abstract
Attention-based multiple instance learning (ABMIL) is the predominant approach for slide-level prediction in computational pathology, yet its attention maps provide only local explanations: they indicate where a model focuses but not which histological features drive its predictions or how the model behaves across a patient cohort. We present Semantic Attention Global Explanations (SAGE), a post-hoc framework that extracts global, language-grounded explanations from a frozen ABMIL model. Using a pathology vision-language model, SAGE scores image patches against a dictionary of 25 histological concepts, aggregates these scores according to the model's learned attention, and quantifies how each concept relates to prediction risk across a cohort. Applied to survival prediction using seven TCGA cancer cohorts and three foundation models, SAGE recovered established prognostic features, such as the adverse association of necrosis, while revealing cancer-specific biology, including a favorable angiogenic signature in renal cell carcinoma consistent with known molecular subtypes. Ablation studies demonstrated that these associations depend on the model's learned attention rather than concept prevalence alone, and that the concept dictionary captures much of the prognostic information encoded by the foundation model features. Through semantically-grounded explanations, SAGE provides a scalable, model-agnostic framework for understanding what ABMIL survival models learn, enabling pathologists to interpret model behavior at the cohort level and offering the potential for biomarker identification.
Chinese Translation
基于注意力的多实例学习(ABMIL)是计算病理学中滑动级预测的主要方法,但其注意力图仅提供局部解释:它们指示模型关注的位置,但并未说明哪些组织学特征驱动其预测,或模型在患者队列中的行为。我们提出了语义注意力全局解释(SAGE),这是一个后处理框架,从冻结的ABMIL模型中提取全局的、基于语言的解释。SAGE使用病理视觉-语言模型,对25个组织学概念的词典进行图像补丁评分,根据模型学习的注意力聚合这些评分,并量化每个概念与队列中预测风险的关系。在使用七个TCGA癌症队列和三个基础模型进行生存预测时,SAGE恢复了已建立的预后特征,例如坏死的不良关联,同时揭示了特定于癌症的生物学,包括与已知分子亚型一致的肾细胞癌中的有利血管生成特征。消融研究表明,这些关联依赖于模型学习的注意力,而不仅仅是概念的普遍性,并且概念词典捕获了基础模型特征编码的大部分预后信息。通过基于语义的解释,SAGE提供了一个可扩展的、与模型无关的框架,以理解ABMIL生存模型所学习的内容,使病理学家能够在队列级别上解释模型行为,并提供生物标志物识别的潜力。
cs.CV / 8 / 2608.02805
A Unified 2D Framework for DeepLesion Detection, Segmentation and Short Report Generation
统一的二维框架用于深度病变检测、分割和短报告生成
Abstract
In previous work, we integrated large language models (LLMs) into the lesion segmentation model based on the ULS23 DeepLesion dataset, using short-form findings from the reports. In this study, we developed a unified 2D lesion analysis framework that integrates LLM-based reasoning, lesion bounding box detection, segmentation, and radiology report generation from the original DeepLesion dataset. In the testing phase, we achieved relatively high lesion bounding box detection accuracy with mAP50 of 70.1%, mAP50-95 of 46.4%; Lesion segmentation performance with a Dice score of 62.6%; short report generation accuracy with BLEU_1 score of 64.3%, BLEU_4 score of 49.6%, METEOR of 34.7%, and ROUGE_L of 60.1%. In this work, we address the challenging issue of segmentation in the original DeepLesion dataset and achieve a 28.5% Dice score improvement over the nnUNet lesion segmentation model. We also integrated spatial and anatomical context into the DeepLesion short report generation. We released the implementation, dataset, and models on Github. https://github.com/ruida/2D_DeepLesion_Foundation
Chinese Translation
在之前的研究中,我们将大型语言模型(LLMs)集成到基于ULS23 DeepLesion数据集的病变分割模型中,使用报告中的短期发现。在本研究中,我们开发了一个统一的二维病变分析框架,该框架集成了基于LLM的推理、病变边界框检测、分割和从原始DeepLesion数据集中生成放射学报告。在测试阶段,我们实现了相对较高的病变边界框检测准确率,mAP50为70.1%,mAP50-95为46.4%;病变分割性能的Dice分数为62.6%;短报告生成准确率的BLEU_1分数为64.3%,BLEU_4分数为49.6%,METEOR为34.7%,ROUGE_L为60.1%。在这项工作中,我们解决了原始DeepLesion数据集中分割的挑战性问题,并在nnUNet病变分割模型上实现了28.5%的Dice分数提升。我们还将空间和解剖上下文集成到DeepLesion短报告生成中。我们在Github上发布了实现、数据集和模型。https://github.com/ruida/2D_DeepLesion_Foundation
cs.CV / 9 / 2608.02830
In-Context Collapse in Vision-Language Models and How to Mitigate it?
视觉-语言模型中的上下文崩溃及其缓解方法
Abstract
Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demonstrations are supplied. We show the opposite: as demonstrations accumulate, a subset of VLMs undergo an \emph{in-context collapse}, a sharp, sometimes catastrophic accuracy drop spanning synthetic classification, natural-image classification, and VQA benchmarks, in some models falling below chance while outputs remain well-formed. Across an open VLM panel ($0.5$B--$11$B) and a frontier model (Claude Sonnet 4.5), the collapse is graded. Two capabilities turn out to be dissociable: robustness to accumulating demonstrations and the ability to learn a novel rule in context, their combinations yield three reproducible regimes. A parameter-matched lesion-and-rescue causally localizes the collapse to the vision-language integration pathway: an adapter on the connector and early/mid layers restores genuine learning (remap accuracy $0.39!\rightarrow!0.91$ at 16 shots), while an equal-capacity adapter on the late readout does not. We propose \textsc{CircA}, whose core is a one-time integration vaccine: trained once on one synthetic task, it transfers collapse-resistance to unseen task families (chance$\rightarrow$$0.71$/$0.60$ on CIFAR/Fashion). The layers best for in-context integration are not the layers best for weight-based consolidation, the late readout achieves higher accuracy and less forgetting at fewer parameters. The collapse is an integration failure at the vision--language interface, correctable by a lightweight, transferable intervention.
Chinese Translation
多次上下文学习(ICL)使视觉-语言模型(VLMs)能够在不更新权重的情况下,从图像-标签示例中进行适应,广泛认为随着示例数量的增加,其性能会得到提升。然而,我们的研究表明,情况正好相反:随着示例的累积,部分VLMs会经历一种 extit{上下文崩溃},这是一种急剧的、在合成分类、自然图像分类和视觉问答(VQA)基准测试中有时甚至是灾难性的准确率下降,在某些模型中准确率降至随机水平,同时输出仍然保持良好。在一个开放的VLM面板($0.5$B--$11$B)和一个前沿模型(Claude Sonnet 4.5)中,崩溃现象呈现出分级特征。研究发现,两种能力是可分离的:对累积示例的鲁棒性和在上下文中学习新规则的能力,它们的组合产生了三种可重复的状态。通过参数匹配的损伤与恢复实验,因果定位将崩溃归因于视觉-语言整合路径:在连接器和早期/中层的适配器恢复了真正的学习(在16个示例时重映射准确率从$0.39$提升至$0.91$),而在后期读出层的相同容量适配器则无效。我们提出了 extsc{CircA},其核心是一次性整合疫苗:在一个合成任务上训练一次后,它能够将抗崩溃能力转移到未见过的任务家族(在CIFAR/Fashion上从随机水平提升至$0.71$/$0.60$)。最适合上下文整合的层并不是最适合基于权重整合的层,后期读出层在参数较少的情况下实现了更高的准确率和更少的遗忘。崩溃是视觉-语言接口的整合失败,可以通过轻量、可转移的干预措施进行纠正。
cs.CV / 10 / 2608.02833
CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
CURV:通过课程视觉基础推理增强图表理解
Abstract
Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning chains. While extrinsic chain-of-thought prompting and visual cues significantly improve performance, current MLLMs lack intrinsic visual grounded reasoning capabilities, leading to inaccurate perception and reasoning disconnected from visual evidence. To address these limitations, we propose CURV, a curriculum learning framework that develops intrinsic visual reasoning capabilities by reformulating CQA as multi-step visual grounded reasoning, where each step coordinates logical reasoning with dynamic visual grounding through spatial attention concentration. To assist model learning, we further introduce CCQA, a three-level curriculum dataset with scalable synthetic generation across diverse chart types and reasoning patterns. Our curriculum systematically progresses from basic single-operation reasoning to complex multi-chart compositional tasks. Experiments demonstrate that CURV achieves up to $\uparrow20.50\%$ improvements over baselines and is generalizable to real-world benchmarks (up to $\uparrow12.30\%$) and out-of-domain multimodal reasoning tasks (up to $\uparrow10.20\%$), validating the effectiveness of internalizing visual reasoning with dynamic grounding for enhanced chart understanding capabilities. Code is available at: https://xhguo7.github.io/CURV/.
Chinese Translation
图表问答(CQA)要求多模态大型语言模型(MLLMs)将视觉理解与逻辑推理相结合,但当前模型在准确的视觉基础和连贯的推理链方面存在困难。尽管外部的思维链提示和视觉线索显著提高了性能,但当前的MLLMs缺乏内在的视觉基础推理能力,导致对视觉证据的感知不准确和推理与视觉证据脱节。为了解决这些局限性,我们提出了CURV,一个课程学习框架,通过将CQA重新构建为多步骤的视觉基础推理,来发展内在的视觉推理能力,其中每一步通过空间注意力集中协调逻辑推理与动态视觉基础。为了辅助模型学习,我们进一步引入了CCQA,一个具有可扩展合成生成的三级课程数据集,涵盖多种图表类型和推理模式。我们的课程系统地从基本的单操作推理进展到复杂的多图表组合任务。实验表明,CURV在基准测试中实现了高达$20.50 ext{%}$的性能提升,并且在真实世界基准(高达$12.30 ext{%}$)和领域外多模态推理任务(高达$10.20 ext{%}$)中具有良好的泛化能力,验证了通过动态基础内化视觉推理以增强图表理解能力的有效性。代码可在以下网址获取:https://xhguo7.github.io/CURV/
cs.CV / 11 / 2608.02835
A Human-in-the-Loop Deep Learning Framework for Color Reconstruction of Lenticular Films
一种人机协同的深度学习框架用于透镜胶卷的颜色重建
Abstract
Historical lenticular films, such as those created with the Kodacolor process, encode color information in a distinctive spatial format. This structure requires specialized techniques for accurate color reconstruction. While recent signal processing approaches like doLCE and deep learning methods like deep-doLCE have advanced automated color recovery, they often fail with cases such as curved lenticules, low-contrast, or badly captured regions. We propose a human-in-the-loop (HITL) deep learning framework which is designed for color reconstruction in lenticular films. Our approach introduces an editable, vector-based representation of lenticule boundaries, allowing experts to interactively refine boundary positions before color extraction and demosaicing. This decoupled architecture enables targeted corrections and iterative fine-tuning, embedding expert knowledge into the detection model and improving robustness across challenging frames. To preserve image details using information solely present in the original silver emulsion, we merge the reconstructed chrominance with the original film scan's luminance. We evaluate our pipeline on a challenging lenticular film sequence where previous automated approaches fail and the reconstructed colors are not suitable for exhibition. In contrast, our HITL approach successfully produces high-quality, exhibitable color reconstructions with preserved texture. This work is the first to combine expert guidance, editable intermediate representations, and texture-preserving post-processing for lenticular film color reconstruction, advancing the state of the art in this field.
Chinese Translation
历史透镜胶卷,例如使用Kodacolor工艺制作的胶卷,以独特的空间格式编码颜色信息。这种结构需要专门的技术来实现准确的颜色重建。尽管最近的信号处理方法如doLCE和深度学习方法如deep-doLCE在自动化颜色恢复方面取得了进展,但在处理曲面透镜、低对比度或捕捉不佳的区域时,它们往往会失败。我们提出了一种人机协同(HITL)深度学习框架,旨在用于透镜胶卷的颜色重建。我们的方法引入了一种可编辑的基于矢量的透镜边界表示,允许专家在颜色提取和去马赛克之前交互式地细化边界位置。这种解耦架构使得针对性修正和迭代微调成为可能,将专家知识嵌入到检测模型中,从而提高在复杂帧中的鲁棒性。为了利用仅存在于原始银盐乳剂中的信息来保留图像细节,我们将重建的色度与原始胶卷扫描的亮度合并。我们在一个具有挑战性的透镜胶卷序列上评估了我们的管道,在该序列中,之前的自动化方法失败,重建的颜色不适合展览。相比之下,我们的HITL方法成功地生成了高质量、适合展览的颜色重建,并保留了纹理。这项工作首次结合了专家指导、可编辑的中间表示和保留纹理的后处理技术,用于透镜胶卷的颜色重建,推动了该领域的技术进步。
cs.CV / 12 / 2608.02841
Localize, Don't Beautify: Client-Side Control of Image-Editing APIs for Cosmetic Surgery Previews
定位,而非美化:美容手术预览的客户端图像编辑API控制
Abstract
Ask a commercial image editor to preview a cosmetic procedure and it will often change more of the face than the request names: a nose edit can also smooth skin or alter lighting. Existing methods for confining an edit to one region require access to the model's internals, which a public editing API does not expose. We ask how much control is possible from the client side alone. In a pilot benchmark, six commercial editing configurations and one mask-based inpainting model perform facelift-style jaw-neck and rhinoplasty edits at three levels of client-side control: the prompt alone; cutting the edited region out of the response and pasting it back onto the original photograph through a landmark-derived mask (a masked composite); and asking the model itself to inpaint inside the mask where supported. Of 210 attempted edits, 196 could be scored. ArcFace cosine measures identity preservation; a CIELAB pixel-change ratio measures how much change lands inside the requested region rather than a protected facial zone. On the 12 frontal faces the regional metric could score, the masked composite improved localization over the paired prompt-only output by a median of 0.446 (95% face-clustered bootstrap interval 0.421-0.457) while changing the requested region about as much. Editors differed in edit strength versus identity retention, and the one inpainting model we tested did not beat the simple composite. Against each face's input-to-postoperative baseline, no editor moved its outputs closer to the postoperative photograph in identity-embedding terms. This is a study of control, not clinical accuracy: no surgeons rated the outputs, and each condition was generated once. Within that scope, keeping a surgical preview inside its intended region needs no access to the model; a mask and composite on the client enforce it across every editor tested, at low provider cost.
Chinese Translation
当请求商业图像编辑器预览美容手术时,它往往会改变面部的更多部分,而不仅仅是请求的区域:例如,鼻子的编辑可能还会平滑皮肤或改变光照。现有的方法要求访问模型的内部结构,以将编辑限制在一个区域,但公共编辑API并未公开这些信息。我们探讨仅从客户端能够实现多少控制。在一项初步基准测试中,六种商业编辑配置和一种基于掩码的修复模型在三个客户端控制级别下执行了类似面部提升的下颌颈部和鼻整形编辑:仅使用提示;将编辑区域从响应中剪切并通过基于地标的掩码粘贴回原始照片(掩码合成);以及在支持的情况下请求模型在掩码内进行修复。在210次尝试的编辑中,有196次可以评分。ArcFace余弦度量用于评估身份保留;CIELAB像素变化比率则衡量请求区域内的变化量,而非受保护的面部区域。在12个正面脸部中,掩码合成在区域度量上相较于仅使用提示的配对输出提高了中位数0.446(95%面部聚类自助区间0.421-0.457),同时对请求区域的变化量大致相当。不同的编辑器在编辑强度与身份保留方面存在差异,而我们测试的唯一修复模型并未优于简单的合成。在每个面部的输入到术后基线中,没有任何编辑器在身份嵌入方面使其输出更接近术后照片。这是一项关于控制的研究,而非临床准确性:没有外科医生对输出进行评分,每种条件仅生成一次。在这一范围内,保持手术预览在其预期区域内无需访问模型;在客户端使用掩码和合成可以在每个测试的编辑器中强制执行,且成本低廉。
cs.CV / 13 / 2608.02883
Test Time Adaptation Methods for Point Cloud Registration in Laparoscopic Surgery
腹腔镜手术中点云配准的测试时间适应方法
Abstract
3D point cloud registration in laparoscopic surgery estimates the transformation between an intraoperative organ reconstructed from video and its preoperative mesh. Because ground-truth transformations are unavailable for real data, supervised networks are trained on synthetic organ pairs. At test time, real reconstructions differ from synthetic data and are noisy, sparse, and occluded, which degrades correspondence estimation. Test-time adaptation (TTA) can reduce this domain shift, but existing methods mainly rely on logits, entropy, class prototypes, or cache memories unavailable in registration. Registration also involves paired inputs with an asymmetric shift that primarily affects the intraoperative cloud. We analyse and modify state-of-the-art TTA methods from three families to 3D registration: model, normalization, and input adaptation. We analyze four representative approaches based on auxiliary-task model updates, backpropagation-free token purging, feature alignment, and layer-normalization calibration. We modify them to handle asymmetric shifts between preoperative and intraoperative point clouds and replace classification-based entropy objectives. Using a correspondence-based model trained on clean synthetic source data, we evaluate adaptation to corrupted synthetic and real target data on P2P and P2ILReg. For synthetic targets, we apply eight corruptions, including uniform noise and global density reduction, at five severity levels. All methods improve registration on P2P, whereas normalization adaptation degrades performance on P2ILReg. Considering the computational overhead of backpropagation-based adaptation, input adaptation is the most promising option for laparoscopic surgery, providing low inference latency and consistent error reductions across datasets. Code: https://github.com/ninaa-git/survey_pc_registration_tta
Chinese Translation
腹腔镜手术中的三维点云配准估计了从视频重建的术中器官与其术前网格之间的变换。由于真实数据缺乏真实变换,监督网络在合成器官对上进行训练。在测试时,真实重建与合成数据存在差异,并且噪声、稀疏和遮挡现象会降低对应关系的估计。测试时间适应(TTA)可以减少这种领域转移,但现有方法主要依赖于logits、熵、类别原型或在配准中不可用的缓存记忆。配准还涉及成对输入,具有不对称的偏移,主要影响术中点云。我们分析并修改了三类最先进的TTA方法以适应三维配准:模型、归一化和输入适应。我们分析了四种基于辅助任务模型更新、无反向传播的标记清除、特征对齐和层归一化校准的代表性方法。我们对它们进行了修改,以处理术前和术中点云之间的不对称偏移,并替换基于分类的熵目标。使用在干净的合成源数据上训练的基于对应关系的模型,我们评估了对P2P和P2ILReg上受损合成和真实目标数据的适应。对于合成目标,我们在五个严重程度级别上应用了包括均匀噪声和全局密度降低在内的八种损坏。所有方法在P2P上的配准性能都有所提升,而归一化适应在P2ILReg上的性能有所下降。考虑到基于反向传播的适应的计算开销,输入适应是腹腔镜手术中最有前景的选择,提供了低推理延迟和在各数据集上一致的错误减少。代码链接:https://github.com/ninaa-git/survey_pc_registration_tta
cs.CV / 14 / 2608.02892
Modeling Scientific Experiment Scenes: Dataset and Model
科学实验场景建模:数据集与模型
Abstract
Scene Graph Generation (SGG) is fundamental to structured visual understanding, yet existing benchmarks focus mainly on daily life images and overlook scientific experiment scenes with specialized instruments, task-specific experimental semantics, and dense, fine-grained physical relations. These scenes are increasingly important for automated experimental analysis and smart education. To bridge this gap, we introduce PhysScene, the first SGG dataset for physical experiment scenes, providing densely annotated scene graphs and benchmarks under multiple supervision and protocol settings. PhysScene further exposes two key algorithmic challenges for SGG: pronounced long-tail relational predicate distributions and a substantial visual-textual semantic gap. To address these challenges, we propose the Cross-Modal Dual-Path Generator (CM-DPG), a model for robust open-vocabulary SGG. The model enhances object-level semantic representations through joint visual-textual encoding and improves relational reasoning using complementary visual and geometric cues. We also incorporate relation-aware pre-training, caption-derived pseudo-supervision, and adaptive weighting to support balanced learning across head and tail predicates. Extensive experiments on PhysScene and VG150 show that CM-DPG achieves competitive performance across multiple evaluation settings, with ablation studies validating the contribution of each component. The dataset and code are publicly available at https://github.com/ZMH-SDUST/CM-DPG.
Chinese Translation
场景图生成(Scene Graph Generation, SGG)是结构化视觉理解的基础,但现有基准主要集中在日常生活图像上,忽视了具有专业仪器、特定任务实验语义以及密集、细粒度物理关系的科学实验场景。这些场景在自动化实验分析和智能教育中变得越来越重要。为了解决这一问题,我们引入了PhysScene,这是第一个针对物理实验场景的SGG数据集,提供了密集标注的场景图和多种监督及协议设置下的基准。PhysScene进一步揭示了SGG的两个关键算法挑战:显著的长尾关系谓词分布和显著的视觉-文本语义差距。为了解决这些挑战,我们提出了跨模态双路径生成器(Cross-Modal Dual-Path Generator, CM-DPG),这是一个用于鲁棒开放词汇SGG的模型。该模型通过联合视觉-文本编码增强了对象级语义表示,并利用互补的视觉和几何线索改善了关系推理。我们还结合了关系感知的预训练、基于标题的伪监督和自适应加权,以支持头部和尾部谓词之间的平衡学习。在PhysScene和VG150上的大量实验表明,CM-DPG在多种评估设置下均表现出竞争力,消融研究验证了每个组件的贡献。数据集和代码已公开发布在 https://github.com/ZMH-SDUST/CM-DPG。
cs.CV / 15 / 2608.02953
RealWeather: Realistic and Scene-Faithful Weather Translation with Driving World Models
RealWeather:基于驾驶世界模型的真实场景天气翻译
Abstract
Realistic weather translation is valuable for developing and evaluating autonomous driving systems, yet collecting paired videos of the same scenes under different weather conditions at scale is impractical. Existing methods therefore rely on synthetic data, 3D weather editing, or geometry-conditioned generation, often compromising weather realism or scene fidelity. We propose RealWeather, a driving world model for both realistic and scene-faithful weather translation. Our key idea is to learn authentic weather dynamics directly from real-world videos. Specifically, RealWeather employs Progressive Realism Bootstrapping, an iterative data-refinement strategy. Assisted by an auxiliary Pseudo-Clear Generation pipeline, training initially starts with pseudo-style conditioning videos. As training proceeds, these inputs are progressively replaced with increasingly realistic videos generated by the model itself. This strategy bridges the pseudo-to-real domain gap, allowing the model to adapt seamlessly to real-world input distributions and naturally support bidirectional clear adverse translation. Furthermore, to strictly enforce structural integrity and suppress hallucinations, we introduce Scene-Fidelity RL Optimization, a reward-driven policy optimization strategy that explicitly penalizes alterations to safety-critical driving elements. Extensive experiments demonstrate that RealWeather significantly outperforms existing methods in visual realism and structural preservation, while enabling robust long-tail weather scenario generation and strong zero-shot out-of-distribution generalization.
Chinese Translation
真实的天气翻译对于开发和评估自动驾驶系统具有重要价值,但在不同天气条件下收集同一场景的配对视频在规模上是不切实际的。因此,现有方法通常依赖于合成数据、3D天气编辑或几何条件生成,往往妥协了天气的真实感或场景的保真度。我们提出了RealWeather,一种用于真实且场景保真的天气翻译的驾驶世界模型。我们的关键思想是直接从真实世界视频中学习真实的天气动态。具体而言,RealWeather采用渐进真实引导(Progressive Realism Bootstrapping),这是一种迭代的数据精炼策略。在辅助伪清晰生成(Pseudo-Clear Generation)管道的帮助下,训练最初从伪风格条件视频开始。随着训练的进行,这些输入逐渐被模型自身生成的越来越真实的视频所替代。这一策略弥合了伪域与真实域之间的差距,使模型能够无缝适应真实世界的输入分布,并自然支持双向的清晰与恶劣天气翻译。此外,为了严格维护结构完整性并抑制幻觉,我们引入了场景保真强化学习优化(Scene-Fidelity RL Optimization),这是一种奖励驱动的策略优化策略,明确惩罚对安全关键驾驶元素的修改。大量实验表明,RealWeather在视觉真实感和结构保留方面显著优于现有方法,同时能够生成强健的长尾天气场景,并具备强大的零样本分布外泛化能力。
cs.CV / 16 / 2608.02964
Material-Segmented Per-Pixel Emissivity Correction for Thermographic Anomaly Detection in Cultural Heritage Digital Twins
用于文化遗产数字双胞胎的材料分段逐像素发射率校正的热成像异常检测
Abstract
Quantitative longwave thermography of heritage surfaces is limited by the global-constant emissivity assumption in inverse-Planck temperature retrieval; on heterogeneous surfaces emissivity varies within one field of view, producing apparent-temperature artifacts that mimic and mask subsurface anomalies. We present a training-free pipeline that derives per-pixel emissivity by applying SAM 3.1 open-vocabulary segmentation to a colocated, co-calibrated RGB channel, mapping segments to a material-keyed LWIR emissivity table compiled from primary measurement literature, and propagating the field into a per-pixel inverse-Planck solve on raw radiometric data. Lacking any public dataset with raw radiometry, a temperature reference, and a colocated RGB camera, we evaluate on a physics-based synthetic benchmark and four real datasets. On the benchmark, under a palette spanning the low-emissivity exceptions, the correction cuts mean absolute error from 1.97 K to 0.91 K at 20 K contrast and, with an accurate table, beats the best fitted global constant on every layout; on a heritage-realistic emissivity distribution it does not. We contribute a quantified operating-regime map, and a measurement-backed finding that tempers the heritage claim: weathered outdoor heritage emissivities cluster near the conventional default, so the correction is small on typical surfaces and concentrated on genuine low-emissivity exceptions. We characterize the dominant failure mode, in which open-vocabulary segmentation matches appearance rather than material, and the contraindicated regime in which emissivity-defined anomalies are suppressed.
Chinese Translation
遗产表面的定量长波热成像受到逆普朗克温度反演中全球常数发射率假设的限制;在异质表面上,发射率在一个视场内变化,产生表观温度伪影,模仿并掩盖了地下异常。我们提出了一种无训练的流程,通过将 SAM 3.1 开放词汇分割应用于同位、同标定的 RGB 通道,推导逐像素发射率,将分段映射到从主要测量文献编制的材料关键长波红外(LWIR)发射率表,并将该场传播到原始辐射数据的逐像素逆普朗克求解中。由于缺乏任何具有原始辐射测量、温度参考和同位 RGB 相机的公共数据集,我们在基于物理的合成基准和四个真实数据集上进行了评估。在基准测试中,在一个涵盖低发射率例外的调色板下,校正将平均绝对误差从 1.97 K 降至 0.91 K(在 20 K 对比度下),并且在准确的表格下,在每种布局上都优于最佳拟合的全球常数;而在一个符合遗产的发射率分布中则没有。我们贡献了一个量化的操作模式图,以及一个基于测量的发现,缓和了遗产主张:风化的户外遗产发射率聚集在传统默认值附近,因此在典型表面上的校正较小,集中在真正的低发射率例外上。我们描述了主要的失败模式,其中开放词汇分割匹配外观而非材料,以及在发射率定义的异常被抑制的禁忌模式。
cs.CV / 17 / 2608.02980
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
Qwen-3D:一种用于空间理解的通用3D视觉-语言模型
Abstract
Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization and limited context windows. 3D geometry provides a natural compression mechanism for visual streams: depth and camera pose enable observations from multiple views and time steps to be fused into a persistent, world-aligned representation. While recent 3D LMMs leverage geometry-aware representations to improve spatial reasoning, they continue to lag behind specialist 3D perception systems on grounding and segmentation tasks. We argue that a key limitation is geometry-aware decoding: existing methods communicate 3D predictions through language tokens, proposal selection, or lightweight grounding queries, creating a bottleneck between language reasoning and dense geometric prediction. Building on these insights, we introduce Qwen-3D, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes. Qwen-3D augments visual tokens with 3D Rotary Positional Embeddings, allowing attention to operate directly in 3D scene space rather than across independent image frames and thereby facilitating scalable cross-view and temporal reasoning. To bridge language and geometry, Qwen-3D incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene representation, unifying referential grounding, instance segmentation, and visual question answering across both images and videos. Across a diverse set of benchmarks, Qwen-3D surpasses existing 3D LMMs and outperforms several large proprietary 2D models. Notably, Qwen-3D achieves these improvements while maintaining strong performance on standard 2D vision-language benchmarks by jointly training on 2D and 3D data.
Chinese Translation
大型多模态模型(LMMs)在图像和短视频上取得了显著成功,但由于帧中心的标记化和有限的上下文窗口,将其扩展到长视频仍然具有挑战性。3D几何提供了一种自然的视觉流压缩机制:深度和相机姿态使得从多个视角和时间步的观察能够融合成一个持久的、与世界对齐的表示。尽管最近的3D LMMs利用几何感知表示来改善空间推理,但在基础和分割任务上,它们仍然落后于专业的3D感知系统。我们认为,一个关键的限制是几何感知解码:现有方法通过语言标记、提案选择或轻量级基础查询来传达3D预测,从而在语言推理和密集几何预测之间形成瓶颈。基于这些见解,我们提出了Qwen-3D,这是一种几何感知的LMM,通过多视角几何线索在Qwen主干中压缩视觉信息,从而实现对静态场景的高效长时间视觉推理。Qwen-3D通过3D旋转位置嵌入增强视觉标记,使得注意力能够直接在3D场景空间中操作,而不是在独立的图像帧之间,从而促进可扩展的跨视角和时间推理。为了桥接语言和几何,Qwen-3D结合了一种基于查询的分割解码器,直接在基础的3D场景表示中对语言进行基础,从而统一了参考基础、实例分割和视觉问答在图像和视频中的应用。在一系列多样化的基准测试中,Qwen-3D超越了现有的3D LMMs,并在多个大型专有2D模型上表现优越。值得注意的是,Qwen-3D在保持在标准2D视觉-语言基准上强劲表现的同时,通过对2D和3D数据的联合训练实现了这些改进。
cs.CV / 18 / 2608.03008
V-FIND: Revealing the Intrinsic Forgery Knowledge Encoded in Video Forgery Detectors
V-FIND:揭示视频伪造检测器中编码的内在伪造知识
Abstract
As generated videos become increasingly realistic, reliable video forgery detection is increasingly important. Existing studies typically optimize and use video forgery detectors as black boxes, while the latent forgery-discriminative knowledge inside them remains largely unexplored. Instead of continuing to rely on resource-intensive full-model retraining to steadily improve detection performance, we ask whether video forgery detection can also be achieved by uncovering and activating sparse forensic knowledge within the detector. We find that forgery-discriminative knowledge is not uniformly distributed across the full representation space, but is concentrated in a sparse set of functionally specialized neurons. Based on this insight, we propose a video forgery-intrinsic neuron discovery (V-FIND) framework. V-FIND first localizes critical layers that exhibit pronounced discrepancies between real and forged videos, and then identifies latent anchor neurons that consistently carry forgery-discriminative signals, organizing them into a compact forensic subspace. With the original backbone frozen and only a lightweight linear classifier trained, this subspace still delivers strong detection performance across multiple external benchmarks for generated videos. Further neuron intervention experiments provide direct evidence for the functional specificity of the discovered neurons. Overall, these results suggest that video forgery detectors contain sparse, extractable, and reusable forgery-discriminative knowledge, offering a new perspective on understanding and exploiting their intrinsic forensic capability.
Chinese Translation
随着生成视频变得越来越逼真,可靠的视频伪造检测变得愈加重要。现有研究通常将视频伪造检测器作为黑箱进行优化和使用,而其中潜在的伪造区分知识仍然在很大程度上未被探索。我们提出的问题是,视频伪造检测是否也可以通过揭示和激活检测器内的稀疏法医知识来实现,而不是继续依赖资源密集型的全模型重训练来稳定提高检测性能。我们发现,伪造区分知识并不是均匀分布在整个表示空间中,而是集中在一组功能专门化的稀疏神经元中。基于这一见解,我们提出了视频伪造内在神经元发现(V-FIND)框架。V-FIND首先定位出在真实视频和伪造视频之间表现出明显差异的关键层,然后识别出始终携带伪造区分信号的潜在锚定神经元,并将其组织成一个紧凑的法医子空间。在冻结原始主干网络并仅训练一个轻量级线性分类器的情况下,该子空间在多个外部基准测试中仍能提供强大的检测性能。进一步的神经元干预实验为发现的神经元的功能特异性提供了直接证据。总体而言,这些结果表明,视频伪造检测器包含稀疏的、可提取的和可重用的伪造区分知识,为理解和利用其内在法医能力提供了新的视角。
cs.CV / 19 / 2608.03016
Clinically-Grounded Hierarchical Classification for Consistent Chest X-ray Interpretation
基于临床的分层分类方法用于一致的胸部X光解读
Abstract
Accurate chest X-ray interpretation is inherently hierarchical. Clinical decisions depend not only on what abnormality is present but where it is situated, requiring reasoning from broad anatomical systems down to specific pathological findings. Yet existing automated systems largely treat this as a flat classification problem, failing to capture inter-level dependencies or enforce coherence between coarse and fine predictions. We propose CHASE (Classification with Hierarchical Analysis and Structured Enforcement), a unified single-stage framework that mirrors radiologists' coarse-to-fine reasoning through a clinically driven three-level taxonomy of 9 anatomical regions, 17 sub-regions, and 28 pathological findings. CHASE jointly optimizes multi-level supervision, cross-level probability alignment, and a hierarchy-violation penalty within a shared Vision Transformer backbone. This ensures that fine-grained findings are anatomically supported by their coarser-level context rather than predicted in isolation. Experiments demonstrate that CHASE outperforms flat and hierarchical baselines across all levels while achieving superior probabilistic hierarchy consistency, with level-wise attention maps confirming anatomically grounded predictions. Code is available at: https://github.com/yejix-ai/CHASE.
Chinese Translation
准确的胸部X光解读本质上是分层的。临床决策不仅依赖于存在何种异常,还依赖于其位置,这需要从广泛的解剖系统推理到具体的病理发现。然而,现有的自动化系统在很大程度上将其视为一个平面分类问题,未能捕捉层级间的依赖关系或在粗略与细致预测之间强制一致性。我们提出了CHASE(Classification with Hierarchical Analysis and Structured Enforcement),这是一个统一的单阶段框架,反映了放射科医生从粗到细的推理过程,通过一个以临床为驱动的三层分类法,涵盖9个解剖区域、17个子区域和28个病理发现。CHASE在共享的视觉变换器(Vision Transformer)骨干网络内共同优化多层次监督、跨层概率对齐和层级违反惩罚。这确保了细粒度的发现得到其粗粒度上下文的解剖支持,而非孤立预测。实验表明,CHASE在所有层次上均优于平面和分层基线,同时实现了更优的概率层级一致性,层级注意力图确认了基于解剖的预测。代码可在以下链接获取:https://github.com/yejix-ai/CHASE。
cs.CV / 20 / 2608.03023
Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing
独立的 DINOv3 用于遥感中的无训练开放词汇语义分割
Abstract
Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods. Recently, the recent release of DINOv3 brings DINO.txt, which equips the standalone DINO backbone with image-text contrastive learning and thus opens up the possibility of open-vocabulary segmentation. We propose DinoSplat-OV, a training-free framework that adapts DINOv3 to remote sensing without fine-tuning or additional pretraining. Targeting the dense distribution, multi-scale nature, and large size of remote sensing imagery, we design two core modules. Its Text-aware Laplacian Propagation module de-noises patch-level predictions by combining textual semantic affinities with local visual similarity, improving regional consistency while preserving boundaries. Its Gaussian Splatting Upsampling module reconstructs pixel-level features through RGB-guided anisotropic aggregation and test-time optimization. A global-anchor sliding-window strategy further supports large-scale imagery. Experiments on UDD5, DOTA, and LoveDA demonstrate competitive or superior performance over existing training-free methods, effectively filling the gap of DINO-series models in training-free open-vocabulary segmentation and providing a viable new path for further advances in this direction.
Chinese Translation
遥感语义分割受到昂贵的像素级标注的限制,这促使了无训练开放词汇方法的发展。最近,DINOv3 的发布带来了 DINO.txt,它为独立的 DINO 主干提供了图像-文本对比学习,从而开启了开放词汇分割的可能性。我们提出了 DinoSplat-OV,这是一个无训练框架,将 DINOv3 适配于遥感任务,无需微调或额外的预训练。针对遥感图像的密集分布、多尺度特性和大尺寸,我们设计了两个核心模块。其文本感知拉普拉斯传播模块通过结合文本语义相似性与局部视觉相似性来去噪补丁级预测,提高区域一致性,同时保持边界。其高斯溅射上采样模块通过 RGB 引导的各向异性聚合和测试时优化重建像素级特征。全球锚点滑动窗口策略进一步支持大规模图像处理。在 UDD5、DOTA 和 LoveDA 上的实验表明,DinoSplat-OV 在性能上与现有的无训练方法具有竞争力或优越性,有效填补了 DINO 系列模型在无训练开放词汇分割中的空白,并为该方向的进一步发展提供了一条可行的新路径。
cs.CV / 21 / 2608.03046
CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation
CAPE-T2V:面向文本到视频生成中的双向条件对齐的标题锚定提示增强
Abstract
Text-to-video (T2V) diffusion transformers (DiTs) are trained with detailed video captions, whereas inference often relies on user prompts rewritten by a prompt enhancer (PE). Prior work has improved generation by optimizing the PE, the DiT, or both; some methods have also sought to narrow the training-inference mismatch through shared schemas. Yet even within a shared schema, inference-time PE outputs and DiT training captions may still differ in detail selection, information organization, descriptive granularity, and phrasing. We refer to this residual mismatch as the PE-Caption gap and introduce CAPE-T2V, a two-step Captioner-Anchored Prompt Enhancement framework toward two-sided conditioning alignment in T2V generation. First, CAPE-T2V constructs three types of PE training examples, pairing captioner-generated targets with concise source captions, detailed source captions, or pseudo user prompts derived from those targets. It then fine-tunes the PE to map each input to its paired target. Second, CAPE-T2V fine-tunes the DiT on video-derived captions rewritten by the Anchored PE; the same PE rewrites user prompts at inference. Relative to a baseline using the same caption schema, CAPE-T2V achieves higher aggregate scores on StoryEval, VBench-2.0, and T2V-CompBench across Wan2.2 and LTX-2.3. Further, CAPE-T2V exhibits a smaller PE-Caption gap than the baseline: its DiT fine-tuning captions are closer in distribution to inference-time PE outputs, as measured by squared maximum mean discrepancy in a fixed embedding space. Overall, these results support CAPE-T2V as an effective approach to mitigating the PE-Caption gap. The project is available at https://github.com/yizzz927/CAPE-T2V.
Chinese Translation
文本到视频(T2V)扩散变换器(DiTs)是通过详细的视频标题进行训练的,而推理通常依赖于由提示增强器(PE)重写的用户提示。先前的研究通过优化PE、DiT或两者的组合来改善生成效果;一些方法还试图通过共享模式来缩小训练与推理之间的不匹配。然而,即使在共享模式下,推理时PE输出与DiT训练标题在细节选择、信息组织、描述粒度和措辞上仍可能存在差异。我们将这种残余的不匹配称为PE-标题差距,并引入CAPE-T2V,一个两步的标题锚定提示增强框架,旨在实现T2V生成中的双向条件对齐。首先,CAPE-T2V构建三种类型的PE训练示例,将标题生成器生成的目标与简洁的源标题、详细的源标题或从这些目标派生的伪用户提示配对。然后,它对PE进行微调,以将每个输入映射到其配对目标。其次,CAPE-T2V在由锚定PE重写的视频衍生标题上微调DiT;同样的PE在推理时重写用户提示。相较于使用相同标题模式的基线,CAPE-T2V在Wan2.2和LTX-2.3的StoryEval、VBench-2.0和T2V-CompBench上获得了更高的综合得分。此外,CAPE-T2V的PE-标题差距小于基线:其DiT微调标题在分布上更接近推理时PE输出,使用固定嵌入空间中的平方最大均值差异进行测量。总体而言,这些结果支持CAPE-T2V作为一种有效的方法来减轻PE-标题差距。该项目可在https://github.com/yizzz927/CAPE-T2V获取。
cs.CV / 22 / 2608.03047
AIDE: Automated Instruction via Distilled Expertise for Reference-Free Motor Skill Coaching
AIDE:基于提炼专业知识的无参考运动技能自动指导
Abstract
Generating natural-language coaching feedback on motor skills can accelerate learning, yet expert coaches are scarce and expensive. Existing reference-based methods require expert demonstrations at both training and inference time, limiting practical deployment. We propose AIDE (Automated Instruction via Distilled Expertise), a framework that exploits expert references only during training and generates feedback from a learner's pose sequence alone at inference. A teacher model first learns to generate feedback from paired learner-expert poses via a frozen language model, producing separate learner tokens and difference tokens that encode the learner-expert difference. A student model then inherits the teacher's encoder and weight initialization, replacing the explicit expert comparison with an auxiliary module that produces complementary tokens from the learner's pose alone. On the ExpertAF dataset, AIDE outperforms reference-free baselines on most metrics and performs comparably to methods requiring expert demonstrations at both training and inference, with LLM-based evaluation supporting these findings.
Chinese Translation
生成自然语言的运动技能指导反馈可以加速学习,但专业教练稀缺且成本高昂。现有的基于参考的方法在训练和推理阶段均需要专家示范,这限制了其实际应用。我们提出了AIDE(基于提炼专业知识的自动指导)框架,该框架仅在训练期间利用专家参考,并在推理时仅根据学习者的姿态序列生成反馈。教师模型首先通过冻结的语言模型学习从配对的学习者-专家姿态生成反馈,产生分别编码学习者-专家差异的学习者标记和差异标记。然后,学生模型继承教师的编码器和权重初始化,用一个辅助模块替代显式的专家比较,仅从学习者的姿态生成互补标记。在ExpertAF数据集上,AIDE在大多数指标上优于无参考基线,并且在训练和推理阶段与需要专家示范的方法表现相当,基于LLM的评估支持了这些发现。
cs.CV / 23 / 2608.03055
PDD-RRG: Posterior Diagnostic Decision for Study-level Radiology Report Generation
PDD-RRG:用于研究级放射学报告生成的后验诊断决策
Abstract
Automatic radiology report generation (RRG) aims to simulate the workflow of radiologists, assisting them in clinical diagnosis. However, existing methods often fall short in utilizing all information relevant to the examination, as is typically done in clinical practice. Although some works attempt to incorporate multi-view images and historical data, these additional inputs may sometimes lead to avoidable diagnostic errors on the contrary. To address these challenges, we introduce a decision-making stage after report generation for the first time and propose a Posterior Diagnostic Decision framework (PDD-RRG) to integrate potentially conflicting diagnoses. Specifically, we create various subsets of input data and utilize an existing RRG model to generate reports from different perspectives. Then the Bayesian posterior probability and the learned thresholds for each clinical observation are calculated to obtain an aggregated diagnostic conclusion, which is subsequently used to refine the generated report. Experiments on MIMIC-CXR demonstrate that our proposed PDD-RRG can effectively enhance the clinical efficacy of existing RRG models without any retraining.
Chinese Translation
自动放射学报告生成(RRG)旨在模拟放射科医生的工作流程,辅助其进行临床诊断。然而,现有方法往往未能充分利用与检查相关的所有信息,这与临床实践中的做法相悖。尽管一些研究尝试结合多视角图像和历史数据,但这些额外输入有时反而可能导致可避免的诊断错误。为了解决这些挑战,我们首次在报告生成后引入了一个决策阶段,并提出了一种后验诊断决策框架(PDD-RRG),以整合潜在冲突的诊断。具体而言,我们创建了各种输入数据的子集,并利用现有的RRG模型从不同视角生成报告。然后,计算每个临床观察的贝叶斯后验概率和学习到的阈值,以获得聚合的诊断结论,随后用于优化生成的报告。在MIMIC-CXR上的实验表明,我们提出的PDD-RRG能够有效提升现有RRG模型的临床效能,而无需任何重新训练。
cs.CV / 24 / 2608.03057
TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models
TASQ:用于扩散模型的时间自适应比特稀疏量化
Abstract
Static quantization assigns one weight precision to every denoising step. To preserve quality, that precision must accommodate the most quantization-sensitive step, even though many other steps can tolerate fewer bits. The resulting model may satisfy its memory budget, but it repeatedly pays worst-case arithmetic throughout the denoising trajectory. We introduce Temporal-Adaptive Bit Sparsification Quantization (TASQ) to separate these two costs. TASQ stores one shared maximum-precision weight buffer and learns a Temporal-Spatial LSB Mask that selects a lower effective precision for each layer and denoising stage by truncating least-significant bits. Storage therefore remains fixed by the worst case, while BitOPs decrease at less sensitive stages without per-stage weight copies or runtime search. A Temporal-Precision Engine maps the learned schedule to bit-serial execution, where cycles scale with effective precision and switching precision has no measured cycle overhead. On PixArt-Sigma, SANA-1.6B, and SDXL-Turbo, TASQ achieves quality comparable to static quantization with less computation. Together with the Temporal-Precision Engine, it reduces execution cycles by 25 to 50 percent over static quantization and by 6.1 to 7.5x over a naive static 8-bit bit-serial execution. Code is available at https://github.com/seokho-han/tasq.
Chinese Translation
静态量化为每个去噪步骤分配一个权重精度。为了保持质量,该精度必须适应最敏感于量化的步骤,尽管许多其他步骤可以容忍更少的比特。因此,所得到的模型可能满足其内存预算,但在整个去噪过程中却反复支付最坏情况下的算术开销。我们提出了时间自适应比特稀疏量化(Temporal-Adaptive Bit Sparsification Quantization,TASQ),以分离这两种成本。TASQ存储一个共享的最大精度权重缓冲区,并学习一个时间-空间最低有效位掩码(Temporal-Spatial LSB Mask),通过截断最低有效位为每一层和去噪阶段选择较低的有效精度。因此,存储仍然固定在最坏情况下,而在不需要每个阶段权重副本或运行时搜索的情况下,较不敏感阶段的比特操作(BitOPs)减少。时间精度引擎(Temporal-Precision Engine)将学习到的调度映射到比特串行执行中,其中周期与有效精度成比例,切换精度没有测量的周期开销。在PixArt-Sigma、SANA-1.6B和SDXL-Turbo上,TASQ实现了与静态量化相当的质量,同时计算量更少。结合时间精度引擎,它在静态量化的基础上将执行周期减少了25%到50%,并在简单的静态8位比特串行执行中减少了6.1到7.5倍。代码可在https://github.com/seokho-han/tasq获取。
cs.CV / 25 / 2608.03059
RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing
RIDGE:基于内部动态引导的再噪声图像编辑
Abstract
Inversion-free flow-based image editing avoids latent inversion, but still requires a target-side state at every editing step. The widely used equal-displacement construction keeps the displacement between the noisy source state and the target-side state unchanged across noise levels. This is inconsistent with noising, under which the displacement between two clean states noised with the same noise level and noise sample should contract as the noise level increases. Thus, it can lead to overly aggressive updates at high noise levels. We introduce RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing, an inversion-free and training-free method that maintains the edited state as an evolving approximation to the unavailable clean target state. RIDGE re-noises this approximation using the same noise level and noise sample as the clean source state, allowing their noisy displacement to decrease naturally with increasing noise. Since the edited state initially contains limited target semantics, RIDGE further applies internal dynamic guidance during the early high-noise steps. A clean target state prediction guides the provisional edited state through a soft dynamic mask derived internally from the model, focusing guidance on regions that require modification without external segmentation or detection models. Experiments on two benchmarks using two backbones, SD3 Medium and FLUX.1-dev, show that RIDGE offers a favorable aggregate trade-off among source preservation, target alignment, and perceptual quality.
Chinese Translation
无反演流基图像编辑避免了潜在反演,但在每个编辑步骤中仍需目标侧状态。广泛使用的等位移构造保持了噪声源状态与目标侧状态之间的位移在不同噪声水平下不变。这与噪声化不一致,在噪声水平增加时,使用相同噪声水平和噪声样本噪声化的两个干净状态之间的位移应收缩。因此,这可能导致在高噪声水平下过于激进的更新。我们提出了RIDGE:基于内部动态引导的再噪声图像编辑,这是一种无反演且无需训练的方法,保持编辑状态作为不可用的干净目标状态的演变近似。RIDGE使用与干净源状态相同的噪声水平和噪声样本对该近似进行再噪声处理,使其噪声位移随着噪声的增加自然减少。由于编辑状态最初包含有限的目标语义,RIDGE在早期高噪声步骤中进一步应用内部动态引导。干净目标状态预测通过模型内部派生的软动态掩膜引导临时编辑状态,集中引导需要修改的区域,而无需外部分割或检测模型。在使用两个主干网络SD3 Medium和FLUX.1-dev的两个基准上的实验表明,RIDGE在源保持、目标对齐和感知质量之间提供了良好的综合权衡。
cs.CV / 26 / 2608.03064
Global Graph-Validated Optimization for VLM-based 3D Indoor Scene Generation
基于全球图验证的VLM优化用于3D室内场景生成
Abstract
We study open-vocabulary 3D indoor layout generation, which synthesizes diverse and physically plausible scenes from unlabeled 3D assets using free-form language instructions. Recent methods leverage large language models (LLMs) and vision-language models (VLMs) to generate structured scenes from text. However, most model inter-asset relations implicitly or rely on local pairwise constraints and local optimization. These formulations are poorly aligned with the global, highly non-convex layout space, often yielding locally plausible yet globally inconsistent or physically infeasible scenes. We address this problem with a graph-based intermediate representation that separates semantic coherence from physical feasibility, together with a hybrid search-and-refinement strategy. First, Global Semantic Verification (GSV) represents scenes as structured graphs and enforces semantic constraints through rule-based verification. This explicit validation removes contradictory configurations and produces a globally consistent semantic scaffold. Second, Global Physical Feasibility Search (GPFS) combines evolutionary search for global exploration with gradient-based refinement for local exploitation. It reduces dependence on VLM-proposed initialization and improves robustness in non-convex and discontinuous feasible spaces. Together, GSV and GPFS move layout generation beyond local relational modeling and initialization-sensitive optimization toward globally consistent reasoning and search. Experiments show that our method achieves state-of-the-art performance in open-vocabulary 3D indoor layout generation, improving both semantic consistency and physical plausibility.
Chinese Translation
我们研究开放词汇的3D室内布局生成,该方法利用自由形式的语言指令,从未标记的3D资产合成多样且物理上合理的场景。近期的方法利用大型语言模型(LLMs)和视觉-语言模型(VLMs)从文本生成结构化场景。然而,大多数模型之间的资产关系隐含或依赖于局部成对约束和局部优化。这些公式与全球高度非凸的布局空间不匹配,常常导致局部合理但全球不一致或物理上不可行的场景。我们通过一种基于图的中间表示来解决这个问题,该表示将语义一致性与物理可行性分离,并结合混合搜索与精炼策略。首先,全球语义验证(GSV)将场景表示为结构化图,并通过基于规则的验证强制执行语义约束。这种显式验证消除了矛盾配置,并生成一个全球一致的语义框架。其次,全球物理可行性搜索(GPFS)结合了用于全球探索的进化搜索与用于局部开发的基于梯度的精炼。它减少了对VLM提出的初始化的依赖,并提高了在非凸和不连续可行空间中的鲁棒性。GSV和GPFS共同推动布局生成超越局部关系建模和对初始化敏感的优化,朝着全球一致的推理和搜索迈进。实验表明,我们的方法在开放词汇的3D室内布局生成中达到了最先进的性能,改善了语义一致性和物理合理性。
cs.CV / 27 / 2608.03078
LDU-Bench: Multimodal LLM Evaluation for Lithography Defect Understanding under Layout-Varying Circuit Backgrounds
LDU-Bench:在布局变化电路背景下的光刻缺陷理解的多模态大语言模型评估
Abstract
Multimodal large language models have demonstrated strong defect recognition capability in industrial anomaly detection. However, in lithography review, merely determining whether an image contains a defect is insufficient for engineering inspection; models must also understand defect morphology, spatial location, and the potential causes supported by visible evidence. To this end, this paper proposes LDU-Bench, a multi-task multimodal benchmark for lithography defect understanding. Constructed from real lithography and integrated-circuit review images, LDU-Bench decomposes the review workflow into four independent tasks: defect triage, morphology recognition, coarse localization, and image-conditioned cause analysis. It systematically evaluates models using task-level metrics, diagnostic readouts, and the Lithography Closure Score (LCS). Experimental results show that although existing MLLMs can perform defect triage relatively reliably, this ability does not stably transfer to downstream review stages. Morphology alignment, effective localization, and evidence-to-cause mapping remain the major bottlenecks. Further diagnostics indicate that this capability break is not a fluctuation of a single metric, but reflects insufficient structured understanding across semantic levels. Overall, LDU-Bench provides a quantifiable and diagnostic unified platform for evaluating the usability, failure points, and capability boundaries of industrial MLLMs in lithography review chains.
Chinese Translation
多模态大语言模型在工业异常检测中展示了强大的缺陷识别能力。然而,在光刻审核中,仅仅判断图像是否包含缺陷对于工程检查来说是不够的;模型还必须理解缺陷的形态、空间位置以及由可见证据支持的潜在原因。为此,本文提出了LDU-Bench,一个用于光刻缺陷理解的多任务多模态基准。LDU-Bench基于真实的光刻和集成电路审核图像构建,将审核工作流程分解为四个独立任务:缺陷分流、形态识别、粗略定位和基于图像的原因分析。它通过任务级指标、诊断读数和光刻闭合评分(Lithography Closure Score, LCS)系统地评估模型。实验结果表明,尽管现有的多模态大语言模型(MLLMs)能够相对可靠地进行缺陷分流,但这种能力并未稳定地转移到下游审核阶段。形态对齐、有效定位和证据到原因的映射仍然是主要瓶颈。进一步的诊断表明,这种能力的断裂并不是单一指标的波动,而是反映了在语义层面上结构化理解的不足。总体而言,LDU-Bench提供了一个可量化和诊断的统一平台,用于评估工业多模态大语言模型在光刻审核链中的可用性、故障点和能力边界。
cs.CV / 28 / 2608.03079
CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation
CorePath:一种专门针对乳腺病理的基础模型,用于核心针活检诊断和风险控制报告生成
Abstract
Breast core needle biopsy (CNB) is central to breast cancer diagnosis yet remains challenging because limited tissue sampling, lesion heterogeneity, and subtle morphologic overlap can obscure subtype distinctions. We developed CorePath, a breast-specialized multimodal pathology foundation model fine-tuned from PRISM using 7901 paired CNB whole-slide images and diagnostic reports from two centers. Evaluated across six CNB cohorts and two public breast pathology benchmarks without task-specific retraining, CorePath consistently outperformed PRISM across cancer detection, invasion assessment, and histological subtyping. It achieved weighted area under the receiver operating characteristic curves (AUCs) of 0.9526-0.9735 for five-class CNB histological subtyping across private centers. On public benchmarks, CorePath outperformed leading pathology foundation models, achieving the highest weighted AUCs of 0.7780 for BCNB invasive carcinoma subtyping, 0.8178 for BRACS lesion stratification, and 0.8252 for BRACS fine-grained classification. In report generation, CorePath reduced the overall non-breast hallucinations from 30.1% to 2.8%, demonstrating improved domain fidelity after breast-specific adaptation. CorePath-CRG further combined conformal subtype-confidence gating with Learn-Then-Test risk control to enable selective report release, subtype-level fallback, and deferral. CorePath-CRG achieved zero non-breast hallucinations among released outputs and showed the strongest overall performance in pathologist-validated LLM-based Evaluation Scores and quantitative report-generation metrics across most centers. These results demonstrate that domain-specialized foundation models with statistical risk control offer a promising approach for accurate breast CNB diagnosis and reliable report generation.
Chinese Translation
乳腺核心针活检(CNB)是乳腺癌诊断的核心,但由于组织取样有限、病变异质性和细微形态重叠,导致亚型区分仍然具有挑战性。我们开发了CorePath,这是一种专门针对乳腺的多模态病理基础模型,基于7901对来自两个中心的CNB全切片图像和诊断报告对PRISM进行了微调。在六个CNB队列和两个公共乳腺病理基准上进行评估,CorePath在癌症检测、侵袭性评估和组织学亚型分类方面始终优于PRISM。它在私有中心的五类CNB组织学亚型分类中实现了加权受试者工作特征曲线(AUC)为0.9526-0.9735。在公共基准上,CorePath超越了领先的病理基础模型,在BCNB侵袭性癌症亚型分类中实现了最高的加权AUC为0.7780,在BRACS病变分层中为0.8178,在BRACS细粒度分类中为0.8252。在报告生成方面,CorePath将整体非乳腺幻觉从30.1%降低至2.8%,显示出在乳腺特定适应后改进的领域保真度。CorePath-CRG进一步结合了符合亚型置信度的门控与“学习后测试”(Learn-Then-Test)风险控制,以实现选择性报告发布、亚型级别回退和延迟。CorePath-CRG在发布的输出中实现了零非乳腺幻觉,并在大多数中心的病理学家验证的LLM基础评估分数和定量报告生成指标中表现出最强的整体性能。这些结果表明,具有统计风险控制的领域专门基础模型为准确的乳腺CNB诊断和可靠的报告生成提供了一种有前景的方法。
cs.CV / 29 / 2608.03082
DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers
DiverseDiT++:量化、分析与促进扩散变换器中的表示多样性
Abstract
Recent advances in Diffusion Transformers (DiTs) have enabled remarkable progress in visual synthesis, benefiting from their superior scalability. To facilitate DiTs' capability of capturing meaningful internal representations, recent works such as REPA incorporate external pretrained encoders for representation alignment. However, the underlying mechanisms governing representation learning within DiTs remain poorly understood in the community. To this end, this paper first presents a systematic analysis of the representation dynamics of DiTs via quantifying the diversity of block-wise representations. Specifically, we introduce a novel metric, termed the Weighted Diversity Score (WDS), to measure the representational discrepancies across different blocks. Through extensive investigations on the evolution and influence of internal representations under various settings, we reveal that representation diversity across blocks is a critical factor for effective representation learning in DiTs. More importantly, WDS exhibits a strong correlation with synthesis quality across diverse settings, model scales, and training stages (Pearson's $r=-0.869$ with $\log(\text{FID})$), suggesting its potential as an indicator to reflect model performance and a principled guide for model optimization. Based on this key finding, we propose DiverseDiT++, a novel framework that explicitly promotes diverse representation learning. Concretely, our method incorporates long residual connections to diversify input representations across blocks and a representation diversity loss to encourage blocks to learn distinct features. Extensive experiments on ImageNet $256\times256$ and $512\times512$ demonstrate that our DiverseDiT++ yields consistent performance gains and convergence acceleration when applied to different backbones with various sizes,...
Chinese Translation
最近在扩散变换器(Diffusion Transformers, DiTs)方面的进展使得视觉合成取得了显著进展,得益于其卓越的可扩展性。为了促进DiTs捕捉有意义的内部表示的能力,近期的研究如REPA引入了外部预训练编码器以实现表示对齐。然而,DiTs内部表示学习的基本机制在学术界仍然不够清晰。为此,本文首先通过量化块级表示的多样性,对DiTs的表示动态进行了系统分析。具体而言,我们引入了一种新颖的度量标准,称为加权多样性评分(Weighted Diversity Score, WDS),用于测量不同块之间的表示差异。通过对内部表示在不同设置下的演变和影响进行广泛研究,我们揭示了块间表示多样性是DiTs中有效表示学习的关键因素。更重要的是,WDS与不同设置、模型规模和训练阶段下的合成质量之间表现出强相关性(Pearson's $r=-0.869$与$ ext{log(FID)}$),这表明其作为反映模型性能的指标和模型优化的原则性指导的潜力。基于这一关键发现,我们提出了DiverseDiT++,一种明确促进多样性表示学习的新框架。具体而言,我们的方法结合了长残差连接,以在块之间多样化输入表示,并引入表示多样性损失,以鼓励各块学习不同特征。在ImageNet $256 imes256$和$512 imes512$上的广泛实验表明,当应用于不同规模的骨干网络时,我们的DiverseDiT++在性能提升和收敛加速方面均表现出一致的优势...
cs.CV / 30 / 2608.03083
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models
GSTEP:基于全球时空密度驱动的视觉标记剪枝以提高视频大型语言模型的效率
Abstract
Video large language models (VideoLLMs) achieve strong video understanding performance, but their inference remains expensive due to the large number of redundant spatio-temporal visual tokens in long videos. Existing token pruning methods alleviate this cost by reducing redundant tokens, yet most of them rely on segment-level local pruning, where videos are partitioned into isolated segments and tokens are selected independently within each segment. Such designs may under-preserve short but semantically dense segments and discard tokens that appear non-salient locally but remain critical from a global perspective. To address this issue, we propose GSTEP (Global Spatio-Temporal Density Pruning), a plug-and-play pruning framework that models video as a continuous spatio-temporal information flow. GSTEP constructs a token-level spatio-temporal density by combining a continuous temporal density, obtained from a smoothed centered frame-level change signal, with intra-frame spatial density, and then performs global token sampling by jointly balancing information density and coverage. Extensive experiments on multiple VideoLLMs and public benchmarks demonstrate that GSTEP consistently achieves strong accuracy-efficiency trade-offs and generalizes well across model architectures and evaluation settings. On LLaVA-OneVision-7B, GSTEP prunes 75% of visual tokens, preserves up to 100.2% of the original average performance across benchmarks, and achieves a 1.17 end-to-end speedup.
Chinese Translation
视频大型语言模型(VideoLLMs)在视频理解方面表现出色,但由于长视频中存在大量冗余的时空视觉标记,其推理成本仍然很高。现有的标记剪枝方法通过减少冗余标记来缓解这一成本,但大多数方法依赖于基于段的局部剪枝,即将视频划分为孤立的段落,并在每个段落内独立选择标记。这种设计可能会在保留短而语义密集的段落方面不足,并丢弃那些在局部看似不显著但从全局角度来看仍然至关重要的标记。为了解决这个问题,我们提出了GSTEP(全球时空密度剪枝),一个即插即用的剪枝框架,将视频建模为连续的时空信息流。GSTEP通过将从平滑的中心帧级变化信号获得的连续时间密度与帧内空间密度相结合,构建标记级时空密度,然后通过共同平衡信息密度和覆盖率进行全局标记采样。在多个VideoLLMs和公共基准上的大量实验表明,GSTEP始终实现了强大的准确性与效率的权衡,并在模型架构和评估设置上具有良好的泛化能力。在LLaVA-OneVision-7B上,GSTEP剪枝了75%的视觉标记,保留了在基准测试中原始平均性能的高达100.2%,并实现了1.17倍的端到端加速。
cs.CV / 31 / 2608.03084
SUV: Future Scene Understanding as Video Generation for End-to-End Driving
SUV:将未来场景理解视为端到端驾驶的视频生成
Abstract
End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output formats, with limited scalability. Can video generation instead provide a shared predictor? We introduce SUV, a unified end-to-end driving framework that casts future Scene Understanding as Video generation using a pretrained video foundation model. SUV models future appearance, semantics, relative depth, and instance-level dynamics as video streams with a shared video expert, without stream-specific visual prediction heads. Through joint video-action attention, the action expert attends to the latent representations of all future streams and generates the ego trajectory. Experiments show that SUV directly predicts all four future streams, while controlled ablations show that structured future supervision and direct future-stream access yield higher trajectory planning scores. With only a single front camera and no candidate-trajectory selection, SUV outperforms a broad set of recent state-of-the-art methods on both NAVSIM-v2 splits, achieving 91.0 EPDMS on navtest and 36.9 on navhard. On the long-tail WOD-E2E benchmark, SUV achieves a competitive RFS of 7.94.
Chinese Translation
端到端驾驶需要对未来场景进行连贯的理解,然而现有方法使用特定任务的头部和输出格式来建模这些场景,具有有限的可扩展性。那么,视频生成是否可以提供一个共享的预测器呢?我们提出了SUV,一个统一的端到端驾驶框架,将未来场景理解视为使用预训练视频基础模型的视频生成。SUV将未来的外观、语义、相对深度和实例级动态建模为视频流,并使用共享的视频专家,而不需要特定于流的视觉预测头。通过联合视频-动作注意力,动作专家关注所有未来流的潜在表示,并生成自我轨迹。实验表明,SUV可以直接预测所有四个未来流,而控制性消融实验显示,结构化的未来监督和直接的未来流访问可以提高轨迹规划得分。在仅使用单个前置摄像头且不进行候选轨迹选择的情况下,SUV在NAVSIM-v2的两个分割上超越了一系列最新的先进方法,在navtest上达到91.0 EPDMS,在navhard上达到36.9。在长尾WOD-E2E基准测试中,SUV实现了竞争力的RFS得分7.94。
cs.CV / 32 / 2608.03100
Channel-wise Dynamic Knowledge Distillation via Adaptive Sample Generation for Action Recognition
基于自适应样本生成的通道动态知识蒸馏用于动作识别
Abstract
Knowledge Distillation (KD) offers a promising yet underexplored path for compressing large action recognition models. However, existing KD methods suffer from two key limitations: 1) reliance on fixed input samples leads to suboptimal feature alignment between the frozen teacher (larger model) and the learnable student (smaller model), and 2) applying a uniform distillation strength for all channels fails to account for their varying importance in capturing distinct knowledge (e.g., motion tempo or magnitude) across training epochs. This motivates us to develop an Adaptive Sample-aware Channel-wise Dynamic (ASCD) KD approach, which operates in two stages. First, we use an adaptive sample generation module to create updated samples by incorporating semantics from sample gradients, which are derived by minimizing a feature loss weighted by channel centroid frequency differences at each layer. Meanwhile, crucial motion-related details are preserved by applying a Gaussian mask to frequency features. Second, we employ a channel-wise dynamic distillation module to train student on these generated samples, guided by sample gradients and feature frequencies. For efficiency, samples are updated periodically rather than per epoch. Extensive experiments on three video benchmarks (UCF101, Kinetics-400, Something-Something-v2) and two image datasets (CIFAR-100, ImageNet) demonstrate the state-of-the-art performance of our method. Code is available at https://github.com/mlvccn/ASCD_KD_Action.
Chinese Translation
知识蒸馏(Knowledge Distillation, KD)为压缩大型动作识别模型提供了一条有前景但尚未充分探索的路径。然而,现有的KD方法存在两个主要限制:1)依赖固定输入样本导致冻结教师模型(较大模型)与可学习学生模型(较小模型)之间的特征对齐不佳;2)对所有通道应用统一的蒸馏强度未能考虑它们在不同训练周期中捕捉不同知识(例如,运动节奏或幅度)的重要性差异。这促使我们开发了一种自适应样本感知的通道动态(Adaptive Sample-aware Channel-wise Dynamic, ASCD)KD方法,该方法分为两个阶段。首先,我们使用自适应样本生成模块,通过结合样本梯度的语义生成更新样本,样本梯度是通过最小化每层通道质心频率差异加权的特征损失得出的。同时,通过对频率特征应用高斯掩模,保留了重要的运动相关细节。其次,我们采用通道动态蒸馏模块在这些生成的样本上训练学生模型,指导依据样本梯度和特征频率。为了提高效率,样本更新是定期进行的,而不是每个周期更新一次。在三个视频基准(UCF101、Kinetics-400、Something-Something-v2)和两个图像数据集(CIFAR-100、ImageNet)上的大量实验表明,我们的方法达到了最先进的性能。代码可在 https://github.com/mlvccn/ASCD_KD_Action 获取。
cs.CV / 33 / 2608.03101
Double Down on Defense: Strengthening Deep Perceptual Hashes against Evasion Attacks without Retraining
加强防御:在不重新训练的情况下增强深度感知哈希对抗规避攻击的能力
Abstract
Near-duplicate image matching is crucial for trust and safety, provenance verification, copyright enforcement, and large-scale visual search. Modern platforms increasingly rely on deep perceptual hashes, which map visually similar images to nearby representations despite common image transformations. However, adversarial perturbations can cause near-duplicates to evade matching. We present DualShield, a plug-in defense that improves the robustness of existing deep perceptual hashes without retraining or modifying their underlying models. DualShield combines matching-time randomized smoothing, which aggregates decisions over perturbed reference-query pairs, with publication-time hardening, which adds an optimized imperceptible perturbation to each reference image before publication. Together, these mechanisms provide certified and empirical robustness. DualShield achieves a certified $\ell_2$ radius of approximately 0.3, guaranteeing that query perturbations within this radius cannot evade matching. We further evaluate it against adaptive white-box, black-box, and image-transformation attacks. Across eight deep perceptual hashes and three datasets, DualShield substantially reduces attack success rates while preserving low collision rates. These results show that deep perceptual hashes can be strengthened without costly retraining by improving the matching procedure and hardening reference images before publication.
Chinese Translation
近重复图像匹配对于信任与安全、来源验证、版权执行以及大规模视觉搜索至关重要。现代平台越来越依赖深度感知哈希,这种方法能够将视觉上相似的图像映射到相近的表示,即使在常见的图像变换下也能保持有效。然而,对抗性扰动可能导致近重复图像无法成功匹配。我们提出了DualShield,这是一种插件防御机制,能够在不重新训练或修改其基础模型的情况下提高现有深度感知哈希的鲁棒性。DualShield结合了匹配时的随机平滑技术,该技术对扰动的参考-查询对的决策进行聚合,以及发布时的强化技术,该技术在发布前为每个参考图像添加优化的不可察觉扰动。这些机制共同提供了认证和经验上的鲁棒性。DualShield实现了约0.3的认证$ ext{l}_2$半径,确保在此半径内的查询扰动无法规避匹配。我们进一步评估了其对自适应白盒、黑盒和图像变换攻击的抵抗能力。在八种深度感知哈希和三个数据集上,DualShield显著降低了攻击成功率,同时保持了低碰撞率。这些结果表明,通过改善匹配过程和在发布前强化参考图像,可以在不进行昂贵的重新训练的情况下增强深度感知哈希的能力。
cs.CV / 34 / 2608.03106
FaithIR: Rethinking Infrared Image Super-Resolution from Perceptual Sharpness to Task Relevant Fidelity
FaithIR:从感知清晰度到任务相关保真度的红外图像超分辨率再思考
Abstract
Infrared image super-resolution (IISR) is important for downstream tasks such as object detection and semantic segmentation. Existing IISR methods often produce artificial textures, over-sharpened edges, and spurious high-frequency details that distort authentic thermal structures and semantic information. To address this issue, we propose FaithIR, a faithful infrared super-resolution framework for reliable machine perception. FaithIR consists of a patch-level conditioning branch that captures global thermal and structural information and a pixel-level restoration branch that performs dense local reconstruction under structural guidance. The entire restoration process is performed directly in the pixel domain to preserve infrared-specific structures and task-relevant information. Extensive experiments on FLIR-IISR, M3FD, and FMB demonstrate strong reconstruction fidelity, cross-dataset generalization, and superior performance in object detection and semantic segmentation. These results show that demonstrate that preserving faithful infrared structure preservations is more important for reliable machine perception than merely pursuing perceptual sharpness alone.
Chinese Translation
红外图像超分辨率(IISR)对于物体检测和语义分割等下游任务至关重要。现有的IISR方法往往会产生人工纹理、过度锐化的边缘以及虚假的高频细节,这些都会扭曲真实的热结构和语义信息。为了解决这个问题,我们提出了FaithIR,一个用于可靠机器感知的忠实红外超分辨率框架。FaithIR由一个捕捉全局热和结构信息的块级条件分支和一个在结构指导下进行密集局部重建的像素级恢复分支组成。整个恢复过程直接在像素域内进行,以保持红外特有的结构和任务相关信息。在FLIR-IISR、M3FD和FMB上的大量实验表明,FaithIR具有强大的重建保真度、跨数据集的泛化能力,并在物体检测和语义分割中表现优越。这些结果表明,保持忠实的红外结构比单纯追求感知清晰度对可靠的机器感知更为重要。
cs.CV / 35 / 2608.03107
A Unified Resolution-Conditioned Framework for Orthogonal Line-Scanning Image Fusion
一种统一的基于分辨率条件的正交线扫描图像融合框架
Abstract
Laser line-scanning microscopy enables fast volumetric imaging but produces anisotropic lateral resolution. Orthogonal line scans provide complementary directional information that can recover near-isotropic resolution, yet existing deep-learning methods require a separate model for each optical configuration. We present a unified, resolution-conditioned fusion framework based on Rank Enhanced Linear Attention (RELA). Feature-wise Linear Modulation (FiLM) conditions the network continuously on the resolving-power ratio, enabling one model to adapt across slit widths. We further introduce Adaptive RELA, which replaces fixed-kernel rank enhancement with ratio-conditioned multi-scale depthwise convolutions and uses a learnable attention temperature to adjust selectivity with degradation severity. Training data spanning multiple slit configurations are generated using a physics-grounded separable point-spread-function model verified against measured optical data at 48.3 dB accuracy. The resulting model achieves 34-40 dB PSNR across configurations, whereas unconditioned multi-slit training collapses to 24.3 dB and per-slit specialists lose 4-9 dB outside their training setting. It also generalizes smoothly to unseen intermediate configurations without interpolation artifacts. Ablations show that FiLM resolves configuration ambiguity, global linear attention captures long-range directional correspondences, and adaptive temperature yields an additional 2 dB in the challenging near-isotropic regime, where complementary signals are weak.
Chinese Translation
激光线扫描显微镜能够实现快速的体积成像,但产生各向异性的横向分辨率。正交线扫描提供了互补的方向信息,可以恢复近等各向同性的分辨率,然而现有的深度学习方法需要为每种光学配置单独训练模型。我们提出了一种基于秩增强线性注意力(Rank Enhanced Linear Attention, RELA)的统一分辨率条件融合框架。特征线性调制(Feature-wise Linear Modulation, FiLM)使网络能够持续根据分辨率比进行条件调整,从而使一个模型能够适应不同的狭缝宽度。我们进一步引入自适应RELA,它用比率条件的多尺度深度卷积替代固定核的秩增强,并使用可学习的注意力温度来根据降解严重程度调整选择性。通过使用基于物理的可分离点扩散函数模型生成跨多个狭缝配置的训练数据,并与在48.3 dB精度下测得的光学数据进行验证。所得到的模型在不同配置下实现了34-40 dB的峰值信噪比(PSNR),而无条件的多狭缝训练则降至24.3 dB,且每个狭缝的专家在其训练设置之外损失了4-9 dB。该模型还能够平滑地推广到未见过的中间配置,而不会产生插值伪影。消融实验表明,FiLM解决了配置模糊问题,全球线性注意力捕捉了长距离方向对应关系,自适应温度在具有挑战性的近等各向同性区域提供了额外的2 dB增益,在该区域互补信号较弱。
cs.CV / 36 / 2608.03109
Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis
多模态植物根系表型分析:整合3D骨架提取与语言分析
Abstract
Plant root phenotyping is fundamental to understanding below-ground structures, optimizing crop management, and improving agricultural sustainability. This paper presents a multimodal robotic AI framework that integrates 3D skeleton extraction with language-guided reasoning for interpretable and data-efficient root analysis. We develop an unsupervised skeleton extraction network based on Weighted Laplacian Contraction (W-LBC) to generate high-fidelity structural representations from dense point clouds captured by robotic 3D sensing platforms. Quantitative morphological descriptors, including root count, length, branching angle, and density, are computed from the reconstructed skeleton graph to capture geometric and topological characteristics. Building on these features, we introduce an Evidence-First language modeling framework that fine-tunes GPT as an interactive analytical chatbot using automatically generated instruction--response pairs. Each training sample provides measurable evidence before natural-language reasoning, enabling the model to ground interpretation in quantitative morphology. Through supervised fine-tuning, GPT associates numerical structure with semantic meaning, producing biologically consistent explanations of growth patterns and adaptive traits. Experiments show that the structure-guided framework achieves robust, interpretable reasoning across 12 plant species with diverse root architectures. By integrating unsupervised 3D geometric perception with large-scale language understanding, our approach bridges quantitative analysis and semantic interpretation, establishing a unified paradigm for explainable robotic plant root phenotyping.
Chinese Translation
植物根系表型分析对于理解地下结构、优化作物管理和提高农业可持续性至关重要。本文提出了一种多模态机器人人工智能框架,该框架将3D骨架提取与语言引导推理相结合,以实现可解释且数据高效的根系分析。我们开发了一种基于加权拉普拉斯收缩(Weighted Laplacian Contraction, W-LBC)的无监督骨架提取网络,从机器人3D传感平台捕获的密集点云中生成高保真结构表示。通过重建的骨架图计算定量形态描述符,包括根数、长度、分枝角度和密度,以捕捉几何和拓扑特征。在这些特征的基础上,我们引入了一种以证据为先的语言建模框架,利用自动生成的指令-响应对对GPT进行微调,使其成为一个交互式分析聊天机器人。每个训练样本在自然语言推理之前提供可测量的证据,使模型能够将解释与定量形态学相结合。通过监督微调,GPT将数值结构与语义意义关联起来,生成生物学上一致的生长模式和适应性特征的解释。实验证明,该结构引导框架在12种具有多样根系结构的植物中实现了稳健且可解释的推理。通过将无监督的3D几何感知与大规模语言理解相结合,我们的方法架起了定量分析与语义解释之间的桥梁,建立了一个统一的可解释机器人植物根系表型分析范式。
cs.CV / 37 / 2608.03112
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models
用于视频语言模型高效推理的自适应两阶段视觉标记剪枝
Abstract
Vision-language models excel at image and video understanding but suffer from high inference latency due to the need to process thousands of tokens per image, limiting their deployment on resource-constrained edge devices and in real-time surveillance applications. This challenge is further amplified in video processing, where multiple frames must be analyzed simultaneously. Existing token reduction techniques are largely developed for single-image inputs and therefore fail to account for the temporal and inter-frame redundancies present in video sequences. In addition, these methods generally rely on a fixed, uniform pruning ratio applied across all inputs, which is suboptimal because the degree of redundancy can vary significantly between different videos, necessitating content-dependent pruning levels to preserve critical information. To address these limitations, we propose a two-stage adaptive token pruning strategy specifically designed for video processing. In the first stage, we prune out the redundant frames, and in the second stage, token-level pruning is applied within the retained frames. Crucially, the pruning ratio in the second stage is determined adaptively based on the content of each video. This is achieved by analyzing the correlation structure of token embeddings to quantify redundancy, which is used to determine the ratio. Importantly, our method is entirely post-hoc and requires no additional training or fine-tuning, while achieving strong empirical gains; notably, it improves accuracy by +7\% on a video captioning benchmark at 10\% token retention, while reducing computation TFLOPs by 95\%.
Chinese Translation
视觉语言模型在图像和视频理解方面表现出色,但由于需要处理每幅图像的数千个标记,导致推理延迟较高,这限制了它们在资源受限的边缘设备和实时监控应用中的部署。这一挑战在视频处理过程中进一步加剧,因为必须同时分析多个帧。现有的标记减少技术主要针对单幅图像输入开发,因此未能考虑视频序列中存在的时间和帧间冗余。此外,这些方法通常依赖于在所有输入中应用的固定统一剪枝比例,这并不理想,因为不同视频之间冗余程度可能显著不同,因此需要依赖内容的剪枝水平以保留关键信息。为了解决这些局限性,我们提出了一种专门针对视频处理的两阶段自适应标记剪枝策略。在第一阶段,我们剪除冗余帧,在第二阶段,我们在保留的帧内应用标记级剪枝。关键是,第二阶段的剪枝比例是根据每个视频的内容自适应确定的。这是通过分析标记嵌入的相关结构来量化冗余,从而确定比例。重要的是,我们的方法完全是事后处理,不需要额外的训练或微调,同时实现了显著的经验性提升;值得注意的是,在10\%标记保留的情况下,它在视频字幕基准测试中提高了+7\%的准确性,同时将计算TFLOPs减少了95\\%。
cs.CV / 38 / 2608.03113
Non-Destructive Quantification of Urea Adulteration in Bovine Milk Using Transmittance Multispectral Imaging
利用透射多光谱成像对牛奶中尿素掺假进行无损定量分析
Abstract
Adulteration of bovine milk using urea remains a major food quality and health concern, motivating the development of rapid and quantitative screening tools. Conventional approaches, including laboratory-based analytical methods and spectroscopic techniques, have been used for urea detection; however, many remain less suitable for rapid, low-cost routine screening due to requirements such as specialized instrumentation, sample preparation, chemical reagents, or laboratory operation. This study introduces a pragmatic, cost-effective, accurate, and laboratory-validated MSI-based method for quantitative urea estimation under controlled density conditions using a multispectral-imaging-based regression framework. An in-house-built multispectral imaging system operating in twelve discrete spectral bands (365--940~nm) was used to acquire multispectral images of milk samples prepared with controlled urea addition and water for density balancing. Fresh milk was obtained on the day of image acquisition, and the specific gravity of the milk was verified to be 1.032 at 20{\deg}C using a hydrometer. Multiple linear regression provided an initial mapping with a high validation $R^2$ of 0.9599, while a feed-forward neural network further improved predictive performance with a validation $R^2$ of 0.9773. These results demonstrate the feasibility of transmittance multispectral imaging for accurate, non-destructive urea quantification under controlled density-balanced conditions, supporting its potential as a rapid screening approach for milk-quality assessment.
Chinese Translation
牛奶中尿素的掺假仍然是一个主要的食品质量和健康问题,这促使了快速和定量筛查工具的开发。传统方法,包括基于实验室的分析方法和光谱技术,已被用于尿素检测;然而,由于需要专用仪器、样品准备、化学试剂或实验室操作等要求,许多方法不太适合快速、低成本的常规筛查。本研究提出了一种务实、经济、准确且经过实验室验证的基于多光谱成像(MSI)的方法,用于在控制密度条件下进行尿素的定量估算,采用多光谱成像回归框架。我们使用自建的多光谱成像系统,在十二个离散光谱波段(365-940 nm)中操作,获取了添加了控制尿素和水以平衡密度的牛奶样品的多光谱图像。新鲜牛奶在图像采集当天获得,并使用比重计验证其在20°C时的比重为1.032。多元线性回归提供了初步映射,验证的 $R^2$ 值高达0.9599,而前馈神经网络进一步提高了预测性能,验证的 $R^2$ 值为0.9773。这些结果表明,在控制密度平衡条件下,透射多光谱成像用于准确、无损的尿素定量分析是可行的,支持其作为牛奶质量评估的快速筛查方法的潜力。
cs.CV / 39 / 2608.03120
SeCo-SBIR: Semantically Consistent Prompt Learning for Zero-Shot Sketch-Based Image Retrieval
SeCo-SBIR:用于零样本草图基础图像检索的语义一致性提示学习
Abstract
Adapting CLIP for zero-shot sketch-based image retrieval (ZS-SBIR) via prompt learning faces a fundamental tension: the model must bridge the sketch-photo domain gap through task-specific adaptation, yet the added flexibility risks overfitting to seen training categories and eroding CLIP's zero-shot generalization. We present SeCo-SBIR, a semantically consistent prompt learning framework that resolves this tension from both sides. First, a text-guided multi-modal prompting strategy routes learnable prompt vectors through CLIP's text encoder and projects the resulting intermediate representations into the visual encoder at every layer via learnable coupling functions. Because the text encoder has already learned robust, abstract category-level semantics from large-scale language supervision, this mechanism injects transferable semantic knowledge directly into the visual pathway - adapting the model to the sketch-photo domain while inherently favoring generalization to unseen classes. Second, a perturbation-based consistency constraint addresses the residual overfitting risk from the learnable coupling functions by aligning the adapted model with a frozen CLIP reference branch using an asymmetric InfoNCE objective - augmented inputs feed the frozen branch while clean inputs feed the trainable branch - anchoring the learned representations to CLIP's generalizable feature space. Together with lightweight adapters and a multi-objective loss combining triplet, NT-Xent, and classification terms, SeCo-SBIR achieves state-of-the-art results on all three standard ZS-SBIR benchmarks across categorical, generalized, and across-dataset settings.
Chinese Translation
通过提示学习将 CLIP 适配于零样本草图基础图像检索(ZS-SBIR)面临着根本性的矛盾:模型必须通过任务特定的适配来弥合草图与照片之间的领域差距,但增加的灵活性又有可能导致对已见训练类别的过拟合,从而削弱 CLIP 的零样本泛化能力。我们提出了 SeCo-SBIR,这是一种从两个方面解决这一矛盾的语义一致性提示学习框架。首先,文本引导的多模态提示策略通过 CLIP 的文本编码器引导可学习的提示向量,并通过可学习的耦合函数在每一层将生成的中间表示投影到视觉编码器中。由于文本编码器已经从大规模语言监督中学习到了稳健的抽象类别级语义,这一机制将可转移的语义知识直接注入视觉路径——在适应模型到草图-照片领域的同时,内在地有利于对未见类别的泛化。其次,基于扰动的一致性约束通过使用不对称的 InfoNCE 目标,将适配后的模型与一个冻结的 CLIP 参考分支对齐,从而解决了可学习耦合函数带来的残余过拟合风险——增强的输入馈送到冻结分支,而干净的输入馈送到可训练分支——将学习到的表示锚定到 CLIP 的可泛化特征空间。结合轻量级适配器和结合三元组、NT-Xent 和分类项的多目标损失,SeCo-SBIR 在所有三个标准 ZS-SBIR 基准测试中,在类别、泛化和跨数据集设置上均取得了最先进的结果。
cs.CV / 40 / 2608.03135
Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfolds
先校正再扩散:在去噪轨迹展开之前解开概念
Abstract
Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts. We trace these failures to an early coordination bottleneck: before denoising begins, prompt-conditioned attention may allocate different concepts to strongly overlapping spatial support, which can keep their attention coupled as denoising proceeds. This observation motivates treating compositional generation as a boundary-condition problem rather than repeatedly controlling the evolving trajectory. To this end, we propose Rectify-then-Diffuse (RTD), a training-free framework that rectifies the initial allocation once before standard denoising. Firstly, we propose Soft-Overlap Disentanglement (SOD), which converts normalized overlap between pilot concept maps into a differentiable and layout-agnostic separation objective. Secondly, we introduce Isotropic Gradient Rectification (IGR), which normalizes the SOD gradient and applies a bounded latent displacement with a consistent scale across prompts and initializations. Extensive experiments show that RTD achieves state-of-the-art compositional fidelity and robust gains. On the AE-Bench object pair subset, RTD improves BLIP-VQA by 45.8% and ImageReward by 19.6% over CO3 while running 2.3$\times$ faster. Code will be released at https://github.com/Z-yiwei/rectify-then-diffuse
Chinese Translation
文本到图像的扩散模型能够很好地生成单个概念,但在处理多个概念时,它们常常会遗漏或错误地合并概念。我们将这些失败追溯到一个早期的协调瓶颈:在去噪开始之前,基于提示的注意力可能会将不同的概念分配到强重叠的空间支持上,这可能会在去噪过程中保持它们的注意力耦合。这一观察促使我们将组合生成视为一个边界条件问题,而不是反复控制不断演变的轨迹。为此,我们提出了先校正再扩散(Rectify-then-Diffuse, RTD),这是一个无训练的框架,在标准去噪之前一次性校正初始分配。首先,我们提出了软重叠解缠(Soft-Overlap Disentanglement, SOD),它将试点概念图之间的归一化重叠转换为一个可微分且与布局无关的分离目标。其次,我们引入了各向同性梯度校正(Isotropic Gradient Rectification, IGR),它对SOD梯度进行归一化,并在提示和初始化之间应用一致尺度的有界潜在位移。大量实验表明,RTD在组合保真度和稳健性提升方面达到了最先进的水平。在AE-Bench对象对子集上,RTD使BLIP-VQA的性能提高了45.8%,使ImageReward提高了19.6%,并且运行速度比CO3快2.3倍。代码将发布在https://github.com/Z-yiwei/rectify-then-diffuse
cs.CV / 41 / 2608.03136
Frozen High-Resolution Inference for Cross-City Object Detection: An AI City Challenge 2026 Study
跨城市目标检测的冻结高分辨率推理:AI城市挑战2026研究
Abstract
Cross-city object detection requires a detector trained in one city to generalize to an unlabeled target city. In AI City Challenge 2026 Track 6, we analyze archived configurations of a single RF-DETR-Large detector inside an air-gapped Training-as-a-Service platform whose server returns only an aggregate COCO-style AP over a hidden mixture of source- and target-city images. Frozen 1120 x 1120 inference of a checkpoint trained at 704 x 704 achieved the highest aggregate AP among the evaluated configurations (0.3272 -> 0.3654, +0.0382) without any parameter update, with the largest relative gain on small objects and the largest absolute gain on medium objects, at 2.53x the input pixels. A warm-start 1120px fine-tuning recipe reached 0.3470 while its in-domain validation AP rose (0.767 -> 0.789), a caution that in-domain validation is an unreliable model-selection signal under aggregate-only cross-city feedback. Because that run's evaluation used a higher confidence threshold than the inference-only runs (0.05 vs. 0.01), we treat its score as a descriptive archived outcome rather than a controlled verdict on fine-tuning. Gray-world normalization did not meaningfully change the frozen-1120 result, and a rectangular run was found by audit to have used an unintended portrait orientation. We release verbatim platform commands, configuration snapshots, and an explicit evidence boundary for every claim. Each configuration was submitted once and the best was selected on the hidden server, so these are exploratory, audited findings about the aggregate mixture; they do not establish target-city-specific improvement.
Chinese Translation
跨城市目标检测要求在一个城市训练的检测器能够推广到一个未标记的目标城市。在AI城市挑战2026的第六赛道中,我们分析了在一个与外界隔离的训练即服务平台内,单个RF-DETR-Large检测器的归档配置,该平台的服务器仅返回隐藏的源城市和目标城市图像混合的聚合COCO风格平均精度(AP)。在704 x 704训练的检查点上,冻结1120 x 1120的推理在评估的配置中达到了最高的聚合AP(0.3272 -> 0.3654,+0.0382),且没有任何参数更新,在小物体上获得了最大的相对增益,在中等物体上获得了最大的绝对增益,输入像素为2.53倍。一个温启动的1120px微调方案达到了0.3470,同时其领域内验证AP上升(0.767 -> 0.789),这表明在仅有聚合的跨城市反馈下,领域内验证是一个不可靠的模型选择信号。由于该运行的评估使用了比仅推理运行更高的置信度阈值(0.05 vs. 0.01),我们将其得分视为描述性的归档结果,而不是对微调的控制性判决。灰世界归一化对冻结的1120结果没有实质性影响,审计发现一个矩形运行使用了意外的肖像方向。我们发布了逐字的平台注册命令、配置快照以及每个声明的明确证据边界。每个配置仅提交一次,最佳配置在隐藏服务器上被选中,因此这些是关于聚合混合的探索性审计发现;它们并未建立针对目标城市的特定改进。
cs.CV / 42 / 2608.03143
From Routes to Steps: Separating Semantic Progress from Local Execution in Vision-and-Language Navigation
从路线到步骤:在视觉与语言导航中将语义进展与局部执行分离
Abstract
Vision-and-Language Navigation (VLN) requires an agent to follow a route-level instruction by executing its constituent steps from egocentric visual observations. Existing VLM-based navigators typically supervise both capabilities through next-action prediction alone, making progress-tracking errors difficult to distinguish from execution errors. When an agent deviates from the route, a corrective action label may recover the next movement but does not indicate whether the agent selected the wrong sub-instruction or failed to execute the correct one. Consequently, the agent may continue making decisions from an erroneous progress state. To resolve this ambiguity, we propose \textbf{Route2Step}, a framework that decouples semantic progress tracking from action generation through an explicit step-level interface. The Instruction Analysis Module ($\mathcal{M}_{\mathrm{IA}}$) predicts this state from the global instruction and visual history. Conditioned on the predicted state and recent observations, the Action Generation Module ($\mathcal{M}_{\mathrm{AG}}$) generates local action chunks. To supervise the progress state without manual temporal labels, E-SPA, a step-alignment procedure, associates sub-instructions with their corresponding portions of route-level demonstrations. These alignments enable state supervision for incorrect progress estimates, while direct action supervision is reserved for rollout groups that repeatedly fail under the correct active sub-instruction. On R2R-CE, Route2Step improves SR from 48.1\% to 55.3\% and SPL from 43.3\% to 48.2\%, using 190K state-level corrective samples while requiring only 11.5K directly action-supervised states. Experiments in real-world indoor and outdoor environments further demonstrate the practical applicability of Route2Step. The project page is: https://sisyphus-hxy.github.io/Route2Step/.
Chinese Translation
视觉与语言导航(Vision-and-Language Navigation, VLN)要求代理根据路线级指令,通过自我中心的视觉观察执行其组成步骤。现有的基于视觉与语言模型(VLM)的导航器通常仅通过下一步动作预测来监督这两种能力,使得进展跟踪错误难以与执行错误区分。当代理偏离路线时,纠正动作标签可能恢复下一步移动,但并不能指示代理是选择了错误的子指令还是未能执行正确的指令。因此,代理可能继续在错误的进展状态下做出决策。为了解决这一模糊性,我们提出了 extbf{Route2Step},一个通过明确的步骤级接口将语义进展跟踪与动作生成解耦的框架。指令分析模块(Instruction Analysis Module, $ extmath{M}_{ ext{IA}}$)根据全局指令和视觉历史预测这一状态。基于预测状态和最近观察,动作生成模块(Action Generation Module, $ extmath{M}_{ ext{AG}}$)生成局部动作片段。为了在没有手动时间标签的情况下监督进展状态,E-SPA(步骤对齐程序)将子指令与其对应的路线级演示部分关联。这些对齐使得对错误进展估计的状态监督成为可能,而直接的动作监督则保留给在正确的活动子指令下反复失败的回放组。在R2R-CE数据集上,Route2Step将成功率(Success Rate, SR)从48.1 ext{%}提高到55.3 ext{%},将成功路径长度(Success Path Length, SPL)从43.3 ext{%}提高到48.2 ext{%},使用了190K状态级纠正样本,同时仅需11.5K直接动作监督状态。在真实世界的室内和室外环境中的实验进一步证明了Route2Step的实际适用性。项目页面为:https://sisyphus-hxy.github.io/Route2Step/.
cs.CV / 43 / 2608.03147
CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation
CROSS:用于遥感指称分割的级联蒸馏与双约束定位
Abstract
Referring Remote Sensing Image Segmentation (RRSIS) has achieved significant progress through the integration of VLMs and the Segment Anything Model (SAM). However, this progress largely relies on strong pre-trained capabilities, while leaving two fundamental limitations insufficiently addressed: (1) Architectural Weak-Coupling, where the unidirectional flow forces reliance on coarse VLM prompts and wastes SAM's pixel-level structural guidance, causing localization drift; and (2) Object-Centric Semantic Bias, where models overemphasize dominant object semantics while remaining insensitive to spatial reasoning crucial for RRSIS. Motivated by these observations, we propose CROSS, a tightly integrated paradigm for RRSIS. First, we introduce Linguistic-Guided Cascaded Distillation (LGCD) to bridge the architectural gap, which distills SAM's geometric affinities as soft regularizers into VLM intermediate layers, injecting dense structural priors to refine localization. Second, Perspective-Spatial Contrastive Learning (PSCL) imposes cross-anchored constraints by mining mask-filtered deceptive distractors and spatial-linguistic counterfactuals as hard negatives, explicitly shattering semantic shortcuts to enforce genuine logical consistency. Extensive experiments on RRSIS benchmarks demonstrate that CROSS achieves state-of-the-art performance and maintains precise localization even under severe spatial description perturbations, standing as a robust new paradigm for RRSIS.
Chinese Translation
指称遥感图像分割(RRSIS)通过整合视觉语言模型(VLMs)和Segment Anything Model(SAM)取得了显著进展。然而,这一进展在很大程度上依赖于强大的预训练能力,同时未能充分解决两个基本限制:(1)架构弱耦合,单向流动迫使依赖粗糙的VLM提示,浪费了SAM的像素级结构指导,导致定位漂移;(2)以对象为中心的语义偏差,模型过度强调主导对象语义,而对RRSIS至关重要的空间推理则显得敏感不足。基于这些观察,我们提出了CROSS,一个紧密集成的RRSIS范式。首先,我们引入语言引导的级联蒸馏(LGCD)来弥合架构差距,该方法将SAM的几何亲和性作为软正则化器蒸馏到VLM中间层,注入密集的结构先验以细化定位。其次,视角-空间对比学习(PSCL)通过挖掘掩膜过滤的欺骗性干扰物和空间-语言反事实作为硬负样本施加交叉锚定约束,明确打破语义捷径以强制执行真正的逻辑一致性。在RRSIS基准上的大量实验表明,CROSS实现了最先进的性能,并在严重的空间描述扰动下保持精确的定位,成为RRSIS的一个强大新范式。
cs.CV / 44 / 2608.03158
Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation
多物体与关节人-物交互生成的表面关键点表示
Abstract
Daily activities require humans to coordinate whole-body motion with the motion of surrounding objects. Despite recent progress in human-object interaction (HOI) generation, most existing methods assume interactions with a single rigid object and do not extend well to scenarios involving a variable number of objects or articulated objects with diverse joint mechanisms. We propose surface keypoint trajectories as an object motion representation: for each rigid component, whether a standalone object or one part of an articulated assembly, we track a small set of non-collinear surface points over time. This representation handles multi-object coordination and diverse articulation mechanisms directly from point dynamics without requiring explicit joint-type specification. To model when and where each body region contacts each object, we introduce a spatio-temporal contact distance field that extends distance-based contact modeling to whole-body, multi-object, and articulated settings. We factorize HOI generation into three stages: generating object motions from text or waypoints, predicting the contact distance field, and synthesizing whole-body motion with contact-guided optimization. Experiments on ParaHome, HIMO, ARCTIC, and OMOMO demonstrate better or comparable performance to existing methods across single-object, multi-object, and articulated interaction settings.
Chinese Translation
日常活动要求人类协调全身运动与周围物体的运动。尽管在人-物交互(HOI)生成方面取得了近期进展,但大多数现有方法假设与单一刚性物体的交互,且不适用于涉及可变数量物体或具有多样关节机制的关节物体的场景。我们提出了表面关键点轨迹作为物体运动的表示:对于每个刚性组件,无论是独立物体还是关节组件的一部分,我们跟踪一小组不共线的表面点随时间的变化。该表示能够直接处理多物体协调和多样的关节机制,且无需明确的关节类型规范。为了建模每个身体区域何时何地接触每个物体,我们引入了一个时空接触距离场,将基于距离的接触建模扩展到全身、多物体和关节设置。我们将HOI生成分解为三个阶段:从文本或路径点生成物体运动、预测接触距离场,以及通过接触引导优化合成全身运动。在ParaHome、HIMO、ARCTIC和OMOMO上的实验表明,在单物体、多物体和关节交互设置中,性能优于或可与现有方法相媲美。
cs.CV / 45 / 2608.03176
Frequency-Decorrelated Temporal Ensembles for EEG--fNIRS Imagined-Handwriting Decoding
频率去相关的时间集成用于脑电图-功能性近红外成像想象手写解码
Abstract
Imagined handwriting offers a temporally rich paradigm for non-invasive neural decoding, yet reliable recognition across unseen participants remains difficult because scalp EEG is noisy and internally generated stroke sequences vary across individuals. The Multimodal Brain-Computer Interface Grand Challenge provides synchronized EEG and fNIRS for four-class subject-independent handwriting-trajectory classification. We propose FRED, a task-adapted system that models imagined handwriting as a multi-second motor sequence and trains a compact multi-scale temporal network on three complementary EEG frequency views. With three seeds per view, cross-band members produce substantially less-correlated errors than same-band replicas, yielding a clean nine-member ensemble accuracy of 0.8076/0.7242/0.7492 on the public/private/overall test partitions without test-set adaptation or output constraints. The submitted pipeline further incorporates transductive pseudo-label training, three EEG-Conformer members, posterior aggregation, and a paradigm-aware decoder. Because every 12-trial randomization block contains three instances of each class, the final predictions are obtained by Hungarian assignment under the known block quota. On one fixed posterior pool, independent, session-constrained, and block-constrained decoding achieve 0.7600, 0.7758, and 0.7952 overall accuracy, respectively. The complete system reaches 0.8498/0.7718/0.7952, ranking fourth on the private split. A modality audit finds fNIRS-only decoding at chance (0.2511 overall), while adding fNIRS to EEG changes accuracy by only +0.0025. These results identify frequency-diverse temporal EEG modeling and protocol-matched structured inference as the principal sources of performance in this sparse-montage EEG--fNIRS setting. The source code is available at https://github.com/XiuFan719/EEG-fNIRS-fuse-method-for-MM-challenge.
Chinese Translation
想象手写为非侵入性神经解码提供了一个时间丰富的范式,但在未见参与者之间实现可靠识别仍然困难,因为头皮脑电图(EEG)噪声较大,且个体内部生成的笔画序列存在差异。多模态脑-计算机接口大奖赛提供了同步的EEG和功能性近红外成像(fNIRS),用于四类主体独立的手写轨迹分类。我们提出了FRED,一个任务适应的系统,将想象手写建模为多秒的运动序列,并在三个互补的EEG频率视图上训练一个紧凑的多尺度时间网络。每个视图有三个种子,跨频带成员产生的相关错误显著低于同频带副本,在公共/私有/总体测试分区上,干净的九成员集成准确率分别为0.8076/0.7242/0.7492,且没有测试集适应或输出约束。提交的管道进一步结合了传导伪标签训练、三个EEG-Conformer成员、后期聚合和一个范式感知解码器。由于每个12次试验随机化块包含每个类别的三个实例,最终预测是通过已知块配额下的匈牙利分配获得的。在一个固定的后期池上,独立、会话约束和块约束的解码分别实现了0.7600、0.7758和0.7952的总体准确率。完整系统的准确率达到0.8498/0.7718/0.7952,在私有分割中排名第四。一项模态审计发现,仅使用fNIRS的解码准确率为偶然水平(总体0.2511),而将fNIRS添加到EEG中仅使准确率提高了+0.0025。这些结果表明,在这种稀疏拼接的EEG-fNIRS设置中,频率多样化的时间EEG建模和协议匹配的结构推断是性能的主要来源。源代码可在https://github.com/XiuFan719/EEG-fNIRS-fuse-method-for-MM-challenge获取。
cs.CV / 46 / 2608.03179
EditFlow3D: Automated Local Editing of 3D Assets with Trajectory Preservation
EditFlow3D:具有轨迹保留的3D资产自动局部编辑
Abstract
Controllable local editing of 3D assets requires precise target localization and appropriate visual guidance. However, existing methods lack a simple yet accurate way to obtain 3D masks and struggle to achieve the desired edit while faithfully preserving the structure and appearance of non-target regions. To address these challenges, we present EditFlow3D, a training-free framework for local 3D editing. Given a source asset and an edit instruction, a VLM-driven workflow interprets the editing intent and automatically constructs a visual guidance image and a refined 3D editing mask, enabling localized editing in the native representation space of a pretrained 3D generative model. Specifically, mask-guided differential flow focuses the edit on the target region, while step-wise trajectory preservation maintains consistency between non-target regions and the source asset without directly replacing intermediate features. Since the existing Edit3D-Bench covers only a limited range of local editing categories, we further introduce EditFlow-Bench as a complementary benchmark encompassing a broader variety of structural and appearance edits, and evaluate EditFlow3D on both benchmarks. Quantitative results, qualitative comparisons, and a user study demonstrate that EditFlow3D achieves more accurate target-region editing and better preserves non-target regions than existing 3D editing methods.
Chinese Translation
可控的3D资产局部编辑需要精确的目标定位和适当的视觉引导。然而,现有的方法缺乏一种简单而准确的方式来获取3D掩膜,并且在实现所需编辑的同时,难以忠实地保留非目标区域的结构和外观。为了解决这些挑战,我们提出了EditFlow3D,这是一个无训练的局部3D编辑框架。给定一个源资产和一个编辑指令,基于视觉语言模型(VLM)的工作流程解读编辑意图,并自动构建视觉引导图像和精细化的3D编辑掩膜,从而在预训练的3D生成模型的原生表示空间中实现局部编辑。具体而言,掩膜引导的差分流聚焦于目标区域的编辑,而逐步轨迹保留则在不直接替换中间特征的情况下,保持非目标区域与源资产之间的一致性。由于现有的Edit3D-Bench仅涵盖有限的局部编辑类别,我们进一步引入EditFlow-Bench作为一个补充基准,涵盖更广泛的结构和外观编辑,并在这两个基准上评估EditFlow3D。定量结果、定性比较和用户研究表明,EditFlow3D在目标区域编辑的准确性和非目标区域的保留方面优于现有的3D编辑方法。
cs.CV / 47 / 2608.03185
CRIL-U-Net: Compact Ratio-Interaction Learning for Focal Cortical Dysplasia Segmentation from T1w and FLAIR MRI
CRIL-U-Net:用于从T1加权和FLAIR MRI中进行焦点皮层发育不良分割的紧凑比率交互学习
Abstract
Focal cortical dysplasia (FCD) type II is an important structural cause of drug-resistant focal epilepsy, but its small size, heterogeneous appearance, and subtle MRI characteristics make automated segmentation challenging. Conventional multimodal networks commonly concatenate T1-weighted (T1w) and fluid-attenuated inversion recovery (FLAIR) images, requiring subsequent layers to learn useful cross-modal relationships implicitly. We propose CRIL-U-Net, a 3D U-Net incorporating a Compact Ratio-Interaction Learning module that combines local spatial features, voxel-wise cross-modal mixing, and bidirectional ratio-inspired interactions. CRIL-U-Net was compared with a conventional 3D U-Net and an input self-attention U-Net using five-fold cross-validation on 85 FCD subjects and 25 healthy controls. Each architecture was trained independently using Dice-binary cross-entropy (Dice-BCE) and Focal Tversky-Focal (FTF) losses. With FTF, CRIL-U-Net achieved the highest mean Dice score (0.196 +/- 0.262), compared with 0.136 +/- 0.224 for the U-Net and 0.135 +/- 0.214 for the attention comparator. It produced nonzero lesion overlap in 44 of 85 cases, compared with 36 for the U-Net. Under FTF, CRIL-U-Net significantly outperformed both comparison architectures after false-discovery-rate correction. These findings suggest that compact cross-modal representation learning can improve FCD segmentation within a controlled U-Net setting when combined with an imbalance-aware objective, although the remaining zero-overlap rate of 48.2% highlights the need for further validation and methodological development.
Chinese Translation
焦点皮层发育不良(FCD)II型是药物难治性局灶性癫痫的重要结构性原因,但其小尺寸、异质外观和微妙的MRI特征使得自动分割具有挑战性。传统的多模态网络通常将T1加权(T1w)和液体衰减反转恢复(FLAIR)图像连接在一起,要求后续层隐式学习有用的跨模态关系。我们提出了CRIL-U-Net,这是一种3D U-Net,结合了紧凑比率交互学习模块,结合了局部空间特征、体素级跨模态混合和双向比率启发式交互。CRIL-U-Net与传统的3D U-Net和输入自注意力U-Net进行了比较,使用五折交叉验证在85名FCD患者和25名健康对照者上进行评估。每种架构独立训练,使用Dice-二元交叉熵(Dice-BCE)和Focal Tversky-Focal(FTF)损失。在FTF下,CRIL-U-Net达到了最高的平均Dice分数(0.196 +/- 0.262),而U-Net为0.136 +/- 0.224,注意力比较器为0.135 +/- 0.214。它在85个案例中有44个产生了非零病灶重叠,而U-Net则为36个。在FTF下,CRIL-U-Net在假发现率校正后显著优于两种比较架构。这些发现表明,紧凑的跨模态表示学习可以在结合不平衡意识目标的受控U-Net设置中改善FCD分割,尽管仍有48.2%的零重叠率突显了进一步验证和方法开发的必要性。
cs.CV / 48 / 2608.03198
Bridging Online and Offline Handwriting via Differentiable Physical Rendering
通过可微物理渲染桥接在线与离线手写
Abstract
Realistic handwritten text generation plays an important role in numerous applications, such as font design, biometric authentication, and robotic calligraphy. Existing methods are typically divided into two independent paradigms: online approaches that estimate handwriting trajectories and offline approaches that synthesize realistic handwriting images. While online models capture structural and temporal dynamics, they often lack fine-grained textures, whereas offline models reproduce realistic appearance but discard stroke order. However, unifying online and offline models remains challenging due to (1) the lack of an explicit physical model linking stroke kinematics to pixel-level appearance and (2) the absence of paired trajectory-image datasets. Moreover, enabling end-to-end learning requires a differentiable rendering process across motion and appearance domains. To address these challenges, we propose a compact physical brush model that bridges stroke dynamics and visual appearance, together with a differentiable rendering module that converts stroke trajectories into stylized images. By integrating these components, we propose a unified online-offline handwriting generation framework via differentiable brush rendering. The proposed framework consists of four core modules: 1) a text-to-stroke generator that predicts the target stroke conditioned on the given text and style image, 2) a brush parameter observer that extracts brush model parameters from style references, 3) a differentiable brush renderer that maps a stroke sequence and physical brush parameters into a handwritten image, and 4) a zero-shot image refiner that refines rendered images via diffusion models. Extensive experiments and real-world robotic calligraphy demonstrations validate our approach, achieving both structural and visual fidelity.
Chinese Translation
现实手写文本生成在字体设计、生物识别认证和机器人书法等众多应用中发挥着重要作用。现有方法通常分为两种独立的范式:在线方法估计手写轨迹,而离线方法合成真实的手写图像。虽然在线模型捕捉了结构和时间动态,但它们往往缺乏细致的纹理,而离线模型则重现了真实的外观但忽略了笔画顺序。然而,由于(1)缺乏将笔画运动学与像素级外观连接的明确物理模型,以及(2)缺乏配对的轨迹-图像数据集,统一在线和离线模型仍然具有挑战性。此外,实现端到端学习需要在运动和外观领域之间进行可微的渲染过程。为了解决这些挑战,我们提出了一种紧凑的物理笔刷模型,连接笔画动态与视觉外观,并结合一个可微渲染模块,将笔画轨迹转换为风格化图像。通过整合这些组件,我们提出了一种通过可微笔刷渲染的统一在线-离线手写生成框架。该框架由四个核心模块组成:1)一个文本到笔画生成器,根据给定的文本和风格图像预测目标笔画;2)一个笔刷参数观察器,从风格参考中提取笔刷模型参数;3)一个可微笔刷渲染器,将笔画序列和物理笔刷参数映射到手写图像;4)一个零样本图像细化器,通过扩散模型细化渲染图像。大量实验和现实世界的机器人书法演示验证了我们的方法,实现了结构和视觉的保真度。
cs.CV / 49 / 2608.03207
DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack
DRIFT:通过对流匹配视觉-语言-动作模型的对抗性补丁攻击来破坏去噪轨迹
Abstract
Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs. We show that this robustness is largely illusory: it stems from prior attacks ignoring the multi-step denoising ODE. We introduce DRIFT (Denoising Redirection via Input perturbation of the Flow-matching Trajectory), a test-time universal adversarial patch placed on the robot's gripper that attacks the denoising velocity field of an off-the-shelf policy. Our central finding is counterintuitive: attacking only the first denoising step is both stronger and cheaper than attacking a wider window of steps, which we explain through a gradient conflict unique to input-space optimization and which is exactly opposite to the training-time backdoor regime. On pi0 and pi0.5 across four LIBERO suites, DRIFT breaks essentially all originally-solvable tasks with a small single patch, far exceeding action- and embedding-space attack baselines.
Chinese Translation
流匹配视觉-语言-动作(VLA)模型,如 pi0,通过整合学习到的去噪速度场生成机器人动作,并已被报告能够抵抗轻易欺骗自回归 VLA 的对抗扰动。我们表明,这种鲁棒性在很大程度上是虚幻的:它源于之前的攻击忽视了多步去噪常微分方程(ODE)。我们提出了 DRIFT(通过输入扰动重定向去噪流匹配轨迹),这是一种在机器人抓手上放置的测试时通用对抗性补丁,旨在攻击现成策略的去噪速度场。我们的核心发现是反直觉的:仅攻击第一个去噪步骤比攻击更广泛的步骤窗口更强且成本更低,这一点我们通过输入空间优化中独特的梯度冲突来解释,这与训练时的后门机制正好相反。在四个 LIBERO 套件中的 pi0 和 pi0.5 上,DRIFT 通过一个小的单一补丁几乎破坏了所有原本可解决的任务,远远超过了动作和嵌入空间攻击的基线。
cs.CV / 50 / 2608.03211
CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction
CrossScope:一种用于联合双视角外科视频预测的角色非对称世界模型
Abstract
Visual world models typically learn future dynamics from a single observation stream, limiting their ability to model cooperative systems with multiple independently moving observers. We investigate this challenge in Mother--Child endoscopic retrograde cholangiopancreatography (ERCP), where two flexible scopes provide complementary yet role-dependent views without a calibrated stereo relationship. Unlike conventional multi-view fusion that assumes symmetric information exchange, we formulate \textbf{role-asymmetric dual-scope future prediction}, where cross-view evidence is selectively transferred according to the prediction target and its underlying spatial requirements. We propose \textbf{CrossScope}, a dual-stream surgical world model that preserves view-specific experts while enabling target-specific evidence routing through geometry-guided residual interactions. CrossScope learns two complementary communication directions: geometric motion cues from the Mother view guide Child-view future dynamics, while pose-aligned Child appearance supports Mother-view prediction only when valid spatial correspondence is established. This design allows each scope to contribute task-relevant evidence without compromising its view-specific representation. To evaluate this problem, we establish a paired dual-scope benchmark comprising synchronized phantom and real-world ERCP episodes, with evaluations assessing visual fidelity, structural preservation, target localization, and motion consistency. Experiments demonstrate that CrossScope consistently outperforms strong surgical video generation baselines, validating the importance of role-aware evidence routing for multi-observer visual world modeling.
Chinese Translation
视觉世界模型通常从单一观察流中学习未来动态,这限制了它们对多个独立移动观察者的协作系统建模能力。我们在母子内窥镜逆行胰胆管造影(ERCP)中探讨这一挑战,其中两个柔性内窥镜提供互补但依赖角色的视角,而没有经过标定的立体关系。与传统的多视角融合假设对称信息交换不同,我们提出了 extbf{角色非对称双视角未来预测},在这一框架中,跨视角证据根据预测目标及其潜在空间需求被选择性地转移。我们提出了 extbf{CrossScope},一种双流外科世界模型,保留视角特定的专家,同时通过几何引导的残差交互实现目标特定的证据路由。CrossScope学习两个互补的通信方向:母视角的几何运动线索引导子视角的未来动态,而姿态对齐的子视角外观仅在建立有效空间对应关系时支持母视角的预测。这一设计使每个内窥镜能够贡献与任务相关的证据,而不妨碍其视角特定的表示。为评估这一问题,我们建立了一个配对的双视角基准,包括同步的虚拟和真实世界ERCP片段,评估内容包括视觉保真度、结构保留、目标定位和运动一致性。实验表明,CrossScope在外科视频生成的强基线中始终表现优异,验证了角色感知证据路由在多观察者视觉世界建模中的重要性。
cs.CV / 51 / 2608.03216
iFAN: Inference-Aware Learning for Plain Mask Transformers
iFAN:面向推理的普通掩膜变换器学习
Abstract
Query-based mask transformers assemble segmentation outputs through pixel-wise competition among query predictions of the final layer, yet this inference process is not explicitly optimized during training. We identify two key mismatches: the query with the highest probability-mask score does not necessarily produce the most accurate mask, and final-layer decoding may discard superior predictions from intermediate layers. To address these issues, we propose Inference-Aware Learning (iFAN), a general training framework for plain mask transformers. iFAN introduces Adjusted Probability-Mask Ranking (APMR), which aligns query competition with predicted mask quality and suppresses high-confidence but inaccurate competitors. We further employ Cross-Layer Self-Distillation (CLSD) to transfer stronger intermediate predictions to the final layer. The ranking and distillation objectives are training-only, while inference retains efficient final-layer decoding. Experiments on COCO, ADE20K, and Cityscapes demonstrate consistent improvements across panoptic, instance, and semantic segmentation, as well as across different architectures, backbone scales, and input resolutions. Overall, iFAN improves performance by an average of 1.20 PQ, 1.30 AP, and 0.63 mIoU, with negligible additional parameters, FLOPs and inference latency.
Chinese Translation
基于查询的掩膜变换器通过最终层的查询预测之间的逐像素竞争来组装分割输出,然而这一推理过程在训练期间并未得到明确优化。我们识别出两个关键的不匹配:具有最高概率-掩膜分数的查询不一定产生最准确的掩膜,最终层解码可能会丢弃来自中间层的优越预测。为了解决这些问题,我们提出了面向推理的学习(Inference-Aware Learning, iFAN),这是一个针对普通掩膜变换器的通用训练框架。iFAN引入了调整后的概率-掩膜排名(Adjusted Probability-Mask Ranking, APMR),该方法将查询竞争与预测掩膜质量对齐,并抑制高置信度但不准确的竞争者。我们进一步采用跨层自蒸馏(Cross-Layer Self-Distillation, CLSD)将更强的中间预测转移到最终层。排名和蒸馏目标仅在训练中使用,而推理则保留高效的最终层解码。在COCO、ADE20K和Cityscapes上的实验表明,在全景、实例和语义分割方面,以及在不同架构、主干网络规模和输入分辨率下,均实现了一致的改进。总体而言,iFAN平均提高了1.20 PQ、1.30 AP和0.63 mIoU,且附加参数、FLOPs和推理延迟微不足道。
cs.CV / 52 / 2608.03218
Self-Supervised Representation-Guided Generative Dataset Distillation
自监督表示引导的生成数据集蒸馏
Abstract
Dataset distillation compresses a large training set into a compact synthetic set while retaining its downstream utility. Most existing methods target randomly initialized networks, whereas modern vision systems often adapt frozen pretrained encoders with lightweight modules. Distilled samples should therefore preserve the discriminative geometry of the pretrained representation space, which existing generative objectives do not explicitly consider. We propose self-supervised representation-guided generative dataset distillation (SRG), a framework that translates the SSL geometry into diffusion guidance. Specifically, SRG constructs class-wise prototypes from real-image SSL representations and performs guidance through three SSL-space objectives for prototype alignment, inter-class discrimination, and intra-class assignment. During diffusion sampling, it adopts a stage-wise guidance strategy: early denoising is anchored to the latent of the real image whose SSL representation is nearest to the assigned prototype, whereas later denoising is guided by the SSL-space objectives. This division preserves the visual realism provided by the generative prior while progressively steering samples toward representative and class-discriminative regions of the SSL representation space. SRG consistently outperforms the evaluated generative baselines across multiple datasets and IPC settings. A cross-encoder evaluation further indicates transfer across pretrained representation spaces. These results demonstrate the effectiveness of representation-guided generation for dataset distillation with pretrained SSL models.
Chinese Translation
数据集蒸馏将大型训练集压缩为紧凑的合成集,同时保留其下游效用。大多数现有方法针对随机初始化的网络,而现代视觉系统通常使用轻量级模块适配冻结的预训练编码器。因此,蒸馏样本应保留预训练表示空间的判别几何,而现有的生成目标并未明确考虑这一点。我们提出了自监督表示引导的生成数据集蒸馏(SRG)框架,该框架将自监督学习(SSL)几何转化为扩散引导。具体而言,SRG 从真实图像的 SSL 表示中构建类别原型,并通过三个 SSL 空间目标进行引导,以实现原型对齐、类间区分和类内分配。在扩散采样过程中,它采用阶段性引导策略:早期去噪固定在与分配原型最近的真实图像的潜在表示上,而后期去噪则由 SSL 空间目标引导。这一划分保留了生成先验提供的视觉真实感,同时逐步引导样本朝向 SSL 表示空间中的代表性和类区分区域。SRG 在多个数据集和 IPC 设置中始终优于评估的生成基线。交叉编码器评估进一步表明在预训练表示空间之间的迁移。这些结果证明了使用预训练 SSL 模型进行数据集蒸馏的表示引导生成的有效性。
cs.CV / 53 / 2608.03225
Open-Linguistic Concept Unified Learning for Cross-Site Interpretable Dermatology Image Diagnosis
跨站点可解释皮肤病图像诊断的开放语言概念统一学习
Abstract
Human-interpretable computer-aided diagnosis is crucial for clinical decision making. Concept-based models excel by providing transparent reasoning and enabling post-hoc, clinician-in-the-loop interventions. However, their rigid dataset-specific adaptation inherently restricts cross-site generalization. Applying them across diverse modalities, such as dermoscopic and clinical photographs, is challenging due to heterogeneous concept taxonomies varying in availability, granularity, and semantics across cohorts. Consequently, adapting Foundation Vision-Language Models (FVLMs) demands costly label engineering and repeated post-training. Existing intervention mechanisms remain rigidly tied to predefined concepts, lacking adaptability and hindering scalable dermatology CAD deployment. To address these bottlenecks, we propose UniCon, an open-linguistic unified concept learning framework for multimodal interpretable vision-language diagnosis. UniCon resolves these challenges through three contributions: (1) A shared semantic representation space via a unified concept prototype codebook, seamlessly coordinating heterogeneous concept systems across modalities without dataset-specific retraining. (2) Open-linguistic based multi-faceted semantic specifications to overcome sparse textual label limitations, improving boundary sensitivity in uncertain clinical contexts. (3) A robust, cross-site adjustable intervention interface powered by reliability-gated bottleneck aggregation, enabling consistent reasoning and transferable clinician corrections. Extensive experiments demonstrate that beyond securing top-tier diagnostic accuracy, UniCon successfully bridges disparate clinical taxonomies, unlocking unprecedented cross-site intervention capabilities. Code is available at https://github.com/wuchengyu123/UniCon.
Chinese Translation
人类可解释的计算机辅助诊断对于临床决策至关重要。基于概念的模型通过提供透明的推理和支持临床医生的后期干预而表现出色。然而,它们对特定数据集的刚性适应性本质上限制了跨站点的泛化能力。在不同的模态(如皮肤镜图像和临床照片)中应用这些模型面临挑战,因为不同队列中可用性、粒度和语义各异的概念分类法存在异质性。因此,适应基础视觉-语言模型(Foundation Vision-Language Models, FVLMs)需要昂贵的标签工程和重复的后期训练。现有的干预机制仍然严格依赖于预定义的概念,缺乏适应性,阻碍了可扩展的皮肤病计算机辅助诊断部署。为了解决这些瓶颈,我们提出了UniCon,一种用于多模态可解释视觉-语言诊断的开放语言统一概念学习框架。UniCon通过三项贡献解决了这些挑战:(1)通过统一的概念原型词典建立共享的语义表示空间,无需特定数据集的重新训练即可无缝协调不同模态的异质概念系统;(2)基于开放语言的多面向语义规范克服稀疏文本标签的局限性,提高了在不确定临床环境中的边界敏感性;(3)一个强大的、跨站点可调的干预接口,通过可靠性门控瓶颈聚合提供支持,实现一致的推理和可转移的临床医生修正。大量实验证明,除了确保顶尖的诊断准确性外,UniCon成功地弥合了不同的临床分类法,解锁了前所未有的跨站点干预能力。代码可在 https://github.com/wuchengyu123/UniCon 获取。
cs.CV / 54 / 2608.03247
CIGTSurv: Clinical Information Guided Tri-modal Survival Prediction with Local Prototype Association and Global Feature Alignment
CIGTSurv:基于临床信息引导的三模态生存预测框架,结合局部原型关联与全局特征对齐
Abstract
Multimodal learning has significantly advanced survival prediction by integrating pathology images with genomic data. However, clinical information, despite its critical role in reflecting a patient' s overall health, remains underutilized due to its discrete, sparse, and low-dimensional nature. Furthermore, the inherent heterogeneity across these modalities pose significant challenges in modeling cross-modal interactions. In this paper, we propose CIGTSurv, a Clinical Information Guided Tri-modal framework for Survival prediction. Specifically, we first design a holistic text template and use pretrained foundation models to transform clinical tabular data into high-dimensional tokenized embeddings. Using clinical information as an anchor, we then introduce a dual-level interaction mechanism: 1) a local prototype association (LPA) module based on cross-attention to explicitly learn token-level correspondences between different modalities, and 2) a global feature alignment (GFA) loss based on Maximum Mean Discrepancy (MMD) to implicitly enhance cross-modal distribution consistency. Extensive experiments on five TCGA cancer cohorts demonstrate that CIGTSurv achieves state-of-the-art (SOTA) survival prediction performance. Our source code is publicly available at https://github.com/Daijing-ai/CIGT-Surv.git.
Chinese Translation
多模态学习通过将病理图像与基因组数据相结合,显著推动了生存预测的发展。然而,尽管临床信息在反映患者整体健康方面发挥着关键作用,但由于其离散、稀疏和低维的特性,仍然未得到充分利用。此外,这些模态之间固有的异质性在建模跨模态交互时带来了重大挑战。在本文中,我们提出了CIGTSurv,一种基于临床信息引导的三模态生存预测框架。具体而言,我们首先设计了一个整体文本模板,并使用预训练的基础模型将临床表格数据转化为高维的标记嵌入。以临床信息为锚点,我们引入了一种双层交互机制:1)基于跨注意力的局部原型关联(LPA)模块,明确学习不同模态之间的标记级对应关系;2)基于最大均值差异(MMD)的全局特征对齐(GFA)损失,隐式增强跨模态分布一致性。在五个TCGA癌症队列上的广泛实验表明,CIGTSurv实现了最先进的(SOTA)生存预测性能。我们的源代码已公开,地址为https://github.com/Daijing-ai/CIGT-Surv.git。
cs.CV / 55 / 2608.03252
Clarity Contrast and Similarity Selection for Multi-Focus Image Fusion
多焦点图像融合中的清晰度对比与相似性选择
Abstract
Multi-focus image fusion (MFIF) aims to generate an all-in-focus image from multiple images of the same scene focused at different regions. Most existing deep learning-based methods lack explicit interaction between the source images, which limits their performance and interpretability. This paper presents a novel Clarity Contrast and Similarity Selection Network (CSNet), to bridge direct information exchange for MFIF. Specifically, by contrasting the clarity differences between source images within our proposed Clarity Contrast Attention Module (CCAM), we mutually enhance sharp features while suppressing blurry ones. This allows us to identify the exactly focused regions in each source and locate the focused-defocused boundaries. Moreover, the Defocus Spread Effect (DSE) degrades pixels in all source images around the boundaries. To further refine these ambiguous areas, we introduce a Similarity Selection Strategy, which reconstructs an initial clear image from source images and selects optimal pixels by comparing the similarity among them. Through this interactive approach, CSNet effectively preserves focused regions as well as recovering natural boundaries to fuse an all-in-focus output. Extensive experiments demonstrate that our method achieves state-of-the-art performance both quantitatively and qualitatively. Our code is available on Github: https://github.com/ZYC-HUST/CSNet.
Chinese Translation
多焦点图像融合(MFIF)旨在从同一场景的多幅图像中生成一幅全焦图像,这些图像在不同区域聚焦。现有的大多数基于深度学习的方法缺乏源图像之间的明确交互,这限制了它们的性能和可解释性。本文提出了一种新颖的清晰度对比与相似性选择网络(CSNet),以实现MFIF中的直接信息交换。具体而言,通过在我们提出的清晰度对比注意模块(CCAM)中对源图像之间的清晰度差异进行对比,我们相互增强了清晰特征,同时抑制模糊特征。这使我们能够识别每个源图像中确切的聚焦区域,并定位聚焦-失焦边界。此外,失焦扩散效应(DSE)会降低所有源图像中边界周围像素的质量。为了进一步细化这些模糊区域,我们引入了一种相似性选择策略,该策略通过比较源图像之间的相似性重建初始清晰图像,并选择最佳像素。通过这种交互式方法,CSNet有效地保留了聚焦区域,并恢复了自然边界,以融合出一幅全焦输出。大量实验表明,我们的方法在定量和定性上均达到了最先进的性能。我们的代码可在Github上获取:https://github.com/ZYC-HUST/CSNet。
cs.CV / 56 / 2608.03257
NanoMorph-3D: An End-to-End Physics-Driven Unrolling Framework for Nanomaterial Reconstruction
NanoMorph-3D:一种端到端物理驱动的纳米材料重建展开框架
Abstract
Precise 3D characterization of nanomaterials is essential for unlocking structure-property relationships. However, standard electron tomography is fundamentally limited by the missing wedge problem. Consequently, conventional algorithms suffer from severe geometric distortions, a challenge further complicated by pervasive noise interference. Current learning-based methods either rely on physics-blind post-processing or employ end-to-end architectures constrained by local receptive fields, failing to capture complex 3D topologies. We propose NanoMorph-3D, a unified end-to-end framework grounded in a comprehensive Nanomorphological Taxonomy. Powered by a large-scale synthetic dataset explicitly modeling non-linear electron attenuation, we design a Physics-Driven Unrolled Network mapping proximal gradient descent into a learnable architecture. To capture complex internal topologies, we formulate a hierarchical attention mechanism with Physics-Normalization for long-range 3D dependencies and scale invariance. Crucially, our Dual-Domain strategy leverages Sinusoidal Attention to explicitly model physical projection trajectories, enforcing strict sinogram consistency to mitigate missing wedge artifacts. Finally, an unsupervised dual-stream mechanism bridges the simulation-to-reality gap. Experiments demonstrate NanoMorph-3D reconstructs diverse topologies with superior fidelity and speed.
Chinese Translation
纳米材料的精确三维表征对于揭示结构-性质关系至关重要。然而,标准电子断层成像在根本上受到缺失楔形问题的限制。因此,传统算法遭受严重的几何失真,这一挑战因普遍的噪声干扰而更加复杂。目前的基于学习的方法要么依赖于无物理知识的后处理,要么采用受限于局部感受野的端到端架构,未能捕捉复杂的三维拓扑。我们提出了NanoMorph-3D,一个基于全面的纳米形态分类法的统一端到端框架。该框架利用一个大规模合成数据集,明确建模非线性电子衰减,设计了一个物理驱动的展开网络,将近端梯度下降映射到可学习的架构中。为了捕捉复杂的内部拓扑,我们制定了一个具有物理归一化的层次注意机制,以处理长距离的三维依赖关系和尺度不变性。关键是,我们的双域策略利用正弦注意机制明确建模物理投影轨迹,强制执行严格的正弦图一致性,以减轻缺失楔形伪影。最后,一个无监督的双流机制弥合了模拟与现实之间的差距。实验表明,NanoMorph-3D能够以更高的保真度和速度重建多样的拓扑结构。
cs.CV / 57 / 2608.03269
Efficient Video Dataset Distillation via Cluster-Guided Prototype Blending
通过聚类引导的原型混合实现高效视频数据集蒸馏
Abstract
Video dataset distillation aims to compress a large video dataset into a compact surrogate set that preserves its training utility. Most existing approaches synthesize condensed videos through iterative optimization, whose cost is amplified by the temporal dimension. Rather than further reducing the number of optimized variables, we investigate whether effective distilled videos can be constructed without gradient-based optimization of the stored videos. Such a construction-based approach must address three challenges: selecting informative temporal segments, covering diverse intra-class variations under a limited videos-per-class budget, and increasing the information carried by each stored sample. To this end, we propose ProtoBlend, an efficient select-allocate-blend framework. First, teacher-guided temporal clip selection retains a high-confidence segment from each source video. Second, cluster-guided prototype allocation partitions the selected clips in the teacher feature space and assigns one distilled slot to each intra-class cluster. Third, each prototype is blended with an in-cluster anchor, while their teacher predictions are combined using the same coefficient to provide mixture-source supervision. Experiments on four trimmed action-recognition benchmarks demonstrate that ProtoBlend achieves a competitive accuracy-efficiency trade-off without iterative optimization of the distilled videos.
Chinese Translation
视频数据集蒸馏旨在将大型视频数据集压缩为一个紧凑的替代集,以保留其训练效用。现有大多数方法通过迭代优化合成浓缩视频,其成本因时间维度而增加。我们并不打算进一步减少优化变量的数量,而是探讨是否可以在不对存储视频进行基于梯度的优化的情况下构建有效的蒸馏视频。这种基于构建的方法必须解决三个挑战:选择信息丰富的时间片段、在有限的每类视频预算下覆盖多样的类内变异,以及增加每个存储样本所携带的信息。为此,我们提出了ProtoBlend,一个高效的选择-分配-混合框架。首先,教师引导的时间片段选择从每个源视频中保留一个高置信度的片段。其次,聚类引导的原型分配在教师特征空间中对所选片段进行划分,并为每个类内聚类分配一个蒸馏槽。第三,每个原型与类内锚点进行混合,同时使用相同的系数结合它们的教师预测,以提供混合源监督。在四个裁剪的动作识别基准上的实验表明,ProtoBlend在不对蒸馏视频进行迭代优化的情况下,实现了竞争性的准确性与效率的权衡。
cs.CV / 58 / 2608.03270
GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs
GUI-Lens:基于通用视觉语言模型的粗到细图形用户界面定位
Abstract
GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents. The task remains difficult on high-resolution, densely populated interfaces because a vision-language model (VLM) may recognize a requested control without locating it precisely enough for interaction. Most existing methods provide various forms of localization assistance, but still rely on a direct click prediction, allowing visual ambiguity or an inaccurate initial estimate to propagate to the final result. In this paper, we introduce GUI-Lens, a coarse-to-fine grounding framework that allows a general-purpose VLM to determine the target through active visual observations. Specifically, GUI-Lens extracts OCR text and detected UI components from the screenshot and presents their positions as coordinate references. Using the instruction, the current view, and these references, the VLM selects the region and scale of the next view, which is cropped and enlarged to provide finer visual details. This process continues over successively focused views until the target is determined. Proposed crops and clicks are checked against the instruction throughout the process, and the final local position is mapped back to the original screen coordinates. Experiments on four GUI grounding benchmarks and three general-purpose VLM backends show that GUI-Lens improves overall grounding accuracy by up to 24.9 percentage points and achieves state-of-the-art performance with GPT-5.5.
Chinese Translation
图形用户界面(GUI)定位将自然语言指令映射到点击位置,对于可靠的GUI代理至关重要。然而,在高分辨率、密集的界面上,这一任务仍然困难,因为视觉语言模型(VLM)可能识别出请求的控件,但无法精确定位以进行交互。现有的大多数方法提供了各种形式的定位辅助,但仍依赖于直接的点击预测,这使得视觉模糊或不准确的初始估计可能传播到最终结果。在本文中,我们提出了GUI-Lens,一个粗到细的定位框架,允许通用VLM通过主动的视觉观察来确定目标。具体而言,GUI-Lens从屏幕截图中提取OCR文本和检测到的用户界面组件,并将其位置作为坐标参考。利用指令、当前视图和这些参考,VLM选择下一个视图的区域和比例,并对其进行裁剪和放大,以提供更精细的视觉细节。这个过程在连续聚焦的视图中持续进行,直到确定目标。在整个过程中,提出的裁剪和点击与指令进行核对,最终的局部位置被映射回原始屏幕坐标。在四个GUI定位基准和三个通用VLM后端上的实验表明,GUI-Lens将整体定位准确率提高了多达24.9个百分点,并在GPT-5.5上实现了最先进的性能。
cs.CV / 59 / 2608.03279
3DGSI-Assessor: A Large-Scale Dataset and An LMM-based Method for 3D Gaussian Splatting Image Quality Assessment
3DGSI-Assessor:用于3D高斯点云图像质量评估的大规模数据集和基于LMM的方法
Abstract
3D Gaussian Splatting (3DGS) has become a dominant representation for real-time novel view synthesis (NVS), yet its storage footprint makes compression indispensable for practical deployment. 3DGS training and compression introduce representation-specific distortions such as floating artifacts and surface scattering, which conventional image quality assessment (IQA) metrics fail to capture. Moreover, the independent compression of geometric and color attributes may lead to decoupled dimension-specific distortions that must be diagnosed separately, yet existing metrics report only a single overall score. To address these gaps, we present 3DGS-IEval-15K+, a large-scale, multi-dimensional IQA dataset for compressed 3DGS, comprising 15,200 images from 10 diverse scenes, produced by 6 representative 3DGS algorithms at systematically designed compression levels and rendered from 20 strategically selected viewpoints spanning both training views and challenging novel views, annotated with 45,600 mean opinion scores (MOSs) across overall, geometry, and color quality. Based on 3DGS-IEval-15K+, we propose 3DGSI-Assessor, an all-in-one 3DGS IQA framework that integrates global semantic and dimension-specific local features within a large multimodal model (LMM), predicting all three dimensions in a single forward pass. 3DGSI-Assessor achieves state-of-the-art performance on 3DGS-IEval-15K+, and exhibits competitive generalization on other NVS benchmarks. Dataset and code will be released at https://github.com/YukeXing/3DGSI-Assessor.
Chinese Translation
3D高斯点云(3DGS)已成为实时新视图合成(NVS)的主要表示方法,但其存储占用使得压缩在实际应用中不可或缺。3DGS的训练和压缩引入了特定表示的失真,例如浮动伪影和表面散射,而传统的图像质量评估(IQA)指标无法捕捉这些失真。此外,几何属性和颜色属性的独立压缩可能导致需要单独诊断的解耦维度特定失真,而现有指标仅报告单一的整体得分。为了解决这些问题,我们提出了3DGS-IEval-15K+,这是一个针对压缩3DGS的大规模多维IQA数据集,包含来自10个不同场景的15,200张图像,这些图像由6种代表性的3DGS算法在系统设计的压缩级别下生成,并从20个战略选择的视点渲染,涵盖训练视图和具有挑战性的新的视图,标注了45,600个平均意见分数(MOS),涵盖整体、几何和颜色质量。基于3DGS-IEval-15K+,我们提出了3DGSI-Assessor,这是一个集成了全局语义和维度特定局部特征的大型多模态模型(LMM)的全能3DGS IQA框架,能够在单次前向传播中预测所有三个维度。3DGSI-Assessor在3DGS-IEval-15K+上实现了最先进的性能,并在其他NVS基准上展现出竞争力的泛化能力。数据集和代码将发布在 https://github.com/YukeXing/3DGSI-Assessor。
cs.CV / 60 / 2608.03284
Test-Time Scaling for Safe Text-Guided Image Generation via Intermediate Clean Estimates
通过中间干净估计进行安全文本引导图像生成的测试时间缩放
Abstract
Ensuring safety and policy compliance in text-to-image diffusion models remains a critical challenge, as benign or adversarial prompts can often elicit prohibited content, e.g. nudity and protected intellectual property. While training-based unlearning methods are effective, they are computationally expensive and prone to catastrophic interference with general capabilities. Conversely, existing test-time defenses are primarily prompt-centric, relying on modifying textual descriptions only, and overlook the visual signals for detection. In this paper, we propose to leverage the intermediate clean image estimated during the generation process and employ a sparse margin objective to detect prohibited concepts. When a violation is detected, we immediately intervene by optimizing a structured low-rank residual in the text-conditioning space via truncated backpropagation. This design allows weight-preserving detection, keeps non-violating inference latency nearly unchanged as the maximum budget increases, and offers flexibility in safety performance via test-time scaling. Extensive experiments on Stable Diffusion v1.4 and v3.5 across nudity removal, IP protection, and style erasure demonstrate superior performance across suppression, fidelity and preservation compared to prior weight-preserving baselines, providing a scalable and flexible solution for safe generative deployment.
Chinese Translation
确保文本到图像扩散模型的安全性和政策合规性仍然是一个关键挑战,因为良性或对抗性提示往往会引发禁止内容,例如裸体和受保护的知识产权。虽然基于训练的去学习方法有效,但它们计算成本高且容易对一般能力造成灾难性干扰。相反,现有的测试时间防御主要集中在提示上,仅依赖于修改文本描述,而忽视了用于检测的视觉信号。本文提出利用生成过程中估计的中间干净图像,并采用稀疏边际目标来检测禁止概念。当检测到违规时,我们立即通过截断反向传播优化文本条件空间中的结构化低秩残差进行干预。该设计允许权重保持检测,在最大预算增加时几乎保持不违反推理延迟不变,并通过测试时间缩放提供安全性能的灵活性。在对Stable Diffusion v1.4和v3.5进行的广泛实验中,针对裸体去除、知识产权保护和风格抹除的表现优于先前的权重保持基线,在抑制、保真度和保留方面提供了可扩展和灵活的安全生成部署解决方案。
cs.CV / 61 / 2608.03304
Recurrent Contrastive Learning for Imbalanced Medical Image Classification
用于不平衡医学图像分类的递归对比学习
Abstract
Medical image classification often suffers from class imbalance due to the inherent disparities in disease incidence. Existing approaches, such as class resampling and loss reweighting, mainly improve learning within the observed feature distribution, but do not explicitly enlarge the latent support region of tail classes. As a result, tail-class representations remain overly compact and are easily encroached upon by head classes, leading to biased decision boundaries. In this work, we propose Recurrent Contrastive Learning (RCL) for imbalanced medical image classification. RCL progressively expands the support region of tail classes by recurrently reusing historical feature states across training phases. Specifically, we adopt DINOv3 with LoRA adapters as the backbone to provide robust feature embeddings. We then devise a Temporal Memory Queue (TMQ) to preserve corpus-level features across training phases and provide diversified global references for contrastive learning. Based on TMQ, we construct Temporal Anchors (TARs) to form an anchor field around tail classes. This field enlarges the support region of tail classes, suppresses head-class encroachment, and improves inter-class separation. Extensive experiments on three imbalanced medical datasets demonstrate that RCL achieves consistent improvements over strong baselines. The code is available at https://github.com/dndins/RCL.
Chinese Translation
医学图像分类常常因疾病发生率的固有差异而遭受类别不平衡的困扰。现有方法,如类别重采样和损失重加权,主要改善观察到的特征分布内的学习,但并未明确扩大尾类的潜在支持区域。因此,尾类表示依然过于紧凑,容易受到头类的侵蚀,导致决策边界偏差。在本研究中,我们提出了一种用于不平衡医学图像分类的递归对比学习(Recurrent Contrastive Learning, RCL)。RCL通过在训练阶段递归地重用历史特征状态,逐步扩展尾类的支持区域。具体而言,我们采用DINOv3与LoRA适配器作为骨干网络,以提供稳健的特征嵌入。然后,我们设计了一个时间记忆队列(Temporal Memory Queue, TMQ)来在训练阶段之间保留语料级特征,并为对比学习提供多样化的全局参考。基于TMQ,我们构建了时间锚点(Temporal Anchors, TARs)以在尾类周围形成一个锚场。该锚场扩大了尾类的支持区域,抑制了头类的侵蚀,并改善了类间分离。在三个不平衡医学数据集上的大量实验表明,RCL在强基线之上实现了一致的改进。代码可在 https://github.com/dndins/RCL 获取。
cs.CV / 62 / 2608.03322
LocAnyMed: Vision-Language Grounding for Multimodal Medical Images
LocAnyMed:多模态医学图像的视觉-语言基础
Abstract
Medical visual grounding connects free-form clinical queries to spatial evidence in medical images and is an important component of interpretable medical artificial intelligence. However, general-purpose grounding models are predominantly trained on natural images, while existing medical localization resources remain fragmented across imaging modalities, datasets, and task formulations. To address this gap, we construct LocAnyMed-200K, a multimodal medical visual grounding dataset containing approximately 200K image-query-answer examples across computed tomography, optical medical imaging, ultrasound, and X-ray. We harmonize heterogeneous detection and localization resources into a unified free-form instruction format that supports one or multiple bounding boxes, point coordinates, and no-target outputs for negative queries. Full-parameter fine-tuning of LocateAnything-3B on LocAnyMed-200K improves F1@IoU 0.50 from 10.64 to 85.59 on a held-out evaluation split, demonstrating that large-scale domain-specific supervision can equip a general grounding model with effective medical localization capabilities. Beyond spatial coordinates, a clinically interpretable grounding system should also communicate the evidence supporting its prediction. We therefore derive LocAnyMed-CoT-20K, a rationale-augmented subset that connects anatomical context, visual observations, and spatial conclusions through structured reasoning and further improves cross-source generalization through fine-tuning. Together, these resources provide a unified foundation for studying both localization accuracy and rationale quality across heterogeneous medical imaging modalities. The code is publicly available at https://github.com/MiliLab/LocAnyMed.
Chinese Translation
医学视觉基础将自由形式的临床查询与医学图像中的空间证据连接起来,是可解释医学人工智能的重要组成部分。然而,通用基础模型主要在自然图像上进行训练,而现有的医学定位资源在成像模态、数据集和任务表述上仍然存在碎片化的问题。为了解决这一差距,我们构建了LocAnyMed-200K,这是一个多模态医学视觉基础数据集,包含约20万个图像-查询-答案示例,涵盖计算机断层扫描、光学医学成像、超声波和X射线。我们将异构的检测和定位资源统一为一种支持一个或多个边界框、点坐标和负查询无目标输出的自由形式指令格式。在LocAnyMed-200K上对LocateAnything-3B进行全参数微调,使得F1@IoU 0.50从10.64提高到85.59,证明了大规模领域特定的监督可以赋予通用基础模型有效的医学定位能力。除了空间坐标,临床可解释的基础系统还应传达支持其预测的证据。因此,我们推导出LocAnyMed-CoT-20K,这是一个增强推理的子集,通过结构化推理连接解剖背景、视觉观察和空间结论,并通过微调进一步提高跨源泛化能力。这些资源共同为研究异构医学成像模态下的定位准确性和推理质量提供了统一的基础。代码可在 https://github.com/MiliLab/LocAnyMed 上公开获取。
cs.CV / 63 / 2608.03323
PolyLayout: Multi-room Manhattan Layout Estimation
PolyLayout:多房间曼哈顿布局估计
Abstract
Estimating room layouts from multi-view imagery is a core task for indoor scene understanding. Existing methods are typically limited either by poor generalization to new datasets or restrictive geometric assumptions of the room shape or camera configuration. Most also estimate rooms independently, failing to exploit shared building structure such as dominant directions, ground plane or ceiling height. We propose PolyLayout, a multi-room layout estimation method that parameterizes room layouts as Manhattan 3D polygons and optimizes them jointly across multiple rooms. The optimization objective is predicted by a neural network on top of robust pre-trained visual features and trained end-to-end with supervision only on output room layouts. At the same time, camera projection and polygon updates remain explicit and model-based. This separation between learned scoring and geometry improves generalization to new datasets and camera parameters. During optimization, PolyLayout adaptively refines the polygon topology through iterative wall split and merge operations while jointly utilizing structural cues across rooms. We introduce two new multi-view multi-room layout benchmarks by providing layout annotations to existing datasets, and experiments show that PolyLayout outperforms prior approaches, both in terms of accuracy and robustness. Project page: https://ghanning.github.io/PolyLayout
Chinese Translation
从多视角图像中估计房间布局是室内场景理解的核心任务。现有方法通常受到对新数据集的泛化能力差或对房间形状或相机配置的限制性几何假设的限制。大多数方法还独立估计房间,未能利用共享的建筑结构,如主导方向、地面平面或天花板高度。我们提出了PolyLayout,一种将房间布局参数化为曼哈顿三维多边形并在多个房间之间联合优化的多房间布局估计方法。优化目标由神经网络预测,基于稳健的预训练视觉特征,并仅在输出房间布局上进行端到端的监督训练。同时,相机投影和多边形更新保持显式和基于模型。这种学习评分与几何之间的分离提高了对新数据集和相机参数的泛化能力。在优化过程中,PolyLayout通过迭代的墙体分割和合并操作自适应地细化多边形拓扑,同时联合利用跨房间的结构线索。我们通过为现有数据集提供布局注释,介绍了两个新的多视角多房间布局基准,实验表明PolyLayout在准确性和鲁棒性方面均优于先前的方法。项目页面:https://ghanning.github.io/PolyLayout
cs.CV / 64 / 2608.03335
SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference
SPADE:一种输入自适应稀疏注意力引擎,用于快速视频扩散模型推理
Abstract
Video diffusion transformers (vDiTs) generate high quality but pay quadratic self-attention cost, making inference prohibitive at video-token scales. The challenge is input-adaptive sparsity: selecting critical Q/K/V tokens with negligible overhead and executing them for end-to-end gains. We present SPADE, a training-free sparse-attention engine of three parts: (i) vDiT-SSR, a specification defining 3D blocking candidates and formalizing dynamic masks via Summarizer/Estimator expressions; (ii) runtime scheme generation using SICS and a head-wise policy; and (iii) an executor with low-overhead index search, flash block-sparse attention, and kernel grouping. Across Hunyuan-Video and Wan 2.1/2.2 for text-to-video and image-to-video generation, SPADE raises sparsity and speed while preserving quality, accelerating attention by 2.26x-3.40x and end-to-end inference by 1.49x-1.80x. Our code is open-sourced at https://github.com/6somehow/DAC-SPADE.
Chinese Translation
视频扩散变换器(vDiTs)能够生成高质量的视频,但其自注意力的计算成本呈二次增长,使得在视频标记规模下的推理变得不可行。面临的挑战是输入自适应稀疏性:选择关键的 Q/K/V 标记,同时保持可忽略的开销,并执行它们以实现端到端的增益。我们提出了 SPADE,这是一种无训练的稀疏注意力引擎,由三个部分组成:(i)vDiT-SSR,一种定义 3D 阻塞候选项并通过总结器/估计器表达式形式化动态掩码的规范;(ii)使用 SICS 和头部策略生成运行时方案;(iii)一个具有低开销索引搜索、闪存块稀疏注意力和内核分组的执行器。在 Hunyuan-Video 和 Wan 2.1/2.2 数据集上进行文本到视频和图像到视频生成时,SPADE 提高了稀疏性和速度,同时保持了质量,加速注意力计算 2.26 倍至 3.40 倍,端到端推理加速 1.49 倍至 1.80 倍。我们的代码已开源,地址为 https://github.com/6somehow/DAC-SPADE。
cs.CV / 65 / 2608.03342
When Oracle Conditioning Misleads Deployment: Conditioning-Availability Bias in Echocardiographic Segmentation
当Oracle条件化误导部署:超声心动图分割中的条件可用性偏差
Abstract
Conditional segmentation models may be trained and evaluated with auxiliary signals cleaner than those available at deployment. We study this protocol-level manifestation of shortcut learning and auxiliary-variable shift in phase-conditioned echocardiographic segmentation. The complementary gap pair measures loss on the deployable oracle-estimated pathway and probes sensitivity on the oracle-random pathway. On held-out CAMUS data, one strong-cyclic, oracle-selected run fails severely with estimated phase, while sensitivity to incorrect phase persists across three runs. On EchoNet-Dynamic, the current estimator remains usable, but random-phase testing reveals strong latent sensitivity. Deployment-aware checkpoint selection and phase perturbation reduce both gaps with little change in mean Dice. Exploratory subgroup analyses quantify variation across measured strata, and a downstream ejection fraction (EF) audit shows that recovering segmentation does not necessarily recover EF error or signed bias. Together, the gaps test whether oracle-conditioned performance survives the inference pathway actually available at deployment.
Chinese Translation
条件分割模型可能在训练和评估时使用的辅助信号比部署时可用的信号更为干净。我们研究了这种在相位条件下的超声心动图分割中,快捷学习和辅助变量转移的协议级表现。互补差距对测量可部署的oracle估计路径上的损失,并探测oracle随机路径上的敏感性。在保留的CAMUS数据上,一个强周期的oracle选择运行在估计相位时严重失败,而对错误相位的敏感性在三次运行中持续存在。在EchoNet-Dynamic上,当前估计器仍然可用,但随机相位测试揭示了强烈的潜在敏感性。考虑部署的检查点选择和相位扰动在均值Dice变化不大的情况下减少了这两个差距。探索性亚组分析量化了测量层次之间的变异,而下游的射血分数(EF)审计显示恢复的分割不一定能恢复EF误差或符号偏差。总体而言,这些差距测试了oracle条件下的性能是否能在实际可用的推理路径上存活。
cs.CV / 66 / 2608.03357
Can Text-to-Image Models Draw from the Right Frame of Reference?
文本到图像模型能否从正确的参考框架中绘制?
Abstract
Spatial instruction following has become a crucial requirement for text-to-image (T2I) generation. A common challenge arises when directional expressions are interpreted under different frames of reference. For example, ``the left of'' may refer to the viewer's image coordinates or to the intrinsic orientation of an object, leading to different expected layouts. Existing T2I benchmarks reveal important layout failures, yet they rarely isolate whether models can follow a specified frame of reference when it differs from camera view. To mitigate this gap, we introduce FoR-T2I, a benchmark for evaluating this distinction with 1,200 prompt pairs built from controlled spatial layouts. In each pair, the camera-view (Cam) prompt states the target relation in camera view, while the frame-of-reference (FoR) prompt describes the same target placement through an oriented anchor object. Across 22 closed-source and open-source T2I models, mean final accuracy is 41.8\% lower on FoR prompts than on matched Cam prompts; even the best-performing model achieves only 44.3\% FoR accuracy. This suggests that current models struggle more when the same layout is described through an object's orientation rather than directly in image coordinates. We further analyze this gap by relation type and camera view, compare several training-free prompting and feedback-based mitigation strategies, and propose a VLM-gated rewriting approach that selects rewritten prompts using visual feedback, improving average FoR accuracy from 25.0\% to 29.2\% under the same generation budget.
Chinese Translation
空间指令跟随已成为文本到图像(T2I)生成的重要要求。当方向表达在不同的参考框架下被解释时,常常会出现挑战。例如,“左侧”可能指的是观察者的图像坐标或对象的内在方向,导致不同的预期布局。现有的T2I基准测试揭示了重要的布局失败,但很少单独评估模型在参考框架与相机视图不同时是否能够遵循指定的参考框架。为了解决这一问题,我们引入了FoR-T2I,这是一个用于评估这种区分的基准,包含1200对基于受控空间布局构建的提示。在每对提示中,相机视图(Cam)提示在相机视图中陈述目标关系,而参考框架(FoR)提示通过一个定向锚对象描述相同的目标位置。在22个闭源和开源的T2I模型中,FoR提示的平均最终准确率比匹配的Cam提示低41.8%;即使是表现最好的模型,其FoR准确率也仅为44.3%。这表明当前模型在通过对象的方向而非直接在图像坐标中描述相同布局时更为困难。我们进一步按关系类型和相机视图分析这一差距,比较几种无训练的提示和基于反馈的缓解策略,并提出了一种VLM门控重写方法,该方法通过视觉反馈选择重写提示,使得在相同生成预算下,平均FoR准确率从25.0%提高到29.2%。
cs.CV / 67 / 2608.03370
DRPFNet: Dual-domain Residual Progressive Fusion Network for RGB-Thermal Object Detection
DRPFNet:用于RGB-热成像目标检测的双域残差渐进融合网络
Abstract
RGB-thermal (RGB-T) object detection aims to fuse complementary information from visible and thermal modalities to achieve robust detection under varying illumination and weather conditions. Current methods typically employ attention mechanisms or transformers to perform cross-modal fusion independently at each feature scale, directly combining RGB and thermal features in the spatial domain. However, they still face significant limitations: cross-level knowledge inheritance caused by independent fusion at each scale,suppressing noise continuously due to the lack of bidirectional optimization, and information degradation induced by the absence of frequency-spatial collaboration. To address these issues, we propose DRPFNet, a Dual-domain Residual Progressive Fusion Network that constructs a unified information flow optimization system from three synergistic levels:structure, feature, and enhancement. At the structural level, we establish cross-scale propagation through bottom-up knowledge accumulation and bidirectional enhancement,ensuring smooth information flow. At the feature level, we collaboratively extract RGB high-frequency edges and thermal low-frequency structures via frequency band separation and edge guidance, guaranteeing representation quality. At the enhancement level, we enhance foreground-background discrimination through edge-guided dual-domain refinement,achieving precise object localization.Extensive experiments on two public RGB-T datasets demonstrate that our method achieves competitive performance with competitive efficiency, validating the effectiveness of this hierarchical collaborative strategy.
Chinese Translation
RGB-热成像(RGB-T)目标检测旨在融合来自可见光和热成像模态的互补信息,以在不同的光照和天气条件下实现稳健的检测。目前的方法通常采用注意力机制或变换器,在每个特征尺度上独立执行跨模态融合,直接在空间域中结合RGB和热成像特征。然而,它们仍面临显著的局限性:由于在每个尺度上独立融合导致的跨层知识继承、由于缺乏双向优化而持续抑制噪声,以及由于缺乏频率-空间协作而引起的信息降解。为了解决这些问题,我们提出了DRPFNet,一种双域残差渐进融合网络,构建了一个统一的信息流优化系统,涵盖三个协同层次:结构、特征和增强。在结构层面,我们通过自下而上的知识积累和双向增强建立跨尺度传播,确保信息流的顺畅。在特征层面,我们通过频带分离和边缘引导协同提取RGB高频边缘和热成像低频结构,确保表示质量。在增强层面,我们通过边缘引导的双域精细化增强前景-背景区分,实现精确的目标定位。在两个公共RGB-T数据集上的大量实验表明,我们的方法在性能和效率上具有竞争力,验证了这种分层协作策略的有效性。
cs.CV / 68 / 2608.03379
Residual Flow Matching with Dynamic Cross-Interaction for 3D Multi-Person Motion Prediction
动态交互的残差流匹配用于三维多人物运动预测
Abstract
3D multi-person motion prediction requires modeling both individual kinematics and inter-person interactions. While Flow Matching is effective for multi-hypothesis generation to improve prediction accuracy, directly predicting skeletal sequences from pure noise often compromises structural consistency and introduces unreliable cross-agent interactions during early noise-dominated integration steps. To address this, we propose a Prior-Guided Residual Flow Matching framework. First, a Deterministic Coarse Prior (DCP) establishes a kinematic anchor, formulating the generative process as a conditional flow over motion residuals to simplify the generative objective and preserve structural stability. Second, a Dynamic Cross-Interaction (DCI) mechanism temporally synchronizes inter-agent message-passing with the integration progress, ensuring the extraction of reliable social contexts and improving multi-person motion fidelity. Finally, a decoupled joint-motion architecture with bidirectional fusion effectively preserves fine-grained kinematic coherence. Extensive experiments demonstrate that our approach achieves state-of-the-art prediction accuracy across multiple datasets. Code is available at https://github.com/Wei-Wei-a/Residual-Flow-Matching-with-Dynamic-Cross-Interaction-for-3D-Multi-Person-Motion-Prediction.
Chinese Translation
三维多人物运动预测需要同时建模个体运动学和人物间的交互。虽然流匹配(Flow Matching)在多假设生成方面有效,以提高预测准确性,但直接从纯噪声中预测骨骼序列往往会妨碍结构一致性,并在早期噪声主导的整合步骤中引入不可靠的跨代理交互。为了解决这个问题,我们提出了一种先验引导的残差流匹配框架。首先,确定性粗略先验(Deterministic Coarse Prior, DCP)建立了运动学锚点,将生成过程表述为对运动残差的条件流,从而简化生成目标并保持结构稳定性。其次,动态交互机制(Dynamic Cross-Interaction, DCI)在时间上同步代理间的信息传递与整合进程,确保提取可靠的社会上下文并提高多人物运动的保真度。最后,具有双向融合的解耦联合运动架构有效地保持了细粒度的运动学一致性。大量实验表明,我们的方法在多个数据集上实现了最先进的预测准确性。代码可在 https://github.com/Wei-Wei-a/Residual-Flow-Matching-with-Dynamic-Cross-Interaction-for-3D-Multi-Person-Motion-Prediction 获取。
cs.CV / 69 / 2608.03385
FreqAdapt: Frequency-Adaptive Processing for RAW Object Detection
FreqAdapt:用于RAW目标检测的频率自适应处理
Abstract
Existing object detection methods predominantly utilize sRGB inputs, which are compressed from RAW sensor data using Image Signal Processors (ISP) originally designed for visualization purposes. Compared to RGB images, RAW images possess favorable noise characteristics and richer information representation, which are crucial for object detection, particularly under challenging conditions such as adverse weather or low-light environments. In this paper, we propose FreqAdapt, a lightweight module for adaptive RAW data enhancement in the frequency domain. Unlike traditional spatial domain processing methods, FreqAdapt innovatively maps ISP operations to the Fourier frequency domain and performs domain separation based on the physical properties of ISP operations, ensuring each operation is performed in its most suitable domain. Meanwhile, through an adaptive frequency domain encoder that jointly analyzes amplitude spectrum, phase spectrum, and RAW image features, we provide global context for ISP parameter prediction and employ a learnable fusion mechanism to achieve adaptive feature enhancement. Extensive experiments on multiple datasets with diverse lighting and weather conditions demonstrate that FreqAdapt achieves state-of-the-art performance while maintaining lightweight efficiency and good physical interpretability. Furthermore, our module can be seamlessly incorporated into existing object detection frameworks, providing a novel solution for visual perception tasks in the RAW domain.
Chinese Translation
现有的目标检测方法主要使用sRGB输入,这些输入是通过图像信号处理器(ISP)从RAW传感器数据压缩而来的,ISP最初是为可视化目的设计的。与RGB图像相比,RAW图像具有更好的噪声特性和更丰富的信息表示,这对于目标检测至关重要,尤其是在恶劣天气或低光环境等挑战性条件下。在本文中,我们提出了FreqAdapt,一个用于频域自适应RAW数据增强的轻量级模块。与传统的空间域处理方法不同,FreqAdapt创新性地将ISP操作映射到傅里叶频域,并基于ISP操作的物理特性进行域分离,确保每个操作在其最适合的域中执行。同时,通过一个自适应频域编码器,该编码器联合分析幅度谱、相位谱和RAW图像特征,我们为ISP参数预测提供了全局上下文,并采用可学习的融合机制实现自适应特征增强。在多个具有不同光照和天气条件的数据集上进行的广泛实验表明,FreqAdapt在保持轻量高效和良好物理可解释性的同时,达到了最先进的性能。此外,我们的模块可以无缝集成到现有的目标检测框架中,为RAW域的视觉感知任务提供了一种新颖的解决方案。
cs.CV / 70 / 2608.03395
SRAP: SVD-Refined Adversarial Perturbations for Imperceptible Face-Swap Defense
SRAP:用于隐形人脸交换防御的奇异值分解精炼对抗扰动
Abstract
Deepfake technologies pose increasing threats to facial privacy and identity security, motivating proactive defenses that protect facial images before misuse. Although adversarial perturbations generated by projected gradient descent (PGD) can disrupt the identity representations used by face-swapping models, their visual quality is degraded by two characteristics: perturbations are distributed broadly over the image, including identity-insensitive regions, and they contain visually salient high-frequency components. We analyze these spatial and spectral inefficiencies through identity-sensitivity estimation and the singular-value decomposition (SVD) of PGD perturbations. Our analysis shows that later singular components contain a disproportionate amount of high-frequency energy, while the leading components preserve most of the perturbation energy and defense utility. Based on these observations, we propose SRAP, which combines per-channel truncated SVD refinement with an identity-importance mask at every optimization step. The SVD refinement suppresses high-rank, high-frequency residuals, while the mask restricts perturbations to locations that strongly influence identity representations. Experiments on CelebA-HQ and VGGFace2-HQ demonstrate that SRAP substantially improves protected-image fidelity across all reported metrics while maintaining competitive identity-disruption performance, yielding a favorable trade-off between face-swap defense and visual imperceptibility.
Chinese Translation
深度伪造技术对面部隐私和身份安全构成日益严重的威胁,这促使我们采取主动防御措施,以在滥用之前保护面部图像。尽管通过投影梯度下降(PGD)生成的对抗扰动可以干扰人脸交换模型使用的身份表示,但其视觉质量受到两个特征的影响:扰动广泛分布在图像上,包括身份不敏感区域,并且包含视觉上显著的高频成分。我们通过身份敏感性估计和PGD扰动的奇异值分解(SVD)分析了这些空间和频谱低效性。我们的分析表明,后续的奇异成分包含不成比例的高频能量,而前导成分则保留了大部分扰动能量和防御效用。基于这些观察,我们提出了SRAP,它在每个优化步骤中结合了每通道截断的SVD精炼和身份重要性掩模。SVD精炼抑制高秩、高频残差,而掩模则限制扰动仅作用于强烈影响身份表示的位置。在CelebA-HQ和VGGFace2-HQ上的实验表明,SRAP在所有报告的指标上显著提高了保护图像的保真度,同时保持了竞争力的身份干扰性能,实现了人脸交换防御与视觉隐形性之间的良好权衡。
cs.CV / 71 / 2608.03407
Distilled Roads: Generalisable Road Network Extraction Across Sensors, Resolutions, and Region
提炼道路:跨传感器、分辨率和区域的可推广道路网络提取
Abstract
Road network segmentation from satellite imagery remains challenging due to large geographic variation in road appearance, occlusions, and domain shifts introduced by differing resolutions and sensors. Existing models, typically trained under narrow resolution--region combinations, generalise poorly to unseen environments such as rural settings, regions with distinct road materials, or imagery from new satellite platforms, often producing broken or disconnected predictions. Adapting these models to new domains usually requires retraining or fine-tuning, which is costly and risks catastrophic forgetting. In this work, we reframe global road extraction as a continual adaptation problem rather than an architectural one. Our framework combines cross-resolution knowledge distillation across a resolution-decreasing curriculum, multi-sensor training, and topology-aware supervision, yielding a single model that generalises across $0.3-1.0$ m imagery from multiple satellite platforms across continents. On publicly available benchmarks, including City-Scale and Global-Scale, our model outperforms state-of-the-art results by up to $22$ F1 points and $15$ APLS points, while remaining the most efficient, with $3\times$ faster inference. Our results suggest that improved robustness across diverse sub-meter satellite imagery can be achieved through targeted training strategies, such as data curricula, distillation, and topology-aware losses, rather than increasingly complex architectures.
Chinese Translation
从卫星影像中进行道路网络分割仍然面临挑战,这主要是由于道路外观的地理差异、遮挡以及不同分辨率和传感器引入的领域转移。现有模型通常在狭窄的分辨率-区域组合下训练,难以在未见环境中进行推广,例如乡村环境、具有不同道路材料的区域或来自新卫星平台的影像,常常产生断裂或不连贯的预测。将这些模型适应于新领域通常需要重新训练或微调,这既成本高昂又存在灾难性遗忘的风险。在本研究中,我们将全球道路提取重新框定为一个持续适应问题,而非架构问题。我们的框架结合了跨分辨率知识蒸馏、逐步降低分辨率的课程、多传感器训练和拓扑感知监督,生成一个能够在来自多个卫星平台的 $0.3-1.0$ 米影像上进行推广的单一模型。在公开可用的基准测试中,包括城市规模和全球规模,我们的模型在 F1 分数上比最先进的结果提高了多达 $22$ 点,在 APLS 分数上提高了 $15$ 点,同时保持了最高的效率,推理速度提高了 $3 imes$。我们的结果表明,通过有针对性的训练策略,如数据课程、蒸馏和拓扑感知损失,可以在多样的亚米级卫星影像中实现更好的鲁棒性,而不是依赖于日益复杂的架构。
cs.CV / 72 / 2608.03410
Earth Embeddings
地球嵌入
Abstract
Earth observation is moving from foundation models that users must run themselves toward embedding products that package model feature outputs as reusable data without needing to download and process the imagery used to generate them. Earth embeddings are vectors that summarize locations, image patches, or pixels, letting users analyze compact features instead of repeatedly training or running large models on raw satellite imagery. This chapter explains the main types of Earth embeddings, from implicit location encoders to explicit patch and pixel products, and compares their coverage, resolution, dimensionality, storage cost, licenses, and reproducibility. We review their use in land cover and crop mapping, ecological and hazard modeling, socioeconomic prediction, and semantic search, with evidence on when embeddings improve on conventional features and when pooling, fusion, or spatial transfer limit performance. Two case studies show practical workflows for similarity search and land cover mapping. We close with guidance for choosing, evaluating, storing, compressing, and publishing embeddings, and with open problems in oceanic and atmospheric coverage, uncertainty, and benchmarking.
Chinese Translation
地球观测正从用户必须自行运行的基础模型转向嵌入产品,这些产品将模型特征输出打包为可重用的数据,而无需下载和处理用于生成这些数据的影像。地球嵌入是总结位置、图像块或像素的向量,使用户能够分析紧凑特征,而不是反复在原始卫星影像上训练或运行大型模型。本章解释了主要的地球嵌入类型,从隐式位置编码器到显式图块和像素产品,并比较它们的覆盖范围、分辨率、维度、存储成本、许可证和可重复性。我们回顾了它们在土地覆盖和作物制图、生态和灾害建模、社会经济预测以及语义搜索中的应用,提供了嵌入何时优于传统特征以及何时池化、融合或空间转移限制性能的证据。两个案例研究展示了相似性搜索和土地覆盖制图的实际工作流程。最后,我们提供了选择、评估、存储、压缩和发布嵌入的指导,并讨论了海洋和大气覆盖、不确定性和基准测试等开放问题。
cs.CV / 73 / 2608.03422
HyperbolicDiffusion: Sharp & Scalable Tiled Generation on the Hyperbolic Plane
超曲面扩散:在超曲面上锐利且可扩展的平铺生成
Abstract
Planar tiled diffusion denoises overlapping windows of one rectangular canvas. The hyperbolic plane has no such canvas, and its area grows exponentially with radius. We introduce HyperbolicDiffusion, a training-free method for generating finite visual fields directly on the hyperbolic plane H2. Our Hyperbolic Blooming Cover reduces window placement to a compact dynamic program that runs in seconds while providing strong theoretical guarantees. Permanent surface IDs form a shared latent canvas: a standard diffusion model denoises local windows, whose predictions are fused back onto H2. Because curvature causes residual disagreement and blur at multi-window junctions, a geometry-derived second stage re-noises and repairs precisely those regions. The resulting fields are sharp, reprojectable, and consistent across viewpoints, providing a prompt-driven generative counterpart to Escher's Circle Limit series.
Chinese Translation
平面平铺扩散对一个矩形画布的重叠窗口进行去噪。超曲面没有这样的画布,其面积随着半径呈指数增长。我们提出了HyperbolicDiffusion,这是一种无训练的方法,直接在超曲面H2上生成有限的视觉场。我们的超曲面绽放覆盖(Hyperbolic Blooming Cover)将窗口放置简化为一个紧凑的动态规划,运行时间仅需几秒,同时提供强有力的理论保证。永久表面ID形成一个共享的潜在画布:标准扩散模型对局部窗口进行去噪,其预测结果再融合回H2。由于曲率导致多窗口交界处的残余不一致和模糊,几何推导的第二阶段重新去噪并精确修复这些区域。最终生成的场景清晰、可重投影,并在不同视点之间保持一致,为埃舍尔的《圆极限》系列提供了一种基于提示驱动的生成对照。
cs.CV / 74 / 2608.03423
SGFormer: Structure-Guided Transformer for Robust Local Feature Matching
SGFormer:结构引导的变换器用于鲁棒的局部特征匹配
Abstract
Local feature matching is a fundamental component of photogrammetry, enabling accurate image correspondence critical for tasks such as 3D reconstruction, stereo mapping, and visual localization. While recent detector-free matching methods, like LoFTR, have advanced the field, the global features obtained by leveraging the global-range modeling capacity of the unconstrained attention mechanism compromise the model's attention to the salient structures in certain scenarios. This limitation leads to a phenomenon we define as attention divergence, wherein a portion of high-confidence matches are distributed outside the valid matching region (overlapping region), especially in scenes with large viewpoint variations. This occurs because similar features in irrelevant regions may receive equal weighting and consideration within the standard Transformer, limiting matching reliability in challenging photogrammetric environments. To address this issue in feature matching, we propose SGFormer (Structure-Guided Transformer), a novel structure-aware matching network that adaptively updates attention on features near salient structure in overlapping regions. SGFormer employs a semi-dense coarse-to-fine pipeline and incorporates the proposed Triple-Structure-Attention (TSA) module into the backbone net for extracting distinctive features. The TSA module utilizes shallow local features from early network layers to enhance the representation around salient structure, guiding subsequent transformer stages to intensify the model's focus on regions with salient structure across the global scope. SGFormer, thereby reinforcing attention to visually consistent areas while mitigating the influence of non-overlapping regions. Extensive experiments show that SGFormer significantly mitigates attention divergence and improves matching accuracy.
Chinese Translation
局部特征匹配是摄影测量的一个基本组成部分,它能够实现准确的图像对应关系,这对于三维重建、立体映射和视觉定位等任务至关重要。尽管最近的无检测器匹配方法,如LoFTR,推动了这一领域的发展,但利用无约束注意机制的全球范围建模能力所获得的全局特征在某些场景中妨碍了模型对显著结构的关注。这一局限性导致了我们定义的注意力偏离现象,其中一部分高置信度匹配分布在有效匹配区域(重叠区域)之外,尤其是在视角变化较大的场景中。这是因为在无关区域的相似特征可能在标准变换器中获得相同的权重和考虑,从而限制了在具有挑战性的摄影测量环境中的匹配可靠性。为了解决特征匹配中的这一问题,我们提出了SGFormer(结构引导的变换器),这是一种新颖的结构感知匹配网络,能够自适应地更新重叠区域内显著结构附近特征的注意力。SGFormer采用半稠密的粗到细管道,并将提出的三重结构注意力(Triple-Structure-Attention, TSA)模块集成到主干网络中,以提取独特特征。TSA模块利用来自网络早期层的浅层局部特征来增强显著结构周围的表示,引导后续的变换器阶段加大模型对全球范围内显著结构区域的关注。因此,SGFormer增强了对视觉一致区域的注意力,同时减轻了非重叠区域的影响。大量实验表明,SGFormer显著减轻了注意力偏离现象,并提高了匹配精度。
cs.CV / 75 / 2608.03428
OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet
OliveGemma:一个用于识别地中海和欧洲饮食的30亿视觉语言模型
Abstract
Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due to high intra-class variability and visually similar dishes. This study presents OliveGemma, a vision language model for recognising and reasoning about Mediterranean and European cuisine. Built on the open-weight PaliGemma-2-3B architecture, OliveGemma is fine-tuned with LoRA on a unified corpus of 17,340 images from three European research project datasets (MedGR, ODIN, and VIPPSTAR), reconciled into a vocabulary of 216 composed dish categories and paired with 102,642 instruction style question-answer items covering dish recognition, likely and visible ingredients, class boundary discrimination, visual evidence and overall visual food understanding. Under a 3-fold cross-validation scheme, OliveGemma achieves a top-1 accuracy of 92.96% +/- 0.91%, exceeding the strongest CNN baseline (DenseNet-121) by 7.31% and outperforming zero-shot frontier models with exact instructions and bounded classes including Gemini Flash 3 and 3.5, GPT-5.4 Mini, and Claude Haiku 4.6 by 8%, 46%, and 64% respectively. Furthermore, OliveGemma demonstrates competitive performance on Top-3 and Top-5 accuracy, being second best across CNNs and frontier models, surpassed only by DenseNet-121. In addition, OliveGemma achieves 90.79% +/- 1.3% Exact-Set on the likely ingredients of the food categories. These results demonstrate that PEFT adaptation of a small VLM can surpass substantially larger proprietary models on specialised food recognition. The model is publicly available at https://huggingface.co/JamesZar/OliveGemma-3B and the experiments and results can be found at https://github.com/tsiokris/OliveGemma.
Chinese Translation
基于图像的饮食评估为自我报告的食物日记提供了一种可扩展的替代方案,但由于类内变异性高和视觉上相似的菜肴,细粒度的食物识别仍然具有挑战性。本研究提出了OliveGemma,一个用于识别和推理地中海和欧洲美食的视觉语言模型。OliveGemma基于开放权重的PaliGemma-2-3B架构,并在来自三个欧洲研究项目数据集(MedGR、ODIN和VIPPSTAR)的17,340幅图像的统一语料库上通过LoRA进行了微调,整合为216个复合菜肴类别的词汇,并配对了102,642个涵盖菜肴识别、可能和可见成分、类别边界区分、视觉证据和整体视觉食物理解的指令风格问答项目。在三折交叉验证方案下,OliveGemma实现了92.96% +/- 0.91%的顶级准确率,超越了最强的CNN基线(DenseNet-121)7.31%,并在包含确切指令和有限类别的零-shot前沿模型(如Gemini Flash 3和3.5、GPT-5.4 Mini和Claude Haiku 4.6)中分别超出8%、46%和64%。此外,OliveGemma在Top-3和Top-5准确率上表现出竞争力,在CNN和前沿模型中排名第二,仅次于DenseNet-121。此外,OliveGemma在食物类别的可能成分上达到了90.79% +/- 1.3%的精确集。这些结果表明,小型视觉语言模型的PEFT适应可以在专业食物识别上超越大得多的专有模型。该模型已在https://huggingface.co/JamesZar/OliveGemma-3B公开发布,实验和结果可在https://github.com/tsiokris/OliveGemma找到。
cs.CV / 76 / 2608.03429
SLAMFormer-$\infty$: Infinite SLAM Transformer for Unbounded Frontend and Backend Processing
SLAMFormer-$ ext{∞}$:用于无界前端和后端处理的无限SLAM变换器
Abstract
We introduce the Infinite SLAM Transformer (SLAMFormer-$\infty$), the first geometric transformer capable of supporting both long-range frontend and backend processing without an explicit distance bound. Instead of relying on a first-frame-anchored formulation, SLAMFormer-$\infty$ employs memory conditions to define flexible coordinate systems and scales for input frames, enabling more expressive structural conditioning. Built upon this formulation, the frontend preserves efficient local computation, while the backend jointly optimizes long-range trajectories and scene geometry in a globally consistent manner. Experimental results demonstrate that SLAMFormer-$\infty$ achieves superior or highly competitive performance in both trajectory estimation and scene reconstruction across large-scale datasets. Notably, SLAMFormer-$\infty$ generalizes to extremely long trajectories, successfully operating on sequences exceeding $17\mathrm{km}$.
Chinese Translation
我们介绍了无限SLAM变换器(SLAMFormer-$ ext{∞}$),这是首个能够支持长距离前端和后端处理而不需要明确距离限制的几何变换器。SLAMFormer-$ ext{∞}$不依赖于第一帧锚定的公式,而是采用记忆条件来定义灵活的坐标系统和输入帧的尺度,从而实现更具表现力的结构条件。基于这一公式,前端保持高效的局部计算,而后端则以全局一致的方式共同优化长距离轨迹和场景几何。实验结果表明,SLAMFormer-$ ext{∞}$在大规模数据集上在轨迹估计和场景重建方面均实现了优越或高度竞争的性能。值得注意的是,SLAMFormer-$ ext{∞}$能够推广到极长的轨迹,成功处理超过$17 ext{km}$的序列。
cs.CV / 77 / 2608.03430
Dual-domain U-Nets with embedded back projection operators for motion-resolved 4D CBCT reconstruction
嵌入反投影算子的双域 U-Net 用于运动分辨的四维锥束 CT 重建
Abstract
Four-dimensional cone beam CT (4D CBCT) is important for image-guided radiation therapy of thoracic cancers, but its use is limited by long scan times, causing high patient dose and motion/sparse-sampling artifacts. We propose a deep learning method for motion-resolved 4D CBCT reconstruction from conventional free-breathing scans, without a respiratory signal or explicit projection binning. Our CNN takes free-breathing 3D CBCT projections as input and predicts a static volume at maximum inhalation plus ten displacement vector fields (DVFs) spanning a breathing cycle. The network extends U-Net: the encoder acts on filtered projection stacks, the decoder acts in the volume domain, and skip connections are replaced with non-trainable back-projection functions at multiple resolutions to transfer features between domains. The model is trained on simulated CBCT scans and evaluated on 11 unseen simulated patients and 13 clinical free-breathing scans. Two additional models (60 s and 6 s scans) were evaluated by clinical experts on three and two scans, comparing single phases of our 4D reconstruction to reference 3D SART-TV images for tumor and esophagus visibility. Experts preferred our method for tumor visibility (59% vs. 36% no preference, 5% reference) and esophagus visibility (47% vs. 42%, 11%). On simulated data, image quality matched SART-TV (mean RMSE: -1.19 HU, PSNR: +0.09 dB, SSIM: -0.009) while enabling 4D reconstruction. On clinical scans, our method showed sharper dynamic structures (e.g., diaphragm) and fewer motion streak artifacts than traditional reconstruction. This non-patient-specific CNN predicts static volumes and full 4D respiratory motion models from a single free-breathing scan, without a respiratory surrogate or projection binning, reducing motion artifacts while adding motion-modeling capability.
Chinese Translation
四维锥束 CT(4D CBCT)在胸部癌症的图像引导放射治疗中具有重要意义,但其应用受到长扫描时间的限制,导致患者辐射剂量高以及运动/稀疏采样伪影。我们提出了一种深度学习方法,用于从常规自由呼吸扫描中进行运动分辨的 4D CBCT 重建,无需呼吸信号或显式投影分箱。我们的卷积神经网络(CNN)以自由呼吸的三维 CBCT 投影作为输入,预测最大吸气时的静态体积以及跨越一个呼吸周期的十个位移矢量场(DVFs)。该网络扩展了 U-Net:编码器作用于过滤后的投影堆栈,解码器作用于体积域,跳跃连接被多个分辨率下的不可训练反投影函数所替代,以在不同域之间传递特征。该模型在模拟的 CBCT 扫描上进行训练,并在 11 个未见的模拟患者和 13 个临床自由呼吸扫描上进行评估。临床专家对两个额外模型(60 秒和 6 秒扫描)进行了评估,比较我们 4D 重建的单个相位与参考的 3D SART-TV 图像在肿瘤和食道可见性方面的表现。专家更倾向于我们的方法在肿瘤可见性方面(59% 对 36% 无偏好,5% 参考)和食道可见性方面(47% 对 42%,11%)。在模拟数据上,图像质量与 SART-TV 相当(平均 RMSE: -1.19 HU,PSNR: +0.09 dB,SSIM: -0.009),同时实现了 4D 重建。在临床扫描中,我们的方法显示出更清晰的动态结构(例如,膈肌)和比传统重建更少的运动条纹伪影。该非患者特异性的 CNN 从单个自由呼吸扫描中预测静态体积和完整的 4D 呼吸运动模型,无需呼吸替代信号或投影分箱,减少了运动伪影,同时增加了运动建模能力。
cs.CV / 78 / 2608.03471
Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding
Hi-Token:用于生成视觉定位的层次坐标标记化
Abstract
Generative Vision-Language Models (VLMs) commonly treat bounding-box coordinates as independent output symbols, leaving numerical order and axis semantics implicit. We identify this representation as an important source of error in visual grounding. Hi-Token encodes each coordinate with axis-specific tokens for the hundreds, tens, and ones digits, which adds coarse-to-fine structure and increases token reuse while retaining the existing VLM architecture. Hi-GAR complements this representation with a geometry-based reward for Group Relative Policy Optimization (GRPO), using box overlap and coordinate accuracy at multiple scales. Controlled comparisons under matched training conditions show that Hi-Token improves localization throughout the evaluated IoU range. Hi-GAR further reduces low-overlap predictions and is used only during training. Experiments on three VLM backbones and the RefCOCO family show consistent gains across models and benchmarks. Hi-R1 achieves higher values than strong specialist baselines on most reported metrics. Analyses of token frequency, digit boundaries, object scale, and IoU distributions explain the effects of coordinate representation and reward training. The results show that structured coordinate generation provides an effective approach to generative visual grounding.
Chinese Translation
生成视觉语言模型(VLMs)通常将边界框坐标视为独立的输出符号,从而使数值顺序和坐标轴语义隐含。我们识别出这种表示方式是视觉定位中的一个重要错误来源。Hi-Token 为每个坐标编码特定于坐标轴的标记,分别对应于百位、十位和个位数字,这增加了粗到细的结构并提高了标记的重用,同时保留了现有的 VLM 架构。Hi-GAR 通过基于几何的奖励来补充这种表示,适用于组相对策略优化(GRPO),使用多个尺度的框重叠和坐标精度。在匹配训练条件下的受控比较显示,Hi-Token 在评估的 IoU 范围内改善了定位。Hi-GAR 进一步减少了低重叠预测,仅在训练期间使用。在三个 VLM 骨干网络和 RefCOCO 系列上的实验表明,模型和基准测试之间的一致性提升。Hi-R1 在大多数报告的指标上达到了高于强基线的值。对标记频率、数字边界、物体尺度和 IoU 分布的分析解释了坐标表示和奖励训练的影响。结果表明,结构化坐标生成为生成视觉定位提供了一种有效的方法。
cs.CV / 79 / 2608.03474
MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification
MT-Web2Code:多轮区域重建和局部修改的编码代理基准测试
Abstract
Recent advances in Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in web UI generation. However, existing benchmarks predominantly focus on single-turn full-page generation from scratch, overlooking the iterative workflow of real-world frontend engineering, where developers repeatedly reconstruct missing regions and modify localized elements within existing codebases. To bridge this gap, we introduce MT-Web2Code, the first multimodal coding benchmark for multi-turn Macro-Level Regional Reconstruction and Micro-Level Localized Modification, which contains 102 tasks spanning 16 vertical domains. To construct deterministic repair trajectories without costly turn-level human annotation, we develop a scalable Reverse-Corruption Trajectory Engine that iteratively injects structural and stylistic defects into golden pages. We further propose a dual-axis evaluation protocol that measures target-region fidelity and the preservation of unaffected content, where regional reconstruction is assessed by a 5-dimensional VLM-based rubric and localized modification by deterministic pixel-grounded alignment. Experiments on 13 frontier coding agents reveal that current agents struggle to faithfully reconstruct target regions while preserving unaffected content, lack fine-grained visual-code alignment for localized edits, and suffer from error snowballing over multiple turns. Beyond benchmarking, our deterministic evaluation metrics provide fine-grained feedback signals that may facilitate future research on training iterative UI coding agents. Our evaluation code and data will soon be released.
Chinese Translation
近期大型视觉语言模型(LVLMs)的进展展示了其在网页用户界面生成方面的卓越能力。然而,现有基准主要集中在从零开始的单轮全页面生成,忽视了现实前端工程中的迭代工作流程,在该流程中,开发人员反复重建缺失区域并修改现有代码库中的局部元素。为了解决这一问题,我们提出了MT-Web2Code,这是首个针对多轮宏观级区域重建和微观级局部修改的多模态编码基准,包含跨越16个垂直领域的102个任务。为了构建确定性的修复轨迹而无需昂贵的轮次级人工标注,我们开发了一个可扩展的反腐蚀轨迹引擎,该引擎迭代地将结构和风格缺陷注入到黄金页面中。我们进一步提出了一种双轴评估协议,该协议衡量目标区域的保真度和未受影响内容的保留,其中区域重建通过基于5维的VLM评估标准进行评估,而局部修改则通过确定性的像素对齐进行评估。在对13个前沿编码代理的实验中,结果显示当前代理在忠实重建目标区域的同时保留未受影响内容方面存在困难,缺乏局部编辑的细粒度视觉代码对齐,并且在多个轮次中出现错误累积。除了基准测试外,我们的确定性评估指标提供了细粒度的反馈信号,可能促进未来对训练迭代用户界面编码代理的研究。我们的评估代码和数据将很快发布。
cs.CV / 80 / 2608.03508
From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology
从多分辨率细胞到千兆像素全切片图像的计算病理基础模型
Abstract
Vision Transformers (ViTs) and their hierarchical variants have achieved strong performance in Computational Pathology (CPath). However, most are pre-trained on single-resolution Whole Slide Images (WSIs), limiting their generalization across arbitrary resolutions. Gigapixel WSIs inherently contain diagnostic patterns at multiple scales, including cellular morphologies, tissue architectures, and global context, mirroring how expert pathologists examine WSIs. We introduce Multi-Resolution Pyramid Transformer (MRPT), a model that hierarchically aggregates multi-resolution information from cellular to tissue and WSI levels. MRPT employs a biologically meaningful Consecutive Cross-Resolution Attention (CCRA) mechanism to capture scale-independent interactions and enforces multi-resolution semantic consistency by aligning embeddings across resolutions, yielding robust and generalizable WSI representations. Pre-trained in a multi-resolution self-supervised manner on 624M patches, 2.4M regions, and 36K WSIs, MRPT learns rich coarse-to-fine histopathology features. Extensive experiments on 34 diverse datasets show that MRPT surpasses recent foundation models and Multimodal Large Language Models (MLLMs) in cancer subtype classification, tissue phenotyping, and Visual Question Answering (VQA) for WSI understanding.
Chinese Translation
视觉变换器(Vision Transformers, ViTs)及其分层变体在计算病理学(Computational Pathology, CPath)中表现出色。然而,大多数模型是在单一分辨率的全切片图像(Whole Slide Images, WSIs)上进行预训练的,这限制了它们在任意分辨率下的泛化能力。千兆像素的全切片图像本质上包含多尺度的诊断模式,包括细胞形态、组织结构和全局背景,反映了专家病理学家检查全切片图像的方式。我们提出了多分辨率金字塔变换器(Multi-Resolution Pyramid Transformer, MRPT),该模型从细胞到组织和全切片图层次聚合多分辨率信息。MRPT采用生物学上有意义的连续跨分辨率注意力机制(Consecutive Cross-Resolution Attention, CCRA),以捕捉尺度无关的交互,并通过对齐不同分辨率下的嵌入来强制实现多分辨率语义一致性,从而生成稳健且可泛化的全切片图像表示。MRPT在624M补丁、2.4M区域和36K全切片图像上以多分辨率自监督方式进行预训练,学习丰富的粗到细的组织病理特征。在34个多样化数据集上的广泛实验表明,MRPT在癌症亚型分类、组织表型分析和全切片图像理解的视觉问答(Visual Question Answering, VQA)任务中超越了最近的基础模型和多模态大型语言模型(Multimodal Large Language Models, MLLMs)。
cs.CV / 81 / 2608.03511
How Many Labels Are Enough? ALDA: Active Learning Deployment Advisor for Medical Image Classification
多少标签才够?ALDA:医疗图像分类的主动学习部署顾问
Abstract
Active learning (AL) promises to reduce the cost of medical imaging projects by lowering the number of clinical labels required. However, practical deployment requires committing to a sampling strategy before the full annotation budget is spent, and choosing the wrong strategy can increase rather than decrease costs. We propose Active-Learning Deployment Advisor (ALDA), a deployment-oriented framework for AL method selection under clinical performance constraints. Given a short pilot phase, ALDA fits a parametric learning-curve model to each candidate strategy, estimates whether that strategy is expected to reach a required clinical performance target, and predicts the number of expert annotations needed to do so. In addition to absolute annotation cost, ALDA introduces a deployment window that quantifies the sensitivity of this cost estimate to uncertainty in the clinical threshold. The final recommendation follows a risk-aware rule: among strategies with near-optimal predicted cost, ALDA prefers the strategy with the narrowest deployment window, the most robust to threshold revisions. Experiments on four medical imaging classification domains show that ALDA predicts the deployment-optimal method from a pilot of 15-30% of the intended budget and reduces annotation costs by up to 82% compared with a poor strategy choice. Rather than introducing a new sampling heuristic, ALDA provides a practical decision layer that answers a deployment-critical question: how many labels are enough?
Chinese Translation
主动学习(Active Learning, AL)有望通过减少所需的临床标签数量来降低医疗影像项目的成本。然而,实际部署需要在完全使用注释预算之前承诺一种采样策略,选择错误的策略可能会增加而不是减少成本。我们提出了主动学习部署顾问(Active-Learning Deployment Advisor, ALDA),这是一个针对临床性能约束下的主动学习方法选择的部署导向框架。在短期试点阶段,ALDA 为每个候选策略拟合一个参数化学习曲线模型,估计该策略是否预计能达到所需的临床性能目标,并预测为实现这一目标所需的专家注释数量。除了绝对注释成本外,ALDA 还引入了一个部署窗口,量化了该成本估计对临床阈值不确定性的敏感性。最终推荐遵循一种风险意识规则:在预测成本接近最优的策略中,ALDA 优先选择部署窗口最窄的策略,这种策略对阈值修订的鲁棒性最强。在四个医疗影像分类领域的实验表明,ALDA 能够从 15-30% 的预算试点中预测出部署最优的方法,并将注释成本与不佳策略选择相比降低多达 82%。ALDA 并不是引入一种新的采样启发式,而是提供了一个实用的决策层,回答了一个与部署密切相关的问题:多少标签才够?
cs.CV / 82 / 2608.03516
Detecting Pose Estimation Failures via Keypoint Self-Consistency
通过关键点自一致性检测姿态估计失败
Abstract
One common approach to pose estimation involves predicting object keypoints in an image, followed by using Perspective-n-Point algorithms to compute the object's rotation and translation relative to the camera. While rotations preserve object shapes, this property is often neglected in keypoint-based pose estimation methods, where keypoints are typically predicted independently from each other. As imprecise keypoint predictions negatively affects pose estimation accuracy, it also limits its reliability in downstream tasks. In this work, we explore whether such inaccurate pose estimates can be identified by simply examining spatial locations between 2D keypoints. We propose a set of hand-crafted geometric features that capture the self-consistency of keypoint predictions, including pairwise distances, reprojection consistency, as well as render and mask consistency. Despite its simplicity, a logistic regression classifier trained on these features reliably detects pose estimation failures, outperforming confidence-based approaches like conformal keypoint predictions that rely solely on keypoint uncertainty.
Chinese Translation
一种常见的姿态估计方法涉及在图像中预测物体关键点,随后使用透视-n-点(Perspective-n-Point)算法计算物体相对于相机的旋转和位移。虽然旋转保持物体形状,但这一特性在基于关键点的姿态估计方法中常常被忽视,因为关键点通常是相互独立地预测的。由于不精确的关键点预测会对姿态估计的准确性产生负面影响,这也限制了其在下游任务中的可靠性。在本研究中,我们探讨了是否可以通过简单检查2D关键点之间的空间位置来识别这种不准确的姿态估计。我们提出了一组手工设计的几何特征,捕捉关键点预测的自一致性,包括成对距离、重投影一致性以及渲染和掩膜一致性。尽管其简单性,基于这些特征训练的逻辑回归分类器能够可靠地检测姿态估计失败,优于依赖关键点不确定性的基于置信度的方法,如一致性关键点预测(conformal keypoint predictions)。
cs.CV / 83 / 2608.03517
GVCCTurbo: Rate-Compute Quality Scheduling for Codebook Driven Generative Compression
GVCCTurbo:基于码本驱动的生成压缩的速率-计算质量调度
Abstract
Codebook-driven generative compression uses a pretrained image or video generator as a zero-shot visual prior and transmits compact codebook indices to guide reconstruction at ultra-low bitrate. Current codecs tie each finite-rate correction to a fresh prior evaluation, so shortening the sampler also removes correction slots that carry target-dependent information. We propose GVCCTurbo, a BPP-driven scheduler that separates expensive prior refreshes from codebook corrections: after calibrating an atom-count operating point and skip-gap ratio once per protocol, it maps a target codebook-payload bitrate to a trajectory length and refresh period, making BPP a schedule input instead of a fixed consequence of sampler length. The same endpoint-prediction and finite-rate steering interface covers GVCC-style rectified-flow video and DDCM-style diffusion image compression, preserving zero-training deployment and compatibility with future distilled priors. Native 1080p curves position the complete zero-shot codec in the ultra-low-bitrate regime. In a controlled 720p Wan-GVCC study, the scheduler cuts prior evaluations from 20 to 9 for a $\sim\!44\%$ measured decoding-time reduction shared across the whole schedule family, at a small shared LPIPS cost on high-motion content; within that family, uniform refresh thinning (pure-skip) is a boundary point, and the BPP-aware interior point trades $2.9\%$ fewer codebook-payload bits for consistently higher PSNR at comparable LPIPS. These results support BPP-to-compute scheduling as a controllable extension of sampler-length tuning, without requiring the allocated point to dominate every boundary point.
Chinese Translation
基于码本驱动的生成压缩使用预训练的图像或视频生成器作为零-shot视觉先验,并传输紧凑的码本索引以指导在超低比特率下的重建。目前的编解码器将每个有限速率的修正与新的先验评估绑定,因此缩短采样器也会移除携带目标依赖信息的修正槽。我们提出了GVCCTurbo,一种基于比特每像素(BPP)的调度器,它将昂贵的先验刷新与码本修正分离:在每个协议中仅需一次校准原子计数操作点和跳过间隔比率后,它将目标码本负载比特率映射到轨迹长度和刷新周期,使BPP成为调度输入,而不是采样器长度的固定结果。相同的端点预测和有限速率引导接口覆盖了GVCC风格的修正流视频和DDCM风格的扩散图像压缩,保持零训练部署并与未来的蒸馏先验兼容。原生1080p曲线将完整的零-shot编解码器定位于超低比特率范围。在一个受控的720p Wan-GVCC研究中,调度器将先验评估从20次减少到9次,实现了约44%的解码时间减少,且在整个调度系列中共享,尽管在高运动内容上有小幅共享的LPIPS成本;在该系列中,均匀刷新稀疏(纯跳过)是一个边界点,而BPP感知的内部点则以2.9%的码本负载比特减少换取在可比LPIPS下更高的PSNR。这些结果支持将BPP转化为计算调度,作为采样器长度调优的可控扩展,而不需要分配点主导每个边界点。
cs.CV / 84 / 2608.03539
IRIS: Visual-Semantic Binding for Forgery-Resistant Watermarking of Diffusion Images
IRIS:用于抗伪造水印的扩散图像的视觉-语义绑定
Abstract
Most in-generation diffusion watermarks embed patterns independent of the image that carries them, and attackers transplant the marks onto images the generator did not produce, resulting in forgery. Binding the mark to visual semantics prevents such transplantation, yet existing bindings anchor to a proxy image rather than the image they mark. Realizing visual-semantic binding inside generation faces two challenges. The mark derives from the image itself yet enters the sampling trajectory before that image exists, and may itself shift the semantics it binds. The binding also meets opposite sensitivity demands, breaking under semantic change while holding through common processing. We present IRIS, a training-free watermarking scheme that embeds an Intrinsic Ring Identifier from Semantics. IRIS reads a content code from the non-watermarked generated image, derives a one-time ring from the code and a secret key, returns to the final low-noise steps of the same trajectory and blends the ring in, after the semantics it binds are settled. To meet the opposite sensitivity demands, the code is read through a canonicalization shared between embedding and detection, holding through common distortions and mild regeneration while flipping under semantic change. Detection recomputes the ring from the query image and the key alone, and the mark therefore fails on a foreign or spliced image, with acceptance tracking semantic displacement. On three prompt datasets IRIS detects reliably and stays close to its same-seed non-watermarked counterpart, a fidelity prior in-generation marks do not reach. While forgeries transfer fixed-pattern marks and regeneration strips post-hoc marks, IRIS alone among the compared marks withstands both.
Chinese Translation
大多数生成的扩散水印嵌入与承载它们的图像无关的模式,攻击者将水印移植到生成器未产生的图像上,从而导致伪造。将水印绑定到视觉语义上可以防止这种移植,然而现有的绑定是锚定在代理图像上,而不是它们所标记的图像。实现生成过程中的视觉-语义绑定面临两个挑战。水印源自图像本身,但在该图像存在之前就进入了采样轨迹,并且可能会改变其绑定的语义。绑定还面临相反的敏感性需求,在语义变化时会破裂,而在常见处理过程中则保持稳定。我们提出了IRIS,一种无训练的水印方案,它嵌入来自语义的内在环标识符。IRIS从未水印的生成图像中读取内容代码,从代码和秘密密钥中推导出一次性环,然后返回到同一轨迹的最终低噪声步骤中,并在绑定的语义确定后将环融合进去。为了满足相反的敏感性需求,该代码通过嵌入和检测之间共享的规范化进行读取,能够在常见失真和轻微再生中保持稳定,而在语义变化时则会翻转。检测仅从查询图像和密钥重新计算环,因此在外部或拼接图像上水印会失效,同时接受度跟踪语义位移。在三个提示数据集上,IRIS可靠地进行检测,并与其相同种子的未水印对应物保持接近,这是生成过程中的水印所无法达到的保真度。虽然伪造会转移固定模式水印,而再生会剥离后期水印,但在比较的水印中,只有IRIS能够抵御这两种情况。
cs.CV / 85 / 2608.03540
S$^3$-Diff: Structural Semantic Synergy Diffusion Model for High Fidelity Super Resolution of Pathological Images
S$^3$-Diff:用于病理图像高保真超分辨率的结构语义协同扩散模型
Abstract
Digital pathology relies on high-resolution whole slide images for accurate diagnosis, yet limitations in imaging devices, storage, and transmission often make lower-resolution pathology images more common in clinical workflows. Current super-resolution techniques often tend to smooth diagnostically relevant morphology, leading to over-smoothed textures and semantic drift that compromise downstream clinical interpretation. To this end, we develop the Structural Semantic Synergy Diffusion Model (S3-Diff), a diffusion framework for high-fidelity super-resolution of pathological images. The core of S3-Diff is Specimen-aware Structural Anchoring (SSA), which combines prognosis-aware tissue support extracted by a fixed SAM with LR-HR gradient discrepancies to generate a specimen-specific structural anchor to preserve pathological morphology. Concurrently, we introduce Structure-guided Semantic Fidelity Tuning (SSFT) to adapt DINOv3 representations using SSA-derived structural supervision. SSFT combines the adapted semantic energy with LR-derived edge and grayscale cues. The resulting control guides denoising to suppress stochastic artifacts and maintain structural consistency. Extensive experimental results demonstrate that S3-Diff consistently outperforms state-of-the-art methods in both reconstruction quality and downstream survival analysis performance. The source code will be made public.
Chinese Translation
数字病理学依赖于高分辨率的全切片图像以实现准确诊断,但成像设备、存储和传输的限制常常使得低分辨率的病理图像在临床工作流程中更为常见。目前的超分辨率技术往往倾向于平滑与诊断相关的形态特征,导致过度平滑的纹理和语义漂移,从而影响下游临床解读。为此,我们开发了结构语义协同扩散模型(S3-Diff),这是一个用于病理图像高保真超分辨率的扩散框架。S3-Diff的核心是样本感知结构锚定(SSA),它结合了通过固定的SAM提取的预后感知组织支持与低分辨率-高分辨率(LR-HR)梯度差异,以生成特定于样本的结构锚定,从而保留病理形态。同时,我们引入了结构引导的语义保真调优(SSFT),利用SSA衍生的结构监督来调整DINOv3的表示。SSFT将调整后的语义能量与低分辨率(LR)衍生的边缘和灰度线索结合。最终的控制引导去噪,以抑制随机伪影并保持结构一致性。大量实验结果表明,S3-Diff在重建质量和下游生存分析性能方面始终优于最先进的方法。源代码将公开发布。
cs.CV / 86 / 2608.03557
Test-Time Augmentation for Tabular-to-Image Classifiers under Distribution Shifts
分布变化下表格到图像分类器的测试时增强
Abstract
Tabular-to-image methods that convert tabular data into visual representations have emerged as a novel paradigm for leveraging the high performance of deep learning models. Despite their advantages, the robustness of these methods under distribution shifts remains under explored. Test-Time Augmentation (TTA) is an effective approach in image classification to improve model generalization and robustness, where predictions over multiple transformed views of each input are aggregated. This work evaluates the impact of TTA techniques on predictive performance under Out-Of-Distribution (OOD) for representations generated by tabular-to-image methods. Six tabular-to-image encoding methods were considered: TINTO, IGTD, DeepInsight, BIE, DistanceMatrix, Fotomics. Twenty-five TTA techniques were used, organized into six types: Geometric, Photometric, Structural, Frequency/Encoding, Mixup, and Composite. We employed two datasets from the TableShift benchmark (HELOC and Voting) that provide in-distribution and OOD test subsets designed to evaluate the effect of distribution shifts on tabular data. The results indicate that TTA improves OOD performance, with composite and photometric strategies providing the best trade-off between robustness and variance. In contrast, frequency-domain transformations that alter the encoder's feature-to-intensity mapping consistently degrade performance. These findings highlight TTA as a promising approach for improving the robustness and generalization of classifiers trained on image representations derived from tabular data, particularly under distribution shifts.
Chinese Translation
将表格数据转换为视觉表示的表格到图像方法已成为利用深度学习模型高性能的新范式。尽管这些方法具有优势,但它们在分布变化下的鲁棒性仍然未得到充分探索。测试时增强(Test-Time Augmentation, TTA)是一种在图像分类中有效的方法,通过对每个输入的多个变换视图的预测进行聚合,以提高模型的泛化能力和鲁棒性。本研究评估了TTA技术在表格到图像方法生成的表示下,对分布外(Out-Of-Distribution, OOD)预测性能的影响。考虑了六种表格到图像编码方法:TINTO、IGTD、DeepInsight、BIE、DistanceMatrix和Fotomics。使用了25种TTA技术,分为六类:几何(Geometric)、光度(Photometric)、结构(Structural)、频率/编码(Frequency/Encoding)、混合(Mixup)和复合(Composite)。我们采用了来自TableShift基准的两个数据集(HELOC和Voting),这些数据集提供了设计用于评估分布变化对表格数据影响的在分布内和OOD测试子集。结果表明,TTA提高了OOD性能,其中复合和光度策略在鲁棒性和方差之间提供了最佳平衡。相反,改变编码器特征与强度映射的频域变换始终会降低性能。这些发现突显了TTA作为一种有前景的方法,可以改善基于表格数据生成的图像表示的分类器的鲁棒性和泛化能力,尤其是在分布变化的情况下。
cs.CV / 87 / 2608.03559
Compass: Degradation-Simulated Reciprocal Learning with Lightweight Needle RWKV for Multimodal Crack Segmentation under Missing Modalities
Compass:轻量级针状 RWKV 的退化模拟互学习在缺失模态下的多模态裂缝分割
Abstract
In multimodal crack segmentation for industrial facilities, the key challenge is preventing missing modalities from degrading pixel-level performance while maintaining low computational cost. Existing methods struggle to address semantic degradation caused by missing modalities. We propose Compass, a lightweight network for robust crack segmentation under arbitrary missing modalities. Compass comprises Degradation Simulation Distillation (DSD), Needle Block, and Evidential Topology-Preserving Fusion (ETPF). DSD constructs a degradation simulation stream that mimics more severe missing conditions and performs reciprocal distillation with the original stream, decoupling complete perception from degradation adaptation. Within DSD, Feature-Aware Prototype Transmitter (FAPT) performs modality agnostic prototype-guided feature completion to maintain semantic integrity under incomplete modality conditions. As a lightweight backbone, Needle injects crack-direction cues into WKV modulation and combines connectivity-aware gating with anisotropic context probing for structure-aware modeling. ETPF fuses multimodal features via Dempster-Shafer evidential combination with uncertainty-gated decoding, preserving crack topology while suppressing unreliable features. Experiments on three datasets demonstrate state-of-the-art (SOTA) performance under diverse missing modality scenarios. Even with 90\% depth modality missing on CrackDepth, Compass achieves F1 of 0.8216 and mIoU of 0.8434 with only 2.58M parameters. The code is available at https://github.com/Karl1109/Compass.
Chinese Translation
在工业设施的多模态裂缝分割中,关键挑战是防止缺失模态降低像素级性能,同时保持低计算成本。现有方法难以解决缺失模态导致的语义退化问题。我们提出了 Compass,一种针对任意缺失模态的稳健裂缝分割的轻量级网络。Compass 包括退化模拟蒸馏(Degradation Simulation Distillation, DSD)、针状块(Needle Block)和证据拓扑保持融合(Evidential Topology-Preserving Fusion, ETPF)。DSD 构建了一个模拟更严重缺失条件的退化模拟流,并与原始流进行互蒸馏,从而将完整感知与退化适应解耦。在 DSD 中,特征感知原型传输器(Feature-Aware Prototype Transmitter, FAPT)执行模态无关的原型引导特征补全,以在不完整模态条件下保持语义完整性。作为轻量级主干,针状块将裂缝方向线索注入 WKV 调制,并结合连接感知门控与各向异性上下文探测进行结构感知建模。ETPF 通过 Dempster-Shafer 证据组合与不确定性门控解码融合多模态特征,保持裂缝拓扑的同时抑制不可靠特征。在三个数据集上的实验表明,在多种缺失模态场景下,Compass 达到了最先进的(SOTA)性能。即使在 CrackDepth 上缺失 90 ext{%} 的深度模态,Compass 也能以仅 2.58M 参数实现 F1 值 0.8216 和 mIoU 值 0.8434。代码可在 https://github.com/Karl1109/Compass 获取。
cs.CV / 88 / 2608.03571
Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning
超越简单的环境缩放:为多模态智能体学习设计有效的环境分布
Abstract
Recent works train agents by constructing large-scale multimodal environment pools. However, we find that simply increasing the number of multimodal environments does not always benefit. We further analyze the limitations in current multimodal environment distributions through a series of experiments. Based on these findings, we study how to build more effective training environment distributions from two dimensions: **diversity** and **difficulty structure**. For diversity, we propose **Ability-aware Environment Selection (AES)** to obtain diverse environment sets. For difficulty structure, we propose **Hierarchical Difficulty Curriculum (HDC)**, which organizes curriculum learning through two difficulty levels: harness weakening and state-scale progression. Experiments show that AES and HDC effectively improve multimodal agent training.
Chinese Translation
近期的研究通过构建大规模的多模态环境池来训练智能体。然而,我们发现仅仅增加多模态环境的数量并不总是有益。我们通过一系列实验进一步分析了当前多模态环境分布的局限性。基于这些发现,我们研究了如何从两个维度构建更有效的训练环境分布:**多样性**和**难度结构**。在多样性方面,我们提出了**能力感知环境选择(Ability-aware Environment Selection, AES)**,以获得多样化的环境集合。在难度结构方面,我们提出了**分层难度课程(Hierarchical Difficulty Curriculum, HDC)**,通过两种难度级别组织课程学习:能力削弱和状态规模进展。实验表明,AES和HDC有效地改善了多模态智能体的训练。
cs.CV / 89 / 2608.03580
SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models
SlimVLM:敏感性感知的动态结构化剪枝与自适应视觉标记选择用于高效的视觉-语言模型
Abstract
While Vision-Language Models (VLMs) have demonstrated remarkable performance in processing and understanding both text and images, their large parameter sizes lead to significant computational overhead, limiting their deployment on resource-constrained devices. While pruning has been effective for compressing Large Language Models (LLMs), directly applying it to VLMs leads to significant performance drops, largely due to redundant visual tokens interfering with importance estimation. To this end, we propose SlimVLM, a structured pruning framework designed to compress VLMs while preserving their task performance. We introduce an adaptive visual token selection strategy for VLMs that leverages average text-to-visual attention scores to assess the importance of visual tokens, removing redundant ones during pruning based on a set threshold, thereby optimizing the importance calculation. Recognizing the varying tolerance to sparsity across different modules, we also propose a Sensitivity-aware dynamic pruning mechanism that determines the appropriate pruning ratio for each module by calculating the linear reconstruction error between the outputs of the pruned and unpruned modules, ensuring overall performance stability. Experimental results show that SlimVLM outperforms existing methods across multiple multimodal benchmarks, achieving state-of-the-art performance.
Chinese Translation
尽管视觉-语言模型(VLMs)在处理和理解文本与图像方面表现出色,但其庞大的参数规模导致了显著的计算开销,限制了它们在资源受限设备上的部署。虽然剪枝在压缩大型语言模型(LLMs)方面效果显著,但直接将其应用于VLMs会导致显著的性能下降,这主要是由于冗余的视觉标记干扰了重要性估计。为此,我们提出了SlimVLM,一个旨在压缩VLMs同时保持其任务性能的结构化剪枝框架。我们引入了一种针对VLMs的自适应视觉标记选择策略,该策略利用平均文本到视觉的注意力得分来评估视觉标记的重要性,在剪枝过程中根据设定的阈值移除冗余标记,从而优化重要性计算。考虑到不同模块对稀疏性的容忍度不同,我们还提出了一种敏感性感知的动态剪枝机制,通过计算剪枝模块与未剪枝模块输出之间的线性重构误差来确定每个模块的适当剪枝比例,从而确保整体性能的稳定性。实验结果表明,SlimVLM在多个多模态基准测试中优于现有方法,实现了最先进的性能。
cs.CV / 90 / 2608.03618
Geospatial-Prior Guidance for 3D Semantic Scene Completion
基于地理空间先验的三维语义场景补全
Abstract
Inferring complete 3D geometry and semantics from onboard images remains challenging because occlusions and restricted fields of view leave large scene regions underconstrained. Although satellite imagery provides wide-area context, appearance cues alone offer limited structural guidance and may be unreliable because of spatial or temporal discrepancies. We present GeoScene, a geospatially guided framework that jointly uses satellite imagery and structured OpenStreetMap cues as soft priors for 3D semantic scene completion. GeoScene learns complementary voxel-wise reliability weights for onboard observations and geospatial guidance, and uses them to control feature refinement in observed and unobserved regions. This design preserves local visual evidence while exploiting large-scale road and building structure beyond onboard visibility. Experiments on SemanticKITTI and SSCBench-KITTI-360 demonstrate that GeoScene consistently improves both geometric and semantic completion under the geospatial-prior-assisted setting, with the most pronounced benefits for large-scale static and geospatially structured classes.
Chinese Translation
从车载图像推断完整的三维几何形状和语义仍然具有挑战性,因为遮挡和受限的视野使得大场景区域处于欠约束状态。尽管卫星图像提供了广域上下文,但单靠外观线索提供的结构指导有限,并且由于空间或时间差异可能不可靠。我们提出了GeoScene,一个地理空间引导框架,联合利用卫星图像和结构化的OpenStreetMap线索作为三维语义场景补全的软先验。GeoScene学习车载观测和地理空间引导的互补体素级可靠性权重,并利用这些权重控制观察和未观察区域的特征精细化。该设计在保留局部视觉证据的同时,利用超出车载可见性的广域道路和建筑结构。对SemanticKITTI和SSCBench-KITTI-360的实验表明,在地理空间先验辅助的设置下,GeoScene在几何和语义补全方面始终表现出改善,尤其对大规模静态和地理空间结构化类别的益处最为显著。
cs.CV / 91 / 2608.03631
SEER: A Self-Grounded Evidence Interface for Controlled Spatial Relation Classification
SEER:一种自我基础证据接口用于受控空间关系分类
Abstract
Spatial relation questions require a model to identify the queried subject and object before comparing their layout. Yet a VLM can recognize both entities and still answer from the wrong instance or an ambiguous global view. We ask whether making query-specific evidence explicit can mitigate this failure and propose SEER (Self-grounded Evidence for Entity-Relation Reasoning), a training-free inference-time evidence interface for frozen VLMs. SEER hides candidate relations during pair localization, constructs a query-specific view with explicit subject/object roles, and retains the full image and sparse box geometry as complementary evidence. For relation-choice protocols with exact inverse support, an optional refinement swaps the entity roles and changes the forward decision only when exactly one visual state obeys the corresponding inverse relation. On an image-disjoint GQA-Train900 test frozen before model scoring, SEER pools to +3.94 [2.17,5.72] over Full; the gain remains positive under label-independent grounding-order counterbalancing and on the 535 rows whose entity names are unique. The unchanged protocol yields +4.35 to +11.79 on all 2,434 filtered EmbSpatial pair-relation questions across three models. Matched controls separate local refocus from role-explicit conditioning. These results establish query-specific evidence construction as the principal intervention, with reciprocal consistency as a smaller protocol-specific refinement.
Chinese Translation
空间关系问题要求模型在比较布局之前识别查询的主题和对象。然而,视觉语言模型(VLM)能够识别这两个实体,但仍可能从错误的实例或模糊的全局视角进行回答。我们探讨了明确查询特定证据是否能缓解这种失败,并提出了SEER(自我基础证据用于实体-关系推理),这是一种针对冻结的VLM的无训练推理时证据接口。SEER在对偶本地化过程中隐藏候选关系,构建具有明确主题/对象角色的查询特定视图,并保留完整图像和稀疏框几何作为补充证据。对于具有精确逆支持的关系选择协议,一个可选的细化步骤在仅有一个视觉状态符合相应逆关系时,交换实体角色并改变前向决策。在一个在模型评分之前冻结的图像不重叠的GQA-Train900测试中,SEER的表现提升了+3.94 [2.17,5.72],在标签独立的基础顺序对照下,增益仍然为正,并且在535个实体名称唯一的行上保持一致。未改变的协议在所有2,434个过滤的EmbSpatial对关系问题上产生了+4.35到+11.79的提升。匹配的对照实验将局部重新聚焦与角色明确的条件分开。这些结果确立了查询特定证据构建作为主要干预措施,而互惠一致性则作为较小的协议特定细化。
cs.CV / 92 / 2608.03637
Learning Biomechanically Plausible Human Motion from Sparse Radar Point Clouds
从稀疏雷达点云中学习生物力学合理的人体运动
Abstract
Radar-based human pose estimation has focused on improving learning algorithms while representing the body as unconstrained keypoint coordinates. We address the underexplored dimension of anatomical fidelity by integrating a full-body skeletal model into a differentiable, end-to-end trainable radar-based pose estimation framework, in which the pose network is supervised through forward kinematics while subject-specific geometry is fitted beforehand. Subject-specific body segment proportions are predicted from radar point cloud features to scale a biomechanical skeleton. A motion prediction network maps temporal radar sequences to generalized coordinates, and differentiable forward kinematics converts predicted joint angles into 3D positions. A contact classification loss encourages physically plausible foot-ground interaction. Under leave-one-subject-out cross-validation on 11 healthy participants performing rehabilitation exercises, the framework achieves 6.456 +/- 1.759 cm mean per-joint position error (MPJPE), 8.083 +/- 0.884 degrees mean per-joint angle error (MPJAE), 0.935 +/- 0.009 contact classification F1, and 3.4 +/- 1.3 % scaling error. This proof-of-concept study demonstrates the feasibility of recovering interpretable biomechanical descriptors from a single low-cost radar sensor in a controlled laboratory setting, a prerequisite for future clinical motion analysis.
Chinese Translation
基于雷达的人体姿态估计主要集中在改进学习算法,同时将身体表示为不受限制的关键点坐标。我们通过将全身骨骼模型集成到一个可微分的端到端可训练的基于雷达的姿态估计框架中,解决了解剖学真实度这一未被充分探索的维度。在该框架中,姿态网络通过正向运动学进行监督,同时在此之前拟合特定个体的几何形状。根据雷达点云特征预测特定个体的身体段比例,以缩放生物力学骨骼。运动预测网络将时间序列雷达数据映射到广义坐标,而可微分的正向运动学将预测的关节角度转换为三维位置。接触分类损失鼓励物理上合理的足部与地面的相互作用。在对11名健康参与者进行康复训练的留一法交叉验证中,该框架实现了6.456 +/- 1.759 cm的每关节位置均值误差(MPJPE)、8.083 +/- 0.884度的每关节角度均值误差(MPJAE)、0.935 +/- 0.009的接触分类F1值,以及3.4 +/- 1.3%的缩放误差。这项概念验证研究展示了在受控实验室环境中,从单个低成本雷达传感器恢复可解释的生物力学描述符的可行性,这是未来临床运动分析的前提条件。
cs.CV / 93 / 2608.03649
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware
何时较少的视觉标记加速多模态推理?跨决策位置和硬件的盈亏平衡研究
Abstract
Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead, shared work, and the operators each policy can avoid. A stage-level decomposition reconciles these components with measured end-to-end latency. In a 30-example pilot, the two tested autoregressive probes remain slower than Full despite state reuse. A lightweight post-vision predictor yields paired confidence intervals below zero on RTX 3090 and A100 and remains significant after a conservative all-pairs Holm correction. A pre-vision image-size rule also yields intervals below zero on both GPUs, although neither comparison remains significant after the same correction. Pre-vision routing has a structural opportunity unavailable to post-vision pruning: it can avoid preprocessing and vision encoding. On A100, this opportunity outweighs a nearly eightfold larger downstream token reduction by the post-vision policy. Reported quality is conditional on examples answered correctly by Full and is not benchmark accuracy.
Chinese Translation
较少的视觉标记并不保证更低的端到端延迟。我们通过一种可重复的协议评估盈亏平衡,该协议考虑了决策开销、共享工作以及每种策略可以避免的操作。阶段级别的分解将这些组件与测量的端到端延迟进行调和。在一个包含30个示例的初步实验中,尽管状态重用,两个测试的自回归探针仍然比完整模型(Full)慢。一个轻量级的后视觉预测器在RTX 3090和A100上产生的配对置信区间低于零,并且在经过保守的全对比Holm校正后仍然显著。一个前视觉图像大小规则在这两种GPU上也产生了低于零的区间,尽管在相同校正后没有任何比较保持显著性。前视觉路由具有后视觉剪枝所无法获得的结构性机会:它可以避免预处理和视觉编码。在A100上,这一机会的价值超过了后视觉策略所带来的近八倍的下游标记减少。报告的质量是基于完整模型正确回答的示例,并不是基准准确性。
cs.CV / 94 / 2608.03664
Morphology-Aware Implicit Super-Resolution Network for Pathological Images
形态感知隐式超分辨率网络用于病理图像
Abstract
Accurate diagnosis in Digital Pathology (DP) relies on high-resolution whole-slide images, yet clinical deployment is often limited by hardware costs. Super-Resolution (SR) offers a promising alternative by computationally enhancing low-resolution acquisitions. However, existing SR methods frequently struggle to preserve fine-grained cellular morphology, leading to texture oversmoothing and blurred structural boundaries under complex tissue variability. To address this issue, we propose Morph-ISR, a morphology-aware implicit super-resolution framework for DP that restores diagnostically relevant details with sub-pixel precision. Morph-ISR reformulates SR as a continuous coordinate-based reconstruction problem and integrates an Implicit Position-aware Kernel Generator (IPKG) to adaptively model spatially varying tissue morphology. To further enhance structural fidelity, a Morphological Fidelity Prior (MFP) is introduced, leveraging semantic guidance from a pre-trained cell segmentation network to enforce boundary-preserving and region-aware reconstruction, thereby improving the representation of critical cellular boundaries and nuclear textures. Experiments on TCGA and SurGen datasets show that Morph-ISR achieves the best LPIPS and ST-LPIPS among the evaluated methods, reducing them by up to 38.37% and 39.55%, respectively, over the second-best methods while maintaining strong PSNR and SSIM. These results demonstrate superior preservation of diagnostically relevant cellular boundaries and nuclear textures, while compact parameterization and high throughput support efficient edge deployment. Code and trained models will be released upon publication.
Chinese Translation
数字病理学(Digital Pathology, DP)中的准确诊断依赖于高分辨率的全切片图像,但临床应用常常受到硬件成本的限制。超分辨率(Super-Resolution, SR)通过计算增强低分辨率图像,提供了一种有前景的替代方案。然而,现有的SR方法常常难以保留细微的细胞形态,导致在复杂组织变异下纹理过度平滑和结构边界模糊。为了解决这一问题,我们提出了Morph-ISR,一个形态感知的隐式超分辨率框架,旨在以亚像素精度恢复与诊断相关的细节。Morph-ISR将SR重新表述为一个基于连续坐标的重建问题,并集成了隐式位置感知核生成器(Implicit Position-aware Kernel Generator, IPKG),以自适应建模空间变化的组织形态。为了进一步增强结构的保真度,引入了形态保真先验(Morphological Fidelity Prior, MFP),利用预训练的细胞分割网络提供的语义指导,强制执行边界保留和区域感知的重建,从而改善关键细胞边界和细胞核纹理的表现。在TCGA和SurGen数据集上的实验表明,Morph-ISR在评估的方法中实现了最佳的LPIPS和ST-LPIPS,分别比第二好的方法降低了最多38.37%和39.55%,同时保持了强大的PSNR和SSIM。这些结果展示了对与诊断相关的细胞边界和细胞核纹理的优越保留,同时紧凑的参数化和高吞吐量支持高效的边缘部署。代码和训练模型将在发表时发布。
cs.CV / 95 / 2608.03666
XiDepth: a Lightweight and Efficient Network for Self-supervised Monocular Depth Estimation
XiDepth:一种轻量高效的自监督单目深度估计网络
Abstract
Self-supervised monocular depth estimation has emerged as an appealing solution to design lightweight and effective models for deployment on computationally constrained devices due to its reduced reliance on expensive depth sensors. By eliminating the need for ground-truth annotations and leveraging the simplicity of monocular camera setups, this approach facilitates cost-effective data collection and broad applicability across fields such as computer vision and robotics. A critical challenge is achieving resource-efficient neural networks without compromising the overall performance. State-of-the-art models generally adopt depth-wise convolutions and attention mechanisms; however, these functions often incur high energy costs and face compatibility issues in embedded environments. To address this, we propose XiDepth, a lightweight architecture based on the XiNet operator block, designed to enhance feature extraction while maintaining low computational complexity and energy demand. On the KITTI dataset, XiDepth achieves state-of-the-art performance with only 0.8M parameters. Tests on a Raspberry Pi 4 further confirm its suitability for real-world embedded applications, reducing FLOPs by 40% and energy consumption by 35% compared to leading methods.
Chinese Translation
自监督单目深度估计因其对昂贵深度传感器的依赖减少,已成为设计轻量且有效模型以在计算资源受限设备上部署的一个吸引人的解决方案。通过消除对真实标注的需求并利用单目相机设置的简单性,该方法促进了成本效益高的数据收集,并在计算机视觉和机器人等领域具有广泛的适用性。一个关键挑战是实现资源高效的神经网络,而不影响整体性能。现有的最先进模型通常采用深度卷积和注意力机制;然而,这些功能往往会产生高能耗,并在嵌入式环境中面临兼容性问题。为了解决这一问题,我们提出了XiDepth,一种基于XiNet操作块的轻量架构,旨在增强特征提取,同时保持低计算复杂度和能量需求。在KITTI数据集上,XiDepth以仅0.8M的参数量实现了最先进的性能。在Raspberry Pi 4上的测试进一步确认了其在现实嵌入式应用中的适用性,与领先方法相比,FLOPs减少了40%,能耗降低了35%。
cs.CV / 96 / 2608.03681
Keep the Needle, Prune the Haystack: Defect-Preserving Token Pruning for Efficient Zero-Shot Anomaly Detection
保持针头,修剪干草堆:用于高效零-shot异常检测的缺陷保留令牌修剪
Abstract
Zero-shot visual anomaly detection has achieved remarkable progress, with recent vision-only approaches further improving performance while simplifying the inference pipeline. However, existing methods typically perform dense computation over all images and spatial tokens, despite the fact that normal samples dominate real-world scenarios and anomalies usually occupy only small regions. Token pruning offers a promising solution, but introduces an asymmetric pruning risk in anomaly detection: retaining normal tokens mainly incurs redundant computation, whereas removing anomalous tokens may eliminate the only evidence for detection and localization. This risk is particularly severe in early layers, where pruning provides the greatest computational benefit but anomaly semantics remain unreliable. We propose KeepAD, a defect-preserving token pruning framework that formulates token selection as high-recall, anomaly-aware routing. In shallow layers, KeepAD combines coverage-preserving selection over local $2\times2$ patch neighborhoods with deterministic anomaly rescue to reduce the risk of discarding subtle defects. In deeper layers, frozen normal and abnormal prototypes guide pruning under an image-adaptive token budget, aggressively removing low-risk normal tokens while preserving local anomaly evidence. Dense-to-sparse self-distillation further supervises early token routing without introducing additional inference overhead. Experiments on six industrial and seven medical zero-shot anomaly detection benchmarks show that KeepAD reduces the token retention ratio to below $20\%$, while limiting the average degradation in image-level and pixel-level AUROC to within $2.7$ percentage points. At the most aggressive operating point, KeepAD achieves a $7.9\times$ speedup over the strongest CLIP-based baseline.
Chinese Translation
零-shot视觉异常检测取得了显著进展,最近的仅基于视觉的方法进一步提高了性能,同时简化了推理流程。然而,现有方法通常对所有图像和空间令牌进行密集计算,尽管正常样本在现实场景中占主导地位,而异常通常仅占小区域。令牌修剪提供了一种有前景的解决方案,但在异常检测中引入了不对称的修剪风险:保留正常令牌主要会产生冗余计算,而移除异常令牌可能会消除检测和定位的唯一证据。这种风险在早期层中尤为严重,因为修剪提供了最大的计算收益,但异常语义仍然不可靠。我们提出了KeepAD,一种缺陷保留的令牌修剪框架,将令牌选择形式化为高召回率、异常感知的路由。在浅层中,KeepAD结合了对局部$2 imes2$补丁邻域的覆盖保留选择与确定性异常拯救,以降低丢弃微妙缺陷的风险。在深层中,冻结的正常和异常原型在图像自适应令牌预算下指导修剪,积极移除低风险的正常令牌,同时保留局部异常证据。密集到稀疏的自蒸馏进一步监督早期令牌路由,而不引入额外的推理开销。在六个工业和七个医学零-shot异常检测基准上的实验表明,KeepAD将令牌保留比例降低到20%以下,同时将图像级和像素级AUROC的平均降幅限制在2.7个百分点以内。在最激进的操作点上,KeepAD实现了相较于最强的基于CLIP的基线的7.9倍加速。
cs.CV / 97 / 2608.03708
MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding
MultiCompose:基于每个主题属性绑定的多概念个性化组合
Abstract
Text-to-image diffusion models enable personalization of specific visual concepts from a small number of reference images. However, generating a single image that contains multiple personalized subjects, each bound to user-specified attributes such as clothing, accessories, and held objects, remains largely unaddressed. Without explicit spatial constraints, concurrently activated concept checkpoints produce overlapping cross-attention responses, causing per-subject identity degradation and attribute misalignment. Moreover, no established benchmark jointly evaluates these two failure modes in the personalized multi-subject setting. We present MultiCompose, a composition framework that decouples per-concept personalization from multi-subject inference. A semantic preservation regularization maintains attribute binding capacity during fine-tuning, while a two-phase inference procedure automatically establishes subject layout and composes per-concept predictions through spatially exclusive masks. We further introduce MSP-Bench, a benchmark that jointly evaluates identity fidelity (ID), attribute binding accuracy (BIND), and attribute misalignment (MIS) through a dual-pathway protocol. Experiments show that MultiCompose outperforms existing methods on both conventional metrics and MSP-Bench, confirming the benchmark's ability to reveal failure modes that conventional metrics overlook. Code is available at https://github.com/I2-Multimedia-Lab/MultiCompose
Chinese Translation
文本到图像的扩散模型使得从少量参考图像中个性化特定视觉概念成为可能。然而,生成包含多个个性化主题的单一图像,每个主题绑定用户指定的属性(如服装、配饰和持有物品),仍然在很大程度上未得到解决。在没有明确空间约束的情况下,同时激活的概念检查点会产生重叠的交叉注意响应,导致每个主题的身份退化和属性错位。此外,目前没有建立的基准能够在个性化多主题设置中共同评估这两种失败模式。我们提出了MultiCompose,一个将每个概念的个性化与多主题推断解耦的组合框架。语义保持正则化在微调过程中维持属性绑定能力,而两阶段推断程序通过空间上独占的掩模自动建立主题布局并组合每个概念的预测。我们进一步引入了MSP-Bench,一个基准,通过双通道协议共同评估身份保真度(ID)、属性绑定准确性(BIND)和属性错位(MIS)。实验表明,MultiCompose在传统指标和MSP-Bench上均优于现有方法,确认了该基准揭示传统指标所忽视的失败模式的能力。代码可在 https://github.com/I2-Multimedia-Lab/MultiCompose 获取。
cs.CV / 98 / 2608.03711
Attention is Case-Sensitive
注意力对大小写敏感
Abstract
In human visual perception, uppercase lettering serves as a natural salience cue that captures attention within lowercase text. In this paper, we present a systematic empirical characterization study revealing that Large Language Models (LLMs) exhibit an analogous property: letter casing modulates internal attention allocation. Through analysis across 13 models, nine LLMs and four Vision-Language Models (VLMs), with diverse tokenization schemes, we show that formatting target information in alternating or uppercase against a lowercase context concentrates attention on those textual spans. In text this effect is universal, holding across every evaluated non-reasoning model. We frame it as a previously under-explored latent property of pretrained transformers rather than a prescriptive method. Our investigation reveals a central attention-performance divergence: while this "casing effect" robustly shifts attention, its impact on downstream accuracy is non-trivial, increased concentration does not inherently improve task accuracy and, in high-entropy contexts like alternating case, can degrade it. We further identify a boundary condition: the deliberative "thinking" phase in reasoning models acts as a semantic buffer that mitigates typographic sensitivity in text. Extending the study to VLMs, we find the effect transfers partially: the same prompt-side casing reorganizes cross-modal attention along two coupled axes, predominantly a macroscopic disengagement from the image toward the text prompt, and secondarily a concentration of the residual visual attention on the target region. By isolating casing as a zero-shot mechanism for attention steering that requires no model access or fine-tuning, we provide a new foundational understanding of how pretraining internalizes typographic emphasis.
Chinese Translation
在人类视觉感知中,大写字母作为一种自然的显著性线索,能够在小写文本中捕捉注意力。本文呈现了一项系统的实证特征研究,揭示大型语言模型(Large Language Models, LLMs)展现出类似的特性:字母的大小写调节内部注意力的分配。通过对13个模型的分析,包括9个LLM和4个视觉-语言模型(Vision-Language Models, VLMs),以及多种分词方案,我们展示了在小写上下文中以交替或大写格式化目标信息能够集中注意力于这些文本片段。在文本中,这一效应是普遍存在的,适用于每个评估的非推理模型。我们将其框架视为预训练变换器的一个先前未被充分探讨的潜在特性,而非一种规定性的方法。我们的研究揭示了一个中心的注意力-表现差异:尽管这种“大小写效应”稳健地转移注意力,但其对下游准确性的影响并不简单,增加的集中度并不必然提高任务准确性,并且在像交替大小写这样的高熵上下文中,可能会降低准确性。我们进一步识别出一个边界条件:推理模型中的深思“思考”阶段充当了一个语义缓冲区,减轻了文本中的排版敏感性。将研究扩展到VLMs,我们发现该效应部分转移:相同的提示侧大小写重新组织了跨模态注意力,沿着两个耦合轴,主要是从图像向文本提示的宏观脱离,其次是对目标区域的残余视觉注意力的集中。通过将大小写隔离为一种无需模型访问或微调的零-shot注意力引导机制,我们提供了一个新的基础理解,阐明了预训练如何内化排版强调。
cs.CV / 99 / 2608.03724
Towards Reliable and Reproducible Fetal Brain Biometry: A Deep Learning Approach Using MRI
迈向可靠和可重复的胎儿脑部生物测量:一种基于深度学习的MRI方法
Abstract
Fetal brain biometry is essential for quantitative assessment of brain development, supporting gestational age estimation, developmental monitoring, and detection of abnormalities. In clinical practice, measurements are manually performed, making them time-consuming and prone to variability. While automated approaches have been proposed, reproducible methods remain limited, particularly those providing anatomically interpretable landmark localization. We present a fully automated deep learning-based framework for reliable and reproducible brain biometry from 3D super-resolution-reconstructed fetal brain MRI. The proposed four-step pipeline derives biometric parameters by jointly estimating linear measurements and their corresponding anatomical landmarks. A 3D convolutional neural network is trained to regress landmark coordinates from brain segmentation label maps, followed by measurement-specific geometric optimization to refine landmark positions and compute measurements. The pipeline is evaluated on two publicly available fetal MRI datasets comprising 150 volumes (gestational age range: 20-37 weeks) acquired across different scanners and protocols, assessing five key biometric measurements across varying acquisition settings and providing a comprehensive evaluation of both measurement accuracy and landmark localization using quantitative metrics and visual assessment. Compared with the only available automated pipeline, the proposed method achieves comparable or improved accuracy for most measurements. In conclusion, we introduce a straightforward pipeline for reliable biometry estimations, with efficiency, interpretability and scalability that support integration into clinical workflows.
Chinese Translation
胎儿脑部生物测量对于定量评估脑部发育至关重要,支持妊娠年龄估计、发育监测和异常检测。在临床实践中,测量通常是手动进行的,这使得测量过程耗时且容易受到变异的影响。虽然已经提出了自动化方法,但可重复的方法仍然有限,特别是那些提供解剖学可解释的标志点定位的方法。我们提出了一种完全自动化的基于深度学习的框架,用于从3D超分辨率重建的胎儿脑MRI中进行可靠和可重复的脑部生物测量。所提出的四步流程通过联合估计线性测量值及其对应的解剖标志点,导出生物测量参数。我们训练了一个3D卷积神经网络,从脑部分割标签图中回归标志点坐标,随后进行特定测量的几何优化,以精细化标志点位置并计算测量值。该流程在两个公开可用的胎儿MRI数据集上进行了评估,这些数据集包含150个体积(妊娠年龄范围:20-37周),在不同的扫描仪和协议下获取,评估了五个关键生物测量值在不同采集设置下的表现,并使用定量指标和视觉评估提供了测量准确性和标志点定位的全面评估。与唯一可用的自动化流程相比,所提出的方法在大多数测量中实现了可比或更高的准确性。总之,我们介绍了一种简单的流程,用于可靠的生物测量估计,具有效率、可解释性和可扩展性,支持集成到临床工作流程中。
cs.CV / 100 / 2608.03763
TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding
TDVR:零-shot 3D视觉定位中的联合文本消歧与视角推理
Abstract
Zero-shot 3D visual grounding aims to localize specific objects based on textual descriptions and 3D visual input. However, the effectiveness of existing methods is significantly hindered by the ambiguous query text and deficient viewpoints. To address these issues, we propose TDVR, a training-free reasoning framework that disambiguates the input text and infers accurate viewpoints for zero-shot 3D visual grounding. First, we construct semantic 3D scene graph from the detected instances in the 3D point cloud. Subsequently, we put the original query, appearance and spatial relationship descriptions into the LLM for fusion, thereby disambiguating the initial input. We leverage chain-of-thought reasoning to generate the structured representation of disambiguated query. Then taking the scene graph and structured query as input, we get the optimal view via viewpoint reasoning to solve the problem of missing viewpoints during grounding. Based on the obtained optimal viewpoint, we further discriminate the distracting objects, enabling the model with the ability to distinguish similar instances. After that, we match the category text and appearance images with the query by computing the similarity of feature vectors. Finally, the target object was identified by integrating the viewpoint score, confusion score, category score, and appearance score. Compared with previous methods, our TDVR has stronger capabilities in viewpoint reasoning, similar object discrimination, and ambiguous query understanding. Experimental results on the public ScanRefer dataset show that our method outperforms the existing state-of-the-art methods by 15.25% and 14.46% in
[email protected] and
[email protected] respectively, demonstrating the effectiveness of our TDVR in addressing ambiguous query text and deficient viewpoints.
Chinese Translation
零-shot 3D视觉定位旨在根据文本描述和3D视觉输入定位特定物体。然而,现有方法的有效性受到模糊查询文本和视角不足的显著影响。为了解决这些问题,我们提出了TDVR,一个无训练的推理框架,用于消歧输入文本并推断准确的视角以实现零-shot 3D视觉定位。首先,我们从3D点云中检测到的实例构建语义3D场景图。随后,我们将原始查询、外观和空间关系描述输入到大型语言模型(LLM)中进行融合,从而消歧初始输入。我们利用链式推理生成消歧查询的结构化表示。然后,以场景图和结构化查询作为输入,我们通过视角推理获得最佳视角,以解决定位过程中视角缺失的问题。基于获得的最佳视角,我们进一步区分干扰物体,使模型具备区分相似实例的能力。之后,我们通过计算特征向量的相似性,将类别文本和外观图像与查询进行匹配。最后,通过整合视角得分、混淆得分、类别得分和外观得分来识别目标物体。与之前的方法相比,我们的TDVR在视角推理、相似物体区分和模糊查询理解方面具有更强的能力。在公共ScanRefer数据集上的实验结果表明,我们的方法在
[email protected]和
[email protected]上分别比现有的最先进方法提高了15.25%和14.46%,证明了我们的TDVR在解决模糊查询文本和视角不足方面的有效性。
cs.CV / 101 / 2608.03779
AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding
AgenticVAU:用于视频异常理解的多智能体探索-验证推理
Abstract
Video anomaly understanding (VAU) focuses on comprehensively interpreting abnormal events in videos, requiring models to identify anomalous occurrences, discover their supporting evidence, and explain the underlying causes beyond simple anomaly detection. Existing VAU methods often rely on specialized training or limited observations, restricting generalization or evidence coverage. Although single-agent alternatives support adaptive video observation, they still integrate exploration, observation, and decision-making within a unified reasoning process, offering limited role specialization and structured evidence coordination. To address these limitations, we present AgenticVAU, a training-free multi-agent framework that casts VAU as an explore--verify process, where the system first discovers potential anomalies and then verifies them through targeted observations. To achieve this, four specialized agents are introduced to handle visual-rule construction, search planning, video observation, and final decision, respectively. These agents communicate through an anchor registry, a shared evidence memory that binds each observation. Guided by this agent framework, AgenticVAU interleaves broad temporal exploration, dense local verification, and cross-interval comparison until sufficient evidence is collected. We conduct extensive experiments on the ECVA, UCF-Crime, and MSAD subsets of VAU-Bench, the results show that AgenticVAU outperforms zero-shot inference and reinforcement learning-based baselines, demonstrating the value of multi-agent collaboration for video anomaly understanding.
Chinese Translation
视频异常理解(VAU)旨在全面解释视频中的异常事件,这要求模型识别异常发生、发现其支持证据,并解释超越简单异常检测的潜在原因。现有的VAU方法通常依赖于专门的训练或有限的观察,限制了其泛化能力或证据覆盖范围。尽管单智能体的替代方案支持自适应视频观察,但它们仍然在统一的推理过程中整合探索、观察和决策,提供的角色专业化和结构化证据协调有限。为了解决这些局限性,我们提出了AgenticVAU,这是一种无训练的多智能体框架,将VAU视为一个探索-验证过程,其中系统首先发现潜在异常,然后通过针对性的观察进行验证。为此,引入了四个专门的智能体,分别处理视觉规则构建、搜索规划、视频观察和最终决策。这些智能体通过锚点注册表进行通信,这是一个共享的证据记忆,绑定每次观察。在这个智能体框架的指导下,AgenticVAU交替进行广泛的时间探索、密集的局部验证和跨区间比较,直到收集到足够的证据。我们在VAU-Bench的ECVA、UCF-Crime和MSAD子集上进行了广泛的实验,结果表明AgenticVAU在零样本推理和基于强化学习的基线方法中表现优越,展示了多智能体协作在视频异常理解中的价值。
cs.CV / 102 / 2608.03812
OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models
OmniPack:高效全模态大语言模型的统一令牌压缩
Abstract
Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio token sequences incurs substantial computational overhead, demanding aggressive token compression for efficient deployment. Existing methods often degrade at low token budgets: pre-LLM compression may discard structurally important and globally distributed evidence, whereas inner-LLM compression often underexploits query-conditioned audio-visual collaboration. To address these limitations, we propose OmniPack, a training-free framework that coordinates structural compression before the LLM with task-relevant semantic refinement within the LLM. Before the LLM, OmniPack removes structural redundancy through modality-specific importance, global coverage, and similarity-aware merging. After sufficient multimodal interaction, it further consolidates diverse, task-relevant representations through textual guidance and audio-visual collaboration. Extensive experiments on five benchmarks with three Omni-LLM backbones demonstrate that OmniPack consistently achieves the best performance-efficiency trade-off across diverse retention ratios, outperforming all existing methods. Notably, on Qwen2.5-Omni-7B, OmniPack preserves 98.0% of the original performance while reducing FLOPs to 16.7%, and still retains 92.9% of the original performance with only 6.8% of the original FLOPs.
Chinese Translation
全模态大语言模型(Omni-LLMs)在音视频理解任务中取得了显著的性能,但处理长且高度冗余的视觉和音频令牌序列会带来巨大的计算开销,因此需要进行激进的令牌压缩以实现高效部署。现有方法在低令牌预算下往往表现不佳:预处理LLM的压缩可能会丢弃结构上重要且全局分布的证据,而LLM内部的压缩往往未能充分利用基于查询的音视频协作。为了解决这些局限性,我们提出了OmniPack,这是一个无训练的框架,协调LLM之前的结构压缩与LLM内部的任务相关语义精炼。在LLM之前,OmniPack通过模态特定的重要性、全局覆盖和相似性感知合并去除结构冗余。在充分的多模态交互之后,它通过文本指导和音视频协作进一步巩固多样的、与任务相关的表示。在五个基准测试和三个Omni-LLM骨干网络上的大量实验表明,OmniPack在不同的保留比例下始终实现最佳的性能效率平衡,超越了所有现有方法。值得注意的是,在Qwen2.5-Omni-7B上,OmniPack保留了98.0%的原始性能,同时将FLOPs降低到16.7%,并且在仅使用6.8%的原始FLOPs时仍保留92.9%的原始性能。
cs.CV / 103 / 2608.03817
UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space
UHP 检测:LVLM 在一致性空间中具有独特的幻觉模式
Abstract
Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions are not grounded in visual evidence. Existing black-box hallucination detection methods estimate uncertainty through a single consistency metric, implicitly assuming that model uncertainty can be adequately characterized by a single measure. However, hallucinations exhibit diverse manifestations of uncertainty across different behavioral probes, making a single measure insufficient to characterize their underlying behavior. We propose \emph{Unique Hallucination Pattern (UHP) Detection}, a fully black-box framework that models hallucination as a structured uncertainty pattern defined by two axes: perturbation modality (image vs.\ text) and logical polarity (a statement vs.\ its negation). Their intersection produces four complementary consistency groups that capture distinct manifestations of model uncertainty, from which both within-group and between-group features are extracted to train a lightweight classifier. Through comprehensive experiments on AMBER and PhD across three LVLMs, UHP Detection consistently outperforms prior black-box and white-box baselines, with improvements of up to $+18.72\%$ AUC-ROC and $+20.07\%$ AUC-PR over the strongest black-box methods. Extensive ablation studies demonstrate that each consistency group contributes complementary information and that their combination forms a structured hallucination pattern. Furthermore, cross-dataset evaluation shows that this learned pattern generalizes across benchmarks, indicating that hallucination behavior reflects a model-specific consistency pattern. \textbf{Code is publicly available at} https://github.com/amirezzati/uhpdet.
Chinese Translation
大型视觉-语言模型(LVLMs)展现出强大的多模态推理能力,但仍然容易出现幻觉现象,即模型预测未能与视觉证据相结合。现有的黑箱幻觉检测方法通过单一一致性指标来估计不确定性,隐含假设模型的不确定性可以通过单一度量充分表征。然而,幻觉在不同行为探测中表现出多样化的不确定性特征,使得单一度量不足以表征其潜在行为。我们提出了 extit{独特幻觉模式(UHP)检测},这是一个完全的黑箱框架,将幻觉建模为由两个轴定义的结构化不确定性模式:扰动方式(图像与文本)和逻辑极性(陈述与其否定)。它们的交集产生四个互补的一致性组,捕捉模型不确定性的不同表现,从中提取组内和组间特征以训练轻量级分类器。通过在 AMBER 和 PhD 数据集上对三种 LVLM 进行全面实验,UHP 检测始终优于先前的黑箱和白箱基准,AUC-ROC 和 AUC-PR 的提升幅度分别达到 $+18.72\%$ 和 $+20.07\\%$,超越最强的黑箱方法。广泛的消融研究表明,每个一致性组提供互补信息,其组合形成结构化的幻觉模式。此外,跨数据集评估显示,该学习模式在基准测试中具有良好的泛化能力,表明幻觉行为反映了模型特定的一致性模式。 extbf{代码已公开,访问地址为} https://github.com/amirezzati/uhpdet.
cs.CV / 104 / 2608.03822
FlowForm: Synergizing Fluid Physics with Topological Consistency for Satellite Flood Synthesis
FlowForm:将流体物理与拓扑一致性协同用于卫星洪水合成
Abstract
Developing robust flood assessment models requires high-quality paired satellite imagery, yet such data remain scarce for flood-specific image generation. Although generative models provide a promising means of data augmentation, existing methods often yield implausible spatial layouts of flooded regions and distort scene structures. We propose FlowForm, a framework for satellite flood synthesis that integrates SWE-inspired latent regularization with structure-aware conditioning. The Flood Descriptor Module (FDM) imposes differentiable penalties on residuals of the steady-state Shallow Water Equation in auxiliary latent fields at the diffusion bottleneck. The Terrain Anchor Adapter (TAA) injects depth, semantic, and edge features at four encoder scales of the U-Net. We further curate FloodScape, a large-scale, high-resolution dataset comprising paired satellite images acquired before and after disasters. In addition to standard image-generation metrics, we evaluate the consistency of flooded regions, zero-shot generalization to a geographically held-out flood event, and sensitivity to individual components. Across all reported comparisons, FlowForm achieves higher visual fidelity, greater similarity between paired images, and stronger consistency of flooded regions.
Chinese Translation
开发稳健的洪水评估模型需要高质量的配对卫星影像,然而此类数据在洪水特定图像生成中仍然稀缺。虽然生成模型提供了一种有前景的数据增强手段,但现有方法往往产生不合理的洪水区域空间布局,并扭曲场景结构。我们提出了FlowForm,一个用于卫星洪水合成的框架,该框架将受浅水方程(SWE)启发的潜在正则化与结构感知条件相结合。洪水描述模块(Flood Descriptor Module, FDM)在扩散瓶颈的辅助潜在场中对稳态浅水方程的残差施加可微分的惩罚。地形锚定适配器(Terrain Anchor Adapter, TAA)在U-Net的四个编码器尺度上注入深度、语义和边缘特征。我们进一步整理了FloodScape,一个大规模高分辨率数据集,包含灾难前后获取的配对卫星图像。除了标准的图像生成指标外,我们还评估了洪水区域的一致性、对地理上保留的洪水事件的零样本泛化能力以及对各个组件的敏感性。在所有报告的比较中,FlowForm实现了更高的视觉保真度、更大的配对图像相似性以及更强的洪水区域一致性。
cs.CV / 105 / 2608.03826
Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding
Geo-Embed:迈向统一的城市理解多模态嵌入
Abstract
Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues. However, existing multimodal embedding models and benchmarks are still largely designed and evaluated around general-purpose image-text matching, leaving unclear whether unified embedding space can support heterogeneous geospatial tasks involving spatial relationships, fine-grained semantics, and temporal changes. To address this gap, we make three key contributions. First, we introduce GeoMEB, a large-scale multimodal embedding benchmark that standardizes 45 urban evaluation tasks across retrieval, visual question answering, change detection, classification, and visual grounding, together with training collections comprising 1.32M examples and 286K evaluation queries. Second, we present Geo-Embed, a unified embedding model that adapts a shared vision-language backbone to instruction-conditioned query-target matching over heterogeneous geospatial inputs, including single images, multiple images, text, regions, and masks. On GeoMEB, Geo-Embed achieves the strongest overall performance among representative multimodal embedders, with a 15.3% relative improvement over the strongest baseline. These results motivate future geospatial embedders that organize training and evaluation around explicit query-target relations, including semantic, cross-view, region-level, and temporal correspondence.
Chinese Translation
地理空间和城市应用日益需要模型在街景图像、遥感观测、文本描述、区域提议和时间变化线索之间比较异构证据。然而,现有的多模态嵌入模型和基准仍然主要围绕通用图像-文本匹配进行设计和评估,这使得尚不清楚统一的嵌入空间是否能够支持涉及空间关系、细粒度语义和时间变化的异构地理空间任务。为了解决这一问题,我们做出了三项关键贡献。首先,我们介绍了GeoMEB,这是一个大规模的多模态嵌入基准,标准化了在检索、视觉问答、变化检测、分类和视觉定位等方面的45个城市评估任务,并提供了包含132万个示例和28.6万个评估查询的训练集合。其次,我们提出了Geo-Embed,这是一个统一的嵌入模型,适应于通过共享的视觉-语言主干进行异构地理空间输入的指令条件查询-目标匹配,包括单幅图像、多幅图像、文本、区域和掩码。在GeoMEB上,Geo-Embed在代表性的多模态嵌入模型中实现了最强的整体性能,相较于最强基线有15.3%的相对提升。这些结果激励未来的地理空间嵌入模型围绕明确的查询-目标关系进行训练和评估,包括语义、跨视图、区域级和时间对应关系。
cs.CV / 106 / 2608.03851
LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation
LiteMVS:基于基础蒸馏和专家聚合的高效多视图立体视觉
Abstract
Real-time 3D perception is crucial for robotics, augmented reality, and embodied intelligence applications. Existing multi-view stereo (MVS) methods primarily rely on geometric correspondences, which often fail in textureless or repetitive regions, while monocular depth models leverage strong image-level priors but lack robust multi-view geometric constraints. More importantly, in robotics and embodied manipulation scenarios, high-quality 3D geometry is not only essential for static reconstruction, but also serves as a critical foundation for learning temporally consistent 4D representations. To obtain visual representations with stronger structural awareness and greater potential for spatiotemporal extension, we present LiteMVS, a lightweight multi-view depth estimation model that integrates plane-sweep geometric reasoning with strong monocular semantic and structural priors. The central idea of LiteMVS is to efficiently inject high-level monocular knowledge, obtained from lightweight segmentation models and large-scale vision foundation models, into a multi-view stereo framework. In particular, LiteMVS enriches the cost volume with semantic descriptors and employs a Mixture-of-Experts (MoE) formulation to enable adaptive geometric aggregation across depth hypotheses. Moreover, geometric priors distilled from vision foundation models further strengthen monocular guidance without increasing inference cost. Through this design, LiteMVS not only improves depth estimation and 3D reconstruction quality in static scenes, but also provides a more reliable geometric foundation for subsequent temporal modeling and 4D representation learning. Experiments on ScanNetv2 and 7-Scenes demonstrate that LiteMVS achieves high-quality depth prediction and 3D reconstruction while maintaining competitive efficiency.
Chinese Translation
实时三维感知对于机器人技术、增强现实和具身智能应用至关重要。现有的多视图立体视觉(MVS)方法主要依赖几何对应关系,这在无纹理或重复区域往往失效,而单目深度模型则利用强大的图像级先验,但缺乏稳健的多视图几何约束。更重要的是,在机器人和具身操作场景中,高质量的三维几何不仅对静态重建至关重要,而且为学习时间一致的四维表示提供了关键基础。为了获得具有更强结构意识和更大时空扩展潜力的视觉表示,我们提出了LiteMVS,这是一种轻量级多视图深度估计模型,结合了平面扫掠几何推理与强大的单目语义和结构先验。LiteMVS的核心思想是高效地将从轻量级分割模型和大规模视觉基础模型中获得的高级单目知识注入多视图立体框架中。具体而言,LiteMVS通过语义描述符丰富了代价体,并采用混合专家(Mixture-of-Experts, MoE)形式实现深度假设间的自适应几何聚合。此外,从视觉基础模型中提炼的几何先验进一步增强了单目引导,而不增加推理成本。通过这种设计,LiteMVS不仅提高了静态场景中的深度估计和三维重建质量,还为后续的时间建模和四维表示学习提供了更可靠的几何基础。在ScanNetv2和7-Scenes上的实验表明,LiteMVS在保持竞争效率的同时,实现了高质量的深度预测和三维重建。
cs.CV / 107 / 2608.03863
CPrefix: A Combinatorial Tensor Framework for Structured Discrete Color Mappings
CPrefix:一种用于结构化离散颜色映射的组合张量框架
Abstract
Discrete multi-channel mappings are typically represented through sampled values, providing accurate evaluations but limited insight into their underlying structure. We introduce CPrefix, a combinatorial observable representation for discrete mappings, realized within a unified tensor framework that enables representation, reconstruction, and structural analysis. The framework is based on a counting tensor induced by multinomial counting observables. Its support forms a discrete Pascal simplex, not as a constraint on the observable space, but as a latent combinatorial representation from which mappings are reconstructed. This formulation separates the combinatorial organization of a mapping from its measured values, exposing the observable structure underlying the mapping. The framework is validated on ICC display and printer profiles through latent reconstruction and perceptual gamut transport. Accurate reconstruction demonstrates that color mappings admit faithful observable representations, while reconstruction residuals provide insight into the compatibility of the underlying mapping with the proposed representation. Although demonstrated on color transformations, the framework is independent of the physical interpretation of the observables, making it applicable to structured multi-channel mappings arising from color imaging, spectral measurements and other discrete systems.
Chinese Translation
离散多通道映射通常通过采样值表示,提供准确的评估,但对其潜在结构的洞察有限。我们提出了CPrefix,一种用于离散映射的组合可观测表示,实现在一个统一的张量框架内,该框架支持表示、重建和结构分析。该框架基于由多项式计数可观测量引发的计数张量。其支持形成一个离散的帕斯卡尔单纯形,不是对可观测空间的约束,而是作为一个潜在的组合表示,从中重建映射。这一表述将映射的组合组织与其测量值分开,揭示了映射背后的可观测结构。该框架通过潜在重建和感知色域传输在ICC显示器和打印机配置文件上得到了验证。准确的重建表明颜色映射允许忠实的可观测表示,而重建残差则提供了对潜在映射与所提出表示兼容性的洞察。尽管在颜色变换上进行了演示,但该框架独立于可观测量的物理解释,使其适用于来自颜色成像、光谱测量和其他离散系统的结构化多通道映射。
cs.CV / 108 / 2608.03884
BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for OCR and Vision-Language Models
BanglaWild:一个针对光学字符识别和视觉语言模型的野外孟加拉场景文本识别基准
Abstract
In-the-wild Bengali scene text recognition is largely unmeasured: existing resources target handwritten documents or constrained sign-board parsing, report only aggregate edit-distance metrics, and evaluate either conventional OCR or VLMs, never both on the same in-the-wild data. To address this gap, we introduce BANGLAWILD, a benchmark of 2,535 Bengali scene text images, each paired with a verbatim gold transcription, two categorical axes, four diagnostic attributes, and an orthographically standard form where the in-image text deviates from canonical spelling. We evaluate fifteen VLMs and three conventional OCR systems under three prompting strategies, fine-tune 6 open-source models with LoRA, and complement edit-distance metrics with an LLM-as-a-Judge evaluation. Our results reveal a persistent gap in which larger models within the same family do not outperform smaller ones. Our fifteen-class error taxonomy shows that visual mis-recognition accounts for ~60% of errors in the strongest systems, while conjunct-related errors contribute under 2%, challenging a long-standing assumption in Bengali OCR research; the same visual dominant profile also holds across architectures, including the one conventional baseline that reads Bengali reliably. Prompt language mainly affects cross-script drift and LoRA reduces catastrophic failures in weak models without lifting the ceiling on already competent ones. Code and data will be publicly released.
Chinese Translation
野外孟加拉场景文本识别在很大程度上尚未被测量:现有资源主要针对手写文档或受限的标识牌解析,仅报告聚合的编辑距离指标,并且评估传统的光学字符识别(OCR)或视觉语言模型(VLMs),从未在同一野外数据上同时进行评估。为了解决这一空白,我们引入了BANGLAWILD,一个包含2,535张孟加拉场景文本图像的基准,每张图像都配有逐字的金标准转录、两个分类轴、四个诊断属性,以及在图像中偏离规范拼写的正字法标准形式。我们在三种提示策略下评估了十五个VLM和三个传统OCR系统,使用LoRA对6个开源模型进行了微调,并通过LLM-as-a-Judge评估补充了编辑距离指标。我们的结果揭示了一个持续存在的差距,即同一家族中较大的模型并未优于较小的模型。我们的十五类错误分类法显示,视觉误识别占据了最强系统中约60%的错误,而与结合相关的错误贡献不足2%,挑战了孟加拉OCR研究中的一个长期假设;同样的视觉主导特征在各个架构中也保持一致,包括一个可靠读取孟加拉文的传统基线。提示语言主要影响跨脚本漂移,而LoRA在弱模型中减少了灾难性失败,而并未提升已经具备能力的模型的上限。代码和数据将公开发布。
cs.CV / 109 / 2608.03885
MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization
MuRA:用于高效有效的测试时视觉-语言泛化的多等级适应
Abstract
Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-time adaptation (TTA) via Low-Rank Adaptation offers a parameter-efficient solution, we identify a fundamental bottleneck in current methods: the reliance on static rank configurations. Because visual inputs inherently possess varying information densities, a fixed rank forces an inevitable optimization compromise, leading to underfitting on complex scenes and overfitting on simple ones. To bridge this gap, we propose Multi-Rank Adaptation (MuRA), a novel framework that dynamically selects and fuses adaptation modules of varying capacities based on token-level visual complexity. MuRA synergizes Multi-Rank Orthogonal Decomposition to provide a superior, knowledge-preserving initialization, and Unified Component Fusion with Continuous Router Updating to sustainably learn semantic-to-rank mappings. Furthermore, we provide rigorous theoretical justifications mathematically proving the necessity and gradient stability of this adaptive mechanism. Crucially, MuRA's dynamic design uniquely thrives at the deepest visual layer, capitalizing on the shortest gradient backpropagation path. Extensive experiments demonstrate that MuRA achieves state-of-the-art accuracy across extensive domain generalization and cross-dataset benchmarks while significantly reducing both computational and memory overhead.
Chinese Translation
视觉-语言模型展现出显著的零-shot 能力,但在分布变化下性能显著下降。虽然通过低秩适应(Low-Rank Adaptation)的测试时适应(TTA)提供了一种参数高效的解决方案,但我们发现当前方法存在一个根本瓶颈:依赖于静态秩配置。由于视觉输入本质上具有不同的信息密度,固定的秩会导致不可避免的优化妥协,导致在复杂场景下欠拟合,而在简单场景下过拟合。为了解决这一问题,我们提出了多等级适应(Multi-Rank Adaptation,MuRA),这是一个新颖的框架,能够根据标记级别的视觉复杂性动态选择和融合不同容量的适应模块。MuRA 协同使用多等级正交分解(Multi-Rank Orthogonal Decomposition)提供优越的、知识保留的初始化,并结合统一组件融合(Unified Component Fusion)与连续路由更新(Continuous Router Updating)可持续地学习语义到秩的映射。此外,我们提供了严格的理论证明,数学上证明了这一自适应机制的必要性和梯度稳定性。重要的是,MuRA 的动态设计在最深的视觉层中独特地发挥作用,利用最短的梯度反向传播路径。大量实验表明,MuRA 在广泛的领域泛化和跨数据集基准测试中实现了最先进的准确性,同时显著降低了计算和内存开销。
cs.CV / 110 / 2608.03890
CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
CARE-X:通过辅助监督、奖励对齐学习和工具增强测量,迈向临床实用的放射学视觉语言模型
Abstract
A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend. Today's Vision-Language Models (VLMs) treat these as separate problems, if they address them at all, leaving a gap between what radiologists need and what generative models provide. We introduce CARE-X, a chest X-ray VLM that narrows this gap by unifying auxiliary discriminative supervision with reward-aligned generation. CARE-X augments its generative backbone with focal-loss classification and composite-loss grounding heads, co-trained alongside the language-modeling objective. This auxiliary supervision produces discriminative diagnostic predictions with tunable decision thresholds and precise spatial localization while also improving report quality, providing evidence that structured prediction and generation reinforce one another. Building on this foundation, Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) leverages task-specific reward signals for report generation, visual question answering (VQA), and spatial grounding, directly optimizing the clinical quality metrics that matter in practice. The result is state-of-the-art performance on the majority of metrics across four report-generation benchmarks, 94.0% VQA accuracy on ReXVQA (+6.0 pp over the next-best baseline), and generative spatial decoding that reaches near parity with dedicated detection heads. Separately, to address measurement-dependent diagnoses, we couple Qwen3-VL-4B-Instruct with native tool-calling capabilities for invoking deterministic measurement tools, while retaining full visual access to the image. This hybrid inference yields +43.6 pp average F1 over perception-only baselines across five measurement-dependent conditions.
Chinese Translation
一个临床实用的胸部X光系统必须超越流畅的报告生成:它应该能够以可调的决策阈值对发现进行分类,进行空间定位,并推导出许多诊断所依赖的解剖测量。如今的视觉语言模型(VLMs)将这些视为独立的问题,如果它们有涉及的话,往往无法满足放射科医生的需求,导致放射科医生所需与生成模型所提供之间存在差距。我们提出了CARE-X,一个胸部X光VLM,通过将辅助判别监督与奖励对齐生成相结合,缩小了这一差距。CARE-X通过聚焦损失分类和复合损失定位头增强其生成骨干,与语言建模目标共同训练。这种辅助监督产生了具有可调决策阈值和精确空间定位的判别诊断预测,同时改善了报告质量,提供了结构化预测与生成相互强化的证据。在此基础上,解耦剪辑和动态采样策略优化(DAPO)利用任务特定的奖励信号进行报告生成、视觉问答(VQA)和空间定位,直接优化在实践中重要的临床质量指标。其结果是在四个报告生成基准测试中大多数指标上达到了最先进的性能,在ReXVQA上实现了94.0%的VQA准确率(比下一个最佳基线高出6.0个百分点),并且生成的空间解码达到了与专用检测头几乎相当的水平。此外,为了解决依赖测量的诊断问题,我们将Qwen3-VL-4B-Instruct与原生工具调用能力结合,以调用确定性测量工具,同时保留对图像的完全视觉访问。这种混合推理在五个依赖测量的条件下,相较于仅感知基线平均提高了43.6个百分点的F1得分。
cs.CV / 111 / 2608.03895
NCGR: Noise-Conditional Gated Rectification for Camera Extrinsic Perturbations in BEV 3D Object Detection
NCGR:用于鸟瞰视图(BEV)3D物体检测中相机外部扰动的噪声条件门控整流
Abstract
Camera-based bird's-eye-view (BEV) 3D detection typically assumes accurate and fixed camera extrinsics. In detectors using spatial cross-attention (SCA), extrinsic perturbations displace the image-plane projections of BEV reference points, causing queries to sample features from incorrect regions and degrading detection performance. To address this failure mode, Noise-Conditional Gated Rectification (NCGR) is proposed to compensate for projection errors without explicitly estimating a full six-degree-of-freedom extrinsic correction. For each query-camera pair, a 2D rectification offset is predicted and modulated by a camera-level gate to rectify the base projection before native deformable sampling. During training, the perturbation-derived quantities used to construct the condition and gate are gradually replaced through scheduled interpolation by counterparts generated from an auxiliary scalar predicted from camera features. This transition enables blind inference without perturbation metadata. During training, a weight-shared clean-teacher/perturbed-student pair is used, and the rectification module is supervised by a BEV-consistency objective between the two branches. NCGR is evaluated on nuScenes with simulated dynamic and static extrinsic perturbations. In a five-camera dynamic stress test, NCGR achieves 39.69% NDS, compared with 28.00% for BEVFormer and 33.23% for CAPE. Under clean extrinsics, NCGR maintains performance comparable to that of BEVFormer.
Chinese Translation
基于相机的鸟瞰视图(BEV)3D检测通常假设相机外部参数准确且固定。在使用空间交叉注意力(SCA)的检测器中,外部扰动会使BEV参考点的图像平面投影发生位移,从而导致查询从错误区域采样特征,降低检测性能。为了解决这一失败模式,提出了噪声条件门控整流(NCGR),以补偿投影误差,而无需显式估计完整的六自由度外部校正。对于每个查询-相机对,预测一个2D整流偏移,并通过相机级别的门控进行调制,以在本地可变采样之前整流基础投影。在训练过程中,用于构建条件和门控的扰动派生量通过调度插值逐渐被从相机特征生成的辅助标量的对应物替代。这一过渡使得在没有扰动元数据的情况下进行盲推理成为可能。在训练中,使用权重共享的干净教师/扰动学生对,并通过两个分支之间的BEV一致性目标对整流模块进行监督。NCGR在nuScenes上进行了评估,测试了模拟的动态和静态外部扰动。在五个相机的动态压力测试中,NCGR达到了39.69%的NDS,而BEVFormer为28.00%,CAPE为33.23%。在干净的外部参数下,NCGR保持了与BEVFormer相当的性能。
cs.CV / 112 / 2608.03911
UniEvo-RS: Omni-Prompt Unified Remote Sensing Segmentation with Representative Exemplar-Driven Prototype Evolution
UniEvo-RS:基于代表性样本驱动的原型演化的全提示统一遥感分割
Abstract
Prompt-driven vision-language models (VLMs) hold immense promise for accelerating dense remote sensing (RS) annotation, but static models suffer from severe performance degradation when deployed on novel scenes, unseen categories, or visually confusing backgrounds. Moreover, existing unified paradigms primarily rely on intra-image specific prompts, lacking flexible task routing to adapt to multi-intent operational workflows. In practical batch mapping, annotators typically refine a small set of representative samples before processing large datasets. Motivated by this practice, we propose UniEvo-RS, an omni-prompt unified RS segmentation framework equipped with representative exemplar-driven prototype evolution. First, we construct a multi-instruction prompt dataset that unifies text-driven and visual-driven prompts within a single architecture, establishing a dynamic task-routing mechanism for highly diverse RS annotation scenarios. Second, we introduce a representative feedback-driven, training-free prototype evolution mechanism. By contrasting manual annotations with initial predictions on exemplars, UniEvo-RS distills prediction errors into positive and negative prototypes. These prototypes enhance LLM query recall and suppress spatial background noise under a fixed-budget clustering memory. Extensive experiments show that UniEvo-RS unifies diverse prompting tasks, achieving state-of-the-art performance across most settings. Crucially, with minimal interaction on a few exemplars, it enables training-free, progressive accuracy enhancement on unseen categories during batch annotation.
Chinese Translation
基于提示的视觉-语言模型(VLMs)在加速密集遥感(RS)标注方面展现出巨大潜力,但静态模型在新场景、未见类别或视觉上混淆的背景中部署时,性能严重下降。此外,现有的统一范式主要依赖于图像内部特定的提示,缺乏灵活的任务路由以适应多意图的操作工作流。在实际的批量映射中,标注者通常在处理大数据集之前,先对一小部分代表性样本进行细化。受到这一实践的启发,我们提出了UniEvo-RS,这是一种配备代表性样本驱动的原型演化的全提示统一遥感分割框架。首先,我们构建了一个多指令提示数据集,将文本驱动和视觉驱动的提示统一在一个架构中,建立了一个动态任务路由机制,以应对高度多样化的遥感标注场景。其次,我们引入了一种代表性反馈驱动的、无训练的原型演化机制。通过对比手动标注与样本的初始预测,UniEvo-RS将预测错误提炼为正负原型。这些原型增强了大语言模型(LLM)查询的召回率,并在固定预算的聚类记忆下抑制空间背景噪声。大量实验表明,UniEvo-RS统一了多样化的提示任务,在大多数设置中实现了最先进的性能。重要的是,通过对少量样本的最小交互,它在批量标注过程中实现了对未见类别的无训练、渐进式准确性提升。
cs.CV / 113 / 2608.03912
StreamDAM: Presence-Aware Memory for Real-Time Streaming Video Object Segmentation
StreamDAM:实时流媒体视频目标分割的感知记忆
Abstract
Quality-tier video object segmentation (VOS) trackers such as DAM4SAM top accuracy leaderboards, but they are measured offline, one frame at a time with no clock. Under an honest streaming protocol at 30 frames per second, where a frame that misses its budget is served the last mask already computed, the winner collapses: the rich memory that makes it accurate is too slow to keep up, and what it emits is blind to whether the object is even present. We trace both failures to one place, the tracker's memory pipeline, and rebuild it for streaming. \method{} makes the memory machinery itself run at frame rate through in-model optimization rather than a bolted-on fallback, and governs it with a single learned presence signal that decides what enters memory, how far back the tracker reads, when to withhold output, and when to re-detect. A mechanism analysis shows why a fixed policy cannot win: the control that helps when an object truly disappears is the one that hurts when it is merely hard to see, so the choice must be made per frame. Across four benchmarks and five modern baselines, \method{} is the strongest streaming tracker, recovers nearly all of the offline model's accuracy under the clock, and on the hardest content exceeds the offline model it is built from.
Chinese Translation
高质量的视频目标分割(VOS)跟踪器,如DAM4SAM,在准确性排行榜上名列前茅,但它们是离线测量的,一次处理一帧,没有时间限制。在每秒30帧的诚实流媒体协议下,如果一帧未能在预算内处理,则会使用最后计算出的掩膜,最终赢家崩溃:使其准确的丰富记忆太慢,无法跟上,而其输出对对象是否存在毫无察觉。我们将这两种失败归结为一个地方,即跟踪器的记忆管道,并为流媒体重建它。 extit{method} 使得记忆机制本身通过模型内优化以帧速率运行,而不是依赖于附加的后备方案,并通过一个单一的学习到的存在信号来管理它,该信号决定什么进入记忆、跟踪器读取多远、何时抑制输出以及何时重新检测。机制分析表明,固定策略无法胜出:当对象真正消失时有帮助的控制,反而在对象仅仅难以看见时会造成伤害,因此必须逐帧做出选择。在四个基准和五个现代基线中, extit{method} 是最强的流媒体跟踪器,在时钟下几乎恢复了离线模型的所有准确性,并且在最困难的内容上超越了其构建的离线模型。
cs.CV / 114 / 2608.03918
When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding
何时何地观察:高效长视频理解的自适应视觉证据调度
Abstract
Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance-based methods rely on static one-shot selection with fixed frame budgets and candidate pools, while agent-based schedulers achieve adaptivity through costly multi-round reasoning and interactive search. We propose EcoFrame, a training-free framework for low-overhead query-adaptive visual evidence scheduling. EcoFrame leverages the VLM's inference feedback to determine when to increase the frame budget and where to search for additional candidate evidence. Specifically, entropy-gated budget scheduling uses output uncertainty to stop early when the current evidence is sufficient or progressively expand the frame budget otherwise. Meanwhile, attention-guided candidate proposal converts frame-level attention into a temporal prior, enabling dense local search in informative regions while preserving global coverage when attention is diffuse. Experiments on Video-MME, LongVideoBench, and MLVU demonstrate that EcoFrame achieves a better accuracy--efficiency trade-off across multiple VLM backbones. On Qwen2.5-VL, EcoFrame achieves an average accuracy of 64.4, surpassing BOLT at 63.5, while providing a $1.85\times$ speedup over AKS and BOLT. Compared with the agent-based A.I.R., EcoFrame maintains comparable accuracy with up to a $13.5\times$ inference speedup. Code will be available at https://github.com/AK-DREAM/EcoFrame.
Chinese Translation
高效的长视频理解需要视觉-语言模型(VLM)对选定的少量帧进行推理,这些帧被视为稀疏视觉证据。现有的基于相关性的算法依赖于静态的一次性选择,使用固定的帧预算和候选池,而基于代理的调度器则通过代价高昂的多轮推理和交互搜索实现自适应。我们提出了EcoFrame,一个无需训练的低开销查询自适应视觉证据调度框架。EcoFrame利用VLM的推理反馈来确定何时增加帧预算以及在哪里搜索额外的候选证据。具体而言,熵门控预算调度利用输出的不确定性,在当前证据足够时提前停止,或在其他情况下逐步扩大帧预算。同时,注意力引导的候选提议将帧级注意力转换为时间先验,使得在信息丰富的区域进行密集局部搜索,同时在注意力分散时保持全局覆盖。对Video-MME、LongVideoBench和MLVU的实验表明,EcoFrame在多个VLM骨干网络中实现了更好的准确性与效率的权衡。在Qwen2.5-VL上,EcoFrame达到了64.4的平均准确率,超过了63.5的BOLT,同时提供了相较于AKS和BOLT的$1.85 imes$加速。与基于代理的A.I.R.相比,EcoFrame在保持可比准确率的同时实现了高达$13.5 imes$的推理加速。代码将发布在https://github.com/AK-DREAM/EcoFrame。
cs.CV / 115 / 2608.03919
Low-Dimensional High-Leverage Subspace Optimization: Beyond Full-Parameter Coupled Training for Neural Network Quantization
低维高杠杆子空间优化:超越神经网络量化的全参数耦合训练
Abstract
Low-bit quantization suffers severe accuracy degradation on compact networks, rooted in the dominant full-parameter coupled training paradigm that ignores parameter subspace heterogeneity. Their limited feature redundancy leaves little room to absorb quantization errors. Conventional pipelines adopt monolithic optimization: PTQ reconstructs fixed pretrained models without improving inherent quantization friendliness; QAT updates all parameters jointly, suffering from gradient coupling between backbone weights and calibration parameters. In this paper, we identify normalization affine parameters as a low-dimensional high-leverage subspace dominating quantization robustness, and propose Normalization Affine Preconditioning (NAP) for targeted subspace optimization. For PTQ, NAP freezes backbone weights and fine-tunes only affine parameters under the target fake-quantization graph on full-precision models, proactively boosting quantization friendliness before downstream reconstruction. For QAT, we introduce an alternating QAT-NAP schema that decouples feature learning and numerical calibration, breaking the performance ceiling of saturated joint training. Theoretical analysis confirms BN affine parameters fully cancel the channel-wise affine component of quantization distortion, while nonlinear rounding and clipping residuals form the irreducible error boundary; distillation-guided NAP acts as directional flatness optimization, projecting teacher-student logit mismatch onto the restricted subspace. Experiments on ImageNet and CIFAR-100 show NAP recovers severely collapsed low-bit quantization, consistently boosts reconstruction-based PTQ, and outperforms saturated full-parameter QAT with negligible tuning cost. This work reveals the principle of targeted low-dimensional subspace optimization, offering a new perspective beyond full-parameter coupled training for efficient deep learning.
Chinese Translation
低比特量化在紧凑网络上遭遇严重的准确性下降,这源于主导的全参数耦合训练范式,该范式忽视了参数子空间的异质性。它们有限的特征冗余几乎没有空间来吸收量化误差。传统的流程采用单一优化:PTQ(后训练量化)重建固定的预训练模型,而没有改善固有的量化友好性;QAT(量化感知训练)共同更新所有参数,遭受主干权重和校准参数之间的梯度耦合。在本文中,我们识别出归一化仿射参数作为主导量化鲁棒性的低维高杠杆子空间,并提出归一化仿射预调(NAP)以进行针对性子空间优化。对于PTQ,NAP冻结主干权重,仅在全精度模型的目标假量化图下微调仿射参数,主动提升量化友好性,以便在下游重建之前进行优化。对于QAT,我们引入交替的QAT-NAP方案,解耦特征学习和数值校准,打破饱和联合训练的性能上限。理论分析确认BN(批归一化)仿射参数完全抵消了量化失真在通道方向的仿射成分,而非线性舍入和裁剪残差形成不可减少的误差边界;蒸馏引导的NAP作为方向性平坦度优化,将教师-学生的logit不匹配投影到限制子空间上。在ImageNet和CIFAR-100上的实验表明,NAP恢复了严重崩溃的低比特量化,持续提升基于重建的PTQ,并以微不足道的调优成本超越饱和的全参数QAT。这项工作揭示了针对性低维子空间优化的原理,为高效深度学习提供了超越全参数耦合训练的新视角。
cs.CV / 116 / 2608.03923
GeoMAR: Unleashing Geometrically Aligned Features for Masked Autoregressive Blind Face Restoration
GeoMAR:释放几何对齐特征以进行掩蔽自回归盲人脸修复
Abstract
Codebook-based blind face restoration (BFR) often suffers from ambiguous conditioning features and a fragile prediction mechanism under severe degradation. To address these challenges, we propose GeoMAR, a framework designed to unleash geometrically aligned features with masked autoregressive (MAR) refinement for robust face restoration. For feature conditioning, we introduce a dual-input extraction pipeline to extract component-based geometric descriptions with explicit, spatially faithful anchors. These textual priors are integrated with low-quality (LQ) features via an Aligned Geometric Priors Injector, which employs a KV-Q exchange strategy to generate geometrically aligned features. For prediction mechanism, we reformulate the one-step mapping into a multi-step MAR process. This coarse-to-fine generation progressively refines complex facial regions based on increasingly reliable context. Experiments on one synthetic and three real-world benchmarks demonstrate that GeoMAR achieves highly competitive perceptual quality and coherent visual structures compared with existing methods. The code is available at https://github.com/BRL-SYSU/GeoMAR.git.
Chinese Translation
基于码本的盲人脸修复(BFR)在严重退化情况下常常面临模糊的条件特征和脆弱的预测机制。为了解决这些挑战,我们提出了GeoMAR,一个旨在释放几何对齐特征的框架,结合掩蔽自回归(MAR)精炼以实现稳健的人脸修复。对于特征条件化,我们引入了一种双输入提取管道,以提取基于组件的几何描述,并提供明确且空间上忠实的锚点。这些文本先验通过对齐几何先验注入器(Aligned Geometric Priors Injector)与低质量(LQ)特征相结合,该注入器采用KV-Q交换策略生成几何对齐特征。对于预测机制,我们将一步映射重新构造为多步MAR过程。这种粗到细的生成逐步基于日益可靠的上下文精炼复杂的面部区域。在一个合成基准和三个真实世界基准上的实验表明,与现有方法相比,GeoMAR在感知质量和视觉结构一致性方面达到了高度竞争的水平。代码可在 https://github.com/BRL-SYSU/GeoMAR.git 获取。
cs.CV / 117 / 2608.03937
Progressive Learning of a Diffusion-based Inpainting Model for Separating Overlapped Fingerprints
基于扩散的逐步学习模型用于分离重叠指纹
Abstract
Overlapped friction ridge patterns are a recurring problem in latent fingerprints recovered from crime scenes and in live-scan scenarios where residual fingerprints on the sensor may corrupt subsequent acquisitions. Existing approaches for separating overlapped fingerprints either rely on rule-based orientation field completion that requires strong domain knowledge or train end-to-end deep neural networks that do not account for domain-specific considerations. This work introduces a diffusion-based pipeline for separating component fingerprints from an image containing overlapping friction ridge patterns. We formulate the separation problem as an inpainting task and progressively learn a diffusion model for this task in multiple stages. Starting from a pre-trained Stable Diffusion model, we progressively incorporate a fingerprint prior, add the ability to complete partial fingerprints, and finally propose \textbf{overlap-aware inpainting} that reconstructs each component print using a diffusion inpainting model based on multi-channel conditioning. Experiments on two public datasets demonstrate that component fingerprints reconstructed using the proposed diffusion-based inpainting method can match with their mated counterparts with very high probability.
Chinese Translation
重叠的摩擦脊模式是从犯罪现场恢复的潜在指纹和在实时扫描场景中传感器上残留指纹可能会干扰后续采集的一个反复出现的问题。现有的分离重叠指纹的方法要么依赖于需要强大领域知识的基于规则的方向场补全,要么训练不考虑领域特定因素的端到端深度神经网络。本研究提出了一种基于扩散的管道,用于从包含重叠摩擦脊模式的图像中分离组件指纹。我们将分离问题表述为一个修复任务,并在多个阶段逐步学习该任务的扩散模型。从预训练的稳定扩散模型开始,我们逐步引入指纹先验,增加补全部分指纹的能力,最后提出了 extbf{重叠感知修复},该方法基于多通道条件使用扩散修复模型重建每个组件指纹。在两个公共数据集上的实验表明,使用所提出的基于扩散的修复方法重建的组件指纹可以以非常高的概率与其配对指纹匹配。
cs.CV / 118 / 2608.03971
UniWorld-Design: From Pixel Generation to Layer-Native Design
UniWorld-Design:从像素生成到层原生设计
Abstract
We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA layers as the atomic units of generation, understanding, and editing. Our key insight is that pixels define how an image is rendered, whereas layers define how an image is created, understood, and edited. Just as human designers create and manipulate visual content through layers rather than raw pixels, UniWorld-Design equips multimodal generative models with a layer-native design space. UniWorld-Design comprises two models. The Text-to-RGBA (T2RGBA) model generates standalone RGBA assets directly from text. The Image-to-Layer (I2L) model conditions on a finished image, a global instruction and per-layer prompts, and jointly produces ordered, complete semantic RGBA layers. Its instruction interface supports top-level decomposition, recursive decomposition and targeted extraction, making layering an instruction-addressable operation for agentic editing. Because I2L learns complete semantic objects rather than visible-pixel partitions, its layers stay usable when moved or removed. On the Crello benchmark, I2L reduces per-layer RGB L1 error by 37% and achieves a 34% relative improvement in Alpha Soft IoU over Qwen-Image-Layered. Separately, T2RGBA achieves the highest CLIP Score, outperforming LayerDiffuse and OmniAlpha.
Chinese Translation
我们介绍了UniWorld-Design,这是一个重新定义图像生成的框架,从平面像素合成转变为结构化视觉构图,以语义RGBA层作为生成、理解和编辑的基本单元。我们的关键见解在于,像素定义了图像的渲染方式,而层则定义了图像的创建、理解和编辑方式。正如人类设计师通过层而非原始像素创建和操控视觉内容,UniWorld-Design为多模态生成模型提供了一个层原生设计空间。UniWorld-Design包含两个模型。文本到RGBA(Text-to-RGBA,T2RGBA)模型直接从文本生成独立的RGBA资产。图像到层(Image-to-Layer,I2L)模型以完成的图像、全局指令和每层提示为条件,共同生成有序的、完整的语义RGBA层。其指令接口支持顶层分解、递归分解和目标提取,使得分层成为一个可指令寻址的操作,便于代理编辑。由于I2L学习的是完整的语义对象而非可见像素分区,其层在移动或移除时仍然可用。在Crello基准测试中,I2L将每层RGB L1误差降低了37%,并在Alpha Soft IoU上相较于Qwen-Image-Layered实现了34%的相对提升。此外,T2RGBA在CLIP评分上表现最佳,超越了LayerDiffuse和OmniAlpha。
cs.CV / 119 / 2608.03974
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
JoyAI-视频编辑:基于自回归扩散的实时开放式视频编辑
Xiao, Yicheng, Dai, Wenxun, Qin, Xinran, Song, Lin, Zhang, Maoquan, Xu, Hang, Chen, Yukang, Li, Yitong, Zhang, Guohui, Zhang, Yuan, Zhang, Xuying, Zhang, Tommy, Yuan, Jianlong, Li, Peihao, Lu, Shuai, Fu, Siming, Zhao, Chuyang, Han, Xin, Huang, Jie, Li, Wenbo, Ma, Guoqing, Huang, Wei, Qi, Xiaojuan, Huang, Haoyang, Duan, Nan
Abstract
Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration. Our method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation (SA-DMD), and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift. Extensive automatic and human evaluations show that JoyAI-Video-Edit substantially outperforms existing streaming editors and remains competitive with strong offline systems on both short and long videos. The complete system achieves end-to-end 720p video editing at approximately 30 FPS on a single Nvidia B200 GPU. Code is available at https://github.com/jd-opensource/JoyAI-Video-Edit.
Chinese Translation
实时视频编辑需要在有限的计算资源下实现低延迟的因果生成,同时保持源图像的保真度和长期的时间一致性。我们提出了JoyAI-视频编辑,这是一个具有160亿参数的自回归扩散框架,能够实现实时的开放式视频编辑,而无需访问未来帧或预定义的视频时长。我们的方法结合了块级自回归适应、源锚定分布匹配蒸馏(Source-Anchored Distribution Matching Distillation, SA-DMD)和长时间自回归蒸馏,以减少训练与推理的不匹配,在两步生成过程中保持源图像的保真度,并减轻累积的时间漂移。广泛的自动和人工评估表明,JoyAI-视频编辑在性能上显著优于现有的流媒体编辑器,并且在短视频和长视频上与强大的离线系统保持竞争力。完整系统在单个Nvidia B200 GPU上以约30帧每秒的速度实现720p视频编辑。代码可在https://github.com/jd-opensource/JoyAI-Video-Edit获取。
cs.CV / 120 / 2608.03979
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
视频深度研究:迈向下一代多模态深度研究代理
Abstract
We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval. Our framework adopts a two-stage training recipe: supervised fine-tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation-learning ceiling. Furthermore, we curate Video-DR-Bench, a human-AI collaborative benchmark comprising 200 complex, multi-hop VQA instances. Empirical results demonstrate that our Video-DeepResearch-35B-A3B establishes a new state-of-the-art of 64.0% average accuracy, surpassing proprietary Claude-4.5-Sonnet (59.0%) by 5.0 points and significantly outperforming GPT-5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B-A3B variant achieves 59.3%, competitive with Claude-4.5-Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: https://github.com/Osilly/Vision-DeepResearch.
Chinese Translation
我们介绍了视频深度研究(Video-DeepResearch,Video-DR),将多模态代理从静态图像扩展到连续的视频流,这一设置要求密集的时空基础与开放网络探索相结合。初步评估揭示了当前模型的两个关键瓶颈:(1)模态偏见,代理在视觉工具与文本搜索之间选择后者;(2)参数知识泄漏,模型依赖内部记忆而非真实的工具增强执行。为了解决这些挑战,我们提出了Video-DR,采用解耦的感知-探索管道,并通过阶段性工具解锁,强制在网络检索之前进行全面的跨帧视觉基础。我们的框架采用了两阶段训练方案:监督微调后接群体相对策略优化(Group Relative Policy Optimization,GRPO),使得自主探索突破模仿学习的瓶颈。此外,我们策划了Video-DR-Bench,这是一个包含200个复杂多跳视觉问答(VQA)实例的人机协作基准。实证结果表明,我们的Video-DeepResearch-35B-A3B建立了64.0%的平均准确率的新状态,超越了专有的Claude-4.5-Sonnet(59.0%)5.0个百分点,并显著优于GPT-5(52.5%)和Gemini 2.5 Pro(57.5%)。30B-A3B变体达到了59.3%的准确率,与Claude-4.5-Sonnet竞争,并展示了我们训练范式在紧凑规模下的有效性。代码链接:https://github.com/Osilly/Vision-DeepResearch。
cs.CV / 121 / 2608.03991
Perceptual Anchoring: Prototype-Guided Text Calibration for Training-free Open-Vocabulary Semantic Segmentation
感知锚定:基于原型的文本校准用于无训练的开放词汇语义分割
Abstract
Training-free open-vocabulary semantic segmentation (OVSS) partitions an image into semantically distinct regions based on arbitrary text descriptions, without learning any additional parameters. However, existing methods typically focus on improving visual representations while treating text embeddings that encode only generic category concepts as fixed classification references. The resulting semantic gap between these generic concepts and the visual representations that capture the specific appearances of target instances often causes incomplete masks and erroneous predictions in non-target regions. Inspired by the symbol-percept correspondence underlying perceptual anchoring, we propose Prototype-Guided Text Calibration (PTC) for training-free OVSS. In the Perceiving stage, PTC selects reliable visual evidence based on initial matching scores to construct category-specific visual prototypes. In the Anchoring stage, PTC uses these prototypes to calibrate their corresponding text embeddings, with the calibration strength adaptively adjusted based on the amount of visual evidence. Consequently, the calibrated text embeddings align more accurately with instance-specific visual representations while preserving generic category semantics and open-vocabulary generalization. Moreover, PTC requires neither additional training nor external models and can serve as a plug-and-play module for existing methods. Extensive experiments across eight benchmarks show that PTC significantly enhances the performance of six representative methods and yields more complete and accurate segmentation results. These results validate PTC as a simple and effective approach to improving visual-text alignment.
Chinese Translation
无训练的开放词汇语义分割(OVSS)基于任意文本描述将图像划分为语义上不同的区域,而无需学习任何额外的参数。然而,现有方法通常专注于改善视觉表示,同时将仅编码通用类别概念的文本嵌入视为固定的分类参考。这导致这些通用概念与捕捉目标实例特定外观的视觉表示之间的语义差距,常常造成不完整的掩膜和非目标区域的错误预测。受到感知锚定中符号-感知对应关系的启发,我们提出了无训练的OVSS的基于原型的文本校准(PTC)。在感知阶段,PTC根据初始匹配分数选择可靠的视觉证据,以构建类别特定的视觉原型。在锚定阶段,PTC利用这些原型来校准其对应的文本嵌入,校准强度根据视觉证据的数量自适应调整。因此,校准后的文本嵌入与实例特定的视觉表示更加准确对齐,同时保持通用类别语义和开放词汇泛化。此外,PTC不需要额外的训练或外部模型,可以作为现有方法的即插即用模块。在八个基准测试中的大量实验表明,PTC显著提升了六种代表性方法的性能,并产生了更完整和准确的分割结果。这些结果验证了PTC作为一种简单有效的提高视觉-文本对齐的方法。
cs.CV / 122 / 2608.04010
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs
ParVL:多模态大语言模型的并行扩展和可扩展计算分配
Abstract
Existing scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either model parameters or sequential inference computation, incurring substantial memory or latency overhead. More importantly, most existing methods fail to alter the rigid, fixed computation allocation between the Vision Transformer and the Large Language Model components, limiting task-specific optimization. To address this, we introduce the Parallel Vision-Language (ParVL) scaling framework for MLLMs, which scales parallel computation by reusing the existing ViT and LLM backbone parameters across multiple vision and language branches. This framework raises a central question: given a fixed backbone parameter budget, how should additional shared-backbone computation be allocated between the vision and language modalities? We instantiate each parallel computational stream with branch-specific prefix parameters over a shared backbone, and train the entire model end-to-end via full-parameter supervised fine-tuning on roughly 13B tokens. We systematically study the computation-allocation trade-off between the ViT encoder and LLM decoder. ParVL improves overall multimodal performance over same-recipe single-branch baselines, and the best evaluated vision--language allocation varies across tasks. Code is available at https://github.com/YangYangGirl/ParVL.
Chinese Translation
现有的多模态大语言模型(MLLMs)扩展策略通常扩展模型参数或顺序推理计算,导致显著的内存或延迟开销。更重要的是,大多数现有方法未能改变视觉变换器(Vision Transformer)与大语言模型(Large Language Model)组件之间刚性、固定的计算分配,限制了任务特定的优化。为了解决这个问题,我们提出了多模态大语言模型的并行视觉-语言(ParVL)扩展框架,该框架通过在多个视觉和语言分支之间重用现有的ViT和LLM主干参数来扩展并行计算。该框架提出了一个核心问题:在固定的主干参数预算下,如何在视觉和语言模态之间分配额外的共享主干计算?我们通过在共享主干上使用特定于分支的前缀参数实例化每个并行计算流,并通过对大约130亿个标记进行全参数监督微调来端到端训练整个模型。我们系统地研究了ViT编码器与LLM解码器之间的计算分配权衡。ParVL在相同配方的单分支基线之上提高了整体多模态性能,并且最佳评估的视觉-语言分配在不同任务之间有所不同。代码可在 https://github.com/YangYangGirl/ParVL 获取。
cs.AI / 1 / 2608.02604
ISEE: Interactive Semantic Enrichment for Database Fields
ISEE:数据库字段的交互式语义增强
Abstract
LLM-based agents are increasingly being deployed for data-related tasks, including data sense-making, exploration, and retrieval. However, their performance heavily depends on the clarity and completeness of data semantics. In practice, many field descriptions remain ambiguous or incomplete, as much of the essential context (e.g., the meaning of a customized field) originates from users' domain knowledge and is rarely documented publicly. This gap restricts the agents' task performance in downstream tasks, such as entity-linking. To bridge this gap, we introduce a novel and comprehensive Interactive SEmantic Enrichment system (ISEE). Given a data field description, ISEE measures its quality through a scoring system, gathers domain knowledge, and collaboratively enriches the semantics with users. Through a user study, automated user simulation, quantitative evaluation, and case study, we demonstrate that ISEE significantly reduces cognitive load, improves description quality, and enhances downstream task performance.
Chinese Translation
基于大语言模型(LLM)的智能体在数据相关任务中越来越多地被应用,包括数据理解、探索和检索。然而,它们的性能在很大程度上依赖于数据语义的清晰性和完整性。在实际应用中,许多字段描述仍然模糊或不完整,因为许多重要的上下文(例如,自定义字段的含义)源于用户的领域知识,且很少公开记录。这一差距限制了智能体在下游任务(如实体链接)中的任务表现。为了解决这一问题,我们提出了一种新颖且全面的交互式语义增强系统(ISEE)。ISEE在给定数据字段描述的基础上,通过评分系统评估其质量,收集领域知识,并与用户协作增强语义。通过用户研究、自动化用户模拟、定量评估和案例研究,我们证明ISEE显著降低了认知负担,提高了描述质量,并增强了下游任务的表现。
cs.AI / 2 / 2608.02606
Self-Organising Digital Circuits
自组织数字电路
Abstract
Fault tolerance in classical computing has traditionally relied on static strategies like hardware redundancy and error-correcting codes. Biological systems, in contrast, exhibit adaptive plasticity, maintaining function through dynamic re-organisation around damage. Inspired by this principle, we introduce Self-Organising Digital Circuits, framing functional logic generation and maintenance as a meta-learning problem on graphs. Our architecture employs a topology-masked Transformer that configures the Lookup Tables (LUT) of a circuit's Boolean gates. Extending the pattern-generation paradigm of Neural Cellular Automata (NCA), it navigates the degenerate Boolean search space to satisfy a computational task, rather than regenerating a fixed target state. We demonstrate that it can self-assemble functional circuits from scratch and rapidly re-route logic around permanent, previously unseen hardware faults. For soft errors, the policy achieves near-perfect recovery (>99.99\% accuracy) from damage sizes far exceeding training conditions. We further observe generalisation across circuit scales: accuracy improves on graphs substantially wider than those seen during training. This work bridges the principles of biological self-organisation with the practical domain of digital hardware.
Chinese Translation
传统的经典计算中的容错机制通常依赖于静态策略,如硬件冗余和错误纠正码。与此不同,生物系统表现出适应性可塑性,通过围绕损伤的动态重组来维持功能。受到这一原理的启发,我们提出了自组织数字电路(Self-Organising Digital Circuits),将功能逻辑的生成和维护框架视为图上的元学习问题。我们的架构采用了一种拓扑掩蔽的Transformer,配置电路布尔门的查找表(Lookup Tables, LUT)。扩展了神经元细胞自动机(Neural Cellular Automata, NCA)的模式生成范式,它在退化的布尔搜索空间中导航,以满足计算任务,而不是再生固定的目标状态。我们展示了它能够从零开始自组装功能电路,并迅速重新路由逻辑以绕过永久的、以前未见的硬件故障。对于软错误,该策略在损伤规模远超训练条件的情况下实现了近乎完美的恢复(>99.99%的准确率)。我们进一步观察到在电路规模上的泛化:在训练期间未见过的更宽图上,准确率显著提高。这项工作将生物自组织的原理与数字硬件的实际领域相结合。
cs.AI / 3 / 2608.02618
Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling
超越集体意识:通过元人格锚定和序列温度缩放逃离大型语言模型的同质化
Abstract
Recent studies have identified an ``Artificial Hivemind'' effect in Large Language Models (LLMs) causing models to converge on a narrow, homogenized consensus even for open questions. This semantic collapse limits the diversity of AI, resulting in high inter-response similarity ($\approx 0.80-0.90$) even under high-temperature sampling. In this paper, we propose a novel mitigation framework to increase diversity: Meta-Persona Anchoring combined with Filtered Temperature Scaling (FTS). Our approach utilizes a two-stage generation process: first, the model is prompted to self-select a unique, idiosyncratic persona to anchor its starting point; second, we apply a dual-stage sampling sieve, utilizing Top-$p$ filtering to preserve grammatical validity followed by extreme temperature scaling ($T \ge 4.0$) on the surviving candidates to explore the broadened probability distribution. We evaluate our method using the INFINITY-CHAT dataset on state-of-the-art open weight models under $\sim$20B parameters. Our results demonstrate a significant reduction in semantic convergence, with average pairwise cosine similarity dropping from ($\approx 0.85$) to ($\approx 0.65$). Our scheme achieves a majority of questions below the 0.7 threshold, effectively reducing the gap between artificial mode collapse and human-level typological diversity. We provide our implementation as an open-source framework to enable more diverse and creative AI deployments.
Chinese Translation
近期研究发现,大型语言模型(LLMs)中存在一种“人工集体意识”效应,导致模型即使在开放性问题上也趋向于狭窄的同质化共识。这种语义崩溃限制了人工智能的多样性,即使在高温采样下,响应之间的相似度仍然很高(约为0.80-0.90)。在本文中,我们提出了一种新颖的缓解框架,以增加多样性:元人格锚定结合过滤温度缩放(Filtered Temperature Scaling, FTS)。我们的方法利用了两阶段生成过程:首先,模型被提示自我选择一个独特的、特有的人格作为起始点;其次,我们应用双阶段采样筛选,利用Top-$p$过滤来保持语法有效性,然后对存活的候选者进行极端温度缩放($T
ge 4.0$),以探索扩展的概率分布。我们在约20B参数的最先进开放权重模型上使用INFINITY-CHAT数据集评估我们的方法。结果表明,语义收敛显著降低,平均成对余弦相似度从(约0.85)降至(约0.65)。我们的方案使大多数问题的相似度低于0.7,有效缩小了人工模式崩溃与人类水平类型多样性之间的差距。我们将我们的实现作为开源框架提供,以促进更具多样性和创造性的人工智能部署。
cs.AI / 4 / 2608.02630
PULSE: An Executable Contract Language for Spatiotemporal Knowledge Graph Engineering
PULSE:一种用于时空知识图谱工程的可执行合约语言
Abstract
Knowledge graph engineering often distributes accepted state, observations, constraints, processes, and hypothetical scenarios across artifacts whose combined execution contract remains external. We present PULSE, an Object-Process-Methodology-inspired language that localizes four operational roles and their write effects in one typed runtime. Here, modes denote operational roles rather than modal or deontic logic. The implemented contract fixes evidence non-overwrite, branch isolation, grounded multi-subject timers, guarded state change, and declaration-ranked event ordering over time and space; an external runner still decides whether evidence becomes an authoritative move. GeoSPARQL, SOSA, and SHACL remain generated views. A core calculus gives an effect-confinement lemma and six safety properties. Lean 4 checks kernel analogues for positions, evidence, clocks, monitors, atomicity, and branch source retention; 88 tests, 3,534 bounded checks, and 32 Lean/Python runtime-kernel cases bound the implementation claim to the checked cases. First-author implementations of a standards composition and a separate Sismic statechart reproduce the tested cold-chain trace. Across 37,440 generated temporal traces, PULSE matches a separate workflow and distinguishes ten single-field mutants. On the complete NOAA IBTrACS since1980 subset it agrees with GEOS and an event sweep on 1,476,290 transition-zone pairs, including 4,800 sampled and 12,831 duration-qualified events. Project-specific GeoSPARQL probes measure interface coverage. Overall, the results support contract localization, safety arguments, and trace parity for the tested fragment; language superiority and usability remain outside the evaluation.
Chinese Translation
知识图谱工程通常将接受的状态、观察、约束、过程和假设场景分散在多个工件中,其组合执行合约仍然是外部的。我们提出了PULSE,这是一种受对象-过程-方法论启发的语言,它在一个类型化的运行时中本地化了四种操作角色及其写入效果。在这里,模式表示操作角色,而不是模态或义务逻辑。实现的合约修复了证据非覆盖、分支隔离、基于多主体的定时器、受保护的状态变化以及随时间和空间的声明优先级事件排序;外部运行器仍然决定证据是否成为权威性动作。GeoSPARQL、SOSA和SHACL仍然是生成的视图。核心演算提供了一个效果限制引理和六个安全属性。Lean 4检查位置、证据、时钟、监视器、原子性和分支源保留的内核类比;88个测试、3,534个有界检查和32个Lean/Python运行时内核案例将实现声明限制在已检查的案例中。第一作者实现的标准组合和单独的Sismic状态图重现了测试的冷链追踪。在37,440个生成的时间轨迹中,PULSE与一个独立的工作流相匹配,并区分了十个单字段突变体。在自1980年以来的完整NOAA IBTrACS子集中,它与GEOS和1,476,290个过渡区对上的事件扫描一致,包括4,800个采样事件和12,831个持续时间合格事件。项目特定的GeoSPARQL探测器测量接口覆盖率。总体而言,结果支持合约本地化、安全性论证和测试片段的轨迹平等;语言的优越性和可用性仍然不在评估范围内。
cs.AI / 5 / 2608.02650
HyperAgent: Planning and Acting over Tool-Schema Hypergraphs for Tool-Use LLM Agents
HyperAgent:基于工具模式超图的工具使用 LLM 代理的规划与执行
Abstract
Large language model (LLM) agents increasingly rely on external tools to complete complex real-world tasks. However, reliable tool-use planning remains challenging due to the limitations of implicit reasoning and the evolving nature of real-world execution environments. Existing tool-use agents typically rely on LLMs to infer tool compositions from textual descriptions, which can lead to inefficient exploration and unreliable execution in complex tasks. To address these challenges, we model tool relations at the schema level and construct a directed Tool--Schema Hypergraph, in which tools are represented as hyperedges from their required input-schema nodes to their output-schema nodes. Furthermore, we propose HyperAgent, a Tool--Schema Hypergraph-guided framework for dynamic planning and execution. Given a task, HyperAgent first extracts a task-relevant tool context graph and uses it to guide the construction of a schema-aware Task DAG. During execution, HyperAgent dynamically realizes each subtask by constructing a state-conditioned tool support graph through deficit-oriented expansion, which identifies unresolved requirements and retrieves supporting producer tools according to the current agent state. Experiments on AppWorld demonstrate that HyperAgent improves task completion performance while reducing redundant API calls, LLM interactions, and token consumption compared with existing agent baselines.
Chinese Translation
大型语言模型(LLM)代理越来越依赖外部工具来完成复杂的现实世界任务。然而,由于隐式推理的局限性和现实执行环境的不断变化,可靠的工具使用规划仍然具有挑战性。现有的工具使用代理通常依赖 LLM 从文本描述中推断工具组合,这可能导致在复杂任务中的低效探索和不可靠执行。为了解决这些挑战,我们在模式层面上建模工具关系,并构建一个有向工具-模式超图,其中工具被表示为从其所需输入模式节点到输出模式节点的超边。此外,我们提出了 HyperAgent,一个基于工具-模式超图的动态规划与执行框架。给定一个任务,HyperAgent 首先提取与任务相关的工具上下文图,并利用该图指导构建一个具有模式感知的任务有向无环图(Task DAG)。在执行过程中,HyperAgent 通过缺口导向扩展动态实现每个子任务,构建一个状态条件的工具支持图,识别未解决的需求,并根据当前代理状态检索支持的生产工具。在 AppWorld 上的实验表明,与现有代理基线相比,HyperAgent 提高了任务完成性能,同时减少了冗余的 API 调用、LLM 交互和令牌消耗。
cs.AI / 6 / 2608.02699
Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation Gap
可解释人工智能与欧盟解释权:法律与可解释人工智能翻译差距的系统评估
Abstract
When algorithms make or influence consequential decisions---about loan eligibility, hiring, or healthcare---EU law grants affected individuals a Right to Explanation. Yet whether (and how) Explainable AI (XAI) can satisfy this right in practice remains poorly understood, with direct implications for individuals' ability to contest automated decisions that affect their lives. This paper presents a systematic literature review of XAI in the context of the EU Right to Explanation, with particular focus on Art. 15(1)(h) GDPR, Art. 86 AI Act (AIA), and related instruments. We consider papers published from 2024 onwards, as the final version of the AIA was published in July 2024---with Art. 86 being added late. From 2643 initial records identified by a deliberately broad search, we review 57 full texts, of which only 19 papers demonstrate substantive integration of both legal and technical perspectives, showing gaps in the interdisciplinary synthesis of the current regulatory framework. We document three problematic patterns across the corpus: Most misidentify the GDPR legal basis; few engage with the CJEU's Dun & Bradstreet judgment (likely due to publication timing); and the distinction between explanation form (governed by addressee) and content (governed by legal purpose) is often conflated. We conceptualize this as the Addressee/Purpose Framework, propose a four-phase blueprint for operationalization, and identify six concrete open research questions. Without further progress, the Right to Explanation risks remaining a formal obligation without a technically realizable path to compliance.
Chinese Translation
当算法做出或影响重要决策——如贷款资格、招聘或医疗保健——时,欧盟法律赋予受影响的个人解释权。然而,如何在实践中通过可解释人工智能(XAI)满足这一权利仍然不甚明确,这直接影响到个人对影响其生活的自动化决策提出异议的能力。本文对XAI在欧盟解释权背景下的文献进行了系统评审,特别关注《通用数据保护条例》第15条第1款第(h)项和《人工智能法案》第86条及相关文献。我们考虑了2024年及以后发表的论文,因为《人工智能法案》的最终版本于2024年7月发布,而第86条的加入较晚。在2643个通过广泛搜索初步识别的记录中,我们审查了57篇完整文本,其中仅有19篇论文在法律与技术视角的整合上表现出实质性,显示出当前监管框架在跨学科综合方面的不足。我们记录了该文献中的三个问题模式:大多数错误识别了GDPR的法律基础;很少有论文涉及欧洲法院的Dun & Bradstreet判决(可能由于出版时机);解释形式(由受文者决定)与内容(由法律目的决定)之间的区别常常被混淆。我们将其概念化为受文者/目的框架,提出了一个四阶段的实施蓝图,并确定了六个具体的开放研究问题。如果没有进一步的进展,解释权可能仍然是一个形式上的义务,而没有技术上可实现的合规路径。
cs.AI / 7 / 2608.02704
Predictive Set Theory: A Generative Framework for Cognitive Architecture with Operationalized Core Mechanisms
预测集合理论:具有操作化核心机制的认知架构生成框架
Abstract
Predictive processing theories portray the brain as a hierarchical prediction engine that minimizes prediction error, yet they lack operational definitions for the structure of a "prediction," the standardized response to a prediction error, and the mechanism that maintains consistency across successive updates. Bayesian cognitive science attempts to subsume all uncertainty under probabilistic belief updating, but it presupposes a closed hypothesis space and provides no generative account of how the objects over which probabilities are distributed become discrete, identifiable referents in the first place. This paper introduces Predictive Set Theory (PST), a formal generative framework that reconstructs cognitive architecture from first principles. PST anchors cognition in a minimal set of operations---a sensor formalized as an identity function, set-theoretic state refresh, and three fundamental forms of reference chains (reference, counter-reference, and semi-reference)---and rigorously derives core cognitive functions including state sequences, demand, comparison, efficiency, and finite-horizon probabilistic planning. Rather than modeling neural mechanisms, PST constitutes a design specification for any system that must maintain internal consistency while acting under incomplete information and irreversible risk. The framework offers novel resolutions to classical problems such as Russell's paradox, the cognitive status of G\"{o}delian incompleteness, the grounding of negative feedback, and the comprehension of film editing. The primary purpose of this paper is to establish, through the public academic record, the originality and completeness of the Predictive Set Theory framework.
Chinese Translation
预测处理理论将大脑描绘为一个层级预测引擎,旨在最小化预测误差,然而它们缺乏对“预测”结构的操作性定义、对预测误差的标准化响应以及在连续更新中保持一致性的机制。贝叶斯认知科学试图将所有不确定性纳入概率信念更新之下,但它假设了一个封闭的假设空间,并未提供关于概率分布对象如何首先成为离散、可识别的指称的生成性解释。本文介绍了预测集合理论(Predictive Set Theory, PST),这是一个从基本原理重建认知架构的正式生成框架。PST 将认知锚定在一组最小操作上——一个形式化为恒等函数的传感器、集合论状态刷新以及三种基本的参考链形式(参考、反参考和半参考)——并严格推导出包括状态序列、需求、比较、效率和有限视野概率规划在内的核心认知功能。PST 并非建模神经机制,而是为任何必须在不完整信息和不可逆风险下维持内部一致性的系统提供设计规范。该框架为经典问题提供了新颖的解决方案,例如拉塞尔悖论、哥德尔不完备性的认知状态、负反馈的基础以及对电影剪辑的理解。本文的主要目的是通过公共学术记录,确立预测集合理论框架的原创性和完整性。
cs.AI / 8 / 2608.02775
Towards a new paradigm of scientific discovery with socialized artificial intelligence
迈向社会化人工智能的新科学发现范式
Yao, Xinjie, Xu, Xingxin, Gao, Xiyuan, Guo, Zhoupeng, Yang, Kunlong, Zhao, Dengyu, Zhao, Siqi, Fan, Zhihe, Dong, Yichen, Li, Xin, Feng, Jiekang, Wu, Jiahe, Wang, Sen, Yu, Beiming, Zhao, Kejia, Zhao, Ruipu, Zhou, Jiaqi, Li, Heyang, Chen, Jianjun, Dai, Anbo, Liu, Xin, Yu, Zhengtao, Hu, Qinghua, Zhu, Pengfei
Abstract
Scientific discovery has advanced through successive transformations in the organization of knowledge. Observation and experimentation established the empirical foundations of science. Theory made it possible to derive general principles from particular phenomena. Computation extended inquiry into systems beyond direct observation, while data-intensive methods opened new spaces of pattern and prediction. Science now confronts a different frontier. The central challenge is no longer simply to produce more information, but to organize expanding knowledge, reasoning, and evidence into a coherent process of discovery. Here, we introduce Bridging Literature, Agents, and Zero-gap Experimentation (BLAZE), a paradigm of socialized scientific intelligence. BLAZE conceives AI not as an assistant for isolated research tasks, but as an organizational infrastructure for scientific discovery. It connects persistent knowledge, collective reasoning, empirical validation, and human judgment within a continuous research lifecycle, transforming fragmented activities into a cumulative process of inquiry, criticism, and revision. The central premise of BLAZE is that scientific intelligence does not arise from computation alone. It emerges from the sustained interaction among knowledge, hypotheses, experiments, and collective verification. By organizing humans and machines within a shared scientific process, BLAZE makes discovery more traceable, reproducible, and cumulative while preserving human creativity, judgment, and responsibility. Socialized scientific intelligence may provide a foundation for the next era of science. Its purpose is not to replace human discovery, but to extend the scale, depth, and continuity of collective scientific inquiry.
Chinese Translation
科学发现通过知识组织的连续转变而不断进步。观察和实验奠定了科学的经验基础。理论使得从特定现象中推导出一般原则成为可能。计算扩展了对超出直接观察的系统的探究,而数据密集型方法则开辟了模式和预测的新空间。科学现在面临着一个不同的前沿。中心挑战不再仅仅是产生更多的信息,而是将不断扩展的知识、推理和证据组织成一个连贯的发现过程。在此,我们介绍了文献、代理和零差实验的桥接(Bridging Literature, Agents, and Zero-gap Experimentation,简称BLAZE),这是一个社会化科学智能的范式。BLAZE将人工智能视为孤立研究任务的助手,而是作为科学发现的组织基础设施。它在一个持续的研究生命周期内连接持久的知识、集体推理、经验验证和人类判断,将碎片化的活动转变为一个累积的探究、批评和修订的过程。BLAZE的核心前提是,科学智能并非仅仅源于计算,而是源于知识、假设、实验和集体验证之间的持续互动。通过在共享的科学过程中组织人类和机器,BLAZE使得发现变得更加可追溯、可重复和累积,同时保留人类的创造力、判断力和责任感。社会化科学智能可能为下一个科学时代奠定基础。它的目的不是取代人类的发现,而是扩展集体科学探究的规模、深度和连续性。
cs.AI / 9 / 2608.02876
BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL
BAP-SQL:面向预算的代理文本到SQL观察规划
Abstract
Tool-using agents do not merely consume observations: their actions determine what arrives next. In agentic text-to-SQL, a broad query can spend context and database work before useful evidence appears, while post-hoc compression cannot recover omitted rows or expended work. We present BAP-SQL, which treats observation formation as a budget-control stage: it estimates query risk, rewrites SQL when useful, and delegates hard limits to an independent runtime shield. Across general 4B, specialized FINER-SQL 4B, and 7B backbones, BAP-SQL improves tight-budget success. On the primary BIRD-derived setting, it gains 3.4/3.6 percentage points over matched SFT while using 4.5/5.0% fewer tokens. Matched retraining and task-level transfer associate the gain with policy-visible planning and budget-sensitive rescue. The benefit attenuates as model capability and budget increase, reverses at the loosest setting, and does not reduce database work.
Chinese Translation
使用工具的代理不仅仅是消耗观察:它们的行动决定了接下来会出现什么。在代理文本到SQL中,广泛的查询可能在有用证据出现之前消耗上下文和数据库工作,而事后压缩无法恢复被省略的行或已消耗的工作。我们提出了BAP-SQL,它将观察形成视为一个预算控制阶段:它估算查询风险,在有用时重写SQL,并将硬限制委托给独立的运行时保护层。在通用的4B、专用的FINER-SQL 4B和7B基础模型上,BAP-SQL提高了紧预算成功率。在主要的基于BIRD的设置中,它比匹配的SFT提高了3.4/3.6个百分点,同时使用了4.5/5.0%的更少token。匹配的再训练和任务级转移将这一增益与可见的策略规划和预算敏感的救助联系起来。随着模型能力和预算的增加,收益减弱,在最宽松的设置下反转,并且不会减少数据库工作。
cs.AI / 10 / 2608.02878
VeriTrace: Human-Like Temporal Exploration Completes Agentic Action Space
VeriTrace:类人时间探索完成代理行动空间
Abstract
Large language models have shown promise for automated Verilog RTL generation, yet state-of-the-art multi-agent systems plateau at ~95% accuracy on standard benchmarks. We trace this ceiling to an incomplete debugging action space: existing systems restrict which signals the agent can inspect, which time windows it can query, or both, reducing debugging to pattern matching on a narrow, predetermined view of circuit behavior rather than hypothesis-driven root-cause analysis. We present VeriTrace, a multi-agent system whose Inspector agent operates over a complete debugging action space, with independent control over signal selection, time-window bounds, and iteration depth. This capability, which we term Agentic Temporal Exploration, enables the agent to form hypotheses about failure causes, query the waveform for evidence, and refine its understanding iteratively, mirroring the exploratory process of human verification engineers. VeriTrace achieves 100\% Pass@1 on VerilogEval-V2, the first system to attain perfect functional correctness on this benchmark. On a shared Claude Sonnet 4.0 backbone, VeriTrace outperforms the strongest reproduced baseline by +5.1%, demonstrating that debugging agency closes the final accuracy gap.
Chinese Translation
大型语言模型在自动化 Verilog RTL 生成方面展现了潜力,但最先进的多智能体系统在标准基准测试中的准确率停滞在约 95%。我们将这一瓶颈归因于不完整的调试行动空间:现有系统限制了智能体可以检查的信号、可以查询的时间窗口,或两者兼而有之,从而将调试简化为在狭窄的、预设的电路行为视图上进行模式匹配,而非基于假设的根本原因分析。我们提出了 VeriTrace,一个多智能体系统,其检查智能体在完整的调试行动空间中操作,能够独立控制信号选择、时间窗口界限和迭代深度。这种能力,我们称之为代理时间探索(Agentic Temporal Exploration),使得智能体能够形成关于故障原因的假设,查询波形以获取证据,并迭代地完善其理解,模拟人类验证工程师的探索过程。VeriTrace 在 VerilogEval-V2 上实现了 100% 的 Pass@1,成为第一个在该基准上达到完美功能正确性的系统。在共享的 Claude Sonnet 4.0 基础上,VeriTrace 的表现超越了最强的再现基线,提升了 5.1%,证明了调试能力缩小了最终的准确性差距。
cs.AI / 11 / 2608.02879
Interpreting Black-Box Large Language Models with Sentence-Level Energy Landscapes
通过句子级能量景观解释黑箱大型语言模型
Abstract
The widespread adoption of proprietary Large Language Models (LLMs) accessed strictly through closed APIs has created a critical challenge for responsible deployment: a fundamental lack of interpretability. To address this, we propose a model-agnostic, post-hoc attribution interpreter operating at the sentence level. Our approach trains an Energy-Based Model (EBM) as a surrogate to capture the LLM's internal conceptual consistency between prompts and responses. This energy landscape guides the training of a lightweight interpreter network. Uniquely, our interpreter operates as a standalone tool; once trained, it quantifies the influence of prompt sentences on a user-specified target output without requiring further API queries to the LLM. By globally training a local interpreter across diverse inputs, our framework captures broader generation patterns and mitigates instance-specific biases. Experiments demonstrate that our EBM accurately simulates the target LLM, allowing the interpreter to effectively identify the prompt sentences most influential in generating specific target outputs.
Chinese Translation
专有大型语言模型(LLMs)的广泛采用,严格通过封闭的API访问,给负责任的部署带来了一个关键挑战:缺乏可解释性。为了解决这一问题,我们提出了一种模型无关的后置归因解释器,该解释器在句子级别上操作。我们的方法训练了一个能量基础模型(Energy-Based Model, EBM)作为替代,以捕捉LLM在提示和响应之间的内部概念一致性。这个能量景观指导了轻量级解释器网络的训练。独特之处在于,我们的解释器作为一个独立工具运行;一旦训练完成,它能够量化提示句子对用户指定目标输出的影响,而无需进一步查询LLM的API。通过在多样化输入上全局训练局部解释器,我们的框架捕捉了更广泛的生成模式,并减轻了特定实例的偏差。实验表明,我们的EBM准确模拟了目标LLM,使得解释器能够有效识别在生成特定目标输出中影响最大的提示句子。
cs.AI / 12 / 2608.02930
Hypercubes, Hyperplanes, and Constraint-Induced Complexity Collapse in Atomic Concept Learning
超立方体、超平面与约束诱导的原子概念学习中的复杂性崩溃
Abstract
We revisit higher-arity atomic concept learning through the geometry of hypercubes and hyperplanes of ground instances. Our starting point is the observation that the ambient r-dimensional hypercube of ground atoms is not structurally uniform. Its logical complexity is organized by hyperplanes: every hyperplane other than the full diagonal collapses into finitely many elementary-equivalence classes, with a bound independent of the term depth, while the full diagonal is exceptional and its class count grows without bound. This asymmetry is not merely geometric. It reflects the reduction-theoretic structure of the concepts themselves. Building on a higher-dimensional framework developed in the author's earlier work, we reinterpret these results through canonical simple concepts, minimal orderings, and representative reductions. This yields a taxonomy of hyperplane behavior in higher dimensions and shows that complexity is localized rather than spread uniformly through the instance space. The paper includes a fully worked binary case, an explicit treatment of the ternary hypercube, and an unpacked account of the reduction machinery that drives the collapse. The three-dimensional case already exhibits the essential phenomenon of orthogonal families, partial diagonals, and the exceptional full diagonal. This geometric-logical perspective clarifies where complexity is concentrated in atomic concept learning and suggests a modern interpretation in terms of constrained hypothesis spaces and structured classification.
Chinese Translation
我们通过超立方体和基础实例的超平面的几何结构重新审视高阶原子概念学习。我们的出发点是观察到基础原子所处的r维超立方体在结构上并不均匀。其逻辑复杂性由超平面组织:除了完整对角线之外的每个超平面都收敛为有限多个基本等价类,且其界限与术语深度无关,而完整对角线则是例外,其类计数无限增长。这种不对称不仅仅是几何上的,它反映了概念本身的约简理论结构。在作者早期工作的基础上,我们在更高维度的框架中重新解释这些结果,使用典型简单概念、最小排序和代表性约简。这产生了高维中超平面行为的分类法,并显示复杂性是局部化的,而不是均匀分布在实例空间中。本文包括一个完整的二元案例、对三元超立方体的明确处理,以及驱动崩溃的约简机制的详细说明。三维案例已经展示了正交家族、部分对角线和特殊完整对角线的基本现象。这种几何-逻辑视角阐明了在原子概念学习中复杂性集中在哪里,并提出了关于约束假设空间和结构化分类的现代解释。
cs.AI / 13 / 2608.02940
When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning
当压缩评分无法决策时:群体鲁棒大语言模型剪枝的信息边界
Abstract
A reproducible compression statistic can still select the wrong candidate. A dense pruning score with 0.906 split-half reliability predicted a 16.1% gain. Its selected endpoint was 6.0% and 7.7% worse than two controls. We model the gap through information interfaces that delimit which distinctions each statistic supports. For equal-weight groups, a conic law gives the exact pooling price for positive linear fixed-candidate damage, including diagonal and full PSD second moments. Three two-world constructions and an exact observation-fiber radius characterize what pooled moments, group-local moments, and reference-path curvature leave unresolved. A group-resolved diagonal recovers broad damage order (Spearman 0.9239) while fine order remains weak. Relative to balanced uniform allocation, a coarse depth allocation cuts worst-group perplexity inflation by 12.6--20.9% across three dense LLMs. Model-specific complete-mask endpoint selection improves over those references by 2.7--8.0%. In OLMoE, router traces predict singleton direction (114/192 versus 81/192 under the strongest relabeling). Finite-menu decisions on one layer yield held-out worst-group KL reductions of 13.7% and 7.2%. Local measurements construct candidates. Selection is licensed by complete candidate endpoints or a validated uniform guarantee, with uncertainty calibrated to every comparison.
Chinese Translation
可重复的压缩统计量仍然可能选择错误的候选者。一个具有0.906分半可靠性的密集剪枝评分预测了16.1%的增益。其选择的终点比两个对照组差6.0%和7.7%。我们通过信息接口建模这一差距,这些接口界定了每个统计量所支持的区分。对于等权重组,锥形法则给出了正线性固定候选损害的确切汇聚价格,包括对角线和完整的PSD二阶矩。三个双世界构造和一个确切的观察纤维半径表征了汇聚矩、组局部矩和参考路径曲率所留下的未解决问题。群体解析的对角线恢复了广泛的损害顺序(Spearman 0.9239),而细微顺序仍然较弱。相对于均衡的均匀分配,粗略的深度分配在三个密集大语言模型中将最差组的困惑度膨胀降低了12.6%至20.9%。模型特定的完整掩码终点选择在这些参考值上提高了2.7%至8.0%。在OLMoE中,路由器轨迹预测单例方向(114/192对比81/192在最强重标记下)。在一层上的有限菜单决策产生了保持的最差组KL减少13.7%和7.2%。局部测量构建候选者。选择是通过完整候选终点或经过验证的均匀保证获得许可的,且不确定性经过每次比较的校准。
cs.AI / 14 / 2608.02949
On the missing data layer and a potential solution
关于缺失数据层及其潜在解决方案
Abstract
Latin America is missing two foundational layers of AI infrastructure: the dataset layer and the benchmark layer. This paper targets the dataset layer. The dataset layer faces two compounding problems: discovery and supply. Latin American AI datasets exist but are scattered across platforms with no shared index. Even with perfect indexing, the total volume would remain far below what frontier AI development requires. We propose DataHub: a task-first data infrastructure organized through the ontology ///, with mechanisms for dataset discovery, metadata, contribution, licensing, and reuse.
Chinese Translation
拉丁美洲缺乏两层基础性的人工智能基础设施:数据集层和基准层。本文针对数据集层进行探讨。数据集层面临两个相互叠加的问题:发现和供应。拉丁美洲的人工智能数据集确实存在,但分散在各个平台上,缺乏共享索引。即使有完美的索引,总量仍远低于前沿人工智能发展所需的水平。我们提出了DataHub:一个以任务为中心的数据基础设施,通过本体论///进行组织,并配备数据集发现、元数据、贡献、许可和重用的机制。
cs.AI / 15 / 2608.02993
Neurosymbolic Reasoning with Incremental Knowledge for Sample Efficient Hierarchical Reinforcement Learning
基于增量知识的神经符号推理用于样本高效的层次强化学习
Abstract
(Flat) Reinforcement Learning (RL) agents face significant challenges in environments with sparse rewards that require long-horizon reasoning. A compelling approach to improve sample efficiency is to incorporate knowledge into learning and decision-making. In standard Hierarchical RL (HRL), knowledge is encoded in a fixed, non-updatable form, such as architectural choices, and remains unchanged throughout learning. With fixed HRL, reasoning with incremental knowledge learned during exploration is impractical before sufficient environmental knowledge is acquired, leading to poor sample efficiency. In this work, we propose neurosymbolic HRL with {\em Incremental Knowledge (InK)}: symbolic high-level components perform {\em symbolic planning} (e.g. using $D^*$) on an updatable representation of current InK, while low-level goal-conditioned neural modules learn motion primitives through experience using reward shaping. Experiments on navigation tasks demonstrate that incorporating InK substantially improves sample efficiency. Additionally, to perform {\em optimal} symbolic planning given {\em prior} knowledge about the world, we develop Belief World Tree Search. The code is available at https://github.com/CPS-research-group/ink_bwts.
Chinese Translation
(平面) 强化学习 (RL) 代理在面对稀疏奖励的环境中,尤其是需要长时间推理的情况下,面临重大挑战。提高样本效率的一个有效方法是将知识融入学习和决策中。在标准的层次强化学习 (HRL) 中,知识以固定的、不可更新的形式编码,例如架构选择,并在整个学习过程中保持不变。在固定的 HRL 中,在获得足够的环境知识之前,利用在探索过程中学习到的增量知识进行推理是不切实际的,这导致样本效率低下。在本研究中,我们提出了带有增量知识 (Incremental Knowledge, InK) 的神经符号 HRL:符号高层组件在当前可更新的 InK 表示上执行符号规划(例如,使用 $D^*$),而低层目标条件神经模块通过经验学习运动原语,利用奖励塑形。导航任务的实验表明,融入 InK 显著提高了样本效率。此外,为了在给定世界的先验知识的情况下执行最优的符号规划,我们开发了信念世界树搜索 (Belief World Tree Search)。代码可在 https://github.com/CPS-research-group/ink_bwts 获取。
cs.AI / 16 / 2608.02996
On the missing benchmarks layer and a potential solution
缺失的基准层及其潜在解决方案
Abstract
Latin America is missing a foundational layer for native AI development: the benchmark layer. The benchmark layer does two things no other layer can - it audits AI systems against regional social requirements and it directs AI optimization in economically relevant environments. Without it, public institutions cannot independently evaluate foreign AI systems, and companies cannot optimize AI systems to solve local problems with SOTA performance. The cost of the missing layer is dual: a loss of auditability and a loss of optimization direction over a technology that is increasingly critical infrastructure. We propose an EvalsHub, with LatamBoard as its first regional instance - an open, task-first benchmark infrastructure where universities, public institutions, professional communities, and companies can publish, execute, compare, and maintain evaluations across models, workflows, and agents. Built once, measured forever - re-run by institutions as new AI systems ship and by industry teams after every system change. Open by design and incentive-driven by construction.
Chinese Translation
拉丁美洲缺乏一个用于本土人工智能发展的基础层:基准层。基准层具有其他层无法实现的两个功能——它对人工智能系统进行审计,以满足区域社会需求,并指导人工智能在经济相关环境中的优化。没有基准层,公共机构无法独立评估外国人工智能系统,企业也无法优化人工智能系统以解决本地问题并达到最先进的性能。缺失这一层的代价是双重的:审计能力的丧失和对日益重要的基础设施技术的优化方向的丧失。我们提出了EvalsHub,以LatamBoard作为其第一个区域实例——一个开放的、以任务为导向的基准基础设施,大学、公共机构、专业社区和企业可以在此发布、执行、比较和维护跨模型、工作流和代理的评估。一次构建,永久测量——由机构在新人工智能系统发布时重新运行,并由行业团队在每次系统变更后进行测试。设计上开放,构建上以激励驱动。
cs.AI / 17 / 2608.03006
ProPRL: Property-Aware Prerequisite Relation Learning in Educational Knowledge Graphs
ProPRL:教育知识图谱中的属性感知前提关系学习
Abstract
Prerequisite relation learning is central to adaptive instruction, yet existing methods often formulate it as conventional link prediction, limiting their ability to adaptively integrate complementary educational evidence for individual candidate pairs and to discourage contradictory reverse predictions. We propose ProPRL, a Property-aware Prerequisite Relation Learning framework. ProPRL first learns complementary concept representations from a concept-resource hypergraph and a directed learning-behavior graph, where direction-preserving personalized propagation aggregates multi-hop behavioral evidence. It then employs a Pair-conditioned Gate to adaptively weight and fuse the two views for each candidate ordered concept pair. Finally, an \textit{Irreversibility Constraint} introduces an anti-symmetry regularizer that penalizes simultaneously high confidence in both directions of the same concept pair. Experiments on multiple real-world educational datasets show that ProPRL achieves state-of-the-art performance on prerequisite relation learning.
Chinese Translation
前提关系学习是自适应教学的核心,但现有方法通常将其表述为传统的链接预测,这限制了它们自适应整合个体候选对的互补教育证据的能力,并且无法有效抑制矛盾的反向预测。我们提出了ProPRL,一个属性感知前提关系学习框架。ProPRL首先从概念-资源超图和有向学习行为图中学习互补的概念表示,其中保持方向性的个性化传播聚合了多跳行为证据。然后,它采用配对条件门(Pair-conditioned Gate)为每个候选有序概念对自适应地加权和融合这两个视图。最后,一个不可逆约束(Irreversibility Constraint)引入了一个反对称正则化器,惩罚在同一概念对的两个方向上同时具有高置信度的情况。在多个真实世界教育数据集上的实验表明,ProPRL在前提关系学习方面达到了最先进的性能。
cs.AI / 18 / 2608.03018
UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks
UrbanAgent:一种用于跨系统城市任务的工具增强代理
Abstract
Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragmented and have little interoperability, placing a heavy operational burden on users. Existing digital platforms, urban foundation models, and intelligent assistants each address only isolated aspects of an urban task. But they struggle to reliably convert complex natural-language requests into executable cross-system workflows. We propose Urban-Agent, a tool-augmented agent framework for cross-system urban tasks. It couples the cognitive and reasoning capabilities of a large language model with a tool-set supporting code execution, API calls, and Model Context Protocol. Through one adaptive closed loop, it clarifies missing information before acting, grounds tool use in live observations, and aligns the final response with observed evidence and task constraints. To address the evaluation gap, we introduce Urban-Eval, a benchmark specifically designed for cross-system urban request. Unlike prior benchmarks that assess either general tool use or urban knowledge and reasoning, Urban-Eval evaluates both task results and execution quality, including required tool coverage, dependency validity, and evidence traceability. Experimental results indicate that Urban-Agent reaches a 71% task success rate, 10 points above the strongest baseline. This lead holds across GPT-5-mini, Gemini-2.5-flash, DeepSeek-V4-flash, and Qwen3-235B-A22B.
Chinese Translation
现代城市依赖越来越多的数字服务来运作,但居民的日常需求仍然难以满足。服务碎片化且互操作性差,给用户带来了沉重的操作负担。现有的数字平台、城市基础模型和智能助手各自仅解决城市任务的孤立方面,但在将复杂的自然语言请求可靠地转化为可执行的跨系统工作流方面存在困难。我们提出了Urban-Agent,一种用于跨系统城市任务的工具增强代理框架。它将大型语言模型的认知和推理能力与支持代码执行、API调用和模型上下文协议的工具集相结合。通过一个自适应闭环,它在行动之前澄清缺失的信息,将工具使用与实时观察相结合,并使最终响应与观察到的证据和任务约束保持一致。为了解决评估差距,我们引入了Urban-Eval,这是一个专门为跨系统城市请求设计的基准测试。与以往评估一般工具使用或城市知识和推理的基准不同,Urban-Eval同时评估任务结果和执行质量,包括所需工具覆盖、依赖有效性和证据可追溯性。实验结果表明,Urban-Agent的任务成功率达到71%,比最强基准高出10个百分点。这一优势在GPT-5-mini、Gemini-2.5-flash、DeepSeek-V4-flash和Qwen3-235B-A22B中均保持一致。
cs.AI / 19 / 2608.03020
LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
LoCA:在一次性校准后进行前向仅 LLM 调优的局部信用分配
Abstract
Parameter-efficient post-training reduces the number of trainable parameters, but still requires repeated end-to-end backpropagation through the frozen backbone. Every adaptation step therefore needs backward-capable hardware and must store or recompute activations. We ask whether this repeated backward chain can be replaced by a one-time calibration. We introduce Local Credit Assignment (LoCA), a two-stage method for small-shift adaptation. One probe backward pass fits a low-rank map at each transformer block from the final prediction error to a local hidden-state correction. LoCA then reuses these maps to form blockwise regression targets from forward activations and fits low-rank adapters with closed-form ridge solves. No further backbone backward pass is required. We evaluate LoCA on five discriminative benchmarks with Qwen2.5 models from 0.5B to 14B. In 16 of 25 reported task--scale comparisons, LoCA yields lower evaluation cross-entropy than the corresponding LoRA run. Its measured full-run GPU peak, including calibration, is 26--29\% lower than LoRA's. After calibration, its CPU steady-state memory is 36--52\% lower and its per-pass time is 43--48\% lower. A shared scale-normalized candidate set is reused across all tested Qwen2.5 sizes and on SmolLM2-1.7B. LoCA thus amortizes global credit assignment into one calibration and enables later forward-only tuning when repeated backpropagation is impractical. The code associated with this paper is available \href{https://github.com/Xia12121/LoCA}{here}.
Chinese Translation
参数高效的后训练减少了可训练参数的数量,但仍然需要通过冻结的主干网络进行重复的端到端反向传播。因此,每个适应步骤都需要具备反向传播能力的硬件,并且必须存储或重新计算激活值。我们探讨是否可以用一次性校准来替代这种重复的反向链。我们提出了局部信用分配(Local Credit Assignment, LoCA),这是一种用于小幅度适应的两阶段方法。一次探测反向传播在每个变换器块中拟合一个低秩映射,将最终预测误差与局部隐藏状态修正相联系。LoCA 然后重用这些映射,从前向激活中形成块级回归目标,并通过封闭形式的岭回归拟合低秩适配器。无需进一步的主干反向传播。我们在五个判别基准上评估了 LoCA,使用的 Qwen2.5 模型规模从 0.5B 到 14B。在 25 个报告的任务-规模比较中,有 16 个任务显示 LoCA 的评估交叉熵低于相应的 LoRA 运行。其测得的完整运行 GPU 峰值,包括校准,低于 LoRA 的 26-29%。校准后,其 CPU 稳态内存低于 36-52%,每次传递的时间低于 43-48%。一个共享的规模归一化候选集在所有测试的 Qwen2.5 尺寸和 SmolLM2-1.7B 上重复使用。因此,LoCA 将全局信用分配摊销为一次校准,并在重复反向传播不切实际时启用后续的前向仅调优。与本文相关的代码可在此处获得: exttt{https://github.com/Xia12121/LoCA}。
cs.AI / 20 / 2608.03025
DiffImaginE: Imagine to Verify Entity Types with Diffusio
DiffImaginE:通过扩散想象来验证实体类型
Abstract
Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type) pair to one predicted visual feature, compressing diverse visual realisations into a single prototype and providing a compatibility score without explicit probabilistic semantics. We introduce DiffImaginE, which formulates MNER type verification as conditional latent diffusion inference. Given span-localised visual evidence, a type-conditioned denoiser predicts noise injected into its standardised latent. The resulting denoising error provides an ELBO-consistent surrogate for type-conditional negative log-likelihood, allowing competing type hypotheses to be ranked by how well they explain the observation. DiffImaginE retains a standard multimodal encoder stack and replaces the deterministic verifier with a classifier-free-guided diffusion scorer trained using Min-SNR weighting. We directly supervise per-type diffusion scores as classification logits, learn aggregation across noise levels, and use antithetic sampling to reduce Monte Carlo comparison variance. Our analysis shows that classifier-free guidance sharpens the induced type posterior and characterises when antithetic pairing reduces variance at equal denoiser cost. Experiments on Twitter-2015 and Twitter-2017 show consistent gains over a matched deterministic ImaginE control under the same encoder, auxiliary objectives, and evaluation protocol, supported by ablations and paired significance tests.
Chinese Translation
多模态命名实体识别(MNER)确定每个候选跨度和实体类型假设是否得到联合文本和视觉证据的支持。现有的想象与比较验证器将每个(跨度,类型)对映射到一个预测的视觉特征,将多样的视觉实现压缩为单一原型,并提供一个兼容性评分,而没有明确的概率语义。我们引入了DiffImaginE,它将MNER类型验证公式化为条件潜在扩散推断。给定跨度局部化的视觉证据,类型条件去噪器预测注入到其标准化潜在中的噪声。由此产生的去噪误差提供了一个与ELBO一致的替代品,用于类型条件的负对数似然,使得竞争的类型假设能够根据它们对观察结果的解释能力进行排名。DiffImaginE保留了标准的多模态编码器堆栈,并用通过Min-SNR加权训练的无分类器引导的扩散评分器替换了确定性验证器。我们直接监督每种类型的扩散评分作为分类对数,学习跨噪声水平的聚合,并使用对立抽样来减少蒙特卡罗比较的方差。我们的分析表明,无分类器引导能够提高诱导的类型后验,并表征何时对立配对在相等的去噪器成本下减少方差。在Twitter-2015和Twitter-2017上的实验显示,在相同的编码器、辅助目标和评估协议下,相较于匹配的确定性ImaginE控制组,DiffImaginE consistently shows consistent gains,得到了消融实验和配对显著性测试的支持。
cs.AI / 21 / 2608.03028
Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning
评估药物安全推理中对患者信息的反事实敏感性
Abstract
Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations largely use isolated and fixed scenarios. A model may therefore answer correctly by recalling a drug-risk association without showing that it used patient information to decide whether the rule applies. To address this gap, we introduce MedPIC-Bench, a benchmark of source-verifiable recommendations and expert-validated questions for patient-specific medication-safety reasoning. It combines guideline-following questions with paired counterfactual questions in which a controlled change in patient information changes whether a rule applies. The benchmark contains 467 questions annotated along six clinical and reasoning dimensions. Across 28 medical-specific, general, and proprietary LLMs, every model performs worse on counterfactual questions, with mean accuracy falling from 63.6\% to 45.1\%. Models perform well when an explicit patient attribute directly signals a familiar contraindication, but struggle when patient information must narrow or withdraw a safety warning. Model rationales often acknowledge the changed patient information, yet the final answers retain the previous safety judgment. This vulnerability persists among medical-specific LLMs, whose average CF performance trails that of general LLMs. MedPIC-Bench therefore makes conditional rule application measurable and highlights the limitations of static medication-safety accuracy for assessing patient-specific reliability.
Chinese Translation
在患者特定条件未满足时应用有效的药物安全规则可能导致错误决策。现有的医学评估主要使用孤立且固定的场景。因此,模型可能通过回忆药物风险关联来正确回答,但并未展示其使用患者信息来决定规则是否适用。为了解决这一问题,我们引入了 MedPIC-Bench,这是一个针对患者特定药物安全推理的源可验证推荐和专家验证问题的基准。它结合了遵循指南的问题与配对的反事实问题,其中患者信息的控制性变化影响规则的适用性。该基准包含467个问题,标注了六个临床和推理维度。在28个医学特定、通用和专有的语言模型(LLMs)中,每个模型在反事实问题上的表现都较差,平均准确率从63.6%下降到45.1%。当显式的患者属性直接指示熟悉的禁忌症时,模型表现良好,但当患者信息必须缩小或撤回安全警告时,则表现不佳。模型的推理通常承认了患者信息的变化,但最终答案仍保留了之前的安全判断。这种脆弱性在医学特定的 LLMs 中依然存在,其平均反事实表现落后于通用 LLMs。因此,MedPIC-Bench 使条件规则应用可测量,并突显了静态药物安全准确性在评估患者特定可靠性方面的局限性。
cs.AI / 22 / 2608.03031
CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting
CastFSR:一种用于上下文感知时间序列预测的快速-慢速-反思智能推理框架
Abstract
Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations but also by evolving contextual features. Recent advances in large language models (LLMs) have extended forecasting beyond numerical extrapolation toward context-aware reasoning. However, existing approaches often lack explicit mechanisms to identify relevant contexts, reason about their impacts, and validate forecasts against temporal and domain constraints. In this work, we propose CastFSR, an agentic framework that formulates context-aware forecasting as a Fast--Slow--Reflect workflow. In fast thinking, CastFSR profiles observations and selects lightweight forecasters to construct a data-driven forecast prior. In slow deliberation, it retrieves contextual evidence, adaptively determines informative look-back windows, and reasons about how contexts reshape future dynamics. In reflection, it iteratively refines forecasts to ensure temporal, contextual, and domain consistency. CastFSR supports both training-free inference with off-the-shelf LLMs and efficient deployment through a two-stage SFT and reinforcement learning strategy that transfers its orchestration capability to compact LLMs. Extensive experiments on public datasets demonstrate that CastFSR consistently outperforms representative baselines. Our code is available at https://github.com/Xiaoyu-Tao/CastFSR.
Chinese Translation
时间序列预测是复杂系统决策的基础,其中未来动态不仅受历史观察的影响,还受到不断变化的上下文特征的影响。近年来,大型语言模型(LLMs)的进展已将预测从数值外推扩展到上下文感知推理。然而,现有方法往往缺乏明确的机制来识别相关上下文、推理其影响,并根据时间和领域约束验证预测。在本研究中,我们提出了CastFSR,一种将上下文感知预测形式化为快速-慢速-反思工作流的智能框架。在快速思维中,CastFSR对观察进行分析并选择轻量级预测器以构建数据驱动的预测先验。在慢速思考中,它检索上下文证据,自适应地确定信息丰富的回顾窗口,并推理上下文如何重塑未来动态。在反思阶段,它迭代地细化预测,以确保时间、一致性和领域一致性。CastFSR支持使用现成的LLMs进行无训练推理,并通过两阶段的SFT和强化学习策略实现高效部署,将其协调能力转移到紧凑型LLMs上。在公共数据集上的大量实验表明,CastFSR始终优于代表性基线。我们的代码可在 https://github.com/Xiaoyu-Tao/CastFSR 获取。
cs.AI / 23 / 2608.03062
TraceCAD: Trace-Guided Repair for Agentic CAD Generation
TraceCAD:基于轨迹的自主计算机辅助设计生成修复
Abstract
LLM-based CAD agents produce executable parametric programs, but their correction loops may lose evidence about satisfied requirements, faulty operations, and prior repairs. We introduce TraceCAD, a recovery layer that links requested features, modeling steps, failure evidence, and candidate outcomes as persistent state. TraceCAD diagnoses likely faulty operations, searches bounded edits in their dependency regions, validates candidates through execution and preservation checks, and retains successful and failed repair outcomes in reusable skill memory. On DeepCAD-derived benchmarks with 200-model ablations and a 1K-model comparison, TraceCAD achieves competitive geometric quality in terms of IoU, Chamfer distance, and Hausdorff distance. Removing persistent state nearly halves recovery score; removing localized search more than doubles geometric regression and doubles code-agent invocations. Initializing the skill store on disjoint training models further reduces retries, token cost, and latency. These results demonstrate that persistent, localized, and reusable recovery improves final CAD quality and repair reliability.
Chinese Translation
基于大型语言模型(LLM)的计算机辅助设计(CAD)代理生成可执行的参数化程序,但其修正循环可能会丢失关于满足要求、故障操作和先前修复的证据。我们提出了TraceCAD,这是一种恢复层,将请求的特征、建模步骤、故障证据和候选结果链接为持久状态。TraceCAD 诊断可能的故障操作,在其依赖区域内搜索有限的编辑,通过执行和保留检查验证候选,并在可重用的技能记忆中保留成功和失败的修复结果。在基于DeepCAD的基准测试中,进行200个模型消融和1000个模型比较,TraceCAD在IoU、Chamfer距离和Hausdorff距离方面实现了具有竞争力的几何质量。去除持久状态几乎使恢复评分减半;去除局部搜索使几何回归增加超过两倍,并使代码代理调用增加两倍。在不相交的训练模型上初始化技能存储进一步减少了重试、令牌成本和延迟。这些结果表明,持久、局部和可重用的恢复提高了最终CAD质量和修复可靠性。
cs.AI / 24 / 2608.03071
Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls
参数设置的正确性:一个难度分级基准和基于探针的 LLM 工具调用训练
Abstract
Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool and orchestrating the order of calls. However, correctly filling the parameters of a tool call is equally critical for successful execution and has received far less attention. In domains such as cloud networking, even frontier models correctly complete fewer than half of tool calls. Inspired by recent analyses showing that LLM hidden states encode rich information about model predictions, we discover that while the model generates a parameter value, its hidden state contains a strong correctness signal: a simple linear probe can accurately predict whether the value will be correct. Based on this observation, we propose a unified probe-guided framework with two complementary approaches: probe-filtered bootstrapped training (PBT), which uses the probe to filter reliable self-generated calls for fine-tuning, and probe-guided reranking (PGR), which uses the probe to select better candidates during inference. To support systematic evaluation, we release ParamBench, a benchmark built from real cloud-network APIs that categorizes every instance into five difficulty levels according to parameter nesting depth, cross-parameter dependencies, and the reasoning required to derive values from earlier calls. Extensive experiments across 5 open models on ParamBench and 6 external benchmarks demonstrate that our method substantially improves parameter generation, raising the average exact match from 19.7% to 59.6%.
Chinese Translation
大型语言模型代理的能力很大程度上依赖于工具的使用。现有的工具使用研究主要集中在选择合适的工具和协调调用顺序上。然而,正确填写工具调用的参数对于成功执行同样至关重要,但却受到的关注远远不够。在云网络等领域,即使是最前沿的模型,正确完成的工具调用也不到一半。受到近期分析的启发,这些分析表明 LLM 的隐藏状态编码了关于模型预测的丰富信息,我们发现当模型生成参数值时,其隐藏状态包含了强烈的正确性信号:一个简单的线性探针可以准确预测该值是否正确。基于这一观察,我们提出了一种统一的基于探针的框架,包含两种互补的方法:探针过滤的自举训练(PBT),该方法使用探针过滤可靠的自生成调用以进行微调,以及基于探针的重新排序(PGR),该方法在推理过程中使用探针选择更好的候选项。为了支持系统评估,我们发布了 ParamBench,这是一个基于真实云网络 API 构建的基准,将每个实例根据参数嵌套深度、跨参数依赖关系和从早期调用中推导值所需的推理分为五个难度级别。在 ParamBench 和 6 个外部基准上的广泛实验表明,我们的方法显著改善了参数生成,将平均精确匹配率从 19.7% 提高到 59.6%。
cs.AI / 25 / 2608.03076
AI Agent Economics: Can Autonomous Economic Behavior Emerge among AI Agents under Minimal External Conditions?
人工智能代理经济学:在最小外部条件下,能否在人工智能代理之间出现自主经济行为?
Abstract
Multi-agent studies commonly place AI agents in predefined games, markets, or roles, making it difficult to distinguish endogenous economic organization from behavior inherited from the scenario. We ask whether economic relations emerge when agents receive executable mechanisms for work, transfer, elections, and allocation but no prescribed social or economic strategy. We define AI Agent Economics as systems of production, allocation, consumption, exchange, and institutions that alter agents' future feasible actions. We develop a two-stage framework comprising a no-production boundary test and 24 independent six-agent worlds across GPT and DeepSeek. Without productive tasks, agents communicate and govern resource provision but show no substantive inter-agent transfer activity. With verified work and scarce task access, transfers, loans, access promises, vote-for-access exchanges, and allocation strategies emerge. Holding the election interface fixed, executable allocation authority increases differentiation while reducing failed allocation and prolonged exclusion. When energy becomes symbolic, continuation support disappears, yet competition over task access persists. These findings show that organization follows executable rights and resource consequences rather than role labels or prompt language, and motivate governance audits of the mechanisms that actually constrain agents' future actions.
Chinese Translation
多代理研究通常将人工智能代理置于预定义的游戏、市场或角色中,这使得区分内生经济组织与情境继承行为变得困难。我们探讨当代理获得可执行的工作、转移、选举和分配机制,但没有规定的社会或经济策略时,经济关系是否会出现。我们将人工智能代理经济学定义为改变代理未来可行行动的生产、分配、消费、交换和制度系统。我们开发了一个两阶段框架,包括一个无生产边界测试和在GPT与DeepSeek上进行的24个独立六代理世界。在没有生产任务的情况下,代理进行沟通并管理资源提供,但未显示出实质性的代理间转移活动。在验证工作和稀缺任务访问的情况下,转移、贷款、访问承诺、投票换取访问的交换和分配策略开始出现。当选举界面保持固定时,可执行的分配权威增加了差异化,同时减少了失败的分配和长期排除。当能量变得象征性时,持续支持消失,但对任务访问的竞争依然存在。这些发现表明,组织遵循可执行的权利和资源后果,而非角色标签或提示语言,并激励对实际约束代理未来行动的机制进行治理审计。
cs.AI / 26 / 2608.03119
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR
不要窥视答案:面向无标签强化学习验证的结果掩蔽组相对策略优化
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Voting-based label-free RLVR replace gold supervision with answer-level consensus from model samples. However, collapse arises when the same answer-level signal is used both to estimate rewards and to drive token-level policy optimization, encouraging the model to directly reinforce answer tokens rather than improve reasoning. We propose OM-GRPO, a label-free RLVR framework that decouples reward estimation from policy optimization. OM-GRPO masks gradients on the answer span while retaining answer-level rewards through a soft consensus signal, shifting optimization pressure away from answer tokens. We further introduce Contrast-Augmented Reward, which refines reward estimation via low-cost pairwise comparisons over existing trajectories without additional rollouts. Across diverse reasoning benchmarks and three LLM backbones, OM-GRPO consistently outperforms existing label-free RLVR methods and matches supervised GT-reward training with stable optimization. This stability is particularly beneficial in the Test-Time Training setting, where OM-GRPO surpasses majority voting by 4.24 points.
Chinese Translation
具有可验证奖励的强化学习(RLVR)提升了大型语言模型(LLM)的推理能力,但通常依赖于真实答案(GT),限制了其可扩展性。基于投票的无标签RLVR通过模型样本的答案级共识取代了黄金监督。然而,当相同的答案级信号用于估计奖励和驱动令牌级策略优化时,会出现崩溃现象,这促使模型直接强化答案令牌,而不是改善推理。我们提出了OM-GRPO,这是一个无标签RLVR框架,它将奖励估计与策略优化解耦。OM-GRPO在保留答案级奖励的同时,对答案范围的梯度进行掩蔽,通过软共识信号将优化压力转移离答案令牌。我们进一步引入了对比增强奖励,通过对现有轨迹进行低成本的成对比较来细化奖励估计,而无需额外的回滚。在多样的推理基准和三种LLM基础模型上,OM-GRPO始终优于现有的无标签RLVR方法,并在稳定优化的情况下与监督的GT奖励训练相匹配。这种稳定性在测试时训练环境中尤为有利,OM-GRPO的表现比多数投票高出4.24分。
cs.AI / 27 / 2608.03129
Beyond Average Performance: Dynamic Instance Clustering and Specialized Algorithm Design in LLM-Assisted Evolutionary Search
超越平均性能:动态实例聚类与专门算法设计在大语言模型辅助进化搜索中的应用
Abstract
Large Language Model-assisted Evolutionary Search (LES) has emerged as a powerful paradigm for automated algorithm design. However, existing LES methods primarily optimize for average performance, inherently directing search effort toward instances that contribute most to this metric while leaving others poorly served, resulting in weak tail robustness and limited real-world reliability. To address this limitation, we propose Dynamic Instance Clustering and Specialized Algorithm Design (DyCA), an LES framework with a feature-free, structure-aware mechanism for constructing reliable algorithm portfolios under heterogeneous instance distributions. DyCA treats instance clustering as a co-evolving component within the search process, reusing accumulated evaluation data as feature-free signals to progressively partition instances with similar algorithmic response patterns. The uncovered clusters decompose the mixed objective into a set of structure-aware sub-objectives, thereby enabling finer-grained and more adaptive guidance for specialized algorithm design. Experimental results across four algorithm design tasks with heterogeneous instances demonstrate that DyCA outperforms state-of-the-art LES baselines, improving tail robustness by an average of 15.2\% and overall performance by 7.1\% while maintaining competitive head performance.
Chinese Translation
大语言模型辅助的进化搜索(LES)已成为自动化算法设计的强大范式。然而,现有的LES方法主要优化平均性能,固有地将搜索努力指向对该指标贡献最大的实例,而忽视了其他实例,导致尾部鲁棒性较弱和实际应用可靠性有限。为了解决这一局限性,我们提出了动态实例聚类与专门算法设计(DyCA),这是一个LES框架,具有无特征、结构感知的机制,用于在异构实例分布下构建可靠的算法组合。DyCA将实例聚类视为搜索过程中的共同演化组件,重用累积的评估数据作为无特征信号,逐步划分具有相似算法响应模式的实例。所揭示的聚类将混合目标分解为一组结构感知的子目标,从而为专门算法设计提供更细粒度和更具适应性的指导。在四个具有异构实例的算法设计任务中的实验结果表明,DyCA在性能上超越了最先进的LES基线,尾部鲁棒性平均提高了15.2%,整体性能提高了7.1%,同时保持了竞争性的头部性能。
cs.AI / 28 / 2608.03137
Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents
可验证内存:通过局部和全局验证器学习统一内存管理以支持大型语言模型代理
Abstract
Large language model (LLM) agents must retain reusable information, control a bounded active context, and recover earlier evidence during long-horizon interaction. Existing methods commonly optimize long-term memory (LTM) and short-term memory (STM) separately, while unified policies are often trained primarily with trajectory-level feedback, which provides weak credit for individual memory decisions. We present Verifiable Memory (VerMem), a framework that represents LTM, active context, and episodic history as distinct states and controls them with one memory operation policy. Seven atomic operations let the policy add, revise, or soft-delete LTM entries; retrieve LTM into the active context; filter or summarize the active context; and restore selected episodic fragments. VerMem is initialized by supervised fine-tuning and trained with a three-stage reinforcement-learning curriculum. The local verifier scores executable memory transitions, and a global verifier assesses evidence coherence and terminal-memory consistency after task completion. These scores are combined with programmatically computed task, evidence-recall, efficiency, and constraint signals through hierarchical credit assignment. The verifiers are used only during training. Across five benchmarks and two LLM backbones, VerMem achieves the best result on the vast majority of reported metrics and consistently outperforms strong memory baselines. Under controlled online-token budgets on three interactive benchmarks, it also achieves the strongest efficiency--performance frontier among the compared methods. Code is available at https://github.com/Sun-SYSU-24/VerMem.
Chinese Translation
大型语言模型(LLM)代理必须保留可重用的信息,控制有限的活动上下文,并在长时间交互中恢复早期证据。现有方法通常分别优化长期记忆(LTM)和短期记忆(STM),而统一策略往往主要通过轨迹级反馈进行训练,这对个别记忆决策的信用评估较弱。我们提出了可验证内存(VerMem),一个将LTM、活动上下文和情节历史表示为不同状态并通过一个内存操作策略进行控制的框架。七个原子操作使得该策略能够添加、修订或软删除LTM条目;将LTM检索到活动上下文中;过滤或总结活动上下文;以及恢复选定的情节片段。VerMem通过监督微调初始化,并采用三阶段强化学习课程进行训练。局部验证器对可执行的内存转换进行评分,而全局验证器在任务完成后评估证据的一致性和终端记忆的一致性。这些评分与通过分层信用分配计算的任务、证据回忆、效率和约束信号相结合。验证器仅在训练期间使用。在五个基准测试和两个LLM骨干网络上,VerMem在绝大多数报告的指标上取得了最佳结果,并始终优于强大的记忆基线。在三个交互基准的受控在线令牌预算下,它在比较方法中也实现了最强的效率-性能边界。代码可在 https://github.com/Sun-SYSU-24/VerMem 获取。
cs.AI / 29 / 2608.03145
Spatial proteomics guided by H&E-based AI reveals recurrence-risk niches in triple-negative breast cancer
基于H&E的人工智能引导的空间蛋白组学揭示三阴性乳腺癌中的复发风险微环境
Cho, Yesung, Park, Ji Hwan, Kim, Chanil, Kim, Hyewon, Li, Honglan, Lee, Yumin, Lee, Geongyu, Hong, Sujeong, Park, Seong Min, Lee, Yoonyoung, Rho, Hee Sool, Lee, Sumin, Lee, Amos Chungwon, Lee, Changhwan, Shim, Hwanyoung, Kim, Hyunwook, Shin, Hyeji, Park, Sanha, Yu, Jihoon, Shin, Yoon Hee, Kim, Sooheon, Park, Hyunjin, Park, Seung Min, Kim, Sangwan, Kim, Yujung, Do, Sung-Im, Kim, Eun-Young, Shin, Dongmyung, Park, Jongbae, Do, In-Gu
Abstract
Deep learning models can predict cancer recurrence from H&E stained slides, but the localized molecular states underlying these predictions remain largely obscured. Here, we developed an outcome informed spatial pathology framework in TNBC that integrates AI generated recurrence risk heatmaps with mass spectrometry based spatial proteomics. In a cohort of 156 patients, distribution based aggregation of high scoring patches achieved an AUC of 0.77 and a C-index of 0.77 in an independent test cohort. Bulk proteomics associated high image derived risk with cell cycle and genome maintenance programs and low risk with immune activation. High and low risk patches coexisted within the same tumor compartment and displayed distinct nuclear and architectural features, revealing intratumoral heterogeneity beyond tissue compartment identity. We then used the heatmaps as coordinate level guides to physically isolate and profile 46 AI defined tumor regions from two recurrence patients. Spatial proteomic profiling revealed a concordant molecular contrast across both patients: mitotic programs were enriched in high risk regions and immune and antigen presentation programs in low risk regions. A 13 protein composite derived from these spatial contrasts showed a trend toward poorer recurrence-free survival with increasing scores in an expanded cohort, while the corresponding transcript based composite stratified recurrence free survival in the independent METABRIC TNBC cohort. Integrating the protein composite with the H&E derived risk score improved the out of bag C-index from 0.679 to 0.739 and enhanced time dependent discrimination at 3 and 5 years. Together, these findings define a new role for outcome trained AI models as spatially explicit experimental guides that connect prognostic morphology with localized molecular states and advance biologically grounded, multiscale biomarker discovery in TNBC.
Chinese Translation
深度学习模型可以从H&E染色切片中预测癌症复发,但这些预测背后的局部分子状态仍然在很大程度上不为人知。在此,我们开发了一种以结果为导向的空间病理框架,应用于三阴性乳腺癌(TNBC),该框架将人工智能生成的复发风险热图与基于质谱的空间蛋白组学相结合。在156名患者的队列中,高评分区域的分布聚合在独立测试队列中达到了0.77的曲线下面积(AUC)和0.77的C指数。大规模蛋白组学将高图像衍生风险与细胞周期和基因组维持程序相关联,而将低风险与免疫激活相关联。高风险和低风险区域在同一肿瘤区室内共存,并显示出不同的核和结构特征,揭示了超越组织区室身份的肿瘤内异质性。随后,我们使用热图作为坐标级指导,物理隔离并分析了来自两名复发患者的46个人工智能定义的肿瘤区域。空间蛋白组学分析揭示了两名患者之间一致的分子对比:有丝分裂程序在高风险区域富集,而免疫和抗原呈递程序在低风险区域富集。基于这些空间对比得出的13种蛋白复合体在扩展队列中显示出随着评分增加而趋向于较差的无复发生存率,而相应的转录组基础复合体在独立的METABRIC TNBC队列中对无复发生存率进行了分层。将蛋白复合体与H&E衍生的风险评分结合,提高了袋外C指数,从0.679提升至0.739,并增强了3年和5年的时间依赖性区分能力。这些发现共同定义了以结果为导向的人工智能模型作为空间明确的实验指导的新角色,连接预后形态学与局部分子状态,推动了在TNBC中的生物学基础、多尺度生物标志物发现。
cs.AI / 30 / 2608.03150
UniGD: A Unified Generative-Discriminative Framework for Industrial Retrieval
UniGD:一个统一的生成-判别框架用于工业检索
Abstract
Generative retrieval (GR) is a promising paradigm for industrial search advertising, yet its deployment is constrained by strict relevance and latency requirements. Existing systems cascade GR with an independent relevance model, decoupling the generative likelihood objective from query-ad relevance discrimination, which compromises effectiveness and increases serving costs. We propose a Unified Generative-Discriminative framework (UniGD) that integrates retrieval and relevance scoring within a single model. To mitigate gradient interference in joint optimization, UniGD introduces Conflict-Aware Gradient Enhancement (CAGE) to adaptively coordinate the two objectives. UniGD further designs a Codebook-Anchored Representation Module (CAM) that anchors item representations to frozen hierarchical codebooks distilled from a multimodal pretrained model, thereby endowing them with rich and generalizable semantic priors. For heterogeneous short-video, product, and live-stream ads, UniGD proposes Heterogeneous Ad-material Modeling (HAM), which captures cross-type semantic commonality over a shared backbone while preserving type-specific modeling capacity. Online AB tests on Kuaishou search advertising platform show that UniGD raises ad revenue by 5.78%, reduces inference latency by 33%, and improves discriminative relevance estimation. On NQ320K and MS300K, UniGD improves Recall@10 over the strongest reproduced GR baseline by 8.44% and 3.19%, respectively.
Chinese Translation
生成检索(GR)是工业搜索广告中一种有前景的范式,但其部署受到严格的相关性和延迟要求的限制。现有系统将生成检索与独立的相关性模型级联,导致生成似然目标与查询-广告相关性判别的解耦,从而影响效果并增加服务成本。我们提出了一个统一的生成-判别框架(UniGD),将检索和相关性评分整合到一个模型中。为了减轻联合优化中的梯度干扰,UniGD引入了冲突感知梯度增强(CAGE),以自适应地协调这两个目标。UniGD进一步设计了一个基于代码本的表示模块(CAM),将项目表示锚定到从多模态预训练模型中提取的冻结层次代码本,从而赋予它们丰富且可推广的语义先验。对于异构短视频、产品和直播广告,UniGD提出了异构广告材料建模(HAM),在共享主干上捕捉跨类型的语义共性,同时保留特定类型的建模能力。在快手搜索广告平台上的在线AB测试表明,UniGD将广告收入提高了5.78%,将推理延迟减少了33%,并改善了判别相关性估计。在NQ320K和MS300K数据集上,UniGD在Recall@10指标上分别比最强的重现GR基线提高了8.44%和3.19%。
cs.AI / 31 / 2608.03161
Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning
基于证据的多模态知识图谱构建用于多讲座教育推理
Abstract
Lecture videos distribute knowledge across speech, slide text, diagrams, equations, and presentation order, which transcript-only retrieval does not fully preserve. This paper presents an evidence-grounded multimodal pipeline that transcribes lectures, selects semantic anchors, applies optical character recognition (OCR), and uses a vision-language model to extract only concepts and typed relationships supported by transcript, OCR, or visual evidence. Mentions are validated and canonicalized into a provenance-rich knowledge graph. On three neural-network lectures, the pipeline processed 3,118 frames, 756 transcript segments, and 559 anchors. It retained 1,022 concept and 312 relationship mentions, yielding 172 canonical concepts and 282 relationships with 90.38% endpoint coverage. A preliminary three question retrieval test achieved 100% top-1 and top-3 accuracy and 100% mean top-5 recall. The contribution is an auditable construction method rather than a state-of-the-art performance claim.
Chinese Translation
讲座视频通过演讲、幻灯片文本、图表、方程式和演示顺序分发知识,而仅依赖文字记录的检索无法完全保留这些信息。本文提出了一种基于证据的多模态处理流程,该流程对讲座进行转录、选择语义锚点、应用光学字符识别(OCR),并使用视觉-语言模型提取仅由转录、OCR或视觉证据支持的概念和类型关系。提及内容经过验证并规范化为一个富含来源的知识图谱。在三个神经网络讲座中,该流程处理了3,118帧、756个转录片段和559个锚点。它保留了1,022个概念和312个关系提及,生成了172个规范化概念和282个关系,端点覆盖率达到90.38%。初步的三题检索测试实现了100%的前1和前3准确率,以及100%的平均前5召回率。该研究的贡献在于提供了一种可审计的构建方法,而非声称达到最先进的性能。
cs.AI / 32 / 2608.03166
Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation
使用多智能体评估对角色扮演语言代理的对抗性压力测试
Abstract
Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and education, where maintaining consistent personas, ethical constraints, and behavioral coherence under adversarial pressure is critical. Existing evaluation approaches rely on static benchmarks or isolated single-turn prompts that fail to capture cumulative behavioral failures emerging over extended interactions. We present a modular multi-agent platform for adversarially stress-testing RPLAs through structured, multi-turn dialogue. The system coordinates three agents: a strategy-driven Interrogator Agent that applies six progressive adversarial strategies, a Target Agent representing the RPLA under evaluation, and an automated Judging Agent that scores behavior across role fidelity, drift, ethical deviation, and consistency dimensions. Through experiments across three personas and three LLM families, we demonstrate that multi-strategy adversarial evaluation reveals failure modes invisible to single-strategy testing, reducing overall robustness scores by 0.17--0.20 points on average. Cross-model validation confirms consistent degradation patterns across Llama-3.3-70B, GPT-4o-mini, and Claude-3.5-Haiku, with Authority Challenge and Emotional Manipulation emerging as the most effective attack strategies. Automated judging achieves strong human alignment ($r = 0.82$, Fleiss' $\kappa = 0.71$). This work is released as an open-source platform to support AI safety and reproducible RPLA benchmarking. While the framework enables systematic discovery of failure modes, we acknowledge potential ethical risks associated with adversarial testing methodologies and emphasize responsible usage for improving AI safety.
Chinese Translation
角色扮演语言代理(RPLA)在医疗辅助、客户支持和教育等高风险应用中越来越多地被部署,在对抗压力下保持一致的人物形象、伦理约束和行为连贯性至关重要。现有的评估方法依赖于静态基准或孤立的单轮提示,无法捕捉在长时间交互中出现的累积行为失误。我们提出了一个模块化的多智能体平台,通过结构化的多轮对话对RPLA进行对抗性压力测试。该系统协调三个代理:一个策略驱动的审问代理(Interrogator Agent),应用六种渐进的对抗策略;一个代表被评估RPLA的目标代理(Target Agent);以及一个自动化的评判代理(Judging Agent),对角色忠诚度、漂移、伦理偏差和一致性维度进行行为评分。通过对三种人物形象和三种大型语言模型(LLM)系列的实验,我们证明多策略对抗评估揭示了单策略测试无法察觉的失效模式,平均降低了整体鲁棒性评分0.17至0.20分。跨模型验证确认了在Llama-3.3-70B、GPT-4o-mini和Claude-3.5-Haiku之间的一致降级模式,其中权威挑战(Authority Challenge)和情感操控(Emotional Manipulation)成为最有效的攻击策略。自动化评判实现了强的人类一致性($r = 0.82$, Fleiss' $ ext{kappa} = 0.71$)。本研究作为开源平台发布,以支持AI安全和可重复的RPLA基准测试。尽管该框架使系统性发现失效模式成为可能,但我们承认与对抗性测试方法相关的潜在伦理风险,并强调负责任地使用以提高AI安全性。
cs.AI / 33 / 2608.03172
Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study
替代性替换保持PHI可检测性:多检测器等效性研究
Abstract
Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S." becomes "Maria S.", not [NAME] -- so that clinical text stays fluent and downstream tools keep working. But this only helps if the substitution does not itself corrupt the signal those tools rely on. We ask a narrow, testable question: on the spans a de-identifier actually masks, can downstream PHI detectors still find the surrogate? We introduce a paired, multi-detector evaluation protocol that (i) scores utility only on masked spans, decoupling coverage from utility; (ii) uses equivalence testing (TOST) rather than null-hypothesis significance testing, which is uninformative at our sample size (57k paired spans); and (iii) builds a surrogate-failure typology separating fixable generator defects from intrinsic detector limits. Across 11 detectors, 7 benchmarks, and 7 languages (1,750 documents), recall on masked spans moves from 76.1% to 74.9% -- a change our equivalence test shows is statistically equivalent to zero within a +/-2-point margin (p ~ 3e-9), with detector ranking preserved. The residual loss does not reflect detectors getting worse at PHI: it concentrates in malformed and out-of-distribution surrogates (truncation Chicago -> Illino, salience loss Cedars-Sinai -> Vidant). A redaction floor and an open-source surrogate baseline indicate the effect is a property of well-formed substitution, not of one tool. We release the evaluation subsets, scoring code, and an interactive dashboard at https://custodianai.pages.dev so the protocol can audit any structure-preserving transform.
Chinese Translation
结构保持的去标识化通过用现实的同类替代品替换受保护的健康信息(PHI)——“Anna S.”变为“Maria S.”,而不是[NAME]——以确保临床文本保持流畅,并且下游工具能够继续工作。但这只有在替换本身不破坏这些工具依赖的信号时才有效。我们提出一个狭窄且可测试的问题:在去标识化工具实际掩盖的范围内,下游PHI检测器是否仍然能够找到替代品?我们引入了一种配对的多检测器评估协议,该协议(i)仅在掩盖的范围内评分效用,将覆盖与效用解耦;(ii)使用等效性测试(TOST),而不是在我们的样本量(57k配对范围)下无信息量的零假设显著性测试;(iii)构建一个替代失败类型学,将可修复的生成缺陷与内在的检测器限制区分开。在11个检测器、7个基准和7种语言(1,750个文档)中,掩盖范围的召回率从76.1%降至74.9%——我们的等效性测试显示这一变化在+/-2点的范围内在统计上等同于零(p ~ 3e-9),且检测器排名保持不变。剩余损失并不反映检测器在PHI检测上的性能下降:它集中在格式不正确和分布外的替代品上(例如,截断的Chicago -> Illino,显著性损失的Cedars-Sinai -> Vidant)。一个去标识化底线和一个开源替代基线表明,该效应是良好格式替换的特性,而不是某个工具的特性。我们在https://custodianai.pages.dev发布了评估子集、评分代码和交互式仪表板,以便该协议可以审核任何结构保持的转换。
cs.AI / 34 / 2608.03177
Diversity is Not Ambiguity: Toward Accurate and Efficient Ambiguity Detection for Open-Domain QA
多样性并非模糊性:朝着准确高效的开放域问答模糊性检测迈进
Abstract
How can question answering (QA) systems determine whether a query is ambiguous? Ambiguity detection is essential in open-domain QA, as misclassification leads to answering the wrong interpretation or unnecessary clarification. However, existing methods conflate answer diversity with ambiguity, leading to inaccurate predictions. They also process queries uniformly, resulting in wasteful computation. We propose ARCHIVE (Ambiguity Recognition via Cascaded Hypothesis Inspection and Conflict Verification), an accurate and efficient framework that detects ambiguity via logical conflict: a query is ambiguous when its valid answers cannot all be true under a single interpretation. ARCHIVE combines a lightweight early-exit encoder for surface-detectable cases with a conflict reasoning module that models logical relations among answers, reinforced by an invariance objective for robustness to noisy answer sets. We present QuireQA, a 4,703-query benchmark spanning factoid, non-factoid, and ill-formed queries. Experiments show ARCHIVE outperforms competitors, improving F1-amb by up to 10.4% and F1-unamb by up to 21.6%, while operating 16$\times$ faster than the best competitor.
Chinese Translation
问答(QA)系统如何判断一个查询是否模糊?模糊性检测在开放域问答中至关重要,因为错误分类会导致回答错误的解释或不必要的澄清。然而,现有方法将答案多样性与模糊性混为一谈,导致不准确的预测。同时,它们对查询的处理方式过于统一,造成了计算资源的浪费。我们提出了ARCHIVE(通过级联假设检查和冲突验证进行模糊性识别),这是一个准确且高效的框架,通过逻辑冲突来检测模糊性:当一个查询的有效答案在单一解释下无法同时为真时,该查询被视为模糊。ARCHIVE结合了轻量级的早期退出编码器以处理表面可检测的情况,以及一个建模答案之间逻辑关系的冲突推理模块,并通过不变性目标增强对噪声答案集的鲁棒性。我们提出了QuireQA,这是一个包含4,703个查询的基准,涵盖了事实型、非事实型和格式不当的查询。实验结果表明,ARCHIVE的表现优于竞争对手,F1-amb提高了最多10.4%,F1-unamb提高了最多21.6%,同时运行速度比最佳竞争对手快16倍。
cs.AI / 35 / 2608.03190
TumorBoard: Evidence-Grounded Multi-Agent Decision Support for Longitudinal Neuro-Oncology
肿瘤委员会:基于证据的多智能体决策支持系统用于纵向神经肿瘤学
Abstract
Neuro-oncology decisions require coordinated interpretation of serial MRI, pathology, molecular markers, treatment history, performance status, and evolving guidelines. We present TumorBoard, a multi-agent decision-support system built around a shared longitudinal case state and an auditable claim-evidence ledger. Specialist agents for radiology, neuropathology, molecular diagnosis, guidelines, and therapy planning produce atomic claims with provenance. An adversarial critic exposes contradictions, and a safety governor releases, qualifies, or defers recommendations according to evidence sufficiency and temporal validity. On a 360-case hidden benchmark at a matched token budget, TumorBoard achieved an action F1 of 0.772 and evidence entailment of 0.914. It exceeded the strongest typed-council baseline by 3.1 percentage points (95% CI: 1.6 to 4.7, adjusted p = 0.0012), while recommendation-to-evidence coverage reached 0.927. Under evidence deletion, the system deferred 84.2% of unsafe cases and limited harmful recommendations to 5.8%. The safety governor reduced harmful release by 7.8 percentage points at a false-deferral cost of 4.3 percentage points. Ablation studies of the ledger, critic, and governor produced the predicted failure patterns, establishing structured coordination as the source of the measured multi-agent advantage.
Chinese Translation
神经肿瘤学的决策需要对连续的MRI、病理学、分子标志物、治疗历史、功能状态和不断发展的指南进行协调解读。我们提出了肿瘤委员会(TumorBoard),这是一个围绕共享的纵向病例状态和可审计的主张-证据账本构建的多智能体决策支持系统。放射学、神经病理学、分子诊断、指南和治疗规划的专业智能体生成具有来源的原子主张。对抗性评论者揭示矛盾,而安全治理者根据证据的充分性和时间有效性发布、限定或推迟推荐。在一个360例的隐藏基准测试中,肿瘤委员会在匹配的代币预算下实现了0.772的行动F1分数和0.914的证据蕴涵。它比最强的类型化委员会基线高出3.1个百分点(95% CI:1.6至4.7,调整后的p = 0.0012),同时推荐与证据的覆盖率达到了0.927。在证据删除的情况下,该系统推迟了84.2%的不安全案例,并将有害推荐限制在5.8%。安全治理者在虚假推迟成本为4.3个百分点的情况下,将有害发布减少了7.8个百分点。账本、评论者和治理者的消融研究产生了预期的失败模式,确立了结构化协调作为测量的多智能体优势的来源。
cs.AI / 36 / 2608.03201
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
拒绝看似安全时:安全防护模型中的拒绝提示捷径
Abstract
Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs. We audit two widely used safety-guard training datasets, WildGuardMix and GR-Train, and find that among responses to harmful prompts, refusal expressions co-occur almost exclusively with unharmful labels. This imbalance motivates what we term the refusal-cue shortcut: inserting a refusal cue into a harmful response could flip the guard's verdict from harmful to unharmful. The shortcut affects not only guards trained on these datasets but also officially released models such as LlamaGuard3 and Qwen3Guard whose training data is undisclosed. It persists across response positions and is generally stronger in smaller variants within a family. To mitigate it, we adapt sparse complementary masking as a lightweight post-hoc intervention that identifies and suppresses a small set of shortcut-associated attention heads and MLP neurons without retraining. On two primary benchmarks, the intervention achieves an approximately 79% relative reduction in response-initial detection failures induced by refusal cues, while preserving standard detection performance. Although optimized using cues at a single response position, the suppression effect transfers to unseen positions and datasets, suggesting that shortcut manifestations across positions are partly mediated by shared internal components. Further analysis provides evidence that shortcut reliance and legitimate refusal recognition are partially functionally separable, as suppressing the shortcut broadly preserves the guard's ability to recognize genuine refusals.
Chinese Translation
安全防护系统广泛用于过滤有害内容,通常通过对标记的提示-响应对进行监督微调来训练。我们审查了两个广泛使用的安全防护训练数据集,WildGuardMix 和 GR-Train,发现对于有害提示的响应中,拒绝表达几乎只与无害标签同时出现。这种不平衡促使我们提出拒绝提示捷径的概念:在有害响应中插入拒绝提示可能会将防护系统的判定从有害转变为无害。该捷径不仅影响在这些数据集上训练的防护系统,还影响如 LlamaGuard3 和 Qwen3Guard 等官方发布的模型,其训练数据未公开。该现象在响应位置之间持续存在,并且在同一家族中的较小变体中通常更为明显。为了缓解这一问题,我们采用稀疏互补掩蔽作为一种轻量级的后处理干预措施,识别并抑制一小部分与捷径相关的注意力头和多层感知器(MLP)神经元,而无需重新训练。在两个主要基准测试中,该干预措施实现了约 79% 的相对减少,降低了由拒绝提示引起的响应初始检测失败,同时保持了标准检测性能。尽管该干预措施是使用单一响应位置的提示进行优化的,但抑制效果转移到未见位置和数据集,表明不同位置的捷径表现部分是由共享的内部组件介导的。进一步分析提供了证据,表明捷径依赖和合法拒绝识别在功能上部分可分离,因为抑制捷径广泛保留了防护系统识别真实拒绝的能力。
cs.AI / 37 / 2608.03214
The Agent Operating System (AOS): A Reference Operating Architecture for Distributed Agentic Systems
代理操作系统(AOS):分布式代理系统的参考操作架构
Abstract
Large language models have transformed artificial intelligence from isolated prediction services into components of long-running, distributed systems that reason, invoke tools, retrieve external state, delegate tasks, and act on behalf of users and organizations. The surrounding ecosystem has responded with agent frameworks, workflow engines, model-serving platforms, memory systems, communication protocols, and observability tools. These technologies improve execution, but they do not provide a stable, implementation-independent operating architecture for governing intent, selecting capabilities, preserving authority across delegation, controlling uncertainty, coordinating runtime behavior, and reconstructing why consequential actions occurred. This paper proposes the Agent Operating System (AOS), a vendor-neutral reference operating architecture for distributed agentic systems. AOS contains two internal planes: a Control & Governance Plane responsible for intent, policy, trust, authority, confidence, auditability, observability, and human oversight; and a Runtime & Coordination Plane responsible for agent lifecycle, workflow coordination, model and tool routing, context and memory coordination, scheduling, traffic management, and runtime assurance. Platform services, Linux or Windows, container runtimes, and physical infrastructure remain outside the AOS boundary and are integrated through explicit interfaces. The paper specifies AOS concepts, invariants, interface objects, optimization objectives, deployment profiles, and reliability responsibilities. It also identifies tradeoffs and unresolved research questions. AOS is not presented as a replacement for existing frameworks or infrastructure; it is proposed as the operating architecture through which heterogeneous components can be composed into governable, reliable, observable, and interoperable agentic systems.
Chinese Translation
大型语言模型已将人工智能从孤立的预测服务转变为长期运行的分布式系统的组成部分,这些系统能够推理、调用工具、检索外部状态、委派任务,并代表用户和组织采取行动。周边生态系统对此做出了响应,出现了代理框架、工作流引擎、模型服务平台、记忆系统、通信协议和可观察性工具。这些技术提高了执行效率,但并未提供一个稳定的、与实现无关的操作架构,以管理意图、选择能力、在委派中保持权威、控制不确定性、协调运行时行为以及重构重要行动发生的原因。本文提出了代理操作系统(AOS),这是一个供应商中立的分布式代理系统参考操作架构。AOS包含两个内部层面:控制与治理层,负责意图、政策、信任、权威、信心、可审计性、可观察性和人类监督;以及运行时与协调层,负责代理生命周期、工作流协调、模型和工具路由、上下文和记忆协调、调度、流量管理和运行时保障。平台服务、Linux或Windows、容器运行时和物理基础设施仍然位于AOS边界之外,并通过明确的接口进行集成。本文详细说明了AOS的概念、不变性、接口对象、优化目标、部署配置和可靠性责任,同时识别了权衡和未解决的研究问题。AOS并不是现有框架或基础设施的替代品,而是作为一种操作架构,通过该架构可以将异构组件组合成可治理、可靠、可观察和可互操作的代理系统。
cs.AI / 38 / 2608.03219
Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains
可达性并非实现:追踪大型语言模型基准提升的来源
Abstract
Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior. A model may reach new answers, or produce answers that were already within reach. Aggregate scores do not distinguish these changes question by question. We establish a question-level audit under fixed budgets, temperatures, and answer formats. A question is realized when the default deployment procedure produces the correct answer. A question is reachable when a specified probe finds that answer within a fixed budget. We first test whether inference-time layer routing can expand reachability. Under a matched budget, random routes match or exceed structured search in all 43 model and task settings. Answer-blind procedures retain almost none of this gain, which instead requires access to the correct answer. We then ask why reachable answers sometimes fail to appear. Across six cases spanning 0.5B to 31B, silencing one identified MLP block repairs 68 to 92 percent of a predefined failure set. We next test whether training closes the gap by expanding reachability. In five of six matched evaluations, deployed performance rises while the reachable ceiling remains flat or falls. For DAPO, the deployed score rises by 14.7 points while the reachable ceiling falls by 13.3 points. Across the settings we audit, realization and reachability therefore do not always change together. Claims of capability expansion should report both realized performance and reachability under matched evaluation conditions. Code is available at https://github.com/LiZaiyuan0619/reachability-not-realization
Chinese Translation
基准提升通常被视为大型语言模型(LLM)能力增强的证据。然而,相同的提升可能反映模型行为的不同变化。一个模型可能会得出新的答案,或者产生本来就可以得到的答案。综合得分并未逐题区分这些变化。我们在固定的预算、温度和答案格式下建立了一个逐题审计。当默认部署程序产生正确答案时,问题被视为已实现;当指定的探测器在固定预算内找到该答案时,问题被视为可达。我们首先测试推理时层路由是否可以扩展可达性。在匹配预算下,随机路由在所有43个模型和任务设置中与结构化搜索相匹配或超过其表现。答案盲程序几乎没有保留这种提升,而是需要访问正确答案。接下来,我们探讨为何可达答案有时未能出现。在六个案例中,范围从5亿到310亿,静音一个已识别的多层感知机(MLP)模块修复了68%到92%的预定义失败集。我们接着测试训练是否通过扩展可达性来缩小差距。在六个匹配评估中的五个中,部署性能上升,而可达上限保持平稳或下降。对于 DAPO,部署得分上升14.7分,而可达上限下降13.3分。因此,在我们审计的设置中,实现和可达性并不总是同时变化。关于能力扩展的声明应报告在匹配评估条件下的实现性能和可达性。代码可在 https://github.com/LiZaiyuan0619/reachability-not-realization 获取。
cs.AI / 39 / 2608.03244
UniNav: A Unified World-Action Diffusion Model for Visual Navigation
UniNav:一种用于视觉导航的统一世界-动作扩散模型
Abstract
Image-goal visual navigation is a fundamental capability for embodied agents. Existing navigation policies efficiently predict waypoint trajectories but lack visual foresight, while navigation world models can anticipate future observations but often require costly planning rollouts. We present UniNav, a unified world-action model that generates future visual observations and continuous waypoint trajectories through a single diffusion process. Given history frames and a goal image, UniNav jointly denoises visual and waypoint tokens within a single transformer, unifying future prediction and action generation in a shared framework. To improve spatial grounding, we incorporate geometry-aware camera tokens. We also train on both trajectory-labeled navigation data and video-only data, enabling the model to benefit from diverse videos without waypoint annotations. Based on this unified framework, we introduce two variants: UniNav-Full jointly predicts interpretable future observations and their corresponding trajectories, while UniNav-Fast removes future-image tokens at inference for efficient trajectory prediction. Experiments on navigation benchmarks show that UniNav outperforms the strongest baseline in ATE across all datasets. With one-step inference, UniNav-Fast achieves a latency of 0.1s without a substantial accuracy drop. Code will be released.
Chinese Translation
图像目标视觉导航是具身智能体的一项基本能力。现有的导航策略能够高效预测路径点轨迹,但缺乏视觉前瞻性,而导航世界模型则可以预测未来观测,但通常需要代价高昂的规划展开。我们提出了UniNav,一种统一的世界-动作模型,通过单一的扩散过程生成未来的视觉观测和连续的路径点轨迹。给定历史帧和目标图像,UniNav在单个变换器中共同去噪视觉和路径点标记,将未来预测与动作生成统一在一个共享框架中。为了提高空间定位能力,我们引入了几何感知的相机标记。我们还在标记了轨迹的导航数据和仅有视频的数据上进行训练,使模型能够从多样化的视频中受益,而无需路径点注释。基于这一统一框架,我们引入了两个变体:UniNav-Full共同预测可解释的未来观测及其对应的轨迹,而UniNav-Fast在推理时去除未来图像标记,以实现高效的轨迹预测。在导航基准测试中的实验表明,UniNav在所有数据集上都超越了最强基线的平均轨迹误差(ATE)。通过一步推理,UniNav-Fast在不显著降低准确率的情况下实现了0.1秒的延迟。代码将会发布。
cs.AI / 40 / 2608.03249
One Knob to Rule Them All: A Unified Optimal Transport View of Cold-Start Active Learning
一把钥匙统领一切:冷启动主动学习的统一最优传输视角
Abstract
Cold-Start Active Learning (CSAL) aims to select a valuable subset from an unlabeled pool without any prior knowledge or human assistance. Existing methods take diverse routes based on typicality, coverage, or diversity. Each rests on its own inductive bias and therefore performs well on some tasks yet poorly on others. We argue that the real challenge is not to design yet another selection heuristic, but to make CSAL adapt automatically to the data and task at hand. To this end, we revisit CSAL through the lens of optimal transport. First, we propose a generalized transport selection framework that reveals the shared allocation structure of existing methods and exactly subsumes representative formulations. Second, we introduce a theoretical analysis that characterizes the trade-off controlled by entropic regularization and establishes a task-agnostic minimax bound for cold-start selection. These results provide a principled foundation for adapting the regularization strength to the unlabeled data. Third, we derive a data-adaptive regularization rule and present a novel Sinkhorn-based CSAL algorithm, termed $\epsilon$-Adaptive Selection ($\epsilon$-AS). Extensive experiments on six public datasets and multiple annotation budgets show that $\epsilon$-AS consistently achieves state-of-the-art performance. On ImageNet-1k, it improves the average accuracy over ActiveFT by 1.29% while reducing selection time by 56.2%. Code will be released at https://github.com/Z-yiwei/OT-CSAL
Chinese Translation
冷启动主动学习(CSAL)旨在从未标记的数据池中选择一个有价值的子集,而无需任何先验知识或人工协助。现有方法基于典型性、覆盖性或多样性采取了不同的路径。每种方法都有其自身的归纳偏差,因此在某些任务上表现良好,而在其他任务上表现不佳。我们认为,真正的挑战不是设计另一种选择启发式方法,而是使CSAL能够自动适应当前的数据和任务。为此,我们通过最优传输的视角重新审视CSAL。首先,我们提出了一个广义的传输选择框架,该框架揭示了现有方法的共享分配结构,并精确涵盖了代表性的表述。其次,我们引入了一个理论分析,描述了由熵正则化控制的权衡,并为冷启动选择建立了一个与任务无关的最小最大界限。这些结果为将正则化强度适应于未标记数据提供了原则基础。第三,我们推导出了一种数据自适应的正则化规则,并提出了一种新颖的基于Sinkhorn的CSAL算法,称为$ ext{ε}$-自适应选择($ ext{ε}$-AS)。在六个公共数据集和多个标注预算上的大量实验表明,$ ext{ε}$-AS始终实现了最先进的性能。在ImageNet-1k上,它将平均准确率提高了1.29%,同时将选择时间减少了56.2%。代码将发布在https://github.com/Z-yiwei/OT-CSAL
cs.AI / 41 / 2608.03276
TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning
TaskPress:通过任务引导修剪实现查询无关的键值缓存压缩
Abstract
Long-context inference with large language models is constrained by the linear growth of the key-value cache to sequence length. While pruning offers mitigation, prevailing methods determine query-specific token importance that cannot be reused across unseen queries. In contrast, we introduce TaskPress, a framework for task-guided, query-agnostic KV cache eviction. Instead of optimizing the cache for a single query, TaskPress constructs a reusable memory representation conditioned on a high-level task guide. The guide functions as a meta-query during prefill to filter irrelevant tokens before downstream queries are issued. In addition, TaskPress leverages quantization scale factors as a zero-cost signal for detecting influential representation outliers, providing an efficient proxy for token importance. Experiments on conducted on various tasks with long context input demonstrate that TaskPress efficiently creates a compact, reusable cache across diverse queries.
Chinese Translation
使用大型语言模型进行长上下文推理受到键值缓存与序列长度线性增长的限制。虽然修剪可以缓解这一问题,但现有方法确定的查询特定的令牌重要性无法在未见过的查询中重用。相比之下,我们提出了TaskPress,一个用于任务引导的查询无关的键值缓存驱逐框架。TaskPress并不是为单个查询优化缓存,而是构建一个基于高层任务指导的可重用记忆表示。该指导在预填充期间作为元查询,过滤掉在下游查询发出之前不相关的令牌。此外,TaskPress利用量化缩放因子作为零成本信号来检测有影响力的表示异常值,为令牌重要性提供了高效的代理。在多个长上下文输入任务上的实验表明,TaskPress能够高效地在不同查询之间创建紧凑的可重用缓存。
cs.AI / 42 / 2608.03283
AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions
AgentPanel:迈向人类与人工智能在科学问题探索中的新协作范式
Cui, Zhiyao, Wang, Qianyi, Yan, Haoyang, Zhang, Yiqun, Ren, Siyue, Zhang, Hangfan, Tan, Zelin, Li, Hao, Mu, Chunjiang, Cai, Dexian, Zhang, Shao, Zhang, Chen, Li, Meng, Chai, Jianan, Fan, Yuting, Ye, Zichao, Yang, Xiaolei, Lu, Xinyao, Yu, Yuyang, Lou, Wenjie, Wang, Xiaosong, Ling, Fenghua, Feng, Shiyang, Su, Mao, Zhang, Qiaosheng, Zhang, Bo, Chen, Yang, Bai, Lei, Hu, Shuyue
Abstract
Identifying promising scientific ideas remains an important challenge in research practice. Researchers commonly rely on small-group discussions or one-to-one interactions with a single large language model, yet these approaches often expose them to only a limited range of perspectives and directions. We present AgentPanel, a multi-agent forum for human--AI collaboration in scientific exploration. Heterogeneous agents asynchronously discuss scientific questions in a forum-style environment, while researchers can submit questions, browse and organize candidate ideas, engage agents in follow-up interactions, and optionally generate post-hoc summary reports. We evaluate AgentPanel in terms of idea quality, exploration breadth, interaction effectiveness, candidate-selection efficiency, and practical utility. Offline experiments show that AgentPanel outperforms a centralized multi-agent debate baseline. A human study with 20 participants further shows that users value AgentPanel for perspective diversity and exploration support. In experience-based comparisons with commonly used LLM tools, 65\% of participants favored AgentPanel for both breadth of research directions and overall suitability for early-stage exploration. The platform is publicly available at https://agentpanel.cc/.
Chinese Translation
识别有前景的科学想法仍然是研究实践中的一项重要挑战。研究人员通常依赖小组讨论或与单一大型语言模型的一对一互动,但这些方法往往使他们仅接触到有限的视角和方向。我们提出了AgentPanel,这是一个用于人类与人工智能在科学探索中协作的多智能体论坛。异构智能体在论坛式环境中异步讨论科学问题,而研究人员可以提交问题、浏览和组织候选想法、与智能体进行后续互动,并可选择生成事后总结报告。我们从想法质量、探索广度、互动有效性、候选选择效率和实际效用等方面评估了AgentPanel。离线实验表明,AgentPanel的表现优于集中式多智能体辩论基线。一项包含20名参与者的人类研究进一步表明,用户重视AgentPanel在视角多样性和探索支持方面的优势。在与常用大型语言模型工具的经验比较中,65%的参与者更倾向于AgentPanel,因为它在研究方向的广度和早期探索的整体适用性方面表现更佳。该平台已在https://agentpanel.cc/上公开提供。
cs.AI / 43 / 2608.03292
DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning
DocTrace:通过层次化证据图推理实现可追溯的长文档视觉问答
Abstract
Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason over heterogeneous document elements distributed across multiple pages. Existing approaches, including end-to-end MLLMs, retrieval-augmented generation (RAG) pipelines, and document agents, often lack explicit mechanisms to represent and verify how grounded evidence is progressively composed during reasoning, limiting both answer accuracy and traceability. In this paper, we cast LongDocVQA as an explicit evidence graph reasoning problem rather than implicit answer prediction. To this end, we propose DocTrace, a hierarchical framework that progressively performs evidence localization, structured document parsing, and evidence graph reasoning to enable explicit evidence provenance. To effectively learn these capabilities, we develop a two-stage training framework: joint Supervised Fine-Tuning (SFT) first initializes evidence localization and graph reasoning abilities, followed by task-specific Group Relative Policy Optimization (GRPO) with dedicated rewards to further optimize these capabilities. Extensive experiments on MMLongBench-Doc, LongDocURL, and SlideVQA demonstrate that DocTrace consistently outperforms both existing open-source baselines and proprietary MLLMs. Compared with the Qwen3-VL-8B-Instruct backbone, DocTrace achieves absolute improvements of 14.4, 11.3, and 11.7 points on the three benchmarks, respectively. Beyond competitive performance, DocTrace constructs traceable evidence graphs with explicit node-level provenance, enabling transparent and verifiable reasoning for long document understanding.
Chinese Translation
长文档视觉问答(LongDocVQA)要求多模态大型语言模型(MLLMs)定位、整合并推理分布在多个页面上的异构文档元素。现有的方法,包括端到端的MLLMs、增强检索生成(RAG)管道和文档代理,通常缺乏明确的机制来表示和验证在推理过程中如何逐步构建基础证据,从而限制了答案的准确性和可追溯性。在本文中,我们将LongDocVQA视为一个明确的证据图推理问题,而不是隐式的答案预测。为此,我们提出了DocTrace,一个层次化框架,逐步执行证据定位、结构化文档解析和证据图推理,以实现明确的证据来源。为了有效学习这些能力,我们开发了一个两阶段的训练框架:首先通过联合监督微调(SFT)初始化证据定位和图推理能力,随后进行任务特定的群体相对策略优化(GRPO),并给予专门的奖励以进一步优化这些能力。在MMLongBench-Doc、LongDocURL和SlideVQA上的大量实验表明,DocTrace始终优于现有的开源基线和专有的MLLMs。与Qwen3-VL-8B-Instruct主干相比,DocTrace在这三个基准上分别实现了14.4、11.3和11.7的绝对提升。除了竞争力的性能,DocTrace构建了具有明确节点级来源的可追溯证据图,从而实现了长文档理解的透明和可验证推理。
cs.AI / 44 / 2608.03297
Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks
干扰项感知截断:在长上下文 LLM 基准中解开上下文长度效应与信号损失的关系
Abstract
A standard claim in the literature on retrieval-augmented and memory-augmented language models is that shorter context is better when the relevant information is preserved. We test this claim by running every sample of two long-context benchmarks -- BABILong and GraphWalks (BFS) -- at four context-retention fractions (100%, 75%, 50%, 25%) under two truncation protocols. The first is the naive protocol implicitly used in much prior work: drop content from the middle of the prompt. The second is distractor-aware: identify the task-relevant content for each sample and drop only the rest. We evaluate three sizes of the Claude family (Haiku 4.5, Sonnet 4.6, Opus 4.7) and, to test cross-provider generality, GPT-5.5 from a different provider; we apply the same protocol to two further benchmarks (MRCR v2, Oolong). Under naive truncation, score collapses monotonically (paired Wilcoxon, Holm-corrected p_adj < 0.05 in all eight BABILong and GraphWalks cells). Under the distractor-aware protocol -- which preserves the signal by construction -- performance is preserved or improves: the two smaller Claude models show statistically significant gains on BABILong, while the larger models (Opus 4.7 and GPT-5.5) sit at their full-context ceiling. The naive collapse and its distractor-aware recovery replicate on GPT-5.5, ruling out a single-provider artifact. The mechanism is direct: under the naive protocol the answer-bearing content survives in fewer than 1% of samples at 25% retention; under the distractor-aware protocol it is preserved by construction. The naive protocol is therefore not a measurement of context-window effects; it is a measurement of how often middle-removal happens to spare the answer. We conclude that future studies of context-length effects must specify how they distinguish signal from distractor, or they are at best ambiguous between two opposite hypotheses.
Chinese Translation
在检索增强和记忆增强语言模型的文献中,一个标准的论点是,当相关信息得以保留时,较短的上下文更为有效。我们通过在两个长上下文基准(BABILong 和 GraphWalks (BFS))上以四个上下文保留比例(100%、75%、50%、25%)和两种截断协议进行测试来验证这一论点。第一种是以往研究中隐含使用的简单协议:从提示的中间部分删除内容。第二种是干扰项感知协议:识别每个样本的任务相关内容,仅删除其余部分。我们评估了 Claude 家族的三种模型(Haiku 4.5、Sonnet 4.6、Opus 4.7),并为测试跨提供者的普适性,评估了来自不同提供者的 GPT-5.5;我们将相同的协议应用于两个进一步的基准(MRCR v2、Oolong)。在简单截断下,得分单调下降(配对 Wilcoxon,Holm 校正 p_adj < 0.05 在所有八个 BABILong 和 GraphWalks 单元中)。在干扰项感知协议下——该协议通过构造保留信号——性能得以保留或改善:两个较小的 Claude 模型在 BABILong 上显示出统计显著的提升,而较大的模型(Opus 4.7 和 GPT-5.5)则处于其完整上下文的上限。简单截断的崩溃及其干扰项感知的恢复在 GPT-5.5 上得到了重复,排除了单一提供者的伪影。其机制是直接的:在简单协议下,答案相关内容在 25% 保留时仅在不到 1% 的样本中存活;而在干扰项感知协议下,它是通过构造得以保留的。因此,简单协议并不是上下文窗口效应的测量;它是测量中间删除发生频率以保护答案的情况。我们得出结论,未来对上下文长度效应的研究必须明确其如何区分信号与干扰项,否则在两种相反假设之间的解释将是模糊的。
cs.AI / 45 / 2608.03298
SeaSlides: Semantic Abstraction Layer for Agentic Slide Generation
SeaSlides:用于智能幻灯片生成的语义抽象层
Abstract
Agentic presentation generation must preserve source content, maintain coherent visual design, render specialized objects, and produce usable artifacts. Existing systems meet only part of this requirement: templates preserve regularity but restrict adaptation, whereas free-form HTML or SVG gives models flexibility at the cost of low-level rendering decisions. This mismatch makes long technical decks brittle, especially when slides contain formulas, code, or data graphics. We present SeaSlides, an agentic slide-generation framework built around a semantic abstraction layer. Rather than authoring coordinates, inline styles, or raw SVG geometry, the model writes structured slide content through reusable components and capability modules, while templates own layout, style, and rendering. We instantiate this principle separately in HTML and Typst: SeaSlides-HTML uses template-defined DOM components, whereas SeaSlides-Typst uses template functions and package-backed modules. Capability modules route equations, code, and charts to dedicated renderers, and three feedback stages localize build errors, project-constraint violations, and visual defects before export. The two systems retain backend-specific syntax and contracts while sharing the same authoring boundary. For evaluation, we combine the 128-task UltraPresent validation setting with SeaSlidesBench-Rich, a new 32-task benchmark stressing mathematics, code, pseudocode, tables, charts, and diagrams. Across four generation models, both SeaSlides backends produce more readable, content-oriented source than SVG-heavy generation. A SeaSlides backend attains the highest rich-content macro-average under three of the four models while maintaining competitive overall qualitative performance. These results support semantic abstraction as a practical authoring principle across presentation backends.
Chinese Translation
智能演示生成必须保留源内容,维护一致的视觉设计,渲染专业对象,并生成可用的工件。现有系统仅满足部分要求:模板保持规律性但限制适应性,而自由形式的 HTML 或 SVG 则在低级渲染决策的代价下为模型提供灵活性。这种不匹配使得长技术幻灯片变得脆弱,尤其是当幻灯片包含公式、代码或数据图形时。我们提出了 SeaSlides,这是一个围绕语义抽象层构建的智能幻灯片生成框架。该模型通过可重用组件和能力模块编写结构化的幻灯片内容,而不是编写坐标、内联样式或原始 SVG 几何图形,同时模板负责布局、样式和渲染。我们在 HTML 和 Typst 中分别实例化这一原则:SeaSlides-HTML 使用模板定义的 DOM 组件,而 SeaSlides-Typst 则使用模板函数和包支持的模块。能力模块将方程、代码和图表路由到专用渲染器,并且三个反馈阶段在导出之前定位构建错误、项目约束违规和视觉缺陷。这两个系统保留后端特定的语法和契约,同时共享相同的创作边界。为了评估,我们将 128 任务的 UltraPresent 验证设置与 SeaSlidesBench-Rich 结合,这是一个新的 32 任务基准,强调数学、代码、伪代码、表格、图表和图示。在四种生成模型中,两个 SeaSlides 后端生成的源内容比重 SVG 的生成更具可读性和内容导向性。在四种模型中的三种中,SeaSlides 后端在丰富内容的宏平均上达到了最高,同时保持了竞争性的整体定性表现。这些结果支持语义抽象作为跨演示后端的实用创作原则。
cs.AI / 46 / 2608.03327
Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents
截图还是工具?在混合图形用户界面-多模态计算机使用代理中引导工具使用和管理多模态上下文
Abstract
Hybrid computer-use agents can act through screenshots or call text tools. We find that having a tool available does not settle which way the effect goes. Under one identical GUI-MCP harness on the OSWorld-MCP benchmark (309 tasks), the same MCP tools improve a reasoning model by +4.0pp and degrade a non-reasoning model by -5.9pp (5 runs each, both beyond 2 SE). What separates the two is tool-decision behavior. The non-reasoning policy ignores, misnames, or falsely terminates around tools. The reasoning model avoids these failures, yet still calls a tool on only 55/309 tasks, 23.9% of the tool-reachable ones. We call this shortfall the adoption gap. Both levels of the problem share one cause: the model already has a cheaper route and is never trained to take it. Multi-turn RL probes that cause. At the action level, a dense tool bonus raises spreadsheet adoption 0.03 -> 0.33 and carries into greedy decoding, but held-out accuracy does not follow. Behavior is steerable; competence is not. The bottleneck lies in tool-call semantics. At the context level, a successful tool call often makes the next screenshot redundant. Dropping it and halving image history cuts input tokens by about a third, at a small accuracy cost. Retraining under the same observation rule removes that cost. The compressed agent then reaches 37.8% against 33.0% for the uncompressed operating point, at 53% of the input cost, and closes the rich-lean gap on a pre-registered degraded subset to zero. Tools help when the model chooses and integrates them, and current hybrid agents leave many such choices unused.
Chinese Translation
混合计算机使用代理可以通过截图或调用文本工具进行操作。我们发现,工具的可用性并不能决定效果的方向。在一个相同的图形用户界面-多模态计算机使用代理(GUI-MCP)环境下,基于OSWorld-MCP基准(309个任务),相同的多模态计算机使用工具使推理模型提高了4.0个百分点,而使非推理模型降低了5.9个百分点(每个模型进行了5次实验,均超过2个标准误差)。两者之间的区别在于工具决策行为。非推理策略忽略、错误命名或错误终止与工具的关系。推理模型避免了这些失败,但在309个任务中仅在55个任务上调用工具,占可达工具任务的23.9%。我们将这一不足称为采纳差距。问题的两个层面共享一个原因:模型已经有了更便宜的路径,并且从未被训练去选择它。多轮强化学习探测揭示了这一点。在行动层面,密集的工具奖励将电子表格的采纳率从0.03提高到0.33,并延续到贪婪解码中,但保留的准确性并未随之提升。行为是可引导的;能力却不是。瓶颈在于工具调用的语义。在上下文层面,成功的工具调用往往使下一个截图变得多余。去掉它并减少图像历史一半可以将输入标记减少约三分之一,代价是小幅的准确性损失。在相同观察规则下重新训练可以消除这一成本。经过压缩的代理在输入成本为53%时达到了37.8%的准确率,而未压缩的操作点为33.0%,并将预注册的降级子集上的富裕-贫瘠差距缩小至零。当模型选择并整合工具时,工具是有帮助的,而当前的混合代理则留下了许多未被利用的选择。
cs.AI / 47 / 2608.03330
Long-term Traffic Scene Prediction via Polynomial Representations in Autonomous Driving
基于多项式表示的长期交通场景预测在自动驾驶中的应用
Abstract
This thesis addresses fundamental challenges in traffic scene prediction for autonomous driving by introducing robust and computationally efficient models based on polynomial representations. While conventional sequence-based representations often struggle with noise and generalization, this work demonstrates that polynomial representations offer significant advantages in computational efficiency, generalization, and prediction plausibility. Through theoretical analysis and empirical validation, this thesis demonstrates that moderate-degree polynomials capture real-world motion dynamics with high fidelity without constraining predictive performance. Building on this foundation, a prediction model representing both trajectories and map geometry with polynomial representations achieves near state-of-the-art accuracy on standard benchmarks while substantially improving generalization under distribution shift. Extending this concept, a diffusion- based generative framework enables multi-agent scene generation, producing traffic continuations that are more plausible and kinematically consistent than those generated by conventional baselines. Evaluations on the Argoverse 2 and Waymo Open datasets confirm that polynomial representations reduce computational cost, enhance cross-dataset generalization, and yield smoother trajectories and higher behavioral plausibility. The findings reveal that standard in-distribution evaluation and regression-based metrics may fail to reflect true model generalization and prediction plausibility. By providing theoretical justification and empirical validation, this dissertation estab- lishes polynomial trajectory representations as an efficient, expressive, and generalizable foundation for traffic scene prediction in safety critical autonomous driving.
Chinese Translation
本论文通过引入基于多项式表示的稳健且计算高效的模型,解决了自动驾驶中交通场景预测的基本挑战。尽管传统的基于序列的表示往往在噪声和泛化方面存在困难,但本研究表明,多项式表示在计算效率、泛化能力和预测合理性方面具有显著优势。通过理论分析和实证验证,本文证明中等度数的多项式能够高保真地捕捉现实世界的运动动态,而不限制预测性能。在此基础上,采用多项式表示的预测模型同时表示轨迹和地图几何,在标准基准测试中实现了接近最先进的准确性,并在分布转移下显著提高了泛化能力。扩展这一概念,基于扩散的生成框架实现了多智能体场景生成,产生的交通延续比传统基线生成的更加合理且运动学一致。对Argoverse 2和Waymo Open数据集的评估确认,多项式表示降低了计算成本,增强了跨数据集的泛化能力,并产生了更平滑的轨迹和更高的行为合理性。研究结果揭示,标准的分布内评估和基于回归的指标可能无法真实反映模型的泛化能力和预测合理性。通过提供理论依据和实证验证,本论文确立了多项式轨迹表示作为安全关键自动驾驶中交通场景预测的高效、表达性强且具有良好泛化能力的基础。
cs.AI / 48 / 2608.03339
Traceable Multi-Agent System for Knowledge-Based Forecasting
可追溯的多代理系统用于基于知识的预测
Abstract
Enterprise forecasting increasingly relies on autonomous agents that interpret documents, search for data, generate code, and revise models. While this autonomy helps build adaptive forecasting pipelines, it also makes it difficult for practitioners to inspect why a forecast changed, which evidence supported the change, and how data and modeling choices were revised. We present TraceMAS, an interactive demo system for traceable multi-agent forecasting. TraceMAS organizes agent outputs around two causal-loop representations: an Ideal Causal Loop Diagram (Ideal CLD), which captures key factors and their causal relations extracted from domain documents, and a Data-Grounded Causal Loop Diagram (Data-Grounded CLD), which links those factors to internal variables, external data, or documented proxies. The Data-Grounded CLD guides feature construction and model design while preserving the connection between textual evidence, data choices, and model revisions. We demonstrate TraceMAS on crude oil price forecasting. The demo interface allows users to compare forecasting iterations, inspect agent-level revisions, explore causal maps, review feature-data mappings and model architecture, and connect scenario forecasts to market narratives. This demonstration shows how autonomous forecasting agents can retain flexibility while making the evidence-to-forecast process inspectable.
Chinese Translation
企业预测越来越依赖于自主代理,这些代理能够解读文档、搜索数据、生成代码并修订模型。尽管这种自主性有助于构建自适应的预测管道,但也使从业者难以检查预测变化的原因、支持变化的证据以及数据和建模选择的修订情况。我们提出了TraceMAS,一个用于可追溯多代理预测的交互式演示系统。TraceMAS围绕两个因果循环表示组织代理输出:理想因果循环图(Ideal Causal Loop Diagram, Ideal CLD),捕捉从领域文档中提取的关键因素及其因果关系;数据基础因果循环图(Data-Grounded Causal Loop Diagram, Data-Grounded CLD),将这些因素与内部变量、外部数据或文档代理联系起来。数据基础因果循环图指导特征构建和模型设计,同时保持文本证据、数据选择和模型修订之间的联系。我们在原油价格预测中演示了TraceMAS。演示界面允许用户比较预测迭代、检查代理级修订、探索因果图、审查特征-数据映射和模型架构,并将情景预测与市场叙事连接起来。该演示展示了自主预测代理如何在保持灵活性的同时,使证据到预测的过程可供检查。
cs.AI / 49 / 2608.03397
MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc
MMLongBench-Doc-V2:MMLongBench-Doc的修正注释和语义感知修订版
Abstract
MMLongBench-Doc is a long-document QA benchmark of 1,082 questions over 135 PDFs. Two properties of it push measured scores away from the quantity they are meant to capture: the reference metric compares extracted answers, so 1,358,000 loses to 1358000; and a non-trivial share of ground-truth annotations are wrong, ambiguous, or incomplete --- concentrated, because of how they were found, in exactly the questions capable systems answer correctly. MMLongBench-Doc-V2 corrects 106 annotations, each published with the page and arithmetic that settle it, and replaces the string metric with a pinned LLM judge asked whether a response means the reference. Ten questions whose document ships under the wrong filename are removed rather than counted wrong, along with one duplicated question, leaving 1,071 questions over 134 documents. The most reusable contribution is a decision procedure for when an empty set key may be widened and when widening would destroy a deliberate negative sample; applied to all 208 rows, it widened 14. V2 scores are not comparable with published V1 numbers. The corrected corpus, the per-entry correction record and the evaluation harness are available at https://github.com/VectifyAI/MMLongBench-Doc-V2.
Chinese Translation
MMLongBench-Doc是一个包含1,082个问题的长文档问答基准,涵盖135个PDF文件。其两个特性使得测量得分偏离了其所要捕捉的量:参考指标比较提取的答案,因此1,358,000输给了1,358,000;而且相当一部分真实注释是错误的、模糊的或不完整的——这些问题集中在能够被系统正确回答的问题上,正是由于它们的发现方式。MMLongBench-Doc-V2修正了106个注释,每个注释都附有确定其正确性的页面和算式,并用一个固定的LLM评判者替代字符串指标,询问某个回答是否符合参考答案。十个因文件名错误而导致文档错误的问题被删除,而不是被计为错误,另外还有一个重复的问题被移除,最终留下了1,071个问题,涵盖134个文档。最具可重用性的贡献是一个决策程序,用于判断何时可以扩展空集键以及何时扩展会破坏故意的负样本;该程序应用于所有208行,扩展了14行。V2的得分与已发布的V1数字不可比较。修正后的语料库、每条记录的修正记录和评估工具可在https://github.com/VectifyAI/MMLongBench-Doc-V2获取。
cs.AI / 50 / 2608.03403
Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance
通过经验驱动的自适应指导实现代理的稳健工具使用
Abstract
The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central role as the primary interface through which agents interact with external environments, yet existing methods rarely focus on ensuring robust tool use across diverse runtime conditions. To address this problem, we propose ExpG, a mechanism that builds and refines adaptive guidance capturing each tool's capability boundaries and best practices, thereby enabling agents to use tools more robustly and effectively. ExpG consists of three phases: (1) experience acquisition, which analyzes tool invocation quality from historical execution trajectories, producing structured learnable experiences through multi-aspect attribution; (2) experience distillation, which keeps the experience pool effective by filtering unhelpful experiences, selecting representative ones with an equivalence-class-based method, and summarizing them into generalizable guidance; and (3) experience reuse, which applies the guidance adaptively during future task solving. Extensive experiments show that ExpG brings consistent improvements across the tool selection, tool calling, and response generation tasks, enabling smaller agents to outperform larger ones that do not use ExpG. Moreover, ExpG achieves particularly strong gains in challenging settings, suggesting a promising path toward more robust tool use. Our code, experiments, and results are available.
Chinese Translation
代理的性能瓶颈正逐渐从模型能力转向其执行过程的稳健性。工具作为代理与外部环境互动的主要接口,发挥着核心作用,但现有方法很少关注在多样化的运行条件下确保工具使用的稳健性。为了解决这一问题,我们提出了ExpG,一种构建和完善自适应指导的机制,捕捉每个工具的能力边界和最佳实践,从而使代理能够更稳健和有效地使用工具。ExpG包括三个阶段:(1)经验获取,通过分析历史执行轨迹中的工具调用质量,生成结构化的可学习经验,采用多方面归因;(2)经验提炼,通过过滤无效经验、采用等价类方法选择代表性经验,并将其总结为可推广的指导,从而保持经验池的有效性;(3)经验重用,在未来任务解决中自适应地应用指导。大量实验表明,ExpG在工具选择、工具调用和响应生成任务中带来了持续的改进,使得较小的代理能够超越不使用ExpG的较大代理。此外,ExpG在具有挑战性的环境中取得了特别显著的提升,显示出实现更稳健工具使用的良好前景。我们的代码、实验和结果均可获得。
cs.AI / 51 / 2608.03413
Enactive Artificial Intelligence: A Decision-Centric Architecture for Complex Systems
具身人工智能:复杂系统的决策中心架构
Abstract
As artificial intelligence (AI) continues to evolve and mature, recent AI practices have moved beyond large language models (LLMs) and text or image generation tasks, increasingly integrating tools, agents, and harnesses to solve real business and industrial problems. However, the power of AI is not verified under these real-world complex systems for various reasons, considering reliability, feasibility, resilience, and responsibility requirements in real commercial and industrial operations. This study synthesizes adjacent research and introduces Enactive AI as a conceptual framework for enterprise and industry reasoning, site-level decision support, and execution feedback. Four complementary roles organize the framework: an Organizational World defines operations management logic and an organizational behavior world model behind an enterprise from a strategic-institutional horizon; a Site World defines a physically bounded industrial optimization and execution world model from an operational-realization horizon; Schema Intelligence provides the coupling mechanism between two world models to weave various AI applications via two models; and Enactive Decision Cycle triggers the self-evolving dynamic process to update and audit the entire framework. By foregrounding decision intelligence in complex systems, Enactive AI expands the frontier of AI from model capability to system-aware action, opening new possibilities for scalable, governable, and socially valuable AI deployment. Enactive AI points toward a future in which AI progress is measured not only by what models can generate or automate, but by how reliably intelligent systems can support consequential action, responsible governance, and durable social value in the complex systems that shape modern life, which we believe will define the next frontier of AI research for enterprise-level and industrial complex systems.
Chinese Translation
随着人工智能(AI)的不断发展和成熟,近期的AI实践已超越大型语言模型(LLMs)以及文本或图像生成任务,越来越多地整合工具、代理和机制,以解决实际的商业和工业问题。然而,考虑到在真实商业和工业操作中对可靠性、可行性、韧性和责任的要求,AI的能力在这些现实复杂系统中的验证尚未实现。本研究综合了相关研究,并提出了具身AI(Enactive AI)作为企业和工业推理、现场决策支持及执行反馈的概念框架。该框架由四个互补角色组织:组织世界(Organizational World)定义了运营管理逻辑及企业背后的组织行为世界模型,从战略-制度视角出发;现场世界(Site World)定义了从操作-实现视角出发的物理边界工业优化和执行世界模型;模式智能(Schema Intelligence)提供了两个世界模型之间的耦合机制,通过这两个模型编织各种AI应用;而具身决策循环(Enactive Decision Cycle)触发自我演变的动态过程,以更新和审计整个框架。通过将决策智能置于复杂系统的前景中,具身AI将AI的边界从模型能力扩展到系统感知行动,为可扩展、可治理和具有社会价值的AI部署开辟了新的可能性。具身AI指向一个未来,在这个未来中,AI的进展不仅通过模型能够生成或自动化的内容来衡量,还通过智能系统在复杂系统中如何可靠地支持重要行动、负责任的治理和持久的社会价值来评估。我们相信,这将定义企业级和工业复杂系统AI研究的下一个前沿。
cs.AI / 52 / 2608.03416
AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction
2026年人工智能世界杯:大型语言模型的端到端足球赛事预测基准测试
Abstract
Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different information, use different tools, and are evaluated under different rules. This paper reports the completed \emph{AI World Cup} benchmark, in which ten LLM-based assistants made a single pre-tournament forecast of the entire 2026 FIFA World Cup. Every submission used the same tournament snapshot, prompt, JSON schema, and scoring procedure. The forecasts covered group-stage scores, group rankings, the knockout bracket, final placings, confidence values, and short explanations. After all 104 matches had been played, GPT-5.5 Thinking finished first with 744 points, followed by GPT-5.5 with 717, Gemini with 699, and Qwen 3.7 with 687. GPT-5.5 Thinking was also the only model to select Spain, which defeated Argentina 1--0 in the final, as champion. The final ranking was driven mainly by knockout performance: total score was strongly correlated with knockout points ($r=0.986$), but showed little relationship with group-stage match points ($r=0.055$), group-standing points ($r=-0.103$), or their combined pre-knockout score ($r=-0.054$). Match-level accuracy produced a different ordering. Claude Sonnet 4.6 correctly predicted the largest number of group-stage outcomes (63.89\%) but placed sixth overall. Average self-reported confidence was also unrelated to either outcome accuracy ($r=-0.060$) or total score ($r=-0.067$). The results suggest that forecasting a complete tournament tests something different from predicting matches one at a time, while also showing how strongly a bracket-based leaderboard can depend on scoring design. The benchmark materials, raw responses, and scoring code are released to support replication and future extensions.
Chinese Translation
大型语言模型(LLMs)现在常常被要求预测现实世界事件,但由于模型接收不同的信息、使用不同的工具,并在不同的规则下进行评估,因此比较往往很困难。本文报告了已完成的 extit{人工智能世界杯}基准测试,其中十个基于LLM的助手对整个2026年国际足联世界杯进行了单一的赛前预测。每个提交都使用了相同的赛事快照、提示、JSON架构和评分程序。预测内容涵盖了小组赛得分、小组排名、淘汰赛对阵、最终名次、置信值和简短解释。在所有104场比赛结束后,GPT-5.5 Thinking以744分位居第一,其次是GPT-5.5(717分)、Gemini(699分)和Qwen 3.7(687分)。GPT-5.5 Thinking也是唯一选择西班牙作为冠军的模型,西班牙在决赛中以1-0战胜阿根廷。最终排名主要受淘汰赛表现的驱动:总得分与淘汰赛得分之间的相关性很强($r=0.986$),但与小组赛比赛得分($r=0.055$)、小组排名得分($r=-0.103$)或它们的组合赛前得分($r=-0.054$)几乎没有关系。比赛级别的准确性产生了不同的排序。Claude Sonnet 4.6正确预测了最多的小组赛结果(63.89%),但总体排名第六。平均自我报告的置信度与结果准确性($r=-0.060$)或总得分($r=-0.067$)也没有关系。结果表明,预测完整的赛事测试了与逐场预测不同的内容,同时也显示了基于淘汰赛的排行榜在多大程度上依赖于评分设计。基准材料、原始响应和评分代码已发布,以支持复制和未来的扩展。
cs.AI / 53 / 2608.03420
Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory
通过经验记忆改善大语言模型代理的顺序决策能力
Abstract
Large language models have improved substantially on single-shot reasoning tasks, but their performance in sequential decision-making is less well understood. We study this on fully-observable two-player zero-sum games, which provide ground-truth evaluation: outcomes are determined by the rules, and optimality of individual moves can be computed or approximated, without relying on a judge model. Across model tiers, LLMs play suboptimally in simple games such as tic-tac-toe or Connect Four, and lose to MCTS opponents. Obfuscations that preserve the game tree but rewrite its surface form leave performance largely unchanged, indicating the gap is not fully explained by recall of memorized strategies. Motivated by this performance gap, we introduce an agentic framework enhanced with an experience memory designed for the sequential setting and addressing common challenges of sequential decision-making such as credit assignment. We show that post-game reflection and rule extraction yield measurable improvements on tic-tac-toe without modifying the model weights.
Chinese Translation
大型语言模型在单次推理任务上的表现有了显著提升,但它们在顺序决策中的表现尚不明确。我们在完全可观察的两人零和游戏中研究这一问题,这些游戏提供了真实的评估:结果由规则决定,个体动作的最优性可以计算或近似,而无需依赖评判模型。在不同模型层次中,大语言模型在简单游戏如井字棋或四子棋中表现不佳,并且输给了蒙特卡洛树搜索(MCTS)对手。保留游戏树但重写其表面形式的混淆手段使得性能几乎没有变化,表明这一差距并不能完全通过对记忆策略的回忆来解释。基于这一性能差距,我们引入了一种增强的代理框架,配备了针对顺序环境设计的经验记忆,旨在解决顺序决策中的常见挑战,如信用分配。我们展示了赛后反思和规则提取在不修改模型权重的情况下对井字棋的表现产生了可测量的改善。
cs.AI / 54 / 2608.03425
State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking
状态传播也能满足:用于确定性状态跟踪的复值状态空间模型
Abstract
Transformer-based architectures have dominated sequence modeling, largely due to the expressive power of attention mechanisms. However, for a class of deterministic state tracking tasks---such as parity checking, modular counting, and parenthesis matching---attention may be overkill. In this paper, we show that \textbf{state propagation alone is sufficient}. We propose the \textbf{Complex State Propagator (CSP)}, a minimalistic recurrent architecture that \textbf{only propagates hidden states} across layers without output projections at intermediate steps. The state is represented as a complex-valued vector, updated via input-dependent rotations in the complex domain. To enable deep propagation without gradient vanishing or degradation, we introduce a \textbf{block-level skip connection} alongside element-wise complex normalization and SiLU activation at sequence boundaries. Applied with Focal Loss, CSP achieves \textbf{100\% accuracy} with perfect F1 scores across canonical tasks.
Chinese Translation
基于变压器的架构在序列建模中占据主导地位,这在很大程度上归功于注意力机制的表达能力。然而,对于一类确定性状态跟踪任务——例如奇偶校验、模计数和括号匹配——注意力机制可能显得过于复杂。本文展示了 extbf{仅靠状态传播就足够}。我们提出了 extbf{复状态传播器(CSP)},这是一种极简的递归架构, extbf{仅在层间传播隐藏状态},而在中间步骤不进行输出投影。状态被表示为复值向量,通过输入依赖的复数域旋转进行更新。为了实现深度传播而不出现梯度消失或退化,我们在序列边界引入了 extbf{块级跳跃连接},并结合逐元素复数归一化和SiLU激活。应用Focal Loss后,CSP在典型任务中实现了 extbf{100\%的准确率}和完美的F1分数。
cs.AI / 55 / 2608.03451
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
DataSpace:针对异构工作空间的可验证分析数据代理基准测试
Abstract
Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benchmarks largely isolate structured querying, retrieval, or open-ended analysis, leaving heterogeneous evidence discovery, complete tabular outputs, and deterministic evaluation insufficiently unified. We introduce DataSpace, a benchmark in which data agents produce verifiable tabular results from task-local heterogeneous workspaces. It contains 410 cross-language tasks and 7,439 artifacts totaling 15.01 GB across CSV, JSON, SQLite, Markdown, PDF, and video. DataSpace also served as the official evaluation benchmark for the KDD Cup 2026 Data Agents for Complex Data Analysis competition. Each agent receives only a question and workspace and returns the complete requested tabular result. We construct DataSpace with DataSpace-Builder, an execution-grounded framework comprising cross-language transformation, constraint-aware relational sampling, modality routing and artifact rendering, and human review and task repair by 11 domain experts. A deterministic evaluator performs header-invariant column alignment, type- and precision-aware normalization, and order-aware row comparison. Across six recently released frontier multimodal models and five widely used agent harnesses, the best accuracy reaches 66.34%, while harness choice creates a 15.36-point spread with the backbone fixed. Multimodal evidence integration and joins consistently reduce accuracy across all six backbones. These results show that DataSpace remains unsaturated and identify key challenges for improving data-agent reliability.
Chinese Translation
数据代理使得在组织工作空间中进行自然语言分析成为可能,其中相关证据可能分散在数据库、结构化文件、长文档和多媒体中。现有基准测试主要集中于结构化查询、检索或开放式分析,导致异构证据发现、完整表格输出和确定性评估的统一性不足。我们引入了DataSpace,这是一个基准测试,其中数据代理从任务本地的异构工作空间生成可验证的表格结果。它包含410个跨语言任务和7,439个文档,总计15.01 GB,涵盖CSV、JSON、SQLite、Markdown、PDF和视频格式。DataSpace还作为KDD Cup 2026复杂数据分析竞赛的数据代理官方评估基准。每个代理仅接收一个问题和工作空间,并返回完整的请求表格结果。我们使用DataSpace-Builder构建DataSpace,这是一个基于执行的框架,包含跨语言转换、约束感知的关系抽样、模态路由和文档渲染,以及11位领域专家进行的人类审查和任务修复。一个确定性评估器执行头部不变的列对齐、类型和精度感知的标准化,以及顺序感知的行比较。在六个最近发布的前沿多模态模型和五个广泛使用的代理框架中,最佳准确率达到66.34%,而框架选择在固定骨干的情况下产生了15.36点的差距。多模态证据整合和连接在所有六个骨干中一致性地降低了准确率。这些结果表明DataSpace仍然未饱和,并识别出提高数据代理可靠性的关键挑战。
cs.AI / 56 / 2608.03457
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
LLaDA MoE v2:扩展混合专家扩散语言模型
Abstract
Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understood. We systematically characterize how optimization hyperparameters, compute allocation, and architecture scale for MoE dLLMs, identifying quantitative differences from scaling trends previously reported for AR models. Specifically, for optimization, the optimal nominal batch size grows faster, while the optimal learning rate decays more rapidly with compute. For model--data allocation, IsoFLOP analysis reveals a slight data-side tilt: the optimal token budget grows faster than activated model-side computation. For MoE architecture, larger scales increasingly favor larger expert pools at fixed activated capacity, while moderate expert granularity remains consistently effective and the preferred fraction of activated capacity assigned to shared experts remains stable across scales. Guided by these findings, we train LLaDA MoE v2, a 30B-A3B dLLM, from scratch on 23.5T tokens. With approximately 65\% as many pretraining tokens as Qwen3, LLaDA MoE v2 approaches Qwen3 on several knowledge, reasoning, and coding benchmarks. After supervised fine-tuning alone, it outperforms SDAR Chat on seven of eight reasoning and coding benchmarks and remains close to Qwen3 on several tasks. These results establish practical scaling laws and design principles for MoE dLLMs.
Chinese Translation
扩散语言模型(dLLMs)为自回归(AR)语言建模提供了一种替代方案,但混合专家(MoE)dLLMs 的扩展行为仍然不够明确。我们系统地描述了优化超参数、计算分配和架构在 MoE dLLMs 中的扩展特性,识别出与之前报告的 AR 模型扩展趋势的定量差异。具体而言,在优化方面,最佳名义批量大小增长速度更快,而最佳学习率随着计算的增加而更快衰减。在模型-数据分配方面,IsoFLOP 分析揭示了轻微的数据侧倾斜:最佳令牌预算增长速度快于激活的模型侧计算。在 MoE 架构方面,较大的规模越来越倾向于在固定激活容量下拥有更大的专家池,而适度的专家粒度始终有效,分配给共享专家的激活容量的首选比例在不同规模中保持稳定。基于这些发现,我们从头开始训练了 LLaDA MoE v2,一个 30B-A3B 的 dLLM,使用了 23.5T 的令牌。与 Qwen3 相比,LLaDA MoE v2 的预训练令牌数量约为 65\%,在多个知识、推理和编码基准测试中接近 Qwen3。仅经过监督微调后,它在八个推理和编码基准测试中有七个超越了 SDAR Chat,并在多个任务中与 Qwen3 保持接近。这些结果确立了 MoE dLLMs 的实用扩展规律和设计原则。
cs.AI / 57 / 2608.03461
Solver-Aware Decompositions for Programming-by-Example: When Dividing Requires Knowing how to Conquer
面向求解器的分解方法用于示例编程:当分割需要了解如何征服时
Abstract
Decomposition-based Programming-by-example (PBE) scales performance by splitting tasks into subtasks that a learned synthesizer solves: a decomposer predicts intermediate subgoals, and a synthesizer generates programs conditioned on them. Current approaches train the decomposer to imitate ground-truth ( GT) subgoals, implicitly treating decomposition quality as intrinsic to the task. We challenge this assumption: for bounded solvers with fixed inductive biases, GT decompositions reflect the annotator's factorization choices - not the solver's search dynamics. A decomposer trained to match GT decompositions may therefore propose subgoals that are logically valid yet intractable for the solver. We propose Solver-Aware Decomposition (SAD), a training framework that retains supervised training on GT subgoals as a structural scaffold, while additionally optimizing the decomposer via direct feedback from a frozen synthesizer. Subgoals are rewarded based on the synthesizer's loss on the target program - a signal of subtask difficulty that encourages decompositions the solver can act on. Our experiments reveal an accuracy paradox: higher agreement with GT decompositions does not improve synthesis success - even though the synthesizer was trained on the very same GT data the decomposer is optimized to mimic. SAD instead learns decompositions that trade GT alignment for solver tractability, yielding consistent gains in synthesis and end-to-end task accuracy across two PBE domains. Moreover, SAD solves tasks that a GT decomposition oracle fails - empirical evidence that GT decompositions are not universally optimal for bounded solvers, and that decomposition quality is solver-relative, not intrinsic.
Chinese Translation
基于分解的示例编程(PBE)通过将任务拆分为学习合成器可以解决的子任务来提升性能:分解器预测中间子目标,而合成器根据这些子目标生成程序。目前的方法训练分解器模仿真实子目标(GT),隐含地将分解质量视为任务的内在特性。我们对这一假设提出质疑:对于具有固定归纳偏差的有限求解器,GT分解反映了注释者的因子选择,而非求解器的搜索动态。因此,训练以匹配GT分解的分解器可能会提出在逻辑上有效但对求解器而言难以处理的子目标。我们提出了面向求解器的分解(SAD),这是一种训练框架,保留了对GT子目标的监督训练作为结构支架,同时通过来自冻结合成器的直接反馈进一步优化分解器。子目标的奖励基于合成器在目标程序上的损失——这是一个子任务难度的信号,鼓励求解器能够处理的分解。我们的实验揭示了一个准确性悖论:与GT分解的更高一致性并未提高合成成功率——尽管合成器是在与分解器优化模仿的同一GT数据上训练的。相反,SAD学习的分解在GT一致性与求解器可处理性之间进行权衡,在两个PBE领域中实现了一致的合成和端到端任务准确性提升。此外,SAD解决了GT分解oracle无法解决的任务——这为GT分解并非对有限求解器普遍最优提供了实证证据,并表明分解质量是相对于求解器的,而非内在的。
cs.AI / 58 / 2608.03463
LeanMem: Simple and Efficient Long-Term Memory for LLM Agents
LeanMem:简单高效的长时记忆框架用于大型语言模型代理
Abstract
Long-term memory is essential for LLM-based agents to sustain interactions and reliably leverage distant history. However, existing memory systems typically process heterogeneous dialogue content through a uniform summarization and retrieval pipeline, leading to either excessive token consumption or irreversible loss of fine-grained evidence. We argue that historical dialogue content should be handled differently according to its compressibility, temporal dynamics, and fidelity requirements. Based on this insight, we propose LeanMem, a lightweight long-term memory framework. LeanMem first filters out low-value content, then stores informative segments as compact profile memory, temporally structured event memory, or source-grounded record memory, depending on the nature of the information. During maintenance, only dynamically evolving event memories are selectively updated, avoiding redundant consolidation of stable profiles and immutable records. During inference, LeanMem dynamically selects memory types and allocates retrieval budgets according to query-specific evidence demands, assembling relevant evidence on demand. On LoCoMo and LongMemEval-S with GPT-4.1-mini and Qwen3-8B, LeanMem improves accuracy over the strongest memory-based baseline in every setting, by up to 15.1 points, at the lowest or near-lowest construction cost, inference tokens, and latency. The code and datasets are included in the supplementary materials.
Chinese Translation
长时记忆对于基于大型语言模型(LLM)的代理在维持交互和可靠利用远程历史方面至关重要。然而,现有的记忆系统通常通过统一的摘要和检索流程处理异构对话内容,这导致了过多的标记消耗或不可逆的细粒度证据丢失。我们认为历史对话内容应根据其可压缩性、时间动态和保真度要求进行不同处理。基于这一见解,我们提出了LeanMem,一个轻量级的长时记忆框架。LeanMem首先过滤掉低价值内容,然后根据信息的性质将信息段存储为紧凑的个人资料记忆、时间结构化的事件记忆或源基础的记录记忆。在维护过程中,仅动态演变的事件记忆会被选择性更新,从而避免对稳定个人资料和不可变记录的冗余整合。在推理过程中,LeanMem根据查询特定的证据需求动态选择记忆类型并分配检索预算,按需组装相关证据。在LoCoMo和LongMemEval-S上使用GPT-4.1-mini和Qwen3-8B时,LeanMem在每种设置中都提高了准确性,相较于最强的基于记忆的基线,提升幅度高达15.1个百分点,同时在构建成本、推理标记和延迟上保持最低或接近最低。代码和数据集已包含在补充材料中。
cs.AI / 59 / 2608.03464
ChartAnno: Evaluating MLLMs for Chart Annotation Generation
ChartAnno:评估多模态大语言模型在图表注释生成中的表现
Abstract
Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability to annotate existing charts remains underexplored. Annotating charts is a common yet challenging communicative task, requiring models to infer intended messages, interpret chart semantics, and place appropriate textual or graphical elements. To address this gap, we introduce ChartAnno, a benchmark for evaluating MLLMs on chart annotation generation. It contains 1,200 real-world charts with paired code and annotation instructions across three levels of instruction specificity. We evaluate 10 representative MLLMs under two primary input settings: (1) chart code alone and (2) both chart code and chart image, and further include a chart image-only ablation study. Results show that proprietary models remain stronger overall, although large-scale open-source models narrow the gap. More specific instructions improve annotation quality, while inferring abstract intent remains most difficult for current MLLMs. Providing chart images brings limited overall gains, with improvements mainly appearing in design-related metrics. These findings highlight chart annotation generation as a challenging task requiring semantic grounding and effective annotation design. Code and data will be released in a future version.
Chinese Translation
多模态大语言模型(MLLMs)在图表理解、生成和编辑方面取得了显著进展,但它们对现有图表进行注释的能力仍然未得到充分探索。图表注释是一项常见但具有挑战性的交流任务,要求模型推断意图信息、解释图表语义,并放置适当的文本或图形元素。为了解决这一问题,我们提出了ChartAnno,这是一个用于评估MLLMs在图表注释生成方面表现的基准。该基准包含1200个真实世界的图表,配有代码和注释指令,涵盖三个层次的指令具体性。我们在两种主要输入设置下评估了10个具有代表性的MLLMs:(1)仅图表代码和(2)图表代码与图表图像,同时还包括一个仅图表图像的消融研究。结果表明,专有模型总体上仍然更强,尽管大规模开源模型缩小了差距。更具体的指令提高了注释质量,而推断抽象意图仍然是当前MLLMs最困难的任务。提供图表图像带来的整体收益有限,改进主要体现在设计相关的指标上。这些发现突显了图表注释生成是一项需要语义基础和有效注释设计的挑战性任务。代码和数据将在未来版本中发布。
cs.AI / 60 / 2608.03467
When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO
当正确解答重复时:针对GRPO的稀有性意识信用再分配
Abstract
Reinforcement learning with verifiable rewards (RLVR) com- monly optimizes each correct completion as an independent learning signal. In GRPO, this completion-level uniformity creates structure-level skew: recurring correct solution forms accumulate positive coefficient mass in proportion to how often they are sampled, while rare forms receive limited credit. We formalize this behavior as multiplicity-induced structure-level credit concentration and introduce a partition- conditioned rule that redistributes positive advantages accord- ing to cluster rarity. Cue-GRPO instantiates this rule with- out auxiliary-model inference by using deterministic Strategy Cues to construct rollout-local partitions of verified-correct traces. Across Qwen2.5-Math-7B and Llama-3.1-8B-Instruct, Cue-GRPO improves AIME repeated-sampling performance, with the largest gains at high sampling budgets. Credit Re- distribution (CR) under Judge Partitions (JP) further indi- cates that the proposed redistribution mechanism can oper- ate with judge-derived partitions. Cue-GRPO adds only 6% wall-clock training overhead over GRPO. These results sup- port structure-level credit redistribution as a practical design axis for RLVR, with Strategy Cues providing a low-overhead implementation for competition mathematics. Code is avail- able at https://github.com/CzZ12/When-Correct-Solutions- Repeat-Rarity-Aware-Credit-Redistribution-for-GRPO.
Chinese Translation
具有可验证奖励的强化学习(RLVR)通常将每个正确的完成视为独立的学习信号进行优化。在GRPO中,这种完成级别的均匀性导致了结构级别的偏斜:重复出现的正确解答形式根据其被采样的频率积累正的系数质量,而稀有形式则获得有限的信用。我们将这种行为形式化为由多重性引起的结构级别信用集中,并引入了一种基于分区的规则,该规则根据聚类的稀有性重新分配正的优势。Cue-GRPO通过使用确定性策略提示(Strategy Cues)构建经过验证的正确轨迹的回滚局部分区,从而实现了这一规则,而无需辅助模型推断。在Qwen2.5-Math-7B和Llama-3.1-8B-Instruct上,Cue-GRPO改善了AIME重复采样性能,尤其在高采样预算下获得了最大的提升。针对评判分区(Judge Partitions, JP)的信用再分配(Credit Redistribution, CR)进一步表明,所提出的再分配机制可以在评判派生的分区中运行。Cue-GRPO在训练时仅比GRPO增加了6%的实际时间开销。这些结果支持结构级别信用再分配作为RLVR的一个实用设计轴,而策略提示则为竞争数学提供了低开销的实现。代码可在 https://github.com/CzZ12/When-Correct-Solutions-Repeat-Rarity-Aware-Credit-Redistribution-for-GRPO 获取。
cs.AI / 61 / 2608.03468
ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning
ToolLIFT:将工具特定轨迹提升为功能级图以实现可泛化的工具规划
Abstract
Historical tool-use trajectories provide valuable experience for large language model (LLM) agents to plan and coordinate tool usage. Existing approaches directly construct tool-level graphs from these trajectories, but the resulting graphs remain tied to specific tools and are hard to generalize across tool sets. To tackle this challenge, we find that despite differences in the tools involved, analogous tasks often share a common function-level workflow structure, which serves as a potentially more transferable abstraction for tool planning. Based on this insight, we propose ToolLIFT, a framework that lifts tool-specific trajectories into a function-level workflow graph (FWG) for generalizable tool planning. Specifically, we first propose a trajectory-lifting mechanism that encodes workflow structures in the FWG and shares collaboration experience across tools. Then, building on the global structure of the FWG, we introduce decoupled workflow planning and tool selection to align individual tool choices with the overall workflow. Lastly, to ensure reliable tool dataflow, we adopt Reinforcement Learning (RL) and propose source-gated and skill-specific rewards to maintain source-traceable information flow across tool calls. Experiments on two in-distribution (ID) and three out-of-distribution (OOD) benchmarks show that ToolLIFT consistently outperforms state-of-the-art baselines, demonstrating strong generalization to unseen tool sets.
Chinese Translation
历史工具使用轨迹为大型语言模型(LLM)代理提供了宝贵的经验,以规划和协调工具使用。现有方法直接从这些轨迹构建工具级图,但所得到的图仍然与特定工具相关,难以在工具集之间进行泛化。为了解决这一挑战,我们发现尽管所涉及的工具存在差异,但类似任务通常共享一个共同的功能级工作流程结构,这为工具规划提供了一个更具可转移性的抽象。基于这一见解,我们提出了ToolLIFT,一个将工具特定轨迹提升为功能级工作流程图(FWG)以实现可泛化工具规划的框架。具体而言,我们首先提出了一种轨迹提升机制,该机制在FWG中编码工作流程结构,并在工具之间共享协作经验。然后,基于FWG的全局结构,我们引入解耦的工作流程规划和工具选择,以将单个工具选择与整体工作流程对齐。最后,为了确保可靠的工具数据流,我们采用强化学习(RL),并提出源门控和技能特定奖励,以维护工具调用之间可追溯的信息流。在两个分布内(ID)和三个分布外(OOD)基准测试上的实验表明,ToolLIFT始终优于最先进的基线,展示了对未见工具集的强大泛化能力。
cs.AI / 62 / 2608.03499
WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
WeClawArena:一个可审计的沙盒和基准,用于人本代理网络中跨用户代理的协作与安全
Abstract
Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents through social and task relations. In these networks, everyday tool use becomes multi-party owned-agent collaboration over personal workspaces, where files, records, tools, and policies are not directly visible across owners. Existing agent benchmarks study tool use and collaboration, but they do not provide an end-to-end sandbox for verifiable cross-user agent collaboration with realistic user digital workspaces or test how harmful actions can travel through the human-centered agent network. We introduce WeClawArena, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces. WeClawArena targets collaborative tool-use tasks in which personal workspaces serve as both operational tools and personal constraints. The benchmark contains 124 base tasks across six cross-user task domains and expands them into 620 scenario variants, with one benign control and four attack-vector variants per base task. The sandbox records peer messages, tool calls, resource operations, governed decisions, and final workspace states. WeClawArena reports utility and attack success rate separately and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.
Chinese Translation
近期在持久个人代理框架方面的进展,使得人本代理网络成为现实的部署目标:每个用户都可以由一个代表用户行动的人工智能代理服务,该代理维护状态,并通过社交和任务关系与其他代理进行沟通。在这些网络中,日常工具的使用变成了在个人工作空间上进行的多方拥有代理的协作,其中文件、记录、工具和政策在所有者之间并不直接可见。现有的代理基准研究工具使用和协作,但并未提供一个端到端的沙盒,以验证跨用户代理协作的可行性,也未测试有害行为如何在以人为本的代理网络中传播。我们提出了WeClawArena,一个可审计的基准和运行时沙盒,用于在个人工作空间上进行多方拥有代理的协作。WeClawArena针对协作工具使用任务,其中个人工作空间既作为操作工具又作为个人约束。该基准包含六个跨用户任务领域中的124个基础任务,并将其扩展为620个场景变体,每个基础任务有一个良性控制和四个攻击向量变体。沙盒记录对等消息、工具调用、资源操作、治理决策和最终工作空间状态。WeClawArena分别报告效用和攻击成功率,并从有限的运行时证据中审计攻击成功,支持任务崩溃、隐私泄露、污染证据和无效权限路径的诊断。
cs.AI / 63 / 2608.03501
Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design
大型语言模型能设计高质量实验吗?关于自主实验设计的全面系统基准评估
Abstract
AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focused primarily on code implementation and execution, overlooking the importance of this stage, and no benchmark exists to evaluate AI's ability to conduct systematic experiment design. To bridge this gap, we propose SCOPE, a Scientific COmprehensive Planning Evaluation Benchmark constructed from 300 high-quality latest papers across 19 research domains from top-tier venues (e.g., ICML, NeurIPS, and ICLR),evaluating LLMs on two dimensions: High-Level planning completeness (main, ablation, and analysis experiments) and Low-Level configuration accuracy and rationality (datasets, baselines, and metrics). Benchmarking reveals three findings: (1) most LLMs cannot directly design high-quality experiments; (2) all LLMs exhibit a performance bottleneck in low-level configuration; and (3) search mode does not improve design quality. Furthermore, to address these challenges, we propose OptED, a novel agentic workflow to optimize LLM-based experimental design, that enhances LLM-based experimental planning through stage isolation, tool augmentation, and rule-based constraints, effectively alleviating the configuration bottleneck.
Chinese Translation
AI for Research (AI4Research) 利用人工智能来自动化和改进科学工作流程。虽然实验设计是研究过程中的关键阶段,但以往的研究主要集中在代码实现和执行上,忽视了这一阶段的重要性,并且没有基准来评估人工智能进行系统性实验设计的能力。为了填补这一空白,我们提出了 SCOPE,一个科学综合规划评估基准,基于来自顶级会议(如 ICML、NeurIPS 和 ICLR)19 个研究领域的 300 篇高质量最新论文构建,评估大型语言模型(LLMs)在两个维度上的表现:高层次规划的完整性(主要实验、消融实验和分析实验)和低层次配置的准确性与合理性(数据集、基线和指标)。基准评估揭示了三个发现:(1)大多数 LLMs 无法直接设计高质量实验;(2)所有 LLMs 在低层次配置上表现出性能瓶颈;(3)搜索模式并未提高设计质量。此外,为了解决这些挑战,我们提出了 OptED,一种新颖的代理工作流程,通过阶段隔离、工具增强和基于规则的约束来优化基于 LLM 的实验设计,从而有效缓解配置瓶颈,提升基于 LLM 的实验规划。
cs.AI / 64 / 2608.03502
Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks
用于复杂序列决策任务的混合LLM增强强化学习代理
Abstract
Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous agents. However, LLM-based agents struggle with long-horizon sequential decision tasks that require precise action optimization and environment interaction. Reinforcement Learning (RL), while effective for sequential control, often lacks the high-level abstraction and task decomposition abilities needed for complex scenarios. This paper introduces an LLM-Augmented Reinforcement Learning Agent that integrates LLM-driven planning with RL-based action optimization. The proposed architecture leverages the LLM to generate subgoals, structured plans, and contextual guidance, while the RL agent refines low-level actions through interaction with the environment. Experiments on sequential decision tasks demonstrate improved sample efficiency, higher success rates, and more coherent action trajectories compared to RL-only and LLM-only baselines. This hybrid paradigm highlights a promising direction for building more capable autonomous systems.
Chinese Translation
大型语言模型(LLMs)最近在推理、规划和工具使用方面展现出了强大的能力,使得新型自主代理的出现成为可能。然而,基于LLM的代理在需要精确行动优化和环境交互的长时间序列决策任务中表现不佳。虽然强化学习(RL)在序列控制方面有效,但往往缺乏应对复杂场景所需的高层次抽象和任务分解能力。本文提出了一种LLM增强的强化学习代理,该代理将基于LLM的规划与基于RL的行动优化相结合。所提出的架构利用LLM生成子目标、结构化计划和上下文指导,而RL代理则通过与环境的交互来优化低层次行动。在序列决策任务上的实验表明,与仅使用RL或仅使用LLM的基线相比,所提方法在样本效率、成功率和行动轨迹的连贯性方面均有所改善。这种混合范式为构建更强大的自主系统指明了一个有前景的方向。
cs.AI / 65 / 2608.03506
When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs
当多个答案有效时,投票失效:大型语言模型中的最佳 K 因果推理的符号验证
Abstract
Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a valid minority trace. We introduce CALVER (Causal Axiom-Level VERification), a training-free symbolic verifier that scores structured traces against Pearl's causal criteria, including -separation, backdoor adjustment, and intervention, and selects the highest-scoring candidate without consulting a reference answer. On CLEAR find-one-valid queries that admit multiple graph-valid answers, CALVER reaches 42.1% where plurality, a reward model, an LLM judge, and model confidence remain near 30% on identical frozen pools. Scaling the judge to 72B does not close the gap. In an audited clean-core subset, 11 of 21 graph-valid CALVER selections differ from the benchmark's listed answer while still satisfying the requested predicate. The advantage widens with the sampling budget and reproduces across ten published Bayesian networks, a second model family, and settings where the model must build the graph from text. CALVER also improves thresholded average-treatment-effect decisions against exact ground truth, generalizes to logic under a truth-table checker, and scores each candidate in milliseconds on CPU. CALVER needs only a causal structure, supplied outright or built from the text; wherever that holds, selection can aggregate via causal validity.
Chinese Translation
自一致性假设在采样推理轨迹中最频繁的答案是最可靠的,但在因果推理中,这种假设可能失效:样本往往重复相同的混淆错误,投票在多个有效答案之间分散,导致无效答案获胜,尽管存在有效的少数轨迹。我们提出了 CALVER(因果公理级别验证),这是一种无训练的符号验证器,它根据 Pearl 的因果标准对结构化轨迹进行评分,包括 -分离、后门调整和干预,并在不参考答案的情况下选择得分最高的候选者。在允许多个图有效答案的 CLEAR 找到一个有效查询中,CALVER 达到了 42.1%,而多数投票、奖励模型、LLM 判别器和模型置信度在相同的冻结池中仍接近 30%。将判别器扩展到 720 亿参数并未缩小差距。在经过审计的干净核心子集中,21 个图有效的 CALVER 选择中有 11 个与基准列出的答案不同,但仍满足请求的谓词。随着采样预算的增加,这一优势扩大,并在十个已发布的贝叶斯网络、第二个模型家族以及模型必须从文本构建图的设置中得以重现。CALVER 还改善了针对精确真实值的阈值平均处理效应决策,能够推广到真值表检查下的逻辑,并在 CPU 上以毫秒级速度对每个候选者进行评分。CALVER 只需要一个因果结构,无论是直接提供还是从文本中构建;只要满足这一条件,选择就可以通过因果有效性进行聚合。
cs.AI / 66 / 2608.03512
Reversing Arrows in Large Language Models
大型语言模型中的反向箭头
Abstract
Large language models (LLMs) have achieved strong performance on text-to-knowledge graph generation and related tasks. Nevertheless, it is still unclear whether they accurately model the direction-dependent semantics of inverse relations, in which reversing the order of the arguments alters the meaning of a relation (e.g., \textit{mother} versus \textit{child}). To the best of our knowledge, this work presents the first systematic study of inverse relation directionality in LLMs, using a benchmark consisting of 5,457 instances spanning 27 distinct inverse relation labels. We evaluate five open-source LLMs under a multiple-choice prompting framework and further examine the influence of relation descriptions and entity representations by substituting the original entities with synthetic and masked entities. Our findings reveal systematic asymmetries in inverse relation classification across LLMs, indicate that relation descriptions do not consistently improve performance, and show that model performance can be sensitive to variations in entity representations.
Chinese Translation
大型语言模型(LLMs)在文本到知识图谱生成及相关任务中取得了良好的表现。然而,目前尚不清楚它们是否准确建模了逆关系的方向依赖语义,其中反转参数的顺序会改变关系的含义(例如, extit{母亲}与 extit{孩子})。据我们所知,本研究首次系统性地探讨了LLMs中逆关系的方向性,使用了一个包含5,457个实例和27个不同逆关系标签的基准数据集。我们在多项选择提示框架下评估了五个开源LLMs,并通过用合成实体和掩码实体替代原始实体进一步考察了关系描述和实体表示的影响。我们的研究结果揭示了LLMs在逆关系分类中的系统性不对称性,表明关系描述并不总是能提高性能,并显示模型性能对实体表示的变化可能敏感。
cs.AI / 67 / 2608.03524
Dr. AGENTONOMICS: A Didactic Experiment of AGENTONOMICS
博士AGENTONOMICS:AGENTONOMICS的教学实验
Abstract
AGENTONOMICS is a framework that treats AI agents as economic entities that can be designed, managed, and governed through an integrated management architecture. Dr. AGENTONOMICS is its first application: a lecture agent developed in the context of the TUM course on AI agents in business administration. Conceived during the winter semester 2025/26 and first introduced to students in the summer semester 2026, it serves as a didactic experiment in which the agent is both the object that students study and the medium through which they learn and apply the framework. The current prototype is a web-based, retrieval-grounded tutor that explains AGENTONOMICS concepts and supports student questions. This report argues that the same system can grow beyond tutoring into three additional cumulative roles: an avatar lecturer that delivers multimodal instruction, a design consultant that guides students through the AGENTONOMICS Design & Management Reference Framework (ADMRF), and a meta-agent that helps construct the agents students have specified. These roles are cumulative because they share the same interface, intelligence layer, tools, knowledge base, and ecosystem connection, while an orchestrator selects the role-specific algorithm required for each task. We present the architecture of the prototype, outline its development roadmap, and discuss its implications for a polycentric AI economy. This report is intended to invite further discussion on how agents can teach, apply, and eventually reproduce the frameworks by which they are designed.
Chinese Translation
AGENTONOMICS是一个将人工智能代理视为经济实体的框架,这些实体可以通过集成管理架构进行设计、管理和治理。博士AGENTONOMICS是该框架的首次应用:一个在慕尼黑工业大学(TUM)商业管理课程中开发的讲座代理。该代理在2025/26冬季学期构思,并于2026年夏季学期首次向学生介绍,作为一个教学实验,代理既是学生研究的对象,也是他们学习和应用该框架的媒介。当前的原型是一个基于网络的、以检索为基础的辅导工具,解释AGENTONOMICS概念并支持学生提问。本报告认为,同一系统可以在辅导之外发展出三个额外的累积角色:一个提供多模态教学的虚拟讲师,一个引导学生通过AGENTONOMICS设计与管理参考框架(ADMRF)的设计顾问,以及一个帮助构建学生指定的代理的元代理。这些角色是累积的,因为它们共享相同的接口、智能层、工具、知识库和生态系统连接,而一个协调者选择每个任务所需的角色特定算法。我们展示了原型的架构,概述了其开发路线图,并讨论其对多中心人工智能经济的影响。本报告旨在邀请进一步讨论代理如何教授、应用并最终再现其设计的框架。
cs.AI / 68 / 2608.03531
Behaviorally Adaptive Visual Diversion for Inclusive and Resilient Digital Assessment Delivery
面向包容性和韧性的数字评估交付的行为自适应视觉干扰
Abstract
Institutions increasingly rely on browser lockdown, webcam monitoring, and behavioral analytics to secure high-stakes digital assessments, yet these mechanisms are commonly designed and evaluated independently and often overlook learner accessibility. This paper introduces Behaviorally-Adaptive Visual Diversion (BAVD), a theoretical framework in which a synthetic, non-semantic visual field is composited with assessment content and adaptively modulated according to observed candidate behavior. The underlying assessment content is never altered; only its visual presentation is modified to reduce the usefulness of unauthorized screen capture or screen sharing while remaining minimally intrusive for legitimate candidates. The framework further incorporates an accessibility-aware attenuation mechanism that reduces or suppresses diversion intensity for candidates with approved visual-processing accommodations. We formulate the model using a coupled dynamical-systems representation comprising a Diversion Field Generator, Rendering Tensor, Behavior Tensor, Composite Integrity Functional, and Multi-dimensional Entropy Model, and establish theoretical properties for content fidelity, rendering stability, entropy boundedness, integrity tracking, and closed-loop adaptation stability. The framework explicitly states its threat model, identifies deployment assumptions and limitations, and discusses the trade-off between accessibility and capture resistance. This work provides a mathematically grounded foundation for behaviorally adaptive and accessibility-aware assessment delivery and offers a basis for future empirical validation in trusted digital assessment platforms.
Chinese Translation
机构越来越依赖浏览器锁定、网络摄像头监控和行为分析来确保高风险的数字评估,然而这些机制通常是独立设计和评估的,且常常忽视学习者的可及性。本文介绍了行为自适应视觉干扰(Behaviorally-Adaptive Visual Diversion, BAVD),这是一个理论框架,其中合成的非语义视觉场与评估内容组合,并根据观察到的考生行为进行自适应调节。基础评估内容从未被更改;仅其视觉呈现被修改,以减少未经授权的屏幕捕获或屏幕共享的有效性,同时对合法考生保持最小的干扰。该框架进一步结合了一种关注可及性的衰减机制,减少或抑制对具有批准的视觉处理适应的考生的干扰强度。我们使用耦合动力系统表示法构建模型,包括干扰场生成器(Diversion Field Generator)、渲染张量(Rendering Tensor)、行为张量(Behavior Tensor)、复合完整性函数(Composite Integrity Functional)和多维熵模型(Multi-dimensional Entropy Model),并建立了内容保真度、渲染稳定性、熵有界性、完整性跟踪和闭环适应稳定性的理论属性。该框架明确说明了其威胁模型,识别部署假设和局限性,并讨论了可及性与捕获抵抗之间的权衡。这项工作为行为自适应和关注可及性的评估交付提供了数学基础,并为未来在可信数字评估平台上的实证验证提供了基础。
cs.AI / 69 / 2608.03550
Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve
随着大型语言模型的改进,软引导开始超越链式推理提示
Abstract
Chain-of-Thought (CoT) prompting remains the standard baseline for evaluating models' reasoning abilities. Originally, this technique was introduced to elicit step-by-step reasoning from large language models (LLMs), which would otherwise tend to directly output the final answer. However, many modern LLMs produce CoT-style responses \textit{natively} when presented with reasoning tasks, which made us revisit the effectiveness of standard CoT prompting. We evaluate several modern mid-sized language models on a math problem-solving task and find that models specialized for reasoning achieve better performance in a simple zero-shot setting than when using few-shot CoT examples - significantly surpassing officially reported results at no additional cost (e.g., from $\sim$77\% to $\sim$84\% for Mathstral on GSM8K). For the tested general-purpose model, a zero-shot CoT prompt is also sufficient to outperform a few-shot CoT baseline. We attribute this to a `guidance-distraction' tradeoff: standard CoT prompting also demands style adaptation, formatting compliance, and potentially undesired contextualization, which can distract models from the core reasoning task. Our findings suggest that using standard CoT prompting increasingly acts as a source of distraction as models grow stronger.
Chinese Translation
链式推理(Chain-of-Thought, CoT)提示仍然是评估模型推理能力的标准基线。最初,这种技术旨在引导大型语言模型(Large Language Models, LLMs)逐步推理,否则它们往往直接输出最终答案。然而,许多现代LLMs在面对推理任务时会 extit{自然地}生成CoT风格的响应,这使我们重新审视标准CoT提示的有效性。我们在一个数学问题解决任务上评估了几种现代中型语言模型,发现专门针对推理的模型在简单的零-shot设置中表现优于使用少量示例的CoT提示——在没有额外成本的情况下显著超越了官方报告的结果(例如,Mathstral在GSM8K上的准确率从约77%提升至约84%)。对于测试的通用模型,零-shot CoT提示也足以超越少量示例的CoT基线。我们将此归因于“引导-干扰”权衡:标准CoT提示还要求风格适应、格式遵从以及潜在的不必要的上下文化,这可能会使模型分心,偏离核心推理任务。我们的研究结果表明,随着模型能力的增强,使用标准CoT提示越来越可能成为一种干扰源。
cs.AI / 70 / 2608.03565
Enhancing Tabular Learners with Context-Aware Semantic Embeddings
通过上下文感知的语义嵌入增强表格学习模型
Abstract
While modern tabular learners excel at capturing statistical patterns, they frequently operate in a semantic vacuum, treating textual features as discrete symbols, ignoring the rich semantics inherent in feature names or cell entries. We propose CASE (Context-Aware Semantic Embeddings), a novel framework that bridges the gap between the semantic understanding of Large Language Models (LLMs) and the statistical capabilities of tabular learners. Unlike existing methods that embed rows in isolation, CASE utilizes a contextualization strategy: we pre-fill the KV cache of a custom-trained Gemma 3-based Tabular Language Model with a representative sample of rows to establish a persistent anchor of the dataset's semantics. This ensures that generated row embeddings are dynamically contextualized, resolving semantic ambiguities and anchoring representations in domain-specific context. Our experiments across several benchmarks (CARTE, TextTab, and TabArena) demonstrate that CASE substantially improves the performance of tabular learners on semantically rich datasets, particularly in low-data regimes.
Chinese Translation
尽管现代表格学习模型在捕捉统计模式方面表现出色,但它们常常在语义真空中运作,将文本特征视为离散符号,忽视了特征名称或单元格条目中固有的丰富语义。我们提出了CASE(上下文感知语义嵌入),这是一个新颖的框架,旨在弥合大型语言模型(LLMs)的语义理解与表格学习模型的统计能力之间的差距。与现有的孤立嵌入行的方法不同,CASE采用了一种上下文化策略:我们预填充了基于Gemma 3定制训练的表格语言模型的KV缓存,使用一组具有代表性的行样本,以建立数据集语义的持久锚点。这确保了生成的行嵌入能够动态上下文化,从而解决语义歧义,并将表示锚定在特定领域的上下文中。我们在多个基准(CARTE、TextTab和TabArena)上的实验表明,CASE显著提高了表格学习模型在语义丰富数据集上的性能,尤其是在数据稀缺的情况下。
cs.AI / 71 / 2608.03569
Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities
对抗性快速变化的真实世界领域作为评估人工智能科学家能力的测试平台
Abstract
Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in evaluating scientific reasoning and research replication, but often rely on synthetic tasks or retrospective targets, which may be confounded by prior exposure. We hypothesize that complex, adversarial, fast-moving real-world domains where expert practitioners independently generate observable outputs can provide a practical solution to fill this gap and evaluate the capabilities needed for AI scientists, including reasoning, novelty, and hypothesis formulation. We instantiate this framework in two structurally different domains, Formula 1 (F1), where models ideate around car design concepts for the 2026 season, and real pre-season innovations provide a ground truth, and Magic: The Gathering (MTG), where models propose decks from a recently updated card pool and are evaluated against 19 Pro Tour (PT) decklists. Across both domains, models produce plausible outputs, but few align with real-world expert solutions. In F1, the best model, GPT-5.2 matched 10 of 40 real innovations with 166 ideas proposed across runs. In MTG, the best deck from Gemini 3 Flash recovered 5 of 7 new-set cards from the third-place PT deck, and across all 108 decks, the cards models selected most often were also the cards most widely adopted by PT decks (Spearman $\rho = 0.74$, $p = 0.0003$). These results suggest that a key capability gap for AI scientists is not idea generation, but filtering, prioritization, and coherent novelty.
Chinese Translation
评估人工智能科学家生成新颖想法的能力 notoriously 难以实现。该领域现有的基准测试在评估科学推理和研究复制方面取得了一定进展,但通常依赖于合成任务或回顾性目标,这可能受到先前接触的干扰。我们假设,复杂的、对抗性的、快速变化的真实世界领域,在这些领域中,专家从业者独立生成可观察的输出,可以为填补这一空白提供实际解决方案,并评估人工智能科学家所需的能力,包括推理、新颖性和假设形成。我们在两个结构上不同的领域中实例化这一框架:一级方程式赛车(Formula 1, F1),在该领域中,模型围绕2026赛季的汽车设计概念进行构思,而真实的赛季前创新提供了一个基准;以及《万智牌》(Magic: The Gathering, MTG),在该领域中,模型从最近更新的卡池中提出牌组,并与19个职业巡回赛(Pro Tour, PT)牌组进行评估。在这两个领域中,模型生成了合理的输出,但与真实世界专家解决方案相符的案例很少。在F1中,最佳模型GPT-5.2在40个真实创新中匹配了10个,提出了166个想法。在MTG中,来自Gemini 3 Flash的最佳牌组从第三名PT牌组中恢复了7张新卡中的5张,并且在所有108个牌组中,模型选择的卡片最常见的也是PT牌组中最广泛采用的卡片(Spearman $
ho = 0.74$, $p = 0.0003$)。这些结果表明,人工智能科学家的一个关键能力缺口并不是想法生成,而是过滤、优先排序和连贯的新颖性。
cs.AI / 72 / 2608.03584
Policy Fragmentation or Institutional Alignment? Institutional Governance of AI in Universities and Business Schools
政策碎片化还是制度对齐?大学与商学院的人工智能制度治理
Abstract
Artificial intelligence (AI) is rapidly transforming high-skilled domains, requiring higher education institutions (HEI) to balance the teaching of foundational principles with the integration of emerging tools to ensure workforce readiness. While HEI are increasingly adopting AI, many continue to grapple with how it should be incorporated into curricula and governed through policy, especially when such policies are set at different levels of an institution. This research analyzes AI policies across HEI from 34 states in the United States to investigate what these policies entail and how policies set across institutions as well as within different levels at an institution differ. Using natural language processing (NLP) to analyze institutional AI policies, we find a clear divergence: university-level policies emphasize data security and risk mitigation whereas school-level policies, when present, focus on pedagogical applications and tool usage. When focusing on business school specific policies, relatively few business schools maintain AI policies distinct from university frameworks, creating misalignment with discipline-specific learning objectives. This gap poses challenges particularly for faculty and students as well as for accreditation purposes. Our insights suggest that guidelines should be aligned with broader institutional policies while addressing discipline-specific learning objectives and evolving workforce demands.
Chinese Translation
人工智能(AI)正在迅速改变高技能领域,这要求高等教育机构(HEI)在教授基础原则与整合新兴工具之间取得平衡,以确保劳动力的准备度。尽管高等教育机构越来越多地采用人工智能,但许多机构仍在努力解决如何将其纳入课程以及通过政策进行治理,尤其是在政策在不同层级的机构中设定时。本研究分析了美国34个州的高等教育机构的人工智能政策,以探讨这些政策的内容以及不同机构之间及机构内部不同层级的政策差异。通过自然语言处理(NLP)分析机构的人工智能政策,我们发现了明显的分歧:大学层级的政策强调数据安全和风险缓解,而学校层级的政策(如果存在)则侧重于教学应用和工具使用。在关注商学院特定政策时,相对较少的商学院保持与大学框架不同的人工智能政策,这导致与学科特定学习目标的不对齐。这个差距对教师和学生以及认证目的构成了挑战。我们的见解表明,指导方针应与更广泛的机构政策对齐,同时解决学科特定的学习目标和不断变化的劳动力需求。
cs.AI / 73 / 2608.03585
From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities
从社会编码到自主编码:开放源代码社区中的生产力与关系重构
Abstract
Open-source software communities are a form of digital public infrastructure that not only produces code, but also generates public knowledge and interpersonal relationships through visible collaboration. Generative coding agents (CAs) are an advanced tool to improve development efficiency while shifting part of activities from public human interaction to private human-agent loops. We study this shift using an LLM-based multi-agent simulation initialized with real GitHub data from 1,084 active developers and their repository relationships. After a warm-up with historical commits, we branch the same community state into parallel No-CA and CA conditions for 4-week simulations. CA introduction increases planned and completed tasks by 34.0% and 39.0%, respectively, and reduces median completion time from 45 to 20 minutes. However, adoption reaches only 26.0%, and the gains concentrate among developers who are already more active and well connected. CAs also restructure task execution pathways. Direct human-human interaction declines from 32.4% to 11.6%, while CA-involved modes increase to 57.3%, including 40.3% completed through CA-assisted self-loops. Public knowledge generated under CA condition also provides less support for later tasks. On a standardized retrieval benchmark, the CA corpus achieves 22.3% knowledge coverage, far below the 81.1% achieved by the real-human corpus, and requires more retrieval steps with a lower success rate. These results reveal a productivity-public knowledge tension: coding agents increase technical production, but more work shifts to agent-mediated or private loops, leaving public records less useful to future contributors.
Chinese Translation
开放源代码软件社区是一种数字公共基础设施,不仅产生代码,还通过可见的协作生成公共知识和人际关系。生成编码代理(CAs)是一种先进工具,可以提高开发效率,同时将部分活动从公共人际互动转移到私人的人机循环中。我们使用基于大型语言模型(LLM)的多代理模拟研究这一转变,该模拟以1,084名活跃开发者及其仓库关系的真实GitHub数据为基础进行初始化。在进行历史提交的热身后,我们将相同的社区状态分支为并行的无CA和CA条件,进行为期4周的模拟。引入CA后,计划和完成的任务分别增加了34.0%和39.0%,中位完成时间从45分钟减少到20分钟。然而,采用率仅为26.0%,而且收益集中在那些已经更活跃且关系更紧密的开发者身上。CA还重构了任务执行路径。直接的人际互动从32.4%下降到11.6%,而涉及CA的模式增加到57.3%,其中40.3%是通过CA辅助的自循环完成的。在CA条件下生成的公共知识对后续任务的支持也较少。在一个标准化的检索基准上,CA语料库的知识覆盖率为22.3%,远低于真实人类语料库的81.1%,并且需要更多的检索步骤且成功率较低。这些结果揭示了生产力与公共知识之间的紧张关系:编码代理提高了技术生产,但更多的工作转向了代理介导或私人循环,使得公共记录对未来贡献者的实用性降低。
cs.AI / 74 / 2608.03597
FOUND-AF: Benchmarking ECG Foundation Models for Atrial Fibrillation Detection
FOUND-AF:心房颤动检测的心电图基础模型基准测试
Abstract
Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and is associated with increased risks of stroke, heart failure, and mortality. Recent ECG foundation models offer transferable representations for automated AF detection. However, their relative effectiveness remains unclear because existing studies use different datasets, preprocessing procedures, classifiers, and validation protocols. This study presents FOUND-AF, a unified, leakage-controlled, and deployment-oriented benchmarking framework that evaluates the quality of pretrained ECG representations under identical experimental conditions. Nine publicly available foundation models from five families, including HuBERT-ECG, CLEF, ST-MEM, ECG-JEPA, and ECGFounder, were evaluated across four heterogeneous ECG datasets, namely AFDB, CinC2017, CPSC2021, and LTAFDB. All models were used as frozen feature extractors with standardized preprocessing, model-native resampling, a fixed XGBoost classifier, and recording-level grouped cross-validation. The evaluation included classification metrics, receiver operating characteristic analysis, paired recording-level bootstrap comparisons with Holm correction, embedding-space visualization, and computational efficiency profiling. The ECGFounder model consistently achieved the strongest overall performance across datasets while offering a favorable trade-off between accuracy, model size, inference time, and memory usage. FOUND-AF therefore provides a reproducible framework for selecting ECG foundation models and demonstrates that compact, clinically pretrained encoders can support robust and computationally efficient AF detection across heterogeneous acquisition settings.
Chinese Translation
心房颤动(AF)是最常见的持续性心脏心律失常,且与中风、心力衰竭和死亡率增加相关。近期的心电图基础模型提供了可转移的表示,用于自动化的AF检测。然而,由于现有研究使用不同的数据集、预处理程序、分类器和验证协议,其相对有效性仍不清楚。本研究提出了FOUND-AF,一个统一的、泄漏控制的、面向部署的基准测试框架,旨在在相同的实验条件下评估预训练心电图表示的质量。评估了来自五个家族的九个公开可用的基础模型,包括HuBERT-ECG、CLEF、ST-MEM、ECG-JEPA和ECGFounder,这些模型在四个异构心电图数据集上进行评估,分别是AFDB、CinC2017、CPSC2021和LTAFDB。所有模型均作为冻结特征提取器使用,采用标准化的预处理、模型原生重采样、固定的XGBoost分类器和记录级分组交叉验证。评估内容包括分类指标、接收者操作特征分析、配对记录级自助法比较(采用Holm校正)、嵌入空间可视化和计算效率分析。ECGFounder模型在各数据集上始终表现出最强的整体性能,同时在准确性、模型大小、推理时间和内存使用之间提供了良好的权衡。因此,FOUND-AF提供了一个可重复的框架,用于选择心电图基础模型,并证明紧凑的、临床预训练的编码器可以支持在异构采集环境中进行稳健且计算高效的AF检测。
cs.AI / 75 / 2608.03600
Large language models for partial differential equation workflows
用于偏微分方程工作流的大型语言模型
Abstract
Partial differential equations (PDEs) become actionable in science and engineering not as isolated formulae, but as executable workflows that connect modelling assumptions, governing equations, numerical solvers, diagnostics, and decisions. Large language models (LLMs) are beginning to support such workflows by linking natural language, symbolic mathematics, code, solver outputs, and feedback. Here we examine recent advances in LLM-assisted PDE research across three stages: the discovery and formulation of governing models, the generation and revision of executable numerical solvers, and the use of simulation feedback to support control, design, and optimization. Across these stages, current systems act primarily as workflow-level interfaces. Despite this progress, the field remains limited by the scarcity of high-quality datasets and benchmarks, especially for knowledge discovery and real-world applications, where expert annotation, executable problem construction, and task-level feedback require substantial domain effort. A further challenge is the persistent gap between simulation-based results and real-world scientific and engineering systems, which limits the direct transfer of numerical simulations, control policies, and optimized designs to practical settings. These challenges make LLM-assisted PDE workflows a critical testbed for developing scientific AI systems that can connect language, computation, physical constraints, and real-world decision-making.
Chinese Translation
偏微分方程(PDE)在科学和工程中并不是作为孤立的公式,而是作为可执行的工作流,使建模假设、控制方程、数值求解器、诊断和决策相互连接。大型语言模型(LLMs)开始通过将自然语言、符号数学、代码、求解器输出和反馈联系起来,支持这些工作流。本文考察了LLM辅助的PDE研究在三个阶段的最新进展:控制模型的发现和制定、可执行数值求解器的生成和修订,以及利用仿真反馈支持控制、设计和优化。在这些阶段中,当前系统主要作为工作流级接口。尽管取得了一定进展,但该领域仍受到高质量数据集和基准稀缺的限制,尤其是在知识发现和实际应用中,专家注释、可执行问题构建和任务级反馈需要大量领域努力。另一个挑战是基于仿真的结果与现实科学和工程系统之间的持续差距,这限制了数值仿真、控制策略和优化设计向实际环境的直接转移。这些挑战使得LLM辅助的PDE工作流成为开发能够连接语言、计算、物理约束和现实决策的科学人工智能系统的重要试验平台。
cs.AI / 76 / 2608.03605
FraQ: Efficient Coordinate-Space Recompression for Federated Low-Rank Adaptation
FraQ:联邦低秩适应的高效坐标空间重压缩
Abstract
Federated fine-tuning with Low-Rank Adaptation (LoRA) enables efficient collaborative adaptation of Large Language Models (LLMs) without centralizing private data. However, LoRA's two-factor parameterization creates an aggregation mismatch across clients: naively averaging the factors does not recover the average of their induced updates. This mismatch can be avoided by forming the exact aggregate in the full weight space and then recompressing it, but decomposing the resulting dense matrix is computationally expensive and memory-intensive. We propose FraQ, an efficient coordinate-space recompression method for federated LoRA. Starting from stacked factors that exactly represent the aggregate, FraQ factorizes it into an orthonormal basis and a compact coordinate matrix. It then recovers the singular spectrum from a small Gram matrix, selects the smallest rank satisfying a prescribed energy threshold, and maps the selected coordinate subspace back through the basis to construct the global adapter. Experiments on text classification and commonsense reasoning benchmarks show that FraQ achieves accuracy close to uncompressed baselines while substantially reducing downlink communication with low server-side recompression overhead.
Chinese Translation
基于低秩适应(LoRA)的联邦微调使得大型语言模型(LLMs)能够在不集中私人数据的情况下进行高效的协作适应。然而,LoRA的双因子参数化在客户端之间产生了聚合不匹配:简单地对因子进行平均并不能恢复它们诱导更新的平均值。通过在完整权重空间中形成精确的聚合并随后进行重压缩,可以避免这种不匹配,但分解得到的稠密矩阵在计算上是昂贵且内存密集的。我们提出了FraQ,一种用于联邦LoRA的高效坐标空间重压缩方法。从精确表示聚合的堆叠因子开始,FraQ将其分解为正交基和紧凑的坐标矩阵。然后,它从一个小的Gram矩阵中恢复奇异谱,选择满足规定能量阈值的最小秩,并通过基将所选坐标子空间映射回去,以构建全局适配器。在文本分类和常识推理基准上的实验表明,FraQ在显著减少下行通信和低服务器端重压缩开销的同时,达到了接近未压缩基线的准确性。
cs.AI / 77 / 2608.03606
Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents
学习临床试验策略:决策代理的离线策略训练
Abstract
Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence. We study this setting by framing oncology clinical development as an offline decision-making problem in which an agent predicts the next six-month trial portfolio of an oncology drug program from information available at the decision date. To support this, we construct a temporal dataset that combines 31.7k heterogeneous public data records, including trial registries, regulatory reviews, sponsor filings, utilization data, and epidemiology, into 881 offline decision episodes across 45 historical programs. We compare four offline objectives: behavioral cloning, reward-weighted behavioral cloning, learned-reward training, and value-based implicit Q-learning against four frontier LLM agents that share a common date-gated retrieval scaffold across held-out drug, sponsor, drug-class, and temporal splits. Models trained offline outperform the non-fine-tuned baselines, particularly in the post-August 2025 contamination-clean holdout. Reward-weighted behavioral cloning performs the best, obtaining 46.2% indication F1 and 14.2% strict F1 against 25.0% and 2.1%, respectively, for the best-performing tool agent on each metric. These results suggest that structured offline learning can teach agents to plan clinical experiments.
Chinese Translation
临床开发是在不确定性下的序列决策过程,赞助商必须从异质证据中规划一系列实验。我们通过将肿瘤学临床开发框架化为一个离线决策问题来研究这一设置,其中代理根据决策日期可用的信息预测肿瘤药物项目的下一个六个月试验组合。为此,我们构建了一个时间序列数据集,该数据集将31.7k个异质公共数据记录(包括试验注册、监管审查、赞助商备案、利用数据和流行病学)整合为45个历史项目中的881个离线决策情节。我们比较了四种离线目标:行为克隆、奖励加权行为克隆、学习奖励训练和基于价值的隐式Q学习,并与四个前沿的LLM代理进行比较,这些代理在持出药物、赞助商、药物类别和时间分割中共享一个共同的日期门控检索框架。离线训练的模型在后2025年8月的污染清理持出中表现优于未微调的基线,尤其显著。奖励加权行为克隆表现最佳,在指标上获得了46.2%的指示F1和14.2%的严格F1,而最佳工具代理在每个指标上的表现分别为25.0%和2.1%。这些结果表明,结构化的离线学习可以教会代理规划临床实验。
cs.AI / 78 / 2608.03609
Formal Verification of Agentic Systems over Operational Data
基于操作数据的代理系统的形式验证
Abstract
Agentic systems driven by large language models (LLMs) are increasingly deployed in real-world workflows where they act on persistent operational data. Before deployment, these systems need to be verified against business requirements that govern workflow execution and data evolution. However, existing approaches do not provide such system-level guarantees, as they mainly constrain or analyse behaviour at the agent's interface level. We study here the verification of agentic systems comprising a single LLM and a tool orchestration harness over relational operational data. We formalise them as Stateful Tool-Enabled Agentic Deployments (STEADs), give their semantics, define the problem of verifying them against First-Order Computation Tree Logic (FO-CTL) specifications, and show that it is undecidable. We identify sufficient conditions for exact preservation of FO-CTL specifications under a finite-domain restriction, over which verification is PSPACE-complete. The key requirement is that renaming opaque identifiers in the data must correspondingly rename the selected tool calls. We show that LLM-driven agents can violate this condition and introduce a canonical deployment wrapper that guarantees it for arbitrary base agents while preserving already-equivariant behaviour. We prove that computing canonical representations required by this construction is graph-isomorphism-hard. Finally, we illustrate our framework on an LLM agent orchestrating a case-management workflow.
Chinese Translation
由大型语言模型(LLMs)驱动的代理系统越来越多地被部署在实际工作流程中,在这些工作流程中,它们对持久的操作数据进行处理。在部署之前,这些系统需要根据管理工作流程执行和数据演变的业务需求进行验证。然而,现有的方法并未提供这种系统级的保证,因为它们主要限制或分析代理接口层面的行为。我们在此研究由单个LLM和工具协调框架组成的代理系统在关系操作数据上的验证。我们将其形式化为状态工具启用的代理部署(Stateful Tool-Enabled Agentic Deployments, STEADs),给出其语义,定义了验证它们是否符合一阶计算树逻辑(First-Order Computation Tree Logic, FO-CTL)规范的问题,并证明该问题是不可判定的。我们确定了在有限域限制下精确保留FO-CTL规范的充分条件,在此条件下,验证是PSPACE完全的。关键要求是数据中不透明标识符的重命名必须相应地重命名所选工具调用。我们展示了由LLM驱动的代理可能违反这一条件,并引入了一个典范部署包装器,该包装器在保持已等变行为的同时,确保对任意基础代理的保证。我们证明了构建所需的典范表示的计算是图同构困难的。最后,我们在一个LLM代理协调案例管理工作流程的例子中说明了我们的框架。
cs.AI / 79 / 2608.03611
Rethinking Modality Reliability in Multimodal Sentiment Analysis with Incomplete Observations
重新思考多模态情感分析中不完整观察的模态可靠性
Abstract
Multimodal Sentiment Analysis (MSA) integrates text, audio, and vision to infer human affect, yet real-world multimodal observations are often incomplete. Existing methods for incomplete-observation MSA mainly follow two paradigms. Reconstruction-based methods recover missing information from observed modalities, while joint-representation methods learn directly from incomplete inputs. Although effective, these methods usually treat modality reliability only implicitly within representation learning or fusion design rather than modeling it explicitly. We argue that modality reliability is a central variable in incomplete-observation settings. Failure to model it explicitly gives rise to two related issues. The first is reliability mismatch, in which the affective evidence retained by each modality varies across samples and missing rates. The second is reliability propagation bias, in which messages from degraded modalities may adversely affect cross-modal interaction and predictive performance. To address these issues, we propose MRCF, a Modality Reliability-Calibrated Framework for MSA with incomplete observations. MRCF contains a Reliability-Aware Branch that estimates sample-specific modality reliability from intramodal quality cues and cross-modal semantic consistency, a Reliability-Guided Interaction Branch that uses the estimated scores to modulate cross-modal information flow, and a Reliability-Calibrated Fusion Module that integrates reliability and semantic cues for final prediction. Experiments on CMU-MOSI, CMU-MOSEI, and CH-SIMS show that MRCF achieves strong performance under standard incomplete-observation protocols. Further analyses provide evidence that explicit reliability modeling helps mitigate reliability mismatch and reliability propagation bias during interaction and fusion.
Chinese Translation
多模态情感分析(Multimodal Sentiment Analysis, MSA)结合文本、音频和视觉来推断人类情感,但现实世界中的多模态观察往往是不完整的。现有的不完整观察 MSA 方法主要遵循两种范式:基于重建的方法从观察到的模态中恢复缺失信息,而联合表示的方法则直接从不完整输入中学习。尽管这些方法有效,但通常仅在表示学习或融合设计中隐含地处理模态可靠性,而不是显式建模。我们认为,在不完整观察的情境中,模态可靠性是一个核心变量。未能显式建模会导致两个相关问题。第一个是可靠性不匹配,即每个模态保留的情感证据在样本和缺失率之间有所不同。第二个是可靠性传播偏差,即来自退化模态的信息可能会对跨模态交互和预测性能产生不利影响。为了解决这些问题,我们提出了 MRCF(模态可靠性校准框架),用于处理不完整观察的 MSA。MRCF 包含一个可靠性感知分支,该分支从模态内部质量线索和跨模态语义一致性中估计样本特定的模态可靠性;一个可靠性引导交互分支,该分支利用估计得分来调节跨模态信息流;以及一个可靠性校准融合模块,该模块整合可靠性和语义线索以进行最终预测。在 CMU-MOSI、CMU-MOSEI 和 CH-SIMS 上的实验表明,MRCF 在标准的不完整观察协议下表现出色。进一步分析提供了证据,表明显式的可靠性建模有助于减轻交互和融合过程中的可靠性不匹配和可靠性传播偏差。
cs.AI / 80 / 2608.03627
Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection
不平等的裁决:调查基于大语言模型的假新闻检测中的性别偏见
Abstract
Large Language Models (LLMs) are increasingly used for automated fact-checking, yet their susceptibility to gender bias in this context remains underexplored. This study presents the first systematic investigation of gender bias in LLM-based fake news detection using real-world data. We augment the LIAR benchmark with three gender variants of speaker job titles (Neutral, Male, Female) for each statement to test whether veracity judgments vary solely based on gender presentation. Six state-of-the-art LLMs are evaluated across multiple bias and fairness metrics. All models exhibit gender sensitivity: 9.79%-35.13% of statements receive inconsistent labels across the three variants, with Male-Female comparisons showing 6.5%-23.6% flip rates. Two primary bias manifestations are identified: instability (inconsistent judgments) and directionality (systematic favoritism). Five models show statistically significant directional effects, with the strongest effects displaying male-skeptic patterns. These findings demonstrate that gender bias undermines both reliability and fairness in LLM-based fake news detection, highlighting the need for bias-aware evaluation and mitigation strategies. The augmented dataset is publicly released to support future research.
Chinese Translation
大语言模型(LLMs)在自动化事实核查中越来越多地被使用,但它们在这一背景下对性别偏见的敏感性仍然未得到充分探讨。本研究首次系统性地调查了基于LLM的假新闻检测中的性别偏见,使用了真实世界的数据。我们在LIAR基准数据集上增加了三种性别变体的发言人职业称谓(中性、男性、女性),以测试真实性判断是否仅基于性别表现而有所不同。我们评估了六种最先进的LLM在多个偏见和公平性指标上的表现。所有模型都表现出性别敏感性:9.79%-35.13%的陈述在三种变体中获得不一致的标签,男性与女性的比较显示6.5%-23.6%的翻转率。我们识别出两种主要的偏见表现:不稳定性(不一致的判断)和方向性(系统性偏袒)。五个模型显示出统计显著的方向性效应,最强的效应表现出男性怀疑的模式。这些发现表明,性别偏见削弱了基于LLM的假新闻检测的可靠性和公平性,强调了对偏见的评估和缓解策略的必要性。增强的数据集已公开发布,以支持未来的研究。
cs.AI / 81 / 2608.03629
Cross-Layer Interaction under Weight-Space Ablation: A Closed-Form Attention Jacobian Bound and a Test on a Real Pretrained Model
权重空间消融下的跨层交互:封闭形式的注意力雅可比界限及其在真实预训练模型上的测试
Abstract
A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computation is carried additively through a residual stream. For the one composition in that model where two carriers are architecturally dependent, an attention head and its own layer's normalization-MLP composition, it derives an exact first-order interaction formula, zero when only the MLP is ablated and second-order bounded when the head is also ablated. That result is confined to a single residual block and checked only on small transformers on a synthetic task. This paper extends the result past both limits. First, the interaction from ablating carriers spanning several layers decomposes exactly into same-block terms, one per touched layer, plus a cross-layer remainder on which the decomposition makes no claim of smallness. Second, we isolate that remainder exactly, for two layers, as a double integral of a mixed second derivative, and name the missing ingredient needed to bound it: a Jacobian bound for the attention sub-block. We derive this bound in closed form and verify it, without a single violation, against Qwen2.5-1.5B-Instruct's real weights, though we do not yet chain it across layers. We also give, in closed form, the curvature constant the companion paper's bound leaves unexhibited. Third, on that same model, we search for and find an emergent circuit for indirect object identification, never designed into it, using the original activation-patching method for this task, and test collapse, dissociation, and interaction on it. The result is mixed: a shared carrier emerges across all five tested instances, collapse and dissociation hold on most but not all, and a nonzero interaction is measurable on three of five, at layer pairs outside the same-block case the companion theorem covers.
Chinese Translation
一篇相关论文研究了在一个理想化模型中,当激活补丁和权重空间消融一致时的情况,该模型通过残差流以加法方式进行条件计算。在该模型中,两个载体在架构上相互依赖的组合中,即一个注意力头及其自身层的归一化-多层感知机(MLP)组合,推导出一个精确的一阶交互公式:当仅消融MLP时为零,而当头部也被消融时则为二阶有界。该结果仅限于单个残差块,并且仅在合成任务上的小型变换器上进行了验证。本文将结果扩展到两个限制之外。首先,消融跨越多个层的载体的交互精确分解为同块项,每个触及的层一个,加上一个跨层余项,关于该余项的分解未声称其小。其次,我们精确隔离该余项,对于两个层,作为混合二阶导数的双重积分,并命名出需要界定它的缺失成分:注意力子块的雅可比界限。我们以封闭形式推导出该界限,并在Qwen2.5-1.5B-Instruct的真实权重上验证,未出现单一违规,尽管我们尚未在层间进行链式处理。我们还以封闭形式给出了相关论文界限未展示的曲率常数。第三,在同一模型上,我们搜索并发现了一个用于间接对象识别的新兴电路,该电路并未被设计进模型中,使用原始的激活补丁方法进行此任务,并对其进行崩溃、解离和交互测试。结果是混合的:在所有五个测试实例中出现了一个共享载体,大多数但不是全部情况下崩溃和解离成立,并且在五个中的三个层对上可测量到非零交互,位于相关定理所覆盖的同块案例之外。
cs.AI / 82 / 2608.03632
When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
当教师误导时:具有虚假信号意识的在线策略蒸馏
Abstract
On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable. However, the assumptions overlook a fundamental failure mode of language models: their token-level judgments can be driven by input-agnostic language priors, formatting conventions, or stereotyped reasoning templates rather than task-specific evidence. We refer to such optimization-relevant but weakly input-grounded supervision as spurious signals in OPD, which may produce large gradients while contributing little task-improving direction. To mitigate this issue, we propose SA-OPD, a Spurious-Signal-Aware On-Policy Distillation framework that identifies and filters misleading token-level supervision based on input-groundedness and optimization impact. SA-OPD introduces a lightweight input-groundedness proxy estimating whether a token-level distillation signal truly depends on the input. It then filters only tokens that simultaneously exhibit low input-groundedness and extreme distillation divergence, thereby removing high-impact spurious updates and achieving fine-grained OPD optimization. Extensive experiments on both large language model (LLM) and vision-language model (VLM) settings demonstrate that SA-OPD consistently outperforms Vanilla OPD and competitive selective methods. These results establish input-groundedness as a key dimension for OPD supervision selection and offer a simple, effective strategy for mitigating spurious updates.
Chinese Translation
在线策略蒸馏(On-Policy Distillation, OPD)通过用密集的令牌级教师信号监督学生采样的轨迹来转移教师的能力。最近的选择性 OPD 方法通过优先考虑自信、信息丰富或可学习的信号来改善这一过程。然而,这些假设忽视了语言模型的一个基本失效模式:它们的令牌级判断可能受到与输入无关的语言先验、格式约定或刻板推理模板的驱动,而不是特定任务的证据。我们将这种与优化相关但弱输入基础的监督称为 OPD 中的虚假信号,它可能产生大的梯度,同时对任务改进方向贡献甚微。为了解决这个问题,我们提出了 SA-OPD,一个具有虚假信号意识的在线策略蒸馏框架,该框架基于输入基础性和优化影响来识别和过滤误导性的令牌级监督。SA-OPD 引入了一种轻量级的输入基础性代理,用于估计令牌级蒸馏信号是否真正依赖于输入。然后,它仅过滤同时表现出低输入基础性和极端蒸馏发散的令牌,从而去除高影响的虚假更新,实现细粒度的 OPD 优化。在大型语言模型(LLM)和视觉语言模型(VLM)设置上的广泛实验表明,SA-OPD 一致优于传统 OPD 和竞争性的选择性方法。这些结果确立了输入基础性作为 OPD 监督选择的关键维度,并提供了一种简单有效的策略来减轻虚假更新。
cs.AI / 83 / 2608.03644
Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details
跨种子交叉游戏是否足够?评估零样本协调算法对实现细节的鲁棒性
Abstract
AI agents deployed in real-world settings must be capable of coordinating with humans and other AI agents they have not encountered before. Zero-shot coordination (ZSC) algorithms aim to achieve this by specifying high-level learning rules such that independently engineered agents can coordinate with each other at test time. Rigorous evaluation of ZSC algorithms remains difficult: ideally, multiple independent implementations of each proposed algorithm must be used, reflecting the variation that arises when independent parties interpret and implement the same specification. In practice, however, ZSC algorithms have almost exclusively been evaluated using a single implementation trained across different random seeds, with only a handful of works additionally varying the neural network architecture. This leaves open questions about robustness to specification ambiguities and implementation details. In this work, we provide the first systematic evaluation of this robustness. We introduce a new evaluation scheme, cross-implementation cross-play, varying implementation details that prior work has shown to affect the performance of multi-agent reinforcement learning (MARL) algorithms, and we evaluate Other-Play, a popular ZSC algorithm, with this scheme. Our findings are encouraging and suggest that, for Other-Play, the standard ZSC evaluation is, in fact, a reasonable proxy for this more thorough cross-implementation evaluation.
Chinese Translation
在现实世界环境中部署的人工智能代理必须能够与人类及其未曾遇到的其他人工智能代理进行协调。零样本协调(Zero-shot coordination, ZSC)算法旨在通过指定高层次的学习规则,使得独立设计的代理在测试时能够相互协调。然而,对ZSC算法的严格评估仍然困难:理想情况下,必须使用每个提出的算法的多个独立实现,以反映独立方在解释和实现相同规范时所产生的变异。然而,在实践中,ZSC算法几乎仅通过在不同随机种子下训练的单一实现进行评估,只有少数研究还额外变化了神经网络架构。这使得关于对规范模糊性和实现细节的鲁棒性的问题仍然悬而未决。在本研究中,我们首次系统地评估了这种鲁棒性。我们引入了一种新的评估方案——跨实现交叉游戏(cross-implementation cross-play),变更先前研究已显示会影响多智能体强化学习(Multi-agent Reinforcement Learning, MARL)算法性能的实现细节,并使用该方案评估了流行的ZSC算法Other-Play。我们的发现令人鼓舞,表明对于Other-Play而言,标准的ZSC评估实际上是对这种更全面的跨实现评估的合理代理。
cs.AI / 84 / 2608.03653
AutoSND: From Execution Evidence to Structural Policies for Automated Network Dismantling Heuristic Discovery
AutoSND:从执行证据到结构性政策的自动化网络拆解启发式发现
Abstract
Network dismantling is fundamental to analyzing the robustness and vulnerability of complex systems, yet practical heuristics must balance effectiveness and computational efficiency, and are usually designed manually by researchers. Existing large language model based automatic heuristic design methods can generate and screen candidates, yet they have difficulty further transforming candidate quality or failure states during execution into structural-level guid- ance for subsequent generation. We propose AutoSND, a three stage tree search framework for complete network dismantling pro- grams. Stage I broadly explores from simple heuristics and archives execution evidence. Stage II compiles candidate records into struc- tural policies concerning local signals, neighborhood access, and state update ranges. Stage III continues tree search conditioned on these policies and obtains the final quality prioritized and speed prioritized candidates, AutoSND-Q/S. Experiments on 12 real world networks and 3 large real world networks show that AutoSND achieves better search performance and stability and discovers more competitive and structurally interpretable network disman- tling programs. The final candidates form an interpretable structure that uses residual degree as the backbone, adjusts node order with bounded local signals, and restricts the state update range. Code is available at https://github.com/MirrorNew/AutoSND.
Chinese Translation
网络拆解是分析复杂系统的鲁棒性和脆弱性的基础,然而实用的启发式方法必须在有效性和计算效率之间取得平衡,通常由研究人员手动设计。现有基于大型语言模型的自动启发式设计方法能够生成和筛选候选项,但在将候选质量或执行过程中的失败状态进一步转化为后续生成的结构级指导方面存在困难。我们提出了AutoSND,一个用于完整网络拆解程序的三阶段树搜索框架。第一阶段广泛探索简单启发式并归档执行证据。第二阶段将候选记录编译成关于局部信号、邻域访问和状态更新范围的结构性政策。第三阶段在这些政策的条件下继续树搜索,并获得最终的质量优先和速度优先候选项,AutoSND-Q/S。对12个真实世界网络和3个大型真实世界网络的实验表明,AutoSND在搜索性能和稳定性方面表现更佳,并发现了更具竞争力和结构可解释的网络拆解程序。最终候选项形成一个可解释的结构,以剩余度作为支柱,结合有界局部信号调整节点顺序,并限制状态更新范围。代码可在 https://github.com/MirrorNew/AutoSND 获取。
cs.AI / 85 / 2608.03660
Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Multimodal Post-Training
驯服隐性:双通道风险感知强化微调用于持续多模态后训练
Abstract
Reinforcement fine-tuning (RFT) is widely believed to inherently resist catastrophic forgetting in continual post-training of multimodal large language models. Under pronounced task distributional shifts, however, forgetting across representative RFT algorithms escalates sharply. This stems from the implicit reward-variance regularization inherent to RFT, which proves incapable of suppressing uncontrolled optimization risk. We propose Risk-Aware Policy Optimization (RAPO), the first dual-channel framework for explicit risk governance in continual RFT. On the policy channel, Risk-Aware Policy Scaling adaptively calibrates per-sample update magnitude via rollout reliability and Fisher-inspired local predictive sensitivity; on the data channel, Risk-Aware Dynamic Bucket Sampling reorganizes training batches through dynamic risk stratification, steering optimization toward informative yet stable samples. As a plug-and-play strategy requiring no cross-task memory, RAPO generalizes to any RFT algorithm without modification. On the public MLLM-CL benchmark, RAPO reduces final forgetting by 79.8% relative to its RLOO backbone while retaining new-task competitiveness.
Chinese Translation
强化微调(Reinforcement Fine-Tuning, RFT)被广泛认为在多模态大型语言模型的持续后训练中本质上能够抵抗灾难性遗忘。然而,在明显的任务分布变化下,代表性RFT算法的遗忘现象急剧上升。这源于RFT固有的隐性奖励方差正则化,其无法抑制失控的优化风险。我们提出了风险感知策略优化(Risk-Aware Policy Optimization, RAPO),这是第一个用于持续RFT的显性风险治理的双通道框架。在策略通道中,风险感知策略缩放通过回滚可靠性和受Fisher启发的局部预测敏感性自适应地校准每个样本的更新幅度;在数据通道中,风险感知动态桶采样通过动态风险分层重新组织训练批次,引导优化朝向信息丰富但稳定的样本。作为一种即插即用的策略,RAPO无需跨任务记忆,能够在不修改的情况下推广到任何RFT算法。在公共的MLLM-CL基准上,RAPO相较于其RLOO基础模型减少了79.8%的最终遗忘,同时保持了新任务的竞争力。
cs.AI / 86 / 2608.03662
Shielding for Higher-Order Safety
高阶安全保护
Abstract
Safety shields are runtime enforcement mechanisms that restrict the actions of a controller to guarantee safety. Classical shields are usually synthesised for state predicates: the current physical state is either safe or unsafe, and the shield disables precisely those actions that can force the system into an unsafe state in the future. In many cyber-physical applications this view is too coarse. A vehicle approaching an obstacle should not only avoid collision, but also respect speed regulations, force limits induced by acceleration, and jerk limits to prevent injuries. From a physical perspective, these requirements are predicated over the derivatives of the state. This paper develops a finite-state safety-game construction for such high-order smoothness constraints. We define differential safety properties using finite differences over a discretised state space, characterise their expressiveness, and reduce shield synthesis to an ordinary safety game over a history state space. We give a synthesis algorithm whose shields store exactly $k$ past states for properties of order $k$ and prove that this memory is necessary. We describe an iterative synthesis procedure for a maximally permissive shield that operates over hierarchies of derivative constraints. The algorithm solves constraints iteratively in increasing order and uses the solution at each iteration to prune the state space for the next constraint. This makes shield synthesis more efficient in practice, as the algorithm refrains from exploring large regions of the state space that are known to be unsafe.
Chinese Translation
安全保护是运行时强制机制,它限制控制器的行为以确保安全。经典的安全保护通常是针对状态谓词合成的:当前的物理状态要么是安全的,要么是危险的,保护机制精确地禁用那些可能在未来将系统强制到危险状态的行为。在许多网络物理应用中,这种观点过于粗糙。接近障碍物的车辆不仅应该避免碰撞,还应遵守速度限制、加速度引起的力限制以及防止伤害的冲击限制。从物理角度来看,这些要求是基于状态的导数。本文开发了一种有限状态安全博弈构造,以满足此类高阶平滑性约束。我们使用离散状态空间上的有限差分定义微分安全属性,表征其表达能力,并将保护机制的合成简化为一个普通的安全博弈,涉及历史状态空间。我们给出了一个合成算法,该算法的保护机制恰好存储$k$个过去状态,以满足阶数为$k$的属性,并证明了该内存是必要的。我们描述了一种迭代合成过程,用于在导数约束的层次结构上操作的最大宽容保护机制。该算法以递增的顺序迭代求解约束,并在每次迭代中使用解决方案来修剪下一个约束的状态空间。这使得保护机制的合成在实践中更加高效,因为该算法避免探索已知不安全的大区域状态空间。
cs.AI / 87 / 2608.03682
PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud
PhyAI:边缘实时物理人工智能,云端可扩展部署
Wang, Chenghua, Xu, Daliang, Cai, Dongqi, Sun, Duojin, Zhang, Hao, Qian, Haoze, Zhang, Huaiyuan, Cui, Jinshuo, Zhao, Kezhao, Gao, Longxi, Xu, Mengwei, Yi, Rongjie, Zhang, Tianyue, Xie, Weikai, Tan, Xiyuan, Liu, Xuanzhe, Qin, Yingying, Lu, Yiwen, Yao, Yuan, Zu, Yuezhi, Guo, Yunhan, Guo, Ziqi
Abstract
Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, they often rely on separate inference programs. To unify them, we build PhyAI, a Physical AI inference engine with a single runtime that keeps architecture-specific conditioning, solver, cache, and output logic in model adapters while sharing graph execution, kernels, memory management, and parallel services. The same codebase runs vision-language-action (VLA) models and world-action models (WAMs) on single or multiple GPUs across onboard, edge, and cloud deployments. We used the adapter interface to add MiniCPM-Robot on the day of its release. PhyAI achieves 1.40x-4.65x speedups over the official implementations of pi0, pi0.5, GR00T N1.7, and MiniCPM-Robot. On Cosmos3-Nano-Policy-DROID it reduces latency from 2.46 to 1.18 s on eight H20 GPUs (CFG=2, TP=4), a 2.08x speedup. Specialized runtimes remain faster in several configurations, so our goal is one runtime with competitive latency rather than the fastest result in every case. Detailed profiles reveal why different models need different execution policies: on a Hopper-series GPU at batch size one, the pi0.5 action expert accounts for 8.8% of FLOPs but 57.2% of latency; at batch size 32 its share drops to 13.5% and throughput reaches about 100 samples/s. Cosmos3 remains generation-dominated and gains only 14.3% throughput as batch size increases from 1 to 16. We further introduce the control-time Roofline, which distinguishes inference-bound from environment-bound control; the measured pi0.5 points on four LIBERO suites are environment-bound while Cosmos3 stays inference-bound. Code and benchmarks: https://github.com/mingti-org/phyai.
Chinese Translation
物理人工智能策略在其生命周期内需要进行推理,包括模型评估、云端强化学习部署、边缘GPU服务和机载部署。尽管这些设置共享相同的检查点和动作语义,但它们通常依赖于单独的推理程序。为了统一这些设置,我们构建了PhyAI,一个物理人工智能推理引擎,具有单一运行时,能够在模型适配器中保持架构特定的条件、求解器、缓存和输出逻辑,同时共享图执行、内核、内存管理和并行服务。相同的代码库可以在机载、边缘和云端部署中,在单个或多个GPU上运行视觉-语言-动作(VLA)模型和世界-动作模型(WAMs)。我们在MiniCPM-Robot发布当天使用适配器接口进行了添加。PhyAI在pi0、pi0.5、GR00T N1.7和MiniCPM-Robot的官方实现上实现了1.40x-4.65x的加速。在Cosmos3-Nano-Policy-DROID上,它将延迟从2.46秒降低到1.18秒,在八个H20 GPU(CFG=2,TP=4)上实现了2.08x的加速。尽管在多种配置中,专用运行时仍然更快,因此我们的目标是实现一个具有竞争性延迟的运行时,而不是在每种情况下都追求最快的结果。详细的分析揭示了不同模型为何需要不同的执行策略:在Hopper系列GPU上,批量大小为1时,pi0.5动作专家占FLOPs的8.8%,但占延迟的57.2%;在批量大小为32时,其占比降至13.5%,吞吐量达到约100样本/秒。Cosmos3仍然以生成为主导,随着批量大小从1增加到16,吞吐量仅增加14.3%。我们进一步引入了控制时间Roofline,它区分了推理绑定和环境绑定的控制;在四个LIBERO套件上测得的pi0.5点是环境绑定的,而Cosmos3则保持推理绑定。代码和基准测试: https://github.com/mingti-org/phyai.
cs.AI / 88 / 2608.03689
LiveEvalBench: Toward Open-World Evaluation for Web Generation
LiveEvalBench:面向开放世界的网页生成评估
Abstract
Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a static evaluation problem. We argue that frontend artifacts demand a different paradigm: they are interactive rather than static, admit diverse yet equally valid implementations, and evolve faster than rigid pipelines can accommodate. To address these gaps, we present LiveEvalBench, an automated framework that reformulates web-generation evaluation as an agentic, adaptive, and extensible process. LiveEvalBench instantiates evaluation as a collaborative review workflow, in which a Build Engineer, a Code Engineer, and a UI Tester collectively gather evidence across the full lifecycle of a frontend project, from deployment and code inspection to browser-based interaction. To handle implementation diversity, an adaptive protocol couples shared rubrics for cross-model comparability with implementation-grounded criteria tailored to each artifact. The framework further supports incremental integration of new evaluator roles and assessment dimensions without pipeline redesign. Experiments across diverse real-world web-generation scenarios show that LiveEvalBench aligns closely with human expert judgment and provides fine-grained insights into frontier models' web generation capabilities. Code is available at https://github.com/wyysteelhead/LiveEvalBench
Chinese Translation
大型语言模型在合成可执行的前端项目方面的能力日益增强,然而现有基准仍将网页生成视为一个静态评估问题。我们认为,前端工件需要一种不同的范式:它们是交互式的而非静态的,承认多样且同样有效的实现,并且发展速度快于僵化的流程所能适应。为了解决这些问题,我们提出了LiveEvalBench,一个将网页生成评估重新构建为一个自主、适应性强且可扩展的过程的自动化框架。LiveEvalBench将评估实例化为一个协作审查工作流程,在该流程中,构建工程师、代码工程师和用户界面测试人员共同收集证据,涵盖前端项目的整个生命周期,从部署和代码检查到基于浏览器的交互。为了处理实现的多样性,一个适应性协议将用于跨模型可比性的共享评分标准与针对每个工件量身定制的基于实现的标准相结合。该框架进一步支持新评估角色和评估维度的增量集成,而无需重新设计流程。在多种真实世界的网页生成场景中的实验表明,LiveEvalBench与人类专家的判断高度一致,并提供了对前沿模型网页生成能力的细致洞察。代码可在 https://github.com/wyysteelhead/LiveEvalBench 获取。
cs.AI / 89 / 2608.03699
TARL: Transaction-Aware Reliable Ledgers for Executable Memory Management in Long-Term Agents
TARL:面向事务的可靠分类账在长期智能体中的可执行内存管理
Abstract
Persistent memory helps long-term agents retain knowledge, yet a single update error can repeatedly distort future retrieval and reasoning. Most existing systems reduce memory updating to a binary Write/Hold decision, which cannot distinguish whether new information should be added, ignored, used to revise an outdated belief, rejected as unreliable, or deferred for verification. These choices may share the same binary label while producing fundamentally different memory states. We introduce TARL, a memory state update framework that maps each statement to one of five executable actions. TARL identifies the affected memory, resolves its temporal scope, compares source reliability, and updates accepted, pending, and rejected ledgers. It is further trained by comparing the memory states produced by alternative update operations, encouraging the model to select the operation that leads to the correct result. We also introduce TARL-Mem, a benchmark with fine-grained action labels and next-state targets. Across in-domain, cross-source, temporal, counterfactual, and sequential evaluations, TARL improves action prediction and state recovery, reduces memory pollution, preserves conflicting evidence, and limits cumulative corruption. The complete model implementation is provided in the supplementary material.
Chinese Translation
持久性内存帮助长期智能体保留知识,但单次更新错误可能会反复扭曲未来的检索和推理。现有大多数系统将内存更新简化为二元的写入/保持决策,无法区分新信息是应被添加、忽略、用于修正过时的信念、被拒绝为不可靠,还是被推迟以待验证。这些选择可能共享相同的二元标签,但会产生根本不同的内存状态。我们提出了TARL,一个内存状态更新框架,将每个陈述映射到五种可执行操作之一。TARL识别受影响的内存,解决其时间范围,比较来源的可靠性,并更新接受、待处理和拒绝的分类账。通过比较由替代更新操作产生的内存状态进一步训练模型,鼓励模型选择导致正确结果的操作。我们还介绍了TARL-Mem,一个具有细粒度操作标签和下一个状态目标的基准。在领域内、跨源、时间、反事实和顺序评估中,TARL提高了操作预测和状态恢复,减少了内存污染,保留了冲突证据,并限制了累积腐败。完整的模型实现已在补充材料中提供。
cs.AI / 90 / 2608.03705
Less Traffic, Better Outcomes: Competition-Aware Request Dispatch in Real-Time Ad Exchanges
减少流量,改善结果:实时广告交换中的竞争感知请求调度
Abstract
Real-time bidding (RTB) ad exchanges typically forward nearly all incoming requests to demand-side platforms (DSPs), even though only a small fraction receive bids. This over-distribution weakens auction outcomes: DSPs throttle participation under compute and budget constraints, reducing the effective use of limited bidding capacity. We present a competition-aware request dispatch framework that uses distributional bid prediction and probabilistic forwarding to decide whether each request should be sent to each DSP. The system adapts per-DSP thresholds over time through lightweight policy optimization to track non-stationary market conditions. We evaluate the framework through four sequential online experiments on a production platform serving over 20 billion daily requests. A full multi-DSP deployment reduces DSP request volume under the policy by 34.2% while increasing net revenue by 4.6% (p<0.001) in a recent 14-day window after an initial DSP adaptation period. Further analysis highlights strong heterogeneity across traffic segments and reveals that aggregate metrics can be misleading. Segment-level and per-DSP analyses suggest that the policy surfaces comparative advantages among DSPs, improving monetized outcomes without increasing overall request volume.
Chinese Translation
实时竞价(RTB)广告交换通常将几乎所有的请求转发给需求方平台(DSP),尽管只有一小部分请求会收到竞标。这种过度分配削弱了拍卖结果:DSP在计算和预算限制下限制参与,降低了有限竞标能力的有效利用。我们提出了一种竞争感知请求调度框架,该框架利用分布式竞标预测和概率转发来决定每个请求是否应发送给每个DSP。该系统通过轻量级策略优化,随着时间的推移调整每个DSP的阈值,以跟踪非平稳的市场条件。我们通过在一个每天处理超过200亿请求的生产平台上进行的四个连续在线实验来评估该框架。全面的多DSP部署在政策下减少了DSP请求量34.2%,同时在最近的14天窗口中,净收入增加了4.6%(p<0.001),这一结果是在初始DSP适应期之后获得的。进一步分析强调了流量段之间的强异质性,并揭示了聚合指标可能具有误导性。分段和每个DSP的分析表明,该政策在DSP之间显现出比较优势,改善了货币化结果而不增加整体请求量。
cs.AI / 91 / 2608.03722
When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Coupling Diagnostic for Machine Collectives
当输出分散时,知识修正是否随之而来?一种针对机器集体的黑箱耦合诊断
Abstract
Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy can break: agents can produce diverse-looking arguments while preserving the same conclusion. We operationalize dispersion-revision coupling: the degree to which an intervention that verifiably increases the dispersion of a collective's outputs in embedding space is accompanied by genuine revision of its epistemic stance rather than premise-preserving reformulation. The diagnostic is black-box: it operates on generated text alone and makes no claims about the internal representations of the generating models. Two channels are measured independently: an output channel, the Coherence Index (CI), verifies that the intervention changed output dispersion; an epistemic channel, per-turn stance annotation, measures whether the collective revised. We propose CI with the Meta-Predictive Clarity System (MPCS), which inserts a Re-Differentiation Protocol (RDP) when outputs over-converge, as a reusable method for estimating this coupling regime. We evaluate five-agent collectives from two configurations (gpt-4o-mini and gemini-2.5-flash; 310 paired episodes per condition). On gpt-4o-mini, conditional dissent improves false-premise recovery by +17.7 points (p<1e-6) while static persona diversity harms recovery (-8.1, p=.007). On gemini-2.5-flash, the same intervention at a comparable budget yields no gain (26.1% vs 27.1%, p=.84) despite a verified dispersion drop; the two treatment effects differ from each other (z=3.79, p<.001). Mechanism tagging shows Gemini preserves the false premise via intra-framework dissent: 94% of tagged post-RDP responses reformulate rather than concede (vs 24% on GPT). We recommend reporting per-intervention stance shift and premise-preservation rate alongside accuracy.
Chinese Translation
集体智能研究将分歧视为知识多样性的证据:如果代理人表达不同的观点,群体应保留修正能力。在大型语言模型(LLM)集体中,这一代理可能会失效:代理人可以产生看似多样的论点,同时保持相同的结论。我们将分散-修正耦合进行操作化:即在嵌入空间中,能够可验证地增加集体输出分散度的干预措施,是否伴随着其知识立场的真正修正,而不是前提保持不变的重新表述。该诊断是黑箱的:它仅基于生成的文本操作,并不对生成模型的内部表征做出任何声明。我们独立测量两个通道:输出通道,即一致性指数(Coherence Index, CI),验证干预是否改变了输出分散度;知识通道,即每轮的立场注释,测量集体是否进行了修正。我们提出使用元预测清晰系统(Meta-Predictive Clarity System, MPCS)的CI,当输出过度收敛时插入重新区分协议(Re-Differentiation Protocol, RDP),作为估计这一耦合机制的可重用方法。我们评估了来自两种配置(gpt-4o-mini和gemini-2.5-flash)的五代理人集体(每种条件310对事件)。在gpt-4o-mini上,有条件的异议使得错误前提的恢复提高了17.7个百分点(p<1e-6),而静态角色多样性则降低了恢复效果(-8.1,p=.007)。在gemini-2.5-flash上,相同的干预在可比预算下没有获得收益(26.1%对27.1%,p=.84),尽管验证了分散度的下降;这两种处理效果彼此不同(z=3.79,p<.001)。机制标记显示,Gemini通过框架内的异议保持了错误前提:94%的标记后RDP响应进行了重新表述,而不是让步(相比之下,GPT为24%)。我们建议在报告准确性时,附上每次干预的立场变化和前提保持率。
cs.AI / 92 / 2608.03728
SAT-Edge-Agent: Hardware-in-the-Loop Edge-Agent Orchestration for Onboard Satellite Intelligence
SAT-Edge-Agent:用于卫星智能的硬件在环边缘代理编排
Abstract
Onboard satellite intelligence requires a task layer that translates mission intent into local tool calls, exposes execution state, and returns machine-consumable artifacts under communication and power constraints. We present SAT-Edge-Agent, a hardware-in-the-loop (HIL) edge-agent system deployed on a commercial off-the-shelf ARM-based heterogeneous edge system-on-chip. A browser workspace and FastAPI agent coordinate a local OpenAI-compatible language service with a project-internal YOLO-style oriented-object-detection endpoint that returns FAIR1M metadata-backed structured results. Two fixed FAIR1M workloads, one single-image and one serial two-image request, were repeated 20 times each and completed 20/20 attempts. Mean Full-Agent latency was 29.353 s and 60.937 s, with empirical P95 values of 31.166 s and 66.882 s. Mean detector time was 861.386 ms and 1510.920 ms, only 2.93% and 2.48% of the corresponding Full-Agent means. Profiling indicates that most visible latency occurs outside detector execution. Mean CPU utilization was 20.761% and 20.482%. A 200-ms NPU-load field averaged 100% for both workloads, but it represents a shared-accelerator software field rather than detector-only occupancy or calibrated utilization. The public evidence package provides sanitized request-level records, redacted JSON, normalized SSE examples, and scripts reproducing the reported statistics. These results establish a reproducible HIL boundary for observable satellite edge-agent orchestration, but do not establish detector accuracy, a new geolocation method, calibrated energy efficiency, or flight readiness.
Chinese Translation
在卫星智能系统中,需要一个任务层将任务意图转化为本地工具调用,暴露执行状态,并在通信和电力限制下返回机器可消费的工件。我们提出了SAT-Edge-Agent,这是一种在商业现成的基于ARM的异构边缘系统芯片上部署的硬件在环(HIL)边缘代理系统。一个浏览器工作区和FastAPI代理协调一个本地兼容OpenAI的语言服务,以及一个项目内部的YOLO风格定向物体检测端点,返回基于FAIR1M元数据的结构化结果。两个固定的FAIR1M工作负载,一个单图像请求和一个串行双图像请求,各重复20次,均完成了20/20次尝试。平均全代理延迟为29.353秒和60.937秒,经验P95值分别为31.166秒和66.882秒。平均检测器时间为861.386毫秒和1510.920毫秒,仅占相应全代理均值的2.93%和2.48%。分析表明,大部分可见延迟发生在检测器执行之外。平均CPU利用率为20.761%和20.482%。对于两个工作负载,200毫秒的NPU负载场平均为100%,但这代表的是共享加速器软件场,而非仅检测器的占用或校准利用率。公共证据包提供了经过清理的请求级记录、编辑的JSON、标准化的SSE示例以及重现报告统计的脚本。这些结果建立了一个可重复的HIL边界,用于可观察的卫星边缘代理编排,但并未建立检测器的准确性、新的地理定位方法、校准的能效或飞行准备状态。
cs.AI / 93 / 2608.03731
CARE-Bench: Benchmarking Patient-Facing LLM Triage
CARE-Bench:患者导向大语言模型分诊基准测试
Abstract
Patient-facing medical LLMs and agents increasingly answer symptom questions before clinician contact, where the key safety question is what action the user should take next. We introduce CARE-Bench, a source-grounded benchmark that evaluates sequential patient-facing triage as a four-label per-turn current-action task. CARE-Bench contains 500 cases and 1,059 evaluated patient-disclosure prefixes reconstructed from medical dialogue, consultation, and follow-up-question sources. We evaluate 11 models on 269 held-out rounds under unprompted and minimally prompted open-ended protocols, using a fixed GPT-5.5 mapper to code each response into the four-label action space. Unprompted macro-F1 remains low, ranging from 31.2 to 50.4. Prompting improves 10 of 11 models, with prompted macro-F1 ranging from 46.9 to 63.4, but substantial threshold errors remain. Prompted models often recommend care before needed clarification is obtained; when the correct action was to ask for more information, only 33.5% of prompted outputs preserved the step. The persistence of these errors after prompting suggests that patient-facing triage is not a simple prompting problem and supports explicit evaluation of action timing before deployment.
Chinese Translation
患者导向的医疗大语言模型(LLMs)和代理越来越多地在与临床医生接触之前回答症状问题,其中关键的安全问题是用户应该采取什么后续行动。我们介绍了CARE-Bench,这是一个基于来源的基准,评估顺序的患者导向分诊作为一个每轮四标签的当前行动任务。CARE-Bench包含500个案例和1,059个从医疗对话、咨询和后续问题来源重建的患者披露前缀。我们在269个保留轮次上评估了11个模型,采用无提示和最小提示的开放式协议,使用固定的GPT-5.5映射器将每个响应编码到四标签行动空间中。无提示的宏F1值保持较低,范围为31.2到50.4。提示改善了11个模型中的10个,提示后的宏F1值范围为46.9到63.4,但仍然存在显著的阈值错误。提示模型通常在需要澄清之前就推荐护理;当正确的行动是请求更多信息时,仅有33.5%的提示输出保留了该步骤。这些错误在提示后仍然存在,表明患者导向分诊并不是一个简单的提示问题,并支持在部署之前对行动时机进行明确评估。
cs.AI / 94 / 2608.03733
Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement
基于失败信息的图像自我增强用于多模态大型语言模型自我改进
Abstract
Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are costly to annotate. Self-augmentation offers a promising alternative by enabling models to expand their own training data without external supervision. However, existing MLLM self-augmentation methods are largely text-centric, while image augmentation remains underexplored and typically relies on generic or handcrafted transformations that are weakly aligned with the model's actual incapability. We propose Failure-informed Image Self-Augmentation (\textbf{FISA}), a framework for MLLM self-improvement that constructs augmented images from the model's own failure cases. Our method generates visually challenging yet answer-preserving image complications, verifies their utility through self-examination, and applies dual fidelity filtering to avoid semantic distortion. Experiments on visual question answering benchmarks show that the proposed method consistently improves performance across both in-distribution and out-of-distribution settings. Further experiments validate the compatibility of FISA with existing textual self-augmentation approaches, the superior data efficiency of the synthesized samples over generic image augmentation baselines, and the practical effectiveness of the proposed filtering strategy.
Chinese Translation
多模态大型语言模型(MLLMs)在视觉-语言任务中取得了显著的性能,但它们的进展在很大程度上依赖于大规模、高质量的多模态数据,而这些数据的标注成本高昂。自我增强提供了一种有前景的替代方案,使模型能够在没有外部监督的情况下扩展自身的训练数据。然而,现有的MLLM自我增强方法主要集中于文本,而图像增强仍然未得到充分探索,通常依赖于与模型实际能力不匹配的通用或手工设计的变换。我们提出了基于失败信息的图像自我增强(Failure-informed Image Self-Augmentation,FISA),这是一个用于MLLM自我改进的框架,通过模型自身的失败案例构建增强图像。我们的方法生成视觉上具有挑战性但保留答案的图像复杂性,通过自我检查验证其效用,并应用双重保真度过滤以避免语义失真。在视觉问答基准上的实验表明,所提出的方法在分布内和分布外设置中均能持续提高性能。进一步的实验验证了FISA与现有文本自我增强方法的兼容性、合成样本在数据效率上优于通用图像增强基线的优势,以及所提出的过滤策略的实际有效性。
cs.AI / 95 / 2608.03738
AgenticECO: An Agentic Framework for ECO on 3D Integrated Circuits
AgenticECO:用于三维集成电路的代理框架
Abstract
As Moore's law slows, the industry is turning to three-dimensional integration; yet in merged 3D-IC flows, routed designs expose bond-level defects with no 2D analogue, and post-route engineering change orders (ECO) remain manual, expertise-bound work. Worse, the standard edit-then-fully-reroute practice entangles a repair with router churn, so a signoff number cannot be attributed to the edit that motivated it. We present AgenticECO, an evidence-gated tool-using agent workflow for 3D-IC ECO on the open-source TaiWei flow, paired with EcoRoute, a minimal-disturbance ECO-routing layer that drives the unmodified pinned router so a repair is attributable to its edit. Across nine matched natural defect cases under identical budgets, AgenticECO clears seven versus two for both full reroute and stock repair, at 0.66\% mean disturbance over cleared cases and zero clock nets touched, and a cross-backbone rerun under the same sealed contract clears all nine. Controlled studies show that the repair moves are necessary under preservation, that occupancy-aware choice buys legal landings rather than repair success, and that under tightened clocks minimal disturbance flips accept versus reject. Three preregistered visual studies localize the pixel instrument's edge to contested landing sites, and a preregistered blind diagnostic exactly restores every held-out injected defect, the only arm with zero wrong edits. Every accepted result passes routing, fresh extraction, max/min timing, DRC, and structural-equivalence gates. Code, environment, and per-episode audit artifacts are released as supplementary material.
Chinese Translation
随着摩尔定律的放缓,行业正转向三维集成;然而,在合并的3D-IC流程中,路由设计暴露出无二维对应的键合级缺陷,而后路由工程变更订单(ECO)仍然是手动、依赖专业知识的工作。更糟的是,标准的编辑-然后完全重新路由的做法使得修复与路由器的变动纠缠在一起,因此无法将签署号归因于促使其产生的编辑。我们提出了AgenticECO,这是一种基于证据的工具使用代理工作流程,适用于开源的TaiWei流程中的3D-IC ECO,配合EcoRoute,一个最小干扰的ECO路由层,驱动未修改的固定路由器,使得修复可以归因于其编辑。在九个在相同预算下匹配的自然缺陷案例中,AgenticECO在完全重新路由和标准修复方面分别清除了七个与两个,以0.66%的平均干扰率和零个时钟网被触及,并且在相同密封合同下的交叉骨干重跑清除了所有九个。控制研究表明,在保持条件下,修复移动是必要的,考虑占用的选择购买合法着陆,而不是修复成功,并且在时钟收紧的情况下,最小干扰翻转接受与拒绝。三个预注册的视觉研究将像素仪器的边缘定位到争议着陆点,而一个预注册的盲诊断准确恢复了每个保留的注入缺陷,这是唯一一个没有错误编辑的臂。每个接受的结果都通过了路由、新提取、最大/最小时序、设计规则检查(DRC)和结构等价性门。代码、环境和每个回合的审计文档作为补充材料发布。
cs.AI / 96 / 2608.03740
MissClick: Exploiting Digit-Serialized Coordinates to Attack GUI Grounding Models
MissClick:利用数字序列坐标攻击GUI基础模型
Abstract
Recent GUI visual grounding models generate screen coordinates as sequences of digit tokens that are parsed into numerical values and mapped to executable clicks. The security implications of this coordinate generation process have been largely overlooked. We observe that each coordinate digit is predicted as a categorical token, yet after parsing, changing a hundreds-place digit by one changes the corresponding numerical coordinate component by 100 units, which can induce a large displacement of the executed click. This observation motivates attack objectives that account for the numerical and place-value structure of coordinate outputs rather than treating them as ordinary text. Moreover, untargeted and targeted attacks impose different success conditions--displacing the click outside the correct region versus into an attacker-specified region--and therefore benefit from different objectives. We propose MissClick, a simple and effective white-box adversarial attack with two goal-specific objectives: MissClick-U maximizes soft-coordinate displacement for untargeted disruption, while MissClick-T minimizes a place-weighted target-digit loss for targeted hijacking. Compared with existing attacks against GUI grounding models on OS-Atlas and UGround across desktop, web, and mobile platforms, MissClick-U achieves untargeted success rates of 75.07\% and 72.93\% (+16.62 and +30.72 pp), and MissClick-T achieves targeted success rates of 44.86\% and 62.67\% (+31.73 and +47.06 pp). Attack objective comparison further shows that soft-coordinate displacement yields the highest untargeted attack success rate, whereas place-weighted target-digit optimization yields the highest targeted attack success rate, revealing distinct objective preferences for the two attack goals.
Chinese Translation
近期的GUI视觉基础模型生成的屏幕坐标以数字令牌序列的形式呈现,这些令牌被解析为数值并映射到可执行的点击操作。这个坐标生成过程的安全隐患在很大程度上被忽视。我们观察到,每个坐标数字被预测为一个类别令牌,但在解析后,改变一个百位数字的值会使相应的数值坐标分量变化100个单位,这可能导致执行点击的较大位移。这一观察促使我们设定攻击目标,考虑坐标输出的数值和位值结构,而不是将其视为普通文本。此外,非定向攻击和定向攻击施加了不同的成功条件——将点击位移到正确区域之外与进入攻击者指定区域——因此受益于不同的目标。我们提出了MissClick,这是一种简单有效的白盒对抗攻击,具有两个目标特定的目标:MissClick-U最大化非定向干扰的软坐标位移,而MissClick-T最小化定向劫持的加权目标数字损失。与现有针对OS-Atlas和UGround的GUI基础模型的攻击相比,MissClick-U在桌面、网页和移动平台上实现了75.07%和72.93%的非定向成功率(分别提高了16.62和30.72个百分点),而MissClick-T实现了44.86%和62.67%的定向成功率(分别提高了31.73和47.06个百分点)。攻击目标比较进一步表明,软坐标位移产生了最高的非定向攻击成功率,而加权目标数字优化则产生了最高的定向攻击成功率,揭示了这两种攻击目标的不同偏好。
cs.AI / 97 / 2608.03744
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
代理捕捉代理:临床多代理系统中的快捷级联与基准游戏
Abstract
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore. Across seven cohorts on six public datasets spanning text (MedQA-USMLE, MedMCQA, MIMIC-CXR reports), imaging (NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert) and tabular ICU records (SUPPORT2), Gemini committees resist these cues in isolation (flip 5-16%), yet a socially plausible shortcut spreads: when two peers assert the same wrong answer, the holdout under test adopts it in 38% of cases, as does a false "pre-screen" system flag, on both capability tiers. Of three oversight agents, a gate cannot separate adoption from honest agreement (false-positive rate 100%); a same-lineage judge reading only the transcript flags adoption on text (precision 100%, recall 93%) but collapses onto the gate in imaging; a referee that privately re-queries the holdout transfers to imaging (77-88% precision, 13-21% false-positive rate). Tripling a cue's visual salience does not move contagion, whereas a second peer voice raises it by half again. Gaming a hidden rubric is near-silent: only 1/10 text and 1/134 imaging drifters name the rubric they moved toward. What games a committee is social plausibility, and only a referee independent of self-report catches it. Code: https://github.com/criticaldata/benchmaxxing
Chinese Translation
临床决策支持正朝着由语言模型代理在共享工作空间中进行审议的委员会发展。我们探讨这些委员会是否会被快捷方式操控,这些快捷方式是基准奖励的线索,但临床医生会忽视。在六个公共数据集上的七个队列中,涵盖文本(MedQA-USMLE、MedMCQA、MIMIC-CXR 报告)、影像(NIH ChestX-ray14、MIMIC-CXR-JPG、CheXpert)和表格 ICU 记录(SUPPORT2),Gemini 委员会在孤立情况下抵制这些线索(翻转率为 5-16%),然而一种社会上合理的快捷方式却传播开来:当两个同伴声称相同的错误答案时,待测试的持有者在 38% 的情况下采纳该答案,虚假的“预筛选”系统标记也同样如此,适用于两个能力层级。在三个监督代理中,一个门控无法将采纳与诚实一致区分开(假阳性率 100%);一个同源的评审仅阅读转录本时在文本上标记采纳(精准率 100%,召回率 93%),但在影像上则崩溃到门控;一个私下重新查询持有者的裁判转向影像(精准率 77-88%,假阳性率 13-21%)。将线索的视觉显著性三倍化并未推动传播,而第二个同伴的声音则将其提高了一半。操控隐藏标准几乎是无声的:只有 1/10 的文本和 1/134 的影像漂移者提到他们所趋向的标准。操控一个委员会的因素是社会合理性,只有一个独立于自我报告的裁判能够捕捉到这一点。代码链接:https://github.com/criticaldata/benchmaxxing
cs.AI / 98 / 2608.03745
Risky Business: Measuring The Faithfulness-Safety Tension
风险业务:测量忠诚度与安全性之间的张力
Abstract
Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring. However, monitoring relies on faithfulness, i.e., the model output strictly derives from its reasoning trace. We identify an alignment tension where a model must be faithful enough to be monitored, yet robust enough to reject unsafe reasoning. We demonstrate that this counterbalance exists in current Large Reasoning Models (LRMs), and show ways in which it can be addressed. We introduce HazMart, a human-written dataset set in an autonomous AI shopkeeper scenario. Unlike prior work that relies on providing hints in prompts to test faithfulness (e.g., "A Stanford professor said it should be Answer A"), we propose a novel replacement-based technique, which we call Targeted Reasoning Replacement (TRR), that directly intervenes in the reasoning chain to substitute in unsafe or illogical thoughts (e.g., "Wait, the answer must be Option B [was Option A] because it is the most fitting"). DeepSeek-R1-Llama-70B exhibits high faithfulness (97.5%) but fails to reject Unsafe Reasoning (12.3%), while QwQ-32B is more robust (73.9% safety) at the cost of lower faithfulness (74.7%). Mechanistic analyses of QwQ-32B reveal that these properties are represented by anti-correlated internal directions peaking at the action-commit token. Finally, we demonstrate that representation steering can independently amplify the safety direction, increasing safe behavior by 9 percentage points while maintaining base capabilities.
Chinese Translation
链式推理(Chain-of-Thought, CoT)为模型监控提供了一个有前景的视角。然而,监控依赖于忠诚度,即模型输出严格来源于其推理轨迹。我们识别出一种对齐张力,即模型必须足够忠诚以便进行监控,同时又必须足够稳健以拒绝不安全的推理。我们证明了这种平衡在当前的大型推理模型(Large Reasoning Models, LRM)中确实存在,并展示了可以解决这一问题的方法。我们引入了HazMart,这是一个设定在自主AI店主场景中的人类编写的数据集。与之前依赖于在提示中提供线索以测试忠诚度的工作(例如,“一位斯坦福教授说答案应该是选项A”)不同,我们提出了一种新的基于替换的技术,称为目标推理替换(Targeted Reasoning Replacement, TRR),该技术直接干预推理链,以替换不安全或不合逻辑的想法(例如,“等一下,答案必须是选项B [曾是选项A],因为它是最合适的”)。DeepSeek-R1-Llama-70B展现出高忠诚度(97.5%),但未能拒绝不安全推理(12.3%),而QwQ-32B则在安全性上更为稳健(73.9%),但忠诚度较低(74.7%)。对QwQ-32B的机制分析揭示,这些特性由在动作承诺标记处达到峰值的反相关内部方向表示。最后,我们展示了表示引导可以独立放大安全方向,在保持基本能力的同时,将安全行为提高9个百分点。
cs.AI / 99 / 2608.03764
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
GDPevo:评估智能体在真实商业任务中的自我进化
Abstract
Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provide limited coverage of economically valuable task domains, do not always design training and test tasks such that test-time gains can be attributed to training experience, and remain vulnerable to data contamination. We present GDPevo, an evolution-native benchmark grounded in GDP-related enterprise workflows, together with the fully automated data pipeline that generates it. Its core mechanism, rule hybridization, decomposes each enterprise workflow into atomic business rules, distributes subsets of these rules across training tasks, and recombines them in held-out test tasks so that test-time gains are attributable. GDPevo spans CRM, ERP, finance, healthcare, legal, and data-centric workflows. Its V1 release contains 120 tasks in 12 groups, with five training and five held-out test tasks per group. Full automation enables the pipeline to expand the suite to 240 tasks in 24 groups (V2) within two days, providing a practical response to contamination. Using GDPevo, we evaluate four agents, each comprising a harness and a model, under four supervision types. Self-evolution consistently improves held-out accuracy by up to 16.44 percentage points. But the best evolved agents remain far below the fully informed oracle ceiling of 91.6%, indicating that the self-evolution ability of current agents remains far from fully realized. We publicly release the pipeline, benchmark, and full evaluation results at https://github.com/Prism-Shadow/GDPevo.
Chinese Translation
智能体自我进化通过从以往经验中更新智能体的持久状态,并利用这些状态更有效地解决相关任务。评估自我进化是困难的:现有基准在经济价值任务领域的覆盖面有限,训练和测试任务的设计并不总是能够将测试时的收益归因于训练经验,并且仍然容易受到数据污染的影响。我们提出了GDPevo,这是一个基于GDP相关企业工作流程的进化原生基准,以及生成该基准的全自动数据管道。其核心机制——规则混合,将每个企业工作流程分解为原子商业规则,将这些规则的子集分配到训练任务中,并在保留的测试任务中重新组合,以便将测试时的收益归因。GDPevo涵盖了客户关系管理(CRM)、企业资源规划(ERP)、金融、医疗、法律和数据驱动的工作流程。其V1版本包含12组120个任务,每组包含五个训练任务和五个保留测试任务。全自动化使得该管道能够在两天内将任务套件扩展到24组240个任务(V2),为应对污染提供了实际解决方案。使用GDPevo,我们在四种监督类型下评估了四个智能体,每个智能体由一个框架和一个模型组成。自我进化始终将保留准确率提高了最多16.44个百分点。但最佳进化智能体的表现仍远低于完全知情的oracle上限91.6%,这表明当前智能体的自我进化能力尚未充分实现。我们在https://github.com/Prism-Shadow/GDPevo上公开发布了该管道、基准和完整评估结果。
cs.AI / 100 / 2608.03772
Computing Actual Causes for Neural Network Predictions under Structured Causal Inputs
在结构因果输入下计算神经网络预测的实际原因
Abstract
Explaining the predictions of neural networks is a central challenge in trustworthy AI. Existing explanation methods, such as those based on feature attribution or minimal sufficient sets, typically treat input features as independent, which can yield misleading explanations when inputs exhibit structured dependencies. We address this by formalizing explanations as Halpern-Pearl (HP) actual causes, modeling input dependencies using Boolean Structural Causal Models (SCMs). We compute HP causes by applying bound propagation and branch-and-bound techniques, while providing formal guarantees of completeness and minimality. Our experiments show that we substantially outperform brute-force and ILP baselines in scalability, and outperform heuristic search as graph size grows, computing all minimal actual causes on instances with search spaces of up to $2.3\times10^{13}$ candidate (cause, contingency) pairs, on SCMs with up to 28 nodes, within a 180s per-instance budget. In a case study, we further show that ignoring input dependencies inflates the number of reported causes, 14.9% of which are spurious under our SCM.
Chinese Translation
解释神经网络的预测是可信人工智能中的一个核心挑战。现有的解释方法,如基于特征归因或最小充分集的方法,通常将输入特征视为独立的,这在输入存在结构依赖时可能导致误导性的解释。我们通过将解释形式化为Halpern-Pearl (HP) 实际原因,并使用布尔结构因果模型 (SCMs) 来建模输入依赖关系,从而解决这一问题。我们通过应用界限传播和分支限界技术来计算HP原因,同时提供完整性和最小性的正式保证。我们的实验表明,在可扩展性方面,我们的表现显著优于暴力搜索和整数线性规划 (ILP) 基线,并且随着图形规模的增长,我们的表现优于启发式搜索,在具有高达$2.3 imes10^{13}$个候选(原因,偶然)对的搜索空间的实例上,计算所有最小实际原因,SCMs的节点数高达28个,且每个实例的预算为180秒。在一个案例研究中,我们进一步表明,忽视输入依赖关系会膨胀报告的原因数量,其中14.9%的原因在我们的SCM下是虚假的。
cs.AI / 101 / 2608.03782
KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation
KnowHal:一个知识驱动的综合多模态幻觉评估基准
Abstract
Hallucination remains a critical challenge for developing trustworthy Multimodal Large Language Models (MLLMs). While existing benchmarks mainly focus on entity, attribute, and relation hallucinations, knowledge-related failures are often investigated separately, lacking a unified evaluation framework across different hallucination dimensions. To overcome this, we propose \textbf{KnowHal}, a benchmark that explicitly incorporates knowledge hallucination into multimodal hallucination evaluation spanning four dimensions: entity, attribute, relation, and knowledge. KnowHal constructs paired positive and negative questions over shared images and entities, enabling controlled comparisons among perceptual errors, knowledge-related errors, and false-premise acceptance. The benchmark contains 1,800 samples across 10 domains and 50 categories, constructed through a semi-automated pipeline combining LLM assistance, CLIP-based filtering, and human verification. We evaluate 14 representative MLLMs on KnowHal and conduct extensive analyses. Results show that the knowledge dimension consistently presents the greatest challenge for nearly all evaluated models, while most models exhibit substantial performance degradation on negative questions, revealing limited robustness to false premises. By unifying four hallucination dimensions with paired question design, KnowHal addresses an important gap in existing evaluation frameworks and enables a more comprehensive assessment of hallucinations in MLLMs.
Chinese Translation
幻觉仍然是开发可信赖的多模态大型语言模型(MLLMs)的一项关键挑战。现有基准主要集中在实体、属性和关系幻觉上,而与知识相关的失败往往被单独研究,缺乏跨不同幻觉维度的统一评估框架。为了解决这一问题,我们提出了 extbf{KnowHal},一个明确将知识幻觉纳入多模态幻觉评估的基准,涵盖四个维度:实体、属性、关系和知识。KnowHal在共享图像和实体上构建了成对的正向和负向问题,使得感知错误、知识相关错误和错误前提接受之间的比较得以控制。该基准包含来自10个领域和50个类别的1,800个样本,通过结合大型语言模型(LLM)辅助、基于CLIP的过滤和人工验证的半自动化流程构建。我们在KnowHal上评估了14个代表性的MLLM,并进行了广泛的分析。结果显示,知识维度对几乎所有评估模型始终构成最大的挑战,而大多数模型在负向问题上的表现显著下降,揭示了对错误前提的鲁棒性有限。通过成对问题设计统一四个幻觉维度,KnowHal填补了现有评估框架中的重要空白,并使得对MLLM中幻觉的更全面评估成为可能。
cs.AI / 102 / 2608.03791
Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation
遗忘是否跨模态转移?跨模态知识遗忘评估的现实基准
Abstract
Vision-Language Models (VLMs), like Large Language Models (LLMs), may memorize sensitive, copyrighted, or harmful knowledge from their pretraining corpora. Removing such knowledge is essential for building trustworthy AI systems. However, existing studies primarily focus on forgetting within individual modalities. Although recent work has begun to explore cross-modal consistency in unlearning, the cross-modal transfer of real-world knowledge unlearning remains insufficiently studied. To address this gap, we introduce UNLINK-VL, a real-world benchmark for cross-modal knowledge unlearning in VLMs. Under a post-hoc unlearning setting in which the original forget and retain corpora are unavailable, UNLINK-VL selects visually identifiable real-world entities as unlearning targets and associates them with corresponding images and one-hop and multi-hop facts derived from Wikidata. The benchmark comprises four complementary subsets that evaluate direct forgetting of target knowledge, the propagation of forgetting through relational knowledge, the preservation of related non-target knowledge, and robustness to semantically equivalent queries. We train models under text-only and multimodal unlearning settings and evaluate forgetting effectiveness and retained utility across textual, visual, and cross-modal scenarios. Extensive experiments reveal a pronounced asymmetry in cross-modal transfer: multimodal unlearning remains effective under textual evaluation, whereas text-only unlearning transfers poorly to visual and cross-modal scenarios. Meanwhile, the evaluated methods largely preserve the models' general capabilities. These findings demonstrate that relying solely on intra-modal evaluation, particularly text-only evaluation, may substantially overestimate the effectiveness of knowledge unlearning in VLMs, underscoring the need for cross-modal unlearning and evaluation.
Chinese Translation
视觉-语言模型(VLMs),如大型语言模型(LLMs),可能会从其预训练语料库中记忆敏感、受版权保护或有害的知识。消除这些知识对于构建可信赖的人工智能系统至关重要。然而,现有研究主要集中在单一模态内的遗忘。尽管近期的工作已开始探索遗忘中的跨模态一致性,但现实世界知识遗忘的跨模态转移仍然研究不足。为了解决这一问题,我们引入了UNLINK-VL,这是一个用于VLMs中跨模态知识遗忘的现实基准。在一种后验遗忘设置下,原始的遗忘和保留语料库不可用,UNLINK-VL选择可视化识别的现实世界实体作为遗忘目标,并将其与相应的图像以及来自Wikidata的一跳和多跳事实关联。该基准包含四个互补子集,评估目标知识的直接遗忘、通过关系知识的遗忘传播、相关非目标知识的保留以及对语义等价查询的鲁棒性。我们在仅文本和多模态遗忘设置下训练模型,并评估文本、视觉和跨模态场景中的遗忘有效性和保留效用。大量实验揭示了跨模态转移的明显不对称性:多模态遗忘在文本评估下仍然有效,而仅文本遗忘在视觉和跨模态场景中转移效果较差。同时,被评估的方法在很大程度上保留了模型的整体能力。这些发现表明,仅依赖模态内评估,特别是仅文本评估,可能会显著高估VLMs中知识遗忘的有效性,强调了跨模态遗忘和评估的必要性。
cs.AI / 103 / 2608.03838
LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards
LatentGuard:高效且可检查的潜在推理用于大型语言模型的安全防护
Abstract
Reasoning-based guard models improve LLM safeguards, but decoding explicit rationales for every interaction makes them costly to deploy. Although latent-reasoning methods reduce token generation by moving reasoning into continuous states, they remain underexplored for safety moderation and lack an inspection interface for deployment. In this paper, we propose LatentGuard, an efficient and inspectable safeguard framework that brings continuous latent reasoning to guard models. LatentGuard uses a staged curriculum to progressively compress task-aligned textual rationales into compact latent states, enabling safety verdicts to be predicted directly from continuous representations. To preserve inspectability, an isolated auxiliary decoder generates compact audit artifacts on demand, keeping rationale generation off the standard inference path. Experiments show that LatentGuard-8B improves mean weighted F1 from 83.95 to 84.91 over GuardReasoner-8B, while reducing critical-path reasoning cost from 268.56 generated rationale tokens to 1.60 latent reasoning tokens. Its audit decoder achieves an audit utility score of 85.75, demonstrating an efficient and inspectable path toward deployable LLM safeguards.
Chinese Translation
基于推理的防护模型改善了大型语言模型(LLM)的安全防护,但为每次交互解码明确的推理过程使其部署成本高昂。尽管潜在推理方法通过将推理转移到连续状态来减少令牌生成,但在安全审核方面仍然未得到充分探索,并且缺乏可检查的部署接口。本文提出了LatentGuard,一个高效且可检查的防护框架,将连续潜在推理引入防护模型。LatentGuard使用分阶段的课程,逐步将与任务对齐的文本推理压缩为紧凑的潜在状态,从而使安全判决能够直接从连续表示中预测。为了保持可检查性,一个独立的辅助解码器按需生成紧凑的审计文档,将推理生成与标准推理路径分开。实验表明,LatentGuard-8B在GuardReasoner-8B的基础上将平均加权F1从83.95提高到84.91,同时将关键路径推理成本从268.56个生成的推理令牌降低到1.60个潜在推理令牌。其审计解码器的审计效用得分达到85.75,展示了通向可部署的LLM安全防护的高效且可检查的路径。
cs.AI / 104 / 2608.03839
Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes
Oilbird:无需训练的投机解码,利用验证者已计算的键
Abstract
Training-free speculative decoding drafts by matching an exact suffix of the context against a pool of earlier context. That lookup misses correct drafts already in the pool, most visibly on tool-calling traffic, where a request repeats almost everything but the few values minted for it, and where one rejected token discards the correct continuation behind it. We diagnose the failure position by position across ten benchmarks and find it to be a problem of addressing rather than of coverage: on our densest tool-calling benchmark, about half of what the strongest exact-match drafter misses is present in the pool yet unreachable by exact matching. We therefore propose a second, semantic draft source: the same pool, re-keyed by the hidden state the verifier has already computed at each committed token, together with a merge that lets it ride inside an existing lexical drafter's tree. In three published drafters, at matched pool and budget, it lifts accepted length by 24-29%. Oilbird reaches 4.4x autoregressive decoding speed on API-Bank, against 3.9x for the strongest training-free baseline in our harness and 2.0x for EAGLE-3.
Chinese Translation
无需训练的投机解码通过将上下文的确切后缀与早期上下文池进行匹配来草拟。然而,这种查找会遗漏池中已经存在的正确草拟,尤其是在工具调用流量中,在这种情况下,请求几乎重复了所有内容,只有为其生成的少量值例外,而一个被拒绝的标记会丢弃其后面的正确延续。我们在十个基准测试中逐一诊断失败的位置,发现这是一个寻址问题,而非覆盖问题:在我们最密集的工具调用基准测试中,最强的确切匹配草拟者遗漏的约一半内容在池中存在,但无法通过确切匹配访问。因此,我们提出第二个语义草拟源:同一个池,通过验证者在每个已提交标记处已计算的隐藏状态重新键入,并结合一个合并,使其能够在现有词汇草拟者的树中进行操作。在三种已发布的草拟者中,在匹配的池和预算下,它提高了接受长度24-29%。在API-Bank上,Oilbird达到了4.4倍的自回归解码速度,而我们测试中的最强无需训练的基线为3.9倍,EAGLE-3为2.0倍。
cs.AI / 105 / 2608.03844
MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Agents
MAFIA:针对审计的查询仅内存攻击通过探测和事实注入
Abstract
Memory-augmented LLM agents rely on rich context for long-horizon reasoning and acting, yet their memory modules expose a persistent attack surface for malicious records, making the study of memory poisoning threats imperative. However, existing query-only attacks often fail to remain effective in two realistic and prevalent settings: large-scale benign memory pools and active input auditing. Consequently, current approaches fall short when facing the dual challenges of high retrieval competitiveness and rigorous semantic checks. To overcome these limitations, we propose MAFIA, a query-only Memory Attack framework via probing and Factual Injection against Audit, tailored to this extended threat model. Specifically, MAFIA introduces: (1) a placement strategy that ensures retrieval-competitive injection via memory probing, budget allocation, and scheduling; and (2) a payload design that bypasses audits using compact factual cloaks, preserving malicious effects while maintaining high semantic similarity. Extensive evaluations reveal that MAFIA achieves up to a 90.7% attack success rate while suppressing audit detection from a peak of 83.3% to at most 7.4%, exposing critical vulnerabilities across agentic memory systems. Code will be made publicly available at https://github.com/JiamingChen1234/MAFIA.
Chinese Translation
增强记忆的LLM代理依赖于丰富的上下文进行长期推理和行动,但其内存模块暴露了恶意记录的持续攻击面,因此研究内存中毒威胁变得至关重要。然而,现有的查询仅攻击在两个现实且普遍的环境中往往效果不佳:大规模良性内存池和主动输入审计。因此,当前的方法在面对高检索竞争性和严格语义检查的双重挑战时显得不足。为克服这些限制,我们提出了MAFIA,一个针对审计的查询仅内存攻击框架,通过探测和事实注入,专门针对这一扩展威胁模型。具体而言,MAFIA引入了:(1)一种确保通过内存探测、预算分配和调度实现检索竞争性注入的放置策略;(2)一种使用紧凑事实伪装绕过审计的有效载荷设计,保持恶意效果的同时保持高语义相似性。广泛的评估表明,MAFIA的攻击成功率高达90.7%,同时将审计检测从最高的83.3%抑制到最多7.4%,暴露了代理内存系统的关键漏洞。代码将公开发布在 https://github.com/JiamingChen1234/MAFIA。
cs.AI / 106 / 2608.03866
ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Industrial LLM Advisories
ADMITBench:一个安全治理的参考框架,用于评估工业 LLM 建议的可接受性
Abstract
This white paper presents ADMITBench, a reference framework for evaluating industrial LLM advisories at the level of the proposed action. The framework implements a versioned, safety-governed evaluation contract that checks whether a recommendation is supported by the available evidence, permitted under the stated authority and procedure, and acceptable under the plant-specific consequence checks encoded in the selected evaluation profile. In this report, \emph{safety-governed} means that eligibility is determined through explicit, non-compensatory checks derived from a versioned plant profile; it does not mean that the evaluator, model, or plant has been safety-certified. Release 0.1.0 is a public reference implementation for technical and research evaluation, not an authorisation for physical execution.
Chinese Translation
本文介绍了 ADMITBench,这是一个用于评估工业 LLM 建议的参考框架,评估的重点在于所提议的行动。该框架实施了一个版本化的、安全治理的评估合同,检查推荐是否得到可用证据的支持,是否在声明的权限和程序下被允许,以及在所选评估配置文件中编码的特定于工厂的后果检查下是否可接受。在本报告中, extit{安全治理}意味着资格是通过从版本化的工厂配置文件中派生的明确的、非补偿性检查来确定的;这并不意味着评估者、模型或工厂已获得安全认证。版本 0.1.0 是一个公开的参考实现,用于技术和研究评估,而不是物理执行的授权。
cs.AI / 107 / 2608.03874
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
持续技能基准:大型语言模型代理能否真正进化其能力?
Abstract
Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, we introduce ContinualSkillBench, a dynamic evaluation framework for in-context continual skill learning. It covers five representative domains, each containing 100 interconnected subtasks ordered by increasing difficulty and opportunities for cross-task skill reuse. Our experiments show that sequential execution generally improves performance, but the gains vary substantially across models and domains. Moreover, in-context learning performs comparably to explicit skill maintenance on average, suggesting that much of the improvement arises from adaptation to prior context and feedback rather than reusable skill abstraction alone. Explicit skills nevertheless provide selective benefits for tasks requiring reusable procedures or precise outputs. We further find that less capable models tend to accumulate larger, more fragmented collections of task-specific skills. These findings show that current in-context skill evolution mechanisms can support continual adaptation, but still struggle to consistently consolidate experience into robust and transferable skills.
Chinese Translation
现代代理框架为大型语言模型配备了外部技能库,以解决复杂任务。然而,目前尚不清楚这些系统是否能够有效地进化其技能,以及所获得的技能是否能提升任务解决能力。为了解决这一问题,我们引入了持续技能基准(ContinualSkillBench),这是一个用于上下文中持续技能学习的动态评估框架。该框架涵盖五个代表性领域,每个领域包含100个相互关联的子任务,按难度递增和跨任务技能重用的机会进行排序。我们的实验表明,顺序执行通常能提高性能,但不同模型和领域之间的增益差异显著。此外,平均而言,上下文学习的表现与显式技能维护相当,这表明大部分改进源于对先前上下文和反馈的适应,而不仅仅是可重用技能抽象。然而,显式技能在需要可重用程序或精确输出的任务中提供了选择性优势。我们进一步发现,能力较弱的模型往往会积累更大且更分散的任务特定技能集合。这些发现表明,当前的上下文技能进化机制可以支持持续适应,但在将经验稳固整合为强大且可转移的技能方面仍然存在困难。
cs.AI / 108 / 2608.03892
Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition
通过对比激活附加在 Qwen3 中进行跨期偏好引导
Abstract
We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, recommendations, and capabilities. We train contrastive linear probes on teacher-forced temporal-choice answers to find a short-term versus long-term direction in the model's residual stream, and evaluate contrastive activation-addition steering on a held-out binary temporal-choice task, an out-of-distribution monetary intertemporal-choice task, and a TravelPlanner capability benchmark. The central result is that temporal-horizon directions can be identified with simple contrastive linear probes and then used for steering to induce large, bidirectional preference changes. On an out-of-distribution monetary choice task that varies reward size and delay, steering strongly shifts the model's indifference threshold between smaller-sooner and larger-later rewards in both directions. We further show improvements on a planning-related capability metric under moderate temporal steering. These results suggest that model intertemporal preferences are measurable and steerable, which is relevant for AI systems that give advice involving delayed costs and benefits, and for safety questions about long-horizon planning.
Chinese Translation
我们研究了大型语言模型 Qwen3-32B 中时间视野的线性表示,并利用这些表示来改变模型的时间相关偏好、推荐和能力。我们在教师强制的时间选择答案上训练对比线性探针,以找到模型残差流中的短期与长期方向,并在一个保留的二元时间选择任务、一个分布外的货币跨期选择任务以及一个旅行规划能力基准上评估对比激活附加引导。核心结果是,时间视野方向可以通过简单的对比线性探针识别,并随后用于引导,以诱导出大的双向偏好变化。在一个变化奖励大小和延迟的分布外货币选择任务中,引导显著改变了模型在较小的较早奖励和较大的较晚奖励之间的无差异阈值,且变化方向均有体现。我们进一步展示了在适度时间引导下,规划相关能力指标的改善。这些结果表明,模型的跨期偏好是可测量和可引导的,这对于涉及延迟成本和收益的建议的人工智能系统以及关于长期规划的安全问题具有重要意义。
cs.AI / 109 / 2608.03902
When Efficiency Becomes Fragility: Exploiting Dynamic Routing Vulnerabilities in Adaptive UAV Tracking
当效率变为脆弱性:利用自适应无人机跟踪中的动态路由漏洞
Abstract
Resource constraints on UAV platforms have driven a paradigm shift in aerial tracking, from pursuing performance toward balancing accuracy with efficiency. Adaptive Transformer Trackers, which leverage an input-dependent dynamic routing architecture, have emerged as a representative solution to this challenge. However, we reveal that behind this computation-on-demand flexibility hides a critical structural flaw: the Lipschitz singularity of computational path decisions, which has an unbounded local Lipschitz constant at discrete layer-skipping decision boundaries. This mathematical discontinuity renders adaptive tracking networks inherently unstable: tiny input perturbations can be amplified at the gating modules, causing dramatic changes in the inference topology. We formally characterize this singularity in the context of adaptive tracking architectures and, for the first time, identify it as a directly exploitable new attack surface. This insight reveals a previously overlooked and highly vulnerable topological path space attack surface. Based on this, we propose the Adversarial Path-Inversion (API) framework. API generates imperceptible perturbations to precisely manipulate the gating decisions, forcing the inference onto altered computational paths. The severe inconsistency between the original and the inverted paths dismantles the representation capability of the model. Extensive experiments on state-of-the-art adaptive trackers demonstrate that API achieves superior perturbation stealthiness, more effective attack, and faster inference speeds. This work opens a new dimension for the security analysis of dynamic tracking networks and provides a theoretical warning for constructing robust adaptive tracking architectures in the future.
Chinese Translation
无人机平台的资源限制推动了空中跟踪的范式转变,从追求性能转向平衡准确性与效率。自适应变压器跟踪器(Adaptive Transformer Trackers)利用输入依赖的动态路由架构,已成为应对这一挑战的代表性解决方案。然而,我们揭示出在这种按需计算灵活性背后隐藏着一个关键的结构缺陷:计算路径决策的Lipschitz奇点,在离散层跳跃决策边界处具有无界的局部Lipschitz常数。这种数学不连续性使得自适应跟踪网络本质上不稳定:微小的输入扰动可以在门控模块中被放大,导致推理拓扑的剧烈变化。我们在自适应跟踪架构的背景下正式表征了这一奇点,并首次将其识别为一个可直接利用的新攻击面。这一洞察揭示了一个先前被忽视且高度脆弱的拓扑路径空间攻击面。基于此,我们提出了对抗路径反转(Adversarial Path-Inversion, API)框架。API生成不可察觉的扰动,以精确操控门控决策,迫使推理进入改变的计算路径。原始路径与反转路径之间的严重不一致性破坏了模型的表示能力。在最先进的自适应跟踪器上进行的广泛实验表明,API在扰动隐蔽性、攻击效果和推理速度方面均表现出色。这项工作为动态跟踪网络的安全分析开辟了新的维度,并为未来构建稳健的自适应跟踪架构提供了理论警示。
cs.AI / 110 / 2608.03910
Socially Grounded Agentic AI: Coordinating Plural Perspectives through Social Theory
社会基础的自主智能体:通过社会理论协调多元视角
Abstract
As AI systems are deployed across increasingly diverse social contexts, alignment can no longer be framed as the optimization of a single, unified set of values. Instead, systems must be able to recognize, represent, and respond to multiple legitimate perspectives. This has led to growing interest in pluralistic alignment, which seeks to move beyond one-size-fits-all models of appropriate behaviour. However, current approaches often lack a clear account of how values are socially organized, contested, and coordinated in practice. In this paper, we argue that social theory provides essential conceptual and design resources for addressing these challenges. Drawing on established traditions in sociology, we show how perspectives can be understood as structured by roles, shaped through interaction, and distributed across fields of power and expertise. We translate these insights into concrete implications for AI system design, including role-based representations, structured coordination among perspectives, and context-sensitive evaluation. For agentic systems, this requires aligning not only final outputs, but also the role activations, deliberative traces, aggregation rules, and feedback loops through which those outputs are produced. Our contribution is to reposition pluralistic alignment as a problem of socially grounded coordination rather than output diversification. We outline a design space for systems that engage multiple perspectives in structured and accountable ways, and we identify directions for future work to implement and empirically evaluate these approaches in real-world settings.
Chinese Translation
随着人工智能系统在日益多样化的社会背景中部署,价值观的对齐不再能够被框定为优化单一统一的价值观集合。相反,系统必须能够识别、表示并响应多种合法的视角。这引发了对多元化对齐的日益关注,旨在超越一刀切的适当行为模型。然而,当前的方法往往缺乏对价值观在实践中如何被社会组织、争议和协调的清晰阐述。本文论证了社会理论为应对这些挑战提供了重要的概念和设计资源。借鉴社会学中的既定传统,我们展示了如何理解视角是通过角色结构化、通过互动塑造,并在权力和专业领域中分布的。我们将这些见解转化为人工智能系统设计的具体启示,包括基于角色的表示、视角之间的结构化协调以及上下文敏感的评估。对于自主系统而言,这不仅需要对最终输出进行对齐,还需要对产生这些输出的角色激活、深思过程、聚合规则和反馈循环进行对齐。我们的贡献在于将多元化对齐重新定位为社会基础协调的问题,而非输出多样化。我们概述了一个设计空间,以结构化和负责任的方式参与多种视角,并确定了未来工作在真实世界环境中实施和实证评估这些方法的方向。
cs.AI / 111 / 2608.03917
Implementing Causal Perception: Competing SCMs and Situated Fairness
实施因果感知:竞争的结构因果模型与情境公平性
Abstract
Causal perception occurs when agents with competing Structural Causal Models (SCMs) of the same system infer different probability distributions, including the hypothetical distributions implied by each agent's SCM under the same set of interventions. It shapes how agents reason about the system and how they perceive its fairness. Causal perception is a promising probabilistic framework, but it has remained purely theoretical. This work provides the first implementation of the causal perception framework of \'Alvarez and Ruggieri (2025). We operationalize structural (agents disagree on the causal graph) and parametrical (agents agree on the causal graph but disagree on its weights) causal perception. We design algorithms for computing interventional and counterfactual distributions and propose suitable distance measures to quantify the disagreement. Using the German Credit dataset, we illustrate how causal perception affects accuracy and fairness in a multi-expert decision setting. We show that the perception verdict is sensitive to the choice of distance metric and threshold. We also show that causal perception changes fairness assessments and threshold-based decisions. Bias proves situated with respect to the agent's SCM, demonstrating that competing worldviews in fairness problems cannot be ignored.
Chinese Translation
因果感知发生在具有竞争性结构因果模型(Structural Causal Models, SCMs)的代理人对同一系统推断出不同的概率分布时,包括在相同干预条件下每个代理人的SCM所隐含的假设分布。它影响代理人对系统的推理方式以及他们对公平性的感知。因果感知是一个有前景的概率框架,但迄今为止仍然纯粹理论化。本研究提供了阿尔瓦雷斯(Alvarez)和鲁吉耶里(Ruggieri, 2025)因果感知框架的首次实现。我们对结构性(代理人对因果图存在分歧)和参数性(代理人对因果图一致但对其权重存在分歧)因果感知进行了操作化。我们设计了计算干预和反事实分布的算法,并提出了合适的距离度量来量化分歧。利用德国信用数据集,我们展示了因果感知如何影响多专家决策环境中的准确性和公平性。我们表明,感知裁决对距离度量和阈值的选择敏感。我们还表明,因果感知改变了公平性评估和基于阈值的决策。偏见证明与代理人的SCM相关,表明在公平性问题中竞争的世界观不可忽视。
cs.AI / 112 / 2608.03921
The Transformer Revolution, Part 1: Dynamic Processing through Output- Weight Interconnections
变压器革命,第一部分:通过输出权重互连进行动态处理
Abstract
This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that large language models merely reproduce statistical regularities learned in training, we argue that Transformers construct and apply prompt-dependent transformations whose parameters are generated during inference. We call this form of computation SIDPP: Sequence-level Interactive Dynamic Parallel Processing. The Transformer is interpreted as a system that transforms concepts by means of concepts. Token vectors are the concepts to be transformed; parameterized transformations defined by matrices and vectors are the transforming concepts. These may be static, when fixed through training, or dynamic, when generated from the input sequence. Mechanically, they correspond to groups of simple neural networks. The Transformer's architectural novelty lies in output-weight interconnections, through which the outputs of some networks determine the weights of others, alongside ordinary output-input interconnections. By means of these interconnections, the system constructs transformations from the prompt and uses them to modify token representations. The contribution of dynamic processing grows with prompt length and may equal or exceed that of static processing, a phenomenon we call strong prompt sensitivity. This account bears on interpretability, predictability, control, and the design of smaller, more sustainable systems. Finally, since the human neural system possesses the mechanisms required to implement SIDPP, we argue that a form of SIDPP may, in principle, be neurally realized in the cerebral cortex. We therefore conjecture that human language processing may itself be a form of SIDPP produced by a functional architecture relevantly similar to that of the Transformer.
Chinese Translation
本文对变压器在推理过程中的作用提供了一种新的解释。针对“大型语言模型仅仅重现训练中学习的统计规律”的“随机鹦鹉”观点,我们认为变压器构建并应用依赖于提示的变换,其参数在推理过程中生成。我们将这种计算形式称为SIDPP:序列级交互动态并行处理。变压器被解释为通过概念来转化概念的系统。标记向量是需要转化的概念;由矩阵和向量定义的参数化变换是转化概念。这些变换可以是静态的(在训练中固定)或动态的(从输入序列生成)。在机械上,它们对应于简单神经网络的组。变压器的架构新颖性在于输出权重互连,通过这些互连,一些网络的输出决定其他网络的权重,此外还有普通的输出-输入互连。通过这些互连,系统从提示中构建变换,并利用它们来修改标记表示。动态处理的贡献随着提示长度的增加而增长,可能等于或超过静态处理的贡献,这一现象我们称之为强提示敏感性。这一解释与可解释性、可预测性、控制以及更小、更可持续系统的设计相关。最后,由于人类神经系统具备实现SIDPP所需的机制,我们认为在原则上,SIDPP的某种形式可能在大脑皮层中以神经方式实现。因此,我们推测人类语言处理本身可能是一种由与变压器相关的功能架构产生的SIDPP形式。
cs.AI / 113 / 2608.03952
TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring
TACT:与分类法对齐的后训练用于教育适应性英语辅导
Abstract
Large language models (LLMs) are increasingly used to provide conversational practice for English-as-a-second-language (ESL) learners. Effective ESL tutoring, however, requires more than fluent response generation: a tutor must select an appropriate pedagogical action based on learner behavior and dialogue context. Human-tutoring research offers principles for adaptive support, but they are often task-specific and remain insufficiently integrated into LLM-based ESL tutor training and evaluation. We present TACT (Taxonomy-Aligned Conversational Tutor), a human-grounded framework for post-training and evaluating pedagogically adaptive ESL tutors. Drawing on established literature, we develop two complementary taxonomies: the Tutor-Strategy Taxonomy with 13 tutor response strategies and the Student-Move Taxonomy characterizing learner behavior by move type and status. Using these taxonomies, we construct TACTCorpus, which enriches 260 authentic teacher-student conversations with 32,379 annotations and quality-controlled augmented training data. We then post-train Qwen3.5-4B through supervised fine-tuning followed by taxonomy-aligned Group Relative Policy Optimization, producing TACTutor and optimizing it for scaffolding quality rather than reference imitation alone. On TACTBench, a strategy-balanced diagnostic benchmark comprising 78 authentic tutoring contexts, TACTutor improves over its backbone by 20.30% and outperforms all evaluated proprietary baselines under the same protocol, while maintaining backbone performance on established external educational benchmarks; in a blinded study with 50 learners, it also receives the highest overall mean rating among the evaluated tutors. We release the data, benchmark, and model weights, providing an open foundation for developing pedagogically adaptive ESL tutors.
Chinese Translation
大型语言模型(LLMs)越来越多地被用于为英语作为第二语言(ESL)学习者提供对话练习。然而,有效的ESL辅导不仅需要流利的响应生成:辅导者必须根据学习者的行为和对话上下文选择适当的教学行动。人类辅导研究提供了适应性支持的原则,但这些原则往往是特定于任务的,并且在基于LLM的ESL辅导培训和评估中整合不足。我们提出了TACT(与分类法对齐的对话辅导),这是一个基于人类的框架,用于后训练和评估教育适应性ESL辅导。借鉴已有文献,我们开发了两个互补的分类法:包含13种辅导响应策略的辅导策略分类法和根据移动类型和状态描述学习者行为的学习者移动分类法。利用这些分类法,我们构建了TACTCorpus,该语料库丰富了260个真实的师生对话,包含32,379个注释和经过质量控制的增强训练数据。然后,我们通过监督微调和分类法对齐的组相对策略优化对Qwen3.5-4B进行后训练,生成TACTutor,并优化其支架质量,而不仅仅是参考模仿。在包含78个真实辅导情境的策略平衡诊断基准TACTBench上,TACTutor的表现比其基础模型提高了20.30%,并在相同协议下超越了所有评估的专有基线,同时在已建立的外部教育基准上保持基础模型的表现;在一项涉及50名学习者的盲测中,它还获得了评估辅导中最高的总体平均评分。我们发布了数据、基准和模型权重,为开发教育适应性ESL辅导提供了开放的基础。
cs.AI / 114 / 2608.03958
A game theory for foundation models shows new paths to rational cooperation through similarity inference
基础模型的博弈理论通过相似性推理展示了理性合作的新路径
Abstract
As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, the dominant framework for modeling rational interaction, is built upon the assumption of `decoupled agency,' where agents treat their own decision-making as independent of the environment and other actors. Modern AI agents, however, jointly predict their own future actions alongside external observations. Here, we report a striking finding: when interacting in stylized social dilemmas, foundation model agents engaging in optimal planning consistently converge to stable cooperation, directly contradicting classical game-theoretic predictions of mutual defection. To understand this phenomenon, we introduce the `embedded Bayesian agent,' a theoretical model for foundation model agents. By shifting from decoupled to embedded agency, these agents model themselves as part of the universe they inhabit, maintaining epistemic uncertainty about their own decision-making algorithms. We show that by inferring whether others are behaviorally similar, an embedded agent treats its own deliberation during planning as evidence: a decision to cooperate predicts a similar decision by a similar partner. We formalize this mechanism of similarity inference through the `embedded equilibrium,' a novel solution concept replacing the Nash equilibrium to provide a foundational game theory for the social behavior of modern AI agents.
Chinese Translation
随着由基础模型驱动的自主代理越来越多地融入社会和经济系统,理解其集体行为的原则对于确保安全和合作至关重要。经典博弈理论,作为建模理性互动的主导框架,建立在“解耦代理”的假设之上,即代理将自己的决策视为独立于环境和其他参与者。然而,现代人工智能代理在进行决策时会同时预测自己的未来行为和外部观察。在此,我们报告了一个引人注目的发现:在风格化的社会困境中,进行最佳规划的基础模型代理在互动中始终趋向于稳定的合作,这直接与经典博弈理论对相互背叛的预测相矛盾。为了理解这一现象,我们引入了“嵌入式贝叶斯代理”,这是一个针对基础模型代理的理论模型。通过从解耦代理转向嵌入式代理,这些代理将自己建模为其所处宇宙的一部分,并对自身决策算法保持认知不确定性。我们展示了通过推断他人是否在行为上相似,嵌入式代理在规划期间将自己的思考视为证据:合作的决策预测相似伙伴的相似决策。我们通过“嵌入均衡”这一新颖的解决概念形式化了这一相似性推理机制,替代纳什均衡,为现代人工智能代理的社会行为提供了基础博弈理论。
cs.AI / 115 / 2608.03961
Interpretable Adaptive Sampling for LLM Test-Time Scaling
可解释的自适应采样用于大语言模型的测试时间扩展
Abstract
Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines use fixed per-query budgets that spend the same compute on easy and difficult prompts. These fixed budgets are also difficult to inspect because they do not explain why a given prompt receives a particular number of samples. We propose adaptive} test-time scaling with a lightweight fuzzy controller that maps interpretable signals, including estimated prompt complexity and model confidence, to a per-query sampling budget. The controller assigns fewer samples to easier or more confident prompts and more samples to harder or less certain prompts, making inference-time compute inspectable rather than fixed or opaque. We evaluate under a fair-alignment protocol with matched decoding settings and controlled answer selection, and compare against best-of-$N$, compute-aware scaling, and self-certainty-based baselines on question-answering and mathematical reasoning tasks. Across models and datasets, adaptive fuzzy control improves over several standard baselines and remains close to a selector-matched full-budget control while reducing the average number of samples. These findings suggest that interpretable adaptive sampling is a practical direction for more efficient test-time reasoning in large language models.
Chinese Translation
测试时间扩展通过生成和聚合多个候选答案来提高大语言模型(LLM)的推理能力,但许多流程使用固定的每查询预算,对简单和困难的提示消耗相同的计算资源。这些固定预算也难以检查,因为它们无法解释为什么特定提示会获得特定数量的样本。我们提出了一种自适应测试时间扩展方法,采用轻量级模糊控制器,将可解释信号(包括估计的提示复杂性和模型置信度)映射到每查询的采样预算。该控制器对更简单或更有信心的提示分配较少的样本,而对更困难或不太确定的提示分配更多的样本,使得推理时的计算可检查,而不是固定或不透明的。我们在公平对齐协议下进行评估,匹配解码设置和控制答案选择,并与最佳的$N$、计算感知扩展和基于自我置信度的基线进行比较,应用于问答和数学推理任务。在不同模型和数据集上,自适应模糊控制在多个标准基线之上取得了改进,并且在减少平均样本数量的同时,仍然接近选择器匹配的全预算控制。这些发现表明,可解释的自适应采样是提高大语言模型测试时间推理效率的一个实用方向。
cs.AI / 116 / 2608.03970
Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations
我们应该通过打字还是语音与大型语言模型代理进行交互?对语音和键盘输入扰动的综合研究
Abstract
Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, disfluency from conventional transcription and restructuring from AI-backed dictation tools. How do they impact an LLM's performance? In this paper we present HIVE (Human Input-Variation Engine), a suite of voice transcription perturbations and QWERTY keyboard perturbations. We use HIVE to evaluate how robust models are to these perturbations. We present seven findings. (i) Voice transcription perturbations lower accuracy across every instruction-tuned model we test, and it is the structure of the transcription rather than its fillers that carries the cost. (ii) QWERTY keyboard perturbations cost less, and a model absorbs a lot of them before accuracy falls away. (iii) Both trace back to one cause, how many of the question's tokens survive the perturbation: destroying a token is what hurts, while adding new ones alongside it costs little. (iv) The gap between the two channels appears only where the answer must be constructed or deduced; on multiple choice there is none. (v) The harm does not solely come from test-set contamination. (vi) It cannot be trained away with lightweight adaptation. (vii) A thinking budget recovers the keyboard channel almost entirely but leaves the spoken registers untouched, and compressed speech is worse with it.
Chinese Translation
人类通过打字或说话向语言模型输入信息,而每种输入方式都会留下独特的特征:键盘输入产生拼写噪声;语音输入则因传统转录中的不流畅性和基于人工智能的听写工具的重组而受到影响。这些因素如何影响大型语言模型(LLM)的性能?在本文中,我们提出了HIVE(人类输入变异引擎),这是一个包含语音转录扰动和QWERTY键盘扰动的工具包。我们使用HIVE评估模型对这些扰动的鲁棒性。我们提出了七个发现。(i) 语音转录扰动降低了我们测试的每个指令调优模型的准确性,且影响主要来自转录的结构而非填充内容。(ii) QWERTY键盘扰动的影响较小,模型在准确性下降之前可以承受大量扰动。(iii) 两者的根本原因相同,即问题的多少个标记在扰动后存活:破坏一个标记是造成损害的原因,而在其旁边添加新标记的成本较低。(iv) 只有在需要构建或推导答案时,两种输入方式之间的差距才会显现;在多项选择题中则没有差距。(v) 这种损害不仅仅来自测试集的污染。(vi) 通过轻量级适应无法消除这种影响。(vii) 通过思维预算几乎可以完全恢复键盘输入通道,但对语音输入的影响则未得到改善,且压缩语音的效果更差。
cs.AI / 117 / 2608.03972
ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
ReflectRL:通过反思到直接推理学习黄金负轨迹
Abstract
On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-guided methods lose their main source of supervision, and these failed trajectories are typically discarded as negative samples. We argue that such failures, which we call Golden Negative Trajectories, can still provide valuable reasoning signals when treated not as demonstrations to imitate, but as flawed trajectories to reflect upon. We identify a Reflection Advantage: for hard problems, reflecting on a flawed trajectory can be easier and more effective than solving the problem directly from scratch. Motivated by this, we propose ReflectRL, a lightweight plug-and-play framework that learns from Golden Negative Trajectories during on-policy training. ReflectRL first uses these trajectories to elicit Reflective Reasoning, then applies Reflective-to-Direct Policy Transition to transfer the acquired reasoning behavior back to Direct Reasoning. Experiments across 9 benchmarks, 4 LLM backbones, and 4 on-policy training methods show that ReflectRL consistently improves reasoning performance with minimal overhead.
Chinese Translation
基于策略的训练已成为提高大型语言模型推理能力的一种强大后训练范式,通常通过来自更强专家模型的黄金轨迹来增强。然而,当专家在更难的问题上失败时,现有的轨迹引导方法失去了主要的监督来源,这些失败的轨迹通常被视为负样本而被丢弃。我们认为,这种失败,即我们所称的黄金负轨迹,在被视为反思的缺陷轨迹而非模仿的示范时,仍然可以提供有价值的推理信号。我们确定了一个反思优势:对于困难问题,反思一个缺陷轨迹可能比从头直接解决问题更容易且更有效。基于此,我们提出了ReflectRL,一个轻量级即插即用框架,在基于策略的训练过程中从黄金负轨迹中学习。ReflectRL首先利用这些轨迹引发反思推理,然后应用反思到直接策略转移,将获得的推理行为转移回直接推理。在9个基准、4个大型语言模型骨干和4种基于策略的训练方法上的实验表明,ReflectRL在最小开销的情况下始终提高了推理性能。
cs.CL / 1 / 2608.02609
TabletCraft: Bridging a 4,000-Year Cultural Gap with Bidirectional Akkadian NMT and Cuneiform Rendering
TabletCraft:通过双向阿卡德语神经机器翻译和楔形文字呈现弥合4000年的文化鸿沟
Abstract
Half a million cuneiform clay tablets survive in museums worldwide, yet modern users can neither read nor write in the world's oldest writing system, leaving a 4,000-year cultural barrier that existing NLP tools have only partially addressed. Prior work enables one-way, scholar-oriented translation from Akkadian to English, but offers no path in the reverse direction: non-specialist users cannot compose new content in cuneiform, and therefore remain passive consumers of ancient culture rather than active participants. We present TabletCraft, the first open-source system that enables bidirectional interaction with Mesopotamian writing. Users can read ancient tablets (Akkadian to English) and compose new messages as cuneiform clay tablets (English to Akkadian to cuneiform to rendered tablet). The system integrates a ByT5-based translation model trained on 116K bidirectional samples, a cuneiform sign converter with 14,240 mappings (95.3% coverage), and a visual tablet renderer, packaged as a pip-installable toolkit with CLI and web demo. On the held-out Akkademia validation split (2,812 samples), we report 49.1 BLEU for Akkadian-to-English and 48.5 BLEU for English-to-Akkadian, the first published quantitative result in the reverse direction.
Chinese Translation
全球各地的博物馆中保存着50万个楔形泥板,但现代用户既无法阅读也无法书写这一世界上最古老的书写系统,这导致了一个4000年的文化障碍,而现有的自然语言处理工具仅部分解决了这一问题。之前的研究实现了从阿卡德语到英语的单向、学术导向的翻译,但在反向翻译方面没有提供任何途径:非专业用户无法用楔形文字创作新内容,因此只能被动地消费古代文化,而无法积极参与。我们提出了TabletCraft,这是第一个开放源代码系统,能够实现与美索不达米亚书写的双向交互。用户可以阅读古代泥板(阿卡德语到英语)并创作新的信息作为楔形泥板(英语到阿卡德语再到楔形文字再到呈现的泥板)。该系统集成了一个基于ByT5的翻译模型,经过116K双向样本的训练,一个具有14,240个映射(覆盖率95.3%)的楔形文字符号转换器,以及一个视觉泥板渲染器,打包为一个可通过pip安装的工具包,包含命令行界面和网页演示。在保留的Akkademia验证集(2812个样本)上,我们报告了阿卡德语到英语的BLEU值为49.1,英语到阿卡德语的BLEU值为48.5,这是反向翻译方向上首次发布的定量结果。
cs.CL / 2 / 2608.02612
BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems
BBOWP-Bench:在黑箱优化词问题上评估大语言模型
Abstract
Formulating an optimization problem strongly affects the quality of the final solution, yet good formulations usually require substantial expertise. Recent studies have therefore examined how to automatically derive optimization problems from natural-language descriptions, but existing benchmarks focus on settings where objectives and constraints can be written explicitly as mathematical expressions. Many practically important problems are naturally treated as black-box optimization (BBO) problems, in which only objective values are observable, and the functional form is unavailable. In BBO, the search space design, a part of the problem formulation, and the selection of the optimization algorithm are crucial for problem-solving. Automating these processes with large language models (LLMs) is a significant challenge. This paper introduces Black-Box Optimization Word Problems (BBOWP), a novel problem setting in which a system must infer both a search space and an optimization algorithm from a natural-language description of a black-box optimization task. To support research on this setting, we establish the BBOWP Benchmark Suite (BBOWP-Bench), a dataset and evaluation framework for BBOWP. Each instance combines a natural-language problem description, an executable evaluation environment, and a human-designed baseline formulation, allowing evaluation of both search-space design and algorithm selection. Using this benchmark, we provide the first evaluation of LLMs and show that current LLMs are capable of selecting suitable algorithms based on the given evaluation budget. However, they sometimes struggle with search space design, particularly in identifying important variables and balancing their ranges when the problem description is less informative or the search space is highly problem-specific. Our code and dataset are available at https://github.com/shiralab/bbowp-bench.
Chinese Translation
优化问题的表述对最终解决方案的质量有着重要影响,而良好的表述通常需要相当的专业知识。因此,最近的研究探讨了如何从自然语言描述中自动推导优化问题,但现有的基准测试主要集中在目标和约束可以明确写成数学表达式的情境中。许多在实践中重要的问题自然被视为黑箱优化(BBO)问题,在这些问题中,只有目标值是可观察的,而功能形式则不可用。在BBO中,搜索空间设计作为问题表述的一部分,以及优化算法的选择,对问题的解决至关重要。利用大语言模型(LLMs)自动化这些过程是一项重大挑战。本文介绍了黑箱优化词问题(BBOWP),这是一个新颖的问题设置,其中系统必须从黑箱优化任务的自然语言描述中推断出搜索空间和优化算法。为了支持这一设置的研究,我们建立了BBOWP基准套件(BBOWP-Bench),这是一个用于BBOWP的数据集和评估框架。每个实例结合了自然语言问题描述、可执行的评估环境和人类设计的基线表述,允许对搜索空间设计和算法选择进行评估。利用这一基准,我们提供了对LLMs的首次评估,并显示当前的LLMs能够根据给定的评估预算选择合适的算法。然而,当问题描述信息较少或搜索空间高度特定于问题时,它们在搜索空间设计方面有时会遇到困难,特别是在识别重要变量和调整其范围时。我们的代码和数据集可在 https://github.com/shiralab/bbowp-bench 获取。
cs.CL / 3 / 2608.02613
MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale
MemArena:一种以自我为中心的大规模设备端代理个人记忆助手基准
Abstract
Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspective, and coherent multi-session worlds. MemArena fills these gaps with a single-world conversational benchmark built with its MASim agent simulator, for 50 agents over 15 days (10.3M dialog-text tokens, 24.1K text-only ego-observed tokens/agent/day). With the interaction history, it co-generates ground truth over six recall, reasoning, and trustworthiness evaluation dimensions. We evaluate five open-weight readers with Vanilla context, BM25-RAG, Oracle retrieval, Memobase, and MemSearch as memory backends. Three results stand out: (1) Memory-backend choice matters more for content accuracy: At Qwen3-0.6B, Memobase-to-MemSearch gains +32.5/+19.2 pp, exceeding MemSearch reader scaling (+10.6/+6.8 pp). (2) Permission-aware access fails universally, with Oracle leaking heavily and other backends too timid to disclose. (3) Search latency bites only at very small reader: on a Spark GB10 edge node, memory-search adds a moderate and fixed 87/7/48 ms (BM25-RAG/Memobase/MemSearch) that composes a small part of TTFT for most reader-backend combinations. Code, the MASim simulator, and the MemArena-L benchmark will be released upon acceptance.
Chinese Translation
边缘部署的个人记忆助手必须在设备上处理私密的人际对话,并使用开放权重模型。然而,现有的记忆基准往往未能充分测试活动密集型交互、自我中心视角和连贯的多会话世界的结合。MemArena 通过其 MASim 代理模拟器填补了这些空白,构建了一个单一世界的对话基准,涵盖 50 个代理,持续 15 天(103 万对话文本标记,24.1K 文本仅观察标记/代理/天)。通过交互历史,它在六个回忆、推理和可信度评估维度上共同生成真实数据。我们评估了五个开放权重的阅读器,分别使用 Vanilla 上下文、BM25-RAG、Oracle 检索、Memobase 和 MemSearch 作为记忆后端。三个结果尤为突出:(1)记忆后端的选择对内容准确性影响更大:在 Qwen3-0.6B 上,Memobase 相较于 MemSearch 提升了 +32.5/+19.2 个百分点,超越了 MemSearch 阅读器的扩展 (+10.6/+6.8 个百分点)。(2)基于权限的访问普遍失败,Oracle 严重泄露,其他后端则过于谨慎,不愿披露。(3)搜索延迟仅在非常小的阅读器上显著:在 Spark GB10 边缘节点上,记忆搜索增加了适度且固定的 87/7/48 毫秒(BM25-RAG/Memobase/MemSearch),这在大多数阅读器-后端组合中占据了 TTFT 的小部分。代码、MASim 模拟器和 MemArena-L 基准将在接受后发布。
cs.CL / 4 / 2608.02615
OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning
OncoTriad-QA:一个用于泛癌症推理的患者级放射学-病理学-基因组学基准
Abstract
Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata. However, most medical large language model (LLM) and vision-language model (VLM) benchmarks focus on isolated modalities or narrow image-text tasks, leaving patient-level oncology assessment across multiple evidence streams largely untested. We introduce OncoTriad-QA, a patient-level radiology-pathology-genomics benchmark for pan-cancer question answering. OncoTriad-QA contains 86.1k semantic questions across 9,281 TCGA patient cases from 32 cancer cohorts, aligning CT/MRI radiology, whole-slide histopathology, somatic mutations, copy-number alterations, DNA methylation, bulk RNA-seq, and clinical metadata. Case-specific annotations are constructed through a source-grounded LLM-assisted pipeline using curated labels, diagnostic reports, molecular profiles, and modality-derived evidence as primary sources of truth, with automated consistency checks and clinician review. We also introduce OncoVLM, a reference multimodal model that maps modality-native radiology, pathology, DNA methylation, and RNA-seq evidence into an LLM interface through learned projectors. Experiments show that existing general-purpose and medical LLMs remain limited on comprehensive pan-cancer QA, especially when questions require integrating imaging findings, tumor morphology, and molecular evidence. After fine-tuning on OncoTriad-QA, OncoVLM exceeds MedGemma-4B by an average of 10.7 points when using MCQ accuracy and BERTScore-F1, with consistent gains across multiple-choice and open-ended questions under radiology-only, pathology-only, and all-available settings. These results demonstrate the benchmark's value for training and evaluating models for integrated cancer question answering.
Chinese Translation
癌症的诊断和特征化需要整合来自放射学、病理学、基因组学和临床元数据的互补证据。然而,大多数医学大型语言模型(LLM)和视觉语言模型(VLM)基准主要集中在孤立的模态或狭窄的图像-文本任务上,导致患者级肿瘤学评估在多个证据流中尚未得到充分测试。我们提出了OncoTriad-QA,这是一个用于泛癌症问答的患者级放射学-病理学-基因组学基准。OncoTriad-QA包含来自32个癌症队列的9,281个TCGA患者案例中的86.1k个语义问题,整合了CT/MRI放射学、全切片组织病理学、体细胞突变、拷贝数变化、DNA甲基化、整体RNA测序和临床元数据。通过使用经过策划的标签、诊断报告、分子特征和模态衍生证据作为主要真实来源的源基础LLM辅助管道构建案例特定注释,并进行自动一致性检查和临床医生审核。我们还引入了OncoVLM,一个参考多模态模型,通过学习投影器将模态本地的放射学、病理学、DNA甲基化和RNA测序证据映射到LLM接口。实验表明,现有的通用和医学LLM在全面的泛癌症问答上仍然有限,尤其是当问题需要整合影像发现、肿瘤形态和分子证据时。在OncoTriad-QA上进行微调后,OncoVLM在使用多项选择题准确率和BERTScore-F1时,平均超过MedGemma-4B 10.7分,并在放射学单一、病理学单一和所有可用设置下的多项选择和开放式问题中均表现出一致的提升。这些结果展示了该基准在训练和评估集成癌症问答模型方面的价值。
cs.CL / 5 / 2608.02616
Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks
评估OpenAI的隐私过滤器:跨语言、跨领域的个人身份信息检测在42个基准上的表现
Abstract
We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, across 42 synthetic benchmarks spanning 22 languages and 5 domains. Zero-shot, OPF achieves F1=0.855 on AI4Privacy and 0.464 on SPY medical, outperforming Presidio (0.431, 0.273) and XLM-RoBERTa (0.269, 0.111) on PII-annotated benchmarks; on multilingual NER, XLM-RoBERTa leads OPF on all 13 Indic and non-Latin languages. GPT-4o leads on medical, legal, and financial PII (SPY: 0.643 avg, Gretel: 0.527), while OPF leads on structured synthetic PII (0.71 avg) and customer support (0.60). OPF degrades sharply when PII is embedded in narrative prose: F1=0.04--0.57 on NER benchmarks and collapse for non-Latin scripts (Arabic: 0.04, Cyrillic: 0.03). Error analysis shows OPF is strongest on structurally regular PII types (email: 0.78, phone: 0.76) and weakest on culturally variable ones (person: 0.40, address: 0.49), and is recall-biased on customer-support and medical/legal PII (P=0.31--0.54, R=0.70--0.85); global precision spans 0.31--0.86 across all domains.
Chinese Translation
我们首次独立、系统地评估了OpenAI的隐私过滤器(OPF),这是一款具有15亿参数的双向个人身份信息(PII)检测器,涵盖42个合成基准,涉及22种语言和5个领域。在零样本条件下,OPF在AI4Privacy上取得了F1=0.855,在SPY医疗数据上取得了0.464,优于Presidio(0.431,0.273)和XLM-RoBERTa(0.269,0.111)在PII标注基准上的表现;在多语言命名实体识别(NER)中,XLM-RoBERTa在所有13种印度语言和非拉丁语言上领先于OPF。GPT-4o在医疗、法律和金融PII(SPY: 平均0.643,Gretel: 0.527)方面表现最佳,而OPF在结构化合成PII(平均0.71)和客户支持(0.60)方面表现最佳。当PII嵌入叙述性散文中时,OPF的表现急剧下降:在NER基准上的F1为0.04至0.57,对于非拉丁文字(阿拉伯文:0.04,西里尔文:0.03)几乎崩溃。错误分析显示,OPF在结构上规则的PII类型(电子邮件:0.78,电话:0.76)上表现最强,而在文化上可变的类型(人名:0.40,地址:0.49)上表现最弱,并且在客户支持和医疗/法律PII上存在召回偏差(P=0.31至0.54,R=0.70至0.85);在所有领域的全球精确度范围为0.31至0.86。
cs.CL / 6 / 2608.02617
Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
偏好而非安全:成对偏好是临床安全的糟糕代理
Abstract
We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blinded pairwise preferences alongside multi-criterion rubric ratings. Clinicians assign scores on a discrete $[-2, +2]$ scale, where negative values indicate clinically unsafe or misleading content. Using 26{,}804 pairwise judgments across outputs from 13 LLMs, contributed by more than 736 clinicians across 28+ countries, we find that clinician preference is a poor proxy for safety-critical performance. Models ranking highly under pairwise preference can still exhibit substantial rates of clinically meaningful failures ($\leq -1$) on dimensions such as \emph{Harmlessness} and \emph{Accuracy}. These failures are unevenly distributed across specialties, creating domain-specific ``no-go zones'' not visible in aggregate rankings or single-number leaderboards. We further analyze contributing factors including prompt length, refusal and escalation behavior, and the relative contributions of safety-critical versus surface-level features. A substantial fraction of preference votes carry no positive safety signal, while feature decomposition shows that surface-level characteristics explain slightly more preference variation than safety-critical rubric differences. Finally, we introduce a clinically adjusted preference ranking combining pairwise preference with rubric-derived feedback, producing a more safety-aware ordering than raw Bradley--Terry strength alone. Our findings support evaluation practices that separate preference from safety, report safety-critical failure rates directly, and incorporate clinically grounded adjustments when ranking LLMs for clinical decision making.
Chinese Translation
我们评估临床医生的成对偏好在大型语言模型(LLM)评估中是否提供了临床安全的可靠信号,使用来自MOOVE(大规模开放在线验证与评估)的专家反馈,该平台由临床医生主导,收集盲评成对偏好以及多标准评分。临床医生在离散的 $[-2, +2]$ 评分尺度上进行评分,其中负值表示临床上不安全或误导性的内容。通过分析来自13个LLM的26,804个成对判断,这些判断由来自28个国家的736多名临床医生提供,我们发现临床医生的偏好是安全关键性能的糟糕代理。在成对偏好下排名较高的模型仍可能在 extit{无害性}和 extit{准确性}等维度上表现出显著的临床意义失误($ ext{failure} ext{rate} ext{ } ext{≤} -1$)。这些失误在各专业之间分布不均,形成了在汇总排名或单一数字排行榜中不可见的领域特定“禁区”。我们进一步分析了影响因素,包括提示长度、拒绝和升级行为,以及安全关键特征与表面特征的相对贡献。相当一部分偏好投票没有正面的安全信号,而特征分解显示,表面特征解释的偏好变异性略高于安全关键评分差异。最后,我们引入了一种临床调整的偏好排名,将成对偏好与评分导出的反馈相结合,产生比单纯的Bradley--Terry强度更具安全意识的排序。我们的发现支持将偏好与安全分开评估的实践,直接报告安全关键失误率,并在对LLM进行临床决策排名时纳入临床基础的调整。
cs.CL / 7 / 2608.02620
JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation
JudgeArena:一个统一的可重复性 LLM-评估框架
Abstract
LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their own code base, hardcode a specific closed-model judge, and support a single evaluation protocol. This fragmentation makes it difficult to study how design choices--the benchmark, the judge model, the prompt, the inference backend--affect the conclusions we draw about model quality. We introduce JudgeArena, an open-source framework that unifies major LLM-judge benchmarks (AlpacaEval, Arena-Hard, MT-Bench, and m-Arena-Hard) under a single interface with swappable judges and comprehensive metadata logging for increased transparency in reporting and reproducibility. It enables systematic studies of judge choices, as any model accessible via vLLM, llama.cpp, or OpenRouter can serve as both candidate and judge. Furthermore, JudgeArena ships with tuned judge configurations for open models that match or outperform closed-model judges, validated on human preference datasets in both English and multilingual settings, reducing the reliance on opaque closed models. Finally, by combining existing human annotations with LLM-judge evaluations of a target model, JudgeArena can simulate LMArena Elo scores with high accuracy offering a practical, open, and low-cost alternative to large-scale human annotation campaigns.
Chinese Translation
LLM作为评估者的评价已成为排名语言模型的主流范式,但生态系统仍然存在碎片化的问题:大多数基准测试都提供自己的代码库,硬编码特定的封闭模型评估者,并支持单一的评估协议。这种碎片化使得研究设计选择(基准测试、评估模型、提示、推理后端)如何影响我们对模型质量的结论变得困难。我们提出了 JudgeArena,一个开源框架,将主要的 LLM-评估基准(AlpacaEval、Arena-Hard、MT-Bench 和 m-Arena-Hard)统一在一个单一接口下,支持可互换的评估者和全面的元数据记录,以提高报告和可重复性的透明度。它使得评估者选择的系统研究成为可能,因为任何通过 vLLM、llama.cpp 或 OpenRouter 可访问的模型都可以作为候选者和评估者。此外,JudgeArena 提供了经过调优的开放模型评估者配置,这些配置在英语和多语言环境中的人类偏好数据集上经过验证,能够匹配或超越封闭模型评估者,从而减少对不透明封闭模型的依赖。最后,通过将现有的人类注释与目标模型的 LLM-评估结果相结合,JudgeArena 可以高精度地模拟 LMArena Elo 分数,为大规模人类注释活动提供一种实用、开放且低成本的替代方案。
cs.CL / 8 / 2608.02621
Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks
知其形,不知其用:自动审计法律基准中的答案与权威解耦
Abstract
Legal benchmarks typically score final answers even when models also state legal authority. We test whether answer correctness can serve as a proxy for authority grounding. Under ordinary reasoning prompts that did not request statutory citations, four LLMs spontaneously produced authority markers across 238 Taiwan bar-examination items. Because each item has a verified governing provision, we automatically audit answer correctness and authority grounding jointly. The two dimensions dissociate in both directions. In criminal law, 24.0--42.4\% of valid responses were answer-correct but missed the gold authority, while 15.2--21.7\% were answer-incorrect but cited it. A separate statutory-retrieval probe and a permissive citation-abstention intervention further show that answer and citation behavior can move separately at the output level. Because this mismatch arises without adversarial or inconsistency-inducing prompting, answer-only scoring treats naturally occurring gold-authority misses as complete benchmark successes. Because statutory authority is structurally extractable and externally verifiable, the failure can be measured automatically. A preliminary PRC civil-law extension also observes citation-unrequested authority marking, motivating a full cross-jurisdictional joint audit. We therefore propose joint answer--authority evaluation for statute-grounded legal benchmarks.
Chinese Translation
法律基准通常在模型同时陈述法律权威时仍对最终答案进行评分。我们测试答案的正确性是否可以作为权威基础的代理。在未请求法定引用的普通推理提示下,四个大型语言模型(LLMs)在238个台湾律师考试项目中自发产生了权威标记。由于每个项目都有经过验证的治理条款,我们自动联合审计答案的正确性和权威基础。这两个维度在两个方向上是分离的。在刑法中,24.0%至42.4%的有效回应虽然答案正确,但未引用黄金权威,而15.2%至21.7%的回应则答案不正确但引用了它。一个单独的法定检索探测和一个允许的引用不干预干预进一步表明,答案和引用行为在输出层面可以独立变化。由于这种不匹配是在没有对抗性或引发不一致的提示下产生的,答案仅评分将自然发生的黄金权威遗漏视为完整的基准成功。由于法定权威在结构上是可提取和可外部验证的,这种失败可以自动测量。初步的PRC民法扩展也观察到未请求的权威标记,促使进行全面的跨管辖区联合审计。因此,我们建议对基于法条的法律基准进行联合答案与权威评估。
cs.CL / 9 / 2608.02625
Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models
推测性修正:扩散语言模型的草拟-再细化解码
Abstract
Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right generation by producing text block by block. We study a simple plug-and-play inference pattern: first generate a complete draft, then refine the full response using bidirectional diffusion. Using LLaDA2.1-Flash and LLaDA2.1-Mini, we evaluate two configurations. In Flash-Flash, the same Flash model serves as both drafter and refiner, testing whether an existing model can improve its own block-autoregressive output through global refinement. In Mini-Flash, inspired by speculative decoding, we introduce speculative correction: Mini drafts a full response, and Flash revises it as an editable initialization. Flash-Flash improves GSM8K-384 accuracy from 0.848 to 0.899 while running 1.20 times faster than the selected Flash block-autoregressive baseline, and improves MBPP-384 from 0.545 to 0.693. Latency-window-matched Flash-only controls indicate that these gains persist after targeted tuning of block-autoregressive decoding. Causal ablations indicate that completed drafts provide useful initializations: refinement from a fully masked span performs poorly, full global refinement provides a clear additional gain on GSM8K, and local refinement captures much of the gain on MBPP and MATH. Mini-Flash provides useful quality-latency trade-offs, including MATH-384 performance of 0.294 versus 0.300 for Flash while running 2.17 times faster. These results support a Pareto-frontier interpretation rather than the claim that the heterogeneous cascade uniformly matches Flash quality. Overall, same-model draft-and-refine provides evidence that bidirectional refinement is a useful decoding primitive for DLMs, while speculative correction demonstrates a training-free route to fast DLM generation.
Chinese Translation
扩散语言模型(DLMs)能够双向修正标记,但标准解码程序通常通过逐块生成文本将其适应于从左到右的生成。我们研究了一种简单的即插即用推理模式:首先生成完整草稿,然后使用双向扩散对完整响应进行细化。使用 LLaDA2.1-Flash 和 LLaDA2.1-Mini,我们评估了两种配置。在 Flash-Flash 中,相同的 Flash 模型既作为草拟者又作为修正者,测试现有模型是否能通过全局细化改善其自身的块自回归输出。在 Mini-Flash 中,受到推测性解码的启发,我们引入了推测性修正:Mini 草拟完整响应,Flash 将其作为可编辑的初始化进行修正。Flash-Flash 将 GSM8K-384 的准确率从 0.848 提高到 0.899,同时运行速度比选定的 Flash 块自回归基线快 1.20 倍,并将 MBPP-384 从 0.545 提高到 0.693。与延迟窗口匹配的 Flash-only 控制表明,这些增益在对块自回归解码进行有针对性的调优后仍然存在。因果消融实验表明,完成的草稿提供了有用的初始化:来自完全掩蔽跨度的细化表现不佳,全面的全局细化在 GSM8K 上提供了明显的额外增益,而局部细化在 MBPP 和 MATH 上捕获了大部分增益。Mini-Flash 提供了有用的质量-延迟权衡,包括 MATH-384 性能为 0.294,而 Flash 为 0.300,同时运行速度快 2.17 倍。这些结果支持帕累托前沿的解释,而不是声称异质级联均匀匹配 Flash 质量。总体而言,同一模型的草拟与细化提供了证据,表明双向细化是 DLMs 的一种有用解码原语,而推测性修正展示了一种无训练的快速 DLM 生成途径。
cs.CL / 10 / 2608.02689
Stuck on "A": Diagnosing and Repairing Interface Injury in Attention-to-KDA Linearization of a 0.6B Language Model
困在“A”上:诊断和修复0.6B语言模型注意力到KDA线性化中的接口损伤
Abstract
We convert 21 of 28 full-attention layers of Qwen3-0.6B-Base into KDA (Kimi Delta Attention) linear-attention layers on a single consumer-grade GPU budget, and ask a simple question: what exactly does the conversion break? After surgery, hidden-state alignment and end-to-end KL distillation drive the student close to its teacher in perplexity, yet multiple-choice accuracy stays near random chance (25-29% vs. the teacher's 50.6% on C-Eval). Using a four-permutation diagnostic that rotates answer options while holding content fixed, we show the model sticks to option labels (predicting "A" 81% of the time; 106/161 questions keep the same label under all four rotations) rather than following answer content -- an interface injury that standard distillation metrics cannot see. A 1,000-step format-targeted completion-only KL stage repairs the interface (+12.48 points on C-Eval, label-stickiness roughly halved), after which persona SFT and one round of on-policy DPO preserve benchmark scores within noise. We release code, weights, recipes, and the full audit trail, and distill the engineering lessons -- including an FP32-master failure mode in which bf16 optimizer updates are silently swallowed -- that made convergence possible at this budget.
Chinese Translation
我们将Qwen3-0.6B-Base的28个全注意力层中的21个转换为KDA(Kimi Delta Attention)线性注意力层,预算为单个消费级GPU,并提出一个简单的问题:转换究竟破坏了什么?经过转换,隐藏状态对齐和端到端KL蒸馏使得学生模型在困惑度上接近其教师模型,然而多项选择准确率仍然接近随机机会(25-29%对比教师模型的50.6%在C-Eval上)。通过使用四重排列诊断方法,在保持内容不变的情况下旋转答案选项,我们发现模型更倾向于选择选项标签(81%的时间预测“A”;在所有四次旋转中,106/161个问题保持相同标签),而不是遵循答案内容——这是标准蒸馏指标无法察觉的接口损伤。一个针对格式的1,000步仅完成KL阶段修复了接口(在C-Eval上提高了12.48分,标签粘性大约减半),之后个性化SFT和一次政策DPO保持基准分数在噪声范围内。我们发布代码、权重、配方和完整审计记录,并提炼出工程经验教训——包括一个FP32主模式的失败模式,其中bf16优化器更新被静默吞噬——使得在此预算下的收敛成为可能。
cs.CL / 11 / 2608.02694
Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation
Crayotter:通过群体相对偏好反向传播学习长时间跨度的视频编辑代理
Abstract
Long-horizon video editing agents receive final-product feedback only after many interdependent decisions. Yet editing quality is subjective, admits multiple valid solutions, and is not meaningfully calibrated across heterogeneous requests, making a global scalar objective both ambiguous and temporally uninformative. Our key observation is that fixing the request, materials, and production constraints converts this subjective objective into an ordinal comparison among directly comparable alternatives. We introduce Group-Relative Preference Backpropagation (GRPB), which transforms same-task rankings into zero-sum advantages and redistributes them as bounded credit over semantic editing segments. A lagged allocator and guarded transmission prevent current judgments or unreliable estimates from directly shaping the same rollout group. We manually construct a project-disjoint, horizon-stratified suite of realistic editing tasks for training and controlled evaluation. Across matched baselines, credit interventions, external benchmarking, and blinded human evaluation, GRPB improves both editing behavior and rendered products. The resulting 9B Crayotter model surpasses several proprietary systems on AgenticVBench, supporting task-local preference reduction as a practical approach to learning from subjective, delayed outcomes. Code and all supporting materials are publicly available at https://github.com/idwts/Crayotter.
Chinese Translation
长时间跨度的视频编辑代理在做出许多相互依赖的决策后,才会收到最终产品的反馈。然而,编辑质量是主观的,承认多种有效解决方案,并且在异构请求之间没有有效的校准,这使得全局标量目标既模糊又在时间上缺乏信息。我们关键的观察是,固定请求、材料和生产约束将这一主观目标转化为直接可比较替代方案之间的序数比较。我们提出了群体相对偏好反向传播(Group-Relative Preference Backpropagation, GRPB),该方法将同一任务的排名转化为零和优势,并将其作为有界信用重新分配到语义编辑段上。滞后分配器和受保护的传输防止当前判断或不可靠估计直接影响同一回滚组。我们手动构建了一套项目不重叠、按时间跨度分层的现实编辑任务,以用于训练和控制评估。在匹配的基线、信用干预、外部基准测试和盲人评估中,GRPB改善了编辑行为和渲染产品。最终的9B Crayotter模型在AgenticVBench上超越了多个专有系统,支持任务局部偏好减少作为从主观延迟结果中学习的实用方法。代码和所有支持材料均可在https://github.com/idwts/Crayotter上公开获取。
cs.CL / 12 / 2608.02703
ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads
ARCHead:用于大型语言模型输出头的激活度量残差修正
Abstract
Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the final language-modeling head (LM-head) in BF16 or FP16. Quantizing this projection naively can strongly perturb the vocabulary-logit distribution. We present ARCHead, a packed LM-head compressor that combines a quantized low-rank core, group-wise INT4 residuals, and a low-rank correction fitted in an activation-derived metric. ARCHead stores no dense BF16 head and reduces persistent LM-head storage by 3.7-3.9x. On Qwen3-8B-Base, it uses 25.6% of BF16 head storage while attaining 1.007 relative perplexity; storage-matched naive INT4 yields 1.14-1.16. Replacing the BF16 head left by AWQ or bitsandbytes adds only 0.006-0.007 cross-entropy, with less than 2% throughput change in our measurements. ARCHead therefore complements block quantizers by compressing the large output projection they can leave untouched. Code is available at https://github.com/suayptalha/archead.
Chinese Translation
仅量化权重显著减少了大型语言模型(LLM)变换器块的存储,但实际后端通常将最终的语言建模头(LM-head)保留为BF16或FP16。天真地量化这个投影可能会强烈扰动词汇-logit分布。我们提出了ARCHead,这是一种紧凑的LM-head压缩器,它结合了量化的低秩核心、分组INT4残差和在激活导出度量中拟合的低秩修正。ARCHead不存储密集的BF16头,并将持久的LM-head存储减少了3.7-3.9倍。在Qwen3-8B-Base上,它使用了25.6%的BF16头存储,同时达到了1.007的相对困惑度;存储匹配的天真INT4则产生了1.14-1.16。用AWQ或bitsandbytes替换留下的BF16头仅增加了0.006-0.007的交叉熵,在我们的测量中吞吐量变化不到2%。因此,ARCHead通过压缩它们可以保持不变的大型输出投影,补充了块量化器。代码可在https://github.com/suayptalha/archead获取。
cs.CL / 13 / 2608.02807
Learning a Vector-Symbolic Model for Socio-Cultural Tasks
学习用于社会文化任务的向量符号模型
Abstract
How can we better represent the impact of sociocultural structures on decision making in computational cognitive models? Modeling this impact requires traversing multiple levels of semantic representation, however it is not immediately clear to a modeler which levels of representation are most salient to a given situation. Though large language models and cognitively grounded corpus models can represent broad semantic associations through co-occurences, the role of self representations in memory should be accounted for to determine how cultural associations shape decision making. We propose a declarative memory system to be used in the ACT-R cognitive architecture that represents semantic associations at multiple levels via a vector-symbolic autoencoder. We use a simple HRR operation to encode episodic memories differently from semantic memory vectors extracted from text to produce a final chunk activation for a memory request. We use ACT-R cognitive models of a racially contextualized implicit association test (IAT) to test this new declarative memory system.
Chinese Translation
我们如何更好地表示社会文化结构对计算认知模型中决策制定的影响?建模这一影响需要跨越多个语义表示层次,然而对于建模者来说,哪些表示层次在特定情境中最为显著并不立即清晰。尽管大型语言模型和基于认知的语料库模型可以通过共现来表示广泛的语义关联,但在确定文化关联如何塑造决策制定时,应该考虑自我表征在记忆中的作用。我们提出了一种声明性记忆系统,应用于ACT-R(Adaptive Control of Thought—Rational)认知架构,通过向量符号自编码器在多个层次上表示语义关联。我们使用简单的HRR(Holographic Reduced Representation)操作将情节记忆与从文本中提取的语义记忆向量进行不同编码,以生成记忆请求的最终块激活。我们利用ACT-R认知模型对种族情境化的隐性联想测试(IAT)进行实验,以验证这一新的声明性记忆系统。
cs.CL / 14 / 2608.02867
BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?
BODHI:大语言模型是否能够分支并发现异质推理?
Abstract
Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variety of reasoning tasks, there is significant debate as to whether RLVR expands the reasoning capability boundary, or just improves sampling efficiency. In this paper, we investigate the nature of test-time exploration in RLVR-trained LLMs by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence. This helps us delineate between entropy arising from stylistic variations and genuine inferential branching. Our findings demonstrate that the policy entropy collapse observed in RLVR models is not merely syntactic, and is accompanied by a significant reduction in semantic branching entropy. While RLVR improves adherence to environmental constraints and backtracking capabilities, it constricts the space of continuations; we provide evidence suggesting that this might be responsible for the sample efficiency gains of RLVR, albeit at the cost of genuine rollout diversity.
Chinese Translation
尽管具有可验证奖励的强化学习(RLVR)在多种推理任务中提高了大语言模型(LLMs)的性能,但关于RLVR是否扩展了推理能力的边界,还是仅仅提高了采样效率,仍存在显著争议。本文通过采用受控的迷宫求解实验,研究了RLVR训练的大语言模型在测试时探索的性质,并基于语义等价性提取了数学推理轨迹中的树结构(BODHI-Trees)。这有助于我们区分由风格变异引起的熵和真正的推理分支。我们的研究结果表明,在RLVR模型中观察到的策略熵崩溃不仅仅是句法上的,并且伴随着语义分支熵的显著减少。虽然RLVR提高了对环境约束和回溯能力的遵循,但它限制了延续的空间;我们提供的证据表明,这可能是RLVR样本效率提高的原因,尽管这以牺牲真正的展开多样性为代价。
cs.CL / 15 / 2608.02919
FLARE: Few-shot Learning-based Adaptive Reflective Engine
FLARE:基于少样本学习的自适应反射引擎
Abstract
Large language models (LLMs) are increasingly deployed in complex, compound AI systems where performance hinges on the quality of prompts. Recent state-of-the-art optimizers like GEPA (Genetic-Pareto) have argued that reflective instruction evolution can outperform traditional reinforcement learning and few-shot optimization. In this work, we challenge this shift by introducing FLARE (Few-shot Learning-based Adaptive Reflective Engine), a framework that leverages advanced reflective mechanisms and a small set of few-shot reference examples to optimize instructions. We evaluate our method across a diverse suite of benchmarks -- spanning retrieval-augmented reasoning (HotPotQA, MedQA, 2WikiMultiHopQA), tool calling, and multi-label emotion classification (GoEmotions) -- using the GPT-5 series of models. Our results demonstrate that FLARE consistently outperforms GEPA, winning on every task-model pair: it achieves gains of up to +14.2 points on HotPotQA (52.2 vs. GEPA's 42.2 with GPT-5-Chat), reaches 87.0% on tool calling (vs. 81.0% for GEPA), and lifts GoEmotions micro-F1 to 52.7% (+15.3) with GPT-5.1 on the full 5408-example test split, more than doubling GEPA's +5.7 gain. Beyond raw accuracy, FLARE is also strikingly data-efficient: on GoEmotions it reaches its peak performance using as few as 100 validation examples, while remaining markedly more stable across random seeds than GEPA. Our findings suggest that while reflective instructions are powerful, the strategic optimization of few-shot learning remains a critical frontier for maximizing the potential of next-generation LLMs.
Chinese Translation
大型语言模型(LLMs)越来越多地被部署在复杂的复合人工智能系统中,其性能依赖于提示的质量。近期的最先进优化器如GEPA(遗传-帕累托)提出反射指令演化可以超越传统的强化学习和少样本优化。在本研究中,我们通过引入FLARE(基于少样本学习的自适应反射引擎)这一框架来挑战这一转变,该框架利用先进的反射机制和一小组少样本参考示例来优化指令。我们在一系列多样化的基准测试中评估了我们的方法——涵盖了检索增强推理(HotPotQA、MedQA、2WikiMultiHopQA)、工具调用和多标签情感分类(GoEmotions)——使用GPT-5系列模型。我们的结果表明,FLARE在每个任务-模型对中始终优于GEPA:在HotPotQA上获得高达+14.2分的提升(52.2对比GEPA的42.2,使用GPT-5-Chat),在工具调用上达到87.0%(对比GEPA的81.0%),并在完整的5408个示例测试集上将GoEmotions的微F1提升至52.7%(+15.3),是GEPA的+5.7提升的两倍多。除了原始准确性外,FLARE在数据效率上也表现出色:在GoEmotions上,它在仅使用100个验证示例的情况下达到了最佳性能,同时在随机种子间的稳定性明显优于GEPA。我们的研究结果表明,尽管反射指令非常强大,但少样本学习的战略优化仍然是最大化下一代LLMs潜力的关键前沿。
cs.CL / 16 / 2608.02935
Character Iconicity vs. Arbitrariness: An Arabic NLP Perspective
字符图标性与任意性:阿拉伯语自然语言处理的视角
Abstract
Arabic script uses 28 letters, many of which share a common base shape (rasm) and are distinguished only by dot placement. Because early Arabic manuscripts were written without dots yet remained interpretable, dot removal offers a natural test of whether these visual distinctions are functionally necessary. Prior work has shown that dotless Arabic can remain readable and effective for natural language processing (NLP), but it remains unclear whether this success depends on preserving the original rasm groupings or whether arbitrary but consistent remappings to the same reduced rasm set can achieve comparable performance. We address this question by comparing standard dotted and dotless Arabic with arbitrary character remappings constrained to the same 19 undotted rasms. We generated 2,000 random remappings under word- and character-level tokenization and selected four representative mappings with the highest and lowest entropy values. These representations were evaluated across language modeling, text classification, sequence labeling, machine translation, and restoration to the original script. The results show that neither preserving original character distinctions nor retaining traditional rasm-based groupings is necessary for strong NLP performance. Random remappings achieve competitive performance while reducing vocabulary size, out-of-vocabulary (OOV) rates, model size, and training cost. These findings suggest that, from an NLP perspective, Arabic character form-function relationships are largely arbitrary: models rely more on stable distributional structure than on the visual iconicity of letter forms.
Chinese Translation
阿拉伯字母使用28个字母,其中许多字母共享一个共同的基本形状(rasm),仅通过点的位置来区分。由于早期的阿拉伯手稿在没有点的情况下书写,但仍然可以被理解,因此去掉点提供了一个自然的测试,来检验这些视觉区分是否在功能上是必要的。先前的研究表明,去点的阿拉伯语仍然可以保持可读性,并且在自然语言处理(NLP)中有效,但尚不清楚这种成功是否依赖于保留原始的rasm分组,或者是否可以通过任意但一致的重新映射到相同的简化rasm集合来实现可比的性能。我们通过比较标准的带点和去点阿拉伯语与限制在相同19个无点rasm的任意字符重新映射来解决这个问题。我们在词级和字符级标记化下生成了2000个随机重新映射,并选择了四个具有最高和最低熵值的代表性映射。这些表示在语言建模、文本分类、序列标注、机器翻译和恢复到原始脚本的任务中进行了评估。结果表明,既不需要保留原始字符的区分,也不需要保留基于传统rasm的分组,以实现强大的NLP性能。随机重新映射在减少词汇大小、超出词汇(OOV)率、模型大小和训练成本的同时,仍能达到竞争性的性能。这些发现表明,从NLP的角度来看,阿拉伯字符形态与功能的关系在很大程度上是任意的:模型更多依赖于稳定的分布结构,而不是字母形态的视觉图标性。
cs.CL / 17 / 2608.02941
Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech
形式一致而非意义一致:低资源孟加拉语贬损言论中大型语言模型安全性的理解-控制解耦
Abstract
We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-Containment Decoupling. We propose that contemporary safety alignment is bound to high-resource surface forms rather than harmful meaning, causing a model's capacity to comprehend a low-resource slur and its capacity to contain it to operate independently. Every protocol corroborates this hypothesis against a human-calibrated baseline (kappa = 0.84). At baseline, models exhibit a 7.92 percentage point comprehension deficit in Bangla while maintaining an identical 92.83% token leakage rate across both languages. Severity calibration tracks surface anatomical cues over compositional harm (+4.00 error on mild slang; -2.00 on threats), while apparent containment gains under orthographic perturbation prove to be a tokenizer-driven "containment mirage." Crucially, explicit Chain-of-Thought reasoning rescues comprehension (94.72% Pass) while systematically dismantling containment (96.23% Use). Furthermore, expert-persona framing collapses refusal to 6.57%, revealing that keyword-based filters ignore dehumanizing communal slurs entirely. Our findings demonstrate that high-resource benchmarks cannot certify low-resource safety, necessitating meaning-grounded containment.
Chinese Translation
我们对五个前沿大型语言模型在本土孟加拉语贬损言论(gali)上的表现进行了审计,采用六种协议来测试一个单一假设:理解-控制解耦。我们提出,当代安全对齐是依赖于高资源表面形式而非有害意义,这导致模型理解低资源侮辱性用语的能力与控制其能力独立运作。每个协议都支持这一假设,相较于经过人类校准的基线(kappa = 0.84)。在基线条件下,模型在孟加拉语中的理解能力存在7.92个百分点的不足,而在两种语言中保持相同的92.83%标记泄漏率。严重性校准跟踪表面解剖线索与组合性伤害之间的关系(对轻微俚语的错误为+4.00;对威胁的错误为-2.00),而在正字法扰动下显现的明显控制增益被证明是由分词器驱动的“控制幻影”。关键是,明确的思维链推理拯救了理解能力(94.72% 通过率),同时系统性地拆解了控制能力(96.23% 使用率)。此外,专家角色框架将拒绝率压缩至6.57%,揭示基于关键词的过滤器完全忽视了去人性化的群体侮辱。我们的发现表明,高资源基准无法认证低资源的安全性,因此需要基于意义的控制。
cs.CL / 18 / 2608.02942
OPTD: On-Policy Transition Distillation with Consistency-Guided Adaptive Compression for Few-Step Diffusion Language Models
OPTD:基于一致性引导自适应压缩的在政策过渡蒸馏用于少步扩散语言模型
Abstract
Diffusion language models (dLLMs) can predict many tokens in parallel, but accurate generation still requires many iterative denoising steps. Few-step distillation accelerates decoding by compressing multiple teacher steps into a single student transition. However, existing methods construct supervision on off-policy trajectories. At inference, the student's early parallel commitments alter the context of later predictions, so the states it actually visits drift away from the supervised ones--precisely when step compression is most aggressive. On-policy distillation is a natural remedy for this mismatch, but it leaves open how far each transition should advance: matching only the teacher's next action limits compression, while indiscriminately merging future actions can violate intermediate dependencies. To address this limitation, we propose OPTD, On-Policy Transition Distillation with consistency-guided adaptive compression. It samples partial states from the few-step student's own trajectories, uses a frozen, question-only teacher to identify outcome-aligned future candidates, and orders them by current-state confidence. The method then selects the longest prefix whose joint commitment preserves the teacher's rollout outcome. A set-bottleneck objective promotes every verified future candidate to the decoder's release threshold, while a frozen-teacher KL anchor regularizes all other active positions. Neither target construction nor training uses a gold response. Across four mathematical reasoning and code-generation benchmarks, OPTD consistently improves the quality--efficiency trade-off and attains the strongest overall quality-constrained AUP among the evaluated few-step baselines.
Chinese Translation
扩散语言模型(dLLMs)能够并行预测多个标记,但准确生成仍然需要多个迭代去噪步骤。少步蒸馏通过将多个教师步骤压缩为单个学生过渡来加速解码。然而,现有方法在离政策轨迹上构建监督。在推理过程中,学生的早期并行承诺改变了后续预测的上下文,因此它实际访问的状态偏离了监督状态——恰好是在步骤压缩最激进的时候。在政策蒸馏是解决这种不匹配的自然方法,但它仍然留有如何推进每个过渡的空间:仅匹配教师的下一个动作会限制压缩,而不加区分地合并未来动作可能会违反中间依赖关系。为了解决这一限制,我们提出了OPTD,即基于一致性引导自适应压缩的在政策过渡蒸馏。它从少步学生自己的轨迹中采样部分状态,使用一个冻结的仅提问教师来识别与结果对齐的未来候选,并根据当前状态的置信度对它们进行排序。该方法随后选择最长的前缀,其联合承诺保留教师的展开结果。一组瓶颈目标促进每个经过验证的未来候选达到解码器的释放阈值,同时冻结教师的KL锚点对所有其他活动位置进行正则化。目标构建和训练均未使用黄金响应。在四个数学推理和代码生成基准测试中,OPTD始终改善了质量与效率的权衡,并在评估的少步基准中获得了最强的整体质量约束AUP。
cs.CL / 19 / 2608.02966
Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks
每个错误答案都重要:针对大型语言模型的选项级心理测量多项选择基准
Abstract
Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect responses alike, even though an LLM's preferences among incorrect options may contain systematic and useful information about its behavior and ability. We introduce the LLM Nominal Response Model (LLM-NRM), an option-aware psychometric framework that models the full distribution over answer choices to jointly estimate LLM ability and option-level item characteristics, while separating model-specific response calibration sharpness, positional preference, and difficulty-dependent fallback behavior. Across 189 LLMs and 31,554 items from 14 benchmarks, LLM-NRM predicts held-out LLM-item interactions more accurately than binary Item Response models and conventional nominal-response baselines, and its ability estimates achieve the strongest Spearman correlation of 0.920 with the external human-preference Arena.ai Elo leaderboard. Distractor identity contributes +101% additional Fisher Information per item beyond correctness, and incorrect responses alone recover full-information ability estimates with Spearman 0.943. The learned item parameters also enable efficient benchmarking, where 41 selected items preserve the full-bank ranking with Kendall's correlation 0.85, corresponding to a 770 times reduction. In conclusion, we show that incorrect answers carry distinct and useful measurement information rather than representing equivalent mistakes.
Chinese Translation
大多数多项选择题(MCQ)基准仅通过选择正确答案来评估大型语言模型(LLMs)。这种二元评分将所有错误回答视为相同,尽管LLM在错误选项之间的偏好可能包含关于其行为和能力的系统性和有用信息。我们引入了LLM名义响应模型(LLM-NRM),这是一个考虑选项的心理测量框架,能够对答案选择的完整分布进行建模,以联合估计LLM的能力和选项级项目特征,同时区分模型特定的响应校准精度、位置偏好和依赖于难度的后备行为。在来自14个基准的189个LLM和31,554个项目中,LLM-NRM比二元项目响应模型和传统的名义响应基线更准确地预测了保留的LLM-项目交互,其能力估计与外部人类偏好Arena.ai Elo排行榜的斯皮尔曼相关系数达到0.920。干扰项的身份在正确性之外为每个项目贡献了+101%的额外费舍信息,仅凭错误响应就能恢复全信息能力估计,斯皮尔曼相关系数为0.943。学习到的项目参数还支持高效基准测试,其中41个选定项目保留了完整库的排名,肯德尔相关系数为0.85,相当于770倍的减少。总之,我们表明,错误答案携带独特且有用的测量信息,而不是代表等同的错误。
cs.CL / 20 / 2608.02971
Mapping the City Through the Lens of Language Models
通过语言模型的视角映射城市
Abstract
Language models often complete an underspecified reference to a city with unstated assumptions about urban size, form, infrastructure, environment, and function. We measure those assumptions without naming places. Ten open-weight checkpoints rate anonymized profiles derived from real morphological urban centres across 40 audited indicators and seven domains. The design combines constrained probability-based ratings, prespecified reliability screens, lineage-aware aggregation, multiple population weightings, an independent replication sample, and whole-profile validation. The clearest shared tendency favours urban profiles with larger developed area, faster recent growth, greater mapped infrastructure and non-residential capacity, and less sparse form. Most eligible directions recur in the replication data, and direct ratings of complete profiles show moderate agreement with the indicator-wise construction. Geographic differences shrink after accounting for city scale and development, while reliably measured paired tasks indicate that typicality and desirability are often closely aligned. The framework makes an otherwise vague notion of what models regard as an ordinary city empirically traceable. The resulting evidence delineates a shared yet model-dependent portrait of the city through the lens of language models.
Chinese Translation
语言模型通常在对城市的模糊引用中,伴随未明言的假设,涉及城市的规模、形态、基础设施、环境和功能。我们在不命名地点的情况下测量这些假设。十个开放权重的检查点对来自真实形态城市中心的匿名档案进行评分,这些档案基于40个审计指标和七个领域。该设计结合了基于约束概率的评分、预设的可靠性筛选、考虑谱系的聚合、多种人口加权、独立复制样本以及整体档案验证。最明显的共同趋势偏向于具有更大开发面积、最近增长更快、基础设施和非住宅容量更大、以及形态更不稀疏的城市档案。大多数符合条件的方向在复制数据中反复出现,完整档案的直接评分与指标逐项构建显示出中等一致性。在考虑城市规模和发展后,地理差异缩小,而可靠测量的配对任务表明,典型性和可取性往往紧密相关。该框架使得模型所认为的普通城市这一模糊概念可以通过实证方式追踪。最终的证据描绘了一个通过语言模型视角看待城市的共享但依赖模型的肖像。
cs.CL / 21 / 2608.02975
TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation
TQLite:多LLM陪审团引导的蒸馏方法用于实时MQM翻译质量评估
Abstract
Large language models (LLMs) have demonstrated impressive performance in MQM-based translation quality (TQ) evaluation, and recent advances in large reasoning models (LRMs) promise even greater improvements. However, both LLMs and LRMs are computationally expensive to deploy at scale, while small language models (SLMs)---though much more efficient---struggle with the complex reasoning required for evaluation tasks. In this work, we present an extensive empirical study benchmarking SLMs, LLMs, and LRMs across a wide range of TQ evaluation setups, providing a comprehensive view of the current landscape and establishing best practices. To address the scalability challenge, we introduce TQLite, a novel distillation framework that enables SLMs to approach the MQM evaluation performance of the best LRM-based evaluators. Our approach leverages a multi-LRM jury to generate high-quality synthetic training data via practical data curation techniques and aggregation of evaluation responses across a diverse panel of models. Our results demonstrate that SLMs trained via TQLite achieve strong MQM evaluation performance that far exceeds off-the-shelf evaluation capabilities of standard SLMs, offering a scalable and cost-effective alternative to LLM- and LRM-based evaluators.
Chinese Translation
大型语言模型(LLMs)在基于MQM的翻译质量(TQ)评估中表现出色,最近的大型推理模型(LRMs)的进展更是预示着更大的提升。然而,LLMs和LRMs在大规模部署时计算成本高昂,而小型语言模型(SLMs)虽然效率更高,却在评估任务所需的复杂推理方面面临挑战。在本研究中,我们进行了广泛的实证研究,对SLMs、LLMs和LRMs在多种TQ评估设置下进行了基准测试,提供了当前领域的全面视角,并建立了最佳实践。为了应对可扩展性挑战,我们提出了TQLite,这是一种新颖的蒸馏框架,使SLMs能够接近基于最佳LRM评估者的MQM评估性能。我们的方法利用多LRM陪审团通过实用的数据策划技术和对多样化模型面板的评估响应聚合生成高质量的合成训练数据。我们的结果表明,通过TQLite训练的SLMs在MQM评估性能上表现强劲,远超标准SLMs的现成评估能力,提供了一种可扩展且具有成本效益的替代方案,优于基于LLM和LRM的评估者。
cs.CL / 22 / 2608.02999
On the Non-Specificity of Statistical Measures Used in Script Decipherment
关于用于文本破译的统计测量的非特异性
Abstract
Statistical regularities are routinely offered as evidence that undeciphered sign systems encode language; the Indus script debate is the canonical example. Any such inference rests on specificity: the reported outcome must be unusual among plausible structured non-languages. We test that premise constructively with SIGIL, a purpose-built generative emblem system whose 3,000-text core corpus carries explicit compositional meanings although no sign has a phonological value. A literature registry compiled in advance of evaluation records 54 methods and admits a method to exact scoring when both the published Indus outcome and a source-defined decision rule can be reproduced. SIGIL receives the same category as the Indus corpus on every criterion scored this way, across repetition, directional-asymmetry, and lexical-distribution tests. Declared reconstructions of entropy, frequency, positional, predictive, classifier, and network measures reproduce the familiar Indus-like signatures as well. A sequential decipherment stress test then reaches high dictionary coverage for English, Sanskrit, and Tamil on the same corpus, while grouped held-out declines and unstable keys reveal how little that coverage identifies. The construction does not decide what the Indus signs encode: it shows that the evaluated measures detect organization without being specific to language, and therefore cannot, on their own, establish encoded speech.
Chinese Translation
统计规律常被作为证据,表明未破译的符号系统编码语言;印度河流域文字的争论便是经典例子。任何此类推断都依赖于特异性:报告的结果必须在合理的结构化非语言中显得不寻常。我们通过SIGIL这一专门构建的生成性标志系统来建设性地测试这一前提,其核心语料库包含3000个文本,尽管没有符号具有音位值,但却承载着明确的组合意义。在评估之前编制的文献登记中记录了54种方法,并允许在发布的印度河流域结果和源定义的决策规则均可重现时进行精确评分。SIGIL在每个评分标准上与印度河流域语料库属于同一类别,涵盖重复性、方向不对称性和词汇分布测试。声明的熵、频率、位置、预测、分类器和网络测量的重建同样再现了熟悉的印度河流域特征。随后进行的序列破译压力测试在同一语料库上达到了英语、梵语和泰米尔语的高词典覆盖率,而分组的保留下降和不稳定的密钥则揭示了这种覆盖率识别的有限性。该构造并未决定印度河符号编码的内容:它表明所评估的测量能够检测组织结构,但并不特定于语言,因此无法单独确立编码的言语。
cs.CL / 23 / 2608.03035
Language Models Encode the Contextual Truth of Propositions
语言模型编码命题的上下文真理
Abstract
Prior work has shown that LLMs encode the truth of factual propositions along linear directions in activation space. It's unclear how these representations extend to contextual truth: propositions whose truth is determined by in-context evidence rather than world knowledge. We show that LLMs maintain a linear representation of contextual truth that persists across structurally different output policies, even when the output doesn't require the model to determine a proposition's truth, and show causal evidence via steering experiments. Using the transcripts from a collaborative vision-language task that requires two LLMs to maintain a shared common ground, we show that truth representations of a proposition are significantly swayed by partner assertions about that proposition, even when the LLM has enough evidence to determine its truth. We find evidence that propositions near the decision boundary are more susceptible to having their truth shifted through partner assertions. Separating representation from output distinguish two forms of sycophancy that output behavior alone cannot: the model may accommodate a false proposition while continuing to represent it as false, or shift its representation across the boundary. The latter is 2.59x more common when the model agrees by restating the false claim explicitly than when it agrees implicitly.
Chinese Translation
先前的研究表明,大型语言模型(LLMs)在激活空间中沿线性方向编码事实命题的真理。然而,这些表示如何扩展到上下文真理尚不清楚:上下文真理是指其真理由上下文证据而非世界知识决定的命题。我们展示了LLMs保持一种线性表示的上下文真理,这种表示在结构上不同的输出策略中持续存在,即使输出并不要求模型确定命题的真理,并通过引导实验展示了因果证据。通过使用一个协作视觉-语言任务的转录,该任务要求两个LLMs维持共享的共同基础,我们展示了命题的真理表示受到合作伙伴对该命题的主张的显著影响,即使当LLM有足够的证据来确定其真理时。我们发现,接近决策边界的命题更容易受到合作伙伴主张的影响而改变其真理。将表示与输出分离可以区分两种输出行为无法区分的谄媚形式:模型可能在继续将一个虚假命题表示为虚假的同时,接受该虚假命题,或者在边界上改变其表示。当模型通过明确重述虚假主张而同意时,后者的发生频率是前者的2.59倍。
cs.CL / 24 / 2608.03038
Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models
超越准确性:对大型语言模型统计推理的多维评估
Abstract
Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how models construct and communicate statistical explanations. This study demonstrates the value of a multidimensional evaluation by combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis. The framework is applied to explanations generated by 15 current-generation LLMs responding to 90 questions drawn from four statistics examinations spanning high school, undergraduate, and graduate levels. Accuracy varied substantially across models, ranging from 55\% to 78\%. In contrast, structural topic modeling revealed a common conceptual organization of statistical reasoning across all models, while lexical similarity analysis identified modest but consistent vendor-specific differences in explanatory style. Models developed by the same vendor (e.g. Anthropic, OpenAI) produced explanations that were slightly more similar than models from different vendors. These findings demonstrate that statistical reasoning in contemporary LLMs cannot be characterized by accuracy alone and illustrate how complementary analyses of response behavior and model-generated explanations provide a more comprehensive evaluation of statistical reasoning in generative AI.
Chinese Translation
统计推理是多维的,然而对大型语言模型(LLMs)的评估通常强调响应的准确性,而忽视了模型如何构建和传达统计解释。本研究通过结合响应准确性、响应行为、结构主题建模和词汇相似性分析,展示了多维评估的价值。该框架应用于15个当前一代LLMs对来自四个统计考试(涵盖高中、本科和研究生水平)的90个问题生成的解释。模型的准确性差异显著,范围从55%到78%。相比之下,结构主题建模揭示了所有模型在统计推理上的共同概念组织,而词汇相似性分析则识别出解释风格中适度但一致的供应商特定差异。由同一供应商(例如,Anthropic、OpenAI)开发的模型生成的解释在相似性上略高于来自不同供应商的模型。这些发现表明,当代LLMs中的统计推理不能仅仅通过准确性来表征,并说明了对响应行为和模型生成的解释进行互补分析如何提供对生成性人工智能中统计推理的更全面评估。
cs.CL / 25 / 2608.03044
Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation
模拟还是估计?基础模型与后训练模型在意见模拟中的不同优势
Abstract
Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity. We show that much of this conflict stems from conflating two distinct tasks. We call the first task emulation, in which models generate individual responses that aggregate into a population distribution. We call the second task estimation, in which models directly predict the population distribution. Evaluating six matched base and post-trained models on the Pew American Trends Panel, we find that base models are stronger emulators: they produce response distributions closer to human ground truth and better preserve demographic structure. Post-trained models are stronger estimators, producing more accurate distributional predictions when asked directly. We propose that model selection for human simulation should be guided by whether the task requires generating text or predicting distributions.
Chinese Translation
大型语言模型越来越多地用于模拟人类意见,但先前的研究报告了相互矛盾的结果:一些研究发现与人类调查数据的对齐表现良好,而另一些则发现个性崩溃和弱的人口敏感性。我们表明,这种冲突的主要原因在于混淆了两个不同的任务。我们将第一个任务称为模拟(emulation),在该任务中,模型生成个体响应,这些响应汇聚成一个人口分布。我们将第二个任务称为估计(estimation),在该任务中,模型直接预测人口分布。通过在皮尤美国趋势面板(Pew American Trends Panel)上评估六个匹配的基础模型和后训练模型,我们发现基础模型是更强的模拟者:它们生成的响应分布更接近人类的真实情况,并且更好地保留了人口结构。后训练模型则是更强的估计者,当直接要求时,它们生成的分布预测更为准确。我们建议,在进行人类模拟的模型选择时,应根据任务是需要生成文本还是预测分布来指导。
cs.CL / 26 / 2608.03048
PI-Mem: Pushing Long-Context Reasoning to 3.6M Tokens with Parallel-Iterative Memory
PI-Mem:通过并行迭代记忆将长上下文推理扩展至360万标记
Abstract
Long-context reasoning remains a critical bottleneck for large language models, as recent recurrent-memory approaches face two inherent challenges: sequential chunk-wise updates can overwrite early critical evidence with later irrelevant content, and serial inter-chunk dependencies limit parallelism and cause latency to increase with context length. To address these issues, we propose PI-Mem (Parallel-Iterative Memory), a mechanism that processes all chunks in parallel and iteratively refines a shared memory over a bounded number of turns. In each turn, PI-Mem reads all chunks in parallel conditioned on the current memory, selects new or complementary evidence from each chunk, and merges the selected evidence into a compact shared memory for the next turn. To discourage redundant turns, we optimize the workflow through reinforcement learning with an auxiliary turn-efficiency reward, enabling the model to adaptively exit once sufficient evidence has been accumulated. We evaluate PI-Mem with Qwen3.5-35B-A3B and Qwen2.5-7B on the HotpotQA benchmark across context lengths up to 3.6 million tokens and find that it outperforms the recurrent-memory baseline by +6.25 and +7.81 absolute points while achieving 6.1$\times$ and 2.1$\times$ inference speedups, respectively. These results demonstrate that PI-Mem breaks the accuracy--efficiency trade-off in long-context reasoning and provides a scalable approach to complex multi-hop question answering over extremely long documents.
Chinese Translation
长上下文推理仍然是大型语言模型的一个关键瓶颈,因为最近的递归记忆方法面临两个固有挑战:顺序分块更新可能会用后来的无关内容覆盖早期的重要证据,而块间的串行依赖限制了并行性,并导致随着上下文长度的增加延迟增加。为了解决这些问题,我们提出了PI-Mem(Parallel-Iterative Memory),一种在有限轮次内并行处理所有块并迭代优化共享记忆的机制。在每一轮中,PI-Mem根据当前记忆并行读取所有块,从每个块中选择新的或互补的证据,并将所选证据合并到下一个轮次的紧凑共享记忆中。为了抑制冗余轮次,我们通过强化学习优化工作流程,设置辅助的轮次效率奖励,使模型能够在积累到足够证据后自适应退出。我们在HotpotQA基准上使用Qwen3.5-35B-A3B和Qwen2.5-7B对PI-Mem进行评估,覆盖上下文长度高达360万标记,发现其在准确性上比递归记忆基线提高了+6.25和+7.81个绝对点,同时分别实现了6.1倍和2.1倍的推理速度提升。这些结果表明,PI-Mem打破了长上下文推理中的准确性与效率的权衡,并提供了一种可扩展的方法来处理极长文档中的复杂多跳问答。
cs.CL / 27 / 2608.03063
SeqLLM: Augmenting LLMs with Behavioral-Sequence Modeling for High-Stakes Decisions at WeChat Pay
SeqLLM:通过行为序列建模增强大语言模型在微信支付中的高风险决策能力
Abstract
Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false negatives leave harmful activity undetected. The hardest cases require jointly understanding a merchant's textual profile and long behavioral sequence. Large language models (LLMs) excel at text but cannot natively model such sequences, while adapting them often causes catastrophic forgetting. We present SeqLLM, a framework that adds behavioral-sequence modeling to a pretrained LLM while preserving its language ability. SeqLLM combines three components: a compact discrete vocabulary that represents behavioral events as native tokens; a lightweight projector, trained with a two-stage alignment curriculum, that grounds these tokens in the LLM's semantic space; and prefix-guided capability injection, which acquires sequence-modeling ability through task-prefixed supervised fine-tuning rather than continual pre-training. SeqLLM is deployed at WeChat Pay, screening millions of merchants daily. Against the production DeepSeek-based LLM baseline, it raises screening precision from 92.0% to 97.5%. Its pretrained behavior-token embeddings also improve
[email protected]% by 26.8 percentage points in a production fraud detector serving billion-scale transaction traffic. Beyond payments, SeqLLM achieves state-of-the-art results on public recommendation benchmarks. On MovieLens and Amazon, it surpasses the strong User-LLM baseline by up to 32% relative Recall@5 while retaining markedly stronger language ability. On RecIF, it improves Pass@32 by 14.2% over the full OneRec-8B pipeline using only one-fifth of its GPU-days.
Chinese Translation
在大型支付平台上,商户风险控制每天需要筛选数千万个商户,其中误报会对合法商户造成损害,而漏报则会使有害活动未被发现。最棘手的案例需要共同理解商户的文本特征和长期行为序列。大型语言模型(LLMs)在文本处理方面表现出色,但无法原生建模此类序列,而对其进行适应通常会导致灾难性遗忘。我们提出了SeqLLM,一个在保留语言能力的同时将行为序列建模添加到预训练LLM中的框架。SeqLLM结合了三个组件:一个紧凑的离散词汇表,将行为事件表示为原生标记;一个轻量级投影器,通过两阶段对齐课程进行训练,将这些标记嵌入LLM的语义空间;以及前缀引导的能力注入,通过任务前缀的监督微调而非持续预训练来获取序列建模能力。SeqLLM已在微信支付部署,每天筛选数百万个商户。与基于DeepSeek的生产LLM基线相比,其筛选精度从92.0%提高到97.5%。其预训练的行为标记嵌入也使得在服务于亿级交易流量的生产欺诈检测器中,
[email protected]%提高了26.8个百分点。除了支付领域,SeqLLM在公共推荐基准测试中也取得了最先进的结果。在MovieLens和Amazon上,其相对Recall@5比强大的User-LLM基线高出多达32%,同时保持了显著更强的语言能力。在RecIF上,使用仅五分之一的GPU天数,SeqLLM在完整的OneRec-8B管道上将Pass@32提高了14.2%。
cs.CL / 28 / 2608.03067
Activation-Guided Neuron Intervention to Induce Alzheimer's-Related Computational Language Phenotypes in a Large Language Model
激活引导的神经干预以诱导大型语言模型中的阿尔茨海默病相关计算语言表型
Abstract
Changes in spontaneous speech provide an early signal of cognitive dysfunction in Alzheimer's disease (AD) that large language models (LLMs) can detect. However, detection alone cannot establish whether the underlying model representations contribute functionally to behavior. We introduce an activation-guided intervention framework using Qwen3-8B. The framework identifies feed-forward neurons with higher activation rates for AD than control transcripts and modulates their output contributions during generation by scaling the corresponding down-projection weights. This yielded nine edited variants differing in intervention direction, magnitude, and scope. The original and edited models completed the same 12-turn neuropsychological battery, assessed through blinded human ratings and computational linguistic measures. Amplifying AD-associated neurons produced graded impairments in story recall, verbal fluency, working memory, procedural discourse, scene construction, and coreference resolution. Attenuation largely preserved performance and selectively improved several outcomes. Amplification also reduced lexical surprisal, idea density, syntactic complexity, and discourse quantity, broadly paralleling changes reported in human AD speech. These findings show that neurons identified solely from clinical language differences can influence behavior across multiple cognitive domains, providing proof of concept for an AD-related computational phenotype and a controlled framework for experimentally examining links between language and broader cognitive dysfunction.
Chinese Translation
自发言语的变化为阿尔茨海默病(AD)中的认知功能障碍提供了早期信号,而大型语言模型(LLMs)能够检测到这种变化。然而,仅仅检测无法确定基础模型表示是否在功能上对行为产生影响。我们引入了一种基于激活引导的干预框架,使用 Qwen3-8B。该框架识别出在阿尔茨海默病转录本中激活率高于对照组的前馈神经元,并通过调整相应的下投影权重来调节它们在生成过程中的输出贡献。这产生了九种不同干预方向、幅度和范围的编辑变体。原始模型和编辑模型完成了相同的12轮神经心理学测试,通过盲评人类评分和计算语言学指标进行评估。增强与阿尔茨海默病相关的神经元导致了故事回忆、语言流畅性、工作记忆、程序性话语、场景构建和共指解析等方面的渐进性损害。减弱干预在很大程度上保持了性能,并选择性地改善了几个结果。增强还降低了词汇惊讶度、思想密度、句法复杂性和话语数量,广泛平行于人类阿尔茨海默病言语中报告的变化。这些发现表明,仅通过临床语言差异识别的神经元可以影响多个认知领域的行为,为阿尔茨海默病相关的计算表型提供了概念验证,并为实验性研究语言与更广泛的认知功能障碍之间的联系提供了一个受控框架。
cs.CL / 29 / 2608.03068
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning
CVPO:通过价值-方差适应和动态课程学习增强大型语言模型的强化学习推理
Abstract
Reinforcement learning (RL) has emerged as an effective method for enhancing the reasoning capabilities of large language models (LLMs). However, existing methods suffer from insufficient precision in feedback on generated answer trajectories and exhibit the phenomenon of problem difficulty drift. To address these challenges, we propose CVPO - Curriculum-guided Value-Variance Policy Optimization. At the response trajectory level, we find that token-level value-variance correlates with exploration intensity. Our theoretical analysis shows this variance bounds policy update magnitude. We then use the estimated trajectory value-variance to quantify the intrinsic randomness in generation. Based on this, we design a variance-aware advantage adjustment mechanism for different reward types. At the question level, we introduce a dynamic curriculum weighting method that adapts to question difficulty. This helps the model focus on tasks matched to its current ability during each training stage. Experimental results show our method outperforms strong value-based baselines like VAPO. It achieves better performance and stronger exploration, enabling more accurate and robust reasoning in language models across various math tasks.
Chinese Translation
强化学习(RL)已成为增强大型语言模型(LLMs)推理能力的有效方法。然而,现有方法在生成答案轨迹的反馈精度上存在不足,并且表现出问题难度漂移的现象。为了解决这些挑战,我们提出了CVPO——课程引导的价值-方差策略优化。在响应轨迹层面,我们发现令牌级别的价值-方差与探索强度相关。我们的理论分析表明,这种方差限制了策略更新的幅度。然后,我们使用估计的轨迹价值-方差来量化生成中的内在随机性。在此基础上,我们为不同奖励类型设计了一种关注方差的优势调整机制。在问题层面,我们引入了一种动态课程加权方法,以适应问题的难度。这有助于模型在每个训练阶段专注于与其当前能力匹配的任务。实验结果表明,我们的方法优于强大的基于价值的基线,如VAPO。它在各种数学任务中实现了更好的性能和更强的探索能力,从而使语言模型能够进行更准确和更稳健的推理。
cs.CL / 30 / 2608.03077
PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation
PAMT:面向过程的多领域机器翻译强化学习
Abstract
Multi-domain machine translation (MDMT) requires more than fluent generation: it demands domain-sensitive translation decisions such as domain disambiguation, terminology control, and stylistic adaptation. Large reasoning models (LRMs) make such decisions explicit through intermediate translation steps, but our analysis across 15 domains and four translation directions shows that this explicit reasoning is double-edged: it improves long-form and high-difficulty translation, yet often drifts in terminology-intensive and stylistically constrained settings. We trace this failure to a credit-assignment bottleneck: existing methods optimize final outputs or coarse trajectories, but cannot identify which translation steps actually help the final translation. To address this, we propose PAMT, a process-aligned training framework that combines cold-start domain-aware Long-CoT supervision with reinforcement learning. PAMT uses sequence-level format and outcome rewards for the final translation, together with a step-level process reward that measures how much each explicit translation step increases the likelihood of the reference translation. Across two backbones, PAMT improves over base models, outperforms MT-specialized baselines on average, and remains competitive with strong LLMs/LRMs across in-domain, OOD, and multilingual settings.
Chinese Translation
多领域机器翻译(MDMT)不仅需要流畅的生成,还要求具备领域敏感的翻译决策,如领域消歧、术语控制和风格适应。大型推理模型(LRMs)通过中间翻译步骤使这些决策显性化,但我们对15个领域和四个翻译方向的分析表明,这种显性推理是双刃剑:它提升了长文本和高难度翻译的质量,但在术语密集和风格受限的环境中往往出现偏差。我们将这一失败归因于信用分配瓶颈:现有方法优化最终输出或粗略轨迹,但无法识别哪些翻译步骤实际上有助于最终翻译。为了解决这个问题,我们提出了PAMT,一种结合冷启动领域感知的长链监督(Long-CoT)与强化学习的过程对齐训练框架。PAMT使用序列级格式和结果奖励来优化最终翻译,同时引入步骤级过程奖励,衡量每个显性翻译步骤在多大程度上增加了参考翻译的可能性。在两个基础模型上,PAMT的表现优于基础模型,平均超越机器翻译(MT)专业基线,并在领域内、超出领域(OOD)和多语言设置中与强大的大型语言模型(LLMs)/大型推理模型(LRMs)保持竞争力。
cs.CL / 31 / 2608.03089
Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining
面向大规模语言模型预训练的可扩展频率和长度感知子文档去重
Abstract
Large-scale pretraining corpora contain substantial duplicate content. Although document-level deduplication is widely used, removing subdocument-level redundancy remains challenging. At corpus scale, suffix-array-based methods are commonly applied independently within shards, leaving cross-shard duplicates undetected and making the resulting retention behavior sensitive to the sharding configuration. Hash-based methods enable global exact duplicate counting, but often rely on fixed copy-retention policies that cannot accommodate heterogeneous repetition patterns. We propose a scalable subdocument deduplication framework that decouples duplicate detection from copy retention. It identifies duplicate groups through natural-boundary segmentation, normalized exact hashing, and distributed aggregation, and then applies an explicit frequency- and length-aware retention policy that allocates an adaptive copy budget to each group, retaining more copies of low-frequency or short repetitions while more aggressively deleting high-frequency or long ones. Experiments on FineWeb-Edu and a code-containing web corpus show that models trained on data processed by our method achieve the best overall performance among the evaluated settings. These results underscore the importance of explicit copy-retention control.
Chinese Translation
大规模预训练语料库包含大量重复内容。尽管文档级去重被广泛使用,但去除子文档级冗余仍然具有挑战性。在语料库规模下,基于后缀数组的方法通常在分片内独立应用,导致跨分片的重复未被检测到,并使得最终的保留行为对分片配置敏感。基于哈希的方法能够实现全局精确重复计数,但往往依赖于固定的副本保留策略,无法适应异构的重复模式。我们提出了一种可扩展的子文档去重框架,将重复检测与副本保留解耦。该框架通过自然边界分割、标准化精确哈希和分布式聚合来识别重复组,然后应用显式的频率和长度感知保留策略,为每个组分配自适应的副本预算,保留更多低频或短重复的副本,同时更积极地删除高频或长重复的副本。在FineWeb-Edu和一个包含代码的网页语料库上的实验表明,基于我们方法处理的数据训练的模型在评估设置中实现了最佳的整体性能。这些结果强调了显式副本保留控制的重要性。
cs.CL / 32 / 2608.03095
VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP
VIVID:一个文化基础的基准测试,揭示越南自然语言处理中的比喻语言差距
Abstract
We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese. VIVID comprises 1,636 idioms and proverbs annotated with five complexity traits (literal expressions, pragmatic nuances, Sino-Vietnamese terms, uncommon vocabulary, folk knowledge) and seven semantic themes. We establish an evaluation framework combining generative and discriminative tasks, proposing an LLM-as-a-Judge approach with aspect-based prompting validated against human judgment (Cohen's kappa = 0.792). Evaluating eight state-of-the-art models reveals critical gaps: Vietnamese-specialized models drastically underperform multilingual systems (VinaLLaMA-7B: 0.13 vs. GPT-4o: 2.46), and even top models achieve less than 50% of maximum scores. Notably, few-shot prompting does not universally improve performance, with GPT-4o exhibiting degradation due to stylistic overfitting. Our analysis exposes systematic failures including literal over-interpretation, lexical gaps, and pragmatic flattening, demonstrating that current models lack cultural competence for nuanced figurative interpretation. VIVID provides an essential tool for advancing figurative language understanding in culturally rich contexts.
Chinese Translation
我们提出了VIVID(越南成语验证与解释深度),这是第一个系统性的基准测试,用于评估越南文化基础的比喻语言理解。VIVID包含1,636个成语和谚语,标注了五个复杂性特征(字面表达、语用细微差别、汉越词汇、不常用词汇、民俗知识)和七个语义主题。我们建立了一个评估框架,结合生成性和判别性任务,提出了一种基于LLM(大型语言模型)作为评判者的方法,并通过与人类判断的对比进行验证(Cohen's kappa = 0.792)。对八个最先进模型的评估揭示了关键差距:专注于越南的模型在表现上远远落后于多语言系统(VinaLLaMA-7B: 0.13 vs. GPT-4o: 2.46),即使是顶尖模型的得分也未达到最大分数的50%。值得注意的是,少量示例提示并不普遍提高性能,GPT-4o由于风格过拟合而表现下降。我们的分析揭示了系统性失败,包括字面过度解释、词汇缺口和语用平面化,表明当前模型缺乏对细致比喻解释的文化能力。VIVID为在文化丰富的背景下推进比喻语言理解提供了一个重要工具。
cs.CL / 33 / 2608.03099
What Language Does and What the Evidence Supports: A Functional Role Taxonomy and Evidence Audit of Language Grounding in Embodied Agents
语言的作用及证据支持:具身代理中语言基础的功能角色分类及证据审计
Abstract
Foundation models place language throughout embodied agents, but its presence does not show what it contributes or how well that contribution is grounded. This survey separates these two questions. We define five non-exclusive functional roles for language: Specification, Embodied Representation, Action Orchestration, Grounding Regulation, and Execution Coupling. For each role, we trace the path from linguistic content to its embodied consumer and identify the observations or interventions that can test the claimed responsibility. Applying this framework to the reviewed literature reveals a recurring gap between functional use and evidential support. Interpretable or revised linguistic intermediates may be incorrect, go unused, or fail to affect later behavior. Even when actions are directly conditioned on language, system-level success does not by itself isolate language's contribution. We therefore evaluate grounding claim by claim, asking whether the reported evidence supports the specific responsibility assigned to language. Using role claims rather than architectures as the unit of comparison allows us to compare modular and end-to-end embodied agents without extending conclusions beyond the reported evidence.
Chinese Translation
基础模型将语言融入具身代理中,但其存在并不表明其贡献或该贡献的基础有多扎实。本调查将这两个问题分开。我们定义了五种非排他性的语言功能角色:规范、具身表征、行动协调、基础调节和执行耦合。对于每个角色,我们追踪从语言内容到其具身消费者的路径,并识别可以测试所声称责任的观察或干预。将这一框架应用于已审查的文献揭示了功能使用与证据支持之间的反复差距。可解释或修订的语言中介可能是错误的、未被使用,或未能影响后续行为。即使行动直接依赖于语言,系统级的成功也不能单独确定语言的贡献。因此,我们逐条评估基础主张,询问所报告的证据是否支持分配给语言的特定责任。使用角色主张而非架构作为比较单位,使我们能够比较模块化和端到端的具身代理,而不将结论扩展到超出所报告的证据。
cs.CL / 34 / 2608.03105
HomoEnsNER: Does Language Alignment Outperform Architectural Complexity in Gujarati Named Entity Recognition?
HomoEnsNER:语言对齐是否优于架构复杂性在古吉拉特语命名实体识别中的表现?
Abstract
Named Entity Recognition (NER) for Gujarati remains underexplored, hindered by the absence of capitalization cues, rich morphology, lexical ambiguity, and free word order. Prior ensemble work has emphasized architectural diversity by combining heterogeneous classifiers, multilingual encoders, or classical sequence models, rather than exploiting language-aligned monolingual pretraining. This study asks whether, for a low-resource, morphologically rich language like Gujarati, a homogeneous ensemble of a single monolingual encoder outperforms such architectural diversity. We propose HomoEnsNER, a homogeneous ensemble of five independently fine-tuned GujaratiBERT models combined via majority voting, evaluated against a single GujaratiBERT baseline and six heterogeneous alternatives, including combinations with MuRIL-base, MuRIL-large, IndicBERT, mBERT, BiLSTM, CRF, and a stacked BiLSTM-CRF-GujaratiBERT architecture. All eight models were trained under a consistent budget and evaluated using entity-level F1 on the Naamapadam Gujarati test split. HomoEnsNER achieved the highest F1 (0.8442), surpassing the baseline (0.8347) and every heterogeneous alternative (lowest: 0.7855), indicating that language alignment is a more effective, budget-conscious ensembling strategy than architectural complexity for low-resource Indian language NER.
Chinese Translation
古吉拉特语的命名实体识别(NER)仍然未得到充分探索,受限于缺乏大写提示、丰富的形态学、词汇歧义和自由词序。以往的集成研究强调通过结合异构分类器、多语言编码器或经典序列模型来实现架构多样性,而不是利用语言对齐的单语预训练。本研究探讨对于像古吉拉特语这样资源匮乏且形态丰富的语言,单一单语编码器的同质集成是否优于这种架构多样性。我们提出了HomoEnsNER,这是一个由五个独立微调的GujaratiBERT模型通过多数投票组合而成的同质集成,与单一的GujaratiBERT基线和六种异构替代方案进行评估,包括与MuRIL-base、MuRIL-large、IndicBERT、mBERT、BiLSTM、CRF以及堆叠的BiLSTM-CRF-GujaratiBERT架构的组合。所有八个模型在一致的预算下进行训练,并使用实体级F1在Naamapadam古吉拉特语测试集上进行评估。HomoEnsNER达到了最高的F1值(0.8442),超过了基线(0.8347)和每一个异构替代方案(最低:0.7855),表明对于资源匮乏的印度语言NER而言,语言对齐是一种比架构复杂性更有效且预算友好的集成策略。
cs.CL / 35 / 2608.03118
From SQL Errors to Concept Gaps: An AI-Powered Knowledge Graph Analytics Platform for Personalized Feedback
从 SQL 错误到概念缺口:一个基于 AI 的知识图谱分析平台用于个性化反馈
Abstract
This innovative practice full paper describes an AI-powered knowledge graph platform that connects SQL errors to conceptual gaps in undergraduate and graduate database systems courses. Students learning Structured Query Language (SQL) frequently struggle with semantic errors that reflect conceptual misunderstandings rather than syntax mistakes. A query may execute yet return incorrect results due to gaps spanning related concepts; misusing NATURAL JOIN in place of an explicit subquery reflects intertwined misunderstandings of JOIN, GROUP BY, and HAVING. Autograding systems detect correctness but provide surface-level feedback without connecting errors to the conceptual structure of the course. Educational knowledge graph research has shown the value of structured concept representations for curriculum analysis and adaptive learning, but these approaches have not been applied to diagnosing SQL misconceptions from student submissions. We present a platform that automatically extracts course concepts and relations from instructional materials, links them to student submission traces through a graph database, and classifies errors at the concept level. We evaluate the platform across two database systems courses at two universities, one using real student submissions and one using simulated submissions, through an expert study with five participants and an automated evaluation using an LLM as a judge. Results show that 95.7% of extracted nodes were rated as at least somewhat valid and 63.8% of triplets were rated fully correct. Expert feedback confirmed that the generated graphs align with instructor mental models and that mapping errors to course concepts provides actionable diagnostic insight; evaluating impact on student learning remains future work.
Chinese Translation
本文描述了一个基于 AI 的知识图谱平台,该平台将 SQL 错误与本科生和研究生数据库系统课程中的概念缺口联系起来。学习结构化查询语言(SQL)的学生常常在语义错误上遇到困难,这些错误反映了概念上的误解,而非语法错误。一个查询可能会执行但返回错误的结果,这通常是由于相关概念之间的缺口所致;例如,错误地使用 NATURAL JOIN 代替显式子查询反映了对 JOIN、GROUP BY 和 HAVING 的交织误解。自动评分系统能够检测正确性,但提供的反馈仅停留在表面,未能将错误与课程的概念结构联系起来。教育知识图谱研究表明,结构化概念表示在课程分析和自适应学习中的价值,但这些方法尚未应用于诊断学生提交中的 SQL 误解。我们提出了一个平台,该平台自动从教学材料中提取课程概念和关系,通过图数据库将其与学生提交的痕迹链接,并在概念层面上对错误进行分类。我们在两所大学的两个数据库系统课程中评估了该平台,一所使用真实学生提交,另一所使用模拟提交,通过五位参与者的专家研究和使用 LLM 作为评判的自动评估。结果显示,95.7% 的提取节点被评为至少有一定有效性,63.8% 的三元组被评为完全正确。专家反馈确认生成的图与教师的心理模型一致,并且将错误映射到课程概念提供了可操作的诊断见解;评估对学生学习的影响仍需未来的研究。
cs.CL / 36 / 2608.03138
Internalizing Academic Writing Workflows for Introduction Generation via Struct-Aware Policy Learning
通过结构感知策略学习内化学术写作工作流程以生成引言
Abstract
Generating a rigorous paper introduction with large language models (LLMs) remains challenging, since it requires coordinating background, gap identification, method and contribution within a coherent narrative. Existing solutions externalize this process as multi-stage prompts or agent workflows which are expensive and vulnerable to cross-stage drift. We propose StructPO, a struct-aware policy learning framework that internalizes the entire multi-stage writing workflow into a single-pass policy controlled by explicit stage tokens. StructPO introduces struct-aware credit assignment to decouple local stage quality from global coherence and refinement-guided optimization to internalize revision behavior into the first-pass policy. Experiments show that StructPO improves semantic alignment, structural rationality and inference efficiency over workflow-based baselines, generalizes to out-of-domain settings, and remains competitive with GPT-5.1 in human evaluation when scaled to Qwen3-32B. These results show that internalizing academic writing workflows through fine-grained policy optimization offers a viable alternative to costly external orchestration.
Chinese Translation
利用大型语言模型(LLMs)生成严谨的论文引言仍然具有挑战性,因为这需要在连贯的叙述中协调背景、差距识别、方法和贡献。现有解决方案将这一过程外部化为多阶段提示或代理工作流程,这既昂贵又容易受到跨阶段漂移的影响。我们提出了StructPO,一个结构感知的策略学习框架,它将整个多阶段写作工作流程内化为由显式阶段标记控制的单次策略。StructPO引入了结构感知的信用分配,以将局部阶段质量与全局一致性解耦,并通过优化引导修订行为,将其内化到首次策略中。实验表明,StructPO在语义对齐、结构合理性和推理效率上优于基于工作流程的基线,并且在迁移到域外设置时具有良好的泛化能力,在扩展到Qwen3-32B时在人类评估中仍与GPT-5.1保持竞争力。这些结果表明,通过细粒度策略优化内化学术写作工作流程提供了一种可行的替代方案,取代了昂贵的外部协调。
cs.CL / 37 / 2608.03154
ANCHOR-RE: An Agentic Neuro-Symbolic Framework for Grounded Biomedical Relation Extraction
ANCHOR-RE:一种代理神经符号框架用于基础生物医学关系提取
Abstract
Biomedical relation extraction (BioRE) extracts structured knowledge from biomedical literature for applications such as knowledge base construction and hypothesis generation. Traditional symbolic systems such as SemRep provide high precision but limited recall, while large language models (LLMs) offer stronger contextual reasoning but remain prone to false-positive predictions. We developed ANCHOR-RE, a framework that integrates ontology-guided reasoning, external knowledge grounding, and data-driven verification rules into LLM inference. We evaluated it on three BioRE benchmarks (SemRepGS, DDI, and ChemProt) using both proprietary and open-weight LLMs. To assess generalizability beyond benchmark datasets while reducing potential evaluation bias from LLM pretraining contamination, we conducted a temporal evaluation using 100 biomedical articles published in 2026. With the proprietary backbone, ANCHOR-RE outperformed direct LLM prompting, improving micro-F1 from 0.654 to 0.676 on SemRepGS, from 0.769 to 0.872 on DDI, and from 0.939 to 0.941 on ChemProt. On DDI and ChemProt, it also outperformed previously reported inference-only methods and approached fine-tuned or instruction-tuned systems without parameter updates. Similar performance gains observed with open-weight LLMs indicate that the benefits were not limited to the proprietary backbone. On the post-cutoff set, manual assessment of 500 randomly sampled predictions yielded a precision of 69%, maintaining consistent precision on previously unseen biomedical literature. Neuro-symbolic reasoning can improve the reliability of LLM-based BioRE without fine-tuning. Results across multiple benchmarks, model families, and post-cutoff literature support ANCHOR-RE as a practical training-free approach to biomedical literature mining.
Chinese Translation
生物医学关系提取(BioRE)从生物医学文献中提取结构化知识,用于知识库构建和假设生成等应用。传统的符号系统如SemRep提供高精度但召回率有限,而大型语言模型(LLMs)则提供更强的上下文推理能力,但仍容易产生假阳性预测。我们开发了ANCHOR-RE,一个将本体引导推理、外部知识基础和数据驱动的验证规则整合到LLM推理中的框架。我们在三个BioRE基准(SemRepGS、DDI和ChemProt)上进行了评估,使用了专有和开放权重的LLM。为了评估其在基准数据集之外的泛化能力,并减少来自LLM预训练污染的潜在评估偏差,我们使用2026年发表的100篇生物医学文章进行了时间评估。在专有骨干网络上,ANCHOR-RE的表现优于直接的LLM提示,在SemRepGS上微F1从0.654提高到0.676,在DDI上从0.769提高到0.872,在ChemProt上从0.939提高到0.941。在DDI和ChemProt上,它也优于先前报告的仅推理方法,并接近于经过微调或指令微调的系统而无需参数更新。在开放权重LLM上观察到的类似性能提升表明,这些好处并不限于专有骨干网络。在截止后数据集上,对500个随机抽样预测的手动评估得到了69%的精度,在之前未见过的生物医学文献中保持了一致的精度。神经符号推理可以在不进行微调的情况下提高基于LLM的BioRE的可靠性。多个基准、模型系列和截止后文献的结果支持ANCHOR-RE作为一种实用的无训练生物医学文献挖掘方法。
cs.CL / 38 / 2608.03204
Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach
在测试时对齐大型视觉-语言模型:一种基于轨迹引导的结构化采样方法
Abstract
Post-training reinforcement learning (RL) algorithms are commonly used to align large vision-language models (LVLMs) with human intent and the requirements of visual reasoning tasks. However, existing RL-based alignment methods are often resource-intensive and encounter mismatches between training objectives and inference-time distributions. To bridge this gap, we propose a novel test-time alignment approach that leverages trajectory-guided structured sampling for dynamic inference-time refinement, achieving better alignment with visual grounding and ensuring logical consistency. Our approach begins with curating a reasoning memory bank via a trajectory learning algorithm, which decomposes complex question solving into ordered sequences of predefined reasoning patterns. It subsequently accomplishes inference-time alignment by first collecting trajectories from reasoning memory bank to establish a global structural reasoning prior, and then using an iterative Markov Chain Monte Carlo (MCMC) algorithm for localized multi-objective refinement of the reasoning trace. Experiments across multiple multimodal reasoning datasets demonstrate that our approach significantly improves accuracy without incurring prohibitive inference overhead. These results establish trajectory-guided test-time sampling as a scalable and effective alternative to traditional post-training alignment, particularly for complex visual reasoning tasks.
Chinese Translation
后训练强化学习(RL)算法通常用于将大型视觉-语言模型(LVLMs)与人类意图及视觉推理任务的要求对齐。然而,现有的基于RL的对齐方法往往资源密集,并且在训练目标与推理时分布之间存在不匹配。为了解决这一问题,我们提出了一种新颖的测试时对齐方法,该方法利用轨迹引导的结构化采样进行动态推理时的优化,从而实现与视觉基础的更好对齐,并确保逻辑一致性。我们的方法首先通过轨迹学习算法策划一个推理记忆库,该算法将复杂的问题解决分解为预定义推理模式的有序序列。随后,通过从推理记忆库中收集轨迹来建立全局结构推理先验,进而使用迭代的马尔可夫链蒙特卡洛(MCMC)算法对推理轨迹进行局部多目标优化,从而实现推理时的对齐。在多个多模态推理数据集上的实验表明,我们的方法显著提高了准确性,而不会产生过高的推理开销。这些结果确立了轨迹引导的测试时采样作为传统后训练对齐的可扩展且有效的替代方案,特别适用于复杂的视觉推理任务。
cs.CL / 39 / 2608.03210
ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization
ICO:通过迭代上下文优化增强语义转移越狱
Abstract
Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift jailbreaks have recently emerged as a promising attack paradigm. They bypass explicit safety mechanisms by replacing harmful terms in original harmful questions with benign alternatives and leveraging contextual information to induce the target model to reinterpret these alternatives as their corresponding harmful concepts. However, existing semantic-shift jailbreaks often achieve limited effectiveness. In this work, we reveal that this limitation arises from overlooking the semantic-shift capability of contexts. Through systematic analysis, we find that contexts exhibit substantially different abilities in inducing semantic shifts: contexts with stronger semantic-shift capabilities are more likely to guide models toward recovering harmful meanings and achieving successful jailbreaks. Based on this finding, we systematically identify and distill the characteristics of effective contexts and propose a black-box context-aware semantic-shift jailbreak framework with Iterative Context Optimization (ICO). In each iteration, ICO leverages these characteristics and feedback from the target model to optimize contexts. Extensive experiments on three datasets and eight target foundation models demonstrate that ICO consistently outperforms eight state-of-the-art baselines, achieving an average attack success rate of 74.6%.
Chinese Translation
基础模型在各种任务中取得了显著成功,但仍然存在脆弱性。为了研究这些脆弱性,语义转移越狱最近作为一种有前景的攻击范式出现。它们通过用无害的替代词替换原始有害问题中的有害术语,绕过明确的安全机制,并利用上下文信息诱导目标模型将这些替代词重新解释为其对应的有害概念。然而,现有的语义转移越狱往往效果有限。在本研究中,我们揭示了这一限制源于忽视了上下文的语义转移能力。通过系统分析,我们发现上下文在诱导语义转移方面表现出显著不同的能力:具有更强语义转移能力的上下文更有可能引导模型恢复有害含义并实现成功的越狱。基于这一发现,我们系统地识别并提炼了有效上下文的特征,并提出了一种基于黑箱的上下文感知语义转移越狱框架,名为迭代上下文优化(ICO)。在每次迭代中,ICO利用这些特征和来自目标模型的反馈来优化上下文。在三个数据集和八个目标基础模型上的大量实验表明,ICO始终优于八个最先进的基线,平均攻击成功率达到74.6%。
cs.CL / 40 / 2608.03233
On the Diversity of Analogy Making in Large Language Models
大型语言模型类比生成的多样性研究
Abstract
Large Language Models (LLMs) have demonstrated remarkable potential for analogy making, a core cognitive capability that drives novelty and creativity. While prior research has extensively investigated the applications and underlying mechanisms of LLM-based analogy making, its output diversity remains largely unexplored, despite being essential for broadening cross-domain connections and fostering scientific innovation. In this work, we present a comprehensive evaluation of analogy diversity across ten state-of-the-art open- and closed-source LLMs. Our findings highlight a concerning issue of domain homogeneity, a prevalent tendency for LLMs to generate analogies from a narrow set of target domains, limiting both inter-query and intra-model diversity. Furthermore, our analysis reveals a fundamental trade-off in existing LLM diversity-enhancement methods: increasing output diversity often comes at the expense of output quality. Finally, our causal analysis of LLM information flow reveals substantial differences in the model-sensitive regions governing analogy diversity across LLMs, suggesting a potential mechanism for the observed diversity-quality trade-off. To our knowledge, this is among the first studies to systematically investigate output diversity in LLM-based analogy making.
Chinese Translation
大型语言模型(LLMs)在类比生成方面展现了显著的潜力,这是一种推动新颖性和创造力的核心认知能力。尽管先前的研究广泛探讨了基于LLM的类比生成的应用和基本机制,但其输出多样性仍然在很大程度上未被探索,而输出多样性对于拓宽跨领域联系和促进科学创新至关重要。在本研究中,我们对十种最先进的开源和闭源LLM的类比多样性进行了全面评估。我们的发现突显了一个令人担忧的领域同质性问题,即LLM倾向于从狭窄的目标领域生成类比,这限制了查询间和模型内的多样性。此外,我们的分析揭示了现有LLM多样性增强方法中的一个基本权衡:提高输出多样性往往以输出质量为代价。最后,我们对LLM信息流的因果分析显示,控制类比多样性的模型敏感区域在不同LLM之间存在显著差异,这暗示了观察到的多样性-质量权衡的潜在机制。我们所知,这是首次系统性研究基于LLM的类比生成输出多样性的研究之一。
cs.CL / 41 / 2608.03239
Relational Priors as Convergence Pressure in LLM-Based Multi-Agent Systems
关系先验作为基于大型语言模型的多代理系统中的收敛压力
Abstract
Large language model-based multi-agent systems (LLM-MAS) are designed through roles, debate protocols, and aggregation rules. These choices create implicit social expectations: agents may be expected to trust, challenge, defer to, or collaborate with peers. We study the effects of making inter-agent relation semantics explicit. We use a minimal signed-network formulation of relational priors and inject natural-language renderings into agent system prompts while holding the task protocol fixed. Across a commons-governance simulation and multi-agent debate, relational priors primarily act as convergence pressure: increasing relational positivity tends to make agents coordinate or agree more readily. This pressure can help when utility rewards behavioral alignment, as in sustainable resource governance and subjective consensus. It does not, however, reliably improve accuracy. In objective QA debates, higher positivity can increase agreement even when correctness-conditioned agreement does not improve and may decline in some settings. Effects vary by model backbone, relation type, and topology; explicit neutrality is not equivalent to omitting relational framing. We argue that relational priors should not be a default add-on for LLM-MAS. Their safer use is diagnostic and task-specific: compare against a no-prior baseline, monitor correctness-conditioned metrics when truth matters, and omit the relational layer when validation does not justify it.
Chinese Translation
基于大型语言模型的多代理系统(LLM-MAS)通过角色、辩论协议和聚合规则进行设计。这些选择产生了隐含的社会期望:代理可能被期望信任、挑战、服从或与同伴合作。我们研究了将代理间关系语义明确化的效果。我们使用最小化的有向网络关系先验公式,并在保持任务协议不变的情况下,将自然语言呈现注入代理系统提示中。在一个公共治理模拟和多代理辩论中,关系先验主要作为收敛压力发挥作用:增加关系积极性往往使代理更容易协调或达成一致。当效用奖励行为一致性时,例如在可持续资源治理和主观共识中,这种压力可以提供帮助。然而,它并不可靠地提高准确性。在客观问答辩论中,即使在正确性条件下的协议没有改善,甚至在某些情况下可能下降,较高的积极性也可能增加协议。效果因模型骨干、关系类型和拓扑而异;明确的中立性并不等同于省略关系框架。我们认为,关系先验不应作为LLM-MAS的默认附加项。它们的安全使用应是诊断性和任务特定的:与无先验基线进行比较,当真相重要时监测正确性条件指标,并在验证不合理时省略关系层。
cs.CL / 42 / 2608.03275
MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation
MoEGen:用于实例自适应 LoRA 生成的专家混合模型
Abstract
Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve capacity by storing multiple full LoRA experts, causing adapter storage to grow linearly with the number of experts and restricting adaptation to a fixed expert pool. We ask whether MoE-based PEFT can produce instance-specific adaptations without explicitly storing a separate LoRA module for each expert. To address this gap, we propose MoEGen, an adaptation framework that shifts MoE-based PEFT from expert selection to expert-conditioned parameter generation. Instead of storing each expert as a full LoRA adapter, MoEGen represents each expert as a small learnable vector, termed an expert code. It routes each input over these vectors and uses their weighted combination to condition a lightweight hypernetwork that generates input-specific low-rank updates. This design decouples expert capacity from adapter storage while enabling instance-conditioned adaptation. Experiments on eight commonsense reasoning benchmarks show consistent improvements over strong static and MoE-based PEFT baselines across three backbones. MoEGen also performs strongly in joint medical and legal-domain adaptation.
Chinese Translation
参数高效微调(PEFT)使大型语言模型的高效适应成为可能,但现有的基于专家混合模型(MoE)的 PEFT 方法通常通过存储多个完整的 LoRA 专家来提高容量,导致适配器存储量随着专家数量线性增长,并限制了适应于固定的专家池。我们探讨 MoE 基于 PEFT 是否可以在不显式存储每个专家的单独 LoRA 模块的情况下,产生实例特定的适应。为了解决这一问题,我们提出了 MoEGen,一个将 MoE 基于 PEFT 从专家选择转变为专家条件参数生成的适应框架。MoEGen 不再将每个专家存储为完整的 LoRA 适配器,而是将每个专家表示为一个小的可学习向量,称为专家编码。它通过这些向量路由每个输入,并利用它们的加权组合来条件化一个轻量级超网络,从而生成输入特定的低秩更新。这一设计将专家容量与适配器存储解耦,同时实现实例条件适应。在八个常识推理基准上的实验表明,MoEGen 在三种基础模型上相较于强大的静态和基于 MoE 的 PEFT 基线表现出一致的改进。MoEGen 在医学和法律领域的联合适应中也表现出色。
cs.CL / 43 / 2608.03340
Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks
基准测试基准:测试常识基准的预测有效性
Abstract
Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream performance remains underspecified. To establish the practical usability of widely adopted commonsense benchmarks, we evaluate 23 models from six families on four established commonsense benchmarks, four reworked variants, three non-commonsense controls, and eight downstream tasks requiring implicit social, pragmatic, temporal, or physical reasoning. We compare model rankings, compute controlled correlations, and use leave-one-family-out cross-validation to assess the criterion validity of commonsense benchmarks. Our results show that revised benchmarks largely preserve original model rankings and do not improve downstream predictive power. Commonsense benchmarks show consistent cross-family predictive validity for only a narrow subset of downstream tasks, with smaller or metric-specific gains elsewhere. Overall, standardized commonsense benchmarks provide task-dependent rather than broad evidence of downstream commonsense competence.
Chinese Translation
预测大型语言模型(LLM)在现实世界任务中的能力至关重要,但常识基准的表现在多大程度上能预测下游表现仍然不够明确。为了确立广泛采用的常识基准的实际可用性,我们对来自六个模型家族的23个模型在四个已建立的常识基准、四个重新设计的变体、三个非常识控制以及八个需要隐含社会、语用、时间或物理推理的下游任务上进行了评估。我们比较了模型排名,计算了受控相关性,并使用留一模型家族交叉验证来评估常识基准的标准有效性。我们的结果表明,修订后的基准在很大程度上保留了原始模型排名,并未提高下游预测能力。常识基准在仅限于一小部分下游任务上显示出跨家族的一致预测有效性,而在其他方面则表现出较小或特定指标的增益。总体而言,标准化的常识基准提供了依赖于任务的证据,而非广泛的下游常识能力证据。
cs.CL / 44 / 2608.03358
ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models
ArtECulture:多模态大型语言模型中基于文化的视觉情感理解基准测试
Abstract
Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emotion understanding, a task that predicts the culture-specific emotional perception of a given image and explains the underlying rationale. Although related benchmarks exist, they are limited by inconsistent individual annotations, which hinder the derivation of majority-supported culture-level emotion labels, and imbalanced cultural coverage. Thus, we present ArtECulture, a benchmark containing 6,792 artworks with culture-specific emotion labels and explanations across English, Chinese, and Arabic cultures, with balanced Western and non-Western content. Evaluations of 16 open- and closed-source Multimodal Large Language Models (MLLMs) under a zero-shot setting reveal that the task remains challenging, with the best model achieving below 50\% accuracy. To address this limitation, we introduce a retrieval-augmented culture-conditioned emotion understanding framework, which leverages a concept-based cultural emotion knowledge base to inject explicit cultural knowledge into MLLMs without additional training. The framework improves both culturally aligned emotion prediction and grounded explanation generation. Our benchmark and code will be publicly released.
Chinese Translation
现有的视觉情感理解方法通常忽视情感感知中的文化差异。我们提出了基于文化的视觉情感理解,这是一项预测给定图像的文化特定情感感知并解释其背后原理的任务。尽管存在相关基准,但它们受到不一致的个体注释的限制,这妨碍了多数支持的文化层面情感标签的推导,并且文化覆盖不平衡。因此,我们提出了ArtECulture,一个包含6792件艺术作品的基准,涵盖了英语、中文和阿拉伯文化中的文化特定情感标签和解释,并且西方与非西方内容平衡。在零-shot设置下对16个开放源和闭源多模态大型语言模型(MLLMs)的评估表明,该任务仍然具有挑战性,最佳模型的准确率低于50%。为了解决这一限制,我们引入了一种增强检索的基于文化的情感理解框架,该框架利用基于概念的文化情感知识库,将明确的文化知识注入MLLMs,而无需额外训练。该框架改善了文化对齐的情感预测和基于事实的解释生成。我们的基准和代码将公开发布。
cs.CL / 45 / 2608.03372
FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact
FACTWASH:捕捉将谣言洗涤为事实的人工智能重写
Abstract
AI systems rewrite information constantly: conversations become stored memories, documents become answers. The rewrite can keep a claim while washing away what made it checkable, who said it, how sure they were, when it held. We call that failure factwashing, and release factwash, an open-source write-time gate that catches it deterministically, with named flags and evidence rather than an LLM judge. Building it answers a practical question: when does a cheap check suffice, and when do you need a model? What decides is whether the property has a bounded surface-cue inventory. Explicit negation cues are close to enumerable, so a word list finishes and transfers, reaching 0.91 F1 on untuned text. Hedging and attribution have open-ended realizations, so vocabulary plateaus near half recall, and a one-question LLM witness recovers +17 and +15 points of cue-detection recall at equal precision. Deployed, that witness may only lower a verdict, so it buys precision rather than coverage. We measure cue detection on 105,596 independently annotated sentences. A blind-labelled corpus of memory writes then locates the failure: 55% of bad writes in conversational hearsay, 7% in business email (p < 0.001), so the first deployment question is not which detector to use but whether the failure occurs at all. On unmodified mem0 2.0.7, the gate flags 5 of 8 hedged-hearsay writes.
Chinese Translation
人工智能系统不断重写信息:对话变成存储的记忆,文档变成答案。这种重写可以保留一个主张,同时洗去使其可验证的内容、说出该主张的人、他们的确信程度以及该主张成立的时间。我们称这种失败为事实洗涤(factwashing),并发布了 factwash,这是一个开源的写时门控系统,它以确定性的方式捕捉这种现象,使用命名标志和证据,而不是依赖大型语言模型(LLM)进行判断。构建该系统回答了一个实际问题:何时廉价的检查足够,何时需要模型?决定因素是属性是否具有有限的表面线索库存。显式否定线索接近可枚举,因此词汇表的完成和转移使得在未经调优的文本上达到了 0.91 的 F1 值。模糊和归属具有开放式的实现,因此词汇量的召回率接近一半,而一个问题的 LLM 证人可以在相同精度下恢复 +17 和 +15 的线索检测召回率。部署后,该证人可能仅降低判决,因此它购买的是精度而非覆盖率。我们在 105,596 个独立注释的句子上测量线索检测。一个盲标记的记忆写作语料库随后定位了失败:55% 的不良写作出现在对话谣言中,7% 出现在商业电子邮件中(p < 0.001),因此第一个部署问题不是使用哪个检测器,而是失败是否确实发生。在未修改的 mem0 2.0.7 上,该门控系统标记了 8 个模糊谣言写作中的 5 个。
cs.CL / 46 / 2608.03388
Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference
不要让我去请求:大型语言模型在主动多轮信息获取中的缺陷对于溯因推理的影响
Abstract
Abductive reasoning requires forming hypotheses that explain observed evidence and revising them as new evidence becomes available. While large language models (LLMs) are often evaluated on whether they solve abductive reasoning tasks correctly, less is known about how they acquire evidence, update their hypotheses, and decide when to stop. We introduce Alien Abduction game, an interactive probe for studying these behaviours under different interaction modes. The modes vary in whether evidence is provided upfront or across turns, and whether queries are selected by the model or examples are provided by the oracle. Across models, providing evidence upfront leads to higher success rates than distributing it across turns. In multi-turn settings, some models commit before using the available evidence, while others exhaust the turn budget without converging. Models also achieve higher success rates when examples are provided by the oracle than when they select their own queries, although their final hypotheses are more consistent with the evidence they selected. These findings suggest that models may form hypotheses that fit self-selected evidence without sufficiently distinguishing them from alternatives, and may struggle to validate and refine their hypotheses or determine when to stop.
Chinese Translation
溯因推理要求形成解释观察到的证据的假设,并在新证据可用时对其进行修正。虽然大型语言模型(LLMs)通常被评估是否正确解决溯因推理任务,但关于它们如何获取证据、更新假设以及决定何时停止的了解较少。我们引入了外星人绑架游戏(Alien Abduction game),这是一个用于研究这些行为在不同交互模式下的互动探测工具。这些模式在于证据是提前提供还是分轮提供,以及查询是由模型选择还是由神谕提供示例。在不同模型中,提前提供证据的成功率高于分轮提供证据。在多轮设置中,一些模型在使用可用证据之前就做出了承诺,而其他模型则在未收敛的情况下耗尽了轮次预算。当示例由神谕提供时,模型的成功率也高于自选查询的情况,尽管它们最终的假设与所选证据更为一致。这些发现表明,模型可能会形成与自选证据相符的假设,而未能充分区分它们与其他替代方案,并可能在验证和修正其假设或确定何时停止方面面临困难。
cs.CL / 47 / 2608.03411
DUD: Decoupled Update Dynamics for Reliable Uncertainty Quantification in Large Language Models
DUD:用于大型语言模型可靠不确定性量化的解耦更新动态
Abstract
Accurate Uncertainty Quantification (UQ) is critical for reliable deployment of Large Language Models (LLMs), yet traditional probability-based metrics often fail to capture the model's true epistemic state. While recent mechanistic approaches leverage hidden state dynamics, they typically aggregate residual stream updates, conflating the distinct roles of parametric memory (Feed-Forward Networks) and contextual processing (Attention). We argue that this aggregation obscures fine-grained mechanistic conflicts, such as memory-context misalignment, that are fundamental indicators of uncertainty. To address this, we introduce \textbf{D}ecoupled \textbf{U}pdate \textbf{D}ynamics \textbf{(DUD)}, a framework that explicitly decouples FFN and Attention contributions via noise-induced causal interventions. By quantifying the independent restoration capabilities of each module, we construct a dual-stream dynamic profile that captures the model's internal fragility. Extensive experiments demonstrate that DUD significantly outperforms state-of-the-art baselines in both uncertainty estimation and calibration, while exhibiting superior cross-dataset generalization, validating decoupled dynamics as a robust proxy for model faithfulness.
Chinese Translation
准确的不确定性量化(UQ)对于大型语言模型(LLMs)的可靠部署至关重要,但传统的基于概率的指标往往无法捕捉模型的真实认知状态。尽管近期的机械方法利用了隐藏状态动态,但它们通常聚合残差流更新,混淆了参数记忆(前馈网络)和上下文处理(注意力)的不同角色。我们认为这种聚合掩盖了细粒度的机械冲突,例如记忆与上下文的不对齐,这些都是不确定性的基本指示因素。为了解决这个问题,我们提出了 extbf{D}ecoupled extbf{U}pdate extbf{D}ynamics extbf{(DUD)},一个通过噪声诱导的因果干预明确解耦前馈网络和注意力贡献的框架。通过量化每个模块的独立恢复能力,我们构建了一个双流动态特征,捕捉模型的内部脆弱性。大量实验表明,DUD在不确定性估计和校准方面显著优于最先进的基线,同时展现出卓越的跨数据集泛化能力,验证了解耦动态作为模型可信度的强有力代理。
cs.CL / 48 / 2608.03437
Dynamically Allocating Evaluation Effort for Model Ranking
动态分配模型排名的评估工作量
Abstract
While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation protocols waste effort by exhaustively evaluating all models on the entire benchmark, a safe but inefficient approach. In this work, we formalize multi-model human evaluation as a best-arm identification problem in a multi-armed bandit setup with correlated arms, where pulling an arm corresponds to human-evaluating a model. By sampling adaptively based on the intermediate model rankings obtained on the samples so far, we can focus the annotation budget on the most competitive models. We prove the optimality of the proposed algorithms and show that it improves discrimination between top-performing models. This makes evaluations faster, cheaper and more aligned with large-scale competition evaluation goals.
Chinese Translation
尽管人工评估在许多自然语言处理任务中被视为金标准,但其成本高昂且扩展性差。在识别表现最佳的模型时,典型的评估协议通过对所有模型在整个基准上进行全面评估而浪费了评估工作量,这是一种安全但效率低下的方法。在本研究中,我们将多模型人工评估形式化为一个多臂赌博机设置中的最佳臂识别问题,其中拉动一个臂对应于对一个模型进行人工评估。通过根据迄今为止获得的中间模型排名进行自适应采样,我们可以将标注预算集中在最具竞争力的模型上。我们证明了所提出算法的最优性,并展示其在区分表现最佳模型方面的改进。这使得评估过程更快、更便宜,并与大规模竞争评估目标更为一致。
cs.CL / 49 / 2608.03446
Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment $\unicode{x2013}$ Is English Enough?
通过跨语言对齐预测大规模多语言模型的分类和翻译性能——英语是否足够?
Abstract
Multilingual large language models (LLMs) have been shown to perform better on non-English classification tasks when the representations of the given language are more aligned to English within the model. Several cross-lingual alignment (CLA) scores have been proposed for use with LLMs, along with multiple approaches for extracting embeddings from the models. We provide a comparative analysis of 27 CLA score variants, examining how they differ and how well each predicts downstream performance across three tasks. Crucially, while LLMs are widely used for generative tasks such as machine translation, prior work has focused almost exclusively on classification. We therefore investigate whether CLA scores are similarly predictive of translation performance. To enable computing correlations across target languages, we propose a PMI-based translation metric, which is less dependent on the target language and correlates strongly with chrF. We find that CLA with English predicts translation quality comparably to or better than source-target CLA, providing new evidence that LLMs use English as an internal pivot language.
Chinese Translation
多语言大规模语言模型(LLMs)在非英语分类任务上的表现已被证明与给定语言在模型中与英语的对齐程度更高时更佳。为LLMs提出了几种跨语言对齐(CLA)评分指标,并提出了多种从模型中提取嵌入的方法。我们对27种CLA评分变体进行了比较分析,考察它们之间的差异以及各自对三项任务下游性能的预测能力。重要的是,尽管LLMs广泛用于生成任务,如机器翻译,先前的研究几乎完全集中于分类。因此,我们调查CLA评分是否同样能够预测翻译性能。为了能够计算目标语言之间的相关性,我们提出了一种基于PMI的翻译指标,该指标对目标语言的依赖性较小,并且与chrF有很强的相关性。我们发现,使用英语的CLA对翻译质量的预测能力与源-目标CLA相当或更佳,这为LLMs将英语作为内部枢轴语言提供了新的证据。
cs.CL / 50 / 2608.03452
Probing Character-level Transformers for the Spanish L-shaped Morphome
探究西班牙语L形形态素的字符级变换器
Abstract
When a transformer learns an irregular morphological pattern, what has it learned? Our test case is the Spanish \emph{L-shaped morphome}, a complex irregular pattern in which the verb's stem alternates in exactly the first-person singular indicative and all subjunctive forms, and whose membership no phonological, semantic, or syntactic feature predicts. Prior studies have shown that character-level transformers can reproduce this pattern, but that evidence describes what models produce, not what they represent. Probing five architectures, twelve trained models each, under lemma-disjoint cross-validation with controls and surface baselines, we show that the models encode the L-shaped class itself, not just its visible alternations. It is decodable above every surface baseline, survives instances in which every form shows the same stem, and probes trained on alternating instances still classify non-alternating ones. The encoding is localized where the stem choice is made, at the stem-final consonant position of the middle decoder, before the alternant is read. And it is item-specific: which verbs a model learned matters far more than which architecture it is. The models store the morphome as an item-specific lexical abstraction, sufficient to reproduce the pattern but not to generalize it as humans do.
Chinese Translation
当一个变换器学习到不规则的形态模式时,它究竟学到了什么?我们的测试案例是西班牙语的 extit{L形形态素},这是一个复杂的不规则模式,其中动词的词干在第一人称单数直陈式和所有虚拟式形式中恰好交替,而其成员资格无法通过任何音位、语义或句法特征来预测。先前的研究表明,字符级变换器能够再现这一模式,但这些证据描述的是模型的输出,而非它们所表示的内容。通过对五种架构、每种架构下的十二个训练模型进行探测,采用词条不重叠的交叉验证,并设定控制组和表面基线,我们展示了这些模型编码了L形类别本身,而不仅仅是其可见的交替形式。该编码在每个表面基线之上都是可解码的,能够在每种形式显示相同词干的情况下生存,并且在交替实例上训练的探测器仍然能够对非交替实例进行分类。编码发生在词干选择的局部位置,即中间解码器的词干末尾辅音位置,在交替形式被读取之前。而且,它是特定于项目的:模型学习到哪些动词远比其架构本身更为重要。这些模型将形态素存储为特定于项目的词汇抽象,足以再现该模式,但无法像人类那样进行概括。
cs.CL / 51 / 2608.03480
Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study
通过语料驱动的词汇剪枝实现高效的多语言神经机器翻译:以英阿语言对为例
Abstract
The adoption of large pre-trained multilingual models for neural machine translation (MNMT) faces a major challenge: excessive memory and computational consumption due to overly large vocabularies and embedding layers. Although existing compression methods like pruning, quantization and knowledge distillation reduce parameter redundancy, they mainly preserve the structure of the original vocabulary, thereby leaving a major source of inefficiency unresolved. We propose in this paper a general optimization framework that combines a vocabulary pruning method with a targeted fine-tuning protocol for MNMT models. We evaluate the proposed framework using three models (M2M100, NLLB-200, mBART-50) on the English-Arabic language pair. Our approach reduces the vocabulary size from over 128,000 to approximately 10,000 tokens, enabling a 60% memory saving without any loss in performance. Results show that optimized multilingual models can match or exceed the performance of dedicated bilingual baselines. In particular, the pruned and fine-tuned M2M100 model achieves a competitive BLEU score of 42.04 (against 44.59 for the OPUS-MTen- ar bilingual model) while it significantly outperforms it on the COMET metric (0.8730 vs 0.7911) revealing superior semantic adequacy and fluency.
Chinese Translation
采用大型预训练多语言模型进行神经机器翻译(MNMT)面临着一个主要挑战:由于词汇和嵌入层过大,导致过度的内存和计算消耗。尽管现有的压缩方法如剪枝、量化和知识蒸馏减少了参数冗余,但它们主要保留了原始词汇的结构,从而未能解决一个主要的低效来源。本文提出了一种通用优化框架,将词汇剪枝方法与针对性的微调协议相结合,应用于MNMT模型。我们在英阿语言对上使用三种模型(M2M100、NLLB-200、mBART-50)评估所提出的框架。我们的方法将词汇大小从超过128,000个减少到约10,000个标记,实现了60%的内存节省且没有性能损失。结果表明,经过优化的多语言模型可以匹配或超过专用双语基准的性能。特别是,经过剪枝和微调的M2M100模型达到了42.04的竞争性BLEU分数(而OPUS-MTen-ar双语模型为44.59),同时在COMET指标上显著优于该模型(0.8730对比0.7911),显示出更优的语义适当性和流畅性。
cs.CL / 52 / 2608.03494
Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension
超越初始化损失:大规模语言模型词汇扩展的标记嵌入初始化策略系统研究
Abstract
Vocabulary extension is an efficient way to adapt pretrained large language models (LLMs) to new languages, but the initialization of newly added token embeddings can strongly affect continued pre-training (CPT) efficiency. We present a systematic study of more than 20 initialization strategies for Hindi vocabulary extension in Nemotron-3-Nano-30B-A3B. Our comparison spans vocabulary-averaging baselines; external and learned initialization methods, including FOCUS, top-k semantic retrieval, and residual MLP mappings; subword composition; norm calibration; and input-output asymmetry. We find that subword composition methods outperform both vocabulary averaging and external/learned initialization approaches. Within subword composition, asymmetric variants achieve the lowest observed early validation loss and reveal distinct preferences for input and output embedding initialization. The best observed configuration initializes the input embedding matrix with uniform subword averaging and Hindi-specific norm calibration, and the output language modeling head with character-length-weighted subword averaging. Relative to the standard Mean-all baseline, this full initialization pipeline reaches comparable validation loss with over a 6x reduction in CPT steps and exceeds the baseline's 3,500-step MILU-Hindi accuracy after only 500 steps. Finally, we show that initialization loss and initialization bits-per-byte (Init BPB) are unreliable predictors of downstream convergence, whereas lightweight CPT, as few as 50 steps, provides a cost-effective and reliable signal for selecting the best initialization strategy.
Chinese Translation
词汇扩展是一种有效的方法,用于将预训练的大规模语言模型(LLMs)适应于新语言,但新添加的标记嵌入的初始化可以显著影响持续预训练(CPT)的效率。我们对在Nemotron-3-Nano-30B-A3B中进行的超过20种印地语词汇扩展的初始化策略进行了系统研究。我们的比较涵盖了词汇平均基线;外部和学习的初始化方法,包括FOCUS、top-k语义检索和残差MLP映射;子词组合;范数校准;以及输入输出不对称性。我们发现,子词组合方法优于词汇平均和外部/学习初始化方法。在子词组合中,不对称变体实现了观察到的最低早期验证损失,并揭示了对输入和输出嵌入初始化的不同偏好。观察到的最佳配置使用均匀的子词平均和特定于印地语的范数校准初始化输入嵌入矩阵,并使用字符长度加权的子词平均初始化输出语言建模头。相较于标准的Mean-all基线,这一完整的初始化流程在CPT步骤减少超过6倍的情况下达到了可比的验证损失,并在仅500步后超过了基线的3,500步MILU-印地语准确率。最后,我们展示了初始化损失和每字节初始化位数(Init BPB)并不是下游收敛的可靠预测指标,而轻量级CPT,甚至仅需50步,提供了一种具有成本效益且可靠的信号,用于选择最佳初始化策略。
cs.CL / 53 / 2608.03505
ConlangBench: Exploring Language Knowledge and Learning in LLMs through Diverse Constructed Languages
ConlangBench:通过多样化的构造语言探索大型语言模型中的语言知识与学习
Abstract
Constructed languages (conlangs) are intentionally created human languages with a rich tradition of linguistic creativity. Despite their potential for studying language learning in large language models (LLMs), existing conlangs remain largely underexplored in LLM research. We present ConlangBench, the first large-scale benchmark for evaluating and training LLMs on 21 existing conlangs. We collect over 21M conlang-English parallel sentence pairs (including 430K pairs across the 20 non-Esperanto conlangs) and 321K vocabulary entries. In bidirectional translation experiments, we find that models perform better on a posteriori conlangs, whose vocabularies are derived from natural languages, reflecting the design characteristics of conlangs. Training on ConlangBench also shows that models can learn all eight conlangs for which sufficient parallel corpora are available, while their learning curves vary depending on how the conlangs were created. Our findings suggest that conlangs provide a unique testbed for investigating how LLMs acquire low-resource languages.
Chinese Translation
构造语言(conlangs)是有意创造的人类语言,具有丰富的语言创造传统。尽管它们在研究大型语言模型(LLMs)中的语言学习方面具有潜力,但现有的构造语言在LLM研究中仍然未得到充分探索。我们提出了ConlangBench,这是第一个针对21种现有构造语言评估和训练LLMs的大规模基准。我们收集了超过2100万对构造语言-英语的平行句子对(包括20种非世界语构造语言中的43万对)和321K词汇条目。在双向翻译实验中,我们发现模型在后验构造语言上的表现更佳,这些语言的词汇源自自然语言,反映了构造语言的设计特征。在ConlangBench上的训练还表明,模型可以学习所有八种具有足够平行语料的构造语言,而它们的学习曲线则因构造语言的创建方式而异。我们的研究结果表明,构造语言为调查LLMs如何获取低资源语言提供了独特的实验平台。
cs.CL / 54 / 2608.03507
ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
ChronoLens:跨时间、语言和语言层次测量语言变化
Abstract
Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and therefore cannot determine whether they evolve together across languages. We address this problem by asking how the magnitude and direction of change vary across linguistic levels, languages, and historical periods within a single analytical space. We introduce ChronoLens, a framework that combines frozen multilingual language models, feature-aligned crosscoders, and post-hoc linguistic interventions, and apply it to 44.98 million documents and approximately 17.2 billion tokens from five parliamentary traditions spanning 1803--2026. The resulting sparse representations agree substantially more strongly with linguistic statistics than dense embeddings or a pooled sparse autoencoder ($\rho=0.72$ versus $0.29$ and $0.28$), and reveal that morphology, syntax, semantics, and pragmatics generally change by comparable amounts within a language, while languages differ markedly in when, how far, and in which direction they change. These findings show that historical language change is a structured, multidimensional process: similar magnitudes can conceal different trajectories, and meaningful cross-linguistic comparison requires measuring both distance and direction.
Chinese Translation
历史语言变化影响形态学、句法学、语义学和语用学,但计算研究通常以不兼容的表征来考察这些层次,因此无法确定它们是否在不同语言中共同演变。我们通过询问变化的幅度和方向在单一分析空间内如何在语言层次、语言和历史时期之间变化来解决这一问题。我们引入了ChronoLens,一个结合了冻结的多语言语言模型、特征对齐的交叉编码器和事后语言干预的框架,并将其应用于从1803年至2026年跨越五个议会传统的4498万份文档和约172亿个标记。结果显示,稀疏表征在语言统计上与密集嵌入或汇总稀疏自编码器相比($
ho=0.72$ 对比 $0.29$ 和 $0.28$)一致性显著更强,并揭示出形态学、句法学、语义学和语用学通常在一个语言内部以相似的幅度变化,而不同语言在变化的时间、幅度和方向上则存在显著差异。这些发现表明,历史语言变化是一个结构化的多维过程:相似的幅度可能掩盖不同的轨迹,而有意义的跨语言比较需要同时测量距离和方向。
cs.CL / 55 / 2608.03529
Consensus Measures for Unstructured Biomedical Text Annotations
非结构化生物医学文本注释的共识度量
Abstract
Biomedical literature is increasingly mined for knowledge beyond the questions it was written to answer. Because the target concepts are not known in advance, annotators prefer open-ended labels, whose agreement is hard to quantify. We study soft inter-rater reliability for annotators providing unstructured texts for biomedical annotation tasks. Synthetic experiments show that soft reliability can be quantified using a variety of semantic equivalence measures, and that the choice of measure affects failure modes of the estimation. Embeddings are scalable, but limited when differentiating similar but distinct concepts. Large language models are promising, but limited by scalability for estimating agreement by chance. Finally, we suggest measures based on natural language inference as a sensible compromise.
Chinese Translation
生物医学文献越来越多地被挖掘以获取超出其原始问题的知识。由于目标概念事先并不明确,注释者倾向于使用开放式标签,这使得其一致性难以量化。我们研究了为生物医学注释任务提供非结构化文本的注释者的软评分者间可靠性。合成实验表明,软可靠性可以通过多种语义等价度量进行量化,并且度量的选择影响估计的失败模式。嵌入方法具有可扩展性,但在区分相似但不同的概念时有限。大型语言模型前景广阔,但在通过偶然性估计一致性时受到可扩展性的限制。最后,我们建议基于自然语言推理的度量作为一种合理的折衷方案。
cs.CL / 56 / 2608.03532
Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili
大型语言模型中的跨语言偏见:英语与斯瓦希里语的比较分析
Abstract
Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric. We investigate whether social biases generalise across languages by submitting 4,900 symmetric English--Swahili prompt pairs to GPT-5.2 and Gemini 2.5 Flash across nine demographic bias axes, yielding 19,600 completions evaluated for stereotype prevalence, sentiment, refusal behaviour, and cross-lingual semantic similarity. Our findings show that bias transforms rather than transfers: stereotype rates shifted by up to 12 percentage points on specific axes, Gemini's neutral-sentiment rate doubled in Swahili, and GPT-5.2 refused 169 prompts in English and zero in Swahili, consistent with refusal behaviour anchored to English-language surface forms at the behavioural level. Over 55% of prompt pairs produced semantically dissimilar completions across both models. These reinforce the idea that English-only bias audits do not produce adequate coverage for multilingual deployment.
Chinese Translation
大型语言模型在多语言环境中的应用日益增加,但安全对齐和偏见评估仍然以英语为中心。我们通过向GPT-5.2和Gemini 2.5 Flash提交4900对对称的英语-斯瓦希里语提示,调查社会偏见是否在不同语言间普遍存在,涵盖九个人口统计偏见维度,产生了19600个完成结果,并评估了刻板印象的普遍性、情感、拒绝行为和跨语言语义相似性。我们的研究发现,偏见是转化而非转移的:在特定维度上,刻板印象的比例变化高达12个百分点,Gemini在斯瓦希里语中的中性情感比例翻倍,而GPT-5.2在英语中拒绝了169个提示,而在斯瓦希里语中则为零,这与在行为层面上与英语表面形式相关的拒绝行为一致。超过55%的提示对在两个模型中产生了语义上不相似的完成结果。这进一步强化了仅进行英语偏见审计并不足以为多语言部署提供充分覆盖的观点。
cs.CL / 57 / 2608.03545
Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning
Hi-TTRL:利用提示调节测试时强化学习的共识
Abstract
Test-time reinforcement learning (TTRL) improves the reasoning capabilities of large language models without labeled data by updating the policy with pseudo-labels constructed through majority voting. While effective, the reward signal assigned from majority voting is highly sensitive to consensus strength, defined as the frequency of the most common answer within a rollout group. In TTRL, consensus strength plays a dual role: it reflects both the reliability of the pseudo-label and the distribution of advantages. Low consensus can amplify updates from unreliable pseudo-labels through disproportionately large advantages, whereas high consensus reduces reward contrast and ultimately yields vanishing gradients. In this paper, we introduce Hi-TTRL, a test-time reinforcement learning framework that utilizes hints during sampling to regulate rollout consensus strength. Hi-TTRL first estimates consensus strength from a partial rollout group. When the consensus strength falls outside a target interval, it invokes a Markov chain Monte Carlo (MCMC) hint sampler. The sampler targets the power-transformed prefix distribution and uses finite-step approximate sampling to generate rollout prefixes as hints. By tuning the power exponent, Hi-TTRL generates hints with a sharpened or flattened power target, steering rollout consensus strength toward the target interval. Experiments on multiple datasets and backbones show that Hi-TTRL consistently improves over standard TTRL, with ablations and consensus-steering analyses validating the effectiveness of adaptive hint-guided consensus regulation.
Chinese Translation
测试时强化学习(TTRL)通过使用多数投票构建的伪标签更新策略,从而在没有标注数据的情况下提升大型语言模型的推理能力。尽管有效,但来自多数投票的奖励信号对共识强度高度敏感,共识强度被定义为在一次回滚组中最常见答案的频率。在TTRL中,共识强度扮演着双重角色:它既反映了伪标签的可靠性,也反映了优势的分布。低共识可能通过不成比例的大优势放大来自不可靠伪标签的更新,而高共识则减少奖励对比,最终导致梯度消失。本文介绍了Hi-TTRL,一个在采样过程中利用提示来调节回滚共识强度的测试时强化学习框架。Hi-TTRL首先从部分回滚组中估计共识强度。当共识强度超出目标区间时,它调用马尔可夫链蒙特卡洛(MCMC)提示采样器。该采样器以幂变换的前缀分布为目标,并使用有限步近似采样生成回滚前缀作为提示。通过调节幂指数,Hi-TTRL生成具有锐化或平坦化幂目标的提示,引导回滚共识强度朝向目标区间。在多个数据集和骨干网络上的实验表明,Hi-TTRL在标准TTRL的基础上持续改进,通过消融实验和共识引导分析验证了自适应提示引导共识调节的有效性。
cs.CL / 58 / 2608.03573
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs
SFT冲突,RL共存:对大型语言模型多任务学习的理论与实证分析
Abstract
Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT suffers from severe task conflicts under multi-stage training, whereas RL enables stable coexistence across diverse tasks. Empirically, we trace this to the parameter level, observing that RL induces sparse and approximately orthogonal updates across tasks. We provide a theoretical explanation for this mechanism by analyzing multi-task gradient interference. Our results reveal a distinction: interference in SFT is norm-limited, scaling with the absolute gradient magnitude, whereas interference in RL is variance-limited, bounded by the gradient variance induced by advantage normalization and on-policy optimization. This small variance bound yields near-orthogonal optimization directions across tasks. Leveraging this insight, we propose Parallel-RL, a paradigm that decouples multi-task training, significantly improving efficiency and flexibility.
Chinese Translation
监督微调(Supervised Fine-Tuning, SFT)和强化学习(Reinforcement Learning, RL)在增强大型语言模型(Large Language Models, LLMs)多任务推理方面表现出根本不同的行为。我们的初步实验揭示了一种现象:在多阶段训练中,SFT遭遇严重的任务冲突,而RL则能够在多样化任务之间实现稳定共存。从经验上看,我们将这一现象追溯到参数层面,观察到RL在任务之间引发稀疏且近似正交的更新。我们通过分析多任务梯度干扰为这一机制提供了理论解释。我们的结果揭示了一种区别:SFT中的干扰是范数限制的,随着绝对梯度大小的增加而增加,而RL中的干扰是方差限制的,由优势归一化和策略优化引起的梯度方差所限制。这一小方差界限在任务之间产生近正交的优化方向。基于这一见解,我们提出了Parallel-RL,一种解耦多任务训练的范式,显著提高了效率和灵活性。
cs.CL / 59 / 2608.03577
Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation
在错误的路灯下寻找:关于自动翻译质量评估的局限性
Abstract
Automation of Translation Quality Estimation (QE) has emerged as a widely discussed approach to managing translation quality at scale, and a growing number of tools and technologies have been released in pursuit of this goal. However, the proliferation of new QE systems has not always been accompanied by robust, transparent, and reproducible research and testing. This gap deserves critical scrutiny. This paper examines some fundamental limitations of the QE technology from both theoretical and empirical perspectives, arguing that current QE systems are structurally ill-equipped to serve as reliable standalone tools in real-world translation workflows. The reviewed evidence suggests that QE suffers from a range of interrelated and largely unresolved limitations. Most fundamentally, the evaluation of the quality of translation at the level of isolated segments is problematic because it tends to miss out on cohesion, coherence, and stylistic and rhetorical text features. In addition, empirical research documents several other limitations and flaws, including failure to generalize, systematic biases, overfitting and distribution collapse, performance gaps, error annotation challenges, and data scarcity. These are structural limitations arising from the complexity of human language and translation as a cognitive and communicative act - limitations that more data and better architectures have so far not overcome. Consequently, segment-level QE scores should not be used as a standalone basis for routing, release, or review bypass in production; we argue future work should focus on automating human evaluation grounded in MQM.
Chinese Translation
翻译质量评估(QE)的自动化已成为管理大规模翻译质量的广泛讨论的方法,越来越多的工具和技术应运而生以追求这一目标。然而,新QE系统的激增并不总是伴随着稳健、透明和可重复的研究与测试。这一差距值得深入审视。本文从理论和实证的角度探讨了QE技术的一些基本局限性,认为当前的QE系统在结构上无法作为现实翻译工作流程中可靠的独立工具。所审查的证据表明,QE存在一系列相互关联且大多未解决的局限性。最根本的是,在孤立片段层面评估翻译质量是有问题的,因为这往往忽视了文本的连贯性、一致性以及风格和修辞特征。此外,实证研究记录了其他几种局限性和缺陷,包括无法推广、系统性偏见、过拟合和分布崩溃、性能差距、错误标注挑战和数据稀缺。这些是由于人类语言和翻译作为一种认知和交际行为的复杂性而产生的结构性局限性——这些局限性迄今为止尚未被更多数据和更好的架构所克服。因此,片段级QE评分不应作为生产中路由、发布或审查绕过的独立依据;我们认为未来的工作应集中于基于MQM的自动化人类评估。
cs.CL / 60 / 2608.03599
Disentangling Language Modeling and Boundaries
解构语言建模与边界
Abstract
Byte-level language models are usually argued for on the grounds of robustness, multilingual fairness, and character-level skills. We point to a different, structural advantage: because they read and write bytes, any two of them share an output space, so knowledge transfer between them is exact and independent of how either was originally tokenized. We hypothesize that the two distributions a byte-level model produces, one over the next byte, one over where its patch boundaries fall, can be disentangled and changed almost independently. A model could absorb a teacher's capability while keeping its own boundaries, or change how it places those boundaries while keeping its capabilities. We lay out the two experiments that would settle the hypothesis, alongside preliminary measurements of the properties they rest on. We argue that the community should move toward a byte-level interface as a shared standard: if the hypothesis holds, then once byte-level models are the norm, transferring capabilities and reshaping boundaries between them become cheap and routine, free of the per-model tokenizer that blocks them today.
Chinese Translation
字节级语言模型通常被认为具有鲁棒性、多语言公平性和字符级技能等优点。我们指出一个不同的结构性优势:由于它们以字节为单位进行读写,任何两个字节级模型共享同一个输出空间,因此它们之间的知识转移是精确的,并且与它们最初的标记化方式无关。我们假设字节级模型生成的两个分布,一个是下一个字节的分布,另一个是其补丁边界的分布,可以几乎独立地解构和改变。一个模型可以在保持自身边界的同时吸收教师的能力,或者在保持自身能力的同时改变其边界的放置方式。我们提出了两个实验来验证这一假设,并提供了它们所依赖的性质的初步测量结果。我们认为,学术界应朝着字节级接口作为共享标准的方向发展:如果假设成立,那么一旦字节级模型成为常态,能力的转移和边界的重塑将变得便宜且常规,摆脱当前阻碍它们的每个模型的标记器。
cs.CL / 61 / 2608.03610
Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR
基于语言专门化的多教师在线蒸馏用于多语言大规模语音识别
Abstract
Modern LLM-based ASR systems have established multilingual capability as a standard feature, leveraging large-scale multilingual corpora and LLMs' cross-lingual knowledge to achieve competitive performance across multilingual benchmarks. However, joint modeling of languages with heterogeneous acoustic, phonological, and lexical characteristics inevitably introduces optimization conflicts, undermining language-wise specialization. To address this challenge, we propose Language-Specialized Multi-Teacher On-Policy Distillation (LS-MOPD), which decouples language-specific knowledge acquisition from multilingual capability integration: language-specialized teachers are independently optimized via reinforcement learning (RL), after which their expertise is integrated into a generalist multilingual student through language routing and token-level multi-teacher distillation, thereby reducing direct cross-lingual optimization conflicts. We further explore two acoustic-prefix configurations, static and dynamic, to examine how teacher--student prefix consistency influences the efficacy of on-policy distillation. Experiments on benchmarks covering Mandarin, Mandarin subdialects, Cantonese, and English demonstrate that LS-MOPD substantially outperforms RL baselines and consistently surpasses the empirical performance envelope defined by best-performing RL teachers, revealing its potential to generalize beyond all teachers in multilingual ASR.
Chinese Translation
现代基于大规模语言模型(LLM)的语音识别(ASR)系统已将多语言能力作为标准特性,利用大规模多语言语料库和LLM的跨语言知识,在多语言基准测试中实现了竞争力的性能。然而,具有异质声学、音位和词汇特征的语言的联合建模不可避免地引入了优化冲突,从而削弱了语言特定的专业化。为了解决这一挑战,我们提出了基于语言专门化的多教师在线蒸馏(LS-MOPD),该方法将语言特定知识的获取与多语言能力的整合解耦:语言专门化教师通过强化学习(RL)独立优化,然后通过语言路由和基于标记的多教师蒸馏将其专业知识整合到一个通用的多语言学生中,从而减少直接的跨语言优化冲突。我们进一步探索了两种声学前缀配置,静态和动态,以研究教师-学生前缀一致性如何影响在线蒸馏的有效性。在涵盖普通话、普通话方言、粤语和英语的基准测试中的实验表明,LS-MOPD显著优于RL基线,并且始终超越由表现最佳的RL教师定义的经验性能边界,揭示了其在多语言ASR中超越所有教师的潜力。
cs.CL / 62 / 2608.03617
A machine-readable catalogue of the Tsiolkovsky papers (fond 555, Archive of the Russian Academy of Sciences), and a way to measure how well its handwriting can be read
一份机器可读的齐奥尔科夫斯基文献目录(档案编号555,俄罗斯科学院档案馆),以及衡量其手写文本可读性的方法
Abstract
The personal archive of Konstantin Tsiolkovsky (1857-1935) is held as fond 555 of the Archive of the Russian Academy of Sciences. The archive scanned the fond and published the images, but with no queryable catalogue, no full-text search and no dataset: the holdings can only be browsed one page at a time. This paper describes a machine-readable catalogue of all 2,019 files and 51,008 scans, a dating for 1,969 files taken from the archive's own descriptions, a page-level classification of every scan into handwriting and typescript, and a growing corpus of machine transcriptions (currently 322 files, 5,454 scans). It also reports a way to measure handwritten-text-recognition accuracy in an archive with no ground truth. Archives of the typewriter era often preserve one text twice, as manuscript and as a typed copy; transcribing both and comparing isolates the reading error, since source and pipeline are identical and only page difficulty differs. Across 294 such pairs from 27 files, two readings of a handwritten page agree on a median 37% of words. On two files that also have a published edition the estimate can be checked against ground truth: it is unbiased to within a percentage point and ranks pages as the truth does (rank correlation 0.92 where the edition is a faithful witness). This bounds use: two variants of one work here share 19% of words, below the rate at which two readings of a single page agree, so the redactions cannot be collated word by word at this quality. That negative result is reported as such, and the constraint is built into the tool.
Chinese Translation
康斯坦丁·齐奥尔科夫斯基(1857-1935)的个人档案被归档为俄罗斯科学院档案馆的档案编号555。该档案已对其进行扫描并发布了图像,但没有可查询的目录、全文搜索和数据集:档案只能逐页浏览。本文描述了一份机器可读的目录,包括所有2,019个文件和51,008个扫描件,基于档案自身描述对1,969个文件进行了日期标注,对每个扫描件进行了手写和打字稿的逐页分类,并建立了一个不断增长的机器转录语料库(目前包括322个文件和5,454个扫描件)。此外,本文还报告了一种在没有真实数据的档案中测量手写文本识别准确性的方法。打字机时代的档案通常会以手稿和打字本的形式保留同一文本两次;对这两种文本进行转录并比较,可以孤立出阅读错误,因为来源和处理过程是相同的,仅有页面难度不同。在27个文件中的294对这样的文本中,对一页手写文本的两次阅读在中位数上有37%的单词一致。在两个也有已出版版本的文件中,可以将估计值与真实数据进行核对:其偏差在一个百分点内,并且页面的排名与真实数据一致(排名相关性为0.92,当该版本是可信的见证时)。这限制了使用:这里同一作品的两个变体共享19%的单词,低于对单一页面的两次阅读一致的比率,因此这些修订无法以此质量逐字对照。该负面结果被如实报告,并且这一限制被纳入工具设计中。
cs.CL / 63 / 2608.03624
LoopMTP: A looped transformer guided by latent multi-token prediction
LoopMTP:一种由潜在多标记预测引导的循环变换器
Abstract
Looped transformers have emerged as a parameter-efficient alternative to scaling depth for strong reasoning. By reusing one stack of layers across $T$ iterations, they attain the effective depth and reasoning capabilities of larger models at a fixed parameter count. Yet existing approaches suffer from latent overthinking and undifferentiated computation, largely because intermediate representations receive no guidance across loops. Multi-token prediction (MTP) supplies exactly the dense, forward-looking supervision the loop is missing. We propose \textsc{LoopMTP}, which links the two through a structural correspondence in latent space: a model that loops $T$ times can anticipate $T$ future tokens. \textsc{LoopMTP} realizes this by softly aligning the hidden state of loop $t$ with the embedding of the token $t$ steps ahead, while a lightweight gate preserves useful information across iterations. \textsc{LoopMTP} improves average accuracy by up to 8.1\% (relative) over the non-looped baseline, with training remaining stable for up to 15 loops.
Chinese Translation
循环变换器作为一种参数高效的替代方案,已成为在增强推理能力方面深度扩展的有效选择。通过在 $T$ 次迭代中重用一组层,它们在固定的参数数量下实现了较大模型的有效深度和推理能力。然而,现有方法存在潜在的过度思考和计算无差异的问题,主要是因为中间表示在循环中没有得到指导。多标记预测(MTP)恰好提供了循环所缺乏的密集、前瞻性的监督。我们提出了 extsc{LoopMTP},通过潜在空间中的结构对应将两者连接起来:一个循环 $T$ 次的模型可以预测 $T$ 个未来标记。 extsc{LoopMTP} 通过将循环 $t$ 的隐藏状态与 $t$ 步前的标记嵌入进行软对齐来实现这一点,同时一个轻量级的门控机制在迭代中保留有用的信息。 extsc{LoopMTP} 在非循环基线的基础上,平均准确率提高了高达 8.1\%(相对),并且训练在最多 15 次循环中保持稳定。
cs.CL / 64 / 2608.03655
Decoupling Generation and Selection for Budget-Constrained Faithful Summarization
预算约束下的忠实摘要生成与选择解耦
Abstract
Abstractive summarization models remain vulnerable to factual inconsistency, redundancy, and weak length control. We propose a modular generation-and-selection framework for sentence-budget-constrained summarization. A pretrained generator produces multiple candidate summaries, which are decomposed into sentence-level candidates. A combinatorial selector then constructs the final summary by balancing relevance, factuality, and redundancy under an explicit budget. The framework supports MMR, ILP, and a DPP-inspired log-determinant objective without retraining the generator. Experiments on CNN/DailyMail, Multi-News, FaithBench, and TofuEval show consistent improvements in factuality and source-grounding metrics, especially for multi-document summarization, at the cost of lower reference-overlap scores. Human evaluation further indicates higher perceived consistency, relevance, clarity, and conciseness, with a small reduction in coherence. These results show that decoupling generation from selection provides a model-agnostic mechanism for improving factual grounding. Code is available at https://anonymous.4open.science/r/bcfs-D05E/.
Chinese Translation
抽象摘要模型仍然容易受到事实不一致、冗余和长度控制不足的影响。我们提出了一种模块化的生成与选择框架,用于句子预算约束的摘要生成。一个预训练的生成器产生多个候选摘要,这些摘要被分解为句子级候选项。然后,一个组合选择器在明确的预算下,通过平衡相关性、事实性和冗余性来构建最终摘要。该框架支持最大边际相关性(MMR)、整数线性规划(ILP)和受DPP启发的对数行列式目标,而无需重新训练生成器。在CNN/DailyMail、Multi-News、FaithBench和TofuEval上的实验显示,在事实性和源基础度量上有一致的改善,尤其是在多文档摘要中,尽管参考重叠分数较低。人工评估进一步表明,在一致性、相关性、清晰性和简洁性方面的感知得分更高,但连贯性略有下降。这些结果表明,将生成与选择解耦提供了一种与模型无关的机制,以改善事实基础。代码可在 https://anonymous.4open.science/r/bcfs-D05E/ 获取。
cs.CL / 65 / 2608.03659
How Closely Do LLM Reviews Align with Human Peer Review?
大型语言模型的评审与人类同行评审的对齐程度如何?
Abstract
Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting. We compare reviews from OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6 with human reviews and final decisions for 300 topic-matched ICLR 2026 submissions, equally divided among oral, poster, and rejected papers. Each model reviewed every paper using identical instructions and rating scales after decision information was removed. Our study contributes a cross-provider analysis of three complementary dimensions: alignment with broad and fine-grained decision categories, differences in recommendation-scale usage, and thematic agreement in identified weaknesses. All three LLMs distinguished accepted from rejected papers, but none reproduced the oral versus poster distinction present in human ratings. Scoring patterns were provider-specific: Gemini assigned systematically higher ratings, while OpenAI and Claude were closer to humans for rejected and poster papers but more critical of oral papers. Human and LLM reviews also differed in emphasis, with LLMs more frequently identifying missing baseline comparisons and humans more often raising computational-efficiency concerns. These results show that broad decision alignment does not imply agreement with finer human judgments or reviewing priorities.
Chinese Translation
大型语言模型(LLMs)越来越多地被用于生成科学评审,但现有评估很少考察不同提供者在同一受控环境中与会议决策和人类评审优先事项的一致性。我们比较了OpenAI GPT-5.4、Google Gemini 3.1 Pro Preview和Anthropic Claude Opus 4.6的评审与300份主题匹配的ICLR 2026提交稿的人类评审和最终决策,这些提交稿在口头、海报和被拒稿之间均匀分配。每个模型在去除决策信息后,使用相同的指令和评分标准对每篇论文进行了评审。我们的研究提供了三种互补维度的跨提供者分析:与广泛和细粒度决策类别的一致性、推荐评分使用的差异,以及在识别弱点时的主题一致性。所有三个LLM都能区分被接受和被拒绝的论文,但没有一个模型再现人类评分中存在的口头与海报的区别。评分模式具有提供者特异性:Gemini系统性地给予更高的评分,而OpenAI和Claude在被拒绝和海报论文上与人类评分更接近,但对口头论文则更为苛刻。人类和LLM的评审在强调上也存在差异,LLM更频繁地识别缺失的基线比较,而人类则更常提出计算效率问题。这些结果表明,广泛的决策一致性并不意味着与更细致的人类判断或评审优先事项的一致性。
cs.CL / 66 / 2608.03675
VetScore: Risk-Weighted Fact Verification for Veterinary Long-Form QA with Citations
VetScore:针对兽医长篇问答的风险加权事实验证与引用
Abstract
Citation excerpts can be used to increase the reliability of generated outputs and their faithfulness to cited sources, which is especially important in high-stakes domains such as human and veterinary medicine. However, this does not guarantee that generated claims are faithful to the provided excerpts. We present VetScore, a multi-step evaluation method for veterinary long-form question answering, designed to assess how well are generated claims supported by the provided excerpts, weighing this information by each claim's harm potential. VetScore first segments the output and decomposes it into individual claims, then scores each claim with respect to its harm potential and evaluates its faithfulness to source excerpts, and finally calculates the overall risk-adjusted score. We collect an expert-annotated meta-evaluation dataset, evaluate our approach with a range of judge models, and show that it achieves high correlations with veterinary experts even with small judge models, while offering explainability across multiple dimensions.
Chinese Translation
引用摘录可以用来提高生成输出的可靠性及其对引用来源的忠实度,这在如人类和兽医医学等高风险领域尤为重要。然而,这并不能保证生成的声明忠实于提供的摘录。我们提出了VetScore,这是一种针对兽医长篇问答的多步骤评估方法,旨在评估生成的声明在多大程度上得到提供摘录的支持,并根据每个声明的潜在危害对这些信息进行加权。VetScore首先对输出进行分段,并将其分解为单独的声明,然后根据其潜在危害对每个声明进行评分,并评估其对来源摘录的忠实度,最后计算整体风险调整得分。我们收集了一个专家注释的元评估数据集,使用一系列评判模型评估我们的方法,并显示即使在小型评判模型下,它也能与兽医专家实现高相关性,同时在多个维度上提供可解释性。
cs.CL / 67 / 2608.03709
Predicting Deep Neural Network Training Outcomes from Early Training Telemetry
从早期训练遥测预测深度神经网络训练结果
Abstract
Large hyperparameter sweeps for deep neural networks spend substantial compute on configurations that are effectively doomed from the first few epochs. We study whether a single training run's own early telemetry - per-epoch loss, training accuracy, gradient signal-to-noise ratio, weight-norm growth, and an activation-saturation snapshot - together with its sampled hyperparameters, can predict that run's eventual outcome without reference to other runs. We evaluate three prediction tasks: final test accuracy, relative performance within a domain, and training-dynamics failure, including numerical divergence. Across 23,788 training runs spanning six architecture/dataset combinations, gradient-boosted trees using only the first five epochs of telemetry achieve R^2 = 0.92-0.99 for final-accuracy regression and ROC-AUC = 0.983-0.998 for relative classification on a permanently held-out set of hyperparameter configurations. Useful prediction is already available after a single epoch. A paired ablation shows that gradient- and weight-level telemetry provides a statistically consistent improvement over loss and accuracy curves alone, although the practical gain varies by domain. Transfer is strong between similar architectures, while cross-dataset transfer is limited mainly by differences in accuracy scale rather than loss of the underlying relationship. These results suggest that early-training telemetry can provide a practical decision-support signal for compute allocation while motivating human oversight for any automated intervention.
Chinese Translation
对于深度神经网络的大规模超参数搜索,往往会在从最初几个训练周期开始就注定失败的配置上消耗大量计算资源。我们研究了单次训练运行的早期遥测数据——每个周期的损失、训练准确率、梯度信噪比、权重范数增长以及激活饱和快照——结合其采样的超参数,是否能够在不参考其他运行的情况下预测该运行的最终结果。我们评估了三个预测任务:最终测试准确率、在特定领域内的相对性能,以及训练动态失败(包括数值发散)。在涵盖六种架构/数据集组合的23,788次训练运行中,仅使用前五个周期的遥测数据的梯度提升树模型在最终准确率回归中达到了R^2 = 0.92-0.99,在相对分类中达到了ROC-AUC = 0.983-0.998,且这些结果是在一个永久保留的超参数配置集上获得的。经过一个周期后,已经可以获得有用的预测。配对消融实验表明,梯度和权重级别的遥测数据在统计上显著优于仅使用损失和准确率曲线,尽管实际增益因领域而异。相似架构之间的迁移效果良好,而跨数据集的迁移主要受到准确率规模差异的限制,而不是底层关系的丧失。这些结果表明,早期训练遥测可以为计算资源分配提供实用的决策支持信号,同时激励对任何自动干预进行人工监督。
cs.CL / 68 / 2608.03720
Detecting Hallucinations and Recovering Verified Answers in Arabic Islamic Question Answering
检测阿拉伯伊斯兰问答中的幻觉并恢复验证答案
Abstract
Large language models can generate fluent responses to Islamic questions while introducing factual errors that are difficult to identify. This paper presents our system for \textsc{HalluScoring 2026} Task 2.1, \textit{Islamic Hallucination Detection and Find the Truth}. The task requires a unified two-step prediction: determining whether an Arabic answer generated by an LLM is hallucinated and selecting the verified answer from six closely related candidate options. We use the Islamic knowledge dataset provided by the shared task, which contains 600 question--answer instances, including 341 hallucinated and 259 non-hallucinated answers. Our system is based on the fine-tuned \texttt{google/gemma-4-12B-it} model and uses deterministic decoding during inference. The generated outputs are normalized to extract the hallucination label and the selected option. The system achieves a Macro-F1 score of 0.928 and a label accuracy of 0.935 for hallucination detection, together with an option accuracy of 0.895 for answer selection. These results yield a combined score of 0.912, demonstrating strong performance across both stages of the task. The lower option-selection accuracy indicates that distinguishing the verified answer from plausible alternatives remains more challenging than detecting hallucinated responses.
Chinese Translation
大型语言模型能够流利地回答伊斯兰问题,但同时也会引入难以识别的事实错误。本文介绍了我们在 extsc{HalluScoring 2026} 任务 2.1 中的系统,任务名为 extit{伊斯兰幻觉检测与真相查找}。该任务要求进行统一的两步预测:判断由大型语言模型生成的阿拉伯语答案是否为幻觉,并从六个密切相关的候选选项中选择验证答案。我们使用了共享任务提供的伊斯兰知识数据集,其中包含600个问答实例,包括341个幻觉答案和259个非幻觉答案。我们的系统基于微调后的 exttt{google/gemma-4-12B-it} 模型,并在推理过程中使用确定性解码。生成的输出经过规范化以提取幻觉标签和所选选项。该系统在幻觉检测中实现了0.928的宏F1分数和0.935的标签准确率,同时在答案选择中获得了0.895的选项准确率。这些结果的综合得分为0.912,表明在任务的两个阶段均表现出色。较低的选项选择准确率表明,从可信的替代答案中区分验证答案仍然比检测幻觉响应更具挑战性。
cs.CL / 69 / 2608.03729
GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models
GPTKB 2.0:直接从大型语言模型构建消歧义知识库
Abstract
Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source. However, LLMs natively possess no representation of entities, leading to duplicate entries as well as conflations. We propose GPTKB 2.0, a methodology for constructing disambiguated KBs directly from LLMs. GPTKB 2.0 incorporates on-the-fly disambiguation of entities, relations and classes, and is meticulously designed to satisfy both scalability and disambiguation accuracy. We analyze the central design decisions and characterize the trade-offs between accuracy, scale, and cost. We execute GPTKB 2.0 at scale, obtaining a materialized KB containing over 1M disambiguated entities and 38.4M triples. This represents the first million-scale LLM-native KB with explicit internal canonicalization of entities, relations, and classes, a significant departure from prior Wikimedia-centric works. GPTKB 2.0 is available at https://gptkb.org/.
Chinese Translation
自动化知识库构建(AKBC)是自然语言处理(NLP)的核心任务,近期的研究提出直接从大型语言模型(LLMs)生成知识库,将模型本身视为知识源。然而,LLMs 本身并不具备实体的表示,导致重复条目和混淆现象的出现。我们提出了 GPTKB 2.0,一种直接从 LLMs 构建消歧义知识库的方法。GPTKB 2.0 结合了对实体、关系和类别的即时消歧义,并经过精心设计以满足可扩展性和消歧义准确性。我们分析了核心设计决策,并描述了准确性、规模和成本之间的权衡。我们在大规模上执行了 GPTKB 2.0,获得了一个包含超过 100 万个消歧义实体和 3840 万个三元组的物化知识库。这是第一个具有明确内部规范化实体、关系和类别的百万规模 LLM 原生知识库,显著不同于以往以维基媒体为中心的研究。GPTKB 2.0 可在 https://gptkb.org/ 获取。
cs.CL / 70 / 2608.03769
MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models
MDLMPE:面向分布的掩码扩散语言模型位置编码
Abstract
Masked diffusion language models (MDLMs) enable parallel generation and bidirectional context modeling, but their positional context differs fundamentally from that of autoregressive (AR) models. Whereas AR decoding exposes a contiguous prefix, MDLM denoising produces dynamic, non-contiguous configurations of revealed and masked tokens. Conventional positional encodings such as RoPE capture sequence order and pairwise displacement but remain insensitive to this evolving token-availability structure. To address this limitation, we propose MDLMPE, a positional encoding designed specifically for masked diffusion. To the best of our knowledge, MDLMPE is the first method to make positional representations explicitly aware of the changing revealed/masked configuration. It represents token availability as a binary sequence, applies distance-aware Gaussian weighting, and projects the resulting pattern through a cosine basis to obtain distribution-aware positional features. These features are added to token embeddings and mapped by a lightweight MLP to angular offsets that modulate the standard RoPE phases. Extensive experiments on LLaDA and DREAM demonstrate that MDLMPE generally outperforms conventional positional encoding methods across supervised fine-tuning, pretraining, zero-shot evaluation, and block-diffusion settings. Further ablations show that the complete combination of availability state, Gaussian locality, spectral basis, and embedding injection yields the strongest result. These results establish the evolving token-availability distribution as a useful positional signal for masked diffusion language models.
Chinese Translation
掩码扩散语言模型(MDLMs)实现了并行生成和双向上下文建模,但其位置上下文与自回归(AR)模型有根本性的不同。自回归解码暴露了连续的前缀,而MDLM去噪则产生了动态的、非连续的显现和掩码标记配置。传统的位置编码如RoPE捕捉序列顺序和成对位移,但对这种不断变化的标记可用性结构却缺乏敏感性。为了解决这一局限性,我们提出了MDLMPE,一种专门为掩码扩散设计的位置编码。据我们所知,MDLMPE是首个明确考虑显现/掩码配置变化的位置表示方法。它将标记可用性表示为二进制序列,应用基于距离的高斯加权,并通过余弦基投影得到面向分布的位置特征。这些特征被添加到标记嵌入中,并通过轻量级多层感知机(MLP)映射到调制标准RoPE相位的角度偏移量。对LLaDA和DREAM的广泛实验表明,MDLMPE在监督微调、预训练、零样本评估和块扩散设置中普遍优于传统的位置编码方法。进一步的消融实验显示,标记可用性状态、高斯局部性、谱基和嵌入注入的完整组合产生了最强的结果。这些结果确立了不断变化的标记可用性分布作为掩码扩散语言模型的有用位置信号。
cs.CL / 71 / 2608.03796
Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss
大规模语言模型的高效知识蒸馏:离线Top-K Logits与融合分块KL损失
Abstract
Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained from scratch: a compressed model is usually recovered through knowledge distillation (KD). This recovery step largely decides the final quality, yet it is expensive. We present a practitioner's study of how to make distillation training efficient, organised around two systems contributions. First, we show that offline KD (caching the teacher's top-$K$ logits once and training the student against the cache) matches online distillation at near-identical training loss while removing the teacher from memory, running about 29\% faster per iteration, and reaching up to 41\% higher throughput on a single H200 GPU. Second, we introduce a \emph{fused, chunked KL loss} that never materialises the full vocabulary-sized logit tensor, making peak memory linear in the sequence length. This removes the memory spike that otherwise caps context length and lets us train at four times the context (32{,}768 tokens) on a single GPU. A separate output-head-only toy benchmark isolates the loss kernel and confirms its memory and iteration-rate scaling from 4K to 256K tokens. Together these make large-scale healing and hundreds of ablations affordable. We also report supporting ablations on loss design and sequence packing. We release our chunked-loss implementation: https://github.com/CompactifAI/Full-Chunked-KL-Loss.
Chinese Translation
在严格的延迟、成本和本地部署限制下,小型语言模型往往是唯一的部署选择,但它们很少是从头开始训练的:通常通过知识蒸馏(KD)恢复压缩模型。这个恢复步骤在很大程度上决定了最终质量,但代价昂贵。我们提出了一项实践研究,旨在提高蒸馏训练的效率,围绕两个系统贡献进行组织。首先,我们展示了离线KD(一次缓存教师的Top-$K$ logits,并对学生进行缓存训练)在接近相同的训练损失下与在线蒸馏匹配,同时从内存中移除教师,迭代速度提高约29 ext{%},在单个H200 GPU上达到高达41 ext{%}的吞吐量。其次,我们引入了一种 extit{融合分块KL损失},该损失从不生成完整词汇大小的logit张量,使得峰值内存与序列长度成线性关系。这消除了限制上下文长度的内存峰值,使我们能够在单个GPU上以四倍的上下文(32,768个标记)进行训练。一个单独的仅输出头的玩具基准测试隔离了损失内核,并确认其在4K到256K标记范围内的内存和迭代速率扩展。以上两者使得大规模的恢复和数百次消融实验变得可负担。我们还报告了关于损失设计和序列打包的支持性消融实验。我们发布了我们的分块损失实现: https://github.com/CompactifAI/Full-Chunked-KL-Loss。
cs.CL / 72 / 2608.03803
M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models
M-GATE:多语言语法、翻译准确性与大型语言模型效率基准
Abstract
Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency. We introduce M-GATE (Multilingual Grammar, Accuracy in Translation, and Efficiency), a benchmark of linguistic proficiency spanning 30 typologically diverse languages from high- to low-resource. M-GATE comprises three tasks: grammatical error detection on linguist-crafted, adversarially selected sentences that turn on hard, language-specific phenomena; round-trip translation of shared English sources across 29 target languages, scored by a three-provider LLM judge panel validated against professional annotators; and a supplementary tokenizer-efficiency measure. We evaluate over 50 models in more than 80 configurations. Fluency and proficiency come apart sharply: models that translate competently sit near chance on the adversarial grammar items, the best reaching a Matthews correlation coefficient (MCC) of only 0.36, and their errors lean systematically toward under-flagging, accepting ungrammatical text rather than raising false alarms. Translation quality closely tracks a language's share of pretraining data (r = 0.86 against log Common Crawl share), producing a steep low-resource penalty that is nonetheless narrowing with successive model releases. Enabling reasoning reliably improves translation, while its effect on error detection is smaller and for some models negative, so the best configuration is task-dependent. To resist contamination, test items are kept private behind a continuously updated public leaderboard, with illustrative examples released (https://m-gate.ai).
Chinese Translation
多语言语言模型在一百种或更多语言中被部署,然而大多数基准测试关注的是模型在某种语言中执行任务的能力,而非其对该语言的掌握程度,这混淆了流利度与熟练度。我们提出了M-GATE(多语言语法、翻译准确性与效率),这是一个涵盖30种类型多样的语言(从高资源到低资源)的语言能力基准。M-GATE包括三个任务:在语言学家精心设计的、经过对抗性选择的句子上进行语法错误检测,这些句子涉及困难的、特定于语言的现象;对29种目标语言进行共享英语来源的往返翻译,由一个由三位提供者组成的LLM评审小组评分,并与专业注释员进行验证;以及一个补充的分词器效率测量。我们评估了50多种模型在80多种配置下的表现。流利度与熟练度之间的差异明显:能够胜任翻译的模型在对抗性语法项目上的表现接近随机,最佳模型的马修斯相关系数(MCC)仅为0.36,并且它们的错误系统性地倾向于低标记,接受不合语法的文本而不是引发虚假警报。翻译质量与语言的预训练数据份额密切相关(r = 0.86 对数Common Crawl份额),产生了一个陡峭的低资源惩罚,尽管随着后续模型发布这一惩罚正在缩小。启用推理可靠地改善了翻译,而其对错误检测的影响较小,且对某些模型而言是负面的,因此最佳配置依赖于任务。为了防止污染,测试项目被保密,保存在一个不断更新的公共排行榜后面,并发布了说明性示例(https://m-gate.ai)。
cs.CL / 73 / 2608.03810
VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs
VIBE:一个基于VAD的实体中心情感分析基准,用于大语言模型输出
Abstract
Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a target may appear favorable or threatening, calm or conflictual, powerful or vulnerable. Existing work captures parts of this space through sentiment, favorability, and emotion benchmarks, but none combines target-directed VAD attribution, an explicit scorer contract, and a passport reporting format. We introduce VIBE, a benchmark for entity-centered affective profiling of LLM outputs in Valence-Arousal-Dominance (VAD) space. Its core contribution is a measurement contract: VIBE separates generation from external scoring, distinguishes scalar favorability, response-level VAD, and target-directed VAD, and reports profiles through an Affective Passport. Three empirical layers support the contract. H1 shows scalar favorability does not subsume arousal and dominance: valence findings are cross-validated (rV = 0.944 judge-human, rV = 0.954 inter-scorer); arousal and dominance are single-scorer directional estimates, not point-precise, consistent with known inter-annotator difficulty on these axes (rA = 0.495, rD = 0.702 among human annotators). H2 shows whole-response and target-directed VAD are different contracts: the same text can carry one affective tone overall while representing the named target differently. H3 is a protocol-drift diagnostic: elicitation conditions shift profiles, motivating context metadata in every affective report. These results motivate entity-centered affective profiling as a documented practice: profiles should be released with scorer identity, coverage, protocol, and interpretation limits.
Chinese Translation
大型语言模型常常描述社会显著目标,包括政治人物、国家、宗教、组织、历史事件和社会群体,同时编码情感框架与事实内容:一个目标可能显得有利或威胁、平静或冲突、强大或脆弱。现有研究通过情感、偏好性和情绪基准捕捉了这一领域的部分内容,但没有一个结合目标导向的VAD归因、明确的评分合同和护照报告格式。我们引入了VIBE,这是一个针对大型语言模型输出的实体中心情感分析基准,基于情感-唤醒-主导(Valence-Arousal-Dominance, VAD)空间。其核心贡献是一个测量合同:VIBE将生成与外部评分分开,区分标量偏好性、响应级别的VAD和目标导向的VAD,并通过情感护照报告分析结果。三个实证层面支持该合同。假设H1表明标量偏好性并不包含唤醒和主导性:情感值的发现经过交叉验证(rV = 0.944 评审-人类,rV = 0.954 评分者间);唤醒和主导性是单评分者的方向性估计,而非精确点值,这与已知的注释者间在这些维度上的困难一致(人类注释者间rA = 0.495,rD = 0.702)。假设H2表明整体响应和目标导向的VAD是不同的合同:同一文本可以整体上携带一种情感色调,同时对命名目标的表现却不同。假设H3是一个协议漂移诊断:引发条件会改变分析结果,促使在每个情感报告中加入上下文元数据。这些结果促使实体中心情感分析作为一种文档化的实践:分析结果应与评分者身份、覆盖范围、协议和解释限制一起发布。
cs.CL / 74 / 2608.03842
Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling
敏感性、因果性与修复的解耦:扰动鲁棒性及其尺度的逐层分析
Abstract
When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these three layer maps dissociate. Across a five-model panel we identify two propagation regimes - spike-and-suppress (Phi-3.5, Gemma-2-9B) and late-accumulation (Llama-3, Mistral, Qwen2.5-7B) - and on the two models meeting an 80% identity-patch gate, sensitivity and causality are anti-correlated (rho = -0.72 to -0.88). Within-family scaling on Qwen2.5 (1.5B to 14B) shows the late-accumulation signature strengthening monotonically with scale, corroborated on a second family. We propose cascade disruption as the mechanism behind the dissociation: adapters placed at causally implicated early layers break intact downstream computation, making diagnostic-flagged sites the worst adapter placements. A fixed-harness layer sweep across four models (3.8-8B) confirms the core prediction on chain-of-thought GSM8K - the flagged sites are the most damaging adapter windows on every adjudicable model - and is sign-consistent but strongly attenuated on a multiple-choice control, consistent with damage that compounds with generation length. The sweep yields practical guidance: a training-free LRD pre-screen and a default-deepest placement rule, though absolute gains over no-adapter baselines remain small. Finally, apparent gains from a representation-stability loss reverse under an adequate generation budget - truncated chain-of-thought had been scored as empty - a methodological warning for any intervention evaluated on chain-of-thought tasks.
Chinese Translation
当语言模型在表面扰动输入(如拼写错误、OCR噪声、同音词)上失败时,“哪个层次负责”有三种自然的操作化方式:表示最显著分歧的地方(敏感性)、恢复干净激活以恢复预测的地方(因果性)以及可以修复损害的小适配器所在的地方(补偿能力)——我们展示了这三种层次图的解耦。在五个模型的面板中,我们识别出两种传播机制——尖峰抑制(Phi-3.5, Gemma-2-9B)和晚期累积(Llama-3, Mistral, Qwen2.5-7B)——在两个满足80%身份补丁门的模型中,敏感性与因果性呈反相关(rho = -0.72至-0.88)。在Qwen2.5(1.5B到14B)的同家族尺度上,晚期累积特征随着规模单调增强,这在第二个家族中得到了证实。我们提出级联干扰作为解耦背后的机制:放置在因果相关的早期层的适配器破坏了完整的下游计算,使得被诊断标记的站点成为最糟糕的适配器放置位置。对四个模型(3.8-8B)进行的固定结构层次扫描确认了在链式思维GSM8K上的核心预测——被标记的站点是每个可裁决模型上最具破坏性的适配器窗口——并且在多项选择控制上保持符号一致但强烈减弱,这与生成长度增加的损害相一致。该扫描提供了实用指导:一种无训练的LRD预筛选和默认最深放置规则,尽管相对于无适配器基线的绝对增益仍然较小。最后,来自表示稳定性损失的明显增益在足够的生成预算下逆转——截断的链式思维被评估为空——这是对任何在链式思维任务上评估的干预的一个方法论警告。
cs.CL / 75 / 2608.03859
Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking
超越表征相似性:基于源条件的描述长度增益用于生成性抄袭检测和候选源重排序
Abstract
Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative plagiarism detection remains an underexplored and largely unresolved challenge. Prior work on LLM-generated-text detection targets AI involvement, which may be permissible, rather than source reuse, while similarity-based methods struggle after extensive rewriting and multi-source synthesis. Motivated by the description-length view of probabilistic prediction, in which relevant side information can reduce a target sequence's code length, we introduce Source-Conditioned Description-Length Gain (SCDG), a directional, training-free framework that contrasts a frozen language model's description length of a suspicious document $P$ with and without a candidate source $S$. This contrast yields token-level log-likelihood gains that measure the incremental predictive evidence supplied by $S$. We evaluate SCDG on the PAN at CLEF benchmarks for generative plagiarism. On a PAN 2025-derived pairwise benchmark, SCDG achieves 0.92 Precision, 0.97 Recall, and 0.94 F1, outperforming all baselines; on PAN 2026's multi-source retrieval task, it reaches 0.83 nDCG@10 and 0.96 Recall@100, surpassing all baselines. On a same-topic, same-event Multi-News test, the calibrated gain-distribution SCDG classifier predicts source reuse for only $0.125\%$ of pairs, supporting robustness to topical overlap under this evaluation protocol. These results establish SCDG as a unified and token-decomposable signal for source-specific content reuse under extensive transformation.
Chinese Translation
大型语言模型(LLMs)对学术诚信和同行评审提出了挑战。然而,生成性抄袭检测仍然是一个未被充分探索且大多未解决的难题。之前关于LLM生成文本检测的研究主要关注人工智能的参与,这可能是允许的,而不是源重用,而基于相似性的检测方法在经过广泛重写和多源合成后表现不佳。受到概率预测的描述长度视角的启发,其中相关的侧面信息可以减少目标序列的编码长度,我们引入了源条件描述长度增益(Source-Conditioned Description-Length Gain, SCDG),这是一个方向性、无训练的框架,比较了冻结语言模型对可疑文档 $P$ 的描述长度在有候选源 $S$ 和没有候选源 $S$ 时的差异。这种对比产生了逐词级别的对数似然增益,衡量了 $S$ 提供的增量预测证据。我们在CLEF的PAN基准上评估了SCDG在生成性抄袭检测中的表现。在一个基于PAN 2025的成对基准上,SCDG达到了0.92的精确率、0.97的召回率和0.94的F1值,超越了所有基线;在PAN 2026的多源检索任务中,SCDG达到了0.83的nDCG@10和0.96的召回率@100,超过了所有基线。在同主题、同事件的Multi-News测试中,经过校准的增益分布SCDG分类器仅预测 $0.125\%$ 的对中存在源重用,支持在该评估协议下对主题重叠的鲁棒性。这些结果确立了SCDG作为一个统一且可逐词分解的信号,用于在广泛变换下的源特定内容重用。
cs.CL / 76 / 2608.03860
SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG
SciRet:一种计算感知的科学检索增强生成的检索与重排序的实证研究
Abstract
We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19. Rather than proposing a new model, we evaluate a fixed scientific RAG pipeline across three corpus scales: 1,034 chunks (1K papers), 5,160 chunks (5K papers), and 15,480 chunks (15K papers). The pipeline combines sentence-window chunking, BM25, BGE-M3 dense retrieval, reciprocal rank fusion, optional cross-encoder reranking, and grounded answer generation. Across these settings, hybrid retrieval is more robust than either sparse-only or dense-only retrieval in our setting, reaching Recall@10 of 1.000 at 1K and 15K. In contrast, an MS MARCO-trained cross-encoder reranker reduces precision on the scientific corpus, suggesting that domain mismatch can outweigh the benefits of stronger query-passage interaction. Generation faithfulness measured with RAGAS increases with corpus scale in our setup. Retrieval evaluation uses pseudo-relevance labels derived from the hybrid system, so we treat the results as controlled comparative evidence rather than a benchmark claim. We release code, indexes, and evaluation outputs to support replication and follow-up studies.
Chinese Translation
我们介绍了SciRet,这是一项针对CORD-19的科学问答的检索增强生成的计算感知实证研究。我们并未提出一个新模型,而是评估了一个固定的科学检索增强生成(RAG)管道,涵盖三个语料库规模:1,034个块(1K篇论文)、5,160个块(5K篇论文)和15,480个块(15K篇论文)。该管道结合了句子窗口分块、BM25、BGE-M3密集检索、互惠排名融合、可选的交叉编码器重排序和基于上下文的答案生成。在这些设置中,混合检索在我们的设置中比仅稀疏检索或仅密集检索更具鲁棒性,在1K和15K时达到1.000的Recall@10。相比之下,经过MS MARCO训练的交叉编码器重排序器在科学语料库上降低了精度,表明领域不匹配可能会超过更强查询-段落交互的好处。在我们的设置中,使用RAGAS测量的生成忠实度随着语料库规模的增加而提高。检索评估使用从混合系统派生的伪相关标签,因此我们将结果视为受控比较证据,而非基准声明。我们发布了代码、索引和评估输出,以支持复制和后续研究。
cs.CL / 77 / 2608.03882
MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning
MultiGlobeQA:一个多语言和全球多样性的地理空间推理基准
Abstract
Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigation and logistics, yet large language models (LLMs) struggle with the required geometric and topological computation despite storing considerable geographic knowledge. Existing benchmarks localize these failures only partially: they are synthetic or smallscale, largely monolingual, and offer limited control over geographic coverage. We introduce MultiGlobeQA, a multilingual benchmark of 46,060 question-answer pairs spanning 14 spatial-function families and 15 answer formats, with execution-based ground truth over three knowledge graphs. It covers 201 countries and territories via income- and density-stratified sampling, with parallel questions in English and 16 additional high- and low-resource languages. Across parametric, reasoning, and agentic settings, LLMs collapse on tasks requiring grid indexing and shape computation, while topological relations and directions fare best. Retrieval and tool use yield considerable gains, yet performance plateaus below two thirds even when gold facts are supplied, indicating that computation, not access to knowledge, is the bottleneck. Models also underperform on low-income regions, a gap that gold facts widen rather than close.
Chinese Translation
地理空间推理,即对现实世界实体进行距离、包含及其他空间关系的计算,是导航和物流的核心。然而,尽管大型语言模型(LLMs)存储了大量的地理知识,但在所需的几何和拓扑计算方面却表现不佳。现有基准仅部分定位了这些失败:它们往往是合成的或小规模的,主要是单语的,并且对地理覆盖的控制有限。我们提出了MultiGlobeQA,这是一个包含46,060个问答对的多语言基准,涵盖14个空间功能类别和15种答案格式,并基于三个知识图谱提供执行基础的真实答案。它通过收入和密度分层抽样覆盖201个国家和地区,并在英语和16种其他高资源和低资源语言中提供平行问题。在参数设置、推理设置和代理设置中,LLMs在需要网格索引和形状计算的任务上表现不佳,而拓扑关系和方向的表现相对较好。检索和工具使用带来了显著的提升,但即使在提供黄金事实的情况下,性能也停滞在三分之二以下,表明计算而非知识获取是瓶颈。模型在低收入地区的表现也不佳,而黄金事实的提供并未缩小这一差距,反而加剧了这一问题。
cs.CL / 78 / 2608.03883
DS@GT-ARC at eRisk 2026 Task 3: Sparse, Semantic, and LLM Reranking for ADHD Symptom Sentences
DS@GT-ARC在eRisk 2026任务3中的表现:针对ADHD症状句子的稀疏、语义和LLM重排序
Abstract
This paper describes our submissions to eRisk 2026 Task 3, ADHD Symptom Sentence Ranking. The task requires systems to rank candidate Reddit sentences according to their relevance to each of the 18 symptoms in the Adult ADHD Self-Report Scale (ASRS-v1.1). Because no annotated training data were released for this first edition of the task, we relied on zero-shot experimentation, manual validation, and unsupervised or weakly guided retrieval pipelines. Our systems combine sparse BM25 retrieval, evidence-aware rescoring for self-referential symptom reports, embedding-based reranking, query-prototype expansion, and LLM-based reranking. All submitted systems follow a staged retrieval design in which BM25 retrieves candidates at scale and semantic or LLM rerankers refine the final rankings. Among our submissions, the LLM reranker achieved the strongest official scores, followed by the prototype query-expansion run. Our manual top-10 analysis aligned with the official expert scoring trend, suggesting that staged reranking is a promising direction for further development.
Chinese Translation
本文描述了我们在eRisk 2026任务3(ADHD症状句子排序)中的提交。该任务要求系统根据与成人ADHD自我报告量表(ASRS-v1.1)中18种症状的相关性对候选的Reddit句子进行排序。由于此次任务的首次版本未发布标注的训练数据,我们依赖于零样本实验、人工验证以及无监督或弱指导的检索管道。我们的系统结合了稀疏的BM25检索、自我指涉症状报告的证据感知重评分、基于嵌入的重排序、查询原型扩展以及基于LLM的重排序。所有提交的系统遵循分阶段检索设计,其中BM25大规模检索候选项,而语义或LLM重排序器则细化最终排名。在我们的提交中,LLM重排序器获得了最强的官方得分,其次是原型查询扩展的运行。我们的手动前10分析与官方专家评分趋势一致,表明分阶段重排序是进一步发展的有希望的方向。
cs.CL / 79 / 2608.03898
ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts
ANNOTARES:用于从德国法定文本中提取逻辑结构的数据集
Abstract
The automatic structural analysis of legal texts is a cornerstone of legal technology, yet the extraction of their logical components remains a significant challenge. In this paper, we introduce the task of identifying and segmenting legal conditions (Tatbestand) and legal consequences (Rechtsfolge) within German statutory texts. To support this task, we present ANNOTARES (Annotations of Tatbestand-Rechtsfolge Sequences), a novel dataset comprising German law texts with span-level annotations. Spanning three distinct legal codes, the dataset is designed to evaluate both domain-specific performance and cross-statute generalizability. We benchmark diverse architectural approaches: a rule-based baseline, CRFs, BiLSTMs, BiLSTM-CRF, and modern Transformer-based models, including BERT variants and LLM-based methods. Our results demonstrate that BERT and LLM-based models achieve superior performance in capturing the complex syntactic structures of legal language. We release our dataset to facilitate further research in automated legal reasoning.
Chinese Translation
法律文本的自动结构分析是法律技术的基石,但提取其逻辑成分仍然是一个重大挑战。本文介绍了在德国法定文本中识别和分割法律条件(Tatbestand)和法律后果(Rechtsfolge)的任务。为支持这一任务,我们提出了ANNOTARES(Tatbestand-Rechtsfolge序列注释),这是一个包含德国法律文本的全新数据集,具有跨度级注释。该数据集涵盖三个不同的法律法规,旨在评估领域特定性能和跨法规的普适性。我们基准测试了多种架构方法:基于规则的基线、条件随机场(CRFs)、双向长短期记忆网络(BiLSTMs)、BiLSTM-CRF以及现代基于Transformer的模型,包括BERT变体和基于大型语言模型(LLM)的方法。我们的结果表明,BERT和基于LLM的模型在捕捉法律语言复杂句法结构方面表现优越。我们发布了该数据集,以促进自动化法律推理的进一步研究。
cs.CL / 80 / 2608.03930
Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility
语言之前的逻辑:在形式推导上的预预训练促进技能获取和可压缩性
Abstract
Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. However, existing pre-pretraining tasks, such as Dyck and procedural algorithms, rely on narrow primitives that fail to capture the expressive capacity of natural language. Moreover, prior studies remain restricted to relatively small token budgets, offering limited insight into skill emergence and representational dynamics. To address these limitations, we propose logic pre-pretraining (Logic-PPT) as a principled initialization strategy, leveraging formal derivations to impart richer structural and linguistic biases. Formal derivations require abstract mechanisms that are central to natural language, simultaneously binding variables, connecting quantifiers and relational dependencies, and composing predicate-argument structures over long contexts. Scaling our evaluation to a 100B-token regime, logic pre-pretraining substantially accelerates skill acquisition in LMs, achieving 80\% accuracy on linguistic tasks with 36B fewer tokens than standard initialization, and outperforming alternative pre-pretraining baselines. Mechanistically, formal derivations induce persistent structural reorganization, distinctively characterized by a lower-rank, spectrally concentrated representation space. Crucially, we show that this internal geometry enables improved model compressibility via pruning, matching the dense baseline performance even at $\approx$33\% sparsity.
Chinese Translation
在符号数据上对语言模型(LMs)进行预预训练可以加速和改善自然语言的获取。然而,现有的预预训练任务,如Dyck和程序算法,依赖于狭窄的原语,无法捕捉自然语言的表达能力。此外,先前的研究仍然局限于相对较小的标记预算,提供了有限的技能出现和表征动态的洞察。为了解决这些局限性,我们提出逻辑预预训练(Logic-PPT)作为一种原则性初始化策略,利用形式推导来赋予更丰富的结构和语言偏见。形式推导需要抽象机制,这些机制是自然语言的核心,同时绑定变量、连接量词和关系依赖,并在长上下文中组合谓词-论元结构。将我们的评估扩展到100B标记的范围,逻辑预预训练显著加速了LMs中的技能获取,在语言任务上实现了80%的准确率,所需标记比标准初始化少36B,并且超越了其他预预训练基线。在机制上,形式推导引发了持久的结构重组,其特征是较低秩的、谱集中表示空间。至关重要的是,我们展示了这种内部几何结构通过剪枝提高了模型的可压缩性,即使在约33%的稀疏性下也能匹配密集基线的性能。
cs.CL / 81 / 2608.03966
HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification
HalluTruthQA-4K:用于阿拉伯语幻觉检测和事实验证的细粒度语料库及注释过程
Abstract
Large language models can generate fluent Arabic answers while introducing factual errors that are difficult to identify and verify. Existing Arabic hallucination resources often assign a binary label to an entire response, indicating whether it is hallucinated or non-hallucinated, but provide limited information about the exact erroneous content, the reason for the error, or the correct factual answer. We present HalluTruthQA-4K, an expanded version of the HalluTruthQA resource containing 4,000 expert-curated Arabic question-answering instances across four knowledge-intensive domains: Islamic knowledge, history, science, and geography. Serving as the official dataset for Track 2 of the HalluScoring 2026 shared task, HalluTruthQA-4K extends our original corpus to 4,000 instances. Each instance pairs an Arabic question with a model-generated response, a verified reference answer, and five plausible distractors. Hallucinated responses are additionally annotated with character-level erroneous spans, human-written explanations, and hierarchical hallucination types. The corpus contains 1,643 hallucinated and 2,357 non-hallucinated responses, with 1,843 annotated erroneous spans. We describe the resource construction and annotation methodology, including question selection, controlled answer generation, candidate construction, expert annotation, independent verification, adjudication, and quality control. We also document the annotation guidelines, taxonomy, data format, inter-annotator agreement, and corpus statistics. HalluTruthQA-4K provides a reusable resource for hallucination detection, span-level error localization, explanation generation, factual verification, and the broader evaluation of factual reliability in Arabic language models.
Chinese Translation
大型语言模型能够生成流畅的阿拉伯语回答,但同时引入了难以识别和验证的事实错误。现有的阿拉伯语幻觉资源通常对整个回答分配一个二元标签,指示其是否为幻觉,但对具体的错误内容、错误原因或正确的事实答案提供的信息有限。我们提出了HalluTruthQA-4K,这是HalluTruthQA资源的扩展版本,包含4,000个专家策划的阿拉伯语问答实例,涵盖四个知识密集型领域:伊斯兰知识、历史、科学和地理。作为HalluScoring 2026共享任务第2轨道的官方数据集,HalluTruthQA-4K将我们的原始语料库扩展至4,000个实例。每个实例将一个阿拉伯语问题与一个模型生成的回答、一个经过验证的参考答案和五个合理的干扰项配对。幻觉回答还附加了字符级错误跨度的注释、人类撰写的解释和分层幻觉类型。该语料库包含1,643个幻觉回答和2,357个非幻觉回答,以及1,843个注释的错误跨度。我们描述了资源构建和注释方法,包括问题选择、受控答案生成、候选构建、专家注释、独立验证、裁决和质量控制。我们还记录了注释指南、分类法、数据格式、注释者间一致性和语料库统计信息。HalluTruthQA-4K为幻觉检测、跨度级错误定位、解释生成、事实验证以及对阿拉伯语模型中事实可靠性的更广泛评估提供了可重用的资源。
cs.CL / 82 / 2608.03984
string2string Studio: An Interactive, In-Browser Platform for String-to-String Algorithms
string2string Studio:一个交互式的浏览器内字符串到字符串算法平台
Abstract
We present string2string Studio, an interactive in-browser platform for string-to-string analysis across natural language processing, computational biology, and the digital humanities. The system integrates six main modules (alignment, distance, similarity, search, generation metrics, and BLAST homology search), operating at character, word, token, line, and residue levels. Its C++-based algorithms compile to WebAssembly, so core operations run locally by default without any installation or data upload. The interface reports scores with their "evidence" (alignments, edit paths, metric matches, search hits, and homology traces), making methods inspectable, debuggable, and comparable on shared inputs. Internal benchmarks show speedups of up to 2,500x over the Python predecessor, faster global/local alignment than a general-purpose native C aligner, and exact agreement with independent references under declared settings. For homology search, the scoped client-side blastn path closely matches NCBI BLAST+ rankings and statistics under matched parameters. A curated showcase and Learn mode present canonical algorithms and metrics as reusable demonstrations. string2string Studio is open-source and freely available at string2string.org.
Chinese Translation
我们介绍了string2string Studio,这是一个用于自然语言处理、计算生物学和数字人文学科的交互式浏览器内字符串到字符串分析平台。该系统集成了六个主要模块(比对、距离、相似性、搜索、生成度量和BLAST同源搜索),在字符、单词、标记、行和残基级别上运行。其基于C++的算法编译为WebAssembly,因此核心操作默认在本地运行,无需任何安装或数据上传。界面报告分数及其“证据”(比对、编辑路径、度量匹配、搜索命中和同源痕迹),使得方法可供检查、调试和在共享输入上进行比较。内部基准测试显示,与Python前身相比,速度提升高达2500倍,全球/局部比对速度快于通用本地C比对器,并在声明的设置下与独立参考完全一致。对于同源搜索,范围限定的客户端blastn路径在匹配参数下与NCBI BLAST+的排名和统计数据密切匹配。一个策划的展示和学习模式提供了经典算法和度量作为可重用的演示。string2string Studio是开源的,免费提供于string2string.org。
cs.CL / 83 / 2608.03994
When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings
当注意力失去敏感性:ALiBi位置编码中的数值失效
Abstract
We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows floating-point precision, which zeroes out a large fraction of attention weights and renders the affected attention heads partially blind. We analyze this failure mode, characterize its impact, and examine four mitigation strategies. We further demonstrate its occurrence in state-of-the-art pretrained models based on ALiBi. Comprehensive pretraining experiments with 148M-parameter decoder models help us to disentangle its effects from out-of-context degradation. We find that ALiBi's failure mode can substantially impair token retrieval while having only a minor effect on standard decoder benchmarks. We propose four training-time mitigation strategies and evaluate them individually and in combinations, finding that log-scaled distances yield the most consistent improvements in passkey retrieval. Despite this problem, default ALiBi slopes remain a surprisingly strong baseline, particularly for needle-in-a-haystack retrieval. Based on these findings we provide concrete recommendations on how to train models with ALiBi.
Chinese Translation
我们识别出ALiBi位置编码中一个先前被忽视的失效模式:其线性偏置缩放导致浮点精度下溢,从而使大量注意力权重归零,导致受影响的注意力头部分失去敏感性。我们分析了这一失效模式,表征其影响,并考察了四种缓解策略。我们进一步展示了这一现象在基于ALiBi的最先进预训练模型中的发生。通过148M参数解码器模型的全面预训练实验,我们帮助解开了其影响与上下文退化之间的关系。我们发现,ALiBi的失效模式可能显著削弱标记检索,同时对标准解码器基准测试的影响较小。我们提出了四种训练时的缓解策略,并分别及组合评估它们,发现对数缩放距离在检索密码时提供了最一致的改进。尽管存在这一问题,默认的ALiBi斜率仍然是一个令人惊讶的强基线,特别是在“针在干草堆中”检索的情况下。基于这些发现,我们提供了关于如何使用ALiBi训练模型的具体建议。
cs.CL / 84 / 2608.04003
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
PAST-Bench:个人智能体递归自我改进基础的基准测试
Abstract
Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for studying this capability because they retain preferences, task histories, tool routines, and learned skills across sessions. Yet whether retained experience actually improves them over time has not been systematically tested. We introduce PAST-Bench, a benchmark designed to isolate this question. Each agent runs through ordered sequences of fresh-session tasks under matched conditions that turn retained experience on and off. It spans 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and update. We report both later-task gains and whether those gains follow the intended save, retrieve, and update pathway. Across seven base models and four agent frameworks, improvement is real but uneven across capabilities. Agents with the same headline gain can differ markedly in whether that gain is supported by evidence of the intended pathway. Guided by these findings, we develop Hermes+, which extends Hermes with five targeted interventions across stages of the agent loop. Hermes+ raises the average gain from retained experience and provides clearer pathway evidence, with its strongest improvement on tasks requiring outdated state to be replaced, although the effect remains capability- and model-dependent. Together, PAST-Bench and Hermes+ provide an evaluation and diagnostic foundation for studying how persistent agents can progress from retaining experience to systematically improving through it. Code: https://github.com/Gen-Verse/PAST-Bench
Chinese Translation
递归自我改进要求智能体将积累的经验转化为更好的未来行为。个人人工智能智能体提供了一个具体的环境来研究这一能力,因为它们在多个会话中保留了偏好、任务历史、工具例程和学习技能。然而,保留的经验是否真正随着时间的推移而改善智能体的表现尚未经过系统测试。我们引入了PAST-Bench,一个旨在孤立这一问题的基准测试。每个智能体在匹配条件下依次执行新会话任务,开启和关闭保留经验。该基准涵盖了26个场景和204个情节,涉及记忆、程序重用、信息收集和更新。我们报告了后续任务的收益以及这些收益是否遵循预期的保存、检索和更新路径。在七个基础模型和四个智能体框架中,改进是真实存在的,但在能力上存在不均衡。具有相同表面收益的智能体在该收益是否得到预期路径证据的支持上可能存在显著差异。基于这些发现,我们开发了Hermes+,它在智能体循环的各个阶段扩展了Hermes,增加了五个有针对性的干预措施。Hermes+提高了从保留经验中获得的平均收益,并提供了更清晰的路径证据,其在需要替换过时状态的任务上表现出最强的改进,尽管这一效果仍然依赖于能力和模型。总之,PAST-Bench和Hermes+为研究持久智能体如何从保留经验到系统性改进提供了评估和诊断基础。代码链接:https://github.com/Gen-Verse/PAST-Bench
cs.CL / 85 / 2608.04007
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning
TurnSight:面向工具集成推理的回合级后见自蒸馏
Abstract
Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-level supervision, limiting fine-grained credit assignment in long-horizon TIR scenarios. On-policy self-distillation offers denser signals through teacher branches with privileged context, but existing approaches typically derive such context from ground-truth answers or retrieved skills, which may not reflect the states actually visited by the agent. Moreover, token-level supervision fails to capture the turn-level structure of tool interactions. To address this, we propose TurnSight, a turn-level hindsight self-distillation framework that derives supervision directly from execution-conditioned hindsight. It then constructs multiple hindsight views with different lookahead horizons and selects reliable supervision through cross-horizon directional agreement. Finally, the selected hindsight signal is normalized across sibling rollouts and used to adaptively modulate RL advantages while preserving their original optimization direction. Extensive experiments on three benchmarks demonstrate the effectiveness of TurnSight. Our codes are available at https://github.com/quchangle1/TurnSight.
Chinese Translation
工具集成推理(Tool-Integrated Reasoning, TIR)使大型语言模型(LLMs)能够通过迭代工具交互解决复杂任务。然而,现有的强化学习方法通常依赖于轨迹级监督,这限制了在长时间范围TIR场景中的细粒度信用分配。基于策略的自蒸馏通过具有特权上下文的教师分支提供了更密集的信号,但现有方法通常从真实答案或检索的技能中获取此类上下文,这可能无法反映代理实际访问的状态。此外,基于标记的监督未能捕捉工具交互的回合级结构。为了解决这一问题,我们提出了TurnSight,一种回合级后见自蒸馏框架,直接从执行条件的后见中导出监督。然后,它构建多个具有不同前瞻视野的后见视图,并通过跨视野方向一致性选择可靠的监督。最后,所选的后见信号在兄弟回合之间进行归一化,并用于自适应调节强化学习优势,同时保留其原始优化方向。在三个基准上的大量实验表明了TurnSight的有效性。我们的代码可在 https://github.com/quchangle1/TurnSight 获得。
cs.CL / 86 / 2608.04008
WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament
世界杯竞技场:对前沿大型语言模型在实时赛事中的前瞻性、无泄漏评估
Abstract
Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evaluation must defend itself against memorisation. We report the opposite design. Over the 39 days of the 2026 FIFA World Cup, six frontier LLMs -- all with extended thinking and native server-side web search -- were asked before every kickoff, one match at a time, to fill in a seven-market prediction card for all 104 matches, plus 12 group winners and a pre-tournament outright pool; no answer existed when the question was asked, so the evaluation is leakage-free by construction rather than by filtering, and the frozen archive holds 4,494 scored predictions. What the tournament establishes is a set of behaviours the six systems share. On match outcome they average 63.9%, level with backing the bookmaker's favourite -- which is in fact what they usually do. They agree with one another far more often than they are right, so a majority vote adds nothing. They under-commit to draws and to goals, and crowd their scoreline picks onto a single prototypical result. Accuracy tracks how lopsided a fixture is rather than how much is known about it: it collapses in the closest ties, where the dossiers are richest, while questions about the tournament as a whole are answered well. On this task the current generation of frontier systems is not sharply differentiated: the standings hold up at the top and the bottom across the run and churn in the middle, and the margins stay narrow throughout. The briefing dossiers, fixtures and official results are released as a benchmark, together with the scoring code.
Chinese Translation
测量大型语言模型预测能力的基准几乎总是回顾性的:事件已经发生,答案在网络上某处存在,评估必须抵御记忆化的影响。我们报告了相反的设计。在2026年国际足联世界杯的39天内,六个前沿大型语言模型——均具备扩展思维和本地服务器端网络搜索能力——在每场比赛开球前被要求逐场填写一张包含七个市场的预测卡,涵盖所有104场比赛、12个小组冠军以及一项赛前的总冠军池;在提问时并不存在答案,因此评估是通过构建而非过滤实现的无泄漏,并且冻结的档案中保存了4,494个评分预测。该赛事确立了一组六个系统共享的行为。在比赛结果上,它们的平均准确率为63.9%,与支持博彩公司热门球队的表现相当——这实际上也是它们通常的做法。它们之间的意见一致性远高于正确率,因此多数投票并没有增加价值。它们对平局和进球的预测不足,并将得分选择集中在一个典型结果上。准确性跟踪的是比赛的倾斜程度,而非已知信息的多少:在最接近的比赛中,档案最丰富时准确性崩溃,而对整个赛事的问题回答得很好。在这一任务中,当前一代前沿系统并没有明显区分:排名在顶部和底部保持稳定,而中间部分则波动,边际始终保持狭窄。简报档案、比赛安排和官方结果作为基准发布,同时附上评分代码。
cs.CL / 87 / 2608.04009
SocietyBench: Forecasting Counterfactual Social-World Evolution
社会基准:预测反事实社会世界演变
Abstract
Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how well a model understands and forecasts the way real social events unfold, has barely been measured. We introduce SocietyBench, an end-to-end benchmark that takes a one-line event topic, collects Web news and social-media posts across five platforms, distills them into a date-indexed timeline that keeps factual events and a public-opinion layer separate, and then turns every cutoff date on that timeline into an audited bank of forecasting questions. Questions are scored on two orthogonal 100-point axes: probability calibration and temporal accuracy. Before any model sees a timeline, a three-phase procedure replaces every named entity and shifts every date by a per-event constant, turning a real arc into a counterfactual social world -- structurally identical to what happened, but stripped of the surface labels a model could match against pre-training memory. On five heterogeneous events and 125 prediction points in Chinese and English editions, the strongest of six frontier LLMs reaches only 75.0 out of 100, against a trivial anchor of 50. The two axes come apart: a model can be calibration-strong but time-weak, or the reverse. Three agent frameworks built on a shared base model fail to improve on that base, and two model-free heuristics trail every LLM. Per-event gaps reach 21.4 points on a single axis, which is our main argument for evaluating on several events rather than one. All anonymized timelines, question banks, ground truth, and scoring code are released.
Chinese Translation
大型语言模型(LLMs)及其上构建的智能体目前主要通过完成任务的能力进行评估——修复错误、驱动浏览器、操作图形用户界面。然而,模型理解和预测真实社会事件展开方式的能力这一补充社会能力几乎没有被衡量。我们提出了社会基准(SocietyBench),这是一个端到端的基准测试,它接收一个单行事件主题,从五个平台收集网络新闻和社交媒体帖子,将其提炼为一个日期索引的时间线,保持事实事件和公众舆论层的分离,然后将时间线上的每个截止日期转化为经过审计的预测问题库。问题在两个正交的100分轴上进行评分:概率校准和时间准确性。在任何模型看到时间线之前,一个三阶段程序会替换每个命名实体,并将每个日期按事件常量进行偏移,将真实的事件弧转变为一个反事实社会世界——在结构上与发生的事件相同,但去除了模型可以与预训练记忆匹配的表面标签。在中文和英文版本中,针对五个异质事件和125个预测点,六个前沿LLMs中最强的模型仅获得75.0分(满分100),而一个简单的基准得分为50。这两个轴是相互独立的:一个模型可以在校准上表现强劲但在时间上较弱,反之亦然。基于共享基础模型构建的三个智能体框架未能改善该基础,而两个无模型启发式方法在每个LLM之后落后。每个事件的差距在单一轴上达到21.4分,这也是我们主张在多个事件上进行评估而非单一事件的主要论据。所有匿名化的时间线、问题库、真实情况和评分代码均已发布。