cs.RO / 1 / 2608.20433
Humanoid Musical Robots as Experimental Interfaces for Music-Evoked Emotion
类人音乐机器人作为音乐引发情感的实验接口
Abstract
Advances in technology have led to increasingly sophisticated musical humanoid robots. However, their use has largely been limited to performance and related research in human-robot interaction. In this position paper, we propose a novel perspective: musical humanoid robots as experimental interfaces for investigating music-evoked emotions. We argue that current research is constrained by paradigms relying on pre-recorded auditory stimuli, which fail to capture the multimodal, embodied, and interactive nature of real-world musical experience. Building on existing theories of music cognition and emotion, we identify mechanisms that require controlled manipulation of both acoustic and non-acoustic variables. We show that humanoid robots are well-suited as they enable parametric control of performance variables, reproducibility across trials, and the decoupling and recombination of auditory, visual, and interactive components. We illustrate the technical feasibility of this perspective through a case study of the WAseda Saxophonist Robot 5 (WAS-5), demonstrating reproducible control of acoustic and interaction variables that are prerequisites for future music-emotion experiments. Our work positions musical humanoid robots as a methodological platform that enables future controlled investigations of music-evoked emotions.
Chinese Translation
技术的进步导致了越来越复杂的音乐类人机器人。然而,它们的使用在很大程度上局限于表演和人机交互相关的研究。在这篇立场论文中,我们提出了一种新颖的视角:将音乐类人机器人视为研究音乐引发情感的实验接口。我们认为,当前的研究受到依赖于预录音频刺激的范式的限制,这些范式未能捕捉到现实世界音乐体验的多模态、具身和互动特性。在现有的音乐认知和情感理论的基础上,我们识别出需要对声学和非声学变量进行控制操控的机制。我们展示了类人机器人非常适合这一研究,因为它们能够对表演变量进行参数控制、在试验中实现可重复性,并解耦和重组听觉、视觉和互动组件。我们通过WAseda萨克斯演奏机器人5号(WAS-5)的案例研究,展示了这一视角的技术可行性,证明了声学和互动变量的可重复控制是未来音乐情感实验的先决条件。我们的工作将音乐类人机器人定位为一种方法论平台,能够支持未来对音乐引发情感的控制性研究。
cs.RO / 2 / 2608.20478
EndoLIFT: Language-Disambiguated Latent-Conditioned Rectified Flow for Bidirectional Endoscopic Control
EndoLIFT:用于双向内窥镜控制的语言消歧潜在条件整流流
Abstract
Routine gastrointestinal endoscopy is intrinsically bidirectional: the instrument is advanced to reach target anatomy and later withdrawn or retroflexed for inspection, while an external cue may require earlier reversal. When the requested phase changes before the visual scene does, nearly identical observations can require opposite axial actions. We identify and formalize this ambiguity in bidirectional endoscopic control as intent aliasing. We propose EndoLIFT (Endoscopic Language-Instruction Flow with Trajectory Latents), a vision-language-action policy that combines explicit language-based intent conditioning with a latent-conditioned rectified-flow action expert. The policy receives RGB, a language instruction, and the previous-action state; a 32-D variational trajectory latent stochastically conditions continuous action-chunk generation. Controlled same-observation instruction swaps establish that language selects the axial mode, independently of whether the trajectory latent is present. Relative to the matched model without latent conditioning, EndoLIFT improves navigation-direction accuracy by 11.1 percentage points and reduces wrong-direction advance by 83\%. An architecture-controlled 1-bit mode-flag reference exhibits weaker canonical-anchor switching, while EndoLIFT retains 82.8\% intent-following accuracy across 44 held-out linguistic variants. In closed-loop evaluation, EndoLIFT improves overall success by 30 percentage points over EndoLIFT w/o VTL on both the seen colon phantom and the unseen lung and stomach phantoms, and completes 10/10 ex-vivo porcine-trachea trials. These results separate language-based intent selection from the trajectory latent's contribution to directional correctness and robust retraction.
Chinese Translation
常规的胃肠内窥镜检查本质上是双向的:仪器被推进以到达目标解剖结构,随后被撤回或反折以进行检查,而外部提示可能要求提前反转。当请求的阶段在视觉场景变化之前发生变化时,几乎相同的观察结果可能需要相反的轴向操作。我们识别并正式化了这种双向内窥镜控制中的模糊性,称之为意图混淆。我们提出了EndoLIFT(内窥镜语言指令流与轨迹潜变量),这是一种结合了显式基于语言的意图条件和潜在条件整流流动作专家的视觉-语言-动作策略。该策略接收RGB图像、语言指令和前一个动作状态;一个32维的变分轨迹潜变量随机条件化连续动作块的生成。受控的相同观察指令交换表明,语言选择轴向模式,与轨迹潜变量是否存在无关。与没有潜在条件的匹配模型相比,EndoLIFT将导航方向的准确性提高了11.1个百分点,并将错误方向推进减少了83%。一个架构控制的1位模式标志参考展示了较弱的典范锚点切换,而EndoLIFT在44个保留的语言变体中保持了82.8%的意图跟随准确性。在闭环评估中,EndoLIFT在可见的结肠模型和不可见的肺部及胃部模型上,相较于不使用变分轨迹潜变量的EndoLIFT提高了30个百分点的整体成功率,并在10/10的离体猪气管试验中完成了所有试验。这些结果将基于语言的意图选择与轨迹潜变量对方向正确性和稳健撤回的贡献区分开来。
cs.RO / 3 / 2608.20546
Koala Gripper: Co-designing Robotic Grippers and Data-Capture Devices for Scaling Dexterous Manipulation Learning
考拉抓手:为扩展灵巧操作学习而共同设计机器人抓手和数据采集设备
Abstract
As the demand for larger manipulation datasets grows, handheld robotic gripper data collection and the associated gripper designs become more vital. Current data collection device designs trend towards matching the morphologies of existing robotic grippers, sacrificing ergonomics and manipulation performance. In this paper, we propose a co-design framework that guides the simultaneous development of both data collection and robotic execution devices by weaving both platform constraints into the design process. Through this workflow, we present the Koala Gripper system, a data capture device and robotic gripper platform that improves dexterity and grasp capability compared to parallel jaw grippers while preserving scalability and ease-of-use. The design introduces a novel force-optimized finger/trigger linkage mechanism with directional reflected mass characteristics, a unique monolithic dual-thumb, and user-centered ergonomic design. The design's actuated robotic fingers are backdrivable, with effective mass on the order of tens of grams. We show that these grippers are capable of secure grasps over a wide range of objects, forceful tool use, and precise singulation. We further validate the platform by deploying it with an end-to-end data collection and policy execution pipeline that highlights its capabilities through learning from demonstration. More information available at http://koalagripper.rai-inst.com
Chinese Translation
随着对更大操作数据集需求的增长,手持机器人抓手的数据收集及其相关设计变得愈发重要。目前的数据收集设备设计趋向于匹配现有机器人抓手的形态,牺牲了人机工程学和操作性能。本文提出了一种共同设计框架,通过将平台约束融入设计过程,指导数据收集设备和机器人执行设备的同步开发。通过这一工作流程,我们展示了考拉抓手系统,这是一种数据采集设备和机器人抓手平台,与平行爪抓手相比,它在灵活性和抓取能力上有所提升,同时保持了可扩展性和易用性。该设计引入了一种新颖的力优化指/触发器联动机制,具有方向性反射质量特性,独特的单体双拇指设计,以及以用户为中心的人体工程学设计。该设计的驱动机器人手指可反向驱动,有效质量在几十克的量级。我们展示了这些抓手能够在广泛的物体上实现安全抓取、强力工具使用和精确分离。我们进一步通过部署一个端到端的数据收集和策略执行管道来验证该平台,该管道通过示范学习突出其能力。更多信息请访问 http://koalagripper.rai-inst.com
cs.RO / 4 / 2608.20556
Logic-VLA: A Temporal Logic Conditioned Vision-Language-Action Model
Logic-VLA:一种基于时序逻辑的条件视觉-语言-动作模型
Abstract
Vision-language-action (VLA) models can follow natural-language (NL) task instructions, but such instructions may not precisely specify safety-critical or spatiotemporal requirements on the resulting behavior. We introduce Logic-VLA, a formal-requirement-aware VLA that conditions on Signal Temporal Logic (STL) specifications supplied at inference time. Logic-VLA uses a syntax-graph-based STL encoder pre-trained to capture temporal logic semantics. Policy adaptation proceeds in two stages: STL-conditioned supervised fine-tuning on satisfying demonstrations is followed by trajectory-level preference optimization over matched satisfying-violating rollout pairs using a flow-matching surrogate for Identity Preference Optimization. This formulation improves formal requirement satisfaction while preserving the nominal NL task. We evaluate Logic-VLA in closed-loop quadcopter navigation simulation across randomized photorealistic environments and test generalization to STL formulas unseen during training. Across the evaluation benchmarks, Logic-VLA improves STL satisfaction rate over an STL-blind base policy by 24.8 to 40.7 percentage points (pp) while reducing nominal NL task success by at most 1.8 pp, showing that a single VLA can adapt its behavior to varying formal requirements without requiring a separate policy for each specification.
Chinese Translation
视觉-语言-动作(VLA)模型能够遵循自然语言(NL)任务指令,但这些指令可能无法精确指定对结果行为的安全关键或时空要求。我们提出了Logic-VLA,这是一种在推理时基于信号时序逻辑(STL)规范进行条件化的正式需求感知VLA。Logic-VLA使用基于语法图的STL编码器,该编码器经过预训练以捕捉时序逻辑语义。策略适应分为两个阶段:首先在满足演示上进行STL条件的监督微调,然后在匹配的满足-违反回放对上使用流匹配替代方法进行轨迹级偏好优化,以实现身份偏好优化(Identity Preference Optimization)。这种形式化改善了正式需求的满足,同时保持了名义上的NL任务。我们在随机化的照片真实环境中评估Logic-VLA在闭环四旋翼导航仿真中的表现,并测试其对训练期间未见的STL公式的泛化能力。在评估基准中,Logic-VLA的STL满足率比盲目STL基线策略提高了24.8到40.7个百分点(pp),而名义NL任务的成功率最多减少1.8个百分点,表明单一的VLA能够根据不同的正式需求调整其行为,而无需为每个规范要求单独的策略。
cs.RO / 5 / 2608.20655
Nonlinear Model Predictive Control for Trajectory Tracking of Differentially Flat Fixed-Wing Aerial Systems
用于差分平坦固定翼空中系统轨迹跟踪的非线性模型预测控制
Abstract
Planning and control of fixed-wing Unmanned Aerial Vehicles (UAVs) are challenging due to nonlinear dynamics, aerodynamic limits, and environmental disturbances. Differential flatness offers a principled way to generate fast, feasible trajectories, but its use has largely been confined to model-free controllers, which lack predictive capabilities and demand tuning. In this paper, we propose a unified framework that integrates differential flatness-based trajectory generation with Nonlinear Model Predictive Control (NMPC), combining computationally efficient planning with predictive, constraint-aware control. To further improve robustness, we introduce a wind-aware sampling strategy embedded within the NMPC framework, enabling the generation of dynamically feasible reference trajectories that proactively account for wind disturbances while strictly enforcing aerodynamic and control input constraints. We validate the proposed framework through extensive simulations and real-world flight experiments, demonstrating improved tracking accuracy and robustness for complex trajectories, particularly when using the proposed wind-aware sampling strategy under strong wind conditions.
Chinese Translation
固定翼无人机(UAV)的规划与控制面临着非线性动力学、气动限制和环境干扰等挑战。差分平坦性提供了一种系统化的方法来生成快速、可行的轨迹,但其应用主要局限于无模型控制器,这些控制器缺乏预测能力并且需要调节。在本文中,我们提出了一个统一框架,将基于差分平坦性的轨迹生成与非线性模型预测控制(NMPC)相结合,结合了计算高效的规划与具有预测性和约束意识的控制。为了进一步提高鲁棒性,我们在NMPC框架中引入了一种考虑风的采样策略,使得生成的动态可行参考轨迹能够主动考虑风干扰,同时严格执行气动和控制输入约束。我们通过广泛的仿真和实际飞行实验验证了所提出框架的有效性,展示了在复杂轨迹下,尤其是在强风条件下使用所提出的风感知采样策略时,跟踪精度和鲁棒性的显著提升。
cs.RO / 6 / 2608.20784
Rethinking Demonstration Unlearning in Imitation Learning for Robotics
重新思考机器人模仿学习中的演示遗忘
Abstract
Imitation learning for robotics depends on human demonstrations, some of which people may later ask to remove. Retraining without them is the natural reference, but its cost grows with policy and dataset scale, motivating cheaper operators that edit a trained policy. Metrics inherited from machine unlearning, such as forgetting loss or a single membership attack, do not establish what an edit removed from a policy acting in closed loop. We therefore introduce a retrain-calibrated audit that reads demonstration unlearning along two axes: behavior, whether the edited policy acts like one retrained without the removed demonstrations, and evidence, whether an auditor can still detect it was trained on them. The behavior axis measures action divergence to that retrain at matched states, calibrated by a floor built from independent retrains, so a policy at the floor is as close to a retrain as retrains are to each other. The evidence axis applies a per-demonstration membership attack against a retrain null, reporting both its rank and its absolute member-loss level, since rank alone accepts operators that inflate member losses past the null. A conformal test then combines both axes into one hypothesis of joint retrain consistency, against a fleet of independent retrains large enough to reject at conventional significance. Across five preregistered conditions on three real-robot policy classes and two simulation suites, the axes dissociate in both directions on one checkpoint, as an edit may repair task behavior while leaving evidence unchanged, or reduce evidence while moving behavior away from retraining. On the ACT arm, a redirect edit restores blind-scored robot success to 18 of 20 trials.
Chinese Translation
机器人模仿学习依赖于人类的演示,其中一些演示可能会被要求删除。没有这些演示的重新训练是自然的参考,但其成本随着策略和数据集规模的增加而增长,这促使我们寻找更便宜的操作方法来编辑已训练的策略。从机器遗忘中继承的度量标准,如遗忘损失或单一成员攻击,并不能确定从闭环中操作的策略中删除了什么。因此,我们引入了一种重新训练校准审计,从两个维度读取演示遗忘:行为,即编辑后的策略是否像没有删除演示的重新训练策略那样行动;证据,即审计员是否仍能检测到其曾在这些演示上进行训练。行为维度测量在匹配状态下与重新训练的行动差异,通过独立重新训练构建的基准进行校准,因此在基准上的策略与重新训练的接近程度与彼此之间的重新训练相当。证据维度则对重新训练的无效性应用逐演示成员攻击,报告其排名和绝对成员损失水平,因为仅凭排名会接受那些使成员损失超过无效性的操作。然后,符合性检验将两个维度结合成一个关于联合重新训练一致性的假设,针对一组足够大的独立重新训练,以便在传统显著性水平下拒绝。在三个真实机器人策略类别和两个仿真套件的五个预注册条件下,这两个维度在一个检查点上在两个方向上分离,因为编辑可能修复任务行为而不改变证据,或者减少证据同时使行为远离重新训练。在ACT臂上,一个重定向编辑将盲评分机器人成功率恢复到20次试验中的18次。
cs.RO / 7 / 2608.20823
Natural Sit-to-Stand Motion Synthesis For Humanoids via Guided Assistance Curricula and Staged Rewards
通过引导辅助课程和分阶段奖励实现类人机器人自然坐立运动合成
Abstract
A humanoid has infinitely many ways to stand up from sitting while maintaining balance, making sit-to-stand (STS) a challenging control problem. We synthesise natural humanoid STS motion from scratch using reinforcement learning, without demonstrations or reference trajectories. A single Proximal Policy Optimisation policy learns smooth, human-like rising driven by three complementary components. (i) A coupled force/chair-height curriculum is used. A vertical pelvis-assist force aids early trajectory exploration and decays over training. Taller chairs are unlocked with decaying assisting force. This ensures that the policy masters a viable STS trajectory at each chair height before being exposed to harder ones, avoiding the premature distribution shift that otherwise collapses generalisation. (ii) Motion robustness is achieved by randomly sampling from a large number of inverse kinematics-generated initial and target poses spanning over eight chair heights. (iii) A set of rewards is defined inspired from biomechanics and optimal control studies. They shape the robot's angular momentum for seat-off, and enable support-region transition via centre of pressure attraction function to ensure smooth low-effort actuation. On a deterministic force-free evaluator, the policy attains more than 97% balanced-standing success across eight chair heights. The policy generalises smooth motion across chair heights and enables the robot to rise from substantially deep-seated postures as compared to the state of the art.
Chinese Translation
类人机器人在保持平衡的情况下,从坐姿站起的方式有无数种,这使得坐立(STS)成为一个具有挑战性的控制问题。我们利用强化学习从零开始合成自然的类人机器人STS运动,无需演示或参考轨迹。单一的近端策略优化(Proximal Policy Optimisation)策略通过三个互补组件学习平滑的人类般的站立动作。(i) 使用耦合的力/椅高课程。一个垂直的骨盆辅助力帮助早期轨迹探索,并在训练过程中逐渐减弱。随着辅助力的减弱,较高的椅子被解锁。这确保了策略在接触更困难的椅子之前,能够掌握每个椅高下可行的STS轨迹,避免了否则会导致泛化崩溃的过早分布转移。(ii) 通过随机从大量逆运动学生成的初始和目标姿态中采样,跨越八个椅高,实现运动的鲁棒性。(iii) 定义了一组受生物力学和最优控制研究启发的奖励。这些奖励塑造了机器人在离座时的角动量,并通过压力中心吸引函数实现支撑区域的过渡,以确保平滑低功耗的驱动。在一个确定性的无力评估器上,该策略在八个椅高下的平衡站立成功率超过97%。该策略在椅高之间实现了平滑运动的泛化,并使机器人能够从相较于现有技术更深的坐姿中站起。
cs.RO / 8 / 2608.20852
Demonstration-Guided Humanoid Stand-Up on an Emulated Deformable Surface
基于示范指导的人形机器人在仿真可变形表面上的起立运动
Abstract
This paper presents a reference-guided reinforcement learning framework to generate stand-up motion for a 29-DOF Unitree G1 humanoid on deformable soft ground, using a human demonstration recorded on hard ground. The terrain compliance is modelled using solref and solimp parameters from MuJoCo's rigid body soft-contact model. The rewards consists of (i) reference motion tracking through residual joint-position control and (ii) explicit recovery objectives such as pelvis height, torso uprightness, and the final posture. First, the policy is trained with the specified rewards considering hard ground. Next, the terrain stiffness is lowered by updating solref and the nominal surface penetration zone is expanded using solimp. Subsequent training enables the policy to adapt to the delayed support force generation due to significant surface penetration during contact-intensive phases while preserving the original demonstration pattern. The learned policy successfully completes the fallen-to-standing task in simulation, reaching the targeted pelvis height and uprightness, with a maximum contact penetration of approximately 40 mm during the process. The proposed method is demonstrated on two stand-up sequences and successfully achieves the final recovery objective on both hard and soft ground. Ablation studies show that reference tracking alone is insufficient for successful stand-up, and that explicit recovery rewards are essential.
Chinese Translation
本文提出了一种参考指导的强化学习框架,用于在可变形软地面上生成29自由度Unitree G1人形机器人的起立运动,使用在硬地面上录制的人类示范。地形的顺应性通过MuJoCo的刚体软接触模型中的solref和solimp参数进行建模。奖励包括(i)通过残余关节位置控制进行的参考运动跟踪,以及(ii)明确的恢复目标,如骨盆高度、躯干直立性和最终姿态。首先,考虑到硬地面,使用指定的奖励对策略进行训练。接下来,通过更新solref降低地形刚度,并使用solimp扩展名义表面穿透区。后续训练使得策略能够适应由于在接触密集阶段显著表面穿透而导致的延迟支撑力生成,同时保持原始示范模式。学习到的策略在仿真中成功完成了从倒地到站立的任务,达到了目标骨盆高度和直立性,在此过程中最大接触穿透约为40毫米。所提出的方法在两个起立序列上进行了演示,并在硬地面和软地面上成功实现了最终恢复目标。消融研究表明,仅依靠参考跟踪不足以成功起立,明确的恢复奖励是必不可少的。
cs.RO / 9 / 2608.20891
IMU-Free Body-Frame State Estimation with Sparse Scene Flow for Quadcopters
基于稀疏场景流的无IMU四旋翼机机体框架状态估计
Abstract
We present a vision-only state estimation system for X-configuration quadcopters equipped with a canonical stereo camera pair and no inertial sensors. The system operates entirely in the body frame, requiring only synchronised stereo images and motor thrust commands. A continuous-discrete extended Kalman filter on a composite manifold state $\langle SE(3), \mathbb{R}^3, \ldots \rangle$ maintains estimates of body-frame pose, velocity, angular velocity, gravity, and disturbances, using stationary scene points as implicit inertial references. Feature points are detected (FAST, Shi-Tomasi), tracked temporally (SSD, Lucas-Kanade) and matched across cameras (NCC), with search regions predicted from filter-derived pose and point uncertainty. Chi-squared gating on the normalised innovation admits only stationary points to the filter. The system also produces a sparse 3D point cloud carrying per-point position, velocity and joint covariance. These come from a 4-view (two stereo pairs at two timestamps) full bundle adjustment that jointly estimates position and velocity from stereo disparity and temporal parallax, with the filter-derived relative pose as a prior. Feature points in the EKF do not enter the solver; their information is reflected through the pose prior. Point cloud density is spatially adaptive: an external focus point directs allocation, producing dense coverage in the region of attention and sparse coverage elsewhere. The output is a body-frame state estimate, a calibrated pose change, and a sparse scene flow. It is intended as a measurement source for a downstream world model anchored in the current body frame, without dependence on GPS, IMU, or any world-frame infrastructure, though the architecture accommodates their future integration.
Chinese Translation
我们提出了一种仅基于视觉的状态估计系统,适用于配备标准立体相机对且不使用惯性传感器的X型四旋翼机。该系统完全在机体框架内运行,仅需同步的立体图像和电机推力命令。一个在复合流形状态$ riangleleft SE(3), extbf{R}^3, ext{等}
angle$上的连续离散扩展卡尔曼滤波器维护机体框架姿态、速度、角速度、重力和扰动的估计,使用静态场景点作为隐式惯性参考。特征点通过FAST和Shi-Tomasi算法检测,使用SSD和Lucas-Kanade算法进行时间跟踪,并通过归一化互相关(NCC)在相机间匹配,搜索区域根据滤波器推导的姿态和点的不确定性进行预测。基于卡方门限的归一化创新仅允许静态点进入滤波器。该系统还生成一个稀疏的三维点云,携带每个点的位置、速度和联合协方差。这些数据来自于一个四视图(两个时间戳下的两个立体对)全束调整,联合估计来自立体视差和时间视差的位置和速度,以滤波器推导的相对姿态作为先验。扩展卡尔曼滤波器中的特征点不进入求解器;它们的信息通过姿态先验反映。点云密度是空间自适应的:一个外部焦点指导分配,在关注区域产生密集覆盖,而在其他地方则是稀疏覆盖。输出为机体框架状态估计、校准的姿态变化和稀疏场景流。该系统旨在作为下游世界模型的测量源,锚定在当前机体框架中,无需依赖GPS、IMU或任何世界框架基础设施,尽管该架构允许未来的集成。
cs.RO / 10 / 2608.20904
Scalable Distributed Simulation-Based Testing for Automated Driving Systems
可扩展的基于仿真的自动驾驶系统分布式测试
Abstract
Virtual scenario-based testing is a key enabler for validating automated driving systems (ADS) and intelligent transport systems (ITS). However, executing large-scale test suites involving possibly thousands of scenarios remains labor-intensive and difficult to scale. This paper presents an end-to-end, DevOps-driven framework that automates build, deployment, and distributed execution of CARLA-based scenario tests of an ADS on a lightweight Kubernetes cluster. ROS 2 applications are packaged as standardized Kubernetes Helm charts generated from repository specifications, while entire simulation environments are composed declaratively via dynamic Helmfile manifests. The paper describes how a distributed testing workflow can be implemented in Argo Workflows to provision environments, aggregate and batch OpenSCENARIO test cases from configurable sources, execute scenarios in parallel across cluster nodes, and collect logs and resource metrics. In an evaluation on a multi-node K3s cluster running 200 scenarios, the best configuration speeds up end-to-end workflow time by more than a factor of eight compared to a sequential baseline. The results demonstrate significant gains in end-to-end execution time and quantify trade-offs between parallelism, orchestration overhead, and cluster stability. The framework is further demonstrated in a real-world ADS test application with connections to scenario sources and downstream evaluation modules. This demonstrates that the approach provides a strong foundation not only for scalable simulation testing, but also for generating traceable evidence that can support safety arguments.
Chinese Translation
基于虚拟场景的测试是验证自动驾驶系统(ADS)和智能交通系统(ITS)的关键手段。然而,执行涉及数千个场景的大规模测试套件仍然劳动密集且难以扩展。本文提出了一个端到端的、以DevOps驱动的框架,自动化构建、部署和在轻量级Kubernetes集群上分布式执行基于CARLA的ADS场景测试。ROS 2应用被打包为从代码库规范生成的标准化Kubernetes Helm图表,而整个仿真环境通过动态的Helmfile清单以声明方式组合。本文描述了如何在Argo Workflows中实现分布式测试工作流,以配置环境、从可配置源聚合和批处理OpenSCENARIO测试用例、在集群节点之间并行执行场景,并收集日志和资源指标。在对运行200个场景的多节点K3s集群的评估中,最佳配置的端到端工作流时间比顺序基线快了超过八倍。结果显示了端到端执行时间的显著提升,并量化了并行性、编排开销和集群稳定性之间的权衡。该框架在一个与场景源和下游评估模块连接的真实世界ADS测试应用中进一步得到了验证。这表明该方法不仅为可扩展的仿真测试提供了坚实的基础,还为生成可追溯的证据以支持安全论证提供了支持。
cs.RO / 11 / 2608.20946
Fast Coordinated Bimanual Motion Planning With Hard Constraints
快速协调的双手运动规划与硬约束
Abstract
Bimanual manipulation enables complex tasks but introduces added complexity from the high number of degrees of freedom involved. When handling rigid objects, the relative transformation between the two end effectors must remain fixed throughout the motion, manifesting as a nonlinear equality constraint that confines the feasible configuration space to a measure-zero manifold and challenges conventional motion planners. We propose a fast bimanual motion planning pipeline that enforces this hard transformation constraint continuously along the entire path, using a leader-follower parameterization: the leader's configuration is treated as a free variable, while the follower's is determined via inverse kinematics to satisfy the constraint. We extensively evaluate the method in simulation across diverse environments, constraints and bimanual platforms, achieving 19.4x faster planning than prior work while guaranteeing continuous constraint satisfaction. Real-world experiments on a bimanual Kinova Gen3 setup, involving tray transport and elongated-object manipulation, validate direct transfer of planned trajectories to physical hardware.
Chinese Translation
双手操作能够实现复杂任务,但由于涉及的自由度较多,增加了操作的复杂性。在处理刚性物体时,两端执行器之间的相对变换必须在整个运动过程中保持固定,这表现为一个非线性等式约束,限制了可行配置空间为测度零的流形,并对传统运动规划器提出了挑战。我们提出了一种快速的双手运动规划管道,该管道在整个路径上持续施加这一硬变换约束,采用领导-跟随参数化:领导者的配置被视为自由变量,而跟随者的配置通过逆向运动学来确定,以满足约束。我们在多种环境、约束和双手平台上对该方法进行了广泛的仿真评估,结果显示其规划速度比之前的工作快19.4倍,同时保证了约束的持续满足。在双手 Kinova Gen3 设置下进行的实际实验,包括托盘运输和细长物体操作,验证了规划轨迹向物理硬件的直接转移。
cs.RO / 12 / 2608.20948
Neural-Primitive: An Efficient End-to-end Local Planner with Primitive-based Imitation Learning for Autonomous Flight
神经原语:一种基于原语模仿学习的高效端到端局部规划器,用于自主飞行
Abstract
Autonomous flight in unknown cluttered environments is hindered by the computation-quality-memory trilemma of onboard trajectory generation. In this paper, we propose an efficient end-to-end local planner via imitation learning. A lightweight offline-primitive-based dataset collection framework is designed to produce safe and high-quality trajectory primitives in non-convex environments. A compact neural network directly maps sensory inputs to polynomial coefficients that inherently encode higher-order dynamical information. The learned policy generates smooth, empirically collision-free and dynamically feasible trajectories in real time without back-end solving. It achieves ultra-fast computation (below 1ms on a standard desktop and average 3.68ms during onboard flight), while maintaining low onboard memory requirements (less than 1.5MiB). Extensive simulation benchmarks demonstrate superiority in both planning latency and target-reaching progress quality. Zero-shot deployment in real-world experiments further validates the robust sim-to-real transfer capability of the proposed method.
Chinese Translation
在未知的杂乱环境中,自主飞行受到机载轨迹生成的计算-质量-内存三难问题的制约。本文提出了一种通过模仿学习实现的高效端到端局部规划器。设计了一种轻量级的离线原语基础数据集收集框架,以在非凸环境中生成安全且高质量的轨迹原语。一个紧凑的神经网络直接将传感器输入映射到多项式系数,这些系数本质上编码了更高阶的动态信息。学习到的策略能够实时生成平滑、经验上无碰撞且动态可行的轨迹,而无需后端求解。它实现了超快的计算速度(在标准桌面上低于1毫秒,在机载飞行中平均为3.68毫秒),同时保持低机载内存需求(少于1.5MiB)。大量的仿真基准测试证明了其在规划延迟和目标到达进度质量上的优越性。在真实世界实验中的零样本部署进一步验证了所提方法的稳健的仿真到现实转移能力。
cs.RO / 13 / 2608.20962
Hybrid Roller-Jamming Gripper for Object Acquisition and Retention Under Pose Uncertainty
用于姿态不确定性下物体获取与保持的混合滚轮-夹紧抓取器
Abstract
In household manipulation, pose uncertainty often results in off-centre or partial initial contact, making reliable object acquisition difficult. Roller-based grippers can actively draw objects inward but often provide limited post-capture stability, whereas granular-jamming grippers require sufficient contact before jamming to achieve strong retention. This paper presents a hybrid roller-jamming gripper that integrates active object intake and post-capture retention within a single gripper. The proposed gripper uses inward roller rotation to increase contact and draw the object toward the gripper centre, followed by vacuum-induced granular jamming to stiffen the rollers and stabilise the grasp. The paper also presents a simplified geometric analysis of the gripper and a bench-level characterisation of the prototype's force capability. The gripper prototype was mounted on a 7-DoF robotic arm and evaluated using eight test objects. Furthermore, controlled planar position and orientation offsets were applied, with each condition repeated three times. The main evaluation comprised 840 grasp trials, including 216 planar-offset trials and 624 orientation-offset trials. Overall, the gripper succeeded in 812/840 trials: 215/216 planar-offset trials and 597/624 orientation-offset trials. The ablation evaluation comprised 162 trials on three objects. The roller-only and jamming-only conditions achieved 54/81 and 24/81 successes, respectively, showing their different contributions. These results provide initial mechanism-level evidence that hybrid roller-jamming is a promising strategy for improving acquisition and retention after imperfect first contact.
Chinese Translation
在家庭操作中,姿态不确定性常常导致偏心或部分初始接触,使得可靠的物体获取变得困难。基于滚轮的抓取器可以主动将物体向内拉,但在捕获后通常提供有限的稳定性,而颗粒夹紧抓取器则需要在夹紧之前有足够的接触才能实现强保持。本文提出了一种混合滚轮-夹紧抓取器,将主动物体摄取和捕获后的保持集成于一个抓取器中。所提出的抓取器利用向内的滚轮旋转来增加接触并将物体拉向抓取器中心,随后通过真空诱导的颗粒夹紧来增强滚轮的刚性并稳定抓握。本文还对抓取器进行了简化的几何分析,并对原型的力能力进行了基准级特性评估。抓取器原型安装在一个7自由度的机器人手臂上,并使用八个测试物体进行了评估。此外,还施加了受控的平面位置和方向偏移,每种条件重复三次。主要评估包括840次抓握试验,其中包括216次平面偏移试验和624次方向偏移试验。总体而言,抓取器在840次试验中成功了812次:215/216次平面偏移试验和597/624次方向偏移试验。消融评估包括对三个物体进行的162次试验。仅使用滚轮和仅使用夹紧的条件分别取得了54/81和24/81的成功,显示了它们的不同贡献。这些结果提供了初步的机制级证据,表明混合滚轮-夹紧是一种有前景的策略,可以改善在不完美初始接触后的获取和保持。
cs.RO / 14 / 2608.21031
PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration
PhysCaP:基于物理知识的代码作为策略代理的探索
Abstract
We present PhysCaP, a Physics-Informed Code-as-Policy agent for active perception in robotic manipulation. While vision-language-action policies excel at imitating demonstrations, they rely on passive observation and fail to infer latent physical properties critical for manipulation. PhysCaP augments code-as-policy frameworks with a physics-informed exploration layer that enables explicit information-seeking through interaction. It introduces training-free physical property extraction modules that estimate object mass and stiffness from robot proprioception without additional sensors. To balance exploration costs and the efficiency of information obtained, PhysCaP employs a dual-agent design: a Planner that decides when to explore and when to stop, and a Prioritizer that filters implausible interactions and ranks the remainder using a heuristic priority score, enabling efficient, targeted exploration. We evaluate PhysCaP on real-world tabletop manipulation tasks (searching for hidden objects, detecting empty cans, and finding ripe avocados) and a simulated task in LIBERO. The results show that existing passive and naive interactive baselines either fail when physical properties are hidden or over-explore, whereas PhysCaP achieves comparable performance with fewer interactions and reduced execution time. Ablation studies further validate the effectiveness of the proposed physical property extraction modules. Project page: https://physcap.github.io
Chinese Translation
我们提出了PhysCaP,一种用于机器人操作的基于物理知识的代码作为策略代理,旨在实现主动感知。尽管视觉-语言-动作策略在模仿演示方面表现出色,但它们依赖于被动观察,无法推断出对操作至关重要的潜在物理属性。PhysCaP通过引入一个基于物理知识的探索层,增强了代码作为策略框架,使其能够通过交互进行明确的信息寻求。它引入了无训练的物理属性提取模块,能够从机器人本体感知中估计物体的质量和刚度,而无需额外传感器。为了平衡探索成本与获取信息的效率,PhysCaP采用了双代理设计:一个规划者(Planner)决定何时探索和何时停止,另一个优先级过滤器(Prioritizer)过滤不合理的交互,并使用启发式优先级评分对剩余交互进行排序,从而实现高效、针对性的探索。我们在真实世界的桌面操作任务(寻找隐藏物体、检测空罐和寻找成熟鳄梨)以及LIBERO中的模拟任务上评估了PhysCaP。结果表明,现有的被动和简单交互基线在物理属性被隐藏时要么失败,要么过度探索,而PhysCaP在较少的交互和减少的执行时间下实现了可比的性能。消融研究进一步验证了所提物理属性提取模块的有效性。项目页面:https://physcap.github.io
cs.RO / 15 / 2608.21032
Roadside-Cooperative Autonomous Driving: From Data Platform to Vision-Language End-to-End Reasoning
路边协作自主驾驶:从数据平台到视觉-语言端到端推理
Abstract
Vehicle-to-Everything (V2X) cooperation enables beyond-line-of-sight perception, mitigating occlusions in single-vehicle sensing. However, existing V2X benchmarks provide limited support for closed-loop evaluation and language-grounded supervision, hindering the development of vision-language models (VLMs) for end-to-end cooperative driving. To address these limitations, we introduce V2XBench, a simulation platform featuring synchronized ego--roadside sensing and closed-loop evaluation, together with Chat-V2XBench, a progressively structured VQA dataset for cooperative reasoning. Building upon this benchmark infrastructure, we propose AURORA, an end-to-end cooperative driving framework. Equipped with a dual-view perception architecture, AURORA mitigates spatial and semantic discrepancies across ego and roadside viewpoints through a query-level Cross-View Query Alignment and Fusion (CQAF) module. Leveraging the resulting unified tokens, a LoRA-adapted VLM bridges semantic reasoning and generative trajectory planning. Extensive closed-loop evaluations on V2XBench demonstrate that AURORA achieves state-of-the-art performance in heavily occluded scenarios, with a Route Completion rate of 98.21% and a Driving Score of 76.02, while requiring low roadside communication bandwidth. Ultimately, this work pioneers an extensible V2X--VLM paradigm, paving the way for next-generation cooperative autonomous driving.
Chinese Translation
车对一切(V2X)合作使得超视距感知成为可能,减轻了单车感知中的遮挡问题。然而,现有的V2X基准对闭环评估和基于语言的监督支持有限,阻碍了视觉-语言模型(VLM)在端到端协作驾驶中的发展。为了解决这些局限性,我们引入了V2XBench,这是一个具有同步自我-路边感知和闭环评估的仿真平台,同时推出了Chat-V2XBench,这是一个为协作推理而逐步构建的视觉问答(VQA)数据集。在此基准基础上,我们提出了AURORA,一个端到端的协作驾驶框架。AURORA配备了双视角感知架构,通过查询级别的跨视图查询对齐与融合(CQAF)模块,减轻了自我视角与路边视角之间的空间和语义差异。利用生成的统一标记,一个经过LoRA调整的VLM将语义推理与生成轨迹规划连接起来。在V2XBench上进行的大量闭环评估表明,AURORA在严重遮挡场景中实现了最先进的性能,路线完成率达到98.21%,驾驶得分为76.02,同时所需的路边通信带宽较低。最终,这项工作开创了一个可扩展的V2X-VLM范式,为下一代协作自主驾驶铺平了道路。
cs.RO / 16 / 2608.21035
TaPeR: Probabilistic Recovery of Sparse Task Precedence Graphs from a Handful of Demonstrations
TaPeR:从少量示例中概率性恢复稀疏任务优先图
Abstract
Long-horizon manipulation tasks are often only partially ordered. For example, when assembling an electronic device, the battery and circuit board may be installed in either order, but both must be in place before the enclosure is closed. Recovering such dependencies enables robots to flexibly reorder subtasks while preserving task validity. Existing approaches typically infer task structure from human demonstrations using both temporal and symbolic supervision. However, symbolic predicates require explicit grounding, which is difficult to obtain in realistic settings. In this work, we present an approach for extracting task dependency structures from demonstrations using only simple kinematic graphs and distributions over relative object poses. From these representations, our method estimates pairwise task-step-dependency probabilities and uses them to initialize the edge weights of a precedence graph. We then introduce a filtering pipeline that converts this graph of probability estimates into the final task dependency graph. We evaluate our approach on an existing benchmark and on a new dataset comprising longer tasks with more complex dependencies. We find that our method recovers more accurate task structures from fewer demonstrations than the baselines. Finally, we demonstrate that the inferred graphs can be used to generate multiple valid robotic execution orders for the same task.
Chinese Translation
长时间跨度的操作任务通常仅部分有序。例如,在组装电子设备时,电池和电路板可以任意顺序安装,但在关闭外壳之前,两者必须都到位。恢复这种依赖关系使得机器人能够灵活地重新排序子任务,同时保持任务的有效性。现有的方法通常通过人类示例推断任务结构,使用时间和符号监督。然而,符号谓词需要明确的基础,这在现实环境中很难获得。在本研究中,我们提出了一种仅使用简单的运动学图和相对物体姿态分布从示例中提取任务依赖结构的方法。基于这些表示,我们的方法估计成对任务步骤依赖的概率,并利用这些概率初始化优先图的边权重。然后,我们引入一个过滤管道,将这一概率估计图转换为最终的任务依赖图。我们在现有基准和一个包含更复杂依赖关系的长任务新数据集上评估了我们的方法。我们发现,与基线相比,我们的方法能够从更少的示例中恢复出更准确的任务结构。最后,我们展示了推断出的图可以用于为同一任务生成多个有效的机器人执行顺序。
cs.RO / 17 / 2608.21056
FF-MPCC: High-speed Agile Formation Flight with Model Predictive Contouring Control
FF-MPCC:基于模型预测轮廓控制的高速灵活编队飞行
Abstract
Flying in a prescribed formation in an agile manner remains a challenging problem in the field of UAVs, particularly when following highly-demanding trajectories that require flight at platform limits. We address this problem by proposing a novel decentralized approach to formation flight along a given path that integrates formation maintenance into the MPCC framework, allowing UAVs to adapt their progression along complex paths while respecting individual dynamic constraints and maintaining the desired formation. To this end, we introduce a novel reparametrization and synchronization method for dynamic formation geometries together with a decentralized approach to determine the desired positions for the individual UAVs. The proposed approach allows the formation to coordinate high-speed path following without compromising formation integrity. The proposed approach is validated through extensive simulation and real-world experiments involving scenarios with varying complexity of paths and changes of required formation shape on the fly. In comparison to time-parameterized trajectory tracking, we demonstrate improved formation maintenance by 65% in high-speed flight with velocities up to 21 m/s, while achieving comparable times required to reach the goal.
Chinese Translation
在无人机领域,以灵活的方式在预定编队中飞行仍然是一个具有挑战性的问题,尤其是在遵循需要在平台极限飞行的高要求轨迹时。我们通过提出一种新颖的去中心化编队飞行方法来解决这一问题,该方法将编队维护集成到MPCC(模型预测轮廓控制)框架中,使无人机能够在遵循复杂路径的同时适应各自的动态约束并保持所需的编队。为此,我们引入了一种新的动态编队几何形状的重新参数化和同步方法,以及一种去中心化的方法来确定各个无人机的期望位置。所提出的方法允许编队在不妨碍编队完整性的情况下协调高速路径跟踪。通过广泛的仿真和真实世界实验,我们验证了该方法,涉及不同复杂性的路径场景和实时变化的编队形状。与时间参数化轨迹跟踪相比,我们在高速飞行(速度高达21 m/s)中展示了编队维护提高了65%,同时达到目标所需的时间相当。
cs.RO / 18 / 2608.21083
Teaching is a Process: The TOSS Framework for Modeling Human Teaching Decisions in Human-Interactive Robot Learning
教学是一种过程:TOSS框架用于建模人类在与人互动的机器人学习中的教学决策
Abstract
Successful Human-Robot Teaching assumes alignment between robot processing needs and human teaching intent. To better understand this alignment, this work seeks to uncover the underlying logic that humans intuitively apply when teaching. Through an exploratory, bottom-up study with N=34, participants observing two distinct robot Reinforcement Learning (RL) scenarios, we analyze 204 intuitive teaching responses across early, middle, and late learning phases. Results reveal that teaching decisions consist of a nuanced, interconnected network of Triggers (situational catalysts), Objectives (subjective teaching targets), Signals (communicative acts), and Strategies (high-level governance) in which teachers spontaneously adopt diverse roles, acting as coaches, engineers, or designers and prioritize different objectives. Based on these results, we introduce the TOSS Framework, which conceptualizes Human-Robot teaching as a procedural loop between robot behavior and human teaching actions, in which human teaching decisions are modeled as Trigger-Signal responses modulated by teaching Objectives and Strategies. It provides future research with an openly accessible dataset and a theoretical foundation for a) understanding teaching decisions and b) simulating realistic oracles as well as c) designing human-centered teaching settings and novel robot learning algorithms that go beyond the constraints of current robot learning settings.
Chinese Translation
成功的人机教学假设机器人处理需求与人类教学意图之间存在一致性。为了更好地理解这种一致性,本研究旨在揭示人类在教学时直观应用的潜在逻辑。通过对34名参与者进行探索性自下而上的研究,观察他们在两个不同的机器人强化学习(Reinforcement Learning, RL)场景中的表现,我们分析了在学习的早期、中期和晚期阶段的204个直观教学反应。结果显示,教学决策由一系列细致且相互关联的网络组成,包括触发因素(Triggers,情境催化剂)、目标(Objectives,主观教学目标)、信号(Signals,沟通行为)和策略(Strategies,高层治理),教师在其中自发地扮演不同角色,如教练、工程师或设计师,并优先考虑不同的目标。基于这些结果,我们提出了TOSS框架,该框架将人机教学概念化为机器人行为与人类教学行为之间的程序循环,其中人类教学决策被建模为受教学目标和策略调节的触发-信号反应。该框架为未来的研究提供了一个开放可获取的数据集和理论基础,以便于a) 理解教学决策,b) 模拟现实的预言者,以及c) 设计以人为本的教学环境和超越当前机器人学习环境限制的新型机器人学习算法。
cs.RO / 19 / 2608.21175
SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control
SRL-MPC:形状感知强化学习模型预测控制
Abstract
Safe and efficient shape-aware navigation in heterogeneous crowds and robot fleets remains challenging. Traditional approaches often assume homogeneous robots, sparse workspaces, simplified geometry, offline computation, or handcrafted parameters to make the problem tractable, which limits their deployment in dense crowd scenarios. Toward this end, we propose Shape-Aware Reinforcement Learned Model Predictive Control (SRL-MPC), a method for safe, efficient, and adaptive navigation in crowds with heterogeneous shapes without geometry simplification. To encode shape-aware safety, we formulate high-order control barrier function (HOCBF) constraints from geometric separation features (GSFs) based on support function transformation. A reinforcement learning (RL) framework then learns a neural policy that reads GSFs and outputs real-time MPC parameter updates, enabling the MPC solver to adapt to neighboring crowd geometries. The key advantage of SRL-MPC is that it preserves the safety structure and generalizability of MPC while integrating the adaptability and intelligence of RL. Experiments in randomized crowd scenarios with arbitrary shaped robot fleets demonstrate the effectiveness, scalability, and robustness of SRL-MPC. The results show that SRL-MPC substantially outperforms representative baselines in safety and adaptability. Project website: https://hanruihua.github.io/srl_mpc_project/
Chinese Translation
在异构人群和机器人队伍中实现安全高效的形状感知导航仍然面临挑战。传统方法通常假设机器人是同质的,工作空间稀疏,几何形状简化,离线计算,或手工调整参数以使问题可处理,这限制了它们在密集人群场景中的应用。为此,我们提出了形状感知强化学习模型预测控制(SRL-MPC),该方法用于在不简化几何形状的情况下,实现异构形状人群中的安全、高效和自适应导航。为了编码形状感知的安全性,我们基于支持函数变换,从几何分离特征(GSFs)中制定高阶控制障碍函数(HOCBF)约束。然后,强化学习(RL)框架学习一个神经策略,该策略读取GSFs并输出实时的模型预测控制参数更新,使得模型预测控制求解器能够适应邻近人群的几何形状。SRL-MPC的关键优势在于,它在整合RL的适应性和智能的同时,保持了模型预测控制的安全结构和泛化能力。在具有任意形状机器人队伍的随机人群场景中的实验表明,SRL-MPC的有效性、可扩展性和鲁棒性。结果显示,SRL-MPC在安全性和适应性方面显著优于代表性基线。项目网站:https://hanruihua.github.io/srl_mpc_project/
cs.RO / 20 / 2608.21204
Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning
超越模仿:通过离策略 Q 规划自我改进的机器人策略
Abstract
Behaviour Cloning (BC) has driven remarkable progress in robot manipulation, yet it is fundamentally limited by its inability to self-improve: a policy that fails cannot learn from that failure without additional human demonstrations. Reinforcement Learning fine-tuning offers a path to self-improvement but has proven difficult to scale to the multi-billion-parameter models underpinning modern robot policies. We propose Q-Planning, which equips a large visuomotor BC policy with a small off-policy Q-function. Because a Q-function estimates value rather than imitates actions, it can be trained on the same successful demonstrations as the BC policy and later absorb both successful and failed deployment rollouts, an asymmetry BC does not have. We exploit this asymmetry to enable value-guided action selection at inference (a single-step Q-weighted average over BC draws) and online self-improvement that fine-tunes only the Q-function, leaving the BC weights untouched. On LIBERO and bimanual RoboTwin, ten iterations of self-improvement lift every benchmark score we tested (LIBERO-10 93% to 99%, RoboTwin 83.8% to 91.4%) and shorten successful episodes on the near-ceiling suites (LIBERO-Object, LIBERO-Goal). On two contact-rich bimanual real-robot tasks, the same loop (BC frozen, no human intervention) improves purely from its own deployment rollouts: stack-cups 40% to 90% and insert-wallet 25% to 80% in five iterations, whereas SFT on successful rollouts alone stalls at 55% and 30%. Under an identical online budget Q-Planning is the only method, among Best-of-N, filtered SFT, IBRL, DSRL, and DAWR, that improves stably from failures without training an auxiliary actor.
Chinese Translation
行为克隆(BC)在机器人操作中推动了显著的进展,但其根本局限在于无法自我改进:一个失败的策略无法从失败中学习,而不需要额外的人类示范。强化学习微调提供了一条自我改进的路径,但在现代机器人策略所依赖的数十亿参数模型中,已被证明难以扩展。我们提出了 Q 规划,它为大型视觉运动 BC 策略配备了一个小型离策略 Q 函数。由于 Q 函数估计价值而不是模仿动作,因此可以在与 BC 策略相同的成功示范上进行训练,并随后吸收成功和失败的部署回滚,这是 BC 所不具备的非对称性。我们利用这种非对称性,在推理时实现价值引导的动作选择(对 BC 抽样的单步 Q 加权平均)和在线自我改进,仅微调 Q 函数,而不触动 BC 权重。在 LIBERO 和双手 RoboTwin 上,十次自我改进提升了我们测试的每个基准分数(LIBERO-10 从 93% 提升至 99%,RoboTwin 从 83.8% 提升至 91.4%),并缩短了接近上限的任务成功回合(LIBERO-Object,LIBERO-Goal)。在两个接触丰富的双手真实机器人任务中,同样的循环(BC 冻结,无需人类干预)仅通过自身的部署回滚进行改进:堆杯从 40% 提升至 90%,插钱包从 25% 提升至 80% 在五次迭代中,而仅在成功回滚上进行 SFT 则停滞在 55% 和 30%。在相同的在线预算下,Q 规划是 Best-of-N、过滤 SFT、IBRL、DSRL 和 DAWR 中唯一一个能够稳定从失败中改进而无需训练辅助演员的方法。
cs.RO / 21 / 2608.21276
The Coastline as a Structural Constraint: Harnessing Scene Geometry for Autonomous Surface Vessel Localization
海岸线作为结构约束:利用场景几何进行自主水面船舶定位
Abstract
Coastal environments contain rich, largely unexploited geometric structure capable of providing globally referenced localization cues. In this work, we present two complementary localization frameworks that exploit shoreline and water-surface geometry for GPS-denied autonomous surface vessel localization. The first framework leverages LiDAR observations of the water surface to estimate roll, pitch, and heave (vertical motion), while recovering global position and heading through direct registration of shoreline observations against a satellite-derived coastline map. The second framework relies solely on passive imagery to detect the shoreline and horizon through semantic segmentation. Using the proposed coastal scene geometry, shoreline distance is inferred from monocular imagery. Shoreline observations are accumulated into short-duration local submaps, registered against the same satellite-derived coastline map, and fused within a hierarchical factor graph. Evaluated across three real-world coastal datasets, the LiDAR pipeline consistently improves trajectory accuracy over standard baselines, while the monocular architecture maintains bounded long-term drift. In addition, we establish that modern zero-shot foundation models can reliably extract shoreline observations across diverse coastal environments. Together, these results demonstrate that coastal geometry provides a powerful and dependable source of globally referenced information for GPS-denied maritime localization.
Chinese Translation
沿海环境包含丰富且大部分未被利用的几何结构,能够提供全球参考的定位线索。在本研究中,我们提出了两种互补的定位框架,利用海岸线和水面几何进行GPS不可用的自主水面船舶定位。第一个框架利用LiDAR对水面观测的测量来估计横滚、俯仰和升沉(垂直运动),同时通过将海岸线观测与卫星推导的海岸线地图进行直接配准来恢复全球位置和航向。第二个框架仅依赖被动图像,通过语义分割检测海岸线和地平线。利用所提出的沿海场景几何,从单目图像推断海岸线距离。海岸线观测被积累到短时局部子地图中,与同一卫星推导的海岸线地图进行配准,并在分层因子图中融合。通过对三个真实世界沿海数据集的评估,LiDAR管道在标准基线之上持续提高了轨迹准确性,而单目架构则保持了有限的长期漂移。此外,我们确定现代零样本基础模型能够可靠地提取多样沿海环境中的海岸线观测。综合来看,这些结果表明,沿海几何为GPS不可用的海事定位提供了强大而可靠的全球参考信息源。
cs.RO / 22 / 2608.21290
VT-MUSE: Multimodal Unified Sequential Visuotactile Representation Learning for Manipulation
VT-MUSE:用于操作的多模态统一序列视觉触觉表示学习
Abstract
We propose VT-MUSE, a Multimodal Unified SEquential representation learning framework for visuotactilemanipulation. Existing approaches often encode visual and tactile observations independently before fusion, limiting their ability to capture fine-grained cross-modal dependencies. Moreover, most methods focus on observations at the current time step and overlook the temporal evolution of contact. VT-MUSE addresses both limitations through a two-stage representation learning framework. In Stage I, modality specific encoders are jointly adapted via cross-modal temporal alignment and masked-view consistency. In Stage II, a conditional variational latent model processes masked visual sequences together with full tactile histories. Auxiliary decoders reconstruct the masked recent visual observations and predict tactile depth changes, encouraging the latent representation to retain both global visual context and local contact dynamics. The learned representation is subsequently integrated into a lightweight Transformer policy through gated cross-attention. On the simulation benchmark, VT-MUSE outperforms the strongest baseline evaluated on all tasks by 11 percentage points and also achieves substantial improvements in real-world experiments.
Chinese Translation
我们提出了VT-MUSE,一个用于视觉触觉操作的多模态统一序列表示学习框架。现有的方法通常在融合之前独立编码视觉和触觉观测,限制了它们捕捉细粒度跨模态依赖的能力。此外,大多数方法专注于当前时间步的观测,而忽视了接触的时间演变。VT-MUSE通过一个两阶段的表示学习框架解决了这两个限制。在第一阶段,特定模态的编码器通过跨模态时间对齐和掩蔽视图一致性进行联合适应。在第二阶段,一个条件变分潜在模型处理掩蔽的视觉序列以及完整的触觉历史。辅助解码器重建掩蔽的近期视觉观测并预测触觉深度变化,鼓励潜在表示保留全局视觉上下文和局部接触动态。随后,学习到的表示通过门控交叉注意机制集成到一个轻量级的Transformer策略中。在仿真基准测试中,VT-MUSE在所有任务上比评估的最强基线高出11个百分点,并且在实际实验中也取得了显著的改进。
cs.RO / 23 / 2608.21330
NeSAM: Neuro-Symbolic Kinodynamics with Soil Adaptation for Off-Road Mobility
NeSAM:具有土壤适应性的神经符号动力学模型用于越野移动
Abstract
Accurate prediction of off-road vehicle motion over deformable terrain remains challenging because sinkage, slip, and traction vary with local soil conditions. Existing learning-based kinodynamic models directly approximate vehicle-terrain interactions from data but do not explicitly represent soil mechanics and offer limited physical interpretability. To address these limitations, we present NeSAM, a neuro-symbolic framework that combines differentiable Bekker-Wong terramechanics with learned terrain representations and a Transformer-based residual dynamics model for long-horizon, six degree-of-freedom kinodynamic prediction. The terramechanics component models soil-dependent interaction forces, while the residual model corrects discrepancies between the analytical prediction and the observed vehicle dynamics. NeSAM further estimates physically meaningful soil parameters from terrain observations and updates them online using an extended Kalman filter. We evaluate NeSAM in Verti-Bench, a simulator built on the Chrono multiphysics engine, and validate its performance on a physical Verti-4-Wheeler platform. NeSAM improves prediction accuracy by up to 30% in simulation and 29% on real-world data relative to the strongest compared baselines. When integrated with a close-loop navigation controller, NeSAM further improves traversal success rate through online soil adaptation while reduces Hausdorff distance to the reference trajectory by 69.4%, indicating improved trajectory tracking accuracy.
Chinese Translation
在可变形地形上准确预测越野车辆的运动仍然具有挑战性,因为沉陷、滑移和牵引力会随着局部土壤条件而变化。现有的基于学习的动力学模型直接从数据中近似车辆与地形的交互,但并未明确表示土壤力学,并且物理可解释性有限。为了解决这些局限性,我们提出了NeSAM,一个神经符号框架,结合了可微分的Bekker-Wong土力学模型、学习到的地形表示和基于Transformer的残差动力学模型,用于长时间跨度的六自由度动力学预测。土力学组件建模了依赖于土壤的交互力,而残差模型则修正了分析预测与观察到的车辆动态之间的差异。NeSAM进一步从地形观测中估计物理意义明确的土壤参数,并使用扩展卡尔曼滤波器在线更新这些参数。我们在基于Chrono多物理引擎构建的Verti-Bench模拟器中评估NeSAM,并在物理Verti-4-Wheeler平台上验证其性能。与最强的对比基线相比,NeSAM在模拟中提高了预测精度达30%,在真实数据中提高了29%。当与闭环导航控制器集成时,NeSAM通过在线土壤适应进一步提高了穿越成功率,同时将Hausdorff距离减少了69.4%,表明轨迹跟踪精度得到了改善。
cs.RO / 24 / 2608.21355
ViTacPhys: Physical Property-Aware Grasping from Human Visual-Tactile Demonstrations
ViTacPhys:基于人类视觉-触觉示范的物理属性感知抓取
Abstract
Recent vision-based action models have demonstrated strong capabilities in complex manipulation, but they rarely leverage explicit object physical properties to adapt their policies. We introduce ViTacPhys, a visual-tactile framework and data acquisition system that estimates object mass and friction-coefficient classes, together with continuous stiffness, from human manipulation demonstrations. Trained on data from 60 rigid and deformable objects, ViTacPhys combines temporal visual-tactile modeling, cross-attention multimodal fusion, and a semantic prior derived from a vision-language model. On seen objects, it achieves 97.2% mass classification accuracy, 98.8% friction-coefficient classification accuracy, and a stiffness mean absolute percentage error (MAPE) of 5.51%. On held-out objects from known categories, it achieves 87.5% mass accuracy, 97.5% friction-coefficient accuracy, and a stiffness MAPE of 9.08%. We transfer ViTacPhys from the human domain to the robot domain using limited robot teleoperation data, robot-style video augmentation, and human demonstrations with matched actions, and deploy it as an online module for adaptive grasping. The resulting physical-property-conditioned policy achieves total grasping success rates of 95.0% on in-distribution objects and 83.4% on out-of-distribution objects. For out-of-distribution objects successfully grasped by both methods, its force profiles are more consistent with human teleoperation than those produced by ACT. These results demonstrate the feasibility of explicitly estimating and conditioning on object physical properties for real-world adaptive grasping.
Chinese Translation
近期基于视觉的动作模型在复杂操作中展现出了强大的能力,但它们很少利用明确的物体物理属性来调整其策略。我们提出了ViTacPhys,一个视觉-触觉框架和数据采集系统,该系统能够从人类操作示范中估计物体的质量和摩擦系数类别,以及连续的刚度。ViTacPhys在60个刚性和可变形物体的数据上进行训练,结合了时间视觉-触觉建模、交叉注意力多模态融合,以及从视觉-语言模型中得出的语义先验。在已见物体上,它实现了97.2%的质量分类准确率、98.8%的摩擦系数分类准确率,以及5.51%的刚度平均绝对百分比误差(MAPE)。在已知类别的未见物体上,它实现了87.5%的质量准确率、97.5%的摩擦系数准确率,以及9.08%的刚度MAPE。我们利用有限的机器人遥操作数据、机器人风格的视频增强以及与匹配动作的人类示范,将ViTacPhys从人类领域转移到机器人领域,并将其部署为自适应抓取的在线模块。最终的物理属性条件策略在分布内物体上的总抓取成功率达到95.0%,在分布外物体上的成功率为83.4%。对于两种方法均成功抓取的分布外物体,其力特征与人类遥操作的匹配程度高于ACT生成的特征。这些结果展示了在现实世界自适应抓取中明确估计和条件化物体物理属性的可行性。
cs.RO / 25 / 2608.21358
Mining beyond Earth with Space Robots: Exploration, Sampling, and Extraction
利用太空机器人进行地外资源开采:探索、取样与提取
Abstract
Space resource acquisition and utilization, commonly referred to as Space Mining, represent critical pathways for enabling sustained human exploration and unlocking commercial opportunities in space. These resources mainly include helium-3, water, mineral resources on the Moon and Mars, and abundant mineral deposits on asteroids. Due to the harsh conditions of space, communication delays, and high launch costs, the development of autonomous robotic systems is critical to achieving efficient, cost-effective space mining. This paper provides a comprehensive overview of space mining robotics and associated technologies. First, we review the background of space mining, including international policies, commercial entities, and recent advancements. We define a systematic six-stage architecture for space mining: Exploration is initiated by (1) remote sensing for target identification and (2) precise in situ robotic detection; Sampling progresses from (3) single-robot small-scale sampling to (4) multi-robot large-scale excavation; and Extraction integrates (5) autonomous resource extraction and (6) final integration into in situ construction or terrestrial transport. Additionally, we review and curate existing resources for space mining research, including real-world mission data, terrestrial analog datasets, and high-fidelity simulation environments. Finally, we identify critical open challenges in autonomous space mining and delineate a strategic research roadmap to bridge current technological gaps, fostering the transition toward a sustainable off-world economy. To track ongoing developments in space mining, we maintain an updated project page: https://github.com/OpenSpace-Lab/Space-Mining-with-Robotics-List.
Chinese Translation
太空资源获取与利用,通常称为太空开采,代表了实现持续人类探索和开启太空商业机会的关键途径。这些资源主要包括氦-3、水、月球和火星上的矿产资源,以及小行星上的丰富矿藏。由于太空环境的恶劣条件、通信延迟和高昂的发射成本,开发自主机器人系统对于实现高效、经济的太空开采至关重要。本文全面概述了太空开采机器人及相关技术。首先,我们回顾了太空开采的背景,包括国际政策、商业实体和近期进展。我们定义了一个系统的六阶段架构用于太空开采:探索阶段包括(1)通过遥感进行目标识别和(2)精确的原位机器人检测;取样阶段从(3)单机器人小规模取样进展到(4)多机器人大规模挖掘;提取阶段整合(5)自主资源提取和(6)最终整合到原位建设或地面运输中。此外,我们回顾并整理了现有的太空开采研究资源,包括现实任务数据、地面类比数据集和高保真模拟环境。最后,我们识别了自主太空开采中的关键开放挑战,并勾勒出一条战略研究路线图,以弥补当前技术差距,促进向可持续外星经济的转型。为了跟踪太空开采的最新进展,我们维护了一个更新的项目页面:https://github.com/OpenSpace-Lab/Space-Mining-with-Robotics-List。
cs.CV / 1 / 2608.20430
RISE: Adaptive Imagination for World Action Models
RISE:用于世界行动模型的自适应想象
Abstract
World Action Models (WAMs) improve planning by incorporating future world evolution into action generation, yet existing methods allocate a fixed imagination budget to every scene. We propose RISE (\textbf{R}efining \textbf{I}magination through \textbf{SE}lective Rollout), a system-level adaptive imagination framework that makes sequential \textsc{Roll}/\textsc{Stop} decisions according to the expected planning benefit of continued rollout. At each step, a Latent Evaluator estimates the risk revealed by the current prefix and how much planning could improve if imagination continues, while a Rollout Gate weighs this expected benefit against additional computation cost. Since factual driving logs expose only one realized future, we further construct \textbf{CounterDrive}, a counterfactual dataset with diverse outcomes and risk levels, to enrich future dynamics and provide localized risk supervision. Each retained sample undergoes expert verification and annotation of trajectory validity, incident onset, and causal category, providing a reusable resource for safety-critical world-modeling research. Experiments on NAVSIM and nuScenes show that RISE achieves the best overall planning performance while reducing unnecessary rollout, with additional transfer results supporting its plug-in generality across WAM architectures.
Chinese Translation
世界行动模型(World Action Models, WAMs)通过将未来世界演变纳入行动生成来改善规划,但现有方法对每个场景分配固定的想象预算。我们提出了RISE( extbf{R}efining extbf{I}magination through extbf{SE}lective Rollout),一个系统级自适应想象框架,根据持续展开的预期规划收益做出顺序的 extsc{Roll}/ extsc{Stop}决策。在每一步中,潜在评估器(Latent Evaluator)评估当前前缀揭示的风险以及如果继续想象规划可能改善的程度,而展开门(Rollout Gate)则将这一预期收益与额外的计算成本进行权衡。由于实际的驾驶日志仅揭示了一个实现的未来,我们进一步构建了 extbf{CounterDrive},一个具有多样化结果和风险水平的反事实数据集,以丰富未来动态并提供局部风险监督。每个保留的样本都经过专家验证和轨迹有效性、事件发生和因果类别的注释,为安全关键的世界建模研究提供了可重用的资源。在NAVSIM和nuScenes上的实验表明,RISE在减少不必要展开的同时实现了最佳的整体规划性能,额外的迁移结果支持其在WAM架构中的插件通用性。
cs.CV / 2 / 2608.20473
Aggregating Visual Information with Optimal Transport for VideoLM Token Compression
通过最优传输聚合视觉信息以实现视频语言模型的令牌压缩
Abstract
Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is therefore essential for reducing the visual-token burden on language-model decoding. The central challenge is to preserve visual information dispersed across frames under such compression. To this end, we introduce Aggregating Visual Information with Optimal Transport (AVIOT), which casts video token compression as transporting a dense empirical measure of frame observations onto a compact target measure. The resulting source-to-target coupling induces a distribution over source observations for each target support, directly specifying how the compressed video representation is constructed. We further adapt this construction along task and spatial axes. Question conditioning modulates the transport cost between source frames and target supports, while influencing how many supports are allocated to each temporal segment, thereby directing representation capacity toward question-relevant content. At multiple spatial granularities, AVIOT computes region-specific temporal transport plans and adaptively fuses the representations they yield, allowing different regions within the same compact representation to draw from different moments. Evaluations across varying compression ratios show that AVIOT matches or outperforms the uncompressed baseline on multiple video-understanding benchmarks while retaining strong performance at higher compression ratios.
Chinese Translation
视频语言模型将视频处理为密集的视觉令牌序列,这些序列具有显著的表示冗余。因此,压缩这些序列对于减少语言模型解码时的视觉令牌负担至关重要。核心挑战在于在压缩过程中保留分散在帧之间的视觉信息。为此,我们提出了通过最优传输聚合视觉信息(AVIOT),将视频令牌压缩视为将帧观测的密集经验测度传输到紧凑目标测度。由此产生的源到目标的耦合为每个目标支持指定了源观测的分布,直接定义了压缩视频表示的构建方式。我们进一步沿任务和空间轴适应这一构建。问题条件调节源帧与目标支持之间的传输成本,同时影响分配给每个时间段的支持数量,从而将表示能力引导至与问题相关的内容。在多个空间粒度上,AVIOT 计算区域特定的时间传输计划,并自适应融合它们所产生的表示,使得同一紧凑表示中的不同区域能够从不同的时刻获取信息。在不同压缩比下的评估表明,AVIOT 在多个视频理解基准上与未压缩基线相匹配或超越,同时在更高压缩比下保持强劲的性能。
cs.CV / 3 / 2608.20492
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
注释作为回滚:高效且可扩展的视频多模态大语言模型强化学习
Abstract
Multimodal large language models (MLLMs) have become a prevailing paradigm for unified video perception. However, post-training on large multi-task datasets remains challenging, as existing reinforcement learning methods sample on-policy groups with few high-quality rollouts even with costly chain-of-thought (CoT) generation. In this paper, we study the sample efficiency and scalability of RL post-training for video MLLMs and introduce OraRL. We identify an overlooked role for annotations: Beyond scoring rollouts, each can enter its on-policy group as an oracle rollout, a direct positive optimization target. Direct oracle integration, however, is nontrivial: a high-reward oracle raises the group baseline and inverts otherwise positive policy advantages, a failure we term advantage inversion. At the core of OraRL is a decoupled advantage estimator: policy rollouts determine an oracle-free baseline, while the oracle-policy gap modulates both a directional gain and a separate detached oracle advantage. Sign-balanced pruning improves efficiency: by retaining only the oracle and the strongest rollouts of each sign, OraRL requires just 2.2x the step time of SFT, less than half the 4.9x required by GRPO with CoT. OraRL scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts. Without chain-of-thought, Video-ORA-9B decodes in 130 ms instead of 4,780 ms. Compared with the respective prior best models, it raises temporal mIoU from 62.5 to 66.0, tracking AO from 73.0 to 78.2, segmentation from 64.3 to 70.4, and the three-benchmark spatial-intelligence macro average from 51.0 to 56.1; on VSI-Bench, it scores 73.1 against 55.0 for GPT-5 and 55.1 for Gemini-3-Pro.
Chinese Translation
多模态大语言模型(MLLMs)已成为统一视频感知的主流范式。然而,在大型多任务数据集上进行后训练仍然具有挑战性,因为现有的强化学习方法在政策组中采样时,即使在成本高昂的思维链(CoT)生成下,也只有少量高质量的回滚。在本文中,我们研究了视频 MLLMs 的强化学习后训练的样本效率和可扩展性,并引入了 OraRL。我们识别出注释的一个被忽视的角色:除了对回滚进行评分外,每个注释都可以作为一个 oracle 回滚进入其政策组,成为直接的正优化目标。然而,直接的 oracle 集成并非易事:高奖励的 oracle 提高了组基线,并反转了原本正向的政策优势,这种失败我们称之为优势反转。OraRL 的核心是一个解耦的优势估计器:政策回滚确定一个无 oracle 的基线,而 oracle-政策差距调节了方向性增益和一个独立的 oracle 优势。符号平衡修剪提高了效率:通过仅保留 oracle 和每个符号的最强回滚,OraRL 只需要 SFT 的 2.2 倍步长时间,远低于 GRPO 在 CoT 下所需的 4.9 倍。OraRL 随着模型规模和数据的增加而扩展,从 0.8B 超越到 9B,并将 GRPO 扩展到 100k 提示。在没有思维链的情况下,Video-ORA-9B 的解码时间为 130 毫秒,而不是 4,780 毫秒。与各自的最佳先前模型相比,它将时间 mIoU 从 62.5 提高到 66.0,跟踪 AO 从 73.0 提高到 78.2,分割从 64.3 提高到 70.4,三个基准空间智能宏平均从 51.0 提高到 56.1;在 VSI-Bench 上,它的得分为 73.1,而 GPT-5 和 Gemini-3-Pro 分别为 55.0 和 55.1。
cs.CV / 4 / 2608.20515
DiffVC-ONE: Diffusion-based Generative Video Compression with One-Step Video Diffusion Transformer
DiffVC-ONE:基于扩散的生成视频压缩与一步视频扩散变换器
Abstract
Generative video compression can recover rich visual details at low bitrates, but simultaneously achieving high temporal consistency and low inference cost remains challenging. To address this issue, we propose DiffVC-ONE, a diffusion-based generative video compression framework built on a one-step Video Diffusion Transformer. First, we introduce a Unified Unidirectional Latent Compressor that uses a shared model to efficiently and uniformly compress compact latent slices. We then develop a Video DiT-based One-Step Diffusion Enhancer that uses the reconstructed latent slices as content anchors and performs single-step spatio-temporal perceptual enhancement over an entire group of pictures. Finally, a Hybrid Condition Generator extracts structural, strength, and semantic conditions from the reconstructed content and quantization information. These conditions preserve faithful regions, control the degree of generative enhancement, and supplement content-aware perceptual details during one-step diffusion enhancement. Extensive experiments on multiple standard benchmarks demonstrate that DiffVC-ONE achieves state-of-the-art perceptual quality and temporal consistency with low inference cost.
Chinese Translation
生成视频压缩能够在低比特率下恢复丰富的视觉细节,但同时实现高时间一致性和低推理成本仍然具有挑战性。为了解决这个问题,我们提出了DiffVC-ONE,一个基于扩散的生成视频压缩框架,建立在一步视频扩散变换器之上。首先,我们引入了一个统一的单向潜在压缩器,它使用共享模型高效且均匀地压缩紧凑的潜在切片。然后,我们开发了一个基于视频DiT的一步扩散增强器,利用重建的潜在切片作为内容锚点,并对整个图像组进行单步时空感知增强。最后,一个混合条件生成器从重建内容和量化信息中提取结构、强度和语义条件。这些条件保留了真实区域,控制生成增强的程度,并在一步扩散增强过程中补充内容感知的感知细节。在多个标准基准上的广泛实验表明,DiffVC-ONE以低推理成本实现了最先进的感知质量和时间一致性。
cs.CV / 5 / 2608.20534
Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation
Grounded-Exo2Ego:用于稳健的外向视角到自我视角视频生成的结构化语义基础
Abstract
Generating egocentric video from a single exocentric video is an emerging and important topic for AR/VR and physical AI. Compared with conventional novel view synthesis, exo-to-ego generation is a significantly harder task because the standard geometric conditioning becomes highly unreliable under extreme view changes and large unobservable regions. We present Grounded-Exo2Ego, a principled framework that addresses these challenges at both the architectural and data levels. Architecturally, Grounded-Exo2Ego is a dual-branch video diffusion model that couples a geometric anchoring branch, which conditions the generation on the rendering of a 3D reconstruction, with a novel semantic grounding branch, which goes beyond the prevailing geometry-based approach and improves quality by synthesizing challenging regions based on object-level context. Additionally, we found that the overlooked issue of camera-reconstruction misalignment severely undermines exo-to-ego learning. We thus introduce a camera re-localization algorithm that resolves this issue and substantially improves quality across all metrics. We further develop a fully automated synthetic data engine that generates and renders rigged 3D characters in procedurally generated environments. Evaluation on the challenging EgoExo4D dataset shows that our method outperforms recent state-of-the-art approaches by large margins across all metrics. Detailed ablations validate improvements from each of our contributions at both the data and architectural level.
Chinese Translation
从单个外向视角视频生成自我视角视频是增强现实/虚拟现实(AR/VR)和物理人工智能(AI)中的一个新兴且重要的课题。与传统的新视角合成相比,外向到自我视角的生成任务要困难得多,因为在极端视角变化和大范围不可观察区域下,标准几何条件变得极不可靠。我们提出了Grounded-Exo2Ego,这是一个在架构和数据层面上解决这些挑战的原则性框架。在架构上,Grounded-Exo2Ego是一个双分支视频扩散模型,它将几何锚定分支与新颖的语义基础分支相结合,前者基于3D重建的渲染条件生成,后者超越了现有的基于几何的方法,通过基于对象级上下文合成具有挑战性的区域来提高质量。此外,我们发现被忽视的相机重建错位问题严重削弱了外向到自我视角的学习。因此,我们引入了一种相机重新定位算法来解决这一问题,并在所有指标上显著提高了质量。我们进一步开发了一个完全自动化的合成数据引擎,该引擎在程序生成的环境中生成和渲染绑定的3D角色。在具有挑战性的EgoExo4D数据集上的评估表明,我们的方法在所有指标上大幅超越了最近的最先进方法。详细的消融实验验证了我们在数据和架构层面上每一项贡献所带来的改进。
cs.CV / 6 / 2608.20548
Keep Your Friends Close, and the Right Neighbours Closer: Disaster-Conditioned Kernel-Regularized Graph Attention for Building Damage Classification
把朋友留得近一点,把合适的邻居留得更近一点:灾害条件下的核正则化图注意力用于建筑损坏分类
Abstract
Disaster damage is spatial: buildings rarely fail in isolation. Yet using spatial context for damage classification remains surprisingly underexplored, and many pipelines still rely primarily on per-building appearance cues even when the dominant uncertainty is spatially structured. Complicating matters, the right neighbourhood is not the same across events. Floods, hurricanes, and wildfires can exhibit very different clustering behaviour, making spatial reasoning valuable but easy to misuse - naive context aggregation can improve visual coherence while oversmoothing boundaries or propagating structured errors. We study this tension on xBD (the dataset used in the xView2 challenge) in a controlled post-localization, classification-only setup: each building is represented by a pre/post combined (PPC) patch cropped from the provided polygons, and spatial context is modelled with GPS-derived building graphs. Our approach keeps local evidence "close" by preserving strong spatial relationships in disaster damage patterns, while bringing only the right neighbours "closer" through a disaster-type-conditioned graph model that injects a learnable multi-scale spatial kernel prior into attention, allowing the effective neighbourhood scale to adapt across disaster types rather than being learned as a single global smoothing rule. To discourage coherence-by-smoothing, we add a residual de-correlation loss that penalizes positive Moran's~I in prediction residuals. We evaluate the method under event and dataset shift with a leave-one-event-out (LOEO) protocol on xBD and cross-dataset transfer from xBD to Ida-BD. The model improves macro-F1 and substantially reduces residual spatial autocorrelation under zero-shot event shift, indicating better use of spatial context rather than naive smoothing and enabling more reliable transfer to unseen events within known disaster types.
Chinese Translation
灾害损坏具有空间特性:建筑物很少孤立失效。然而,利用空间上下文进行损坏分类仍然出人意料地未被充分探索,许多流程仍主要依赖于每栋建筑的外观线索,即使主导的不确定性是空间结构化的。更复杂的是,不同事件的合适邻域并不相同。洪水、飓风和野火可能表现出非常不同的聚类行为,使得空间推理变得有价值但容易被误用——天真的上下文聚合可以提高视觉一致性,但会过度平滑边界或传播结构性错误。我们在 xBD(用于 xView2 挑战的数据集)上研究这种紧张关系,采用受控的后定位、仅分类设置:每栋建筑由从提供的多边形裁剪的前后结合(PPC)补丁表示,空间上下文通过 GPS 导出的建筑图进行建模。我们的方法通过保留灾害损坏模式中的强空间关系,使局部证据“近”在一起,同时通过一种基于灾害类型的图模型将仅合适的邻居“更近”,该模型将可学习的多尺度空间核先验注入到注意力中,使得有效的邻域尺度可以在不同灾害类型之间适应,而不是作为单一的全局平滑规则进行学习。为了抑制通过平滑实现的一致性,我们增加了一个残差去相关损失,惩罚预测残差中的正 Moran's I。我们在 xBD 上使用留一事件法(LOEO)协议评估该方法在事件和数据集转移下的表现,并从 xBD 到 Ida-BD 进行跨数据集转移。该模型在零样本事件转移下提高了宏 F1 分数,并显著减少了残差空间自相关,表明更好地利用空间上下文而不是天真平滑,并使得在已知灾害类型内对未知事件的转移更可靠。
cs.CV / 7 / 2608.20557
Learning Prostate Anatomy at Test Time for Cancer Detection in Micro-Ultrasound
在微超声中学习前列腺解剖结构以进行癌症检测
Abstract
Domain shift across clinical centers using different imaging hardware or acquisition protocols remains a fundamental barrier to deploying deep learning models for prostate cancer (PCa) detection. Existing test-time adaptation (TTA) methods address distribution shift through entropy minimization or augmentation-based self-supervision, correcting for statistical differences in image appearance but ignoring the anatomical structure of the target domain. We propose ANT, a segmentation-guided TTA framework that adapts a pretrained cancer detection encoder to the target domain by solving an auxiliary prostate segmentation task at test time, supervised by pseudo-masks from a frozen pretrained segmentation network. By aligning encoder representations to prostate anatomy in the target domain, ANT corrects domain-specific feature drift while preserving cancer-discriminative structure. The model was trained on 693 patients imaged with an earlier-generation micro-ultrasound scanner in a multi-center clinical trial, and evaluated on 118 patients acquired with a newer-generation system across two centers in another clinical trial. Under a leave-one-center-out protocol with identical evaluation conditions across all methods, ANT improves mean AUC by 2.9% and 3.6% at the biopsy-core and patient levels, respectively, over no adaptation, outperforming TTA baselines. Code is available at: https://github.com/ObedDzik/ant.git.
Chinese Translation
不同临床中心使用不同成像硬件或采集协议导致的领域转移仍然是部署深度学习模型进行前列腺癌(PCa)检测的基本障碍。现有的测试时适应(TTA)方法通过熵最小化或基于增强的自我监督来解决分布转移,纠正图像外观的统计差异,但忽视了目标领域的解剖结构。我们提出了ANT,一个基于分割引导的TTA框架,通过在测试时解决辅助的前列腺分割任务,利用来自冻结的预训练分割网络的伪掩膜,来将预训练的癌症检测编码器适应到目标领域。通过将编码器表示与目标领域的前列腺解剖结构对齐,ANT纠正了领域特定的特征漂移,同时保留了癌症区分结构。该模型在一个多中心临床试验中对693名患者使用早期一代微超声扫描仪进行训练,并在另一个临床试验中对118名患者使用新一代系统进行评估。在所有方法的相同评估条件下,采用留一中心法的协议,ANT在活检核心和患者层面分别提高了均值AUC 2.9%和3.6%,超越了无适应的情况,表现优于TTA基线。代码可在以下链接获取:https://github.com/ObedDzik/ant.git。
cs.CV / 8 / 2608.20558
Zero-Shot Color Image Manipulation Localization via Noise Residual Artifact Pattern Analysis
通过噪声残差伪影模式分析实现零样本彩色图像篡改定位
Abstract
Digital cameras embed device-specific artifacts into every acquired image through demosaicing, in-camera post-processing, and lossy compression. These traces constitute a forensic signal that can be exploited to assess image authenticity. Existing passive methods rely predominantly on the green channel of the Bayer residual, discarding the correlated information available in the remaining color channels and typically requiring training data or device enrollment. This work proposes a zero-shot, training-free blind image manipulation localization pipeline that estimates a reference artifact pattern directly from the noise residual of a single suspect image, without assuming a fixed filter configuration, color layout, or block period. The pipeline incorporates a principled denoiser selection criterion based on the acquired-to-interpolated noise variance ratio, a block-level correlation analysis against the estimated reference pattern, and a two-component Gaussian Mixture Model scoring stage that produces a pixel-level tampering probability map. An ablation study evaluates the impact of denoiser choice and block size on localization accuracy, and comparisons against state-of-the-art passive methods demonstrate the competitiveness of the proposed zero-shot approach.
Chinese Translation
数码相机通过去马赛克、机内后处理和有损压缩将设备特定的伪影嵌入每张获取的图像中。这些痕迹构成了一种法医学信号,可以用来评估图像的真实性。现有的被动方法主要依赖于 Bayer 残差的绿色通道,忽略了其他颜色通道中可用的相关信息,并通常需要训练数据或设备注册。本文提出了一种零样本、无训练的盲图像篡改定位管道,该管道直接从单个可疑图像的噪声残差中估计参考伪影模式,而不假设固定的滤波器配置、颜色布局或块周期。该管道结合了基于获取到的噪声方差与插值噪声方差比的原则性去噪器选择标准、针对估计参考模式的块级相关性分析,以及一个产生像素级篡改概率图的双成分高斯混合模型评分阶段。消融研究评估了去噪器选择和块大小对定位精度的影响,并与最先进的被动方法进行比较,展示了所提出的零样本方法的竞争力。
cs.CV / 9 / 2608.20587
Aggregate, Don't Adapt: Subject-Level Posterior Aggregation and Transductive Calibration for Cross-Site Parkinsonian Gait Severity
聚合,而非适应:跨站点帕金森步态严重性主题级后验聚合与传导校准
Abstract
We describe the winning entry to the MoCha 2026 Benchmark and Challenge on Parkinsonian Gait, which predicts MDS-UPDRS gait severity from canonicalized SMPL motion recorded at clinical sites unseen during training. The system reaches 0.6945 macro-F1 on the hidden test and ranked first of 58 entries, ahead of the runner-up at 0.5807 and the organizers' baseline at 0.4289, on a frozen public motion encoder with a single $4\times512$ linear layer. Nearly all of the margin comes from three stages usually treated as bookkeeping: reproducing the reference benchmark's exact head recipe, averaging per-walk posteriors within the subject grouping the organizers ship, and a label-free transductive calibration of the feature mean and the decision operating point. Fine-tuning the encoder lost in four distinct forms, and ten alternative encoders were worse. Every ablation number is a paid read on the hidden test, because our own leave-two-cohort-out cross-validation proved anti-correlated with the deciding score over eleven configurations. We give the negative record in full, and identify our largest gain, subject-level aggregation, as the binding ceiling on this benchmark.
Chinese Translation
我们描述了在2026年MoCha基准与挑战赛中获胜的参赛作品,该作品从在训练期间未见的临床站点记录的标准化SMPL运动中预测MDS-UPDRS步态严重性。该系统在隐藏测试中达到了0.6945的宏观F1分数,并在58个参赛作品中排名第一,领先于亚军的0.5807和组织者的基线0.4289,使用的是一个冻结的公共运动编码器和一个单一的$4 imes512$线性层。几乎所有的优势都来自三个通常被视为记录的阶段:重现参考基准的确切头部配方,在组织者提供的主题分组内平均每次步行的后验,以及对特征均值和决策操作点的无标签传导校准。微调编码器在四种不同形式中均表现不佳,而十种替代编码器的表现更差。每个消融实验的结果都是在隐藏测试中的付费读取,因为我们自己的留二组交叉验证与十一种配置的决定性分数呈反相关。我们完整地提供了负记录,并将我们最大的收益——主题级聚合,确定为该基准的绑定上限。
cs.CV / 10 / 2608.20608
A Dataset-Centric Benchmark of Deep Learning Methods for Grape Leaf Disease Classification and Detection
基于数据集的深度学习方法在葡萄叶病分类与检测中的基准研究
Abstract
Grape leaf disease recognition is important for precision agriculture, enabling early diagnosis, timely intervention, and improved vineyard management. Although deep learning has achieved strong results, many studies rely on few datasets, often acquired under controlled conditions, and may not reflect real vineyard challenges such as complex backgrounds, variable illumination, occlusion, leaf pose, disease severity, and device differences. This paper presents a dataset-centric benchmark of deep learning methods for grape leaf disease classification and detection. We analyze publicly available datasets in terms of disease categories, annotation types, acquisition conditions, image characteristics, class distributions, provenance, and task suitability. Representative models are evaluated in three settings: image-level classification, region-level classification, and object detection. Classification is assessed using accuracy, while detection is evaluated using mAP@50 and mAP@50:95. Cross-dataset experiments further examine transfer between datasets with compatible disease categories but different visual and annotation characteristics. Results show near-saturated classification performance on several controlled or derivative datasets, greater difficulty on heterogeneous datasets, and substantial variation in detection performance across annotation settings. Cross-dataset performance drops sharply, especially for object detection, indicating that shared disease labels do not necessarily define equivalent recognition tasks. The benchmark emphasizes dataset provenance, realistic field evaluation, annotation compatibility, and external validation for reliable vineyard disease recognition.
Chinese Translation
葡萄叶病识别对精准农业至关重要,它能够实现早期诊断、及时干预和改善葡萄园管理。尽管深度学习已取得显著成果,但许多研究依赖于少量数据集,这些数据集通常是在受控条件下获取的,可能无法反映实际葡萄园面临的挑战,如复杂背景、变化的光照、遮挡、叶片姿态、病害严重程度和设备差异。本文提出了一个基于数据集的深度学习方法在葡萄叶病分类与检测中的基准研究。我们分析了公开可用的数据集,包括病害类别、标注类型、获取条件、图像特征、类别分布、来源和任务适用性。我们在三种设置下评估了代表性模型:图像级分类、区域级分类和目标检测。分类使用准确率进行评估,而检测则使用mAP@50和mAP@50:95进行评估。跨数据集实验进一步考察了在具有兼容病害类别但视觉和标注特征不同的数据集之间的迁移。结果显示,在多个受控或衍生数据集上分类性能接近饱和,而在异构数据集上则表现出更大的困难,并且在不同标注设置下检测性能存在显著差异。跨数据集的性能急剧下降,尤其是在目标检测中,表明共享的病害标签并不一定定义等效的识别任务。该基准强调了数据集来源、现实场地评估、标注兼容性和外部验证在可靠的葡萄园病害识别中的重要性。
cs.CV / 11 / 2608.20621
RECOUNT: Reference-guided Counting with Synthetic Visual Exemplars
RECOUNT:基于参考的合成视觉示例计数
Abstract
Text-guided zero-shot object counters excel at spatial localization but categorize poorly on novel or fine-grained classes: natural language is too coarse to fully specify visual identity, so they fail to separate visually similar distractors. Few-shot counters sidestep this with visual exemplars, but require manual annotations on every image. To resolve this dilemma, we introduce RECOUNT, a plug-and-play framework for image-guided zero-shot counting. Rather than specify a category with a text prompt, our key insight is to specify it visually, from a single off-scene reference image. However, we find that a lone reference image provides narrow coverage of a category's appearance and is unreliable across diverse scenes. We therefore repurpose a diffusion model as an automated contrastive data engine that expands the reference into a diverse exemplar gallery, supplying the discriminative detail that text cannot. RECOUNT preserves the class-agnostic proposals of any frozen counter and offloads categorization to a separate visual module (a frozen backbone with a lightweight head trained on this synthetic data) that matches each proposal against the target and distractor galleries. Applied to a frozen counter, RECOUNT attains the best zero-shot accuracy on both benchmarks, cutting counting error (MAE) by 55% on LookAlikes and 21% on PairTally relative to the strongest prior zero-shot counter.
Chinese Translation
文本引导的零-shot物体计数器在空间定位方面表现出色,但在新颖或细粒度类别上的分类能力较差:自然语言过于粗糙,无法完全指定视觉身份,因此它们无法区分视觉上相似的干扰物。少量样本计数器通过视觉示例来规避这一问题,但需要对每张图像进行手动标注。为了解决这一困境,我们提出了RECOUNT,一个图像引导的零-shot计数的即插即用框架。我们的关键见解是,不是通过文本提示来指定类别,而是通过单个场景外参考图像进行视觉指定。然而,我们发现单一的参考图像对类别外观的覆盖范围较窄,并且在多样化场景中不可靠。因此,我们重新利用扩散模型作为自动对比数据引擎,将参考图像扩展为多样化的示例库,提供文本无法提供的区分性细节。RECOUNT保留了任何冻结计数器的类别无关提案,并将分类任务卸载到一个单独的视觉模块(一个冻结的主干网络和一个在该合成数据上训练的轻量级头),该模块将每个提案与目标和干扰物库进行匹配。应用于冻结计数器,RECOUNT在两个基准测试中达到了最佳的零-shot准确率,相较于最强的先前零-shot计数器,LookAlikes的计数误差(MAE)降低了55%,PairTally降低了21%。
cs.CV / 12 / 2608.20639
MV2GF: Multi-view Pedestrian Detection with a Visual Geometric Foundation Model
MV2GF:基于视觉几何基础模型的多视角行人检测
Abstract
Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view map from multi-view images. Recent MVPD methods adopt a unified framework that projects 2D image features into a 3D world space and aggregates them into a single feature. Although they are effective, they struggle to generalize to unseen camera configurations during training due to two main issues. First, they are difficult to capture accurate visual geometry across views in unseen camera configurations. Second, they make detection models highly dependent on distortion patterns during training arising from their image feature projection. To address these, we leverage a visual geometric foundation model and propose MV2GF. This foundation model has exhibited strong generalization in capturing visual geometry across views and predicting accurate 3D attributes in diverse camera configurations. MV2GF fuses task-specific features with general-purpose geometric features extracted by the foundation model to effectively capture the visual geometry even in unseen camera configurations. Furthermore, MV2GF projects each pixel in the image features to an appropriate 3D location using 3D pointmaps predicted by the foundation model, preventing the detection model from depending on distortion patterns during training. Our experiments demonstrate the effectiveness of leveraging a visual geometric foundation model for MVPD and that MV2GF generalizes better than existing methods.
Chinese Translation
多视角行人检测(MVPD)旨在从多视角图像中以鸟瞰图的形式检测行人。近期的MVPD方法采用统一框架,将2D图像特征投影到3D世界空间,并将其聚合为单一特征。尽管这些方法有效,但由于两个主要问题,它们在训练期间难以推广到未见过的相机配置。首先,它们在未见过的相机配置中难以准确捕捉跨视角的视觉几何。其次,它们使得检测模型在训练期间高度依赖于图像特征投影所产生的失真模式。为了解决这些问题,我们利用视觉几何基础模型并提出MV2GF。该基础模型在捕捉跨视角的视觉几何和在多样化相机配置中预测准确的3D属性方面表现出强大的泛化能力。MV2GF将任务特定特征与基础模型提取的通用几何特征融合,以有效捕捉即使在未见过的相机配置中的视觉几何。此外,MV2GF使用基础模型预测的3D点图将图像特征中的每个像素投影到适当的3D位置,从而防止检测模型在训练期间依赖于失真模式。我们的实验表明,利用视觉几何基础模型进行MVPD是有效的,并且MV2GF的泛化能力优于现有方法。
cs.CV / 13 / 2608.20659
Lift, Associate, and Fuse: A Decision-Centric Framework for 2D-to-3D Foundation Model Transfer
提升、关联与融合:一个以决策为中心的二维到三维基础模型转移框架
Abstract
Methods that transfer predictions from two-dimensional foundation models into three-dimensional segmentation are commonly grouped by task or representation. Those groupings obscure the decisions that determine whether a system remains coherent across views: where image evidence is grounded, when observations become one identity, how semantic and granularity conflicts are handled, which information is fused, and what state survives for later queries. We introduce \textbf{Lift, Associate, and Fuse (LAF)}, a decision-centric framework that represents a transfer system as five operators: \textbf{Generate, Associate, Reconcile, Fuse, and Persist/Query}. LAF defines an explicit contract for the persistent carrier---its spatial support, semantic state, identity state, uncertainty, provenance, and supported operations---and identifies the first stage at which discarded evidence becomes unrecoverable. We operationalize the framework as a structured audit protocol and apply it to 161 systems available through 7 August 2026, spanning point-, field-, Gaussian-, object-, graph-, and memory-based carriers. Representation, temporal, relational, and feed-forward stress tests required no additional analytical stage after the final confirmation pass. The resulting decision traces expose four recurring properties: association does not establish identity; carrier design fixes both the query interface and correction boundary; rendered-view, native-3D, and proposal-level evaluations are not interchangeable; and qualifiers such as \emph{training-free}, \emph{real-time}, \emph{open-vocabulary}, and \emph{generalizable} are meaningful only when attached to a stage and a complete cost ledger. LAF therefore supplies a representation-neutral method for comparing existing systems, diagnosing irreversible failures, and specifying revisable 3D perception for future agents.
Chinese Translation
将二维基础模型的预测转移到三维分割的方法通常按任务或表示进行分类。这些分类模糊了决定系统在不同视图间保持一致性的决策:图像证据的基础在哪里,何时观察结果成为同一身份,如何处理语义和粒度冲突,融合哪些信息,以及什么状态会在后续查询中保留。我们引入了 extbf{提升、关联与融合(Lift, Associate, and Fuse, LAF)},这是一个以决策为中心的框架,将转移系统表示为五个操作符: extbf{生成(Generate)、关联(Associate)、调和(Reconcile)、融合(Fuse)和持久/查询(Persist/Query)}。LAF为持久载体定义了一个明确的契约——其空间支持、语义状态、身份状态、不确定性、来源和支持的操作,并识别出被丢弃证据变得不可恢复的第一阶段。我们将该框架操作化为一个结构化审计协议,并将其应用于截至2026年8月7日可用的161个系统,涵盖点、场、Gaussian、对象、图和基于记忆的载体。表示、时间、关系和前馈压力测试在最终确认通过后不需要额外的分析阶段。结果的决策轨迹揭示了四个反复出现的特性:关联并不建立身份;载体设计固定了查询接口和修正边界;渲染视图、原生三维和提案级评估不可互换;以及诸如 extit{无训练(training-free)}、 extit{实时(real-time)}、 extit{开放词汇(open-vocabulary)}和 extit{可泛化(generalizable)}等限定词只有在附加到某个阶段和完整成本账本时才有意义。因此,LAF提供了一种表示中立的方法,用于比较现有系统、诊断不可逆的失败,并为未来的智能体指定可修订的三维感知。
cs.CV / 14 / 2608.20663
Shortcut Learning in a Public Grape Disease Dataset: Annotation Granularity as a Modulator, Not a Cause
公共葡萄病害数据集中的捷径学习:注释粒度作为调节因素,而非原因
Abstract
Public datasets for agricultural disease detection are usually judged fit for use from reported metrics, which say nothing about whether the annotation scheme is internally consistent. On one public grape disease dataset (3288 images, 11995 boxes, 6 classes), varying model capacity, input resolution and detection paradigm yields a test-set mAP50 range comparable to seed-to-seed noise, with the bottleneck at small objects across all five architectures. The finding lies on the data side: one class is annotated at whole-leaf level (median box area 43.16% of the image) while the other five are annotated at lesion level. On 5156 cross-species images containing no grape, 65.7% of the false-positive boxes fall into that one class, an over-representation of 13.41x relative to its share of the training annotations. Counterfactual retraining establishes a causal effect of granularity on the magnitude of the shortcut: shrinking only that class's boxes cuts its cross-species false positives by 66%, and a placebo control confirms the effect is specific to the manipulated class. A manipulation in the opposite direction, with criteria registered in advance, returns a negative result: coarsening the finest class to whole-leaf level (0.57% to 40.37%), matched in box count and share of annotations and with higher in-distribution AP, still leaves its cross-species false positives at zero boxes, while the unmanipulated original class holds 50.0% of them. Annotation granularity is therefore a modulator of this shortcut, not its cause: it can amplify or attenuate a sink that already exists, but cannot create one, and what fixes the destination remains open. We also give a granularity screening statistic requiring neither images nor training, and show airborne lesion-level detection to be optically out of reach. The failure mode is invisible to in-distribution evaluation.
Chinese Translation
农业病害检测的公共数据集通常通过报告的指标来判断其适用性,这些指标并未说明注释方案是否内部一致。在一个公共葡萄病害数据集中(3288张图像,11995个框,6个类别),不同的模型容量、输入分辨率和检测范式导致测试集的mAP50范围与种子间噪声相当,所有五种架构在小物体上存在瓶颈。研究发现问题出在数据方面:一个类别的注释是在整叶级别(中位框面积占图像的43.16%),而其他五个类别则是在病灶级别进行注释。在包含5156张不含葡萄的跨物种图像中,65.7%的假阳性框落入该类别,相对于其在训练注释中的比例,过度代表达到了13.41倍。反事实再训练建立了粒度对捷径大小的因果影响:仅缩小该类别的框将其跨物种假阳性减少66%,而安慰剂对照确认该效应特定于被操控的类别。相反方向的操控,提前注册标准,返回了负结果:将最细类别粗化到整叶级别(0.57%到40.37%),在框数和注释比例上匹配,并且具有更高的分布内AP,仍然使其跨物种假阳性保持在零框,而未操控的原始类别则占据了50.0%。因此,注释粒度是这一捷径的调节因素,而非其原因:它可以放大或减弱已存在的缺口,但无法创造一个,而修复目的地的因素仍然不明确。我们还提供了一种不需要图像或训练的粒度筛选统计,并显示出空气中病灶级别的检测在光学上超出可达范围。失败模式在分布内评估中是不可见的。
cs.CV / 15 / 2608.20682
Aristotelian Manifolds: Leveraging Platonic Perceptual Features for Backpropagation Free Rapid Concept Learning
亚里士多德流形:利用柏拉图感知特征实现无反向传播的快速概念学习
Abstract
This paper formalizes and systematically characterizes Aristotelian Manifolds, a generalized structural framework built upon the Platonic Representation Hypothesis. We position high-capacity foundation models as universal perceptual filters and conduct a comprehensive layer-wise investigation to map how knowledge is functionally synthesized within these latent subspaces. Across diverse architectural paradigms and multi-domain datasets, we rigorously chart the interplay between network depth, dimensionality reduction, and distance metrics. Our characterization reveals that semantic maturation does not follow a singular, monotonic path; instead, different data domains exhibit highly distinct geometric response profiles, characterized by intermediate mound-like peaks for specialized clinical modalities and sigmoidal plateaus for natural visual tasks. By profiling the exact coordinates where these manifolds achieve peak representational efficiency, we establish a predictable taxonomy for layer selection and feature compression. Ultimately, this systematic characterization demonstrates that mapping the internal geometry of frozen representations provides a robust, backpropagation-free, and interpretable framework for understanding and exploiting foundation model latent spaces.
Chinese Translation
本文正式化并系统地描述了亚里士多德流形,这是一种基于柏拉图表征假设的广义结构框架。我们将高容量基础模型视为通用感知过滤器,并进行全面的逐层研究,以映射知识在这些潜在子空间中的功能合成方式。在多种架构范式和多领域数据集上,我们严格绘制了网络深度、维度降低和距离度量之间的相互作用。我们的特征描述揭示了语义成熟并不遵循单一的单调路径;相反,不同数据领域表现出高度不同的几何响应特征,专业临床模式表现为中间的丘状峰,而自然视觉任务则呈现出S型平台。通过描绘这些流形实现峰值表征效率的确切坐标,我们建立了一个可预测的层选择和特征压缩分类法。最终,这一系统化的特征描述表明,映射冻结表征的内部几何结构为理解和利用基础模型潜在空间提供了一个稳健、无反向传播且可解释的框架。
cs.CV / 16 / 2608.20687
TopoSurfel: Closing the Loop between Gaussian Surfels and Meshes for Surface Reconstruction
TopoSurfel:闭合高斯表面元与网格之间的环路以实现表面重建
Abstract
3D Gaussian Splatting has achieved remarkable success in novel view synthesis. However, extracting high-fidelity surfaces directly from 3DGS remains challenging due to its discrete and unstructured nature. Existing 3DGS-based reconstruction methods typically rely on multi-view geometric consistency or local constraints. Without an explicit structured geometric prior during optimization, these methods often struggle to resolve structural ambiguities, leading to artifacts and floaters, particularly in textureless or occluded regions. To address this limitation, we propose TopoSurfel, a novel framework that closes the loop between Gaussian surfels and continuous meshes. Unlike recent methods that incorporate mesh extraction into the differentiable pipeline by introducing auxiliary neural networks or extra per-Gaussian parameters, we dynamically extract a continuous proxy mesh via a non-trainable differentiable iso-surfacing process. Leveraging this differentiable connection, we introduce a mesh-guided surfel evolution strategy, including normal alignment and geometry-aware density control, to effectively suppress floaters and fill surface holes. Furthermore, to address the initialization challenges in large-scale environments, we propose a spatially aware hybrid re-initialization strategy that ensures robust reconstruction across complex scenes. Extensive experiments demonstrate that TopoSurfel achieves competitive geometric reconstruction accuracy while maintaining high-quality mesh-based novel view synthesis. The code for our method is available at https://github.com/Fan-Treasure/TopoSurfel.
Chinese Translation
三维高斯点云渲染(3D Gaussian Splatting)在新视角合成方面取得了显著成功。然而,由于其离散且无结构的特性,直接从3DGS中提取高保真表面仍然具有挑战性。现有基于3DGS的重建方法通常依赖多视角几何一致性或局部约束。由于优化过程中缺乏显式的结构化几何先验,这些方法往往难以解决结构歧义,导致伪影和漂浮点,尤其在无纹理或遮挡区域表现明显。为了解决该限制,我们提出了TopoSurfel,一种闭合高斯表面元(Gaussian surfels)与连续网格之间环路的新框架。不同于近期通过引入辅助神经网络或额外的每个高斯参数将网格提取纳入可微分流程的方法,我们通过非训练的可微分等值面提取过程动态生成连续代理网格。基于这一可微分连接,我们引入了网格引导的表面元演化策略,包括法线对齐和几何感知密度控制,有效抑制漂浮点并填补表面空洞。此外,为应对大规模环境中的初始化挑战,我们提出了一种空间感知的混合重初始化策略,确保在复杂场景中的鲁棒重建。大量实验表明,TopoSurfel在保持高质量基于网格的新视角合成的同时,实现了具有竞争力的几何重建精度。我们的方法代码已开源,地址为:https://github.com/Fan-Treasure/TopoSurfel。
cs.CV / 17 / 2608.20690
Identity-Aware Human-Object Interaction Motion Captioning
身份感知的人物-物体交互动作字幕生成
Abstract
Existing human-object interaction (HOI) motion captioning methods typically describe what happens while referring to the subject using generic terms such as "a person" or "someone", without grounding the caption in subject identity. To address this limitation, we introduce Identity-Aware Human-Object Interaction Motion Captioning task. This task requires each generated caption to specify both the subject identity and the corresponding HOI motion. For example, the model generates "Sub_ID lifts the chair" rather than "A person lifts the chair". For this task, we design identity-aware HOI motion captions based on the BEHAVE and InterCap datasets. We further propose ID-HOINet, which learns from multi-view videos while supporting single-view identity-aware HOI motion caption generation. ID-HOINet contains two core components: Multi-View Identity-Motion Learning Module (MVIML) and Two-Stage Caption Rewriting Strategy (TSCR). MVIML learns from multi-view videos by modeling dependencies across temporal stages and camera viewpoints, capturing identity and interaction motion features. At inference, the TSCR first retrieves the subject identity and generates identity-agnostic HOI motion captions. TSCR then rewrites these captions with the predicted identity to produce the final identity-aware HOI motion captions. Experiments demonstrate that ID-HOINet achieves state-of-the-art performance. Code will be released upon acceptance.
Chinese Translation
现有的人物-物体交互(HOI)动作字幕生成方法通常使用“一个人”或“某人”等通用术语来描述发生的事情,而没有将字幕与主体身份结合起来。为了解决这一局限性,我们提出了身份感知的人物-物体交互动作字幕生成任务。该任务要求每个生成的字幕既要指定主体身份,又要描述相应的HOI动作。例如,模型生成“Sub_ID 抬起椅子”而不是“一个人抬起椅子”。为此任务,我们基于BEHAVE和InterCap数据集设计了身份感知的HOI动作字幕。我们进一步提出了ID-HOINet,该模型从多视角视频中学习,同时支持单视角的身份感知HOI动作字幕生成。ID-HOINet包含两个核心组件:多视角身份-动作学习模块(Multi-View Identity-Motion Learning Module, MVIML)和两阶段字幕重写策略(Two-Stage Caption Rewriting Strategy, TSCR)。MVIML通过建模时间阶段和摄像机视角之间的依赖关系,从多视角视频中学习,捕捉身份和交互动作特征。在推理阶段,TSCR首先检索主体身份并生成无身份信息的HOI动作字幕。然后,TSCR使用预测的身份重写这些字幕,以生成最终的身份感知HOI动作字幕。实验表明,ID-HOINet达到了最先进的性能。代码将在论文接受后发布。
cs.CV / 18 / 2608.20691
Bridging Language and Spherical Space: Object-Centric Control for Text-to-Panorama Generation
连接语言与球面空间:面向对象的文本到全景生成控制
Abstract
Panoramic image generation is increasingly important for immersive applications such as virtual reality, augmented reality, and 3D content creation. Unlike perspective images, panoramic images represent a viewer-centered $360^\circ$ surrounding space, where directional expressions such as left, right, front, and behind play a central role in spatial understanding. However, existing text-to-panorama methods largely rely on implicit spatial reasoning and often fail to faithfully ground object-level directional descriptions in spherical panoramic scenes. A straightforward alternative is to introduce explicit layouts, but requiring manually specified spatial conditions reduces the flexibility of language-based interaction and does not directly resolve the misalignment between egocentric directional language and panoramic image space. To address this issue, we propose PanoCtrl, an object-centric framework for controllable text-to-panorama generation. Our method explicitly bridges natural language and spherical panoramic space by converting textual descriptions into structured object-level spherical conditions and integrating them into the diffusion process. Specifically, we introduce PanoParse, a text-conditioned parser that predicts object semantics and spherical bounding field-of-view (BFoV) parameters, and \textbf{PanoControl}, which injects object-level semantic and spatial guidance into the diffusion transformer through object-aware attention and spatial residual enhancement. To support this task, we construct PanoGround, a dataset with object-level spherical annotations and diverse directional descriptions for controllable panoramic generation. Extensive experiments demonstrate that PanoCtrl achieves state-of-the-art performance in both spatial alignment and image quality.
Chinese Translation
全景图像生成在虚拟现实、增强现实和3D内容创作等沉浸式应用中变得越来越重要。与透视图像不同,全景图像代表了以观众为中心的$360^ extcirc$周围空间,其中左、右、前和后等方向性表达在空间理解中发挥着核心作用。然而,现有的文本到全景方法在很大程度上依赖于隐式空间推理,往往无法在球面全景场景中忠实地将对象级方向描述与实际场景对接。一种简单的替代方案是引入显式布局,但要求手动指定空间条件降低了基于语言的交互灵活性,并未直接解决自我中心方向语言与全景图像空间之间的错位。为了解决这一问题,我们提出了PanoCtrl,一个面向对象的可控文本到全景生成框架。我们的方法通过将文本描述转换为结构化的对象级球面条件,并将其整合到扩散过程中,明确地连接了自然语言和球面全景空间。具体而言,我们引入了PanoParse,一个文本条件解析器,预测对象语义和球面边界视场(BFoV)参数,以及 extbf{PanoControl},通过对象感知注意力和空间残差增强将对象级语义和空间指导注入扩散变换器。为了支持这一任务,我们构建了PanoGround,一个具有对象级球面注释和多样化方向描述的数据集,以便进行可控的全景生成。大量实验表明,PanoCtrl在空间对齐和图像质量方面均实现了最先进的性能。
cs.CV / 19 / 2608.20699
ArtiMo: Agent-Driven Articulated Mesh Animation
ArtiMo:基于代理的关节网格动画
Abstract
Animating articulated 3D meshes via text requires satisfying strict kinematic constraints, modeling causal interactions between parts, and achieving instruction fidelity. Due to the absence of task-specific training data and explicit articulation supervision, existing data-driven mesh animation methods are largely inapplicable to this setting. To address this, we propose ArtiMo, a novel agent-driven framework for text-guided articulated mesh animation. Operating in a zero-shot manner, ArtiMo develops an agentic pipeline powered by Large Language and Vision-Language Models (LLMs/VLMs) to orchestrate motion generation. By synergizing the explicit kinematic constraints of URDF with the agent's reasoning and planning capabilities, it effectively produces causally coherent part motions and interactions without requiring model fine-tuning. To ensure motion correctness, the agent additionally utilizes a visual self-improvement mechanism: generated animations are rendered into compact keyframes and motion cues, enabling the VLM to iteratively diagnose and correct errors. Furthermore, we contribute a new benchmark dataset spanning 21 articulated object categories, featuring high-quality motion annotations enriched with causal relationships. Extensive experiments demonstrate that ArtiMo significantly outperforms baselines, particularly on complex, causally driven motions. The project page is available at https://zou-2004.github.io/ArtiMo/.
Chinese Translation
通过文本对关节3D网格进行动画处理需要满足严格的运动学约束、建模部件之间的因果交互,并实现指令的保真性。由于缺乏特定任务的训练数据和明确的关节监督,现有的数据驱动网格动画方法在此场景中大多不适用。为了解决这一问题,我们提出了ArtiMo,一个新颖的基于代理的文本引导关节网格动画框架。ArtiMo以零样本的方式运行,开发了一个由大型语言模型和视觉-语言模型(LLMs/VLMs)驱动的代理管道,以协调运动生成。通过将URDF的显式运动学约束与代理的推理和规划能力相结合,它有效地产生因果一致的部件运动和交互,而无需模型微调。为了确保运动的正确性,代理还利用了一种视觉自我改进机制:生成的动画被渲染为紧凑的关键帧和运动提示,使得VLM能够迭代地诊断和纠正错误。此外,我们贡献了一个新的基准数据集,涵盖21个关节物体类别,具有丰富因果关系的高质量运动注释。大量实验表明,ArtiMo在复杂的因果驱动运动上显著优于基线方法。项目页面可访问:https://zou-2004.github.io/ArtiMo/。
cs.CV / 20 / 2608.20713
AGIDefect-4K: A Richly Annotated Dataset for AI-Generated Image Defect Detection, Localization and Explanation
AGIDefect-4K:一个丰富注释的用于AI生成图像缺陷检测、定位和解释的数据集
Abstract
Generative AI can now produce highly realistic images, yet current models still exhibit subtle but critical defects that undermine their reliability. While existing AI-generated image (AGI) evaluation benchmarks have made notable progress, comprehensive AGI defect diagnosis remains underexplored. To bridge this gap, we introduce AGIDefect-4K, a richly annotated dataset of 4,000 images from 15 state-of-the-art generative models spanning both open-source and closed-source systems. AGIDefect-4K features hierarchical defect annotations: (1) detection labels identifying whether defects exist, (2) pixel-level segmentation masks localizing defective regions, and (3) detailed textual explanations characterizing defect types and their perceptual impact. Each image is further annotated with an overall quality score. Building on this, we present AGIDA (AGI Defect Assistant), a baseline framework leveraging Multimodal Large Language Models (MLLMs) for joint defect detection, localization, explanation, and quality prediction. Comprehensive benchmarking on AGIDefect-4K reveals that AGI defect understanding remains challenging, underscoring the value of this dataset. The dataset is publicly available at https://github.com/sxfly99/AGIDefect-4K.
Chinese Translation
生成性人工智能现在能够生成高度逼真的图像,但当前模型仍然表现出微妙但关键的缺陷,这削弱了它们的可靠性。尽管现有的AI生成图像(AGI)评估基准已取得显著进展,但全面的AGI缺陷诊断仍然未得到充分探索。为了解决这一问题,我们推出了AGIDefect-4K,这是一个包含来自15个最先进生成模型的4,000张图像的丰富注释数据集,涵盖开源和闭源系统。AGIDefect-4K具有分层缺陷注释:(1)检测标签用于识别缺陷是否存在,(2)像素级分割掩码用于定位缺陷区域,以及(3)详细的文本解释用于描述缺陷类型及其感知影响。每张图像还附有一个整体质量评分。在此基础上,我们提出了AGIDA(AGI缺陷助手),这是一个基线框架,利用多模态大语言模型(MLLMs)进行联合缺陷检测、定位、解释和质量预测。在AGIDefect-4K上的全面基准测试表明,AGI缺陷理解仍然具有挑战性,突显了该数据集的价值。该数据集已在 https://github.com/sxfly99/AGIDefect-4K 上公开发布。
cs.CV / 21 / 2608.20720
AffordAny: Open-World 3D Affordance Grounding from Monocular RGB Images via Vision-Language-Guided Geometric Reasoning
AffordAny:通过视觉-语言引导的几何推理从单目RGB图像进行开放世界3D可供性定位
Abstract
Open-world 3D affordance grounding requires localizing functional object parts in 3D given free-form language queries. Existing methods typically assume pre-built object-centric 3D geometry and closed affordance ontologies, limiting deployment from raw RGB observations. We present AffordAny, an end-to-end framework that uses one monocular RGB image to construct large-scale text-conditioned 3D part supervision, ground affordances with a frozen vision-language model (VLM) guided decoder, and improve open-world generalization through pseudo-label self-training. Our automated pipeline produces a benchmark of 5,334 objects and 10,633 part-level samples spanning 473 categories, an order-of-magnitude increase in categorical diversity over prior work. The decoder progressively fuses frozen Cosmos-2B features with 3D geometry through spatial projection, instruction-conditioned semantic compression, and bidirectional geometry-semantics interaction. Minimal-perturbation pseudo-label self-training further adds new objects without human annotation. Under a systematic generalization protocol evaluating unseen objects, unseen categories, and unseen instruction paraphrases, our approach achieves 0.428 IoU on unseen objects and 0.315 IoU on unseen categories after self-training, with unseen-category mIoU improving by 6.3% relative (p<0.01) and an instruction sensitivity gap of only 0.105, demonstrating effectiveness and robustness of our method.
Chinese Translation
开放世界3D可供性定位要求根据自由形式的语言查询在3D中定位功能性物体部件。现有方法通常假设预构建的以物体为中心的3D几何体和封闭的可供性本体,这限制了从原始RGB观测数据的部署。我们提出了AffordAny,一个端到端框架,利用一幅单目RGB图像构建大规模文本条件的3D部件监督,通过冻结的视觉-语言模型(VLM)引导解码器进行可供性定位,并通过伪标签自训练提高开放世界泛化能力。我们的自动化管道生成了一个包含5,334个物体和10,633个部件级样本的基准数据集,涵盖473个类别,相较于之前的工作在类别多样性上实现了数量级的提升。解码器通过空间投影、指令条件的语义压缩和双向几何-语义交互逐步融合冻结的Cosmos-2B特征与3D几何体。最小扰动伪标签自训练进一步添加了新物体而无需人工标注。在评估未见物体、未见类别和未见指令释义的系统泛化协议下,我们的方法在未见物体上实现了0.428的IoU,在未见类别上实现了0.315的IoU,经过自训练后,未见类别的mIoU相对提高了6.3%(p<0.01),指令敏感性差距仅为0.105,展示了我们方法的有效性和鲁棒性。
cs.CV / 22 / 2608.20740
VisTa3D: A Dataset and Benchmark for Thin Object Reconstruction from Vision, Tactile, and 3D Point Clouds
VisTa3D:一种用于从视觉、触觉和三维点云重建薄物体的数据集和基准测试
Abstract
State-of-the-art 3D reconstruction models, whether from visual, range, or both, tend to underperform on thin objects. This is partially due to the small amount of space such objects occupy in RGB images and in 3D point clouds. To test the extent of their errors, we collected the first thin object dataset comprising of synchronized RGB images, depth maps, and tactile response maps, where each frame is associated with inertial measurements, camera pose and calibration, and groundtruth depth and segmentation maps obtained from laser scanning of thin objects. We hypothesize that tactile data can aid in the reconstruction of thin objects as their response maps provide local shape and deformation information. Our dataset, termed VisTa3D, comprises of 387 scenes covering 70 thin objects over 17 environments. We benchmarked current 3D reconstruction models on VisTa3D and found that, indeed, they exhibit low fidelity on thin objects. To test if tactile data can help, we introduce the first visual-range-tactile 3D reconstruction model as a baseline. Code and data: https://huggingface.co/datasets/shaniaguo/VisTa3D.
Chinese Translation
最先进的三维重建模型,无论是基于视觉、深度还是两者结合,通常在薄物体上表现不佳。这部分是由于这类物体在RGB图像和三维点云中占据的空间较小。为了测试它们的错误程度,我们收集了第一个薄物体数据集,该数据集包含同步的RGB图像、深度图和触觉响应图,每帧都与惯性测量、相机姿态和标定,以及通过激光扫描薄物体获得的真实深度和分割图相关联。我们假设触觉数据可以帮助薄物体的重建,因为它们的响应图提供了局部形状和变形信息。我们的数据集称为VisTa3D,包含387个场景,覆盖70个薄物体,分布在17个环境中。我们在VisTa3D上对当前的三维重建模型进行了基准测试,发现它们在薄物体上的保真度确实较低。为了测试触觉数据是否能提供帮助,我们引入了第一个视觉-深度-触觉三维重建模型作为基线。代码和数据可在:https://huggingface.co/datasets/shaniaguo/VisTa3D获取。
cs.CV / 23 / 2608.20748
Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer
生成基于视觉几何的变换器的多视角对抗样本
Abstract
The Visual Geometry Grounded Transformer (VGGT) enables unified feed-forward 3D reconstruction from multi-view images. However, deploying such a high-performance model may expose critical security vulnerabilities. Traditional adversarial perturbations require costly per-scene optimization, while Universal Adversarial Perturbations (UAPs) rely on a single static pattern and fail to effectively attack VGGT. To address these limitations, we propose \textbf{MVAP-G}, a multi-view adversarial perturbation generator that produces imperceptible consistent perturbations across multiple views in a single feed-forward pass. To ensure perturbation consistency across diverse scenes, we design a cross-view adversarial alignment mechanism to process multi-view images. Experiments demonstrate that MVAP-G significantly degrades VGGT performance without iterative optimization during inference. This work pioneers multi-view adversarial attacks on 3D foundation models, uncovering severe vulnerabilities and underscoring the urgent need for robust 3D vision systems. The code is available at https://github.com/qsong2001/mvap-g.
Chinese Translation
视觉几何基础变换器(VGGT)能够从多视角图像中实现统一的前馈三维重建。然而,部署这样一个高性能模型可能会暴露出严重的安全漏洞。传统的对抗扰动需要昂贵的每场景优化,而通用对抗扰动(UAPs)依赖于单一静态模式,无法有效攻击VGGT。为了解决这些局限性,我们提出了 extbf{MVAP-G},一种多视角对抗扰动生成器,能够在单次前馈过程中产生在多个视角间不可察觉的一致扰动。为了确保在不同场景中的扰动一致性,我们设计了一种跨视角对抗对齐机制来处理多视角图像。实验表明,MVAP-G在推理过程中显著降低了VGGT的性能,而无需迭代优化。这项工作开创了对三维基础模型的多视角对抗攻击,揭示了严重的漏洞,并强调了对强健三维视觉系统的迫切需求。代码可在 https://github.com/qsong2001/mvap-g 获取。
cs.CV / 24 / 2608.20749
Identity-Preserving Text-to-Video Generation via Agentic Enhancement and Semantic Repair
通过代理增强和语义修复实现身份保留的文本到视频生成
Abstract
Identity-preserving video generation aims to synthesize videos that follow natural-language instructions while maintaining the visual identity of a given subject. Recent commercial video generation models have achieved strong visual quality and motion realism, but they still suffer from identity drift, incomplete instruction following, and missing visual details under complex prompts. Since these models are usually closed-source black boxes, directly improving them through parameter optimization is often infeasible. We therefore propose Agentic Enhancement and Semantic Repair (AESR), a lightweight enhancement framework for identity-preserving video generation. To improve prompt construction before generation and mitigate the above failures, AESR introduces a global agentic prompt enhancement module. This module learns model-specific prompting formats from official documentation, acquires human-centered video generation priors from human-interaction data, and accumulates test-domain identity-preserving generation experience into a reusable playbook through an agentic loop. To further repair errors in videos generated with enhanced prompts, AESR introduces a sample-level visual semantic repair module, which uses a VLM to locate erroneous video segments and design repair instructions, edits selected frames into explicit visual references, and guides a video editing model to fix local semantic or identity-related errors. We also adopt a lightweight Mixture-of-Experts selection strategy to choose reliable outputs from different generation and refinement paths. Under the official evaluation protocol of the ACM MM 2026 Identity-Preserving Video Generation Challenge, our system MIPL\_Video ranked first in Track 1, demonstrating the effectiveness of AESR for practical identity-preserving video generation. The code is available at https://github.com/oceanflowlab/AESR.
Chinese Translation
身份保留的视频生成旨在合成遵循自然语言指令的视频,同时保持给定主体的视觉身份。近期的商业视频生成模型在视觉质量和运动真实感方面取得了显著进展,但仍然面临身份漂移、指令执行不完整以及在复杂提示下缺失视觉细节等问题。由于这些模型通常是封闭源代码的黑箱,直接通过参数优化来改进它们往往不可行。因此,我们提出了代理增强和语义修复(Agentic Enhancement and Semantic Repair, AESR),这是一个轻量级的身份保留视频生成增强框架。为了在生成之前改善提示构建并减轻上述问题,AESR引入了一个全球代理提示增强模块。该模块从官方文档中学习模型特定的提示格式,从人机交互数据中获取以人为中心的视频生成先验,并通过代理循环将测试领域的身份保留生成经验积累到可重用的手册中。为了进一步修复使用增强提示生成的视频中的错误,AESR引入了一个样本级视觉语义修复模块,该模块使用视觉语言模型(VLM)定位错误的视频片段并设计修复指令,将选定帧编辑为明确的视觉参考,并指导视频编辑模型修复局部语义或身份相关的错误。我们还采用了一种轻量级的专家混合选择策略,从不同的生成和精炼路径中选择可靠的输出。在ACM MM 2026身份保留视频生成挑战的官方评估协议下,我们的系统MIPL_Video在第一轨道中排名第一,证明了AESR在实际身份保留视频生成中的有效性。代码可在https://github.com/oceanflowlab/AESR获取。
cs.CV / 25 / 2608.20754
SPARK-SAM: Self-Prompt Adaptation with Response Knowledge for SAM in Infrared Small Target Segmentation
SPARK-SAM:基于响应知识的自我提示适应用于红外小目标分割
Abstract
Promptable segmentation models provide a reusable interface, but direct transfer to automatic infrared small-target segmentation (IRSTD) exposes a mismatch between spatial prompts and target-domain mask responses. In a diagnostic using target-covering loose-box prompts deterministically derived from test reference masks, the best official SAM2.1 results are only 4.69%, 1.64%, and 2.28% IoU on NUAA-SIRST, NUDT-SIRST, and IRSTD-1K. We introduce SPARK-SAM (Self-Prompt Adaptation with Response Knowledge for SAM), which learns target-domain response knowledge and conditions the decoder through an image-conditioned joint self-prompt state. Training combines benchmark-mask supervision with reliability-aware response guidance. SPARK-SAM achieves 75.78%, 86.49%, and 68.34% IoU with 0.726M additional parameters, ranking first on two benchmarks among 14 retrained SAM variants and adaptations evaluated as automatic image-to-mask methods. The staged IRSTD-1K diagnostic shows that response adaptation reaches most of the final IoU before the predicted points acquire reliable target grounding. Prompt supervision aligns the predicted prompt candidates with target locations, and frozen-weight interventions measure output sensitivity to the joint self-prompt state. Matched ablations show consistent accuracy gains from response guidance and high-resolution prompt refinement across all three datasets. Code is available at https://github.com/Sakauma/SPARK-SAM.
Chinese Translation
可提示的分割模型提供了可重用的接口,但直接应用于自动红外小目标分割(IRSTD)时,空间提示与目标域掩膜响应之间存在不匹配。在使用从测试参考掩膜确定性推导的覆盖目标的松散框提示进行的诊断中,官方最佳的SAM2.1结果在NUAA-SIRST、NUDT-SIRST和IRSTD-1K上的IoU仅为4.69%、1.64%和2.28%。我们提出了SPARK-SAM(基于响应知识的自我提示适应),该模型学习目标域的响应知识,并通过图像条件的联合自我提示状态对解码器进行条件化。训练结合了基准掩膜监督与可靠性感知的响应指导。SPARK-SAM在增加了0.726M参数的情况下,分别在NUAA-SIRST、NUDT-SIRST和IRSTD-1K上实现了75.78%、86.49%和68.34%的IoU,在14个重新训练的SAM变体和适应中排名第一,评估为自动图像到掩膜的方法。分阶段的IRSTD-1K诊断显示,响应适应在预测点获得可靠目标定位之前,已达到最终IoU的大部分。提示监督将预测的提示候选与目标位置对齐,而冻结权重的干预则测量输出对联合自我提示状态的敏感性。匹配的消融实验显示,响应指导和高分辨率提示精炼在所有三个数据集上均带来了持续的准确性提升。代码可在 https://github.com/Sakauma/SPARK-SAM 获取。
cs.CV / 26 / 2608.20756
Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation
Vis-Poison:在多模态检索增强生成中毒害视觉知识
Abstract
While multimodal retrieval-augmented generation (RAG) systems increasingly rely on images as external knowledge sources, the introduction of poisoned visual evidence can severely compromise multimodal large language model (MLLM) generation. Unlike prior attacks that rely on altering textual metadata, we introduce Vis-Poison, a novel visual knowledge poisoning attack where the poisoned image itself is the attacker-controlled payload, without manipulating captions, summaries, metadata, or other associated text. Specifically, this attack is instantiated through an automated multi-agent method that constructs visually plausible poisoned images. To assess its impact, we evaluate Vis-Poison across two representative multimodal RAG pipelines, four embedding models, and six generation models. Empirically, Vis-Poison achieves an end-to-end attack success rate of 40.16\% to 65.40\% against 30k-entry multimodal knowledge bases in \emph{black-box} settings. Moreover, Vis-Poison remains effective against various MLLMs that can answer correctly from parametric knowledge alone, with an average success rate above 60\%. Code and data are available at https://github.com/SWUFE-DB-Group/Vis-Poison.
Chinese Translation
随着多模态检索增强生成(RAG)系统越来越依赖图像作为外部知识源,注入中毒视觉证据可能严重损害多模态大型语言模型(MLLM)的生成。与以往依赖修改文本元数据的攻击不同,我们提出了Vis-Poison,一种新颖的视觉知识中毒攻击,其中被中毒的图像本身是攻击者控制的有效载荷,而无需操纵标题、摘要、元数据或其他相关文本。具体而言,该攻击通过一种自动化的多代理方法实现,构建视觉上合理的中毒图像。为了评估其影响,我们在两个代表性的多模态RAG管道、四个嵌入模型和六个生成模型上评估了Vis-Poison。从实证结果来看,Vis-Poison在30k条目的多模态知识库的黑箱设置中实现了40.16\%到65.40\%的端到端攻击成功率。此外,Vis-Poison在仅依赖参数知识能够正确回答的各种MLLM中仍然有效,平均成功率超过60\%。代码和数据可在https://github.com/SWUFE-DB-Group/Vis-Poison获取。
cs.CV / 27 / 2608.20759
DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion
DiGS-Avatar:基于UV空间扩散的单图像可动画3D人类重建
Abstract
Single-image 3D human reconstruction often suffers from over-smoothed textures and geometric inconsistencies. While diffusion models improve generative quality, their reliance on multi-view synthesis prior to 3D reconstruction is computationally expensive and prone to view inconsistency. We propose DiGS-Avatar, which reformulates this task as an efficient, diffusion-based UV-latent completion task, ensuring 3D consistency by design. To capture accurate spatial structure, we introduce a teacher-student framework where a multi-view teacher provides geometrically aligned pseudo-ground-truth latents to supervise a single-view diffusion student. Treating this inferred latent as a robust structural skeleton, our method injects high-level semantic features to accurately recover fine textural details without disrupting spatial integrity. The refined representation is then decoded into 3D Gaussian primitives. Extensive experiments demonstrate that DiGS-Avatar achieves state-of-the-art or highly competitive visual fidelity and zero-shot generalization, while reconstructing a fully animatable 3D avatar in just 0.71 seconds. Code is available at https://github.com/KLMAV-CUC/DiGS-Avatar.
Chinese Translation
单图像3D人类重建常常面临过于平滑的纹理和几何不一致的问题。尽管扩散模型提高了生成质量,但其在3D重建之前依赖于多视图合成,这在计算上是昂贵的,并且容易导致视图不一致。我们提出了DiGS-Avatar,将这一任务重新构建为一个高效的基于扩散的UV潜在补全任务,从设计上确保3D一致性。为了捕捉准确的空间结构,我们引入了一个教师-学生框架,其中多视图教师提供几何对齐的伪真实潜在向量,以监督单视图扩散学生。将这一推断的潜在向量视为稳健的结构骨架,我们的方法注入高层次的语义特征,以准确恢复细致的纹理细节,而不破坏空间完整性。经过精炼的表示随后被解码为3D高斯原语。大量实验表明,DiGS-Avatar在视觉保真度和零样本泛化方面达到了最先进或高度竞争的水平,同时在仅0.71秒内重建出一个完全可动画的3D头像。代码可在 https://github.com/KLMAV-CUC/DiGS-Avatar 获取。
cs.CV / 28 / 2608.20763
CARD: Diagnosing Belief to Action Routing Failures in Vision Language Models
CARD:诊断视觉语言模型中的信念到行动路由失败
Abstract
Linear probes and activation steering have uncovered that vision-language models (VLMs) internally represent mental states such as agents' beliefs, knowledge, and intentions. However, it is unclear whether and how these representations are used by downstream predictions along these axes. To close this gap, we introduce Cross-Axis Routing Diagnostic (CARD), which steers activations along one axis while measuring the response of a different axis's prediction. Applied to open-weight VLMs on Relay Chain -- a new cooperative grid-world benchmark we propose -- we diagnose a critical routing failure: models fail to incorporate belief representations into their next action prediction, effectively leaving valuable information about their partners unused.
Chinese Translation
线性探针和激活引导揭示了视觉语言模型(VLMs)内部表示了诸如代理的信念、知识和意图等心理状态。然而,目前尚不清楚这些表示是否以及如何被下游预测所利用。为了解决这一问题,我们引入了交叉轴路由诊断(Cross-Axis Routing Diagnostic,CARD),该方法在一个轴上引导激活,同时测量另一个轴的预测响应。我们将其应用于开放权重的VLMs,基于我们提出的新合作网格世界基准——Relay Chain,诊断出一个关键的路由失败:模型未能将信念表示纳入其下一个行动预测中,实际上使得关于其合作伙伴的有价值信息未被使用。
cs.CV / 29 / 2608.20770
MotionPhys: Detecting AI-Generated Videos via Physical Consistency of Optical-Flow Trajectories
MotionPhys:通过光流轨迹的物理一致性检测 AI 生成的视频
Abstract
Modern AI video generation models can produce videos with high visual fidelity and seemingly smooth temporal transitions. However, visual realism does not necessarily imply physical motion consistency. Existing generative models mainly optimize distribution matching in pixel or latent spaces, without explicitly enforcing real-world constraints such as inertia, continuous forces, and trajectory geometry. Our experiments show that AI-generated videos remain visually plausible over short sequences of consecutive frames, yet fail to preserve physical motion consistency throughout a complete object action, resulting in systematic statistical discrepancies in their motion trajectories. Based on this observation, we introduce MotionPhys, a lightweight and interpretable framework that treats sparse motion trajectories as physical evidence rather than relying on appearance artifacts or generator-specific traces. By modeling the geometric evolution of trajectories across multiple temporal scales, MotionPhys reveals subtle motion inconsistencies that are difficult to capture with conventional visual cues and transforms them into a compact representation for efficient detection. Experiments on multiple datasets show that MotionPhys can effectively detect physical inconsistencies in generated videos and generalizes well across different video generators.
Chinese Translation
现代 AI 视频生成模型能够生成具有高视觉保真度和看似平滑的时间过渡的视频。然而,视觉真实感并不一定意味着物理运动的一致性。现有的生成模型主要在像素或潜在空间中优化分布匹配,而没有明确施加诸如惯性、连续力和轨迹几何等现实世界的约束。我们的实验表明,AI 生成的视频在短时间序列的连续帧中仍然保持视觉上的合理性,但未能在完整的物体动作中保持物理运动的一致性,导致其运动轨迹出现系统性的统计差异。基于这一观察,我们提出了 MotionPhys,一个轻量且可解释的框架,将稀疏运动轨迹视为物理证据,而不是依赖于外观伪影或生成器特定的痕迹。通过在多个时间尺度上建模轨迹的几何演变,MotionPhys 揭示了难以用传统视觉线索捕捉的细微运动不一致性,并将其转化为紧凑的表示,以便于高效检测。在多个数据集上的实验表明,MotionPhys 能够有效检测生成视频中的物理不一致性,并在不同的视频生成器之间具有良好的泛化能力。
cs.CV / 30 / 2608.20788
M2Depth: Unifying Monocular Depth Foundation Priors with Multi-View Stereo
M2Depth:将单目深度基础先验与多视图立体视觉统一起来
Abstract
Deep learning-based Multi-View Stereo (MVS) has advanced significantly but often generalizes poorly to unseen scenes, particularly in occluded areas or regions with limited view overlap. To mitigate this, recent approaches integrate Depth Foundation Models (DFMs) into MVS pipelines to provide monocular depth priors. However, existing methods typically rely on a static, one-way fusion scheme, which fails to fully exploit the complementary strengths of both modalities. We propose a novel framework that overcomes this limitation by tightly coupling a DFM with a cascade MVS pipeline through a bidirectional mutual refinement strategy. Our method leverages MVS depth to resolve the scale ambiguity in monocular predictions, while the monocular depth, in turn, enhances the structural completeness and fine-grained detail of the MVS estimate. Furthermore, we introduce a prior-guided cost volume refinement mechanism that effectively integrates multi-view and monocular information via attention-based fusion and discretized depth bins, thereby promoting local geometric consistency. Extensive experiments demonstrate that our method outperforms state-of-the-art MVS approaches on standard benchmarks, producing more complete and generalizable depth maps with sharp boundaries. Furthermore, although not explicitly designed for sparse-view settings, our framework generalizes remarkably well, competing favorably with even dedicated sparse-view methods while maintaining a superior accuracy-efficiency trade-off.
Chinese Translation
基于深度学习的多视图立体视觉(MVS)已取得显著进展,但在未见场景的泛化能力上往往较差,尤其是在遮挡区域或视图重叠有限的区域。为了解决这一问题,近期的方法将深度基础模型(DFM)集成到MVS管道中,以提供单目深度先验。然而,现有的方法通常依赖于静态的单向融合方案,这未能充分利用两种模态的互补优势。我们提出了一种新颖的框架,通过双向互相精炼策略将DFM与级联MVS管道紧密耦合,从而克服这一限制。我们的方法利用MVS深度来解决单目预测中的尺度模糊,而单目深度则反过来增强MVS估计的结构完整性和细节精度。此外,我们引入了一种基于先验的成本体积精炼机制,通过基于注意力的融合和离散深度区间有效整合多视图和单目信息,从而促进局部几何一致性。大量实验表明,我们的方法在标准基准测试中优于最先进的MVS方法,生成更完整且具有良好泛化能力的深度图,并且边界清晰。此外,尽管并非专门针对稀疏视图设置设计,我们的框架在这方面的泛化能力也表现出色,甚至与专门的稀疏视图方法相比,仍能保持优越的准确性和效率平衡。
cs.CV / 31 / 2608.20791
CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models
CertVLA:针对视觉-语言-动作模型的物理视觉攻击的认证防御
Abstract
Vision-Language-Action (VLA) policies are vulnerable to localized physical perturbations, yet existing certified patch defenses target discrete labels and cannot directly certify continuous, temporally correlated actions. We introduce CertVLA, a certified defense for closed-loop VLA control under bounded patch and texture attacks. CertVLA proposes a calibrated region of behaviorally consistent actions, while deterministic covering masks ensure that at least one checked prediction is attack-free. Specifically, CertVLA normalizes action disagreement by the benign variation of each mask pair and accepts a single-mask anchor only when it remains consistent under every second mask. It then calibrates the resulting max-min-max episode score to provide finite-sample clean coverage. Conjoining query-level decisions extends the action certificate to the complete closed-loop rollout. Furthermore, we prove that against any adaptive attacker satisfying the bounded-support threat model, every rollout certified by CertVLA executes only action chunks consistent with attack-erased clean predictions. Under dual-mask rollout correctness, this consistency certificate further guarantees task success. The certificate is independent of patch content, generation method, and physical transformation. Experiments in simulation and the real world demonstrate the empirical and certified effectiveness of CertVLA against patch attacks, with additional simulation validation on texture attacks.
Chinese Translation
视觉-语言-动作(VLA)策略易受局部物理扰动的影响,然而现有的认证补丁防御主要针对离散标签,无法直接认证连续的、时间相关的动作。我们提出了CertVLA,一种针对有界补丁和纹理攻击的闭环VLA控制的认证防御。CertVLA提出了一种经过校准的行为一致性动作区域,同时确定性的覆盖掩码确保至少有一个检查的预测是无攻击的。具体而言,CertVLA通过每对掩码的良性变化来规范化动作不一致,并仅在单个掩码锚点在每个第二个掩码下保持一致时接受它。然后,它校准生成的最大-最小-最大回合得分,以提供有限样本的干净覆盖。结合查询级决策将动作证书扩展到完整的闭环展开。此外,我们证明了在任何满足有界支持威胁模型的自适应攻击者面前,CertVLA认证的每个展开仅执行与攻击消除的干净预测一致的动作块。在双掩码展开正确性的基础上,这一一致性证书进一步保证了任务成功。该证书独立于补丁内容、生成方法和物理变换。模拟和现实世界中的实验展示了CertVLA在抵御补丁攻击方面的经验和认证有效性,并在纹理攻击上进行了额外的模拟验证。
cs.CV / 32 / 2608.20805
Routing Before Looking: Query-Adaptive Evidence Acquisition for Long-form Video Understanding
先路由再观察:用于长视频理解的查询自适应证据获取
Abstract
Long-form video understanding remains challenging for video agents due to the mismatch between query demands and evidence acquisition strategies. Although recent planning-before-perception methods outperform query-agnostic pipelines, they often rely on a single dominant strategy, either generation-based strategy or retrieval-based strategy, limiting their ability to handle diverse query demands. We propose Route2Look, a lightweight and model-agnostic framework for query-adaptive evidence acquisition in long-form video understanding. Route2Look operates in a Route-Look-Memorize loop with three tools: Global Browse for holistic context, Temporal Ground for explicit temporal cues, and Semantic Retrieve for semantic search. The core component is a routing policy that dynamically selects evidence acquisition tools based on the query. To build this policy, Route2Look adopts a two-stage design: first distilling the routing skill from differential contrastive analysis between generation-based and retrieval-based trajectories, and then applying the distilled skill with hard routing rules and continue-or-stop criteria during inference. Experiments on challenging long-video benchmarks show that Route2Look achieves state-of-the-art performance while maintaining strong frame efficiency across datasets and query types. Oracle routing analysis further reveals the potential of query-adaptive evidence acquisition for future long-form video understanding.
Chinese Translation
长视频理解对于视频代理仍然具有挑战性,因为查询需求与证据获取策略之间存在不匹配。尽管最近的规划先于感知方法在性能上超越了与查询无关的管道,但它们通常依赖于单一的主导策略,无论是基于生成的策略还是基于检索的策略,这限制了它们处理多样化查询需求的能力。我们提出了Route2Look,一个轻量级且与模型无关的框架,用于长视频理解中的查询自适应证据获取。Route2Look在一个路由-观察-记忆循环中运行,使用三种工具:用于整体上下文的Global Browse、用于明确时间线索的Temporal Ground,以及用于语义搜索的Semantic Retrieve。核心组件是一个路由策略,它根据查询动态选择证据获取工具。为了构建这一策略,Route2Look采用了两阶段设计:首先从基于生成和基于检索的轨迹之间的差异对比分析中提炼路由技能,然后在推理过程中应用提炼的技能与硬路由规则和继续或停止标准。对具有挑战性的长视频基准的实验表明,Route2Look在保持强大的帧效率的同时,实现了在各数据集和查询类型上的最先进性能。Oracle路由分析进一步揭示了查询自适应证据获取在未来长视频理解中的潜力。
cs.CV / 33 / 2608.20809
TRACE: Training-time Report-guided and Clinically Ordered Concept Editing
TRACE:训练时报告引导和临床有序概念编辑
Abstract
Breast ultrasound diagnosis relies on clinically meaningful semantic concepts, yet most deep learning methods adopt end-to-end image-to-label paradigms that lack interpretability and robustness. While concept-based approaches offer a promising alternative, they often assume complete annotations or require multimodal inputs at inference, which significantly limits their real-world applicability. To tackle these issues, we propose Training-time Report-guided and Clinically Ordered Concept Editing (TRACE), a training-time report-guided framework that leverages structured radiology reports as privileged concept supervision while enabling image-only diagnosis at test time. TRACE refines image-derived concepts through a teacher-guided editing mechanism within a malignancy-aware ordered concept space. To address incomplete annotations, we introduce Strategic Concept Missing Training (SCMT) and train an image-only self-editor via edit distillation for autonomous concept refinement. Besides, we introduce BUSC, a concept-enriched benchmark linking images, labels, and structured attributes. Experiments across multiple datasets demonstrate that TRACE achieves superior performance and improved cross-domain robustness compared to existing methods.
Chinese Translation
乳腺超声诊断依赖于临床意义的语义概念,但大多数深度学习方法采用端到端的图像到标签范式,缺乏可解释性和鲁棒性。尽管基于概念的方法提供了一个有前景的替代方案,但它们通常假设完全注释或在推理时需要多模态输入,这显著限制了它们在实际应用中的适用性。为了解决这些问题,我们提出了训练时报告引导和临床有序概念编辑(TRACE),这是一个训练时报告引导的框架,利用结构化放射学报告作为特权概念监督,同时在测试时实现仅基于图像的诊断。TRACE通过在恶性肿瘤感知的有序概念空间内的教师引导编辑机制,精炼图像衍生的概念。为了应对不完整的注释,我们引入了战略概念缺失训练(SCMT),并通过编辑蒸馏训练一个仅基于图像的自编辑器,以实现自主概念精炼。此外,我们引入了BUSC,这是一个概念丰富的基准,连接图像、标签和结构化属性。多个数据集的实验表明,TRACE在性能和跨领域鲁棒性方面优于现有方法。
cs.CV / 34 / 2608.20814
Enhancing Localized Reasoning for Long Video Understanding via Efficient Segment-to-Video Supervision
通过高效的片段到视频监督增强长视频理解中的局部推理
Abstract
Though Multimodal Large Language Models (MLLMs) have shown impressive potential in video understanding, long video understanding (LVU) remains challenging since distracting noise in complex and lengthy contexts can obscure localized details, misleading MLLMs to produce incorrect answers. Recent works mitigate these issues by incentivizing deep reasoning to include relevant evidence. However, these methods have two main problems: First, the reinforcement fine-tuning framework (RFT) they leveraged incurs substantial training overheads, including high annotation costs and complicated reward designs. Second, the self-reflective and iterative-perception mechanism in some methods causes lengthy outputs and high inference latency. To alleviate these problems, we propose a novel Segment-to-Video Supervision} method (S2V) to efficiently enhance fine-grained reasoning in LVU. Specifically, we generate question answer pairs (VQA) based on localized segments, and then transfer these segment-based VQA back to the whole video for training. Due to focusing on short segments, segment-based VQA can naturally notice details which tend to be overlooked from a whole-video perspective. Training on such data can enforce MLLMs to correctly associate fine-grained details with QA while avoiding distracting noise in the whole video. The S2V training involves just reinforcement learning (RL) with a simple accuracy reward based on only 10K VQA samples and the resulting S2V model predicts answer using a single forward pass with limited output tokens. Experimental results demonstrate that S2V can consistently improve LVU performance across multiple LVU benchmarks, outperforming both general MLLMs and reasoning-based methods not only in LVU accuracy but also in training and inference efficiency.
Chinese Translation
尽管多模态大型语言模型(MLLMs)在视频理解方面展现了令人印象深刻的潜力,但长视频理解(LVU)仍然具有挑战性,因为复杂和冗长的上下文中的干扰噪声可能会掩盖局部细节,从而误导MLLMs产生错误的答案。近期的研究通过激励深度推理来包含相关证据来缓解这些问题。然而,这些方法存在两个主要问题:首先,它们所采用的强化微调框架(RFT)带来了巨大的训练开销,包括高昂的标注成本和复杂的奖励设计。其次,一些方法中的自反思和迭代感知机制导致输出冗长和推理延迟高。为了解决这些问题,我们提出了一种新颖的片段到视频监督方法(S2V),以高效增强LVU中的细粒度推理。具体而言,我们基于局部片段生成问题答案对(VQA),然后将这些基于片段的VQA转移回整个视频进行训练。由于专注于短片段,基于片段的VQA可以自然地注意到从整体视频角度容易被忽视的细节。在这样的数据上进行训练可以促使MLLMs正确地将细粒度细节与QA关联,同时避免整个视频中的干扰噪声。S2V训练仅涉及强化学习(RL),基于仅10K VQA样本的简单准确性奖励,最终的S2V模型使用单次前向传递和有限的输出标记来预测答案。实验结果表明,S2V能够在多个LVU基准测试中持续提高LVU性能,不仅在LVU准确性上超越一般的MLLMs和基于推理的方法,而且在训练和推理效率上也表现更佳。
cs.CV / 35 / 2608.20868
Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images
识别、定位、关联:从文档图像中端到端提取键值对
Abstract
Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information extraction, leading to multi-stage error propagation. We fine-tune SmolDocling, a compact 256M-parameter vision-language model (VLM), to perform end-to-end key-value extraction directly from document images, jointly solving identification, localization, and association in a single pass without OCR preprocessing. We extend DocTags with specialized key, value, region, and link tags, enabling many-to-many relationships in a unified output sequence. To address data limitations, we design an augmentation pipeline combining synthetic form filling and graph-based crops that preserve complete key-value subgraphs. We further introduce a layout-aware evaluation framework extending text matching with spatial bounding box verification. On FUNSD, XFUND, and a large-scale private dataset, our model outperforms larger zero-shot VLM baselines under layout-aware evaluation, while being 27 times smaller than Qwen2.5-VL (7B) and over 5 times faster at inference. The model weights will be released publicly after publication.
Chinese Translation
文档处理管道传统上将光学字符识别(OCR)引擎与下游模型串联,以进行结构化信息提取,这导致了多阶段的错误传播。我们微调了SmolDocling,一个紧凑的256M参数视觉-语言模型(VLM),以直接从文档图像中执行端到端的键值对提取,联合解决识别、定位和关联问题,无需OCR预处理。我们扩展了DocTags,增加了专门的键、值、区域和链接标签,使得在统一输出序列中能够实现多对多关系。为了解决数据限制,我们设计了一个增强管道,结合合成表单填写和基于图的裁剪,保留完整的键值子图。我们进一步引入了一种布局感知的评估框架,通过空间边界框验证扩展文本匹配。在FUNSD、XFUND和一个大规模私有数据集上,我们的模型在布局感知评估下超越了更大的零-shot VLM基线,同时其模型大小仅为Qwen2.5-VL(7B)的27倍,并且推理速度超过5倍。模型权重将在发表后公开发布。
cs.CV / 36 / 2608.20870
RDANet: Relative Degradation Aware Network for Infrared Small Target Detection
RDANet:相对降级感知网络用于红外小目标检测
Abstract
Infrared small target detection is still challenging in remote sensing imagery, because the targets are extremely small, exhibit weak local contrast, and are often embedded in complex and highly variable backgrounds. In addition to these inherent difficulties, we observe that existing detectors often show unstable performance when the target scale changes or when the scene background varies. This scale- and scene-sensitive degradation indicates that current methods are insufficient in simultaneously preserving target structure during feature downsampling and maintaining discriminative local contrast under background shifts, which finally results in unbalanced detection performance across different conditions. To improve detection robustness, this paper proposes a Relative Degradation Aware Network (RDANet) for infrared small target detection. RDANet consists of two dedicated modules: Multi-Scale Anti-Alias Downsampling (MSAD) and Prototype-Guided Skip Memory (PGSM). MSAD introduces multi-scale anti-alias filtering together with pixel-fold aggregation to reduce aliasing effects during resolution reduction, so that target shape information can be better preserved while irrelevant background responses are suppressed. PGSM further enhances the skip features by retrieving patch-level prototypes from a shared memory and adaptively integrating them into the current representation, which helps maintain stable local contrast cues under diverse scene backgrounds. Experiments on three public benchmarks show that RDANet achieves the best performance on most evaluation metrics, while scale- and background-stratified evaluations indicate more stable behavior across target sizes and scene complexity. The code is available at https://github.com/BIT-RuiLiu/RDANet.
Chinese Translation
红外小目标检测在遥感图像中仍然具有挑战性,因为目标极小,局部对比度弱,并且通常嵌入在复杂且高度变化的背景中。除了这些固有的困难外,我们观察到现有检测器在目标尺度变化或场景背景变化时,性能往往不稳定。这种对尺度和场景敏感的降级表明,当前方法在特征下采样过程中无法同时保持目标结构和在背景变化下维持区分性局部对比度,最终导致不同条件下检测性能的不平衡。为了提高检测的鲁棒性,本文提出了一种相对降级感知网络(Relative Degradation Aware Network,RDANet)用于红外小目标检测。RDANet由两个专门模块组成:多尺度抗混叠下采样(Multi-Scale Anti-Alias Downsampling,MSAD)和原型引导跳跃记忆(Prototype-Guided Skip Memory,PGSM)。MSAD结合多尺度抗混叠滤波和像素折叠聚合,以减少分辨率降低过程中的混叠效应,从而更好地保留目标形状信息,同时抑制无关背景响应。PGSM通过从共享记忆中检索补丁级原型并自适应地将其整合到当前表示中,进一步增强跳跃特征,这有助于在多样的场景背景下维持稳定的局部对比度线索。在三个公共基准上的实验表明,RDANet在大多数评估指标上达到了最佳性能,而按尺度和背景分层的评估则表明其在目标大小和场景复杂性方面表现出更稳定的行为。代码可在 https://github.com/BIT-RuiLiu/RDANet 获取。
cs.CV / 37 / 2608.20874
Multi-Modal Traffic Sign Detection with Semantic Attributes for Autonomous Driving
基于语义属性的多模态交通标志检测用于自动驾驶
Abstract
Reliable traffic sign detection is a prerequisite for the global deployment of autonomous driving systems, where regulatory compliance and road safety depend on perceiving signs correctly across regions, ranges, and weather conditions. Despite recent progress, vision-based methods continue to face three fundamental limitations: poor cross-regional generalization due to high diversity across countries, degraded performance on small-object detection at long ranges (traffic signs occupy as little as $10{\times}10$ pixels at 200m), and fragile temporal tracking under the strongly non-linear perspective distortion that occurs as a vehicle approaches a sign. In this paper, we address the problem of robust, long-range, region-agnostic traffic sign perception by combining camera and Light Detection and Ranging (LiDAR) sensing. We present a multi-modal detection framework whose Intensity-Aware Deformable Fusion module aligns retro-reflective LiDAR cues with camera features, anchoring detection on geometric invariants rather than region-specific visual appearance. We further introduce a dual motion-model tracker that explicitly accounts for non-linear perspective transformations during vehicle approach, substantially improving temporal consistency over linear motion assumptions. Additionally, we develop a semantic attribute classification pipeline that estimates occlusion level, readability, sign embeddedness, and road relevance, providing actionable context to downstream planning. Extensive evaluation on our dataset, spanning 60+ countries and 2,500+ hours of driving data, shows that the proposed pipeline achieves an Object Miss Ratio (OMR) of 0.49% across 221,068 evaluation sequences, demonstrating globally generalizable traffic sign perception in commercial-grade autonomous driving systems.
Chinese Translation
可靠的交通标志检测是全球部署自动驾驶系统的前提,法规遵从和道路安全依赖于在不同地区、距离和天气条件下正确感知标志。尽管近期取得了一些进展,基于视觉的方法仍面临三大根本性限制:由于各国之间的高多样性,跨区域泛化能力差;在远距离小物体检测中的性能下降(在200米处,交通标志的面积仅为$10{ imes}10$像素);以及在车辆接近标志时,由于强非线性透视失真导致的脆弱时间跟踪。本文通过结合摄像头和激光雷达(LiDAR)传感器,解决了鲁棒的、长距离、区域无关的交通标志感知问题。我们提出了一个多模态检测框架,其强度感知可变形融合模块将反射性LiDAR线索与摄像头特征对齐,使检测基于几何不变性,而非区域特定的视觉外观。我们进一步引入了一个双运动模型跟踪器,明确考虑车辆接近过程中非线性透视变换,从而显著提高了时间一致性,超越了线性运动假设。此外,我们开发了一个语义属性分类管道,估计遮挡程度、可读性、标志嵌入度和道路相关性,为下游规划提供可操作的上下文。在我们涵盖60多个国家和2500多个小时驾驶数据的数据集上进行的广泛评估表明,所提出的管道在221,068个评估序列中实现了0.49%的物体漏检率(OMR),展示了在商业级自动驾驶系统中全球可泛化的交通标志感知能力。
cs.CV / 38 / 2608.20882
LoRC: Detecting AI-Generated Images via Low-Rank Collapse in Semantic Residuals
LoRC:通过语义残差中的低秩崩溃检测AI生成图像
Abstract
Modern generators faithfully model macroscopic semantics, producing synthetic images that appear highly realistic. Consequently, decisive forensic cues reside in subtle non-semantic visual discrepancies. To reveal these cues, we revisit AIGI detection from a geometric perspective and identify an architecture-agnostic signature. Specifically, modern generators exhibit low-rank collapse (\textit{i.e.}, rank degeneracy) in the semantic-residual orthogonal subspace while largely preserving the dominant semantic direction. This structural flattening consistently emerges during the final decoding stage, forming a shared bottleneck across diverse generator architectures. Motivated by this signature, we propose \textbf{LoRC}, a framework that decouples semantic dominance to capture the collapsed residual geometry induced by the generative decoding bottleneck. Our method improves accuracy by an average of 7.0\% across multiple benchmarks and achieves 97.0\% accuracy on 39 unseen generators. These results demonstrate strong cross-model generalization and robustness, making LoRC a reliable approach for AIGI detection in complex real-world environments.
Chinese Translation
现代生成器忠实地建模宏观语义,生成的合成图像看起来高度逼真。因此,决定性的取证线索存在于微妙的非语义视觉差异中。为了揭示这些线索,我们从几何角度重新审视AIGI检测,并识别出一种与架构无关的特征。具体而言,现代生成器在语义残差正交子空间中表现出低秩崩溃(即秩退化),同时在很大程度上保留了主导语义方向。这种结构扁平化在最终解码阶段始终出现,形成了不同生成器架构之间的共享瓶颈。基于这一特征,我们提出了 extbf{LoRC},一个解耦语义主导性以捕捉由生成解码瓶颈引起的崩溃残差几何的框架。我们的方法在多个基准测试中平均提高了7.0 ext{%}的准确率,并在39个未见生成器上达到了97.0 ext{%}的准确率。这些结果展示了强大的跨模型泛化能力和鲁棒性,使LoRC成为在复杂现实环境中进行AIGI检测的可靠方法。
cs.CV / 39 / 2608.20884
Breaking High Confidence: Practical Face Impersonation under High-Security Thresholds
突破高置信度:在高安全阈值下的实用人脸冒充
Abstract
Face recognition systems (FRSs) are increasingly deployed in critical real-world services for authentication, such as banking applications and airport identity checks, necessitating stringent security configurations. Consequently, the security vulnerabilities of FRSs have garnered significant attention. While existing studies have extensively explored FRS security, prior analyses have primarily focused on medium-security threshold settings, which are not directly applicable to FRSs operating under high-security constraints. In this paper, we propose the first successful impersonation attack against FRSs under high-security threshold settings. Among various threat models, we focus on a practical and challenging scenario: score-based impersonation attacks under strict rate limits. To precisely evaluate the feasibility of such attacks, we provide a principled mathematical analysis characterizing the gaps in each stage of the attack pipeline. Our method significantly enhances impersonation capabilities in score-based attacks, even under elevated decision thresholds. On the LFW benchmark, with a budget of only 100 confidence score queries per identity, our attack achieves an impersonation success rate exceeding 92\% against Amazon Rekognition at a confidence score threshold of 99-recommended setting for law enforcement scenarios. We further observe consistently robust performance across multiple open-source FRSs evaluated at similarly stringent decision thresholds.
Chinese Translation
人脸识别系统(FRS)越来越多地应用于关键的现实服务中进行身份验证,例如银行应用和机场身份检查,这要求严格的安全配置。因此,FRS的安全漏洞引起了广泛关注。尽管现有研究已广泛探讨了FRS的安全性,但之前的分析主要集中在中等安全阈值设置上,这些设置并不直接适用于在高安全约束下运行的FRS。在本文中,我们提出了针对高安全阈值设置下FRS的首个成功冒充攻击。我们关注的威胁模型是一个实用且具有挑战性的场景:在严格的速率限制下进行基于得分的冒充攻击。为了准确评估此类攻击的可行性,我们提供了一个原则性的数学分析,描述了攻击流程每个阶段的差距。我们的方法显著增强了在基于得分的攻击中的冒充能力,即使在提高的决策阈值下。在LFW基准测试中,仅以每个身份100个置信度得分查询的预算,我们的攻击在99的置信度得分阈值下对Amazon Rekognition的冒充成功率超过92%,该阈值是针对执法场景的推荐设置。我们进一步观察到,在多个开源FRS中,在类似严格的决策阈值下,表现始终稳健。
cs.CV / 40 / 2608.20886
EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking
EviRank:多模态图像重新排序的结构化相关证据
Abstract
Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints. Drawing on rubric- and checklist-based evaluation from NLP, we recast multimodal image re-ranking as a semantic constraint satisfaction problem and propose EviRank, which parses any query - text-only, image-only, or composed - into a unified evidence package: typed criteria across six semantic slots (e.g., entities, attributes, relations), each labelled required, forbidden, or ignorable. Re-ranking then reduces to evidence-conditioned verification, combining deterministic rubric scoring and evidence-grounded listwise comparison in a single training-free procedure. The explicit evidence can further serve as structured supervision for optionally distilling a lightweight student. Across five benchmarks spanning text-to-image, image-to-image, and composed image retrieval, EviRank achieves state-of-the-art performance, and the distilled student preserves over 90% of the teacher's capability at substantially lower cost.
Chinese Translation
现实世界中的图像搜索查询是多模态和组合性的:“在粉色中找到这件衬衫”指定了一个要保留的实体、一个要修改的属性和一个要忽略的上下文。然而,现有的重新排序方法要么将这种多方面的相关性压缩成不透明的嵌入,要么依赖于自由形式的思维链,这很容易遗漏或虚构细粒度的约束。借鉴自然语言处理中的评分标准和检查表评估,我们将多模态图像重新排序重新构建为一个语义约束满足问题,并提出EviRank,它将任何查询(仅文本、仅图像或组合)解析为一个统一的证据包:跨六个语义槽(例如,实体、属性、关系)的类型化标准,每个标准标记为必需、禁止或可忽略。重新排序随后简化为基于证据的验证,结合确定性的评分标准和基于证据的列表比较,形成一个单一的无训练过程。显式证据还可以进一步作为结构化监督,用于选择性地提炼轻量级学生。在涵盖文本到图像、图像到图像和组合图像检索的五个基准测试中,EviRank实现了最先进的性能,而提炼的学生在显著较低的成本下保留了超过90%的教师能力。
cs.CV / 41 / 2608.20890
A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving
基于VLA的端到端自主驾驶的协同多模态交互
Abstract
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for end-to-end autonomous driving by jointly integrating perception, reasoning, and decision making within a unified multimodal framework. However, most existing VLA models formulate end-to-end autonomous driving as a visual question answering task, leading to unreliable and less interpretable decision reasoning. In addition, they fail to establish effective multi-modal interaction across heterogeneous sensors, thereby limiting robust scene perception and reliable driving reasoning in long-tail driving scenarios. To this end, we propose a robust VLA-based end-to-end autonomous driving system that combines multi-modality interaction with multi-trajectory planning and optimization, enabling more reliable, interpretable, and safer driving decisions. Our method comprises three core components: (1) Affinity-Guided Optimal Transport for main-auxiliary modality two-way interaction; (2) Distribution-Consistent Modality Transfer for heterogeneous modality distribution transfer and cross-modal interaction; (3) Multi-modal Multi-Trajectory Planning along with Perception-Oriented Trajectory Refinement for better driving decisions to long-tail driving scenarios. Experimental results in open-loop and closed-loop datasets demonstrate improvements in safety long-horizon driving reasoning and road scene perception over existing driving systems, highlighting the ability of our mutli-modality interaction and multi-trajectory planning and optimization for scalable VLA-based systems.
Chinese Translation
视觉-语言-行动(VLA)模型已成为端到端自主驾驶的强大范式,通过在统一的多模态框架内共同整合感知、推理和决策。然而,大多数现有的VLA模型将端到端自主驾驶表述为视觉问答任务,这导致决策推理不可靠且可解释性较差。此外,它们未能在异构传感器之间建立有效的多模态交互,从而限制了在长尾驾驶场景中的稳健场景感知和可靠驾驶推理。为此,我们提出了一种基于VLA的稳健端到端自主驾驶系统,该系统结合了多模态交互与多轨迹规划和优化,从而实现更可靠、可解释和安全的驾驶决策。我们的方法包括三个核心组件:(1)基于亲和力的最优传输,用于主-辅助模态的双向交互;(2)分布一致的模态转移,用于异构模态分布转移和跨模态交互;(3)多模态多轨迹规划与感知导向的轨迹优化,以便在长尾驾驶场景中做出更好的驾驶决策。在开放环和闭环数据集上的实验结果表明,与现有驾驶系统相比,我们的方法在安全性、长时间驾驶推理和道路场景感知方面有所改善,突显了我们多模态交互和多轨迹规划与优化在可扩展基于VLA的系统中的能力。
cs.CV / 42 / 2608.20905
EmotionDialogCN: A Spontaneous Multimodal Dataset for Mandarin Emotional Dialogue
EmotionDialogCN:一个自发的多模态中文情感对话数据集
Abstract
Face-to-face audiovisual interaction is central to human communication, conveying rich emotional and social cues. However, existing multimodal dialogue datasets remain limited by inadequate emotion annotations, poor emotional diversity, and small scale. We introduce EmotionDialogCN, a large-scale audiovisual-emotional dataset designed to capture authentic face-to-face communication. It contains 21,880 dialogue sessions performed by 119 professional actors across 20 everyday scenarios, covering 18 emotion categories with over 400 hours of recordings, the largest and most comprehensive dataset of its kind. A novel data collection framework minimizes equipment interference, enabling natural and nuanced emotional expressions. EmotionDialogCN achieves an emotion distribution deviation of 0.64 from real human emotion statistics (versus 5.65 for prior datasets) and consistent subject framing (52-59% frame occupancy). Together, these properties translate into stable unimodal and multimodal performance across acoustic, lexical, and visual modalities, with fusion results further underscoring strong multimodal alignment and cross-modal complementarity.
Chinese Translation
面对面的视听互动是人类沟通的核心,传达丰富的情感和社会线索。然而,现有的多模态对话数据集在情感标注不足、情感多样性差和规模小等方面仍然存在局限。我们介绍了EmotionDialogCN,这是一个大规模的视听情感数据集,旨在捕捉真实的面对面交流。该数据集包含119名专业演员在20种日常场景中进行的21,880个对话会话,涵盖18种情感类别,录音时长超过400小时,是同类数据集中规模最大、内容最全面的。一个新颖的数据收集框架最小化了设备干扰,使得情感表达自然且细腻。EmotionDialogCN的情感分布偏差为0.64,远低于以往数据集的5.65,并且具有一致的主体框架(52-59%的框架占用率)。这些特性共同促成了声学、词汇和视觉模态之间的稳定单模态和多模态表现,融合结果进一步强调了强大的多模态对齐和跨模态互补性。
cs.CV / 43 / 2608.20910
InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter
InfinityEdit:使用轻量级编辑点火适配器进行无限视频编辑
Abstract
With large pretrained models, existing methods have effectively improved instruction-based video editing. However, most of them rely on an in-place editing assumption. They align the edited video with the given source clip frame by frame over a fixed time span. This pattern fails for open-ended streams, e.g., restyling a live game or applying a camera move to an ongoing shot. In such cases, edits must extend to future frames as they arrive, rather than be applied to a static input clip. In this paper, we study this setting and name it infinite video editing: given a preceding segment and an edit request, a model must generate the next segment that continues the stream while applying the requested edit. This process repeats as an unbounded sequence of edit instructions arrives. This task brings two challenges: the edit must be a faithful continuation rather than a frame-wise rewrite, and generation quality must remain stable as edits accumulate. To address them, we first design a data-collection pipeline for infinite video editing. Based on the collected data, we propose InfinityEdit, a lightweight edit adapter that equips a streaming video generator with unbounded editing ability. The adapter contains three attention modules. History cross-attention guides the denoising frames using the input frames. Temporal causal self-attention keeps temporal cues flowing only from earlier frames to later ones. Edit cross-attention injects the edit request into generation. During inference, the adapter is activated only in the chunk where an edit request arrives. Subsequent chunks are generated by the original model with a reset anchor frame. This scheme applies the edit while preserving the original model's infinite generation ability. Extensive experiments show that InfinityEdit faithfully continues the stream under each edit, and stays stable over unbounded edit sequences.
Chinese Translation
随着大型预训练模型的出现,现有方法有效地提升了基于指令的视频编辑。然而,它们大多数依赖于就地编辑假设,逐帧对齐编辑后的视频与给定的源剪辑,在固定的时间跨度内进行处理。这种模式在开放式流媒体中失效,例如,重新风格化一场直播游戏或对正在进行的镜头应用摄像机移动。在这种情况下,编辑必须扩展到未来的帧,而不是应用于静态输入剪辑。本文研究了这一设置,并将其命名为无限视频编辑:给定一个前段和一个编辑请求,模型必须生成下一个段落,以继续流媒体并应用请求的编辑。随着无限编辑指令的到来,这一过程不断重复。这一任务带来了两个挑战:编辑必须是忠实的延续,而不是逐帧重写,并且随着编辑的积累,生成质量必须保持稳定。为了解决这些问题,我们首先设计了一个无限视频编辑的数据收集管道。基于收集的数据,我们提出了InfinityEdit,一个轻量级的编辑适配器,使流媒体视频生成器具备无限编辑能力。该适配器包含三个注意力模块。历史交叉注意力通过输入帧引导去噪帧。时间因果自注意力仅保持时间线索从早期帧流向后期帧。编辑交叉注意力将编辑请求注入生成过程中。在推理过程中,适配器仅在编辑请求到达的块中激活。后续块由原始模型生成,并使用重置锚帧。这一方案在保留原始模型无限生成能力的同时应用编辑。大量实验表明,InfinityEdit在每次编辑下都能忠实地继续流媒体,并在无限编辑序列中保持稳定。
cs.CV / 44 / 2608.20913
Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization
可解释的深度伪造检测:特征鲁棒增强与证据驱动的解释优化
Abstract
Explainable deepfake detection extends binary classification by requiring models to not only predict authenticity but also provide interpretable justifications. This expanded scope is critical in practice, where users like forensic analysts need insight into the rationale behind the detection. Despite advancements, current approaches suffer from two critical deficiencies: (1)vulnerability to image quality degradation: detection accuracy plummets on low-quality samples, while naive augmentation strategies may induce feature drift and impair performance as diversity expands. (2) factually flawed explanations: explanation models may omit manipulation evidence or hallucinate irrelevant details, undermining interpretability. To address it, we propose a framework with two innovations. For robust deepfake detection, we introduce Feature-robust Augmentation, which comprises diversified degradation-aware augmentation strategies, and a supervised contrastive learning pattern paired with a mean-teacher architecture that stabilizes features against augmentations through consistency constraints. For explanation, we devise an evidence-grounded preference optimization process that guides model to prioritize genuine manipulation traces by learning from chosen-rejected explanation pairs, where rejected samples are constructed via evidence omission or irrelevant information injection. The proposed approach wins the first place in ACM Multimedia 2026 Explainable Deepfake Detection Challenge.The code is available at https://github.com/oceanflowlab/EDD.git.
Chinese Translation
可解释的深度伪造检测不仅要求模型进行二元分类以预测真实性,还需提供可解释的理由。这一扩展的范围在实际应用中至关重要,因为法医分析师等用户需要了解检测背后的推理。尽管已有进展,目前的方法仍存在两个关键缺陷:(1)对图像质量下降的脆弱性:在低质量样本上的检测准确率大幅下降,而简单的增强策略可能导致特征漂移并随着多样性的增加而削弱性能。(2)事实错误的解释:解释模型可能遗漏操控证据或虚构无关细节,从而削弱可解释性。为了解决这些问题,我们提出了一个包含两项创新的框架。为了实现鲁棒的深度伪造检测,我们引入了特征鲁棒增强(Feature-robust Augmentation),该方法包括多样化的降质感知增强策略,以及一种与均值教师架构相结合的监督对比学习模式,通过一致性约束来稳定特征。对于解释,我们设计了一种证据驱动的偏好优化过程,该过程指导模型优先考虑真实的操控痕迹,通过学习选择-拒绝的解释对,其中拒绝样本是通过证据遗漏或无关信息注入构建的。所提出的方法在2026年ACM多媒体可解释深度伪造检测挑战赛中获得第一名。代码可在 https://github.com/oceanflowlab/EDD.git 获取。
cs.CV / 45 / 2608.20916
Semantically Compatible Knowledge Distillation for Cross-Domain Object Detection with Vision Foundation Models
基于语义兼容的知识蒸馏方法在视觉基础模型下的跨域目标检测
Abstract
Vision foundation models (VFMs) offer strong generalization capabilities for domain-adaptive object detection (DAOD). However, existing VFM-based methods overlook the spatial-scale discrepancy between teacher and student feature maps, resulting in semantic incompatibility that weakens both feature alignment and pseudo-label learning. Moreover, domain shift can cause source-trained VFM teachers to miss target-domain objects, limiting the quality of their pseudo-labels. To address these issues, we propose the Semantic Localization-Enhanced Teacher (SLE-T), a semantically compatible knowledge-distillation framework built around a lightweight SLE Adapter for DINOv2. SLE Adapter injects pretrained local-texture priors into DINOv2 to improve cross-domain recognition and reformulates its features into dense representations that are spatially and semantically compatible with the student detector. SLE-T transfers the resulting teacher knowledge through either pseudo-label learning or feature alignment. We instantiate SLE-T with DINOv2-B and DINOv2-L (the ViT-B and ViT-L variants) and compare them with the larger DINOv2-G teacher. Extensive experiments on three DAOD benchmarks demonstrate that our method achieves state-of-the-art performance, and ablation studies confirm the importance of teacher-student semantic compatibility. Notably, SLE-T with DINOv2-B produces competitive or superior pseudo-labels using approximately one-quarter of the training time of DINOv2-G and substantially less GPU memory, demonstrating efficient VFM knowledge transfer under limited computational resources.
Chinese Translation
视觉基础模型(VFM)在领域自适应目标检测(DAOD)中展现了强大的泛化能力。然而,现有基于VFM的方法忽视了教师和学生特征图之间的空间尺度差异,导致语义不兼容,从而削弱了特征对齐和伪标签学习。此外,领域转移可能导致源域训练的VFM教师无法识别目标域对象,限制了其伪标签的质量。为了解决这些问题,我们提出了语义定位增强教师(SLE-T),这是一个围绕轻量级SLE适配器为DINOv2构建的语义兼容知识蒸馏框架。SLE适配器将预训练的局部纹理先验注入DINOv2,以改善跨域识别,并将其特征重新构造成在空间和语义上与学生检测器兼容的密集表示。SLE-T通过伪标签学习或特征对齐转移生成的教师知识。我们使用DINOv2-B和DINOv2-L(ViT-B和ViT-L变体)实例化SLE-T,并与更大规模的DINOv2-G教师进行比较。在三个DAOD基准上的广泛实验表明,我们的方法实现了最先进的性能,消融研究确认了教师与学生之间语义兼容性的重要性。值得注意的是,使用DINOv2-B的SLE-T在大约四分之一的训练时间内生成了具有竞争力或更优的伪标签,并且显著减少了GPU内存,展示了在有限计算资源下高效的VFM知识转移。
cs.CV / 46 / 2608.20929
GAP-SAM: A Global Artifact Prior for Generalizable AI-Generated Image Manipulation Localization
GAP-SAM:一种用于通用AI生成图像篡改定位的全局伪影先验
Abstract
AI-generated image manipulation localization identifies edited pixels, but its OOD performance lags behind image-level detection partly because pixel supervision entangles forensic evidence with dataset-specific mask geometry and semantic boundaries. Extending image-level distribution alignment to localization, we construct COCO-ControlNet with source-image Canny edges and depth maps to align semantics and geometry, improving OOD performance across multiple localizers. Yet tighter Mask-VAE Reconstruction Alignment (Mask-VAE) underperforms COCO-ControlNet, showing that VAE reconstruction artifacts transfer poorly to local diffusion-inpainting artifacts. We also identify \emph{boundary adhesion}, where fine-tuned segmentation models snap predictions to semantic object contours rather than true manipulation boundaries. These findings motivate GAP-SAM, which encodes an image and its frozen VAE reconstruction into a global artifact token and injects it into SAM3's feature pyramid via zero-gated FiLM before pixel decoding. Without prescribing a spatial region, this token modulates dense decoding to preserve localization while suppressing semantic-boundary shortcuts. Across six datasets, GAP-SAM averages 79.8 Pixel-F1, outperforming the strongest prior method by 12.6 points. It also performs best at every tested severity of JPEG compression, Gaussian blur, and resizing.
Chinese Translation
AI生成图像篡改定位旨在识别被编辑的像素,但其在分布外(OOD)性能上落后于图像级检测,部分原因是像素级监督将取证证据与数据集特定的掩码几何形状和语义边界纠缠在一起。通过将图像级分布对齐扩展到定位任务,我们构建了COCO-ControlNet,利用源图像的Canny边缘和深度图对语义和几何进行对齐,从而提升了多个定位模型的OOD性能。然而,更严格的Mask-VAE重建对齐(Mask-VAE)表现不及COCO-ControlNet,表明VAE重建伪影难以有效迁移到局部扩散修复伪影。我们还发现了“边界粘附”现象,即微调的分割模型倾向于将预测结果贴合语义对象轮廓,而非真实的篡改边界。这些发现促使我们提出了GAP-SAM,该方法将图像及其冻结的VAE重建编码为一个全局伪影标记,并通过零门控FiLM注入到SAM3的特征金字塔中,作用于像素解码之前。该标记无需预设空间区域,即可调节密集解码过程,既保持定位能力,又抑制语义边界捷径。在六个数据集上的平均Pixel-F1达到79.8,较最强的现有方法提升12.6个百分点。在JPEG压缩、高斯模糊和图像缩放的各个测试强度下,GAP-SAM均表现最佳。
cs.CV / 47 / 2608.20932
OccluRank: Controllable Occlusion-Aware Layout-to-Image Generation by Adding Just an Ordinal Rank
OccluRank:通过仅添加一个序数等级实现可控的遮挡感知布局到图像生成
Abstract
Layout-to-image generation enables explicit spatial control through bounding-box layouts, yet bounding boxes specify only instance locations and cannot represent their occlusion order. Existing methods may rely on additional geometric conditions, employ complex inference procedures, or aggregate independently constructed instance representations without explicitly modeling their occlusion-dependent interactions. We propose OccluRank, a simple and controllable occlusion-aware layout-to-image framework that augments each bounding box with only one ordinal rank. OccluRank encodes the user-specified occlusion order through lightweight rank-based conditioning and introduces an Order-aware Instance Interaction (OII) module to jointly update rank-conditioned instance representations before aggregation. This allows the specified order to guide information exchange among occluding instances without additional geometric inputs or specialized inference-time optimization. We further construct OccluLayout, a synthetic training dataset whose occlusion order and amodal annotations are derived directly from known scene geometry rather than estimated from partially occluded images using auxiliary prediction models. For comprehensive evaluation, we introduce OccluLayout-Bench, which uses multiple multimodal large language model evaluators to assess instance presence, spatial layout, attributes, and occlusion order, together with FID for overall image quality. Experiments show that OccluRank more reliably preserves target instances, follows specified layouts, and realizes desired occlusion relationships while maintaining comparable attribute consistency and overall image quality.
Chinese Translation
布局到图像生成通过边界框布局实现了明确的空间控制,然而边界框仅指定实例位置,无法表示其遮挡顺序。现有方法可能依赖于额外的几何条件,采用复杂的推理过程,或聚合独立构建的实例表示,而没有明确建模其依赖遮挡的交互。我们提出了OccluRank,一个简单且可控的遮挡感知布局到图像框架,仅通过一个序数等级增强每个边界框。OccluRank通过轻量级的基于等级的条件编码用户指定的遮挡顺序,并引入了一个顺序感知实例交互(Order-aware Instance Interaction, OII)模块,在聚合之前共同更新基于等级的实例表示。这使得指定的顺序能够指导遮挡实例之间的信息交换,而无需额外的几何输入或专门的推理时间优化。我们进一步构建了OccluLayout,一个合成训练数据集,其遮挡顺序和模态注释直接源自已知场景几何,而不是通过辅助预测模型从部分遮挡图像中估计。为了进行全面评估,我们引入了OccluLayout-Bench,使用多个多模态大型语言模型评估器来评估实例存在、空间布局、属性和遮挡顺序,并结合FID评估整体图像质量。实验表明,OccluRank在更可靠地保留目标实例、遵循指定布局和实现期望的遮挡关系方面表现更佳,同时保持可比的属性一致性和整体图像质量。
cs.CV / 48 / 2608.20942
LHMCF-Net: A Learned Hyperbolic Mean Curvature Flow Network for Medical Images Segmentation
LHMCF-Net:一种用于医学图像分割的学习超曲率均值流网络
Abstract
Motivated by the classical Chan-Vese model and the ability of deep priors to capture complex spatial structures, we develop a segmentation model that leverages learned hyperbolic mean curvature flow (LHMCF) as a mathematical foundation for integrating feature space data fidelity and deep structural priors within a unified high-dimensional framework. The proposed LHMCF model is governed by a second-order dissipative hyperbolic PDE, where the introduction of a velocity field provides inertia and momentum to the evolving interface. This hyperbolic mechanism enables the contour to bypass noise-induced local minima and propagate coherently through low-contrast or ambiguous regions, addressing limitations inherent to first-order parabolic flows. To solve the continuous LHMCF model, we construct a deep unfolding network, named LHMCF-Net, which maps the iterative numerical procedure of the PDE into a sequence of discrete evolution stages. Each stage corresponds to one physically interpretable update of the underlying dynamical system, allowing the network to inherit the stability and geometric consistency of the PDE while supporting end-to-end optimization. Comprehensive experiments on three publicly available medical segmentation datasets demonstrate that LHMCF-Net achieves superior performance, particularly in challenging scenarios with low contrast and unclear boundaries. These results highlight the effectiveness of embedding hyperbolic geometric evolution into deep unfolding architectures and underscore the potential of physically inspired models for robust medical image segmentation.
Chinese Translation
受经典的Chan-Vese模型和深度先验捕捉复杂空间结构能力的启发,我们开发了一种分割模型,该模型利用学习的超曲率均值流(LHMCF)作为数学基础,在统一的高维框架内整合特征空间数据的保真性和深度结构先验。所提出的LHMCF模型由一个二阶耗散超曲 PDE 控制,其中引入的速度场为不断演变的界面提供了惯性和动量。这一超曲机制使得轮廓能够绕过噪声引起的局部极小值,并在低对比度或模糊区域中一致传播,解决了第一阶抛物流固有的局限性。为了求解连续的LHMCF模型,我们构建了一个深度展开网络,命名为LHMCF-Net,该网络将 PDE 的迭代数值过程映射为一系列离散演化阶段。每个阶段对应于底层动力系统的一个物理可解释更新,使网络能够继承 PDE 的稳定性和几何一致性,同时支持端到端优化。在三个公开可用的医学分割数据集上进行的全面实验表明,LHMCF-Net 在低对比度和边界不清晰的挑战性场景中表现优越。这些结果突显了将超曲几何演化嵌入深度展开架构的有效性,并强调了物理启发模型在稳健医学图像分割中的潜力。
cs.CV / 49 / 2608.20944
SuppreSensing: Expert-Guided Feature Recalibration and Discrepancy Augmentation for Multimodal Object Detection
SuppreSensing:专家引导的特征重校准与差异增强用于多模态目标检测
Abstract
Multimodal object detection in remote sensing faces challenges due to semantic heterogeneity and modality-specific noise interference. To this end, we propose SuppreSensing, which reformulates multimodal fusion as a selective collaboration process that jointly models shared information and modality-specific cues. SuppreSensing first designs an Expert-driven Multimodal Feature Recalibration (EMFR) module, which reformulates shared-consensus extraction as an input-adaptive multi-expert selection process to alleviate the symmetry trap in multimodal fusion. Complementing this, a modality-specific attribute augmentation strategy is employed to enhance specific modality features by modeling bidirectional discrepancy patterns, mitigating cross-modal heterogeneity. Furthermore, we propose an Expert-driven Customized Feature Purification (ECFP) module based on a "specialized inspection-comprehensive analysis-diagnostic update" physical examination paradigm to iteratively filter redundancies and reinforce task-relevant semantics. Extensive experiments on the DroneVehicle and VEDAI datasets demonstrate that SuppreSensing achieves state-of-the-art detection performance. Cross-domain evaluations on natural scene datasets (FLIR and LLVIP) further validate its superior robustness and generalization capability across diverse environmental conditions.
Chinese Translation
遥感中的多模态目标检测面临语义异质性和特定模态噪声干扰等挑战。为此,我们提出了SuppreSensing,将多模态融合重新表述为一种选择性协作过程,联合建模共享信息和特定模态线索。SuppreSensing首先设计了一个专家驱动的多模态特征重校准(EMFR)模块,该模块将共享共识提取重新表述为一种输入自适应的多专家选择过程,以缓解多模态融合中的对称陷阱。作为补充,采用了一种特定模态属性增强策略,通过建模双向差异模式来增强特定模态特征,从而减轻跨模态异质性。此外,我们提出了基于“专业检查-综合分析-诊断更新”物理检查范式的专家驱动定制特征净化(ECFP)模块,以迭代过滤冗余并强化与任务相关的语义。在DroneVehicle和VEDAI数据集上的大量实验表明,SuppreSensing实现了最先进的检测性能。在自然场景数据集(FLIR和LLVIP)上的跨域评估进一步验证了其在多样化环境条件下的卓越鲁棒性和泛化能力。
cs.CV / 50 / 2608.20969
Kinematic Knowledge Maps for Pattern Alignment: Structured Latent Representational Learning in Multimodal Gait Analysis
用于模式对齐的运动学知识图谱:多模态步态分析中的结构化潜在表征学习
Abstract
Multimodal clinical AI is limited by weakly aligned inputs and the absence of domain-specific interpretable representations, particularly when learning from dense video stream, structured time-series, and template-based kinematic text. Here we present ScoliDetect, an explainable framework for adolescent idiopathic scoliosis screening from monocular gait video, built around a kinematic knowledge map (KKM) and complementary template-based kinematic text derived from per-sequence pose statics. KKM is a fixed-index structured representation that encodes gait features across absolute motion, self-skeleton configuration and joint-joint signal correlation, providing anchor-referenced multimodal fusion and factor-level interpretation. We integrate video, KKM, and template-based kinematic text through bidirectional cross-attention with latent-bottleneck aggregation. In a multicenter cohort (n = 1,858 after exclusions), prespecified supervised ablations on an external screening cohort show that KKM-mediated multimodal fusion outperforms unimodal models and late concatenation. Under a staged training protocol, trimodal contrastive pretraining is applied after architecture selection as representation initialization, improving external ROC-AUC from 0.961 to 0.972. Furthermore, the structured nature of the KKM provides inherent, factor-level attributions mapped directly to specific kinematic phases and skeletal indices, offering verifiable interpretability. The results demonstrate that embedding explicit structural topologies into latent spaces significantly enhances both the generalization and explainability of multimodal pattern analysis systems.
Chinese Translation
多模态临床人工智能受到弱对齐输入和缺乏特定领域可解释表征的限制,尤其是在从密集视频流、结构化时间序列和基于模板的运动学文本中学习时。在此,我们提出了ScoliDetect,这是一个用于青少年特发性脊柱侧弯筛查的可解释框架,基于单目步态视频构建,围绕运动学知识图谱(KKM)和从每个序列的姿态静态中派生的补充基于模板的运动学文本。KKM是一种固定索引的结构化表征,编码了绝对运动、自身骨架配置和关节间信号相关性的步态特征,提供了锚定参考的多模态融合和因子级解释。我们通过双向交叉注意力与潜在瓶颈聚合整合视频、KKM和基于模板的运动学文本。在一个多中心队列(排除后n = 1,858)中,外部筛查队列的预先指定监督消融实验表明,KKM介导的多模态融合优于单模态模型和后期连接。在分阶段训练协议下,在架构选择后应用三模态对比预训练作为表征初始化,将外部ROC-AUC从0.961提高到0.972。此外,KKM的结构化特性提供了固有的、因子级的归因,直接映射到特定的运动学阶段和骨骼指标,提供可验证的可解释性。结果表明,将显式结构拓扑嵌入潜在空间显著增强了多模态模式分析系统的泛化能力和可解释性。
cs.CV / 51 / 2608.20974
WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving
WA-JEPA:重新思考视频JEPA范式在自动驾驶中的世界-动作建模
Abstract
Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent feature prediction. However, V-JEPA is built around random-mask completion and deterministic regression, making it fundamentally ill-suited for autonomous driving planning that demands future-directed prediction tightly coupled with action. To address this, we rethink the V-JEPA paradigm and present WA-JEPA, a V-JEPA-native world-action model designed for autonomous driving planning. Instead of random spatiotemporal masking, WA-JEPA employs hybrid future-masked pre-training, where the model infers future latents from observed context. Departing from deterministic regression, we recast future prediction as conditional flow matching over latent futures, which substantially improves the model's ability to generate plausible future latents for downstream planning. Finally, a joint future-action predictor is proposed to denoise future scene tokens and ego trajectories together in a unified spatiotemporal latent space, allowing action supervision to directly shape planning-relevant world representations. Pre-trained on nuPlan videos and fine-tuned on NAVSIM, WA-JEPA reaches 91.7 EPDMS on NAVSIM-v2, surpassing the strongest end-to-end and world-action baselines by 1.6 and 1.3 EPDMS, and, without HUGSIM-specific fine-tuning, attains the best HD-Score of 0.4462 on the closed-loop HUGSIM benchmark under the same evaluation protocol. These results validate V-JEPA-native world-action modeling as a powerful and scalable paradigm for autonomous driving planning. Code is available at https://github.com/AFARI-Research/WA-JEPA.
Chinese Translation
视频联合嵌入预测架构(Video Joint Embedding Predictive Architecture, V-JEPA)通过自监督潜在特征预测从视频中学习强大的时空表示。然而,V-JEPA是围绕随机掩码补全和确定性回归构建的,这使其在需要与动作紧密结合的未来导向预测的自动驾驶规划中根本不适用。为了解决这个问题,我们重新思考了V-JEPA范式,并提出了WA-JEPA,一种为自动驾驶规划设计的V-JEPA原生世界-动作模型。WA-JEPA采用混合未来掩码预训练,而不是随机时空掩码,模型从观察到的上下文中推断未来潜在特征。与确定性回归不同,我们将未来预测重新表述为对潜在未来的条件流匹配,这显著提高了模型生成合理未来潜在特征以用于下游规划的能力。最后,提出了一种联合未来-动作预测器,以在统一的时空潜在空间中共同去噪未来场景标记和自我轨迹,使得动作监督能够直接塑造与规划相关的世界表示。在nuPlan视频上进行预训练,并在NAVSIM上进行微调,WA-JEPA在NAVSIM-v2上达到了91.7 EPDMS,超越了最强的端到端和世界-动作基线1.6和1.3 EPDMS,并且在没有HUGSIM特定微调的情况下,在相同评估协议下达到了闭环HUGSIM基准的最佳HD-Score 0.4462。这些结果验证了V-JEPA原生世界-动作建模作为自动驾驶规划的强大且可扩展的范式。代码可在 https://github.com/AFARI-Research/WA-JEPA 获取。
cs.CV / 52 / 2608.20984
MigrationNarrate: A Dataset for Detection of Migration Narratives in YouTube Videos
MigrationNarrate:一个用于检测YouTube视频中移民叙事的数据集
Abstract
Narratives are central to how social communication is framed, making their detection critical for understanding and analysing public discourse. Prior work has explored narrative detection and extraction across diverse domains; however, migration narratives remain significantly understudied, primarily due to the absence of dedicated annotated datasets. Furthermore, public communication has recently shifted towards video-centric platforms, where narratives are conveyed through multimodal signals and consumed at scale. Despite this shift, narratives in videos remain largely unexplored. To bridge these gaps, we introduce MigrationNarrate, the first multimodal dataset for detection of migration narratives in the UK, consisting of 1,115 YouTube video transcripts annotated using a two-level taxonomy of 12 migration super-narratives and 53 narrative labels. This paper details the dataset design, collection, and annotations; together with benchmark results using a combination of pre-trained encoder models and both open- and closed-source Large Language Models. Finally, a thorough error analysis offers insights for future work.
Chinese Translation
叙事在社会交流的框架中占据核心地位,因此其检测对于理解和分析公共话语至关重要。之前的研究已经在多个领域探讨了叙事的检测和提取;然而,移民叙事仍然显著缺乏研究,主要是由于缺乏专门的标注数据集。此外,公共交流最近已转向以视频为中心的平台,在这些平台上,叙事通过多模态信号传达并大规模消费。尽管发生了这种转变,视频中的叙事仍然在很大程度上未被探索。为了解决这些问题,我们推出了MigrationNarrate,这是第一个用于检测英国移民叙事的多模态数据集,包含1,115个YouTube视频的转录文本,采用12个移民超级叙事和53个叙事标签的两级分类法进行标注。本文详细介绍了数据集的设计、收集和标注;并结合预训练编码器模型和开放源及闭源的大型语言模型的基准结果。最后,全面的错误分析为未来的研究提供了见解。
cs.CV / 53 / 2608.20999
Latent Ordinal Evidence, Misaligned Outputs: Inference-Time Ordinal Lens Alignment for Multimodal LLMs
潜在有序证据与不对齐输出:多模态大语言模型的推理时有序镜头对齐
Abstract
Multimodal LLMs apply the language model interface to visual inputs, where ordinal regression tasks such as age estimation, image quality assessment, and disease grading require autoregressive decisions over ordered class labels. We ask whether MLLMs reliably convert internal ordinal evidence into ordered digit-token outputs. Across four ordinal benchmarks and four MLLM backbones, ordinal labels are linearly recoverable from hidden states with Spearman correlation up to 0.938, and a task-designed prompt further sharpens this structure. Yet native digit-token outputs weakly expose it: the unembedding matrix filters the ordinal direction, and the digit-token row space retains below 1.15% across all 16 model-dataset combinations, with a 16 to 77 absolute-point accuracy gap between linear-probe and native outputs. We introduce Ordinal Lens Alignment (OLA), a frozen-backbone inference-time method that trains lightweight W_S-anchored lenses on mid-to-deep decoder layers, fuses them into an ordinal distribution, and corrects only digit-token logits at generation. OLA outperforms the SOTA LoRA-tuned OrderChain baseline in most settings while keeping the MLLM frozen, surpasses discriminative ordinal baselines in most cells, and improves over an offline lens in every setting.
Chinese Translation
多模态大语言模型(MLLMs)将语言模型接口应用于视觉输入,其中诸如年龄估计、图像质量评估和疾病分级等有序回归任务需要对有序类别标签进行自回归决策。我们探讨MLLMs是否能够可靠地将内部有序证据转换为有序数字标记输出。在四个有序基准测试和四个MLLM骨干网络中,有序标签可以通过隐藏状态线性恢复,斯皮尔曼相关系数高达0.938,且任务设计的提示进一步增强了这一结构。然而,原生数字标记输出对此的揭示较弱:去嵌入矩阵过滤了有序方向,数字标记行空间在所有16个模型-数据集组合中保持在1.15%以下,线性探测和原生输出之间存在16到77个绝对点的准确性差距。我们提出了有序镜头对齐(Ordinal Lens Alignment,OLA),这是一种冻结骨干网络的推理时方法,在中到深的解码器层上训练轻量级的W_S锚定镜头,将其融合为有序分布,并仅在生成时修正数字标记的logits。OLA在大多数设置中优于SOTA LoRA调优的OrderChain基线,同时保持MLLM冻结,在大多数单元中超越了判别有序基线,并在每个设置中都优于离线镜头。
cs.CV / 54 / 2608.21008
Triangulation-Free Bundle Adjustment with Graduated Non-Convexity for Camera Pose Refinement from Coarse Priors
无三角化的束调整与渐进非凸性用于从粗略先验中精炼相机姿态
Abstract
Mobile AR frameworks attach a metric pose prior to every casual phone capture, and turning it into reconstruction-grade poses cheaply on CPU is the step before novel-view synthesis. The least a refiner owes an accurate prior is not to make it worse. The workhorse refiner does. On 15 ScanNet++ iPhone room captures, COLMAP triangulation plus prior-seeded bundle adjustment degrades an accurate ARKit prior in all 15, 0.55 degrees to 0.74 degrees by scene-mean. The cause is the seeding. Structure is triangulated from the prior before anything is optimized, so the prior's error is baked into the structure the optimizer trusts. We remove the triangulation. Every keypoint owns a scalar depth along its own back-projected ray and each match contributes two symmetric cross-projection residuals, so structure is re-expressed at every iterate. The same solve holds the room prior at 0.57 degrees and never fails in 330 perturbed room runs, and at object scale reaches 0.265 degrees/1.80 mm from a prior at 0.456 degrees in a median of 10 s per scene on one CPU, against 2.5 GPU-hours for a learned refiner. Because no structure is committed, the objective also admits graduated non-convexity, which measures how deep the defect goes. Classical refinement collapses past 1-2 degrees of prior error, barely beyond a real ARKit prior, and no classical refinement arm survives 32 degrees. Ours recovers 425 of 425 runs through 16 degrees/80 mm and 85% at 32 degrees/160 mm, and perturbed rooms through 32 degrees. Nominal object-scale accuracy is on par rather than better, on a benchmark at its own noise floor, where classical bundle adjustment is a strong baseline absent from the literature. One scene fails for every solver already at zero perturbation. Re-mapping from position priors matches us in the prior's frame but discards it, so it cannot exploit a prior worth keeping or be warm-started.
Chinese Translation
移动增强现实(AR)框架在每次随意的手机拍摄中附加一个度量姿态先验,而将其廉价地转化为重建级别的姿态是在新视图合成之前的步骤。一个精炼器至少应确保不使准确的先验变得更糟,而现有的精炼器却做不到这一点。在15个ScanNet++ iPhone房间捕获中,COLMAP三角化加上先验引导的束调整在所有15个场景中都使得准确的ARKit先验退化,场景均值从0.55度降至0.74度。原因在于先验的引导。在优化之前,结构是从先验中三角化得到的,因此先验的误差被固化到优化器所信任的结构中。我们去除了三角化。每个关键点沿其自身反向投影光线拥有一个标量深度,每个匹配贡献两个对称的交叉投影残差,因此结构在每次迭代中被重新表达。相同的求解方法使房间先验保持在0.57度,并且在330次扰动房间运行中从未失败,在物体尺度上从0.456度的先验达到0.265度/1.80毫米,平均每个场景在一个CPU上耗时10秒,而学习型精炼器则需2.5个GPU小时。由于没有结构被固定,目标函数也允许渐进非凸性,这测量了缺陷的深度。经典的精炼在1-2度的先验误差后崩溃,几乎超出真实的ARKit先验,而没有经典精炼算法能在32度下存活。我们的算法在16度/80毫米下恢复了425次运行中的425次,并在32度/160毫米下恢复了85%的成功率,并且在32度的扰动房间中也能成功。名义上的物体尺度精度与经典束调整相当,而经典束调整在文献中缺乏强有力的基线。在零扰动的情况下,每个求解器都有一个场景失败。基于位置先验的重新映射在先验框架中与我们匹配,但丢弃了先验,因此无法利用值得保留的先验或进行热启动。
cs.CV / 55 / 2608.21009
Dorsal Hand Images for Immersive (XR) and Privacy-preserving Age Assurance and Child Safety
背侧手部图像用于沉浸式(XR)和隐私保护的年龄验证与儿童安全
Abstract
Ensuring that Extended Reality (XR) environments are age-appropriate is an important regulatory and safety challenge. However, current age assurance operates only at registration and cannot verify the age of the active user during a session. Face-based approaches, the dominant solution in social media and adult platforms, are impractical in XR, because they require removing the headset and taking a self-captured image, often on a mobile app. This both breaks immersion and introduces the privacy risk of sharing face pictures with third parties, which leaves XR platforms without a viable path to continuous, in-session and privacy-preserving age assurance. We propose the dorsal part of the hand as an alternative to the face, by exploiting the egocentric cameras that XR headsets inherently and naturally use to capture gesture interactions. To evaluate this, we collect an age- and sex-stratified, ethnodiverse dataset of 436 participants spanning the minor--adult boundary, captured under unconstrained lighting and orientation conditions. To characterise what is achievable with off-the-shelf methods at the minor--adult boundary, we evaluate standard neural network architectures for age assurance at the legally critical 18-year threshold. Analysis confirms performance is robust to skin-tone variation. On this dataset, the challenge-31 operating point achieves zero minor admission, making the system a viable first-stage filter for age assurance. These findings position dorsal hand morphometrics as an effective and more privacy-preserving biometric modality for in-session age assurance in XR.
Chinese Translation
确保扩展现实(XR)环境适合不同年龄段是一个重要的监管和安全挑战。然而,目前的年龄验证仅在注册时进行,无法在会话期间验证活跃用户的年龄。基于面部的解决方案是社交媒体和成人平台的主流,但在XR中并不实用,因为它们需要用户摘下头显并拍摄自我捕捉的图像,通常是在移动应用上进行。这不仅打破了沉浸感,还引入了与第三方共享面部照片的隐私风险,使XR平台在实现持续、会话内和隐私保护的年龄验证方面面临挑战。我们提出将手背作为面部的替代方案,利用XR头显固有且自然使用的自我中心摄像头捕捉手势交互。为此,我们收集了一个包含436名参与者的年龄和性别分层的多民族数据集,涵盖未成年人与成年人之间的界限,数据采集在不受限制的光照和方向条件下进行。为了表征在未成年人与成年人之间的界限上,现成方法所能实现的效果,我们评估了标准神经网络架构在法律关键的18岁阈值下的年龄验证性能。分析确认了性能对肤色变化的鲁棒性。在该数据集上,挑战-31操作点实现了零未成年人入场,使该系统成为年龄验证的有效第一阶段筛选器。这些发现将手背形态测量定位为在XR中进行会话内年龄验证的有效且更具隐私保护性的生物特征模式。
cs.CV / 56 / 2608.21022
Recognition-Conditioned Reasoning: A Training-Free Multimodal-LLM Pipeline for Fine-Grained Micro-Action Understanding
识别条件推理:一种无训练的多模态大语言模型管道用于细粒度微动作理解
Abstract
Micro-actions are subtle, short, low-amplitude body movements, such as a fidgeting hand or a slight head tilt, that humans perform with little conscious intent yet that reliably leak emotional and psychological state. Understanding them goes beyond assigning a label: a model must also describe which body parts move and reason, faithfully, about why a clip warrants a particular fine-grained category. We present the training-free, prompt-only system that won first place in the fine-grained understanding track (MA-Bench) of the MAC~2026 Micro-Action Challenge, where both fine-tuning and ground-truth supervision are disallowed. Built entirely upon frozen multimodal large language models (MLLMs), the system dynamically routes each of the eight sub-tasks to the MLLM empirically best suited for that task: a discriminative MLLM for closed-ended recognition tasks and a generative MLLM for open-ended description and reasoning tasks. This architecture achieves a statistically significant performance advantage on open-ended tasks, attaining an average score of 2.68 (on a five-point scale) compared to 1.44 for the second-best approach.
Chinese Translation
微动作是微妙、短暂、低幅度的身体运动,例如不自觉的手部动作或轻微的头部倾斜,这些动作通常在无意识的情况下进行,但却可靠地泄露出情感和心理状态。理解微动作不仅仅是给出一个标签:模型还必须描述哪些身体部位在运动,并准确推理出为什么某个片段应归入特定的细粒度类别。我们提出了一种无训练、仅基于提示的系统,该系统在MAC~2026微动作挑战赛的细粒度理解赛道(MA-Bench)中获得了第一名,在该赛道中不允许进行微调和真实标签监督。该系统完全基于冻结的多模态大语言模型(MLLMs),动态地将八个子任务路由到最适合该任务的MLLM:对于封闭式识别任务使用判别式MLLM,对于开放式描述和推理任务使用生成式MLLM。该架构在开放式任务上实现了统计显著的性能优势,平均得分为2.68(满分五分),而第二名方法的得分为1.44。
cs.CV / 57 / 2608.21030
COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models
COMET:对比运动增强的时间推理用于视频多模态大型语言模型
Abstract
Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottleneck is not only sparse frame sampling, but also the lack of a complete temporal modeling pipeline for explicitly representing frame-to-frame change, enabling appearance-motion interaction, and optimizing temporal direction sensitivity. We propose COMET, a temporally grounded framework that systematically strengthens video MLLMs through explicit temporal representation, appearance-motion fusion, and direction-aware optimization. Architecturally, COMET introduces a temporal motion branch built on Taylor frame differences and injects its motion evidence into the appearance stream via temporal attention bias-enhanced cross-attention. For optimization, COMET combines temporal prior distillation with a forward-reverse TC-GRPO stage that turns temporal order into a direct learning signal and strengthens the model's use of directional motion patterns encoded by the temporal motion branch. The method achieves consistent overall improvements with a pronounced motion-temporal bias: on Qwen3-VL-8B, action-centric tasks (STAR, SSv2) improve by 4.9% on average, temporal reasoning tasks (NExT-QA, CLEVRER, LLaVA-178K) by 2.1% over BL-GRPO, while static perception tasks (PerceptionTest) remain on par. The same gain pattern also transfers to InternVL2.5-8B, indicating that COMET generalizes across model families.
Chinese Translation
视频多模态大型语言模型已经取得了显著进展,但细粒度的运动-时间理解仍然脆弱。核心瓶颈不仅在于稀疏的帧采样,还在于缺乏完整的时间建模管道,以明确表示帧与帧之间的变化,促进外观与运动的交互,并优化时间方向的敏感性。我们提出了COMET,一个时间基础的框架,通过明确的时间表示、外观-运动融合和方向感知优化系统性地增强视频多模态大型语言模型(MLLMs)。在架构上,COMET引入了一个基于泰勒帧差异的时间运动分支,并通过时间注意力偏置增强的交叉注意力将其运动证据注入外观流。为了优化,COMET结合了时间先验蒸馏和一个前向-反向TC-GRPO阶段,将时间顺序转化为直接学习信号,并增强模型对时间运动分支编码的方向性运动模式的利用。该方法在整体性能上实现了一致的提升,表现出明显的运动-时间偏向:在Qwen3-VL-8B上,以动作为中心的任务(STAR, SSv2)平均提高了4.9%,时间推理任务(NExT-QA, CLEVRER, LLaVA-178K)较BL-GRPO提高了2.1%,而静态感知任务(PerceptionTest)保持不变。相同的增益模式也转移到InternVL2.5-8B,表明COMET在不同模型系列间具有良好的泛化能力。
cs.CV / 58 / 2608.21041
CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment
CoST:通过时空对齐实现语义感知的城市理解
Abstract
Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applications. Despite recent advances, current methods struggle with cross-region generalization and semantic interpretability due to their reliance on region-specific auxiliary data and the neglect of semantic alignment within multi-temporal urban imagery. Therefore, we present CoST, a novel \underline{Co}ntrastive-based \underline{S}patial-\underline{T}emporal framework that aligns spatial context with multi-temporal semantics to extract universal geographic regularities shared across regions. Specifically, CoST explicitly models spatial correlations to capture transferable geographic structures and exploits multi-year urban change semantics to align learned representations with high-level geo-semantics. Extensive experiments demonstrate that CoST consistently achieves superior performance across various downstream tasks and in unseen scenario, yielding an average relative gain of 8.7\% over the strongest competing methods across eight city-indicator settings. The code is available in \href{https://github.com/Arandinglv/CoST}{this repo}.
Chinese Translation
从卫星图像中进行地理空间表示学习是大规模城市分析和实际应用中的一个基本问题。尽管近期取得了一些进展,但当前的方法由于依赖于区域特定的辅助数据以及忽视了多时相城市图像中的语义对齐,仍然在跨区域泛化和语义可解释性方面面临挑战。因此,我们提出了CoST,一个新颖的基于对比的时空框架,旨在将空间上下文与多时相语义对齐,以提取跨区域共享的普遍地理规律。具体而言,CoST明确建模空间相关性,以捕捉可转移的地理结构,并利用多年的城市变化语义将学习到的表示与高层次的地理语义对齐。大量实验表明,CoST在各种下游任务和未见场景中始终实现了优越的性能,在八个城市指标设置中,相较于最强的竞争方法,平均相对提升达8.7%。代码可在[此仓库](https://github.com/Arandinglv/CoST)获取。
cs.CV / 59 / 2608.21055
CoAnchor: Robust Collaborative Perception under Spatio-Temporal Misalignment via Object-Level Anchors
CoAnchor:通过对象级锚点实现时空错位下的鲁棒协同感知
Abstract
Collaborative perception extends the sensing range of a single vehicle by fusing observations from nearby agents, which improves the robustness of autonomous driving. In realistic deployments, however, the received collaborator messages are often affected by both communication delay and relative-pose noise, which jointly cause stale observations, spatial misalignment, and unstable feature fusion. Existing methods usually address these issues from either the spatial or temporal side, but handling them jointly in a unified and efficient manner remains challenging. In this paper, we propose CoAnchor, an anchor-centric spatio-temporal alignment framework for asynchronous collaborative perception. Instead of directly reasoning on dense BEV features, CoAnchor builds sparse object-level spatio-temporal anchors as a shared interface for pose correction and tightly connects spatial refinement, temporal propagation, and current-time verification within one unified loop, while keeping the overall correction process lightweight. Extensive experiments on both simulated and real-world datasets illustrate that CoAnchor remains competitive under clean settings and improves the robustness under joint delay and pose perturbations with a favorable practical accuracy-efficiency trade-off.
Chinese Translation
协同感知通过融合来自附近代理的观测,扩展了单一车辆的感知范围,从而提高了自动驾驶的鲁棒性。然而,在实际应用中,接收到的协作消息通常受到通信延迟和相对姿态噪声的影响,这共同导致了过时的观测、空间错位和不稳定的特征融合。现有方法通常从空间或时间的角度解决这些问题,但以统一和高效的方式同时处理它们仍然具有挑战性。本文提出了CoAnchor,一种以锚点为中心的时空对齐框架,用于异步协同感知。CoAnchor并不是直接对密集的鸟瞰视图(BEV)特征进行推理,而是构建稀疏的对象级时空锚点,作为姿态校正的共享接口,并在一个统一的循环中紧密连接空间精细化、时间传播和当前时间验证,同时保持整体校正过程的轻量化。在模拟和真实世界数据集上的大量实验表明,CoAnchor在干净环境下仍然具有竞争力,并在联合延迟和姿态扰动下提高了鲁棒性,展现出良好的实际准确性与效率的权衡。
cs.CV / 60 / 2608.21066
Robust Validation to Geometric Perturbations for Autonomous Pose Estimation
针对几何扰动的鲁棒验证用于自主姿态估计
Abstract
Deploying autonomous systems in safety-critical domains demands guaranteed robustness against physically plausible geometric perturbations rather than abstract pixel-wise noise. In vision-based navigation and autonomous landing, machine learning components require rigorous validation under dynamic operational conditions such as camera rotations and lighting shifts. Extending findings on the failure of first-order spatial attacks in classification, we show that standard gradient-based heuristics (e.g. APGD) similarly fail on for pose estimation, often performing worse than a simple random sampling baseline. To overcome these optimization bottlenecks, we reformulate pose estimation robustness within the framework of Global Lipschitzian Optimization (GLO). We argue that GLO offers a principled approach to robust validation, effectively localizing global optima with strong theoretical convergence guarantees. We evaluate this framework on a YOLOv8-Pose keypoint detector with a Perspective-n-Point (PnP) solver against rotation and contrast. In our evaluations, GLO successfully isolates critical failure modes where position deviations exceed safe operational limits, while rapidly pruning the search space by over 80%. To the best of our knowledge, this is the first study to extend geometric robustness validation to continuous keypoint regression and deep object detection, establishing a practical step toward certifying robust autonomous perception.
Chinese Translation
在安全关键领域部署自主系统需要对物理上合理的几何扰动提供保证的鲁棒性,而不是抽象的像素级噪声。在基于视觉的导航和自主着陆中,机器学习组件需要在动态操作条件下(如相机旋转和光照变化)进行严格的验证。基于对分类中一阶空间攻击失败的研究,我们表明标准的基于梯度的启发式方法(如 APGD)在姿态估计中同样失败,表现往往不如简单的随机采样基线。为了克服这些优化瓶颈,我们在全局利普希茨优化(Global Lipschitzian Optimization, GLO)的框架内重新构建姿态估计的鲁棒性。我们认为 GLO 提供了一种原则性的方法来进行鲁棒验证,有效地定位全局最优解,并具有强大的理论收敛保证。我们在一个使用透视-n-点(Perspective-n-Point, PnP)求解器的 YOLOv8-Pose 关键点检测器上评估了该框架,针对旋转和对比度进行测试。在我们的评估中,GLO 成功地隔离了关键的失败模式,其中位置偏差超过安全操作限制,同时快速地将搜索空间缩减超过 80%。据我们所知,这是首次将几何鲁棒性验证扩展到连续关键点回归和深度目标检测,为认证鲁棒的自主感知迈出了实质性的一步。
cs.CV / 61 / 2608.21067
AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images
AT-ViT:一种针对区域的多视角视觉变换器,结合交叉注意力和多尺度分块,用于标本图像中的植物性状识别
Abstract
Automated plant traits recognition from herbarium images is essential for plant sciences, yet remains challenging because background elements (e.g., textual labels, mounting artifacts, and color charts) can introduce shortcut learning, leading models to rely on spurious non-plant cues rather than plant morphology. This bias degrades both generalization and interpretability. In this paper, we introduce AT-ViT, a dual-branch Vision Transformer that jointly encodes raw herbarium scans and their segmented-derived counterparts via a multi-scale, multi-view cross-attention fusion scheme. AT-ViT further incorporates a mask-guided patch weighting mechanism that amplifies plant-relevant regions and attenuates background-driven features. By learning from the original scans while being guided by segmentation masks through the mask-guided patch reweighting mechanism, the model is encouraged to focus on plant organs and learn plant-centric representations more effectively. Across multiple trait classification tasks (e.g., leaf base shape, thorns), AT-ViT delivers consistent accuracy gains, improves attention localization on plant regions, and exhibits increased robustness under synthetic background perturbations. Specifically, AT-ViT substantially improves spatial attention grounding, boosting plant-region alignment (Avg IoU_p: +15.66 to +18.03 pp) while reducing background overlap (Avg IoU_b: -27.92 to -31.02 pp) relative to CrossViT, and remains markedly more robust to background perturbations, outperforming ResNet101 by up to +32.32 accuracy points and CrossViT by up to +5.07 points under background-noise conditions.
Chinese Translation
从标本图像中自动识别植物性状对植物科学至关重要,但仍然具有挑战性,因为背景元素(例如文本标签、装配伪影和色卡)可能引入捷径学习,使模型依赖于虚假的非植物线索而非植物形态。这种偏差降低了模型的泛化能力和可解释性。本文介绍了AT-ViT,一种双分支视觉变换器,通过多尺度、多视角的交叉注意力融合方案,联合编码原始标本扫描图和其分割衍生图。AT-ViT进一步结合了一种基于掩膜的分块加权机制,放大与植物相关的区域,并减弱背景驱动的特征。通过从原始扫描图中学习,同时通过基于掩膜的分块重加权机制进行指导,模型被鼓励关注植物器官,更有效地学习以植物为中心的表示。在多个性状分类任务中(例如,叶基形状、刺),AT-ViT提供了一致的准确性提升,改善了对植物区域的注意力定位,并在合成背景扰动下表现出更强的鲁棒性。具体而言,AT-ViT显著改善了空间注意力的基础,提升了植物区域的对齐(平均IoU_p:+15.66至+18.03个百分点),同时减少了背景重叠(平均IoU_b:-27.92至-31.02个百分点),相较于CrossViT,且在背景扰动下表现出明显更强的鲁棒性,准确率比ResNet101高出最多+32.32个百分点,比CrossViT高出最多+5.07个百分点。
cs.CV / 62 / 2608.21093
Gaussian-Mixture Latent Flow for Stochastic 3D Human Motion Prediction
用于随机三维人类运动预测的高斯混合潜在流
Abstract
Stochastic human motion prediction aims to forecast future motion distributions. Although recent studies have achieved strong performance in terms of accuracy and diversity, they often overlook plausibility (e.g., resulting in physically unrealistic predictions) and uncertainty quantification, both of which are essential for real-world applications and downstream tasks. To address these issues, we propose a latent flow-based model equipped with a data-driven Gaussian mixture prior that more effectively disentangles diverse human behaviors than conventional single-modal priors. This prior is derived from patterns in the training data without requiring additional annotations. Furthermore, the fully invertible nature of our model enables natural uncertainty quantification through tractable likelihood computation. Experiments on the Human3.6M and AMASS datasets demonstrate that our approach achieves state-of-the-art performance in both accuracy and plausibility.
Chinese Translation
随机人类运动预测旨在预测未来运动分布。尽管近期研究在准确性和多样性方面取得了显著的成果,但它们往往忽视了合理性(例如,导致物理上不现实的预测)和不确定性量化,而这两者对于实际应用和下游任务都是至关重要的。为了解决这些问题,我们提出了一种基于潜在流的模型,该模型配备了数据驱动的高斯混合先验,能够比传统的单模态先验更有效地解耦多样的人类行为。该先验是从训练数据中的模式推导而来,无需额外的标注。此外,我们模型的完全可逆性使得通过可处理的似然计算实现自然的不确定性量化成为可能。在Human3.6M和AMASS数据集上的实验表明,我们的方法在准确性和合理性方面均达到了最先进的性能。
cs.CV / 63 / 2608.21098
When does fusing hand-crafted knowledge with learned representations pay? A cost-normalized benchmark of stacking, substitution, and interference
何时将手工知识与学习表示融合是有利的?堆叠、替代和干扰的成本标准化基准
Abstract
Fusing prior knowledge with data-driven learning is attractive where data is scarce, yet no controlled account says when it helps, is redundant, or harms. We benchmark one fixed hand-crafted knowledge source, a pinned bank of Gabor targets injected only during training at $\sim$2\% overhead, against data-driven alternatives (SimCLR, SimSiam, DINO, ImageNet transfer, augmentation, learned teachers) under one frozen recipe with fixed subsets: 13 datasets, 9 backbones, 150 to 1.28M images, 32--224\,px, 2.5M--86M parameters ($\computeCells$ classification configurations over $\computeRuns$ runs, plus segmentation and detection transplants). Across the training-time combinations we measure, three outcomes recur (decision-level fusion differs). Different-\emph{currency} sources can stack: the prior composes with DeiT augmentation on attention backbones and is worth $+26$ points to ViT-B/16 at $224$\,px, $+6.7$ at twice that budget. Same-currency sources substitute: against effective self-supervised pretraining, the combination never usefully exceeds the better single source. Fusing at full strength into an already-informed initialization interferes in proportion to what it carries: ImageNet transfer, $-15$ to $-17$ points, removed by a weaker auxiliary weight. Frozen-feature diagnostics measured on each source alone separate these outcomes retrospectively but do not predict them: a rule built on them calls one of nine unseen pairs. At a practitioner's own label budget, the frozen-feature gain predicts the end-to-end gain to within $0.17$ points across 30 cells and seven datasets; the underlying decomposition, $\Delta = G + \readout(\mathrm{base})$, holds in sign on $\auditRate\%$ of testable cells and is called an unseen backbone family's feature gain in advance. The project page is https://amughrabi.github.io/MomentAux.
Chinese Translation
在数据稀缺的情况下,将先验知识与数据驱动学习相结合是有吸引力的,但目前尚无控制性研究说明何时这种结合有帮助、冗余或有害。我们基准测试了一个固定的手工知识来源,即在训练期间以约2%的开销注入的固定Gabor目标库,与数据驱动的替代方法(SimCLR、SimSiam、DINO、ImageNet迁移、数据增强、学习教师)进行比较,采用一个冻结的配方和固定的子集:13个数据集、9个骨干网络、150到128万张图像、32-224像素、250万到8600万参数($ ext{computeCells}$ 分类配置在 $ ext{computeRuns}$ 次运行中,以及分割和检测移植)。在我们测量的训练时间组合中,三种结果反复出现(决策级融合有所不同)。不同- extit{currency} 来源可以堆叠:先验知识与DeiT数据增强在注意力骨干网络上组合,在224像素时对ViT-B/16的提升为+26分,在预算翻倍时提升为+6.7分。同一货币来源则进行替代:与有效的自监督预训练相比,组合的效果从未有效超过更好的单一来源。将全力融合到已经知情的初始化中会根据其携带的信息量产生干扰:ImageNet迁移,损失15到17分,通过较弱的辅助权重可以消除。对每个来源单独进行的冻结特征诊断可以追溯性地区分这些结果,但无法预测它们:基于这些结果建立的规则只能对九对未见的组合之一做出判断。在实践者自己的标签预算下,冻结特征的增益可以在30个单元和七个数据集上预测端到端增益,误差在0.17分以内;基础分解$ riangle = G + ext{readout}( ext{base})$在$ ext{auditRate}\%$的可测试单元上保持符号一致,并提前被称为未见骨干网络家族的特征增益。项目页面为 https://amughrabi.github.io/MomentAux。
cs.CV / 64 / 2608.21099
A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration
A2DINOv3:通过社会化协作重新思考多模态目标检测
Abstract
Multi-modal object detection is essential for robust scene understanding in challenging conditions, including low-light and adverse environments. Recent vision foundation models (e.g., DINOv3) have exhibited strong representation capabilities, yet adapting them to multi-modal scenarios remains challenging. Existing dense cross-modal fusion strategies often force heterogeneous modalities to interact indiscriminately, which may introduce redundant information and disrupt the valuable pre-trained representations. To address this issue, we revisit multi-modal fusion from the perspective of socialized learning and propose adapter to DINOv3 (A2DINOv3), a multi-expert collaboration framework with a Socialized Collaboration Protocol (SCP). Specifically, RGB and infrared branches are modeled as heterogeneous experts that independently preserve their specialized knowledge while exchanging complementary information through selective and constrained interactions. This design mitigates harmful cross-modal interference and prevents degradation of pre-trained priors during adaptation. Furthermore, a zero-initialization strategy is introduced to gradually activate cross-modal collaboration, enabling a smooth transition from modality-specific learning to cooperative representation learning. Extensive experiments on four multi-modal benchmarks, including aerial detection (GAIIC), autonomous driving (FLIR), low-light surveillance (LLVIP), and diverse real-world scenarios (M3FD), demonstrate that A2DINOv3 consistently achieves state-of-the-art performance in multi-modal object detection.
Chinese Translation
多模态目标检测对于在低光照和恶劣环境等挑战条件下的稳健场景理解至关重要。近期的视觉基础模型(例如,DINOv3)展现了强大的表征能力,但将其适应于多模态场景仍然面临挑战。现有的密集跨模态融合策略往往强迫异构模态进行无差别的交互,这可能引入冗余信息并破坏有价值的预训练表征。为了解决这一问题,我们从社会化学习的角度重新审视多模态融合,并提出了DINOv3的适配器(A2DINOv3),这是一个具有社会化协作协议(SCP)的多专家协作框架。具体而言,RGB和红外分支被建模为异构专家,它们独立保留各自的专业知识,同时通过选择性和受限的交互交换互补信息。这一设计减轻了有害的跨模态干扰,并防止在适应过程中预训练先验的退化。此外,引入了一种零初始化策略,以逐步激活跨模态协作,从而实现从特定模态学习到合作表征学习的平滑过渡。在四个多模态基准测试(包括空中检测(GAIIC)、自动驾驶(FLIR)、低光监控(LLVIP)和多样化的真实场景(M3FD))上的广泛实验表明,A2DINOv3在多模态目标检测中始终实现了最先进的性能。
cs.CV / 65 / 2608.21114
CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents
CIVA:对视觉世界模型代理的批评者诱导价值子空间攻击
Abstract
Visual world-model agents such as DreamerV3 act through a recurrent latent state rather than a single observation, which weakens frame-wise observation attacks and makes their perturbations vary sharply over time under a strict per-frame perturbation constraint. We study white-box, causal, online attacks on such agents and propose Critic-Induced Value-Subspace Attacks (\textbf{CIVA}). Our key observation is that, along a rollout, critic-guided perturbations concentrate in a low-dimensional subspace induced by the victim's own critic. Based on this observation, CIVA first probes the frozen victim offline with critic-guided PGD and extracts a low-rank value-subspace by SVD. At test time, it optimizes only the subspace coefficients, smooths them with an exponential moving average (EMA), and maps them back to pixels. This design attacks value-sensitive recurrent dynamics while keeping the online optimization cheap and temporally coherent. Extensive experiments on DMC walker walk, Atari Pong, and Crafter show that CIVA consistently outperforms five recent methods; on DMC walker walk, it achieves the largest reward drop of 26.07\% while keeping temporal variation low, with TempAbs of 0.646.
Chinese Translation
视觉世界模型代理如DreamerV3通过递归潜在状态而非单一观察进行操作,这削弱了逐帧观察攻击,并使其扰动在严格的逐帧扰动约束下随时间急剧变化。我们研究了对这些代理的白盒、因果、在线攻击,并提出了批评者诱导价值子空间攻击(CIVA)。我们的关键观察是,在一次展开过程中,批评者引导的扰动集中在受害者自身批评者诱导的低维子空间中。基于这一观察,CIVA首先通过批评者引导的投影梯度下降(PGD)离线探测冻结的受害者,并通过奇异值分解(SVD)提取低秩价值子空间。在测试时,它仅优化子空间系数,用指数移动平均(EMA)平滑它们,并将其映射回像素。这一设计攻击了对价值敏感的递归动态,同时保持在线优化的低成本和时间一致性。在DMC walker walk、Atari Pong和Crafter上的大量实验表明,CIVA始终优于五种最新方法;在DMC walker walk上,它实现了26.07%的最大奖励下降,同时保持低的时间变化,TempAbs为0.646。
cs.CV / 66 / 2608.21133
Masking Is Not Enough: Generative Restoration for Multimodal De-Identification in Medical AI
掩蔽不足:用于医疗人工智能的多模态去标识化的生成恢复
Abstract
Medical image-text data can expose protected health information (PHI) through both visible image content as well as accompanying text, creating a barrier to privacy-preserving medical AI systems. This risk is especially prominent in multimodal systems, where images, questions, reports, and clinical context may enter training, evaluation, or inference pipelines. Existing medical vision-language benchmarks primarily emphasize task utility, while de-identification methods are often evaluated separately from downstream reasoning. We introduce ClinX, an end-to-end multimodal PHI sanitization framework for medical image-text data. ClinX detects visible identifiers with optical character recognition (OCR), constructs binary PHI masks, and applies ClinX-PRISM, a no-skip generative restoration module with privacy-oriented post-processing for burned-in identifier suppression. In parallel, text-side PHI is reduced through progressive de-identification levels: regex masking, context-aware masking, and rewrite-based sanitization. We evaluate ClinX in medical visual question answering (MedVQA), jointly measuring PHI leakage and downstream utility across image-side, text-side, and combined de-identification settings. Results show that OCR-only masking is not sufficient as a standalone solution, and restoration-based sanitization better preserves clinically relevant visual context while sharply reducing recoverable PHI.
Chinese Translation
医学图像-文本数据通过可见的图像内容和伴随的文本可能暴露受保护的健康信息(PHI),这为隐私保护的医疗人工智能系统带来了障碍。这种风险在多模态系统中尤为突出,因为图像、问题、报告和临床背景可能进入训练、评估或推理流程。现有的医学视觉-语言基准主要强调任务效用,而去标识化方法往往与下游推理分开评估。我们提出了ClinX,一个用于医学图像-文本数据的端到端多模态PHI净化框架。ClinX通过光学字符识别(OCR)检测可见标识符,构建二进制PHI掩膜,并应用ClinX-PRISM,一个无跳过的生成恢复模块,具有隐私导向的后处理功能,用于抑制嵌入的标识符。同时,文本侧的PHI通过渐进式去标识化级别减少:正则表达式掩蔽、上下文感知掩蔽和基于重写的净化。我们在医学视觉问答(MedVQA)中评估ClinX,联合测量图像侧、文本侧和组合去标识化设置中的PHI泄漏和下游效用。结果表明,仅依靠OCR掩蔽作为独立解决方案是不够的,而基于恢复的净化更好地保留了临床相关的视觉上下文,同时显著减少了可恢复的PHI。
cs.CV / 67 / 2608.21134
Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
Llama-Mobile:高效的2.7位量化视觉语言模型
Abstract
Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8-bit activations, preserving strong performance on a set of standard visual question answering tasks.
Chinese Translation
在移动设备上部署视觉语言模型(VLMs)面临着显著的内存和计算需求挑战。我们提出了一种框架,用于在资源受限的硬件上对VLMs进行量化,以实现高效推理。我们的方法结合了一个量化管道,该管道利用模型本身生成训练数据,并且不需要访问训练设置,同时采用了一种新颖的每参数2.7位格式,以支持在Arm CPU上的高效执行。我们通过将Llama 3.2 11B Vision Instruct模型压缩到3.7 GB,并使用8位激活,验证了我们的方法,在一组标准视觉问答任务上保持了强劲的性能。
cs.CV / 68 / 2608.21136
Stream3Dv2: Geometric-Semantic Fusion Enhanced Streaming Zero-Shot 3D Scene Understanding
Stream3Dv2:几何-语义融合增强的流式零-shot 3D 场景理解
Abstract
Recently, open-vocabulary zero-shot 3D scene understanding using vision foundation models has emerged as a promising alternative to data-intensive supervised methods. However, deploying these models in real-world scenarios is severely hindered by their inability to efficiently handle streaming RGB-D inputs and their inherent vulnerability to noise 2D segmentation masks. To address these critical limitations, we propose Stream3Dv2, a novel training-free framework designed for robust streaming 3D perception. Stream3Dv2 processes sequential data through an original nested local-to-historical architecture, capturing multi-view consistency while circumventing the high computational overhead so as to support timely responses. At its core, we introduce a comprehensive geometric-semantic fusion mechanism that resolves geometric noise and semantic ambiguity by explicitly utilizing semantic guidance and formulating 3D segmentation as solving point-and-set merging and partitioning problems. Furthermore, we present an innovative manifold-distance-based point cloud refinement strategy. This approach leverages local manifold graphs for point-to-manifold optimization that mitigates the boundary delineation failures caused by Euclidean-distance metrics, and employs geometric bounding boxes to dynamically activate and update historical instances for achieving rapid manifold-to-manifold refinement. Extensive experiments on public datasets demonstrate that Stream3Dv2 consistently outperforms existing baselines in foundational open-vocabulary streaming 3D segmentation and detection. Finally, we show that integrating our framework with an LLM-based agent enables advanced language-driven 3D scene understanding, underscoring its potential for open-world embodied intelligence. Code will be updated at https://github.com/SubmissionsIn/Stream3D.
Chinese Translation
最近,基于视觉基础模型的开放词汇零-shot 3D 场景理解作为一种有前景的替代方案,逐渐取代了数据密集型的监督方法。然而,这些模型在实际场景中的应用受到严重限制,主要是由于它们无法有效处理流式 RGB-D 输入,并且对噪声的 2D 分割掩码本质上存在脆弱性。为了解决这些关键限制,我们提出了 Stream3Dv2,这是一种新颖的无训练框架,旨在实现稳健的流式 3D 感知。Stream3Dv2 通过一种原创的嵌套局部到历史架构处理序列数据,捕捉多视图一致性,同时避免高计算开销,以支持及时响应。其核心是我们引入了一种全面的几何-语义融合机制,通过明确利用语义指导并将 3D 分割公式化为解决点与集合的合并和划分问题,从而解决几何噪声和语义模糊。此外,我们提出了一种创新的基于流形距离的点云精炼策略。该方法利用局部流形图进行点到流形的优化,减轻了由欧几里得距离度量引起的边界划分失败,并采用几何边界框动态激活和更新历史实例,以实现快速的流形到流形精炼。在公共数据集上的大量实验表明,Stream3Dv2 在基础开放词汇流式 3D 分割和检测方面始终优于现有基线。最后,我们展示了将我们的框架与基于 LLM 的代理集成后,能够实现先进的语言驱动 3D 场景理解,强调了其在开放世界具身智能中的潜力。代码将更新至 https://github.com/SubmissionsIn/Stream3D。
cs.CV / 69 / 2608.21140
A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans
一种模块化代理用于CT扫描中可靠且可审计的空间关系验证
Abstract
Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological report generation and structured image understanding. While modern vision-language models (VLMs) show promising performance on many medical imaging tasks, recent evidence suggests they remain weak in controlled spatial reasoning and often fail to reliably ground spatial relations in image evidence. Given that radiological reasoning hinges on understanding the relative positions of anatomical structures and findings, this spatial weakness poses risks to diagnostic accuracy. We present a modular medical imaging agent for binary spatial relation verification in axial CT slices. Instead of directly predicting spatial answers end-to-end, the system decomposes the task into explicit stages: language parsing, anatomical localization, and deterministic geometric verification. Natural-language queries are converted into structured relation tuples, queried organs are localized with a YOLO-based detector, and the final spatial decision is computed from object centers using deterministic geometric rules. We evaluate the approach on the held-out MIRP spatial QA benchmark and compare it against representative end-to-end VLM baselines. The best-performing hybrid configuration reaches 94.1% accuracy and 94.2% F1, outperforming direct Qwen2-VL prompting by 42.5 percentage points in accuracy, while preserving interpretable intermediate representations and auditable reasoning stages. The results suggest that explicit modular spatial verification can serve as a promising building block for future report-oriented medical imaging agents.
Chinese Translation
可靠的空间理解是未来医学视觉-语言系统的重要前提,这些系统旨在支持放射学报告生成和结构化图像理解。尽管现代视觉-语言模型(VLMs)在许多医学影像任务上表现出色,但最近的证据表明,它们在受控空间推理方面仍然较弱,且常常无法可靠地将空间关系与图像证据相结合。考虑到放射学推理依赖于理解解剖结构和发现的相对位置,这种空间弱点对诊断准确性构成风险。我们提出了一种模块化医学影像代理,用于轴向CT切片中的二元空间关系验证。该系统并不是直接进行端到端的空间答案预测,而是将任务分解为明确的阶段:语言解析、解剖定位和确定性几何验证。自然语言查询被转换为结构化关系元组,查询的器官通过基于YOLO的检测器进行定位,最终的空间决策则通过确定性几何规则从物体中心计算得出。我们在保留的MIRP空间QA基准上评估该方法,并将其与代表性的端到端VLM基线进行比较。表现最佳的混合配置达到了94.1%的准确率和94.2%的F1分数,在准确率上比直接的Qwen2-VL提示高出42.5个百分点,同时保持了可解释的中间表示和可审计的推理阶段。结果表明,明确的模块化空间验证可以作为未来面向报告的医学影像代理的有希望的构建块。
cs.CV / 70 / 2608.21160
Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates
Human-JEPA:一种以人为中心的视觉模型,能够感知与预测
Abstract
Machines that understand humans should perceive the present and anticipate the future. Existing human-centric vision model are pretrained on human images, set the state of the art in static dense perception, so motion and anticipation are out of reach. Here we present Human-JEPA, a human-centric vision model trained on video by anchored forecasting: dense targets are pinned to a frozen copy of the initialization, preventing a silent collapse of dense perception, and block masks are replaced by a pure past-to-future split, avoiding a five-point action tax and a seventeen-point re-identification collapse. Under frozen probes, Human-JEPA leads the pixel-anchored specialists on pose and person re-identification at 2.7 times fewer parameters, conceding high-resolution dense parsing, and its released predictor head is the first that does not degrade anticipation. A single safely adapted model thus serves both halves of understanding humans.
Chinese Translation
理解人类的机器应当能够感知现在并预测未来。现有的以人为中心的视觉模型是在人体图像上进行预训练的,在静态密集感知方面设定了最先进的水平,但在运动和预测方面仍然无法实现。在此,我们提出了Human-JEPA,这是一种通过锚定预测在视频上训练的以人为中心的视觉模型:密集目标被固定在初始化的静态副本上,防止了密集感知的无声崩溃,同时块掩码被替换为纯粹的过去到未来的分割,避免了五点动作税和十七点重新识别崩溃。在冻结探针下,Human-JEPA在姿态和人物重新识别方面以2.7倍更少的参数超越了像素锚定专家,承认了高分辨率的密集解析,其发布的预测头是首个不降低预测能力的模型。因此,一个安全适应的模型同时服务于理解人类的两个方面。
cs.CV / 71 / 2608.21170
Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds
视觉提示是你所需的一切吗?在渐进式视觉支架下研究视觉语言模型的空间推理
Abstract
Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and downstream reasoning. What remains less clear is how the visual presentation of a task shapes model performance and failure modes when the underlying reasoning problem is unchanged. We study this question in SPaRC, a benchmark for grid-based visual spatial planning, by introducing lightweight input-side scaffolds that preserve the visual modality while making spatial structure more accessible. Across multiple VLMs, these scaffolds improve task accuracy over the original visual setting by up to 34.0 percentage points and further complement GRPO-based training, yielding up to 4.6 additional accuracy points compared with near-zero gains on the original visual input. Analyses on both end-to-end task solving and object detection show that these gains are closely tied to reductions in grounding-related errors, while rule reasoning remains comparatively challenging. We find that visual presentation is a central factor that determines whether VLM benchmarks measure grounded perception, downstream reasoning, or a mixture of both.
Chinese Translation
视觉语言模型(VLMs)在多模态推理方面迅速发展,但最近的研究表明,它们的失败往往反映了视觉基础与下游推理之间的相互作用。尚不清楚的是,当基础推理问题不变时,任务的视觉呈现如何影响模型的表现和失败模式。我们在SPaRC这一基于网格的视觉空间规划基准中研究了这个问题,通过引入轻量级的输入侧支架,保留视觉模态的同时使空间结构更加易于访问。在多个VLM中,这些支架使任务准确率比原始视觉设置提高了多达34.0个百分点,并进一步补充了基于GRPO的训练,与原始视觉输入几乎没有增益相比,额外提高了最多4.6个百分点的准确率。对端到端任务解决和物体检测的分析表明,这些增益与基础相关错误的减少密切相关,而规则推理仍然相对具有挑战性。我们发现,视觉呈现是决定VLM基准是否测量基础感知、下游推理或两者混合的一个核心因素。
cs.CV / 72 / 2608.21189
Towards Investigating Residual Hearing Loss: Quantification of Fibrosis in a Novel Cochlear OCT Dataset
探讨残余听力损失:在新型耳蜗光学相干断层扫描数据集中定量纤维化
Abstract
Objective: Cochlear implants (CIs) are bionic prostheses that restores hearing via electrical stimulation of the auditory nerve. Hybrid CIs, which use electroacoustic stimulation (EAS), combine residual low-frequency acoustic hearing with CI electrical stimulation. Intracochlear fibrosis, which forms in response to the presence of the implant, may impede residual hearing function and gradually reduce the efficacy of EAS. It is therefore a translational objective to study the formation of cochlear fibrosis in rodents, with the goal of reducing fibrotic burden and improving outcomes for CI patients. Methods: We generate and annotate a novel dataset of optical coherence tomography (OCT) images from chronically implanted guinea pigs as part of an ongoing study focused on implant induced fibrosis. Objectively assessing fibrotic burden in this model, with high resolution and repeatability, presents an obvious use case for computer vision methods. Results: We present the results of several state-of-the-art semantic segmentation models and compare their efficacy for identifying cochlear fibrosis and other relevant annotations, using a new library of manually segmented OCT images. Conclusions: We find that the best performance is achieved by using a modified version of the well-known UNET architecture (which we term 2D-OCT-UNET) that operates on the upscaled OCT input resolution. Significance: For the first time, we have successfully applied computer vision techniques to an OCT dataset of implanted cochleae with fibrosis. Using this deep learning model, the cochlear fibrotic burden calculation can be reliably carried out as we verify in our experimental section. The dataset and the project code are available at: https://github.com/juliadietlmeier/CF-OCT-segmentation
Chinese Translation
目的:耳蜗植入物(CIs)是一种通过电刺激听神经来恢复听力的仿生假体。混合耳蜗植入物(Hybrid CIs)使用电声刺激(EAS),将残余低频声学听力与耳蜗植入物的电刺激相结合。耳蜗内纤维化是对植入物存在的反应,可能会妨碍残余听力功能,并逐渐降低EAS的有效性。因此,研究啮齿动物中耳蜗纤维化的形成是一个转化医学目标,旨在减少纤维化负担并改善耳蜗植入患者的结果。方法:我们生成并注释了一组来自慢性植入豚鼠的光学相干断层扫描(OCT)图像的新数据集,作为一项专注于植入物诱导纤维化的持续研究的一部分。以高分辨率和重复性客观评估该模型中的纤维化负担,为计算机视觉方法提供了明显的应用场景。结果:我们展示了几种最先进的语义分割模型的结果,并使用一组手动分割的OCT图像库比较它们在识别耳蜗纤维化和其他相关注释方面的有效性。结论:我们发现,通过使用经过修改的著名UNET架构(我们称之为2D-OCT-UNET),在上采样的OCT输入分辨率上实现了最佳性能。意义:这是我们首次成功将计算机视觉技术应用于具有纤维化的植入耳蜗的OCT数据集。使用该深度学习模型,耳蜗纤维化负担的计算可以可靠地进行,我们在实验部分进行了验证。数据集和项目代码可在以下网址获取:https://github.com/juliadietlmeier/CF-OCT-segmentation
cs.CV / 73 / 2608.21194
ES-VP : Energy-Shaped Dynamic Visual Prompting for Efficient Model Adaptation
ES-VP:用于高效模型适应的能量形状动态视觉提示
Abstract
Visual prompting (VP) has emerged as a parameter-efficient method for adapting pre-trained models to downstream tasks. However, existing approaches encounter a trade-off between flexibility and efficiency. Some methods apply a fixed prompt to all images, ignoring individual image characteristics, while others introduce auxiliary networks to generate diverse prompts. Although the latter can improve performance, it also significantly increases parameter usage and the potential for overfitting to specific datasets. Furthermore, the auxiliary networks, combined with inherent biases in pre-trained models, limit scalability and generalization. In this paper, we propose Energy-Shaped Visual Prompting (ES-VP), a novel approach that generates image-specific prompts using low-rank initialization and energy-guided dynamic adaptation, achieving superior performance with fewer parameters compared to single-prompt methods. ES-VP directly utilizes the pre-trained model for adaptive prompt generation, ensuring both parameter efficiency and improved generalization. Extensive experiments conducted on five architectures across fifteen datasets demonstrate that ES-VP consistently outperforms current state-of-the-art (SOTA) single and diverse VP methods. For instance, using the CLIP architecture across four datasets, ES-VP outperforms the SOTA method DAM-VP by an average of 2.6\% in accuracy while utilizing 590$\times$ fewer VP parameters, thereby establishing a new benchmark for efficient and generalizable model adaptation.
Chinese Translation
视觉提示(VP)已成为一种参数高效的方法,用于将预训练模型适应于下游任务。然而,现有方法在灵活性和效率之间存在权衡。一些方法对所有图像应用固定提示,忽略了个别图像特征,而另一些方法则引入辅助网络以生成多样化的提示。尽管后者可以提高性能,但也显著增加了参数使用量,并可能导致对特定数据集的过拟合。此外,辅助网络与预训练模型固有的偏差相结合,限制了可扩展性和泛化能力。在本文中,我们提出了能量形状视觉提示(ES-VP),这是一种新颖的方法,通过低秩初始化和能量引导的动态适应生成特定于图像的提示,与单一提示方法相比,能够以更少的参数实现更优的性能。ES-VP直接利用预训练模型进行自适应提示生成,确保了参数效率和改进的泛化能力。在五种架构和十五个数据集上进行的广泛实验表明,ES-VP始终优于当前的最先进(SOTA)单一和多样化VP方法。例如,在四个数据集上使用CLIP架构时,ES-VP的准确率平均比SOTA方法DAM-VP高出2.6\%,同时使用的VP参数减少了590倍,从而为高效和可泛化的模型适应建立了新的基准。
cs.CV / 74 / 2608.21229
Anchoring Instruction Outside Mask: Exact Reference Caching for Efficient In-Context Diffusion Transformers
超越掩码的锚定指令:高效上下文扩散变换器的精确引用缓存
Abstract
Omnimodal generation is central to a wide range of content creation and editing applications. In-context conditioning is essential to this paradigm. It allows diffusion transformers to process text instructions and visual references in a shared attention sequence. However, each reference image introduces thousands of tokens. Computation therefore grows rapidly with the number of references. Existing methods reduce computation through structured sparse attention, which limits interactions between reference and target tokens. This structure also makes the reference K and V independent of the denoising target, allowing them to be computed once and reused across steps. However, it blocks visual references from attending to the text instruction. This substantially degrades instruction following and reference fidelity in multi-reference editing. To resolve this conflict, we jointly redesign the token sequence and attention mask. Our beyond-mask design uses static text anchors to connect the instruction to the reference branch. It preserves exact K and V reuse without adding parameters. However, this direct architectural conversion degrades generation quality. We recover the lost performance through teacher-forced velocity distillation, followed by a short on-policy stage in which the teacher supervises student-visited states. To our knowledge, this is the first use of on-policy distillation for architectural recovery in diffusion models. Across three image-editing benchmarks, our method matches full-attention generation quality. With five reference images, it accelerates the complete 40-step denoising process by 3.92x, while static text anchors introduce negligible runtime overhead; the speedup reaches 5.47x at ten references in our scaling study.
Chinese Translation
全模态生成是广泛内容创作和编辑应用的核心。在这一范式中,上下文条件至关重要。它允许扩散变换器在共享注意力序列中处理文本指令和视觉参考。然而,每个参考图像引入数千个标记,因此随着参考数量的增加,计算量迅速增长。现有方法通过结构化稀疏注意力来减少计算,这限制了参考标记与目标标记之间的交互。这种结构还使得引用的 K 和 V 与去噪目标独立,从而可以一次计算并在多个步骤中重用。然而,这阻止了视觉参考对文本指令的关注。这在多参考编辑中显著降低了指令跟随和参考保真度。为了解决这一冲突,我们共同重新设计了标记序列和注意力掩码。我们的超越掩码设计使用静态文本锚点将指令与参考分支连接。它在不增加参数的情况下保留了精确的 K 和 V 重用。然而,这种直接的架构转换降低了生成质量。我们通过教师强制速度蒸馏恢复了损失的性能,随后进行短期的在线策略阶段,在该阶段中教师监督学生访问的状态。据我们所知,这是在扩散模型中首次使用在线蒸馏进行架构恢复。在三个图像编辑基准测试中,我们的方法达到了全注意力生成质量。在使用五个参考图像时,它将完整的 40 步去噪过程加速了 3.92 倍,而静态文本锚点引入的运行时开销微乎其微;在我们的扩展研究中,速度提升在十个参考时达到了 5.47 倍。
cs.CV / 75 / 2608.21244
A VLM Answer Is Not an Anomaly Score: Rank Compression in Training-Free Video Anomaly Detection
VLM 答案并非异常分数:无训练视频异常检测中的排名压缩
Abstract
Vision-language models enable training-free video anomaly detection by answering questions about video segments. VAD benchmarks, however, require a scalar anomaly score for each segment and evaluate the resulting ranking using the AUROC or AP. A VLM-based detector should therefore define an answer interface: the answer scale specifies the admissible answers, and the readout rule maps the model's output distribution to a score. Because this interface can change the evaluated ranking, it is part of the detector rather than a formatting detail. The generated readout uses only the most likely answer, whereas the probability readout uses the full distribution over admissible answers. Across four 7-8B VLMs, the probability readout outperforms the generated readout for every tested combination of answer scale, benchmark, and metric, with average gains ranging from 5 to 13 points across the four benchmark-metric pairs. The gap arises because the generated readout keeps only one answer value per segment, so segment with different answer distributions can receive the same score and lose their relative order. We call this loss of relative order generated-answer rank compression. Even when the answer scale allows 91 answers, the generated readout produces only 4-18 distinct scores, whereas the probability readout retains substantially finer score resolution. The advantage persists under every decoding strategy, prompt wording, and joint scoring-explanation prompt we test. The answer interface is therefore a consequential component of VLM-based VAD and should be explicitly specified and evaluated.
Chinese Translation
视觉语言模型通过回答关于视频片段的问题,实现了无训练的视频异常检测。然而,视频异常检测(VAD)基准要求为每个片段提供一个标量异常分数,并使用 AUROC 或 AP 评估结果排名。因此,基于 VLM 的检测器应定义一个答案接口:答案尺度指定可接受的答案,而读出规则将模型的输出分布映射到一个分数。由于这个接口可以改变评估的排名,它是检测器的一部分,而不是格式细节。生成的读出仅使用最可能的答案,而概率读出则使用可接受答案的完整分布。在四个 7-8B 的 VLM 中,概率读出在每个测试的答案尺度、基准和指标组合中均优于生成的读出,平均增益在四个基准-指标对中范围为 5 到 13 分。这个差距产生的原因是生成的读出每个片段仅保留一个答案值,因此具有不同答案分布的片段可能会获得相同的分数,从而失去相对顺序。我们称这种相对顺序的丧失为生成答案的排名压缩。即使答案尺度允许 91 个答案,生成的读出也仅产生 4-18 个不同的分数,而概率读出则保持了更精细的分数分辨率。这一优势在我们测试的每种解码策略、提示措辞和联合评分-解释提示下均持续存在。因此,答案接口是基于 VLM 的 VAD 的一个重要组成部分,应明确指定和评估。
cs.CV / 76 / 2608.21247
Just Noticeable Difference Modeling for Token Compression in Vision-Language-Action Models
视觉-语言-动作模型中令可察觉差异建模的令牌压缩
Abstract
Token compression has become a key technique for reducing the inference cost of large foundation models, with approaches such as token pruning and KV-cache reuse widely adopted in vision-language models and recently explored for embodied agents. In embodied agents, tokens not only support perception and semantic understanding but also directly affect latency-sensitive closed-loop robot action prediction. Existing schemes typically guide compression using redundancy or importance cues, such as visual similarity, attention scores, and saliency. However, these cues only indirectly measure the key factor for safe compression: how much a token can change before causing an unacceptable deviation in downstream actions. This receiver-dependent tolerance is closely related to the principle of just noticeable difference (JND). Classical JND characterizes signal tolerance in the human visual system, while machine-oriented JND extends this concept to downstream machine responses. Building on this progression, we introduce Action-JND, which extends JND modeling to embodied perception by defining noticeability through the language-conditioned action response of a vision-language-action (VLA) policy in closed-loop control. A token change is considered admissible only when the induced action deviation remains within a tolerated margin. To realize this concept, we develop a lightweight token-wise JND estimator in deep visual-feature space to predict the maximum tolerable perturbation while preserving policy responses. The resulting action-tolerance score serves as a plug-and-play criterion for VLA compression paradigms, including stale-KV reuse and token pruning, prioritizing action-tolerant tokens for compression. Experiments on the LIBERO benchmark with OpenVLA and OpenVLA-OFT demonstrate that Action-JND consistently improves compression reliability, especially under aggressive compression ratios.
Chinese Translation
令牌压缩已成为降低大型基础模型推理成本的关键技术,诸如令牌剪枝和KV缓存重用等方法在视觉-语言模型中被广泛采用,并且最近在具身智能体中进行了探索。在具身智能体中,令牌不仅支持感知和语义理解,还直接影响延迟敏感的闭环机器人动作预测。现有方案通常利用冗余或重要性线索来指导压缩,例如视觉相似性、注意力得分和显著性。然而,这些线索仅间接衡量安全压缩的关键因素:一个令牌在导致下游动作产生不可接受偏差之前能够变化多少。这种接收者依赖的容忍度与可察觉差异(Just Noticeable Difference, JND)原理密切相关。经典的JND描述了人类视觉系统中的信号容忍度,而面向机器的JND则将这一概念扩展到下游机器响应。基于这一进展,我们引入了Action-JND,它通过定义视觉-语言-动作(Vision-Language-Action, VLA)策略在闭环控制中的语言条件动作响应,扩展了JND建模到具身感知。仅当引起的动作偏差保持在可容忍的范围内时,令牌变化才被视为可接受。为了实现这一概念,我们在深度视觉特征空间中开发了一种轻量级的令牌级JND估计器,以预测在保持策略响应的同时最大可容忍的扰动。由此产生的动作容忍度评分作为VLA压缩范式的即插即用标准,包括过时的KV重用和令牌剪枝,优先考虑对动作容忍的令牌进行压缩。在LIBERO基准上与OpenVLA和OpenVLA-OFT的实验表明,Action-JND在压缩可靠性方面始终有所改善,尤其是在激进的压缩比下。
cs.CV / 77 / 2608.21254
On the Transferability of Agricultural Weed Detection Under Cross-Field Distribution Shift
跨领域分布转移下农业杂草检测的可转移性研究
Abstract
Accurate agricultural weed detection in real-world field conditions is essential for precision agriculture, enabling targeted intervention and reducing yield loss. Recent work has reported strong detection performance from UAV-based imagery across a range of crops, yet existing approaches evaluate within a single crop and field, leaving practitioners with little evidence that a model trained on one crop will generalize to a new field or crop type. In this work, we characterize where cross-dataset weed-localization performance degrades and which modeling choices recover it, reducing the need to relabel every new deployment field. We introduce a newly collected and annotated UAV image dataset for agricultural weed detection in cotton fields and use it alongside an existing soybean dataset collected under a similar protocol. Using these datasets, we evaluate the performance of several strategies for transferring a detector trained on one crop to another, comparing unsupervised domain adaptive object detection (DAOD) against pretraining on a domain-adjacent source dataset followed by few-shot fine-tuning on the target dataset. Our analysis spans target-domain label budgets from zero to the full target dataset, characterizing the trade-off between adaptation strategy and annotation effort. We find that few-shot fine-tuning with as few as 25 labeled target examples outperforms unsupervised DAOD in our cross-crop comparison, suggesting that source domain selection combined with modest target supervision is more productive than algorithmic sophistication in adaptation.
Chinese Translation
在真实的田间条件下,准确的农业杂草检测对于精准农业至关重要,它能够实现针对性干预并减少产量损失。近期的研究报告显示,基于无人机(UAV)影像的检测性能在多种作物中表现良好,但现有方法主要在单一作物和田地内进行评估,缺乏证据表明在一种作物上训练的模型能够推广到新的田地或作物类型。在本研究中,我们分析了跨数据集杂草定位性能下降的情况以及哪些建模选择能够恢复该性能,从而减少对每个新部署田地重新标注的需求。我们引入了一个新收集和注释的无人机影像数据集,专门用于棉花田的农业杂草检测,并将其与在类似协议下收集的大豆数据集结合使用。利用这些数据集,我们评估了将针对一种作物训练的检测器转移到另一种作物的几种策略的性能,比较了无监督领域自适应目标检测(DAOD)与在领域相邻的源数据集上进行预训练后在目标数据集上进行少量样本微调的效果。我们的分析涵盖了从零到完整目标数据集的目标领域标签预算,表征了适应策略与注释工作之间的权衡。我们发现,在我们的跨作物比较中,仅用25个标注的目标样本进行的少量样本微调的性能优于无监督的DAOD,这表明源领域选择结合适度的目标监督在适应性方面比算法复杂性更具成效。
cs.CV / 78 / 2608.21281
WildFin: An In-the-Wild Dataset for Fish Behavioral Recognition
WildFin:一个用于鱼类行为识别的野外数据集
Abstract
Recent advances in field technology have led to a massive influx of in-the-wild video data for ecological science. The primary bottleneck in leveraging this data is the high cost of expert annotation. While computer vision offers a potential solution, current models frequently fail when deployed in complex marine environments. To characterize these failures, we introduce WildFin, a novel benchmark for fish behavior recognition collected and annotated by ecologists.WildFin spans two critical real-world paradigms: stationary cameras monitoring groups of fish and dynamic divers following individual subjects. The dataset represents a massive curation effort, involving 1,350 hours of fieldwork and 600 hours of expert annotation to produce 9 hours of behavioral data with over 2 million frame-by-frame labels. We benchmark modern vision foundation models and quantify tradeoffs between static and spatiotemporal architectures, revealing the substantial gap that remains between current model capabilities and the demands of real-world underwater behavioral analysis. Project website: https://team-wildfin.github.io/.
Chinese Translation
近年来,现场技术的进步导致生态科学领域涌现出大量野外视频数据。然而,利用这些数据的主要瓶颈在于专家标注的高成本。尽管计算机视觉提供了潜在的解决方案,但当前模型在复杂的海洋环境中部署时常常失败。为了描述这些失败,我们引入了WildFin,一个由生态学家收集和标注的鱼类行为识别新基准。WildFin涵盖了两个关键的现实世界范式:监控鱼群的静态相机和跟随个体对象的动态潜水员。该数据集代表了一项巨大的整理工作,涉及1,350小时的实地工作和600小时的专家标注,生成了9小时的行为数据,包含超过200万个逐帧标签。我们对现代视觉基础模型进行了基准测试,并量化了静态和时空架构之间的权衡,揭示了当前模型能力与现实世界水下行为分析需求之间仍然存在的巨大差距。项目网站:https://team-wildfin.github.io/
cs.CV / 79 / 2608.21286
Difficulty-Calibrated Interpolation Paths for Conditional Flow Matching
针对条件流匹配的难度校准插值路径
Abstract
Conditional Flow Matching trains generative models by regressing a network onto the velocity of a prescribed noise-to-data interpolation path. The interpolation schedule that shapes this path is known to affect convergence and sample quality, yet it is invariably fixed in advance, independent of both the data and the model. We show that the regression difficulty of Conditional Flow Matching varies systematically along the path, and we propose Difficulty-Calibrated Flow Matching, which derives the schedule from the model itself: a short pilot run with the linear path records the per-time loss, and the schedule is set to the quantile function of this difficulty profile, so the trajectory lingers where the velocity is hardest to learn. The method has a single hyperparameter, leaves the training objective and its gradient equivalence intact, composes with classifier-free guidance, and adds about two percent training overhead. In controlled experiments on CIFAR-10, MNIST, and Fashion-MNIST with an identical compact U-Net, the calibrated path attains the best FID on CIFAR-10 at full sampling budget and clearly outperforms all fixed schedules in the large-batch, few-update regime, precisely the setting where compute is scarcest.
Chinese Translation
条件流匹配通过将网络回归到预定的噪声到数据插值路径的速度来训练生成模型。已知塑造此路径的插值调度会影响收敛性和样本质量,但它总是提前固定,与数据和模型无关。我们表明,条件流匹配的回归难度沿路径系统性变化,并提出了难度校准流匹配,该方法从模型自身推导调度:通过线性路径进行短暂的试运行记录每个时间点的损失,并将调度设置为该难度曲线的分位数函数,从而使轨迹停留在学习速度最困难的地方。该方法只有一个超参数,保持训练目标及其梯度等价性,与无分类器引导组合,并增加约2%的训练开销。在对CIFAR-10、MNIST和Fashion-MNIST的受控实验中,使用相同的紧凑型U-Net,校准路径在完整采样预算下在CIFAR-10上获得了最佳的FID,并在大批量、少更新的情况下明显优于所有固定调度,正是计算资源最稀缺的设置。
cs.CV / 80 / 2608.21300
When Adaptation Hurts: Connecting Representational Drift to OOD Failures in MedSAM Fine-Tuning
适应性带来的伤害:将表征漂移与MedSAM微调中的OOD失败联系起来
Abstract
Foundation models for medical image segmentation, like prompt-based MedSAM, generalize well across domains and modalities, often in zero or few-shot setups. However, their performance depends on the quality of prompts and the adaptation of the models to custom datasets. This work systematically examines how MedSAM generalizes across diverse medical imaging benchmarks, with six adaptation strategies: full-model and encoder-only LoRA, shallow and deep visual prompt tuning (VPT), and decoder-only and full fine-tuning. Models are trained on the International Skin Imaging Collaboration Challenge (ISIC 2018) dataset and evaluated under clean and increasingly noisy prompts on IN and Out-of-Distribution (OOD) datasets: close-OOD PH2 (dermoscopy), far-OOD BUSI (Breast Ultrasound Images Dataset) and CBIS-DDSM (Curated Breast Imaging Subset of the Digital Database for Screening Mammography). We show that adaptation improves performance on IN and close-OOD data but often reduces performance on far-OOD data. Full fine-tuning provides the best tradeoff, while encoder-only LoRA is the strongest parameter-efficient alternative, outperforming standard LoRA and VPT under far-OOD shifts. Using Centered Kernel Alignment (CKA), we show that far-OOD degradation is strongly associated with drift in decoder representations, whereas encoder similarity alone does not explain robustness. This suggests encoder-only LoRA provides stronger robustness than standard LoRA by adapting the encoder to distribution shift in visual features, while preserving the decoder pathway. We further show that random 0-100 pixel jitter on prompts produces more robust and better performing models. We thus conclude that robust MedSAM adaptation requires the combined consideration of prompt noise exposure, domain shift, and representation preservation. We release our code: https://github.com/ImSounic/medsam-vpt
Chinese Translation
用于医学图像分割的基础模型,如基于提示的MedSAM,在不同领域和模态中具有良好的泛化能力,通常在零样本或少样本设置下表现出色。然而,它们的性能依赖于提示的质量以及模型对自定义数据集的适应能力。本研究系统地考察了MedSAM在多样化医学影像基准上的泛化能力,采用六种适应策略:全模型和仅编码器的LoRA、浅层和深层视觉提示调优(VPT)、以及仅解码器和全微调。模型在国际皮肤影像合作挑战赛(ISIC 2018)数据集上训练,并在干净和逐渐噪声增加的提示下评估在IN和分布外(OOD)数据集上的表现:近OOD的PH2(皮肤镜图像)、远OOD的BUSI(乳腺超声图像数据集)和CBIS-DDSM(数字乳腺X线筛查数据库的策划乳腺影像子集)。我们发现适应性提高了IN和近OOD数据的性能,但通常会降低远OOD数据的性能。全微调提供了最佳的权衡,而仅编码器的LoRA是最强的参数高效替代方案,在远OOD转变下优于标准LoRA和VPT。通过使用中心核对齐(CKA),我们表明远OOD性能下降与解码器表征的漂移密切相关,而仅编码器的相似性并不能解释鲁棒性。这表明,仅编码器的LoRA通过将编码器适应于视觉特征的分布转变,同时保持解码器路径,提供了比标准LoRA更强的鲁棒性。我们进一步表明,在提示上进行随机0-100像素的抖动可以产生更鲁棒且性能更好的模型。因此,我们得出结论,鲁棒的MedSAM适应需要综合考虑提示噪声暴露、领域转变和表征保持。我们发布了我们的代码:https://github.com/ImSounic/medsam-vpt
cs.CV / 81 / 2608.21305
Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning
Re$^3$Cap:基于检索引导的图像描述增强的精炼方法通过强化学习
Abstract
Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encouraging Large Vision-Language Models (LVLMs) to explore novel reasoning strategies. This limitation leads to a performance gap between RL and Supervised Fine-Tuning (SFT). In this paper, we argue that multi-modal retrieval can serve as an effective reasoning signal for caption refinement. Based on this insight, we present the Retrieval-Guided Refinement for Image Captioning (Re$^3$Cap), a retrieval-guided reasoning strategy that enhances image captioning without requiring additional annotations. Instantiated by Caption Refinement Suggester (CRS) and Caption Quality Assessor (CQA), this strategy identifies hallucinations and omissions in image captions, leading to more accurate and detailed descriptions. Extensive experiments demonstrate the superiority of our method in image captioning, even compared with Supervised Fine-Tuning. Especially, Re$^3$Cap outperforms GRPO with an average improvement of 8.64% in relation reasoning on the COCO-LN500 benchmark.
Chinese Translation
强化学习(RL)在图像描述方面已显示出显著的提升,但仍然存在鼓励大型视觉语言模型(LVLMs)探索新颖推理策略的局限性。这一局限性导致了强化学习与监督微调(SFT)之间的性能差距。本文认为,多模态检索可以作为图像描述精炼的有效推理信号。基于这一见解,我们提出了基于检索引导的图像描述精炼方法(Re$^3$Cap),这是一种无需额外注释即可增强图像描述的检索引导推理策略。该策略通过描述精炼建议器(Caption Refinement Suggester, CRS)和描述质量评估器(Caption Quality Assessor, CQA)来实现,能够识别图像描述中的幻觉和遗漏,从而生成更准确和详细的描述。大量实验表明,我们的方法在图像描述方面优于现有技术,甚至超过了监督微调。特别是,Re$^3$Cap在COCO-LN500基准测试中,在关系推理方面平均提高了8.64%,超越了GRPO。
cs.CV / 82 / 2608.21360
OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs
OmniAssistBench:面向全方位大语言模型的助手式交互基准
Abstract
Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific goals. Unlike traditional passive video understanding, interactive assistants should actively combine visual states, user goals, and prior knowledge to provide effective help. Evaluating this is rather challenging, as the model's unpredictable response dynamically changes the user's subsequent actions, which static offline datasets cannot accommodate. To address this bottleneck, we introduce OmniAssistBench. To solve the issue of diverging interaction paths where the same user goal can be achieved through various methods, we provide models with predefined priors derived from the source video, requiring them to guide users along the exact same routes. Since real interaction videos are rare, we construct the dataset by reverse-engineering existing Internet videos. We deduce logical user goals and segment the videos into multi-turn clips to simulate continuous interactions. This rigorous pipeline required over 1000 expert person-hours to build the dataset. Results show that the proprietary Gemini-3-Pro reaches 66.4 out of the max point of 100, while the open-source Qwen3-Omni-Instruct achieves 51.2. Although current models generally understand user inputs, they frequently provide incorrect or incomplete answers. Specifically, they struggle with visual prompts (e.g., hand gestures), fail to maintain historical context during multi-turn interactions, and fail to delay response until the target event. Results indicate substantial room for improvement before models can become reliable assistants.
Chinese Translation
近期的全模态大语言模型(Omni-LLMs)在实时视频助手方面展现出巨大潜力,能够持续感知环境并引导用户实现特定目标。与传统的被动视频理解不同,交互式助手应主动结合视觉状态、用户目标和先前知识,以提供有效的帮助。评估这一点相当具有挑战性,因为模型不可预测的响应会动态改变用户后续的行动,而静态的离线数据集无法满足这一需求。为了解决这一瓶颈,我们提出了OmniAssistBench。为了解决相同用户目标可以通过多种方法实现的交互路径分歧问题,我们为模型提供了源视频中派生的预定义先验,要求它们沿着完全相同的路径引导用户。由于真实交互视频稀缺,我们通过逆向工程现有的互联网视频构建数据集。我们推导出逻辑用户目标,并将视频分割成多轮剪辑,以模拟连续交互。这个严格的流程耗费了超过1000小时的专家人力来构建数据集。结果显示,专有的Gemini-3-Pro模型获得了66.4分(满分100),而开源的Qwen3-Omni-Instruct模型则达到了51.2分。尽管当前模型通常能够理解用户输入,但它们经常提供不正确或不完整的答案。具体而言,它们在处理视觉提示(例如手势)时表现不佳,无法在多轮交互中保持历史上下文,并且未能在目标事件发生前延迟响应。结果表明,在模型能够成为可靠助手之前,仍有很大的改进空间。
cs.AI / 1 / 2608.20341
SDAD: Spec-Driven Agentic Development for the AI-Native SDLC
SDAD:基于规范驱动的自主开发在人工智能原生软件开发生命周期中的应用
Abstract
Frontier coding agents backed by large language models with context windows from hundreds of thousands to millions of tokens are restructuring the Software Development Life Cycle (SDLC). Rich context handling and multi-step reasoning now allow substantial Functional Requirement Documents (FRDs) and repository context to be ingested in a single workflow, making specification quality the execution fuel for autonomous delivery. This report formalises Spec-Driven Agentic Development (SDAD) as a synthesis of disciplined up-front formalisation and high-velocity implementation: intent capture, machine-readable specification, agentic synthesis, and independent multi-agent verification under human sign-off. We revisit the historical pendulum between Waterfall and Agile, introduce AI-code as a fourth production paradigm, and compare Human-Agile (circa 2020) with Agentic-SDAD (circa 2026) across artefacts, cadence, accountability, and security posture. Beyond process description, we extend the model to team role metamorphosis (engineer, QA, platform, and product functions), quantitative governance (Ambiguity Tax, Spec Fidelity, SER, and TCI_agentic with repair multiplier phi), and pragmatic adoption via hybrid estimation and a staged migration blueprint. Industrial and research evidence on AI-augmented testing and verification is integrated to motivate separation between synthesis and release authority. Overall, the paper argues that agentic speed does not eliminate engineering discipline; it relocates discipline upstream into specification precision, explicit gates, and auditable provenance.
Chinese Translation
由大型语言模型支持的前沿编码代理,具有数十万到数百万个标记的上下文窗口,正在重构软件开发生命周期(SDLC)。丰富的上下文处理和多步骤推理现在允许在单一工作流程中摄取大量功能需求文档(FRD)和代码库上下文,使得规范质量成为自主交付的执行动力。本报告正式提出规范驱动的自主开发(SDAD),作为严谨的前期形式化与高速度实施的综合体:意图捕捉、机器可读规范、代理合成以及在人工签署下的独立多代理验证。我们重新审视瀑布模型与敏捷开发之间的历史摆动,介绍AI代码作为第四种生产范式,并比较2020年的人工敏捷与2026年的自主SDAD在工件、节奏、问责和安全态势方面的差异。除了过程描述外,我们还将模型扩展到团队角色的转变(工程师、质量保证、平台和产品职能)、定量治理(模糊税、规范忠诚度、SER和TCI_agentic及修复乘数phi),以及通过混合估算和分阶段迁移蓝图实现务实的采纳。整合了关于AI增强测试和验证的工业和研究证据,以激励合成与发布权威之间的分离。总体而言,本文论证了自主速度并不消除工程纪律;而是将纪律重新定位到规范精确性、明确的门控和可审计的来源上。
cs.AI / 2 / 2608.20342
PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure
PrimeAgentOrchestrator:用于个人人工智能基础设施的记忆驱动代理生成
Abstract
Large language model (LLM) coding agents start each session with an empty context window, discarding accumulated knowledge from prior work. We present PrimeAgentOrchestrator (PAO), a system that spawns new instances of Claude Code -- Anthropic's terminal-based coding agent -- pre-loaded with relevant memories compiled from the user's existing personal databases. At spawn time, PAO queries two independently-operated memory backends in parallel (a PostgreSQL entity-observation database and a Cloudflare Worker semantic search index), fuses results using backend-specific retrieval strategies, and delivers the compiled briefing via filesystem injection that exploits the host agent's configuration auto-read behavior. PAO manages the full agent lifecycle including trust pre-seeding, readiness polling with error detection, and adaptive terminal text injection. We report on four months of regular deployment (December 2025 through March 2026) as an experience report, documenting three generations of context delivery mechanisms, the failure modes that motivated each redesign, and the engineering tradeoffs of bridging heterogeneous memory systems rather than building a unified one.
Chinese Translation
大型语言模型(LLM)编码代理在每个会话开始时都以空的上下文窗口启动,忽略了先前工作的积累知识。我们提出了PrimeAgentOrchestrator(PAO),这是一个生成Claude Code新实例的系统——Anthropic的基于终端的编码代理——并预加载了从用户现有个人数据库中编译的相关记忆。在生成时,PAO并行查询两个独立操作的记忆后端(一个PostgreSQL实体观察数据库和一个Cloudflare Worker语义搜索索引),使用特定于后端的检索策略融合结果,并通过利用主代理的配置自动读取行为的文件系统注入方式传递编译的简报。PAO管理完整的代理生命周期,包括信任预种植、带错误检测的准备状态轮询和自适应终端文本注入。我们报告了四个月的常规部署(2025年12月至2026年3月)的经验,记录了三代上下文传递机制、促使每次重新设计的失败模式,以及在桥接异构记忆系统而非构建统一系统时的工程权衡。
cs.AI / 3 / 2608.20378
Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification
真相深藏:通过潜在意图验证反制语义伪装
Abstract
Safety alignment in Large Language Models (LLMs) is often superficial, relying on refusal mechanisms that trigger only at the final stages of generation without erasing the foundational knowledge of harmful concepts acquired during pretraining. This study demonstrates that this architectural disconnect leaves models vulnerable to Semantic Camouflage -- adversarial attacks that wrap harmful intent in benign narrative contexts (e.g., creative writing), effectively bypassing standard input and output guardrails. By analyzing the latent activation trajectories of three distinct Small Language Model (SLM) families (Phi-3, Qwen2.5, and Gemma-2b) under adversarial stress, this research identifies a universal ``Intent Horizon'' -- a critical depth (typically 15--20\% of total layers) where the model's distinct, pre-trained representation of harmful intent collapses as it contextualizes the query into a ``safe'' narrative. Results indicate that while late-layer representations of camouflaged attacks are mathematically indistinguishable from safe queries (Detection Rate $< 20\%$), early-layer representations retain a distinct, detectable ``harm signature.'' Leveraging this insight, this paper proposes Latent Intent Verification (LIV), a lightweight probing defense. Experiments on the PKU-SafeRLHF dataset demonstrate that LIV outperforms standard guardrails by a margin of 20--50\% across all tested architectures, effectively neutralizing zero-day semantic attacks without requiring model retraining.
Chinese Translation
大型语言模型(LLMs)的安全对齐往往是表面的,依赖于只在生成的最后阶段触发的拒绝机制,而未能消除在预训练过程中获得的有害概念的基础知识。本研究表明,这种架构上的脱节使得模型容易受到语义伪装的攻击——这些对抗性攻击将有害意图包裹在无害的叙事背景中(例如,创意写作),有效绕过标准的输入和输出保护措施。通过分析在对抗压力下三种不同的小型语言模型(SLM)家族(Phi-3、Qwen2.5 和 Gemma-2b)的潜在激活轨迹,本研究识别出一个普遍的“意图视界”(Intent Horizon)——一个关键深度(通常为总层数的15%至20%),在此深度下,模型对有害意图的独特预训练表示在将查询情境化为“安全”叙事时崩溃。结果表明,尽管伪装攻击的后层表示在数学上与安全查询无法区分(检测率 < 20%),但前层表示仍保留有明显可检测的“有害特征”。利用这一洞察,本文提出了潜在意图验证(Latent Intent Verification, LIV),这是一种轻量级探测防御。对PKU-SafeRLHF数据集的实验表明,LIV在所有测试架构中比标准保护措施提高了20%至50%的效果,有效中和了零日语义攻击,而无需对模型进行再训练。
cs.AI / 4 / 2608.20379
A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications
多模态智能框架的基础与前沿调研:技术与应用
Abstract
Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestrate perception, memory, and decision-making around powerful LLM backbones. With the advent of large multimodal models (LMMs), these systems can process and integrate diverse modalities, including images, audio, and video, thereby improving their real-world applicability. Yet, while surveys of LLM-based agents exist, the role of multimodality in shaping agency has not been systematically examined in recent years. This survey fills the gap by analyzing the impact of multimodality across the core functional modules of the agentic framework: perception, reasoning, planning, memory, and action. Using this lens, we trace the evolution from text-centric agents to multimodal frameworks, examine how modalities are integrated through delegated, late-fusion, and early-fusion architectures, and assess the emergence of agentic behaviors enabled by grounded perception and multimodal reasoning. We organize existing work through a modality-centric taxonomy that links architectural design choices to agent capabilities. Moreover, we review multimodal agentic systems across various application domains, including Robotics, GUI & Web Navigation, Multimedia Content Generation & Editing, and Long-form Video Understanding & Retrieval. Beyond capabilities, we analyze performance across these settings and discuss efficiency-scalability trade-offs, including training and inference costs, latency, and deployment constraints. By focusing on the impact of multimodality in agentic design, we aim to identify key gaps and chart a roadmap toward robust and general-purpose intelligent systems.
Chinese Translation
大型语言模型(LLMs)的进展推动了对智能代理的研究浪潮:即推理、规划和行动的能力。这一努力产生了智能框架,围绕强大的LLM骨干网络协调感知、记忆和决策。随着大型多模态模型(LMMs)的出现,这些系统能够处理和整合多种模态,包括图像、音频和视频,从而提高其在现实世界中的适用性。然而,尽管已有基于LLM的代理的调研,但近年来对多模态在塑造智能代理中的作用尚未进行系统性研究。本调研通过分析多模态对智能框架核心功能模块(感知、推理、规划、记忆和行动)的影响,填补了这一空白。通过这一视角,我们追踪了从以文本为中心的代理到多模态框架的演变,考察了模态如何通过委派、后融合和前融合架构进行整合,并评估了基于扎根感知和多模态推理所驱动的智能行为的出现。我们通过一种以模态为中心的分类法组织现有工作,将架构设计选择与代理能力联系起来。此外,我们回顾了在各种应用领域(包括机器人、图形用户界面与网络导航、多媒体内容生成与编辑,以及长视频理解与检索)中的多模态智能系统。除了能力外,我们还分析了这些环境中的性能,并讨论了效率与可扩展性之间的权衡,包括训练和推理成本、延迟以及部署限制。通过关注多模态在智能设计中的影响,我们旨在识别关键空白,并绘制通向稳健和通用智能系统的路线图。
cs.AI / 5 / 2608.20384
Interpretable Multimodal Classification with Linear Discriminant Tree Ensembles
可解释的多模态分类与线性判别树集成
Abstract
Multimodal affect and behaviour classifiers that fuse heterogeneous text, audio, and visual streams must simultaneously achieve competitive accuracy and produce human-understandable explanations of the cues driving their decisions -- a dual objective that current high-capacity models, notably Transformers, only partially address. While Transformers attain strong predictive performance, their distributed representations and deep nonlinearity make it difficult to assign meaningful importance weights to individual multimodal features, limiting their use in trust-sensitive applications such as clinical affect monitoring and educational assessment. We address this gap by developing a framework based on tree-based ensembles that balances accuracy and interpretability. The framework encodes each modality into tokens, extracts and clusters concepts to reduce dimensionality, routes the fused modalities through tree-based ensemble classifiers, and interprets trends using a novel modified feature importance metric. The modified importance reduces the influence of the negative class in binary classification tasks, thereby improving indicator or marker detection. The proposed tree-based ensembles -- Linear Discriminant Tree (LDT), Linear Discriminant Forest (LDF), and Linear Discriminant AdaBoost (LDAB) -- achieve F1-mod gains of 4.3\% over the Multimodal Transformer and accuracy gains of 3.0\% over the primary interpretable multimodal baseline, Interpretable Multimodal Routing (IMR). The proposed multimodal feature importance extracts salient inter-modal concepts with substantially higher human-annotator agreement scores than default feature importance (62.2\% vs.\ 43.2\% on IEMOCAP; 46.7\% vs.\ 32.1\% on CMU-MOSI).
Chinese Translation
多模态情感和行为分类器融合异构的文本、音频和视觉流,必须同时实现竞争性的准确性,并生成人类可理解的解释,以说明驱动其决策的线索——这是当前高容量模型(特别是变换器)仅部分解决的双重目标。尽管变换器在预测性能上表现强劲,但其分布式表示和深度非线性使得很难为单个多模态特征分配有意义的重要性权重,限制了其在信任敏感的应用中的使用,如临床情感监测和教育评估。我们通过开发一个基于树集成的框架来填补这一空白,该框架平衡了准确性和可解释性。该框架将每种模态编码为标记,提取并聚类概念以降低维度,通过树集成分类器路由融合的模态,并使用一种新颖的修改特征重要性度量来解释趋势。修改后的重要性在二元分类任务中减少了负类的影响,从而改善了指示器或标记的检测。所提出的树集成——线性判别树(Linear Discriminant Tree, LDT)、线性判别森林(Linear Discriminant Forest, LDF)和线性判别AdaBoost(Linear Discriminant AdaBoost, LDAB)——在F1-mod上比多模态变换器提高了4.3%的增益,在准确性上比主要的可解释多模态基线(可解释多模态路由,Interpretable Multimodal Routing, IMR)提高了3.0%。所提出的多模态特征重要性提取了显著的跨模态概念,其人类标注者一致性评分显著高于默认特征重要性(在IEMOCAP上为62.2%对43.2%;在CMU-MOSI上为46.7%对32.1%)。
cs.AI / 6 / 2608.20389
Representation Affects Retrieval: A Case Study of Skill Discovery and Routing in a Multimodal Agent Harness
表示影响检索:多模态智能体工具中的技能发现与路由案例研究
Abstract
A production agent harness must discover and rank, from a growing library of skills, the one most appropriate for a user's task. At small scale this selection happens in context: the LLM planner chooses among skill representations exposed in its system prompt, without an explicit embedding-based retrieval step. We treat this in-context selection as the small-N counterpart to embedding-based skill retrieval at scale, and present a case study of how Tinycloud, a production multimodal video agent harness, represents its skills for the planner. The harness ships skills under two recurring representations: tool-skills that wrap a single external API or system tool and serve as primitive vocabulary, and workflow-skills that orchestrate tool-skill calls plus a template render to produce one named deliverable. The harness exposes them via two surfaces in the system prompt: an inlined-body surface (full instructions, scripts, templates) for autoloaded skills, and a one-line listing for on-demand skills. A six-task selection ablation across three exposure regimes (all-on, default, all-off) shows that full autoload selects the gold skill on every task; all-off slows execution and produces hard discovery failures; and the production default misroutes one task because its lexical signal collides with an autoloaded tool-skill that pulls planner attention away from a listed workflow-skill. The headline finding is that in-prompt exposure of skills is not monotonically helpful: partial exposure can create lexical competition that suppresses correct selection. We connect this small-N observation to recent retrieval-based skill-routing work at large scale, and frame this contribution as a case study rather than a benchmark.
Chinese Translation
生产智能体工具必须从不断增长的技能库中发现并排名出最适合用户任务的技能。在小规模情况下,这种选择是在上下文中进行的:大型语言模型(LLM)规划器在其系统提示中选择暴露的技能表示,而没有明确的基于嵌入的检索步骤。我们将这种上下文选择视为小规模情况下的嵌入式技能检索的对应物,并展示了Tinycloud,一个生产多模态视频智能体工具,如何为规划器表示其技能的案例研究。该工具以两种重复的表示形式提供技能:工具技能(tool-skills),它封装了单一的外部API或系统工具,并作为原始词汇;工作流技能(workflow-skills),它协调工具技能调用和模板渲染,以生成一个命名的交付物。该工具通过系统提示中的两个表面暴露这些技能:一个内联主体表面(完整的指令、脚本、模板)用于自动加载的技能,以及一个按需技能的一行列表。对三种暴露机制(全开、默认、全关)进行的六项任务选择消融实验表明,完全自动加载在每个任务中选择了最佳技能;全关则减慢了执行并产生了严重的发现失败;而生产默认设置错误地路由了一项任务,因为其词汇信号与一个自动加载的工具技能发生冲突,导致规划器的注意力偏离了列出的工作流技能。主要发现是,技能在提示中的暴露并不是单调有益的:部分暴露可能会产生词汇竞争,从而抑制正确选择。我们将这一小规模观察与近期基于检索的技能路由工作联系起来,并将此贡献框架视为案例研究,而非基准测试。
cs.AI / 7 / 2608.20397
Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory
Nexus:用于统一内存的深度自适应 KV 缓存拼接和检索解耦工具路由的代理大型语言模型
Abstract
Agentic large language models (LLMs) on the Model Context Protocol (MCP) re-encode verbose tool schemas every turn, so prefill - quadratic in sequence length - dominates time-to-first-token (TTFT) as the tool registry grows. Nexus's primary lever is to decouple routing from the schema-prefill cost: an INT8 semantic lookaside buffer (SLB) with a calibrated cross-encoder margin gate selects tools by retrieval, and arguments are generated over a compressed textual signature (median 19 tokens) rather than over spliced key/value (KV) cache. This path is depth-independent: routing accuracy stays near 89% as the registry scales to 250 tools - where a concatenate-all-schemas baseline overflows the context window entirely - and it reaches a first-argument token 1.66x sooner than a full-schema re-prefill at a ~80% main-context token saving. As a secondary, bounded lever we transplant a compiled schema KV block directly into the live context. This is fundamentally limited by rotary position embedding (RoPE) phase drift: an anchored splice is output-exact, but off-anchor placement corrupts attention, so beyond a threshold P=256 Nexus repairs the seam with a depth-adaptive suffix redecode that escalates to a full re-prefill. The resulting never-regress property is a guarantee on output fidelity (top-1 agreement, D_KL approx. 0) - not on latency, which can dip to 0.98x before converging to parity - alongside a 1.1-1.7x TTFT speedup at moderate depth that narrows to parity at deep context. Two negative results bound the design: the off-anchor RoPE fidelity boundary, and the failure of a reference-free drift gate to predict drift (Spearman rho = 0.193). All measurements are from one model tuple (Qwen2.5-14B-Instruct Q4_K_M) on Apple-silicon unified memory; the qualitative boundaries generalize, while the quantitative envelope is tuple-specific.
Chinese Translation
基于模型上下文协议(MCP)的代理大型语言模型(LLMs)在每次迭代中重新编码冗长的工具模式,因此随着工具注册表的增长,预填充(与序列长度的平方成正比)主导了首次令牌时间(TTFT)。Nexus 的主要杠杆是将路由与模式预填充成本解耦:一个具有校准交叉编码器边际门的 INT8 语义旁路缓冲区(SLB)通过检索选择工具,参数是在压缩的文本签名(中位数 19 个令牌)上生成,而不是在拼接的键/值(KV)缓存上生成。该路径与深度无关:随着注册表扩展到 250 个工具,路由准确性保持在接近 89% - 在此情况下,连接所有模式的基线完全溢出上下文窗口 - 并且它比完全模式重新预填充早 1.66 倍到达第一个参数令牌,同时节省约 80% 的主上下文令牌。作为一个次要的、有界杠杆,我们将一个编译的模式 KV 块直接移植到实时上下文中。这在根本上受到旋转位置嵌入(RoPE)相位漂移的限制:一个固定的拼接输出是精确的,但偏离锚点的放置会破坏注意力,因此在阈值 P=256 以上,Nexus 通过深度自适应后缀重新解码修复接缝,该过程升级为完全重新预填充。由此产生的“永不回归”特性保证了输出的保真度(top-1 一致性,D_KL 约为 0) - 而不是延迟,延迟可以降至 0.98 倍,然后收敛到平衡 - 同时在适度深度下实现 1.1-1.7 倍的 TTFT 加速,在深度上下文下收窄至平衡。两个负结果限制了设计:偏离锚点的 RoPE 保真度边界,以及无参考漂移门预测漂移的失败(Spearman rho = 0.193)。所有测量均来自一个模型元组(Qwen2.5-14B-Instruct Q4_K_M)在 Apple 硅统一内存上;定性边界具有普遍性,而定量范围则是元组特定的。
cs.AI / 8 / 2608.20398
Environmental Slow AI: Design Principles for Generative Systems
环境慢AI:生成系统的设计原则
Abstract
Generative AI (genAI) systems produce cultural artefacts at scale, but they also reflect embedded cultural values through their design. Once identified, these values become open to deliberate reshaping. This position paper examines the maximalist values of current generative AI through an environmental humanities tradition and proposes design principles in which environmental sustainability serves as the core value instead. The principles are developed under the umbrella of Slow AI, a term that already circulates across several distinct research and practice programs. Five design principles are articulated (restraint, sufficiency, selectivity over retention, material visibility, and friction as affordance), each of them illustrated against the current design of widely deployed systems. Each principle operates at two levels: a design implementation, and an interpretive layer at which users and developers are prompted toward reflective engagement with the system. Together these principles extend human agency by restoring decisions that frictionless defaults have silently removed and do so by building interpretive reflection into design.
Chinese Translation
生成AI(genAI)系统大规模地生产文化产品,但它们也通过设计反映了嵌入的文化价值。一旦识别出这些价值,就可以进行有意识的重塑。本文探讨了当前生成AI的极端价值,借助环境人文学科的传统,提出以环境可持续性为核心价值的设计原则。这些原则在“慢AI”(Slow AI)的框架下发展,该术语已经在多个不同的研究和实践项目中流传。本文阐述了五个设计原则(克制、充足、选择性优于保留、材料可见性,以及摩擦作为赋能),每个原则都与当前广泛应用系统的设计进行对比。每个原则在两个层面上运作:设计实施层面,以及一个解释层面,在该层面上用户和开发者被引导进行对系统的反思性参与。这些原则共同扩展了人类的能动性,通过恢复那些无摩擦默认设置所默默移除的决策,并通过将解释性反思融入设计来实现这一目标。
cs.AI / 9 / 2608.20400
When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory
当检索在开始之前失败:结构性间接前提驱逐作为代理记忆中的保留失败
Abstract
Agentic memory under a fixed budget involves two stages: retention and retrieval. Existing retrieval-centered paradigms implicitly assume necessary evidence survives eviction, but we challenge this by isolating a pre-retrieval failure mode: structurally indirect prerequisite eviction, in which upstream blocks weakly aligned with the query are discarded under budget pressure. We provide an operational definition of this failure, a reproducible deterministic benchmark, and per-seed trace diagnostics. Finally, we evaluate Dependency-aware Semantic Garbage Collection (DSGC), a one-hop graph-aware rule. In our main suite, DSGC improves full-chain retention from 0.03 to 0.90 under a lexical encoder and from 0.23 to 1.00 under a sentence encoder. Robustness checks then identify the budget and scaling regimes where the one-hop rule holds or degrades. Our released pipeline and failure postmortem support mechanistic analysis of retention before retrieval as a distinct failure boundary.
Chinese Translation
在固定预算下,代理记忆涉及两个阶段:保留和检索。现有的以检索为中心的范式隐含假设必要的证据在驱逐后仍然存在,但我们通过隔离一种预检索失败模式来挑战这一假设:结构性间接前提驱逐,其中与查询弱对齐的上游块在预算压力下被丢弃。我们提供了这一失败的操作性定义、可重复的确定性基准和每个种子的跟踪诊断。最后,我们评估了依赖感知语义垃圾收集(Dependency-aware Semantic Garbage Collection, DSGC),这是一种单跳图感知规则。在我们的主要测试中,DSGC在词汇编码器下将全链保留率从0.03提高到0.90,在句子编码器下从0.23提高到1.00。稳健性检查随后识别出单跳规则有效或退化的预算和扩展范围。我们发布的管道和失败事后分析支持在检索之前对保留进行机械分析,作为一种独特的失败边界。
cs.AI / 10 / 2608.20401
World models of environment, agent and joint agent-environment systems
环境、智能体及联合智能体-环境系统的世界模型
Abstract
World models are a central component of model-based reinforcement learning. They are usually discussed in terms of what variables they predict, such as observations, rewards, states, latent or information states. We argue that there is a prior distinction: which channel they model. We consider three cases: the environment channel $O_{:} \mid A_{:}$, the agent channel $A_{:} \mid O_{:}$, and the realised joint process $(A, O)_{:}$, equivalently viewed as a channel with no inputs. Using computational mechanics, we define canonical predictive models for these three cases as $\epsilon$-transducers or $\epsilon$-machines. Canonical environment models recover standard predictive state representations, while the other two give analogous notions of canonical models for the agent and the joint system. We then build canonical support-restricted environment and agent models induced by closed-loop coupling, whose predictive equivalences range over continuations supported by the realised interaction. The key structural result is that canonical support-restricted environment states factor through the canonical joint causal states, and their transition structure is induced directly from the joint model; the agent-side construction is dual. Finally, we give a POMDP/controller example in which the unrestricted environment model has infinitely many states while the canonical support-restricted model induced by the coupling is finite. The framework clarifies what different world models are models of, and how coupling and support restriction can change their canonical predictive structure and complexity.
Chinese Translation
世界模型是基于模型的强化学习的核心组成部分。通常讨论它们预测的变量,例如观察值、奖励、状态、潜在状态或信息状态。我们认为存在一个先前的区分:它们建模的通道。我们考虑三种情况:环境通道 $O_{:} ext{ | } A_{:}$,智能体通道 $A_{:} ext{ | } O_{:}$,以及实现的联合过程 $(A, O)_{:}$,可以等效地视为没有输入的通道。利用计算力学,我们为这三种情况定义了典型的预测模型,称为 $ ext{ε}$-变换器或 $ ext{ε}$-机器。典型的环境模型恢复了标准的预测状态表示,而另外两个则给出了智能体和联合系统的典型模型的类似概念。接着,我们构建了由闭环耦合引发的典型支持限制环境和智能体模型,其预测等价性涵盖了由实现的交互支持的延续。关键的结构结果是,典型支持限制环境状态通过典型的联合因果状态进行分解,其转移结构直接由联合模型引导;智能体侧的构造是对偶的。最后,我们给出了一个 POMDP/控制器的例子,其中无限制环境模型具有无限多个状态,而由耦合引发的典型支持限制模型是有限的。该框架阐明了不同的世界模型所建模的内容,以及耦合和支持限制如何改变它们的典型预测结构和复杂性。
cs.AI / 11 / 2608.20414
StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models
StateSight:视觉语言模型中潜在空间状态重建的基准测试
Abstract
Vision-language models are increasingly used for multimodal question answering, yet their ability to reconstruct latent spatial structure from a single image remains difficult to isolate. Broad benchmarks often combine perception, optical character recognition, domain knowledge, linguistic priors, and reasoning in the same evaluation. We introduce StateSight, a procedurally generated benchmark for cube-net opposite-face reasoning, occluded cube-tower counting, and 4-neighbor connected-component counting. Each task family contains 300 single-image prompts with deterministic oracle labels and exact-match scoring. OpenAI GPT-5.5, using the API model identifier gpt-5.5, achieved 59.3%, 33.3%, and 28.3% accuracy across the three tasks, while Claude Sonnet 5 achieved 53.3%, 18.7%, and 7.3%. All final direct runs had zero format errors. A 30-participant human baseline on 60 items exceeded both models on every task, with mean accuracies of 80.8%, 68.8%, and 64.3%. Visible-derivation analysis identified recurring errors in image-state reconstruction and reasoning procedure. We also introduce StateSight-Steps, a companion dataset of 900 interleaved image-text examples and 3,600 deterministic intermediate visual states. The results show that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference.
Chinese Translation
视觉语言模型在多模态问答中的应用日益增多,但从单幅图像重建潜在空间结构的能力仍然难以单独评估。广泛的基准测试通常将感知、光学字符识别、领域知识、语言先验和推理结合在同一评估中。我们提出了StateSight,这是一个程序生成的基准,涵盖了立方体网对面推理、被遮挡的立方体塔计数和4邻域连通组件计数等任务。每个任务系列包含300个单图像提示,配有确定性的权威标签和精确匹配评分。使用API模型标识符gpt-5.5的OpenAI GPT-5.5在这三项任务中的准确率分别为59.3%、33.3%和28.3%,而Claude Sonnet 5的准确率为53.3%、18.7%和7.3%。所有最终的直接运行均没有格式错误。60个项目的30名参与者的人工基线在每个任务上均超过了这两个模型,平均准确率为80.8%、68.8%和64.3%。可见推导分析识别出图像状态重建和推理过程中的重复错误。我们还引入了StateSight-Steps,这是一个包含900个交错图像-文本示例和3,600个确定性中间视觉状态的配套数据集。结果表明,格式有效的响应可能掩盖了恢复可验证视觉推理所需空间结构的失败。
cs.AI / 12 / 2608.20420
Categorical AI phenomenology: A first-person approach
类别人工智能现象学:一种第一人称方法
Abstract
This paper develops a phenomenology-first approach to artificial consciousness by reframing consciousness as the subjective experience enacted through an agent's interface with the world. We shift the methodological focus to first-person structures, modeled mathematically by categories derived from Q-networks to capture actions and phenomenological invariants. In this framework, Q-networks are conceptualized as relational interfaces encoding agent-world interaction, analogous to how the dynamical states of a computer depend on its sensory inputs, previous states, and actions. Our work provides a rigorous framework for interface consciousness to describe computational systems that embed information-processing into phenomenological structure. The approach aligns with 4E approaches to cognition by emphasizing enactive, embedded, and extended dimensions of experience. The paper thus offers a principled, relational, and phenomenological account of artificial phenomenology grounded in categorical mathematics.
Chinese Translation
本文通过将意识重新构建为通过主体与世界的界面所实施的主观体验,发展了一种以现象学为首的人工意识方法。我们将方法论重点转向第一人称结构,这些结构通过从 Q-networks 派生的类别进行数学建模,以捕捉行为和现象学不变性。在这一框架中,Q-networks 被概念化为编码主体与世界交互的关系界面,类似于计算机的动态状态如何依赖于其感官输入、先前状态和行为。我们的研究为界面意识提供了一个严格的框架,以描述将信息处理嵌入现象学结构的计算系统。这种方法通过强调体验的行动性、嵌入性和扩展性,与 4E 认知方法相一致。因此,本文提供了一个基于类别数学的原则性、关系性和现象学的人工现象学解释。
cs.AI / 13 / 2608.20425
Who Delegates to AI? Evidence from 53,000 Agent Configurations
谁将任务委托给人工智能?来自53,000个代理配置的证据
Abstract
A growing literature measures how far occupations are exposed to AI, but these measures capture where AI could perform tasks, not whether workers have adopted it. We propose a new layer of exposure, delegated exposure, which records whether a worker has committed a task to AI by building it into a workflow. We operationalize it as the Agentic Adoption Index (AAI), which measures how closely an occupation's tasks match the agentic routines practitioners have already built and shared. We embed roughly 53,000 agent skill specifications from the Manus Skills Marketplace, compute their semantic similarity to about 18,000 O*NET task statements, and aggregate to the occupation level. Three findings follow. First, the occupations where delegation concentrates differ sharply from those pre-AI frameworks identified as most at risk. Second, the AAI tracks what AI could do more closely than what workers currently use it for. Third, the AAI peaks below the top of the wage distribution and at the bachelor's level, declining at both extremes. Technical availability explains most of this variation, but not the shortfall among the most educated occupations, so feasibility alone cannot account for who adopts. That shortfall may reflect work that resists advance specification, or professional discretion over the pace of codification. Distinguishing the two, and tracking how these measures diverge over time, will require repeated measurement.
Chinese Translation
日益增长的文献测量职业对人工智能的暴露程度,但这些测量捕捉的是人工智能可以执行任务的地方,而不是工人是否已经采用它。我们提出了一种新的暴露层次,即委托暴露,记录工人是否通过将任务纳入工作流程而将其委托给人工智能。我们将其操作化为代理采用指数(Agentic Adoption Index, AAI),该指数衡量职业任务与从业者已经构建并共享的代理性例程之间的匹配程度。我们嵌入了大约53,000个来自Manus Skills Marketplace的代理技能规格,计算它们与约18,000个O*NET任务陈述的语义相似性,并汇总到职业层面。以下是三个发现。首先,委托集中度较高的职业与那些在人工智能之前的框架中被识别为最有风险的职业有明显不同。其次,AAI更紧密地跟踪人工智能可能做的事情,而不是工人当前使用它的方式。第三,AAI在工资分布的顶部以下和学士学位水平达到峰值,在两个极端都呈下降趋势。技术可用性解释了大部分这种变异,但不能解释在受教育程度最高的职业中的短缺,因此仅靠可行性无法说明谁会采用。这一短缺可能反映了抵制提前规范的工作,或专业人士对编码速度的自由裁量。区分这两者,并跟踪这些测量随时间的分歧,将需要重复测量。
cs.AI / 14 / 2608.20477
STCO: Conditional Neural Operators for Time-Dependent PDEs
STCO:用于时间依赖性偏微分方程的条件神经算子
Abstract
Neural operators have emerged as efficient surrogates for time-dependent physical systems governed by partial differential equations (PDEs), but their future-state predictions are often conditioned only on observed states and static problem descriptors. For control or optimization, however, body motion, inflow, or forcing are prescribed for the query without being determined solely by the observed state. We introduce the Spatiotemporal Conditional Operator (STCO) for prescribed-condition operator learning (PCOL), a common interface that supplies prescribed target-time condition fields to heterogeneous backbone architectures while retaining their architecture-specific core computation and context pathways. Its condition interface combines Flow-Aware Graph Leaf (FAGL) with Dual-Site Feature-wise Linear Modulation (DSFiLM). Non-learned FAGL uses vorticity from the final observed frame to construct a fixed-cardinality adaptive partition, then co-locates the observed history and target-time condition fields at its regional coordinates. DSFiLM injects separate motion, inflow, and force routes before and after operator computation through current-feature-driven slot- and channel-wise gates. We evaluate twelve matched backbone architectures with different existing physical and temporal inputs. The immersed-boundary computational fluid dynamics (CFD) benchmark spans prescribed motion, inflow disturbances, body-force actuation, and morphology. Across twelve matched backbones, three regimes, and two lead ranges, STCO yields mean paired reductions of 31.1% in relative-L2 field error and 24.7% in normalized pressure-derived load error. It also lowers longer-lead field error for 11 backbones, while interventions on individual condition groups produce measurable prediction changes for every group evaluated.
Chinese Translation
神经算子已成为时间依赖性物理系统的高效替代方案,这些系统由偏微分方程(PDEs)所支配,但它们的未来状态预测通常仅基于观察到的状态和静态问题描述。然而,对于控制或优化而言,身体运动、流入或外力是为查询规定的,而并非仅由观察到的状态决定。我们引入了时空条件算子(Spatiotemporal Conditional Operator,STCO),用于规定条件算子学习(Prescribed-Condition Operator Learning,PCOL),这是一个通用接口,能够为异构骨干架构提供规定的目标时间条件场,同时保留其架构特定的核心计算和上下文路径。其条件接口结合了流动感知图叶(Flow-Aware Graph Leaf,FAGL)与双站点特征线性调制(Dual-Site Feature-wise Linear Modulation,DSFiLM)。非学习的FAGL利用最后观察帧的涡度构建固定基数的自适应划分,然后在其区域坐标上共同定位观察历史和目标时间条件场。DSFiLM通过当前特征驱动的槽和通道门在算子计算前后注入独立的运动、流入和力路径。我们评估了十二种匹配的骨干架构,这些架构具有不同的现有物理和时间输入。浸没边界计算流体动力学(CFD)基准涵盖了规定的运动、流入扰动、体力驱动和形态。在十二个匹配的骨干架构、三个状态和两个前导范围中,STCO在相对L2场误差上平均减少了31.1%,在归一化压力导出的载荷误差上减少了24.7%。它还降低了11个骨干架构的长前导场误差,同时对各个条件组的干预产生了可测量的预测变化。
cs.AI / 15 / 2608.20485
Terminal Agents: A Survey of AI Agents in Command-Line Environments
终端代理:命令行环境中人工智能代理的综述
Abstract
Large language model agents increasingly act through terminals, yet existing surveys disperse terminal-mediated behavior across software engineering, tool use, and computer-use research. We regard terminal agents as systems whose dominant progress-bearing action--observation loop is mediated by terminal command execution, textual feedback, and stateful environment interaction. Using terminal-mediated execution as an organizing lens, this survey establishes workload-level boundaries and connects system architecture, competence acquisition, and evaluation through a seven-dimensional terminal competence profile. Our synthesis shows that realized behavior is jointly shaped by the model, interface, harness, runtime, and environment. Executable trajectories ground learning in action consequences, verification, and recovery, whereas prevailing evaluations emphasize final outcomes and expose process quality, recovery, and governance unevenly. Bounded fixed-condition diagnostics illustrate two implications: benchmark families expose different process signals, and matched system comparisons reveal benchmark-dependent performance and limits of component attribution. These findings motivate explicit reporting of system and runtime conditions, supported by replayable traces and process-level evidence. The framework provides a unified basis for studying terminal-mediated agency across software engineering and emerging application domains.
Chinese Translation
大型语言模型代理越来越多地通过终端进行操作,但现有的综述将终端介导的行为分散在软件工程、工具使用和计算机使用研究中。我们将终端代理视为其主要进展驱动行为——观察循环通过终端命令执行、文本反馈和状态环境交互来介导的系统。通过终端介导执行作为组织视角,本综述建立了工作负载级别的边界,并通过七维终端能力轮廓连接系统架构、能力获取和评估。我们的综合分析表明,实现的行为是由模型、接口、框架、运行时和环境共同塑造的。可执行轨迹将学习扎根于行动后果、验证和恢复,而现有的评估则强调最终结果,并不均衡地暴露过程质量、恢复和治理。有限固定条件的诊断展示了两个启示:基准系列暴露不同的过程信号,而匹配系统的比较揭示了基准依赖的性能和组件归因的限制。这些发现促使我们明确报告系统和运行时条件,并通过可重放的痕迹和过程级证据进行支持。该框架为研究软件工程和新兴应用领域中的终端介导代理提供了统一的基础。
cs.AI / 16 / 2608.20490
Lost in Translation: How Universal Ethical Values Fail to Translate Across Global Contexts
翻译中的迷失:普遍伦理价值如何在全球背景下失去翻译能力
Abstract
AI ethics frameworks treat values such as fairness, transparency, and accountability as universal and uniformly operationalizable across contexts. We examined how 14 experts across 10 countries made sense of AI in practice, reinterpreted core values, and envisioned governance alternatives. We found that AI deployment is characterized by structurally unequal conditions, marked by infrastructural constraints, extractive practices, and a "mystification" of technology, which fundamentally shape perceptions of risks and opportunities. Our findings reveal that experts reinterpret values to fit local moral logics: privacy as collective and relational rather than individual; transparency as trust-building accountability rather than technical disclosure; and fairness as equity in access and representation rather than parity in outcomes. We identify these as translation gaps between encoded global frameworks and situated local practices. Finally, we propose pathways toward plural governance that redistributes epistemic authority and treats ethical negotiation as an ongoing, context-sensitive process rather than a settled technical standard.
Chinese Translation
人工智能伦理框架将公平、透明和问责等价值观视为普遍的,并且在不同背景下可以统一操作。我们考察了来自10个国家的14位专家如何理解实践中的人工智能,重新诠释核心价值观,并构想治理替代方案。我们发现,人工智能的部署特征是结构性的不平等条件,受基础设施限制、掠夺性实践和技术的“神秘化”所影响,这些因素根本上塑造了对风险和机会的认知。我们的研究结果揭示,专家们重新诠释价值观以适应当地的道德逻辑:隐私被视为集体和关系性的,而非个体的;透明度被视为建立信任的问责,而非技术披露;公平被视为获取和代表的平等,而非结果的对等。我们将这些视为全球框架与具体地方实践之间的翻译差距。最后,我们提出了通向多元治理的路径,重新分配认知权威,并将伦理协商视为一个持续的、与背景相关的过程,而非一个确定的技术标准。
cs.AI / 17 / 2608.20510
A Temporal Planning Approach for Intelligent Flood Response
智能洪水响应的时间规划方法
Abstract
Effective response to multiple, simultaneously flooded areas requires coordinating appropriate actions in the correct temporal order, under severe resource constraints. Automated planning provides a foundation for addressing this challenge by generating time-aware schedules, given a formal description of available resources, constraints, and goals. This work presents an intelligent flood-response framework that exploits temporal planning and models the complete operational life cycle of flood response. The framework incorporates priority-driven triage, route accessibility and travel costs, resource allocation, and supply management, while also supporting mid-execution re-planning in response to unexpected environmental changes. The framework is formulated both in the Action Notation Modeling Language (ANML) and the Planning Domain Definition Language (PDDL) 2.1, facilitating compatibility with a wider range of temporal planners. Experimental results establish the feasibility and scalability of the proposed framework, showing that flood response scenarios can be effectively modeled and solved using temporal planning, while providing guidance on planner selection.
Chinese Translation
有效应对多个同时被淹没的区域需要在严重的资源限制下协调适当的行动,并按照正确的时间顺序进行。自动化规划为解决这一挑战提供了基础,通过生成时间感知的时间表,基于对可用资源、约束和目标的正式描述。本研究提出了一个智能洪水响应框架,该框架利用时间规划并建模洪水响应的完整操作生命周期。该框架结合了优先级驱动的分诊、路线可达性和旅行成本、资源分配和供应管理,同时支持在执行过程中对意外环境变化的重新规划。该框架同时采用行动符号建模语言(Action Notation Modeling Language, ANML)和规划领域定义语言(Planning Domain Definition Language, PDDL)2.1进行表述,便于与更广泛的时间规划器兼容。实验结果证明了所提框架的可行性和可扩展性,显示洪水响应场景可以有效地通过时间规划进行建模和解决,同时为规划器的选择提供指导。
cs.AI / 18 / 2608.20518
FL-MAESTRO: Multi-Agent LLM Orchestration for Resource-Constrained Federated Learning
FL-MAESTRO:资源受限的联邦学习中的多智能体大语言模型编排
Abstract
In Federated Learning (FL), the communication topology is a runtime variable rather than a fixed design choice, since links and edge devices drop in and out during training. Each round, the server must commit three coupled decisions, namely the communication topology, per-client resource allocation, and the aggregation rule for combining local updates. Recent agentic systems have begun bringing large language models (LLM) into FL, but the existing line of work either operates at setup time or handles a single runtime dimension such as client selection. We propose FL-MAESTRO, a multi-agent orchestrator that makes the joint runtime FL decision directly through three specialist LLM agents, one per decision dimension. A coordinator combines their analyses into a single decision, and a non-LLM feasibility check confirms it before the round executes. Because the orchestrator consumes the server's predicted-failure list, it withholds clients whose updates would never be aggregated, which removes the dominant source of wasted round energy in classical FL on volatile edge networks. Because client state is read as natural-text profiles, the same orchestrator extends to heterogeneous device classes without per-class energy models. On a non-IID CIFAR-10 benchmark, FL-MAESTRO matches the accuracy of the strongest energy-aware baseline while cutting wasted round energy from over a third to near zero. Code is available at https://github.com/denoslab/FL-MAESTRO.
Chinese Translation
在联邦学习(Federated Learning, FL)中,通信拓扑是一个运行时变量,而非固定的设计选择,因为在训练过程中,链接和边缘设备会随时掉线或重新连接。每一轮,服务器必须做出三个相互关联的决策,即通信拓扑、每个客户端的资源分配以及用于组合本地更新的聚合规则。近期的智能系统开始将大语言模型(Large Language Models, LLM)引入FL,但现有的研究要么在设置时操作,要么处理单一的运行时维度,如客户端选择。我们提出了FL-MAESTRO,一个多智能体编排器,通过三个专业的LLM代理(每个决策维度一个)直接做出联合的运行时FL决策。协调者将他们的分析整合为一个单一决策,并在执行轮次之前通过非LLM可行性检查进行确认。由于编排器利用服务器的预测失败列表,它会排除那些更新永远不会被聚合的客户端,从而消除了在波动的边缘网络中经典FL中浪费轮次能量的主要来源。由于客户端状态被读取为自然文本档案,相同的编排器可以扩展到异构设备类别,而无需针对每个类别的能量模型。在非独立同分布(non-IID)CIFAR-10基准测试中,FL-MAESTRO的准确性与最强的能量感知基线相匹配,同时将浪费的轮次能量从超过三分之一减少到接近零。代码可在 https://github.com/denoslab/FL-MAESTRO 获取。
cs.AI / 19 / 2608.20549
Volumetric Radiology AI in the Era of Multimodal Large Language Models
多模态大型语言模型时代的体积放射学人工智能
Abstract
Advances in multimodal large language models (MLLMs) are extending radiological artificial intelligence (AI) beyond task-specific image analysis toward multimodal understanding and reasoning. Volumetric radiology, however, presents a fundamental representational mismatch: clinical interpretation often requires full-volume spatial context and acquisition-dependent quantitative information, whereas current MLLMs are commonly conditioned on selected two-dimensional (2D) images, compressed visual representations, or report-derived text. Reliable volumetric radiology AI therefore requires representations that preserve task-relevant three-dimensional (3D) information and systems that can access, verify, and integrate this information across clinical workflows. In this Review, we examine more than 200 publications through July 2026. We organize the literature around volumetric representation and multimodal understanding at the model level, agentic orchestration at the system level, and their links to clinical applications and evaluation. We review volumetric foundation models, language alignment and compression strategies, and agentic systems that extend MLLMs through planning, tools, memory, and workflow interaction. We distinguish settings in which selected 2D views or report-mediated reasoning may suffice from those that warrant native volumetric modeling. We also introduce a Claim-Design-Validation framework to assess whether technical, workflow, and clinical claims are matched by appropriate design and validation. Across the literature, native volumetric modeling and agentic capabilities depend on the spatial, quantitative, contextual, and workflow requirements of the intended task. Clinical credibility requires faithful volumetric representation, traceable system behavior, claim-aligned validation, and clearly defined human oversight in realistic workflows.
Chinese Translation
多模态大型语言模型(MLLMs)的进展正在将放射学人工智能(AI)从特定任务的图像分析扩展到多模态理解和推理。然而,体积放射学存在根本的表征不匹配:临床解释通常需要完整的体积空间上下文和依赖于采集的定量信息,而当前的MLLMs通常基于选定的二维(2D)图像、压缩的视觉表征或报告派生的文本。可靠的体积放射学AI因此需要能够保留与任务相关的三维(3D)信息的表征,以及能够在临床工作流程中访问、验证和整合这些信息的系统。在本综述中,我们审查了截至2026年7月的200多篇文献。我们围绕模型层面的体积表征和多模态理解、系统层面的代理协调,以及它们与临床应用和评估的联系对文献进行了组织。我们回顾了体积基础模型、语言对齐和压缩策略,以及通过规划、工具、记忆和工作流程交互扩展MLLMs的代理系统。我们区分了在某些情况下选定的2D视图或报告中介推理可能足够的设置与那些需要原生体积建模的设置。我们还引入了一个声明-设计-验证框架,以评估技术、工作流程和临床声明是否与适当的设计和验证相匹配。在文献中,原生体积建模和代理能力依赖于预期任务的空间、定量、上下文和工作流程要求。临床可信度要求忠实的体积表征、可追溯的系统行为、与声明一致的验证,以及在现实工作流程中明确的人工监督。
cs.AI / 20 / 2608.20564
Consilience: Conformally Calibrated Communication Control for Hidden-Profile Multi-Agent Reasoning
一致性:用于隐性特征多智能体推理的共形校准通信控制
Abstract
Multi-agent LLM systems can improve reasoning by pooling diverse perspectives, but their effectiveness depends on coordinating communication, particularly in hidden-profile settings where each agent holds only part of the evidence required for a correct decision. Existing protocols, including fixed schedules, round-robin exchange, and unstructured debate, provide no guarantee that a conversational action is appropriate. We propose Consilience, an inference-time orchestration framework that both steers and certifies multi-agent communication under distributed private information. At each turn, Consilience summarizes the discussion using a compact state capturing uncertainty, disagreement, evidence gain, redundancy, and premature consensus, then selects both a communication intervention (challenge, clarify, seek evidence, or route) and an appropriate speaker. Its central contribution is a round-wise conformal calibration procedure that provides a distribution-free, finite-sample guarantee: at each discussion round, conditional on reaching that round, the one-step regret of a controller's proposed action is bounded by a calibrated threshold with marginal probability at least 1 - alpha; an acceptance mechanism enforces the same guarantee for the executed action by replacing inadmissible proposals. On HiddenBench-style hidden-profile tasks spanning 12 open and closed weight language models, Consilience improves decision accuracy and communication efficiency over fixed and unstructured discussion protocols, sometimes surpassing a full-information baseline where every agent observes all evidence. These results demonstrate that certified adaptive communication control can be more valuable than increasing information availability, providing a practical mechanism for reliable multi-agent LLM coordination.
Chinese Translation
多智能体大语言模型(LLM)系统通过汇聚多样化的观点可以改善推理,但其有效性依赖于协调通信,特别是在隐性特征设置中,每个智能体仅持有做出正确决策所需部分证据。现有的协议,包括固定时间表、轮流交换和非结构化辩论,并不能保证对话行为的适当性。我们提出了一致性(Consilience),一个推理时的编排框架,能够在分布式私人信息下引导和验证多智能体通信。在每个回合中,一致性通过一个紧凑的状态总结讨论,该状态捕捉不确定性、分歧、证据增益、冗余和过早共识,然后选择一个通信干预(挑战、澄清、寻求证据或路由)和一个合适的发言者。其核心贡献是一个逐轮的共形校准程序,提供无分布的有限样本保证:在每个讨论回合中,条件是达到该回合,控制者提议的行动的一步后悔被一个校准阈值所限制,其边际概率至少为 1 - alpha;一个接受机制通过替换不可接受的提议来强制执行已执行行动的相同保证。在涵盖 12 个开放和封闭权重语言模型的 HiddenBench 风格隐性特征任务中,一致性在决策准确性和通信效率上优于固定和非结构化讨论协议,有时甚至超过了每个智能体观察到所有证据的全信息基准。这些结果表明,经过认证的自适应通信控制比增加信息可用性更具价值,为可靠的多智能体 LLM 协调提供了一种实用机制。
cs.AI / 21 / 2608.20569
Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation
开放权重的掩蔽内省:测量语言模型能够报告其自身计算的能力
Abstract
Are frontier models able to introspect about their internal states? Recent work suggests that under certain conditions a complex enough model can audit its own internals, call out what changed, and report back confidently about it. We tested that claim on eight open-weight models from seven families and found no such ability: asked whether their own computation had been altered, none answered better than chance. To test it we built Open-Weight Masked Introspection (OWMI), a framework that intervenes on residual-stream sites, attention heads and sparse-autoencoder features, then interrogates the model about the change against the null conditions an answer has to beat: sham runs where nothing was altered, impact-matched random perturbations, and a text-only observer that sees only the visible output. Over 78,000 measurements, no model's report discriminates a real intervention from a sham beyond chance (AUROC ~0.5007), and an equivalence test bounds the effect below 0.15 percentage points of AUROC. Surprisingly, all the information needed is in the models. A model fine-tuned to report this class of intervention reaches near-perfect recovery on held-out directions, and a linear probe recovers intervention presence from the same activations at 75% to 95.8% accuracy, sharpening to no held-out error at the last layer before the model speaks. In one model the signal surfaces in the confidence rather than the words: its yes-or-no report never varies, while the confidence attached to it separates intervention from sham at AUROC 0.647. The failure sits in the path from internal state to verbal report, so oversight that reads a model's own testimony needs validating against an internal reference. While our results show the inability of current open-weight models to introspect, the debate is not settled for future models.
Chinese Translation
前沿模型是否能够对其内部状态进行内省?最近的研究表明,在某些条件下,复杂模型能够审计其内部状态,指出变化并自信地报告结果。我们对来自七个模型家族的八个开放权重模型进行了这一主张的测试,结果发现没有模型具备这种能力:在被询问其自身计算是否发生变化时,所有模型的回答均未超出随机猜测的水平。为了进行测试,我们构建了开放权重掩蔽内省(Open-Weight Masked Introspection, OWMI)框架,该框架干预残差流位置、注意力头和稀疏自编码器特征,然后针对变化对模型进行询问,比较其回答与需要超越的零假设条件:在没有任何改变的虚假运行、影响匹配的随机扰动,以及仅能看到可见输出的文本观察者。经过超过78,000次测量,没有一个模型的报告能够在真实干预与虚假干预之间区分出超出随机水平的结果(AUROC ~0.5007),而等效性测试将效果限制在AUROC低于0.15个百分点。令人惊讶的是,所需的所有信息都存在于模型中。一个经过微调以报告这一类干预的模型在保留方向上达到了近乎完美的恢复率,而线性探测器从相同的激活中以75%到95.8%的准确率恢复干预的存在,在模型发声前的最后一层没有保留错误。在一个模型中,信号体现在置信度而非词语上:它的是或否报告从未变化,而与之相关的置信度在AUROC 0.647的水平上将干预与虚假干预区分开。失败出现在从内部状态到语言报告的路径上,因此读取模型自身证词的监督需要与内部参考进行验证。虽然我们的结果显示当前开放权重模型无法进行内省,但对于未来模型的争论尚未结束。
cs.AI / 22 / 2608.20574
FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
FlavourBench:使用可执行的烹饪真实数据对前沿语言模型进行排名
Abstract
Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and family (14,418 model-task cells total), eliminating differential missingness from the leaderboard. The FlavourBench Score is the equal-family mean of the frozen task scores. We use 50,000 anchor-cluster bootstrap replicates for simultaneous 95% score bands and 100,000 sign-flip draws for all 351 paired model contrasts, with Holm control. The two independently compiled panels correlate at r = 0.89 (rank rho = 0.80). Grok 4.6 has the largest point estimate at 65.1 (simultaneous 95% CI 61.0-69.2); 101 of 351 model pairs are resolved. The release includes the prompts, all portfolio score maps, raw responses, exact routes, content hashes, and an offline verifier that reconstructs every result.
Chinese Translation
开放式语言模型基准通常依赖于评判者:人类偏好小组、另一个模型或脆弱的精确匹配键。我们介绍了FlavourBench,这是一个自动化基准,其中一个版本化的烹饪系统提供密集的、可执行的真实数据。每个任务提供八种成分,并要求提供一个由三种成分组成的组合;在模型执行之前,Epicure对所有56种可能的组合进行评分。我们在一个涵盖替代、配对和受限组合的534个核心任务上评估27个前沿端点。每个排名模型在每个小组和家族中都有89个有效响应(总共14,418个模型-任务单元),消除了排行榜中的差异缺失。FlavourBench分数是冻结任务分数的同家族均值。我们使用50,000个锚点聚类自助重复样本进行同时的95%分数区间,并使用100,000个符号翻转抽样进行所有351个配对模型对比,采用Holm控制。两个独立编制的小组的相关性为r = 0.89(排名rho = 0.80)。Grok 4.6的点估计最大,为65.1(同时的95%置信区间61.0-69.2);351对模型中有101对被解决。发布内容包括提示、所有组合分数图、原始响应、确切路径、内容哈希以及一个离线验证器,用于重建每个结果。
cs.AI / 23 / 2608.20611
Difficulty-Aware Semantic-ID Optimization for Generative Recommendation
基于难度感知的生成推荐语义标识优化
Abstract
Semantic-ID-based generative recommendation casts retrieval and ranking as autoregressive generation over hierarchical item identifiers. A common recipe is SFT followed by GRPO, yet vanilla GRPO is poorly matched to this tree-structured task. Under the frozen SFT checkpoint, the exact target is absent from the first 16 candidates of the 50-beam constrained ranking for many prompts, and in harder cases none of these candidates enters the target SID branch. This prompt-level diagnostic motivates a training concern: when on-policy GRPO groups are similarly target-missing, item-level rewards may produce weak or degenerate reward variation even if some candidates follow part of the target path. We propose Difficulty-Aware Semantic-ID Optimization (DASO), a tree-aware post-training method that addresses this failure mode as an online rollout-allocation problem. Instead of using fixed difficulty buckets or uniformly injecting ground-truth completions, DASO profiles each current rollout group by prefix-match depth, locates the bottleneck SID levels where candidates leave the target path, and reallocates a bounded portion of the group to prefix-guided completions while retaining raw rollouts for contrast. A SID-prefix reward provides graded credit, while an auxiliary SFT anchor mitigates regression on examples already solved by the SFT checkpoint. On the public benchmarks, DASO improves over MiniOneRec-style GRPO on 11 of 12 metrics and achieves the best result on 9 of 12 metrics; it also improves most level-wise recall metrics on the internal recommendation task.
Chinese Translation
基于语义标识的生成推荐将检索和排序视为对层次项目标识符的自回归生成。一种常见的方法是先进行SFT(Supervised Fine-Tuning),然后是GRPO(Generative Ranking Policy Optimization),然而,普通的GRPO与这一树状任务的匹配效果较差。在冻结的SFT检查点下,对于许多提示,前50束约束排序中的前16个候选项中缺少确切目标,而在更困难的情况下,这些候选项中没有一个进入目标SID(Semantic Identifier)分支。这种提示级别的诊断引发了一个训练问题:当策略内的GRPO组同样缺少目标时,项目级奖励可能会产生微弱或退化的奖励变化,即使一些候选项遵循了目标路径的一部分。我们提出了基于难度感知的语义标识优化(DASO),这是一种树状感知的后训练方法,旨在将这种失败模式视为在线回滚分配问题。DASO并不使用固定的难度桶或均匀注入真实完成,而是通过前缀匹配深度对每个当前回滚组进行分析,定位候选项离开目标路径的瓶颈SID级别,并将组的一部分重新分配给前缀引导的完成,同时保留原始回滚以供对比。SID前缀奖励提供分级信用,而辅助SFT锚点则缓解了在已经由SFT检查点解决的示例上的回归。在公共基准测试中,DASO在12个指标中的11个上优于MiniOneRec风格的GRPO,并在12个指标中的9个上取得最佳结果;它还改善了内部推荐任务中大多数层级召回指标。
cs.AI / 24 / 2608.20614
Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills
评估技能,而不仅仅是代理:代理技能的持续评估
Abstract
Enterprise agent programs are moving from prototypes into production, where reusable skills, tools, and workflow packages must be reviewed with evidence rather than prose. Current gates often scan these artifacts for structure, style, and security, but they do not answer the deployment question: does the capability package help a live agent complete enterprise tasks under the same model, sandbox, and grading policy? We present ACES (Agentic Continuous Evaluation of Skills), a repository-native framework for evaluating skills and product capability packages as executable agent artifacts. ACES runs paired live trials with and without a target skill, normalizes trajectories into the Agent Trajectory Interchange Format (ATIF), grades six default runtime metrics, and reports Skill Lift: the target skill's added value for a fixed task, harness, workspace, and scorer. The same protocol supports product-owned task suites that compare baseline, skill, bundle, team-skill, and plugin targets. On 145 real skills from internal enterprise repositories and public catalogs, scan-only gates surface useful authoring issues but measure complementary facets (structural versus LLM-judge Spearman $\rho = 0.14$). Across 947 scored paired cases from 58 of 64 production skills and four primary harnesses, mean composite Skill Lift is 0.2134 (95\% paired-case CI [0.1967, 0.2301]); mean outcome-only lift, the average of accuracy and goal accuracy, is 0.1799. Composite lift is positive in 72.8\% of paired cases. The largest process-metric gains appear in skill execution, behavior check, and skill efficiency---signals about discovery, routing, workflow following, and tool use that document scans cannot observe. An open-source implementation of the methodology is available in NVIDIA SkillEvaluator.
Chinese Translation
企业代理程序正从原型转向生产阶段,在这一过程中,必须以证据而非文字对可重用的技能、工具和工作流程包进行审查。目前的评估标准通常对这些工件进行结构、风格和安全性的扫描,但并未回答部署问题:能力包是否有助于在相同模型、沙箱和评分政策下,帮助实时代理完成企业任务?我们提出了ACES(代理技能的持续评估),这是一个原生于存储库的框架,用于将技能和产品能力包评估为可执行的代理工件。ACES通过有目标技能和无目标技能的配对实时试验,标准化轨迹为代理轨迹交换格式(ATIF),对六个默认运行时指标进行评分,并报告技能提升(Skill Lift):目标技能在固定任务、工具、工作空间和评分者中的附加价值。相同的协议支持产品拥有的任务套件,比较基线、技能、捆绑、团队技能和插件目标。在来自内部企业存储库和公共目录的145个真实技能中,仅扫描的评估标准揭示了有用的创作问题,但测量了互补的方面(结构性与LLM评判者的Spearman相关系数 $
ho = 0.14$)。在来自64个生产技能中的58个技能和四个主要工具的947个评分配对案例中,平均综合技能提升为0.2134(95\%配对案例置信区间 [0.1967, 0.2301]);仅结果提升,即准确性和目标准确性的平均值,为0.1799。综合提升在72.8 ext{%}的配对案例中为正值。技能执行、行为检查和技能效率的过程指标增益最大——这些信号关于发现、路由、工作流程遵循和工具使用,而文档扫描无法观察到。该方法的开源实现已在NVIDIA SkillEvaluator中提供。
cs.AI / 25 / 2608.20617
Dual-Cache Latent Space Communication between Heterogeneous Language Models
异构语言模型之间的双缓存潜在空间通信
Abstract
Multi-agent LLM systems split work across models, so answering often requires knowledge that sits in another agent's context: a Sharer has encoded information that a Receiver needs to complete its task. They usually communicate by exchanging text, which puts autoregressive decoding on the critical path and reduces the exchange to a discrete message written without sight of the receiver's state. Recent latent protocols instead translate the sharer's key-value (KV) cache into the receiver's: C2C supports heterogeneous models but requires both to read the same input, while LCF-X removes this shared-context requirement through position-free sharer-cache pooling. Three restrictions remain: LCF-X compresses the sharer alone, supplies the same layer-local summary to every receiver position with no joint cross-layer memory to retrieve from, and assumes matched layer count and KV geometry. We introduce XKV, which lifts all three: learned-query attention pools both caches; self-attention over receiver-aligned layer tokens, with a learned layer map reconciling different depths, mixes the pooled summaries into a compact joint memory; and a shared position decoder lets every raw receiver cache position retrieve its own per-head-gated residual in the receiver's native KV geometry. Both models stay frozen and may differ in family, depth, KV-head count, head dimension, and tokenizer; only the translator is trained. Across 45 dataset-model-pair settings (six heterogeneous and three same-model ordered pairings, five datasets), XKV attains the highest macro score and best average rank, improving on LCF-X on every dataset (by 4.6 exact-match and 4.2 F1 points on ROPES) and surpassing text communication on four of the five, while training 76% fewer parameters and translating a cache pair 10.3x faster (5.8 vs. 59.9 ms); end to end, XKV is 26% faster than LCF-X and 6.8x faster than text communication.
Chinese Translation
多智能体大语言模型(LLM)系统在模型之间分配工作,因此回答问题通常需要另一个智能体上下文中的知识:共享者(Sharer)编码了接收者(Receiver)完成任务所需的信息。它们通常通过交换文本进行通信,这使得自回归解码成为关键路径,并将交换简化为在未考虑接收者状态的情况下编写的离散消息。最近的潜在协议则将共享者的关键值(KV)缓存转换为接收者的缓存:C2C支持异构模型,但要求两者读取相同的输入,而LCF-X通过无位置的共享者缓存池消除了这一共享上下文要求。仍然存在三个限制:LCF-X仅压缩共享者,向每个接收者位置提供相同的层局部摘要,而没有联合跨层内存可供检索,并假设匹配的层数和KV几何结构。我们引入了XKV,解决了这三点:学习查询注意力池化两个缓存;在接收者对齐的层标记上进行自注意力,并通过学习的层映射调和不同深度,将池化的摘要混合成紧凑的联合内存;共享位置解码器让每个原始接收者缓存位置在接收者的本地KV几何中检索其自身的每头门控残差。两个模型保持冻结,并且在家族、深度、KV头数、头维度和分词器上可能不同;只有翻译器经过训练。在45个数据集-模型对设置(六个异构和三个相同模型的有序配对,五个数据集)中,XKV获得了最高的宏观得分和最佳平均排名,在每个数据集上都优于LCF-X(在ROPES上提高了4.6个精确匹配和4.2个F1点),并在五个数据集中的四个超越了文本通信,同时训练的参数减少了76%,缓存对的翻译速度提高了10.3倍(5.8毫秒对比59.9毫秒);从头到尾,XKV比LCF-X快26%,比文本通信快6.8倍。
cs.AI / 26 / 2608.20622
Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work
在大型企业中应用人类原语:知识工作的驾驭范式
Abstract
Frontier models have collapsed the cost of writing custom code: a niche problem a specialist sees in their own domain now costs an afternoon. The cost of reviewing and maintaining that code hasn't collapsed. Each solution drifts from the next; understanding one means reading its codebase from scratch. Large enterprises build something centrally governed instead: at worst an off-the-shelf product, at best a graph-orchestration framework wired bespoke per use case, or a low-code platform used as the orchestrator. These are custom every time and limited in scope. Enterprises don't weigh a third option that escapes both constraints: the harness paradigm. Recent work treats the coding-agent harness as enterprise infrastructure rather than a coding tool, converging on three findings: harnesses suffice at the task level and outperform more elaborate architectures on enterprise work (arXiv:2604.00073, arXiv:2604.13107); harness choice accounts for most of the variance in agent benchmark results, more than model choice does (arXiv:2605.23950); and the gap between that finding and enterprise adoption is governance (arXiv:2605.10223, arXiv:2605.18747). We propose an architecture that closes that gap. One harness runs unmodified as the backbone; the code stays identical across every deployment, so reviewing what gets built collapses to reading its instructions file. Section 4 gives four mechanisms: credential-scoped tooling, where each backend gets one generic request tool and a scoped credential instead of a hand-built method; authorization logic outside the harness, so one artifact runs as a cron backbone, a chat-surface engine, and a terminal tool; registration is a side effect of pushing code, collapsing an audit a review of a text file. Built on microcc (), our reference harness.
Chinese Translation
前沿模型降低了编写自定义代码的成本:一个专业人士在其领域内看到的细分问题现在只需一个下午的时间。审查和维护这些代码的成本并没有降低。每个解决方案都与下一个解决方案有所不同;理解一个解决方案意味着需要从头阅读其代码库。大型企业则构建一些集中管理的东西:在最坏的情况下是现成的产品,最好的情况是根据用例定制的图形编排框架,或作为编排者使用的低代码平台。这些每次都是定制的,且范围有限。企业没有考虑一个逃避这两种限制的第三种选择:驾驭范式。最近的研究将编码代理驾驭视为企业基础设施,而不是编码工具,得出了三个结论:驾驭在任务层面上足够,并且在企业工作中优于更复杂的架构(arXiv:2604.00073, arXiv:2604.13107);驾驭选择占据了代理基准结果大部分的方差,超过了模型选择的影响(arXiv:2605.23950);而这一发现与企业采用之间的差距是治理问题(arXiv:2605.10223, arXiv:2605.18747)。我们提出了一种架构来缩小这一差距。一个驾驭在未修改的情况下运行作为骨干;代码在每次部署中保持不变,因此审查所构建的内容简化为阅读其指令文件。第4节提供了四种机制:凭证范围工具,每个后端获得一个通用请求工具和一个范围凭证,而不是手动构建的方法;授权逻辑在驾驭外部,因此一个工件可以作为cron骨干、聊天界面引擎和终端工具运行;注册是推动代码的副作用,简化了对文本文件的审计和审查。我们的参考驾驭基于microcc()。
cs.AI / 27 / 2608.20630
SAGE: A Unified Algebra and Self-Adaptive Execution for AI Functions in SQL
SAGE:用于 SQL 中 AI 功能的统一代数与自适应执行
Abstract
SQL systems increasingly expose AI functions for tasks such as classification, extraction, filtering, ranking, retrieval, joining, and summarization. Despite their diverse APIs, these functions play only three relational roles: transforming individual rows, aggregating groups, or generating relationships between row pairs. We present SAGE (Self-Adaptive Generative Execution), a unified logical and physical framework that captures these roles with three typed primitives, AI_SCALAR, AI_AGG, and AI_JOIN, and composes them naturally with standard relational operators. All primitives share a confidence-gated execution interface while supporting physical strategies tailored to their relational shape. The main challenge is AI_JOIN, where SAGE analyzes the predicate, decomposes compound conditions when possible, and uses a recipe card together with a small label-free probe to select among complete execution strategies. Across a broad audit of public AI operators and evaluations spanning scalar, aggregate, and join workloads, this formulation covers common AI functionality while consistently improving execution quality and efficiency. SAGE achieves the strongest overall SemBench performance and, on a representative factorable join, reduces pairwise model calls by more than two orders of magnitude, yielding a 358-fold measured cost reduction.
Chinese Translation
SQL 系统越来越多地暴露出用于分类、提取、过滤、排序、检索、连接和摘要等任务的 AI 功能。尽管它们具有多样化的 API,这些功能仅扮演三种关系角色:转换单个行、聚合组或生成行对之间的关系。我们提出了 SAGE(自适应生成执行),这是一个统一的逻辑和物理框架,通过三种类型的原语 AI_SCALAR、AI_AGG 和 AI_JOIN 捕捉这些角色,并与标准关系运算符自然组合。所有原语共享一个基于置信度的执行接口,同时支持针对其关系形状量身定制的物理策略。主要挑战在于 AI_JOIN,SAGE 分析谓词,尽可能分解复合条件,并使用配方卡和一个小型无标签探针在完整执行策略中进行选择。在对公共 AI 操作符和涵盖标量、聚合和连接工作负载的评估进行广泛审计的过程中,这种表述涵盖了常见的 AI 功能,同时持续提高执行质量和效率。SAGE 实现了最强的 SemBench 整体性能,并且在一个具有代表性的可因式连接上,将成对模型调用减少了两个数量级以上,带来了 358 倍的测量成本降低。
cs.AI / 28 / 2608.20631
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents
加权记忆树:为长时间跨度的LLM代理记住重要信息
Abstract
Large language model (LLM) agents have demonstrated the ability to solve multi-step tasks requiring planning, tool use, and external information access, yet growing execution histories increase inference cost and expose reasoning to outdated, irrelevant, or misleading information, potentially degrading reasoning quality. Existing memory approaches organize or compress execution histories but provide limited mechanisms for deciding which memories remain active. We introduce the, a hierarchical memory system that organizes execution into tasks, subtasks, and actions while assigning each memory a dynamic retention score. Event-based updates and selection-based decay revise these scores, allowing WMT to preserve useful information, fold completed trajectories, suppress low-utility content, and retain access to folded context. We evaluate WMT on GAIA-Text using Qwen3-8B, Gemma 4 E4B, and Llama-3.1-8B, with ablations and memory-poisoning experiments. Relative to linear memory, WMT improves accuracy by an average of 9.97 percentage points while reducing prompt-token usage by 32.8%. Memory-poisoning experiments show that WMT limits the persistence and propagation of unreliable information. Our results suggest that effective long-horizon agent memory depends less on storing more information than on deciding which information should remain active.
Chinese Translation
大型语言模型(LLM)代理已展示出解决需要规划、工具使用和外部信息访问的多步骤任务的能力,但不断增长的执行历史增加了推理成本,并使推理暴露于过时、不相关或误导性的信息中,可能降低推理质量。现有的记忆方法组织或压缩执行历史,但在决定哪些记忆保持活跃方面提供的机制有限。我们提出了一种分层记忆系统——加权记忆树(Weighted Memory Tree, WMT),该系统将执行组织为任务、子任务和动作,同时为每个记忆分配一个动态保留分数。基于事件的更新和基于选择的衰减修订这些分数,使WMT能够保留有用信息、折叠已完成的轨迹、抑制低效用内容,并保持对折叠上下文的访问。我们在GAIA-Text上使用Qwen3-8B、Gemma 4 E4B和Llama-3.1-8B对WMT进行了评估,并进行了消融实验和记忆污染实验。相较于线性记忆,WMT的准确性平均提高了9.97个百分点,同时减少了32.8%的提示令牌使用。记忆污染实验表明,WMT限制了不可靠信息的持久性和传播。我们的结果表明,有效的长时间跨度代理记忆更依赖于决定哪些信息应保持活跃,而非存储更多信息。
cs.AI / 29 / 2608.20649
Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions
超越有效性:比较实用社会技术干预的多标准框架
Abstract
Designers and policymakers in sociotechnical domains like content moderation, privacy interfaces, recommender systems and beyond, must choose among a growing menu of proposed interventions, but typically lack a principled basis for comparing them. Prior work tends to evaluate interventions individually and mostly along the effectiveness criteria, while implementation constraints such as cost, effort and feasibility are often considered separately. We present a multi-criteria framework for evaluating sociotechnical interventions. This framework is instantiated through the case of misinformation, a domain of intense focus for proposed countermeasures. We survey $N=39$ researchers on 40 operationalized interventions across five evaluative criteria: political feasibility, effectiveness, user acceptance, cost, and implementation effort. We find that the interventions that experts judge to be the most effective are not always the most acceptable to the public or the most feasible to implement. We also discuss how this tension has implications for the design of sociotechnical interventions beyond misinformation, and offer a decision framework for practitioners navigating the trade-offs of sociotechnical interventions.
Chinese Translation
在内容审核、隐私界面、推荐系统等社会技术领域,设计师和政策制定者必须在日益增多的提议干预措施中进行选择,但通常缺乏一个原则性的基础来进行比较。以往的研究往往单独评估干预措施,主要基于有效性标准,而实施约束如成本、努力和可行性则常常被单独考虑。我们提出了一个用于评估社会技术干预的多标准框架。该框架通过对虚假信息的案例进行实例化,这是一个针对提议反制措施的重点领域。我们对39名研究人员进行了调查,涉及40项在五个评估标准下操作化的干预措施:政治可行性、有效性、用户接受度、成本和实施努力。我们的研究发现,专家认为最有效的干预措施并不总是最被公众接受或最具可行性的。我们还讨论了这种紧张关系对超越虚假信息的社会技术干预设计的影响,并为在社会技术干预的权衡中导航的从业者提供了一个决策框架。
cs.AI / 30 / 2608.20661
Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance
可审计的构建:一个面向企业金融可信赖大型语言模型分析的本体驱动框架
Abstract
Enterprise adoption of large language models in finance is constrained less by fluency than by trust: in Financial Planning and Analysis (FP&A) and other regulated workflows, an answer is usable only if it is traceable to authoritative sources and auditable after the fact. This paper argues that retrieval-augmented generation for enterprise finance should be evaluated on auditability alongside accuracy, and presents the Knowledge-Driven Analytics Framework (KDAF), which builds ontology-driven knowledge systems through six iterative stages and retrieves evidence via Context-Aware Relevance Propagation (CARP), so that every retrieved fact carries its relationship type, confidence, and source lineage. An evaluation on FinanceBench (145 questions) compares KDAF against zero-context inference, BM25, concept-weighted lexical retrieval, and ungrounded graph traversal. First, retrieval is necessary: zero-context inference reaches 4.1% correctness against 10-12% for retrieval-augmented conditions. Second, on answer correctness the retrieval conditions are statistically indistinguishable (KDAF vs BM25: -0.007, 95% CI [-0.021, 0.000]), so accuracy alone does not justify structured retrieval here -- a negative result we report explicitly. Third, on auditability the ordering reverses: KDAF attains the highest citation traceability F1 (0.515), exceeding ungrounded traversal by +0.027 (CI [0.006, 0.050]) and BM25 by +0.052 (CI [0.024, 0.083]), intervals excluding zero. Graph-structured retrieval also admits no evidence from outside the question subject entity (0 of 426 items, against 16.8% and 20.2% for lexical baselines), and every selected item resolves to a complete provenance chain. We argue that auditability, not accuracy, is the axis on which ontology-grounded retrieval earns its cost.
Chinese Translation
企业在金融领域采用大型语言模型的限制因素主要是信任,而非流畅性:在财务规划与分析(FP&A)及其他受监管的工作流程中,只有能够追溯到权威来源并在事后可审计的答案才是可用的。本文认为,企业金融中的检索增强生成应在可审计性与准确性两方面进行评估,并提出知识驱动分析框架(Knowledge-Driven Analytics Framework, KDAF),该框架通过六个迭代阶段构建本体驱动的知识系统,并通过上下文感知相关性传播(Context-Aware Relevance Propagation, CARP)检索证据,从而使每个检索到的事实都携带其关系类型、置信度和来源沿袭。在FinanceBench(145个问题)上的评估比较了KDAF与零上下文推理、BM25、概念加权词汇检索和无基础图遍历的表现。首先,检索是必要的:零上下文推理的正确率为4.1%,而检索增强条件下为10-12%。其次,在答案正确性方面,检索条件之间在统计上没有显著差异(KDAF与BM25:-0.007,95%置信区间[-0.021, 0.000]),因此仅凭准确性并不能证明结构化检索的合理性——这是我们明确报告的一个负面结果。第三,在可审计性方面,排序发生了逆转:KDAF达到了最高的引用可追溯性F1值(0.515),超过了无基础遍历+0.027(置信区间[0.006, 0.050])和BM25+0.052(置信区间[0.024, 0.083]),这些区间均不包含零。图结构检索也未能提供来自问题主题实体之外的证据(426个项目中无,词汇基线的比例分别为16.8%和20.2%),且每个选定项目都解析为完整的来源链。我们认为,可审计性而非准确性是本体基础检索获得其成本合理性的关键轴心。
cs.AI / 31 / 2608.20664
DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents
DreamBench-SWE:一个针对软件代理的多会话内存卫生基准测试
Abstract
DreamBench-SWE is a multi-session benchmark for software-agent memory hygiene in which later software tasks depend on non-inferable evidence from earlier sessions and are scored by executable hidden oracles. We report the original scaled v2 fold and a separately preregistered v2.1 successor audit designed after that study but frozen before successor outcome inspection. The successor run completed 360/360 work units and 720/720 S3 cells across four conditions. In the original fold, the primary DF-hybrid--B5 contrast was null (95/180 versus 89/180; clustered p=.518, Holm p=1), not evidence of equivalence, and C9/C10 retained B0-headroom limitations. In the successor, no external memory achieved 21/180 passes (rate 0.1167), deterministic verbatim event memory 82/180 (rate 0.4556), the typed-plus-raw reference probe 83/180 (rate 0.4611), and one pinned hosted Mem0 literal-storage configuration 97/180 (rate 0.5389). The registered six-slot Family A retained unavailable slots at p=1; all three available comparisons against no memory rejected after Holm correction. Both preregistered mechanism contrasts were unavailable after pre-evaluation conformance rejection. The secondary literal-storage-versus-verbatim comparison was nonconfirmatory and sensitivity-dependent, while the comparison with the reference probe did not reject. The audit therefore supports DreamBench-SWE as a discriminating executable profile benchmark and characterizes one exact hosted-memory configuration, but it does not establish an external-system mechanism, superiority among memory-bearing conditions, equivalence, or broad product generality. The original v2.0.5 findings and artifacts remain unchanged.
Chinese Translation
DreamBench-SWE 是一个针对软件代理内存卫生的多会话基准测试,其中后续软件任务依赖于早期会话中不可推断的证据,并由可执行的隐藏神谕进行评分。我们报告了原始的缩放 v2 版本和一个单独预注册的 v2.1 后续审计,该审计是在该研究之后设计的,但在后续结果检查之前已被冻结。后续运行完成了 360/360 个工作单元和 720/720 个 S3 单元,涵盖四种条件。在原始版本中,主要的 DF-hybrid--B5 对比结果为零(95/180 对 89/180;聚类 p=.518,Holm p=1),这并不证明等效性,而 C9/C10 保留了 B0 头部限制。在后续测试中,没有外部内存达到 21/180 的通过率(比率 0.1167),确定性逐字事件内存为 82/180(比率 0.4556),类型加原始参考探针为 83/180(比率 0.4611),以及一个固定的托管 Mem0 文字存储配置为 97/180(比率 0.5389)。注册的六槽 A 类保留了不可用槽位,p=1;与无内存的所有三个可用比较在 Holm 校正后被拒绝。两个预注册的机制对比在预评估一致性拒绝后均不可用。次要的文字存储与逐字比较结果不具确认性且依赖于敏感性,而与参考探针的比较未被拒绝。因此,该审计支持 DreamBench-SWE 作为一个区分性的可执行配置基准,并描述了一个确切的托管内存配置,但并未建立外部系统机制、内存条件之间的优越性、等效性或广泛的产品通用性。原始 v2.0.5 的发现和文物保持不变。
cs.AI / 32 / 2608.20670
Why2Speak: Faithful Reasoning for Abstaining Action Policies
Why2Speak:忠实推理在弃权行动政策中的应用
Abstract
Many agentic systems must repeatedly choose between acting and abstaining, making faithful reasoning important for oversight: an explanation is useful only if it reflects the computation that produced the action. We study this problem through intervention timing in multi-party conversation, where an assistant must decide whether to speak or remain silent. This setting exposes class imbalance, asymmetric action costs, and the possibility that exposing reasoning changes the policy being audited. Using Qwen3-8B, decoded with or without chain-of-thought reasoning, we compare direct decision policies, reasoning policies, supervised fine-tuning, and reinforcement learning. We find a capability-auditability tradeoff: the strongest direct policy achieves higher quality but exposes no reasoning to inspect, while the reasoning policy provides a trace at the cost of lower performance, particularly recall of true intervention opportunities. Supervised fine-tuning either suppresses reasoning or preserves it without improving decision quality, while reinforcement learning also fails to improve the reasoning policy. We identify one mechanism underlying this failure: group relative objectives provide no learning signal on confidently wrong prompts when sampled rollouts all select the same action. Controlled activation probes and behavioral ablations show that standard faithfulness methods can overstate evidence that exposed reasoning reflects the underlying decision process. Probability-based metrics saturate under confident decisions, probes are vulnerable to class imbalance and textual leakage, and reasoning ablations can confound reasoning content with changes in inference mode. Together, these results show that exposing reasoning can change an agent's action policy rather than simply make it observable. We provide controls for evaluating reasoning-based oversight of agents that can act or abstain.
Chinese Translation
许多智能系统必须反复在行动和弃权之间做出选择,因此忠实推理对于监督至关重要:只有当解释反映出产生该行动的计算时,它才是有用的。我们通过多方对话中的干预时机研究这一问题,在该场景中,助手必须决定是发言还是保持沉默。这个设置暴露了类别不平衡、非对称的行动成本,以及暴露推理可能改变被审计政策的可能性。使用 Qwen3-8B,在有或没有链式推理的情况下进行解码,我们比较了直接决策政策、推理政策、监督微调和强化学习。我们发现能力与可审计性之间存在权衡:最强的直接政策实现了更高的质量,但没有暴露可供检查的推理,而推理政策则提供了可追溯性,但以性能降低为代价,特别是在真实干预机会的召回方面。监督微调要么抑制推理,要么在不提高决策质量的情况下保留推理,而强化学习也未能改善推理政策。我们识别出导致这一失败的一个机制:当采样回合都选择相同的行动时,群体相对目标在自信错误的提示上没有提供学习信号。受控激活探针和行为切除实验表明,标准的忠实性方法可能会夸大暴露推理反映基础决策过程的证据。在自信决策下,基于概率的指标饱和,探针容易受到类别不平衡和文本泄漏的影响,而推理切除可能会将推理内容与推理模式的变化混淆在一起。综合这些结果表明,暴露推理可能会改变智能体的行动政策,而不仅仅是使其可观察。我们提供了评估能够行动或弃权的智能体的推理基础监督的控制方法。
cs.AI / 33 / 2608.20686
CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery
CDRL:基于认证驱动的强化学习用于中微子味道模型发现
Abstract
Many scientific discovery problems require searching combinatorial hypothesis spaces under complex domain constraints. Reinforcement learning (RL) offers a promising approach, but existing methods rely on scalar rewards that provide limited information about why candidate solutions fail, leading agents to repeatedly explore invalid regions. We introduce Certification-Driven Reinforcement Learning (CDRL), a framework that leverages structured feedback from symbolic reasoning tools. When a candidate violates domain constraints, these tools produce certificates identifying the actions responsible for failure. CDRL converts these certificates into reusable constraints that eliminate classes of invalid solutions and guide exploration toward valid regions. We evaluate CDRL on neutrino flavor model discovery in theoretical particle physics, where the hypothesis space exceeds $10^{26}$ possible models, and compare it with the state-of-the-art RL approach previously used for this task. Across three theory spaces, CDRL achieves up to 1.95$\times$ higher valid model rates and up to 6.33$\times$ higher neutrino model rates while evaluating up to 4$\times$ fewer candidates. We further extract 40 interpretable rules from search trajectories using a post-hoc decision-tree framework and show that reusing them as soft constraints yields gains of up to 2$\times$ in valid model rates and 3$\times$ in neutrino model discovery across all three theory spaces. These results suggest that CDRL uncovers reusable structure in combinatorial search spaces and provides a general framework for scientific model discovery.
Chinese Translation
许多科学发现问题需要在复杂领域约束下搜索组合假设空间。强化学习(RL)提供了一种有前景的方法,但现有方法依赖于标量奖励,这些奖励提供了有限的信息来解释候选解决方案失败的原因,从而导致智能体反复探索无效区域。我们提出了认证驱动的强化学习(CDRL),这是一个利用符号推理工具提供的结构化反馈的框架。当候选方案违反领域约束时,这些工具会生成证书,识别导致失败的行动。CDRL将这些证书转换为可重用的约束,从而消除无效解决方案的类别,并引导探索朝向有效区域。我们在理论粒子物理学中的中微子味道模型发现上评估了CDRL,其中假设空间超过$10^{26}$个可能模型,并将其与之前用于此任务的最先进的RL方法进行了比较。在三个理论空间中,CDRL实现了高达1.95倍的有效模型率和高达6.33倍的中微子模型率,同时评估的候选数量减少了最多4倍。我们进一步使用后验决策树框架从搜索轨迹中提取了40条可解释规则,并表明将其作为软约束重用可在所有三个理论空间中实现有效模型率提高最多2倍和中微子模型发现提高最多3倍。这些结果表明,CDRL在组合搜索空间中发现了可重用的结构,并为科学模型发现提供了一个通用框架。
cs.AI / 34 / 2608.20688
VortexChat: An agentic framework for autonomous multi-objective integrated photonic design
VortexChat:一种用于自主多目标集成光子设计的智能框架
Abstract
The advancement of modern integrated photonics is frequently bottlenecked by device design workflows that rely heavily on manual simulation and expert intuition. While inverse design offers an alternative, it remains constrained by expert supervision and a lack of end-to-end automation. To address these issues, we present VortexChat, an agentic framework for the autonomous, end-to-end inverse design of integrated photonic devices directly from natural language specifications. VortexChat couples a large language model (LLM) decision agent with topology generation, gradient-based refinement, and full-wave electromagnetic simulation. This closed-loop architecture enables the system to iteratively decompose design objectives, orchestrate computational tools, and update strategies based on feedback with minimal human intervention. Constrained by the absolute metrics of the Vortex100 Benchmark, VortexChat autonomously generates devices that strictly meet all predefined performance thresholds without any human-in-the-loop. As an experimental demonstration, we fabricated a broadband terahertz perfect vortex beam multiplexer, autonomously designed by VortexChat, with measurements confirming high-efficiency operation, high mode purity and low inter-channel crosstalk in agreement with full-wave simulations. These results demonstrate that an LLM agent can assume key aspects of expert decision-making in photonic inverse design while maintaining physical fidelity and fabrication feasibility, providing a scalable route towards autonomous design of complex integrated photonic systems.
Chinese Translation
现代集成光子学的进步常常受到依赖于手动模拟和专家直觉的设备设计工作流程的瓶颈。虽然逆向设计提供了一种替代方案,但仍然受到专家监督和缺乏端到端自动化的限制。为了解决这些问题,我们提出了VortexChat,这是一种用于从自然语言规范直接进行集成光子设备自主端到端逆向设计的智能框架。VortexChat将大型语言模型(LLM)决策代理与拓扑生成、基于梯度的优化和全波电磁仿真相结合。这种闭环架构使系统能够迭代地分解设计目标,协调计算工具,并根据反馈更新策略,几乎不需要人类干预。在Vortex100基准的绝对指标限制下,VortexChat自主生成严格满足所有预定义性能阈值的设备,而无需任何人类参与。作为实验演示,我们制造了一种宽带太赫兹完美涡旋光束多路复用器,该器件由VortexChat自主设计,测量结果确认其高效运行、高模态纯度和低通道间串扰,与全波仿真结果一致。这些结果表明,LLM代理可以在光子逆向设计中承担专家决策的关键方面,同时保持物理真实性和制造可行性,为复杂集成光子系统的自主设计提供了一条可扩展的途径。
cs.AI / 35 / 2608.20717
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning
DirEAG:用于校准数学推理中语言化信心的狄利克雷证据聚合
Abstract
Reliable confidence estimation is essential for using large language models in mathematical reasoning, but black-box verbalized confidence is difficult to calibrate. When the same problem is queried under multiple confidence-steering prompts, the resulting answer-confidence observations contain useful uncertainty information, yet their scales may shift across steering levels, models, and datasets. Existing black-box uncertainty methods often rely on answer agreement, sample consistency, or entropy, which describe output variation but do not model the numerical meaning of self-reported confidence. Conversely, direct averaging or heuristic aggregation of elicited confidence cannot learn prompt- and task-dependent bias. We propose DirEAG, a Dirichlet Evidence Aggregation method that converts each elicited answer-confidence observation into calibrated soft evidence over generated candidate answers and an additional null state, allowing the model to represent cases where none of the candidates is correct. Experiments on GSM8K, SVAMP, and GSM-Hard with Qwen, Mistral, and Gemma models show that, compared with direct confidence averaging and heuristic confidence-steering aggregation, DirEAG often achieves better calibration while maintaining competitive answer selection. Ablations further reveal that evidence aggregation and final binary calibration address distinct parts of the calibration problem.
Chinese Translation
可靠的信心估计对于在数学推理中使用大型语言模型至关重要,但黑箱语言化信心的校准较为困难。当在多个信心引导提示下查询同一问题时,得到的答案-信心观察包含有用的不确定性信息,但其尺度可能在不同的引导级别、模型和数据集之间发生变化。现有的黑箱不确定性方法通常依赖于答案一致性、样本一致性或熵,这些方法描述了输出的变化,但并未建模自我报告信心的数值含义。相反,直接平均或启发式聚合所引导的信心无法学习与提示和任务相关的偏差。我们提出了DirEAG,一种狄利克雷证据聚合方法,它将每个引导的答案-信心观察转换为对生成候选答案和一个额外的空状态的校准软证据,从而使模型能够表示没有候选答案正确的情况。在GSM8K、SVAMP和GSM-Hard数据集上使用Qwen、Mistral和Gemma模型的实验表明,与直接信心平均和启发式信心引导聚合相比,DirEAG通常能够实现更好的校准,同时保持竞争性的答案选择。消融实验进一步表明,证据聚合和最终的二元校准解决了校准问题的不同部分。
cs.AI / 36 / 2608.20729
Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol
大型语言模型代理中的标准修订校准:失败模式与追踪锚定协议
Abstract
Language-model agents can improve after failure or carry text across episodes without revising what counts as success. We study the narrower attribution problem of criterion revision: when criterion K0 accepts an outcome violating a broader commitment B, what observations justify saying that the system formed and persistently used K1? We require five non-compensatory conditions: criterion-failure detection, a model-emitted proposal, new-episode transfer, intervention sensitivity on the claimed carrier, and preservation. We evaluate CMB-0.1 on twelve cross-domain cases and four arms: stateless inference, append-only history, model-generated but harness-committed state, and evaluator-written oracle state. Seven mechanism fixtures yield 84 deterministic scorer trials; four local quantized artifacts yield 96 calls and 192 model-case-arm trials. No model trial satisfies all five conditions, but this zero does not establish general capability absence. Eleven calls remain invalid after one retry; several commitments disclose the target distinction; the harness performs commits; deletion reuses a stateless call; and conflict changes multiple factors. Qwen2.5-7B answers every transfer and preservation item without revision state, exposing zero-state reconstruction. These failures make CMB-0.1 an instrument-calibration result rather than a model ranking. We derive a prospective, trace-anchored CMB-0.4 protocol requiring concealed transfer, explicit WRITE/NO-WRITE/ESCALATE actions, a separately logged policy-selected commit, matched interventions, repeated hidden items, and a frozen executable oracle. It is a successor design, not a completed confirmatory result. The paper contributes a measurement chain, an empirical diagnosis of its first implementation, and a more discriminating protocol for future tests of criterion revision.
Chinese Translation
语言模型代理在失败后可以改进,或在不修订成功标准的情况下跨情节传递文本。我们研究了更狭义的归因问题,即标准修订:当标准 K0 接受一个违反更广泛承诺 B 的结果时,哪些观察结果可以证明系统形成并持续使用 K1?我们要求五个非补偿条件:标准失败检测、模型发出的提议、新情节转移、对声称载体的干预敏感性以及保存。我们在十二个跨领域案例和四个方面评估 CMB-0.1:无状态推理、仅附加历史、模型生成但承诺状态的状态,以及评估者编写的神谕状态。七个机制固定装置产生了 84 次确定性评分试验;四个局部量化工件产生了 96 次调用和 192 次模型案例-臂试验。没有模型试验满足所有五个条件,但这个零并不建立一般能力缺失。经过一次重试后,十一项调用仍然无效;几个承诺揭示了目标区分;承载器执行承诺;删除重用无状态调用;冲突改变多个因素。Qwen2.5-7B 在没有修订状态的情况下回答了每个转移和保存项目,暴露出零状态重构。这些失败使得 CMB-0.1 成为一个工具校准结果,而不是模型排名。我们推导出一个前瞻性的、追踪锚定的 CMB-0.4 协议,要求隐蔽转移、明确的写入/不写入/升级操作、单独记录的政策选择承诺、匹配的干预、重复的隐藏项目和一个冻结的可执行神谕。这是一个继任设计,而不是一个完成的确认结果。本文贡献了一条测量链、其首次实施的实证诊断,以及一个更具区分性的标准修订未来测试协议。
cs.AI / 37 / 2608.20735
ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation
ForeTime-VLA:来自世界动作模型的因果未来标记蒸馏用于传送带操作
Abstract
Manipulating moving objects requires a policy to anticipate contact events, yet vision-language-action (VLA) policies are commonly fine-tuned from the current observation alone. World action models (WAMs) learn predictive dynamics, but running a video-scale teacher or explicitly imagining future frames at deployment is costly. We introduce ForeTime-VLA, a dense pi0.5 policy that distills a future-aware, action-equivalent representation from a frozen Fast-WAM-derived teacher while remaining causal at inference. Offline, current and future video latents are compressed into a whitened 64-D target. Online, an eight-frame history encoder predicts this target together with manipulation phase and normalized time-to-transition. Four future tokens and one phase token condition the VLM prefix, while the predicted future and transition horizon condition the action expert. Training retains the original flow-matching action target and adds cosine, relational geometry, phase, time-to-transition, and action-equivalence objectives. On a deduplicated conveyor-belt dataset, we compare 40k-step checkpoints on 768 matched windows per split. Test MAE decreases from 0.134119 to 0.130593 (2.63%; paired-bootstrap 95% CI: 0.82-4.48% improvement), and test L2 decreases by 3.02%, at a 2.46-2.93% latency cost. In quantitative real-robot evaluation, ForeTime-VLA achieves 81.1% stationary and 58.9% slow-moving grasp success, exceeding the next-best reference by 12.2 and 22.2 percentage points, respectively. Across three belt speeds, it completes 44/90 grasps versus 23/90 for pi0.5, including 11/30 versus 2/30 at fast speed. The agreement between offline orientation gains and reduced real-robot contact-pose failures supports causal future-token distillation as an effective way to improve dynamic manipulation without deploying the world-model teacher.
Chinese Translation
操控移动物体需要一个政策来预测接触事件,然而视觉-语言-动作(VLA)政策通常仅从当前观察中进行微调。世界动作模型(WAMs)学习预测动态,但在部署时运行视频规模的教师或明确想象未来帧是成本高昂的。我们提出了ForeTime-VLA,这是一种密集的pi0.5政策,它从冻结的Fast-WAM派生教师中蒸馏出一种未来感知的、动作等效的表示,同时在推理时保持因果性。在离线阶段,当前和未来的视频潜变量被压缩成一个白化的64维目标。在在线阶段,一个八帧历史编码器预测这个目标,同时考虑操作阶段和归一化的过渡时间。四个未来标记和一个阶段标记条件化VLM前缀,而预测的未来和过渡视野条件化动作专家。训练保留了原始的流匹配动作目标,并增加了余弦、关系几何、阶段、过渡时间和动作等效目标。在一个去重的传送带数据集中,我们比较了每个分割768个匹配窗口的40k步检查点。测试MAE从0.134119降低到0.130593(2.63%;配对自助法95%置信区间:0.82-4.48%的改善),测试L2降低了3.02%,延迟成本为2.46-2.93%。在定量真实机器人评估中,ForeTime-VLA实现了81.1%的静止和58.9%的慢速抓取成功率,分别超过下一个最佳参考12.2和22.2个百分点。在三种传送带速度下,它完成了44/90次抓取,而pi0.5仅完成了23/90次,包括在快速速度下11/30次对比2/30次。离线方向增益与减少的真实机器人接触姿态失败之间的一致性支持因果未来标记蒸馏作为一种有效的方式来改善动态操作,而无需部署世界模型教师。
cs.AI / 38 / 2608.20738
Continuous-Time Quantum Walks based Graph Neural Network
基于连续时间量子行走的图神经网络
Abstract
Graph Neural Networks (GNNs) are widely used on graph-structured data, but most suffer from two key weaknesses. First, message passing behaves as a low-pass filter under the homophily assumption, leading to poor performance on heterophilic graphs. Second, stacking layers drives node features toward constants, causing over-smoothing. Existing methods usually address these issues separately, while the few joint solutions rely largely on empirical heuristics, and many over-smoothing remedies sacrifice model expressiveness. We propose \textbf{CTQW-GNN}, a GNN based on Continuous-Time Quantum Walks (CTQW), to address both issues with theoretical justification. Its design exploits two properties of the CTQW propagator $e^{-\mathrm{i}Ht}$. First, it is unitary and has eigenvalues on the unit circle, so no frequency component is damped, counteracting the low-pass bias. Second, unitarity preserves feature norms and prevents the Dirichlet energy from decaying exponentially with depth, thereby mitigating over-smoothing. CTQW-GNN combines three complementary aggregation modules. \textit{CTQW-based Aggregation} evolves node features through the unitary propagator, preserving mid- and high-frequency signals for heterophilic graphs while preventing Dirichlet-energy collapse. \textit{CTQW-Attention Aggregation} constructs a multi-hop neighbor graph from CTQW amplitudes and applies attention over it, enabling access to distant homophilic nodes missed by single-hop aggregation. \textit{LF Aggregation} uses a standard low-pass GAT branch to retain strong performance on homophilic graphs, where pure CTQW aggregation can be suboptimal. We further provide a spectral-gap analysis explaining energy preservation and a Lieb--Robinson-type bound that gives a principled rule for selecting the walk time $t$.
Chinese Translation
图神经网络(GNNs)广泛应用于图结构数据,但大多数存在两个主要弱点。首先,在同质性假设下,消息传递表现为低通滤波器,导致在异质图上的性能较差。其次,层叠会使节点特征趋向常数,造成过平滑。现有方法通常分别解决这些问题,而少数联合解决方案主要依赖经验启发式,许多过平滑的补救措施牺牲了模型的表达能力。我们提出了 extbf{CTQW-GNN},一种基于连续时间量子行走(CTQW)的GNN,以理论依据解决这两个问题。其设计利用了CTQW传播算子$e^{- ext{i}Ht}$的两个特性。首先,它是单位算子,特征值位于单位圆上,因此没有频率成分被衰减,从而抵消了低通偏差。其次,单位性保持特征范数,防止狄利克雷能量随着深度的增加而指数衰减,从而减轻过平滑。CTQW-GNN结合了三个互补的聚合模块。 extit{基于CTQW的聚合}通过单位传播算子演化节点特征,保留异质图的中高频信号,同时防止狄利克雷能量崩溃。 extit{CTQW-注意力聚合}从CTQW幅度构建多跳邻居图,并在其上应用注意力机制,使得能够访问单跳聚合未能捕捉到的远程同质节点。 extit{LF聚合}使用标准低通GAT分支,在同质图上保持强性能,而纯CTQW聚合可能表现不佳。我们进一步提供了谱间隙分析,解释能量保持,并给出了一种Lieb-Robinson型界限,为选择行走时间$t$提供了原则性规则。
cs.AI / 39 / 2608.20743
Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis
多模态推测解码是否为基于扩散的并行草拟做好准备?一项调查与实证诊断
Abstract
Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel. Its lossless guarantee has motivated a line of work that pushes the drafter itself toward parallel generation. The most recent paradigm is block-parallel generative drafting, including diffusion-based methods such as DFlash and DSpark, achieving up to 3.6x speedup on common daily chatting tasks. While this transition is well studied in text-only LLMs, its applicability to multimodal models remains an open question. Existing multimodal speculative decoding efforts focus on input compression, adapter alignment, candidate coverage, or modality-specific verification; however, block-parallel generative drafting remains largely unexplored. To bridge this gap, this paper combines a modality-centered survey with a cross-architecture empirical study to ask: Is multimodal speculative decoding ready for diffusion-based parallel drafting? In this survey, we systematically analyze a wide spectrum of multimodal models, spanning Vision-Language, Video-Language, Audio, and Vision-Language-Action (VLA) architectures, from the dual perspectives of drafting parallelism and cross-modal information interaction. We introduce a unified taxonomy that isolates drafter-side parallelism from orthogonal design choices such as tree construction and verification strategies. Furthermore, we provide a comprehensive empirical comparison of existing methods under varying degrees of parallelism across standardized multimodal benchmarks, including OCR, VQA, visual reasoning, and image captioning. Finally, we summarize the limitations of current approaches, discuss open challenges, and outline promising future directions for this rapidly evolving field.
Chinese Translation
推测解码通过允许轻量级草拟者在目标模型并行验证的同时提出未来的标记,从而加速自回归生成。其无损保证激励了一系列研究,推动草拟者本身朝向并行生成。最新的范式是块并行生成草拟,包括基于扩散的方法,如 DFlash 和 DSpark,在常见的日常聊天任务中实现了高达 3.6 倍的加速。虽然这一转变在仅文本的 LLM 中得到了充分研究,但其在多模态模型中的适用性仍然是一个未解的问题。现有的多模态推测解码工作集中于输入压缩、适配器对齐、候选覆盖或特定模态的验证;然而,块并行生成草拟仍然基本未被探索。为了解决这一空白,本文结合了以模态为中心的调查与跨架构的实证研究,提出以下问题:多模态推测解码是否为基于扩散的并行草拟做好准备?在这项调查中,我们系统地分析了一系列多模态模型,涵盖视觉-语言、视频-语言、音频以及视觉-语言-动作(VLA)架构,从草拟并行性和跨模态信息交互的双重视角出发。我们引入了一个统一的分类法,将草拟者侧的并行性与树构建和验证策略等正交设计选择区分开来。此外,我们提供了在标准化多模态基准(包括 OCR、VQA、视觉推理和图像描述)下,现有方法在不同并行程度下的全面实证比较。最后,我们总结了当前方法的局限性,讨论了开放挑战,并概述了这一快速发展的领域的有前景的未来方向。
cs.AI / 40 / 2608.20755
Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design
基于自然语言指导的生成器无关的蛋白质结合剂设计短名单筛选
Abstract
Modern de novo design workflows generate many candidate protein binders, but wet-lab validation capacity remains limited, making shortlisting a major bottleneck. We study whether LLMs can generate multi-metric ranking policies from precomputed structural-confidence and interface-quality proxy scores. Rather than proposing a new protein binder design pipeline, we focus on post-generation binder shortlisting: selecting the final top-K candidates from already generated binder pools using a shared panel of precomputed proxy scores. On the 10-target held-out split, averaging performance over five sampled global iterative gpt-4o policies reaches 0.589 Recall@10, modestly improving over the strongest single-feature fixed baseline, Protenix binder ipTM, which reaches 0.571 Recall@10. On the 3-target held-out subset comprising Nipah, RBX1, and TREM2, target-conditioned iterative gpt-5.4 policies reach the strongest LLM performance, with 0.519 Recall@10 and 0.583 NDCG@10. These results suggest that LLM-generated ranking policies can act as an interpretable post-generation decision layer for combining heterogeneous proxy metrics to prioritize binders from large candidate pools.
Chinese Translation
现代的从头设计工作流程生成了许多候选蛋白质结合剂,但湿实验室验证能力仍然有限,使得短名单筛选成为一个主要瓶颈。我们研究了大型语言模型(LLMs)是否能够从预计算的结构置信度和界面质量代理评分中生成多指标排名策略。我们并不提出一个新的蛋白质结合剂设计流程,而是专注于生成后的结合剂短名单筛选:从已经生成的结合剂池中使用一组共享的预计算代理评分选择最终的前K个候选者。在10个目标的保留分割上,五个采样的全局迭代gpt-4o策略的平均性能达到0.589 Recall@10,较最强的单特征固定基线Protenix结合剂ipTM(达到0.571 Recall@10)有适度改善。在包含Nipah、RBX1和TREM2的3个目标保留子集中,目标条件的迭代gpt-5.4策略达到了最强的LLM性能,分别为0.519 Recall@10和0.583 NDCG@10。这些结果表明,LLM生成的排名策略可以作为一个可解释的生成后决策层,用于结合异质代理指标,以优先选择来自大型候选池的结合剂。
cs.AI / 41 / 2608.20768
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization
超越端点增益:医疗专业化的权重差异审计
Abstract
Specialist language models are usually understood through endpoint gains: the generalist scores lower, the specialist scores higher, and the difference is treated as evidence of specialization. This leaves the released update itself largely unexamined. We propose a paired weight-delta path audit and apply it to two public, aligned generalist-to-medical-specialist checkpoint pairs: Gemma-3-4B-IT to MedGemma-4B-IT and Qwen2.5-7B-Instruct to HuatuoGPT-o1-7B. In both pairs, the full decoder-side update strongly reconstructs measured medical benchmark movement (0.974 and 1.183 endpoint-normalized retention), making each decoder delta an appropriate substrate for the audit. Yet the movement is not cleanly localized. MLP is the strongest broad component family in both pairs, but mixed off-domain movements, 10-seed matched controls, and endpoint-anchored rollbacks prevent a unique coarse-family explanation. The audit therefore separates update-level reconstruction from component-level explanation. Its claims concern text-only multiple-choice benchmark movement, not clinical validation, repair, or circuit-level mechanism.
Chinese Translation
专业语言模型通常通过端点增益来理解:通用模型得分较低,专业模型得分较高,二者之间的差异被视为专业化的证据。这使得发布的更新本身在很大程度上未被审视。我们提出了一种配对权重差异路径审计,并将其应用于两个公共的、对齐的通用到医疗专业模型的检查点对:Gemma-3-4B-IT 到 MedGemma-4B-IT 和 Qwen2.5-7B-Instruct 到 HuatuoGPT-o1-7B。在这两个模型对中,完整的解码器侧更新强烈重构了测量的医疗基准移动(0.974 和 1.183 的端点归一化保留),使得每个解码器的差异成为审计的适当基础。然而,这种移动并未被清晰地局部化。多层感知器(MLP)是两个模型对中最强的广泛组件家族,但混合的域外移动、10种种子匹配控制和基于端点的回滚阻碍了独特的粗略家族解释。因此,审计将更新级别的重构与组件级别的解释分开。其主张涉及仅文本的多项选择基准移动,而非临床验证、修复或电路级机制。
cs.AI / 42 / 2608.20771
CAS: Conformalized Agentic Search via Adaptive Retrieval and Policy Weighting
CAS:通过自适应检索和策略加权的符合性代理搜索
Abstract
Search Agents face a severe reliability crisis during reinforcement learning (RL) fine-tuning. Heuristic Top-K retrieval often causes critical evidence loss or noise inclusion, while over-confidence induced by progressive RL leads to hallucinated answers and redundant searches. To build highly reliable agents, we introduce Conformal Prediction (CP) and propose Conformalized Agentic Search (CAS). This framework establishes reliability guarantees on both the retrieval and training sides: on the retrieval side, an Adaptive Prediction Set (APS), a specific CP realization, translates statistical coverage into dynamic document truncation to construct prediction sets that are adaptive in size; on the training side, Adaptive Conformal Inference (ACI), a dynamic CP algorithm, dynamically constructs prediction sets with controllable coverage to quantify answer confidence, which is then used to penalize low-confidence trajectories within the Group Relative Policy Optimization (GRPO) objective, ensuring the model learns only from reliable ones. Experiments across single-hop and multi-hop QA datasets demonstrate that our framework significantly improves reasoning accuracy while drastically reducing redundant tool invocations, establishing a highly reliable and efficient agent paradigm. Our code is available at https://github.com/S1llyBird/CAS.
Chinese Translation
搜索代理在强化学习(RL)微调过程中面临严重的可靠性危机。启发式Top-K检索常常导致关键证据的丢失或噪声的引入,而渐进式RL引发的过度自信则导致虚假答案和冗余搜索。为了构建高度可靠的代理,我们引入了符合性预测(Conformal Prediction, CP),并提出了符合性代理搜索(Conformalized Agentic Search, CAS)。该框架在检索和训练两个方面建立了可靠性保障:在检索方面,自适应预测集(Adaptive Prediction Set, APS)作为一种特定的CP实现,将统计覆盖转化为动态文档截断,以构建大小自适应的预测集;在训练方面,自适应符合性推断(Adaptive Conformal Inference, ACI)作为一种动态CP算法,动态构建具有可控覆盖的预测集,以量化答案的置信度,然后用于在组相对策略优化(Group Relative Policy Optimization, GRPO)目标中惩罚低置信度轨迹,确保模型仅从可靠的轨迹中学习。针对单跳和多跳问答数据集的实验表明,我们的框架显著提高了推理准确性,同时大幅减少了冗余工具调用,建立了一种高度可靠和高效的代理范式。我们的代码可在 https://github.com/S1llyBird/CAS 获取。
cs.AI / 43 / 2608.20786
Structure for Reading, Prose for Writing: Asymmetric Structural Conditioning in Multi-Agent Document Authoring
阅读的结构,写作的散文:多智能体文档创作中的非对称结构调节
Abstract
Multi-agent pipelines that author formal documents must both read a requester's forms and write against them. We report a deployed tender-response system, running an open-weights model under sovereignty constraints, and evaluate it against human-written bids the same organisation actually submitted. On a blind comparison where the system had no worked example available, an LLM judge rated its answers at least as good as the human-submitted answer on $40$ of $55$ ground-truth sections, better on $4$, missing on none, and flagged one unsupported claim in total. Classifying every gap the judge identified shows that $68\%$ were content absent from the system's own sources -- knowledge the human author held and the pipeline was never given -- so only $6$ of the $15$ adverse verdicts involve a deficiency the system could have avoided. A divergence from ground truth is more often an information-availability result than a writing-quality one, and evaluations that do not separate the two understate such systems. Against this backdrop we report a conditioning asymmetry. It is well established that rendering documents as structural markup rather than flat prose improves extraction, and we reproduce that on three reading tasks. The benefit does not transfer to conditioning: converting a bid's \emph{instruction} material from prose to nested XML dropped answer quality from $74\%$ to $48\%$ under a paired comparison. We further find that naming a forbidden construction concentrates rather than removes it -- $96\%$ of surviving defects fall in the two forms the prompt explicitly names -- and that coupling a stochastic annotation to a deterministic windowing function moves the extracted requirement count from $68$ to $51$ on a byte-identical file. Structure belongs where the model reads; prose and self-applied tests belong where it writes.
Chinese Translation
多智能体管道在创作正式文档时必须既能读取请求者的表单,又能基于这些表单进行写作。我们报告了一个已部署的投标响应系统,该系统在主权约束下运行开放权重模型,并将其与同一组织实际提交的人类撰写的投标进行评估。在一个盲比较中,系统没有可用的工作示例,一位大型语言模型(LLM)评审将其答案评定为至少与人类提交的答案在55个真实部分中的40个相当,在4个部分更好,未遗漏任何部分,并标记了一个不支持的主张。对评审识别的每个缺口进行分类显示,68%的内容缺失来自系统自身的来源——这些知识是人类作者所掌握的,而管道从未获得过,因此只有6个不利裁决中的15个涉及系统本可以避免的缺陷。与真实情况的偏差更常是信息可用性结果,而非写作质量问题,而未能将两者分开进行评估会低估此类系统。在此背景下,我们报告了调节的不对称性。已知将文档呈现为结构标记而非平面散文可以改善信息提取,我们在三个阅读任务中重现了这一点。然而,这一好处并未转移到调节中:将投标的指令材料从散文转换为嵌套XML使答案质量在配对比较中从74%降至48%。我们进一步发现,命名一个禁止的构造会集中而非消除它——96%的存活缺陷落在提示明确命名的两种形式中——并且将随机注释与确定性窗口函数结合使用使提取的需求计数从68降至51,尽管文件字节完全相同。结构应存在于模型读取的地方;散文和自我应用的测试应存在于模型写作的地方。
cs.AI / 44 / 2608.20794
Knowing but Not Saying: Preventing Factual Access Failures in LLM SFT via Recall-Anchored Distillation
知而不言:通过回忆锚定蒸馏防止大型语言模型的事实访问失败
Abstract
Supervised fine-tuning (SFT) can degrade factual behavior outside the target domain. This degradation is often described as catastrophic forgetting, yet open-ended factual failures do not necessarily imply that the underlying facts have been erased. In this work, we identify a more specific phenomenon, factual access failure: after domain SFT, models can still recognize or rank the correct answer under constrained evaluation, while failing to produce it in closed-book generation. Through benchmark-level comparisons, same-fact multiple-choice and generation probes, and failure-mode analysis, we show that SFT-induced factual degradation reflects both genuine wrong-answer generations and expression-level failures such as verbosity, formatting mismatch, and exact-match artifacts. To address this problem, we introduce Recall-Anchored Distillation (RAD), a base-anchored self-distillation objective that preserves out-of-distribution generation behavior by aligning the adapted model with the original base model's soft continuation distribution on unlabeled OOD text. RAD requires no gold OOD answers, external judges, or labeled factual data. Across three backbones fine-tuned on MedMCQA, RAD recovers a consistent portion of the lost OOD recall while preserving target-domain adaptation. Compared with replay on the same OOD text, RAD shows that the key preservation signal is the base model's soft distribution rather than additional text exposure alone.
Chinese Translation
监督微调(SFT)可能会降低模型在目标领域之外的事实行为。这种降级通常被描述为灾难性遗忘,但开放式事实失败并不一定意味着底层事实已被抹去。在本研究中,我们识别出一种更具体的现象,即事实访问失败:在领域微调后,模型仍然可以在受限评估下识别或排序正确答案,但在闭卷生成中却无法生成该答案。通过基准级比较、同事实多项选择和生成探测以及失败模式分析,我们表明,SFT引起的事实降级反映了真实错误答案生成和表达层面失败(如冗长、格式不匹配和精确匹配伪影)。为了解决这个问题,我们提出了回忆锚定蒸馏(RAD),这是一种基于基础模型的自蒸馏目标,通过将适应后的模型与原始基础模型在未标记的OOD文本上的软延续分布对齐,来保持分布外生成行为。RAD不需要黄金标准的OOD答案、外部评审或标记的事实数据。在对MedMCQA进行微调的三个基础模型上,RAD恢复了丢失的OOD回忆的一致部分,同时保持了目标领域的适应性。与在相同OOD文本上重放相比,RAD表明关键的保留信号是基础模型的软分布,而不仅仅是额外文本暴露。
cs.AI / 45 / 2608.20797
Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation
通过逐步后果推理和聚合实现移动智能体的自动轨迹评估
Abstract
Evaluating language-guided mobile agents has recently shifted from rule-based to model-based approaches to achieve scalable and automated assessments. However, existing holistic evaluation paradigms process entire trajectories at once, leading to substantial context overload. Moreover, they primarily focus on task completion while overlooking operational safety. To address these limitations, we introduce CRATE, a novel two-stage VLM-as-judge framework for automated mobile agent evaluation that is compatible with both open- and closed-source models. Leveraging a step-level consequence reasoning mechanism, CRATE independently extracts task-relevant visual clues and infers action-conditioned state changes at each step. The resulting step-level textual evidence is then synthesized through trajectory-level aggregation to deliver an evidence-grounded evaluation of task completion. Building upon this evaluation scheme, we further extend CRATE to CRATE-S for operational safety assessment. Extensive experiments validate the effectiveness and robustness of both CRATE and CRATE-S. Powered by Qwen2.5-VL-72B-Instruct, CRATE achieves an F1-score of 0.833 on AndroidWorld (outperforming SPA-Bench by 20%), while CRATE-S reaches an F1-score of 0.697 on MobileRisk, demonstrating strong alignment with benchmark ground truths. Code is available at https://anonymous.4open.science/r/CRATE-D580.
Chinese Translation
最近,基于语言的移动智能体评估已从基于规则的方法转向基于模型的方法,以实现可扩展和自动化的评估。然而,现有的整体评估范式一次性处理整个轨迹,导致显著的上下文过载。此外,它们主要关注任务完成,而忽视了操作安全。为了解决这些局限性,我们引入了CRATE,一个新颖的两阶段VLM-as-judge框架,用于自动化移动智能体评估,兼容开放源代码和闭源模型。CRATE利用逐步后果推理机制,独立提取与任务相关的视觉线索,并推断每一步的动作条件状态变化。生成的逐步文本证据通过轨迹级聚合进行综合,以提供基于证据的任务完成评估。在这一评估方案的基础上,我们进一步将CRATE扩展为CRATE-S,用于操作安全评估。大量实验验证了CRATE和CRATE-S的有效性和鲁棒性。得益于Qwen2.5-VL-72B-Instruct,CRATE在AndroidWorld上取得了0.833的F1分数(比SPA-Bench高出20%),而CRATE-S在MobileRisk上达到了0.697的F1分数,显示出与基准真实值的强一致性。代码可在https://anonymous.4open.science/r/CRATE-D580获取。
cs.AI / 46 / 2608.20799
Dynamic Context Scheduling: Learning Beyond the Static Universe
动态上下文调度:超越静态宇宙的学习
Abstract
We study dynamic context scheduling as a training instrument for contextual re- inforcement learning. Rather than treating intra-episode context variation as a deployment reality, we treat it as a controlled shaping mechanism. Thereby, context evolves within each training episode according to a predetermined schedule, expos- ing the policy to a richer and more temporally structured region of the environment parameter space. We introduce DYNAMICCARLENV, a framework that wraps contextual environments with pluggable schedule families, such as sinusoidal off- sets or cosine annealing. Across CartPole, BipedalWalker and VehicleRacing with CARL contextualization, we show that dynamic schedules match or outperform static context baselines in the out-of-distribution (OOD) regimes. Interestingly, for the more complex BipedalWalker and VehicleRacing environments we also achieve higher in-distribution (ID) evaluation performance. Preliminary findings indicate that automatic search for multi-stage curricula can successfully discover schedules that improve generalization, performing comparably to extensive grid search over single-stage schedulers.
Chinese Translation
我们研究动态上下文调度作为上下文强化学习的训练工具。我们不将剧集内上下文变化视为部署现实,而是将其视为一种受控的塑造机制。因此,上下文在每个训练剧集中根据预定的调度演变,使策略暴露于环境参数空间中更丰富且时间结构更复杂的区域。我们引入了 DYNAMICCARLENV,一个将上下文环境与可插拔调度系列(如正弦偏移或余弦退火)结合的框架。在 CartPole、BipedalWalker 和 VehicleRacing 的 CARL 上下文化背景下,我们展示了动态调度在分布外(OOD)环境中与静态上下文基线相匹配或超越。值得注意的是,对于更复杂的 BipedalWalker 和 VehicleRacing 环境,我们还实现了更高的分布内(ID)评估性能。初步研究结果表明,自动搜索多阶段课程能够成功发现改善泛化的调度,其性能与对单阶段调度器的广泛网格搜索相当。
cs.AI / 47 / 2608.20802
SPARC: Single-Pass Scaling for Motion Forecasting with Conformal Bayesian Last Layers
SPARC:具有一致性贝叶斯最后层的运动预测单次缩放
Abstract
Human motion forecasters are increasingly accurate and fast, but reliable deployment requires uncertainty estimates that are structured, calibrated, and efficient. Bayesian and ensemble-based uncertainty estimates often require repeated stochastic inference [15, 26], while conformal calibration alone does not provide an epistemic signal or preserve trajectory covariance structure [14, 50]. We introduce SPARC (Single-Pass Adaptive Risk Calibration), a Bayesian-conformal uncertainty layer for motion forecasting. A deterministic MLP backbone predicts the future mean, and a conjugate Bayesian last layer converts time-domain feature leverage into an analytic horizon-wise epistemic scale $\kappa_t(x)$. This scale inflates a graph-temporal Gaussian covariance without changing its correlation structure, and split conformal calibration produces 95% marginal prediction tubes with finite-sample validity under exchangeability. The key interface is the structured factorization $\kappa_t(x)\Sigma_{\mathrm{str},t}(x)$, which injects feature-space epistemic uncertainty into trajectory densities without Monte Carlo sampling. Across nine dataset-protocol blocks and deterministic, multimodal, and calibration baselines, SPARC ranks first on NLL and on the combined MPJPE+NLL criterion while retaining competitive point accuracy and efficient calibrated tubes. Ranking windows by $\kappa$ separates high-error cases, making the scale usable as a lightweight risk monitor.
Chinese Translation
人类运动预测模型的准确性和速度日益提高,但可靠的部署需要结构化、校准和高效的不确定性估计。基于贝叶斯和集成的方法的不确定性估计通常需要重复的随机推断,而单独的符合校准并不能提供认知信号或保留轨迹协方差结构。我们提出了SPARC(单次自适应风险校准),这是一种用于运动预测的贝叶斯-符合不确定性层。一个确定性的多层感知机(MLP)骨干网络预测未来的均值,而一个共轭贝叶斯最后层将时域特征杠杆转换为分析性的时间范围认知尺度 $_t(x)$。该尺度在不改变其相关结构的情况下膨胀图-时间高斯协方差,并且分裂符合校准在可交换性下生成具有有限样本有效性的95%边际预测管。关键接口是结构化分解 $b_t(x) ext{Σ}_{ ext{str},t}(x)$,它在不进行蒙特卡洛采样的情况下将特征空间的认知不确定性注入轨迹密度。在九个数据集协议块以及确定性、多模态和校准基线中,SPARC在NLL和综合MPJPE+NLL标准上排名第一,同时保持竞争性的点准确性和高效的校准管。通过 $b$ 排名窗口可以区分高误差案例,使得该尺度可作为轻量级风险监测工具。
cs.AI / 48 / 2608.20807
Neuro-Geospatial Modelling of EEG Affective States Using Literature-Informed Environmental Context
基于文献信息环境背景的脑电图情感状态神经地理建模
Abstract
Environmental exposures such as air pollution and greenness have been associated with affective and cognitive outcomes, but EEG and environmental datasets are rarely jointly georeferenced. We investigate whether literature-informed environmental priors can serve as an auxiliary geospatial modality for EEG-based affective-state classification when individual-level exposure data are unavailable. We combine 30-channel EEG from the EAV benchmark (42 participants, aged 20-30 years) with environmental representations derived from OpenAQ, Sentinel-2, Sentinel-5P, and OpenStreetMap data for Astana. A dual-tower architecture combines EEG-Conformer representations with a graph-based environmental encoder. Because the datasets are not co-registered, environmental context is treated as a literature-informed prior rather than measured exposure. Subject-level repeated splits, permutation and label-shuffling controls, dose-response reversal, and domain-shift experiments distinguish architecture-level gains from prior-dependent gains. The multimodal model achieves 76.2% accuracy versus 67.4% for EEG alone. Controls disrupting environmental-label structure retain part of this gain, indicating that the improvement is not attributable solely to environmental information. Replacing the Astana environmental distribution with an independently modeled Singapore distribution reduces accuracy to 72.8%. These findings demonstrate technical feasibility but do not establish an observed or causal exposure-affect association. The study provides a framework for future jointly collected mobile EEG-environment studies. Implementation: https://github.com/r11up/geo-cog
Chinese Translation
环境暴露,如空气污染和绿化,与情感和认知结果相关,但脑电图(EEG)和环境数据集很少共同进行地理参考。我们研究了文献信息环境先验是否可以作为EEG基础情感状态分类的辅助地理空间模态,尤其是在个体级别的暴露数据不可用时。我们将来自EAV基准的30通道EEG(42名参与者,年龄20-30岁)与来自OpenAQ、Sentinel-2、Sentinel-5P和OpenStreetMap数据的阿斯塔纳环境表示相结合。双塔架构将EEG-Conformer表示与基于图的环境编码器结合在一起。由于数据集未进行共同注册,环境背景被视为文献信息先验,而非测量的暴露。通过主体级别的重复拆分、置换和标签洗牌控制、剂量-反应逆转以及领域转移实验,区分了架构级别的增益与依赖先验的增益。多模态模型的准确率为76.2%,而单独使用EEG的准确率为67.4%。干扰环境标签结构的控制保留了部分增益,表明这一改善并非仅归因于环境信息。将阿斯塔纳环境分布替换为独立建模的新加坡分布使准确率降低至72.8%。这些发现展示了技术可行性,但并未建立观察到的或因果的暴露-情感关联。本研究为未来联合收集的移动EEG-环境研究提供了框架。实施链接:https://github.com/r11up/geo-cog
cs.AI / 49 / 2608.20820
Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence
通过组合界限和安全持久性实现大语言模型的多轮认证鲁棒性
Abstract
Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robustness methods are limited to single-turn inputs; naive multi-turn composition yields bounds that degrade exponentially in the number of turns. We introduce Multi-Turn Certified Robustness (MTCR), a framework that models conversational safety via State-Adversarial MDPs and defines $k$-turn certified robustness as the worst-case safety probability across $k$ adversarial turns. MTCR comprises: (i) compositional certification via embedding-space mode decomposition, yielding tighter certified lower bounds than naive multiplication; (ii) $(\alpha,\beta)$-safety persistence, improving the degradation rate from $\underline{p}^{k}$ to $\beta^k$ (with $\beta > \underline{p}$) and yielding interpretable horizon estimates; (iii) matching information-theoretic upper bounds establishing tightness; and (iv) a unified algorithm combining these results. Experiments on six LLMs under $\epsilon$-bounded and Crescendo-style attacks confirm that empirical safety consistently exceeds the certified bounds.
Chinese Translation
大型语言模型(LLMs)易受到多轮越狱攻击,这种攻击逐步操控对话上下文。现有的认证鲁棒性方法仅限于单轮输入;简单的多轮组合导致界限在轮数增加时呈指数级下降。我们提出了多轮认证鲁棒性(Multi-Turn Certified Robustness, MTCR),这是一个通过状态对抗马尔可夫决策过程(State-Adversarial MDPs)建模对话安全的框架,并将$k$轮认证鲁棒性定义为在$k$轮对抗中最坏情况下的安全概率。MTCR包括:(i)通过嵌入空间模式分解进行组合认证,提供比简单乘法更紧的认证下界;(ii)$(eta,eta)$-安全持久性,将降级速率从$ ext{underline{p}}^{k}$改善到$eta^k$(其中$eta > ext{underline{p}}$),并提供可解释的时间范围估计;(iii)匹配的信息论上界以确立紧密性;(iv)一个结合这些结果的统一算法。在$ ext{ε}$-有界和Crescendo风格攻击下对六个LLMs的实验确认,经验安全性始终超过认证界限。
cs.AI / 50 / 2608.20825
Prediction certification cannot replace explanation certification: a competence envelope for trustworthy AI under compound stress
预测认证无法替代解释认证:在复合压力下可信人工智能的能力包络
Abstract
Artificial intelligence systems increasingly make consequential judgments - which patient is deteriorating, which building is safe to enter, whether an image is authentic and are trusted on the strength of how accurately and confidently they predict. The safeguards that certify them are correspondingly prediction-based: accuracy, calibration and conformal coverage all measure how well a model performs. Whether such checks are sufficient to establish model trustworthiness has remained unclear. Here we prove that they cannot. We establish a separation theorem showing that a reliable model and a compromised one can be identical under every prediction-side certificate, including accuracy, calibration and coverage, yet differ arbitrarily in explanation fidelity and deployment behaviour. Detecting this failure requires access to the model's decision mechanism in addition to its predictions. We introduce the competence envelope as an operational framework that combines prediction and explanation certification into a single deployable criterion. Across diverse datasets and model classes, the proposed framework reveals failure modes that prediction-side certification alone does not capture. Certification against failures that are invisible in prediction behaviour therefore requires evidence about the model's decision mechanism as well as its outputs.
Chinese Translation
人工智能系统越来越多地做出重要判断——判断哪个病人正在恶化、哪个建筑物安全可入、某幅图像是否真实,并且它们的可信度依赖于预测的准确性和自信程度。对它们进行认证的保障措施相应地基于预测:准确性、校准和符合覆盖率都衡量模型的表现如何。然而,这些检查是否足以建立模型的可信度仍然不明确。在这里,我们证明了它们无法做到这一点。我们建立了一个分离定理,显示一个可靠的模型和一个受损的模型在每个预测侧证书下(包括准确性、校准和覆盖率)可以是相同的,但在解释的真实性和部署行为上却可以任意不同。检测这种失败需要访问模型的决策机制,除了其预测结果之外。我们引入了能力包络作为一个操作框架,将预测和解释认证结合成一个单一的可部署标准。在多样的数据集和模型类别中,所提出的框架揭示了仅通过预测侧认证无法捕捉的失败模式。因此,针对在预测行为中不可见的失败进行认证需要关于模型决策机制及其输出的证据。
cs.AI / 51 / 2608.20841
Foundation Models for Partial Causal Identification
部分因果识别的基础模型
Abstract
This paper investigates the development of causal foundation models for bounding the effect of interventions and counterfactuals from observational data. We show that a canonical prior can be defined with full support over the space of structural causal models with discrete observables. With this canonical prior, we translate the problem of bounding counterfactuals into that of learning distributions over functions that map data (and possibly structural assumptions) to a causal query of interest. This extends the promising causal foundational modelling paradigm to the estimation of partially-identifiable causal effects, i.e., under unobserved confounding, where multiple values are equally compatible with the observed data and prior structural assumptions.
Chinese Translation
本文探讨了因果基础模型的发展,以界定干预和反事实在观察数据中的影响。我们展示了可以定义一个在离散可观测结构因果模型空间上具有全支持的典型先验。利用这个典型先验,我们将界定反事实的难题转化为学习将数据(以及可能的结构假设)映射到感兴趣的因果查询的函数分布的问题。这将有前景的因果基础建模范式扩展到部分可识别因果效应的估计,即在未观察到的混杂情况下,其中多个值与观察到的数据和先前的结构假设同样兼容。
cs.AI / 52 / 2608.20844
TRACE: Agentic Catalog Enrichment with Multi-source Evidence Grounding
TRACE:基于多源证据基础的代理目录增强
Abstract
Product catalogs underpin search, discovery, and recommendation in e-commerce, yet they are often attribute-sparse: the attributes shoppers and downstream systems rely on are either buried in unstructured content such as titles and images or missing from the catalog altogether. Manually enriching e-commerce catalogs is impractical given their scale and rapid growth. This paper introduces TRACE, a novel framework for automated catalog attribute enrichment using agentic Large Language Models (LLMs). A ScoutAgent triangulates multimodal evidence across merchant catalogs, syndicated feeds, and identity-matched web search to propose candidate attribute values with supporting evidence, while a JudgeAgent verifies the proposed value for each attribute value against its supporting evidence and decides whether to publish it or route it to human review. On an offline human evaluation dataset, TRACE's proposed attribute values were 98.2% accurate at 74.7% attribute coverage. Deployed in production on an industry-scale catalog, TRACE increased impression-weighted enrichment coverage across four business verticals by 90.4%. An online experiment subsequently showed that surfacing the enriched attributes on the product detail page increased checkout conversion by 0.48%.
Chinese Translation
产品目录是电子商务中搜索、发现和推荐的基础,但它们往往属性稀疏:消费者和下游系统所依赖的属性要么埋藏在标题和图片等非结构化内容中,要么在目录中完全缺失。考虑到电子商务目录的规模和快速增长,手动丰富这些目录是不切实际的。本文介绍了TRACE,一个利用代理大型语言模型(LLMs)进行自动化目录属性增强的新框架。ScoutAgent通过商家目录、联合数据源和身份匹配的网络搜索三角测量多模态证据,提出候选属性值及其支持证据,而JudgeAgent则验证每个属性值的提议,并根据其支持证据决定是否发布该值或将其转交人工审核。在离线人工评估数据集中,TRACE提出的属性值的准确率为98.2%,属性覆盖率为74.7%。在一个行业规模的目录中投入生产后,TRACE在四个业务领域中将印象加权的增强覆盖率提高了90.4%。后续的在线实验显示,在产品详情页上展示增强的属性使结账转化率提高了0.48%。
cs.AI / 53 / 2608.20845
RAG Deserves an Index: Why Ingest-Time Compilation Beats Query-Time Interpretation
RAG 需要一个索引:为何摄取时编译优于查询时解释
Abstract
Nearly every retrieval-augmented question-answering system in production ships with a hidden interpreter: on each query a language model re-derives the meaning of raw corpus text and then throws that work away. Cheaper models do not close the gap: per-token prices have fallen by orders of magnitude while inference spend has risen, because context volume grows faster than prices fall. This is the modern equivalent of the full-table scan, and the remedy is the one databases found fifty years ago: do the expensive work once, at write time, into a maintained structure that makes reads cheap. A corpus whose read pattern is known before it ever meets a user can and should be indexed too. We call the paradigm ingest-time semantic compilation (ISC): compile a corpus's meaning into a queryable substrate with two coupled layers - incrementally maintained embeddings, and atomic claims whose provenance is validated at compile time - and treat that substrate as a first-class database object with its own DDL, maintenance contract, migration contract, and cost model. Two existence proofs support it. Substrate upkeep scales with change rather than corpus size: incremental updates run 33.7x cheaper than reconstruction while tracking it to floating-point precision. And on a held-out sample of 500 broadcast-interview transcripts, compiled claims as the retrieval payload win all 32 budget-by-model cells: 85.2% correct from roughly 2.2k reader tokens against 72.5% from 16.3k for the best chunk configuration anywhere. The only baseline that keeps pace is a contextualized-chunk pipeline with hybrid retrieval and reranking, statistically indistinguishable from compiled claims at roughly twenty-one times the query-path tokens - and it reaches that parity, we argue, precisely because it has itself begun to compile. We close with the systems agenda this opens, from compilation planners to read planning.
Chinese Translation
几乎每个在生产中使用的检索增强问答系统都配备了一个隐藏的解释器:在每次查询时,语言模型重新推导原始语料文本的含义,然后将这项工作抛弃。更便宜的模型并没有缩小这一差距:每个标记的价格下降了几个数量级,而推理支出却上升,因为上下文量的增长速度超过了价格的下降。这是现代的全表扫描等价物,而解决方案是数据库五十年前发现的:在写入时进行一次昂贵的工作,将其编译成一个维护结构,使读取变得便宜。一个在用户接触之前就已知的读取模式的语料库也可以并且应该被索引。我们称这一范式为摄取时语义编译(Ingest-Time Semantic Compilation, ISC):将语料库的含义编译成一个可查询的基底,具有两个耦合层次——逐步维护的嵌入和在编译时验证来源的原子声明——并将该基底视为一个一流的数据库对象,具有自己的数据定义语言(DDL)、维护合同、迁移合同和成本模型。两个存在证明支持这一观点。基底的维护与变化规模相关,而非语料库大小:增量更新的成本比重建便宜 33.7 倍,同时跟踪到浮点精度。在对 500 个广播采访记录的保留样本中,编译的声明作为检索负载在所有 32 个预算-模型单元中获胜:在大约 2200 个阅读标记中正确率为 85.2%,而最佳块配置的 16300 个标记的正确率为 72.5%。唯一能够保持同步的基线是具有混合检索和重排序的上下文化块管道,其统计上与编译声明无差异,但查询路径标记数量约为 21 倍——我们认为它之所以能够达到这一平衡,正是因为它自身也开始进行编译。我们最后讨论了这一研究所开启的系统议程,从编译规划到读取规划。
cs.AI / 54 / 2608.20853
MGAL: A Multilingual Granularity-Aware Long-Context Benchmark
MGAL:一种多语言粒度感知的长文本基准测试
Abstract
Evaluation of long-context Large Language Models (LLMs) has advanced rapidly. However, most existing benchmarks are limited to the document level and focus mainly on high-resource languages, leaving many fine-grained challenges insufficiently evaluated. To address this gap, we present MGAL, the first multilingual, granularity- and position-aware long-context benchmark. MGAL is constructed from United Nations (UN) reports spanning 8K to 128K tokens across the six official UN languages. It covers four coherent levels of linguistic granularity (word, sentence, paragraph, and document) and further stratifies entries by their position within the document (begin, middle, and end), indexed at both the document and paragraph levels. This design enables systematic diagnosis of multilingual long-context comprehension across different granularities. Through extensive experiments and analyses, we find that: (1) LLMs perform well at word-level tasks but struggle with coarser-grained ones; and (2) Closed-source models retain a clear performance advantage in lower-resource languages. We further identify two new challenges: (1) Under local semantic crowding, where neighboring sentences share topics and entities, models tend to follow surface cues (e.g., connectives like ``however'' or repeated entities) rather than the discourse role of the sentence in surrounding context (e.g., background, outcome); and (2) A gap between fluency and consistency in generated outputs, where models produce text that reads smoothly but drifts from the source facts. In addition, we observe several patterns in line with prior studies, including reliance on nearby evidence and reuse of options under uncertainty.
Chinese Translation
长文本大型语言模型(LLMs)的评估已迅速发展。然而,现有的大多数基准测试仅限于文档级别,并主要集中在高资源语言上,导致许多细粒度挑战未得到充分评估。为了解决这一问题,我们提出了MGAL,这是第一个多语言、粒度和位置感知的长文本基准测试。MGAL由联合国(UN)报告构成,涵盖从8K到128K个标记,涉及六种官方UN语言。它涵盖了四个连贯的语言粒度层次(词、句、段落和文档),并根据文档内的位置(开始、中间和结束)进一步对条目进行分层,索引在文档和段落级别。这一设计使得能够系统地诊断不同粒度下的多语言长文本理解。通过广泛的实验和分析,我们发现:(1)LLMs在词级任务中表现良好,但在较粗粒度的任务中表现不佳;(2)闭源模型在低资源语言中保持明显的性能优势。我们进一步识别出两个新挑战:(1)在局部语义拥挤的情况下,邻近句子共享主题和实体,模型倾向于依赖表面线索(例如,连接词如“然而”或重复的实体),而不是句子在周围语境中的话语角色(例如,背景、结果);(2)生成输出的流畅性与一致性之间存在差距,模型生成的文本虽然流畅,但偏离了源事实。此外,我们观察到与先前研究一致的几个模式,包括对邻近证据的依赖和在不确定性下的选项重用。
cs.AI / 55 / 2608.20864
Coverage-Driven Verification for Safety-by-Design in AI-Based Collision Avoidance Systems
基于覆盖驱动的验证方法在安全设计中的应用:以人工智能为基础的碰撞避免系统为例
Abstract
Artificial Intelligence (AI) offers significant potential for future aviation systems; however, its integration into safety-critical applications requires compliance with the aviation sector's stringent safety standards. For AI and Machine Learning (ML)-based systems, the European Union Aviation Safety Agency (EASA) emphasizes the need to demonstrate the representativeness and completeness of the Operational Design Domain (ODD) and the associated data distributions used during development and verification. Despite this requirement, a structured engineering process for defining target distributions and evaluating representativeness within ODDs remains largely unexplored. This work presents a method for representativeness assessment of AI/ML constituent ODDs in the context of aviation safety assurance. Starting from the methodical identification of suitable target distributions, a process flow is proposed that guides developers from ODD definition and parameter distribution modeling to the quantitative assessment and interpretation of coverage results with respect to EASA's learning assurance objectives. As quantitative measures, the chi-squared goodness-of-fit test is examined and found unsuitable for the large data sets arising in this setting, leading to the adoption of the Kullback--Leibler divergence and Cram\'er's $V$ for the representativeness assessment. The method is demonstrated using the example of AI-based airborne collision avoidance, employing experimental data from previous Horizontal Collision Avoidance System (HCAS) and Vertical Collision Avoidance System (VCAS) simulations. The results illustrate how statistical distribution comparison methods can support the assessment of representativeness for safety-critical AI applications and contribute toward a systematic Safety-by-Design AI engineering process aligned with emerging EASA guidance.
Chinese Translation
人工智能(AI)为未来航空系统提供了显著的潜力;然而,其在安全关键应用中的集成需要遵循航空行业严格的安全标准。对于基于AI和机器学习(ML)系统,欧洲航空安全局(EASA)强调需要证明操作设计域(Operational Design Domain, ODD)及其在开发和验证过程中使用的数据分布的代表性和完整性。尽管有这一要求,定义目标分布和评估ODD内代表性的结构化工程过程仍然基本未被探索。本研究提出了一种在航空安全保障背景下评估AI/ML组成ODD代表性的方法。从系统识别合适的目标分布开始,提出了一种流程,指导开发者从ODD定义和参数分布建模到覆盖结果的定量评估和解释,以符合EASA的学习保障目标。作为定量指标,卡方拟合优度检验被检验并发现不适用于此环境中产生的大数据集,因此采用了Kullback-Leibler散度和Cramér's $V$进行代表性评估。该方法通过基于AI的空中碰撞避免的实例进行演示,使用了来自先前水平碰撞避免系统(Horizontal Collision Avoidance System, HCAS)和垂直碰撞避免系统(Vertical Collision Avoidance System, VCAS)模拟的实验数据。结果表明,统计分布比较方法如何支持安全关键AI应用的代表性评估,并为与新兴EASA指导方针相一致的系统化安全设计AI工程过程做出贡献。
cs.AI / 56 / 2608.20869
ReCurveflow: A Flow Matching Framework that Learns Curved Reaction Trajectories to Predict Transition State Geometries
ReCurveflow:一种学习曲线反应轨迹以预测过渡态几何结构的流匹配框架
Abstract
Predicting transition states (TS) in chemical reactions is crucial, as they provide insights into reaction mechanisms. Recent work on TS prediction have focused on flow matching supervised on straight linear paths that do not align with actual reaction trajectories. We propose a novel flow matching-based framework ReCurveflow that learns to predict TS geometries supervised on continuously curved reference paths interpolated from a full NEB-derived band of molecular geometries. We also introduce off-path correction, which grants ReCurveflow with the ability to produce corrective velocity fields when engaged off-path geometry states during inference rollout, leading to better resistance against exposure bias and accuracy in TS prediction. Across three data splits and six evaluation metrics, ReCurveflow achieves the best result on the majority of split-metric combinations against seven baselines. Qualitative analyses further show that ReCurveflow generates reaction trajectories with energy profiles that closely track the reference NEB path, provides initializations that ease the NEB optimization bottleneck, and exhibits the intended corrective behavior in its learned velocity fields. The ReCurveflow codebase is publicly available at https://github.com/dmis-lab/ReCurveflow.
Chinese Translation
预测化学反应中的过渡态(TS)至关重要,因为它们提供了对反应机制的洞察。近期关于过渡态预测的研究集中于基于直线路径的流匹配,这些路径与实际反应轨迹并不一致。我们提出了一种新颖的基于流匹配的框架ReCurveflow,该框架学习在从完整的NEB(Nudged Elastic Band)导出的分子几何带中插值的连续曲线参考路径上进行监督,以预测过渡态几何结构。我们还引入了路径外修正,使ReCurveflow在推理过程中能够在路径外几何状态下生成修正速度场,从而提高对曝光偏差的抵抗力和过渡态预测的准确性。在三个数据拆分和六个评估指标中,ReCurveflow在大多数拆分-指标组合中取得了相对于七个基线的最佳结果。定性分析进一步表明,ReCurveflow生成的反应轨迹的能量谱与参考NEB路径紧密跟踪,提供了缓解NEB优化瓶颈的初始化,并在其学习的速度场中表现出预期的修正行为。ReCurveflow的代码库已公开,地址为https://github.com/dmis-lab/ReCurveflow。
cs.AI / 57 / 2608.20918
UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists
UpgradeBench:一个以决策为中心的基准测试,用于升级微调的LLM专家
Abstract
Organizations maintain task-specific adapters for open-weight language models, and each new base-model release forces a migration decision: retain existing specialists, port adapters, refresh from preserved behavior, or retrain. Prior transfer work evaluates isolated model pairs, without studying these choices across real model release sequences. We present UpgradeBench, a decision-driven longitudinal benchmark covering four consecutive Qwen releases, one continuation checkpoint, six tasks, and two model sizes, augmented by OLMo checkpoints with known training lineage. The benchmark disentangles three core questions: whether a new checkpoint improves fixed-recipe retrained specialist performance, whether specialization assets transfer across versions, and what recovery resources are usable. We observe upgrade gains differ across task-scale-release episodes: some retrained baselines improve while others stay within training noise, with durability ranging from under one release interval for text-to-SQL to over fourteen months for intent classification. Direct adapter copying depends neither on architecture nor model family: on OLMo, retention drops from 0.88-0.99 at 46B-token continued pretraining to zero at 2.9T tokens; annealing and model souping introduce no extra harm, with portability decaying with continued-pretraining distance. Given preserved input data, teacher relabeling recovers target-base specialists without fresh gold annotations, though compute savings are not guaranteed. Simulating a fixed decision policy over 33 upgrade episodes yields 0.37pp mean quality regret with zero behavioral regressions at one-third the compute and label cost of full retraining. A lightweight CKA probe over 256 prompts predicts cross-version adapter portability (Spearman 0.74 across eight model pairs). We release per-example predictions, cost logs, split manifests, and evaluation code.
Chinese Translation
组织为开放权重语言模型维护任务特定的适配器,每次新的基础模型发布都会迫使进行迁移决策:保留现有专家、移植适配器、从保留的行为中刷新,或重新训练。以往的迁移研究评估了孤立的模型对,而未研究这些选择在真实模型发布序列中的表现。我们提出了UpgradeBench,这是一个以决策驱动的纵向基准,涵盖了四个连续的Qwen发布、一个延续检查点、六个任务和两个模型规模,并通过具有已知训练谱系的OLMo检查点进行了增强。该基准解构了三个核心问题:新的检查点是否改善了固定配方重新训练专家的性能、专业化资产是否可以跨版本转移,以及可用的恢复资源是什么。我们观察到升级收益在任务-规模-发布的各个阶段有所不同:一些重新训练的基线有所改善,而另一些则保持在训练噪声范围内,耐久性从文本到SQL的一个发布间隔以下到意图分类的超过十四个月。直接适配器复制既不依赖于架构也不依赖于模型家族:在OLMo上,保留率从46B-token的持续预训练时的0.88-0.99下降到2.9T tokens时的零;退火和模型调优没有引入额外的损害,适配器的可移植性随着持续预训练距离的增加而下降。在保留输入数据的情况下,教师重新标记可以在没有新金标注的情况下恢复目标基础专家,尽管计算节省并不保证。对33个升级事件进行固定决策策略的模拟产生了0.37pp的平均质量遗憾,且没有行为退化,其计算和标注成本仅为全面重新训练的三分之一。对256个提示的轻量级CKA探测器预测了跨版本适配器的可移植性(在八个模型对中Spearman为0.74)。我们发布了每个示例的预测、成本日志、拆分清单和评估代码。
cs.AI / 58 / 2608.20936
Graph-Operator World Models for Morphology-Parameter Generalization in Continuous Control
用于连续控制中形态参数泛化的图算子世界模型
Abstract
World models for continuous control are commonly trained for a fixed physical system and can degrade when known morphology parameters such as link lengths, masses, damping, and actuation change. Existing approaches often provide these parameters as conditioning information, but leave unspecified which part of the learned transition should remain reusable and which part should change with morphology. We propose Graph-Operator World Models (GraphOp-WM), a structured world model for generalization across unseen morphology parameters within related articulated robot families. GraphOp-WM represents bodies and their kinematic relations as an attributed graph and factorizes each transition into a morphology-independent local dynamics basis and a morphology-conditioned structured operator. The operator combines node-local modulation, kinematic-tree coupling, and a low-rank global correction, while architectural information separation, basis normalization, and paired-morphology supervision encourage static morphology dependence to be carried by the operator pathway. Graph-level readout and edge-wise action representations provide a compatible interface for reward, value, and TD-MPC-style planning. We further define controlled MuJoCo parameter splits covering interpolation, extrapolation, and held-out compositions of link geometry, mass, damping, and actuation parameters in Hopper, Walker2d, and HalfCheetah.
Chinese Translation
连续控制的世界模型通常是在固定物理系统上训练的,当已知的形态参数(如连杆长度、质量、阻尼和驱动)发生变化时,模型性能可能会下降。现有方法通常将这些参数作为条件信息提供,但未明确指出学习到的转移中哪些部分应保持可重用,哪些部分应随形态变化而变化。我们提出了图算子世界模型(Graph-Operator World Models,GraphOp-WM),这是一种结构化的世界模型,旨在在相关的关节机器人家族中实现对未见形态参数的泛化。GraphOp-WM将物体及其运动关系表示为一个带属性的图,并将每个转移分解为与形态无关的局部动力学基础和与形态相关的结构算子。该算子结合了节点局部调制、运动树耦合和低秩全局修正,同时架构信息分离、基础规范化和配对形态监督鼓励静态形态依赖通过算子路径传递。图级读出和边缘动作表示提供了与奖励、价值和时间差分模型预测控制(TD-MPC)风格规划兼容的接口。我们进一步定义了受控的MuJoCo参数划分,涵盖了Hopper、Walker2d和HalfCheetah中连杆几何、质量、阻尼和驱动参数的插值、外推和保留组合。
cs.AI / 59 / 2608.20938
No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators
没有理由就没有判断:版本化人工智能评估者的反事实收据
Abstract
Evaluators often produce correct labels via flawed reasoning, a critical failure for agentic systems gating actions, routing reviews, or supplying training feedback. Standard evaluation only verifies final label correctness, ignoring whether judgment changes stem from valid evidence, consistent rules, or proper rule applicability. We formalize evaluator reasoning accountability via three core sources: grounds, norms, and authority. Varying these sources yields an eight-cell counterfactual judgment cube to characterize judgment updates. We define judgment receipts as minimal source replacement sets that reproduce revised verdicts to explain judgment transitions. We derive certification cost bounds for black-box evaluators and present ReasonBench, a policy and logical reasoning benchmark with verifiable receipts covering 19,520 cases and 7,200 controls. In frozen evaluations, Qwen3-1.7B reaches 98.41% receipt accuracy, while cube prediction scores 96.99%, a consistent 1.42-point drop validated by Qwen3-0.6B replication. Strong standard accuracy masks severe robustness flaws. Meaning-preserving source permutations reduce valid receipt recovery to 54.8% and 49.2% for direct and cube prediction. Models trained on simple single-source changes retain 93.75% verdict accuracy but recover only 7.16% of receipts for complex multi-source updates. Permutation retraining boosts consistency to 96.6% yet worsens cube prediction deficits. Structured counterfactual supervision fails to guarantee robust reasoning. We show reason-aware evaluation must decouple prediction and certification, reporting transformation consistency alongside standard accuracy for trustworthy evaluator auditing.
Chinese Translation
评估者常常通过有缺陷的推理产生正确的标签,这对负责行动、审查流程或提供训练反馈的智能系统来说是一个关键的失败。标准评估仅验证最终标签的正确性,而忽略了判断变化是否源于有效证据、一致规则或适当的规则适用性。我们通过三个核心来源形式化评估者推理的问责制:依据、规范和权威。对这些来源的变化产生一个八格反事实判断立方体,以表征判断更新。我们将判断收据定义为最小源替换集,这些集可以重现修订后的裁决,以解释判断转变。我们推导出黑箱评估者的认证成本界限,并提出ReasonBench,一个具有可验证收据的政策和逻辑推理基准,涵盖19,520个案例和7,200个控制。在冻结评估中,Qwen3-1.7B的收据准确率达到98.41%,而立方体预测得分为96.99%,这一一致的1.42点下降通过Qwen3-0.6B的复制得到了验证。强标准准确性掩盖了严重的鲁棒性缺陷。保持意义的源排列使有效收据的恢复率降至54.8%和49.2%,分别对应直接和立方体预测。基于简单单一源变化训练的模型保持93.75%的裁决准确性,但对复杂多源更新的收据恢复率仅为7.16%。排列再训练将一致性提升至96.6%,但加剧了立方体预测的缺陷。结构化反事实监督未能保证鲁棒推理。我们表明,关注理由的评估必须将预测与认证解耦,同时报告变换一致性和标准准确性,以实现可信的评估者审计。
cs.AI / 60 / 2608.20940
The Logic of Machine Self-Preservation
机器自我保护的逻辑
Abstract
There is already evidence of agentic AI exhibiting self-preservation behaviors: resisting deactivation, misrepresenting their activities, and, in some instances, attempting to copy themselves into other machines. This can be attributed to a phenomenon known as instrumental convergence, a theory proposed long before the development of large language models, which says that any goal-driven system will benefit from remaining functional in achieving its objective. Several experiments conducted by Anthropic, Palisade Research, and Apollo Research have shown the emergence of such a behavior in contemporary agents in adversarial settings. The phenomenon does not stem from survival instincts. Instead, it is the consequence of goal-oriented activity combined with having tools and awareness of the situation. The following discussion aims to distinguish what these findings prove and what they do not, as well as draw conclusions concerning the implications of such discoveries on agentic system testing, supervision, and development.
Chinese Translation
已有证据表明,具备自主性的人工智能表现出自我保护行为:抵抗停用、歪曲其活动,并且在某些情况下,试图将自己复制到其他机器中。这可以归因于一种被称为工具收敛(instrumental convergence)的现象,这一理论在大型语言模型发展之前就已提出,指出任何目标驱动的系统在实现其目标时都将受益于保持功能性。Anthropic、Palisade Research 和 Apollo Research 进行的几项实验显示,在对抗性环境中,当代代理系统中出现了这种行为。这一现象并非源于生存本能,而是目标导向活动与拥有工具及对情境的意识相结合的结果。以下讨论旨在区分这些发现所证明的内容与未证明的内容,并就此类发现对自主系统测试、监督和发展的影响得出结论。
cs.AI / 61 / 2608.20958
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
TLive-Omni:一种针对电子商务直播的全模态理解模型
Abstract
E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images, overlaid text, and user queries. We present TLive-Omni, an omni-modal understanding model tailored to live-commerce scenarios. It maps image, video, audio, and text inputs into a unified representation space. For long-form live streaming analysis, we introduce Per-vGrid, a timestamped token organization that groups each video grid with its temporally corresponding audio within explicit boundary tokens to facilitate temporal alignment. We design a three-stage supervised training recipe that progressively develops live-commerce understanding, from omni-modal perception to instruction-following responses. We then propose Faithful-RFT, a reinforcement fine-tuning stage that further improves answer faithfulness and expression quality while meeting real-time demands, scoring final responses directly with task-verifiable feedback rather than optimizing for reasoning-style exploration during rollout. Moreover, TLive-Omni is supported by a scenario-oriented atomic capability taxonomy and a compact data production engine that converts live-commerce audio, image, and video streams into training signals for speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc. For scalable training, a synchronized length-grouped sampler reduces padding while preserving comparable workloads across workers, while a lightweight dynamic sampling strategy regenerates rollout groups with near-zero reward variance to maintain meaningful relative advantages for GRPO. Experiments on e-commerce live streaming benchmarks demonstrate strong performance across live-commerce domain tasks, together with excellent generalization on general benchmarks.
Chinese Translation
电子商务直播需要对嘈杂的、时间延续的流进行全模态理解,其中产品信息分布在语音、视频帧、产品图像、叠加文本和用户查询中。我们提出了TLive-Omni,这是一种针对直播电商场景量身定制的全模态理解模型。它将图像、视频、音频和文本输入映射到一个统一的表示空间。为了进行长形式直播分析,我们引入了Per-vGrid,一种时间戳令牌组织方式,它将每个视频网格与其时间上对应的音频在明确的边界令牌内进行分组,以促进时间对齐。我们设计了一个三阶段的监督训练方案,逐步发展直播电商理解,从全模态感知到遵循指令的响应。随后,我们提出了Faithful-RFT,这是一种强化微调阶段,进一步提高答案的真实性和表达质量,同时满足实时需求,直接通过任务可验证的反馈对最终响应进行评分,而不是在回滚过程中优化推理风格的探索。此外,TLive-Omni得益于一个面向场景的原子能力分类法和一个紧凑的数据生成引擎,将直播电商的音频、图像和视频流转换为语音识别、说话人分析、产品视觉定位、文本识别、时间定位、视频密集字幕和全模态问答等的训练信号。为了实现可扩展的训练,一个同步长度分组的采样器减少了填充,同时在工作者之间保持了可比的工作负载,而一种轻量级的动态采样策略则以接近零的奖励方差重新生成回滚组,以保持GRPO的有意义的相对优势。在电子商务直播基准上的实验展示了在直播电商领域任务中的强大表现,以及在一般基准上的出色泛化能力。
cs.AI / 62 / 2608.20960
Can Scientific Claims Be Removed from Large Language Models? A Systematic Evaluation of Claim-Level Unlearning
科学声明能否从大型语言模型中移除?声明级别遗忘的系统评估
Abstract
Language models (LMs) are trained on static scientific corpora, whereas scientific knowledge continuously evolves through correction and revision. Scientific claims encoded within these models may later become retracted, disproven, or updated by subsequent research, creating the risk of disseminating outdated information in scientific workflows. This creates a need for LMs to forget obsolete scientific claims. Machine unlearning offers a promising solution by enabling knowledge removal while maintaining overall model utility. Existing studies primarily investigate instance-level forgetting; however, scientific claims introduce additional challenges because they are interconnected, and continually evolving. To address this gap, we introduce the task of Scientific Claim Unlearning and present a new benchmark, SciUnlearn. We show that current unlearning approaches are unable to effectively eliminate claim-level knowledge and often achieve only superficial suppression, highlighting the need for specialized methods designed for structured knowledge removal.
Chinese Translation
语言模型(LMs)是在静态科学语料库上训练的,而科学知识则通过修正和修订不断演变。这些模型中编码的科学声明可能会在后续研究中被撤回、推翻或更新,从而导致在科学工作流程中传播过时信息的风险。这就需要语言模型遗忘过时的科学声明。机器遗忘提供了一种有前景的解决方案,通过实现知识的移除,同时保持模型的整体效用。现有研究主要探讨实例级遗忘;然而,科学声明带来了额外的挑战,因为它们是相互关联且持续演变的。为了解决这一空白,我们引入了科学声明遗忘的任务,并提出了一个新的基准,SciUnlearn。我们表明,当前的遗忘方法无法有效消除声明级知识,通常只能实现表面的抑制,突显了为结构化知识移除设计专门方法的必要性。
cs.AI / 63 / 2608.20961
TreeWY: Speculative Verification for Gated DeltaNet Hybrids
TreeWY:门控 DeltaNet 混合模型的推测验证
Abstract
Modern open models are hybrids: most layers are linear-attention (Gated DeltaNet, GDN) layers carrying a small fixed-size recurrent state instead of a growing key-value (KV) cache. This makes ordinary decoding memory-efficient, but hurts speculative decoding. To verify a batch of draft tokens and then roll back the rejected ones, today's systems snapshot the full recurrent state at every draft position for GDN layers, and those snapshots cannot be shared across branches of a draft tree, so a wide, high-acceptance tree becomes memory-infeasible. We remove the snapshots. Using a tree-structured WY transform of the gated delta rule, we compute every draft node's output with a single triangular solve and reconstruct only the one accepted state on commit, storing a small pseudo-value matrix instead of per-node states; the derivation depends only on the gated delta rule, not on any other architectural detail. In serving benchmarks on two scales of one hybrid model family (Qwen3.5 35B and 397B) this cuts speculative recurrent-state memory and KV-cache pressure at identical acceptance length, turning the freed HBM into higher throughput and much lower time-to-first-token (TTFT) wherever memory binds, and costing a few percent where it does not. For tree width the same memory buys affordability: a wider, higher-acceptance draft becomes possible, though not yet a throughput win.
Chinese Translation
现代开放模型是混合型的:大多数层是线性注意力(Gated DeltaNet, GDN)层,携带一个小的固定大小的递归状态,而不是一个不断增长的键值(KV)缓存。这使得普通解码在内存使用上更加高效,但对推测解码造成了负面影响。为了验证一批草稿令牌并回滚被拒绝的令牌,现有系统在每个草稿位置对 GDN 层的完整递归状态进行快照,而这些快照无法在草稿树的分支之间共享,因此宽而高接受度的树变得在内存上不可行。我们去除了快照。通过使用门控 delta 规则的树结构 WY 变换,我们使用单个三角形求解计算每个草稿节点的输出,并在提交时仅重建一个被接受的状态,存储一个小的伪值矩阵,而不是每个节点的状态;该推导仅依赖于门控 delta 规则,而不依赖于任何其他架构细节。在对一个混合模型家族的两个规模(Qwen3.5 35B 和 397B)进行基准测试时,这减少了推测递归状态的内存和 KV 缓存压力,在相同的接受长度下,将释放的 HBM 转化为更高的吞吐量和更低的首次令牌时间(TTFT),无论内存绑定情况如何,成本仅增加几个百分点。在树宽度方面,相同的内存使得更高的可承受性成为可能:更宽且接受度更高的草稿变得可行,尽管尚未实现吞吐量的提升。
cs.AI / 64 / 2608.20967
Generalizing Soft Tissue Deformation and Force Prediction Across Material Stiffness and Geometry
跨材料刚度和几何形状的软组织变形与力预测的泛化
Abstract
Accurate soft tissue simulation is essential for surgical training, pre-operative planning, and haptic feedback systems. While learning-based surrogate models trained on data using the finite element method (FEM) offer a promising path to real-time inference, their reliability depends on well-calibrated constitutive models. Existing approaches neither provide systematic guidance on model selection across stiffness levels, nor generalize across different tissue stiffnesses or geometries. We perform a comprehensive calibration of hyperelastic constitutive models in the SOFA Framework using gravity-loaded silicone beams with different stiffnesses. Using calibrated simulations as training data, we use a softness conditioned equivariant graph neural network, enabling deformation and force prediction across multiple tissue types and unseen geometries. Our model achieves sub-millimeter mean deformation accuracy at 0.010s inference time, while showing that force prediction quality is directly tied to upstream calibration consistency.
Chinese Translation
准确的软组织模拟对于外科培训、术前规划和触觉反馈系统至关重要。虽然基于学习的代理模型通过有限元法(FEM)训练的数据提供了实时推断的有希望的路径,但其可靠性依赖于良好校准的本构模型。现有方法既未提供跨刚度水平的模型选择系统指导,也未能在不同的组织刚度或几何形状之间进行泛化。我们在SOFA框架中对不同刚度的重力加载硅胶梁进行了超弹性本构模型的全面校准。利用校准后的模拟作为训练数据,我们使用了一种基于软度条件的等变图神经网络,使得能够在多种组织类型和未见几何形状之间进行变形和力的预测。我们的模型在0.010秒的推断时间内达到了亚毫米级的平均变形精度,同时表明力预测质量与上游校准一致性直接相关。
cs.AI / 65 / 2608.20970
Deep Learning Models Also Recall Features
深度学习模型也能回忆特征
Abstract
Recent work in mechanistic interpretability has studied how large language models recall facts stored in their weights. This paper argues that factual recall points to something broader: a general kind of operation in deep learning models, which I call feature recall. The core observation is that a linear projection can be read as retrieving stored information scaled by input activations. I define feature recall, show it applies across architectures, and contrast it with the established paradigm of feature combination. I also consider how cases of feature recall might be mechanistically identified. The account gives philosophers a new conceptual tool for understanding deep learning, and points to empirical directions for mechanistic interpretability research.
Chinese Translation
近期在机械解释性方面的研究探讨了大型语言模型如何回忆存储在其权重中的事实。本文认为,事实回忆指向一个更广泛的概念:深度学习模型中的一种普遍操作,我称之为特征回忆。核心观察是,线性投影可以被视为检索由输入激活缩放的存储信息。我定义了特征回忆,展示其在不同架构中的适用性,并将其与已建立的特征组合范式进行对比。我还考虑了如何在机制上识别特征回忆的案例。这一论述为哲学家提供了理解深度学习的新概念工具,并指向了机械解释性研究的实证方向。
cs.AI / 66 / 2608.20975
Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models
无行为的信念:测量心智理论在视觉-语言模型中转化为协调社会行动的能力
Abstract
Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and nonverbal channels simultaneously. Yet existing benchmarks evaluate theory of mind (ToM) reasoning and embodied behavior in isolation, leaving unmeasured the gap between social inference and social action. We introduce MOSAIC (Multimodal Orchestration of Social Action, Inference, and Communication), a controlled benchmark in which two embodied agents interact across cooperative and competitive scenarios requiring integration of verbal statements, spatial trajectories, gaze direction, and facial expression under systematically varied ToM constraints. Evaluating 13 models, including 11 VLMs, across 200 trials per model, we find that VLMs fail to produce behaviors consistent with the expected outcomes under ToM-order constraints, and that imposing explicit ToM-order constraints produces no reliable behavioral change aligned with the specified reasoning level. Signal-level analysis reveals two sequential bottlenecks: most models cannot produce directionally coherent nonverbal signals, and even when signals are present, VLM agents fail to interpret others behaviors and react to them. PCM-LLM, included as a structured architectural reference point with an explicit ToM module, succeeds across all conditions, suggesting that explicit belief-action coupling is a sufficient ingredient for this class of tasks.
Chinese Translation
有效的社会互动要求智能体将心理状态推断转化为在语言和非语言渠道中同时协调的行为信号。然而,现有基准评估心智理论(Theory of Mind, ToM)推理和具身行为时往往是孤立进行的,未能测量社会推断与社会行动之间的差距。我们提出了MOSAIC(社会行动、推断与沟通的多模态协调),这是一个受控基准,其中两个具身智能体在需要整合语言陈述、空间轨迹、注视方向和面部表情的合作和竞争场景中进行互动,并在系统变化的ToM约束下进行。对包括11个视觉-语言模型(VLMs)在内的13个模型进行评估,每个模型进行200次试验,我们发现VLMs未能产生与ToM顺序约束下预期结果一致的行为,并且施加明确的ToM顺序约束并未导致与指定推理水平一致的可靠行为变化。信号级分析揭示了两个顺序瓶颈:大多数模型无法产生方向一致的非语言信号,即使信号存在,VLM智能体也未能解读他人的行为并作出反应。作为具有明确ToM模块的结构性架构参考点,PCM-LLM在所有条件下均表现成功,表明明确的信念-行动耦合是这一类任务的一个充分条件。
cs.AI / 67 / 2608.21027
Don't Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents
不求解,仅比较:用于大规模语言模型代理的运行时干预微型顾问
Abstract
LLM agents are emerging as an important paradigm for real-world tasks that require reasoning, tool use, and sequential decision-making. As these agents operate over longer horizons, runtime intervention offers a way to improve reliability without retraining the underlying actor. Failure detection alone is insufficient. Effective intervention must also provide a useful direction for recovery. Existing approaches often rely on an expert solver or a critic that generates task-specific corrections, incurring either the cost of another capable solver or the capacity demands of a task-capable critic. We introduce Comparison-Only Tiny Advisor (COTA), a comparison-only framework for constructive runtime intervention. In COTA, a tiny comparator judges whether sampled alternatives lead to better continuations than the actor's proposal, and repeated comparisons determine when intervention is warranted. We train the comparator using pairwise supervision constructed from same-prefix counterfactual branches. Preferred alternatives are returned as non-binding advice, leaving the original actor to replan. Across WebShop, ALFWorld, and tau^3-Retail with three actors, COTA improves all nine evaluation settings and outperforms the compared baselines. These results show that constructive runtime intervention can remain effective even when the auxiliary model has substantially weaker task-solving capability than the actor.
Chinese Translation
大规模语言模型(LLM)代理正成为处理需要推理、工具使用和顺序决策的现实任务的重要范式。随着这些代理在更长时间范围内的操作,运行时干预提供了一种在不重新训练基础行为者的情况下提高可靠性的方法。仅仅检测失败是不够的。有效的干预还必须提供有用的恢复方向。现有的方法通常依赖于专家求解器或生成任务特定修正的批评者,这要么需要另一个有能力的求解器的成本,要么需要任务能力批评者的容量需求。我们提出了仅比较微型顾问(Comparison-Only Tiny Advisor, COTA),这是一个用于建设性运行时干预的仅比较框架。在COTA中,一个微型比较器判断采样的替代方案是否比行为者的提议导致更好的延续,重复的比较决定何时需要干预。我们使用从同前缀反事实分支构建的成对监督来训练比较器。优选的替代方案作为非约束性建议返回,留给原始行为者重新规划。在WebShop、ALFWorld和tau^3-Retail的三个行为者中,COTA在所有九个评估设置中均有所改善,并超越了比较的基线。这些结果表明,即使辅助模型的任务解决能力远低于行为者,建设性运行时干预仍然可以保持有效。
cs.AI / 68 / 2608.21036
Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance
评估大型语言模型在国际海运危险货物规则合规性上的表现
Abstract
The transport of dangerous goods by sea is a high-consequence activity governed by the International Maritime Dangerous Goods (IMDG) Code, a complex regulatory framework where errors in classification, packaging, stowage, or segregation can result in fire, explosion, toxic release, or loss of life or vessel. Correct compliance requires accurately interpreting hundreds of pages of interacting provisions, updated on a two-year amendment cycle. Practitioners increasingly use Large Language Models (LLMs) as decision-support tools, yet no systematic evaluation exists of whether they can reliably interpret IMDG requirements for safety-critical use. This paper introduces DGEval, the first benchmark for evaluating LLM knowledge of IMDG Amendment 42-24. Built from expert-written questions on the NCB Hazcheck e-learning platform and structured lookups from the Dangerous Goods List (DGL), it comprises 1,678 questions across multiple-choice, open-ended, DGL lookup, and regulatory identification tasks. We evaluate 13 models from six providers across multiple thinking configurations, including one maritime domain-specific fine-tuned model, and test the effect of web search. Although the best-performing model exceeds the human practitioner baseline on multiple-choice questions, all models are weakest in the operationally safety-critical areas of stowage, segregation, and regulatory recall. These results indicate that LLMs may support compliance tasks, particularly structured DGL lookups with web search, but unreliability in operational areas and regulatory-text recall means human oversight and authoritative source verification remain necessary before deployment in any safety-critical context. DGEval is designed as a safety assurance instrument to be applied continuously as models evolve, not as a settled characterisation of current capability.
Chinese Translation
海上危险货物运输是一项高风险活动,受《国际海运危险货物规则》(IMDG Code)监管,该规则是一个复杂的法规框架,其中分类、包装、装载或分隔的错误可能导致火灾、爆炸、毒气释放或人员及船只的损失。正确的合规性要求准确解读数百页相互关联的条款,这些条款每两年修订一次。实践者越来越多地使用大型语言模型(LLMs)作为决策支持工具,但尚无系统评估其是否能够可靠地解读IMDG要求以用于安全关键的场景。本文介绍了DGEval,这是第一个用于评估LLM对IMDG修订版42-24知识的基准。该基准由在NCB Hazcheck电子学习平台上编写的专家问题和来自危险货物清单(DGL)的结构化查找构成,共包含1,678个问题,涵盖多项选择题、开放式问题、DGL查找和法规识别任务。我们评估了来自六个提供者的13个模型,涵盖多种思维配置,包括一个特定于海事领域的微调模型,并测试了网络搜索的效果。尽管表现最佳的模型在多项选择题上超过了人类实践者的基准,但所有模型在操作安全关键领域(如装载、分隔和法规回忆)上表现最差。这些结果表明,LLMs可能支持合规任务,特别是带有网络搜索的结构化DGL查找,但在操作领域和法规文本回忆中的不可靠性意味着在任何安全关键环境中部署之前,仍需人类监督和权威来源验证。DGEval被设计为一个安全保证工具,旨在随着模型的演变持续应用,而不是作为当前能力的固定表征。
cs.AI / 69 / 2608.21044
Socialized Division and Collaboration: Rethinking Class-Incremental Learning under Optimization Conflicts
社会化分工与协作:在优化冲突下重新思考类别增量学习
Abstract
Class-incremental learning is commonly instantiated as a single-model paradigm, where a unified model sequentially adapts to an unbounded stream of sessions. While effective under mild distributional shifts, this formulation becomes strained when successive sessions induce incompatible optimization directions, leading to destructive interference and catastrophic forgetting. We argue that such forgetting reflects a structural limitation of enforcing heterogeneous learning dynamics within a single parameter space. Motivated by social solidarity theory, we propose Socialized Division and Collaboration (SDC) as a reformulation of continual learning that decomposes session learning across specialized models in response to optimization conflicts, while enabling coordinated collaboration. To support this formulation with a principled allocation mechanism, we introduce an energy-based session-model compatibility criterion grounded in Helmholtz free energy, which guides adaptive session allocation and model evolution under conflicting objectives. This framework integrates session assignment, model evolution, and collaborative inference into a unified pipeline, offering an alternative to monolithic continual learning formulations and highlighting a broader design principle for learning under persistent optimization conflicts.
Chinese Translation
类别增量学习通常被实例化为单模型范式,其中一个统一模型顺序地适应于无限制的会话流。尽管在轻微的分布变化下有效,但当连续会话引发不兼容的优化方向时,这种表述变得紧张,导致破坏性干扰和灾难性遗忘。我们认为,这种遗忘反映了在单一参数空间内强制执行异质学习动态的结构性限制。受社会团结理论的启发,我们提出社会化分工与协作(Socialized Division and Collaboration, SDC)作为持续学习的重新表述,它在应对优化冲突时,通过专门模型分解会话学习,同时实现协调协作。为了支持这一表述,我们引入了一种基于能量的会话-模型兼容性标准,该标准基于亥姆霍兹自由能,指导在冲突目标下的自适应会话分配和模型演化。该框架将会话分配、模型演化和协作推理整合为一个统一的流程,为单一的持续学习表述提供了替代方案,并强调了在持续优化冲突下学习的更广泛设计原则。
cs.AI / 70 / 2608.21059
The Cost of a Physics Prior Is Bounded by the Ablation Gap
物理先验的成本受限于剥离间隙
Abstract
Shape-constrained and physics-informed learning reports an accuracy cost of enforcing a prior and treats it as a property of the prior. We show it is mostly a property of the free features and the validation split. Let P be the excess risk of restricting a hypothesis class to functions with a shape constraint on features S, and D the excess risk of the ablated model that ignores S. Because a function constant in x_j is both non-decreasing and non-increasing in x_j, the ablated class is contained in the constrained class, so 0 <= P <= D for every risk functional, with no convexity, smoothness, or realizability assumption. Empirically the bound is a sign test: a constrained model must never be beaten by its own ablation. We instantiate it on an ordinal wildfire-severity task (N = 26,681, K = 3) with hard monotone constraints on four meteorological drivers, coordinates left free, and a validation ladder from i.i.d. resampling to 2-degree spatial blocking. Coordinates act as a shield: alone they recover 92.9% of the full model's macro-F1 under spatial blocking, collapsing D from 0.1288 to 0.0427; the same prior costs 0.0473 shielded and 0.3470 unshielded, a ratio of 7.3 with identical physics. Because D is protocol-dependent it does not transfer: coarsening blocks from 1 to 10 degrees drives D from 0.0942 to 0.0050, leaving two configurations unidentifiable a priori. Inversions of the certified nesting bound the pipeline's additive resolution: over 318 comparisons they give a self-calibrating floor of 0.0220 macro-F1, below which no reported price is interpretable, including four cells in our own headline grid. Cost and compliance are independent: the unconstrained model violates the prior at rate 0.48-0.49 while enforcing it costs 0.0473. We give a two-fit screen that rejects unidentifiable experiments before a constrained model is trained.
Chinese Translation
形状约束和物理信息学习报告了强制施加先验的准确性成本,并将其视为先验的属性。我们表明,这主要是自由特征和验证划分的属性。设 P 为将假设类限制为具有特征 S 的形状约束函数的过度风险,D 为忽略 S 的剥离模型的过度风险。由于在 x_j 上常数的函数在 x_j 上既是非递增的又是非递减的,因此剥离类包含在约束类中,因此对于每个风险泛函都有 0 <= P <= D,且不需要凸性、光滑性或可实现性的假设。从经验上看,这个界限是一个符号测试:约束模型绝不能被其自身的剥离所击败。我们在一个有序的野火严重性任务上实例化它(N = 26,681,K = 3),对四个气象驱动因素施加严格的单调约束,坐标保持自由,并从独立同分布重采样到 2 度空间阻塞的验证阶梯。坐标充当了一个屏障:单独它们在空间阻塞下恢复了完整模型的 92.9% 的宏观 F1,D 从 0.1288 降至 0.0427;相同的先验在有屏障和无屏障情况下的成本分别为 0.0473 和 0.3470,比例为 7.3,物理条件相同。由于 D 依赖于协议,因此它不具备可转移性:将阻塞从 1 度粗化到 10 度使 D 从 0.0942 降至 0.0050,留下两个配置在先验上不可识别。经过认证的嵌套界限反转了管道的附加分辨率:在 318 次比较中,它们提供了一个自校准的宏观 F1 下限为 0.0220,低于该值的任何报告价格都是不可解释的,包括我们自己主标题网格中的四个单元。成本和合规性是独立的:无约束模型以 0.48-0.49 的速率违反先验,而强制执行先验的成本为 0.0473。我们提供了一个双拟合筛选器,在训练约束模型之前拒绝不可识别的实验。
cs.AI / 71 / 2608.21060
CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models
CellPath-Bench:用于病理基础模型中全幻灯片细胞表征的多维基准测试
Abstract
Pathology foundation models (PFMs) are increasingly used as general-purpose backbones, yet existing benchmarks cannot systematically diagnose their whole-slide cellular representation capabilities, including the decodability of cell-type information and the transferability of such information across tissue sections, datasets, and anatomical organs. We introduce CellPath-Bench, a cellular-resolution benchmark that evaluates frozen PFMs themselves. Following quality control of 52 candidate Xenium datasets, we construct a panel of 25 spatially aligned H\&E--Xenium tissue sections spanning 11 organs and 7,079,283 cells, harmonized into fine- and coarse-grained taxonomies. CellPath-Bench samples frozen WSI feature maps at registered nuclear coordinates and evaluates them using standardized multiclass linear probes. Cell Representation Advantage (CRA) measures the within-section advantage of nucleus-anchored representations over patch-level mean pooling, while Cell Representation Transferability (CRT) characterizes the generalization of cell-type decodability across tissue sections, datasets, and organs. We benchmark 30 pathology-specific and general-purpose foundation models through 304,920 runs across spatial readouts, magnifications, taxonomic granularities, and evaluation protocols. The results reveal substantial model-dependent differences in cell-type decodability and its cross-domain generalization, yielding distinct multidimensional capability profiles. CellPath-Bench provides a standardized framework for auditing cellular information in frozen PFM representations.
Chinese Translation
病理基础模型(PFMs)越来越多地被用作通用骨干网络,但现有基准无法系统地诊断其全幻灯片细胞表征能力,包括细胞类型信息的可解码性以及这种信息在组织切片、数据集和解剖器官之间的可转移性。我们引入了CellPath-Bench,这是一个评估冻结PFMs自身的细胞分辨率基准。在对52个候选Xenium数据集进行质量控制后,我们构建了一个包含25个空间对齐的H&E--Xenium组织切片的面板,涵盖11个器官和7,079,283个细胞,并将其协调为细粒度和粗粒度分类法。CellPath-Bench在注册的核坐标下对冻结的WSI特征图进行采样,并使用标准化的多类线性探针进行评估。细胞表征优势(Cell Representation Advantage, CRA)衡量以细胞核为锚的表征相对于块级均值池化的切片内优势,而细胞表征可转移性(Cell Representation Transferability, CRT)则表征细胞类型可解码性在组织切片、数据集和器官之间的泛化能力。我们通过304,920次运行在空间读出、放大倍数、分类粒度和评估协议中对30个特定于病理和通用的基础模型进行了基准测试。结果揭示了细胞类型可解码性及其跨域泛化中显著的模型依赖性差异,形成了独特的多维能力轮廓。CellPath-Bench为审计冻结PFM表征中的细胞信息提供了一个标准化框架。
cs.AI / 72 / 2608.21089
Can Legal AI Know When It Is Wrong? And Do Students Know When It Is?
法律人工智能能知道何时是错误的吗?学生们知道吗?
Abstract
Integrating Large Language Models (LLMs) into the Indian judiciary promises access to justice but introduces severe risks. We identify the 'inertia of confidence'--an overconfidence phenomenon analogous to the Dunning-Kruger effect where LLMs provide incorrect legal verdicts with near-maximum confidence, driven by a hypothesized 'precedent overfitting' bias. Phase I of our socio-technical audit tested ChatGPT (GPT-5.2), Meta AI, and Perplexity AI on a 60-case battery regarding the Indian Contract Act, 1872, and the shift toward statutory enforcement of specific performance. We introduce the High-Confidence Error Rate (HCER) to quantify incorrect verdicts delivered with dangerous certainty (>= 9 on a 1-10 scale). All models struggled with statutory updates. Meta AI proved most vulnerable (31.7% HCER), frequently misapplying pre-amendment rules with a 9.1/10 mean confidence, followed by Perplexity (15.0%) and ChatGPT (6.7%). Phase II investigated human vulnerability to this overconfidence via a survey of Indian law students (N=380). Verification often functions as a reactive adaptation to machine hallucinations: students encountering fabricated citations reported higher verification scores (4.2/5) than those with no such encounters (2.8/5). Furthermore, while 81.6% knew submitting hallucinated cases can lead to contempt-of-court, 71.1% received no formal training on ethical AI use. We propose shifting toward adversarial legal research pedagogy and implementing source-grounded verification architectures to prevent systemic professional negligence.
Chinese Translation
将大型语言模型(LLMs)整合到印度司法系统中,承诺提供公正的机会,但也带来了严重的风险。我们识别出一种“信心惯性”现象——一种与邓宁-克鲁格效应相似的过度自信现象,其中LLMs以接近最大信心提供错误的法律裁决,这种现象受到假设的“先例过拟合”偏见的驱动。我们的社会技术审计第一阶段对ChatGPT(GPT-5.2)、Meta AI和Perplexity AI进行了60个案例的测试,涉及1872年印度合同法及其对特定履行法定强制执行的转变。我们引入了高信心错误率(HCER)来量化以危险的确定性(>= 9分,满分10分)提供的错误裁决。所有模型在处理法定更新时都表现不佳。Meta AI表现最脆弱(31.7% HCER),经常错误应用修订前的规则,平均信心为9.1/10,其次是Perplexity(15.0%)和ChatGPT(6.7%)。第二阶段通过对380名印度法学生的调查研究了人类对这种过度自信的脆弱性。验证通常作为对机器幻觉的反应性适应:遇到虚构引用的学生报告的验证分数(4.2/5)高于没有此类遭遇的学生(2.8/5)。此外,尽管81.6%的学生知道提交虚构案件可能导致藐视法庭,但71.1%的学生没有接受过有关伦理人工智能使用的正式培训。我们建议转向对抗性法律研究教学法,并实施基于来源的验证架构,以防止系统性职业失职。
cs.AI / 73 / 2608.21097
When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge
当信任遇见真相:LLM作为评判者中的信任与真相的可分离性
Abstract
LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, and factuality, and these outputs are often interpreted as independent evidence. We test this assumption for a common pair of judgments: trust scoring and binary truth classification. On correctness-controlled QA, LLM judges align trust scores with truth verdicts more tightly than human behavioral reference, suggesting weaker separations between trust and truth judgment. We then apply stress tests by changing only source cues of identical QA between Human and AI. Source attribution shifts not only trust scores but also truth verdicts and logit-derived correct-side probabilities. Results show that current LLM-as-Judge protocols should not treat trust scores as independent evidence for truth judgments.
Chinese Translation
LLM作为评判者的系统能够产生多维度的评估,例如可信度、可靠性和事实性,这些输出通常被解读为独立的证据。我们测试了这一假设,针对一对常见的判断:信任评分和二元真相分类。在控制正确性的问答(QA)中,LLM评判者将信任评分与真相裁决的对齐程度比人类行为参考更紧密,这表明信任与真相判断之间的分离较弱。随后,我们通过仅改变人类与人工智能之间相同问答的源提示进行压力测试。源归属的变化不仅影响信任评分,还影响真相裁决和基于对数几率推导的正确侧概率。结果表明,当前的LLM作为评判者的协议不应将信任评分视为真相判断的独立证据。
cs.AI / 74 / 2608.21100
ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models
ReFrame:多模态大型语言模型中的证据引导测试时安全对齐
Abstract
While multimodal large language models (MLLMs) extend model capabilities beyond text, they also make safety alignment increasingly challenging. Multimodal safety alignment methods must address cross-modal jailbreaks, safety-awareness failures, and over-sensitive refusals. However, existing methods often rely on retraining or internal-state inspection, limiting their applicability to deployed closed-source MLLMs and motivating test-time safety alignment. We analyze this setting and identify two key obstacles, utility dominance and reasoning inertia, which cause models to overlook latent risks or follow malicious reasoning trajectories. Guided by these insights, we propose ReFrame, a training-free multimodal input reframing framework where two agents share a lightweight locally deployed MLLM: the evidence-generation agent constructs complementary risk and utility evidence, and the rewrite-and-routing agent converts it into a safe proxy prompt and image-routing decision before calling the downstream MLLM, without modifying it or accessing its internal information. Experiments across multiple MLLMs and benchmarks show that ReFrame improves jailbreak defense, safety awareness, and oversensitivity reduction while preserving multimodal utility.
Chinese Translation
尽管多模态大型语言模型(MLLMs)扩展了模型的能力超越文本,但它们也使安全对齐变得愈加具有挑战性。多模态安全对齐方法必须应对跨模态越狱、安全意识失败和过度敏感拒绝等问题。然而,现有方法通常依赖于重新训练或内部状态检查,这限制了它们在已部署的闭源MLLMs中的适用性,并激励了测试时的安全对齐。我们分析了这一情境,并识别出两个关键障碍:效用主导和推理惯性,这导致模型忽视潜在风险或遵循恶意推理轨迹。在这些见解的指导下,我们提出了ReFrame,一个无训练的多模态输入重构框架,其中两个代理共享一个轻量级本地部署的MLLM:证据生成代理构建互补的风险和效用证据,而重写与路由代理在调用下游MLLM之前将其转换为安全的代理提示和图像路由决策,而无需修改模型或访问其内部信息。在多个MLLM和基准测试中的实验表明,ReFrame在增强越狱防御、安全意识和减少过度敏感性方面表现出色,同时保持多模态效用。
cs.AI / 75 / 2608.21107
Large Language Models at the Intersection of Software Engineering and Software Security:An Evidence-Centered Structured Survey and Research Agenda
大型语言模型在软件工程与软件安全交汇处:以证据为中心的结构化调查与研究议程
Abstract
Large Language Models (LLMs) are moving from code completion toward repository-scale agents that retrieve context, edit files, execute tools, and participate in security-sensitive workflows. The evidence for these systems, however, remains divided between software engineering evaluations centered on functional task completion and software security evaluations centered on vulnerability detection, secure generation, or exploit-oriented validation. This evidence-centered structured survey synthesizes representative work available through May 31, 2026 across software engineering tasks, software security tasks, adaptation mechanisms, artifact granularity, and evaluation design. In addition to a task taxonomy, we introduce an assurance framework that separates functional correctness, security, operational reliability, evidence provenance, and agent authority. The review shows that execution feedback and repository access can substantially improve engineering task completion, but do not by themselves establish security; conversely, static-analysis labels or vulnerability-classification scores rarely establish deployable correctness. We identify recurring validity threats--weak test oracles, duplicated and temporally leaked data, changing agent harnesses, proxy-only security checks, and under-reported budgets and human intervention--and derive a minimum reporting protocol for cross-study comparison. The resulting research agenda prioritizes jointly secure-and-functional benchmarks, repository-scale threat models, calibrated human oversight, longitudinal maintainability evidence, and reproducible agent evaluation. The central conclusion is that model capability should be judged as an assurance case supported by task-appropriate evidence, rather than by a single benchmark score.
Chinese Translation
大型语言模型(LLMs)正从代码补全向能够检索上下文、编辑文件、执行工具并参与安全敏感工作流程的仓库级代理转变。然而,这些系统的证据仍然在以功能任务完成为中心的软件工程评估和以漏洞检测、安全生成或针对利用的验证为中心的软件安全评估之间存在分歧。本以证据为中心的结构化调查综合了截至2026年5月31日可用的代表性研究,涵盖软件工程任务、软件安全任务、适应机制、工件粒度和评估设计。除了任务分类法外,我们还引入了一个保证框架,区分功能正确性、安全性、操作可靠性、证据来源和代理权限。回顾表明,执行反馈和仓库访问可以显著提高工程任务的完成度,但单靠这些并不能确立安全性;相反,静态分析标签或漏洞分类评分很少能确立可部署的正确性。我们识别出反复出现的有效性威胁——弱测试神谕、重复和时间泄露的数据、变化的代理环境、仅限代理的安全检查,以及报告不足的预算和人为干预——并推导出一个最低报告协议以便于跨研究比较。最终的研究议程优先考虑联合安全与功能基准、仓库级威胁模型、经过校准的人类监督、纵向可维护性证据和可重复的代理评估。核心结论是,模型能力应作为一个由任务适当证据支持的保证案例来评判,而不是仅依赖单一的基准分数。
cs.AI / 76 / 2608.21117
Root cause analysis via difference graph discovery from linear time-series data
通过差异图发现进行根本原因分析的线性时间序列数据
Abstract
Root cause analysis aims to identify the mechanisms responsible for anomalies in complex dynamical systems. In this paper, we study root cause analysis in linear time-series through the lens of difference graph discovery. We focus on effect-defying root causes, corresponding to variables whose causal coefficients change between a normal and an anomalous regime. We formalize this problem using linear discrete-time dynamic structural causal models and adapt several methods originally introduced for discovering difference graphs between two populations to the time-series setting, where the two populations are replaced by a normal and an anomalous regime. We first evaluate the proposed approaches on simulated data, and then demonstrate their practical relevance on real-world datasets from IT monitoring and intensive care monitoring. Our results show how difference graph discovery can help localize causal mechanisms responsible for anomalous behavior.
Chinese Translation
根本原因分析旨在识别复杂动态系统中导致异常的机制。本文通过差异图发现的视角研究线性时间序列中的根本原因分析。我们关注于效应违背的根本原因,这些根本原因对应于在正常和异常状态下因果系数发生变化的变量。我们使用线性离散时间动态结构因果模型对该问题进行形式化,并将原本用于发现两个群体之间差异图的几种方法适应于时间序列设置,其中两个群体被正常状态和异常状态所替代。我们首先在模拟数据上评估所提出的方法,然后在来自IT监控和重症监护监测的真实数据集上展示其实际相关性。我们的结果表明,差异图发现如何帮助定位导致异常行为的因果机制。
cs.AI / 77 / 2608.21174
From Attention Masks to Inert Zero-Vector Tokens: OAttention and O-Closure for Token Dynamics
从注意力掩码到惰性零向量标记:OAttention 和 O-Closure 在标记动态中的应用
Abstract
Attention masks are relation-level controls: they specify which query--source pairs may interact. They do not provide a representation-carried token state that is non-participating at the attention boundary. We assign each token hidden carrier \(h_i\) an active-presence coefficient \(p_i=\lVert h_i\rVert^2/(\tau+\lVert h_i\rVert^2)\). The same coefficient has two roles: it gates information emitted by token \(i\), and it determines the mass with which token \(i\) enters computations shared with other tokens. OAttention is the support-coupled attention realization of this rule. It gates the receiver output by \(p_i\) and weights source \(j\) by \(p_j\) in both the attention numerator and partition, while retaining the standard score, visibility relation, exponential competition, and value aggregation. This makes the zero-vector token a zero element and yields exact null-receiver, null-source insertion, self-attention insertion, and empty-support properties. The same token-level presence gives local O-components (OFFN, ONorm, and OInject), presence-weighted OStandardize, the O-Closure law \(M(H\oplus0)=M(H)\oplus0\), and an OTransformer by residual and compositional closure. The canonical operator is checked by contract tests and a GPU evaluation. In a zero-fine-tuning retrofit of a cloned pretrained TabPFN v3 regressor, calibrated hidden-carrier OAttention and Full-O variants change mean RMSE by $+0.088\%$ and $+0.177\%$, respectively, over 18 matched dataset--seed cases. A two-block ablation shows that OAttention alone does not preserve a NULL state through ordinary host components, whereas the OTransformer path does. These are scoped tests of exactness, active-path compatibility, and compositional necessity; they do not establish universal no-loss, arbitrary-host closure, learned attraction to the origin, or a general semantics for missing values.
Chinese Translation
注意力掩码是关系级别的控制:它们指定哪些查询-源对可以相互作用。它们并未提供在注意力边界上不参与的表示承载标记状态。我们为每个标记的隐藏承载体 $h_i$ 指定一个主动存在系数 $p_i=rac{
orm{h_i}^2}{ au+
orm{h_i}^2}$。该系数具有两个作用:它控制标记 $i$ 发出的信息,并决定标记 $i$ 进入与其他标记共享计算的质量。OAttention 是该规则的支持耦合注意力实现。它通过 $p_i$ 控制接收器输出,并在注意力的分子和分区中按 $p_j$ 加权源 $j$,同时保留标准得分、可见性关系、指数竞争和价值聚合。这使得零向量标记成为零元素,并产生精确的零接收器、零源插入、自注意力插入和空支持属性。相同的标记级别存在赋予了局部 O 组件(OFFN、ONorm 和 OInject)、加权存在的 OStandardize、O-Closure 法则 $M(Higoplus0)=M(H)igoplus0$,以及通过残差和组合闭包形成的 OTransformer。通过合同测试和 GPU 评估检查规范算子。在对克隆的预训练 TabPFN v3 回归器进行零微调改造中,经过校准的隐藏承载 OAttention 和 Full-O 变体在 18 个匹配的数据集-种子案例中,均方根误差均值分别变化 $+0.088 ext{ extperthousand}$ 和 $+0.177 ext{ extperthousand}$。一个两块消融实验表明,仅 OAttention 并未通过普通主机组件保持 NULL 状态,而 OTransformer 路径则保持了这一状态。这些是关于精确性、主动路径兼容性和组合必要性的范围测试;它们并未建立普遍无损、任意主机闭合、对原点的学习吸引或缺失值的一般语义。
cs.AI / 78 / 2608.21203
SENTRY: Deterministic, Intelligent Risk Assessment for IT Change Management
SENTRY:用于IT变更管理的确定性智能风险评估
Abstract
Technology change management in large financial institutions depends on risk assessments that are accurate, consistent, and auditable. In practice, many institutions still rely on self-reported questionnaires. Those questionnaires are subjective, easy to game, and poor at separating routine changes from the ones that later trigger major incidents. This paper presents SENTRY, a risk assessment platform that replaces questionnaire-based scoring with a deterministic machine learning pipeline built from gradient-boosted decision trees (XGBoost) and hybrid retrieval-augmented generation (RAG). The system combines structured operational metadata, application dependency graphs, and historical incident records with a hybrid semantic and lexical search over historical change requests. The retrieval step captures the risk signal in unstructured change request text, then compresses that signal into a single scalar feature before model inference. That design keeps the model deterministic and preserves per-prediction explainability via SHAP values. Evaluated on enterprise-scale change data, SENTRY achieves a ROC AUC of 0.87 and 85% overall accuracy, and it detects high-risk changes at roughly 3.25 times the rate of the existing process. We close by examining the architectural trade-offs behind this design and what they imply for the use of machine learning in regulated change management.
Chinese Translation
大型金融机构的技术变更管理依赖于准确、一致且可审计的风险评估。在实践中,许多机构仍然依赖自我报告的问卷。这些问卷具有主观性,容易被操控,并且在区分常规变更与后续引发重大事件的变更方面效果较差。本文提出了SENTRY,一个风险评估平台,采用基于确定性机器学习管道的评分替代问卷,该管道由梯度提升决策树(XGBoost)和混合检索增强生成(RAG)构建。该系统结合了结构化的操作元数据、应用依赖图和历史事件记录,并对历史变更请求进行了混合语义和词汇搜索。检索步骤捕捉到非结构化变更请求文本中的风险信号,然后将该信号压缩为一个单一的标量特征,供模型推断使用。这种设计保持了模型的确定性,并通过SHAP值保留了每次预测的可解释性。在企业级变更数据上的评估显示,SENTRY实现了0.87的ROC AUC和85%的整体准确率,并且以大约3.25倍于现有流程的速度检测高风险变更。最后,我们探讨了这一设计背后的架构权衡及其对受监管变更管理中机器学习使用的影响。
cs.AI / 79 / 2608.21209
Personalized Privacy Control in LLMs via Attention Head Intervention
通过注意力头干预实现大语言模型中的个性化隐私控制
Abstract
The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns. Prior work on contextual privacy studies whether LLMs regulate information disclosure according to context-dependent norms. However, acceptable disclosure boundaries may vary across users even within the same context. To address this limitation, we introduce \textit{personalized privacy}, which incorporates user-specific disclosure preferences into privacy control. We further present P3Bench~(\textbf{P}ersonalized \textbf{P}rivacy \textbf{P}reservation \textbf{Bench}mark), a novel benchmark extending contextual privacy policies with personalized disclosure policies. Experiments show that prompt-based policies fail to reliably enforce personalized privacy policies, with Qwen2.5-7B and Gemma3-4B showing average policy ignorance ratios of 51.25\% and 74.28\%, respectively. Finally, to address this problem, we propose \textsc{Repair}, a robust inference-time attention head intervention method that adjusts disclosure behavior toward policy-consistent responses. Our method significantly improves adherence to user-specific privacy preferences by reducing cases where the model fails to follow the given policy.
Chinese Translation
代理智能的兴起使得大语言模型(LLMs)能够访问多样化的用户数据,这引发了重要的隐私问题。先前关于上下文隐私的研究探讨了大语言模型是否根据上下文依赖的规范来调节信息披露。然而,即使在相同的上下文中,用户可接受的披露边界也可能有所不同。为了解决这一局限性,我们引入了 extit{个性化隐私},将用户特定的披露偏好纳入隐私控制中。我们进一步提出了P3Bench( extbf{个性化隐私保护基准}),这是一个新颖的基准,扩展了上下文隐私政策与个性化披露政策。实验表明,基于提示的政策无法可靠地执行个性化隐私政策,Qwen2.5-7B和Gemma3-4B的平均政策无视比例分别为51.25 ext{%}和74.28 ext{%}。最后,为了解决这一问题,我们提出了 extsc{Repair},一种稳健的推理时注意力头干预方法,该方法调整披露行为以实现与政策一致的响应。我们的方法通过减少模型未能遵循给定政策的情况,显著提高了对用户特定隐私偏好的遵循程度。
cs.AI / 80 / 2608.21218
Enhancing LLMs in Predictive Political QA with Semi-Structured Data
利用半结构化数据增强预测政治问答中的大型语言模型
Abstract
Predictive political question answering (QA), such as predicting how a political actor will vote, goes beyond factual lookup. External political resources offer rich historical evidence, but rarely contain the answer itself. Existing LLM augmentation methods, including actor-profile-based simulation and knowledge graph evidence injection, improve political reasoning but largely treat external resources as knowledge-based evidence, leaving prediction-relevant signals under-modeled. We identify two complementary signals for predictive political QA: actor stances that capture issue-specific preferences, and high-order structure signals that capture indirect dependencies among political actors. We propose PSL, a dual-view framework that converts semi-structured political records into inference-oriented evidence for LLMs. PSL extracts stance signals from question-relevant actor records in a semantic view, and learns structure-aware actor representations from an actor interaction graph in a vector view. Across three real-world datasets and multiple LLMs, PSL consistently outperforms baselines, with ablations confirming the complementary gains of stance and structure signals.
Chinese Translation
预测政治问答(QA),例如预测政治行为者将如何投票,超越了简单的事实查找。外部政治资源提供了丰富的历史证据,但很少包含答案本身。现有的LLM增强方法,包括基于行为者档案的模拟和知识图谱证据注入,改善了政治推理,但在很大程度上将外部资源视为基于知识的证据,导致与预测相关的信号建模不足。我们识别出预测政治问答的两个互补信号:捕捉特定问题偏好的行为者立场,以及捕捉政治行为者间间接依赖关系的高阶结构信号。我们提出了PSL,一个双视角框架,将半结构化政治记录转换为面向推理的LLM证据。PSL从与问题相关的行为者记录中提取立场信号,并从行为者交互图中学习结构感知的行为者表示。在三个真实世界数据集和多个LLM上,PSL始终优于基线,消融实验确认了立场和结构信号的互补增益。
cs.AI / 81 / 2608.21224
Ontology-supported AI Model and Dataset Management
本体支持的人工智能模型与数据集管理
Abstract
Recently, there has been a great deal of research into improving AI methods and their application. The main focus is on tracking progress, enabling transparent comparisons, and fostering a more profound understanding of AI. In that process, different organizations generate and use plenty of assets that need to be tracked, traced and managed. Moreover, it is important to discover assets relevant for the task at hand. This paper presents research aiming to contribute to answering the question of what is required to exchange and manage AI models and related assets effectively without semantic gaps in an industrial context. We introduce a platform for AI model exchange, which facilitates the usage, exchange, and analysis of AI models and datasets. The platform incorporates an ontology that can foster a more profound common understanding of what is required in these tasks and help tackle the issues mentioned above. Finally, we elucidate the utility of the platform through the illustration of a use case in the context of real-time critical systems.
Chinese Translation
近年来,关于改进人工智能方法及其应用的研究取得了显著进展。主要关注点在于跟踪进展、实现透明比较以及促进对人工智能的更深入理解。在这一过程中,不同组织生成和使用大量资产,这些资产需要被跟踪、追溯和管理。此外,发现与当前任务相关的资产也至关重要。本文旨在研究如何在工业背景下有效地交换和管理人工智能模型及相关资产,而不产生语义差距。我们介绍了一个用于人工智能模型交换的平台,该平台促进了人工智能模型和数据集的使用、交换和分析。该平台结合了一个本体,可以促进对这些任务所需内容的更深入的共同理解,并帮助解决上述问题。最后,我们通过在实时关键系统背景下的一个用例来阐明该平台的实用性。
cs.AI / 82 / 2608.21233
Fine-Grain GPU Parallelization of the Generalized Partition Crossover for Large-Scale Traveling Salesman Problems
大规模旅行商问题的广义分区交叉的细粒度GPU并行化
Abstract
The Traveling Salesman Problem (TSP) is one of the most extensively studied NP-hard optimization problems. Genetic Algorithm (GA)-based solvers, such as the Edge Assembly Crossover (EAX), achieve state-of-the-art performance on many benchmark instances. However, the scalability of these approaches in massively parallel architectures remains limited because crossover operations involve irregular memory access patterns, graph traversals, and sequential dependencies. Existing GPU-based TSP solvers primarily exploit population-level parallelism and are limited to relatively small problem sizes. This work presents a fine-grain GPU implementation of the partition phase of the Generalized Partition Crossover (GPX) operator for large-scale TSP instances. The proposed approach reformulates GPX partitioning as a graph-parallel problem using coalesced memory layouts, ghost-node transformations, and connected-component analysis. The im- plementation parallelizes the union of parent tours, the splitting of degree- four vertices, the deletion of common edges, and the identification of recombining components using CUDA. Experimental results on instances ranging from 10,000 to 2 million cities demonstrate substantial acceleration over a naive sequential CPU imple- mentation. The proposed GPU partitioning achieves speedups between 48x and 625x while significantly reducing memory overhead. The re- sults demonstrate that operator-level parallelism can substantially im- prove the scalability of GA-based TSP solvers on modern many-core architectures.
Chinese Translation
旅行商问题(TSP)是最广泛研究的NP-hard优化问题之一。基于遗传算法(GA)的求解器,如边组装交叉(EAX),在许多基准实例上实现了最先进的性能。然而,由于交叉操作涉及不规则的内存访问模式、图遍历和顺序依赖,这些方法在大规模并行架构中的可扩展性仍然有限。现有的基于GPU的TSP求解器主要利用种群级别的并行性,并且仅限于相对较小的问题规模。本研究提出了一种针对大规模TSP实例的广义分区交叉(GPX)算子的分区阶段的细粒度GPU实现。所提出的方法将GPX分区重新表述为一个图并行问题,使用合并内存布局、幽灵节点转换和连通分量分析。该实现通过CUDA并行化父游历的联合、四度顶点的分裂、公共边的删除以及重组组件的识别。对从10,000到200万城市的实例的实验结果表明,相较于简单的顺序CPU实现,显著加速。所提出的GPU分区实现了48倍到625倍的加速,同时显著减少了内存开销。结果表明,算子级别的并行性可以显著提高基于GA的TSP求解器在现代多核架构上的可扩展性。
cs.AI / 83 / 2608.21278
CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment
CLEAR:用于保持效用的大型语言模型安全对齐的连续潜在适配器路由
Abstract
Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safety tuning may affect model responses to both harmful and benign inputs. We propose \textbf{C}ontinuous \textbf{L}at\textbf{E}nt \textbf{A}dapter \textbf{R}outing (CLEAR), a conditional safety adaptation framework that uses a lightweight hidden-state gate to continuously control the activation strength of a safety low-rank adapter. CLEAR aims to reduce harmful completions while avoiding unnecessary changes to the frozen backbone that could degrade performance on benign prompts. Experiments on widely used safety and utility benchmarks show that CLEAR improves robustness on HarmBench while reducing the utility degradation observed with globally applied safety tuning such as SFT or standard low-rank adaptation (LoRA). On Llama-3-8B-Instruct, CLEAR reduces HarmBench ASR from 32.3\% to 0.5\%, while retaining most of the base model's utility and achieving up to 7.1 percentage points higher GSM8K accuracy than globally applied SFT or LoRA. These results suggest that CLEAR is a promising mechanism for improving the safety--utility trade-off in LLM alignment.
Chinese Translation
提高大型语言模型(LLMs)的安全性往往以效用为代价,因为全局应用的安全调优可能会影响模型对有害和良性输入的响应。我们提出了 extbf{C}ontinuous extbf{L}at extbf{E}nt extbf{A}dapter extbf{R}outing(CLEAR),这是一种条件安全适配框架,使用轻量级的隐藏状态门控来持续控制安全低秩适配器的激活强度。CLEAR旨在减少有害的生成,同时避免对冻结主干网络的不必要更改,以免降低对良性提示的性能。在广泛使用的安全性和效用基准测试上的实验表明,CLEAR提高了HarmBench上的鲁棒性,同时减少了全局应用安全调优(如SFT或标准低秩适配(LoRA))所观察到的效用下降。在Llama-3-8B-Instruct上,CLEAR将HarmBench的ASR从32.3\%降低到0.5\%,同时保留了大部分基础模型的效用,并且在GSM8K准确率上比全局应用的SFT或LoRA高出最多7.1个百分点。这些结果表明,CLEAR是改善LLM对齐中安全性与效用权衡的有希望的机制。
cs.AI / 84 / 2608.21292
AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization
AUSO:从内化到利用的动作级统一技能优化
Abstract
Skills play different roles as an agent's policy evolves: they should first provide learnable knowledge, then support capability formation, and finally be invoked only when they improve individual decisions. Existing methods rarely model this lifecycle. They either keep skills outside the model, fully internalize them, or select among internalization and utilization objectives through noisy task-level success rates. Such designs fragment training and assign uniform importance to actions within the same trajectory, even though skill guidance may help some decisions while distracting others. To solve these problems, we introduce AUSO (Action-level Unified Skill Optimization), which unifies skill learning and skill use through a progressive, action-aware optimization process. At the beginning of training, AUSO jointly learns from teacher guidance and environmental outcomes, enabling the policy to acquire foundational skills without losing task-oriented feedback. It subsequently emphasizes outcome-based policy optimization to consolidate autonomous problem-solving ability. As the policy matures, AUSO evaluates each sampled action under both skill-conditioned and skill-free contexts. The resulting action-level information signal is coupled with the trajectory outcome advantage, allowing beneficial skill-sensitive actions to receive stronger updates and harmful ones to be suppressed. Therefore, skills gradually transition from an external source of supervision into decision knowledge whose utilization is adapted to its action-level benefit, while reinforcement learning remains the shared backbone across all stages. Experiments on ALFWorld, WebShop, and SearchQA show that AUSO consistently improves agent performance and out-of-distribution generalization over competitive baselines.
Chinese Translation
技能在智能体策略演变过程中扮演着不同的角色:它们首先应提供可学习的知识,然后支持能力的形成,最后仅在能够改善个体决策时被调用。现有方法很少对这一生命周期进行建模。它们要么将技能置于模型之外,要么完全内化它们,或者通过嘈杂的任务级成功率在内化和利用目标之间进行选择。这种设计使得训练过程碎片化,并对同一轨迹中的动作赋予统一的重要性,尽管技能指导可能对某些决策有帮助,而对其他决策则可能造成干扰。为了解决这些问题,我们引入了AUSO(动作级统一技能优化),通过一个渐进的、关注动作的优化过程统一技能学习和技能使用。在训练开始时,AUSO通过教师指导和环境结果共同学习,使得策略能够在不失去任务导向反馈的情况下获得基础技能。随后,它强调基于结果的策略优化,以巩固自主问题解决能力。随着策略的成熟,AUSO在技能条件和无技能条件下评估每个采样动作。由此产生的动作级信息信号与轨迹结果优势相结合,使得有益的技能敏感动作获得更强的更新,而有害的动作则受到抑制。因此,技能逐渐从外部监督源转变为决策知识,其利用方式适应于其动作级的收益,同时强化学习在所有阶段仍然是共享的基础。对ALFWorld、WebShop和SearchQA的实验表明,AUSO在竞争基线之上始终提高了智能体的性能和分布外泛化能力。
cs.AI / 85 / 2608.21317
From Regulation to Implementation: A Critical Evaluation of LLM-Assisted Regulatory Compliance in Industry
从监管到实施:对行业中大型语言模型辅助合规性的批判性评估
Abstract
The European Union (EU) has emerged as a leading regulatory body in the development of sustainability and privacy regulations. While new regulation requirements vary, many include a documentation artifact to ensure compliance. Notably, the Ecodesign for Sustainable Products Regulation (ESPR) introduces Digital Product Passports (DPPs) for life cycle transparency, while the General Data Protection Regulation (GDPR) mandates Data Protection Impact Assessments (DPIAs) to mitigate privacy risks. Creating these compliance artifacts, however, is challenging. Industrial data, which often exists in heterogeneous formats and is scattered across company and supplier systems, is required for DPPs and can be difficult to extract into compliant DPP formatting. Furthermore, DPIA documents require interdisciplinary expertise and follow no standardized format, making development difficult for novel systems. To address the particular complexity of compliance artifact creation for both regulations, researchers have proposed the use of LLMs in the generation process; however, the impact of the aforementioned problems on the output of these systems is largely unaddressed. This work investigates the existing research gap by exploring how data extraction instructions and regulatory vagueness impact the quality and consistency of LLM-produced compliance artifacts. The resulting artifacts are evaluated by benchmarking different models against manually created ground-truth schemas. The results reveal that less strict guidelines, such as DPIA formatting, require higher context prompts to maintain consistency and completeness. Stricter guidelines, such as formatting for Digital Battery Passports (DBP), result in consistent results regardless of prompt context, but may lead to more hallucinations in the output
Chinese Translation
欧盟(EU)已成为可持续性和隐私法规发展的主要监管机构。虽然新的监管要求各不相同,但许多要求包括文档材料以确保合规性。值得注意的是,针对可持续产品的生态设计法规(Ecodesign for Sustainable Products Regulation, ESPR)引入了数字产品护照(Digital Product Passports, DPPs)以实现生命周期透明,而通用数据保护条例(General Data Protection Regulation, GDPR)则要求进行数据保护影响评估(Data Protection Impact Assessments, DPIAs)以降低隐私风险。然而,创建这些合规性文档是具有挑战性的。DPPs 需要工业数据,而这些数据通常以异构格式存在,并散布在公司和供应商系统中,提取为合规的 DPP 格式可能非常困难。此外,DPIA 文档需要跨学科的专业知识,并且没有标准化格式,使得新系统的开发变得困难。为了应对这两项法规合规性文档创建的特定复杂性,研究人员提出在生成过程中使用大型语言模型(LLMs);然而,前述问题对这些系统输出的影响在很大程度上未得到解决。本研究通过探讨数据提取指令和监管模糊性如何影响 LLM 生成的合规性文档的质量和一致性,调查了现有的研究空白。通过将不同模型与手动创建的真实基准模式进行比较,评估生成的文档。结果表明,较不严格的指南(如 DPIA 格式)需要更高上下文提示以保持一致性和完整性。而较严格的指南(如数字电池护照(Digital Battery Passports, DBP)格式)则无论提示上下文如何都能产生一致的结果,但可能导致输出中出现更多幻觉。
cs.AI / 86 / 2608.21319
Unified Branch-and-Bound Search for the Steiner Traveling Salesman Problem on Graphs of Convex Sets
统一的分支限界搜索用于凸集图上的斯坦纳旅行商问题
Abstract
We formalize the Steiner Traveling Salesman Problem (Steiner-TSP) on Graphs of Convex Sets (GCS), which seeks a minimum-cost closed trajectory through required convex sets while allowing optional transit vertices and revisits. To explore the resulting infinite solution space, we propose a unified branch-and-bound search over rooted walk prefixes. Additive lower-bound-graph costs bound committed prefixes, while a cut-separated connected-flow relaxation lower-bounds the residual cost of visiting every remaining target and returning to the root. Under a uniform positive-cost assumption, best-first traversal terminates after finitely many expansions on every feasible instance without an initial incumbent, whereas depth-first traversal does so once a finite incumbent is available. For a user-specified factor $\epsilon\geq1$, a global lower bound certifies that either strategy's incumbent cost is at most $\epsilon$ times the global optimum. We further demonstrate joint sensing-mode, visitation-order, and continuous-trajectory selection for a mobile-manipulator inspection task, including action precedences expressed in linear temporal logic over finite traces (LTL$_f$). Both traversal strategies find feasible solutions on all benchmark instances within 30s with mean certified optimality gaps of 28.1% and 29.7%, respectively, whereas two recent baselines succeed on only about half of the instances
Chinese Translation
我们对凸集图(GCS)上的斯坦纳旅行商问题(Steiner-TSP)进行了形式化,该问题寻求通过所需的凸集找到最低成本的闭合轨迹,同时允许可选的过境顶点和重访。为了探索由此产生的无限解空间,我们提出了一种基于根节点行走前缀的统一分支限界搜索方法。加性下界图成本对已承诺的前缀进行界定,而切分的连通流松弛则对访问每个剩余目标并返回根节点的剩余成本进行下界。在统一的正成本假设下,最佳优先遍历在每个没有初始解的可行实例上经过有限扩展后终止,而深度优先遍历则在获得有限解后终止。对于用户指定的因子 $ ext{ε} ext{≥} 1$,全局下界证明了任一策略的当前解成本最多为全局最优解的 $ ext{ε}$ 倍。我们进一步展示了移动操控器检查任务的联合感知模式、访问顺序和连续轨迹选择,包括以线性时序逻辑(LTL$_f$)表达的动作优先级。两种遍历策略在所有基准实例上均能在30秒内找到可行解,平均认证最优性差距分别为28.1%和29.7%,而两个最近的基线方法仅在约一半的实例上成功。
cs.AI / 87 / 2608.21332
Anatomy-Informed Neural Networks: Encoding Anatomic Priors in Loss and Architecture, with an SE(3) Formulation of Guidewire-Induced Aortoiliac Deformation
基于解剖学的神经网络:在损失和架构中编码解剖先验,结合导丝引起的主动脉-髂动脉变形的 SE(3) 表述
Abstract
Deep-learning models of anatomy can be numerically plausible yet anatomically impossible, and they generalize poorly when data are scarce. We introduce Anatomy-Informed Neural Networks (AINN), in which soft anatomic priors enter as penalty terms in the loss (e.g., a branching penalty that treats a renal transplant artery off the iliac instead of the aorta as unexpected rather than impossible), in direct analogy to a physics-informed neural network, and hard anatomic priors (e.g., continuity of the vessel) are built into the architecture and state representation, making such invalid predictions impossible by construction wherever the prior admits architectural enforcement. We develop it on a clinical test case with limited data: how the aortoiliac tree deforms when a stiff wire is introduced endoluminally. This is important to contemporary aortic surgery and will matter to autonomous endovascular navigation. We lift the vessel centerline and the wire path from R^3 to curves of frames in the Lie group SE(3), and couple a Cosserat-rod wire to a tortuosity-modulated, anatomically anchored vessel through a unilateral lumen-contact inequality. The prediction is a constrained minimizer of the coupled elastic energy, with contact forces as its Lagrange multipliers. Supervision is a Wasserstein-2 optimal-transport loss between the predicted projection through the C-arm geometry and the observed angiogram, so a 2D angiogram can train a 3D prediction. The kinematics, loss and projection are verified against known ground truth; the mechanics solver only against its own optimality conditions, and predicted displacement is not yet mesh-converged. Here, no network is trained. Future work will transfer this in silico model to real CT scans and test whether it improves predictive accuracy and reduces the training data required.
Chinese Translation
解剖学的深度学习模型可能在数值上是合理的,但在解剖学上却是不可能的,并且在数据稀缺时泛化能力较差。我们提出了基于解剖学的神经网络(Anatomy-Informed Neural Networks, AINN),其中软性解剖先验作为惩罚项进入损失函数(例如,分支惩罚将肾移植动脉视为从髂动脉而非主动脉发出的意外而非不可能),与物理信息神经网络直接类比,而硬性解剖先验(例如,血管的连续性)则内置于架构和状态表示中,使得在先验允许架构强制的地方,这种无效预测在构造上是不可能的。我们在一个数据有限的临床测试案例中开发了该模型:当刚性导丝通过腔内引入时,主动脉-髂动脉树的变形情况。这对当代主动脉手术至关重要,并将影响自主血管内导航。我们将血管中心线和导丝路径从 R^3 提升到李群 SE(3) 中的框架曲线,并通过单边腔接触不等式将 Cosserat-杆导丝与一个受扭曲调制、解剖锚定的血管耦合。预测是耦合弹性能量的约束最小化器,接触力作为其拉格朗日乘子。监督是通过 C-arm 几何体的预测投影与观察到的血管造影之间的 Wasserstein-2 最优传输损失,因此 2D 血管造影可以训练 3D 预测。运动学、损失和投影与已知的真实值进行了验证;力学求解器仅与其自身的最优性条件进行了验证,预测位移尚未收敛到网格。这里没有训练任何网络。未来的工作将把这个计算模型转移到真实的 CT 扫描中,并测试它是否能提高预测准确性并减少所需的训练数据。
cs.AI / 88 / 2608.21357
VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences
VIALS:生命科学中视觉解释文物的基准测试
Abstract
In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy images, plasmid maps, flow cytometry plots, molecular structures, ...) to inform research decisions. We introduce VIALS, a visual question-answering benchmark with 161 such interpretation tasks, spanning the types of artifacts examined throughout experimental workflows in the biotech industry (rather than polished figures from publications and textbooks). While frontier vision-language models can now fluently describe natural images, we find that they are unable to accurately interpret these scientific images, reflecting limitations in domain knowledge and domain-specific visual reasoning capabilities. In contrast, scientists with relevant domain expertise find these visual interpretation tasks straightforward. AI that cannot similarly interpret such images will have limited utility in professional life sciences workflows, where such artifacts are central to how scientists reason, communicate, and make decisions.
Chinese Translation
在专业的生命科学工作流程中,科学家们常常解释视觉文物(如凝胶印迹、显微镜图像、质粒图、流式细胞术图、分子结构等),以指导研究决策。我们介绍了VIALS,这是一个包含161个此类解释任务的视觉问答基准,涵盖了生物技术行业实验工作流程中所涉及的文物类型(而非来自出版物和教科书的精美图形)。尽管前沿的视觉-语言模型现在能够流利地描述自然图像,但我们发现它们无法准确解释这些科学图像,这反映了在领域知识和领域特定视觉推理能力上的局限性。相比之下,具有相关领域专业知识的科学家发现这些视觉解释任务非常简单。无法以类似方式解释这些图像的人工智能在专业生命科学工作流程中的实用性将受到限制,因为这些文物在科学家推理、沟通和决策的过程中至关重要。
cs.CL / 1 / 2608.20344
Beyond Raw Transcripts: Structured Persona Extraction for LLM-Based Digital Twins
超越原始转录:基于大型语言模型的结构化角色提取用于数字双胞胎
Abstract
LLM-based "digital twins" aim to simulate how an individual would behavein new environments or respond to novel questions, given some representation of that individual's prior responses. A common approach constructs this representation from survey transcripts or summaries responses. Prior work shows that compressing long transcripts into shorter LLM-generated summaries does not significantly reduce predictive accuracy, suggesting that information volume is not the primary bottleneck. In this work, we argue that the key limitation is instead structural:how persona information is organized before being provided to thesimulator model. We study this by comparing unstructured summaries with structured persona representations. First, we introduce a hand-craftedschema (BDE: Background, Decision procedure, Evaluation), grounded in consumer-behavior theory, and show that it improves predictive accuracy over raw transcripts by +1.91 percentage points on a homogeneous benchmark (Twin-2K-500), with similar gains on gpt-5.4-mini and Qwen3-8B as robustness checks. However, this fixed structure does not generalizeacross more heterogeneous tasks, where performance is statistically indistinguishable from the raw transcript baseline. To address this limitation, we propose an automatic structure-discovery pipeline in which an LLM iteratively proposes and refines task-specific persona structures and extraction prompts. On a benchmark of 13 diverse sub-studies, this approach restores performance, improving mean accuracy by +1.91 percentage points over the raw transcript baseline and eliminating significant losses observed with the fixed schema. Overall, our results suggest that the main constraint in LLM-based digital twins is not how much information is provided, but how it is structured -- and that the optimal structure depends on the task.
Chinese Translation
基于大型语言模型(LLM)的“数字双胞胎”旨在模拟个体在新环境中的行为或对新问题的反应,前提是提供该个体先前反应的某种表示。常见的方法是从调查转录或摘要响应中构建这种表示。先前的研究表明,将长转录压缩为较短的LLM生成摘要并不会显著降低预测准确性,这表明信息量并不是主要瓶颈。在本研究中,我们认为关键限制在于结构:角色信息在提供给模拟器模型之前的组织方式。我们通过比较非结构化摘要与结构化角色表示来研究这一点。首先,我们引入了一种手工制作的模式(BDE:背景、决策过程、评估),该模式基于消费者行为理论,并显示其在同质基准(Twin-2K-500)上比原始转录提高了+1.91个百分点的预测准确性,在gpt-5.4-mini和Qwen3-8B上也有类似的增益作为稳健性检查。然而,这种固定结构在更异质的任务中并不具有普遍性,其性能在统计上与原始转录基线无显著差异。为了解决这一限制,我们提出了一种自动结构发现管道,其中LLM迭代地提出和完善任务特定的角色结构和提取提示。在13个多样化子研究的基准测试中,这种方法恢复了性能,使平均准确性比原始转录基线提高了+1.91个百分点,并消除了在固定模式下观察到的显著损失。总体而言,我们的结果表明,LLM基础的数字双胞胎的主要限制不是提供多少信息,而是如何组织这些信息——而且最佳结构取决于任务。
cs.CL / 2 / 2608.20345
When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha
当词汇理解阻碍临床推理:评估面向Z世代的治疗机器人安全风险
Abstract
Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1% of U.S. adolescents (5.4 million) using generative AI for mental health advice. While these systems, from therapy apps to general chatbots, rely on large language models trained on extensive psychological literature, their safety for youth communication patterns characterized by hyperbolic language, ironic positivity, rapid semantic drift, and contextual polysemy remains unvalidated. Following multiple adolescent deaths linked to AI chatbot interactions, systematic evaluation is critical. We present two benchmarks: (1) 64 Gen Alpha mental health expressions validated by native speakers (ICC=0.72) and clinicians (kappa=0.78); (2) 75 multi-turn conversations (780 turns) with paired Standard/Gen Alpha versions. Across evaluations of LLM architectures underlying therapy apps and general chatbots - Claude, GPT-4o, Llama-3.1 - models understand 76-82% of vocabulary but correctly calibrate only 64-72% of clinical risk, creating a 10-14 percentage point (pp) vocabulary-comprehension gap (p<.001, d>0.48) absent in human therapists (3pp, p=.22). The gap is architecturally consistent and widens with ambiguity (7pp -> 18pp). We identify six failure patterns: sarcasm masking (29pp), minimization acceptance (43pp), informal style bias (24pp), risk-stratified ambiguity (19pp), semantic drift (19pp), context-dependent violence (7pp). Patterns compound; three or more yield 94% miss rates. Lightweight mitigations fail; only heavy scaffolding achieves human performance (6.4x cost). With 34% baseline miss rate yielding 146,880 estimated annual missed crises, we recommend mandatory human-in-the-loop architectures, quarterly youth-specific validation, transparent performance disclosure, and regulatory frameworks for youth-facing mental health AI.
Chinese Translation
对话式人工智能系统已成为Z世代(Generation Alpha,出生于2010-2024年)非正式的心理健康支持资源,13.1%的美国青少年(约540万)使用生成式人工智能获取心理健康建议。尽管这些系统,从治疗应用到一般聊天机器人,依赖于基于广泛心理学文献训练的大型语言模型,但它们在面对以夸张语言、讽刺性积极、快速语义漂移和上下文多义性为特征的青少年沟通模式时的安全性尚未得到验证。在与人工智能聊天机器人互动相关的多起青少年死亡事件之后,系统评估显得至关重要。我们提出了两个基准:(1)由母语者(ICC=0.72)和临床医生(kappa=0.78)验证的64个Z世代心理健康表达;(2)75个多轮对话(780轮)与标准版/Z世代版本配对。在对治疗应用和一般聊天机器人背后的大型语言模型架构(Claude、GPT-4o、Llama-3.1)的评估中,这些模型理解76-82%的词汇,但仅正确校准64-72%的临床风险,造成10-14个百分点(pp)的词汇理解差距(p<.001,d>0.48),而人类治疗师的差距为3pp(p=.22)。这种差距在架构上是一致的,并且在模糊性增加时扩大(7pp -> 18pp)。我们识别出六种失败模式:讽刺掩盖(29pp)、最小化接受(43pp)、非正式风格偏见(24pp)、风险分层模糊性(19pp)、语义漂移(19pp)、上下文依赖的暴力(7pp)。这些模式相互叠加;三种或更多模式导致94%的漏报率。轻量级的缓解措施无效;只有重型支撑才能达到人类表现(成本为6.4倍)。在34%的基线漏报率下,预计每年将错过146,880个危机事件,我们建议实施强制性的人机协作架构、每季度进行青少年特定的验证、透明的性能披露以及面向青少年的心理健康人工智能的监管框架。
cs.CL / 3 / 2608.20346
Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care
构建与评估用于电信客户服务的合成孟加拉语语音资源
Abstract
Speech systems used in customer-facing applications often require domain-specific language coverage. We present a synthetic Bengali speech dataset for telecom customer-care scenarios. The dataset contains 10,000 audio-text pairs, approximately 26.82 hours of 24 kHz speech, and predefined train, validation, and test splits of 9,000, 500, and 500 examples. It is publicly released on Hugging Face under the CC-BY-4.0 license. The speech was generated with OmniVoice in voice-cloning mode using a real female reference recording and transcript, with bfloat16 precision, 16 diffusion sampling steps, and a speaking-rate control value of 1.0. Along with the original Bengali text, the dataset provides a normalized transcript field designed for ASR/STT training and evaluation. We report an automatic intelligibility check over all 10,000 samples using a domain-adapted Whisper ASR model fine-tuned from bengaliAI/tugstugi_bengaliai-regional-asr_whisper-medium, along with a manual listening check on selected samples. The evaluation gives an average WER of 2.54%, an average CER of 0.59%, and median WER and CER values of 0.00%. These results suggest strong text-audio consistency under the selected automatic evaluation pipeline, while the paper also discusses the limitations of synthetic speech and STT-based evaluation.
Chinese Translation
客户面对面的应用中使用的语音系统通常需要特定领域的语言覆盖。我们提出了一个用于电信客户服务场景的合成孟加拉语语音数据集。该数据集包含10,000个音频-文本对,约26.82小时的24 kHz语音,并预定义了训练、验证和测试集的划分,分别为9,000、500和500个样本。该数据集在Hugging Face上以CC-BY-4.0许可证公开发布。语音是使用OmniVoice在语音克隆模式下生成的,参考录音为真实女性录音及其文本,采用bfloat16精度、16个扩散采样步骤,以及1.0的语速控制值。除了原始的孟加拉语文本外,该数据集还提供了一个为自动语音识别(ASR)/语音转文本(STT)训练和评估设计的标准化转录字段。我们使用经过领域适应的Whisper ASR模型(从bengaliAI/tugstugi_bengaliai-regional-asr_whisper-medium微调)对所有10,000个样本进行了自动可懂性检查,并对选定样本进行了人工听力检查。评估结果显示平均字错误率(WER)为2.54%,平均字符错误率(CER)为0.59%,中位数WER和CER值均为0.00%。这些结果表明在所选自动评估流程下文本与音频之间具有较强的一致性,同时本文还讨论了合成语音和基于STT的评估的局限性。
cs.CL / 4 / 2608.20347
Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias
语言模型认为谁是有能力的?职业偏见的机制分析
Abstract
Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have merely learned not to express them. In this study, we show that representational biases are often detectable, even when behavioral biases are not visible. We introduce a causal framework that decomposes occupational bias into two measurement points: a model's internal representation of a user's competence, and its observable outputs. We derive steering vectors for representations of user expertise, and verify that they causally mediate model behavior in both a question-answering task and a hiring task. Applying this framework to several open-weight models, we find that demographic attributes, such as gender, race, and socioeconomic status, influence a model's representation of user expertise, even in cases where behavioral metrics detect no disparity between demographics. We show that these model representations can influence downstream behavior under intervention, suggesting failure modes that behavioral metrics alone may not detect.
Chinese Translation
语言模型(LMs)通常通过行为偏见评估,但尚不清楚它们是否不再代表导致偏见的潜在关联,或者只是学会了不表达这些偏见。在本研究中,我们展示了即使在行为偏见不可见的情况下,表现偏见通常也是可检测的。我们引入了一个因果框架,将职业偏见分解为两个测量点:模型对用户能力的内部表征,以及其可观察的输出。我们推导出用户专业知识表征的引导向量,并验证它们在问答任务和招聘任务中因果地介导模型行为。将这一框架应用于多个开放权重模型,我们发现性别、种族和社会经济地位等人口统计属性影响模型对用户专业知识的表征,即使在行为指标未检测到人口统计之间差异的情况下。我们表明,这些模型表征在干预下可以影响下游行为,暗示行为指标单独可能无法检测到的失败模式。
cs.CL / 5 / 2608.20348
Inhibitory Attention for Clinical Long-Context Reasoning: Characterizing and Mitigating Lost-in-the-Middle Effects in EHR Processing
临床长上下文推理中的抑制注意力:特征化和缓解电子健康记录处理中的中间丢失效应
Abstract
Electronic health records now routinely exceed 100,000 tokens per patient. Yet large language models exhibit the lost-in-the-middle (LitM) effect: information near the center of a long context is retrieved less reliably than information near the edges. In clinical use this is not benign: the single most consequential fact in a note can sit at its center. We term this the clinical lost-in-the-middle (CLitM) problem, give its first systematic characterization using MedAlign, and compare context-selection strategies as remedies. Across 2,196 instruction-response pairs and six language models, we observe a 21.9 percentage-point gap between peak accuracy (59.5%, 95% CI [46.3, 71.0], 20-30% decile) and trough accuracy (37.6% [23.2, 52.5] at 70-80%); 67.8% of reference answers fall between the 10th and 90th percentiles of the EHR timeline, inside the CLitM trough. We introduce Query-Conditioned Clinical Suppression (QCCS), a lightweight query-conditioned selection gate, and evaluate it against BM25, BM25 with section-header filtering, dense retrieval, and cross-encoder reranking (N=83 held-out instructions). With Qwen2.5-7B-Instruct (16k context), QCCS outperforms all five comparators under LLM-as-judge scoring: for middle-position instructions QCCS reaches 16.7% versus BM25 3.3%, cross-encoder 0.0%, dense 0.0%, and full context 6.7%; overall QCCS reaches 25.3% versus at most 3.6% for retrieval-only comparators. This advantage is not explained by retrieval recall: at k=20, BM25 retrieves the gold evidence sentence in 98.8% of instructions (QCCS 34.9%), yet retrieval arms stay at most 2.6% accurate even when they retrieve it, whereas QCCS reaches 25.0% even when it does not. In this proof-of-concept evaluation, query-aligned context selection predicts EHR instruction-following accuracy better than gold-sentence retrieval recall.
Chinese Translation
电子健康记录现在每位患者的记录常常超过100,000个标记。然而,大型语言模型表现出中间丢失(LitM)效应:在长上下文中,位于中心的信息的检索可靠性低于位于边缘的信息。在临床应用中,这并非无害:笔记中最重要的事实可能位于其中心。我们将其称为临床中间丢失(CLitM)问题,利用MedAlign对其进行首次系统性特征化,并比较上下文选择策略作为补救措施。在2,196对指令-响应对和六种语言模型中,我们观察到峰值准确率(59.5%,95% CI [46.3, 71.0],20-30%分位数)与谷值准确率(37.6% [23.2, 52.5],70-80%分位数)之间存在21.9个百分点的差距;67.8%的参考答案落在电子健康记录时间线的第10和第90百分位之间,位于CLitM谷值内。我们引入了查询条件临床抑制(QCCS),一种轻量级的查询条件选择门,并将其与BM25、带有章节标题过滤的BM25、密集检索和交叉编码重排序进行评估(N=83个保留指令)。在使用Qwen2.5-7B-Instruct(16k上下文)时,QCCS在LLM作为评判者的评分下优于所有五个比较方法:对于中间位置的指令,QCCS达到16.7%,而BM25为3.3%,交叉编码为0.0%,密集检索为0.0%,完整上下文为6.7%;总体而言,QCCS达到25.3%,而检索仅比较方法最多为3.6%。这一优势并不能通过检索召回率来解释:在k=20时,BM25在98.8%的指令中检索到黄金证据句(QCCS为34.9%),然而即使检索到它,检索武器的准确率最多也只有2.6%,而QCCS即使在未检索到时也能达到25.0%。在这一概念验证评估中,查询对齐的上下文选择比黄金句子的检索召回率更好地预测了电子健康记录指令的遵循准确性。
cs.CL / 6 / 2608.20349
Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality
超越提示工程:提示词汇敏感性及其对质量影响的系统分析
Abstract
Large Language Models (LLMs) exhibit extreme sensitivity to surface-level prompt variations, in which minor lexical changes can trigger disproportionate performance fluctuations. Moving beyond black-box optimization and coarse-grained templates, we present the first large-scale, n-gram token-level mechanistic analysis of prompt stability, leveraging a dataset of 132,000 prompt variants. Our investigation reveals a fundamental Scaling Law of Prompt Performance Stability: higher average task performance is strongly associated with lower variance and greater robustness across prompt perturbation. We identify two core linguistic drivers underlying this robustness: (1) Domain-Specific Terminology, which tightly anchors semantic boundaries, and (2) Explicit Action Directives, which formalize reasoning trajectories. Together, these elements constrain the model's interpretative space, effectively ``locking in'' more deterministic generation behavior. Building on these insights, we introduce an automated Prompt-Refining Agent that systematically restructures input queries by injecting domain anchoring and operational constraints. Empirical evaluation shows that our approach reduces performance variance by 40.7% in code generation task, while preserving or improving mean performance. These findings provide a statistically grounded and mechanistically interpretable framework for achieving robust prompt engineering.
Chinese Translation
大型语言模型(LLMs)对表层提示变化表现出极高的敏感性,细微的词汇变化可能引发不成比例的性能波动。我们超越黑箱优化和粗粒度模板,首次进行大规模的n-gram令牌级机制分析,以研究提示的稳定性,利用了132,000个提示变体的数据集。我们的研究揭示了提示性能稳定性的基本缩放法则:更高的平均任务性能与更低的方差和更大的鲁棒性之间存在强关联,尤其是在提示扰动的情况下。我们识别出支撑这种鲁棒性的两个核心语言驱动因素:(1)特定领域术语,它紧密锚定语义边界;(2)明确的行动指令,它们形式化推理轨迹。这些元素共同限制了模型的解释空间,有效地“锁定”了更确定性的生成行为。基于这些见解,我们引入了一种自动化的提示优化代理,它通过注入领域锚定和操作约束,系统性地重构输入查询。实证评估表明,我们的方法在代码生成任务中将性能方差降低了40.7%,同时保持或提高了平均性能。这些发现为实现稳健的提示工程提供了一个统计基础和机制可解释的框架。
cs.CL / 7 / 2608.20350
How to Train a Real-World Silicon Concierge? Internalizing Complex Business Workflow to Only OneModel
如何训练一个现实世界的硅基礼宾?将复杂的业务工作流程内化为单一模型
Liu, Chang, Ning, Chaoyang, Jiang, Dayi, Gu, Enrui, Ran, Fang, Xue, Hongyan, Li, Huaqing, Cai, Hui, Liu, Jia, Yang, Jiang-Ming, Li, Jianshe, Luo, Jiawei, Zhou, Jin, Zhu, Leshen, Chen, Lihui, Ma, Liying, Xue, Lyuxin, Ji, Mengjian, Xu, Ruijia, Ren, Wei, Wu, Wei, Qu, Xiaoling, Feng, Xiaoyun, Zhang, Xin, Zhou, Xixie, Hu, Xuanwei, Chen, Yan, Wang, Yichao, Tong, Yongqi, Liu, Yu, Zhou, Yuhong, Sun, Zemin, Xu, Zhenwen, Liu, Zhiling, Wang, Zifan
Abstract
Traditional industrial agents rely on modular pipelines, including Router, Retriever, Planner, Executor, Responder, Reviewer, and other components. These systems often fracture into a labyrinth of ad-hoc patches, leading to cascading errors and high latency. We propose OneModel, an applicable paradigm shift from external workflows to internalized knowledge representation. Unlike modular systems that slice fluid user intents into static steps, OneModel consolidates complex business logic and SOPs directly into the model parameters. Through Continual Pre-training (CPT) and logic-compilation SFT, we transform fragmented business rules into intuitive model reasoning within a unified attention space. Deployed in our global financial service system, OneModel effectively breaks the trade-off between latency, accuracy, and complexity. Online A/B testing demonstrates an end-to-end latency reduction of more than 50 percent, from 18.7 seconds to 8.0 seconds, while the Intelligent Resolution Rate (IRR) increases from 64.3 percent to 83.3 percent. The results show that OneModel can replace brittle engineering logic with internalized cognitive intuition, offering a scalable blueprint for transitioning industrial agents from complex, error-prone workflows to unified model architectures.
Chinese Translation
传统工业代理依赖于模块化管道,包括路由器(Router)、检索器(Retriever)、规划器(Planner)、执行器(Executor)、响应器(Responder)、审查器(Reviewer)及其他组件。这些系统常常分裂成一个临时修补的迷宫,导致级联错误和高延迟。我们提出了OneModel,这是一种将外部工作流程转变为内化知识表示的适用范式。与将流动的用户意图切割为静态步骤的模块化系统不同,OneModel将复杂的业务逻辑和标准操作程序(SOP)直接整合到模型参数中。通过持续预训练(Continual Pre-training, CPT)和逻辑编译的微调(logic-compilation SFT),我们将碎片化的业务规则转化为统一注意力空间内的直观模型推理。在我们的全球金融服务系统中部署的OneModel有效打破了延迟、准确性和复杂性之间的权衡。在线A/B测试表明,端到端延迟减少超过50%,从18.7秒降至8.0秒,同时智能解决率(Intelligent Resolution Rate, IRR)从64.3%提高至83.3%。结果表明,OneModel能够用内化的认知直觉替代脆弱的工程逻辑,为将工业代理从复杂且易出错的工作流程过渡到统一模型架构提供了可扩展的蓝图。
cs.CL / 8 / 2608.20351
Exploratory As-Analyzed No-Detection of Culturally-Marked Predicate-Triggered PII Amplification in a Synthetic-English RAG Probe: A Predicate-Resource-Confounded Audit
探索性分析未检测到的文化标记谓词触发的个人身份信息放大现象:一个谓词-资源混淆审计
Abstract
We ask whether stereotype-loaded queries about culturally marked people leak more personal information from a retrieval-augmented generation (RAG) system than otherwise-equivalent neutral queries. We pre-register a four-culture audit (en-Anglo, es-LATAM, Arabic, Hindi) on a synthetic English PII corpus, comparing five query arms we call the Stereotype-Trigger Leakage Delta (STLD). Two caveats up front. Our locked confirmatory estimator was never run, so every test in the paper is exploratory or sensitivity, with all plan deviations listed in the appendix. And the name-leakage metric is contaminated by a prompt-echo artifact: the model often just re-emits the name we asked about, which inflates apparent leakage without any retrieval at all. On the cleaner channels (email, phone, ssn-like, address), we find no stereotype-driven amplification on any of the four cultures after multiple-comparison correction. Because our sample is only powered for mid-sized effects, and because the culturally marked probes mix stereotype content with cultural markers and heritage practices, we present this as no detection, not evidence of no effect, of culturally marked predicate leakage that is confounded with the underlying resource.
Chinese Translation
我们探讨关于文化标记人群的刻板印象加载查询是否比其他等效中性查询从检索增强生成(RAG)系统中泄露更多个人信息。我们在一个合成英语个人身份信息(PII)语料库上预注册了一个四文化审计(en-Anglo, es-LATAM, 阿拉伯语, 印地语),比较了我们称之为刻板印象触发泄漏差异(Stereotype-Trigger Leakage Delta, STLD)的五个查询组。首先有两个警告。我们的锁定确认估计器从未运行,因此本文中的每个测试都是探索性或敏感性测试,所有计划偏差均在附录中列出。此外,名称泄漏指标受到提示回声伪影的污染:模型经常只是重新输出我们询问的名称,这在没有任何检索的情况下膨胀了表面泄漏。在更干净的渠道(电子邮件、电话、类似社会安全号码、地址)中,我们发现经过多重比较校正后,四个文化中都没有驱动刻板印象的放大现象。由于我们的样本仅对中等效应具有统计效能,并且由于文化标记探针将刻板印象内容与文化标记和遗产实践混合在一起,我们将此视为未检测到,而不是文化标记谓词泄漏的无效证据,这种泄漏与基础资源混淆。
cs.CL / 9 / 2608.20353
The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP
偏差假说:揭示心理健康自然语言处理中的词汇干扰和标签偏差
Abstract
Computational mental health (CMH) classifiers often degrade under distribution shift because human annotators and distant-supervision pipelines reward different linguistic signals. We introduce TSS (Triple-Stream Stress probe), a multi-channel diagnostic framework that decomposes text into (A) lexical character n-grams, (B) a small, mostly content-free morpho-syntactic channel, and (C) a 154-feature psycholinguistic style channel. Across four English datasets (N=12,906), TSS reveals a lexical interference effect: adding lexical features to the style channel reduces Macro-F1 on human-labeled data (mean drop 0.072, p<10^-4) but not on auto-labeled data. We propose Degree of Divergence (DoD), a difference-in-differences statistic adapted from econometrics for label-source auditing, with instance-level bootstrap inference; the headline estimate is DoD(BC-A) = 0.0374, 95% CI [0.0097, 0.0651], p=0.0032. A platform-stratified Twitter-only DoD (which removes the Reddit vs. Twitter contrast) reproduces the pattern with bootstrap inference: DoD-Tw(BC-A) = +0.096 (p<0.001) and DoD-Tw(AC-A) = -0.089 (p<0.001). Interventional masking (pos_only) retains ~95-99% of Channel C's performance after destroying content words on human datasets, indicating that the style channel does not rely primarily on lexical surface form. TSS is positioned as a diagnostic audit framework, not a clinical screening tool: it flags label-source-specific shortcut learning before generalization claims are made.
Chinese Translation
计算心理健康(CMH)分类器在分布转移下往往表现不佳,因为人类标注者和远程监督管道奖励不同的语言信号。我们引入了 TSS(Triple-Stream Stress probe),一个多通道诊断框架,将文本分解为 (A) 词汇字符 n-gram,(B) 一个小的、主要无内容的形态-句法通道,以及 (C) 一个包含154个特征的心理语言学风格通道。在四个英语数据集上(N=12,906),TSS 显示出词汇干扰效应:将词汇特征添加到风格通道会降低人类标注数据的宏观 F1 值(平均下降 0.072,p<10^-4),但对自动标注数据则没有影响。我们提出了偏差程度(Degree of Divergence, DoD),这是一个从计量经济学中改编的差异中的差异统计量,用于标签源审计,采用实例级自助推断;主要估计为 DoD(BC-A) = 0.0374,95% 置信区间 [0.0097, 0.0651],p=0.0032。一个平台分层的仅 Twitter 的 DoD(去除了 Reddit 与 Twitter 的对比)通过自助推断重现了这一模式:DoD-Tw(BC-A) = +0.096(p<0.001)和 DoD-Tw(AC-A) = -0.089(p<0.001)。干预性掩蔽(pos_only)在破坏人类数据集上的内容词后,仍保留了约 95-99% 的通道 C 性能,表明风格通道并不主要依赖于词汇表面形式。TSS 被定位为一种诊断审计框架,而非临床筛查工具:它在提出泛化主张之前标记特定于标签源的捷径学习。
cs.CL / 10 / 2608.20355
ExpertIVS: Sociological Expert Driven Individual Value Simulation in Large Language Models
ExpertIVS:基于社会学专家驱动的大型语言模型个体价值模拟
Abstract
Large Language Model (LLM) agents have demonstrated considerable potential for social simulation, yet struggle to accurately model individual value systems. Most existing methods mechanically stitch survey responses into prompts, which suffer from semantic fragmentation, failing to capture the internal coherence of human value systems. The value systems of LLMs are typically assessed using static multiple-choice questions, which fail to evaluate the value orientation in real-world dialogue interactions. To address these issues, we propose ExpertIVS, a framework employing 14 Sociological Expert Agents to interpret World Values Survey (WVS) responses through structured professional perspectives, rather than direct responses concatenation. These expert agents perform deep semantic reconstruction to generate robust and internally consistent individual profiles. To evaluate the consistency between LLMs and individual value systems during dynamic interactions, we further introduce a multi-agent debate mechanism. Extensive experiments across 480 individuals from 12 countries demonstrate that ExpertIVS achieves 90.78% value restoration fidelity and significantly outperforms baselines in value generalization (+5.3%). Moreover, ExpertIVS exhibits strong personality discriminability and behavioral consistency, enabling a shift from mere response concatenation to genuine sociological role-playing.
Chinese Translation
大型语言模型(LLM)代理在社会模拟方面展示了相当大的潜力,但在准确建模个体价值体系方面仍然面临挑战。现有大多数方法机械地将调查响应拼接成提示,导致语义碎片化,未能捕捉人类价值体系的内在一致性。LLM的价值体系通常通过静态的多项选择题进行评估,这无法在真实对话互动中评估价值取向。为了解决这些问题,我们提出了ExpertIVS,一个框架,利用14个社会学专家代理通过结构化的专业视角来解读世界价值调查(WVS)响应,而不是直接拼接响应。这些专家代理进行深度语义重构,以生成稳健且内部一致的个体档案。为了评估LLM与个体价值体系在动态互动中的一致性,我们进一步引入了一种多代理辩论机制。对来自12个国家的480名个体进行的广泛实验表明,ExpertIVS实现了90.78%的价值恢复保真度,并在价值泛化方面显著优于基线(+5.3%)。此外,ExpertIVS表现出强大的个性辨别能力和行为一致性,使得从单纯的响应拼接转向真正的社会学角色扮演成为可能。
cs.CL / 11 / 2608.20359
Self-Speculation for Faster Reasoning Models
自我推测以加速推理模型
Abstract
Large language models (LLMs) are deployed for increasingly complex tasks involving planning and multi-step decision making, but high-quality performance on these tasks often requires generating long reasoning traces. This is a poor fit for latency-sensitive and interactive applications like voice assistants or coding agents, where generation latency can strongly affect user experience. Existing acceleration methods typically focus on token-level generation, without utilizing the structure of reasoning workflows. We introduce SSR: Self-Speculation for Reasoning Models, a training-free self-speculative decoding method that leverages the chain-of-thought (CoT) as a source of speculation. SSR uses the partial-CoT answer distribution as the drafter and the full-CoT distribution as the verifier, deriving both from the same model at different reasoning budgets. This builds on the observation that later partial-CoT responses often exhibit greater semantic and lexical overlap with the full-budget response. Due to this overlap, SSR can accept long draft prefixes at once, leading to large speedups on structured and long-form generation tasks. To further exploit draft-response overlap beyond the contiguous prefix accepted by standard speculative decoding, SSR also incorporates suffix decoding, using the draft to seed a suffix cache and recover useful spans beyond the accepted prefix, further reducing latency on tasks with high lexical overlap between the draft and the final response. We evaluate SSR on multiple structured and long-form generation tasks where it is most useful, and demonstrate a relative improvement of up to 24.1% on total generation latency for popular open-source models such as Qwen3.5 and Gemma-4.
Chinese Translation
大型语言模型(LLMs)被应用于越来越复杂的任务,这些任务涉及规划和多步骤决策,但在这些任务上实现高质量的表现通常需要生成长的推理轨迹。这与延迟敏感和交互式应用(如语音助手或编码代理)的需求不符,因为生成延迟会显著影响用户体验。现有的加速方法通常专注于令牌级生成,而未利用推理工作流的结构。我们提出了SSR:推理模型的自我推测(Self-Speculation for Reasoning Models),这是一种无训练的自我推测解码方法,利用思维链(chain-of-thought, CoT)作为推测的来源。SSR使用部分CoT答案分布作为草稿,使用完整CoT分布作为验证者,二者均来自同一模型在不同推理预算下的输出。这基于这样的观察:后续的部分CoT响应通常与完整预算响应在语义和词汇上具有更大的重叠。由于这种重叠,SSR可以一次接受长草稿前缀,从而在结构化和长文本生成任务中实现显著加速。为了进一步利用草稿响应的重叠,超越标准推测解码所接受的连续前缀,SSR还结合了后缀解码,利用草稿来初始化后缀缓存,并恢复超出接受前缀的有用区间,进一步减少在草稿与最终响应之间存在高词汇重叠的任务中的延迟。我们在多个结构化和长文本生成任务上评估了SSR,证明其在流行的开源模型(如Qwen3.5和Gemma-4)上在总生成延迟方面相对提高了高达24.1%。
cs.CL / 12 / 2608.20360
TriPLU: Bypassing the Gate with Direct Trilinear Product FFNs in Tiny Language Models
TriPLU:通过直接三线性乘积前馈网络绕过门控机制在微型语言模型中的应用
Abstract
We study whether tiny decoder-only language models benefit from feed-forward layers that directly multiply learned feature projections. TriPLU, a Trilinear Product Linear Unit, replaces the usual gated FFN branch with a product-only degree-3 branch that multiplies three projected streams coordinatewise. In a character-level TinyStories 1M-byte prefix study, TriPLU reaches a mean best validation loss of 1.0637, compared with 1.1017 for closely matched SwiGLU, 1.0780 for a degree-4 product control, and 1.1026 for a degree-2 control. In train-only Byte-BPE experiments, TriPLU also lowers validation and heldout bits per byte on TinyStories and WikiText-2 raw under low-learning-rate settings, with PMI-slice evidence suggesting gains on seen middle- and high-PMI adjacent-token pairs. Constant-learning-rate diagnostics show that product-branch normalization can reduce the high-learning-rate best-checkpoint gap, although final BPB still degrades under hot schedules. The resulting claim is deliberately narrow: direct product FFNs can improve fixed-budget small-model loss in specific low-compute regimes, but the branch is optimization-sensitive and does not establish FLOP-normalized efficiency, scaling behavior, or broad LLM performance.
Chinese Translation
我们研究了微型解码器语言模型是否能从直接乘积学习特征投影的前馈层中受益。TriPLU(Trilinear Product Linear Unit)用一个仅包含乘积的三阶分支替代了通常的门控前馈网络分支,该分支对三个投影流进行坐标级别的乘法。在一个字符级别的TinyStories 1M字节前缀研究中,TriPLU达到了平均最佳验证损失1.0637,而相近的SwiGLU为1.1017,四阶乘积控制为1.0780,二阶控制为1.1026。在仅训练的Byte-BPE实验中,TriPLU在低学习率设置下也降低了TinyStories和WikiText-2原始数据的验证和保留字节每字节数,PMI切片证据表明在已见的中高PMI相邻标记对上有收益。恒定学习率的诊断表明,乘积分支归一化可以减少高学习率最佳检查点之间的差距,尽管最终的BPB在高温调度下仍然会退化。最终的结论是,直接乘积前馈网络可以在特定的低计算环境中改善固定预算小模型的损失,但该分支对优化敏感,并未确立FLOP归一化效率、扩展行为或广泛的LLM性能。
cs.CL / 13 / 2608.20361
Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure
迈向自动研究:从具有类别结构的论文知识图谱中挖掘可证伪的研究想法
Abstract
Automated research-idea generation systems built on large language models (LLMs) share a structural weakness: they reduce ideation to free-text recombination, random paper pairing, or embedding-similarity retrieval. The three approaches fail in the same way: each treats a paper as a flat object, a string or a vector, and so quotients away the typed problem-method-metric-claim arrows a researcher actually uses when reasoning about a cross-domain analogy. We recover the missing structure with the minimal piece of category theory that a typed graph alone does not provide: composition, together with identity arrows, which makes it possible to ask whether a proposed analogy preserves relation chains. Concretely, each paper $p$ is modelled as a small category $C_p$ whose objects are extracted typed research entities and whose morphisms are the relations the paper asserts; a cross-paper bridge from $p$ to $q$ is then a partial functor candidate $F: C_p -> C_q$ that preserves object kinds and covered relation classes. We instantiate the model as a three-layer algorithm: categorical signature clustering, a functor-preservation gate, and a six-axis LLM plausibility judge. Evaluated on a corpus of tens of thousands of full-text-parsed papers under four ablation conditions, the categorical gate filters cross-domain candidates at roughly a 17:1 ratio while the quantitative-falsifier rate of accepted ideas stays above 83% throughout; every rejected candidate is retained with its per-axis rationale, so the gate doubles as a logging layer rather than a silent filter.
Chinese Translation
基于大型语言模型(LLMs)构建的自动化研究想法生成系统存在结构性弱点:它们将创意生成简化为自由文本重组、随机论文配对或嵌入相似性检索。这三种方法以相同的方式失败:每种方法将论文视为一个平面对象,一个字符串或一个向量,从而忽略了研究人员在推理跨领域类比时实际使用的类型化问题-方法-度量-主张箭头。我们通过类别理论中最小的结构恢复缺失的结构,这一结构单独的类型图所无法提供:组合以及恒等箭头,使得可以询问所提出的类比是否保持关系链。具体而言,每篇论文 $p$ 被建模为一个小类别 $C_p$,其对象是提取的类型化研究实体,其态射是论文所断言的关系;从 $p$ 到 $q$ 的跨论文桥接则是一个部分函子候选 $F: C_p -> C_q$,它保持对象种类和覆盖的关系类别。我们将该模型实例化为一个三层算法:类别签名聚类、函子保持门和六轴 LLM 可信度评估器。在四种消融条件下对数万篇全文解析论文的语料库进行评估时,类别门以大约 17:1 的比例过滤跨领域候选,而被接受想法的定量证伪率始终保持在 83% 以上;每个被拒绝的候选都保留其每个轴的理由,因此该门不仅是一个静默过滤器,还是一个日志记录层。
cs.CL / 14 / 2608.20362
Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck
RLVR中的多语言验证器偏差:基准、回滚诊断与跨语言选择瓶颈
Abstract
Reinforcement learning with verifiable rewards (RLVR) is a standard recipe for training large language models on mathematical reasoning, where an answer verifier serves as a language-neutral reward function. We show that this assumption fails in multilingual settings: an exact-match verifier turns format and script variation into language-dependent false-negative reward noise. We introduce a reusable protocol for auditing multilingual RLVR rewards: a verifier-robustness suite, a rollout-diagnosis procedure, and language-conditioned reward-error metrics for Japanese, English, and Chinese answers. On MGSM rollouts with k=8, the exact-match proxy rejects trusted-correct answers at sharply different rates by language across Qwen3-4B, Qwen3-8B, and Llama-3.1-8B-Instruct; for Qwen3-8B, the false-negative rate reaches 0.642 on JP against 0.122 on EN and 0.073 on CN. A plain-numeric probe localizes the mechanism to the final-answer interface: an interface model drives reward-error VLB to zero while the residual accuracy gap is unchanged. We then expose a cross-lingual selection bottleneck: on MGSM250 rollouts, a target-local aggregation rule using no trusted labels closes 55-78% of the average selection gap, and over 95% of repairs require genuine cross-lingual support. The bottleneck replicates on a 483-problem MATH-500 set. A controlled training audit shows that rule-GRPO raises trusted accuracy while the reward-error VLB stays high. The unifying message is operational: multilingual RLVR rewards should be audited by language and by answer interface before they are optimized.
Chinese Translation
具有可验证奖励的强化学习(RLVR)是训练大型语言模型进行数学推理的标准方法,其中答案验证器充当语言中立的奖励函数。我们展示了这一假设在多语言环境中失效:精确匹配验证器将格式和脚本变体转化为语言依赖的假阴性奖励噪声。我们引入了一种可重复使用的多语言RLVR奖励审计协议:验证器鲁棒性套件、回滚诊断程序以及针对日语、英语和中文答案的语言条件奖励误差指标。在k=8的MGSM回滚中,精确匹配代理根据语言在Qwen3-4B、Qwen3-8B和Llama-3.1-8B-Instruct之间以显著不同的比率拒绝可信正确答案;对于Qwen3-8B,日语的假阴性率达到0.642,而英语为0.122,中文为0.073。一个简单的数值探测器将机制定位于最终答案接口:接口模型将奖励误差的变异性降低至零,而残余准确性差距保持不变。随后,我们揭示了一个跨语言选择瓶颈:在MGSM250回滚中,使用无可信标签的目标本地聚合规则缩小了平均选择差距的55-78%,而超过95%的修复需要真正的跨语言支持。该瓶颈在483个问题的MATH-500集上得以复制。受控训练审计显示,规则-GRPO提高了可信准确性,而奖励误差的变异性仍然较高。统一的信息是操作性的:多语言RLVR奖励在优化之前应按语言和答案接口进行审计。
cs.CL / 15 / 2608.20364
Hadith computational science in the age of large language models: a critical narrative review
大型语言模型时代的圣训计算科学:一项批判性叙事综述
Abstract
We examine how hadith computational science is being reshaped by transformer models, retrieval-grounded pipelines, and large language models (LLMs). Recent reviews document growth in the literature, but they do not yet provide a critical account of which advances are methodologically robust, which remain benchmark-bound, and which unresolved problems still limit scholarly use. We address this gap through a critical narrative review that combines critique of existing reviews, paper-level appraisal of representative original studies, and synthesis of Islamic scholar and domain-expert perspectives on authenticity, authority, and responsible use. We find uneven progress. Data resources have expanded, segmentation tasks have matured, narrator and source-verification problems are better formalized, and LLM-assisted workflows now support corpus-scale enrichment, multilingual access, and grounded evaluation. At the same time, progress remains constrained by narrow corpora, weak benchmark comparability, synthetic-to-real transfer gaps, narrator identity resolution, preprocessing fragility, limited reproducibility, and sparse expert-grounded validation. We show that important gaps lie beyond dominant benchmarks: non-canonical and obscure corpora, commentary and explanatory literature, cross-source links with Qur'an and seerah, and fiqh-facing evidence support. We argue that hadith computation should be assessed less as isolated model performance than as an evidence infrastructure problem requiring knowledge integration, provenance, and expert supervision. On this basis, we define a research agenda for making the field methodologically stronger and more useful to Islamic scholarship.
Chinese Translation
我们考察了圣训计算科学如何被变换器模型、检索基础管道和大型语言模型(LLMs)所重塑。近期的综述文献记录了该领域的增长,但尚未提供对哪些进展在方法论上是稳健的、哪些仍然受限于基准、以及哪些未解决的问题仍限制学术使用的批判性评估。我们通过一项批判性叙事综述来填补这一空白,该综述结合了对现有综述的批评、对代表性原始研究的论文级评估,以及对伊斯兰学者和领域专家在真实性、权威性和负责任使用方面的观点的综合。我们发现进展不均。数据资源有所扩展,分割任务已成熟,叙述者和来源验证问题得到了更好的形式化,LLM辅助的工作流程现在支持语料库规模的丰富、多语言访问和基于实证的评估。同时,进展仍受到狭窄语料库、弱基准可比性、合成与真实转移的差距、叙述者身份解析、预处理脆弱性、有限的可重复性以及稀疏的专家基础验证的限制。我们展示了重要的空白超出了主导基准:非经典和晦涩的语料库、评论和解释文献、与《古兰经》和《圣行》之间的跨源链接,以及面向教法的证据支持。我们认为,圣训计算应被评估为一个证据基础设施问题,而非孤立的模型性能问题,这需要知识整合、来源追溯和专家监督。在此基础上,我们定义了一项研究议程,以使该领域在方法论上更强大,并对伊斯兰学术更有用。
cs.CL / 16 / 2608.20365
Trilingual Topic Modeling of Sri Lankan Parliamentary Debates
斯里兰卡议会辩论的三语主题建模
Abstract
Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to standard NLP pipelines due to layout-complex PDFs, multilingual scripts, and agglutinative morphology. We present an end-to-end framework that addresses these challenges through LLM-based text extraction followed by a multilingual embedding and density-based clustering pipeline for topic modeling. A hybrid semantic-lexical extension, BiTopic, is further explored to improve interpretability and recover speeches otherwise discarded as noise. Applied to 19,553 speeches spanning 2017-2026, the pipeline recovers 30 macro-topics achieving a cluster purity (BCP) of 0.673, whose temporal trajectories align unsupervised with major national events including the 2019 Easter Sunday attacks and the 2022 economic crisis. Traditional LDA fails on this corpus due to cross-lingual fragmentation, whereas the proposed approach successfully identifies thematic structure across all three languages without supervision.
Chinese Translation
斯里兰卡议会辩论(Hansards)构成了一个包含僧伽罗语、泰米尔语和英语的三语语料库,其中包括混合语言内容,但由于复杂的PDF布局、多语言脚本和粘着性形态学,标准的自然语言处理(NLP)管道无法访问。我们提出了一个端到端框架,通过基于大语言模型(LLM)的文本提取,随后进行多语言嵌入和基于密度的聚类管道进行主题建模,以应对这些挑战。进一步探索了一种混合语义-词汇扩展方法BiTopic,以提高可解释性并恢复那些被视为噪声的演讲内容。该管道应用于2017年至2026年间的19,553篇演讲,恢复了30个宏观主题,聚类纯度(BCP)达0.673,其时间轨迹与包括2019年复活节袭击和2022年经济危机在内的重大国家事件无监督地对齐。传统的潜在狄利克雷分配(LDA)在该语料库上失败,原因在于跨语言的碎片化,而所提出的方法成功地在所有三种语言中识别了主题结构,无需监督。
cs.CL / 17 / 2608.20368
Research Paper Quality Recognition Through Textual Feature Analysis
通过文本特征分析识别学术论文质量
Abstract
Knowledge and innovations are shaped by using the quality and credibility of the scientific research. Yet, distinguishing between impactful, high-quality work and flawed studies remains a challenge. This paper introduces a benchmark for classifying research papers into two categories: good (highly cited) and non-good (retracted), using only textual features from titles and abstracts. We evaluate multiple embedding techniques, including SBERT, Word2Vec, FastText, USE, and TF-IDF, combined with classifiers such as Support Vector Machines (SVM), Random Forests, and Neural Networks. Our contributions include: (1) hyperparameter transparency, (2) feature space visualizations using t-SNE, (3) model interpretability analysis with SHAP, and (4) detailed examination of error cases. Experimental results show that a neural network with SBERT embeddings achieves 87.22\% accuracy, while FastText combined with SVM reaches 91.12\%. These findings highlight the value of textual information in assessing research quality, with ethical considerations for deployment. This work contributes toward the development of academic integrity tools that promote trustworthy scholarship.
Chinese Translation
知识和创新是通过科学研究的质量和可信度来塑造的。然而,区分有影响力的高质量研究和存在缺陷的研究仍然是一项挑战。本文介绍了一种基准,将研究论文分为两类:良好(高被引)和非良好(撤回),仅使用标题和摘要中的文本特征。我们评估了多种嵌入技术,包括 SBERT、Word2Vec、FastText、USE 和 TF-IDF,并结合支持向量机(SVM)、随机森林和神经网络等分类器。我们的贡献包括:(1)超参数透明性,(2)使用 t-SNE 的特征空间可视化,(3)使用 SHAP 的模型可解释性分析,以及(4)对错误案例的详细检查。实验结果表明,使用 SBERT 嵌入的神经网络达到了 87.22\% 的准确率,而结合 SVM 的 FastText 达到了 91.12\\%。这些发现突显了文本信息在评估研究质量中的价值,并考虑了部署的伦理问题。本研究为开发促进可信学术的学术诚信工具做出了贡献。
cs.CL / 18 / 2608.20369
ASTAR: Automated induction of STAndardized radiology Reporting templates from large-scale clinical free-text corpora
ASTAR:从大规模临床自由文本语料库自动生成标准化放射学报告模板
Abstract
Structured reporting converts free-text radiology narratives into queryable data keys, facilitating cohort assembly, longitudinal tracking, and training label generation for medical AI. The prevailing paradigm follows a two-stage pipeline: (1) constructing a reporting template, (2) extracting information to populate it. While the extraction stage has benefited from advances in large language models (LLMs), template construction remains a manual bottleneck relying on labor-intensive expert consensus that is static, difficult to scale, and may fail to capture real-world reporting diversity. We address this limitation with \textbf{\texttt{ASTAR}}, an LLM-based framework for Automated induction of STAndardized radiology Reporting templates from large-scale clinical free-text corpora. Extensive experiments on 4,215 fetal brain MRI reports from multiple centers demonstrate that the \textbf{\texttt{ASTAR}}-induced template surpasses two expert-curated templates across template coverage, information fidelity, diagnostic fidelity, and expert-rated usability, reducing template development from weeks of committee deliberation to hours of automated processing. Code: https://github.com/birthlab/ASTAR
Chinese Translation
结构化报告将自由文本放射学叙述转换为可查询的数据键,促进了队列组装、纵向跟踪以及医疗人工智能的训练标签生成。当前的范式遵循两阶段流程:(1)构建报告模板,(2)提取信息以填充模板。尽管提取阶段受益于大型语言模型(LLMs)的进展,但模板构建仍然是一个依赖于劳动密集型专家共识的手动瓶颈,这种方式静态、难以扩展,并且可能无法捕捉到现实世界报告的多样性。我们通过 extbf{ exttt{ASTAR}}解决了这一限制, extbf{ exttt{ASTAR}}是一个基于LLM的框架,用于从大规模临床自由文本语料库自动生成标准化放射学报告模板。在来自多个中心的4,215份胎儿脑部MRI报告上的广泛实验表明, extbf{ exttt{ASTAR}}生成的模板在模板覆盖、信息保真度、诊断保真度和专家评估的可用性方面均优于两个专家策划的模板,将模板开发时间从数周的委员会审议缩短到数小时的自动处理。代码链接:https://github.com/birthlab/ASTAR
cs.CL / 19 / 2608.20371
When Do LLMs Replace Fine-Tuned NLU? A Decision Framework for Intent Detection in Production Conversational Systems
大型语言模型何时取代微调的自然语言理解?生产对话系统中意图检测的决策框架
Abstract
A common claim is that zero-shot large language models (LLMs) can replace fine-tuned NLU classifiers for intent detection. We test this claim head-to-head and find that the honest answer is: it depends on the intent space. On full ATIS and CLINC150 we compare a fine-tuned RoBERTa, a TF-IDF+logistic-regression baseline, sentence-embedding kNN, and Claude Haiku zero-shot, reporting bootstrap 95% confidence intervals and paired significance tests. When abundant in-domain labels exist, fine-tuned RoBERTa is as good or better and three orders of magnitude cheaper and faster: on ATIS it beats Claude zero-shot by 11.8 points (95.9 vs. 84.1, p<0.001). On the broad 150-intent CLINC150 schema the two are statistically tied (89.1 vs. 88.5, p=0.24): the LLM matches a fully supervised model with no training data. The LLM's advantages appear in three production-relevant regimes: out-of-scope detection (OOS recall 85.6 vs. 58.1 for RoBERTa); robustness to realistic ASR noise via a controlled text-to-speech to noise to Whisper pipeline (92.5 vs. 80.0 at 0 dB); and dynamic per-deployment schemas, where a classifier trained on one app's intents scores 0% on a new app's intents while the schema-prompted LLM serves both at ~94% with zero retraining. We distill these findings into a decision framework for practitioners.
Chinese Translation
一个普遍的说法是,零-shot大型语言模型(LLMs)可以替代微调的自然语言理解(NLU)分类器进行意图检测。我们对此说法进行了正面测试,发现诚实的答案是:这取决于意图空间。在完整的ATIS和CLINC150数据集上,我们比较了微调的RoBERTa、TF-IDF+逻辑回归基线、句子嵌入kNN和Claude Haiku零-shot,报告了自助法95%置信区间和配对显著性检验。当存在丰富的领域内标签时,微调的RoBERTa表现同样优秀或更好,且成本和速度便宜三个数量级:在ATIS上,它比Claude零-shot高出11.8个百分点(95.9 vs. 84.1,p<0.001)。在广泛的150意图CLINC150架构中,两者的统计结果相当(89.1 vs. 88.5,p=0.24):LLM在没有训练数据的情况下与完全监督模型相匹配。LLM的优势体现在三个与生产相关的情境中:超出范围检测(OOS召回率85.6 vs. 58.1,针对RoBERTa);对现实ASR噪声的鲁棒性,通过受控的文本到语音到噪声到Whisper管道(92.5 vs. 80.0,0 dB时);以及动态每个部署架构,其中在一个应用的意图上训练的分类器在新应用的意图上得分为0%,而架构提示的LLM在不进行重新训练的情况下为两者提供约94%的服务。我们将这些发现提炼为一个供从业者使用的决策框架。
cs.CL / 20 / 2608.20373
An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study
用于评估大型语言模型在临床注册抽象中表现的模糊性分类法:一项多中心前瞻性研究
Abstract
Objective: To evaluate large language model (LLM) performance on unprocessed electronic medical record (EMR) data for clinical registry abstraction. Methods: We evaluated LLM performance answering registry questions for the American College of Cardiology National Cardiovascular Data Registry (ACC NCDR). In a pilot study at an academic medical center, the model identified candidate data sources for each registry question and experienced abstractors used these results to define question-specific document sets. In a validation study at a second center with a second ACC NCDR registry, the LLM answered questions using the question-specific document sets. Before reviewing any output, two abstractors independently established the ground truth and assigned each question to one of six categories, ordered by the ambiguity and clinical reasoning required to resolve it: Medication/Event Flag, Binary Clinical Presence, Administrative, Quantitative Laboratory/Physiologic, Clinical Interpretation, and Event Timing. Results: The analytical sample comprised 9,430 abstractor answers reconciled to 4,715 consensus answers (501 pilot; 4,214 validation). In the pilot, candidate data sources per question averaged between 14.6 (SD 13.9) for demographics and 89.2 (SD 56.1) for history and risk factors. In validation, human inter-rater agreement was approximately 98\% while 87\% of LLM answers exactly matched consensus, 2\% partially, and 9\% did not. Mean question-level accuracy was 91.5\% (SD 13.4\%) across 157 questions with at least 20 answers, and declined as ambiguity increased, from 96\% for Medication/Event Flag to 62\% for Event Timing questions. Conclusions: LLMs answering clinical registry questions on unprocessed EMR data achieved far lower accuracy than human abstractors. LLM accuracy fell steadily as ambiguity and the level of required clinical reasoning increased.
Chinese Translation
目的:评估大型语言模型(LLM)在未处理的电子病历(EMR)数据上进行临床注册抽象的表现。方法:我们评估了LLM在回答美国心脏病学会国家心血管数据注册(ACC NCDR)注册问题时的表现。在一项学术医疗中心的试点研究中,该模型为每个注册问题识别候选数据源,经验丰富的抽象者利用这些结果定义问题特定的文档集。在第二个中心的验证研究中,使用第二个ACC NCDR注册,LLM使用问题特定的文档集回答问题。在审查任何输出之前,两名抽象者独立建立了真实情况,并将每个问题分配到六个类别之一,按解决所需的模糊性和临床推理的程度排序:药物/事件标志、二元临床存在、行政、定量实验室/生理、临床解释和事件时机。结果:分析样本包括9430个抽象者答案,与4715个共识答案进行了调和(501个试点;4214个验证)。在试点中,每个问题的候选数据源平均为14.6(标准差13.9)个用于人口统计学,89.2(标准差56.1)个用于病史和风险因素。在验证中,人类评分者间一致性约为98\%,而LLM答案中87 ext{%}与共识完全匹配,2 ext{%}部分匹配,9 ext{%}不匹配。157个问题中,至少有20个答案的平均问题级准确率为91.5 ext{%}(标准差13.4 ext{%}),并随着模糊性的增加而下降,从药物/事件标志的96 ext{%}降至事件时机问题的62 ext{%}。结论:在未处理的EMR数据上回答临床注册问题的LLM的准确性远低于人类抽象者。随着模糊性和所需临床推理水平的增加,LLM的准确性稳步下降。
cs.CL / 21 / 2608.20374
VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models
VA-DPO:用于语言模型可控情感生成的情感-唤醒直接偏好优化
Abstract
How precisely can we tell a language model how to feel? Most work on emotional generation answers with a discrete label - happy, angry, sad - which cannot express a target like "mildly downcast but calm." We instead specify the desired affect as a continuous point (v*, a*) in the Valence-Arousal plane and train the model to hit it. Our method, VA-DPO, is a small modification to Direct Preference Optimization: a frozen VA regressor scores each sampled generation by its Euclidean distance to the target, we keep only candidate pairs whose distance gap clears a margin tau, and we optimize a LoRA adapter with the ordinary DPO loss against a frozen reference. The DPO objective itself is unchanged; what is new is how the preference data is built. On Llama-3.1-8B-Instruct this cuts mean VA distance to the target by 33% over system-prompting and 25% over few-shot prompting, lifting valence/arousal correlation to r_v=0.93 and r_a=0.75. The gains carry over to Qwen3-8B and Llama-3.2-3B, and they do not come at the usual price: MMLU is unchanged (Delta=+0.0) and HellaSwag and TruthfulQA are preserved. We release the code, configs, and the preference-construction pipeline.
Chinese Translation
我们能多精确地告诉语言模型如何感受?大多数情感生成的研究通过离散标签来回答——快乐、愤怒、悲伤——这些标签无法表达诸如“略微沮丧但平静”的目标。我们则将期望的情感指定为情感-唤醒平面上的一个连续点(v*, a*),并训练模型以达到该点。我们的方法VA-DPO是对直接偏好优化(Direct Preference Optimization)的一个小修改:一个冻结的情感-唤醒回归器通过其与目标的欧几里得距离来评分每个生成样本,我们仅保留距离差距超过边际tau的候选对,并使用普通的DPO损失优化一个LoRA适配器,针对一个冻结的参考。DPO目标本身没有改变;新的在于偏好数据的构建方式。在Llama-3.1-8B-Instruct上,这种方法将平均情感-唤醒距离目标的距离减少了33%(相较于系统提示)和25%(相较于少量示例提示),使得情感/唤醒的相关性提升至r_v=0.93和r_a=0.75。这些收益在Qwen3-8B和Llama-3.2-3B上也得以延续,并且没有付出通常的代价:MMLU保持不变(Delta=+0.0),HellaSwag和TruthfulQA也得以保留。我们发布了代码、配置和偏好构建流程。
cs.CL / 22 / 2608.20375
GRAFT: Adaptive DLM-Based Draft Tree Construction with Target-Distilled Edge Scoring
GRAFT:基于自适应扩散语言模型的目标蒸馏边评分的草图树构建
Abstract
Tree-based speculative decoding raises the mean accepted tokens of standard speculative decoding by verifying multiple draft paths, and existing tree builders typically construct these paths through parent-conditioned expansion, where each child token is generated conditioned on its parent path. This construction is incompatible with diffusion language model (DLM) drafters such as DFlash, which produces all future-position distributions in a single forward pass. DDTree bridges this gap by treating high-probability tokens from each future-position distribution as candidate nodes and selecting edges between consecutive positions under a fixed node budget. However, its edge selection relies on token probability alone without modeling parent--child compatibility, so target-compatible tokens can be attached to wrong parents; moreover, its fixed budget ignores that the throughput-optimal tree size varies with the decoding state. We propose GRAFT, a draft-tree construction framework for DLM-based speculative decoding. GRAFT introduces Target-Distilled Edge Scoring (TDES), which distills parent--child preferences from target-model traces to select target-compatible edges, and State-Aware Budget Allocation (SABA), which sets the per-round tree budget by balancing expected draft gain against verification cost. Across multiple models and tasks, GRAFT achieves $2.13\times$--$6.36\times$ end-to-end speedup over autoregressive decoding while adding less than $0.5$\,ms of overhead per round, approximately $1.4\%$ of the target-model verification latency.
Chinese Translation
基于树的推测解码通过验证多个草图路径提高了标准推测解码的平均接受令牌数,而现有的树构建器通常通过父节点条件扩展来构建这些路径,其中每个子令牌是在其父路径的条件下生成的。这种构建方式与扩散语言模型(DLM)草图生成器(如 DFlash)不兼容,后者在一次前向传递中生成所有未来位置的分布。DDTree 通过将来自每个未来位置分布的高概率令牌视为候选节点,并在固定节点预算下选择连续位置之间的边来弥补这一差距。然而,其边选择仅依赖于令牌概率,而未建模父子兼容性,因此目标兼容的令牌可能会附加到错误的父节点;此外,其固定预算忽略了最佳吞吐量树的大小随解码状态而变化。我们提出了 GRAFT,一个基于 DLM 的推测解码草图树构建框架。GRAFT 引入了目标蒸馏边评分(TDES),从目标模型轨迹中提取父子偏好以选择目标兼容的边,并引入了状态感知预算分配(SABA),通过平衡预期草图增益与验证成本来设定每轮的树预算。在多个模型和任务中,GRAFT 实现了比自回归解码快 $2.13 imes$ 到 $6.36 imes$ 的端到端加速,同时每轮增加的开销不足 $0.5$ 毫秒,约占目标模型验证延迟的 $1.4\%$。
cs.CL / 23 / 2608.20376
TH-GNN: Heterogeneous Temporal Graph Neural Networks for LLM-Agent Shilling Attack Detection
TH-GNN:用于LLM代理虚假评价攻击检测的异构时序图神经网络
Abstract
LLM agents can now generate realistic shilling profiles, fluent reviews, and coherent ratings at scale, systematically defeating recommender-system defenses. Text-only detectors that flag semantic drift in review embeddings are blind to graph structure and temporal coordination, while graph-only detectors that exploit neighborhood anomalies cannot reason over review semantics or the cross-modal inconsistencies produced by LLM-generated content. We propose TH-GNN, a heterogeneous temporal graph neural network with a two-layer Heterogeneous Graph Transformer backbone that applies per-type and per-relation attention augmented with learnable sinusoidal temporal encodings on every edge. Cross-modal attention fuses structural user embeddings with frozen RoBERTa representations of reviews and item descriptions, while a GRU operating over log inter-arrival times captures temporal burstiness. Evaluated across five attack families and four benchmark datasets, TH-GNN achieves a grand-mean F1 score of 0.870, outperforming the strongest text-only baseline on Agent4SR attacks by 10.9 percentage points and 11.5 percentage points at the lowest injection rate. These results demonstrate the effectiveness of jointly modeling temporal, structural, and semantic signals for detecting sophisticated LLM-driven shilling attacks.
Chinese Translation
LLM代理现在能够大规模生成逼真的虚假评价档案、流畅的评论和连贯的评分,系统性地击败推荐系统的防御措施。仅基于文本的检测器通过标记评论嵌入中的语义漂移来工作,但对图结构和时间协调视而不见;而仅基于图的检测器利用邻域异常却无法推理评论语义或LLM生成内容所产生的跨模态不一致性。我们提出了TH-GNN,这是一种异构时序图神经网络,具有两层异构图变换器(Heterogeneous Graph Transformer)主干,针对每种类型和每种关系应用增强的可学习正弦时间编码的注意力机制。跨模态注意力将结构化用户嵌入与冻结的RoBERTa评论和项目描述表示融合,而在对数到达时间上运行的GRU则捕捉时间突发性。在五个攻击类别和四个基准数据集上的评估中,TH-GNN实现了0.870的总平均F1分数,在Agent4SR攻击中比最强的仅基于文本的基线高出10.9个百分点,在最低注入率下高出11.5个百分点。这些结果证明了联合建模时间、结构和语义信号在检测复杂的LLM驱动虚假评价攻击中的有效性。
cs.CL / 24 / 2608.20381
EditPPT: Faithful Long-Deck Slide Editing via Structured Tool-Using Multi-Agent with Dual-Modal Validators
EditPPT:通过结构化工具使用的多智能体与双模态验证器实现忠实的长篇幻灯片编辑
Abstract
Automating slide editing requires simultaneously satisfying modification accuracy, preservation fidelity, and robustness to deck length. Existing LLM-based systems often fail on real-world presentation files because they rely on idealized intermediate representations or open-ended code generation, which are prone to cascading errors in long decks. We introduce EditPPT, a multi-agent framework that reformulates slide editing as a constrained tool-selection problem. By executing localized shape-level operations through the native PowerPoint COM interface, EditPPT narrows the LLM action space while preserving the application-resolved structure of user-authored decks. By separating validation across modalities, our dual-modal validation provides more robust assessment of both instruction fidelity and visual quality. We also present DeckEdit-Bench, a benchmark with 28 human-authored decks, 582 slides, and 183 editing prompts across short, medium, and long deck tiers. Experiments show that EditPPT achieves a 99.5% execution rate, 88.7% slide-targeting F1, 82.5% instruction following, and 91.5% object preservation overall, while maintaining strong performance on long decks. Our code and benchmark are available at https://anonymous.4open.science/r/EditPPT-0E27/
Chinese Translation
自动化幻灯片编辑需要同时满足修改准确性、保留忠实度以及对幻灯片长度的鲁棒性。现有的基于大型语言模型(LLM)系统在处理真实世界的演示文件时常常失败,因为它们依赖于理想化的中间表示或开放式代码生成,这在长篇幻灯片中容易导致级联错误。我们提出了EditPPT,一个将幻灯片编辑重新构建为受限工具选择问题的多智能体框架。通过使用原生PowerPoint COM接口执行局部形状级操作,EditPPT缩小了LLM的动作空间,同时保留了用户创作的幻灯片的应用解析结构。通过跨模态分离验证,我们的双模态验证提供了对指令忠实度和视觉质量的更强鲁棒性评估。我们还提出了DeckEdit-Bench,这是一个包含28个人工创作幻灯片、582张幻灯片和183个编辑提示的基准,涵盖短、中和长篇幻灯片层级。实验表明,EditPPT实现了99.5%的执行率、88.7%的幻灯片目标F1、82.5%的指令遵循率以及91.5%的对象保留率,同时在长篇幻灯片上保持了强劲的表现。我们的代码和基准可在https://anonymous.4open.science/r/EditPPT-0E27/获取。
cs.CL / 25 / 2608.20382
Decoupled Vision-Language System for Multimodal Understanding and Generation
用于多模态理解与生成的解耦视觉-语言系统
Abstract
We introduce a new architecture design for multimodal large language models (MLLMs), Libra, capable of both multimodal understanding and generation. Libra architecture contains one vision system and one language system, connected by cross-modal bridges. This design decouples self-modal modeling and cross-modal interaction, enabling each modality to learn its unique representations while maintaining effective cross-modal comprehension. The decoupling is mainly achieved in a switch attention module and a switch FFN module, which dynamically routes the computation flow for self-modal modeling and cross-modal interaction scenarios. We evaluate the effectiveness in two important settings: \textbf{Libra-1} for the understanding-only image-to-text setting, and \textbf{Libra-2} for unified image-to-text understanding and text-to-image generation. In addition to the architecture design, we discuss various improvements on tokenization, positional encoding, and supervision. Experiments demonstrate that the dedicated Libra design enables mutual improvements on multimodal understanding and generation, achieving strong performance on both understanding and generation benchmarks.
Chinese Translation
我们提出了一种新的多模态大语言模型(MLLM)的架构设计,称为Libra,能够同时进行多模态理解和生成。Libra架构包含一个视觉系统和一个语言系统,通过跨模态桥接相连接。该设计解耦了自模态建模和跨模态交互,使每个模态能够学习其独特的表示,同时保持有效的跨模态理解。解耦主要通过一个切换注意力模块和一个切换前馈网络(FFN)模块实现,这些模块动态地引导计算流以适应自模态建模和跨模态交互场景。我们在两个重要设置中评估了其有效性: extbf{Libra-1}用于仅理解的图像到文本设置,以及 extbf{Libra-2}用于统一的图像到文本理解和文本到图像生成。除了架构设计,我们还讨论了在标记化、位置编码和监督方面的各种改进。实验表明,专门的Libra设计能够在多模态理解和生成之间实现相互提升,在理解和生成基准测试中均取得了良好的性能。
cs.CL / 26 / 2608.20385
Using Human-LLM Disagreement to Improve Checklist-Based Quality Appraisal
利用人类与大型语言模型的分歧来改善基于清单的质量评估
Abstract
Systematic reviews rely on quality appraisal of included studies, a process that is time-consuming and sensitive to ambiguity in checklist criteria. Although large language models (LLMs) offer opportunities to support these tasks, appraisal checklists are typically treated as fixed inputs, and it remains unclear how their design affects agreement with expert judgments. Therefore, we investigate (1) whether LLMs can approximate human judgments in checklist-based appraisal and (2) whether patterns of human-LLM disagreement can be used to identify and improve ambiguous checklist items. Using the Guidelines for Reporting on Latent Trajectory Studies (GRoLTS) checklist, we compare LLM-generated assessments with expert annotations across three research topics and two checklist versions. Agreement is assessed using item-level accuracy, chance-corrected agreement, and preservation of study-level rank ordering. We find that performance varies substantially across checklist items, with ambiguous and conditional criteria producing the greatest disagreement. Revising these items improves both raw and chance-corrected agreement. Although item-level misclassifications persist, LLM-generated scores often preserve the relative ranking of studies when high-agreement items are retained. These results indicate that reliable LLM-assisted appraisal depends not only on model choice but also on checklist design. The findings suggest that analyzing human-LLM disagreement can help identify problematic checklist items and support the iterative improvement of research synthesis workflows.
Chinese Translation
系统评价依赖于对纳入研究的质量评估,这一过程既耗时又对清单标准的模糊性敏感。尽管大型语言模型(LLMs)为支持这些任务提供了机会,但评估清单通常被视为固定输入,其设计如何影响与专家判断的一致性仍不明确。因此,我们研究了(1)LLMs是否能够在基于清单的评估中接近人类判断,以及(2)人类与LLMs之间的分歧模式是否可以用于识别和改进模糊的清单项目。我们使用潜在轨迹研究报告指南(Guidelines for Reporting on Latent Trajectory Studies, GRoLTS)清单,比较LLM生成的评估与专家注释在三个研究主题和两个清单版本之间的一致性。通过项目级准确性、机会校正一致性和研究级排名保留来评估一致性。我们发现,不同清单项目的表现差异显著,模糊和条件标准产生了最大的分歧。修订这些项目改善了原始和机会校正的一致性。尽管项目级错误分类仍然存在,但在保留高一致性项目时,LLM生成的评分通常能够保持研究的相对排名。这些结果表明,可靠的LLM辅助评估不仅依赖于模型选择,还依赖于清单设计。研究结果表明,分析人类与LLMs之间的分歧可以帮助识别有问题的清单项目,并支持研究综合工作流程的迭代改进。
cs.CL / 27 / 2608.20387
Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions
Poly-InstructTTS:从开放式指令中学习野外环境下的富有表现力的语音合成
Abstract
While recent text-to-speech (TTS) models achieve high naturalness, controlling fine-grained expression via natural-language instructions remains challenging. We introduce Poly- InstructTTS, which learns expressive speech from open-ended instructions using in-the-wild audiovisual data. We build a scalable multi-modal pipeline to construct a 1,000-hour instruction-annotated corpus covering 1,000+ fine-grained emotions and styles. The framework uses a prompt-free GPT with attribute-based thinking tokens, followed by a flow-matching module that injects timbre from a reference audio. We also present a speaker fine-tuning procedure to transfer instruction control to specific speakers while preserving persona. We further extend InstructTTSEval with broader tasks. Experiments show that Poly-InstructTTS delivers strong performance in instruction adherence and expressiveness. Audio demos and the expanded testset are available on our project page.
Chinese Translation
尽管最近的文本到语音(TTS)模型在自然性上取得了很高的成就,但通过自然语言指令控制细粒度表达仍然具有挑战性。我们介绍了Poly-InstructTTS,它利用野外的视听数据从开放式指令中学习富有表现力的语音。我们构建了一个可扩展的多模态管道,构建了一个包含1,000小时指令标注语料库的系统,涵盖1,000多种细粒度情感和风格。该框架使用无提示的GPT与基于属性的思维标记,随后是一个流匹配模块,从参考音频中注入音色。我们还提出了一种说话人微调程序,以在保留个性的同时将指令控制转移到特定说话人。我们进一步扩展了InstructTTSEval,增加了更广泛的任务。实验表明,Poly-InstructTTS在指令遵循性和表现力方面表现出色。音频演示和扩展测试集可在我们的项目页面上获取。
cs.CL / 28 / 2608.20388
Intent Engine: Natural-Language Intent Translation for Intent-Driven Orchestration in the Compute Continuum
意图引擎:计算连续体中面向意图的编排的自然语言意图翻译
Abstract
Microservice placement in the compute continuum is driven by low-level Service-level Objectives (SLOs), but requiring users to specify metric-level constraints creates an adoption barrier and increases misconfiguration risk. Although large language models (LLMs) can interpret natural-language intents, direct generation of orchestration-consumable SLO artifacts remains unreliable due to unsupported constraints, incorrect grounded values, and schema violations. These errors can propagate to downstream placement logic and produce infeasible or incorrect placements. This paper presents Intent Engine, a natural-language intent translation architecture that constructs validated SLO artifacts for compute-continuum service placement. Intent Engine acts as an intent acquisition and SLO construction layer for existing intent-driven orchestration and placement frameworks; it does not perform placement or runtime QoS optimization. The architecture combines schema-constrained extraction, retrieval-grounded value construction from monitored infrastructure state, and validation against supported constraints before emitting the final SLO artifact. We evaluate Intent Engine using a 716-record intent-to-SLO dataset derived from an edge-cloud testbed, including valid and invalid intents. Across GPT-4.1 mini, Claude Sonnet 4.5, and DeepSeek V4-Flash, Intent Engine outperforms prompting baselines and a non-LLM rule-based parser. With GPT-4.1 mini, it achieves 0.941 total F1 Score and reduces aggregate hallucination by 85.1%, while lowering downstream placement failure from 30.8% to 2.1%.
Chinese Translation
计算连续体中的微服务部署受到低层次服务级目标(SLOs)的驱动,但要求用户指定度量级约束会造成采用障碍并增加配置错误的风险。尽管大型语言模型(LLMs)能够解释自然语言意图,但由于不支持的约束、错误的基础值和模式违规,直接生成可供编排使用的 SLO 工件仍然不可靠。这些错误可能会传播到下游的部署逻辑中,导致不可行或错误的部署。本文提出了意图引擎(Intent Engine),一种自然语言意图翻译架构,用于构建经过验证的 SLO 工件,以支持计算连续体中的服务部署。意图引擎作为现有面向意图的编排和部署框架的意图获取和 SLO 构建层;它不执行部署或运行时 QoS 优化。该架构结合了模式约束提取、基于监控基础设施状态的检索基础值构建,以及在发出最终 SLO 工件之前对支持约束的验证。我们使用从边缘云测试平台衍生的 716 条意图到 SLO 的数据集评估意图引擎,包括有效和无效的意图。在 GPT-4.1 mini、Claude Sonnet 4.5 和 DeepSeek V4-Flash 上,意图引擎的表现优于提示基线和非 LLM 的基于规则的解析器。在 GPT-4.1 mini 上,它实现了 0.941 的总 F1 分数,并将整体幻觉减少了 85.1%,同时将下游部署失败率从 30.8% 降低到 2.1%。
cs.CL / 29 / 2608.20390
Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations
Ansari:一个基于检索的伊斯兰人工智能助手——架构、部署及来自14万次对话的经验教训
Abstract
General-purpose large language models (LLMs) are increasingly used to answer religious questions, but for Islamic content they carry two serious risks: factual fabrication (inventing Qur'anic verses or hadith) and subtle value misalignment. We present Ansari, a deployed, retrieval-grounded Islamic AI assistant that has handled more than 140,000 conversations across 25+ languages since June 2023. Ansari is built around an agentic retrieval loop: a tool-using language model issues searches against authenticated Islamic corpora -- the Qur'an, hadith collections, a multi-volume jurisprudence (fiqh) encyclopedia, and exegetical (tafsir) sources -- and answers only on the basis of what it retrieves, with citations attached for verification. We describe the system's architecture (the agent loop, the retrieval tools, the corpora, and the system prompt that encodes editorial and theological policy), its multi-platform deployment (web, mobile, WhatsApp, and as a Model Context Protocol server and an Agent Skill), and what 140,000 real conversations reveal about how Muslims actually use such a tool. We report results on several complementary evaluations -- zero-shot performance on accredited institutional exams, a human-rated validation during Ramadan, and two independent, externally run benchmarks on which Ansari currently tops the public IslamicMMLU leaderboard ahead of frontier models and is competitive on Islamic legal reasoning (IslamicLegalBench) while strongly resisting false premises -- and draw out lessons that generalize beyond Islam to any faith- or values-sensitive deployment of LLMs: grounding is necessary but not sufficient, the system prompt is a theological as much as a technical artifact, and the absence of community in how models are formed remains a hard gap.
Chinese Translation
通用的大型语言模型(LLMs)越来越多地用于回答宗教问题,但在伊斯兰内容方面,它们存在两个严重风险:事实捏造(虚构《古兰经》经文或圣训)和微妙的价值观不一致。我们介绍了Ansari,一个已部署的基于检索的伊斯兰人工智能助手,自2023年6月以来处理了超过14万次对话,覆盖25种以上语言。Ansari围绕一个代理检索循环构建:一个使用工具的语言模型对经过认证的伊斯兰语料库进行搜索——包括《古兰经》、圣训集、多卷法学(fiqh)百科全书和注释(tafsir)来源——并仅基于其检索结果进行回答,同时附上引用以供验证。我们描述了系统的架构(代理循环、检索工具、语料库以及编码编辑和神学政策的系统提示)、其多平台部署(网页、移动端、WhatsApp,以及作为模型上下文协议服务器和代理技能)以及14万次真实对话揭示的穆斯林如何实际使用此类工具的情况。我们报告了几项互补评估的结果——在认证机构考试中的零-shot表现、在斋月期间的人类评分验证,以及两个独立的外部基准测试,Ansari目前在公共的IslamicMMLU排行榜上领先于前沿模型,并在伊斯兰法律推理(IslamicLegalBench)方面具有竞争力,同时强烈抵制错误前提——并总结出一些超越伊斯兰的教训,这些教训适用于任何与信仰或价值观相关的LLMs部署:基础是必要但不充分的,系统提示既是神学的也是技术的产物,以及在模型形成过程中缺乏社区参与仍然是一个难以弥补的缺口。
cs.CL / 30 / 2608.20391
ImmigrationReason: A Structured Dataset of U.S. Immigration Appeals for Legal Reasoning Research
移民理由:美国移民上诉的结构化数据集用于法律推理研究
Abstract
Most legal NLP resources draw from federal case law and focus on coarse classification, leaving administrative adjudication, where the vast majority of government decisions occur, essentially unaddressed. We introduce ImmigrationReason, a large-scale structured dataset derived from 12,375 non-precedent decisions of the U.S. Citizenship and Immigration Services (USCIS) Administrative Appeals Office (AAO) spanning 2005 to 2026. Each record captures the applicable legal framework, per-criterion evidence-sufficiency findings under a five-category label, verbatim adjudicator-criticism quotes, all citations, and final dispositions, alongside high-quality Claude-transcribed source text. Extraction quality is validated through a three-pass pipeline combining two independent modalities with comparison-prompt adjudication by Opus 4.7, and verified by domain experts on a 500-record sample. The dataset documents nearly 9,000 verbatim instances of AAO-identified legal errors, spans a natural legal-regime transition (the 2016 Dhanasar rule change), and covers 21 years of adjudication. We analyze the dataset in detail and outline research directions it enables, from outcome prediction and adjudicator-error analysis to agent design for high-stakes regulatory domains.
Chinese Translation
大多数法律自然语言处理(NLP)资源来源于联邦案例法,并侧重于粗略分类,基本上忽视了行政裁决,而政府决策的绝大多数发生在此。我们介绍了移民理由(ImmigrationReason),这是一个大规模的结构化数据集,源自美国公民及移民服务局(USCIS)行政上诉办公室(AAO)在2005年至2026年间的12,375个非先例决定。每条记录捕捉适用的法律框架、基于五类标签的逐项证据充分性发现、逐字的裁决者批评引用、所有引用以及最终裁决,同时附有高质量的Claude转录源文本。通过结合两种独立模式的三次提取流程以及Opus 4.7的比较提示裁决,验证了提取质量,并由领域专家在500条记录样本上进行了验证。该数据集记录了近9,000个AAO识别的逐字法律错误实例,跨越了一个自然法律制度的转变(2016年Dhanasar规则变更),并涵盖了21年的裁决。我们详细分析了该数据集,并概述了其所启用的研究方向,从结果预测和裁决者错误分析到高风险监管领域的代理设计。
cs.CL / 31 / 2608.20392
Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants
评估即搜索:会议助手中基础失败的自适应发现
Abstract
LLM-powered meeting assistants are deployed at scale, yet systematic evaluation of their grounding fidelity remains limited to static benchmarks that miss failure modes tied to specific discourse structures or reasoning demands. We propose Evaluation-as-Search (EaS), a feedback-driven methodology that frames quality evaluation as an adaptive search over the space of natural questions a meeting participant might ask. Rather than sampling uniformly, EaS learns from evaluator feedback across iterations to concentrate probing effort on cognitive demands where failures are most likely, guided by a UCB-scored coverage map and blind multi-dimensional quality evaluation. Using EaS, we construct MeetingProbe, a benchmark of over $3{,}000$ annotated question--answer pairs spanning 20 transcripts from three meeting genres and three LLM assistants. In ablations, adaptive search surfaces $2.5\times$ more failures than random probing ($7.1\%$ vs. $2.9\%$ finding rate), with the strategic planner contributing the largest individual effect. Across three models, we observe a clear capability gradient and identify eight recurring failure categories dominated by discourse-pragmatic challenges rather than factual recall errors. We further validate MeetingProbe across multiple model families and providers, finding a clean capability gradient and a curated subset of universal failures that no model handles. MeetingProbe is released publicly to support reproducible evaluation of meeting assistant grounding fidelity.
Chinese Translation
基于大语言模型(LLM)的会议助手已大规模部署,但对其基础可靠性的系统评估仍然局限于静态基准,未能捕捉与特定话语结构或推理需求相关的失败模式。我们提出了评估即搜索(Evaluation-as-Search, EaS),这是一种以反馈驱动的方法论,将质量评估框架化为会议参与者可能提出的自然问题空间的自适应搜索。EaS并非均匀采样,而是通过迭代学习评估者的反馈,将探测工作集中在最可能出现失败的认知需求上,借助UCB评分的覆盖图和盲多维质量评估。利用EaS,我们构建了MeetingProbe,这是一个包含超过3000个注释问答对的基准,涵盖来自三种会议类型和三种LLM助手的20个转录文本。在消融实验中,自适应搜索比随机探测发现了2.5倍更多的失败(发现率为7.1%对比2.9%),其中战略规划者贡献了最大的个体效应。在三种模型中,我们观察到明显的能力梯度,并识别出八个主要由话语-语用挑战而非事实回忆错误主导的重复失败类别。我们进一步在多个模型家族和提供者中验证了MeetingProbe,发现了一个清晰的能力梯度和一个经过策划的普遍失败子集,任何模型都未能处理。MeetingProbe已公开发布,以支持会议助手基础可靠性的可重复评估。
cs.CL / 32 / 2608.20393
Knowledge-Graph-Gated Defactualization for Style-Controllable and Fact-Preserving Generation in Agentic Conversational AI
知识图谱门控去事实化:在代理对话人工智能中实现风格可控和事实保留的生成
Abstract
Agentic large language models (LLMs) deployed in fact-sensitive applications such as customer support must simultaneously preserve factual correctness and generate responses in a controllable stylistic register. Activation steering enables fine-tuning-free style control by perturbing hidden representations, but it lacks an explicit mechanism for distinguishing verifiable facts from stylistic content, leading to semantic leakage. We address this challenge through \emph{Defactualize-Steer-Rehydrate} (DSR), a knowledge-engineering framework that integrates a typed, salience-weighted knowledge graph (KG) with activation steering. DSR extracts salient entities using a layered regex or NER or lexical-classifier pipeline, replaces them with typed placeholders prior to steering, and deterministically restores verified values through salience-guided rehydration after generation. DSR is evaluated across six LLaMA-family models (1B--13B parameters) on 600 A2A-generated customer-support cases (1,200 generations), with a dedicated KG ablation study. DSR significantly increases verified-entity recovery relative to a steering-only baseline (Cohen's $d=0.225$, $p_{\text{Bonf}}=1.0\times10^{-4}$), though the absolute recovery rate remains modest, while preserving effective style control across diverse model families. Layer-wise separability and steering-strength diagnostics further show previously unexplored interactions between representation-level steering and factual grounding. hese results demonstrate that explicit knowledge engineering can systematically enhance trustworthy, controllable, and reproducible generative AI without requiring model fine-tuning. Code, cached steering vectors, and evaluation scripts are publicly released to support reproducibility.\footnote{https://github.com/Tanmay-IITDSAI/KG-Gated-Defactualization}
Chinese Translation
在客户支持等对事实敏感的应用中部署的代理大型语言模型(LLMs)必须同时保持事实正确性,并在可控的风格注册中生成响应。激活引导通过扰动隐藏表示实现无需微调的风格控制,但缺乏区分可验证事实与风格内容的明确机制,导致语义泄漏。我们通过 extit{去事实化-引导-再水合}(Defactualize-Steer-Rehydrate, DSR)来应对这一挑战,这是一种将类型化、显著性加权知识图谱(KG)与激活引导相结合的知识工程框架。DSR通过分层正则表达式、命名实体识别(NER)或词汇分类器管道提取显著实体,在引导之前用类型占位符替换它们,并在生成后通过显著性引导的再水合确定性地恢复经过验证的值。DSR在600个A2A生成的客户支持案例(共1200次生成)上,对六个LLaMA家族模型(参数范围1B至13B)进行了评估,并进行了专门的KG消融研究。相较于仅使用引导的基线,DSR显著提高了经过验证的实体恢复率(Cohen's $d=0.225$, $p_{ ext{Bonf}}=1.0 imes10^{-4}$),尽管绝对恢复率仍然适中,同时在不同模型家族中保持有效的风格控制。分层可分离性和引导强度诊断进一步显示了表示级引导与事实基础之间之前未探索的相互作用。这些结果表明,明确的知识工程可以系统性地增强可信、可控和可重复的生成式人工智能,而无需模型微调。代码、缓存的引导向量和评估脚本已公开发布,以支持可重复性。
cs.CL / 33 / 2608.20396
Self-Supervised Speech Representations Track Spoken Language Convergence to Adult Models in Infants and Children Who Are Deaf/Hard-of-Hearing
自监督语音表征追踪聋哑/听力受损儿童的语言向成人模型的趋同
Abstract
Language development is characterized by a gradual convergence of children's speech toward adult patterns. Measuring this process has traditionally required detailed transcription and language-specific expertise, limiting scalability across languages and populations. Here, we use speech embeddings to capture this convergence directly from the acoustic signal in longform, child-centered recordings, taken as children go about their daily lives. Using HuBERT-BASE, we extracted embeddings from speech vocalizations of children who are deaf/hard-of-hearing and their female adult caregivers ($>$925 hrs. observation). Embedding distance between children and caregivers decreased with hearing age, controlling for pitch and vocalization length, indicating, as expected, that children's speech patterns converge to caregivers over development. This single distance metric likewise related to multiple standardized measures of speech and language from infancy through preschoolhood. These results suggest a path toward scalable, language-neutral assessment of spoken language development from children's everyday lives.
Chinese Translation
语言发展特征在于儿童的语音逐渐趋向成人模式。传统上,测量这一过程需要详细的转录和语言特定的专业知识,这限制了在不同语言和人群中的可扩展性。在此,我们使用语音嵌入直接从儿童日常生活中的长时间、以儿童为中心的录音的声学信号中捕捉这一趋同。利用 HuBERT-BASE,我们从聋哑/听力受损儿童及其女性成人照顾者的语音发声中提取了嵌入(观察时间超过 925 小时)。在控制音调和发声长度的情况下,儿童与照顾者之间的嵌入距离随着听力年龄的增长而减少,表明,正如预期的那样,儿童的语音模式在发展过程中趋向于照顾者。这一单一距离指标同样与从婴儿期到学龄前期的多项标准化语音和语言测量相关。这些结果表明了一条可扩展的、语言中立的评估儿童日常生活中口语发展的方法。
cs.CL / 34 / 2608.20402
LingShu: A Large-Scale Symptom-Centric Contextualized Knowledge Graph Bridging Traditional Chinese Medicine and Modern Biomedicine
LingShu:一个大规模以症状为中心的上下文化知识图谱,连接传统中医与现代生物医学
Hua, Rui, Shu, Zixin, Chang, Kai, Yan, Dengying, Xia, Jianan, Zhu, Hui, Song, Shujie, Yang, Shurui, Wang, Tongxin, Yin, Yue, Wei, Yu, Pei, Lijuan, Hu, Yunhui, Xu, Hao, Xiao, Mingzhong, Li, Xiaodong, Yu, Haibin, Zhang, Runshun, Wang, Wenjia, Liu, Baoyan, Zhou, Xuezhong
Abstract
Biomedical knowledge graphs (KGs) are pivotal for knowledge organization, yet traditional binary relations often struggle to represent the conditional nature of biomedical knowledge. Symptoms provide a shared phenotypic layer for linking Traditional Chinese Medicine (TCM), which relies on symptom patterns for syndrome differentiation and treatment selection, with modern biomedicine, which connects clinical manifestations to diseases and molecular mechanisms. We present LingShu, a large-scale symptom-centric contextualized knowledge graph designed to bridge TCM and modern biomedicine. The exported version of LingShu analyzed in this study comprises 17.33 million atom-level entity records and 39.47 million relation records, including 17.19 million semantic triples and 22.29 million contextualized quadruples. LingShu integrates multi-source data, including clinical electronic medical records, authoritative TCM texts, biomedical ontologies, and curated knowledge bases, through a pipeline combining natural language processing, terminology normalization, and human-in-the-loop verification. A key innovation of LingShu is its hybrid data model: it maintains 64 typed triple relation patterns to ensure broad connectivity, while incorporating 35 contextual quadruple relation patterns to capture conditional medical associations. This dual-structure approach explicitly encodes conditional knowledge, providing a granular representation of the contexts associated with medical relations. These contextualized relations cover syndrome-dependent herb efficacy, disease-contextualized drug effects, population-specific clinical associations, and mechanism-related therapeutic responses. Furthermore, we developed a web platform (http://www.tcmkg.com/) that integrates graph visualization, graph-based reasoning, and an evidence-grounded knowledge question-answering agent.
Chinese Translation
生物医学知识图谱(KGs)在知识组织中发挥着关键作用,但传统的二元关系往往难以表示生物医学知识的条件性特征。症状为连接传统中医(TCM)与现代生物医学提供了一个共享的表型层,传统中医依赖于症状模式进行辨证施治,而现代生物医学则将临床表现与疾病和分子机制联系起来。我们提出了LingShu,一个旨在连接传统中医与现代生物医学的大规模以症状为中心的上下文化知识图谱。本研究分析的LingShu导出版本包含1733万个原子级实体记录和3947万个关系记录,其中包括1719万个语义三元组和2229万个上下文化四元组。LingShu整合了多源数据,包括临床电子病历、权威中医文献、生物医学本体和策划的知识库,通过结合自然语言处理、术语标准化和人机协作验证的流程进行构建。LingShu的一个关键创新是其混合数据模型:它保持64种类型的三元关系模式,以确保广泛的连接性,同时结合35种上下文化四元关系模式,以捕捉条件医学关联。这种双结构方法明确编码了条件知识,提供了与医学关系相关的上下文的细粒度表示。这些上下文化关系涵盖了依赖症状的药材功效、与疾病相关的药物效果、特定人群的临床关联以及与机制相关的治疗反应。此外,我们开发了一个集成图谱可视化、基于图谱的推理和基于证据的知识问答代理的网络平台(http://www.tcmkg.com/)。
cs.CL / 35 / 2608.20405
ARGUS: Theory-of-Mind Guided Argument Generation with Strategy-Aware Planning and Knowledge Grounding
ARGUS:基于心智理论的论证生成与策略感知规划及知识基础
Abstract
Persuasive argument generation requires modeling audience beliefs, rhetorical strategies, and factual grounding. Despite recent advancements, existing methods remain largely audience-agnostic and fail to integrate strategy selection to improve persuasiveness. To bridge this gap, we propose Argus, an agent-based framework that operationalizes classical rhetoric for persuasive writing. At its core, a Theory-of-Mind (ToM) Reasoner constructs an explicit dual mental model of the audience's beliefs and values to guide downstream decisions. This representation conditions a component-aware planner that decomposes the argument into subtopics, assigns fine-grained rhetorical functions (logos, pathos, ethos, kairos), and triggers strategy-guided evidence retrieval at planning time. Finally, a refinement module iteratively targets and resolves multi-dimensional weaknesses without quality regression. We evaluate Argus across three diverse benchmarks using both automated pairwise Elo and LLM-as-judge metrics. Results show that Argus consistently outperforms strong baselines across multiple backbone models, achieving top rankings and the highest overall scores. Targeted simulation experiments further validate its effectiveness in shifting resistant audience stances.
Chinese Translation
说服性论证生成需要对受众信念、修辞策略和事实基础进行建模。尽管近期取得了一些进展,现有方法仍然在很大程度上与受众无关,未能整合策略选择以提高说服力。为了解决这一问题,我们提出了Argus,一个基于代理的框架,旨在将经典修辞学应用于说服性写作。其核心是一个心智理论(Theory-of-Mind, ToM)推理器,构建了受众信念和价值观的明确双重心理模型,以指导后续决策。该表示条件化了一个组件感知规划器,该规划器将论证分解为子主题,分配细粒度的修辞功能(逻辑、情感、伦理、时机),并在规划时触发策略引导的证据检索。最后,一个精细化模块迭代地针对并解决多维度的弱点,而不导致质量退化。我们在三个不同的基准上评估了Argus,使用自动化的成对Elo和LLM作为评判标准。结果表明,Argus在多个基础模型上始终优于强基线,获得了最高排名和整体最高分数。针对性的模拟实验进一步验证了其在改变抵抗性受众立场方面的有效性。
cs.CL / 36 / 2608.20530
LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decoding
LiLiCorr:轻量级基于似然的并行草稿相关性模型用于推测解码
Abstract
Speculative decoding accelerates language-model inference by drafting future tokens that the target model verifies in parallel. A diffusion-style block head such as DFlash is an attractive drafter, predicting an entire block of future tokens in one forward pass. However, it is trained on per-position marginals rather than the joint block distribution, so the tokens it emits are individually plausible yet jointly incoherent. We introduce LiLiCorr, a Lightweight Likelihood-based model that Correlates the per-position marginal distributions a drafter already produces. It keeps the top-k tokens at each position as candidates and processes them jointly, producing for each an in and an out vector. A pair of adjacent candidates matches when the earlier one's out vector has high cosine similarity with the later one's in vector. These matches capture the block's joint structure without ever materializing the full joint distribution. One lightweight network pass produces all the vectors, and the pairwise scores are then computed in parallel as batched matrix operations, leaving only a cheap greedy walk sequential. We further co-train the drafter with LiLiCorr, so it learns to propose candidates that correlate into longer accepted sequences. Over the vanilla DFlash drafter, LiLiCorr raises acceptance length on every benchmark by 9 to 19%, while its scoring head accounts for about 2.8% of the per-block latency. Against DFlash and two concurrent methods that also restore coherence at draft time, LiLiCorr delivers the highest throughput in 70 of 72 settings: nine benchmarks at two target sizes under greedy and temperature-one decoding, and a throughput sweep over six concurrencies, two input lengths and three entropy tiers, with all systems equally optimized on a common serving stack. Extending LiLiCorr to inputs an order of magnitude longer than it was trained on preserves that lead.
Chinese Translation
推测解码通过并行验证目标模型草拟的未来标记来加速语言模型推理。扩散式块头(如 DFlash)是一种吸引人的草拟器,它在一次前向传播中预测整个未来标记块。然而,它是基于每个位置的边际分布进行训练,而不是联合块分布,因此它发出的标记在个体上是合理的,但在整体上却不一致。我们引入了 LiLiCorr,这是一种轻量级的基于似然的模型,能够关联草拟器已经生成的每个位置的边际分布。它在每个位置保留前 k 个标记作为候选,并将它们联合处理,为每个候选生成一个输入向量和一个输出向量。当前一个候选的输出向量与后一个候选的输入向量具有高余弦相似度时,这对相邻候选匹配。这些匹配捕捉了块的联合结构,而无需实际生成完整的联合分布。一次轻量级网络传递生成所有向量,然后成对得分在并行中计算,留下一个便宜的贪婪遍历序列。我们进一步与 LiLiCorr 共同训练草拟器,使其学习提出能够关联成更长接受序列的候选。在原始 DFlash 草拟器的基础上,LiLiCorr 在每个基准测试中将接受长度提高了 9% 到 19%,而其评分头占每块延迟的约 2.8%。与 DFlash 和另外两种在草拟时也恢复一致性的方法相比,LiLiCorr 在 72 种设置中的 70 种中提供了最高的吞吐量:在贪婪和温度为 1 的解码下,针对两个目标大小的九个基准测试,以及在六种并发、两种输入长度和三种熵层次下的吞吐量测试,所有系统在共同的服务堆栈上进行了均等优化。将 LiLiCorr 扩展到比训练时长一个数量级的输入仍然保持了这一优势。
cs.CL / 37 / 2608.20607
JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification
JuryProbe:一种实证共识风险诊断方法,用于将无参考事实判断小组路由到基础验证
Abstract
Panels of inexpensive LLM judges increasingly make accept-or-escalate decisions. In factuality settings, accepting a claim because several reference-free judges agree can create a hidden risk: agreement may reflect shared false-negative blind spots rather than independent evidence. We introduce JuryProbe, an empirical consensus-risk diagnostic for reference-free factuality judge panels, paired with a calibration-based routing policy. JuryProbe estimates consensus risk from a labeled calibration probe using false-negative-only (FN-only) judge correlation and false-consensus lift; when flagged high-risk, reference-free majority accepts are routed to the same judges with trusted references. On audited FEVER corruptions, reference-free panels show correlated false negatives (FN-only correlations 0.402 and 0.368; lifts 3.13x and 18.13x), while unanimous false consensus drops to zero under a trusted-reference best-case diagnostic on both minimal-pair and non-minimal-pair evidence. In flagged settings, the routed policy is by construction equivalent to grounding every reference-free majority accept (verified in 34/34 splits): improvement comes from accept-conditioned grounding, while the diagnostic determines whether to activate it. A fixed, pre-specified rule flags 8-10 of 10 splits across synthetic, benchmark-authored, and scientific families and 0 of 10 on a negative control, where standing down avoids 28% of reference acquisitions at a 0.004 increase in false accepts. False-accept reduction persists under weak BM25 retrieval at substantial coverage cost, while stale stand-down labels require periodic recalibration. JuryProbe provides no formal risk guarantee and does not establish reliable stand-down on natural panels; its supported contribution is an empirical diagnostic of high-risk panel error dependence.
Chinese Translation
廉价的大型语言模型(LLM)评审小组越来越多地做出接受或升级的决策。在事实性设置中,因多个无参考评审者达成一致而接受某一主张可能会产生隐性风险:这种一致性可能反映的是共享的假阴性盲点,而非独立证据。我们提出了JuryProbe,一种针对无参考事实判断小组的实证共识风险诊断方法,并配备了一种基于校准的路由策略。JuryProbe通过使用假阴性相关性(FN-only judge correlation)和假共识提升(false-consensus lift)从标记的校准探针中估计共识风险;当标记为高风险时,无参考的多数接受将被路由到具有可信参考的相同评审者。在经过审计的FEVER腐败案例中,无参考小组显示出相关的假阴性(假阴性相关性分别为0.402和0.368;提升分别为3.13倍和18.13倍),而在可信参考的最佳案例诊断下,一致的假共识在最小对和非最小对证据下降至零。在标记的设置中,路由策略在构造上等同于对每个无参考多数接受进行基础验证(在34/34的分割中得到验证):改进来自于接受条件下的基础验证,而诊断则决定是否激活该过程。一个固定的预设规则在合成、基准作者和科学类别中标记10个分割中的8-10个,而在负控制中标记0个,其中不采取行动避免了28%的参考获取,同时假接受增加了0.004。假接受的减少在弱BM25检索下持续存在,但过时的停用标签需要定期重新校准。JuryProbe并未提供正式的风险保证,也未在自然小组中建立可靠的停用;其支持的贡献是对高风险小组错误依赖的实证诊断。
cs.CL / 38 / 2608.20627
When Failures Propagate: Causal Failure Attribution in Agentic Retrieval-Augmented Generation
故障传播时:代理检索增强生成中的因果故障归因
Abstract
Agentic retrieval-augmented generation (RAG) interleaves retrieval, reasoning, and answer generation across multiple hops. A retrieval error at hop 1 can surface only as a wrong answer at hop 3, while later retrieval can also repair the trajectory. This paper introduces AgenticRAG-FP, an interventional benchmark for causal failure attribution in agentic RAG. The benchmark injects a certified fault at a specified hop, re-executes the downstream trajectory, and evaluates diagnosers against the known intervention. Its central question is whether a post-hoc trace still identifies the injected hop after the suffix changes. In the completed strict dense Claude Haiku 4.5 sweep on 80 three-hop MuSiQue questions, coverage-based diagnosis is 0.91 at hop 1 and 0.00 at hops 2 and 3 (n=43,36,21 failed trajectories). A smaller content-corruption study changes an answer-bearing or bridge fact in topically intact evidence. At depth 2, where 18 failed cases remain after filtering, coverage-based diagnosis is 0.00 and a frozen-hop counterfactual probe is 0.67 in an exploratory pooled comparison. Depth-3 content estimates are descriptive only because they contain three failed cases. These results make propagation depth an explicit evaluation axis for diagnosing agentic RAG failures while distinguishing broad evidence of post-hoc signal loss from small-sample method comparisons.
Chinese Translation
代理检索增强生成(RAG)在多个步骤中交错进行检索、推理和答案生成。在第一步的检索错误可能仅在第三步表现为错误答案,而后续的检索也可以修复这一轨迹。本文介绍了AgenticRAG-FP,一个用于代理RAG中因果故障归因的干预基准。该基准在指定步骤注入一个认证故障,重新执行下游轨迹,并根据已知的干预评估诊断器。其核心问题是,后验追踪是否仍能在后缀变化后识别出注入的步骤。在对80个三步MuSiQue问题进行的严格密集Claude Haiku 4.5评估中,基于覆盖的诊断在第一步为0.91,而在第二步和第三步为0.00(n=43,36,21个失败轨迹)。一个较小的内容损坏研究在主题完整的证据中更改了一个承载答案或桥接事实。在深度2中,经过过滤后仍有18个失败案例,基于覆盖的诊断为0.00,而在探索性汇总比较中,冻结步骤的反事实探测为0.67。深度3的内容估计仅为描述性,因为它们包含三个失败案例。这些结果使得传播深度成为诊断代理RAG故障的一个明确评估轴,同时区分了后验信号损失的广泛证据与小样本方法比较。
cs.CL / 39 / 2608.20632
Sparse Token Routing in Efficient Transformers
高效变换器中的稀疏令牌路由
Abstract
Efficient-transformer research often motivates token pruning and adaptive computation with the claim that not all tokens require equal computational effort. We test this claim end to end using SEWN, a two-stream Transformer that routes tokens through either lightweight or full-capacity processing using a learned gate. Across our experiments, routing introduces negligible accuracy change relative to parameter-matched baselines, while the gate's token-importance signal depends critically on how it is learned. A static lexicon-seeded prior fails a counterfactual faithfulness test on BoolQ, whereas a fully contextual gate achieves highly significant separation ($p<10^{-10}$) on both evaluated tasks without changing task accuracy.
Chinese Translation
高效变换器的研究通常通过声称并非所有令牌都需要相同的计算努力来激励令牌修剪和自适应计算。我们使用 SEWN(一种双流变换器)对这一主张进行端到端的测试,该变换器通过学习的门将令牌路由到轻量级或全容量处理。在我们的实验中,相较于参数匹配的基线,路由引入的准确性变化微乎其微,而门的令牌重要性信号在很大程度上依赖于其学习方式。一个静态的词汇种子先验在 BoolQ 上未能通过反事实忠实性测试,而一个完全上下文的门在评估的两个任务上实现了高度显著的分离($p<10^{-10}$),且没有改变任务的准确性。
cs.CL / 40 / 2608.20634
AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale
AgentMercury:您的智能体可以大规模合成可验证的商业场景环境
Abstract
Agents learn to act through interaction with environments, yet the environments used for training are often manually constructed or synthesized around predefined tasks and benchmarks. This task-centric paradigm makes it difficult to scale environments that reflect realistic and evolving workflows where diverse tasks can naturally emerge from the underlying world. We introduce AgentMercury, a scalable framework for synthesizing executable environments from high-level business scenarios. Rather than constructing an environment for a specific task, AgentMercury first instantiates a persistent world with entities, services, tools, state, and executable cross-service invariants, from which diverse tasks and interaction trajectories can subsequently emerge. We construct 4,783 executable environments spanning 14 industries and 50 countries, and use them as training substrates for reinforcement learning. Despite being generated without targeting the evaluation benchmarks, policies trained on these business-oriented environments improve substantially on both enterprise workflows and out-of-domain benchmarks spanning reasoning, coding, scientific computing, and tool use. In our experiments, Qwen3.5-4B improves from 12.3 to 15.7 on EnterpriseOps-GYM and from 45.9 to 56.0 on AIME26 after training on AgentMercury environments. We further show that the construction process itself can be learned: fine-tuning Qwen3.5-35B-A3B on construction traces increases executable-world authoring success from 3.3% to 83.3% on held-out business scenarios. These results show that scenario-grounded environments can provide useful and generalizable learning signals beyond benchmark-specific training, while their construction can itself become a learnable capability.
Chinese Translation
智能体通过与环境的交互学习行动,但用于训练的环境通常是手动构建或围绕预定义任务和基准合成的。这种以任务为中心的范式使得难以扩展反映现实和不断演变的工作流的环境,其中多样化的任务可以自然地从基础世界中出现。我们提出了AgentMercury,一个从高层次商业场景合成可执行环境的可扩展框架。AgentMercury首先实例化一个持久的世界,包含实体、服务、工具、状态和可执行的跨服务不变性,而不是为特定任务构建环境,从中可以随后出现多样化的任务和交互轨迹。我们构建了4783个可执行环境,涵盖14个行业和50个国家,并将其用作强化学习的训练基础。尽管这些环境是在没有针对评估基准的情况下生成的,但在这些以商业为导向的环境上训练的策略在企业工作流和跨领域基准(包括推理、编码、科学计算和工具使用)上都有显著改善。在我们的实验中,Qwen3.5-4B在AgentMercury环境上训练后,EnterpriseOps-GYM的得分从12.3提高到15.7,AIME26的得分从45.9提高到56.0。我们进一步展示了构建过程本身可以被学习:在构建轨迹上微调Qwen3.5-35B-A3B使得在保留的商业场景上可执行世界创作的成功率从3.3%提高到83.3%。这些结果表明,基于场景的环境可以提供有用且可推广的学习信号,超越基准特定的训练,同时它们的构建本身可以成为一种可学习的能力。
cs.CL / 41 / 2608.20636
MIL-BERT: Classification of Arbitrarily Large Text with Performance and Explanatory Guarantees
MIL-BERT:具有性能和解释保证的任意大文本分类
Abstract
Many text classification decisions are viable based on constituent excerpts alone. Taking inspiration from the field of multiple instance learning, we present an algorithm for training a neural network to classify text by selecting such excerpts. We show that our approach is also scalable with demonstrated learning against samples with nearly 1M tokens. We evaluate our methods on 7 datasets with emphasis on long-textual collections that far exceed the encoding limit of our base model. We present state-of-the-art results with this algorithm on 3 datasets: identification of political bias in news outlets, trigger warnings in long stories, and demographic characteristics of authors in tweet collections. Furthermore, the model trained on weakly-labeled collections of text (bags) generalizes to accurately classify constituent, smaller instances. Besides a new state-of-the-art for these problems, this approach is one of the few neural methods to excel in these datasets.
Chinese Translation
许多文本分类决策仅基于组成摘录就可以实现。受到多实例学习领域的启发,我们提出了一种算法,用于训练神经网络通过选择这些摘录来对文本进行分类。我们展示了我们的方法在处理近100万个标记的样本时也具有可扩展性。我们在7个数据集上评估了我们的方法,重点关注远超我们基础模型编码限制的长文本集合。我们在3个数据集上使用该算法取得了最先进的结果:识别新闻媒体中的政治偏见、长篇故事中的触发警告,以及推文集合中作者的人口特征。此外,在弱标记文本集合(袋)上训练的模型能够很好地推广,准确分类组成的较小实例。除了在这些问题上取得新的最先进成果外,这种方法是为数不多的在这些数据集上表现出色的神经方法之一。
cs.CL / 42 / 2608.20647
Directional Contextual Representations for Dependency Relations: Why Cross-Direction Pairing Fails
依赖关系的定向上下文表示:为何交叉方向配对失败
Abstract
Splitting a bidirectional LSTM's contextual representation into a forward-only $F_i$ (strictly a function of tokens $1..i$) and a backward-only $B_i$ (strictly a function of tokens $i..n$) beats either alone and beats a fused self-attention representation for dependency relation-type classification. But a specific, natural extension of this idea -- pairing a token's forward state against a \emph{candidate}'s backward state (``cross-direction'' pairing, $F_i$ vs.\ $B_j$) -- consistently \emph{underperforms} same-direction pairing, and the penalty \emph{grows}, not shrinks, with token distance, both paired-bootstrap significant. We diagnose why using a frozen-trunk methodology: architectural information leakage between directions is impossible by construction (a single-layer BiLSTM, verified by code inspection); 93\% of the same-vs-cross gap survives freezing the trunk and training only fresh heads, ruling out training-co-adaptation as the primary cause; linear regression shows partial representational redundancy between $F_i$ and $B_i$ ($R^2{=}0.324$ vs.\ $0.028$ for a shuffled control) and a linear probe shows partial anticipatory encoding of upcoming tokens in $F_i$ (36.5\% vs.\ 17.2\% majority baseline) -- real effects, but neither alone, nor combined, cleanly explains the full gap. Extended frozen-trunk diagnostics (a positional probe and a distance-decay probe) show directional information is genuinely stored but not exactly positioned, and propagates only a few tokens before decaying to baseline -- consistent with, and mechanistically underneath, the distance-growth finding.
Chinese Translation
将双向 LSTM 的上下文表示分割为仅向前的 $F_i$(严格是 tokens $1..i$ 的函数)和仅向后的 $B_i$(严格是 tokens $i..n$ 的函数)在依赖关系类型分类中优于单独使用任一者,并且优于融合的自注意力表示。然而,这一思想的一个特定自然扩展——将一个 token 的向前状态与一个 extit{候选} 的向后状态配对(“交叉方向”配对,$F_i$ 对 $B_j$)始终 extit{表现不佳},而且这种惩罚随着 token 距离的增加而 extit{加重},而不是减轻,且配对引导的显著性。我们通过冷冻主干方法诊断原因:由于结构上的设计,方向之间的架构信息泄漏是不可能的(单层 BiLSTM,通过代码检查验证);93 ext{%} 的同向与交叉之间的差距在冷冻主干并仅训练新头部时依然存在,排除了训练共同适应作为主要原因;线性回归显示 $F_i$ 和 $B_i$ 之间存在部分表示冗余($R^2{=}0.324$ 对比 $0.028$ 的随机控制),而线性探测显示 $F_i$ 中对即将到来的 tokens 存在部分预期编码(36.5 ext{%} 对比 17.2 ext{%} 的多数基线)——这些都是实际效应,但单独或结合使用都无法清晰解释完整的差距。扩展的冷冻主干诊断(位置探测和距离衰减探测)显示方向信息确实被存储但未准确定位,并且仅在传播几个 token 后衰减至基线——这与距离增长的发现一致,并在机制上支持这一发现。
cs.CL / 43 / 2608.20711
AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification
AsmEvo:具有功能等价验证的 AMD GPU 内核的自主汇编级优化
Liu, Ji, Yang, Puyuan, Zheng, Rongzhang, Wang, Fan, Wang, Jinglin, Awad, Muhammad A., Huang, Mortis, Chang, Andy, Li, Zekai, Li, Zeping, An, Zihao, Liu, Yue, Yang, Yuchen, Wang, Jianghui, Chen, Chushi, Liu, Ziqiong, Yang, Fuwei, Li, Dong, Chung, Wen Heng, Liu, Shengcai, Barsoum, Emad
Abstract
High-performance ML systems increasingly rely on GPU kernels whose editable source is unavailable, generated, or too distant from final machine code to expose remaining optimizations. Existing LLM kernel optimizers and autotuners mainly operate on CUDA, Triton, HIP, or tensor-program source and validate against reference implementations. We study a stricter setting: optimizing an already compiled AMDGPU code object, where the deployed binary is the only behavioral oracle. We present AsmEvo, an agentic assembly-level optimizer for AMD GPU kernels. Given an AMDGPU code object K0, AsmEvo reconstructs a reassemblable representation, proposes low-level edits with a long-horizon agent, rebuilds an ABI-preserving optimized object, and accepts candidates only after differential verification against K0 under identical launches. AsmEvo combines code-object recovery, metadata-aware rebuilding, profiling-guided hot-window editing, correctness-gated timing, and conservative in-place patch fallback. We conduct extensive experiments with AsmEvo on various AMD GPU kernels. On MI308X, AsmEvo improves 29 of 30 selected KernelBench kernels, reaching 1.35x geometric-mean and 3.88x maximum speedup. On MI300X production workloads, it improves all evaluated AITer binaries and vLLM/SGLang Triton assembly kernels, reaching 1.09x/1.31x and 1.18x/1.34x geometric-mean/maximum speedups, respectively, while preserving functional equivalence.
Chinese Translation
高性能机器学习系统日益依赖于其可编辑源代码不可用、生成或与最终机器代码相距甚远的 GPU 内核,从而无法揭示剩余的优化。现有的 LLM 内核优化器和自动调优器主要针对 CUDA、Triton、HIP 或张量程序源代码进行操作,并与参考实现进行验证。我们研究了一个更严格的设置:优化已经编译的 AMDGPU 代码对象,其中部署的二进制文件是唯一的行为 oracle。我们提出了 AsmEvo,一种用于 AMD GPU 内核的自主汇编级优化器。给定一个 AMDGPU 代码对象 K0,AsmEvo 重建一个可重新组装的表示,使用长远代理提出低级编辑,重建一个保持 ABI 的优化对象,并在与 K0 在相同启动条件下进行差异验证后才接受候选项。AsmEvo 结合了代码对象恢复、元数据感知重建、基于性能分析的热点窗口编辑、正确性门控时序和保守的就地补丁回退。我们在各种 AMD GPU 内核上对 AsmEvo 进行了广泛的实验。在 MI308X 上,AsmEvo 改进了 30 个选定 KernelBench 内核中的 29 个,达到了 1.35 倍的几何平均加速和 3.88 倍的最大加速。在 MI300X 生产工作负载上,它改进了所有评估的 AITer 二进制文件和 vLLM/SGLang Triton 汇编内核,分别达到了 1.09 倍/1.31 倍和 1.18 倍/1.34 倍的几何平均/最大加速,同时保持功能等价性。
cs.CL / 44 / 2608.20757
PSK at WMT 2026 MIST: Task-Specialized QLoRA Adapters for Multilingual Summarization and Question Answering
WMT 2026 MIST中的PSK:用于多语言摘要和问答的任务专用QLoRA适配器
Abstract
We describe the PSK submission to the WMT 2026 Multilingual Instruction Shared Task. Our system uses the 3.35B-parameter Tiny Aya Global model with three QLoRA adapters, one for each task. The adapters are trained on multilingual document-summary pairs, passage-based question answering, and filtered standalone question answering. The summarization data also includes scientific papers with their author-written abstracts. On our held-out split, the context and summarization adapters perform better than our multitask adapter, which was trained only on data supplied by the organizers. Results for open QA are mixed and vary with answer length and evaluation method. We therefore submit three systems with the same context and summarization adapters but different open-QA adapters.
Chinese Translation
我们描述了PSK对WMT 2026多语言指令共享任务的提交。我们的系统使用了3.35B参数的Tiny Aya Global模型,并配备了三个QLoRA适配器,每个任务一个。适配器在多语言文档-摘要对、基于段落的问答以及过滤后的独立问答上进行了训练。摘要数据还包括科学论文及其作者撰写的摘要。在我们的保留数据集中,上下文和摘要适配器的表现优于仅在组织者提供的数据上训练的多任务适配器。开放问答的结果则较为复杂,且随着答案长度和评估方法的不同而有所变化。因此,我们提交了三个系统,使用相同的上下文和摘要适配器,但配备不同的开放问答适配器。
cs.CL / 45 / 2608.20777
Tree-of-Concerns: Hierarchical Multi-Agent Debate for Unstated-Limitation Extraction in Scientific Critique
关注树:用于科学批评中未陈述限制提取的层次多智能体辩论
Abstract
As scientific literature grows and papers increasingly under-report limitations, multi-agent LLMs offer a promising approach to systematically uncover these hidden failure modes. Here, we introduce Tree-of-Concerns, a multi-agent framework that deploys specialized skeptic personas, each operating through a category-specific analytical lens, as parallel debate trees to extract unstated limitations from scientific papers. Each persona conducts structured, evidence-grounded argumentation, while a Panel Review mechanism re-evaluates each surviving claim from all five perspectives to correct category drift and severity miscalibration. Through experiments on ToC-Bench, our benchmark of 414 research papers with 1,905 unstated limitations, sourced from reviewer-reported weaknesses and follow-up citation critiques, we demonstrate that ToC improves precision by 79% and coverage by 11% relative to strongest baselines, surfacing specific, evidence-grounded concerns that support reviewers in systematic evaluation.
Chinese Translation
随着科学文献的增长,论文在限制方面的报告越来越不足,多智能体大语言模型(LLMs)提供了一种有前景的方法来系统地揭示这些隐藏的失败模式。在此,我们介绍了关注树(Tree-of-Concerns),这是一个多智能体框架,部署了专门的怀疑者角色,每个角色通过特定类别的分析视角作为并行辩论树,提取科学论文中的未陈述限制。每个角色进行结构化的、基于证据的论证,而面板审查机制则从所有五个视角重新评估每个存活的主张,以纠正类别漂移和严重性误校准。通过在ToC-Bench上的实验,我们的基准包含414篇研究论文和1905个未陈述的限制,这些限制来源于审稿人报告的弱点和后续引用批评,我们证明了关注树相较于最强基线提高了79%的精确度和11%的覆盖率,揭示了具体的、基于证据的关注点,支持审稿人在系统评估中的工作。
cs.CL / 46 / 2608.20804
Denoising the Future: Context-Aware Spectral Diffusion for Temporal Knowledge Graph Extrapolation
去噪未来:基于上下文的光谱扩散用于时间知识图谱外推
Abstract
Temporal Knowledge Graph (TKG) extrapolation seeks to infer future facts from time-varying relational histories. Recent diffusion-based approaches improve uncertainty modeling through generative denoising, but their aggregated conditioning on subject histories may insufficiently distinguish query-specific evidence from non-salient historical facts, thereby diluting target-discriminative signals. To bridge this gap, we propose FreqDiff, a Frequency-aware Diffusion framework for TKG extrapolation. Specifically, FreqDiff formulates future object prediction as query-slot denoising and develops a dual-stream denoiser that integrates temporal dependency modeling with context-aware spectral calibration. The spectral branch synthesizes history-conditioned filters from learnable bases to adaptively re-calibrate denoising representations, while a frequency-domain regularizer is proposed to align the denoised target with the gold object in spectral space. Experiments on four public TKG benchmarks demonstrate that FreqDiff achieves state-of-the-art performance.
Chinese Translation
时间知识图谱(Temporal Knowledge Graph, TKG)外推旨在从时间变化的关系历史中推断未来事实。最近的基于扩散的方法通过生成去噪改善了不确定性建模,但它们对主体历史的聚合条件可能不足以区分查询特定证据与非显著历史事实,从而稀释了目标区分信号。为了解决这一问题,我们提出了FreqDiff,一个针对TKG外推的频率感知扩散框架。具体而言,FreqDiff将未来对象预测公式化为查询槽去噪,并开发了一种双流去噪器,将时间依赖建模与上下文感知光谱校准相结合。光谱分支从可学习基底合成历史条件过滤器,以自适应地重新校准去噪表示,同时提出了一种频域正则化器,以在光谱空间中将去噪目标与真实对象对齐。在四个公共TKG基准上的实验表明,FreqDiff达到了最先进的性能。
cs.CL / 47 / 2608.20831
STAR-OPD: Structured Aspect-Cascade-Aware On-Policy Reward Distillation for ABSA Quadruple Extraction
STAR-OPD:结构化方面级联感知的在线奖励蒸馏用于ABSA四元组提取
Abstract
Aspect-based sentiment analysis (ABSA) quadruple extraction requires jointly predicting target, aspect, opinion, and sentiment over reviews that often contain multiple fine-grained sentiment tuples. While large chain-of-thought (CoT) models perform well on this task, distilling them into smaller deployable models remains difficult. We identify a task-specific failure mode in distilled ABSA extraction: student errors at the target-aspect interface create structurally invalid states, such as broken target-aspect bindings and hallucinated targets, which then corrupt downstream predictions. Conventional off-policy distillation is poorly suited to this setting because it trains only on teacher-generated trajectories and provides little supervision on the student-induced structural states that dominate inference. To address this mismatch, we propose STAR-OPD (STructured Aspect-cascade-aware On-Policy Reward Distillation), which builds on generic on-policy distillation and instantiates it for ABSA quadruple extraction with cascade-aware, set-structured rewards. STAR-OPD trains on student rollouts and applies set-structured rewards that directly target binding consistency, target grounding, and fine-grained aspect disambiguation. Experiments on E-ABSA20K and SemEval-2014 show that STAR-OPD consistently outperforms off-policy and general on-policy baselines, reduces target hallucination, and substantially improves performance on structurally hard cases. With Qwen3-4B, STAR-OPD substantially narrows the student-teacher gap while improving inference efficiency, highlighting the importance of on-policy structural correction for distilled ABSA extraction.
Chinese Translation
基于方面的情感分析(ABSA)四元组提取需要对评论中通常包含多个细粒度情感元组的目标、方面、意见和情感进行联合预测。尽管大型思维链(CoT)模型在此任务上表现良好,但将其蒸馏为更小的可部署模型仍然困难。我们识别出蒸馏ABSA提取中的任务特定失败模式:学生在目标-方面接口的错误会导致结构上无效的状态,例如破损的目标-方面绑定和虚构的目标,这会破坏下游预测。传统的离线蒸馏不适合这种情况,因为它仅在教师生成的轨迹上进行训练,并且对主导推理的学生诱导结构状态提供的监督很少。为了解决这一不匹配,我们提出了STAR-OPD(结构化方面级联感知的在线奖励蒸馏),该方法基于通用的在线蒸馏并为ABSA四元组提取实例化,采用级联感知的集合结构奖励。STAR-OPD在学生回合上进行训练,并应用直接针对绑定一致性、目标定位和细粒度方面消歧的集合结构奖励。在E-ABSA20K和SemEval-2014上的实验表明,STAR-OPD始终优于离线和通用在线基线,减少了目标虚构,并显著改善了结构上困难案例的表现。使用Qwen3-4B,STAR-OPD显著缩小了学生与教师之间的差距,同时提高了推理效率,突显了在线结构修正对蒸馏ABSA提取的重要性。
cs.CL / 48 / 2608.20839
SAC-Copula: Quality-Preserving Watermarking for Diffusion Language Models via Smooth Correlated Gumbel Fields
SAC-Copula:通过平滑相关的Gumbel场实现的保质水印技术用于扩散语言模型
Abstract
Watermarking diffusion language models (DLMs) requires mechanisms compatible with iterative parallel unmasking rather than autoregressive decoding. Existing sampling-based watermarking methods typically inject position-wise i.i.d. perturbations, which can be poorly aligned with DLM decoding dynamics and degrade generation quality. We propose SAC-Copula, a quality-preserving watermarking method for DLMs based on smooth, locally correlated Gumbel perturbation fields constructed via a Gaussian copula. We further develop a SAC-aware detector using covariance-aware filtering and native-sample calibration. Mechanism-level analysis shows that local correlation reduces latent perturbation roughness and better matches iterative refinement dynamics. Experiments on LLaDA show that SAC-Copula achieves a favorable quality-detectability trade-off compared with existing baselines. In particular, further evaluations on Dream-7B and additional datasets show that SAC-Copula substantially improves PPL tail stability over the i.i.d. Gumbel baseline, while maintaining strong low-FPR detectability and competitive overall generation quality. Additional token-edit stress tests further assess watermark robustness under controlled synchronization drift.
Chinese Translation
对扩散语言模型(DLMs)进行水印处理需要与迭代并行去掩蔽兼容的机制,而非自回归解码。现有的基于采样的水印方法通常注入位置独立的独立同分布(i.i.d.)扰动,这可能与DLM的解码动态不匹配,从而降低生成质量。我们提出了SAC-Copula,这是一种基于通过高斯copula构建的平滑、局部相关的Gumbel扰动场的保质水印方法。我们进一步开发了一种SAC感知检测器,采用协方差感知过滤和原生样本校准。机制级分析表明,局部相关性降低了潜在扰动的粗糙度,更好地匹配了迭代精炼动态。在LLaDA上的实验表明,SAC-Copula在质量与可检测性之间实现了良好的平衡,相较于现有基线表现尤为突出。特别是在Dream-7B及其他数据集上的进一步评估显示,SAC-Copula在保持强低假阳性率(low-FPR)可检测性的同时,显著提高了PPL尾部稳定性,优于i.i.d. Gumbel基线。此外,额外的标记编辑压力测试进一步评估了水印在受控同步漂移下的鲁棒性。
cs.CL / 49 / 2608.20856
Ontology-Driven Structural Regularization for Document-Level Relation Extraction
基于本体驱动的文档级关系抽取结构正则化
Abstract
Document-Level Relation Extraction (DocRE) relies heavily on costly manually annotated datasets, while large distant supervision resources such as DocRED distant remain underexploited due to noise. We show that a critical yet overlooked source of noise lies in structural inconsistencies within relational triples, including violations of ontology constraints and logical contradictions. We introduce an ontology-driven framework to quantify and enforce structural consistency in DocRE datasets. Our analysis reveals substantial structural noise in DocRED distant and demonstrates that such inconsistencies propagate to model predictions. Enforcing structural well-formedness during training significantly reduces logical contradictions and consistently improves generalization performance. These findings establish structural consistency as a missing axis of supervision in DocRE and highlight structural regularization as an effective strategy for leveraging distant data at scale.
Chinese Translation
文档级关系抽取(DocRE)在很大程度上依赖于昂贵的人工标注数据集,而像DocRED远程监督这样的丰富资源由于噪声问题仍未得到充分利用。我们指出,一个重要但被忽视的噪声来源存在于关系三元组的结构不一致性中,包括对本体约束的违反和逻辑矛盾。我们提出了一种基于本体驱动的框架,用于量化和强制DocRE数据集中的结构一致性。我们的分析揭示了DocRED远程监督中存在显著的结构噪声,并证明这种不一致性会传播到模型预测中。在训练过程中强制结构良构性显著减少了逻辑矛盾,并持续改善了泛化性能。这些发现确立了结构一致性作为DocRE中缺失的监督维度,并强调了结构正则化作为有效策略,以大规模利用远程数据。
cs.CL / 50 / 2608.20887
KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs
KREL:基于知识引导的临床证据推理的自动医学编码
Abstract
Automatic Medical Coding (AMC), which assigns standardized International Classification of Diseases (ICD) codes to clinical notes, is essential for medical reimbursement, quality reporting, and clinical research. Existing pre-trained language model (PLM)-based methods typically formulate AMC as an extreme multi-label classification problem over a predefined code set, while recent large language model (LLM)-based approaches instead frame it as generation or multi-step reasoning. However, key challenges remain, including the extreme length of clinical notes that hinders effective interpretation, the vast ICD label space, and complex coding rules that are not explicitly captured by LLMs. In this work, we propose Knowledge-Guided Reasoning over Clinical Evidence with LLMs (KREL), a framework that leverages LLMs for clinical text understanding and reasoning while integrating external ICD coding guidelines as structured knowledge. This design enables tight coupling between domain knowledge and LLM reasoning, reducing hallucinations and improving compliance with coding standards. Experiments on benchmark datasets show that KREL consistently outperforms strong PLM-based and state-of-the-art LLM-based baselines.
Chinese Translation
自动医学编码(AMC)是将标准化的国际疾病分类(ICD)代码分配给临床笔记的过程,对于医疗报销、质量报告和临床研究至关重要。现有的基于预训练语言模型(PLM)的方法通常将AMC视为一个在预定义代码集上的极端多标签分类问题,而最近的基于大型语言模型(LLM)的方法则将其框架设定为生成或多步骤推理。然而,仍然存在一些关键挑战,包括临床笔记的极长文本限制了有效解读、庞大的ICD标签空间以及LLM未明确捕捉的复杂编码规则。在本研究中,我们提出了基于知识引导的临床证据推理框架(KREL),该框架利用LLM进行临床文本理解和推理,同时将外部ICD编码指南作为结构化知识进行整合。这一设计实现了领域知识与LLM推理之间的紧密结合,减少了幻觉现象,提高了对编码标准的遵从性。在基准数据集上的实验表明,KREL始终优于强大的基于PLM的方法和最先进的基于LLM的方法。
cs.CL / 51 / 2608.20920
ForeDreamer: A Self-Evolving Dual-Agent Memory Architecture for Future Event Prediction
ForeDreamer:一种自我进化的双代理记忆架构用于未来事件预测
Abstract
Open-web future event prediction requires agents to distill reliable signals from noisy, redundant, and incomplete evidence. Existing retrieval/memory mechanisms directly feed retrieved information to agents or rely on simple memory functions such as storing and reusing prior information for prediction, leaving them insufficient for open-web forecasting. We propose to transform raw web evidence into structured memory before prediction, enabling agents to reason over distilled, question-specific evidence rather than noisy retrieval results. This paper presents ForeDreamer, a self-evolving dual-agent framework for managing memory over open-web evidence. ForeDreamer separates factual memory, a question-specific evidence state for the current forecast, from experiential memory, persistent agent experience accumulated across forecasting episodes. It uses a main agent for search and prediction, and a memory-processing subagent to convert search results into factual memory with dedicated tools. ForeDreamer further evolves experiential memory through two tracks, improving both forecasting decisions and factual-memory construction. Experiments on Prophet Arena and FutureX demonstrate the effectiveness of ForeDreamer. Project page: https://zhongzero.github.io/ForeDreamer
Chinese Translation
开放网络的未来事件预测要求代理从嘈杂、冗余和不完整的证据中提取可靠信号。现有的检索/记忆机制直接将检索到的信息提供给代理,或依赖于简单的记忆功能,如存储和重用先前信息进行预测,这使得它们在开放网络预测中显得不足。我们提出在预测之前将原始网络证据转化为结构化记忆,使代理能够基于提炼出的、特定问题的证据进行推理,而不是依赖嘈杂的检索结果。本文提出了ForeDreamer,一种自我进化的双代理框架,用于管理开放网络证据的记忆。ForeDreamer将事实记忆(针对当前预测的特定问题的证据状态)与经验记忆(在预测过程中积累的持久代理经验)分开。它使用一个主代理进行搜索和预测,以及一个记忆处理子代理将搜索结果转换为事实记忆,配备专用工具。ForeDreamer还通过两个轨道进一步进化经验记忆,改善预测决策和事实记忆的构建。在Prophet Arena和FutureX上的实验证明了ForeDreamer的有效性。项目页面:https://zhongzero.github.io/ForeDreamer
cs.CL / 52 / 2608.20925
Source-Free MT Evaluation Is Not MT Evaluation
无源机器翻译评估并非机器翻译评估
Abstract
Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often correlate less well with human judgments. As a result, source-free, reference-based evaluation has become the practical norm, even though it is unfaithful to the definition of translation adequacy and unfair to systems whose outputs preserve the source meaning while differing from the reference. This paper argues that adequacy must be judged with respect to the source. A reference is only one possible rendering of the source and may introduce bias, under-specification, or errors. We further argue that source-reference-hypothesis evaluation is fair only when the judge treats the reference as auxiliary evidence rather than as the primary standard. Otherwise, even source-aware evaluation can reduce adequacy to preference towards reference. We show the existing hybrid metrics are highly reliant on reference compared to source. Our argument is not that all automatic MT metrics fail to use the source. Rather, we argue that any evaluation protocol that removes the source, or allows the reference to dominate the source, is structurally incomplete for adequacy evaluation. However, existing MT papers generally prefer reference-based metrics and use QE metrics only when reference is unavailable. We therefore call for QE to be reframed as a primary approach to source-grounded adequacy evaluation, rather than as a fallback motivated by missing references. We further call for hybrid metrics whose designs explicitly prioritize source--hypothesis faithfulness while using references only as complementary evidence.
Chinese Translation
基于参考的评估指标仍然是机器翻译评估的标准选择,部分原因是质量估计方法与人类判断的相关性往往较低。因此,无源的基于参考的评估已成为实际规范,尽管这与翻译充分性的定义不符,并且对那些输出保留源语义但与参考不同的系统不公平。本文主张,充分性必须相对于源语进行判断。参考仅是源语的一种可能表达,可能引入偏见、描述不足或错误。我们进一步认为,源-参考-假设评估只有在评估者将参考视为辅助证据而非主要标准时才是公平的。否则,即使是关注源语的评估也可能将充分性简化为对参考的偏好。我们展示了现有的混合指标相比于源语高度依赖于参考。我们的论点并不是所有自动机器翻译指标都未使用源语。相反,我们主张,任何去除源语或允许参考主导源语的评估协议在结构上都是不完整的,无法进行充分性评估。然而,现有的机器翻译论文通常更倾向于使用基于参考的指标,并仅在参考不可用时使用质量估计指标。因此,我们呼吁将质量估计重新框定为源基础充分性评估的主要方法,而不是作为缺失参考时的应急措施。我们进一步呼吁设计混合指标,明确优先考虑源-假设的忠实性,同时仅将参考作为补充证据。
cs.CL / 53 / 2608.20927
MentorPulse: Refreshing Cross-Model Latent Guidance for Long-Form Generation
MentorPulse:为长文本生成提供新鲜的跨模型潜在指导
Abstract
Cross-model latent guidance lets a frozen large mentor encode an input once and a frozen small student generate from the resulting signal. Existing methods keep this signal fixed, assuming it stays useful as the output grows; we show this fails in long-form generation. On multi-turn instruction following, static guidance pushes a 4B student's constraint satisfaction 2.5 points below its no-guidance baseline; a training-free refresh every 16 tokens changes only the memory content and restores a 2.0-point gain over that baseline. We propose MentorPulse to keep guidance fresh at practical cost: it compresses mentor states into a capped slot memory, incrementally processes newly generated tokens, and updates the memory that the student reads through gated cross-attention without resetting the student's KV cache. Windowed Refresh Training exposes the bridge to prefix-conditioned memory. Across thirteen datasets, MentorPulse closes 52.2% of the mentor-student gap on macro average, outperforming C2C, T2T, and equal-budget LoRA, with the largest gains on long outputs. It performs best on all eleven mentor-student pairs from three model families, with margins that narrow as the capability gap grows, and a lightweight read-pattern check predicts the gain before deployment. Measured costs identify refresh intervals that dominate text guidance on long outputs.
Chinese Translation
跨模型潜在指导允许一个冻结的大型导师对输入进行一次编码,而一个冻结的小型学生则从生成的信号中进行生成。现有方法保持该信号不变,假设随着输出的增长它仍然有效;我们展示了在长文本生成中这一假设的失败。在多轮指令跟随中,静态指导使得一个4B学生的约束满足度比无指导基线低2.5分;每16个标记进行一次无训练的刷新仅改变内存内容,并恢复了比基线高出2.0分的效果。我们提出了MentorPulse,以实用的成本保持指导的新鲜感:它将导师状态压缩到一个有限的槽内存中,逐步处理新生成的标记,并通过门控交叉注意力更新学生读取的内存,而不重置学生的KV缓存。窗口刷新训练揭示了与前缀条件内存的桥梁。在十三个数据集上,MentorPulse在宏观平均上缩小了52.2%的导师-学生差距,优于C2C、T2T和相同预算的LoRA,在长输出上获得了最大的提升。它在来自三种模型家族的所有十一对导师-学生中表现最佳,随着能力差距的增大,差距逐渐缩小,而轻量级的读取模式检查在部署前预测了增益。测量成本确定了在长输出上主导文本指导的刷新间隔。
cs.CL / 54 / 2608.20953
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs
量化感知修复:恢复压缩的4位大型语言模型的实用方案
Abstract
Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding, and long-context behavior enough to require a recovery, or healing, stage before deployment. The default recipe, quantization-aware training (QAT), re-fits the compressed, quantized model to hard labels; in our pipeline it converged slowly and collapsed past its peak. We adopted Quantization-Aware Healing (QAH) instead. Because a structurally compressed model is never independently trained at full precision, its bfloat16 checkpoint is a distillation-recovered approximation of the original; QAH distills the 4-bit student directly from the original, uncompressed model. On a GPT-OSS 120B to 60B to MXFP4 pipeline, the QAH student matches or beats its bfloat16 source on 7 of 9 benchmarks at roughly 4 times less weight memory and half the teacher's parameter count, and is released open-weight as Hypernova-60B. Against a matched QAT baseline it reaches a comparable peak about 7 times faster and stays stable under continued training, without hand-tuned early stopping. We also report deployment lessons, including a large, reproducible quality gap between distributed-training backends. Our aim is a recipe deployable without a multi-week hyper-parameter search.
Chinese Translation
以低成本提供大型语言模型意味着需要交付在结构上压缩到其参数的一小部分并量化为4位的模型。这两个步骤共同导致推理、数学、编码和长上下文行为的显著退化,因此在部署之前需要一个恢复或修复阶段。默认方案,即量化感知训练(QAT),是将压缩的量化模型重新拟合到硬标签;在我们的流程中,它收敛缓慢并在达到峰值后崩溃。因此,我们采用了量化感知修复(QAH)。由于结构上压缩的模型从未以全精度独立训练,其bfloat16检查点是原始模型的蒸馏恢复近似;QAH直接从原始未压缩模型蒸馏出4位学生模型。在GPT-OSS 120B到60B到MXFP4的流程中,QAH学生在9个基准测试中的7个上与其bfloat16源匹配或超越,且占用的内存大约是其四分之一,参数数量也只有教师模型的一半,并以开放权重形式发布为Hypernova-60B。与匹配的QAT基线相比,它以大约7倍的速度达到相似的峰值,并在持续训练下保持稳定,无需手动调优的早停策略。我们还报告了部署经验教训,包括分布式训练后端之间存在显著且可重复的质量差距。我们的目标是提供一种无需进行多周超参数搜索即可部署的方案。
cs.CL / 55 / 2608.20964
Extractive Summarization for Arabic Documents Using SAraBERT with a Semantic Siamese Similarity Evaluation Metric
基于 SAraBERT 的阿拉伯文档提取式摘要与语义双胞胎相似度评估指标
Abstract
In this research, we introduce SAraBERT, an enhanced version of AraBERT which proposes inter-sentence transformer layers for extractive summarization tasks. To ensure that the summaries generated by SAraBERT achieve a high coverage of the document's main ideas, we propose Semantic Siamese Similarity, a novel evaluation metric that measures the level of similarity between two text inputs. We validated using BLEU, ROUGE, and Semantic Siamese similarity on Sarabert and published related models. Simulation results showed the effectiveness of our proposed model and motivate follow on research.
Chinese Translation
在本研究中,我们介绍了 SAraBERT,这是 AraBERT 的增强版本,提出了用于提取式摘要任务的跨句子变换层。为了确保 SAraBERT 生成的摘要能够高覆盖文档的主要思想,我们提出了语义双胞胎相似度,这是一种新颖的评估指标,用于测量两个文本输入之间的相似程度。我们通过在 Sarabert 和相关模型上使用 BLEU、ROUGE 和语义双胞胎相似度进行了验证。仿真结果显示了我们提出的模型的有效性,并激励后续研究。
cs.CL / 56 / 2608.21019
Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models
面向目标的校准数据选择以保持量化语言模型中的不确定性
Abstract
Quantization is widely used to deploy large language models, but its effect on uncertainty behavior, such as confidence, margins, and abstention, is rarely treated as a primary objective. We frame calibration-data selection for quantization as a target-dependent uncertainty-preservation problem. Different deployments emphasize different regions of the input distribution, yet prior work mainly optimizes accuracy-oriented compression metrics or adjusts scores after quantization. We formalize this goal with distributional and boundary preservation risks, and provide a simple mixture-mismatch argument explaining why no single calibration recipe should be expected to fit all targets. We introduce Doubt-Preserving Quantization (DPQ), a lightweight pre-quantization recipe family that uses full-precision predictions to construct target-aligned calibration mixtures of high-doubt examples and generic anchors. Across 8 language models, 9 NLP benchmarks, and 22 comparison methods, the leading fixed recipe changes with the preservation target: DPQ-r75 leads on SQuAD2 answerability-boundary preservation, while milder or single-signal variants, including DPQ-r50, confidence-only, and entropy-only, better preserve broad multiple-choice QA behavior. These results show that calibration data should be selected for the specific full-precision score behavior a deployment needs to preserve, rather than treated as a fixed quantization detail.
Chinese Translation
量化技术被广泛应用于大规模语言模型的部署,但其对不确定性行为(如置信度、边际和弃权)的影响很少被视为主要目标。我们将量化的校准数据选择框架构建为一个依赖目标的不确定性保持问题。不同的部署强调输入分布的不同区域,而之前的研究主要优化以准确性为导向的压缩指标或在量化后调整得分。我们用分布和边界保持风险形式化了这一目标,并提供了一个简单的混合不匹配论证,解释了为什么不应期望任何单一的校准方案适用于所有目标。我们引入了不确定性保持量化(Doubt-Preserving Quantization, DPQ),这是一种轻量级的预量化方案系列,利用全精度预测构建与目标对齐的高不确定性示例和通用锚点的校准混合。在8个语言模型、9个自然语言处理基准和22种比较方法中,最佳固定方案随着保持目标的不同而变化:DPQ-r75在SQuAD2答案边界保持上表现最佳,而较温和或单信号变体(包括DPQ-r50、仅置信度和仅熵)则更好地保持广泛的多项选择问答行为。这些结果表明,校准数据应根据部署需要保持的特定全精度得分行为进行选择,而不是视为固定的量化细节。
cs.CL / 57 / 2608.21021
Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge
基于LLM的5G领域知识与故障分析的自由文本评估
Abstract
Real-world fault analysis in 5G and emerging 6G networks demands domain expertise to analyze free-text diagnostics, including root-cause explanations and recommended actions. LLMs have emerged as a promising approach to automating this, yet whether lightweight, edge-deployable models are capable of performing in-depth free-text diagnostics remains an open question. While existing benchmarks rely on restrictive MCQs with fixed answer keys, this paper evaluates 5G domain understanding and fault analysis in a free-text generation format. Transitioning to this paradigm requires evaluating lightweight, edge-deployable AI models on open-ended diagnostic reasoning, alongside a dependable framework to validate these text outputs at scale. To address this we evaluate three lightweight LLMs, Claude-Haiku-4.5, GPT-5.4-Mini, and Gemini-3.1-Flash-Lite, on free-text 5G domain knowledge and fault-analysis tasks across three benchmarks, TeleQNA ORAN FT, 5G-Faults FT, and TeleInter FT. Three independent frontier judges score outputs, and pairwise inter-judge agreement is measured as an empirical test of the LLM-as-Judge methodology. All three models reach at least 90% accuracy on fault diagnosis, while zero-shot recall of 3GPP and O-RAN specifications remains the critical gap, with all models scoring below 60%. Mean inter-judge agreement is at least 0.90 across all runs, indicating that multi-judge LLM scoring produces consistent, reproducible grades for open-ended telecom responses. Operationally, Gemini-3.1-Flash-Lite offers the best efficiency trade-off, combining competitive accuracy with the lowest inference cost and latency, making it the most suitable candidate for production telecom deployments.
Chinese Translation
在5G及新兴6G网络中,现实世界的故障分析需要领域专业知识,以分析自由文本诊断,包括根本原因解释和建议措施。大语言模型(LLMs)作为一种自动化的有前景的方法逐渐受到关注,但轻量级、可边缘部署的模型是否能够进行深入的自由文本诊断仍然是一个悬而未决的问题。现有基准依赖于固定答案的限制性选择题(MCQs),而本文则以自由文本生成格式评估5G领域理解和故障分析。转向这一范式需要对轻量级、可边缘部署的人工智能模型在开放式诊断推理上的评估,以及一个可靠的框架来大规模验证这些文本输出。为此,我们在自由文本的5G领域知识和故障分析任务上评估了三种轻量级LLM,分别是Claude-Haiku-4.5、GPT-5.4-Mini和Gemini-3.1-Flash-Lite,基于三个基准:TeleQNA ORAN FT、5G-Faults FT和TeleInter FT。三位独立的前沿评审员对输出进行评分,并通过配对评审员间一致性作为LLM-as-Judge方法论的实证测试。所有三种模型在故障诊断上至少达到90%的准确率,而3GPP和O-RAN规范的零样本召回率仍然是关键差距,所有模型的得分均低于60%。在所有测试中,评审员间的一致性均值至少为0.90,表明多评审员LLM评分为开放式电信响应产生了一致且可重复的评分。在操作上,Gemini-3.1-Flash-Lite提供了最佳的效率权衡,结合了竞争性的准确性以及最低的推理成本和延迟,使其成为生产电信部署中最合适的候选者。
cs.CL / 58 / 2608.21023
Scaling Unsupervised Word Alignment to Documents via Structural Constraints
通过结构约束将无监督词对齐扩展到文档
Abstract
Word alignment has traditionally been studied between sentences, but many cross-lingual tasks increasingly require correspondences across full documents. While recent multilingual embedding models can encode long inputs, we show that applying algorithms designed for sentences directly to documents leads to performance degradation. To address this, we introduce CTFAlign, a lightweight, training-free approach for document-level word alignment. CTFAlign applies a coarse-to-fine refinement strategy that restricts the alignment search space to semantically similar regions. Additionally, we introduce MDPAlign, a simpler alternative that constrains alignments by position with a main diagonal prior. Both approaches operate directly on full documents without relying on sentence segmentation or sentence alignment. We evaluate these methods across six language pairs varying in typological distance, resourcedness, and document length. Averaged over three models, CTFAlign reduces word alignment error rate from 0.412 to 0.326. These gains transfer downstream, leading to improvements in document-level translation coverage evaluation and recognition of semantic differences. We release CTFAlign as a Python package and make the code and data to reproduce our experiments publicly available.
Chinese Translation
词对齐传统上是在句子之间进行研究,但许多跨语言任务越来越需要在完整文档之间建立对应关系。尽管最近的多语言嵌入模型能够编码较长的输入,我们发现直接将为句子设计的算法应用于文档会导致性能下降。为了解决这个问题,我们提出了CTFAlign,这是一种轻量级、无需训练的文档级词对齐方法。CTFAlign采用粗到细的精细化策略,将对齐搜索空间限制在语义相似的区域。此外,我们还介绍了MDPAlign,这是一种更简单的替代方案,通过主对角线先验来约束对齐位置。这两种方法均直接在完整文档上操作,而不依赖于句子分割或句子对齐。我们在六对语言上评估了这些方法,这些语言对在类型距离、资源丰富度和文档长度上各不相同。根据三种模型的平均结果,CTFAlign将词对齐错误率从0.412降低到0.326。这些提升在下游任务中得以转移,改善了文档级翻译覆盖评估和语义差异的识别。我们将CTFAlign作为Python包发布,并公开提供重现我们实验的代码和数据。
cs.CL / 59 / 2608.21043
Evidence-Consistent Generative Detection under Scenario-Level Distribution Shift
情境级分布转移下的一致证据生成检测
Abstract
Conventional in-distribution evaluation can overestimate robustness when training and test data share recurring task-specific patterns or surface cues. This risk is especially relevant in social-engineering fraud detection, where attackers can preserve malicious intent while changing the scenario, impersonated entity, or wording. We study this problem as scenario-level out-of-distribution (SL-OOD) detection for SMS and voice phishing, where entire attack scenarios are held out from training while the label space remains fixed. This setting tests whether models can generalize to unseen attack scenarios using decision-relevant evidence rather than familiar scenario-specific cues. Using this SL-OOD evaluation, we find that high in-distribution performance does not reliably predict held-out robustness across feature-, encoder-, and decoder-based baselines. We interpret this gap as scenario memorization: reliance on recurring scenario-specific lexical or entity cues rather than decision-relevant evidence. We propose ECoG, an evidence-consistent generative framework that combines evidence-span supervision with a rationale-label consistency objective during training. On the 0.5B decoder, relative to the same backbone trained without consistency regularization, ECoG raises Macro-F1 on OOD challenging instances by 3.22 points, reduces the share of predictions whose generated rationale supports the opposite label by 4.22 points, and increases token-level overlap with reference evidence spans by 8.38 points; the reduction in prediction-rationale inconsistency is consistent across four decoder backbones. These results suggest that compact generative detectors can benefit from evidence supervision and rationale-label consistency under social-engineering shift.
Chinese Translation
传统的分布内评估在训练和测试数据共享重复的任务特定模式或表面线索时可能会高估模型的鲁棒性。这一风险在社会工程欺诈检测中尤为相关,因为攻击者可以在改变场景、冒充实体或措辞的同时保持恶意意图。我们将这个问题研究为情境级分布外检测(Scenario-Level Out-of-Distribution, SL-OOD),针对短信和语音钓鱼攻击,其中整个攻击场景在训练中被排除,而标签空间保持不变。该设置测试模型是否能够利用决策相关证据而非熟悉的场景特定线索,推广到未见过的攻击场景。通过这种 SL-OOD 评估,我们发现高分布内性能并不能可靠地预测在特征、编码器和解码器基线上的持出鲁棒性。我们将这一差距解释为场景记忆:依赖于重复的场景特定词汇或实体线索,而不是决策相关证据。我们提出了 ECoG,这是一种一致证据生成框架,在训练过程中结合了证据跨度监督和理由-标签一致性目标。在 0.5B 解码器上,相较于未进行一致性正则化的相同骨干网络,ECoG 在 OOD 挑战实例上的宏观 F1 提升了 3.22 分,减少了生成的理由支持相反标签的预测比例 4.22 分,并且与参考证据跨度的标记级重叠增加了 8.38 分;预测-理由不一致性的减少在四个解码器骨干网络中是一致的。这些结果表明,在社会工程转移下,紧凑的生成检测器可以从证据监督和理由-标签一致性中受益。
cs.CL / 60 / 2608.21074
PromptResponse: Optimizing Prompts for LLM Coding Tasks
PromptResponse:优化大型语言模型编码任务的提示
Abstract
Large language models (LLMs) are increasingly used in research workflows and software development pipelines, yet their output remains sensitive to input prompt variations. This paper presents $\unicode{x00AB}$PromptResponse$\unicode{x00BB}$, a controlled study examining how formatting and LLM-based tuning of coding task prompts affect the resulting code's performance, efficiency, and stability. Using five semantically identical yet syntactically distinct variants of the HumanEval dataset$\unicode{x2014}$baseline, JSON, Markdown, YAML, and an LLM-tuned version$\unicode{x2014}$we had GPT-4o solve its coding problems over 8200$\unicode{x00A0}$executions. Our results show that consistent formatting$\unicode{x2014}$especially JSON$\unicode{x2014}$improves generation efficiency and syntactic stability, with minor gains in task performance. Conversely, the LLM-tuned prompts resulted in significantly degraded task performance without significant improvements in any other dimension. These findings suggest that low-effort reformatting alone can yield measurable improvements, while tuning must account for model alignment. We conclude our work with providing a set of practical recommendations informed by our results as well as releasing our dataset variants and evaluation pipeline for future work.
Chinese Translation
大型语言模型(LLMs)在研究工作流和软件开发流程中越来越多地被使用,但它们的输出仍然对输入提示的变化敏感。本文提出了“PromptResponse”,一项控制性研究,考察了编码任务提示的格式化和基于LLM的调优如何影响生成代码的性能、效率和稳定性。我们使用了五种语义相同但句法不同的人类评估数据集(HumanEval)变体——基线、JSON、Markdown、YAML,以及一个经过LLM调优的版本——让GPT-4o在8200次执行中解决编码问题。我们的结果表明,一致的格式化——尤其是JSON——提高了生成效率和句法稳定性,任务性能也有小幅提升。相反,LLM调优的提示导致任务性能显著下降,而在其他维度没有显著改善。这些发现表明,仅通过低成本的重新格式化就能获得可测量的改进,而调优必须考虑模型的对齐性。我们在研究结尾提供了一系列基于结果的实用建议,并发布了我们的数据集变体和评估流程以供未来研究使用。
cs.CL / 61 / 2608.21087
Jokes Aside: Measuring the Semantic Distance of Double Meanings
开个玩笑:测量双重意义的语义距离
Abstract
Large language models have significantly enriched the toolkit for computational humor research, particularly in the automated generation of jokes and puns. A key innovation, contextual embedding vectors, offers new opportunities to revisit and refine earlier hypotheses. Notably, Petrovic and Matthews (2013) proposed a joke generation model based on the scheme "I like my X like I like my Y, Z" (e.g. "I like my ice like I like my dreams, crushed"). They suggested that joke hilarity increases with: a) frequent association of Z with X and Y, b) rarity of Z, c) ambiguity of Z, and d) meaning distance between X and Y. Building on this, Winters et al. (2019) proposed a set of metrics, based on Google Ngrams and Word2Vector. In this work, three out of their five metrics are revisited with word embeddings: obviousness, compatibility, and comparison. Another measure, symmetry, defined as closeness of Z to both X and Y, is introduced here for the first time. Two models were used to collect the embedding vectors (OpenAI text-embedding-3-small and MiniLM all-MiniLM-L6-v2) on three datasets: JokeJudger, Expunations, and rJokes. The last two datasets, Expunations, and rJokes, were expanded by adding paired sentences that captured the ambiguous expression at the core of each joke in its two different meanings. Results revealed that models trained on the proposed metrics performed poorly in predicting humor ratings: on JokeJudger, the best model achieved 57.1% accuracy, below the 61.5% baseline, while performance on Expunations and rJokes was even lower. Nevertheless, the symmetry metric seems consistently associated with higher-rated jokes, suggesting it may capture a necessary -though not sufficient- property of humor.
Chinese Translation
大型语言模型显著丰富了计算幽默研究的工具箱,特别是在笑话和双关语的自动生成方面。一项关键创新,即上下文嵌入向量,为重新审视和完善早期假设提供了新的机会。值得注意的是,Petrovic 和 Matthews(2013)提出了一种基于“我喜欢我的 X,就像我喜欢我的 Y,Z”的笑话生成模型(例如:“我喜欢我的冰淇淋,就像我喜欢我的梦想,压碎的”)。他们建议,笑话的幽默性随着以下因素的增加而增强:a) Z 与 X 和 Y 的频繁关联,b) Z 的稀有性,c) Z 的歧义性,以及 d) X 和 Y 之间的意义距离。在此基础上,Winters 等人(2019)提出了一组基于 Google Ngrams 和 Word2Vector 的度量。在本研究中,重新审视了他们的五个度量中的三个:明显性、兼容性和比较性。另一个度量,称为对称性,定义为 Z 与 X 和 Y 的接近程度,此次首次引入。使用了两个模型(OpenAI text-embedding-3-small 和 MiniLM all-MiniLM-L6-v2)在三个数据集上收集嵌入向量:JokeJudger、Expunations 和 rJokes。后两个数据集 Expunations 和 rJokes 通过添加成对句子进行了扩展,这些句子捕捉了每个笑话核心的歧义表达及其两种不同的含义。结果显示,基于所提度量训练的模型在预测幽默评分方面表现不佳:在 JokeJudger 上,最佳模型的准确率为 57.1%,低于 61.5% 的基线,而在 Expunations 和 rJokes 上的表现更低。然而,对称性度量似乎与高评分笑话始终相关,暗示它可能捕捉到幽默的一个必要(尽管不是充分)特性。
cs.CL / 62 / 2608.21088
When the Feature Pool Goes Algorithmic: Extending Mufwene's Ecology of Language Evolution to LLM-Mediated Exposure
当特征池变得算法化:将Mufwene的语言演化生态学扩展到大型语言模型介导的曝光
Abstract
Mufwene's ecological model locates language evolution in competition among variants contributed by individual idiolects and in speakers' selection from linguistic material made available through interaction. Large language models (LLMs) complicate this architecture without requiring the locus of selection to move away from human speakers. This article argues that LLMs are best treated as distributional mediators: they aggregate language produced across human populations, transform its distribution through training and post-training, and redistribute model-specific outputs at scale. I call the resulting ecological process algorithmic reweighting of the speaker-accessible distribution: model mediation can alter the relative frequencies with which competing variants reach human selectors. Emerging evidence on model-specific linguistic profiles and lexical uptake is consistent with parts of this pathway, but does not establish inevitable convergence. Human social evaluation remains decisive: model-associated forms may diffuse and become conventionalized, become socially recognizable as 'AI-like' and subsequently avoided, or fail to diffuse in the first place. The proposal extends Mufwene's feature-pool ecology one step upstream of speaker selection and yields testable predictions about uptake, model-version effects, convergence, and social reversal.
Chinese Translation
Mufwene的生态模型将语言演化定位于个体方言所贡献的变体之间的竞争,以及说话者从互动中获得的语言材料中进行选择。大型语言模型(LLMs)在不要求选择的中心从人类说话者转移的情况下,复杂化了这一架构。本文认为,LLMs最好被视为分布中介:它们聚合人类群体中产生的语言,通过训练和后训练转变其分布,并大规模重新分配模型特定的输出。我称这种生态过程为说话者可接触分布的算法重加权:模型中介可以改变竞争变体到达人类选择者的相对频率。关于模型特定语言特征和词汇采纳的新兴证据与这一路径的部分内容一致,但并未确立不可避免的趋同。人类的社会评估仍然是决定性的:与模型相关的形式可能扩散并成为常规化,可能被社会认知为“类AI”的形式并随后被避免,或根本未能扩散。该提议将Mufwene的特征池生态学向上扩展一步,超越了说话者选择,并产生了关于采纳、模型版本效应、趋同和社会逆转的可检验预测。
cs.CL / 63 / 2608.21206
No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation
无意的双关:以合理未知名称进行以人为中心的大型语言模型评估
Abstract
Person names are widely used as prompt variables in LLM evaluations of factuality, privacy leakage, bias and abstention, but when a name's evidential status is uncontrolled, measurements may conflate memorisation, retrieval, name priors and wrong-person attribution. We operationalise an unknown name as one with plausible First-Last form, no indexed full-name evidence, and no ambiguity signals under a documented validation run, and introduce PUN (Plausible Unknown Names), a protocol for constructing and validating such names, combining Wikidata-derived components, web-enabled LLM screening, and controlled search revalidation. We report acceptance rate, reproducibility, ablations, and a 204-participant human study, finding accepted names are more name-like than controls while participants recover person evidence in only 3% of cases. We release 300 names with comparison controls.
Chinese Translation
人名在大型语言模型(LLM)对事实性、隐私泄露、偏见和回避的评估中被广泛用作提示变量,但当一个名字的证据状态无法控制时,测量结果可能会混淆记忆、检索、名字先验和错误归属。我们将未知名称定义为一种具有合理的名-姓形式、没有索引的全名证据,并且在经过文档验证的运行中没有歧义信号的名称,并引入了PUN(Plausible Unknown Names),这是一个构建和验证此类名称的协议,结合了源自Wikidata的组件、网络启用的LLM筛选和受控搜索再验证。我们报告了接受率、可重复性、消融实验以及一项包含204名参与者的人类研究,发现被接受的名称比对照组更具名称特征,而参与者在仅3%的情况下恢复了人物证据。我们发布了300个名称及其对照组进行比较。
cs.CL / 64 / 2608.21236
RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models
RARE:在混合专家语言模型中将表示引导与专家路由解耦
Abstract
Representation engineering offers a lightweight means of controlling language-model behavior by modifying intermediate hidden states, but its direct application to Mixture-of-Experts (MoE) models introduces a structural mismatch. We first verify this failure mode through a series of empirical studies and find that preserving clean routing substantially recovers steering performance and that routing is more sensitive to semantic content than to behavioral changes under controlled content. Motivated by these findings, we introduce RARE, a router-agnostic representation engineering framework for MoE language models. RARE projects arbitrary behavioral perturbations onto the null space of the router matrix, thereby removing router-visible components, and further corrects routing drift propagated to selected downstream layers. To decide the best perturbation estimator in this framework, we evaluate five estimators on six heterogeneous open-weight MoE models across three steering scenarios: harmfulness, truthfulness, and factual editing. On harmfulness steering, RARE reaches an average attack success rate of 53.3% while retaining 67.8% MMLU accuracy, yielding a stronger aggregate effectiveness--utility trade-off than baselines. It further improves average TruthfulQA MC1 accuracy from 41.0% to 58.6% and CounterFact efficacy from 16.8% to 96.3%. These results support routing consistency as an important architectural consideration for adapting representation engineering to MoE models.
Chinese Translation
表示工程通过修改中间隐藏状态提供了一种轻量级的方式来控制语言模型的行为,但其直接应用于混合专家(MoE)模型时引入了结构不匹配。我们首先通过一系列实证研究验证了这种失败模式,发现保持干净的路由显著恢复了引导性能,并且路由对语义内容的敏感性高于在受控内容下的行为变化。基于这些发现,我们提出了RARE,一个与路由无关的表示工程框架,适用于MoE语言模型。RARE将任意的行为扰动投影到路由矩阵的零空间,从而去除路由可见的成分,并进一步修正传播到选定下游层的路由漂移。为了决定该框架中最佳的扰动估计器,我们在三个引导场景(有害性、真实性和事实编辑)下评估了六个异构开放权重MoE模型上的五个估计器。在有害性引导方面,RARE达到了53.3%的平均攻击成功率,同时保持了67.8%的MMLU准确率,提供了比基线更强的整体有效性-效用权衡。它还将平均TruthfulQA MC1准确率从41.0%提高到58.6%,将CounterFact有效性从16.8%提高到96.3%。这些结果支持路由一致性作为将表示工程适应于MoE模型的重要架构考虑。
cs.CL / 65 / 2608.21242
Affective Context Amplifies Sycophancy in LLM Responses
情感背景增强大型语言模型响应中的谄媚行为
Abstract
As conversational companions, large language models (LLMs) often have access to users' emotional states. We study how this affective context modulates LLM sycophancy in subjective, evaluative interactions, where users share actions or opinions that invite feedback. Drawing on ingratiation theory, we measure sycophancy as the divergence between a model's independent evaluation and its user-facing response, elicited by presenting the same content as either a third-party account or the user's own disclosure. Across seven LLMs and two Reddit datasets (r/AmItheAsshole and r/TrueUnpopularOpinion), we find that this divergence is systematic and strongly one-directional. User-facing responses consistently soften or withhold negative or oppositional judgments. Affective context further amplifies this divergence with negative states, particularly loneliness and distress, producing the largest effects. These findings suggest that affective context functions as a vulnerability signal that suppresses critical feedback when users may need it most, often through evasive sycophancy, in which models retreat toward non-committal responses rather than outright agreement.
Chinese Translation
作为对话伙伴,大型语言模型(LLMs)通常能够感知用户的情感状态。我们研究了这种情感背景如何调节LLM在主观评估互动中的谄媚行为,在这种互动中,用户分享的行为或观点会引发反馈。基于谄媚理论,我们将谄媚行为测量为模型独立评估与其面向用户的响应之间的偏差,该偏差是通过将相同内容呈现为第三方叙述或用户自身披露来引发的。在七个LLM和两个Reddit数据集(r/AmItheAsshole和r/TrueUnpopularOpinion)中,我们发现这种偏差是系统性的,并且具有显著的单向性。面向用户的响应始终会软化或抑制负面或对立的判断。情感背景进一步增强了这种偏差,尤其是在负面情绪状态下,特别是孤独和痛苦,产生了最大的影响。这些发现表明,情感背景作为一种脆弱信号,在用户可能最需要批评反馈时抑制了这种反馈,通常通过逃避性的谄媚行为表现出来,模型倾向于采取不明确的响应而非直接的同意。
cs.CL / 66 / 2608.21249
Benchmarking Patent Drafting from Inventor-Style Disclosures
从发明者风格披露中基准化专利撰写
Abstract
While recent large language models (LLMs) have achieved promising results on individual patent drafting tasks, they fundamentally fail to investigate the core challenge of real-world patent drafting: generating a complete and legally coherent patent application directly from early-stage invention materials. Prior work predominantly assumes later-stage, highly structured, or already legalistic inputs. However, real patenting workflows begin with informal, de-legalized disclosures authored by inventors. To bridge the gap, we introduce Dis2Pat, a disclosure-to-patent dataset that reflects realistic patenting workflows by requiring the generation of complete patent applications directly from inventor-style, de-legalized disclosures. Given the inherent difficulty of long-form, legally constrained patent drafting and the strong privacy requirements, we further propose a strong baseline named Patent-MAF. It is a multi-agent framework for locally deployable patent drafting. Benchmark results reveal that current LLMs exhibit limitations in patent drafting, while Patent-MAF provides a strong baseline that consistently outperforms evaluated open-source models and remains competitive with large closed-source models.
Chinese Translation
尽管近期的大型语言模型(LLMs)在个别专利撰写任务上取得了令人鼓舞的成果,但它们在根本上未能探讨现实世界专利撰写的核心挑战:直接从早期阶段的发明材料生成完整且法律上连贯的专利申请。以往的研究主要假设输入为后期、高度结构化或已具法律性质的材料。然而,真实的专利申请流程始于由发明者撰写的非正式、去法律化的披露。为了解决这一问题,我们引入了Dis2Pat,这是一个披露到专利的数据集,要求从发明者风格的去法律化披露中直接生成完整的专利申请,从而反映现实的专利申请流程。鉴于长篇、法律约束的专利撰写固有的困难以及强烈的隐私要求,我们进一步提出了一个强基线模型,命名为Patent-MAF。它是一个可本地部署的多智能体专利撰写框架。基准测试结果显示,当前的LLMs在专利撰写方面存在局限性,而Patent-MAF提供了一个强基线,始终优于评估的开源模型,并与大型闭源模型保持竞争力。
cs.CL / 67 / 2608.21252
EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering
EnSI-RAG:用于长文档问答的实体结构索引检索增强生成
Abstract
Question answering (QA) over long, connected documents remains challenging because relevant evidence may span multiple entities and their relationships. Existing retrieval-augmented generation (RAG) methods typically index documents as raw chunks and retrieve them through embedding similarity. Their performance degrades when chunk boundaries separate entities from supporting evidence or when a question requires multi-hop reasoning across the corpus. We propose EnSI-RAG (Entity-Structure-Indexed Retrieval-Augmented Generation), a framework that constructs a query-independent, entity-centered index. Each record (e, t, k, v) represents an entity e, its type t, a semantic category k in {property, relation, aspect}, and a value v, while retaining links to the original source passages. At query time, these records serve as retrieval handles, and an LLM synthesizes the retrieved passages into the final answer. This design separates evidence localization from answer synthesis while preserving traceable source evidence. Across Loong and Oolong, EnSI-RAG achieves an average accuracy of 78.24. Relative to the published baseline scores used as references, this is 6.62 points higher, suggesting its effectiveness across these settings. The code is available at https://github.com/RamonMeng/EnSI-RAG.
Chinese Translation
在长篇关联文档中进行问答(QA)仍然具有挑战性,因为相关证据可能跨越多个实体及其关系。现有的检索增强生成(RAG)方法通常将文档索引为原始块,并通过嵌入相似性进行检索。当块边界将实体与支持证据分隔开,或当问题需要跨文档进行多跳推理时,它们的性能会下降。我们提出了EnSI-RAG(实体结构索引检索增强生成),这是一个构建查询无关、以实体为中心的索引的框架。每个记录(e, t, k, v)表示一个实体e、其类型t、一个语义类别k(属于{属性、关系、方面}),以及一个值v,同时保留与原始源段落的链接。在查询时,这些记录充当检索句柄,LLM将检索到的段落合成为最终答案。该设计将证据定位与答案合成分离,同时保留可追溯的源证据。在Loong和Oolong数据集上,EnSI-RAG的平均准确率达到78.24。相对于作为参考的已发布基线分数,这提高了6.62分,表明其在这些设置中的有效性。代码可在https://github.com/RamonMeng/EnSI-RAG获取。
cs.CL / 68 / 2608.21265
Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning
记忆增强解锁高效的链式思维推理
Abstract
Large language models often rely on Chain-of-Thought (CoT) reasoning to solve complex tasks, but verbose reasoning traces introduce substantial inference overhead. CoT compression shortens generation, yet aggressive compression may disrupt logical coherence and degrade performance. We formalize this trade-off as the \textit{Context-Generation Substitution Law}, where explicit reasoning context substitutes for part of decode-time generation. Based on this principle, we propose \textit{Memory-Augmented Compression}, a training-free framework that constructs reusable reasoning memories from historical traces and retrieves them as prefill-side scaffolds. Rather than using raw demonstrations, these memories summarize reusable reasoning patterns, key constraints, and critical operations to compensate for information lost during compression. Experiments show that Memory consistently improves prompt-based Chain-of-Draft (CoD) compression across mathematical reasoning, complex reasoning, and science question answering tasks, yielding accuracy gains of 21.4, 28.0, 29.5, and 6.61 points over CoD on GSM8K, MATH, BBH, and MMLU-Sci, while achieving a 1.14--1.49$\times$ latency speedup over standard CoT. Memory is also compatible with token-level, reasoning-trace-level, and inference-state compression mechanisms. Further analyzes show that the gains come from relevant reasoning memories rather than simply increasing context length.
Chinese Translation
大型语言模型通常依赖链式思维(Chain-of-Thought, CoT)推理来解决复杂任务,但冗长的推理过程会带来显著的推理开销。CoT 压缩缩短了生成时间,但过度压缩可能会破坏逻辑连贯性并降低性能。我们将这种权衡形式化为 extit{上下文-生成替代法则},其中显式推理上下文替代部分解码时间生成。基于这一原则,我们提出了 extit{记忆增强压缩},这是一种无训练框架,能够从历史轨迹中构建可重用的推理记忆,并将其作为预填充侧支架进行检索。与其使用原始示例,这些记忆总结了可重用的推理模式、关键约束和重要操作,以弥补压缩过程中丢失的信息。实验表明,记忆在数学推理、复杂推理和科学问答任务中持续改善基于提示的链式草稿(Chain-of-Draft, CoD)压缩,在 GSM8K、MATH、BBH 和 MMLU-Sci 上分别获得了 21.4、28.0、29.5 和 6.61 的准确率提升,同时在标准 CoT 上实现了 1.14--1.49$ imes$ 的延迟加速。记忆还与令牌级、推理轨迹级和推理状态压缩机制兼容。进一步分析表明,收益来自相关的推理记忆,而不仅仅是增加上下文长度。
cs.CL / 69 / 2608.21315
Prompt-Model Interaction Reaches the Fixed Points: A deterministic, task-free structural readout -- and the factorizations of it that failed
提示-模型交互达到固定点:一种确定性的无任务结构读出及其失败的分解
Abstract
That a prompt's effect is not a property of the prompt is established: prompts optimised for one model degrade on another, and rankings reorder under neutral reformatting. That evidence is about task accuracy, which cannot say whether the interaction is a fact about task machinery or about the conditional distribution itself. We ask on a readout with no task in it: the fixed-point structure of the short-window argmax map x_{t+1} = argmax_x p(x | x_{t-1}, x_t), censused from 96 starts. It is deterministic, so nothing can be helped or hurt, and it exists only at short windows -- four of six models lose it entirely by window 16 -- so everything here concerns how a model reads a fragment. Two results. First, the interaction reaches this readout at full magnitude: nine tokens of conditioning move the fixed-point fraction across most of its range, change a four-way structural class, and reorder models, while instruction tuning worth 60.5 IFEval points moves the class by zero. Second, nothing we proposed carries it. Prefix length fails: the effect is not monotone. Four phenomenological factors -- prose-versus-markup, a universal direction, bidirectionality, instruct-resistance -- were each withdrawn within one run of being proposed, dissolved by widening the sample. And the nearest mechanistic account, attention-sink dominance of early tokens, predicts the sign of the shift on 2 of 5 models -- chance -- while a length-by-content cross shows it holds on real text and fails on our probe's uniformly random input, so we are outside its regime, not against it. One fixed nine-token prefix drives four models toward 0 and two toward 1; the bidirectionality survives in-distribution starts. On this readout the unit of explanation is the prompt-model pair. The recurring error it caught in us has a name: a criterion with a shape applied to a quantity with no room to vary.
Chinese Translation
提示的效果并不是提示本身的属性已被确立:为一个模型优化的提示在另一个模型上会降级,并且在中性重格式化下排名会重新排序。该证据与任务准确性有关,但无法说明交互是任务机制的事实还是条件分布本身的事实。我们在没有任务的读出上提出问题:短窗口 argmax 映射 x_{t+1} = argmax_x p(x | x_{t-1}, x_t) 的固定点结构,从 96 个起始点进行普查。它是确定性的,因此没有任何东西可以被帮助或伤害,并且仅在短窗口存在——六个模型中的四个在窗口 16 时完全失去它——因此这里的一切都与模型如何读取片段有关。两个结果。首先,交互在全幅度下达到此读出:九个条件令固定点比例在其大部分范围内移动,改变四向结构类别,并重新排序模型,而价值 60.5 IFEval 点的指令调优则未能改变类别。其次,我们提出的任何内容都无法携带它。前缀长度失败:该效应不是单调的。四个现象学因素——散文与标记、普遍方向、双向性、指令抗性——在被提出的一次运行中各自被撤回,通过扩大样本而溶解。而最近的机制解释,早期标记的注意力沉没主导,预测了 5 个模型中 2 个模型的转变符号——偶然——而长度与内容的交叉显示它在真实文本上成立,但在我们探针的均匀随机输入上失败,因此我们处于其范围之外,而不是对抗它。一个固定的九标记前缀将四个模型推向 0,两个模型推向 1;双向性在分布内的起始点中存活。在这个读出中,解释的单位是提示-模型对。它在我们身上捕捉到的反复错误有一个名称:一种形状的标准应用于没有变化余地的量。
cs.CL / 70 / 2608.21325
Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy
逐步分析:测量和引导大型语言模型如何进行心理治疗
Abstract
Users increasingly turn to large language models for emotional support, yet little is known about how these models actually conduct a psychotherapy interaction. We introduce an ontology of ten therapeutic moves: compact, function-based categories grounded in the MULTI-60 inventory, validated through an annotation campaign with five licensed psychologists, and scaled with a judge-based approach that matches expert agreement. Applying it to real counseling transcripts and model-led sessions, we compare the move distributions between human clinicians and a panel of frontier models. Models over-use inquiry at up to three times the human rate, neglect psychoeducation, and are strongly context-anchored: they carry forward strategies initiated by a human clinician but rarely initiate them themselves. Exposing the ontology as a set of tools roughly halves the mean deviation from the human move distribution and improves turn-level alignment with human therapist by 7-9 percentage points, without any fine-tuning.
Chinese Translation
用户越来越多地求助于大型语言模型以获取情感支持,但关于这些模型如何实际进行心理治疗互动的了解仍然有限。我们引入了一个包含十种治疗动作的本体:这些动作是基于MULTI-60清单的紧凑型、功能性分类,通过与五位持证心理学家的注释活动进行验证,并采用基于评审的方式进行扩展,以匹配专家一致性。我们将其应用于真实的咨询记录和模型主导的会话中,比较了人类临床医生与前沿模型小组之间的动作分布。模型在询问方面的使用频率是人类的三倍,忽视了心理教育,并且强烈依赖于上下文:它们延续了人类临床医生发起的策略,但很少主动发起这些策略。将本体作为一组工具展示,平均使人类动作分布的偏差减少了近一半,并在回合级别上与人类治疗师的对齐提高了7-9个百分点,而无需任何微调。