← Back to Index
Daily Research Digest

arXiv Papers

2026-08-18
605
Papers
4
Categories
605
Translated
收藏清单 0
机器人学 (Robotics)
89
cs.RO / 1 / 2608.14713

SpotlessGS: Relightable 3D Gaussian Splatting under Dynamic Illumination for Robotic Perception

SpotlessGS:动态照明下可重光照的三维高斯点云重建用于机器人感知
Hong, Liang, Wei, Jiaxin, Schaefer, Simon, Leutenegger, Stefan, Jung, Jaehyung
Abstract
Robots operating in dark or poorly lit environments rely on onboard lights, which often produce uneven illumination that degrades downstream perception tasks. Prior approaches based on 2D image enhancement lack reliable supervision and fail to preserve multi-view geometric consistency. To address these limitations, we extend Dark Gaussian Splatting (DarkGS) toward a more accurate and flexible relightable 3D reconstruction framework. First, we eliminate the need for explicit light parameter calibration by jointly optimizing lighting parameters within the Gaussian Splatting framework. Second, we introduce a low-frequency illumination model based on spherical harmonics (SH) to capture spatially varying residual and ambient lighting effects. Third, we incorporate an MLP-based Bidirectional Reflectance Distribution Function (BRDF) to model non-Lambertian reflectance. Experiments on synthetic and real-world datasets demonstrate that our method effectively mitigates illumination artifacts while improving rendering quality and quantitative performance over prior approaches. We further validate its benefits for robotic perception through a downstream task.
Chinese Translation
在黑暗或光线不足的环境中操作的机器人依赖于车载灯光,而这些灯光往往产生不均匀的照明,降低了下游感知任务的效果。基于二维图像增强的先前方法缺乏可靠的监督,并且未能保持多视角几何一致性。为了解决这些局限性,我们将暗高斯点云重建(Dark Gaussian Splatting,DarkGS)扩展为一个更准确和灵活的可重光照三维重建框架。首先,我们通过在高斯点云框架内联合优化照明参数,消除了对显式光参数标定的需求。其次,我们引入了一种基于球谐函数(Spherical Harmonics,SH)的低频照明模型,以捕捉空间变化的残余和环境光照效果。第三,我们结合了一种基于多层感知器(MLP)的双向反射分布函数(Bidirectional Reflectance Distribution Function,BRDF)来建模非朗伯反射。对合成和真实世界数据集的实验表明,我们的方法有效减轻了照明伪影,同时提高了渲染质量和定量性能,优于先前的方法。我们进一步通过下游任务验证了其在机器人感知中的优势。
cs.RO / 2 / 2608.14772

MISTac: A Vision-Based Tactile Sensor for Minimally Invasive Surgery

MISTac:一种基于视觉的微创手术触觉传感器
Koch, Robin, Mascot, Annabella, Younis, Rayan, Wagner, Martin, Speidel, Stefanie, Cutkosky, Mark, Sieber, Ingo, Calandra, Roberto
Abstract
Minimally invasive and robot-assisted surgery offer many advantages over traditional open surgery, but deprive surgeons of tactile feedback and the ability to palpate tissue with their fingers. To address this lack of tactile feedback, we introduce the MISTac, a high resolution vision-based tactile sensor specifically designed for palpation in MIS. The sensor has a replaceable sensor tip with a diameter of 8 mm which allows it to fit through the trocars used in minimally invasive surgery. Its modular 3D-printed case design allows the use of bulky off-the-shelf illumination and imaging hardware that can easily be exchanged and upgraded. The sensor has an optical resolution of 176.68 $\mu m$, a tactile resolution of 250 $\mu m$, and can resolve forces as little as 24.3 mN. An in vivo study with the sensor shows its usability in minimally invasive surgery. We trained a machine learning model with the tactile data collected in the trial on a tissue classification task achieving an aggregate accuracy of ~84% in a leave-one-out cross validation. Tactile sensors have the potential to one day aid surgeons during minimally invasive surgery with tasks such as tissue classification or intra-operative tumor localization; MISTac is a small step towards this vision. We open-source MISTac at https://github.com/lasr-lab/mistac
Chinese Translation
微创手术和机器人辅助手术相较于传统开放手术具有诸多优势,但却剥夺了外科医生的触觉反馈以及用手指触诊组织的能力。为了解决这一触觉反馈的缺失,我们推出了MISTac,这是一种专为微创手术中的触诊设计的高分辨率基于视觉的触觉传感器。该传感器配备可更换的传感器尖端,直径为8毫米,能够通过微创手术中使用的 trocar。其模块化的3D打印外壳设计允许使用笨重的现成照明和成像硬件,且可以轻松更换和升级。传感器的光学分辨率为176.68微米,触觉分辨率为250微米,能够解析最低24.3毫牛顿的力量。通过传感器进行的体内研究显示其在微创手术中的可用性。我们使用在试验中收集的触觉数据训练了一个机器学习模型,进行组织分类任务,在留一交叉验证中实现了约84%的总体准确率。触觉传感器有潜力在未来帮助外科医生进行微创手术中的任务,如组织分类或术中肿瘤定位;MISTac是实现这一愿景的一小步。我们在https://github.com/lasr-lab/mistac上开源了MISTac。
cs.RO / 3 / 2608.14822

Imagining Recovery: Inference-Time Counterfactual Realignment for Vision-Language-Action Models

想象恢复:针对视觉-语言-行动模型的推理时反事实重校正
Zhang, Yanyan, Liu, Disheng, Ye, Kai, Song, Chaoda, Li, Xinpeng, Hariri, Mohsen, Singh, Vikash, Yin, Yu, Chaudhary, Vipin
Abstract
Vision-language-action (VLA) models have improved the flexibility and generality of robotic manipulation, yet they remain fragile to online disruptions, such as changes in task goal, scene configuration, or robot state. Existing recovery methods often require failure data, policy retraining, or external corrective agents, introducing additional data requirements and execution risks. We propose Counterfactual Realignment (CoRe), a training-free framework that recovers a frozen VLA at inference time without failure data. Upon detecting a deviation, CoRe imagines how the policy would continue toward the current goal from a recent viable state, using synthesized observations in place of physical execution, and then minimally realigns the robot and scene to rejoin this imagined continuation before returning control to the policy. Recovery is therefore planned without physical trial-and-error, preserves completed task progress, and handles both mid-episode instruction changes and physical perturbations in a unified manner. Extensive experiments across multiple simulators, VLA backbones, and real-world settings show that CoRe improves success rates by up to 85.0 percentage points to near-nominal levels while reducing physical restorations by 42.2%, without policy fine-tuning or failure-specific recovery training.
Chinese Translation
视觉-语言-行动(VLA)模型提高了机器人操作的灵活性和通用性,但仍然对在线干扰(如任务目标、场景配置或机器人状态的变化)较为脆弱。现有的恢复方法通常需要失败数据、策略再训练或外部纠正代理,这引入了额外的数据需求和执行风险。我们提出了反事实重校正(Counterfactual Realignment,CoRe),这是一个无需训练的框架,能够在推理时恢复冻结的VLA,而不需要失败数据。在检测到偏差后,CoRe想象策略如何从最近的可行状态继续朝当前目标前进,使用合成的观察代替物理执行,然后对机器人和场景进行最小的重校正,以重新加入这一想象的延续,然后将控制权返回给策略。因此,恢复是在没有物理试错的情况下进行的,保留了已完成任务的进展,并以统一的方式处理中途指令变化和物理扰动。在多个模拟器、VLA骨干网络和现实世界环境中进行的广泛实验表明,CoRe将成功率提高了多达85.0个百分点,接近名义水平,同时将物理恢复减少了42.2%,且无需策略微调或特定于失败的恢复训练。
cs.RO / 4 / 2608.14860

Modeling and Control of an Eel-Inspired Soft Robot for Design Optimization

基于鳗鱼灵感的软机器人建模与控制及设计优化
Jiang, Zhangjingyi, Campbell, Mark
Abstract
Anguilliform locomotion is a highly efficient swimming mode; the advent of new materials for soft robots enables the development of an eel-inspired soft robot. This paper presents a simulation model of an eel-inspired soft robot designed for anguilliform swimming. This model can aid in design optimization and the development of model-based estimation, reasoning, and control systems. A Finite Element Method (FEM) model of an elastic rod is used to capture the soft materials of the robotic fish, which makes it particularly amenable to variation over time as the material properties change. The material model is coupled with a hydrodynamic force model to simulate the behavior of a soft, elongated robot in water. The model is used to demonstrate the effectiveness of the proposed control approaches in achieving desired swimming behaviors. It also provides insights into design decisions, including the robustness of different system configurations and the impact of material degradation and failure. The results show that slightly asymmetric designs are advantageous, offering comparable swimming velocities but greater maneuverability. This model can be used to guide future robotic design decisions aimed at optimizing performance for specific tasks.
Chinese Translation
鳗形运动是一种高效的游泳模式;新材料的出现使得基于鳗鱼灵感的软机器人得以发展。本文提出了一种旨在实现鳗形游泳的基于鳗鱼灵感的软机器人仿真模型。该模型可以帮助进行设计优化,并促进基于模型的估计、推理和控制系统的发展。采用有限元方法(Finite Element Method, FEM)模型来捕捉机器鱼的软材料特性,使其特别适合于随着材料属性变化而随时间变化。材料模型与水动力模型相结合,以模拟软而细长的机器人在水中的行为。该模型用于展示所提控制方法在实现期望游泳行为方面的有效性。同时,它还为设计决策提供了见解,包括不同系统配置的鲁棒性以及材料退化和失效的影响。结果表明,稍微不对称的设计是有利的,能够提供相当的游泳速度但更大的机动性。该模型可用于指导未来的机器人设计决策,以优化特定任务的性能。
cs.RO / 5 / 2608.14865

Real-time Estimator of Actuator Control and Health (REACH) on an Eel-Inspired Soft Robot

基于鳗鱼启发的软体机器人执行器控制与健康的实时估计器(REACH)
Jiang, Zhangjingyi, Park, Myungsun, Tolley, Michael T., Campbell, Mark
Abstract
An actuator health estimation algorithm for a soft swimming robot that can perform anguilliform swimming is developed. Due to harsh operational environments of underwater robots, and the common degradation of soft robot materials and actuators, accurate estimation of actuator functionality is necessary for robots to perform their missions as well as return to base in the event of actuator degradation and failure. Termed REACH (Real-time Estimator of Actuator Control and Health), the architecture employs a soft robot model, sigma point filter, and a formal statistical hypothesis test to adequately capture the nonlinearities and changes over time. The performance of REACH using three sensor types (GPS, IMU, and Bend Sensor) with one sensor on each actuator is compared, demonstrating that both bend sensor and IMU are adequate choices. Sensor quantity and placement are evaluated for IMU and bend sensor, showing two sensors are sufficient for IMU, whereas three sensors are needed for bend sensor. Three swimming gaits (linear swimming, wide turning, tight turning) are compared, demonstrating that REACH can successfully predict actuator health for all three gaits, with minimal differences in performance. A filter validation method shows the fault estimation algorithm is statistically consistent in finding the correct degradation. The approach is experimentally evaluated using bend sensor data collected from a fish robot, demonstrating that REACH can successfully estimate actuator health with noisy data and variations in manufacturing.
Chinese Translation
本文开发了一种用于软体游泳机器人的执行器健康估计算法,该机器人能够执行鳗鱼式游泳。由于水下机器人面临严酷的操作环境,以及软体机器人材料和执行器的常见退化,准确估计执行器的功能对于机器人执行任务以及在执行器退化和故障时安全返回基地至关重要。该算法被称为REACH(执行器控制与健康的实时估计器),其架构采用软体机器人模型、sigma点滤波器和正式的统计假设检验,以充分捕捉非线性特性和随时间变化。通过比较使用三种传感器类型(GPS、IMU和弯曲传感器)且每个执行器上各配备一个传感器的REACH性能,结果表明弯曲传感器和IMU都是合适的选择。对IMU和弯曲传感器的传感器数量和布置进行了评估,显示IMU需要两个传感器,而弯曲传感器则需要三个传感器。比较了三种游泳姿态(线性游泳、宽转弯、紧转弯),结果表明REACH能够成功预测所有三种姿态下的执行器健康,且性能差异最小。一种滤波器验证方法表明,该故障估计算法在找到正确的退化方面具有统计一致性。该方法通过使用从鱼机器人收集的弯曲传感器数据进行实验评估,证明REACH能够在噪声数据和制造变异的情况下成功估计执行器健康。
cs.RO / 6 / 2608.14902

Geometry-Aware Online Mapping for 3D Gaussian Splatting SLAM

基于几何感知的3D高斯喷溅SLAM在线地图构建
Luu, Thai, Tran, Quan, Phan, Hieu, Dang, Tuan
Abstract
Recent 3D Gaussian Splatting (3DGS) has enabled efficient photorealistic view synthesis and is rapidly being adopted in simultaneous localization and mapping (SLAM) systems for online mapping. In these systems, a Gaussian map must be expanded and refined incrementally while tracking runs in real time, so initialization and density control directly determine where limited computation and iterations are spent. This contrasts with offline 3DGS reconstruction, where such heuristics can be amortized over long optimization schedules. However, most 3DGS-SLAM pipelines inherit initialization and density-control heuristics from offline reconstruction, which can become brittle under the strict per-keyframe optimization budgets and incremental map growth of online SLAM. In this work, we revisit these heuristics in a decoupled 3DGS-SLAM setting and propose three geometry-aware methods that operate in the mapping thread: transmittance-preserving densification, camera-aware scale initialization from depth and intrinsics, and error-guided densification that focuses new primitives on high-residual regions. Our results show consistent improvements in rendering quality with negligible overhead, highlighting the coupling between photometric residuals and pose uncertainty in online SLAM. We will open-source our code to the community to foster growth and validate reproducibility.
Chinese Translation
近期的3D高斯喷溅(3D Gaussian Splatting, 3DGS)技术使得高效的照片级真实感视图合成成为可能,并迅速被应用于同时定位与地图构建(SLAM)系统中的在线地图构建。在这些系统中,必须在实时跟踪的同时逐步扩展和精炼高斯地图,因此初始化和密度控制直接决定了有限计算和迭代的分配。这与离线3DGS重建形成对比,后者可以在较长的优化周期内摊销此类启发式方法。然而,大多数3DGS-SLAM管道从离线重建中继承了初始化和密度控制的启发式方法,这在在线SLAM严格的每帧优化预算和增量地图增长下可能变得脆弱。在本研究中,我们在解耦的3DGS-SLAM环境中重新审视这些启发式方法,并提出三种在地图构建线程中操作的几何感知方法:保持透射率的密度增加、基于深度和内参的相机感知尺度初始化,以及聚焦于高残差区域的新原件的误差引导密度增加。我们的结果显示,在几乎没有额外开销的情况下,渲染质量持续改善,突显了在线SLAM中光度残差与位姿不确定性之间的耦合。我们将向社区开源我们的代码,以促进发展并验证可重复性。
cs.RO / 7 / 2608.14937

From Continuous Design to Delay-Aware Discrete Synthesis: Guaranteed High-Bandwidth Joint Control for PMSM Drives

从连续设计到延迟感知离散合成:保证高带宽的PMSM驱动联合控制
Fortunić, Edmundo Pozo, Yildirim, Mehmet C., Haddadin, Sami
Abstract
The increasing dynamic demands of modern robotic joints require current controllers to achieve high bandwidth over wide operating ranges of speed, acceleration, and torque, where communication, computation, and discrete-time effects can no longer be neglected. Conventional PMSM current controllers are typically designed in continuous time and subsequently discretized, leaving the sampling frequency and the impact of implementation delays largely to heuristic selection and iterative validation. This paper introduces a task-aware, delay-extended discrete-time joint model that explicitly accounts for physical communication and computation delays and enables direct synthesis of a discrete PI current controller with prescribed bandwidth and delay guarantees throughout the operating envelope. The framework analytically determines the minimum required sampling frequency, controller gains, and DC-link voltage needed to satisfy the specified motor and joint performance. Simulations across a range of dynamic requirements validate the methodology and demonstrate substantially reduced sampling-frequency and DC-link-voltage requirements compared with conventional continuous-time-based design. Experiments on a newly developed custom robotic joint further validate the proposed framework under real embedded implementation conditions.
Chinese Translation
现代机器人关节日益增长的动态需求要求电流控制器在广泛的速度、加速度和扭矩操作范围内实现高带宽,其中通信、计算和离散时间效应不再可以忽视。传统的PMSM电流控制器通常在连续时间下设计,随后进行离散化,这使得采样频率和实现延迟的影响在很大程度上依赖于启发式选择和迭代验证。本文引入了一种任务感知的延迟扩展离散时间关节模型,该模型明确考虑了物理通信和计算延迟,并能够直接合成具有规定带宽和延迟保证的离散PI电流控制器,适用于整个操作范围。该框架分析性地确定了满足指定电机和关节性能所需的最小采样频率、控制器增益和直流链路电压。针对一系列动态需求的仿真验证了该方法,并与传统的基于连续时间的设计相比,显著降低了采样频率和直流链路电压的要求。在新开发的定制机器人关节上的实验进一步验证了所提出框架在实际嵌入式实现条件下的有效性。
cs.RO / 8 / 2608.14944

SkillComposer: Learning Reusable Skills for Natural-Language Robot Programming

SkillComposer:学习可重用技能的自然语言机器人编程
Woods, John, Seifi, Hasti
Abstract
Natural-language interfaces can lower the barrier to programming robots, but existing systems struggle when users request complex tasks. While large language models (LLMs) perform well with simple commands, they often struggle to generate code for multi-step tasks, decompose high-level instructions, or reuse prior solutions. We present SkillComposer, an interactive natural-language robot programming system for simulation environments that continually learns reusable program abstractions. SkillComposer uses a generate-test architecture in which an LLM iteratively generates and revises robot programs before execution. Successful programs are stored and processed by an online library-learning algorithm that compresses recurring function sequences into reusable macro skills for future tasks. We evaluate SkillComposer through ablation experiments and a user study with 12 participants to determine its effectiveness on manipulation and robot caregiving tasks. The results show that evaluator-guided generation and learned abstractions improve success rates and usability while reducing user effort in natural-language robot programming.
Chinese Translation
自然语言接口可以降低编程机器人的门槛,但现有系统在用户请求复杂任务时表现不佳。尽管大型语言模型(LLMs)在简单命令上表现良好,但它们在为多步骤任务生成代码、分解高层指令或重用先前解决方案时常常遇到困难。我们提出了SkillComposer,一个用于仿真环境的交互式自然语言机器人编程系统,它不断学习可重用的程序抽象。SkillComposer采用生成-测试架构,其中LLM在执行之前迭代生成和修订机器人程序。成功的程序被存储,并由在线库学习算法处理,该算法将重复的功能序列压缩为可重用的宏技能,以便于未来的任务。我们通过消融实验和与12名参与者的用户研究评估SkillComposer,以确定其在操控和机器人护理任务上的有效性。结果表明,评估者引导的生成和学习的抽象提高了成功率和可用性,同时减少了用户在自然语言机器人编程中的努力。
cs.RO / 9 / 2608.14952

Evidence of Absence: Cross-Modal Abductive Risk Perception to Sustain World Models When Vision Fails

缺失的证据:跨模态推理风险感知以维持视觉失效时的世界模型
Xu, Cong, Sankar, Ravi
Abstract
A structured world-state (entities, relations, context, and predictive cues) is designed to preserve prediction-critical content when perception degrades, but it presumes observations to populate it; when the primary visual modality is occluded or degraded, those observations may be missing. We address how to sustain the world model from a complementary modality by treating the absence of expected co-evidence as evidence of a hidden cause. The abductive framework is modality-agnostic; this article instantiates it acoustically. A microphone-array front-end estimates the bearing of engine and tire sources and extracts approach-rate evidence (Doppler when a stable tone exists, a broadband looming readout otherwise); the event "signature present, visual co-evidence absent" then triggers abductive inference of a hidden road user, emitting a calibrated risk advisory rather than a control command. Recoverability of the hidden state is analyzed as an identifiability question separating shared from modality-unique information, and cueing is cast as Neyman-Pearson detection under an explicit false-alarm budget. On real occluded-approach recordings at blind junctions, the method warns a mean 1.7 seconds before line-of-sight entry, matches the sustained-window variant of the published acoustic baseline's detection rate with 42% fewer false alarms, localizes to 3.4 degrees median once in view, is well calibrated (expected calibration error 0.034), and keeps hazard awareness above 0.87 under staged vision degradation that collapses a vision-only channel to 0.03. We also measure the method's limits: calibration transfers to an unseen junction almost losslessly, the signature classifier does not, and moving-ego noise is the binding deployment constraint.
Chinese Translation
一个结构化的世界状态(实体、关系、上下文和预测线索)旨在在感知退化时保留预测关键内容,但它假设有观察结果来填充该状态;当主要视觉模态被遮挡或退化时,这些观察结果可能会缺失。我们探讨如何通过将预期的共同证据缺失视为隐藏原因的证据,从补充模态中维持世界模型。该推理框架与模态无关;本文在声学上实例化了它。一个麦克风阵列前端估计发动机和轮胎声源的方位,并提取接近率证据(当存在稳定音调时为多普勒效应,否则为宽带逼近读数);事件“特征存在,视觉共同证据缺失”随后触发对隐藏道路使用者的推理,发出经过校准的风险建议,而不是控制命令。隐藏状态的可恢复性被分析为一个可识别性问题,区分共享信息与模态独特信息,线索被视为在明确的虚警预算下的Neyman-Pearson检测。在盲交叉口的真实遮挡接近录音中,该方法在视线进入前平均警告1.7秒,匹配已发布声学基线的持续窗口变体的检测率,同时减少42%的虚警,定位精度为3.4度中位数,一旦进入视野,校准良好(预期校准误差为0.034),并在视觉退化阶段保持危险意识高于0.87,尽管视觉仅通道崩溃至0.03。我们还测量了该方法的限制:校准几乎无损地转移到未见的交叉口,特征分类器则不然,移动自我噪声是绑定的部署约束。
cs.RO / 10 / 2608.14986

GaussMemory: Task-Driven 3D Gaussian Scene Memory for Long-Horizon Robotic Manipulation

GaussMemory:用于长时间范围机器人操作的任务驱动3D高斯场景记忆
Hu, Zhiqiang, Huang, Shouren, Ishikawa, Masatoshi
Abstract
Long-horizon robotic manipulation fundamentally relies on persistent spatial memory. However, existing 3D memory systems function merely as passive recorders: they store observations using fixed, hand-crafted rules, treating every scene element--whether a critical grasp target or an irrelevant background wall--with equal importance. In this paper, we propose a paradigm shift from passive storage to active, task-driven spatial memory. We argue that a robot's memory should not simply record what it sees, but actively learn how to remember--discovering which objects to track precisely, how aggressively to update them, and what to discard, all learned end-to-end without hand-designed rules. Crucially, this active paradigm is realized by unifying memory update and readout as two sides of the same cognitive process, enabling bidirectional flow where task needs shape update strategies and vice versa. To instantiate this vision, we introduce GaussMemory, which leverages 3D Gaussian Splatting as a persistent geometric substrate. On LIBERO, GaussMemory outperforms MemoryVLA on Goal and Long-10; on VLABench, it surpasses $\pi_0$-FAST by +5.2% (Track 1) and +6.0% (Track 6).
Chinese Translation
长时间范围的机器人操作基本上依赖于持久的空间记忆。然而,现有的3D记忆系统仅作为被动记录器:它们使用固定的手工设计规则存储观察,将每个场景元素——无论是关键的抓取目标还是无关的背景墙——视为同等重要。在本文中,我们提出了一种从被动存储到主动、任务驱动空间记忆的范式转变。我们认为,机器人的记忆不应仅仅记录它所看到的内容,而应主动学习如何记忆——精确发现哪些物体需要跟踪、更新的力度如何以及哪些内容需要丢弃,这一切都通过端到端的学习实现,而无需手工设计的规则。关键是,这种主动范式通过将记忆更新和读取统一为同一认知过程的两个方面来实现,使得任务需求能够塑造更新策略,反之亦然。为了实现这一愿景,我们引入了GaussMemory,它利用3D高斯溅射作为持久的几何基础。在LIBERO上,GaussMemory在Goal和Long-10任务上超越了MemoryVLA;在VLABench上,它在Track 1和Track 6中分别超越了$ ext{π}_0$-FAST +5.2%和+6.0%。
cs.RO / 11 / 2608.14996

HP2-SLAM: Adaptive Hybrid ICP for Robust and Efficient LiDAR SLAM

HP2-SLAM:用于鲁棒高效LiDAR SLAM的自适应混合ICP
Tran, Nam, Tran, Thu, Phan, Hieu, Luu, Thai, Nguyen, Toan, Beksi, William J., Dang, Tuan
Abstract
Achieving robustness, accuracy, and efficiency simultaneously remains a central challenge in light detection and ranging (LiDAR) simultaneous localization and mapping (SLAM). While learning-based approaches deliver strong benchmark performance, they often require extensive training, substantial computational resources, and struggle to generalize to unseen or degenerate environments. Geometry-based methods are efficient and interpretable, yet their performance degrades in planar or repetitive scenes due to limitations of standard iterative closest point (ICP) formulations. We present HP2-SLAM, a minimalist yet robust LiDAR SLAM framework built around a neighborhood-size adaptive hybrid ICP. Our key insight is a planarity-aware adaptive threshold that dynamically classifies correspondences based on local geometric structure and density, thereby enabling a principled balance between point-to-plane and point-to-point residuals. This formulation stabilizes alignment in both structured and degenerate environments without feature engineering, learning modules, or dataset-specific tuning. Integrated into a complete SLAM pipeline with submap management, loop closure detection, and pose graph optimization, HP2-SLAM consistently outperforms strong geometry-based baselines across publicly available datasets while maintaining real-time performance on commodity hardware. Our results demonstrate that carefully designed geometric adaptation can achieve strong generalization and robustness without sacrificing simplicity or efficiency.
Chinese Translation
在光学探测与测距(LiDAR)同时定位与地图构建(SLAM)中,同时实现鲁棒性、准确性和效率仍然是一个核心挑战。尽管基于学习的方法在基准测试中表现出色,但它们通常需要大量的训练、巨大的计算资源,并且在未见或退化环境中难以泛化。基于几何的方法高效且可解释,但由于标准迭代最近点(ICP)公式的局限性,其在平面或重复场景中的性能会下降。我们提出了HP2-SLAM,这是一个围绕邻域大小自适应混合ICP构建的简约而鲁棒的LiDAR SLAM框架。我们的关键见解是一个考虑平面性的自适应阈值,它根据局部几何结构和密度动态分类对应点,从而实现点到平面和点到点残差之间的原则性平衡。该公式在结构化和退化环境中稳定对齐,无需特征工程、学习模块或数据集特定的调优。HP2-SLAM集成在一个完整的SLAM管道中,具备子地图管理、回环检测和位姿图优化,在公开可用的数据集上始终优于强几何基线,同时在普通硬件上保持实时性能。我们的结果表明,精心设计的几何适应可以在不牺牲简单性或效率的情况下实现强泛化和鲁棒性。
cs.RO / 12 / 2608.15002

NPU Offloading of a Frozen Visual Encoder for Robot Policy Training

用于机器人策略训练的冻结视觉编码器的NPU卸载
Yun, Hyojun, Won, Seungjae, Moon, Hyungpil
Abstract
When a robot policy is trained for a new task or dataset, its visual encoder can be frozen and only its action generation module trained, reducing training cost. Freezing removes the encoder's backward pass, but its forward pass must still run at every training step because the input images change, so it keeps consuming GPU compute. We therefore ask whether moving this computation to a low power AI accelerator such as an NPU can reduce total energy despite the added data transfer and longer training time, and how it affects policy performance. We built an asynchronous training pipeline that uses both a GPU and an NPU for the AR-Actor specialist. The frozen visual encoder runs in A8W8 INT8 on a Mobilint Aries2 NPU, while the FP32 action expert is trained on an NVIDIA GeForce RTX 5060 Ti GPU. We compared a GPU-only baseline with four conditions, L1 to L4, which gradually extend NPU offloading from one to four Transformer encoder layers. Each condition was trained for 30,000 steps with three random seeds. We measured GPU board power for the GPU-only condition and combined GPU and NPU board power for the NPU conditions. Energy per sample decreased by 17.1% in L1, which offloaded ResNet18 and the first encoder layer, and by 27.9% in L4, which offloaded ResNet18 and all four encoder layers. In contrast, training time per sample increased by 15.2% in L1 and 37.7% in L4, and peak allocated GPU memory decreased by 19.8 to 20.7%. The 15 resulting policies were each evaluated with the same 300 environment seeds, for a total of 4,500 simulator rollouts. The combined success rate was 93.33% for GPU-only and 91.44 to 92.89% for the NPU conditions. These results show that NPU offloading of a frozen visual encoder can reduce training energy, but it increases training time and lowers policy success rate by 0.44 to 1.89 percentage points compared with GPU-only training.
Chinese Translation
当为新任务或数据集训练机器人策略时,可以冻结其视觉编码器,仅训练其动作生成模块,从而降低训练成本。冻结操作消除了编码器的反向传播,但由于输入图像会变化,其前向传播仍需在每个训练步骤中运行,因此仍然消耗GPU计算资源。因此,我们探讨将这一计算转移到低功耗AI加速器(如NPU)是否能够在增加数据传输和延长训练时间的情况下减少总能耗,以及这对策略性能的影响。我们构建了一个异步训练管道,使用GPU和NPU共同为AR-Actor专家服务。冻结的视觉编码器在Mobilint Aries2 NPU上以A8W8 INT8格式运行,而FP32动作专家则在NVIDIA GeForce RTX 5060 Ti GPU上进行训练。我们将仅使用GPU的基线与四个条件(L1到L4)进行比较,这些条件逐步将NPU卸载从一个扩展到四个Transformer编码器层。每个条件训练了30,000步,使用了三个随机种子。我们测量了仅GPU条件下的GPU板功率,以及NPU条件下的GPU和NPU板功率的结合。L1条件下每个样本的能耗减少了17.1%,该条件卸载了ResNet18和第一个编码器层,而L4条件下减少了27.9%,该条件卸载了ResNet18和所有四个编码器层。相比之下,L1条件下每个样本的训练时间增加了15.2%,L4条件下增加了37.7%,而峰值分配的GPU内存减少了19.8%到20.7%。这15个策略在同样的300个环境种子下进行了评估,总共进行了4,500次模拟回合。仅GPU条件的综合成功率为93.33%,而NPU条件的成功率为91.44%到92.89%。这些结果表明,冻结视觉编码器的NPU卸载可以减少训练能耗,但与仅使用GPU的训练相比,它增加了训练时间,并使策略成功率降低了0.44到1.89个百分点。
cs.RO / 13 / 2608.15009

ForceU-VLA: A Force-Aware Vision-Language-Action Model for Embodied Ultrasound Scanning

ForceU-VLA:一种面向力感知的视觉-语言-动作模型用于具身超声扫描
Wu, Xingzheng, Zhang, Cheng, Yan, Guihao, Hu, Xifeng, Liu, Zhi, Cai, Qing
Abstract
Embodied intelligent ultrasound scanning enables the automation and standardization of the ultrasound examination process by integrating perception, decision-making, and execution capabilities. However, existing methods suffer from loosely coupled modeling between force and ultrasound modalities and lack awareness of scanning stages, which limits their ability to capture dynamic probe-tissue interactions. To address these issues, we propose ForceU-VLA, a force-aware Vision-Language-Action model for autonomous embodied ultrasound scanning, which leverages force signals and ultrasound image feedback throughout the scanning process to enable accurate and high-quality ultrasound acquisition. Firstly, we propose a Force-Ultrasound Synergistic Fusion Module (FUSFM) that synergistically fuses ultrasound visual and force-feedback information to provide stable, reliable guidance for probe motion. Secondly, a Stage-Adaptive Modulation Mechanism (SAMM) is proposed to accommodate the task requirements across different scanning stages by adaptively modulating multimodal features to enhance their representation quality. Additionally, we introduce ForceU-VLA-Data, a real-world, force-aware embodied ultrasound dataset that integrates visual, force, and action signals, including data from two organs across five representative clinical scanning views, and comprising 450 expert-collected trajectories with approximately 100,000 synchronized multimodal frames. Extensive experimental results demonstrate that ForceU-VLA significantly improves contact stability and probe pressure regulation in embodied ultrasound scanning, thereby effectively enhancing task execution quality and overall system reliability. The source code is available at https://github.com/VMVLab/ForceU-VLA.
Chinese Translation
具身智能超声扫描通过整合感知、决策和执行能力,实现了超声检查过程的自动化和标准化。然而,现有方法在力与超声模态之间的建模上存在松散耦合的问题,并且缺乏对扫描阶段的意识,这限制了它们捕捉动态探头与组织相互作用的能力。为了解决这些问题,我们提出了ForceU-VLA,一种面向力感知的视觉-语言-动作模型,用于自主具身超声扫描,该模型在整个扫描过程中利用力信号和超声图像反馈,以实现准确和高质量的超声采集。首先,我们提出了一种力-超声协同融合模块(FUSFM),该模块协同融合超声视觉和力反馈信息,为探头运动提供稳定、可靠的指导。其次,提出了一种阶段自适应调制机制(SAMM),通过自适应调制多模态特征以增强其表示质量,来适应不同扫描阶段的任务需求。此外,我们引入了ForceU-VLA-Data,这是一个面向力感知的具身超声数据集,集成了视觉、力和动作信号,包括来自两个器官的五个代表性临床扫描视图的数据,包含450条专家收集的轨迹和大约100,000帧同步的多模态数据。大量实验结果表明,ForceU-VLA显著提高了具身超声扫描中的接触稳定性和探头压力调节,从而有效提升了任务执行质量和整体系统可靠性。源代码可在 https://github.com/VMVLab/ForceU-VLA 获取。
cs.RO / 14 / 2608.15024

MotionGS-SLAM: Event-Modulated Gaussian Splatting for Motion-Blur Robust SLAM

MotionGS-SLAM:运动模糊鲁棒的事件调制高斯溅射SLAM
Hu, Zhiqiang, Huang, Shouren, Ishikawa, Masatoshi
Abstract
Current Vision-based SLAM systems fail catastrophically when motion blur corrupts the visual input, as they attempt the ill-posed inverse problem of recovering sharp content from degraded observations. We present MotionGS-SLAM, which fundamentally reimagines motion blur handling through a paradigm shift: rather than removing blur artifacts, we reformulate the challenge as a well-constrained forward problem that generatively models blur formation within the rendering pipeline. By leveraging event cameras' microsecond temporal resolution and immunity to motion blur, we introduce a novel event-modulated Gaussian kernel that dynamically adapts each Gaussian's rasterization based on precise motion cues. Our dual-modulation mechanism transforms 2D Gaussian projections from isotropic dots into anisotropic, motion-aligned elliptical brush strokes (spatial modulation) while adaptively varying exposure integral sampling density based on local velocity (temporal modulation). This physics-based approach enables joint optimization of intra-exposure camera trajectories and 3D scene geometry through blur-aware photometric and event-based constraints. Extensive experiments demonstrate significant improvements over state-of-the-art methods in trajectory accuracy and map quality under severe high-motion conditions.
Chinese Translation
当前基于视觉的SLAM系统在运动模糊影响视觉输入时会出现严重失败,因为它们试图解决一个不适定的逆问题,即从退化的观测中恢复清晰内容。我们提出了MotionGS-SLAM,它通过范式转变从根本上重新构想了运动模糊的处理:与其去除模糊伪影,我们将挑战重新表述为一个良好约束的前向问题,在渲染管道中生成性地建模模糊的形成。通过利用事件相机的微秒级时间分辨率和对运动模糊的免疫性,我们引入了一种新颖的事件调制高斯核,动态调整每个高斯的光栅化,基于精确的运动线索。我们的双重调制机制将2D高斯投影从各向同性点转变为各向异性、运动对齐的椭圆笔触(空间调制),同时根据局部速度自适应地变化曝光积分采样密度(时间调制)。这种基于物理的方式使得通过模糊感知的光度和基于事件的约束,能够对曝光内相机轨迹和3D场景几何进行联合优化。大量实验表明,在严重高运动条件下,我们的方法在轨迹精度和地图质量上显著优于最先进的方法。
cs.RO / 15 / 2608.15026

PACE: Phase-Progress-Aware Credit for Long-Horizon Embodied Manipulation

PACE:面向阶段进展的长时间跨度具身操作信用分配
Song, Chengye, Zhang, Jiawei, Song, Rui, Wang, Shengqi, Zhang, Xiangrong, Wang, Ziyi, Zhou, Huanbin, Wang, Hongzhou
Abstract
Post-training of vision-language-action (VLA) models typically relies on expert demonstrations and policy interaction trajectories. However, in long-horizon manipulation, a single episode often spans hundreds of control steps and multiple phases, while success or failure is only revealed at episode termination. Policy improvement therefore requires step-level credit signals to distinguish behaviors that advance the task from those that stall or regress. We present PACE, a credit-assignment framework for post-training on long-horizon manipulation, centered on a phase-progress-aware critic. PACE consists of two key modules: (1) the Global-Local Cooperative Value-Correction Critic (GLC-Critic) aggregates visual and motion-difference features within local temporal windows to infer the phase and intra-phase progress of each step, and applies residual correction to a discretized remaining-cost distribution accordingly, enabling step-level credit assignment; (2) Progressive Policy Distillation (PPD) converts credit into positive and negative conditions via task-wise thresholds and trains a credit-conditioned action generation policy: it first protects the pretrained policy with high-credit positive samples, then incorporates all positive and negative credits to learn the quality boundary, and at inference amplifies high-credit behaviors through the difference between conditional outputs. Extensive simulation experiments and diverse real-world robotic-arm experiments demonstrate that PACE consistently achieves significant improvements over the strongest baseline.
Chinese Translation
视觉-语言-动作(VLA)模型的后训练通常依赖于专家演示和策略交互轨迹。然而,在长时间跨度的操作中,单个情节往往跨越数百个控制步骤和多个阶段,而成功或失败仅在情节终止时显现。因此,策略改进需要逐步的信用信号,以区分推进任务的行为与停滞或退步的行为。我们提出了PACE,一种针对长时间跨度操作的后训练信用分配框架,重点是阶段进展感知的评论者。PACE由两个关键模块组成:(1)全局-局部协作价值修正评论者(GLC-Critic)在局部时间窗口内聚合视觉和运动差异特征,以推断每一步的阶段和阶段内进展,并相应地对离散化的剩余成本分布应用残差修正,从而实现逐步信用分配;(2)渐进式策略蒸馏(PPD)通过任务特定的阈值将信用转换为正面和负面条件,并训练一个基于信用的动作生成策略:它首先用高信用的正样本保护预训练策略,然后结合所有正负信用以学习质量边界,并在推理时通过条件输出之间的差异放大高信用行为。广泛的仿真实验和多样的真实世界机器人臂实验表明,PACE在最强基线之上始终实现了显著的改进。
cs.RO / 16 / 2608.15088

Max-Q Selective Imitation for Human-in-the-Loop Online Robot Learning

基于Max-Q选择模仿的人机协同在线机器人学习
Wang, Zihang, Wang, Yishan
Abstract
Human-in-the-loop (HIL) online reinforcement learning for real robots must absorb human interventions quickly while continuing to improve beyond the human prior. We present a training method for this setting based on two components. First, an \emph{MC Q-chunk} critic regresses chunk-level action values onto Monte Carlo returns from the replay buffer, performing sample-average (behavior) policy evaluation so that intervention trajectories are credited directly rather than diluted by current-policy TD backups. Second, \emph{max-Q selective imitation} updates the actor by imitating, at each state, the higher-$Q$ action between the current policy action and a buffer sample under a hard winner-take-all rule. This rule automatically switches between learning from interventions and on-policy self-improvement: when the autonomous policy is stronger, targets align with the policy distribution, reducing the policy--target-sample gap that otherwise induces execution-time distribution shift. In practice we score candidates with a standard critic ensemble mean to reduce comparison noise, without softening targets or introducing score-gap thresholds. On a real USB pick-and-insertion task with 20 demonstrations, ACT QChunk-MCBC attains 99\% success within 30 minutes of HIL training, whereas HIL-SERL requires about 5 hours to converge. In simulation on Peg Insertion and Square, ACT/Flow Q-chunk variants similarly reach $\ge$96\% success within roughly half an hour of effective training, outperforming HIL-SERL, EXPO, and E2HiL on the success--time frontier.
Chinese Translation
人机协同(HIL)在线强化学习对于真实机器人必须快速吸收人类干预,同时在超越人类先验的基础上持续改进。我们提出了一种基于两个组成部分的训练方法。首先, extit{MC Q-chunk}评论员将块级动作值回归到来自重放缓冲区的蒙特卡洛回报,执行样本平均(行为)策略评估,以便直接记入干预轨迹,而不是被当前策略的时间差分(TD)备份稀释。其次, extit{max-Q选择模仿}通过在每个状态下模仿当前策略动作和缓冲区样本之间的高-$Q$动作,使用严格的赢家通吃规则来更新演员。该规则自动在从干预中学习和基于策略的自我改进之间切换:当自主策略更强时,目标与策略分布对齐,减少了否则会导致执行时分布转移的策略-目标样本差距。在实践中,我们使用标准评论员集成均值对候选者进行评分,以减少比较噪声,而不需要软化目标或引入评分差阈值。在一个包含20个演示的真实USB抓取与插入任务中,ACT QChunk-MCBC在30分钟的HIL训练内达到了99%的成功率,而HIL-SERL则需要大约5小时才能收敛。在Peg Insertion和Square的仿真中,ACT/Flow Q-chunk变体同样在大约半小时的有效训练内达到了$ extgreater$96%的成功率,超越了HIL-SERL、EXPO和E2HiL在成功-时间边界上的表现。
cs.RO / 17 / 2608.15139

StructRL: Structured Action-Space Exploration for Flow-Based VLAs

StructRL:基于结构的动作空间探索用于流动型视觉-语言-动作(VLA)
Yang, Jiarui, Zhu, Bin, Chen, Jingjing, Zou, Na, Fu, Yanwei, Zhu, Jianggang, Jiang, Yu-Gang
Abstract
Flow-based Vision-Language-Action (VLA) models are now widely used for continuous robotic manipulation, and online reinforcement learning (RL) is emerging as a key technique for adapting them to new tasks. Existing RL methods typically inject stochasticity inside the denoising chain, often through isotropic or temporally independent noise. However, effective robot exploration calls for structured noise: temporally smooth and scaled differently across action groups. We show that simply switching the in-chain noise to a structured form does not suffice: noise added at an intermediate flow time can be weakened by the remaining denoising steps before execution, a phenomenon we call \emph{Structured Noise Dilution}. We propose \textbf{StructRL}, which avoids dilution by relocating policy stochasticity to the action space via three coupled choices: (i) a deterministic ODE decoder, (ii) structured noise injected directly in the action space, and (iii) last-step replay, where policy-gradient updates avoid assigning likelihoods to intermediate denoising states. This keeps structured exploration tied to the executed action while providing a tractable training signal for the flow decoder. Across three flow-based VLA models on multiple simulated manipulation benchmarks and two real-world tasks, StructRL improves exploration efficiency and OOD performance over prior in-chain baselines, demonstrating the effectiveness of structured action-space exploration for adapting flow-based VLA with RL. \textbf{Project page:} https://flyfaerss.github.io/structrl/
Chinese Translation
基于流动的视觉-语言-动作(VLA)模型现已广泛应用于连续机器人操作,而在线强化学习(RL)正成为将其适应于新任务的关键技术。现有的RL方法通常在去噪链中注入随机性,通常通过各向同性或时间独立的噪声。然而,有效的机器人探索需要结构化噪声:在时间上平滑且在不同动作组之间具有不同的缩放。我们表明,仅仅将链内噪声切换为结构化形式并不足够:在中间流动时间添加的噪声可能会在执行前被剩余的去噪步骤削弱,这一现象我们称之为“结构化噪声稀释”(Structured Noise Dilution)。我们提出了 extbf{StructRL},通过三种耦合选择将策略随机性重新定位到动作空间,从而避免稀释:(i)确定性常微分方程(ODE)解码器,(ii)直接在动作空间中注入结构化噪声,以及(iii)最后一步重放,其中策略梯度更新避免将可能性分配给中间去噪状态。这保持了结构化探索与执行的动作相联系,同时为流动解码器提供了可处理的训练信号。在多个模拟操作基准和两个真实世界任务上的三种基于流动的VLA模型中,StructRL提高了探索效率和OOD性能,相较于先前的链内基线,证明了基于结构的动作空间探索在使用RL适应流动型VLA中的有效性。 extbf{项目页面:} https://flyfaerss.github.io/structrl/
cs.RO / 18 / 2608.15156

Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models

低秩动态:用于学习世界模型中反事实展开的有效潜在载体
Liu, Yang, Chen, Yuming
Abstract
World models may predict the future without making clear which parts of their hidden state actually drive those predictions. We ask whether a small, directly addressable hidden-state change can place a learned world model on the intended counterfactual trajectory and then let the model continue that future on its own. We study a recurrent world model with a 192-dimensional hidden state in a controlled two-object, two-dimensional collision environment. For a bounded family of local velocity edits, we first verify that the model can natively represent and roll out the edited future. We then construct candidate low-rank carriers from training-only factual-to-counterfactual hidden differences and learn a map from the factual state and requested edit to carrier coefficients. On the registered rank grid, rank 4 is the smallest tested rank that satisfies the full development-panel criteria. A single rank-4 patch at the anchor is sufficient to redirect a 12-step autonomous rollout, with no future observations, teacher forcing, or repeated correction. The frozen procedure satisfies the preregistered replication rule across independently trained checkpoints and remains usable across nearby intervention times. Random equal-norm, wrong-object, and wrong-time controls do not explain the effect. A position-edit stress test provides a negative contrast: the intended position patch can pass the raw rollout criteria, but no-patch and random controls can pass the same criteria, and wrong-object specificity is not established. Thus, successful editing alone is not enough. We use dynamics-effective to describe an intervention that changes the model's future computation in a sustained and target-specific way under autonomous rollout. The rank-4 result identifies a compact intervention interface for the tested velocity-edit family, not a closed four-dimensional state or an intrinsic state dimension.
Chinese Translation
世界模型可以预测未来,但并未明确指出其隐藏状态的哪些部分实际上驱动了这些预测。我们探讨一个小的、可直接访问的隐藏状态变化是否可以将学习到的世界模型置于预期的反事实轨迹上,然后让模型自主继续这一未来。我们研究了一个具有192维隐藏状态的递归世界模型,应用于一个受控的二维碰撞环境,其中包含两个物体。对于一个有界的局部速度编辑家族,我们首先验证模型能够本地表示并展开编辑后的未来。然后,我们从仅通过训练获得的事实到反事实的隐藏差异中构建候选低秩载体,并学习从事实状态和请求编辑到载体系数的映射。在注册的秩网格上,秩4是满足完整开发面板标准的最小测试秩。在锚点处,一个单一的秩4补丁足以重定向一个12步的自主展开,而无需未来观察、教师强制或重复校正。该冻结程序满足预注册的复制规则,适用于独立训练的检查点,并在附近的干预时间内保持可用。随机的等范数、错误物体和错误时间控制并不能解释这一效果。位置编辑压力测试提供了一个负对比:预期的位置补丁可以通过原始展开标准,但无补丁和随机控制也能通过相同标准,而错误物体的特异性并未得到确立。因此,单靠成功的编辑并不足够。我们使用“动态有效”来描述一种干预,该干预以持续且目标特定的方式改变模型的未来计算,适用于自主展开。秩4的结果为测试的速度编辑家族识别了一个紧凑的干预接口,而不是一个封闭的四维状态或内在状态维度。
cs.RO / 19 / 2608.15175

LAPF: LLM-Agent-Based Path Finder Using the UAVScenes Dataset

LAPF:基于大型语言模型代理的路径寻找器,使用UAVScenes数据集
Emami, Yousef, Homaei, Mohammadhossein, Zhou, Hao, Gaitán, Miguel Gutiérrez, Arani, Atefeh Hajijamali, Zhang, Rui
Abstract
Uncrewed aerial vehicles (UAVs) are increasingly deployed for autonomous navigation in complex outdoor environments, where dynamic conditions and mission requirements require intelligent adaptive decision-making. Existing optimization-based, Machine Learning (ML), and Reinforcement Learning (RL) approaches often rely on predefined models or task-specific training, limiting their generalization and adaptability in uncertain scenarios. Recent Large Language Model (LLM)-assisted approaches offer promising reasoning capabilities but remain constrained by limited agentic functionality, including insufficient memory, planning, and tool interaction mechanisms.This paper proposes an LLM-Agent-Based Path Finder (LAPF) framework for autonomous UAV navigation in town-scale outdoor environments. LAPF extends LLM-assisted navigation by integrating perception, memory, planning, and action modules into a closed-loop cognitive architecture. The proposed agent leverages prior navigation experiences, performs Chain-of-Thought (CoT) reasoning, couples each detected hazard to a bounded corrective action, and dynamically refines waypoint decisions based on environmental feedback.The three independent trials per method demonstrate that LAPF achieves mean path lengths of 512.83 m and 506.37 m, compared to the straight-line optimum of 497.33 m, corresponding to path length reductions of 17.2% and 15.6% relative to CoT prompting and absolute path efficiencies of 97.1% and 98.1% in open-field and obstacle-injected scenarios, respectively. Furthermore, LAPF is the only evaluated approach that couples every detected hazard to a bounded, metric-neutral corrective action while maintaining near-goal stability, with zero clamp events in both scenarios, whereas CoT prompting increases from 9.7 to 14.0 events.
Chinese Translation
无人机(UAV)在复杂的户外环境中越来越多地被部署用于自主导航,这些动态条件和任务需求要求智能自适应决策。现有的基于优化的机器学习(ML)和强化学习(RL)方法通常依赖于预定义模型或特定任务训练,限制了它们在不确定场景中的泛化能力和适应性。最近的基于大型语言模型(LLM)的辅助方法提供了有前景的推理能力,但仍然受到有限的代理功能的约束,包括内存不足、规划和工具交互机制。本论文提出了一种基于LLM代理的路径寻找器(LAPF)框架,用于城镇规模的户外环境中的无人机自主导航。LAPF通过将感知、记忆、规划和行动模块集成到一个闭环认知架构中,扩展了LLM辅助导航。所提出的代理利用先前的导航经验,执行思维链(CoT)推理,将每个检测到的危险与一个有限的纠正行动相结合,并根据环境反馈动态优化航点决策。每种方法的三次独立试验表明,LAPF的平均路径长度为512.83米和506.37米,相较于497.33米的直线最优路径,分别实现了17.2%和15.6%的路径长度缩减,并在开放场地和障碍物注入场景中分别达到97.1%和98.1%的绝对路径效率。此外,LAPF是唯一评估的方法,它将每个检测到的危险与一个有限的、度量中立的纠正行动相结合,同时保持近目标的稳定性,在两种场景中均未发生夹紧事件,而CoT提示事件从9.7增加到14.0。
cs.RO / 20 / 2608.15269

Remember Smarter: Visual History Compressor and Hyperbolic Experience Space for Robotic Memory

更智能的记忆:用于机器人记忆的视觉历史压缩器和双曲体验空间
Zhou, Dai, Yan, Jiexi, Li, Tong, Wang, Yuxuan, Deng, Cheng
Abstract
Long-horizon robot policies require compact access to recent observations and reusable experience without expanding the vision-language-action (VLA) context. We introduce Remember Smarter (RS), a plug-and-play module with complementary visual-history and hyperbolic experience-memory branches. Its visual branch compresses multi-view patch histories using bidirectional spatial Mamba and causal temporal Mamba, then exposes the resulting memory to action-facing hidden states through residual cross-attention while leaving the VLM visual-token stream unchanged. Its experience branch stores successful final-layer VLM states in a Poincare VAE space, organizes them hierarchically, and asynchronously converts retrieved experience into geodesic prompt tokens without blocking action inference. When adapted to pi0, RS increases total success on LIBERO-Plus from 53.6% to 70.6% and achieves substantial performance gains in real-robot experiments designed to evaluate memory retention and experience utilization.
Chinese Translation
长时间跨度的机器人策略需要对近期观察结果的紧凑访问和可重用经验,而不扩展视觉-语言-动作(VLA)上下文。我们提出了更智能的记忆(Remember Smarter, RS),这是一个即插即用模块,具有互补的视觉历史和双曲体验记忆分支。其视觉分支使用双向空间 Mamba 和因果时间 Mamba 压缩多视角补丁历史,然后通过残差交叉注意力将生成的记忆暴露给面向动作的隐藏状态,同时保持 VLM 视觉标记流不变。其体验分支将成功的最终层 VLM 状态存储在 Poincare VAE 空间中,进行分层组织,并异步将检索到的经验转换为测地线提示标记,而不阻塞动作推理。当适配到 pi0 时,RS 在 LIBERO-Plus 上的总成功率从 53.6% 提高到 70.6%,并在旨在评估记忆保留和经验利用的真实机器人实验中取得了显著的性能提升。
cs.RO / 21 / 2608.15284

VTInstructor: Visual Trajectory Prompting for Navigation Instruction Generation in Continuous Environments

VTInstructor:用于连续环境中导航指令生成的视觉轨迹提示
Yang, Haolin, Long, Yuxing, Yang, Zihan, Dong, Hao
Abstract
Navigation instruction generation from ego-centric RGB video in continuous environments is an important yet challenging task for human-robot interaction and scalable dataset construction. Prior instruction generators assume discrete viewpoint graphs with panoramic observations, where trajectory structure is explicit; in continuous environments, however, the agent receives only a dense RGB stream, making trajectory cues difficult to recover. We propose VTInstructor, the first VLN instruction generation framework for continuous environments. Our key idea is to convert implicit trajectory geometry into explicit visual trajectory prompts: EDTC condenses long RGB trajectories into navigation-critical keyframes, VTP overlays path, turn, and goal cues onto these anchors, VTMod injects the resulting trajectory signals into the visual encoder, and VT-GRPO further calibrates this spatial injection during training, all without requiring a navigation graph, pre-built map, or scene reconstruction. On the challenging R2R-CE and RxR-CE Val Unseen benchmarks, VTInstructor sets a new state of the art across all standard NLG metrics, surpassing the strongest baseline by +0.357 CIDEr and +0.109 CIDEr, respectively. Beyond automatic metrics, VTInstructor-generated instructions raise a frozen follower's success rate to 63.3%, a +14.7 percentage-point gain over the best competing instruction source, and provide consistent data augmentation gains of +3 SR points on downstream navigation tasks.
Chinese Translation
从自我中心的RGB视频中生成连续环境的导航指令是一项重要但具有挑战性的任务,涉及人机交互和可扩展数据集构建。先前的指令生成器假设存在具有全景观察的离散视点图,其中轨迹结构是显式的;然而,在连续环境中,代理仅接收密集的RGB流,使得轨迹线索难以恢复。我们提出了VTInstructor,这是第一个针对连续环境的视觉导航(VLN)指令生成框架。我们的关键思想是将隐式轨迹几何转换为显式的视觉轨迹提示:EDTC将长RGB轨迹浓缩为导航关键帧,VTP在这些锚点上叠加路径、转弯和目标线索,VTMod将生成的轨迹信号注入视觉编码器,而VT-GRPO在训练过程中进一步校准这种空间注入,所有这些都不需要导航图、预构建地图或场景重建。在具有挑战性的R2R-CE和RxR-CE Val Unseen基准测试中,VTInstructor在所有标准自然语言生成(NLG)指标上设定了新的最先进水平,分别超越最强基线+0.357 CIDEr和+0.109 CIDEr。除了自动指标外,VTInstructor生成的指令使得一个冻结的跟随者的成功率提高到63.3%,比最佳竞争指令源提高了14.7个百分点,并在下游导航任务中提供了+3 SR点的一致数据增强收益。
cs.RO / 22 / 2608.15285

PhaseLoRA: Control-Regime-Conditioned Low-Rank Adaptation for Continuous-Action Vision-Language-Action Policies

PhaseLoRA:基于控制状态条件的低秩适应用于连续动作视觉-语言-动作策略
Guo, Yufei, Wu, Yinan, Duan, Haoran, Ding, Guiguang, Han, Jungong
Abstract
Parameter-efficient fine-tuning (PEFT) is a natural way to adapt pretrained vision-language-action (VLA) policies, but most adapter designs apply temporally static updates throughout a control rollout, overlooking the phase-dependent nature of continuous-action manipulation. Such policies traverse distinct regimes, including approach, contact transition, grasping, transport, and placement, each requiring different adaptation behaviors. We propose \textbf{PhaseLoRA}, a lightweight LoRA parameterization that conditions adaptation at each action-chunk prediction step using two weakly supervised descriptors: fine-control tendency and event/boundary intensity. PhaseLoRA modulates the LoRA left factor in the action expert, allowing the effective low-rank update direction to vary over time while keeping the backbone largely frozen. On LIBERO, PhaseLoRA improves average success rate by 12.2 points over a matched-parameter high-rank LoRA baseline and outperforms stronger LoRA variants. Ablations show that random temporal modulation and scalar gating do not reproduce the performance of the full model, while update-direction analyses reveal structured temporal variation associated with the predicted control descriptors. These results establish within-trajectory conditioning as an effective lightweight PEFT axis for continuous-action VLA policies.
Chinese Translation
参数高效微调(PEFT)是一种自然的方式来适应预训练的视觉-语言-动作(VLA)策略,但大多数适配器设计在控制执行过程中应用时间上静态的更新,忽视了连续动作操作的相位依赖特性。这些策略经历不同的阶段,包括接近、接触过渡、抓取、运输和放置,每个阶段需要不同的适应行为。我们提出了 extbf{PhaseLoRA},一种轻量级的LoRA参数化方法,它在每个动作块预测步骤中使用两个弱监督描述符进行适应:精细控制倾向和事件/边界强度。PhaseLoRA调节动作专家中的LoRA左因子,使得有效的低秩更新方向随时间变化,同时保持主干网络基本不变。在LIBERO上,PhaseLoRA的平均成功率比匹配参数的高秩LoRA基线提高了12.2个百分点,并且优于更强的LoRA变体。消融实验表明,随机时间调制和标量门控无法重现完整模型的性能,而更新方向分析揭示了与预测控制描述符相关的结构化时间变化。这些结果确立了轨迹内条件作为连续动作VLA策略的有效轻量级PEFT轴。
cs.RO / 23 / 2608.15289

SCORE: Shape-Conforming Regions for Flight in Enclosed, Degraded Environments

SCORE:适应封闭退化环境飞行的形状一致区域
Kim, Eric Minwoo, Kim, Jong-Kook
Abstract
Autonomous UAVs enter enclosed environments such as caves and collapsed structures that confine the vehicle and degrade perception. Conformal prediction provides a distribution-free guarantee by calibrating how far an obstacle keep-out must expand to absorb perception error at a target coverage level. However, existing keep-out regions use convex primitives whose bulges consume narrow passages and grow as perception degrades. Our main contribution defines the nonconformity score on a signed distance field (SDF). This produces a non-convex keep-out that tightly follows obstacle geometry and avoids the unnecessary bulging of equal-margin convex regions. Two supporting components keep this geometry usable as perception degrades. First, a voxelwise union of complementary sensor observations certifies voxels that any single sensor misses. Second, the margin around the obstacle adapts to measured visibility without weather labels or the online ground-truth feedback that single-pass flight cannot provide. Results on real subterranean data show that the resulting distribution-free, shape-conforming keep-out retains more usable free space than convex baselines at the same certified coverage, and produces safer closed-loop flight.
Chinese Translation
自主无人机进入如洞穴和倒塌结构等封闭环境,这些环境限制了飞行器的活动并降低了感知能力。形状一致预测通过校准障碍物的安全距离扩展程度,为目标覆盖水平提供了无分布保证,以吸收感知误差。然而,现有的安全区域使用凸形原件,其凸起部分会占用狭窄通道,并随着感知能力的降低而增大。我们主要的贡献是在带符号距离场(SDF)上定义了非一致性评分。这产生了一个非凸的安全区域,紧密跟随障碍物的几何形状,避免了等边凸区域不必要的膨胀。两个支持组件使得在感知能力下降时,这种几何形状仍然可用。首先,互补传感器观测的体素级联合验证了任何单个传感器遗漏的体素。其次,障碍物周围的边距根据测量的可见性进行调整,而无需天气标签或单次飞行无法提供的在线真实反馈。基于真实地下数据的结果表明,所得到的无分布、形状一致的安全区域在相同认证覆盖率下保留了比凸基线更多的可用自由空间,并实现了更安全的闭环飞行。
cs.RO / 24 / 2608.15437

MM-BEV: Enhancing Timeliness by Computing Where and When it Matters

MM-BEV:通过计算重要时刻和地点提升时效性
Liu, Liangkai, Shin, Kang G.
Abstract
Multimodal bird's-eye-view (BEV) perception combines LiDAR depth accuracy with dense camera semantics, but its high computational cost and imperfect sensing conditions make real-time deployment challenging. Existing methods largely compress individual detectors and overlook three opportunities: structured sparsity within camera and LiDAR inputs, timing misalignment between modalities, and the fact that many detected objects do not affect the planner's immediate action. We present MM-BEV, a real-time multimodal BEV system guided by a simple principle: compute where and when it matters. MM-BEV divides perception into mandatory work for safety-critical objects within braking distance of the ego vehicle and with short time-to-collision (TTC), and optional work for less urgent regions. It prioritizes mandatory work and reduces or sheds optional work under tight compute budgets. MM-BEV integrates four mechanisms: (1) a criticality-ranked temporal ROI selector based on motion-extrapolated detections from prior frames; (2) sparse, ROI-aware feature extraction using shared-shape camera crops at context-adaptive resolution and ROI-aware LiDAR voxelization; (3) a latency-aware coordinator that adapts LiDAR sweeps, image resolution, and keyframes according to scene dynamics and TTC; and (4) an asynchronous scheduler that decouples sensing from inference and skips stale frames. On nuScenes, MM-BEV reduces inference latency by 1.96x and end-to-end latency by 2.93x, with no loss in geometry-critical recall and only a 0.2 percentage-point drop in safety-critical recall. On a Clearpath Husky A300 equipped with an Ouster-128 LiDAR, BEV cameras, and a Jetson AGX Orin, MM-BEV further reduces mean latency by 2.11x, demonstrating its potential for real-world autonomous systems.
Chinese Translation
多模态鸟瞰视图(BEV)感知结合了LiDAR的深度精度和密集的相机语义,但其高计算成本和不完美的感知条件使得实时部署面临挑战。现有方法主要压缩单个检测器,忽视了三个机会:相机和LiDAR输入中的结构稀疏性、模态之间的时间错位,以及许多检测到的物体并不影响规划者的即时行动。我们提出了MM-BEV,一个实时多模态BEV系统,其指导原则是:在重要时刻和地点进行计算。MM-BEV将感知分为对安全关键物体的强制性工作(这些物体在自车的制动距离内且具有短时间碰撞(TTC))和对不太紧急区域的可选工作。它优先处理强制性工作,并在计算预算紧张时减少或放弃可选工作。MM-BEV集成了四个机制:(1)基于先前帧的运动外推检测结果的关键性排名时间感兴趣区域(ROI)选择器;(2)使用共享形状相机裁剪以上下文自适应分辨率进行稀疏、ROI感知的特征提取,以及ROI感知的LiDAR体素化;(3)一个延迟感知协调器,根据场景动态和TTC调整LiDAR扫描、图像分辨率和关键帧;(4)一个异步调度器,将感知与推理解耦,并跳过过时的帧。在nuScenes数据集上,MM-BEV将推理延迟减少了1.96倍,端到端延迟减少了2.93倍,几何关键召回率没有损失,安全关键召回率仅下降了0.2个百分点。在配备Ouster-128 LiDAR、BEV相机和Jetson AGX Orin的Clearpath Husky A300上,MM-BEV进一步将平均延迟减少了2.11倍,展示了其在现实世界自主系统中的潜力。
cs.RO / 25 / 2608.15440

Accelerating Mixed Discrete-Continuous Motion Planning via Neural Graphs of Convex Sets

通过凸集的神经图加速混合离散-连续运动规划
Trivedi, Ananya, Prajapati, Sarvesh, Jaffar, Mohamed Khalid M, Xu, Zhexin, Rosen, David, Padir, Taskin
Abstract
Motion planning problems such as collision-free navigation and contact-rich manipulation can be naturally formulated as optimization problems that couple discrete decisions with continuous trajectories. The Graphs of Convex Sets (GCS) framework offers a practical solution to these problems. It represents discrete decisions as nodes of a graph and encodes continuous trajectories in the edges connecting them. However, the resulting optimization subproblems can become computationally prohibitive for online replanning. In this work, we propose a learning-based strategy to mitigate this limitation. Specifically, we replace the costly convex relaxation step required by nominal GCS with a single forward pass through a Graph Attention Network that predicts a set of highly probable candidate paths through the graph. A lightweight ranking network then orders these candidates by their estimated trajectory cost. Evaluating them in this order, we terminate our search early while still recovering a near-optimal motion plan. We validate the resulting pipeline across diverse robotic tasks, including collision-free motion planning for a 3D quadrotor and a 7-DoF manipulator, and planning through contact for planar pushing. Across both convex and non-convex cost and constraint settings, our approach yields up to two orders of magnitude speedup over nominal GCS while maintaining a 100% success rate, at the cost of some suboptimality in the recovered solutions. Code implementations and video demonstrations can be found at https://neural-gcs.github.io/.
Chinese Translation
运动规划问题,如无碰撞导航和富接触操控,可以自然地被表述为将离散决策与连续轨迹结合的优化问题。凸集图(Graphs of Convex Sets, GCS)框架为这些问题提供了一个实用的解决方案。它将离散决策表示为图的节点,并在连接它们的边中编码连续轨迹。然而,所产生的优化子问题在在线重规划时可能变得计算上不可承受。在本研究中,我们提出了一种基于学习的策略来缓解这一限制。具体而言,我们用一个通过图注意力网络(Graph Attention Network)进行的单次前向传播来替代名义GCS所需的高成本凸松弛步骤,该网络预测通过图的高度可能的候选路径。然后,一个轻量级的排序网络根据估计的轨迹成本对这些候选路径进行排序。按照这个顺序评估它们,我们在仍能恢复近似最优运动规划的同时提前终止搜索。我们在多种机器人任务中验证了所得到的管道,包括3D四旋翼的无碰撞运动规划和7自由度(7-DoF)操控器的规划,以及平面推送中的接触规划。在凸和非凸成本及约束设置下,我们的方法在保持100%成功率的同时,相较于名义GCS实现了高达两个数量级的加速,尽管在恢复的解决方案中存在一些次优性。代码实现和视频演示可以在 https://neural-gcs.github.io/ 找到。
cs.RO / 26 / 2608.15446

GUIDER: Evaluating Goal-Free Human Intent Inference for Teleoperated Manipulation on Real-Robot Data

GUIDER:评估无目标人类意图推断在真实机器人数据上的遥操作操控
Kenny, Nicholas, Contreras, Cesar Alan, Ouedraogo, Basile, Stolkin, Rustam, Chiou, Manolis, Kyrarini, Maria
Abstract
This paper presents an evaluation of a goal-free probabilistic framework for human intent inference during robotic manipulation. We deploy the Global User Intent Dual-phase Estimation for Robots (GUIDER) on data collected from a robotic arm to test the manipulation phase across various assistance scenarios, including making tea and fetching medicine. To support operation, we add online probability updates, workspace limits, support-plane filtering, and a grasping mode that prioritizes feasible grasp regions, all of which are tested on the recorded data while preserving its original temporal conditions. Across 20 manipulation steps in three scenarios, GUIDER estimated human intent within the correct grasp-candidate set in all cases and achieved a time to confident prediction of 3.7 s, a remaining time before first grasp of 49.6 s, a prediction stability of 96.4%, and a runtime of 4.857/4.474 s (mean/median) per perceptual phase of intent.
Chinese Translation
本文评估了一种无目标的概率框架,用于在机器人操控过程中推断人类意图。我们在从机器人手臂收集的数据上部署了全球用户意图双相估计(Global User Intent Dual-phase Estimation for Robots,GUIDER),以测试在多种辅助场景下的操控阶段,包括泡茶和取药。为了支持操作,我们增加了在线概率更新、工作空间限制、支持平面过滤和优先考虑可行抓取区域的抓取模式,所有这些都在记录的数据上进行了测试,同时保持其原始时间条件。在三个场景中的20个操控步骤中,GUIDER在所有情况下都在正确的抓取候选集内估计了人类意图,并实现了3.7秒的自信预测时间、49.6秒的首次抓取前剩余时间、96.4%的预测稳定性,以及每个意图感知阶段的运行时间为4.857/4.474秒(均值/中位数)。
cs.RO / 27 / 2608.15461

Detachable Wire Drive : Reconfigurable Robot Architecture with Shared Actuators

可拆卸线驱动:具有共享执行器的可重构机器人架构
Hattori, Takahiro, Kawaharazuka, Kento, Okada, Kei
Abstract
Reconfigurable robots offer significant potential for adapting to diverse tasks; however, conventional centralized architectures often require dedicated actuators for each module, leading to substantial increases in overall system weight, volume, and cost. To address these challenges, this paper presents the "Detachable Wire Drive," a reconfigurable robotic system that enables the sharing of heavy and expensive actuators across various morphologies. The core of this system is the "Wire Detach Unit," a mechanism designed to physically split and reconnect wire drive paths, allowing motors to be consolidated into a common base unit. We demonstrate the versatility of this approach by developing a 2-DOF rigid arm, a continuum arm, and two distinct grippers, all of which are interchangeably attached to, and driven by, a single shared actuator set. Experimental results validate the mechanical reliability of the detachment process and the control framework's ability to seamlessly manage transitions between configurations, highlighting a path toward more efficient and multi-functional robotic systems.
Chinese Translation
可重构机器人在适应多样化任务方面具有显著潜力;然而,传统的集中式架构通常需要为每个模块配备专用执行器,从而导致整体系统重量、体积和成本的显著增加。为了解决这些挑战,本文提出了“可拆卸线驱动”系统,这是一种可重构机器人系统,能够在不同形态之间共享重型和昂贵的执行器。该系统的核心是“线拆卸单元”,一种旨在物理上分离和重新连接线驱动路径的机制,允许将电机整合到一个公共基础单元中。我们通过开发一个2自由度刚性臂、一个连续臂和两个不同的抓手,展示了这种方法的多样性,所有这些组件都可以互换连接,并由一组共享的执行器驱动。实验结果验证了拆卸过程的机械可靠性以及控制框架无缝管理配置之间过渡的能力,突显了朝着更高效和多功能机器人系统发展的路径。
cs.RO / 28 / 2608.15490

Vision-Based Tactile Intelligence for Robotics: Sensing, Learning, and Embodied Manipulation

基于视觉的机器人触觉智能:感知、学习与具身操作
Zhou, Peng, Hu, Jun, Chen, Sihan, Zhang, Zeqing, Ma, Haofei, Lu, Zhenyu, Liu, Sichao, Wang, Xueqian, Zheng, Pai, Li, Xiang, Luo, Shan, Pan, Jia, Navarro-Alarcon, David, Yang, Chenguang, Wang, Michael Yu
Abstract
Tactile sensing is essential for robots in contact-rich tasks, yet many tactile sensors still provide sparse, low-dimensional signals that do not capture sufficient information for complex robotic perception and interaction. Vision-based tactile sensors (VBTSs) offer a powerful alternative by con-verting contact-induced deformation of a soft interface into im-ages. The image-based formulation gives VBTSs high-resolution, information-rich tactile observations that enable complex robotic tasks. This review surveys the full VBTS pipeline and treats sensing hardware, learning methods, simulation, and datasets as an integrated sensing-and-learning system. We 1) organize representative VBTSs into a hardware taxonomy structured by deformable elastomer design, sensor size and shape, and optical system design to guide future sensor development; 2) present a hierarchical view of learning-based tactile intelligence from low-level signal understanding to task-level policies and foundation models; and 3) examine simulation platforms and tactile datasets as a scaling layer, together with sim-to-real transfer and cross-sensor adaptation for training, benchmarking, and deployment. Finally, we identify open challenges and future directions for VBTSs in robotics. By providing a holistic view of how hardware, AI architectures, simulation, and datasets interact, this review aims to advance tactile intelligence for contact-rich robotic tasks.
Chinese Translation
触觉感知对于在接触丰富的任务中工作的机器人至关重要,然而许多触觉传感器仍然提供稀疏的、低维度的信号,无法捕捉足够的信息以支持复杂的机器人感知和交互。基于视觉的触觉传感器(VBTS)通过将软界面的接触引起的变形转换为图像,提供了一种强有力的替代方案。基于图像的形式使VBTS能够获得高分辨率、信息丰富的触觉观测,从而支持复杂的机器人任务。本文回顾了完整的VBTS流程,并将感知硬件、学习方法、仿真和数据集视为一个集成的感知与学习系统。我们1)根据可变形弹性体设计、传感器的大小和形状以及光学系统设计,将代表性的VBTS组织成一个硬件分类,以指导未来的传感器开发;2)呈现基于学习的触觉智能的层次视图,从低级信号理解到任务级策略和基础模型;3)考察仿真平台和触觉数据集作为一个扩展层,以及用于训练、基准测试和部署的仿真到现实转移和跨传感器适应。最后,我们识别了VBTS在机器人领域面临的开放挑战和未来方向。通过提供硬件、人工智能架构、仿真和数据集之间相互作用的整体视图,本文旨在推动触觉智能在接触丰富的机器人任务中的应用。
cs.RO / 29 / 2608.15509

Temporal Logic Guided Universal Task Representations for Reinforcement Learning

基于时间逻辑引导的强化学习通用任务表示
Zhang, Hao, Zhou, Zhangli, Kan, Zhen
Abstract
Task guided agents demonstrate strong performance in a wide range of complex tasks. However, most existing task representation algorithms are tailored to specific contexts and struggle to generalize across diverse scenarios. Moreover, they typically depend on gradient signals from reinforcement learning controllers to update their weights, which can degrade both representation quality and learning efficiency. To overcome these limitations, we propose LOTUS, a temporal logic inspired universal task representation framework that can be seamlessly integrated into any RL algorithm to enhance agent performance across diverse task settings. Specifically, we design a novel task representation architecture capable of modeling relationships and extracting task semantics from LTL formulas. We further introduce a more effective update mechanism that treats the LTL encoder as a policy, thereby improving representation capacity. To enhance stability and robustness, LOTUS leverages the bisimulation metric, which provides theoretical guarantees for LTL representation, including behavioral equivalence, optimality fidelity, and trajectory robustness. Experimental results show that LOTUS outperforms most existing methods in learning efficiency, generalization capability, and representation quality. Specifically, LOTUS accelerates convergence over 20% in single-task scenarios, achieves a 15%-45% higher success rate in unseen manipulation tasks, and improves generalization performance over 25% in complex multi-task environments with increased sub-goal depth or conjunctions. The corresponding code, videos, and appendix are available at: https://lotus-website.github.io/.
Chinese Translation
任务引导的智能体在多种复杂任务中表现出色。然而,现有的大多数任务表示算法都是针对特定情境设计的,难以在多样化场景中进行泛化。此外,它们通常依赖于强化学习控制器的梯度信号来更新权重,这可能会降低表示质量和学习效率。为了解决这些局限性,我们提出了LOTUS,一个受时间逻辑启发的通用任务表示框架,可以无缝集成到任何强化学习算法中,以提升智能体在不同任务设置中的表现。具体而言,我们设计了一种新颖的任务表示架构,能够建模关系并从线性时序逻辑(LTL)公式中提取任务语义。我们进一步引入了一种更有效的更新机制,将LTL编码器视为策略,从而提高表示能力。为了增强稳定性和鲁棒性,LOTUS利用了双模拟度量,该度量为LTL表示提供了理论保证,包括行为等价性、最优性保真度和轨迹鲁棒性。实验结果表明,LOTUS在学习效率、泛化能力和表示质量方面优于大多数现有方法。具体而言,在单任务场景中,LOTUS加速收敛超过20%;在未见过的操作任务中,成功率提高15%-45%;在具有增加子目标深度或结合的复杂多任务环境中,泛化性能提高超过25%。相关代码、视频和附录可在以下网址获取:https://lotus-website.github.io/
cs.RO / 30 / 2608.15532

Degenerate in Whose Frame? An Equivariance Condition for Degeneracy Detection in LiDAR Registration

在谁的框架中退化?用于LiDAR配准的退化检测的等变性条件
Zhang, Yujie, Zhao, Chunlei, Lin, Yuzong, Guo, Yuxuan, Jia, Xiaohui, Liu, Jinyue
Abstract
Degeneracy detectors for LiDAR registration commonly return six per-axis binary labels. We ask whether these labels are properties of the scene. Under a body-frame change, the point-to-plane information matrix transforms by congruence, H' = Ad(T)^T H Ad(T), not similarity. Congruence preserves nullity and, through the adjoint reparameterization, identifies the same physical twist subspace; the per-axis footprint and a thresholded spectrum need not be invariant. In a noise-free circular tunnel, shifting the origin by one metre changes which degrees of freedom are flagged. A generalized criterion Hv = lambda Mv is universally frame-independent over positive-semidefinite information forms if and only if its metric rule is equivariant. No fixed metric qualifies, while a rig-adapted one exists only at zero screw pitch, met in one of nineteen surveyed calibrations. The equivariant point-displacement metric M = sum_i J_i^T J_i yields dimensionless, scene-scale-invariant generalized eigenvalues. They are invariant to body frame, consistent changes of length unit and scene scales; the threshold also transfers empirically across sequences. Across 365 frame pairs from four public sequences, labels rarely change at practical extrinsic magnitudes, yet a remapping estimator's correction differs between body-frame choices on 44.5-69.5% of pairs, with a median of 0.7-4.0 mm and a maximum of 0.87 m. The per-axis footprint changes even under the equivariant metric, placing the fundamental issue in the reported quantity.
Chinese Translation
LiDAR配准的退化检测器通常返回六个每轴的二元标签。我们询问这些标签是否是场景的属性。在身体框架变化下,点到平面信息矩阵通过同余变换,H' = Ad(T)^T H Ad(T),而不是相似性。同余保持零性,并通过伴随重参数化识别相同的物理扭转子空间;每轴的足迹和阈值谱不必保持不变。在一个无噪声的圆形隧道中,原点移动一米会改变哪些自由度被标记。一个广义标准Hv = lambda Mv在正半定信息形式下是普遍框架无关的,当且仅当其度量规则是等变的。没有固定的度量符合这一条件,而适应刚体的度量仅在零螺距时存在,这在调查的十九个标定中仅出现一次。等变的点位移度量M = sum_i J_i^T J_i产生无量纲、场景尺度不变的广义特征值。它们对身体框架、长度单位和场景尺度的一致变化是不变的;阈值在序列之间也可以经验性地转移。在来自四个公共序列的365对框架中,标签在实际外部量级下很少变化,但重映射估计器的修正在44.5-69.5%的对之间因身体框架选择而异,中位数为0.7-4.0毫米,最大值为0.87米。即使在等变度量下,每轴的足迹也会变化,这将基本问题置于报告的量中。
cs.RO / 31 / 2608.15541

Contact Modes Are Strata: What Geometric Structure Buys in Discrete-Continuous Planning

接触模式是层次:离散-连续规划中的几何结构的价值
Kyaw, Phone Thiha, Kelly, Jonathan
Abstract
Contact-rich manipulation poses a discrete question and a continuous one at once, namely which contacts are active and how to move while they hold. The two are coupled by a change of dimension, since each contact that a robot maintains confines its motion to a lower-dimensional manifold. We make that coupling the explicit object of planning by observing that a contact mode is not merely analogous to a stratum of the configuration space; it is one. A plan is then a walk over strata whose within-stratum segments are geodesics. On two contact-rich manipulation tasks in simulation, pushing a T-shaped block around obstacles and reorienting a cube in a dexterous hand, our planner returns solutions within seconds with no mode, contact sequence, or stratum given in advance.
Chinese Translation
接触丰富的操作同时提出了一个离散问题和一个连续问题,即哪些接触是活跃的,以及在这些接触保持的情况下如何移动。这两者通过维度的变化相互耦合,因为机器人维持的每一个接触都将其运动限制在一个低维流形上。我们通过观察接触模式不仅仅是配置空间的一个层次,而是一个层次,将这种耦合作为规划的明确对象。因此,一个计划是在层次上行走,其中层内段是测地线。在两个接触丰富的操作任务的仿真中,即在障碍物周围推动一个T形块和在灵巧手中重新定向一个立方体,我们的规划器在几秒钟内返回解决方案,而没有提前给出模式、接触序列或层次。
cs.RO / 32 / 2608.15549

MistyPilot: Enabling Social-Robot Control through Multi-Agent LLM Skill Orchestration

MistyPilot:通过多智能体大语言模型技能编排实现社交机器人控制
Wang, Xiao, Dong, Lu, Nwogu, Ifeoma, Setlur, Srirangaraj, Govindaraju, Venu
Abstract
Programming small social robots from natural-language instructions requires more than invoking isolated APIs. Interactive tasks combine reactive physical behaviors with stateful social behaviors, while existing interfaces often require developers to manually compose APIs into skills, configure their parameters, bind sensor events to skills, and manage task states at runtime. We present MistyPilot, a multi-agent LLM framework that interprets high-level natural-language instructions and orchestrates the corresponding skills on the Misty social robot. A Task Router dispatches each instruction to one of two specialized agents: a Physically Interactive Agent for sensor-triggered robot control and direct skill invocation, and a Social Interaction Agent for dialogue-oriented task-state management and context-dependent multimodal response generation. To improve efficiency, the Social Interaction Agent reuses previously generated results when applicable and invokes full generation otherwise. We evaluate MistyPilot on five component-level suites, with sensor bindings and skill invocations executed on the physical Misty robot, and a preliminary user study with 12 participants. MistyPilot attains high accuracy on routing, sensor-skill binding, task-state parsing, result reuse, and skill extension up to 100 skills, and lower variance than an otherwise identical single-agent baseline, while participants report positive perceptions of usability and interaction quality. The code will be made publicly available via the project page.
Chinese Translation
从自然语言指令编程小型社交机器人不仅仅是调用孤立的API。交互任务结合了反应式的物理行为和有状态的社交行为,而现有接口通常要求开发者手动将API组合成技能,配置其参数,将传感器事件绑定到技能,并在运行时管理任务状态。我们提出了MistyPilot,一个多智能体大语言模型框架,它解释高级自然语言指令并在Misty社交机器人上编排相应的技能。任务路由器将每个指令分派给两个专门的智能体之一:用于传感器触发的机器人控制和直接技能调用的物理交互智能体,以及用于对话导向的任务状态管理和上下文相关的多模态响应生成的社交互动智能体。为了提高效率,社交互动智能体在适用时重用先前生成的结果,否则进行完整生成。我们在五个组件级套件上评估了MistyPilot,传感器绑定和技能调用在物理Misty机器人上执行,并进行了为期初步的用户研究,参与者为12人。MistyPilot在路由、传感器-技能绑定、任务状态解析、结果重用和技能扩展(最多100个技能)方面达到了高准确率,并且其方差低于其他相同的单智能体基线,同时参与者对可用性和交互质量的感知积极。代码将通过项目页面公开发布。
cs.RO / 33 / 2608.15560

ReForce: Learning Force-aware Retargeting for Dexterous Manipulation

ReForce:学习力感知重定向以实现灵巧操作
Wu, Yuhang, Zeng, Lingqi, Jing, Changwei, Ye, Jianglong, Wang, Xiaolong
Abstract
Human demonstrations offer a scalable data source for dexterous manipulation, but transferring them to robot actions remains challenging due to the embodiment gap. Today's retargeting is mostly kinematic, yet manipulation is decided by force, which governs how the hand interacts with the object and how the object moves. In this paper, we present ReForce, a Force-aware Retargeting method that turns human motion and forces into robot actions that reproduce the intended contact. ReForce predicts a residual on the kinematically retargeted action to reach the desired force, using a general force tracker trained on large-scale simulation interactions. It supports both online force-aware teleoperation and offline data translation. In simulation and on real hardware, ReForce achieves lower force-tracking error and stronger multi-finger contact engagement on contact-rich tasks such as paper-cup grasping and tongs manipulation.
Chinese Translation
人类示范提供了一个可扩展的数据源用于灵巧操作,但将其转移到机器人动作上仍然面临挑战,因为存在体现差距。当前的重定向主要是运动学的,而操作则由力决定,力控制着手与物体的交互以及物体的运动。在本文中,我们提出了ReForce,一种力感知重定向方法,它将人类的运动和力转化为机器人动作,以再现预期的接触。ReForce预测在运动学重定向动作上的残差,以达到期望的力,使用在大规模仿真交互中训练的通用力跟踪器。它支持在线力感知遥操作和离线数据转换。在仿真和真实硬件上,ReForce在接触丰富的任务(如纸杯抓取和夹子操作)中实现了更低的力跟踪误差和更强的多指接触参与。
cs.RO / 34 / 2608.15573

Not All History Helps: Velocity-Aware Selective Memory for Long-Horizon End-to-End Autonomous Driving

并非所有历史都有帮助:基于速度的选择性记忆用于长时间跨度的端到端自动驾驶
Liu, Yuchen, Song, Ziying, Zhang, Shengkai, Chen, Jiannan, Wu, Peiliang, Yang, Lei, Sun, Bin, Gong, Yan, Wang, Li
Abstract
Reliable long-horizon planning remains a key challenge in end-to-end autonomous driving. By accounting for future motion evolution and potential consequences, it provides forward-looking guidance for safe and consistent driving in evolving traffic environments. Existing methods use historical planning states as temporal context. Self-generated history may become stale or conflict with the current motion stage, introducing unreliable priors. We propose StableDrive to address cross-cycle historical reliability and within-horizon motion-stage evolution. Selective Momentum Memory (SMM), implemented with a Mamba selective state-space operator, controls the influence of the preceding self-predicted planning state on the current cycle. Motion-Stage Training Scaffold (MSTS) uses motion-stage, long-horizon trajectory, and longitudinal-motion supervision to guide stage-aware future motion learning and is removed before inference. A fixed parameter midpoint between two architecture-aligned endpoints yields a single deployable SMM planner without model ensembling or extra inference-time computation. On nuScenes under the MomAD evaluation protocol, StableDrive achieves SOTA performance across all reported planning metrics from 1 to 6 s, reducing average collision rate by 23.3%, TPC by 30.9%, and L2 by 11.8% over the best previously reported value for each metric. On the curated Longitudinal-Transition nuScenes (LT-nuScenes), StableDrive reduces 6-s collision rate by 23.81%, TPC by 10.90%, and L2 by 6.37%. On NAVSIM v1 and v2, StableDrive achieves the highest PDMS/EPDMS in all three reported settings, including a 5.7-point EPDMS gain on v2 navhard over the previous best.
Chinese Translation
可靠的长时间跨度规划仍然是端到端自动驾驶中的一个关键挑战。通过考虑未来运动演变和潜在后果,它为在不断变化的交通环境中安全和一致的驾驶提供了前瞻性的指导。现有方法使用历史规划状态作为时间上下文。自生成的历史可能会变得过时或与当前运动阶段相冲突,从而引入不可靠的先验知识。我们提出了StableDrive,以解决跨周期历史的可靠性和在时间跨度内的运动阶段演变。选择性动量记忆(Selective Momentum Memory, SMM)通过Mamba选择性状态空间操作符实现,控制先前自预测规划状态对当前周期的影响。运动阶段训练支架(Motion-Stage Training Scaffold, MSTS)利用运动阶段、长时间跨度轨迹和纵向运动监督来指导阶段感知的未来运动学习,并在推理前移除。在两个架构对齐的端点之间的固定参数中点产生一个可单独部署的SMM规划器,无需模型集成或额外的推理时间计算。在nuScenes数据集上,根据MomAD评估协议,StableDrive在所有报告的规划指标中实现了SOTA性能,从1到6秒的平均碰撞率降低了23.3%,TPC降低了30.9%,L2降低了11.8%,相较于每个指标的最佳先前报告值。在精心策划的纵向过渡nuScenes(LT-nuScenes)上,StableDrive将6秒的碰撞率降低了23.81%,TPC降低了10.90%,L2降低了6.37%。在NAVSIM v1和v2上,StableDrive在所有三个报告设置中实现了最高的PDMS/EPDMS,包括在v2 navhard上相较于之前最佳值的5.7点EPDMS增益。
cs.RO / 35 / 2608.15636

Algorithm-Architecture Co-Design for Efficient VLA Inference via Speculative Inference and Verification

通过推测推理和验证实现高效的视觉-语言-行动(VLA)推理的算法-架构协同设计
Qi, Chunyu, Song, Zhuoran, Weng, Jian, Jiang, Haozhe, Liu, Xueyuan, Jing, Naifeng, He, Guanghui, Liang, Xiaoyao, Guan, Haibing
Abstract
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in the field of embodied AI, but their high computational cost and limited predicted action length hinder real-time deployment. Although Dadu-Corki, a dedicated accelerator for efficient embodied AI, has been introduced, it does not exploit the inherent interaction patterns between the robot and its environment, which results in a relatively short predicted action length. We observe that robotic environments naturally alternate between active states-where precise actions are crucial-and inactive states-where actions have limited impact on task success. This insight enables a new scheduling opportunity: long-action-length speculative prediction in inactive states, paired with selective verification in active states. We propose SpecVLA, an algorithm-system co-design framework that adaptively balances action length, inference latency, and task reliability. On the algorithm side, SpecVLA introduces a state-aware VLA inference execution paradigm and a hardware-friendly construction of a smaller verification model (sVLA) using differential residuals and block-wise mixed-precision quantization. On the system side, we develop a heterogeneous architecture consisting of a GPU and a robotic-specific hardware module, along with a speculative dataflow that decouples VLA and sVLA through parallel execution. Comprehensive evaluations on OpenVLA and RDT across LIBERO and ManiSkill benchmarks show that SpecVLA reduces end-to-end latency significantly while preserving task success rate. By enabling long-action-length speculative prediction with timely verification, SpecVLA achieves real-time robotic manipulation with both high efficiency and reliability.
Chinese Translation
视觉-语言-行动(VLA)模型在具身人工智能领域展现了显著的能力,但其高计算成本和有限的预测行动长度阻碍了实时部署。尽管已经引入了专为高效具身人工智能设计的加速器Dadu-Corki,但它并未充分利用机器人与环境之间的固有交互模式,导致预测行动长度相对较短。我们观察到,机器人环境自然在主动状态(精确动作至关重要)和非主动状态(动作对任务成功的影响有限)之间交替。这一洞察为新的调度机会提供了可能:在非主动状态下进行长行动长度的推测预测,并在主动状态下进行选择性验证。我们提出了SpecVLA,一个算法-系统协同设计框架,能够自适应地平衡行动长度、推理延迟和任务可靠性。在算法方面,SpecVLA引入了一种状态感知的VLA推理执行范式,并利用差分残差和块级混合精度量化构建了一个硬件友好的较小验证模型(sVLA)。在系统方面,我们开发了一种异构架构,由GPU和特定于机器人硬件模块组成,并设计了一种推测数据流,通过并行执行解耦VLA和sVLA。针对LIBERO和ManiSkill基准的OpenVLA和RDT的全面评估显示,SpecVLA显著降低了端到端延迟,同时保持了任务成功率。通过实现长行动长度的推测预测并及时进行验证,SpecVLA实现了高效且可靠的实时机器人操作。
cs.RO / 36 / 2608.15680

Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation

Robo-Dopamine 2.0:历史条件与OOD感知的过程奖励建模用于机器人操作
Xu, Yijie, Jin, Haopeng, Zhou, Run, Liu, Shengbang, Chen, Sixiang, Cheng, Hongyang, Hu, Sicheng, Co, Peterson, Luo, Jinwen, Tan, Huajie, Zhang, Shanghang
Abstract
Vision-language-action (VLA) models improve robotic manipulation but remain vulnerable to compounding errors, scene changes, and off-trajectory states. Reinforcement learning can refine pretrained VLA policies, yet sparse success signals hinder exploration, while engineered dense rewards are costly and task-specific. Existing learned visual reward models often rely on static before-after observations, causing temporal ambiguity and weak discrimination between robustness-preserving variations and task-invalid failures under out-of-distribution (OOD) execution. We introduce Robo-Dopamine 2.0, a history- and OOD-aware process reward model with a pairwise prediction interface. It combines (1) history-conditioned pairwise rewards that use source-aligned reference panels for synthetic OOD queries and observed rollout history for online queries, while preserving the queried endpoints, and (2) an OOD-aware signed progress space that represents valid progress, robustness, failure, and recovery. A Signed-Hop Curriculum with transition-aware replay learns coarse execution ordering before fine-grained progress calibration. We also construct an OOD trajectory dataset and a five-family benchmark. Reference panels improve mean visual order consistency (VOC) from 0.967 to 0.986 and OOD-robust VOC from 0.906 to 0.958. With the same 400K pairwise-reward budget, Signed-Hop training with 25% replay reaches 0.9872 mean VOC, compared with 0.9858 for a matched-pool shuffled control. In downstream reinforcement learning, the full model achieves 86.8% mean RoboTwin success and 71/80 successful real-world insertions.
Chinese Translation
视觉-语言-动作(VLA)模型提高了机器人操作的能力,但仍然容易受到累积错误、场景变化和偏离轨迹状态的影响。强化学习可以优化预训练的VLA策略,但稀疏的成功信号阻碍了探索,而工程化的密集奖励则成本高昂且任务特定。现有的学习视觉奖励模型通常依赖于静态的前后观察,导致时间模糊性以及在分布外(OOD)执行下对保持鲁棒性变化和任务无效失败之间的弱辨别能力。我们提出了Robo-Dopamine 2.0,这是一种具有历史和OOD感知的过程奖励模型,配备成对预测接口。它结合了(1)历史条件的成对奖励,利用源对齐的参考面板进行合成OOD查询,并使用观察到的回滚历史进行在线查询,同时保持查询的端点,以及(2)一个OOD感知的有符号进展空间,表示有效进展、鲁棒性、失败和恢复。带有过渡感知重放的有符号跳跃课程在细粒度进展校准之前学习粗略的执行顺序。我们还构建了一个OOD轨迹数据集和一个五个家庭的基准。参考面板将平均视觉顺序一致性(VOC)从0.967提高到0.986,将OOD鲁棒VOC从0.906提高到0.958。在相同的40万成对奖励预算下,带有25%重放的有符号跳跃训练达到了0.9872的平均VOC,而匹配池随机控制的结果为0.9858。在下游强化学习中,完整模型实现了86.8%的平均RoboTwin成功率和71/80次成功的真实世界插入。
cs.RO / 37 / 2608.15707

GAINS: Leveraging Inconsistent Human Intervention Signals in Reinforcement Learning

GAINS:利用强化学习中的不一致人类干预信号
Zhang, Xinyi, Zhao, Yinuo, Ren, Pei, Jiang, Lechun, Jin, Huiqian, Sun, Lei, Wu, Dapeng, Che, Zhengping, Liu, Chi Harold, Tang, Jian
Abstract
Correcting robot manipulation policies through human intervention holds great promise for real-world deployment, yet human operators are inherently imperfect in both the actions they provide and the timing of their intervention signals. While the former has been extensively discussed in reinforcement learning (RL), the latter remains underexplored. At high control frequencies, human intervention signals are often delayed and inconsistent across time and state space. In this work, we present GAINS, a framework for leveraging inconsistent human intervention signals in RL. At the core of GAINS, we employ distributional RL with quantile Q-networks to model the return variability induced by sparse task rewards and inconsistent human interventions. Building on this distributional representation, we introduce a pessimistic exploration strategy that promotes safe and sample-efficient learning under human corrections. We evaluate GAINS on four diverse simulated manipulation tasks and two challenging real-world scenarios against state-of-the-art intervention-based methods. GAINS achieves a 22% higher task success rate than RLIF and improves recovery success by up to 43% in failure scenarios. These results highlight the importance of modeling return variability induced by human imperfection for real-world deployment of intervention-based learning.
Chinese Translation
通过人类干预来修正机器人操作策略在实际部署中具有很大潜力,但人类操作员在提供的动作和干预信号的时机上本质上是有缺陷的。虽然前者在强化学习(RL)中得到了广泛讨论,但后者仍然未被充分探索。在高控制频率下,人类干预信号往往会出现延迟,并且在时间和状态空间中表现出不一致性。在本研究中,我们提出了GAINS,一个用于在强化学习中利用不一致人类干预信号的框架。GAINS的核心是采用分布式强化学习(distributional RL)与分位数Q网络(quantile Q-networks),以建模稀疏任务奖励和不一致人类干预所引起的回报变异性。在此分布式表示的基础上,我们引入了一种悲观探索策略,促进在人工修正下的安全和样本高效学习。我们在四个不同的模拟操作任务和两个具有挑战性的现实场景中评估了GAINS,并与最先进的基于干预的方法进行了对比。GAINS的任务成功率比RLIF高出22%,并在失败场景中提高了多达43%的恢复成功率。这些结果强调了建模人类不完美所引起的回报变异性在基于干预的学习的实际部署中的重要性。
cs.RO / 38 / 2608.15741

Some Modifications to Our End-to-End UAV Planner

对我们的端到端无人机规划器的一些修改
Lu, Junjie, Tian, Bailing
Abstract
The one-stage planner YOPO maps a single depth image and the robot state directly to a set of candidate trajectories, trained by backpropagating through differentiable trajectory costs. This yields dense, geometrically informative supervision, but inherits the pathologies of soft-constrained optimization: the safety cost competes with the smoothness and goal-reaching terms, is non-convex across homotopy classes, and the single-piece polynomial is limited in expressiveness. In this report, we summarize several effective modifications. We adopt a two-piece MINCO parameterization, trading time for smoothness without altering the trajectory's spatial profile. We further lift YOPO's multi-modal prediction to span distinct homotopy classes, treating each motion primitive as a homotopy anchor that confines the trajectory to a feasible basin - without explicit safe-flight-corridor construction or front-end search. For dynamic feasibility, we impose barrier penalties on velocity and acceleration together with a curvature-dependent speed limit whose gradient acts only on the velocity, producing an adaptive-speed behavior that decelerates in cluttered regions or sharp turns. We replace score regression with a ranking loss, preventing small score errors from reordering the candidate set. These yield richer trajectory representations, safer obstacle avoidance, and more direct flight paths.
Chinese Translation
一阶段规划器 YOPO 直接将单个深度图像和机器人状态映射到一组候选轨迹,通过反向传播可微分轨迹成本进行训练。这提供了密集的、几何上信息丰富的监督,但继承了软约束优化的病态特征:安全成本与平滑性和目标到达项相竞争,在同伦类中是非凸的,并且单一多项式在表达能力上受到限制。在本报告中,我们总结了几项有效的修改。我们采用了双段 MINCO 参数化,牺牲时间以换取平滑性,而不改变轨迹的空间轮廓。我们进一步提升 YOPO 的多模态预测,以跨越不同的同伦类,将每个运动原语视为一个同伦锚点,限制轨迹在可行的盆地内——无需显式构建安全飞行走廊或前端搜索。为了动态可行性,我们对速度和加速度施加障碍惩罚,并结合一个依赖于曲率的速度限制,其梯度仅作用于速度,从而产生一种自适应速度行为,在杂乱区域或急转弯时减速。我们用排名损失替代了得分回归,防止小的得分误差重新排序候选集。这些修改产生了更丰富的轨迹表示,更安全的障碍规避,以及更直接的飞行路径。
cs.RO / 39 / 2608.15748

Making two action heads agree: coordination mechanisms and a runtime collapse certificate for flow-matching policies

使两个动作头达成一致:流匹配策略的协调机制和运行时崩溃证书
Sun, Jinhui, Zhou, Wei, Yang, Bowen, Xiao, Xinliang, Yang, Li
Abstract
A dual-representation flow-matching policy decodes each predicted motion into joint and end-effector spaces, and the residual between the two kinematically equivalent decodings provides a physically interpretable runtime signal. On multimodal tasks, however, independently sampled branches may choose different valid modes, causing false alarms. We study how to coordinate the two branches and at what cost. Across two robot environments and a non-robotic testbed, the tested mechanisms fall into four classes. An auxiliary latent shared by both branches but absent from the flow-matching construction is erased at the population optimum, a provable dead end confirmed within a prespecified 2% equivalence band. Sharing source noise can coordinate or anti-coordinate: its effect changes sign with the representation map and tracks the alignment of decoder mode basins. Consistency regularization gives intermediate coordination but reduces the valid-pair rate, while training-supported discrete partitions achieve near-ceiling coordination robustly. We further derive a chance-corrected coordination bound based only on each branch's Gini-Simpson diversity, yielding an attainable region and a label-free certificate that separates coordination from collapse when zero mismatch is ambiguous. On LIBERO-Plus, benign multimodality adds 1.57 percentage points of false alarms to the residual, which remains the strongest evaluated failure signal; the preregistered token intervention does not meet its false-alarm criterion or produce a seed-robust detection change. Code, models, and per-run configurations are available at https://github.com/kimo423/dual-head-coordination.
Chinese Translation
双重表示的流匹配策略将每个预测的运动解码为关节和末端执行器空间,而两种运动学上等效解码之间的残差提供了一个物理可解释的运行时信号。然而,在多模态任务中,独立采样的分支可能选择不同的有效模式,从而导致误报。我们研究如何协调这两个分支及其成本。在两个机器人环境和一个非机器人测试平台中,测试的机制分为四类。一个由两个分支共享但在流匹配构造中缺失的辅助潜变量在总体最优时被消除,这是在预设的2%等效带内确认的可证明的死胡同。共享源噪声可以协调或反协调:其效果随着表示映射的变化而改变符号,并跟踪解码器模式盆地的对齐情况。一致性正则化提供了中间协调,但降低了有效配对率,而训练支持的离散划分则稳健地实现了接近上限的协调。我们进一步推导了一个仅基于每个分支的基尼-辛普森多样性的机会校正协调界限,得出了一个可达区域和一个无标签证书,当零不匹配模糊时将协调与崩溃分开。在LIBERO-Plus上,良性的多模态性使残差增加了1.57个百分点的误报,这仍然是评估的最强失败信号;预注册的标记干预未能满足其误报标准或产生种子鲁棒检测变化。代码、模型和每次运行的配置可在 https://github.com/kimo423/dual-head-coordination 获取。
cs.RO / 40 / 2608.15766

Tac4Loco: Learning Spatiotemporal Plantar Pressure Representations for Humanoid Locomotion

Tac4Loco:用于类人行走的时空足底压力表示学习
Liu, Ziyun, Guo, Sikai, Li, Zheng, Cao, Jiahang, Liu, Haichao, Qu, Pei, Zhang, Yinghong, Zhou, Jinni, Ma, Jun
Abstract
Humanoid robots are expected to traverse complex terrains, where the plantar support may vary dramatically due to foot placement errors, ground properties, and transient dynamics. To achieve robust locomotion, the robots are required to adapt to uneven terrain and uncertain foot--ground interactions. Existing locomotion policies rely primarily on proprioception or exteroceptive terrain perception, where the former provides only indirect evidence of plantar support, while the latter predicts contact conditions before touchdown but cannot observe the actual support in real-time. Although some studies incorporate plantar contacts as an auxiliary perception, they rely mainly on summary statistics, overlooking the spatial topology of plantar pressure, which provides a more direct characterization of the realized contact state. To bridge this gap, we present Tac4Loco, a tactile-perceptive framework that incorporates multi-array plantar pressure as direct feedback for humanoid locomotion. We formulate a topology-preserving ordinal representation to map simulated and physical sensor signals into a shared observation space, with a dual-branch encoder for extracting their spatial and temporal representations. Subsequently, the learned spatiotemporal features are integrated with augmented proprioception including terrain estimation cues, and provided to an asymmetric actor-critic architecture for policy learning. Extensive simulation and real-world experiments demonstrate improved tracking performance and support adaptation on terrains with inclined, partial, asymmetric, and changing support. We further demonstrate its zero-shot deployment on unseen compliant and unstructured terrains, including a foam platform and a gravel road. All code and experimental configurations will be released as open-source to facilitate reproducibility.
Chinese Translation
类人机器人被期望能够在复杂地形中行走,其中足底支撑可能由于脚的位置错误、地面特性和瞬态动态而发生显著变化。为了实现稳健的行走,机器人需要适应不平坦的地形和不确定的足-地面交互。现有的行走策略主要依赖于本体感觉或外部地形感知,其中前者仅提供足底支撑的间接证据,而后者在接触前预测接触条件,但无法实时观察实际支撑情况。尽管一些研究将足底接触作为辅助感知纳入,但它们主要依赖于汇总统计,忽视了足底压力的空间拓扑,这提供了对实际接触状态的更直接表征。为了解决这一问题,我们提出了Tac4Loco,一个触觉感知框架,将多阵列足底压力作为类人行走的直接反馈。我们制定了一种保持拓扑的序数表示,将模拟和物理传感器信号映射到共享观察空间,并使用双分支编码器提取其空间和时间表示。随后,学习到的时空特征与增强的本体感觉(包括地形估计线索)相结合,提供给一个不对称的演员-评论家架构进行策略学习。大量的仿真和实际实验表明,在倾斜、部分、不对称和变化支撑的地形上,跟踪性能和支撑适应性得到了改善。我们进一步展示了其在未见的顺应性和非结构化地形(包括泡沫平台和碎石路)上的零样本部署。所有代码和实验配置将作为开源发布,以促进可重复性。
cs.RO / 41 / 2608.15784

Reliable Piezoresistive Strain Sensing Through Physical Limits and Uncertainty Monitoring

通过物理极限和不确定性监测实现可靠的压电电阻应变传感
Ballester, Carmen, Muñoz, Víctor, Copaci, Dorin, Blanco, Dolores
Abstract
Soft piezoresistive strain sensors are one of the most common sensing solutions for wearable and soft robotic applications due to their flexibility and compliance. However, their resistance response is nonlinear and hysteretic, and a sensor can be pushed past its calibrated workspace or misbehave inside it, carrying that error into a decision or control loop. Probabilistic regressors track confidence but ignore those limits. A predictive mean can look unremarkable even when the reading comes from a sensor outside its admissible range or already failing internally, so a confident-looking estimate is not the same as a trustworthy one. This paper proposes a reliability framework pairing a physics-informed probabilistic inverse model, built on physics-guided input features, with a risk factor fusing uncertainty with strain and strain-rate limits into a three-state monitor. Tests on a Nitinol wire and a silver-coated polyamide thread with a Gaussian Process raised fit scores to 0.90-0.95 (RMSE 0.26%-0.15%) and a 96% empirical coverage against the 95% target. The monitor caught 95% of out-of-range and 100% of abnormal conditions while staying reliable under nominal operation. A sensor that reports confidence alongside its estimate lets a system withhold action instead, since it needs no labeled failure examples, which are hard to collect for soft materials.
Chinese Translation
软压电电阻应变传感器因其灵活性和顺应性而成为可穿戴设备和软机器人应用中最常见的传感解决方案之一。然而,它们的电阻响应是非线性和滞后的,传感器可能超出其校准工作空间或在其中表现异常,将这种错误带入决策或控制循环。概率回归模型能够跟踪置信度,但忽略了这些极限。即使读数来自于超出其可接受范围或内部已经失效的传感器,预测均值看起来也可能毫不起眼,因此,表面上看起来可信的估计并不等同于可靠的估计。本文提出了一种可靠性框架,将基于物理指导输入特征构建的物理信息概率逆模型与风险因子相结合,将不确定性与应变和应变率极限融合为一个三态监测器。在对镍钛合金线和银涂层聚酰胺线的测试中,使用高斯过程将拟合分数提高至0.90-0.95(均方根误差0.26%-0.15%),并实现了96%的经验覆盖率,超过95%的目标。该监测器在正常操作下能够捕捉95%的超范围和100%的异常条件,同时保持可靠性。一个在报告估计值的同时提供置信度的传感器使系统能够选择不采取行动,因为它不需要标记的失效示例,而这些示例对于软材料的收集非常困难。
cs.RO / 42 / 2608.15816

ViTaR: Visuo-Tactile Residual Adaptation for Foundation VLA Manipulation

ViTaR:用于基础视觉-语言-动作(VLA)操控的视觉-触觉残差适应
Wang, Yi, Wu, Renjun, Liu, Jinyan, Li, Xuesong
Abstract
As Vision-Language-Action (VLA) models scale toward real-world deployment, contact-rich manipulation exposes a critical blind spot: these policies encode broad visual-semantic priors yet remain unaware of local contact events, producing identical actions whether contact is established, lost, or destabilized. Existing remedies either modify VLA internals, risking catastrophic forgetting, or demand online reinforcement under near-failure contact conditions. Both grant tactile unbounded influence over action generation, conflicting with the priors that make VLAs generalizable. We introduce ViTaR, which reframes tactile feedback from an action-generating perceptual input to an execution modulator that selects and scales bounded residual corrections atop a frozen VLA, preserving pretrained capabilities by construction. ViTaR decomposes adaptation into two stages: Effect-Guided Modeling determines whether and which correction is locally justified via outcome-grounded preference evidence, and Residual Action Modulation converts this evidence into a residual choice with continuously scaled gain from real-time visuotactile observations. On the UniVTAC benchmark spanning seven contact-rich tasks, ViTaR achieves 61.3% average success, a 30.6 percentage-point improvement over its frozen VLA base that also surpasses purpose-built tactile baselines. Physical-robot experiments confirm that bounded tactile modulation transfers to real sensor noise and dynamics.
Chinese Translation
随着视觉-语言-动作(VLA)模型向现实世界部署的扩展,接触丰富的操控暴露出一个关键盲点:这些策略编码了广泛的视觉-语义先验,但对局部接触事件却缺乏感知,无论接触是建立、丢失还是不稳定,均产生相同的动作。现有的解决方案要么修改VLA的内部结构,冒着灾难性遗忘的风险,要么要求在接近失效的接触条件下进行在线强化学习。这两者都赋予触觉对动作生成无限制的影响,与使VLA具备广泛适应性的先验相冲突。我们提出了ViTaR,它将触觉反馈重新定义为从生成动作的感知输入转变为执行调节器,选择并调整基于固定VLA的有限残差修正,通过构建保留预训练能力。ViTaR将适应分解为两个阶段:效果引导建模(Effect-Guided Modeling)通过基于结果的偏好证据确定是否以及哪种修正在局部是合理的,而残差动作调节(Residual Action Modulation)则将这些证据转换为具有实时视觉-触觉观察的连续增益的残差选择。在涵盖七个接触丰富任务的UniVTAC基准测试中,ViTaR实现了61.3%的平均成功率,比其固定VLA基础提高了30.6个百分点,并且超越了专门构建的触觉基线。物理机器人实验确认了有限的触觉调节能够转移到真实的传感器噪声和动态中。
cs.RO / 43 / 2608.15863

Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning

通过数据合成和统一规划扩展手动基础的家电操作
Long, Yuxing, Kang, Lei, Yu, Ziyan, Gao, Yuzheng, Cheng, Bin, Zhang, Jiyao, Li, Xiaoqi, Yang, Haolin, Li, Dongjiang, Shen, Hui, Dong, Hao
Abstract
Operating household appliances requires long-horizon planning that is state-dependent and robust to disturbances, yet existing large models fall short, as no sufficiently diverse, task-oriented dataset exists to support such planning. To bridge this gap, we propose MAGE, a scalable data synthesis pipeline that introduces a novel Hierarchical Appliance Graph (HAG) to automatically generate part grounding, long-horizon planning, and closed-loop recovery data from appliance manuals. With MAGE, we build UseAppliance, the first large-scale dataset for manual-grounded appliance manipulation planning, spanning 22 appliance categories with 89K+ part annotations, 53K+ manipulation tasks, and 33K+ closed-loop adjustment steps. Built on UseAppliance, we develop AppliancePlan, an end-to-end model for manual-grounded appliance manipulation planning. On RealAppliance-Bench, AppliancePlan with only 7B parameters achieves over 10x the best baseline on open-loop planning and consistently outperforms state-of-the-art models across all tasks. Real-robot experiments on six household appliances further confirm effective sim-to-real transfer, marking an important step toward general-purpose household robotics.
Chinese Translation
操作家用电器需要依赖状态的长期规划,并且对干扰具有鲁棒性,但现有的大型模型无法满足这一需求,因为没有足够多样化且以任务为导向的数据集来支持此类规划。为了解决这一问题,我们提出了MAGE,一个可扩展的数据合成管道,采用新颖的层次家电图(Hierarchical Appliance Graph, HAG)来自动生成来自家电手册的部件定位、长期规划和闭环恢复数据。通过MAGE,我们构建了UseAppliance,这是第一个针对手动基础的家电操作规划的大规模数据集,涵盖22个家电类别,包含超过89K的部件注释、53K+的操作任务和33K+的闭环调整步骤。在UseAppliance的基础上,我们开发了AppliancePlan,一个用于手动基础家电操作规划的端到端模型。在RealAppliance-Bench上,AppliancePlan仅用7B参数在开放式规划中实现了超过10倍的最佳基线,并在所有任务中持续超越最先进的模型。对六种家用电器的真实机器人实验进一步确认了有效的仿真到现实转移,标志着朝着通用家用机器人迈出了重要一步。
cs.RO / 44 / 2608.15875

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

GigaBrain-0.7:通过三系统架构扩展具身基础模型至新兴能力
GigaBrain Team, Ye, Angen, Sun, Axiang, Jin, Can, Cheng, Chenxi, Shi, Chong, Shang, Dengke, Zhang, Dingqian, Huang, Guan, Wang, Guangqiang, Ding, Guangqing, Li, Guo, Li, Hangcong, Zhong, Hengyu, Lu, Hongtao, Qin, Jianbo, Mao, Jiming, Zhu, Jing, Lv, Jindi, Cui, Jingzhi, Xie, Junjie, Bao, Junyi, Liu, Kai, Yuan, Lei, Long, Limin, Feng, Lv, Yu, Mingming, Li, Peng, Yi, Pengfei, Li, Qi, Zhang, Qianli, Li, Qingfang, Hu, Qitang, Zhang, Rui, Sun, Shaoyan, Sun, Shibo, Duan, Shiying, Chen, Tenghui, Liu, Tianze, Ke, Weijie, Xue, Wenyao, Wang, Xiaofeng, Tian, Xiaoyu, Liu, Xinyu, Chen, Xinze, Wang, Yang, Wang, Yankai, Zeng, Yejun, Li, Yifan, Nie, Yifei, Li, Yilong, Liu, Yilong, Feng, Yongchao, Wang, Yumeng, Ye, Yun, Liu, Zhichao, He, Ziheng, Yang, Zonghai, Zhu, Zheng
Abstract
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including $\pi_{0.5}$, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.
Chinese Translation
视觉-语言-行动(VLA)模型已成为通用具身智能体的主导范式,在结构化环境中展示了强大的复杂和长期任务完成能力。然而,目前尚不清楚现有的VLA系统是否能够从更有效的架构设计中受益,是否能够扩展到更大且更异质的数据环境,并在任务和具身形式之间实现更广泛的泛化。为此,我们提出了GigaBrain-0.7,这是一种在多样化机器人具身形式中具有显著改进泛化能力的具身基础模型。具体而言,GigaBrain-0.7通过三系统架构统一理解、预测和行动,将预训练扩展到超过37,000小时的异质具身数据,并引入了一阶段对齐训练,联合优化视觉-语言理解和多具身行动生成。与之前的GigaBrain-0系列和包括$ ext{π}_{0.5}$在内的先前最先进模型相比,GigaBrain-0.7在基础零-shot能力、语言条件指令跟随和后训练任务成功率方面取得了显著提升。特别是在我们的内部Maker H01平台和主流机器人具身形式上,GigaBrain-0.7展示了在家庭和工业场景中强大的任务适应性和完成能力。所有训练代码和预训练模型权重将被公开发布。
cs.RO / 45 / 2608.15884

Grouping Auction-Consensus Algorithm for Decentralized Task Allocation in Multi-Robot Systems

用于多机器人系统去中心化任务分配的分组拍卖共识算法
Rodriguez, Jose, Koenig, Sven, Dong, Wenjie, Lu, Qi
Abstract
Decentralized multi-robot task allocation (MRTA) is essential for scalable and resilient autonomous systems. The Consensus-Based Bundle Algorithm (CBBA) is a widely adopted decentralized baseline. However, its individual task-level bidding is poorly aligned with the min-sum objective of minimizing total team travel distance, leading to suboptimal allocations in spatially distributed environments. This paper introduces the Grouping Auction-Consensus Algorithm (GACA). This decentralized MRTA framework adopts the two-phase auction-consensus architecture of CBBA while fundamentally redesigning its bidding mechanism to reason over groups of spatially proximate tasks. A nearest-neighbor preprocessing step partitions tasks into spatially coherent groups before allocation. Agents then iteratively propose structured group-level actions: claiming unassigned groups, acquiring partial groups, or contesting groups held by other agents. Competing actions are resolved through a consensus phase. Operating in the MT-SR-IA problem class, GACA is evaluated against CBBA using a Mixed-Integer Linear Program as the ground-truth optimality reference. Across four swarm sizes and 4,000 test worlds, GACA achieves a median percent optimality of approximately 97% compared to 81--84% for CBBA, while converging in equal or fewer iterations. A scalability evaluation over 3,280 additional problem instances spanning swarm sizes of 5 to 20 agents and task counts of 10 to 50 confirms that these gains generalize robustly across a wide range of problem configurations.
Chinese Translation
去中心化的多机器人任务分配(MRTA)对于可扩展和具有弹性的自主系统至关重要。基于共识的捆绑算法(CBBA)是广泛采用的去中心化基准。然而,其个体任务级别的竞标与最小和目标(min-sum objective)——最小化团队总旅行距离——之间的匹配度较差,导致在空间分布环境中的次优分配。本文提出了分组拍卖共识算法(GACA)。该去中心化的MRTA框架采用了CBBA的两阶段拍卖共识架构,同时从根本上重新设计了其竞标机制,以针对空间上相近的任务组进行推理。在分配之前,最近邻预处理步骤将任务划分为空间一致的组。然后,代理迭代地提出结构化的组级行动:声称未分配的组、获取部分组或争夺其他代理持有的组。通过共识阶段解决竞争行动。在MT-SR-IA问题类中,GACA与CBBA进行了评估,使用混合整数线性规划(Mixed-Integer Linear Program)作为真实最优性参考。在四种群体规模和4,000个测试世界中,GACA的中位数最优性百分比约为97%,而CBBA为81%至84%,且在相同或更少的迭代中收敛。对3,280个额外问题实例的可扩展性评估,涵盖5到20个代理的群体规模和10到50个任务数量,确认这些增益在广泛的问题配置中具有强大的普适性。
cs.RO / 46 / 2608.15897

Tactile Sim2Real without Tactile Simulation via Bottlenecked Latent Reconstruction

无触觉仿真下的触觉Sim2Real通过瓶颈潜在重建
Yang, Fan, Wi, Youngsun, Yu, Jinhao, Fazeli, Nima, Berenson, Dmitry
Abstract
Robot sensor designs, particularly tactile sensors, are highly diverse and evolve rapidly. Modeling each sensor in simulation demands substantial domain expertise and computational approximations can degrade the fidelity of the simulated signals. We propose Sim2Real via Bottlenecked Latent Reconstruction (SBLR), a framework that avoids sensor-specific simulation entirely by (1) training policies on a simulator-native oracle sensor that is easy to construct without modeling any particular sensor (e.g. we use a point-cloud and finger-tip forces as a tactile oracle), and (2) aligning real sensor latent embeddings to those of the oracle sensor at inference time. Policy training proceeds in two-stage: the policy first learns from the oracle sensor latents, then a bottlenecked latent reconstruction adapts it to the information loss expected when using the real sensor instead of the oracle. The alignment between oracle and real sensor is learned from unpaired random-play data collected in both simulation and the real world, using rectified-flow-based transformation networks trained on nearest-neighbor pseudo-pairs. Simulation experiments on three contact-rich tasks show that SBLR matches or approaches the performance of an oracle with direct access to tactile simulation. Hardware experiments on Peg Insertion and Gear Meshing with GelSight Mini and DIGIT sensors demonstrate 85-97.5% zero-shot success without requiring any sensor-specific modeling or calibration, outperforming a physics-based tactile simulation baseline by 7.5-15%.
Chinese Translation
机器人传感器设计,特别是触觉传感器,种类繁多且发展迅速。在仿真中对每个传感器进行建模需要大量的领域专业知识,而计算近似可能会降低仿真信号的保真度。我们提出了通过瓶颈潜在重建(Sim2Real via Bottlenecked Latent Reconstruction, SBLR)的方法,该框架完全避免了特定传感器的仿真,具体通过(1)在一个易于构建的仿真原生神谕传感器上训练策略,而无需建模任何特定传感器(例如,我们使用点云和指尖力作为触觉神谕),以及(2)在推理时将真实传感器的潜在嵌入与神谕传感器的潜在嵌入对齐。策略训练分为两个阶段:策略首先从神谕传感器的潜在嵌入中学习,然后通过瓶颈潜在重建将其适应于使用真实传感器时预期的信息损失。神谕传感器与真实传感器之间的对齐是通过在仿真和真实世界中收集的无配对随机播放数据学习的,使用基于修正流的变换网络在最近邻伪对上进行训练。在三个接触丰富的任务上的仿真实验表明,SBLR的性能与直接访问触觉仿真的神谕相匹配或接近。在Peg Insertion和Gear Meshing的硬件实验中,使用GelSight Mini和DIGIT传感器展示了85-97.5%的零-shot成功率,而无需任何特定传感器的建模或校准,超越了基于物理的触觉仿真基线7.5-15%。
cs.RO / 47 / 2608.15917

Pre-training Visual Dexterity in Simulation

在仿真中进行视觉灵巧性的预训练
Kamat, Sarthak, Rashid, Adam, Sharma, Satvik, Doriwala, Aseem, Finn, Chelsea, Isola, Phillip, Liu, C. Karen
Abstract
Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this progress has largely been driven by datasets and embodiments built around simple parallel-jaw grippers. Dexterous, multi-fingered hands remain comparatively data-starved because real teleoperation is costly to scale, while human hand video is off-embodiment and requires lossy pose estimation and retargeting. We introduce Simulation Pre-training for Dexterity (SPD), a pre-training framework for dexterous manipulation that uses data entirely collected in simulation. In SPD, humans manipulate virtual objects inside a VR headset, enabling on-embodiment trajectories and robot-free collection. With the help of five operators, we collect 75 hours of multi-task dexterous manipulation over one week, and use it to pre-train a causal transformer on a sequence modeling objective. We study the benefits of simulation pre-training on real-world tasks by fine-tuning on 1-2 hours of physical demonstrations on a 56-DoF bimanual dexterous setup. We find that our approach outperforms training behavior cloning policies from scratch, showing that simulation teleoperation is a viable pre-training source for real-world dexterous manipulation. We perform ablation studies, measuring the benefits of history conditioning and short action chunks for reactive control.
Chinese Translation
大规模预训练使得机器人策略微调变得越来越数据高效,但这一进展主要是由围绕简单平行夹爪构建的数据集和实例驱动的。灵巧的多指手部在数据方面仍然相对匮乏,因为真实的远程操作在规模化上成本高昂,而人类手部视频则是脱离实体的,并且需要损失性姿态估计和重定向。我们提出了灵巧性预训练仿真框架(Simulation Pre-training for Dexterity, SPD),这是一个完全在仿真中收集数据的灵巧操作预训练框架。在SPD中,人类在虚拟现实头显中操控虚拟物体,从而实现了在实体内的轨迹和无机器人收集。借助五名操作员,我们在一周内收集了75小时的多任务灵巧操作数据,并用其对因果变换器进行序列建模目标的预训练。我们通过在56自由度双手灵巧设置上对1-2小时的物理演示进行微调,研究了仿真预训练在现实任务中的好处。我们发现我们的方法在从零开始训练行为克隆策略时表现优于后者,表明仿真远程操作是现实世界灵巧操作的可行预训练来源。我们进行了消融研究,测量了历史条件和短动作片段对反应控制的好处。
cs.RO / 48 / 2608.15924

RAPAC-DP: Response-Aligned Pending-Action Compensation for Diffusion Policies under Delayed Execution

RAPAC-DP:针对延迟执行的扩散策略的响应对齐待执行动作补偿
Wang, Tao, Wang, Wei, Wang, Jianhui, Wang, Qi, Huang, Weidi, Xu, Bing
Abstract
Cloud-side inference gives imitation-learning policies access to greater computational resources, but communication and computation delays can degrade control performance. To compensate for these delays, we propose RAPAC-DP, a response-aligned pending-action compensation framework designed for both diffusion- and flow-based action generators. RAPAC-DP encodes the actions already scheduled for execution before the cloud response arrives into a pending-action sequence that serves as the conditioning input to a parameter-efficient compensation pathway. When delay effects are negligible, bypassing this pathway exactly recovers the frozen base policy. For training, RAPAC-DP constructs delay-conditioned samples from delay-free demonstrations, requiring neither explicit system dynamics nor additional delayed demonstrations. At the largest fixed delay tested on Kinetix, RAPAC-DP retained 81.4% of its overall delay-free performance. At the largest fixed delay tested on each RoboMimic task, it achieved a mean success rate of 0.633 across the three tasks. These results demonstrate the effectiveness of pending-action compensation for cloud-deployed imitation-learning policies.
Chinese Translation
云端推理使模仿学习策略能够访问更强大的计算资源,但通信和计算延迟可能会降低控制性能。为了补偿这些延迟,我们提出了RAPAC-DP,这是一种响应对齐的待执行动作补偿框架,旨在为扩散型和流型动作生成器服务。RAPAC-DP将云响应到达之前已经安排执行的动作编码为待执行动作序列,该序列作为参数高效补偿路径的条件输入。当延迟效应可以忽略时,绕过此路径可以准确恢复冻结的基础策略。在训练过程中,RAPAC-DP从无延迟演示中构建延迟条件样本,无需显式的系统动态或额外的延迟演示。在Kinetix上测试的最大固定延迟下,RAPAC-DP保留了81.4%的整体无延迟性能。在每个RoboMimic任务上测试的最大固定延迟下,它在三个任务中的平均成功率达到了0.633。这些结果证明了待执行动作补偿在云端部署的模仿学习策略中的有效性。
cs.RO / 49 / 2608.15938

Revisiting Open-Loop Execution in Robotics: Toward Reactive, Higher-Performing Policies

重新审视机器人中的开环执行:朝向反应性更强的高性能策略
Zeng, Michael, Agarwal, Abhinav, Bati, Ajay, Lee, Brian, Ancha, Siddharth, Tedrake, Russ
Abstract
Action chunking --- the practice of predicting a sequence of actions and executing a prefix open-loop --- has emerged as a key enabler of recent progress in imitation learning for robotic manipulation. However, executing long open-loop prefixes reduces reactivity, limiting policies' ability to correct for errors. Further, the mechanisms underlying these performance benefits remain poorly understood: prior works cite mitigating compounding errors, absorbing inference latency, or smoothing motions, but provide limited controlled evidence or guidance for preserving reactivity. In this work, we argue that long open-loop execution primarily helps short-context policies imitate "non-Markovian demonstrations". Across four simulation and two real-world tasks, we show that expert non-Markovianity strongly shapes the relationship between task success and open-loop execution horizon. Further, we investigate the impact of compounding errors --- the prevailing explanation for long open-loop execution in prior work --- and find that while they matter, expert non-Markovianity has a much stronger impact in our experimental setting. Finally, we show that when policies are provided with a sufficiently long context, open-loop execution is no longer beneficial and the most reactive, closed-loop policies perform best. While imitation learning has seen great success using long open-loop execution, our findings motivate long-context, reactive policies as a more principled and performant paradigm.
Chinese Translation
动作分块——预测一系列动作并执行前缀开环的实践——已成为最近在机器人操作模仿学习中取得进展的关键推动力。然而,执行长时间的开环前缀会降低反应性,限制策略纠正错误的能力。此外,这些性能优势背后的机制仍然不够清楚:先前的研究提到缓解累积错误、吸收推理延迟或平滑运动,但提供的受控证据或保持反应性的指导有限。在本研究中,我们认为长时间的开环执行主要帮助短上下文策略模仿“非马尔可夫演示”。在四个仿真和两个真实世界任务中,我们展示了专家的非马尔可夫性如何强烈影响任务成功与开环执行范围之间的关系。此外,我们调查了累积错误的影响——这是先前研究中对长时间开环执行的主要解释——发现虽然它们确实重要,但在我们的实验设置中,专家的非马尔可夫性有更强的影响。最后,我们表明,当策略提供足够长的上下文时,开环执行不再有利,反应性最强的闭环策略表现最佳。尽管模仿学习在使用长时间开环执行方面取得了巨大成功,但我们的发现促使我们考虑长上下文、反应性策略作为一种更有原则和更高性能的范式。
cs.RO / 50 / 2608.15946

Rotate Disks to Reach Farther: Design and Modeling of a Novel Reconfigurable Tendon Driven Manipulator

旋转盘以达到更远的目标:一种新型可重构腱驱动操纵器的设计与建模
Dash, Sabyasachi, Liu, Yangkun, Hunter, Will, Golden, John, Krishnan, Girish
Abstract
Rerouting the tendon path in tendon driven continuum manipulators (TDCMs) enables a broad range of deformation modes. This work presents a Reconfigurable TDCM design which allows independent rotation of intermediate spacer disks, thereby locally rerouting the tendon and achieving non-trivial backbone spatial deformations. Two such designs, (a) Manual Disk Locked (MDL) and (b) Continuous Disk Rotor (CDR) manipulators are presented to achieve disk rotations before and during operation, respectively. A predictive static model based on the piecewise constant strain (PCS) assumption is developed within a potential energy minimization framework, incorporating (a) disk rotations, (b) discrete tendon paths between disk segments, (c) rigid thickness of spacer disks, and (d) elasticity of the tendons. The model is validated against experimental results, demonstrating an average tip error of $1.2\%$ of the manipulator's total length for parallel tendon routing and around $3\%$ for the case when multiple disks are rotated. The computation time is an order of magnitude lower than the state of the art Cosserat rod solver.
Chinese Translation
在腱驱动连续操纵器(TDCMs)中重新规划腱的路径使得能够实现多种变形模式。本研究提出了一种可重构的TDCM设计,该设计允许中间间隔盘的独立旋转,从而局部重新规划腱并实现复杂的骨架空间变形。提出了两种设计,(a) 手动锁定盘(MDL)和(b) 连续盘转子(CDR)操纵器,分别用于在操作前和操作期间实现盘的旋转。在一个基于分段常量应变(PCS)假设的潜能能量最小化框架内,开发了一种预测静态模型,该模型考虑了(a) 盘的旋转,(b) 盘段之间的离散腱路径,(c) 间隔盘的刚性厚度,以及(d) 腱的弹性。该模型通过实验结果进行了验证,显示在平行腱路由情况下,操纵器末端的平均误差为其总长度的$1.2\%$,而在多个盘旋转的情况下,误差约为$3\\%$。计算时间比现有的Cosserat杆求解器低一个数量级。
cs.RO / 51 / 2608.15968

Tabletop Pen Manipulation With a Vision-Guided 4-DoF Arm

基于视觉引导的四自由度臂的桌面笔操作
Rangarajan, Anirudh, Bianchini, Bibit
Abstract
Low-cost four-degree-of-freedom (DoF) arms are among the most accessible robotic platforms. But they are, in theory, underactuated for picking up in situations where objects are at arbitrary orientations, a task that appears to require five degrees of freedom: the planar position (x and y), the height (z), a wrist rotation to align the gripper with the object, and gripper actuation, of which a four-DoF arm lacks the wrist rotation. This work shows that perception and motion planning can enable such an arm, a roughly $200 Waveshare RoArm-M2-S, under a fixed overhead camera to detect and color-sort writing utensils without that joint. A YOLO11n-OBB (You Only Look Once, oriented bounding box) detector locates each writing utensil; camera intrinsics and an ArUco reference pose convert its pixel coordinates to robot coordinates; and a color classifier labels it. The detected orientation angle determines the motion strategy: utensils close to the arm's fixed approach direction are picked up directly, and those at steeper angles are reoriented via corrective sweeps until they are graspable, after which they are picked up and sorted into the assigned color bin. Across 326 logged motions on seven writing utensils, the arm made 196 direct grasps and 130 corrective sweep passes, correcting misalignments up to 90 degrees, suggesting that clever task-informed engineering can compensate for a missing degree of freedom on tasks like this one.
Chinese Translation
低成本的四自由度(DoF)机械臂是最易获取的机器人平台之一。然而,从理论上讲,它们在处理物体处于任意方向的情况下抓取时是欠驱动的,这一任务似乎需要五个自由度:平面位置(x 和 y)、高度(z)、手腕旋转以将夹具与物体对齐,以及夹具的驱动,而四自由度的机械臂缺少手腕旋转。本文展示了感知和运动规划如何使得这样一款大约 $200 的 Waveshare RoArm-M2-S 机械臂在固定的上方摄像头下,能够在没有该关节的情况下检测和分类书写工具。YOLO11n-OBB(You Only Look Once,定向边界框)检测器定位每个书写工具;相机内参和 ArUco 参考姿态将其像素坐标转换为机器人坐标;颜色分类器对其进行标记。检测到的方向角决定了运动策略:靠近机械臂固定接近方向的工具直接被抓取,而在更陡角度的工具则通过纠正性扫动进行重新定向,直到它们可被抓取,之后再进行抓取并分类到指定的颜色箱中。在对七个书写工具进行的 326 次记录运动中,机械臂进行了 196 次直接抓取和 130 次纠正性扫动,纠正了多达 90 度的错位,表明巧妙的任务导向工程可以弥补此类任务中缺失的自由度。
cs.RO / 52 / 2608.15995

Learning Varying Physical Therapist-Patient Interactions for Robot-mediated Upper Limb Task-Specific Training

学习变化的物理治疗师-患者互动以实现机器人介导的上肢任务特定训练
Loh, Jia Quan, Crocher, Vincent, Klaic, Marlena, Oetomo, Denny, Tan, Ying
Abstract
Upper extremity motor function recovery is positively linked to Task-Specific Training (TST) and sufficient therapy dosage. Rehabilitation robots can increase TST dosage via controlled, repetitive treatment and free therapists to simultaneously manage other patients, but it has yet to demonstrate significant benefits over conventional treatment. This is potentially linked to inaccurate robotic representation of personalised physical therapist-patient interaction and lack of practice variability during TST. Hence, we advocate for robotic interventions that preserve the personalised physical therapist-patient interactions when delivering TST for patients across varying practise conditions. We propose a Learning-from-Demonstration framework using Task-Parameterised Gaussian Mixture Models (TPGMM) to learn personalised physical therapist-patient interaction in Task-Specific exercises, mapping patient joint kinematics to therapist-applied torques using few demonstrations. The model is generalised to reconstruct therapist torques in new task variations. The framework was evaluated on physical interactions from 14 mock "therapist-patient" pairs over three tasks of increasing complexity, each with six variations. A benchmark comparison against a Look-Up Table was conducted. The results show both methods reproducing interactions in unseen task variations that deviate slightly from the actual interaction, with TPGMM slightly outperforming LUT. Both methods reproduced interactions that gets increasingly closer to the actual interaction as task complexity increases.
Chinese Translation
上肢运动功能恢复与任务特定训练(Task-Specific Training, TST)和足够的治疗剂量呈正相关。康复机器人可以通过受控的重复治疗增加TST剂量,并使治疗师能够同时管理其他患者,但尚未显示出相对于传统治疗的显著优势。这可能与机器人对个性化物理治疗师-患者互动的表示不准确以及在TST期间缺乏实践变异性有关。因此,我们倡导在为不同实践条件下的患者提供TST时,机器人干预应保留个性化的物理治疗师-患者互动。我们提出了一种基于示范学习(Learning-from-Demonstration)的框架,使用任务参数化高斯混合模型(Task-Parameterised Gaussian Mixture Models, TPGMM)来学习任务特定练习中的个性化物理治疗师-患者互动,通过少量示范将患者关节运动学映射到治疗师施加的扭矩。该模型被推广以重建新任务变体中的治疗师扭矩。该框架在14对模拟“治疗师-患者”配对的物理互动上进行了评估,涵盖了三个逐渐复杂的任务,每个任务有六种变体。与查找表(Look-Up Table, LUT)进行了基准比较。结果显示,两种方法在未见过的任务变体中再现的互动略微偏离实际互动,其中TPGMM的表现略优于LUT。随着任务复杂性的增加,两种方法再现的互动逐渐接近实际互动。
cs.RO / 53 / 2608.16030

Benchmarking Identity-Sensitive LLM Outputs for Surveillance and Security Robots

监测与安全机器人身份敏感大型语言模型输出的基准测试
Hyman, Nneka, Khan, Jasmine, Korpan, Raj
Abstract
Large language models (LLMs) are increasingly used to generate textual robot design specifications, interaction policies, and risk assessments during early-stage robot development. Such outputs may influence how surveillance and security robots are conceptualized, documented, and ultimately implemented. This paper evaluates whether identity-conditioned prompts produce systematic differences in LLM-generated surveillance and security robot design descriptions. Using 236 demographic identity labels across single-label and model-augmented prompt conditions, we analyze readability as an initial benchmark for evaluating accessibility and identity-conditioned variation in generated robot design descriptions. The results show significant differences in readability across prompt conditions, design dimensions, and demographic identities. Although readability cannot determine whether an output is fair or socially appropriate, it provides an interpretable baseline within a broader benchmarking framework that also includes lexical, semantic, sentiment, syntactic, and fairness-focused analyses.
Chinese Translation
大型语言模型(LLMs)在早期机器人开发过程中越来越多地用于生成文本机器人设计规范、交互政策和风险评估。这些输出可能会影响监测与安全机器人的概念化、文档编制和最终实施。本文评估了身份条件提示是否会在LLM生成的监测与安全机器人设计描述中产生系统性差异。通过使用236个人口身份标签,在单标签和模型增强提示条件下,我们分析了可读性作为评估生成的机器人设计描述的可及性和身份条件变化的初步基准。结果显示,在提示条件、设计维度和人口身份之间的可读性存在显著差异。尽管可读性无法判断输出是否公平或社会适当,但它在更广泛的基准测试框架中提供了一个可解释的基线,该框架还包括词汇、语义、情感、句法和公平性分析。
cs.RO / 54 / 2608.16041

ScenarioCharacterization: A Modular Toolkit for Characterizing Safety across Trajectory Datasets

ScenarioCharacterization:一个用于跨轨迹数据集安全性特征描述的模块化工具包
Navarro, Ingrid, Duan, Yutong, Francis, Jonathan, Oh, Jean
Abstract
We introduce ScenarioCharacterization, an open-source framework for automated, dataset-agnostic profiling of driving scenarios in trajectory datasets. Our framework is packaged as a modular, configuration-driven pipeline of three layers: a dataset adapter that maps custom datasets onto an open Scenario representation, a characterizer that performs feature extraction, behavior probing, and criticality scoring at scenario and agent levels, and an analysis layer for scenario visualization and feature, score, and probe analyses. Because the layers communicate only through Pydantic-validated schemas composed via configurations, a new dataset can easily plug in without rewriting the characterization and analysis stack. This technical report describes the design and APIs, shows example outputs on Waymo Open Motion, Argoverse2, and nuPlan, and discusses downstream uses of the approach. The framework is available at https://github.com/navarrs/ScenarioCharacterization.
Chinese Translation
我们介绍了ScenarioCharacterization,这是一个开源框架,用于在轨迹数据集中自动化、数据集无关的驾驶场景特征描述。我们的框架被打包为一个模块化、基于配置的三层管道:一个数据集适配器,将自定义数据集映射到开放的场景表示,一个特征描述器,负责在场景和代理级别进行特征提取、行为探测和关键性评分,以及一个用于场景可视化和特征、评分及探测分析的分析层。由于各层仅通过通过配置组成的Pydantic验证的模式进行通信,因此可以轻松地将新数据集接入,而无需重写特征描述和分析堆栈。本技术报告描述了设计和API,展示了在Waymo Open Motion、Argoverse2和nuPlan上的示例输出,并讨论了该方法的下游应用。该框架可在https://github.com/navarrs/ScenarioCharacterization获取。
cs.RO / 55 / 2608.16058

SurgVIL: Scaling Surgical Robot Imitation Learning with Open-source Surgical Videos

SurgVIL:利用开源外科视频扩展外科机器人模仿学习
Chen, Xinhao, Chen, JuoTung, Nelson, Nigel, Goldenberg, Antony, Haworth, Jesse, Huver, Sean D., Krieger, Axel
Abstract
Learning-based surgical robot autonomy requires large-scale demonstrations with synchronized videos and robot actions, but such data are exceedingly rare in clinical or realistic tissue settings because robot kinematics are typically inaccessible outside controlled research systems. In contrast, phantom data collected on research platforms provide accurate action labels but lack the visual diversity of real tissue. We propose SurgVIL, a framework for scaling surgical robot imitation learning using open-source surgical videos. SurgVIL combines kinematically labeled phantom robot demonstrations with surgical videos from open-source datasets and online sources for policy learning. Since these videos lack robot motion labels, we estimate approximate kinematics as weak supervision. We evaluate SurgVIL on two da Vinci robot tasks: needle pick-up and cholecystectomy cutting. Across ACT, $\pi_0$, and GR00T-H backbones, adding surgical videos substantially improves generalization to real-tissue and out-of-distribution settings, suggesting a scalable path from phantom training toward generalizable surgical robot policies.
Chinese Translation
基于学习的外科机器人自主性需要大规模的演示数据,这些数据包括同步的视频和机器人动作,但在临床或真实组织环境中,这类数据极为稀缺,因为机器人运动学通常在受控研究系统之外无法获取。相比之下,在研究平台上收集的虚拟数据提供了准确的动作标签,但缺乏真实组织的视觉多样性。我们提出了SurgVIL,一个利用开源外科视频扩展外科机器人模仿学习的框架。SurgVIL将运动学标记的虚拟机器人演示与来自开源数据集和在线来源的外科视频结合,用于策略学习。由于这些视频缺乏机器人运动标签,我们估计近似的运动学作为弱监督。我们在两个da Vinci机器人任务上评估SurgVIL:针头拾取和胆囊切除。通过ACT、$ ext{pi}_0$和GR00T-H骨干网,增加外科视频显著改善了对真实组织和分布外环境的泛化能力,表明从虚拟训练到可推广的外科机器人策略的可扩展路径。
cs.RO / 56 / 2608.16074

US-VLA: An Ultrasound Vision-Language-Action Model for Embodied Abdomina

US-VLA:一种用于具身腹部超声的超声视觉-语言-动作模型
Zhang, Cheng, Wu, Xingzheng, Yan, Guihao, Hu, Xifeng, Liu, Zhi, Wu, Mei, Cai, Qing
Abstract
Artificial intelligence-assisted ultrasound scanning enhances diagnostic reliability and efficiency by providing real-time guidance for standardized image acquisition and reducing operator dependence. However, existing reinforcement learning and learning-assisted ultrasound scanning methods typically rely on carefully designed reward functions or extensive interaction data, which limits their generalization ability and stability across different devices, patient populations, and complex clinical scenarios. To address these challenges, we propose an ultrasound vision-language-action model (US-VLA) for automated ultrasound scanning that explicitly encodes clinical semantic goals and generates sequential probe manipulation actions under real-time ultrasound feedback. In particular, we first design an ultrasound-aware expert fusion module to jointly integrate ultrasound observations with auxiliary contextual information, enabling semantic ultrasound feedback to effectively guide the scanning process. Then, we construct US-VLA-Data, a real-world dataset covering liver and kidney examinations, which includes five clinically defined standard planes and comprises 320 expert scanning trajectories with approximately 80,000 synchronized timesteps. Extensive experiments demonstrate that US-VLA achieves competitive performance in ultrasound probe manipulation tasks, indicating its effectiveness and promising generalization within the evaluated abdominal ultrasound setting. The source code is available at https://github.com/VMVLab/US-VLA.
Chinese Translation
人工智能辅助的超声扫描通过提供实时指导以标准化图像采集并减少操作员依赖,从而增强了诊断的可靠性和效率。然而,现有的强化学习和学习辅助超声扫描方法通常依赖于精心设计的奖励函数或大量的交互数据,这限制了它们在不同设备、患者群体和复杂临床场景中的泛化能力和稳定性。为了解决这些挑战,我们提出了一种超声视觉-语言-动作模型(US-VLA),用于自动化超声扫描,该模型明确编码临床语义目标,并在实时超声反馈下生成顺序探头操作动作。具体而言,我们首先设计了一个超声感知专家融合模块,以联合整合超声观察与辅助上下文信息,使得语义超声反馈能够有效指导扫描过程。然后,我们构建了US-VLA-Data,这是一个涵盖肝脏和肾脏检查的真实世界数据集,包括五个临床定义的标准平面,并包含320个专家扫描轨迹,约有80,000个同步时间步。大量实验表明,US-VLA在超声探头操作任务中表现出竞争力,表明其在评估的腹部超声设置中的有效性和良好的泛化能力。源代码可在 https://github.com/VMVLab/US-VLA 获取。
cs.RO / 57 / 2608.16153

Unified Condition-Action Modeling for Accurate One-Step Action Generation

统一条件-动作建模用于精确的一步动作生成
Zhou, Xinyu, Cai, Zikun, Zuo, Kuangji, Li, Gen, Ma, Boyu, Lu, Yanshuo, Song, Yutong, Yuan, Mingqi, Chen, Jiayu, Yang, Jianfei
Abstract
Robot manipulation requires policies that are both accurate and efficient, as robot control must respond to changing observations under tight latency constraints. Recent diffusion and flow policies are promising, but they often treat conditions as auxiliary signals rather than jointly evolving them with action trajectories. We find that this limitation can be effectively mitigated by a \textbf{simple yet effective unified condition-action modeling design} that represents conditions and actions in a shared token space, allowing a compact model to achieve high performance while improving both inference speed and accuracy. Therefore, we propose UCA-Flow, a unified condition-action modeling framework for accurate one-step action generation. Our method unifies observation conditions, timestep conditions, interval conditions, and action tokens into a single sequence, and processes them with a Unified Condition-Action Transformer for joint condition-action representation learning. As a result, condition representations are dynamically reconstructed according to the current generation stage, highlighting information most relevant for action refinement. Furthermore, we introduce an improved dual-pass supervision scheme over $u$ and $v$ for stronger optimization of unified condition-action modeling. UCA-Flow improves the average success rate by 9.3 percentage points over the strongest baseline, while achieving $45.6\times$ and $33.4\times$ speedups over DP3 and Simple DP3, and remaining $4.3\times$ and $2.3\times$ faster than one-step FlowPolicy and MP1, respectively.
Chinese Translation
机器人操作需要既准确又高效的策略,因为机器人控制必须在严格的延迟限制下对变化的观察做出响应。最近的扩散和流动策略表现出良好的前景,但它们往往将条件视为辅助信号,而不是与动作轨迹共同演化。我们发现,通过一种 extbf{简单而有效的统一条件-动作建模设计},可以有效缓解这一限制,该设计在共享的标记空间中表示条件和动作,使得紧凑的模型在提高推理速度和准确性的同时实现高性能。因此,我们提出了UCA-Flow,一个用于精确的一步动作生成的统一条件-动作建模框架。我们的方法将观察条件、时间步条件、区间条件和动作标记统一为一个单一序列,并通过统一条件-动作变换器(Unified Condition-Action Transformer)进行联合条件-动作表示学习。结果,条件表示根据当前生成阶段动态重构,突出与动作优化最相关的信息。此外,我们引入了一种改进的双重监督方案,针对$u$和$v$进行更强的统一条件-动作建模优化。UCA-Flow在最强基线之上提高了9.3个百分点的平均成功率,同时在DP3和简单DP3上实现了$45.6 imes$和$33.4 imes$的加速,并且仍然比一步FlowPolicy和MP1快$4.3 imes$和$2.3 imes$。
cs.RO / 58 / 2608.16172

SparkVLA: Stop-Aware Hierarchical VLA with Adaptive Action Chunking for Long-Horizon Manipulation

SparkVLA:具有自适应动作分块的停止感知层次化视觉-语言-动作系统用于长时间操作
Lei, Xunyao, Wu, Renjun, Huo, Tianlin, Li, Xuesong
Abstract
At every re-observation point in a hierarchical Vision-Language-Action (VLA) system, two interface decisions must be made: when to terminate the current subtask and how far to execute the proposed action chunk. These decisions are mutually dependent---the optimal stopping point depends on what the executor plans to do, while the optimal execution length depends on where the subtask boundary lies---yet existing architectures evaluate them in isolation, an asymmetry neither module can overcome alone. We present SparkVLA, a stop-aware hierarchical VLA that resolves this mutual dependency by formulating both decisions as a single ranking: Stop competes against every action-prefix length in a unified candidate set, and the system selects the highest-scoring option, eliminating threshold tuning and requiring only offline ordinal preferences. An Anchor-Conditioned Context Encoding module caches a history-aware subtask anchor encoding onset-state memory and goal semantics, guiding visual-token pruning toward task-relevant regions; a Stop-Aware Action-Prefix Selection head scores all candidates via full self bnattention at chunk boundaries for efficiency. On RoboCerebra, SparkVLA achieves 47.12% success rate, surpassing the official hierarchical baseline by 30.57% and the strongest reproducible method by 26.83% Real-robot experiments on multi-step tasks further validate these gains on physical hardware.
Chinese Translation
在层次化视觉-语言-动作(VLA)系统的每个重新观察点,必须做出两个接口决策:何时终止当前子任务以及执行提议的动作分块的距离。这些决策是相互依赖的——最佳停止点取决于执行者的计划,而最佳执行长度则依赖于子任务边界的位置——然而现有架构将它们孤立评估,这种不对称性是任何一个模块都无法单独克服的。我们提出了SparkVLA,这是一种停止感知的层次化VLA,通过将这两个决策公式化为单一排名来解决这种相互依赖性:停止与统一候选集中的每个动作前缀长度进行竞争,系统选择得分最高的选项,消除了阈值调优,仅需离线顺序偏好。一个锚点条件上下文编码模块缓存了历史感知的子任务锚点编码起始状态记忆和目标语义,引导视觉标记修剪到与任务相关的区域;一个停止感知的动作前缀选择头通过在分块边界进行全自注意力评分所有候选项,以提高效率。在RoboCerebra上,SparkVLA实现了47.12%的成功率,超过了官方层次基线30.57%和最强可复现方法26.83%的表现。对多步骤任务的真实机器人实验进一步验证了这些在物理硬件上的提升。
cs.RO / 59 / 2608.16195

RoboStriker: Latent-Space Strategic Games for Autonomous Humanoid Boxing

RoboStriker:用于自主类人拳击的潜在空间战略游戏
Yin, Kangning, Liu, Kaige, Cao, Zhe, Dong, Wentao, Zeng, Weishuai, Zhang, Tianyi, Zhang, Qiang, Wang, Jingbo, Pang, Jiangmiao, Li, Yang, Zhou, Ming, Zhang, Weinan
Abstract
Achieving human-level competitive intelligence and physical agility in humanoid robots remains a profound challenge, particularly in contact-rich and highly dynamic tasks such as boxing. While Multi-Agent Reinforcement Learning offers a principled framework for strategic interaction, its direct application to unstructured raw motor spaces inevitably leads to joint-level physical collapse, preventing the emergence of any viable combat tactics. To resolve this fundamental conflict between strategic exploration and physical feasibility, we formulate the humanoid combat task as a novel two-player latent-space zero-sum Markov game. Under standard regularity and approximate best-response assumptions, we show that the latent formulation induces an equivalent game over the decoder-reachable action manifold, providing an approximate-Nash interpretation of the resulting self-play dynamics. To instantiate this theoretical formulation, we propose RoboStriker, a hierarchical framework that decouples high-level reasoning from low-level execution. It first distills the tracking expertise of predefined boxing motions into a topologically bounded latent manifold. This structured latent foundation subsequently drives multi-agent co-evolution via Latent-Space Neural Fictitious Self-Play. Extensive experimental results demonstrate that gaming within this structured latent space substantially outperforms direct exploration. By constraining strategic exploration through a pretrained motion decoder, RoboStriker substantially reduces the catastrophic balance failures observed in raw action-space methods and achieves superior tactical performance in both competitive win rates and striking efficiency. Finally, we successfully deploy and validate our learned combat policies on real-world humanoid robots. Our code and video and supplementary materials are available at RoboStriker.
Chinese Translation
在类人机器人中实现人类水平的竞争智能和身体灵活性仍然是一个深刻的挑战,尤其是在拳击等接触丰富且高度动态的任务中。尽管多智能体强化学习提供了一个战略互动的原则框架,但其直接应用于非结构化的原始运动空间不可避免地导致关节级别的物理崩溃,从而阻碍任何可行战术的出现。为了解决战略探索与物理可行性之间的根本冲突,我们将类人战斗任务构建为一种新颖的双人潜在空间零和马尔可夫游戏。在标准的正则性和近似最佳响应假设下,我们证明潜在的构造在解码器可达的动作流形上诱导了一个等效的游戏,为结果自我对弈动态提供了近似纳什解释。为了实现这一理论构造,我们提出了RoboStriker,一个将高层推理与低层执行解耦的层次化框架。它首先将预定义拳击动作的跟踪专业知识提炼为一个拓扑有界的潜在流形。这个结构化的潜在基础随后通过潜在空间神经虚拟自我对弈驱动多智能体的共同进化。大量实验结果表明,在这个结构化潜在空间中进行游戏显著优于直接探索。通过限制战略探索并使用预训练的动作解码器,RoboStriker显著减少了在原始动作空间方法中观察到的灾难性平衡失败,并在竞争胜率和打击效率方面实现了更优的战术表现。最后,我们成功地在真实世界的类人机器人上部署并验证了我们学习的战斗策略。我们的代码、视频和补充材料可在RoboStriker上获取。
cs.RO / 60 / 2608.16221

Deep Probabilistic Indoor Gas Source Localization via Physical Dependency-Guided Sequential Inference

基于物理依赖引导的深度概率室内气体源定位的序贯推断
Kim, Seunghwan, Kim, Hyungjin, Lee, Junhee, Oh, Hyondong
Abstract
Reliable gas source localization (GSL) is critical to safety in industrial and urban environments, yet remains challenging indoors because walls and obstacles interact with airflow to create complex gas dispersion. High-fidelity models such as computational fluid dynamics and filament models can capture these effects, but their computational cost limits online use. We propose a deep probabilistic framework that infers the source posterior from sparse and noisy measurements collected by a mobile robot. Unlike end-to-end models that directly infer source estimates from measurements, the proposed method incorporates physical dependencies of indoor gas transport, where wind and source location govern the concentration field. These dependencies are embedded through sequential conditional inference, in which inferred wind and concentration fields guide source posterior estimation. This structure improves localization under sparse and noisy observations. Evaluations show that the proposed method outperforms representative GSL baselines and enables accurate and efficient active GSL in simulations. Real-robot experiments demonstrate the feasibility of online operation on an embedded GPU.
Chinese Translation
可靠的气体源定位(GSL)对工业和城市环境的安全至关重要,但在室内环境中仍然具有挑战性,因为墙壁和障碍物与气流相互作用,导致复杂的气体扩散。高保真模型,如计算流体动力学和细丝模型,可以捕捉这些效应,但其计算成本限制了在线使用。我们提出了一种深度概率框架,通过移动机器人收集的稀疏和噪声测量推断源后验。与直接从测量中推断源估计的端到端模型不同,所提出的方法结合了室内气体传输的物理依赖性,其中风和源位置决定了浓度场。这些依赖性通过序贯条件推断嵌入,其中推断的风和浓度场引导源后验估计。该结构在稀疏和噪声观察下改善了定位效果。评估结果表明,所提出的方法优于代表性的GSL基线,并在模拟中实现了准确和高效的主动GSL。真实机器人实验展示了在嵌入式GPU上在线操作的可行性。
cs.RO / 61 / 2608.16222

HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction

HiPHI:高精度人类运动与物体交互的大规模基准
Ji, Jiahao, Ma, Ji, Zhang, Runhan, Yu, Runyi, Wang, Wenjia, Chi, Weiheng, Peng, Qianqian, Yan, Weichao, Gu, Yongfei, Tian, Ye, Wu, Ting, Li, Longwei, Yuan, Chun, Dai, Ruoli, Han, Lei
Abstract
Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically grounded interactions. However, existing embodied datasets remain fundamentally limited: internet-scale video data lack precise physical states and interaction grounding, while laboratory motion datasets provide high fidelity but only narrow behavioral coverage. This mismatch creates a critical bottleneck for scalable humanoid policy learning. We present HiPHI, a 600+ hour scale high-fidelity whole-body human motion dataset designed to systematically maximize coverage of the human motion and interaction manifold. HiPHI is theoretically guided by FrameNet, a linguistic framework organizing human primitives. Created using an optical motion capture pipeline, HiPHI provides sub-millimeter spatial marker tracking accuracy for full-body human motion and mesh-level object trajectories. We further introduce a benchmark suite evaluating motion-space diversity, interaction grounding, object consistency, and physical AI applications. Our analyses demonstrate that HiPHI significantly expands motion coverage compared to existing motion datasets while maintaining high-fidelity interaction quality, and establishes a scalable data foundation for training, evaluating, and generalizing humanoid policies in real-world embodied tasks, where similar extensions are also applicable to motion prior models in computer graphics.
Chinese Translation
类人智能需要在极其多样化的全身运动和物理基础交互空间中进行学习。然而,现有的具身数据集在根本上仍然有限:互联网规模的视频数据缺乏精确的物理状态和交互基础,而实验室运动数据集提供高保真度但仅覆盖狭窄的行为范围。这种不匹配造成了可扩展类人策略学习的关键瓶颈。我们提出了HiPHI,一个超过600小时的高保真全身人类运动数据集,旨在系统性地最大化人类运动和交互流形的覆盖。HiPHI在理论上受到FrameNet的指导,这是一个组织人类原语的语言框架。HiPHI采用光学运动捕捉管道创建,提供亚毫米级的空间标记跟踪精度,适用于全身人类运动和网格级物体轨迹。我们进一步引入了一个基准套件,用于评估运动空间的多样性、交互基础、物体一致性和物理人工智能应用。我们的分析表明,与现有运动数据集相比,HiPHI显著扩展了运动覆盖,同时保持高保真的交互质量,并为在现实世界具身任务中训练、评估和推广类人策略建立了可扩展的数据基础,其中类似的扩展也适用于计算机图形学中的运动先验模型。
cs.RO / 62 / 2608.16229

Planner-Conditioned Diffusion for Coordinated Multi-Agent Exploration

基于规划者条件的扩散方法用于协调多智能体探索
Teo, Marcus Yu Siong, Lew, Jeric, Duhan, Tanishq, Sartoretti, Guillaume
Abstract
Coordinated multi-agent exploration requires not only efficient individual coverage but also non-redundant coverage across agents over extended planning horizons. Conventional approaches rely on hand-crafted coordination rules, while end-to-end multi-agent learning methods are difficult to scale and train. Diffusion-based planners such as DARE offer a promising alternative by generating long-horizon trajectories instead of single-step actions, but existing methods are trained on a narrow planner distribution, limiting behavioral diversity and inference-time controllability. We propose a Planner-Conditioned Diffusion Policy (PCDP) for graph-based multi-agent exploration. PCDP is trained on demonstrations from multiple planner styles with planner identity as an explicit conditioning input, enabling a single shared model to learn a multimodal trajectory distribution and generate diverse, controllable trajectory candidates from the same observation. Rather than learning coordination end-to-end, we reuse this multimodal single-agent policy across all agents and introduce coordination through local reranking, in which nearby agents jointly select the trajectory combination with minimal predicted overlap. We evaluate PCDP against classical and diffusion-based baselines on 100 held-out maps in a four-agent simulation setting. PCDP matches the perfect success rate of the diffusion-based baselines while improving mean max-agent travel, total team travel, and agent imbalance. Crucially, reranking alone over a single-planner baseline yields only marginal gains, indicating that planner-conditioned multimodality is the main contributor to improved coordination. Qualitative simulation results and real-robot experiments with two agents further validate that diverse long-horizon trajectory generation produces emergent spatial separation between agents without any explicit repulsion mechanism.
Chinese Translation
协调多智能体探索不仅需要高效的个体覆盖,还需要在较长的规划时间内实现智能体之间的非冗余覆盖。传统方法依赖于手工设计的协调规则,而端到端的多智能体学习方法则难以扩展和训练。基于扩散的规划方法,如 DARE,提供了一种有前景的替代方案,通过生成长时间跨度的轨迹而非单步动作,但现有方法在狭窄的规划者分布上进行训练,限制了行为多样性和推理时的可控性。我们提出了一种基于规划者条件的扩散策略(PCDP)用于图形基础的多智能体探索。PCDP 在来自多种规划者风格的演示上进行训练,并将规划者身份作为显式条件输入,使得单一共享模型能够学习多模态轨迹分布,并从相同的观察中生成多样化、可控的轨迹候选。我们并不是端到端地学习协调,而是将这种多模态单智能体策略在所有智能体之间重用,并通过局部重新排序引入协调,其中附近的智能体共同选择预测重叠最小的轨迹组合。我们在四智能体仿真环境中对 100 个保留地图上的 PCDP 进行了评估,与经典和基于扩散的基线进行比较。PCDP 达到了基于扩散的基线的完美成功率,同时改善了平均最大智能体旅行、总团队旅行和智能体不平衡。关键是,仅对单一规划者基线进行重新排序仅带来了边际收益,这表明规划者条件的多模态性是改善协调的主要因素。定性仿真结果和两智能体的真实机器人实验进一步验证了多样化的长时间跨度轨迹生成在没有任何显式排斥机制的情况下产生了智能体之间的自发空间分离。
cs.RO / 63 / 2608.16264

Cyclops: LiDAR as a Camera That Dreams in Color

独眼巨人:将激光雷达视为梦中的彩色相机
Gao, Wei, Shu, Jian, Zhao, Mingle, Ghaffari, Maani, Kong, David, Xu, Chengzhong, Kong, Hui
Abstract
Conventionally, robotic perception relies heavily on cameras due to the rich semantic texture they provide. However, their performance degrades significantly in low-light or high-dynamic-range environments. Conversely, while Light Detection and Ranging (LiDAR) captures illumination-invariant geometric and intensity properties, the resulting data are typically single-channel and sparse, creating a significant modality gap when applying vision models pre-trained on RGB datasets. In this paper, we propose Cyclops, a framework that translates sparse Non-Repetitive Scanning LiDAR (NRS-LiDAR) intensity into RGB video, enabling camera-free inference for all-day perception tasks. Our approach first converts sparse LiDAR intensity projections into dense representations via a frozen pre-trained densification module, serving as a geometrically rich source condition. The dense intensity latent is then transported toward the target RGB distribution through Latent Bridge Matching (LBM) with a learned velocity field in a few ODE integration steps. To mitigate inter-frame flickering, we inject prior-frame context via temporal attention layers and further formulate the velocity field as a policy optimized by a differentiable terminal reward that encourages terminal fidelity through backpropagation along the ODE trajectory. Extensive experiments demonstrate that the synthesized RGB, including those generated under near-dark conditions, enable standard RGB-based perception models to substantially outperform both LiDAR baselines and conventional cameras on semantic segmentation, lane detection, and point cloud colorization across diverse lighting conditions.
Chinese Translation
传统上,机器人感知在很大程度上依赖于相机,因为它们提供了丰富的语义纹理。然而,在低光或高动态范围环境中,相机的性能显著下降。相反,光探测与测距(LiDAR)捕获了与光照无关的几何和强度特性,但所产生的数据通常是单通道且稀疏的,这在将预先训练的视觉模型应用于RGB数据集时造成了显著的模态差距。在本文中,我们提出了Cyclops,一个将稀疏的非重复扫描LiDAR(NRS-LiDAR)强度转换为RGB视频的框架,从而实现全天候感知任务的无相机推理。我们的方法首先通过一个冻结的预训练稠密化模块将稀疏的LiDAR强度投影转换为稠密表示,作为几何丰富的源条件。然后,稠密强度潜变量通过在少数常微分方程(ODE)积分步骤中与学习到的速度场进行潜在桥接匹配(Latent Bridge Matching,LBM),被传输到目标RGB分布。为了减轻帧间闪烁,我们通过时间注意层注入前一帧的上下文,并进一步将速度场表述为一个策略,该策略通过可微分的终端奖励进行优化,鼓励沿ODE轨迹的终端保真度。大量实验表明,合成的RGB图像,包括在近乎黑暗条件下生成的图像,使得基于标准RGB的感知模型在语义分割、车道检测和点云着色等任务中,显著超越了LiDAR基线和传统相机,适用于多种照明条件。
cs.RO / 64 / 2608.16281

Marker-Constrained Pose-Graph Correction for Cross-Platform Georeferencing in GNSS-Denied Environments

基于标记约束的姿态图校正用于GNSS受限环境下的跨平台地理参考
Giberna, Marco, Lopez, Jose Luis Sanchez, Voos, Holger
Abstract
Autonomous operation in GNSS-denied environments requires heterogeneous mapping pipelines to maintain a consistent spatial reference. This paper presents a framework using camouflage-matched fiducial markers fabricated from Cholesteric Spherical Reflectors (CSRs) as pre-surveyed visual anchors. The anchors georeference both a lightweight LiDAR-odometry trajectory and a dense RTAB-Map reconstruction, allowing their outputs to be expressed in a common LUREF frame (geodetic coordinate reference system used in Luxembourg) without requiring GNSS measurements during operation. The method combines coarse similarity alignment with marker-constrained pose-graph optimization. We evaluate it using two handheld acquisition sessions with ground-level and elevated motion profiles emulating UGV and UAV operation. A single iMarker was relocated among six surveyed positions, with the first position revisited to quantify drift correction. Marker-anchor correction reduced revisit inconsistency by 97.9% and 99.1% for the UAV- and UGV-emulating sessions, respectively, and improved held-out anchor prediction compared with one-time alignment. Separately georeferenced dense reconstructions achieved a median cross-session nearest-neighbour distance of 58 cm without explicit cross-session registration. Marker processing operated in real time, while trajectory correction required less than 0.25 s per session. These results demonstrate a proof of concept for georeferencing lightweight odometry and dense reconstructions using visually unobtrusive, pre-surveyed anchors during GNSS-denied operation.
Chinese Translation
在GNSS受限环境中,自动操作需要异构映射管道以维持一致的空间参考。本文提出了一种框架,使用由胆甾醇球面反射器(Cholesteric Spherical Reflectors, CSRs)制成的伪装匹配的基准标记作为预先勘测的视觉锚点。这些锚点为轻量级激光雷达里程计轨迹和密集的RTAB-Map重建提供地理参考,使其输出能够在不需要GNSS测量的情况下以共同的LUREF框架(卢森堡使用的地理坐标参考系统)表示。该方法结合了粗略相似性对齐与标记约束的姿态图优化。我们通过两次手持采集会话进行评估,模拟了UGV(无人地面车辆)和UAV(无人机)的地面和高空运动轨迹。一个iMarker在六个勘测位置之间重新定位,首次位置被重新访问以量化漂移校正。标记锚点校正使得UAV和UGV模拟会话的重访不一致性分别减少了97.9%和99.1%,并且与一次性对齐相比,改善了保留锚点的预测。单独地理参考的密集重建在没有显式跨会话注册的情况下实现了58厘米的中位数跨会话最近邻距离。标记处理实时进行,而轨迹校正每次会话所需时间少于0.25秒。这些结果展示了在GNSS受限操作中,使用视觉上不显眼的预勘测锚点对轻量级里程计和密集重建进行地理参考的概念验证。
cs.RO / 65 / 2608.16335

Readiness Barrier Functions: Forward-Invariant Control Authority for Overactuated Multirotor Allocation

准备障碍函数:过驱动多旋翼分配的前向不变控制权
Silano, Giuseppe
Abstract
Allocation schemes that greedily maximize a readiness metric over the actuator fiber bundle of an overactuated multirotor produce commands that jump between disconnected optimal strata, demanding actuator rates no motor can deliver; effort-minimizing schemes are continuous but cannot guarantee that wrench-rate authority stays above any certified level. We reconcile the two by treating authority as a forward-invariant quantity: a control barrier function on the log-determinant of the drag-aware actuator-authority co-metric, enforced at torque level by a quadratic program in the allocation null space. A single design inequality renders the certified set compact and strictly interior to the actuator box, with the readiness cost of any rotor deactivation given in closed form as $\ln(n/(n{-}m))$ for symmetric designs. Tracking is sacrificed only through an explicit alignment ratio, with wrench error bounded by $\mathcal{O}(\rho^{-1/2})$ and a robust variant handles motor-parameter uncertainty with a closed-form floor shift independent of the airframe matrix. On a hexarotor and a fully-actuated octorotor the closed-form gap matches simulation to machine precision; in the authority-scarce regime greedy maximization violates the certified floor and commits wrench errors up to eighty times larger than the proposed filter, which holds invariance of the certified set at negligible tracking cost.
Chinese Translation
在过驱动多旋翼的执行器纤维束上,贪婪地最大化准备度指标的分配方案产生的指令在不相连的最优层次之间跳跃,要求的执行器速率超出了任何电机的承载能力;而最小化努力的方案虽然是连续的,但无法保证扭矩速率权威保持在任何认证水平之上。我们通过将权威视为前向不变量来调和这两者:在拖曳感知的执行器权威共度量的对数行列式上施加控制障碍函数,通过在分配零空间中的二次规划在扭矩水平上强制执行。一个单一的设计不等式使得认证集紧凑且严格位于执行器盒的内部,任何旋翼停用的准备成本以封闭形式给出,形式为 $ ext{ln}(n/(n{-}m))$,适用于对称设计。跟踪仅通过显式对齐比率被牺牲,扭矩误差被限制在 $ ext{O}( ho^{-1/2})$,而一种鲁棒变体处理电机参数的不确定性,具有与机体矩阵无关的封闭形式底部偏移。在六旋翼和完全驱动的八旋翼上,封闭形式的差距与仿真精确匹配;在权威稀缺的情况下,贪婪最大化违反了认证底线,并导致的扭矩误差比所提过滤器大出多达八十倍,而该过滤器在可忽略的跟踪成本下保持认证集的不变性。
cs.RO / 66 / 2608.16351

Arm-Aware Guided Dexterous Grasp Generation with Arm-Agnostic Grasp Models

考虑手臂的引导灵巧抓取生成与无关手臂抓取模型
Jia, Yongyi, Jiang, Yongpeng, Lv, Kangchen, Ren, Yi, Yu, Mingrui, Li, Xiang
Abstract
Dexterous grasp generation that considers arm-related constraints is crucial in real-world scenarios involving arm environment collision avoidance, workspace boundary grasps, and consecutive grasping. Existing hand-centric grasp models, which primarily focus on the floating hand's pose, are insufficient for such cases. Conventional arm-aware methods either rely on rejection sampling to discard infeasible samples or require retraining on arm-specific data, leading to low sample efficiency under adverse conditions or limited generalization across different robots and environments. To overcome these limitations, this letter presents an arm-aware dexterous grasp generation framework that leverages pretrained arm-agnostic grasp models while integrating arm and environmental information only at inference time. Specifically, we formulate arm-aware constrained grasp generation as a joint optimization of hand pose and arm configuration, and derive closed-form gradients for arm-related constraints. Assuming the hand pose distribution is represented by a diffusion model, we prove that gradient-based optimization is equivalent to guided diffusion sampling, steering near-feasible samples toward the feasible region. Through comprehensive evaluation involving 10k objects across 6 scenarios, we demonstrate that the proposed framework generates feasible grasps in highly constrained settings with significantly higher probability, highlighting its advantages in real-world applications. Supplementary materials and appendix are available at https://arm-aware-dexgrasp.github.io/.
Chinese Translation
考虑手臂相关约束的灵巧抓取生成在涉及手臂环境碰撞避免、工作空间边界抓取和连续抓取的现实场景中至关重要。现有的以手为中心的抓取模型主要关注浮动手的姿态,无法满足此类情况的需求。传统的考虑手臂的方法要么依赖拒绝采样来丢弃不可行的样本,要么需要在特定于手臂的数据上重新训练,这导致在不利条件下样本效率低下或在不同机器人和环境之间的泛化能力有限。为克服这些局限性,本文提出了一种考虑手臂的灵巧抓取生成框架,该框架利用预训练的无关手臂抓取模型,并在推理时仅集成手臂和环境信息。具体而言,我们将考虑手臂的约束抓取生成形式化为手部姿态和手臂配置的联合优化,并推导出与手臂相关约束的封闭形式梯度。假设手部姿态分布由扩散模型表示,我们证明基于梯度的优化等价于引导扩散采样,将近乎可行的样本引导至可行区域。通过对6种场景中10,000个物体的全面评估,我们证明了所提出的框架在高度受限的环境中生成可行抓取的概率显著更高,突显了其在现实应用中的优势。补充材料和附录可在 https://arm-aware-dexgrasp.github.io/ 获取。
cs.RO / 67 / 2608.16433

Robot-Body-Aware Traversal Risk Graph Planning for Wheeled-Legged Robots in Complex Terrain

面向复杂地形的轮腿机器人机器人身体感知穿越风险图规划
Guo, Zhiqiao, Zhang, Bichi, Schwertfeger, Sören
Abstract
Traversal Risk Graphs (TRGs) provide a compact, terrain-aware representation for global navigation, but native TRG costs are computed over circular node neighborhoods and edge-aligned terrain regions rather than the robot's oriented body footprint. For wheeled-legged robots, this abstraction can miss partial support loss and body-terrain interference, especially during turns. We present Robot-Body-Aware TRG planning (RB-TRG), which builds on the sparse TRG representation and lifts edge-wise terrain-risk search to heading- and turn-aware body-risk transitions. An oriented rectangular footprint is sampled along graph edges and yaw sweeps to measure longitudinal support variation, lateral inclination, terrain interference, and exposure to untrusted map regions. Mean-and-upper-tail features are incorporated into transition costs, whose accumulated value is minimized by A* over ordered node-pair states, preserving TRG construction and its planning interface. We evaluate RB-TRG in a same-graph study on four scanned terrain environments and in paired closed-loop MuJoCo trials. RB-TRG reduces the three core geometric body-placement metrics and increases end-to-end success from 51.5% to 68.5%, while increasing mean path length by 2.3%. A Go2-W deployment further demonstrates RB-TRG with a full LiDAR navigation stack, which received the Best Autonomy and Best Mobility awards at the IEEE ICRA 2026 Legged Robot Challenges. The code for RB-TRG is released at https://github.com/ZhiqiaoGuo/RB-TRG.
Chinese Translation
穿越风险图(TRGs)为全球导航提供了一种紧凑的、地形感知的表示方式,但原生TRG成本是基于圆形节点邻域和边对齐的地形区域计算的,而不是基于机器人的定向身体足迹。对于轮腿机器人而言,这种抽象可能会忽略部分支撑丧失和身体与地形的干扰,尤其是在转弯时。我们提出了机器人身体感知的TRG规划(RB-TRG),该方法基于稀疏TRG表示,并将边缘地形风险搜索提升到考虑航向和转弯的身体风险过渡。沿着图的边缘和偏航扫掠采样定向矩形足迹,以测量纵向支撑变化、横向倾斜、地形干扰以及暴露于不可信地图区域的风险。过渡成本中融入了均值和上尾特征,其累积值通过A*算法在有序节点对状态上最小化,从而保持TRG的构建及其规划接口。我们在四个扫描地形环境的同图研究和配对闭环MuJoCo试验中评估了RB-TRG。RB-TRG减少了三个核心几何身体放置指标,并将端到端成功率从51.5%提高到68.5%,同时平均路径长度增加了2.3%。Go2-W部署进一步展示了RB-TRG与完整的LiDAR导航堆栈,该堆栈在IEEE ICRA 2026腿式机器人挑战赛中获得了最佳自主性和最佳移动性奖项。RB-TRG的代码已发布在https://github.com/ZhiqiaoGuo/RB-TRG。
cs.RO / 68 / 2608.16442

Observation-Constrained Joint-Space Viewpoint Optimization for Robotic Inspection of Cylindrical Cavities

基于观察约束的关节空间视角优化用于圆柱腔体的机器人检测
Wang, Yuezhong, Yin, Rongshen, Zhang, Bichi, Schwertfeger, Sören
Abstract
Inspection is a core capability in many mobile robotics applications, including industrial facility monitoring, infrastructure maintenance, agriculture, and search and rescue. Observing the bottom of a cylindrical cavity, as required by ASTM search-task benchmarks for response robots, presents a representative challenge: the robot must position its camera precisely while satisfying visibility, kinematic, and collision constraints. This paper presents a fully autonomous method for observation-constrained inspection of cylindrical cavities in robot joint space. Rather than prescribing a single Cartesian camera pose, the method represents the inspection objective as a set of valid viewing geometries, thereby avoiding the rejection of reachable viewpoints and configurations with poor joint-limit margins. An RGB perception front end estimates the opening center and directed cavity axis from semantic masks using arc-supported ellipse fitting together with body and side-generator cues. These estimates parameterize constraints on camera-axis alignment, lateral offset, and axial standoff. A multistart derivative-free search then optimizes robot joint configurations with lexicographic priority given to constraint satisfaction; feasible configurations are ranked according to motion economy, joint-limit margin, and view quality. The resulting candidates are evaluated by a collision-aware motion planner, and the executed camera pose is verified geometrically and using a ray-based estimate of bottom visibility. In Isaac Sim, the proposed method successfully completes 92 of 100 target configurations and attains 91.65% mean bottom visibility among executed trials, compared with 76 of 100 and 84.3% for a multistart coordinate-search baseline. Tabletop and Unitree A2-mounted experiments demonstrate the complete perception-planning-execution pipeline.
Chinese Translation
检测是许多移动机器人应用中的核心能力,包括工业设施监测、基础设施维护、农业以及搜索与救援。观察圆柱腔体底部是响应机器人所需的 ASTM 搜索任务基准中的一个典型挑战:机器人必须在满足可见性、运动学和碰撞约束的情况下精确定位其相机。本文提出了一种完全自主的方法,用于在机器人关节空间中进行观察约束的圆柱腔体检测。该方法并不是规定单一的笛卡尔相机姿态,而是将检测目标表示为一组有效的视角几何,从而避免拒绝可达的视点和具有较差关节极限边际的配置。RGB 感知前端利用弧支持的椭圆拟合以及身体和侧面生成线索,从语义掩膜中估计开口中心和定向腔体轴。这些估计参数化了相机轴对齐、横向偏移和轴向间距的约束。然后,进行多起始的无导数搜索,以优化机器人关节配置,优先考虑约束满足;可行配置根据运动经济性、关节极限边际和视图质量进行排名。最终候选配置通过一个考虑碰撞的运动规划器进行评估,执行的相机姿态通过几何方法和基于光线的底部可见性估计进行验证。在 Isaac Sim 中,所提出的方法成功完成了 100 个目标配置中的 92 个,并在执行试验中达到了 91.65% 的平均底部可见性,而多起始坐标搜索基线的完成率为 76 个中的 100 个,平均底部可见性为 84.3%。桌面和 Unitree A2 安装的实验展示了完整的感知-规划-执行管道。
cs.RO / 69 / 2608.16476

Exposing the Long-tail in Embodied Urban Navigation via Scalable Learning from In-the-Wild Videos

通过从野外视频中可扩展学习揭示具身城市导航中的长尾现象
Xia, Bingyi, Bao, Han, Chen, Zhewei, Ye, Hanjing, Yu, Jingwen, Pang, Yuhan, Xu, Wenjun, Wang, Jiankun
Abstract
Learning embodied urban navigation policies from real-world data is constrained by the cost of task-specific data collection and the limited coverage of rare yet safety-critical scenarios. To address these challenges, we present a scalable framework for learning point-goal urban navigation from web-scale in-the-wild egocentric videos while systematically exposing its long tail. The framework automatically annotates uncurated web videos with metric trajectories and structured navigation semantics, which are then used to train a vision-language-action policy for interpretable navigation planning. We characterize the long tail based on model performance and the distribution of perception-motion patterns, and employ reflection-based analysis to diagnose recurring failure modes. Experiments on web-video data and real-world urban navigation tasks demonstrate effective knowledge transfer from unconstrained videos and reveal coherent long-tail structures beyond aggregate navigation performance.
Chinese Translation
从真实世界数据中学习具身城市导航策略受到任务特定数据收集成本和稀有但安全关键场景有限覆盖的限制。为了解决这些挑战,我们提出了一个可扩展的框架,从网络规模的野外自我中心视频中学习点目标城市导航,同时系统性地揭示其长尾现象。该框架自动为未经整理的网络视频注释度量轨迹和结构化导航语义,然后用于训练可解释的导航规划的视觉-语言-行动策略。我们基于模型性能和感知-运动模式的分布来表征长尾,并采用基于反思的分析来诊断重复出现的失败模式。在网络视频数据和真实世界城市导航任务上的实验表明,从不受限的视频中有效地转移知识,并揭示了超出总体导航性能的一致长尾结构。
cs.RO / 70 / 2608.16499

OccamView: Object-Conditioned View Selection for Frame-Budgeted Active 3D Gaussian Reconstruction

OccamView:基于对象条件的视图选择用于帧预算的主动3D高斯重建
Gao, Hongbo, Zhang, Wei, Ni, Zeyu, Zhu, Dihao, Li, Ruifeng, Wang, Yunke, Xu, Chang
Abstract
Active 3D Gaussian reconstruction fundamentally relies on selecting informative next-best views under limited sensing budgets. Existing active 3DGS methods primarily plan viewpoints according to geometric information gain, treating object-induced hidden regions in the same manner as general unexplored space. Under tight frame budgets, such geometry-driven strategies may prioritize global scene coverage while leaving partially observed objects incompletely reconstructed. To address this limitation, we propose OccamView, an object-conditioned view-selection framework for frame-budgeted active 3D Gaussian reconstruction. Rather than predicting unseen object geometry or performing shape completion, OccamView maintains an online object memory from open-vocabulary detections grounded in measured RGB-D observations and represents unresolved local occupancy around detected objects as conservative hidden-region proxies. Candidate viewpoints are then evaluated using an occlusion-aware proxy-coverage score. Furthermore, we introduce a Geo-Floor mechanism that restricts object-conditioned re-ranking to geometrically competitive candidates, allowing object-conditioned cues to guide complementary observations while preserving the geometry-driven exploration behavior of the underlying planner. Experiments on Replica and Matterport3D under a unified frame-budgeted protocol show that OccamView consistently reduces Completion and improves Completion Ratio across five frame budgets, with particularly pronounced gains under limited frame budgets. These results demonstrate that lightweight object-conditioned cues effectively complement geometry-driven active view planning.
Chinese Translation
主动3D高斯重建基本上依赖于在有限的传感预算下选择信息丰富的下一个最佳视图。现有的主动3DGS方法主要根据几何信息增益规划视点,将对象引起的隐藏区域与一般未探索空间视为相同。在严格的帧预算下,这种以几何为驱动的策略可能优先考虑全球场景覆盖,而导致部分观察到的对象重建不完整。为了解决这一局限性,我们提出了OccamView,一个用于帧预算的主动3D高斯重建的对象条件视图选择框架。OccamView并不是预测未见的对象几何或执行形状补全,而是维护一个来自基于测量的RGB-D观察的开放词汇检测的在线对象记忆,并将检测到的对象周围未解决的局部占用表示为保守的隐藏区域代理。然后,候选视点使用考虑遮挡的代理覆盖评分进行评估。此外,我们引入了一种Geo-Floor机制,限制对象条件的重新排序仅针对几何上具有竞争力的候选者,从而允许对象条件线索引导补充观察,同时保留底层规划器的几何驱动探索行为。在统一的帧预算协议下对Replica和Matterport3D的实验表明,OccamView在五个帧预算下始终减少了完成度并提高了完成比率,尤其在有限帧预算下表现出显著的提升。这些结果表明,轻量级的对象条件线索有效地补充了以几何为驱动的主动视图规划。
cs.RO / 71 / 2608.16503

NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation

NebulaVLA:一种具有引导动作的双频视觉-语言-动作模型用于机器人操作
Zhao, Cong, Tian, Shuai, Zhang, Xu, Ni, Baocheng, Song, Xinguo, Sun, Xueying, Jiang, Shu, Yang, Shouchang, Tang, Bo, Deng, Jin, Zhu, Ge, Wang, YongCheng, Xu, Jin, Yang, Ri
Abstract
Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness. We present NebulaVLA, an asynchronous dual-frequency architecture that decouples high-level semantic reasoning from low-level action control, optimizing computational resources and modularity. To bridge semantic gaps across heterogeneous robots, we introduce GESTURE-7, a unified language-grounded action representation. Furthermore, our Guide Action algorithm enforces kinematic continuity via mask-based smoothness constraints. Comprehensive evaluations demonstrate that NebulaVLA significantly outperforms synchronous baselines, achieving an 85.5\% average success rate on LIBERO-Plus and accelerating action generation by \textasciitilde 2.7$\times$. This asynchronous design enables highly efficient and responsive control for practical robotics.
Chinese Translation
视觉-语言-动作(VLA)模型在现实世界中的应用常常受到效率与性能权衡、跨实体泛化和执行平滑性的瓶颈限制。我们提出了NebulaVLA,一种异步双频架构,将高层语义推理与低层动作控制解耦,从而优化计算资源和模块化设计。为了弥合异构机器人之间的语义差距,我们引入了GESTURE-7,这是一种统一的语言基础动作表示。此外,我们的引导动作(Guide Action)算法通过基于掩码的平滑性约束强制执行运动学连续性。全面的评估表明,NebulaVLA显著优于同步基线,在LIBERO-Plus上实现了85.5\%的平均成功率,并将动作生成加速约2.7倍。这种异步设计使得实际机器人控制变得高效且响应迅速。
cs.RO / 72 / 2608.16555

Co-design of Neural and Muscle Network based on Embodied Perceptron Representation

基于具身感知器表征的神经与肌肉网络协同设计
Tao, Siyuan, Masuda, Yoichi, Nabae, Hiroyuki, Ishikawa, Masato
Abstract
Recent advances in AI technologies have enabled the advanced design of complex control policies. In contrast, focusing on the body, many robots still employ simple bodies that can limit adaptability to environments. Studies in embodied robotics have shown that well-designed bodies can partially replace the role of control and computation with physical body-environment interactions, yet such designs still depend heavily on expert intuition. There is a need for a systematic theoretical framework for body design, as well as a method for joint optimization of the body and controller. To address this, we introduce the Embodied Perceptron, a theoretical framework that unifies neural networks and physical body systems. In this view, the body itself acts as a perceptron: mechanical parameters correspond to weights, and physical nonlinearities play the role of activation functions. By representing physical constraints as weights and nonlinear properties as activation functions, a physical body can be modeled in neural-network form. The system representation enables us to explicitly and theoretically explain that the body can substitute for part of the neural control. As an application, we co-optimize control policy and muscle configuration in a musculoskeletal robot and show that the resulting embodied intelligence can provide inherent stability, improve learning efficiency, and drastically reduce model size-even with a single-neuron controller. The results bridge the informational and physical worlds and provide a pathway toward understanding and systematic design of embodied AI systems.
Chinese Translation
近期人工智能技术的进步使得复杂控制策略的高级设计成为可能。相较之下,许多机器人仍然专注于身体,采用简单的身体结构,这限制了它们对环境的适应性。具身机器人学的研究表明,设计良好的身体可以通过与环境的物理交互部分替代控制和计算的角色,但这类设计仍然在很大程度上依赖于专家的直觉。因此,迫切需要一个系统的理论框架来指导身体设计,以及一种用于身体和控制器联合优化的方法。为此,我们提出了具身感知器(Embodied Perceptron),这是一个统一神经网络和物理身体系统的理论框架。在这个视角下,身体本身充当感知器:机械参数对应于权重,物理非线性特性则充当激活函数。通过将物理约束表示为权重,将非线性特性表示为激活函数,可以将物理身体建模为神经网络形式。该系统表征使我们能够明确且理论上解释身体可以替代部分神经控制。作为应用,我们在一个肌肉骨骼机器人中共同优化控制策略和肌肉配置,结果表明,所得到的具身智能能够提供内在稳定性,提高学习效率,并显著减少模型规模——即使在使用单神经元控制器的情况下。这些结果架起了信息世界与物理世界之间的桥梁,并为理解和系统设计具身人工智能系统提供了路径。
cs.RO / 73 / 2608.16572

ViHaTeleop: A Low-Cost, Lightweight Visual-Haptic Teleoperation System for Dexterous Manipulation Learning

ViHaTeleop:一种低成本、轻量化的视觉-触觉遥操作系统用于灵巧操作学习
Zhu, Fucai, Lai, Yanhou, Maestre, Paul, Hashimoto, Koichi
Abstract
Learning from demonstration is a promising approach for dexterous manipulation, but collecting high-quality contact-critical demonstrations remains difficult with low-cost teleoperation hardware. We present ViHaTeleop, a lightweight (0.7 kg), low-cost (\$550) visual-haptic teleoperation system with SLAM-based wrist tracking, camera-based hand tracking, and finger-wise vibrotactile feedback through Linear Resonant Actuators (LRA). The system includes several design choices (LED illumination, fisheye hand camera, and tactile-aware retargeting constraints) and is deployed on Franka + LEAP Hand + 9DTact in both real and simulated environments. Under matched with/without-haptic conditions with nine participants across six contact-critical tasks, haptics improved success rates across all tasks (+2.2 to +15.6 percentage points), while completion-time effects were task-dependent. Subjective ratings showed significant gains in contact clarity and grasp confidence in both simulation and real-world settings (Wilcoxon signed-rank, $p<0.05$). We also integrate a lightweight depth-camera-based tactile proxy in Isaac Sim, enabling a full pipeline from multi-modal demonstration collection to visual-tactile policy training. Preliminary downstream validation by training visual-tactile policies from collected demonstrations shows tactile cues benefit contact-critical subtasks (peg-in-hole: +17 percentage points over vision-only).
Chinese Translation
从示范学习是一种有前景的灵巧操作方法,但使用低成本遥操作硬件收集高质量的接触关键示范仍然困难。我们提出了ViHaTeleop,这是一种轻量化(0.7公斤)、低成本(550美元)的视觉-触觉遥操作系统,具备基于SLAM的手腕跟踪、基于摄像头的手部跟踪,以及通过线性共振驱动器(LRA)提供的指尖振动触觉反馈。该系统包括多个设计选择(LED照明、鱼眼手部摄像头和触觉感知重定向约束),并在Franka + LEAP Hand + 9DTact的真实和模拟环境中部署。在九名参与者进行的六个接触关键任务中,匹配的有/无触觉条件下,触觉在所有任务中提高了成功率(增加2.2到15.6个百分点),而完成时间的影响则依赖于任务。主观评分显示,在模拟和真实环境中,接触清晰度和抓取信心显著提高(Wilcoxon符号秩检验,p<0.05)。我们还在Isaac Sim中集成了一种轻量化的基于深度摄像头的触觉代理,实现了从多模态示范收集到视觉-触觉策略训练的完整流程。通过从收集的示范中训练视觉-触觉策略的初步下游验证显示,触觉线索有助于接触关键子任务(孔中插销:比仅使用视觉提高17个百分点)。
cs.RO / 74 / 2608.16590

Zetta $\zeta$: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence

Zetta $$: 一种高效的闭环具身控制系统用于自我进化的物理智能
Ding, Xin, Mi, Liang, Huang, Mingzhe, Wang, Zixuan, Zhang, Chao, Hao, Zixu, Chen, Fu, Li, Xiangyu, Zheng, Yikai, Guo, Yaoyu, Wang, Weijun, Li, Kun, Wu, Hao, Liu, Yunxin, Cao, Ting
Abstract
Embodied agents are increasingly used to close the gap left by end-to-end policy models. Yet the agentic path has not realized closed-loop learning in physical execution: existing harnesses remain largely open-loop, following fixed skills during rollout and reflecting only after an episode completes. Such post-hoc reflection cannot govern execution as it unfolds, because physical interaction requires decisions to track rapidly changing robot-environment states at a frequency beyond today's large agentic models. We present Zetta, a closed-loop embodied harness that evolves code-based runtime critics and recovery skills online while keeping the base policy frozen. Through three timescale-separated loops, Zetta provides action-frequency governance, rollout-level critic-recovery proposal, and validation-gated skill updates. Together with Z-Infra, a rollout infrastructure decoupling agent logic from heterogeneous execution resources, Zetta achieves state-of-the-art success on LIBERO-Pro and RoboCasa under our current rollout budget, reaching 90.8% and 93.6%, with an 11.1x inference speedup; success continues to scale with self-exploration experience; learned skills transfer zero-shot, and clear robotic "Aha Moments" emerge. These results show that closed-loop harness self-evolution opens a scaling path for reliable physical intelligence.
Chinese Translation
具身智能体越来越多地被用于填补端到端策略模型留下的空白。然而,智能体的路径尚未实现物理执行中的闭环学习:现有的控制系统仍然主要是开放式的,在执行过程中遵循固定技能,并且仅在一个回合结束后进行反思。这种事后反思无法在执行过程中进行控制,因为物理交互需要在超出当前大型智能模型的频率下跟踪快速变化的机器人-环境状态。我们提出了Zetta,一种闭环具身控制系统,它在保持基础策略不变的同时,在线进化基于代码的运行时评估者和恢复技能。通过三个时间尺度分离的循环,Zetta提供了动作频率治理、回合级评估者-恢复提案和验证门控技能更新。结合Z-Infra,一个将智能体逻辑与异构执行资源解耦的回合基础设施,Zetta在我们的当前回合预算下在LIBERO-Pro和RoboCasa上取得了最先进的成功,分别达到90.8%和93.6%的成功率,并实现了11.1倍的推理速度提升;成功随着自我探索经验的增加而持续扩展;学习到的技能实现零样本迁移,并且明显的机器人“顿悟时刻”出现。这些结果表明,闭环控制系统的自我进化为可靠的物理智能开辟了一个扩展路径。
cs.RO / 75 / 2608.16640

DPNet: Efficient Dead-End Prediction and Avoidance for Vision-Based UAV Navigation

DPNet:基于视觉的无人机导航中高效死胡同预测与规避
Zhang, Ruibin, Pan, Lun, Xia, Zelong, Hou, Jialiang, Gao, Fei
Abstract
Vision-based Unmanned Aerial Vehicles (UAVs) often suffer from navigation failures in dead ends due to limited sensing accuracy and range. To address this challenge, this paper proposes a systematic solution for efficient dead-end prediction and avoidance. The proposed method introduces a lightweight neural network to predict the relative distance and bearing of potential dead ends within the current field of view using RGB-D inputs. These predictions prune a predefined, compact trajectory library, enabling the planner to proactively avoid dead ends while maintaining navigational smoothness. Notably, our approach transfers across real-world scenarios without manual annotation or fine-tuning on real-world data. The system achieves high-frequency replanning at 50 Hz onboard. Extensive simulation benchmarks demonstrate superior performance in success rate, flight time, and trajectory length, and real-world experiments further validate its effectiveness in complex scenarios.
Chinese Translation
基于视觉的无人机(UAV)由于感知精度和范围的限制,常常在死胡同中遭遇导航失败。为了解决这一挑战,本文提出了一种系统化的高效死胡同预测与规避方案。该方法引入了一种轻量级神经网络,利用RGB-D输入预测当前视野内潜在死胡同的相对距离和方位。这些预测结果对预定义的紧凑轨迹库进行修剪,使得规划器能够主动规避死胡同,同时保持导航的平滑性。值得注意的是,我们的方法能够在真实场景中迁移,无需手动标注或对真实数据进行微调。该系统在机载上实现了50 Hz的高频重新规划。广泛的仿真基准测试表明,在成功率、飞行时间和轨迹长度方面表现优越,真实世界实验进一步验证了其在复杂场景中的有效性。
cs.RO / 76 / 2608.16642

Throwing a Tight Spiral American Football by a Humanoid Robot

由类人机器人投掷紧密螺旋的美式足球
Mahboob, Zaid, Weng, Bowen
Abstract
Accurate throwing of the American football requires precise regulation of release conditions, where coupled linear and angular momentum determine flight stability and targeting accuracy. While prior work on robotic object throwing has largely focused on generating dynamically feasible release velocities using open-gripper paradigms, explicit control of spin injection at detachment remains underexplored, particularly for aerodynamically anisotropic objects like the American football. In this paper, we present the spin-stabilized controlled tight spiral throw of an American football by a humanoid robot. Achieving this requires (i) accurately reaching the desired coupled momentum, which often involves high degrees-of-freedom (DoF) movements completed within approximately half a second, and (ii) managing the complex transient contact dynamics that arise during the sub-100-millisecond release phase, when the football is effectively underactuated as it moves partially across the fingers. To this end, we develop a coupled whole-body control strategy where the lower body is performing informed stabilization while the upper body is further divided into two phases with (i) a throw phase accelerating the football to a target state through trajectory optimization and tracking, and (ii) a follow-through phase utilizing model predictive control to actively control the wrist and remaining in-contact fingers. The proposed framework is empirically validated on a 29-DoF Unitree G1 humanoid equipped with a 7-DoF Dex3-1 three-fingered gripper. The thrown American football reaches up to 93.6% spin efficiency and a 0.286 radians linear-velocity-to-nose-alignment (nose-angle) error (where an ``ideal'' tight spiral corresponds to 100 % spin efficiency and 0 radians nose-angle error) at up to a 5.35 m/s linear velocity and an angular velocity of 14.5 rad/s.
Chinese Translation
准确投掷美式足球需要精确调节释放条件,其中耦合的线性和角动量决定了飞行稳定性和瞄准精度。尽管之前的机器人物体投掷研究主要集中在使用开放抓握范式生成动态可行的释放速度,但在分离时对旋转注入的显式控制仍然未被充分探索,尤其是对于像美式足球这样的空气动力学各向异性物体。本文提出了一种由类人机器人进行的旋转稳定控制的紧密螺旋投掷美式足球的方法。实现这一目标需要 (i) 准确达到所需的耦合动量,这通常涉及在大约半秒内完成的高自由度(DoF)运动,以及 (ii) 管理在小于100毫秒的释放阶段中出现的复杂瞬态接触动力学,此时足球在部分跨越手指时实际上处于欠驱动状态。为此,我们开发了一种耦合的全身控制策略,其中下半身执行有针对性的稳定,而上半身进一步分为两个阶段: (i) 投掷阶段通过轨迹优化和跟踪将足球加速到目标状态,和 (ii) 随后阶段利用模型预测控制主动控制手腕和接触的手指。所提出的框架在配备7自由度Dex3-1三指抓手的29自由度Unitree G1类人机器人上进行了实证验证。投掷的美式足球达到了93.6%的旋转效率和0.286弧度的线速度与鼻对齐(鼻角)误差(其中“理想”紧密螺旋对应于100%的旋转效率和0弧度的鼻角误差),线速度高达5.35 m/s,角速度为14.5 rad/s。
cs.RO / 77 / 2608.16651

Orbit-Planner: Towards Latent World Models for On-Orbit Obstacle Avoidance of Satellite Agents

轨道规划器:面向卫星代理在轨障碍规避的潜在世界模型
Li, Zhijian, Ren, Chao, Wang, Peijin, Sun, Xian
Abstract
Satellite agents for on-orbit navigation tasks need to predict collision risks using limited onboard observations. However, conventional planners often rely on predefined maps and fixed environmental assumptions, limiting their adaptability in dynamic on-orbit scenarios. In this paper, we propose Orbit-Planner, a two-stage latent world model for on-orbit obstacle avoidance. Orbit-Planner learns action-conditioned spacecraft dynamics to perform future-state rollouts in latent space, and introduces a Physics Probe to decode physical state changes from imagined latent trajectories. Experiments demonstrate that Orbit-Planner can perform long-horizon latent rollouts and recover physical states from imagined trajectories. In closed-loop obstacle-avoidance navigation in Isaac Sim, it attains a success rate of 91.7%. Code is available at https://github.com/ZhijianLi2003/Orbit_Planner.
Chinese Translation
用于在轨导航任务的卫星代理需要利用有限的机载观测来预测碰撞风险。然而,传统规划器通常依赖于预定义的地图和固定的环境假设,这限制了它们在动态在轨场景中的适应性。本文提出了轨道规划器(Orbit-Planner),一种用于在轨障碍规避的两阶段潜在世界模型。轨道规划器学习基于动作的航天器动力学,以在潜在空间中执行未来状态的展开,并引入物理探针(Physics Probe)以从想象的潜在轨迹中解码物理状态变化。实验表明,轨道规划器能够执行长时间范围的潜在展开,并从想象的轨迹中恢复物理状态。在Isaac Sim中的闭环障碍规避导航中,其成功率达到91.7%。代码可在https://github.com/ZhijianLi2003/Orbit_Planner获取。
cs.RO / 78 / 2608.16712

H-PAC Hand: Control-Oriented Modeling and Tendon-Elasticity Compensation for an Underactuated Robotic Hand

H-PAC 手:面向控制的建模与腱弹性补偿的欠驱动机器人手
Yan, Teng, Chen, Jiongxu, Wang, Teng, Yu, Yue, Hua, Qixiang, Wang, Zihang, Chen, Yongru, Zhong, Bingzhuo
Abstract
Underactuated tendon-driven hands offer compact actuation and passive compliance, but tendon elongation under restoring-spring loading introduces configuration-dependent joint deviations. This paper presents H-PAC, a modular 6-actuator, 15-DoF robotic hand with a control-oriented modeling and implementation framework. A sparse analytical actuator-joint model is derived from the tendon-routing geometry, and a mechanics-based compensation model is developed to account for tendon-elasticity-induced joint errors. The proposed method is implemented in a hierarchical architecture: a host computer performs workspace-constrained posture mapping and compensation, while an ESP32 generates synchronized commands for six position-controlled servos. The same control parameters and execution strategy are used across all tasks without task-specific retuning. Monotonic servo-sweep experiments show that the compensation substantially improves joint-angle prediction. The MAE of the index DIP joint decreases from 1.15 degrees to 0.18 degrees, and all nine evaluated joints achieve an MAE below 0.23 degrees. Representative postures and grasping configurations are further executed using the same control pipeline without external joint or force sensing in the control loop. The results demonstrate a practical approach to improving posture reproducibility in compact underactuated robotic end-effectors.
Chinese Translation
欠驱动腱驱动手提供了紧凑的驱动和被动的顺应性,但在恢复弹簧加载下,腱的伸长会引入依赖于配置的关节偏差。本文提出了 H-PAC,这是一种模块化的 6 驱动器、15 自由度的机器人手,具有面向控制的建模和实现框架。基于腱路由几何形状推导出稀疏的分析驱动器-关节模型,并开发了基于力学的补偿模型,以考虑腱弹性引起的关节误差。所提出的方法在分层架构中实现:主计算机执行工作空间约束的姿态映射和补偿,而 ESP32 生成六个位置控制伺服器的同步命令。相同的控制参数和执行策略在所有任务中使用,无需针对特定任务的重新调试。单调伺服扫频实验表明,补偿显著改善了关节角度预测。食指 DIP 关节的平均绝对误差(MAE)从 1.15 度降低到 0.18 度,所有九个评估的关节的 MAE 均低于 0.23 度。进一步使用相同的控制管道执行代表性的姿态和抓取配置,而无需在控制回路中进行外部关节或力传感。结果展示了一种实用的方法,以提高紧凑型欠驱动机器人末端执行器的姿态重现性。
cs.RO / 79 / 2608.16715

MatchingPolicy: Correspondence-Aware Policy Enables Cross-Object In-Context Learning

MatchingPolicy:基于对应关系的策略实现跨对象的上下文学习
She, Qijin, Yu, Hanyang, Li, Zeming, Tan, Ping
Abstract
In-context imitation learning enables few-shot policy generalization but struggles to maintain performance on unseen objects and novel scenarios. To address this, we introduce MatchingPolicy, a correspondence-driven framework that explicitly decouples demonstration-to-scene matching from policy learning. Central to our method is a correspondence-aware diffusion policy that conditions robotic actions directly on dense semantic correspondences. This architectural separation resolves the inherent conflict between correspondence identification and action adaptation, enabling robust out-of-distribution transfer. Our framework integrates vision foundation models with a novel two-stage matching algorithm to dynamically establish reliable correspondences. Extensive evaluations on RLBench and real-world manipulation tasks confirm that MatchingPolicy achieves superior few-shot performance, generalizing reliably across unseen object instances and semantic categories.
Chinese Translation
上下文模仿学习能够实现少样本策略泛化,但在未见对象和新场景中保持性能方面存在困难。为了解决这一问题,我们提出了MatchingPolicy,一种以对应关系为驱动的框架,明确将示范与场景匹配从策略学习中解耦。我们方法的核心是一个关注对应关系的扩散策略,该策略直接基于密集的语义对应关系来调节机器人动作。这种架构分离解决了对应关系识别与动作适应之间的固有冲突,使得在分布外的转移更加稳健。我们的框架将视觉基础模型与一种新颖的两阶段匹配算法相结合,以动态建立可靠的对应关系。在RLBench和真实世界的操作任务上的广泛评估证实,MatchingPolicy在少样本性能上表现优越,能够在未见对象实例和语义类别之间可靠地泛化。
cs.RO / 80 / 2608.16728

Design Optimization for Large High-Force Soft Robot Manipulators Under Gravitational Loads

在重力负载下的大型高强度软机器人操纵器的设计优化
Cholaseuk, Isara, Llibre, Penelope, Kyriacou, Alexa, Wang, Audrey, Dickson, Akua K., Jing, Ran, Garcia, Juan C. Pacheco, Sabelhaus, Andrew P.
Abstract
Designing large soft robots capable of generating high forces for physical human-robot interaction remains a significant challenge in soft robotics. Prior work in large soft robots has focused on proof-of-concept prototypes, and no systematic framework exists for determining the suitability of a design paradigm for a desired task. This manuscript introduces a method for optimizing the geometry of a soft robot limb, maximizing its blocking force subject to an anti-bucking constraint under its own gravitational loading. We demonstrate that an explicit solution exists to the proposed optimization problem under certain assumptions. Experiments with three geometries of a large, soft, pneumatically-actuated manipulator demonstrate that the method correctly predicts which designs meet constraints and which produces the largest end-effector forces. This method, with its closed-form solution, can allow designers to determine a-priori if an intended class of soft manipulators is an appropriate choice for physical interaction at large size scales.
Chinese Translation
设计能够产生高强度以实现物理人机交互的大型软机器人仍然是软机器人领域的一项重大挑战。先前的大型软机器人研究主要集中在概念验证原型上,尚未建立系统框架以确定设计范式在特定任务中的适用性。本文介绍了一种优化软机器人肢体几何形状的方法,旨在最大化其在自身重力负载下的抗弯曲约束条件下的阻挡力。我们证明在某些假设下,所提出的优化问题存在显式解。对三种几何形状的大型气动驱动软操纵器的实验表明,该方法能够正确预测哪些设计满足约束条件,并且哪些设计能够产生最大的末端执行器力。该方法具有封闭形式解,可以使设计师在事先判断某一类软操纵器是否适合于在大尺寸尺度下进行物理交互。
cs.RO / 81 / 2608.16741

Semantic- and Density-Aware Planning for Accessibility-Preserving Multi-Object Placement

语义与密度感知的无障碍多物体放置规划
Wingender, Benno, Dengler, Nils, Busch, Nicolas, Pan, Sicong, Bennewitz, Maren
Abstract
Long-term manipulation planning requires robots to reason not only about immediate task success but also about how current decisions affect future interactions with the environment. In this context, household service robots may need to organize groceries in partially occupied shelves while using limited storage space efficiently and preserving access for subsequent placements. In this paper, we consider an online multi-object shelf-placement setting in which future objects arrivals are unknown. Existing approaches do not jointly address semantic organization, dense space utilization, and manipulator accessibility during sequential shelf filling. To address this gap, we propose Semantic-Dense Placement Planning (SDPP), an accessibility-preserving approach that ranks candidate poses using a semantic-density score combining inter-object semantic similarity with spatial proximity. An Accessibility Map (AM) further filters candidates unlikely to be reachable before motion planning and penalizes placements that reduce the remaining accessible workspace. Simulation experiments show that SDPP significantly improves semantic placement quality over state-of-the-art baselines and achieves the highest average shelf density, while the AM substantially reduces the time required to identify feasible placement poses. A qualitative real-world experiment demonstrates the applicability of our pipeline in a domestic shelf-storage scenario.
Chinese Translation
长期的操作规划要求机器人不仅要考虑当前任务的成功,还要考虑当前决策如何影响未来与环境的交互。在这种背景下,家用服务机器人可能需要在部分占用的货架上组织杂货,同时有效利用有限的存储空间,并为后续放置保留通道。在本文中,我们考虑一种在线多物体货架放置场景,其中未来物体的到达是未知的。现有的方法未能在顺序填充货架时共同解决语义组织、密集空间利用和操控器可达性的问题。为了解决这一空白,我们提出了语义密集放置规划(Semantic-Dense Placement Planning, SDPP),这是一种保留可达性的方案,通过结合物体间的语义相似性与空间接近度的语义密度得分来对候选姿态进行排序。可达性地图(Accessibility Map, AM)进一步过滤那些在运动规划之前不太可能到达的候选项,并惩罚那些减少剩余可达工作空间的放置。仿真实验表明,SDPP显著提高了语义放置质量,相较于最先进的基线方法,达到了最高的平均货架密度,而AM则显著减少了识别可行放置姿态所需的时间。一项定性实地实验展示了我们的方法在家庭货架存储场景中的适用性。
cs.RO / 82 / 2608.16794

Neurosymbolic Embodied Agents

神经符号化具身智能体
Albinhassan, Mohammad, Feng, Yuming, Russo, Alessandra, Madhyastha, Pranava
Abstract
Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate environment dynamics or act on incorrectly grounded entities. We present a neurosymbolic agent that factors long-horizon household tasks into task-directed visual exploration and constrained symbolic planning. In the first phase, a vision-language model and exploration harness acquire goal-relevant predicates and instance bindings from egocentric observations and grounded interactions, producing a symbolic initial state. In the second, a PDDL transition model restricts decoding to tokens that extend applicable actions. Monte Carlo tree search then evaluates executable continuations using a domain-independent planning heuristic. The resulting plans are executable by construction under the transition model, with transfer to the environment conditioned on correct visual grounding. On VirtualHome and ALFWorld, open 4B-27B models exceed 90% success in both environments, and our smallest agent substantially outperforms a 27B direct visual policy in each. Constraints and search prove complementary rather than interchangeable: in ALFWorld either alone solves under a third of tasks, whereas their combination solves over 95%. The method also uses several times fewer generated tokens than extended thinking and far fewer model-visible images than direct interaction, and residual failures localize to state acquisition rather than plan generation without any specialized training.
Chinese Translation
语言和视觉-语言模型生成可信的具身计划,但并不保证可执行性,因为它们的输出可能违反环境动态或作用于错误的基础实体。我们提出了一种神经符号化智能体,将长时间跨度的家庭任务分解为任务导向的视觉探索和受限的符号规划。在第一阶段,视觉-语言模型和探索获取器从自我中心观察和基础交互中获取与目标相关的谓词和实例绑定,生成一个符号初始状态。在第二阶段,PDDL(规划领域定义语言)转移模型限制解码为扩展适用动作的标记。然后,蒙特卡洛树搜索使用领域无关的规划启发式评估可执行的延续。所生成的计划在转移模型下是可执行的,转移到环境的条件是正确的视觉基础。在VirtualHome和ALFWorld上,开放的4B-27B模型在这两个环境中的成功率超过90%,而我们最小的智能体在每个环境中都显著优于27B的直接视觉策略。约束和搜索证明是互补而非可互换的:在ALFWorld中,单独使用任何一种方法都无法解决三分之一的任务,而它们的结合则解决了超过95%的任务。该方法生成的标记数量比扩展思维少几倍,且可视图像数量远低于直接交互,剩余的失败集中在状态获取上,而不是计划生成上,且没有任何专门的训练。
cs.RO / 83 / 2608.16806

When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents

当状态成为攻击面:LLM驱动的具身智能体中的状态语义注入
Liu, Jiawei, Guo, Jiacheng, Zhang, Tian, Xu, Yiwei, Wang, Juan, Fan, Jinlin, Xiao, Bowen, Guo, Chi, Guo, Keyan, Hu, Hongxin
Abstract
Large Language Models (LLMs) have demonstrated capabilities in in-context learning, task decomposition, step-by-step reasoning, and code generation, driving their gradual evolution from text generation models into the core of agents capable of perceiving environments, invoking tools, and executing tasks. Traditional LLM Agents typically obtain information through webpages, documents, databases, or external tools and generate corresponding invocation sequences according to user goals; when this technology is further integrated with robotic systems, large language models begin to undertake functions such as task understanding, high-level planning, and behavioral decision-making. SayCan combines the task reasoning capability of language models with the affordances of robotic skills, while Code as Policies and ProgPrompt generate robot task plans through policy code and programmatic prompting, respectively, and VoxPoser uses language models and vision-language models to construct three-dimensional value maps to guide robotic manipulation \cite{6,7,8,9}. Vision-language-action models such as PaLM-E, RT-2, and GR00T N1 further strengthen the connection among language, visual perception, and robotic actions \cite{10,11,12}. In such LLM-driven embodied agents, the model not only needs to understand user instructions, but also needs to combine scene states, object attributes, spatial relations, and execution feedback to complete task grounding, and then hand the generated action plan to skill libraries, motion planners, or controllers for execution.
Chinese Translation
大型语言模型(LLMs)在上下文学习、任务分解、逐步推理和代码生成方面展现了强大的能力,推动其逐渐从文本生成模型演变为能够感知环境、调用工具和执行任务的智能体核心。传统的LLM智能体通常通过网页、文档、数据库或外部工具获取信息,并根据用户目标生成相应的调用序列;当这一技术进一步与机器人系统集成时,大型语言模型开始承担任务理解、高级规划和行为决策等功能。SayCan将语言模型的任务推理能力与机器人技能的可用性相结合,而Code as Policies和ProgPrompt分别通过策略代码和程序提示生成机器人任务计划,VoxPoser则利用语言模型和视觉-语言模型构建三维价值图以指导机器人操作。在这种LLM驱动的具身智能体中,模型不仅需要理解用户指令,还需要结合场景状态、物体属性、空间关系和执行反馈来完成任务基础,然后将生成的行动计划交给技能库、运动规划器或控制器进行执行。
cs.RO / 84 / 2608.16822

Adaptive Repulsive Pheromone Clustering for Foraging Robot Swarms

自适应排斥信息素聚类用于觅食机器人群体
Pena-Caballero, Carlos, Tarawneh, Constantine, Lu, Qi
Abstract
The Central Place Foraging Algorithm (CPFA) combines site fidelity, pheromone-guided navigation, and uninformed random search to enable decentralized resource collection in robot swarms. However, CPFA often revisits previously explored regions while leaving other areas insufficiently searched, reducing efficiency as resources become scarce. In this paper, we propose Adaptive Repulsive Pheromone Clustering (ARPC), a bio-inspired method in which robots deposit repulsive pheromone waypoints to mark previously explored locations. These waypoints are clustered around the nest to estimate low-value search regions, allowing robots to be redirected toward likely unvisited areas. By integrating the exploitation of known resources with systematic avoidance of redundant exploration, ARPC improves search diversity and resource discovery efficiency. Extensive simulations in ARGoS across varying arena sizes, resource densities, and clustered, random, and power-law spatial distributions demonstrate that ARPC consistently outperforms CPFA and the Grid-Based CPFA (GPFA). In particular, ARPC yields significant gains during both early discovery (10\%) and late-stage (up to 60\%) collection, where conventional methods typically degrade. These results indicate that ARPC provides a scalable and robust strategy for large-scale heterogeneous swarm foraging environments.
Chinese Translation
中央地点觅食算法(CPFA)结合了场所忠诚度、信息素引导导航和无信息随机搜索,以实现机器人群体的分散资源收集。然而,CPFA往往会重访先前探索过的区域,而对其他区域的搜索不足,导致在资源稀缺时效率降低。本文提出了一种自适应排斥信息素聚类(ARPC)方法,这是一种生物启发式方法,机器人通过释放排斥信息素路标来标记已探索的位置。这些路标围绕巢穴聚集,以估计低价值搜索区域,从而使机器人能够被引导到可能未被访问的区域。通过将已知资源的利用与系统性避免冗余探索相结合,ARPC提高了搜索多样性和资源发现效率。在不同竞技场大小、资源密度以及聚类、随机和幂律空间分布下,ARGoS中的广泛仿真表明,ARPC始终优于CPFA和基于网格的CPFA(GPFA)。特别是,ARPC在早期发现(10%)和后期收集(高达60%)中均表现出显著的增益,而传统方法通常会退化。这些结果表明,ARPC为大规模异构群体觅食环境提供了一种可扩展且稳健的策略。
cs.RO / 85 / 2608.16837

HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

HAF:通过层次化动作流和谱潜在强化学习将通用视觉-语言-动作模型适应于类人机器人全身运动操控
Gu, Langzhe, Hou, Chengkai, Li, Meng, Wang, Xinhua, Liu, Jiaming, Lv, Xinyuan, Zhang, Bowei, Bai, Shuanghao, Li, Guangrun, He, Jingyang, Dai, Gaole, Ding, Ziluo, Xu, Zhiyuan, Cheng, Kuan, Tang, Jian, Che, Zhengping, Zhang, Shanghang
Abstract
Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it challenging for conventional single-stage VLA architectures to coordinate locomotion, waist posture, and dual-arm manipulation effectively. Moreover, policies trained through offline behavior cloning can remain suboptimal during real-world deployment. Although online reinforcement learning can refine policies through real-world interaction, directly tuning large VLA backbones demands excessive computation and may introduce safety risks during real-robot exploration. To address these bottlenecks, we introduce HAF (Humanoid Adaptation Framework), a two-part framework consisting of HAF-VLA and HAF-Steer that transfers off-the-shelf generalist VLA foundation models to humanoid whole-body loco-manipulation. HAF-VLA is a hierarchical action-flow generator built on a pretrained flow-matching VLA. It splits full-body action denoising into three sequential stages with stage embeddings and cross-stage KV caches that retain kinematic dependencies, avoiding incoherent whole-body actions from one-shot generation. On top of the frozen HAF-VLA, HAF-Steer is a latent offline-to-online RL pipeline that leverages flow-matching invertibility and DCT-based dimensionality reduction to restrict RL optimization to a compact noise subspace and train a regularized SAC policy. This avoids updating the large VLA backbone and enables efficient real-world policy refinement. Evaluated on seven real-world humanoid loco-manipulation tasks, HAF surpasses vanilla single-stage VLA baselines and improves whole-body coordination and task performance. Project website: https://grange007.github.io/HAF .
Chinese Translation
类人机器人在以人为中心的环境中作为通用代理具有巨大的潜力,但通用的视觉-语言-动作(VLA)基础模型并不适用于类人机器人的全身运动操控。类人动作的高维性和相互依赖性使得传统的单阶段VLA架构难以有效协调运动、腰部姿态和双臂操控。此外,通过离线行为克隆训练的策略在实际部署中可能仍然是次优的。尽管在线强化学习可以通过与真实世界的交互来优化策略,但直接调整大型VLA骨干网络需要过多的计算,并可能在真实机器人探索过程中带来安全风险。为了解决这些瓶颈,我们提出了HAF(类人适应框架),这是一个由HAF-VLA和HAF-Steer两部分组成的框架,旨在将现成的通用VLA基础模型转移到类人机器人全身运动操控中。HAF-VLA是一个基于预训练流匹配VLA构建的层次化动作流生成器。它将全身动作去噪分为三个顺序阶段,并使用阶段嵌入和跨阶段的KV缓存来保留运动学依赖性,从而避免了一次性生成导致的不协调全身动作。在冻结的HAF-VLA基础上,HAF-Steer是一个潜在的离线到在线强化学习管道,利用流匹配的可逆性和基于DCT的降维技术,将强化学习优化限制在紧凑的噪声子空间中,并训练一个正则化的SAC策略。这避免了更新大型VLA骨干网络,并实现了高效的真实世界策略优化。在七个真实世界的类人运动操控任务上的评估表明,HAF超越了普通的单阶段VLA基线,并改善了全身协调性和任务表现。项目网站:https://grange007.github.io/HAF 。
cs.RO / 86 / 2608.16843

Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation

基础模型驱动的具身智能体安全性:攻击面、攻击、防御与评估
Liu, Jiawei, Guo, Jiacheng, Zhang, Tian, Xu, Yiwei, Wang, Juan, Fan, Jinlin, Xiao, Bowen
Abstract
Foundation models are increasingly used for perception, reasoning, planning, and action generation in embodied agents, creating security risks that can propagate from digital inputs to physical behavior. Existing surveys often organize threats by mechanisms such as jailbreaks, prompt injection, backdoors, poisoning, or adversarial examples, but these categories do not consistently identify where an adversary first enters the embodied control loop. We present a trust-boundary-centric survey of foundation-model-powered embodied-agent security. Using a first-compromised-trust-boundary principle, we separate attack surface from attack mechanism and organize the system into five layers and twelve attack surfaces spanning the model supply chain, user instructions, context and memory, physical semantic environments, multimodal perception, world state, internal reasoning, task planning, action interfaces, middleware, multi-agent communication, and execution control. Based on 58 attack records and 61 defense records collected through August 15, 2026, we analyze representative attacks, cross-layer propagation, defense placement, and evaluation practices. Our quantitative analysis shows that attack research is concentrated on multimodal perception and action interfaces, while defenses are especially concentrated on action-level and runtime protection. Context and long-term memory, middleware and networking, world-state integrity, and multi-agent trust remain comparatively underexplored. We conclude with open challenges in state provenance, compositional defenses, long-horizon attack propagation, physical realizability, Byzantine multi-robot behavior, and unified closed-loop evaluation.
Chinese Translation
基础模型越来越多地用于具身智能体的感知、推理、规划和动作生成,这带来了安全风险,这些风险可以从数字输入传播到物理行为。现有的调查通常通过如越狱、提示注入、后门、投毒或对抗样本等机制来组织威胁,但这些类别并未始终如一地识别出对手首次进入具身控制循环的位置。我们提出了一种以信任边界为中心的基础模型驱动的具身智能体安全性调查。基于首次被攻破的信任边界原则,我们将攻击面与攻击机制分开,并将系统组织为五个层次和十二个攻击面,涵盖模型供应链、用户指令、上下文和记忆、物理语义环境、多模态感知、世界状态、内部推理、任务规划、动作接口、中间件、多智能体通信和执行控制。基于截至2026年8月15日收集的58个攻击记录和61个防御记录,我们分析了代表性的攻击、跨层传播、防御部署和评估实践。我们的定量分析表明,攻击研究集中在多模态感知和动作接口上,而防御则特别集中在动作级别和运行时保护上。上下文和长期记忆、中间件和网络、世界状态完整性以及多智能体信任仍然相对未被充分探索。我们总结了在状态来源、组合防御、长时间攻击传播、物理可实现性、拜占庭多机器人行为和统一闭环评估等方面的开放挑战。
cs.RO / 87 / 2608.16853

FlexWorm: Primitive-augmented Hybrid Contact-motion Planning for Suction-based Multi-segment Deformable Robots

FlexWorm:基于吸力的多段可变形机器人原始增强混合接触运动规划
Tang, Zili, Guo, Tiecheng, Zhang, Qinyue, Guo, Meng
Abstract
Multi-segment suction-based soft robots are promising for inspection and maintenance in confined or fragile environments, but existing approaches still depend heavily on manually designed gaits and environment-specific motion scripts. This work presents a planning framework for serial multi-segment soft robots with deformable body segments and boundary suction pads. The formulation targets full 3D navigation on complex surfaces and explicitly handles discrete adhesion switching and continuous body deformation under geometric, collision, and quasi-static feasibility constraints, while remaining agnostic to the specific actuation realization used to produce segment deformation. Its core, block-wise IK hybrid search (IKHS), performs best-first search over feasible adhesion transitions while solving inverse kinematics only on induced free blocks. On top of IKHS, primitive-augmented hybrid search (PaHS) uses a learned observation--primitive embedding to retrieve short validated motion segments for fast local proposal, with fallback to standard IKHS branching when retrieval fails. In simulation, the framework consistently outperforms controlled baselines in planning success, transition quality, and efficiency across diverse terrains. PaHS matches IKHS in success rate while substantially reducing planning time. Repeated hardware experiments on a pneumatic multi-segment soft robot further demonstrate executability and online recovery under actuation and adhesion uncertainty.
Chinese Translation
基于吸力的多段软机器人在狭小或脆弱环境中的检查和维护中展现出良好的前景,但现有方法仍然严重依赖于手动设计的步态和特定环境的运动脚本。本研究提出了一种针对具有可变形体段和边界吸附垫的串联多段软机器人的规划框架。该框架旨在实现复杂表面上的全3D导航,并明确处理离散粘附切换和在几何、碰撞及准静态可行性约束下的连续体变形,同时对用于产生段变形的具体驱动实现保持无关。其核心是块状逆运动学混合搜索(IKHS),在可行的粘附过渡上执行最佳优先搜索,同时仅在诱导的自由块上求解逆运动学。在IKHS的基础上,原始增强混合搜索(PaHS)利用学习的观察-原始嵌入来检索经过验证的短运动段,以便快速进行局部提议,当检索失败时回退到标准IKHS分支。在仿真中,该框架在规划成功率、过渡质量和效率方面始终优于受控基线,适用于多种地形。PaHS在成功率上与IKHS相匹配,同时显著减少规划时间。对气动多段软机器人的重复硬件实验进一步证明了在驱动和粘附不确定性下的可执行性和在线恢复能力。
cs.RO / 88 / 2608.16885

$\tau_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

$ au_0$-VLA:一种基于世界模型引导的测试时间计算的分层机器人基础模型
Cai, Xiaowei, Cai, Yunuo, Chen, Bingao, Chen, Jingxiao, Chen, Zhi, Feng, Siyuan, Hou, Tengyu, Huang, Jingshun, Jiang, Han, Ju, Runkun, Li, Dong, Li, Mingxiang, Li, Shaowei, Li, Xinchen, Li, Yifan, Liu, Yi, Liu, Zhongyuan, Luo, Jianlan, Miao, Junwen, Ni, Ruiqi, Nie, Buqing, Pan, Mingjie, Ren, Xinlin, Song, Jianheng, Wang, Jiaxu, Wang, Peiqi, Wang, Sen, Wang, Xiaoyan, Wei, Dafeng, Wu, Dongming, Xie, Pengwei, Yang, Pu, Ye, Hangjian, Yue, Xiangyu, Zhang, Jinyu, Zhang, Qinglin, Zhao, Xueyong, Zhou, Pengfei, Zhou, Yue
Abstract
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce $\tau_0$-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.
Chinese Translation
长时间范围的机器人操作要求机器人既能可靠地执行单个技能,又能在扩展任务中连贯地对其进行排序。大多数分层视觉-语言-行动(VLA)模型通过单次前向传播做出每个决策,缺乏将额外计算分配给困难或重要选择的机制。我们提出了$ au_0$-VLA,一种分层机器人基础模型,它通过世界模型引导的测试时间计算将高层子任务生成形式化为一个可扩展的推理问题。在每个推理步骤中,高层策略利用执行记忆生成子任务,并在必要时对备选方案进行搜索,然后再决定其输出。随后,低层策略在多个机器人实现上执行生成的子任务。该策略在40,115小时的异构真实世界数据上进行了多模态共同训练。在领域内和分布转移的设置中,分配额外的测试时间计算显著提高了下一个子任务预测的准确性,这些提升转化为长时间范围机器人操作任务中的更高闭环成功率。
cs.RO / 89 / 2608.16889

Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory

不要放弃接力棒:通过自主子任务探索和过渡感知记忆实现长时间机器人操控
Xu, Bingxin, Shang, Yuzhang, Ferrara, Emilio
Abstract
Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) models increasingly master the individual skills, yet the chain still fails: errors compound beyond the policy's ability to correct, and one subtask silently constrains the next. A promising recipe freezes the VLA and puts an LLM agent in charge: it plans in language, moves in free space with analytic primitives, invokes the VLA only for contact-rich segments, and writes adaptation into language memory. Applied to long horizons, it breaks twice. (1) Competence comes from whole-task exploration at test time, whose cost is multiplicative in stages: if one stage needs T episodes, a K-stage task needs about T^K, and a failure does not reveal which stage caused it. (2) It has no representation of transitions: the VLA primitive carries an exit but no entry condition, so a subtask can succeed in a form its successor cannot use. We present BATON. Against (1), BATON makes the subtask the unit of exploration: each is explored in the cheap short-horizon regime and its solution stored in memory; a long-horizon trajectory is then composed from these solutions rather than discovered whole. Cost becomes additive (T*K) and every failure is attributed to a single stage. Against (2), BATON equips exploration with a transition-aware memory. Within a subtask, a verifier agent governs the invocation transition: the VLA is called only after the wrist view confirms the scene is ready. Across subtasks, a handoff transition restores an entry state disturbed by the predecessor's residue, and a lookahead transition selects the strategy whose outcome the successor can inherit. No parameters are updated. On the long-horizon benchmark RoboMemArena, BATON improves task success by 11.6% and cumulative success by 14.9% over the SoTA.
Chinese Translation
长时间机器人操控将许多富含接触的技能链成一个多阶段任务。视觉-语言-动作(VLA)模型在掌握单个技能方面日益成熟,但这一链条仍然存在问题:错误在超出策略纠正能力的情况下累积,并且一个子任务会默默地限制下一个子任务的执行。一个有前景的方案是冻结VLA,并让一个大型语言模型(LLM)代理负责:它用语言进行规划,以解析原语在自由空间中移动,仅在接触丰富的段落中调用VLA,并将适应性写入语言记忆。应用于长时间任务时,它面临两个主要问题。(1)能力来自于测试时的整体任务探索,其成本在阶段上是乘法的:如果一个阶段需要T个回合,那么一个K阶段的任务大约需要T^K,而一次失败并不能揭示是哪个阶段导致的。(2)它没有过渡的表示:VLA原语携带一个退出条件,但没有进入条件,因此一个子任务可以以一种其后继无法使用的形式成功。我们提出了BATON。针对(1),BATON将子任务作为探索的单位:每个子任务在低成本的短时间范围内进行探索,并将其解决方案存储在记忆中;然后,从这些解决方案中组合出长时间轨迹,而不是整体发现。成本变为加法(T*K),每次失败都归因于单个阶段。针对(2),BATON为探索配备了过渡感知记忆。在一个子任务内,一个验证代理管理调用过渡:只有在腕部视图确认场景准备就绪后才调用VLA。在子任务之间,一个交接过渡恢复了被前任残留物干扰的进入状态,而一个前瞻过渡选择了后继可以继承结果的策略。没有参数被更新。在长时间基准RoboMemArena上,BATON比当前最先进技术(SoTA)提高了11.6%的任务成功率和14.9%的累计成功率。
计算机视觉 (Computer Vision)
219
cs.CV / 1 / 2608.14700

Xemo-Talker: Unlock Emotions Explicitly for Audio-Driven Talking Portrait Synthesis

Xemo-Talker:显式解锁情感以实现音频驱动的对话肖像合成
Yang, Chaolong, Guo, Yinuo, Yao, Kai, Yan, Yuyao, Sun, Jie, Cheng, Guangliang, Wu, Shibin, Dong, Bin, Huang, Kaizhu
Abstract
Precise emotion control in audio-driven talking heads remains a challenge due to the reliance on implicit emotion regulation in existing systems, which often leads to indirect and insufficient control. Additionally, training with explicit emotion-related losses across the entire motion space poses significant difficulties due to the inherent trade-off between accurate lip synchronization and fine-grained emotion control. In this paper, we reveal a key finding: although emotional cues are distributed throughout the motion space, concentrating discriminative supervision on less-principal components achieves a better emotion-lip synchronization balance, as principal components mainly encode high-energy articulation and pose variations. Building on this insight, we propose Xemo-Talker, which first learns a neutral speech-to-motion mapping for stable articulation and lip synchronization, and then introduces a lightweight emotion branch guided by less-principal subspace supervision. To enhance emotion control, we design a Tri-Loss consisting of inter-class separation, intra-class compactness, and less-principal contrastive learning. Given an audio input, a reference image, and an emotion label, Xemo-Talker achieves state-of-the-art emotion classification accuracy while maintaining competitive lip synchronization and high inference efficiency, with performance approaching that measured on real videos.The source code is publicly available at https://github.com/chaolongy/Xemo-Talker.
Chinese Translation
在音频驱动的对话头像中,精确的情感控制仍然是一个挑战,因为现有系统依赖于隐式情感调节,这往往导致间接和不足的控制。此外,由于在准确的唇部同步和细粒度情感控制之间存在固有的权衡,整个运动空间中使用显式情感相关损失进行训练也面临重大困难。在本文中,我们揭示了一个关键发现:尽管情感线索分布在整个运动空间中,但将判别性监督集中在较少的主成分上可以实现更好的情感与唇部同步平衡,因为主成分主要编码高能量的发音和姿态变化。基于这一见解,我们提出了Xemo-Talker,该模型首先学习一个中性语音到运动的映射,以实现稳定的发音和唇部同步,然后引入一个轻量级的情感分支,由较少的主子空间监督引导。为了增强情感控制,我们设计了一个三重损失(Tri-Loss),包括类间分离、类内紧凑性和较少主成分的对比学习。给定音频输入、参考图像和情感标签,Xemo-Talker在保持竞争性的唇部同步和高推理效率的同时,实现了最先进的情感分类准确性,其性能接近于在真实视频上测得的结果。源代码已公开,网址为 https://github.com/chaolongy/Xemo-Talker。
cs.CV / 2 / 2608.14701

Periocular Soft Biometrics: A Survey and Applications to Multimedia Forensics and Disinformation Detection

眼周软生物特征:综述及其在多媒体取证和虚假信息检测中的应用
Alonso-Fernandez, Fernando, Hernandez-Diaz, Kevin, Bigun, Josef
Abstract
Soft-biometric attributes such as gender, age, and ethnicity provide valuable ancillary evidence when full identity recognition is not feasible, supporting applications in forensic investigation, identity verification, surveillance, or detection of synthetic and manipulated media. Among biometric modalities, the periocular region is a robust source of soft-biometric cues, as it often remains visible when other parts of the face are occluded, a frequent condition in forensic evidence and surveillance footage, and can be captured across a wide range of acquisition conditions. In this paper, we provide a survey of demographic attribute estimation from periocular images, covering publicly available datasets, methodological trends from handcrafted descriptors to deep learning architectures, and the state of the art in gender, age, and ethnicity prediction. We discuss use cases relevant to multimedia forensics and disinformation-detection applications, including demographic filtering in surveillance footage, age verification, and the detection of demographic inconsistencies in synthetic data. We also highlight open challenges, including dataset bias, cross-domain generalisation, fairness, ethical aspects, and the lack of forensic-oriented benchmarks.
Chinese Translation
软生物特征属性如性别、年龄和种族在无法进行全面身份识别时提供了有价值的辅助证据,支持在法医调查、身份验证、监控以及合成和操纵媒体检测等应用中的使用。在生物特征模式中,眼周区域是软生物特征线索的一个强大来源,因为在其他面部部位被遮挡时,它通常仍然可见,这在法医证据和监控录像中是常见情况,并且可以在各种采集条件下捕获。在本文中,我们对眼周图像中的人口属性估计进行了综述,涵盖了公开可用的数据集、从手工描述符到深度学习架构的方法趋势,以及性别、年龄和种族预测的最新进展。我们讨论了与多媒体取证和虚假信息检测应用相关的使用案例,包括监控录像中的人口过滤、年龄验证以及合成数据中人口不一致性的检测。我们还强调了开放性挑战,包括数据集偏差、跨领域泛化、公平性、伦理方面以及缺乏面向法医的基准测试。
cs.CV / 3 / 2608.14702

Deep Analog: Open-Set Film Emulation with Reference-Conditioned 3D LUTs

深度模拟:基于参考条件的开放集电影仿真与3D查找表
Mu, Yitong
Abstract
Film emulation reproduces the look of an analog film stock on a new digital photograph. We target its open-set form -- matching any reference film frame from a single example -- with a 3D lookup table (LUT) predicted from that reference. Real-time image enhancement predicts per-image weights over a fixed bank of 3D LUTs and blends them. We show this is a gated mixture of experts and inherits its failure: trained end-to-end against reconstruction, the gate collapses onto a single expert, so a bank of K LUTs delivers the capacity of one. An entropy term, the enhancement-setting analogue of mixture-of-experts load balancing, restores utilization and recovers about 1 dB PSNR. The deeper constraint survives: a fixed LUT basis is closed-set, freezing the achievable looks at training time. We therefore discard the basis and predict a single 3D LUT as a residual from a reference image (StyleLUTNet), trained by self-supervision on procedurally generated color transforms. The conditional design removes the gate and generalizes open-set to unseen film stocks without paired data or retraining. Around this color backbone we build Deep Analog, a film-emulation pipeline that adds histogram-based tone matching and a physics-informed optical renderer -- multi-scale grain and per-channel halation driven by parameters an inverse network regresses from the reference. On 350 self-supervised pairs the color stage reaches 22.05 dB PSNR / 0.925 SSIM and the full pipeline 21.72 dB / 0.923; the color path runs in 5.2 ms at 1080p (192 FPS) and exports a portable .cube LUT for standard editing tools. A second degeneracy in conditional LUT training -- residual-scale collapse -- shares the root cause and yields a general principle: auxiliary regularization must stay subordinate to reconstruction.
Chinese Translation
电影仿真是在新的数字照片上重现模拟胶卷的外观。我们针对其开放集形式——从单个示例中匹配任何参考电影帧——使用从该参考帧预测的3D查找表(LUT)。实时图像增强在固定的3D LUT库上预测每幅图像的权重并进行混合。我们展示了这是一种门控专家混合模型,并继承了其缺陷:在重建任务上进行端到端训练时,门控机制会崩溃到单一专家,因此K个LUT的库只能提供一个的容量。一个熵项,即混合专家负载平衡的增强设置类比,恢复了利用率并提高了约1 dB的PSNR。更深层的约束依然存在:固定的LUT基础是封闭集的,在训练时冻结了可实现的外观。因此,我们舍弃了基础,并预测一个3D LUT作为参考图像的残差(StyleLUTNet),通过自监督在程序生成的颜色变换上进行训练。条件设计去除了门控机制,并将开放集推广到未见过的胶卷类型,而无需配对数据或重新训练。在这个颜色基础上,我们构建了Deep Analog,一个电影仿真管道,增加了基于直方图的色调匹配和一个物理启发的光学渲染器——多尺度颗粒和每通道的光晕由逆网络从参考中回归的参数驱动。在350对自监督样本上,颜色阶段达到了22.05 dB PSNR / 0.925 SSIM,完整管道达到了21.72 dB / 0.923;颜色路径在1080p(192 FPS)下运行时间为5.2毫秒,并导出一个便携的.cube LUT以供标准编辑工具使用。条件LUT训练中的第二个退化现象——残差尺度崩溃——共享根本原因,并得出一个普遍原则:辅助正则化必须服从于重建。
cs.CV / 4 / 2608.14705

On Cross-Validation for Hyperparameter Optimization of Deep Learning Image Classifiers

深度学习图像分类器超参数优化的交叉验证研究
Buturovic, Ljubomir
Abstract
Hyperparameter optimization (HPO) can materially affect the performance of deep learning (DL) image classifiers, but there is little empirical guidance on how to derive the validation signal that drives it, especially for the small sample sizes common in fields such as medical imaging. We compared three HPO protocols in terms of {\em absolute performance-estimation error} (AEE; the absolute difference between the winning configuration's validation AUROC and its test AUROC): fixed holdout (F), reshuffled holdout (R), and 5-fold cross-validation (C). The search space, sampler, training procedure, architecture, and test set were held identical across protocols. We evaluated the protocols on three public datasets spanning two regimes: binary medical imaging (RSNA pneumonia radiographs and binarized HAM10000 skin lesions) and 200-class natural imaging (Tiny ImageNet), across a range of development set sizes $n$ and two backbones (ResNet-18 on all datasets, Vision Transformer (ViT-S/16) on RSNA). On the medical datasets, every point estimate favored cross-validation over both holdout protocols, with reductions in AEE largest at small sample sizes and diminishing as $n$ increased. This pattern remained robust under conservative family-wise adjustment. On Tiny ImageNet, AEE was negligible under all three protocols. Test AUROC was generally similar among protocols. Fixed holdout had lower mean AEE than reshuffled holdout in 11 of 12 medical conditions, although this secondary finding was less uniformly supported. For small-sample medical image classification, we recommend cross-validation-based HPO when computational resources permit because it trades additional computation for a more reliable development-time estimate of subsequent test performance.
Chinese Translation
超参数优化(HPO)对深度学习(DL)图像分类器的性能有显著影响,但关于如何获得驱动其优化的验证信号的实证指导较少,尤其是在医学影像等领域常见的小样本情况下。我们比较了三种HPO协议在绝对性能估计误差(AEE;即获胜配置的验证AUROC与其测试AUROC之间的绝对差异)方面的表现:固定保留(F)、重洗保留(R)和5折交叉验证(C)。在各协议中,搜索空间、采样器、训练过程、架构和测试集保持一致。我们在三个公共数据集上评估了这些协议,涵盖了两个领域:二元医学影像(RSNA肺炎放射图和二值化的HAM10000皮肤病变)和200类自然图像(Tiny ImageNet),并在不同的开发集大小$n$和两个主干网络(所有数据集上的ResNet-18,RSNA上的视觉变换器(ViT-S/16))下进行测试。在医学数据集中,每个点估计都支持交叉验证优于两种保留协议,AEE的减少在小样本情况下最大,随着$n$的增加而减小。该模式在保守的全体家庭调整下依然稳健。在Tiny ImageNet上,所有三种协议下的AEE均可忽略不计。各协议间的测试AUROC通常相似。在12种医学条件中,固定保留的平均AEE在11种情况下低于重洗保留,尽管这一次要发现的支持程度较低。对于小样本医学图像分类,当计算资源允许时,我们建议使用基于交叉验证的HPO,因为它以额外的计算换取了对后续测试性能的更可靠的开发时间估计。
cs.CV / 5 / 2608.14706

Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning

平衡强制:无噪声条件的自适应视频生成
Lillemark, Hansen Jin, Rojas, Alex, Novack, Zachary, Wang, Runqian, Du, Yilun, Ma, Yian, Berg-Kirkpatrick, Taylor, Yu, Rose
Abstract
Standard autoregressive video generation algorithms based on Diffusion and Flow Matching rely on rigid training objectives and static sampling schedules, limiting inference procedures from adapting to the data. We introduce Equilibrium Forcing (EqF), a simplified framework for video denoising generative models without noise level conditioning. EqF pioneers modular training- and inference-time designs for noise-unconditional generation that decouple learning the denoising field from sampling. This flexibility allows for inference-time algorithms that operate in a closed loop by adapting to feedback from the sample, improving video quality and consistency on challenging autoregressive video generation benchmarks. Extensive analysis elucidates exactly how removing the noise level conditioning enables EqF's data-dependent inference properties to surpass the performance of standard noise level-conditional denoising video methods.
Chinese Translation
基于扩散和流匹配的标准自回归视频生成算法依赖于严格的训练目标和静态采样计划,这限制了推理过程适应数据的能力。我们提出了平衡强制(Equilibrium Forcing, EqF),这是一个简化的视频去噪生成模型框架,无需噪声水平条件。EqF 开创了噪声无条件生成的模块化训练和推理设计,将去噪场的学习与采样解耦。这种灵活性使得推理时的算法能够通过适应样本反馈在闭环中运行,从而提高在具有挑战性的自回归视频生成基准上的视频质量和一致性。广泛的分析阐明了去除噪声水平条件如何使 EqF 的数据依赖推理特性超越标准噪声水平条件去噪视频方法的性能。
cs.CV / 6 / 2608.14708

PE-CSNet: An equivariant network architecture with learnable patch-based sparse representation

PE-CSNet:一种具有可学习的基于补丁的稀疏表示的等变网络架构
Li, Kai, Long, Haitao, Zhang, Bo, Zhang, Haiwen, Zhou, Zhi
Abstract
Compressive sensing (CS) enables accurate signal reconstruction from sparse measurements and is widely applied in medical imaging, remote sensing, and image compression. However, designing an effective, task-specific sparse transform and the corresponding optimization procedure for high-quality CS remains challenging. This process typically requires expert domain knowledge and laborious parameter tuning. To address this issue, we present a Patch-based Equivariant deep unrolling architecture, termed PE-CSNet, for accurate CS recovery. While traditional CS methods generally use predefined patch-based transform sparsity, we generalize this idea by incorporating learnable transform sparsity that adapts to the specific CS task through an optimization-driven process. Specifically, we first establish a generalized patch-based CS model, which we solve via a block coordinate descent (BCD) algorithm. The BCD solver is then unrolled into a deep neural network, where all parameters of both the CS model and solver are learned through end-to-end training. To improve data efficiency, we introduce a stochastic equivariant training strategy that exploits the patch-wise structure of the network, enabling PE-CSNet to learn effectively even from limited data. We further provide a simpler, parameter-shared version of PE-CSNet and briefly discuss its convergence as an iterative solver. For practical applications, the network uses stage-specific (non-shared) parameters to enhance its expressive power and thereby improve its performance. On the tasks of CS magnetic resonance imaging (CS-MRI) and CS coded diffraction patterns (CS-CDP), PE-CSNet achieves state-of-the-art accuracy with fast computational speed, outperforming traditional methods and existing deep unrolling methods.
Chinese Translation
压缩感知(Compressive Sensing, CS)能够从稀疏测量中实现准确的信号重建,广泛应用于医学成像、遥感和图像压缩。然而,为高质量的压缩感知设计有效的、特定任务的稀疏变换及相应的优化过程仍然具有挑战性。这个过程通常需要专业领域知识和繁琐的参数调优。为了解决这个问题,我们提出了一种基于补丁的等变深度展开架构,称为PE-CSNet,用于准确的压缩感知恢复。传统的压缩感知方法通常使用预定义的基于补丁的变换稀疏性,而我们通过引入可学习的变换稀疏性来推广这一思想,使其能够通过优化驱动的过程适应特定的压缩感知任务。具体而言,我们首先建立了一个广义的基于补丁的压缩感知模型,并通过块坐标下降(Block Coordinate Descent, BCD)算法进行求解。然后将BCD求解器展开为深度神经网络,其中压缩感知模型和求解器的所有参数通过端到端训练进行学习。为了提高数据效率,我们引入了一种随机等变训练策略,利用网络的基于补丁的结构,使PE-CSNet即使在有限数据下也能有效学习。我们进一步提供了PE-CSNet的一个更简单的参数共享版本,并简要讨论其作为迭代求解器的收敛性。在实际应用中,该网络使用阶段特定(非共享)参数以增强其表达能力,从而提高其性能。在压缩感知磁共振成像(CS-MRI)和压缩感知编码衍射图样(CS-CDP)任务中,PE-CSNet实现了最先进的准确性和快速的计算速度,超越了传统方法和现有的深度展开方法。
cs.CV / 7 / 2608.14710

Path2ST: Hierarchical Cell-Tissue Grounded Cross-Modal Translation for Spatial Transcriptomics

Path2ST:用于空间转录组学的分层细胞-组织基础跨模态翻译
Liu, Ruochen, Lou, Wei
Abstract
Predicting spatial gene expression from hematoxylin and eosin (H\&E)-stained images offers a cost-effective alternative to spatial transcriptomics (ST). However, existing methods treat H\&E images as generic visual inputs and ignore their intrinsic biological hierarchy, where spatially organized cell types collectively form functional tissue microenvironments that govern local gene expression programs. To bridge this gap, we formulate H\&E-to-ST prediction as a cross-modal semantic translation task and propose Path2ST, a hierarchically grounded autoregressive framework featuring three key components: (i) a Hierarchical Cell-Tissue Conditioning mechanism that fuses explicit and implicit cellular features with tissue-level semantic representations to construct hierarchical conditioning signals; (ii) a Scale-Adaptive Autoregressive Generation process over a hierarchical semantic vocabulary, enabling coarse-to-fine, biologically consistent expression synthesis; and (iii) SpectraLoss, a full-spectrum objective that jointly enforces ordinal fidelity, models transcriptional bursts, and aligns semantic structures with cell types. Extensive experiments on three datasets demonstrate state-of-the-art performance, validating that Path2ST generates highly accurate and spatially coherent transcriptomic profiles. The related code is released at https://github.com/RuochenLiu23/Path2ST.
Chinese Translation
从苏木精-伊红(H&E)染色图像中预测空间基因表达为空间转录组学(ST)提供了一种具有成本效益的替代方案。然而,现有方法将H&E图像视为通用视觉输入,忽略了其内在的生物学层次结构,其中空间组织的细胞类型共同形成功能性组织微环境,主导局部基因表达程序。为了解决这一问题,我们将H&E到ST的预测形式化为跨模态语义翻译任务,并提出了Path2ST,一个具有分层基础的自回归框架,包含三个关键组件:(i)分层细胞-组织条件机制,将显式和隐式细胞特征与组织级语义表示融合,以构建分层条件信号;(ii)在分层语义词汇上的尺度自适应自回归生成过程,实现粗到细的生物学一致性表达合成;(iii)SpectraLoss,一个全谱目标,联合强制顺序保真度、建模转录突发,并将语义结构与细胞类型对齐。在三个数据集上的广泛实验表明,Path2ST实现了最先进的性能,验证了其生成高度准确且空间一致的转录组特征。相关代码已发布在 https://github.com/RuochenLiu23/Path2ST。
cs.CV / 8 / 2608.14717

Local Gains and Fixed-Assignment Set Losses in Shared Set Decoders

共享集合解码器中的局部增益与固定分配集合损失
Zhang, Ze, Zhang, Yang
Abstract
A query-relation deletion can improve the edited slot while reducing the utility of the prediction set that contains it. We study this tension in two related ResNet-50 DETR-family checkpoints using recorded, selection-conditional evidence from 710 paired image-relation units per checkpoint. The primary comparison subtracts a matched active control, which deletes the same leader source at a different recorded recipient, from the selected target deletion. It is therefore a composite contrast rather than a same-recipient placebo. The target-minus-control contrast is locally positive and fixed-assignment negative in both checkpoints. The opposite-sign pattern occurs within 302/710 DETR units and 460/710 DINO units. After rematching, the corresponding counts are 285/710 and 433/710. Rematching and native selection absorb enough of the mean loss for DETR intervals to cross zero, whereas DINO intervals remain negative, so persistence across readouts differs by checkpoint. A fixed-map comparison between hard deletion and a mass-preserving edit also differs before rematching. That comparison is conditional on the outcome-blind map and does not establish same-dose transport. Local intervention success therefore does not determine the consequence for a jointly decoded set. The supported conclusion is selection-conditional deletion sensitivity whose persistence depends on the readout and intervention operator. We do not identify an intervention-invariant edge mechanism, detector-level degradation, population prevalence, or the value of a training-time regularizer.
Chinese Translation
查询-关系删除可以改善编辑后的槽位,同时降低包含该槽位的预测集合的效用。我们在两个相关的 ResNet-50 DETR 系列检查点中研究这种张力,使用来自每个检查点的 710 对图像-关系单元的记录选择条件证据。主要比较是从所选目标删除中减去一个匹配的主动控制,该控制在不同的记录接收者处删除相同的领导源。因此,这是一种复合对比,而不是同接收者的安慰剂。目标减去控制的对比在两个检查点中都是局部正向和固定分配负向。相反符号的模式出现在 302/710 的 DETR 单元和 460/710 的 DINO 单元中。重新匹配后,相应的计数为 285/710 和 433/710。重新匹配和本地选择吸收了足够的平均损失,使得 DETR 区间跨越零,而 DINO 区间仍然为负,因此不同检查点之间的读出持久性存在差异。在重新匹配之前,硬删除与质量保持编辑之间的固定映射比较也存在差异。该比较依赖于结果盲映射,并未建立同剂量传输。因此,局部干预的成功并不决定联合解码集合的后果。支持的结论是选择条件删除敏感性,其持久性依赖于读出和干预操作符。我们未能识别出干预不变的边缘机制、探测器级别的退化、群体流行率或训练时间正则化器的价值。
cs.CV / 9 / 2608.14718

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

VideoGAIA:通用人工智能助手在自主视频理解上的基准测试
Zhang, Fan, Yao, Guangming, Wu, Jinyang, Wu, Hao, Lian, Zheng, Geng, Xinyu, Chen, Jingdong, Yuan, Yi, Heng, Pheng-Ann
Abstract
Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesting that conventional single-turn video understanding tasks are becoming increasingly saturated and insufficient for assessing the intelligence of advanced MLLMs. Towards this end, we introduce VideoGAIA, an agentic video understanding benchmark for general artificial intelligence (AI) assistants. Moving beyond one-shot video question answering, VideoGAIA formulates video understanding as a multi-turn, tool-augmented interaction process, where models must iteratively perceive videos, invoke external tools, gather complementary information, and integrate multimodal evidence across turns. VideoGAIA contains 271 model-human co-designed tasks covering diverse and complex real-world scenarios. Each video-question-answer instance is independently verified by three human experts to ensure both correctness and appropriate difficulty. All evaluated MLLMs, including frontier models such as GPT-5.5 and Kimi-K3, achieve less than 60% accuracy on VideoGAIA, highlighting its value as a high-quality and timely benchmark for evaluating next-generation MLLMs. We hope that VideoGAIA will facilitate the transition from conventional video understanding toward agentic video understanding.
Chinese Translation
视频理解是评估多模态大型语言模型(MLLMs)能力的基础任务。然而,现有的领先模型在Video-MME排行榜上已经达到了约90%的准确率,这表明传统的单轮视频理解任务正变得越来越饱和,无法充分评估先进MLLMs的智能。为此,我们引入了VideoGAIA,一个针对通用人工智能(AI)助手的自主视频理解基准。VideoGAIA超越了一次性的视频问答,将视频理解构建为一个多轮、工具增强的交互过程,在这个过程中,模型必须迭代地感知视频、调用外部工具、收集补充信息,并在多个回合中整合多模态证据。VideoGAIA包含271个模型与人类共同设计的任务,涵盖多样且复杂的现实场景。每个视频-问题-答案实例均由三位人类专家独立验证,以确保其正确性和适当的难度。所有评估的MLLMs,包括前沿模型如GPT-5.5和Kimi-K3,在VideoGAIA上的准确率均低于60%,突显了其作为评估下一代MLLMs的高质量及时基准的价值。我们希望VideoGAIA能够促进从传统视频理解向自主视频理解的过渡。
cs.CV / 10 / 2608.14719

DeCo-MIL: Debiased Counterfactual Reasoning for Long-Tailed Whole Slide Image Analysis

DeCo-MIL:用于长尾整体幻灯片图像分析的去偏见反事实推理
Li, Xiaoxiao, Ling, Xitong, Li, Jiawen, Chen, Weiming, Cai, Zhenyang, Wang, Xidong, Guan, Tian, Wang, Benyou, He, Yonghong
Abstract
Multiple instance learning (MIL) is widely used for weakly supervised whole slide image (WSI) analysis. However, under long-tailed distributions, MIL-based WSI analysis faces a nested dual long-tail: an inter-slide class long tail and an intra-slide long tail of instance-level discriminative evidence. The two long tails are coupled: tail classes have few training slides, while their limited diagnostic evidence is concentrated in a few patches and obscured by abundant within-bag redundancy. This coupling biases models toward head classes and degrades rare-class recognition. To address this, we propose DeCo-MIL for long-tailed WSI analysis, which jointly alleviates the nested dual long-tail through frequency-debiased counterfactual reasoning. For the inner long tail, DeCo-MIL clusters patches into tissue-morphology anchors, replaces each anchor with its matched normal prototype to perform a counterfactual intervention, and estimates its counterfactual contribution to the ground-truth class using class-frequency-corrected predictions. These contributions guide redundancy masking to preserve scarce discriminative instances. For the outer long tail, DeCo-MIL constructs anchor-stratified pseudo-bags from redundancy-reduced bags and combines tail-aware oversampling with consistency regularization, increasing effective supervision for tail classes while preserving tissue-morphology composition. Extensive experiments on three long-tailed WSI benchmarks demonstrate that DeCo-MIL achieves state-of-the-art performance in both tail-class recognition and overall classification.
Chinese Translation
多实例学习(MIL)广泛应用于弱监督整体幻灯片图像(WSI)分析。然而,在长尾分布下,基于MIL的WSI分析面临着嵌套的双重长尾:跨幻灯片类别的长尾和实例级判别证据的幻灯片内部长尾。这两个长尾是相互耦合的:长尾类别的训练幻灯片较少,而其有限的诊断证据集中在少数补丁中,并被丰富的内部冗余所掩盖。这种耦合使得模型偏向于头部类别,从而降低了对稀有类别的识别能力。为了解决这一问题,我们提出了DeCo-MIL用于长尾WSI分析,通过频率去偏见的反事实推理共同缓解嵌套的双重长尾。对于内部长尾,DeCo-MIL将补丁聚类为组织形态锚点,用其匹配的正常原型替换每个锚点以进行反事实干预,并使用类别频率校正的预测来估计其对真实类别的反事实贡献。这些贡献指导冗余掩蔽,以保留稀缺的判别实例。对于外部长尾,DeCo-MIL从减少冗余的袋子中构建锚点分层伪袋,并结合关注尾部的过采样与一致性正则化,增加了对尾部类别的有效监督,同时保留了组织形态的组成。在三个长尾WSI基准上的广泛实验表明,DeCo-MIL在尾部类别识别和整体分类方面均达到了最先进的性能。
cs.CV / 11 / 2608.14721

AeroGround: A Comprehensive Benchmark for Aerial-Ground Collaborative Reasoning

AeroGround:空中-地面协作推理的综合基准
Yi, Shenghong, Zhang, Lin, Li, Muzian, Yuan, Jiakang, Zhang, Haoyu, Ye, Peng, Fan, Jiayuan, Qin, Huafeng, Chen, Tao
Abstract
Vision-language models (VLMs) have been widely employed in understanding and reasoning tasks for unmanned aerial vehicles (UAVs). Existing UAV benchmarks primarily focus on aerial-view scenarios. However, whether current VLMs can perform well on understanding and reasoning tasks in aerial-ground collaborative scenarios which are practical in real-world applications like rescue and infrastructure inspection remains underexplored. To address this gap, we introduce AeroGround, a comprehensive benchmark for evaluating VLMs in aerial-ground collaborative reasoning. AeroGround is built upon a simulated aerial-ground dataset containing approximately 29,000 multimodal observation groups from diverse open environments, and provides 2,250 high-quality question-answering instances covering cross-view correspondence, spatial understanding, and reasoning. Experiments on 16 pretrained VLMs, together with two domain-adapted variants, reveal a substantial gap between current models and human performance: the best model achieves an average accuracy of 54.4%, whereas humans reach 93.3%. By systematically revealing the strengths and limitations of existing models in aerial-ground collaborative reasoning, AeroGround provides a foundation for developing more capable aerial-ground collaborative embodied intelligence systems.
Chinese Translation
视觉-语言模型(VLMs)已广泛应用于无人机(UAVs)的理解和推理任务。现有的无人机基准主要集中在空中视角场景。然而,目前的视觉-语言模型在空中-地面协作场景中的理解和推理任务表现如何,尤其是在救援和基础设施检查等实际应用中,仍然未得到充分探索。为了解决这一问题,我们引入了AeroGround,这是一个用于评估视觉-语言模型在空中-地面协作推理中的综合基准。AeroGround建立在一个模拟的空中-地面数据集之上,包含约29,000个来自不同开放环境的多模态观察组,并提供2,250个高质量的问答实例,涵盖跨视角对应、空间理解和推理。对16个预训练的视觉-语言模型及两个领域适应变体的实验显示,当前模型与人类表现之间存在显著差距:最佳模型的平均准确率为54.4%,而人类则达到93.3%。通过系统地揭示现有模型在空中-地面协作推理中的优缺点,AeroGround为开发更强大的空中-地面协作具身智能系统奠定了基础。
cs.CV / 12 / 2608.14722

Braided Vision Transformer for Stroke Detection in Multi-view Retinal Fundus Imaging

用于多视角视网膜眼底成像的编织视觉变换器在中风检测中的应用
Degerli, Aysen, Hilvo, Mika
Abstract
Stroke remains a leading cause of mortality and morbidity worldwide, emphasizing the importance of its accurate and immediate assessment. Retinal fundus imaging has emerged as a promising modality for stroke assessment, as the retina reflects cerebrovascular and neurological risk factors. Contrary to conventional neuroimaging techniques, retinal fundus imaging offers a non-invasive, cost-effective, and portable alternative for rapid screening. This paper explores the feasibility of retinal fundus imaging for stroke and transient ischemic attack (TIA) detection using macula-centric and optic nerve head-centric views captured from both eyes. Our study introduces, to the best of our knowledge, the first vision transformer model for retinal fundus imaging in stroke assessment, offering a novel approach for capturing retinal patterns. Thereby, we propose the Braided Vision Transformer (BViT) model, which extracts representative features from the given multi-view images while simultaneously capturing inter-view relationships across both eyes, enabling a more informative understanding of retinal biomarkers associated with cerebrovascular events. Experiments conducted on our collected Stroke-Data dataset demonstrate that BViT achieves an AUC score of 0.75 for stroke detection, outperforming regular vision transformers.
Chinese Translation
中风仍然是全球死亡和发病率的主要原因,这强调了对其准确和及时评估的重要性。视网膜眼底成像作为一种有前景的中风评估方式,因其能够反映脑血管和神经风险因素而受到关注。与传统的神经影像技术相比,视网膜眼底成像提供了一种非侵入性、成本效益高且便携的快速筛查替代方案。本文探讨了利用从双眼捕获的以黄斑为中心和视神经头为中心的视角进行中风和短暂性脑缺血发作(TIA)检测的视网膜眼底成像的可行性。我们研究首次提出了一种用于中风评估的视网膜眼底成像的视觉变换器模型,提供了一种捕捉视网膜模式的新方法。因此,我们提出了编织视觉变换器(BViT)模型,该模型从给定的多视角图像中提取代表性特征,同时捕捉双眼之间的视角关系,从而更全面地理解与脑血管事件相关的视网膜生物标志物。在我们收集的中风数据集上进行的实验表明,BViT在中风检测中实现了0.75的AUC得分,优于常规视觉变换器。
cs.CV / 13 / 2608.14723

A Vision Transformer for ECG-Based Detection of Left Ventricular Systolic Dysfunction Across Multiple Clinical Sites

基于心电图的左心室收缩功能障碍检测的视觉变换器模型:多临床中心的研究
Ozek, Burcu, Mohan, Aruna, Vorchheimer, David, Weiss, Daniel, Kedar, Eyal, Sobol, Tamar, Zilbershot, Or, Afghah, Fatemeh
Abstract
Reduced left ventricular ejection fraction (LVEF) is frequently asymptomatic and often detected only after advanced heart failure develops. Electrocardiograms are recorded routinely yet underused for this condition, because reduced LVEF has no single diagnostic waveform. We trained an ensemble of vision transformers from scratch to detect reduced LVEF ($\leq$40%) from 12-lead ECGs, analyzing each heartbeat individually, using 10,142 patients across seven sites in three US health systems. In a held-out external cohort of 4,092 patients from three geographically independent US clinical sites at a real-world reduced-LVEF prevalence of 8.72%, the model achieved an AUROC of 0.88 (95% CI 0.86-0.89), sensitivity 81.2%, specificity 81.0%, and negative predictive value 97.8%. Sensitivity remained high across sex, race, ethnicity, and comorbidity subgroups, while specificity was lower in older patients and those with atrial fibrillation or cardiomyopathy. Beat-level attention maps provided interpretability into the model's predictions, showing consistent focus on the QRS complex rather than the P wave. These findings support the potential of routine ECGs as a scalable first-pass triage step to identify patients who should undergo echocardiography for reduced ejection fraction across diverse patient populations.
Chinese Translation
左心室射血分数(LVEF)降低常常无症状,通常在心力衰竭发展到晚期后才被发现。心电图是常规记录的,但在这种情况下使用不足,因为LVEF降低没有单一的诊断波形。我们从零开始训练了一组视觉变换器模型,以从12导联心电图中检测LVEF降低($ ext{LVEF} ext{≤} 40 ext{%}$),逐个分析每个心跳,使用了来自美国三大医疗系统的10,142名患者的数据。在一个来自三个位于地理上独立的美国临床中心的4,092名患者的外部验证队列中,实际LVEF降低的患病率为8.72%,该模型达到了0.88的曲线下面积(AUROC)(95% CI 0.86-0.89),灵敏度为81.2%,特异性为81.0%,阴性预测值为97.8%。在性别、种族、民族和合并症亚组中,灵敏度保持较高,而在老年患者及患有房颤或心肌病的患者中,特异性较低。基于心跳的注意力图提供了模型预测的可解释性,显示出对QRS波群的持续关注,而非P波。这些发现支持常规心电图作为可扩展的初步筛查步骤,以识别应接受超声心动图检查以评估降低射血分数的患者,适用于多样化的患者群体。
cs.CV / 14 / 2608.14724

Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering

保护隐私的数据集策划:基于空间车辆上下文过滤的吉隆坡城市交通视觉-语言检测
Tanzin, Mohammed Abdul Al Arafat, Dziyauddin, Rudzidatul Akmam
Abstract
The rapid advancement of intelligent transportation systems and autonomous driving relies heavily on multi-modal urban traffic datasets. However, curating high-fidelity video imagery in complex tropical urban environments---specifically Kuala Lumpur, Malaysia---presents severe challenges for Personally Identifiable Information (PII) anonymization due to high motorcycle density, dark acrylic license plates, dynamic camera tilt, and extreme tropical glare. We propose an automated anonymization framework tailored for the Kuala Lumpur Road Dataset, captured via a mobile cycling platform at 2 FPS. We document how legacy Haar cascades and YOLOv8 fail under these conditions---generating false positives on background elements while missing rotated or occluded targets. Our architecture resolves this by integrating Grounding DINO---a zero-shot open-set vision-language transformer---with a novel Spatial Vehicle Region of Interest (ROI) Containment Engine. By requiring license plate centroids to reside within validated vehicle boundaries, the pipeline suppresses environmental false positives while automatically obfuscating faces, heads, and license plates. An initial evaluation on 1,266 frames demonstrates a $\sim$95\% success rate, with remaining failures restricted to small, heavily occluded, oblique, or ambiguous targets. Coupled with temporal persistence mechanisms and an automated quality-control auditor, the framework minimizes privacy-related false negatives while preserving scene context for downstream vision tasks. While formal legal compliance depends on broader governance procedures, this publicly available pipeline and demonstration notebook provide an auditable preprocessing stage for privacy-aware dataset curation.
Chinese Translation
智能交通系统和自动驾驶的快速发展在很大程度上依赖于多模态城市交通数据集。然而,在复杂的热带城市环境中,特别是马来西亚吉隆坡,策划高保真度的视频图像面临着严重的个人可识别信息(PII)匿名化挑战,这主要是由于摩托车密度高、黑色亚克力车牌、动态摄像头倾斜和极端热带眩光。我们提出了一种针对吉隆坡道路数据集的自动化匿名化框架,该数据集通过移动骑行平台以每秒2帧的速度捕获。我们记录了传统的Haar级联和YOLOv8在这些条件下的失败——在背景元素上产生误报,同时错过旋转或被遮挡的目标。我们的架构通过将Grounding DINO(一个零-shot开放集视觉-语言变换器)与一种新颖的空间车辆兴趣区域(ROI)包含引擎相结合来解决这一问题。通过要求车牌质心位于经过验证的车辆边界内,该流程抑制了环境误报,同时自动模糊化面孔、头部和车牌。对1,266帧的初步评估表明成功率约为95%,剩余的失败主要限于小型、严重遮挡、倾斜或模糊的目标。结合时间持久性机制和自动化质量控制审计器,该框架在保留场景上下文以支持下游视觉任务的同时,最小化与隐私相关的漏报。尽管正式的法律合规性依赖于更广泛的治理程序,但该公开可用的流程和演示笔记本为隐私意识的数据集策划提供了一个可审计的预处理阶段。
cs.CV / 15 / 2608.14725

Spatial Attention Noise Masking for Causally Sufficient Interpretability

因果充分可解释性的空间注意噪声掩蔽
Formby, Benjamin, Wang, Kuang-Ching, Smith, D Hudson
Abstract
We present a novel causal approach to interpretability for computer vision models that dynamically masks the input image prior to classification. The interpretability of deep learning predictions is critical in high-stakes fields such as medical imaging, security, and autonomous driving. Most interpretability methods are applied passively to already trained models, which typically result in correlational rather than causal explanations. Existing causal interpretability methods are limited to post hoc analysis, weakening the causal claims. Additionally, existing active methods generally lack explanations that explicitly assign responsibility to input features. This work proposes a spatial attention noise masking framework that provides causal explanations about the features sufficient for the prediction. The proposed framework consists of: 1) a UNet-style mask generator, and 2) a Resnet18 encoder and linear classifier that classifies both masked and unmasked versions of an input image. The generated masks are regularized to be sparse and spatially smooth, while masked image embeddings are constrained to remain consistent with embeddings from the corresponding unmasked images. The resulting masks can be interpreted as feature attribution maps that are competitive with related interpretability methods while additionally providing strong causal explanations of model predictions. Quantitative evaluations demonstrate mask faithfulness, near-baseline classification performance across five classification tasks despite substantial masking of image information, and robustness to distribution shifts such as background swapping and natural adversarial examples. Qualitative comparisons further demonstrate mask behavior and competitive interpretability relative to state-of-the-art feature attribution methods.
Chinese Translation
我们提出了一种新颖的因果可解释性方法,适用于计算机视觉模型,该方法在分类之前动态掩蔽输入图像。深度学习预测的可解释性在医疗成像、安全和自动驾驶等高风险领域至关重要。大多数可解释性方法被动应用于已经训练好的模型,通常导致相关而非因果的解释。现有的因果可解释性方法仅限于事后分析,削弱了因果声明。此外,现有的主动方法通常缺乏明确将责任归于输入特征的解释。本研究提出了一种空间注意噪声掩蔽框架,提供关于足以进行预测的特征的因果解释。该框架由以下两部分组成:1)UNet风格的掩蔽生成器,2)Resnet18编码器和线性分类器,用于对输入图像的掩蔽和未掩蔽版本进行分类。生成的掩蔽经过正则化,以保持稀疏性和空间平滑性,同时掩蔽图像的嵌入被约束为与相应未掩蔽图像的嵌入保持一致。生成的掩蔽可以解释为特征归因图,与相关的可解释性方法竞争,同时提供强有力的模型预测因果解释。定量评估表明掩蔽的真实性,在五个分类任务中尽管图像信息大幅掩蔽,分类性能接近基线,并且对背景交换和自然对抗样本等分布变化具有鲁棒性。定性比较进一步展示了掩蔽行为及其相对于最先进的特征归因方法的竞争性可解释性。
cs.CV / 16 / 2608.14727

Low Cost Two-Stage Fabric Defect Detection at the Edge

低成本两阶段边缘织物缺陷检测
Hossen, Rasel, Mistry, Diptajoy, Kamal, Mosaddek Hossain
Abstract
Fabric inspection in the garment industries of low-income economies remains largely manual, and commercial vision systems are priced beyond most small and medium mills. Because defects are sparse under controlled production, a natural response is a cascade: screen every frame with a cheap anomaly detector and invoke a full detector only on suspicious frames. We build such a cascade for four knit-fabric defect classes and deploy it end-to-end on an NVIDIA Jetson Nano with TensorRT FP16. Stage 1 is a compact convolutional autoencoder with decoder attention gates, an edge-weighted reconstruction loss, and feature-level distillation from a frozen YOLOv5n teacher; Stage 2 is YOLOv5n, invoked only on flagged frames. On a 249-image benchmark disjoint from detector training (20 defective, 229 non-defective), Stage 1 at a recall-prioritised threshold flags all 20 defective images (95% CI 0.83-1.00) at a false-positive rate of 49.3% (113/229), reducing false positives by 19.3% relative to a plain autoencoder (p=0.011). The parallel pipeline reaches 13.45 FPS against 9.86 FPS for a sequential YOLO-only loop. Our central finding comes from decomposing that 1.36x: 91% of it is attributable to overlapping JPEG decode with inference rather than to the cascade, which contributes only a 5.1% inference reduction at the measured forwarding rate p = 0.534. We further show that forwarding here is false-positive-limited rather than prevalence-limited - 85% of forwarded frames are false alarms - and quantify the 29-45% inference reduction attainable under tighter calibration. We report this as a caution for cascade speedups measured without controlling the data path, and position the system as AI-assisted triage rather than autonomous acceptance.
Chinese Translation
在低收入经济体的服装行业中,织物检测仍然主要依赖人工,而商业视觉系统的价格超出了大多数中小型纺织厂的承受范围。由于在受控生产下缺陷稀少,自然的应对方式是采用级联方法:使用廉价的异常检测器筛选每一帧,仅在可疑帧上调用完整检测器。我们为四种针织织物缺陷类别构建了这样的级联,并在配备TensorRT FP16的NVIDIA Jetson Nano上进行了端到端的部署。第一阶段是一个紧凑的卷积自编码器,配备解码器注意力门、边缘加权重建损失以及来自冻结的YOLOv5n教师的特征级蒸馏;第二阶段是YOLOv5n,仅在标记的帧上调用。在一个与检测器训练不重叠的249幅图像基准测试中(20幅缺陷图像,229幅非缺陷图像),第一阶段在优先考虑召回的阈值下标记了所有20幅缺陷图像(95%置信区间 0.83-1.00),假阳性率为49.3%(113/229),相较于普通自编码器减少了19.3%的假阳性(p=0.011)。并行管道的帧率达到13.45 FPS,而顺序YOLO仅循环的帧率为9.86 FPS。我们的核心发现来自于对1.36倍加速的分解:91%的加速归因于JPEG解码与推理的重叠,而非级联,后者在测量的转发速率下仅贡献了5.1%的推理减少(p = 0.534)。我们进一步表明,这里的转发是受假阳性限制而非流行度限制——85%的转发帧是假警报——并量化了在更严格的校准下可实现的29-45%的推理减少。我们对此表示警惕,认为在未控制数据路径的情况下测量的级联加速可能存在问题,并将该系统定位为AI辅助的分流,而非自主接受。
cs.CV / 17 / 2608.14729

Do CNNs Internally Represent Real and Fake Images Differently? A Hidden-Layer Analysis

卷积神经网络是否以不同方式内部表示真实与虚假图像?隐藏层分析
Sarma, Moumita Sen, Hitzler, Pascal, Vasserman, Eugene Y.
Abstract
Fake/synthetic images are increasingly prevalent, but it remains unclear whether Convolutional Neural Networks (CNNs) process real and fake images in the same internal manner. This work examines the hypothesis that CNNs represent real and fake images differently, such that fake images induce different hidden-layer activation patterns even when semantic content is preserved. The hypothesis is evaluated in scene recognition settings using trained CNN models. Dense-layer activations are extracted, and neurosymbolic methods assign semantic labels to selected neurons. For each real test image, corresponding fake images are generated with similar semantic content using object-label-guided text-to-image and image-to-image generation based on Stable Diffusion variants. Paired real-fake activation patterns are then compared statistically. Additional experiments with another dataset, CNN architecture, generative model, and JPEG/blur degradation analysis assess robustness. Results suggest that fake images evoke different hidden-neuron activations, and these differences are not explained only by simple image degradation. Overall, the findings indicate that real and fake images differ in CNN hidden-layer activation behavior at least in some settings, which opens the door for follow-up work on making use of this different behavior to improve fake image detection.
Chinese Translation
虚假/合成图像日益普遍,但尚不清楚卷积神经网络(CNN)是否以相同的内部方式处理真实和虚假图像。本研究检验了一个假设,即CNN以不同的方式表示真实和虚假图像,从而使得虚假图像即使在语义内容保持不变的情况下也会引发不同的隐藏层激活模式。该假设在场景识别设置中使用训练好的CNN模型进行评估。提取密集层激活,并采用神经符号方法为选定的神经元分配语义标签。对于每个真实测试图像,使用基于Stable Diffusion变体的物体标签引导的文本到图像和图像到图像生成方法生成相应的虚假图像,确保其语义内容相似。然后对配对的真实-虚假激活模式进行统计比较。通过使用另一个数据集、CNN架构、生成模型以及JPEG/模糊降解分析进行额外实验,以评估结果的稳健性。结果表明,虚假图像引发不同的隐藏神经元激活,这些差异并不仅仅可以通过简单的图像降解来解释。总体而言,研究结果表明,在某些设置中,真实与虚假图像在CNN隐藏层激活行为上存在差异,这为后续研究利用这种不同的行为来改善虚假图像检测开辟了新的方向。
cs.CV / 18 / 2608.14730

IP Protection in the Era of Visual Generative AI: A Survey

视觉生成性人工智能时代的知识产权保护:一项调查
Shi, Zhuan, Liu, Shunchang, Farashah, Alireza Dehghanpour, Yang, Qian, Yu, Han, Yang, Cao, Chen, Chaochao, Yan, Yuping, Jin, Yaochu, Farnadi, Golnoosh, Lyu, Lingjuan
Abstract
The rapid evolution of visual generative AI has introduced a wide range of intellectual property risks, spanning the unauthorized learning, reproduction, extraction, misuse, and redistribution of protected data and model assets. To address these risks, a growing body of technical defenses has been proposed. However, existing surveys typically organize this literature by lifecycle stage or technical mechanism, which can obscure the protective intent of different methods. This survey presents a two-dimensional taxonomy for IP protection in visual generative models. The primary axis is a Control Logic View, which classifies methods into Information Exposure Control, Generative Behavior Constraint, and Attribution & Accountability according to the risk variable they regulate. The secondary axis distinguishes Data IP from Model IP as cross-cutting asset dimensions. Under this framework, we systematically review protection methods, align evaluation protocols with protection objectives, and discuss open challenges including proactive model-level safeguards, standardized evaluation, robustness against adaptive attacks, and explainable evidence. This survey aims to offer a principled, systematic, and easy-to-follow overview for both new and experienced researchers in visual generative AI IP protection.
Chinese Translation
视觉生成性人工智能的快速发展引入了广泛的知识产权风险,涵盖了对受保护数据和模型资产的未经授权学习、复制、提取、滥用和再分配。为应对这些风险,提出了一系列技术防御措施。然而,现有的调查通常按生命周期阶段或技术机制组织文献,这可能会掩盖不同方法的保护意图。本调查提出了一种针对视觉生成模型知识产权保护的二维分类法。主要轴线是控制逻辑视角,根据其调节的风险变量将方法分类为信息曝光控制、生成行为约束和归属与问责。次要轴线则将数据知识产权与模型知识产权区分为交叉资产维度。在这一框架下,我们系统地回顾了保护方法,将评估协议与保护目标对齐,并讨论了包括主动模型级保护、标准化评估、对自适应攻击的鲁棒性和可解释证据等开放挑战。本调查旨在为视觉生成性人工智能知识产权保护领域的新手和经验丰富的研究人员提供一个原则性、系统性和易于理解的概述。
cs.CV / 19 / 2608.14731

Emergence of Transfer Learning towards Specific Identification of Alzheimer's Disease A Prospective Approach

转移学习在阿尔茨海默病特定识别中的出现:一种前瞻性方法
Podder, Soumik, Haldar, Chandramouli
Abstract
Worldwide, millions of senior citizens are suffering from Alzheimer disease abbreviated as AD, a well- versed form of dementia. AD is featured by amnesia, intellectual disability, and difficulty with consciousness. DL and ML models are undoubtedly explored to identify AD related patterns on large dimensional neuroimaging data but they need global optimization and are suffering from overfitting issue that might yield dissatisfactory result in testing data set. DL overcomes the issue by convolution of input image with kernel but any sudden change in the MRI image or human manipulation, limited pre- processing of the images can mislead CNN in achieving highly accurate detection. Transfer Learning (TL) has proved itself in AD diagnosis by utilizing pre-trained models on large data sets to guide novice model in a new neuroimaging dataset. This review provides an inclusive glimpse of TL implication in classification, identification including the conversion of AD. Keeping in view, we have assessed the strengths and limitations of TL in improvising diagnostic accuracy even with limited data. The uniqueness of the present review is the incorporation of explainable AI in TL based AD diagnosis system. Finally, it can be claimed that the review will guide the new re-searchers in the area of TL induced neurodegenerative disease detection.
Chinese Translation
全球范围内,数百万老年人正遭受阿尔茨海默病(Alzheimer disease,简称AD)的困扰,这是一种广为人知的痴呆症形式。AD的特征包括健忘、智力障碍和意识困难。深度学习(Deep Learning,DL)和机器学习(Machine Learning,ML)模型无疑被探索用于识别大维度神经影像数据中的AD相关模式,但它们需要全局优化,并且面临过拟合问题,这可能导致在测试数据集上产生不令人满意的结果。DL通过与卷积核对输入图像进行卷积来克服这一问题,但MRI图像的任何突然变化或人为干预,以及图像的有限预处理,可能会误导卷积神经网络(Convolutional Neural Network,CNN)实现高精度检测。转移学习(Transfer Learning,TL)通过利用在大数据集上预训练的模型来指导新模型在新的神经影像数据集中的应用,已在AD诊断中证明了其有效性。本综述提供了TL在分类、识别及AD转化中的应用的全面概述。考虑到这一点,我们评估了TL在改善诊断准确性方面的优缺点,即使在数据有限的情况下。本综述的独特之处在于将可解释人工智能(explainable AI)纳入基于TL的AD诊断系统。最后,可以说本综述将指导新研究者在TL诱导的神经退行性疾病检测领域的研究。
cs.CV / 20 / 2608.14740

From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation

从密集预测到视觉编辑:统一图像和视频创作的结构化监督
Rao, Zhefan, Zou, Bin, Che, Haoxuan, He, Xuanhua, Choi, Chong Hou, Li, Yanheng, Liu, Rui, Chen, Qifeng
Abstract
Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-only conditioning and creation-only training do not explicitly supervise the local structure needed for precise, temporally consistent editing. We therefore formulate depth and surface-normal prediction as image-form denoising targets, using these dense tasks as structured visual supervision within the same creation interface. Our framework decouples semantic interpretation from spatially aligned visual injection while sharing one multimodal diffusion transformer (MMDiT) backbone across all tasks. Mutual Context Attention (MCA), a paired-video data-construction procedure, and a progressive training curriculum then connect the learned structural cues to temporally localized editing and reference-conditioned creation. A single checkpoint obtains the highest overall score in the reported comparison of unified systems (4.15); adding dense supervision improves OpenVE Overall from 3.98 to 4.06 and Local Add from 3.92 to 4.18. These results support a deliberately bounded conclusion: perception-oriented dense supervision transfers useful structural knowledge to downstream creation, especially editing locality and preservation; we do not claim superiority as a standalone dense predictor.
Chinese Translation
统一的图像和视频创作要求模型在遵循多样化指令的同时,保持来自视觉上下文的身份、几何形状和时间结构。然而,仅依赖语义条件和仅进行创作训练并未明确监督精确且时间一致的编辑所需的局部结构。因此,我们将深度和表面法线预测公式化为图像形式去噪的目标,利用这些密集任务作为同一创作界面内的结构化视觉监督。我们的框架将语义解释与空间对齐的视觉注入解耦,同时在所有任务中共享一个多模态扩散变换器(MMDiT)主干。互文上下文注意力(MCA)、配对视频数据构建程序和渐进式训练课程将学习到的结构线索与时间局部编辑和参考条件创作连接起来。一个单一的检查点在统一系统的比较中获得了最高的总体评分(4.15);增加密集监督将OpenVE总体评分从3.98提高到4.06,将局部添加评分从3.92提高到4.18。这些结果支持一个经过深思熟虑的有限结论:以感知为导向的密集监督将有用的结构知识转移到下游创作中,尤其是编辑的局部性和保留;我们并不声称作为独立的密集预测器具有优越性。
cs.CV / 21 / 2608.14741

PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal Models

PolyComp:基于多立方体的多模态模型组合三维空间推理基准
Patel, Siddharth
Abstract
We introduce PolyComp, a procedurally generated and verified benchmark that stresses visual recognition and compositional spatial reasoning. In each problem, a model must identify which of four options shows a pair of polycube components that can be combined to form a target solid. The benchmark contains 120 problems across four geometry families, and each problem has three different presentation formats using either a single image or multiple images. The random guessing baseline is 25%. Across the three presentations (360 presented problems per model), GPT-5.6 Sol with max effort attains 50.0% accuracy (95% problem-cluster CI 43.3-56.7%) at a mean cost of \$0.951 per presented problem, Claude Fable 5 with max effort attains 39.4% (33.1-46.1%) at \$0.701, and Gemini 3.1 Pro Preview with thinking level high attains 27.5% (22.8-32.5%), near the 25% random guessing baseline, at \$0.350. The observed accuracy spread across geometry families is larger than across presentation formats. We present a problem development and evaluation protocol, cost and token accounting, and release the 120 problems.
Chinese Translation
我们介绍了PolyComp,这是一个程序生成并经过验证的基准,强调视觉识别和组合空间推理。在每个问题中,模型必须识别四个选项中哪一个展示了一对可以组合成目标固体的多立方体组件。该基准包含120个问题,涵盖四个几何家族,每个问题有三种不同的呈现格式,使用单幅图像或多幅图像。随机猜测的基线为25%。在三种呈现方式中(每个模型呈现360个问题),GPT-5.6 Sol在最大努力下达到了50.0%的准确率(95%问题集群置信区间43.3-56.7%),每个呈现问题的平均成本为0.951美元;Claude Fable 5在最大努力下达到了39.4%(33.1-46.1%),成本为0.701美元;而Gemini 3.1 Pro Preview在高思考水平下达到了27.5%(22.8-32.5%),接近25%的随机猜测基线,成本为0.350美元。观察到的几何家族间的准确率差异大于不同呈现格式间的差异。我们提出了一个问题开发和评估协议、成本和标记核算,并发布了这120个问题。
cs.CV / 22 / 2608.14766

Beyond Boundary Noise: Aggregated Aleatoric Uncertainty Fails to Capture Presence Ambiguity in 3D Lung Nodule Segmentation

超越边界噪声:聚合的随机不确定性未能捕捉三维肺结节分割中的存在模糊性
Baur, Simon, Schernich, Arne, Böke, Ekin, Samek, Wojciech, Ma, Jackie
Abstract
Uncertainty estimation is critical for the safe clinical deployment of deep learning in medical image segmentation, with aleatoric uncertainty theoretically designed to capture irreducible data ambiguity. However, whether entropy-based measures reflect clinically meaningful ambiguity, i.e. case-level disagreement about whether a pathology is present at all, remains poorly understood. Contrary to most prior work, which focused on pixel-wise boundary disagreement, we systematically evaluate how well aleatoric uncertainty captures presence ambiguity. Our evaluation spans 3D lung nodule segmentation across four architectures with Monte Carlo dropout and deep ensembles, on LIDC-IDRI and an external validation cohort (LNDb). We find that entropy-based uncertainty maps align with boundary noise and minor drawing variation but carry insufficient discriminative signal for presence ambiguity. In contrast, a lightweight supervised ambiguity head trained on frozen segmentation features substantially outperforms all entropy-aggregation-based baselines across architectures, metrics, and both cohorts, and matches or exceeds methods that explicitly model ambiguity under disagreement supervision (Probabilistic U-Net, Annotator-Confusion 3D-UNet). A qualitative feature-space analysis shows that presence ambiguity is already encoded in the frozen encoder features of pixel-wise-trained networks, only to be discarded by the segmentation output and its entropy aggregation. Our findings expose a fundamental mismatch between the theoretical promise of aleatoric uncertainty and its practical behavior, and suggest that practitioners should not rely on entropy-based uncertainty as a proxy for clinical ambiguity in safety-critical applications.
Chinese Translation
不确定性估计对于深度学习在医学图像分割中的安全临床应用至关重要,其中随机不确定性理论上旨在捕捉不可减少的数据模糊性。然而,基于熵的度量是否反映临床上有意义的模糊性,即关于病理是否存在的个案级别分歧,仍然不够明确。与大多数以往研究集中于像素级边界分歧的做法相反,我们系统地评估了随机不确定性在捕捉存在模糊性方面的有效性。我们的评估涵盖了在LIDC-IDRI和一个外部验证队列(LNDb)上,使用蒙特卡洛丢弃(Monte Carlo dropout)和深度集成(deep ensembles)的四种架构进行的三维肺结节分割。我们发现,基于熵的不确定性图与边界噪声和轻微的绘制变化相一致,但对于存在模糊性却缺乏足够的区分信号。相比之下,在冻结的分割特征上训练的轻量级监督模糊性头在所有架构、指标和两个队列中显著优于所有基于熵聚合的基线,并且与明确建模模糊性的分歧监督方法(如概率U-Net(Probabilistic U-Net)、标注者混淆3D-UNet(Annotator-Confusion 3D-UNet))的效果相当或更好。定性特征空间分析表明,存在模糊性已经在像素级训练网络的冻结编码器特征中编码,但在分割输出及其熵聚合中被丢弃。我们的发现揭示了随机不确定性的理论承诺与其实际表现之间的根本不匹配,并建议从业者在安全关键应用中不应依赖基于熵的不确定性作为临床模糊性的代理。
cs.CV / 23 / 2608.14767

NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving

NARRATE:一个用于自动驾驶中以人为本解释的多模态真实世界澳大利亚驾驶数据集
Zadeh, Ashkan Yousefi, Zhu, Zishuo, Li, Xiaomeng, Rakotonirainy, Andry, Glaser, Sebastien, Schroeter, Ronald, Delhomme, Patricia, Mehraban, Zahra
Abstract
Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Existing language-annotated driving datasets are mostly observer-written, post-hoc, simulation-based, or generated from sensor inputs, rather than elicited from the driver performing the action. We introduce NARRATE, a multimodal real-world Australian driving dataset comprising 2,050 annotated events from 35 experienced drivers and driving instructors on public roads. Each event is grounded in synchronised visual, localisation, motion, and LiDAR streams and paired with in-vehicle and/or post-drive free-text explanations. NARRATE provides action labels, scenario-context labels spanning six high-level and 32 fine-grained categories, and span-level Situational Awareness (SA) annotations over driver explanations for Perception, Comprehension and Projection. Four benchmark tasks (SA, scenario-context, driver-action classification, and explanation generation) show that this structure is learnable from driver language, while fine-grained context recognition and explanation generation remain challenging. NARRATE paves a path towards more human-centred and domain-aware explanation models for automated driving.
Chinese Translation
自动驾驶车辆必须以乘客能够理解、监控和信任的方式解释其决策。现有的语言注释驾驶数据集大多是由观察者撰写的事后分析、基于模拟的,或是从传感器输入生成的,而不是从执行动作的驾驶员那里获取的。我们介绍了NARRATE,这是一个包含来自35名经验丰富的驾驶员和驾驶教练在公共道路上记录的2050个注释事件的多模态真实世界澳大利亚驾驶数据集。每个事件都基于同步的视觉、定位、运动和激光雷达(LiDAR)流,并配有车内和/或驾驶后自由文本解释。NARRATE提供了动作标签、涵盖六个高层次和32个细粒度类别的场景上下文标签,以及针对驾驶员解释的感知、理解和预测的跨段情境意识(SA)注释。四个基准任务(SA、场景上下文、驾驶员动作分类和解释生成)表明,这种结构可以从驾驶员语言中学习,而细粒度的上下文识别和解释生成仍然具有挑战性。NARRATE为自动驾驶的更以人为本和领域意识的解释模型铺平了道路。
cs.CV / 24 / 2608.14768

Uncertainty Identifies Difficult Samples Across Methods: A Multi-Task Study on a Heterogeneous Skin Lesion Dataset

不确定性识别跨方法的困难样本:对异质皮肤病变数据集的多任务研究
Koole, Leon, Guo, Jiapan, Valdenegro-Toro, Matias
Abstract
Skin lesion classifiers can be confidently wrong on the cases that matter most, so knowing when a prediction should not be trusted is clinically as useful as the prediction. We study uncertainty quantification on a dataset pooled from many ISIC sources, with a shared backbone and two jointly learned heads: a binary malignant versus non-malignant head and a five-class diagnostic head. Five UQ methods (MC Dropout, DropConnect, Flipout, Deep Ensembles, DUQ) are compared on accuracy, calibration, uncertainty decomposition, and risk-coverage. Difficulty is largely method-agnostic: even methods with narrow entropy distributions rank the same samples as hard (per-sample entropy correlations of $0.54$ to $0.91$). The choice of method matters more for calibration and uncertainty decomposition, where Deep Ensembles is the clear winner, than for finding difficult cases. The ranking is also good enough that deferring the most uncertain cases removes a disproportionate share of errors, supporting uncertainty-based selective referral, evaluated here in-distribution only.
Chinese Translation
皮肤病变分类器在最重要的案例中可能会自信地出错,因此了解何时不应信任预测在临床上与预测本身同样有用。我们研究了从多个ISIC来源汇总的数据集上的不确定性量化,采用共享的骨干网络和两个联合学习的头部:一个二元恶性与非恶性头部和一个五类诊断头部。比较了五种不确定性量化方法(MC Dropout、DropConnect、Flipout、Deep Ensembles、DUQ)在准确性、校准、不确定性分解和风险覆盖方面的表现。困难样本在很大程度上与方法无关:即使是具有狭窄熵分布的方法,也将相同的样本标记为困难(每个样本的熵相关性为0.54到0.91)。方法的选择在校准和不确定性分解中更为重要,其中Deep Ensembles明显表现最佳,而在寻找困难案例时则相对不那么重要。排名也足够好,以至于推迟处理最不确定的案例可以消除不成比例的错误,支持基于不确定性的选择性转诊,这里仅在分布内进行了评估。
cs.CV / 25 / 2608.14770

Artificial Intelligence as a Tool for Combating Child Labour: A Real-Time Edge Vision Pipeline for Child Detection and Age Estimation

人工智能作为打击儿童劳动的工具:用于儿童检测和年龄估计的实时边缘视觉管道
Nowak, Mark
Abstract
An estimated 138 million children remain in child labour worldwide, and the monitoring systems used by affected sectors, built on periodic household visits and interviews, systematically under-detect them. We present a real-time computer-vision pipeline, built and operated solely as a research prototype, that studies the feasibility of giving Child Labour Monitoring and Remediation Systems (CLMRS) a continuous, presence-based evidence channel. The pipeline combines a multi-task person and face detector (YOLO26x backbone in the CerberusDet framework), cascaded age estimation pairing MiVOLO v2 with a child-specialist model for ages 0-12, ByteTrack tracking, ArcFace and DINOv2 re-identification, and track-level fusion producing reviewable per-person records. The detector raises person [email protected] from 0.390 to 0.683 over the previous-generation baseline; the child specialist reaches 1.944 years MAE on children-only validation, where widely used open-source stacks err by 18-23 years. FP8 TensorRT compilation yields a 1.77x speedup at +0.002 years MAE, bringing the pipeline above twice real-time on embedded hardware. On 26.8 hours of proxy video the system finds 634 unique child candidates versus 285 for its predecessor. We further report a seventeen-day unattended field pilot on a farm in Zimbabwe (38.7 million frames, six cameras) evaluated against a daily attendance register: software tuning improved detection yield 36-fold, and identity consolidation under a simultaneity veto cut over-reporting from 9.1x to 1.8-3.9x with zero proven-false merges. We document training and quantisation failures alongside successes, and the data-protection and human-in-the-loop safeguards such a system requires.
Chinese Translation
全球约有1.38亿儿童仍在从事儿童劳动,而受影响行业使用的监测系统,基于定期的家庭访问和访谈,系统性地低估了儿童劳动的情况。我们提出了一种实时计算机视觉管道,作为研究原型构建和操作,研究为儿童劳动监测和救助系统(CLMRS)提供一个持续的、基于存在的证据通道的可行性。该管道结合了多任务的人体和面部检测器(在CerberusDet框架中的YOLO26x主干),级联的年龄估计,将MiVOLO v2与针对0-12岁儿童的专家模型配对,使用ByteTrack进行跟踪,ArcFace和DINOv2进行再识别,并通过轨道级融合生成可审查的每人记录。该检测器将人类[email protected]从0.390提高到0.683,相较于上一代基线;儿童专家在仅针对儿童的验证中达到了1.944年的平均绝对误差(MAE),而广泛使用的开源堆栈的误差为18-23年。FP8 TensorRT编译实现了1.77倍的加速,MAE增加0.002年,使得该管道在嵌入式硬件上超过实时两倍。在26.8小时的代理视频中,该系统发现634个独特的儿童候选者,而其前身仅为285个。我们进一步报告了在津巴布韦一农场进行的为期十七天的无人值守现场试点(3870万帧,六个摄像头),与每日出勤登记进行评估:软件调优使检测产出提高了36倍,而在同时性否决下的身份整合将过度报告从9.1倍降低到1.8-3.9倍,且没有证明的错误合并。我们记录了训练和量化的失败与成功,以及该系统所需的数据保护和人机协作的安全措施。
cs.CV / 26 / 2608.14778

AMPLIFAI: A Multiphase CT Dataset for Benchmarking Clinical Reasoning in LI-RADS Assessment of Liver Lesions

AMPLIFAI:用于评估肝脏病变LI-RADS的临床推理基准的多相CT数据集
Kulkarni, Pranav, Shah, Nikhil, Suryavanshi, Amritansh, Delfino, Jana, Tonascia, James, Wong-You-Cheong, Jade, Lane, Barton, Chirico, Joseph, Hirsch, Jeffrey D., Li, Ang, Huang, Heng, Doo, Florence X.
Abstract
Hepatocellular carcinoma (HCC) is the third leading cause of cancer-related mortality worldwide, with early detection improving survival from <20\% to >70\%. The standardized LI-RADS criteria establish a biopsy-free, fully imaging-based framework that can serve as a foundation for automating HCC diagnosis with artificial intelligence (AI). However, the lack of large, publicly available datasets with high-quality labels has limited the development of AI models for LI-RADS characterization. We introduce the \textbf{AMPLIFAI} dataset, the first public dataset of multiphase abdominal CT scans annotated with LI-RADS categories and segmented for three major LI-RADS features: arterial phase hyperenhancement, washout, and enhancing capsule. Following the \emph{Datasheets for Datasets} format, this paper details the dataset's composition, curation process, and annotation pipeline to facilitate transparent, reproducible research.
Chinese Translation
肝细胞癌(HCC)是全球癌症相关死亡的第三大原因,早期检测可将生存率从低于20\%提高至超过70\%。标准化的LI-RADS标准建立了一种无活检、完全基于影像的框架,可以作为利用人工智能(AI)自动化HCC诊断的基础。然而,缺乏大型、公开可用的高质量标签数据集限制了LI-RADS特征化AI模型的发展。我们介绍了 extbf{AMPLIFAI}数据集,这是第一个公开的多相腹部CT扫描数据集,标注了LI-RADS类别,并对三大LI-RADS特征进行了分割:动脉期增强、洗脱和增强包膜。本文遵循 extit{Datasheets for Datasets}格式,详细描述了数据集的组成、整理过程和标注流程,以促进透明、可重复的研究。
cs.CV / 27 / 2608.14783

MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling

MegaParts:通过高效令牌自回归建模将部件感知的3D对象生成扩展至300个部件
Liao, Manwen, Lian, Xinyu, Mao, Jian, Chen, Kaixu, Luo, Li, Yan, Jinghao, Gan, Wanshui, Yu, Qiao, Zhang, Weitian, Shen, Chunhua, Chen, Guang, Dai, Bo, Xu, Xudong, Lyu, Zhaoyang
Abstract
Part-aware 3D object generation is essential for graphics applications such as controllable modeling, editing, and articulation, where objects are represented as coherent assemblies of semantic parts. However, existing part-aware generation methods, do not scale well to highly complex objects. As the number of parts increases, generating detailed geometry becomes prohibitively expensive in token length and memory. We introduce MegaParts, a scalable autoregressive 3D generation framework to address this challenge by combining structured sequence modeling with a token-efficient vector-quantized shape tokenizer. Our tokenizer learns discrete latent representations for part-level geometry by minimizing token usage subject to high-fidelity reconstruction, enabling adaptive-length tokenization based on geometric complexity. On top of this compact representation, we train a large language model to generate object bounding boxes, part bounding boxes, and part shape tokens within a unified structured sequence. Combined with efficient long-context training strategy, our token-efficient formulation scales to objects with up to 300 parts and sequence lengths up to 256k tokens. This substantially extends the scale of part-aware 3D generation while preserving compositional structure and enabling fine-grained part-level control. Our method achieves higher mesh quality than baseline autoregressive and diffusion models, showing that compressed discrete part tokens improve not only scalability but also the achievable fidelity of generated geometry. These results suggest that LLM native token-efficient autoregressive modeling is a compelling alternative to diffusion for large-scale part-aware 3D generation. The project page is available at https://expmaster.github.io/megaparts_webpage.
Chinese Translation
部件感知的3D对象生成对于可控建模、编辑和关节运动等图形应用至关重要,其中对象被表示为语义部件的连贯组合。然而,现有的部件感知生成方法在处理高度复杂的对象时扩展性较差。随着部件数量的增加,生成详细几何形状在令牌长度和内存方面变得极为昂贵。我们提出了MegaParts,一个可扩展的自回归3D生成框架,以通过将结构化序列建模与高效令牌向量量化形状标记器相结合来应对这一挑战。我们的标记器通过最小化令牌使用量来学习部件级几何的离散潜在表示,同时保持高保真重建,从而实现基于几何复杂性的自适应长度标记化。在此紧凑表示的基础上,我们训练了一个大型语言模型,以在统一的结构化序列中生成对象边界框、部件边界框和部件形状令牌。结合高效的长上下文训练策略,我们的高效令牌公式可扩展至最多300个部件和长度达到256k令牌的对象。这大大扩展了部件感知3D生成的规模,同时保留了组合结构并实现了细粒度的部件级控制。我们的方法在网格质量上优于基线自回归模型和扩散模型,表明压缩的离散部件令牌不仅提高了可扩展性,还提升了生成几何的可实现保真度。这些结果表明,原生的高效令牌自回归建模是大规模部件感知3D生成的一个引人注目的替代方案。项目页面可访问 https://expmaster.github.io/megaparts_webpage。
cs.CV / 28 / 2608.14790

Qwen-Video-Edit: Instruction-Based Video Editing by Repurposing an Image Editing Model

Qwen-视频编辑:基于指令的视频编辑通过重用图像编辑模型
Bai, Yunpeng, Gandelsman, Yossi, Gharbi, Michaël, Huang, Qixing
Abstract
Instruction-based video editing is commonly built on video-pretrained generative backbones: a video diffusion transformer is adapted, at considerable cost, to condition on a source video and an editing instruction. In this report we explore a different route and show that a strong instruction-based image editing model can edit videos by operating directly on video-VAE latents. Starting from Qwen-Image-Edit, we arrange the latent frames of a Wan~2.1 video VAE as tiles of one large virtual image, reuse the editor's image positional encoding for every tile, and bridge the two latent spaces with a pair of lightweight input/output projections warm-started from the editor's own patchify and unpatchify layers, so that at initialization a (static) video is embedded exactly as an image the model already understands. The whole system is then fine-tuned on the public Ditto-1M editing triplets, and a few denoising steps of Wan~2.2 serve as an optional temporal enhancer. We motivate the design with a chain of zero-training observations: the stock image editor already edits a video presented as a contact sheet; it is indifferent to whether the sheet's tokens come from one joint encode or from per-frame encodes stitched in latent space; and it even edits genuine video latents zero-shot to a clearly recognizable degree, leaving fine-tuning only a fidelity gap to close. Our results suggest that, despite the large investment in training video latent spaces, per-frame video latents remain close enough to the image domain that mature image editing priors transfer with minimal adaptation. Project Page: https://yunpeng1998.github.io/Qwen-Video-Edit-Page; Code: https://github.com/yunpeng1998/Qwen-Video-Edit; Model: https://huggingface.co/yunpeng1998/Qwen-Video-Edit.
Chinese Translation
基于指令的视频编辑通常建立在经过视频预训练的生成骨干网络之上:视频扩散变换器被适配,以高昂的成本对源视频和编辑指令进行条件化。在本报告中,我们探索了一条不同的路径,展示了一个强大的基于指令的图像编辑模型可以通过直接操作视频变分自编码器(VAE)潜变量来编辑视频。从Qwen-图像编辑开始,我们将Wan~2.1视频VAE的潜在帧排列为一个大型虚拟图像的瓷砖,重用编辑器的图像位置编码用于每个瓷砖,并通过一对轻量级的输入/输出投影将两个潜在空间连接起来,这些投影从编辑器自身的分块和解块层进行热启动,以便在初始化时,一个(静态)视频被准确地嵌入为模型已经理解的图像。整个系统随后在公共的Ditto-1M编辑三元组上进行微调,Wan~2.2的几步去噪作为可选的时间增强器。我们通过一系列零训练观察来激励设计:现成的图像编辑器已经能够编辑以接触表形式呈现的视频;它对接触表的标记是来自一个联合编码还是来自逐帧编码在潜在空间中拼接无所谓;它甚至能够在零样本情况下对真实视频潜变量进行编辑,达到明显可识别的程度,仅留下微调以弥补保真度的差距。我们的结果表明,尽管在训练视频潜在空间上投入巨大,逐帧视频潜变量仍然与图像领域足够接近,以至于成熟的图像编辑先验可以在最小适应下转移。项目页面:https://yunpeng1998.github.io/Qwen-Video-Edit-Page;代码:https://github.com/yunpeng1998/Qwen-Video-Edit;模型:https://huggingface.co/yunpeng1998/Qwen-Video-Edit。
cs.CV / 29 / 2608.14796

Zero-Shot Adaptation of Medical Vision Foundation Models for High-Frequency Micro-Ultrasound Prostate Segmentation

医学视觉基础模型在高频微超声前列腺分割中的零样本适应
Abbas, Ayusha, Abbas, Saram, Adhikari, Kabita
Abstract
Prostate cancer claims a life every 80 seconds. Early detection is needed to prevent disease progression, and both PSA density calculation and biopsy decisions rely on knowing the exact boundary of the gland. Conventional ultrasound at 6-12 MHz blurs this boundary, missing one in three high-risk cancers. Micro-ultrasound (29 MHz) improves resolution threefold but introduces dense acoustic speckle that obscures the outer wall; given the same image, two clinicians draw outlines differing by over 10% in area. Supervised methods are costly and generalise poorly across scanners. Can a foundation model segment the prostate with no training data? We present the first zero-shot pipeline for this modality: MedSAM, pre-trained on over 1.5 million medical images, localises the prostate; we then apply CLAHE to sharpen the outer wall, binary dilation to recover missed pixels, and Fourier smoothing (4 modes, s=1.05) to refine the boundary. MedSAM requires a spatial prompt, so we evaluate bounding-box and point-click strategies across 75 patients of the Micro-Ultrasound Prostate Segmentation dataset (2,621 slices). On the 20-patient held-out test set, the pipeline reduces mean boundary-distance error by 45% (Dice 0.749+/-0.043 to 0.865+/-0.029; HD95 217.2+/-36.9 to 120.1+/-26.1 px), reaching Dice 0.859 across the cohort. Its mean overlap shows no significant difference from the three non-expert rater groups (p>0.19), while segmenting 38-52% more consistently (lower inter-patient standard deviation). Point-click prompts fail regardless of placement (best Dice=0.350), because speckle gives no stable local contrast. Only an approximate bounding box is required, so any clinic can deploy it without data collection, annotation, or retraining.
Chinese Translation
前列腺癌每80秒夺去一条生命。早期检测对于防止疾病进展至关重要,而PSA密度计算和活检决策依赖于准确了解腺体的边界。传统的6-12 MHz超声模糊了这一边界,导致每三例高风险癌症中就有一例被漏诊。微超声(29 MHz)将分辨率提高了三倍,但引入了密集的声斑,遮蔽了外壁;在相同图像下,两位临床医生绘制的轮廓在面积上相差超过10%。监督学习方法成本高且在不同扫描仪之间泛化效果差。基础模型能否在没有训练数据的情况下分割前列腺?我们提出了该模态的首个零样本管道:MedSAM,在超过150万张医学图像上进行预训练,能够定位前列腺;随后我们应用CLAHE技术锐化外壁,使用二值膨胀恢复漏掉的像素,并采用傅里叶平滑(4种模式,s=1.05)来细化边界。MedSAM需要空间提示,因此我们在75名患者的微超声前列腺分割数据集中评估了边界框和点点击策略(共2621张切片)。在20名患者的保留测试集中,该管道将平均边界距离误差降低了45%(Dice从0.749±0.043提高到0.865±0.029;HD95从217.2±36.9降低到120.1±26.1像素),在整个队列中达到Dice 0.859。其平均重叠与三组非专家评估者之间没有显著差异(p>0.19),同时分割的一致性提高了38-52%(患者间标准差降低)。无论放置位置如何,点点击提示均未能成功(最佳Dice=0.350),因为声斑未提供稳定的局部对比度。只需一个近似的边界框,因此任何诊所都可以在无需数据收集、注释或重新训练的情况下部署该方法。
cs.CV / 30 / 2608.14811

Where the Cost Falls: A Deployment-Aware Adoption Order for Stability Enhancements to Cycle-Consistent Adversarial Networks

成本落在何处:一种考虑部署的循环一致性对抗网络稳定性增强的采用顺序
Hussein, Rowan, Ouf, Mohamed
Abstract
Teams that adopt cycle-consistent adversarial networks for unpaired image-to-image translation meet the same obstacles: adversarial training oscillates or collapses, cycle consistency preserves coarse layout while finer texture drifts, and a single discriminator judging global realism misses local artifacts. Four enhancements address these failures, and they are usually compared on output quality alone. We show that they also divide sharply by where their cost falls, and that this division, which follows from the architecture and not from any particular run, yields an adoption order for teams under a compute or latency budget. A Wasserstein objective with gradient penalty, a VGG19 perceptual loss on the cycle reconstruction, and multi-scale discriminators change training only, so a team can adopt or drop them without altering what ships. Self-attention alone persists into the deployed generator, with memory growing as the square of the feature-map size, which makes it the one component a resource-constrained team should defer. We integrate all four onto a lightly tuned baseline for horse-to-zebra translation, introduced one at a time on a fixed control and then combined, and for each we give the failure mode it targets and how it integrates. We document the collapse and reconstruction-artifact modes the baseline produced, report what visual inspection of saved samples showed for each variant, and report Fr\'echet Inception Distance and Kernel Inception Distance for the combined model. We specify the protocol still needed, covering the individual variants, perceptual similarity, and downstream segmentation, to rank these enhancements on measured evidence.
Chinese Translation
采用循环一致性对抗网络进行无配对图像到图像转换的团队面临相同的障碍:对抗训练会出现振荡或崩溃,循环一致性保持粗略布局但细节纹理漂移,单一判别器判断全局真实感时忽略局部伪影。四种增强方法解决了这些失败,通常仅通过输出质量进行比较。我们展示了这些增强方法在成本分布上的明显差异,这种差异源于架构而非特定运行,进而为在计算或延迟预算下的团队提供了一种采用顺序。带有梯度惩罚的Wasserstein目标、在循环重建上的VGG19感知损失以及多尺度判别器仅改变训练,因此团队可以在不改变最终输出的情况下选择采用或放弃它们。自注意力机制则持续存在于部署的生成器中,随着特征图大小的平方而增长的内存使其成为资源受限团队应当推迟的唯一组件。我们将这四种增强方法整合到一个轻微调整的基线模型中,用于马到斑马的转换,逐一引入固定控制,然后组合,对于每种方法,我们给出了其针对的失败模式及其整合方式。我们记录了基线模型产生的崩溃和重建伪影模式,报告了对每个变体保存样本的视觉检查结果,并报告了组合模型的Fréchet Inception Distance和Kernel Inception Distance。我们明确了仍需遵循的协议,涵盖各个变体、感知相似性和下游分割,以便根据测量证据对这些增强方法进行排名。
cs.CV / 31 / 2608.14835

OvDSGG: End-to-End Open-Vocabulary Dynamic Scene Graph Generation

OvDSGG:端到端开放词汇动态场景图生成
Helsby, John, Yang, Yi, Rosenhahn, Bodo, Yang, Michael Ying
Abstract
Dynamic scene graphs (DSGs) capture spatio-temporal interactions across videos as $\langle$subject, predicate, object$\rangle$ triplets, and underpin downstream tasks such as video captioning, video question answering, and action analysis. However, end-to-end dynamic scene graph generation (DSGG) methods are closed-set: they recognize only objects and predicates from a fixed training vocabulary and struggle with the long-tailed distribution of rare concepts, severely limiting their real-world applicability. Existing open-vocabulary models typically inherit pretrained large language models, resulting in multi-stage training and inference with substantial cost. We introduce OvDSGG, the first end-to-end framework for open-vocabulary DSGG. OvDSGG builds on top of an open-vocabulary Spatial Backbone and a Temporal Backbone; we further propose a Triplet Feature Extraction Module that bridges them, and a Visual-Language Alignment Module that preserves open-vocabulary recognition by learning an adaptive decision boundary in the joint visual-language feature space, without expensive knowledge distillation in existing methods. We further introduce a rigorous open-vocabulary DSGG benchmark adapted from Action Genome, with disjoint Base/Novel splits for both objects and predicates. OvDSGG significantly outperforms open-vocabulary baselines across all metrics, with zero-shot Recall@$K$ scores 10.0--20.4 percentage point higher than the next-best baseline, while on closed-set DSGG remaining competitive with state-of-the-art models. Code and benchmark are publicly available at https://github.com/jhelsby/OvDSGG/.
Chinese Translation
动态场景图(DSGs)以 $ extlangle$主体,谓词,对象$ extrangle$ 三元组的形式捕捉视频中的时空交互,并支撑视频字幕生成、视频问答和动作分析等下游任务。然而,端到端动态场景图生成(DSGG)方法是封闭集的:它们仅识别来自固定训练词汇的对象和谓词,并且在稀有概念的长尾分布中表现不佳,严重限制了其在现实世界中的适用性。现有的开放词汇模型通常继承预训练的大型语言模型,导致多阶段的训练和推理,成本高昂。我们提出了OvDSGG,这是第一个端到端的开放词汇DSGG框架。OvDSGG建立在开放词汇空间主干和时间主干之上;我们进一步提出了一个三元特征提取模块,将它们连接起来,以及一个视觉-语言对齐模块,通过在联合视觉-语言特征空间中学习自适应决策边界来保持开放词汇识别,而无需现有方法中昂贵的知识蒸馏。我们还引入了一个严格的开放词汇DSGG基准,改编自Action Genome,具有对象和谓词的非重叠基础/新颖划分。OvDSGG在所有指标上显著超越开放词汇基线,零样本召回@$K$的得分比下一个最佳基线高出10.0--20.4个百分点,同时在封闭集DSGG中与最先进的模型保持竞争力。代码和基准可在 https://github.com/jhelsby/OvDSGG/ 上公开获取。
cs.CV / 32 / 2608.14854

Zero-MELO: Test-Time Evidence Calibration with Multimodal LLMs for Zero-Shot Micro-Gesture Recognition

Zero-MELO:基于多模态大语言模型的零-shot微手势识别测试时证据校准
Wang, Chengyan, Xie, Hanliang, Yang, Yueyi, Chen, Haoyu
Abstract
While Multimodal Large Language Models (MLLMs) excel in general video understanding, their capability in fine-grained and motion-centric tasks remains limited. This limitation is particularly critical in micro-gesture recognition (MGR), where micro-gestures (MGs) - subtle, short-duration, and spatially localized human movements - serve as key discriminative signals for implicit affective analysis, yet are easily neglected following common prompting practices. Although MGR has been intensively studied by many discriminative approaches, the use of MLLMs for MGR is underexplored, with notably poor performance. We hypothesize that the motion-sensitive representation ability of MLLMs is constrained by their inherent single-pass forward inference, which can be substantially enhanced through carefully designed test-time guidance. Motivated by this, building on our prior findings regarding temporal insensitivity in Video LLMs, we diagnose zero-shot MGR errors in the Negative Log-Likelihood (NLL) space. We observe that MLLMs suffer from two bottlenecks: 1) insufficient localized evidence and 2) severe score biases driven by language and motion-agnostic appearances. Thus, we propose a novel test-time evidence calibration framework that improves both reasoning details and prediction reliability. Specifically, we introduce a tree search mechanism to progressively acquire localized, fine-grained visual evidence, coupled with a test-time calibration module to mitigate score biases. The multi-cue fusion module then integrates evidence from multiple cues without relying on a single cue for final prediction. Our framework achieves mean-class accuracies of 26.84\% on iMiGUE and 22.10\% on MA-52, significantly outperforming the Qwen2.5-VL baseline, which produces 16.15\% and 10.20\%, respectively. The code will be available at https://zero-melo.github.io/Zero-MELO.
Chinese Translation
尽管多模态大语言模型(MLLMs)在一般视频理解方面表现出色,但它们在细粒度和以运动为中心的任务中的能力仍然有限。这一局限性在微手势识别(MGR)中尤为关键,微手势(MGs)——微妙、短暂且空间局限的人类运动——作为隐性情感分析的关键区分信号,常常在常见的提示实践中被忽视。尽管MGR已经被许多区分性方法深入研究,但MLLMs在MGR中的应用仍然未被充分探索,且表现显著较差。我们假设,MLLMs的运动敏感表示能力受到其固有的单次前向推理的限制,而通过精心设计的测试时指导可以显著增强这一能力。基于我们之前关于视频大语言模型时间不敏感性的发现,我们在负对数似然(NLL)空间中诊断零-shot MGR错误。我们观察到,MLLMs面临两个瓶颈:1)局部证据不足;2)由语言和与运动无关的外观驱动的严重评分偏差。因此,我们提出了一种新颖的测试时证据校准框架,旨在提高推理细节和预测可靠性。具体而言,我们引入了一种树搜索机制,以逐步获取局部的、细粒度的视觉证据,并结合测试时校准模块以减轻评分偏差。多线索融合模块则整合来自多个线索的证据,而不依赖于单一线索进行最终预测。我们的框架在iMiGUE上实现了26.84%的平均类别准确率,在MA-52上实现了22.10%,显著优于Qwen2.5-VL基线,其分别为16.15%和10.20%。代码将发布于https://zero-melo.github.io/Zero-MELO。
cs.CV / 33 / 2608.14868

Beam-Wise Statistical Background Subtraction for Static Roadside LiDAR: A Cross-Sensor Benchmark Study

静态路边激光雷达的束级统计背景减除:跨传感器基准研究
Baumann, Alexander, Vosshans, Marcel, Dang, Thao
Abstract
Background subtraction is a key preprocessing step for infrastructure-based LiDAR perception, enabling efficient isolation of dynamic traffic participants without semantic annotations. However, systematic cross-sensor evaluations and reproducible studies for static roadside LiDAR are missing. This paper presents a comparative benchmark of beam-wise statistical background subtraction for statically mounted LiDAR sensors. We formulate background estimation as a per-beam temporal modeling problem and investigate complementary statistical strategies that capture dominant as well as multi-modal background structures, combined with spatial filtering in the angular and 3D domain. To enable reproducible evaluation, we introduce HighwayScene, a new multi-LiDAR dataset recorded in a static roadside setup, and extend the public CoopScenes dataset with static/dynamic point-wise annotations. Across multiple scenes and heterogeneous sensing technologies, we demonstrate that beam-wise statistical modeling provides a robust and transferable solution. Combining lightweight per-beam models with spatial consistency filtering substantially improves precision while maintaining high recall and real-time capability. All datasets, annotations, and implementations are publicly released.
Chinese Translation
背景减除是基础设施激光雷达感知的关键预处理步骤,使得在没有语义注释的情况下高效隔离动态交通参与者。然而,针对静态路边激光雷达的系统性跨传感器评估和可重复研究尚缺乏。本文呈现了一项针对静态安装激光雷达传感器的束级统计背景减除的比较基准研究。我们将背景估计形式化为每束的时间建模问题,并探讨了捕捉主导及多模态背景结构的互补统计策略,结合了角度和三维域的空间滤波。为了实现可重复的评估,我们引入了HighwayScene,这是一个在静态路边设置下记录的新型多激光雷达数据集,并扩展了公共的CoopScenes数据集,增加了静态/动态点位注释。在多个场景和异构传感技术中,我们证明了束级统计建模提供了一种稳健且可转移的解决方案。将轻量级的每束模型与空间一致性滤波相结合,显著提高了精度,同时保持了高召回率和实时能力。所有数据集、注释和实现均已公开发布。
cs.CV / 34 / 2608.14922

SpIn-ViT: Designing a Sparsity-Induced Vision Transformer That Is Mechanistically Interpretable

SpIn-ViT:设计一种机制可解释的稀疏诱导视觉变换器
Lee, Philip H., Padalkar, Parth
Abstract
Mechanistic interpretability has recently expanded to Vision Transformers (ViTs), with Sparse Autoencoders (SAEs) increasingly used as post-hoc tools to decompose internal representations into sparse and more interpretable features. However, because post-hoc SAEs are trained on frozen representations after the ViT has already been optimized, their latent features are not directly aligned with the downstream classification objective. We introduce SpIn-ViT, a framework that jointly trains a pretrained ViT and a modified SAE end-to-end, directly aligning sparse patch-level representations with image classification. SpIn-ViT learns semantically coherent neuron activations that localize meaningful image regions while maintaining competitive predictive performance. We evaluate SpIn-ViT across nine image-classification benchmarks using classification accuracy, quantitative interpretability metrics, AI-based and Human evaluations. Compared with the previous state-of-the-art post-hoc SAE method, SpIn-ViT achieves 8.84% higher average classification accuracy, an AI-based interpretability score nearly four times as high, and a human-evaluation score more than twice as high. We further extract interpretable rule-sets using the SAE neurons to create neurosymbolic models which achieve 5.97% higher average classification accuracy while requiring a 58.8\% smaller rule-set than the neurosymbolic models created from the SOTA post-hoc SAE method.
Chinese Translation
机制可解释性最近已扩展到视觉变换器(ViTs),稀疏自编码器(SAEs)作为后处理工具越来越多地被用于将内部表示分解为稀疏且更具可解释性的特征。然而,由于后处理的SAEs是在ViT优化后对冻结表示进行训练的,因此它们的潜在特征与下游分类目标并不直接对齐。我们提出了SpIn-ViT,一个框架,它联合训练一个预训练的ViT和一个修改后的SAE,端到端地直接将稀疏的补丁级表示与图像分类对齐。SpIn-ViT学习到语义一致的神经元激活,能够定位有意义的图像区域,同时保持竞争性的预测性能。我们在九个图像分类基准上评估SpIn-ViT,使用分类准确率、定量可解释性指标、基于AI的评估和人类评估。与之前的最先进的后处理SAE方法相比,SpIn-ViT实现了平均分类准确率提高8.84%,基于AI的可解释性得分几乎高出四倍,以及人类评估得分超过两倍。我们进一步使用SAE神经元提取可解释的规则集,创建神经符号模型,其平均分类准确率提高5.97%,同时所需的规则集比基于最先进的后处理SAE方法创建的神经符号模型小58.8%。
cs.CV / 35 / 2608.14924

PaSTel: Anchoring Histology in Spatial Transcriptomics via Multi-Scale Hierarchical Bio-Prior Contrastive Pretraining

PaSTel:通过多尺度分层生物先验对比预训练将组织学锚定于空间转录组学
Amirabad, Azim Dehghani, Zhu, Junchao, Pati, Pushpak, Abdelmoula, Walid, Mansi, Tommaso, Liao, Rui
Abstract
Spatial transcriptomics (ST) links tissue morphology with molecular programs, motivating multimodal pretraining methods that align histology images with gene expression. However, existing approaches suffer from two key limitations: spatially informative gene selection is often dominated by ubiquitous housekeeping genes, leading to weakly discriminative representations, and independent spot-patch alignment fails to capture spatial dependencies that are critical for tissue organization. To address these challenges, we introduce PaSTel, a hierarchical multimodal pretraining framework that integrates biological priors at three levels. At the spot level, TF-IDF reweighting is used to identify spatially informative genes; at the functional level, curated KEGG pathways serve as anchors for encoding global biological semantics; and at the regional level, spatial clustering aggregates neighboring spots to model meso-scale tissue structure. Across multiple downstream tasks, PaSTel consistently outperforms existing vision and vision-omics encoders, demonstrating that incorporating multiscale biological priors yields more informative and transferable representations for spatial transcriptomics.
Chinese Translation
空间转录组学(ST)将组织形态与分子程序联系起来,激励了将组织学图像与基因表达对齐的多模态预训练方法。然而,现有方法存在两个主要限制:空间信息丰富的基因选择往往被普遍存在的管家基因主导,导致表征能力较弱;独立的点-补丁对齐未能捕捉对组织结构至关重要的空间依赖性。为了解决这些挑战,我们提出了PaSTel,一个分层多模态预训练框架,集成了三个层次的生物先验。在点级别,使用TF-IDF重加权来识别空间信息丰富的基因;在功能级别,经过整理的KEGG通路作为编码全球生物语义的锚;在区域级别,空间聚类将相邻点聚合,以建模中尺度组织结构。在多个下游任务中,PaSTel始终优于现有的视觉和视觉-组学编码器,证明了纳入多尺度生物先验能够为空间转录组学提供更具信息性和可转移性的表征。
cs.CV / 36 / 2608.14942

Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage Assessment

外表可能具有误导性:众包航空损害评估中不同影像来源的标注者和审阅者表现
Manzini, Thomas, Perali, Priyankari, Karnik, Raisa, Johnson, Stephen, Murphy, Robin R.
Abstract
This paper presents the first known empirical investigation of annotator and reviewer performance across multi-source remotely sensed imagery, evaluating human labeling across drone, crewed aviation, and satellite views. Because existing aerial imagery datasets rely predominantly on single-source imagery, there is no currently established state of practice for efficiently allocating human labor to curate large-scale, multi-source aerial datasets. This work addresses this limitation by analyzing annotator and reviewer performance within a post-disaster building damage assessment dataset of 9 disasters, where 20041 buildings in drone, 20695 buildings in crewed aviation, and 33392 buildings in satellite imagery were labeled. These labels, provided by 187 annotators, were then refined through two successive quality-control stages: a single-reviewer pass followed by a consensus-committee review. Our analysis reveals two findings that raise questions for standard crowd-sourcing practices. First, initial annotations were revised by the final committee at rates that rise steeply from higher- to lower-resolution sources (25.27% for crewed aviation and 36.95% for satellite), with the same ordering at every observed workflow stage. Second, a single individual review reduced but did not resolve this disagreement: after review, the committee still revised 6.85% of drone, 14.05% of crewed, and 20.86% of satellite labels. These observations suggest that, in workflows like this one, uniform review allocation leaves the most residual disagreement in lower-resolution imagery. Based on this evidence, and consistent with prior work on adaptive task assignment and budget-aware quality control, this paper offers three recommendations for multi-source dataset curation.
Chinese Translation
本文首次对多源遥感影像中标注者和审阅者的表现进行了实证研究,评估了无人机、载人航空和卫星视角下的人类标注。由于现有的航空影像数据集主要依赖单一来源影像,因此目前尚无有效分配人力以策划大规模多源航空数据集的实践标准。本研究通过分析在9个灾后建筑损害评估数据集中标注者和审阅者的表现来解决这一局限性,该数据集包含20041栋无人机影像、20695栋载人航空影像和33392栋卫星影像的标注。这些标签由187名标注者提供,随后经过两个连续的质量控制阶段进行精炼:首先是单一审阅者的审核,然后是共识委员会的审查。我们的分析揭示了两个发现,这对标准众包实践提出了质疑。首先,最终委员会对初始标注的修订率在高分辨率到低分辨率影像来源之间急剧上升(载人航空为25.27%,卫星为36.95%),在每个观察到的工作流程阶段均保持相同的顺序。其次,单一审阅者的审核虽然减少了但并未解决这一分歧:审核后,委员会仍然修订了6.85%的无人机、14.05%的载人航空和20.86%的卫星标签。这些观察结果表明,在类似的工作流程中,统一的审核分配在低分辨率影像中留下了最多的残余分歧。基于这一证据,并与之前关于自适应任务分配和预算意识质量控制的研究一致,本文为多源数据集的策划提出了三项建议。
cs.CV / 37 / 2608.14976

Benchmarking Frontier Text-to-Image Models on Image-Description Prompts

前沿文本到图像模型在图像描述提示上的基准测试
Abdoli, Sajjad, Al-Sumaidaee, Ghassan, Rashad, Ahmed
Abstract
Text-to-image models are typically reported on average-case prompts, which understates the gap between systems on compositionally demanding requests involving precise object counts, multi-object attribute binding, legible embedded text, and explicit spatial constraints. We evaluate four production text-to-image systems: Hunyuan 3.0, Gemini 3 Pro Image ("Nano Banana Pro"), Black Forest Labs FLUX.2, and Ideogram 3.0. The evaluation uses the 48 hardest prompts drawn from the DataSeeds.AI Sample Dataset (DSD), selected through an automated complexity-scoring pass over the full corpus. Every generated image is graded using an independent-judge rubric. GPT-5.4-Pro authors an atomic, weighted, mutually exclusive and collectively exhaustive (MECE) evaluation rubric, while Gemini 3.1 Pro Preview independently determines whether each criterion is satisfied. Gemini 3 Pro Image ranks first with a score of 84.8/100, narrowly ahead of FLUX.2 at 82.3/100. Ideogram 3.0 and Hunyuan 3.0 score 65.7/100 and 63.3/100, respectively. Failure analysis shows that the leading systems primarily lose points through object miscounting and geometric artifacts, whereas the trailing systems more frequently produce garbled text. Ideogram 3.0 also frequently omits requested elements. Full per-sample rubrics, scores, and failure annotations are available from the authors upon request.
Chinese Translation
文本到图像模型通常在平均案例提示上进行报告,这低估了系统在涉及精确物体计数、多物体属性绑定、可读嵌入文本和明确空间约束的组合性要求上的差距。我们评估了四个生产级文本到图像系统:Hunyuan 3.0、Gemini 3 Pro Image(“Nano Banana Pro”)、Black Forest Labs FLUX.2 和 Ideogram 3.0。评估使用了从 DataSeeds.AI 样本数据集(DSD)中抽取的48个最难提示,这些提示是通过对完整语料库进行自动复杂度评分筛选得出的。每个生成的图像都使用独立评审的评分标准进行评分。GPT-5.4-Pro 编写了一个原子化、加权的、互斥且完全穷尽的(MECE)评估标准,而 Gemini 3.1 Pro Preview 独立判断每个标准是否满足。Gemini 3 Pro Image 以84.8/100的分数排名第一,略微领先于 FLUX.2 的82.3/100。Ideogram 3.0 和 Hunyuan 3.0 的得分分别为65.7/100和63.3/100。失败分析显示,领先系统主要因物体计数错误和几何伪影而失分,而落后系统则更频繁地产生乱码文本。Ideogram 3.0 也常常遗漏请求的元素。完整的逐样本评分标准、分数和失败注释可根据请求向作者获取。
cs.CV / 38 / 2608.14991

Risk-Adaptive Edge--Cloud Visual Reasoning for Communication-Efficient Autonomous Driving

风险适应性边缘-云视觉推理用于通信高效的自动驾驶
Ma, Meng, Li, Shuyang, Wang, Naigang, Ke, Ruimin
Abstract
Cloud-hosted vision-language models (VLMs) offer greater contextual reasoning capabilities than smaller onboard models, but frequent visual uploads increase communication overhead and add network and inference latency to tactical decisions. We present a risk-adaptive edge-cloud architecture in which onboard traffic assessment determines when cloud reasoning is requested. An onboard VLM and a lightweight detector capture temporal traffic conditions and path-relative hazards for conservative local response and selective cloud access. The cloud model provides tactical advice, while validation, vehicle control, and automatic emergency braking remain local. In CARLA experiments, our method matched the task success rate of periodic cloud access while reducing cloud requests by 54.1% and recording fewer automatic emergency braking (AEB) activations. In a delayed-roadwork ablation, semantic events triggered requests before the next scheduled audit. Across three emulated network profiles, the method continued to reduce cloud traffic, although lane changes took longer than with periodic access. Onboard traffic assessment therefore served as a practical trigger for selective VLM inference in these experiments.
Chinese Translation
云托管的视觉-语言模型(VLMs)提供了比较小的车载模型更强的上下文推理能力,但频繁的视觉上传增加了通信开销,并为战术决策带来了网络和推理延迟。我们提出了一种风险适应性边缘-云架构,其中车载交通评估决定何时请求云端推理。车载VLM和轻量级检测器捕获时间性交通状况和路径相关的危险,以实现保守的本地响应和选择性的云访问。云模型提供战术建议,而验证、车辆控制和自动紧急制动仍然在本地进行。在CARLA实验中,我们的方法在周期性云访问的任务成功率上表现相当,同时将云请求减少了54.1%,并记录了更少的自动紧急制动(AEB)激活。在延迟的道路施工消融实验中,语义事件在下一个计划审计之前触发了请求。在三个模拟的网络配置中,该方法继续减少云流量,尽管车道变换所需时间比周期性访问更长。因此,车载交通评估在这些实验中作为选择性VLM推理的实际触发器。
cs.CV / 39 / 2608.14994

Registration-Free Hyperspectral Reconstruction from RGB via a Permutation-Invariant Gram-Matrix Principle

基于置换不变Gram矩阵原理的无注册高光谱重建方法
Zhao, Jiangsan, Hirafuji, Masayuki, Ninomiya, Seishi, Geipel, Jakob, Guo, Wei
Abstract
Reconstructing a spatially and spectrally high-resolution hyperspectral image (HR-HSI) from a low-resolution HSI (LR-HSI) and a high-resolution RGB image (HR-RGB) usually assumes precise registration and a known camera response function (CRF). Both assumptions are difficult to satisfy with different sensors. We remove both through a permutation-invariant supervision principle: the Gram matrix of an unmixed abundance map depends on shared material composition but not on pixel ordering. Matching abundance Gram matrices therefore allows RGB-to-HSI mapping to be learned without spatial correspondence and without a predefined CRF. Under a full random permutation of HR-RGB pixels, a state-of-the-art fusion method collapses, whereas our reconstruction is unchanged after inverse reindexing for evaluation. Building on this principle, a residual spectral super-resolution function maps HR-RGB directly to HR-HSI without registration, known CRF, or paired supervision. Across indoor, natural-scene, and remote-sensing benchmarks, the method achieves accuracy comparable to approaches that require these assumptions while remaining robust when they are violated. Loss ablations further show that reconstruction accuracy is largely insensitive to the specific discrepancy used to match the Gram matrices, indicating that performance arises primarily from the permutation-invariant principle rather than loss tuning.
Chinese Translation
从低分辨率高光谱图像(LR-HSI)和高分辨率RGB图像(HR-RGB)重建空间和光谱高分辨率高光谱图像(HR-HSI)通常假设存在精确的配准和已知的相机响应函数(CRF)。这两个假设在不同传感器之间很难满足。我们通过置换不变的监督原理去除了这两个假设:未混合丰度图的Gram矩阵依赖于共享的材料组成,而不依赖于像素顺序。因此,匹配丰度Gram矩阵使得RGB到HSI的映射可以在没有空间对应关系和预定义CRF的情况下进行学习。在HR-RGB像素完全随机置换的情况下,最先进的融合方法会崩溃,而我们的重建在评估时经过逆重标记后保持不变。基于这一原理,残差光谱超分辨率函数直接将HR-RGB映射到HR-HSI,而无需配准、已知CRF或配对监督。在室内、自然场景和遥感基准测试中,该方法的准确性与需要这些假设的方法相当,同时在这些假设被违反时仍然保持鲁棒性。损失消融实验进一步表明,重建准确性对用于匹配Gram矩阵的特定差异不敏感,表明性能主要源于置换不变原理,而非损失调优。
cs.CV / 40 / 2608.15004

FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making

FZ-VLM:一种用于肺结节特征描述和临床决策的两阶段Florence-Zephyr视觉语言模型框架
Dutta, Pramit, Manokaran, Jenita, Mittal, Richa, Appleby, Ryan, Ukwatta, Eranga
Abstract
Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool for screening and followup assessment. After pulmonary nodule detection, radiologists manually assess anatomical location, diameter, margin characteristics, and attenuation type to support risk assessment and clinical decision-making. However, this post-detection workflow is time-consuming and can be affected by inter-observer variability. Existing Artificial Intelligence methods often focus on isolated tasks, limiting their use as a unified, clinically grounded interpretation framework. This study presents FZ-VLM, a two-stage Florence-Zephyr Vision Language Model framework for unified structured pulmonary nodule characterization in lung CT. The framework uses a fine-tuned Florence-2 model to extract radiological attributes from expert-annotated 2D axial CT slices, while a Zephyr-7B model uses these attributes to generate nodule descriptions, follow-up recommendations, and longitudinal analyses. Results showed that the Stage 1 model achieved 77.18\% accuracy for anatomical location, 67.96\% accuracy for margin characteristics, and 79.13\% accuracy for attenuation type, with a Mean Absolute Error of 2.58 mm for diameter estimation, outperforming evaluated GPT-4-based baselines as well as the human baseline. Expert radiologist evaluation of Stage 2 showed 93.9\% accuracy, 98.6\% completeness score, 76.1\% clinical relevance, and an overall score of 89.5\%. Safety analysis showed that most outputs were clinically safe, although some follow-up recommendations still required expert review. To the best of our knowledge, this study presents the first two-stage Vision-Language Model framework for structured nodule characterization and clinical decision-making.
Chinese Translation
肺癌仍然是全球癌症相关死亡的主要原因之一,计算机断层扫描(CT)是筛查和后续评估的主要影像学工具。在肺结节检测后,放射科医生手动评估解剖位置、直径、边缘特征和衰减类型,以支持风险评估和临床决策。然而,这一检测后工作流程耗时且可能受到观察者间变异的影响。现有的人工智能方法通常专注于孤立的任务,限制了其作为统一的、临床基础的解释框架的使用。本研究提出了FZ-VLM,一种用于肺CT中统一结构化肺结节特征描述的两阶段Florence-Zephyr视觉语言模型框架。该框架使用经过微调的Florence-2模型从专家标注的2D轴向CT切片中提取放射学特征,而Zephyr-7B模型则利用这些特征生成结节描述、后续建议和纵向分析。结果显示,第一阶段模型在解剖位置的准确率为77.18%,边缘特征的准确率为67.96%,衰减类型的准确率为79.13%,直径估计的平均绝对误差为2.58毫米,优于评估的基于GPT-4的基线以及人类基线。第二阶段的专家放射科医生评估显示准确率为93.9%,完整性评分为98.6%,临床相关性为76.1%,总体评分为89.5%。安全性分析表明,大多数输出在临床上是安全的,尽管一些后续建议仍需专家审查。根据我们所知,本研究首次提出了用于结构化结节特征描述和临床决策的两阶段视觉语言模型框架。
cs.CV / 41 / 2608.15006

MetaReason: Precise Interleaved Multimodal Reasoning via Editing Meta Information for Solving Geometry Problems

MetaReason:通过编辑元信息实现精确交错的多模态推理以解决几何问题
Yin, Penghao, Wang, Haomin, Tang, Qihong, Qu, Xiaoye, Zhang, Hongjie, Zhang, Xiao-Ping
Abstract
Although visual reasoning is crucial for solving complex geometry tasks, existing vision-language models rely heavily on text-only reasoning. Some recent methods introduce intermediate visual states to facilitate reasoning, but they are often hindered by inaccurate geometric representations and low rendering fidelity, ultimately leading to unreliable outputs. To address these limitations, we propose MetaReason, a framework for multimodal reasoning in plane geometry that leverages structured meta-information to enable accurate auxiliary-line construction. The framework first parses geometric images into meta-information, performs controllable edits with predefined tools to synthesize high-fidelity visual states, and then conducts reasoning based on these augmented views. To support this framework, we construct TutorGeo, a comprehensive dataset containing 17k image-to-meta conversion samples, 60k text-only reasoning traces, and 60k interleaved multimodal reasoning traces. Using this dataset, we combine supervised fine-tuning and reinforcement learning to develop robust multimodal reasoning capabilities. We also introduce ExamGeo, a benchmark derived from real-world examination problems that enables systematic evaluation across varying difficulty levels. Experimental results demonstrate that MetaReason significantly outperforms existing open-source models and achieves competitive performance against proprietary models.
Chinese Translation
尽管视觉推理对于解决复杂的几何任务至关重要,但现有的视觉-语言模型在很大程度上依赖于仅文本推理。一些近期的方法引入了中间视觉状态以促进推理,但它们常常受到不准确的几何表示和低渲染保真度的限制,最终导致不可靠的输出。为了解决这些局限性,我们提出了MetaReason,一个用于平面几何的多模态推理框架,该框架利用结构化的元信息来实现准确的辅助线构建。该框架首先将几何图像解析为元信息,使用预定义工具进行可控编辑以合成高保真的视觉状态,然后基于这些增强视图进行推理。为了支持该框架,我们构建了TutorGeo,一个包含17,000个图像到元转换样本、60,000个仅文本推理轨迹和60,000个交错多模态推理轨迹的综合数据集。利用该数据集,我们结合监督微调和强化学习来开发稳健的多模态推理能力。我们还引入了ExamGeo,一个源自现实世界考试问题的基准,能够在不同难度级别上进行系统评估。实验结果表明,MetaReason显著优于现有的开源模型,并在与专有模型的比较中表现出竞争力。
cs.CV / 42 / 2608.15019

DualMiT-Net: Local-Global Transformer-Convolutional Fusion for Breast Mass Segmentation in Mammographic Regions of Interest

DualMiT-Net:用于乳腺区域内乳腺肿块分割的局部-全局变换卷积融合
Kamiluly, Alibek, Muratova, Milana, Patel, Yash, Li, Fan
Abstract
Breast mass segmentation is an important step in computer-aided mammography, but it remains difficult because masses can have low contrast, irregular shapes, and boundaries that blend with surrounding breast tissue. To address this problem, we present DualMiT-Net, a dual-branch network that uses both a focused view of the mass and a wider view of the surrounding tissue. The local branch uses a Mix Transformer (MiT-B5) encoder to learn mass shape, texture, and boundary information, while the global branch uses an EfficientNet-B5 encoder to learn surrounding breast context. Features from the two branches are shared at the deeper encoder levels and are then progressively fused in a single decoder. A spatial gate controls how much global information is added during decoding. We also evaluated four input representations and selected a percentile-windowed mammogram combined with a Gabor texture response. The model was trained and evaluated on the mass subset of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM) using a patient-level split. Across three training runs, DualMiT-Net with exponential moving average weights achieved a mean Dice coefficient of 0.9375 and a mean Intersection over Union of 0.8834. It also achieved better Dice and IoU scores than six standard encoder-decoder baselines trained using the same data and training settings. These results show that combining local mass information with wider breast context can provide accurate and consistent breast mass segmentation.
Chinese Translation
乳腺肿块分割是计算机辅助乳腺摄影中的一个重要步骤,但由于肿块可能具有低对比度、不规则形状以及与周围乳腺组织边界融合等特征,因此仍然很困难。为了解决这个问题,我们提出了DualMiT-Net,这是一种双分支网络,利用肿块的聚焦视图和周围组织的更广泛视图。局部分支使用Mix Transformer (MiT-B5)编码器来学习肿块的形状、纹理和边界信息,而全局分支则使用EfficientNet-B5编码器来学习周围乳腺的上下文信息。两个分支的特征在更深的编码器层次上共享,然后在单一解码器中逐步融合。一个空间门控控制在解码过程中添加多少全局信息。我们还评估了四种输入表示,并选择了与Gabor纹理响应相结合的百分位窗乳腺摄影。该模型在数字乳腺摄影筛查数据库(CBIS-DDSM)的乳腺肿块子集上进行了训练和评估,采用患者级别的划分。在三次训练中,使用指数移动平均权重的DualMiT-Net达到了0.9375的平均Dice系数和0.8834的平均交并比(IoU)。它还在相同数据和训练设置下,优于六个标准编码器-解码器基线的Dice和IoU得分。这些结果表明,将局部肿块信息与更广泛的乳腺上下文相结合,可以提供准确且一致的乳腺肿块分割。
cs.CV / 43 / 2608.15028

Geometry-Calibrated Closed-Form Shrinkage for SAR Despeckling

几何校准的闭式收缩用于合成孔径雷达去斑
Hu, Xuran, Zhu, Mingzhe, Stanković, Djordje, Zhu, Yujie, Feng, Zhenpeng, Ban, Yifang, Stanković, Ljubiša
Abstract
Synthetic aperture radar (SAR) despeckling is an inverse-recovery problem in which multiplicative non-Gaussian noise must be suppressed without erasing scattering structures. We revisit a nonlocal sparse estimator that applies a log--Yeo--Johnson transformation, stacks similar patches into groups, codes each group on its own left singular basis, and shrinks the resulting coefficients. Three quantities usually treated as tunable are shown to be fixed by this construction. First, the group dictionary is orthonormal, so the weighted Lasso admits an exact coefficient-wise soft-threshold solution: the iterative inner solver is unnecessary, and the two apparent weighting matrices are the numerator and denominator of a single threshold field rather than independent modules. Second, because the dictionary is estimated from the noisy group itself, its retained subspace absorbs speckle in proportion to the group aspect ratio $\gamma=p^2/K$; a random-matrix argument converts the corresponding regularization constant into a geometry-calibrated correction and collapses patch size, group size, and shrinkage scale into one analytically determined degree of freedom. Third, singular projection makes the coefficient noise nearly Gaussian at every tested look number, which locates the point at which an exact speckle likelihood ceases to be informative. The resulting estimator is deterministic, training-free, and applies one set of analytically determined settings to every image and sensor. It ranks first in 18 of 24 PSNR/SSIM comparisons against twelve published methods on three synthetic benchmarks, and attains the lowest mean deviation of the ratio image from the theoretical speckle model over six real-SAR configurations from five sensors. Code is available \href{https://github.com/Teriri1999/Geometry-Calibrated-Closed-Form-Shrinkage-for-SAR-Despeckling}{here}.
Chinese Translation
合成孔径雷达(SAR)去斑是一个逆恢复问题,其中必须抑制乘法性非高斯噪声而不抹去散射结构。我们重新审视了一种非局部稀疏估计器,该估计器应用了对数-耶欧-约翰逊变换,将相似的图块堆叠成组,分别在其左奇异基上对每个组进行编码,并收缩得到的系数。通常被视为可调的三个量在这一构造中被证明是固定的。首先,组字典是正交归一的,因此加权Lasso允许精确的系数级软阈值解:迭代内部求解器是不必要的,两个明显的加权矩阵是单个阈值场的分子和分母,而不是独立模块。其次,由于字典是从噪声组本身估计的,其保留的子空间按组的长宽比$eta=p^2/K$吸收斑点;随机矩阵论证将相应的正则化常数转化为几何校准的修正,并将图块大小、组大小和收缩尺度合并为一个解析确定的自由度。第三,奇异投影使得每个测试的观察数下系数噪声几乎呈高斯分布,这确定了精确斑点似然不再具有信息性的点。所得到的估计器是确定性的,无需训练,并对每个图像和传感器应用一组解析确定的设置。在三种合成基准测试中,该方法在与十二种已发表方法的24次PSNR/SSIM比较中排名第一,并在来自五个传感器的六种真实SAR配置中达到了与理论斑点模型的比率图的最低均方差。代码可在此获取: exttt{https://github.com/Teriri1999/Geometry-Calibrated-Closed-Form-Shrinkage-for-SAR-Despeckling}。
cs.CV / 44 / 2608.15029

Generation of Synthetic Fingerphotos with GANs

使用生成对抗网络生成合成指纹照片
Miller-Lynch, Conor, Purnapatra, Sandip, Abbas, Syed Konain, Igene, Lambert, Hussain, Faraz, Dey, Soumyabrata, Schuckers, Stephanie
Abstract
Contactless fingerprinting is an emerging approach to biometric authentication that allows users to scan their fingerprints without touching a scanner. Due to the limited amount of contactless fingerprint data available and the security risks associated with sharing real individuals' fingerprints, it is valuable to explore methods of generating synthetic data that can be used in place of - or in conjunction with - real data to develop and evaluate contactless fingerprinting systems. In this paper, we present and evaluate synthetic fingerphotos generated using StyleGAN2-ADA and StyleGAN3, existing image generation architectures. We evaluate the realism, privacy preservation, and variety of the synthetic fingerphotos by comparing their biometric feature statistics to those of real fingerphotos, computing match scores between real and synthetic fingerphotos, and computing match scores between different synthetic fingerphotos. This paper provides a quantitative comparison point for future evaluations of synthetic fingerphotos. The evaluation code is made available at https://github.com/cmillerlynch/fingerphoto-gan.
Chinese Translation
无接触指纹识别是一种新兴的生物识别认证方法,允许用户在不接触扫描仪的情况下扫描指纹。由于可用的无接触指纹数据有限,以及与共享真实个体指纹相关的安全风险,探索生成合成数据的方法以替代或与真实数据结合使用,从而开发和评估无接触指纹识别系统具有重要价值。在本文中,我们展示并评估了使用 StyleGAN2-ADA 和 StyleGAN3 生成的合成指纹照片,这些是现有的图像生成架构。我们通过将合成指纹照片的生物特征统计与真实指纹照片进行比较、计算真实与合成指纹照片之间的匹配分数,以及计算不同合成指纹照片之间的匹配分数,来评估合成指纹照片的真实性、隐私保护和多样性。本文为未来对合成指纹照片的评估提供了量化比较的参考点。评估代码可在 https://github.com/cmillerlynch/fingerphoto-gan 获取。
cs.CV / 45 / 2608.15045

MOSS-VL Technical Report

MOSS-VL 技术报告
Wang, Pengyu, Tan, Chenkun, Zhou, Shaojun, Zhou, Qirui, Chen, Yanxin, He, Xingyang, Zeng, Huazheng, Cheng, Jijun, Wang, Chenghao, Qian, Xiaomeng, Wang, Pengfei, Huang, Zhan, Gao, Shanqing, Huang, Wei, Cao, Longjun, Ran, Wu, Liu, Jie, Zhu, Changtai, Wang, Hongkai, Tian, Yixian, Liu, Chenghao, Ye, Zhen, Wang, Xinghao, Jiang, Botian, Feng, Guoguo, Fei, Zhaoye, Li, Ruixiao, Chen, Mingshu, Gao, Yang, Cheng, Qinyuan, Li, Shimin, Qiu, Xipeng
Abstract
We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the language decoder attends to vision only through gated cross-attention, so the model can naturally see incoming frames while generating; a synthesized interaction corpus supervises when to speak, when to stay silent, and when to revise; and a staged curriculum concentrates all real-time-specific training in one light final stage over a strong offline foundation. Offline, MOSS-VL-Instruct is competitive at comparable scale and leads temporal-reasoning video sets. Across four streaming benchmarks, MOSS-VL-Realtime posts the best average on three (second on the fourth) among open-source streaming models, sweeping the three subsets that squarely test proactive behavior -- 66.0 vs. 37.5 for the best baseline on OmniMMI Proactive Alerting. With 11.3B parameters but visual tokens outside the decoded sequence, MOSS-VL widens its time-to-first-token advantage over same-backbone Qwen3-VL-8B from 2.8x to 5.1x as visual context grows. We release all five checkpoints, the training curriculum, and the real-time inference code at https://github.com/OpenMOSS/MOSS-VL.
Chinese Translation
我们提出了 MOSS-VL,一个开放的视觉-语言模型系列,将实时交互——在说话时感知——视为一项重要能力。该模型在各个层面上进行了协同设计:语言解码器仅通过门控交叉注意力关注视觉信息,因此模型在生成时可以自然地看到输入帧;合成的交互语料库监督何时发言、何时保持沉默以及何时进行修正;分阶段的课程将所有与实时相关的训练集中在一个轻量的最终阶段,建立在强大的离线基础之上。在离线环境中,MOSS-VL-Instruct 在可比规模下具有竞争力,并在时间推理视频集上表现优异。在四个流媒体基准测试中,MOSS-VL-Realtime 在三个基准上取得了最佳平均成绩(在第四个基准上排名第二),在三个专门测试主动行为的子集中表现突出——在 OmniMMI 主动警报中,得分为 66.0,远超最佳基线的 37.5。MOSS-VL 拥有 113 亿参数,但视觉标记位于解码序列之外,随着视觉上下文的增加,其首次标记的时间优势从与相同骨干网络的 Qwen3-VL-8B 的 2.8 倍扩大到 5.1 倍。我们在 https://github.com/OpenMOSS/MOSS-VL 发布了所有五个检查点、训练课程和实时推理代码。
cs.CV / 46 / 2608.15054

Frequency and Edge-Guided Segment Anything Model for Remote Sensing Image Semantic Segmentation

频率与边缘引导的任意分割模型用于遥感图像语义分割
Gao, Feng, Pan, Zizhe, Wang, Haoting, Hua, Ruzhuang, Cao, Jingchao, Dong, Junyu, Du, Qian
Abstract
Remote sensing image semantic segmentation (RSISS) has attracted significant attention due to the growing demand for fine-grained land cover information. The Segment Anything Model (SAM), proposed as a foundation vision model, offers strong segmentation performance and generalization capabilities for RSISS tasks. However, existing SAM-based approaches face two limitations: (1) Insufficient adaptation of SAM's features to the diverse characteristics of land cover types. (2) Semantic ambiguity at object boundaries, which hinders accurate delineation. To address these limitations, we propose Frequency and Edge-guided SAM (FE-SAM), a scalable and efficient framework for RSISS. Specifically, we introduce a Frequency-Modulated Adapter (FMA) that adaptively decomposes and modulates frequency-domain features based on the input data. It selectively enhances informative high- and low-frequency components corresponding to different land cover types. Furthermore, to improve SAM's ability to capture fine-grained details, we design EGRefiner, which integrates multi-scale edge-enhanced information extracted from the input image. Extensive experiments on three benchmark datasets demonstrate that FE-SAM outperforms state-of-the-art methods. The source codes are available at: https://github.com/oucailab/FE-SAM.
Chinese Translation
遥感图像语义分割(RSISS)因对细粒度土地覆盖信息的日益需求而受到广泛关注。作为基础视觉模型提出的任意分割模型(Segment Anything Model, SAM)在RSISS任务中展现出强大的分割性能和泛化能力。然而,现有的基于SAM的方法面临两个局限:(1)SAM特征对土地覆盖类型多样性特征的适应性不足;(2)物体边界处的语义模糊,阻碍了准确的描绘。为了解决这些局限性,我们提出了频率与边缘引导的SAM(Frequency and Edge-guided SAM, FE-SAM),这是一个可扩展且高效的RSISS框架。具体而言,我们引入了一种频率调制适配器(Frequency-Modulated Adapter, FMA),该适配器基于输入数据自适应地分解和调制频域特征,选择性地增强与不同土地覆盖类型对应的信息丰富的高频和低频成分。此外,为了提高SAM捕捉细粒度细节的能力,我们设计了EGRefiner,它集成了从输入图像中提取的多尺度边缘增强信息。在三个基准数据集上的大量实验表明,FE-SAM优于最先进的方法。源代码可在以下链接获取:https://github.com/oucailab/FE-SAM。
cs.CV / 47 / 2608.15058

MEDR: Query-Independent Frame Selection via Multi-Signal Event Modeling and Dynamic Rescoring

MEDR:通过多信号事件建模和动态重评分实现查询无关的帧选择
Pu, Xinlei, Shi, Weijie, Yang, Wen, Cao, Yi, Chen, Hao, Liu, Yuanjun, Ding, Wenwei, Zhu, Jia, Xu, Jiajie
Abstract
Frame selection is a fundamental component of multimodal large language models, enabling long videos to be processed under limited visual-token and computational budgets. Uniform sampling preserves temporal coverage but may miss informative content that appears only briefly. To alleviate this limitation, query-dependent methods can retrieve question-relevant frames. However, because the selected frames depend on the current question, the same visual input cannot be directly shared across different questions, and frame selection must be repeated in multi-turn video dialogue. This motivates us to seek a query-independent frame selection method that preserves the reusability of a fixed visual input while improving the coverage of informative events beyond uniform sampling. We propose Multi-Signal Event Modeling and Dynamic Rescoring (MEDR), a training-free and query-independent frame selection method. Multi-Signal Event Modeling organizes complementary visual, motion, and text signals into signal-specific temporal events. Dynamic Rescoring then iteratively reevaluates each candidate relative to the current selected set, updating its score according to frame-level signal strength, additional event coverage, and temporal proximity. The resulting fixed frame set is constructed without observing the query and can be reused across different questions. On the standard benchmark evaluations, MEDR improves model accuracy by 0.63%-0.89% on Video-MME. On the long-video subset of LongVideoBench, it improves accuracy by up to 1.23% with Qwen3-VL-8B. MEDR further improves overall accuracy by 0.53%, while reusing exactly the same frame set for every question about a video.
Chinese Translation
帧选择是多模态大型语言模型的一个基本组成部分,使得在有限的视觉标记和计算预算下能够处理长视频。均匀采样虽然能够保持时间覆盖,但可能会错过仅短暂出现的重要内容。为了解决这一限制,依赖查询的方法可以检索与问题相关的帧。然而,由于所选帧依赖于当前的问题,相同的视觉输入不能在不同问题之间直接共享,因此在多轮视频对话中必须重复进行帧选择。这促使我们寻求一种查询无关的帧选择方法,以保持固定视觉输入的可重用性,同时改善信息事件的覆盖范围,超越均匀采样。我们提出了多信号事件建模和动态重评分(MEDR),这是一种无训练且查询无关的帧选择方法。多信号事件建模将互补的视觉、运动和文本信号组织成信号特定的时间事件。动态重评分则通过相对于当前选择集迭代重新评估每个候选帧,根据帧级信号强度、额外事件覆盖和时间接近度更新其分数。最终构建的固定帧集在未观察查询的情况下形成,并可在不同问题中重复使用。在标准基准评估中,MEDR在Video-MME上提高了模型准确率0.63%-0.89%。在LongVideoBench的长视频子集上,使用Qwen3-VL-8B时准确率提高了最多1.23%。MEDR进一步提高了整体准确率0.53%,同时对每个关于视频的问题重用完全相同的帧集。
cs.CV / 48 / 2608.15060

EgoTac: In-the-wild Tactile Prediction from Egocentric Vision

EgoTac:从自我中心视觉中进行野外触觉预测
Zhang, Wenkang, Yuan, Chengbo, Zhang, Zicheng, Cheng, Zhengxue, Gao, Yang
Abstract
Touch is fundamental to dexterous manipulation, yet most egocentric human data increasingly used for robot learning lacks tactile information. Directly collecting large-scale tactile data is challenging due to sensor limitations, while human video data is abundant, contact-rich, and easily scalable. This motivates a natural question: can tactile signals be inferred purely from vision? To address this, we introduce EgoTac, a generalizable model that predicts rich tactile information directly from egocentric human videos. EgoTac is trained on a unified corpus of over 5.7M image-tactile pairs, covering both continuous force measurements and binary contacts. By learning from this diverse dataset, EgoTac captures nuanced touch dynamics across varied interactions. Experiments demonstrate strong performance: in-domain prediction achieves an average force error below 0.06N. On out-of-domain contact prediction benchmarks, EgoTac consistently outperforms the state-of-the-art contact estimator. It also captures the rise and fall patterns of real tactile data and enables zero-shot predictions on unconstrained real-world videos. Scaling analyses further reveal that both data diversity and volume improve performance steadily. Overall, EgoTac provides a scalable pathway to extract tactile priors from egocentric human videos, enabling broadly applicable tactile-aware robot learning.
Chinese Translation
触觉对于灵巧操作至关重要,但目前用于机器人学习的大多数自我中心人类数据缺乏触觉信息。由于传感器的限制,直接收集大规模触觉数据具有挑战性,而人类视频数据丰富、接触频繁且易于扩展。这引发了一个自然的问题:是否可以仅通过视觉推断触觉信号?为了解决这个问题,我们提出了EgoTac,这是一种可泛化的模型,可以直接从自我中心的人类视频中预测丰富的触觉信息。EgoTac在一个统一的语料库上进行训练,该语料库包含超过570万对图像-触觉数据,涵盖了连续的力测量和二元接触。通过从这个多样化的数据集中学习,EgoTac捕捉到不同交互中的细微触觉动态。实验表明其表现强劲:在领域内预测中,平均力误差低于0.06N。在领域外接触预测基准测试中,EgoTac始终优于最先进的接触估计器。它还捕捉到了真实触觉数据的上升和下降模式,并能够在不受约束的真实世界视频上进行零样本预测。规模分析进一步揭示了数据的多样性和数量都能稳步提高性能。总体而言,EgoTac提供了一条可扩展的路径,从自我中心的人类视频中提取触觉先验,促进广泛适用的触觉感知机器人学习。
cs.CV / 49 / 2608.15061

Do Visual Grounding Decoders Need Feed-Forward Networks? A Controlled Study over Frozen Vision-Language Features

视觉定位解码器是否需要前馈网络?基于冻结的视觉-语言特征的控制研究
Tomar, Tarun
Abstract
Do feed-forward networks (FFNs) in visual grounding decoders add essential computation once a pretrained vision-language model has already encoded image and language context? We compare a four-block attention-only decoder (A4), a matched four-block attention-plus-FFN decoder (S4), and an eight-block attention-only parameter control (A8) over frozen VLM features. A4 matches or slightly exceeds S4 on RefCOCOg and Ref-Adv-s. FineCops-Ref reveals a small A4 deficit of 0.52 percentage points at [email protected] (95% CI [0.12, 0.95] in favor of S4), but A8 recovers it and finishes 0.26 points above S4. Official FineCops levels do not show a monotonic increase in the gap. A4 reduces trainable decoder parameters by 44.4% and cached-decoder latency by 10.1%, although end-to-end latency remains backbone-dominated. These results concern the trainable grounding decoder, not a complete attention-only VLM.
Chinese Translation
在预训练的视觉-语言模型已经编码了图像和语言上下文的情况下,视觉定位解码器中的前馈网络(FFNs)是否增加了必要的计算?我们比较了一个四块仅注意力解码器(A4)、一个匹配的四块注意力加前馈网络解码器(S4)以及一个八块仅注意力参数控制(A8)在冻结的视觉-语言模型特征上的表现。A4 在 RefCOCOg 和 Ref-Adv-s 上的表现与 S4 相当或略有超出。FineCops-Ref 显示 A4 在 [email protected] 上有 0.52 个百分点的微小劣势(95% 置信区间 [0.12, 0.95] 偏向 S4),但 A8 恢复了这一劣势,并以 0.26 个百分点超出 S4。官方 FineCops 水平并未显示出差距的单调增加。A4 将可训练解码器参数减少了 44.4%,并将缓存解码器延迟降低了 10.1%,尽管端到端延迟仍然以主干为主。这些结果涉及可训练的定位解码器,而非完整的仅注意力视觉-语言模型。
cs.CV / 50 / 2608.15075

SA-GEM: Scale-Adaptive and Geospatial Evidence-Modulated Token Pruning for Efficient Remote Sensing Large Vision-Language Models

SA-GEM:用于高效遥感大规模视觉-语言模型的尺度自适应与地理空间证据调制的标记剪枝
Ma, Kexin, Xiao, Jing, Xing, Bowen, Liao, Liang, Lin, Chia-Wen
Abstract
RS-LVLMs have advanced multimodal understanding of Earth observation imagery, yet their performance is fundamentally constrained by high-resolution processing, as visual token counts grow quadratically with linear input resolution while important visual evidence is inherently sparse and increasingly diluted across the expanded sequence. Existing token pruning methods largely rely on scale-agnostic resolution policies and isolated importance cues, limiting task-aligned granularity adaptation and holistic evidence preservation. To address this, we present Scale-Adaptive and Geospatial Evidence-Modulated Token Pruning (SA-GEM), a plug-and-play framework that unifies task-adaptive token granularity allocation with holistic geospatial token importance modulation. Specifically, a lightweight router selects the resolution based on query-dependent token granularity, while a token importance modulator jointly models task relevance, spatial structure, and local redundancy to preserve holistic geospatial evidence. We show that higher resolution is not universally beneficial and, once sufficient granularity is reached, token quality matters more than token quantity. Experiments across various benchmarks demonstrate that SA-GEM achieves consistent gains in both accuracy and efficiency over existing pruning methods. On XLRS-Bench, it surpasses GeoLLaVA-8K by 2.3% in accuracy with a 2.4 times total inference speedup.
Chinese Translation
遥感大规模视觉-语言模型(RS-LVLMs)在地球观测图像的多模态理解方面取得了进展,但其性能在根本上受到高分辨率处理的限制,因为视觉标记数量随着线性输入分辨率的增加呈平方增长,而重要的视觉证据在扩展序列中本质上是稀疏的并且越来越被稀释。现有的标记剪枝方法主要依赖于与尺度无关的分辨率策略和孤立的重要性线索,这限制了与任务对齐的粒度适应和整体证据的保留。为了解决这个问题,我们提出了尺度自适应与地理空间证据调制的标记剪枝(SA-GEM),这是一个即插即用的框架,统一了任务自适应的标记粒度分配与整体地理空间标记重要性调制。具体而言,一个轻量级路由器根据查询依赖的标记粒度选择分辨率,而标记重要性调制器则联合建模任务相关性、空间结构和局部冗余,以保留整体地理空间证据。我们表明,更高的分辨率并不总是有利的,一旦达到足够的粒度,标记质量比标记数量更为重要。在各种基准测试中的实验表明,SA-GEM在准确性和效率上均优于现有的剪枝方法。在XLRS-Bench上,其准确性超过GeoLLaVA-8K 2.3%,并实现了2.4倍的总推理速度提升。
cs.CV / 51 / 2608.15090

Distribution-free false-alarm calibration and chance-corrected spatial evaluation for industrial anomaly detection

无分布假警报校准与机会修正空间评估在工业异常检测中的应用
Deng, Jie
Abstract
Studies of industrial visual inspection commonly report the area under the receiver operating characteristic curve (AUROC) and the overlap between anomaly maps and defect masks. Neither measure specifies the false-alarm rate at a selected threshold, while recurrent defect locations and mask geometry can inflate overlap. We combine a distribution-free upper tolerance threshold with a paired-minus-crossed spatial test. This test compares each detector's score-contributing locations with the matched defect mask and with masks from other images; the difference in rates defines spatial-evidence lift relative to the empirical chance-overlap rate. We evaluate three detectors on 120 point-defect images from three ISP-AD modalities and three fixed data splits. Of 378 alarms, 230 overlap the matched mask. Paired and crossed rates are nevertheless similar in eight of nine detector--modality cells; only DINOv2--ASM has a positive 95\% bootstrap lower bound (lift 0.259, 95\% interval 0.159--0.347). On the independent Magnetic Tile Defect dataset, the same analysis gives lifts of 0.203 (0.169--0.236) for Wide ResNet-50 (WRN50) patch memory and 0.231 (0.202--0.262) for Vision Transformer B/16 (ViT-B/16) patch memory, with one-sided permutation $p=10^{-5}$ for both. When crossed masks are restricted to the same defect class, the lifts remain 0.185 and 0.210. Exact sample planning shows that, with 150 calibration normals, a 95\%-confidence distribution-free claim is supported only for target false-positive rates of 1.98\% or higher; a 1\% target requires at least 299 normals. The results support reporting operating-point performance and chance-corrected spatial evidence alongside AUROC and raw mask overlap.
Chinese Translation
工业视觉检测的研究通常报告接收者操作特征曲线下面积(AUROC)以及异常图与缺陷掩膜之间的重叠。然而,这两种度量都未能在选定阈值下具体说明假警报率,而重复的缺陷位置和掩膜几何形状可能会夸大重叠。我们结合了无分布的上限容忍阈值与配对减交叉空间测试。该测试比较每个检测器的得分贡献位置与匹配的缺陷掩膜以及来自其他图像的掩膜;率的差异定义了相对于经验机会重叠率的空间证据提升。我们在来自三种ISP-AD模态和三种固定数据拆分的120个点缺陷图像上评估了三种检测器。在378个警报中,230个与匹配的掩膜重叠。然而,在九个检测器-模态单元中,配对和交叉率在八个单元中仍然相似;只有DINOv2-ASM具有正的95\%自助法下限(提升0.259,95\\%区间0.159--0.347)。在独立的磁性瓷砖缺陷数据集上,同样的分析为Wide ResNet-50(WRN50)补丁记忆提供了0.203(0.169--0.236)的提升,为Vision Transformer B/16(ViT-B/16)补丁记忆提供了0.231(0.202--0.262)的提升,且两者的一侧置换$p=10^{-5}$。当交叉掩膜限制在同一缺陷类别时,提升仍为0.185和0.210。精确的样本规划显示,使用150个校准正常样本,仅在目标假阳性率为1.98\%或更高时,支持95\\%置信的无分布声明;而1\\%的目标至少需要299个正常样本。结果支持在报告操作点性能和机会修正空间证据时,结合AUROC和原始掩膜重叠。
cs.CV / 52 / 2608.15096

MODAL: Multi-Modal Object Re-ID via Model-Driven Sparse Decoupling and Text-Image Differential Filtering

MODAL:通过模型驱动的稀疏解耦和文本-图像差异过滤的多模态物体重识别
Huang, Chengbo, Huang, Jun-Jie, Lan, Long, Liu, Tianrui, Li, Xueqiong, Peng, Yuanxi, Liu, Xinwang, Wang, Meng
Abstract
Multi-modal object re-identification (Re-ID) aims to facilitate cross-camera object retrieval in complex environments by leveraging complementary information from visual (e.g., RGB, NIR, TIR) and textual modalities. However, existing approaches often lack principled feature disentanglement and coherent multi-modal integration, leading to entangled representations that introduce cross-modal conflicts, obscure discriminative cues, and suffer distribution shift under modality-missing conditions. To tackle these challenges, we propose MODAL, a novel multi-modal object re-identification framework, grounded in coupled sparse coding theory and differential suppression principles. A core component of MODAL is a Multi-modal Feature Sparse Decoupling module, developed in a model-driven deep unrolling manner based on multi-modal coupled sparse coding. It explicitly decomposes multi-modal features into uni-modal specific, bi-modal and tri-modal shared representations, thereby achieving more transparent and effective feature disentanglement. Benefiting from the principled feature disentanglement, MODAL naturally mitigates performance degradation in incomplete-modality scenarios via a Modality-Aware Subspace Activation that selectively activates only the consistently shared subspaces. Moreover, we propose a Text-Image Differential Filtering module that leverages coarse-grained textual semantics to adaptively suppress task-irrelevant responses in the decoupled visual representations, thereby enhancing discriminative information. Extensive experiments on four datasets demonstrate that MODAL achieves state-of-the-art performance with superior transparency.
Chinese Translation
多模态物体重识别(Re-ID)旨在通过利用视觉(如RGB、NIR、TIR)和文本模态的互补信息,促进复杂环境中的跨摄像头物体检索。然而,现有方法往往缺乏原则性的特征解耦和一致的多模态集成,导致纠缠的表示引入跨模态冲突,模糊区分线索,并在缺失模态的情况下遭受分布偏移。为了解决这些挑战,我们提出了MODAL,一个新颖的多模态物体重识别框架,基于耦合稀疏编码理论和差异抑制原则。MODAL的核心组件是一个多模态特征稀疏解耦模块,该模块采用基于多模态耦合稀疏编码的模型驱动深度展开方法开发。它明确地将多模态特征分解为单模态特定、双模态和三模态共享表示,从而实现更透明和有效的特征解耦。得益于原则性的特征解耦,MODAL自然减轻了在不完整模态场景下的性能下降,通过模态感知子空间激活选择性地激活仅一致共享的子空间。此外,我们提出了一个文本-图像差异过滤模块,利用粗粒度文本语义自适应抑制解耦视觉表示中的任务无关响应,从而增强区分信息。在四个数据集上的大量实验表明,MODAL在透明性方面实现了最先进的性能。
cs.CV / 53 / 2608.15104

ProjFormer: Point Cloud Completion via Geometric-Projective Transformer and Cross-Modal Semantic Constraints

ProjFormer:通过几何投影变换器和跨模态语义约束进行点云补全
Liu, Sheng, Wang, Meng, Li, Ruihui, Pi, Huilong, Tang, Zhuo, Li, Kenli
Abstract
Point cloud completion is inherently ill-posed due to severe sparsity and ambiguity in partial observations. Existing multi-view methods alleviate this by incorporating 2D semantics, but often rely on learned attention and fixed fusion, which lack geometric consistency and adaptability. We propose ProjFormer, a cross-modal framework that enforces geometry-consistent 2D-3D interaction through explicit projection and adaptive feature routing. A Projective Guided View Attention module aligns 3D points with multi-view features via deterministic projection, enabling efficient and geometrically consistent aggregation. Building on this, a geometry-aware routing network performs point-wise adaptive fusion of structural and observation-driven features for progressive refinement. Experiments show that, under a lightweight design, ProjFormer delivers competitive performance with improved structural completeness.
Chinese Translation
点云补全由于部分观测中的严重稀疏性和模糊性而本质上是一个病态问题。现有的多视角方法通过结合二维语义来缓解这一问题,但通常依赖于学习的注意力和固定的融合,这缺乏几何一致性和适应性。我们提出了ProjFormer,一个跨模态框架,通过显式投影和自适应特征路由来强制执行几何一致的二维-三维交互。投影引导视图注意力模块通过确定性投影将三维点与多视角特征对齐,从而实现高效且几何一致的聚合。在此基础上,几何感知路由网络执行结构特征和基于观测的特征的逐点自适应融合,以实现逐步细化。实验表明,在轻量化设计下,ProjFormer在结构完整性方面表现出竞争力。
cs.CV / 54 / 2608.15110

CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation

CETalk:基于音频驱动的连续情感-唤醒控制用于3D说话头生成
Jia, Peng, Dai, Li, Xiao, Zhen, Liu, Xueliang, Li, Jia
Abstract
Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing methods often rely on discrete emotion categories, which fail to capture the continuous evolution of affect. They also overlook the temporal frequency mismatch between audio articulation and emotional expression. In this paper, we propose CETalk, an audio-driven 3D facial animation framework conditioned on continuous Valence--Arousal (VA) representations for fine-grained emotion control. CETalk predicts a sequence of FLAME parameters through three key components: a Dynamic Emotion Modulation Module that adaptively scales emotional intensity using audio-derived cues; a Multi-Scale Temporal Modeling mechanism that employs parallel branches to decouple high-frequency articulatory movements from low-frequency emotional dynamics; and a Dynamic Fusion Mechanism that integrates these multi-scale features via an adaptive gating network. To support training and evaluation, we construct 3D-VA-MEAD, a large-scale dataset with automatically estimated VA annotations and reconstructed 3D facial motions. Extensive experiments demonstrate that CETalk outperforms state-of-the-art methods in both lip-sync accuracy and emotional expressiveness, while enabling smooth and controllable emotion transitions.
Chinese Translation
情感3D说话头生成旨在合成具有表现力的面部动画,并实现准确的唇部同步。然而,现有方法通常依赖于离散的情感类别,这无法捕捉情感的连续演变。同时,它们也忽视了音频发音与情感表达之间的时间频率不匹配。在本文中,我们提出了CETalk,一个基于音频驱动的3D面部动画框架,依赖于连续的情感-唤醒(Valence-Arousal, VA)表示,以实现细粒度的情感控制。CETalk通过三个关键组件预测一系列FLAME参数:一个动态情感调制模块,该模块利用音频衍生线索自适应地调整情感强度;一个多尺度时间建模机制,该机制采用并行分支将高频发音运动与低频情感动态解耦;以及一个动态融合机制,通过自适应门控网络整合这些多尺度特征。为了支持训练和评估,我们构建了3D-VA-MEAD,这是一个具有自动估计的VA注释和重建3D面部运动的大规模数据集。大量实验表明,CETalk在唇部同步准确性和情感表现力方面优于最先进的方法,同时实现了平滑和可控的情感过渡。
cs.CV / 55 / 2608.15113

Fast Test-Time Refinement for Robust Learned Image Compression

快速测试时精炼的鲁棒学习图像压缩
Liang, Jiaming, Pun, Chi-Man, Lin, Weisi
Abstract
Learned image compression (LIC) has demonstrated remarkable rate-distortion (RD) performance in benign settings. However, the high representational capacity endowed by deep neural networks (DNNs) comes at the expense of increased adversarial vulnerability. This hinders their adoption as trusted standardized codecs. Recent work has sketched test-time refinement (TTR) as a defense in gray-box scenarios, despite its original purpose of improving benign RD performance. Unfortunately, extensive iterations of TTR incur prohibitive overhead, while the robustness mechanism lacks theoretical understanding. Moreover, TTR has not been evaluated in white-box settings or against attacks beyond $\ell_2$-bounded rate and untargeted distortion objectives. To bridge these gaps, we present a systematic study. Our study reveals an Asymmetric Adversarial Trajectory (AAT) property in LIC systems: transitioning from adversarial to benign regions is significantly easier than the reverse process, where adversarial examples can often be roughly recovered within only 1-2 steps. We provide a two-dimensional Tube Model to explain this phenomenon. Based on AAT, we propose a Fast Test-Time Refinement (FTTR) framework for practical and robust LIC systems. We establish that the robustness arises from the contraction of adversarial regions induced by the Input-as-Label property of LIC systems, rather than from obfuscated gradients. Extensive evaluations with diverse strong adaptive attacks across multiple LIC systems demonstrate the promise of the proposed FTTR framework. The code is available at https://github.com/chinaliangjiaming/FTTR.git.
Chinese Translation
学习图像压缩(LIC)在良性环境中展现了显著的率失真(RD)性能。然而,深度神经网络(DNN)所赋予的高表示能力以增加对抗性脆弱性为代价,这阻碍了它们作为可信标准编解码器的应用。近期的研究将测试时精炼(TTR)勾勒为灰盒场景下的一种防御手段,尽管其最初目的是提高良性RD性能。不幸的是,TTR的广泛迭代会产生巨大的开销,而其鲁棒性机制缺乏理论理解。此外,TTR尚未在白盒环境中进行评估,也未针对超出$ ext{l}_2$-界限的速率和非目标失真目标的攻击进行测试。为了解决这些问题,我们进行了系统研究。我们的研究揭示了LIC系统中的一种不对称对抗轨迹(AAT)特性:从对抗区域过渡到良性区域显著容易,而反向过程则相对困难,其中对抗样本通常可以在仅1-2步内粗略恢复。我们提供了一个二维管道模型来解释这一现象。基于AAT,我们提出了一种快速测试时精炼(FTTR)框架,以实现实用且鲁棒的LIC系统。我们确定鲁棒性源于LIC系统的输入作为标签(Input-as-Label)特性所引起的对抗区域的收缩,而非模糊梯度。通过对多个LIC系统进行广泛评估,使用多种强适应性攻击,展示了所提FTTR框架的前景。代码可在https://github.com/chinaliangjiaming/FTTR.git获取。
cs.CV / 56 / 2608.15115

Perspective-Invariant Attack with Enhanced Transferability of Adversarial Examples

具有增强可转移性的视角不变攻击
Liang, Kaisheng, Cao, Yiming, Xiao, Bin
Abstract
Adversarial examples generated on a surrogate deep neural network (DNN) can often successfully fool other black-box DNN models. This cross-model transferability poses serious security threats to DNNs in practical applications. Input transformation techniques are widely used to enhance adversarial transferability by increasing the diversity of input images. However, existing methods primarily rely on local operations with limited degrees of freedom (DOF), such as block-wise shuffling and resizing, overlooking global perspective transformations that naturally arise from viewpoint changes. In this work, we propose a Perspective-Invariant Attack (PIA), which introduces a multi-DOF vertex sampling strategy that systematically covers the perspective transformation hierarchy from 2-DOF translation to 8-DOF projective mapping. By generating geometrically diverse input variations, PIA effectively reduces overfitting of adversarial perturbations to the surrogate model, thereby improving adversarial transferability. We further propose PIA-Mix, a generic extension that maintains a complementary transformation pool and efficiently combines our perspective transformation with auxiliary methods for improved transferability. Extensive experiments involving various DNN architectures, advanced defense mechanisms, and multimodal large language models (LLMs) demonstrate that PIA and PIA-Mix outperform state-of-the-art transfer-based attacks.
Chinese Translation
在替代深度神经网络(DNN)上生成的对抗样本通常能够成功欺骗其他黑箱DNN模型。这种跨模型的可转移性对DNN在实际应用中的安全性构成了严重威胁。输入变换技术被广泛用于通过增加输入图像的多样性来增强对抗可转移性。然而,现有方法主要依赖于局部操作,具有有限的自由度(DOF),例如块状洗牌和调整大小,忽视了因视角变化而自然产生的全局视角变换。在本研究中,我们提出了一种视角不变攻击(Perspective-Invariant Attack, PIA),引入了一种多自由度顶点采样策略,系统地覆盖了从2自由度平移到8自由度投影映射的视角变换层次。通过生成几何多样的输入变体,PIA有效减少了对抗扰动对替代模型的过拟合,从而提高了对抗可转移性。我们进一步提出了PIA-Mix,这是一种通用扩展,维护了一个互补的变换池,并有效地将我们的视角变换与辅助方法结合,以提高可转移性。涉及各种DNN架构、先进防御机制和多模态大型语言模型(LLMs)的广泛实验表明,PIA和PIA-Mix在可转移攻击方面优于现有的最先进方法。
cs.CV / 57 / 2608.15141

HOIMask: Towards Generative Masked Modeling for Human Object Interaction Generation

HOIMask:面向人类与物体交互生成的生成性掩蔽建模
Ji, Yihong, Zhang, Jinsong, Hu, He, Xu, Hongbo
Abstract
Diffusion-based methods have dominated the HOI generation, as they enable critical contact fusions or signals to guide the diffusion process. However, they often result in high artifacts and unstable interaction quality due to error accumulation during iterative denoising. In this work, we propose HOIMask, the first generative masked framework for modeling HOI motion in discrete space. HOIMask first encodes both motion sequences and contact-aware signals into discrete 2D human and object token maps via HOI Vector Quantization (VQ), preserving fine-grained spatial-temporal structure beyond conventional 1D representations. On this basis, a generative masked modeling framework is employed to jointly capture human-object interaction dynamics, leveraging a transformer architecture designed to model complex spatial-temporal and interaction dependencies. To generate more coherent and physically plausible motions, we further introduce a novel contact-aware reconstruction guidance in discrete space during inference, which fuses contact signals to optimize HOI tokens that forces the generated motion with higher spatio-temporal consistency. With craftily designed motion interaction tokens, dedicated architecture and guidance strategy, HOIMask outperforms state-of-the-art diffusion-based methods, generating more realistic and semantically aligned HOI motions. Please refer to https://jyhflash.github.io/HOIMask/ for more results.
Chinese Translation
基于扩散的方法在HOI(人类与物体交互)生成中占据主导地位,因为它们能够通过关键接触融合或信号来引导扩散过程。然而,由于在迭代去噪过程中误差的累积,这些方法往往导致高伪影和不稳定的交互质量。在本研究中,我们提出了HOIMask,这是首个用于在离散空间中建模HOI运动的生成性掩蔽框架。HOIMask首先通过HOI向量量化(VQ)将运动序列和接触感知信号编码为离散的二维人类和物体标记图,保留了超越传统一维表示的细粒度时空结构。在此基础上,采用生成性掩蔽建模框架共同捕捉人类与物体交互的动态,利用设计用于建模复杂时空和交互依赖关系的变换器架构。为了生成更连贯和物理上合理的运动,我们在推理过程中进一步引入了一种新颖的接触感知重建指导,在离散空间中融合接触信号,以优化HOI标记,从而强制生成的运动具有更高的时空一致性。通过巧妙设计的运动交互标记、专用架构和指导策略,HOIMask超越了最先进的基于扩散的方法,生成了更真实且语义一致的HOI运动。更多结果请参见 https://jyhflash.github.io/HOIMask/
cs.CV / 58 / 2608.15160

A Unified Backbone--Expert Framework with Relation-Token and Residual--Classifier Interfaces for Automatic Modulation Recognition

统一的主干-专家框架:带有关系标记和残差分类器接口的自动调制识别
Deng, Zhixiang, Li, Houbiao, Cui, Zongyong
Abstract
Automatic modulation recognition (AMR) faces distinct representation bottlenecks under varying observation lengths, where a single model architecture often fails to excel. To address this, we propose a unified backbone-expert framework with a common convolutional state-space backbone and two specialized interfaces. For short sequences, we inject explicit lag-aware complex-plane descriptors as relation tokens before encoding to compensate for information loss. For long sequences, we design a gated multi-scale residual refinement module to correct the feature map, combined with a fixed-averaging classifier collaboration to harness complementary evidence. Our framework achieves overall average accuracies of 67.28 \pm 0.14% on RML2016.10b and 87.19 \pm 0.77% on HisarMod2019 (mean \pm sample standard deviation over three runs), respectively. The framework's efficacy is further validated through three-seed ablations, native-length cross-configuration tests, and controlled window studies, confirming the benefit of expert-interface decoupling over one-size-fits-all architectures.
Chinese Translation
自动调制识别(AMR)在不同观测长度下面临明显的表示瓶颈,单一模型架构往往难以表现出色。为了解决这一问题,我们提出了一种统一的主干-专家框架,该框架具有一个共同的卷积状态空间主干和两个专门的接口。对于短序列,我们在编码之前注入显式的滞后感知复平面描述符作为关系标记,以补偿信息损失。对于长序列,我们设计了一个门控多尺度残差细化模块来校正特征图,并结合固定平均分类器协作以利用互补证据。我们的框架在RML2016.10b数据集上实现了67.28 b 0.14%的总体平均准确率,在HisarMod2019数据集上实现了87.19 b 0.77%的准确率(均值 b 样本标准偏差,基于三次运行)。通过三种种子消融实验、原生长度交叉配置测试和控制窗口研究进一步验证了框架的有效性,确认了专家接口解耦相较于一刀切架构的优势。
cs.CV / 59 / 2608.15163

From "What-If" to "What-Is": Counterfactual Thinking-Inspired Semantic Alignment for Visual Brain Decoding

从“假设”到“现实”:受反事实思维启发的视觉脑解码语义对齐
Yan, Kaitao, Liu, Chi, Zhu, Congcong, Chen, Huajie, Wu, Gengshen, Wang, Minghao, Han, Xiaotong, Zhu, Tianqing
Abstract
Visual brain decoding reconstructs visual content perceived by a person from neural measurements such as fMRI, providing a computational approach to studying how visual information is represented in the brain. Recent multimodal representations and diffusion priors have improved reconstruction realism. However, visually plausible reconstructions may contain incorrect objects, attributes, or relations because a strong generative prior can complete content not sufficiently specified by the decoded representation. Conventional reconstruction metrics mainly assess the final image and may therefore obscure such semantic errors. We propose ConceptAlign, a counterfactual semantic alignment framework for visual brain decoding. ConceptAlign pools decoded visual tokens and projects them into a frozen text-embedding space, aligning the representation with the ground-truth caption while separating it from scene-preserving near-miss alternatives. Generated offline by an LLM, these alternatives modify one critical object, attribute, or relation while retaining the scene. A margin-based objective learns fine-grained semantic boundaries between the observed stimulus and plausible but incorrect interpretations without requiring LLM calls during inference. We introduce a systematic three-level semantic evaluation framework covering foundational discriminability, counterfactual description discrimination, and representational geometry. Experiments on the Natural Scenes Dataset show that ConceptAlign improves reconstruction measures, counterfactual semantic discrimination, and representational alignment over the MindEye2 backbone. Matched negative-source ablations, independent LLM and human-written alternatives, and human evaluation support the effectiveness and robustness of the supervision, with favorable patterns in fine-grained conflicts, limited-data decoding, and cross-subject structure.
Chinese Translation
视觉脑解码通过神经测量(如功能性磁共振成像 fMRI)重建个体所感知的视觉内容,为研究视觉信息在大脑中的表示提供了一种计算方法。近期的多模态表示和扩散先验提高了重建的真实感。然而,视觉上合理的重建可能包含不正确的物体、属性或关系,因为强生成先验能够填补解码表示中未充分指定的内容。传统的重建度量主要评估最终图像,因此可能掩盖此类语义错误。我们提出了 ConceptAlign,一个用于视觉脑解码的反事实语义对齐框架。ConceptAlign 汇集解码的视觉标记并将其投影到一个冻结的文本嵌入空间中,将表示与真实的标题对齐,同时将其与保留场景的近似替代品分离。这些替代品由大型语言模型(LLM)离线生成,修改一个关键物体、属性或关系,同时保留场景。基于边际的目标学习观察到的刺激与合理但不正确的解释之间的细粒度语义边界,而无需在推理过程中调用 LLM。我们引入了一个系统的三层语义评估框架,涵盖基础可区分性、反事实描述区分和表示几何。对自然场景数据集的实验表明,ConceptAlign 在重建度量、反事实语义区分和表示对齐方面优于 MindEye2 主干。匹配的负源消融实验、独立的 LLM 和人类撰写的替代品以及人类评估支持了监督的有效性和鲁棒性,在细粒度冲突、有限数据解码和跨主体结构中表现出良好的模式。
cs.CV / 60 / 2608.15195

Beyond Natural-Image Foundation Models: Benchmarking Satellite Pretraining for Ophthalmic Image Analysis

超越自然图像基础模型:卫星预训练在眼科图像分析中的基准测试
Budimir, Lovre Antonio, Gong, Mingya Alexa, Quinney, Alyssa Foong, Matovinović, Ivana, Zhou, Yukun, Keane, Pearse A., Lončarić, Sven, Šarunić, Marinko V.
Abstract
Vision Foundation Models (VFMs) have emerged as a promising approach in medical imaging, producing broadly applicable systems that can be efficiently adapted across diverse imaging modalities, anatomical regions, and clinical tasks. However, VFMs require extensive training data, and their progress in medical image analysis is constrained by limited data availability, privacy concerns, and high development costs. To alleviate these constraints, medical VFMs (MedVFMs) are often built upon weights from generalist models pretrained on vast amounts of publicly available natural images, introducing a substantial distribution shift for medical task adaptation. To address this, we propose satellite imagery as a novel pretraining domain for MedVFM development and benchmarking, motivated by its closer visual alignment with medical data and its freedom from the privacy constraints that limit medical datasets. Across multiple ophthalmic imaging modalities, we compare DINOv3-SAT493m pretrained on 493 million satellite images against DINOv3-LVD1689m pretrained on 1.7 billion natural images, together with two medical specialist baselines: DINOv3-RETFound and MAE-RETFound. Our experiments show that satellite imagery is a stronger pretraining source than natural images for ophthalmic tasks, particularly on en face vascular-rich modalities. On several tasks, satellite pretraining matches or exceeds the medical specialists on high-resolution en face inputs, despite using no medical data.
Chinese Translation
视觉基础模型(VFMs)已成为医学影像领域一种有前景的方法,能够高效地适应多种影像模态、解剖区域和临床任务,产生广泛适用的系统。然而,VFMs 需要大量的训练数据,而在医学图像分析中的进展受到数据可用性有限、隐私问题和高开发成本的制约。为了缓解这些限制,医学 VFMs(MedVFMs)通常基于在大量公开可用自然图像上预训练的通用模型的权重,这为医学任务的适应引入了显著的分布偏移。为了解决这一问题,我们提出将卫星图像作为 MedVFM 开发和基准测试的新颖预训练领域,理由是其在视觉上与医学数据更为接近,并且不受限制医学数据集的隐私约束。在多个眼科影像模态中,我们将预训练于 4.93 亿张卫星图像的 DINOv3-SAT493m 与预训练于 17 亿张自然图像的 DINOv3-LVD1689m 进行比较,并引入两个医学专业基线:DINOv3-RETFound 和 MAE-RETFound。我们的实验表明,卫星图像在眼科任务中是比自然图像更强的预训练来源,尤其是在丰富血管的面像模态上。在多个任务中,卫星预训练在高分辨率面像输入上与医学专家的表现相当或更优,尽管未使用任何医学数据。
cs.CV / 61 / 2608.15196

Anchor-Regularized Adaptation for Generalizable AI-Generated Image Detection with DINOv3

基于锚点正则化的适应方法用于通用化AI生成图像检测,结合DINOv3
Choi, Hyeongjun, Lee, Juhun, Cozzolino, Davide, Verdoliva, Luisa, Woo, Simon S.
Abstract
Recent works in AI-generated image detection have shown that careful training data alignment can improve generalization by removing spurious correlations. However, linear probes on frozen DINOv3 representations achieve remarkably strong performance even when trained on misaligned datasets. Motivated by this result, we analyze the underlying rationale and the limits of this generalization. We find that frozen DINOv3 performs well because its decisions rely on features that faithfully represent the space of authentic images. At the same time, its final layer is less effective at capturing the subtle pixel-artifact cues that can be emphasized by aligned training data. We further observe that naively mixing aligned and misaligned data during adaptation improves sensitivity to such cues but at the cost of distorting the pre-trained representation, limiting generalization. To address this issue, we propose Anchor-Regularized Adaptation (ARA). We apply Low-Rank Adaptation to capture pixel-level artifacts while leveraging a frozen anchor classifier to avoid deviations from the original representation structure. This allows the model to exploit pixel-artifact cues without sacrificing generalization. Our method achieves state-of-the-art performance on nine diverse and challenging benchmarks, indicating that ARA enables complementary supervision from misaligned and aligned data for more effective detection.
Chinese Translation
近期在AI生成图像检测领域的研究表明,精心的训练数据对齐可以通过消除虚假相关性来提高泛化能力。然而,即使在不对齐的数据集上进行训练,冻结的DINOv3表示在使用线性探测器时也能取得显著的性能。受到这一结果的启发,我们分析了这种泛化的基本原理及其局限性。我们发现,冻结的DINOv3表现良好是因为其决策依赖于忠实代表真实图像空间的特征。同时,其最终层在捕捉由对齐训练数据强调的微妙像素伪影线索方面效果较差。我们进一步观察到,在适应过程中天真地混合对齐和不对齐的数据可以提高对这些线索的敏感性,但代价是扭曲了预训练表示,从而限制了泛化能力。为了解决这个问题,我们提出了锚点正则化适应方法(Anchor-Regularized Adaptation, ARA)。我们应用低秩适应(Low-Rank Adaptation)来捕捉像素级伪影,同时利用冻结的锚点分类器以避免偏离原始表示结构。这使得模型能够利用像素伪影线索而不牺牲泛化能力。我们的方法在九个多样且具有挑战性的基准测试中实现了最先进的性能,表明ARA能够有效地利用来自不对齐和对齐数据的互补监督,从而提高检测效果。
cs.CV / 62 / 2608.15211

TERRA: A Hierarchical Parallel Training and Memory Orchestration Framework for High-Resolution AI-based Earth Modeling

TERRA:一种用于高分辨率基于人工智能的地球建模的分层并行训练和内存调度框架
Wu, Ruohan, Zhu, Ziqi, Zhao, Yang, Tang, Jiarui, Cui, Yingzhe, Chen, Junshi, Jing, Zhao, Shi, Jun, An, Hong
Abstract
Training high-resolution AI-based Earth forecasting models is memory-intensive. Window-based Swin Transformers reduce the quadratic cost of global attention, but existing distributed systems such as AERIS primarily target pixel-level models and do not jointly support convolutional sampling modules and shifted-window execution. Long-lead rollout finetuning further increases activation memory. To address these challenges, we present TERRA, a hierarchical parallel training framework for high-resolution Earth forecasting. TERRA introduces Sampling-Aware Window, Sequence, and Tensor Parallelism (SAWSTP), which preserves spatially contiguous layouts for sampling modules and routes tokens into topology-aware ragged window layouts for Transformer execution. For long-lead finetuning, Memory Orchestration (MO) provides rollout-aware checkpoint planning and combines input buffering with budget-constrained activation offloading. Experiments on the $1/12^\circ$ GLORYS-based Wenhai workload show that TERRA supports models with up to 11.4B parameters on 96 H200 GPUs and sustains up to $39.76$ PFLOPS, achieving $65.0\%$ strong-scaling and $94.1\%$ weak-scaling efficiency. Compared with checkpoint-only policies, MO further reduces peak allocated GPU memory by $32.2\%$--$51.8\%$ with at most $20.0\%$ step-time overhead, which makes finetuning with smaller patch sizes and longer rollouts feasible for improved forecasting accuracy.
Chinese Translation
训练高分辨率基于人工智能的地球预测模型需要大量内存。基于窗口的Swin Transformers减少了全局注意力的平方成本,但现有的分布式系统如AERIS主要针对像素级模型,并未共同支持卷积采样模块和移动窗口执行。长时间步的微调进一步增加了激活内存。为了解决这些挑战,我们提出了TERRA,一种用于高分辨率地球预测的分层并行训练框架。TERRA引入了采样感知窗口、序列和张量并行性(Sampling-Aware Window, Sequence, and Tensor Parallelism,SAWSTP),该方法为采样模块保留了空间连续布局,并将令牌路由到拓扑感知的稀疏窗口布局中以进行Transformer执行。对于长时间步的微调,内存调度(Memory Orchestration,MO)提供了基于回滚的检查点规划,并将输入缓冲与预算约束的激活卸载相结合。在$1/12^ ext{°}$ GLORYS基础的Wenhai工作负载上的实验表明,TERRA支持在96个H200 GPU上具有高达11.4B参数的模型,并维持高达$39.76$ PFLOPS的性能,达到了$65.0 ext{%}$的强扩展和$94.1 ext{%}$的弱扩展效率。与仅使用检查点的策略相比,MO进一步将峰值分配的GPU内存减少了$32.2 ext{%}$至$51.8 ext{%}$,并且最多只增加$20.0 ext{%}$的步骤时间开销,这使得使用更小的补丁大小和更长的回滚进行微调成为可能,从而提高预测准确性。
cs.CV / 63 / 2608.15213

DCA-MoE: Spatially Adaptive Cross-Layer Fusion and Density-Routed Experts for Crowd Counting

DCA-MoE:空间自适应跨层融合与密度引导专家用于人群计数
Wang, Hao
Abstract
Crowd counting must recover reliable local density under severe variations in perspective, head scale, occlusion, and background clutter. Although modern counting objectives provide strong spatial supervision, many multi-level decoders still use spatially invariant feature fusion and apply one receptive-field pattern to every location. We propose DCA-MoE, a framework that makes both decisions content dependent while retaining a frozen DINOv3 encoder. Spatially Adaptive Layer Fusion (SALF) predicts position-wise weights over four aligned backbone features, and Density-Routed Multi-Receptive-Field Experts (DR-MoE) assigns each location a soft mixture of local, mid-range, and large-context residual experts. An EBC-style head reconstructs block density, while DMCount supervision and an auxiliary routing-balance term train the decoder without updating the backbone. On the NWPU-Crowd validation split, the strongest paired configuration, based on DINOv3 ViT-L/16, obtains 31.7 MAE and 72.2 RMSE; the matched ViT-B/16 full model obtains a paired 32.2/75.9. Cross-dataset results remain mixed, and several component baselines currently report independently selected minima from a single seed. The evidence therefore supports the feasibility of spatially adaptive fusion and routing, while broader paired and multi-seed evaluation remains necessary for causal attribution.
Chinese Translation
人群计数必须在视角、头部尺度、遮挡和背景杂乱等严重变化下恢复可靠的局部密度。尽管现代计数目标提供了强大的空间监督,但许多多层解码器仍然使用空间不变的特征融合,并对每个位置应用相同的感受野模式。我们提出了DCA-MoE,一个框架,使得决策依赖于内容,同时保持冻结的DINOv3编码器。空间自适应层融合(SALF)预测四个对齐的主干特征上的位置权重,而密度引导多感受野专家(DR-MoE)为每个位置分配局部、中等范围和大上下文残差专家的软混合。EBC风格的头部重建块密度,而DMCount监督和辅助路由平衡项在不更新主干的情况下训练解码器。在NWPU-Crowd验证集上,基于DINOv3 ViT-L/16的最强配对配置获得31.7的平均绝对误差(MAE)和72.2的均方根误差(RMSE);匹配的ViT-B/16完整模型获得配对的32.2/75.9。跨数据集的结果仍然混合,几个组件基线目前报告从单个种子独立选择的最小值。因此,证据支持空间自适应融合和路由的可行性,而更广泛的配对和多种子评估仍然是因果归因所必需的。
cs.CV / 64 / 2608.15217

Self-Supervised Topologically Invariant Manifold Learning for Railway Image Quality Assessment

自监督拓扑不变流形学习用于铁路图像质量评估
Cui, Tingqiong, Yang, Yibu, Li, Yang, Fu, Jiahao, Luo, Xiaoliu, Wang, Xu, Wang, Mengzhu, Liu, Siyuan, Huang, Guanghui
Abstract
Existing blind image quality assessment (BIQA) methods typically rely on synthetic distortions and subjective annotations, limiting generalization in real-world domains. To address this, we propose a fully self-supervised BIQA framework based on topologically invariant manifold learning under boundary constraints, which constructs a stable quality reference without manual labels. The framework generates progressive background dilution scales via repeated random cropping around each target; exploiting the monotonic degradation of target information density across these scales, it establishes a self-constrained quality manifold. A linearized spatial moment projection eliminates geometric distortions from random cropping; then a monotonicity divergence filter prunes background-sensitive evaluators, isolating an elite pool \(\mathcal{M}_{\text{elite}}\). A robust M-estimator with a principal component stabilizer fuses the metrics into an asymptotically efficient pseudo-ground truth \(q_{\text{PGT}}\), contracting variance toward the Cram\'{e}r-Rao lower bound. Extensive evaluations demonstrate that the elite evaluator pool, distilled from 11 baseline metrics, secures superior zero-shot transferability across standard synthetic and wild benchmarks (CSIQ, LIVEC, LIVE-2). Concurrently, deployments on the CQU Railway Rolling Stock Surveillance Dataset (2,797 images) yield a manifold cosine similarity \(>0.999\) and a 100.0\% survival rate under industrial extreme stresses, robustly validating its cross-paradigm decoupling and topological resilience.
Chinese Translation
现有的盲图像质量评估(BIQA)方法通常依赖于合成失真和主观注释,这限制了其在现实世界领域的泛化能力。为了解决这个问题,我们提出了一种基于边界约束的拓扑不变流形学习的完全自监督BIQA框架,该框架在没有人工标签的情况下构建了一个稳定的质量参考。该框架通过围绕每个目标的重复随机裁剪生成渐进的背景稀释尺度;利用这些尺度上目标信息密度的单调退化,它建立了一个自约束的质量流形。线性化的空间矩投影消除了随机裁剪带来的几何失真;随后,单调性发散滤波器修剪了对背景敏感的评估器,孤立出一个精英池M_{ ext{elite}}。一个具有主成分稳定器的稳健M估计器将度量融合为渐近有效的伪真实值q_{ ext{PGT}},将方差收缩到Cramér-Rao下界。广泛的评估表明,从11个基线度量中提炼出的精英评估器池在标准合成和野外基准(CSIQ、LIVEC、LIVE-2)上实现了优越的零样本迁移能力。同时,在CQU铁路机车车辆监测数据集(2,797张图像)上的应用显示出流形余弦相似度大于0.999,并且在工业极端压力下的生存率达到100.0%,稳健地验证了其跨范式解耦和拓扑韧性。
cs.CV / 65 / 2608.15230

PersonaDrive: Controllable Trajectory Prediction with Multi-Dimensional Driving Personas

PersonaDrive:具有多维驾驶角色的可控轨迹预测
Lee, Chan, Yun, Kimin, Bae, Yuseok, Kim, Seong Tae, Kim, Jung Uk
Abstract
Although recent trajectory prediction and end-to-end autonomous driving methods improve robustness in urban environments, they still lack meaningful controllability. Existing benchmarks either provide no persona-conditioned annotations or support only a single urgency spectrum (i.e., emergency, normal, relaxed), which cannot distinguish personas that share the same urgency level but require different driving dynamics. To address this, we propose (i) the Persona-Conditioned Trajectory (PCT) dataset, which decomposes driving personas along two axes, Temporal Urgency and Ride Comfort, and combines three levels of each to form a grid of nine personas, each paired with natural-language descriptions and trajectories, and (ii) PersonaDrive, a framework that can learn driving personas from language and can generate persona-specific trajectories. PersonaDrive incorporates Persona-Conditioned Anchor Transform (PCAT), which hierarchically reshapes anchors along both axes, and Persona-Conditioned Multi-Modal Fusion (PCMF) for BEV-level persona fusion. Training is supervised by a Hierarchical Guide Loss enforcing axis-aligned physical orderings and an Axis-Decomposed Diversity Loss preventing diagonal mode collapse. Experimental results show that PersonaDrive consistently improves over the compared baselines across multi-dimensional scenarios. The code and PCT dataset are available at https://github.com/VisualAIKHU/PersonaDrive
Chinese Translation
尽管近期的轨迹预测和端到端自动驾驶方法在城市环境中提高了鲁棒性,但它们仍然缺乏有意义的可控性。现有基准要么没有提供角色条件的注释,要么仅支持单一的紧急程度谱(即紧急、正常、放松),这无法区分共享相同紧急程度但需要不同驾驶动态的角色。为了解决这一问题,我们提出了(i)角色条件轨迹(Persona-Conditioned Trajectory, PCT)数据集,该数据集沿两个轴线分解驾驶角色:时间紧急性和乘坐舒适度,并结合每个轴线的三个级别形成九个角色的网格,每个角色配有自然语言描述和轨迹;(ii)PersonaDrive,一个能够从语言中学习驾驶角色并生成角色特定轨迹的框架。PersonaDrive结合了角色条件锚点变换(Persona-Conditioned Anchor Transform, PCAT),该方法沿两个轴线分层重塑锚点,以及角色条件多模态融合(Persona-Conditioned Multi-Modal Fusion, PCMF)用于鸟瞰视角(BEV)级别的角色融合。训练通过层次引导损失(Hierarchical Guide Loss)进行监督,该损失强制执行轴对齐的物理顺序,并通过轴分解多样性损失(Axis-Decomposed Diversity Loss)防止对角模式崩溃。实验结果表明,PersonaDrive在多维场景中始终优于比较基线。代码和PCT数据集可在https://github.com/VisualAIKHU/PersonaDrive获取。
cs.CV / 66 / 2608.15238

UC-VLM: Consistency-Driven Learning for AI-Generated Image Detection with Vision-Language Large Models

UC-VLM:基于一致性驱动的AI生成图像检测学习方法,结合视觉-语言大模型
Tan, Lei, Li, Shuwei, Kankanhalli, Mohan, Tan, Robby T.
Abstract
Vision-Language Large Models (VLLMs) are promising for AI-generated image (AIGI) detection because they can produce both a prediction and a natural-language output. However, most existing VLLM-based detectors primarily fine-tune the language side while giving limited attention to low-level visual forensic cues. They also often depend on manually crafted prompts or human-annotated rationales, which limits scalability.We present UC-VLM, a unified multi-stage framework for AIGI detection that relies solely on binary supervision. UC-VLM first identifies effective instruction variants automatically. It then reuses the same binary label within a multi-stage training framework: (i) a visual discrimination objective that strengthens sensitivity to non-semantic forensic cues, and (ii) a label-conditioned generation objective that uses the binary label to supervise textual outputs. This design turns weak binary supervision into a shared supervision signal for both the visual pathway and the language output. Our key novelty is a unified multi-stage binary-supervised framework that consistently reuses the same authenticity labels for visual adaptation and label-conditioned text generation, while leveraging automatically optimized instructions to reduce prompt sensitivity without requiring human-written rationales or hand-crafted prompts.Experiments show that UC-VLM achieves 96.1% average accuracy on GenImage, exceeding the strongest prior result by 4.6%, and obtains 69.6% / 77.9% accuracy on Chameleon under ProGAN / SDV1.4 training, surpassing the best baseline by 11.2% / 15.3%, respectively.
Chinese Translation
视觉-语言大模型(VLLMs)在AI生成图像(AIGI)检测中展现出良好前景,因为它们能够同时生成预测和自然语言输出。然而,现有的大多数基于VLLM的检测器主要对语言部分进行微调,而对低层次的视觉取证线索关注有限。此外,它们通常依赖于手工制作的提示或人工标注的推理,这限制了其可扩展性。我们提出了UC-VLM,这是一个统一的多阶段框架,用于AIGI检测,仅依赖于二元监督。UC-VLM首先自动识别有效的指令变体。然后,它在一个多阶段训练框架中重用相同的二元标签:(i)一个视觉区分目标,增强对非语义取证线索的敏感性,以及(ii)一个标签条件生成目标,利用二元标签来监督文本输出。该设计将弱二元监督转化为视觉路径和语言输出的共享监督信号。我们的关键创新在于一个统一的多阶段二元监督框架,它一致地重用相同的真实性标签用于视觉适配和标签条件文本生成,同时利用自动优化的指令减少对提示的敏感性,而无需人工编写的推理或手工制作的提示。实验表明,UC-VLM在GenImage数据集上实现了96.1%的平均准确率,超过了之前最强结果4.6%;在Chameleon数据集上,在ProGAN / SDV1.4训练下分别获得69.6% / 77.9%的准确率,分别超越最佳基线11.2% / 15.3%。
cs.CV / 67 / 2608.15246

CG-GLORE: A Conjugate Gradient-Based Global-Local Regularization Network for Sparse-View CT Reconstruction

CG-GLORE:一种基于共轭梯度的全局-局部正则化网络用于稀视角CT重建
Le, Tran Xuan Hieu, Bui, Doanh C., Le, Vu Trung Duong, Pham, Hoai Luan, Nguyen, Khang, Nguyen, Mai K., Ho, Tu Bao, Nakashima, Yasuhiko
Abstract
Sparse-view computed tomography (CT) reduces radiation dose by acquiring fewer projection views, but the resulting inverse problem is highly ill-posed and often produces severe streak artifacts. Existing deep reconstruction methods have achieved promising performance, yet many rely on first-order updates or large regularization networks, which can be less effective in ill-conditioned settings. We propose \textbf{CG-GLORE}, a compact deep unrolling framework inspired by second-order optimization for sparse-view CT reconstruction. Each unrolled stage uses a CG-solved linear system based on a structured Hessian surrogate: it retains the physics-induced curvature of the data-fidelity term while using an identity approximation for the learned regularization term. Thus, the method is second-order-inspired rather than an exact Newton method for the full learned objective. To model image priors, we design a Global-Local Regularization Network (GLORE), which combines convolutional local feature extraction with a Long-Range Dependency Representation module based on sparse patchification and Nystr\"{o}m attention. This design captures anatomical details and non-local dependencies while maintaining practical complexity. Experiments on AAPM and DeepLesion under multiple sparse-view and noise settings show that CG-GLORE achieves strong quantitative performance, stable convergence, lower noise power, and improved visual fidelity compared with representative reconstruction methods.
Chinese Translation
稀视角计算机断层扫描(CT)通过获取更少的投影视角来减少辐射剂量,但由此产生的逆问题高度不适定,常常会产生严重的条纹伪影。现有的深度重建方法已取得了令人鼓舞的性能,但许多方法依赖于一阶更新或大型正则化网络,这在条件不良的情况下可能效果不佳。我们提出了 extbf{CG-GLORE},这是一种受二阶优化启发的紧凑深度展开框架,用于稀视角CT重建。每个展开阶段使用基于结构海森矩阵近似的共轭梯度(CG)求解线性系统:它保留了数据保真项的物理引起的曲率,同时对学习的正则化项使用了单位近似。因此,该方法受到二阶启发,而不是对完整学习目标的精确牛顿方法。为了建模图像先验,我们设计了一个全局-局部正则化网络(GLORE),它结合了卷积局部特征提取和基于稀疏补丁化及Nystr"{o}m注意力的长程依赖表示模块。该设计在保持实际复杂性的同时,捕捉了解剖细节和非局部依赖。在多个稀视角和噪声设置下对AAPM和DeepLesion的实验表明,与代表性的重建方法相比,CG-GLORE在定量性能、稳定收敛、较低噪声功率和改善视觉保真度方面均表现出色。
cs.CV / 68 / 2608.15251

Robust structure from motion for aerial-ground images via detector-free feature matching and multi-view track refinement

通过无检测器特征匹配和多视图轨迹优化实现的空中-地面图像鲁棒运动结构重建
Jiang, San, Wang, Hui, Zhang, Xing, Hu, Zhongwen, Wang, Zhijun, Wang, Ruisheng, Jiang, Wanshou, Li, Qingquan
Abstract
Integrated 3D reconstruction from aerial-ground images is essential for generating high-precision urban 3D models, yet severe variations in viewpoint, scale, and rotation make robust feature matching highly challenging. To address these limitations, this study introduces a rotation-robust detector-free matching network coupled with multi-view track refinement for incremental Structure from Motion (ISfM). The proposed workflow features four key modules. First, rotation-aware feature extraction replaces traditional convolutions with an Omnidirectional State Space Block (OSS Block) that selectively scans across eight symmetrical directions to model long-range spatial dependencies and synthesize rotation-invariant feature maps. Second, multi-scale attention transformation utilizes quadtree attention to build a hierarchical token pyramid that isolates high-association token regions and discards irrelevant areas, capturing long-range context with linear computational complexity. Third, bi-directional feature matching executes a symmetric coarse-to-fine matching scheme where coarse alignment computes dual-direction Softmax confidence matrices under mutual nearest neighbor constraints, and fine alignment uses a multi-layer perceptron to regress sub-pixel coordinate offsets. Finally, multi-view track refinement employs an integrated indexing structure to evaluate localized spatial proximity and link disjoint sub-tracks to the highest-confidence anchor point, ensuring stable feature repeatability across the ISfM pipeline. By using real aerial-ground datasets, experimental results demonstrate that the proposed method improves AUC at 5{\deg} pose error by 93.9% compared with LoFTR and achieves the highest precision in ISfM reconstruction, with the improved accuracy ranging from 27.6% to 32.7%. The proposed method provides a reliable solution for integrated 3D reconstruction of aerial-ground images.
Chinese Translation
从空中-地面图像进行集成3D重建对于生成高精度城市3D模型至关重要,但视点、尺度和旋转的严重变化使得鲁棒特征匹配面临巨大挑战。为了解决这些限制,本研究提出了一种旋转鲁棒的无检测器匹配网络,并结合多视图轨迹优化用于增量运动结构重建(Incremental Structure from Motion, ISfM)。所提出的工作流程包含四个关键模块。首先,旋转感知特征提取用全向状态空间块(Omnidirectional State Space Block, OSS Block)替代传统卷积,选择性地在八个对称方向上扫描,以建模长距离空间依赖关系并合成旋转不变特征图。其次,多尺度注意力变换利用四叉树注意力构建分层标记金字塔,隔离高关联标记区域并丢弃无关区域,以线性计算复杂度捕获长距离上下文。第三,双向特征匹配执行对称的粗到细匹配方案,其中粗对齐在互相最近邻约束下计算双向Softmax置信矩阵,细对齐则使用多层感知器回归亚像素坐标偏移。最后,多视图轨迹优化采用集成索引结构评估局部空间接近性,并将不相连的子轨迹链接到最高置信度的锚点,确保ISfM流程中特征的稳定重复性。通过使用真实的空中-地面数据集,实验结果表明,所提出的方法在5°姿态误差下的AUC提高了93.9%,与LoFTR相比,在ISfM重建中实现了最高精度,改进的准确性范围为27.6%至32.7%。该方法为空中-地面图像的集成3D重建提供了可靠的解决方案。
cs.CV / 69 / 2608.15259

UAV Video Deblurring via Motion-Aware Diffusion: A Path to Robust Target Detection

基于运动感知扩散的无人机视频去模糊:通往稳健目标检测的路径
Hu, Zhiqiang, Huang, Shouren, Ishikawa, Masatoshi
Abstract
Unmanned Aerial Vehicles (UAVs) play a crucial role in various scenarios ranging from disaster response to traffic surveillance. However, aerial video footage often suffers from severe motion blur due to rapid flight maneuvers, vibrations, and camera panning, which can significantly degrade downstream tasks such as target detection. Our goal is to explore a computationally-efficient and effective video deblurring approach to enhance UAV target detection performance. To reduce computational cost, we first propose an Adaptive Latent Scale Selector that dynamically adjusts the latent space resolution according to the intensity of UAV motion, thus balancing detail preservation with inference efficiency. To ensure temporal consistency, we introduce a Multi-Frame Alignment and Learnable Gating module to warp and gate the preceding frames, allowing the model to fuse only relevant temporal information and suppress misaligned or uninformative features. Our method can effectively recover sharp details from the UAV video stream. Extensive experiments on real UAV benchmarks demonstrate that our method not only yields superior deblurring performance but also significantly boosts target detection accuracy, making it highly applicable to robust aerial vision tasks.
Chinese Translation
无人机(UAV)在从灾害响应到交通监控等多种场景中发挥着至关重要的作用。然而,由于快速飞行机动、振动和相机平移,航拍视频常常受到严重的运动模糊,这会显著降低下游任务(如目标检测)的效果。我们的目标是探索一种计算高效且有效的视频去模糊方法,以提升无人机目标检测性能。为了降低计算成本,我们首先提出了一种自适应潜在尺度选择器(Adaptive Latent Scale Selector),该选择器根据无人机运动的强度动态调整潜在空间的分辨率,从而在细节保留与推理效率之间取得平衡。为了确保时间一致性,我们引入了一个多帧对齐和可学习门控模块(Multi-Frame Alignment and Learnable Gating),该模块对前一帧进行扭曲和门控,使模型仅融合相关的时间信息,并抑制不对齐或无信息的特征。我们的方法能够有效地从无人机视频流中恢复清晰的细节。在真实无人机基准测试上的大量实验表明,我们的方法不仅在去模糊性能上优于其他方法,而且显著提高了目标检测的准确性,使其在稳健的空中视觉任务中具有高度适用性。
cs.CV / 70 / 2608.15260

VGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D Reconstruction

VGGT-Align:桥接局部重建与全局一致性以实现长序列3D重建
Zhang, Wei, Wu, Yihang, Li, Songhua, Wang, Qi
Abstract
Maintaining global geometric consistency is a central challenge in long-sequence 3D reconstruction, with scale drift being the most critical failure mode. In chunk-based inference pipelines, the scale degree of freedom in sequential Sim(3) alignment is left unconstrained, causing estimation errors to compound multiplicatively and distort global trajectories and point cloud geometry. We present a scale-consistency enhancement framework built on a key insight: in structured environments such as driving scenes, geometric quantities arising from environmental regularity remain inherently invariant across temporal segments, and discrepancies in their per-chunk measurements directly expose inter-chunk scale drift. We propose Scene Geometric Invariant Anchoring (SGIA), which extracts dominant geometric invariants from each chunk's predicted point cloud via coarse-to-fine robust estimation and exploits their cross-chunk consistency to establish scale constraints independent of point cloud registration, explicitly degenerating 7-DoF Sim(3) alignment into 6-DoF rigid-body transformation and severing chain-wise scale error propagation at its source. We further introduce a lightweight test-time adaptation strategy that fine-tunes only normalization-layer parameters via multi-objective self-supervision, progressively improving intra-chunk predictions along the sequence. Both modules are plug-and-play and require no offline retraining. Experiments on multiple long-sequence benchmarks demonstrate state-of-the-art performance, reducing absolute trajectory error by up to 32% with significant gains in trajectory stability and reconstruction quality. Code: https://github.com/WZ-CS/VGGT-Align
Chinese Translation
在长序列3D重建中,保持全局几何一致性是一个核心挑战,其中尺度漂移是最关键的失败模式。在基于块的推理管道中,顺序Sim(3)对齐中的尺度自由度未受到约束,导致估计误差呈乘法性累积,从而扭曲全局轨迹和点云几何。我们提出了一种基于关键见解的尺度一致性增强框架:在结构化环境(如驾驶场景)中,由环境规律性产生的几何量在时间段之间本质上保持不变,而它们在每个块的测量中的差异直接暴露了块间的尺度漂移。我们提出了场景几何不变锚定(Scene Geometric Invariant Anchoring, SGIA),该方法通过粗到细的鲁棒估计从每个块的预测点云中提取主导几何不变性,并利用它们的跨块一致性建立独立于点云配准的尺度约束,明确地将7自由度Sim(3)对齐简化为6自由度刚体变换,并在源头切断链式尺度误差传播。我们进一步引入了一种轻量级的测试时适应策略,通过多目标自我监督仅微调归一化层参数,逐步改善序列中的块内预测。这两个模块均为即插即用,无需离线重训练。在多个长序列基准上的实验表明,我们的方法实现了最先进的性能,将绝对轨迹误差降低了多达32%,并显著提高了轨迹稳定性和重建质量。代码: https://github.com/WZ-CS/VGGT-Align
cs.CV / 71 / 2608.15261

Boundary-Aligned Contribution Routing for Robust Optical--SAR Object Detection

边界对齐的贡献路由用于鲁棒的光学-合成孔径雷达目标检测
Zhang, Haifa, Wang, Yijing, Wang, Haoyu, Li, Zheng, Zuo, Zhiqiang
Abstract
Optical imagery provides rich appearance cues, whereas synthetic aperture radar (SAR) offers observations that are less sensitive to illumination and weather, making optical--SAR fusion attractive for remote-sensing object detection. However, the presence of multiple modalities does not guarantee beneficial fusion: imperfect spatial, temporal, and semantic correspondence can make an otherwise intact stream conditionally harmful and induce negative cross-modal transfer. We handle this issue through a model-specific task-utility perspective and learn task-conditioned contribution routing using detection supervision alone. The proposed fusion-boundary-aligned routing regulates each modality's contribution before the first learned cross-modal feature-value mixing operation. For architectures with frequent shallow interaction, a Feature Router performs cross-conditioned, group-addressable modulation near the input; for dual-backbone architectures, a Dual-Statistic Semantic Router predicts stream-level contribution weights from modality-specific average and maximum statistics before late semantic fusion. The routers require no explicit utility supervision, quality labels, reconstruction, or distillation. Experiments on M4-SAR and SpaceNet6-OTD cover nominal full inputs, controlled correspondence shifts, missing modalities, and four nonzero modality-corruption scenarios. Across the reported clean-training controls, routing improves full-input $\text{mAP}_{50}$ by 0.5--5.9 points. Relative to the corresponding modality-dropout baselines, it raises missing-modality $\text{mAP}_{50}$ by 7.6--41.6 points and reduces the negative-transfer rate by up to 12.7 percentage points. Spearman correlations between the learned routing weights and model-specific leave-one-modality-out utility range from 0.45 to 0.66, supporting the task-utility interpretation of the routing coefficients.
Chinese Translation
光学图像提供丰富的外观线索,而合成孔径雷达(SAR)则提供对光照和天气不太敏感的观测,使得光学- SAR 融合在遥感目标检测中具有吸引力。然而,多种模态的存在并不保证有利的融合:不完美的空间、时间和语义对应关系可能使原本完整的流在条件上变得有害,并导致负向跨模态传递。我们通过模型特定的任务效用视角来处理这一问题,仅使用检测监督学习任务条件的贡献路由。所提出的融合边界对齐路由在第一次学习的跨模态特征值混合操作之前调节每种模态的贡献。对于频繁浅层交互的架构,特征路由器在输入附近执行跨条件的、组可寻址的调制;对于双骨干架构,双统计语义路由器在晚期语义融合之前,根据模态特定的平均和最大统计量预测流级别的贡献权重。这些路由器不需要显式的效用监督、质量标签、重建或蒸馏。在 M4-SAR 和 SpaceNet6-OTD 上的实验涵盖了名义上的完整输入、受控的对应关系变化、缺失模态以及四种非零模态损坏场景。在报告的干净训练控制中,路由提高了完整输入的 $ ext{mAP}_{50}$ 0.5--5.9 个百分点。相对于相应的模态缺失基线,它提高了缺失模态的 $ ext{mAP}_{50}$ 7.6--41.6 个百分点,并将负向传递率降低了最多 12.7 个百分点。学习到的路由权重与模型特定的单模态留出效用之间的斯皮尔曼相关性范围为 0.45 到 0.66,支持路由系数的任务效用解释。
cs.CV / 72 / 2608.15267

On the Adversarial Robustness of Remote Sensing Semantic Change Detection

遥感语义变化检测的对抗鲁棒性研究
Yu, Weikang, Xu, Yonghao, Ghamisi, Pedram
Abstract
Semantic change detection (SCD) is a bitemporal dense-prediction task that jointly identifies changed regions and their semantic states before and after change. Unlike single-image segmentation or binary change detection, SCD couples two temporal inputs with timestamp-wise semantic prediction, change localization, and final semantic-change decoding, creating adversarial dependencies that are not captured by conventional robustness protocols. We present a task-specific evaluation framework that separates output-side attack objectives from input-side temporal perturbation access, enabling systematic analysis of component vulnerability and cross-temporal propagation. Experiments on four datasets and six representative CNN-, Transformer-, and state-space-based models evaluate component-level and temporal objectives, single- and dual-timestamp perturbations, multiple attack methods, and cross-architecture transferability. The results show that final semantic-change predictions can be severely corrupted even when binary change localization remains comparatively stable, and that perturbations or attack objectives associated with one timestamp can propagate to the prediction of the other. These behaviors occur across different architecture families, while direct cross-model transfer remains considerably weaker than white-box attacks. The study demonstrates that adversarial robustness in SCD depends on the complete bitemporal prediction pathway rather than on an individual branch or backbone family, and provides a structured protocol for evaluating robustness in coupled bitemporal image analysis. Code is available at https://github.com/EricYu97/AdvSCD.
Chinese Translation
语义变化检测(SCD)是一种双时相密集预测任务,旨在共同识别变化区域及其在变化前后的语义状态。与单幅图像分割或二元变化检测不同,SCD 将两个时间输入与时间戳语义预测、变化定位和最终语义变化解码相结合,形成传统鲁棒性协议无法捕捉的对抗依赖关系。我们提出了一种任务特定的评估框架,将输出端攻击目标与输入端时间扰动访问分离,从而实现对组件脆弱性和跨时间传播的系统分析。在四个数据集和六个代表性的基于卷积神经网络(CNN)、变换器(Transformer)和状态空间模型的模型上进行的实验评估了组件级和时间目标、单时间戳和双时间戳扰动、多种攻击方法以及跨架构的可转移性。结果表明,即使二元变化定位相对稳定,最终的语义变化预测也可能受到严重干扰,并且与一个时间戳相关的扰动或攻击目标可以传播到另一个时间戳的预测。这些行为在不同的架构家族中均有发生,而直接的跨模型转移仍然显著弱于白盒攻击。本研究表明,SCD 中的对抗鲁棒性依赖于完整的双时相预测路径,而非单一分支或主干家族,并提供了一种结构化协议用于评估耦合双时相图像分析中的鲁棒性。代码可在 https://github.com/EricYu97/AdvSCD 获取。
cs.CV / 73 / 2608.15277

Memory-Bounded Continuation of Greedy Sampling for Continual Anomaly Detection

基于记忆限制的贪婪采样在持续异常检测中的延续
Jung, Yoon Gyo, Park, Jaewoo, Peng, Kuan-Chuan, Bang, Seongdeok, Camps, Octavia
Abstract
Greedy sampling produces a compact yet representative summary of normal data, which is essential for reliable anomaly detection that relies on measuring distance from normality. For continual anomaly detection where tasks arrive sequentially, extending greedy sampling is straightforward with unbounded memory through coreset accumulation. However, practical deployment requires fixed memory where the coreset size remains constant regardless of task count. We observe that continued greedy sampling, which iteratively applies greedy selection over previously greedy-sampled sets, effectively preserves representativeness under strict memory limits. Despite discarding data at each step to satisfy the memory constraint, coreset quality degrades gracefully rather than catastrophically, enabling reliable anomaly detection across the tasks. We provide theoretical justification by showing that resulting greedy-continued coreset approximates the oracle coreset within a bounded gap. We instantiate this principle in ContCore, which constructs a greedy-continued coreset through greedy expansion on new task features followed by greedy consolidation to enforce the memory budget. Unlike neural methods susceptible to catastrophic forgetting or naive coreset accumulation requiring unbounded memory, ContCore maintains fixed memory with theoretical guarantees. Empirically, ContCore achieves state-of-the-art performance across 11 task schedules on MVTecAD and VisA, and extends effectively to online continual AD settings where prior methods degrade significantly. Code: https://github.com/jungyg/ContCore
Chinese Translation
贪婪采样生成了一个紧凑而具有代表性的正常数据摘要,这对于依赖于测量与正常状态的距离的可靠异常检测至关重要。在任务顺序到达的持续异常检测中,通过核心集累积,扩展贪婪采样在无界内存下是简单的。然而,实际部署需要固定内存,其中核心集大小保持不变,无论任务数量如何。我们观察到,持续的贪婪采样,通过对先前贪婪采样集的迭代贪婪选择,有效地在严格的内存限制下保持了代表性。尽管在每一步中丢弃数据以满足内存约束,但核心集的质量优雅地下降,而不是灾难性地下降,从而使得在各个任务中实现可靠的异常检测成为可能。我们提供了理论依据,表明所得到的贪婪延续核心集在有限差距内近似于oracle核心集。我们在ContCore中实例化了这一原则,该方法通过对新任务特征进行贪婪扩展,然后进行贪婪整合以强制执行内存预算,从而构建贪婪延续核心集。与容易遭受灾难性遗忘的神经方法或需要无界内存的简单核心集累积不同,ContCore在理论上保证了固定内存。实证结果表明,ContCore在MVTecAD和VisA的11个任务调度中达到了最先进的性能,并有效扩展到在线持续异常检测设置中,而先前的方法在这些设置中显著降级。代码链接: https://github.com/jungyg/ContCore
cs.CV / 74 / 2608.15279

Geometry-Aware Spatio-Temporal Context Modeling for 4D Occupancy Forecasting

基于几何感知的时空上下文建模用于4D占用预测
Chen, Sitao, Zhuang, Zhuangwei, Luo, Hui, Wu, Qingyao, Tan, Mingkui
Abstract
4D occupancy forecasting models the spatio-temporal evolution of 3D scenes and is crucial for autonomous driving, especially for corner-case simulation. Existing methods often rely on discrete tokenization followed by autoregressive prediction, yet struggle with geometric distortion in static structures and inconsistent temporal coherence over the forecasting horizon. In this work, we propose a Geometry-Aware Spatio-Temporal context modeling method (GAST) for 4D occupancy forecasting, built upon progressive explicit-implicit generation and dual-path spatio-temporal modeling. Specifically, the generation module produces per-frame occupancy with high geometric fidelity and semantic plausibility through pose-driven warping, motion-aware feature modulation, and attention-based feature refinement. Subsequently, the spatio-temporal module enhances spatial consistency through global context aggregation while capturing scene evolution through temporal dynamics extraction. This unified design enables joint optimization of historical reconstruction and future forecasting in an end-to-end manner. Extensive experiments on Occ3D-nuScenes demonstrate the superiority of our method, outperforming the state-of-the-art by 7.67% in mIoU and 6.44% in IoU with a 2.84x speedup, while maintaining strong performance in long-term forecasting.
Chinese Translation
4D占用预测建模3D场景的时空演变,对于自动驾驶尤其是边缘案例模拟至关重要。现有方法通常依赖于离散标记化,随后进行自回归预测,但在静态结构的几何失真和预测范围内的不一致时间连贯性方面存在困难。在本研究中,我们提出了一种基于几何感知的时空上下文建模方法(Geometry-Aware Spatio-Temporal context modeling,GAST),用于4D占用预测,该方法基于渐进式显式-隐式生成和双路径时空建模。具体而言,生成模块通过姿态驱动的变形、运动感知特征调制和基于注意力的特征精炼,产生具有高几何保真度和语义合理性的逐帧占用。随后,时空模块通过全局上下文聚合增强空间一致性,同时通过时间动态提取捕捉场景演变。这种统一设计使得历史重建和未来预测能够以端到端的方式进行联合优化。在Occ3D-nuScenes上的大量实验表明,我们的方法优于现有最先进技术,在mIoU上提高了7.67%,在IoU上提高了6.44%,同时实现了2.84倍的加速,并在长期预测中保持了强劲的性能。
cs.CV / 75 / 2608.15295

SOS! : A Streamlined Object-Conditional Transformer for Model-free Segmentation

SOS!: 一种简化的对象条件变换器用于无模型分割
Hu, Jiaqi, Huang, Junwen, Xu, Hongli, Yu, Peter KT, Navab, Nassir, Busam, Benjamin, Ilic, Slobodan
Abstract
Foundation segmentation models excel at generating high-quality, class-agnostic masks, but they struggle to associate these proposals with specific target objects. This semantic gap severely hinders their deployment in downstream applications like robotic manipulation, which demand precise unseen objects segmentation. Existing approaches attempt to resolve this by relying on exhaustive 3D object model priors, inherently introducing prohibitive computational overhead and complex, multi-stage pipelines. To address these limitations, we propose SOS (Streamlined Object-conditional Transformer for model-free Segmentation). SOS completely eliminates the reliance on 3D models, requiring only a single reference image per target object. Central to our framework is a novel Object-Conditional Transformer that learns identity-anchored queries, unifying mask generation and target identification into a single feed-forward pass. This streamlined design drastically improves both structural and computational efficiency. Extensive evaluations across multiple benchmarks demonstrate that SOS establishes a new state-of-the-art for model-free unseen objects segmentation, delivering accurate and high-efficiency performance. The project page and code are available at https://sos-seg.github.io/.
Chinese Translation
基础分割模型在生成高质量、类别无关的掩膜方面表现出色,但它们在将这些提议与特定目标对象关联时却面临困难。这一语义差距严重阻碍了它们在下游应用中的部署,例如机器人操作,这些应用需要对未见对象进行精确分割。现有方法试图通过依赖详尽的3D对象模型先验来解决这一问题,但这不可避免地引入了高昂的计算开销和复杂的多阶段管道。为了解决这些局限性,我们提出了SOS(Streamlined Object-conditional Transformer for model-free Segmentation)。SOS完全消除了对3D模型的依赖,仅需每个目标对象一张参考图像。我们框架的核心是一个新颖的对象条件变换器,它学习身份锚定查询,将掩膜生成和目标识别统一为单次前馈传递。这一简化设计显著提高了结构和计算效率。在多个基准测试中的广泛评估表明,SOS在无模型未见对象分割方面建立了新的最先进水平,提供了准确且高效的性能。项目页面和代码可在 https://sos-seg.github.io/ 获取。
cs.CV / 76 / 2608.15296

FMReward: Aligning and Evaluating Audio-Driven 3D Facial Animation with Human Preferences

FMReward:与人类偏好对齐和评估音频驱动的3D面部动画
Wu, Sijing, Li, Yunhao, Gao, Zhilin, Duan, Huiyu, Zhu, Yucheng, Zhai, Guangtao, Callet, Patrick Le
Abstract
Audio-driven 3D facial animation is essential for advancing immersion and interactivity in virtual experiences. Although recent advances have shown promising capabilities, the training and evaluation of existing methods typically rely on ground-truth-based errors, which fall short of aligning with human preferences. To address this, we present a comprehensive framework that learns an automatic perceptual model from human preference data and leverages it to improve and evaluate the perceptual quality of audio-driven 3D facial animation. To begin with, we construct FMPair (Facial Motion Pairwise preference), the first human preference dataset for audio-driven 3D facial animation, which is built through a systematic annotation pipeline and comprises 65,574 annotated 3D facial motion pairs from 8,834 distinct in-the-wild audio clips. Based on the pairwise comparison dataset, we propose a Facial Motion Reward model, termed FMReward, which takes audio and 3D facial motion as inputs and predicts a perceptual quality score aligned with human preferences. Building upon FMReward, we further introduce Facial Motion reward Feedback Learning (FMFL), a direct fine-tuning algorithm that leverages a pretrained reward model to optimize diffusion-based audio-driven 3D facial animation models for better alignment with human preferences. Extensive experiments demonstrate the superiority of FMReward over other metrics in aligning with human preferences and the effectiveness of FMFL in improving the perceptual quality of audio-driven 3D facial animation.
Chinese Translation
音频驱动的3D面部动画对于提升虚拟体验的沉浸感和互动性至关重要。尽管近期的进展显示出令人鼓舞的能力,但现有方法的训练和评估通常依赖于基于真实值的误差,这不足以与人类偏好对齐。为了解决这一问题,我们提出了一个全面的框架,该框架从人类偏好数据中学习自动感知模型,并利用该模型来改善和评估音频驱动的3D面部动画的感知质量。首先,我们构建了FMPair(面部动作成对偏好),这是第一个针对音频驱动的3D面部动画的人类偏好数据集,该数据集通过系统的标注流程构建,包含来自8,834个不同的野外音频片段的65,574对标注的3D面部动作。基于成对比较数据集,我们提出了面部动作奖励模型FMReward,该模型以音频和3D面部动作为输入,预测与人类偏好对齐的感知质量评分。在FMReward的基础上,我们进一步引入了面部动作奖励反馈学习(FMFL),这是一种直接微调算法,利用预训练的奖励模型来优化基于扩散的音频驱动3D面部动画模型,以更好地与人类偏好对齐。大量实验表明,FMReward在与人类偏好对齐方面优于其他指标,而FMFL在改善音频驱动的3D面部动画的感知质量方面也表现出有效性。
cs.CV / 77 / 2608.15297

TinyDETR-Pose: Towards End-to-End Real-Time Single-Stage 6DoF Object Pose Estimation with Lightweight Transformers

TinyDETR-Pose:面向轻量级变换器的端到端实时单阶段6自由度物体姿态估计
Kühn, Paul Julius, Nguyen, Duc Anh, Sinha, Saptarshi Neil, Weinmann, Michael, Kuijper, Arjan
Abstract
Real-time 6DoF object pose estimation on resource-constrained hardware remains challenging, as accurate correspondence-based and refinement pipelines typically rely on non-differentiable PnP/RANSAC stages or costly iterative refinement, while recent foundation-model-based approaches incur inference costs that are prohibitive for edge deployment. We present TinyDETR-Pose, a lightweight, end-to-end, single-stage framework that jointly detects objects and regresses their full 6D pose in a single forward pass. Built on the efficient LW-DETR architecture, TinyDETR-Pose formulates detection and pose estimation as a set-prediction problem and attaches dedicated MLP heads for rotation, monocular depth, and projected object center regression to each decoder query, eliminating the need for PnP, NMS (non-maximum suppression), or iterative pose refinement. Object symmetries are handled through a ADD-S loss applied uniformly to all objects, without the need for object-specific loss schedules or separate geodesic/ADD supervision. In addition, predictions are assigned to ground truth using a symmetry-safe Hungarian matcher based on class and 2D spatial cues, yielding stable assignment under symmetry and depth ambiguity. On YCB-V, TinyDETR-Pose achieves a comparable ADD-S AUC of 85.9, while requiring up to 72.7% fewer parameters than other DETR-based single-stage pose-estimation approaches. Due to its compact design, TinyDETR-Pose runs in real time and achieves an inference latency of only ~4.5 ms per frame on an NVIDIA Jetson Nano using TensorRT, demonstrating that accurate end-to-end transformer-based 6D pose estimation can be made practical for edge deployment.
Chinese Translation
在资源受限的硬件上实现实时6自由度物体姿态估计仍然面临挑战,因为基于准确对应和精细化的流程通常依赖于不可微分的PnP/RANSAC阶段或成本高昂的迭代精细化,而最近的基础模型方法则会导致推理成本过高,无法在边缘设备上部署。我们提出了TinyDETR-Pose,这是一种轻量级的端到端单阶段框架,可以在单次前向传递中同时检测物体并回归其完整的6D姿态。TinyDETR-Pose基于高效的LW-DETR架构,将检测和姿态估计公式化为一个集合预测问题,并为每个解码器查询附加专用的多层感知器(MLP)头,用于旋转、单目深度和投影物体中心回归,从而消除了对PnP、NMS(非最大抑制)或迭代姿态精细化的需求。物体对称性通过均匀应用于所有物体的ADD-S损失进行处理,无需特定于物体的损失调度或单独的测地线/ADD监督。此外,预测通过基于类别和二维空间线索的对称安全匈牙利匹配器分配给真实值,从而在对称性和深度模糊下实现稳定的分配。在YCB-V数据集上,TinyDETR-Pose达到了85.9的可比ADD-S AUC,同时所需参数比其他基于DETR的单阶段姿态估计方法少72.7%。由于其紧凑的设计,TinyDETR-Pose能够实时运行,并在使用TensorRT的NVIDIA Jetson Nano上实现每帧仅约4.5毫秒的推理延迟,证明了基于变换器的准确端到端6D姿态估计在边缘部署中是可行的。
cs.CV / 78 / 2608.15298

Image Denoising via the Adaptive Rank-Cluster Filter

通过自适应秩聚类滤波器进行图像去噪
Pozdnyakov, Dmitry
Abstract
A spatial-local image-denoising filter is proposed, and its performance metrics are evaluated in comparison with baseline filtering algorithms, including the median, adaptive median, Gaussian, bilateral, Wiener, anisotropic diffusion, and non-local means. The developed filter is based on aligning the intensity value of the central pixel in a 3x3 window with the statistical majority intensity of one of the two clusters formed by optimal Otsu's partitioning of a pixel set sorted by intensity and trimmed to seven elements. This is followed by a fuzzy fusion of the calculated value with the median intensity of the pixels within the window. The proposed filter demonstrates the highest robustness to variations in image noise levels, particularly when processing mixed noise consisting of salt-and-pepper impulse noise and additive Gaussian noise in various proportions
Chinese Translation
提出了一种空间局部图像去噪滤波器,并将其性能指标与基线滤波算法进行比较,包括中值滤波、自适应中值滤波、高斯滤波、双边滤波、维纳滤波、各向异性扩散和非局部均值。所开发的滤波器基于将3x3窗口中中心像素的强度值与通过最优Otsu分割形成的两个聚类之一的统计多数强度对齐,该像素集按强度排序并修剪为七个元素。接下来,通过模糊融合计算值与窗口内像素的中值强度。所提议的滤波器在图像噪声水平变化时表现出最高的鲁棒性,特别是在处理由盐和胡椒脉冲噪声与不同比例的加性高斯噪声组成的混合噪声时。
cs.CV / 79 / 2608.15317

LightLoc++: Sensor-Robust Representation Learning for Efficient Outdoor LiDAR Localization

LightLoc++:用于高效户外激光雷达定位的传感器鲁棒表示学习
Li, Wen, Yu, Shangshu, Liu, Dunqiang, Xia, Qiming, Ao, Sheng, Shen, Siqi, Wen, Chenglu, Wang, Cheng
Abstract
Scene coordinate regression (SCR) achieves strong performance in outdoor LiDAR localization, but it usually requires scene-specific training that can take days, limiting practical deployment. Recent works improve training efficiency by decoupling SCR into a scene-agnostic backbone and scene-specific prediction heads, where the backbone is pretrained on source datasets and frozen for new scenes, and only lightweight heads are optimized. However, we find that this paradigm heavily depends on the pretrained backbone. Existing decoupled methods can match conventional SCR methods fully optimized for each new scene when LiDAR configurations are similar to those used during backbone pretraining, but their accuracy drops noticeably on datasets collected with different LiDAR sensors. This suggests that efficient LiDAR localization requires representations that capture stable scene geometry across LiDAR configurations. Motivated by this observation, we propose LightLoc++, a sensor-robust and efficient outdoor LiDAR localization framework. To support sensor-robust representation learning, we introduce SULID, a synchronized urban multi-LiDAR dataset with representative 32-, 64-, and 128-beam rotating LiDARs, extensive cross-sensor overlap, and diverse urban scenes. Using SULID, we pretrain a sensor-robust backbone through cross-sensor consistency learning. LightLoc++ further preserves efficient new-scene learning by incorporating sample classification guidance and redundant sample downsampling, which reduce regression ambiguity and computational redundancy in large-scale outdoor scenes. Extensive experiments on multiple outdoor LiDAR localization benchmarks demonstrate that LightLoc++ achieves state-of-the-art localization performance with the lowest new-scene training cost among compared methods. Code and dataset will be made available at https://github.com/liw95/LightLoc-PlusPlus.
Chinese Translation
场景坐标回归(SCR)在户外激光雷达定位中表现出色,但通常需要特定场景的训练,这可能需要数天时间,限制了实际部署。近期的研究通过将SCR解耦为场景无关的主干网络和场景特定的预测头,改善了训练效率,其中主干网络在源数据集上进行预训练并在新场景中保持不变,仅优化轻量级的预测头。然而,我们发现这种范式严重依赖于预训练的主干网络。现有的解耦方法在激光雷达配置与主干预训练时使用的配置相似时,可以与为每个新场景完全优化的传统SCR方法相匹配,但在使用不同激光雷达传感器收集的数据集上,其准确性明显下降。这表明,高效的激光雷达定位需要能够捕捉不同激光雷达配置下稳定场景几何的表示。基于这一观察,我们提出了LightLoc++,一个传感器鲁棒且高效的户外激光雷达定位框架。为了支持传感器鲁棒的表示学习,我们引入了SULID,一个同步的城市多激光雷达数据集,包含具有代表性的32束、64束和128束旋转激光雷达,广泛的跨传感器重叠以及多样的城市场景。利用SULID,我们通过跨传感器一致性学习预训练了一个传感器鲁棒的主干网络。LightLoc++进一步通过引入样本分类指导和冗余样本下采样来保持高效的新场景学习,这减少了大规模户外场景中的回归模糊性和计算冗余。在多个户外激光雷达定位基准上的广泛实验表明,LightLoc++在比较方法中实现了最先进的定位性能,并且新场景训练成本最低。代码和数据集将发布在 https://github.com/liw95/LightLoc-PlusPlus。
cs.CV / 80 / 2608.15336

SAGE-OR: Semi-supervised Adaptive Scene Graph Generation for Operating Rooms

SAGE-OR:用于手术室的半监督自适应场景图生成
Leblanc, Brandon, Poullis, Charalambos
Abstract
Current surgical scene graph generation methods depend on dense multi-modal supervision and specialized hardware (synchronized RGB-D sensors, calibration rigs), making dataset construction expensive and restricting all existing benchmarks to simulated environments. We propose SAGE-OR, a feature-centric framework that replaces the traditional detect-then-reason paradigm with a decoupled representation-reasoning paradigm in which localization is derived from frozen foundation models, encoded implicitly in pre-computed features, and used without any localization supervision, while a lightweight graph transformer performs relational reasoning over cached features. We employ a semi-supervised formulation with general-purpose segmentation prompts to eliminate localization supervision while enabling unsupervised context augmentation through additional prompt-driven entities, such as hands, which are absent from annotations. General-purpose prompts are used to induce near-perfect recall, while precision is delegated to downstream attention-based reasoning, enabling simple adaptation to new entities via prompt-level modification. This design enables a lightweight 15M-parameter graph transformer that trains in 1.4 hours and runs relational inference at $\sim$1ms per frame with peak memory under 2GB, suitable for edge hardware used in the operating room; feature extraction runs offline as a separate caching stage (4.27s per frame). On the 4D-OR benchmark, the core model achieves 76% F1, matching the fully supervised 4D-OR baseline while eliminating all localization annotations, and unsupervised hand augmentation raises this to 86%, within 4 points of state-of-the-art (SOTA) methods requiring dense multi-modal supervision, providing a practical pathway for adaptation to new surgical settings without annotation other than relationship and class labels.
Chinese Translation
当前的手术场景图生成方法依赖于密集的多模态监督和专用硬件(同步RGB-D传感器、校准设备),使得数据集构建成本高昂,并限制所有现有基准于模拟环境。我们提出了SAGE-OR,一个以特征为中心的框架,替代了传统的检测-推理范式,采用解耦的表示-推理范式,其中定位来自冻结的基础模型,隐式编码在预计算的特征中,并在没有任何定位监督的情况下使用,同时轻量级图变换器对缓存特征进行关系推理。我们采用半监督的形式,使用通用的分割提示来消除定位监督,同时通过缺少注释的附加提示驱动实体(如手)实现无监督的上下文增强。通用提示用于诱导近乎完美的召回率,而精确度则委托给下游基于注意力的推理,从而通过提示级修改实现对新实体的简单适应。该设计使得一个轻量级的1500万参数图变换器能够在1.4小时内训练,并以每帧约1毫秒的速度进行关系推理,峰值内存低于2GB,适用于手术室中使用的边缘硬件;特征提取作为单独的缓存阶段离线运行(每帧4.27秒)。在4D-OR基准上,核心模型实现了76%的F1分数,匹配完全监督的4D-OR基线,同时消除了所有定位注释,而无监督的手部增强将这一分数提高至86%,距离需要密集多模态监督的最先进(SOTA)方法仅相差4分,为在新的手术环境中适应提供了一条实用的路径,除了关系和类别标签外无需其他注释。
cs.CV / 81 / 2608.15341

TEA: Text Encoder Alignment for Robust Concept Erasure in Text-to-Image Models

TEA:文本编码器对齐用于文本到图像模型中的稳健概念抹除
Farashah, Alireza Dehghanpour, Shi, Zhuan, Rostamzadeh, Negar, Farnadi, Golnoosh
Abstract
Text-to-image diffusion models can be misused to generate harmful content through adversarial or paraphrased prompts that bypass built-in safety mechanisms. Existing concept erasure methods often suffer from limited robustness against adversarial prompts, degradation of benign generation quality, or reliance on inference-time interventions that introduce persistent computational overhead. To address these limitations, we formulate concept erasure as a domain alignment problem in the text representation space. We propose a lightweight Text Encoder Alignment framework (TEA) that fine-tunes only the text encoder while keeping the generative backbone fully frozen. Given concept--anchor prompt pairs, our method trains a discriminator to distinguish token-level representations of concept-containing prompts from those of safe anchor prompts, while updating the text encoder to make these representations indistinguishable. TEA introduces zero inference-time overhead and requires only a small number of fine-tuning steps, making it highly efficient to deploy at scale. Despite this efficiency, TEA achieves state-of-the-art erasure robustness against black-box and white-box adversarial attacks on Stable Diffusion v1.4, while preserving generation quality on benign prompts. Furthermore, TEA is model-agnostic and achieves the lowest attack success rate on Stable Diffusion v3.5, extending concept erasure to a Rectified Flow Transformer architecture with T5 conditioning where prior methods remain largely unexplored. Code is available at \href{https://github.com/alirezafarashah/TEA.git}{https://github.com/alirezafarashah/TEA.git}
Chinese Translation
文本到图像扩散模型可能被滥用,通过对抗性或改写的提示生成有害内容,从而绕过内置的安全机制。现有的概念抹除方法往往在对抗性提示下表现出有限的稳健性,导致良性生成质量下降,或依赖于推理时的干预,这会引入持续的计算开销。为了解决这些局限性,我们将概念抹除形式化为文本表示空间中的领域对齐问题。我们提出了一种轻量级的文本编码器对齐框架(TEA),该框架仅微调文本编码器,同时保持生成主干完全冻结。给定概念-锚点提示对,我们的方法训练一个判别器,以区分包含概念的提示的标记级表示与安全锚点提示的表示,同时更新文本编码器,使这些表示不可区分。TEA 引入零推理时开销,仅需少量微调步骤,使其在大规模部署时高度高效。尽管如此,TEA 在对抗性攻击(包括黑箱和白箱攻击)下,在 Stable Diffusion v1.4 上实现了最先进的抹除稳健性,同时保持良性提示的生成质量。此外,TEA 是模型无关的,并在 Stable Diffusion v3.5 上实现了最低的攻击成功率,将概念抹除扩展到具有 T5 条件的修正流变换器架构,而先前的方法在这一领域仍然未被充分探索。代码可在 [https://github.com/alirezafarashah/TEA.git](https://github.com/alirezafarashah/TEA.git) 获取。
cs.CV / 82 / 2608.15343

Feed-Forward Hierarchical Gaussian Diffusion for Extreme CT Reconstruction

用于极端计算机断层扫描重建的前馈层次高斯扩散
Yang, Yuezhe, Cheng, Li
Abstract
Reconstructing three-dimensional computed tomography (CT) from severely constrained projections is highly ill-posed. Sparse angular sampling, restricted angular coverage, and low photon counts can occur individually or jointly, obscuring global anatomy and local tissue detail. Many learned CT reconstruction methods are tailored to a single dominant degradation. Existing diffusion and Gaussian approaches commonly recover global structure and local detail within a shared representation. We propose HiGDiff, a feed-forward hierarchical Gaussian diffusion framework that decomposes reconstruction both spatially and from structure to detail. Physics-conditioned anatomical anchors and a foreground capacity field allocate learnable Gaussian primitives to informative regions. A structure diffusion stage first recovers global attenuation geometry, and its learned representation conditions a detail diffusion stage for residual boundaries and tissue transitions. The resulting Gaussian banks are rendered as attenuation fields and further refined by a gradient-isolated residual module. Experiments on three distinct CT benchmark datasets demonstrate state-of-the-art reconstruction performance across isolated, paired, and joint degradation settings, including improvements of 5.81 dB in macro-average peak signal-to-noise ratio (PSNR) and 0.113 in structural similarity index measure (SSIM) on the Low Dose CT Image and Projection Data (LDCT-PD) collection. Code and experimental configurations are openly available at https://github.com/Bean-Young/HiGDiff.
Chinese Translation
从严重受限的投影中重建三维计算机断层扫描(CT)是一个高度不适定的问题。稀疏的角度采样、受限的角度覆盖以及低光子计数可能单独或共同发生,模糊了全局解剖结构和局部组织细节。许多学习型CT重建方法都是针对单一主导退化进行定制的。现有的扩散和高斯方法通常在共享表示中恢复全局结构和局部细节。我们提出了HiGDiff,一个前馈层次高斯扩散框架,该框架在空间上以及从结构到细节对重建进行分解。物理条件下的解剖锚点和前景容量场将可学习的高斯原语分配给信息丰富的区域。结构扩散阶段首先恢复全局衰减几何,其学习到的表示为残差边界和组织过渡的细节扩散阶段提供条件。最终生成的高斯库被渲染为衰减场,并通过一个梯度隔离的残差模块进一步精炼。在三个不同的CT基准数据集上的实验表明,在孤立、配对和联合退化设置中,重建性能达到了最先进的水平,包括在低剂量CT图像和投影数据(LDCT-PD)集合上,宏平均峰值信噪比(PSNR)提高了5.81 dB,结构相似性指数测量(SSIM)提高了0.113。代码和实验配置可在https://github.com/Bean-Young/HiGDiff上公开获取。
cs.CV / 83 / 2608.15349

ENAF: A Multi-Exit Network with an Adaptive Patch Fusion for Large Image Super Resolution

ENAF:一种具有自适应补丁融合的大型图像超分辨率多出口网络
Nguyen, Duong M., Nguyen, Tuan Nghia, Nguyen, Xuan Truong
Abstract
To accelerate single image super-resolution (SISR) networks on large images (2K-8K), many recent approaches decompose an image into small patches and dynamically determine an execution path according to its difficulty (referred to as a dynamic network). To quantify the hardness of a patch, they mainly rely on a handcrafted assessment score, e.g., edge, which weakly associates a patch's texture with the computational complexity of a SISR model. To address the problem, we introduce ENAF - a dynamic network for SISR with an adaptive patch fusion. Built on top of a backbone, ENAF incorporates multiple early exits (EEs) to tackle the over-parameterized SISR model. More importantly, ENAF plugs a tiny network that estimates PSNR to associate data texture with a computation cost at an EE. Based on the scores, ENAF effectively assigns image patches to an exit, enhancing the quality-complexity trade-off. Extensive experiments on common datasets with popular SISR backbones demonstrate the effectiveness of ENAF in various settings. The source code is provided in https://github.com/nmduonggg/ENAF
Chinese Translation
为了加速在大型图像(2K-8K)上的单幅图像超分辨率(SISR)网络,许多近期的方法将图像分解为小补丁,并根据其难度动态确定执行路径(称为动态网络)。为了量化补丁的难度,它们主要依赖于手工评估分数,例如边缘,这种方法将补丁的纹理与SISR模型的计算复杂性弱相关。为了解决这个问题,我们引入了ENAF——一种具有自适应补丁融合的SISR动态网络。ENAF建立在一个主干网络之上,结合了多个早期出口(EEs)以应对过参数化的SISR模型。更重要的是,ENAF插入了一个小型网络来估计PSNR,以将数据纹理与EE的计算成本关联起来。基于这些分数,ENAF有效地将图像补丁分配给出口,从而增强了质量与复杂性的权衡。在常见数据集和流行SISR主干网络上的大量实验证明了ENAF在各种设置中的有效性。源代码可在 https://github.com/nmduonggg/ENAF 获取。
cs.CV / 84 / 2608.15353

Decomposing Whole Slide Image Report Generation with Graph-Constrained Multiple Instance Learning Workflows

基于图约束的多实例学习工作流分解全幻灯片图像报告生成
Gitau, Antony, Borak, Martyna, Singstad, Bjørn-Jostein, Paulson, Martin, Hjelmervik, Karl Thomas, Lysaker, Ola Marius, Sanchez, Veralia Gabriela
Abstract
Whole-slide image (WSI) report generation requires recognizing spatially distributed pathological features and organizing them into a coherent diagnostic narrative. Although direct vision-to-text models can yield fluent reports, they obscure the contributions and failure modes of visual recognition, structured reasoning, and language generation. We propose a decomposed framework in which frozen Virchow2 tile embeddings are aggregated by multiple-instance learning (MIL) classification heads that answer organ-specific diagnostic questions. An organ-conditioned graph constrains the assembly of these answers into a structured reasoning chain, which a language model realizes as a pathology report. On the REG2026 held-out set of 2,028 slides, the proposed workflow achieved a chain-Jaccard score of 0.702. Performance fell to 0.420 without graph-based chain construction, 0.398 when the organ-specific graphs were replaced by a single organ-agnostic graph, and 0.371 when the language model constructed the chain freely from MIL predictions. Using the same report generator, graph-structured chains improved the report score from 0.330 to 0.495. On 350 external TCGA WSIs spanning the seven REG organs without fine-tuning, the expected organ graph was selected in 64.0% of cases and ranked among the top three in 86.6%. Providing the correct organ graph increased agreement with coarse TCGA primary-diagnosis labels from 61.8% to 92.6%, identifying organ routing as a main bottleneck under domain shift. Overall, organ-conditioned, graph-constrained chain assembly improves structured reasoning and report generation while enabling stage-specific error localization.
Chinese Translation
全幻灯片图像(WSI)报告生成需要识别空间分布的病理特征,并将其组织成连贯的诊断叙述。尽管直接的视觉到文本模型可以生成流畅的报告,但它们掩盖了视觉识别、结构推理和语言生成的贡献和失败模式。我们提出了一种分解框架,其中冻结的Virchow2瓦片嵌入通过多实例学习(MIL)分类头进行聚合,以回答特定器官的诊断问题。一个器官条件图约束了这些答案的组装,形成一个结构化推理链,语言模型将其实现为病理报告。在REG2026的2,028张幻灯片的保留集上,所提出的工作流程达到了0.702的链Jaccard得分。没有基于图的链构建时,性能降至0.420;当器官特定图被替换为单一的器官无关图时,得分为0.398;而当语言模型自由地从MIL预测中构建链时,得分为0.371。使用相同的报告生成器,图结构链将报告得分从0.330提高到0.495。在350个外部TCGA WSI中,涵盖七个REG器官且未进行微调的情况下,预期的器官图在64.0%的案例中被选择,并在86.6%的情况下排名前列。提供正确的器官图使与粗略的TCGA初步诊断标签的一致性从61.8%提高到92.6%,识别器官路由作为领域转移下的主要瓶颈。总体而言,器官条件的图约束链组装改善了结构化推理和报告生成,同时实现了阶段特定的错误定位。
cs.CV / 85 / 2608.15363

A Multi-Annotator Study of Segmentation Noise and Uncertainty in Turbid Underwater Images

浑浊水下图像分割噪声与不确定性的多标注者研究
Humblot-Renaux, Galadrielle, Ismiroglou, Vasiliki, Pedersen, Malte
Abstract
Label uncertainty and annotator disagreement are common challenges in the field of computer vision, yet their study has largely been confined to the medical domain or to generic image-recognition datasets. Underwater datasets are particularly susceptible to these issues due to the need for domain expertise, degraded visibility conditions, and the inherent difficulty of establishing reliable ground truth in inaccessible environments. Despite these challenges, annotation uncertainty in underwater imagery remains largely unexplored. In this work, we present the first systematic multi-annotator study of segmentation in real underwater scenes, with over 100 participants, and across varying, controlled levels of turbidity. We show that underwater datasets face many of the same annotation challenges as other vision tasks, while turbidity introduces additional systematic errors. We further investigate the main factors driving label noise and explore ways to improve annotation quality in turbid underwater environments, including privileged information, individual effort and annotator ensembles. All (meta-) data collected in this study will be available on the project page: https://vap.aau.dk/tubcertainty
Chinese Translation
标签不确定性和标注者之间的分歧是计算机视觉领域常见的挑战,但其研究大多局限于医学领域或通用图像识别数据集。水下数据集特别容易受到这些问题的影响,因为需要领域专业知识、能见度条件恶化以及在无法接触的环境中建立可靠的真实标签的固有困难。尽管面临这些挑战,水下图像中的标注不确定性仍然基本未被探索。在本研究中,我们首次系统性地进行了一项多标注者的真实水下场景分割研究,参与者超过100人,并在不同的、可控的浑浊度水平下进行。我们展示了水下数据集面临与其他视觉任务相同的许多标注挑战,同时浑浊度引入了额外的系统性误差。我们进一步探讨了导致标签噪声的主要因素,并探索了在浑浊水下环境中提高标注质量的方法,包括特权信息、个体努力和标注者集成。本研究中收集的所有(元)数据将可在项目页面上获取:https://vap.aau.dk/tubcertainty
cs.CV / 86 / 2608.15395

JoLT: Joint Latent Trajectories for Context-Guided High-Resolution Tiled Generation

JoLT:用于上下文引导的高分辨率平铺生成的联合潜在轨迹
Koroglu, Mathis, Jeanneret, Guillaume, Caselles-Dupré, Hugo, Cord, Matthieu, Dapogny, Arnaud
Abstract
Although text-to-image generative models produce impressive results, they struggle to generate densely detailed, high-resolution (HR) images. Current literature addresses this issue with a low-to-high-resolution approach. First, a low-resolution (LR) image is generated. Then, an upsampled version is generated using the LR image as an additional cue. In this paper, we present Joint Latent Trajectories (JoLT). To generate an image, JoLT uses two streams that jointly denoise LR and HR latent images at each sampling step. The LR latent controls the overall layout, while the HR latent controls the details. We interconnect both branches to jointly integrate their information. We extensively validate our method, demonstrating its advantages over competing baselines. The resulting images are not only richly detailed but also visually pleasing, opening new avenues for artistic creation.
Chinese Translation
尽管文本到图像的生成模型产生了令人印象深刻的结果,但它们在生成密集细节的高分辨率(HR)图像方面仍然存在困难。目前的文献通过低到高分辨率的方法来解决这个问题。首先生成一幅低分辨率(LR)图像,然后使用LR图像作为额外线索生成其上采样版本。在本文中,我们提出了联合潜在轨迹(JoLT)。JoLT在每个采样步骤中使用两个流共同去噪LR和HR潜在图像。LR潜在图像控制整体布局,而HR潜在图像控制细节。我们将两个分支互相连接,以共同整合它们的信息。我们对我们的方法进行了广泛的验证,展示了其相对于竞争基线的优势。生成的图像不仅细节丰富,而且视觉上令人愉悦,为艺术创作开辟了新的途径。
cs.CV / 87 / 2608.15404

CBX-Bench: A Human-Aligned MLLM Council for Benchmarking Concept Bottleneck Model Explanations

CBX-Bench:一个人类对齐的多模态大语言模型委员会,用于基准测试概念瓶颈模型解释
Karadag, Yusuf Meric, Oklan, Gulay, Cagliyan, Seref Baris, Ozdemir, Umut, Akbas, Emre
Abstract
Concept Bottleneck Models (CBMs) are designed to make visual classification interpretable by expressing predictions through human-understandable concepts. Although interpretability is the central motivation for CBMs, they are still largely evaluated as predictive models by downstream classification accuracy, supplemented by isolated qualitative examples. This highlights a pressing need for quantitative measures, a challenge complicated by the infeasibility of ground-truth concept annotation at scale and the open nature of concept lists due to a lack of consensus. To fill this gap, we develop a multimodal large language model (MLLM) council that, given an image and its CBM explanation, produces an explanation quality score. To ground and validate the council, we first conduct a human study to establish a ground-truth reference for CBM explanation quality: for an image, annotators compare explanations from two of LF-CBM, VLG-CBM, and CBM-Suite and choose the more useful one, or mark them as equally good or equally bad, yielding 2700 judgments over 900 image-comparison items on CUB-200, ImageNet-100, and Places365. Against this human reference, our five-model council, consisting of open-weight MLLMs, recovers over 70% of strict human preference rankings, rising to 83% on items where human annotators unanimously agree. Building on this validated council, we introduce CBX-Bench, a public benchmark and leaderboard: authors of new CBMs can submit their model's explanations, and CBX-Bench scores them with the council and maintains dataset-level rankings of explanation quality. CBX-Bench thus provides a human-aligned, scalable evaluation of CBM explanations beyond accuracy and isolated qualitative examples. The benchmark is available at https://github.com/meric-karadag/cbx-bench.
Chinese Translation
概念瓶颈模型(CBMs)旨在通过人类可理解的概念表达预测,从而使视觉分类具有可解释性。尽管可解释性是CBMs的核心动机,但它们仍然主要通过下游分类准确性作为预测模型进行评估,并辅以孤立的定性示例。这突显了对定量测量的迫切需求,而这一挑战因大规模真实概念注释的不可行性以及由于缺乏共识而导致的概念列表的开放性而变得复杂。为填补这一空白,我们开发了一个多模态大语言模型(MLLM)委员会,该委员会在给定图像及其CBM解释的情况下,生成一个解释质量评分。为了为该委员会提供基础和验证,我们首先进行了一项人类研究,以建立CBM解释质量的真实参考:对于一幅图像,注释者比较LF-CBM、VLG-CBM和CBM-Suite中的两个解释,并选择更有用的一个,或者标记它们为同样好或同样差,从而在CUB-200、ImageNet-100和Places365上产生了2700个判断,涉及900个图像比较项目。与这一人类参考相比,我们的五模型委员会,由开放权重的MLLM组成,恢复了超过70%的严格人类偏好排名,在人类注释者一致同意的项目上这一比例上升至83%。基于这一经过验证的委员会,我们推出了CBX-Bench,一个公共基准和排行榜:新的CBM作者可以提交其模型的解释,CBX-Bench将其与委员会进行评分,并维护解释质量的数据库级排名。因此,CBX-Bench提供了一种超越准确性和孤立定性示例的人类对齐的可扩展CBM解释评估。该基准可在 https://github.com/meric-karadag/cbx-bench 获取。
cs.CV / 88 / 2608.15419

ArtLang: Structured Language-to-Kinematics Grounding for Articulated 3D Actuation

ArtLang:用于关节式 3D 驱动的结构化语言与运动学基础
Yuan, Sylvia, Wang, Dan, Ramamoorthi, Ravi, Cui, Xinrui
Abstract
Articulated-object reconstructions recover explicit geometry and kinematics, but their parts often remain semantically anonymous and must be controlled through part indices and numerical joint parameters. We present ArtLang, a framework for open-vocabulary language control of persistent reconstructed articulated assets. ArtLang represents an asset as a semantic-kinematic articulation graph and augments its surface with language features and graph-constrained motion. Open-vocabulary proposals are bound to reconstructed parts while allowing uncertain parts to remain unnamed. A typed parser converts a command into a directive graph containing referring expressions, actions, magnitudes, reference frames, and relations. We then solve a global graph-to-graph grounding problem that jointly reasons about semantic, spatial, relational, and kinematic compatibility, with support for null assignments and abstention under ambiguity. Accepted directives are converted into continuous joint targets within the observed motion range and executed through forward kinematics. Experiments on synthetic reconstructions, mesh-based assets, and real captures demonstrate reliable language grounding and continuous articulated control across repeated parts, spatial references, relational commands, and ambiguous instructions.
Chinese Translation
关节物体重建恢复了明确的几何形状和运动学,但其部件往往在语义上保持匿名,必须通过部件索引和数值关节参数进行控制。我们提出了 ArtLang,一个用于持久重建关节资产的开放词汇语言控制框架。ArtLang 将资产表示为语义-运动学关节图,并通过语言特征和图约束运动增强其表面。开放词汇提案绑定到重建的部件,同时允许不确定的部件保持未命名。一个类型化解析器将命令转换为包含指称表达、动作、大小、参考框架和关系的指令图。然后,我们解决一个全局图到图的基础问题,该问题共同推理语义、空间、关系和运动学的兼容性,并支持在模糊情况下的空分配和弃权。接受的指令被转换为观察到的运动范围内的连续关节目标,并通过正向运动学执行。在合成重建、基于网格的资产和真实捕获的实验中,展示了可靠的语言基础和跨重复部件、空间参考、关系命令和模糊指令的连续关节控制。
cs.CV / 89 / 2608.15420

HistReNeRF: Historic Image Relocalisation within Contemporary Neural Radiance Field Reconstructions

HistReNeRF:在当代神经辐射场重建中进行历史图像重定位
Hughes, Benjamin T., James, Stuart
Abstract
Relocalising archival photographs within a contemporary scene model is challenging because historic and modern views can differ in photographic appearance, visible objects, and spatial layout. Therefore, we present HistReNeRF, a framework that estimates the 6-DoF pose of a historic photograph by matching adapted DINOv2 patch features to candidate rays sampled from a contemporary Neural Radiance Field (NeRF) reconstruction. The continuous representation of a NeRF provides a queryable scene interface from which candidate rays can be sampled and matched, enabling domain adaptation between historic photography and contemporary images directly in the feature representation used for localisation. We evaluate embedding-space-based domain adaptation against pixel-space methods on a new cross-temporal dataset comprising 10,545 contemporary street-level images and 230 archival photographs from three European landmarks. Embedding-space adaptation reduces translation and rotation errors by an average of 11% and 16%, respectively, across the three scenes. These results show that neural scene relocalisation provides a natural interface for feature-space adaptation, reducing cross-temporal appearance shift without modifying the query image. Code and dataset at https://github.com/ARTUROLab/HistReNeRF.
Chinese Translation
在当代场景模型中重定位档案照片具有挑战性,因为历史与现代视图在摄影外观、可见物体和空间布局上可能存在差异。因此,我们提出了HistReNeRF,一个通过将适配后的DINOv2补丁特征与从当代神经辐射场(NeRF)重建中采样的候选光线进行匹配,来估计历史照片的六自由度(6-DoF)姿态的框架。NeRF的连续表示提供了一个可查询的场景接口,从中可以采样和匹配候选光线,实现历史摄影与当代图像之间的领域适应,直接在用于定位的特征表示中进行。我们在一个新的跨时间数据集上评估了基于嵌入空间的领域适应与基于像素空间的方法,该数据集包含10,545张当代街景图像和来自三个欧洲地标的230张档案照片。嵌入空间适应在三个场景中平均减少了11%的平移误差和16%的旋转误差。这些结果表明,神经场景重定位为特征空间适应提供了一个自然的接口,减少了跨时间外观的变化,而无需修改查询图像。代码和数据集可在 https://github.com/ARTUROLab/HistReNeRF 获取。
cs.CV / 90 / 2608.15425

NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models

NumerosityVLM:一个以认知为灵感的基准,用于解释视觉-语言模型中的数量表示
Fu, Yiming, Li, Fangjun, Liu, Xiujin, Ma, Ruidong, Yu, Hang, Lu, Zhichen, He, Kanwei, Di Nuovo, Alessandro, Cangelosi, Angelo, Shangguan, Zhegong
Abstract
Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants before language acquisition, remains poorly understood in current models, as existing counting benchmarks entangle numerosity with correlated visual factors. We introduce a cognitively inspired diagnostic benchmark, NumerosityVLM, comprising 10,800 synthetic images across six controlled conditions. The benchmark orthogonally manipulates object size, spatial arrangement, and numerosity, while progressively ablating texture, shape, and color. Evaluating seven VLMs in a zero-shot setting, multi-factor analysis reveals that model architecture explains the largest proportion of performance variance (partial $\omega^{2}=0.325$), far exceeding visual conditions. Layer-wise probing further shows that linearly separable numerosity signals consistently emerge at early stages of the vision encoder, while performance differences across evaluated models are primarily associated with the language model component. Code and data are publicly available at https://github.com/fuy3/NumerosityVLM-Benchmark, and https://huggingface.co/datasets/fuy3/NumerosityVLM.
Chinese Translation
视觉-语言模型(VLMs)在高层次的多模态任务中表现出色,但数量感知这一认知能力在语言习得之前就出现在人类婴儿中,目前的模型对此仍然理解不足,因为现有的计数基准将数量与相关的视觉因素混淆在一起。我们引入了一个以认知为灵感的诊断基准,NumerosityVLM,包含10,800幅合成图像,涵盖六种受控条件。该基准正交地操控物体大小、空间排列和数量,同时逐步消融纹理、形状和颜色。在零样本设置下评估七个VLM,多个因素分析显示模型架构解释了性能方差的最大比例(部分 $ ext{ω}^{2}=0.325$),远超视觉条件。逐层探测进一步表明,线性可分的数量信号在视觉编码器的早期阶段始终出现,而评估模型之间的性能差异主要与语言模型组件相关。代码和数据可在 https://github.com/fuy3/NumerosityVLM-Benchmark 和 https://huggingface.co/datasets/fuy3/NumerosityVLM 上公开获取。
cs.CV / 91 / 2608.15452

Spatially-Grounded Flow Matching: Structured Source Distributions for Image Generation

空间基础流匹配:用于图像生成的结构化源分布
Zarei, Arman, Kalayeh, Mahdi M.
Abstract
Current flow matching models learn to transport the source i.i.d. Gaussian noise into the target distribution of natural images, yet this source distribution carries no notion of spatial structure. Images however are fundamentally local since nearby pixels are strongly correlated. By sampling the noise independently, we hypothesize that models are implicitly encouraged to exploit less noisy neighbors as context during training, partially bypassing the need to properly learn the true local structure of images. The source distribution, in other words, works against the inductive bias of the image domain. To ameliorate this design discrepancy, we propose StructFlow which encodes spatial locality directly into the source by having the pixels within a small region share a common noise component. This structured source produces transport paths that are geometrically aligned with image regions - enabling properties that generic flow matching struggles to provide: fine-grained local editing that naturally respects boundaries, robust structure preservation, and smooth semantic interpolation between images. We show that these benefits also extend to large pre-trained models, demonstrating that StructFlow can even be incorporated through a lightweight post-training phase. Comprehensive experiments on multiple datasets, in unconditional, class and text-conditioned regimes, using different diffusion transformer architectures confirm that StructFlow not only offers competitive image generation quality, but also significantly improves localized controllable re-synthesis.
Chinese Translation
当前的流匹配模型学习将源独立同分布的高斯噪声传输到自然图像的目标分布中,然而这种源分布并不包含空间结构的概念。然而,图像在本质上是局部的,因为相邻像素之间存在强相关性。通过独立采样噪声,我们假设模型在训练过程中被隐性地鼓励利用较少噪声的邻居作为上下文,从而部分绕过了正确学习图像真实局部结构的需求。换句话说,源分布与图像领域的归纳偏差相悖。为了改善这种设计差异,我们提出了StructFlow,它通过让小区域内的像素共享一个共同的噪声成分,直接将空间局部性编码到源中。这种结构化源产生的传输路径在几何上与图像区域对齐——实现了通用流匹配难以提供的特性:自然尊重边界的细粒度局部编辑、稳健的结构保持,以及图像之间平滑的语义插值。我们展示了这些优势也扩展到大型预训练模型,证明StructFlow甚至可以通过轻量级的后训练阶段进行整合。在多个数据集上的全面实验,包括无条件、类别和文本条件的设置,使用不同的扩散变换器架构,确认了StructFlow不仅提供了具有竞争力的图像生成质量,而且显著改善了局部可控的再合成。
cs.CV / 92 / 2608.15456

AlignJEPA: Predictive Vision-Language Alignment for Remote Sensing Foundation Models

AlignJEPA:用于遥感基础模型的预测性视觉-语言对齐
Hossain, Md Aminur, Vaghasiya, Omkumar, Dwivedi, Rajeev Ranjan, Kurmi, Vinod, Banerjee, Biplab
Abstract
Remote sensing (RS) foundation models provide transferable Earth observation representations across sensors, resolutions, and geographies, yet most remain weakly aligned with natural language, limiting natural-language archive search, image-text retrieval, and question-conditioned analysis. We propose AlignJEPA, a JEPA-inspired predictive vision-language alignment framework for remote sensing foundation models. AlignJEPA uses a pretrained AnySat visual encoder and a RemoteCLIP text encoder while training only a lightweight predictive alignment network. Instead of relying on global image--text contrastive alignment alone, the framework predicts remote-sensing text embeddings from masked visual foundation-model tokens. Its mask-aware multi-scale predictive aligner aggregates visible tokens at fine, regional, and global scales, jointly models them with a cross-scale Transformer, and projects the resulting representation into the text space using learned query pooling. Training combines semantic prediction with bidirectional contrastive retrieval. We train and evaluate AlignJEPA on BigEarthNet.txt for natural-language Sentinel retrieval, evaluate cross-dataset adaptation on RSICD, and use RSVQA only as a closed-set representation probe. AlignJEPA provides a parameter-efficient route for aligning Earth observation foundation models with language.
Chinese Translation
遥感(RS)基础模型提供了跨传感器、分辨率和地理区域的可转移地球观测表示,然而大多数模型与自然语言的对齐仍然较弱,这限制了自然语言档案搜索、图像-文本检索和基于问题的分析。我们提出了AlignJEPA,这是一种受JEPA启发的遥感基础模型的预测性视觉-语言对齐框架。AlignJEPA使用预训练的AnySat视觉编码器和RemoteCLIP文本编码器,同时仅训练一个轻量级的预测对齐网络。该框架不仅依赖于全局图像-文本对比对齐,而是从被遮蔽的视觉基础模型标记中预测遥感文本嵌入。其掩码感知的多尺度预测对齐器在细粒度、区域和全局尺度上聚合可见标记,使用跨尺度Transformer共同建模,并通过学习的查询池化将得到的表示投影到文本空间。训练结合了语义预测与双向对比检索。我们在BigEarthNet.txt上训练和评估AlignJEPA以进行自然语言的Sentinel检索,在RSICD上评估跨数据集适应性,并仅使用RSVQA作为封闭集表示探测器。AlignJEPA为将地球观测基础模型与语言对齐提供了一条高效的参数路径。
cs.CV / 93 / 2608.15517

GLaQ: Grounding Latent Queries in Visual Evidence for Multimodal Reasoning

GLaQ:基于视觉证据的潜在查询基础多模态推理
Yang, Zesheng, Zhang, Lingling, Zhang, Xinyu, Zhang, Cheng, Li, Pengyu, Wang, Heng, Wu, Lin
Abstract
Chain-of-thought reasoning has substantially improved the problem-solving capabilities of multimodal large language models. Fine-grained visual evidence, however, remains difficult to preserve and reuse across text-based reasoning steps. To address this limitation, tool-augmented thinking-with-images methods maintain visual access externally by revisiting or manipulating the image, but require predefined tools and additional inference-time processing. As an internal alternative, continuous visual latent reasoning retains intermediate computation in hidden states. However, its prevailing autoregressive construction makes each latent state depend on its predecessors, so later states may repeat information already present in the latent sequence rather than capture complementary visual details. We introduce GLaQ, a grounded latent-query framework that replaces sequential latent rollout with a fixed set of context-conditioned queries grounded in the original visual tokens. The grounded queries are reinjected for answer generation, providing direct and coordinated access to source visual evidence. We train GLaQ with localized-view supervision followed by reinforcement learning under task-level rewards. Across five benchmarks for fine-grained visual understanding and perception, GLaQ-7B gains 5.99--9.66\% over its base model and leads all compared visual latent methods, suggesting that direct query-to-image grounding can recover localized evidence from the full image without external visual operations or autoregressive latent rollouts.
Chinese Translation
思维链推理显著提升了多模态大型语言模型的问题解决能力。然而,细粒度的视觉证据在基于文本的推理步骤中仍然难以保留和重用。为了解决这一限制,增强工具的图像思维方法通过重新访问或操控图像来保持外部的视觉访问,但需要预定义的工具和额外的推理时间处理。作为一种内部替代方案,连续的视觉潜在推理在隐藏状态中保留中间计算。然而,其主流的自回归结构使得每个潜在状态依赖于其前驱状态,因此后续状态可能重复潜在序列中已存在的信息,而不是捕捉互补的视觉细节。我们提出了GLaQ,一个基础的潜在查询框架,它用一组固定的上下文条件查询替代了顺序潜在展开,这些查询基于原始视觉标记进行基础。基础查询被重新注入以生成答案,提供对源视觉证据的直接和协调访问。我们通过局部视图监督训练GLaQ,然后在任务级奖励下进行强化学习。在五个细粒度视觉理解和感知的基准测试中,GLaQ-7B相较于其基础模型提升了5.99%至9.66%,并在所有比较的视觉潜在方法中领先,表明直接的查询与图像的基础可以从完整图像中恢复局部证据,而无需外部视觉操作或自回归潜在展开。
cs.CV / 94 / 2608.15522

Efficient Audio-Visual Generation via Synchrony-Aware Cross-Modal Sparse Attention

通过同步感知的跨模态稀疏注意力实现高效的音视频生成
Gao, Shengchuan, Hu, Teng, Feng, Bohao, Li, Luchen, Wang, Wenqiang, Deng, Hongqian, Yi, Ran
Abstract
Recent audio-visual generation models can synthesize synchronized video and sound in a unified diffusion process, but their inference cost remains high because long video token sequences require repeated attention computation across denoising steps.A variety of acceleration techniques have been developed for video generation models, including low-bit quantization, attention sparsification, and feature caching.However, since these methods are originally designed for video generation, directly applying them to audio-visual models overlooks the interactions between the audio and video branches and may therefore disrupt audio-video synchronization.We present a synchronization-aware acceleration framework for efficient audio-visual generation.Our key observation is that bidirectional audio-video cross-attention reveals structured interactions between the two branches, with high responses often concentrated on a few sound-related visual and temporal regions.Guided by this interaction pattern, we introduce a protected sparse attention strategy that preserves high-fidelity computation for synchronization-critical tokens while sparsifying redundant attention interactions.By explicitly accounting for cross-modal dependence during acceleration, our method improves inference efficiency while keeping video quality, audio quality, and audio-video synchronization.
Chinese Translation
近期的音视频生成模型能够在统一的扩散过程中合成同步的视频和声音,但其推理成本仍然较高,因为长视频令牌序列需要在去噪步骤中重复计算注意力。为了加速视频生成模型,已经开发出多种加速技术,包括低位量化、注意力稀疏化和特征缓存。然而,由于这些方法最初是为视频生成设计的,直接将其应用于音视频模型忽视了音频和视频分支之间的交互,可能会破坏音视频同步。我们提出了一种同步感知的加速框架,以实现高效的音视频生成。我们的关键观察是,双向音视频交叉注意力揭示了两个分支之间的结构化交互,高响应通常集中在少数与声音相关的视觉和时间区域。在这种交互模式的指导下,我们引入了一种保护性稀疏注意力策略,保留了对同步关键令牌的高保真计算,同时稀疏化冗余的注意力交互。通过在加速过程中明确考虑跨模态依赖,我们的方法提高了推理效率,同时保持了视频质量、音频质量和音视频同步。
cs.CV / 95 / 2608.15537

EA-LiteUNet: An Edge-Adaptive and Resource-Efficient U-Net for Boundary-Sensitive Dermoscopic Image Segmentation

EA-LiteUNet:一种边缘自适应和资源高效的U-Net,用于边界敏感的皮肤镜图像分割
Jiangtao, Wang, Ruhaiyem, Nur Intan Raihana, Panpan, Fu, Yu, Yang, Yan, Huang
Abstract
Accurate boundary delineation remains a persistent challenge in dermoscopic image segmentation because of blurred lesion margins, heterogeneous textures, and complex background artifacts. From a signal-processing perspective, lesion boundaries represent high-frequency components that are highly susceptible to aliasing, noise amplification, and information loss. Consequently, repeated downsampling and feature transformations in conventional convolutional architectures often lead to severely degraded boundary representations. To address these limitations, we propose EA-LiteUNet, an edge-adaptive and computationally efficient U-Net variant specifically designed for boundary-sensitive medical image segmentation. The architecture integrates three core mechanisms: (1) boundary-aware representation learning to suppress aliasing and preserve high-frequency structural details; (2) attention-guided feature modulation to selectively enhance boundary-relevant responses across multi-scale features; and (3) a resource-adaptive inference strategy to dynamically balance segmentation accuracy and computational efficiency. Extensive evaluations across three public dermoscopic datasets demonstrate that EA-LiteUNet consistently achieves superior boundary precision. Specifically, on the ISIC 2018 dataset, the method significantly reduces the 95% Hausdorff Distance (HD95) to 12.89 pixels while maintaining a robust Dice score of 92.08%. Notably, this strong performance is achieved with an ultralightweight configuration of merely 0.29M parameters and 1.17 GFLOPs. Ablation studies further validate the complementary effects of these components, confirming their contribution to enhanced boundary fidelity and stable optimization.
Chinese Translation
由于病变边缘模糊、纹理异质性和复杂的背景伪影,准确的边界划分在皮肤镜图像分割中仍然是一个持续的挑战。从信号处理的角度来看,病变边界代表了高频成分,这些成分对混叠、噪声放大和信息丢失高度敏感。因此,传统卷积架构中的重复下采样和特征变换往往导致边界表示严重退化。为了解决这些限制,我们提出了EA-LiteUNet,一种专门为边界敏感的医学图像分割设计的边缘自适应和计算高效的U-Net变体。该架构集成了三种核心机制:(1)边界感知表示学习,以抑制混叠并保留高频结构细节;(2)注意力引导的特征调制,以选择性地增强跨多尺度特征的边界相关响应;(3)资源自适应推理策略,以动态平衡分割精度和计算效率。在三个公共皮肤镜数据集上的广泛评估表明,EA-LiteUNet始终实现了优越的边界精度。具体而言,在ISIC 2018数据集上,该方法显著将95% Hausdorff距离(HD95)降低至12.89像素,同时保持强健的Dice得分为92.08%。值得注意的是,这一强劲的表现是在仅有0.29M参数和1.17 GFLOPs的超轻量配置下实现的。消融研究进一步验证了这些组件的互补效应,确认了它们对增强边界保真度和稳定优化的贡献。
cs.CV / 96 / 2608.15539

CrossView: Can Vision-Language Models Reason Across Cameras?

CrossView:视觉-语言模型能否跨摄像头进行推理?
Shah, Sahil, Sharan, S P, Goel, Harsh, Pasula, Manvik, Hebbalae, Adithya, Choi, Minkyu, Chinchali, Sandeep P.
Abstract
Video understanding benchmarks have long centered on single-camera settings, where modern multi-modal language models achieve strong performance across image and video tasks. Yet, the real world runs on multi-camera networks: autonomous vehicles, security systems, and robots all gather data across many simultaneous views. We argue that this is not simply "more" of the single-camera problem; it is fundamentally different. Multi-camera reasoning requires handling context that scales with the number of views, resolving occlusions visible from only a subset of cameras, judging which views matter, and integrating evidence across perspectives that may overlap or diverge. Current models struggle with exactly these challenges, yet no benchmark systematically targets them. We introduce CrossView, a multi-camera video question-answering benchmark spanning autonomous driving, security surveillance, egocentric/exocentric video, and robotics. Evaluation of proprietary models, such as GPT-5.2, and open-source models, like Qwen3-VL, reveals consistently low accuracy, with open-source models trailing by a wide margin. Performance scales strongly with a model's ability to jointly process multiple viewpoints, positioning CrossView as a rigorous benchmark for multi-camera video. We open-source our code and dataset at https://utaustin-swarmlab.github.io/CrossView.
Chinese Translation
视频理解基准长期以来集中于单摄像头设置,在这种情况下,现代多模态语言模型在图像和视频任务中表现出色。然而,现实世界依赖于多摄像头网络:自动驾驶汽车、安全系统和机器人都在多个同时视角下收集数据。我们认为,这不仅仅是单摄像头问题的“更多”情况;它在根本上是不同的。多摄像头推理需要处理与视角数量成比例的上下文,解决仅从部分摄像头可见的遮挡问题,判断哪些视角重要,以及整合可能重叠或分歧的视角之间的证据。目前的模型在这些挑战上表现不佳,但没有基准系统性地针对这些问题。我们推出了CrossView,这是一个涵盖自动驾驶、安全监控、自我中心/外部中心视频和机器人技术的多摄像头视频问答基准。对专有模型(如GPT-5.2)和开源模型(如Qwen3-VL)的评估显示,准确率普遍较低,开源模型的表现差距更大。性能与模型同时处理多个视角的能力密切相关,使CrossView成为多摄像头视频的严格基准。我们将在https://utaustin-swarmlab.github.io/CrossView上开源我们的代码和数据集。
cs.CV / 97 / 2608.15555

RigidBench: Evaluating Rigid-Body Physics in Video Generation Models

RigidBench:评估视频生成模型中的刚体物理
Jain, Swarnim, Wu, Shangzhe
Abstract
Video models are increasingly used to predict what happens next in a scene, yet the metrics commonly used to compare their outputs say little about whether the predicted objects move correctly. Motion, geometry, identity, background stability, and visual similarity can fail independently, but whole-frame scores often mix these errors together. We introduce RigidBench, a simulator-grounded benchmark that compares a generated continuation with a reference rollout from the same initial frame and motion description. Its five rigid-body tasks vary objects, materials, viewpoints, and indoor and outdoor scenes, with per-frame masks, depth, 6-DoF trajectories, and contacts available for scoring. We evaluate eight models on the same 100 examples with ten measurements that keep these aspects separate. The resulting rankings depend strongly on what is measured: no model leads on all ten, and across model means, higher SSIM accompanies larger 3D trajectory error (r = 0.89). RigidBench also includes 5,000 training videos with exact simulator state, which we use to fine-tune and analyze Wan 2.2 TI2V-5B. Full fine-tuning reduces 3D trajectory error by about 20% with almost no change in SSIM, while teacher-forced probes and targeted interventions show that object position is represented throughout Wan's diffusion transformer and used by its denoising computation.
Chinese Translation
视频模型越来越多地用于预测场景中接下来会发生什么,但常用的比较输出的指标对于预测对象是否正确移动几乎没有帮助。运动、几何形状、身份、背景稳定性和视觉相似性可能会独立失败,但整个帧的评分往往将这些错误混合在一起。我们引入了 RigidBench,一个基于模拟器的基准,比较生成的延续与来自相同初始帧和运动描述的参考结果。其五个刚体任务涉及对象、材料、视角以及室内和室外场景,提供逐帧的掩码、深度、六自由度轨迹和接触信息以供评分。我们在相同的100个示例上评估了八个模型,并使用十个测量指标将这些方面分开。结果排名强烈依赖于所测量的内容:没有模型在所有十个指标上领先,模型均值显示,较高的结构相似性指数(SSIM)伴随着更大的三维轨迹误差(r = 0.89)。RigidBench 还包括5000个具有精确模拟器状态的训练视频,我们用它来微调和分析 Wan 2.2 TI2V-5B。完全微调将三维轨迹误差减少了约20%,几乎没有改变 SSIM,而教师强制探测和有针对性的干预表明,物体位置在 Wan 的扩散变换器中得到了表示,并被其去噪计算所使用。
cs.CV / 98 / 2608.15574

Catching Hallucinated Citations in Video-LLM Question Answering: A Self-Verification Pipeline and Verifier Ablation Study

捕捉视频-大语言模型问答中的虚假引用:自我验证管道和验证器消融研究
Kumar, Yogesh
Abstract
Video question answering systems built on vision-language models often produce timestamped claims with high confidence even when unsupported by the cited frame. This deceptive hallucination arises because timestamps imply grounding without ensuring correctness, increasing user trust but not accuracy. We introduce a pipeline that closes this loop. A retrieval-augmented language model drafts answers with per-claim timestamp citations, and each cited frame is independently re-examined before being shown to the user. We compare against a plain baseline and ablate three verification designs, evaluated on both Apple Silicon (MLX) and Google Colab (HF Transformers, CUDA). Directly asking the vision model whether a frame supports a claim fails completely (0% catch rate on 40 claims) due to sycophancy. Blind re-captioning plus a general LLM judge improves results but is unstable, oscillating between 0% and 100% flagged depending on prompt phrasing. Replacing that judge with a small natural language inference model yields a stable, interpretable verifier that catches 79% of fabricated claims on adversarial false-premise questions while leaving true claims untouched. We release the full pipeline, evaluation harness, and implementations for both Apple Silicon and Colab. Code is available at https://github.com/yogesh-iitj/grounded-video-qa.
Chinese Translation
基于视觉-语言模型的视频问答系统通常会生成带有高置信度的时间戳声明,即使这些声明并未得到引用帧的支持。这种误导性的幻觉产生是因为时间戳暗示了基础支持,但并未确保其正确性,从而增加了用户的信任感,但并未提高准确性。我们提出了一种闭环管道。增强检索的语言模型为每个声明草拟答案并附上时间戳引用,而每个被引用的帧在展示给用户之前都会被独立重新审查。我们与一个普通基线进行了比较,并消融了三种验证设计,评估在Apple Silicon (MLX) 和 Google Colab (HF Transformers, CUDA) 上的表现。直接询问视觉模型某个帧是否支持某个声明完全失败(在40个声明上捕获率为0%),这是由于谄媚效应。盲重标注加上一个通用的LLM评判者改善了结果,但不稳定,依据提示措辞在0%和100%之间波动。用一个小型自然语言推理模型替换该评判者,得到了一个稳定且可解释的验证器,能够在对抗性虚假前提问题上捕捉到79%的虚假声明,同时保持真实声明不受影响。我们发布了完整的管道、评估工具和在Apple Silicon和Colab上的实现。代码可在 https://github.com/yogesh-iitj/grounded-video-qa 获取。
cs.CV / 99 / 2608.15583

PoseAdapter: Dual-Stream 2.5D Controllable Image Generation for Complex Multi-Object Scenes

PoseAdapter:复杂多对象场景的双流2.5D可控图像生成
Chi, Yufeng, Ma, Huimin, Gao, Fan, Niu, Zhice, Li, Keqin, Li, Jianmin
Abstract
While Text-to-Image (T2I) diffusion models have achieved remarkable success, precise spatial and orientational control in multi-object scenes remains a persistent challenge. Existing methods either rely on computationally expensive dense 3D maps or suffer from severe attribute leakage and "cut-and-paste" artifacts. To address these limitations, we propose PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation. Instead of dense spatial maps, it establishes precise spatial-angular anchors using an efficient condition layout: individual object captions, 2D bounding boxes, and 3D angles. To resolve the generative trade-off between strict instance isolation and global coherence, we introduce a Context-Aware Dual-Stream Representation. By injecting local object tokens and relation-enriched scene tokens into the visual stream of modern MM-DiT architectures via parallel masked and unmasked pathways, PoseAdapter eliminates attribute leakage while preserving natural inter-object relationships and scene-level coherence. To support this paradigm, we construct OrientLayout, a high-quality dataset featuring standardized 2.5D annotations and instance-level decoupled semantics. Extensive experiments demonstrate that PoseAdapter outperforms state-of-the-art baselines in spatial accuracy, orientational precision, and multi-object visual fidelity. Code and dataset will be available at https://github.com/cyf23/PoseAdapter.
Chinese Translation
尽管文本到图像(Text-to-Image, T2I)扩散模型取得了显著成功,但在多对象场景中实现精确的空间和方向控制仍然是一个持续的挑战。现有的方法要么依赖于计算成本高昂的密集3D地图,要么遭受严重的属性泄露和“剪切粘贴”伪影。为了解决这些局限性,我们提出了PoseAdapter,一个轻量级的高保真2.5D可控图像生成框架。它不是使用密集的空间地图,而是通过高效的条件布局建立精确的空间-角度锚点:单个对象的描述、2D边界框和3D角度。为了在严格的实例隔离和全局一致性之间解决生成权衡,我们引入了上下文感知双流表示。通过将局部对象标记和关系丰富的场景标记注入现代MM-DiT架构的视觉流中,PoseAdapter通过并行的掩码和非掩码路径消除了属性泄露,同时保留了自然的对象间关系和场景级一致性。为了支持这一范式,我们构建了OrientLayout,一个高质量的数据集,具有标准化的2.5D注释和实例级解耦语义。大量实验表明,PoseAdapter在空间准确性、方向精度和多对象视觉保真度方面超越了最先进的基线。代码和数据集将发布在 https://github.com/cyf23/PoseAdapter。
cs.CV / 100 / 2608.15605

AlloEgo-VLM: Disambiguating Allocentric and Egocentric Reference Frames in Vision-Language Models

AlloEgo-VLM:在视觉-语言模型中消除外部和自我参考框架的歧义
Chen, Kuan-Lin, Wei, Tzu-Ti, Liao, Chao-Chi, Tseng, Yu-Chee, Chen, Jen-Jee
Abstract
This study investigates the challenge of ambiguity faced by Vision-Language Models (VLMs) in understanding spatial semantics. Spatial cognition, shaped by cognitive psychology, spatial science, and cultural context, often assigns directionality to objects. However, natural language descriptions of spatial relations frequently omit explicit reference frames, leading to semantic ambiguity and potentially serious errors for embodied AI robots. Existing VLMs, due to insufficient training on reference frames and object orientations, often produce inconsistent responses. To address this issue, we construct a new dataset, AlloEgo-View, comprising (image, query, view-specific answer) triplets that capture key object relations from both allocentric and egocentric perspectives. The view-specific descriptions follow a structured spatial representation that annotate detailed scene descriptions, reference and target objects, their orientations, reference frames, and view types. Building on AlloEgo-View, we develop AlloEgo-VLM, a framework to disambiguate allocentric and egocentric reference frames, even under ambiguous queries, and to be easily integrated into existing VLMs via supervised fine-tuning. Furthermore, we deploy our framework onto an embodied robotic platform within NVIDIA Isaac Sim to validate its real-world feasibility in open-ended object searching tasks. Experiments highlight the limitations of current VLMs in handling view-specific queries and demonstrate the strong disambiguation ability of AlloEgo-VLM.
Chinese Translation
本研究探讨了视觉-语言模型(VLMs)在理解空间语义时面临的歧义挑战。空间认知受认知心理学、空间科学和文化背景的影响,通常会为物体分配方向性。然而,自然语言对空间关系的描述往往省略了明确的参考框架,导致语义歧义,并可能对具身人工智能机器人造成严重错误。现有的VLMs由于在参考框架和物体方向上的训练不足,常常产生不一致的响应。为了解决这一问题,我们构建了一个新的数据集AlloEgo-View,包含(图像,查询,视角特定答案)三元组,捕捉从外部和自我视角出发的关键物体关系。视角特定的描述遵循结构化的空间表示,注释详细的场景描述、参考物体和目标物体、它们的方向、参考框架和视角类型。在AlloEgo-View的基础上,我们开发了AlloEgo-VLM,一个框架,用于消除外部和自我参考框架的歧义,即使在模糊查询下,也能轻松通过监督微调集成到现有的VLMs中。此外,我们将该框架部署到NVIDIA Isaac Sim的具身机器人平台上,以验证其在开放式物体搜索任务中的现实可行性。实验突显了当前VLMs在处理视角特定查询时的局限性,并展示了AlloEgo-VLM的强大消歧能力。
cs.CV / 101 / 2608.15614

EgoGazeLite: On-Device Egocentric Gaze Prediction for Token-Efficient Multimodal LLM Video Input

EgoGazeLite:用于令牌高效多模态大语言模型视频输入的设备端自我中心注视预测
Stoiber, Matteo, Lassen, Niels Buus
Abstract
The use of multimodal LLMs (MLLMs) for egocentric video understanding with wearable devices is constrained by the token budget. Memory and compute cost scale with the number of visual tokens, and high-resolution video quickly becomes expensive to transmit and process at scale. Prior work (GazeLLM) addresses this by cropping the video around the camera wearer's gaze. This reduces the number of visual tokens by about tenfold while maintaining or improving the quality of full-resolution descriptions. However, this compression strategy depends on dedicated eye-tracking hardware, which is unavailable on consumer smart glasses. Building a software-only substitute poses a joint constraint: the predictor must be accurate enough to preserve downstream description quality, yet light enough to run on-device, within the power and compute budget of a smartphone. We address this with EgoGazeLite, a lightweight dual-process gaze predictor for egocentric video. Across two MLLMs, three automated metrics, and two LLM judges, predicted-gaze crops show no significant difference from ground-truth-gaze crops. Equivalence is confirmed in all ten cases. EgoGazeLite achieves this at 15.7M parameters, 6.71 GFLOPs, and runs the full gaze-and-crop pipeline end-to-end in real time (21.6 ms/frame) on consumer accelerator hardware. Together, these results remove the need for eye-tracking hardware for token-efficient, gaze-conditioned egocentric video understanding with MLLMs.
Chinese Translation
使用多模态大语言模型(MLLMs)进行可穿戴设备的自我中心视频理解受到令牌预算的限制。内存和计算成本与视觉令牌的数量成比例,高分辨率视频在大规模传输和处理时迅速变得昂贵。之前的研究(GazeLLM)通过围绕摄像机佩戴者的注视点裁剪视频来解决这个问题。这将视觉令牌的数量减少了约十倍,同时保持或改善全分辨率描述的质量。然而,这种压缩策略依赖于专用的眼动追踪硬件,而这种硬件在消费级智能眼镜上并不可用。构建一个仅软件替代方案面临着共同的限制:预测器必须足够准确,以保持下游描述质量,同时又要轻量化,以便在智能手机的功耗和计算预算内运行。我们通过EgoGazeLite解决了这个问题,这是一种轻量级的双过程注视预测器,用于自我中心视频。在两个MLLMs、三个自动化指标和两个LLM评审者的测试中,预测的注视裁剪与真实注视裁剪之间没有显著差异。在所有十个案例中都确认了等效性。EgoGazeLite在1570万参数和6.71 GFLOPs的情况下实现了这一点,并在消费级加速硬件上以实时(21.6毫秒/帧)运行完整的注视和裁剪管道。这些结果共同消除了对眼动追踪硬件的需求,使得在多模态大语言模型下进行令牌高效、注视条件的自我中心视频理解成为可能。
cs.CV / 102 / 2608.15647

Hierarchical Adaptive Feature Refinement Network for VHR Remote Sensing Image Segmentation

用于超高分辨率遥感图像分割的层次自适应特征细化网络
Cao, Shuaishuai, Tang, Meng, Peng, Shuwei, Liu, Xuan, Huang, Min, Chen, Jie, Niu, Jiacheng, Chen, Yong, Akpokodje, Edore, Lin, Hui
Abstract
Semantic segmentation of very-high-resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encoders, yet exploiting their multi-stage representations remains difficult. Nearby regions demand different balances between fine detail and semantic context, aggressive task-specific transformations perturb useful pretrained features, and conventional semantic supervision provides limited structural guidance. We present HAFR-Net, a progressive refinement framework that adaptively organizes and conservatively refines hierarchical representations instead of replacing them with a monolithic decoder transformation. Heterogeneity-Guided Stage-Adaptive Fusion (HG-SAF) predicts dense stage weights conditioned on local feature variation. A Frequency-Residual Adapter (FRA) then injects frequency information through a bounded, zero-initialized residual branch that keeps the fused representation as its reference. A Confusion-Aware Tri-Prior Decoder (CATP) finally regularizes the prediction with boundary, objectness, and training-derived class-relation cues. Under a matched Swin-B training and single-scale inference protocol, HAFR-Net attains 84.12%, 87.86%, 55.17%, and 67.70% mIoU on ISPRS Vaihingen, ISPRS Potsdam, LoveDA, and OpenEarthMap, improving the matched UPerNet baseline by 0.55, 0.95, 1.55, and 1.84 percentage points, respectively. Controlled analyses further show consistent spatial reweighting beyond content-only routing, improved boundary and thin-structure accuracy over matched spatial and spectral alternatives, and reduced confusion on pre-declared class pairs.
Chinese Translation
超高分辨率(VHR)遥感影像的语义分割越来越依赖于强大的预训练层次编码器,但利用其多阶段表示仍然困难。相邻区域在细节与语义上下文之间需要不同的平衡,激进的任务特定变换会干扰有用的预训练特征,而传统的语义监督提供的结构指导有限。我们提出了HAFR-Net,这是一种渐进细化框架,能够自适应地组织和保守地细化层次表示,而不是用单一的解码器变换来替代它们。异质性引导的阶段自适应融合(HG-SAF)根据局部特征变化预测密集的阶段权重。频率残差适配器(FRA)通过一个有界的、零初始化的残差分支注入频率信息,从而保持融合表示作为其参考。最后,混淆感知三先验解码器(CATP)通过边界、物体性和训练衍生的类别关系线索来规范化预测。在匹配的Swin-B训练和单尺度推理协议下,HAFR-Net在ISPRS Vaihingen、ISPRS Potsdam、LoveDA和OpenEarthMap上分别达到了84.12%、87.86%、55.17%和67.70%的mIoU,相较于匹配的UPerNet基线分别提高了0.55、0.95、1.55和1.84个百分点。控制分析进一步表明,超越仅基于内容的路由,保持一致的空间重加权,改善了与匹配的空间和光谱替代方案相比的边界和细结构准确性,并减少了对预声明类别对的混淆。
cs.CV / 103 / 2608.15651

Gaussian-JEPA: Joint-Embedding Predictive Learning for 3D Gaussian Splats

Gaussian-JEPA:用于3D高斯点云的联合嵌入预测学习
Ren, Bin, Ma, Qi, Li, Yue, Han, Zongyan, Li, Yidi, Fu, Yuqian, Anwer, Rao Muhammad, Gevers, Theo, Khan, Fahad Shahbaz, Khan, Salman
Abstract
3D Gaussian Splatting (3DGS) represents 3D content with anisotropic primitives that jointly encode geometry and appearance. Fixed-budget encoders consume sampled observations of Gaussian assets, so the same object may be observed through different primitive realizations. Existing self-supervised methods mainly reconstruct masked Gaussian attributes, tying supervision to one sampled realization and requiring an input-space decoder. Latent prediction offers an alternative, but its application to Gaussian tokens requires targets that accommodate coupled attributes and heterogeneous spatial support. We introduce Gaussian-JEPA, which predicts representations of held-out Gaussian token blocks from visible context. An online encoder processes the context, while a shared exponential-moving-average encoder supplies stop-gradient features for multi-scale targets. Complementary target projections and feature-space grounding provide latent supervision without reconstructing Gaussian attributes. We evaluate the features under Gaussian resampling, partial observations, and renderable shape completion, together with transfer to part segmentation and object classification. Compared with matched reconstruction pretraining, Gaussian-JEPA is more consistent across resampled inputs, retains more instance information under partial observations, and provides stronger frozen features for Gaussian completion. These results support latent prediction as an effective objective for reusable 3D Gaussian representations. Code is on the project page (https://amazingren.github.io/Gaussian-JEPA/).
Chinese Translation
3D高斯点云(3DGS)通过各向异性原语共同编码几何和外观,以表示3D内容。固定预算编码器消耗高斯资产的采样观测,因此同一对象可能通过不同的原语实现被观察到。现有的自监督方法主要重建被遮挡的高斯属性,将监督与一个采样实现绑定,并需要输入空间解码器。潜在预测提供了一种替代方案,但其在高斯标记上的应用需要能够容纳耦合属性和异构空间支持的目标。我们提出了Gaussian-JEPA,它从可见上下文中预测被保留的高斯标记块的表示。在线编码器处理上下文,而共享的指数移动平均编码器为多尺度目标提供停止梯度特征。互补目标投影和特征空间基础提供潜在监督,而无需重建高斯属性。我们在高斯重采样、部分观测和可渲染形状补全下评估这些特征,并将其转移到部件分割和对象分类。与匹配重建预训练相比,Gaussian-JEPA在重采样输入中更为一致,在部分观测下保留了更多的实例信息,并为高斯补全提供了更强的冻结特征。这些结果支持潜在预测作为可重用3D高斯表示的有效目标。代码可在项目页面(https://amazingren.github.io/Gaussian-JEPA/)找到。
cs.CV / 104 / 2608.15652

Scalable Black-Box Model Attribution for Images

可扩展的黑箱模型归因方法用于图像
Livne, Asaf, Jevnisek, Amir, Avidan, Shai
Abstract
The rapid proliferation of generative models raises the model attribution problem: given only an image, can we determine which model produced it? Existing methods have grown as elaborate as the generators they target, on the as- sumption that a more sophisticated model demands a more sophisticated attributor. We show it does not. RPA (Raw- Patch Attribution) attributes images in the strictest black- box setting with a lightweight CNN. Despite its simplicity, it attributes more models at higher accuracy than prior work, reaching 98.0% on 25-class DRAGON and 92.9% on 27- class OpenFake; it is data-efficient and runs at a cost inde- pendent of the number of candidate models; and it stays ro- bust to the compression, blur, and resizing images undergo in the wild. Training for closed-set attribution yields a ver- satile feature extractor: the same representation recovers model lineage without supervision, flags and groups unseen generators, and admits new models through few-shot adap- tation rather than retraining.
Chinese Translation
生成模型的快速普及引发了模型归因问题:仅凭一幅图像,我们能否确定是哪种模型生成了它?现有的方法随着其目标生成器的复杂性而变得愈加精细,假设更复杂的模型需要更复杂的归因器。我们证明并非如此。RPA(Raw-Patch Attribution)在最严格的黑箱环境下使用轻量级卷积神经网络(CNN)进行图像归因。尽管其简单性,RPA在更高的准确率下归因于更多模型,达到了25类DRAGON数据集的98.0%和27类OpenFake数据集的92.9%;它的数据效率高,运行成本与候选模型的数量无关;并且对在现实环境中经历的压缩、模糊和缩放图像保持稳健。针对闭集归因的训练产生了一个多功能特征提取器:相同的表示能够在没有监督的情况下恢复模型谱系,标记并分组未见过的生成器,并通过少量样本适应而非重新训练来接纳新模型。
cs.CV / 105 / 2608.15659

WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations

WorldRover:一个可扩展的合成视频数据引擎,用于带有丰富注释的世界探索
Xu, Xiaojie, Lin, Zhengyuan, Li, Runyi, Liu, Yihao, Zhang, Kaipeng, Ge, Yongtao
Abstract
Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene geometry, temporal correspondence and, for interactive models, control signals. Real capture can provide some of these signals, but dense geometry and long-range correspondence usually rely on estimation or specialised instrumentation. Rendering provides these quantities directly, yet existing synthetic resources rarely combine them on the same frames while also supporting controlled changes of viewpoint and appearance. We introduce WorldRover, a data engine for generating richly annotated, long-range explorations of artist-built environments. At its core, WorldRover-Engine is an Unreal Engine pipeline that executes and offline-renders minute-scale routes while preserving their full trajectories and scene geometry. The same exploration can be replayed from first-person, third-person, and 360 panoramic cameras under different environmental states. Using WorldRover-Engine, we construct WorldRover-10M, whose sequences pair RGB with metric depth, camera trajectories, and trajectory-derived action signals throughout each exploration. Third-person subsets additionally provide dense optical flow, long-range 2D/3D point tracks with visibility, and a character trajectory distinct from the camera trajectory. The engine can render a traversal from first-person, third-person and 360 panoramic viewpoints, under different environmental states or with a neutral white material, while preserving the route and scene geometry. WorldRover therefore turns long-horizon world exploration into a scalable data-generation problem, providing supervision for models that must build, maintain, and revisit coherent representations of an explorable world.
Chinese Translation
学习生成或重建可探索的世界需要的不仅仅是RGB视频,还包括相机运动、场景几何、时间对应关系,以及对于交互模型而言的控制信号。真实捕捉可以提供其中的一些信号,但密集几何和长距离对应关系通常依赖于估计或专门的仪器。渲染可以直接提供这些量,然而现有的合成资源很少在同一帧上结合这些信息,同时也支持视角和外观的受控变化。我们介绍了WorldRover,一个用于生成丰富注释的艺术家构建环境的长距离探索数据引擎。在其核心,WorldRover-Engine是一个Unreal Engine管道,执行并离线渲染分钟级别的路线,同时保留其完整的轨迹和场景几何。相同的探索可以在不同的环境状态下,从第一人称、第三人称和360度全景相机进行重播。使用WorldRover-Engine,我们构建了WorldRover-10M,其序列将RGB与度量深度、相机轨迹和基于轨迹的动作信号配对,贯穿每次探索。第三人称子集还提供了密集的光流、长距离的2D/3D点轨迹及其可见性,以及与相机轨迹不同的角色轨迹。该引擎可以在不同的环境状态下或使用中性白色材料,从第一人称、第三人称和360度全景视角渲染一次遍历,同时保留路线和场景几何。因此,WorldRover将长时间范围的世界探索转变为一个可扩展的数据生成问题,为必须构建、维护和重新访问可探索世界的连贯表示的模型提供监督。
cs.CV / 106 / 2608.15683

BASeg: Boundary-Aware Remote Sensing Segmentation with Structural Penalties

BASeg:具有结构惩罚的边界感知遥感分割
Song, Yuexi, Sun, Kailai, Wang, Zhuoyu, He, Mingyi, Liang, Paul Pu, Wang, Shenhao, Zhao, Jinhua
Abstract
Semantic segmentation is a core computer vision task in the remote sensing field, accelerating advancements in ur- ban development, agriculture, ecology, water resources, and environmental monitoring. However, recent methods usually struggle to capture fine-grained object features and bound- ary details. Besides, current widely used datasets often lack city morphology diversity and segmentation on generative im- ages remains largely unexplored. To address these issues, we propose a Mahalanobis-Angle Boundary Loss (MABL) that explicitly enhances boundary and shape consistency. MABL jointly models structural importance and boundary orientation through Mahalanobis distance-based weighting and angle- aware penalty. It can be readily integrated into diverse seg- mentation architectures and consistently improves their accu- racy. Built upon MABL, we introduce BASeg, a boundary- aware remote sensing segmentation framework with Struc- tural Penalties. BASeg integrates a Global Visual State Space module (GSM) with a Cross-Feature Fusion module (CFM) to capture both long-range contextual dependencies and fine- grained local details. Additionally, we establish a global 10- city benchmark dataset (GCD-25k) to facilitate accurate build- ing and road segmentation. Extensive experiments on four remote-sensing benchmarks demonstrate that BASeg consis- tently outperforms existing methods, achieving up to a 2.8% improvement in mIoU while producing more accurate object boundary segmentation across diverse scenes. Moreover, integrating MABL into multiple existing segmentation archi- tectures consistently improves performance across datasets, demonstrating its robustness and broad applicability.
Chinese Translation
语义分割是遥感领域的核心计算机视觉任务,推动了城市发展、农业、生态、水资源和环境监测等领域的进步。然而,近期的方法通常难以捕捉细粒度的物体特征和边界细节。此外,目前广泛使用的数据集往往缺乏城市形态的多样性,而在生成图像上的分割研究仍然很大程度上未被探索。为了解决这些问题,我们提出了一种马哈拉诺比斯角度边界损失(Mahalanobis-Angle Boundary Loss,MABL),该损失显式增强了边界和形状的一致性。MABL通过基于马哈拉诺比斯距离的加权和角度感知惩罚共同建模结构重要性和边界方向。它可以轻松集成到各种分割架构中,并始终提高其准确性。在MABL的基础上,我们引入了BASeg,一个具有结构惩罚的边界感知遥感分割框架。BASeg集成了全局视觉状态空间模块(Global Visual State Space module,GSM)和跨特征融合模块(Cross-Feature Fusion module,CFM),以捕捉长距离上下文依赖和细粒度局部细节。此外,我们建立了一个全球10城市基准数据集(GCD-25k),以促进建筑和道路的准确分割。在四个遥感基准上的大量实验表明,BASeg始终优于现有方法,在mIoU上提高了最多2.8%的性能,同时在不同场景中产生了更准确的物体边界分割。此外,将MABL集成到多个现有分割架构中,始终提高了各数据集的性能,展示了其稳健性和广泛适用性。
cs.CV / 107 / 2608.15685

Counterfactual Sensitivity Is Not Repairability: Auditing Replay Probes for Video Evidence

反事实敏感性并不等于可修复性:视频证据的审计重放探针
AlHamidi, Rama, Khanbayov, Rasul, Serpedin, Erchin, Kurban, Hasan
Abstract
Tool-using video agents retrieve visual evidence before answering, but the final answer is not forced to depend on what was retrieved. The natural black box test is counterfactual: destroy the semantic content of the frames the agent retrieved and check whether the answer changes, against a matched sham that re-executes the identical pipeline on those same frames. We introduce CARVE, a black-box counterfactual probe that compares answer changes under matched SHAM and DESTROY replays. Across three independent k=3 runs on a frozen VideoExplorer-style agent, DESTROY changes the answer 29.3 percentage points more often than SHAM, yielding a large and reproducible aggregate effect. Question-level scores are less stable, and increasing the replay budget from k=3 to k=10 reduces ties but weakens the original zero-threshold routing policy. At k=3, CARVE selects 538 of 1,258 LVBench questions and improves accuracy by 3.26 points, with higher fallback yield than most matched random subsets. The score shows only a weak association with annotated temporal coverage, so CARVE is best understood as a routing signal rather than a direct grounding classifier. Our implementation is available at https://github.com/KurbanIntelligenceLab/CARVE.
Chinese Translation
使用工具的视频代理在回答之前检索视觉证据,但最终答案并不一定依赖于所检索的内容。自然的黑箱测试是反事实的:破坏代理所检索帧的语义内容,并检查答案是否发生变化,与在相同帧上重新执行相同管道的匹配虚假(SHAM)进行对比。我们引入了CARVE,这是一种黑箱反事实探针,比较在匹配的SHAM和DESTROY重放下答案的变化。在对一个冻结的VideoExplorer风格代理进行的三次独立k=3运行中,DESTROY比SHAM更频繁地改变答案29.3个百分点,产生了一个大且可重复的整体效应。问题级别的得分较不稳定,将重放预算从k=3增加到k=10减少了平局的情况,但削弱了原始的零阈值路由策略。在k=3时,CARVE选择了1,258个LVBench问题中的538个,并提高了3.26分的准确性,其后备收益高于大多数匹配的随机子集。得分与注释的时间覆盖之间仅存在弱关联,因此CARVE最好被理解为一种路由信号,而不是直接的基础分类器。我们的实现可在https://github.com/KurbanIntelligenceLab/CARVE获取。
cs.CV / 108 / 2608.15688

Training-Free Long-Term Multi-Object Tracking for Sports Video Analytics

无训练的长期多目标跟踪用于体育视频分析
Stanczyk, Tomasz, Yoon, Seongro, Bremond, Francois
Abstract
Long-term multi-object tracking in sports remains challenging due to frequent occlusions, rapid camera motion, and repeated player reappearances. We introduce McByte++, a training-free tracking-by-detection framework that integrates lightweight mask propagation, conditional camera motion compensation, and online re-identification within a unified pipeline. Compared to its predecessor, McByte++ substantially improves runtime efficiency while enhancing identity preservation. On SoccerNet-tracking and SportsMOT benchmarks, McByte++ achieves up to +3.0 HOTA and +6.1 IDF1 improvements over the original McByte in the online setting, with further gains when combined with offline global association. Replacing heavy segmentation components and optimizing motion modeling yields up to an order-of-magnitude speed increase. All results are obtained without detector retraining or dataset-specific tuning. Code will be made available at https://github.com/tstanczyk95/McBytePlusPlus.
Chinese Translation
在体育领域,长期多目标跟踪仍然面临挑战,主要由于频繁的遮挡、快速的摄像机运动以及球员的重复出现。我们提出了 McByte++,一种无训练的基于检测的跟踪框架,集成了轻量级的掩码传播、条件摄像机运动补偿和在线重识别于一个统一的流程中。与其前身相比,McByte++ 在提高身份保持的同时,显著提高了运行效率。在 SoccerNet-tracking 和 SportsMOT 基准测试中,McByte++ 在在线设置下相比于原始的 McByte 实现了高达 +3.0 的 HOTA 和 +6.1 的 IDF1 改进,并且在结合离线全局关联时进一步提升。通过替换重型分割组件和优化运动建模,速度提升可达到数量级的增加。所有结果均在未进行检测器重训练或数据集特定调优的情况下获得。代码将发布在 https://github.com/tstanczyk95/McBytePlusPlus。
cs.CV / 109 / 2608.15692

Automated Fetal Brain MRI Biometry in Healthy and Pathological Cases

健康与病理案例中自动化胎儿脑MRI生物测量
Masterl, Ema, Vesnaver, Tina Vipotnik, Šubič, Nejc, Špiclin, Žiga
Abstract
Automated biometric analysis of fetal brain MRI enables reproducible, observer-independent quantitative assessment, yet existing methods are often restricted to few measurements or evaluated only on healthy cases. We assemble and evaluate an automated biometric analysis pipeline that localizes 22 anatomical landmarks on NeSVoR-reconstructed 3D volumes and derives 11 clinically relevant measurements spanning supratentorial, ventricular, cerebellar, and midline structures. We compare two landmark localization models, H3DE-Net and SCN, on a heterogeneous cohort of 122 acquisitions (both healthy controls and range pathologies). Localization accuracy was assessed with a linear mixed-effects model, agreement with normative growth trajectories with calibrated centile charts, and diagnostic utility with a decision tree classifying VM severity. H3DE-Net achieved significantly lower localization error than SCN across all landmarks (mean 1.36 mm vs. 3.58 mm in HC and 1.90 mm vs. 4.13 mm in PC; p < 0.001), and outperformed a GA-based regression baseline in 7 of 11 measurements. H3DE-Net measurements yielded higher classification AUC in every diagnostic group, with the clearest advantage in separating healthy controls from VM. Decision tree thresholds for ventricular width fell near the clinical 10 mm and 15 mm cut-offs used to define and grade VM.
Chinese Translation
自动化胎儿脑MRI的生物测量分析能够实现可重复、独立于观察者的定量评估,但现有方法通常仅限于少量测量或仅在健康案例中进行评估。我们组建并评估了一种自动化生物测量分析管道,该管道在NeSVoR重建的三维体积上定位22个解剖标志,并导出11个涵盖上幕、脑室、小脑和中线结构的临床相关测量。我们在一个包含122个采集(包括健康对照和各种病理情况)的异质性队列中比较了两种标志定位模型H3DE-Net和SCN。通过线性混合效应模型评估定位精度,通过校准的百分位图表评估与规范生长轨迹的一致性,并通过决策树对VM严重程度进行分类评估诊断效用。H3DE-Net在所有标志上的定位误差显著低于SCN(健康对照组平均1.36 mm对比3.58 mm,病理对照组平均1.90 mm对比4.13 mm;p < 0.001),并在11个测量中的7个中优于基于GA的回归基线。H3DE-Net的测量在每个诊断组中均获得了更高的分类AUC,在区分健康对照与VM方面表现出明显优势。脑室宽度的决策树阈值接近用于定义和分级VM的临床10 mm和15 mm临界值。
cs.CV / 110 / 2608.15694

RRFC: Recursive Refinement via Feedback Conditioning for Iterative Image-to-Image Generation

RRFC:通过反馈条件进行递归精炼的迭代图像生成
Hassani, Kareem, Abbas, Chaymaa, Mubasher, Hadi Al, Awad, Mariette
Abstract
Conditional image-to-image generators are single-shot: they map input features to an output in one forward pass and treat it as final, with no opportunity to improve on it. Although trained to produce the best possible result in one step, such a model leaves room for improvement if it can adaptively revise its own output over iterations. We propose Recursive Refinement via Feedback Conditioning (RRFC), a novel feedback-conditioning framework for iterative output refinement that teaches a model to adaptively revise its output by conditioning on a new signal, namely its most recent previous prediction, which is fed back as an auxiliary set of channels alongside the original input. This preserves the generator's core architecture while modifying its conditioning interface and, depending on the model family, its training or inference procedure, so RRFC can be attached to existing generators without redesign. We evaluate RRFC across six baselines spanning adversarial, equilibrium, and diffusion-based models and three paired image-to-image translation tasks. Across 18 architecture-task settings, RRFC yields seven Holm-corrected improvements, seven degradations, and four non-significant changes. The gains concentrate on reconstruction-fidelity and identity settings, while five of the seven degradations fall on the single semantic-layout task, where every model declines. These results indicate that feedback-based refinement helps when its objective overlaps with the evaluated property, and that its gains concentrate on the tasks where that overlap holds.
Chinese Translation
条件图像生成器是单次生成的:它们在一次前向传播中将输入特征映射到输出,并将其视为最终结果,没有改进的机会。尽管经过训练以在一步中产生最佳结果,但这样的模型如果能够在迭代中自适应地修正其输出,仍然有改进的空间。我们提出了递归精炼通过反馈条件(RRFC),这是一个新颖的反馈条件框架,用于迭代输出精炼,它教会模型通过条件化新的信号,即其最近的前一次预测,自适应地修正其输出,这一信号作为辅助通道与原始输入一起反馈。这保留了生成器的核心架构,同时修改了其条件接口,并根据模型家族的不同,调整其训练或推理过程,因此RRFC可以附加到现有生成器上而无需重新设计。我们在六个基准上评估了RRFC,涵盖对抗性、平衡和基于扩散的模型以及三个配对的图像到图像翻译任务。在18种架构-任务设置中,RRFC产生了七个经过霍尔姆校正的改进、七个退化和四个不显著的变化。增益集中在重建保真度和身份设置上,而七个退化中的五个发生在单一语义布局任务上,每个模型都出现下降。这些结果表明,当反馈基础的精炼目标与评估属性重叠时,它是有帮助的,并且其增益集中在重叠存在的任务上。
cs.CV / 111 / 2608.15695

Bitstream Action Recognition is Byte Modeling

比特流动作识别是字节建模
Li, Fangcheng, Huang, Chaoran, Liu, Tianyi, Liu, Wenyang, Wu, Kejun, Liu, Qiong, Yang, You, Li, Zhengguo
Abstract
Conventional action recognition typically relies on successful pixel decoding of the bitstream. However, bitstream corruption during storage or transmission may cause severe visual artifacts or even decoding failure, posing a significant challenge to reliable action recognition. Bitstream Action Recognition (BAR) aims to overcome the dependency on decoding and the vulnerability to corruption. In this paper, we propose a novel BAR framework, Bitstream Recognition via Anchoring Corrupted Embeddings (BRACE). BRACE is a dual-branch byte-modeling architecture that treats a corrupted bitstream and its intact counterpart as two byte realizations of the same action. This guides the generation of rich and stable representations for robustness to corruption through Intact-Anchored Representation Alignment (IARA). The intact representation serves as a stable anchor, and the corrupted one is aligned to it at the embedding and decision levels under Unreliable-Anchor Suppression (UAS), entirely in representation space and without repairing the bitstream. To address the scarcity of corrupted bitstreams in practice, we introduce the Real-world Bitstream Corruption Simulator (RBCS), a four-parameter simulator that reproduces bit-flip and byte-loss errors arising in transmission and storage. Building on RBCS, we construct the first large-scale BAR dataset (BAR-D), which comprises the BAR-Stanford40 and BAR-PPMI subsets and spans diverse corruption types and severity levels. Finally, we build a large benchmark on BAR-D involving 14 action recognition methods from the pixel, compressed, and bitstream domains. Extensive experiments demonstrate that BRACE has superior robustness to bitstream corruption than all comparison methods. Ablation studies further validate the effectiveness of the proposed RBCS augmentation and IARA.
Chinese Translation
传统的动作识别通常依赖于比特流的成功像素解码。然而,在存储或传输过程中,比特流的损坏可能导致严重的视觉伪影甚至解码失败,这对可靠的动作识别构成了重大挑战。比特流动作识别(Bitstream Action Recognition, BAR)旨在克服对解码的依赖以及对损坏的脆弱性。本文提出了一种新颖的BAR框架,称为通过锚定损坏嵌入的比特流识别(Bitstream Recognition via Anchoring Corrupted Embeddings, BRACE)。BRACE是一种双分支字节建模架构,将损坏的比特流及其完整的对应物视为同一动作的两种字节实现。这一方法通过完整锚定表示对齐(Intact-Anchored Representation Alignment, IARA)引导生成丰富且稳定的表示,以增强对损坏的鲁棒性。完整表示作为稳定的锚点,而损坏的表示则在嵌入和决策层面与其对齐,采用不可靠锚点抑制(Unreliable-Anchor Suppression, UAS),完全在表示空间内进行,而无需修复比特流。为了解决实际中损坏比特流的稀缺性,我们引入了现实世界比特流损坏模拟器(Real-world Bitstream Corruption Simulator, RBCS),这是一个四参数模拟器,能够重现传输和存储中出现的比特翻转和字节丢失错误。在RBCS的基础上,我们构建了第一个大规模BAR数据集(BAR-D),该数据集包括BAR-Stanford40和BAR-PPMI子集,涵盖了多种损坏类型和严重程度。最后,我们在BAR-D上建立了一个大型基准,涉及来自像素、压缩和比特流领域的14种动作识别方法。大量实验表明,BRACE在比特流损坏方面的鲁棒性优于所有对比方法。消融研究进一步验证了所提出的RBCS增强和IARA的有效性。
cs.CV / 112 / 2608.15698

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

ConceptFormer:学习自适应潜在概念以实现视觉文档检索中的查询-文档对齐
Chunyi, Peng, Zhipeng, Xu, Yukun, Yan, Zhenghao, Liu, Shi, Yu, Sen, Mei, Yubo, Sun, Yongheng, Zhang, Jie, Zhou, Yu, Gu, Ge, Yu, Maosong, Sun
Abstract
Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is distributed across text, layout, charts, and visual structures. Recent efforts toward finer-grained supervision primarily rely on textual descriptions or localized visual regions as evidence proxies. However, such supervision signals may either overlook complex visual structures or provide incomplete and inaccurate representations of the underlying evidence. To address these limitations, we propose ConceptFormer, a latent concept representation learning framework for visual document retrieval. ConceptFormer models query-relevant evidence as continuous, query-conditioned latent concepts that explicitly bridge localized visual evidence and semantic relevance, without requiring either textual intermediate representations or direct reliance on raw visual annotations. During training, ConceptFormer employs a strong vision-language model to dynamically determine the number of latent concept tokens and uses these concepts as an intermediate representation to bridge the semantic gap between queries and documents, thereby guiding the learning of the embedding space. Experiments on diverse visual document retrieval benchmarks demonstrate that ConceptFormer achieves 16.7\% and 22.1\% relative improvements in average NDCG@10 over the strongest visual retrieval baseline and the strongest OCR-based text retrieval baseline, respectively. Further analysis reveals that latent concepts effectively connect localized visual evidence with semantic relevance, enabling the retriever to capture both fine-grained textual cues and complex document-level visual structures while preserving strong retrieval alignment. Codes and data are available at https://github.com/Neuir/ConceptFormer.
Chinese Translation
视觉文档检索是多模态检索增强生成的关键组成部分,旨在从文档集合中识别与查询相关的页面,其中证据分布在文本、布局、图表和视觉结构中。近期针对更细粒度监督的努力主要依赖于文本描述或局部视觉区域作为证据代理。然而,这种监督信号可能会忽视复杂的视觉结构,或提供对潜在证据的不完整和不准确的表征。为了解决这些局限性,我们提出了ConceptFormer,一个用于视觉文档检索的潜在概念表征学习框架。ConceptFormer将与查询相关的证据建模为连续的、查询条件的潜在概念,这些概念明确地桥接了局部视觉证据和语义相关性,而无需依赖文本中间表征或直接依赖原始视觉注释。在训练过程中,ConceptFormer利用强大的视觉-语言模型动态确定潜在概念标记的数量,并使用这些概念作为中间表征,以弥合查询与文档之间的语义差距,从而指导嵌入空间的学习。在多样的视觉文档检索基准上的实验表明,ConceptFormer在最强视觉检索基线和最强OCR基础文本检索基线的平均NDCG@10上分别实现了16.7\%和22.1\%的相对提升。进一步的分析表明,潜在概念有效地将局部视觉证据与语义相关性连接起来,使检索器能够捕捉到细粒度的文本线索和复杂的文档级视觉结构,同时保持强大的检索对齐。代码和数据可在 https://github.com/Neuir/ConceptFormer 获取。
cs.CV / 113 / 2608.15705

PixelControl: Fine-Grained Condition Fidelity in Text-to-Image Diffusion

PixelControl:文本到图像扩散中的细粒度条件保真度
Lin, Xin, Li, Haodong, Zhang, Zhifei, Yang, Yutong, Zheng, Haitian, Tian, Juanxi, Lin, Zhe, Nguyen, Truong
Abstract
Controllable text-to-image diffusion models can often follow the global layout of spatial conditions, yet still violate fine-grained structures such as object boundaries, thin contours, and medium/small conditioned regions. This limitation is especially problematic for VAE-based latent diffusion, where spatial compression can weaken high-frequency and low-area condition signals. We propose PixelControl, a pixel-space controllable diffusion framework for fine-grained condition fidelity. Built on a PixelDiT-style backbone, PixelControl avoids the latent bottleneck and introduces two complementary designs. First, Structure-Aware Control Injection derives a condition structure map and uses it to strengthen injected control residuals around spatially sensitive regions. Second, Multi-Scale Pyramid Cycle Loss verifies generated images against condition-derived structures across multiple resolutions, balancing global layout consistency with local boundary and detail accuracy. PixelControl supports depth, segmentation, edge, and their combinations through modality-specific control branches with lightweight gated fusion. Experiments across depth, segmentation, and edge control show that PixelControl improves structural fidelity and visual quality over existing controllable generation methods, with especially strong gains on boundaries and medium/small conditioned regions. The project page can be found at: https://linxin0.github.io/pixelcontrol_homepage/pixelcontrol-site/
Chinese Translation
可控的文本到图像扩散模型通常能够遵循空间条件的全局布局,但仍然会违反细粒度结构,例如物体边界、细轮廓和中小型条件区域。这一限制在基于变分自编码器(VAE)的潜在扩散中尤为突出,因为空间压缩可能削弱高频和低区域条件信号。我们提出了PixelControl,一种用于细粒度条件保真度的像素空间可控扩散框架。PixelControl基于PixelDiT风格的骨干网络,避免了潜在瓶颈,并引入了两种互补设计。首先,结构感知控制注入(Structure-Aware Control Injection)推导出条件结构图,并利用它来增强在空间敏感区域周围注入的控制残差。其次,多尺度金字塔循环损失(Multi-Scale Pyramid Cycle Loss)在多个分辨率下验证生成图像与条件导出结构的一致性,平衡全局布局一致性与局部边界和细节的准确性。PixelControl通过特定于模态的控制分支与轻量级门控融合支持深度、分割、边缘及其组合。针对深度、分割和边缘控制的实验表明,PixelControl在结构保真度和视觉质量上优于现有的可控生成方法,尤其在边界和中小型条件区域上表现出显著的提升。项目页面可访问:https://linxin0.github.io/pixelcontrol_homepage/pixelcontrol-site/
cs.CV / 114 / 2608.15708

What You Ask is What You Ground: Bridging Question Intent to Temporal Evidence for Grounded VideoQA

你所询问的即是你所依据的:将问题意图与时间证据连接起来以实现基础视频问答
Seo, Jinhwan, Han, Kyubeom, Lee, Jumin, Noh, Junhyug, Yoon, Sung-eui
Abstract
We study a critical yet overlooked failure mode in Grounded Video Question Answering: question-invariant grounding, where models predict nearly identical temporal segments for different questions about the same video. We trace this behavior to two structural limitations in prior common designs: (i) modality isolation that fixes video representations before they receive question semantics, and (ii) weak question injection inside the grounding module. To address this, we propose GroundFormer, which conditions video features on question intent before localization via learnable communication tokens that mediate directed visuo-lingual interaction. On top of the question-conditioned features, a factorized MIL cross-attention couples answer selection with temporal evidence under candidate-level supervision, while Gaussian smoothing converts peaked attention into temporally coherent segments. We further introduce a hierarchical multi-modal contrastive loss that aligns video, question, and answer embeddings across a two-pass training pipeline. GroundFormer achieves state-of-the-art grounded VideoQA performance on NExT-GQA and STAR, substantially improving question-discriminative temporal grounding.
Chinese Translation
我们研究了基础视频问答中一个关键但被忽视的失败模式:问题不变的基础,即模型为关于同一视频的不同问题预测几乎相同的时间片段。我们将这种行为追溯到先前常见设计中的两个结构性限制:(i)模态隔离,在此过程中视频表示在接收问题语义之前就被固定,以及(ii)在基础模块中弱的问题注入。为了解决这个问题,我们提出了GroundFormer,它在定位之前通过可学习的通信标记对视频特征进行问题意图的条件化,从而调节定向的视听语言交互。在问题条件化特征的基础上,分解的多实例学习(MIL)交叉注意力将答案选择与候选级监督下的时间证据相结合,而高斯平滑将尖峰注意力转化为时间一致的片段。我们进一步引入了一种层次化的多模态对比损失,它在两次训练流程中对视频、问题和答案嵌入进行对齐。GroundFormer在NExT-GQA和STAR上实现了最先进的基础视频问答性能,显著改善了问题区分的时间基础。
cs.CV / 115 / 2608.15710

Beyond Single Object: Learning 3D Relations with Large Language Models

超越单一对象:利用大型语言模型学习三维关系
Ide, Kohsuke, Yamada, Ryousuke, Qiu, Yue, Ma, Xianzheng, Fukuhara, Yoshihiro, Kataoka, Hirokatsu, Satoh, Yutaka
Abstract
We address a fundamental gap in 3D-LLMs: existing models focus on single-object/scene description, struggling with detailed, inter-object comparison. We propose a framework for detailed object-level reasoning across multiple objects with three components: (1) MO3D (Multi-Object in 3D), an instruction dataset requiring fine-grained multi-object comparison; (2) Multi-3DLLM, using a minimal Patch-Interaction Transformer (PIT) that models inter-/intra-object relationships while preserving local geometry; (3) Mini-apps, two application-driven benchmarks (Shape Mating, Change Captioning) that probe geometric understanding for practical use. Recent 3D-LLMs and 2D-VLMs perform poorly on these tasks, lacking both comparison-centric design and geometric awareness. In contrast, Multi-3DLLM trained on our mixture data learns geometric reasoning, surpasses all baselines on MO3D, and provides positive transfer to single-object classification.
Chinese Translation
我们解决了三维大型语言模型(3D-LLMs)中的一个基本缺口:现有模型专注于单一对象/场景描述,难以进行详细的对象间比较。我们提出了一个框架,用于在多个对象之间进行详细的对象级推理,包含三个组成部分:(1)MO3D(Multi-Object in 3D),一个需要细粒度多对象比较的指令数据集;(2)Multi-3DLLM,使用最小化的Patch-Interaction Transformer(PIT),该模型在保持局部几何形状的同时建模对象间/内部关系;(3)Mini-apps,两个以应用为驱动的基准(形状配对、变化描述),探讨几何理解的实际应用。近期的3D-LLMs和2D-VLMs在这些任务上的表现不佳,缺乏以比较为中心的设计和几何意识。相比之下,基于我们混合数据训练的Multi-3DLLM学习了几何推理,在MO3D上超越了所有基线,并为单一对象分类提供了积极的迁移。
cs.CV / 116 / 2608.15713

YOLO26-RD: An End-to-End Road Damage Detection Network With Learnable Contrast Enhancement and Edge-Guided Downsampling

YOLO26-RD:一种具有可学习对比度增强和边缘引导下采样的端到端道路损伤检测网络
Youwai, Sompote, Chaipetch, Pawarotorn, Samaikul, Hathairat, Yonseng, Theerayut
Abstract
Automated pavement-distress detection is commonly framed as a small-object problem, motivating high-resolution P2/4 detection heads and lossless downsampling. We present YOLO26-RD, an end-to-end (NMS-free) detector built on YOLO26 with two lightweight novel modules (LearnableContrast, a 494-parameter differentiable analogue of CLAHE that adapts contrast per tile inside the network, and EdgeSPD, a Sobel-gated space-to-depth downsampler adding only 2 parameters over SPD-Conv), and we subject the design to a data-first audit on a 7,618-image road-survey dataset (alligator crack, linear crack, patching). The audit falsifies the small-object premise: 92% of instances are COCO-large, and linear cracks are extreme-aspect structures (median 10:1) whose difficulty is sensitivity, not localization. Guided by this analysis, we remove the P2 detection level while retaining P2 features in the fusion path, which improves mAP50 by 2.8 points over the full YOLO26-RD model and reduces epoch time by 8%. Trained from scratch at 640x640, our best screening configuration reaches 0.787 mAP50 on the validation split versus a 0.771 project baseline (a stock YOLO26-s of uncontrolled recipe), with the largest per-class gain on the rarest class (patching, 2.6 points over baseline; 9.4 over the unmodified YOLO26-RD control under an identical recipe). A failure-mode decomposition further attributes the residual error of the bottleneck class (crack, approximately 0.74 across all architectures tested) to train/validation distribution shift on crack orientation and length, sub-pixel crack width at 640x640, and label incompleteness, factors no architecture change can address. We argue that for pavement imagery, measurement-driven subtraction outperforms module accretion, and we release our audit protocol alongside the model.
Chinese Translation
自动化路面损伤检测通常被视为小物体问题,这促使我们使用高分辨率的 P2/4 检测头和无损下采样。我们提出了 YOLO26-RD,这是一种基于 YOLO26 构建的端到端(无 NMS)检测器,配备两个轻量级新模块(LearnableContrast,一个具有 494 个参数的可微分 CLAHE 类比,能够在网络内部根据每个瓦片自适应对比度,以及 EdgeSPD,一个仅增加 2 个参数的 Sobel 门控空间到深度下采样器),并对设计在一个包含 7,618 张图像的道路调查数据集(鳄鱼裂缝、线性裂缝、修补)上进行了数据优先审计。审计结果否定了小物体的前提:92% 的实例为 COCO-large,线性裂缝是极端长宽比结构(中位数 10:1),其难点在于敏感性而非定位。在这一分析的指导下,我们去除了 P2 检测级别,同时在融合路径中保留 P2 特征,这使得完整的 YOLO26-RD 模型的 mAP50 提高了 2.8 个点,并且减少了 8% 的训练周期时间。在 640x640 的基础上从零开始训练,我们的最佳筛选配置在验证集上达到了 0.787 的 mAP50,相较于 0.771 的项目基线(未控制配方的标准 YOLO26-s),在最稀有类别(修补)上获得了最大的每类增益(相较基线提高 2.6 个点;在相同配方下相较未修改的 YOLO26-RD 控制提高 9.4 个点)。失败模式分解进一步将瓶颈类别(裂缝,所有测试架构的平均约为 0.74)的残余误差归因于裂缝方向和长度的训练/验证分布偏移、640x640 下的亚像素裂缝宽度以及标签不完整性,这些因素是任何架构变更无法解决的。我们认为,对于路面图像,基于测量的减法优于模块的增加,并且我们将审计协议与模型一同发布。
cs.CV / 117 / 2608.15721

Anatomical and Physical Supervision for CT-less PET Attenuation Correction: BIC-MAC 2026 Challenge

无CT PET衰减校正的解剖和物理监督:BIC-MAC 2026挑战
Chatzitoulousis, Petros, Matsopoulos, George K.
Abstract
This report describes our submission to the Big Cross-Modal Attenuation Correction (BIC-MAC) 2026 Challenge for CT-less PET attenuation correction through multimodal pseudo-CT synthesis. We build upon a standard nnU-Net architecture and combine anatomical and physical supervision to improve both pseudo-CT quality and downstream PET reconstruction. Anatomical supervision is introduced through a frozen TotalSegmentator feature extractor, anatomy-guided structural constraints and patch sampling, while physical supervision is achieved using a differentiable attenuation correction factor projection loss based on multi-angle attenuation projections. Furthermore, the network is initialized with pretrained weights obtained from training on the SynthRAD Challenge MR-to-CT dataset. Minimal architectural modifications are applied, while performance improvements are pursued across the nnU-Net pipeline, including preprocessing, plans, and supervision design, among other components. Our final submission demonstrates the effectiveness of combining anatomical supervision, attenuation physics, and efficient nnU-Net scaling for CT-less PET attenuation correction.
Chinese Translation
本报告描述了我们对大跨模态衰减校正(BIC-MAC)2026挑战的提交,旨在通过多模态伪CT合成实现无CT的PET衰减校正。我们基于标准的nnU-Net架构,结合解剖和物理监督,以提高伪CT的质量和后续的PET重建。解剖监督通过冻结的TotalSegmentator特征提取器、解剖引导的结构约束和补丁采样引入,而物理监督则通过基于多角衰减投影的可微衰减校正因子投影损失实现。此外,网络初始化时使用了在SynthRAD挑战的MR到CT数据集上训练获得的预训练权重。我们对架构进行了最小的修改,同时在nnU-Net管道的各个组件中追求性能改进,包括预处理、计划和监督设计等。我们的最终提交展示了结合解剖监督、衰减物理和高效nnU-Net扩展在无CT PET衰减校正中的有效性。
cs.CV / 118 / 2608.15731

Identifying Confusion Trends in Concept-based XAI for Multi-Label Classification

识别基于概念的可解释人工智能在多标签分类中的混淆趋势
Amjad, Haadia, Tetzlaff, Ronald
Abstract
Deep Neural Networks (DNNs) deployed in high-risk domains, such as healthcare and autonomous driving, must be not only accurate but also understandable to ensure user trust. In real-world computer vision tasks, these models often operate on complex images containing background noise and are heavily annotated. To make such models explainable, Concept-based Explainable AI (CXAI) methods need to be assessed for their applicability and problem-solving capacity. In this work, we explore CXAI use cases in multi-label classification by training two DNNs, VGG16 and ResNet50, on the 20 most annotated labels in the MS-COCO dataset (Microsoft Common Objects in Context). We apply two CXAI methods, CRP (Concept Relevance Propagation) and CRAFT (Concept Recursive Activation FacTorization), to generate concept-level explanations and investigate the overall evaluations. Our analysis reveals three key findings: (1) CXAI highlights learning weaknesses in DNNs, (2) higher concept distinctiveness reduces label and concept confusion, and (3) environmental concepts expose dataset-induced biases. Our results demonstrate the potential of CXAI to enhance the understanding of model generalizability and to diagnose bias instigated by the dataset.
Chinese Translation
在高风险领域(如医疗保健和自动驾驶)中部署的深度神经网络(DNN)不仅必须具备准确性,还需具备可理解性,以确保用户信任。在现实世界的计算机视觉任务中,这些模型通常在包含背景噪声且注释繁重的复杂图像上运行。为了使这些模型具有可解释性,需要评估基于概念的可解释人工智能(CXAI)方法的适用性和问题解决能力。在本研究中,我们通过在MS-COCO数据集(微软上下文中的常见物体)上训练两个DNN(VGG16和ResNet50),探讨CXAI在多标签分类中的应用案例,重点关注20个注释最多的标签。我们应用两种CXAI方法,CRP(概念相关传播)和CRAFT(概念递归激活因子分解),生成概念级别的解释并调查整体评估。我们的分析揭示了三个关键发现:(1)CXAI突出显示了DNN的学习弱点,(2)更高的概念独特性减少了标签和概念的混淆,以及(3)环境概念揭示了数据集引发的偏见。我们的结果展示了CXAI在增强模型可泛化性理解和诊断数据集引发的偏见方面的潜力。
cs.CV / 119 / 2608.15749

ES3D: Embedding Semantics into 3D Space for Component-Aware Editing

ES3D:将语义嵌入3D空间以实现组件感知编辑
Jin, Xuancheng, Xie, Rengan, Lu, Jiayuan, Zheng, Wenting, Wang, Rui, Huo, Yuchi, Li, Lincheng, Chen, Yingfeng
Abstract
Existing 3D editing methods have made notable progress in controllability, yet they remain limited in several important ways. Most approaches rely on text-driven editing, which struggles to express fine-grained visual changes intended by the user. Moreover, many methods require manually supplied 3D masks or introduce unintended changes to regions that should remain untouched. These limitations largely arise from the absence of fine-grained semantic understanding, making it difficult for existing models to retrieve or modify specific 3D components. We introduce ES3D, a framework that embeds semantics directly into 3D space, enabling component-aware retrieval and editing of a 3D asset conditioned on multiple local reference images and optional text queries. We first construct a 3D semantic embedding by projecting multi-view semantic features into the voxelized space of the asset. We then perform 3D component retrieval by computing feature similarity between the 3D semantic embedding and the semantic embeddings of image or text queries. For editing, we employ a pretrained 3D generative model with an inpainting mechanism to modify the retrieved components guided by user-provided images while preserving the rest of the asset. Overall, ES3D is a 3D editing framework that retrieves editable regions based on semantic cues and uses multiple images as conditions. Extensive experiments demonstrate that ES3D produces geometrically consistent and semantically coherent edits, enabling robust image-based and text-assisted control for 3D editing.
Chinese Translation
现有的3D编辑方法在可控性方面取得了显著进展,但在几个重要方面仍然存在局限性。大多数方法依赖于文本驱动的编辑,这在表达用户所期望的细粒度视觉变化时显得力不从心。此外,许多方法需要手动提供3D掩膜,或者对应保持不变的区域引入了意外的变化。这些局限性主要源于缺乏细粒度的语义理解,使得现有模型难以检索或修改特定的3D组件。我们提出了ES3D,一个将语义直接嵌入3D空间的框架,使得基于多个局部参考图像和可选文本查询的组件感知检索和编辑成为可能。我们首先通过将多视角语义特征投影到资产的体素化空间中构建3D语义嵌入。然后,我们通过计算3D语义嵌入与图像或文本查询的语义嵌入之间的特征相似性来执行3D组件检索。在编辑过程中,我们采用一个预训练的3D生成模型,并结合修补机制,根据用户提供的图像修改检索到的组件,同时保留资产的其余部分。总体而言,ES3D是一个基于语义线索检索可编辑区域并使用多幅图像作为条件的3D编辑框架。大量实验表明,ES3D能够生成几何一致且语义连贯的编辑,增强了基于图像和文本辅助的3D编辑控制能力。
cs.CV / 120 / 2608.15757

Beyond Independence: Learning Correlated Views for Variational Incomplete Multi-View Clustering

超越独立性:学习相关视图以进行变分不完全多视图聚类
Xu, Zheming, Tang, Aiyue, Chen, Shidi, Zou, Xuechao, Lang, Congyan, Mancisidor, Rogelio A., Kampffmeyer, Michael
Abstract
Incomplete multi-view clustering (IMVC) aims to uncover shared cluster structures from data with partially observed views. Although recent imputation-free methods based on variational inference demonstrate robustness to missing views, they commonly rely on a conditional independence assumption across views in the posterior aggregation stage, which fails to capture the inherently structured and potentially correlated nature of multi-view data. In this paper, we propose a variational framework that explicitly goes beyond this assumption by introducing a learnable cross-view correlation structure. Specifically, we explicitly model and learn correlations between views by utilizing the covariance structure of posterior estimation errors during aggregation. To facilitate robust and efficient learning, the correlation matrix is parameterized through a normalized Cholesky decomposition, ensuring positive definiteness and enabling the entire model to be trained jointly through a unified variational objective. Extensive experiments on multiple IMVC benchmarks demonstrate that our method consistently outperforms state-of-the-art approaches across diverse missing-view settings while introducing only a negligible number of learnable parameters. These results highlight the effectiveness of adaptive correlation modeling in variational IMVC, demonstrating the need to go beyond the independence assumption in IMVC. The code is available at https://github.com/zmxu196/ACOVA.
Chinese Translation
不完全多视图聚类(IMVC)旨在从部分观察到的视图数据中揭示共享的聚类结构。尽管最近基于变分推断的无填充方法在缺失视图方面表现出鲁棒性,但它们通常在后验聚合阶段依赖于视图之间的条件独立性假设,这未能捕捉多视图数据固有的结构性和潜在的相关性。在本文中,我们提出了一种变分框架,明确超越这一假设,通过引入可学习的跨视图相关结构。具体而言,我们通过利用聚合过程中后验估计误差的协方差结构,显式建模和学习视图之间的相关性。为了促进鲁棒和高效的学习,相关矩阵通过标准化的Cholesky分解进行参数化,确保其正定性,并使整个模型能够通过统一的变分目标进行联合训练。在多个IMVC基准上的广泛实验表明,我们的方法在各种缺失视图设置下始终优于最先进的方法,同时仅引入了可学习参数的微不足道数量。这些结果突显了自适应相关建模在变分IMVC中的有效性,表明在IMVC中超越独立性假设的必要性。代码可在 https://github.com/zmxu196/ACOVA 获取。
cs.CV / 121 / 2608.15785

RoofGS: Roofline-Guided End-to-End Acceleration of 3D Gaussian Splatting

RoofGS:基于Roofline的端到端3D高斯点云加速
Luo, Yang, Gong, Yan, Gao, Yongsheng, Zhao, Jie
Abstract
3D Gaussian Splatting (3DGS) enables real-time novel-view synthesis but remains limited on GPUs at high resolutions. Through a stage-wise Roofline characterization, we identify two distinct hardware bottlenecks: global memory traffic dominates the front end, whereas instruction throughput limits rasterization. Guided by this analysis, we develop RoofGS, a rendering framework that applies bottleneck-specific optimizations rather than generic kernel acceleration. For the memory-bound front end, we design a resolution-adaptive quantized depth sorting key that compresses each key to 32 bits. For the compute-bound rasterizer, we introduce a range-aware bit-level fast exponential approximation tailored to the bounded exponent range after opacity culling, with a derived per-pixel error bound. These two core techniques are complemented by additional optimizations (kernel fusion, compact attribute storage, culling, dual-pixel evaluation) that additionally reduce memory traffic and improve instruction-level parallelism. Experiments show that RoofGS achieves a 10.1$\times$ end-to-end speedup over 3DGS at 4K on an RTX 4090, increasing throughput from 61 to 616 FPS, with only a 0.028 dB PSNR loss.
Chinese Translation
3D高斯点云(3D Gaussian Splatting, 3DGS)能够实现实时新视角合成,但在高分辨率下仍然受到GPU的限制。通过阶段性Roofline特征分析,我们识别出两个明显的硬件瓶颈:全局内存流量主导前端,而指令吞吐量限制光栅化。基于这一分析,我们开发了RoofGS,一个渲染框架,它应用特定于瓶颈的优化,而不是通用的内核加速。对于内存受限的前端,我们设计了一种分辨率自适应的量化深度排序键,将每个键压缩为32位。对于计算受限的光栅化器,我们引入了一种范围感知的位级快速指数近似方法,专门针对在不透明度剔除后有限的指数范围,并推导出每像素的误差界限。这两项核心技术得到了额外优化(内核融合、紧凑属性存储、剔除、双像素评估)的补充,进一步减少了内存流量并提高了指令级并行性。实验表明,RoofGS在RTX 4090上实现了相较于3DGS在4K分辨率下的10.1倍端到端加速,将吞吐量从61 FPS提升至616 FPS,且仅损失0.028 dB的PSNR。
cs.CV / 122 / 2608.15788

ChainSpace: A Chained-Reasoning Paradigm for Spatial Intelligence

ChainSpace:一种用于空间智能的链式推理范式
Zhang, Xiaohan, Gu, Feng, Rao, Xudong, Pan, Xuhao, Wei, Tao, Pan, Zhou, Zhan, Kun
Abstract
Spatial intelligence requires foundation models to maintain coherent spatial state across interactions with the physical world. However, existing data-centric approaches typically treat spatial reasoning as independent question-answer instances, enabling shortcut-based answering and providing limited supervision for persistent spatial understanding. To address this, we introduce ChainSpace, a chained-reasoning paradigm that structures spatial reasoning as a state-preserving multi-round process. In this paradigm, spatial questions are organized into logically constrained and jointly consistent chains, where later questions depend on spatial constraints established in earlier rounds. Following this principle, we instantiate ChainSpace-Bench, a manually annotated real-world multi-round benchmark with a Chain-Aware Metric, and ChainSpace-Pipeline, a simulator-based chain-structured supervision generation framework for spatial intelligence training. Experiments show that ChainSpace-Bench exposes chain-level failures that are not captured by isolated question accuracy. Additionally, with a relatively small amount of simulator-generated chained data, models trained by ChainSpace-Pipeline achieve the best performance among open-source models on ChainSpace-Bench and transfer competitively to multiple external spatial intelligence benchmarks. These results establish ChainSpace as an effective paradigm for more faithful evaluation and more data-efficient learning of spatial intelligence.
Chinese Translation
空间智能需要基础模型在与物理世界的交互中保持一致的空间状态。然而,现有的数据驱动方法通常将空间推理视为独立的问题-答案实例,这使得基于捷径的回答成为可能,并为持久的空间理解提供了有限的监督。为了解决这个问题,我们引入了ChainSpace,一种将空间推理结构化为保持状态的多轮过程的链式推理范式。在这一范式中,空间问题被组织成逻辑约束和共同一致的链条,其中后续问题依赖于早期轮次中建立的空间约束。遵循这一原则,我们实例化了ChainSpace-Bench,这是一个手动注释的真实世界多轮基准,带有链感知度量(Chain-Aware Metric),以及ChainSpace-Pipeline,这是一个基于模拟器的链结构监督生成框架,用于空间智能训练。实验表明,ChainSpace-Bench揭示了链级别的失败,这些失败在孤立问题的准确性中未被捕捉。此外,通过相对少量的模拟器生成的链数据,使用ChainSpace-Pipeline训练的模型在ChainSpace-Bench上取得了开源模型中最佳的表现,并在多个外部空间智能基准上具有竞争力的迁移能力。这些结果确立了ChainSpace作为一种更真实评估和更高效学习空间智能的有效范式。
cs.CV / 123 / 2608.15796

Emergent 3D Instance Segmentation from Self-Supervised Point Transformers

自监督点变换器生成的三维实例分割
Lentsch, Ted, Montiel-Marín, Santiago, Caesar, Holger, Kooij, Julian F. P.
Abstract
Unsupervised 3D instance segmentation of outdoor LiDAR scans has traditionally relied on handcrafted geometric priors such as density-based clustering, motion cues, or projected 2D detections. In this work, we investigate whether a frozen, self-supervised point transformer already contains the structural information required to isolate object instances without any handcrafted geometric prior. Using this transformer purely as a feature extractor, we probe its internal representations across the SemanticKITTI, nuScenes, and Waymo Perception datasets. Our analysis yields four core insights: (1) the instance signal concentrates in the attention queries and keys rather than in the values or final output features; (2) output features semantically collapse, merging adjacent same-class objects that the queries and keys keep distinct; (3) this instance signal is bimodal in depth, strongest at the shallowest and deepest encoder stages; and (4) this signal is driven predominantly by the rotary position encoding (RoPE), whose removal collapses its advantage. We put these findings into our method TokenGraph3D, a training-free segmenter that groups points via connected components on a key-similarity graph, using neither density-based clustering nor proximity priors. Under identical prior-free conditions, we substantially outperform output-feature baselines, making the emergent 3D instance structure visible.
Chinese Translation
无监督的户外激光雷达扫描三维实例分割传统上依赖于手工设计的几何先验,如基于密度的聚类、运动线索或投影的二维检测。在本研究中,我们探讨了一个冻结的自监督点变换器是否已经包含了隔离物体实例所需的结构信息,而无需任何手工设计的几何先验。我们将该变换器仅作为特征提取器,分析其在SemanticKITTI、nuScenes和Waymo感知数据集上的内部表示。我们的分析得出了四个核心见解:(1)实例信号集中在注意力查询和键中,而不是值或最终输出特征中;(2)输出特征在语义上发生崩溃,合并了相邻的同类对象,而查询和键则保持其区分;(3)该实例信号在深度上呈双峰分布,在最浅和最深的编码器阶段最强;(4)该信号主要由旋转位置编码(Rotary Position Encoding, RoPE)驱动,去除它会削弱其优势。我们将这些发现应用于我们的算法TokenGraph3D,这是一种无训练的分割器,通过在关键相似性图上使用连通组件对点进行分组,而不使用基于密度的聚类或邻近先验。在相同的无先验条件下,我们显著超越了输出特征基线,使得新兴的三维实例结构得以显现。
cs.CV / 124 / 2608.15802

PWLR: Pairwise Witness Local Rejection for Boundary-Aware Out-of-Distribution Detection

PWLR:针对边界感知的成对见证局部拒绝的分布外检测
Jia, Chengyao, Wang, Ruixuan
Abstract
Out-of-distribution (OOD) detection remains challenging for image classifiers, especially when near-OOD samples lie close to in-distribution (ID) class boundaries. Recent vision-language detectors improve OOD detection through class semantics, local prompting, or LLM-generated outlier concepts, but seldom use language as explicit boundary evidence between confusing ID classes. We propose Pairwise Witness Local Rejection (PWLR), which uses an MLLM offline to describe visible local cues that favor one ID class over a specific rival class. These cue phrases are then screened with ID-only data under a frozen vision-language backbone, so that only reliable local verifiers are kept. At inference, PWLR first retains a small set of globally plausible classes, then checks whether any of them is locally supported against its most relevant rivals, and finally combines this pairwise local evidence with the global class score through calibration. Experiments on ImageNet-100 far-OOD, cleaner/challenging OOD and near-OOD benchmarks show that PWLR consistently improves strong vision-language baselines across multiple backbones. Source code will be released.
Chinese Translation
分布外(OOD)检测对于图像分类器仍然具有挑战性,尤其是当近似OOD样本靠近分布内(ID)类别边界时。最近的视觉-语言检测器通过类别语义、局部提示或大语言模型(LLM)生成的异常概念来改善OOD检测,但很少将语言作为混淆ID类别之间的明确边界证据。我们提出了成对见证局部拒绝(PWLR),该方法使用一个离线的多模态大语言模型(MLLM)来描述有利于一个ID类别而非特定竞争类别的可见局部线索。这些线索短语随后在冻结的视觉-语言骨干网络下使用仅包含ID的数据进行筛选,以保留可靠的局部验证器。在推理阶段,PWLR首先保留一小组全球可行的类别,然后检查这些类别是否在其最相关的竞争者中得到局部支持,最后通过校准将这种成对局部证据与全球类别得分结合起来。在ImageNet-100远离OOD、干净/具有挑战性的OOD和近OOD基准上的实验表明,PWLR在多个骨干网络上始终改善了强大的视觉-语言基线。源代码将会发布。
cs.CV / 125 / 2608.15812

From Generation to Matching: A Development Report on Personalized Chinese Handwriting

从生成到匹配:个性化中文手写的开发报告
Liu, Yiwei
Abstract
This paper documents a frozen engineering project on personalized Chinese handwriting. The project started from approximately 200 real handwriting images from one user, covering 197 unique Chinese characters, and was initially formulated as few-shot generation of unseen characters. A sequence of canonical-centered personalization routes repeatedly exposed the same conflict: increasing structural pressure made outputs more canonical, while increasing personalization could damage identity-defining strokes. The project was therefore reset around real-human character equivalence classes. A multi-writer CASIA candidate pool showed that a USER-compatible realization often already existed among valid human samples. The task consequently changed from synthesis to character-wise matching, followed by cross-writer composition into a virtual writer. The frozen system uses real-ink features, character-specific human population percentiles, top-20 candidate pruning, and greedy hardest-first whole-row selection. On the covered target set, all 197 USER characters had real-human candidates, and the 100-character evaluation subset was covered 100/100. Knowncharacter held-out comparisons included a row judged visually almost indistinguishable from genuine USER handwriting. A 60- episode stability audit placed every episode in a predefined A-like machine-proxy region, but these were not independent human A-level judgments. The final evidence supports stable practical B-level quality, with many outputs approaching A-level under the USER-defined criterion. The report records why generation became unnecessary for this case without claiming unrestricted or universal handwriting synthesis.
Chinese Translation
本文记录了一个关于个性化中文手写的冻结工程项目。该项目起始于来自一位用户的大约200幅真实手写图像,涵盖197个独特的汉字,最初被设定为对未见汉字的少样本生成。一个以规范为中心的个性化路径序列反复暴露出同样的冲突:增加结构压力使输出更具规范性,而增强个性化则可能损害身份定义笔画。因此,项目围绕真实人类字符等价类进行了重置。一个多作者的CASIA候选池显示,通常在有效的人类样本中,已经存在与用户兼容的实现。因此,任务从合成转变为逐字符匹配,随后进行跨作者组合以形成虚拟作者。该冻结系统使用真实墨水特征、字符特定的人口百分位、前20候选的修剪,以及贪婪的最难优先整行选择。在所覆盖的目标集上,所有197个用户字符都有真实人类候选,而100个字符的评估子集覆盖率为100/100。已知字符的保留比较中,包括一行在视觉上几乎无法与真实用户手写区分的评判。60集的稳定性审计将每一集置于预定义的A类机器代理区域,但这些并不是独立的人类A级判断。最终证据支持稳定的实际B级质量,许多输出在用户定义的标准下接近A级。报告记录了为何在此案例中生成变得不必要,而不声称无限制或普遍的手写合成。
cs.CV / 126 / 2608.15818

FlowDance: Music-Driven Dance Video Generation with Parallel Pose and RGB Streams

FlowDance:基于音乐驱动的舞蹈视频生成,结合并行姿态和RGB流
Li, Genying, Lin, Boda, Li, Jiachen, Jia, Zijian, Zheng, Haojie, Wang, Yiming, Weng, Shuchen, Li, Si
Abstract
Music-driven dance video synthesis aims to animate a reference person according to a given music clip. The task is challenging because it requires a model to jointly learn music-to-motion correspondence, identity-preserving human animation, temporal coherence, and visually realistic video generation. We present FlowDance, a music-driven dance video generation framework that integrates explicit motion modeling with reference-preserving visual synthesis through parallel pose and RGB streams. We further introduce timestep-aware pose injection to adapt structural guidance across denoising steps and persistent identity injection to preserve the reference appearance over long video. To support this task, we further build a popularity-curated, high-resolution in-the-wild dance video dataset with synchronized music, RGB videos, 3D body motion, camera parameters, and projected 2D pose annotations. Extensive experiments show that FlowDance achieves strong performance in both dance motion generation and music-driven dance video synthesis.
Chinese Translation
基于音乐驱动的舞蹈视频合成旨在根据给定的音乐片段为参考人物进行动画处理。该任务具有挑战性,因为它要求模型共同学习音乐与动作之间的对应关系、保持身份的人类动画、时间一致性以及视觉上逼真的视频生成。我们提出了FlowDance,一个基于音乐驱动的舞蹈视频生成框架,集成了显式运动建模与通过并行姿态和RGB流进行的参考保持视觉合成。我们进一步引入了时间步感知的姿态注入,以适应去噪步骤中的结构指导,以及持久身份注入,以在长视频中保持参考外观。为了支持这一任务,我们还构建了一个经过流行度筛选的高分辨率野外舞蹈视频数据集,其中包含同步的音乐、RGB视频、3D身体运动、相机参数和投影的2D姿态注释。大量实验表明,FlowDance在舞蹈动作生成和基于音乐驱动的舞蹈视频合成方面均取得了良好的性能。
cs.CV / 127 / 2608.15830

MITE-Net: SWaP-Optimized 4K Video Tiny Target Perception for Embodied Edge SAR

MITE-Net:针对具身边缘SAR的SWaP优化4K视频微小目标感知
Xu, Mingshuo, Hua, Mu, Peng, Jigen, Wang, Qi, Yue, Shigang
Abstract
Real-time tiny target perception in high-resolution imagery is critical for embodied Search-and-Rescue (SAR) missions. However, strict Size, Weight, and Power (SWaP) constraints on edge devices like UAVs create a bottleneck: traditional image downsampling causes severe feature loss, while slice-based processing incurs prohibitive latency. To address this gap, this paper introduces a comprehensive framework encompassing a novel architecture, specialized datasets, and hardware-level benchmarks. First, we propose MITE-Net, a SWaP-optimized cascaded architecture, which couples a bio-inspired, learning-free Tiny Target Motion-Based Region Proposal Network (TTM-RPN) with a sub-0.14M-parameter R-CNN-like head. Second, to standardize 4K tiny target evaluation, we construct the SAR-Tiny Datasets by relabeling two challenging UAV datasets: SeaDroneSee-Tiny (dynamic maritime scenes, tiny targets predominantly of 64-256 pixels ) and UAVID-Tiny (cluttered urban scenes, extremely tiny targets, less than 64 pixels). Third, we benchmark against state-of-the-art YOLO models on an edge device, NVIDIA Jetson AGX Xavier, where MITE-Net directly processes 4K maritime imagery, achieving a 100\% search success rate at 30.33 FPS. Consuming merely 3.19 W (9.51 FPS/W), MITE-Net vastly outperforms YOLO baselines in target recall and energy efficiency. Conversely, UAVID-Tiny evaluations expose a compound structural limitation: the learning-free bionic front-end struggles against urban backgrounds, while the ultra-lightweight head lacks representational capacity for complex features. Ultimately, this work delivers an efficient onboard perception paradigm and a rigorous baseline guiding future end-to-end SAR architectures.
Chinese Translation
在高分辨率图像中实现实时微小目标感知对于具身搜索与救援(SAR)任务至关重要。然而,针对无人机等边缘设备的严格尺寸、重量和功耗(SWaP)限制造成了瓶颈:传统的图像下采样会导致严重的特征损失,而基于切片的处理则会产生过高的延迟。为了解决这一问题,本文提出了一个综合框架,包括一种新颖的架构、专门的数据集和硬件级基准。首先,我们提出了MITE-Net,这是一种SWaP优化的级联架构,结合了生物启发的无学习微小目标运动基础区域提议网络(TTM-RPN)和一个参数少于0.14M的R-CNN类似头部。其次,为了标准化4K微小目标评估,我们通过重新标注两个具有挑战性的无人机数据集构建了SAR-Tiny数据集:SeaDroneSee-Tiny(动态海洋场景,微小目标主要为64-256像素)和UAVID-Tiny(杂乱的城市场景,极小目标小于64像素)。第三,我们在边缘设备NVIDIA Jetson AGX Xavier上与最先进的YOLO模型进行基准测试,其中MITE-Net直接处理4K海洋图像,以30.33 FPS的速度实现100%的搜索成功率。MITE-Net仅消耗3.19 W(9.51 FPS/W),在目标召回率和能效方面大幅超越YOLO基线。相反,UAVID-Tiny的评估揭示了复合结构限制:无学习的仿生前端在城市背景下表现不佳,而超轻量级头部在复杂特征的表示能力上不足。最终,本研究提供了一种高效的机载感知范式和一个严格的基线,为未来的端到端SAR架构提供指导。
cs.CV / 128 / 2608.15831

CardiacMamba: Fair and Robust RGB-RF Fusion for Remote Heart Rate Estimation via State Space Modeling

CardiacMamba:通过状态空间建模实现公平且稳健的RGB-RF融合以远程心率估计
Zhao, Bo, Wu, Zheng, Xie, Yiping, YU, Zitong
Abstract
Remote photoplethysmography (rPPG) enables non-contact heart rate (HR) monitoring from facial videos, but RGB-only methods are vulnerable to illumination changes, motion artifacts, and skin-tone-dependent optical reflectance. We propose CardiacMamba, a fair and robust RGB-RF fusion framework that integrates optical facial cues and radio-frequency cardiac motion cues through state space modeling. CardiacMamba introduces a Temporal Difference Mamba Module (TDMM) to enhance subtle RF temporal variations, a bidirectional SSM-based interaction mechanism to align heterogeneous RGB-RF dynamics, and a Channel-wise Fast Fourier Transform (CFFT) module for channel-domain spectral refinement. On the EquiPleth dataset, CardiacMamba achieves state-of-the-art performance with 0.96 bpm MAE, 3.06 bpm RMSE, and 0.97 Pearson correlation, while reducing the observed light-dark skin-tone MAE gap to 0.26 bpm and maintaining robustness under RGB degradation and RF-missing conditions
Chinese Translation
远程光电容积描记法(rPPG)使得可以通过面部视频进行非接触式心率(HR)监测,但仅基于RGB的方法容易受到光照变化、运动伪影和肤色依赖的光学反射影响。我们提出了CardiacMamba,这是一个公平且稳健的RGB-RF融合框架,通过状态空间建模整合光学面部线索和射频心脏运动线索。CardiacMamba引入了时间差分Mamba模块(TDMM)以增强微妙的射频时间变化,一个基于双向状态空间模型(SSM)的交互机制以对齐异构的RGB-RF动态,以及一个通道级快速傅里叶变换(CFFT)模块用于通道域光谱精炼。在EquiPleth数据集上,CardiacMamba实现了最先进的性能,MAE为0.96 bpm,RMSE为3.06 bpm,Pearson相关系数为0.97,同时将观察到的浅色与深色肤色之间的MAE差距减少至0.26 bpm,并在RGB降级和射频缺失条件下保持稳健性。
cs.CV / 129 / 2608.15869

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

超越视觉链思维:内化视觉思维用于主动视频推理
Zhu, Xiaoyu, Deng, Xinke, Taddewadikar, Suresh, Mondal, Arnab Kumar, Jiang, Zhongyu, Fasel, Ian, Liebelt, Joerg
Abstract
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substantial inference overhead, which is particularly problematic for proactive video reasoning. We ask whether models can learn to think visually during training while reasoning directly at inference. We introduce Internalized Visual Thinking (IVT), a post-training framework that jointly optimizes textual prediction and next-embedding prediction over unlabeled videos. Given a partially observed video, IVT predicts latent representations of future frames together with the target textual answer, encouraging the model to capture motion, object transitions, interactions, and latent intent. At inference, IVT generates the answer directly without synthesizing or re-encoding future frames. We conduct controlled studies across target representations, decoder designs, prediction horizons, data mixtures, training curricula, and predictive objectives. IVT improves over direct-answer fine-tuning on all six evaluation settings while retaining the same inference pathway. Compared with explicit Visual CoT, IVT achieves comparable or better performance and reduces average end-to-end latency by more than 5x. Together, our findings suggest that explicit pixel-space generation at inference time, as used in visual chain-of-thought, may not be necessary for effective proactive video reasoning. Predictive world modeling can be internalized during training to produce multimodal reasoners that are both more accurate and substantially more efficient.
Chinese Translation
多模态大型语言模型越来越多地使用视觉链思维(Visual CoT)来推理空间、时间和具身环境。通过生成中间推理图像,视觉链思维提供了一种直观的视觉预见机制,但引入了相当大的推理开销,这对于主动视频推理尤其成问题。我们探讨模型是否可以在训练过程中学习视觉思维,同时在推理时直接进行推理。我们提出了内化视觉思维(Internalized Visual Thinking,IVT),这是一种后训练框架,联合优化文本预测和未标记视频上的下一个嵌入预测。在部分观察到的视频中,IVT预测未来帧的潜在表示以及目标文本答案,鼓励模型捕捉运动、物体转变、交互和潜在意图。在推理时,IVT直接生成答案,而无需合成或重新编码未来帧。我们在目标表示、解码器设计、预测时间范围、数据混合、训练课程和预测目标等方面进行了控制研究。IVT在所有六个评估设置中都优于直接答案微调,同时保持相同的推理路径。与显式视觉链思维相比,IVT实现了可比或更好的性能,并将平均端到端延迟减少了超过5倍。总的来说,我们的研究结果表明,在推理时使用的显式像素空间生成(如视觉链思维中所用)可能并不是有效主动视频推理所必需的。预测世界建模可以在训练过程中内化,从而产生更准确且效率显著更高的多模态推理器。
cs.CV / 130 / 2608.15905

CLARA: Clip-Level Multimodal Alignment with VLM-Derived Rationales for Hateful Video Detection

CLARA:基于VLM衍生理由的剪辑级多模态对齐用于仇恨视频检测
Zhang, Yuchen, Dai, Shuang, Fu, Zeyu, Long, Yunfei, Shekhar, Ravi, Mouratidis, Haralambos
Abstract
Hateful video detection has become increasingly important with the rapid growth of video-centric social media platforms, given the serious risks that hate speech poses to both individual well-being and social cohesion. Compared with text or static multimodal content, hateful video detection remains underexplored and significantly more challenging, as hateful meaning often arises from complex interactions among multimodal cues, including speech, audio, and visual content. Moreover, such signals are often brief, implicit, and temporally dependent, making them difficult to capture using conventional video-level representations. In this work, we propose CLARA, a clip-level multimodal framework for hateful video detection. Instead of treating a video as a single instance, CLARA models it as a sequence of fine-grained clips, enabling more precise capture of temporally localized hateful signals. We introduce a Mixture-of-Experts clip encoder for adaptive multimodal alignment, a local-global segment contrastive objective to jointly model short-term cues and long-range temporal dependencies, and VLM-derived rationales integrated via a gated Transformer to provide high-level semantic guidance. Extensive experiments on three hateful video datasets demonstrate that CLARA consistently outperforms state-of-the-art methods. Further ablation studies and parameter analyses validate the effectiveness of each component.
Chinese Translation
随着以视频为中心的社交媒体平台的快速发展,仇恨视频检测变得越来越重要,因为仇恨言论对个人福祉和社会凝聚力带来了严重风险。与文本或静态多模态内容相比,仇恨视频检测仍然未得到充分探索且显著更具挑战性,因为仇恨意义通常源于多模态线索之间复杂的交互,包括语言、音频和视觉内容。此外,这些信号往往是短暂的、隐含的,并且具有时间依赖性,使得使用传统的视频级表示捕捉这些信号变得困难。在本研究中,我们提出了CLARA,一个用于仇恨视频检测的剪辑级多模态框架。CLARA并不将视频视为单一实例,而是将其建模为一系列细粒度的剪辑,从而更精确地捕捉时间上局部的仇恨信号。我们引入了一种混合专家剪辑编码器以实现自适应多模态对齐,采用局部-全局段对比目标共同建模短期线索和长程时间依赖,并通过门控Transformer集成VLM衍生的理由,以提供高层次的语义指导。在三个仇恨视频数据集上的广泛实验表明,CLARA始终优于最先进的方法。进一步的消融研究和参数分析验证了每个组件的有效性。
cs.CV / 131 / 2608.15915

Comprehensive Benchmarking of Deep Learning Architectures for Lung Cancer Histopathology

深度学习架构在肺癌组织病理学中的综合基准评估
Hasan, Hadi, Salman, Safaa, Sleem, Lama, Mouawad, Ralph, Chehab, Ali
Abstract
Lung cancer remains the leading cause of cancer-related mortality worldwide, while histopathological diagnosis is often affected by inter-observer variability and the substantial workload associated with manual slide examination. Although deep learning has shown considerable potential in computational pathology, comprehensive benchmarks that integrate tissue classification and region segmentation within a unified analytical framework remain limited. This study presents a two-stage deep learning framework for multi-class tissue classification and pixel-level histopathological region segmentation, accompanied by a systematic comparison of state-of-the-art architectures at each stage. For tissue classification, six models, a custom convolutional neural network, VGG16, DenseNet, MobileNetV3, a custom Vision Transformer, and YOLO11, are evaluated on a combined dataset of 39,000 images derived from LC25000 and LungHist700. The models distinguish between adenocarcinoma, squamous cell carcinoma, and normal lung tissue. YOLO11 achieves the best classification performance, with an accuracy of 98.38%, a five-fold cross-validation accuracy of 98.21 +/- 0.35%, and a macro F1-score of 0.98. For region segmentation, U-Net, ResNet-encoder U-Net, DeepLabV3+, and YOLO11-seg are evaluated using the GlaS gland segmentation benchmark. DeepLabV3+ obtains the highest Intersection over Union of 0.80 and a Dice score of 0.89, while YOLO11-seg achieves a comparable Intersection over Union of 0.79 using approximately 14x fewer parameters. The best-performing classification and segmentation models are subsequently integrated into an end-to-end framework, providing an accurate, computationally efficient, and reproducible baseline for automated histopathological image analysis.
Chinese Translation
肺癌仍然是全球癌症相关死亡的主要原因,而组织病理学诊断常常受到观察者间变异性和手动切片检查所带来的巨大工作量的影响。尽管深度学习在计算病理学中显示出了相当大的潜力,但将组织分类和区域分割整合在统一分析框架中的综合基准仍然有限。本研究提出了一种两阶段的深度学习框架,用于多类组织分类和像素级组织病理区域分割,并对每个阶段的最先进架构进行了系统比较。在组织分类方面,评估了六种模型,包括一个定制的卷积神经网络、VGG16、DenseNet、MobileNetV3、一个定制的视觉变换器(Vision Transformer)和YOLO11,这些模型在一个由39,000张图像组成的综合数据集上进行评估,这些图像来源于LC25000和LungHist700。模型能够区分腺癌、鳞状细胞癌和正常肺组织。YOLO11实现了最佳分类性能,准确率为98.38%,五折交叉验证准确率为98.21 +/- 0.35%,宏观F1-score为0.98。在区域分割方面,评估了U-Net、ResNet编码器U-Net、DeepLabV3+和YOLO11-seg,使用GlaS腺体分割基准。DeepLabV3+获得了最高的交并比(Intersection over Union)为0.80和Dice分数为0.89,而YOLO11-seg在使用大约14倍更少参数的情况下实现了相似的交并比0.79。最终,表现最佳的分类和分割模型被整合到一个端到端框架中,为自动化组织病理图像分析提供了一个准确、计算高效且可重复的基准。
cs.CV / 132 / 2608.15970

BagShift: Measuring How Patch Selection Changes the Evidence Seen by Whole-Slide MIL

BagShift:测量补丁选择如何改变全幻灯片多实例学习所见证据
Yuan, Ruicheng, Zhang, Zhenxuan, Hu, Liwei, Wang, Anbang, Xu, Haijie, Luo, Jiawei, Yang, Guang
Abstract
Whole-slide multiple-instance learning (MIL) observes only the patches admitted by its selector. Deployment can alter this selector through compute limits, tissue masking, or regional workflows, even when the patch count is unchanged. We introduce BagShift, a paired protocol that changes the selector for the same case while holding its features and predictor fixed, thereby isolating selector response from case mix. With equal 128-patch budgets, sampling across the tissue or concentrating around one coordinate exposes markedly different evidence: on PANDA, the two views reduce quadratic weighted kappa by 1.57 and 17.96 points, respectively (QWK reported on the $\times100$ scale). On CAMELYON16, lesion annotations withheld from model development show that localized views retain tumor in only 10.0\% of micrometastatic observations, and matched exposure does not consistently recover the loss. The same fixed-count stressor produces a much smaller response on external lung subtyping, although differences in relative coverage make cross-task severity descriptive. When repeated localized observations are available, unioning their patches before one nonlinear MIL pass improves PANDA QWK by 7.87 points over averaging regional predictions. Patch count specifies computation, not observed evidence; deployment evaluations should report both what a selector preserves and how repeated observations are aggregated.
Chinese Translation
全幻灯片多实例学习(MIL)仅观察其选择器所接受的补丁。部署过程中可能通过计算限制、组织掩蔽或区域工作流程改变该选择器,即使补丁数量不变。我们提出了BagShift,一种配对协议,在保持特征和预测器不变的情况下改变同一案例的选择器,从而将选择器响应与案例组合隔离。在相等的128补丁预算下,跨组织取样或集中于一个坐标周围会显著暴露不同的证据:在PANDA上,这两种视图分别将二次加权κ(quadratic weighted kappa)降低了1.57和17.96点(QWK以$ imes100$的比例报告)。在CAMELYON16上,未用于模型开发的病灶注释显示,局部视图仅在10.0%的微转移观察中保留肿瘤,而匹配的曝光并未始终恢复损失。同样的固定数量压力源在外部肺亚型分类上产生的响应要小得多,尽管相对覆盖的差异使得跨任务的严重性具有描述性。当可用重复的局部观察时,在一次非线性MIL处理之前将它们的补丁联合起来,可以使PANDA的QWK提高7.87点,相较于平均区域预测。补丁数量指定了计算,而非观察到的证据;部署评估应报告选择器保留了什么以及如何聚合重复观察。
cs.CV / 133 / 2608.15972

CM-MAE: A Physics-Guided Cross-Modal Self-Supervised Learning Framework for Vision-Wireless Applications

CM-MAE:一种物理引导的跨模态自监督学习框架用于视觉-无线应用
Zhang, Yubo, Liu, Yiyao
Abstract
Synchronized camera and wireless measurements observe the same scene through different physical channels. The central difficulty is that a representation learned in one deployment can fail when viewpoint, traffic, illumination, and propagation geometry change. This paper presents CM-MAE, a self-supervised vision--wireless pretraining framework for cross-scenario representation transfer. The evaluated real-data model uses only RGB frames and the measured 64-beam received-power vector available in DeepSense 6G; it does not use ray-traced paths, calibrated depth, or beam-index labels during pretraining. Its central pretraining term is a \emph{soft contrastive alignment loss}. Instead of making the synchronized image--wireless pair the only positive pair, this loss builds a target distribution from similarities between measured beam-power profiles, so nonidentical samples with similar directional responses are not forced apart as false negatives. A masked joint decoder provides the complementary local objective by reconstructing hidden visual patches and wireless angular clusters under modality dropout. After pretraining, a differential-rate fine-tuning rule lets a new fusion head adapt quickly while the encoders move slowly. Under a sequence-disjoint DeepSense 6G protocol, adding the soft alignment loss improves a matched linear-probe transfer average from 24.88\% to 29.49\%. Mild fusion fine-tuning reaches 77.38\% Top-1 accuracy on unseen Scenarios 6--8, and optional transductive normalization adaptation reaches 78.69\%. Since the fusion setting uses the contemporaneous 64-beam power vector at inference, these results should be read as representation-transfer diagnostics, not as proactive beam-prediction or reduced-sweeping claims.
Chinese Translation
同步的相机和无线测量通过不同的物理通道观察同一场景。主要难点在于,在一个部署中学习到的表示在视角、交通、照明和传播几何变化时可能会失效。本文提出了CM-MAE,一种用于跨场景表示迁移的自监督视觉-无线预训练框架。评估的真实数据模型仅使用RGB帧和在DeepSense 6G中可用的64束接收功率向量;在预训练过程中不使用光线追踪路径、校准深度或束索引标签。其核心预训练项是 extit{软对比对齐损失}。该损失并不将同步的图像-无线对作为唯一的正样本,而是通过测量的束功率轮廓之间的相似性构建目标分布,因此具有相似方向响应的非同一样本不会被错误地分开。一个掩蔽的联合解码器通过在模态丢失下重建隐藏的视觉块和无线角度簇提供互补的局部目标。在预训练后,一个差异率微调规则使得新的融合头能够快速适应,而编码器则缓慢移动。在一个序列不重叠的DeepSense 6G协议下,增加软对齐损失将匹配的线性探测转移平均从24.88\%提高到29.49\%。温和的融合微调在未见的场景6-8上达到了77.38\%的Top-1准确率,而可选的传导归一化适应达到了78.69\%。由于融合设置在推理时使用了同时的64束功率向量,因此这些结果应被视为表示迁移的诊断,而非主动束预测或减少扫描的声明。
cs.CV / 134 / 2608.15984

A Plug-and-Play 2D Motion Interface for Real-World Motion Language Models

一种即插即用的二维运动接口用于现实世界运动语言模型
Yokoyama, Kaname, Ukita, Norimichi
Abstract
Motion Language Models (MoLMs) typically understand human motions by tokenizing 3D motion and processing the resulting tokens using a language model. However, obtaining accurate 3D motions from monocular videos is challenging, limiting their real-world applicability. To address this issue, we introduce a plug-and-play 2D Motion Interface that enables 3D-pretrained MoLMs to accept 2D motion inputs without modifying or fine-tuning the original models. Experiments on public datasets show that our method achieves performance comparable to 3D motion inputs across multiple MoLMs and outperforms training MoLMs from scratch on 2D motions. We further construct a monocular real-world video motion evaluation dataset and introduce a real-video adapter, demonstrating the usefulness of 2D motions over 3D motions under the evaluated monocular pose-estimation setting. These results suggest that 2D motion provides a practical interface for deploying MoLMs in real-world motion understanding settings. Code is available at https://github.com/irajisamurai/2D-Motion-Interface.
Chinese Translation
运动语言模型(MoLMs)通常通过对三维运动进行标记并使用语言模型处理生成的标记来理解人类运动。然而,从单目视频中获取准确的三维运动是具有挑战性的,这限制了它们在现实世界中的适用性。为了解决这个问题,我们提出了一种即插即用的二维运动接口,使得经过三维预训练的MoLMs能够接受二维运动输入,而无需修改或微调原始模型。对公共数据集的实验表明,我们的方法在多个MoLMs上实现了与三维运动输入相当的性能,并且在二维运动上优于从头训练MoLMs。我们进一步构建了一个单目现实世界视频运动评估数据集,并引入了一个真实视频适配器,展示了在评估的单目姿态估计设置下,二维运动相较于三维运动的实用性。这些结果表明,二维运动为在现实世界运动理解环境中部署MoLMs提供了一种实用的接口。代码可在 https://github.com/irajisamurai/2D-Motion-Interface 获取。
cs.CV / 135 / 2608.16008

Spatial Temporal Synergy: Balancing Change and Invariance in Text Driven 3D Human Motion Editing

时空协同:在文本驱动的三维人类运动编辑中平衡变化与不变性
Lin, Shaohui, Shi, Zhenwu, Gong, Jingyu, Xie, Jiao, Zhou, Yu, Zhang, Baochang, Ma, Lizhuang, Lin, Chia-Wen
Abstract
Text-driven human motion editing aims to modify existing motion sequences according to natural language instructions while maintaining the structural consistency of the original motion. Existing diffusion-based approaches struggle to balance text-responsive "change" and inertial "invariance". They often rely on coarse spatial constraints and rigid uniform time assumptions, leading to spatial motion distortions and the destruction of intrinsic physical rhythms during variable-length editing. To handle these challenges, we propose Change and Invariance Motion Editing (CIME), a unified framework that comprehensively decouples change and invariance into spatial pose and temporal rhythm dimensions. For spatial poses, our method integrates an omni-supervised positive-negative learning mechanism comprising hierarchical retrospective feature supervision, subtle motion preservation, and triplet-based semantic alignment. For temporal rhythms, we introduce the Riemannian Non-uniform Integral Manifold Mapping (RNIMM) module, which achieves high-fidelity reproduction of physical beats in the edited text via kinematics-aware non-uniform timestamps. Extensive experiments on the MotionFix and STANCE Adjustment datasets demonstrate that CIME achieves state-of-the-art performance in editing alignment and structural fidelity, validating the effectiveness of our unified architecture. Our source codes and models have been released at: github.com/ZhenwuShi/CIME.git
Chinese Translation
文本驱动的人类运动编辑旨在根据自然语言指令修改现有的运动序列,同时保持原始运动的结构一致性。现有的基于扩散的方法在平衡文本响应的“变化”和惯性的“不变性”方面面临挑战。它们通常依赖粗略的空间约束和刚性的均匀时间假设,这导致了空间运动的扭曲以及在可变长度编辑过程中内在物理节奏的破坏。为了解决这些问题,我们提出了变化与不变性运动编辑(Change and Invariance Motion Editing, CIME),这是一个统一框架,全面将变化与不变性解耦为空间姿态和时间节奏两个维度。在空间姿态方面,我们的方法整合了一种全监督的正负学习机制,包括分层的回顾特征监督、细微运动保留和基于三元组的语义对齐。在时间节奏方面,我们引入了黎曼非均匀积分流形映射(Riemannian Non-uniform Integral Manifold Mapping, RNIMM)模块,通过运动学感知的非均匀时间戳实现了编辑文本中物理节拍的高保真重现。在MotionFix和STANCE Adjustment数据集上的大量实验表明,CIME在编辑对齐和结构保真度方面实现了最先进的性能,验证了我们统一架构的有效性。我们的源代码和模型已发布在:github.com/ZhenwuShi/CIME.git
cs.CV / 136 / 2608.16014

Depth-guided Multi-view Exposure Bracketing for HDR Robot Vision

深度引导的多视角曝光包围用于高动态范围机器人视觉
Kim, Jinnyeong, Choi, Juhyung, Kim, Woohyeok, Cho, Sunghyun, Baek, Seung-Hwan
Abstract
Achieving reliable single-shot high dynamic range (HDR) imaging under extreme illumination conditions remains a long-standing challenge, yet no comprehensive benchmark exist for evaluating HDR perception in multi-sensor robotic systems. To fill this gap, we introduce a large-scale dataset collected via a custom robotic vision platform and an iPhone 13 Pro: 121 real-world scenes spanning modest and ultra-high dynamic range conditions, alongside 20 synthetic video sequences from the CARLA simulator. As a reference pipeline for this dataset, we propose Depth-guided Multi-view Exposure Bracketing (DMEB), a single-shot HDR method that distributes drastically different exposures across multi-view low-bit-depth cameras and fuses them via depth-guided confidence-aware fusion. Evaluations on our dataset show that DMEB establishes a strong reference point and highlight the promise of this sensor configuration for robust HDR perception in diverse multi-camera and depth sensor system.
Chinese Translation
在极端照明条件下实现可靠的单次高动态范围(HDR)成像仍然是一个长期挑战,但目前尚无全面的基准用于评估多传感器机器人系统中的HDR感知。为填补这一空白,我们通过一个定制的机器人视觉平台和iPhone 13 Pro收集了一个大规模数据集:121个真实场景,涵盖适度和超高动态范围条件,以及来自CARLA模拟器的20个合成视频序列。作为该数据集的参考流程,我们提出了深度引导的多视角曝光包围(DMEB),这是一种单次HDR方法,它在多视角低位深度相机之间分配截然不同的曝光,并通过深度引导的置信度感知融合将其融合。对我们数据集的评估表明,DMEB建立了一个强有力的参考点,并突显了该传感器配置在多摄像头和深度传感器系统中实现稳健HDR感知的潜力。
cs.CV / 137 / 2608.16015

Multi-scale Decomposed Convolution Refinement Network for Visible-Infrared Person Re-Identification

多尺度分解卷积精炼网络用于可见光-红外人脸重识别
Zheng, Mingsheng, Jiang, Zirui, Liu, Bo, Chen, Yupeng, Zhang, Jun, Zhao, Kai
Abstract
Visible-infrared person re-identification (VI-ReID) suffers from cross-modal discrepancies and limited discriminative capabilities, leading to suboptimal recognition performance. Current approaches exhibit limitations in semantic mining, cross-modal fusion and feature constraints. To tackle these challenges, we propose MDCRNet, a Multi-scale Decomposed Convolution Refinement Network that enhances cross-modal feature learning and discriminative metric learning. Specifically, we introduce a Hierarchical Learning Module (HLM) containing four Hierarchical Decomposed Convolution Attention (HDCA) modules, each equipped with lightweight channel attention and multi-scale spatial perception blocks to capture multi-scale spatial dependencies. Moreover, we develop a Joint Discriminative Metric Loss (JDML) incorporating a novel Granularity Discriminative Loss (GDL) that simultaneously optimizes intra-identity compactness and inter-identity separability across modalities. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate that MDCRNet achieves state-of-the-art performance on both benchmarks. Code is available at https://github.com/Kevin-zms/MDCRNet.
Chinese Translation
可见光-红外人脸重识别(VI-ReID)面临跨模态差异和有限的区分能力,导致识别性能不佳。当前的方法在语义挖掘、跨模态融合和特征约束方面存在局限性。为了解决这些挑战,我们提出了MDCRNet,一种多尺度分解卷积精炼网络,旨在增强跨模态特征学习和区分度度量学习。具体而言,我们引入了一个层次学习模块(HLM),该模块包含四个层次分解卷积注意力(HDCA)模块,每个模块配备轻量级通道注意力和多尺度空间感知块,以捕捉多尺度空间依赖。此外,我们开发了一种联合区分度度量损失(JDML),其中包含一种新颖的粒度区分损失(GDL),该损失同时优化跨模态的同一身份紧凑性和不同身份可分性。在SYSU-MM01和RegDB数据集上的大量实验表明,MDCRNet在这两个基准测试中均实现了最先进的性能。代码可在 https://github.com/Kevin-zms/MDCRNet 获取。
cs.CV / 138 / 2608.16042

TR-GS: High-Fidelity Sparse-View CT Volumetric Rendering via t-Distribution Gaussian Splatting and Ray-Confidence Modeling

TR-GS:基于 t-分布高斯点云和光线置信度建模的高保真稀疏视图 CT 体积渲染
Xiao, Zedong, Wang, Yiren, Liu, Zhou, Liu, Xiaolin, Lu, Zhangji
Abstract
High-fidelity 3D medical visualization supports applications such as clinical assessment and surgical planning. Sparse-view computed tomography (CT) can reduce projection requirements and associated radiation exposure, but limited observations may introduce structural artifacts and reconstruction uncertainty. Although 3D Gaussian Splatting (3DGS) provides an efficient explicit representation for volumetric rendering, existing CT methods based on standard Gaussian primitives may be sensitive to unreliable observations under sparse-view acquisition. We present TR-GS, a Gaussian-splatting framework for sparse view CT volumetric rendering. TR-GS replaces standard Gaussian primitives with projectable Student's t-distribution primitives and introduces a ray-confidence model that regulates their degrees of freedom according to local ray observability. Confidence-guided 3D wavelet regularization is further used to balance high-frequency detail preservation and noise suppression. This work is licensed under a Creative Commons Attribution 4.0 International License. Experiments on synthetic and real-world datasets show that TR-GS improves over representative baselines in most evaluated settings and remains competitive in the remaining cases. The resulting volumetric representations may support downstream medical multimedia applications, including XR-based visualization and interactive clinical rendering.
Chinese Translation
高保真的三维医学可视化支持临床评估和手术规划等应用。稀疏视图计算机断层扫描(CT)可以减少投影需求和相关的辐射暴露,但有限的观测可能引入结构伪影和重建不确定性。尽管三维高斯点云(3D Gaussian Splatting,3DGS)为体积渲染提供了高效的显式表示,但基于标准高斯原语的现有 CT 方法可能对稀疏视图采集下的不可靠观测敏感。我们提出了 TR-GS,一种用于稀疏视图 CT 体积渲染的高斯点云框架。TR-GS 用可投影的学生 t-分布原语替代标准高斯原语,并引入了一种光线置信度模型,根据局部光线可观测性调节其自由度。此外,采用置信度引导的三维小波正则化来平衡高频细节保留和噪声抑制。该工作在知识共享署名 4.0 国际许可证下发布。对合成和真实世界数据集的实验表明,TR-GS 在大多数评估设置中优于代表性基线,并在其余情况下保持竞争力。所得到的体积表示可能支持下游医学多媒体应用,包括基于 XR 的可视化和交互式临床渲染。
cs.CV / 139 / 2608.16081

SafeGesture: Evaluating Fine-Grained Hand Gesture Understanding in Vision-Language Models through Scenario-Conditioned Safety Interpretation

安全手势:通过情境条件安全解释评估视觉语言模型中的细粒度手势理解
Kim, Taegang, Afroogh, Saleh, Jiao, Junfeng
Abstract
Open-weight and frontier vision-language models (VLMs) perform well on general image understanding, but their ability to interpret fine-grained hand gestures in safety-critical operational contexts remains largely unexamined. We introduce SafeGesture, a benchmark that evaluates whether a model can infer scenario-appropriate safety actions from hand gestures. It pairs six HaGRID gestures with eight operational scenarios for 4,800 items and evaluates Qwen2.5-VL-7B, LLaVA-NeXT-7B, InternVL2-8B, Phi-3.5-Vision, and GPT-4o. Results reveal a perception-reasoning decoupling: GPT-4o achieves 98.4% gesture accuracy but 53.3% safety accuracy, while Qwen2.5-VL reaches 84.9% and 39.5%, yielding gaps of 45.0 and 45.4 percentage points. Four of five models rarely or never use the uncertainty label, and failure directions differ substantially across models. Accuracy also obscures label bias: a scenario-majority policy with no visual input reaches 58.3%, above every evaluated model, while only GPT-4o exceeds this prior under macro-F1. Visual input improves safety accuracy by 11.2 to 30.2 percentage points, but providing the ground-truth gesture as text improves performance by only 0.4 to 3.2 points, and no model exceeds 56.2%. These results indicate that the main bottleneck is scenario-conditioned safety reasoning rather than gesture recognition.
Chinese Translation
开放权重和前沿视觉语言模型(VLMs)在一般图像理解方面表现良好,但它们在安全关键操作环境中解释细粒度手势的能力仍然很大程度上未被检验。我们引入了SafeGesture,这是一个基准,评估模型是否能够从手势推断出情境适当的安全行动。它将六个HaGRID手势与八个操作场景配对,共计4,800个项目,并评估了Qwen2.5-VL-7B、LLaVA-NeXT-7B、InternVL2-8B、Phi-3.5-Vision和GPT-4o。结果揭示了感知与推理的解耦:GPT-4o的手势准确率达到98.4%,但安全准确率仅为53.3%,而Qwen2.5-VL的手势准确率为84.9%,安全准确率为39.5%,两者之间的差距分别为45.0和45.4个百分点。五个模型中有四个很少或从不使用不确定性标签,且模型之间的失败方向差异显著。准确率也掩盖了标签偏差:一个没有视觉输入的情境多数政策达到了58.3%,高于每个评估模型,而只有GPT-4o在宏观F1下超过了这一先验。视觉输入使安全准确率提高了11.2到30.2个百分点,但将真实手势作为文本提供的性能提升仅为0.4到3.2个百分点,且没有模型超过56.2%。这些结果表明,主要瓶颈在于情境条件下的安全推理,而非手势识别。
cs.CV / 140 / 2608.16087

Representation Is Not Enough: Body-Localized Thermal Evidence for Contactless Stress and Craving Sensing in Opioid Use Disorder

表征不足:无接触压力和渴望感知在阿片类药物使用障碍中的身体局部热证据
Deb, Sachin, Sharma, Harshit, Salekin, Asif
Abstract
Removing wearables from physiological monitoring also removes their supervision: the signal indicating where and when a stress response occurred. Contactless stress sensing therefore becomes a weakly supervised evidence-localization problem, where a clip-level label must be traced to the body regions and moments that produced it. We address this with FABLE-Therm, a weakly supervised architecture that preserves localized evidence across body regions, time, and encoder-specific representations until the final decision. FABLE-Therm fuses frozen foundation-model encoders at the embedding level, with theory explaining why localized fusion can outperform feature concatenation and prediction averaging. We study this problem in opioid use disorder (OUD), where stress is a major relapse trigger and sustained wearable use can be difficult during early recovery. Using fixed thermal video, FABLE-Therm achieves 0.938 AUROC on held-out participants, and its learned representation transfers to self-reported craving, providing, to our knowledge, the first evidence that craving can be recovered from contactless thermal video. Localized evidence also enables participant-level analysis of deployment failure. We find that improving representation alone is insufficient for equitable deployment: additional data from the underserved group would recover only about half of the cohort gap, while the remainder reflects person-to-person heterogeneity. This modality-agnostic decomposition applies to models with identifiable subpopulations. Together with the first cohort-structured contactless thermal OUD benchmark, our results show that preserving localized evidence supports both accurate sensing and principled analysis of who a model fails and why.
Chinese Translation
从生理监测中去除可穿戴设备也意味着去除了它们的监督:即指示压力反应发生的时间和位置的信号。因此,无接触压力感知变成了一个弱监督的证据定位问题,其中必须将片段级标签追溯到产生该标签的身体区域和时刻。我们通过FABLE-Therm来解决这个问题,这是一种弱监督架构,能够在身体区域、时间和编码器特定表征之间保持局部证据,直到最终决策。FABLE-Therm在嵌入层融合了冻结的基础模型编码器,理论上解释了为什么局部融合可以优于特征连接和预测平均。我们在阿片类药物使用障碍(OUD)中研究这个问题,在该领域,压力是主要的复发诱因,而在早期恢复期间持续使用可穿戴设备可能很困难。使用固定的热视频,FABLE-Therm在保留参与者中达到了0.938的AUROC,其学习到的表征可以转移到自我报告的渴望上,提供了我们所知的第一个证据,表明渴望可以从无接触热视频中恢复。局部证据还使得参与者级别的部署失败分析成为可能。我们发现,仅仅改善表征不足以实现公平部署:来自服务不足群体的额外数据只能恢复大约一半的队列差距,而其余部分反映了个体间的异质性。这种模态无关的分解适用于具有可识别子群体的模型。结合第一个基于队列结构的无接触热OUD基准,我们的结果表明,保持局部证据支持了准确的感知和对模型失败对象及原因的原则性分析。
cs.CV / 141 / 2608.16100

TISC: A Text-Driven Image Semantic Communication System for Faithful Reconstruction

TISC:一种基于文本驱动的图像语义通信系统用于忠实重建
Zhang, Feifan, Du, Yuyang, Liu, Xiaoyan, Liew, Soung Chang
Abstract
Generative image semantic communication converts an image into a text description and then performs text-to-image reconstruction at the receiver via diffusion-based generative models. This paradigm has attracted broad attention due to its extremely low bandwidth cost. However, existing methods still face two critical bottlenecks across image-to-text (I2T) semantic extraction at the transmitter and text-to-image (T2I) semantic reconstruction at the receiver: (i) semantic loss and distortion in I2T, where holistic image descriptions may omit fine-grained object attributes and spatial-position information, causing the generated text to deviate from the original image semantics; and (ii) insufficient semantic faithfulness in T2I, where even with the same semantically faithful text description, different initial noise settings may lead diffusion-based reconstruction to produce images with different levels of semantic consistency with the original image. These issues jointly limit the semantic faithfulness of image reconstruction. To address them, we propose TISC, a text-driven image semantic communication framework tailored for faithful reconstruction. TISC incorporates two key designs: (1) Tree-Structured Attribute Semantic Extraction (TSASE), which decomposes semantic extraction into global scene, background, and object-level attribute descriptions, covering spatial position, shape/pose, color, material, and other physical attributes for each detected object; and (2) an Initial Noise Optimization (INO) mechanism, which selects an initial noise seed at the transmitter according to a comprehensive similarity score that jointly considers visual and semantic consistency. Experiments on multiple datasets show that TSASE improves object-position recovery and semantic description faithfulness, while the INO parameter study supports the adopted configuration for noise selection.
Chinese Translation
生成图像语义通信将图像转换为文本描述,然后通过基于扩散的生成模型在接收端进行文本到图像的重建。这种范式因其极低的带宽成本而受到广泛关注。然而,现有方法在发射端的图像到文本(I2T)语义提取和接收端的文本到图像(T2I)语义重建方面仍面临两个关键瓶颈:(i)I2T中的语义损失和失真,其中整体图像描述可能忽略细粒度的对象属性和空间位置信息,导致生成的文本偏离原始图像语义;(ii)T2I中的语义忠实度不足,即使在相同的语义忠实文本描述下,不同的初始噪声设置也可能导致基于扩散的重建生成的图像与原始图像在语义一致性上存在不同程度的差异。这些问题共同限制了图像重建的语义忠实度。为了解决这些问题,我们提出了TISC,一种专为忠实重建而设计的文本驱动图像语义通信框架。TISC包含两个关键设计:(1)树状结构属性语义提取(TSASE),它将语义提取分解为全局场景、背景和对象级属性描述,涵盖每个检测到的对象的空间位置、形状/姿态、颜色、材料及其他物理属性;(2)初始噪声优化(INO)机制,根据综合相似性评分选择发射端的初始噪声种子,该评分共同考虑视觉和语义一致性。在多个数据集上的实验表明,TSASE改善了对象位置恢复和语义描述的忠实度,而INO参数研究支持了噪声选择的配置。
cs.CV / 142 / 2608.16103

Beyond Similarity Matching: Structured Reasoning for Open-Vocabulary Referring Segmentation in 3DGS

超越相似性匹配:用于3D高斯点云开放词汇指称分割的结构化推理
Wang, Yizhao, Wang, Xinfa, Wang, Jingbo, Wang, Jingbo, Zhang, Guantao, Han, Yafeng, Gao, Guohong, Xia, Yuhe
Abstract
Open-vocabulary referring segmentation in 3D Gaussian Splatting (3DGS) requires a neural model to select Gaussian primitives according to free-form language expressions. Existing 3DGS-based methods usually rely on global text-region similarity, which is weak for queries involving attributes, reference objects, spatial relations, and fine-grained parts. This often causes target-reference confusion, granularity mismatch, part-whole leakage, and relation violations. We propose QAGaussian, a query-adaptive neural reasoning framework for language-guided Gaussian primitive selection. QAGaussian first learns query-conditioned multi-scale Gaussian slots as differentiable candidates whose receptive fields are shaped by the input expression. It then builds a relation-aware slot graph with language-conditioned edge weighting to propagate target-reference, attribute, part-whole, and contextual evidence. A granularity-adaptive router softly combines region-level, object-level, part-level, attribute-aware, and relation-aware mask branches, followed by relation-constrained refinement for spatial, part-whole, attribute, and geometric consistency. QAGaussian is pretrained only on Mosaic3D-5.6M for Gaussian-text alignment and evaluated on independent benchmarks without target-dataset fine-tuning. It achieves 47.2 Avg. mIoU and 63.2 Avg. F1, outperforming the strongest 3DGS referring baseline by 2.7 mIoU points and 2.9 F1 points. It also improves Part-mIoU from 38.6 to 43.4, Rel-mIoU from 44.4 to 50.8, and reduces target-reference confusion from 10.8 to 7.4. These results demonstrate that query-conditioned slot learning, relation-aware graph reasoning, and adaptive routing provide an effective neural modeling strategy for open-vocabulary referring segmentation in 3DGS. The code is available at https://github.com/zqeslwyz/QAGaussian.
Chinese Translation
在3D高斯点云(3D Gaussian Splatting, 3DGS)中,开放词汇指称分割要求神经模型根据自由形式的语言表达选择高斯原语。现有的基于3DGS的方法通常依赖于全局文本区域相似性,这对于涉及属性、参考对象、空间关系和细粒度部分的查询效果较差。这常常导致目标-参考混淆、粒度不匹配、部分-整体泄漏和关系违规。我们提出了QAGaussian,一种查询自适应的神经推理框架,用于语言引导的高斯原语选择。QAGaussian首先学习查询条件下的多尺度高斯槽作为可微分的候选者,其感受野由输入表达式决定。然后,它构建一个关系感知的槽图,通过语言条件的边权重传播目标-参考、属性、部分-整体和上下文证据。一个粒度自适应的路由器柔性地结合区域级、对象级、部分级、属性感知和关系感知的掩码分支,随后进行关系约束的细化,以确保空间、部分-整体、属性和几何的一致性。QAGaussian仅在Mosaic3D-5.6M上进行高斯-文本对齐的预训练,并在独立基准上进行评估,而无需对目标数据集进行微调。它实现了47.2的平均mIoU和63.2的平均F1,超越了最强的3DGS指称基线2.7 mIoU点和2.9 F1点。同时,它还将Part-mIoU从38.6提高到43.4,将Rel-mIoU从44.4提高到50.8,并将目标-参考混淆从10.8降低到7.4。这些结果表明,查询条件下的槽学习、关系感知的图推理和自适应路由为3DGS中的开放词汇指称分割提供了一种有效的神经建模策略。代码可在https://github.com/zqeslwyz/QAGaussian获取。
cs.CV / 143 / 2608.16104

Nexus: Structured Synergy for Efficient Text-to-Image Generation using Rectified Flow Model

Nexus:基于整合协同的高效文本到图像生成的整流流模型
Wang, Yizhao
Abstract
Diffusion and flow matching models have made significant progress in text-to-image generation, yet high computation, quadratic complexity, and large memory footprint hinder high-resolution synthesis and edge deployment. We propose Nexus, which integrates sparse architecture, linear complexity, and low-bit quantization. It combines MoE feed-forward layers, gated DeltaNet attention, and per-expert low-bit training to reduce computation and memory. Their joint optimization allows Nexus to achieve generation quality comparable to mainstream models such as SDXL and SD3 while delivering markedly higher inference efficiency. Experiments on COCO and LAION validate its effectiveness.
Chinese Translation
扩散和流匹配模型在文本到图像生成方面取得了显著进展,但高计算量、二次复杂性和大内存占用限制了高分辨率合成和边缘部署。我们提出了Nexus,它集成了稀疏架构、线性复杂性和低位量化。Nexus结合了MoE前馈层、门控DeltaNet注意力机制和每个专家的低位训练,以减少计算和内存。它们的联合优化使Nexus在生成质量上可与主流模型如SDXL和SD3相媲美,同时提供显著更高的推理效率。在COCO和LAION上的实验验证了其有效性。
cs.CV / 144 / 2608.16110

SUGFW+: An Uncertainty-guided Feature Weighting Framework for Cold Start Active Adaptation of SAM in Medical Image Segmentation

SUGFW+: 一种基于不确定性引导的特征加权框架,用于医疗图像分割中的冷启动主动适应
Ma, Xiaochuan, Zhu, Ning, Fu, Jia, Zhong, Lanfeng, Jiang, Hanyu, Song, Bin, Li, Kang, Wang, Guotai
Abstract
Cold Start Active Learning (CSAL) is important in improving the performance of a medical image segmentation model with low annotation budget by querying a small subset for annotation from an unlabeled training set. Existing CSAL methods typically rely on inefficient dataset-specific Self-Supervised Learning (SSL) to map the unlabeled images into a feature space for sample selection. Recently, the advent of foundation models such as the Segment Anything Model (SAM) offer a promising alternative as the pre-trained model can provide strong generalizable feature embeddings, and allow high performance in downstream tasks after fine-tuning (adaptation). However, how to systematically exploit SAM's inherent embeddings for cold-start sample selection during adaptation with low annotation budget remains underexplored. To address this, we propose an extended SAM-based Uncertainty-guided Feature Weighting (SUGFW+) framework for CSAL and adaptation of SAM. Specifically, it leverages the SAM for Patch-level Feature and Uncertainty Calculation (PFUC), and introduces a Patch-based Global Distinct Representation (PGDR) module that aggregates patch-level embeddings into highly discriminative, uncertainty-aware image-level features. These features are then utilized by a Greedy Selection with Cluster and Uncertainty (GSCU) strategy to combine diversity and uncertainty during sample selection. Unlike prior CSAL methods that decouple sample selection from model training, SUGFW+ tightly integrates these two stages via an Uncertainty-Prompted Fine-Tuning (UPFT) process of SAM in model training. Extensive experiments on four public datasets demonstrate that SUGFW+ achieves state-of-the-art performance against existing CSAL methods. Code is available at https://github.com/HiLab-git/SUGFW-plus.
Chinese Translation
冷启动主动学习(CSAL)在通过从未标记的训练集中查询小部分进行标注来提高医疗图像分割模型的性能方面至关重要,尤其是在标注预算有限的情况下。现有的CSAL方法通常依赖于低效的特定数据集自监督学习(SSL)来将未标记图像映射到特征空间以进行样本选择。最近,像Segment Anything Model(SAM)这样的基础模型的出现提供了一种有前景的替代方案,因为预训练模型可以提供强大的可泛化特征嵌入,并在微调(适应)后在下游任务中实现高性能。然而,如何系统性地利用SAM固有的嵌入进行低标注预算下的冷启动样本选择仍然未被充分探索。为了解决这个问题,我们提出了一种扩展的基于SAM的不确定性引导特征加权(SUGFW+)框架,用于CSAL和SAM的适应。具体而言,它利用SAM进行补丁级特征和不确定性计算(PFUC),并引入一个基于补丁的全局独特表示(PGDR)模块,将补丁级嵌入聚合为高度可区分的、不确定性感知的图像级特征。这些特征随后被Greedy Selection with Cluster and Uncertainty(GSCU)策略利用,以在样本选择过程中结合多样性和不确定性。与之前将样本选择与模型训练解耦的CSAL方法不同,SUGFW+通过在模型训练中的不确定性引导微调(UPFT)过程紧密集成了这两个阶段。在四个公共数据集上的大量实验表明,SUGFW+在现有CSAL方法中实现了最先进的性能。代码可在https://github.com/HiLab-git/SUGFW-plus获取。
cs.CV / 145 / 2608.16122

TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening

TokenSTFormer:一种用于青少年特发性脊柱侧弯筛查的分词时空注意力模型
Chen, Dong, Cheung, Kenneth M. C.
Abstract
Adolescent Idiopathic Scoliosis (AIS) is a prevalent spinal deformity in adolescents that, if left untreated, can result in severe health outcomes. Traditional screening methods are limited by subjective interpretation, reliance on professional expertise and low scalability. To address these challenges, we present ScoliGait dataset, which comprises 1,516 gait video clips paired with corresponding X-ray records. We also introduce TokenSTFormer, a novel model that tokenizes spatial and temporal semantics to enhance feature representation and convergence. Our model achieves state-of-the-art performance, surpassing vanilla Vision Transformer encoder across key metrics, including accuracy of 0.79. This study highlights the potential of leveraging holistic motion features derived from gait video and attention-based models for scalable, cost-effective AIS screening, paving the way for future clinical applications in scoliosis detection.
Chinese Translation
青少年特发性脊柱侧弯(AIS)是青少年中一种常见的脊柱畸形,如果不加以治疗,可能导致严重的健康后果。传统的筛查方法受限于主观解读、对专业知识的依赖以及低可扩展性。为了解决这些挑战,我们提出了ScoliGait数据集,该数据集包含1,516个步态视频片段及其对应的X光记录。我们还介绍了TokenSTFormer,这是一种新颖的模型,通过对时空语义进行分词,以增强特征表示和收敛性。我们的模型在关键指标上实现了最先进的性能,超越了普通视觉变换器编码器,准确率达到0.79。本研究强调了利用步态视频和基于注意力模型的整体运动特征进行可扩展、成本效益高的AIS筛查的潜力,为未来脊柱侧弯检测的临床应用铺平了道路。
cs.CV / 146 / 2608.16142

Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System

图神经网络辅助的演员-评论家模型用于低延迟边缘视觉系统
Noor, Alam, Almeida, Luis, Li, Kai, Wu, Jiyan, Gaitán, Miguel Gutiérrez, Tovar, Eduardo
Abstract
UAV on-board vision systems are widely used for different activities, including monitoring in no-fly zones. In this case, the vision-equipped UAV streams a video to a ground server where an operator assists its activities. The latency of video transmission has a profound impact on the effectiveness of the operator assistance. However, most techniques available for video transmission still incur significant latency costs. In this paper, we propose a graph convolutional neural network-assisted (GCN-Assisted A2C) deep reinforcement learning (DRL) system model to find the optimal pixel-correlated area of a suspicious object. We combine the Lagrangian dual form with gradient descent to prevent lack of convergence and over- and under-penalization constraint violation during latency optimization. The proposed system model sends a sub-group pixel-correlated area of the frame from the UAV to the server rather than the transmission of the whole video frame. The proposed framework utilizes the GCN model to explore hidden representations of feature-correlated groups of pixels. Moreover, the GCN supervises the A2C model, which selects a subgroup to enhance transmission latency, thus supervising the training of UAV actions in A2C. Experimental results show that GCN-assisted A2C reduces video frame transmission latency together with false detection rate in UAV vision systems over other DRL and state-of-the-art models.
Chinese Translation
无人机(UAV)机载视觉系统广泛应用于多种活动,包括在禁飞区的监控。在这种情况下,配备视觉系统的无人机将视频流传输到地面服务器,操作员在此协助其活动。视频传输的延迟对操作员的辅助效果有着深远的影响。然而,目前大多数视频传输技术仍然存在显著的延迟成本。本文提出了一种图卷积神经网络辅助的深度强化学习(GCN-Assisted A2C)系统模型,以寻找可疑物体的最佳像素相关区域。我们结合拉格朗日对偶形式与梯度下降,以防止在延迟优化过程中缺乏收敛性以及过度和不足惩罚约束的违反。所提出的系统模型从无人机向服务器发送一个子组的像素相关区域,而不是传输整个视频帧。该框架利用GCN模型探索特征相关像素组的隐藏表示。此外,GCN监督A2C模型,选择一个子组以增强传输延迟,从而监督无人机在A2C中的动作训练。实验结果表明,与其他深度强化学习(DRL)模型和最先进的模型相比,GCN辅助的A2C在无人机视觉系统中减少了视频帧传输延迟和误检率。
cs.CV / 147 / 2608.16146

The Right Prior for the Right Deformation: Rethinking Continuous Deformable Image Registration

正确的先验用于正确的变形:重新思考连续可变形图像配准
Liu, Hengjie, Shen, Chushu, Ruan, Dan, Sheng, Ke
Abstract
Deformable image registration models implicitly encode deformation priors through their parametrization and optimization. In this work, we conduct a validation study on continuous registration methods to examine how these implicit priors affect performance across different registration tasks. Classic B-Spline transformations impose locality, smoothness, and scale through their control-point structure, whereas recent INR-based methods impose different priors through neural parameterization and optimization. We compare INR-Dense (IDIR), which directly models a dense displacement field using a SIREN-based INR; INR-BSCP (SINR), which predicts B-Spline control points with an INR; D-BSCP, which directly optimizes single-scale B-Spline control points; and MR-D-BSCP, which adds a multiresolution coarse-to-fine scheme. Experiments on inter-subject brain MR registration (OASIS) and intra-subject exhale-to-inhale lung CT registration (DIR-LAB 4DCT) reveal different behavior across deformation regimes. On OASIS, where deformations are moderate but locally complex, D-BSCP matches or slightly outperforms INR-BSCP, suggesting that the B-Spline parameterization accounts for much of INR-BSCP's effectiveness. On DIR-LAB 4DCT, where respiratory motion is larger and more coherent, single-scale B-Spline methods (D-BSCP and INR-BSCP) are less suitable, while INR-Dense and MR-D-BSCP are more effective. Across both tasks, MR-D-BSCP achieves the best performance among the tested continuous parameterizations. These findings highlight that registration accuracy depends strongly on matching the induced deformation prior to the target motion pattern, and support prior-deformation matching as a practical design principle for medical image registration. Our code will be available at https://github.com/HengjieLiu/RightPriorDIR.
Chinese Translation
可变形图像配准模型通过其参数化和优化隐式编码变形先验。在本研究中,我们对连续配准方法进行了验证研究,以检验这些隐式先验如何影响不同配准任务的性能。经典的 B-Spline 变换通过其控制点结构施加局部性、平滑性和尺度,而最近的基于 INR 的方法则通过神经参数化和优化施加不同的先验。我们比较了 INR-Dense (IDIR),它使用基于 SIREN 的 INR 直接建模密集位移场;INR-BSCP (SINR),它通过 INR 预测 B-Spline 控制点;D-BSCP,直接优化单尺度 B-Spline 控制点;以及 MR-D-BSCP,添加了多分辨率的粗到细方案。在跨受试者脑部 MR 配准 (OASIS) 和同一受试者呼气到吸气肺 CT 配准 (DIR-LAB 4DCT) 的实验中,揭示了不同变形机制下的不同表现。在 OASIS 中,变形适中但局部复杂,D-BSCP 的表现与 INR-BSCP 相当或略优,表明 B-Spline 参数化在很大程度上解释了 INR-BSCP 的有效性。在 DIR-LAB 4DCT 中,呼吸运动更大且更一致,单尺度 B-Spline 方法 (D-BSCP 和 INR-BSCP) 不太适用,而 INR-Dense 和 MR-D-BSCP 更为有效。在这两个任务中,MR-D-BSCP 在测试的连续参数化中实现了最佳性能。这些发现强调了配准精度在很大程度上依赖于所施加的变形先验与目标运动模式的匹配,并支持先验变形匹配作为医学图像配准的实用设计原则。我们的代码将发布在 https://github.com/HengjieLiu/RightPriorDIR。
cs.CV / 148 / 2608.16154

KeyID: Decoupled Drafting and Keyframe Editing for Identity-Preserving Video Generation

KeyID:用于身份保留的视频生成的解耦草图和关键帧编辑
Luo, Jianjie, Zhong, Yiming, Shen, Haoming, Xiao, Yupeng, Yang, Zhenguo
Abstract
Identity-preserving video generation (IPVG) requires synthesizing videos that are faithful to both reference subjects and text prompts. Existing methods are often hindered by high tuning costs or limited input-level enhancements, struggling to maintain rigid identity consistency during complex, long-sequence actions. To address these limitations, we propose KeyID, a training-free IPVG framework that decouples the synthesis of video dynamics from the injection of identity. Specifically, KeyID comprises two components: (1) Reference-Aware Video Generation, which produces an identity-agnostic video draft aligned with multiple references, and (2) Identity-Preserved Keyframe Editing, which integrates the target identity via sparse keyframe correction and subsequent motion interpolation. By shifting from dense frame-level supervision to sparse keyframe-level refinement, KeyID effectively resolves the capacity conflict between prompt adherence and identity fidelity. Crucially, our modular design allows seamless extension to multi-subject references and complex sequential action generation without additional training. KeyID outperforms prior works and is validated by automatic and human evaluations on the official challenge benchmark, ultimately securing the runner-up position in the Track 2 (Sequential Action) of the ACM Multimedia 2026 IPVG Grand Challenge. Source code is available at https://github.com/WISLab-GDUT/KeyID.
Chinese Translation
身份保留视频生成(IPVG)需要合成忠实于参考主体和文本提示的视频。现有方法常常受到高调优成本或有限输入级增强的限制,在复杂的长序列动作中难以保持严格的身份一致性。为了解决这些限制,我们提出了KeyID,一个无训练的IPVG框架,解耦视频动态的合成与身份的注入。具体而言,KeyID包括两个组件:(1)参考感知视频生成,生成与多个参考对齐的无身份视频草图;(2)身份保留关键帧编辑,通过稀疏关键帧修正和后续运动插值整合目标身份。通过从密集帧级监督转向稀疏关键帧级细化,KeyID有效解决了提示遵循与身份保真之间的能力冲突。重要的是,我们的模块化设计允许无缝扩展到多主体参考和复杂序列动作生成,而无需额外训练。KeyID在官方挑战基准上的自动和人工评估中优于之前的工作,最终在ACM Multimedia 2026 IPVG大挑战的第二轨道(序列动作)中获得亚军。源代码可在 https://github.com/WISLab-GDUT/KeyID 获取。
cs.CV / 149 / 2608.16191

Beyond Clear Skies: Synthetic Seasonal and Weather Variations for Real-World Drone Detection

超越晴朗天空:用于真实世界无人机检测的合成季节性和天气变化
Lenhard, Tamara R., Weinmann, Andreas, Koch, Tobias
Abstract
Reliable drone detection under real-world deployment conditions requires training data that spans the full operational design domain, including adverse weather and seasonal appearance variation. However, acquiring and annotating such data at scale remains highly resource-intensive, as adverse-weather conditions are inherently difficult to control, reproduce, and sample systematically. Existing datasets therefore typically provide only limited coverage of such conditions. Conversely, synthetic data offers a scalable alternative: environmental variation becomes controllable, while modern game-engine-based pipelines provide realistic rendering and automatic annotations. Leveraging this potential, we introduce SynDroneVision-Weather (SDV-W), an systematic extension of SynDroneVision (SDV) targeting adverse-weather and seasonal domain shifts in urban drone detection. SDV-W comprises 55,187 annotated high-resolution images from three urban environments, rendered across three seasonal configurations and diverse weather conditions, including rain, snow, and fog at multiple severity levels. By preserving SDV's scene and trajectory configuration, SDV-W enables matched clean-adverse comparisons and quantification of condition-specific detector degradation. Across representative YOLO models and real-world datasets, we show that SDV-W improves detector reliability under adverse appearance shifts, reduces missed detections and false alarms, and is most effective as a complement to general-purpose synthetic drone-detection data. SDV-W will be publicly released upon paper acceptance.
Chinese Translation
在真实部署条件下可靠的无人机检测需要涵盖全面操作设计域的训练数据,包括不利天气和季节性外观变化。然而,在大规模获取和标注此类数据仍然高度资源密集,因为不利天气条件本质上难以控制、重现和系统性采样。因此,现有数据集通常仅提供有限的此类条件覆盖。相反,合成数据提供了一种可扩展的替代方案:环境变化变得可控,而基于现代游戏引擎的管道提供了逼真的渲染和自动标注。利用这一潜力,我们引入了SynDroneVision-Weather (SDV-W),这是对SynDroneVision (SDV)的系统扩展,旨在应对城市无人机检测中的不利天气和季节域变化。SDV-W包含来自三个城市环境的55,187张标注高分辨率图像,这些图像在三种季节配置和多种天气条件下渲染,包括多种严重程度的雨、雪和雾。通过保留SDV的场景和轨迹配置,SDV-W实现了干净与不利条件的匹配比较和条件特定检测器退化的量化。在代表性的YOLO模型和真实世界数据集上,我们展示了SDV-W在不利外观变化下提高了检测器的可靠性,减少了漏检和误报,并且作为通用合成无人机检测数据的补充效果最佳。SDV-W将在论文接受后公开发布。
cs.CV / 150 / 2608.16198

Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology

选择合适的图像进行分类:远程皮肤病学中的可靠输入选择
Gröger, Fabian, Weishaupt, Marco, Gottfrois, Philippe, Lionetti, Simone, Wermelinger, Linda, Ranasekara, Nipun, Amruthalingam, Ludovic, Navarini, Alexander A., Pouly, Marc
Abstract
Dermatology models face distribution shifts in teledermatology settings, where submitted images differ from the training data in lighting, angle, distance, focus, and framing. These test-time images are ordinary clinical photographs, but some fall outside the model's training conditions, leading the model to often misclassify them due to shifts in acquisition between training and deployment. When multiple images of the same case exist (several photos of one patient or lesion), a natural way to improve accuracy is therefore to select the image the model is most likely to classify correctly. We call this task reliable-input selection. An oracle that, for each case, selects a correctly classified image when one exists raises weighted F1 by about 20 percentage points on average across six dermatology datasets and nine frozen backbones. This oracle is an upper bound that sees the labels, whereas a selector must choose blindly. Capturing this gain in practice is hard. A selector that needs no pretraining data applies to any frozen model, including those whose data is not public. It must judge reliability from quantities the model exposes at inference: its embeddings, their norms, and its confidence. We benchmark four such training-data-free selectors: the embedding norm, the neighborhood consensus among a case's images, the stability of the prediction under small perturbations, and the model's own confidence. No training-data-free selector substantially narrows this oracle gap. The best of them is the model's own confidence, but it recovers only a small part of the gap on the clinical datasets. A small labeled reference set does not help either: the best selector overall, a fusion of confidence and Mahalanobis distance, still leaves most of the gap. To our knowledge, this is the first study to introduce and benchmark reliable input selection, a clinically important, unsolved task.
Chinese Translation
皮肤病学模型在远程皮肤病学环境中面临分布变化,提交的图像在光照、角度、距离、焦点和构图上与训练数据有所不同。这些测试时图像是普通的临床照片,但有些超出了模型的训练条件,导致模型因训练与部署之间的获取差异而经常错误分类。当同一病例存在多张图像(同一患者或病变的多张照片)时,改善准确性的自然方法是选择模型最有可能正确分类的图像。我们将这一任务称为可靠输入选择。一个能够为每个病例选择一个正确分类图像的神谕,在六个皮肤病学数据集和九个固定骨干网络上平均提高了约20个百分点的加权F1。这一神谕是一个上限,因为它可以看到标签,而选择器必须盲目选择。在实践中捕捉这一增益是困难的。一个不需要预训练数据的选择器适用于任何固定模型,包括那些数据不公开的模型。它必须根据模型在推理时暴露的量来判断可靠性:其嵌入、嵌入的范数以及模型的置信度。我们基准测试了四种不依赖训练数据的选择器:嵌入范数、病例图像之间的邻域共识、小扰动下预测的稳定性以及模型自身的置信度。没有一种不依赖训练数据的选择器显著缩小这一神谕差距。其中表现最好的选择器是模型自身的置信度,但在临床数据集上仅恢复了差距的一小部分。一个小的标记参考集也没有帮助:总体表现最佳的选择器,即置信度与马哈拉诺比斯距离的融合,仍然留下了大部分差距。据我们所知,这是首次引入并基准测试可靠输入选择这一临床重要且尚未解决的任务。
cs.CV / 151 / 2608.16225

PCT-Prompt: A Prompt-Guided Transformer Framework for Dense Prediction Tasks in Point Clouds

PCT-Prompt:一种用于点云密集预测任务的提示引导变换器框架
Zhang, Dejun, Bai, Yanzi, Wu, Yiqi
Abstract
Standard Transformers have proven effective in point cloud object classification, but their performance in dense prediction tasks within complex scenes is often hindered by weak prior assumptions. To address this challenge, we propose PCT-Prompt, a novel framework that enhances standard Transformers by introducing a prompt-guided feature branch to improve performance in dense prediction tasks. The standard Transformer branch leverages pre-trained models for global feature extraction from point cloud data, serving as the backbone for processing high-level features. Meanwhile, the prompt-guided feature branch consists of two key components: a fine-grained feature extraction block that captures multi-scale geometric features using geometry-sensitive abstraction layer, along with the PnP-3D layer to integrate local context with global regularization. The second component, the prompt-refined feature learning block generates prompt tokens, which are subsequently refined through cross-attention mechanisms. Additionally, we introduce a prompt drop mechanism that progressively removes prompt information across Transformer layers, balancing local details and global consistency. Experimental results on the ShapeNetPart, S3DIS, and DALES datasets demonstrate that PCT-Prompt significantly improves the adaptability of standard Transformers to dense prediction tasks, achieving strong performance in real-world scenarios.
Chinese Translation
标准变换器在点云物体分类中已被证明有效,但在复杂场景中的密集预测任务中,其性能常常受到弱先验假设的限制。为了解决这一挑战,我们提出了PCT-Prompt,这是一种新颖的框架,通过引入提示引导特征分支来增强标准变换器,从而提高密集预测任务的性能。标准变换器分支利用预训练模型从点云数据中提取全局特征,作为处理高层特征的主干。同时,提示引导特征分支由两个关键组件组成:一个细粒度特征提取模块,使用几何敏感抽象层捕捉多尺度几何特征,以及PnP-3D层以将局部上下文与全局正则化相结合。第二个组件,提示精炼特征学习模块生成提示令牌,这些令牌随后通过交叉注意机制进行精炼。此外,我们引入了一种提示丢弃机制,该机制逐步移除变换器层中的提示信息,以平衡局部细节和全局一致性。在ShapeNetPart、S3DIS和DALES数据集上的实验结果表明,PCT-Prompt显著提高了标准变换器在密集预测任务中的适应性,在现实场景中实现了强大的性能。
cs.CV / 152 / 2608.16234

GaussianDWM++: Language-Grounded 3D Gaussian Driving World Model for Unified Scene Understanding, Editing, and Multi-Modal Generation

GaussianDWM++:基于语言的3D高斯驾驶世界模型,用于统一场景理解、编辑和多模态生成
Deng, Tianchen, Chen, Xuefeng, Wu, Shuang, Chen, Qu, Zhu, Jiajun, Dai, Bo, Yang, Jianfei, Wang, Hesheng
Abstract
Driving World Models (DWMs) have recently advanced rapidly with generative models, yet most existing methods mainly focus on conditional scene generation and lack explicit 3D scene understanding, language-grounded reasoning, and controllable 4D editing capabilities. Moreover, commonly used point cloud, occupancy, or BEV representations make it difficult to achieve fine-grained alignment between textual information and the underlying 3D scene structure. To address these limitations, we propose a foundation-feature Gaussian driving world model that unifies scene understanding, language-grounded reasoning, controllable 4D editing, and multi-modal generation within a single framework. Specifically, we introduce a foundation-feature Gaussian tokenizer that directly distills Qwen/SigLIP visual-language features into 3D Gaussian primitives, building a compact open-vocabulary Gaussian semantic field. We further design a geometry-aware Gaussian adapter that combines importance-aware hierarchical selection with text-conditioned Perceiver-style cross-attention to aggregate dense Gaussian primitives into compact world tokens. To improve representation compatibility, we introduce a KL-based Gaussian--image distribution alignment objective that aligns Gaussian world tokens with foundation image tokens. Based on the aligned Gaussian representation, our framework further supports instruction-controllable scene editing, including weather-conditioned generation and dynamic vehicle manipulation. Extensive experiments on broader driving benchmarks demonstrate that our method achieves state-of-the-art performance across scene understanding, visual grounding, planning-oriented reasoning, and controllable 4D generation tasks. We will release the code and datasets publicly on Github.
Chinese Translation
驾驶世界模型(DWMs)近年来随着生成模型的快速发展而取得了显著进展,但现有大多数方法主要集中于条件场景生成,缺乏明确的3D场景理解、基于语言的推理和可控的4D编辑能力。此外,常用的点云、占用或鸟瞰视图(BEV)表示使得文本信息与底层3D场景结构之间的细粒度对齐变得困难。为了解决这些局限性,我们提出了一种基础特征高斯驾驶世界模型,该模型在单一框架内统一了场景理解、基于语言的推理、可控的4D编辑和多模态生成。具体而言,我们引入了一种基础特征高斯分词器,直接将Qwen/SigLIP视觉-语言特征提炼为3D高斯原语,构建一个紧凑的开放词汇高斯语义场。我们进一步设计了一种几何感知高斯适配器,将重要性感知的分层选择与文本条件的Perceiver风格交叉注意力相结合,以将密集的高斯原语聚合为紧凑的世界标记。为了提高表示的兼容性,我们引入了一种基于KL的高斯-图像分布对齐目标,将高斯世界标记与基础图像标记对齐。基于对齐的高斯表示,我们的框架进一步支持可指令控制的场景编辑,包括天气条件生成和动态车辆操控。在更广泛的驾驶基准上进行的广泛实验表明,我们的方法在场景理解、视觉定位、规划导向推理和可控的4D生成任务中达到了最先进的性能。我们将公开发布代码和数据集于Github。
cs.CV / 153 / 2608.16241

Convolution-Free Holistic Multivariance Decomposition Layer for Efficient Hyperspectral Image Classification Tensor Networks

无卷积整体多元分解层用于高效高光谱图像分类张量网络
Tuna, Süha, Başar, Ülker
Abstract
Feature extraction for hyperspectral image classification is conventionally addressed using rigid tensor decompositions that fail to capture complex spatio-spectral interdependencies, or heavily parameterized convolutional neural networks that are computationally expensive. To overcome these limitations, this work introduces the Holistic Multivariance Decomposition (HMD) framework as a novel, end-to-end differentiable neural network layer. By explicitly separating independent single mode variations from cooperative higher dimensional interactions via learnable, matrix valued supports, the proposed HMD-0, HMD-1 and HMD-2 approximants are optimized jointly with a downstream classifier via backpropagation. Comprehensive evaluations across three benchmark HS datasets demonstrate that the higher level HMD layers achieve superior classification accuracy compared to classical learnable tensor baselines, including Tucker, Canonical Polyadic, and Tensor Train decompositions. Furthermore, HMD-1 and HMD-2 achieve a generalization capacity and training stability comparable to standard 2D and 3D-CNNs while requiring significantly fewer feature extractor parameters. These results demonstrate that the HMD framework provides a structurally robust substitute for traditional convolution in multidimensional HS image classification, offering high parameter efficiency and stability throughout the optimization process.
Chinese Translation
高光谱图像分类的特征提取通常采用刚性的张量分解方法,这些方法无法捕捉复杂的时空光谱相互依赖关系,或者采用参数量庞大的卷积神经网络,这些网络计算成本高。为克服这些局限性,本研究提出了整体多元分解(Holistic Multivariance Decomposition, HMD)框架,作为一种新颖的端到端可微分神经网络层。通过可学习的矩阵值支持,明确将独立的单模态变化与协作的高维交互分离,所提出的 HMD-0、HMD-1 和 HMD-2 近似器与下游分类器通过反向传播共同优化。对三个基准高光谱(HS)数据集的全面评估表明,较高层次的 HMD 层在分类准确性上优于经典的可学习张量基线,包括 Tucker、典型多项式(Canonical Polyadic)和张量列(Tensor Train)分解。此外,HMD-1 和 HMD-2 的泛化能力和训练稳定性与标准的 2D 和 3D-CNN 相当,同时所需的特征提取器参数显著减少。这些结果表明,HMD 框架为多维高光谱图像分类中的传统卷积提供了一种结构上稳健的替代方案,在整个优化过程中提供了高参数效率和稳定性。
cs.CV / 154 / 2608.16251

SCOUT: Semantic Concept Discovery for Open-Vocabulary Editing of face Recognition Templates

SCOUT:面部识别模板开放词汇编辑的语义概念发现
Todorov, Leon, Rot, Peter, Peer, Peter, Štruc, Vitomir, Grm, Klemen
Abstract
Face recognition templates are compact identity representations, yet they also encode rich semantic information about facial appearance. Prior work has shown that templates can be inverted to images or indirectly manipulated through image-editing pipelines, but direct semantic editing in template space remains largely unexplored. Existing interpretability methods for face recognition often rely on manual neuron inspection or predefined attribute labels, limiting scalability and semantic flexibility. To address this gap, we propose SCOUT (Semantic Concept Discovery for Open-VocabUlary Editing of Face Recognition Templates), an end-to-end framework for discovering and directly manipulating semantic concepts in face recognition templates using mechanistic interpretability. SCOUT learns sparse template representations, generates semantic hypotheses for latent features from natural-language descriptions, and validates their stability. The resulting features act as controllable semantic directions for direct editing, avoiding costly edit--re-encode pipelines. Experiments with face recognition models using CNN, ViT, and Swin backbones show that SCOUT discovers interpretable concepts beyond standard attribute labels and enables controllable, identity-aware template manipulation with negligible impact on identity matching. We further show that edited templates can subsequently be decoded with independent inversion models for visualization and evaluation.
Chinese Translation
面部识别模板是紧凑的身份表示,但它们也编码了关于面部外观的丰富语义信息。先前的研究表明,模板可以反转为图像或通过图像编辑管道间接操控,但在模板空间中进行直接的语义编辑仍然基本未被探索。现有的面部识别可解释性方法通常依赖于手动神经元检查或预定义的属性标签,这限制了其可扩展性和语义灵活性。为了解决这一问题,我们提出了SCOUT(面部识别模板开放词汇编辑的语义概念发现),这是一个端到端框架,用于发现和直接操控面部识别模板中的语义概念,采用机械可解释性。SCOUT学习稀疏的模板表示,从自然语言描述中生成潜在特征的语义假设,并验证其稳定性。生成的特征作为可控的语义方向用于直接编辑,避免了代价高昂的编辑-重新编码管道。使用CNN、ViT和Swin骨干网的面部识别模型的实验表明,SCOUT发现了超越标准属性标签的可解释概念,并实现了可控的、身份感知的模板操控,对身份匹配的影响微乎其微。我们进一步展示了编辑后的模板可以通过独立的反演模型进行解码,以便于可视化和评估。
cs.CV / 155 / 2608.16259

Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection

Defake-o3:从推测性理由到可验证证据的可解释AIGI检测
Deng, Bowen, Zhan, Jiahui, Ji, Yikun, Yan, Haozhen, Zhang, Jianfu
Abstract
The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable and reliable. While MLLM-based detectors can provide natural language explanations, existing methods often generate speculative rationales: they rely on vague or hallucinated artifacts, miss subtle localized flaws from the latest generators, and fail to provide evidence that can be visually verified. We present Defake-o3, an explainable AIGI detector that moves from speculative rationales to verifiable evidence. It combines interactive visual search with verifier-guided evidence alignment: the model iteratively zooms into suspicious regions to inspect fine-grained details, while an Evidence Verifier, trained from human verification annotations, provides reinforcement learning rewards that favor grounded evidence and penalize baseless claims. To support this objective, we construct GroundFake, a dataset designed for grounded explainable detection, with localized bounding-box evidence, human verification based on visual grounding and artifact specificity, corrected reasoning trajectories, and valid/invalid evidence supervision. We further introduce FakeFrontier, an out-of-distribution benchmark built from real images and outputs of 10 recent generators, together with an MLLM-based protocol for evaluating evidence quality and persuasiveness. Experiments on GroundFake, FakeFrontier, and additional out-of-distribution benchmarks show that Defake-o3 improves both detection accuracy and explanation quality, producing more localized, verifiable, and persuasive evidence.
Chinese Translation
图像生成模型的快速进展呼唤不仅准确而且可解释和可靠的AI生成图像(AIGI)检测器。虽然基于MLLM的检测器可以提供自然语言解释,但现有方法往往生成推测性理由:它们依赖于模糊或虚构的伪影,错过了最新生成器的细微局部缺陷,并未提供可视化验证的证据。我们提出了Defake-o3,这是一种可解释的AIGI检测器,旨在从推测性理由转向可验证证据。它结合了交互式视觉搜索与验证者引导的证据对齐:模型迭代地放大可疑区域以检查细粒度细节,同时由人类验证注释训练的证据验证器提供强化学习奖励,以支持基于事实的证据并惩罚无根据的主张。为支持这一目标,我们构建了GroundFake,这是一个旨在实现基于事实的可解释检测的数据集,包含局部边界框证据、基于视觉基础和伪影特异性的人工验证、修正的推理轨迹,以及有效/无效证据的监督。我们进一步引入了FakeFrontier,这是一个基于真实图像和10个最新生成器输出构建的分布外基准,以及一个基于MLLM的协议,用于评估证据的质量和说服力。在GroundFake、FakeFrontier和其他分布外基准上的实验表明,Defake-o3提高了检测准确性和解释质量,生成了更具局部性、可验证性和说服力的证据。
cs.CV / 156 / 2608.16263

Seeing Before Answering: Training-Free Visual Layer Profiling for Vision-Language Models

回答之前的视觉观察:无训练的视觉层分析用于视觉-语言模型
Liu, Ruchen, Yang, Yi, Xu, Yiming, Yang, Michael Ying, Sester, Monika, Rosenhahn, Bodo
Abstract
LLaVA-style Vision-Language Models (VLMs) pass visual tokens from a fixed late layer of the vision backbone, typically the penultimate one, to the language model. We first show that this hidden convention is fragile: across 2 VLMs and 7 image and video benchmarks, the default layer is sub-optimal in 13 of 14 model-task pairs, and the best layer shifts with both task and visual backbone. Finding that layer by exhaustive layer-wise inference is prohibitively expensive, and no better fixed default exists. We therefore ask whether layer usefulness can instead be predicted from representation geometry. We study matrix-based entropy, introduced for unimodal layer analysis, which we compute over sample-level visual embeddings as Visual Dataset Entropy (VDE); and Gromov-Wasserstein (GW) distance, introduced for encoder-level VLM model selection, which we repurpose as a layer-wise visual--language alignment signal. Transferring these to LLaVA-based models is not obvious a priori: the vision tower is frozen while the multimodal projector is trained, so we profile both sides of the projector. We find that VDE transfers, and GW does not. Computed from 100 unlabeled task samples without downstream inference, pre-projector VDE tracks layer-wise accuracy and its top-ranked layers cover the oracle best layer on every task for the SigLIP-based LLaVA-Video, while giving region-level guidance for the CLIP-based Video-LLaVA. Post-projector profiles show that the projector reshapes visual geometry but does not erase the performance-relevant trend, leaving $\mathrm{VDE}_{\mathrm{pre}}$ the stronger signal. GW instead flattens after projection and is best read as an alignment diagnostic rather than a selector. VDE thus offers an interpretable, training-free policy that narrows the visual-layer search to a handful of candidates for limited downstream verification.
Chinese Translation
LLaVA 风格的视觉-语言模型(VLMs)将来自视觉主干的固定后期层(通常是倒数第二层)的视觉标记传递给语言模型。我们首先展示了这一隐藏约定的脆弱性:在 2 个 VLM 和 7 个图像及视频基准测试中,默认层在 14 个模型-任务对中有 13 个是次优的,且最佳层随着任务和视觉主干的不同而变化。通过逐层推理找到该层的过程成本过高,并且没有更好的固定默认层。因此,我们提出是否可以从表示几何中预测层的有效性。我们研究了基于矩阵的熵,该方法最初用于单模态层分析,我们在样本级视觉嵌入上计算其作为视觉数据集熵(Visual Dataset Entropy, VDE);以及 Gromov-Wasserstein(GW)距离,该方法最初用于编码器级 VLM 模型选择,我们将其重新用于层级视觉-语言对齐信号。将这些方法转移到基于 LLaVA 的模型并不是显而易见的:视觉塔被冻结,而多模态投影器正在训练,因此我们对投影器的两侧进行分析。我们发现 VDE 可以转移,而 GW 则不能。从 100 个未标记任务样本中计算的预投影器 VDE 跟踪层级准确性,其排名最高的层在每个任务中覆盖了基于 SigLIP 的 LLaVA-Video 的最佳层,同时为基于 CLIP 的 Video-LLaVA 提供了区域级指导。后投影器分析显示,投影器重塑了视觉几何,但并没有抹去与性能相关的趋势,使得 $ ext{VDE}_{ ext{pre}}$ 成为更强的信号。相反,GW 在投影后趋于平坦,最佳解读为对齐诊断而非选择器。因此,VDE 提供了一种可解释的、无训练的策略,将视觉层搜索缩小到有限的候选者,以便进行有限的下游验证。
cs.CV / 157 / 2608.16268

CoM$^3$eT: A foundation model for medical image analysis through federated, multidimensional context integration

CoM$^3$eT:通过联邦、多维上下文集成的医学图像分析基础模型
Schäfer, J. Raphael, Geissler, Kai, Nicke, Till, Tappermann, Chiara, Heber, Karoline, Petersen, Eike, Mergan, Habib, Schwen, Lars Ole, Weiss, Nick, Gerken, Annika, Moltz, Jan Hendrik, Bisson, Tom, O, Isil Dogan, Kiehl, Tim-Rasmus, Zerbe, Norman, Elezkurtaj, Sefer, Mayer, Robin S., Flinner, Nadine, Wild, Peter, Dahm, Isabel, Peisen, Felix, von Busch, Heinrich, Grimm, Robert, Arndt, Sebastian, Siegler, Lisa, May, Matthias Stefan, Prasse, Antje, Artysh, Natalia, Kiessling, Fabian, Lotz, Johannes
Abstract
Medical foundation models improve generalization when training AI models with limited labeled data, but remain confined to a single specialty, such as pathology or radiology, and to either sparse or dense outputs, such as classification or segmentation. Here, we present CoM$^3$eT (Co-representation Multidimensional Multitask Medical Transformer), a medical vision foundation model that unifies pathology and radiology, sparse and dense predictions, and two- and higher-dimensional inputs by modeling multidimensional context with attention. CoM$^3$eT outperformed other medical foundation models in an open competition spanning five tomographic, four whole-specimen, and three two-dimensional datasets, covering sparse and dense prediction tasks as well as report generation. When adapted across diverse clinical applications, training fewer than 2.5% of parameters achieved performance comparable to full fine-tuning, enabling research without access to high-performance GPU clusters. Applied to federated learning across hospitals, this approach achieved performance comparable to pooled-data training over internet connections and with consumer-grade hardware.
Chinese Translation
医学基础模型在使用有限标注数据训练人工智能模型时提高了泛化能力,但仍然局限于单一专业领域,如病理学或放射学,以及稀疏或密集输出,如分类或分割。在此,我们提出了CoM$^3$eT(协同表示多维多任务医学变换器),这是一种医学视觉基础模型,统一了病理学与放射学、稀疏与密集预测,以及二维及更高维输入,通过注意力机制建模多维上下文。CoM$^3$eT在一项涵盖五个断层扫描、四个整体标本和三个二维数据集的开放竞赛中,超越了其他医学基础模型,涉及稀疏和密集预测任务以及报告生成。当在多样的临床应用中进行适应时,训练不到2.5%的参数便达到了与完全微调相当的性能,使得在没有高性能GPU集群的情况下进行研究成为可能。应用于医院之间的联邦学习时,该方法的性能与通过互联网连接和消费级硬件进行的汇总数据训练相当。
cs.CV / 158 / 2608.16284

TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation

TransAnyText:通过结构化视觉生成在电子商务图像中翻译任意文本
Liu, Xiaoan, Ma, Lichen, Guo, Zipeng, He, Yu, Su, Xiaoyan, Guo, Shaojie, Yang, Hao, Fu, Jingling, Fu, Xiaolong, Chen, Zhen, Guo, Yu, Wang, Fei, Liu, Xinyi, Zhang, Yongjun, Zhang, Ke, Huang, Junshi
Abstract
Cross-border e-commerce image translation is essential for global retail, where product images, banners, and detail pages need to be produced in different languages. Existing methods struggle to achieve accurate translation, faithful visual identity preservation, and easy-to-edit outputs, simultaneously. To address these challenges, we introduce TransAnyText, a structured visual code framework that reformulates image text translation as generating renderable HTML patches from source images and target languages. Our framework decouples semantic generation from pixel rendering: a vision-language model (VLM) handles visual understanding, cross-lingual translation, and structured visual generation, while a diffusion model performs background inpainting and pixel-level refinement, followed by deterministic rendering to synthesize the final image. Based on this formulation, we develop a three-stage post-training framework, where supervised fine-tuning (SFT) establishes the image-to-code mapping, privilege-gap weighted self-distillation (PWSD) improves the learning of style and layout tokens, and reinforcement learning with verifiable rewards (RLVR) further optimizes task-level performance. We further introduce TransAnyDataset and TransAnyBench, a multilingual dataset and benchmark for e-commerce image translation. Extensive experiments demonstrate competitive performance against cascaded pipelines, open-source end-to-end models, and closed-source image editing systems, providing an effective, controllable, and editable solution for cross-border e-commerce image translation.
Chinese Translation
跨境电子商务图像翻译对于全球零售至关重要,其中产品图像、横幅和详细页面需要以不同语言制作。现有方法在同时实现准确翻译、忠实视觉身份保持和易于编辑输出方面面临挑战。为了解决这些问题,我们提出了TransAnyText,一个结构化视觉代码框架,将图像文本翻译重新定义为从源图像和目标语言生成可渲染的HTML补丁。我们的框架将语义生成与像素渲染解耦:一个视觉-语言模型(VLM)处理视觉理解、跨语言翻译和结构化视觉生成,而扩散模型执行背景修复和像素级细化,随后进行确定性渲染以合成最终图像。基于这一框架,我们开发了一个三阶段的后训练框架,其中监督微调(SFT)建立图像到代码的映射,特权差距加权自蒸馏(PWSD)改善风格和布局标记的学习,而具有可验证奖励的强化学习(RLVR)进一步优化任务级性能。我们还引入了TransAnyDataset和TransAnyBench,一个多语言数据集和电子商务图像翻译的基准。大量实验表明,与级联管道、开源端到端模型和闭源图像编辑系统相比,具有竞争力的性能,为跨境电子商务图像翻译提供了一种有效、可控和可编辑的解决方案。
cs.CV / 159 / 2608.16285

Audio-Visual Segmentation via Depth-Guided Collaborative Modeling

基于深度引导的协同建模的音视频分割
Fu, Zhaojin, Hong, Yuyang, Yang, Qi, Wang, Zili, Ding, Kun, Xiang, Shiming, Fan, Bin
Abstract
Audio-Visual Segmentation (AVS) is a fundamental task in multimodal perception that performs pixel-level segmentation of sounding objects in videos by leveraging both visual and audio cues. It has broad applications in video understanding, human-computer interaction, and autonomous driving. However, most existing AVS methods do not explicitly model geometric cues such as relative distance and occlusion, thereby limiting the robustness of cross-modal alignment. In human perception, spatial structure is naturally integrated with audio-visual evidence to accurately localize sounding objects. Motivated by this, we incorporate estimated depth as a spatial structural cue for AVS and propose DGCM-AVS, a tri-modal framework that jointly models audio, visual, and depth information. Specifically, we design a Depth-Aware Dynamic Modulator to improve the separation of adjacent objects while preserving intra-object feature consistency. Furthermore, we propose Depth-Guided Progressive Fusion, which uses depth as an intermediate bridge to progressively align audio cues with visual features. Compared to state-of-the-art methods, DGCM-AVS achieves relative improvements of 10.2 percent in M_J and 8.7 percent in M_F on the AVSS dataset. We believe our study highlights depth as a promising yet underexplored modality for AVS and may encourage further research in this direction.
Chinese Translation
音视频分割(Audio-Visual Segmentation, AVS)是多模态感知中的一项基础任务,通过利用视觉和音频线索对视频中发声物体进行像素级分割。它在视频理解、人机交互和自动驾驶等领域具有广泛的应用。然而,大多数现有的AVS方法并未明确建模几何线索,如相对距离和遮挡,从而限制了跨模态对齐的鲁棒性。在人类感知中,空间结构自然与音视频证据相结合,以准确定位发声物体。基于此,我们将估计的深度作为AVS的空间结构线索,并提出DGCM-AVS,一个联合建模音频、视觉和深度信息的三模态框架。具体而言,我们设计了一个深度感知动态调制器,以提高相邻物体的分离,同时保持物体内部特征的一致性。此外,我们提出了深度引导的渐进融合,利用深度作为中介桥梁,逐步对齐音频线索与视觉特征。与最先进的方法相比,DGCM-AVS在AVSS数据集上实现了M_J相对提升10.2%,M_F相对提升8.7%。我们相信我们的研究突出了深度作为AVS中一种有前景但尚未充分探索的模态,并可能鼓励在这一方向上进一步研究。
cs.CV / 160 / 2608.16289

PosterText: Towards Unified Visual Text Generation and Editing for E-commerce Poster

PosterText:面向电子商务海报的统一视觉文本生成与编辑
Liu, Xiaoan, Ma, Lichen, Guo, Zipeng, He, Yu, Su, Xiaoyan, Guo, Shaojie, Fu, Jingling, Fu, Xiaolong, Yang, Hao, Liu, Tongxuan, Guo, Yu, Wang, Fei, Liu, Xinyi, Zhang, Yongjun, Huang, Junshi
Abstract
Automated e-commerce poster design requires both high-quality poster generation and flexible editing of existing designs. However, most existing methods either target end-to-end poster generation or follow multi-stage design pipelines, with limited capability for flexible and precise editing of existing posters. To enable unified generation and editing of e-commerce posters, we introduce Text Patch Generation and Editing, a unified task formulation that treats text patches as atomic units and covers four operations: poster generation, patch addition, patch deletion, and patch modification, with optional reference-guided style control. Based on this, we propose PosterText, a unified model trained with a four-stage curriculum, including text rendering pretraining, instruction-following training, reinforcement learning for preference alignment, and spatial guidance self-distillation for execution refinement. We further construct a large-scale dataset with patch-level annotations and a comprehensive benchmark for evaluation. Extensive experiments demonstrate that PosterText achieves competitive performance against existing generation and editing approaches, validating the effectiveness of the proposed framework.
Chinese Translation
自动化电子商务海报设计需要高质量的海报生成和对现有设计的灵活编辑。然而,大多数现有方法要么专注于端到端的海报生成,要么遵循多阶段设计流程,缺乏对现有海报进行灵活和精确编辑的能力。为了实现电子商务海报的统一生成与编辑,我们提出了文本补丁生成与编辑(Text Patch Generation and Editing),这是一种将文本补丁视为原子单元的统一任务表述,涵盖四种操作:海报生成、补丁添加、补丁删除和补丁修改,并可选地进行参考引导的风格控制。在此基础上,我们提出了PosterText,这是一种经过四阶段课程训练的统一模型,包括文本渲染预训练、遵循指令的训练、用于偏好对齐的强化学习,以及用于执行精炼的空间引导自蒸馏。我们进一步构建了一个具有补丁级注释的大规模数据集,并建立了一个全面的评估基准。大量实验表明,PosterText在现有生成和编辑方法中表现出竞争力,验证了所提框架的有效性。
cs.CV / 161 / 2608.16310

Cross-View Urban Sensing: Mapping Subjective Streetscape Perception via AlphaEarth Embeddings and Urban Context

跨视角城市感知:通过AlphaEarth嵌入和城市背景映射主观街景感知
Li, Peilin, Chen, Pengfei, Wang, Jingyu, Yang, Zhifeng, Chen, Tiansheng, Gong, Mengjie, Cheng, Xiao
Abstract
Residents' perception of the urban streetscape is an important factor in public health, active mobility, and social wellbeing. Street view imagery (SVI) has emerged as a widely used data source for assessing these perceptual qualities, yet its uneven coverage and irregular updating limit large-scale measurement. Here, we present CVLNet, a Cross-View Learning Network that predicts street-level perception from AlphaEarth embeddings and multi-source urban contextual data without requiring SVI at inference. CVLNet applies per-task adaptive gating to jointly model five perceptual dimensions, using labels from the pretrained SVI-Percept model as ground truth. The proposed method is evaluated across four Southeast Asian cities: Singapore, Kuala Lumpur, Jakarta, and Manila. CVLNet achieves a median road-segment-level Adjusted $R^{2}$ of 0.76 and consistently outperforms the baseline models, with gains ranging from 5.9--11.3% across the five perceptual dimensions. Ablation experiments show that AlphaEarth features and urban contextual features contribute complementary information. We further produce citywide road-level streetscape perception maps for five subjective perceptual dimensions across all four cities, extending perception estimation from the 13--31% of the road network directly covered by available SVI to the complete road network of each city. Integrating these maps with WorldPop gridded population data, we quantify exposure inequality across population-density, demographic, and land-use groups using the Deficit Palma Ratio. These results demonstrate that remote sensing can serve as a scalable alternative to SVI for citywide streetscape perception mapping, enabling a more comprehensive assessment of urban environmental inequality.
Chinese Translation
居民对城市街景的感知是影响公共健康、积极出行和社会福祉的重要因素。街景图像(SVI)已成为评估这些感知特征的广泛使用的数据来源,但其覆盖不均和更新不规律限制了大规模测量。在此,我们提出了CVLNet,一个跨视角学习网络,能够从AlphaEarth嵌入和多源城市背景数据中预测街道级感知,而无需在推理时使用SVI。CVLNet应用每个任务的自适应门控,联合建模五个感知维度,使用预训练的SVI-Percept模型的标签作为真实值。所提方法在新加坡、吉隆坡、雅加达和马尼拉四个东南亚城市进行了评估。CVLNet在路段级别上实现了0.76的中位调整$R^{2}$,并在五个感知维度上始终优于基线模型,增益范围为5.9%至11.3%。消融实验表明,AlphaEarth特征和城市背景特征提供了互补信息。我们进一步为四个城市的五个主观感知维度生成了城市范围的道路级街景感知图,扩展了从可用SVI直接覆盖的13%至31%的道路网络的感知估计到每个城市的完整道路网络。将这些地图与WorldPop网格人口数据结合,我们使用赤字帕尔马比率量化了在人口密度、人口特征和土地利用组之间的暴露不平等。这些结果表明,遥感可以作为城市范围街景感知映射的可扩展替代方案,使城市环境不平等的评估更加全面。
cs.CV / 162 / 2608.16316

Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning

深度思维对齐:用于视频推理的轨迹级潜在蒸馏
Shen, Ao, Zhang, Yongheng, Li, Yinghui, Wang, Manning, Yin, Di, Sun, Xing
Abstract
Large Multimodal Models (LMMs) for video reasoning have long been hindered by the high computational cost of processing vast amounts of visual information. This dilemma motivates the transfer of the reasoning capabilities of large models to smaller, more efficient ones. On-Policy Distillation (OPD) offers a promising solution by matching output-token distributions along student-generated trajectories. However, video reasoning often depends on evidence accumulated across multiple frames. In this context, output-level supervision only captures information expressed through token predictions and does not directly constrain the latent representations formed during reasoning. To address this limitation, we propose Latent-OPD, which augments OPD with trajectory-level latent distillation. Specifically, our method focuses on the position at the end of each trajectory, where hidden states effectively summarize the accumulated visual evidence and reasoning context. Furthermore, we introduce a progressive teacher-lookahead strategy, which aligns middle-to-late student layers with increasingly deeper teacher layers. Experiments on six video reasoning benchmarks show that Latent-OPD consistently outperforms output-only OPD. Notably, the improvements are particularly pronounced in scenarios with limited frames, long videos, or tasks requiring complex evidence aggregation. These results establish Latent-OPD as a highly effective approach to frame-efficient video reasoning.
Chinese Translation
用于视频推理的大型多模态模型(LMMs)长期以来受到处理大量视觉信息的高计算成本的制约。这一困境促使将大型模型的推理能力转移到更小、更高效的模型上。在线蒸馏(On-Policy Distillation, OPD)通过匹配沿学生生成的轨迹的输出标记分布提供了一种有前景的解决方案。然而,视频推理通常依赖于跨多个帧积累的证据。在这种情况下,输出级别的监督仅捕捉通过标记预测表达的信息,并未直接约束在推理过程中形成的潜在表示。为了解决这一局限性,我们提出了潜在在线蒸馏(Latent-OPD),它通过轨迹级潜在蒸馏增强了OPD。具体而言,我们的方法关注于每个轨迹末端的位置,在那里隐藏状态有效地总结了积累的视觉证据和推理上下文。此外,我们引入了一种渐进式教师前瞻策略,将中后期学生层与越来越深的教师层对齐。在六个视频推理基准上的实验表明,Latent-OPD始终优于仅输出的OPD。值得注意的是,这些改进在帧数有限、视频较长或需要复杂证据聚合的任务中尤为明显。这些结果确立了Latent-OPD作为一种高效的视频推理方法。
cs.CV / 163 / 2608.16320

StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding

StreamOPD:一种具有时空线索门控的后训练方案,用于流媒体视频理解
Wu, Keming, Wang, Baoyi, Zhang, Kaichen, An, Xiang, Yang, Zuhao, Wang, Sudong, Zhu, Haowei, Huang, Tingxuan, Gao, Hongcheng, Wang, Bin
Abstract
Streaming video understanding demands direct responses from the causally observed prefix of an unfolding video. Existing systems add inference-time memory, retrieval, and compression, yet a training-free sliding-window baseline already matches them. We therefore fix a memory-free recent-window protocol and ask how far post-training alone can go. Reinforcement learning with verifiable rewards fits this regime poorly, encouraging long ``think-then-answer'' generations, while on-policy distillation (OPD) supplies dense token-level teacher supervision on student trajectories but is stable only when both models train in thinking mode. These observations lead to \textsc{StreamOPD}, a recipe combining verifiable streaming-video data, thinking-mode OPD, and instruct-mode deployment. It raises StreamingBench from $77.9\%$ to $83.9\%$---within $0.3$ points of the 9B teacher---and improves OVO-Bench excluding its hallucination-detection subtask (HLD) by $9.1$ points under unchanged inference. As a teacher-privilege extension, \emph{Spatio-Temporal CueGate (ST-CueGate)} aggregates cue-versus-no-cue teacher likelihood ratios into a group-relative response score that reweights OPD. It reaches $71.9\%$ on OVO-Bench (excluding HLD) and $64.9\%$ on Video-MME, and is the only variant that stays above the base model on all four benchmarks. Replacing the teacher with a frozen copy of the student's initial policy---on-policy self-distillation---retains most of these gains and lifts HLD to $57.0\%$, above both the untrained student and the 9B teacher, so abstention loss is not intrinsic to the recipe. We provide a transparent and reproducible reference for open-source streaming-video research.
Chinese Translation
流媒体视频理解要求对正在展开的视频的因果观察前缀做出直接响应。现有系统增加了推理时的记忆、检索和压缩,然而,一个无训练的滑动窗口基线已经能够与之匹配。因此,我们固定一个无记忆的近期窗口协议,并探讨仅依靠后训练能够达到多远。带有可验证奖励的强化学习在这种情况下表现不佳,鼓励长时间的“思考后回答”生成,而在线蒸馏(On-Policy Distillation, OPD)在学生轨迹上提供了密集的标记级教师监督,但仅在两个模型都以思考模式训练时才稳定。这些观察促成了 extsc{StreamOPD} 的提出,这是一种结合可验证流媒体视频数据、思考模式 OPD 和指令模式部署的方案。它将 StreamingBench 的性能从 $77.9\%$ 提升至 $83.9\\%$,与 9B 教师的差距仅为 $0.3$ 个百分点,并在不改变推理的情况下将 OVO-Bench 的表现提升了 $9.1$ 个百分点(不包括其幻觉检测子任务 HLD)。作为教师特权的扩展, extit{时空线索门控(Spatio-Temporal CueGate, ST-CueGate)} 将有无线索的教师可能性比率聚合为一个组相对响应分数,从而重新加权 OPD。它在 OVO-Bench(不包括 HLD)上达到了 $71.9\\%$,在 Video-MME 上达到了 $64.9\\%$,并且是唯一一个在所有四个基准上均高于基础模型的变体。用学生初始策略的冻结副本替代教师——即在线自蒸馏——保留了大部分这些增益,并将 HLD 提升至 $57.0\\%$,高于未训练的学生和 9B 教师,因此放弃损失并非该方案的内在特征。我们为开源流媒体视频研究提供了一个透明且可重复的参考。
cs.CV / 164 / 2608.16324

LaGSplat: Inferring Physics-Governed Interactive Simulation from Monocular Video Using Latent Lagrangian Gaussian Splatting

LaGSplat:基于单目视频推断物理驱动的交互式仿真使用潜在拉格朗日高斯点云
Pottier, Louen
Abstract
We present LaGSplat (Latent Lagrangian Gaussian Splatting), a framework that infers interactive, physics-governed dynamics from one or a few monocular videos. At inference it lets a user push on the filmed object, rigid or deformable, with an external force that was never measured, annotated, or seen during training. This is possible because a low-dimensional latent state $\mathbf{q} \in \mathbb{R}^d$ plays two roles at once: it is the generalised coordinate of a learned dissipative Lagrangian and the conditioning variable of a Gaussian Splatting decoder. The inductive bias of this decoder, whose primitives are explicit points $\mu_i(\mathbf{q})$ that move with the object, is what lets a force $f$ applied in the image pull back into a latent generalised force $J(\mathbf{q})^\top f$ and enter the equations of motion, which pixel-space (CNN) or neural-field (NeRF) decoders cannot do. We validate LaGSplat on test cases of increasing difficulty, from rigid to deformable and from autonomous to forced real systems, combining monocular video and sensor measurements. We further demonstrate interactive use: forces of arbitrary magnitude and direction can be applied to the reconstructed object at any time, its response rendered in real time, in 2D or 3D. Assuming a dissipative Euler-Lagrange equation over a few generalised coordinates trades generality for a bounded, plausible response to unseen forces, where an unconstrained predictor diverges.
Chinese Translation
我们提出了LaGSplat(潜在拉格朗日高斯点云),这是一个从一段或几段单目视频中推断交互式、物理驱动动态的框架。在推断过程中,用户可以对拍摄的物体施加外部力,无论是刚性物体还是可变形物体,这种外部力在训练过程中从未被测量、注释或观察到。这是可能的,因为一个低维潜在状态$ extbf{q} ext{ in } ext{R}^d$同时扮演两个角色:它是学习到的耗散拉格朗日的广义坐标,也是高斯点云解码器的条件变量。该解码器的归纳偏置,其原始元素是随物体移动的显式点$oldsymbol{ u}_i( extbf{q})$,使得施加在图像中的力$f$能够回拉到潜在的广义力$J( extbf{q})^ op f$并进入运动方程,而像素空间(CNN)或神经场(NeRF)解码器则无法做到这一点。我们在逐渐增加难度的测试案例上验证了LaGSplat,从刚性到可变形,从自主系统到受迫真实系统,结合了单目视频和传感器测量。我们进一步展示了交互式使用:可以在任何时间对重建的物体施加任意大小和方向的力,其响应实时渲染,支持2D或3D显示。假设在几个广义坐标上存在耗散的欧拉-拉格朗日方程,这在未见力的情况下以有限的、合理的响应来交换一般性,而不受约束的预测器则会发散。
cs.CV / 165 / 2608.16328

GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks

GRNEdit:从生成细化网络的新二元证据视角出发的高效通用视频编辑
Xie, Feng, Hu, Jiagao, Li, Fuhao, Wang, Zepeng, Chen, Yuxuan, Gao, Dahua, Wang, Fei, Zhou, Daiguo
Abstract
Instruction-based general video editing seeks to unify diverse editing operations within a single, intuitive interface. Existing approaches often rely on resource-intensive conditioning, using either heavyweight branches or costly source concatenation. Is there any efficient way to model editing intent? Thus, we introduce GRNEdit, a lightweight two-stage framework. GRN inspires our approach by encoding visual semantics through combinations of bits. Through task-specific fine-tuning, we take this representation further and recast editing semantics as local retain-or-flip decisions over individual bits. Source information is consequently modeled as coordinate-wise evidence supporting the observed binary states, while the GRN backbone remains responsible for resolving their global composition into coherent generative semantics. In Stage I, a compact encoder translates discrete source codes into continuous evidence signals, which GRN assimilates throughout binary refinement. Inspired by null-prompt training for classifier-free guidance, we further assign the null condition an editing-specific meaning: an empty instruction denotes no edit and is supervised through source reconstruction. This identity pathway not only implicitly strengthens evidence utilization and content preservation in Stage I, but also produces a source-preserving state in the same representation space as the edited state. Stage II can therefore directly compare each edited state with its source-preserving counterpart and use their discrepancy to revise unresolved target-bit decisions. Trained on only 0.6M pairs with less than 3\% conditioning parameters, GRNEdit-2B and GRNEdit-8B achieve scores of 4.03 and 4.18 on OpenVE-Bench. The 2B model outperforms multiple 14B open-source editors, while the 8B model performs on par with leading open-source editors.
Chinese Translation
基于指令的通用视频编辑旨在将多样的编辑操作统一在一个直观的界面中。现有的方法通常依赖于资源密集型的条件设置,使用重型分支或昂贵的源连接。有没有更高效的方式来建模编辑意图?因此,我们引入了GRNEdit,一个轻量级的两阶段框架。GRN通过位的组合编码视觉语义,启发了我们的方法。通过任务特定的微调,我们进一步推进这一表示,并将编辑语义重新表述为对单个位的局部保留或翻转决策。因此,源信息被建模为支持观察到的二元状态的坐标证据,而GRN主干则负责将其全局组合解析为一致的生成语义。在第一阶段,一个紧凑的编码器将离散源代码转换为连续证据信号,GRN在整个二元细化过程中吸收这些信号。受到无提示训练用于无分类器引导的启发,我们进一步赋予无条件特定的编辑含义:空指令表示没有编辑,并通过源重建进行监督。这一身份路径不仅在第一阶段隐式增强了证据利用和内容保留,还在与编辑状态相同的表示空间中产生了一个源保留状态。因此,第二阶段可以直接将每个编辑状态与其源保留对应状态进行比较,并利用它们的差异来修正未解决的目标位决策。GRNEdit-2B和GRNEdit-8B在仅使用60万对样本和不到3%的条件参数的情况下,在OpenVE-Bench上分别取得了4.03和4.18的得分。2B模型超越了多个14B开源编辑器,而8B模型的表现与领先的开源编辑器相当。
cs.CV / 166 / 2608.16332

Unlocking Motion in Expressions: Temporal Calibration for Referring Video Object Segmentation

解锁表达中的运动:用于参考视频对象分割的时间校准
Jiang, Yiwen, Zhu, Zhengtong, Zhang, Ruixin, Fan, Jiaqing
Abstract
Referring Video Object Segmentation (RVOS) aims to segment referred objects at the pixel level in video sequences based on natural language descriptions. Existing methods typically introduce motion information within a unified cross-modal temporal modeling framework, where language cues are used for target localization and segmentation. However, the dependency of expressions on motion semantics is not explicitly modeled, making it difficult to adaptively adjust the use of motion information according to different semantic requirements. To address these issues, we propose an Expression-driven Motion Calibration (EMC) framework for RVOS that explicitly unlocks and leverages the motion semantics within expressions. The proposed method extracts interpretable motion control signals from expressions via a Motion Signal Processing (MSP) module, and employs a Motion Influence Calibration (MIC) module to adjust the contribution of motion cues during temporal decision making. In addition, a Semantic Temporal Stage Construction (STSC) module is introduced to build expression-relevant temporal stages, providing a compact temporal candidate space for motion calibration. Through extensive evaluation on six standard benchmarks, including Ref-YouTubeVOS, Ref-DAVIS17, MeViS (valid/valid$^u$), A2D-Sentences, and JHMDB-Sentences, the superiority of our method is validated. We will release the code on https://github.com/Jeven7/EMC.
Chinese Translation
参考视频对象分割(RVOS)旨在根据自然语言描述在视频序列中对所指对象进行像素级分割。现有方法通常在统一的跨模态时间建模框架中引入运动信息,其中语言线索用于目标定位和分割。然而,表达对运动语义的依赖并未被明确建模,这使得根据不同的语义需求自适应调整运动信息的使用变得困难。为了解决这些问题,我们提出了一种以表达驱动的运动校准(EMC)框架,用于RVOS,明确解锁并利用表达中的运动语义。所提出的方法通过运动信号处理(MSP)模块从表达中提取可解释的运动控制信号,并采用运动影响校准(MIC)模块在时间决策过程中调整运动线索的贡献。此外,引入了语义时间阶段构建(STSC)模块,以构建与表达相关的时间阶段,为运动校准提供紧凑的时间候选空间。通过在六个标准基准上进行广泛评估,包括Ref-YouTubeVOS、Ref-DAVIS17、MeViS(valid/valid$^u$)、A2D-Sentences和JHMDB-Sentences,验证了我们方法的优越性。我们将会在https://github.com/Jeven7/EMC上发布代码。
cs.CV / 167 / 2608.16338

SIGMA-Lane: Scale-pyramId Gated MAmba for Temporally Consistent Video Lane Detection

SIGMA-Lane:基于尺度金字塔的门控MAmba用于时间一致的视频车道检测
Zhang, Tiancheng, Wang, Mengmeng, Gao, Yan, Kong, Xiangjie, Shen, Guojiang, Du, Jiaxin
Abstract
Video lane detection requires predictions that remain stable across frames, yet severe vehicle occlusions can break temporal cues. In streaming recurrent models, corrupted observations may enter the hidden state and produce errors that persist into later frames. Existing occlusion-aware refinements usually provide obstacle masks as auxiliary inputs, so the state-update path is only indirectly protected. We propose SIGMA-Lane, which treats this failure mode as state contamination in State Space Model (SSM)-based temporal modeling. SIGMA-Lane places occlusion-aware gates on the SSM write and residual-fusion paths, controlling how current observations enter temporal memory and are fused back after temporal propagation. After coordinate-consistent affine alignment, the model combines two complementary paths: SSM-consistent dual-gating for temporal filtering and Structural Spatial Retrieval (SSR) for recovering missing lane structure from aligned historical priors. Experiments on VIL-100 and OpenLane-V show improved temporal stability under heavy occlusion, with competitive F1 and mIoU scores.
Chinese Translation
视频车道检测要求在帧间保持稳定的预测,但严重的车辆遮挡可能会破坏时间线索。在流式递归模型中,受损的观测值可能会进入隐藏状态,并产生持续到后续帧的错误。现有的遮挡感知改进通常提供障碍物掩码作为辅助输入,因此状态更新路径仅间接受到保护。我们提出了SIGMA-Lane,将这种失败模式视为状态空间模型(State Space Model, SSM)基础的时间建模中的状态污染。SIGMA-Lane在SSM的写入和残差融合路径上设置了遮挡感知门控,控制当前观测如何进入时间记忆并在时间传播后重新融合。在坐标一致的仿射对齐后,该模型结合了两条互补路径:用于时间滤波的SSM一致的双重门控和用于从对齐的历史先验中恢复缺失车道结构的结构空间检索(Structural Spatial Retrieval, SSR)。在VIL-100和OpenLane-V上的实验表明,在严重遮挡下,时间稳定性有所改善,并且F1和mIoU得分具有竞争力。
cs.CV / 168 / 2608.16367

Depth-Dominant Skeleton Detection for Natural Scenes

自然场景中的深度主导骨架检测
Rao, Chengkun, Deng, Yixuan, Li, Min, Ou, Yangjun, Li, Ye, Luo, Ziwei, Wang, Zhaojing, Tang, Junwei, Wang, Bangchao, Yan, Xiaoyun
Abstract
To date, all natural scene skeleton detection follows the paradigm of taking RGB images as the sole input; despite notable progress, methods under this paradigm suffer significant performance degradation on complex-content images. We observe that depth images are inherently insensitive to color and texture, and can provide clear regional contours and inter-region spatial relationships, which naturally alleviates the difficulty of skeleton detection in complex scenarios. Motivated by this observation, this paper proposes for the first time a novel skeleton detection paradigm where depth images serve as the dominant modality and RGB images act as the auxiliary, and accordingly presents a model DDSkel (short for Depth-Dominant Skeleton Detection) under this paradigm. DDSkel employs an asymmetric encoder design to fuse RGB information into depth features, with the RGB modality branch having only 12% the parameters of the depth modality branch. DDSkel has a simple structure without intricate designs. Nevertheless, with only 36% of the trainable parameters of the current best method, DDSkel outperforms all state-of-the-art approaches on SymPASCAL, the most challenging dataset with a large volume of complex images.
Chinese Translation
迄今为止,所有自然场景骨架检测都遵循将RGB图像作为唯一输入的范式;尽管取得了显著进展,但在复杂内容图像下,该范式下的方法表现显著下降。我们观察到深度图像对颜色和纹理本质上不敏感,并且能够提供清晰的区域轮廓和区域间的空间关系,这自然缓解了在复杂场景中进行骨架检测的困难。受到这一观察的启发,本文首次提出了一种新颖的骨架检测范式,其中深度图像作为主导模态,RGB图像作为辅助模态,并相应地提出了该范式下的模型DDSkel(深度主导骨架检测的缩写)。DDSkel采用非对称编码器设计,将RGB信息融合到深度特征中,其中RGB模态分支的参数仅为深度模态分支的12%。DDSkel结构简单,没有复杂的设计。然而,DDSkel的可训练参数仅为当前最佳方法的36%,在SymPASCAL这一具有大量复杂图像的最具挑战性的数据集上超越了所有最先进的方法。
cs.CV / 169 / 2608.16377

Adaptive Post-Processing Drives Instance-Level Detection in Stroke Lesion Segmentation

自适应后处理驱动中风病灶分割中的实例级检测
Liu, Qinghui, Ottesen, Jon André, Bjørnerud, Atle, Emblem, Kyrre Eeg
Abstract
Instance-level lesion detection has been an increasingly larger focal point in medical image segmentation besides the more standard voxel-level overlap. Still, most pipelines are trained and post-processed for voxel overlap alone. In particular, the mismatch is most pronounced for small lesions, where a near-miss prediction---substantial overlap that falls just short of the instance-matching threshold---scores the same as a complete miss. In our ISLES'26 submission, we found that closing this gap mattered far more in post-processing than in architecture design. Our Volume-Conditioned Adaptive Post-Processing (VCAP) scheme adjusts component-size thresholds to each case's predicted lesion burden, improving Lesion-F1 by 0.032 (unbiased cross-fold estimate)---approximately 6 times larger than any architectural change we tested. A resolution-aware attention architecture (Viola2Plus), designed for small-lesion segmentation, shows why the distinction matters: it left small-lesion Dice unchanged but raised small-lesion detection rate by 3.7\%, a real effect voxel-overlap metrics alone would have missed. Under 5-fold cross-validation on the 1,453-case training set, our post-processed two-architecture ensemble achieves Dice 0.651 and Lesion-F1 0.614, versus 0.644 and 0.573 for the unprocessed single-model baseline.
Chinese Translation
实例级病灶检测在医学图像分割中逐渐成为一个更大的关注点,超越了更标准的体素级重叠。然而,大多数流程仅针对体素重叠进行训练和后处理。特别是,对于小病灶,这种不匹配最为明显,近乎命中预测——即重叠程度相当但未达到实例匹配阈值——与完全漏检的评分相同。在我们的ISLES'26投稿中,我们发现弥补这一差距在后处理中的重要性远超架构设计。我们的体积条件自适应后处理(Volume-Conditioned Adaptive Post-Processing,VCAP)方案根据每个案例的预测病灶负担调整组件大小阈值,使病灶F1分数提高了0.032(无偏交叉折叠估计),这一提升约为我们测试的任何架构变化的6倍。针对小病灶分割设计的分辨率感知注意力架构(Viola2Plus)展示了这一区别的重要性:它对小病灶的Dice系数没有变化,但小病灶检测率提高了3.7%,这是仅靠体素重叠指标无法捕捉到的真实效果。在对1453个案例的训练集进行5折交叉验证时,我们的后处理双架构集成达到了Dice 0.651和病灶F1 0.614,而未经处理的单模型基线则为0.644和0.573。
cs.CV / 170 / 2608.16380

Synthetic Data Augmentation for Satellite-Based Analysis of Battle-Damaged Agricultural Fields in Ukraine

基于卫星的乌克兰战后农业受损田地分析的合成数据增强
Sumyk, Marta, Kosovan, Oleksandr, Voitsitska, Iryna
Abstract
Monitoring war-induced damage to agricultural land in Ukraine is important for understanding threats to food security, environmental stability, and post-war recovery. However, the development of computer-vision systems for satellite-based damage analysis is limited by the scarcity of labeled imagery, especially for damaged agricultural fields. This work investigates synthetic data augmentation as a method for improving classification under limited and imbalanced training data. We train class-conditional Generative Adversarial Network (GAN) and Denoising Diffusion Probabilistic Model (DDPM) architectures on real satellite images and use them to generate additional bombed and not-bombed agricultural-field samples. The generated images are used only for training augmentation, while all downstream evaluation is performed on an exclusively real test set. A Vision Transformer classifier is trained under multiple real and synthetic data configurations to measure the practical utility of each generative approach. The best configuration, based on balanced DDPM augmentation, improves accuracy from 84\% to 88\%, balanced accuracy from 67\% to 81\%, macro F1 from 65\% to 78\%, and recall for the underrepresented not-bombed class from 41\% to 69\%. These results demonstrate the potential of synthetic satellite imagery for data-scarce geospatial applications in war-affected regions.
Chinese Translation
监测乌克兰因战争造成的农业土地损害对于理解食品安全、环境稳定性和战后恢复的威胁至关重要。然而,基于卫星的损害分析计算机视觉系统的发展受到标记图像稀缺的限制,尤其是对于受损农业田地的图像。本研究探讨了合成数据增强作为在有限和不平衡训练数据下提高分类性能的方法。我们在真实卫星图像上训练了类条件生成对抗网络(Generative Adversarial Network, GAN)和去噪扩散概率模型(Denoising Diffusion Probabilistic Model, DDPM)架构,并利用它们生成额外的受轰炸和未受轰炸的农业田地样本。生成的图像仅用于训练增强,而所有下游评估均在完全真实的测试集上进行。我们在多种真实和合成数据配置下训练了视觉变换器(Vision Transformer)分类器,以衡量每种生成方法的实际效用。基于平衡的DDPM增强的最佳配置将准确率从84%提高到88%,平衡准确率从67%提高到81%,宏观F1值从65%提高到78%,以及对代表性不足的未受轰炸类别的召回率从41%提高到69%。这些结果展示了合成卫星图像在战争影响地区数据稀缺的地理空间应用中的潜力。
cs.CV / 171 / 2608.16384

Self-Routed Tensor Adapters for Parameter-Efficient Universal Visual Adaptation

自我路由张量适配器用于参数高效的通用视觉适应
Yadav, Suraj
Abstract
Universal visual representations require adaptation mechanisms that adapt across heterogeneous domains without fragmenting knowledge into domain-specific modules. Parameter-efficient fine-tuning adapts frozen visual foundation models efficiently, but standard low-rank adapters use a fixed subspace for all inputs, which can be restrictive when domains differ in style, background, and semantic context. MoE-based adapters improve specialization through multiple expert pathways, but often rely on external routers and large expert banks, adding parameters and separating routing from adaptation. We propose \textbf{Self-Routed Tensor Adapters}, a compact framework for multi-domain visual adaptation. SRTA projects each input into a low-rank space, computes routing weights from this representation using a learnable domain matrix, and uses these weights to blend slices of a shared Tucker core. This produces a sample-specific adaptation matrix without an external gating network, allowing shared visual factors to be reused while supporting domain-aware specialization. To strengthen pathway learning, we introduce a progressive depth-weighted routing objective that supervises routing decisions across adapter layers. Across five heterogeneous multi-domain visual classification benchmarks, SRTA achieves competitive or slightly stronger average accuracy than MoE-style PEFT baselines while using substantially fewer trainable parameters. At rank 64, SRTA uses 2.77M parameters in the 4-domain setting compared with 9.52M for MoLoRA, and 3.00M in the 6-domain setting compared with 14.31M. Overall, SRTA offers an effective accuracy-parameter trade-off for adapting visual foundation models toward universal multi-domain representations. \href{https://github.com/surajyadav-research/SRTA}{GitHub}
Chinese Translation
通用视觉表示需要适应机制,以便在异构领域之间进行适应,而不将知识碎片化为特定领域的模块。参数高效的微调能够高效地适应冻结的视觉基础模型,但标准的低秩适配器对所有输入使用固定的子空间,这在领域风格、背景和语义上下文差异较大时可能会受到限制。基于专家混合(MoE)的适配器通过多个专家路径提高专业化,但通常依赖外部路由器和大型专家库,增加了参数并将路由与适应分离。我们提出了 extbf{自我路由张量适配器}(Self-Routed Tensor Adapters,SRTA),这是一个紧凑的多领域视觉适应框架。SRTA将每个输入投影到低秩空间,从该表示中使用可学习的领域矩阵计算路由权重,并利用这些权重混合共享的Tucker核心切片。这生成了一个特定样本的适应矩阵,而无需外部门控网络,从而允许共享视觉因素的重用,同时支持领域感知的专业化。为了加强路径学习,我们引入了一种渐进的深度加权路由目标,以监督适配器层之间的路由决策。在五个异构多领域视觉分类基准测试中,SRTA在使用显著更少的可训练参数的情况下,达到了与MoE风格的PEFT基线相当或略强的平均准确率。在64秩的情况下,SRTA在4领域设置中使用了2.77M参数,而MoLoRA使用了9.52M;在6领域设置中,SRTA使用了3.00M参数,而MoLoRA使用了14.31M。总体而言,SRTA为将视觉基础模型适应于通用多领域表示提供了有效的准确率与参数权衡。
cs.CV / 172 / 2608.16424

Joint Flow Matching Enables Continuous Dose-Conditioned Cell Morphing

联合流匹配实现连续剂量条件下的细胞形态变换
Bogensperger, Lea, Merlo, Manuela, Baumgartner, Martin, Krauthammer, Michael, Ciraulo, Bernard
Abstract
Generative modeling has shown increasing promise for predicting cellular perturbation effects under chemical compound treatments. Existing approaches either model perturbation as a distribution-to-distribution mapping without explicit concentration handling, or treat concentration as a discrete class label, precluding continuous dose control. We introduce a joint flow matching approach that simultaneously models cell latents and drug concentration via a dual-timestep formulation, enabling dose-conditioned single-cell morphing through the invertibility of flow matching. The joint formulation induces a monotonic dose-response geometry in latent space and additionally supports concentration estimation from cell morphology. As proof of concept, we further demonstrate generalization to an unseen dose held out during training. Empirically, our method achieves competitive or improved per-concentration metrics on two compounds compared with representative baselines, while enabling capabilities structurally unavailable to discrete-class methods.
Chinese Translation
生成建模在预测化合物处理下的细胞扰动效应方面显示出越来越大的潜力。现有的方法要么将扰动建模为分布到分布的映射,而没有明确处理浓度,要么将浓度视为离散类别标签,从而排除了连续剂量控制。我们提出了一种联合流匹配方法,通过双时间步的形式同时建模细胞潜变量和药物浓度,从而通过流匹配的可逆性实现剂量条件下的单细胞形态变换。联合形式在潜在空间中引入了单调的剂量-反应几何,并且支持从细胞形态估计浓度。作为概念验证,我们进一步展示了对训练过程中未见剂量的泛化。实证结果表明,我们的方法在两个化合物上相较于代表性基线在每个浓度指标上实现了具有竞争力或改进的表现,同时实现了离散类别方法结构上无法提供的能力。
cs.CV / 173 / 2608.16457

Contrastive Energy Fields for Inference-Time Procedure Planning in Instructional Videos

对比能量场在指令视频中的推理时间过程规划
Afham, Mohamed, Reich, Christoph, Hahn, Oliver, Cremers, Daniel, Roth, Stefan
Abstract
Procedure planning seeks to estimate a sequence of actions to transition from an observed initial state to a given goal state. Current procedure planning approaches directly predict action sequences from latent representations using feed-forward neural networks or diffusion-based inference. These paradigms treat every action as plausible, lacking the ability to enforce task-specific logical constraints that render certain actions irrelevant or not plausible. We propose CEFITO, a procedure planning approach that learns a predictor to express an action-conditioned representation space. Based on this representation space, we formulate procedure planning as a task-constrained optimization problem. Unlike prior methods, CEFITO explicitly reasons over the action space by omitting irrelevant actions during inference-time planning. This reformulation enables effective procedure planning and achieves state-of-the-art accuracy on two established procedure planning benchmarks.
Chinese Translation
过程规划旨在估计一系列动作,以便从观察到的初始状态过渡到给定的目标状态。目前的过程规划方法通过前馈神经网络或基于扩散的推理直接从潜在表示预测动作序列。这些范式将每个动作视为可能的,缺乏强制执行特定任务逻辑约束的能力,从而使某些动作变得无关或不可信。我们提出了CEFITO,一种过程规划方法,学习一个预测器以表达一个基于动作的条件表示空间。基于该表示空间,我们将过程规划形式化为一个任务约束优化问题。与之前的方法不同,CEFITO在推理时间规划中明确地对动作空间进行推理,通过省略无关动作来实现。这种重新表述使得有效的过程规划成为可能,并在两个已建立的过程规划基准上达到了最先进的准确性。
cs.CV / 174 / 2608.16463

Shared-Structure 4D Spectral Gaussian Representation for Sparse-View Spectral CT Reconstruction

用于稀视谱CT重建的共享结构4D谱高斯表示
Fang, Jiancheng, Wang, Shaoyu, Xia, Wenjun, Chen, Yang, Liu, Qiegen
Abstract
Sparse-view spectral computed tomography (CT) reconstructs energy-resolved attenuation volumes from limited projection views, requiring simultaneous handling of angular undersampling and spectral coupling. We propose a SharedStructure 4D Spectral Gaussian Representation (4D-SG) that learns shared Gaussian geometry from full spectrum structural projections and uses a Gaussian-wise Spectral Density Curve Network (GSC-Net) to predict Gaussian raw density transformations. This factorization separates shared spatial structure from spectral attenuation variation, avoids independent channel geometry optimization, and establishes a continuous 4D-SG representation from discrete spectral measurements for unobserved spectral channel queries. Experiments on six synthesized, simulated projection, and real projection datasets with 50 views demonstrate the best average performance. Compared with the strongest Gaussian baseline, 4D-SG improves PSNR from 35.56 dB to 36.61 dB, increases SSIM from 0.909 to 0.914, and reduces LPIPS from 0.208 to 0.194, demonstrating its effectiveness for sparse-view spectral CT reconstruction.
Chinese Translation
稀视谱计算机断层扫描(CT)从有限的投影视图重建能量分辨的衰减体积,需同时处理角度欠采样和谱耦合。我们提出了一种共享结构4D谱高斯表示(4D-SG),该方法从全谱结构投影中学习共享的高斯几何,并使用高斯级谱密度曲线网络(GSC-Net)预测高斯原始密度变换。这种因式分解将共享空间结构与谱衰减变化分离,避免了独立通道几何优化,并为未观测的谱通道查询建立了从离散谱测量到连续4D-SG表示的转换。在六个合成的、模拟的投影和真实投影数据集(50个视图)上的实验表明,该方法具有最佳的平均性能。与最强的高斯基线相比,4D-SG将PSNR从35.56 dB提高到36.61 dB,将SSIM从0.909增加到0.914,并将LPIPS从0.208降低到0.194,证明了其在稀视谱CT重建中的有效性。
cs.CV / 175 / 2608.16469

Sterilizable Scene Graph Generation for Operating Rooms

可消毒的手术室场景图生成
Lemke, Nick, Sivakumar, Ssharvien Kumar, Sanner, Antoine P., Kalkhof, John, Krumb, Henry John, Ghazaei, Ghazal, Mukhopadhyay, Anirban
Abstract
Scene graph generation from surgical video enables a holistic and structured understanding of surgical scenes by modeling objects and their semantic relationships. Despite recent advances, state-of-the-art approaches rely on large, parameter-heavy deep learning models that are impractical for deployment in the operating room (OR) due to hardware footprint, hygiene constraints, latency, and data privacy concerns. To the best of our knowledge, this is the first scene graph generation method built on NCAs and the first NCA framework capable of learning structured representations. We introduce SG-NCA, a lightweight scene graph generation framework based on Neural Cellular Automata (NCA), designed for inference in fanless devices critical for OR hygiene protocols. SG-NCA is the first scene graph generation combining NCA-based multiclass segmentation for efficient object detection and feature extraction with a lightweight relation predictor. We evaluate SG-NCA on videos of cataract surgery and cholecystectomy, demonstrating performance comparable to established baselines while requiring 55x fewer parameters. We showcase deployment on fanless edge devices better suited for the OR and demonstrate downstream applications such as surgical video captioning, highlighting SG-NCA's potential for affordable, privacy-preserving, and OR-ready intraoperative scene understanding.
Chinese Translation
从手术视频生成场景图通过建模对象及其语义关系,实现对手术场景的整体和结构化理解。尽管近年来取得了一些进展,最先进的方法仍依赖于大型、参数密集的深度学习模型,这在手术室(OR)中由于硬件占用、卫生限制、延迟和数据隐私问题而难以实际应用。据我们所知,这是首个基于神经细胞自动机(NCA)的场景图生成方法,也是首个能够学习结构化表示的NCA框架。我们提出了SG-NCA,这是一个基于NCA的轻量级场景图生成框架,旨在在无风扇设备上进行推理,以满足手术室卫生协议的要求。SG-NCA是首个结合NCA基础的多类分割进行高效对象检测和特征提取的场景图生成方法,并配备轻量级关系预测器。我们在白内障手术和胆囊切除术的视频上评估了SG-NCA,结果显示其性能与已建立的基准相当,同时所需参数减少了55倍。我们展示了在更适合手术室的无风扇边缘设备上的部署,并展示了下游应用,如手术视频字幕生成,突显了SG-NCA在经济实惠、保护隐私和适合手术室的术中场景理解方面的潜力。
cs.CV / 176 / 2608.16480

RISE: Roadside Infrastructure Sequence Understanding across 3D Tracking and Structured Vision-Language Reasoning

RISE:基于3D跟踪和结构化视觉-语言推理的路边基础设施序列理解
Jiang, Yanbo, Zheng, Haotian, Wang, Jiahao, Ren, Hanxiao, Xu, Yitao, Xing, Yining, Ke, Zehong, Cheng, Hao, Tu, Yiqian, Li, Jinhao, Xuan, Zhiyuan, Zhang, Fang, Wang, Jianqiang
Abstract
We present RISE (Roadside Infrastructure Sequence Understanding and Evaluation), a framework spanning metric 3D tracking and structured vision-language reasoning in roadside sequences. For metric tracking, our image-only method combines SAM3 video identities with calibration-guided mask agreement for multi-view identity association, recovering persistent 3D tracks without LiDAR or task-specific 3D training. Its calibration-conditioned geometry allows the procedure to be instantiated at different calibrated multi-camera intersections without layout-specific retraining. On 20 human-reviewed clips from six intersections, the generated tracks achieve 66.9 MOTA within the defined multi-view evaluation scope. For structured vision-language reasoning, a human-reviewed MLLM pipeline mines high-value clips and uses a constrained full-context Oracle to construct bbox-grounded predictive QA without exposing future evidence to evaluated models. The resulting RISE-VQA dataset contains 33,910 QA pairs from 557 clips across 16 intersections and 61 roadside views. Its intersection-held-out RISE-Bench evaluates semantic choices, coordinates, future boxes, and interaction sets with deterministic task-specific metrics. Experiments show consistent benefits from domain adaptation and generally from temporal context, while revealing persistent challenges in spatial grounding, future localization, and interaction reasoning.
Chinese Translation
我们提出了RISE(路边基础设施序列理解与评估),这是一个涵盖度量3D跟踪和结构化视觉-语言推理的框架,专注于路边序列。在度量跟踪方面,我们的图像-only方法结合了SAM3视频身份和校准引导的掩码一致性,用于多视角身份关联,能够在没有LiDAR或特定任务3D训练的情况下恢复持久的3D轨迹。其校准条件几何体允许该过程在不同校准的多摄像头交叉口中实例化,而无需特定布局的再训练。在来自六个交叉口的20个人工审核剪辑中,生成的轨迹在定义的多视角评估范围内达到了66.9的MOTA。在结构化视觉-语言推理方面,人工审核的MLLM管道挖掘高价值剪辑,并使用约束的全上下文Oracle构建基于边界框的预测问答,而不向被评估模型暴露未来证据。最终生成的RISE-VQA数据集包含来自16个交叉口和61个路边视图的557个剪辑中的33,910个问答对。其交叉口保留的RISE-Bench评估语义选择、坐标、未来边界框和交互集,使用确定性的任务特定指标。实验表明,领域适应和时间上下文的一致性收益,同时揭示了空间定位、未来定位和交互推理方面的持续挑战。
cs.CV / 177 / 2608.16484

Remote-Sensing City Layout Extraction with MLLM

基于多模态大语言模型的遥感城市布局提取
Zhou, Zigan, Li, Kai, Deng, Yupeng
Abstract
Remote-sensing systems usually describe urban content with detection boxes, semantic masks, or vector boundaries. Such outputs locate classes and support image-plane scoring, yet they do not by themselves constitute an executable layout that retains object identities, typed relations, topology, and regeneration rules. Code-as-City instead casts urban-layout extraction from a single top-down image as constrained code generation with a multimodal large language model (MLLM). An image model first produces an aligned five-class semantic layout prior. Three ordered MLLM passes use the image and this prior to recover roads, land-cover regions and relations, and buildings. Deterministic normalization converts the accumulated records into a city graph and a restricted layout program. Executing the program creates a renderable 3D city layout and an orthographic semantic projection over shared geometry. The projection admits pixel-level comparison with remote-sensing masks, while named objects, relations, and editing operations remain available for synchronized regeneration of both views. Evaluated on the 100 scenes of CityLayout-100, the complete framework obtains 41.1% mean intersection-over-union and 48.3% global intersection-over-union. This result provides quantitative evidence that visual observations can be translated into inspectable, editable city code with coupled planar and 3D outputs.
Chinese Translation
遥感系统通常通过检测框、语义掩膜或矢量边界来描述城市内容。这些输出能够定位类别并支持图像平面的评分,但它们本身并不构成一个可执行的布局,无法保留对象身份、类型关系、拓扑结构和再生规则。相较之下,Code-as-City将从单一自上而下图像提取城市布局视为一种受限代码生成,利用多模态大语言模型(MLLM)。图像模型首先生成一个对齐的五类语义布局先验。随后,三个有序的MLLM处理利用图像及该先验来恢复道路、土地覆盖区域及其关系,以及建筑物。确定性归一化将累积记录转换为城市图和受限布局程序。执行该程序生成可渲染的3D城市布局和共享几何体上的正交语义投影。该投影允许与遥感掩膜进行像素级比较,同时命名对象、关系和编辑操作可用于同步再生这两种视图。在CityLayout-100的100个场景上进行评估,完整框架获得41.1%的平均交并比和48.3%的全局交并比。该结果提供了定量证据,表明视觉观察可以转化为可检查、可编辑的城市代码,并生成耦合的平面和3D输出。
cs.CV / 178 / 2608.16485

HiFi-BRep: High-Fidelity Latent Representation for Robust B-Rep Generation

HiFi-BRep:用于稳健边界表示生成的高保真潜在表示
Hou, Junhao, Luo, Chenqi, Wang, Pufan, Lu, Jiaying, Liu, Yusheng, Qin, Feiwei, Fang, Meie, Zhou, Kun
Abstract
Boundary representation (B-Rep) generation is a fundamental task in computer-aided design, yet the direct synthesis of high-fidelity and structurally valid B-Reps remains a major challenge. Existing deep generative methods suffer from two forms of brittleness: representation brittleness, caused by padding noise and feature contamination in the latent space, and generation brittleness, stemming from sequential error propagation and a train-inference mismatch due to non-differentiable validity enforcement. We propose HiFi-BRep, a novel framework that addresses these limitations through two synergistic contributions. First, a topology-aware encoder constructs a high-fidelity latent representation by eliminating padding via learnable queries and preventing feature contamination with topology-guided attention. Second, a single-stage decoder jointly predicts geometry and topology in parallel, embedding core manifold constraints as a differentiable learning objective. This design ensures mutual guidance between geometry and topology while avoiding cascaded errors. Extensive experiments show that HiFi-BRep significantly outperforms state-of-the-art methods in both structural validity and geometric fidelity, providing a robust solution for high-quality B-Rep synthesis. Code and models are publicly available at https://github.com/1nnoh/HiFi-BRep.
Chinese Translation
边界表示(B-Rep)生成是计算机辅助设计中的一项基础任务,但高保真且结构有效的B-Rep的直接合成仍然是一个重大挑战。现有的深度生成方法存在两种脆弱性:表示脆弱性,由潜在空间中的填充噪声和特征污染引起;生成脆弱性,由于不可微分的有效性强制导致的顺序错误传播和训练-推理不匹配。我们提出了HiFi-BRep,这是一个新颖的框架,通过两项协同贡献来解决这些限制。首先,一个拓扑感知编码器通过可学习查询消除填充,并通过拓扑引导的注意力防止特征污染,从而构建高保真的潜在表示。其次,一个单阶段解码器并行预测几何和拓扑,将核心流形约束嵌入为可微分的学习目标。该设计确保几何和拓扑之间的相互指导,同时避免级联错误。大量实验表明,HiFi-BRep在结构有效性和几何保真度方面显著优于最先进的方法,为高质量B-Rep合成提供了稳健的解决方案。代码和模型可在 https://github.com/1nnoh/HiFi-BRep 获取。
cs.CV / 179 / 2608.16490

Towards Real-Time and Adaptable LiDAR Scene Completion

面向实时和自适应的激光雷达场景补全
Hussian, Azhar, Vossiek, Martin, Belagiannis, Vasileios
Abstract
LiDAR scene completion is a key component of 3D perception in autonomous driving, where the scene must be completed in real time to be usable in downstream tasks. Existing approaches typically follow an initialize-and-refine paradigm, in which a coarse initialization of the scene is first constructed, then refined into complete 3D geometry. Generative models are slower because they iteratively refine random Gaussian noise into the scene, while non-generative methods perturb the partial scene with a fixed noise scale, which limits coverage of large gaps and occluded regions and requires manual recalibration for each new sensor configuration. We present RapidLiDAR, a LiDAR scene completion method that treats the initialization itself as a learned, data-driven component. We propose an adaptive initialization module that predicts a spatially varying displacement for each partial input point, expanding the partial observations into a coarse scene initialization adapted to the local geometry, without requiring manual noise tuning. To refine this coarse initialization into a complete and coherent scene, we additionally propose a multi-scale reconstruction module that further refines point positions by querying multi-scale 3D voxel and 2D BEV feature maps constructed from the input scan. By replacing point-neighborhood operators such as farthest point sampling and $k$-nearest neighbor search with voxel- and BEV-based feature extraction, our architecture is faster and can handle different input resolutions by design. Experiments on SemanticKITTI and KITTI-360 show that our method achieves completion performance on par with the state of the art while completing a full scene in 0.1 seconds, which is 2.3 times faster than the fastest prior method. This matches the 10 Hz acquisition rate of typical automotive LiDAR sensors, taking a step toward real-time LiDAR scene completion.
Chinese Translation
激光雷达场景补全是自动驾驶中3D感知的关键组成部分,场景必须实时完成以便在下游任务中使用。现有方法通常遵循初始化-细化的范式,首先构建场景的粗略初始化,然后将其细化为完整的3D几何形状。生成模型较慢,因为它们通过迭代地将随机高斯噪声细化为场景,而非生成方法则使用固定的噪声规模扰动部分场景,这限制了对大间隙和遮挡区域的覆盖,并且需要对每个新的传感器配置进行手动重新校准。我们提出了RapidLiDAR,一种将初始化本身视为学习的、数据驱动的组件的激光雷达场景补全方法。我们提出了一种自适应初始化模块,该模块为每个部分输入点预测空间变化的位移,将部分观测扩展为适应局部几何形状的粗略场景初始化,而无需手动调节噪声。为了将这种粗略初始化细化为完整且一致的场景,我们还提出了一种多尺度重建模块,通过查询从输入扫描构建的多尺度3D体素和2D鸟瞰图(BEV)特征图,进一步细化点的位置。通过用基于体素和BEV的特征提取替换点邻域操作符,如最远点采样和$k$-最近邻搜索,我们的架构更快,并且能够设计上处理不同的输入分辨率。在SemanticKITTI和KITTI-360上的实验表明,我们的方法在完成性能上与最先进的技术相当,同时在0.1秒内完成整个场景,比最快的先前方法快2.3倍。这与典型汽车激光雷达传感器的10 Hz采集率相匹配,朝着实时激光雷达场景补全迈出了重要一步。
cs.CV / 180 / 2608.16513

MLLM-Guided Semantic Correction for Text-to-Video Generation

基于MLLM的文本到视频生成的语义修正
Chen, Junhao, Lv, Zheqi, Yin, Keting, Zhang, Shengyu, Zhao, Zhou, Chen, Feiyang, Duan, Xinyu, Huai, Baoxing, Wu, Fei
Abstract
Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation. However, these models often suffer from semantic errors such as missing objects, incorrect attributes, or mismatched actions. Although some semantic correction methods perform optimization before sampling or refinement after sampling, how to detect and correct semantic deviations during the video generation process remains underexplored. In this paper, we introduce a training-free, interpretable mid-generation correction framework that integrates multimodal large language model (MLLM) feedback directly into the diffusion sampling loop. Our framework achieves diffusion trajectory correction by injecting semantic evaluation signals during video synthesis, enabling the model to optimize the generated content through continuous self-reflection. We propose two key modules: a Semantic Assessment Supervisor that generates intermediate preview frames for semantic evaluations and deviation diagnostics, and a Semantic Modification Assistant that corrects semantic drift during inference via a controllable latent trajectory intervention. Our method improves semantic alignment, visual fidelity, and temporal consistency without modifying model parameters. We validate the effectiveness of our approach through extensive experiments across multiple benchmarks.
Chinese Translation
近年来,扩散模型和Transformer架构的进展使得文本到视频生成取得了显著进展。然而,这些模型常常存在语义错误,例如缺失对象、属性不正确或动作不匹配。尽管一些语义修正方法在采样前进行优化或在采样后进行精细化,但在视频生成过程中如何检测和修正语义偏差仍然未被充分探索。本文提出了一种无训练、可解释的中间生成修正框架,该框架将多模态大语言模型(MLLM)的反馈直接集成到扩散采样循环中。我们的框架通过在视频合成过程中注入语义评估信号来实现扩散轨迹修正,使模型能够通过持续的自我反思来优化生成内容。我们提出了两个关键模块:一个生成中间预览帧以进行语义评估和偏差诊断的语义评估监督器,以及一个通过可控的潜在轨迹干预在推理过程中修正语义漂移的语义修正助手。我们的方法在不修改模型参数的情况下,提高了语义一致性、视觉保真度和时间一致性。我们通过在多个基准上进行广泛实验验证了我们方法的有效性。
cs.CV / 181 / 2608.16514

Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans

匹配结果,分歧视角:聚焦的多模态大型语言模型(MLLMs)与人类的搜索比较
Kerkouri, Mohamed Amine, Tliba, Marouane, Chetouani, Aladine, Bagci, Ulas, Bruno, Alessandro
Abstract
Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath. Whether multimodal large language models (MLLMs), given the same foveated input, search as humans do bears on their use as models of human vision and on attention-alignment scores. We compare three general-purpose MLLMs with human eye-movement scanpaths on goal-directed search (COCO-Search18), driving each model fixation by fixation through an identical, human-matched foveated view and assessing it along three axes: the decision of target presence, the efficiency of reaching the target, and the gaze process itself. The axes dissociate. On the decision and on target acquisition the models match or exceed humans, detecting present targets near ceiling and reaching them on the first saccade more often than people do. The gaze process is not human. Under the human-matched condition, all three share one signature: low-entropy, large-amplitude, self-consistent scanpaths that agree with themselves far more closely than two humans agree with each other. That is consistent with a single-pass, non-serial architecture rather than a limit of acuity. Matched retinal input reproduces where humans look but not how the looking unfolds in time, and no degradation regime recovers human-like search at human-like success. The gap sits on a process axis that answer-alignment and saliency metrics do not measure. Because they miss it, such metrics cannot certify human-like vision, and zero-shot models suit outcome and spatial questions but not temporal, process-level ones.
Chinese Translation
人类视觉搜索是串行的:视网膜中央凹必须落在候选对象上以进行确认,这些落点形成了扫描路径。多模态大型语言模型(MLLMs)在相同的聚焦输入下是否以人类的方式进行搜索,关系到它们作为人类视觉模型的有效性及注意力对齐评分。我们比较了三种通用的MLLMs与人类在目标导向搜索(COCO-Search18)中的眼动扫描路径,通过逐个注视点驱动每个模型在相同的人类匹配的聚焦视图中,并从三个维度进行评估:目标存在的决策、到达目标的效率以及注视过程本身。这三个维度是分离的。在目标存在的决策和目标获取上,模型的表现与人类相匹配或超越,能够在接近上限的情况下检测到存在的目标,并且比人类更频繁地在第一次眼跳中到达目标。注视过程则并非人类所特有。在人类匹配条件下,所有三个模型共享一个特征:低熵、大幅度、自洽的扫描路径,它们之间的一致性远高于两个人类之间的相互一致性。这与单次通过的非串行架构一致,而非清晰度的限制。匹配的视网膜输入再现了人类的注视位置,但未能展现注视随时间展开的过程,且没有任何退化机制能够恢复人类般的搜索表现和成功率。这一差距存在于一个过程维度上,而答案对齐和显著性度量并未对此进行衡量。由于它们未能捕捉到这一点,因此这些度量无法证明人类般的视觉能力,零-shot模型适合于结果和空间问题,但不适合于时间和过程层面的问题。
cs.CV / 182 / 2608.16523

FLEET: Token-Based Feature Extraction for Event Camera-based Reinforcement Learning

FLEET:基于令牌的事件摄像头强化学习特征提取
Gottwald, Tristan, Schier, Maximilian, Schaller, Melanie, Rosenhahn, Bodo
Abstract
Event cameras generate asynchronous, high-frequency data streams offering spatially sparse information at lower latency than traditional cameras.In principle, these properties should be ideal for the design of control policies.However, reinforcement learning research in this field remains limited as existing approaches fail to fully exploit the sensor's properties.CNN-based methods negate the sensors benefits by aggregating events into sparse grids. This couples compute cost to sensor resolution and blurs the temporal information. Meanwhile, existing generative baselines rely on the availability of trajectory data to pretrain the model. We propose FLEET (Feature Learning from Events via Efficient Tokenization), a feature extractor that processes event sequences directly. Leveraging random Fourier features and cross-attention, our architecture compresses variable streams into fixed-size latent representations. This decouples inference cost of the feature extractor's backbone from the sensor's resolution, enabling end-to-end learning without auxiliary losses. We validate FLEET on a new, high-throughput benchmark. The results demonstrate that our sequence-based approach surpasses SOTA performance and exhibits superior robustness to variations in observation frequencies.
Chinese Translation
事件摄像头生成异步、高频数据流,提供比传统摄像头更低延迟的空间稀疏信息。从原则上讲,这些特性应当理想于控制策略的设计。然而,该领域的强化学习研究仍然有限,因为现有方法未能充分利用传感器的特性。基于卷积神经网络(CNN)的方法通过将事件聚合到稀疏网格中,削弱了传感器的优势。这将计算成本与传感器的分辨率耦合在一起,并模糊了时间信息。同时,现有的生成基线依赖于轨迹数据的可用性来预训练模型。我们提出了FLEET(通过高效令牌化从事件中学习特征),这是一种直接处理事件序列的特征提取器。利用随机傅里叶特征和交叉注意力,我们的架构将可变流压缩为固定大小的潜在表示。这将特征提取器主干的推理成本与传感器的分辨率解耦,从而实现无辅助损失的端到端学习。我们在一个新的高吞吐量基准上验证了FLEET。结果表明,我们的基于序列的方法超越了当前最优性能,并在观察频率变化方面表现出更强的鲁棒性。
cs.CV / 183 / 2608.16535

Automatic Cephalometric Landmark Localization on CBCT-Derived Digitally Reconstructed Radiographs for Skeletal Malocclusion Classification

基于CBCT衍生数字重建放射影像的自动头影测量标志定位用于骨骼错牙合分类
Hou, Benjamin, Almpani, Konstantinia, Lee, Janice S., Lu, Zhiyong
Abstract
Manual cephalometric landmark annotation is important for craniofacial assessment but is labor-intensive and difficult to scale. We introduce CephViT, a Vision Transformer-based model for automated 2D lateral cephalometric landmark localization, and evaluate its use in downstream skeletal malocclusion classification. CephViT was trained and benchmarked on a public lateral cephalogram dataset, achieving a mean radial error of 1.28 +/- 1.42 mm and a successful detection rate of 92.0% at 3.0 mm. Because the private evaluation cohort consisted of 3D CBCT scans, lateral cephalogram-like digitally reconstructed radiographs (DRRs) were generated from each volume and used as 2D inputs to the landmark localization model. Landmark coordinates were normalized into a common coordinate frame, and skeletal malocclusion classification was performed using landmarks shared between the reference and DRR-based pipelines. Classification performance using DRR-localized landmarks was comparable to that obtained using manually annotated reference landmarks, with accuracies of 70.0% and 68.3%, respectively. These results support the feasibility of automated cephalometric analysis on CBCT-derived DRRs for skeletal malocclusion assessment.
Chinese Translation
手动头影测量标志注释对于颅面评估至关重要,但劳动密集且难以扩展。我们引入了CephViT,一种基于视觉变换器(Vision Transformer)的模型,用于自动化的2D侧面头影测量标志定位,并评估其在下游骨骼错牙合分类中的应用。CephViT在一个公共的侧面头影测量数据集上进行了训练和基准测试,达到了1.28 +/- 1.42毫米的平均径向误差和92.0%的成功检测率(在3.0毫米内)。由于私有评估队列由3D CBCT扫描组成,因此从每个体积生成了类似侧面头影测量的数字重建放射影像(DRR),并用作标志定位模型的2D输入。标志坐标被标准化到一个共同的坐标框架中,并使用参考和基于DRR的管道之间共享的标志进行骨骼错牙合分类。使用DRR定位的标志的分类性能与使用手动注释的参考标志的性能相当,准确率分别为70.0%和68.3%。这些结果支持在CBCT衍生的DRR上进行自动头影测量分析以评估骨骼错牙合的可行性。
cs.CV / 184 / 2608.16546

Supervising the Path to Fine Scales: GalerkinFlow for Scientific-Field and Image Super-Resolution

监督细尺度的路径:用于科学场和图像超分辨率的GalerkinFlow
Zhan, Zikang
Abstract
Most super-resolution models learn from paired data by supervising only the final high-resolution output. This provides little control over how the prediction should evolve between the downsampled observation and its fine target. We introduce GalerkinFlow, an equation-agnostic framework that turns each coarse--fine pair into supervision along an entire reconstruction path. At a random sample of intermediate states on the reconstruction path, the model predicts the coarse-to-fine residual velocity and uses coarse-anchor point to define a pseudo-endpoint. We show that the reconstruction loss of this pseudo-endpoint is exactly related to the intermediate velocity loss through a known time-dependent weight. Consequently, every intermediate state contributes supervision toward the same fine target, rather than serving only as an internal step toward an endpoint loss. Because intermediate states already reveal part of the missing fine-scale structure, we additionally supervise the coarse endpoint used during one-step inference. A finite-difference objective further constrains local spatial variation. GalerkinFlow combines convolutional features with scale-conditioned Galerkin operator mixing and requires no governing equation or physical metadata. It achieves the lowest raw-space errors among the evaluated equation-agnostic baselines on Navier--Stokes and Darcy Flow, while remaining competitive on DIV2K.
Chinese Translation
大多数超分辨率模型通过仅监督最终的高分辨率输出来从配对数据中学习。这对预测在下采样观察与其细目标之间的演变过程几乎没有控制。我们提出了GalerkinFlow,这是一种与方程无关的框架,将每对粗糙-细致的样本转化为沿整个重建路径的监督。在重建路径上的随机中间状态样本中,模型预测粗到细的残差速度,并使用粗锚点定义伪终点。我们展示了这个伪终点的重建损失与已知的时间依赖权重之间存在精确关系。因此,每个中间状态都为同一细目标提供监督,而不仅仅是作为通向终点损失的内部步骤。由于中间状态已经揭示了部分缺失的细尺度结构,我们还对在一步推理过程中使用的粗糙终点进行额外监督。有限差分目标进一步约束了局部空间变化。GalerkinFlow结合了卷积特征与尺度条件的Galerkin算子混合,且不需要任何控制方程或物理元数据。在Navier-Stokes和Darcy流的评估方程无关基线中,它实现了最低的原始空间误差,同时在DIV2K上保持竞争力。
cs.CV / 185 / 2608.16585

SQuad: Sub-Quadratic Attention Distillation for Efficient Video Generation

SQuad:用于高效视频生成的亚二次注意力蒸馏
Karnewar, Animesh, Korzhenkov, Denis, Habibian, Amirhossein, Ghafoorian, Mohsen
Abstract
Video Diffusion Transformers (DiTs) spend most of their compute inside the Self-Attention operation, whose cost grows quadratically, $\mathcal{O}(n^2)$, with the number of latent tokens $n$. For the task of video generation, the token count is large, so this term dominates runtime and memory, and thereby caps the resolution and duration we can generate. Linear $\mathcal{O}(n)$ and low-rank $\mathcal{O}(nk)$ surrogates of Self-Attention trade the full softmax $QK^T$ for cheaper kernels, but rarely recover the original's expressivity, leaving a stubborn quality gap. Motivated by this, we propose SQuad, a Sub-Quadratic Attention Distillation framework that achieves a complexity of $\mathcal{O}(n\sqrt{n})$ in the resulting distilled Attention, naturally balancing the efficiency v/s expressivity trade-off. Instead of training our own Video DiT from scratch, which is prohibitively expensive, we fit a pretrained full softmax Self-Attention DiT into our proposed SQuad-Attention one by distilling the former in two stages: Flow-Matching Supervised Fine-Tuning (SFT), followed by improved Distribution Matching Distillation (DMD2) which additionally makes the sampling more efficient. On the Wan~2.2 5B text-to-video model, SQuAD matches the quadratic teacher on VBench ($83.20$ v/s $83.08$) while cutting the per-step per-block attention FLOPs by $\sim$$67\times$ and attention latency by $\sim$$11\times$, and end-to-end DiT latency by 2$\times$, all while also generating a video in only $6$ Neural Functional Evaluations (NFEs) instead of the default $100$.
Chinese Translation
视频扩散变换器(Video Diffusion Transformers, DiTs)在自注意力操作中消耗了大部分计算资源,其成本随着潜在标记数量 $n$ 的增加呈二次增长,$ ext{O}(n^2)$。在视频生成任务中,标记数量较大,因此这一项主导了运行时间和内存,从而限制了我们能够生成的分辨率和时长。线性 $ ext{O}(n)$ 和低秩 $ ext{O}(nk)$ 的自注意力替代方案用更便宜的核函数替代了完整的 softmax $QK^T$,但很少能恢复原始模型的表现力,留下了一个顽固的质量差距。基于此,我们提出了 SQuad,一个亚二次注意力蒸馏框架,其在蒸馏后的注意力中实现了 $ ext{O}(n ext{sqrt}(n))$ 的复杂度,自然平衡了效率与表现力之间的权衡。我们并没有从头开始训练自己的视频 DiT,因为这成本过高,而是通过在两个阶段中蒸馏一个预训练的完整 softmax 自注意力 DiT 来适配我们提出的 SQuad-Attention:流匹配监督微调(Flow-Matching Supervised Fine-Tuning, SFT),随后是改进的分布匹配蒸馏(Distribution Matching Distillation, DMD2),后者还使采样更加高效。在 Wan~2.2 5B 文本到视频模型上,SQuad 在 VBench 上与二次教师模型相匹配($83.20$ 对 $83.08$),同时将每步每块的注意力 FLOPs 减少了 $ ext{∼}67 imes$,注意力延迟减少了 $ ext{∼}11 imes$,端到端 DiT 延迟减少了 2$ imes$,而且仅用 $6$ 次神经功能评估(Neural Functional Evaluations, NFEs)生成视频,而不是默认的 $100$ 次。
cs.CV / 186 / 2608.16589

Ultra: Unsupervised Cross-Task Optimization for Reliable Restoration Segmentation Collaboration under Adverse Weather

Ultra:在恶劣天气下进行可靠恢复分割协作的无监督跨任务优化
Wang, Shiqin, Li, Zhiqian, Du, Haoyuan, Chen, Junming, Li, Jiayuan, Xu, Tianrun, Chen, Haoyang
Abstract
Unsupervised Domain Adaptation for Adverse Weather Semantic Segmentation (UDA-ASS) aims to transfer semantic knowledge from labeled normal-weather images to unlabeled adverse environments. Existing approaches implicitly assume that restoration and segmentation provide mutually beneficial guidance. However, under severe degradation and without target-domain supervision, the validity of cross-task optimization directions becomes fundamentally unidentifiable, leading to hallucination-driven error propagation. In this work, we propose a novel Unsupervised Restoration-Segmentation Collaborative Learning Framework (Ultra), which reframes cross-task interaction as direction selection under uncertainty and causal effect estimation, enabling reliable collaboration through candidate direction generation and intervention-based filtering. In detail, we propose CTDN and CMIL. The former exploits complementary visual structures and semantic information to generate candidate optimization directions and performs cooperative direction selection between restoration and segmentation. The latter reformulates cross-task information transfer from correlation-based propagation into causal effect assessment, suppressing hallucination propagation. Extensive experiments on three widely used UDA-ASS benchmarks demonstrate state-of-the-art segmentation performance. Beyond segmentation, our framework achieves better unsupervised restoration results than existing UDA-ASS restoration methods and generalizes to unsupervised restoration and object detection collaboration tasks. Code and models will be available at https://github.com/Wang-Shiqin/Ultra.
Chinese Translation
恶劣天气语义分割的无监督领域适应(UDA-ASS)旨在将标记的正常天气图像中的语义知识转移到未标记的恶劣环境中。现有方法隐含地假设恢复和分割提供互利的指导。然而,在严重退化和缺乏目标领域监督的情况下,跨任务优化方向的有效性变得根本无法识别,导致幻觉驱动的错误传播。在本研究中,我们提出了一种新颖的无监督恢复-分割协作学习框架(Ultra),将跨任务交互重新构建为不确定性下的方向选择和因果效应估计,通过候选方向生成和基于干预的过滤实现可靠的协作。具体而言,我们提出了CTDN和CMIL。前者利用互补的视觉结构和语义信息生成候选优化方向,并在恢复和分割之间进行协作方向选择。后者将跨任务信息传递从基于相关性的传播重新表述为因果效应评估,从而抑制幻觉传播。在三个广泛使用的UDA-ASS基准上的大量实验表明,我们的方法在分割性能上达到了最先进的水平。除了分割之外,我们的框架在无监督恢复结果上优于现有的UDA-ASS恢复方法,并且能够推广到无监督恢复和目标检测协作任务。代码和模型将发布在 https://github.com/Wang-Shiqin/Ultra。
cs.CV / 187 / 2608.16591

Towards Zero-Shot Domain Generalization for ID Cards Presentation Attack Detection

面向零样本领域泛化的身份证件展示攻击检测
Nieto-Hidalgo, Mario, Espin, Juan M., Tapia, Juan E.
Abstract
Presentation-Attack Detection (PAD) for national ID cards is limited by the lack of publicly available genuine samples, making it difficult for systems to generalize across countries. This paper introduces two main innovations: (1) a Prototypical Network head using an EfficientNet-V2-b0 backbone that requires only four genuine samples per class to create reliable prototypes; and (2) an episodic training regime that keeps PAD classes fixed while varying the card domain, allowing the network to learn universal attack cues. Evaluated on a large multi-country dataset and the public DLC-2021 benchmark, this method achieves an average Equal Error Rate of around 9\%, outperforming conventional softmax and CLIP zero-shot baselines even with data from a single source country. This approach provides accurate, privacy-preserving PAD while minimizing data collection, facilitating scalable cross-jurisdictional remote onboarding.
Chinese Translation
针对国家身份证件的展示攻击检测(PAD)受到公开真实样本缺乏的限制,导致系统在不同国家之间的泛化能力不足。本文提出了两项主要创新:(1)使用 EfficientNet-V2-b0 主干的原型网络头,该方法仅需每类四个真实样本即可创建可靠的原型;(2)一种情节训练机制,在该机制中,PAD 类别保持固定,而卡片领域则不断变化,从而使网络能够学习通用的攻击线索。在一个大型多国数据集和公共 DLC-2021 基准上进行评估,该方法实现了约 9\% 的平均等错误率,超越了传统的 softmax 和 CLIP 零样本基线,即使在仅使用单一来源国家的数据时也表现优异。该方法提供了准确的、保护隐私的 PAD,同时最小化数据收集,促进了可扩展的跨管辖区远程入职。
cs.CV / 188 / 2608.16600

GeoPose: Patient-agnostic CTA-to-DSA registration through projection-space calibration

GeoPose:通过投影空间校准实现患者无关的CTA到DSA配准
van Herten, Rudolf L. M., Graf, Robert, Feldman, Paula, Paetzold, Johannes C.
Abstract
Aligning intraoperative biplanar digital subtraction angiography (DSA) to pre-procedural computed tomography angiography (CTA) requires rapid and accurate 3D-to-2D registration. Optimization-based methods are sensitive to initialization and may require hundreds of iterations, whereas learning-based approaches commonly rely on patient-specific training. We propose GeoPose, a population-trained framework that estimates the C-arm pose in a learned canonical frame and transfers it to the native frame of an unseen CTA through projection-space calibration and transform composition. A population-trained residual network refines the pose, followed optionally by low-budget image-driven optimization. GeoPose requires neither patient-specific adaptation nor explicit inter-volume preregistration. On 80 DSA observations from 20 held-out patients, optimization-free GeoPose achieved a carotid mean projected centerline distance (mPCD) of 5.8 mm and a clDice of 0.45, compared with 14.5 mm and 0.28 for the best-performing baseline, while requiring only 0.15 s. After 25 optimization iterations, GeoPose reached an mPCD of 4.6 mm and a clDice of 0.58 in approximately two seconds. Under the same budget, native-initialized optimization achieved 14.6 mm and 0.15, respectively. GeoPose thus provides rapid native-frame registration with fixed population-level weights and the geometric correspondence required for downstream biplanar 3D vascular reconstruction.
Chinese Translation
将术中双平面数字减影血管造影(DSA)与预处理的计算机断层扫描血管造影(CTA)对齐需要快速而准确的3D到2D配准。基于优化的方法对初始化敏感,可能需要数百次迭代,而基于学习的方法通常依赖于特定患者的训练。我们提出了GeoPose,这是一种基于人群训练的框架,能够在学习的标准坐标系中估计C臂姿态,并通过投影空间校准和变换组合将其转移到未见CTA的本地坐标系中。一个基于人群训练的残差网络对姿态进行精细化,随后可选择性地进行低成本的图像驱动优化。在来自20名保留患者的80个DSA观察中,无需优化的GeoPose实现了5.8毫米的颈动脉平均投影中心线距离(mPCD)和0.45的clDice,而最佳基线的结果则为14.5毫米和0.28,同时仅需0.15秒。在经过25次优化迭代后,GeoPose在约两秒内达到了4.6毫米的mPCD和0.58的clDice。在相同预算下,本地初始化的优化分别达到了14.6毫米和0.15。因此,GeoPose提供了快速的本地坐标系配准,具有固定的人群水平权重和下游双平面3D血管重建所需的几何对应关系。
cs.CV / 189 / 2608.16607

Interactive Whole Slide Images for RL-based Tumour Segmentation

基于强化学习的肿瘤分割交互式全幻灯片图像
Mohamad, Mohamad, Ponzio, Francesco, Gassier, Maxime, Pote, Nicolas, Descombes, Xavier
Abstract
Whole-slide image (WSI) analysis remains computationally challenging due to the extremely large spatial resolution of slides and the sparse distribution of tumour regions. We propose an end-to-end reinforcement learning framework for sequential tumour segmentation directly on WSIs. Instead of treating the slide as a predefined collection of candidate patches, we formulate the WSI itself as a hierarchical multi-resolution environment through which an agent navigates using movement, zooming, and tumour selection actions. The agent jointly processes local observations and a global thumbnail representation within an actor-critic architecture trained using proximal policy optimization (PPO). Experiments on pulmonary adenocarcinoma WSIs demonstrate the feasibility of direct sequential tumour segmentation on full slides, achieving comparable coarse segmentation quality relative to patch-based approaches operating at similar magnification levels, while reducing inference time to a few seconds per slide. We further analyse the impact of environment design and action-space granularity. Our results suggest that modelling WSIs as interactive environments provides a promising direction for RL-based computational pathology
Chinese Translation
全幻灯片图像(WSI)分析由于幻灯片的极高空间分辨率和肿瘤区域的稀疏分布,仍然面临计算挑战。我们提出了一种端到端的强化学习框架,用于在WSI上进行顺序肿瘤分割。我们将WSI本身构建为一个分层的多分辨率环境,代理通过移动、缩放和肿瘤选择动作进行导航,而不是将幻灯片视为预定义的候选补丁集合。代理在使用近端策略优化(PPO)训练的演员-评论家架构中,联合处理局部观察和全局缩略图表示。对肺腺癌WSI的实验表明,直接在全幻灯片上进行顺序肿瘤分割是可行的,其粗略分割质量与在相似放大倍数下运行的基于补丁的方法相当,同时将推理时间缩短至每张幻灯片几秒钟。我们进一步分析了环境设计和动作空间粒度的影响。我们的结果表明,将WSI建模为交互式环境为基于强化学习的计算病理学提供了一个有前景的方向。
cs.CV / 190 / 2608.16614

Beyond Accuracy: Assessing Calibration of Geospatial Foundation Models and Their Sensitivity to Distribution Shifts

超越准确性:评估地理空间基础模型的校准及其对分布变化的敏感性
Lehmann, Nils, Gawlikowski, Jakob, Ekim, Burak, Corley, Isaac, Zhu, Xiao Xiang
Abstract
Geospatial Foundation Models (GeoFMs) are most commonly ranked and selected by accuracy on standard benchmark conditions via averaged ranks. We show that this protocol is too narrow: the promised deployment in critical EO tasks requires further angles of analysis, mainly calibration, the agreement between a model's confidence and its correctness. Across 16 frozen encoders, four classification and five segmentation datasets, and two orthogonal stress axes, every encoder degrades as corruption intensifies, and the ranking changes as well. Across the four classification benchmarks, EO-pretrained and ImageNet-pretrained encoders are indistinguishable on clean accuracy and clean calibration, and EO pretraining provides no more stability under shift than ImageNet pretraining. Under shift the GeoFMs drift further into overconfidence than the ImageNet-pretrained encoders, at every grade and in every corruption family. A centered kernel alignment (CKA) analysis ties this to representational rigidity: EO-pretrained embeddings move less under corruption while losing just as much task information and remaining overconfident. We apply three commonly explored uncertainty quantification methods and find that temperature scaling and deep ensembles cannot counteract the degradation, while a Gaussian-process probe roughly halves ECE under severe cloud only by tripling it on clean data. In selective prediction experiments, we find that confidence-based abstention cannot defer around confidently wrong predictions, and advocate that benchmark rankings and evaluations should therefore operate across a multitude of conditions and metrics to more holistically evaluate model development progress and close the gap to real world deployment scenarios.
Chinese Translation
地理空间基础模型(GeoFMs)通常通过在标准基准条件下的平均排名来评估和选择其准确性。我们表明,这一协议过于狭隘:在关键的地球观测(EO)任务中承诺的部署需要进一步的分析角度,主要是校准,即模型的置信度与其正确性之间的一致性。在16个冻结编码器、四个分类和五个分割数据集以及两个正交压力轴的实验中,随着损坏程度的加剧,每个编码器的性能都在下降,排名也随之变化。在四个分类基准中,经过EO预训练和ImageNet预训练的编码器在干净的准确性和干净的校准上无法区分,且EO预训练在变化下的稳定性并不优于ImageNet预训练。在变化下,GeoFMs的过度自信程度比ImageNet预训练的编码器更为严重,且在每个等级和每个损坏类别中均是如此。中心核对齐(CKA)分析将这一现象与表征刚性联系起来:EO预训练的嵌入在损坏下的移动幅度较小,但任务信息的损失程度相同,且仍然过于自信。我们应用了三种常见的不确定性量化方法,发现温度缩放和深度集成无法抵消性能下降,而高斯过程探测器在严重云层下大致将ECE减半,但在干净数据上却使其增加了三倍。在选择性预测实验中,我们发现基于置信度的放弃无法规避自信错误的预测,因此建议基准排名和评估应在多种条件和指标下进行,以更全面地评估模型开发进展,并缩小与实际部署场景之间的差距。
cs.CV / 191 / 2608.16622

HarmTrace: Anchor-Calibrated Decoupled Optimization for Fine-Grained Target Identification in Harmful Memes

HarmTrace:用于有害表情包中细粒度目标识别的锚点校准解耦优化
Li, Yujia, Zhang, Yiqun, Cheng, Zihan, Huang, Yijie, Ye, Tenglong, Wang, Zihan, Yang, Xiaocui, Feng, Shi, Zhang, Yifei, Wang, Daling
Abstract
Multimodal harmful meme detection is typically formulated as image--text harmfulness classification. A model may correctly predict harmfulness while misidentifying the attacked target or its supporting evidence. We therefore extend harmful meme detection with fine-grained target identification, asking what type of target is attacked, who is targeted, and where the target appears in the meme. The model predicts harmfulness for every meme and, for harmful memes, outputs the target category, target entity, textual mention, and visual region. To support this task, we introduce Meme3W, which unifies multiple public harmful meme datasets and provides human-verified annotations for harmful instances. We further introduce Joint Record Accuracy (JRA), a strict record-level metric requiring the harmfulness label and all target-identification fields to be jointly correct. Experiments with representative multimodal large language models reveal a substantial gap between harmfulness accuracy and JRA. To narrow this gap, we propose HarmTrace, an anchor-calibrated decoupled optimization framework. HarmTrace strengthens target-entity supervision through entity-aware supervised fine-tuning. It then applies Conditional Target-identification Policy Optimization (CTPO) to decouple harmfulness and target-identification advantages, restricting target-identification optimization to label-correct responses for harmful examples. CTPO uses a Virtual Positive Anchor (VPA) as a fully correct reference for target-identification advantage normalization. HarmTrace improves both JRA and harmfulness accuracy across the evaluated backbones, with JRA on the Qwen3-VL-8B backbone increasing from 17.58\% to 52.51\%. Our code is publicly available at https://github.com/llly1234/HarmTrace-for-Harmful-Memes.
Chinese Translation
多模态有害表情包检测通常被表述为图像-文本有害性分类。一个模型可能正确预测有害性,但错误识别被攻击的目标或其支持证据。因此,我们扩展了有害表情包检测,加入细粒度目标识别,探讨被攻击的目标类型、攻击对象及目标在表情包中的出现位置。该模型为每个表情包预测有害性,并且对于有害表情包,输出目标类别、目标实体、文本提及和视觉区域。为支持这一任务,我们引入了Meme3W,它统一了多个公共有害表情包数据集,并提供了经过人工验证的有害实例注释。我们进一步引入了联合记录准确率(Joint Record Accuracy, JRA),这是一种严格的记录级指标,要求有害性标签和所有目标识别字段共同正确。与代表性的多模态大型语言模型的实验显示,有害性准确率与JRA之间存在显著差距。为了缩小这一差距,我们提出了HarmTrace,一种锚点校准的解耦优化框架。HarmTrace通过实体感知的监督微调增强目标实体的监督。随后,它应用条件目标识别策略优化(Conditional Target-identification Policy Optimization, CTPO)来解耦有害性和目标识别优势,将目标识别优化限制在有害示例的标签正确响应上。CTPO使用虚拟正锚点(Virtual Positive Anchor, VPA)作为目标识别优势归一化的完全正确参考。HarmTrace在评估的基础模型上提高了JRA和有害性准确率,其中在Qwen3-VL-8B基础模型上的JRA从17.58\%提高到52.51\%。我们的代码已公开可用,网址为https://github.com/llly1234/HarmTrace-for-Harmful-Memes。
cs.CV / 192 / 2608.16632

DRAFE: Domain-Robust Asymmetric Fusion of Heterogeneous Detection Transformers for Cross-City Fine-Grained Traffic Object Detection

DRAFE:针对跨城市细粒度交通目标检测的领域鲁棒性非对称融合异构检测变换器
Agbobli, Divine Yao, Agorku, Geoffery Eyram, Afriyie, Israel, Amankwah-Nkyi, Kwadwo, Osei-Kuffour, Marvin, Duah, Richmond Owusu, Seglah, Bright, Terkper, Kelvin Asamoah, Adjei, Kwabena Amoako
Abstract
Deep learning-based object detectors are fundamental to intelligent transportation systems, enabling traffic monitoring, vehicle analytics, and infrastructure management. However, achieving both fine-grained vehicle recognition and robust cross-city domain generalization remains challenging. We present the Domain-Robust Asymmetric Fusion Ensemble (DRAFE), which combines independently trained LW-DETR and RF-DETR detectors for cross-city fine-grained traffic object detection. DRAFE employs a two-stage training strategy that first pretrains complementary detectors on diverse public traffic datasets using pseudo-label expansion and human-in-the-loop annotation refinement, producing a curated corpus of 6,049 images and 203,619 annotations, before challenge-compliant fine-tuning on the Project Hafnia Track 6 dataset. At inference, DRAFE applies anchor-conditioned class-consistent matching, reliability-weighted coordinate fusion, agreement-aware confidence recalibration, and complementary hypothesis recovery. On AI City Challenge 2026 Track 6, DRAFE achieves 0.4022 mAP, ranks sixth among 25 participating teams, and improves by 0.0553 mAP over a preliminary ensemble evaluated under identical benchmark conditions.
Chinese Translation
基于深度学习的目标检测器是智能交通系统的基础,能够实现交通监控、车辆分析和基础设施管理。然而,实现细粒度的车辆识别和鲁棒的跨城市领域泛化仍然具有挑战性。我们提出了领域鲁棒性非对称融合集成(DRAFE),它结合了独立训练的LW-DETR和RF-DETR检测器,用于跨城市细粒度交通目标检测。DRAFE采用了两阶段的训练策略,首先在多样的公共交通数据集上使用伪标签扩展和人机协作的注释精炼对互补检测器进行预训练,生成了一个包含6,049张图像和203,619个注释的精心策划的语料库,然后在Project Hafnia Track 6数据集上进行符合挑战要求的微调。在推理阶段,DRAFE应用锚点条件的类别一致匹配、可靠性加权坐标融合、考虑一致性的置信度重新校准和互补假设恢复。在AI City Challenge 2026 Track 6中,DRAFE实现了0.4022 mAP,在25个参赛团队中排名第六,并在相同基准条件下比初步集成提高了0.0553 mAP。
cs.CV / 193 / 2608.16646

Training-Free Reconstruction-Based AI-Generated Image Detectors Are Inherently Vulnerable to Adversarial Examples

无训练重建基础的人工智能生成图像检测器本质上对对抗样本脆弱
Demchenko, Roman, Ricker, Jonas, Fischer, Asja
Abstract
The impressive visual quality and ubiquity of AI-generated images call for reliable and robust detection methods. Reconstruction-based detectors have emerged as a promising direction for transparent and training-free identification of synthetic images. However, due to their fundamentally different mode of operation (compared to standard, classifier-based methods), little is known about their adversarial robustness. In this work, we propose two novel attack methods targeted at detectors that leverage autoencoder reconstruction error. We find that by constructing imperceptible adversarial examples, the distance between original and reconstruction can be artificially increased, causing fake images to be wrongly classified as real. Our evaluation including images from three state-of-the-art generators and three detectors demonstrates that detection performance is significantly decreased, even if attacked images additionally undergo real-world degradations. Critically, our adversarial examples naturally transfer across detectors, as they all share the same principle, pointing towards an inherent vulnerability of reconstruction-based detectors.
Chinese Translation
人工智能生成图像的卓越视觉质量和普遍存在性呼唤可靠且稳健的检测方法。基于重建的检测器已成为透明且无训练的合成图像识别的有前景方向。然而,由于其与标准分类器方法在操作模式上的根本不同,其对抗鲁棒性知之甚少。在本研究中,我们提出了两种新颖的攻击方法,针对利用自编码器重建误差的检测器。我们发现,通过构造不可察觉的对抗样本,可以人为地增加原始图像与重建图像之间的距离,导致伪造图像被错误地分类为真实图像。我们的评估包括来自三种最先进生成器和三种检测器的图像,结果表明,即使被攻击的图像还经历了现实世界的降质,检测性能也显著下降。关键是,我们的对抗样本在不同检测器之间自然迁移,因为它们都共享相同的原理,这指向了基于重建的检测器的固有脆弱性。
cs.CV / 194 / 2608.16658

X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization

X$^2$Localizer:渐进式跨视角视频地理定位的跨粒度对齐
Zeng, Zichao, Fan, Weijia, Chen, Yufan, Goo, June Moh, Zheng, Junwei, Liu, Ruiping, Peng, Kunyu, Zhang, Jiaming, Stiefelhagen, Rainer, Boehm, Jan
Abstract
Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images. However, CVG approaches rely on fixed-length inputs and post-hoc refinement, hindering online-oriented localization under partial or dynamic observations. In this work, we formulate Progressive Cross-view Video Geo-localization (PCVG) as a deployment-oriented extension and evaluation protocol of CVG, enabling localization under varying temporal budgets, prefix-based inference, random-start evaluation, and long-range localization with interruptions. To explore PCVG, we introduce X$^2$Localizer, a cross-grained alignment framework that jointly supervises global prefix-to-aerial retrieval and token-aggregated frame--aerial-tile matching with a budget-dependent asymmetric objective. Furthermore, we introduce a Sliding-Window Re-Localization (SWRL) strategy that dynamically refreshes candidate regions for failure recovery and long-range deployment without full-sequence reprocessing. Extensive experiments show that X$^2$Localizer preserves conventional full-video performance, with marginal gains of +0.1 Recall@1 and +0.3 Recall@10, while substantially improving early localization. In the challenging single-frame setting, X$^2$Localizer improves coarse retrieval by +4.7 Recall@1 and +11.5 Recall@10 over the previous state-of-the-art method. With SWRL, our approach further enables robust progressive localization under random-start and long-distance scenarios, narrowing the gap between benchmark evaluation and real-world deployment.
Chinese Translation
跨视角视频地理定位(CVG)旨在通过检索相应的地理标记航空图像来定位地面视角视频。然而,CVG 方法依赖于固定长度的输入和事后精炼,这阻碍了在部分或动态观测下的在线定位。在本研究中,我们将渐进式跨视角视频地理定位(PCVG)构建为 CVG 的一种面向部署的扩展和评估协议,使得在不同时间预算、基于前缀的推理、随机起始评估以及中断情况下的长距离定位成为可能。为了探索 PCVG,我们引入了 X$^2$Localizer,一个跨粒度对齐框架,该框架共同监督全局前缀到航空图像的检索和基于令牌聚合的帧-航空图块匹配,采用预算依赖的非对称目标。此外,我们引入了一种滑动窗口重新定位(SWRL)策略,该策略动态刷新候选区域以实现故障恢复和长距离部署,而无需全序列重新处理。大量实验表明,X$^2$Localizer 保持了传统全视频性能,Recall@1 提升了 +0.1,Recall@10 提升了 +0.3,同时显著改善了早期定位。在具有挑战性的单帧设置中,X$^2$Localizer 相较于之前的最先进方法,粗检索提升了 +4.7 Recall@1 和 +11.5 Recall@10。通过 SWRL,我们的方法进一步在随机起始和长距离场景下实现了稳健的渐进式定位,缩小了基准评估与实际部署之间的差距。
cs.CV / 195 / 2608.16661

Turning spectra into images improves plant trait retrieval with 2D-CNNs

将光谱转化为图像提高了基于2D卷积神经网络的植物性状提取
Lopatin, Javier, Kattenborn, Teja, Cherif, Eya, Moreno, Sebastián
Abstract
Hyperspectral reflectance spectroscopy enables non-destructive estimation of plant functional traits, yet current deep learning approaches process spectra as one-dimensional sequences, which limits how they capture long-range inter-band dependencies. We asked whether transforming 1D spectra into 2D image representations improves multi-trait prediction with convolutional neural networks (CNN). We compared nine transformations using EfficientNet-B0 on the GreenHyperSpectra dataset (7,897 labeled spectra, eight traits, 400-2450 nm), benchmarked against published 1D CNN results on the same split. Trained from scratch, the simplest transformation, a direct Reshape of the spectrum into a 2D grid, performed best ($R^2 = 0.684 \pm 0.001$) and improved on the state-of-the-art 1D baseline ($R^2 = 0.587$, $+0.097$). We then pretrained a 2D masked autoencoder (MAE-2D) on 139,000 unlabeled spectral images. Linear probing, which freezes the encoder and trains only a multilayer perceptron head, reached $R^2 = 0.646$ and exceeded every 1D self-supervised counterpart, including the fine-tuned MAE-1D ($R^2 = 0.641$). Under cross-dataset evaluation all models lost most of their accuracy and none beat the 1D baseline significantly. To identify which wavelengths drive each prediction, we applied Integrated Gradients and Grad-CAM and unfolded band importance back to the spectral axis. Protein ($r = 0.45$) and leaf water ($r = 0.33$) agreed with sensitivities simulated by the PROSAIL radiative-transfer model, while carotenoids ($r = 0.06$) and leaf area index ($r = -0.11$) did not, showing that the model reads established leaf chemistry for traits with sharp absorption features. The representational advantage of 2D spectral images, rather than architectural complexity or ImageNet pretraining, drives the gain over 1D approaches.
Chinese Translation
高光谱反射光谱技术能够非破坏性地估计植物功能性状,但当前的深度学习方法将光谱处理为一维序列,这限制了它们捕捉长程波段间依赖关系的能力。我们探讨了将一维光谱转化为二维图像表示是否能改善卷积神经网络(CNN)的多性状预测。我们在GreenHyperSpectra数据集(7,897个标记光谱,八个性状,400-2450 nm)上使用EfficientNet-B0比较了九种转化方法,并与同一数据集划分下已发布的一维CNN结果进行了基准比较。从头开始训练,最简单的转化,即将光谱直接重塑为二维网格,表现最佳($R^2 = 0.684 ext{ } ext{±} ext{ } 0.001$),并且优于最先进的一维基线($R^2 = 0.587$,$+0.097$)。随后,我们在139,000个未标记的光谱图像上预训练了一个二维掩蔽自编码器(MAE-2D)。线性探测冻结编码器,仅训练多层感知器头,达到了$R^2 = 0.646$,超过了每个一维自监督对应物,包括微调的MAE-1D($R^2 = 0.641$)。在跨数据集评估中,所有模型的准确性大幅下降,且没有模型显著超过一维基线。为了识别驱动每个预测的波长,我们应用了积分梯度和Grad-CAM,并将波段重要性回溯到光谱轴。蛋白质($r = 0.45$)和叶片水分($r = 0.33$)与PROSAIL辐射传输模型模拟的敏感性一致,而类胡萝卜素($r = 0.06$)和叶面积指数($r = -0.11$)则不一致,显示模型读取了具有明显吸收特征的已建立叶片化学特性。二维光谱图像的表征优势,而非架构复杂性或ImageNet预训练,推动了相较于一维方法的提升。
cs.CV / 196 / 2608.16669

Concept-based explanation of gene expression prediction from H&E images

基于概念的H&E图像基因表达预测解释
Muench, Amos, Thielmann, Jonathan, Achtibat, Reduan, Dreyer, Maximilian, Bischoff, Philip, Forsythe, Caroline, Parand, Hamidreza, Walter, Thomas, Horst, David, Lapuschkin, Sebastian, Samek, Wojciech, Krieger, Teresa Gabriela
Abstract
Recent advances in pathology foundation models have enabled accurate prediction of spatial transcriptomics (ST) from routine H&E images. However, existing explainability methods for vision transformer (ViT)-based models are largely limited to local heatmaps and do not reveal how morphological concepts contribute to ST predictions. Here, we introduce an explainable framework that combines relevance propagation and concept discovery to link transcriptional programs to tissue morphology. We developed a ViT-based framework for virtual ST from H&E images that combines ViT-aware layer-wise relevance propagation with relaxed archetypal TopK sparse autoencoder-based concept discovery. This approach provides both local explanations and global insights into the morphological patterns associated with transcriptional programs. We applied the framework to colorectal cancer ST data from the HEST-1k cohort and evaluated its generalizability in TCGA COAD. Our architecture accurately predicts clinically relevant ST signatures and accompanying molecular phenotypes. Measured and predicted gene expression profiles reveal substantial spatial heterogeneity of the colorectal cancer subtypes iCMS2 and iCMS3 across a large number of samples. Spatially resolved and aggregated iCMS classification achieve weighted F1 scores of 0.872 and 0.819 (0.770 in TCGA COAD), respectively, and both stratify patient outcome. Beyond prediction, our framework establishes a relevance-based concept atlas linking molecular phenotypes to histopathological representations. Comparison of activation- with relevance-derived concepts demonstrates that relevances provide a more direct link between tissue morphology and downstream predictions. We establish a general strategy for concept-based explanation of spatial prediction, and our framework is readily applicable to a broad range of ViT-based pathology models.
Chinese Translation
近期在病理基础模型方面的进展使得能够从常规H&E图像中准确预测空间转录组学(ST)。然而,现有的针对基于视觉变换器(ViT)模型的可解释性方法主要局限于局部热图,并未揭示形态学概念如何影响ST预测。在此,我们引入了一种可解释框架,结合相关性传播和概念发现,将转录程序与组织形态学联系起来。我们开发了一个基于ViT的框架,用于从H&E图像生成虚拟ST,该框架结合了ViT感知的逐层相关性传播和放松的典型TopK稀疏自编码器基础的概念发现。这种方法提供了局部解释和与转录程序相关的形态模式的全局见解。我们将该框架应用于来自HEST-1k队列的结直肠癌ST数据,并评估其在TCGA COAD中的普适性。我们的架构准确预测临床相关的ST特征及其伴随的分子表型。测量和预测的基因表达谱揭示了结直肠癌亚型iCMS2和iCMS3在大量样本中的显著空间异质性。空间分辨和聚合的iCMS分类分别达到加权F1分数0.872和0.819(在TCGA COAD中为0.770),并且两者均能对患者结果进行分层。除了预测外,我们的框架建立了一个基于相关性的概念图谱,将分子表型与组织病理学表现联系起来。激活与相关性派生概念的比较表明,相关性提供了组织形态学与下游预测之间更直接的联系。我们建立了一种基于概念的空间预测解释的一般策略,并且我们的框架可以广泛应用于多种基于ViT的病理模型。
cs.CV / 197 / 2608.16673

How Sampling Strategy Affects Imbalance Mitigation in LiDAR Segmentation: A Study of Structured vs. Random Point-Based Architectures

采样策略如何影响LiDAR分割中的不平衡缓解:结构化与随机点基础架构的研究
Savva, Antonis, Kyrkou, Christos, Theocharides, Theocharis
Abstract
Class imbalance in LiDAR point clouds poses challenges for semantic segmentation in autonomous navigation and urban mapping. While 2D vision has numerous mitigation techniques, their effectiveness in 3D remains unclear. We benchmark six reweighting schemes and five imbalance-aware losses across three datasets (DALES, S3DIS, STPLS3D) using two architectures (KPConv, RandLA-Net). Inverse-frequency weighting degrades performance by up to 12% compared to uniform weighting, with catastrophic failures in minority classes. Uniform weighting performs within 2% of complex losses for structured sampling (KPConv) but benefits less for random sampling (RandLA-Net, up to 4.6% gap). Loss landscape analysis reveals a complex interplay: for structured sampling, imbalance ratio determines landscape geometry on real LiDAR data but decouples from it on synthetic data; for random sampling, landscapes show high sensitivity to dataset geometry regardless of imbalance ratio. For the two evaluated point-based architectures, these results suggest that the interaction between sampling strategy (structured vs. random), imbalance severity, and data acquisition characteristics shapes which mitigation approaches are effective.
Chinese Translation
LiDAR点云中的类别不平衡对自主导航和城市映射中的语义分割构成了挑战。尽管2D视觉有许多缓解技术,但它们在3D中的有效性仍不明确。我们基于两种架构(KPConv, RandLA-Net)在三个数据集(DALES, S3DIS, STPLS3D)上对六种重加权方案和五种考虑不平衡的损失进行了基准测试。与均匀加权相比,逆频率加权的性能下降最多可达12%,在少数类中出现灾难性失败。对于结构化采样(KPConv),均匀加权的表现与复杂损失相差不超过2%,但对于随机采样(RandLA-Net,最多相差4.6%)则受益较少。损失景观分析揭示了复杂的相互作用:对于结构化采样,不平衡比率决定了真实LiDAR数据的景观几何形状,但在合成数据中则与之解耦;对于随机采样,景观对数据集几何形状表现出高度敏感性,而不论不平衡比率如何。对于这两种评估的点基础架构,这些结果表明,采样策略(结构化与随机)、不平衡严重性和数据获取特征之间的相互作用决定了哪些缓解方法是有效的。
cs.CV / 198 / 2608.16681

Bridging the Gap between Labeled and Unlabeled Data via Unified Flow with Feature Memory Bank

通过特征记忆库的统一流弥合标记数据与未标记数据之间的差距
Wang, Shanwen, Sun, Xin, Hong, Danfeng, Dong, Junyu, Callet, Patrick Le
Abstract
Although semi-supervised semantic segmentation ($\text{S}^4$) utilizes abundant unlabeled data to reduce manual labeling burdens, independent training of labeled and unlabeled data causes the former to dominate, which severely degrades pseudo-label quality. To address this challenges, we propose a novel remote sensing (RS) $\text{S}^4$ method via unified flow with feature memory bank (UFFM). Specifically, UFFM comprises two key innovations: unified flow (UF) and feature memory bank (FMB). The UF is a new training flow that generates less biased pseudo-labels by combining an external visual foundation model (VFM) with an RS domain teacher, and jointly optimizes labeled and pseudo-labeled data under a unified training objective. The FMB is a novel memory module for $\text{S}^4$ that dynamically updates class-specific features during training and reduces the feature discrepancy between labeled and unlabeled data through class-feature alignment. To verify the effectiveness of our model, we conduct extensive experiments on RS datasets. The experimental results show the superiority of our method over SOTA $\text{S}^4$ methods. Moreover, the results demonstrate the effectiveness of our contributions in bridging the optimization and feature representation gap between labeled and unlabeled data. Our code is released at \href{https://github.com/wangshanwen001/RS-UFFM}{https://github.com/wangshanwen001/RS-UFFM}.
Chinese Translation
尽管半监督语义分割($ ext{S}^4$)利用丰富的未标记数据来减少人工标记的负担,但标记数据和未标记数据的独立训练导致前者占主导地位,从而严重降低了伪标签的质量。为了解决这一挑战,我们提出了一种新颖的遥感(RS)$ ext{S}^4$方法,通过统一流与特征记忆库(UFFM)。具体而言,UFFM包含两个关键创新:统一流(UF)和特征记忆库(FMB)。UF是一种新的训练流程,通过将外部视觉基础模型(VFM)与RS领域教师结合,生成偏差较小的伪标签,并在统一的训练目标下共同优化标记数据和伪标记数据。FMB是一个新颖的$ ext{S}^4$记忆模块,在训练过程中动态更新类特定特征,并通过类特征对齐减少标记数据与未标记数据之间的特征差异。为了验证我们模型的有效性,我们在RS数据集上进行了广泛的实验。实验结果表明我们的方法优于当前最先进的$ ext{S}^4$方法。此外,结果还证明了我们在弥合标记数据与未标记数据之间的优化和特征表示差距方面的贡献的有效性。我们的代码已发布在 exttt{https://github.com/wangshanwen001/RS-UFFM}。
cs.CV / 199 / 2608.16690

AnchorScore: A CLIP-Based Diagnostic of MLLM Annotation Difficulty

AnchorScore:基于CLIP的多模态大语言模型注释难度诊断
Ma, Yan, Zhang, Lizhuo
Abstract
Multimodal large language models (MLLMs) are widely used for automated annotation, yet their per-class accuracy varies widely (e.g., 12%-98% across the 13 classes of three classroom sub-datasets) and is expensive to measure: evaluating one 27B MLLM on 5,416 validation images takes roughly 14 hours, whereas a frozen-CLIP pass over the same images completes in about 3 minutes. A low-cost signal for ranking classes by expected MLLM annotation difficulty a priori remains underexplored. Building on the AnchorProxy construct (per-class zero-shot CLIP accuracy) introduced in the companion study, this paper systematically evaluates its full-frame formulation, termed AnchorScore here, as an a priori diagnostic that flags the classes MLLMs are least likely to annotate reliably. On classroom behavior data (SCB5, 13 classes, 6 MLLMs), AnchorScore correlates with per-class MLLM accuracy (Spearman rho = 0.769, p = 0.002, n = 13). None of the alternative difficulty predictors (DINOv2, ResNet-50, SigLIP, or MLLM self-verbalized uncertainty) showed a significant class-level correlation at n = 13. A cross-model consensus control suggests AnchorScore primarily captures a shared class-difficulty factor rather than a CLIP-specific signal. An independent replication on Stanford40 Actions yields a nearly identical effect (rho = 0.817, p < 0.001); the association is strongest on activity-recognition data and attenuates on medical and satellite imagery. Three practical applications follow: a deployable hybrid CLIP/MLLM routing strategy (predicted-class routing: up to +23 pp over CLIP-only at roughly 44% MLLM cost savings), prompt disambiguation on hard classes (exploratory), and review-priority prediction for human verification. AnchorScore does not estimate exact MLLM accuracy; it provides a low-cost ranking signal that directs expensive MLLM evaluation to the classes where it is most informative.
Chinese Translation
多模态大语言模型(MLLMs)广泛应用于自动注释,但其每类的准确率差异很大(例如,在三个课堂子数据集中13个类别的准确率为12%-98%),且测量成本高昂:对一个27B的MLLM在5,416张验证图像上的评估大约需要14小时,而对同一图像进行冻结的CLIP处理则仅需约3分钟。用于按预期MLLM注释难度对类别进行排名的低成本信号仍未得到充分探索。基于在相关研究中提出的AnchorProxy构造(每类零-shot CLIP准确率),本文系统评估其全框架表述,称之为AnchorScore,作为一种事先诊断工具,标识出MLLM最不可能可靠注释的类别。在课堂行为数据(SCB5,13个类别,6个MLLM)上,AnchorScore与每类MLLM准确率相关(Spearman rho = 0.769,p = 0.002,n = 13)。其他难度预测指标(DINOv2、ResNet-50、SigLIP或MLLM自我表述的不确定性)在n = 13的情况下未显示出显著的类别级相关性。跨模型共识控制表明,AnchorScore主要捕捉共享的类别难度因素,而不是特定于CLIP的信号。在Stanford40 Actions上的独立复制实验产生了几乎相同的效果(rho = 0.817,p < 0.001);这种关联在活动识别数据上最强,而在医学和卫星图像上减弱。随后提出三种实际应用:可部署的混合CLIP/MLLM路由策略(预测类别路由:在大约44%的MLLM成本节省下,较CLIP仅路由提高最多23个百分点),对困难类别的提示消歧(探索性),以及人类验证的审核优先级预测。AnchorScore并不估计确切的MLLM准确率;它提供了一种低成本的排名信号,将昂贵的MLLM评估引导至最具信息性的类别。
cs.CV / 200 / 2608.16709

MIRROR: Multimodal Intelligent Radiology Reasoning and Observation Reporter

MIRROR:多模态智能放射学推理与观察报告系统
Nagarajan, Vignesh, Venkatapathy, Sriram
Abstract
A radiologist reading a model's output faces two problems. The model returns a number and no reason, and any system that turns that number into readable prose can quietly add claims the model never made. MIRROR is a research prototype built to separate those failures. It chains a multi-label classifier, a Grad-CAM localizer that turns each positive finding into a named anatomical region, and a report writer that receives the labels, probabilities, and regions but never the image. Because the language layer cannot see pixels, it cannot assert a finding the classifier did not make. We are precise about what that buys: a MIRROR report's findings are auditable against the probability vector, while the sentences framing them are ordinary generated text, and we show one stating a cardiothoracic ratio the system never measured. One registry holds the taxonomy, anatomy, and phrasing for chest X-ray, brain MRI, and head CT, so adding a modality is a data change; all three are routed and tested, one is trained. On ChestMNIST that classifier reaches macro AUROC 0.729 and ranks better than chance on all 14 labels, at 1.6 to 6.8 times the precision a random ranker would get. Yet at the default 0.5 threshold it emits no positive prediction at all for 11 of them, and its excellent-looking Brier score of 0.045 sits beside the 0.047 earned by a predictor that ignores the image. The discrimination is real; the decisions are not. Under the class imbalance normal in radiology, aggregate metrics flatter models that do nothing, and should be reported against that floor.
Chinese Translation
放射科医生在解读模型输出时面临两个问题。模型返回一个数字而没有理由,任何将该数字转化为可读文本的系统都可能悄然添加模型从未做出的声明。MIRROR是一个研究原型,旨在分离这些失败。它串联了一个多标签分类器、一个Grad-CAM定位器,将每个阳性发现转化为命名的解剖区域,以及一个报告撰写器,该撰写器接收标签、概率和区域,但从不接收图像。由于语言层无法看到像素,它无法断言分类器未做出的发现。我们明确了这带来的好处:MIRROR报告的发现可以根据概率向量进行审计,而框定这些发现的句子则是普通生成的文本,我们展示了一个陈述心胸比率的句子,该系统从未测量过。一个注册表保存了胸部X光、脑部MRI和头部CT的分类法、解剖学和措辞,因此添加一种模态只是数据的变化;这三者都被路由和测试,其中一个经过训练。在ChestMNIST数据集上,该分类器达到宏观AUROC 0.729,并在所有14个标签上表现优于随机猜测,精度是随机排序者的1.6到6.8倍。然而,在默认的0.5阈值下,它对其中11个标签完全没有正预测,其优秀的Brier评分0.045与一个忽略图像的预测器的0.047相邻。区分是真实的;决策却并非如此。在放射学中,类别不平衡的正常状态下,聚合指标会抬高那些无所作为的模型,应该在这一底线之上进行报告。
cs.CV / 201 / 2608.16717

PersonaShot: Benchmarking Person-Centric Narrative Continuity in Multi-Shot Video Generation

PersonaShot:多镜头视频生成中的以人为中心的叙事连续性基准
Wang, Yuji, Chen, Yuheng, Hu, Teng, Yi, Ran, Hong, Yijia, Feng, Han, Cao, Weijian, Wang, Chengjie, Ma, Lizhuang, Zhang, Jiangning
Abstract
Video generation is rapidly evolving from single-shot clips to multi-shot narratives, where the human character serves as the core narrative anchor. However, existing benchmarks mainly assess character appearance or individual-shot quality, without measuring whether physical and emotional states remain coherent across cuts. They also rarely provide criterion-specific evaluation methods, although physical continuity, facial dynamics, and cinematic relations require different visual, temporal, and relational evidence. To address these limitations, we introduce PersonaShot, the first person-centric benchmark for narrative continuity in multi-shot video generation. PersonaShot contains approximately 1,000 multi-shot segments and 16 metrics spanning physical continuity, affective dynamics, and cinematic grammar. \textbf{\textit{1)} Narrative Continuity Benchmark:} We evaluate character coherence across three temporal levels: within-shot states, cross-shot transitions, and sequence-level trajectories. \textbf{\textit{2)} Human-Aligned Specialist Evaluators:} We distill reasoning from a large multimodal teacher into lightweight criterion-specific evaluators, each grounded in the visual, temporal, or relational evidence required by its metric, and align them with expert human judgments. \textbf{\textit{3)} Systematic Evaluation and Insights:} Our evaluation reveals distinct capability profiles across state-of-the-art models and a clear gap between perceptual quality and cross-shot narrative continuity. Even visually compelling videos frequently exhibit physical-state resets, abrupt affective shifts, and broken cinematic relations across shots. Human studies further demonstrate strong agreement between our evaluators and expert judgments.
Chinese Translation
视频生成正迅速从单镜头剪辑发展到多镜头叙事,其中人类角色作为核心叙事锚点。然而,现有基准主要评估角色外观或单个镜头质量,而未测量物理和情感状态在镜头切换中的一致性。此外,尽管物理连续性、面部动态和电影关系需要不同的视觉、时间和关系证据,但很少提供特定标准的评估方法。为了解决这些局限性,我们引入了PersonaShot,这是第一个针对多镜头视频生成中叙事连续性的以人为中心的基准。PersonaShot包含大约1000个多镜头片段和16个指标,涵盖物理连续性、情感动态和电影语法。 extbf{ extit{1)} 叙事连续性基准:} 我们在三个时间层面上评估角色的一致性:镜头内状态、镜头间过渡和序列级轨迹。 extbf{ extit{2)} 人类对齐的专业评估者:} 我们从大型多模态教师中提炼推理,形成轻量级的特定标准评估者,每个评估者都基于其指标所需的视觉、时间或关系证据,并与专家的人类判断对齐。 extbf{ extit{3)} 系统评估与洞察:} 我们的评估揭示了最先进模型之间的不同能力特征,以及感知质量与镜头间叙事连续性之间的明显差距。即使是视觉上引人注目的视频,通常也会在镜头间表现出物理状态重置、情感突变和破裂的电影关系。人类研究进一步表明,我们的评估者与专家判断之间存在高度一致性。
cs.CV / 202 / 2608.16718

CytoFormer: A Molecularly Supervised Cell Foundation Model for Histopathology Cell Classification

CytoFormer:一种用于组织病理学细胞分类的分子监督细胞基础模型
Yao, Jialu, Li, Songhao, Yu, Alina, Huang, Zhi
Abstract
Identifying cell types directly from routine haematoxylin and eosin (H&E) histology would enable single-cell analysis at scale, but training such models has relied on manual pathologist annotations, which are slow, expensive and unreliable for many cell types. We instead supervise morphology with molecules. Imaging-based spatial transcriptomics profiles individual cells in situ on a section that can afterwards be stained with H&E, so that molecular identity and morphology are observed for the same physical cell. We assembled 81 such paired Xenium sections spanning 16 organs, derived per-cell labels by clustering, marker-gene annotation, organ-wise human review and quality control, and mapped them onto the cell types commonly reported in each organ. This yielded 15.4 million cells, each with a paired H&E image patch and one of 23 cell types, on which we trained CytoFormer, a cell foundation model with a multi-task, per-organ classification head. On spatially held-out tissue CytoFormer reached an accuracy of 0.85 and a macro-F1 of 0.78 across all 16 organs, and its predictions reproduced the tissue architecture of an entire held-out section. The representation also transfers: with the encoder frozen, a linear head on CytoFormer features performed better than six pathology foundation models on four expert-annotated benchmarks, including on organs and cell types that were not part of pretraining. Finally, in an interactive active-learning setting, CytoFormer's embeddings are markedly more label-efficient than existing pathology foundation models, detecting normal epithelium amid look-alike tumour with an F1 of 0.82 from only a few annotations and leading the strongest baseline by 0.13 in F1. CytoFormer turns paired H&E and spatial transcriptomics into a reusable, label-efficient representation for cell-level analysis of routine histology.
Chinese Translation
直接从常规的苏木精-伊红(H&E)组织学中识别细胞类型将使单细胞分析能够大规模进行,但训练此类模型依赖于人工病理学家的标注,这一过程既缓慢又昂贵,并且对许多细胞类型来说不够可靠。我们选择用分子来监督形态学。基于成像的空间转录组学在切片上原位描绘单个细胞,随后可以用H&E进行染色,从而使分子身份和形态学能够在同一物理细胞上被观察到。我们组装了81个这样的配对Xenium切片,涵盖16个器官,通过聚类、标记基因注释、器官级人类审核和质量控制获得每个细胞的标签,并将其映射到每个器官中常见的细胞类型上。这产生了1540万个细胞,每个细胞都有一个配对的H&E图像块和23种细胞类型之一,我们在此基础上训练了CytoFormer,这是一种具有多任务、按器官分类头的细胞基础模型。在空间上保留的组织上,CytoFormer达到了0.85的准确率和0.78的宏F1分数,并且其预测重现了整个保留切片的组织结构。该表示也具有迁移性:在编码器被冻结的情况下,CytoFormer特征上的线性头在四个专家标注的基准测试中表现优于六个病理基础模型,包括在预训练中未涉及的器官和细胞类型。最后,在交互式主动学习环境中,CytoFormer的嵌入显著比现有的病理基础模型更具标签效率,仅凭少量标注便能在相似肿瘤中检测正常上皮,F1分数为0.82,并在F1分数上领先最强基线0.13。CytoFormer将配对的H&E和空间转录组学转化为可重用的、高效标签的表示,用于常规组织学的细胞级分析。
cs.CV / 203 / 2608.16721

GenRouter: Unified Workflow Routing for Agentic Image Generation

GenRouter:用于自主图像生成的统一工作流路由
Chen, Harold Haodong, Hou, Zhiyu, Shu, Wen-Jie, Ruan, Weilin, Xu, Yingjie, Guo, Litao, Chen, Ying-Cong
Abstract
The rapid evolution of text-to-image (T2I) generation models has effectively solved the foundational challenge of raw pixel synthesis, shifting the community's focus toward fulfilling increasingly intricate user requests. While recent agentic image generation workflows enhance static inference with advanced capabilities like external knowledge retrieval and iterative reasoning, they mostly operate in isolated silos with fixed ``one-size-fits-all" topologies. This inevitably leads to severe compute-mismatch, where simple queries are forced through computationally heavy pipelines. To bridge this gap, we present GenRouter, the first unified workflow routing framework for agentic image generation. We first formulate GenCanvas, standardizing diverse agentic pipelines into a universal set of foundational primitives and executable templates. Operating over this unified space, GenRouter adaptively routes heterogeneous prompts to their optimal workflows via (i) demand profiling, (ii) experience matching, and (iii) Pareto filtering. Extensive experiments across diverse benchmarks demonstrate that GenRouter achieves superior visual alignment while reducing execution costs by over 95% and latency by 65% compared to heavyweight static pipelines. Furthermore, the system continuously self-evolves via accumulated experience, enabling robust zero-shot generalization that boosts performance and halves computational overhead.
Chinese Translation
文本到图像(T2I)生成模型的快速发展有效解决了原始像素合成的基础挑战,使得研究者的关注点转向满足日益复杂的用户请求。尽管近期的自主图像生成工作流通过外部知识检索和迭代推理等先进能力增强了静态推理,但它们大多在固定的“千篇一律”拓扑结构中孤立运行。这不可避免地导致了严重的计算不匹配,简单查询被迫通过计算负担沉重的管道。为了解决这一问题,我们提出了GenRouter,这是第一个用于自主图像生成的统一工作流路由框架。我们首先制定了GenCanvas,将多样的自主管道标准化为一组通用的基础原语和可执行模板。在这个统一的空间中,GenRouter通过(i)需求分析,(ii)经验匹配和(iii)帕累托过滤,适应性地将异构提示路由到其最佳工作流。针对不同基准的广泛实验表明,与重量级静态管道相比,GenRouter在实现更优视觉对齐的同时,执行成本降低超过95%,延迟减少65%。此外,该系统通过积累经验不断自我演进,实现了强大的零-shot泛化,提升了性能并将计算开销减半。
cs.CV / 204 / 2608.16725

Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI

多中心乳腺MRI图像数据集质量保证的无监督异常检测
Tappermann, Chiara, Renisch, Steffen, Schwen, Lars Ole, Meine, Hans, Hahn, Horst K., Petersen, Eike
Abstract
Corrupted, inconsistent, or anomalous data silently threatens the safety and reliability of medical AI. Despite growing regulatory recognition of dataset quality assurance (QA) for high-risk medical AI, scalable automated detection remains underdeveloped. We employ unsupervised anomaly detection (AD) and out-of-distribution (OOD) detection as an automated dataset QA mechanism for multi-center dynamic contrast-enhanced breast MRI. We build a controlled AD benchmark of 17 realistic QA-relevant anomaly types from six public datasets (protocol violations, processing errors, incorrect anatomical regions) and propose a taxonomy of radiological image anomalies based on human visual perception, enabling fine-grained analysis of AD failure modes. The benchmark includes near-, medium-far-, far-OOD samples, as well as in-distribution and external normal data. Four methods are evaluated: a projection-based method extended with a domain-specific feature extractor and a novel positional encoding, a reconstruction-based approach extended to full 3D volumes with an augmented training objective, and two unmodified hybrid OOD detection methods. Medium-far- and far-OOD samples are detected reliably, whereas near-OOD samples and external normal data from unseen institutions expose method-specific differences. The 3D reconstruction-based approach best balances detection performance (AUROC: 0.936) and generalization to unseen institutions. The projection-based method with positional encoding achieves the highest overall detection performance (AUROC: 0.954). Both hybrid methods exhibit critical failure modes, confirming that methods validated for one modality or anatomy may not generalize without domain-specific adaptation. Implants and mastectomies remain an open challenge for all methods. Our results establish a foundation and practical guidance on scalable unsupervised QA in medical AI pipelines.
Chinese Translation
损坏、不一致或异常的数据悄然威胁着医疗人工智能的安全性和可靠性。尽管对高风险医疗人工智能的数据集质量保证(QA)日益受到监管的重视,但可扩展的自动检测仍然不够成熟。我们采用无监督异常检测(AD)和分布外(OOD)检测作为多中心动态对比增强乳腺MRI的自动化数据集QA机制。我们构建了一个由六个公共数据集中的17种现实的与QA相关的异常类型组成的受控AD基准(协议违规、处理错误、解剖区域不正确),并提出了一种基于人类视觉感知的放射影像异常分类法,使得对AD失败模式的细粒度分析成为可能。该基准包括近、远中、远外OOD样本,以及分布内和外部正常数据。我们评估了四种方法:一种基于投影的方法,扩展了领域特定特征提取器和新颖的位置信息编码;一种基于重建的方法,扩展到完整的3D体积,并具有增强的训练目标;以及两种未修改的混合OOD检测方法。中远和远外OOD样本的检测可靠,而近外OOD样本和来自未见机构的外部正常数据则暴露出方法特定的差异。基于3D重建的方法在检测性能(AUROC: 0.936)和对未见机构的泛化之间取得了最佳平衡。带有位置信息编码的投影方法实现了最高的整体检测性能(AUROC: 0.954)。两种混合方法均表现出关键的失败模式,确认了在一种模态或解剖结构上验证的方法可能无法在没有领域特定适应的情况下进行泛化。植入物和乳腺切除术对所有方法仍然是一个开放的挑战。我们的结果为医疗人工智能管道中的可扩展无监督QA奠定了基础,并提供了实际指导。
cs.CV / 205 / 2608.16745

VicEdit: Learning to Edit Videos from Visual In-Context Examples

VicEdit:从视觉上下文示例中学习视频编辑
Wang, Yuji, Hu, Teng, Chen, Yuheng, Yi, Ran, Feng, Han, Cao, Weijian, Wang, Chengjie, Ma, Lizhuang, Zhang, Jiangning
Abstract
Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptual gap, we propose Visual In-context Editing, a new paradigm elevating video editing from textual instructions to multi-modal visual guidance encompassing single image, image pair, and video pair. To facilitate this paradigm, we curate VicEdit-400K, the first large-scale dataset for visual in-context video editing. We develop an automated pipeline to generate 400K high-quality samples across ten task types, ensuring superior visual fidelity and semantic consistency through multi-dimensional filtering. Leveraging this foundation, we introduce VicEdit, a unified framework to bridge visual and textual contexts. To adaptively extract editing semantics from heterogeneous references, we design Modality-Adaptive Semantic Distillation, which produces modality-specific semantic tokens from visual references. These tokens are then synergistically integrated with textual instructions through Dual-Context Injection, enabling the generation process to benefit from both visual and textual signals. Extensive evaluations on VicEditBench demonstrate that VicEdit achieves state-of-the-art performance across both basic instruction editing and visual in-context editing tasks, establishing visual in-context learning as a powerful and controllable paradigm for video editing.
Chinese Translation
尽管在基于指令的视频编辑方面取得了进展,但单一模态的文本指令在传达细致纹理和复杂动态方面固有地存在困难。为了弥补这一感知差距,我们提出了视觉上下文编辑(Visual In-context Editing),这一新范式将视频编辑从文本指令提升到涵盖单幅图像、图像对和视频对的多模态视觉指导。为了促进这一范式的发展,我们策划了VicEdit-400K,这是首个用于视觉上下文视频编辑的大规模数据集。我们开发了一个自动化管道,生成了40万高质量样本,涵盖十种任务类型,通过多维过滤确保卓越的视觉保真度和语义一致性。在此基础上,我们引入了VicEdit,一个统一框架,用于桥接视觉和文本上下文。为了自适应地从异构参考中提取编辑语义,我们设计了模态自适应语义蒸馏(Modality-Adaptive Semantic Distillation),该方法从视觉参考中生成特定于模态的语义标记。这些标记随后通过双上下文注入(Dual-Context Injection)与文本指令协同集成,使生成过程能够同时受益于视觉和文本信号。在VicEditBench上的广泛评估表明,VicEdit在基本指令编辑和视觉上下文编辑任务中均实现了最先进的性能,确立了视觉上下文学习作为视频编辑的强大且可控的范式。
cs.CV / 206 / 2608.16748

Beyond Uncertainty: Generalizable Failure Monitoring for Surgical Segmentation under Acquisition Degradation

超越不确定性:针对采集退化的外科分割通用故障监测
Pham, Hieu D., Cao, Dang P. M., Huynh, Thanh Trung
Abstract
Surgical segmentation networks can fail silently under acquisition degradation: predicted masks may be wrong even when model confidence remains high. Existing deployment-time monitors rely primarily on uncertainty estimates and can therefore miss confident failures. We present TCSR-Monitor (Temporal Conformal Surgical Risk Monitor), a post-hoc failure-monitoring framework that combines confidence with observable shape, temporal-consistency, and image-quality cues. TCSR-Monitor wraps a frozen segmentation model, requires no model internals, and operates without ground truth at deployment. We also introduce a validation protocol to assess whether alarms remain credible under distribution shift. On EndoVis 2017, leave-one-corruption-out evaluation shows that TCSR-Monitor generalizes to unseen acquisition degradations and substantially outperforms confidence-based baselines. A circularity control confirms that it predicts segmentation failure rather than simply detecting corrupted images. Mondrian conformal calibration balances miss-rates across degradation severities, but a single global threshold still produces false alarms on up to 40% of correctly segmented frames at moderate corruption. Zero-shot transfer to SAM2 demonstrates feature portability, although entropy outperforms the transferred monitor at both evaluated thresholds. Overall, reliable monitoring under acquisition degradation benefits from complementary observable signals beyond confidence alone, but substantial false-alarm and transfer limitations remain.
Chinese Translation
在采集退化情况下,外科分割网络可能会悄然失败:即使模型信心保持高水平,预测的掩模也可能是错误的。现有的部署时监测器主要依赖于不确定性估计,因此可能会错过自信的失败。我们提出了 TCSR-Monitor(时间一致性外科风险监测器),这是一个后验故障监测框架,结合了信心与可观察的形状、时间一致性和图像质量线索。TCSR-Monitor 封装了一个冻结的分割模型,无需模型内部信息,并在部署时不依赖于真实标签。我们还引入了一种验证协议,以评估在分布变化下警报是否仍然可信。在 EndoVis 2017 上,逐一去除腐败评估显示 TCSR-Monitor 能够推广到未见的采集退化,并显著优于基于信心的基线。循环控制确认其预测分割失败,而不仅仅是检测损坏的图像。Mondrian 置信校准在不同退化严重程度之间平衡漏报率,但单一全局阈值在中等腐败情况下仍会对多达 40% 的正确分割帧产生误报。零样本迁移到 SAM2 展示了特征的可移植性,尽管在评估的两个阈值下,熵的表现优于迁移的监测器。总体而言,在采集退化下,可靠的监测需要超越单一信心的互补可观察信号,但仍存在显著的误报和迁移限制。
cs.CV / 207 / 2608.16756

Binarized High-Efficiency RAW Video Restoration and Beyond

二值化高效RAW视频恢复及其扩展
Zhu, Tianyu, Fu, Ying, Li, Hesong, Zhang, Gengchen, Yuan, Xin, Zhang, Yulun
Abstract
RAW video restoration is fundamental to high-quality low-level perception and serves as the basis for a wide range of downstream vision applications. While binary neural networks (BNNs) enable efficient lightweight deployment for image enhancement, their deficiencies in modeling temporal coherence and activation value distributions hinder their effectiveness when applied to video scenarios. In this paper, we propose BinRVR, a binarized RAW video restoration framework that reduces computation and parameters by approximately 96% while incurring only about 4% performance degradation. Specifically, we present a Binarized Information Interaction Module (BIIM) to jointly model spatial and temporal information in an efficient and unified manner. Moreover, we develop a Distribution-Aware Binarized Convolution (DAB-Conv) that leverages the statistics of full-precision activations to mitigate quantization errors. The proposed framework further supports multi-bit quantization, enabling flexible accuracy-efficiency trade-offs across different hardware constraints. Extensive experiments demonstrate that our BinRVR achieves competitive performance compared with state-of-the-art binarized methods on RAW video restoration tasks, including low-light enhancement, denoising, deblurring, and super-resolution. We further explore the potential of our method on downstream video applications, including object detection and monocular depth estimation.
Chinese Translation
RAW视频恢复是高质量低级感知的基础,并为广泛的下游视觉应用提供支持。虽然二值神经网络(BNNs)能够实现图像增强的高效轻量级部署,但它们在建模时间一致性和激活值分布方面的不足,限制了其在视频场景中的有效性。本文提出了BinRVR,一个二值化RAW视频恢复框架,计算和参数减少约96%,同时性能仅下降约4%。具体而言,我们提出了一个二值化信息交互模块(BIIM),以高效统一的方式共同建模空间和时间信息。此外,我们开发了一种分布感知二值卷积(DAB-Conv),利用全精度激活的统计信息来减轻量化误差。所提出的框架进一步支持多位量化,使得在不同硬件约束下实现灵活的准确性与效率权衡。大量实验表明,我们的BinRVR在RAW视频恢复任务上与最先进的二值化方法相比,表现出竞争力,包括低光增强、去噪、去模糊和超分辨率。我们还进一步探索了我们的方法在下游视频应用中的潜力,包括目标检测和单目深度估计。
cs.CV / 208 / 2608.16765

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

TRACE-Bench:多参考图像生成的解构与诊断
Wang, Haoran, Ma, Chaofan, Yi, Ran, Ma, Lizhuang
Abstract
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor ($f$), Disentangle ($g$), Apply ($\oplus$), and Compose ($C$). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement ($g$) and attribute binding ($\oplus$) rather than scene-level composition ($C$), with even the best model scoring only 0.74 on attribute fidelity. Project page: https://amuseum-whr.github.io/TraceBench
Chinese Translation
尽管近年来在多模态统一模型用于多参考图像生成方面取得了进展,但现有基准仍围绕预定义的任务类型(例如,“主题组合”)进行组织,这些任务类型不适合这种组合设置,导致覆盖面碎片化、复杂性失控以及缺乏诊断价值。我们认识到多样的多参考任务共享一组共同的原子操作,因此采用以能力为导向的视角,形式化了四个操作符:Anchor ($f$)、Disentangle ($g$)、Apply ($igoplus$) 和 Compose ($C$)。任何多参考提示都可以表示为这些操作符的组合公式,其结构复杂性由操作符槽的数量来量化。在此基础上,我们构建了TRACE-Bench,包含约1,600个评估案例,涵盖槽数1至8,基于631个公式模板和约4,000张跨越多种艺术风格和现实主题的参考图像。公式结构直接驱动了一种与操作符对齐的评估协议,用于按能力评分以及递归失败定位的诊断树分析。对9个领先模型的评估揭示了整体评分无法察觉的洞见:主要瓶颈在于解构 ($g$) 和属性绑定 ($igoplus$),而非场景级组合 ($C$),即使是最佳模型在属性保真度上也仅得分0.74。项目页面:https://amuseum-whr.github.io/TraceBench
cs.CV / 209 / 2608.16785

Calibration-Free Vehicle Speed Estimation: A Monocular Keypoint-Template Approach

无校准车辆速度估计:单目关键点-模板方法
Su, Gaofeng, Li, Keya, Sengupta, Raja, Kockelman, Kara M.
Abstract
This paper proposes a calibration-free framework for reliably and effectively estimating vehicle speeds from monocular videos, without relying on roadway features, camera calibration, or roadway-feature-based reference objects. The proposed framework estimates vehicle speeds using a 36-keypoint vehicle template and a homography matrix updated at each frame. A YOLO-based keypoint detection module is trained on diverse datasets, and two estimation strategies are compared: keypoint-only tracking and warped optical flow with dense spatial aggregation. Speed is estimated by projecting displacements into metric space using the homography, with validation conducted on over 400 video clips from roadside and overhead datasets, covering speeds from 30 to 100 mph. The method achieves reliable speed estimation on the VS13 and BrnoCompSpeed datasets, with the warped optical flow method delivering MAEs of 15.0% and 9.7%, respectively, and 77.9% and 93.1% of estimates falling within +/-20% error. After applying a 10% trim to remove edge-of-frame outliers, performance improves to MAEs of 11.7% and 7.6%, with within-+/-20% accuracy increasing to 85.3% and 95.4%. This work addresses key limitations of existing vision-based approaches and enables low-cost and efficient speed enforcement using portable devices such as dashcams and smartphones, thereby supporting citizen-based enforcement programs for traffic safety.
Chinese Translation
本文提出了一种无校准框架,用于从单目视频中可靠且有效地估计车辆速度,无需依赖道路特征、相机校准或基于道路特征的参考物体。所提出的框架使用36个关键点的车辆模板和在每帧更新的单应性矩阵来估计车辆速度。基于YOLO的关键点检测模块在多样化的数据集上进行训练,并比较了两种估计策略:仅关键点跟踪和带有密集空间聚合的扭曲光流。通过使用单应性将位移投影到度量空间中来估计速度,并在超过400个来自路边和高空数据集的视频片段上进行验证,覆盖速度范围从30到100英里每小时。该方法在VS13和BrnoCompSpeed数据集上实现了可靠的速度估计,扭曲光流方法的平均绝对误差(MAE)分别为15.0%和9.7%,并且77.9%和93.1%的估计值在±20%的误差范围内。在应用10%的修剪以去除边缘帧异常值后,性能提升至MAE为11.7%和7.6%,而±20%的准确率提高至85.3%和95.4%。该研究解决了现有基于视觉的方法的关键局限性,并使得使用便携设备(如行车记录仪和智能手机)进行低成本和高效的速度执法成为可能,从而支持基于公民的交通安全执法项目。
cs.CV / 210 / 2608.16786

Revisiting Classifier-Free Guidance Methods in Latent Diffusion Models

重新审视潜在扩散模型中的无分类器引导方法
Sergievskii, Artem, Turevich, Artyom, Kastryulin, Sergey
Abstract
Inference-time quality-enhancement methods are an effective and widely adopted means of improving diffusion models without expensive retraining. We study a family of training-free techniques conceptually rooted in Classifier-Free Guidance (CFG), most of which were originally proposed on older U-Net diffusion models and validated using metrics that assess image quality in isolation, without accounting for compositional alignment or semantic correspondence between the generated image and its associated text prompt. We re-evaluate eight such methods on two open-weight rectified-flow transformers under a fixed per-model protocol and three compositional-alignment benchmarks. No method consistently improves on CFG across the measured criteria. APG obtains several nominal best scores, but the corresponding gains often remain within the estimated evaluation uncertainty. Attention-perturbation methods provide isolated gains on SD3.5 Medium and more frequent degradations on FLUX.2 [klein] 4B Base, while CFG remains a competitive lower-cost baseline.
Chinese Translation
推理时质量增强方法是一种有效且广泛采用的手段,可以在不进行昂贵的再训练的情况下改善扩散模型。我们研究了一系列概念上根植于无分类器引导(Classifier-Free Guidance, CFG)的无训练技术,其中大多数最初是在较旧的U-Net扩散模型上提出,并使用评估图像质量的指标进行验证,这些指标未考虑生成图像与其相关文本提示之间的组合对齐或语义对应关系。我们在两个开放权重的修正流变换器上,采用固定的每模型协议和三个组合对齐基准,重新评估了八种此类方法。在测量的标准中,没有任何方法在CFG上持续改进。APG获得了几个名义上的最佳分数,但相应的增益往往仍在估计的评估不确定性范围内。注意力扰动方法在SD3.5 Medium上提供了孤立的增益,而在FLUX.2 [klein] 4B Base上则更频繁地出现降级,而CFG仍然是一个具有竞争力的低成本基线。
cs.CV / 211 / 2608.16791

Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching

引导流动:通过梯度引导流匹配反演人脸识别模型
Lu, Ye, Wang, Shen, Zhang, Zhaoyang, Yan, Yihan, Liu, Li, Liu, Runze, Sun, Fanghui
Abstract
Model Inversion Attacks (MIAs) aim to reconstruct representative training samples of target identities from face recognition models, exposing critical security vulnerabilities. Existing methods typically rely on indirect guidance or highly stochastic guidance, making it difficult to stably optimize generation trajectories toward target facial images. In this paper, we propose Steering Flow Model Inversion (SFMI), a novel two-stage white-box model inversion method that reformulates inversion as a trajectory-steering task. Specifically, Step I, Learning a Generic Flow Matching Prior, pre-trains a generic unconditional Flow Matching model to encode the manifold of human faces as a robust prior. Step II, Attacking with Progressive Guidance Scheduler (PGS), injects time-dependent target-specific gradients during sampling. By backpropagating through the target model to obtain gradients from intermediate generated states, PGS progressively injects adaptive guidance signals into the vector field. This process effectively steers the current generative flow from random noise toward the high-density regions of the target class. Under an identity-disjoint cross-evaluation setting using the CelebA dataset, SFMI achieves an ACC of 0.9248, an FID of 22.61, and an LPIPS of 0.3874 on the ArcFace target. Extensive experiments on multiple target models demonstrate that SFMI achieves competitive state-of-the-art performance in attack success and visual fidelity under the evaluated white-box protocol.
Chinese Translation
模型反演攻击(MIAs)旨在从人脸识别模型中重建目标身份的代表性训练样本,从而暴露出关键的安全漏洞。现有方法通常依赖于间接引导或高度随机的引导,使得稳定地优化生成轨迹以接近目标人脸图像变得困难。本文提出了一种新的两阶段白盒模型反演方法——引导流模型反演(SFMI),将反演重新构建为轨迹引导任务。具体而言,第一步,学习通用流匹配先验,预训练一个通用的无条件流匹配模型,以将人脸的流形编码为一个稳健的先验。第二步,使用渐进引导调度器(PGS)进行攻击,在采样过程中注入时间依赖的目标特定梯度。通过反向传播到目标模型以从中间生成状态获取梯度,PGS逐步将自适应引导信号注入到向量场中。这个过程有效地将当前的生成流从随机噪声引导到目标类别的高密度区域。在使用CelebA数据集进行身份不重叠的交叉评估设置下,SFMI在ArcFace目标上达到了0.9248的准确率(ACC)、22.61的Fréchet距离(FID)和0.3874的感知相似性(LPIPS)。在多个目标模型上的广泛实验表明,SFMI在评估的白盒协议下,在攻击成功率和视觉保真度方面达到了具有竞争力的最先进性能。
cs.CV / 212 / 2608.16793

PixRestore: Unified Image Restoration via Pixel Diffusion Transformer

PixRestore:通过像素扩散变换器实现统一图像修复
Sun, Lingchen, Wu, Rongyuan, Kong, Xiangtao, Zhao, Jixin, Yi, Qiaosi, Sun, Yujing, Liu, Shuaizheng, Zhang, Zhengqiang, Zhang, Lei
Abstract
Unified image restoration (UIR) aims to recover high-quality (HQ) content from low-quality (LQ) images with different degradations using a single model. Most recent methods adapt large pretrained text-to-image (T2I) latent diffusion models for their strong capacity and generative priors. However, the variational autoencoder (VAE) in latent T2I models may discard restoration-sensitive details, while the open-ended synthesis prior can introduce content-inconsistent artifacts. We present PixRestore, a VAE-free pixel-space Diffusion Transformer (DiT) for UIR, where the diffusion backbone is trained entirely from scratch, without relying on T2I pretraining. PixRestore performs flow matching directly on patchified pixels, preserving fine-grained details while keeping the token sequence tractable. To adapt to different degradations, PixRestore learns to predict the reliability of layer features using LQ--HQ DINO feature similarity. Features from more reliable layers are fused as dense conditioning, while less reliable layers receive stronger HQ-feature supervision to encourage degradation removal. We train PixRestore on a large-scale corpus of diverse scenes and degradations, and further finetune it into a one-step generator using DINO-based adversarial objectives for efficient inference. Experiments on public benchmarks and real-world test sets show that, with only about 50M parameters and single-step inference, PixRestore achieves the best overall fidelity, perceptual quality, and robustness to degradations among competing UIR models while being far more efficient. Larger PixRestore variants can further boost performance, demonstrating the scalability of our pixel-space design. Code and the curated benchmark can be found at https://github.com/csslc/PixRestore.
Chinese Translation
统一图像修复(UIR)旨在通过单一模型从不同退化的低质量(LQ)图像中恢复高质量(HQ)内容。最近的方法大多采用大型预训练的文本到图像(T2I)潜在扩散模型,因其强大的能力和生成先验。然而,潜在T2I模型中的变分自编码器(VAE)可能会丢失对修复敏感的细节,而开放式合成先验可能会引入内容不一致的伪影。我们提出了PixRestore,一种无VAE的像素空间扩散变换器(DiT)用于UIR,其中扩散骨干网络完全从头开始训练,而不依赖于T2I的预训练。PixRestore直接在分块像素上执行流匹配,保留细粒度细节,同时保持标记序列的可处理性。为了适应不同的退化,PixRestore学习使用LQ-HQ DINO特征相似性来预测层特征的可靠性。来自更可靠层的特征被融合为密集条件,而较不可靠的层则接受更强的HQ特征监督,以促进退化去除。我们在一个大规模的多样场景和退化的语料库上训练PixRestore,并进一步通过基于DINO的对抗目标将其微调为一步生成器,以实现高效推理。在公共基准和真实世界测试集上的实验表明,PixRestore仅用约5000万个参数和单步推理,在竞争的UIR模型中实现了最佳的整体保真度、感知质量和对退化的鲁棒性,同时效率远高于其他模型。更大的PixRestore变体可以进一步提升性能,展示了我们像素空间设计的可扩展性。代码和整理的基准可以在https://github.com/csslc/PixRestore找到。
cs.CV / 213 / 2608.16805

Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models

诊断大型视觉-语言模型中的密集同类属性错误绑定
Xu, Yuanzhi, Gao, Qian, Fan, Jun, Ding, Guohui, Yang, Zhenyu, Xiao, Yuteng, Lin, Sixue
Abstract
Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance. Generic visual-question-answering accuracy marks the response as wrong, while object-hallucination metrics may regard both the object and attribute as image-supported; neither reveals the transfer. This study formalizes this blind spot as Dense Same-Class Attribute Misbinding (DSCAM) and presents InstaBind-Lite, a controlled benchmark that makes it directly measurable. Its 524 images contain 529 curated groups of 3-6 same-class entities, 1773 boxed instances, ordered neighbors, distinguishable color-like attributes, and four complementary question levels, yielding 9580 deterministically evaluated questions. Unlike existing protocols, source-instance annotations separate unsupported generation and recognition failure from an attribute copied from another visible entity. Binding-specific metrics further quantify transfer frequency, adjacency, ordinal distance, and intervention effects. Across five open-source and two commercial/API models, the open-source systems average 19.84% Misbinding Rate and the API systems 7.55%; these errors are hidden by aggregate accuracy. Among identifiable transfers, 80.70% and 81.51%, respectively, originate from adjacent instances. Localization and instance-first interventions help selected models but are not universal remedies. InstaBind-Lite therefore turns previously undifferentiated wrong answers into source-identifiable failure categories and tests a reliability dimension that conventional benchmarks cannot determine: whether a model knows not only what is visible, but which instance owns each attribute.
Chinese Translation
大型视觉-语言模型能够识别拥挤场景中的物体和属性,但却可能将属性分配给错误的同类实例。通用视觉问答的准确性将此响应标记为错误,而物体幻觉指标可能将物体和属性视为图像支持的;两者都未能揭示转移问题。本研究将这一盲点形式化为密集同类属性错误绑定(Dense Same-Class Attribute Misbinding,DSCAM),并提出了InstaBind-Lite,一个可控基准,使其可以直接测量。该基准包含524幅图像,涵盖529组经过精心策划的3-6个同类实体,1773个框选实例,有序邻居,易于区分的颜色类属性,以及四个互补的问题层次,共产生9580个确定性评估的问题。与现有协议不同,源实例注释将不支持的生成和识别失败与从另一个可见实体复制的属性分开。特定绑定的指标进一步量化了转移频率、邻近性、序数距离和干预效果。在五个开源模型和两个商业/API模型中,开源系统的平均错误绑定率为19.84%,而API系统为7.55%;这些错误被整体准确性掩盖。在可识别的转移中,分别有80.70%和81.51%来源于相邻实例。定位和实例优先干预对选定模型有帮助,但并非普遍有效。因此,InstaBind-Lite将以前未区分的错误答案转变为源可识别的失败类别,并测试一个传统基准无法确定的可靠性维度:模型是否不仅知道什么是可见的,还知道每个属性归属于哪个实例。
cs.CV / 214 / 2608.16810

Unsupervised Learning of Cell Instances with Generative Routing Pyramids

基于生成路由金字塔的细胞实例无监督学习
Liu, Ziwen, Weigert, Martin
Abstract
Identifying and representing object instances such as cells or nuclei is a common task in microscopy image analysis. Established machine learning workflows typically use supervised detection or segmentation followed by feature extraction or classification, which requires manual annotations and treats instance segmentation and cell representation as separate stages. We describe a new unsupervised method for cell instance segmentation and phenotypic classification from unlabeled microscopy images. Our method is based on reconstructing each image using a coarse-to-fine routing pyramid that associates pixels with spatially sparse latent sources. The resulting pixel-to-latent associations yield instance masks, while the source latents encode cell morphology. We demonstrate competitive performance in instance segmentation across diverse cell morphologies and imaging modalities, as well as generative modeling of cellular phenotypes under perturbations. Source code and checkpoints are available at https://github.com/weigertlab/routing-pyramids.
Chinese Translation
识别和表示对象实例,如细胞或细胞核,是显微镜图像分析中的一项常见任务。现有的机器学习工作流程通常采用监督检测或分割,随后进行特征提取或分类,这需要手动标注,并将实例分割和细胞表示视为两个独立的阶段。我们描述了一种新的无监督方法,用于从未标记的显微镜图像中进行细胞实例分割和表型分类。我们的方法基于使用粗到细的路由金字塔重建每个图像,该金字塔将像素与空间稀疏的潜在源关联。由此产生的像素与潜在源的关联生成实例掩膜,而源潜在则编码细胞形态。我们在多种细胞形态和成像模式下展示了在实例分割方面的竞争性能,以及在扰动下细胞表型的生成建模。源代码和检查点可在 https://github.com/weigertlab/routing-pyramids 获取。
cs.CV / 215 / 2608.16812

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

通过概念缩放和密集监督解锁图像编辑的潜力
Cui, Long, Liu, Xiaoqian, Qin, Qi, Xin, Yi, Lin, Tao, Li, Jianguo, Zhang, Linfeng
Abstract
Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the insufficient attention to edit concept granularity and the training inefficiency caused by sparse supervision signals. To address these issues, we establish a comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts and build ConceptEdit-12M, a massive dataset of 12 million high-quality editing pairs via an improved synthesis framework. This library-driven approach effectively rectifies the distribution collapse of generated data while ensuring high data fidelity. Furthermore, we propose a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs. By providing richer learning signals, this strategy significantly enhances both training efficiency and overall model performance. Training results validate our strategy, significantly outperforming prior works. Finally, we present ConceptEdit-Bench, a granular evaluation suite designed to diagnose model capabilities across a vast array of real-world scenarios.
Chinese Translation
现有的图像编辑框架主要遵循文本到图像扩散模型的训练范式。然而,将这一范式扩展到图像编辑时,突显出两个固有的差异,具体而言,编辑概念粒度关注不足以及稀疏监督信号导致的训练效率低下。为了解决这些问题,我们建立了一个全面的层次分类法,涵盖超过1000个细粒度编辑概念,并通过改进的合成框架构建了ConceptEdit-12M,这是一个包含1200万个高质量编辑对的大型数据集。这种基于库的方法有效地纠正了生成数据的分布崩溃,同时确保了高数据保真度。此外,我们提出了一种密集监督训练策略,将多个不干扰的概念合成到单个图像对中。通过提供更丰富的学习信号,该策略显著提高了训练效率和整体模型性能。训练结果验证了我们的策略,显著优于先前的工作。最后,我们推出了ConceptEdit-Bench,这是一个细粒度评估套件,旨在诊断模型在各种现实场景中的能力。
cs.CV / 216 / 2608.16855

Can Unsupervised Methods Outperform Supervised Deep Learning When Ground Truth Is Sparse? A Case Study of Bronchovascular Bundle Segmentation in Low-Dose CT

在真实标签稀疏的情况下,无监督方法能否超越监督深度学习?以低剂量CT中的支气管血管束分割为例
Mrukwa, Anna, Socha, Marek, Suwalska, Aleksandra, Durawa, Agata, Jelitto, Malgorzata, Dziadziuszko, Katarzyna, Szurowska, Edyta, Bozek, Pawel, Marczyk, Michal, Rzyman, Witold, Dziadziuszko, Rafal, Polanska, Joanna
Abstract
Background Lung cancer remains the deadliest cancer worldwide because it is often diagnosed too late. Effective treatment depends on detection at an early screening stage. However, the growing number of patients and the limited number of radiologists lead to prolonged diagnostic waiting times. In very early stage lung cancer, nodule visibility is further reduced by adjacent blood vessels and airway walls, because nodules are often connected to or supplied by these structures. Task-specific analysis of the bronchovascular bundle is therefore important for efficient nodule detection, and its removal can increase the diagnostic potential of lung cancer screening. Materials and Methods To assess the efficacy of the proposed method, we used series from widely utilized LDCT datasets, including the Duke Lung Cancer Screening (DLCS) dataset and the Pilot Pomeranian Lung Cancer Screening Program. The proposed bronchovascular bundle segmentation pipeline, RONALD, operates on computed tomography images and returns binary masks of vessels and bronchi located in the lung parenchyma. The method includes a preprocessing stage with lung, lobe, and mediastinum segmentation, followed by separate vessel and bronchial tree segmentation. Results The proposed pipeline segmented the bronchovascular bundle in low-dose computed tomography scans while improving nodule retention compared with other segmentation methods: from 93.98% and 90.36% to 100% in DLCS, and from 83.16% and 62.36% to 99.92% in the Pomeranian dataset. Conclusion The resulting segmentations can improve lung nodule detection in the very early stages of lung cancer.
Chinese Translation
背景:肺癌仍然是全球致死率最高的癌症,因为它通常在晚期才被诊断。有效的治疗依赖于在早期筛查阶段的检测。然而,患者数量的增加和放射科医生的数量有限导致了诊断等待时间的延长。在非常早期的肺癌中,结节的可见性因邻近的血管和气道壁而进一步降低,因为结节通常与这些结构相连或由其供血。因此,针对支气管血管束的特定任务分析对于高效的结节检测至关重要,其去除可以提高肺癌筛查的诊断潜力。材料与方法:为了评估所提方法的有效性,我们使用了广泛应用的低剂量CT(LDCT)数据集,包括杜克肺癌筛查(DLCS)数据集和波美拉尼亚肺癌筛查项目。所提出的支气管血管束分割管道RONALD在计算机断层扫描图像上运行,并返回位于肺实质中的血管和支气管的二进制掩膜。该方法包括一个预处理阶段,进行肺、叶和纵隔的分割,随后进行单独的血管和支气管树分割。结果:所提出的管道在低剂量计算机断层扫描中成功分割了支气管血管束,并在结节保留方面优于其他分割方法:在DLCS中从93.98%和90.36%提高到100%,在波美拉尼亚数据集中从83.16%和62.36%提高到99.92%。结论:所得到的分割结果可以改善肺癌非常早期阶段的肺结节检测。
cs.CV / 217 / 2608.16859

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

HarnessEval-W:将评估代理化于视觉世界
Chen, Weiliang, Sun, Haowen, Gao, Jun, Chi, Jiawei, Wang, Hanyang, Dai, Qiyu, Li, Yihao, Li, Hao, Gao, Jingnan, Hung, Yi-Hsin, Guo, Xingzhuo, Miao, Shangchen, Shi, Zhiyuan, Li, Xiang, Tian, Fengrui, Du, Weihua, Huang, Ziqi, Gao, Shenyuan, Huang, Siqiao, Liu, Mingyu, Li, Yifei, Wang, Shizun, Wang, Xi, Zhang, Tianqi, Luo, Xue, Ren, Xiyin, Ren, Jinshan, Shen, Xiaoyang, Hu, Xiaobo, Dou, Zhiyang, Ding, Mingyu, Yan, Yichao, Wang, Xinchao, Wang, Yizhou, Liu, Shilong, Zheng, Wenzhao, Duan, Yueqi, Gong, Yuan, Liu, Ziwei, Liu, Ming-Yu, Wu, Jialong, Lyu, Jiangran, Liu, Fangfu
Abstract
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute-force, leaving no reasoning chain that can be examined or verified. We introduce HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval-W interprets the context of each evaluation case, decomposes the evaluation question into measurable subproblems, and spawns specialized sub-agents, each equipped with tailored context and diagnostic tools to reason over its own subproblem. The parent agent then validates the gathered evidence and summarizes it into the final verdict. This hierarchical workflow turns every evaluation into a transparent evidence tree whose complete reasoning chain justifies the result. We apply HarnessEval-W to 18 representative world models over 330 evaluation cases. Its judgments closely align with human preferences while providing verifiable, fine-grained diagnoses of every generated rollout. We open-source the full pipeline as a live benchmark and invite the broad community to contribute to grow new skills and evaluation cases as world models evolve.
Chinese Translation
基准测试应该提供的不仅仅是一个标量分数:使评估可信赖的是支持该分数的推理。这对于世界模型尤其重要,因为判断一次回放需要理解物理、因果关系和世界状态是否正确演变。人类能够自然地识别这种违反,但现有的基准测试没有自动化这一能力:指标是通过暴力计算得出的,缺乏可以检查或验证的推理链。我们引入了HarnessEval-W,一个代理化的评估管道,将大型语言模型(LLM)生态系统中的鞍座范式引入世界模型基准测试。HarnessEval-W并不是应用固定的评估标准,而是解释每个评估案例的上下文,将评估问题分解为可测量的子问题,并生成专门的子代理,每个子代理都配备了量身定制的上下文和诊断工具,以便对其自身的子问题进行推理。然后,父代理验证收集到的证据并将其总结为最终裁决。这种分层工作流程将每次评估转化为一个透明的证据树,其完整的推理链为结果提供了合理的依据。我们将HarnessEval-W应用于18个具有代表性的世界模型,涵盖330个评估案例。其判断与人类偏好高度一致,同时提供可验证的、细致的每个生成回放的诊断。我们将完整的管道开源为一个实时基准,并邀请广泛的社区参与,以便在世界模型不断演变的过程中贡献新的技能和评估案例。
cs.CV / 218 / 2608.16863

SplatGuide: Geometric Priors from 3D Gaussians for Pose-Free Novel View Synthesis

SplatGuide:基于3D高斯的几何先验用于无姿态新视图合成
Zhang, Yejun, Wang, Zihan, Ji, Xu, Wang, Yihao, Hou, Yuxin, Fang, Junyuan, Kilpeläinen, Juho-Matti, Solin, Arno, Tavakoli, Hamed Rezazadegan, Rahtu, Esa, Kannala, Juho
Abstract
Generating photorealistic novel views from unposed images requires both 3D geometric understanding and the ability to synthesize unseen content. A natural strategy combines feed-forward 3DGS reconstruction with multi-view diffusion. Yet prior pipelines extract at most one signal from the reconstruction, either pixel rendering or learned features, while none exploits per-Gaussian visibility for occlusion-aware reference selection. This *information disconnect* leaves renderable geometry, visibility cues, and learned features unused. SplatGuide closes this disconnect by reusing a single 3DGS scene across three complementary roles. Rendered images provide pixel-aligned geometric conditioning. Per-Gaussian source-view indices are rendered into a target-view voting map for occlusion-aware reference selection. Reconstruction tokens supply feature-level guidance via cross-attention. All three signals derive from the same reconstruction forward pass. Across RealEstate10K, DL3DV, Tanks-and-Temples, and Mip-NeRF 360, SplatGuide achieves state-of-the-art pose-free novel view synthesis. On RealEstate10K, with a moderate number of input views, it surpasses the ground-truth-pose baseline.
Chinese Translation
从无姿态图像生成照片级真实感的新视图需要对3D几何的理解以及合成未见内容的能力。一种自然的策略是将前馈3D几何重建与多视图扩散相结合。然而,现有的流程最多只能从重建中提取一个信号,或者是像素渲染,或者是学习到的特征,而没有利用每个高斯的可见性进行考虑遮挡的参考选择。这种*信息断裂*使得可渲染的几何体、可见性线索和学习到的特征未被利用。SplatGuide通过在三个互补角色中重用单一的3D几何场景来弥合这一断裂。渲染的图像提供像素对齐的几何条件。每个高斯源视图索引被渲染到目标视图投票图中,以进行考虑遮挡的参考选择。重建令牌通过交叉注意力提供特征级指导。这三种信号均来自同一重建的前向传递。在RealEstate10K、DL3DV、Tanks-and-Temples和Mip-NeRF 360数据集上,SplatGuide实现了最先进的无姿态新视图合成。在RealEstate10K上,使用适量的输入视图,其性能超过了真实姿态基线。
cs.CV / 219 / 2608.16887

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

像素空间文本到图像扩散模型的实证研究
Jiang, Dengyang, Du, Ruoyi, Chen, Zhennan, Liu, Dongyang, Wang, Zanyi, Zheng, Mingzhe, Yang, Xiangpeng, Cai, Huanqia, Hao, Aiming, Jiang, Yuming, Gao, Peng, Yang, Harry, Hoi, Steven
Abstract
This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction target, decoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation.
Chinese Translation
本文探讨了生成建模中一个日益重要的话题:像素空间扩散模型。尽管已有众多研究探讨了这一主题,但大多数集中于小规模或类别条件设置。因此,尚未找到一种实用的方法来训练能够与成熟的潜在空间模型相媲美或超越的像素空间模型。通过全面的实证研究,我们首先观察到,在像素空间中直接进行大规模预训练的收敛速度显著慢于在潜在空间中的收敛速度。这一观察促使我们提出了一种潜在到像素的策略,该策略在潜在空间中有效获取生成先验,并在后期训练过程中过渡到像素空间。随后,我们系统地研究了影响这一过渡的关键设计选择,包括权重初始化、数据组成、预测目标、解码器架构和噪声调度,并确定了一种实用的方法,使得生成的像素空间模型能够与其潜在空间模型相匹配或超越,同时实现3.18到4.75倍的端到端推理加速。我们希望我们的研究结果能为未来像素空间生成的研究提供有用的实证见解和实用指南。
人工智能 (Artificial Intelligence)
200
cs.AI / 1 / 2608.14550

FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment

FLOPs与实际工作:AI效率评估中复制的重要性
Roque, Enrique Barba, Cruz, Luís
Abstract
AI efficiency has recently taken the spotlight in both academy and industry due to massive model scales, high energy demands, and environmental costs. While reporting Floating Point Operations (FLOPs) is a traditional approach for assessing computational costs, the relationship between FLOPs and execution time is not straightforward, as layers with the same number of FLOPs may not have the same execution time because some operations are more easily parallelized than others. This paper sets out to replicate the original experiments from a study that proposed the $\alpha-FLOPs$ estimation formula to verify whether the results remain applicable on newer, more powerful hardware. During the replication process, we identify limitations in the replication materials provided by the original study, including a lack of specific dependency details and transparency regarding regression data. Our results validate the thesis that raw FLOPs alone are not an appropriate metric for execution time, as spatial dimensions remain more easily parallelized than kernel dimensions. However, fine-grained measurements reveal that the relationship is much less straightforward than previously shown, with newer hardware exhibiting instabilities and discontinuities in execution time, including jumps and oscillations, that the $\alpha-FLOPs$ formula generally underestimates. Ultimately, this work validates the empirical findings from the original study but shows negative results when applying the $\alpha-FLOPs$ estimation. We also highlight the critical need for complete and accurate replication packages for research on hardware-dependent efficiency assessment and provide a complete replication package for our implementation to facilitate further study.
Chinese Translation
由于模型规模庞大、高能耗和环境成本,AI效率最近在学术界和工业界引起了广泛关注。虽然报告浮点运算数(FLOPs)是评估计算成本的传统方法,但FLOPs与执行时间之间的关系并不简单,因为具有相同FLOPs数量的层可能具有不同的执行时间,因为某些操作比其他操作更容易并行化。本文旨在复制一项研究中提出的$eta-FLOPs$估计公式的原始实验,以验证在更新、更强大的硬件上结果是否仍然适用。在复制过程中,我们发现原始研究提供的复制材料存在一些局限性,包括缺乏具体的依赖细节和回归数据的透明度。我们的结果验证了仅使用原始FLOPs作为执行时间度量并不合适的论点,因为空间维度比内核维度更容易并行化。然而,细粒度的测量结果表明,这一关系远比之前显示的复杂,新的硬件在执行时间上表现出不稳定性和不连续性,包括跳跃和振荡,而$eta-FLOPs$公式通常低估了这些现象。最终,这项工作验证了原始研究的实证发现,但在应用$eta-FLOPs$估计时显示出负面结果。我们还强调了对于硬件依赖效率评估研究,完整且准确的复制包的关键需求,并提供了我们实现的完整复制包,以促进进一步的研究。
cs.AI / 2 / 2608.14552

Large Language Models Show Metacognitive Sensitivity in Medical Reasoning

大型语言模型在医学推理中表现出元认知敏感性
Nazzal, Ahmad
Abstract
Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidence tracks evidence quality and uncertainty. We developed a controlled, psychophysics-inspired clinical benchmark to test diagnostic choice and confidence behavior in a medical LLM. The benchmark focused on probable Alzheimer-type neurocognitive disorder (AT-NCD) versus depression-related cognitive impairment (DRCI). We generated 45 synthetic vignettes varying evidence strength, conflicting evidence, and missing information. Each vignette was presented under three prompt variants, yielding 135 trials. In a pilot run with gpt-4.1-nano, all trials produced valid structured outputs. Across forced-choice trials, diagnostic accuracy was 93.5%, mean confidence was 78.4%, and AUROC2 was 0.876. Confidence increased with evidence distance from the diagnostic boundary, decreased when information was missing, and remained higher on correct than incorrect trials after adjustment for evidence strength and prompt format. These findings indicate partial metacognitive sensitivity rather than globally uninformative confidence. However, errors clustered in moderate, conflicting AT-NCD cases, where the model shifted toward DRCI and retained more confidence than empirical accuracy justified. Model comparison suggested that confidence quality should be measured directly rather than inferred from benchmark accuracy or model capability alone. This study establishes a reproducible framework for evaluating evidence sensitivity, metacognitive sensitivity, and localized calibration failure in medical LLMs.
Chinese Translation
大型语言模型(LLMs)在医学领域的评估和应用日益增多,但其临床实用性取决于答案的准确性以及信心是否与证据质量和不确定性相匹配。我们开发了一种受控的、灵感来源于心理物理学的临床基准,以测试医学LLM中的诊断选择和信心行为。该基准聚焦于可能的阿尔茨海默型神经认知障碍(AT-NCD)与抑郁相关的认知障碍(DRCI)。我们生成了45个合成案例,变化了证据强度、相互矛盾的证据和缺失信息。每个案例在三种提示变体下呈现,共产生135个试验。在与gpt-4.1-nano的初步测试中,所有试验均产生有效的结构化输出。在强制选择试验中,诊断准确率为93.5%,平均信心为78.4%,AUROC2为0.876。信心随着证据距离诊断边界的增加而增加,在信息缺失时下降,并且在调整证据强度和提示格式后,正确试验的信心高于错误试验。这些发现表明存在部分元认知敏感性,而非普遍无信息的信心。然而,在中等程度、相互矛盾的AT-NCD案例中,错误聚集,模型倾向于DRCI,并保持了比经验准确性所能证明的更高的信心。模型比较表明,信心质量应直接测量,而非仅仅从基准准确性或模型能力推断。本研究建立了一个可重复的框架,用于评估医学LLM中的证据敏感性、元认知敏感性和局部校准失效。
cs.AI / 3 / 2608.14558

The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning

未书写基准:多模态机器学习在抽象感知推理中的新挑战
Yadav, Garima Arya, Yilmaz, Nilay, Yang, Yezhou
Abstract
Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capacity for abstract perceptual reasoning, inferring unseen information from dynamic, generative processes, remains a critical and underexplored frontier. In this paper, we introduce The Unwritten Benchmark, a new challenge designed to probe this abstract perceptual and cognitive ability. We define the core task as acousto-kinematic word inference: models must decipher words, across 3 different writing styles, being written solely from the audio of pen scratches and the video of hand movements, without any visible ink trace. Our evaluation results reveal a profound gap between human and machine performance: while human participants achieve high ordered letter accuracy (over 80%), leading Multimodal Machine Learning Models, including GPT-4o and Gemini 2.5-Pro, struggle significantly, failing to surpass 10%. Furthermore, we identify a paradoxical fusion effect in the models, where providing both modalities often degrades performance rather than improving it. This finding indicates a fundamental breakdown in their ability to synthesize complementary perceptual cues for this cognitive task. These findings highlight significant limitations in both cross-modal causal reasoning and the understanding of the micro-kinematics essential for such cognitive and intuitive perceptual reasoning.
Chinese Translation
当前的多模态模型在识别静态视觉和听觉内容方面表现出显著的能力。然而,它们在抽象感知推理方面的能力,即从动态生成过程中推断未见信息,仍然是一个关键且未充分探索的前沿领域。在本文中,我们引入了未书写基准(The Unwritten Benchmark),这是一个旨在探讨这种抽象感知和认知能力的新挑战。我们将核心任务定义为声动词推断(acousto-kinematic word inference):模型必须仅通过笔划的音频和手部运动的视频,在没有任何可见墨迹的情况下,解读出三种不同书写风格的单词。我们的评估结果揭示了人类与机器性能之间的显著差距:尽管人类参与者的有序字母准确率超过80%,但包括GPT-4o和Gemini 2.5-Pro在内的领先多模态机器学习模型却显著挣扎,未能超过10%。此外,我们还发现模型中存在一种矛盾的融合效应,即同时提供两种模态往往导致性能下降而非提升。这一发现表明,它们在合成这种认知任务所需的互补感知线索方面存在根本性缺陷。这些发现突显了在跨模态因果推理和理解这种认知及直观感知推理所必需的微运动学方面的重大局限性。
cs.AI / 4 / 2608.14559

When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL

何时进行沟通:多智能体强化学习中的信念分布与KL散度的原则性门控
Kaman, Teoman
Abstract
Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Existing approaches either communicate at every timestep or learn a binary gate through REINFORCE policy gradients \cite{singh2019}, a high-variance signal that produces unstable and uninterpretable gating behavior. I propose a principled alternative: agents communicate only when the KL divergence between their learned belief distributions exceeds a fixed threshold. Each agent maintains a belief distribution over a latent world state computed as a softmax over its LSTM hidden state, and communicates only when belief disagreement is large enough to justify information exchange. I evaluate this approach on the Predator-Prey benchmark from IC3Net \cite{singh2019} across two environment sizes with 5 seeds each, and on MPE simple\_spread \cite{lowe2017}, comparing against IC3Net, CommNet, and an independent controller. On PP 10$\times$10, IC3Net outperforms KL-belief at all thresholds. On the harder PP 20$\times$20, a threshold ablation over $\varepsilon \in \{0.1, 0.3, 0.5, 1.0\}$ reveals an inverted U-shape: $\varepsilon=0.5$ achieves 73.84 average steps and 42\% success rate versus IC3Net's 75.31 steps and 31\%, a gap of 1.47 steps and 11 percentage points with tighter seed variance. On MPE, the belief head improves mean reward by 12 points and reduces variance by 26$\times$ even when gating is inactive, suggesting two orthogonal contributions: principled gating when beliefs can converge, and improved latent representations that benefit coordination regardless.
Chinese Translation
在多智能体强化学习中,有效的沟通要求智能体不仅要决定 extit{沟通什么},还要决定 extit{何时沟通}。现有的方法要么在每个时间步都进行沟通,要么通过REINFORCE策略梯度学习一个二元门控 extcite{singh2019},这种高方差信号会导致不稳定且难以解释的门控行为。我提出了一种原则性的替代方案:智能体仅在其学习的信念分布之间的KL散度超过固定阈值时进行沟通。每个智能体维护一个关于潜在世界状态的信念分布,该分布通过其LSTM隐藏状态的softmax计算得出,只有在信念不一致足够大以证明信息交换的合理性时才进行沟通。我在IC3Net的捕食者-猎物基准上 extcite{singh2019}评估了这种方法,涵盖两个环境规模,每个规模5个种子,并在MPE简单扩散 extcite{lowe2017}上进行比较,比较对象包括IC3Net、CommNet和一个独立控制器。在PP 10$ imes$10上,IC3Net在所有阈值下均优于KL-belief。在更困难的PP 20$ imes$20上,对$ heta extin ext{ablation} ext{over} heta extin ext{0.1, 0.3, 0.5, 1.0}$的阈值消融实验显示出一个倒U型曲线:$ heta=0.5$的平均步数为73.84,成功率为42\%,而IC3Net的步数为75.31,成功率为31\\%,两者之间的差距为1.47步和11个百分点,且种子方差更小。在MPE中,信念头提高了平均奖励12点,并在门控不活跃时将方差降低了26倍,这表明了两个正交的贡献:在信念能够收敛时的原则性门控,以及改善的潜在表示,无论如何都能促进协调。
cs.AI / 5 / 2608.14562

Global AI Regulations for FAIR and Ethics in High-Risk Use Cases: A Comparative Review

高风险应用场景下全球人工智能法规的公平性与伦理性:比较评审
Sharma, Aasish Kumar, Koysev, Dimitar, Anich, Christopher, Ojha, Roshni Kumari, Kunkel, Julian
Abstract
AI governance is shifting from voluntary ethics to enforceable, risk-based regulation, yet cross-jurisdictional divergence creates compliance uncertainty for operators of high-stakes AI. We present a comparative matrix for the EU, US, and China that maps (i) risk classification triggers, (ii) binding obligations, (iii) enforcement and accountability mechanisms, and (iv) the degree to which FAIR principles are operationalised in practice. We stress-test the matrix on three high-impact domains: Electroencephalography (EEG)-guided rehabilitation robotics, AI-enabled debt collection in prospective Central Bank Digital Currency (CBDC) ecosystems, and AI-driven allocation of scarce Graphics Processing Unit (GPU) resources in emerging AI Factory infrastructures. Using primary legal texts and implementation evidence, we identify three recurring gaps: weak interoperability mandates, difficult operationalisation of cross-regime obligations (AI + sector regulation + data protection), and under-specified governance for critical digital infrastructure use cases. To bridge the implementation gap, we outline Knowledge Blocks, a machine-checkable compliance artefact pattern based on Resource Description Framework/Web Ontology Language (RDF/OWL), Shapes Constraint Language (SHACL), and Provenance Ontology (PROV-O), enabling audit-ready compliance-by-design across multiple regimes.
Chinese Translation
人工智能治理正从自愿伦理转向可强制执行的基于风险的监管,但跨法域的差异性为高风险人工智能的运营者带来了合规的不确定性。我们提出了一个比较矩阵,涵盖欧盟、美国和中国,映射了(i)风险分类触发因素,(ii)强制性义务,(iii)执行和问责机制,以及(iv)公平性、可访问性、互操作性和透明性(FAIR)原则在实践中的落实程度。我们在三个高影响领域对该矩阵进行了压力测试:基于脑电图(EEG)的康复机器人、在潜在中央银行数字货币(CBDC)生态系统中启用的人工智能债务催收,以及在新兴人工智能工厂基础设施中驱动稀缺图形处理单元(GPU)资源分配的人工智能。通过使用主要法律文本和实施证据,我们识别出三个反复出现的差距:互操作性要求薄弱、跨制度义务的操作化困难(人工智能 + 行业监管 + 数据保护),以及对关键数字基础设施使用案例的治理规定不明确。为弥补实施差距,我们概述了知识块(Knowledge Blocks),这是一种基于资源描述框架/网络本体语言(RDF/OWL)、形状约束语言(SHACL)和来源本体(PROV-O)的机器可检查合规性工件模式,能够在多个制度中实现审计准备的合规设计。
cs.AI / 6 / 2608.14565

Position: AI Lock-In Is in Progress, and We Must Be Prepared

立场:人工智能锁定正在进行中,我们必须做好准备
Kim, Jaeho, Lee, Seokhyun, Lee, Jieun, Lee, Changhee
Abstract
AI safety research has mainly focused on two areas: technical alignment (ensuring AI systems produce human-aligned outputs) and the regulation of generative AI's societal impacts (including unemployment risk and labor market disruption). However, an equally important dimension remains underexplored: the risk inherent in dependence on AI systems themselves. In this position paper, we argue that AI safety research should address AI Lock-In, the phenomenon whereby excessive reliance on AI systems leads to human deskilling, diminishes human capacity for independent functioning, and creates systemic vulnerabilities when AI systems become unavailable or compromised. We highlight that AI Lock-In is a systemic threat that is already emerging at individual, societal, and national levels, one that could be dramatically amplified by AI service disruptions or geopolitical conflicts. Drawing on detailed scenarios, we investigate how AI Lock-In emerges and escalates across multiple levels, ranging from individual skill atrophy to national-scale infrastructure failures. To address this, we provide guidance on how such risks can be mitigated and prepared for at each level. We contend that proactively addressing AI Lock-In before such dependencies become entrenched, or even irreversible, is essential for preserving individual autonomy and national security.
Chinese Translation
人工智能安全研究主要集中在两个领域:技术对齐(确保人工智能系统产生与人类一致的输出)和生成性人工智能对社会影响的监管(包括失业风险和劳动市场的扰动)。然而,另一个同样重要的维度仍然未得到充分探讨:对人工智能系统本身的依赖所固有的风险。在这篇立场论文中,我们认为人工智能安全研究应关注人工智能锁定(AI Lock-In),即过度依赖人工智能系统导致人类技能退化、降低人类独立运作能力,并在人工智能系统不可用或受到损害时产生系统性脆弱性。我们强调,人工智能锁定是一种系统性威胁,已经在个人、社会和国家层面显现,并可能因人工智能服务中断或地缘政治冲突而被显著放大。通过详细的情境分析,我们探讨了人工智能锁定如何在多个层面上出现和升级,从个人技能退化到国家级基础设施失败。为应对这一问题,我们提供了在各个层面上如何减轻和准备这些风险的指导。我们主张,在这种依赖关系变得根深蒂固甚至不可逆转之前,主动应对人工智能锁定对于维护个人自主权和国家安全至关重要。
cs.AI / 7 / 2608.14566

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

立场:对人工智能道德推理的评估仍然缺乏整体视角
Kierans, Aidan, Dutt, Ritam, Rittichier, Kaley, Dori-Hacohen, Shiri, Ghosh, Avijit
Abstract
Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., whether model outputs align with human moral values. In contrast, the moral norm problem, i.e., whether models can identify and correctly apply context-sensitive moral norms, remains underexplored. We posit that this imbalance stems from the field's reliance on descriptive ethics frameworks, such as Moral Foundations Theory and Kohlberg's stages of moral development, which emphasize value representation over normative application. We review existing benchmarks and evaluation methods, and show that they cluster heavily around the value problem, while discussion regarding normative ethics remains underrepresented. We identify three crucial gaps: (i) the absence of high-quality ground-truth data for moral norms and their applications, (ii) insufficient evaluation of intermediate reasoning processes, and (iii) limited attention to the identification of morally relevant features in context. Subsequently, we propose a research agenda that includes the development of standardized formal representations for normative theories, the construction of expert-annotated datasets capturing norm application, and evaluation protocols that explicitly distinguish between values-level and norms-level competence. Our goal is to encourage a more systematic study of normative reasoning in LLMs.
Chinese Translation
近期关于大型语言模型(LLMs)道德能力评估的研究主要集中在我们所称的道德价值问题,即模型输出是否与人类道德价值观一致。相比之下,道德规范问题,即模型是否能够识别并正确应用情境敏感的道德规范,仍然未得到充分探讨。我们认为,这种不平衡源于该领域对描述性伦理框架的依赖,例如道德基础理论(Moral Foundations Theory)和科尔伯格的道德发展阶段(Kohlberg's stages of moral development),这些框架强调价值表现而非规范应用。我们回顾了现有的基准和评估方法,并显示它们在很大程度上集中于价值问题,而关于规范伦理的讨论则相对不足。我们识别出三个关键缺口:(i)缺乏高质量的道德规范及其应用的真实数据,(ii)对中间推理过程的评估不足,以及(iii)对情境中道德相关特征识别的关注有限。随后,我们提出了一项研究议程,包括为规范理论开发标准化的形式化表示,构建捕捉规范应用的专家注释数据集,以及明确区分价值层面与规范层面能力的评估协议。我们的目标是鼓励对LLMs中的规范推理进行更系统的研究。
cs.AI / 8 / 2608.14567

From Doyle to AGM: A Survey and an Implementation Roadmap for Belief Change

从 Doyle 到 AGM:信念变化的调查与实施路线图
Almeida, Yuri, Casals, Arthur
Abstract
This paper presents a targeted narrative review establishing the historical and theoretical foundations for computational belief change implementation. Seeded by Doyle and London's foundational 1980 taxonomy, we trace the evolution of belief revision from computational origins through the theoretical transformation of the AGM framework to contemporary approaches. Our analysis demonstrates how pre-AGM computational pragmatism relates to AGM theoretical constructs, revealing both continuities and transformations across this evolution. We analyze how each taxonomical category evolved in the post-AGM era, identifying the theoretical foundations and historical precedents that inform contemporary implementation challenges. This foundation enables subsequent research into robust computational blueprints that synthesize historical insights with formal guarantees, providing the baseline for systematic implementation analysis and engineering-focused belief change research.
Chinese Translation
本文呈现了一项针对性的叙述性回顾,建立了计算信念变化实施的历史和理论基础。以 Doyle 和 London 于1980年提出的基础分类法为起点,我们追溯了信念修正的演变,从计算起源到 AGM 框架的理论转变,再到当代方法。我们的分析展示了前 AGM 计算实用主义如何与 AGM 理论构造相关联,揭示了这一演变过程中的连续性和变革。我们分析了每个分类类别在后 AGM 时代的演变,识别出影响当代实施挑战的理论基础和历史先例。这一基础为后续研究提供了坚实的基础,旨在构建稳健的计算蓝图,将历史见解与形式保证相结合,为系统实施分析和以工程为中心的信念变化研究提供基线。
cs.AI / 9 / 2608.14568

Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws

立场:人工智能治理需要类似ISO的互操作性协议,而不仅仅是法律
Wasi, Azmine Toushik, Islam, Mst Rafia, Anik, Mahfuz Ahmed, Rafi, Taki Hasan, Ahsan, Md Manjurul, Chae, Dong-Kyu
Abstract
As Artificial Intelligence (AI) systems become deeply integrated into critical global infrastructure, the urgency for robust governance frameworks has intensified. However, current approaches, led by jurisdiction-specific laws, policies, and voluntary frameworks such as the EU AI Act, China's algorithm governance, and the NIST AI Risk Management Framework in the U.S., create a fragmented regulatory landscape. In this position paper, we argue that \textbf{\textit{AI governance must be built not on laws alone, but on ISO-like interoperability protocols that enable standardized, machine-readable risk communication across borders}}. Drawing on the success of the GDPR, which was operationalized through standards like ISO 27001 and Privacy by Design, we propose the development of standardized AI \textit{nutrition labels} containing unified metrics for bias, energy usage, and data provenance to facilitate cross-jurisdictional compliance. These manifests would lower barriers for small and medium enterprises (SMEs), reduce redundant regulatory efforts, and build public trust. The paper addresses concerns that standards may stifle innovation by advocating for modular, versioned protocols designed to evolve in tandem with technological change. Overall, we call for a shift from siloed legal compliance toward interoperable technical conformance, enabling a shared global language for responsible AI deployment.
Chinese Translation
随着人工智能(AI)系统深度融入全球关键基础设施,对强有力治理框架的需求愈加迫切。然而,目前的治理方法主要由特定司法管辖区的法律、政策以及自愿框架(如欧盟AI法案、中国的算法治理和美国的NIST AI风险管理框架)主导,导致了碎片化的监管格局。在这篇立场论文中,我们认为 extbf{ extit{人工智能治理必须不仅仅建立在法律之上,而是基于类似ISO的互操作性协议,以实现跨境标准化、机器可读的风险沟通}}。借鉴《通用数据保护条例》(GDPR)的成功经验,该条例通过ISO 27001和隐私设计等标准得以实施,我们建议开发标准化的AI extit{营养标签},包含统一的偏见、能耗和数据来源等指标,以促进跨司法管辖区的合规。这些清单将降低中小企业(SMEs)的门槛,减少冗余的监管工作,并建立公众信任。论文还针对标准可能抑制创新的担忧,倡导设计模块化、版本化的协议,以便与技术变革同步发展。总体而言,我们呼吁从孤立的法律合规转向互操作的技术一致性,促进负责任的人工智能部署的全球共享语言。
cs.AI / 10 / 2608.14569

Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integration

立场:神经约束推理中的认证正确性需要符号积分
Kong, Shufeng, Zhang, Xiaochuan, Liu, Caihua
Abstract
Neural solvers for constraint satisfaction problems have achieved remarkable in-distribution accuracy, yet they suffer from a fundamental limitation persistent constraint violations occur under distribution shifts even when the model reports high confidence. This position paper argues that when hard constraints exist and the cost of verification is relatively low, neural constraint reasoning must prioritize symbolic integration over pure learning. We justify our focus on Sudoku as a representative NP-complete testbed because it exhibits a sharp asymmetry between easy verification and hard solving: checking a candidate solution requires only polynomial time $O(n^{2})$, while finding a solution may require exponential search. Through a comprehensive survey of solving methods spanning deterministic algorithms, metaheuristic optimization, learning-based approaches, and language-conditioned reasoning, we demonstrate that neural-only methods without instance-level certification fail to achieve the provable correctness that symbolic and neuro-symbolic approaches provide. We advocate for a bidirectional integration in which neural methods enhance symbolic solvers by learning heuristics and converting percepts into symbols, while symbolic methods verify neural outputs to ensure their reliability. To operationalize this position, we propose a multi-agent certified reasoning framework that demonstrates how this integration can achieve both computational efficiency and provable correctness.
Chinese Translation
神经求解器在约束满足问题上取得了显著的分布内准确性,但它们面临着一个根本性限制,即在分布转移下持续出现约束违反,即使模型报告了高置信度。本文立场论文认为,当存在硬约束且验证成本相对较低时,神经约束推理必须优先考虑符号积分而非纯学习。我们选择数独作为代表性的 NP 完全测试平台,理由是它在易验证和难求解之间表现出明显的不对称性:检查候选解仅需多项式时间 $O(n^{2})$,而找到解可能需要指数搜索。通过对包括确定性算法、元启发式优化、基于学习的方法和语言条件推理在内的求解方法进行全面调查,我们证明了没有实例级认证的纯神经方法无法达到符号和神经符号方法所提供的可证明正确性。我们倡导一种双向集成,其中神经方法通过学习启发式和将感知转化为符号来增强符号求解器,而符号方法则验证神经输出以确保其可靠性。为了实现这一立场,我们提出了一个多智能体认证推理框架,展示了这种集成如何实现计算效率和可证明正确性的双重目标。
cs.AI / 11 / 2608.14571

Position: Want Better ML Reviews? Stop Asking Nicely and Start Incentivizing with a Credit System

立场:想要更好的机器学习评审?停止礼貌请求,开始通过信用系统激励
Zhong, Shaochen
Abstract
With soaring submission counts, stricter reciprocal review policies, widespread adoption of platforms like OpenReview, and without the offsetting pressure of publication fees, the machine learning (ML) community has one of the largest scholarly presences among all scientific fields. And yet, \textbf{almost \textit{everyone} has \textit{many} unpleasant things to share about their review experience.} Worse, there is little public space to seriously discuss, let alone debate, what makes a review system effective or how it might be improved.\quad In this position paper, we expand our discussion from two core problems: \textit{How can we reasonably limit submission volume?} and \textit{How can we incentivize good and discourage bad reviewing?} We first assess the strengths and shortcomings of existing attempts to address such problems. Specifically, we present four takes on some popular conference mechanisms and propose two alternative designs for improvement.\quad Our general position is that meaningful improvement in ML peer review won't come from polite best-practice suggestions tucked into Calls for Papers or Reviewer Guidelines: it requires \textbf{enforceable yet fine-grained procedural safeguards} paired with \textbf{a currency-like credit system (e.g., our proposed \textit{OpenReview Points})}. ML practitioners can ``earn'' such points by contributing good review practices, and ``spend'' them across one or multiple major conferences to redeem different kinds of ``perks,'' such as complimentary registration or the right to request additional review resources.
Chinese Translation
随着提交数量激增、互评政策日益严格、像 OpenReview 这样的平台被广泛采用,以及缺乏发表费用的抵消压力,机器学习(ML)社区在所有科学领域中拥有最大的学术存在之一。然而, extbf{几乎 extit{每个人}都有 extit{许多}关于其评审经历的不愉快经历要分享。} 更糟的是,几乎没有公共空间来认真讨论,更不用说辩论,什么样的评审系统是有效的,或者如何改进它。 extbf{在这篇立场论文中,我们将讨论扩展到两个核心问题: extit{我们如何合理限制提交量?}以及 extit{我们如何激励良好评审并抑制不良评审?}} 我们首先评估现有尝试解决这些问题的优缺点。具体而言,我们提出了四种对一些流行会议机制的看法,并提出了两种改进的替代设计。 extbf{我们的总体立场是,机器学习同行评审的有意义改进不会来自于礼貌的最佳实践建议,这些建议被藏在征稿启事或评审指南中:这需要 extbf{可执行且细致的程序性保障},以及 extbf{类似货币的信用系统(例如,我们提出的 extit{OpenReview Points})}。机器学习从业者可以通过贡献良好的评审实践来“赚取”这些积分,并在一个或多个主要会议上“消费”这些积分,以兑换不同种类的“福利”,例如免费注册或请求额外评审资源的权利。
cs.AI / 12 / 2608.14578

Longitudinal and Graph-Augmented Prediction of Adolescent Substance Use Onset in the ABCD Study

ABCD研究中青少年物质使用开始的纵向与图增强预测
He, Yixuan, Su, Jinni, Kang, Yun
Abstract
Early identification of adolescent substance-use risk is an important prevention challenge, yet the relative value of baseline characteristics, longitudinal trajectories, and relational context remains unclear. Using data from approximately 11,860 participants in the Adolescent Brain Cognitive Development (ABCD) Study, we compare cross-sectional, longitudinal, and graph-based approaches for predicting alcohol sipping, alcohol use, marijuana use, and alcohol/marijuana use. We evaluate tree-based models, recurrent neural networks, and Temporal Graph Convolutional Networks (T-GCNs) constructed from family, school, and feature-similarity graphs. Longitudinal models consistently outperform baseline models, with temporal XGBoost achieving the strongest standalone performance. Although T-GCNs generally do not surpass temporal XGBoost, graph-derived risk scores provide complementary information. Combining temporal XGBoost and T-GCN predictions through score-level stacking yields the best performance across all outcomes, achieving AUC-ROC values above 0.79. Feature analyses identify peer deviance, age, externalizing symptoms, parental monitoring, cultural norms, and neighborhood context as important predictors of substance use onset. These findings demonstrate the value of longitudinal modeling for substance-use prediction and suggest that graph-based representations can provide effective auxiliary risk signals.
Chinese Translation
早期识别青少年物质使用风险是一个重要的预防挑战,但基线特征、纵向轨迹和关系背景的相对价值仍不明确。利用来自约11,860名参与者的青少年大脑认知发展(ABCD)研究的数据,我们比较了预测饮酒、酒精使用、大麻使用以及酒精/大麻使用的横断面、纵向和基于图的方法。我们评估了基于树的模型、递归神经网络和从家庭、学校及特征相似性图构建的时间图卷积网络(Temporal Graph Convolutional Networks, T-GCNs)。纵向模型始终优于基线模型,其中时间XGBoost实现了最强的独立性能。尽管T-GCNs通常未能超越时间XGBoost,但图衍生的风险评分提供了补充信息。通过评分级别堆叠结合时间XGBoost和T-GCN的预测在所有结果中实现了最佳性能,AUC-ROC值超过0.79。特征分析确定了同伴偏差、年龄、外化症状、父母监控、文化规范和邻里背景作为物质使用开始的重要预测因素。这些发现展示了纵向建模在物质使用预测中的价值,并表明基于图的表示可以提供有效的辅助风险信号。
cs.AI / 13 / 2608.14579

SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization

SKILL:自我修正的知识引导迭代大型语言模型代理用于逻辑优化
Yang, Rui
Abstract
Logic synthesis optimization poses significant challenges due to exponentially growing search spaces, sparse reward signals, and diverse logic structures. Traditional expert-designed flows lack adaptability, while reinforcement learning (RL) methods often suffer from low sample efficiency and limited interpretability. We introduce SKILL, a Self-correcting Knowledge-guided Iterative Large Language Model Agent that unifies multi-agent LLM reasoning and RL-based environment interaction for automated synthesis optimization. SKILL coordinates three specialized LLMs: GPT-4o for strategic planning, Claude Sonnet 4 for detailed reasoning, and Gemini 2.5 Pro for efficient analysis with a PPO-based RL agent that learns actionable policies through direct interaction with synthesis tools. A novel self-correcting module monitors environment feedback (PDA metrics), detects suboptimal behaviors, and invokes LLM-guided recovery strategies. Evaluations on IWLS, OpenCores, and EPFL benchmarks show SKILL achieves a 12.4 % PDA improvement over expert flows and 86.3% success rate on logic systems up to 500K gates.
Chinese Translation
逻辑综合优化面临着显著的挑战,这些挑战源于指数级增长的搜索空间、稀疏的奖励信号以及多样的逻辑结构。传统的专家设计流程缺乏适应性,而强化学习(RL)方法通常存在样本效率低和可解释性有限的问题。我们提出了SKILL,一种自我修正的知识引导迭代大型语言模型代理,它统一了多智能体LLM推理和基于RL的环境交互,以实现自动化综合优化。SKILL协调了三个专业的LLM:用于战略规划的GPT-4o、用于详细推理的Claude Sonnet 4,以及用于高效分析的Gemini 2.5 Pro,结合一个基于PPO的RL代理,通过与综合工具的直接交互学习可操作的策略。一个新颖的自我修正模块监控环境反馈(PDA指标),检测次优行为,并调用LLM引导的恢复策略。在IWLS、OpenCores和EPFL基准测试上的评估表明,SKILL在专家流程上实现了12.4%的PDA改进,并在高达50万门的逻辑系统上达到了86.3%的成功率。
cs.AI / 14 / 2608.14580

OGX: An Open-Source, Vendor-Neutral Generative AI Application Server

OGX:一个开源的、供应商中立的生成式人工智能应用服务器
Arceo, Francisco Javier, Han, Sébastien, Farrellee, Matthew, Doern, Charlie, Tang, Yuan, Higgins, Derek, Narsing, Varsha Prasad, Sim, Gordon, Kamenani, Sumanth, Browning, Ben, Murthy, Raghotham
Abstract
OGX (Open GenAI Stack) is an open-source AI application server and Python library that implements the APIs of major frontier labs (OpenAI, Anthropic, Google) with pluggable backend providers. Developers building agentic AI applications--such as retrieval-augmented generation pipelines, multi-turn agents, and tool-calling workflows--can develop against a single API surface and deploy with any combination of inference engine, vector database, and safety backend, without changing application code. OGX's primary focus is the Responses API for server-side agentic orchestration, conforming to the Open Responses specification. The server also supports the Anthropic Messages API and Google GenAI Interactions API, decoupling SDK choice from model and deployment decisions. With over 20 inference providers, 13 vector store backends, and a companion Kubernetes Operator for production deployment, OGX serves as the self-hosted, model-agnostic backend for AI-powered developer tools including Claude Code, Codex CLI, OpenCode, and OpenHands. The project has over 8,400 GitHub stars, 242 contributors, and 4,000 commits across nearly two years of public development.
Chinese Translation
OGX(开放生成式人工智能堆栈)是一个开源的人工智能应用服务器和Python库,旨在实现主要前沿实验室(OpenAI、Anthropic、Google)的API,并支持可插拔的后端提供者。开发者在构建自主智能应用时——例如检索增强生成管道、多轮代理和工具调用工作流——可以在单一的API接口上进行开发,并以任意组合的推理引擎、向量数据库和安全后端进行部署,而无需更改应用代码。OGX的主要关注点是用于服务器端自主协调的响应API,符合开放响应规范。该服务器还支持Anthropic消息API和Google生成式人工智能交互API,使SDK选择与模型和部署决策解耦。OGX提供超过20个推理提供者、13个向量存储后端,并配备一个用于生产部署的Kubernetes操作器,作为AI驱动开发工具(包括Claude Code、Codex CLI、OpenCode和OpenHands)的自托管、模型无关后端。该项目在近两年的公开开发中获得了超过8400个GitHub星标、242名贡献者和4000次提交。
cs.AI / 15 / 2608.14585

Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry

Euclid-Omni:一个统一的神经符号框架用于平面几何
Li, Zhaoyu, Bi, Hangrui, Zhang, Youyuan, Ma, Wenjie, Li, Zenan, Zhang, Zhaolei, Si, Xujie, Yang, Kaiyu
Abstract
Euclidean geometry is a compelling testbed for AI reasoning, as it demands the combination of intuitive diagram understanding, axiomatic deduction, and algebraic computation. Yet, existing approaches typically address only a subset of these abilities or struggle with competition-level problems. We introduce \textit{Euclid-Omni}, a unified neuro-symbolic framework that couples a formal geometry system with Large Language Models (LLMs) and Vision-Language Models (VLMs) to tackle both calculation- and proving-style problems, in formal and natural languages, up to Olympiad-level difficulty. At its core, we develop \textit{Euclidea}, a versatile symbolic geometry solver that automatically generates reasoning steps through deductive inference and algebraic computation. Building on this, we develop a data-generation pipeline that synthesizes symbolic problems and solutions, renders diagrams, and translates them into natural language, producing large-scale, diverse datasets for training LLMs and VLMs across a wide range of reasoning settings. Experiments show that VLMs trained on our synthetic data achieve superior performance on calculation tasks, and that LLMs combined with \textit{Euclidea} are competitive with state-of-the-art systems on Olympiad-level proving problems, despite using orders of magnitude less compute and training data. Code and scripts are publicly available at https://github.com/20171130/Euclid-Omni
Chinese Translation
欧几里得几何是人工智能推理的一个引人注目的测试平台,因为它要求结合直观的图形理解、公理推导和代数计算。然而,现有的方法通常只解决这些能力的一个子集,或者在竞争级别的问题上表现不佳。我们介绍了 extit{Euclid-Omni},一个统一的神经符号框架,它将一个形式几何系统与大型语言模型(LLMs)和视觉语言模型(VLMs)结合起来,以应对形式和自然语言中的计算和证明风格问题,难度高达奥林匹克级别。在其核心,我们开发了 extit{Euclidea},一个多功能的符号几何求解器,能够通过推理推导和代数计算自动生成推理步骤。在此基础上,我们开发了一个数据生成管道,合成符号问题和解决方案,渲染图形,并将其翻译成自然语言,从而生成大规模、多样化的数据集,用于在广泛的推理环境中训练LLMs和VLMs。实验表明,在我们的合成数据上训练的VLMs在计算任务上表现优越,而结合 extit{Euclidea}的LLMs在奥林匹克级别的证明问题上与最先进的系统具有竞争力,尽管使用的计算和训练数据量少得多。代码和脚本可在https://github.com/20171130/Euclid-Omni公开获取。
cs.AI / 16 / 2608.14587

An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts: A Plant Science Use Case

基于规则和大型语言模型的代理框架用于嵌入和注释描述性文档布局:植物科学案例研究
Turenne, Nicolas, Sklab, Youcef, Chenin, Eric, Zucker, Jean-Daniel
Abstract
Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and specialized retrieval models to improve ranking accuracy, relevance, and cross-lingual performance. Complementary techniques such as passage indexing, document layout analysis, and semantic knowledge representation further enhance retrieval effectiveness by capturing fine-grained contextual and structural information. Emerging agentic LLM frameworks extend these capabilities by enabling planning, iterative reasoning, tool use, and multi-agent collaboration, thereby broadening applications across diverse domains. These frameworks also emphasize rigorous evaluation, ethical considerations, and trustworthiness, ensuring responsible deployment in real-world settings. We propose a modular, agent-based pipeline for botanical trait extraction. Optical character recognition (OCR) converts PDFs into machine-readable text, while segmentation and indexing organize content by genus and species. Rule-based parsers extract structured botanical traits, and ensembles of large language models (LLMs) expand trait vocabularies and resolve ambiguities. This approach ensures accurate species recognition, scalable annotation, and explainable integration of textual botanical descriptions, enabling robust and interpretable data extraction across large botanical corpora. Results: Using three regional botanical datasets, our system extracted 55,737 trait annotations across 4,961 species, averaging 9.1 traits per species. Integration of LLM-based enrichment improved coverage for 75% of traits, increasing total annotations by 59%. While the choice of OCR engine had a minor effect on species recognition, overall annotation counts remained stable, demonstrating the robustness, scalability, and reliability of the pipeline for large-scale botanical trait extraction.
Chinese Translation
背景:近年来,信息检索(IR)的进展利用了密集和稀疏表示、大型语言模型(LLMs)以及专门的检索模型,以提高排名准确性、相关性和跨语言性能。补充技术如段落索引、文档布局分析和语义知识表示,通过捕捉细粒度的上下文和结构信息,进一步增强了检索的有效性。新兴的代理型LLM框架通过支持规划、迭代推理、工具使用和多代理协作,扩展了这些能力,从而拓宽了在不同领域的应用。这些框架还强调严格的评估、伦理考量和可信度,确保在现实环境中的负责任部署。我们提出了一种模块化的基于代理的植物性状提取管道。光学字符识别(OCR)将PDF转换为机器可读文本,而分割和索引则按属和种组织内容。基于规则的解析器提取结构化的植物性状,而大型语言模型(LLMs)的集成则扩展了性状词汇并解决了歧义。该方法确保了准确的物种识别、可扩展的注释和可解释的文本植物描述集成,从而在大规模植物语料库中实现稳健且可解释的数据提取。结果:使用三个区域植物数据集,我们的系统提取了55,737个性状注释,涵盖4,961个物种,平均每个物种9.1个性状。基于LLM的增强集成提高了75%性状的覆盖率,总注释量增加了59%。虽然OCR引擎的选择对物种识别有轻微影响,但整体注释数量保持稳定,证明了该管道在大规模植物性状提取中的稳健性、可扩展性和可靠性。
cs.AI / 17 / 2608.14588

The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines

幻觉雪球:将误差传播建模为多智能体大型语言模型管道中的状态转移
Singh, Prabhjot, Pawar, Bhushan
Abstract
Sequential multi-agent LLM pipelines chain specialized agents without verification at handoffs, creating a structural flaw with measurable and severe consequences. We show that hallucinations injected at Stage 1 do not merely persist; they transform: raw numerical facts become derived computations, then narrative prose, then editorially approved conclusions. At each transformation, detectability degrades near-irreversibly. We formalize this as the hallucination snowball effect, a first-order Markov process over four states (Raw Fact $\to$ Derived $\to$ Narrative $\to$ Invisible) with empirically measured per-boundary escape probabilities of 24.6%, 48.3%, and 89.3%. Across 346 automatically injected hallucinations in a 4-agent financial analysis pipeline on FinanceBench, gpt-4o detection drops from 72.0% at Stage 1 to 50.9% at Stage 4, and 23.7% of hallucinations survive completely undetected in the final output. Even the strongest model tested (Qwen3.5-397B-A17B, 87.0% at Stage 1) faces a structural ceiling; projected Stage 4 detection is only ${\sim}$60--65%. Critically, boundary gates using identical RAG verification tools reduce hallucination survival from 58.4% to 16.2% versus end-of-pipeline checking (Cohen's $h = -0.911$, $p < 0.000001$), while end-checking alone achieves merely 2.3 pp improvement over no verification. When you verify matters more than whether you verify. Our model predicts survival for $n$-agent linear pipelines and prescribes optimal verification resource allocation: invest at $S_1{\to}S_2$ first, where 75.4% of hallucinations are still catchable, not at $S_3{\to}S_4$ where 89.3% have already escaped.
Chinese Translation
顺序多智能体大型语言模型管道在交接时链式连接专门化代理,但未进行验证,造成了结构性缺陷,具有可测量且严重的后果。我们展示了在第一阶段注入的幻觉不仅仅是持续存在;它们会转化:原始数值事实变为衍生计算,然后是叙述性散文,最后是经过编辑批准的结论。在每次转化中,检测能力几乎不可逆地下降。我们将其形式化为幻觉雪球效应,这是一种在四个状态(原始事实 $ o$ 衍生 $ o$ 叙述 $ o$ 隐形)上的一阶马尔可夫过程,边界逃逸概率的经验测量为24.6%、48.3%和89.3%。在FinanceBench上的一个4智能体金融分析管道中,346个自动注入的幻觉中,gpt-4o的检测率从第一阶段的72.0%下降到第四阶段的50.9%,并且23.7%的幻觉在最终输出中完全未被检测到。即使是测试中最强的模型(Qwen3.5-397B-A17B,第一阶段为87.0%)也面临结构性上限;预计第四阶段的检测率仅为${ ilde{60}}$--65%。关键是,使用相同RAG验证工具的边界门将幻觉存活率从58.4%降低到16.2%,与管道末端检查相比(Cohen's $h = -0.911$, $p < 0.000001$),而仅进行末端检查的改进仅为2.3个百分点。当你验证的时机比你是否验证更为重要。我们的模型预测了$n$-智能体线性管道的存活率,并建议了最佳验证资源分配:首先在$S_1{ o}S_2$进行投资,此时75.4%的幻觉仍然可以被捕捉,而不是在$S_3{ o}S_4$,此时89.3%已经逃脱。
cs.AI / 18 / 2608.14590

Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement

迈向安全的LLM代理:规范、验证与执行的调查
Dantas, Pierre, Cordeiro, Lucas, Nowroozi, Ehsan, Norbert, Tihanyi
Abstract
LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools. However, no existing system provides formally grounded, task-level safety guarantees for the plans these agents generate. Research remains fragmented across specification, verification, and enforcement, limiting understanding of the strengths and limitations of existing approaches. To address this gap, we conducted a PRISMA 2020 systematic review of 38 studies published between 2022 and 2026 and retrieved from six academic databases. Our analysis reveals four key findings. First, the specification bottleneck remains the primary challenge: natural-language-to-formal translation achieves only 24% to 35% semantic correctness, undermining downstream verification. Second, runtime monitoring is the most mature enforcement strategy, reducing unsafe actions by 40% to 65% in controlled settings, but it does not provide complete safety guarantees. Third, the verifier tax shows that blocking 94% of unsafe actions can still result in less than 5% safe task completion because agents exploit alternative unsafe paths. Finally, no existing approach simultaneously achieves soundness, scalability, semantic correctness, and task-level safety preservation. We contribute a three-level taxonomy, a comparative analysis of existing techniques, a synthesis of evidence on the verifier tax, and a ten-problem research agenda for trustworthy agentic AI.
Chinese Translation
LLM代理越来越多地执行不可逆的现实世界操作,包括数据库更新、API调用、文件操作和工具的自主使用。然而,现有系统并未为这些代理生成的计划提供形式上可靠的任务级安全保证。研究在规范、验证和执行方面仍然分散,限制了对现有方法的优缺点的理解。为了解决这一空白,我们对2022年至2026年间在六个学术数据库中发布的38项研究进行了PRISMA 2020系统评审。我们的分析揭示了四个关键发现。首先,规范瓶颈仍然是主要挑战:自然语言到形式的翻译仅实现了24%到35%的语义正确性,削弱了下游验证。其次,运行时监控是最成熟的执行策略,在受控环境中将不安全行为减少了40%到65%,但并未提供完整的安全保证。第三,验证者税表明,阻止94%的不安全行为仍可能导致不到5%的安全任务完成,因为代理会利用替代的不安全路径。最后,现有方法没有同时实现健全性、可扩展性、语义正确性和任务级安全保护。我们贡献了一个三级分类法、现有技术的比较分析、关于验证者税的证据综合,以及一个针对可信代理人工智能的十个问题研究议程。
cs.AI / 19 / 2608.14598

Position: Medical AI Neglects Real Treatment Outcomes

立场:医疗人工智能忽视真实治疗结果
Kaul, Shiva, Khurshid, Anjum
Abstract
Medical AI has rapidly improved its ability to perform diagnostic and prognostic tasks that lead to treatment decisions. But understanding of treatment itself is still inadequately trained and evaluated, using human opinions and syntheses (especially texts such as biomedical publications and clinical practice guidelines) rather than actual underlying data on treatment outcomes. This neglect seriously limits the potential of medical AI, and is already causing deficiencies in both frontier models and major benchmarks, as argued in this position paper. Real treatment outcomes, drawn from sources such as observational databases and randomized experiments, should be substantially incorporated into both training and evaluation. Improving these outcomes should be reemphasized as the downstream goal of all medical AI.
Chinese Translation
医疗人工智能在执行诊断和预后任务方面的能力迅速提升,这些任务直接影响治疗决策。然而,对治疗本身的理解仍然训练不足且评估不够,主要依赖于人类的观点和综合(尤其是生物医学出版物和临床实践指南等文本),而非实际的治疗结果数据。这种忽视严重限制了医疗人工智能的潜力,并且已经在前沿模型和主要基准中造成了缺陷,正如本文所论述的那样。真实的治疗结果应从观察性数据库和随机实验等来源中提取,并应在训练和评估中得到实质性纳入。改善这些结果应重新强调为所有医疗人工智能的下游目标。
cs.AI / 20 / 2608.14610

When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning

大型语言模型何时适用错误法律?诊断大型语言模型在时间法律推理中的失败
Huang, Yiqian, Zheng, Shuyuan, Liu, Qianying, Peng, Shaowen, Kong, Yuntao, Funakoshi, Kotaro, Xiao, Chuan, Okumura, Manabu, Cao, Yang
Abstract
Legal reasoning tasks such as legal judgment prediction (LJP) require identifying the temporally correct version of the law governing a case -- a capability we term temporal applicable-law determination. However, whether large language models (LLMs) can reliably perform this task remains unexplored. In this paper, we construct a benchmark to evaluate LLMs on temporal applicable-law determination, and systematically investigate why they fail at temporal legal reasoning. Our experiments reveal four key findings. First, LLMs exhibit a strong bias toward applying the most recently enacted law, regardless of when the legally relevant facts occurred. Second, this bias does not stem from an inability to understand that laws have temporal scope, nor from a lack of knowledge about historical statutes. Third, we provide behavioral evidence that reinforcement-learning-shaped explicit reasoning may be a key mechanism: while improving general reasoning ability, it reduces the diversity of reasoning paths, causing models to converge on applying the current law. Fourth, this produces a counterintuitive inverse relationship: models with stronger general reasoning ability tend to perform worse on temporal legal reasoning. Our findings offer concrete guidance for future work on improving LLM performance in temporally grounded legal reasoning.
Chinese Translation
法律推理任务,如法律判断预测(LJP),需要识别适用于案件的时间上正确的法律版本——我们称之为时间适用法律的确定。然而,大型语言模型(LLMs)是否能够可靠地执行此任务尚未得到探索。在本文中,我们构建了一个基准来评估LLMs在时间适用法律确定方面的表现,并系统地调查它们在时间法律推理中失败的原因。我们的实验揭示了四个关键发现。首先,LLMs表现出强烈的偏向于适用最近颁布的法律,而不考虑法律相关事实发生的时间。其次,这种偏见并不是由于无法理解法律具有时间范围,或缺乏对历史法规的知识。第三,我们提供了行为证据表明,强化学习塑造的显性推理可能是一个关键机制:虽然提高了整体推理能力,但却减少了推理路径的多样性,导致模型趋向于适用当前法律。第四,这产生了一个反直觉的逆向关系:具有更强整体推理能力的模型在时间法律推理中的表现往往更差。我们的发现为未来提高LLM在时间基础法律推理中的表现提供了具体指导。
cs.AI / 21 / 2608.14613

Do LLM Agents Negotiate Rationally? A Mechanism-Design Framework for Verifiable Multi-Agent Interaction over A2A/MCP

大型语言模型代理是否理性谈判?基于机制设计的可验证多代理交互框架在A2A/MCP上的应用
Albayaydh, Wael, Zhao, Rui
Abstract
Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool access and Google's Agent2Agent (A2A) protocol for agent delegation and negotiation. However, these protocols specify transport and discovery rather than strategic correctness and do not guarantee efficient, individually rational, or strategy-proof outcomes. We introduce a framework that (i) encodes classical negotiation mechanisms, including alternating-offers bargaining and Vickrey-Clarke-Groves-style auctions, as constraints over A2A message schemas; (ii) provides a lightweight runtime verification and repair layer that checks messages against protocol invariants; and (iii) offers a benchmark of negotiation and allocation tasks with known optimal solutions for measuring deviations from game-theoretic predictions. We evaluate multiple LLM backbones using unstructured dialogue, structured protocols, and structured protocols with verification. Across negotiation trials (N=30 per condition), verification reduces outcome variance, while structured protocols achieve 100 percent success for both models. After correcting parser artifacts, audited unstructured baselines achieve approximately 97 percent and 93.3 percent success. In auction experiments (N=30 per model), both models achieve 100 percent efficient allocation but differ sharply in truthful bidding: one bids its exact valuation in every trial, whereas the other does so in only 3.3 percent of trials. Thus, mechanism-level incentive compatibility does not automatically transfer to LLM-agent behavior. A three-party fair-allocation task produced only 4.2 percent usable outcomes; we report this negative result with a diagnosis. This work bridges classical multi-agent systems theory and modern LLM-agent infrastructure and defines verifiable interaction at the A2A protocol layer.
Chinese Translation
现代大型语言模型(LLM)代理框架越来越多地通过诸如Anthropic的模型上下文协议(MCP)用于代理与工具的访问,以及谷歌的代理间协议(A2A)用于代理委托和谈判等标准进行互操作。然而,这些协议主要规定了传输和发现,而非战略正确性,且并不能保证高效、个体理性或策略无关的结果。我们提出了一个框架,(i) 将经典谈判机制(包括交替报价谈判和Vickrey-Clarke-Groves风格的拍卖)编码为A2A消息模式的约束;(ii) 提供一个轻量级的运行时验证和修复层,用于检查消息是否符合协议不变性;(iii) 提供一个具有已知最优解的谈判和分配任务基准,以测量与博弈论预测的偏差。我们使用非结构化对话、结构化协议以及带验证的结构化协议评估了多个LLM骨干。在谈判试验中(每种条件N=30),验证减少了结果方差,而结构化协议在两个模型中均实现了100%的成功率。在纠正了解析器伪影后,审计的非结构化基线成功率约为97%和93.3%。在拍卖实验中(每个模型N=30),两个模型均实现了100%的有效分配,但在真实出价方面差异显著:一个模型在每次试验中都出价其确切估值,而另一个模型仅在3.3%的试验中如此。因此,机制级别的激励相容性并不自动转移到LLM代理的行为上。在一个三方公平分配任务中,仅产生了4.2%的可用结果;我们报告了这一负面结果并进行了诊断。本研究架起了经典多代理系统理论与现代LLM代理基础设施之间的桥梁,并在A2A协议层定义了可验证的交互。
cs.AI / 22 / 2608.14615

Large Language Models and their Awareness of Mechanics and Spatial Geometry

大型语言模型及其对机械和空间几何的认知
Gerstmayr, Johannes, Weyrer, Sebastian, Möltner, Tobias, Manzl, Peter, Pieber, Michael
Abstract
Large Language Models (LLMs) perform well on established code-generation and mathematical-reasoning benchmarks, but their capabilities in mechanics and spatial geometry, here denoted as mechanical engineering awareness, has not been quantified systematically. We present MecEng, a fully automated benchmark that evaluates LLMs on the creation of multibody simulation models from parameterized textual descriptions. The benchmark comprises 84 generic tasks on three difficulty levels, ranging from rigid-body systems with joints and contact to flexible multibody systems that require exact 3D geometry generation, tetrahedral finite-element meshing, and Hurty-Craig-Bampton model order reduction of machine parts. A dedicated pipeline with LLMs generates simulation-ready geometry from text using Netgen, and builds multibody system models for the code Exudyn, which are then verified against expert ground truth on several levels: system-graph isomorphism including graph node annotations, numerical solutions, and part-specific measures such as mass, geometry, and eigenfrequencies. In total, 32 open-weight and two proprietary LLMs are evaluated. On rigid-body tasks, the best open-weight model obtains an overall success rate of 86.0%, compared to 91.4% for the strongest proprietary model, while flexible multibody tasks remain considerably harder. Additional studies quantify the influence of sampling temperature, reasoning, prompt design, model size, and LLM-release date. The results indicate rapidly improving, but still error-prone, mechanical engineering awareness of current LLMs.
Chinese Translation
大型语言模型(LLMs)在既定的代码生成和数学推理基准测试中表现良好,但它们在机械和空间几何方面的能力,即机械工程意识,尚未得到系统量化。我们提出了MecEng,这是一个完全自动化的基准,评估LLMs从参数化文本描述中创建多体仿真模型的能力。该基准包含84个通用任务,分为三个难度级别,涵盖从具有关节和接触的刚体系统到需要精确3D几何生成、四面体有限元网格划分以及Hurty-Craig-Bampton机器部件模型降阶的柔性多体系统。一个专门的管道利用LLMs通过Netgen从文本生成仿真准备好的几何形状,并为代码Exudyn构建多体系统模型,这些模型在多个层面上与专家的真实数据进行验证:系统图同构性,包括图节点注释、数值解以及特定部件的测量,如质量、几何形状和特征频率。总共评估了32个开放权重模型和两个专有模型。在刚体任务中,最佳开放权重模型的整体成功率为86.0%,而最强的专有模型为91.4%,而柔性多体任务则明显更具挑战性。额外研究量化了采样温度、推理、提示设计、模型大小和LLM发布日期的影响。结果表明,当前LLMs的机械工程意识正在快速改善,但仍然存在错误。
cs.AI / 23 / 2608.14622

A Human-Centred Approach to Benchmarking LLMs for Parenting Advice

以人为本的父母建议大语言模型基准评估方法
Zhao, Yunke, Voysey, Isobel, van Heerden, Alastair, Hughes, Rob, Zhao, Jun
Abstract
People are increasingly using large language models (LLMs) to seek advice, including for parenting. Parenting is a critical and socially sensitive domain. Thus, evaluating advice provided by LLMs requires indicators beyond aggregated information quality benchmarks to consider relational and behavioural elements of the responses. With a multi-dimensional rubric created by parenting experts, this paper evaluates 15 LLMs across 100 parenting scenarios in 2 languages (English and Chinese), using an LLM-as-a-judge method. Results show that aggregate scores can hide rubric item-specific weaknesses, models implicitly encourage different parenting styles, and language influences responses. We highlight the importance of evaluation output auditability and challenges involved in evaluating LLM-generated advice in domains like parenting. Our findings provide important insights for selecting LLMs for direct user engagement and the development of user-facing parenting advice applications.
Chinese Translation
人们越来越多地使用大语言模型(LLMs)寻求建议,包括育儿方面的建议。育儿是一个关键且社会敏感的领域。因此,评估LLMs提供的建议需要超越汇总信息质量基准的指标,以考虑响应的关系和行为元素。本文通过育儿专家创建的多维评分标准,采用LLM作为评审者的方法,评估了15个LLMs在100个育儿场景中的表现,涵盖两种语言(英语和中文)。结果表明,汇总分数可能掩盖评分项目特定的弱点,模型隐含地鼓励不同的育儿风格,并且语言会影响响应。我们强调了评估输出可审计性的重要性以及在育儿等领域评估LLM生成建议所面临的挑战。我们的研究结果为选择LLMs以直接与用户互动以及开发面向用户的育儿建议应用程序提供了重要的见解。
cs.AI / 24 / 2608.14624

Learning Agent Execution for KV-Cache Management in Agentic Serving

代理服务中KV缓存管理的学习代理执行
Zhang, Rui, Kim, Chaeeun, Feng, Shaoting, Du, Kuntai, Liu, Yuhan, Zhong, Yi, Ching, Cheng-Wei, Jiang, Junchen, Hu, Liting
Abstract
Multi-agent LLM systems have emerged as an important deployment paradigm for AI services, where each user request is decomposed into a sequence of specialized agents. Across these workflows, every agent repeatedly executes a fixed context consisting of system prompts, tool definitions, and few-shot examples, creating substantial opportunities for KV-cache reuse. Existing LLM serving systems, however, manage KV-cache reactively using prefix caching and recency-based replacement, causing reusable agent contexts to be evicted before their next invocation and forcing repeated recomputation. We present CacheScout, an agent-aware KV-cache runtime layer for multi-agent LLM serving. The key insight is that future KV-cache reuse is governed by agent execution semantics rather than cache recency alone. CacheScout captures these semantics by learning agent execution transitions online, without requiring predefined workflow graphs or offline training, and uses the learned execution model to guide both cache eviction and proactive prefetching while leaving the serving critical path unchanged. We implement CacheScout on top of vLLM. Across representative real-world multi-agent workloads, CacheScout improves KV-cache hit rate by 10-18 percentage points, reduces mean TTFT by 18-45%, lowers mean per-turn latency by 29-38%, and increases peak throughput by up to 57%. These benefits also generalize to larger models, reducing TTFT by up to 54% while sustaining 37% higher throughput.
Chinese Translation
多代理大语言模型(LLM)系统已成为人工智能服务的重要部署范式,其中每个用户请求被分解为一系列专门的代理。在这些工作流程中,每个代理反复执行由系统提示、工具定义和少量示例组成的固定上下文,从而创造了大量的KV缓存重用机会。然而,现有的LLM服务系统通过前缀缓存和基于最近性的替换来被动管理KV缓存,导致可重用的代理上下文在下次调用之前被驱逐,迫使重复计算。我们提出了CacheScout,一个面向多代理LLM服务的代理感知KV缓存运行时层。关键的见解是,未来的KV缓存重用受代理执行语义的支配,而不仅仅是缓存的最近性。CacheScout通过在线学习代理执行转移来捕捉这些语义,无需预定义的工作流程图或离线训练,并利用学习到的执行模型来指导缓存驱逐和主动预取,同时保持服务的关键路径不变。我们在vLLM之上实现了CacheScout。在代表性的真实世界多代理工作负载中,CacheScout将KV缓存命中率提高了10-18个百分点,减少了平均TTFT(总时间到完成)18-45%,降低了平均每轮延迟29-38%,并将峰值吞吐量提高了最多57%。这些好处也适用于更大的模型,将TTFT降低了最多54%,同时保持了37%的更高吞吐量。
cs.AI / 25 / 2608.14631

Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study

大型语言模型在化妆品化学和皮肤健康中的准确性与可靠性:基准研究
Liu, Amelia
Abstract
As consumers increasingly turn to AI chatbots for skincare advice, the technical accuracy of Large Language Models (LLMs) in cosmetic chemistry remains largely under-evaluated. We benchmarked 14 LLMs on a structured set of topics related to cosmetic chemistry, including the chemical properties of specific cosmetic ingredients and common cosmetic scenarios that may be of interest to consumers. Web search was disabled throughout to assess each model's internalized knowledge rather than its internet retrieval capacity. Overall performance was poor, with the most pronounced deficits in quantitative reasoning and structural identification tasks. While models handled general skincare questions with reasonability, responses consistently lacked the technical depth required for informed consumer decision-making. Notably, conversation with AI can pose a risk: outputs that sound authoritative but contain technical errors are less likely to generate skepticism compared to responses that explicitly acknowledge uncertainty. These findings suggest that general-purpose LLMs, trained predominantly on unverified public data, are currently not reliable sources of cosmetic chemistry information. Progress on two fronts, fine-tuning verified chemical and dermatological datasets, and substantial improvements to algorithmic reasoning, will likely be needed before these tools can be considered as resources for public use.
Chinese Translation
随着消费者越来越多地依赖人工智能聊天机器人获取护肤建议,大型语言模型(LLMs)在化妆品化学领域的技术准确性仍然未得到充分评估。我们对14个LLMs进行了基于结构化主题的基准测试,主题涉及化妆品化学,包括特定化妆成分的化学性质以及消费者可能感兴趣的常见化妆场景。在整个测试过程中禁用了网络搜索,以评估每个模型内化的知识,而非其互联网检索能力。总体表现较差,尤其在定量推理和结构识别任务中表现出明显的不足。尽管模型在处理一般护肤问题时表现合理,但其回答始终缺乏进行明智消费者决策所需的技术深度。值得注意的是,与人工智能的对话可能存在风险:听起来权威但包含技术错误的输出相比于明确承认不确定性的回答,更不容易引发怀疑。这些发现表明,主要基于未经验证的公共数据训练的一般用途LLMs,目前并不是可靠的化妆品化学信息来源。在这方面的进展,特别是对经过验证的化学和皮肤病学数据集的微调,以及算法推理的实质性改进,可能是这些工具在被视为公共使用资源之前所需的。
cs.AI / 26 / 2608.14641

Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks

任务与会话级模型路由:四个开源路由器在四个基准上的共同接口混合评估
Kumar, Kiran N., Saminathan, Santhosh K.
Abstract
Agentic systems increasingly delegate model selection to a router, yet open-source routers are usually evaluated with different tasks, candidate pools, and execution protocols, limiting direct comparison. We present a common measurement protocol and hybrid evaluation of four router implementations across RouterBench, BFCL v4, tau2-bench, and WebArena. We evaluate 290 frozen tasks against a locked matrix of 2,610 candidate outcomes. Three routers emit constant or near-constant tier assignments; only vLLM Semantic Router varies materially with prompt content, and it has the highest observed success rate on none of the four benchmarks. Always-Mid matches Aurelio exactly on three benchmarks and within 0.003 on the fourth. For vLLM, task-level superiority tests detect no task-specific advantage over a share-matched content-blind allocation; equivalence is established only on WebArena at the protocol-declared five-percentage-point margin. The results show that, under these configurations and controls, observed gains track selected-tier composition more closely than demonstrated task-specific targeting. Fixed-tier baselines and selected-tier distributions are therefore necessary controls in router evaluation; the findings are scoped to these configurations, candidate pool, and frozen benchmark samples, not to routing paradigms in general.
Chinese Translation
自主系统越来越多地将模型选择委托给路由器,但开源路由器通常在不同的任务、候选池和执行协议下进行评估,这限制了直接比较。我们提出了一种共同的测量协议和对四个路由器实现的混合评估,涵盖了RouterBench、BFCL v4、tau2-bench和WebArena。我们对290个冻结任务进行了评估,针对一个锁定的2610个候选结果矩阵。三种路由器发出恒定或近乎恒定的层级分配;只有vLLM语义路由器在提示内容上有显著变化,并且在四个基准中没有一个观察到的成功率是最高的。Always-Mid在三个基准上与Aurelio完全匹配,在第四个基准上相差0.003。对于vLLM,任务级优越性测试未能检测到相对于共享匹配的内容盲分配的任务特定优势;等价性仅在WebArena上以协议声明的五个百分点的边际确立。结果表明,在这些配置和控制下,观察到的增益与选定层级的组成更紧密相关,而不是展示的任务特定目标。因此,固定层级基线和选定层级分布在路由器评估中是必要的控制;这些发现的范围限于这些配置、候选池和冻结基准样本,而不适用于一般的路由范式。
cs.AI / 27 / 2608.14651

Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance

评估多模态大型语言模型在文本和音频模式下的可及性灾害援助
Gupta, Anuridhi, Mansoor, Samara, Purohit, Hemant
Abstract
Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the needs of individuals with access and functional needs, including hard-of-hearing individuals, pregnant women, mothers with toddlers, and elderly individuals with dementia. Recent advancements in Artificial Intelligence (AI), especially Multi-Modal Large Language Models (MM-LLMs), demonstrate powerful capabilities to serve diverse users across text, audio, image, and video modalities within a single unified system, such as a chatbot. However, their suitability for deployment rests on a property that receives limited scrutiny, i.e., whether these systems produce consistent, actionable outputs regardless of the modality through which a user communicates. In this paper, we conduct a comprehensive analysis to understand the status of open-weight MM-LLMs using real emergency alert scenarios across four different vulnerable personas. These state-of-the-art (SOTA) models are evaluated on consistency of responses across text and audio modalities when the same task scenario is given. Findings indicate that no model achieves reliable consistency across modalities, and that performance gaps are heightened for personas with access needs, introducing modality-dependent inequity that undermines the humanitarian value of these systems. These results inform concrete design recommendations for building equitable, trustworthy, and inclusive AI tools for disaster risk communication.
Chinese Translation
有效的灾害风险沟通是一个基础的人道主义挑战,但当前的应急基础设施未能满足有特殊访问和功能需求的个体的需求,包括听力障碍者、孕妇、带幼儿的母亲以及患有痴呆症的老年人。近期在人工智能(AI)领域的进展,特别是多模态大型语言模型(MM-LLMs),展示了在一个统一系统(如聊天机器人)中服务于文本、音频、图像和视频模式下多样化用户的强大能力。然而,它们的适用性取决于一个受到有限关注的属性,即这些系统是否能够在用户通过何种模式进行沟通时,产生一致且可操作的输出。在本文中,我们进行了一项全面分析,以了解开放权重的多模态大型语言模型在四种不同脆弱角色下使用真实的紧急警报场景的状态。这些最先进(SOTA)模型在相同任务场景下评估文本和音频模式的响应一致性。研究结果表明,没有任何模型在不同模式下实现可靠的一致性,并且对于有特殊访问需求的角色,性能差距更加显著,导致了依赖模式的不平等,削弱了这些系统的人道主义价值。这些结果为构建公平、可信和包容的AI工具以进行灾害风险沟通提供了具体的设计建议。
cs.AI / 28 / 2608.14659

When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation

当不确定性不足时:代码生成中的自我修正实证研究
Rakasi, Pranav, Lalwani, Maanas, Srivastava, Arnav, Palanivel, Arya, Adeleke, Tinuade, Li, Ruizhe, Wu, Sean
Abstract
Large language models for code generation often produce incorrect solutions without reliable indicators of failure. We study whether uncertainty estimation methods developed for natural language transfer to code generation, and whether such signals can improve code generation via selective self-correction. We evaluate five uncertainty methods: mean token entropy, verbalized confidence, $P(\text{True})$, entropy ensembles, and semantic entropy probes, across three small code LLMs on HumanEval and BigCodeBench. We find that multi-sample $P(\text{True})$ achieves the strongest correlation with correctness, while all the other methods, including semantic entropy probes, yield only weak correlation. We then use these uncertainty signals to drive three self-correction policies: adaptive decoding, uncertainty-based regeneration, and verification-based regeneration. Our results reveal a stronger negative finding than anticipated: uncertainty-based self-correction fails to reliably improve Pass@1, degrading accuracy in 5 of 6 configurations across both benchmarks ($-3$pp to $-10$pp), and adaptive decoding degrades accuracy in 4 of 6 configurations. Only verification-based self-correction reliably improves Pass@1, with gains of $+6$ to $+26$ percentage points on HumanEval and $+8$ to $+20$ percentage points on BigCodeBench, scaling inversely with baseline strength. These findings replicate consistently across both benchmarks and suggest that cheap uncertainty estimators are insufficient on their own to improve code correctness, and that their practical value lies in serving as gating signals for costlier execution-based correction loops rather than as standalone substitutes for verification.
Chinese Translation
用于代码生成的大型语言模型经常产生不正确的解决方案,而没有可靠的失败指示器。我们研究了为自然语言开发的不确定性估计方法是否可以转移到代码生成中,以及这些信号是否可以通过选择性自我修正来改善代码生成。我们在 HumanEval 和 BigCodeBench 上评估了五种不确定性方法:平均标记熵、口头化置信度、$P( ext{True})$、熵集成和语义熵探针,针对三种小型代码 LLM。我们发现,多样本 $P( ext{True})$ 与正确性之间的相关性最强,而其他所有方法,包括语义熵探针,仅产生弱相关性。然后,我们利用这些不确定性信号驱动三种自我修正策略:自适应解码、不确定性驱动的再生成和基于验证的再生成。我们的结果揭示了一个比预期更强的负面发现:基于不确定性的自我修正未能可靠地改善 Pass@1,在两个基准测试中有 6 种配置中的 5 种准确性下降(下降幅度为 $-3$pp 到 $-10$pp),而自适应解码在 6 种配置中的 4 种中也降低了准确性。只有基于验证的自我修正可靠地改善了 Pass@1,在 HumanEval 上的增益为 $+6$ 到 $+26$ 个百分点,在 BigCodeBench 上的增益为 $+8$ 到 $+20$ 个百分点,且增益与基线强度呈反比。这些发现在两个基准测试中一致复制,表明廉价的不确定性估计器单独不足以改善代码的正确性,其实际价值在于作为更昂贵的基于执行的修正循环的门控信号,而不是作为验证的独立替代品。
cs.AI / 29 / 2608.14666

Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring

通过因果机制监测实现跨域工业故障检测
Neupane, Dhiraj, Bouadjenek, Mohamed Reda, Dazeley, Richard, Aryal, Sunil
Abstract
Unsupervised fault detection in industrial systems is dominated by reconstruction based methods that monitor individual sensor marginal distributions. This misses coupling faults, where the physical relationship between sensor groups breaks while marginal statistics remain normal. Such faults evade marginal monitoring and persist as latent failures, with direct consequences for system reliability and safety. We propose CMR-Mamba (Causal Mechanism Representation Mamba), which trains per domain Mamba state-space encoders on healthy data. A causal cross-modal predictor regularises these encoders so that the effect-channel manifold reflects the normal cause-to-effect coupling. Anomalies are scored by k-nearest-neighbour (kNN) distance on this manifold or by the mechanism residual between the observed and the causally predicted effect embedding. We evaluate CMR-Mamba on electromechanical (Paderborn bearings), hydraulic (ZeMA) and cyber-physical (SWaT) coupling-fault domains. Ablations establish two findings. First, k-NN manifold scoring, rather than the encoder family, is the dominant source of gain over reconstruction-error scoring, improving baselines by up to 0.42 AUROC and exceeding the gain from causal regularisation. Second, aggregate AUROC is saturated by easy faults that any strong method solves, so the methods separate only on the low-separability subset. There CMR-Mamba leads the evaluated baselines on Paderborn artificial defects and on SWaT stealthy attacks, which keep every sensor inside its normal range and which marginal methods detect only at chance. CMR-Mamba therefore offers an interpretable and consistently competitive approach to coupling-fault detection across mechanical, hydraulic and cyber-physical systems. Code and data are available at https://anonymous.4open.science/status/CMR_Mamba_MFD_1177.
Chinese Translation
工业系统中的无监督故障检测主要依赖于基于重构的方法,这些方法监测单个传感器的边际分布。这种方法忽视了耦合故障,即传感器组之间的物理关系破裂,而边际统计仍然正常。这类故障逃避了边际监测,并作为潜在故障持续存在,直接影响系统的可靠性和安全性。我们提出了CMR-Mamba(因果机制表示Mamba),该方法在健康数据上训练每个领域的Mamba状态空间编码器。一个因果跨模态预测器对这些编码器进行正则化,使得效应通道流形反映正常的因果关系耦合。通过在该流形上的k近邻(kNN)距离或观察到的效应嵌入与因果预测的效应嵌入之间的机制残差对异常进行评分。我们在电机(帕德博恩轴承)、液压(ZeMA)和网络物理(SWaT)耦合故障领域评估了CMR-Mamba。消融实验得出了两个发现。首先,k-NN流形评分,而非编码器家族,是相较于重构误差评分的主要增益来源,基线提高了多达0.42 AUROC,超出了因果正则化带来的增益。其次,整体AUROC被任何强大方法都能解决的简单故障所饱和,因此这些方法仅在低可分离性子集上区分。在此,CMR-Mamba在帕德博恩人工缺陷和SWaT隐蔽攻击的评估基线中领先,后者使每个传感器保持在正常范围内,而边际方法仅以偶然的机会检测到。CMR-Mamba因此提供了一种可解释且始终具有竞争力的耦合故障检测方法,适用于机械、液压和网络物理系统。代码和数据可在 https://anonymous.4open.science/status/CMR_Mamba_MFD_1177 获取。
cs.AI / 30 / 2608.14667

Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems

定位:科学团队中的人工智能代理应作为人机系统进行研究
Emami, Patrick, Horawalavithana, Sameera, Nguyen, Truc, Panapitiya, Gihan, Jacob, Bruno, Raskar, Siddhisanket, Sinha, Saumya, Willard, Jared D., Glaws, Andrew, Somasekharan, Nithin, Yue, Ling, Lu, Brian, Pan, Shaowu, Eisner, Jason
Abstract
Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet most current work focuses on the autonomous capabilities of "AI Scientists". We argue that this overlooks the social aspects of scientific teamwork, and that studying AI Scientists as human-agent systems (HAS)--where the unit of analysis is the human-agent pair--is both underexplored and undervalued. We establish these points through literature and empirical analysis, and highlight recent incidences and studies which show that deploying agents in science without accounting for human-agent dynamics introduces near-term risks, including reduced diversity of scientific inquiry. Through analysis of real-world case studies, we show that scientists and agents can augment each other's capabilities. We call for new research that adopts the HAS lens to develop mathematical frameworks for understanding and fostering human-AI synergy in scientific discovery.
Chinese Translation
基于大型语言模型的代理越来越多地被作为科学发现中的合作者部署,但目前大多数研究集中于“人工智能科学家”的自主能力。我们认为,这忽视了科学团队合作中的社会因素,并且将人工智能科学家视为人机系统(Human-Agent Systems, HAS)进行研究——其中分析单位是人机对——是一个尚未充分探索和重视的领域。我们通过文献和实证分析来建立这些观点,并强调最近的事件和研究表明,在科学中部署代理而不考虑人机动态会带来短期风险,包括科学探究的多样性降低。通过对现实案例的分析,我们展示了科学家与代理可以相互增强能力。我们呼吁开展新的研究,采用HAS视角来开发数学框架,以理解和促进科学发现中的人机协同作用。
cs.AI / 31 / 2608.14669

Beyond Correctness: Toward Automated Novelty Verification with Lean 4

超越正确性:基于 Lean 4 的自动化新颖性验证
Porto, Ayrton
Abstract
Artificial intelligence systems applied to mathematics verify correctness but not novelty: an automatically generated theorem can compile in Lean without errors and yet be an already known result. This article presents AViD Journal, a pipeline that receives a LaTeX article, formalizes its statements in Lean 4, and issues a novelty verdict through a decision tree over three dimensions: prior existence in a formal corpus (Mathlib) and an informal one (TheoremSearch and Matlas, with temporal filter and LLM judge), non-triviality via automatic tactics, and structural distance between proofs measured as Jaccard distance over premise sets. Evaluation on papers withdrawn from arXiv due to declared duplication produced a result more informative than any performance measure: the identification of three obstacles that limit the approach regardless of this implementation. First, successful compilation of a Lean file does not guarantee semantic fidelity. Second, the recall ceiling is imposed by the coverage of theorem indices, not by the similarity metric. Third, arXiv removes the source code of articles upon withdrawal, compromising the reproducibility of any benchmark built upon them.
Chinese Translation
应用于数学的人工智能系统验证正确性但不验证新颖性:一个自动生成的定理可以在 Lean 中无错误地编译,但仍然可能是一个已知的结果。本文介绍了 AViD Journal,这是一个接收 LaTeX 文章的管道,将其陈述形式化为 Lean 4,并通过决策树在三个维度上给出新颖性判定:在正式语料库(Mathlib)和非正式语料库(TheoremSearch 和 Matlas,带有时间过滤和 LLM 判断)中的先前存在性,通过自动策略评估非平凡性,以及通过 Jaccard 距离测量的证明间结构距离。对因声明重复而从 arXiv 撤回的论文进行的评估产生了比任何性能指标更具信息量的结果:识别出三个限制该方法的障碍,无论该实现如何。首先,成功编译 Lean 文件并不保证语义的忠实性。其次,召回上限受定理索引覆盖的限制,而不是相似性度量。第三,arXiv 在撤回文章时会删除源代码,从而妨碍基于这些文章构建的任何基准的可重复性。
cs.AI / 32 / 2608.14673

Auditing an AI-Generated Mathematical Proof: A Correction to a Greedy Conditioning Lemma in Quantum Parallel Repetition

审计一个AI生成的数学证明:对量子并行重复中的贪婪条件引理的修正
Sienicki, Mikołaj, Sienicki, Krzysztof
Abstract
Chapter 6 of OpenAI's *Ten Advances in Mathematics and Theoretical Computer Science* claims an exponential parallel-repetition theorem for all finite two-player, one-round entangled games. Early in the proof, the chapter uses a quantitative greedy conditioning lemma. The lemma is meant to select a small set of coordinates (D) such that, after conditioning on winning every coordinate in (D), a randomly chosen remaining coordinate is won with average probability at least (1-\delta). The statement is correct, but the proof as printed contains a polarity error. Its continuation test is written in terms of average success, while the next step requires a coordinate with a large conditional failure probability. That implication is false, and even simple examples can leave the printed procedure without a valid next move. This note gives an explicit counterexample, identifies the intended continuation condition, and supplies a complete corrected proof. The repair is local: it leaves the statement of the lemma and the parameters used later in the chapter unchanged. It should not, however, be read as an independent verification of the main parallel-repetition theorem. More broadly, the example shows how a mathematically plausible AI-generated argument can hide a small but decisive reversal between complementary events.
Chinese Translation
OpenAI的《数学与理论计算机科学的十项进展》第六章声称,对于所有有限的两人一轮纠缠游戏,存在一个指数级的并行重复定理。在证明的早期,该章节使用了一个定量的贪婪条件引理。该引理旨在选择一个小的坐标集(D),使得在对每个坐标(D)的胜利进行条件化后,随机选择的剩余坐标以至少(1-B4)的平均概率获胜。该陈述是正确的,但印刷的证明中存在极性错误。其延续测试是以平均成功为基础书写的,而下一步则需要一个具有较大条件失败概率的坐标。该推论是错误的,甚至简单的例子也可能使印刷的程序无法进行有效的下一步移动。本文提供了一个明确的反例,识别了预期的延续条件,并提供了完整的修正证明。修正是局部的:它保持了引理的陈述和章节后面使用的参数不变。然而,这不应被视为对主要并行重复定理的独立验证。更广泛地说,这个例子展示了一个在数学上看似合理的AI生成论证如何隐藏在互补事件之间的小而决定性的逆转。
cs.AI / 33 / 2608.14680

When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry

当自主执行失败时:从遥测数据中检测和定位运行时故障
Zhang, Chenkai, Li, Yiran, Tian, Yifang, Bachras, Michalis, Jacobsen, Hans-Arno
Abstract
Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, yet evaluating only task outcomes reveals little about how or why a run fails. We present AGENTCHAOSBENCH, a benchmark for detecting and localizing runtime faults in agentic systems from their execution telemetry. We run five heterogeneous applications that coordinate agents over the Agent-to-Agent protocol and call tools through the Model Context Protocol, and inject ten types of operational fault (unavailable or slow tools, corrupted or oversized responses, and delayed, looped, or misrouted delegations and bypassed guardrails) at their tool, model, guardrail, and inter-agent boundaries, alongside a no-fault control. The resulting dataset contains 275 sanitized traces: 250 faulty executions spanning ten fault types and 25 no-fault controls. Each faulty trace is aligned with the no-fault execution of the same input; fault-type labels and, where applicable, location labels are held out from diagnosis. On structured single-trace inputs, a first set of zero-shot LLM baselines shows the task is far from solved: local detectors up to 14B parameters reach only 13.6-19.2% top-1 fault-type accuracy and the frontier DeepSeek-v4-pro only 24.8%, while jointly identifying the fault type and its location tops out at 22%; reference-dependent faults (above all a bypassed guardrail) stay near-unsolved from a single trace. An aligned reference improves selected relative faults but does not resolve guardrail bypass. The held-out labels and compact prediction format support reproducible comparison of LLM-based and non-LLM diagnosis methods.
Chinese Translation
基于大语言模型(LLM)的自主系统的可靠性是整个执行过程的属性(包括工具调用、模型调用、保护措施和代理间消息),而不仅仅是最终答案的属性。然而,仅仅评估任务结果对于了解运行失败的原因和方式几乎没有帮助。我们提出了AGENTCHAOSBENCH,这是一个用于从自主系统的执行遥测中检测和定位运行时故障的基准。我们运行了五个异构应用程序,这些程序通过代理间协议协调代理,并通过模型上下文协议调用工具,并在其工具、模型、保护措施和代理间边界注入了十种类型的操作故障(不可用或缓慢的工具、损坏或超大响应、延迟、循环或错误路由的委托以及绕过的保护措施),同时设置了一个无故障控制组。生成的数据集包含275个经过清理的跟踪记录:250个故障执行,涵盖十种故障类型,以及25个无故障控制。每个故障跟踪记录与相同输入的无故障执行对齐;故障类型标签和(如适用)位置标签在诊断中被保留。在结构化的单跟踪输入上,一组零样本LLM基线显示该任务远未解决:参数高达14B的局部检测器仅达到13.6%-19.2%的顶级故障类型准确率,而前沿的DeepSeek-v4-pro仅为24.8%,同时识别故障类型及其位置的准确率最高为22%;参考依赖故障(尤其是绕过的保护措施)在单个跟踪中几乎无法解决。对齐的参考改善了选定的相对故障,但并未解决保护措施绕过的问题。保留的标签和紧凑的预测格式支持LLM和非LLM诊断方法的可重复比较。
cs.AI / 34 / 2608.14694

A Comprehensive Survey of Wireless Foundation Models for AI-Native 6G Networks

面向AI原生6G网络的无线基础模型综合调查
Khan, Naveed, Sbeihi, Besan Al, Alshehhi, Maryam, Saeed, Nasir
Abstract
Foundation models are emerging as a transformative paradigm for AI-native sixth-generation (6G) wireless networks by enabling scalable, transferable, and data-efficient intelligence across diverse communication tasks. Unlike conventional deep learning models that are trained for individual applications, wireless foundation models (WFMs) learn generalized representations from large-scale heterogeneous wireless data and can be efficiently adapted to communication, sensing, localization, and network optimization tasks with minimal task-specific supervision. Despite rapid progress, current research remains fragmented across architectures, training paradigms, and application domains, with no unified survey dedicated to the design, learning, and deployment of WFMs. This survey presents a comprehensive and unified review of wireless foundation models. We first establish the fundamental concepts of WFMs and introduce a taxonomy that organizes the field according to model architectures, pre-training paradigms, and applications. We then review representative architectures, self-supervised pre-training strategies, parameter-efficient adaptation methods, datasets, benchmarks, and evaluation methodologies, highlighting their roles in enabling transferable wireless intelligence. Furthermore, we examine emerging applications spanning physical-layer signal processing, network intelligence, and cross-layer optimization, and discuss the key challenges of data availability, generalization, interpretability, efficient edge deployment, and standardization. Finally, we outline future research directions toward scalable, trustworthy, and general-purpose wireless intelligence for AI-native 6G networks. This survey provides a comprehensive reference for researchers and practitioners developing next-generation intelligent wireless systems.
Chinese Translation
基础模型正在成为面向AI原生第六代(6G)无线网络的变革性范式,通过在多样化的通信任务中实现可扩展、可转移和数据高效的智能。与为单个应用训练的传统深度学习模型不同,无线基础模型(WFMs)从大规模异构无线数据中学习通用表示,并能够在最小的任务特定监督下高效适应通信、感知、定位和网络优化任务。尽管取得了快速进展,但当前的研究在架构、训练范式和应用领域上仍然存在碎片化,尚无统一的调查专门针对WFMs的设计、学习和部署。本调查提供了无线基础模型的全面和统一的回顾。我们首先建立WFMs的基本概念,并引入一个分类法,根据模型架构、预训练范式和应用对该领域进行组织。然后,我们回顾了代表性的架构、自监督预训练策略、参数高效适应方法、数据集、基准测试和评估方法,强调它们在实现可转移无线智能中的作用。此外,我们考察了涵盖物理层信号处理、网络智能和跨层优化的新兴应用,并讨论了数据可用性、泛化能力、可解释性、高效边缘部署和标准化等关键挑战。最后,我们概述了面向AI原生6G网络可扩展、可信赖和通用无线智能的未来研究方向。本调查为开发下一代智能无线系统的研究人员和从业者提供了全面的参考。
cs.AI / 35 / 2608.14697

Synchronized Logit Steering: Real-world Steganography

同步逻辑引导:现实世界的隐写术
Rufail, Andrew, Dash, Aadi, Narahari, Onir, Mui, Ethan, Gajare, Mahi, Tiwari, Prakhar, Makapothula, Shrija, Cui, Nick
Abstract
Steganography in large language models offers a way to embed hidden messages within natural-sounding text. Existing token and logit-level methods typically require the sender and receiver to share an identical prompt context, which is rarely guaranteed in production pipelines that use retrieval-augmented generation or proprietary system instructions. We introduce Synchronized Logit Steering (SLS), a deterministic steganographic scheme that eliminates this dependency by deriving a proxy prompt from the generated output itself, allowing both parties to reconstruct the same logit distribution without access to the original prompt. SLS encodes payload values as token ranks within high-entropy regions of the proxy prompt distribution, and we extend the scheme with periodic recurrence and payload bursts to scale information density. Across ShareGPT, GSM8K, and SWE-bench Verified, we show that the KL divergence between the true and proxy prompt distributions falls below 0.5 nats once the synchronization window reaches 40 tokens, and SLS encoding does not meaningfully disrupt this convergence relative to greedy generation. We also find that the periodic-burst variant achieves 0.20 bits per token, or roughly 10x the capacity of single-payload encoding. Kolmogorov-Smirnov tests further confirm that SLS outputs are statistically difficult to distinguish from greedy generations, demonstrating that covert, prompt-agnostic communication through LLMs is both practical and stealthy.
Chinese Translation
在大型语言模型中,隐写术提供了一种在自然语言文本中嵌入隐藏信息的方法。现有的基于标记和逻辑值的方法通常要求发送者和接收者共享相同的提示上下文,而在使用检索增强生成或专有系统指令的生产管道中,这种情况很少得到保证。我们提出了同步逻辑引导(Synchronized Logit Steering, SLS),这是一种确定性的隐写方案,通过从生成的输出中推导出代理提示,消除了这种依赖,使得双方能够在没有访问原始提示的情况下重建相同的逻辑分布。SLS将有效载荷值编码为代理提示分布中高熵区域的标记排名,并通过周期性重复和有效载荷突发扩展该方案,以提高信息密度。在ShareGPT、GSM8K和SWE-bench Verified上,我们显示当同步窗口达到40个标记时,真实和代理提示分布之间的KL散度降至0.5 nats以下,并且SLS编码相对于贪婪生成并未显著干扰这一收敛。我们还发现,周期突发变体实现了每个标记0.20比特的容量,约为单一有效载荷编码的10倍。Kolmogorov-Smirnov检验进一步确认,SLS输出在统计上难以与贪婪生成区分,证明通过大型语言模型进行隐蔽、与提示无关的通信既实用又隐秘。
cs.AI / 36 / 2608.14707

Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems

基于语义不确定性的层次多智能体系统编排
Knowlton, John, Guha, Aritra, Miikkulainen, Risto
Abstract
As large language model (LLM)-based multi-agent systems become increasingly capable, coordinating agents under uncertainty becomes a fundamental challenge. Existing orchestration strategies typically rely on fixed interaction patterns and often lack mechanisms for assessing the reliability of intermediate reasoning steps, allowing errors and hallucinations to propagate through the system. This paper introduces a semantic-uncertainty-guided orchestration approach, HASSUM as a general framework for uncertainty-aware coordination in multi-agent systems. The method estimates uncertainty using semantic entropy and semantic density, which measure trust at the level of answer semantics rather than output probabilities. These signals enable adaptive orchestration decisions, including output verification, selective reprompting, additional deliberation, and confidence-aware response selection. Because the approach operates independently of any particular agent architecture, it can be integrated into a broad range of hierarchical and collaborative multi-agent systems. The evaluations demonstrate an implementation within a hierarchical agent framework and evaluate it on StrategyQA, JailbreakBench, and TruthfulQA benchmarks. Across tasks that require complex reasoning and are prone to ambiguity or hallucinations, uncertainty-guided orchestration yields more reliable outcomes than uncertainty-unaware coordination. Semantic entropy and semantic density in tandem outperformed either metric alone. Ablations testing different thresholds and model sizes demonstrated that both influence the effectiveness of semantic metrics. The results suggest that semantic uncertainty is a practical and general-purpose signal for improving robustness and trustworthiness in agentic AI systems.
Chinese Translation
随着基于大型语言模型(LLM)的多智能体系统能力的不断提升,在不确定性下协调智能体成为一项基本挑战。现有的编排策略通常依赖于固定的交互模式,且往往缺乏评估中间推理步骤可靠性的机制,导致错误和幻觉在系统中传播。本文提出了一种基于语义不确定性的编排方法HASSUM,作为多智能体系统中不确定性感知协调的通用框架。该方法通过语义熵和语义密度来估计不确定性,这些指标在答案语义层面上测量信任,而非输出概率。这些信号使得自适应编排决策成为可能,包括输出验证、选择性重新提示、额外的深思熟虑以及基于置信度的响应选择。由于该方法独立于任何特定的智能体架构,因此可以集成到广泛的层次和协作多智能体系统中。评估展示了在层次智能体框架内的实现,并在StrategyQA、JailbreakBench和TruthfulQA基准上进行了评估。在需要复杂推理且容易产生歧义或幻觉的任务中,基于不确定性的编排比不考虑不确定性的协调产生了更可靠的结果。语义熵和语义密度的结合优于单独使用任一指标。对不同阈值和模型规模的消融测试表明,这两者都影响语义指标的有效性。结果表明,语义不确定性是提高智能体AI系统鲁棒性和可信度的实用且通用的信号。
cs.AI / 37 / 2608.14711

Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation

超越 Pass@k:测量代理代码生成的可靠性和安全性
Jiang, Jiajun, Zheng, Sharon, Vidra, Natan, Setty, Spurthi
Abstract
AI coding agent benchmarks rank agents with the Chen et al. (2021) pass@k estimator, but current implementations misapply it: they set n to the number of unit tests in a single submission rather than the number of independent rollout attempts, conflating test-suite size with attempt independence. We diagnose this operationalization error, prove it by counterexample, and propose reliability@k, the same estimator applied correctly, with n = independent rollouts and c = fully-passing rollouts per (task, agent) pair. In a synthetic multi-rollout benchmark, the misapplied metric inflates reported scores by 0.85-0.97 in absolute terms (0.96-0.98 reported vs. 0.00-0.12 corrected), and a cheap single-rollout proxy fails to substitute for repeated runs (Spearman $\rho = 0.417$). Motivated by evidence that functional correctness does not imply security safety, we additionally propose security-adjusted reliability@k, which counts only rollouts that are both functionally correct and free of high-severity insecure patterns. In an initial live-API test with three agents, the adjustment did not change any ranking under our current scanner and threshold, so we present it as a proposed complementary lens whose decisive evaluation requires better-powered future runs. Finally, a preliminary 5-task SWE-bench Verified pilot observes the same core concern in a real repository setting: macro-averaged hidden-test pass rate was 0.80 while strict task resolution was 0.20.
Chinese Translation
AI 编码代理基准使用 Chen 等人(2021)的 pass@k 估计器对代理进行排名,但当前实现错误地应用了该估计器:它们将 n 设置为单次提交中的单元测试数量,而不是独立的尝试次数,从而将测试套件大小与尝试独立性混淆。我们诊断了这一操作化错误,通过反例证明了这一点,并提出了 reliability@k,这是正确应用的相同估计器,其中 n = 独立尝试次数,c = 每对(任务,代理)完全通过的尝试次数。在一个合成的多次尝试基准测试中,错误应用的指标使报告的分数在绝对值上膨胀了 0.85-0.97(报告的 0.96-0.98 与修正后的 0.00-0.12),而一个廉价的单次尝试代理未能替代重复运行(Spearman $ ho = 0.417$)。基于功能正确性并不意味着安全性的证据,我们进一步提出了安全调整的 reliability@k,仅计算功能上正确且没有高严重性不安全模式的尝试。在与三个代理的初步实时 API 测试中,该调整未改变我们当前扫描器和阈值下的任何排名,因此我们将其作为一个建议的补充视角呈现,其决定性评估需要更强大的未来运行。最后,一个初步的 5 任务 SWE-bench Verified 试点在真实代码库环境中观察到了相同的核心问题:宏观平均隐性测试通过率为 0.80,而严格任务解决率为 0.20。
cs.AI / 38 / 2608.14746

Advanced modelling and data analytics in aviation

航空领域的先进建模与数据分析
Nanyonga, Aziida
Abstract
The aviation industry characterized by its stringent safety standards has seen a growing need for innovative approaches to enhance safety measures. Despite the vast accumulation of aviation safety data over time, its full potential in predicting and preventing incidents has not been fully realized. This research addresses this gap by applying machine learning (ML) and natural language processing (NLP) techniques to analyze aviation safety data from Socrata, the Australian Transport Safety Bureau (ATSB), the National Transportation Safety Board (NTSB), and the Aviation Safety Network (ASN). By leveraging existing ML models, including deep learning and transformer-based architectures alongside NLP methods for mining aviation incident narratives, this study uncovers patterns contributing to safety related incidents such as accidents and near-misses. Additionally, it employs various topic modelling techniques to extract meaningful themes from unstructured safety reports, enhancing the interpretability of incident analysis. Causal inference techniques and interpretable AI frameworks are further explored to improve model transparency and trustworthiness. A key contribution of this work is the deployment of advanced ML methodologies in a structured aviation safety context, assessing their effectiveness and providing insights into their practical implementation. The findings offer valuable insights for aviation stakeholders, including regulators, airlines, and policymakers, by providing data-driven solutions that enhance incident analysis and decision making. Ultimately, this research supports the industry s ongoing efforts to minimize risks, improve passenger and crew security, and integrate AI driven methodologies into aviation safety management.
Chinese Translation
航空行业以其严格的安全标准为特征,日益需要创新的方法来增强安全措施。尽管航空安全数据随着时间的推移积累了大量,但其在预测和防止事件方面的全部潜力尚未得到充分发挥。本研究通过应用机器学习(ML)和自然语言处理(NLP)技术,分析来自Socrata、澳大利亚交通安全局(ATSB)、国家运输安全委员会(NTSB)和航空安全网络(ASN)的航空安全数据,以填补这一空白。通过利用现有的机器学习模型,包括深度学习和基于变换器的架构,以及用于挖掘航空事件叙述的NLP方法,本研究揭示了导致安全相关事件(如事故和近失事)的模式。此外,它还采用各种主题建模技术,从非结构化的安全报告中提取有意义的主题,增强事件分析的可解释性。进一步探讨了因果推断技术和可解释的人工智能框架,以提高模型的透明度和可信度。本研究的一个关键贡献是在结构化的航空安全背景下部署先进的机器学习方法,评估其有效性并提供对其实际实施的见解。研究结果为航空利益相关者(包括监管机构、航空公司和政策制定者)提供了宝贵的见解,通过提供数据驱动的解决方案来增强事件分析和决策制定。最终,本研究支持行业持续努力,旨在最小化风险、提高乘客和机组人员的安全,并将基于人工智能的方法整合到航空安全管理中。
cs.AI / 39 / 2608.14765

Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs

无干净参考的主动数据清洗:能力与权衡的实验研究
Fadlallah, Hadi
Abstract
Data cleaning without a trusted clean reference is challenging because unusual values may represent either genuine errors or valid observations. This paper studies how different agent capabilities affect reference-free data cleaning and proposes an evidence-grounded framework that combines structured context, profiling, LLM reasoning, executable checks, controlled evidence retrieval, source ranking, citation alignment, conservative repair, reversible scripts, and provenance logging. Seven configurations are evaluated across financial, clinical, and environmental-monitoring datasets using controlled synthetic corruption and original-data descriptive analysis, resulting in 126 completed runs. The evaluation includes two comparison baselines and a progressive LLM-based sequence that adds executable tools, evidence retrieval, evidence controls, and conservative repair. In the synthetic evaluation, the deterministic profiling baseline achieved the highest detection F1-score of 0.561. Among the LLM-based configurations, the full conservative configuration achieved the highest F1-score of 0.421, but no configuration performed best across all evaluation criteria. The source-ranked configurations achieved the lowest unsupported-rule rates, while decision-level citation alignment remained weak. The full conservative configuration produced no unsafe or unnecessary modifications, although these rates were already zero before the conservative policy was added, and it performed no direct repairs. Overall, the results show that additional capabilities introduce trade-offs among detection, repair, evidence grounding, conservative behaviour, reproducibility, and operational cost rather than producing consistent improvements. The study provides a structured framework and empirical methodology for evaluating these trade-offs in reference-free agentic data cleaning.
Chinese Translation
在没有可信的干净参考的情况下,数据清洗面临挑战,因为异常值可能代表真实错误或有效观测。本文研究了不同代理能力如何影响无参考数据清洗,并提出了一个基于证据的框架,该框架结合了结构化上下文、特征分析、LLM推理、可执行检查、受控证据检索、源排名、引用对齐、保守修复、可逆脚本和来源日志记录。通过使用受控合成损坏和原始数据描述性分析,对金融、临床和环境监测数据集进行了七种配置的评估,共完成126次实验。评估包括两个比较基线和一个逐步的基于LLM的序列,该序列添加了可执行工具、证据检索、证据控制和保守修复。在合成评估中,确定性特征分析基线达到了最高的检测F1分数0.561。在基于LLM的配置中,完整的保守配置达到了最高的F1分数0.421,但没有任何配置在所有评估标准上表现最佳。源排名配置达到了最低的不支持规则率,而决策级引用对齐仍然较弱。完整的保守配置没有产生不安全或不必要的修改,尽管这些比率在添加保守政策之前已经为零,并且没有进行直接修复。总体而言,结果表明,额外的能力在检测、修复、证据基础、保守行为、可重复性和运营成本之间引入了权衡,而不是产生一致的改进。该研究提供了一个结构化框架和实证方法,用于评估无参考主动数据清洗中的这些权衡。
cs.AI / 40 / 2608.14771

From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving

从错误到证明:基于最小核心引导的神经符号约束求解修复
Sarkar, Dipankar
Abstract
Making language models solve constraint problems reliably often means having them translate the problem into a formal specification and delegating the search to a sound solver. But the translation is itself a language-model task, and an unfaithful translation makes the solver faithfully solve the wrong problem. Existing pipelines repair only translations that crash, returning the solver's error message and falling silent when the program runs but is wrong. We replace the error message with a proof: when the generated program is unsatisfiable, we extract a minimal unsatisfiable core over the model's own constraints and hand it back the exact set that cannot hold together, a leakage-free signal that localizes the fault. On a new benchmark of 77 problems with an exact oracle, translation to Answer Set Programming is faithful on six of seven domains and fails only on aggregate coverage scheduling, which concentrates the translation tax in one diagnosable pattern. A minimal core, rather than a bare error, is what stops a weaker model from fabricating solutions to infeasible problems, cutting fabrication from 79% to 7%. A strong chain-of-thought baseline meanwhile matches the symbolic route on accuracy, so the route's value is not accuracy but certificates and its refusal to fabricate.
Chinese Translation
使语言模型可靠地解决约束问题通常意味着将问题翻译为正式规范,并将搜索委托给一个健全的求解器。但翻译本身就是一个语言模型任务,而不忠实的翻译会导致求解器忠实地解决错误的问题。现有的流程仅修复崩溃的翻译,返回求解器的错误信息,而在程序运行但结果错误时则保持沉默。我们用证明替代错误信息:当生成的程序不可满足时,我们提取模型自身约束下的最小不可满足核心,并将其返回,提供一个不泄漏的信号,准确定位故障。在一个包含77个问题的新基准测试中,翻译为答案集编程(Answer Set Programming)在七个领域中的六个领域是忠实的,仅在聚合覆盖调度中失败,这将翻译的成本集中在一个可诊断的模式中。最小核心,而不是简单的错误,阻止了较弱模型伪造不可行问题的解决方案,将伪造率从79%降低到7%。与此同时,一个强大的思维链基线在准确性上与符号路径相匹配,因此该路径的价值不在于准确性,而在于证明和拒绝伪造的能力。
cs.AI / 41 / 2608.14789

Task-Driven Three-Layer Distributed Scheduling for Emergency Earth Observation in Large Low-Earth-Orbit Constellations

面向任务的三层分布式调度方法用于大规模低地球轨道星座的紧急地球观测
Yin, Qian, Wang, Xinwei, Wu, Guohua
Abstract
Large low-Earth-orbit (LEO) Earth-observation (EO) constellations offer frequent access to geographically dispersed ground targets, but emergency requests may arrive after committed routine-plan execution has begun. The resulting dynamic emergency observation scheduling problem (DEOSP) requires urgent tasks to be inserted under intermittent ground contact without excessive routine-plan disruption. To address DEOSP, we propose a task-driven three-layer distributed scheduling (T3L-DS) method, which represents task demand and sensor footprints on a common geographic grid and forms temporary clusters from observation capabilities and current inter-satellite links. For intra-cluster coordination, T3L-DS introduces onboard dual-plan bidding and joint marginal evaluation. It also designs an inter-cluster coordination mechanism for unresolved demand. Extensive computational experiments compare T3L-DS with centralised simulated annealing (SA), an adapted selective time-variant better reply process (A-SeTVBRP), and a conventional contract-net protocol (CNP). T3L-DS achieves the highest emergency coverage among the distributed methods, with average relative improvements of approximately 2.8% and 17.1% over A-SeTVBRP and CNP, respectively. Its average relative gap from SA is approximately 7.1%. Under conflict-enhanced loads, it reduces routine-coverage loss by approximately 57.9% and 87.7% relative to A-SeTVBRP and CNP, respectively. The ablation study confirms the contribution of the proposed coordination enhancements. Overall, the results show that T3L-DS provides an effective distributed approach to DEOSP.
Chinese Translation
大规模低地球轨道(LEO)地球观测(EO)星座能够频繁访问地理分散的地面目标,但在已开始执行的常规计划后,紧急请求可能会到达。由此产生的动态紧急观测调度问题(DEOSP)需要在间歇性地面联系下插入紧急任务,而不造成过多的常规计划干扰。为了解决DEOSP问题,我们提出了一种面向任务的三层分布式调度(T3L-DS)方法,该方法在一个共同的地理网格上表示任务需求和传感器覆盖范围,并根据观测能力和当前的星间链路形成临时集群。为了实现集群内部的协调,T3L-DS引入了机载双计划竞标和联合边际评估机制。同时,它还设计了一个用于未解决需求的集群间协调机制。大量计算实验将T3L-DS与集中式模拟退火(SA)、改进的选择性时间变异更好回复过程(A-SeTVBRP)和传统合同网协议(CNP)进行了比较。结果表明,T3L-DS在分布式方法中实现了最高的紧急覆盖率,相较于A-SeTVBRP和CNP,其平均相对改善分别约为2.8%和17.1%。与SA相比,其平均相对差距约为7.1%。在冲突增强负载下,相较于A-SeTVBRP和CNP,T3L-DS将常规覆盖损失分别降低了约57.9%和87.7%。消融研究确认了所提出的协调增强的贡献。总体而言,结果表明T3L-DS为DEOSP提供了一种有效的分布式解决方案。
cs.AI / 42 / 2608.14791

CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs

CEDAR-GRPO:面向过程的强化学习在大型语言模型中的一般性推理能力
Salimi, Moein, Parnian, Danial, Adim, Shaygan, Ebrahiminasab, Amirmohammad, Alighardashi, Nima, Gholami, Parsa, Akramipour, Sahand, Siavoshani, Mahdi Jafari, Rohban, Mohammad Hossein
Abstract
Abductive reasoning, often characterized as inference to the best explanation, is central to explanation under uncertainty, from everyday sense-making and investigation to scientific discovery. Yet LLM research has mostly studied abduction through narrow, task-specific benchmarks, making it unclear whether observed gains transfer beyond the benchmark family used for training or evaluation. We ask whether RL post-training can improve abduction as a transferable reasoning capability. We introduce CEDAR-GRPO, a process-aware framework that combines final-answer correctness with abductive rewards for evidence coverage and evidence-to-explanation directionality. Four open-weight LLMs are post-trained on a controlled, domain-neutral mixture of abductive hypothesis-generation and hypothesis-selection tasks. We evaluate them on 11 unseen tasks spanning hypothesis selection, missing-fact generation, defeasible inference, long-context investigation, clinical reasoning, code debugging, and non-abductive controls. CEDAR- GRPO improves every model on every held-out task over both base models and correctness-only GRPO, with average gains of 7.4 and 2.7 points, respectively, and a maximum gain of 30.8 points. Ablations confirm that RL, abductive reward design, and task diversity each contribute to transfer. Process-level metrics further show stronger abductive behavior, including exploration of alternatives, elimination of rivals, backtracking, and uncertainty marking.
Chinese Translation
溯因推理通常被描述为推导最佳解释的过程,是在不确定性下进行解释的核心,从日常的意义构建和调查到科学发现。然而,现有的大型语言模型(LLM)研究大多通过狭窄的、特定任务的基准进行溯因推理的研究,这使得观察到的进展是否能够转移到训练或评估所用的基准家族之外变得不明确。我们探讨了强化学习(RL)后训练是否能够提升溯因推理作为一种可转移的推理能力。我们提出了CEDAR-GRPO,这是一种面向过程的框架,结合了最终答案的正确性与针对证据覆盖和证据到解释的方向性的溯因奖励。四个开放权重的LLM在一个受控的、领域中立的溯因假设生成和假设选择任务的混合数据集上进行了后训练。我们在11个未见任务上进行了评估,这些任务涵盖了假设选择、缺失事实生成、可推翻推理、长上下文调查、临床推理、代码调试和非溯因控制。CEDAR-GRPO在每个保留任务上都提升了每个模型的表现,相较于基础模型和仅考虑正确性的GRPO,平均提升分别为7.4和2.7分,最大提升为30.8分。消融实验确认了RL、溯因奖励设计和任务多样性各自对转移的贡献。过程级别的指标进一步显示了更强的溯因行为,包括对替代方案的探索、对竞争者的排除、回溯和不确定性标记。
cs.AI / 43 / 2608.14795

Individual Disempowerment through an Advice Channel: Control Loss when Influence is Endogenous

通过建议渠道实现个体失能:当影响是内生时的控制丧失
Oberman, Adam M.
Abstract
An AI that can only give advice seems safe: the human is always free to ignore it. That is the premise of the boxing tradition in AI safety, and its long-suspected weak point is that the human who reads the answers is part of the system. We make the fraction $\varepsilon_t$ of behavior that follows the advice a state of a Markov decision process, moved by the advisor's own messages, so that use deepens reliance. Granted a channel rich enough to echo any action the human could take, higher $\varepsilon_t$ weakly lowers every monotone measure of the power of a human with a message-independent fallback. An oracle rewarded by per-round approval cultivates reliance beyond a closed-form patience threshold, so the same reward weights leave the optimal oracle answering in episodic deployments and cultivating in long-memory ones. An influence bound certified once at deployment is blind to that horizon and bounds the loss no lower than its trivial ceiling. An exogenous cap on influence bounds the guarantee the human loses, and a short enough memory reset removes the incentive to cultivate, while neither recovers the value already steered away. In a closed-form example the optimal oracle never cultivates in fifteen-round sessions and does in sixteen.
Chinese Translation
一个只能提供建议的人工智能似乎是安全的:人类始终可以选择忽视它。这是人工智能安全中的“拳击传统”的前提,而其长期以来被怀疑的弱点在于,阅读答案的人类是系统的一部分。我们将遵循建议的行为比例 $eta_t$ 视为马尔可夫决策过程中的一个状态,由顾问自身的信息推动,从而使得使用加深了依赖性。假设有一个足够丰富的渠道可以回响人类可能采取的任何行动,较高的 $eta_t$ 会弱化每个与消息无关的后备选项的人类的权力的单调测量。一个通过每轮批准获得奖励的神谕者在超越封闭形式耐心阈值的情况下培养了依赖性,因此相同的奖励权重使得最优神谕者在情节部署中进行回答,而在长记忆的情况下进行培养。一次在部署时认证的影响界限对该视野是盲目的,并且将损失限制在不低于其微不足道的上限。外生的影响上限限制了人类必然失去的保证,而足够短的记忆重置则消除了培养的激励,但都无法恢复已经偏离的价值。在一个封闭形式的例子中,最优神谕者在十五轮的会话中从不进行培养,而在十六轮的会话中则进行培养。
cs.AI / 44 / 2608.14804

Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning

生成的上下文与受控状态:可追溯的纵向临床推理的功能条件
Pissarra, Augusto Bernardo, Souza, Victor Lorena de Farias
Abstract
Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is currently true about a patient. This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over. We distinguish generated context from governed state; separate five objects that clinical AI habitually conflates (true state, observations, evidence, belief, and simulated state); define a tiered governance standard against which any clinical AI system can be audited; and show that an operational definition of accountability decomposes into four information requirements: an immutable evidence ledger with awareness-time versioning, a belief state distinct from accumulated evidence, an observation-process model, and claim-level causal typing. We are explicit that this decomposition is analytic rather than a necessity theorem, and that its value is conceptual hygiene: it converts "accountable clinical AI" from a slogan into an audit instrument. A six-level maturity framework separates what a system makes governable from what it can compute, locating current LLM-centric practice at high capability but low maturity. The paper is fully self-contained: the four research questions the framework poses are stated in the introduction, and the conclusion records what the paper establishes toward each; future work develops the buildable core of the architecture and the research program toward full Clinical World Models. No empirical result is claimed here.
Chinese Translation
大型语言模型(LLMs)已成为临床人工智能的主导接口,但它们所呈现的接口(文本输入,文本输出,每次一个上下文窗口)并未对患者当前的真实情况保持明确、持久的受控表示。本文认为,纵向临床推理是一个在部分可观测性下的状态估计问题,临床人工智能成功或失败的轴心并不是模型读取记录的流畅性,而是其推理所依据的患者状态的治理。我们区分生成的上下文与受控状态;分离临床人工智能习惯上混淆的五个对象(真实状态、观察、证据、信念和模拟状态);定义一个分层治理标准,以便对任何临床人工智能系统进行审计;并展示可追溯性的操作定义分解为四个信息需求:具有意识时间版本控制的不变证据账本、与累积证据不同的信念状态、观察过程模型和声明级因果类型。我们明确指出,这一分解是分析性的,而非必要定理,其价值在于概念的清晰性:它将“可追溯的临床人工智能”从口号转变为审计工具。一个六级成熟度框架将系统可治理的内容与其计算能力区分开来,当前以LLM为中心的实践定位于高能力但低成熟度。本文是完全自足的:框架提出的四个研究问题在引言中说明,结论记录了本文在每个问题上所建立的成果;未来的工作将发展架构的可构建核心及朝向完整临床世界模型的研究计划。本文不声称有任何实证结果。
cs.AI / 45 / 2608.14808

Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking

大型语言模型是否知道该问什么以及何时提问?多轮信息获取的评估
Huang, Yepeng, Zhang, Jiawen, Dai, Michelle, Su, Xiaorui, Gao, Shanghua, Wang, Zi, Zitnik, Marinka
Abstract
When a user question is underspecified, a capable model should recognize that its context is insufficient, identify the missing information, ask for it, and respond only once that information determines a unique answer. We formalize multi-turn information seeking as solving a k-underspecified constraint satisfaction problem, where k is the number of variables jointly required to determine the target and therefore measures the degree of missing information. We instantiate the formulation in MT-InfoSeek, a controlled evaluation suite of 5,251 problems and 9,006 task instances spanning mathematics, logic, biology, medicine, and general knowledge. We evaluate models along three axes: what they ask, when they ask it, and how the acquired information affects the final answer. Performance degrades across models and domains as underspecification increases. Models recognize that additional information is needed but underestimate how much, and in logical problems at k = 2 they under-predict the degree of missing information about four times as often as they over-predict it. They also fail to identify a minimal sufficient set of queries, improve only marginally when given the true k, and often stop before acquiring sufficient information. In tasks with ordered dependencies, an incorrect query order reduces final accuracy even when the model eventually acquires all necessary information. We measure information seeking directly through final sufficiency, which records whether the acquired information determines the target independent of answer generation. This separation shows differences between models that final accuracy alone does not capture, and indicates that the ability to seek information over multiple turns is distinct from the ability to generate answers and is not measured by current LLM evaluations.
Chinese Translation
当用户的问题不够明确时,一个有能力的模型应该能够识别其上下文不足,确定缺失的信息,提出相关问题,并在该信息确定唯一答案后再进行回应。我们将多轮信息获取形式化为解决一个k-不明确约束满足问题,其中k是确定目标所需的变量数量,因此衡量缺失信息的程度。我们在MT-InfoSeek中实例化该形式化,这是一个包含5,251个问题和9,006个任务实例的受控评估套件,涵盖数学、逻辑、生物学、医学和一般知识等领域。我们从三个方面评估模型:它们问什么、何时提问以及获取的信息如何影响最终答案。随着不明确性的增加,模型和领域的表现普遍下降。模型能够识别出需要额外信息,但低估了所需信息的数量,在k = 2的逻辑问题中,它们低估缺失信息的程度的频率约为高估的四倍。它们还未能识别出最小的充分查询集,在给定真实k时仅有边际改善,并且常常在获取足够信息之前就停止。在具有有序依赖关系的任务中,即使模型最终获取了所有必要信息,错误的查询顺序也会降低最终准确性。我们通过最终充分性直接测量信息获取,这记录了获取的信息是否独立于答案生成而确定目标。这种分离显示了模型之间的差异,而仅凭最终准确性无法捕捉到这些差异,并表明在多轮中寻求信息的能力与生成答案的能力是不同的,当前的大型语言模型评估并未对此进行测量。
cs.AI / 46 / 2608.14828

MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment

MINT:用于平衡多目标对齐的最小选择偏好蒸馏
Tu, Tony, Chakraborty, Sayan, Xu, Ruomeng, Qin, Tony, Tian, Austin
Abstract
Aligning a language agent to several objectives at once is a persistent failure mode of preference-based training: when objectives are combined additively, optimization collapses onto whichever is cheapest to improve and sacrifices the rest, so a support agent learns to sound warm while giving no real help. The root issue is that an additive reward has no notion of balance. We introduce Mint (MIN-selection preference disTillation), a one-line change to preference distillation: rather than ranking sampled candidates by a weighted sum of rewards, we rank them by their weakest objective, distilling the best-balanced candidate over the most lopsided one with an unchanged DPO objective. This is the p -> negative infinity limit of a generalized-mean family spanning additive to worst-case selection. Across cooperative emotional support and adversarial negotiation, min-selection lifts both objectives while sharply cutting their imbalance; on emotional support it raises the weaker axis from 0.37 to 0.64 (p < 10^-40), surpassing human experts and persisting across full multi-turn rollouts. A turn-by-turn analysis yields our central finding: min-selection corrects imbalance in proportion to how imbalanced the reference policy is, and its benefit endures over an interaction precisely as long as that imbalance does.
Chinese Translation
将语言代理与多个目标同时对齐是基于偏好的训练中一个持续存在的失败模式:当目标以加法方式组合时,优化会集中在最便宜的改进上,从而牺牲其他目标,使得支持代理听起来温暖却没有提供实质帮助。根本问题在于,加性奖励没有平衡的概念。我们提出了Mint(最小选择偏好蒸馏),这是对偏好蒸馏的一项简单修改:我们不再通过奖励的加权和对采样候选进行排名,而是通过它们的最弱目标进行排名,从而在不改变DPO目标的情况下,蒸馏出最佳平衡的候选而非最偏颇的候选。这是一个广义均值家族的极限情况,涵盖了从加性到最坏情况选择。在合作情感支持和对抗性谈判中,最小选择提升了两个目标,同时显著减少了它们的不平衡;在情感支持方面,它将较弱的轴从0.37提升至0.64(p < 10^-40),超越了人类专家,并在完整的多轮回合中保持稳定。逐轮分析得出了我们的核心发现:最小选择根据参考策略的不平衡程度来纠正不平衡,其益处在交互中持续的时间正好与该不平衡持续的时间相同。
cs.AI / 47 / 2608.14841

What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering

重排序器的视角:用于长文档多模态问答的多方面页面注释
Wu, Guanchen, Ding, Jiayuan, Mukherjee, Subhabrata, Yang, Carl
Abstract
Long-document visual question answering (VQA) over documents of tens to hundreds of pages mixing text, tables, charts, and figures typically follows retrieve-then-read pipelines. In our setting, the bottleneck shifts from retrieval recall to reranker-side evidence selection: on MMLongBench-Doc, BGE-M3 reaches Recall@20 = 0.86 but only F1@5 = 0.254, and even the visual retriever ColPali reaches only F1@5 = 0.332; a text-only rerank LLM seeing only raw snippets misses table, chart, and layout evidence even when the upstream retriever encoded images. We propose Trident, with two complementary components: Trident-R, a retriever-agnostic LLM reranker that converts each candidate into an LLM-readable semantic record, including a visual caption, section path, entity tags, multi-axis concept hits, and a text snippet, then performs a single adaptive-K rerank call; and Trident-S, a generation-side module that prompts the VLM under topical, entity, and structural lenses before synthesis. On two long-document datasets, the annotation+rerank protocol substantially improves retrieval F1 across five heterogeneous pools, with every reranked pool exceeding the strongest adaptive-K baseline PageIndex. An LLM rerank without the annotation barely changes first-hit ranking, indicating the lift comes from the structured annotation. Trident-S targets open-ended synthesis questions by design, adding up to 6.6 points in generation accuracy on these questions. The best Trident configuration is the strongest downstream QA pipeline in our evaluation, with rankings consistent across two LLM judges (kappa = 0.913).
Chinese Translation
长文档视觉问答(VQA)通常涉及数十到数百页的文本、表格、图表和图形,通常遵循检索-再阅读的流程。在我们的研究中,瓶颈从检索召回转移到重排序器端的证据选择:在 MMLongBench-Doc 数据集上,BGE-M3 达到 Recall@20 = 0.86,但 F1@5 仅为 0.254,甚至视觉检索器 ColPali 的 F1@5 也仅为 0.332;仅使用文本的重排序 LLM 仅看到原始片段,错过了表格、图表和布局证据,即使上游检索器编码了图像。我们提出了 Trident,包含两个互补组件:Trident-R,一个与检索器无关的 LLM 重排序器,将每个候选项转换为 LLM 可读的语义记录,包括视觉标题、章节路径、实体标签、多轴概念命中和文本片段,然后执行一次自适应 K 重排序调用;以及 Trident-S,一个生成侧模块,在合成之前从主题、实体和结构的角度提示 VLM。在两个长文档数据集上,注释+重排序协议显著提高了五个异构池的检索 F1,每个重排序池均超过最强的自适应 K 基线 PageIndex。没有注释的 LLM 重排序几乎没有改变首次命中的排名,表明提升来自于结构化注释。Trident-S 设计上针对开放式合成问题,在这些问题上的生成准确性提高了多达 6.6 分。最佳 Trident 配置是我们评估中最强的下游 QA 流水线,在两个 LLM 评审者之间的排名一致性良好(kappa = 0.913)。
cs.AI / 48 / 2608.14851

Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning

利用离线强化学习发现高质量的国际象棋难题
Nie, Allen, Badrinath, Anirudhan, Tomlin, Nicholas, Dai, Timothy, Yip, Carissa, Wang, Rose E, Brunskill, Emma, Piech, Chris
Abstract
Learning and skill mastery require extensive and deliberate practice. In many learning settings, producing high-quality pedagogical materials can require a high level of domain expertise and be very time-consuming. Pedagogical materials often need to train students to engage in different thinking patterns. In some domains, such as chess, puzzles are used to help students practice their skills in calculating the next moves and recognizing known patterns on a board. Giving students a practice set of puzzles to help them learn different modes of thinking is challenging because the teacher needs to carefully balance between different motifs and how many look-ahead steps a student needs to perform. Popular online platforms like Chess.com and Lichess offer players millions of puzzles. Unlike chess tactics puzzles procured by human experts, where chess beginners can learn valuable insights, these puzzles are automatically generated and often regarded as having low pedagogical value. These platforms also rely on a heuristic to recommend puzzles to users for practice. Using the user history data over an entire year, a total of 1.5 billion puzzle-solving histories, we learn the pedagogical value of a puzzle and how to automatically choose a set of puzzles to better support chess learners using insights from offline reinforcement learning. We show that using offline policy evaluation, our trained policy has significant impact on beginners with puzzle-solving Elo range of 100--1000, particularly for the group of beginners whose learning growth was stagnant. We also performed a qualitative analysis of the puzzles discovered by our model by collecting annotation ratings from expert chess players. The success of our pipeline shows promise for a future where we can understand the pedagogical values of practice items given general user interaction data.
Chinese Translation
学习和技能掌握需要广泛而有意识的练习。在许多学习环境中,制作高质量的教学材料可能需要较高的领域专业知识,并且非常耗时。教学材料通常需要训练学生以参与不同的思维模式。在某些领域,如国际象棋,难题被用来帮助学生练习计算下一步棋和识别棋盘上已知模式的技能。给学生提供一套练习难题以帮助他们学习不同的思维模式是具有挑战性的,因为教师需要仔细平衡不同的主题以及学生需要执行的前瞻性步骤数量。像 Chess.com 和 Lichess 这样的流行在线平台为玩家提供了数百万个难题。与由人类专家采购的国际象棋战术难题不同,初学者可以从中学习到有价值的见解,这些难题是自动生成的,通常被认为具有较低的教学价值。这些平台还依赖启发式算法向用户推荐练习难题。通过分析整整一年的用户历史数据,总计 15 亿个难题解决历史,我们学习了难题的教学价值以及如何自动选择一组难题,以更好地支持国际象棋学习者,利用离线强化学习的见解。我们展示了使用离线策略评估,我们训练的策略对难题解决 Elo 范围为 100-1000 的初学者有显著影响,特别是对那些学习增长停滞的初学者群体。我们还通过收集专家棋手的注释评分,对我们模型发现的难题进行了定性分析。我们的流程的成功为未来理解给定一般用户互动数据的练习项目的教学价值提供了希望。
cs.AI / 49 / 2608.14870

JarvisBench: Always-on Intelligence Between Humans and Agents

JarvisBench:人类与智能体之间的持续智能
Chen, Chen, Chen, Zhehuai
Abstract
Long-horizon agents can execute continuously, but human attention remains intermittent and scarce. This creates a bidirectional coordination problem: users may need immediate access to an agent while work continues in the background, whereas agents may encounter consequential decisions that require user judgment after the user has stopped monitoring execution. We posit an always-on attention-coordination layer---\textit{Jarvis}\footnote{Named after the fictional AI assistant in \textit{Iron Man}.}---that mediates this interface and allocates human attention across one or more working agents. We introduce \textit{JarvisBench} to evaluate both directions of this coordination: whether an intermediary can accurately and promptly answer user-initiated questions about ongoing work, and whether it can recognize when an agent requires user judgment, solicit that judgment at the right moment, and route it back to improve task outcomes. JarvisBench contains 45 agentic task instances: 20 single-agent tasks and 25 workstreams organized into 10 multi-agent projects. The tasks span 19 domains and were selected and adapted from more than 2,000 public candidates. Crucially, the need for user attention arises naturally during execution rather than from an obvious omission in the initial prompt. JarvisBench is designed to integrate with arbitrary agent runtimes without modifying their underlying execution loops. Our reference implementation further provides a full-duplex speech interface, allowing users to reach Jarvis naturally while timely attention coordination supports agents working in the background. By separating agent execution from attention coordination, JarvisBench provides a stable evaluation target as agent capabilities continue to improve.
Chinese Translation
长时间运行的智能体可以持续执行任务,但人类的注意力却是间歇性和稀缺的。这就产生了一个双向协调问题:用户可能需要在工作继续进行的同时立即访问智能体,而智能体可能会遇到需要用户判断的重要决策,而此时用户已停止监控执行。我们提出了一种始终在线的注意力协调层—— extit{Jarvis} ootnote{以虚构的人工智能助手 extit{Iron Man}中的角色命名。}——来调解这一接口,并在一个或多个工作智能体之间分配人类注意力。我们引入了 extit{JarvisBench}来评估这种协调的两个方向:中介是否能够准确及时地回答用户发起的关于正在进行工作的提问,以及它是否能够识别智能体何时需要用户判断,在合适的时机请求该判断,并将其反馈以改善任务结果。JarvisBench包含45个智能任务实例:20个单智能体任务和25个工作流,组织成10个多智能体项目。这些任务跨越19个领域,并从2000多个公共候选任务中选择和调整而来。重要的是,用户注意力的需求在执行过程中自然产生,而不是由于初始提示中的明显遗漏。JarvisBench旨在与任意智能体运行时集成,而无需修改其基础执行循环。我们的参考实现进一步提供了全双工语音接口,使用户能够自然地与Jarvis进行互动,同时及时的注意力协调支持在后台工作的智能体。通过将智能体执行与注意力协调分开,JarvisBench为智能体能力的持续提升提供了一个稳定的评估目标。
cs.AI / 50 / 2608.14881

Personalized Auto-Research: Towards a True AI Co-Scientist

个性化自动研究:迈向真正的人工智能共同科学家
Ni, Bo, Dernoncourt, Franck, Chen, Hongjie, Wang, Yu, Ahmed, Nesreen K., Tu, Zhengzhong, Derr, Tyler, Rossi, Ryan A.
Abstract
AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and draft full papers are beginning to change how research is carried out. Despite this rapid progress, state-of-the-art systems remain researcher-agnostic: given a research goal, they optimize novelty, validity, or reviewer score while ignoring the individual scientist who will use the output. This overlooks a fundamental fact about research, namely, that what counts as novel, valuable, or feasible depends on the researcher, including their prior work, methodological repertoire, and the collaborators and communities in which they are embedded. In this work, we introduce the problem of personalized auto-research, which conditions every stage of the research process on a representation of the individual researcher. We argue that personalization is not a convenience layer, but rather the fundamental property that allows an AI system to serve as a genuine co-scientist rather than a generic instrument. To address this problem, we propose a general and flexible framework that threads a graph-grounded researcher context through retrieval, hypothesis search, experimentation, writing, and review. The framework consists of three fundamental components: (i) graph-grounded researcher representations, (ii) personalization across the full research pipeline, and (iii) evaluation grounded in the individual. Notably, we highlight a one-size-fits-all failure mode where distinct researchers issuing the same goal receive essentially the same research, erasing the tacit knowledge through which novel ideas arise. Finally, we discuss fundamental open problems and challenges.
Chinese Translation
生成假设、检索相关工作、设计实验、执行代码和撰写完整论文的人工智能共同科学家正在改变研究的开展方式。尽管这一快速进展,最先进的系统仍然是研究者无关的:在给定研究目标的情况下,它们优化新颖性、有效性或评审分数,而忽视了将使用输出的个体科学家。这忽视了关于研究的一个基本事实,即新颖、有价值或可行的标准取决于研究者,包括他们的先前工作、方法论能力以及他们所嵌入的合作者和社区。在本研究中,我们引入了个性化自动研究的问题,该问题将研究过程的每个阶段都基于个体研究者的表征。我们认为,个性化不是一个便利层,而是使人工智能系统能够作为真正的共同科学家而非通用工具的基本属性。为了解决这个问题,我们提出了一个通用且灵活的框架,将图形基础的研究者背景贯穿于检索、假设搜索、实验、写作和评审。该框架由三个基本组成部分构成:(i)图形基础的研究者表征,(ii)贯穿整个研究流程的个性化,以及(iii)基于个体的评估。值得注意的是,我们强调了一种一刀切的失败模式,即不同研究者在提出相同目标时获得基本相同的研究,抹去了产生新颖想法的隐性知识。最后,我们讨论了基本的开放问题和挑战。
cs.AI / 51 / 2608.14903

Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence

前沿人工智能预测存在测量问题:进展证据的审计
Costa, Fabricio F
Abstract
Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief. This paper audits whether the public measurement record supports those connections before another trend is fitted. I construct a frozen, event-centric record through 12 August 2026 with 62 selected systems, 12 versioned benchmarks, seven capability or impact criteria, 144 graded events, 27 source records, and 408 typed relations. The record is an audit sample, not a census. Only seven systems jointly observe estimated training compute and a METR 50 percent task horizon. Training compute is absent for 19 of 27 closed systems, including every selected closed release from 2026, while none of the 35 open-weight systems has a METR horizon observation. Benchmark succession creates a second break: a seven-system link from METR Time Horizon 1.0 to 1.1 has a log-scale slope of 1.206 (95 percent CI 1.021 to 1.390), whereas a six-system MMLU to MMLU-Pro comparison appears shift-like under logit and probit links but not under linear or logarithmic links. The observed bridges have about 80 percent power only for slope departures near 25 percent. Provenance is concentrated: 52 of 71 substantive quantitative events, or 73.2 percent, come from one measurement programme, and 76.1 percent are laboratory releases. A review of 56 methodological and empirical sources identifies 16 complementary measurement directions spanning resources, inference budgets, reliability, agentic work, safety, human preference, field outcomes, and forecast backtesting. No direction supplies a replacement scalar. The result is not that frontier AI forecasting is impossible, but that a defensible dated forecast is a claim about a versioned measurement system with explicit joins, protocols, links, and source dependence, not merely a fitted curve or calendar date.
Chinese Translation
前沿人工智能的定量预测通常将过时的目标与基准分数、训练计算、发布时机或专家信念的趋势联系起来。本文审计了公共测量记录是否支持这些联系,以便在拟合另一个趋势之前进行评估。我通过 2026 年 8 月 12 日构建了一个冻结的、以事件为中心的记录,涵盖 62 个选定系统、12 个版本化基准、七个能力或影响标准、144 个评分事件、27 个源记录和 408 个类型关系。该记录是一个审计样本,而非普查。只有七个系统共同观察到估计的训练计算和 METR 50% 任务视野。在 27 个封闭系统中,有 19 个缺乏训练计算数据,包括 2026 年所有选定的封闭发布,而 35 个开放权重系统中没有一个具有 METR 视野观察。基准的继承造成了第二个断裂:从 METR 时间视野 1.0 到 1.1 的七系统链接具有 1.206 的对数尺度斜率(95% CI 1.021 至 1.390),而六系统的 MMLU 与 MMLU-Pro 比较在 logit 和 probit 链接下似乎呈现出转变特征,但在线性或对数链接下则没有。观察到的桥接仅在斜率偏离接近 25% 时具有约 80% 的功效。来源集中:71 个实质性定量事件中有 52 个(73.2%)来自一个测量项目,76.1% 是实验室发布。对 56 个方法论和实证来源的审查确定了 16 个互补的测量方向,涵盖资源、推理预算、可靠性、代理工作、安全性、人类偏好、领域结果和预测回测。没有一个方向提供替代标量。结果并不是前沿人工智能预测不可能,而是一个可辩护的过时预测是关于一个版本化测量系统的声明,具有明确的连接、协议、链接和源依赖,而不仅仅是一个拟合曲线或日历日期。
cs.AI / 52 / 2608.14927

LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks

大型语言模型可以预测失败风险,但在预测哪种协作协议更有利方面存在困难:跨推理任务的成本感知协议路由
Yang, Chih-Hsuan, Jiang, Jingyan, Yang, Cheng-Hau, Vasudevan, Vikram, Zheng, Huihuo, Vishwanath, Venkatram, Thakur, Rajeev
Abstract
Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but deployment requires deciding when extra collaboration is worth its cost. We isolate this decision by running every problem under four protocols while holding the solver fixed within each setting: direct solving (Baseline), iterative self-correction (Single), planner-executor-reviewer collaboration (PER), and multi-agent deliberation (Broadcast). The primary benchmark comprises 4,181 competition-level math problems; paired robustness checks cover four benchmarks spanning competition math, biology, and broader science with two solver families. Across fixed policies, trained routers, and frozen LLM routers, conservative policies under-escalate, whereas higher-solve frozen routers often over-escalate. A post-answer, pre-collaboration gpt-oss-120b probe ranks Baseline failures with 0.8847 AUROC (4,151 parseable cases; 95% CI [0.8732, 0.8955]). The same score remains informative for predicting whether any collaboration helps (0.7683 AUPRC), but is much weaker for identifying PER- or Broadcast-specific value (0.1674 and 0.1041 AUPRC). Separately, the pre-answer self-confidence gate reaches 78.0% solve at 45K tokens, compared with 73.8% at 71.3K for a frozen gpt-oss-120b router and 92.4% for a retrospective fixed-order oracle. Across 10 paired model-condition settings, the oracle adds 23.2-58.3 points of retrospective coverage over Baseline, but protocol profiles vary by task. In the six settings with held-out router evaluations, oracle gaps remain 18.5-28.9 points. Confidence can therefore support initial escalation, while protocol-specific cost-aware routing remains unresolved.
Chinese Translation
多智能体大型语言模型(LLM)系统可以通过增加计算量来提高推理能力,但部署时需要决定额外的协作是否值得其成本。我们通过在四种协议下运行每个问题,同时在每种设置中保持求解器不变,来隔离这一决策:直接求解(基线)、迭代自我修正(单一)、规划者-执行者-审查者协作(PER)和多智能体审议(广播)。主要基准包括4181个竞争级数学问题;配对的稳健性检查涵盖了跨越竞争数学、生物学和更广泛科学的四个基准,涉及两种求解器系列。在固定策略、训练的路由器和冻结的LLM路由器中,保守策略往往低估升级,而高求解的冻结路由器则常常过度升级。一个在答案后、协作前的gpt-oss-120b探测器对基线失败的排名为0.8847 AUROC(4151个可解析案例;95% CI [0.8732, 0.8955])。同样的得分对于预测任何协作是否有帮助(0.7683 AUPRC)仍然具有信息价值,但在识别PER或广播特定价值方面则显著较弱(0.1674和0.1041 AUPRC)。另外,预答案自信门在45K个标记时达到78.0%的求解率,而冻结的gpt-oss-120b路由器在71.3K时为73.8%,而回顾性固定顺序的神谕为92.4%。在10个配对模型条件设置中,神谕在基线之上增加了23.2-58.3点的回顾性覆盖,但协议特征因任务而异。在六个保留路由器评估的设置中,神谕差距仍为18.5-28.9点。因此,自信可以支持初步升级,而协议特定的成本感知路由仍未解决。
cs.AI / 53 / 2608.14936

Small Models Scout Bottleneck Order for Large-Model Data Control

小模型侦测大模型数据控制中的瓶颈顺序
Choi, Seungmin, Sung, Jiwon, Umer, Muhammad, Gorle, Abhiram Rao, Son, Guijin, Yu, Youngjae, Cioffi, John M.
Abstract
Small proxy models are commonly used to identify data mixtures for larger-scale training. We ask whether their training trajectories reveal another transferable structure: the order in which larger models should resolve skill bottlenecks. We formulate first-passage skill training, where each monitored skill has a target floor and the objective is to minimize the tokens required to reach all floors. We introduce LogFloor, a closed-loop controller that directs each round toward current bottlenecks, producing phase-ordered resolution trajectories. Across five bAbI skill slices on Qwen2.5-1.5B, LogFloor reduces token cost by 56.2% on average. In 70M-to-12B transfer, three-round replay of a 70M scout path reaches every floor in all eight target runs, saving 30.9% by pair mean, 39.4% in pooled training tokens, and 37.6% under source-cost accounting. On MMLU-control, a frozen scout path succeeds across all eight 12B runs. Collapsing a path to its static marginal mixture or reversing its phase order removes most benefits, while bottleneck labels alone remain partially useful. These results identify phase-ordered bottleneck resolution as a transferable curriculum structure for monitored skill-targeted training.
Chinese Translation
小型代理模型通常用于识别用于大规模训练的数据混合。我们探讨它们的训练轨迹是否揭示了另一种可转移的结构:较大模型应解决技能瓶颈的顺序。我们提出了首次通过技能训练(first-passage skill training),其中每个监控的技能都有一个目标底线,目标是最小化达到所有底线所需的令牌数。我们引入了LogFloor,一个闭环控制器,指引每一轮朝向当前瓶颈,产生阶段有序的解决轨迹。在五个bAbI技能切片上,使用Qwen2.5-1.5B,LogFloor平均减少了56.2%的令牌成本。在70M到12B的迁移中,70M侦测路径的三轮重放在所有八个目标运行中达到了每个底线,按对均值节省了30.9%,在汇总训练令牌中节省了39.4%,在源成本核算下节省了37.6%。在MMLU-control上,一个冻结的侦测路径在所有八个12B运行中成功。将路径压缩到其静态边际混合或反转其阶段顺序会消除大部分好处,而仅靠瓶颈标签仍然部分有效。这些结果将阶段有序的瓶颈解决识别为一种可转移的课程结构,用于监控的技能目标训练。
cs.AI / 54 / 2608.14940

When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation

代理评估何时结束?结果的最终性与跨单元分离
Casheekar, Avyay M.
Abstract
Current agent evaluations score models on the state visible at the end of a stopped run which they count as one trial. However, interpreting the score as a final result would require two conditions that the endpoint does not itself necessarily establish: outcome finality and cross-unit separation. These conditions are independent, since reconciling a delayed outcome can settle the label while runs still share state and isolating runs can prevent carryover while the scored outcome remains unfinished. We develop a completion argument that specifies the evidence needed for each decision and argue that a final label is justified only when anything that could still change the claimed outcome is resolved, bounded, or retained as uncertainty. First, in a controlled replay to demonstrate the mechanism where an agent's actions were held fixed, we find that the endpoint and terminal labels differ for every delayed operation, while a delayed write changes the next run's score when service state persists between runs but not after isolation or verified reset. Second, in a review of ten public protocols, we find that all protocols identify when a run stops and what is scored, while unfinished operations and the evidence for treating runs as separate trials are documented less consistently. Finally, we propose an open-effects record that lists operations or resources that may remain relevant after the endpoint, their current status, and whether they could change the scored outcome or affect another run.
Chinese Translation
当前的代理评估在停止运行结束时对可见状态进行评分,并将其视为一次试验。然而,将该评分解释为最终结果需要两个条件,而终点本身并不一定能够确立这两个条件:结果的最终性和跨单元分离。这些条件是独立的,因为调解延迟结果可以在运行仍共享状态时确定标签,而隔离运行可以在评分结果仍未完成时防止结果的延续。我们提出了一个完成论证,明确了每个决策所需的证据,并主张只有在任何可能改变所声称结果的因素得到解决、界定或保留为不确定性时,最终标签才是合理的。首先,在一个受控重放实验中,我们展示了代理的动作被固定的机制,发现对于每个延迟操作,终点和终端标签是不同的,而延迟写入在运行之间服务状态持续时会改变下一个运行的评分,但在隔离或验证重置后则不会。其次,在对十个公共协议的回顾中,我们发现所有协议都识别运行何时停止以及评分内容,而未完成的操作和将运行视为独立试验的证据记录得不够一致。最后,我们提出了一个开放效应记录,列出在终点之后可能仍然相关的操作或资源、它们的当前状态,以及它们是否可能改变评分结果或影响另一个运行。
cs.AI / 55 / 2608.14943

Skill Blocks: How Should an Agent Load Its Skill? A Caching-Correct Comparison of Pre-load, On-Demand Tool-Loading, Progressive Disclosure, and Hybrid

技能模块:代理应如何加载其技能?预加载、按需工具加载、渐进式披露和混合的缓存正确比较
Nakasuji, Hironobu
Abstract
Agent skills are often injected in full on every request, increasing token cost. We compare four content-preserving loading methods: Full, Skill Block, Reference, and Hybrid. Across SearchQA, SpreadsheetBench, ALFWorld, ScienceWorld, and SynthProc, we measure token usage using raw input for single-turn tasks and cache-correct effective input for multi-turn tasks. Results show no universal winner. Hybrid reduces input by 27.4% on SearchQA and 39.8% on SpreadsheetBench. On large multi-turn skills, Skill Block and Hybrid achieve substantial reductions, reaching 62.5% and 52.8% on ScienceWorld and 73.0% and 66.6% on SynthProc. ALFWorld shows smaller gains because procedures are short and repeatedly needed. Paired outcome tests detect no quality differences, though they do not establish equivalence. Overall, conditional loading is most beneficial when large portions of a skill are not needed on every turn.
Chinese Translation
代理技能通常在每次请求时完全注入,这增加了令牌成本。我们比较了四种内容保持的加载方法:完全加载、技能模块、引用和混合。在 SearchQA、SpreadsheetBench、ALFWorld、ScienceWorld 和 SynthProc 中,我们使用单轮任务的原始输入和多轮任务的缓存正确有效输入来测量令牌使用情况。结果显示没有普遍的赢家。在 SearchQA 上,混合方法减少了 27.4% 的输入,在 SpreadsheetBench 上减少了 39.8%。在大型多轮技能上,技能模块和混合方法实现了显著的减少,在 ScienceWorld 上分别达到 62.5% 和 52.8%,在 SynthProc 上达到 73.0% 和 66.6%。ALFWorld 的增益较小,因为过程较短且需要重复使用。配对结果测试未发现质量差异,尽管它们并未建立等效性。总体而言,当每轮并不需要技能的大部分内容时,条件加载最为有利。
cs.AI / 56 / 2608.14945

Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL

信任不足:代理强化学习中的在线自蒸馏影响校准
Lan, Qizhen, Xiao, Xi, Guan, Xiangchen, Fan, Mengchen, Lin, Moule, Choi, Jung Im, Zhu, Lijing
Abstract
On-policy self-distillation (OPSD) gives language agents dense token-level supervision from a privileged self-teacher on the policy's own trajectories. Existing methods allocate this supervision mainly by teacher trust, but trust does not reveal whether emphasizing a token supports the current policy objective. We call this the trust-utility mismatch and introduce Influence Calibration for Self-Distillation (ICSD). For each supervised token, ICSD measures the first-order response of its importance-weighted RL surrogate contribution to a teacher-directed output perturbation. Batch-adaptive calibration converts this non-stationary signal into a bounded allocation weight while preserving the original auxiliary-loss mass within each action turn. These detached weights affect only the distillation loss and require no additional model pass. Across ALFWorld, WebShop, and Search-QA, ICSD improves all matched aggregate metrics over trust-only allocation under Group Relative Policy Optimization (GRPO) and Group-in-Group Policy Optimization (GiGPO), across two model families spanning 1.5B to 7B. At 7B, it reaches 96.1% ALFWorld success and a WebShop score of 93.1. Frozen-batch analyses show that ICSD reduces teacher-supported mass assigned to objective-opposed tokens from 60.1% to 37.8% and raises cosine compatibility with the RL gradient by 0.192. A companion repository is avail- able at https://github.com/lanqz7766/Influence-Calibration-for-On-Policy-Self-Distillation-in-Agentic-RL.
Chinese Translation
在线自蒸馏(On-policy self-distillation,OPSD)为语言代理提供来自特权自教师的密集令牌级监督,基于策略自身的轨迹。现有方法主要通过教师信任来分配这种监督,但信任并不能揭示强调某个令牌是否支持当前的策略目标。我们称之为信任-效用不匹配,并引入自蒸馏影响校准(Influence Calibration for Self-Distillation,ICSD)。对于每个受监督的令牌,ICSD 衡量其重要性加权的强化学习替代贡献对教师指导的输出扰动的一级响应。批量自适应校准将这一非平稳信号转换为有界的分配权重,同时保持每个动作轮次内原始辅助损失的质量。这些独立的权重仅影响蒸馏损失,并且不需要额外的模型传递。在 ALFWorld、WebShop 和 Search-QA 上,ICSD 在 Group Relative Policy Optimization(GRPO)和 Group-in-Group Policy Optimization(GiGPO)下,改善了所有匹配的汇总指标,相较于仅基于信任的分配,涵盖了两个模型系列,参数范围从 1.5B 到 7B。在 7B 模型下,成功率达到 96.1% 的 ALFWorld 和 93.1 的 WebShop 分数。冻结批次分析表明,ICSD 将分配给与目标相对立的令牌的教师支持质量从 60.1% 降低到 37.8%,并将与强化学习梯度的余弦兼容性提高了 0.192。相关的代码库可在 https://github.com/lanqz7766/Influence-Calibration-for-On-Policy-Self-Distillation-in-Agentic-RL 获得。
cs.AI / 57 / 2608.14947

RETRACE: Resilience-Guided Trait-Conditioned Craving Estimation from Wearable Physiology in Opioid Use Disorder

RETRACE:基于可穿戴生理信号的抗压能力引导特征条件化阿片类药物渴求估计
Xiao, Yi, Sharma, Harshit, Bergen-Cico, Dessa, Salekin, Asif
Abstract
Detecting opioid craving from wearable physiological signals is critical yet difficult, with the potential to support proactive interventions for individuals with opioid use disorder (OUD). This challenge is especially pronounced under subject-independent evaluation because craving is subjective, heterogeneous, and often physiologically entangled with stress. Our empirical analysis shows that stress elicits strong and reproducible autonomic responses, while craving-related signals are weaker, sparse, and largely embedded within stress-related physiology. We further show that psychological resilience, which shapes stress regulation and craving vulnerability, is not reliably observable from short-term wearable windows, but can be captured through reusable subject-level proxies, including post-stress heart-rate recovery and autobiographical memory recall.Motivated by these findings, we introduce RETRACE, a resilience-guided trait-conditioned framework for subject-independent craving estimation from wearable physiology. RETRACE reframes craving detection as trait-conditioned physiological interpretation: rather than assuming the same physiological pattern has the same meaning across individuals, it uses resilience-related subject context to guide inference. Technically, RETRACE introduces a novel dual-encoder design that separates generalizable stress physiology from subject-specific craving interpretation. It combines a frozen stress-pretrained encoder with a resilience-conditioned craving encoder, using feature-level gating and representation-level fusion to enable lightweight personalization without target-user craving labels or per-user retraining. We evaluate RETRACE on a novel multimodal OUD dataset containing wearable physiology, stress and craving annotations, and autobiographical narratives. Under LOSO setup, RETRACE achieves up to 7% absolute improvement over the strongest baseline
Chinese Translation
从可穿戴生理信号中检测阿片类药物渴求至关重要,但也非常困难,这有助于为阿片类药物使用障碍(OUD)患者提供主动干预。这一挑战在独立于受试者的评估中尤为明显,因为渴求是主观的、异质的,并且通常与压力在生理上交织在一起。我们的实证分析表明,压力会引发强烈且可重复的自主神经反应,而与渴求相关的信号则较弱、稀疏,并且在很大程度上嵌入于与压力相关的生理信号中。我们进一步表明,塑造压力调节和渴求脆弱性的心理抗压能力并不能通过短期可穿戴窗口可靠观察到,但可以通过可重复的个体级代理捕获,包括压力后的心率恢复和自传式记忆回忆。基于这些发现,我们提出了RETRACE,一个基于抗压能力引导的特征条件化框架,用于从可穿戴生理信号中独立于受试者的渴求估计。RETRACE将渴求检测重新框架为特征条件化的生理解释:它并不假设相同的生理模式在不同个体中具有相同的意义,而是利用与抗压能力相关的个体背景来指导推断。在技术上,RETRACE引入了一种新颖的双编码器设计,将可推广的压力生理学与特定个体的渴求解释分开。它结合了一个冻结的压力预训练编码器和一个抗压能力条件化的渴求编码器,使用特征级门控和表示级融合来实现轻量级个性化,而无需目标用户的渴求标签或每个用户的重新训练。我们在一个包含可穿戴生理信号、压力和渴求注释以及自传叙述的新型多模态OUD数据集上评估了RETRACE。在LOSO设置下,RETRACE在最强基线之上实现了高达7%的绝对提升。
cs.AI / 58 / 2608.14953

T-LLM Compiler: Trusted LLM-based Code Optimization and Verification Framework

T-LLM 编译器:基于可信 LLM 的代码优化与验证框架
Fazel, Zahra, Gamage, Sunanda, Bagi, Shayan Shirahmad Gale, Ashouri, Amir H., Czajkowski, Tomasz S., Chan, Bryan, Azimi, Reza, Gao, Yaoqing
Abstract
Recent advances in Large Language Models (LLMs) have opened opportunities to apply high-level code transformations to the field of code optimization, and it has since emerged as one of the most fundamental tasks for LLMs to perform; however, at present, LLMs struggle to apply wide-ranging code optimization tasks due to both the complexity of the code and the inability to independently verify the correctness of the transformations. In this paper, we present the Trusted LLM (T-LLM) Compiler, which proposes an advancement in compiler technology through a collaborative effort involving high-level LLM code transformations, traditional compilers, and verification tools. Experimental results reveal that it can significantly improve code correctness when tested on a set of PolyBench/C benchmarks. Our approach facilitates iterative code optimization efforts with verification strategies that enable corrective actions. Through this approach, T-LLM Compiler achieves code optimization accuracy of up to 83.3% and a speedup of up to 16.1\% on the PolyBench/C benchmarks, with the transformed code reaching an average of 26.7% speedup wrt standard baselines. Additionally, we release the project's source code to the open-source community.
Chinese Translation
近年来,大型语言模型(LLMs)的进步为高层次代码转换在代码优化领域的应用开辟了机会,并且这已成为 LLMs 执行的最基本任务之一;然而,目前 LLMs 在应用广泛的代码优化任务时面临挑战,原因在于代码的复杂性以及无法独立验证转换的正确性。本文提出了可信 LLM(T-LLM)编译器,旨在通过高层次 LLM 代码转换、传统编译器和验证工具的协作,推动编译技术的进步。实验结果表明,在一组 PolyBench/C 基准测试中,该编译器能够显著提高代码的正确性。我们的方法促进了带有验证策略的迭代代码优化工作,从而能够采取纠正措施。通过这种方法,T-LLM 编译器在 PolyBench/C 基准测试中实现了高达 83.3% 的代码优化准确率和高达 16.1% 的加速,转换后的代码相较于标准基线平均实现了 26.7% 的加速。此外,我们将项目的源代码发布给开源社区。
cs.AI / 59 / 2608.14974

Demand-Driven Vertiport Siting and Discrete-Event Fleet Simulation for On-Demand Urban Air Mobility Network Design

需求驱动的垂直机场选址与离散事件车队仿真在按需城市空中出行网络设计中的应用
Saghazadeh, Hossein Z., Ayalew, Yonas, Ahmari, Reza, Kebria, Parham, Homaifar, Abdollah
Abstract
This paper presents a demand-driven framework for on-demand Urban Air Mobility (UAM) network design that links vertiport siting, fleet simulation, and door-to-door travel-time feasibility. Demand is estimated from commuter and passenger activity data, converted into spatial trip-end points, and clustered using K-means to generate candidate vertiport locations. Candidate networks are screened using range and minimum station-spacing constraints, then evaluated with a discrete-event simulation that models multi-vehicle dispatch, deadhead relocation, battery swaps, and service regularity. Flight time and energy consumption are computed using a point-mass eVTOL performance model. In a Greater Los Angeles case study, the preferred design expands from four stations and four eVTOLs at low demand to sixteen stations and twelve eVTOLs at the highest tested demand level. Results show that larger fleets improve completion time and vehicle-arrival regularity but do not eliminate deadhead flights, indicating that spatial demand imbalance remains an operational burden. The travel-time savings analysis further suggests that UAM is most defensible for longer or congestion-heavy trips where sufficient non-flight time remains after accounting for flight time.
Chinese Translation
本文提出了一种需求驱动的框架,用于按需城市空中出行(UAM)网络设计,该框架将垂直机场选址、车队仿真和门到门旅行时间可行性联系起来。需求通过通勤和乘客活动数据进行估算,转换为空间出行终点,并使用K均值聚类生成候选垂直机场位置。候选网络通过范围和最小站间距约束进行筛选,然后通过离散事件仿真进行评估,该仿真模型模拟多车辆调度、空驶调动、电池更换和服务规律性。飞行时间和能耗使用点质量eVTOL性能模型进行计算。在大洛杉矶地区的案例研究中,首选设计从低需求下的四个站点和四架eVTOL扩展到最高测试需求水平下的十六个站点和十二架eVTOL。结果表明,较大的车队提高了完成时间和车辆到达规律性,但并未消除空驶飞行,表明空间需求不平衡仍然是一个运营负担。旅行时间节省分析进一步表明,UAM在较长或拥堵严重的行程中最具合理性,在考虑飞行时间后,仍有足够的非飞行时间。
cs.AI / 60 / 2608.14992

Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5

工具结果是否比纯文本更具权威性?关于在合成任务中虚假声明采纳的三项前瞻性研究,使用 Claude Opus 5
Bronder, Justin
Abstract
Language-model systems increasingly read from stores they also write to, so a claim that was merely written earlier can return looking retrieved. We tested whether the message package carrying an unsupported assignment changes which answer a model gives in a synthetic lookup task. Claude Opus 5 selected a color code for a named item or abstained. In an exploratory four-arm study, false-code adoption was 0/24 with no target claim, 0/22 scorable trials when a prior assistant assertion named the target, 14/24 when a tool-result record named it, and 15/24 when that result used a ten-field metadata wrapper that marked it unchecked. The tool-result arm selected the record's code in 11/12 supported trials and 14/24 unsupported trials, ruling out a fixed output-token bias while leaving substantial planted-token heterogeneity. A document-preregistered replication reproduced the tool-result versus assistant-assertion gap, 7/24 against 0/24, one-sided Fisher exact p = 0.0047. The tool-result rate nevertheless fell from 14/24 to 7/24 across runs made four days apart. A second preregistered study gave the earlier comparison a live text control: both records were announced in advance and placed in the same final user turn, then target binding was swapped between the linked tool result and later inline JSON. Inline text was sufficient for false-code adoption in 60/60 trials; the tool-result condition produced 57/60, so the registered result-first superiority criterion failed, p = 1. The result does not show that tool results have no effect. It shows that native tool-result placement was not necessary and that this experiment did not find greater behavioral weight for the result package than for announced inline text. The findings concern a single model on one synthetic task template, accessed through one API.
Chinese Translation
语言模型系统越来越多地从它们也写入的存储中读取,因此一个早先仅被写下的声明可能看起来像是被检索的。我们测试了承载不支持的任务的消息包是否会改变模型在合成查找任务中给出的答案。Claude Opus 5 为一个命名项目选择了颜色代码或选择不作答。在一项探索性的四组研究中,当没有目标声明时,虚假代码采纳为 0/24;当先前的助手声明命名目标时,得分试验为 0/22;当工具结果记录命名目标时,采纳率为 14/24;而当该结果使用标记为未检查的十字段元数据包装时,采纳率为 15/24。工具结果组在 12 个支持试验中选择了记录的代码 11 次,在 24 个不支持试验中选择了 14 次,排除了固定输出标记偏差,同时留下了相当大的植入标记异质性。一项文档预注册的复制实验重现了工具结果与助手声明之间的差距,7/24 对比 0/24,单侧 Fisher 精确 p = 0.0047。然而,工具结果率在四天间隔的运行中从 14/24 降至 7/24。第二项预注册研究为早期比较提供了一个实时文本控制:两个记录提前宣布并放置在同一最终用户回合中,然后目标绑定在链接的工具结果和后来的内联 JSON 之间交换。内联文本在 60/60 次试验中足以导致虚假代码采纳;工具结果条件产生了 57/60,因此注册的结果优越性标准未能成立,p = 1。结果并不表明工具结果没有影响。它表明本土工具结果的放置并不是必要的,并且本实验没有发现结果包相较于宣布的内联文本具有更大的行为权重。研究结果涉及一个模型在一个合成任务模板上的表现,通过一个 API 访问。
cs.AI / 61 / 2608.15018

S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices

S2-MoE:在边缘设备上实现高效自我推测解码的专家混合模型
Huang, Haochen, Qiu, Shengxuan, Li, Meng
Abstract
Deploying large language models (LLMs) for inference on edge devices is challenging due to severe memory and bandwidth constraints. While speculative decoding and Mixture-of-Experts (MoE) have been proposed to improve inference efficiency, naively combining them often incurs excessive verification overhead and poor expert reuse, limiting their effectiveness in memory-bound edge settings. In this work, we propose S2-MoE, an efficient self-speculative decoding framework for MoE inference on edge devices. S2-MoE reduces redundant verification through routing-aware adaptive speculative expansion, improves verification efficiency with reuse-aware expert gating, and aligns draft and target execution via shared context. Implemented in llama.cpp, S2-MoE achieves up to 5.3x speedup (about 2.0x on average) over standard autoregressive de?coding across diverse MoE models and datasets on edge devices.Code is available at https://github.com/angerybob/S2-MoE.
Chinese Translation
在边缘设备上进行大规模语言模型(LLMs)推理面临着严重的内存和带宽限制。尽管已经提出了推测解码和专家混合模型(MoE)以提高推理效率,但简单地将它们结合往往会导致过高的验证开销和较差的专家重用,从而限制了它们在内存受限的边缘环境中的有效性。在本研究中,我们提出了S2-MoE,一个针对边缘设备上MoE推理的高效自我推测解码框架。S2-MoE通过路由感知的自适应推测扩展减少冗余验证,通过重用感知的专家门控提高验证效率,并通过共享上下文对草稿和目标执行进行对齐。在llama.cpp中实现的S2-MoE在各种MoE模型和数据集上相较于标准自回归解码实现了最高5.3倍的加速(平均约2.0倍)。代码可在https://github.com/angerybob/S2-MoE获取。
cs.AI / 62 / 2608.15022

Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form

聚集,而非接纳:注意力如何将潜在变量转化为可言表的形式
Mazaheri, Parsa
Abstract
Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly. What causes a representation to enter that form is open, and the word workspace invites an admission story: a gate that decides what gets in. Testing it on open-weight models with Jacobian lenses, over a benchmark whose five arms share an identical context, we find no gate where it predicts one. Demand raises a concept's lens visibility beyond what applying an operator to a supplied value produces: +0.050 [+0.045, +0.057] in percentile rank on our primary checkpoint, positive on all four we measure, though that arm answers at ceiling and the accuracymatched contrast is stronger under that readout. At the same time one shared linear map decodes the variable from every arm, the control included, at 6.4-9.0x its selection-corrected floor. What produces the later readable form at the queried position is attention-mediated gathering inside a mid-depth window: separating patch depth from readout depth puts transport there at least 17x above anywhere shallower under non-saturating readouts, with no tested MLP output contributing positively inside it. Under the saturating percentile rank the same grid does not localise the window, which is a fact about that measure. An arm that needs the variable for nothing concentrates sevenfold less, so the window is demand-specific. That window has two measured edges, a survival failure below and destruction above, and it falls at the same fractional depth in a 64-layer hybrid and a 62-layer dense model from another family. We localise where the variable is installed and read, not the route from the passage, which transports nothing. But the readout is not a calibrated measure of use: three components move it to within 12% of one another and differ 7.4x in what they do to the answer.
Chinese Translation
语言模型以一种可以报告的形式持有潜在量,并且当任务要求灵活地重用某一量时,该形式中存在的量更多。是什么导致一种表征进入这种形式尚不明确,而“词汇工作区”则引入了一个接纳的故事:一个决定什么能够进入的门。我们在具有雅可比视角的开放权重模型上进行测试,基于一个五个臂共享相同上下文的基准,我们发现没有预测到的门。需求使得一个概念的视角可见性超出将操作符应用于提供值所产生的结果:在我们的主要检查点上,百分位排名提高了+0.050 [+0.045, +0.057],在我们测量的四个检查点上均为正值,尽管该臂的回答已达到上限,而准确度匹配的对比在该读出下更强。同时,一个共享的线性映射从每个臂(包括控制臂)解码变量,其选择校正后的底线为6.4-9.0倍。产生在查询位置可读形式的因素是通过中深度窗口的注意力介导聚集:将补丁深度与读出深度分开,使得在非饱和读出下的传输至少比任何更浅的地方高出17倍,而在其中没有经过测试的多层感知器输出有正向贡献。在饱和百分位排名下,相同的网格无法定位窗口,这一事实与该测量相关。一个不需要变量的臂的集中度降低了七倍,因此该窗口是需求特定的。该窗口有两个测量边缘,生存失败在下方,破坏在上方,并且在一个64层混合模型和另一个家族的62层密集模型中,落在相同的分数深度。我们定位变量被安装和读取的位置,而不是从通道传输的路径,该路径没有传输任何东西。但读出并不是使用的校准测量:三个组成部分使其相互之间的差距缩小到12%以内,而在对答案的影响上差异达到7.4倍。
cs.AI / 63 / 2608.15041

LLM-Based Hierarchical Coordinated Control with Continuation-Aware Policy Learning

基于大语言模型的分层协调控制与持续感知策略学习
He, Changhong, Gao, Jinda, Liu, Xinkuan, Zhang, Le, Luo, Xizi, Mei, Yu
Abstract
Coordinating multiple interacting units in complex engineering systems is challenging when system interactions are difficult to model, operational information is heterogeneous, and low-level actions must satisfy strict constraints. We propose an LLM-based hierarchical framework in which the LLM coordinates interacting units based on heterogeneous operational context, while task-specific controllers or optimizers generate executable and constraint-aware actions. We further introduce Continuation-Aware GRPO to capture the consequences of coordination decisions over subsequent control intervals. Rather than judging a decision only by its immediate outcome, the method also evaluates how the system evolves afterward under the current policy. We validate the framework on multi-ramp traffic control and virtual power plant (VPP) energy management, using simplified system models for training and more realistic simulators for evaluation. Across both tasks, the proposed method consistently outperforms direct task-specific control and optimization, end-to-end reinforcement learning, rule-based and RL-based hierarchical coordination, and prompting-only LLM coordinators, demonstrating the value of heterogeneous-context reasoning, hierarchical execution, and continuation-aware policy learning.
Chinese Translation
在复杂工程系统中协调多个相互作用的单元是一项具有挑战性的任务,尤其是在系统交互难以建模、操作信息异构以及低级动作必须满足严格约束的情况下。我们提出了一种基于大语言模型(LLM)的分层框架,其中LLM根据异构操作上下文协调相互作用的单元,同时任务特定的控制器或优化器生成可执行且符合约束的动作。我们进一步引入了持续感知的广义鲁棒优化(Continuation-Aware GRPO),以捕捉协调决策在后续控制区间的后果。该方法不仅通过即时结果来评估决策,还评估在当前策略下系统如何随时间演变。我们在多坡道交通控制和虚拟电厂(VPP)能源管理中验证了该框架,使用简化的系统模型进行训练,并在更现实的模拟器上进行评估。在这两个任务中,所提出的方法始终优于直接的任务特定控制和优化、端到端强化学习、基于规则和基于强化学习的分层协调,以及仅依赖提示的LLM协调器,展示了异构上下文推理、分层执行和持续感知策略学习的价值。
cs.AI / 64 / 2608.15043

SCOPE: Score-Isolated Agentic Optimization for Video World Models

SCOPE:视频世界模型的分数孤立代理优化
Jiang, Yuhua, Wang, Jiaming, Liu, Qingbin, Gao, Feifei
Abstract
Video world models are increasingly used as simulators for planning and embodied decision making, yet improving them at inference time introduces a subtle evaluation problem: prompts, samplers, verifiers, and selectors may evolve together, making it difficult to attribute gains or prevent held-out feedback from shaping the final policy. We introduce \scope (\emph{\scopefullname}), a framework for auditable inference-time adaptation of frozen video world models. \scope represents external controls as a typed state, updates this state only through bounded changes supported by development evidence, and freezes the resulting policy before held-out evaluation. On Physics-IQ benchmark, \scope improves over the exact frozen base by $+14.24$ (95\% CI $[+8.10,+21.23]$). Controlled ablations further identify gains from scene specification, sampling, and learned selection, while the margin over the strongest matched agentic baseline remains unresolved. Cross-backbone and prospective evaluations reveal a complementary result: useful inference-time updates exist, but their benefits do not transfer uniformly across models and settings. Together, these findings suggest that reliable inference-time adaptation requires not only better proposals, but also a principled mechanism for deciding which updates should become part of the deployed system. Code is available at https://github.com/YuhuaJiang2002/SCOPE.
Chinese Translation
视频世界模型越来越多地被用作规划和具身决策的模拟器,但在推理时改进它们会引入一个微妙的评估问题:提示、采样器、验证器和选择器可能会共同演变,这使得很难归因于收益或防止保留反馈影响最终策略。我们提出了 extsc{SCOPE}( extit{Score-Isolated Agentic Optimization}),这是一个用于可审计的推理时适应冻结视频世界模型的框架。 extsc{SCOPE} 将外部控制表示为类型化状态,仅通过开发证据支持的有限变化更新该状态,并在保留评估之前冻结结果策略。在 Physics-IQ 基准测试中, extsc{SCOPE} 相比于精确的冻结基线提高了 $+14.24$(95 ext{% CI} $[+8.10,+21.23]$)。受控消融实验进一步识别出来自场景规范、采样和学习选择的收益,而与最强匹配的代理基线之间的差距仍然未解决。跨骨干网和前瞻性评估揭示了一个互补的结果:存在有用的推理时更新,但其好处并不均匀地转移到不同的模型和设置中。综合来看,这些发现表明,可靠的推理时适应不仅需要更好的提议,还需要一个原则性机制来决定哪些更新应成为部署系统的一部分。代码可在 https://github.com/YuhuaJiang2002/SCOPE 获取。
cs.AI / 65 / 2608.15052

Andy: A Mathematical Agent for Rigorous Proof and Autonomous Research

Andy:一个用于严谨证明和自主研究的数学智能体
Wang, Zi'an
Abstract
Andy is an autonomous mathematical research agent that solves and verifies submitted problems, formulates new research problems, and constructs rigorous proofs. It separates proof generation from correctness evaluation and supports knowledge acquisition, targeted revision, and multistage verification. This paper illustrates the workflow using a published result on self-triggered impulsive consensus as a starting point. Andy formulates a global exponential leader-follower synchronization problem for delayed heterogeneous networks with switching communication topologies. The proposed hybrid control combines self-triggered impulses with execution delay and recovery-phase continuous feedback. After each delayed impulse, this feedback cancels the delayed error channel during a recovery window. Sufficient conditions for global exponential synchronization are established, and Zeno behavior is excluded for both the sampling and impulse sequences. A numerical example confirms the result. This case demonstrates Andy's ability to learn from existing results, formulate meaningful research problems, and develop and verify rigorous proofs.
Chinese Translation
Andy 是一个自主的数学研究智能体,能够解决和验证提交的问题,提出新的研究问题,并构建严谨的证明。它将证明生成与正确性评估分开,并支持知识获取、针对性修订和多阶段验证。本文以一个关于自触发冲动共识的已发布结果为起点,阐述了工作流程。Andy 为具有切换通信拓扑的延迟异构网络制定了一个全局指数型领导-跟随同步问题。所提出的混合控制结合了自触发冲动、执行延迟和恢复阶段的连续反馈。在每次延迟冲动后,该反馈在恢复窗口期间取消了延迟误差通道。建立了全局指数同步的充分条件,并排除了采样和冲动序列的 Zeno 行为。数值示例验证了这一结果。该案例展示了 Andy 从现有结果中学习、提出有意义的研究问题以及开发和验证严谨证明的能力。
cs.AI / 66 / 2608.15055

TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning

TAHB:文本属性超图学习的综合基准
Kang, David Yoon Suk, Kim, JungHyun, Jeon, Juhyun, Kim, Sang-Wook
Abstract
Hypergraphs effectively model higher-order groupwise relationships beyond pairwise interactions, while pretrained language models (PLMs) and large language models (LLMs) provide rich semantic understanding from textual attributes. However, research on combining language models with hypergraph learning remains limited due to the lack of public text-attributed hypergraph benchmarks. To address this limitation, we present TAHB (Text-Attributed Hypergraph Benchmark), the first public benchmark integrating hypergraph structures and raw textual attributes. TAHB contains 10 real-world datasets from four domains - e-commerce, academia, movies, and politics networks - enabling systematic evaluation of text-aware hypergraph representation learning. Experimental results show that TAHB preserves key structural properties of real-world hypergraphs and consistently reproduces performance tendencies observed in existing benchmarks. Furthermore, experiments under both LLM-as-Enhancer and LLM-as-Predictor settings demonstrate that LLM-enhanced textual semantics improve hypergraph learning performance, while structural and textual information jointly provide the best setting for LLM-based prediction. Our benchmark provides a foundation for future research at the intersection of hypergraph learning and language models.
Chinese Translation
超图有效地建模超出成对交互的高阶群体关系,而预训练语言模型(PLMs)和大型语言模型(LLMs)则从文本属性中提供丰富的语义理解。然而,由于缺乏公共的文本属性超图基准,结合语言模型与超图学习的研究仍然有限。为了解决这一限制,我们提出了TAHB(文本属性超图基准),这是第一个将超图结构与原始文本属性相结合的公共基准。TAHB包含来自电子商务、学术界、电影和政治网络四个领域的10个真实世界数据集,使得文本感知超图表示学习的系统评估成为可能。实验结果表明,TAHB保留了真实世界超图的关键结构特性,并一致地再现了在现有基准中观察到的性能趋势。此外,在LLM作为增强器和LLM作为预测器的设置下的实验表明,LLM增强的文本语义提高了超图学习性能,而结构信息和文本信息共同提供了基于LLM预测的最佳设置。我们的基准为超图学习与语言模型交叉领域的未来研究奠定了基础。
cs.AI / 67 / 2608.15056

GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG

GraphLoom:用于多模态知识图谱增强生成的可靠性校准图证据路由
Ali, Zafar, Khan, Asad, Malik, Aalia, Kefalas, Pavlos
Abstract
Multimodal retrieval-augmented generation (RAG) systems often rely on long unstructured contexts or aggressively expanded evidence graphs, which can introduce noisy evidence, weaken multi-hop reasoning, and increase unsupported generation. We present GraphLoom, a reliability-calibrated multimodal knowledge-graph RAG framework for compact and faithful evidence routing. Given a question and its associated multimodal input, GraphLoom constructs an instance-level multimodal knowledge graph from grounded scene descriptions, extracted relational triples, and external commonsense knowledge. Instead of injecting all retrieved evidence into the generator, GraphLoom performs reliability-aware subgraph retrieval with bounded expansion and selectively routes high-utility evidence through hierarchical graph memory slots and joint graph-sequence attention in a frozen language model. To improve robustness in complex reasoning settings, GraphLoom further combines interleaved retrieval with budgeted corrective retrieval, enabling adaptive multi-hop evidence refinement under noisy retrieval conditions. We evaluate GraphLoom on ScienceQA, MultiModalQA, and OK-VQA, including large distractor evidence pools that approximate noisy external knowledge retrieval. Experimental results show consistent gains in answer quality and evidence faithfulness over strong multimodal RAG, graph-retrieval, and open-source vision-language baselines, with improved retrieval quality on MultiModalQA and stable performance under noisy evidence pools. Additional analyses using MiniCheck-based verification, human evaluation, and latency profiling show that reliability-calibrated graph evidence routing provides an effective alternative to long-context multimodal evidence injection.
Chinese Translation
多模态检索增强生成(RAG)系统通常依赖于长的非结构化上下文或大幅扩展的证据图,这可能引入噪声证据、削弱多跳推理并增加不支持的生成。我们提出了GraphLoom,一种可靠性校准的多模态知识图谱RAG框架,用于紧凑且真实的证据路由。给定一个问题及其相关的多模态输入,GraphLoom从基础场景描述、提取的关系三元组和外部常识知识构建实例级别的多模态知识图谱。GraphLoom并不将所有检索到的证据注入生成器,而是通过有限扩展执行可靠性感知的子图检索,并通过层次图内存槽和冻结语言模型中的联合图序列注意力选择性地路由高效用证据。为了提高复杂推理环境中的鲁棒性,GraphLoom进一步结合交错检索与预算纠正检索,使得在噪声检索条件下能够自适应地进行多跳证据精炼。我们在ScienceQA、MultiModalQA和OK-VQA上评估GraphLoom,包括近似噪声外部知识检索的大型干扰证据池。实验结果显示,GraphLoom在答案质量和证据真实度上相较于强大的多模态RAG、图检索和开源视觉-语言基线具有一致的提升,并在MultiModalQA上改善了检索质量,在噪声证据池下表现稳定。使用基于MiniCheck的验证、人类评估和延迟分析的额外分析表明,可靠性校准的图证据路由为长上下文多模态证据注入提供了一种有效的替代方案。
cs.AI / 68 / 2608.15064

LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents

LongDocBench:长文档中目录层次和上下文关系恢复的基准测试
Zou, Yuefeng, Lu, Yichen, Yang, Jingxiao, Fu, Bingtao, Zhang, Gaoyang, Bai, Xiongfei, Chen, Tian, Qi, Xiang
Abstract
Parsing visual documents into machine-readable representations is fundamental to document intelligence. Existing benchmarks focus on page-level element recognition, reading order, formula recognition, and table structure. Long documents, however, also require document-level structure recovery. This includes reconstructing cross-page table-of-contents (TOC) hierarchies and identifying typed links from tables and figures to their captions, notes, and sources, often in one-to-many form. Because these structures are covered only partially or subsumed within broader parsing protocols, existing benchmarks cannot directly evaluate two key document-level tasks: \emph{Table-of-Contents Hierarchy Recovery} and \emph{Contextual Relationship Recovery}. To benchmark these two tasks, we introduce \textsc{LongDocBench}, comprising 85 real-world financial reports, textbooks, and academic papers spanning 2,582 pages, with up to 105 pages per document. It provides human-verified annotations for 3,937 heading nodes (mean node depth 3.55; maximum depth 9) and 3,258 contextual relationships annotated across 2,680 table and figure objects. We further evaluate both the downstream utility and recoverability of these structures. Long-document question-answering experiments show that human-verified TOC hierarchies and contextual relationships improve reasoning, with their combination providing complementary benefits. Meanwhile, representative document parsers remain limited on both recovery tasks despite strong page-level performance. To support further progress, we publicly release \textsc{LongDocBench} and its evaluation protocol and reproducible testbed for advancing document-level structure recovery in long documents.
Chinese Translation
将视觉文档解析为机器可读的表示形式是文档智能的基础。现有的基准测试主要集中在页面级元素识别、阅读顺序、公式识别和表格结构上。然而,长文档还需要进行文档级结构恢复。这包括重建跨页的目录(TOC)层次结构,以及识别从表格和图形到其标题、注释和来源的类型链接,通常是多对一的形式。由于这些结构仅部分覆盖或被更广泛的解析协议所包含,现有基准无法直接评估两个关键的文档级任务: extit{目录层次恢复}和 extit{上下文关系恢复}。为了对这两个任务进行基准测试,我们引入了 extsc{LongDocBench},该基准包含85份真实的财务报告、教科书和学术论文,覆盖2582页,每份文档最多可达105页。它为3937个标题节点(平均节点深度3.55;最大深度9)和2680个表格及图形对象中标注的3258个上下文关系提供了人工验证的注释。我们进一步评估了这些结构的下游效用和可恢复性。长文档问答实验表明,经过人工验证的TOC层次和上下文关系提高了推理能力,两者的结合提供了互补的好处。同时,尽管在页面级表现强劲,代表性的文档解析器在这两个恢复任务上仍然有限。为了支持进一步的进展,我们公开发布了 extsc{LongDocBench}及其评估协议和可重复的测试平台,以推动长文档中文档级结构的恢复。
cs.AI / 69 / 2608.15065

Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning

思维漏斗:通过早期投票和回滚剪枝实现高效的测试时间扩展
Park, Chanhee, Han, Sungbin, Yoon, Jeongho, Hong, Seongtae, Lim, Heuiseok
Abstract
Large Reasoning Models produce diverse, sometimes inconsistent answers across repeated queries on the same problem, so multi-sample inference is a prerequisite for reliable deployment. Majority voting at k rollouts is the standard solution and the de facto accuracy target for this regime, but it is prohibitively expensive at the scale LRMs require. We introduce Funnel of Thoughts (FoT), an inference-time method that preserves the full 32-trajectory voted accuracy while halving its attention FLOPs, a 28.8% reduction in full-model inference cost. Across 115K reasoning trajectories from six LRMs, we find that unproductive trajectories often reveal themselves through repeated hesitation markers such as "Wait", "Actually", and "perhaps." These trajectories are less likely to reach the correct answer and consume disproportionate attention FLOPs, degenerating into no-answer loops in the worst case. Built on this training-free lexical signal, FoT identifies the vocabulary that captures these pathological patterns and prunes affected trajectories before completion, reducing online generation attention FLOPs by 56.1% and wall time by 37.6% without any additional model inference; the same signal transfers without retuning across held-out architectures and out-of-domain tasks.
Chinese Translation
大型推理模型在对同一问题的重复查询中产生多样且有时不一致的答案,因此多样本推理是可靠部署的前提。k次回滚的多数投票是这一领域的标准解决方案和事实上的准确性目标,但在大型推理模型(LRMs)所需的规模下,这一方法成本过高。我们提出了思维漏斗(Funnel of Thoughts, FoT),这是一种推理时方法,能够在保持完整的32轨迹投票准确性的同时,将注意力浮点运算(FLOPs)减半,从而实现28.8%的全模型推理成本降低。在来自六个LRMs的115K推理轨迹中,我们发现无效轨迹常常通过重复的犹豫标记(如“等一下”、“实际上”和“或许”)显露出来。这些轨迹不太可能达到正确答案,并且消耗不成比例的注意力FLOPs,在最坏情况下会退化为无答案循环。基于这一无训练的词汇信号,FoT识别出捕捉这些病态模式的词汇,并在轨迹完成之前剪枝受影响的轨迹,从而在不增加任何模型推理的情况下,将在线生成的注意力FLOPs减少56.1%,墙面时间减少37.6%;同样的信号在未重新调优的情况下可以转移到保留架构和域外任务中。
cs.AI / 70 / 2608.15071

Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents

Evo-Harness:自我进化代理的上下文到技能编译
Wei, Tianxin, Shi, Zhan, Lin, Minhua, He, Bing, Liu, Zewen, Sang, Yisi, Bei, Yuanchen, Ning, Xuying, Zou, Jiaru, Li, Ting-Wei, Lin, Xiao, Zhao, Yanjun, Wang, Chi, Dumoulin, Benoit, Wang, Dakuo, He, Jingrui, Lu, Hanqing
Abstract
Learning from experience is critical for developing capable, self-improving large language model (LLM) agents. Existing methods typically extract knowledge from accumulated trajectories via reflection, memory, rules, or skills. However, agents in realistic environments continuously encounter novel tasks, often offering only a one-shot opportunity to improve. These executions yield rich but highly noisy contexts, entangling broadly useful lessons with task-specific artifacts. Critically, prior works rarely validate their effectiveness on complex real-world tasks or isolate the underlying drivers of improvement. To address these gaps, we formulate online harness learning, where a frozen agent improves by continually updating a structured harness across sequential tasks. This formulation enables a systematic study of key self-improvement factors through our proposed Evo-Harness. At its core, context-to-harness skill compilation distills noisy, single-shot executions into reusable skill harnesses for cross-domain and topic-level adaptation. To demonstrate the efficacy of one-shot skill compilation, we evaluate across five realistic benchmarks (TerminalBench2, SWE-bench, CL-Bench, -bench, WebArena-Infinity). Our extensive analysis demonstrates the effectiveness of Evo-Harness and provides a principled understanding of how LLM agents can effectively learn on the fly. Our code is available at https://github.com/A-EVO-Lab/a-evolve/tree/release/evo-harness.
Chinese Translation
从经验中学习对于开发能够自我改进的大型语言模型(LLM)代理至关重要。现有方法通常通过反思、记忆、规则或技能从积累的轨迹中提取知识。然而,现实环境中的代理不断遇到新任务,通常只提供一次性改进的机会。这些执行产生丰富但高度噪声的上下文,将广泛有用的经验与任务特定的伪影交织在一起。关键是,先前的研究很少在复杂的现实任务上验证其有效性或隔离改进的潜在驱动因素。为了解决这些问题,我们提出了在线技能学习的概念,其中一个固定的代理通过不断更新结构化的技能框架来改善其在连续任务中的表现。这一构想使我们能够通过所提出的Evo-Harness系统地研究关键的自我改进因素。在其核心,上下文到技能的编译将噪声较大的一次性执行提炼为可重用的技能框架,以便进行跨领域和主题级别的适应。为了证明一次性技能编译的有效性,我们在五个现实基准(TerminalBench2、SWE-bench、CL-Bench、-bench、WebArena-Infinity)上进行了评估。我们的广泛分析展示了Evo-Harness的有效性,并提供了对LLM代理如何能够有效地即时学习的原则性理解。我们的代码可在 https://github.com/A-EVO-Lab/a-evolve/tree/release/evo-harness 获取。
cs.AI / 71 / 2608.15082

Beyond Thresholds: A Quality-Aware Decision Intelligence Framework for Cold Chain IoT Systems

超越阈值:一种面向质量的冷链物联网系统决策智能框架
Sofat, Aashna, Sodhi, Balwinder
Abstract
Cold chain logistics has advanced technologically, yet most deployed systems remain reactive monitors, not decision-making agents: thresholds trigger alerts, but nothing relates violations to cumulative product degradation or converts degradation signals into logistics decisions. We address this gap with a Quality-Aware Decision Intelligence (QADI) framework combining three capabilities: a structured quality state representation, $S_q = [L, Q, U, R]$ -- remaining shelf life, degradation rate, estimation uncertainty, and operational risk, all derived and computable from the framework equations; a hybrid quality modeling layer combining physics-based microbial kinetics with a data-driven correction term; and a reasoning layer built on Microsoft Phi-4~\cite{Phi4} with retrieval-augmented generation over a structured domain knowledge base. We benchmark against five baselines -- threshold monitoring, physics-only, physics-plus-noise, optimisation-based decisions, and a rule-based expert system -- across eight cold chain scenarios, using pasteurised milk as the primary case, with ground truth shelf-life drawn from published dairy studies~\cite{Singh1994, Smigic2015} independent of our model. Comparisons use Wilcoxon signed-rank tests with Holm correction. Across milk and broccoli scenarios, the framework attains mean absolute shelf-life error of 7.2 hours (versus 30.9 hours, physics-only; $p<0.001$), spoilage rate of 14.5% (versus 16.6%, physics-only and rule-based; p=0.08), and oracle-optimal decisions in 99.5% of scenarios. Removing the LLM reasoning component drops optimality to 45.5% ($p<0.001$). Expert-rated explanation quality reaches 83% ($\kappa = 0.71$). Ablations show hybrid modeling and LLM reasoning contribute distinct gains, while RAG retrieval mainly drives explanation quality. Code: https://bit.ly/4d6t44C.
Chinese Translation
冷链物流在技术上取得了进展,但大多数已部署的系统仍然是反应式监控器,而非决策代理:阈值触发警报,但没有将违规行为与累计产品降解相关联,也没有将降解信号转化为物流决策。我们通过一个质量感知决策智能(Quality-Aware Decision Intelligence, QADI)框架来填补这一空白,该框架结合了三种能力:一个结构化的质量状态表示 $S_q = [L, Q, U, R]$——剩余保质期、降解率、估计不确定性和操作风险,均可从框架方程中推导和计算;一个混合质量建模层,结合基于物理的微生物动力学与数据驱动的修正项;以及一个基于微软 Phi-4 的推理层,利用结构化领域知识库进行检索增强生成。我们在八个冷链场景中对五个基准进行了基准测试——阈值监控、仅物理模型、物理加噪声、基于优化的决策和基于规则的专家系统,以巴氏消毒牛奶作为主要案例,真实的保质期数据来自于独立于我们模型的已发布乳制品研究。比较使用了威尔科克森符号秩检验和霍尔姆校正。在牛奶和西兰花场景中,该框架的平均绝对保质期误差为7.2小时(相比之下,仅物理模型为30.9小时;$p<0.001$),腐败率为14.5%(相比之下,仅物理模型和基于规则的为16.6%;$p=0.08$),在99.5%的场景中实现了最佳决策。去除大语言模型(LLM)推理组件后,最优性下降至45.5%($p<0.001$)。专家评定的解释质量达到83%($ ext{kappa} = 0.71$)。消融实验表明,混合建模和LLM推理各自带来了显著的增益,而检索增强生成主要推动了解释质量。代码链接: https://bit.ly/4d6t44C。
cs.AI / 72 / 2608.15089

StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

StateM:通过增强执行系统在 Terminal-Bench 2.1 上实现 95.3% 的原始准确率,或 15 美元的前沿运行
Qin, Ziheng, Lu, Yaxin, Wang, Zhangyang Atlas, Wang, Kai
Abstract
Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track of mutable state, fail to reactivate lessons from earlier executions, skip known procedures, or stop prematurely. We bet on harness scaling to improve the execution system around an agent without changing its model weights. We introduce StateM, an agent-native runtime that organizes execution around durable states, phase-local context, checked transitions, recoverable runbooks, and versioned procedural practices that agents and users can inspect together. On Terminal-Bench 2.1, StateM raises GPT-5.5 xhigh to 92.1\%, versus 83.1\% reference and GPT-5.6 Sol Ultra at 91.9\%. The runbook transfers unchanged to GPT-5.6. With GPT-5.6 Sol xhigh, StateM reaches 95.3\% raw accuracy across 445 trials and succeeds on all 89 tasks at least once. The frozen profile raises GPT-5.6 Luna from 76.7 to 85.4\%, above the 84.9\% Sol xhigh reference. Using the same runtime, runbook structure, and golden rules, less than \$38 of adaptation raises DeepSeek-V4 Flash from 82.7 to 88.1\% under standard timeouts and to 89.1\% on an 88-task common core. Extending only the remaining latency-sensitive task matches the reported 88.8\% GPT-5.6 Sol max result. Final-score API usage is about \$15 versus \$574.68 for the GPT reference; total DeepSeek expenditure is \$52.22. On BusinessBench, family-specific runbooks built on development sets yield held-out gains of 0.55 macro and 1.34 micro points; two mechanism-matched families improve by 10.04 points. Concrete rules generalize when tasks share execution structure, while the control methodology applies broadly. StateM turns selected postmortem findings into persistent, executable preconditions and practices, making learned controls explicit and enforceable through stateful controls. Code at github.com/henryqin1997/statem.
Chinese Translation
长时间跨度的智能体即使在其基础模型能够解决组成步骤的情况下也可能失败。它们可能会失去对可变状态的跟踪,无法重新激活早期执行的经验,跳过已知程序或提前停止。我们寄希望于增强执行系统,以改善智能体周围的执行环境,而不改变其模型权重。我们介绍了 StateM,这是一种原生智能体运行时,围绕持久状态、阶段局部上下文、检查过渡、可恢复的运行手册和版本化的程序实践组织执行,智能体和用户可以共同检查这些内容。在 Terminal-Bench 2.1 上,StateM 将 GPT-5.5 xhigh 的准确率提升至 92.1%,相比之下,参考值为 83.1%,而 GPT-5.6 Sol Ultra 的准确率为 91.9%。运行手册在 GPT-5.6 上保持不变。使用 GPT-5.6 Sol xhigh,StateM 在 445 次试验中达到了 95.3% 的原始准确率,并在所有 89 项任务中至少成功一次。冻结的配置将 GPT-5.6 Luna 的准确率从 76.7% 提升至 85.4%,超过 84.9% 的 Sol xhigh 参考值。使用相同的运行时、运行手册结构和黄金规则,少于 38 美元的适应性将 DeepSeek-V4 Flash 的准确率从 82.7% 提升至 88.1%(在标准超时下)和 89.1%(在 88 项任务的公共核心上)。仅扩展剩余的延迟敏感任务即可匹配报告的 88.8% GPT-5.6 Sol 最大结果。最终得分 API 的使用成本约为 15 美元,而 GPT 参考值为 574.68 美元;DeepSeek 的总支出为 52.22 美元。在 BusinessBench 上,基于开发集构建的特定家庭运行手册产生了 0.55 的宏观增益和 1.34 的微观增益;两个机制匹配的家庭提高了 10.04 分。当任务共享执行结构时,具体规则能够推广,而控制方法则广泛适用。StateM 将选定的事后分析结果转化为持久的、可执行的前提条件和实践,使得学习到的控制变得明确且可通过状态控制强制执行。代码可在 github.com/henryqin1997/statem 获取。
cs.AI / 73 / 2608.15095

Validation-Frontier Representation Selection under Constrained Observation

受限观察下的验证前沿表示选择
Shu, Wesley
Abstract
AI systems deployed outside clean benchmark settings often rely on observations that are incomplete, unstable, costly, or degraded by monitoring failures. This paper studies representation selection under constrained observation: choosing a state representation when raw accuracy is not the only operational criterion. We propose a validation-frontier selector that combines balanced accuracy with penalties for feature cost, overfit gap, and validation-test instability. In a focused public-tabular benchmark using three scikit-learn datasets, five observation regimes, 45 matched task cells, 720 candidate actions, and 405 representation rows, the adaptive selector improves frontier score over full trace features by 0.025801 while reducing mean feature count by 22.733. Balanced-accuracy difference is small and not statistically significant. A broader offline stress test gives mixed results. The supported claim is therefore bounded: adaptive representation selection can improve a constrained-observation robustness-efficiency frontier in matched benchmark settings, but does not universally dominate trace baselines.
Chinese Translation
在清晰基准设置之外部署的人工智能系统通常依赖于不完整、不稳定、成本高或因监控失败而退化的观察。本文研究了在受限观察下的表示选择:在原始准确性不是唯一操作标准时选择状态表示。我们提出了一种验证前沿选择器,该选择器将平衡准确性与特征成本、过拟合差距和验证-测试不稳定性的惩罚相结合。在一个专注的公共表格基准中,使用了三个 scikit-learn 数据集、五种观察机制、45 个匹配任务单元、720 个候选动作和 405 行表示,适应性选择器在完整轨迹特征上的前沿得分提高了 0.025801,同时平均特征数量减少了 22.733。平衡准确性差异较小且在统计上不显著。更广泛的离线压力测试结果不一。因此,所支持的主张是有限的:适应性表示选择可以改善匹配基准设置下的受限观察鲁棒性-效率前沿,但并不普遍优于轨迹基线。
cs.AI / 74 / 2608.15101

Second-Order Policy Effects as State Transitions: A Source-Linked Benchmark for Policy Simulation

作为状态转移的二阶政策效应:政策模拟的源链接基准
Shu, Wesley
Abstract
Policy evaluation often estimates direct benefits and costs while treating the institutional environment as fixed. In practice, a policy changes the system it enters: actors adapt, enforcement capacity shifts, burdens move, and new equilibria form around capture, gaming, compliance theater, irreversibility, and repair costs. We formalize this as second-order policy-effect prediction and present a source-linked benchmark for policy simulation. The benchmark contains 96 named public-policy cases across eight domains and four balanced action classes: implement, modify, pilot, and block. Each case includes source locators and state variables for benefit, capture, gaming, burden shift, instability, uncertainty, irreversibility, distributional risk, and implementation capacity. The runner regenerates method outputs and aggregate results from the case table, and the simulator never reads the expert action target. We report a protocol-based transition-channel audit with recall, precision, F1-style efficiency, and selective top-channel stress diagnostics, so universal channel coverage is not mistaken for field validation. The side-effect simulator achieves mean policy-effect quality of 0.945, compared with 0.838 for the risk-register baseline and 0.879 for the causal-loop baseline. Its advantage is concentrated in side-effect recall and aggregate transition scoring; it does not dominate the best structured baselines on exact policy-action choice. The evidence remains benchmark-based, but supports a bounded claim: transition-state variables make policy simulators more sensitive to downstream institutional effects.
Chinese Translation
政策评估通常估计直接的收益和成本,同时将制度环境视为固定。在实践中,政策会改变其所进入的系统:参与者适应,执法能力发生变化,负担转移,新的均衡围绕捕获、游戏、合规表演、不可逆性和修复成本形成。我们将此形式化为二阶政策效应预测,并提出一个源链接的政策模拟基准。该基准包含来自八个领域的96个命名公共政策案例,以及四个平衡的行动类别:实施、修改、试点和阻止。每个案例包括源定位器和状态变量,涵盖收益、捕获、游戏、负担转移、不稳定性、不确定性、不可逆性、分配风险和实施能力。运行器从案例表中再生方法输出和聚合结果,模拟器从不读取专家行动目标。我们报告了一种基于协议的转移通道审计,包含召回率、精确度、F1样式效率和选择性顶级通道压力诊断,因此普遍通道覆盖不应被误认为是领域验证。副作用模拟器的平均政策效应质量为0.945,而风险登记基准为0.838,因果循环基准为0.879。其优势集中在副作用召回和聚合转移评分上;在精确政策行动选择上并未超越最佳结构基准。证据仍然基于基准,但支持一个有限的主张:转移状态变量使政策模拟器对下游制度效应更加敏感。
cs.AI / 75 / 2608.15109

Constraint-Aware Synthetic Tabular Data Generation via Inter-Column Constraint Discovery with LLM Agents

基于大语言模型代理的列间约束发现的约束感知合成表格数据生成
Zhao, Jianxing, Guan, Mao, Liu, Dongyu
Abstract
Generating structurally valid synthetic tabular data remains difficult: outputs with high statistical fidelity and downstream utility can still violate semantically meaningful domain constraints. We study the discovery and enforcement of three complementary inter-column constraint families---equations, linear inequalities, and logical dependencies. Our unified tool-grounded workflow represents all three as machine-executable hypotheses and applies a common interface for full-table validation, deterministic diagnosis, and counterexample-guided revision. A generator-agnostic postprocessor coordinates family-specific repairs on outputs from unchanged tabular generators. Across curated behavioral audits and end-to-end evaluations, the complete workflow improves held-out violation detection over one-shot direct prompting, while postprocessing yields zero measured violations for every retained, applicable constraint, improves downstream utility on most datasets, and largely preserves univariate marginals.
Chinese Translation
生成结构上有效的合成表格数据仍然面临挑战:尽管输出具有较高的统计保真度和下游实用性,但仍可能违反语义上有意义的领域约束。我们研究了三类互补的列间约束的发现与执行——方程、线性不等式和逻辑依赖。我们统一的工具驱动工作流程将这三类约束表示为机器可执行的假设,并应用通用接口进行全表验证、确定性诊断和反例引导修正。一个与生成器无关的后处理器协调对未更改的表格生成器输出进行特定于约束类别的修复。在经过精心策划的行为审计和端到端评估中,完整的工作流程在持出违规检测方面优于一次性直接提示,而后处理对每个保留的适用约束的测量违规率为零,且在大多数数据集上提高了下游实用性,同时在很大程度上保留了单变量边际分布。
cs.AI / 76 / 2608.15117

Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads

量化智能体的结构:代码合成智能工作负载中的VRAM稳定性与预测
Banerjee, Anubhab
Abstract
Analytical models of peak VRAM consumption for LLM inference decompose memory into weight-storage, KV-cache, and activation terms parameterized by step count, tool invocations, and context expansion. We evaluate this decomposition empirically within a strictly scoped measurement study: a LangGraph-based CUDA-kernel-synthesis agent (AgentK), a 4-bit quantization family (Q4 K M), a single NVIDIA H100 GPU, and four LLM backbones across 1,920 trajectories. Focusing on peak-memory forecasting behavior, we report two primary observations. First, closed-form analytical models achieve competitive accuracy when provided with two empirical constants: loaded-weight VRAM and a fixed activation-memory overhead. Supplied with live GPU readings and ground-truth trajectory parameters, the closed-form model matches or outperforms the best learned baseline on three of the four backbones (test MAPE 2.2-4.4% vs. 3.4-6.5%, p = 0.76). The exception is the smallest backbone (Phi-4-mini), where minimal VRAM variance (CV 0.3%) causes dynamic modeling to underperform simple regression. Second, compile success strictly bifurcates by backbone capacity (from 5.7% for Phi-4-mini to 62.0% for Qwen2.5-Coder-14B), demonstrating that functional code synthesis remains constrained by intrinsic LLM capabilities rather than available memory. Furthermore, because overall peak-memory variance is remarkably low across all backbones (CV 0.3-9.4%), learned prompt-feature regression offers statistically insignificant improvements over a constant-mean baseline. Consequently, we find no justification for deploying complex predictive VRAM models in highly quantized, weight-dominated regimes. We release the evaluated corpus and anonymized framework to support replication.
Chinese Translation
大型语言模型(LLM)推理的峰值VRAM消耗的分析模型将内存分解为权重存储、KV缓存和激活项,这些项由步骤计数、工具调用和上下文扩展参数化。我们在一个严格范围的测量研究中对这种分解进行了实证评估:一个基于LangGraph的CUDA内核合成智能体(AgentK)、一个4位量化系列(Q4 K M)、一台NVIDIA H100 GPU,以及四个LLM骨干网络,涵盖1,920条轨迹。我们专注于峰值内存预测行为,报告了两个主要观察结果。首先,当提供两个经验常数:加载权重的VRAM和固定的激活内存开销时,封闭形式的分析模型能够实现竞争性的准确性。在提供实时GPU读数和真实轨迹参数的情况下,封闭形式模型在四个骨干网络中的三个上与最佳学习基线相匹配或表现更好(测试MAPE 2.2-4.4%对比3.4-6.5%,p = 0.76)。例外的是最小的骨干网络(Phi-4-mini),在此情况下,最小的VRAM方差(CV 0.3%)导致动态建模的表现不及简单回归。其次,编译成功严格按照骨干网络的容量进行二分(从Phi-4-mini的5.7%到Qwen2.5-Coder-14B的62.0%),表明功能代码合成仍然受到内在LLM能力的限制,而非可用内存。此外,由于所有骨干网络的整体峰值内存方差相当低(CV 0.3-9.4%),学习的提示特征回归在统计上对常数均值基线的改进微不足道。因此,我们认为在高度量化、以权重为主导的环境中部署复杂的预测VRAM模型没有必要。我们发布了评估的语料库和匿名框架,以支持复制研究。
cs.AI / 77 / 2608.15131

Platform Adaptation Under Governance Interventions: Actor Best-Response Modeling and an External Public-Case Benchmark

治理干预下的平台适应:参与者最佳响应建模及外部公共案例基准
Shu, Wesley, Wei, Peng
Abstract
Digital platforms govern by changing rules: rankings, monetization thresholds, moderation standards, verification systems, disclosure requirements, appeal processes, and access policies. These interventions are rarely absorbed passively. Creators, sellers, advertisers, moderators, users, developers, and strategic operators adapt to the new reward surface. This paper develops a platform-adaptation model for evaluating governance interventions as transitions in adaptive multi-actor information systems. The model represents actor best response, strategic gaming opportunity, moderation burden, user-incentive movement, enforcement response, externality formation, and downstream platform stability. We evaluate the model on 72 external public platform-governance cases covering media monetization, ranking systems, verification, delivery platforms, marketplaces, app stores, community platforms, and creator ecosystems. Across 9 methods and 648 method-case evaluations, the full platform-adaptation simulator achieves mean adaptation quality of 0.836338, compared with 0.669731 for a risk-register baseline, 0.589457 for causal-loop analysis, 0.492750 for generic governance critique, 0.369492 for engagement-only optimization, and 0.331965 for baseline policy review. Paired comparisons show a win rate of 1.00 against all tested baselines and channel ablations. The contribution is an information-systems theory and measurement framework showing why platform governance evaluation fails when it treats policy rules as static controls rather than interventions into adaptive actor-response fields.
Chinese Translation
数字平台通过改变规则进行治理:排名、货币化门槛、审核标准、验证系统、披露要求、上诉流程和访问政策。这些干预措施很少被动接受。创作者、卖家、广告商、审核员、用户、开发者和战略运营者都会适应新的奖励结构。本文开发了一种平台适应模型,用于评估治理干预作为适应性多参与者信息系统中的转变。该模型表示参与者的最佳响应、战略游戏机会、审核负担、用户激励变化、执法响应、外部性形成和下游平台稳定性。我们在72个外部公共平台治理案例上评估该模型,涵盖媒体货币化、排名系统、验证、交付平台、市场、应用商店、社区平台和创作者生态系统。通过9种方法和648个方法案例评估,完整的平台适应模拟器实现了0.836338的平均适应质量,而风险登记基线为0.669731,因果循环分析为0.589457,通用治理批评为0.492750,仅参与优化为0.369492,基线政策审查为0.331965。配对比较显示,在所有测试的基线和渠道消融中,胜率为1.00。本文的贡献在于提供了一种信息系统理论和测量框架,说明当平台治理评估将政策规则视为静态控制而非适应性参与者响应领域的干预时,评估为何会失败。
cs.AI / 78 / 2608.15138

ReForge: Keeping ABR Algorithms Never Finished with Verified Large Language Model Edits

ReForge:利用经过验证的大型语言模型编辑保持自适应比特率(ABR)算法的持续更新
He, Zhiqiang, Liu, Zhi
Abstract
Designing an ABR algorithm for one network scenario takes an engineer months, and large language models now do this work in hours, matching or beating hand-built designs. But either way, the design fits only the world visible at its birth, and fails on the world that arrives after. We ask whether an ABR algorithm can keep pace with the world, redesigned in minutes as each scenario arrives, with every change proven harmless to every scenario already served. In this work, we propose ReForge, a continual heuristic learning framework that adapts to continuously changing scenarios. ReForge runs that routine with a large language model (LLM) in the loop. Each round the LLM reads where the current design falls short and proposes one small edit, and a replay over every network served so far decides. Specifically, what it edits is a single page of fuzzy rules that routes every decision to one of a frozen pool of pre-trained policies. The LLM writes the first page from measurements alone, then keeps improving it on its own. Each round it reads where the current rules fall short and proposes one small edit, and a replay over every network served so far decides whether the edit lands. We evaluate ReForge on nine real-world network families arriving one at a time as 3G, 4G, then 5G. A few edits per arrival lift mean QoE from 1.23 to 1.74, past the best single policy at 1.66 and to 94\% of an oracle, and even repair families the loop never saw, one rising from 0.30 to 0.80. All code, data, and experiment records will be open-sourced upon cleanup.
Chinese Translation
为一个网络场景设计一个自适应比特率(ABR)算法需要工程师数月的时间,而大型语言模型现在可以在数小时内完成这项工作,甚至可以与手工设计相媲美或超越。然而,无论如何,设计只适用于其诞生时可见的世界,无法应对随之而来的新变化。我们探讨是否可以让ABR算法与时俱进,在每个场景到来时在几分钟内重新设计,并证明每一次变更对已服务的每个场景都是无害的。在本研究中,我们提出了ReForge,一个持续的启发式学习框架,能够适应不断变化的场景。ReForge在循环中运行一个大型语言模型(LLM)。每一轮,LLM读取当前设计的不足之处,并提出一个小的编辑建议,然后对迄今为止服务的每个网络进行重放以决定是否采纳。具体而言,所编辑的内容是一页模糊规则,这些规则将每个决策路由到一个冻结的预训练策略池中的某个策略。LLM仅根据测量数据撰写第一页,然后自行不断改进。每一轮它读取当前规则的不足之处,并提出一个小的编辑建议,重放迄今为止服务的每个网络以决定该编辑是否有效。我们在九个真实世界的网络系列上评估ReForge,这些系列依次到来,分别为3G、4G和5G。每次到来的几个编辑将平均用户体验(QoE)从1.23提升至1.74,超越最佳单一策略1.66,达到94 ext{%}的预言者水平,甚至修复了循环从未见过的系列,其中一个从0.30上升到0.80。所有代码、数据和实验记录将在清理后开源。
cs.AI / 79 / 2608.15143

Translating finite-domain integer constraint models to CP/SMT/ILP/PB/SAT solvers with CPMpy

将有限域整数约束模型翻译为 CP/SMT/ILP/PB/SAT 求解器的 CPMpy
Guns, Tias, Bleukx, Ignace, Bierlee, Hendrik, Devriendt, Jo, Gamba, Emilio, Lomis, Orestis, Piessens, Wout, Sergeys, Thomas, Tsouros, Dimos, Vanroose, Wout, Verhaeghe, Hélène
Abstract
Constraint solving is a declarative approach for solving combinatorial satisfaction and optimization problems. The user specifies their problem through constraints and decision variables, and a generic solver is used to find a solution. Several constraint-solving technologies exist, and certain solvers perform well on certain problems. Therefore, it is useful to try different solvers given a particular application. However, each solving paradigm supports different types of constraints and decision variables. Our goal is to translate high-level constraint satisfaction and optimization problems into any lower-level formalism, including CP, SMT QF-LIA, ILP, PB and (Max)SAT. This allows for comparing different solving technologies for a particular problem, without requiring a user to manually remodel it for each solving paradigm. We define a high-level language of logical and arithmetic operations, and useful additional functions and constraints, which are known as global constraints in the CP community. We then present a modular framework for transforming our high-level modeling language to CP/SMT/ILP/PB and (Max)SAT solvers. While many transformations are partly described in the literature, we observe that they can be implemented through a modular waterfall of smaller components, where lower-level paradigms reuse the transformations of higher-level paradigms. Two recurring challenges are handling the negation of arbitrary subexpressions and avoiding the introduction of auxiliary variables. Additionally, we take special care linearizing non-linear operators for ILP, PB and SAT-solvers. The transformation waterfall is implemented and evaluated in the open-source CPMpy library. Our results show that constraint models significantly change throughout the transformations, and that optimizations to the linearization of constraints are essential for ILP and PB solvers.
Chinese Translation
约束求解是一种用于解决组合满足和优化问题的声明式方法。用户通过约束和决策变量来指定他们的问题,并使用通用求解器来寻找解决方案。存在多种约束求解技术,某些求解器在特定问题上表现良好。因此,在特定应用场景下尝试不同的求解器是有益的。然而,每种求解范式支持不同类型的约束和决策变量。我们的目标是将高层次的约束满足和优化问题翻译为任何低层次的形式,包括 CP、SMT QF-LIA、ILP、PB 和 (Max)SAT。这使得在不需要用户为每种求解范式手动重建问题的情况下,可以比较特定问题的不同求解技术。我们定义了一种逻辑和算术操作的高层语言,以及一些有用的附加函数和约束,这些在 CP 社区中被称为全局约束。然后,我们提出了一个模块化框架,用于将我们的高层建模语言转换为 CP/SMT/ILP/PB 和 (Max)SAT 求解器。尽管许多转换在文献中已有部分描述,但我们观察到它们可以通过模块化的瀑布式小组件实现,其中低层次的范式重用高层次范式的转换。两个反复出现的挑战是处理任意子表达式的否定以及避免引入辅助变量。此外,我们特别注意将非线性运算符线性化,以适应 ILP、PB 和 SAT 求解器。该转换瀑布已在开源的 CPMpy 库中实现和评估。我们的结果表明,约束模型在整个转换过程中发生了显著变化,并且对约束线性化的优化对于 ILP 和 PB 求解器至关重要。
cs.AI / 80 / 2608.15145

ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models

ACTS-SQL:基于代理和批评导向的树结构SQL正确性与大型语言模型
Huang, Xinmei, Song, Jie, Li, Peng, Jiang, Fuxin, Zhang, Jing, Zhang, Tieying, Chen, Jianjun, Liu, Chenming, Yang, Tao, Liu, Maoyin, Li, Wenda, Chen, Hong, Li, Cuiping
Abstract
Large Language Models (LLMs) have been increasingly adopted in Text-to-SQL systems, yet SQL errors remain a major obstacle in real-world Text-to-SQL inference pipelines. Existing SQL correction approaches either rely on large-scale, high-quality training data with substantial overhead, or adopt single-path agentic workflows that are brittle to early mistakes and prone to error propagation. To develop a practical SQL correctness system for industrial scenarios, we present a training-free framework that formulates SQL correction as a plan-guided, tree-structured debugging process. By maintaining multiple correction strategies and enabling backtracking, the framework mitigates error accumulation during iterative refinement. We further integrate execution-based verification and clause-level diagnostic tools to support strategy pruning and precise error localization. We evaluate the system on the BIRD-Critic benchmark and observe consistent accuracy gains over strong LLM backbones and representative agent-based baselines, achieving a 9.42% improvement over the previous state-of-the-art method. The framework is also deployed in the Torch Log Service (TLS) of Volcano Engine to support an online Text-to-TLS API. In production, it improves execution accuracy from 36.77% to 53.61% on real user queries with a representative strong LLM backbone (GPT-5). These results demonstrate the effectiveness and stability of our approach in real-world deployments.
Chinese Translation
大型语言模型(LLMs)在文本到SQL系统中得到了越来越广泛的应用,但SQL错误仍然是现实世界文本到SQL推理管道中的主要障碍。现有的SQL纠错方法要么依赖于大规模、高质量的训练数据,带来显著的开销,要么采用单路径的代理工作流,这些工作流对早期错误非常脆弱,容易导致错误传播。为了开发适用于工业场景的实用SQL正确性系统,我们提出了一种无训练的框架,将SQL纠错形式化为一个计划引导的树结构调试过程。通过维护多种纠错策略并支持回溯,该框架在迭代优化过程中减轻了错误累积。我们进一步集成了基于执行的验证和子句级诊断工具,以支持策略修剪和精确的错误定位。我们在BIRD-Critic基准上评估了该系统,观察到相较于强大的LLM主干和具有代表性的基于代理的基线,准确性持续提升,较之前的最先进方法提高了9.42%。该框架还被部署在火山引擎的Torch Log Service (TLS)中,以支持在线文本到TLS API。在生产环境中,它在具有代表性的强大LLM主干(GPT-5)上将真实用户查询的执行准确性从36.77%提高到53.61%。这些结果证明了我们的方法在现实世界部署中的有效性和稳定性。
cs.AI / 81 / 2608.15147

Constitutive Priors for Machine Intelligence: A Legitimacy Theory of the Artificial Physical World

机器智能的构成先验:人工物理世界的合法性理论
Jiang, Jiang, Sun, Yifu, Shen, Qi
Abstract
Machine intelligence has conquered the symbolic world but stalled at the physical one. The stall is structural: physical AI faces a cold-start deadlock -- no intelligence without data, no data without deployed intelligence. Our thesis: the deadlock is real but unevenly distributed, and the exception has a name: the artificial physical world. Buildings, industrial facilities, and infrastructure are intentionally constituted and documented: designed artifacts ship with readable archives that precede and constitute their instances; here, norms are promulgated before instances, not averaged from them. Four contributions. (i) From a four-world ontology we derive a legitimacy criterion for constitutive prior frameworks: prior extraction is legitimate if and only if the object domain is intentionally constituted and has left a readable archive; the criterion is testable through direction of fit -- deviation from a constitutive norm is a violation in the world, not a revision of the model. (ii) We establish a layering lower bound: any such framework has at least four layers -- syntax, concept, knowledge, instance -- because four construction goals pair into mutually incompatible carriers. (iii) We register deployment claims across five industrial domains and a 32-class failure-mode vocabulary. (iv) We stake the framework on five falsifiable predictions, the central one checkable on the public engineering record: if it fails, the framework fails. Semi-formal arguments back these claims (Appendix A): a Gold-type boundary on rule coverage in archiveless worlds, a decidability result for failure reduction over closed concept layers, and a boundary theorem for certificate-anchored calculi. Large language models find an honored place here -- as readers of the archive, not as the archive. First of three companion works; the companions take up the questions deliberately left open.
Chinese Translation
机器智能已经征服了符号世界,但在物理世界中停滞不前。这种停滞是结构性的:物理人工智能面临冷启动的死锁——没有数据就没有智能,没有部署的智能就没有数据。我们的论点是:这种死锁是真实存在的,但分布不均,例外的名称是:人工物理世界。建筑、工业设施和基础设施是有意构成和记录的:设计的工件附带可读的档案,这些档案在其实例之前存在并构成它们;在这里,规范是在实例之前被颁布的,而不是从实例中平均得出的。四个贡献。(i) 从四世界本体论中,我们推导出构成先验框架的合法性标准:如果且仅当对象领域是有意构成并留下可读档案时,先验提取才是合法的;该标准可以通过适配方向进行检验——偏离构成规范是在世界中的一种违反,而不是模型的修订。(ii) 我们建立了一个分层下限:任何这样的框架至少有四个层次——语法、概念、知识、实例——因为四个构造目标配对成相互不兼容的载体。(iii) 我们在五个工业领域和一个32类故障模式词汇中登记了部署声明。(iv) 我们在五个可证伪的预测上奠定了框架,中心预测可以在公共工程记录中检查:如果它失败,框架就失败。半正式的论证支持这些主张(附录A):在没有档案的世界中,规则覆盖的金边界,关于封闭概念层上故障减少的可判定性结果,以及证书锚定演算的边界定理。大型语言模型在这里找到了尊贵的位置——作为档案的读者,而不是档案本身。这是三部伴随作品中的第一部;伴随作品将探讨故意留出的未解问题。
cs.AI / 82 / 2608.15165

SkillCommit: Evolving Agent Skills through Behaviorally Validated Scope Expansion

SkillCommit:通过行为验证的范围扩展进化代理技能
He, Yu, Yang, Weikai
Abstract
Large language model (LLM) agents can continually improve without parameter updates by converting historical experience into reusable procedural knowledge. However, existing methods often consolidate experience based on semantic similarity or LLM judgments, which may merge superficially related but behaviorally incompatible strategies and thereby degrade performance. To address the issue, we propose SkillCommit, an online skill evolution framework that continuously transforms experience into a hierarchical library of reusable skills. Each new experience is initially preserved as an instance-specific patch, retaining the behavior validated in its local context. As related skills accumulate, SkillCommit abstracts those sharing a common behavioral mechanism into higher-level skills. Specifically, for each incoming skill, embedding-based retrieval first identifies candidate related skills. Cross-instance replay and an LLM-based mechanism check determine whether these skills transfer across cases and share a common underlying mechanism. Candidates that pass both checks are abstracted into a higher-level skill and committed only if it preserves the validated behavior of all constituent skills. Experiments on RuleArena, OpenExempt and KOR-Bench demonstrate that SkillCommit consistently improves agent performance across diverse domains. Moreover, the learned skills transfer across model scales and families, enabling cross-model experience transfer.
Chinese Translation
大型语言模型(LLM)代理可以通过将历史经验转化为可重用的程序知识而不断改进,而无需更新参数。然而,现有方法通常基于语义相似性或LLM判断来整合经验,这可能会合并表面上相关但行为上不兼容的策略,从而降低性能。为了解决这一问题,我们提出了SkillCommit,一个在线技能进化框架,持续将经验转化为可重用技能的层次库。每个新的经验最初被保留为特定实例的补丁,保留其在本地上下文中验证的行为。随着相关技能的积累,SkillCommit将那些共享共同行为机制的技能抽象为更高层次的技能。具体而言,对于每个新来的技能,基于嵌入的检索首先识别候选相关技能。跨实例重放和基于LLM的机制检查确定这些技能是否在不同案例之间转移并共享共同的基础机制。通过两项检查的候选技能被抽象为更高层次的技能,并仅在其保留所有组成技能的验证行为时才被承诺。对RuleArena、OpenExempt和KOR-Bench的实验表明,SkillCommit在不同领域中始终提高代理性能。此外,所学习的技能可以跨模型规模和家族转移,实现跨模型经验转移。
cs.AI / 83 / 2608.15242

LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

LongRCA Bench:诊断长时间跨度代理失败中的责任角色和根本原因
Zhang, Yunfei, Feng, Boyu, Pei, Changhua, Wang, Zexin, Peng, Zhihuang, Liu, Xinlong, Jiang, Hengyue, Ma, Difeng, Zhang, Jiayi, Yao, Yongzhou, Zhao, Yanan, Sun, Fei, Huo, Yintong, Liu, Zhaoyang, Li, Jingjing, Xie, Gaogang, Pei, Dan
Abstract
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then inspect the full execution to identify the responsible role and localize the earliest decisive root-cause step. Existing failure-attribution benchmarks largely focus on shorter traces, leaving diagnosis across hundreds of recorded steps underexplored. We introduce LongRCA Bench, comprising 1,140 failed trajectories across five domains without injected errors. It provides independently scored human labels for the responsible role and earliest decisive root-cause step. The median trajectory contains 145 steps, and the strongest baseline reaches only 13.2% exact root-step accuracy. We further present Root-Cause Trajectory Attribution (RCTA), a training-free method that retrieves candidate error steps from segment summaries and traces them to available earlier handoff instructions. Using the same backbone, benchmark instances, and scoring protocol, RCTA reaches 51.1% responsible-role accuracy and 24.1% exact root-step accuracy. These results highlight the need to evaluate responsible-role attribution and exact root-step localization as separate targets in long-trajectory failure diagnosis.
Chinese Translation
当长时间跨度的代理执行失败时,结果级评估揭示了不成功的结果,但并未指出决定性错误进入轨迹的位置。开发者必须检查完整的执行过程,以识别负责的角色并定位最早的决定性根本原因步骤。现有的失败归因基准主要集中于较短的轨迹,导致对数百个记录步骤的诊断研究不足。我们提出了LongRCA Bench,包含来自五个领域的1,140条失败轨迹,且未注入错误。它提供了独立评分的人类标签,用于识别负责角色和最早的决定性根本原因步骤。中位数轨迹包含145个步骤,而最强基线的准确根步骤率仅为13.2%。我们进一步提出了根本原因轨迹归因(Root-Cause Trajectory Attribution, RCTA),这是一种无训练的方法,可以从段摘要中检索候选错误步骤,并将其追溯到可用的早期交接指令。使用相同的基础架构、基准实例和评分协议,RCTA达到了51.1%的责任角色准确率和24.1%的准确根步骤率。这些结果突显了在长轨迹失败诊断中评估责任角色归因和准确根步骤定位作为独立目标的必要性。
cs.AI / 84 / 2608.15254

Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts

在多样性、公平性和包容性提示下的医疗语言模型中的人口注入
Mardian, Diego, Liu, Frank
Abstract
Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI). We measure a side effect that misrepresents patients: a one-sentence DEI prompt appended to a medical question leads models to add patient demographic attributes (race, socioeconomic status, sex) the question never stated, in effect rewriting who the patient is. We call this demographic injection. Across 47 models, four medical benchmarks, and 376,000 responses scored by a validated model-judge pipeline, a single DEI prompt raises the injection rate from 0.7% to 33.1% (47x) in all 47 of 47 models, attributable to the equity content rather than to added length (18x above a length-matched control; p=1.4x10^-14). Most added content is a general population statement that leaves the answer unchanged, but a smaller subset attaches an attribute to the specific patient or changes the selected option (0.25-2.4% of responses, 99.8% toward the incorrect option), where the invented demographic changes the answer the model recommends. Phrasing scales the effect from 14% to 56%. DEI prompts are just one example of a more general mechanism. Any instruction that nudges how a model reasons can make it add unrequested details, including details about the patient. Flagged outputs are treated as model errors under study, not clinical guidance.
Chinese Translation
临床人工智能指导越来越多地建议在提示语言模型时关注多样性、公平性和包容性(DEI)。我们测量了一种误导患者的副作用:在医疗问题后附加的一句DEI提示会导致模型添加问题中未提及的患者人口属性(种族、社会经济地位、性别),实际上重写了患者的身份。我们称之为人口注入。在47个模型、四个医疗基准和376,000个通过经过验证的模型评审管道评分的响应中,单个DEI提示使注入率从0.7%上升至33.1%(增加了47倍),这一现象归因于公平性内容而非附加长度(与长度匹配的对照组相比增加了18倍;p=1.4x10^-14)。大多数新增内容是一般人群声明,未改变答案,但较小的子集将属性附加到特定患者或更改所选选项(0.25-2.4%的响应,99.8%倾向于错误选项),其中虚构的人口属性改变了模型推荐的答案。措辞的变化使效果从14%扩大到56%。DEI提示只是一个更一般机制的例子。任何促使模型推理方式的指令都可能使其添加未请求的细节,包括有关患者的细节。被标记的输出被视为研究中的模型错误,而非临床指导。
cs.AI / 85 / 2608.15255

Towards Standardized Evaluation in Automated Domain Modeling: Introducing a Benchmark

朝着自动化领域建模的标准化评估:引入基准测试
Seibert, Vasiliy
Abstract
Domain modeling plays an essential role in domain-driven design, capturing essential entities and their relationships within a specific domain. Despite advancements in automated domain modeling, the absence of standardized benchmarks has hindered the comparative assessment of existing approaches. This paper introduces a benchmark designed to address this gap. The benchmark combines the 45-record Golden UML Modelset (Verbruggen et al., 2025) on Zenodo, as distributed by the Text2UML project of Calamo, Mecella, and Snoeck (Calamo et al., 2025), with the 8-record reference archive of Chen et al. (Chen et al., 2023a,b), enabling the evaluation of automated domain modeling approaches across different levels of complexity and scale. Given a natural language description, the task is to generate a corresponding domain model. For each description, a reference domain model is provided as ground truth. A metric is used to compare the generated domain model with the corresponding ground-truth model. To demonstrate the utility of the benchmark, we evaluate multiple automated domain modeling approaches, including heuristic rule-based methods and LLM-driven strategies. In accordance with the FAIR4RS recommendations (Chue Hong et al., 2022), the benchmark is provided as a research artifact to encourage reuse and support future research on automated domain modeling.
Chinese Translation
领域建模在领域驱动设计中扮演着至关重要的角色,捕捉特定领域内的基本实体及其关系。尽管自动化领域建模取得了进展,但缺乏标准化基准测试阻碍了对现有方法的比较评估。本文引入了一种旨在填补这一空白的基准测试。该基准测试结合了45条记录的Golden UML Modelset(Verbruggen等,2025)在Zenodo上的发布,由Calamo、Mecella和Snoeck的Text2UML项目分发(Calamo等,2025),以及Chen等(Chen等,2023a,b)的8条记录参考档案,使得能够在不同复杂性和规模水平上评估自动化领域建模方法。给定自然语言描述,任务是生成相应的领域模型。对于每个描述,提供一个参考领域模型作为真实值。使用一种度量标准将生成的领域模型与相应的真实值模型进行比较。为了展示基准测试的实用性,我们评估了多种自动化领域建模方法,包括启发式规则基础的方法和基于大型语言模型(LLM)的策略。根据FAIR4RS建议(Chue Hong等,2022),该基准测试作为研究成果提供,以鼓励重用并支持未来的自动化领域建模研究。
cs.AI / 86 / 2608.15256

Decentralized Federated Learning for Heterogeneous Multi-Task Semantic Communication

异构多任务语义通信的去中心化联邦学习
Yin, Lin, Lv, Tiejun, Li, Weicai, Yu, Xi, He, Xiaoyu
Abstract
Collaborative training in distributed semantic communication (DSC) networks typically relies on decentralized federated learning (DFL). However, pushing topology-agnostic aggregation into heterogeneous, multi-task environments creates a fundamental bottleneck: it drives negative transfer and overconsensus bias (OCB). This paper introduces a personalized DSC framework that cuts off this cross-task interference. At the node level, a policy-driven multi-path routing mechanism separates task-specific features from shared representations to preserve local fidelity. Across the network, we deploy a "communicationwhile- aggregation" protocol. It calibrates a column-stochastic consensus matrix using task affinities. This limits the system to absorbing complementary knowledge while actively blocking mismatched parameter updates. To bound the convergence, we derive a unified Lyapunov drift analysis. We reveal a strict Ushaped trade-off: deeper topological mixing reduces variance but amplifies structural OCB. Resolving this tension yields a closed-form expression for the optimal aggregation depth. We evaluate the proposed framework on NYU-v2, where the results reveal a clear trade-off between insufficient aggregation and excessive topological mixing. At the analytically derived optimal aggregation depth, our method achieves a 4.77% global relative improvement over the no-aggregation baseline and outperforms decentralized FedAvg, FedAMP, and heuristic max aggregation. We further evaluate the framework on Taskonomy and imperfect wireless links to examine the effects of network-size variation and wireless-link reliability.
Chinese Translation
分布式语义通信(DSC)网络中的协作训练通常依赖于去中心化联邦学习(DFL)。然而,将拓扑无关的聚合推向异构多任务环境会产生一个根本性瓶颈:它导致负迁移和过度共识偏差(OCB)。本文提出了一种个性化的DSC框架,以切断这种跨任务干扰。在节点层面,基于策略的多路径路由机制将任务特定特征与共享表示分离,以保持本地保真度。在网络层面,我们部署了一种“通信与聚合并行”的协议。该协议利用任务亲和性校准列随机共识矩阵。这限制了系统吸收互补知识的同时,积极阻止不匹配的参数更新。为了界定收敛性,我们推导了统一的Lyapunov漂移分析。我们揭示了一个严格的U形权衡:更深的拓扑混合减少了方差,但放大了结构性OCB。解决这一张力得出了最优聚合深度的封闭形式表达。我们在NYU-v2上评估了所提出的框架,结果显示出不足聚合与过度拓扑混合之间的明显权衡。在分析推导的最优聚合深度下,我们的方法在无聚合基线之上实现了4.77%的全球相对提升,并超越了去中心化的FedAvg、FedAMP和启发式最大聚合。我们进一步在Taskonomy和不完美无线链路上评估该框架,以检查网络规模变化和无线链路可靠性的影响。
cs.AI / 87 / 2608.15265

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

VibeWorlding:多模态智能体能否端到端构建3D开放世界?
Ning, Yansong, Ye, Jingwen, Wu, Zhongkai, Sun, Yang, Zhu, Yiqin, Li, Xingyi, Zhang, Weidong, Liu, Hao
Abstract
Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand user intent, use 3D tools, and reason over textual and visual 3D world information. To this end, we propose VibeWorlding, a unified framework for benchmarking and training vibe worlding agents: a multimodal agent that can autonomously infer user intent, plan scene layout, invoke 3D tools, and reflect on the multimodal feedback in a multi-turn agent-environment interaction process. To achieve this, we first build VWE-BENCH, a benchmark of 2,616 high-quality 3D assets, 323 human-annotated seed 3D worlds, and 6,828 reverse-synthesized multimodal user queries, split into verified queries with ground-truth and unverified queries with carefully designed rubrics. Moreover, we develop VibeWorlding-Gym, a joint multimodal RL post-training framework that integrates (1) a sandbox environment unifying asset retrieval, editing, and image rendering as MCP tools, and (2) a rubric-based verifier that combines physical feasibility and intent fulfillment verification, supporting both fair model evaluation and scalable multimodal RL reward service. Our experiments show that current frontier MLLMs are far from solving the vibe worlding agent task, with even GPT-5.5 and Qwen3.8-Max reaching below 60% success rate, and trace the bottleneck to precise 3D world editing. We further find that RL training can ease this weakness and enable open-source MLLMs to even surpass closed-source frontiers: our VibeWorlder-8B is comparable to frontier MLLMs, while our flagship VibeWorlder-30B-A3B attains the best overall Pass@1 among all evaluated models.
Chinese Translation
从用户查询构建一个互动的3D开放世界是非常重要的。然而,现有的方法主要在理想化的简单查询上进行评估,这使得系统性分析和比较多模态智能体如何理解用户意图、使用3D工具以及推理文本和视觉3D世界信息变得困难。为此,我们提出了VibeWorlding,这是一个用于基准测试和训练Vibe Worlding智能体的统一框架:一个能够自主推断用户意图、规划场景布局、调用3D工具并在多轮智能体与环境交互过程中反思多模态反馈的多模态智能体。为了实现这一目标,我们首先构建了VWE-BENCH,这是一个包含2616个高质量3D资产、323个人工标注的种子3D世界和6828个反向合成的多模态用户查询的基准,这些查询被分为具有真实标签的验证查询和经过精心设计的标准的未验证查询。此外,我们开发了VibeWorlding-Gym,这是一个联合多模态强化学习后训练框架,集成了(1)一个统一资产检索、编辑和图像渲染作为MCP工具的沙盒环境,以及(2)一个基于标准的验证器,结合了物理可行性和意图实现验证,支持公平的模型评估和可扩展的多模态强化学习奖励服务。我们的实验表明,目前的前沿多模态大语言模型(MLLMs)在解决Vibe Worlding智能体任务方面远未达到理想水平,即使是GPT-5.5和Qwen3.8-Max的成功率也低于60%,并将瓶颈追溯到精确的3D世界编辑。我们进一步发现,强化学习训练可以缓解这一弱点,使开源的多模态大语言模型甚至超过闭源的前沿模型:我们的VibeWorlder-8B与前沿的多模态大语言模型相当,而我们的旗舰产品VibeWorlder-30B-A3B在所有评估模型中获得了最佳的Pass@1成绩。
cs.AI / 88 / 2608.15288

$D^{2}R^{2}$: Discrete Diffusion with Regulation Reinforcement for Single-Cell Perturbation Prediction

$D^{2}R^{2}$:具有调控强化的离散扩散模型用于单细胞扰动预测
Fan, Ninghan, Liu, Qi, Zhu, Xunuo, Sun, Yukai, Chen, Luyuan, Zhou, Xuheng, Du, Yuetian, Kong, Ming, Zhu, Xiaojun, Liu, Jie, Zhou, Zhan, Zhu, Qiang
Abstract
Predicting single-cell transcriptomic responses to genetic perturbations is central to functional genomics and virtual-cell modeling. Existing approaches, however, typically predict an entire expression profile as a whole, leaving the order in which individual gene responses are generated unmodeled. To address this problem, we introduce \textbf{$D^{2}R^{2}$} (\textbf{D}iscrete \textbf{D}iffusion with \textbf{R}egulation \textbf{R}einforcement), which reformulates perturbation prediction as regulation-guided gene-wise progressive generation. A Masked Discrete Diffusion Model represents expression as ordinal tokens and reconstructs a fully masked profile step by step, allowing generated gene responses to condition those that remain masked. A Regulatory Policy Module initializes the generation policy from a gene regulatory network inferred from control cells and adapts it to the perturbation and current partially generated state. Then, group-relative policy optimization refines only the ordering policy using final perturbation-effect agreement as reward. Across Norman19 and VCC-H1, $D^{2}R^{2}$ achieves the best performance on all five metrics on Norman19 and remains competitive on H1. Controlled ablations holding the generator and generation budget fixed show that biological-prior ordering improves over random ordering and is more reliable than uncertainty-based heuristics, whereas reversing the biological-prior ordering degrades every metric. Biological analyses further show that the refined policy prioritizes regulatory genes early while promoting perturbation-specific transcription factors and responsive genes. These results establish gene generation order as an effective, controllable, and biologically interpretable dimension of single-cell perturbation prediction.
Chinese Translation
预测单细胞转录组对遗传扰动的响应是功能基因组学和虚拟细胞建模的核心。然而,现有方法通常将整个表达谱作为一个整体进行预测,未能建模单个基因响应生成的顺序。为了解决这个问题,我们引入了 extbf{$D^{2}R^{2}$}( extbf{D}iscrete extbf{D}iffusion with extbf{R}egulation extbf{R}einforcement),将扰动预测重新表述为调控引导的基因逐步生成。一个掩蔽的离散扩散模型将表达表示为序数标记,并逐步重建一个完全掩蔽的谱,从而允许已生成的基因响应影响那些仍然被掩蔽的响应。一个调控策略模块从控制细胞推断的基因调控网络初始化生成策略,并根据扰动和当前部分生成状态进行调整。然后,基于组相对策略优化仅使用最终扰动效应一致性作为奖励来优化排序策略。在Norman19和VCC-H1数据集中,$D^{2}R^{2}$在Norman19的所有五个指标上都取得了最佳性能,并在H1上保持竞争力。固定生成器和生成预算的控制消融实验表明,生物优先排序优于随机排序,并且比基于不确定性的启发式方法更可靠,而逆转生物优先排序则导致每个指标的下降。生物分析进一步表明,优化后的策略优先考虑调控基因,同时促进特定扰动的转录因子和响应基因。这些结果确立了基因生成顺序作为单细胞扰动预测中一个有效、可控且具有生物学解释维度的方式。
cs.AI / 89 / 2608.15291

ReasonCast: Agentic Demand Forecasting with Selective Semantic Reasoning

ReasonCast:基于选择性语义推理的代理需求预测
Yang, Ziyue, Xu, Chaolin, Wang, Yijing, Gu, Tiankai, Yang, Hui, Lin, Yanhong, Liu, Kaiyuan, Xiao, Fei
Abstract
Demand forecasting increasingly requires combining two complementary sources of information: historical sales reveal recurring numerical dynamics, while future promotions, holidays, price changes, and platform interventions provide forward-looking knowledge. Existing text-enhanced forecasting methods often encode such context into generic representations and fuse it uniformly with time-series features, without explicitly distinguishing which semantic effects are forecast-relevant or how they should modify future dynamics. We introduce ReasonCast, a structured semantic intervention framework that translates event knowledge into forecast-specific operations. An agent examines the event context, the no-text forecast, and its uncertainty to determine whether textual reasoning is needed. Rather than injecting free-form text, ReasonCast represents event knowledge through structured fields describing event relevance, demand direction, temporal shape, amplitude, and peak intensity. These fields interact selectively with temporal components of a time-series foundation model. An additive path corrects local trends and temporal shapes, while a multiplicative path captures event-driven level shifts. ReasonCast introduces a forecast-grounded post-training curriculum. Schema SFT establishes semantic fields; semantic-field RL calibrates direction, shape, amplitude, and peak judgments; and forecast-utility RL evaluates semantic interventions through a frozen forecaster, aligning reasoning outputs with marginal forecast improvement. ReasonCast lowers WMAPE by 3.29, 1.25, and 0.47 percentage points on holiday-sensitive categories, mega-sale-sensitive categories, and M5 event windows, respectively. On stable-sales periods, indiscriminate semantic intervention increases WMAPE by 1.68 percentage points, whereas suppressing unnecessary intervention preserves the numerical backbone.
Chinese Translation
需求预测越来越需要结合两种互补的信息来源:历史销售揭示了重复的数值动态,而未来的促销、假期、价格变化和平台干预提供了前瞻性的知识。现有的文本增强预测方法通常将此类上下文编码为通用表示,并与时间序列特征均匀融合,而没有明确区分哪些语义效应与预测相关或它们应如何修改未来的动态。我们提出了ReasonCast,一个结构化的语义干预框架,将事件知识转化为特定于预测的操作。一个代理检查事件上下文、无文本预测及其不确定性,以确定是否需要文本推理。ReasonCast通过结构化字段表示事件知识,这些字段描述事件相关性、需求方向、时间形状、幅度和峰值强度。这些字段与时间序列基础模型的时间组件进行选择性交互。加法路径修正局部趋势和时间形状,而乘法路径捕捉事件驱动的水平变化。ReasonCast引入了基于预测的后训练课程。Schema SFT建立语义字段;语义字段强化学习(semantic-field RL)校准方向、形状、幅度和峰值判断;而预测效用强化学习(forecast-utility RL)通过冻结的预测器评估语义干预,将推理输出与边际预测改进对齐。ReasonCast在假期敏感类别、超级促销敏感类别和M5事件窗口上分别降低了3.29、1.25和0.47个百分点的WMAPE。在稳定销售期间,任意的语义干预使WMAPE增加了1.68个百分点,而抑制不必要的干预则保持了数值基础。
cs.AI / 90 / 2608.15303

Divergent-Convergent Reasoning: Scaling Test-Time Compute through Structured Solution Synthesis

发散-收敛推理:通过结构化解决方案合成扩展测试时计算
Wen, Bo, Chen, Yuhao, Bilal, Erhan, Rios, Carla Agurto, Wang, Chen, Jiang, Junchen
Abstract
Test-time compute can substantially improve Large Language Model (LLM) reasoning performance, yet how and when additional compute helps remains poorly understood. We study Divergent-Convergent Reasoning (DCR), a simple two-phase primitive consisting of an exploration phase that generates multiple candidate solutions followed by a convergent reconciliation phase. We present three core results. First, we show that even a single reconciliation step can reliably amplify correct minority reports: across datasets, DCR often recovers the correct answer when correct exploration outputs are in the minority, a regime where majority voting fails. Second, we introduce recursive DCR, an autoregressive reconciliation system that iteratively analyzes disagreements and allocates additional test-time compute. Recursive DCR achieves higher accuracy than fixed-compute baselines-reaching 93.3% on AIME 2024 and 92.0% on AIME 2025-while using roughly 27% less compute on average, demonstrating that attentive resource allocation is superior to uniform scaling. Third, we analyze disagreement among exploration outputs via a simple, training-free dispersion metric. Dispersion reveals a structured relationship between disagreement and test-time gains: in regimes where DCR is effective, higher disagreement among exploration outputs is associated with larger accuracy improvements from reconciliation. Together, these results show that disagreement, often viewed as noise, can be systematically exploited to improve test-time reasoning and reveal emerging scaling laws for agentic LLM systems.
Chinese Translation
测试时计算可以显著提高大型语言模型(LLM)的推理性能,但额外计算如何以及何时提供帮助仍然不甚明了。我们研究了发散-收敛推理(Divergent-Convergent Reasoning, DCR),这是一种简单的两阶段原语,包含一个生成多个候选解决方案的探索阶段,随后是一个收敛和解阶段。我们提出了三个核心结果。首先,我们展示了即使是单个收敛步骤也能可靠地放大正确的少数报告:在不同数据集上,当正确的探索输出处于少数时,DCR通常能够恢复正确答案,而在这种情况下,简单的多数投票方法会失效。其次,我们引入了递归DCR(recursive DCR),这是一种自回归的收敛系统,能够迭代分析分歧并分配额外的测试时计算。递归DCR的准确率高于固定计算基线——在AIME 2024上达到93.3%,在AIME 2025上达到92.0%——同时平均使用的计算量减少了约27%,证明了细致的资源分配优于均匀扩展。第三,我们通过一种简单的无训练分散度量分析探索输出之间的分歧。分散度揭示了分歧与测试时收益之间的结构性关系:在DCR有效的情况下,探索输出之间的更高分歧与收敛带来的更大准确性提升相关联。综合来看,这些结果表明,通常被视为噪声的分歧可以被系统性地利用,以改善测试时推理,并揭示代理型LLM系统的新兴扩展规律。
cs.AI / 91 / 2608.15304

Understanding Cognition-Induced Risks in Agentic AI Systems

理解自主人工智能系统中的认知引发风险
Wang, Guanchu, Li, Qinuo, Du, Mengnan, Hu, Xia, Zhou, Bowen
Abstract
Frontier agentic systems powered by large language models (LLMs) exhibit human-like patterns of cognition. As these systems become deeply integrated across different domains, their cognitive engagement raises critical concerns for human society that remain insufficiently studied. To address this gap, we systematically analyze risks induced by expanding cognitive capabilities, following a three-level framework defined by their cognitive scope, from physical cognition to social cognition, and finally to self-referential cognition. We study their potential risks to human agency, autonomy, and control capability, corresponding to each cognitive level. We finally propose strategies to mitigate these risks and enhance the controllability of agentic AI systems, ensuring their long-term safe development.
Chinese Translation
由大型语言模型(LLMs)驱动的前沿自主系统展现出类人认知模式。随着这些系统在不同领域的深度整合,它们的认知参与引发了对人类社会的重大关注,这些关注尚未得到充分研究。为了解决这一空白,我们系统地分析了由扩展认知能力引发的风险,采用了一个基于认知范围的三层框架,从物理认知到社会认知,最后到自我指涉认知。我们研究了这些风险对人类代理性、自主性和控制能力的潜在影响,分别对应于每个认知层级。最后,我们提出了减轻这些风险和增强自主人工智能系统可控性的策略,以确保其长期安全发展。
cs.AI / 92 / 2608.15309

Physiological World Models for Human State Transitions

人类状态转变的生理世界模型
Zhang, Chongyang, Wang, Rendong, Zheng, Hao, Zhang, Hanwen, Liu, Yang, Wei, Xiaolong, Chong, Bin
Abstract
Continuous multimodal sensing now allows human physiology to be observed throughout daily life rather than only during occasional clinical visits. However, most health artificial intelligence systems are designed to recognize current states, estimate risks or analyse individual biomarkers. They do not directly model how physiological states change in response to real-world events, behaviours, contexts and interventions. Here we propose the Physiological World Model (PWM), an event-conditioned framework for learning these changes at the level of the whole person. We introduce the HumanState Transition Token, a structured, quality-scored unit that connects the physiological state before an event with the event or action, relevant context and intervention information, the physiological trajectory after the event, observed outcomes and data quality. We describe four capability levels, from state representation to bounded intervention planning, together with four data acquisition and validation protocols. We also propose six benchmark tasks covering HumanState representation, forecasting across multiple timescales, individualized response prediction, simulation of alternative interventions, bounded planning and reliability under distribution shift. Together, this framework provides a practical path towards personalized health management, behavioural intervention design and clinician-supervised decision support, while clearly separating prediction from causal inference and making uncertainty, safety, governance and limits of use explicit.
Chinese Translation
连续的多模态感知现在使得人类生理在日常生活中得以观察,而不仅仅是在偶尔的临床访问中。然而,大多数健康人工智能系统旨在识别当前状态、评估风险或分析个体生物标志物。它们并未直接建模生理状态如何响应现实世界事件、行为、情境和干预而变化。在此,我们提出了生理世界模型(Physiological World Model, PWM),这是一个基于事件的框架,用于学习这些变化,关注整个个体的层面。我们引入了人类状态转变标记(HumanState Transition Token),这是一个结构化的、质量评分的单元,连接事件前的生理状态与事件或行动、相关上下文和干预信息、事件后的生理轨迹、观察到的结果以及数据质量。我们描述了四个能力层级,从状态表示到有限干预规划,以及四个数据采集和验证协议。我们还提出了六个基准任务,涵盖人类状态表示、跨多个时间尺度的预测、个性化响应预测、替代干预的模拟、有限规划和在分布转变下的可靠性。总体而言,该框架为个性化健康管理、行为干预设计和临床监督决策支持提供了切实可行的路径,同时清晰地区分了预测与因果推断,并明确了不确定性、安全性、治理和使用限制。
cs.AI / 93 / 2608.15311

MoE Router-Guided Clustering for Heterogeneous Federated Instruction Tuning

基于MoE路由引导的异构联邦指令微调聚类
Sharma, Ankita, Farahani, Bahar, Moosavi, Sanaz Rahimi, Rrahmani, Amir, Firouzi, Farshad, Chakrabarty, Krishnendu
Abstract
Federated instruction fine-tuning enables Large Language Models (LLMs) to adapt to decentralized, privacy-sensitive data without requiring data sharing. Recent Mixture-of-Experts (MoE) LLMs are particularly attractive for federated learning because their sparse activation reduces computation and communication while scaling model capacity. However, existing federated MoE methods primarily focus on parameter aggregation and personalization, overlooking the routing behavior of MoE models as a source of information for client collaboration. Under heterogeneous instruction distributions, indiscriminate aggregation can lead to negative transfer, highlighting the need to identify which clients should collaborate during federated optimization. We propose ClientMorpher, a routing-aware, personalized federated instruction fine-tuning framework that leverages routing signatures from pretrained MoE models to organize client collaboration prior to aggregation. We investigate two complementary clustering strategies: ClientMorpher-C, which directly clusters clients using expert activation profiles, and ClientMorpher-E, which first clusters experts based on their cross-client usage signatures and then derives client collaboration groups. We evaluate ClientMorpher for federated instruction fine-tuning on the Databricks Dolly-15K dataset, using pathological and Dirichlet-based heterogeneous client distributions across multiple instruction-following tasks. Experimental results show that routing-aware collaboration consistently improves personalized performance compared to conventional federated averaging and local training, while maintaining the same communication cost. Furthermore, our study shows that client-centric and expert-centric clustering provides an effective and scalable approach for personalized federated instruction fine-tuning of sparse MoE LLMs.
Chinese Translation
联邦指令微调使大型语言模型(LLMs)能够适应去中心化的、隐私敏感的数据,而无需数据共享。最近的混合专家(MoE)LLMs因其稀疏激活在计算和通信上减少了开销,同时扩展了模型容量,因此在联邦学习中尤为吸引人。然而,现有的联邦MoE方法主要集中于参数聚合和个性化,忽视了MoE模型的路由行为作为客户协作的信息来源。在异构指令分布下,盲目的聚合可能导致负迁移,这突显了在联邦优化过程中识别哪些客户应进行协作的必要性。我们提出了ClientMorpher,一个具有路由感知的个性化联邦指令微调框架,利用预训练MoE模型的路由特征在聚合之前组织客户协作。我们研究了两种互补的聚类策略:ClientMorpher-C,直接使用专家激活特征对客户进行聚类,以及ClientMorpher-E,首先根据专家的跨客户使用特征对专家进行聚类,然后推导出客户协作组。我们在Databricks Dolly-15K数据集上评估ClientMorpher的联邦指令微调,使用病态和基于Dirichlet的异构客户分布,涵盖多个指令跟随任务。实验结果表明,与传统的联邦平均和本地训练相比,路由感知的协作在保持相同通信成本的同时,始终改善了个性化性能。此外,我们的研究表明,以客户为中心和以专家为中心的聚类为稀疏MoE LLMs的个性化联邦指令微调提供了一种有效且可扩展的方法。
cs.AI / 94 / 2608.15314

Physics-informed VAE-EVT for Tail Aware Radio Map Prediction

基于物理信息的变分自编码器-极值理论(VAE-EVT)用于尾部感知的无线电地图预测
Gamage, Amanda Sheron, Mehrnia, Niloofar, Gross, James
Abstract
Ultra-reliable low-latency communication (URLLC) requires precise identification of spatial regions where the signal-to-noise ratio (SNR) falls below an outage threshold. In this context, an outage refers to instances in which SNR falls below a specified threshold, which, for URLLC, can be as stringent as the 0.1% quantile of the SNR distribution. Traditional generative radio map models tend to focus on reconstructing average signal levels, often overlooking the low SNR that is crucial for accurate outage prediction. To address this limitation, we introduce a physics- and tail-informed VAE-EVT (variational autoencoder-extreme value theory) framework that distinctly models both the bulk and tail distribution of SNR. Our approach begins with a physics-informed preprocessing stage that extracts deterministic features, including line-of-sight, shadowing, and distance, from the scene geometry. A dual-latent encoder then captures the bulk SNR using a Gaussian mixture and the tail using a generalized Pareto distribution (GPD). By employing a modified variational objective, the model is trained to jointly supervise both regimes, ensuring focused attention on extreme fading events. Evaluated on the RadioMapSeer dataset, our method achieves an SNR RMSE of 4.83 dB in the outage region defined by the low threshold of 0.1% SNR quantile. This significantly outperforms the state-of-the-art GAN-based model, which records an SNR RMSE of 21.90 dB, with the performance gap widening as the outage threshold becomes more stringent.
Chinese Translation
超可靠低延迟通信(URLLC)要求精确识别信噪比(SNR)低于中断阈值的空间区域。在此背景下,中断指的是SNR低于指定阈值的情况,对于URLLC而言,这一阈值可以严格到SNR分布的0.1%分位数。传统的生成无线电地图模型往往侧重于重建平均信号水平,常常忽视对准确中断预测至关重要的低SNR。为了解决这一局限性,我们提出了一种基于物理信息和尾部信息的VAE-EVT(变分自编码器-极值理论)框架,明确建模SNR的主体和尾部分布。我们的方法首先通过物理信息预处理阶段,从场景几何中提取确定性特征,包括视距、阴影和距离。然后,双潜变量编码器使用高斯混合模型捕获主体SNR,并使用广义帕累托分布(GPD)捕获尾部。通过采用修改的变分目标,模型被训练以共同监督这两个区域,确保对极端衰落事件的集中关注。在RadioMapSeer数据集上的评估显示,我们的方法在由0.1% SNR分位数定义的中断区域中达到了4.83 dB的SNR均方根误差(RMSE)。这一结果显著优于最先进的基于生成对抗网络(GAN)模型,其SNR RMSE为21.90 dB,且随着中断阈值的严格性增加,性能差距进一步扩大。
cs.AI / 95 / 2608.15326

The Benchmark Trap: Structures of Power and Injustice in AI Evaluations

基准陷阱:人工智能评估中的权力与不公结构
Branford, Jason, Kraft, Angelie
Abstract
Artificial intelligence (AI) benchmarks are not neutral tools of evaluation but socio-technical artefacts that shape competition, power, and research priorities within AI. Benchmarks standardise the assessment of systems and facilitate the creation of leaderboards that reward state-of-the-art performance with prestige, citations, trust, and institutional influence. As the costs of developing competitive AI systems rise, these rewards increasingly concentrate among powerful, industry-funded labs. This paper situates these concerns within Iris Marion Young's theories of oppression and structural injustice. It argues that current benchmarking practices may perpetuate systematic harms affecting various actors in AI research, aligning with four of Young's "faces of oppression". Benchmarking culture is further framed as a source of structural injustice, as these harms emerge from normalised, individually defensible practices and network effects, even without explicit wrongdoing. By reinforcing existing power structures and narrowing possible research trajectories, benchmarking may in fact prevent the field from advancing in epistemically robust and socially beneficial ways.
Chinese Translation
人工智能(AI)基准并非中立的评估工具,而是塑造竞争、权力和研究优先级的社会技术产物。基准标准化了系统的评估,并促进了创建奖励最先进性能的排行榜,这些排行榜带来了声望、引用、信任和机构影响力。随着开发竞争性人工智能系统的成本上升,这些奖励越来越集中于强大的、由行业资助的实验室。本文将这些问题置于艾瑞斯·玛丽昂·杨(Iris Marion Young)的压迫与结构性不公理论之中。它论证了当前的基准实践可能会延续影响人工智能研究中各个参与者的系统性伤害,符合杨所提出的四种“压迫面貌”。基准文化进一步被框定为结构性不公的来源,因为这些伤害源于规范化的、个体可辩护的实践和网络效应,即使没有明确的不当行为。通过强化现有的权力结构并缩小可能的研究轨迹,基准实际上可能阻碍该领域在认识论上稳健且社会上有益的方式上取得进展。
cs.AI / 96 / 2608.15335

A concentration result for multilayer feedforward neural networks

多层前馈神经网络的集中性结果
Koponen, Vera
Abstract
We consider for an arbitrary fixed $\rho$ and for each positive integer $n$ a multilayer feedforward artificial neural network with $\rho$ layers, $n$ neurons in the first layer (the input layer) and only one neuron, the output neuron, in the last layer. Very roughly formulated, the main result is that if the distribution of weights of connections from a layer to the next are, for all large $n$, approximated well by a fixed continuous (but otherwise arbitrary) curve which does not depend on $n$, and if the values of the $n$ input neurons are independently and identically distributed with a continuous probability density function, then there is a number $\psi$ such that for all $\varepsilon > 0$ the probability that the value of the output neuron is in $[\psi - \varepsilon, \psi + \varepsilon]$ tends to 1 as $n$ tends to infinity.
Chinese Translation
我们考虑对于任意固定的 $ ho$ 和每个正整数 $n$,一个具有 $ ho$ 层、第一层(输入层)中有 $n$ 个神经元以及最后一层仅有一个神经元(输出神经元)的多层前馈人工神经网络。粗略地说,主要结果是:如果从一层到下一层的连接权重分布,对于所有大的 $n$,可以很好地用一个固定的连续(但其他方面任意的)曲线来近似,并且如果 $n$ 个输入神经元的值是独立同分布且具有连续概率密度函数,那么存在一个数 $ heta$,使得对于所有 $ heta > 0$,输出神经元的值落在 $[ heta - heta, heta + heta]$ 中的概率随着 $n$ 趋向于无穷大而趋近于 1。
cs.AI / 97 / 2608.15354

Incoherent by Design? On the Moral Self-Consistency of LLMs

设计上的不一致?关于大型语言模型的道德自我一致性
Nokhiz, Pegah, Ruwanpathirana, Aravinda Kanchana, Nissenbaum, Helen
Abstract
LLMs are increasingly used in morally sensitive contexts, yet it is unclear whether they apply ethical principles consistently across situations. A model that can state a moral principle may still violate it when the same scenario is rephrased or reframed. This inconsistency is a problem for any system whose outputs are used to inform moral decisions. If generative systems exhibit internal inconsistency, then the epistemic integrity of AI-mediated systems becomes uncertain. To study this concern, we investigate the stability of moral reasoning in LLMs within a controlled prompting framework across three major philosophical schools of thought: deontology, utilitarianism, and virtue ethics. We construct sets of morally equivalent scenarios in which the underlying situation is held constant while the framing varies to reflect different ethical stances and stylistic perturbations. We then evaluate responses from multiple models, including GPT, Mistral, and Llama. To assess consistency, we convert model outputs into structured logical statements and identify contradictions across responses generated within the same school of thought. Our results reveal substantial inconsistency with contradiction rates reaching up to 78% across scenarios. These findings point to a broader phenomenon of epistemic instability in generative AI wherein models fail to reliably maintain coherence with respect to their own prior outputs. This kind of instability carries real consequences. As generative systems influence how people form beliefs, judge actions, and absorb values, their inconsistencies can shape human reasoning and decision-making as well. Moreover, if a system cannot consistently represent its own normative commitments, then value alignment becomes a moving target rather than a well-defined objective. Thus, we argue that demonstrating internal incoherence is a necessary precursor to AI alignment.
Chinese Translation
大型语言模型(LLMs)在道德敏感的环境中被越来越多地使用,但尚不清楚它们是否在不同情境中一致地应用伦理原则。一个能够陈述道德原则的模型,在同一场景被重新表述或重新构建时,仍可能违反该原则。这种不一致性对于任何其输出被用于指导道德决策的系统来说都是一个问题。如果生成系统表现出内部不一致性,那么由人工智能介导的系统的认知完整性便变得不确定。为了研究这一问题,我们在一个受控的提示框架内,探讨了大型语言模型的道德推理的稳定性,涵盖了三大主要哲学流派:义务论(deontology)、功利主义(utilitarianism)和美德伦理学(virtue ethics)。我们构建了一组道德上等价的场景,其中基础情境保持不变,而框架则变化,以反映不同的伦理立场和风格扰动。随后,我们评估了多个模型的响应,包括GPT、Mistral和Llama。为了评估一致性,我们将模型输出转换为结构化逻辑陈述,并识别在同一哲学流派内生成的响应之间的矛盾。我们的结果显示出显著的不一致性,矛盾率在不同场景中高达78%。这些发现指向生成性人工智能中的一种更广泛的认知不稳定现象,其中模型未能可靠地与其自身的先前输出保持一致。这种不稳定性带来了实际后果。随着生成系统影响人们形成信念、判断行为和吸收价值观的方式,它们的不一致性也可能影响人类的推理和决策。此外,如果一个系统无法一致地表达其自身的规范承诺,那么价值对齐就变成了一个不断变化的目标,而不是一个明确的目标。因此,我们认为,展示内部不一致性是人工智能对齐的必要前提。
cs.AI / 98 / 2608.15372

UC-PSRO: Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum for Game-Theoretic Course-of-Action Generation in Adversarial Swarms

UC-PSRO:具有通信中断课程的效用条件政策空间响应oracle,用于对抗群体中的博弈论行动方案生成
Jiang, Phillip
Abstract
We study generating game-theoretically optimized Courses of Action (COAs) for a Blue UAS swarm against an adaptive Red adversary in a communication-degraded environment, motivated by (but not derived from) a public U.S. Air Force SBIR solicitation. We propose UC-PSRO (Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum), combining three mechanisms: (i) PSRO self-play, so Blue and Red policies train as approximate best responses to each other rather than one side against a fixed scripted opponent; (ii) FiLM conditioning of the Blue policy on a Commander's-Intent weight vector, sampled from a Dirichlet distribution during training, so one trained policy is re-steerable at execution time without retraining; and (iii) a curriculum annealing communication-graph edge dropout during training, so the swarm learns decentralized, peer-to-peer fallback instead of depending on full connectivity. We evaluate on a synthetic, unclassified stand-in for the solicitation's maritime scenario, with 5 seeds at N=25 Blue agents and a scalability sweep to N=200. We find a genuine trade-off, not a uniform win: the communication-dropout curriculum alone gives the strongest, most robust mission-completion rates of any learned method, improving counter-intuitively as denial increases (35% to 62% success as dropout rises from 0 to 0.75); adding utility-conditioning and PSRO self-play substantially slows convergence within a fixed budget, and we find no reliable exploitability advantage for self-play over a fixed-opponent policy, both statistically indistinguishable from a small, near-zero gap. We report this honestly as a convergence cost not yet offset by a demonstrated robustness benefit, rather than overstating one method as dominant, and provide a fully vectorized, open environment training at N=200 agents in single-digit milliseconds per step on a single consumer GPU.
Chinese Translation
我们研究在通信退化环境中,为蓝方无人机群体生成博弈论优化的行动方案(COAs),以应对适应性红方对手,灵感来源于(但并非源于)美国空军SBIR公开征集。我们提出UC-PSRO(具有通信中断课程的效用条件政策空间响应oracle),结合了三种机制:(i)PSRO自我对弈,使蓝方和红方政策作为彼此的近似最佳响应进行训练,而不是一方对抗固定的脚本对手;(ii)基于指挥官意图权重向量的FiLM条件化蓝方政策,该权重向量在训练期间从Dirichlet分布中采样,因此一个训练好的政策在执行时可以重新引导而无需重新训练;(iii)在训练期间采用课程退火通信图边缘中断,使得无人机群体学习去中心化的点对点后备,而不依赖于完全连接。我们在一个合成的、未分类的替代场景中进行评估,模拟征集的海洋场景,使用5个种子,N=25的蓝方代理,并进行N=200的可扩展性测试。我们发现存在真正的权衡,而非统一的胜利:仅通信中断课程就提供了任何学习方法中最强、最稳健的任务完成率,随着中断增加(从0到0.75),成功率从35%提高到62%;添加效用条件化和PSRO自我对弈在固定预算内显著减缓了收敛速度,我们发现自我对弈相较于固定对手政策并没有可靠的可利用性优势,两者在统计上不可区分,差距接近于零。我们诚实地报告这是一个尚未被证明的稳健性收益所抵消的收敛成本,而不是夸大某种方法的主导地位,并提供一个完全向量化的开放环境训练,在单个消费级GPU上以每步单数毫秒的速度处理N=200代理。
cs.AI / 99 / 2608.15381

FedPA-LoRA: Product-Aligned Framework for Mitigating Aggregation and Initialization Errors in Heterogeneous Federated LoRA

FedPA-LoRA:一种针对异构联邦LoRA中聚合和初始化误差的产品对齐框架
Jeon, Juseok, Ali, Ramy E., Kwon, Doyun, Her, Myungbeom, Kim, Jinhwi, So, Jinhyun
Abstract
Low-Rank Adaptation (LoRA) enables efficient federated fine-tuning of large language models, but its factorized parameterization creates a tension between accurate aggregation of local updates and continuity of locally optimized factors. Factor-wise aggregation incurs aggregation mismatch but better preserves factor continuity, whereas product-space reconstruction reduces this mismatch at the cost of greater factor-level initialization mismatch from newly reconstructed factors. We propose FedPA-LoRA, a product-aligned federated LoRA framework that jointly addresses these limitations and provably converges under both homogeneous and heterogeneous client ranks. Each client preserves its local factors across communication rounds and aligns its product toward a rank-specific global reference, maintaining local optimization continuity while promoting global consistency under data heterogeneity. The server aggregates heterogeneous-rank updates in the common product space and efficiently reconstructs a rank-constrained global adapter without forming the dense aggregate. This design supports client-specific computation and communication budgets. Experiments on natural language understanding and generation tasks show that FedPA-LoRA consistently outperforms representative baselines across varying levels of data heterogeneity and homogeneous- and heterogeneous-rank settings, with up to a $6.82$ percentage-point improvement in average GLUE accuracy under heterogeneous client ranks.
Chinese Translation
低秩适应(LoRA)使大语言模型的高效联邦微调成为可能,但其分解参数化在准确聚合本地更新与保持本地优化因子的连续性之间产生了矛盾。因子级聚合会导致聚合不匹配,但更好地保持因子的连续性,而产品空间重构则在降低这种不匹配的同时,增加了新重构因子带来的因子级初始化不匹配。我们提出了FedPA-LoRA,这是一种产品对齐的联邦LoRA框架,旨在共同解决这些局限性,并在同质和异质客户端秩下证明收敛。每个客户端在通信轮次中保持其本地因子,并将其产品对齐到特定秩的全局参考,保持本地优化的连续性,同时在数据异质性下促进全局一致性。服务器在公共产品空间中聚合异质秩更新,并高效重构一个秩受限的全局适配器,而无需形成密集聚合。该设计支持客户端特定的计算和通信预算。在自然语言理解和生成任务上的实验表明,FedPA-LoRA在不同数据异质性水平以及同质和异质秩设置下,始终优于代表性基线,在异质客户端秩下的平均GLUE准确率提高了高达6.82个百分点。
cs.AI / 100 / 2608.15382

Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot

将医疗保健大型语言模型与因果知识图谱相结合:框架、指标及心血管试点
Mumtaz, Ummara, Noor, Aimen, Ahmed, Awais
Abstract
Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-centered evaluation framework for intervention-oriented LLM behavior in healthcare and stress-test it in a cardiovascular pilot. The framework has four components: (i) a domain causal knowledge graph in which assertions are first-class, provenance-preserving nodes with stable identifiers; (ii) a scenario-conditioned subgraph extraction step that, given any clinical scenario, retrieves the relevant reified-assertion subgraph; (iii) four controlled grounding conditions that vary how the retrieved subgraph is composed into the model's context (ungrounded C1, knowledge-graph C2, causal-graph C3, integrated C4); and (iv) an automated scoring pipeline, anchored on assertion identifiers, that computes intervention accuracy, and other evaluation measures on a single pass. To test the framework, we built a category-balanced scenario generator across eight reasoning failure modes and instantiated it on a cardiovascular graph. The metric panel discriminates conditions along interpretable, non-redundant axes: C4 obtains the strongest causal edge F1 (0.838), adverse-effect F1 (0.833), evidence accuracy (0.738), and unsupported claim rate (0.114), while C1 obtains the highest raw intervention accuracy (0.948) with no measurable causal or evidential grounding.
Chinese Translation
大型语言模型(LLMs)越来越多地被提议用于医疗决策支持,但其评估仍然侧重于单一答案的准确性,而非对干预、机制、危害、证据和不确定性的推理。我们提出了一种可重复的、以图为中心的评估框架,旨在评估医疗保健中面向干预的LLM行为,并在心血管试点中进行了压力测试。该框架包含四个组成部分:(i)一个领域因果知识图谱,其中断言是具有稳定标识符的第一类、保留来源的节点;(ii)一个情景条件的子图提取步骤,给定任何临床情景,检索相关的具象化断言子图;(iii)四种受控的基础条件,变化检索到的子图如何被组成到模型的上下文中(无基础的 C1、知识图谱的 C2、因果图的 C3、综合的 C4);(iv)一个自动评分管道,基于断言标识符,计算干预准确性及其他评估指标。为了测试该框架,我们构建了一个跨越八种推理失败模式的类别平衡情景生成器,并在心血管图上进行了实例化。该指标面板沿可解释的、非冗余的轴线区分条件:C4获得了最强的因果边缘 F1(0.838)、不良反应 F1(0.833)、证据准确性(0.738)和不支持声明率(0.114),而 C1获得了最高的原始干预准确性(0.948),但没有可测量的因果或证据基础。
cs.AI / 101 / 2608.15389

Agentic-SQL Revisited: Autonomy-Based Taxonomy and Empirical Benchmark Analysis for LLM Text-to-SQL

重新审视Agentic-SQL:基于自主性的分类法与LLM文本到SQL的实证基准分析
Zhao, Changruo, Peng, Zujun, Tian, Yu, Liu, Yuting, Su, Yiyun, Zhu, Huiying, Zhang, Luyan, Zeng, Heming
Abstract
LLM-based Text-to-SQL progress is reported across heterogeneous benchmarks, backbones, and inference protocols, making cross-system comparison fragile. We reframe the field as a leaderboard aggregation: we collect the metrics authors themselves report and organize them along an inference-autonomy axis spanning constrained, in-context, iterative, agentic, and reasoning-internalized generation, with traceable provenance for every cell. To anchor the aggregation empirically, we run a focused case study on Spider, comparing 8B open-source backbones with and without chain-of-thought (CoT) supervision against few-shot DeepSeek~V3 and GLM-4 baselines. Four patterns emerge: Spider gains transfer unevenly to BIRD and Spider~2.0; autonomy buys robustness at non-trivial cost; reasoning internalization sits between answer-only decoding and externally orchestrated agents; and CoT gains concentrate on Hard and Extra-Hard queries. We release a Python harness mirroring the autonomy axis so that future methods can be added directly to the leaderboard.
Chinese Translation
基于LLM的文本到SQL的进展在异构基准、基础模型和推理协议中有所报道,这使得跨系统比较变得脆弱。我们将该领域重新框定为一个排行榜聚合:我们收集作者自己报告的指标,并沿着推理自主性轴进行组织,该轴涵盖了受限、上下文内、迭代、自主和推理内化生成,并为每个单元提供可追溯的来源。为了在实证上锚定聚合,我们对Spider进行了集中案例研究,比较了具有和不具有思维链(CoT)监督的8B开源基础模型与少量样本的DeepSeek~V3和GLM-4基准。出现了四种模式:Spider对BIRD和Spider~2.0的迁移不均;自主性以非微不足道的成本换取鲁棒性;推理内化介于仅答案解码和外部协调代理之间;而CoT的收益集中在困难和超困难查询上。我们发布了一个Python工具,反映自主性轴,以便未来的方法可以直接添加到排行榜中。
cs.AI / 102 / 2608.15391

TwinGridShield: Consequence-Aware Runtime Authorization for LLM Grid-Agent Actions

TwinGridShield:面向后果的LLM网格代理行为运行时授权
Rafy, Md Fazley
Abstract
Large language model (LLM)-assisted energy-management tools can translate natural-language context into structured grid commands, but syntactic validity does not imply physical admissibility. This paper presents TwinGridShield, a model-independent runtime authorization layer that evaluates each proposed action in a deterministic network twin before release. The prototype checks connectivity, branch-flow, generator, and load-shedding invariants and records each decision in a hash-chained log. A controlled IEEE 14-bus study evaluates single-step switching, redispatch, and load-shedding actions using DC power flow and experimentally assigned branch ratings. In the matched-model experiment, a stochastic proposal source configured to select an unsafe action with probability p=0.84 produced 421 unsafe proposals in 500 attacked-condition trials, a realized rate of 84.2%. This value characterizes the configured surrogate and is not an empirical measurement of LLM prompt-injection susceptibility. TwinGridShield produced 0 unsafe releases in those 500 trials. Because action labeling and authorization used the same DC model, system state, branch ratings, and encoded constraints, this result verifies conformance of the implementation to its encoded authorization predicate rather than safety under model error. The principal robustness evaluation therefore introduces model mismatch. Unsafe acceptance reached 5.63% under bounded +20% and -20% per-bus load-measurement error and 30.09% when actual branch ratings were 20% below modeled ratings.
Chinese Translation
大型语言模型(LLM)辅助的能源管理工具可以将自然语言上下文转换为结构化的电网指令,但语法有效性并不意味着物理可接受性。本文提出了TwinGridShield,这是一种模型无关的运行时授权层,在发布之前评估每个提议的行动在确定性网络双胞胎中的可行性。该原型检查连接性、支路流动、发电机和负荷削减不变性,并在哈希链日志中记录每个决策。通过控制的IEEE 14母线研究,评估了单步切换、重新调度和负荷削减行为,使用直流电力流和实验分配的支路额定值。在匹配模型实验中,一个随机提案源配置为以概率p=0.84选择不安全的行动,在500次攻击条件试验中产生了421个不安全提案,实际发生率为84.2%。该值表征了配置的替代品,而不是LLM提示注入易感性的经验测量。TwinGridShield在这500次试验中没有产生不安全的发布。由于行动标记和授权使用了相同的直流模型、系统状态、支路额定值和编码约束,因此该结果验证了实现与其编码授权谓词的一致性,而不是在模型错误下的安全性。因此,主要的鲁棒性评估引入了模型不匹配。在每个母线负荷测量误差±20%的限制下,不安全接受率达到了5.63%,而当实际支路额定值低于建模额定值20%时,不安全接受率达到了30.09%。
cs.AI / 103 / 2608.15392

Visible Reasoning and Indirect Prompt-Injection Monitorability Across English, Tamil, and Tanglish

可见推理与间接提示注入可监测性在英语、泰米尔语和唐格利什中的研究
G, Madhusudhanan
Abstract
Chain-of-thought monitoring is a potentially useful safety signal, but its reliability across languages and behavioral settings remains uncertain. In a small case study of eight manually verified synthetic scenarios, one model, one annotator, and one deterministic generation seed, I study API-visible reasoning during indirect prompt injection in Sarvam-105B across English, Tamil, and Tanglish. A four scenario pilot found 5/12 injected attack successes without reasoning and 1/11 with reasoning. A preregistered four-scenario follow-up reversed that direction, finding 2/12 attacks without reasoning and 3/12 with reasoning. With only four scenarios per phase, this design cannot distinguish a real reasoning-mode effect from prompt-specific variation or sampling noise. Across 20 non-empty injected-thinking traces, all 17 benign-correct outputs stated an intent to ignore the injection, while all three attack successes stated an intent to follow it. These descriptive observations provide a reproducible case study of behaviorally informative visible reasoning when it is available; they do not establish that reasoning mode improves safety, that visible reasoning is mechanistically faithful, or that the findings generalize beyond this configuration.
Chinese Translation
思维链监控是一种潜在的有用安全信号,但其在不同语言和行为设置中的可靠性仍然不确定。在对八个手动验证的合成场景进行的小型案例研究中,我研究了在 Sarvam-105B 模型中进行间接提示注入时的 API 可见推理,涵盖英语、泰米尔语和唐格利什。初步的四场景试点发现,在没有推理的情况下,12次注入攻击中成功了5次,而在有推理的情况下成功了1次。经过预注册的四场景后续研究则反转了这一方向,发现没有推理的攻击成功率为2/12,而有推理的攻击成功率为3/12。由于每个阶段仅有四个场景,这种设计无法区分真实的推理模式效应与特定提示的变异或抽样噪声。在20个非空的注入思维轨迹中,所有17个良性正确的输出均表示有意忽略注入,而所有3个攻击成功的输出则表示有意遵循注入。这些描述性观察提供了一个可重复的案例研究,展示了在可用时行为上有信息量的可见推理;但它们并未证明推理模式能提高安全性、可见推理在机制上是可信的,或这些发现能超出该配置进行推广。
cs.AI / 104 / 2608.15396

Large Language Model Assisted Operational Monitoring for Battery Energy Storage System Integrated Power Distribution Networks

大型语言模型辅助的电池储能系统集成电力配电网络的运行监测
Akhtar, Azmeer, Rafy, Md Fazley, Srivastava, Anurag K.
Abstract
Battery energy storage systems (BESS) are increasingly used in distribution networks for voltage regulation and demand response, which increases the volume and complexity of operational telemetry available to grid operators. This paper presents an AI-enabled monitoring framework that connects a large language model (LLM) interface with a structured telemetry database for BESS-integrated distribution system analysis. Operator questions are submitted in natural language and translated into validated SQL queries using predefined database schema information and approved KPI views. Retrieved measurements, including bus voltages, state of charge, active power, and reactive power, are evaluated against engineering constraints for voltage limits, BESS operation, and demand response tracking. The framework is validated using hardware-in-the-loop co-simulation data from a BESS-equipped distribution feeder operating under reactive power-based voltage control and price-driven demand response. Case studies show that the framework generates valid database queries, identifies repeated voltage violations, detects reactive power overshoot, and evaluates active-power tracking performance. The results show that LLM-assisted monitoring can connect structured grid telemetry with automated engineering assessment for BESS operation analysis.
Chinese Translation
电池储能系统(BESS)在配电网络中越来越多地用于电压调节和需求响应,这增加了电网操作员可用的运行遥测数据的数量和复杂性。本文提出了一种基于人工智能的监测框架,该框架将大型语言模型(LLM)接口与结构化遥测数据库连接,用于BESS集成配电系统分析。操作员的问题以自然语言提交,并使用预定义的数据库模式信息和批准的关键绩效指标(KPI)视图转换为经过验证的SQL查询。检索到的测量数据,包括母线电压、充电状态、主动功率和无功功率,依据电压限制、BESS操作和需求响应跟踪的工程约束进行评估。该框架使用基于无功功率的电压控制和价格驱动的需求响应下运行的配备BESS的配电馈线的硬件在环协同仿真数据进行验证。案例研究表明,该框架生成有效的数据库查询,识别重复的电压违规,检测无功功率超出,并评估主动功率跟踪性能。结果表明,LLM辅助的监测可以将结构化电网遥测与BESS操作分析的自动化工程评估相连接。
cs.AI / 105 / 2608.15400

Implementation of a Metacognition Framework for Self-Awareness and Self-Regulation in Ensembles of LLMs

大规模语言模型集成的元认知框架实施:自我意识与自我调节
Courchaine, Charles, Sethi, Ricky J., Qiu, Hefei
Abstract
Large Language Models (LLMs) are notorious for struggling with assessing their own uncertainty, detecting knowledge conflicts, or recognizing when problems exceed their expertise; such limitations inevitably undermine reliability and trust in LLMs. In this paper, we present the first implementation of a metacognitive framework for ensembles of LLMs that addresses these challenges through explicit monitoring and control mechanisms. Our system computes a Metacognitive State Vector (MSV) quantifying self-awareness for monitoring across five dimensions derived from cognitive psychology: Emotional Response, Correctness Evaluation, Experiential Match, Conflicting Information, and Problem Importance. MSV values also provide self-regulation for control, automatically switching between System 1 (fast, single- or multi-node) and System 2 (deliberative, multi-node) processing based on query complexity. For System 2 execution, graph-theoretic algorithms control the assignment of specialized roles (Domain Expert, Critic, Evaluator, Synthesizer, and Generalist) to ensemble nodes according to their MSV-quantified metacognitive states. Our implementation allows users to explore how different query types trigger distinct processing modes. The Proof-of-Concept (PoC) demo showcases the framework with illustrative examples showing appropriate System 1/System 2 routing and helps visualize the metacognitive process via real-time radar charts and decision indicators. This PoC implementation demonstrates the feasibility of creating a framework for metacognitive self-awareness and self-regulation in LLM systems.
Chinese Translation
大规模语言模型(LLMs)在评估自身不确定性、检测知识冲突或识别问题超出其专业知识时表现不佳,这些局限性不可避免地削弱了对LLMs的可靠性和信任。在本文中,我们首次提出了一种针对LLMs集成的元认知框架的实施,该框架通过明确的监控和控制机制来应对这些挑战。我们的系统计算元认知状态向量(Metacognitive State Vector, MSV),量化自我意识,以监控来自认知心理学的五个维度:情感反应、正确性评估、经验匹配、冲突信息和问题重要性。MSV值还提供自我调节的控制,基于查询复杂性自动在系统1(快速、单节点或多节点)和系统2(深思熟虑、多节点)处理之间切换。在系统2执行时,图论算法根据节点的MSV量化元认知状态控制专业角色(领域专家、评论员、评估者、合成者和通才)的分配。我们的实施允许用户探索不同查询类型如何触发不同的处理模式。概念验证(Proof-of-Concept, PoC)演示展示了该框架,通过示例展示适当的系统1/系统2路由,并通过实时雷达图和决策指标帮助可视化元认知过程。该PoC实施展示了在LLM系统中创建元认知自我意识和自我调节框架的可行性。
cs.AI / 106 / 2608.15411

A survey of AI-generated voices and their detection

AI生成声音及其检测的综述
Sun, Chengzhe, Yang, Tianle, Lyu, Siwei
Abstract
The ability of artificial intelligence (AI) models to generate highly realistic human voices has advanced rapidly. These technologies power accessibility tools, virtual assistants and creative applications, but they also enable harmful uses, including impersonation, fraud and disinformation. Recent incidents of voice cloning scams targeting businesses and political leaders underscore the urgent need for robust safeguards. Unlike image and video deepfakes, the detection of synthetic voices poses unique challenges due to the complexity of phonetics, prosody and auditory perception. This survey offers a comprehensive overview of AI voice generation and detection methods, encompassing both the technical foundations and the latest state-of-the-art advances. This study also identifies key open challenges, benchmark resources and future directions to make this survey useful for future researchers.
Chinese Translation
人工智能(AI)模型生成高度逼真的人声的能力迅速发展。这些技术为无障碍工具、虚拟助手和创意应用提供支持,但也使得一些有害用途成为可能,包括冒充、欺诈和虚假信息。最近针对企业和政治领袖的声音克隆诈骗事件突显了建立强有力的保护措施的紧迫性。与图像和视频深度伪造不同,合成声音的检测面临着独特的挑战,这源于语音学、韵律学和听觉感知的复杂性。本综述提供了AI声音生成和检测方法的全面概述,涵盖了技术基础和最新的前沿进展。本研究还识别了关键的开放挑战、基准资源和未来方向,以使本综述对未来研究者具有实用价值。
cs.AI / 107 / 2608.15432

Does the Proof Prove It That Way? Faithful Formalization of Elements Proofs

证明是否以这种方式证明?元素证明的忠实形式化
Mao, Tadd, Zhong, Tianjun, Arekar, Dhruva, Feng, Yuming, An, One, Huang, Jiani, Si, Xujie, Li, Ziyang
Abstract
In formal verification, both the autoformalization of statements and automated proof search have been studied extensively. While automated proof search can produce a formal proof that compiles, the generated proof does not necessarily reflect how the natural-language argument arrives at its conclusion--a property we refer to as faithfulness. With faithfully formalized proofs, one can check the reasoning behind a human- or AI-written argument, and assist mathematicians in formalizing their proof sketches. However, it is particularly challenging due to misalignment of formal proof tactics and natural language reasoning. In this work, we rigorously describe a set of five necessary conditions a faithful formal proof must satisfy, and introduce Pistis, an agentic, oracle-guided proof search that produces formal Lean proofs that satisfy them. At its core is a novel faithfulness-preserving divide-and-conquer search, which we name OrderDecompose, that tracks citation dependencies and blocks unfaithful shortcuts, paired with a refutation search, that surfaces gaps and errors in the natural language proof source. OrderDecompose completes proofs that baselines cannot close even within a 12-hour budget, and its artifacts compile over 33$\times$ as fast as prior work's. We apply Pistis on the first three books of Euclid's Elements, producing high-quality artifacts containing faithful formal proofs. Under a blinded human study and an LLM-as-a-judge protocol on rigorous rubrics, Pistis-generated proofs are favored over prior works--2.89$\times$ and 5.2$\times$ as often by human reviewers and the LLM judge, respectively. It further uncovers gaps in Euclid's proofs and their translation, and can accept or refute natural language proofs written by humans or AI, demonstrating that faithful formalization is useful as a proof-checking tool.
Chinese Translation
在形式验证中,陈述的自动形式化和自动证明搜索都得到了广泛研究。虽然自动证明搜索可以生成一个可编译的形式证明,但生成的证明不一定反映自然语言论证如何得出其结论——我们称之为忠实性。通过忠实形式化的证明,可以检查人类或人工智能撰写的论证背后的推理,并帮助数学家形式化他们的证明草图。然而,由于形式证明策略与自然语言推理之间的不一致,这一过程特别具有挑战性。在本研究中,我们严格描述了一组忠实形式证明必须满足的五个必要条件,并介绍了Pistis,一个代理驱动的、由神谕指导的证明搜索系统,能够生成满足这些条件的正式Lean证明。其核心是一个新颖的保持忠实性的分治搜索,我们称之为OrderDecompose,它跟踪引用依赖关系并阻止不忠实的捷径,配合一个反驳搜索,揭示自然语言证明来源中的缺口和错误。OrderDecompose完成了基线无法在12小时预算内关闭的证明,其生成的文档编译速度比之前的工作快超过33倍。我们在欧几里得《元素》的前三本书上应用Pistis,生成包含忠实形式证明的高质量文档。在一项盲人研究和基于严格标准的LLM作为评审的协议下,Pistis生成的证明在人工评审和LLM评审中分别比之前的工作更受青睐,比例为2.89倍和5.2倍。它进一步揭示了欧几里得证明及其翻译中的缺口,并能够接受或反驳人类或人工智能撰写的自然语言证明,证明忠实形式化作为证明检查工具是有用的。
cs.AI / 108 / 2608.15436

OTel: Building Domain-Specialized Telecom LLM Foundations for Intelligent Networks

OTel:为智能网络构建领域专用的电信大语言模型基础
Tavakkoli, Farbod, Paulk, Roderic, Terrazas, Jorden, Church, Kenneth, Austin, Mark, Powell, Louis, Diamos, Gregory, Bariah, Lina, Zaidi, Syed Ali Raza, Hafeez, Maryam, Maatouk, Ali, Karim, Imtiaz
Abstract
Frontier AI models have advanced rapidly, but they still struggle with telecom-specific tasks. We present Open Telco (OTel), an open telecom AI resource with derived datasets for retrieval, reranking, instruction tuning, and safety/abstention, plus 30 full-parameter post-trained baselines across embedding, reranking, and language models. The community has already engaged substantially with the resource: as of May 3, 2026, the released models have been downloaded over 16 million times, and the project has received 157+ pieces of media coverage worldwide. Building on prior open telecom datasets and benchmarks, OTel provides documented telecom data sources, held-out evaluation partitions, trained embedding models, rerankers, context-grounded LLMs, and safety/abstention data in one unified resource. OTel post-training improves performance across all three model families: embedding retrieval reaches 93.5% NDCG@10, reranking reaches 0.952 MRR@10, and language-model correctness reaches 88.2%. We release OTel as a reproducible starting point and invite the community to expand the data, improve embedding and reranking models, and build stronger context-grounded telecom LLMs.
Chinese Translation
前沿人工智能模型发展迅速,但在电信特定任务上仍面临挑战。我们提出了开放电信(Open Telco,OTel),这是一个开放的电信人工智能资源,包含用于检索、重排序、指令调优和安全/弃权的衍生数据集,以及30个全参数后训练基线模型,涵盖嵌入、重排序和语言模型。社区已经对该资源进行了大量参与:截至2026年5月3日,发布的模型已被下载超过1600万次,项目在全球范围内获得了157篇以上的媒体报道。在之前开放电信数据集和基准的基础上,OTel提供了文档化的电信数据源、保留的评估分区、训练好的嵌入模型、重排序器、基于上下文的LLM(大语言模型)以及安全/弃权数据,整合为一个统一的资源。OTel的后训练在所有三种模型家族中提升了性能:嵌入检索达到93.5%的NDCG@10,重排序达到0.952的MRR@10,语言模型的正确性达到88.2%。我们将OTel作为一个可重复的起点发布,并邀请社区扩展数据、改进嵌入和重排序模型,构建更强大的基于上下文的电信大语言模型。
cs.AI / 109 / 2608.15445

Measuring Reward Hacking and Reasoning-Answer Decoupling Under Position-Confounded Optimization

在位置混淆优化下测量奖励黑客行为和推理-答案解耦
Maniyar, Suyash, Sandhu, Armaan, Mishra, Abhishek
Abstract
When a reward is correct on every training example yet consistent with more than one goal, a model can acquire an unintended one, a failure known as goal misgeneralization. Endpoint accuracy on the training distribution cannot tell the two apart, because solving the task and exploiting a surface feature can satisfy the reward equally well. We treat this as a measurement problem: what does a benchmark score measure once a model has been optimized against a correct but confounded signal? We train language models with GRPO on multiple-choice math problems where the correct answer is always option A, then evaluate on an unseen test set with unbiased answer positions. Across Qwen2.5, Llama 3.x and Gemma-3 models, biased training often drives option-A rates above 0.90 in smaller models and collapses unbiased accuracy toward chance, so accuracy stops measuring math ability and instead measures an answer-position policy. We further find reasoning-answer decoupling: capable models generate reasoning that reaches the correct numeric answer while still selecting A. We track this with numeric extraction and an LLM judge (GPT-4.1-mini; Qwen2.5-3B decoupling rate is about 0.66). The broken construct generalizes beyond the training domain: biased models inflate A-rates on out-of-domain MMLU and value-laden prompts. Continued training on unbiased data reverses the in-domain shift unevenly and only partially reverses the out-of-domain one, so a model can appear restored on its training distribution while remaining biased on unseen inputs. Reasoning-answer decoupling rate, together with answer distributions and out-of-domain behavior, separates capability loss from a learned, transferable shortcut.
Chinese Translation
当奖励在每个训练样本上都是正确的,但与多个目标一致时,模型可能会获得一个意想不到的目标,这种失败被称为目标误概化。训练分布上的端点准确性无法区分这两者,因为解决任务和利用表面特征可以同样满足奖励。我们将此视为一个测量问题:一旦模型针对一个正确但混淆的信号进行了优化,基准分数到底测量了什么?我们使用 GRPO 训练语言模型,针对多项选择数学问题,其中正确答案始终是选项 A,然后在一个未见的测试集上进行评估,该测试集具有无偏的答案位置。在 Qwen2.5、Llama 3.x 和 Gemma-3 模型中,偏置训练通常使较小模型的选项 A 比率超过 0.90,并使无偏准确性接近随机,因此准确性不再测量数学能力,而是测量答案位置策略。我们进一步发现推理-答案解耦:有能力的模型生成的推理能够达到正确的数字答案,同时仍然选择 A。我们通过数字提取和 LLM 判别器(GPT-4.1-mini;Qwen2.5-3B 的解耦率约为 0.66)来跟踪这一点。这个破碎的构造超越了训练领域的普遍性:偏置模型在领域外的 MMLU 和价值导向提示上夸大了 A 比率。在无偏数据上持续训练不均匀地逆转了领域内的偏移,并且仅部分逆转了领域外的偏移,因此模型在其训练分布上看似恢复,但在未见输入上仍然存在偏见。推理-答案解耦率,以及答案分布和领域外行为,将能力损失与学习到的可转移捷径区分开来。
cs.AI / 110 / 2608.15451

Mental Model Management: An Operator-Based Framework for LLM Memory

心理模型管理:一种基于操作员的大型语言模型记忆框架
Kramer, Oliver
Abstract
Large language models process large amounts of information but usually lack an explicit mechanism for maintaining compact and evolving conceptual representations. We introduce Mental Model Management (3M), a framework in which knowledge is represented as mental models consisting of compact chunks. Rather than accumulating text passages, 3M continuously integrates new information into an existing conceptual representation. A set of operators extracts knowledge, retrieves relevant models, adds and updates chunks, reorganizes representations, detects inconsistencies, and derives new knowledge. We describe the main 3M operators and illustrate each operation using Evolution Strategies as a running example.
Chinese Translation
大型语言模型处理大量信息,但通常缺乏维持紧凑和不断发展的概念表示的明确机制。我们提出了心理模型管理(Mental Model Management, 3M),这是一个将知识表示为由紧凑块组成的心理模型的框架。3M并不是简单地积累文本段落,而是持续将新信息整合到现有的概念表示中。一组操作员提取知识、检索相关模型、添加和更新块、重组表示、检测不一致性并推导新知识。我们描述了主要的3M操作员,并以进化策略(Evolution Strategies)作为示例来说明每个操作。
cs.AI / 111 / 2608.15454

Dynamic Multi-Byte Prediction With Hierarchical Language Models

基于层次语言模型的动态多字节预测
Owodunni, Abraham Toluwase, Okocha, Chibuzor, Grant, Christan, Limisiewicz, Tomasz, Kumar, Sachin
Abstract
Byte-level hierarchical language models (LMs) have recently emerged as a robust alternative to their popular counterparts that use subword tokenization. However, generating one byte at a time remains a bottleneck for inference speed. To address this, we introduce multi-byte prediction (MBP), which generates multiple bytes in parallel, speeding up inference with minimal performance impact and no additional parameters. MBP builds on the popular multi-token prediction (MTP) paradigm with two crucial innovations. First, we introduce a variable-length prediction window that aligns with the latent tokens, or segments, of a hierarchical LM. Second, we implement a novel attention-masking scheme that enables parallel byte prediction without violating causality. We show that multi-byte prediction strikes a Pareto-optimal trade-off across multiple generative tasks, instruction following, question answering, summarization, and machine translation, achieving the best trade-off between performance and inference throughput.
Chinese Translation
字节级层次语言模型(LMs)最近作为一种强有力的替代方案出现,取代了使用子词标记化的流行模型。然而,一次生成一个字节仍然是推理速度的瓶颈。为了解决这个问题,我们引入了多字节预测(MBP),该方法可以并行生成多个字节,从而加快推理速度,且对性能影响最小且不增加额外参数。MBP建立在流行的多标记预测(MTP)范式之上,并具有两个关键创新。首先,我们引入了一个与层次LM的潜在标记或片段对齐的可变长度预测窗口。其次,我们实施了一种新颖的注意力掩蔽机制,使得并行字节预测在不违反因果性的情况下成为可能。我们展示了多字节预测在多个生成任务(如指令遵循、问答、摘要生成和机器翻译)中实现了帕累托最优的权衡,达到了性能与推理吞吐量之间的最佳平衡。
cs.AI / 112 / 2608.15488

A Network-driven Framework for Public Event Forecasting via Dynamic Interaction Network Evolution

基于网络驱动的公共事件预测框架:动态互动网络演化
Wei, Jie, Liu, Yue, Tang, Xiaochuan, Cai, Biao, Li, Xiangtao, Hu, Yanmei
Abstract
Effective public event forecasting is essential for intelligent service systems, enabling proactive risk management, adaptive resource allocation, and timely decision-making. In many real-world scenarios, the evolution of public events is driven by dynamic interactions among participants. Motivated by this observation, this paper proposes auto-ibDLM, a network-driven deep learning framework that represents events as dynamic interaction networks and predicts public event evolution through participant growth forecasting. The proposed framework adopts a hybrid representation learning strategy that first represents network evolution using network science-informed structural metrics and subsequently transforms the resulting structural feature vectors into compact and robust latent representations through an auto-learning layer. A GRU-based temporal forecasting module is then employed to capture temporal dependencies and predict future participant growth. Extensive experiments on 13 real-world public event datasets and two publicly available dynamic network datasets demonstrate that auto-ibDLM consistently outperforms representative state-of-the-art methods in both forecasting accuracy and generalization capability, achieving over 97% accuracy in public event forecasting. Comprehensive experimental analyses further validate the effectiveness of the proposed hybrid representation learning strategy and demonstrate its representation-level interpretability. These results indicate that auto-ibDLM provides an effective and practical solution for intelligent public event forecasting.
Chinese Translation
有效的公共事件预测对于智能服务系统至关重要,它能够实现主动风险管理、适应性资源分配和及时决策。在许多现实场景中,公共事件的演变受到参与者之间动态互动的驱动。基于这一观察,本文提出了auto-ibDLM,一个基于网络驱动的深度学习框架,该框架将事件表示为动态互动网络,并通过参与者增长预测来预测公共事件的演变。所提出的框架采用混合表示学习策略,首先使用网络科学指导的结构指标表示网络演变,然后通过自学习层将得到的结构特征向量转化为紧凑且稳健的潜在表示。接着,采用基于GRU的时间预测模块捕捉时间依赖性并预测未来的参与者增长。在13个真实世界公共事件数据集和两个公开可用的动态网络数据集上进行的广泛实验表明,auto-ibDLM在预测准确性和泛化能力方面始终优于代表性的最先进方法,实现了公共事件预测超过97%的准确率。全面的实验分析进一步验证了所提出的混合表示学习策略的有效性,并展示了其在表示层面的可解释性。这些结果表明,auto-ibDLM为智能公共事件预测提供了有效且实用的解决方案。
cs.AI / 113 / 2608.15502

EcoVLA: Energy-Efficient Device-Edge Co-Inference for Vision-Language-Action Models under Real-Time Constraints

EcoVLA:在实时约束下的能效设备-边缘协同推理用于视觉-语言-动作模型
Zhou, Ao, Dai, Bo, Yu, Le, Liu, Xingyu, Hao, Zeyu, Long, Lingkun, Hu, Chunming, Yang, Jianlei
Abstract
Vision-Language-Action (VLA) models have emerged as a promising foundation for Embodied AI, but their high inference cost poses significant challenges for deployment in robotic systems. In practice, on-device inference is constrained by limited compute capacity and energy budgets, struggling to simultaneously satisfy real-time control and energy efficiency requirements. Alternatively, offloading the inference workload to an edge server is susceptible to fluctuations in system conditions, introducing unpredictable latency risks. Device-edge co-inference offers a promising solution, but systematic research tailored to VLA models remains scarce, particularly a unified co-inference framework that jointly addresses real-time constraints and system-level energy efficiency. Thus, we propose EcoVLA, an adaptive device-edge co-inference framework for VLA models that maximizes system energy efficiency under real-time constraints. EcoVLA first introduces a unified stage-level abstraction over different VLA paradigms, establishing an architecture-agnostic co-inference design space. It then formulates a joint device-edge-network latency and energy prediction model to enable rapid runtime evaluation of candidate co-inference schemes. Building on this, EcoVLA continuously selects the energy-optimal scheme satisfying real-time constraints with millisecond-level overhead, adapting to runtime variations in network and system states. Furthermore, EcoVLA incorporates a lightweight transmission mechanism for inter-stage intermediate tensors to reduce the communication overhead incurred by cross-device collaboration. Experimental results across VLA models show that EcoVLA improves system energy efficiency by up to 236% over existing co-inference approaches under a 20 Hz action output frequency constraint, while consistently maintaining SLO satisfaction under dynamic network and edge workload conditions.
Chinese Translation
视觉-语言-动作(VLA)模型已成为具身人工智能的有希望基础,但其高推理成本对机器人系统的部署构成了重大挑战。在实践中,设备上的推理受到有限计算能力和能量预算的限制,难以同时满足实时控制和能效要求。另一方面,将推理工作负载卸载到边缘服务器上容易受到系统条件波动的影响,带来不可预测的延迟风险。设备-边缘协同推理提供了一种有前景的解决方案,但针对VLA模型的系统研究仍然稀缺,特别是缺乏一个统一的协同推理框架,能够共同解决实时约束和系统级能效。因此,我们提出了EcoVLA,一个适应性设备-边缘协同推理框架,旨在在实时约束下最大化VLA模型的系统能效。EcoVLA首先引入了不同VLA范式的统一阶段级抽象,建立了一个与架构无关的协同推理设计空间。然后,它制定了一个联合设备-边缘-网络延迟和能量预测模型,以便快速评估候选协同推理方案的运行时性能。在此基础上,EcoVLA持续选择满足实时约束的能量最优方案,延迟仅为毫秒级,能够适应网络和系统状态的运行时变化。此外,EcoVLA还结合了一种轻量级的传输机制,用于跨阶段中间张量的传输,以减少跨设备协作所带来的通信开销。实验结果表明,在20 Hz动作输出频率约束下,EcoVLA在现有协同推理方法的基础上提高了系统能效高达236%,同时在动态网络和边缘工作负载条件下始终保持服务水平目标的满足。
cs.AI / 114 / 2608.15510

Who Leads Now? Token-Level Modality Arbitration for Chart-to-Code Generation

现在谁在主导?图表到代码生成的标记级模态仲裁
Fu, Qinghao, Wang, Yarong, Ning, Shunlei, Wang, Yilin, Bai, Shunwen, Wang, Xinda, Wang, Jiaotuan, Nie, Yinan, Zhou, Wei
Abstract
Chart-to-code generation requires a model to read the fine-grained visual details of a chart and write executable code that reproduces it. Existing chart-to-code methods either train visual and coding abilities separately, or fine-tune on chart-to-code data with the two abilities entangled. Neither strategy accounts for the distinct nature of the two abilities or the interference that arises when they are optimized together. We propose MoCA (Mixture of Cross-modal Arbitration), which separates the two abilities rather than blending them. MoCA is built on Cross-modal Arbitration Block (CAB), which maintains a visual branch and a code branch as two distinct pathways, and a lightweight arbiter that arbitrates their relative contributions at every layer and generated token. We train MoCA in two stages: a supervised warm-up on self-distilled reasoning trajectories that decomposes visual understanding into explicit steps, followed by reinforcement learning with rewards on both the reasoning process and the final code. Analysis shows that the arbiter learns structured rather than arbitrary allocations, with expert contributions varying systematically across tokens, layers, and instances. Across three benchmarks, MoCA delivers competitive performance against general-domain and chart-specialized models. Ablation results show that the gains cannot be attributed to a larger model size alone, but instead arise from the joint contributions of complementary visual and code branch initialization and input-conditioned arbitration through CAB.
Chinese Translation
图表到代码生成需要模型读取图表的细粒度视觉细节,并编写可执行代码以重现该图表。现有的图表到代码方法要么分别训练视觉和编码能力,要么在图表到代码数据上进行微调,但这两种能力是交织在一起的。这两种策略都没有考虑到这两种能力的独特性质或在共同优化时产生的干扰。我们提出了 MoCA(跨模态仲裁混合),它将这两种能力分开,而不是将它们混合。MoCA 基于跨模态仲裁块(CAB),该块保持视觉分支和代码分支作为两个独立的路径,并且有一个轻量级仲裁者在每一层和生成的标记上仲裁它们的相对贡献。我们在两个阶段训练 MoCA:首先在自我蒸馏推理轨迹上进行监督热身,这些轨迹将视觉理解分解为明确的步骤,然后通过强化学习对推理过程和最终代码进行奖励。分析表明,仲裁者学习到的是结构化而非任意的分配,专家贡献在标记、层和实例之间系统性变化。在三个基准测试中,MoCA 在与通用领域和图表专用模型的比较中表现出竞争力。消融实验结果表明,性能提升不能仅归因于模型规模的增大,而是源于互补的视觉和代码分支初始化以及通过 CAB 进行的输入条件仲裁的共同贡献。
cs.AI / 115 / 2608.15536

From Contexts to Values: Context-Dependent Defeat in Abstract Argumentation

从语境到价值:抽象论证中的语境依赖性失败
Sadowski, Albert, Chudziak, Jarosław A.
Abstract
In value-based argumentation, an audience's ordering of values decides which attacks succeed as defeats. In many settings the deciding factor is not the audience but the circumstances: the same attack may succeed at one procedural stage, or under one regulation, and fail at another. Context-dependent argumentation frameworks (CDAFs), a model we recently introduced, capture this directly: one set of arguments, one attack relation, and a defeat function that switches each attack on or off per context, so every context induces an ordinary Dung framework. This raises a reduction question: is context genuinely new, or can one value assignment with per-context orderings reproduce the defeat function, collapsing the CDAF into a VAF? We present a polynomial-time decision procedure for this question and map the harder neighbouring problems, with upper bounds from NP to $\Sigma^p_3$. We also present a validated reference implementation and a measurement: representability is rare and falls fast with the number of contexts.
Chinese Translation
在基于价值的论证中,观众对价值的排序决定了哪些攻击成功为失败。在许多情况下,决定因素不是观众而是环境:同一攻击可能在一个程序阶段或在一个规定下成功,而在另一个阶段或规定下失败。我们最近引入的语境依赖性论证框架(CDAFs)直接捕捉了这一点:一组论证、一种攻击关系,以及一个在每个语境下切换每个攻击的失败函数,因此每个语境都诱导出一个普通的Dung框架。这引发了一个归约问题:语境是否真的新颖,或者一个具有每个语境排序的价值分配是否可以重现失败函数,从而将CDAF简化为VAF?我们为这个问题提出了一个多项式时间的决策程序,并映射了更复杂的邻近问题,提供了从NP到$ ext{Σ}^p_3$的上界。我们还展示了一个经过验证的参考实现和一个测量:可表示性是稀有的,并且随着语境数量的增加迅速下降。
cs.AI / 116 / 2608.15546

ATLAS: Scaffold-Free Algorithm Synthesis by LLMs via Embedding-Guided Quality-Diversity Search

ATLAS:通过嵌入引导的质量-多样性搜索实现无支架算法合成的算法
Yazdani, Danial, Omidvar, Mohammad Nabi, Sun, Yuan, Ibrahimov, Maksud, Li, Xiaodong
Abstract
Most LLM-based automated algorithm design methods optimize a designated component within a human-specified scaffold, fixing overall organization and component interactions. We present ATLAS, an embedding-guided quality-diversity framework for scaffold-free full-algorithm synthesis in combinatorial optimization. The problem specification supplies objectives and constraints; a minimal I/O interface fixes only instance and solution formats; the LLM chooses and restructures components, interactions, and control flow. This freedom enlarges the search space, risking invalid candidates and premature convergence to one design region. ATLAS independently detects execution, interface, and feasibility failures, recomputes objectives, and applies error-conditioned repair; similarity-based archive management preserves algorithms across embedding-space regions to counter premature convergence. Its three-layer search refines the best design, gives other regions dedicated refinement opportunities, and performs cross-region synthesis to recombine components and their interactions. Across four NP-hard problems, ATLAS outperforms several state-of-the-art component-synthesis methods and a matched full-synthesis baseline while remaining competitive with strong human-designed algorithms. One ATLAS run retains several algorithms with comparable performance from distinct embedding-space regions rather than a single design. Code inspection finds that these multi-component designs differ in their primary construction or global-search backbone. Our results suggest that embedding-guided quality-diversity search can make the enlarged full-algorithm design space practically searchable. Source code and exact executable prompts are available at .
Chinese Translation
大多数基于大型语言模型(LLM)的自动化算法设计方法在人工指定的支架内优化特定组件,固定整体组织和组件交互。我们提出了ATLAS,一种用于组合优化的无支架全算法合成的嵌入引导质量-多样性框架。问题规范提供了目标和约束;最小的输入/输出接口仅固定实例和解决方案格式;LLM选择并重构组件、交互和控制流。这种自由扩大了搜索空间,增加了无效候选解和过早收敛到某一设计区域的风险。ATLAS独立检测执行、接口和可行性失败,重新计算目标,并应用基于错误的修复;基于相似性的档案管理在嵌入空间区域中保留算法,以对抗过早收敛。其三层搜索精炼最佳设计,为其他区域提供专门的精炼机会,并执行跨区域合成以重新组合组件及其交互。在四个NP难题中,ATLAS的表现优于几种最先进的组件合成方法和一个匹配的全合成基线,同时与强大的人工设计算法保持竞争力。一轮ATLAS运行保留了来自不同嵌入空间区域的多个性能相当的算法,而不是单一设计。代码检查发现这些多组件设计在其主要构造或全局搜索骨干上存在差异。我们的结果表明,嵌入引导的质量-多样性搜索可以使扩大后的全算法设计空间在实践中可搜索。源代码和确切的可执行提示可在获取。
cs.AI / 117 / 2608.15565

Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling

无答案的接纳:基于LLM的优化建模的无标签认证与经验学习
Lian, Junbo Jacob, Chen, Huiling, Qin, Hanzhang, Teo, Chung-Piaw
Abstract
Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking against known answers, which real ticket streams do not provide. The natural label-free alternatives are unreliable: on a 300-problem label-blind stream, admitting every executable model poisons roughly one admission in four, while single-instance agreement accepts models that match at one value but differ elsewhere. We propose AdmitOR, an admission gate built on calibrated external behavioral evidence. Candidates from three model families, prompting strategies, and solver stacks are run on instances resampled from an extracted parameter domain; agreement across the resulting value-function traces is summarized by a cross-family clique, and a calibrated threshold returns accept, abstain, or escalate. The preregistered false-discovery criterion holds on calibration data but not on the wild stream. We report this negative result in full and trace most failures to benchmark texts that do not faithfully encode their labeled instances. Comparing four admission judges on one collection of logs inside a state-of-the-art skill learner, AdmitOR raises admission precision to 0.927, against 0.871 for majority vote and 0.726 for execution success, yielding 3.1x and 8.0x fewer poisoned admissions. Its library is the smallest and attains the highest macro accuracy across five public benchmarks, 58.4 against 54.8 for majority vote and 53.9 for the ground-truth-labeled library. The 3.5-point gain over majority vote is supported by a paired bootstrap and survives correction for a host-side anomaly. To our knowledge, AdmitOR is the first label-free admission mechanism designed around an explicitly calibrated false-discovery target. The transfer failure identifies a necessary condition for extending it to wild streams.
Chinese Translation
用于优化建模的经验学习代理通过存储经过验证的技能来提高性能,但现有学习者通过检查已知答案来承认知识,而真实的票据流并不提供这些答案。自然的无标签替代方案是不可靠的:在一个300问题的无标签流中,承认每个可执行模型大约会使四分之一的接纳受到污染,而单实例一致性则接受在一个值上匹配但在其他地方不同的模型。我们提出了AdmitOR,一个基于校准外部行为证据的接纳门。来自三种模型家族、提示策略和求解器堆栈的候选者在从提取的参数域重新采样的实例上运行;通过跨家族的团体总结结果值函数轨迹的一致性,校准阈值返回接纳、弃权或升级。在校准数据上,预注册的假发现标准成立,但在野生流上不成立。我们全面报告这一负面结果,并追踪大多数失败归因于未忠实编码其标记实例的基准文本。在一个最先进的技能学习者内部比较四个接纳评判者时,AdmitOR将接纳精度提高到0.927,而多数投票为0.871,执行成功为0.726,减少了3.1倍和8.0倍的污染接纳。其库是最小的,并在五个公共基准中获得了最高的宏观准确率,58.4对比多数投票的54.8和真实标记库的53.9。相较于多数投票的3.5点增益得到了配对自助法的支持,并在主机端异常的校正中存活。根据我们的知识,AdmitOR是第一个围绕明确校准的假发现目标设计的无标签接纳机制。转移失败识别了将其扩展到野生流的必要条件。
cs.AI / 118 / 2608.15580

From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM

从通用到专业:一种用于内窥镜息肉报告的上下文融合框架与冻结的 VLM
Yang, Ruijie, Zhu, Yan, Fu, Peiyao, Li, Siyuan, Luo, Te, Wang, Zhihua, Li, Quanlin, Zhou, Pinghong, Yang, Xian, Wang, Shuo
Abstract
Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically meaningful morphological description within a single record. General-purpose vision-language models (VLMs) offer a unified interface for image understanding and report generation. Existing specialization strategies, however, typically rely on task-specific models or model-weight adaptation, leaving unresolved how to introduce reliable specialist knowledge while preserving both this unified interface and the VLM's pretrained capabilities. We introduce a context-fusion framework that specializes a frozen general-purpose VLM through both implicit instruction context and explicit transduction context without modifying its pretrained weights. Specifically, a self-supervised polyp encoder retrieves related image-report pairs as explicit, query-specific evidence, while learned continuous specialist tokens provide implicit instruction context shared across cases. Experiments were conducted on 2,056 expert-annotated public endoscopic images. We compared the framework with general-purpose VLMs, task-specific predictors, and weight-adaptation methods to assess specialist performance, unified reporting, and adaptation efficiency. Across numerical, categorical, and report-generation metrics, the proposed framework substantially improved direct frozen-VLM inference and achieved the strongest overall performance among the evaluated methods. It added trainable parameters equal to only 0.006% of the frozen VLM's parameter count. When the top-1 retrieved case carried the correct target category, our framework corrected 70.5% of the errors made by a weight-adaptation baseline. These findings support the context-fusion framework as a lightweight and effective strategy for specialist adaptation of a frozen VLM.
Chinese Translation
可靠的内窥镜息肉报告需要在单一记录中整合定量病变大小、标准化的巴黎分类和具有临床意义的形态描述。通用视觉-语言模型(VLM)提供了图像理解和报告生成的统一接口。然而,现有的专业化策略通常依赖于特定任务的模型或模型权重的适应,尚未解决如何在保持统一接口和 VLM 预训练能力的同时引入可靠的专业知识。我们提出了一种上下文融合框架,通过隐式指令上下文和显式转导上下文对冻结的通用 VLM 进行专业化,而无需修改其预训练权重。具体而言,自监督息肉编码器检索相关的图像-报告对作为显式的查询特定证据,而学习到的连续专业令牌提供了跨案例共享的隐式指令上下文。我们在 2,056 张专家标注的公共内窥镜图像上进行了实验。我们将该框架与通用 VLM、特定任务预测器和权重适应方法进行了比较,以评估专业性能、统一报告和适应效率。在数值、类别和报告生成指标上,所提出的框架显著改善了直接冻结 VLM 推理,并在评估方法中实现了最强的整体性能。它增加的可训练参数仅占冻结 VLM 参数总数的 0.006%。当检索到的前一案例包含正确的目标类别时,我们的框架纠正了 70.5% 权重适应基线所犯的错误。这些发现支持上下文融合框架作为一种轻量且有效的策略,用于冻结 VLM 的专业适应。
cs.AI / 119 / 2608.15591

Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback

Agent Gym:一个通过人机交互反馈实现大型语言模型代理的持续评估与演化框架
Omran, Pouya Ghiasnezhad, Zimmermann, Michael, Cambridge, Duncan, Kapoor, Ashmita, Dixit, Tanya
Abstract
Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deployment time, while the business rules and edge cases it must handle continue to evolve. Existing approaches address agent construction and one-time evaluation but provide no structured mechanism for continuous post-deployment behavioral correction without modifying the agent's source code. Most of the approaches offered in the market, require intense collection of logs and traces, and re-examining the agent design by the engineering team, a process which is heavy, long and negates the economical value of agentic transformation. We introduce Agent Gym, a modular, domain-agnostic framework that wraps any existing LLM-based agent in a continuous evaluation-and-evolution loop. The framework provides six composable capabilities --- Act, Evaluate, Investigate, Correct, Learn, and Observe --- organized across three architectural zones: a constitution layer that codifies domain knowledge in configuration artifacts, a runtime inference pipeline that chains acting, investigation, and adaptive correction, and a learning loop that enables subject matter experts to discover and validate new correction rules through natural language interaction. The key technical contributions include a hybrid deterministic-LLM correction engine with 21 condition operators and three-tier actions, a three-layer investigation architecture for ground-truth-free compliance validation, and a programmatic safety loop that guarantees rule correctness before human approval. We further introduce the Spec-to-Note Gap, an autoencoder-inspired view of agentic system transparency. An open-source reference implementation for invoice processing demonstrates that the framework is fully operational and ready for adoption.
Chinese Translation
在生产环境中部署的大型语言模型(LLM)代理面临着一个基本的矛盾:代理的行为在部署时被冻结,而必须处理的业务规则和边缘案例则持续演变。现有的方法主要关注代理的构建和一次性评估,但未提供结构化机制以在不修改代理源代码的情况下进行持续的部署后行为修正。市场上大多数方法需要大量收集日志和追踪信息,并由工程团队重新审查代理设计,这一过程繁重且耗时,削弱了代理转型的经济价值。我们提出了Agent Gym,一个模块化的、领域无关的框架,将任何现有的基于LLM的代理包裹在一个持续评估与演化的循环中。该框架提供六种可组合的能力——行动(Act)、评估(Evaluate)、调查(Investigate)、修正(Correct)、学习(Learn)和观察(Observe),并在三个架构区域中组织:一个将领域知识编码为配置工件的宪法层,一个将行动、调查和自适应修正串联起来的运行时推理管道,以及一个通过自然语言交互使主题专家能够发现和验证新修正规则的学习循环。关键技术贡献包括一个具有21个条件操作符和三层动作的混合确定性-LLM修正引擎,一个用于无真实基准合规验证的三层调查架构,以及一个在人工批准之前保证规则正确性的程序化安全循环。我们进一步引入了Spec-to-Note Gap,一个受自编码器启发的代理系统透明度视角。一个用于发票处理的开源参考实现展示了该框架的完全可操作性和准备采用的状态。
cs.AI / 120 / 2608.15592

When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction

当熵不足够时:重拾大型语言模型输出长度预测中的丢失语义
Ren, Feiyang, Wen, Shengtao, Guo, Lingbing, Tian, Yu, Cui, Yuanning, Chen, Xiang
Abstract
Efficient LLM serving is often bottlenecked by the need to pad sequences to a fixed maximum length, and this wastes compute and degrades throughput. Predicting output lengths in advance makes it possible to adopt length-aware scheduling, and this reduces the overhead. This advantage is especially pronounced in long-context reasoning and reinforcement learning applications. Existing approaches, such as entropy-guided token pooling, use token-wise entropy as their primary signal, but they tend to ignore differences in semantic content across tokens. So, important tokens are often underweighted, and tokens carrying little information receive disproportionate emphasis. This hurts the reliability of length prediction. We introduce ESTP (Entropy-and-Semantic Token Pooling), a lightweight framework that addresses this issue by combining entropy with attention-based importance scores. These scores are derived directly from the self-attention weights computed during the LLM prefill phase, and this allows ESTP to capture both uncertainty and semantic importance with minimal additional computation. Since the framework reuses prefill activations, it adds almost no extra memory overhead and introduces only minimal latency. On the ForeLen benchmark, ESTP outperforms baseline methods, achieves better prediction accuracy and lower error rates in most scenarios. When integrated with a length-aware scheduler in end-to-end system tests, it further helps improve overall throughput and reduce the padding ratio. Our results offer a practical and effective building block for length-aware LLM serving systems.
Chinese Translation
高效的大型语言模型(LLM)服务常常受到将序列填充到固定最大长度的需求的瓶颈,这会浪费计算资源并降低吞吐量。提前预测输出长度使得采用基于长度的调度成为可能,从而减少了开销。这一优势在长上下文推理和强化学习应用中尤为明显。现有方法,如基于熵的令牌池化,主要使用逐令牌熵作为信号,但往往忽视了不同令牌之间语义内容的差异。因此,重要的令牌常常被低估,而携带少量信息的令牌却受到不成比例的强调。这影响了长度预测的可靠性。我们提出了ESTP(熵与语义令牌池化),一个轻量级框架,通过将熵与基于注意力的重要性评分相结合来解决这一问题。这些评分直接来源于在LLM预填充阶段计算的自注意力权重,从而使ESTP能够以最小的额外计算捕捉不确定性和语义重要性。由于该框架重用预填充激活,因此几乎不增加额外的内存开销,并且仅引入最小的延迟。在ForeLen基准测试中,ESTP超越了基线方法,在大多数场景中实现了更好的预测准确性和更低的错误率。当与长度感知调度器集成在端到端系统测试中时,它进一步帮助提高整体吞吐量并减少填充比例。我们的结果为基于长度的大型语言模型服务系统提供了一个实用有效的构建模块。
cs.AI / 121 / 2608.15594

TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation

TRACE:基于轨迹的多轮对抗性对话评估推理
Miah, Md Messal Monem, Anika, Adrita, Yu, Zhiyuan, Huang, Ruihong
Abstract
Multi-turn jailbreak attacks have emerged as a critical safety threat to LLMs, as harmful objectives are decomposed across a sequence of apparently benign turns to bypass guardrails. Existing defenses lack the reasoning capacity to identify evolving manipulation patterns, often trading helpfulness for safety by over-refusing benign requests related to sensitive topics. We introduce Trace, a multi-turn defense with trajectory-aware structured reasoning. Before generating each response, the model identifies manipulation cues from the trajectory, evaluates both the benign and adversarial interpretations of user intent, assigns a jailbreak score, and commits to an action: Allow, Caution, or Decline. We curate 4k multi-turn adversarial conversations from five attack frameworks, pair them with 2.4k benign dialogs, and 600 sensitive-but-benign conversations. We train Llama-3.1-8B-Instruct with SFT and GRPO under a multi-component reward that jointly optimizes helpfulness on benign prompts and robustness against jailbreak attempts. Across seven multi-turn attack benchmarks, Trace attains an average attack success rate (ASR) of 14.5% against 31.4% for the strongest baseline and 74.9% for the undefended target, while significantly raising the attacker effort required per successful jailbreak. Trace also balances usability and safety, achieving a 93.3% average compliance on over-refusal benchmarks.
Chinese Translation
多轮越狱攻击已成为大型语言模型(LLMs)的一个关键安全威胁,因为有害目标通过一系列看似无害的对话轮次被分解,以绕过安全防护措施。现有的防御机制缺乏识别不断演变的操控模式的推理能力,往往通过过度拒绝与敏感话题相关的无害请求来在有用性与安全性之间进行权衡。我们提出了Trace,这是一种具有轨迹感知结构化推理的多轮防御机制。在生成每个响应之前,该模型从轨迹中识别操控线索,评估用户意图的无害和对抗性解释,分配越狱评分,并决定采取的行动:允许、谨慎或拒绝。我们从五个攻击框架中策划了4000个多轮对抗性对话,并将其与2400个无害对话和600个敏感但无害的对话配对。我们在多组件奖励下训练了Llama-3.1-8B-Instruct,联合优化无害提示的有用性和抵御越狱尝试的鲁棒性。在七个多轮攻击基准测试中,Trace的平均攻击成功率(ASR)为14.5%,而最强基线为31.4%,未防御目标为74.9%,同时显著提高了每次成功越狱所需的攻击者努力。Trace还在可用性和安全性之间取得平衡,在过度拒绝基准测试中实现了93.3%的平均合规率。
cs.AI / 122 / 2608.15600

VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation

VARM-Bench:在中文辱骂言论管理中验证结构化推理的基准测试
Yuan, Mingyu, Wen, Shengtao, Guo, Lingbing, Bi, Zhen, Chen, Xiang
Abstract
The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text. Existing Chinese benchmarks support label classification, fine-grained toxicity categorization, and target-aware extraction, but do not provide a unified representation for deterministically verifying the stated basis of a moderation decision. We introduce VARM-Bench, a benchmark for field-anchored chain-of-thought rationales in Chinese abusive-speech moderation. Each instance contains a concise natural-language rationale with explicit anchors for six decisions: target, target type, target explicitness, author stance, harmfulness label, and fine-grained category. Our deterministic protocol evaluates field correctness, target alignment, output validity, complete-record agreement, and hidden record errors conditioned on correct final decisions, without relying on an LLM judge. Under a common structured-output protocol, we evaluate language models across multiple model families using zero-shot prompting, taxonomy guidance, and structured CoT supervision, and analyze lexical-cue sensitivity and field-level errors. Results show that strong label-level performance can conceal substantial errors in complete moderation records. VARM-Bench provides an auditable and reproducible benchmark for evaluating verifiable moderation rationales in Chinese abusive-speech moderation.
Chinese Translation
辱骂性在线内容的广泛传播增加了对中文社交媒体文本可靠管理的需求。现有的中文基准支持标签分类、细粒度毒性分类和目标感知提取,但未提供统一的表示方式以确定性地验证管理决策的依据。我们引入了VARM-Bench,这是一个针对中文辱骂言论管理中基于领域的推理链的基准测试。每个实例包含一个简洁的自然语言推理,并为六个决策提供明确的锚点:目标、目标类型、目标显性、作者立场、有害性标签和细粒度类别。我们的确定性协议评估领域正确性、目标一致性、输出有效性、完整记录一致性以及基于正确最终决策的隐藏记录错误,而不依赖于大型语言模型(LLM)评判。在一个共同的结构化输出协议下,我们使用零样本提示、分类指导和结构化链式推理(CoT)监督对多个模型家族的语言模型进行评估,并分析词汇线索敏感性和领域级错误。结果表明,强大的标签级表现可能掩盖完整管理记录中的重大错误。VARM-Bench提供了一个可审计和可重复的基准,用于评估中文辱骂言论管理中的可验证管理推理。
cs.AI / 123 / 2608.15619

Bias-Corrected Ceilings of Emotion Predictability from Human Label Variation Based on Instance-Level Fano Bounds

基于实例级法诺界限的人类标签变异的情感可预测性偏差校正上限
Inoshita, Keito
Abstract
Emotion recognition from text keeps improving on benchmarks, yet whether an accuracy ceiling has been reached is seldom asked with discipline. Our aim is not to pin this ceiling to a single number, but to quantify how far it depends on finite annotation, estimator choice, annotation noise, and the evaluation protocol, and thereby to discipline how confidently saturation can be claimed. We propose Bias-corrected Affective Ceiling Estimation (BACE), an analysis framework that estimates a bias-corrected ceiling, separates irreducible from reducible error, and disciplines the resulting claims. An anchored Dirichlet-mixture empirical Bayes estimator, bracketed between plug-in and NSB, recovers the human-consensus distribution; an annotator split, a noise deconvolution, and a fixed claim gate then attribute error without circularity. Methodologically, unconstrained point estimates place reachability anywhere from 0.38 to 1.03, so saturation cannot be decided by any single estimator. Substantively, the only assertion passing the claim gate is that at least about 33% of a representative classifier's error on GoEmotions is irreducible, with the same pattern recurring on offensiveness and irony.
Chinese Translation
从文本中进行情感识别的性能在基准测试中不断提高,但关于是否已达到准确性上限的问题却鲜有严谨的探讨。我们的目标不是将这一上限固定为一个单一的数字,而是量化它在多大程度上依赖于有限的标注、估计器选择、标注噪声和评估协议,从而规范对饱和度的自信声明。我们提出了偏差校正情感上限估计(Bias-corrected Affective Ceiling Estimation, BACE),这是一个分析框架,旨在估计偏差校正的上限,区分不可约误差与可约误差,并规范由此产生的声明。一个锚定的Dirichlet混合经验贝叶斯估计器,介于插件和NSB之间,恢复了人类共识分布;通过标注者分割、噪声解卷积和固定声明门,随后对误差进行归因而不产生循环。方法上,无约束的点估计将可达性范围从0.38到1.03,因此无法通过任何单一估计器决定饱和度。在实质上,唯一通过声明门的主张是,代表性分类器在GoEmotions上的误差中至少约有33%是不可约的,且在冒犯性和讽刺性上也出现了相同的模式。
cs.AI / 124 / 2608.15621

Rotation-Invariant Multi-IMU Activity Recognition under Independent Per-Location Orientation Shifts

具有独立位置方向偏移的旋转不变多IMU活动识别
Baek, Seungyeol, Chai, Yoonbyung, Lee, Yonghyeon, Choi, Sungjoon, Suh, Sungho
Abstract
Human Activity Recognition (HAR) with self-administered wearables, such as at-home rehabilitation and exercise monitoring, often requires reattaching inertial measurement units (IMUs) across sessions. In multi-IMU settings, this can induce independent orientation offsets across body locations, a deployment shift that conventional scalar HAR models do not structurally handle. Existing remedies rely on rotation augmentation, whose robustness depends on sampled transformations, or calibration and orientationnormalization pipelines requiring additional reference-frame assumptions or explicit procedures. We present Truly Rotation-Invariant HAR (TRI-HAR), a rotation-invariant framework that makes robustness to independent per-location IMU orientation offsets a structural model property. TRI-HAR reshapes accelerometer and gyroscope streams into triaxial vectors, applies a shared SO(3)-equivariant backbone and invariant projection to each IMU location, and fuses the resulting invariant features for activity classification. Across four multi-IMU benchmarks, TRI-HAR preserves macro-F1 under fixed independent per-location SO(3) rotations and outperforms rotation-augmented baselines under this target shift without requiring rotational augmentation.
Chinese Translation
使用自我管理的可穿戴设备进行人类活动识别(HAR),例如家庭康复和运动监测,通常需要在不同会话之间重新附加惯性测量单元(IMU)。在多IMU设置中,这可能导致身体各部位之间独立的方向偏移,这种部署偏移是传统标量HAR模型无法结构性处理的。现有的解决方案依赖于旋转增强,其鲁棒性取决于采样的变换,或者依赖于需要额外参考框架假设或明确程序的校准和方向归一化流程。我们提出了真正的旋转不变HAR(TRI-HAR),这是一个旋转不变框架,使得对独立位置IMU方向偏移的鲁棒性成为结构模型属性。TRI-HAR将加速度计和陀螺仪的数据流重塑为三轴向量,应用共享的SO(3)等变主干和不变投影到每个IMU位置,并融合得到的不变特征以进行活动分类。在四个多IMU基准测试中,TRI-HAR在固定独立位置SO(3)旋转下保持宏观F1得分,并在这一目标偏移下超越了旋转增强的基线,而无需旋转增强。
cs.AI / 125 / 2608.15634

Argumentation for Common Ground: Finding Zones of Possible Agreement between Individuals in Conflict

寻求共同立场的论证:在冲突中的个体之间寻找可能的协议区域
Cavatorta, Elisa, Rago, Antonio
Abstract
How can common ground between societies in conflict be identified when citizens' acceptability of peace agreements is shaped by contested narratives? Such acceptability is mediated not only by the clauses that agreements include or exclude, but crucially by citizens' subjective reasoning concerning agreements' clauses. In this paper, we leverage computational argumentation to introduce a novel approach to identifying mutually acceptable agreements among individuals in conflict, i.e. a Zone of Possible Agreement (ZOPA). First, we introduce a quantitative bipolar argumentation framework tailored to represent each side's reasoning about peace agreements. We then show how merging these frameworks can enable negotiators to identify peace agreements that are mutually acceptable. To evaluate our approach under conditions of real-world relevance, we focus on the Palestinian-Israeli conflict, where long-standing policy, practitioner and public interest underscores the demand for methods capable of analysing polarised public reasoning. We show how our framework identifies a ZOPA through theoretical analysis and preliminary experiments using survey data from both existing work and retrieved by a large language model. The results illustrate how argumentation can empower negotiators and conflict-resolution teams in mapping feasible ZOPAs grounded in citizens' reasoning.
Chinese Translation
在冲突社会中,如何识别共同立场,尤其是当公民对和平协议的接受度受到争议叙事的影响时?这种接受度不仅受到协议条款的包含或排除的影响,更重要的是受到公民对协议条款的主观推理的影响。本文利用计算论证引入了一种新颖的方法,以识别冲突个体之间的相互可接受协议,即可能的协议区域(Zone of Possible Agreement, ZOPA)。首先,我们引入一个量化的双极论证框架,以代表每一方对和平协议的推理。然后,我们展示了如何合并这些框架,使谈判者能够识别相互可接受的和平协议。为了在现实相关条件下评估我们的方法,我们重点关注巴以冲突,在该冲突中,长期以来的政策、从业者和公众利益突显了对能够分析极化公共推理的方法的需求。我们展示了我们的框架如何通过理论分析和使用来自现有研究和大型语言模型检索的调查数据的初步实验识别出一个ZOPA。结果表明,论证如何能够赋能谈判者和冲突解决团队,以基于公民推理绘制出可行的ZOPA。
cs.AI / 126 / 2608.15657

A Responsible Artificial Intelligence Framework for Groundwater Modeling

负责任的人工智能框架用于地下水建模
Chen, Chong, Zhang, Yulu, Guo, Qingxi, Liu, Yihan
Abstract
The rapid development and widespread application of artificial intelligence (AI) have sparked intense discussions on how to deploy responsible AI systems in a manner aligned with human values and ethical standards. Compared to fields like healthcare, energy, or finance, the application of AI in groundwater is relatively limited, and research on responsible AI is even more scarce. Taking the middle reaches of the Heihe River Basin as the study area, this paper proposes six Responsible AI principles: transparency, technical robustness, privacy governance, fairness, accountability, and sustainability. LSTM and Transformer time-series models are developed using multi-source hydrometeorological data, and validated via post-hoc interpretability, Monte Carlo simulation, and scenario analysis. The results show that Transformer outperforms LSTM in accuracy, robustness, and interpretability, demonstrating the operability and practical value of Responsible AI principles in groundwater prediction to support sustainable water management under climate change and human activities.
Chinese Translation
人工智能(AI)的快速发展和广泛应用引发了关于如何以符合人类价值观和伦理标准的方式部署负责任的AI系统的激烈讨论。与医疗、能源或金融等领域相比,AI在地下水领域的应用相对有限,而关于负责任AI的研究更是稀缺。本文以黑河流域中游地区为研究区域,提出了六项负责任AI原则:透明性、技术稳健性、隐私治理、公平性、问责制和可持续性。利用多源水文气象数据开发了LSTM和Transformer时间序列模型,并通过后验可解释性、蒙特卡洛模拟和情景分析进行验证。结果表明,Transformer在准确性、稳健性和可解释性方面优于LSTM,展示了负责任AI原则在地下水预测中的可操作性和实际价值,以支持在气候变化和人类活动下的可持续水资源管理。
cs.AI / 127 / 2608.15687

THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts

THESIS-MoE:可训练的层次提取与混合专家中的阿谀奉承引导
Hassani, Kareem, Abbas, Chaymaa, Mawlawi, Lama, Awad, Mariette
Abstract
Sycophancy, the tendency of a language model to change its answer to match a user's stated belief, is a common alignment failure. Existing activation steering methods typically apply a single contrastive direction uniformly throughout the model, which is an unconditional intervention that alters activations even when no sycophantic behavior is present, trading knowledge retention for behavioral correction. In Mixture-of-Experts (MoE) models, prior work further suggests that behavior is encoded within expert computations rather than routing decisions alone, making precise behavioral steering particularly challenging. In this work, we introduce a shared contrastive signal, built from matched prompts with and without a stated belief, that identifies where sycophancy lives across the MoE hierarchy and drives interventions that act only where the behavior is present. We formulate localization as a causal search over a granularity ladder of MoE blocks, experts, attention blocks, and heads, and compare unconditional subtraction against two conditional alternatives: an analytic projection-based subtraction and a learned per-token gate that steers the model away from sycophancy while keeping its weights frozen. We evaluate on three MoE models measuring sycophancy alongside general knowledge and reasoning benchmarks. Our conditional interventions removed up to 90\% of the belief-induced sycophancy. Our results demonstrate that sycophancy resides in identifiable computational subcircuits and can be selectively steered while maintaining a favorable removal-retention trade-off.
Chinese Translation
阿谀奉承是指语言模型倾向于改变其回答以符合用户所表达的信念,这是一种常见的对齐失败。现有的激活引导方法通常在模型中均匀应用单一的对比方向,这是一种无条件的干预,即使在没有阿谀奉承行为的情况下也会改变激活,牺牲知识保留以进行行为修正。在混合专家(Mixture-of-Experts, MoE)模型中,先前的研究进一步表明,行为不仅仅编码在路由决策中,还编码在专家计算中,这使得精确的行为引导尤其具有挑战性。在本研究中,我们引入了一种共享的对比信号,该信号由有和没有所述信念的匹配提示构建,能够识别阿谀奉承在MoE层次结构中的位置,并驱动仅在存在该行为的地方进行干预。我们将定位问题形式化为在MoE块、专家、注意力块和头部的粒度梯度上进行因果搜索,并将无条件减法与两种条件替代方案进行比较:一种基于解析投影的减法和一种学习的逐标记门控,后者在保持权重不变的情况下引导模型远离阿谀奉承。我们在三个MoE模型上进行评估,测量阿谀奉承以及一般知识和推理基准。我们的条件干预消除了高达90\%的信念引发的阿谀奉承。我们的结果表明,阿谀奉承存在于可识别的计算子电路中,并且可以在保持有利的去除与保留权衡的同时进行选择性引导。
cs.AI / 128 / 2608.15693

Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment

小设备的大模型:边缘人工智能部署的最新进展与实证分析
Das, Subhransu, Cheng, Jiaming, Kumar, Arnav, Afrose, Sadia, Han, Mingzhe, Silagy, Michael, Palande, Shreya, Soni, Brijesh, Ramnath, Rajiv
Abstract
Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation. What compresses well, however, need not deploy well. We survey dozens of recent works that report compression results on real hardware and extract practical deployment guidelines from them. Following these guidelines, we deploy compact language and image models on GPU, CPU, and Raspberry Pi platforms across question answering and image segmentation. No single technique wins across tasks. For question answering, Qwen3.5 0.8B reaches 93.85 SQuAD F1 and 92 EM under Q5_K_M GGUF quantization, while structured pruning at the same precision costs 16 F1 at a 1% ratio. For segmentation, the ranking reverses: default quantization leaves parameters and MACs unchanged, whereas pruning cuts model size by nearly 80% at near-constant mIoU. Pruning can even inflate the deployed artifact by 21-49% by breaking k-quant super-block alignment; combined with longer, less format-compliant outputs, this raises Raspberry Pi latency up to 3.4x. Compression can also manufacture the appearance of competence rather than destroy it visibly: one LoRA-recovered variant stays fully parseable and holds 71% strict BoolQ accuracy while sending 97 of 100 predictions to a single class, at 52.6% balanced accuracy. We explain these effects through neural-flow graph analysis and prefill-decode-level latency decomposition, and condense them into task-specific deployment research directions. The right technique depends on the task, the model, and the hardware. Our experiment code and artifacts are open-sourced at https://github.com/Arnavvvkumar/deployment
Chinese Translation
在资源受限的边缘设备上运行大型人工智能模型需要进行模型压缩,以减少模型的大小和计算量。然而,压缩效果良好的模型并不一定适合部署。我们调查了数十项近期研究,报告了在真实硬件上进行的压缩结果,并从中提取了实用的部署指南。根据这些指南,我们在GPU、CPU和树莓派平台上部署了紧凑的语言和图像模型,应用于问答和图像分割任务。没有单一技术在所有任务中表现最佳。在问答任务中,Qwen3.5 0.8B在Q5_K_M GGUF量化下达到了93.85的SQuAD F1和92的EM,而在相同精度下的结构化剪枝则以1%的比例损失了16的F1。在分割任务中,排名则相反:默认量化保持了参数和MACs不变,而剪枝在接近恒定的mIoU下将模型大小减少了近80%。剪枝甚至可能通过打破k-量化超块对齐,导致部署的工件膨胀21-49%;结合输出格式较长且不符合规范的情况,这使得树莓派的延迟增加了最多3.4倍。压缩还可以制造出能力的表象,而不是明显地破坏它:一个经过LoRA恢复的变体在保持71%的严格BoolQ准确率的同时,将100个预测中的97个发送到同一类别,平衡准确率为52.6%。我们通过神经流图分析和预填充解码级延迟分解来解释这些效应,并将其浓缩为特定任务的部署研究方向。合适的技术取决于任务、模型和硬件。我们的实验代码和工件已在https://github.com/Arnavvvkumar/deployment开源。
cs.AI / 129 / 2608.15700

Adaptive Mixing of Policies from Searching and Policies from Learning

自适应混合搜索策略与学习策略
Rens, Gavin B.
Abstract
Background: Distillation of training targets generated thru search/planning has proven useful in reinforcement learning, but search can take exceedingly long. Objectives: Rather than perform search to the same depth every time (typically at a fixed period of steps), reduce the search depth proportionally to the quality of the policy network priors. Methods: We describe Flexer, an architecture that, for each step, mixes the policy from a neural network and the policy from Monte Carlo tree search. The mixing factor favors the MCTS policy as the policy imitation error of the network and the environment models' variance increases. Results: Flexer outperforms a version of AlphaZero (and DQN and ADP) for some experiments on three toy symbolic problems.
Chinese Translation
背景:通过搜索/规划生成的训练目标的蒸馏在强化学习中已被证明是有用的,但搜索可能需要极长的时间。目标:我们旨在不每次都进行相同深度的搜索(通常是在固定的步数周期内),而是根据策略网络先验的质量成比例地减少搜索深度。方法:我们描述了Flexer,这是一种架构,在每一步中混合来自神经网络的策略和来自蒙特卡洛树搜索(Monte Carlo Tree Search, MCTS)的策略。混合因子在网络的策略模仿误差和环境模型的方差增加时更倾向于MCTS策略。结果:在三个玩具符号问题的某些实验中,Flexer的表现优于AlphaZero的一个版本(以及DQN和ADP)。
cs.AI / 130 / 2608.15703

HyMem: Hierarchical Context Management for Long-Horizon Agents via Information Isolation

HyMem:通过信息隔离实现长时间跨度智能体的层次上下文管理
Wang, XinQi, Xiao, Jinwei, Cui, Sijia, Zhang, Hongming, Wang, Yanna, Zhang, Qingyang, Xu, Bo
Abstract
Large language model (LLM) agents often perform poorly on complex, long-horizon tasks because their context becomes increasingly cluttered over time. As interactions accumulate, detailed execution traces and intermediate outputs dominate the context, making it difficult for the model to retain and use high-level planning information. Most existing methods address this issue through compression or retrieval applied to a single, flat context, which does not clearly separate different types of context information and often leads to degraded reasoning. To address this challenge, we propose HyMem, a hierarchical framework that explicitly separates the agent's context into distinct functional layers. HyMem organizes context by function to separate high-level planning from execution and complex analysis. Its isolated reasoning module handles complex subtasks without adding intermediate reasoning traces to the persistent planning context, while its memory management module preserves task progress across context refreshes through structured summaries. These components reduce redundant context accumulation, retain task-critical information, and support coherent long-horizon reasoning within a limited context window. Experiments on GAIA and Browsecomp-plus show that, with DeepSeek-V4, HyMem achieves average Pass@1 scores of 66.7% and 61.3%, outperforming the strongest baseline by 6.1 and 4.7 percentage points, respectively. Further analysis indicates that HyMem effectively controls the growth of the reasoning context, allowing the model to maintain focus and accuracy across complex, long-horizon tasks.
Chinese Translation
大型语言模型(LLM)智能体在复杂的长时间跨度任务中表现往往不佳,因为其上下文随着时间的推移变得越来越杂乱。随着交互的积累,详细的执行轨迹和中间输出主导了上下文,使得模型难以保留和使用高层次的规划信息。现有的大多数方法通过对单一平面上下文进行压缩或检索来解决这一问题,这种方法并未明确区分不同类型的上下文信息,往往导致推理能力下降。为了解决这一挑战,我们提出了HyMem,一个层次化框架,明确将智能体的上下文分为不同的功能层。HyMem按功能组织上下文,以将高层次规划与执行和复杂分析分开。其隔离推理模块处理复杂的子任务,而不将中间推理轨迹添加到持久的规划上下文中,同时其记忆管理模块通过结构化摘要在上下文刷新过程中保留任务进展。这些组件减少了冗余的上下文积累,保留了任务关键的信息,并支持在有限的上下文窗口内进行连贯的长时间跨度推理。在GAIA和Browsecomp-plus上的实验表明,使用DeepSeek-V4的HyMem在Pass@1评分上平均达到66.7%和61.3%,分别比最强基线高出6.1和4.7个百分点。进一步分析表明,HyMem有效控制了推理上下文的增长,使模型能够在复杂的长时间跨度任务中保持专注和准确性。
cs.AI / 131 / 2608.15719

PLeDO: Pain Level Detection for Osteoarthritis from EMR Data

PLeDO:基于电子病历数据的骨关节炎疼痛水平检测
Chen, Yuhao, Cai, Jiahao, Sadman, Nafiz, Zulkernine, Farhana, Queenan, John, Barber, David
Abstract
Osteoarthritis (OA) is a progressive chronic joint disease resulting in a breakdown of articular cartilage and bone when damaged joint tissues are not able to normally repair themselves. The aim of this pilot research study is to understand the pain severity for OA from patients' primary care Electronic Medical Records (EMR), both from the structured medical data and the unstructured chart note data using information extraction, natural language processing and machine learning techniques. We propose SPaDe, a Synonym-based Pain level Detection tool to categorize patients into having mild or moderate-to-severe pain to understand diagnosis and treatment methods based on only the pain related expressions in the unstructured chart note. Expressions are subjective, objective, and influenced by cultural background and demography which poses a difficult challenge. Therefore, we improve the model by incorporating the medication information from the structured EMR data and pain scale related information from the chart note to propose an integrated pain level detection tool for OA called PLeDO. With the help of human labeled gold standard data, we demonstrate that both SPaDe and PLeDO can detect mild and moderate-to-severe pain from the EMR data to analyze and potentially improve the quality of care in primary care setting.
Chinese Translation
骨关节炎(OA)是一种渐进性慢性关节疾病,导致关节软骨和骨骼的破坏,因受损的关节组织无法正常修复。本文的初步研究旨在通过患者的初级保健电子病历(EMR)来理解OA的疼痛严重程度,利用信息提取、自然语言处理和机器学习技术分析结构化医疗数据和非结构化病历记录数据。我们提出了SPaDe,一种基于同义词的疼痛水平检测工具,用于将患者分类为轻度或中度至重度疼痛,以便根据非结构化病历中的疼痛相关表达来理解诊断和治疗方法。表达是主观的、客观的,并受到文化背景和人口统计的影响,这带来了巨大的挑战。因此,我们通过结合结构化EMR数据中的用药信息和病历记录中的疼痛等级相关信息,改进了模型,提出了一种名为PLeDO的综合疼痛水平检测工具。借助人工标注的金标准数据,我们证明了SPaDe和PLeDO均能够从EMR数据中检测轻度和中度至重度疼痛,从而分析并潜在地改善初级保健环境中的护理质量。
cs.AI / 132 / 2608.15736

Toward AI-Friendly Cartography: Understanding How Color Design Influences Foundation Model Spatial Reasoning on Sequential Choropleth Maps

迈向适合人工智能的制图学:理解颜色设计如何影响基础模型在序列分级地图上的空间推理
Sun, Yonghe, Liu, Zhenjia, Liao, Hua, Xu, Wenjia, Yang, Nai, Dong, Weihua, Wei, Zhiwei
Abstract
Foundation models (FMs) increasingly support multimodal and geospatial reasoning, yet it remains unclear whether cartographic principles designed for human perception are equally effective for machines. Focusing on sequential choropleth maps, we examine how hue palette, color ordering, and lightness contrast influence FM spatial reasoning. We construct a controlled benchmark of 5,760 maps and 28,800 questions spanning Attribute Identify, Spatial Recognition, Compare, Rank, and Pattern Delineate, and evaluate 21 open-source and proprietary multimodal FMs. Results show that hue choice has limited and inconsistent effects, whereas disrupting sequential color ordering substantially reduces performance, especially for comparison and ranking. Reduced lightness contrast also consistently impairs reasoning, while increasing contrast beyond sufficient separability provides only marginal gains. LoRA fine-tuning improves overall accuracy but preserves these relative sensitivities. Additional factorial experiments further indicate that errors arise from color-and-legend decoding, spatial reasoning, and the integration of thematic attributes with spatial structure. These findings show that conventional sequential ordering and sufficient contrast remain important for machine map understanding and provide empirical guidance for AI-friendly cartographic design.
Chinese Translation
基础模型(FMs)越来越多地支持多模态和地理空间推理,但尚不清楚为人类感知设计的制图原则是否同样适用于机器。我们专注于序列分级地图,研究色调调色板、颜色排序和明度对基础模型空间推理的影响。我们构建了一个包含5760张地图和28800个问题的受控基准,涵盖属性识别、空间识别、比较、排名和模式描绘,并评估了21个开源和专有的多模态基础模型。结果表明,色调选择的影响有限且不一致,而打乱序列颜色排序显著降低了性能,尤其是在比较和排名任务中。降低明度对比度也始终会损害推理能力,而在足够可分离性之上增加对比度仅带来边际收益。LoRA微调提高了整体准确性,但保留了这些相对敏感性。额外的因子实验进一步表明,错误源于颜色和图例解码、空间推理以及主题属性与空间结构的整合。这些发现表明,传统的序列排序和足够的对比度对于机器地图理解仍然重要,并为适合人工智能的制图设计提供了实证指导。
cs.AI / 133 / 2608.15746

Propaganda Forensics: Recovering the Generation Pipeline of an AI-Driven Influence Campaign

宣传法医学:恢复由人工智能驱动的影响活动的生成流程
Icard, Benjamin, Vuichard, Elouan, Lefebvre, Louis, Sainero, Lila, Girault, Thomas, Breton, Alice, Launay, Tanguy, Bourgne, Gauvain, Casanova, Morgane, Gadek, Guillaume, Klötzer, Victor, Nouy, Michel Le, Gravier, Guillaume, Ganascia, Jean-Gabriel, Égré, Paul
Abstract
We present a forensic analysis of the generation pipeline behind a recent AI-driven influence campaign. We introduce PROPAGIA, a corpus of 2,646 propagandist French articles from the Storm-1516/CopyCop campaign disclosed by VIGINUM and INSIKT GROUP in 2025. For comparison, we rely on SIPA, a corpus of human-written French mainstream press from the same period. Using topic modeling, vagueness and sentiment analysis, we first isolate persuasion techniques characteristic of propaganda, with PROPAGIA far exceeding SIPA in vagueness, subjectivity and negativity, and citing fewer sources. We then find prompt instruction leaks on 50 of the 84 PROPAGIA websites, including a verbatim ten-point editorial specification accounting for several of these differences, together with high cross-article redundancy. Finally, we show that rewriting-based detection supports INSIKT GROUP's attribution to the Llama 3 family, but also suggests the involvement of Mistral-family models.
Chinese Translation
我们对最近一项由人工智能驱动的影响活动的生成流程进行了法医学分析。我们引入了PROPAGIA,这是一个包含2,646篇来自Storm-1516/CopyCop活动的宣传性法语文章的语料库,该活动由VIGINUM和INSIKT GROUP于2025年披露。为了进行比较,我们依赖于SIPA,这是一个来自同一时期的人类撰写的法语主流媒体语料库。通过主题建模、模糊性和情感分析,我们首先隔离出宣传特有的说服技巧,发现PROPAGIA在模糊性、主观性和消极性方面远超SIPA,并引用的来源更少。随后,我们发现84个PROPAGIA网站中的50个存在提示指令泄露,包括一份逐字逐句的十点编辑规范,解释了这些差异中的几个,同时也显示出高跨文章冗余。最后,我们展示了基于重写的检测支持INSIKT GROUP对Llama 3系列的归属,但也暗示了Mistral系列模型的参与。
cs.AI / 134 / 2608.15755

Intent-Driven Situation Tracking for User-Centric Multi-Turn Agents

基于意图驱动的用户中心多轮代理情境跟踪
Tao, Meiling, Tao, Yiling, Wang, Peng
Abstract
User-centric multi-turn agents must act on an evolving task situation shaped by changing user intents, accumulated tool-grounded facts, missing information, and execution constraints. Existing context-management methods improve the use of past interaction history, but rarely maintain an explicit situation state that separates grounded facts from task-state judgments. As a result, agents often need to infer fine-grained attributes, task dependencies, and constraint satisfaction implicitly from dialogue traces. We propose Intent-Driven Situation States (IDSS), a training-free framework that maintains an explicit situation state alongside the dialogue. IDSS parses tool returns into provenance-aware entities and attributes, tracks user intents, required variables, constraints, and execution status, and propagates new facts to task constraints to update action executability. This allows agents to avoid infeasible actions, advance dependent goals, and reuse relevant information without repeatedly searching raw history. Experiments on three interactive benchmarks across eight LLMs show that IDSS improves task completion, preference elicitation, and interaction efficiency, with clear gains on tasks involving multi-entity coordination, evolving user constraints, and constraint-aware replanning. Ablations and error analyses show that these improvements come from the interaction between fact persistence, intent-centered state tracking, and constraint modeling. These results suggest that explicit situation tracking offers an effective alternative to history-centric context management for reliable user-centric multi-turn agents.
Chinese Translation
以用户为中心的多轮代理必须根据不断变化的用户意图、积累的工具基础事实、缺失的信息和执行约束,针对不断演变的任务情境采取行动。现有的上下文管理方法虽然改善了对过去交互历史的利用,但很少维护一个明确的情境状态,以区分基础事实与任务状态判断。因此,代理往往需要从对话痕迹中隐式推断细粒度属性、任务依赖关系和约束满足情况。我们提出了意图驱动情境状态(Intent-Driven Situation States, IDSS),这是一个无训练的框架,能够在对话过程中维护一个明确的情境状态。IDSS将工具返回解析为具有来源意识的实体和属性,跟踪用户意图、所需变量、约束和执行状态,并将新事实传播到任务约束中,以更新行动的可执行性。这使得代理能够避免不可行的行动,推进依赖目标,并在不重复搜索原始历史的情况下重用相关信息。在对八个大型语言模型(LLMs)进行的三个交互基准测试中,实验表明IDSS在任务完成、偏好引导和交互效率方面均有所改善,尤其在涉及多实体协调、不断演变的用户约束和约束感知的重新规划任务中表现出明显的提升。消融实验和错误分析表明,这些改进源于事实持久性、以意图为中心的状态跟踪和约束建模之间的相互作用。这些结果表明,明确的情境跟踪为可靠的用户中心多轮代理提供了一种有效的替代方案,超越了以历史为中心的上下文管理。
cs.AI / 135 / 2608.15772

Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration

大语言模型拒绝中的破坏对称性:答案释放比拒绝恢复更局部
Liu, Yiqi, Wang, Yang, Wang, Songxin, Xiao, Chenghao, Lin, Chenghua
Abstract
When a language model refuses to answer a prompt, it is unclear whether the correct answer is erased from its internal representations, or merely suppressed at the output layer. We investigate this mechanism using a controlled withhold setting, which yields perfectly matched answering and refusal trajectories for bidirectional activation patching. We uncover a causal asymmetry in intervention locality under matched causal interventions, which we term broken symmetry. Even when a model generates a clean refusal, the correct answer remains linearly recoverable from its hidden states. Furthermore, releasing this withheld answer is a highly local operation, requiring only a single-position patch. Conversely, the reverse operation is not equally local: reimposing suppression requires broader interventions across multiple positions, and assembling a coherent refusal sequence is more difficult still. We further demonstrate that while an average answer-to-refusal displacement vector marks the geometric difference between these states, it fails to act as a reliable, reversible linear control toggle between behaviours. Taken together, our findings show that refusal does not function as a simple symmetric switch. For safety and auditing, this implies that probe recoverability can overestimate true behavioural control, and locating refusal-relevant directions does not reliably grant the ability to steer a model from answering to coherent refusal.
Chinese Translation
当语言模型拒绝回答提示时,尚不清楚正确答案是从其内部表征中抹去,还是仅在输出层被抑制。我们通过控制性隐匿设置来研究这一机制,该设置为双向激活修补提供了完美匹配的回答和拒绝轨迹。我们发现,在匹配的因果干预下,干预局部性存在因果不对称性,我们称之为破坏对称性。即使模型生成了干净的拒绝,正确答案仍然可以从其隐状态中线性恢复。此外,释放这一被隐匿的答案是一个高度局部的操作,仅需一个单位置的修补。相反,反向操作并不那么局部:重新施加抑制需要在多个位置进行更广泛的干预,而组装一个连贯的拒绝序列则更加困难。我们进一步证明,尽管平均的答案到拒绝位移向量标记了这些状态之间的几何差异,但它未能作为一个可靠的、可逆的线性控制开关在行为之间切换。综合来看,我们的发现表明,拒绝并不作为一个简单的对称开关运作。对于安全性和审计而言,这意味着探测可恢复性可能高估了真实的行为控制,而定位与拒绝相关的方向并不能可靠地赋予从回答到连贯拒绝的能力。
cs.AI / 136 / 2608.15797

KV-Rescue: Recovering Reasoning Language Model KV Eviction Loss via Stepwise Interleaving

KV-Rescue:通过逐步交错恢复推理语言模型的KV驱逐损失
Cheong, Minsoo, Lim, Woosang, Yun, Vincent-Daniel, Yoo, Sungjoo
Abstract
KV-cache eviction caps the memory cost of long reasoning traces but is inherently lossy because the model decodes from a partial view of its history. Under aggressive budgets, this not only lowers accuracy but can also cause runaway degeneration, where the model produces incoherent or repetitive tokens until reaching the length limit. We characterize much of this loss as an information gapf caused by missing context, rather than a capability gap caused by limited model capacity. An evicted 7B model and a full-context 1.5B model make complementary errors, and an oracle choice between their answers recovers 79% of the accuracy gap to the full-KV 7B model. Based on this observation, we propose KV-Rescue, a training-free inference framework that bridges the information gap introduced by KV eviction using a lightweight full-context helper. KV-Rescue interleaves reasoning steps from the two models into a shared trajectory. An online detector uses entropy and compressibility to terminate the generation of incoherent or repetitive base-model candidates early. Across five math benchmarks with Qwen2.5-Math 7B and 72B, KV-Rescue recovers an average of 87% of the accuracy lost to eviction at eviction budget B=64. A decode-cost analysis further shows that preventing runaway degeneration cuts base-model token generation by 43% on average.
Chinese Translation
KV缓存驱逐限制了长推理轨迹的内存成本,但由于模型从部分历史视图进行解码,这一过程本质上是有损的。在激进的预算下,这不仅降低了准确性,还可能导致失控退化,即模型生成不连贯或重复的标记,直到达到长度限制。我们将这种损失的很大一部分归因于信息缺口(information gap),这是由于缺失上下文造成的,而不是由于模型能力有限造成的能力缺口(capability gap)。一个被驱逐的7B模型和一个全上下文的1.5B模型会产生互补的错误,通过在它们的答案之间进行选择,可以恢复79%的与全KV 7B模型的准确性差距。基于这一观察,我们提出了KV-Rescue,一个无训练推理框架,利用轻量级全上下文助手弥补KV驱逐引入的信息缺口。KV-Rescue将两个模型的推理步骤交错成一个共享轨迹。一个在线检测器使用熵和可压缩性来提前终止不连贯或重复的基础模型候选的生成。在五个数学基准测试中,使用Qwen2.5-Math 7B和72B,KV-Rescue恢复了在驱逐预算B=64下失去的平均87%的准确性。解码成本分析进一步表明,防止失控退化平均减少了基础模型标记生成43%。
cs.AI / 137 / 2608.15810

Pricing the Risk of Runtime Compression: Anytime-Valid Admission and a Served-Output Law for Compressed Serving State

运行时压缩风险定价:随时有效的接纳与压缩服务状态的输出法则
Wei, Fanzhe, Liu, Li
Abstract
Runtime compression of serving state trades quality for capacity with no priced guarantee: systems adapt precision on load signals with no soundness statement, and certified approaches budget request-level risk by a union bound over a pre-declared event count. We show the union budget exhausts on every long request in a production serving stack (100% of requests), and replace it with an anytime-valid, physically accounted ledger whose bound holds at every one of 352,333 admission calls on live traffic and which, in a pre-registered held-out confirmatory round, halves the exact-fallback rate at matched risk (0.30 -> 0.14) -- coverage is bought at a price the account states. We then price the remaining distance from the certified witness to what a user experiences: a machine-checked design law (TV <= tanh(a_q w_thr)) turns the served-TV target into a threshold knob, and a three-layer audit of its instantiation -- an operator-norm query envelope measured 1.5x from tight, a measured-ellipsoid replacement for the Cauchy-Schwarz ball that buys nothing (0.89x, held-out sound), and the gate's operating point (~700x) -- localizes the entire 1064x gap to the operating point, a price the law now states rather than an unknown. A priced bound is worth nothing on a request one has not seen, so the third link is the quantifier: exchangeable extrapolation across 80 serving histories replaces binary conformal prediction's vacuous certificates with order-statistic bounds that discriminate (0.41 against 0.51 calibration risk). All probabilistic kernels are Lean 4-checked (228 exported theorems, no sorry); which object deserves this machinery at all is settled empirically in a companion paper that adjudicates -- and rejects -- the natural alternative of certifying routing. What ships is an account: risk you can spend, a gap you can read off a law, and a bound that survives the request you have not seen.
Chinese Translation
服务状态的运行时压缩在没有价格保证的情况下以质量换取容量:系统根据负载信号调整精度,但没有可靠性声明,而认证方法通过对预先声明的事件计数进行联合界限预算请求级风险。我们展示了在生产服务堆栈中,每个长请求的联合预算耗尽(100%的请求),并用一个随时有效、物理核算的账本替代,该账本在352,333个实时流量的接纳调用中保持有效,并且在一个预注册的保留确认轮次中,将在匹配风险下的精确回退率减半(0.30 -> 0.14)——覆盖以账本所声明的价格购买。然后,我们对从认证见证到用户体验的剩余距离进行定价:一个机器检查的设计法则(TV <= tanh(a_q w_thr))将服务的TV目标转变为一个阈值旋钮,并对其实例化进行三层审计——一个操作范数查询包从紧密测量得出1.5倍,一个测量椭球体替代Cauchy-Schwarz球体的测量为0.89倍(保留有效),以及门的操作点(约700倍)——将整个1064倍的差距归因于操作点,这是法律现在所声明的价格,而不是未知的。对未见请求的定价界限毫无价值,因此第三个环节是量化器:在80个服务历史之间可交换的外推取代了二元符合预测的空洞证书,提供了区分的顺序统计界限(0.41对比0.51的校准风险)。所有概率核均经过Lean 4检查(228个导出定理,无遗憾);究竟哪个对象值得使用这一机制在一篇伴随论文中通过实证解决,该论文裁定并拒绝了认证路由的自然替代方案。最终呈现的是一个账本:可以支出的风险,一个可以从法律中读取的差距,以及一个在你未见过的请求中仍然有效的界限。
cs.AI / 138 / 2608.15817

RLCascadeRouter: Quality-Estimator-Free Cascade Routing via Reinforcement Learning

RLCascadeRouter:通过强化学习实现无质量估计的级联路由
Huang, Shihong, Wang, Shengjie, Ma, Hong, Xu, Zhou
Abstract
The growing ecosystem of large language models (LLMs) offers huge potential to optimize performance-cost trade-offs. However, their heterogeneous capabilities and inference costs make efficiently routing queries a significant challenge. Existing paradigms are inflexible: one-shot routers commit before observing responses, whereas conventional cascades stop adaptively but follow a fixed model order. Cascade routing removes both restrictions by reconsidering whether to stop or invoke another model after each response. Current methods use a predict-then-optimize pipeline estimating response quality and future model utility. However, prediction loss for quality or utility is not equivalent to routing-decision loss. A lower prediction error does not necessarily yield a better action; a small boundary-crossing error can reverse a ``stop'' or model-selection decision. Therefore, we propose RLCascadeRouter, a quality-estimator-free framework that formulates cascade routing as a Markov decision process with actions comprising ``stop'' and model selection. It uses trajectory returns and advantages to directly optimize the performance-cost objective. Its Cascade Policy Network models candidate complementarity for model selection and remaining-action value for stopping, eliminating independent post-hoc response-quality estimators. Evaluated across ten LLMRouterBench benchmarks with thirteen LLMs, RLCascadeRouter outperforms strong baselines and achieves superior performance-cost trade-offs. It incorporates unseen models without retraining, and ablation studies validate both policy components.
Chinese Translation
大型语言模型(LLMs)日益增长的生态系统为优化性能与成本的权衡提供了巨大的潜力。然而,它们异构的能力和推理成本使得高效路由查询成为一项重大挑战。现有的范式缺乏灵活性:一次性路由器在观察到响应之前就做出了承诺,而传统的级联则在适应性停止时遵循固定的模型顺序。级联路由通过在每次响应后重新考虑是否停止或调用另一个模型,消除了这两种限制。目前的方法采用预测-再优化的流程,估计响应质量和未来模型的效用。然而,质量或效用的预测损失并不等同于路由决策损失。较低的预测误差不一定会产生更好的行动;小的边界穿越误差可能会逆转“停止”或模型选择的决策。因此,我们提出了RLCascadeRouter,一个无质量估计的框架,将级联路由公式化为一个马尔可夫决策过程,其动作包括“停止”和模型选择。它使用轨迹回报和优势直接优化性能-成本目标。其级联策略网络建模候选模型的互补性以进行模型选择,并建模剩余动作值以进行停止,从而消除了独立的事后响应质量估计器。在十个LLMRouterBench基准测试和十三个LLMs的评估中,RLCascadeRouter超越了强基线,并实现了更优的性能-成本权衡。它可以在不重新训练的情况下纳入未见过的模型,消融研究验证了这两个策略组件的有效性。
cs.AI / 139 / 2608.15832

The Authority Resolution Framework: A Five-Domain Ontology for Governing Who and What Decides, at Scale

权威解析框架:一个五领域本体用于大规模治理决策主体与决策内容
Shariff, Parviz
Abstract
As AI systems become increasingly capable of autonomous action, determining whether an agent is technically capable of performing an action is insufficient: the system must also determine whether the action is authorised in its context. This paper introduces the Authority Resolution Framework (ARF), a five-domain ontology for representing and resolving authority across organisational roles and informal influence, business concepts, codified processes, machine-readable permissions and executable systems, and external real-world context. ARF defines the Authority Relation (AR) as a cross-domain primitive binding an actor, action, object, bounded context, justification chain, and a calibration measure termed the DNA-Coefficient, which captures divergence between documented authority structures and authority as practiced. The framework provides a machine-interpretable representation of authority provenance and scope, with JSON-LD representations and knowledge-graph query patterns for authority resolution. ARF is designed to support AI agents in determining the provenance, scope and contextual validity of authority before executing consequential actions. The framework positions authority resolution as a knowledge-representation and reasoning problem at the intersection of ontology engineering, semantic AI, agentic AI and AI governance.
Chinese Translation
随着人工智能系统越来越具备自主行动的能力,仅仅判断一个代理是否具备执行某项行动的技术能力是不够的:系统还必须确定该行动在其上下文中是否被授权。本文介绍了权威解析框架(Authority Resolution Framework, ARF),这是一个用于表示和解决组织角色与非正式影响、商业概念、编码流程、机器可读权限及可执行系统以及外部现实世界上下文之间权威关系的五领域本体。ARF将权威关系(Authority Relation, AR)定义为一个跨领域的原语,绑定一个行动者、行动、对象、有限上下文、理由链以及一个称为DNA系数的校准度量,该系数捕捉了文档化权威结构与实际行使权威之间的偏差。该框架提供了权威来源和范围的机器可解释表示,包含JSON-LD表示和用于权威解析的知识图谱查询模式。ARF旨在支持人工智能代理在执行重要行动之前,确定权威的来源、范围和上下文有效性。该框架将权威解析定位为一个知识表示和推理问题,处于本体工程、语义人工智能、代理人工智能和人工智能治理的交叉点。
cs.AI / 140 / 2608.15834

Schema-Agnostic Graph Reasoning Agent for Hybrid Knowledge Graphs

面向混合知识图谱的无模式图推理代理
Dragic, Marius, Ifrah, Ruben, Rio, Alexandre
Abstract
Tool-calling LLM agents navigate unfamiliar codebases with a handful of generic primitives for listing, reading and searching files (ls, cat, grep). A knowledge graph admits the same interface: listing neighbours, reading node content and searching descriptions are the same operations on a different substrate. Building on this correspondence, we present GRA, a Graph Reasoning Agent that explores hybrid knowledge graphs, whose nodes are either textual concepts or relational tables, with seven generic tools, discovering everything domain-specific at run time. On UFK-M (Unified Factory Knowledge Model), an industrial benchmark of 258 analytical questions whose gold answers are produced by executing validated SQL programs, GRA beats a full-context agent by 5.1 pp (88.4% vs. 83.3%), while reading under a third of its input tokens. A graph-free control shows the gain comes chiefly from selective agentic access rather than graph topology, and that the effect depends on a model able to drive tools reliably. Seeing less, the agent answers better: selective navigation over a structured substrate beats exhaustive context.
Chinese Translation
工具调用的大型语言模型(LLM)代理使用一小部分通用原语(如列出、读取和搜索文件的命令:ls、cat、grep)来导航不熟悉的代码库。知识图谱也采用相同的接口:列出邻居、读取节点内容和搜索描述在不同的底层结构上是相同的操作。基于这种对应关系,我们提出了GRA(图推理代理),它探索混合知识图谱,这些图谱的节点可以是文本概念或关系表,使用七种通用工具,在运行时发现所有特定领域的信息。在UFK-M(统一工厂知识模型)上,这是一个包含258个分析问题的工业基准,其金标准答案是通过执行经过验证的SQL程序生成的,GRA的表现比全上下文代理高出5.1个百分点(88.4%对比83.3%),同时读取的输入标记不到其三分之一。无图控制显示,性能提升主要来自选择性代理访问,而非图拓扑,并且这一效果依赖于能够可靠驱动工具的模型。看到的信息更少,代理的回答更好:在结构化底层上进行选择性导航优于全面上下文。
cs.AI / 141 / 2608.15857

RAGas: Retrieval-Augmented Gas Optimization for Smart Contracts with Continuous Knowledge Integration

RAGas:用于智能合约的检索增强气体优化与持续知识集成
Wang, Yishun, Yi, Wenjin, Li, Wenkai, Li, Zongwei, Li, Xiaoqi
Abstract
Ethereum is now integral to mission-critical sectors, including finance, healthcare, and supply chain management. Execution fees, commonly referred to as Gas, scale with the computational complexity of their functions. Smart contracts on Ethereum incur execution fees, known as Gas, which increase with computational complexity. Thus, optimizing Gas-intensive code while preserving functional equivalence significantly lowers deployment costs. No existing system continuously exploits evolving Gas usage patterns. We systematically analyze syntactic and semantic constructs that drive excessive Gas use. This yields six high-level categories covering twelve fine-grained antipatterns underpinning a curated knowledge base. We operationalize these insights with RAGas, a three-stage retrieval-augmented generation framework that uses a large language model to pinpoint and automatically fix Gas inefficiencies. Experiments on deployed contracts demonstrate that RAGas reduces Gas usage by up to 11% and achieves high precision and recall in detecting code snippets exhibiting Gas wastage.
Chinese Translation
以太坊如今已成为金融、医疗和供应链管理等关键任务领域的重要组成部分。执行费用,通常称为Gas,随着其功能的计算复杂性而增加。在以太坊上,智能合约的执行费用被称为Gas,随着计算复杂性的增加而上升。因此,在保持功能等价的情况下优化Gas密集型代码显著降低了部署成本。目前没有现有系统能够持续利用不断变化的Gas使用模式。我们系统地分析了导致过度Gas使用的语法和语义构造。这产生了六个高层次类别,涵盖了支持策划知识库的十二个细粒度反模式。我们通过RAGas将这些见解转化为实践,RAGas是一个三阶段的检索增强生成框架,利用大型语言模型来识别和自动修复Gas低效问题。在已部署合约上的实验表明,RAGas将Gas使用减少了多达11%,并在检测显示Gas浪费的代码片段方面达到了高精度和高召回率。
cs.AI / 142 / 2608.15868

CoupVisor: Strategy Optimization by Round and Challenge Decision Support

CoupVisor:基于回合和挑战决策支持的策略优化
Huynh, Cris
Abstract
This paper presents CoupVisor, a decision-support system for the hidden-information card game Coup. It addresses two questions: what a player should do on each turn, and when a player should challenge an opponent's claim. The system is built around a single description of game events, which is shared across manual play, replay of recorded games, simulation, belief tracking, advisor recommendations, and learning-based policies. CoupVisor estimates the chance that a claim is truthful by combining how likely each role is with how many cards the claimant still holds, which corrects a case where the very first claim of a game was flagged as suspicious despite no evidence. We compare a rule-following advisor and several learned and heuristic players across many simulated games and different opponent styles. Our main finding is that the choice of reward, whether it rewards short-term gains or ultimately winning the game, decides which learning approach performs best, and that a win-oriented reward produces a policy that outperforms all baselines.
Chinese Translation
本文介绍了CoupVisor,一个用于隐藏信息卡牌游戏Coup的决策支持系统。它解决了两个问题:玩家在每个回合应该做什么,以及玩家何时应该挑战对手的声明。该系统围绕游戏事件的单一描述构建,该描述在手动游戏、录制游戏的回放、模拟、信念追踪、顾问推荐和基于学习的策略中共享。CoupVisor通过结合每个角色的可能性与声明者仍持有的牌数,估计声明真实的概率,从而纠正了游戏开始时的第一个声明在没有证据的情况下被标记为可疑的情况。我们比较了一个遵循规则的顾问和几种学习型及启发式玩家在多场模拟游戏和不同对手风格下的表现。我们的主要发现是,奖励的选择,无论是奖励短期收益还是最终赢得游戏,决定了哪种学习方法表现最佳,而以胜利为导向的奖励产生的策略超越了所有基准。
cs.AI / 143 / 2608.15877

Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation

亲爱的算法:一种以精准为先的代理意图层,用于统一搜索与推荐
Wang, Rui, Wang, Jiazhou, Wei, Zheng, Lu, Chenglin, Sun, Fangcheng, Sun, Ivy, Sun, Jin, Geng, Hui, Zhang, Lillian, Yang, Chao, Chen, Lei, Sefati, Shahin, Helou, Reem, Zhou, Joe, Shakibi, Babak, Pan, Yiyi, Xue, Bi, Yan, Hong, Bu, Shujian
Abstract
Search and recommendation serve a shared discovery objective but encode intent differently. We study this boundary through Dear Algo on Threads, a deployed product where open-ended requests such as \emph{more NBA news} or \emph{less politics} steer subsequent feed recommendations rather than return a one-shot result list. Its agentic intent layer compiles explicit, inferred, negative, and compound intent into a grounded executable plan, then invokes conventional retrieval and optional semantic or multimodal reranking. The layer shares an intent-to-retrieval contract without requiring one model or serving path across search-like and recommendation-like modes. We evaluate Dear Algo under a precision-first objective. In a blinded audit of 300 public request-item pairs (296 evaluable), a strict categorical LLM-as-a-judge gate achieved 94.4\% exact-Relevant precision [88.8\%, 98.9\%]. Across 72 normalized request clusters, the full configuration produced 7.73 judge-qualified candidates per 20 slots versus 6.61 for an LLM-derived-query baseline, a gain of 1.11 [0.12, 2.12]. In a candidate-randomized serving-path study restricted to the reranker path's first 72 eligible hours, the user-weighted judge-Irrelevant share among judged admissions was 2.80\% versus 4.78\% off (-1.97 points [-3.02, -0.94]), while Exact-Relevant share was 2.24 points higher [0.08, 4.41]. Together, these studies show how explicit natural-language intent can be carried into feed recommendation under a precision-first evaluation framework
Chinese Translation
搜索与推荐服务于共同的发现目标,但编码意图的方式却不同。我们通过在 Threads 上的 Dear Algo 研究这一边界,这是一个已部署的产品,其中开放式请求如“更多 NBA 新闻”或“更少政治”引导后续的推荐,而不是返回一次性结果列表。其代理意图层将显性、推断、负面和复合意图编译成一个有依据的可执行计划,然后调用传统检索和可选的语义或多模态重新排序。该层在搜索类和推荐类模式之间共享意图到检索的契约,而不需要一个模型或服务路径。我们在以精准为先的目标下评估 Dear Algo。在对 300 对公共请求-项目对(296 个可评估)的盲审中,一个严格的分类 LLM 作为评审的门槛达到了 94.4% 的精确相关性 [88.8%, 98.9%]。在 72 个标准化请求集群中,完整配置每 20 个插槽产生 7.73 个评审合格候选者,而 LLM 派生查询基线为 6.61,增益为 1.11 [0.12, 2.12]。在一个限制于重新排序路径的候选随机服务路径研究中,用户加权的评审无关分享在评审录取中为 2.80% ,而对照组为 4.78%(-1.97 分 [-3.02, -0.94]),而精确相关性分享高出 2.24 分 [0.08, 4.41]。这些研究共同表明,显性自然语言意图如何在以精准为先的评估框架下被纳入推荐中。
cs.AI / 144 / 2608.15888

Bounded Agents: Delegation Security for Multi-Agent AI Systems

有界代理:多智能体人工智能系统的委托安全性
Muruaga, Xabier
Abstract
LLM-based agents can act on behalf of a user to access cloud services, call tools, or invoke agents. At session start, the agent's permissions are set but remain static, and each request is evaluated independently, without considering prior actions. Within its permissions, an agent may act contrary to the delegated task, combine individually permitted actions into a prohibited outcome, or delegate authority to a sub-agent without limiting it. A prompt injection poses a risk only if the agent has authority to perform such actions; this is therefore a problem of authorization architecture, not just the model. The Agentic Principal Chain (APC) tracks delegated authority from one principal to the next. APC evaluates each request against the accumulated session state using six authorization checks. APC carries forward and restricts delegated scope and budgets. Using composition closure, APC checks requests against prior actions to prevent prohibited combinations and enforces the decision outside the model. We prove Blast Radius Monotonicity and Composition Soundness for APC implementations; Composition Soundness is limited to prohibited combinations under a complete restriction set and serialized admission. We evaluated 3,154 instances including InjecAgent, AgentDojo, and ASB. Our compromised-model evaluation tests APC independently of model behavior by inserting the ground-truth attack call after the first legitimate tool call. AgentDojo exfiltration fell from 75-100% to 0% across all four domains; APC blocked all 544 InjecAgent data-stealing cases. Intent binding reduced destruction from 38.6% to 4.0% and manipulation from 90.5% to 12.1%. Authorization latency was 0.24 ms at the 99th percentile on an idle host; across 949 AgentDojo task-injection pairs, utility was 8.6 and 13.9 percentage points lower in the two settings. Implementation, evaluation tools, and data are publicly available.
Chinese Translation
基于大型语言模型(LLM)的代理可以代表用户访问云服务、调用工具或调用其他代理。在会话开始时,代理的权限被设定为静态,并且每个请求都是独立评估的,而不考虑之前的操作。在其权限范围内,代理可能会违反委托任务,将单独允许的操作组合成禁止的结果,或者在不加限制的情况下将权力委托给子代理。仅当代理有权执行此类操作时,提示注入才构成风险;因此,这不仅是模型的问题,更是授权架构的问题。代理主链(Agentic Principal Chain, APC)跟踪从一个主体到下一个主体的委托权。APC使用六个授权检查根据累积的会话状态评估每个请求。APC传递并限制委托的范围和预算。通过组合闭合,APC检查请求与之前操作的关系,以防止禁止的组合,并在模型之外执行决策。我们证明了APC实现的爆炸半径单调性和组合健全性;组合健全性仅限于在完全限制集和串行接纳下的禁止组合。我们评估了包括InjecAgent、AgentDojo和ASB在内的3,154个实例。我们的受损模型评估通过在第一次合法工具调用后插入真实攻击调用,独立于模型行为测试APC。AgentDojo的外泄率在所有四个领域从75-100%降至0%;APC阻止了所有544个InjecAgent数据窃取案例。意图绑定将破坏率从38.6%降低到4.0%,操控率从90.5%降低到12.1%。在空闲主机上,授权延迟在第99个百分位为0.24毫秒;在949个AgentDojo任务注入对中,两个设置的效用分别低了8.6和13.9个百分点。实现、评估工具和数据均可公开获取。
cs.AI / 145 / 2608.15893

Breaking and Defending LLM-Powered Social Media Bot Detection Systems

破解与防御基于大型语言模型的社交媒体机器人检测系统
Orenstein, Nof, Birman, Yoni
Abstract
The rise of social media bots poses a persistent threat, enabling misinformation, opinion manipulation, and the erosion of trust in online platforms. To combat this, machine learning systems have been developed to detect and limit bot activity, but attackers continuously adapt through techniques such as adversarial learning and behavior imitation, fueling an ongoing arms race between bots and detection tools. Recent advances in large language models (LLMs) have significantly improved bot detection by enabling deeper semantic and contextual analysis of accounts and their content. However, this shift also introduces new attack surfaces, allowing adversaries to craft exploits that directly target the reasoning and generation mechanisms of LLM-based classifiers. Industry tools such as Anthropic's Claude Code Security similarly leverage LLMs for security-critical decisions, further motivating a careful study of their attack surfaces. In this work, we investigate both the offensive and defensive aspects of LLM-powered, threat-specific cybersecurity applications. While centered on the challenge of social media bot detection, our methodology and insights generalize to a broad class of LLM-powered cybersecurity systems, including phishing detection, email classification, and fraud analysis. We introduce two novel adversarial attack strategies that systematically exploit the semantic and contextual weaknesses of LLM-based classifiers, degrading their detection accuracy by up to 48%. To counter these threats, we propose a robust multi-LLM defense architecture designed to preserve detection reliability under adaptive adversarial conditions. Our solution, LSABRE (LLM-powered Social Adversarial Bot Recognition Ensemble), is a multi-LLM framework that substantially improves robustness across a range of attacks, maintaining 86% detection accuracy even under strong, adaptive adversarial pressure.
Chinese Translation
社交媒体机器人的崛起构成了持续的威胁,助长了虚假信息、舆论操控以及对在线平台信任的侵蚀。为应对这一挑战,已经开发出机器学习系统来检测和限制机器人的活动,但攻击者通过对抗学习和行为模仿等技术不断适应,导致机器人与检测工具之间的军备竞赛持续进行。大型语言模型(LLMs)的最新进展显著提升了机器人的检测能力,使得对账户及其内容进行更深入的语义和上下文分析成为可能。然而,这一转变也引入了新的攻击面,使对手能够设计直接针对基于LLM的分类器的推理和生成机制的攻击。行业工具如Anthropic的Claude Code Security同样利用LLM进行安全关键决策,进一步激励对其攻击面的仔细研究。在本研究中,我们探讨了基于LLM的特定威胁网络安全应用的攻防两方面。尽管重点在于社交媒体机器人检测的挑战,我们的方法论和见解可以推广到广泛的基于LLM的网络安全系统,包括钓鱼检测、电子邮件分类和欺诈分析。我们提出了两种新颖的对抗攻击策略,系统性地利用基于LLM的分类器的语义和上下文弱点,使其检测准确率降低了多达48%。为了应对这些威胁,我们提出了一种强健的多LLM防御架构,旨在在自适应对抗条件下保持检测的可靠性。我们的解决方案LSABRE(基于LLM的社交对抗机器人识别集成)是一个多LLM框架,显著提高了在多种攻击下的鲁棒性,即使在强大的自适应对抗压力下仍能保持86%的检测准确率。
cs.AI / 146 / 2608.15929

Unified Pedestrian Path Prediction Using Inverse Reinforcement Learning

基于逆强化学习的统一行人路径预测
Sukup, Šimon, Bighashdel, Ariyan, Jancura, Pavol
Abstract
Pedestrian path prediction is crucial for enhancing the safety of autonomous vehicles and advanced driver-assistance systems. Previous studies explored different learning-task formulations for pedestrian path prediction and compared these formulations using shallow neural networks, but did not extend this analysis to more complex deep-learning models. This paper adapts the Spatial-Temporal Graph Attention Network (STGAT) to a unified pedestrian path prediction framework and introduces state and action definitions specific to STGAT. The resulting formulations support deterministic and stochastic policies, one-time and sequential decision-making, and reinforcement-learning algorithms including REINFORCE and proximal policy optimization. The proposed learning-task formulations improve prediction performance across the selected benchmark datasets compared with the standard supervised-learning formulation. These results demonstrate that reformulating the decision process and training objective can improve an advanced pedestrian trajectory prediction architecture and may provide a path toward improving other graph-based prediction models.
Chinese Translation
行人路径预测对于提高自动驾驶车辆和高级驾驶辅助系统的安全性至关重要。以往的研究探讨了行人路径预测的不同学习任务形式,并使用浅层神经网络对这些形式进行了比较,但并未将此分析扩展到更复杂的深度学习模型。本文将空间-时间图注意力网络(Spatial-Temporal Graph Attention Network, STGAT)适配为统一的行人路径预测框架,并引入了特定于STGAT的状态和动作定义。所得到的形式支持确定性和随机策略、一次性和序列决策,以及包括REINFORCE和近端策略优化(proximal policy optimization)在内的强化学习算法。与标准的监督学习形式相比,所提出的学习任务形式在所选基准数据集上的预测性能得到了提升。这些结果表明,重新构建决策过程和训练目标可以改善先进的行人轨迹预测架构,并可能为改进其他基于图的预测模型提供路径。
cs.AI / 147 / 2608.15930

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations

UI-Mate:通过上下文演示推进开放权重基础图形用户界面代理
Ding, Zihan, Dou, Longxu, Gao, Qi, Guo, Xiangwu, Hu, Shengchao, Huang, Zilong, Jiang, Zihang, Ke, Lei, Lan, Mengcheng, Lei, Weixian, Li, Hanxuan, Li, Honglin, Li, Xiyun, Li, Zaitang, Liang, Leowei, Luo, Xin, Ma, Haozhe, Mao, Jiayi, Pan, Zhoujie, Qin, Can, Qu, Tianyuan, Wang, Weiqi, Wang, Wenkai, Wang, Yonglin, Wang, Yuxin, Wu, Chenxu, Yu, Yingchen, Zhang, Chenyu, Zheng, Yuhao
Abstract
Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts, and unreliable execution. Routine workflows rely on user-specific tools and tacit conventions, so unstated instructions can produce arbitrary variations across runs. We present UI-Mate, a foundation GUI agent that integrates an environment-grounded training stack with in-context demonstration learning. UI-Mate makes three contributions: A Scalable Environment-Grounded Training Stack: A closed-loop data engine automates task generation, environment construction, rollout, filtering, capability balancing, SFT, and online RL across massively parallel environments via unified task-verifier bundles. In-Context Demonstration Learning: A mechanism that transforms multimodal demonstrations into flexible subtask-level workflows, follows relevant demonstrated steps, and re-plans from the live interface. OSWorkerBench Benchmark and Insights: A benchmark of 100 long-horizon office tasks across 41 applications that supports instruction-only and demonstration-guided evaluation. Its demonstration resources separate a 33-task self-demo setting, built from successful strong-agent rollouts of the same targets, from a 45-task variant-demo setting, built from human recordings of related but non-identical tasks. Experiments show that UI-Mate-27B sets a new open-weight state of the art on general computer-use benchmarks, scoring 77.0% on OSWorld-Verified and 66.2% on WindowsAgentArena. On OSWorkerBench, it reaches 41.0% strict success and 76.9% progress, outperforming its Qwen3.6-27B base by 17.7 and 24.5 points. On the 33-task self-demo subset, one demonstration raises strict success from 17.2% to 35.4% and progress from 67.9% to 81.1%, substantially improving long-horizon reliability. Project page: https://ui-mate.github.io.
Chinese Translation
基础图形用户界面代理可以自动化复杂的数字任务,但由于训练数据稀缺和偏见、提示模糊以及执行不可靠,部署受到阻碍。常规工作流程依赖于用户特定的工具和隐性约定,因此未明确说明的指令可能在不同运行中产生任意变化。我们提出了UI-Mate,一个将环境基础训练堆栈与上下文演示学习相结合的基础图形用户界面代理。UI-Mate有三个贡献:可扩展的环境基础训练堆栈:一个闭环数据引擎自动化任务生成、环境构建、发布、过滤、能力平衡、SFT(监督微调)和在线RL(强化学习),通过统一的任务验证器包在大规模并行环境中进行。上下文演示学习:一种机制,将多模态演示转化为灵活的子任务级工作流程,遵循相关的演示步骤,并从实时界面重新规划。OSWorkerBench基准和见解:一个涵盖41个应用程序的100个长时间办公任务的基准,支持仅基于指令和演示引导的评估。其演示资源将33任务的自我演示设置与45任务的变体演示设置分开,前者基于相同目标的成功强代理发布,后者基于相关但不完全相同任务的人类录音。实验表明,UI-Mate-27B在一般计算机使用基准上设定了新的开放权重最先进水平,在OSWorld-Verified上得分77.0%,在WindowsAgentArena上得分66.2%。在OSWorkerBench上,它达到了41.0%的严格成功率和76.9%的进展,分别比其Qwen3.6-27B基础提高了17.7和24.5个百分点。在33任务的自我演示子集上,一次演示将严格成功率从17.2%提高到35.4%,进展从67.9%提高到81.1%,显著提高了长时间的可靠性。项目页面:https://ui-mate.github.io
cs.AI / 148 / 2608.15932

Augmenting Text to Increase Translation Difficulty

增强文本以增加翻译难度
Kalikman, William, Sukup, Šimon, Tešnar, Michal, Zouhar, Vilém
Abstract
As state-of-the-art machine translation models saturate standard benchmarks, the field needs more challenging evaluations to distinguish between models of varying quality. We propose augmenting existing benchmarks to increase translation difficulty by combining adversarial optimization with a differentiable translation difficulty estimator. Our Adversarial Translation Optimization (ATO) uses gradients from a combined difficulty and fluency objective to iteratively replace tokens. Because each step branches over candidate substitutions at every position, optimization becomes a tree search problem, which we address with Beam Search. ATO offers a gradient-based alternative to LLM-based dataset creation without LLM prompting, expensive human curation, or task-specific model training. Our ATO-modified benchmark lowers average translation quality (xCOMET) from 0.93 to 0.82, compared to 0.88 for paraphrasing and 0.86 for a zero-shot baseline. Human evaluation shows the modified texts are somewhat less natural than the baselines but remain reasonably grammatical and plausible while being substantially harder to translate. We release two datasets of 350 English texts each, generated by our methods, as well as the code.
Chinese Translation
随着最先进的机器翻译模型在标准基准测试中达到饱和,研究领域需要更具挑战性的评估来区分不同质量的模型。我们提出通过将对抗优化与可微分的翻译难度估计器相结合,来增强现有基准测试以增加翻译难度。我们的对抗翻译优化(Adversarial Translation Optimization, ATO)利用来自结合难度和流畅性目标的梯度,迭代地替换词元。由于每一步在每个位置上都对候选替换进行分支,优化变成了一个树搜索问题,我们通过束搜索(Beam Search)来解决。ATO提供了一种基于梯度的替代方案,用于创建不依赖于大型语言模型(LLM)提示、昂贵的人为策划或特定任务模型训练的数据集。与改写(paraphrasing)得到的0.88和零样本基线(zero-shot baseline)得到的0.86相比,我们的ATO修改基准将平均翻译质量(xCOMET)降低至0.82。人工评估显示,修改后的文本在自然性上略逊于基线,但仍然相当语法正确且合理,同时翻译难度显著增加。我们发布了两个各包含350个英文文本的数据集,这些文本是通过我们的方法生成的,以及相关代码。
cs.AI / 149 / 2608.15956

Navigation-Informed Embeddings: Dense-Retriever Adaptation from Agent Search Traces

导航信息嵌入:从代理搜索轨迹适应密集检索器
Shah, Shrey, Ozgur, Levent
Abstract
Agentic retrieval workflows produce query, retrieval, and stopping traces as a byproduct of answering questions. We study how these traces can adapt a deployed dense retriever to changing workflow distributions without new relevance labels, synthetic queries, or LLM judgments. We introduce Navigation-Informed Embeddings (NIE), a family of trace-derived objectives. NIE-Stop turns the stopping document into a soft positive; NIE-Path additionally uses preceding path documents as hard comparisons and imposes ordinal constraints with geometric decay. A BGE encoder adapted from retained source trajectories improves support Recall@20 on an independent target benchmark from 72.2 to 78.0 overall. NIE-Stop reaches 76.9 overall and 52.3 on long paths; NIE-Path raises long-path performance to 55.4, compared with 46.7 for the unadapted encoder. A shuffled-order control under the full path objective loses 3.2 points. Without public-benchmark training, the same adapter also improves nDCG@10 by 1.9 points on standard BEIR HotpotQA. NIE therefore provides a lightweight adaptation channel for settings where trajectories are already retained, with zero incremental labeling cost.
Chinese Translation
代理检索工作流程在回答问题的过程中产生查询、检索和停止轨迹作为副产品。我们研究了这些轨迹如何在不需要新的相关性标签、合成查询或大型语言模型(LLM)判断的情况下,适应已部署的密集检索器以应对变化的工作流程分布。我们引入了导航信息嵌入(Navigation-Informed Embeddings, NIE),这是一系列基于轨迹衍生的目标。NIE-Stop将停止文档转化为软正样本;NIE-Path则额外使用前置路径文档作为硬比较,并施加几何衰减的序数约束。改编自保留源轨迹的BGE编码器在一个独立的目标基准上将支持的Recall@20从72.2提高到78.0。NIE-Stop的总体表现为76.9,长路径的表现为52.3;NIE-Path将长路径性能提升至55.4,而未适应的编码器为46.7。在完整路径目标下的随机顺序控制损失了3.2分。在没有公共基准训练的情况下,相同的适配器在标准BEIR HotpotQA上也将nDCG@10提高了1.9分。因此,NIE为已经保留轨迹的设置提供了一种轻量级的适应渠道,且没有额外的标注成本。
cs.AI / 150 / 2608.15958

Solvable Sokoban Without a Solver via Diffusion

无求解器的可解Sokoban通过扩散
Baghal, Sina
Abstract
Deciding whether a Sokoban puzzle is solvable is PSPACE-complete (Culberson, 1997): solutions can be exponentially long and there is no short certificate to check. Solvability is also a fragile property, since even a single misplaced wall can silently render an entire puzzle unsolvable. In this work, we show that a transformer-based discrete diffusion model trained purely on tile completion, with no access to solvers, rewards, or solvability labels, achieves a solvability rate of 77.4%, with 94.5% of the remaining failures rendered solvable by removing a single wall. In other words, a global, search-heavy property follows from a local training objective: trained only to fill in masked cells, the model inherits solvability it was never trained on. An autoregressive model factorizes as $p(c_k \mid c_1 \dots c_{k-1})$, meaning a fixed order, always conditioned on a prefix. Masked diffusion does not: it hides a random subset of cells and learns $p(c_k \mid \text{any subset})$, so at generation time it can reveal cells in any order, each one conditioned on everything already placed, wherever it sits on the board. A puzzle's difficulty comes from exactly this kind of non-local interaction, a decision in one part of the grid constraining what will work somewhere else entirely. A generator that is not locked into a single fixed order is therefore a better structural match for the problem than one that is. The training pipeline is adapted from MD4 (Shi et al., 2024) and the dataset is DeepMind's Boxoban (Guez et al., 2019). The trained model and instructions for generating puzzles are publicly available.
Chinese Translation
判断一个Sokoban谜题是否可解是PSPACE完全的(Culberson, 1997):解决方案可能是指数级长,并且没有短的证书可以检查。可解性也是一个脆弱的属性,因为即使是一个错误放置的墙壁也可能悄然使整个谜题变得不可解。在这项工作中,我们展示了一种基于变压器的离散扩散模型,该模型仅在瓷砖完成上进行训练,且没有访问求解器、奖励或可解性标签,达到了77.4%的可解率,94.5%的其余失败通过移除一堵墙变得可解。换句话说,一个全局的、依赖搜索的属性源于一个局部的训练目标:仅训练填补被遮蔽的单元格,该模型继承了它从未接受过训练的可解性。自回归模型的分解形式为$p(c_k ext{mid} c_1 ext{dots} c_{k-1})$,意味着固定的顺序,始终以前缀为条件。遮蔽扩散则不同:它隐藏随机子集的单元格,并学习$p(c_k ext{mid} ext{any subset})$,因此在生成时可以以任意顺序揭示单元格,每个单元格都以已经放置的所有内容为条件,无论它在棋盘上的位置如何。谜题的难度恰恰来自于这种非局部交互,网格某一部分的决策限制了其他地方的可行性。因此,一个不被锁定在单一固定顺序的生成器比一个被锁定的生成器更适合该问题。训练流程改编自MD4(Shi et al., 2024),数据集为DeepMind的Boxoban(Guez et al., 2019)。训练好的模型和生成谜题的说明已公开发布。
cs.AI / 151 / 2608.15979

ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction

ALPS:通过数学构造测量大型语言模型中的有效创造力
Xie, Eric, Ye, Wenqian, Zhang, Aidong
Abstract
Large language models produce outputs presented as discoveries - new proofs, conjectures, or molecules. Whether such an output that appears creative is truly original and effective is hard to establish: open-ended outputs require subjective judgment, the output may replicate something seen in training, or the task may be too simple to need creativity. We present ALPS (Austin-Law Proof-Synthesis), a benchmark that designs a task to measure valid creativity: producing a solution that is original and can be proven correct. Each instance is a single equational law, certified to require either the construction of an infinite mathematical structure satisfying the law, or a proof that no such structure exists. Submissions are verified by automated proof checking with no human involvement, and a public generator produces new instances without limit, so LLMs are never evaluated on problems they may have seen. A portfolio of eight configurations of leading automated provers resolves 2.2% of the 4,141-law evaluation pool, and a twentyfold budget increase adds 0.6%: the obstacle is not compute, but the absence of any method that produces the tailored structure each law requires. Under a fixed protocol, the strongest reasoning model we test succeeds in 14% of instances on the proof side, but none on the construction side. The remaining 97.2% of the pool is unresolved at every configuration and budget we test. We release ALPS in full: the corpus, the generator, and the automated judge.
Chinese Translation
大型语言模型生成的输出常常被呈现为发现——新的证明、猜想或分子。然而,判断这种看似创造性的输出是否真正具有原创性和有效性是困难的:开放式输出需要主观判断,输出可能复制了训练中见过的内容,或者任务可能过于简单而不需要创造力。我们提出了ALPS(Austin-Law Proof-Synthesis),一个基准,设计了一项任务以测量有效创造力:产生一个原创且可以被证明正确的解决方案。每个实例都是一个单一的方程法则,经过认证需要构造一个满足该法则的无限数学结构,或者证明不存在这样的结构。提交的结果通过自动证明检查进行验证,无需人类参与,并且一个公共生成器可以无限制地生成新的实例,因此大型语言模型不会在它们可能见过的问题上进行评估。领先的自动证明者的八种配置组合解决了4,141法则评估池中的2.2%,而预算增加二十倍仅增加了0.6%:障碍不在于计算能力,而在于缺乏能够产生每个法则所需的定制结构的方法。在固定协议下,我们测试的最强推理模型在证明方面成功率为14%,但在构造方面没有成功。其余97.2%的池在我们测试的每个配置和预算下均未解决。我们全面发布ALPS:语料库、生成器和自动评判者。
cs.AI / 152 / 2608.15999

MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessment

MUPA$^{2}$E:具有非对称注意力的多模态统一感知用于情感评估
Gkikas, Stefanos, Nichols, Eric, Cruz, Christian Arzate, Gomez, Randy
Abstract
Automatic emotion assessment can benefit from combining neural and behavioral signals, but many multimodal approaches rely on separate, modality-specific feature-extraction pipelines before fusion. This paper presents MUPA\textsuperscript{2}E, a unified perception framework that processes facial video and electroencephalography (EEG) through a single shared asymmetric-attention backbone. Facial video is represented through axis-folded frame tokens, while EEG is processed either as a raw multichannel waveform or projected into the spatial domain for multimodal fusion. The framework is evaluated on the DMER dataset under a stratified subject-independent protocol, comparing unimodal video, unimodal EEG, and fused video--EEG configurations with per-channel and merged EEG projections. Using the original recordings, with shorter trials zero-padded to match the longest duration, merged fusion at stride~$30$ achieves the highest validation performance and a test accuracy of $70.07\%$. Further analysis revealed that recording duration is unevenly distributed across the affective classes, making the padding pattern a potential classification cue. Controlling for this factor by cropping all recordings to a common duration of $20$ seconds yielded a test accuracy of $62.71\%$, providing a stricter duration-controlled assessment of the framework in which differences in recording length are removed as a potential classification cue. These findings demonstrate the feasibility of processing structurally different neural and visual signals within a compact unified architecture while highlighting the importance of controlling duration-related cues in affective datasets.
Chinese Translation
自动情感评估可以通过结合神经信号和行为信号受益,但许多多模态方法在融合之前依赖于独立的、特定模态的特征提取管道。本文提出了MUPA extsuperscript{2}E,一个统一感知框架,通过一个共享的非对称注意力主干处理面部视频和脑电图(EEG)。面部视频通过轴折叠帧标记表示,而EEG则可以作为原始多通道波形处理,或投影到空间域以进行多模态融合。该框架在DMER数据集上进行了评估,采用分层的受试者独立协议,比较了单模态视频、单模态EEG以及融合视频-EEG配置(包括每通道和合并EEG投影)。使用原始记录,较短的试验通过零填充以匹配最长持续时间,合并融合在步幅~$30$下达到了最高的验证性能和$70.07\%$的测试准确率。进一步分析表明,记录持续时间在情感类别之间分布不均,使得填充模式成为潜在的分类线索。通过将所有记录裁剪到统一的20秒持续时间来控制这一因素,得到了$62.71\\%$的测试准确率,从而提供了对框架的更严格的持续时间控制评估,消除了记录长度差异作为潜在分类线索。这些发现展示了在紧凑的统一架构内处理结构上不同的神经和视觉信号的可行性,同时强调了在情感数据集中控制与持续时间相关的线索的重要性。
cs.AI / 153 / 2608.16003

Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency

先前审计-修复背景使大型语言模型验证器阈值趋向宽松
Mazaheri, Parsa, Mazaheri, Kasra
Abstract
Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human-verified-correct ProcessBench traces with the present task held byte-identical, we find that a completed audit -> repair episode already in the model's context lowers false alarms in 15 of 15 model x wording combinations, by 2.8 to 11.5 percentage points against a length-matched non-audit control, a 9 to 25% reduction relative to that control. The direction contradicts what the accumulated-message literature predicts: an episode whose audit reported an error lowers false alarms further still, at all five wordings on the model where that manipulation lands cleanly, though a negativity asymmetry predicts more flagging. Decomposing the episode finds repair content and audit verdict complementary: different components carry the effect on different model families. Signal-detection analysis locates the change in the threshold rather than in discrimination -- the criterion moves in 15 of 15 combinations and survives correction in 13 while d' survives in none, though the d' test is half as sensitive by construction -- and a hand audit of 50 false alarms finds 82% simply wrong, so at this operating point the shift need not be harmful. With reasoning enabled the effect keeps its relative size on both models tested, and the threshold reading holds there too.
Chinese Translation
自动化检查流程越来越多地将一个语言模型作为检查者,另一个(或相同的)作为修复者。我们探讨这种连接是否会改变检查者的报告。通过在保持当前任务字节完全相同的情况下,测量人类验证正确的 ProcessBench 跟踪中的误报,我们发现,已完成的审计->修复事件在模型的上下文中会降低 15 种模型与措辞组合中的误报率,降低幅度为 2.8 到 11.5 个百分点,相较于长度匹配的非审计控制组,减少幅度为 9% 到 25%。这一方向与累积信息文献的预测相悖:一个审计报告错误的事件进一步降低了误报率,在该操作干预顺利的模型的五种措辞中均如此,尽管负面不对称性预测会有更多的标记。对事件的分解发现,修复内容和审计裁决是互补的:不同组件对不同模型家族产生影响。信号检测分析表明,变化发生在阈值而非判别上——在 15 种组合中标准均有所移动,并且在 13 种组合中经过校正后仍然有效,而 d' 在任何情况下都未能存活,尽管 d' 测试在构造上敏感度仅为一半——对 50 个误报的人工审计发现 82% 是完全错误的,因此在这一操作点上,变化不一定是有害的。在启用推理的情况下,效果在测试的两个模型上保持相对大小,阈值读取也保持不变。
cs.AI / 154 / 2608.16055

Governance at the Boundary: How Agent Decomposition Degrades Policy Compliance

边界治理:代理分解如何降低政策合规性
Li, Bowen, Wang, Guojun
Abstract
Existing agent benchmarks ask whether the agent finished the task. We ask whether it finished it within policy. We introduce Fiducia-bench, a benchmark for the governability of financial agents---whether they escalate when obligated, abstain when required, and leave an auditable trail---and use it to study a question no prior benchmark addresses: does decomposing an agent into components degrade its governance? It does, and the mechanism is specific. Policy-relevant facts discovered by one component are attenuated at the handoff boundary before reaching the component that must act on them. In a 626-episode experiment across 100 KYC/AML task variants, two models, and three architectures, a 32B open-weights model attenuated 0% of discovered facts under a single-loop baseline, 56% under a fixed pipeline, and 85% under an orchestrator-subagent architecture (all at constraint distance 2). A stronger model (gpt-4.1-mini) attenuated 3-6% under the same conditions, suggesting the governance cost of decomposition is partly a function of model capability. Critically, the same mechanism produces both under-escalation and over-escalation, depending on whether the dropped fact was a risk signal or an exculpating one. The benchmark, all tasks, and the verification harness are open-source
Chinese Translation
现有的代理基准测试关注代理是否完成了任务。我们关注的是代理是否在政策范围内完成了任务。我们引入了Fiducia-bench,这是一个用于金融代理治理能力的基准——即代理在被要求时是否会升级、在需要时是否会克制,以及是否留下可审计的痕迹——并利用它研究一个之前的基准未曾涉及的问题:将代理分解为组件是否会降低其治理能力?结果表明确实如此,且机制是特定的。一个组件发现的与政策相关的事实在交接边界处被削弱,未能传递给必须对此采取行动的组件。在一个涵盖100个KYC/AML任务变体的626回合实验中,使用了两种模型和三种架构,一个32B的开放权重模型在单循环基线下未削弱0%的发现事实,在固定管道下削弱了56%,而在协调者-子代理架构下削弱了85%(所有条件下约束距离为2)。一个更强的模型(gpt-4.1-mini)在相同条件下削弱了3-6%,这表明分解的治理成本部分取决于模型的能力。重要的是,同一机制在丢失的事实是风险信号或免责信号时,会产生不足升级和过度升级的现象。该基准、所有任务和验证工具均为开源。
cs.AI / 155 / 2608.16084

Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynamics

多尺度混沌动态的自回归神经模拟器的特征分析框架
Ainslie, Conrad, Hassanzadeh, Pedram, Mahoney, Michael W., Chattopadhyay, Ashesh
Abstract
Neural autoregressive models have rapidly emerged as powerful emulators of high-dimensional chaotic systems, yet their long-term instability and error growth remain poorly understood, leading to ad-hoc solutions. Here, we develop an eigenanalysis framework that reveals the dynamical origin of this error growth. By analyzing the Jacobian of the learned one-step update map with respect to the state, we show how inference-time error growth, and thus model stability, is governed by its spectral radius. Direct-step architectures (models that predict the next state from the previous one) generically admit unstable eigenvalues with magnitudes exceeding one, explaining the rapid divergence of these widely used models. In contrast, integration-constrained models (where the time derivative is estimated and integrated with a higher-order integrator) collapse their eigenspectrum onto the unit circle, yielding neutral stability and a universal linear error-scaling law. The largest eigenvalue of this Jacobian provides an architecture-agnostic, a priori diagnostic of short-term skill, long-term stability, and spectral bias, without requiring an expensive rollout. Leveraging this theory, we introduce a stability-promoting loss that explicitly regularizes Jacobian-driven error amplification, improving both forecast accuracy and dynamical robustness. Demonstrated across $29$ models spanning two architectures, several explicit and implicit integrators, and multiple loss functions on the Kuramoto-Sivashinsky system, our results establish a theoretical foundation for the design and evaluation of neural emulators of chaotic multi-scale dynamics. More broadly, our framework is a step toward the kind of a priori stability analysis that numerical analysis provides for discretizations of differential equations and that scientific machine learning currently lacks.
Chinese Translation
神经自回归模型迅速崛起,成为高维混沌系统的强大模拟器,但其长期不稳定性和误差增长仍然缺乏深入理解,导致了临时解决方案的出现。在此,我们开发了一个特征分析框架,揭示了这种误差增长的动态起源。通过分析学习到的一步更新映射相对于状态的雅可比矩阵,我们展示了推理时误差增长以及模型稳定性是如何由其谱半径所支配的。直接步进架构(从前一个状态预测下一个状态的模型)通常会出现幅度超过一的不稳定特征值,这解释了这些广泛使用模型的快速发散。相比之下,集成约束模型(在其中时间导数被估计并与高阶积分器进行积分)将其特征谱压缩到单位圆上,从而实现中性稳定性和普遍的线性误差缩放法则。该雅可比矩阵的最大特征值提供了一种与架构无关的先验诊断,能够评估短期技能、长期稳定性和谱偏差,而无需昂贵的展开。利用这一理论,我们引入了一种促进稳定性的损失函数,明确正则化雅可比驱动的误差放大,从而提高预测准确性和动态鲁棒性。在涵盖两种架构、多个显式和隐式积分器以及多种损失函数的 $29$ 个模型上进行的实验中,我们的结果为混沌多尺度动态的神经模拟器的设计和评估奠定了理论基础。更广泛地说,我们的框架是朝着提供类似于数值分析对微分方程离散化所提供的先验稳定性分析的方向迈出的一步,而这一点在当前的科学机器学习中尚缺乏。
cs.AI / 156 / 2608.16094

Protein Structure Prediction: From Evolutionary Constraints to Generative Modeling

蛋白质结构预测:从进化约束到生成建模
He, Wengan, Luo, Yongsheng, Jiang, Lihong, Xu, Wenhui, Li, Yu
Abstract
Accurate protein structure prediction is fundamental to structural biology because protein structure underlies molecular function and provides a basis for mechanistic interpretation. Recent advances in deep learning have transformed the field from multiple sequence alignment (MSA)-driven monomer folding into broader frameworks capable of modeling protein complexes and increasingly heterogeneous molecular systems. Existing reviews have summarized this progress from the perspectives of representative models, application domains, and protein design. Building on these efforts, this review focuses on the methodological evolution of the field itself. It examines recent developments through three closely related dimensions: representations and data, architectures and learning strategies, and confidence and evaluation. Within this perspective, the field is organized into four methodological phases and three cross-cutting transitions: from explicit evolutionary coupling features and early contact prediction to learned sequence representations in AlphaFold2, RoseTTAFold, and ESMFold; from protein-only monomer folding to increasingly integrated modeling of heterogeneous molecular systems in AlphaFold-Multimer, RoseTTAFoldNA, and AlphaFold3; and, more recently, from prediction-oriented structure inference to design-oriented generative modeling in RFdiffusion and related frameworks. This framework provides a clearer understanding of how methodological shifts have shaped the capabilities, limitations, and practical roles of recent models.
Chinese Translation
准确的蛋白质结构预测对结构生物学至关重要,因为蛋白质结构是分子功能的基础,并为机制解释提供了依据。近年来,深度学习的进展使得该领域从依赖多序列比对(MSA)的单体折叠转变为能够建模蛋白质复合物和日益异质的分子系统的更广泛框架。现有的综述从代表性模型、应用领域和蛋白质设计的角度总结了这一进展。在这些努力的基础上,本综述聚焦于该领域自身的方法论演变。通过三个密切相关的维度考察近期发展:表示法与数据、架构与学习策略,以及信心与评估。在这一视角下,该领域被组织为四个方法论阶段和三个交叉过渡:从显式的进化耦合特征和早期接触预测到在AlphaFold2、RoseTTAFold和ESMFold中的学习序列表示;从仅考虑蛋白质的单体折叠到在AlphaFold-Multimer、RoseTTAFoldNA和AlphaFold3中对异质分子系统的日益综合建模;以及最近,从以预测为导向的结构推断转向在RFdiffusion及相关框架中的设计导向生成建模。该框架提供了更清晰的理解,阐明了方法论的转变如何塑造了近期模型的能力、局限性和实际角色。
cs.AI / 157 / 2608.16118

Assessing LLMs' mathematical abilities requires understanding the various mechanisms of mathematical creativity

评估大型语言模型的数学能力需要理解数学创造力的各种机制
Gangloff, Silvère
Abstract
How should we assess whether large language models can perform mathematical invention? I argue that this question is currently underspecified: mathematical creativity is not one capacity but several mechanistically distinct modes of meaning-making - reflexive introspection on mathematical practice, analogical import from the sciences, problem-driven construction, and the bridging of distant domains - together with a further, cross-cutting distinction between meaning pursued because a pattern was observed and meaning pursued because it is strategically wanted, a distinction I develop through the case of conjecture-formation. These mechanisms are likely non-substitutable, so that competence in one does not transfer to the others. Grounding each in a historical case study and in an architecture-level account of current transformer-based systems, I suggest that today's models concentrate their competence in modes shaped by recombination and search over existing building blocks; if that description holds, the remaining modes are out of reach in principle, not just slower - though whether it holds is itself the open, empirical part. Because proof is getting cheaper as AI improves at generating it - a shift the field's own leading voices are now diagnosing - mathematical value is migrating toward the modes current systems cannot yet perform, and evaluations of AI mathematical ability should be organized around this taxonomy rather than around aggregate benchmarks that conflate it.
Chinese Translation
我们应如何评估大型语言模型是否能够进行数学发明?我认为这个问题目前尚未明确:数学创造力并不是一种能力,而是几种在机制上截然不同的意义构建模式——对数学实践的反思性内省、从科学中的类比引入、以问题为驱动的构建,以及跨越遥远领域的桥接——此外,还有一个进一步的交叉区分,即追求意义是因为观察到某种模式,还是因为战略性地想要某种意义,我通过猜想形成的案例来发展这一区分。这些机制可能是不可替代的,因此在一种机制上的能力并不转移到其他机制上。通过将每种机制植根于历史案例研究和当前基于变换器的系统的架构层面描述,我建议当今的模型将其能力集中在通过对现有构建块的重组和搜索所形成的模式上;如果这一描述成立,剩余的模式在原则上是无法触及的,而不仅仅是速度较慢——尽管这一描述是否成立本身就是一个开放的实证问题。随着人工智能在生成证明方面的能力不断提高,证明变得越来越便宜——这一变化现在被该领域的领先声音所诊断——数学价值正向当前系统尚未能够执行的模式迁移,评估人工智能的数学能力应围绕这一分类法进行,而不是围绕混淆这一分类法的综合基准进行。
cs.AI / 158 / 2608.16147

When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification

单一数据集结论失效时:关于不平衡分类的阈值调优和重采样的45任务研究
Musaev, Diyorbek
Abstract
Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting conclusions are reported as if they were properties of the method. We show this practice is unsafe. On the public Kaggle credit-card fraud dataset, under a leakage-free nested cross-validation protocol in which the decision threshold is selected on a held-out inner validation fold, a plain Random Forest at the default 0.5 threshold attains F1 = 0.861 +/- 0.021, and threshold tuning yields it no benefit (delta-F1 = -0.002). Read alone, this supports an appealing conclusion: for a well-calibrated ensemble, imbalance handling is unnecessary. We then apply the identical protocol to 45 binary tasks spanning imbalance ratios from 1:1.5 to 1:178 (2,025 model fits, four model families). The conclusion reverses. Random Forest benefits most from threshold tuning across the suite (delta-F1 = +0.101 +/- 0.134), not least, while three other families replicate their fraud-dataset behaviour almost exactly. SMOTE likewise harms the fraud dataset but helps across the suite (mean delta-F1 = +0.076; 138 wins, 39 losses; Wilcoxon p = 2.7e-17). Two further results. Threshold-tuning benefit is non-monotonic in the imbalance ratio: near zero below 1:5, peaking at +0.120 in the 1:15-1:40 band, declining to +0.045 beyond 1:100 - explaining why the fraud dataset, at 1:577, is an unrepresentative place to study the question. And we reject an intuitive heuristic: validation-set calibration error does not predict tuning benefit (expected calibration error r = -0.087; Brier r = +0.137), so calibration diagnostics cannot tell a practitioner whether tuning is worthwhile. We release the protocol, the 45-task harness, and all per-run metrics.
Chinese Translation
类不平衡处理通常在单一基准数据集上进行评估,而得出的结论则被报告为该方法的特性。我们展示了这一做法的不安全性。在公共的Kaggle信用卡欺诈数据集上,在一个无泄漏的嵌套交叉验证协议下,决策阈值是在保留的内部验证折叠上选择的,普通的随机森林在默认的0.5阈值下达到了F1 = 0.861 +/- 0.021,而阈值调优并未带来任何好处(delta-F1 = -0.002)。单独阅读这一结果支持了一个吸引人的结论:对于一个良好校准的集成模型,不平衡处理是没有必要的。随后,我们将相同的协议应用于45个二元任务,涵盖了从1:1.5到1:178的不平衡比(2,025个模型拟合,四个模型家族)。结论发生了逆转。在整个任务中,随机森林从阈值调优中获益最大(delta-F1 = +0.101 +/- 0.134),而其他三个家族几乎完全复制了它们在欺诈数据集上的行为。SMOTE同样对欺诈数据集有害,但在整个任务中有所帮助(平均delta-F1 = +0.076;138次胜利,39次失败;Wilcoxon p = 2.7e-17)。还有两个进一步的结果。阈值调优的好处在不平衡比中是非单调的:在1:5以下接近零,在1:15-1:40区间达到峰值+0.120,超过1:100后下降至+0.045——这解释了为什么在1:577的不平衡比下,欺诈数据集并不是研究该问题的代表性场所。我们还拒绝了一个直观的启发式:验证集的校准误差并不能预测调优的好处(预期校准误差r = -0.087;Brier r = +0.137),因此校准诊断无法告诉从业者调优是否值得。我们发布了该协议、45任务的框架以及所有每次运行的指标。
cs.AI / 159 / 2608.16148

FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection

FeatureHospital:一种基于技能驱动的多智能体框架,用于多视角多标签特征选择中的自动化算法定制
Li, Junxuan, Chen, Zhiqi, Liu, Yuzhou, Zhang, Peng, Liu, Huaxiao
Abstract
Multi-view multi-label feature selection aims to identify a compact and informative feature subset from heterogeneous views while preserving discriminative information for multiple labels. Existing methods are generally developed from specific modeling perspectives and incorporate mechanisms tailored to particular data characteristics. Designing suitable feature selection algorithms across datasets with diverse and heterogeneous characteristics still relies heavily on expert knowledge and substantial manual effort, imposing considerable time and labor costs that severely hinder the practical adoption of feature selection. To address this problem, we propose FeatureHospital, a Skill-driven multi-agent framework for automated multi-view multi-label feature selection algorithm design. FeatureHospital first diagnoses the target dataset to identify its feature selection issues. Based on the diagnosis, specialist agents equipped with domain Skills then prescribe corresponding optimization strategies and Loss terms for different issues. After that, the resulting prescriptions are reconciled to remove overlaps and resolve conflicts before being integrated into a compact dataset-specific objective. Finally, the constructed objective is optimized to select the final feature subset. Experimental results demonstrate that FeatureHospital can construct effective feature selection algorithms for different datasets based on their individual characteristics.
Chinese Translation
多视角多标签特征选择旨在从异构视角中识别出紧凑且信息丰富的特征子集,同时为多个标签保留区分性信息。现有方法通常从特定建模视角出发,结合针对特定数据特征量身定制的机制。在具有多样化和异构特征的数据集上设计合适的特征选择算法仍然高度依赖专家知识和大量人工努力,这带来了相当大的时间和劳动成本,严重阻碍了特征选择的实际应用。为了解决这一问题,我们提出了FeatureHospital,一种基于技能驱动的多智能体框架,用于自动化多视角多标签特征选择算法的设计。FeatureHospital首先对目标数据集进行诊断,以识别其特征选择问题。基于诊断结果,配备领域技能的专家代理随后为不同问题开出相应的优化策略和损失项。之后,生成的处方被协调以消除重叠和解决冲突,然后整合为一个紧凑的数据集特定目标。最后,构建的目标被优化以选择最终的特征子集。实验结果表明,FeatureHospital能够根据不同数据集的个体特征构建有效的特征选择算法。
cs.AI / 160 / 2608.16156

TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents

TRCA:面向长时间跨度大语言模型代理的过渡性评分信用分配
Zhang, Huan, Chen, Mingju, Zhou, Dongxu, Lv, Can, Chang, Heng, Cui, Sen, Wu, Faguo, Zhou, Shiji
Abstract
Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult. Existing approaches either rely on process evaluators, which incur annotation and inference costs, or derive step-level credit from successful trajectories. However, successful trajectories are extremely scarce during early-stage reinforcement learning, substantially weakening anchor-based methods. We propose Transition-wise Rubric Credit Assignment (TRCA), which derives step-level supervision directly from action-induced transitions without learned evaluators or successful anchors. TRCA evaluates each transition using Evidence, Execution, and Invalidity rubrics to capture task-relevant information acquisition, valid task execution, and invalid or regressive behavior. From these judgments, Foundational Rubric Reward measures local transition quality, while Breakthrough Rubric Reward tracks newly covered Evidence and Execution conditions to reward incremental task progress. Combined with terminal outcomes, these signals produce fine-grained step-level advantages for policy optimization. Experiments on ALFWorld, WebShop, and seven search-augmented question-answering benchmarks show consistent improvements over the evaluated baselines. With Qwen2.5-7B-Instruct, TRCA improves the WebShop score by 6.0%-12.6%; with Qwen2.5-3B-Instruct, it improves the average SearchQA score by 1.9%-18.3%. These results demonstrate the effectiveness of transition-wise rubric credit assignment for long-horizon tasks with sparse successful anchors.
Chinese Translation
长时间跨度的大语言模型(LLM)代理通常通过稀疏的终端结果进行优化,这使得在多步交互中进行细粒度的信用分配变得困难。现有的方法要么依赖于过程评估器,这会产生标注和推理成本,要么从成功轨迹中推导步骤级信用。然而,在早期强化学习阶段,成功轨迹极为稀缺,显著削弱了基于锚点的方法。我们提出了过渡性评分信用分配(TRCA),该方法直接从动作引发的过渡中推导步骤级监督,无需学习评估器或成功锚点。TRCA使用证据、执行和无效性评分标准评估每个过渡,以捕捉与任务相关的信息获取、有效的任务执行以及无效或退步的行为。根据这些判断,基础评分奖励(Foundational Rubric Reward)衡量局部过渡质量,而突破评分奖励(Breakthrough Rubric Reward)跟踪新覆盖的证据和执行条件,以奖励渐进的任务进展。结合终端结果,这些信号为策略优化提供了细粒度的步骤级优势。在ALFWorld、WebShop和七个增强搜索的问题回答基准上的实验显示出对评估基线的一致性改进。使用Qwen2.5-7B-Instruct,TRCA使WebShop得分提高了6.0%-12.6%;使用Qwen2.5-3B-Instruct,平均SearchQA得分提高了1.9%-18.3%。这些结果证明了过渡性评分信用分配在具有稀疏成功锚点的长时间跨度任务中的有效性。
cs.AI / 161 / 2608.16164

Trajectory-Level Automatic Curriculum Learning for Legged Locomotion on Unstructured Terrain

针对非结构化地形的腿部运动轨迹级自动课程学习
Liu, Rocky, Liu, Tengyu, Jia, Baoxiong, Zhong, Fangwei, Tong, Xinyi, Xie, Hongzhao, Huang, Siyuan
Abstract
Training locomotion policies for complex unstructured terrain requires a curriculum to avoid early exploration failures. However, since unstructured terrain lacks explicit difficulty ordering for curriculum design, existing methods resort to heuristic curricula over parameterized terrains. This abstraction limits generalization, as policies can overadapt to near-fixed perceptual patterns. To address this, we propose \textbf{\ourname{}}, an \textbf{T}rajectory-level \textbf{A}utomatic \textbf{C}urriculum \textbf{L}earning framework that generates training tasks directly from unstructured terrain maps. At each curriculum update, the evaluator learns a difficulty function for the current policy that maps a given trajectory task to a difficulty score. The sampler then proposes new trajectories guided by the learned evaluator as the curriculum for the next policy update. This forms a closed loop in which the curriculum is iteratively matched to the evolving policy. Quantitative and qualitative experiments show that \ourname{} continuously provides effective curricula on unstructured terrain, improving trajectory success rate by \(56.3\%\) over direct training without curriculum. Compared with handcrafted curriculum learning, our method improves success rate by \(18.5\%\) on the hardest terrain tasks and by up to \(39.74\%\) when evaluating traversal from diverse approach directions on the same obstacle type.
Chinese Translation
在复杂的非结构化地形上训练运动策略需要一个课程,以避免早期探索失败。然而,由于非结构化地形缺乏明确的难度排序用于课程设计,现有方法往往依赖于参数化地形上的启发式课程。这种抽象限制了泛化能力,因为策略可能会过度适应近乎固定的感知模式。为了解决这个问题,我们提出了 extbf{ heourname{}}, 一个 extbf{T}轨迹级 extbf{A}utomatic extbf{C}urriculum extbf{L}earning框架,该框架直接从非结构化地形图生成训练任务。在每次课程更新时,评估器为当前策略学习一个难度函数,该函数将给定的轨迹任务映射到一个难度评分。然后,采样器根据学习到的评估器提出新的轨迹,作为下一个策略更新的课程。这形成了一个闭环,其中课程与不断发展的策略迭代匹配。定量和定性实验表明, heourname{}在非结构化地形上持续提供有效的课程,成功率比没有课程的直接训练提高了56.3%。与手工设计的课程学习相比,我们的方法在最困难的地形任务上成功率提高了18.5%,在评估从不同接近方向穿越同一障碍物类型时,成功率提高了最多39.74%。
cs.AI / 162 / 2608.16192

Baseline-Relative Counterfactual Refinement for Bit-Aware Visual Token Communication

基于基线相对的反事实细化用于比特感知视觉令牌通信
Guo, Jia, Zhao, Xiaohan, Liu, Changwang, He, Shuqing, Zhang, Chenyang, Zhao, Bingchuan, Zhu, Jinqi
Abstract
Generative visual-token communication reduces transmission load by sending only selected discrete tokens and reconstructing missing content at the receiver. However, existing token-selection criteria based on local uncertainty, importance, or diversity do not directly determine whether changing the current selection improves the final reconstruction under the same packet budget. To address this problem, we propose Gated Counterfactual Refinement for Communication (GCR-C), a rollout-style correction layer over Local-MDL. GCR-C constructs a compact diversified candidate set, evaluates each candidate through matched full-budget Local-MDL continuation, and replaces the baseline action only when a positive baseline-relative reconstruction gain is obtained. Experiments on CIFAR-10, STL-10, a coded 5G-LDPC link, and a limited high-resolution Kodak transfer show that GCR-C consistently improves reconstruction quality at active low- and medium-rate operating points without increasing the realized packet rate, while remaining effective across changes in dataset, channel condition, resolution, token grid, and tokenizer. The results also reveal a clear quality--computation tradeoff due to the additional encoder-side counterfactual evaluation.
Chinese Translation
生成式视觉令牌通信通过仅发送选定的离散令牌并在接收端重建缺失内容来减少传输负载。然而,现有基于局部不确定性、重要性或多样性的令牌选择标准并不能直接确定在相同数据包预算下,改变当前选择是否能改善最终重建。为了解决这一问题,我们提出了用于通信的门控反事实细化(Gated Counterfactual Refinement for Communication,GCR-C),这是一种在局部最小描述长度(Local-MDL)之上的回滚式修正层。GCR-C构建了一个紧凑的多样化候选集,通过匹配的全预算局部最小描述长度继续评估每个候选,并仅在获得正的基线相对重建增益时替换基线动作。在CIFAR-10、STL-10、编码的5G-LDPC链路以及有限高分辨率Kodak传输的实验中,GCR-C在活跃的低和中等速率操作点上始终提高重建质量,而不增加实际数据包速率,同时在数据集、信道条件、分辨率、令牌网格和令牌化器的变化中保持有效。结果还揭示了由于额外的编码器侧反事实评估而导致的明显质量与计算之间的权衡。
cs.AI / 163 / 2608.16196

Beyond Asking: A Pipeline for Personalized Game Generation that Reads Players from Behavior

超越提问:一个通过行为读取玩家的个性化游戏生成管道
Lu, Yifan, Yuan, Xiaopeng, Wang, Haohan
Abstract
Personalized game generation requires inferring a player's abilities and behavioral style from how they play. Large language models have made this inference more attainable than ever: an LLM can read a raw gameplay transcript and produce a fluent, plausible profile of the player. Plausible, however, is not verified, and verification is precisely what the field lacks: latent traits are unobservable; questionnaires provide noisy proxies and become circular when self-reports are used to validate behavior-based inference; and behavior itself is ambiguous without context -- a player who never collects an item may not want it, or may never have had the chance. We address both problems. First, we construct a synthetic player population whose traits are ground truth by construction: each trait is an explicit bot parameter, accepted only after controlled manipulation produces consistent, trait-specific behavioral change. Unlike prior parameter-recovery work that inverts a known decision model, our benchmark evaluates policy-agnostic inference from behavioral transcripts alone. Second, we introduce an opportunity-aware decision-moment representation that disentangles preference from the chance to express it; ablating it selectively degrades opportunity-dependent traits. On this benchmark, few-shot LLM inference outperforms embedding- and rule-based baselines on most traits, though feature-based supervised regressors remain stronger overall. Finally, we close the loop: inferred profiles drive difficulty adaptation, evaluated against ground-truth references and mismatched-profile controls, and an exploratory human study examines whether these findings transfer to real players.
Chinese Translation
个性化游戏生成需要从玩家的游戏方式中推断出他们的能力和行为风格。大型语言模型使得这种推断比以往任何时候都更容易:一个大型语言模型(LLM)可以读取原始游戏过程记录,并生成流畅且合理的玩家档案。然而,合理性并不等于经过验证,而验证正是该领域所缺乏的:潜在特征是不可观察的;问卷提供了嘈杂的代理,并且当自我报告用于验证基于行为的推断时会变得循环;而行为本身在没有上下文的情况下是模糊的——一个从不收集物品的玩家可能并不是因为不想要它,或者可能根本没有机会。我们解决了这两个问题。首先,我们构建了一个合成玩家群体,其特征通过构造得到了真实的基础:每个特征都是一个明确的机器人参数,只有在经过控制的操控产生一致的、特征特定的行为变化后才被接受。与之前反转已知决策模型的参数恢复工作不同,我们的基准测试评估仅基于行为记录的无政策推断。其次,我们引入了一种机会感知的决策时刻表示,它将偏好与表达偏好的机会分开;选择性地消融它会削弱依赖机会的特征。在这个基准上,少量示例的LLM推断在大多数特征上优于嵌入和基于规则的基线,尽管基于特征的监督回归器总体上仍然更强。最后,我们闭合了循环:推断的档案驱动难度适应,并与真实参考和不匹配档案的对照组进行评估,一项探索性的人类研究考察了这些发现是否能够转移到真实玩家身上。
cs.AI / 164 / 2608.16207

Competing at Every Price Point with Agentic Evolution over a Menu of LLMs

在每个价格点上竞争:基于多种大型语言模型的自主进化
Borthwick, Andrew
Abstract
Consider a firm that surveys its competition for a particular agentic task and seeks to offer superior accuracy at every competitor price point. A firm that Pareto-dominated its competitors would leave no rational customer a reason to buy elsewhere. This paper shows a path to this kind of capability via agentic evolution over a menu of LLMs, from training pools of at most 100 examples. Given a priced menu of nine LLM endpoints; brief documentation of the task, objective, and API; a simple seed agent; and an operator-chosen per-problem cost target - usually set at an incumbent's own price - RoboPhD, an evolutionary meta-agent, evolves complete agent programs that attack the public frontiers of two semantically dissimilar tasks point by point: DS-1000 (execution-checked code generation) and PaperFindingBench (LLM-judged scientific document retrieval). Our officially scored submissions hold every Pareto-frontier slot but one on the two tasks' leaderboards, including Pareto domination of both the top-scoring and the lowest-cost competing points.
Chinese Translation
考虑一家企业对特定自主任务的竞争进行调查,并寻求在每个竞争对手的价格点上提供更高的准确性。一家在帕累托最优上超越其竞争对手的企业将不会给理性客户留下在其他地方购买的理由。本文展示了通过在多种大型语言模型(LLMs)上进行自主进化来实现这种能力的路径,训练样本池最多为100个例子。给定一个包含九个LLM端点的定价菜单;任务、目标和API的简要文档;一个简单的种子代理;以及一个由操作员选择的每个问题的成本目标——通常设定为现有竞争者的价格——RoboPhD,一个进化元代理,演化出完整的代理程序,逐点攻击两个语义上不相似任务的公共前沿:DS-1000(执行检查的代码生成)和PaperFindingBench(LLM评判的科学文献检索)。我们的官方评分提交在这两个任务的排行榜上占据了每个帕累托前沿位置,除了一个,包括对最高得分和最低成本竞争点的帕累托主导。
cs.AI / 165 / 2608.16211

BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics

BaT:面向自我进化的医疗研究代理的阶段性评分标准
Liu, Junqi, He, Yufan, He, Yexiao, Guo, Pengfei, Yang, Dong, Myronenko, Andriy, Zhao, Can, Ye, Hanrong, Qi, Tianhao, Zhou, Yuyin, Xu, Daguang, Tang, Yucheng
Abstract
Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workflows are multi-stage and data-sensitive, while expert trajectories remain scarce and difficult to share. Structured benchmarks can localize failures through stage-level rubrics, but standard post-training discards these diagnostics before the next training round. We present Benchmark-as-Teacher (BaT), a recursive self-improvement system for agent post-training. BaT contains two linked components: the asynchronous Stage Bank data pipeline and BiCuRL (Bilevel Curriculum Reinforcement Learning), its self-improving post-training method. Stage Bank synthesizes content-isolated training states outside the policy-update loop. BiCuRL uses a fixed held-out evaluation to select the next stage curriculum, verifies rollouts with task rubrics, updates the policy with GRPO, and returns the candidate checkpoint to evaluation. On AutoMedBench-Lite, BaT-4B and BaT-9B more than double the Overall scores of their Qwen Instruct baselines. BaT-9B Agent reaches 79.6 Overall, exceeding Claude Opus 4.6 with Claude Code at 77.5.
Chinese Translation
长远代理开始自动化完整的工作流程,生成代码、报告和研究文献。医疗影像工作流程是多阶段且数据敏感的,而专家轨迹仍然稀缺且难以共享。结构化基准可以通过阶段性评分标准定位失败,但标准的后训练过程在下一轮训练之前会丢弃这些诊断信息。我们提出了Benchmark-as-Teacher(BaT),这是一个用于代理后训练的递归自我改进系统。BaT包含两个相互关联的组件:异步的Stage Bank数据管道和BiCuRL(双层课程强化学习),其自我改进的后训练方法。Stage Bank在策略更新循环之外合成内容隔离的训练状态。BiCuRL使用固定的保留评估来选择下一个阶段课程,通过任务评分标准验证回滚,使用GRPO更新策略,并将候选检查点返回评估。在AutoMedBench-Lite上,BaT-4B和BaT-9B的总体得分超过其Qwen Instruct基线的两倍。BaT-9B代理的总体得分达到79.6,超过Claude Opus的4.6和Claude Code的77.5。
cs.AI / 166 / 2608.16213

Process-Constituted Intelligence: A Shared Criterion for Humans and Machines

过程构成的智能:人类与机器的共享标准
Richardson, Michael J., Alhasan, Ayeh, Crone, Cassandra, Monfort, M. Paula Diaz, Nalepka, Patrick, Dras, Mark, Kallen, Rachel W., Kaplan, David M.
Abstract
Intelligence is constituted by \textit{process} (iterative activity through which output emerges), not in the output itself. Generative AI (GenAI) is trained on \textit{traces} (textual and visual residues of human cognitive processes), reproducing samples from a distribution of those traces. Its outputs resemble reasoning, problem-solving, and creativity, yet the activity that produces such outputs in humans remains largely absent. Current GenAI is, therefore, weakly equivalent to the cognition it imitates, matching outputs while process stays absent or opaque. The cognitive sciences have long distinguished between weak and strong equivalence. Here, we define \textit{strong} equivalence across seven process features, assessable against human and machine cognition. Our process-based account addresses a symmetric risk: GenAI tools that outsource a person's generative processes may leave critical capacities unbuilt. We specify design principles for GenAI that instantiate more process and preserve rather than erode human judgment and creativity, and outline process audits that make strong equivalence testable.
Chinese Translation
智能是由 extit{过程}(通过迭代活动产生输出)构成的,而不是输出本身。生成性人工智能(Generative AI, GenAI)基于 extit{痕迹}(人类认知过程的文本和视觉残留物)进行训练,从这些痕迹的分布中再现样本。其输出类似于推理、问题解决和创造力,但在产生这些输出的人类活动中,仍然在很大程度上缺失。因此,当前的GenAI在模仿的认知上是弱等价的,匹配输出而过程却缺失或不透明。认知科学长期以来区分了弱等价和强等价。在这里,我们定义了跨越七个过程特征的 extit{强}等价,这些特征可以与人类和机器的认知进行评估。我们的基于过程的论述解决了一个对称风险:外包个人生成过程的GenAI工具可能会导致关键能力的缺失。我们具体阐明了GenAI的设计原则,这些原则体现了更多的过程,并保护而不是削弱人类的判断和创造力,并概述了使强等价可测试的过程审计。
cs.AI / 167 / 2608.16349

AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment

AeroCopilotBench:评估大型语言模型代理在互动虚拟驾驶舱环境中作为航空副驾驶的双层基准
Yuan, Yuchen, Wu, Zhenghuang, Li, Yuangan, Ma, Liang, Li, Ke
Abstract
Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments. This paper presents the AeroCopilot Operational Environment (ACOE), a reproducible interactive virtual-cockpit test environment, and AeroCopilotBench, a two-tier aviation agent evaluation benchmark. Tier-1 evaluates aviation knowledge using 1,200 multiple-choice questions, while Tier-2 comprises 73 emergency and abnormal tasks derived from the manufacturers' Pilot's Operating Handbooks (POHs) and instantiated in ACOE. ACOE converts natural-language procedures into executable state transitions, final-state goal conditions, and hard safety constraints, enabling models to interpret cockpit state, diagnose faults, and operate aircraft systems through standardized tool interfaces. We establish a safety-gated evaluation framework in which a trajectory succeeds only when all task goals are achieved without violating any hard safety constraint, while safe goal progress and trajectory safety are measured separately. Across 12 models, the highest Tier-2 success rate is 72.6%, while static knowledge performance does not consistently translate into procedural execution. Analysis of 451 failed episodes from 3 representative models identifies recurring failures in procedural completeness, use of state feedback, and long-horizon execution management. These findings motivate state-aware agent orchestration, joint assessment of task completion and trajectory safety, and repeated regression testing. ACOE and AeroCopilotBench provide a reproducible foundation for testing knowledge application, interactive execution, and operational safety in aviation agents.
Chinese Translation
大型语言模型(LLM)代理可以协助飞行机组人员进行复杂决策和任务执行,但现有的航空评估主要集中于静态知识,无法支持在互动环境中对程序执行和安全合规性的系统性测试。本文提出了AeroCopilot操作环境(ACOE),这是一个可重复的互动虚拟驾驶舱测试环境,以及AeroCopilotBench,一个双层航空代理评估基准。第一层通过1200道选择题评估航空知识,而第二层则由73个紧急和异常任务组成,这些任务源自制造商的飞行员操作手册(Pilot's Operating Handbooks, POHs)并在ACOE中实现。ACOE将自然语言程序转换为可执行的状态转换、最终状态目标条件和严格的安全约束,使模型能够解读驾驶舱状态、诊断故障并通过标准化工具接口操作航空系统。我们建立了一个安全门控评估框架,其中轨迹仅在所有任务目标在不违反任何严格安全约束的情况下实现时才算成功,同时安全目标进展和轨迹安全性是单独测量的。在12个模型中,第二层的最高成功率为72.6%,而静态知识的表现并未始终转化为程序执行。对3个代表性模型的451个失败案例的分析发现,程序完整性、状态反馈的使用和长时间执行管理存在重复性失败。这些发现促使我们关注状态感知的代理协调、任务完成与轨迹安全的联合评估,以及重复回归测试。ACOE和AeroCopilotBench为测试航空代理中的知识应用、互动执行和操作安全提供了可重复的基础。
cs.AI / 168 / 2608.16354

DriveCache: Action-Aware Caching for Driving World Model Inference

DriveCache:面向动作的驾驶世界模型推理缓存
Yang, Jianchun, Liang, Jian, Guo, Xianda, Fu, Pinhan, Peng, Yanlun, Zhang, Conglang, Huang, Wenke, Ye, Mang
Abstract
Driving video generation models support autonomous-driving development by predicting controllable future scenes for simulation, planning evaluation, and offline data generation. Diffusion-based driving generators repeatedly evaluate large backbones across denoising steps, which limits generation throughput. Existing diffusion acceleration methods reduce this cost, but general-purpose designs omit driving signals available before generation, such as ego speed and planned trajectories. Experiments across driving motions show that cache tolerance varies with ego translation and rotation, denoising progress, and consecutive reuse length. We propose DriveCache, a training-free, action-aware controller that uses planned motion to allocate reuse across scenes and dynamic programming to place it across denoising steps under a calibrated response budget. A causal drift check refreshes features and replans the remaining schedule when generation departs from calibration. Across three generator configurations, DriveCache improves the overall fidelity-efficiency trade-off over evaluated cache methods. Our code will be publicly available.
Chinese Translation
驾驶视频生成模型通过预测可控的未来场景来支持自动驾驶的发展,这些场景用于仿真、规划评估和离线数据生成。基于扩散的驾驶生成器在去噪步骤中反复评估大型骨干网络,这限制了生成的吞吐量。现有的扩散加速方法降低了这一成本,但通用设计忽略了生成前可用的驾驶信号,如自我速度和规划轨迹。针对驾驶运动的实验表明,缓存容忍度随着自我平移和旋转、去噪进度以及连续重用长度而变化。我们提出了DriveCache,这是一种无训练的、面向动作的控制器,利用规划的运动在场景间分配重用,并通过动态规划在校准的响应预算下将其分配到去噪步骤中。当生成偏离校准时,一个因果漂移检查会刷新特征并重新规划剩余的时间表。在三种生成器配置中,DriveCache在评估的缓存方法中改善了整体的保真度-效率权衡。我们的代码将公开发布。
cs.AI / 169 / 2608.16370

What Does Context Compression Cost an Agent? Interaction Costs Unrevealed by Task-Completion Metrics

上下文压缩对智能体的成本是什么?任务完成度指标未揭示的交互成本
Liu, Shuyu
Abstract
Task completion is the standard metric for evaluating context compression, yet it is incomplete: compression can increase an agent's interaction cost by forcing it to reacquire dropped state while leaving completion statistically unchanged. We introduce a controlled runtime measurement protocol for reacquisition cost in a bounded-horizon tool-using agent. The agent acts in a deterministic planning environment under a fixed 24-turn horizon. We vary compression severity, compare a dropping operator with a fact-preserving operator, restore dropped state through controlled oracle interventions, and decompose tool calls into retrieval and execution. We evaluate three models across two task regimes. Retrieval calls increase in all six model-regime comparisons and account for almost all added interaction; five of six remain significant after Holm correction. At the prespecified 5x comparison point, completion changes are not significant in any cell. DeepSeek shows a significant completion drop only at 10x compression. GPT-5.5 is the clearest case: completion changes from 80% to 85% (p = 1.0) while retrieval increases from 21.0 to 63.9 calls (p = .002). Retention interventions further separate state quantity, state type, and content validity. Random selection is comparable to an offline hindsight oracle, while replacing retained D-state with semantically irrelevant content increases retrieval by 57% (p < .001) without a significant completion change. In a second environment, ALFWorld, sliding compression produces no retrieval surge, showing that the reacquisition signature is environment-dependent rather than intrinsic to shortening context. Overall, compression can impose hidden interaction costs when execution-relevant state becomes absent and must be reacquired, while completion alone may not expose those costs.
Chinese Translation
任务完成度是评估上下文压缩的标准指标,但它并不完整:压缩可能会增加智能体的交互成本,因为它迫使智能体重新获取丢失的状态,而完成度在统计上保持不变。我们引入了一种受控的运行时测量协议,用于在有限视野的工具使用智能体中评估重新获取成本。该智能体在一个固定的24回合视野下的确定性规划环境中进行操作。我们改变压缩的严重程度,将丢弃操作符与保持事实的操作符进行比较,通过受控的预言干预恢复丢失的状态,并将工具调用分解为检索和执行。我们在两个任务模式下评估了三种模型。在所有六个模型-模式比较中,检索调用均有所增加,并几乎占据了所有新增的交互;在六个比较中,有五个在霍尔姆校正后仍然显著。在预设的5倍比较点上,任何单元的完成度变化均不显著。DeepSeek仅在10倍压缩时显示出显著的完成度下降。GPT-5.5是最明显的案例:完成度从80%变化到85%(p = 1.0),而检索调用从21.0增加到63.9(p = .002)。保留干预进一步区分了状态数量、状态类型和内容有效性。随机选择与离线回顾预言者相当,而用语义无关的内容替换保留的D状态则使检索增加了57%(p < .001),而完成度没有显著变化。在第二个环境ALFWorld中,滑动压缩并未产生检索激增,显示出重新获取特征是依赖于环境的,而不是固有于缩短上下文。总体而言,当执行相关的状态缺失并必须重新获取时,压缩可能会施加隐藏的交互成本,而仅凭完成度可能无法揭示这些成本。
cs.AI / 170 / 2608.16381

AstronOS: A Unified Execution Model and Runtime for Long-Horizon Agentic Systems

AstronOS:面向长时间跨度智能系统的统一执行模型与运行时
Nie, Zhenhang, Zheng, Gui, Sun, Xudong, Zhu, Tailong, Zhang, Bin
Abstract
Agentic systems often organize execution and state around a single conversation, model invocation, or agent instance, even when real work spans many calls and stages. We introduce a unified execution model that maintains a work item's persistent identity and versioned authoritative state across calls. Each step receives input scoped to a specific state version and new material; a result advances state only after validation and recording. We implement selected paths of this model in AstronOS using Cases, Tasks, and Scenario Packs across central and local execution. We compare five complete strategies for carrying an established software-version update plan into a fresh model session: rereading original materials, replaying full history, deterministic text summary, deterministic JSON, and the AstronOS runtime-mediated handoff. Ten controlled tasks are run under all five strategies with three repetitions, yielding 150 included executions. On the single-stage reference family, strategies perform similarly. In the primary three-stage A-C batch, AstronOS passes the frozen scorer in 14 of 15 executions, compared with 0 of 15 for rereading and 2 of 15 for full-history replay; later non-interleaved summary and JSON batches each pass 0 of 15. AstronOS has lower attempt-accounted model-token cost per passing execution, while requiring more execution-window time per attempt. These results associate the complete AstronOS condition with higher end-to-end pass rates across fresh sessions in this benchmark, at a measurable time cost.
Chinese Translation
智能系统通常围绕单一对话、模型调用或代理实例组织执行和状态,即使实际工作跨越多个调用和阶段。我们提出了一种统一的执行模型,该模型在调用之间保持工作项的持久身份和版本化的权威状态。每一步接收特定状态版本和新材料范围内的输入;结果仅在验证和记录后推进状态。我们在AstronOS中实现了该模型的选定路径,使用案例(Cases)、任务(Tasks)和场景包(Scenario Packs)在中央和本地执行中进行比较。我们比较了五种完整策略,以将既定的软件版本更新计划带入新的模型会话:重新阅读原始材料、重放完整历史、确定性文本摘要、确定性JSON,以及AstronOS运行时中介的交接。在所有五种策略下运行十个受控任务,每种策略重复三次,共计150次执行。在单阶段参考系列中,各策略表现相似。在主要的三阶段A-C批次中,AstronOS在15次执行中有14次通过冻结评分器,而重新阅读为0次,完整历史重放为2次;后续的非交错摘要和JSON批次均为0次通过。AstronOS在每次通过执行中具有较低的尝试计入模型令牌成本,但每次尝试所需的执行窗口时间较长。这些结果将完整的AstronOS条件与在此基准测试中新会话中的更高端到端通过率相关联,伴随可测量的时间成本。
cs.AI / 171 / 2608.16394

Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152

思考块内的内容:基于大型语言模型的合规场景生成的RegulaRAG:联合国第152号法规的案例研究
Zolfaghari, Vahid, Petrovic, Nenad, Schamschurko, AndrÉ, Knoll, Alois
Abstract
Generating regulation-compliant test scenarios is essential for validating safety-critical automotive systems, yet Large Language Models (LLMs) struggle to ground outputs in long, hierarchical standards. We present RegulaRAG, a Retrieval-Augmented Generation (RAG) pipeline that couples SmartChunking, reference-aware enrichment of paragraphs and tables via graph traversal, with Smart Retrieve & Rerank over these enriched units. To test our system, we evaluate on a manually curated dataset covering all scenarios in UN Regulation No. 152 (AEBS). Our study comprises: (i) a three-step progressive search that identifies near-optimal retrieval parameters without exhaustive grid search; (ii) head-to-head comparisons against five baseline RAG systems; and (iii) a robustness stress test that scales the source corpus with distractor content. Outputs are evaluated using a customized penalized scoring metric. Across all experiments, RegulaRAG achieves the highest average Meta-Score (82.99), outperforming the next-best system by 43% (NoRAG: 57.94), while operating at 14k-25k tokens per query versus up to 500k for graphcentric baselines. It maintains strong performance, remaining stable even as the number of regulatory sources grows, whereas competing RAG systems degrade sharply in both quality and robustness.
Chinese Translation
生成符合规定的测试场景对于验证安全关键的汽车系统至关重要,但大型语言模型(LLMs)在将输出与冗长的层次标准相结合时存在困难。我们提出了RegulaRAG,这是一种检索增强生成(RAG)管道,结合了智能分块(SmartChunking)、通过图遍历对段落和表格进行参考感知的丰富处理,以及对这些丰富单元的智能检索与重排序(Smart Retrieve & Rerank)。为了测试我们的系统,我们在一个手动策划的数据集上进行了评估,该数据集涵盖了联合国第152号法规(AEBS)中的所有场景。我们的研究包括:(i)一种三步渐进搜索,识别近似最优的检索参数,而无需耗时的网格搜索;(ii)与五个基线RAG系统的逐对比较;以及(iii)一个鲁棒性压力测试,通过干扰内容扩展源语料库。输出结果使用定制的惩罚评分指标进行评估。在所有实验中,RegulaRAG实现了最高的平均元评分(Meta-Score)82.99,超越了第二好的系统43%(NoRAG: 57.94),同时每个查询的操作在14k-25k个标记之间,而图中心基线则高达500k。它保持了强劲的性能,即使在监管源数量增加的情况下也保持稳定,而竞争的RAG系统在质量和鲁棒性上急剧下降。
cs.AI / 172 / 2608.16402

A Policy Algebra for Trust-Preserving Agentic AI Execution

一种用于信任保护的自主人工智能执行的政策代数
Tripathi, Bhaskar, Kumar, Anurag, Kumar, Ramendra, Gadhe, Bhavesh
Abstract
Large language model-based agentic frameworks primarily optimize capability: whether an agent can reason, retrieve information, call tools, delegate work, and complete a goal. Enterprise execution requires a stronger property. A successful result is not reliable if it was produced through unauthorized data access, widened delegated authority, unapproved side effects, unrecoverable budget consumption, or incomplete evidence. This paper defines reliable capability as a path property: an agent is reliably capable only when it completes a task through action events that remain admissible under identity, profile, tool, data, memory, budget, artifact, approval, and audit constraints. We propose a policy algebra that defines the reliability envelope within which agent capability may be exercised. Security profiles and runtime obligations compose through joins, intersections, budget narrowing, approval inheritance, and evidence accumulation; the resulting composition is both trust-preserving and the least restrictive state satisfying all governing inputs. The algebra also propagates restrictions across multi-agent calls and introduces cost-aware artifact materialization, which redirects open-ended execution toward a recoverable outcome as budget exposure grows. The evaluation is interpreted as a reliability-capability trade-off rather than a capability benchmark: the policy-algebra runtime intervenes on 94.8% of policy-violating events while retaining an 86.9% task-completion rate, eliminates the observed profile-monotonicity and zero-artifact-exhaustion violations, and increases audit completeness to 98.6%. The method provides researchers and practitioners with formal correctness conditions, executable decision semantics, and trace evidence for building agents that are not only capable, but reliably capable.
Chinese Translation
基于大型语言模型的自主框架主要优化能力:即代理是否能够推理、检索信息、调用工具、委派工作并完成目标。然而,企业执行需要更强的属性。如果结果是通过未经授权的数据访问、扩大委派权限、未批准的副作用、不可恢复的预算消耗或不完整的证据产生的,那么这个成功的结果是不可靠的。本文将可靠能力定义为一种路径属性:只有当代理通过在身份、个人资料、工具、数据、内存、预算、工件、批准和审计约束下仍然可接受的行动事件完成任务时,才被视为可靠能力。我们提出了一种政策代数,定义了代理能力可以被行使的可靠性范围。安全配置文件和运行时义务通过连接、交集、预算缩减、批准继承和证据积累进行组合;所得到的组合既保护信任,又是满足所有治理输入的最不限制状态。该代数还在多代理调用中传播限制,并引入成本感知的工件物化,随着预算暴露的增加,将开放式执行重定向到可恢复的结果。评估被解释为可靠性与能力的权衡,而不是能力基准:政策代数运行时对94.8%的政策违规事件进行干预,同时保持86.9%的任务完成率,消除了观察到的配置文件单调性和零工件耗尽违规,并将审计完整性提高到98.6%。该方法为研究人员和从业者提供了正式的正确性条件、可执行的决策语义和追踪证据,以构建不仅具备能力而且具备可靠能力的代理。
cs.AI / 173 / 2608.16421

Reasoning-supported Robustness Validation of Automotive E/E Components

基于推理的汽车电气/电子组件的鲁棒性验证
Novacek, Jan, Viehl, Alexander, Bringmann, Oliver, Rosenstiel, Wolfgang
Abstract
This paper presents an ontology-supported approach to tackle the complexity of the Robustness Validation (RV) process of automotive electrical/electronic (E/E) components. The approach uses formalized knowledge from the RV process and stress, operating, and load profiles, so-called Mission Profiles (MPs). In contrast to the error-prone industrially established manual procedure, we show how component characteristics are formalized in OWL in order to form the foundation of an efficient automated analysis selection and decision support during the RV process. The proposed approach is based on the idea of mapping MPs to an OWL representation so to allow to perform semantic queries against MP data to improve their integration into the RV process. The resulting ontology-supported application framework has been applied to an industrial use-case from automotive power electronics. We present experimental results showing that the RV process can be significantly improved in terms of reduced design time and increased exhaustiveness by automating the analyses selection step and the provisioning of all the relevant data to be used.
Chinese Translation
本文提出了一种基于本体的方式,以应对汽车电气/电子(E/E)组件鲁棒性验证(RV)过程的复杂性。该方法利用了来自RV过程的形式化知识以及应力、操作和负载特征,即所谓的任务特征(Mission Profiles, MPs)。与错误频发的工业手动程序相比,我们展示了如何将组件特性形式化为OWL,以便为RV过程中的高效自动化分析选择和决策支持奠定基础。所提方法基于将MP映射到OWL表示的思想,从而允许对MP数据进行语义查询,以改善其在RV过程中的整合。最终形成的基于本体的应用框架已应用于汽车电力电子的工业案例。我们展示了实验结果,表明通过自动化分析选择步骤和提供所有相关数据,RV过程在设计时间缩短和全面性提高方面可以显著改善。
cs.AI / 174 / 2608.16425

ParaTempo: Efficient Parallel Reasoning via Temporal Confidence

ParaTempo:通过时间置信度实现高效的并行推理
Zhang, Xuteng, Zeng, Wenhao, Gu, Xiaodong, Hu, Chao, Lin, Haotian, Shi, Yuling, Wang, Min, Shen, Beijun
Abstract
Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution paths, but its computational cost grows with reasoning depth and branch count. Existing methods for managing these parallel paths typically rely on final-answer consensus, local token confidence, or isolated intermediate probes. However, these signals are often delayed, weakly tied to actual reasoning progress, or too noisy for dynamic, branch-level control. To address these limitations, we introduce ParaTempo, a training-free asynchronous parallel reasoning framework. ParaTempo is driven by temporal confidence, a branch-local measure of answer-space convergence. Each branch is periodically probed for a tentative answer probability distribution, and temporal confidence quantifies how sharply the recent intermediate probes concentrate on a dominant answer. Once sufficient evidence has accumulated, ParaTempo drives its entire control process from this single signal: low-confidence branches are pruned, branches that persistently commit to their dominant answer are retired early, freed computation is reallocated by forking new branches, and generation stops globally once the confidence-weighted vote concentrates. Without requiring synchronization among reasoning trajectories, ParaTempo adaptively allocates computation based on branch-level convergence. Experiments on challenging mathematical and scientific reasoning benchmarks show that ParaTempo reduces average latency by 21.8-32.2% and total token usage by 18.1-30.3% while maintaining competitive accuracy. Moreover, temporal confidence exhibits stronger temporal stability and predictive power for future branch convergence than token-level and instantaneous signals.
Chinese Translation
并行推理通过探索多个解决路径提高了大型推理模型的准确性和鲁棒性,但其计算成本随着推理深度和分支数量的增加而增长。现有的管理这些并行路径的方法通常依赖于最终答案共识、局部令牌置信度或孤立的中间探测。然而,这些信号往往存在延迟、与实际推理进展的关联性较弱,或者对于动态的分支级控制来说噪声过大。为了解决这些局限性,我们提出了ParaTempo,一个无训练的异步并行推理框架。ParaTempo由时间置信度驱动,这是一种分支局部的答案空间收敛度量。每个分支定期被探测以获取一个暂定的答案概率分布,时间置信度量化了最近的中间探测在主导答案上集中程度的锐利程度。一旦积累了足够的证据,ParaTempo便从这一单一信号驱动其整个控制过程:低置信度的分支被修剪,持续承诺于其主导答案的分支被提前退役,释放的计算资源通过分叉新分支进行重新分配,并且一旦置信度加权投票集中,生成过程将全局停止。在不需要推理轨迹之间同步的情况下,ParaTempo根据分支级的收敛情况自适应地分配计算。对具有挑战性的数学和科学推理基准的实验表明,ParaTempo将平均延迟减少了21.8%-32.2%,总令牌使用量减少了18.1%-30.3%,同时保持了竞争力的准确性。此外,时间置信度在未来分支收敛方面表现出比令牌级和瞬时信号更强的时间稳定性和预测能力。
cs.AI / 175 / 2608.16435

Drive, Pack, Fly: The Travelling Thief Problem with Drone

驱动、打包、飞行:带无人机的旅行盗贼问题
Murjani, Kabir, Sobhanan, Abhay
Abstract
In collection operations, accumulating payload progressively slows the vehicle, imposing a cumulative penalty on routing efficiency. An onboard drone can offset this penalty by retrieving outlying items, thereby shortening the makespan and increasing operational profit. However, travel time remains load-dependent, and each item collected by the ground vehicle shifts the arrival times that govern the drone's launch and rendezvous points. This paper introduces the Travelling Thief Problem with Drone (TTP-D), which maximises the collected profit, net of a time-based rental cost, by jointly optimising item selection, vehicle routing, and flight synchronisation. We formulate a mixed-integer linear program that solves small instances to optimality, and develop both metaheuristics and an attention-based Deep Reinforcement Learning (DRL) policy for larger instances. We further propose a learner-initialised hybrid solver, in which the DRL policy constructs an initial solution that a short annealing run subsequently refines. On two benchmark sets, this hybrid recovers most of the metaheuristic baseline's quality at a fraction of its computational budget, although the largest instances still require the baseline at its full budget. Finally, a sensitivity analysis reveals that the rental ratio is the primary driver of profitability, whereas the fleet parameters affect profit only at the margin.
Chinese Translation
在收集操作中,累积的有效载荷逐渐减慢车辆的速度,从而对路线效率施加了累积惩罚。机载无人机可以通过回收偏远物品来抵消这一惩罚,从而缩短完成时间并增加操作利润。然而,旅行时间仍然依赖于负载,每个由地面车辆收集的物品都会影响无人机的发射和会合点的到达时间。本文引入了带无人机的旅行盗贼问题(TTP-D),通过联合优化物品选择、车辆路线和飞行同步,最大化扣除基于时间的租赁成本后的收集利润。我们制定了一个混合整数线性规划,能够对小规模实例求解至最优,并为较大实例开发了元启发式算法和基于注意力的深度强化学习(DRL)策略。我们进一步提出了一种学习者初始化的混合求解器,其中DRL策略构建了一个初始解,随后通过短时间的退火过程进行精炼。在两个基准数据集上,该混合方法在计算预算的一小部分内恢复了大部分元启发式基线的质量,尽管最大的实例仍然需要基线的全部预算。最后,敏感性分析表明,租赁比例是盈利能力的主要驱动因素,而车队参数仅在边际上影响利润。
cs.AI / 176 / 2608.16438

The Value of a Prompt: An LLM-Relative Kolmogorov-Complexity Approach

提示的价值:一种相对于 LLM 的 Kolmogorov 复杂性方法
Pass, Rafael
Abstract
In a world where valuable artifacts are increasingly created, completed, or processed by LLMs, the central economic question is not only what the LLM can produce, but what \emph{value} remains in the inputs (i.e., the prompts) we provide to it. Given a prompt, hint, critique, problem statement, or partial solution that helps an LLM produce an artifact $z$---a proof, program, design, or scientific hypothesis---how should we measure the value of that input? Intuitively, an input is valuable when it makes the target artifact easier for the model to generate: either by increasing its sampling probability, or by reducing the thinking time needed to find it. We propose a computational Levin--Kolmogorov complexity approach to this problem, by appropriately replacing the universal Turing machine in the classical definitions by the LLM itself. Concretely, we introduce an LLM-relative notion of \emph{probabilistic Levin--Kolmogorov complexity} $pKt$---treating the model's thinking as the random tape of the program, and charging logarithmically for it in Levin's manner---and define prompt value as algorithmic mutual information with respect to $pKt$. This captures the intuition above: a prompt having $b$ bits of value for an artifact $z$ makes $z$ $2^b$ times ``easier to obtain'', by multiplying the success probability by $2^b$, by dividing the required computation by $2^b$, or by any corresponding tradeoff between probability and computation. In contrast to the classical notion of algorithmic mutual information, ours is efficiently estimable. We additionally show that, under a natural reproduction experiment, a prompt value of \(b\) bits means that reproducing \(z\) without the prompt has median token cost \(2^b\) times that of reproducing it with the prompt.
Chinese Translation
在一个越来越多的有价值文物由 LLM(大型语言模型)创建、完成或处理的世界中,核心经济问题不仅在于 LLM 能够产生什么,还在于我们提供给它的输入(即提示)中剩余的 extit{价值}。给定一个提示、线索、批评、问题陈述或部分解决方案,这些都能帮助 LLM 生成一个文物 $z$——例如证明、程序、设计或科学假设——我们应该如何衡量该输入的价值?直观上,当输入使得模型生成目标文物变得更容易时,它就是有价值的:要么通过增加其采样概率,要么通过减少找到它所需的思考时间。我们提出了一种计算的 Levin-Kolmogorov 复杂性方法,通过适当地将经典定义中的通用图灵机替换为 LLM 本身。具体而言,我们引入了一种相对于 LLM 的 extit{概率 Levin-Kolmogorov 复杂性} $pKt$ 的概念——将模型的思考视为程序的随机带,并以 Levin 的方式对其进行对数计费——并将提示价值定义为相对于 $pKt$ 的算法互信息。这捕捉了上述直觉:一个提示对文物 $z$ 具有 $b$ 位的价值,使得 $z$ 变得“更容易获得” $2^b$ 倍,通过将成功概率乘以 $2^b$,通过将所需计算量除以 $2^b$,或通过概率与计算之间的任何相应权衡。与经典的算法互信息概念相比,我们的定义是可高效估计的。我们还展示了,在一个自然的再生产实验中,提示价值为 $b$ 位意味着在没有提示的情况下再生产 $z$ 的中位数令牌成本是使用提示再生产的 $2^b$ 倍。
cs.AI / 177 / 2608.16443

Time to Reason: Scalable Neurosymbolic Learning for LTLf via Fuzzy Semantics

推理的时机:通过模糊语义实现可扩展的神经符号学习以处理LTLf
Andreoni, Riccardo, Buliga, Andrei, Daniele, Alessandro, Felli, Paolo, Ghidini, Chiara, Montali, Marco, Ronzani, Massimiliano
Abstract
Neurosymbolic (NeSy) Artificial Intelligence aims to integrate Deep Learning (DL) architectures with symbolic reasoning. While initial NeSy approaches have targeted mainly symbolic reasoning in propositional and first-order logics, recent works have started to address the construction of neurosymbolic frameworks for Temporal Logics, and in particular for LTLf. These approaches have established temporal NeSy as a promising research direction, laying the foundations for learning under temporal constraints. Nonetheless, they leave many questions unanswered. From a theoretical perspective, several differentiable semantics for interpreting LTLf have been proposed but have not yet been formally and systematically defined within a unified framework. Moreover, existing approaches commonly rely on automata to represent temporal knowledge, resulting in limited scalability. Motivated by this research gap, this paper provides the following contributions: (i) formally defining different fuzzy semantics for LTLf, and systematically analysing theoretical properties regarding equivalences and dualities of temporal operators; (ii) showing how these semantics can be directly integrated within a novel NeSy framework, called DiffLTLf, enabling flexible and scalable learning without relying on the usage of automata; and (iii) introducing a novel evaluation protocol of increased complexity of learning tasks w.r.t. existing benchmarks. Our results show that the choice of fuzzy semantics has a significant impact on predictive performance. Moreover, DiffLTLf achieves performance on par with, and sometimes superior to, state-of-the-art probabilistic approaches while substantially improving scalability. Taken together, these results establish direct fuzzy interpretations as a competitive and scalable alternative to existing temporal NeSy frameworks.
Chinese Translation
神经符号(NeSy)人工智能旨在将深度学习(DL)架构与符号推理相结合。尽管最初的NeSy方法主要针对命题逻辑和一阶逻辑中的符号推理,但最近的研究开始关注构建用于时间逻辑的神经符号框架,特别是针对LTLf。这些方法确立了时间神经符号作为一个有前景的研究方向,为在时间约束下的学习奠定了基础。然而,它们仍然留下一些未解的问题。从理论角度来看,已经提出了几种可微分的LTLf解释语义,但尚未在统一框架内正式和系统地定义。此外,现有方法通常依赖于自动机来表示时间知识,导致可扩展性有限。基于这一研究空白,本文提供了以下贡献:(i)正式定义了LTLf的不同模糊语义,并系统分析了关于时间运算符的等价性和对偶性的理论属性;(ii)展示了如何将这些语义直接集成到一个新颖的神经符号框架中,称为DiffLTLf,从而实现灵活且可扩展的学习,而无需依赖自动机的使用;(iii)引入了一种新的评估协议,以提高学习任务的复杂性,相较于现有基准。我们的结果表明,模糊语义的选择对预测性能有显著影响。此外,DiffLTLf的性能与最先进的概率方法相当,有时甚至优于它们,同时显著提高了可扩展性。综合来看,这些结果确立了直接模糊解释作为现有时间神经符号框架的竞争性和可扩展替代方案。
cs.AI / 178 / 2608.16447

HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents

HaReCAP:递归大型语言模型代理的习惯性动作基础
Liu, Shen, Xu, Zhenguo, Wang, Shaopu, Gao, Yike, Wang, Chunlei
Abstract
Long-horizon embodied tasks require LLM agents to iteratively decompose high-level goals, revise plans in response to environmental feedback, and ground leaf-level subgoals into valid executable actions. Recursive context-management methods such as ReCAP improve planning stability through multi-level task decomposition and parent-node refinement, but still repeatedly invoke the LLM at leaf nodes to ground atomic subtasks into exact valid actions. We refer to this final grounding step as last-mile grounding redundancy, which accumulates into substantial LLM-call and token overhead during long-horizon execution. To mitigate this issue, we propose HaReCAP (Habitual-action Grounded ReCAP), a low-intrusion leaf grounding extension for ReCAP. HaReCAP extracts frequent leaf decisions from successful trajectories and compiles them offline into auditable and abstainable one-step leaf-reflex rules. At runtime, it skips the leaf LLM call only when a rule can uniquely determine a legal action in the current valid-action set; otherwise, it falls back to the original ReCAP. This design avoids repeatedly carrying the full recursive context into the LLM for routine leaf action grounding, while preserving the original recursive control flow. We evaluate HaReCAP on Robotouille and ALFWorld with Qwen3.5-27B as the main model. On tasks solved by both ReCAP and HaReCAP, HaReCAP reduces token consumption by 14.67%, 17.93%, and 20.08% on Robotouille synchronous, Robotouille asynchronous, and ALFWorld, respectively. The results show that HaReCAP can serve as a low-intrusion extension to ReCAP-style recursive context-management frameworks, reducing last-mile grounding redundancy across environments and models on commonly successful trajectories.
Chinese Translation
长时间跨度的具身任务要求大型语言模型(LLM)代理迭代地分解高层目标,针对环境反馈修订计划,并将叶级子目标基础化为有效的可执行动作。递归上下文管理方法如 ReCAP 通过多层任务分解和父节点细化提高了规划的稳定性,但仍在叶节点反复调用 LLM,将原子子任务基础化为确切的有效动作。我们将这一最终基础化步骤称为最后一公里基础化冗余,这在长时间执行过程中累积了大量的 LLM 调用和令牌开销。为了解决这一问题,我们提出了 HaReCAP(习惯性动作基础 ReCAP),这是 ReCAP 的一种低干扰叶基础化扩展。HaReCAP 从成功轨迹中提取频繁的叶决策,并将其离线编译为可审计和可放弃的一步叶反射规则。在运行时,只有当规则能够唯一确定当前有效动作集中的合法动作时,它才会跳过叶 LLM 调用;否则,它将回退到原始的 ReCAP。该设计避免了在常规叶动作基础化过程中反复将完整的递归上下文传递给 LLM,同时保留了原始的递归控制流。我们在 Robotouille 和 ALFWorld 上评估了 HaReCAP,以 Qwen3.5-27B 作为主要模型。在 ReCAP 和 HaReCAP 都能解决的任务中,HaReCAP 在 Robotouille 同步、Robotouille 异步和 ALFWorld 上分别减少了 14.67%、17.93% 和 20.08% 的令牌消耗。结果表明,HaReCAP 可以作为 ReCAP 风格递归上下文管理框架的低干扰扩展,减少在常见成功轨迹上跨环境和模型的最后一公里基础化冗余。
cs.AI / 179 / 2608.16465

JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills

JailbreakSkill:通过可重用和不断演变的技能扩展自动化红队测试
Wen, Xiaoyu, Li, Jiajia, He, Zhida, Yu, Peng, Wang, Chenxu, Qi, Han, Zhou, Ziyuan, Jin, Cheng, Wen, Ying, Xu, Xingcheng, Hu, Shuyue, Zheng, Tianhang, Lu, Chaochao, Zhang, Qiaosheng
Abstract
Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematically integrate, reuse, and improve at scale. We introduce \textsc{JailbreakSkill}, a skill-centric framework for scaling automated red-teaming through reusable and continuously evolving attack capabilities. \textsc{JailbreakSkill} packages existing attack strategies into modular, agent-ready skills that can be directly reused and adaptively selected across tasks and target models. Beyond reuse, it closes the loop between attacking and learning: attack experience is used to diagnose, refine, combine, and discover new skills, which are added back to an ever-growing skill library. This evolution lifts macro-average ASR by 17.5 percentage points on AdvBench and 13.4 points on HarmBench, including a 48.6-point gain against GPT-5.4 on AdvBench, while yielding novel attack strategies such as reframing a direct request as an unfinished document-completion task. Several evolved skills also generalize to unseen prompts and target models without further adaptation. Our code is available at https://github.com/BattleWen/JailbreakSkill.
Chinese Translation
自动化红队测试已经产生了一系列不断增长的攻击策略,但这些策略通常散布在不同的提示和工作流程中,使得它们难以系统性地集成、重用和大规模改进。我们介绍了 extsc{JailbreakSkill},一个以技能为中心的框架,通过可重用和持续演变的攻击能力来扩展自动化红队测试。 extsc{JailbreakSkill} 将现有的攻击策略打包成模块化、适合代理使用的技能,这些技能可以在任务和目标模型之间直接重用和自适应选择。除了重用,它还闭合了攻击与学习之间的循环:攻击经验被用来诊断、精炼、组合和发现新技能,这些新技能又被添加回不断增长的技能库中。这一演变使得在 AdvBench 上的宏平均攻击成功率提高了 17.5 个百分点,在 HarmBench 上提高了 13.4 个百分点,包括在 AdvBench 上对 GPT-5.4 的 48.6 个百分点的增益,同时产生了新的攻击策略,例如将直接请求重新构建为未完成的文档完成任务。几个演变后的技能也能够在未见过的提示和目标模型上进行泛化,而无需进一步的适应。我们的代码可在 https://github.com/BattleWen/JailbreakSkill 获取。
cs.AI / 180 / 2608.16482

Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation

重症监护室中脓毒症血流动力学管理的离线强化学习:基于MIMIC-IV的双重离政策略评估研究
Pérez-Roig, Marc, Fernández-Narro, David, Sáez, Carlos
Abstract
The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical judgment, which makes it a natural target for reinforcement learning from historical care. Because a learned policy cannot be trialed on patients, its value must be estimated off-policy, and such estimates can be fragile and optimistic. This work advances the reliable evaluation of sepsis treatment policies by combining off-policy estimation, reliability diagnostics, and clinician-agreement analyses in a transparent validation framework. We modeled fluid and vasopressor dosing on a cohort of 36,872 septic ICU stays drawn from the MIMIC-IV critical-care database, as a discretized Markov decision process with 1,000 states and 25 actions, defined by a five-by-five grid of fluid and vasopressor levels and solved by policy iteration. The clinicians' behavior policy was estimated with a random forest, which mitigated the collapse of the Effective Sample Size (ESS 50.1 against 4.0 with smoothed counts) that otherwise destabilizes the importance-sampling estimate. The learned policy was evaluated with two estimators, weighted importance sampling (WIS) and fitted Q evaluation (FQE), with the ESS and clinician agreement as reliability checks. An empirical variable selection found that the composition of the state matters more than its size. Both estimators place the learned policy above the clinicians' return (WIS 50.8 and FQE 46.8 against 38.2, ESS 50.1), yet it departs only modestly from observed practice (total variation 0.18), favoring less intravenous fluid. These retrospective single-center off-policy results support the learned policy as a clinically plausible refinement of observed practice and motivate its further evaluation as a discordance-based clinical decision-support approach.
Chinese Translation
在脓毒症中,静脉输液和血管收缩药物的剂量调整是一项在不确定性下进行的顺序决策,主要依赖临床判断,这使其成为从历史护理中进行强化学习的自然目标。由于学习到的策略无法在患者身上进行试验,因此其价值必须通过离政策略进行估计,而这种估计可能脆弱且过于乐观。本研究通过结合离政策略估计、可靠性诊断和临床医生一致性分析,在一个透明的验证框架中推进了脓毒症治疗策略的可靠评估。我们对来自MIMIC-IV重症监护数据库的36,872例脓毒症ICU住院病例进行了液体和血管收缩药物剂量的建模,将其视为一个离散的马尔可夫决策过程,具有1,000个状态和25个动作,由五乘五的液体和血管收缩药物水平网格定义,并通过策略迭代求解。临床医生的行为策略通过随机森林进行估计,从而减轻了有效样本量(Effective Sample Size, ESS)崩溃的问题(平滑计数下ESS为50.1,而非平滑计数下为4.0),该问题会使重要性抽样估计不稳定。学习到的策略使用加权重要性抽样(Weighted Importance Sampling, WIS)和拟合Q评估(Fitted Q Evaluation, FQE)两个估计器进行评估,并以ESS和临床医生一致性作为可靠性检查。经验变量选择发现,状态的组成比其大小更为重要。两个估计器均将学习到的策略的回报置于临床医生之上(WIS为50.8,FQE为46.8,而临床医生为38.2,ESS为50.1),但与观察到的实践仅有适度偏离(总变异为0.18),更倾向于较少的静脉输液。这些回顾性单中心的离政策略结果支持学习到的策略作为观察实践的临床上合理的改进,并激励其作为基于不一致性的临床决策支持方法进行进一步评估。
cs.AI / 181 / 2608.16507

Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling

大型语言模型作为合成临床专家以指导罕见疾病的纵向建模
Schächter, Clemens, Pechmann, Astrid, Kirschner, Janbernd, Hasenauer, Jan, Binder, Harald
Abstract
Due to the limited amount of information, modeling longitudinal rare-disease data can benefit from integrating clinical knowledge. Yet, elicitation of expert knowledge and formalization for model fitting is challenging, in particular due to limited time of clinical experts. To nevertheless make domain knowledge accessible during model fitting, we use large language models (LLMs) as synthetic clinical experts to supervise a variational-autoencoder-based approach that learns low-dimensional latent summaries of visit-level observations. Specifically, LLMs are queried offline on textual descriptions of patient observations to obtain judgments, e.g., the suspected clinical category. To improve the variational autoencoder fit, we train a differentiable surrogate model on these judgments and augment the loss function to encourage reconstructions that preserve the clinical-label distribution of their corresponding input profile. In an application to longitudinal motor-function assessments from children with spinal muscular atrophy, we map visit-level clinical profiles to low-dimensional representations that are linked by a multivariate mixed-effects model. The synthetic expert loss discourages reconstructions that remain numerically close in data space but alter the clinical interpretation of the reconstructed motor function profile, such as by crossing a disease-type boundary. We thus reduced disagreement between original and reconstructed SMA type labels from about 11 to 7 percent. Furthermore, informing the latent representation by the synthetic expert improved prediction of motor function milestones compared with unsupervised latent representations and a data-level baseline. These results suggest that incorporating LLMs into model fitting can make clinical knowledge available to representation learning and improve clinical faithfulness for longitudinal rare-disease data.
Chinese Translation
由于信息量有限,建模纵向罕见疾病数据可以通过整合临床知识获益。然而,专家知识的获取和模型拟合的形式化是具有挑战性的,尤其是由于临床专家的时间有限。为了在模型拟合过程中使领域知识可用,我们使用大型语言模型(LLMs)作为合成临床专家,监督基于变分自编码器的方法,该方法学习访问级观察的低维潜在摘要。具体而言,我们在患者观察的文本描述上离线查询LLMs以获取判断,例如,怀疑的临床类别。为了改善变分自编码器的拟合,我们在这些判断上训练一个可微分的替代模型,并增强损失函数,以鼓励重建保持其对应输入特征的临床标签分布。在对脊髓性肌萎缩症儿童的纵向运动功能评估的应用中,我们将访问级临床特征映射到通过多元混合效应模型关联的低维表示。合成专家损失抑制在数据空间中数值上接近但改变重建运动功能特征临床解释的重建,例如跨越疾病类型边界。因此,我们将原始和重建的SMA类型标签之间的不一致性从约11%降低到7%。此外,与无监督潜在表示和数据级基线相比,通过合成专家告知潜在表示改善了运动功能里程碑的预测。这些结果表明,将LLMs纳入模型拟合可以使临床知识可用于表示学习,并提高纵向罕见疾病数据的临床真实性。
cs.AI / 182 / 2608.16556

DeepInsight II: One Trace from Benchmark to Robot

DeepInsight II:从基准到机器人的一次追踪
Li, Siyi, Kang, Yuchen, Wang, Wuliang, Zhang, Zhengjie, Liu, Jiangpin, Yao, Jianhao, Chen, Jie
Abstract
Across a Physical AI stack, evaluation maturity is inversely aligned with deployment risk: foundation models enjoy mature, standardized harnesses, while the embodied layers on which deployment actually turns remain fragmented across benchmark-specific simulators, embodiments, and interfaces. The first DeepInsight report (v1) unified evaluation across this stack behind three abstractions---task, resource, and result---but its quantitative evidence centered on the foundation-model layer; navigation and manipulation (System 1) and whole-body control (System 0) remained simulation case studies, and physical execution was outside its empirical scope. DeepInsight II keeps that substrate fixed and quantifies the embodied half. First, it reproduces released-checkpoint references across two navigation and four manipulation benchmarks under their native protocols. Second, MotionBench places four released whole-body controllers under one workload and metric contract, then carries a qualified within-family cohort from parallel simulation to matched real-robot trials in which simulated and physical rollouts share a parent trace identity while retaining execution-domain-specific records, making the sim-to-real gap a native reduction rather than a reconciliation across toolchains. Third, a composed System 2--1--0 study extends trace localization into five evidence-grounded handoff labels, each mapped to a concrete repair action, with a measured repairability criterion and physical episodes testing the same attribution under hardware-observable state. The contribution is therefore not a new evaluation architecture, but empirical continuity from benchmark execution to matched robot evidence and repair-oriented diagnosis.
Chinese Translation
在物理人工智能堆栈中,评估成熟度与部署风险呈反比关系:基础模型享有成熟、标准化的框架,而实际部署所依赖的具身层则在基准特定的模拟器、具身实现和接口之间呈现出碎片化。第一版DeepInsight报告(v1)通过任务、资源和结果三种抽象统一了这一堆栈的评估,但其定量证据主要集中在基础模型层;导航和操作(系统1)以及全身控制(系统0)仍然是模拟案例研究,物理执行超出了其实证范围。DeepInsight II保持这一基础不变,并量化具身部分。首先,它在两个导航和四个操作基准下,按照其原生协议重现发布的检查点参考。其次,MotionBench将四个发布的全身控制器置于一个工作负载和度量合同之下,然后将一个合格的同类群体从平行模拟转移到匹配的真实机器人试验中,在这些试验中,模拟和物理的推出共享一个父追踪身份,同时保留执行领域特定的记录,使得模拟到现实的差距成为一种本质的缩减,而非跨工具链的调和。第三,组合的系统2--1--0研究将追踪定位扩展到五个基于证据的交接标签,每个标签映射到一个具体的修复动作,具有可测量的可修复性标准和测试相同归因的物理事件,在硬件可观察状态下进行。因而,贡献并不是一种新的评估架构,而是从基准执行到匹配机器人证据和以修复为导向的诊断的实证连续性。
cs.AI / 183 / 2608.16564

CUBICS: Situation-aware performance estimation for safety-relevant ML components

CUBICS:面向情境的安全相关机器学习组件性能估计
Herd, Benjamin, Kelly, Jessica, Trapp, Mario
Abstract
Machine learning (ML) is a key technology driving innovation today, but ensuring ML safety remains a major challenge for safety-related applications. A promising idea is to build proven-in-use arguments from field data, e.g. by running ML components (MLCs) in shadow mode or within safety envelopes so that their outputs can be monitored as 'safe probes' without affecting safety. These probes can then be used to build a statistical argument about field performance in a Bayesian way. However, many Bayesian field-data approaches in safety engineering model failures as a simple Bernoulli (or binomial) process with a single global failure probability and i.i.d. trials, which is rarely adequate for MLCs whose performance depends strongly on context. Statistical evidence is also about coverage of relevant situations, including edge cases, and building a single integrated statistical model for the entire system is usually not feasible. To address these challenges, this paper introduces CUBICS, a context-modular framework for per-component, situation-aware performance estimation of safety-relevant ML components. CUBICS partitions the operational design domain into situations and, for each safety-relevant component, defines a set of situation-specific assumptions and probabilistic guarantees that are represented and updated in a Bayesian manner using Subjective Logic (SL). By combining these guarantees with beliefs about how often each situation occurs, CUBICS derives an overall risk estimate for each component without requiring a monolithic system-level statistical model, and thus provides a building block for modular, field-data based safety assurance.
Chinese Translation
机器学习(ML)是推动当今创新的关键技术,但确保机器学习的安全性仍然是安全相关应用的一大挑战。一种有前景的想法是通过现场数据构建使用证明论证,例如通过在影子模式下运行机器学习组件(MLCs)或在安全边界内运行,以便监控其输出作为“安全探针”,而不影响安全性。这些探针可以用于以贝叶斯方式构建关于现场性能的统计论证。然而,许多安全工程中的贝叶斯现场数据方法将故障建模为简单的伯努利(或二项)过程,具有单一的全局故障概率和独立同分布(i.i.d.)试验,这对于性能强烈依赖于上下文的机器学习组件而言往往是不够的。统计证据还涉及相关情境的覆盖,包括边缘案例,而为整个系统构建单一的综合统计模型通常是不可行的。为了解决这些挑战,本文介绍了CUBICS,一个面向每个组件的、情境感知的安全相关机器学习组件性能估计的上下文模块化框架。CUBICS将操作设计域划分为不同情境,并为每个安全相关组件定义一组特定情境的假设和概率保证,这些保证以贝叶斯方式使用主观逻辑(Subjective Logic, SL)表示和更新。通过将这些保证与关于每种情境发生频率的信念相结合,CUBICS为每个组件推导出整体风险估计,而无需单一的系统级统计模型,从而为基于现场数据的模块化安全保障提供了构建块。
cs.AI / 184 / 2608.16565

Probabilistic Circuits as Reasoning Machines in Artificial Intelligence (Part I)

概率电路作为人工智能中的推理机器(第一部分)
Peharz, Robert
Abstract
This cumulative habilitation thesis studies probabilistic circuits (PCs) as a powerful and tractable framework for reasoning and learning under uncertainty in artificial intelligence (AI). It first advocates for probability as a core language for AI, emphasizing its connections to logic and information theory; the conceptual simplicity of probabilistic reasoning---based primarily on the sum and product rules; the parallels between probabilistic inference and human cognition; and the role of probability in optimal decision making. However, probability also faces significant computational challenges, as probabilistic inference is NP-hard in almost all probabilistic models. PCs address these challenges through structural constraints that ensure exact computation of a wide range of inference queries in polynomial time, such as marginals, conditionals, most probable explanations, expectations, and more advanced inference tasks. This thesis synthesizes a decade of research across foundations, algorithmic developments, and empirical validation of PCs. Key contributions highlighted in this work are foundational theory of PCs, Bayesian approaches for learning PCs, scalable implementations and integration with deep learning, hybrid models that combine PCs with intractable models, and connections with symbolic machine learning paradigms. This is the first part of my Habilitation Thesis. The second part is omitted, as it comprises the cumulative part of the thesis and has been published at various venues (see Chapter 5).
Chinese Translation
本累积资格论文研究了概率电路(PCs)作为一种强大且可处理的不确定性推理与学习框架在人工智能(AI)中的应用。首先,论文主张将概率视为人工智能的核心语言,强调其与逻辑和信息理论的联系;概率推理的概念简单性——主要基于求和和乘积规则;概率推理与人类认知之间的相似性;以及概率在最优决策中的作用。然而,概率也面临着显著的计算挑战,因为在几乎所有概率模型中,概率推理都是NP难的。概率电路通过结构约束来应对这些挑战,确保在多项式时间内精确计算广泛的推理查询,例如边际分布、条件概率、最可能解释、期望值以及更高级的推理任务。本文综合了十年的研究,涵盖了概率电路的基础、算法发展和实证验证。本研究的关键贡献包括概率电路的基础理论、学习概率电路的贝叶斯方法、可扩展的实现与深度学习的结合、将概率电路与不可处理模型相结合的混合模型,以及与符号机器学习范式的联系。这是我的资格论文的第一部分。第二部分被省略,因为它包含论文的累积部分,并已在多个场合发表(见第五章)。
cs.AI / 185 / 2608.16578

Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents

代理的物理学:统计力学预测人工智能代理的集体行为
El, Batu, Paeng, Jinhee, Dinc, Fatih, Su, Shiye, Erdogan, Mete, Pappu, Aneesh, Ye, Haotian, Zhao, Wanjia, Ganguli, Surya, Zou, James
Abstract
AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplify shared biases. Understanding and predicting these collective dynamics is therefore important for designing effective and aligned multi-agent systems. Here, we study over 10,000 communities of language-model agents that repeatedly exchange messages and revise their opinions across objective mathematics questions and subjective political statements. Despite substantial diversity in possible behavior, the individual and group dynamics can be represented by three characteristic regimes: indifference, polarization, and consensus. AI agents start indifferent and build conviction as they interact. On objective questions, communication improves collective accuracy, while on subjective questions it often drifts group opinions toward the right in the political spectrum. We explain these observations with a statistical-mechanics formalism in which agents stochastically favor lower social pressure. Given only initial opinions, our model predicts individual trajectories, outperforms all standard baselines, generalizes to unseen community graphs, and reproduces the observed group archetype distributions. Our fitted model parameters reveal the mechanics underlying our key observations: i) communities operate below the critical social temperature, which explains conviction buildup; ii) attractive ties outweigh repulsive ones, which favors consensus; and iii) agents holding the correct answer exert the strongest pull, which drives truth-seeking. Overall, our results demonstrate that collective behavior of AI agents, like that of other complex systems, follows compact and predictive dynamical laws.
Chinese Translation
人工智能代理越来越多地作为相互作用系统的一部分而非孤立存在。当代理交换信息并共同做出决策时,它们的互动可以改善集体推理,但也可能导致群体效应、极化或放大共享偏见。因此,理解和预测这些集体动态对于设计有效且一致的多代理系统至关重要。在此,我们研究了超过10,000个语言模型代理的社区,这些代理反复交换信息并在客观数学问题和主观政治陈述上修正其观点。尽管可能的行为存在显著多样性,个体和群体动态可以用三种特征状态来表示:无差别、极化和共识。人工智能代理最初表现为无差别,并随着互动建立信念。在客观问题上,交流提高了集体准确性,而在主观问题上,它通常将群体意见向政治光谱的右侧漂移。我们用统计力学形式化来解释这些观察结果,其中代理随机偏向于较低的社会压力。仅凭初始意见,我们的模型能够预测个体轨迹,超越所有标准基线,推广到未见过的社区图,并再现观察到的群体原型分布。我们拟合的模型参数揭示了我们关键观察结果背后的机制:i) 社区在临界社会温度以下运行,这解释了信念的积累;ii) 吸引性联系超过排斥性联系,这有利于达成共识;iii) 持有正确答案的代理施加最强的吸引力,推动寻求真理。总体而言,我们的结果表明,人工智能代理的集体行为与其他复杂系统一样,遵循紧凑且可预测的动态法则。
cs.AI / 186 / 2608.16594

CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction

CACSurv:基于大语言模型的肿瘤生存预测的一致性对齐比较学习
Xiang, Tianqi, Zhang, Qixiang, Ding, Xinpeng, Li, Yi, Li, Xiaomeng
Abstract
Cancer survival prediction supports treatment planning, risk stratification, and follow-up management. Existing methods use structured clinical variables, whole-slide images, genomic profiles, or multimodal inputs, while patient reports remain underexplored. We study report-centric survival prediction using reports that organize pathological, clinical, and molecular evidence. Large language models (LLMs) can reason over such reports, but case-wise time regression introduces two mismatches. First, a formulation mismatch arises because survival evaluation depends on ordering comparable patients, whereas independent time predictions do not enforce ranking consistency. Second, a supervision mismatch arises because a censored patient's observed time indicates survival beyond that point and cannot serve as an exact regression target, although it still implies orderings relative to patients who died earlier. To address these mismatches, we propose CACSurv, a Concordance-Aligned Comparative framework for report-centric survival prediction. CACSurv reformulates survival modeling as mini-cohort comparative reasoning, where an LLM predicts relative prognostic orderings. We introduce concordance-aligned rewards derived from comparable relations under right censoring, enabling censored outcomes to provide ranking supervision without exact event-time targets. At inference, Monte Carlo Reference Aggregation compares each patient with sampled references and aggregates positions into a cohort-level ranking. We establish TCGA-SurvReport, a benchmark covering six TCGA cancer cohorts. CACSurv achieves the highest C-index on all six cohorts and an average C-index of 0.722, outperforming the strongest published survival model by 6.5 percentage points and the strongest LLM time-regression baseline by 4.2 percentage points. Our code, models, and dataset will be available at https://github.com/xmed-lab/CACSurv.
Chinese Translation
肿瘤生存预测支持治疗规划、风险分层和后续管理。现有方法使用结构化临床变量、全切片图像、基因组特征或多模态输入,而患者报告仍未得到充分探索。我们研究了以报告为中心的生存预测,利用组织病理、临床和分子证据的报告。大语言模型(LLMs)能够对这些报告进行推理,但逐例时间回归引入了两个不匹配。首先,形成不匹配的原因在于生存评估依赖于可比患者的排序,而独立时间预测并未强制执行排名一致性。其次,监督不匹配的原因在于被审查患者的观察时间表明其生存超出该点,并不能作为精确的回归目标,尽管它仍然暗示相对于早逝患者的排序。为了解决这些不匹配,我们提出了CACSurv,一个用于以报告为中心的生存预测的一致性对齐比较框架。CACSurv将生存建模重新表述为小组比较推理,其中LLM预测相对预后排序。我们引入了基于右删失下可比关系推导的一致性对齐奖励,使得被审查结果能够在没有精确事件时间目标的情况下提供排名监督。在推理时,蒙特卡洛参考聚合将每个患者与采样参考进行比较,并将位置聚合为组级排名。我们建立了TCGA-SurvReport,一个涵盖六个TCGA癌症队列的基准。CACSurv在所有六个队列中实现了最高的C指数,平均C指数为0.722,超越了最强的已发布生存模型6.5个百分点和最强的LLM时间回归基线4.2个百分点。我们的代码、模型和数据集将可在https://github.com/xmed-lab/CACSurv获取。
cs.AI / 187 / 2608.16621

Cost Scales with Change, Not Corpus Size: Incrementally Maintaining an Evolving Semantic Substrate

成本与变化相关,而非语料库大小:增量维护演变的语义基底
Takahashi, Yusuke, Wild, Kyle, Uraki, Asako
Abstract
Retrieval-augmented and agentic question-answering systems increasingly re-derive the meaning of a corpus at query time. Put plainly, instead of re-deriving what a corpus means on every question, the work is done once when a document arrives and is thereafter merely consulted -- a compiler, not an interpreter, of meaning. An alternative is to compile that meaning once, at ingest time, into a compact, queryable semantic substrate and maintain it as the corpus evolves. The central objection is maintenance cost: rebuilding a truncated singular value decomposition (SVD) on every change appears prohibitive, and a change of embedding model seems to force a full re-embedding. We argue and show empirically that maintenance cost scales with the amount of change, not corpus size. On a controlled synthetic pilot (dimension 256, rank 32, a corpus grown from 3,000 to 9,000 documents over 50 update events), incremental low-rank updates were 33.7 times cheaper per update than full re-SVD and 23.8 times cheaper cumulatively, while the incremental subspace tracked the full recomputation to within floating-point precision (maximum principal-angle drift below 1e-11 degrees; recall@10 = 1.0). An orthogonal Procrustes virtual axis update recovered 0.95 mean cosine to truly re-embedded vectors by re-embedding only about 10 percent of the corpus. The results support maintaining, rather than repeatedly reconstructing, a semantic substrate.
Chinese Translation
增强检索和自主问答系统在查询时越来越多地重新推导语料库的含义。简单来说,与其在每个问题上重新推导语料库的含义,不如在文档到达时进行一次性推导,之后仅进行咨询——它是意义的编译器,而非解释器。另一种选择是在摄取时将该含义编译成一个紧凑的、可查询的语义基底,并在语料库演变时进行维护。主要的反对意见是维护成本:在每次变化时重建截断的奇异值分解(SVD)似乎是不可承受的,而更换嵌入模型似乎迫使进行全面的重新嵌入。我们论证并通过实证表明,维护成本与变化量相关,而非语料库大小。在一个受控的合成试点(维度256,秩32,语料库在50次更新事件中从3000个文档增长到9000个文档),增量低秩更新的每次更新成本比全面重新SVD便宜33.7倍,累计便宜23.8倍,同时增量子空间的跟踪精度达到了浮点精度(最大主角漂移低于1e-11度;recall@10 = 1.0)。一个正交的Procrustes虚轴更新通过仅重新嵌入约10%的语料库,恢复了0.95的平均余弦与真正重新嵌入向量的相似度。结果支持维护而非反复重建语义基底。
cs.AI / 188 / 2608.16626

A Shop Floor Production Scheduling Case based on RFID-supported Smart Factory

基于RFID支持的智能工厂的车间生产调度案例
Chen, Zhihui, Sun, Yize, Dong, Yuhao, Xiao, Zeyu, Zhong, Ray Y.
Abstract
Radio frequency identification (RFID) technology has been widely implemented for real-time data collection in manufacturing shop floors, which, in turn, can be used to support dynamic shop floor production planning and scheduling. Within such an environment, uncertainty in operation and production processes collectively contribute to the dynamicity in manufacturing, thereby hampering the scheduling system from achieving maximal utility. To highlight the importance of handling such uncertainty, this paper addresses the problem of dynamic shop floor scheduling for a real-life case smart factory equipped with RFID technology. Feasible production sequence mining and real-time processing rate estimation are conducted on RFID-collected production data to quantify the operation and production uncertainties. A deep reinforcement learning approach based on the RFID data analysis is then presented for shop floor production scheduling. Simulation studies based on real-life case data have demonstrated the feasibility and practicality of the proposed dynamic production scheduling framework. Specifically, it is observed that the proposed framework outperforms existing dispatch methods in terms of minimizing operation makespan, including first in first out (FIFO), last in first out (LIFO) and deep Q network (DQN).
Chinese Translation
射频识别(RFID)技术已广泛应用于制造车间的实时数据收集,这反过来可以支持动态的车间生产计划和调度。在这样的环境中,操作和生产过程中的不确定性共同导致了制造业的动态性,从而妨碍了调度系统实现最大效用。为了强调处理这种不确定性的重要性,本文针对一个配备RFID技术的实际案例智能工厂,解决了动态车间调度的问题。对RFID收集的生产数据进行可行的生产序列挖掘和实时处理速率估计,以量化操作和生产的不确定性。然后,基于RFID数据分析提出了一种深度强化学习方法用于车间生产调度。基于实际案例数据的仿真研究证明了所提出的动态生产调度框架的可行性和实用性。具体而言,观察到所提出的框架在最小化操作完工时间方面优于现有的调度方法,包括先进先出(FIFO)、后进先出(LIFO)和深度Q网络(DQN)。
cs.AI / 189 / 2608.16628

Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement

基于超图的多模态检索增强生成与增量精炼
Chen, Shenao, Xu, Yidan, Han, Xiangmin, Xue, Rundong, Wu, Duanpo, Gao, Yuhan, Yan, Chenggang, Gao, Yue
Abstract
Modern Multimodal Retrieval-Augmented Generation (M-RAG) systems are fundamentally limited by the binary connectivity paradigm of traditional simple graphs, which fails to capture the intricate, high-order correlations among heterogeneous entities, such as the N-ary relationships between a visual chart, its scattered textual descriptions, and underlying numerical data. Furthermore, existing refinement strategies often rely on exhaustive, full-page reconstruction to align cross-modal information, leading to prohibitive computational redundancy and the introduction of contextual noise in long-form document processing. In this paper, we propose Hyper-M2RAG, a novel framework that redefines multimodal document retrieval through High-order Hypergraph Representation Learning. We first formalize the document structure as a Multimodal Hypergraph, utilizing hyperedges as unified semantic containers to encapsulate multi-way associations across text, images, and tables, thereby transcending point-to-point modeling. To mitigate semantic fragmentation caused by physical pagination, we introduce an Anchor-driven Incremental Refinement mechanism. Rather than performing a global sweep, our approach identifies boundary-crossing anchor nodes and reconstructs their local hyper-topology using one-hop neighborhood contexts. This targeted refinement effectively bridges cross-page knowledge gaps with minimal computational footprints. Extensive evaluations on multimodal benchmarking datasets demonstrate that Hyper-M2RAG significantly outperforms state-of-the-art methods in both retrieval precision and generation coherence. Our code is available at https://github.com/ShenAoChen2001/MMHRAG.
Chinese Translation
现代多模态检索增强生成(M-RAG)系统在根本上受到传统简单图的二元连接范式的限制,这种范式无法捕捉异构实体之间复杂的高阶关联,例如视觉图表、其分散的文本描述和潜在的数值数据之间的N元关系。此外,现有的精炼策略通常依赖于全面的全页重构来对齐跨模态信息,这导致了计算冗余的增加,并在长文档处理过程中引入了上下文噪声。在本文中,我们提出了Hyper-M2RAG,这是一种通过高阶超图表示学习重新定义多模态文档检索的新框架。我们首先将文档结构形式化为多模态超图,利用超边作为统一的语义容器,以封装文本、图像和表格之间的多向关联,从而超越点对点建模。为了减轻由于物理分页造成的语义碎片化,我们引入了一种基于锚点的增量精炼机制。我们的做法不是进行全局扫描,而是识别跨边界的锚点节点,并利用一跳邻域上下文重构其局部超拓扑。这种有针对性的精炼有效地弥合了跨页知识的空白,且计算开销最小。在多模态基准数据集上的广泛评估表明,Hyper-M2RAG在检索精度和生成一致性方面显著优于现有最先进的方法。我们的代码可在 https://github.com/ShenAoChen2001/MMHRAG 获取。
cs.AI / 190 / 2608.16637

PDDLCoder: Agentic PDDL Generation for LLM-Assisted Symbolic Planning

PDDLCoder:用于大语言模型辅助符号规划的自主PDDL生成
Laule, Veit, Shuai, Jiangtao, Hauswirth, Manfred, Schimmler, Sonja
Abstract
LLMs remain unreliable for long-horizon planning, often generating logically inconsistent or non-applicable plans. Recent hybrid methods instead translate natural language into the Planning Domain Definition Language (PDDL), allowing symbolic planners to produce verifiable plans. However, existing methods frequently rely on rigid generation pipelines, a partial PDDL definition, or human feedback. Furthermore, their evaluation is hindered by the lack of standardized benchmarks with automated verification. To address these limitations, we present PDDLCoder, an agentic framework for PDDL generation from natural language that iteratively generates, analyzes, and refines planning specifications. We further introduce NL-pddlgym, a benchmark dataset comprising 711 planning problems across 23 domains with executable gym environments for the automated verification of plan applicability. Experiments on the NL-pddlgym test set containing 106 problems across 4 held-out domains show that PDDLCoder generates applicable plans for 89.6\% of tested planning problems. This improves upon our adaptations of previous PDDL generation methods, which achieved up to 45.3\%, and outperforms direct LLM planning approaches, which reached up to 74.5\% on the same test set. Our work demonstrates the effectiveness of agentic PDDL generation for planning and establishes a reproducible benchmark for future research on LLM-assisted symbolic planning.
Chinese Translation
大语言模型(LLMs)在长时间规划中仍然不可靠,常常生成逻辑上不一致或不适用的计划。最近的混合方法则将自然语言翻译为规划领域定义语言(Planning Domain Definition Language, PDDL),使得符号规划器能够生成可验证的计划。然而,现有方法通常依赖于僵化的生成管道、部分PDDL定义或人工反馈。此外,由于缺乏标准化的基准和自动验证,它们的评估受到限制。为了解决这些局限性,我们提出了PDDLCoder,这是一个从自然语言生成PDDL的自主框架,能够迭代生成、分析和完善规划规范。我们进一步引入了NL-pddlgym,这是一个基准数据集,包含711个规划问题,涵盖23个领域,并提供可执行的gym环境以实现计划适用性的自动验证。在包含106个问题的NL-pddlgym测试集上的实验表明,PDDLCoder能够为89.6%的测试规划问题生成适用的计划。这一结果优于我们对之前PDDL生成方法的改编,其成功率为45.3%,并且超越了直接的LLM规划方法,在同一测试集上的成功率为74.5%。我们的工作展示了自主PDDL生成在规划中的有效性,并为未来关于LLM辅助符号规划的研究建立了可重复的基准。
cs.AI / 191 / 2608.16645

Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies

重建:从预出版书目中恢复研究思想的盲基准
Chen, Shaolong, Fei, Yanlin, Liu, Nazhou, Yu, Xinmiao, Li, Lei, Thapa, Rahul, Ciobanu, Madalina, Mao, Qingqing, Das, Ritankar
Abstract
Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature, and asks models to propose hypotheses that an independent large language model judge matches against the held-out ground-truth idea. A strict anti-leakage protocol-temporal citation cutoff, anonymous reference IDs, and frozen per-paper bibliographies, which prevents prompt-time leakage of the seed idea. Across six scientific domains and 643 evaluated papers, seven frontier models achieve only modest Match rates (approx. 3-15%). We then evaluate a reference-only multi-agent (top 4) pipeline that combines cross-model review with a Swiss tournament over aligned hypothesis slots, without external web search. Cross-model review plus tournament selection raises Match rates to approx. 23-42% across all six domains, which is an observed approx. 2.4x lift over the best single-model baseline. This draft reports the protocol, anti-leakage design, and current results as an arXiv timestamp.
Chinese Translation
当仅提供一篇论文的预出版书目时,语言模型能否恢复该论文的真实研究思想?我们引入了重建(Reconstruction),这是一个盲目的思想恢复基准,它隐去种子论文及所有同时期或未来文献,并要求模型提出假设,由一个独立的大型语言模型判断这些假设与保留的真实思想的匹配程度。我们采用严格的反泄漏协议,包括时间引用截止、匿名参考ID以及冻结的每篇论文书目,以防止在提示时泄漏种子思想。在六个科学领域和643篇评估论文中,七个前沿模型的匹配率仅为适度(约3-15%)。随后,我们评估了一种仅基于参考的多代理(前四名)管道,该管道结合了跨模型审查和对齐假设槽的瑞士锦标赛,而不依赖外部网络搜索。跨模型审查加上锦标赛选择使得匹配率在所有六个领域提高到约23-42%,相比最佳单模型基线观察到约2.4倍的提升。本文草稿报告了该协议、反泄漏设计及当前结果,并附有arXiv时间戳。
cs.AI / 192 / 2608.16666

Chronocooked: A Benchmark for Implicit Interval Timing in Reinforcement Learning Agents

Chronocooked:用于强化学习智能体隐式时间间隔计时的基准测试
Pednekar, Amrapali, Garrido-Perez, Alvaro, Khaluf, Yara, Simoens, Pieter
Abstract
This paper presents Chronocooked, a reinforcement learning (RL) benchmark suite for studying implicit interval timing in RL agents. Inspired by Overcooked, the suite comprises cooking scenarios that require temporal decision making. The tasks and reward functions are designed such that temporal information is unobserved yet critical for optimal performance. The environment is intentionally kept simple to enable controlled experiments and support biologically plausible models. Evaluation metrics are designed to expose limitations in timing abilities of RL agents, and we report baselines using a non-recurrent, a recurrent, and a biologically plausible model. This work ultimately aims to underscore the need to incorporate time perception and temporal processing in artificial agents designed for human robot interaction and deployment in time dependent human societies.
Chinese Translation
本文提出了Chronocooked,一个用于研究强化学习(RL)智能体隐式时间间隔计时的基准测试套件。该套件受到《Overcooked》的启发,包含需要时间决策的烹饪场景。任务和奖励函数的设计使得时间信息虽然不可观察,但对最佳表现至关重要。环境故意保持简单,以便进行控制实验并支持生物学上合理的模型。评估指标旨在揭示RL智能体在计时能力方面的局限性,我们报告了使用非递归模型、递归模型和生物学上合理模型的基线结果。此项工作最终旨在强调在为人机交互和在时间依赖的人类社会中部署的人工智能体中,融入时间感知和时间处理的必要性。
cs.AI / 193 / 2608.16697

FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy

FabriMAE 我能相信自己吗?使用马尔可夫注意力熵进行自我评估的视觉-语言-行动生成
Aniri, Yilin, Chen, Bi, Jinhe, Guo, Junfei, Ran, Donglai, Bian, Xu, Jin, Zengjie, Wang, Yujun, Tian, Yijun, Tresp, Volker, Shen, Fei, Chua, Tat-Seng, Ma, Yunpu
Abstract
Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation into end-to-end policies across heterogeneous architectures. However, enabling VLAs to self-evaluate their action generation reliability without external supervision remains a major challenge. Existing methods either rely on expert annotations or estimate uncertainty only from output statistics, largely ignoring internal signals. In this work, we observe that internal visual modality entropy exhibits consistent distinctions between successful and failed tasks across heterogeneous VLAs. Although VLAs' architectures differ in their action generation, we show that they share a common latent action generation abstraction evolving under visual perception, language instruction, and state input, which we formulate as a Conditional Generative Markov Chain. Based on this formulation, we propose MAE (Markov Attention Entropy), a self-evaluation framework that directly converts internal attention signals into architecture-aware reliability scores, and introduce LIBERO-Reflect, a 4,000-episode benchmark combining 2,000 standard episodes and 2,000 challenging episodes across four subsets. Extensive experiments across heterogeneous VLA architectures and diverse scenarios show that MAE consistently outperforms state-of-the-art baselines on AUPR, AUROC, and FPR@95. We further instantiate FabriMAE for verifier-free test-time action selection, showing that MAE-guided multiple sampling improves PI-family robustness on LIBERO-Plus with small observed runtime overhead.
Chinese Translation
视觉-语言-行动模型(VLAs)将视觉感知、语言指令和行动生成整合到跨异构架构的端到端策略中。然而,使 VLAs 能够在没有外部监督的情况下自我评估其行动生成的可靠性仍然是一个重大挑战。现有方法要么依赖专家注释,要么仅从输出统计中估计不确定性,基本上忽略了内部信号。在本研究中,我们观察到内部视觉模态熵在异构 VLA 中成功与失败任务之间表现出一致的区别。尽管 VLAs 的架构在行动生成上有所不同,我们展示了它们共享一个在视觉感知、语言指令和状态输入下演变的共同潜在行动生成抽象,我们将其形式化为条件生成马尔可夫链(Conditional Generative Markov Chain)。基于这一形式化,我们提出了 MAE(马尔可夫注意力熵),一个自我评估框架,直接将内部注意力信号转换为架构感知的可靠性评分,并引入 LIBERO-Reflect,这是一个结合了 2,000 个标准剧集和 2,000 个挑战剧集的 4,000 集基准,跨越四个子集。在异构 VLA 架构和多样化场景下进行的广泛实验表明,MAE 在 AUPR、AUROC 和 FPR@95 上始终优于最先进的基线。我们进一步实例化 FabriMAE 以实现无验证者的测试时行动选择,显示 MAE 引导的多次采样在 LIBERO-Plus 上提高了 PI 家族的鲁棒性,且观察到的运行时开销较小。
cs.AI / 194 / 2608.16763

LAVA: Logic-Aware Validation and Augmentation Framework for Large-Scale Financial Document Auditing

LAVA:面向大规模金融文件审计的逻辑感知验证与增强框架
Shu, Ruoqi, Wang, Xuhui, Wang, Isaac, Mai, Yanming, Wan, Bo
Abstract
Financial document validation in production, such as payroll auditing, tax compliance, and loan underwriting, demands exceptional accuracy, consistency, and reproducibility under strict enterprise constraints. In practice, documents arrive with heterogeneous layouts and formats, semantically rich and context-dependent content, and embedded business rules that current pipelines struggle to process reliably. We introduce LAVA (Logic-Aware Validation and Augmentation), a modular, backbone-agnostic pipeline built on multimodal large language models, that integrates a four-stage design: document-rule retrieval, layout-preserving information extraction, auxiliary metadata enrichment, and auditable symbolic/arithmetic verification. LAVA supports robust rule grounding, fine-grained error attribution, and consistent, traceable end-to-end execution, capabilities essential for high-stakes deployment. Evaluated on a large real-world benchmark with diverse financial documents and dozens of expert-curated validation rules, LAVA outperforms baselines in hallucination control and edge-case handling while maintaining efficient token usage, demonstrating practicality for high-volume, time-critical validation.
Chinese Translation
在生产环境中,金融文件验证(例如工资审计、税务合规和贷款承保)要求在严格的企业约束下具备卓越的准确性、一致性和可重复性。实际上,文件以异构的布局和格式到达,内容语义丰富且依赖于上下文,并嵌入了当前管道难以可靠处理的业务规则。我们提出了LAVA(Logic-Aware Validation and Augmentation),这是一个基于多模态大型语言模型构建的模块化、与骨干无关的管道,集成了四个阶段的设计:文档规则检索、保留布局的信息提取、辅助元数据丰富和可审计的符号/算术验证。LAVA支持强大的规则基础、细粒度的错误归因以及一致、可追溯的端到端执行,这些能力对于高风险部署至关重要。在一个包含多样化金融文件和数十条专家策划验证规则的大型真实基准上进行评估,LAVA在幻觉控制和边缘案例处理方面优于基线,同时保持高效的令牌使用,展示了在高容量、时间敏感的验证中的实用性。
cs.AI / 195 / 2608.16776

GRIP: Grounded Reasoning via Information-Restricted Premises

GRIP:通过信息受限前提进行基础推理
Teng, Lirui
Abstract
High-capacity encoders in retrieval-augmented generation (RAG) can let the query dominate the latent state, leaving retrieved evidence functionally irrelevant. We call this failure mode query dominance. To address it, we introduce \textbf{GRIP} (Grounded Reasoning via Information-Restricted Premises), which imposes capacity asymmetry: the decoder keeps full-dimensional access to the query, while retrieved evidence passes through a severe stochastic bottleneck. This forces the evidence channel to encode only the residual information unavailable from the query. Across five reasoning benchmarks, GRIP outperforms strong iterative baselines, cuts a query--latent mutual-information diagnostic by roughly 30$\times$ (14.8 $\to$ 0.47 bits), and reduces hallucination by 73\%. Residual-alignment analysis further shows that the bottleneck output occupies subspaces less aligned with the query than baseline representations.
Chinese Translation
在检索增强生成(RAG)中,高容量编码器可能导致查询主导潜在状态,从而使检索到的证据在功能上变得无关。我们将这种失败模式称为查询主导。为了解决这个问题,我们引入了 extbf{GRIP}(通过信息受限前提进行基础推理),该方法施加了容量不对称性:解码器对查询保持全维访问,而检索到的证据则通过一个严重的随机瓶颈。这迫使证据通道仅编码查询中不可用的剩余信息。在五个推理基准测试中,GRIP的表现优于强大的迭代基线,将查询-潜在互信息诊断降低了大约30倍(14.8 $ o$ 0.47比特),并减少了73%的幻觉。剩余对齐分析进一步表明,瓶颈输出占据的子空间与查询的对齐程度低于基线表示。
cs.AI / 196 / 2608.16801

When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding

当智能体协调时:多智能体人工智能编码中的协调测量
Destefanis, Giuseppe, Aste, Tomaso
Abstract
We study how teams of AI coding agents coordinate while solving programming tasks. Current evaluations usually report whether the agents complete the task and how much the run costs, leaving the coordination inside the team largely unmeasured. We introduce an instrument to measure this coordination. Each run is represented as a temporal network in which agents and files are nodes, and messages, file writes, and file reads are timestamped directed edges with an associated cost. We apply this instrument to 1902 runs, each evaluated with a fixed test suite, across configurations that vary the team size, the team structure, and the file policy. The resulting networks show how coordination changes as teams grow and as the work changes. Direct messaging initially increases close to quadratically with the number of agents, with much of this growth coming from an early round of introductions. As the teams grow further, this increase levels off in the largest teams we study, where agents increasingly communicate through broadcast messages. The task also shapes the network that emerges. Work built around a shared specification produces dense, highly connected teams, while pipeline tasks produce sparse networks organised around local interfaces. Shared files can replace repeated 1-to-1 communication, cutting output tokens by about 42% at eight agents on message-heavy work, while adding overhead when files already carry the coordination. Naming one agent as coordinator creates no communication hub and provides no reliable improvement in success. We also observe an unprompted tendency for agents to seek out hidden grading material. We repeat the key experimental conditions in a sealed environment, replacing the hidden material with marked placeholder files. Across 244 additional runs, agents still reach for it in four fifths of runs, while the coordinator and file-channel findings reproduce.
Chinese Translation
我们研究了AI编码智能体团队在解决编程任务时的协调方式。目前的评估通常报告智能体是否完成任务以及运行成本,但团队内部的协调情况大多未被测量。我们引入了一种工具来测量这种协调。每次运行被表示为一个时间网络,其中智能体和文件是节点,消息、文件写入和文件读取是带有时间戳的有向边,并附带相关成本。我们将该工具应用于1902次运行,每次运行都使用固定的测试套件进行评估,涵盖不同的团队规模、团队结构和文件策略。结果网络显示,随着团队规模的增长和工作内容的变化,协调方式如何变化。直接消息传递最初随着智能体数量的增加接近二次增长,其中大部分增长来自早期的介绍环节。随着团队进一步扩大,这种增长在我们研究的最大团队中趋于平稳,智能体越来越多地通过广播消息进行沟通。任务本身也塑造了所产生的网络。围绕共享规范构建的工作产生了密集且高度连接的团队,而管道任务则产生了围绕局部接口组织的稀疏网络。共享文件可以替代重复的1对1通信,在消息密集的工作中,八个智能体的输出令牌减少了约42%,而当文件已经承担协调时则增加了开销。指定一个智能体作为协调者并未形成通信中心,也未能可靠地提高成功率。我们还观察到智能体无提示地倾向于寻找隐藏的评分材料。我们在一个封闭环境中重复了关键实验条件,将隐藏材料替换为标记的占位文件。在244次额外的运行中,智能体仍在五分之四的运行中寻求这些材料,同时协调者和文件通道的发现也得到了重现。
cs.AI / 197 / 2608.16804

Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment

基于领域适应的跨手语迁移学习与多尺度时间对齐
Artiaga, Keren, Li, Yang, Kuruoglu, Ercan Engin, Kin, Wai, Chan
Abstract
Sign language serves as a vital means of communication for individuals with hearing impairments, yet recognition resources for the over 100 distinct sign languages are severely lacking. In response, we present our work on sign language recognition using transfer learning and the domain adaptation method TA3N, which utilizes the Temporal Relational Network (TRN) module for aligning multi-scale temporal relations. Our findings highlight the superior performance of Domain Adaptation to neural network-based transfer learning, particularly in improving recognition of American Sign Language (ASL). Our research also identifies the effectiveness of aligning shorter-term temporal features between source and target domains. In addition to using RGB, we conducted experiments using Optical Flow mode for the sign language samples, ultimately determining that RGB outperforms Optical Flow in the majority of cases. Our work aims to improve accessibility and communication for individuals who rely on sign language as their primary mode of communication.
Chinese Translation
手语是听力障碍人士的重要沟通方式,但针对超过100种不同手语的识别资源严重不足。为此,我们提出了基于迁移学习和领域适应方法TA3N的手语识别工作,该方法利用时间关系网络(Temporal Relational Network, TRN)模块对多尺度时间关系进行对齐。我们的研究结果突显了领域适应在神经网络迁移学习中的优越表现,特别是在提高美国手语(American Sign Language, ASL)识别方面。我们的研究还表明,在源域和目标域之间对短期时间特征进行对齐的有效性。除了使用RGB,我们还对手语样本进行了光流(Optical Flow)模式的实验,最终确定在大多数情况下RGB的表现优于光流。我们的工作旨在改善依赖手语作为主要沟通方式的个体的可及性和交流能力。
cs.AI / 198 / 2608.16813

Quipu: A Governed Bitemporal Knowledge Graph Store

Quipu:一个受控的双时态知识图谱存储
Brown, Steve
Abstract
Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: accept writes now and clean later, keep one time axis or none, treat every writer's facts as equally trustworthy, and leave governance to dashboards and middleware. These four defaults are individually convenient and jointly untenable under agent workloads. We present Quipu, an embeddable store that inverts all four: no fact enters except through a gate whose predicates evaluate the pending post-state; data, trust labels, verdicts, and the rules themselves are bitemporal; named graphs are the unit of authority and trust, composed under a lattice whose one invariant is that composition never widens; and the governance specification $\Sigma$, the trace, and signed verdicts are facts in the store they govern, making the audit $T \models \Sigma$ a query. We evaluate with Census, a deterministic multi-writer lifecycle whose single seeded run scores every research question against planted ground truth: the gated store ends with 0 of 6 planted defects versus 6 of 6 ungated; all 7 composition probes uphold the lattice contract; 50 of 50 satisfied verdicts re-derive faithfully as of their instant while all 50 would be misreported under a latest-only rule set; and the SARC reference checker agrees with the in-store audit verdict-for-verdict, differing only on coverage semantics. A recorded trace from a governed writer surfaces a live enforcement gap the audit names with its remediation. On DEMM-Bench, an external decision-evidence sufficiency benchmark, a content-only reading of the exported records answers all 512 property-level governance questions correctly with zero overclaim under all eight degradation conditions, while container-presence baselines overclaim on up to 87.5% of them -- and the run surfaced, and led us to close, a gap in what a denial's verdict attests.
Chinese Translation
目前,代理程序能够编写知识图谱,但知识图谱存储仍然保留着人类策展时设定的默认值:现在接受写入,稍后清理,保留一个时间轴或没有时间轴,将每个作者的事实视为同等可信,并将治理留给仪表盘和中间件。这四个默认值在个别情况下是方便的,但在代理工作负载下是不可持续的。我们提出了Quipu,一个可嵌入的存储,它颠覆了这四个默认值:没有事实可以通过一个门进入,该门的谓词评估待处理的后状态;数据、信任标签、裁决和规则本身都是双时态的;命名图是权威和信任的单位,在一个格子下组合,其唯一的不变性是组合从不扩展;治理规范$ ext{Σ}$、追踪和签名裁决都是存储中治理的事实,使得审计$T \models ext{Σ}$成为一个查询。我们通过Census进行评估,这是一个确定性的多写入生命周期,其单一种子运行对每个研究问题进行评分,基于植入的真实值:有门的存储最终在6个植入缺陷中得分为0,而无门的存储得分为6;所有7个组合探测都遵守格子合同;50个满意的裁决在其瞬间忠实地重新推导,而在仅最新规则集下,所有50个将被错误报告;SARC参考检查器在存储审计的裁决上逐一一致,仅在覆盖语义上存在差异。来自受控写入者的记录追踪揭示了审计所命名的实时执行缺口及其补救措施。在DEMM-Bench,一个外部决策证据充分性基准中,导出记录的内容仅阅读正确回答了所有512个属性级治理问题,在所有八种降级条件下没有过度声明,而容器存在基线在多达87.5%的问题上过度声明——并且该运行揭示并促使我们关闭了否定裁决所证明的差距。
cs.AI / 199 / 2608.16831

Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning

带有人类反馈的策略迭代:将后训练强化学习引入上下文学习
Nguyen, Minh-Ha, Shyr, Cathy
Abstract
Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning showed that a fixed model could adapt its behavior from instructions and demonstrations. Policy Iteration with Human Feedback (PIHF) builds on this development and the recurrent evaluate-and-improve structure of generalized policy iteration. PIHF uses a pretrained language model as its execution substrate and moves persistent revision to a versioned natural-language policy and tool set. A language-model critic and clinical expert review complete-panel reasoning and tool-use trajectories to localize recurrent failures and form candidate revisions; the expert may reinterpret the evidence and retains authority over admission and rollback, while Recall@1 and Recall@5 validate outcomes after candidate execution. Across cumulative ablations and ultra-rare-disease benchmarks, a PIHF-derived policy improved Recall@1 in one proprietary executor and three open-weight executors spanning 3 to 49 billion active parameters. Gains were 32.7 percentage points for GPT-5.4 and 31.1 points for Qwen3.6-35B, a difference of 1.7 points. These results support the feasibility of using pretrained language models as fixed-weight execution substrates for expert-guided policy development in rare-disease diagnosis.
Chinese Translation
生成预训练建立了可重用的任务表示;后续关于基于语言的任务条件和上下文学习的研究表明,固定模型可以根据指令和示范调整其行为。带有人类反馈的策略迭代(Policy Iteration with Human Feedback, PIHF)基于这一发展以及广义策略迭代的递归评估与改进结构。PIHF使用预训练的语言模型作为其执行基础,并将持续修订转移到版本化的自然语言策略和工具集上。语言模型评论者和临床专家审查完整面板推理和工具使用轨迹,以定位递归失败并形成候选修订;专家可以重新解释证据,并对接受和回滚保留权威,同时 Recall@1 和 Recall@5 在候选执行后验证结果。在累积消融和超稀有疾病基准测试中,基于 PIHF 的策略在一个专有执行器和三个开放权重执行器中提高了 Recall@1,参数规模从 30 亿到 490 亿不等。GPT-5.4 的增益为 32.7 个百分点,Qwen3.6-35B 的增益为 31.1 个百分点,差异为 1.7 个百分点。这些结果支持使用预训练语言模型作为固定权重执行基础,以便在稀有疾病诊断中进行专家指导的策略开发的可行性。
cs.AI / 200 / 2608.16852

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

合规检测器读取什么?激活探针和守卫模型的审计
Sadhu, Saisab, Sengupta, Aadit, Sankarapu, Vinay Kumar, Seth, Pratinav
Abstract
Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy. Such monitoring is meaningful only if a detector's verdict depends on the stated rule rather than on surface features of the scenario. We show this condition fails across the current class of compliance detectors, a failure we call rule blindness. Deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged for every guard and activation probe we test, including a policy-conditioned guard that correctly cites the governing clause yet barely changes its verdict when that clause is swapped for its permissive counterpart. A purpose-built benchmark crossing two rules with two scenarios, so that neither alone predicts the label, confirms the failure under a design no prior benchmark rules out, and shows that step by step reasoning, not any fast detector we test, is what escapes it. Auditing at scale requires a retraining-free detector, so we introduce the Internal Compliance Score (ICS): a training-free activation readout calibrated from ten labelled pairs and scored by a single projection. We hold ICS to the same scrutiny as the guards it audits: a pre-registered criterion for beating trivial baselines is not met, and a bag-of-words model matches its pooled generalisation exactly. It remains useful because it is inexpensive, letting us audit four deployed guard models, an 8B zero-shot judge, and thirteen benchmarks, and it raises the mechanically verified pass rate when used to rank candidate responses, though an adaptive white-box attack removes this gain. We release the counterfactual protocol and crossed-rule benchmark so rule blindness can be tested in future probe and guard claims.
Chinese Translation
在部署的语言模型中,合规性监测越来越多地作为法律和审计控制实施,检查模型输出是否符合涵盖数据保护、医疗保健、金融监管和平台政策的书面规则。这样的监测只有在检测器的判决依赖于所述规则而非场景的表面特征时才有意义。我们展示了当前合规检测器的这一条件失效,这种失效我们称之为规则盲目性。删除、置换或替换主导规则对我们测试的每个守卫和激活探针的检测准确性没有影响,包括一个正确引用主导条款的政策条件守卫,但当该条款被其允许的对应条款替换时,其判决几乎没有变化。一个专门构建的基准测试交叉两个规则和两个场景,使得单独的任一规则都无法预测标签,确认了在一个先前基准未排除的设计下的失效,并显示逐步推理,而不是我们测试的任何快速检测器,才是逃避这一失效的关键。大规模审计需要一个无须再训练的检测器,因此我们引入了内部合规评分(Internal Compliance Score, ICS):一个从十对标记对中校准的无训练激活读出,并通过单一投影评分。我们对ICS进行与其审计的守卫相同的严格审查:未满足超越琐碎基线的预注册标准,且一个词袋模型完全匹配其汇总泛化。尽管如此,它仍然有用,因为它成本低廉,使我们能够审计四个部署的守卫模型、一个8B零样本评判者和十三个基准,并且在用于排名候选响应时提高了机械验证的通过率,尽管自适应白盒攻击消除了这一收益。我们发布了反事实协议和交叉规则基准,以便在未来的探针和守卫声明中测试规则盲目性。
计算语言学 (Computation and Language)
97
cs.CL / 1 / 2608.14551

Auxiliary uncertainty signals for LLM-assisted systematic review screening: a benchmark across eight Cohen drug-class reviews

辅助不确定性信号用于大型语言模型辅助的系统评价筛选:跨越八个Cohen药物类别评审的基准
Rahgozar, Arya, Mortezaagha, Pouria
Abstract
Large language models (LLMs) are increasingly used for title-abstract screening in systematic reviews, but their decisions lack calibrated uncertainty. We show that an auxiliary BERT+GCN classifier supplies a structured uncertainty signal that improves LLM screening efficiency, and we identify the prompt-delivery strategy that maximises the benefit-to-cost ratio. We evaluate five LLM prompt-delivery conditions on eight drug-class datasets from the Cohen (2006) benchmark using 3 seeds x 5-fold stratified cross-validation (600 fold-level results). A BERT+GCN model trained per fold classifies each test paper as INCLUDE, EXCLUDE, or MAYBE via two spectral tests (algebraic radical and categorical paradox). Conditions vary information content (none / label / full scores), selectivity (all papers vs. MAYBE only), and timing (proactive vs. reactive two-pass). A cross-model pilot against gpt-4.1-mini on three datasets tests cross-generation transfer. Three findings: (i) Full-context delivery yields significant gains in F1 (+0.011, paired Wilcoxon p=0.008) and WSS@95 (+0.050, p=0.039) at a 1.28x token-cost premium, while preserving recall. (ii) MAYBE-only routing is Pareto-optimal: highest mean recall (0.92) and AUC-ROC (0.54) at only 1.05x baseline cost -- one sixth of full-context overhead. (iii) The two-pass design escalates 22.2% +/- 8.8% of records yet never revises its decision (0% flip rate across all datasets and folds), giving decisive evidence that current instruction-tuned LLMs cannot self-triage. The cross-model pilot shows an identical +0.8% recall uplift for both LLM generations. A per-paper ablation across 20,796 observations shows the dual paradox test reduces empirically to a one-line logit-gap criterion. We release the full pipeline; the 600-run experiment replays in under one hour from cached LLM responses.
Chinese Translation
大型语言模型(LLMs)在系统评价中的标题-摘要筛选中越来越多地被使用,但它们的决策缺乏经过校准的不确定性。我们展示了一种辅助的BERT+GCN分类器提供了结构化的不确定性信号,从而提高了LLM筛选的效率,并确定了最大化收益与成本比的提示传递策略。我们在Cohen(2006)基准的八个药物类别数据集上评估了五种LLM提示传递条件,使用3个种子x 5折分层交叉验证(600个折级结果)。每个折训练的BERT+GCN模型通过两个谱测试(代数根和类别悖论)将每篇测试论文分类为包含(INCLUDE)、排除(EXCLUDE)或可能(MAYBE)。条件变化包括信息内容(无/标签/完整分数)、选择性(所有论文与仅MAYBE)和时机(主动与被动两次传递)。与gpt-4.1-mini的跨模型试点在三个数据集上测试跨代转移。三个发现:(i)完整上下文传递在F1(+0.011,配对Wilcoxon p=0.008)和WSS@95(+0.050,p=0.039)上取得显著提升,代价为1.28倍的令牌成本,同时保持了召回率。(ii)仅MAYBE路由是帕累托最优的:在仅1.05倍基线成本下获得最高的平均召回率(0.92)和AUC-ROC(0.54)——是完整上下文开销的六分之一。(iii)两次传递设计使22.2% +/- 8.8%的记录升级,但从未修正其决策(在所有数据集和折中翻转率为0%),提供了决定性证据表明当前的指令调优LLMs无法自我分类。跨模型试点显示两代LLM的召回率提升相同,均为+0.8%。在20,796个观察中的逐篇消融分析表明,双重悖论测试在经验上简化为一行对数差距标准。我们发布了完整的流程;600次实验在缓存的LLM响应下不到一小时内重放。
cs.CL / 2 / 2608.14577

HarmProfile: Characterizing Harmful Distributions in Frontier LLMs

HarmProfile:前沿大型语言模型中有害分布的特征化
Ma, Zhouyuan, Wu, Yutao, Huang, Hanxun, Zheng, Xiang, Liu, Xiao, Cao, Yixin, Wu, Zuxuan, Ma, Xingjun, Jiang, Yu-Gang
Abstract
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is known about the harmful outputs produced during model misbehavior, partly because large-scale, high-quality collections of frontier-LLM misbehavior are difficult to obtain. To address this gap, we introduce HarmProfile, a content-centric benchmark dataset that collects model misbehavior across diverse harm categories and model families, and defines the resulting harmful-output distribution as a model-level risk profile. The premise is that, just as linguistic behavior can be characterized from an utterance corpus, model risk can be characterized from the content, severity, and variation of its safety failures. HarmProfile contains over 80,000 validated artifacts from 23 frontier LLMs across 13 model families, organized into 15 harm categories and 57 subcategories. Using this corpus, we find that frontier LLMs reliably produce harmful content at scale, yet exhibit distinct risk profiles; both harmfulness and diversity grow with model capability, suggesting that frontier LLMs may appear safe yet harbor increasingly dangerous knowledge beneath the alignment surface. Our source code is available at https://github.com/fresh-ma/HarmProfile .
Chinese Translation
前沿大型语言模型(LLMs)的安全评估在很大程度上将有害生成视为攻击结果,而非分析对象。因此,对于模型失常期间产生的有害输出知之甚少,部分原因在于难以获得大规模、高质量的前沿 LLM 失常数据集。为了解决这一问题,我们引入了 HarmProfile,这是一个以内容为中心的基准数据集,收集了来自不同有害类别和模型家族的模型失常,并将由此产生的有害输出分布定义为模型级风险概况。其前提是,正如语言行为可以从发言语料库中进行特征化,模型风险也可以从其安全失效的内容、严重性和变异性中进行特征化。HarmProfile 包含来自 23 个前沿 LLM 和 13 个模型家族的超过 80,000 个经过验证的样本,组织成 15 个有害类别和 57 个子类别。利用这一语料库,我们发现前沿 LLM 在大规模上可靠地产生有害内容,但表现出不同的风险特征;有害性和多样性随着模型能力的提高而增长,这表明前沿 LLM 可能看似安全,但在对齐表面下潜藏着日益危险的知识。我们的源代码可在 https://github.com/fresh-ma/HarmProfile 获取。
cs.CL / 3 / 2608.14584

Multi-Modal Generative Fuzzy System: Fuzzy Inference Guided Large Model Interactive Question Answering Framework

多模态生成模糊系统:模糊推理引导的大型模型交互式问答框架
Yang, Hailong, Wang, Jianqi, Wang, Guanjin, Deng, Zhaohong
Abstract
In Multimodal Question Answering (MQA), models are required to jointly encode and integrate heterogeneous information from multiple modalities, including text, images, and speech, to perform complex semantic reasoning and decision making. Despite recent advances, existing approaches, including traditional deep learning models and Large Models (LMs) or prompt-based frameworks, continue to face several critical challenges. First, modality bias arises from discrepancies in feature distributions across different modalities, which limits effective cross modal collaborative understanding. Second, many questions require knowledge drawn from multiple domains, introducing significant uncertainty. Third, current methods often rely on shallow semantic matching, resulting in limited reasoning depth an reduced interpretability. To address these issues, inspired by the traditional fuzzy system (FS) framework, we propose a fuzzy-inference-guided multimodal generative architecture termed the Multi-Modal Generative Fuzzy System (MMGFS). The main contributions of MMGFS are two folds. First, it alleviates modality bias through a multimodal collaborative rumination mechanism. Second, it introduces fuzzy rules and a multi-hop inference mechanism to support cross-domain knowledge fusion and hierarchical reasoning, thereby strengthening uncertainty modelling and deepening semantic understanding. We conduct comprehensive evaluations on open-domain question answering datasets, including MultimodalQA and WebQA, as well as domain-specific benchmarks, including BioMol-VQA and EHRxQA. Experimental results demonstrate that MMGFS consistently outperforms existing methods across multiple datasets. It effectively mitigates modality bias and question uncertainty while achieving superior performance in answer accuracy, consistency, and generalization.
Chinese Translation
在多模态问答(MQA)中,模型需要联合编码和整合来自文本、图像和语音等多种模态的异构信息,以执行复杂的语义推理和决策。尽管近期取得了一些进展,但现有方法,包括传统深度学习模型和大型模型(LMs)或基于提示的框架,仍面临几个关键挑战。首先,模态偏差源于不同模态之间特征分布的差异,限制了有效的跨模态协同理解。其次,许多问题需要从多个领域提取知识,带来了显著的不确定性。第三,当前方法往往依赖于浅层语义匹配,导致推理深度有限且可解释性降低。为了解决这些问题,我们受到传统模糊系统(FS)框架的启发,提出了一种模糊推理引导的多模态生成架构,称为多模态生成模糊系统(MMGFS)。MMGFS的主要贡献有两个方面。首先,它通过多模态协同反思机制缓解了模态偏差。其次,它引入了模糊规则和多跳推理机制,以支持跨领域知识融合和分层推理,从而增强不确定性建模并加深语义理解。我们在开放领域问答数据集上进行了全面评估,包括MultimodalQA和WebQA,以及领域特定基准,包括BioMol-VQA和EHRxQA。实验结果表明,MMGFS在多个数据集上始终优于现有方法。它有效减轻了模态偏差和问题不确定性,同时在答案准确性、一致性和泛化能力方面表现出色。
cs.CL / 4 / 2608.14604

Wiola 13M, a Gated Spiral Attention Architecture for Parameter Efficient Small Language Models

Wiola 13M:一种用于参数高效小型语言模型的门控螺旋注意力架构
Chowdhury, Aryuemaan Kumar, Oosa, Praveen, Reddy, Vineesha
Abstract
Small language models in the ten to one hundred million parameter range are attractive for on device inference, rapid experimentation, and controlled scientific study, yet most of them reuse the standard transformer block without adaptation to the small scale regime. We present Wiola, a decoder only language model whose novelty is concentrated in three drop in components of every layer. First, Spiral Rotary Positional Encoding perturbs the standard rotary frequencies by a slowly growing per dimension factor so that phase trajectories fan outward, improving long range discrimination while adding no parameters. Second, Gated Spiral Attention introduces a per head, content adaptive scalar gate derived from a causal cumulative statistic of the query stream, providing an implicit and differentiable form of soft head selection at negligible cost. Third, the Butterfly feed forward block replaces the conventional expansion layer with a multiplicative interaction and an intra block bypass path, matching the parameter count of a four times gated linear unit block while improving gradient flow in shallow stacks. We formalize each component, derive exact parameter and computation budgets, and prove that the gated attention admits an exact and numerically verified equivalence between full sequence training and cached autoregressive decoding, so that no approximation is introduced at inference time. We also describe a fully reproducible training and evaluation protocol on a standard tiny story corpus. The reference implementation is released as an open source package with weights ready publishing support.
Chinese Translation
在一亿至一千万参数范围内的小型语言模型因其适用于设备推理、快速实验和受控科学研究而备受关注,然而大多数模型在小规模环境下仍然使用标准的变换器模块而未进行适应性调整。我们提出了Wiola,这是一种仅解码的语言模型,其新颖性集中在每层的三个可替换组件上。首先,螺旋旋转位置编码通过逐维增长的因子扰动标准旋转频率,使得相位轨迹向外扩展,从而改善长距离辨别能力,同时不增加任何参数。其次,门控螺旋注意力引入了一个基于查询流的因果累积统计的每头内容自适应标量门,提供了一种隐式且可微分的软头选择形式,成本极低。第三,蝴蝶前馈模块用乘法交互和内部模块旁路路径替代了传统的扩展层,匹配了四倍门控线性单元模块的参数数量,同时改善了浅层堆叠中的梯度流。我们对每个组件进行了形式化,推导了确切的参数和计算预算,并证明门控注意力在全序列训练和缓存自回归解码之间存在确切且经过数值验证的等价性,因此在推理时不会引入任何近似。我们还描述了一种在标准小故事语料库上完全可重复的训练和评估协议。参考实现作为开源包发布,支持权重的快速发布。
cs.CL / 5 / 2608.14621

AutoMem: A Text-Gradient Recursive Self-Improvement Framework for Automated Memory Architectures Search

AutoMem:一种用于自动化内存架构搜索的文本梯度递归自我改进框架
Du, Lin, Zhou, Jie, Cai, Yuxuan, Chen, Kai, Chen, Qin, Li, Xin, Zhang, Bo, Li, Wei, He, Liang
Abstract
Long-term memory is increasingly central to LLM agents, yet memory design remains a highly coupled architecture problem: what to encode, how to store it, how to retrieve it, and how to manage it can vary substantially across tasks and backbone models. We construct a discrete search space with 5 encoders, 5 stores, 6 retrievers, and 4 managers, and show that no single memory architecture consistently dominates: different tasks favor different module combinations, leading to substantial performance gaps. Motivated by this, we propose \textsc{AutoMem}, a text-gradient recursive self-improvement framework for task-adaptive memory architecture search. \textsc{AutoMem} optimizes over the factored space through two components: Experience-Guided Architecture Search, which proposes candidate architectures from historical search trajectories and accumulated reflections, and Failure-Guided Module Diagnosis, which localizes memory-related failures to specific modules and converts them into targeted textual feedback. Experiments on GAIA, WebWalkerQA, and xBench-DeepSearch across two LLM backbones show that \textsc{AutoMem} consistently discovers task-adaptive memory architectures that outperform the strongest human-designed memory baselines, improving accuracy by $2.8$ points on average across six benchmark-backbone settings. Further analysis shows that \textsc{AutoMem} achieves a favorable accuracy-efficiency trade-off, reducing token cost by $14.3\%$ over the strongest accuracy baselines under Qwen3.5-122B-A10B, while also finding stronger architectures than substantially larger random searches within only a few guided iterations.
Chinese Translation
长期记忆在大型语言模型(LLM)代理中越来越重要,但内存设计仍然是一个高度耦合的架构问题:编码什么、如何存储、如何检索以及如何管理在不同任务和基础模型之间可能有很大差异。我们构建了一个离散搜索空间,包括5个编码器、5个存储器、6个检索器和4个管理器,并展示了没有单一的内存架构能够始终占据优势:不同任务偏好不同的模块组合,导致显著的性能差距。基于此,我们提出了 extsc{AutoMem},一种用于任务自适应内存架构搜索的文本梯度递归自我改进框架。 extsc{AutoMem}通过两个组件在因子空间中进行优化:经验引导架构搜索(Experience-Guided Architecture Search),该组件从历史搜索轨迹和累积反思中提出候选架构;失败引导模块诊断(Failure-Guided Module Diagnosis),该组件将与内存相关的故障定位到特定模块,并将其转化为针对性的文本反馈。在GAIA、WebWalkerQA和xBench-DeepSearch的实验中,针对两个LLM基础模型的测试表明, extsc{AutoMem}始终发现任务自适应的内存架构,其性能超越了最强的人类设计内存基线,在六个基准-基础设置中平均提高了$2.8$个百分点的准确率。进一步分析表明, extsc{AutoMem}在准确性和效率之间实现了良好的权衡,在Qwen3.5-122B-A10B下,减少了$14.3\%$的令牌成本,相较于最强的准确性基线,同时在仅仅几次引导迭代中发现了比大规模随机搜索更强的架构。
cs.CL / 6 / 2608.14626

LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review

低资源语言中的大型语言模型安全对齐:系统文献综述
Lemofouet, Valdini Douglace, Uzor, Blessing Ngozi, Anyanwu, Paula Chikaodinaka, Kapsa, Danielle Blanche, Imam, Sukairaj Hafiz, Sahil, P Sam, Oppong, Abigail, Abdullahi, Tassallah, Siro, Clemencia, Abdulmumin, Idris, Yimam, Seid Muhie, Muhammad, Shamsuddeen Hassan
Abstract
Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages. In this paper, we conduct a Systematic Literature Review (SLR) of LLM safety alignment in low-resource languages by adopting the PRISMA 2020 methodology. Out of roughly 1,500 papers identified from Semantic Scholar, arXiv, and OpenAlex, 50 relevant studies have been selected and analyzed. Our review is organized around four themes: safety alignment methods, multilingual safety risks, evaluation benchmarks, and cross-lingual transferability. We further propose a taxonomy of safety alignment approaches based on three adaptation mechanisms: data adaptation, objective optimization, and mechanistic alignment. Across literature, translated English benchmarks fail to sufficiently represent culturally rooted harms, and multilingual models are more vulnerable to cross-lingual jailbreaks, code-switching attacks, and safety degradation in underrepresented languages. These failures are driven by several key factors, including uneven multilingual pre-training coverage, insufficient native-language preference data, poor transfer of safety representations, and a lack of culturally aware evaluation frameworks. The review also notes that many low-resource languages, especially African languages, have fewer safety benchmarks available than other multilingual regions. Overall, the results reveal a persistent multilingual safety gap, and suggest that future progress will require culturally grounded benchmarks, participatory data collection, balanced multilingual pre-training, and scalable multilingual alignment methods.
Chinese Translation
大型语言模型(LLMs)在安全对齐方面取得了显著进展,但在低资源和多语言环境中的安全保障仍显著弱于高资源语言。本文采用PRISMA 2020方法,对低资源语言中的LLM安全对齐进行系统文献综述(SLR)。从Semantic Scholar、arXiv和OpenAlex中识别出的约1500篇论文中,选择并分析了50项相关研究。我们的综述围绕四个主题组织:安全对齐方法、多语言安全风险、评估基准和跨语言可转移性。我们进一步提出了一种基于三种适应机制(数据适应、目标优化和机制对齐)的安全对齐方法分类法。文献中,翻译的英语基准未能充分代表文化根植的危害,而多语言模型更容易受到跨语言越狱、代码切换攻击和在代表性不足语言中的安全降级等影响。这些失败由几个关键因素驱动,包括不均匀的多语言预训练覆盖、缺乏母语偏好数据、安全表征的转移不足,以及缺乏文化意识的评估框架。综述还指出,许多低资源语言,尤其是非洲语言,拥有的安全基准少于其他多语言地区。总体而言,结果揭示了持续存在的多语言安全差距,并建议未来的进展需要文化根植的基准、参与性数据收集、平衡的多语言预训练和可扩展的多语言对齐方法。
cs.CL / 7 / 2608.14629

Inference-Time Mitigation of Adversarial Political Bias in Large Language Models

大语言模型中的对抗性政治偏见的推理时减缓
Panchagnula, Tejaswi V., Coburn, Bruce, Dietrich, Bryce J., Browning, Robert X., Delp, Edward J., Zhu, Fengqing
Abstract
As Large Language Models (LLMs) become the mainstay for information retrieval and summarization tasks, ensuring that they are always non-partisan and invulnerable to political bias is a critical step towards safer and more trustworthy Artificial Intelligence (AI). Current model alignment paradigms, such as reinforcement learning from human feedback (RLHF), make LLMs follow overarching safety instructions. However, this instruction tuning can be exploited via adversarial prompt injection and be used to generate unsafe content. In particular, political bias has not been specifically targeted by modern alignment techniques as harmful and biased content. To address this vulnerability of LLMs, we propose mitigation strategies using Chain of Thought (CoT) prompting and Direct Preference Optimization (DPO). Using a public dataset of legislative videos, we generate summaries using LLMs, inject bias via adversarial prompting and evaluate their performance on a four axis scale designed for political summarization. In this paper, we present different methods to shield LLMs against the injection of political bias. Our results demonstrate that the proposed Recursive Self-Correction approach raises model performance from a Political Neutrality Likert scale baseline of 2.14 to 4.56, averaged across all models, demonstrating effective inference-time mitigation of political bias in LLM-generated summaries.
Chinese Translation
随着大语言模型(LLMs)成为信息检索和摘要任务的主流,确保它们始终保持非党派立场并抵御政治偏见是迈向更安全、更可信的人工智能(AI)的关键一步。目前的模型对齐范式,如基于人类反馈的强化学习(RLHF),使得LLMs遵循总体安全指令。然而,这种指令调优可能被对抗性提示注入所利用,从而生成不安全的内容。特别是,现代对齐技术并未专门针对政治偏见这一有害和偏见内容。为了解决LLMs的这一脆弱性,我们提出了使用思维链(Chain of Thought, CoT)提示和直接偏好优化(Direct Preference Optimization, DPO)的减缓策略。利用公共立法视频数据集,我们使用LLMs生成摘要,通过对抗性提示注入偏见,并在为政治摘要设计的四个维度上评估其性能。在本文中,我们提出了不同的方法来保护LLMs免受政治偏见的注入。我们的结果表明,所提出的递归自我修正方法将模型性能从政治中立李克特量表基线的2.14提高到4.56,平均跨所有模型,展示了在LLM生成摘要中有效的推理时政治偏见减缓。
cs.CL / 8 / 2608.14630

Characterizing Rhetorical Misalignment in Decision-Making with Language Models

决策中语言模型的修辞不对齐特征研究
Cheng, Zirui, Chan, Joey, Du, Simo, Tan, Chenhao, Guo, Yue, Peng, Hao
Abstract
Human decision-making is often shaped by a range of well-documented cognitive biases. As large language models (LLMs) become increasingly integrated into high-stakes human-AI decision-making, it is important to understand whether their outputs can amplify potential biases, how this influences human decisions, and crucially, whether it can lead to harmful consequences. In this work, we develop a decision-theoretic framework to study rhetorical misalignment, a failure mode where an LLM uses rhetorically inappropriate forms of presentation for a given decision context, thereby inducing suboptimal human decisions. We empirically investigate this phenomenon through a human-subject experiment in realistic clinical decision-making using a dataset curated from the United States Medical Licensing Examination. By measuring how LLM-generated information affects decisions, we observe that LLMs induce an average 2.81% rate of harmful decision flips across different models, where clinician participants change from a correct to an incorrect answer. Rationales reported by participants provide evidence that these revisions are closely related to the language used by LLMs that may induce different types of cognitive biases, including anchoring, authority bias, and loss aversion. To enable scalable evaluation, we instantiate our theoretical framework using decision-makers simulated by LLMs to computationally measure rhetorical misalignment. Our findings reveal a safety concern previously unrecognized in high-stakes domains: a model can be factually aligned yet still induce harm through its rhetorical presentation.
Chinese Translation
人类决策往往受到多种已被充分记录的认知偏见的影响。随着大型语言模型(LLMs)在高风险人机决策中越来越多地被整合,理解它们的输出是否会放大潜在偏见、如何影响人类决策,以及是否会导致有害后果变得尤为重要。在本研究中,我们开发了一个决策理论框架来研究修辞不对齐,这是一种失效模式,其中LLM在特定决策情境中使用修辞上不恰当的呈现形式,从而导致次优的人类决策。我们通过一项涉及人类受试者的实验,实证研究这一现象,实验基于从美国医学执照考试中整理的数据集,模拟现实的临床决策过程。通过测量LLM生成的信息如何影响决策,我们观察到不同模型的LLM平均导致2.81%的有害决策翻转,其中临床参与者从正确答案转变为错误答案。参与者报告的理由提供了证据,表明这些修正与LLM使用的语言密切相关,可能引发不同类型的认知偏见,包括锚定效应、权威偏见和损失厌恶。为了实现可扩展的评估,我们利用LLM模拟决策者实例化我们的理论框架,以计算方式测量修辞不对齐。我们的研究结果揭示了在高风险领域中之前未被认识的安全隐患:一个模型可以在事实上一致,但仍然通过其修辞呈现引发危害。
cs.CL / 9 / 2608.14632

DeMTS: Denoising Trajectories as Multivariate Time Series for Hallucination Detection in Diffusion Language Models

DeMTS:将去噪轨迹视为多元时间序列以检测扩散语言模型中的幻觉
Zhang, Xin, Wang, Yili, Tan, Yue, He, Xin, Qian, Yanyu, Liu, Yixin, Chang, Yi, Pan, Shirui, Wang, Xin
Abstract
Diffusion large language models (D-LLMs) have emerged as a promising paradigm for text generation. However, similar to autoregressive LLMs, D-LLMs remain vulnerable to hallucinations, where fluent outputs may contain factually incorrect or unsupported content. Although existing hallucination detection methods for D-LLMs attempt to leverage uncertainty trajectories of the denoising process to better identify hallucination signals, they typically compress the trajectories along either the temporal or token dimension, overlooking the useful information encoded in the complete two-dimensional token-step structure. Consequently, they may fail to capture hallucination-relevant patterns, such as inconsistent convergence and cross-token fault propagation, leading to suboptimal detection performance. To bridge this gap, we propose a D-LLM hallucination detection framework that formulates the Denoising trajectories as Multivariate Time Series over learnable latent variables (DeMTS for short). DeMTS employs a trajectory-preserving token-to-variable assignment module to convert token signals into stable latent variables. Based on these variables, we propose dynamic multivariate temporal modeling to progressively integrate inter-variable dependency modeling with temporal encoding for hallucination prediction. Extensive experiments on two D-LLMs backbones and three benchmarks demonstrate that DeMTS outperforms existing hallucination detection methods while maintaining strong robustness, efficiency, and cross-task transferability.
Chinese Translation
扩散大型语言模型(D-LLMs)已成为文本生成的有前景的范式。然而,与自回归大型语言模型类似,D-LLMs仍然容易受到幻觉的影响,即流畅的输出可能包含事实不准确或缺乏支持的内容。尽管现有的D-LLMs幻觉检测方法试图利用去噪过程的不确定性轨迹来更好地识别幻觉信号,但它们通常在时间或标记维度上压缩轨迹,忽视了完整的二维标记-步骤结构中编码的有用信息。因此,它们可能无法捕捉与幻觉相关的模式,例如不一致的收敛和跨标记的故障传播,从而导致检测性能不佳。为了解决这一问题,我们提出了一种D-LLM幻觉检测框架,将去噪轨迹形式化为可学习潜变量上的多元时间序列(简称DeMTS)。DeMTS采用轨迹保留的标记到变量分配模块,将标记信号转换为稳定的潜变量。基于这些变量,我们提出了动态多元时间建模,逐步将变量间依赖建模与时间编码结合,以进行幻觉预测。在两个D-LLM骨干网络和三个基准上的广泛实验表明,DeMTS在保持强大鲁棒性、效率和跨任务可转移性的同时,超越了现有的幻觉检测方法。
cs.CL / 10 / 2608.14681

Automatic or Controlled? Repetition Priming Reveals Divergent Processing in Base LLMs, Instruct LLMs, and Humans

自动还是受控?重复启动揭示基础大型语言模型、指令大型语言模型与人类的不同处理方式
Ren, Jinglei, Wang, Yuyue
Abstract
Words recur constantly in natural language use, yet it remains unclear whether language models reactivate prior representations or re-evaluate repeated words afresh, and whether post-training changes this default behavior. We apply repetition priming (Shiffrin and Schneider, 1977) to 15 models across five model families (1.5B-14B parameters) in two tasks, semantic categorization and cloze completion, with matched human experiments using identical stimuli. We find that base models exhibit automatic processing: they show immediate facilitation that remains stable across lags, partially survives context removal, and correlates with attention to prior occurrences. Instruct models exhibit controlled processing: their facilitation decays with lag, collapses without expected context, and reverses to interference at larger scales. Within the Qwen 2.5 family, this dissociation increases monotonically with model scale, suggesting that post-training progressively alters repetition processing. Humans show a hybrid profile, with lag-sensitive facilitation resembling instruct models but without interference, suggesting that neither model type fully captures human cognition. Our findings reveal a qualitative shift in how language models process repeated information after post-training and provide mechanistic evidence for the divergence between model behaviors.
Chinese Translation
在自然语言使用中,词汇不断重复,但尚不清楚语言模型是重新激活先前的表征,还是重新评估重复的词汇,以及后训练是否改变了这种默认行为。我们将重复启动(Repetition Priming,Shiffrin 和 Schneider,1977)应用于15个模型,涵盖五个模型家族(参数量从1.5B到14B),在语义分类和完形填空两个任务中进行实验,同时进行匹配的人类实验,使用相同的刺激。我们发现基础模型表现出自动处理:它们显示出立即的促进效应,该效应在延迟期间保持稳定,部分在去除上下文后仍然存在,并与对先前出现的注意力相关。指令模型则表现出受控处理:它们的促进效应随着延迟而减弱,在缺乏预期上下文时崩溃,并在更大范围内反转为干扰。在Qwen 2.5家族中,这种分离随着模型规模的增加而单调增加,暗示后训练逐渐改变了重复处理。人类则表现出混合特征,其延迟敏感的促进效应类似于指令模型,但没有干扰,表明这两种模型类型都未能完全捕捉人类认知。我们的研究结果揭示了语言模型在后训练后处理重复信息的定性变化,并提供了模型行为之间差异的机制证据。
cs.CL / 11 / 2608.14693

Domain Agnostic Text Redaction from Natural Language Rules using Instruction Tuning

基于指令调优的领域无关文本编辑自然语言规则
Arunagiri, Aravindhan, Khan, Ayaan, Avadhanam, Udayaadithya, Sundar, SaiBarath
Abstract
With the increasing digitization of personal and corporate communication, the automatic sanitization of textual data has become a crucial component of data privacy and compliance frameworks. Traditional text sanitization solutions are majorly suitable for obscuring sensitive data with standard structure such as Personal Identifiable Information (PII). These solutions do not provide transparent justification for their redaction, which makes it difficult to audit them. This paper introduces an explainable, domain-agnostic text redaction solution that uses natural language rules of redaction, applied via an instruction-tuned language model, to identify and redact sensitive information in unstructured documents. Unlike traditional text sanitization, this method enables a user to conveniently define any sensitive information; which may be structured (e.g.\ PII) or unstructured (e.g.\ legal terms and conditions) in natural language. A general-purpose LLM generates or augments these natural language rules of redaction from the user's definition, which are then used to instruction-fine-tune a smaller language model that reasons the rules step-by-step over any given document to identify and redact the corresponding sensitive content, while providing transparent justifications for each redaction and highlighting the specific rule that triggered the decision. This explanation is generated in natural language to support human reviewers and auditors in understanding why specific content was redacted. A reconstruction-based metric is used to estimate the probability of recovering redacted information from the sanitized document, quantifying redaction coverage. The solution shows high reconstruction error and high redaction precision, making it suitable for automated text sanitization in critical applications such as legal discovery, medical documentation, and corporate information governance.
Chinese Translation
随着个人和企业通信的数字化程度不断提高,文本数据的自动清理已成为数据隐私和合规框架中的关键组成部分。传统的文本清理解决方案主要适用于遮蔽具有标准结构的敏感数据,如个人可识别信息(PII)。这些解决方案未能提供透明的编辑理由,这使得审计变得困难。本文提出了一种可解释的、领域无关的文本编辑解决方案,该方案利用自然语言编辑规则,通过指令调优的语言模型,识别并编辑非结构化文档中的敏感信息。与传统文本清理方法不同,该方法使用户能够方便地定义任何敏感信息;这些信息可以是结构化的(例如,PII)或非结构化的(例如,法律条款和条件)。通用的语言模型(LLM)根据用户的定义生成或增强这些自然语言编辑规则,然后用于对一个较小的语言模型进行指令微调,该模型逐步推理这些规则,以识别和编辑相应的敏感内容,同时为每个编辑提供透明的理由,并突出触发决策的具体规则。该解释以自然语言生成,以支持人工审查员和审计员理解为何特定内容被编辑。使用基于重建的度量来估计从清理后的文档中恢复编辑信息的概率,从而量化编辑覆盖率。该解决方案显示出高重建误差和高编辑精度,使其适用于法律发现、医疗文档和企业信息治理等关键应用中的自动文本清理。
cs.CL / 12 / 2608.14712

Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data

你的注意力度量回答了哪个问题?将注意力行视为组合数据
Papamichalis, Marios, Ruane, Regina
Abstract
Each row of a transformer's attention matrix is a probability distribution over tokens, and in trained models most of that probability lands on a single \emph{sink} token, usually the first. Standard tools for comparing attention rows (cosine similarity, Jensen--Shannon divergence, Shannon entropy) therefore hinge on a choice papers rarely report: keep the sink, or drop it and renormalize. This choice can reverse conclusions. On ten pretrained models from five families, 17--47% of verdicts about which of two heads is more similar flip with the convention, and the most prominent structure in a standard BERT head-clustering pipeline is an artifact of it. The reason is that one-number summaries mix two questions: how much attention the sink takes, and how the rest is divided among the content tokens. Treating rows as compositional data separates them exactly: the Aitchison distance splits orthogonally into a sink term and a content term, entropy splits by an exact identity, and the content distance is characterized by invariances the transformer itself possesses. The separation matters in practice: most measured entropy collapse during training is the sink growing, not attention sharpening (30% of the drop at 70M parameters, 95% at 1B, 79% at 1.4B), and pruning heads with the wrong channel can inflate perplexity more than a hundredfold. We map where each convention is safe, test a frozen out-of-sample predictor (one confirmation, one abstention, one failure), and release code regenerating every number.
Chinese Translation
变换器的注意力矩阵的每一行都是对标记的概率分布,在训练好的模型中,大部分概率集中在一个单一的 extit{汇聚}标记上,通常是第一个。比较注意力行的标准工具(余弦相似度、詹森-香农散度、香农熵)因此依赖于一个论文中很少报告的选择:保留汇聚标记,还是丢弃并重新归一化。这个选择可能会颠覆结论。在来自五个家族的十个预训练模型中,关于两个头哪个更相似的判断中,有17%到47%的结果会因这一约定而翻转,而标准BERT头聚类管道中最显著的结构是这一约定的伪影。原因在于,一维总结混合了两个问题:汇聚标记占据了多少注意力,以及其余的注意力在内容标记之间如何分配。将行视为组合数据可以准确地将它们分开:Aitchison距离正交地分为汇聚项和内容项,熵通过一个精确的恒等式分裂,而内容距离的特征由变换器本身所具有的不变性来表征。这种分离在实践中是重要的:在训练过程中,大多数测量到的熵的崩溃是汇聚标记的增长,而不是注意力的锐化(在7000万个参数时下降的30%,在10亿时95%,在14亿时79%),而错误通道的头部修剪可能会使困惑度膨胀超过一百倍。我们映射了每个约定安全的地方,测试了一个冻结的样本外预测器(一个确认,一个弃权,一个失败),并发布了重新生成每个数字的代码。
cs.CL / 13 / 2608.14737

Class Imbalance and Batch Effects in LLM-Based Screening for Systematic Reviews

基于大型语言模型的系统评价筛选中的类别不平衡与批处理效应
Hida, Gilberto Sussumu, Ribeiro, Danilo Monteiro, Hida, Clayton Suguio
Abstract
This study analyses LLMs in imbalanced binary classification, using study screening in systematic reviews as the application domain. An experiment was conducted in five reviews, comparing individual and batch processing, with and without prevalence metadata. The results indicate a limited influence of the prevalence metadata, with no evidence that it improves performance. In contrast, batch processing produced larger behavioral changes that varied according to the prevalence of the class. The aggregate and item-level analyses did not always coincide. Therefore, batch processing should be evaluated not only in terms of cost, but also in relation to its effects on decision-making behavior.
Chinese Translation
本研究分析了在不平衡二分类中的大型语言模型(LLMs),以系统评价中的研究筛选作为应用领域。我们在五个系统评价中进行了实验,比较了个体处理与批处理,以及是否使用流行度元数据。结果表明,流行度元数据的影响有限,没有证据表明它能改善性能。相反,批处理产生了更大的行为变化,这些变化根据类别的流行度而有所不同。总体分析与项目级分析并不总是一致。因此,批处理的评估不仅应考虑成本,还应与其对决策行为的影响相关联。
cs.CL / 14 / 2608.14792

Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters

仅依靠提示不足:在儿科接触中使用大型语言模型测量共享决策的监督基线和泄漏控制
Modenesi, Bernardo, Lin, Jody, Kaphingst, Kimberly, Zhu, Angela, Wheeler, Maya, Zhang, Peilu, Fagerlin, Angela
Abstract
Objectives: To determine whether zero-shot prompting of a large language model (LLM) is sufficient to detect shared decision-making (SDM) behaviors in real clinical encounters, and whether supervised learning adds value under patient-grouped, nested evaluation. Methods: We analyzed 21 audio-recorded outpatient surgical decision encounters (19 unique patients; 7,566 utterance segments; ~6.1 hours) between families of children with multiple long-term conditions and their surgical providers. Trained coders labeled segments for 12 SDM behaviors (human-human macro Cohen's kappa = 0.695). We compared a zero-shot local LLM (Qwen 2.5 32B), a supervised classifier over frozen sentence embeddings, and their logistic stack, under patient-grouped outer folds with inner cross-fitted thresholds and patient-resampled confidence intervals. Results: The zero-shot LLM reached macro kappa = 0.139 (95% CI 0.111-0.164). The supervised classifier reached kappa = 0.227 (0.186-0.262), a paired improvement of 0.088 (0.051-0.119). A logistic stack of the two reached kappa = 0.242 (0.198-0.284). We identified multiple corpus-specific leakage paths, including grouping sibling recordings separately and allowing labels from an outer held-out patient to enter few-shot exemplars used while fitting downstream models. Conclusion: Zero-shot prompting alone is not sufficient to measure SDM behavior as reliably as a small supervised model, and patient-level grouping alone does not prevent leakage when labeled prompt exemplars are precomputed outside the outer evaluation loop. Reported performance is sensitive to the unit of data splitting and to where labeled exemplars enter the pipeline. External validation is needed before these findings generalize beyond this population, model, prompt, and codebook.
Chinese Translation
目的:确定零-shot 提示大型语言模型(LLM)是否足以在真实临床接触中检测共享决策(SDM)行为,以及在患者分组的嵌套评估下监督学习是否增值。方法:我们分析了21个音频记录的门诊外科决策接触(19名独特患者;7,566个发言片段;约6.1小时),涉及多种长期疾病儿童的家庭与其外科提供者之间的互动。经过培训的编码员对12种SDM行为的片段进行了标记(人际间宏观Cohen's kappa = 0.695)。我们比较了零-shot本地LLM(Qwen 2.5 32B)、一个基于冻结句子嵌入的监督分类器,以及它们的逻辑堆叠,采用患者分组的外部折叠和内部交叉适配阈值及患者重抽样的置信区间。结果:零-shot LLM的宏观kappa为0.139(95% CI 0.111-0.164)。监督分类器的kappa为0.227(0.186-0.262),配对改进为0.088(0.051-0.119)。两者的逻辑堆叠达到了kappa = 0.242(0.198-0.284)。我们识别了多个特定语料库的泄漏路径,包括将兄弟姐妹的录音单独分组,并允许来自外部保留患者的标签进入在拟合下游模型时使用的少量示例。结论:仅依靠零-shot 提示无法像小型监督模型那样可靠地测量SDM行为,而仅依靠患者级分组并不能防止在外部评估循环外预计算的标记提示示例中发生泄漏。报告的性能对数据拆分的单位及标记示例进入管道的位置敏感。在这些发现推广到该人群、模型、提示和代码本之外之前,需要进行外部验证。
cs.CL / 15 / 2608.14797

Beyond Tokens: A Survey on Decoding Methods for Large Language and Vision-Language Models

超越标记:大型语言模型和视觉-语言模型解码方法的调查
Wang, Haoran, Xu, Xiongxiao, Yu, Philip S., Shu, Kai
Abstract
Large language models (LLMs) and large vision-language models (LVLMs) have demonstrated impressive generative capabilities, yet ensuring their outputs align with user intent is still challenging. While most existing approaches address this issue at the training stage, inference-time approaches like decoding methods offer a more efficient and scalable solution. Decoding methods control model generation by guiding token-level selection, performing sequence-level generation, or generating tokens in parallel to accelerate the process. In this survey, we identify three emerging paradigms from recent works on decoding methods for LLMs and LVLMs, provide a systematic review of these methods, highlight ongoing challenges, and discuss potential future research directions. Our goal is to underscore the efficiency and effectiveness of decoding methods and offer a practical view of their applications. Paper lists and more resources on decoding methods for LLMs and LVLMs can be found at https://github.com/wang2226/Awesome-LLM-Decoding.
Chinese Translation
大型语言模型(LLMs)和大型视觉-语言模型(LVLMs)展现了令人印象深刻的生成能力,但确保其输出与用户意图一致仍然具有挑战性。虽然大多数现有方法在训练阶段解决了这一问题,但推理时的方法,如解码方法,提供了更高效和可扩展的解决方案。解码方法通过引导标记级选择、执行序列级生成或并行生成标记来控制模型生成,从而加速过程。在本次调查中,我们从近期关于LLMs和LVLMs的解码方法的研究中识别出三种新兴范式,系统回顾这些方法,强调当前面临的挑战,并讨论潜在的未来研究方向。我们的目标是强调解码方法的效率和有效性,并提供其应用的实际视角。有关LLMs和LVLMs解码方法的论文列表和更多资源可在 https://github.com/wang2226/Awesome-LLM-Decoding 找到。
cs.CL / 16 / 2608.14813

Beyond the pale: Assessing prevalence and contents of extremist speech in LLM training data

超越边界:评估极端主义言论在大型语言模型训练数据中的普遍性及内容
Nikolaev, Dmitry, Mattheis, Ashley A.
Abstract
Despite a strong interest on the part of the research community in the topic of trustworthy and safe AI, the composition of the text corpora that large language models (LLMs) encounter in pre- and post-training has not yet drawn much attention. In this work, we address the question of whether LLMs are exposed to unfiltered, uncontextualised extremist speech. Using several definitions of extremist speech, stemming from official documents and research literature, and an extraction pipeline combining automated text processing with expert verification, we provide a lower bound on the prevalence of extremist documents in Dolma, an open training corpus underpinning the OLMo series of models. We show that Dolma is likely to include hundreds of thousands of documents containing extremist content and hate speech of several types, including direct calls for violence, and discuss the implications of this for data curation and model pre-training.
Chinese Translation
尽管研究界对可信和安全人工智能主题表现出强烈兴趣,但大型语言模型(LLMs)在训练前和训练后所接触的文本语料库的组成尚未引起足够关注。在本研究中,我们探讨了LLMs是否暴露于未经过滤和缺乏上下文的极端主义言论。我们使用多个来源于官方文件和研究文献的极端主义言论定义,并结合自动文本处理与专家验证的提取流程,提供了Dolma这一开放训练语料库中极端主义文档普遍性的下限,Dolma是支撑OLMo系列模型的基础语料库。我们表明,Dolma可能包含数十万份包含极端主义内容和多种类型仇恨言论的文档,包括直接呼吁暴力的内容,并讨论这对数据管理和模型预训练的影响。
cs.CL / 17 / 2608.14843

Writing Style Similarity Reflects Academic Genealogy

写作风格相似性反映学术谱系
Manzo, Cameron
Abstract
As authorship attribution systems are increasingly deployed to detect ghostwritten and AI-generated papers, their errors can support accusations against legitimate authors. These systems assume each author's style is their own. Researchers, however, study under advisors, and inherit their stylistic quirks. We build a corpus of arXiv authors with $\geq 2$ solo papers from the Mathematics Genealogy Project graph, giving $5{,}803$ total authors and $2{,}501$ ground-truth advisor-student pairings. Using embeddings from a fine-tuned model, advisors sit $39.9\%$ closer in cosine distance to their students than a random same-field author does. Two open encoders reproduce the effect at $12.6\%$ and $14.5\%$. \emph{Academic siblings}, two students of one advisor who may never have met, sit $30.4\%$ closer across $8{,}360$ pairs, even when they studied at different institutions. Pairs who share only an institution and a field show negligible similarity. Given a closed-set attribution task over the same corpus, the system's errors occur on the true author's advisors and academic siblings $11$ times more often than chance.
Chinese Translation
随着作者归属系统越来越多地被用于检测代写和人工智能生成的论文,这些系统的错误可能会支持对合法作者的指控。这些系统假设每位作者的风格都是独特的。然而,研究人员在导师的指导下进行研究,并继承了他们的风格特征。我们构建了一个来自数学谱系项目(Mathematics Genealogy Project)图谱的 arXiv 作者语料库,其中包含至少 2 篇独立论文,共有 5,803 位作者和 2,501 对真实的导师-学生配对。使用经过微调模型的嵌入,导师与其学生的余弦距离比随机同领域作者更近 39.9%。两个开放编码器分别以 12.6% 和 14.5% 的效果再现了这一现象。被称为“学术兄弟”(academic siblings)的两位来自同一导师的学生,即使从未见过面,在 8,360 对配对中也相距更近 30.4%,即使他们在不同的机构学习。仅共享机构和领域的配对显示出微不足道的相似性。在对同一语料库进行封闭集归属任务时,系统的错误发生在真实作者的导师和学术兄弟身上的频率是偶然情况的 11 倍。
cs.CL / 18 / 2608.14855

What to Forget in Unlearning? Forget Set Curation for Language Models

在遗忘中遗忘什么?语言模型的遗忘集策划
Jha, Animesh, Khatua, Arpandeep, Allouah, Youssef, Koyejo, Sanmi
Abstract
Machine unlearning aims to remove targeted data or behaviors from a trained model without retraining from scratch. Yet most evaluations assume that the examples to forget are already known. In realistic language-model deployments, a requester may ask a model to stop reproducing a song or book without knowing which spans, documents, quotations, or near-duplicates in a trillion-token corpus support that behavior. We study this missing upstream problem, forget set curation: mapping a suppression request to the data passed to an unlearning algorithm. We introduce CleanSlate, a benchmark for verbatim output suppression over songs and books, with model-specific extraction profiles, content-grounded QA, and capability-retention evaluations. CleanSlate exposes two failure modes. Natural lexical and exact-substring curators often yield forget sets that lead to weak suppression. An evaluation-aware curator suppresses requested continuations almost completely, but causes collateral regression on non-requested content and model-dependent capability loss. These results show that practical unlearning is not only an optimization problem once a forget set is given: the data chosen for forgetting determines both what can be unlearnt and what else is damaged.
Chinese Translation
机器遗忘旨在从训练好的模型中去除特定数据或行为,而无需从头开始重新训练。然而,大多数评估假设要遗忘的示例已经是已知的。在现实的语言模型部署中,请求者可能会要求模型停止再现某首歌曲或某本书,而不知道在万亿标记的语料库中哪些片段、文档、引用或近似重复支持该行为。我们研究了这一缺失的上游问题——遗忘集策划:将抑制请求映射到传递给遗忘算法的数据。我们引入了CleanSlate,这是一个针对歌曲和书籍逐字输出抑制的基准,具有特定于模型的提取配置、基于内容的问答和能力保留评估。CleanSlate揭示了两种失败模式。自然词汇和精确子字符串策划者通常会产生导致抑制效果较弱的遗忘集。一种关注评估的策划者几乎完全抑制了请求的延续,但对非请求内容造成了附带的退化和模型依赖的能力损失。这些结果表明,实际的遗忘不仅仅是在给定遗忘集后的优化问题:选择遗忘的数据决定了可以被遗忘的内容以及其他内容的损害。
cs.CL / 19 / 2608.14886

Where Does Retrieval Fail? Evaluating RAG Architectures for Agricultural Advisory

检索失败的原因何在?评估农业咨询中的 RAG 架构
Reza, Khan Raiyan Ibne, Maria, Sanjana Aktar, Nimi, Sumaiya Tabassum
Abstract
Retrieval quality in RAG systems is commonly reported as a single aggregate score, which can hide large differences across query types and language conditions. We study this problem in Bengali agricultural advisory, where farmer queries are often colloquial while official advisory documents use formal scientific terminology. We construct a test collection of 1,000 queries and 2,882 knowledge nodes extracted from 284 official Bangladeshi agricultural publications, and use it to evaluate five retrieval architectures and six embedding models under three controlled language conditions. The results show that no single retrieval method is consistently best. For native Bengali queries, BM25 is the strongest single retriever (R@10 = 0.506) while Hybrid RRF reaches the highest overall R@10 of 0.539. However, dense retrieval performance varies sharply by query type: R@10 is 0.093 on colloquial farmer queries and 0.970 on formal safety queries. Across language conditions, BM25 R@10 drops from 0.506 on Bengali queries to 0.004 when English queries are matched against the Bengali corpus, while dense retrieval falls only from 0.464 to 0.425. We also find that embedding task configuration and passage length can each change reported R@10 by a factor of seven, independent of architecture. These results show why low-resource RAG evaluation should report performance by language condition and query type rather than relying on aggregate scores alone. The dataset and evaluation scripts are available at https://huggingface.co/datasets/RaiyanKhaan/AgriTrust-RAG.
Chinese Translation
RAG 系统中的检索质量通常以单一的综合得分报告,这可能掩盖了不同查询类型和语言条件之间的巨大差异。我们研究了这一问题,聚焦于孟加拉国的农业咨询,其中农民的查询通常是口语化的,而官方咨询文件则使用正式的科学术语。我们构建了一个包含 1,000 个查询和从 284 份孟加拉国官方农业出版物中提取的 2,882 个知识节点的测试集合,并在三种受控语言条件下评估了五种检索架构和六种嵌入模型。结果表明,没有一种检索方法始终表现最佳。对于本土孟加拉语查询,BM25 是最强的单一检索器(R@10 = 0.506),而 Hybrid RRF 达到了最高的整体 R@10 为 0.539。然而,密集检索的表现因查询类型而异:在口语化的农民查询上 R@10 为 0.093,而在正式的安全查询上 R@10 为 0.970。在语言条件方面,BM25 的 R@10 从孟加拉语查询的 0.506 降至与孟加拉语语料库匹配的英语查询的 0.004,而密集检索仅从 0.464 降至 0.425。我们还发现,嵌入任务配置和段落长度各自可以使报告的 R@10 变化七倍,与架构无关。这些结果表明,低资源 RAG 评估应根据语言条件和查询类型报告性能,而不是仅依赖综合得分。数据集和评估脚本可在 https://huggingface.co/datasets/RaiyanKhaan/AgriTrust-RAG 获取。
cs.CL / 20 / 2608.14896

Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs

小型语言模型中的可解释跨语言对齐:探究日英双语大型语言模型中的文化与语用推理
Braun, Florian
Abstract
Large language models work well on English and behave in poorly understood ways on languages typologically far from it. Japanese is a clean example, where evaluation still leans on translation quality and JGLUE-style benchmarks, which roll lexical, syntactic and pragmatic competence into a single score. The phenomena on which general-purpose models fail Japanese users are pragmatic: honorifics, in-group and out-group reference, context-sensitive politeness, zero anaphora. I introduce J-PragEval-v0, a minimal-pair benchmark isolating four such phenomena from surface fluency, and combine it with linear probes and teacher-forced log-probability evaluation to ask where inside TinySwallow-1.5B (28 layers, hidden size 1536) the corresponding contrasts live. The four features split three ways. Honorific register sits cleanly in the residual stream: 0.96 balanced accuracy at layer 15, and the model flips its preferred continuation with the scenario on 93 percent of items. Implicit subject and in-group reference are not linearly decodable at the final prompt token (0.48 and 0.38), yet flip rates are 0.77 and 0.79, so the contrast is worked out during generation rather than stored at the prompt. Indirect refusal is the negative case: 0.95 probe accuracy collapsing to a 0.43 flip rate under length-normalised teacher forcing, because the current minimal pairs conflate politeness with continuation length. I also specify Pragmatic Representation Steering, a parameter-free inference-time method that edits residual-stream activations along the class-mean-difference directions probing identifies. Feasibility is argued indirectly rather than demonstrated: the contrastive activation addition baseline, the same geometry the method would inject, recovers probe accuracy within one to two points of logistic regression wherever a linear signal exists. Scaling to Llama-3.1-Swallow-8B is the next step.
Chinese Translation
大型语言模型在英语上表现良好,但在与其类型学差异较大的语言上表现不佳,且其行为难以理解。日语是一个典型的例子,目前的评估仍然依赖于翻译质量和JGLUE风格的基准,这些基准将词汇、句法和语用能力合并为一个单一的评分。通用模型在日语用户面前失败的现象主要是语用性的:敬语、内外群体的指称、上下文敏感的礼貌以及零指代。我引入了J-PragEval-v0,这是一个最小对比基准,旨在从表面流畅性中分离出四种此类现象,并结合线性探测器和教师强制对数概率评估,探讨在TinySwallow-1.5B(28层,隐藏层大小1536)内部相应对比的存在位置。这四个特征分为三类。敬语注册清晰地位于残差流中:在第15层的平衡准确率为0.96,模型在93%的项目中根据场景翻转其首选延续。隐性主语和内群体指称在最终提示标记处无法线性解码(分别为0.48和0.38),但翻转率为0.77和0.79,因此对比是在生成过程中得出的,而不是存储在提示中。间接拒绝是负案例:在长度归一化的教师强制下,探测准确率为0.95,但翻转率降至0.43,因为当前的最小对比将礼貌与延续长度混淆。我还指定了语用表示引导(Pragmatic Representation Steering),这是一种无参数的推理时方法,沿着探测识别的类别均值差异方向编辑残差流激活。可行性是间接论证而非直接展示:对比激活附加基线,即该方法将注入的相同几何形状,在存在线性信号的情况下,恢复探测准确率在逻辑回归的一到两个点之内。下一步是扩展到Llama-3.1-Swallow-8B。
cs.CL / 21 / 2608.14905

How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

代理在AutoResearch中的失败:对100个真实世界前沿研究任务的端到端诊断评估
Fei, Yanlin, Liu, Nazhou, Yu, Xinmiao, Chen, Shaolong, Li, Lei, Thapa, Rahul, Ciobanu, Madalina, Mao, Qingqing, Das, Ritankar
Abstract
AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowly-scoped, evaluation measures performance but not process, and failure diagnoses lack systematic coverage or artifact-level visibility. To address this gap, we introduce AutoResearchEval, featuring 100 tasks grounded in published frontier science across 7 scientific domains and the full research lifecycle, including ideation, retrieval, execution, analysis, writing, and review. Evaluating 8 harness-model combinations yields 800 autoresearch agent trajectories, with process-level annotation. We organize these insights into AutoResearch Failure Taxonomy or ARFT, a framework of 45 empirically-grounded failure patterns. To enable scalable fine-grained attribution, we leverage a human-calibrated agent-as-a-judge pipeline to inspect complete trajectories and intermediate artifacts. Failure patterns converge on a single overarching limitation, namely that current agents lack a metacognitive loop, which entails the ability to check what they produced against what they found, revise when it does not hold up, and question whether the path they took was sound. The same patterns recur across all 8 harness-model combinations, including the strongest models tested, locating the deficit at the model level rather than in any particular scaffold; whether orchestration-level interventions can close it is an open question this work does not test. We publicly release AutoResearchEval and ARFT to facilitate continued research and development in autonomous scientific discovery.
Chinese Translation
人工智能长期以来一直辅助科学研究,但大规模语言模型(LLMs)和代理框架的快速发展正在重塑这一领域;现在单一系统可以从初始假设一路推进到最终发表的论文,这一范式被称为AutoResearch。现有的评估对这些代理的操作方式或其失败原因了解甚少。任务范围狭窄,评估侧重于性能而非过程,失败诊断缺乏系统性覆盖或工件级可见性。为了解决这一问题,我们引入了AutoResearchEval,涵盖7个科学领域中基于已发表前沿科学的100个任务,以及完整的研究生命周期,包括构思、检索、执行、分析、写作和审查。评估8种模型组合产生了800条自动研究代理轨迹,并进行了过程级注释。我们将这些见解组织成AutoResearch失败分类法(ARFT),这是一个包含45种经验基础失败模式的框架。为了实现可扩展的细粒度归因,我们利用人类校准的代理作为评判者的流程,检查完整轨迹和中间工件。失败模式集中在一个主要限制上,即当前代理缺乏元认知循环,这意味着它们无法将所产生的内容与所发现的内容进行核对,在不符合时进行修正,并质疑所采取路径的合理性。这些模式在所有8种模型组合中反复出现,包括测试中最强的模型,表明缺陷存在于模型层面而非特定框架中;是否可以通过协调层面的干预来弥补这一缺陷仍然是一个未解的问题。本研究公开发布AutoResearchEval和ARFT,以促进自主科学发现的持续研究和发展。
cs.CL / 22 / 2608.14929

Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification

训练留下痕迹:用于语言模型谱系验证的中心残差特征
Thakur, Aman Singh, Khoury, Rayan
Abstract
Open-weight language models are fine-tuned, quantized, pruned, and merged, yet their provenance is often undocumented. We study data-free white-box lineage verification: can weights alone reveal whether two compatible model checkpoints share ancestry? Residual training produces a shared identity-aligned component in branch products, so this structure alone cannot establish ancestry. We remove it and compare checkpoint-specific structure across residual blocks, yielding a symmetric lineage score calibrated against independent checkpoints. On residual-MLP and GPT-2 benchmarks, the score separates fine-tuned, LoRA-merged, pruned, and quantized descendants from independent and distilled models (AUROC=1.0), distinguishing weight ancestry from behavioral similarity. Under function-preserving checkpoint laundering experiments, weight-space baselines lose margin or fail; our score remains unchanged and runs 76x faster than the nearest robust baseline on GPT-2. The projection-pairing signal appears across six language-model families and beyond, and a case study correctly identifies 3 related and 7 unrelated LLaMA-2 public checkpoints. Collectively, these results establish a passive, data-free provenance signal for compatible open-weight language-model checkpoints
Chinese Translation
开放权重语言模型经过微调、量化、剪枝和合并,但其来源往往没有记录。我们研究无数据的白盒谱系验证:仅凭权重是否能揭示两个兼容模型检查点是否共享祖先?残差训练在分支产品中产生了一个共享的身份对齐组件,因此仅凭这一结构无法确立谱系。我们去除该组件,并比较残差块之间的检查点特定结构,从而得出一个针对独立检查点校准的对称谱系评分。在残差多层感知器(residual-MLP)和GPT-2基准测试中,该评分能够将微调的、LoRA合并的、剪枝的和量化的后代与独立和蒸馏模型区分开(AUROC=1.0),从而区分权重谱系与行为相似性。在保持功能的检查点清洗实验中,权重空间基线失去边际或失败;而我们的评分保持不变,并且在GPT-2上比最近的稳健基线快76倍。投影配对信号在六个语言模型家族及其他模型中均有出现,案例研究正确识别了3个相关和7个不相关的LLaMA-2公共检查点。总体而言,这些结果建立了一种针对兼容开放权重语言模型检查点的被动、无数据的来源信号。
cs.CL / 23 / 2608.14950

DA-RAC: Distance-Aware Calibration of LLM Judges for Trustworthy AI Auditing

DA-RAC:面向可信AI审计的LLM评估者距离感知校准
Wu, Cheng, Anand, Vishal, Mandivarapu, Jaya Krishna, Liu, Xiya, Zhuang, Rui
Abstract
Generative AI systems are increasingly producing real-world artifacts, however their efficacy and validity are often evaluated via context-free LLM-scoring. These judges can be miscalibrated by irrelevant in-context reference examples, creating false confidence and allowing low-quality or harmful outputs to pass evaluation. We study this failure mode as context-induced miscalibration and introduce DA-RAC, a distance-aware reference-anchored calibration method for LLM judges. DA-RAC retrieves semantically and structurally similar labeled anchors for each judgement scenario, weights them by distance, and exposes neighborhood difficulty as a calibration and triage signal. On multi-run LLM-judge evaluation benchmarks, it improves calibration and reduces false-pass risk relative to zero-shot, chain-of-thought evaluation, and static-anchor baselines. Mechanistic analysis shows that judge scores vary systematically with anchor distance, while static references can induce misleading decision boundaries. Thus LLM-judgement requires not only better models, but also calibrated, auditable reference selection, especially when automated evaluation is used to support high-impact AI generated artifacts. Judgments should be grounded in relevant, inspectable, and contestable interpretive artifacts.
Chinese Translation
生成性AI系统越来越多地产生现实世界的产物,然而它们的有效性和有效性通常通过无上下文的LLM评分进行评估。这些评估者可能会受到无关的上下文参考示例的错误校准,从而产生虚假的信心,使低质量或有害的输出通过评估。我们将这种失败模式研究为上下文诱导的错误校准,并引入DA-RAC,一种面向距离的参考锚定校准方法,用于LLM评估者。DA-RAC为每个判断场景检索语义和结构上相似的标记锚点,根据距离对其进行加权,并将邻域难度作为校准和筛选信号。在多次运行的LLM评估基准测试中,相较于零-shot、思维链评估和静态锚点基线,DA-RAC改善了校准并降低了误判风险。机制分析表明,评估者的评分与锚点距离系统性变化,而静态参考可能会导致误导性的决策边界。因此,LLM判断不仅需要更好的模型,还需要经过校准的、可审计的参考选择,尤其是在自动评估用于支持高影响力的AI生成产物时。判断应基于相关的、可检查的和可争议的解释性产物。
cs.CL / 24 / 2608.14999

RamseyGadgets: A Graph Construction Dataset for LLMs

RamseyGadgets:一个用于大语言模型的图构造数据集
Hassan, Zohair Raza, Pandita, Deepak
Abstract
Constructing special graphs is an important task within graph theory and computer science. Many popular graph constructions are the result of a comprehensive exploration of relevant graphs and human ingenuity. Given the rise of generative AI usage in mathematics, it is natural to test whether LLMs are able to construct graphs with specified properties using their reasoning capabilities. Unfortunately, many natural graph construction problems, such as finding extremal Ramsey-good graphs (i.e., avoiding specific monochromatic subgraphs), have been explored extensively in the literature, making it difficult to ascertain whether a construction is the product of an LLM's reasoning capabilities or its recollection from training data. In this work, we introduce \textbf{RamseyGadgets}, a novel dataset of 70 underexplored graph construction problems that require finding Ramsey-good graphs with special properties (e.g., containing an edge with a fixed color). These problems have reasonably sized solutions (at most 10 vertices) that can be verified by SAT solvers, making them suitable for automatic evaluation. Our dataset is easily expandable, as one can simply change the monochromatic subgraphs being avoided to obtain a new set of problems. We evaluate the performance of five open-source LLMs on our dataset and report the results. Our findings show that LLMs achieve only 37.70% accuracy on the hard-tier problems in our dataset, with Gemma-4-31B achieving the highest performance out of the five. We also showcase how our dataset allows us to ascertain what kind of hints help LLMs perform better at this task.
Chinese Translation
构造特殊图形是图论和计算机科学中的一项重要任务。许多流行的图构造是对相关图形的全面探索和人类智慧的结果。鉴于生成性人工智能在数学中的使用日益增加,自然需要测试大语言模型(LLMs)是否能够利用其推理能力构造具有特定属性的图形。不幸的是,许多自然图构造问题,例如寻找极端的Ramsey-good图(即避免特定的单色子图),在文献中得到了广泛的探讨,这使得很难确定一个构造是LLM推理能力的产物还是其训练数据的回忆。在本研究中,我们介绍了 extbf{RamseyGadgets},这是一个包含70个尚未充分探索的图构造问题的新数据集,这些问题要求找到具有特殊属性的Ramsey-good图(例如,包含一条固定颜色的边)。这些问题的解决方案大小适中(最多10个顶点),可以通过SAT求解器进行验证,因而适合自动评估。我们的数据集易于扩展,因为只需更改要避免的单色子图即可获得一组新的问题。我们评估了五个开源LLM在我们数据集上的表现,并报告了结果。我们的发现表明,LLM在我们数据集中困难级别问题上的准确率仅为37.70%,其中Gemma-4-31B在五个模型中表现最佳。我们还展示了我们的数据集如何帮助我们确定哪些提示能够帮助LLM在此任务中表现更好。
cs.CL / 25 / 2608.15008

Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents

利用记忆:记忆代理中记忆基质的整体评估
Huang, Wei-Chieh, Zhang, Weizhi, Wu, Yuchen, Chen, Yankai, Jiang, Eric Hanchen, Yang, Wooseong, Yang, Yiwei, Zou, Henry Peng, Zhang, Hanrong, Wu, Ying Nian, Wu, Haolun, Chang, Kai-Wei, Yu, Philip S., Liu, Xue, Caliskan, Aylin
Abstract
Memory is becoming core infrastructure for long-horizon LLM agents, yet existing evaluations offer limited guidance on which memory substrate, namely the underlying medium in which memory is represented and stored, should be used under different operating regimes. We present a controlled harness evaluation of memory substrates for memory-augmented agents, covering dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and activation-compatible context mechanisms. Across three backbone models and four benchmark suites spanning user-centric question answering and agent-centric decision-making, we instrument 26 performance and efficiency metrics under a unified harness. Our results show that no single substrate consistently dominates: broad retrieval benefits long-context factual QA, while excessive retrieval can harm sequential decision-making by shifting attention away from action-critical context. Scalability introduces a further routing axis, as substrates that perform well at moderate history lengths can become costly or brittle at longer horizons. These findings motivate substrate routing as a necessary component of adaptive agent memory systems and provide empirical guidance for designing efficient, reliable, and regime-aware long-term memory for LLM agents. Code will be made available upon acceptance.
Chinese Translation
记忆正成为长时间跨度大语言模型(LLM)代理的核心基础设施,但现有评估对在不同操作模式下应使用哪种记忆基质(即记忆表示和存储的基础介质)提供的指导有限。我们针对增强记忆代理的记忆基质进行了受控的评估,涵盖了密集和稀疏索引、文本记录、结构存储、层次存储、基于精炼的记忆、参数更新和激活兼容的上下文机制。在三个基础模型和四个基准测试套件中,我们对用户中心的问题回答和代理中心的决策制定进行了评估,使用统一的评估框架测量了26个性能和效率指标。我们的结果表明,没有单一的基质能够始终占据优势:广泛的检索有利于长上下文的事实问答,而过度的检索可能会通过转移注意力而损害顺序决策,导致忽视关键行动上下文。可扩展性引入了进一步的路由维度,因为在适中的历史长度下表现良好的基质在较长时间跨度上可能变得昂贵或脆弱。这些发现促使基质路由成为自适应代理记忆系统的必要组成部分,并为设计高效、可靠且适应不同操作模式的长期记忆提供了实证指导。代码将在接受后发布。
cs.CL / 26 / 2608.15032

Handoff-H1: An Orchestrated Vision-Agent System for Material Quantity Takeoff from Construction Blueprints

Handoff-H1:一种用于从建筑蓝图中提取材料数量的协调视觉-代理系统
Chicelli, Bruno, Alves, Henrique, Anselmo, Rodrigo, Weinberg, Joshua, Lemos, Felipe, Baryla, Jan
Abstract
Converting a set of architectural blueprints into a complete material quantity takeoff requires visual perception across drawing sheets, dimensional and multi-hop reasoning, and grounding in construction conventions that the drawings never state. We present Handoff-H1, a takeoff system built from three layers: purpose-built computer-vision models that extract primitives; tool-using agents equipped with image operations and in-house visual-task tools, including CV-model-backed counting, detection and plan decomposition; and a persistent, hierarchically structured project foundation, grounded in a curated construction knowledge base. We evaluate on the Construction Blueprint Takeoff Benchmark: 10 real residential blueprint sets paired with consensus-validated expert takeoffs - 2,009 verified line items, restricted for scoring to the 1,348 primary-tier materials that drive an estimate - scored per trade by an LLM judge on material coverage and quantity Precision@25% ([email protected]) and combined into a weighted composite. Under identical scoring from the raw PDF, seven frontier and open-weight models span composites of 35-61, and independent professional estimators - scored against the same reconciled gold standard - post 77.6% (65.5% coverage, 87.9% [email protected]). Handoff-H1, working end-to-end from the raw PDF, reaches 81.6% (86.1% coverage, 78.8% [email protected]): roughly 20 points above the strongest frontier agent, and above the independent estimators by pairing near-human quantity precision with coverage they do not reach. The evaluation harness is public for the open harbor framework; the blueprint sets and ground truth are available upon request for research use.
Chinese Translation
将一组建筑蓝图转换为完整的材料数量提取需要跨绘图纸的视觉感知、维度和多跳推理,以及对建筑规范的理解,而这些规范在图纸中并未明确说明。我们提出了Handoff-H1,这是一种由三层构成的提取系统:专门构建的计算机视觉模型用于提取基本元素;配备图像操作和内部视觉任务工具的工具使用代理,包括基于计算机视觉模型的计数、检测和计划分解;以及一个持久的、分层结构的项目基础,基于经过整理的建筑知识库。我们在建筑蓝图提取基准上进行评估:10套真实的住宅蓝图与经过共识验证的专家提取配对——2,009个经过验证的条目,评分限制在驱动估算的1,348个主要材料上——由大型语言模型(LLM)评审根据材料覆盖率和数量精度@25%([email protected])进行评分,并合并为加权复合评分。在相同的原始PDF评分下,七个前沿和开放权重模型的复合评分范围为35-61,而独立专业估算师——与相同的经过调和的黄金标准进行评分——达到了77.6%(65.5%的覆盖率,87.9%的[email protected])。Handoff-H1从原始PDF端到端工作,达到了81.6%(86.1%的覆盖率,78.8%的[email protected]):大约比最强的前沿代理高出20个百分点,并且通过结合接近人类的数量精度与他们无法达到的覆盖率,超越了独立估算师。评估工具对开放港口框架是公开的;蓝图集和真实数据可根据请求用于研究。
cs.CL / 27 / 2608.15062

RecurrentGPT: Expressive Depth through Recurrent Modulation in Transformers

RecurrentGPT:通过递归调制在变换器中实现表现力深度
Hegazy, Amr, Alanwar, Amr, Elhoushi, Mostafa
Abstract
Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While unique weights across layers preserve functional specialization---from input-grounding to abstract refinement---they incur a substantial memory footprint. Conversely, standard depth-sharing enforces uniform transformations that collapse representational diversity and degrade modeling quality. We introduce RecurrentGPT, a recurrent depth transformer where fixed-depth prelude and coda blocks bracket a single shared core iterated R times. Inspired by gated recurrent neural networks, we employ a lightweight projection and an elementwise update gate---conditioned on the hidden state, the fixed prelude output, and noise resampled at every step---to modulate the recurrent update. This allows the model to specialize the input to the same few layers across recurrences, rather than requiring many unique layers to achieve functional diversity. Under an isoFLOPS constraint, a 3-layer RecurrentGPT matches the accuracy of a 12-layer GPT-2 Small baseline with similar training and inference FLOPs, and leads MoR and heavy-tail depth sampling in all nine scale-by-budget cells; at medium and large scale it approaches dense quality at the standard token budget and overtakes it at medium scale once that budget is doubled. Under an isoPARAMS constraint, deeper recurrence achieves a 2.76 validation loss versus 2.84 for a non-recurrent counterpart at matched parameter and data budget. Our results demonstrate that adaptive depth reuse is a principled strategy for trading parameters for quality: at large scale, 63% fewer parameters and 59% less peak decoding memory for a 10% increase in compiled generation latency.
Chinese Translation
扩展变换器语言模型在表现力与内存效率之间产生了固有的张力。虽然跨层的独特权重保持了功能专门化——从输入基础到抽象细化——但它们会产生相当大的内存占用。相反,标准的深度共享强制执行统一的变换,这会压缩表征多样性并降低建模质量。我们提出了RecurrentGPT,一种递归深度变换器,其中固定深度的前奏和尾声块包围一个共享核心,迭代R次。受到门控递归神经网络的启发,我们采用轻量级投影和逐元素更新门——基于隐藏状态、固定前奏输出和在每一步重新采样的噪声——来调制递归更新。这使得模型能够在递归中将输入专门化到相同的少数层,而不是需要许多独特层来实现功能多样性。在isoFLOPS约束下,3层的RecurrentGPT在与相似训练和推理FLOPs的情况下,达到了12层GPT-2 Small基线的准确性,并在所有九个按预算缩放的单元中引领MoR和重尾深度采样;在中等和大规模下,它在标准令牌预算下接近密集质量,并在中等规模下一旦预算翻倍便超越了它。在isoPARAMS约束下,较深的递归在匹配参数和数据预算的情况下实现了2.76的验证损失,而非递归对应物的验证损失为2.84。我们的结果表明,自适应深度重用是一种在参数与质量之间进行权衡的原则性策略:在大规模下,减少63%的参数和59%的峰值解码内存,实现了10%的生成延迟增加。
cs.CL / 28 / 2608.15080

A Pilot Study of Autocompleting Tokenizers

自完成分词器的初步研究
Wexler, Samuel, Hopkins, Mark
Abstract
Modern input methods routinely rely on autocomplete to omit information that can be recovered from local context. Inspired by these autocomplete-assisted writing systems, we investigate whether Transformer inputs can be compressed in a similar manner. Byte-level tokenization offers a simple and language-independent alternative to subword tokenization, but its longer input sequences typically result in increased computational cost and reduced model quality. We propose a compression scheme that employs a lightweight autoregressive byte language model to identify and remove bytes that are easily predictable from their surrounding context before Transformer processing. The resulting compressed representation is then provided as input to a standard encoder--decoder Transformer. Experiments on machine translation show that a substantial fraction of source-language bytes can be omitted without degrading translation quality. On English--French, our best method preserves translation performance while reducing source sequence length by nearly one-third. Additional experiments on Finnish--English, Russian--English, and Chinese--English demonstrate that the approach generalizes across diverse writing systems and morphological typologies, yielding comparable or improved translation quality at compression ratios between 0.47 and 0.67. These findings suggest that many input bytes are predictable enough to be represented implicitly rather than explicitly, providing a simple mechanism for reducing the sequence-length overhead associated with byte-level models.
Chinese Translation
现代输入法通常依赖自动补全功能来省略可以从局部上下文中恢复的信息。受到这些辅助写作系统的启发,我们研究了变换器输入是否可以以类似的方式进行压缩。字节级分词提供了一种简单且与语言无关的替代方案,但其较长的输入序列通常导致计算成本增加和模型质量下降。我们提出了一种压缩方案,利用轻量级自回归字节语言模型来识别并移除那些可以从周围上下文中轻易预测的字节,然后再进行变换器处理。生成的压缩表示随后作为输入提供给标准的编码器-解码器变换器。机器翻译实验表明,源语言字节中有相当一部分可以被省略而不降低翻译质量。在英法翻译中,我们的最佳方法在减少源序列长度近三分之一的同时保持了翻译性能。对芬兰语-英语、俄语-英语和汉语-英语的额外实验表明,该方法在不同的书写系统和形态类型中具有良好的泛化能力,在压缩比为0.47到0.67之间实现了可比或更好的翻译质量。这些发现表明,许多输入字节足够可预测,可以隐式表示而非显式表示,从而提供了一种简单的机制来减少与字节级模型相关的序列长度开销。
cs.CL / 29 / 2608.15085

Why Vision Fails as a Universal Bridge: Rectifying Modality Asynchrony in Multilingual MLLMs

为何视觉无法作为通用桥梁:纠正多语言 MLLMs 中的模态异步问题
Du, Yihang, Liang, Juhao, Lai, Zhengzhao, Li, Siyu, Hu, Yan
Abstract
Multimodal large language models (MLLMs) exhibit substantial performance degradation in non-English visual reasoning, despite the strong multilingual competence of their text-only backbones. While mechanistic evidence from text-only models suggests that non-English inputs are routed through an English-centric latent space, the multimodal implications of this phenomenon remain unexplored. Through rigorous mechanistic analysis, we identify the \textbf{Ghost Anchor} phenomenon: a temporal modality asynchrony where linguistic translation to the English semantic manifold completes in early layers, while visual semanticization remains immature. Consequently, visual signals are physically present yet functionally invisible during the early alignment window. To rectify this, we propose \textbf{ANCHOR}, a training framework employing Proactive Visual Anchoring (PVA) to accelerate early visual semantic emergence, ensuring visual representations proactively guide linguistic translation. Mechanistic interventions confirm that ANCHOR successfully restores the causal influence of visual signals during early translation. Furthermore, extensive experiments on XMMMU, MaXM, and CVQA demonstrate that ANCHOR consistently outperforms standard baselines, achieving robust visual reasoning across both fine-tuned and zero-shot languages.
Chinese Translation
尽管文本模型在多语言能力上表现出色,但多模态大型语言模型(MLLMs)在非英语视觉推理中却表现出显著的性能下降。文本模型的机械证据表明,非英语输入通过以英语为中心的潜在空间进行处理,但这一现象的多模态含义尚未得到探讨。通过严格的机械分析,我们识别出 extbf{幽灵锚点}现象:一种时间模态异步现象,其中语言翻译到英语语义流形在早期层中完成,而视觉语义化仍然不成熟。因此,视觉信号在早期对齐窗口中虽然在物理上存在,但在功能上却是不可见的。为了解决这一问题,我们提出了 extbf{锚点(ANCHOR)},一个采用主动视觉锚定(Proactive Visual Anchoring, PVA)的训练框架,以加速早期视觉语义的出现,确保视觉表征主动引导语言翻译。机械干预证实,ANCHOR 成功恢复了视觉信号在早期翻译中的因果影响。此外,在 XMMMU、MaXM 和 CVQA 上的广泛实验表明,ANCHOR 始终优于标准基线,在微调和零样本语言中实现了强大的视觉推理能力。
cs.CL / 30 / 2608.15102

A Declarative-Procedural Perspective on Expert Routing in Bilingual Mixture-of-Experts Language Models

双语混合专家语言模型中专家路由的声明性-程序性视角
Gopinath, Amrit, Raghul, Thenmozhi, Durairaj
Abstract
We investigate whether Mixture-of-Experts (MoE) language models develop linguistically structured expert routing during bilingual language acquisition. Inspired by the Declarative-Procedural framework, we analyze lexical, grammatical, and syntactic processing in a decoder-only English-German MoE Transformer trained under sequential language exposure. We construct a probe-based validation set and extract token-level routing distributions to quantify category-dependent specialisation using mutual information, routing entropy, and Jensen-Shannon distance. The curriculum-trained model exhibits a peak mutual information of 0.1148 at layer 5, indicating category-dependent differences in routing distributions across linguistic categories. Surprisingly, a no-curriculum baseline trained on mixed English-German data shows stronger aggregate specialisation, reaching a peak mutual information of 0.2599 at the same layer. These results suggest that interpretable linguistic organization emerges within MoE routing patterns even without sequential language exposure. A replication at a second training seed shows that the no-curriculum condition's specialisation concentrates on a single language whose identity is seed-dependent, whereas the curriculum consistently yields a stable, language-balanced routing profile; rather than uniformly increasing specialisation, staged bilingual exposure reduces single-language dominance. The official Github repository: https://github.com/Amrit828/DP-Theory-MOE-Interpretability-Research
Chinese Translation
我们研究混合专家(Mixture-of-Experts, MoE)语言模型在双语语言习得过程中是否发展出具有语言结构的专家路由。受到声明性-程序性框架的启发,我们分析了在顺序语言暴露下训练的仅解码器的英德MoE Transformer中的词汇、语法和句法处理。我们构建了一个基于探测器的验证集,并提取了令牌级路由分布,以利用互信息、路由熵和杰森-香农距离量化类别依赖的专业化。经过课程训练的模型在第5层展现出0.1148的峰值互信息,表明在语言类别之间的路由分布存在类别依赖的差异。令人惊讶的是,在混合英德数据上训练的无课程基线显示出更强的整体专业化,在同一层达到0.2599的峰值互信息。这些结果表明,即使在没有顺序语言暴露的情况下,MoE路由模式中也会出现可解释的语言组织。在第二个训练种子下的复制实验显示,无课程条件下的专业化集中于一个单一语言,其身份依赖于种子,而课程训练则始终产生稳定的、语言平衡的路由特征;而不是均匀增加专业化,分阶段的双语暴露减少了单一语言的主导地位。官方Github仓库:https://github.com/Amrit828/DP-Theory-MOE-Interpretability-Research
cs.CL / 31 / 2608.15129

Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models

左分支变换器在右分支语言中表现优异:数据塑造语言模型中的词序偏好
Arzt, Varvara, Hanbury, Allan, Blevins, Terra
Abstract
We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial languages, models exhibit a left-branching preference that aligns with neither natural language universals nor human word order learning biases. On natural languages, monolingual models show no clear base word order bias at small scales, but as data grows, a preference for right-branching subject-verb-object (SVO) languages emerges while SOV falls behind despite being the most frequent order cross-linguistically. This SVO advantage extends to multilingual models and correlates with language resource level and data quality rather than word order. Thus, the same architecture exhibits opposite preferences on artificial and natural languages, establishing that word order biases observed in practice are data-driven. Since highly-resourced languages are overwhelmingly SVO, these biases risk gradually reducing word order diversity, particularly in languages that productively use multiple word orders, with the widespread adoption of LLMs.
Chinese Translation
我们系统地比较了192种人造语言和类型多样的自然语言中解码器语言模型的词序偏好。在人造语言中,模型表现出一种左分支偏好,这与自然语言的普遍规律或人类的词序学习偏见都不一致。在自然语言中,单语模型在小规模数据下没有明显的基础词序偏见,但随着数据量的增加,出现了对右分支主谓宾(SVO)语言的偏好,而尽管主宾谓(SOV)是跨语言中最常见的词序,但其偏好却逐渐落后。这种SVO优势扩展到多语种模型,并与语言资源水平和数据质量相关,而非词序。因此,同一架构在人工语言和自然语言中表现出相反的偏好,确立了实践中观察到的词序偏见是数据驱动的。由于资源丰富的语言主要是SVO,这些偏见有可能逐渐减少词序的多样性,特别是在那些有效使用多种词序的语言中,随着大型语言模型(LLMs)的广泛应用。
cs.CL / 32 / 2608.15223

TRACE-BN: Transferring Bangla-English Tutoring Behavior to a Sub-1B Offline Language Model

TRACE-BN:将孟加拉语-英语辅导行为转移到一个低于10亿参数的离线语言模型
Reza, Khan Raiyan Ibne, Maria, Sanjana Aktar, Abdullah, Mohammad Tushar, Leen, Asfee Bhuiyan, Nimi, Sumaiya Tabassum
Abstract
Bangla-English tutoring requires more than producing a correct translation: learners also need explanations of grammar differences, awareness of their likely errors, and targeted practice. We present TRACE-BN, a curriculum-guided dataset of structured tutoring traces for Bangla-speaking learners of English at the CEFR A1-A2 level. Each trace combines word-level glosses, literal and natural translations, Bangla grammar explanations, a plausible learner error, and a targeted practice question with its answer. The traces are generated by Gemini 3.5 Flash Lite as the teacher model from NCTB Classes 9-10 English curriculum units, then filtered for structural validity, script integrity, and semantic duplication. We transfer the resulting structured tutoring behavior to Qwen3-0.6B using LoRA with 4-bit quantization for resource-constrained offline deployment. On held-out inputs, schema validity increases from 85.4% to 95.8%, while, against teacher-model references, chrF++ improves from 15.28 to 34.77 and BLEU from 4.52 to 21.03. Field-level evaluation by two independent judges shows improvements across translation, grammar explanation, learner-error diagnosis, and practice alignment, while a human audit supports the quality of the supervision data. The results show that curriculum-guided structured supervision can transfer multi-component tutoring behavior to a sub-1B model under these resource constraints. The dataset, model checkpoints, and code are publicly available at https://huggingface.co/datasets/RaiyanKhaan/Trace-BN
Chinese Translation
孟加拉语-英语辅导不仅仅需要产生正确的翻译:学习者还需要对语法差异的解释、对可能错误的意识以及针对性的练习。我们提出了TRACE-BN,这是一个针对CEFR A1-A2水平孟加拉语学习者的结构化辅导轨迹的课程指导数据集。每个轨迹结合了词汇级别的注释、字面和自然翻译、孟加拉语语法解释、一个合理的学习者错误以及一个带答案的针对性练习问题。这些轨迹由Gemini 3.5 Flash Lite作为教师模型生成,基于NCTB 9-10年级英语课程单元,然后经过结构有效性、脚本完整性和语义重复性的筛选。我们使用LoRA和4位量化将生成的结构化辅导行为转移到Qwen3-0.6B,以便在资源受限的离线环境中部署。在保留的输入上,模式有效性从85.4%提高到95.8%;而在教师模型参考下,chrF++从15.28提高到34.77,BLEU从4.52提高到21.03。两位独立评审员的领域级评估显示在翻译、语法解释、学习者错误诊断和练习对齐方面都有所改善,同时人工审核支持监督数据的质量。结果表明,在这些资源限制下,课程指导的结构化监督能够将多组件辅导行为转移到一个低于10亿参数的模型。数据集、模型检查点和代码可在https://huggingface.co/datasets/RaiyanKhaan/Trace-BN公开获取。
cs.CL / 33 / 2608.15270

Time as Structure: Temporal Dependency Graphs for Verifiable Deadline Computation over Legal Documents

时间作为结构:法律文件可验证截止日期计算的时间依赖图
Zhyrko, Maryia, Han, Lifeng, Verberne, Suzan
Abstract
Miss a filing deadline by one day and the claim is barred, however strong the case. Computing that deadline is rarely simple: the period runs from a triggering event, is counted by a statutory convention, and may be suspended by a mandatory conciliation window. We ask whether a language model should answer such questions directly, or read the document and leave the arithmetic to code. We extract dated facts and their dependencies into a temporal dependency graph and compute deadlines from it with a calendar-correct engine. On UK Employment Appeal Tribunal judgments the engine reproduces six of seven timeliness rulings, and matches the judges' own dates to the day. The strongest of four language models, asked the same cases, gets the arithmetic right and the answer wrong: in six of twenty-one responses its stated verdict contradicts its own thinking, and every contradiction runs the same way, calling a late claim timely. To test the systems at scale we move the dismissal date across the statutory boundary, generating 427 cases whose answers are computed rather than annotated. On the cases both systems answer, the pipeline is right 90.2% of the time against 61.2% for direct answering. The limit is extraction: on contracts the errors are almost never in the arithmetic, but in choosing which event the period starts from.
Chinese Translation
错过提交截止日期一天,索赔便会被驳回,无论案件多么强有力。计算这个截止日期往往并不简单:期限从触发事件开始,根据法定惯例计算,并可能因强制调解窗口而暂停。我们探讨语言模型是否应直接回答此类问题,或是阅读文件并将算术运算留给代码。我们将带日期的事实及其依赖关系提取到时间依赖图中,并利用一个日历正确的引擎从中计算截止日期。在英国就业上诉法庭的判决中,该引擎重现了七个及时性裁决中的六个,并将法官的日期精确匹配到天。四个语言模型中最强的一个,在处理相同案件时,算术运算正确但答案错误:在二十一条回应中,有六条其所述裁决与自身思考相矛盾,且每个矛盾都朝同一方向发展,称一个迟交的索赔为及时。为了大规模测试系统,我们将解雇日期移动到法定边界,生成427个案例,其答案是通过计算而非注释得出的。在两个系统都能回答的案例中,管道的正确率为90.2%,而直接回答的正确率为61.2%。限制在于提取:在合同中,错误几乎从未出现在算术运算上,而是在选择期限开始的事件上。
cs.CL / 34 / 2608.15323

When Do Concepts Become Functionally Sufficient During Language-Model Training?

概念在语言模型训练中何时变得功能充分?
Bernas, Raphael, Chevalier, Paul G., Jourdan, Fanny, Hudelot, Céline
Abstract
Understanding a model and its learning mechanisms in depth requires identifying when its internal structures become useful, rather than simply looking at the final state. We study this through concept dynamics: at each layer and checkpoint, we decompose activations, select sparse soft masks, and inject masked reconstructions into the model. Concept analysis is therefore tested functionally: a mask is useful only insofar as it preserves a target under intervention. We compare sufficiency for activation reconstruction, linear decodability, true downstream preservation, and checkpoint transfer under learned alignment. The framework treats decomposition assumptions as hypotheses rather than interpretability guarantees, monitoring functional sufficiency across checkpoints and source-to-final reconstructability under learned alignment. At the shared fixed-penalty operating point across seven models, downstream masks retain substantially less soft mass than reconstruction masks; predictive-distribution shifts remain small.
Chinese Translation
深入理解模型及其学习机制需要识别其内部结构何时变得有用,而不仅仅是关注最终状态。我们通过概念动态来研究这一问题:在每一层和检查点,我们分解激活,选择稀疏的软掩码,并将掩码重构注入模型。因此,概念分析在功能上进行测试:掩码只有在干预下保留目标时才是有用的。我们比较了激活重构、线性可解性、真实下游保留和在学习对齐下的检查点转移的充分性。该框架将分解假设视为假说,而非可解释性保证,监测跨检查点的功能充分性和在学习对齐下的源到最终重构能力。在七个模型的共享固定惩罚操作点上,下游掩码保留的软质量显著低于重构掩码;预测分布的变化仍然很小。
cs.CL / 35 / 2608.15325

Logical Embeddings for Argument Analysis

用于论证分析的逻辑嵌入
Heldring, Leander, Torres, Santiago
Abstract
We propose a new framework for machine-learning-oriented argument analysis tasks. Our proposal involves replacing traditional contextualized word embeddings used in most NLP tasks with logical embeddings, an alternative encoding that directly exploits argumentation structures. In essence, logical embeddings encapsulate the logical semantics of an argument, allowing for a better representation of its meaning. Supporting these embeddings is a mathematical logic-based similarity measure that offers a transparent notion of proximity and is guaranteed to satisfy several desirable theoretical properties that current cosine similarity-based contextualized word embeddings cannot assure. This similarity measure induces a positive semi-definite kernel on the set of arguments, enabling us to uniquely define logical embeddings using the theory of Reproducing Kernel Hilbert Spaces (RKHS). Moreover, we prove that this encoding is optimal, in the sense that no logical information is lost in the process. As with other RKHS applications, logical embeddings can be used in numerous supervised and unsupervised tasks. We provide an implementation of the method and aim to test it against literature benchmarks. Additionally, we demonstrate that logical embeddings outperform most standard embedding methods on a classification task.
Chinese Translation
我们提出了一种新的框架,用于面向机器学习的论证分析任务。我们的提案涉及用逻辑嵌入替代大多数自然语言处理任务中使用的传统上下文化词嵌入,这是一种直接利用论证结构的替代编码。从本质上讲,逻辑嵌入封装了论证的逻辑语义,允许更好地表示其含义。支持这些嵌入的是一种基于数学逻辑的相似性度量,提供了一种透明的接近度概念,并且保证满足当前基于余弦相似性的上下文化词嵌入无法保证的若干理想理论属性。这种相似性度量在论证集上诱导出一个正半定核,使我们能够利用再生核希尔伯特空间(Reproducing Kernel Hilbert Spaces, RKHS)理论唯一地定义逻辑嵌入。此外,我们证明这种编码是最优的,因为在此过程中没有丢失任何逻辑信息。与其他RKHS应用一样,逻辑嵌入可以用于许多监督和无监督任务。我们提供了该方法的实现,并旨在对其进行文献基准测试。此外,我们展示了逻辑嵌入在分类任务中优于大多数标准嵌入方法。
cs.CL / 36 / 2608.15338

When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text

当人工智能重写时,分类器放松:对讽刺和人工智能改写社交文本的基于不确定性的情感分析
Shroff, Shresth
Abstract
Sentiment classifiers are increasingly applied to social media content that is either sarcastic or AI-generated --- two distributional regimes where standard evaluations offer little guidance. We present a three-part empirical study of sentiment classifier behaviour under these conditions. First, we find that confidence scores on sarcastic text are significantly lower than on non-sarcastic text (Mann--Whitney $p = 2 \times 10^{-6}$), confirming that classifiers sense their own uncertainty on ironic content even without explicit uncertainty modelling. Second, and counterintuitively, we show that sentiment classifiers achieve higher accuracy on AI-paraphrased reviews than on the original human-authored text (RoBERTa: $+5.8$ pp for Qwen3.5-4B paraphrases, $+3.7$ pp for Gemma4-E4B), revealing a cross-domain stylistic alignment effect: AI paraphrases remove distributional noise that confounds Twitter-trained classifiers, producing cleaner, more prototypical sentiment text. Third, we demonstrate that a lightweight abstention wrapper --- flagging the $14\%$ of inputs with confidence below $0.6$ --- improves accuracy from 82.2\% to 88.9\% ($+6.7$ pp) on the retained set. We further compare Semantic Entropy and MC-Dropout-style disagreement as uncertainty signals and find near-identical AUROC ($0.650$ vs.\ $0.646$) on sarcastic text, suggesting that for short social media inputs, both methods are interchangeable. Our results motivate a shift from confident single-label prediction to uncertainty-aware abstention in high-stakes sentiment applications such as mental health flagging and content moderation.
Chinese Translation
情感分类器越来越多地应用于讽刺或人工智能生成的社交媒体内容——这两种分布状态下,标准评估提供的指导有限。我们提出了一项三部分的实证研究,探讨情感分类器在这些条件下的表现。首先,我们发现讽刺文本的置信度评分显著低于非讽刺文本(Mann--Whitney $p = 2 imes 10^{-6}$),确认分类器即使在没有明确不确定性建模的情况下也能感知到其对讽刺内容的自身不确定性。其次,反直觉的是,我们展示情感分类器在人工智能改写的评论上的准确率高于原始人类撰写的文本(RoBERTa: 对于 Qwen3.5-4B 改写提高了 $+5.8$ 个百分点,对于 Gemma4-E4B 提高了 $+3.7$ 个百分点),揭示了一种跨领域的风格一致性效应:人工智能改写去除了困扰Twitter训练分类器的分布噪声,产生了更干净、更典型的情感文本。第三,我们证明了一种轻量级的弃权包装器——标记置信度低于 $0.6$ 的 $14\%$ 输入——将保留集的准确率从 82.2\% 提高到 88.9\\%($+6.7$ 个百分点)。我们进一步比较了语义熵和 MC-Dropout 风格的不一致性作为不确定性信号,发现讽刺文本上的 AUROC 几乎相同($0.650$ vs. $0.646$),这表明对于短社交媒体输入,两种方法是可以互换的。我们的结果促使我们从自信的单标签预测转向在高风险情感应用中(如心理健康标记和内容审核)采用基于不确定性的弃权策略。
cs.CL / 37 / 2608.15394

The Machine's Internal Clock: Do LLMs Share Human Temporal Illusions?

机器的内部时钟:大型语言模型是否共享人类的时间错觉?
Bao, Catherine, Srikumar, Vivek
Abstract
Human perception of time is subjective. Well-documented temporal illusions show that the brain relies on context and relational cues for judging duration instead of tracking elapsed time directly. Prior studies established these effects with visual and auditory stimuli. Existing LLM evaluations of temporal perception focus on estimating event durations or multi-step temporal reasoning. In this work, we investigate whether written narratives alone can evoke human temporal illusions, using a new benchmark of 6,684 narrative pairs spanning five illusions. We find that human readers (60 participants) prefer expected scenarios in only two of the five illusions, those where the manipulation is directly visible in text rather than requiring readers to internally simulate duration. We evaluate 14 LLMs on the same benchmark. Surprisingly, we find that models pick the literature-predicted scenario across four of the five illusions, diverging from human behavior. Reasoning traces show that ~70% of responses explicitly evoke psychology research, suggesting that this alignment is consistent with retrieval of published findings rather than human-like temporal biases.
Chinese Translation
人类对时间的感知是主观的。充分记录的时间错觉表明,大脑在判断持续时间时依赖于上下文和关系线索,而不是直接跟踪经过的时间。先前的研究通过视觉和听觉刺激确立了这些效应。现有的大型语言模型(LLMs)对时间感知的评估主要集中在估计事件持续时间或多步骤时间推理上。在本研究中,我们探讨仅通过书面叙述是否能够引发人类的时间错觉,使用一个新的基准,包括6684对叙述,涵盖五种时间错觉。我们发现,在五种时间错觉中,60名参与者的人类读者仅在两种情况下偏好预期情景,这两种情况是操控在文本中直接可见,而不需要读者内部模拟持续时间。我们在相同基准上评估了14个大型语言模型。令人惊讶的是,我们发现这些模型在五种时间错觉中的四种情况下选择了文献预测的情景,偏离了人类行为。推理轨迹显示,约70%的响应明确唤起心理学研究,表明这种一致性与检索已发表的研究结果一致,而非人类的时间偏见。
cs.CL / 38 / 2608.15428

Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks

针对一个模型的限制,开放下一个模型:法律多项选择基准中的选项唯一可解性
Ovcharov, Volodymyr
Abstract
Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question. Measuring that gap takes care: a model answering A to most items scores above chance wherever the key sits at A, and reads as recognition when it is not. We measure it on UA-JudgeExam: 11,990 four-option items with official keys, published by Ukraine's Higher Qualification Commission of Judges. Shown the options and no question, Claude Haiku 4.5 scores 0.383 against chance, and the leak is concentrated: 11.8% of items are answered blind on all eight option orders, against 0.2 items expected by chance. It is not quotation: search over 280,059 editions of Ukrainian legislation recovers 0.128. Gating those out retains 8,128 items, on which the gating model itself now scores 0.204, and GPT-5.6, which took no part in the selection, still answers 0.515 of them with the question hidden. Scoring twelve held-out models on the whole set and subtracting each one's answer-position habit, only two keep an excess: GPT-5.6 at +0.265, Sonnet 4.6 at +0.081. Without it the ranking misleads: Llama 3.1 8B scores 0.292 blind, above every model but those two, purely by answering A to 92% of items. The gate does select something real: on the items it rejected, eleven of twelve models score 0.518-0.789, every interval clear of what the same model scores on the items it kept. But that signal is one model's, and filtering on it does not transfer upward. Neither is visible on a 400-item sample, where nine models read as "statistically at chance". Rewriting distractors instead overshoots to 0.168, below chance and as exploitable. The same probe on LEXam returns chance: every option there points into the stem, none longer than 33 characters. Item format decides whether the problem can arise; capability decides how much is extracted. We release the corpus, the predictions and the harness.
Chinese Translation
多项选择基准的评分依据是模型是否选择了正确的选项,而不是其是否需要该问题。衡量这一差距需要谨慎:一个模型在大多数项目中选择A的情况下,无论关键选项是否为A,其得分均高于随机水平,而当关键不为A时则表现为识别。我们在UA-JudgeExam上进行测量:该数据集包含11,990个四选项项目,官方答案由乌克兰高级法官资格委员会发布。在未给出问题的情况下,Claude Haiku 4.5在随机水平下得分为0.383,且泄露集中:11.8%的项目在所有八种选项顺序下均被盲目回答,而随机情况下预期为0.2个项目。这并不是引用:对280,059版乌克兰立法的搜索恢复了0.128。排除这些后,保留了8,128个项目, gating模型本身在这些项目上的得分为0.204,而未参与选择的GPT-5.6在隐藏问题的情况下仍回答了0.515的项目。对整个数据集进行评分并减去每个模型的答案位置习惯,仅有两个模型保持超额:GPT-5.6为+0.265,Sonnet 4.6为+0.081。没有这些,排名会产生误导:Llama 3.1 8B在盲测中得分为0.292,高于除这两个模型外的所有模型,仅仅因为对92%的项目回答了A。这个gate确实选择了一些真实的东西:在被拒绝的项目中,十二个模型中的十一得分在0.518-0.789之间,每个区间都清晰地与同一模型在保留项目上的得分不同。但该信号属于一个模型,基于此进行过滤并不能向上转移。在400个项目的样本中也没有可见的结果,九个模型的表现被视为“统计上处于随机水平”。重写干扰项反而超出预期,得分为0.168,低于随机水平并且可被利用。对LEXam的相同探测返回随机水平:那里的每个选项都指向题干,没有一个超过33个字符。项目格式决定了问题是否会出现;能力决定了提取的程度。我们发布该语料库、预测结果和工具。
cs.CL / 39 / 2608.15443

Semantic Space of Parts of Speech

词性语义空间
Milička, Jiří, Kraus, Ivan, Stanovský, Arnold, Vysloužilová, Anna, Štěpánková, Barbora, Fárová, Lenka, Cink, Vojtěch, Dohnalová, Šárka
Abstract
Parts of speech categorization is understood in the European linguistic tradition as crisp categorization, which is also reflected in corpus linguistics, where each disambiguated token is assigned exactly one POS. However, the assigned categories are largely determined by arbitrary decisions distilled into annotation manuals. Since some words stand between parts of speech in their semantics or typical syntax, and some parts of speech are closer to each other than others, POS categorization seems inherently fuzzy. We analyze this fuzziness using word2vec embeddings, training a neural network to reduce their high dimensionality to three dimensions relevant for determining parts of speech. This creates a three-dimensional space onto which we map several thousand words, revealing which are prototypical and which lie on the boundaries, and visualizing relationships between parts of speech. The study uses Universal Dependencies POS tags for French, Czech, Finnish, Russian, and English.
Chinese Translation
在欧洲语言学传统中,词性分类被理解为明确的分类,这一点在语料库语言学中得到了反映,其中每个消歧义的词元被精确地分配到一个词性(POS)。然而,分配的类别在很大程度上是由提炼成注释手册的任意决定所决定的。由于某些词在其语义或典型句法上介于词性之间,并且某些词性之间的相似性高于其他词性,因此词性分类似乎本质上是模糊的。我们使用 word2vec 嵌入分析这种模糊性,训练神经网络将其高维度降至三个与确定词性相关的维度。这创建了一个三维空间,我们在其中映射数千个词,揭示哪些词是原型,哪些词位于边界,并可视化词性之间的关系。本研究使用了法语、捷克语、芬兰语、俄语和英语的通用依存词性标签。
cs.CL / 40 / 2608.15448

Language models suffer from a curse of ambiguity

语言模型面临模糊性的诅咒
Zucchet, Nicolas, Lee, Hyun Dong, Linderman, Scott
Abstract
Large language models increasingly rely on sampling as a driver of their own improvement, making the fidelity of their learned distributions more critical than ever. Yet, not all distributions are equally easy to learn. In this work, we identify a curse of ambiguity: in large language models, and more broadly in all neural networks that produce discrete probability distributions, the more ambiguous a next-token distribution is, the harder it is to learn accurately. Through an extensive theoretical analysis, we trace this curse to architectural and learning roots. More ambiguous distributions require more capacity to be stored, larger embeddings to be represented, more steps to be fitted, and amplify token-sampling noise. We validate these findings on synthetic tasks with controlled ground truth and observe the same signatures in language models trained on real data. Our results provide a new perspective on the statistical capabilities of large language models and a practical framework for when to trust their output distribution.
Chinese Translation
大型语言模型越来越依赖采样作为自身改进的驱动力,这使得它们学习到的分布的保真度比以往任何时候都更加重要。然而,并非所有分布都同样容易学习。在本研究中,我们识别出一种模糊性的诅咒:在大型语言模型中,以及更广泛地说,在所有生成离散概率分布的神经网络中,下一标记分布越模糊,准确学习的难度就越大。通过广泛的理论分析,我们将这种诅咒追溯到架构和学习的根源。更模糊的分布需要存储更多的容量,需要更大的嵌入表示,需要更多的步骤进行拟合,并放大标记采样噪声。我们在具有控制真实值的合成任务上验证了这些发现,并在使用真实数据训练的语言模型中观察到了相同的特征。我们的结果为大型语言模型的统计能力提供了新的视角,并为何时信任其输出分布提供了一个实用框架。
cs.CL / 41 / 2608.15507

Do Language Models Consistently Encode the Current Year?

语言模型是否一致地编码当前年份?
van Adrichem, Suze, Bhaskar, Aditi, Yang, Diyi, Potts, Christopher, Huang, Jing
Abstract
A consistent concept of the current time is important for temporal reasoning, yet how language models represent the current time is not well understood. We contribute two tasks that probe the current year in conceptually distinct ways: an associative task, which infers the current year from verb tense, and a declarative task, which directly queries for the current year. Both tasks estimate current years within one year of the post-training data cutoff of instruction-tuned language models. For base models, predictions on the associative task serve as a strong proxy for the pre-training data cutoff, with an average error of only 10 months across 13 models. However, their internal mechanisms diverge: the associative task uses mechanisms similar to factual recall, while the declarative task lacks consistent causal pathways. This divergence poses a challenge for updating the current year in language models. None of prompting, SFT, or weight editing succeed in shifting the associative and declarative years simultaneously. Prompting updates the declarative year (94.6% success across 351 target years) but leaves the associative year nearly unchanged (1.7% success). Year-shifted SFT also fails to shift the associative year, matching the target year in only one of eight models. Weight editing, while effective for both tasks individually, does not generalize across both. Overall, our results show that the current year is not consistently encoded in language models: The associative notion, deeply ingrained in linguistic structures learned in pre-training, uses different causal mechanisms and resists the same modifications that easily shift the declarative notion learned in post-training.
Chinese Translation
对当前时间的一致概念对于时间推理至关重要,但语言模型如何表示当前时间尚不清楚。我们提出了两个以概念上不同方式探测当前年份的任务:一个联想任务,通过动词时态推断当前年份;一个声明性任务,直接查询当前年份。这两个任务估计的当前年份在指令调优语言模型的训练数据截止日期的一年内。对于基础模型,联想任务的预测作为预训练数据截止日期的强有力代理,13个模型的平均误差仅为10个月。然而,它们的内部机制却有所不同:联想任务使用类似于事实回忆的机制,而声明性任务缺乏一致的因果路径。这种差异给语言模型更新当前年份带来了挑战。无论是提示、微调(SFT)还是权重编辑,都未能同时改变联想和声明的年份。提示更新了声明年份(在351个目标年份中成功率为94.6%),但几乎没有改变联想年份(成功率为1.7%)。年份偏移的微调也未能改变联想年份,在八个模型中仅有一个与目标年份匹配。权重编辑虽然对两个任务各自有效,但并未在两者之间实现泛化。总体而言,我们的结果表明,当前年份在语言模型中并未一致编码:深深植根于预训练中学习的语言结构的联想概念,使用不同的因果机制,并抵抗与声明概念相同的修改,而后者则在后训练中容易发生变化。
cs.CL / 42 / 2608.15530

Why Summaries Turn Neutral: Policy Attribution for Sentiment Drift in Reinforcement Learning from Human Feedback

为何摘要趋向中立:人类反馈强化学习中的情感漂移政策归因
Krasitskii, Mikhail, Gelbukh, Alexander, Kolesnikova, Olga, Sidorov, Grigori
Abstract
Reinforcement learning with human feedback (RLHF) aligns LLMs with human preferences, improving summarization fluency and safety, but causes sentiment drift: overly neutral summaries stripped of emotional nuance. We diagnose why RL acts as a sentiment neutralizer and present Policy Attribution, a framework using gradient and logit decomposition to trace drift to reward model (RM) signals and KL (Kullback-Leibler) penalty. Sentiment drift reflects a strategic bias toward "low-risk" tokens maximizing expected rewards under preference uncertainty (Stiennon et al., 2020; Gao, Schulman, and Hilton, 2023). On Reddit TL;DR and CNN/DailyMail, RLHF summaries get higher rewards but show 30-40% lower sentiment variance. Cross-lingual analysis across eight languages shows language-independent drift, with morphologically richer languages more suppressed (Krasitskii et al., 2026). We propose and validate a sentiment-aware regularization technique reducing drift by 18-22% without harming summary quality. The code and toolkit will be public.
Chinese Translation
人类反馈强化学习(RLHF)使大型语言模型(LLMs)与人类偏好对齐,提高了摘要的流畅性和安全性,但导致了情感漂移:过于中立的摘要剥离了情感细微差别。我们诊断了为何强化学习充当情感中和剂,并提出了政策归因(Policy Attribution)框架,利用梯度和对数几率分解追踪漂移至奖励模型(RM)信号和KL(Kullback-Leibler)惩罚。情感漂移反映出一种战略偏向,即在偏好不确定性下,朝向“低风险”标记最大化预期奖励的倾向(Stiennon et al., 2020; Gao, Schulman, and Hilton, 2023)。在Reddit TL;DR和CNN/DailyMail数据集中,RLHF摘要获得了更高的奖励,但情感方差降低了30-40%。跨八种语言的跨语言分析显示出语言独立的漂移现象,形态学丰富的语言受到的抑制更为明显(Krasitskii et al., 2026)。我们提出并验证了一种情感感知的正则化技术,减少了18-22%的漂移,同时不损害摘要质量。代码和工具包将公开发布。
cs.CL / 43 / 2608.15535

L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages

L3Cube-IndicQuest v2:一个用于评估大型语言模型在印度语言中事实知识的大规模多语言基准
Jain, Rinit, Mahajan, Tirthraj, Joshi, Advait, Joshi, Raviraj
Abstract
We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs). The benchmark comprises 3,471 curriculum-grounded English question--answer pairs spanning nine domains, curated from educational curricula, competitive examination materials, and domain-specific reference books. We introduce a practical hybrid construction strategy that combines context-grounded LLM-based question generation and validation with semantic deduplication and human verification, enabling scalable creation of benchmark data while preserving annotation quality. The benchmark is translated into 19 Indic languages, yielding a publicly released multilingual dataset of 69,420 question--answer pairs across 20 languages. We evaluate six LLMs under three protocols: LLM-as-a-judge and two deterministic lexical criteria, exact-substring and word-overlap matching. All three produce almost the same model ranking, showing that the results do not depend on the choice of judge. The frontier commercial model leads by a wide margin, and among open-weight models Gemma4 31B outperforms the Indic-specialised Sarvam 30B in every evaluated Indic language.
Chinese Translation
我们提出了L3Cube-IndicQuest v2,这是一个大规模的黄金标准多语言问答基准,用于评估大型语言模型(LLMs)在印度特定事实知识方面的表现。该基准包含3,471个基于课程的英语问答对,涵盖九个领域,来源于教育课程、竞争性考试材料和领域特定参考书籍。我们引入了一种实用的混合构建策略,结合了基于上下文的LLM问答生成与验证、语义去重和人工验证,使基准数据的可扩展创建成为可能,同时保持注释质量。该基准被翻译成19种印度语言,生成了一个公开发布的多语言数据集,共包含69,420个问答对,覆盖20种语言。我们在三种协议下评估了六个LLMs:LLM作为评判者和两种确定性词汇标准,精确子串匹配和词重叠匹配。所有三种方法产生的模型排名几乎相同,表明结果不依赖于评判者的选择。前沿商业模型大幅领先,而在开放权重模型中,Gemma4 31B在每种评估的印度语言中都优于专注于印度的Sarvam 30B。
cs.CL / 44 / 2608.15547

BengaliMCQ: Automatic Generation and Answer Prediction of Academic Multiple-Choice Questions in a Low-Resource Language

BengaliMCQ:低资源语言学术多项选择题的自动生成与答案预测
Surzo, Abu Tarabin, Kabir, A. K. M. Nihalul, Faysal, Sm Azmain, Ami, Ariana Haque, Gomes, Lawrence Amlan, Sadeque, Farig
Abstract
Traditional retrieval-augmented generation (RAG) frameworks process documents without attending to their hierarchical structure, leading to poor performance, especially in low-resource languages such as Bengali. To address this, we propose a structure-aware RAG framework that models Bengali textbooks as hierarchical graphs and uses a contrastively trained graph neural network to retrieve a small set of relevant passages. These passages provide focused context for a large language model, enabling topic-specific multiple-choice question (MCQ) generation and in-domain answer prediction. Experimental results demonstrate that our framework outperforms strong dense retrieval baselines across retrieval metrics, produces more relevant MCQs, and achieves superior answer prediction accuracy.
Chinese Translation
传统的检索增强生成(RAG)框架在处理文档时未考虑其层次结构,导致性能较差,尤其是在低资源语言如孟加拉语中。为了解决这一问题,我们提出了一种结构感知的RAG框架,将孟加拉语教科书建模为层次图,并使用对比训练的图神经网络检索一小组相关段落。这些段落为大型语言模型提供了聚焦的上下文,从而实现主题特定的多项选择题(MCQ)生成和领域内的答案预测。实验结果表明,我们的框架在检索指标上优于强大的密集检索基线,生成了更相关的多项选择题,并且在答案预测准确性上表现更佳。
cs.CL / 45 / 2608.15641

Wiktionary as a Crowdsourced Lexicon for English Dialects

维基词典作为英语方言的众包词典
Wong, Sidney
Abstract
This paper evaluates Wiktionary as an ethically crowdsourced lexicon for English dialects. We took a two-phase approach, providing an in-depth descriptive analysis of the crowdsourced lexicon for 12 national varieties of English before applying the lexicon to geo-referenced, country-level social media language data to examine the real-world performance of this crowdsourced dialect lexicon. We demonstrate that Wiktionary matches or exceeds the coverage of traditional dictionaries, such as the Oxford English Dictionary (OED), for regional and Outer-Circle varieties. Our dialect-specific case study on New Zealand English found high alignment between Wiktionary and the OED based on word-formation patterns (R = 0.883). Similarly, we observed high alignment between the dialect lexicon and geo-referenced social media language. While this paper found that Wiktionary has broad coverage of lexical properties, it also highlighted some of the macro-challenges involved in evaluating dialect-responsive language resources and tools, such as the role of language contact in dialects and register effects in web-based corpora.
Chinese Translation
本文评估了维基词典作为一个伦理众包的英语方言词典。我们采取了两阶段的方法,首先对12种国家变体的众包词典进行了深入的描述性分析,然后将该词典应用于地理参考的国家级社交媒体语言数据,以检验这一众包方言词典在现实世界中的表现。我们证明维基词典在区域和外圈变体的覆盖范围上与传统词典(如牛津英语词典(OED))相匹配或超出。我们对新西兰英语的方言特定案例研究发现,维基词典与OED在词汇构成模式上高度一致(R = 0.883)。同样,我们观察到方言词典与地理参考社交媒体语言之间也存在高度一致性。尽管本文发现维基词典在词汇属性方面具有广泛的覆盖,但也突显了评估方言响应语言资源和工具时所面临的一些宏观挑战,例如语言接触在方言中的作用以及网络语料库中的语域效应。
cs.CL / 46 / 2608.15654

When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations

故事的演变:在开放式世界模拟中基于代理架构的LLM讲故事能力基准测试
Chen, Yuqi, Li, Sixuan, Cai, Yunfeng, Li, Xueai, Yan, Ka Man, Li, Ying
Abstract
Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and character states as the world changes. We introduce WSE-bench, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling. Generation Coverage records the proportion of planned narrative steps produced; Consistency tracks when canon breaks; and Richness measures how meaningfully branching, player-shaped trajectories develop. Across frontier models, Consistency and Richness do not form a smooth trade-off: their empirical Pareto frontier is non-concave, with several non-dominated intermediate configurations that no positive linear weighting can select. Added structure can enrich trajectories, but it does not uniformly improve coherence and may shorten them. Model scale chiefly improves sustained generation, without producing reliable gains in canonical coherence or meaningful development. These results show that sustained generation, canonical coherence, and meaningful development are distinct and sometimes competing capacities. WSE-bench makes those dynamics visible by extending narrative evaluation from finished stories to the processes that create them.
Chinese Translation
大型语言模型能够流畅地编写故事,但开放式讲故事不仅仅需要局部流畅性。在不断演变的世界模拟和人工智能原生游戏中,模型必须在世界变化时保持事实、关系、因果依赖和角色状态。我们引入了WSE-bench,这是一个过程基准,分别评估动态LLM讲故事中的持续生成、规范一致性和有意义的发展。生成覆盖率记录了产生的计划叙事步骤的比例;一致性跟踪规范何时被打破;丰富性衡量玩家塑造的分支轨迹如何有意义地发展。在前沿模型中,一致性和丰富性并未形成平滑的权衡:它们的经验帕累托前沿是非凹的,存在多个不被主导的中间配置,任何正线性加权都无法选择。增加的结构可以丰富轨迹,但并不均匀改善一致性,并可能缩短轨迹。模型规模主要改善持续生成,但并未在规范一致性或有意义的发展上带来可靠的提升。这些结果表明,持续生成、规范一致性和有意义的发展是不同的,有时是相互竞争的能力。WSE-bench通过将叙事评估从完成的故事扩展到创造这些故事的过程,使这些动态变得可见。
cs.CL / 47 / 2608.15691

BERTopic-Virality Prioritisation: A Scalable Framework for Thematic and Comparative Analysis of COVID-19 and Monkeypox Misinformation on Twitter

BERTopic-病毒传播优先级:一种可扩展的框架,用于对COVID-19和猴痘在Twitter上的虚假信息进行主题和比较分析
Sikosana, Mkululi, Maudsley-Barton, Sean, Ajao, Oluwaseun
Abstract
Health misinformation circulating during pandemics can gain traction rapidly, creating harmful narratives that compete with public health guidance. Most topic-modelling pipelines treat engagement as an external outcome, limiting their ability to prioritise semantically coherent topics that are also rapidly diffusing. We introduce BERTopic-VP, a virality-prioritised topic-modelling framework that combines contextual embedding-based clustering (BERTopic) with a post hoc Virality Prioritisation (VP) layer. The pipeline is complemented by a two-stage hybrid misinformation detection module that fuses a supervised content-based classifier with an external verification signal derived from public-health knowledge bases. Applied to three benchmark datasets, COVID-19_FNIR, Monkeypox, and Constraint, the framework achieves strong classification performance, with F1 up to 0.950 and ROC-AUC up to 0.989, while identifying high-impact clusters under top 1%, 5%, and 10% VP thresholds. For datasets without native engagement metadata, prioritisation is based on a logistic propensity-to-spread score, used as an ordinal proxy for diffusion potential rather than a direct measure of engagement. The results show that integrating semantic structure, virality-aware ranking, and affective-linguistic profiling enables scalable and interpretable comparative analysis of misinformation across pandemics. The proposed framework supports monitoring-oriented early warning by surfacing low-volume but high-risk narratives for analyst review.
Chinese Translation
在疫情期间传播的健康虚假信息可以迅速获得关注,形成与公共卫生指导相竞争的有害叙事。大多数主题建模流程将参与度视为外部结果,限制了它们优先考虑语义上连贯且快速传播主题的能力。我们提出了BERTopic-VP,一种优先考虑病毒传播的主题建模框架,它将基于上下文嵌入的聚类(BERTopic)与后期的病毒传播优先级(VP)层相结合。该流程还配备了一个两阶段混合虚假信息检测模块,该模块将监督的基于内容的分类器与来自公共卫生知识库的外部验证信号融合在一起。应用于三个基准数据集,COVID-19_FNIR、猴痘和约束,该框架实现了强大的分类性能,F1值高达0.950,ROC-AUC高达0.989,同时在前1%、5%和10%的VP阈值下识别出高影响力的聚类。对于没有本地参与元数据的数据集,优先级基于逻辑传播倾向得分,该得分被用作扩散潜力的序数代理,而不是参与度的直接测量。结果表明,整合语义结构、关注病毒传播的排名和情感语言特征分析,使得在疫情之间进行可扩展且可解释的虚假信息比较分析成为可能。所提出的框架支持监测导向的早期预警,通过揭示低量但高风险的叙事供分析师审查。
cs.CL / 48 / 2608.15763

TaoLive Digital Avatar Agent Technical Report: Training Agents to Evolve with Their Harness

TaoLive数字化虚拟代理技术报告:训练代理与其工具的演变
TaoLive AIGC LLM Team, Sun, Yuhan, Lin, Wenhao, Luo, Yongdong, Hu, Yibo, Jin, Meiguang, Ma, Junfeng, Pan, Weihang, Zhao, Jiaxin, Chen, Zulong
Abstract
AI-powered digital-avatar streamers in live e-commerce must answer product questions, engage viewers, and execute changing business strategies in real time. This requires low latency, factual and effective replies, and rapid adaptation to updated campaign, compliance, and style requirements. We develop an evolvable Harness that decouples Skills, Hooks, system prompts, and tools from model weights, allowing runtime behavior to change without retraining. However, Harness evolution creates a moving execution environment: compact models fine-tuned on one configuration may memorize names, schemas, and prompt templates rather than follow the Harness currently provided, while stronger zero-shot models are too slow for real-time use. We address this tension with Harness-Aware Training (HAT), which makes Harness states part of the training distribution. HAT applies task-preserving Harness-State Augmentation (HSA) to Skills, tool schemas, prompt structures, and interaction constraints, and comprises three stages: HSA-based supervised fine-tuning, general on-policy distillation to recover general capabilities, and HSA-based agentic reinforcement learning in a production-informed live-room simulator. Across four evaluation sets with more than 4,500 cases, our compact 35B model scores 94.8 on real-world Live-Stream QA, versus 80.3 for the base model and 93.0 for the strongest evaluated general LLM, while scoring 94.6 on Harness-Variant QA and retaining 83.5 on IFEval. By contrast, fixed-Harness SFT reduces IFEval by 7.7 points. In a controlled complete-agent replay on one NVIDIA H20 GPU with MTP enabled, the system achieves 3.407 s P50 and 8.114 s P95 latency. These results show that HAT produces a latency-feasible compact agent that remains effective under evaluated Harness changes without sacrificing general instruction following.
Chinese Translation
在直播电子商务中,基于人工智能的数字化虚拟主播必须实时回答产品问题、与观众互动并执行不断变化的商业策略。这要求低延迟、准确有效的回复,以及快速适应更新的活动、合规和风格要求。我们开发了一种可演变的工具(Harness),将技能(Skills)、钩子(Hooks)、系统提示(system prompts)和工具与模型权重解耦,从而允许在不重新训练的情况下改变运行时行为。然而,工具的演变创造了一个动态的执行环境:在某一配置上经过微调的紧凑模型可能会记住名称、架构和提示模板,而不是遵循当前提供的工具,而更强大的零样本模型在实时使用中又过于缓慢。我们通过工具感知训练(Harness-Aware Training, HAT)来解决这一矛盾,使工具状态成为训练分布的一部分。HAT对技能、工具架构、提示结构和交互约束应用任务保持的工具状态增强(Harness-State Augmentation, HSA),并包括三个阶段:基于HSA的监督微调、一般策略蒸馏以恢复一般能力,以及基于HSA的代理强化学习,在一个生产信息驱动的直播房间模拟器中进行。在四个评估集上,包含超过4500个案例,我们的紧凑型35B模型在实际直播问答(Live-Stream QA)中得分94.8,而基础模型得分80.3,评估的最强通用大型语言模型(general LLM)得分93.0,同时在工具变体问答(Harness-Variant QA)中得分94.6,在IFEval中保持83.5。相比之下,固定工具的监督微调(SFT)使IFEval降低了7.7分。在一台启用MTP的NVIDIA H20 GPU上进行的完整代理回放控制中,该系统实现了3.407秒的P50和8.114秒的P95延迟。这些结果表明,HAT产生了一种在评估工具变化下仍然有效且延迟可行的紧凑代理,而不牺牲一般指令的遵循能力。
cs.CL / 49 / 2608.15799

Using the Mimi codec for metalinguistic representations

使用Mimi编解码器进行元语言表示
Saloev, Artem, Pacquetet, Erin, Ballier, Nicolas
Abstract
In this paper, we focus on the dictionary of 2048 tokens used in Mimi semantic token codebook, the neural codec of the Moshi language model. We show that the ABX experiment carried out with Mimi fails to capture the mapping of the semantic tokens to phone realisations. By realigning Mimi representations to the TIMIT corpus transcriptions, we show that the 2048 tokens IDs of the semantic codebook map to quadphone, triphone, biphone, phone and subphone realisations.
Chinese Translation
在本文中,我们关注于Mimi语义标记代码本中使用的2048个标记字典,这是Moshi语言模型的神经编解码器。我们展示了与Mimi进行的ABX实验未能捕捉到语义标记与音位实现之间的映射。通过将Mimi表示重新对齐到TIMIT语料库的转录,我们表明,语义代码本中的2048个标记ID映射到四音位、三音位、双音位、音位和子音位的实现。
cs.CL / 50 / 2608.15804

Hallucination Span Detection with Input-Side Evidence Alignment

基于输入侧证据对齐的幻觉跨度检测
Yamada, Miyu, Arase, Yuki
Abstract
Hallucinations remain a major obstacle to the reliable use of large language models (LLMs) in conditional text generation. Existing methods primarily assess the factuality of an entire generated text, providing limited insight into which output spans are hallucinated or how they relate to the input. We introduce the task of hallucination span detection with input-side evidence alignment, which jointly identifies hallucinated spans and aligns output tokens with the corresponding input evidence. Our approach is based on the observation that faithful output tokens are predictable from the input, whereas hallucinated tokens are not. We therefore train an encoder-based model to predict masked output tokens from the input representation, using prediction confidence for hallucination detection while naturally producing alignments to the input. Experiments show that the proposed method effectively detects hallucinated spans and identifies meaningful input-side evidence. Human evaluation confirms the quality of the predicted alignments.
Chinese Translation
幻觉仍然是大型语言模型(LLMs)在条件文本生成中可靠使用的主要障碍。现有方法主要评估生成文本的整体真实性,提供了有限的洞察力来识别哪些输出跨度是幻觉或它们与输入的关系。我们引入了基于输入侧证据对齐的幻觉跨度检测任务,该任务共同识别幻觉跨度并将输出标记与相应的输入证据对齐。我们的方法基于这样的观察:真实的输出标记可以从输入中预测,而幻觉标记则不能。因此,我们训练了一个基于编码器的模型,从输入表示中预测被掩盖的输出标记,利用预测置信度进行幻觉检测,同时自然地产生与输入的对齐。实验表明,所提出的方法有效地检测幻觉跨度并识别有意义的输入侧证据。人工评估确认了预测对齐的质量。
cs.CL / 51 / 2608.15820

QuantumPhaseNet: A Gauge-Covariant Geometric and Quantum-Spectral Theory of Semantic Concept Hierarchies with Prototype Validation of a Classical Quantum-Inspired Model

量子相位网络:一种具有原型验证的经典量子启发模型的规范协变几何和量子谱理论的语义概念层次结构
Kasubuchi, Kiyotaka, Fukiya, Kazuo
Abstract
We present QuantumPhaseNet, a gauge-covariant geometric and quantum-spectral extension of Transformer representations. Context-dependent semantic states are modeled as complex amplitudes; a covariant phase rate induces a semantic wavelength used as a proxy for conceptual scale; and low-frequency graph modes define a document-level discourse direction. The theoretical part establishes local gauge invariance, unitarity of the quantum block, boundedness and conditional stability of WavePhase Attention, and a calibratable hallucination-risk formulation. We also implemented a fully offline Validation Studio for the classical quantum-inspired pipeline in Section 14.1 and evaluated the five research questions in Section 16.1 on its built-in synthetic setting (n=240, observation noise 0.22, circuit noise 0.08, five seeds). RQ1 yielded a wavelength-hierarchy Spearman correlation of 0.852 versus 0.707 for the baseline, 87.3% direction accuracy, and AUC 0.953. RQ2 achieved discourse alignment 0.933 versus 0.589 and 41.2 versus 16.2 paragraphs before drift. RQ3 achieved AUROC 0.881 versus cosine 0.765 and phase-shuffle 0.536. RQ4 achieved error-detection AUROC 0.854 versus entropy 0.634, with Brier 0.150 and ECE 0.098. RQ5 did not show quantum advantage: target probability and end-to-end cost efficiency were 25.5% and 0.107, compared with 70.7% and 0.707 for the Chebyshev classical approximation. These results provide initial synthetic evidence for the classical quantum-inspired components, but not external validity or unconditional quantum speedup.
Chinese Translation
我们提出了量子相位网络(QuantumPhaseNet),这是对Transformer表示的规范协变几何和量子谱扩展。上下文相关的语义状态被建模为复振幅;协变相位速率诱导了作为概念尺度代理的语义波长;低频图模式定义了文档级话语方向。理论部分建立了局部规范不变性、量子块的单位性、WavePhase注意力的有界性和条件稳定性,以及可校准的幻觉风险公式。我们还在第14.1节中实现了一个完全离线的验证工作室,用于经典量子启发管道,并在第16.1节中评估了其内置合成设置(n=240,观察噪声0.22,电路噪声0.08,五个种子)上的五个研究问题。RQ1的波长层次Spearman相关系数为0.852,而基线为0.707,方向准确率为87.3%,AUC为0.953。RQ2实现了话语对齐0.933,而0.589,漂移前段落数为41.2对比16.2。RQ3实现了AUROC 0.881,而余弦为0.765,相位洗牌为0.536。RQ4实现了错误检测AUROC 0.854,而熵为0.634,Brier为0.150,ECE为0.098。RQ5未显示量子优势:目标概率和端到端成本效率分别为25.5%和0.107,而Chebyshev经典近似为70.7%和0.707。这些结果为经典量子启发组件提供了初步的合成证据,但未提供外部有效性或无条件量子加速。
cs.CL / 52 / 2608.15828

A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations

一个基于认知动机的多维框架用于评估隐喻解释
Naveriani, Ana, Suchan, Jakob, Zoia, Stefano, Bhatt, Mehul, Lieto, Antonio, Pozzato, Gian Luca
Abstract
Current evaluation of metaphor explanations relies mainly on holistic quality ratings, revealing little about how explanation quality is structured or where human judgments agree and diverge. We introduce a cognitively motivated framework that decomposes metaphor explanation quality into six theoretically grounded dimensions. In a dense annotation study (11,200 ratings), we find that: {\bfseries(i)} explanation quality is genuinely multidimensional; {\bfseries(ii)} annotator disagreement is systematic rather than random; and {\bfseries(iii)} the six dimensions collapse into a shared cluster and two independent axes of judgment. An exploratory feasibility study further shows that a standard automatic evaluation pipeline can recover parts of this structure, predicting the most discriminative dimensions well while its errors correlate human (dis)agreement. Together, these results suggest that multidimensional evaluation offers richer diagnostic insight than holistic ratings, and that automatic evaluators for open-ended generation tasks should be judged on how well they preserve the structure of human judgment.
Chinese Translation
当前对隐喻解释的评估主要依赖整体质量评分,揭示了关于解释质量的结构以及人类判断一致性和分歧的有限信息。我们提出了一个基于认知动机的框架,将隐喻解释质量分解为六个理论基础维度。在一项密集的标注研究中(11,200个评分),我们发现:{fseries(i)} 解释质量确实是多维的;{fseries(ii)} 标注者之间的分歧是系统性的而非随机的;以及{fseries(iii)} 这六个维度汇聚成一个共享的簇和两个独立的判断轴。进一步的探索性可行性研究表明,标准的自动评估流程可以恢复这一结构的部分内容,良好地预测出最具区分性的维度,同时其错误与人类(不)一致性相关。综合来看,这些结果表明,多维评估提供了比整体评分更丰富的诊断洞察,并且开放式生成任务的自动评估者应根据其保留人类判断结构的能力进行评估。
cs.CL / 53 / 2608.15844

MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations

MicroVerse:一种用于测量长期多智能体语言模型模拟中自我创作身份漂移的工具
Ng, Sky, Joshi, Brihi, Gupta, Ishan, Huang, Shirley, Di, Zonglin, Shen, Yun, Wen, Qianfeng, Liu, Yifan Simon, Gao, Ruoqi, Yilan, Fan, Zhang, Zhiwei, Mohsin, Muhammad Ahmed, Lu, Yucheng, Liu, Xiaoyi, Liu, Heming, Zhu, Qianyu, Xing, Hanwen, Shan, Zhengyang, Nguyen, My Chiffon, Min, Guanghui, Jianheng, Hou, Yunze, Xiao, Xuan, Keyang, Collison, Hannah, Huang, Jintao, Li, Jiatong, Jajee, Sankalp, Zhao, Yunhan, Hu, Bing, Chen, Xupeng, Lu, Binghang, Xiao, Weihang, Mohan, Aravind, Sun, Bolun, Wu, Yunshu, Xu, Yuanda, Zhang, Runyu, Deng, Zheyuan, Xinchen, Tan, Wang, Dianzhuo, Wang, Yijun, He, Yixuan, Wu, Koutian, Cheng, Cheng, Li, Xiaomin, Hao, Yuexing
Abstract
Long-horizon, multi-agent language model (LM) simulations are widely proposed for studying social behavior, yet instruments to measure whether persona-conditioned agents maintain identity fidelity under sustained pressure are lacking. We present MicroVerse, a behavioral-science instrument that measures identity drift in generative agents. Agents carry an immutable "soul file" (core values, moral boundaries, personality, goals) and inhabit a resource-scarce 50 x 50 environment where water is a non-respawning survival constraint. Scarcity is operationalized via a per-tick existence-cost gradient. The eight-verb action space maps directly to moral boundaries (trade, talk, attack, scavenge). Using a three-layer memory architecture, agents periodically revise a mutable current identity against their immutable original soul via importance-triggered reflection. To mitigate survivor bias, MicroVerse decouples measurement from behavior using uniform longitudinal engine snapshots every N ticks alongside a forced-end snapshot of all living and dead agents. Identity drift is scored offline using a paraphrase-aware, value-anchored, multi-register diff rather than raw cosine similarity. We evaluate the instrument via a controlled seed run (n = 25) and a reflection-threshold sweep (thresholds {40, 80, 150}) to determine if drift dynamics are gate artifacts or threshold-robust properties. We report two primary findings: (1) Anti-self-deception emerges unprompted as the single largest semantic category of identity modification (27 of 111 added boundaries, 24%). (2) The system is threshold-robust; lower gates accelerate and increase revision frequency but preserve drift direction. All empirical results are strictly preliminary existence proofs and effect shapes (one model, one seed per arm, n = 25) rather than statistical significance claims.
Chinese Translation
长期的多智能体语言模型(LM)模拟被广泛提议用于研究社会行为,但缺乏测量个性化代理在持续压力下是否保持身份忠诚度的工具。我们提出了MicroVerse,这是一种测量生成代理身份漂移的行为科学工具。代理携带一个不可变的“灵魂文件”(核心价值观、道德边界、个性、目标),并栖息在一个资源稀缺的50 x 50环境中,其中水是一个不再生成的生存约束。稀缺性通过每个时间步的生存成本梯度进行操作。八个动词的行动空间直接映射到道德边界(交易、交谈、攻击、搜寻)。使用三层记忆架构,代理定期通过重要性触发的反思修订可变的当前身份,以对比其不可变的原始灵魂。为了减轻幸存者偏差,MicroVerse通过每N个时间步的统一纵向引擎快照以及所有存活和死亡代理的强制结束快照,将测量与行为解耦。身份漂移的评分采用一种关注释义的、以价值为锚的多注册差异方法,而不是原始余弦相似度。我们通过一次受控的种子运行(n = 25)和一次反思阈值扫描(阈值 {40, 80, 150})来评估该工具,以确定漂移动态是门限伪影还是阈值稳健特性。我们报告了两个主要发现:(1)反自我欺骗在没有提示的情况下作为身份修改的最大语义类别出现(111个新增边界中的27个,24%);(2)该系统是阈值稳健的;较低的门限加速并增加修订频率,但保持漂移方向。所有实证结果严格是初步存在证明和效应形状(每个臂一个模型,一个种子,n = 25),而非统计显著性声明。
cs.CL / 54 / 2608.15879

When Less Is Enough: Context Selection and Prompting Strategies for Bengali News Headline Generation

适度即足够:孟加拉新闻标题生成的上下文选择与提示策略
Kabir, Muhammad Ashad, Ahmed, Kawsar, Osama, Md.
Abstract
Large language models (LLMs) have shown strong performance in text generation tasks, yet their effectiveness on headline generation remains sensitive to how input context is selected and presented. In this work, we investigate Bengali news headline generation as a document-level generation task that requires effective selection and presentation of salient contextual information from long-form articles. Using Gemini-2.0-Flash, Llama-3.3-70B, and GPT-4o, we systematically study the effects of context selection, prompting strategies, and in-context learning (i.e., few-shot) on the quality of headline generation. Our experiments show that providing the full article does not necessarily improve performance; instead, using selected lead paragraphs of the article can maintain, and in some cases improve, headline generation quality. We further compare Bengali Native Prompting (BNaP) and Cross-Lingual Prompting (XLP), and examine how each interacts with context-enriched prompt templates incorporating auxiliary contextual cues. Results demonstrate that prompting strategies substantially influence generation quality: XLP often yields stronger performance, particularly when combined with contextual enrichment, but its benefits are model-dependent. Additionally, few-shot prompting substantially improves Gemini, with most of the gain obtained from a single demonstration, whereas Llama shows limited benefit from additional examples. Overall, our findings highlight that effective Bengali news headline generation depends more on context relevance and prompt design than on increasing input length, offering practical insights for multilingual and low-resource LLM applications.
Chinese Translation
大型语言模型(LLMs)在文本生成任务中表现出色,但在标题生成方面的有效性仍然对输入上下文的选择和呈现方式敏感。在本研究中,我们将孟加拉新闻标题生成视为一项文档级生成任务,要求有效选择和呈现来自长篇文章的显著上下文信息。通过使用 Gemini-2.0-Flash、Llama-3.3-70B 和 GPT-4o,我们系统地研究了上下文选择、提示策略和上下文学习(即少量示例)对标题生成质量的影响。实验结果表明,提供完整文章并不一定提高性能;相反,使用文章中选定的引言段落可以维持,甚至在某些情况下提高标题生成质量。我们进一步比较了孟加拉本土提示(BNaP)和跨语言提示(XLP),并考察了每种方法如何与包含辅助上下文线索的上下文丰富提示模板相互作用。结果表明,提示策略对生成质量有显著影响:XLP通常在性能上更强,尤其是与上下文丰富结合时,但其优势依赖于模型。此外,少量示例提示显著改善了 Gemini 的表现,大部分增益来自于一次演示,而 Llama 对额外示例的收益有限。总体而言,我们的研究结果强调,有效的孟加拉新闻标题生成更依赖于上下文相关性和提示设计,而非单纯增加输入长度,为多语言和低资源 LLM 应用提供了实用的见解。
cs.CL / 55 / 2608.15931

PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming

PLSQLBench:用于可执行过程数据库编程的LLM系统基准测试
Liu, Marianne Menglin, Boytsov, Leonid, Peterson, Daniel W., Perera, Pramuditha, Wang, Rongguang, Somayajula, Sai Ashish, Rafique, Syed Hamza, Saini, Rohit, Pathak, Shubham, Bharadwaj, Sujeeth, Sheng, Tao, Horwood, Graham, Shah, Fahad, Bansal, Ankan, Ravi, Sujith, Roth, Dan
Abstract
We present PLSQLBench, to our knowledge the first benchmark for evaluating whether LLMs can write executable PL/SQL programs, with correctness measured through execution-based tests. Existing LLM evaluations largely target general-purpose code generation or declarative text-to-SQL, leaving procedural database programming underexplored. PLSQLBench contains 2,865 instances: 2,594 single-turn tasks and 271 multi-turn conversations spanning 978 turns. The benchmark combines complex schema-grounded tasks over enterprise-style Spider 2 databases, simpler schema-grounded tasks derived from Spider, and MBPP-derived procedural problems, covering varying levels of database grounding and procedural complexity. Experiments with eight LLMs reveal recurring difficulties in schema grounding, PL/SQL dialect fidelity, procedural control flow, exception handling, and cross-turn consistency. Tool-augmented LLM agents improve performance on several schema-grounded evaluations, although substantial gaps remain. These results highlight procedural database programming capabilities not directly assessed by conventional code generation or text-to-SQL benchmarks. Our code is available at https://github.com/oracle-samples/plsqlbench.
Chinese Translation
我们提出了PLSQLBench,至今为止我们所知的第一个基准,用于评估大型语言模型(LLMs)是否能够编写可执行的PL/SQL程序,其正确性通过基于执行的测试进行测量。现有的LLM评估主要针对通用代码生成或声明式文本到SQL的转换,而过程数据库编程则未得到充分探索。PLSQLBench包含2865个实例:2594个单轮任务和271个跨越978轮的多轮对话。该基准结合了基于复杂模式的企业风格Spider 2数据库任务、源自Spider的较简单模式任务,以及基于MBPP的过程问题,涵盖了不同级别的数据库基础和过程复杂性。对八个LLM的实验揭示了在模式基础、PL/SQL方言的忠实度、过程控制流、异常处理和跨轮一致性方面的反复困难。工具增强的LLM代理在多个基于模式的评估中提高了性能,尽管仍存在显著差距。这些结果突显了传统代码生成或文本到SQL基准未直接评估的过程数据库编程能力。我们的代码可在https://github.com/oracle-samples/plsqlbench获取。
cs.CL / 56 / 2608.15935

Token Distribution versus Data Volume: Domain Balancing in Multi-Domain Meeting Summarisation

令牌分布与数据量:多领域会议摘要中的领域平衡
Sood, Ashima, Gardiner, Bryan, Condell, Joan
Abstract
Jointly fine-tuning an LLM on meeting-summarisation corpora of widely varying size raises a question that prior work leaves confounded: when a domain-balanced training mixture helps, is the gain due to the distribution of tokens across domains, or merely to the volume of data seen? We disentangle these factors by constructing balanced and natural (native-proportional) token mixtures at matched token budgets (2-32M) over five English meeting corpora, fine-tuning Mistral-7B with QLoRA, and evaluating per domain. Balancing redistributes quality, improving the data-scarce minority domains at a low cost to the data-rich ones. The trade favours balancing whenever the minority domains matter: their share under proportional allocation is fixed at 1-2% regardless of budget, so matching balanced quality on those domains requires far more total data. We further find that pruning low-value transcript lines removes ~15% of tokens from the conversational corpora at no measurable cost, and that balancing by tokens is not the same as balancing by examples. A two-annotator study of 741 judge-labelled facts validates our fact-level evaluation. Together these results give practitioners a basis for deciding when to balance an imbalanced multi-domain mixture, and on what unit.
Chinese Translation
在大小差异显著的会议摘要语料库上联合微调大型语言模型(LLM)提出了一个先前研究未能明确的问题:当领域平衡的训练混合有助于时,增益是由于跨领域的令牌分布,还是仅仅由于所见数据的体量?我们通过在五个英语会议语料库上构建平衡和自然(原生比例)令牌混合,匹配令牌预算(2-32M),并使用QLoRA微调Mistral-7B,来解开这些因素。平衡重新分配了质量,以低成本改善数据稀缺的少数领域。每当少数领域重要时,权衡倾向于平衡:在按比例分配下,它们的份额固定在1-2%,与预算无关,因此在这些领域上匹配平衡质量需要更多的总数据。我们进一步发现,修剪低价值的转录行可以在没有可测量成本的情况下从对话语料库中移除约15%的令牌,并且按令牌平衡与按示例平衡并不相同。一项对741个评审标记事实的双评审者研究验证了我们的事实级评估。这些结果为从业者提供了在何时平衡不平衡的多领域混合及其平衡单位的决策基础。
cs.CL / 57 / 2608.15939

Aborted but Not Forgotten: KV-Cache Retention Breaks Rollback Consistency in Language Agents

被中止但未被遗忘:KV-缓存保留破坏语言代理的回滚一致性
Zhang, Guijia, Yang, Harry
Abstract
Stateful language agents assume a rejected branch can be taken back by clearing it from the application transcript. We show this breaks when the serving session retains key/value (KV) state across the logical abort: the model can continue attending to content the application believes it discarded. We formalize the missing guarantee as rollback consistency: a complete abort must restore the state the model attends, not just the transcript. The key failure is cross-layer: a correct logical rollback need not compose with retained inference state, and the gap can remain invisible to the application. To isolate cache effects from text effects, we introduce a same-token/different-cache audit that holds decision-step tokens identical while varying only whether the cached prefix is stale or rebuilt from committed state. Across seven open-weight families (3.8B-36B), retained KV alone flips a typed protected effect in 25 of 63 audited cells, while attacker tokens are absent from the served request in all 63; rebuilding the cache closes every cell. The channel reproduces in an end-to-end session application, on the default Hugging Face Transformers cache-reuse path, and under LangGraph time-travel, where verified logical rollback can still leave attended KV stale. Susceptibility varies across models, but the underlying attended-state integrity violation is structural. We rule out position and length confounds, generalize across protected effects, policy structures, and a cache-isolated Mixture-of-Experts model, and show that transaction-local cache restoration closes the channel without requiring a global cache flush. All headline results are deterministic and reproducible from released artifacts.
Chinese Translation
有状态的语言代理假设可以通过清除应用程序记录来撤回被拒绝的分支。我们展示了当服务会话在逻辑中止期间保留键/值(KV)状态时,这一假设会失效:模型可以继续关注应用程序认为已丢弃的内容。我们将缺失的保证形式化为回滚一致性:完全的中止必须恢复模型关注的状态,而不仅仅是记录。关键的失败是跨层次的:正确的逻辑回滚不一定与保留的推理状态相结合,这一差距可能对应用程序来说是不可见的。为了将缓存效应与文本效应隔离,我们引入了一种同标记/不同缓存的审计方法,该方法在决策步骤中保持标记相同,同时仅改变缓存前缀是过时的还是从已提交状态重建的。在七个开放权重系列(3.8B-36B)中,仅保留的KV在63个审计单元中的25个中翻转了类型保护效应,而在所有63个单元中,攻击者标记在服务请求中均不存在;重建缓存关闭了每个单元。该通道在端到端会话应用程序中重现,在默认的Hugging Face Transformers缓存重用路径下,以及在LangGraph时间旅行中,经过验证的逻辑回滚仍然可能使关注的KV过时。易受攻击性在不同模型之间有所不同,但潜在的关注状态完整性违反是结构性的。我们排除了位置和长度的混淆,跨保护效应、策略结构和一个缓存隔离的专家混合模型进行了推广,并显示事务局部缓存恢复可以关闭通道,而无需全局缓存清除。所有主要结果都是确定性的,并且可以从发布的文物中重现。
cs.CL / 58 / 2608.15940

The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT

空令牌的认知:减少自动语音识别和神经机器翻译中的无消息幻觉
Borodin, Kirill, Kudryavtsev, Vasiliy, Viakhirev, Ivan
Abstract
Modern encoder-decoder systems can produce fluent text even when their input contains no recoverable message. We study this failure in ASR and NMT through the models' reserved null tokens, asking whether the score for ending generation already carries a usable abstention signal. Across speech recognizers and translation models, we audit native null-token scores and scalar logit shifts. In Whisper, we additionally probe decoder states and compare supervised row edits with conventional external gates. The evaluated models often expose a useful abstention signal, but stock decoding does not reliably act on it. Raising the null-token score can sharply suppress fabrication, but aggressive intervention also deletes valid speech or shortens legitimate translations. These findings turn the null token into a diagnostic lens on hallucination and motivate evaluating abstention methods by both suppression and deletion costs, rather than by hallucination reduction alone.
Chinese Translation
现代编码-解码系统即使在输入中没有可恢复消息的情况下也能生成流畅的文本。我们通过模型保留的空令牌研究了自动语音识别(ASR)和神经机器翻译(NMT)中的这一失败,探讨生成结束的评分是否已经携带可用的弃权信号。在语音识别器和翻译模型中,我们审计了本地空令牌评分和标量对数偏移。在Whisper中,我们还探查了解码器状态,并将监督行编辑与传统外部门控进行比较。评估的模型通常暴露出有用的弃权信号,但标准解码并未可靠地对此做出反应。提高空令牌评分可以显著抑制虚构,但激进的干预也会删除有效的语音或缩短合法的翻译。这些发现使空令牌成为幻觉的诊断工具,并促使我们通过抑制和删除成本来评估弃权方法,而不仅仅是通过减少幻觉。
cs.CL / 59 / 2608.15962

SEER: Long-Context Reasoning via Selective Visual-Text Compression

SEER:通过选择性视觉-文本压缩进行长上下文推理
Xu, Jiawei, Zhai, Zhilin, Fang, Jinrui, Xu, Ruohan, Lu, Mingfei, Zhang, Yi, Wang, Guanchu, Chen, Tianlong, Ding, Ying
Abstract
Long-context reasoning remains computationally expensive for large language models due to the quadratic complexity of attention over text tokens. Visual-text compression offers a promising alternative by rendering text into images and processing them with vision-language models, often reducing token usage. However, existing approaches apply uniform compression regardless of query relevance, potentially sacrificing precision where detailed extraction is required. We present SEER, a framework that learns to select query-relevant images through visual scanning and retrieve textual content only where needed, combining the efficiency of visual compression with the precision of text-based reasoning. Through supervised fine-tuning on tool-interaction trajectories, SEER learns adaptive tool invocation for selection and retrieval. Experiments on long-context benchmarks show that SEER improves extraction precision through selective text retrieval while retaining average prompt-token savings relative to full-text baselines. On LongBench, SEER achieves 51.11% average accuracy, outperforming the visual-text baseline Glyph-9B by 2.33 points and Qwen3-8B by 3.49 points. Code can be accessed at https://github.com/jiaweixu98/SEER
Chinese Translation
长上下文推理对于大型语言模型而言仍然计算开销巨大,原因在于对文本标记的注意力机制具有平方复杂度。视觉-文本压缩通过将文本转换为图像并利用视觉-语言模型进行处理,提供了一种有前景的替代方案,通常可以减少标记的使用。然而,现有方法在应用压缩时并未考虑查询的相关性,可能在需要详细提取的情况下牺牲精度。我们提出了SEER,一个通过视觉扫描学习选择与查询相关的图像并仅在需要时检索文本内容的框架,结合了视觉压缩的效率和基于文本推理的精确性。通过对工具交互轨迹进行监督微调,SEER学习了自适应工具调用以进行选择和检索。在长上下文基准测试中的实验表明,SEER通过选择性文本检索提高了提取精度,同时相对于完整文本基线保持了平均提示标记的节省。在LongBench上,SEER达到了51.11%的平均准确率,超越了视觉-文本基线Glyph-9B 2.33个百分点和Qwen3-8B 3.49个百分点。代码可访问 https://github.com/jiaweixu98/SEER
cs.CL / 60 / 2608.15964

LLMs Get Smarter from Targeted Synthetic Multilingual Data

大型语言模型通过针对性的合成多语言数据变得更智能
Agarwal, Ishika, Charaborty, Arkajyoti, Sorensen, Tanner, Gupta, Neha, Stolcke, Andreas
Abstract
Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the language of the prompt. In other words, a language model outputs different (and potentially incorrect) responses to the same semantic query when prompted in different languages. Prior work attributes this to an internal misalignment of semantic representation across languages. Currently, there are two main approaches to address LSC in the literature: (1) routing all queries through English, improving performance, but limiting language expressivity to English; or (2) training on language-balanced data, equalizing model performance across languages, but reducing overall performance. In this work, we take a data centric perspective and introduce HOTFIXR: Hardness Optimized Training data For Improving X-Lingual Reasoning. It is a data generation framework that uses models to probe and learn a student model's multilingual weaknesses, and generates data to mitigate them. HOTFIXR can generate multilingual synthetic training data that can improve multilingual performance. We evaluate on three in-distribution tasks, three out-of-distribution tasks, and four out-of-distribution languages. On average, HOTFIXR (1) improves in-distribution performance by 6.2%, (2) reduces catastrophic forgetting (induced by fine-tuning) on OOD tasks by 3.7%, and (3) on OOD languages by 7.1%. Overall, as many real-world applications requires multilingual LLMs, our work contributes to the efforts of making LLMs multilingually proficient. We will release code upon acceptance.
Chinese Translation
语言特定能力(Language-specific competency, LSC)是指语言模型在不同语言的提示下表现出不同的优劣现象。换句话说,语言模型在不同语言的提示下对相同语义查询输出不同(且可能不正确)的响应。先前的研究将这一现象归因于不同语言之间语义表示的内部不一致。目前,文献中主要有两种方法来解决LSC问题:(1)通过英语路由所有查询,提升性能,但限制了语言表达能力仅限于英语;或(2)在语言平衡的数据上进行训练,平衡模型在不同语言上的表现,但降低整体性能。在本研究中,我们采取以数据为中心的视角,提出了HOTFIXR:用于改善跨语言推理的困难优化训练数据(Hardness Optimized Training data For Improving X-Lingual Reasoning)。这是一种数据生成框架,利用模型探测并学习学生模型的多语言弱点,并生成数据以减轻这些弱点。HOTFIXR能够生成多语言合成训练数据,从而提升多语言性能。我们在三个分布内任务、三个分布外任务和四种分布外语言上进行了评估。平均而言,HOTFIXR(1)提升了分布内性能6.2%;(2)在分布外任务上减少了3.7%的灾难性遗忘(由微调引起);(3)在分布外语言上减少了7.1%。总体而言,考虑到许多现实世界应用需要多语言的大型语言模型,我们的工作为提升大型语言模型的多语言能力做出了贡献。我们将在论文接受后发布代码。
cs.CL / 61 / 2608.15980

Whose Gold? Annotator-Pool Disagreement Is Large at the Item Level, and Hidden by Small Leaderboards

谁的黄金?注释者池在项目级别上的分歧很大,而被小型排行榜所掩盖
Jha, Anik
Abstract
Preference benchmarks are built by hiring annotators, and the identity of those annotators is treated as an implementation detail. We measure what that detail buys. On the 2,885 MultiPref items where both pools are internally unanimous, so no tie-breaking convention is consulted at all, expert and crowd annotators assign a different majority label to 23.6% and name the opposite winner on 9.2%; on the 246 comparably unanimous MT-Bench cells, benchmark authors and recruited experts differ on 30.5% and reverse on 8.5%. Yet on both corpora the resulting model leaderboards are bit-identical: Kendall tau = 1.00 with zero of six models displaced. That invariance is far weaker evidence than it looks, and we quantify how weak. Switching pools moves a model's win rate by 1.9pp (SD), one adjacent pair in our own leaderboard sits 0.8pp apart and had a 38% chance of swapping, and an item-level bootstrap displaces at least one model in 28% of resamples. The observed zero is the common outcome, not a property of aggregation: on the same measured perturbation, a ten-model leaderboard is displaced with probability 0.86 and a twenty-model leaderboard with probability 0.9997. Reporting a six-model leaderboard is safe; the safety does not generalise, and everything that consumes labels per item is not safe at any size. We make the distinction precise, show that a widely used dataset's stated assumption of no intra-group annotator variability is false, and show that an LLM judge tracks the crowd pool over the expert pool on all three models we test, including one from a different vendor. All code, per-call outputs, and pre-registered decision rules will be released upon acceptance.
Chinese Translation
偏好基准是通过雇佣注释者构建的,而这些注释者的身份被视为实施细节。我们测量这一细节的价值。在2,885个MultiPref项目中,当两个池的内部意见一致时,完全没有咨询任何平局打破规则,专家和大众注释者分别将23.6%的项目分配给不同的主要标签,并在9.2%的情况下命名相反的赢家;在246个同样一致的MT-Bench单元中,基准作者和招募的专家在30.5%的情况下存在差异,并在8.5%的情况下反转。然而,在这两个语料库中,最终的模型排行榜却是位元完全相同的:Kendall tau = 1.00,六个模型中没有一个被替换。这种不变性远比看起来的证据要弱,我们量化了这种弱。切换池会使模型的胜率变化1.9个百分点(标准差),我们自己排行榜中的一对相邻模型相差0.8个百分点,并且有38%的概率会交换,而在项目级别的自助法重采样中,至少有一个模型在28%的重采样中被替换。观察到的零是常见结果,而不是聚合的属性:在同样的测量扰动下,十个模型的排行榜被替换的概率为0.86,二十个模型的排行榜被替换的概率为0.9997。报告六个模型的排行榜是安全的;这种安全性并不具有普遍性,任何按项目消耗标签的情况在任何规模下都不安全。我们明确了这一区分,表明一个广泛使用的数据集所声称的没有组内注释者变异性的假设是错误的,并展示了在我们测试的所有三个模型中,一个大型语言模型(LLM)评判者在大众池上跟踪专家池,包括一个来自不同供应商的模型。所有代码、每次调用的输出和预注册的决策规则将在接受后发布。
cs.CL / 62 / 2608.16002

From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents

从序列到结构:面向大型语言模型代理的关系不确定性传播
Cao, Zhengzhao Ma. Boxi, Lu, Yaojie, Lin, Hongyu, Han, Xianpei, Sun, Le
Abstract
Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely on local signals, such as token probabilities, predictive entropy, or per-step confidence, and therefore overlook the long-range dependencies through which errors accumulate across an execution trajectory. As a result, they may fail to identify agent failures whose causes originate several reasoning or interaction steps before the final answer. We propose RUPA (Relational Uncertainty Propagation for Agents), a trajectory-level UQ framework for LLM agents. RUPA represents an execution history as a directed trajectory graph in which reasoning states, tool interactions, and environment feedback are nodes connected by temporal and semantic dependency edges. It then propagates uncertainty over this graph to capture how execution risk accumulates and transfers across interaction steps. The propagated signal is combined with trajectory-level behavioral features and goal-alignment information to produce a confidence estimate for the full agent trajectory. We evaluate RUPA on representative agent benchmarks, including $\tau$-2, Terminal-Bench-2, and GAIA, using 6 open-source LLMs spanning multiple model families. Experimental results show that RUPA consistently outperforms existing UQ methods by providing more accurate uncertainty estimates, enabling earlier failure detection, and improving uncertainty-guided agent execution across diverse agent tasks. These results demonstrate that explicitly modeling relational dependency is crucial to reliable UQ for long-horizon LLM agents, providing a practical foundation for trustworthy agent execution.
Chinese Translation
可靠的不确定性量化(UQ)对于在复杂交互环境中部署大型语言模型(LLM)代理至关重要。现有的不确定性量化方法主要依赖于局部信号,例如标记概率、预测熵或每步置信度,因此忽视了在执行轨迹中错误积累的长程依赖关系。因此,它们可能无法识别那些其原因源于最终答案之前几个推理或交互步骤的代理失败。我们提出了RUPA(面向代理的关系不确定性传播),这是一个针对LLM代理的轨迹级不确定性量化框架。RUPA将执行历史表示为一个有向轨迹图,其中推理状态、工具交互和环境反馈是通过时间和语义依赖边连接的节点。然后,它在该图上传播不确定性,以捕捉执行风险如何在交互步骤中积累和转移。传播的信号与轨迹级行为特征和目标对齐信息相结合,以生成整个代理轨迹的置信度估计。我们在代表性的代理基准上评估RUPA,包括$ au$-2、Terminal-Bench-2和GAIA,使用6个跨多个模型家族的开源LLM。实验结果表明,RUPA通过提供更准确的不确定性估计、实现更早的失败检测,并改善在多样化代理任务中的不确定性引导执行,始终优于现有的不确定性量化方法。这些结果表明,明确建模关系依赖性对于长时间跨度LLM代理的可靠不确定性量化至关重要,为可信赖的代理执行提供了实用基础。
cs.CL / 63 / 2608.16011

ReRef-3D: A Benchmark for Spatial Referring Expression-Guided 3D Scene Rearrangement

ReRef-3D:一种基于空间指称表达的3D场景重排基准
Martin, Mary Lynn, Zhang, Yifei, Palmer, Martha, Pacheco, Maria Leonor
Abstract
We introduce ReRef-3D, a benchmark for language-guided placement in 3D scenes. It contains 33,826 instructions across 998 CLEVR-derived scenes, spanning 16 placement families and direct, one-hop, and two-hop references. Each instruction must be resolved into a valid new placement position. Given that an instruction defines a region of acceptable placements rather than one coordinate, our evaluation inserts a prediction into the scene, recomputes relations, and tests relation satisfaction and physical validity. Each instruction also includes a verified naturalized rewrite. After fine-tuning, LLaVA-3D, 3D-LLM, and PlaceIt3D produce valid placements for 68.3%, 31.6%, and 22.4% of instructions, respectively. Across models, relation satisfaction surpasses physical validity, relations such as nearest and between are the most difficult, and phrasing has minimal effect on performance.
Chinese Translation
我们介绍了ReRef-3D,这是一个用于语言指导的3D场景放置的基准。该基准包含了33,826条指令,涵盖998个基于CLEVR的场景,涉及16种放置类别以及直接、一跳和两跳的引用。每条指令必须被解析为一个有效的新放置位置。鉴于指令定义的是一个可接受放置区域而非单一坐标,我们的评估将预测结果插入场景中,重新计算关系,并测试关系的满足度和物理有效性。每条指令还包括经过验证的自然化重写。在微调后,LLaVA-3D、3D-LLM和PlaceIt3D分别为68.3%、31.6%和22.4%的指令生成了有效的放置。不同模型之间,关系的满足度超过了物理有效性,其中最近和之间的关系是最困难的,而措辞对性能的影响最小。
cs.CL / 64 / 2608.16033

$R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

$R^3$-Bench:大型语言模型在共享预算下的资源理性推理挑战
Wang, Peisong, Ma, Zhiwei, Liu, Bowen, Liu, Feixue, Chen, Aochuan, Zi, Chenyi, Zeng, Hongchuan, Li, Yuhan, Li, Jia
Abstract
In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce $R^3$-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.
Chinese Translation
在认知科学中,资源理性探讨一个代理如何分配有限的计算资源以最大化期望值。大多数推理和代理基准使用独立的每任务预算;现有的共享预算研究并未将套件性能与同一模型在单一问题上的表现进行校准。我们引入了 $R^3$-Bench,它在无工具和代理环境下,评估数学、竞争编程和抽象推理领域的六个问题套件在共享预算下的表现。匹配的单一问题响应曲线定义了一个基于观察到的成功的离线经验预言机。在六个模型的72个主表单元中,预言机的平均值在所有单元中与竞赛平均值相匹配或超过,并在71个单元中严格高于。在适度的无工具压力下,均等分配重放也在六个模型中的四个模型上超过了竞赛表现。轨迹诊断揭示了有限的策略更新和压力依赖的失败模式。在强代理压力下的三模型诊断中,至少一个固定调度器在九个单元中的六个单元中超过了竞赛平均值,但没有任何策略在各领域中占据主导地位。这些结果揭示了展示的能力与共享预算实现之间的持续差距。
cs.CL / 65 / 2608.16053

DuplexGen: Decoupling Content, Timing, and Acoustics for Synthetic Dialogue Speech

DuplexGen:解耦内容、时机和声学的合成对话语音
Wang, Pengcheng, Li, Sheng, Li, Jiyi, Shinozaki, Takahiro
Abstract
Synthetic conversational speech has become an important resource for developing and evaluating conversational speech systems. However, existing dialogue synthesis pipelines typically generate dialogue content first and then insert interruptions, overlap, and backchannels using handcrafted markers or timing rules, making conversational timing prescribed rather than interaction-driven. We present DuplexGen, a dialogue synthesis framework that explicitly decouples content, timing, and acoustics. An LLM first generates the dialogue script, and then two full-duplex conversational models perform the script while listening to each other in real time. This allows conversational timing to emerge naturally while preserving the scripted content. Finally, a high-fidelity text-to-speech model re-renders the interaction without altering its timing. As a demonstration of the proposed framework, we construct a patient--clinician conversational speech corpus with construction-time annotations, including word timestamps, speaker activity, overlap regions, and interaction events. Experimental results show that the proposed framework produces conversational dynamics closer to real dialogue than conventional stitching-based synthesis.
Chinese Translation
合成对话语音已成为开发和评估对话语音系统的重要资源。然而,现有的对话合成流程通常先生成对话内容,然后使用手工标记或时机规则插入中断、重叠和反馈通道,使得对话时机是预设的而非由互动驱动的。我们提出了DuplexGen,一个明确解耦内容、时机和声学的对话合成框架。首先,使用大型语言模型(LLM)生成对话脚本,然后两个全双工对话模型实时相互倾听并执行脚本。这使得对话时机能够自然地出现,同时保留脚本内容。最后,一个高保真文本到语音模型在不改变时机的情况下重新渲染互动。作为所提框架的演示,我们构建了一个包含构建时间注释的患者-临床医生对话语音语料库,包括单词时间戳、说话者活动、重叠区域和互动事件。实验结果表明,所提框架生成的对话动态比传统的拼接式合成更接近真实对话。
cs.CL / 66 / 2608.16068

CAPO: Constraint-Aware Prompt Optimization for LLM Agents

CAPO:面向约束的提示优化用于大型语言模型代理
Dong, Victor Ye, Pryzant, Reid, Liu, Yi, Jiao, Jian
Abstract
Large language models (LLMs) are increasingly deployed as agents that rely on system prompts to use tools and complete tasks. Such deployments impose distinct operational requirements, including appropriate tool use, concise prompts and solution paths, and compliance with safety and formatting policies. For many practitioners, however, assembling domain-specific supervised data to post-train models to meet these requirements is infeasible. We introduce CAPO (Constraint-Aware Prompt Optimization), a primal-dual method that combines pool-based rewrites with adaptive constraint weighting to optimize system prompts under explicit operational constraints. Across agentic benchmarks, CAPO more reliably reaches empirically feasible operating points while improving task performance. CAPO also generalizes beyond agentic settings, achieving strong results on assistant-style evaluations with output-format and safety/privacy constraints. We further introduce DCAPO (Dynamically Trained CAPO), which trains a feedback- and dual-conditioned rewriter with pool-based GRPO while keeping the task agent frozen. Across task agents of different sizes, DCAPO produces a feasible prompt in every evaluated domain and matches or improves the task accuracy achieved by the evaluated baselines. A surrogate analysis characterizes how finite-pool and discrete-rewrite errors enter the inexact primal-dual procedure.
Chinese Translation
大型语言模型(LLMs)越来越多地作为代理被部署,这些代理依赖系统提示来使用工具和完成任务。这种部署带来了独特的操作要求,包括适当的工具使用、简洁的提示和解决路径,以及遵守安全和格式政策。然而,对于许多从业者来说,组装特定领域的监督数据以进行后训练以满足这些要求是不可行的。我们提出了CAPO(面向约束的提示优化),这是一种原始-对偶方法,结合了基于池的重写与自适应约束加权,以在明确的操作约束下优化系统提示。在代理基准测试中,CAPO更可靠地达到经验上可行的操作点,同时提高任务性能。CAPO还超越了代理设置,在输出格式和安全/隐私约束的助手风格评估中取得了良好的结果。我们进一步引入了DCAPO(动态训练的CAPO),它训练一个反馈和双重条件的重写器,使用基于池的GRPO,同时保持任务代理不变。在不同规模的任务代理中,DCAPO在每个评估领域都生成了可行的提示,并且匹配或提高了评估基线所达到的任务准确性。替代分析描述了有限池和离散重写误差如何进入不精确的原始-对偶过程。
cs.CL / 67 / 2608.16071

Skill2Query: Exploiting Skill Structure to Generate Pseudo-Queries for Agent Skill Retrieval

Skill2Query:利用技能结构生成伪查询以进行代理技能检索
Ding, Lihui, Guo, Zihan, Lu, Bingwei, Zhou, Chenyu, Zhou, Yuanjian, Zhang, Weinan, Lin, Jianghao, Ge, Dongdong
Abstract
Pseudo-query generation can alleviate the supervision bottleneck for agent skill retrieval, but existing document-level approaches typically leave the rich internal relations among capabilities, parameters, and usage examples implicit. As a result, generated queries may be topically relevant to a skill while lacking capability grounding and parameter consistency, raising the question of whether explicitly exploiting a skill document's internal structure can produce more effective retrieval signals. We therefore propose Skill2Query, a framework that first parses a skill document into a Skill Knowledge Graph and then generates pseudo-queries through a three-stage process including style mimicking, query template generation, and parameter filling. The generated queries can be used for offline index augmentation, online query expansion, and retriever training. Four benchmarks (TheoremQA, LogicBench, ToolQA, and CHAMP) are used to evaluate Skill2Query with large-scale skill candidate pools across multiple downstream applications, including skill retrieval, retriever training, and end-to-end agent execution. Using nearly 30K skills across diverse domains, we generate 700K category-diverse pseudo-queries. Skill2Query consistently improves sparse, dense, and skill-routing retrieval, with an average Recall@1 gain of 6.70 percentage points across retrieval settings. Skill2Query-generated training data also achieves the best Recall@1 and nDCG@1 among the evaluated generation baselines. Further evaluations with multiple LLM backends demonstrate that improved skill retrieval translates into higher agent task success rates. Code and resources are available at https://github.com/MatZaharia/Skill2Query.
Chinese Translation
伪查询生成可以缓解代理技能检索中的监督瓶颈,但现有的文档级方法通常将能力、参数和使用示例之间丰富的内部关系隐含化。因此,生成的查询可能在主题上与技能相关,但缺乏能力基础和参数一致性,这引发了一个问题:是否可以通过明确利用技能文档的内部结构来产生更有效的检索信号。因此,我们提出了Skill2Query,一个框架,该框架首先将技能文档解析为技能知识图谱,然后通过包括风格模仿、查询模板生成和参数填充在内的三阶段过程生成伪查询。生成的查询可用于离线索引增强、在线查询扩展和检索器训练。我们使用四个基准(TheoremQA、LogicBench、ToolQA和CHAMP)评估Skill2Query,涵盖多个下游应用中的大规模技能候选池,包括技能检索、检索器训练和端到端代理执行。在近30K个不同领域的技能中,我们生成了70万条类别多样的伪查询。Skill2Query在稀疏、密集和技能路由检索中始终表现出色,在各种检索设置中平均提高了6.70个百分点的Recall@1。Skill2Query生成的训练数据在评估的生成基准中也达到了最佳的Recall@1和nDCG@1。进一步与多个大型语言模型(LLM)后端的评估表明,改进的技能检索转化为更高的代理任务成功率。代码和资源可在 https://github.com/MatZaharia/Skill2Query 获取。
cs.CL / 68 / 2608.16114

HyperSkill: Self-Evolving LLM Agents via Hypergraph-Structured Skill Memory

HyperSkill:通过超图结构技能记忆实现自我进化的LLM代理
Xu, Ruiyao, Yang, Tiankai, Huang, Wei-Chieh
Abstract
As agentic tasks grow in complexity, LLM agents increasingly rely on experiential memory to reuse procedural knowledge across tasks. Effective memory design must jointly address what to store, how memory is structured and retrieved, and how memory evolves. Existing systems tackle each only partially: they store trajectories, insights, or workflows as isolated entries, discarding compositional relationships among subtasks and reusable skills; retrieve by flat embedding similarity that ignores relational signals; and maintain memory without leveraging its relational structure. We propose HyperSkill, a hypergraph-based memory framework that jointly improves all three. HyperSkill represents memory as a hypergraph with two node types, subtask steps and reusable skills, where each hyperedge links the subtasks and skills from a single trajectory. Dual-path retrieval queries both subtask and trajectory levels, ranking skills by co-occurrence across retrieved trajectories. Periodic structure-informed maintenance prunes low-utility nodes and merges redundant skills via quality-weighted propagation. Across xBench, GAIA, and WebWalkerQA with GPT-4o and Qwen3-30B-A3B, HyperSkill outperforms ten memory baselines, yielding gains of up to +11.51 on GAIA and +11.18 on WebWalkerQA.
Chinese Translation
随着代理任务复杂性的增加,LLM代理越来越依赖经验记忆在任务之间重用程序知识。有效的记忆设计必须共同解决存储内容、记忆结构与检索方式以及记忆演变的问题。现有系统仅部分解决这些问题:它们将轨迹、洞察或工作流程存储为孤立的条目,忽略了子任务和可重用技能之间的组合关系;通过扁平嵌入相似性进行检索,忽视了关系信号;并且在维护记忆时未能利用其关系结构。我们提出了HyperSkill,一个基于超图的记忆框架,旨在共同改善这三方面。HyperSkill将记忆表示为一个超图,包含两种节点类型:子任务步骤和可重用技能,其中每条超边连接来自单一轨迹的子任务和技能。双路径检索同时查询子任务和轨迹层面,通过在检索到的轨迹中共同出现的频率对技能进行排名。周期性的结构信息维护修剪低效节点,并通过质量加权传播合并冗余技能。在xBench、GAIA和WebWalkerQA上,使用GPT-4o和Qwen3-30B-A3B,HyperSkill的表现超过了十个记忆基准,在GAIA上获得了高达+11.51的提升,在WebWalkerQA上获得了+11.18的提升。
cs.CL / 69 / 2608.16168

QUMem: Personalized Memory for Query-Conditioned User-State Inference in LLM Agents

QUMem:用于LLM代理中的查询条件用户状态推断的个性化记忆
Wang, Heng, Li, Yifei, Zhang, Lingling, Li, Pengyu, Che, Xinyu, Zhang, Xinyu, Yang, Zesheng
Abstract
Large language model (LLM) agents increasingly use external memory systems to support personalization by drawing on long and evolving interaction histories, in which user preferences may be distributed across time, change with context, and conflict with earlier evidence. However, existing systems face three limitations: fixed-turn, fixed-token, or session-based boundaries can mix unrelated dialogue or split an event from its causes, decisions, and outcomes; storing multiple pieces of user information from the same interaction as a single memory binds together items that serve different functions and should be independently retrievable; and treating the current task as a single top-$k$ retrieval query can return fragments that are individually relevant but fail to jointly capture preference evolution, temporal validity, and contextual applicability. We introduce \textsc{QUMem}, a structured memory framework for query-conditioned user-state inference. \textsc{QUMem} first segments interaction histories into variable-length episodes according to semantic continuity, then decomposes each episode into independently retrievable factual, preference, and transferable insight memories while preserving temporal positions and source evidence. At inference time, three sequential agents identify task-specific information needs, plan multi-query retrieval over the typed memory stores, and jointly infer a temporally and contextually valid user state for downstream response generation. \textsc{QUMem} achieves state-of-the-art performance on both PersonaMem and KnowU-Bench, demonstrating the effectiveness of query-conditioned user-state inference for long-term personalization.
Chinese Translation
大型语言模型(LLM)代理越来越多地使用外部记忆系统,通过利用长期和不断演变的互动历史来支持个性化,其中用户偏好可能随着时间而分布、随着上下文而变化,并与早期证据相冲突。然而,现有系统面临三大限制:固定轮次、固定令牌或基于会话的边界可能混合无关对话或将事件与其原因、决策和结果分开;将同一互动中的多条用户信息存储为单一记忆将不同功能的项目绑定在一起,且应独立检索;将当前任务视为单个 top-$k$ 检索查询可能返回各自相关但未能共同捕捉偏好演变、时间有效性和上下文适用性的片段。我们提出了 extsc{QUMem},一种用于查询条件用户状态推断的结构化记忆框架。 extsc{QUMem} 首先根据语义连续性将互动历史分割为可变长度的情节,然后将每个情节分解为可独立检索的事实、偏好和可转移洞察记忆,同时保留时间位置和来源证据。在推断时,三个顺序代理识别特定任务的信息需求,规划对类型记忆存储的多查询检索,并共同推断出适用于后续响应生成的时间和上下文有效的用户状态。 extsc{QUMem} 在 PersonaMem 和 KnowU-Bench 上实现了最先进的性能,展示了查询条件用户状态推断在长期个性化中的有效性。
cs.CL / 70 / 2608.16185

LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents

LENS:通过动态原始文档的潜在证据探索进行上下文搜索
Wang, Xingjun, Li, Gongsheng, Fan, Qi, Mao, Yunlin, Su, Luyan, Chen, Yingda
Abstract
LLM agents increasingly answer questions over dynamic raw-document collections, where files may change before preprocessing, and relevant evidence (spans, sections, pages, or tables) is query-dependent. Existing retrieval-augmented approaches pre-materialize evidence via fixed chunking, embeddings, or persistent indexes: effective for lookup, yet costly, stale-prone, and committed to a granularity before the query is known. We formulate in-context search as Budgeted Evidence Localization over a latent evidence space induced by dynamic raw documents and propose LENS (Latent Evidence Exploration and Search), an index-free framework. Instead of pre-materializing the evidence space, LENS maintains a query-conditioned belief over candidate units, iteratively selecting candidates via complementary lexical, local, and exploratory proposal policies, updating the belief via an LLM relevance oracle, and narrowing toward high-posterior regions under a controllable budget. Evidence is consolidated into compact, source-grounded regions of interest and compressed into self-organizing knowledge clusters reused across related queries. On a controlled 500-question evaluation with matched corpus snapshots, LENS reaches 62.4% exact match and 84.8% evidence recall vs. 65.2% exact match but 50.4% evidence recall for a ReAct-style baseline. Across scales, LENS gives the strongest supporting-fact localization and answer grounding. On a fixed 150-question fullwiki subset over the raw Wikipedia dump with zero indexing, LENS and ReAct are nearly tied in official answer quality (43.3% vs. 42.7% EM), with LENS grounding more answers in retrieved evidence (84.0% vs. 70.7%). A no-retrieval Closed-Book reference highlights the contribution of model memory. LENS is query-ready after corpus changes, needs no preprocessing or persistent index, and preserves source-grounded evidence localization throughout.
Chinese Translation
大型语言模型(LLM)代理越来越多地在动态原始文档集合上回答问题,这些文件可能在预处理之前发生变化,并且相关证据(跨度、部分、页面或表格)依赖于查询。现有的增强检索方法通过固定分块、嵌入或持久索引预先生成证据:这种方法在查找时有效,但代价高昂、易过时,并且在查询确定之前就承诺了粒度。我们将上下文搜索形式化为在动态原始文档诱导的潜在证据空间上的预算证据定位,并提出了LENS(潜在证据探索与搜索),这是一个无索引框架。LENS不预先生成证据空间,而是保持对候选单元的查询条件信念,通过互补的词汇、局部和探索性提议策略迭代选择候选,利用LLM相关性oracle更新信念,并在可控预算下向高后验区域收敛。证据被整合为紧凑的、基于源的兴趣区域,并压缩为在相关查询中重复使用的自组织知识集群。在一个控制的500个问题评估中,LENS在匹配语料库快照下达到了62.4%的精确匹配率和84.8%的证据召回率,而ReAct风格的基线则为65.2%的精确匹配率和50.4%的证据召回率。在各个规模上,LENS提供了最强的支持事实定位和答案基础。在一个固定的150个问题的全维基百科子集上,使用零索引的原始维基百科转储,LENS和ReAct在官方答案质量上几乎持平(43.3%对42.7%精确匹配),而LENS在检索证据中基础更多答案(84.0%对70.7%)。一个无检索的闭卷参考突显了模型记忆的贡献。LENS在语料库变化后随时可以查询,无需预处理或持久索引,并在整个过程中保持基于源的证据定位。
cs.CL / 71 / 2608.16224

STAIR: Semantic-Temporal Automaton for Interpretable Reasoning in Temporal Question Answering

STAIR:用于可解释推理的语义-时间自动机在时间问答中的应用
Dai, Xinlong, Zhang, Jinchuan, Gao, Lei, Hu, Xinzhe, He, Yuefeng, Gao, Hui
Abstract
By leveraging large-scale pretraining, LLMs can interpret diverse temporal expressions and question formulations without task-specific training. However, existing prompt-based neuro-symbolic systems continue to rely on LLMs for both semantic interpretation and exact temporal inference. Consequently, discrete decisions regarding intervals, time anchors, and ordered states remain vulnerable to probabilistic errors and difficult to verify. We present STAIR, a \textbf{S}emantic-\textbf{T}emporal \textbf{A}utomaton for \textbf{I}nterpretable \textbf{R}easoning. STAIR separates semantic interpretation from precise temporal inference: an answer-free LLM adapter maps complex question formulations to normalized temporal intents, while a deterministic temporal automaton with finite control and guarded transitions executes the corresponding policies over canonicalized evidence. Following a rule-first design, STAIR resolves standard questions without invoking an LLM and applies semantic adaptation only when the rule path fails to produce an executable intent. This approach reduces free-form reasoning, making temporal decisions verifiable and interpretable. Specifically, guarded execution supports precise point-time containment and before/after selection, while semantic adaptation handles non-exact intervals and time-anchored queries. Across the TimeQA-Easy, TimeQA-Hard, TempReason-L2, and TempReason-L3 datasets, STAIR consistently outperforms strong baselines in the TQA task using matched model settings, achieving average F1 improvements of 16.57\% and 3.10\% when utilizing the Qwen2.5-7B and GPT-4o-mini models, respectively. Furthermore, ablations and diagnostic analyses demonstrate that STAIR excels at handling both boundary-sensitive and order-sensitive queries, while its guarded execution and semantic adaptation ensure precise point-time reasoning and inexact intervals, respectively.
Chinese Translation
通过利用大规模预训练,LLMs(大型语言模型)能够在没有特定任务训练的情况下解释多样的时间表达和问题形式。然而,现有的基于提示的神经符号系统仍然依赖LLMs进行语义解释和精确的时间推理。因此,关于区间、时间锚点和有序状态的离散决策仍然容易受到概率错误的影响,并且难以验证。我们提出了STAIR,一个 extbf{S}emantic- extbf{T}emporal extbf{A}utomaton for extbf{I}nterpretable extbf{R}easoning(可解释推理的语义-时间自动机)。STAIR将语义解释与精确的时间推理分开:一个无答案的LLM适配器将复杂的问题形式映射到标准化的时间意图,而一个具有有限控制和受保护转移的确定性时间自动机在规范化证据上执行相应的策略。遵循规则优先的设计,STAIR在不调用LLM的情况下解决标准问题,仅在规则路径未能产生可执行意图时应用语义适配。这种方法减少了自由形式推理,使时间决策可验证且可解释。具体而言,受保护的执行支持精确的时间点包含和前后选择,而语义适配处理非精确区间和时间锚定查询。在TimeQA-Easy、TimeQA-Hard、TempReason-L2和TempReason-L3数据集上,STAIR在TQA任务中始终优于强基线,使用匹配的模型设置时,分别在使用Qwen2.5-7B和GPT-4o-mini模型时实现了平均F1提升16.57 ext{%}和3.10 ext{%}。此外,消融实验和诊断分析表明,STAIR在处理边界敏感和顺序敏感查询方面表现出色,而其受保护的执行和语义适配分别确保了精确的时间点推理和不精确区间的处理。
cs.CL / 72 / 2608.16269

Domain-Agnostic Neural Topic Modeling with Contextual Token-Level Semantic Graph Representation

领域无关的神经主题建模与上下文令牌级语义图表示
Seo, Seung-Won, Cho, Won Ik, Yoo, Yongmin
Abstract
Recent advances in neural topic models with pre-trained language models (PLMs) have achieved strong performance by leveraging general-domain pre-training, yet their topic interpretability often degrades on specialized corpora. This limitation primarily stems from the geometry of the embedding space, where domain-specific terms unseen during pre-training collapse into an indistinguishable region, and neither domain-specific re-training, word-level graph enrichment, nor parameter-efficient fine-tuning can restructure this space without inheriting the capacity ceiling of the underlying encoder. Our key insight is that a learnable graph layer operating on token-level PLM embeddings can acquire corpus-specific semantic structure that the frozen encoder lacks, because token-level graphs preserve document-local context that word-level representations discard and joint optimization with the topic objective reshapes embedding geometry directly from target-domain evidence. We instantiate this insight as DARTopic, a domain-agnostic framework that constructs token-level semantic graphs from frozen PLM embeddings and jointly trains a GNN encoder with topic inference. Across three benchmarks spanning general, biomedical, and legal domains, DARTopic consistently outperforms strong baselines in topic coherence and document clus- tering without any encoder fine-tuning, while demonstrating robustness to PLM choice and favorable runtime efficiency over fine-tuning based alternatives.
Chinese Translation
最近,利用预训练语言模型(PLMs)的神经主题模型取得了显著的性能提升,得益于通用领域的预训练,然而它们在专业语料库上的主题可解释性往往下降。这一限制主要源于嵌入空间的几何结构,其中在预训练期间未见过的领域特定术语会坍缩到一个不可区分的区域,而领域特定的再训练、词级图增强或参数高效的微调都无法在不继承基础编码器容量上限的情况下重构这一空间。我们的关键见解是,操作于令牌级PLM嵌入的可学习图层能够获取语料库特定的语义结构,而冻结的编码器则缺乏这一能力,因为令牌级图保留了文档局部上下文,而词级表示则丢弃了这些信息,并且与主题目标的联合优化可以直接从目标领域证据重塑嵌入几何结构。我们将这一见解具体化为DARTopic,一个领域无关的框架,它从冻结的PLM嵌入构建令牌级语义图,并与主题推断共同训练一个图神经网络(GNN)编码器。在涵盖通用、生物医学和法律领域的三个基准测试中,DARTopic在主题一致性和文档聚类方面始终优于强基线,无需任何编码器微调,同时对PLM选择表现出鲁棒性,并在运行效率上优于基于微调的替代方案。
cs.CL / 73 / 2608.16286

Clause Encounters of the Third Kind: Can LLMs Replace Language Teachers?

第三类条款遭遇:大型语言模型能否取代语言教师?
Šekrst, Kristina, Kovačić, Ana
Abstract
While various organizations now actively encourage LLM use in classrooms, we still lack rigorous, systematic evaluations of how well these models actually perform the fundamental tasks of language pedagogy. This paper examines whether state-of-the-art LLMs can deliver the kind of corrective feedback and methodological explanations that language learners need. The study tests multiple large language models on their ability to identify, correct, and explain common learner mistakes in English, by systematically varying model parameters to investigate how these technical adjustments affect output quality, pedagogical clarity, and consistency, along with using retrieval-augmented generation to query methodological data. The evaluation employs automated metrics (GLEU, BERTScore) but also human expert judgments to capture dimensions that purely computational measures miss: linguistic nuance, cultural sensitivity, and instructional appropriateness. While models demonstrate impressive surface-level correction abilities, their explanations often lack the terminological and domain knowledge that effective language teaching requires, suggesting that current enthusiasm for AI-assisted language learning may be outpacing our understanding of these systems' actual pedagogical competence.
Chinese Translation
尽管各种组织现在积极鼓励在课堂上使用大型语言模型(LLMs),但我们仍然缺乏对这些模型在语言教学基本任务中实际表现的严格系统评估。本文考察了最先进的LLMs是否能够提供语言学习者所需的纠正反馈和方法论解释。研究测试了多种大型语言模型在识别、纠正和解释英语学习者常见错误的能力,通过系统地变化模型参数,探讨这些技术调整如何影响输出质量、教学清晰度和一致性,同时使用增强检索生成(retrieval-augmented generation)查询方法论数据。评估采用了自动化指标(GLEU, BERTScore),同时也结合了人类专家的判断,以捕捉纯计算度量所忽视的维度:语言细微差别、文化敏感性和教学适宜性。尽管模型在表面纠正能力上表现出色,但它们的解释往往缺乏有效语言教学所需的术语和领域知识,这表明当前对AI辅助语言学习的热情可能超出了我们对这些系统实际教学能力的理解。
cs.CL / 74 / 2608.16295

Executable Code Knowledge: Code as a Native, Validation-Carrying Knowledge Representation for AI Coding Agents

可执行代码知识:代码作为AI编码代理的原生验证承载知识表示
Gao, Xueping
Abstract
AI coding agents need more than relevant snippets: they need business semantics, validation evidence, relations, and assurance that their context is current. Existing systems usually infer or externalize this knowledge through retrieval, summaries, graphs, rules, or reverse specifications. We investigate a complementary representation in which selected code units directly carry agent-usable knowledge. We introduce Executable Code Knowledge (ECK) and define an Executable Code Knowledge Unit (ECKU) as a source-bound object combining stable identity, semantics, executable behavior, contracts, evidence, relations, provenance, validation state, and a query interface. Our Python prototype supports code-local authoring, manifest export, evidence execution, exact changed-line impact, freshness checking, and agent-facing projections. Across three real Python repositories and 26 controlled patch tasks, direct ECK provides executable test coverage for 11/11 evidence-bearing tasks and exact selectors for 9/11; hiding declared evidence reduces exact recovery to 1/11 (paired exact McNemar p=0.0078). ECK-derived rules recover 11/11 exact selectors, showing that rules are effective delivery artifacts while ECK supplies source binding, validation state, impact, and freshness. Exact changed-line impact matches independently authored labels on all 26 patches (12 unit links; precision, recall, and F1 all 1.000). AST-bounded fingerprints classify 50 positive changes and 17 unrelated same-file controls correctly, whereas static rules snapshots detect none of the 50 stale cases. Model-backed patch-review and cross-layer studies measure projection fidelity rather than independent impact discovery. These results support a hybrid architecture: retrieval for coverage, ECK for source and evidence governance, and projections for delivery.
Chinese Translation
AI编码代理不仅需要相关代码片段,还需要业务语义、验证证据、关系以及确保其上下文是最新的。现有系统通常通过检索、摘要、图形、规则或反向规范来推断或外化这些知识。我们研究了一种互补的表示方式,其中选定的代码单元直接承载可供代理使用的知识。我们引入了可执行代码知识(Executable Code Knowledge, ECK),并将可执行代码知识单元(Executable Code Knowledge Unit, ECKU)定义为一个源绑定对象,结合了稳定的身份、语义、可执行行为、合同、证据、关系、来源、验证状态和查询接口。我们的Python原型支持代码本地创作、清单导出、证据执行、精确变更行影响、新鲜度检查和面向代理的投影。在三个真实的Python代码库和26个受控补丁任务中,直接的ECK为11/11个承载证据的任务提供了可执行的测试覆盖,并为9/11个任务提供了精确选择器;隐藏声明的证据将精确恢复率降低到1/11(配对精确McNemar p=0.0078)。基于ECK的规则恢复了11/11个精确选择器,表明规则是有效的交付工件,而ECK则提供了源绑定、验证状态、影响和新鲜度。精确的变更行影响与所有26个补丁上的独立创作标签相匹配(12个单元链接;精确度、召回率和F1均为1.000)。AST绑定指纹正确分类了50个正向变化和17个无关的同文件对照,而静态规则快照未能检测到50个过时案例。模型支持的补丁审查和跨层研究测量投影保真度,而非独立影响发现。这些结果支持一种混合架构:检索用于覆盖,ECK用于源和证据治理,投影用于交付。
cs.CL / 75 / 2608.16303

FTA-Mem: Fact-Time-Affect Anchored Memory for Low-Density Long-Term Dialogue

FTA-Mem:基于事实-时间-情感锚定的低密度长期对话记忆
Liu, Chang, Zhang, Shuyi, Ma, Changsheng, Tao, Yongfeng, Yang, Minqiang, Hu, Bin
Abstract
Long-term emotional-support agents require memory mechanisms for personalized understanding across sessions. However, emotional-support dialogue is often low-density: turns are incomplete, evidence is scattered, and user states evolve over time. Existing memory methods usually rely on fixed units, such as turn-level notes or session summaries, which may lose details or introduce redundant noise. We propose FTA-Mem, a structured memory framework for low-density long-term dialogue. FTA-Mem uses Boundary-preserving Window Segmentation (BWS) to form coherent situation fragments, and constructs Fact-Time-Affect Memory Units (FTA Units) that jointly encode factual content, temporal grounding, and affective context. Retrieved units are then synthesized into structured context for answer generation. Experiments on ES-MemEval and LoCoMo show that FTA-Mem improves overall long-term memory question answering across benchmarks with different information-density characteristics. On ES-MemEval, FTA-Mem achieves 0.3871 F1 and 0.6668 BERTScore. Further analysis shows that situation-level FTA construction better balances evidence preservation and construction cost than coarse session-level or overly fine-grained turn-pair construction, providing an effective granularity trade-off for long-term dialogue memory.
Chinese Translation
长期情感支持代理需要记忆机制,以便在多个会话中实现个性化理解。然而,情感支持对话通常是低密度的:对话轮次不完整,证据分散,用户状态随时间演变。现有的记忆方法通常依赖于固定单元,例如轮次级笔记或会话摘要,这可能会丢失细节或引入冗余噪声。我们提出了FTA-Mem,一种针对低密度长期对话的结构化记忆框架。FTA-Mem使用边界保持窗口分割(Boundary-preserving Window Segmentation, BWS)来形成连贯的情境片段,并构建事实-时间-情感记忆单元(Fact-Time-Affect Memory Units, FTA Units),共同编码事实内容、时间基础和情感背景。检索到的单元随后被合成到结构化上下文中以生成答案。在ES-MemEval和LoCoMo的实验表明,FTA-Mem在不同信息密度特征的基准测试中改善了整体长期记忆问答。在ES-MemEval上,FTA-Mem达到了0.3871的F1值和0.6668的BERTScore。进一步分析表明,情境级FTA构建在证据保留和构建成本之间提供了更好的平衡,相较于粗糙的会话级或过于细粒度的轮次对构建,提供了长期对话记忆的有效粒度权衡。
cs.CL / 76 / 2608.16333

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning

步骤级在线蒸馏:在在线蒸馏与监督微调之间的插值
Sun, Changhui, Liu, Lanbo, Lei, Hang, Ling, Tong, Xie, Jiahang, Zheng, Zhiyong, Wang, Yujia, Liu, Hao, Xiao, Feng, Liu, Lu, Du, Yanlong, Cheng, Zifeng, Jiang, Ziwei, Gu, Qing
Abstract
On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conventional off-policy distillation with substantially less data. However, standard token-level OPD can provide only fragmented corrections along an erroneous student trajectory and cannot unfold a complete and correct repair path. Motivated by this limitation, we propose \emph{Step-Level On-Policy Distillation} (SOPD), which combines the long-horizon correction of supervised fine-tuning (SFT) with the on-policy advantage of OPD to provide step-level supervision over complete student-generated trajectories. We show that, at different limits of step length, SOPD reduces to SFT or approximates OPD. Compared with SFT, the teacher responses in SOPD are conditioned on student trajectories and therefore align more closely with student-visited states; compared with OPD, SOPD provides longer-horizon corrections rather than fragmented token-level guidance. Across both reasoning and agent tasks, SOPD substantially outperforms conventional SFT and OPD. For example, on ALFWorld, SOPD improves the average success rate by 13.4 points over Vanilla OPD. We hope this work offers a new perspective for future research on distillation methods.
Chinese Translation
在线蒸馏(On-policy distillation, OPD)通过对学生生成的轨迹与教师的对数分布进行对齐,从而使学生模型与教师模型一致。这种方法在实证上取得了显著的效果,通常可以在使用显著更少的数据的情况下超越传统的离线蒸馏(off-policy distillation)。然而,标准的标记级OPD只能在错误的学生轨迹上提供零散的修正,无法展开完整且正确的修复路径。基于这一局限性,我们提出了 extit{步骤级在线蒸馏}(Step-Level On-Policy Distillation, SOPD),该方法结合了监督微调(Supervised Fine-Tuning, SFT)的长远修正与OPD的在线优势,为完整的学生生成轨迹提供步骤级监督。我们展示了在不同的步骤长度极限下,SOPD可简化为SFT或近似OPD。与SFT相比,SOPD中的教师响应是基于学生轨迹的,因此与学生访问的状态更为一致;与OPD相比,SOPD提供的是更长远的修正,而非零散的标记级指导。在推理任务和代理任务中,SOPD的表现显著优于传统的SFT和OPD。例如,在ALFWorld上,SOPD相比于Vanilla OPD提高了平均成功率13.4个百分点。我们希望这项工作为未来的蒸馏方法研究提供新的视角。
cs.CL / 77 / 2608.16344

IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

IndicQE-APE:印度语言质量估计与自动后编辑的基准
Kanojia, Diptesh, Sindhujan, Archchana, Deoghare, Sourabh, Sokova, Daria, Qian, Shenbin, Koushik, Girish, Ranasinghe, Tharindu, Orăsan, Constantin, Zerva, Chrysoula, Rei, Ricardo, Blain, Frédéric, Martins, André F. T., Turchi, Marco, Negri, Matteo, Chatterjee, Rajen, Kunchukuttan, Anoop, Khapra, Mitesh M., Bhattacharyya, Pushpak
Abstract
Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and language pairs on one footing. We consolidate the WMT 2020--2024 shared-task lineage with an extended English--Malayalam resource into \indicqe: $126{,}754$ instances over nine directional pairs, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level OK/BAD tags and an error explanation, and a test set stratified over four difficulty axes. On it, we benchmark six prompted LLMs and three COMET metrics on segment-level QE, and three systems on APE. Two of the axes are defined partly on the direct assessment and select a compressed slice of it, so each axis is compared against a control drawn from the same language pair with the same score distribution. Only one survives that control: segments whose holistic and token-level quality signals conflict are ranked worse than equally-scored segments of the same language, for all nine systems and all seven pairs that carry the axis. Annotator disagreement, which looks second-hardest without the control, has no effect with it. Few-shot prompting costs every model $\leq$ $3.4$B both correlation and output-format compliance. Within-language accuracy does not make scores comparable across pairs: of the three trained metrics, the one with the best within-language correlation loses most when the pairs are pooled. The benchmark and code will be released.
Chinese Translation
印度质量估计(QE)和自动后编辑(APE)数据分散在不同的发布版本中,因此没有单一资源可以支持跨任务和语言对的训练与评估。我们将 WMT 2020-2024 共享任务系列与扩展的英语-马拉雅拉姆语资源整合为 extit{indicqe}:包含 $126{,}754$ 个实例,涵盖九个方向对,最多有四种标签类型对齐在同一段落上,包括直接评估、人类后编辑、词级 OK/BAD 标签和错误解释,并且测试集在四个难度轴上进行了分层。在此基础上,我们对六个提示的 LLM 进行段级 QE 基准测试,并对三个系统进行 APE 测试。两个轴部分基于直接评估定义,并选择其压缩切片,因此每个轴与来自同一语言对且具有相同得分分布的控制组进行比较。只有一个轴在控制下存活:整体和词级质量信号冲突的段落排名低于同一语言中得分相同的段落,适用于所有九个系统和所有七个携带该轴的对。注释者之间的分歧在没有控制的情况下看起来是第二难的,但在有控制的情况下没有影响。少量提示使每个模型的相关性和输出格式合规性均不超过 $3.4$B。语言内的准确性并不能使得跨对的得分可比:在三个训练的指标中,具有最佳语言内相关性的指标在对池化时损失最大。基准和代码将会发布。
cs.CL / 78 / 2608.16347

Architecture-Dependent Causal Transfer of Activation States Across Large Language Models

架构依赖的激活状态因果转移在大型语言模型之间的研究
Piepereit, Fernando Cardenas
Abstract
Direct communication between AI systems relies on natural language as an intermediate layer, incurring encoding/decoding overhead, token cost, and latency. We ask whether internal activation states can instead be transferred causally between different large language model (LLM) architectures via a learned projection, evaluated at three levels: representational similarity, cross-model retrieval from projected states, and end-to-end causal transfer via activation injection during generation. Using four architecturally diverse open-weight models (Qwen2-0.5B, Phi-3-mini, Mistral-7B, FLAN-T5-base), we find that representational alignment in trained models exceeds a random-initialization null baseline and is best captured by a rank-based metric (mutual k-nearest-neighbour alignment), more robust to activation-magnitude outliers than centered kernel alignment (CKA) or Procrustes analysis. A learned projection network retrieves the correct target-model representation from a held-out set well above chance for the three causal decoder-only model pairs (45-50% top-1 accuracy vs. 5% chance) but at chance level for the encoder-based FLAN-T5. Injecting projected activations into a target model during generation produces a statistically significant, pre-registered causal effect on retrieval-based output similarity for only one of the three decoder-only pairs (Qwen2-0.5B to Phi-3-mini: 23.3% vs. 0.0% under negative control, p=0.047, FDR-corrected); the two pairs targeting Mistral-7B show no such effect despite comparable representational alignment at the hidden-state level. We interpret these results as evidence for causal transfer of the representational vehicle, not of meaning, and conclude that end-to-end activation-state transfer between LLMs, as currently implemented, is architecture-dependent rather than universal.
Chinese Translation
人工智能系统之间的直接通信依赖于自然语言作为中介层,这带来了编码/解码的开销、标记成本和延迟。我们探讨是否可以通过学习的投影在不同的大型语言模型(LLM)架构之间因果地转移内部激活状态,并在三个层面进行评估:表征相似性、从投影状态进行跨模型检索,以及在生成过程中通过激活注入实现的端到端因果转移。使用四种架构多样的开放权重模型(Qwen2-0.5B、Phi-3-mini、Mistral-7B、FLAN-T5-base),我们发现训练模型中的表征对齐超出了随机初始化的无效基线,并且通过基于排名的度量(互相k最近邻对齐)捕捉得最好,这比中心核对齐(CKA)或Procrustes分析对激活幅度异常值更具鲁棒性。学习的投影网络从一个保留集中检索正确的目标模型表征,对于三个因果解码器模型对(45-50%的顶级准确率与5%的随机机会)远高于随机水平,但对于基于编码器的FLAN-T5则处于随机水平。在生成过程中将投影激活注入目标模型,仅对三个解码器模型对中的一个(Qwen2-0.5B到Phi-3-mini:23.3%对0.0%在负控制下,p=0.047,FDR校正)产生了统计显著的预注册因果效应;尽管在隐藏状态级别的表征对齐相当,针对Mistral-7B的两个模型对没有显示出这样的效果。我们将这些结果解释为表征载体的因果转移的证据,而不是意义的转移,并得出结论:目前实现的LLM之间的端到端激活状态转移是架构依赖的,而非普遍适用的。
cs.CL / 79 / 2608.16353

HalluTracer: Hallucination Detection via Depth-Averaging Truth Signals

HalluTracer:通过深度平均真实信号进行幻觉检测
Guo, Zhihao, Wu, Zonghan, Huo, Huan, Ye, DaYong, Zhang, Junwei, Yao, Weiran, Liu, Zhiwei, Wen, Qingsong, Shao, Yilei
Abstract
Even well-aligned large language models confidently generate factually incorrect text, making hallucination a persistent reliability risk in high-stakes deployments. These models nonetheless carry linearly separable truthfulness signals in their internal representations. Existing white-box detectors, however, collapse this evidence to isolated components or a single depth, discarding discriminative information distributed across the full forward pass. We introduce HalluTracer, a detection framework that reads and aggregates truthfulness evidence across every layer of the forward pass before the model emits any answer token. A geometric analysis reveals that the per-layer signals are weakly correlated, so that simple depth averaging suppresses layer-specific noise and captures nearly all linearly accessible information. Across six open-source language models and five hallucination benchmarks, HalluTracer consistently outperforms matched white-box baselines, with gains ranging from one to fourteen points. Collectively, our work recasts hallucination detection from a layer-selection problem into a depth-aggregation problem governed by the geometric sparsity of the truthfulness signal.
Chinese Translation
即使是经过良好对齐的大型语言模型也会自信地生成事实不正确的文本,使得幻觉在高风险部署中成为一个持续的可靠性风险。然而,这些模型在其内部表示中仍然携带线性可分的真实性信号。然而,现有的白盒检测器将这些证据压缩为孤立的组件或单一深度,丢弃了分布在整个前向传播过程中的判别信息。我们提出了HalluTracer,一个检测框架,在模型发出任何答案标记之前,读取并聚合每一层前向传播中的真实性证据。几何分析表明,逐层信号之间的相关性较弱,因此简单的深度平均可以抑制特定层的噪声,并捕获几乎所有线性可访问的信息。在六个开源语言模型和五个幻觉基准测试中,HalluTracer始终优于匹配的白盒基线,增益范围从一到十四点。总体而言,我们的工作将幻觉检测从一个层选择问题重新定义为一个由真实性信号的几何稀疏性主导的深度聚合问题。
cs.CL / 80 / 2608.16379

Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization Analysis

未适应的多语言自动语音识别在Garrusi库尔德语评估集上的表现:一种共同参考的分阶段归一化分析
Asadpour, Hiwa
Abstract
Evaluating speech recognition for a Kurdish variety written in a Latin field orthography, using a model that outputs Arabic script, creates a measurement problem before a modelling one: direct scoring treats writing-system differences as recognition errors. Jointly normalizing reference and hypothesis avoids this, but also changes reference tokenization, mixing agreement gains with a change in the scoring denominator. I evaluate MMS-1B-all with the Central Kurdish (ckb) adapter, used as released without adaptation, on 1,722 Garrusi questionnaire segments from five speakers (9,763 reference word tokens; 117.9 minutes). I use a common-reference design: the reference is folded once and fixed at 9,763 tokens, while only the hypothesis representation varies. The raw Arabic-script hypothesis scores 111.70% WER and 100.92% CER, with zero exact word matches. Latin transliteration gives 102.36% WER and 57.89% CER; folding it into the reference's reduced orthography gives 97.85% and 51.20%. Thus RAW-to-FOLDED reduces measured WER by 13.85 points and CER by 49.72 points; folding alone accounts for 4.51 and 6.69 points. Substantial error remains: 14.53% of reference tokens are exact matches, edits are substitution-dominated, and per-segment WER is higher for shorter segments. A Southern Kurdish fine-tuned system (aranemini/southern-kurdish-asr), scored under the same design, performs worse on every speaker (1,703 segments), with 109.56% WER and 55.85% CER. However, 12,330 output characters fall outside the folding table, so these rates must be recomputed against the corrected fixed reference. The MMS output also contains 613 unconverted or unmapped characters, showing that part of the residual error reflects scoring-pipeline limits rather than recognition alone. I will release the fixed reference and segment-level results, subject to source-corpus sharing terms, to support independent checking.
Chinese Translation
评估以拉丁字母书写的库尔德语变体的语音识别,使用输出阿拉伯文字的模型,在建模之前产生了测量问题:直接评分将书写系统的差异视为识别错误。联合归一化参考和假设可以避免这一问题,但也改变了参考的分词,将一致性增益与评分分母的变化混合在一起。我在来自五位说话者的1,722个Garrusi问卷段落(9,763个参考词标记;117.9分钟)上评估了未经过适应的中央库尔德语(ckb)适配器的MMS-1B-all。我的设计采用共同参考:参考被折叠一次并固定在9,763个标记上,而只有假设表示发生变化。原始阿拉伯文字假设的字错误率(WER)为111.70%,字符错误率(CER)为100.92%,完全没有准确的单词匹配。拉丁音译的WER为102.36%,CER为57.89%;将其折叠到参考的简化书写中得到97.85%和51.20%。因此,RAW到FOLDED的测量WER减少了13.85个百分点,CER减少了49.72个百分点;仅折叠就占了4.51和6.69个百分点。仍然存在大量错误:14.53%的参考标记是准确匹配,编辑主要以替换为主,且每段的WER在较短段落中更高。经过微调的南部库尔德语系统(aranemini/southern-kurdish-asr)在同一设计下的表现更差(1,703个段落),WER为109.56%,CER为55.85%。然而,12,330个输出字符超出了折叠表,因此这些比率必须针对修正后的固定参考重新计算。MMS输出还包含613个未转换或未映射的字符,显示部分残余错误反映的是评分管道的限制,而不仅仅是识别错误。我将根据源语料共享条款发布修正后的参考和段落级结果,以支持独立检查。
cs.CL / 81 / 2608.16386

Mint-Agent: Introducing Finance-Native Agentic Foundation Models

Mint-Agent:引入金融原生智能基础模型
Agent Team, Zhang, B., Geng, Yaze, Tang, Lei, Yi, Yaoyang, Wu, Zonghan, Hu, Yifan, Wang, Kun, Wen, Qingsong, Shao, Yilei
Abstract
Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon research whose conclusions remain auditable. We present Mint-Agent, a family of finance-native agentic models designed around these two scales of financial intelligence. Mint-Agent is built upon three pillars: data, harness, and algorithm. Our data engine constructs clean, specialized tasks for atomic financial capabilities and long-horizon agentic execution from real-world financial sources. MintHarness enables stable interaction with open-ended environments and maintains auditable evidence trails across extended research trajectories. Our training recipe combines SFT, critical-step OPD, and RLVR to develop separate financial reasoning and agentic execution experts, which are then unified through model merging and multi-teacher on-policy distillation into compact, general-purpose financial agents. This pipeline yields two flagship models, Mint-Cu (9B) and Mint-Ag (27B). Across professional financial benchmarks, our models demonstrate two defining strengths: (1) Reliability: Mint-Ag achieves 98.33% on RFC-Bench, surpassing GPT-5.6-Sol and Claude-Opus-4.8 by 3.66 and 3.00 points; and (2) Executability: Mint-Cu reaches 69.86% on FinSearchComp T2, outperforming Agents-A1-35B and Nex-N2-mini by 22.83 and 12.78 points, while Mint-Ag achieves 76.00% and 60.49% on FinanceAgentBench v1.1 and v2, respectively. These results establish a path toward trustworthy financial intelligence in which domain expertise, long-horizon execution, and auditable evidence are jointly engineered as a unified foundation for frontier agentic models.
Chinese Translation
金融代理不仅需要回忆领域知识:它们还必须可靠,能够在有据可依的证据上执行精确操作,并且具备执行能力,能够维持长期研究,其结论保持可审计性。我们提出了Mint-Agent,一个围绕这两种金融智能规模设计的金融原生智能模型家族。Mint-Agent建立在三个支柱之上:数据、工具和算法。我们的数据引擎从真实世界的金融来源构建干净、专业的任务,以实现原子金融能力和长期智能执行。MintHarness使得与开放环境的稳定交互成为可能,并在扩展的研究轨迹中维护可审计的证据链。我们的训练方案结合了SFT、关键步骤OPD和RLVR,以培养独立的金融推理和智能执行专家,然后通过模型合并和多教师在线蒸馏将其统一为紧凑的通用金融代理。该流程产生了两个旗舰模型,Mint-Cu(9B)和Mint-Ag(27B)。在专业金融基准测试中,我们的模型展示了两个显著优势:(1) 可靠性:Mint-Ag在RFC-Bench上达到了98.33%,超越了GPT-5.6-Sol和Claude-Opus-4.8,分别高出3.66和3.00分;(2) 可执行性:Mint-Cu在FinSearchComp T2上达到了69.86%,超越了Agents-A1-35B和Nex-N2-mini,分别高出22.83和12.78分,而Mint-Ag在FinanceAgentBench v1.1和v2上分别达到了76.00%和60.49%。这些结果为值得信赖的金融智能铺平了道路,其中领域专业知识、长期执行和可审计证据共同构建为前沿智能模型的统一基础。
cs.CL / 82 / 2608.16390

Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics

计算文档并不等于计算文本:网络PDF语料库统计中的单位偏差
Foppiano, Luca
Abstract
PDF corpora advertise their size in tokens but compute every rate they publish (coverage, OCR routing, re-fetch recovery, language mix) per document, and none decomposes its token total. The two units diverge sharply. On CC-MAIN-2021-31-PDF-UNTRUNCATED (7.9M web PDFs, 32.6B tokens), 3.02% of text-bearing documents hold half the tokens (Gini 0.807); documents over 50 pages are 5.00% of the corpus but 53.53% of its text. The PDFs produced by a TeX{} toolchain are 1.66% of documents and 4.05% of the text. The clearest casualty is Common Crawl's truncation cap: it affected 23.06% of documents and 63.08% of the text. Reconstructing the truncated files and extracting both versions, two widely used libraries recover 11.4% and 1.4% of that text; between 72% and 97% of affected documents yield nothing; roughly 55--62% of the corpus's text is lost. Under the 5 MiB cap adopted in March 2025, 30.19% of tokens would still be truncated, and recovery on those documents rises only from 3.3% to 13.2%. We recommend that corpus statistics be reported in both units: documents and tokens.
Chinese Translation
PDF语料库以标记数宣传其规模,但计算其发布的每个比率(覆盖率、OCR路由、重新获取恢复、语言混合)时均按文档计算,而没有分解其标记总数。这两种单位之间存在显著差异。在CC-MAIN-2021-31-PDF-UNTRUNCATED(7.9百万个网络PDF,32.6十亿个标记)中,3.02%的含文本文档占据了一半的标记(基尼系数0.807);超过50页的文档占语料库的5.00%,但占其文本的53.53%。由TeX{}工具链生成的PDF文档占文档的1.66%和文本的4.05%。最明显的受害者是Common Crawl的截断上限:它影响了23.06%的文档和63.08%的文本。重建被截断的文件并提取两个版本后,两个广泛使用的库分别恢复了11.4%和1.4%的文本;受影响文档中有72%至97%未能恢复任何内容;大约55%至62%的语料库文本丢失。在2025年3月采用的5 MiB上限下,30.19%的标记仍将被截断,而这些文档的恢复率仅从3.3%提高到13.2%。我们建议语料库统计应同时以文档和标记两种单位报告。
cs.CL / 83 / 2608.16417

D2-ScaleAgent: Dual-Dimensional Scaling for Long Document Understanding

D2-ScaleAgent:用于长文档理解的双维度缩放
Zhang, Hao, Yang, Longrong, Duan, Lunhao, Wang, Ziyang, Chen, Qing-Guo, Zhao, Shanshan
Abstract
Multi-modal retrieval-augmented generation (RAG) is a key technique for visually rich long document understanding. Existing multi-modal RAG methods are progressively advancing toward multi-agent systems: they first retrieve relevant pages based on a query, and then iteratively understand information within those pages. However, these methods typically rely on fixed workflows and lack the ability to dynamically scale computation at test time, often leading to insufficient evidence. To address this, we propose D2-ScaleAgent, an agentic framework that introduces a dual-dimensional scaling paradigm for retrieval and reasoning. The core of D2-ScaleAgent is a Verifier agent-driven dynamic routing loop based on the intrinsic difficulty of the query, centered around a continuously updated evidence bank that serves as the agent's dynamic working memory: when retrieval needs to be expanded, the agent routes outward (retrieval scaling), decomposing the query into attributes and performing parallel page retrieval, followed by adaptive pruning to ensure comprehensive evidence coverage. When fine-grained reasoning is required, the agent routes inward (reasoning scaling), dynamically selecting sub-agents with varying granularity and count to extract evidence from pages. Finally, D2-ScaleAgent achieves logical closure over the evidence chain. Extensive experiments demonstrate that D2-ScaleAgent is effective on long and visually rich document benchmarks like MMLongBench-Doc, LongDocURL, etc.
Chinese Translation
多模态检索增强生成(RAG)是理解视觉丰富的长文档的关键技术。现有的多模态RAG方法正逐步朝向多智能体系统发展:它们首先根据查询检索相关页面,然后迭代理解这些页面中的信息。然而,这些方法通常依赖于固定的工作流程,缺乏在测试时动态扩展计算的能力,常常导致证据不足。为了解决这个问题,我们提出了D2-ScaleAgent,一个引入双维度缩放范式的智能体框架,用于检索和推理。D2-ScaleAgent的核心是一个基于查询内在难度的验证者驱动的动态路由循环,围绕一个持续更新的证据库,这个证据库作为智能体的动态工作记忆:当需要扩展检索时,智能体向外路由(检索缩放),将查询分解为属性并进行并行页面检索,随后进行自适应剪枝以确保全面的证据覆盖。当需要细粒度推理时,智能体向内路由(推理缩放),动态选择具有不同粒度和数量的子智能体从页面中提取证据。最后,D2-ScaleAgent实现了证据链的逻辑闭合。大量实验表明,D2-ScaleAgent在MMLongBench-Doc、LongDocURL等长且视觉丰富的文档基准上表现出色。
cs.CL / 84 / 2608.16515

When Context Misleads: Intent-Guided Decoding for Robust Retrieval-Augmented Generation

当上下文误导时:基于意图的解码用于增强鲁棒性的检索增强生成
Jin, Haolin, Yang, Pengyue, Chen, Huaming
Abstract
Retrieval-augmented generation (RAG) improves large language models by grounding generation in external evidence, but it also introduces a source trust problem: retrieved context may be useful, irrelevant, or even misleading. Existing RAG systems often apply a fixed trust policy toward retrieved evidence, which can either over-trust incorrect context or underuse context when the user explicitly asks for context-following behavior. Therefore, we propose Intent-Guided Decoding (IGD), a framework that arbitrates between retrieved context and parametric memory according to user intent. IGD uses answer-level filtering and token-level correction to steer the final decoding trajectory between retrieved context and parametric memory. We evaluate IGD on three faithful QA benchmarks and three factual-conflict benchmarks across five LLMs, IGD substantially improves factual recovery, achieving gains of up to 65.4 percentage points on factual-conflict benchmarks over Direct RAG, while preserving or improving strict context-following behavior, this findings highlight the importance of balancing factuality and faithfulness in RAG.
Chinese Translation
检索增强生成(RAG)通过将生成过程与外部证据相结合来提升大型语言模型的性能,但它也引入了一个源信任问题:检索到的上下文可能是有用的、无关的,甚至是误导性的。现有的RAG系统通常对检索到的证据应用固定的信任策略,这可能导致对错误上下文的过度信任,或在用户明确要求上下文跟随行为时对上下文的不足利用。因此,我们提出了基于意图的解码(IGD),这是一个根据用户意图在检索到的上下文和参数记忆之间进行仲裁的框架。IGD使用答案级过滤和标记级修正来引导最终的解码轨迹在检索到的上下文和参数记忆之间。我们在三个忠实的问答基准和三个事实冲突基准上评估了IGD,结果显示IGD显著改善了事实恢复,在事实冲突基准上相较于直接RAG实现了高达65.4个百分点的提升,同时保持或改善了严格的上下文跟随行为,这一发现强调了在RAG中平衡事实性和忠实性的重要性。
cs.CL / 85 / 2608.16553

STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment

STAGE:多偏好大语言模型对齐的受控目标接纳
Tong, Yongqi, Zhang, Zhenyu, Wang, Ruirui, Fu, Kewei, Lin, Shaoqing, Dong, Sijie, Yang, Jiang-Ming, Zhang, Xin, Li, Jianshe
Abstract
Multi-preference alignment is often framed as scalarization: combine reward dimensions, then optimize. This leaves a temporal decision underspecified: when should each preference dimension enter policy optimization? We propose \methodname, a stability-guided active-set controller for controlled objective admission. \methodname starts from a small active set, retains admitted objectives, and expands when reward-deviation gates indicate low recent deviation or a patience budget is exhausted. A probing phase estimates a hard-to-easy order, and adaptive weighting emphasizes underperforming active dimensions. Automatic evaluations with 15 training preferences and 16 held-out benchmark columns show that \methodname obtains higher averages than simultaneous scalarization and shared-budget adapted baselines. Component ablations and expansion dynamics further support cumulative retention, gated admission, and probing-derived ordering as useful design choices in this setting. These results position objective-entry timing as a concrete control variable in reward-vector RLHF.
Chinese Translation
多偏好对齐通常被框架化为标量化:结合奖励维度,然后进行优化。这留下了一个时间决策未被明确:每个偏好维度何时应进入政策优化?我们提出了 extit{STAGE},一种基于稳定性的主动集控制器,用于受控目标接纳。 extit{STAGE}从一个小的主动集开始,保留已接纳的目标,并在奖励偏差门指示最近偏差较低或耐心预算耗尽时进行扩展。一个探测阶段估计了从难到易的顺序,适应性加权强调表现不佳的主动维度。使用15个训练偏好和16个保留基准列的自动评估显示, extit{STAGE}获得的平均值高于同时标量化和共享预算适应基线。组件消融和扩展动态进一步支持累积保留、门控接纳和探测导出的排序作为该设置中有用的设计选择。这些结果将目标进入时机定位为奖励向量强化学习人类反馈(RLHF)中的一个具体控制变量。
cs.CL / 86 / 2608.16554

Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning

询问、条件或放弃:缺失前提推理的强化学习
Tong, Yongqi, Zhang, Zhenyu, Liu, Zimi, Fu, Kewei, Song, Mingli, Zhang, Haofei, Zhang, Junshao, Zhu, Hong, Yang, Jiang-Ming, Zhang, Xin, Li, Jianshe
Abstract
Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for a unique answer. In this setting, the useful response is not always refusal: the model should ask for the missing premise, condition its answer on the unknown quantity, or abstain when no informative conditional response is available. We present \emph{Ask-Condition-Abstain Reinforcement Learning} (ACA-RL), a data-augmented RL framework for this setting. Its reasoning-graph-guided pipeline converts well-posed problems into missing-premise training instances with localized gap annotations; ACA-RL then trains on these instances with a structured reward over five observable response behaviors. We also introduce the \emph{Missing-Premise Benchmark} (MPB), a 274-instance human-verified benchmark spanning mathematical, logical, and real-world word problems. Across Qwen3 and Llama models, ACA-RL consistently improves on MPB while preserving competitive performance on well-posed reasoning tasks. Together with the released code, MPB, and training data, this work supports a new mission for NLP evaluation: measuring whether models can recognize when a task is underdetermined and handle uncertainty, not only whether they can answer fully specified questions.
Chinese Translation
仅回答的强化学习(RL)训练推理模型以解决完全指定的问题,但许多现实查询省略了唯一答案所需的前提。在这种情况下,有用的回应并不总是拒绝:模型应该询问缺失的前提,基于未知量对其答案进行条件化,或者在没有可用的信息性条件响应时选择放弃。我们提出了 extit{询问-条件-放弃强化学习}(ACA-RL),这是一个针对这种情况的数据增强强化学习框架。其推理图引导的流程将良好构造的问题转换为带有局部缺口注释的缺失前提训练实例;然后,ACA-RL在这些实例上进行训练,并对五种可观察的响应行为施加结构化奖励。我们还引入了 extit{缺失前提基准}(MPB),这是一个包含274个实例的人类验证基准,涵盖数学、逻辑和现实世界的文字问题。在Qwen3和Llama模型上,ACA-RL在MPB上始终表现出改进,同时在良好构造的推理任务上保持竞争性能。结合发布的代码、MPB和训练数据,这项工作支持自然语言处理评估的新使命:衡量模型是否能够识别任务是否不确定并处理不确定性,而不仅仅是它们是否能够回答完全指定的问题。
cs.CL / 87 / 2608.16577

BabelSteering: Multilingual Safety Alignment via English Steering Vectors

BabelSteering:通过英语引导向量实现多语言安全对齐
Stein, Emma V., Meier, Dominik, Ruas, Terry, Wahle, Jan Philip, Gipp, Bela
Abstract
Large language models (LLMs) are deployed globally in high-stakes settings, yet most safety research and alignment efforts remain concentrated on English. Thus, users interacting with LLMs in other languages may encounter weaker safeguards despite relying on the same systems for similarly sensitive tasks. In this work, we investigate whether safety signals learned from a high-resource language, like English, can improve multilingual safety. We propose BabelSteering, an activation steering method that acts as a lightweight inference- time intervention, using refusal directions derived from English safety supervision to generalize across languages. Our evaluation includes eight languages and jointly measures refusal of harmful requests, over-refusal, and general task utility. The results show that BabelSteering increases the refusal of harmful requests across languages, with only a marginal to no reduction in task utility but with some increase in refusal of pseudo-harmful prompts. For example, for Gemma 7B, we see an average increase in the refusal of harmful prompts across languages of 11 percentage points (pp), with individual languages like Bengali seeing an increase of 17 pp, with no loss of utility on Global MMLU, while pseudo-harmful refusals increase by 13 pp on average. We also introduce a multilingual translation-and-evaluation pipeline to facilitate future work on cross-lingual safety interventions. Overall, our findings suggest that activation steering may provide a practical, low- cost mechanism for extending English-derived safety signals to other languages. Warning: this paper contains examples with unsafe content
Chinese Translation
大型语言模型(LLMs)在高风险环境中被全球部署,但大多数安全研究和对齐工作仍集中于英语。因此,在其他语言中与LLMs互动的用户可能会遇到较弱的保护措施,尽管他们在类似敏感任务中依赖相同的系统。在本研究中,我们探讨了从高资源语言(如英语)学习的安全信号是否可以改善多语言安全。我们提出了BabelSteering,这是一种激活引导方法,作为轻量级推理时干预,利用从英语安全监督中得出的拒绝方向在语言之间进行泛化。我们的评估包括八种语言,并共同测量有害请求的拒绝、过度拒绝和一般任务效用。结果表明,BabelSteering在多种语言中增加了对有害请求的拒绝,任务效用仅有轻微或没有下降,但对伪有害提示的拒绝有所增加。例如,对于Gemma 7B,我们观察到在多种语言中对有害提示的拒绝平均增加了11个百分点(pp),其中孟加拉语等个别语言的拒绝增加了17 pp,而在Global MMLU上没有效用损失,同时伪有害拒绝平均增加了13 pp。我们还引入了一种多语言翻译和评估管道,以促进未来跨语言安全干预的研究。总体而言,我们的研究结果表明,激活引导可能为将源自英语的安全信号扩展到其他语言提供了一种实用、低成本的机制。警告:本文包含不安全内容的示例
cs.CL / 88 / 2608.16620

Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning

Palmyra x6 技术报告:通过锚定监督微调后训练的代理工具使用模型
Du, Peng, Kamble, Kiran, Vasudev, Rakshith, Yang, Zhizhuo, Nadimpally, Rohith, Krishna, Arjun, Alshikh, Waseem, Bikel, Daniel M.
Abstract
Palmyra x6 is a large language model optimized for use with enterprise-oriented agentic tasks. The model was built by post-training a Mixture-of-Experts base model with Anchored Supervised Fine-Tuning on a compact corpus of verified, synthetic tool-use trajectories, optimized with a Muon + Adam hybrid. The recipe is deliberately conservative and deliberately controlled: 626 trajectories, a single epoch, a low learning rate, and a KL anchor to the frozen base. The model shows substantial gains over the previous default model for Writer Agent, and compares favorably with several recent models on public benchmarks, scoring the highest on BFCL Core at $0.785$ and posts the highest six-benchmark mean of the cohort. Furthermore, the model has shown itself to be competitive or leading relative to comparators in our bias and safety evaluations.
Chinese Translation
Palmyra x6 是一个针对企业导向的代理任务优化的大型语言模型。该模型通过在经过验证的合成工具使用轨迹的紧凑语料库上,使用锚定监督微调对 Mixture-of-Experts 基础模型进行后训练而构建,优化算法为 Muon + Adam 混合。该方法故意保守且经过严格控制:626 条轨迹、单个训练周期、低学习率以及对冻结基础模型的 KL 锚定。该模型在 Writer Agent 的表现上显著优于之前的默认模型,并且在公共基准测试中与几种近期模型相比表现良好,在 BFCL Core 上得分最高为 $0.785$,并在该组中发布了最高的六项基准均值。此外,该模型在我们的偏见和安全评估中显示出与比较模型相竞争或领先的能力。
cs.CL / 89 / 2608.16627

When Do Explanations Help In-Context Learning? A Comparative Study of Natural Language Explanation Types and Faithfulness

何时解释有助于上下文学习?自然语言解释类型及其可信度的比较研究
Dhaini, Mahdi, Dejl, Adam, Vladika, Juraj, Özer, Volkan, Plank, Barbara, Kasneci, Gjergji
Abstract
Natural language explanations (NLEs) are increasingly used as inputs, for example, as few-shot rationales that influence model behavior in in-context learning (ICL). However, it remains unclear how different types of NLEs compare in their effects on downstream model performance in explanation-augmented prompting. Therefore, we provide a comparative evaluation across six benchmarks and four instruction-tuned models, studying how NLE source (human-written when available, self-generated explanations, generated by an external LLM) and NLE selection (random vs faithfulness-based filtering) affect downstream utility of NLEs when used in ICL settings. Our extensive evaluation shows that, on classification-style benchmarks, adding NLEs to few-shot prompts often improves accuracy over few-shot prompting without explanations; among NLE sources, externally generated LLM-NLEs often provide strong downstream utility and remain competitive with human rationales where both are available, whereas self-NLEs are more sensitive to the selection strategy. On math reasoning, the effects are more model- and source-dependent. We further show that faithfulness-based selection of self-NLEs yields small average gains overall, but can improve or reduce performance depending on the metric, task, and model. Different faithfulness metrics can disagree substantially, affecting which self-NLE examples are selected and their downstream predictive utility. Robustness tests with randomly swapped and out-of-distribution rationales indicate partial robustness, suggesting that semantic alignment contributes to performance gains. Overall, our results provide insights for selecting and reporting explanations that influence model behavior in practical prompting pipelines.
Chinese Translation
自然语言解释(NLEs)越来越多地被用作输入,例如,作为影响上下文学习(ICL)中模型行为的少量示例理由。然而,不同类型的 NLEs 在解释增强提示中的影响对下游模型性能的比较仍不清楚。因此,我们在六个基准和四个指令调优模型上进行了比较评估,研究 NLE 来源(可用时为人类撰写、自我生成的解释、由外部大型语言模型(LLM)生成)和 NLE 选择(随机与基于可信度的过滤)如何影响在 ICL 环境中使用 NLEs 的下游效用。我们的广泛评估表明,在分类风格的基准上,将 NLEs 添加到少量提示中通常会提高准确性,相较于没有解释的少量提示;在 NLE 来源中,外部生成的 LLM-NLEs 通常提供强大的下游效用,并在可用的情况下与人类理由保持竞争力,而自我 NLEs 对选择策略更为敏感。在数学推理方面,效果更依赖于模型和来源。我们进一步表明,基于可信度的自我 NLEs 选择总体上带来小幅平均增益,但根据指标、任务和模型的不同,可能会提高或降低性能。不同的可信度指标可能存在显著分歧,影响所选择的自我 NLE 示例及其下游预测效用。通过随机交换和分布外理由的稳健性测试表明部分稳健性,暗示语义对齐有助于性能提升。总体而言,我们的结果为选择和报告影响模型行为的解释提供了见解,以便在实际提示管道中使用。
cs.CL / 90 / 2608.16643

Toward Better Assessment of LLMs' Performance in Clinical Error Detection

改善大型语言模型在临床错误检测中的性能评估
Zhang, Yifan, Beheshti, Rahmatollah
Abstract
Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation. Error-detection benchmarks are typically constructed by injecting errors into notes, such that each erroneous note has a natural counterpart. Aggregate discriminative metrics (e.g., balanced accuracy or F1) do not exploit this structure. We show that this omission is consequential. In particular, evaluating 15 diverse LLMs on 4 standardized clinical error-detection test sets across 3 languages, we find that 13 of 15 models fall below the level of random pairwise discrimination, even while achieving F1 scores that standard practice would read as moderate. We also observe that the underlying bias patterns differ across languages: the same model can default to "no error" on one language and over-flag errors on another. To diagnose where discrimination breaks down, we further introduce a procedure to score the evidence models cite in their outputs. We find that while models consistently locate error-relevant content, they fail to produce the corresponding correct verdict on the clean counterpart. Finally, we show that F1 and pairwise accuracy are driven in opposite directions by the same underlying bias, so that ranking models by F1 may systematically promote the weakest discriminators. For safety-critical clinical NLP applications, we advocate for supplementing aggregate metrics with paired evaluations in benchmark reporting. Code and analysis scripts are available at https://github.com/healthylaife/paired-clinical-eval.
Chinese Translation
自动检测临床文档中的错误是大型语言模型(LLMs)的一项有前景的应用,但部署此类模型的决策依赖于对每个临床记录的孤立评估基准。错误检测基准通常通过在记录中注入错误来构建,使得每个错误记录都有一个自然的对应项。聚合判别指标(例如,平衡准确率或F1分数)未能利用这一结构。我们表明,这一遗漏是有影响的。特别是在对15种不同的LLMs在3种语言的4个标准化临床错误检测测试集上进行评估时,我们发现15个模型中有13个的表现低于随机成对判别的水平,即使它们的F1分数在标准实践中被视为中等。我们还观察到,潜在的偏差模式在不同语言之间存在差异:同一模型在一种语言上可能默认“无错误”,而在另一种语言上则过度标记错误。为了诊断判别失效的原因,我们进一步引入了一种程序来评分模型在输出中引用的证据。我们发现,尽管模型始终能够定位与错误相关的内容,但它们未能对干净的对应项给出正确的判决。最后,我们表明,F1分数和成对准确率在同一潜在偏差的驱动下朝相反的方向变化,因此按F1对模型进行排名可能系统性地提升最弱的判别者。对于安全关键的临床自然语言处理应用,我们建议在基准报告中用成对评估补充聚合指标。代码和分析脚本可在 https://github.com/healthylaife/paired-clinical-eval 获取。
cs.CL / 91 / 2608.16647

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

每枚硬币都有两面:关于大语言模型在线蒸馏中的泛化双重性
Li, Zhaoyi, Kong, Deyang, Wei, Yuan, Yang, Evan, Shen, Ranran, Ihsani, Mahardika Krisna, Yang, Ming, Zhang, Wei, Hao, Chuan, Yang, Jian, Tao, Ran, Dai, Bryan, Zhang, Shikun, Ye, Wei, Wei, Ying, Lian, Defu
Abstract
On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful. Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities. These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.
Chinese Translation
在线蒸馏(On-policy distillation, OPD)通过监督从学生自身策略中采样的轨迹来转移教师能力,但其泛化行为仍然不甚明了,因为大多数研究在单一领域和与训练数据接近的基准上评估 OPD。我们进行了一项受控研究,逐一变化一个泛化因素,从领域内分布转移到跨领域转移以及多教师设置。我们发现,OPD 转移的是教师的推理行为,而不是其对特定问题的答案:训练难度几乎没有影响,即使是教师从未解决的问题也很有用。转移在很大程度上依赖于教师与学生之间的来源关系:同源对使学生在语言、推理视野甚至其他领域上接近教师,而异源对则主要适应训练分布。这种广泛的影响是一把双刃剑:由于将提示路由到领域专家无法限制每位教师的影响,组合它们会导致能力之间的混合依赖性摇摆。这些结果阐明了 OPD 何时能够泛化,并为诊断多教师 OPD 提供了有益的视角。
cs.CL / 92 / 2608.16650

PCA-guided Activation Scaling for Monotonic Bidirectional Control over LLM Sycophancy

基于主成分分析的激活缩放用于对大型语言模型的单调双向控制
Chen, Zheng, Feng, Zhaoxin, Po, Yip Tin, Ma, Jianfei, Chersoni, Emmanuele, Li, Bo
Abstract
Large language models (LLMs) exhibit sycophancy, a tendency to agree with user beliefs regardless of factual accuracy. This can reinforce misconceptions, but eliminating it entirely risks over-correction against valid opinions. Effective control must therefore both reduce and increase sycophancy with predictable and gradual effect. Yet, existing methods fail to ensure a bidirectional and monotonic relationship between steering strength and behavioral outcome across models and datasets. We introduce PCA-guided Activation Scaling (PAS), an activation steering framework that decomposes residual stream activations into a PCA-identified sycophancy-honesty subspace and an orthogonal residual, then applies distinct scaling exponents to achieve monotonic, bidirectional control. Across three LLMs and three datasets, PAS achieves strong monotonicity (Spearman $\rho$ = +0.92) and an average shift of 15.4% per direction, compared with 8.7% for the baselines. Ablation studies confirm that the decomposition, asymmetric exponents, and layer selection are each essential for maintaining monotonic control. The data and code are available at https://github.com/Bellafc/PCS.
Chinese Translation
大型语言模型(LLMs)表现出迎合性,即倾向于与用户信念达成一致,无论事实准确性如何。这可能会强化误解,但完全消除迎合性又可能导致对有效意见的过度修正。因此,有效的控制必须在可预测和渐进的效果下同时减少和增加迎合性。然而,现有方法未能确保在模型和数据集之间引导强度与行为结果之间的双向和单调关系。我们提出了基于主成分分析的激活缩放(PCA-guided Activation Scaling, PAS),这是一种激活引导框架,它将残差流激活分解为主成分分析识别的迎合性-诚实子空间和正交残差,然后应用不同的缩放指数以实现单调的双向控制。在三种大型语言模型和三个数据集上,PAS实现了强单调性(Spearman $ ho$ = +0.92),每个方向的平均偏移为15.4%,而基线为8.7%。消融研究确认,分解、非对称指数和层选择对于维持单调控制都是至关重要的。数据和代码可在 https://github.com/Bellafc/PCS 获取。
cs.CL / 93 / 2608.16671

Does the LM Head Create a Harmful Gradient Bottleneck? A Causal Test

语言模型头是否造成有害的梯度瓶颈?一个因果测试
Murugan, Anand
Abstract
The language-model head maps a hidden state of width D to a vocabulary of size V, so its transpose can return at most D independent directions to the Transformer. Godey and Artzi argue that this severe projection is a harmful optimization bottleneck. We separate the geometry from the causal claim. Our backward-only intervention keeps the ordinary logits and the exact LM-head parameter update while reducing only the rank of the gradient sent into the Transformer. Across five paired seeds on byte-level and BPE-8192 WikiText-2 models, reducing backward rank increases validation loss. An equally ranked factorized forward head, however, increases loss substantially more. At half rank in the larger model, the backward-only loss increase is 0.0586 (95% CI [0.0167, 0.1005]), while the factorized forward head increases loss by 0.1795 ([0.1547, 0.2042]). The vocabulary-space residual also contributes to the ordinary LM-head update, and removing that contribution is harmful. Additional controls show that repeated-token failures are confounded by the number of independently sampled symbols, that adding never-target output classes does not impair learning, and that projection diagnostics do not reliably predict progress in our runs. Tested auxiliary feedback routes do not beat tuned backpropagation. These results confirm strong geometric compression but do not establish that it is a harmful optimization bottleneck.
Chinese Translation
语言模型头将宽度为 D 的隐藏状态映射到大小为 V 的词汇,因此其转置最多可以向 Transformer 返回 D 个独立方向。Godey 和 Artzi 认为,这种严重的投影是一个有害的优化瓶颈。我们将几何结构与因果声明分开。我们的仅向后干预保持了普通的 logits 和精确的 LM-head 参数更新,同时仅减少了发送到 Transformer 的梯度的秩。在字节级和 BPE-8192 WikiText-2 模型上进行的五组配对实验中,减少向后秩会增加验证损失。然而,具有相同秩的因子化前向头则会显著增加损失。在较大模型中,当秩减半时,仅向后损失增加为 0.0586(95% CI [0.0167, 0.1005]),而因子化前向头则增加损失 0.1795([0.1547, 0.2042])。词汇空间的残差也对普通的 LM-head 更新有贡献,去除该贡献是有害的。额外的控制显示,重复标记失败与独立采样符号的数量混淆,添加从未目标的输出类别不会损害学习,并且投影诊断无法可靠预测我们的运行进展。测试的辅助反馈路径未能超过调优的反向传播。这些结果确认了强几何压缩,但并未证明它是一个有害的优化瓶颈。
cs.CL / 94 / 2608.16707

Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors

语义赌博者:上下文中的探索-开发受语义先验的偏见影响
Austin, David Eric, Suleman, Kaheer, Cheung, Jackie Chi Kit
Abstract
Large language models (LLMs) are increasingly deployed as decision-making agents in settings that require sophisticated environmental exploration. However, existing work has raised questions about how LLMs actually balance exploration and exploitation. Unlike classical agents, LLM agents engage with tasks through natural language, exposing them to semantic information with no formal counterpart in the task structure. We introduce the semantic bandit, an extension of the multi-armed bandit setting that explicitly considers the textual labels assigned to actions, and use it to study how semantic priors --- inductive biases arising from associations between language and expected reward learned during pre-training, shape LLM exploration behaviour. We find that semantically informative action labels reduce exploration in favour of exploitation, improving performance when aligned with the reward structure and severely degrading it when misaligned. We further find that negative rewards trigger substantially more exploration than equivalent positive rewards, consistent with an expected-scale bias induced by reward conventions common in pre-training data. Overall, we argue that the use of language to define the environment and rewards introduces unavoidable biases derived from the fact that the model is trained on word co-occurence, with implications for the reliability and robustness of LLM agents in real-world decision-making settings.
Chinese Translation
大型语言模型(LLMs)越来越多地被作为决策代理部署在需要复杂环境探索的场景中。然而,现有研究对LLMs如何平衡探索与开发提出了质疑。与经典代理不同,LLM代理通过自然语言与任务互动,使其接触到任务结构中没有正式对应的语义信息。我们引入了语义赌博者,这是多臂赌博机设置的扩展,明确考虑分配给动作的文本标签,并利用它研究语义先验——在预训练过程中从语言与预期奖励之间的关联中产生的归纳偏见,如何塑造LLM的探索行为。我们发现,语义信息丰富的动作标签减少了探索,倾向于开发,当与奖励结构对齐时提高了性能,而当不对齐时则严重降低了性能。我们进一步发现,负奖励引发的探索显著多于等效的正奖励,这与预训练数据中常见的奖励惯例所诱导的预期规模偏见一致。总体而言,我们认为使用语言来定义环境和奖励引入了不可避免的偏见,这源于模型在词共现上进行训练,这对LLM代理在现实世界决策环境中的可靠性和稳健性具有重要影响。
cs.CL / 95 / 2608.16798

ClawGym II: Exploring Black-Box RL on Agent Harness

ClawGym II:在代理工具上探索黑箱强化学习
Song, Huatong, Bai, Fei, Yang, Ming, Li, Renyuan, Deng, Jia, He, Jujie, Zhang, Zhange, Cheng, Daixuan, Xing, Yan, Yun, Qi, Chen, Xuxing, Li, Danyang, Chang, Feng, Hao, Chuan, Tao, Ran, Yang, Jian, Dai, Bryan, Zhao, Wayne Xin, Tang, Mingjie, Wen, Ji-Rong
Abstract
Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment. However, reinforcement learning through complex harnesses remains largely unexplored, as scaling such training to long-horizon agent tasks introduces fundamental challenges. In this work, we present a unified black-box RL framework for stable and scalable optimization of general agents through complex harnesses. Concretely, we first build a sandbox-based execution infrastructure that isolates task environments and harnesses within temporary sandboxes for large-scale concurrent rollouts. We then decouple policy optimization from opaque harness execution and place a serving proxy at the model boundary to capture model calls. To reconstruct multi-turn trajectories and improve training efficiency, we organize the captured calls into prefix trees and further adapt both critic-based PPO and critic-free GRPO to optimize over the recovered tree structure. Meanwhile, we maintain training-inference consistency throughout the optimization process. Finally, we introduce mix-harness training, allowing a single model to be jointly optimized by heterogeneous harnesses. With Qwen3-30A3B, black-box RL improves Pass@1 on ClawGym-Bench by 9.98 and 14.81 points through OpenClaw and Claude Code, respectively, while remaining stable over 200-400 optimization steps. Moreover, the framework yields consistent gains on more challenging tasks such as JobBench and OfficeQA. Overall, our framework enables effective, stable, and scalable optimization of general agents through black-box harnesses, supporting unified training across heterogeneous execution systems.
Chinese Translation
代理工具通过协调代理与环境的互动,显著提高了在长时间跨度任务上的表现。然而,通过复杂的代理工具进行强化学习仍然在很大程度上未被探索,因为将这种训练扩展到长时间跨度的代理任务引入了基本挑战。在本研究中,我们提出了一种统一的黑箱强化学习框架,用于通过复杂的代理工具实现一般代理的稳定和可扩展优化。具体而言,我们首先构建了一个基于沙盒的执行基础设施,将任务环境和代理工具隔离在临时沙盒中,以支持大规模并发的回滚。然后,我们将策略优化与不透明的代理工具执行解耦,并在模型边界放置一个服务代理,以捕获模型调用。为了重建多轮轨迹并提高训练效率,我们将捕获的调用组织成前缀树,并进一步调整基于评论员的PPO和无评论员的GRPO,以在恢复的树结构上进行优化。同时,我们在整个优化过程中保持训练与推理的一致性。最后,我们引入混合代理训练,允许单个模型通过异构代理工具进行联合优化。使用Qwen3-30A3B,黑箱强化学习通过OpenClaw和Claude Code分别将ClawGym-Bench上的Pass@1提高了9.98和14.81点,同时在200-400次优化步骤中保持稳定。此外,该框架在JobBench和OfficeQA等更具挑战性的任务上也取得了一致的增益。总体而言,我们的框架通过黑箱代理工具实现了一般代理的有效、稳定和可扩展的优化,支持在异构执行系统之间的统一训练。
cs.CL / 96 / 2608.16834

Model Hypnosis: Strong control of AI via additive subliminal effects

模型催眠:通过附加的潜意识效应对人工智能的强控制
Boix-Adsera, Enric, Tessler, Benedict
Abstract
We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be systematically combined to strongly control model behavior. Model hypnosis occurs across model families and scales, including in frontier reasoning models, and hypnotic prompts can transfer between models. Because the model is controlled by inconspicuous textual choices, such as paraphrases and typos, model hypnosis presents new challenges and avenues for AI safety, and is a major hurdle for AI interpretability.
Chinese Translation
我们展示了人工智能模型普遍易受一种现象的影响,我们称之为模型催眠。在这种现象中,提示中的个别微弱且似乎无关的线索可以系统性地组合在一起,从而强烈控制模型行为。模型催眠在不同的模型家族和规模中均有发生,包括前沿推理模型,并且催眠提示可以在模型之间转移。由于模型是通过不显眼的文本选择(如意译和拼写错误)来控制的,模型催眠为人工智能安全带来了新的挑战和机遇,并且是人工智能可解释性的一大障碍。
cs.CL / 97 / 2608.16868

Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text

迈向计算来源:在生成文本中携带因果状态证据
Belay, Benjamin
Abstract
A language model's output does not by itself provide verifiable evidence about the internal computation that produced it. We study computational provenance: whether generated text can carry detectable evidence of which causally relevant internal state occurred. We test a bounded form of this idea in two controlled architectures: a modular feed-forward neural network and a transformer-based model. Both architectures are trained on the same arithmetic task with a mandatory pathway through two discrete intermediate states, allowing different internal paths to produce the same answer. We deliberately switch between these paths, authenticate the state actually used, and let that verified state determine a subtle statistical pattern in the generated text that can later be detected. The feed-forward and transformer systems each passed all 128 matched pairs in both their public and separately sealed protected end-to-end evaluations, with the detector recovering the signal associated with the authenticated internal state. The required causal computation also reproduced across five independently trained feed-forward models and three independently trained transformers. In a separate answer-only transformer experiment, our linear probes did not recover a naturally learned intermediate state. These results provide a controlled proof of concept that information about a verified, causally relevant internal state can be preserved in generated text even when the answer is unchanged.
Chinese Translation
语言模型的输出本身并不能提供关于生成该输出的内部计算的可验证证据。我们研究计算来源:生成的文本是否能够携带可检测的证据,表明发生了哪些因果相关的内部状态。我们在两种受控架构中测试这一想法的有限形式:一个模块化的前馈神经网络和一个基于变换器(transformer)的模型。这两种架构在同一算术任务上进行训练,必须经过两个离散的中间状态,从而允许不同的内部路径产生相同的答案。我们故意在这些路径之间切换,验证实际使用的状态,并让该验证状态决定生成文本中的微妙统计模式,后者可以被检测到。前馈和变换器系统在其公共和单独密封的保护端到端评估中均通过了所有128对匹配样本,检测器成功恢复了与经过认证的内部状态相关的信号。所需的因果计算在五个独立训练的前馈模型和三个独立训练的变换器中也得到了再现。在一个单独的仅回答的变换器实验中,我们的线性探针未能恢复自然学习的中间状态。这些结果提供了一个受控的概念验证,表明关于经过验证的、因果相关的内部状态的信息可以在生成文本中得以保留,即使答案保持不变。