← Back to Index
Daily Research Digest

arXiv Papers

2026-07-30
195
Papers
4
Categories
195
Translated
收藏清单 0
机器人学 (Robotics)
27
cs.RO / 1 / 2607.26121

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

迈向可信的具身智能:一个系统框架和分级可信度水平
Yang, Xinyu, Chen, Tianxing, Su, Honghao, Wang, Minxuan, Yu, Chenze, Tu, Zhangzheng, Chen, Yue, Huo, Yuxiao, Zhang, Lingfeng, Huang, Yan, Qin, Yan, Zhu, Shaolong, Liang, Qiwei, Tian, Hekun, Liu, Shujia, Chen, Guangyu, Gong, Junhao, Li, Zixuan, Lin, Wenwei, Lin, Zijian, Zhu, Wenxuan, Chen, Eric J, Yuan, Yue, Yu, Qize, Liang, Jiaqi, Yan, Haowen, Zhao, Hengfei, Wan, Weijie, Xiao, Zikun, Tang, Junyuan, Chen, Baijun, Lei, Kai-Chong, Wang, Kaixuan, Su, Kailun, Chen, Zanxin, Mu, Yao, Xu, Renjing, Lyu, Chuqiao, Xiong, Qi, Luo, Ping, Ding, Wenbo
Abstract
Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures can cause immediate physical or operational harm, task completion alone does not establish trustworthiness. We define trustworthy embodied intelligence as the sustained capacity to execute specified tasks reliably under environmental and system variation while maintaining risk within acceptable bounds. We term this objective sustained safe success. Its supporting mechanisms are organized into four interdependent layers. The model layer generates task-competent action proposals with calibrated uncertainty and explicit safety preferences. The system layer realizes authorized actions dependably through integrated sensing, computation, control, hardware safeguards, fault containment, and fallback. The evidence layer substantiates bounded claims through evaluation, verification, validation, traceability, and structured assurance arguments. The deployment layer maintains claim validity through runtime monitoring, authority management, intervention, incident response, and controlled updates. Because assumptions and failures propagate across these layers, neither model capability, isolated safeguards, nor benchmark performance alone can establish end-to-end trustworthiness. Drawing on embodied AI, robotics, control, dependable computing, distributed systems, and autonomous driving, we further propose a non-normative hierarchy of trustworthiness levels. This hierarchy grades the strength of bounded deployment claims across task capability, safety, system assurance, operational governance, and supporting evidence, providing a basis for bounded deployment, comparative evaluation, research prioritization, and future standardization.
Chinese Translation
具身智能将学习到的感知和决策与实时计算、控制和物理交互相结合。由于故障可能导致即时的物理或操作损害,仅仅完成任务并不能建立可信度。我们将可信的具身智能定义为在环境和系统变化下,持续可靠地执行特定任务的能力,同时将风险保持在可接受的范围内。我们将这一目标称为持续安全成功。其支持机制被组织为四个相互依赖的层次。模型层生成具有校准不确定性和明确安全偏好的任务能力行动建议。系统层通过集成感知、计算、控制、硬件保护、故障隔离和后备措施可靠地实现授权行动。证据层通过评估、验证、确认、可追溯性和结构化保证论证来证实有界声明。部署层通过运行时监控、权限管理、干预、事件响应和受控更新来维护声明的有效性。由于假设和故障在这些层次之间传播,因此单靠模型能力、孤立的保护措施或基准性能都无法建立端到端的可信度。基于具身人工智能、机器人技术、控制、可靠计算、分布式系统和自动驾驶,我们进一步提出了一种非规范性的可信度水平层次结构。该层次结构对任务能力、安全性、系统保证、操作治理和支持证据的有界部署声明的强度进行分级,为有界部署、比较评估、研究优先级和未来标准化提供了基础。
cs.RO / 2 / 2607.26148

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation

具身代理掌控:最小接口零-shot代理在视觉与语言导航中与工业规模策略相媲美
Zhou, Jian, Zhao, Xunyi, Zhou, Gengze, Li, Zerui, Lin, Sihao, Liu, Jiajun, Wu, Qi
Abstract
Autonomous embodied agents must sustain a long decision-making loop that involves perceiving, acting, verifying, and self-correcting over many steps. Current systems sustain this loop through task-specific workflows or embodied policies. We study a third form, agentic embodied control, in which a general-purpose agent holds the loop itself. Using zero-shot navigation as a controlled testbed, we evaluate three software-engineering agent harnesses given only a monocular RGB camera and discrete actions. Under this strictly minimal condition, replicated default-effort configurations reach 70.7$\pm$3.5% success (opus-5, mean over three runs), and fable-5 reaches 78% at maximum effort. When a trained waypoint tool is exposed alongside primitives as an optional capability, the hybrid fable-5 agent reaches 76.7$\pm$0.6% at default effort, using half the environment steps and less than one quarter of the wall time of the maximum-effort primitive run. Controlled interventions show that capability is primarily model-centered: model choice strongly changes success, harness effects are descriptive, and a forced waypoint interface helps weaker models but can hinder stronger ones. Performance nevertheless falls sharply on longer-horizon tasks, while latency and context growth limit sustained operation. These results show that agentic control is already competitive in zero-shot navigation and that models, harnesses, and interfaces offer complementary paths toward autonomous embodied agents.
Chinese Translation
自主具身代理必须维持一个长时间的决策循环,该循环涉及在多个步骤中感知、行动、验证和自我纠正。当前系统通过特定任务的工作流程或具身策略来维持这一循环。我们研究第三种形式,即代理具身控制,其中一个通用代理掌握整个循环。利用零-shot导航作为受控测试平台,我们评估了三种软件工程代理工具,仅使用单目RGB相机和离散动作。在这一严格的最小条件下,复制的默认努力配置达到了70.7$ ext{±}$3.5%的成功率(opus-5,三次运行的平均值),而fable-5在最大努力下达到了78%。当一个训练好的航点工具与原始能力一起作为可选功能暴露时,混合的fable-5代理在默认努力下达到了76.7$ ext{±}$0.6%的成功率,使用的环境步骤仅为最大努力原始运行的一半,墙面时间不到四分之一。受控干预显示,能力主要以模型为中心:模型选择显著改变成功率,工具的效果是描述性的,而强制航点接口有助于较弱的模型,但可能会阻碍较强的模型。然而,性能在较长时间任务上急剧下降,而延迟和上下文增长限制了持续操作。这些结果表明,代理控制在零-shot导航中已经具有竞争力,并且模型、工具和接口提供了通向自主具身代理的互补路径。
cs.RO / 3 / 2607.26279

Multi-Objective Compliance-Integrated Coevolution For Simulated And Real-World Deployment Of Multi-Robot Marine Autonomy

多目标合规集成共进化用于多机器人海洋自主的模拟与现实世界部署
Gonzalez, Everardo, Paine, Tyler M., Vallejo, Manuel Agraz, Dixit, Gaurav, Benjamin, Michael R., Tumer, Kagan
Abstract
Collaborative robots are well-suited to maritime missions that benefit from coordination, such as the exploration of unknown reef structures, inspection of subsea infrastructure, or search-and-rescue operations. These missions typically provide sparse feedback signals for measuring progress and require adherence to safety and regulatory norms, turning a mission into a multi-objective optimization problem. Coevolutionary algorithms can process these sparse feedback signals to generate coordinated behaviors, and in some cases extend behaviors to multiple objectives. However, incorporating high-level team objectives with low-level compliance considerations on the fly to balance norm adherence with team performance remains elusive. This paper introduces a multi-objective framework that blends coevolved behaviors with compliance behaviors to achieve a balance between maximizing team progress and minimizing norm violations. The key insight is to decouple learning from compliance since operational norms are prescribed rather than discovered. We demonstrate that our framework achieves high team performance while avoiding collisions on a collaborative swimmer rescue mission with up to 8 vehicles in a hardware deployment, and 12 vehicles in simulation. The key contribution of this paper is Marine Multi-Objective Compliance-Integrated Coevolution (MMOCIC), a framework that blends team-wide optimization with established norms for real-world deployments of learning-based coordination.
Chinese Translation
协作机器人非常适合需要协调的海洋任务,例如探索未知的珊瑚礁结构、检查海底基础设施或进行搜救行动。这些任务通常提供稀疏的反馈信号来衡量进展,并要求遵循安全和监管规范,从而将任务转化为多目标优化问题。共进化算法能够处理这些稀疏的反馈信号以生成协调行为,并在某些情况下将行为扩展到多个目标。然而,如何在飞行中将高层团队目标与低层合规考虑结合起来,以平衡规范遵循与团队表现仍然是一个难题。本文提出了一种多目标框架,将共进化行为与合规行为相结合,以实现最大化团队进展与最小化规范违反之间的平衡。关键的见解在于将学习与合规解耦,因为操作规范是规定的而非发现的。我们展示了我们的框架在一个协作游泳者救援任务中实现了高团队表现,同时避免了碰撞,在硬件部署中可支持多达8辆车辆,在模拟中可支持12辆车辆。本文的关键贡献是海洋多目标合规集成共进化(Marine Multi-Objective Compliance-Integrated Coevolution,MMOCIC),这是一个将团队优化与既定规范相结合的框架,适用于基于学习的协调在现实世界中的部署。
cs.RO / 4 / 2607.26315

MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization

MoMo:机器人操作中的运动模式与时空动作标记化
Hu, Yuhan, Thomas, Hugues, Huang, Peide, Sivapurapu, Mouli, Landry, Benoit, Kivila, Arto
Abstract
To operate effectively across diverse contexts, robots must not only perform manipulation tasks accurately but also adapt how their actions unfold to the task, object, and interaction setting. We ask whether this execution-level variation can be learned as a reusable behavioral factor shared across tasks. We present \textbf{MoMo}, a two-stage imitation-learning framework consisting of a spatiotemporal action tokenizer and a behavior-cloning transformer that takes task and a continuous motion-mode condition as inputs. Across six real-robot manipulation tasks, varying this condition produces steady, dynamic, and intermediate behaviors that human raters can distinguish and that differ in joint speed, acceleration, and end-effector approach pitch. On tasks demonstrated in only one mode, MoMo transfers the unseen requested mode while largely preserving task success. Together, these results provide evidence of compositional generalization to unseen task--mode combinations and show that motion mode can be reused across tasks to control how a manipulation skill is performed.
Chinese Translation
为了在多样化的环境中有效操作,机器人不仅需要准确执行操作任务,还必须根据任务、物体和交互环境调整其动作展开方式。我们探讨这种执行层面的变化是否可以作为一种可重用的行为因子在任务间共享。我们提出了 extbf{MoMo},一个由时空动作标记器和行为克隆变换器组成的两阶段模仿学习框架,该框架以任务和连续运动模式条件作为输入。在六个真实机器人操作任务中,改变这一条件产生了稳定的、动态的和中间的行为,这些行为可以被人类评估者区分,并且在关节速度、加速度和末端执行器接近角度上存在差异。在仅以一种模式演示的任务中,MoMo能够转移未见的请求模式,同时在很大程度上保持任务成功。这些结果共同提供了对未见任务-模式组合的组合泛化的证据,并表明运动模式可以在任务间重用,以控制操作技能的执行方式。
cs.RO / 5 / 2607.26337

Reeling It In: Flexible Needle Pick Up via Thread Manipulation for Autonomous Suturing

收回针线:通过线材操控实现灵活的针头拾取以进行自主缝合
Huang, Emma, Chiu, Zih-Yun, Joglekar, Neelay, Liu, Shanglei, Yip, Michael C.
Abstract
Suture-needle pickup is necessary for autonomous suturing, as a needle can be unexpectedly dropped or strategically released to adjust the grasping configuration. Current methods for autonomous needle pickup typically guide a robot to move straight toward the needle and grasp it, limited to conditions where the needle is observable and directly approachable. In addition, grasping the needle lying on tissue can lead to the robot pinching nearby tissue or the needle jumping around due to its slippery surface, posing potential safety issues. This work proposes an autonomous framework that uses a suture thread as an assistive tool for indirect needle pickup, avoiding unnecessary tool-tissue contact and enabling pickup even when the needle is occluded or inaccessible. The framework spans the entire workflow, including thread and tissue reconstruction, safe grasp-point selection, stable thread lifting, and bimanual thread-following until securing needle grasping. The robot policies account for visual uncertainty to maximize robustness in real-world environments. We evaluate the proposed framework on a da Vinci Research Kit under various real-world conditions. The results demonstrate robust performance even with a challenging thread configuration or a non-approachable needle, closing the gap in applying autonomous robot policies to unstructured suturing environments.
Chinese Translation
针头拾取是自主缝合所必需的,因为针头可能会意外掉落或被策略性释放以调整抓取配置。目前的自主针头拾取方法通常引导机器人直线移动到针头处并进行抓取,这限制了其在针头可见且可直接接近的条件下使用。此外,抓取躺在组织上的针头可能导致机器人夹住附近的组织,或由于针头表面光滑而导致针头跳动,从而带来潜在的安全问题。本研究提出了一种自主框架,利用缝合线作为辅助工具进行间接针头拾取,避免不必要的工具与组织接触,并在针头被遮挡或无法接近时仍能进行拾取。该框架涵盖了整个工作流程,包括线材和组织重建、安全抓取点选择、稳定的线材提升以及双手跟随线材直至安全抓取针头。机器人策略考虑了视觉不确定性,以最大化在现实环境中的鲁棒性。我们在多种现实条件下对提出的框架进行了评估,结果表明即使在复杂的线材配置或无法接近的针头情况下,仍表现出稳健的性能,缩小了将自主机器人策略应用于非结构化缝合环境的差距。
cs.RO / 6 / 2607.26370

Self-Adaptive Learning and Model Predictive Control for Tracking Unknown Dynamics with No Regret

自适应学习与模型预测控制用于无悔追踪未知动态
Navsalkar, Atharva, Zhou, Hongyu, Tzoumas, Vasileios
Abstract
We propose a self-adaptive online learning for control method for tracking unknown target dynamics. The target dynamics can exhibit switching behavior, particularly, a mixture of structured, random, and/or adversarial motion. Such challenging target tracking scenarios arise in applications of dynamic mapping, traffic control, and pursuit evasion, where robots need to track, pursue, or avoid collision with moving landmarks, objects, humans, etc., whose dynamics are unknown. Our method simultaneously learns multiple predictors from scratch, via self-supervised, one-shot, and computationally efficient learning, and adaptively selects the best one to match the observed target behavior. The method enjoys finite-time near-optimality guarantees in expectation, characterized as a function of the learning error of the target dynamics and the frequency that the target dynamics switch. In the absence of both error and switching, the method asymptotically matches the optimal non-causal control policy that knows a priori the target dynamics, i.e., the method enjoys no regret in expectation. In the presence of learning errors and switching, the method degrades gracefully, \eg when there are errors and no switching, the average regret is proportional to the average learning error and switching times. To prove these guarantees, a novel technical approach is required compared to the existing works that employ RFF-based online learning. We validate our method in Crazyflie simulations and hardware experiments, across target trajectories that vary from structured to random to adversarial, in comparison to non-stochastic, kernel-based, and neural-network-based methods for online learning.
Chinese Translation
我们提出了一种自适应在线控制学习方法,用于追踪未知目标动态。目标动态可能表现出切换行为,特别是结构化、随机和/或对抗性运动的混合。这种具有挑战性的目标追踪场景出现在动态映射、交通控制和追逐规避等应用中,机器人需要追踪、追逐或避免与移动地标、物体、人类等发生碰撞,而这些动态是未知的。我们的方法通过自我监督、一键式和计算效率高的学习,从零开始同时学习多个预测器,并自适应地选择最佳预测器以匹配观察到的目标行为。该方法在期望上享有有限时间近似最优性保证,其特征是目标动态的学习误差和目标动态切换的频率。在没有误差和切换的情况下,该方法渐近匹配已知目标动态的最优非因果控制策略,即该方法在期望上享有无悔性。在存在学习误差和切换的情况下,该方法优雅地降级,例如,当存在误差且没有切换时,平均悔恨与平均学习误差和切换次数成正比。为了证明这些保证,与现有采用基于随机傅里叶特征(RFF)的在线学习的工作相比,需要一种新颖的技术方法。我们在Crazyflie模拟和硬件实验中验证了我们的方法,针对从结构化到随机再到对抗性的目标轨迹,与非随机、基于核的方法和基于神经网络的方法进行比较。
cs.RO / 7 / 2607.26434

Reinforcement Learning on Cost-Constrained Quadrupedal Hardware

在成本受限的四足硬件上进行强化学习
Weddington, Javier C., Ölveczky, Bence P., Baccus, Stephen A.
Abstract
Deploying learned control policies on low-cost robotic platforms introduces transport latencies and noisy motor feedback that systematically widens the sim-to-real gap. The chasm of simulation to deployment in hardware lies in the delay of the actuator reaching the commanded position. On platforms such as the Mini Pupper 2, a measured > $50 ms transport delay transforms the locomotion task from a standard Markov decision process into a partially observable one. In this paper, we take a biologically inspired approach of handling noisy and delayed feedback to close the sim-to-real gap, thereby expanding the capability of reinforcement learning on cost-constrained hardware. Using a low-cost quadrupedal hardware platform, we find that using a forward model of the average actuator delay, paired with a time-aware neural network results in robust locomotion. Additionally, our time-aware neural network learned a central pattern generator (CPG): a self-sustaining rhythmic gait that is robust to +320 ms latency perturbations, mirroring the CPGs found in the spinal cords of vertebrates. We posit that temporal self-organization may be a general strategy for cost-constrained locomotion.
Chinese Translation
在低成本机器人平台上部署学习到的控制策略会引入传输延迟和噪声电机反馈,从而系统性地扩大了模拟与现实之间的差距。硬件部署中的模拟与现实之间的鸿沟在于执行器达到指令位置的延迟。在像 Mini Pupper 2 这样的平台上,测得的超过 50 毫秒的传输延迟使得运动任务从标准的马尔可夫决策过程转变为部分可观察的过程。本文采用一种生物启发的方法来处理噪声和延迟反馈,以缩小模拟与现实之间的差距,从而扩展了在成本受限硬件上进行强化学习的能力。使用低成本的四足硬件平台,我们发现,结合平均执行器延迟的前向模型与时间感知神经网络可以实现稳健的运动。此外,我们的时间感知神经网络学习到了一个中央模式发生器(Central Pattern Generator, CPG):一种自我维持的节律步态,对超过 320 毫秒的延迟扰动具有鲁棒性,类似于脊椎动物脊髓中的 CPG。我们认为,时间自组织可能是一种适用于成本受限运动的普遍策略。
cs.RO / 8 / 2607.26460

RLMM-Flow: A Flow-based Mobile Manipulation Framework with Latent-Space Reinforcement Learning

RLMM-Flow:一种基于流的移动操控框架,结合潜在空间强化学习
Wang, Shuhang, Li, Ziming, Cheng, Hui
Abstract
Mobile manipulation requires generating whole-body action chunks that jointly satisfy goal reaching, collision avoidance, base kinematic constraints, manipulator joint limits, and trajectory smoothness. Flow-based generative policies provide an efficient paradigm for learning multimodal and temporally consistent motion priors from expert demonstrations, but imitation-only training cannot improve policy quality beyond the demonstration distribution. We propose RLMM-Flow, a flow-based mobile manipulation framework that combines expert flow-policy pretraining with latent-space reinforcement learning post-training. The framework first learns a flow policy that captures a multimodal whole-body motion prior from expert demonstrations. The pretrained flow policy is then frozen, while a latent steering network steers its initial noise toward higher-value action chunks. To stabilize high-dimensional latent optimization, we warm up an action-space critic before jointly training the latent critic and latent actor, and introduce coarse-to-fine latent steering that progressively expands control from a horizon-shared latent representation to a full-dimensional residual representation. Experiments on mobile manipulation motion-planning benchmarks show that RLMM-Flow substantially improves task success, collision avoidance, and trajectory quality over imitation-only flow policies and existing reinforcement learning post-training baselines, while preserving fast flow-based inference.
Chinese Translation
移动操控需要生成整体动作片段,以共同满足目标达成、避免碰撞、基础运动学约束、操控器关节限制和轨迹平滑性。基于流的生成策略提供了一种高效的范式,用于从专家演示中学习多模态和时间一致的运动先验,但仅依赖模仿训练无法提高策略质量超出演示分布。我们提出了RLMM-Flow,一种基于流的移动操控框架,结合了专家流策略的预训练与潜在空间强化学习的后训练。该框架首先学习一个流策略,从专家演示中捕捉多模态的整体运动先验。然后,预训练的流策略被冻结,而潜在引导网络将其初始噪声引导至更高价值的动作片段。为了稳定高维潜在优化,我们在联合训练潜在评论者和潜在演员之前,先对动作空间评论者进行预热,并引入粗到细的潜在引导,逐步将控制从共享的潜在表示扩展到全维的残差表示。在移动操控运动规划基准上的实验表明,RLMM-Flow在任务成功率、碰撞避免和轨迹质量方面显著优于仅依赖模仿的流策略和现有的强化学习后训练基线,同时保持快速的基于流的推理。
cs.RO / 9 / 2607.26513

Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models

基于解析概念的显式运动指导用于视觉-语言-动作模型
Sun, Mingyang, Wei, Jiude, Liang, Xiujian, He, Qichen, Wang, Donglin, Lu, Cewu, Sun, Jianhua
Abstract
Current Vision-Language-Action (VLA) models rely mainly on 2D inputs, neglecting the rich object structural information and commonsense knowledge inherent in the 3D physical world. This deficiency restricts their spatial awareness and adaptability for complex, high-precision manipulation. To bridge this crucial gap, we construct a Concept Expert module for VLA to build executable Analytic Concepts that represent objects as explicit, programmatic blueprints. Our mechanism operates in two synergistic phases: First, prior to VLA inference, the Concept Expert leverages 3D information from Vision Foundation Models (VFMs) to estimate the initial kinematic and structural parameters. Second, throughout the manipulation process, the VLA model utilizes its inherent capability to dynamically track the dynamic concept parameters, continuously aligning them with observational changes to ensure persistent accuracy. Once established, the Analytic Concepts provide explicit, high-quality guidance for VLA fine-tuning through (1) dense, programmatic manipulation rewards and (2) precise spatial guidance. This formulation allows VLA models to learn physically grounded interaction behaviors while maintaining end-to-end learning flexibility. Our experimental results show consistent improvements in success rate and learning efficiency across supervised and reinforcement learning settings, demonstrating the effectiveness of structured, concept-based guidance for VLA post-training.
Chinese Translation
当前的视觉-语言-动作(VLA)模型主要依赖于二维输入,忽视了三维物理世界中固有的丰富物体结构信息和常识知识。这一缺陷限制了它们在复杂、高精度操作中的空间意识和适应能力。为了填补这一重要空白,我们为VLA构建了一个概念专家模块,以构建可执行的解析概念,将物体表示为显式的、程序化的蓝图。我们的机制分为两个协同阶段:首先,在VLA推理之前,概念专家利用来自视觉基础模型(VFM)的三维信息来估计初始的运动和结构参数。其次,在整个操作过程中,VLA模型利用其固有能力动态跟踪动态概念参数,持续将其与观察变化对齐,以确保持续的准确性。一旦建立,解析概念通过(1)密集的程序化操作奖励和(2)精确的空间指导,为VLA的微调提供显式的高质量指导。这一构想使VLA模型能够学习基于物理的交互行为,同时保持端到端学习的灵活性。我们的实验结果显示,在监督学习和强化学习环境中,成功率和学习效率均有持续改善,证明了基于结构的概念指导在VLA后训练中的有效性。
cs.RO / 10 / 2607.26567

Speech2Grasp: Data-Efficient Transfer of Text-Conditioned Grasp Detection to Speech in Humanoid Robots

Speech2Grasp:文本条件抓取检测向语音的数据高效转移在类人机器人中的应用
Nguyen, Hung, Nguyen, Kim Nhat Minh, Vu, Van Duc, Le, Van-Danh, Le, Hoang Huy, Nguyen, Dinh Tuan, Le, Pham Tuyen, Nguyen, Van-Truong, Nguyen, Quan
Abstract
Humanoid robots increasingly require multi-modal understanding for natural interaction with humans. Despite the prominence of vision-language models, they generally assume textual rather than the more natural speech inputs. In this paper, we investigate whether a well-established text-conditioned model can be transferred to speech in a data-efficient manner. Using ALBEF as a case study, we conduct diagnostic analyses showing that a lightweight MLP-based projector effectively adapts it to speech, while preserving semantic discrimination and robustness. Motivated by these findings, we introduce Speech2Grasp, a framework for data-efficient transfer of text-conditioned grasp detection to speech. Real-world humanoid robot experiments show that Speech2Grasp outperforms cascaded ASR-based pipeline, while reducing inference latency. Our findings suggest a practical paradigm for extending established text-conditioned systems to speech.
Chinese Translation
类人机器人越来越需要多模态理解,以实现与人类的自然互动。尽管视觉-语言模型日益受到关注,但它们通常假设输入为文本而非更自然的语音。在本文中,我们探讨了一个成熟的文本条件模型是否可以以数据高效的方式转移到语音上。以 ALBEF 为案例研究,我们进行了诊断分析,表明一个轻量级的基于 MLP 的投影器能够有效地将其适配到语音,同时保持语义区分能力和鲁棒性。基于这些发现,我们提出了 Speech2Grasp,一个将文本条件抓取检测高效转移到语音的框架。实际的类人机器人实验表明,Speech2Grasp 超越了级联的基于 ASR 的管道,同时减少了推理延迟。我们的研究结果为将已建立的文本条件系统扩展到语音提供了一种实用的范式。
cs.RO / 11 / 2607.26570

Semi-Decentralized Multi-Spacecraft Collision Avoidance under Communication Constraints

通信约束下的半去中心化多航天器避碰
Kim, Grace Ra, Al-Husseini, Mahdi, Eddy, Duncan, Kochenderfer, Mykel J.
Abstract
Current spacecraft collision-avoidance operations rely on intermittent ground-station contacts, requiring operators to plan with delayed and asynchronously updated information. Consequently, maneuvers must be planned with only intermittent information sharing between operators, raising the question of how much coordination is needed to achieve collision-avoidance performance comparable to centralized planning. Although decision-theoretic approaches such as partially observable Markov decision processes (POMDPs) capture the sequential and uncertain nature of collision avoidance, existing multiagent extensions typically assume either continuous information sharing or communication models that do not reflect operational ground-station constraints. To explicitly model this intermittent information availability, we formulate the spacecraft-to-spacecraft collision avoidance problem as a semi-decentralized POMDP (SDec-POMDP), where we govern information propagation directly by realistic ground-station visibility windows. Joint maneuver policies are computed using approximate Recursive Small-Step Semi-Decentralized A* (RS-SDA*), following the state-of-the-art A*-based lineage for decentralized multiagent planning. Across a representative suite of conjunction scenarios, semi-decentralized planning recovers near-centralized maneuver quality while requiring 28.5% fewer synchronization events than continuous coordination. Comparisons with representative rule-based operator heuristics further show that communication-aware planning more consistently achieves the desired operational miss-distance band while minimizing unnecessary trajectory deviation. Together, these results establish a practical planning framework for autonomous collision avoidance under realistic intermittent communication, bridging the gap between idealized centralized coordination and fully decentralized planning execution.
Chinese Translation
当前的航天器避碰操作依赖于间歇性的地面站联系,这要求操作员在延迟和异步更新的信息下进行规划。因此,机动必须仅在操作员之间进行间歇性的信息共享下进行规划,这引发了一个问题:为了实现与集中式规划相当的避碰性能,需要多少协调。尽管决策理论方法如部分可观测马尔可夫决策过程(POMDPs)捕捉了避碰的顺序性和不确定性,但现有的多智能体扩展通常假设要么是连续的信息共享,要么是未能反映操作地面站约束的通信模型。为了明确建模这种间歇性的信息可用性,我们将航天器间的避碰问题表述为半去中心化的POMDP(SDec-POMDP),在其中我们通过现实的地面站可见窗口直接控制信息传播。使用近似递归小步半去中心化A*(RS-SDA*)计算联合机动策略,遵循基于A*的去中心化多智能体规划的最新进展。在一系列典型的接近场景中,半去中心化规划恢复了接近中心化的机动质量,同时比连续协调减少了28.5%的同步事件。与代表性的基于规则的操作员启发式方法的比较进一步表明,考虑通信的规划更一致地实现了所需的操作失距带,同时最小化了不必要的轨迹偏离。这些结果共同建立了一个实用的规划框架,用于在现实的间歇性通信下实现自主避碰,弥合了理想化的集中协调与完全去中心化规划执行之间的差距。
cs.RO / 12 / 2607.26579

ContactFlow: A video action conditioning that transfers across embodiments

ContactFlow:一种跨体现的动作条件视频
Azirar, Sami, Pallotta, Enrico, Nogga, Jan, Gall, Jürgen, Behnke, Sven, Blum, Hermann
Abstract
World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. However, current video-based world models often struggle to capture the physical constraints that govern manipulation, particularly contact. Further, their action conditioning is often constrained to specific embodiments such as parallel grippers. We propose \emph{Contact Flow}, an embodiment-agnostic action representation that encodes manipulation through the trajectory of 3D contact points between an actor and a target object. By discarding actor-specific appearance and kinematics, Contact Flow provides a shared conditioning signal for both human demonstrations and robotic execution. Therefore, we can train a large-scale video generative model on both human and robotic object interaction videos conditioned on Contact Flow, yielding a world model that predicts physically plausible manipulation outcomes. We integrate this model into a propose-imagine-verify-act pipeline, where generated rollouts are assessed by a vision-language model before execution. Experiments on the DROID dataset and real-world tabletop manipulation tasks demonstrate that Contact Flow enables transfer between human demonstrations and different robotic embodiments.
Chinese Translation
世界模型为机器人规划提供了一条有前景的途径,使得代理能够在执行之前想象和验证动作的后果。然而,当前基于视频的世界模型往往难以捕捉支配操控的物理约束,特别是接触。此外,它们的动作条件通常仅限于特定的体现,例如并行夹爪。我们提出了 extit{Contact Flow},一种与体现无关的动作表示,通过演员与目标物体之间3D接触点的轨迹来编码操控。通过舍弃演员特定的外观和运动学,Contact Flow 为人类示范和机器人执行提供了共享的条件信号。因此,我们可以在基于Contact Flow条件的人类和机器人物体交互视频上训练一个大规模的视频生成模型,从而生成一个预测物理上合理的操控结果的世界模型。我们将该模型集成到一个提议-想象-验证-执行的流程中,其中生成的滚动预测在执行之前由视觉-语言模型进行评估。在DROID数据集和现实世界的桌面操控任务上的实验表明,Contact Flow 使得人类示范与不同机器人体现之间的转移成为可能。
cs.RO / 13 / 2607.26657

Enfold: Folding World-Generator Computation into Predictive Representations for Efficient Embodied Control

Enfold:将世界生成器计算折叠到预测表示中以实现高效的具身控制
Zeng, Weili, Xing, Yitong, Liu, Fulong, Yang, Chengqun, Xiang, Antao, Tian, Feng, Gao, Jingnan, Cai, Jisong, Wang, Xin, Wu, Xiaomin, Mu, Yao, Yan, Yichao
Abstract
World generative models are typically used through what they produce: a rendered future, a video-conditioned action, or latent context computed by a costly generative branch. We argue that their more reusable asset is the computation that constructs a future. As a generator transforms a corrupted future into a coherent trajectory, its intermediate states organize appearance, spatial layout, and interaction across levels of abstraction. Can this future-generative computation be internalized in a representation inferred from the present alone? We present Enfold, which transfers this computation into a representation predicted from the current visual context and language instruction. During training, multi-level states exposed as the generator processes the observed future supervise a current-only encoder. The learned representation is fed back to condition future generation and is read by task heads without allowing task gradients to reshape the encoder. At deployment, action prediction no longer executes the generator. Across LIBERO, RoboTwin2.0, and real-robot tasks, Enfold supports strong control while reducing action latency by $3.7\times$ relative to Fast--WAM, Enfold-Flash reaches $10.1\times$. Representation analyses show that it suppresses nuisance variation and preferentially captures changes that emerge over longer horizons. When the current scene is altered by human intervention, both the generated continuation and the executed actions adapt, which is inconsistent with fixed trajectory replay. These results recast a world generator as a source of predictive control representations: its future need not be materialized at every step if its internal structure can be enfolded into the present.
Chinese Translation
世界生成模型通常通过其生成的内容使用:渲染的未来、视频条件下的动作或由成本高昂的生成分支计算出的潜在上下文。我们认为,它们更可重用的资产是构建未来的计算。随着生成器将一个损坏的未来转化为一个连贯的轨迹,其中间状态组织了外观、空间布局和跨抽象层次的交互。这个未来生成计算能否在仅从当前状态推断的表示中内化?我们提出了Enfold,它将这一计算转移到从当前视觉上下文和语言指令预测的表示中。在训练过程中,生成器处理观察到的未来时暴露出的多层状态监督一个仅基于当前的编码器。学习到的表示被反馈用于条件未来生成,并被任务头读取,而不允许任务梯度重塑编码器。在部署时,动作预测不再执行生成器。在LIBERO、RoboTwin2.0和真实机器人任务中,Enfold支持强大的控制,同时将动作延迟相对于Fast--WAM减少了$3.7 imes$,Enfold-Flash达到了$10.1 imes$。表示分析表明,它抑制了干扰变化,并优先捕捉在更长时间范围内出现的变化。当当前场景受到人类干预时,生成的延续和执行的动作都会适应,这与固定轨迹重放不一致。这些结果将世界生成器重新定义为预测控制表示的来源:如果其内部结构可以被折叠到当前状态中,则其未来不必在每一步都具现化。
cs.RO / 14 / 2607.26712

ActSWM: Action-Sensitive World Models for Long-Horizon Planning in Open-World Games

ActSWM:用于开放世界游戏的长远规划的动作敏感世界模型
Gan, Zhenfeng, Zeng, ZiTong, Cheng, Jiajun, Song, Yeke, Tang, Yongyi, Wang, Xueqian
Abstract
Latent world models support efficient model-predictive control by optimizing future control sequences in latent space and replanning in a receding-horizon manner. However, existing latent predictors often lack stable long-horizon rollout ability, and prediction accuracy alone does not ensure that rollouts remain responsive to the actions being planned. We identify Context Collapse, a failure mode in which autoregressive latent predictors maintain high similarity to future states while producing nearly indistinguishable futures under different action sequences. To address this issue, we propose ActSWM, an action-sensitive latent world model grounded in a transition-separation principle: a planning-useful latent dynamics model should keep alternative-action futures distinguishable and make the action associated with each local transition recoverable. Under this principle, action sensitivity is enforced as a constraint on latent rollouts rather than treated only as an auxiliary prediction target, encouraging predicted futures to preserve action-dependent differences over long horizons. Across step-drift analysis, closed-loop Minecraft planning, and cross-game local action recovery, ActSWM preserves larger action-dependent rollout gaps than existing baselines, improves task success in long-horizon interactive settings, and enables world-model-based action recovery from offline gameplay videos.
Chinese Translation
潜在世界模型通过在潜在空间中优化未来控制序列并以递归视野的方式进行重新规划,支持高效的模型预测控制。然而,现有的潜在预测器往往缺乏稳定的长远展开能力,单靠预测准确性并不能确保展开对计划中的动作保持响应。我们识别出上下文崩溃(Context Collapse),这是一种故障模式,其中自回归潜在预测器在保持与未来状态的高度相似性的同时,在不同的动作序列下产生几乎无法区分的未来。为了解决这个问题,我们提出了ActSWM,一种基于转移分离原则的动作敏感潜在世界模型:一个对规划有用的潜在动态模型应该保持替代动作未来的可区分性,并使与每个局部转移相关的动作可恢复。在这一原则下,动作敏感性被作为潜在展开的约束来强制执行,而不仅仅被视为辅助预测目标,从而鼓励预测的未来在长远视野中保持依赖于动作的差异。在步骤漂移分析、闭环Minecraft规划和跨游戏局部动作恢复的实验中,ActSWM保持了比现有基线更大的依赖于动作的展开差距,提高了长远交互环境中的任务成功率,并使基于世界模型的动作恢复能够从离线游戏视频中实现。
cs.RO / 15 / 2607.26770

Vision-TL-Action: Neuro-Symbolic Trajectory Generation from Visual Observations and Temporal Logic

视觉-时序逻辑-动作:基于视觉观察和时序逻辑的神经符号轨迹生成
Liu, Zezhi, Zheng, Zhiwei, Luo, Hanqian, Qin, Deyun, Wu, Shizhen, Fang, Yongchun
Abstract
Temporal logic (TL) provides a compositional language for the formulation of long horizon robotic tasks, but existing TL-conditioned trajectory generators can sidestep perception-to-symbol binding by encoding exact object geometry in the task graph. We introduce \emph{Vision-TL-Action}, which generates action trajectories from multi-view images, a coordinate-free TL syntax graph, and the robot initial state. TL-node tokens and spatial visual tokens are fused through bidirectional cross-attention, and the resulting representation conditions a flow-matching trajectory generator. Visual tokens are augmented only with normalized image-plane locations and camera-view identifiers, while a training-only predicate-to-region objective encourages grounding to referenced objects. Consistent with prior work in this domain, we evaluate the model using Success@$K$, the fraction of tasks for which at least one of K sampled trajectories satisfies the TL specification. On Panda task, our model achieves 67.45% Success@1024, compared with 59.11% for the oracle-state baseline. On AntMaze task, it achieves 96.35% Success@256, comparable to the oracle result of 96.88%. Resolution and intervention studies show that spatial detail depends on semantic grounding and predicate identity affects both attention and performance. These results demonstrate a direct mapping from visual observations and structured TL goals to action trajectories without requiring object geometry at inference. Code is available at https://github.com/AricLau07/vision-tl-action.
Chinese Translation
时序逻辑(TL)为长时间范围的机器人任务提供了一种组合语言,但现有的基于TL的轨迹生成器通过在任务图中编码精确的物体几何形状,可以绕过感知与符号绑定的问题。我们提出了 extit{视觉-时序逻辑-动作}(Vision-TL-Action),该方法从多视角图像、无坐标的TL语法图和机器人初始状态生成动作轨迹。TL节点令牌和空间视觉令牌通过双向交叉注意力进行融合,生成的表示条件化了流匹配轨迹生成器。视觉令牌仅通过归一化的图像平面位置和相机视图标识符进行增强,同时仅在训练中使用的谓词到区域目标鼓励与参考物体的绑定。与该领域的先前研究一致,我们使用Success@$K$来评估模型,即至少有一个K个采样轨迹满足TL规范的任务比例。在Panda任务中,我们的模型实现了67.45%的Success@1024,而oracle状态基线为59.11%。在AntMaze任务中,模型实现了96.35%的Success@256,与oracle结果的96.88%相当。分辨率和干预研究表明,空间细节依赖于语义绑定,而谓词身份影响注意力和性能。这些结果展示了从视觉观察和结构化TL目标到动作轨迹的直接映射,且在推理时不需要物体几何形状。代码可在https://github.com/AricLau07/vision-tl-action获取。
cs.RO / 16 / 2607.26789

CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation

CheckVLA:基于动作条件的世界模型的执行时间验证用于长时间跨度的移动操控
Liu, Yushan, Sun, Peibo, Chao, Xintao, Yang, Zhenyang, Xie, Yifan, Zhang, Lingfeng, Li, Shoujie, Tang, Chenyu, Chen, Fang, Zhang, Xiao-Ping, Ding, Wenbo
Abstract
Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action chunks, issuing multiple actions without receiving new high-level visual input. A committed chunk therefore implies how observations should evolve, but accidental deviations can violate this expectation while the remaining actions continue to propagate the error: commit-time policy confidence cannot react to a deviation that occurs after dispatch, and observation-only anomaly scores lack an action-conditioned reference for separating expected effects from unexplained changes. We propose CheckVLA, which verifies execution with a separately trained, frozen action-conditioned world model. A conformally calibrated risk threshold bounds the episode-level probability of an unnecessary first intervention and determines when to intervene, its exceedance controls how strongly the rewritten suffix retains the superseded chunk, latency-aware hard prefixing restricts replacement to actions that remain deployable, and an event-driven keyframe bank preserves evidence of prior progress across repairs. On RoboCasa365, under a common training recipe and a matched invocation budget, CheckVLA attains a 36.1% average success rate against 27.6% for periodic replanning (+8.5 points). At a matched 5% episode-level false-alarm target, action conditioning raises timely recall to 77.9%, against 48.6% for an observation-only control and 37.9% for an action-shuffled control. These simulation results support action-conditioned verification as a way to restore feedback during chunked execution while keeping the repair consistent with inference latency.
Chinese Translation
视觉-语言-动作(VLA)策略通常通过开放式动作块执行长时间跨度的移动操控,发出多个动作而不接收新的高层视觉输入。因此,一个已承诺的动作块暗示了观察结果应如何演变,但意外的偏差可能会违反这一期望,而剩余的动作继续传播错误:承诺时的策略置信度无法对调度后发生的偏差作出反应,而仅基于观察的异常评分缺乏一个动作条件的参考来区分预期效果与无法解释的变化。我们提出了CheckVLA,它通过一个单独训练的、冻结的动作条件世界模型来验证执行。一个符合校准的风险阈值限制了不必要的首次干预的情节级概率,并决定何时干预,其超出控制了重写后缀保留被取代块的强度,延迟感知的硬前缀限制了替换为仍可部署的动作,并且事件驱动的关键帧库保留了修复过程中的先前进展证据。在RoboCasa365上,在一个共同的训练方案和匹配的调用预算下,CheckVLA的平均成功率达到了36.1%,而周期性重新规划的成功率为27.6%(提高了8.5个百分点)。在匹配的5%情节级误报目标下,动作条件提高了及时召回率至77.9%,而仅基于观察的对照组为48.6%,动作打乱的对照组为37.9%。这些仿真结果支持基于动作条件的验证作为在分块执行期间恢复反馈的一种方式,同时保持修复与推理延迟的一致性。
cs.RO / 17 / 2607.26802

Risk-Aware Motion Planning with Learned Trajectory Primitives and Probabilistic Safety Assessment

基于学习的轨迹原语和概率安全评估的风险感知运动规划
Kaufeld, Marc, Zhuang, Dian, Betz, Johannes
Abstract
This paper presents a radial basis function network (RBFN)-informed motion planning framework for safe and efficient urban autonomous driving. The proposed approach combines RBFN-based candidate trajectory generation with an analytic collision probability assessment and optimization-based trajectory refinement. The network learns jerk-minimal trajectories, enabling the MPC to operate within a reduced and dynamically consistent search space. Candidate motion primitives are selected based on an accurate probabilistic risk measure. This design decreases solver complexity while preserving safety and constraint satisfaction. The framework is evaluated in numerous urban driving scenarios. Results demonstrate improved risk awareness and fewer vehicle-limit violations compared to benchmark methods. The proposed approach integrates learning-based trajectories into optimization-based motion planning, thereby ensuring safety and interpretability.
Chinese Translation
本文提出了一种基于径向基函数网络(RBFN)的运动规划框架,以实现安全高效的城市自动驾驶。所提出的方法结合了基于RBFN的候选轨迹生成、解析碰撞概率评估和基于优化的轨迹优化。该网络学习最小颠簸的轨迹,使得模型预测控制(MPC)能够在一个减少且动态一致的搜索空间内运行。候选运动原语的选择基于准确的概率风险度量。该设计在保持安全性和约束满足的同时,降低了求解器的复杂性。该框架在多个城市驾驶场景中进行了评估。结果表明,与基准方法相比,风险意识得到了改善,车辆限值违规情况减少。所提出的方法将基于学习的轨迹集成到基于优化的运动规划中,从而确保安全性和可解释性。
cs.RO / 18 / 2607.26807

Route by Kinematics, Act by Observation: Kinematics-Supervised Expert Routing in MoE-Augmented VLA

通过运动学引导,基于观察行动:运动学监督的专家路由在MoE增强的VLA中的应用
Yang, Tianhang, Zheng, Yanze, Wang, Junjie, Kou, Wei-Bin, Li, Ruotong, Yang, Yujiu
Abstract
While MoE augments VLA via expert specialization, router suffers from ineffective expert routing owing to the kinematic heterogeneity of actions across manipulation tasks and, even worse, the unavailability of the kinematic signals at inference time. In this work, we first observe that most semantically distinct manipulation tasks reduce to multiple kinematic archetypes. Motivated by this finding, we propose Kinematics-supervised explicit routing (KinRT), a new paradigm that shifts from implicit, observation-driven expert routing to explicit, kinematics-guided expert dispatching. Specifically, we perform kinematic clustering on action trajectories into multiple kinematically coherent groups, whose IDs serve as ground truth to supervise the training of the router; at inference time, the router dispatches experts only using visual-language observations, without any reliance on action kinematics. KinRT actually introduces an asymmetric bridging mechanism that distills the task kinematics from the action space in training into the observation space at inference. In addition, to assess KinRT's cross-platform generalization, we build an economical, Do-It-Yourself robot (DIYRobot) platform from scratch using 3D-print technology ($<$ 2,000USD). Extensive experiments demonstrate KinRT's superiority over both dense and MoE-featured VLAs by more than 23.26% on RoboTwin benchmark and 20.27% on our introduced DIYRobot platform. Our code and DIYRobot platform will be open-sourced.
Chinese Translation
尽管MoE通过专家专业化增强了VLA,但由于操作任务中动作的运动学异质性,路由器在专家路由方面存在低效的问题,更糟糕的是,在推理时无法获得运动学信号。在本研究中,我们首先观察到大多数语义上不同的操作任务可以归结为多个运动学原型。基于这一发现,我们提出了运动学监督的显式路由(Kinematics-supervised explicit routing, KinRT),这是一种新的范式,它将隐式的、基于观察的专家路由转变为显式的、基于运动学的专家调度。具体而言,我们对动作轨迹进行运动学聚类,将其分为多个运动学一致的组,这些组的ID作为监督路由器训练的真实标签;在推理时,路由器仅使用视觉-语言观察来调度专家,而不依赖于动作的运动学。KinRT实际上引入了一种不对称的桥接机制,将训练中的任务运动学从动作空间提炼到推理时的观察空间。此外,为了评估KinRT的跨平台泛化能力,我们从零开始构建了一个经济型的DIY机器人平台(DIYRobot),使用3D打印技术(成本低于2000美元)。大量实验表明,KinRT在RoboTwin基准测试中比密集型和MoE特征的VLA优越超过23.26%,在我们引入的DIYRobot平台上优越超过20.27%。我们的代码和DIYRobot平台将开源。
cs.RO / 19 / 2607.26809

Practice Makes Policies: Bootstrapping and Consolidating Robotic Capabilities from Zero Human Demonstrations

实践造就政策:从零人类示范中引导和巩固机器人能力
Li, Jialiang, Wang, Yuhan, Li, Haojun, Zhang, Gaojing, Ye, Yangtian, Liu, Qipeng, Liang, Haotian, Lian, Wenzhao
Abstract
General-purpose robotic manipulation requires robots to perform diverse tasks in open-world environments while improving their skills over time. Despite recent progress in robotic manipulation, existing systems still primarily acquire manipulation skills in a static manner, where capabilities are learned for specific tasks or settings rather than adaptively evolving through physical interaction. Resembling how repeated practice enables humans to develop muscle memory, advanced manipulation proficiency requires an autonomous capability evolution mechanism that allows robots to progressively transform interaction experiences into increasingly effective manipulation abilities. To this end, we propose HERO, a self-improving hierarchical embodied agent that enables autonomous capability evolution from zero human demonstrations. HERO organizes heuristic reasoning, exemplar reuse, and reflexive execution into a unified orchestration framework, allowing robots to autonomously bootstrap manipulation experience, rapidly accumulate reusable behaviors through experience transfer, and progressively consolidate recurring interactions into efficient closed-loop visuomotor policies. By tightly coupling autonomous data collection with task execution, HERO continuously expands and dynamically schedules manipulation capabilities according to different stages of experience accumulation and execution requirements. Extensive experiments demonstrate that HERO substantially reduces human intervention during robotic data collection while achieving robust manipulation across diverse tasks, providing a promising path toward self-improving robotic systems.
Chinese Translation
通用机器人操作需要机器人在开放世界环境中执行多样化任务,并随着时间的推移提升其技能。尽管在机器人操作方面取得了近期进展,现有系统仍主要以静态方式获取操作技能,即能力是针对特定任务或设置学习的,而不是通过物理互动自适应演变的。类似于重复练习使人类发展肌肉记忆,先进的操作熟练度需要一种自主能力演变机制,使机器人能够逐步将互动经验转化为越来越有效的操作能力。为此,我们提出了HERO(自我改进的分层具身代理),该代理能够在没有人类示范的情况下实现自主能力演变。HERO将启发式推理、范例重用和反射执行组织成一个统一的协调框架,使机器人能够自主引导操作经验,通过经验转移快速积累可重用行为,并逐步将重复互动巩固为高效的闭环视觉运动策略。通过将自主数据收集与任务执行紧密结合,HERO根据不同的经验积累阶段和执行要求不断扩展和动态调度操作能力。大量实验表明,HERO显著减少了机器人数据收集过程中的人类干预,同时在多样化任务中实现了稳健的操作,为自我改进的机器人系统提供了有希望的路径。
cs.RO / 20 / 2607.26817

From Uncertainty to Determinism: Coarse-to-Fine Visual Floorplan Localization without Ray Matching

从不确定性到确定性:无光线匹配的粗到细视觉平面图定位
Meng, Shiyong, Chen, Bolei, Zhong, Ping, Wan, Yang, Wang, Rongzhi, Xia, Jiazhi, Wang, Jianxin
Abstract
Visual Floorplan Localization (FLoc) has emerged as a promising solution for indoor localization by matching egocentric images against minimalist structural maps. However, due to cross-modal information asymmetry and repetitive indoor layouts, visual FLoc is fundamentally challenged by multimodal pose distributions, where visually identical observations map to distinct, spatially separated locations. Existing ray-matching-based methods tackle this by explicitly predicting sparse geometric or semantic rays, which inherently incur information loss and demand resource-intensive preprocessing alongside exhaustive matching during inference. In this paper, we bypass the intermediate ray-matching paradigm and propose a coarse-to-fine visual FLoc framework that progresses from uncertainty to determinism. In the coarse stage, we design an image-conditioned pose diffusion model to parameterize the continuous multimodal pose distribution, effectively routing stochastically initialized pose particles toward distinct candidate modes. In the refinement stage, we propose a localized refiner that predicts bounded sub-meter pose residuals from candidate-centered floorplan crops, where structural ambiguities are largely eliminated. Our method effectively balances global multi-hypothesis tracking and local sub-meter refinement without requiring any offline map preprocessing or test-time lookup tables. Comprehensive results on the S3D (full) and ZInD benchmarks demonstrate that our approach achieves state-of-the-art accuracy and robustness.
Chinese Translation
视觉平面图定位(FLoc)作为一种有前景的室内定位解决方案,通过将自我中心图像与简约结构图进行匹配而出现。然而,由于跨模态信息不对称和重复的室内布局,视觉FLoc在多模态姿态分布方面面临根本挑战,其中视觉上相同的观察结果映射到不同的、空间上分离的位置。现有的基于光线匹配的方法通过显式预测稀疏的几何或语义光线来应对这一问题,但这固有地导致信息损失,并在推理过程中需要资源密集型的预处理和全面匹配。在本文中,我们绕过中间的光线匹配范式,提出了一种从不确定性到确定性的粗到细视觉FLoc框架。在粗略阶段,我们设计了一种图像条件的姿态扩散模型,以参数化连续的多模态姿态分布,有效地将随机初始化的姿态粒子引导到不同的候选模式。在细化阶段,我们提出了一种局部细化器,从以候选为中心的平面图裁剪中预测有界的亚米级姿态残差,在此过程中大大消除了结构模糊性。我们的方法有效平衡了全局多假设跟踪和局部亚米级细化,而无需任何离线地图预处理或测试时查找表。在S3D(完整)和ZInD基准上的全面结果表明,我们的方法达到了最先进的准确性和鲁棒性。
cs.RO / 21 / 2607.26855

NeoRacer: An Open, Standardized 1:12 Scale Autonomous Race Car for Benchmarking and Education

NeoRacer:一个开放、标准化的1:12比例自主赛车平台,用于基准测试和教育
Bandyopadhyay, Koneshka, Mehta, Ansh, Mabsout, Bassel El, Mancuso, Renato
Abstract
Many scientific fields rely on standard benchmarks and shared platforms to improve review and reproducibility, but autonomous systems research still lacks widely accepted open hardware. Where standardization has emerged, progress has accelerated. This is especially evident in autonomous racing, where teams often build custom systems or buy niche, expensive vehicles, making control and robotics research and education hard to compare and reproduce. High costs also limit access outside well-funded labs, while affordable educational robots are often underpowered. To address this gap, we present NeoRacer, an open-source 1:12 scale autonomous racing platform. It is built around an NVIDIA Jetson Orin Nano (67 TOPS), a 270{\deg} LiDAR, a 120 fps global-shutter camera, and a 9-axis IMU. NeoRacer ships pre-assembled for USD 2,699, offering over 3x the compute of comparable platforms at less than half the cost of the nearest pre-assembled alternative. Co-developed by the Neobotics Foundation and Seeed Studio, and manufactured by Seeed Studio, NeoRacer combines open hardware and software design with scalable, repeatable production. The modular, extensible platform provides a standardized benchmarking environment for autonomous racing algorithms across institutions. We describe the hardware/software architecture, design decisions from two pilot deployments (MIT IAP, 15 students; BU CPS Lab, 10 students), and key cost-performance tradeoffs. Hardware is licensed under CERN-OHL-S v2 and software under GPLv3, with all design files, firmware, and ROS2 packages publicly accessible.
Chinese Translation
许多科学领域依赖标准基准和共享平台来提高评审和可重复性,但自主系统研究仍然缺乏广泛接受的开放硬件。在标准化出现的地方,进展得以加速。这在自主赛车中尤为明显,团队通常构建定制系统或购买小众、昂贵的车辆,使得控制和机器人研究及教育难以比较和重复。高昂的成本也限制了资金不足的实验室的访问,而价格适中的教育机器人往往性能不足。为了解决这一差距,我们推出了NeoRacer,一个开放源代码的1:12比例自主赛车平台。它基于NVIDIA Jetson Orin Nano(67 TOPS)、270° LiDAR、120 fps全局快门相机和9轴IMU构建。NeoRacer以2,699美元的价格预装配出售,提供超过3倍于可比平台的计算能力,而成本不到最近预装配替代品的一半。NeoRacer由Neobotics基金会和Seeed Studio共同开发,并由Seeed Studio制造,结合了开放硬件和软件设计与可扩展、可重复的生产。该模块化、可扩展的平台为各机构的自主赛车算法提供了标准化的基准测试环境。我们描述了硬件/软件架构、两个试点部署(MIT IAP,15名学生;BU CPS实验室,10名学生)的设计决策,以及关键的成本-性能权衡。硬件遵循CERN-OHL-S v2许可证,软件遵循GPLv3许可证,所有设计文件、固件和ROS2包均可公开访问。
cs.RO / 22 / 2607.26914

BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories

BioVLN:生物医学实验室视觉语言导航的仿真平台
Liu, Zhe, Lu, Quan, Du, Zhaohui, Wang, Zhe, Jin, Huanbo, Gu, Jiaming, Wang, Qi, Xiao, Ting, Pan, Minting, Zhou, Dongzhan
Abstract
Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed for household environments and treat a target as an object center or an arbitrary nearby position. This representation is inadequate for laboratory instruments, which must be approached from their operating side while maintaining safe clearance from surrounding equipment. We introduce BioVLN, a simulation platform for developing and evaluating visual-language navigation agents in biomedical laboratories. BioVLN represents each instrument with three regions: its physical body, a surrounding clearance region, and an operation area in front of the usable side. This model is applied consistently to scene generation, target placement, navigation evaluation, and safety analysis, so success depends on reaching a position from which the instrument can be accessed. BioVLN supports procedural scene generation and manually designed environments, producing 47 scenes and 1667 episodes. Standardized navigation and reinforcement-learning interfaces enable trajectory collection and policy training. Experiments show that geometric exploration reaches 74.4--87.5% success, while sampling multiple valid positions in the operation area improves success to 83.3--92.5% and reduces unsafe proximity.
Chinese Translation
生物医学实验室机器人在执行实验程序之前必须导航到仪器位置。现有的具身导航平台主要针对家庭环境,将目标视为物体中心或任意附近位置。这种表示方法对于实验室仪器而言是不够的,因为必须从其操作侧接近,同时保持与周围设备的安全间距。我们提出了BioVLN,一个用于开发和评估生物医学实验室视觉语言导航代理的仿真平台。BioVLN将每个仪器表示为三个区域:其物理主体、周围的安全间隔区域以及可用侧前方的操作区域。该模型在场景生成、目标放置、导航评估和安全分析中一致应用,因此成功依赖于到达可以接触仪器的位置。BioVLN支持程序化场景生成和手动设计的环境,生成了47个场景和1667个实验。标准化的导航和强化学习接口使得轨迹收集和策略训练成为可能。实验表明,几何探索的成功率达到74.4%至87.5%,而在操作区域内采样多个有效位置则将成功率提高到83.3%至92.5%,并减少了不安全的接近。
cs.RO / 23 / 2607.26980

Dense Soft Weighting for Radar Ego-Velocity Estimation

用于雷达自我速度估计的密集软加权
Babgei, Atar, Zhao, Chenyu, Breza, Michael, McCann, Julie A.
Abstract
Sensing ego-velocity estimation is fundamental to state estimation in visually degraded environments, where camera- and LiDAR-based pipelines can become unreliable. Millimetre-wave radar is well suited to these conditions because it provides direct Doppler velocity sensing and remains robust to poor illumination, textureless scenes, and airborne particulates. However, conventional radar ego-velocity pipelines typically apply constant false alarm rate (CFAR) thresholding to convert dense radar spectra into sparse point clouds, prematurely discarding sub-threshold returns that may still retain useful Doppler motion cues. We present Dense Soft Weighting, an analytic radar front-end that maps every range-Doppler cell to a continuous confidence metric rather than enforcing a binary detection threshold. Ego-velocity is then estimated using a deterministic robust weighted least-squares formulation, while the same weighted measurements provide a closed-form, measurement-derived velocity covariance for integration with a shared inertial back-end. The method requires no platform-specific training data or learning-based uncertainty model, supporting transfer across single-chip radar configurations. Across two public datasets and one self-collected dataset, Dense Soft Weighting reduces mean absolute pose error by 31-45% relative to the strongest CFAR point-cloud baseline under an identical inertial back-end, while running in real time on embedded hardware.
Chinese Translation
自我速度估计在视觉退化环境中的状态估计中至关重要,此时基于相机和激光雷达的处理流程可能变得不可靠。毫米波雷达非常适合这些条件,因为它提供直接的多普勒速度感知,并且在光照不足、纹理缺失的场景和空气颗粒物存在的情况下仍然保持稳健。然而,传统的雷达自我速度处理流程通常应用恒定虚警率(CFAR)阈值,将密集的雷达谱转换为稀疏点云,过早地丢弃可能仍然保留有用多普勒运动线索的阈值以下的回波。我们提出了密集软加权(Dense Soft Weighting),这是一种分析性雷达前端,它将每个距离-多普勒单元映射到一个连续的置信度度量,而不是强制执行二元检测阈值。然后使用确定性稳健加权最小二乘法来估计自我速度,同时相同的加权测量提供了一个封闭形式的、基于测量的速度协方差,以便与共享的惯性后端进行集成。该方法不需要特定平台的训练数据或基于学习的不确定性模型,支持在单芯片雷达配置之间的迁移。在两个公共数据集和一个自收集的数据集上,密集软加权相对于在相同惯性后端下最强的CFAR点云基线减少了31-45%的平均绝对姿态误差,同时在嵌入式硬件上实时运行。
cs.RO / 24 / 2607.26985

SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception

SymmGrid:通过并行对称性和自我中心-外部中心视觉感知实现机器人学习的超规模化
Everett, Gabe, Gunter, Brice, Stelt, Ryan Vander, Ruiz-Martinez, Cleiver, Hull, Blake, Rojas, Juan
Abstract
Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock training times. We present SymmGrid, a trajectory level augmentation framework inspired by parallelized symmetries that super-scales group transformations to significantly accelerate on-robot learning in both egocentric and exocentric visual setups. We model a Markov Decision Process (MDP) under a symmetry tree, in which state-action pairs have admissible parallelized invariant transformations that yield a geometric grid structure. The state is modelled with ego- or exocentric images and proprioception information. The latter require special treatment, in the form of homographies, to warp visual scenes in line with their corresponding spatial transformations. These parallelized transformations produce a large set of unique symmetric equivalences that populate the replay buffer with diverse and consistent experiences that speed up learning and improve performance. We present extensive training and evaluations performed directly on real robot manipulation contact tasks including peg-insertions, cable routing, and object relocations. Relative to SOTA, SymmGrid achieved wall-clock training convergence speed-ups of 1.37-2.17x, evaluation success rate improvements of 1.09x-1.27x, fastest training convergence times of 16.6, 10.9, and 79.3 minutes respectively. For trajectory wide assessments, we used normalized area under the curve (nAUC) ratios. SymmGrid achieved improvements of up to 2.59x. These results confirm that simple branch symmetries can have an outsized result due to super-scaling and bring us closer to sub-10 minute on-robot learning training in manipulation tasks suitable for arms and humanoids. The project page is available at symmgrid-robot.github.io
Chinese Translation
在物理机器人中直接进行深度强化策略学习(机器人学习)仍然受到训练时间较慢的瓶颈限制。我们提出了SymmGrid,这是一种受并行对称性启发的轨迹级增强框架,能够超规模化群体变换,从而显著加速自我中心和外部中心视觉设置下的机器人学习。我们在对称树下建模了一个马尔可夫决策过程(MDP),其中状态-动作对具有可接受的并行不变变换,形成几何网格结构。状态通过自我中心或外部中心图像及本体感知信息进行建模。后者需要特殊处理,以同质变换的形式扭曲视觉场景,使其与相应的空间变换一致。这些并行变换产生了一大批独特的对称等价,丰富了重放缓冲区中的多样化和一致性经验,从而加速学习并提高性能。我们进行了大量的训练和评估,直接在真实机器人操作接触任务上进行,包括插销插入、线缆布置和物体重新定位。与现有技术(SOTA)相比,SymmGrid实现了1.37-2.17倍的墙钟训练收敛速度提升,评估成功率提高了1.09-1.27倍,最快的训练收敛时间分别为16.6、10.9和79.3分钟。对于轨迹广泛的评估,我们使用了归一化曲线下面积(nAUC)比率。SymmGrid的提升达到了2.59倍。这些结果确认简单的分支对称性由于超规模化可以产生显著的效果,使我们更接近于在适合机械臂和类人机器人操作任务中实现10分钟以内的机器人学习训练。项目页面可访问symmgrid-robot.github.io
cs.RO / 25 / 2607.26991

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

RL$^2$-VLA:适应性强化学习潜在组合引导与测试时缩放的视觉-语言-动作模型
Tan, Derek Ming Siang, Shailesh, Shailesh, Iyer, Srikrishna, Teo, William Wei Jie, Ju, Yuanliang, Gu, Qiao, Sartoretti, Guillaume
Abstract
Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance often degrades on challenging and out-of-domain tasks. Recent test-time steering and scaling methods improve performance without extensive data collection and retraining, but action samples often remain concentrated around similar behaviors and therefore inherit correlated failure modes. Moreover, existing methods apply the same intervention strategy at every timestep, regardless of whether the base policy is already likely to succeed. To address these limitations, we introduce $RL^2$, an adaptive inference-time steering framework that leverages Reinforcement Learning on VLA Latents. First, we train a lightweight offline RL policy conditioned on expressive latents extracted from the VLA action expert and compose its flow velocity with that of the frozen VLA during inference. This compositional steering strategy combines the behavioral priors of large-scale imitation learning with the action diversity induced by offline RL beyond dominant demonstration modes. We further discover that inference-time steering follows fundamentally different scaling laws under success and failure states, revealing that action diversity is most beneficial when the base VLA is likely to fail, but can unnecessarily perturb already-accurate actions when success is likely. Building on this insight, $RL^2$ activates compositional steering only when failure is predicted. Across the SIMPLER and PolaRiS benchmarks, $RL^2$ improves success rates by up to +17.3% in out-of-domain settings, while ablations and scaling studies demonstrate the importance of latent representations and RL training. Finally, real-world experiments demonstrate that these gains transfer beyond simulation, establishing $RL^2$ as a practical and modular steering framework for VLA deployment.
Chinese Translation
尽管视觉-语言-动作(VLA)模型展现了令人印象深刻的视觉运动能力,但它们在具有挑战性和超领域任务上的表现往往会下降。近期的测试时引导和缩放方法在不需要大量数据收集和重新训练的情况下提高了性能,但动作样本通常仍然集中在相似的行为上,因此继承了相关的失败模式。此外,现有方法在每个时间步都应用相同的干预策略,而不考虑基础策略是否已经可能成功。为了解决这些局限性,我们引入了$RL^2$,一种适应性推理时引导框架,利用强化学习对VLA潜在变量进行引导。首先,我们训练一个轻量级的离线强化学习策略,该策略以从VLA动作专家提取的表达性潜在变量为条件,并在推理时将其流速与冻结的VLA流速进行组合。这种组合引导策略结合了大规模模仿学习的行为先验与离线强化学习所引发的动作多样性,超越了主导演示模式。我们进一步发现,推理时引导在成功和失败状态下遵循根本不同的缩放规律,揭示了当基础VLA可能失败时,动作多样性最为有利,但在成功可能性较高时,可能会不必要地扰动已经准确的动作。基于这一洞察,$RL^2$仅在预测到失败时激活组合引导。在SIMPLER和PolaRiS基准测试中,$RL^2$在超领域设置中将成功率提高了多达17.3%,而消融实验和缩放研究则证明了潜在表示和强化学习训练的重要性。最后,现实世界的实验表明,这些增益超越了模拟,确立了$RL^2$作为VLA部署的实用和模块化引导框架。
cs.RO / 26 / 2607.27085

Controlled Experiments on Lane Changing by Transitional Autonomous Vehicle: Dataset and Behavioral Insights

过渡性自动驾驶车辆的变道控制实验:数据集与行为洞察
Sharma, Abhinav, Hasan, Md Abdullah Al, Chen, Danjue, List, George F.
Abstract
This paper presents the North Carolina Transitional Autonomous Vehicle Lane-Changing (NC-tALC) dataset and uses it to characterize mandatory lane-changing behavior of transitional automated vehicles (tAVs). It quantifies the evolution of lead--lag gaps throughout the lane-change process and examines how potential collision risk develops during the maneuver. A controlled field experiment comprising 78 mandatory lane-change trials was conducted on a public roadway in Apex, North Carolina. Four instrumented vehicles created repeatable traffic conditions while varying the lane changer's initial position within the candidate target gap. High-resolution RTK-GNSS/INS trajectories were processed to identify key timestamps, calculate lead, lag, and lane-change gaps, and estimate interactions using time-gap- and speed-based surrogate safety measures. Despite substantial differences in initial conditions, lead and lag gaps consistently converged toward a relatively narrow range near lane crossing. Potential collision risk increased as the maneuver progressed, peaked near physical lane entry, and was dominated by interactions with the target-lane leader. Lane-change completion did not necessarily coincide with the disappearance of collision risk. This study provides one of the first controlled empirical characterizations of the complete mandatory lane-change process of tAVs using repeatable public-road experiments. The NC-tALC dataset supports analysis of behavioral and safety evolution throughout the maneuver rather than only at the gap-acceptance instant. The dataset and findings provide empirical benchmarks for evaluating automated lane-changing behavior, calibrating behavioral models, and validating simulation and safety assessment methods for mandatory lane-change scenarios.
Chinese Translation
本文介绍了北卡罗来纳州过渡性自动驾驶车辆变道(NC-tALC)数据集,并利用该数据集对过渡性自动化车辆(tAVs)的强制变道行为进行了特征描述。研究量化了变道过程中前后间隙的演变,并考察了在该操作过程中潜在碰撞风险的发展。我们在北卡罗来纳州阿佩克斯的一条公共道路上进行了包含78次强制变道试验的受控实地实验。四辆配备仪器的车辆创造了可重复的交通条件,同时改变了变道者在候选目标间隙中的初始位置。高分辨率RTK-GNSS/INS轨迹被处理以识别关键时间戳,计算前后间隙和变道间隙,并使用基于时间间隙和速度的替代安全措施来估计交互作用。尽管初始条件存在显著差异,前后间隙在接近变道时始终趋向于相对狭窄的范围。随着操作的进行,潜在碰撞风险增加,在物理车道入口附近达到峰值,并主要受目标车道领头车辆的交互影响。变道的完成并不一定与碰撞风险的消失同时发生。本研究提供了对tAVs完整强制变道过程的首次受控实证特征描述,基于可重复的公共道路实验。NC-tALC数据集支持在整个操作过程中分析行为和安全的演变,而不仅仅是在间隙接受的瞬间。该数据集及其研究结果为评估自动变道行为、校准行为模型以及验证强制变道场景的仿真和安全评估方法提供了实证基准。
cs.RO / 27 / 2607.27138

DLAM: Distributional Latent Actions with Temporal Constraints

DLAM:具有时间约束的分布式潜在动作
Tang, Zuojin, Luo, Feifan, Liu, Haoyun, Yuan, Botai, Qi, Dekang, Chen, Ronghan, Yang, Yandan, Lin, Tong, Chang, Xinyuan, Xu, Mu, Liu, Bin, Ma, De, Ma, Zhiheng
Abstract
Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations without the structure required for joint generation with robot actions. Existing structured methods add temporal constraints but retain deterministic transition points, so residual errors in locally inferred transitions may propagate and compound under recursive composition. We introduce DLAM, a distributional latent-action model that represents each transition as a diagonal Gaussian. Reconstruction conditioned on the reference frame grounds the mean in observed visual change, while normalized composition and reversal over equal-gap triplets constrain both the mean and dimension-wise variance. Variance composition uses a lightweight shared-correlation coefficient to account for dependence between adjacent transitions that share an intermediate frame, whereas reversal negates the mean and preserves the variance. For downstream policy learning, we freeze the encoder and train a flow-matching policy to jointly generate mean transition sequences and robot actions. On held-out transitions, DLAM learns more temporally consistent latent dynamics than existing latent-action baselines and achieves stronger direct and cumulative reconstruction on held-out videos. Under the same controlled $\pi_0$ transfer protocol, it also improves policy performance on MetaWorld MT50, LIBERO, and real-world manipulation tasks. Controlled ablations show that normalized mean constraints account for most of the reconstruction gain, while learned variance and correlation-aware composition provide complementary improvements in downstream control.
Chinese Translation
视觉-语言-动作(VLA)模型受到稀缺的动作标记机器人数据的限制,而无动作视频则提供了丰富的物理变化观察。潜在动作模型可以提取这些先验信息,但重构训练的编码可能在没有与机器人动作联合生成所需结构的情况下预测未来观察。现有的结构化方法增加了时间约束,但保留了确定性的过渡点,因此在局部推断的过渡中残余误差可能在递归组合下传播和累积。我们提出了DLAM,一种分布式潜在动作模型,将每个过渡表示为对角高斯分布。基于参考帧的重构将均值固定在观察到的视觉变化上,而归一化的组合和在等间隔三元组上的反转则限制了均值和维度方差。方差组合使用轻量级共享相关系数来考虑共享中间帧的相邻过渡之间的依赖关系,而反转则否定均值并保留方差。对于下游策略学习,我们冻结编码器并训练流匹配策略,以联合生成均值过渡序列和机器人动作。在保留的过渡上,DLAM学习到的潜在动态比现有的潜在动作基线更具时间一致性,并在保留的视频上实现了更强的直接和累积重构。在相同的控制$ ext{π}_0$转移协议下,它还提高了在MetaWorld MT50、LIBERO和现实世界操控任务上的策略性能。控制消融实验表明,归一化均值约束占据了大部分重构增益,而学习的方差和相关性感知组合在下游控制中提供了互补的改进。
计算机视觉 (Computer Vision)
85
cs.CV / 1 / 2607.26097

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition

基于知识引导的原子动作解耦用于动作识别
Wu, Tianci, Cao, Siqi, Zhu, Guangming, Lu, Jiang, Wang, Siyuan, Zhang, Longfei, Huang, Jincai, Sheng, Jun, Zhang, Liang
Abstract
Action recognition in complex scenes often involves multiple concurrent fine-grained actions, making it challenging to model internal action structures. Most existing methods rely on holistic representations, which are insufficient for capturing subtle interactions and fine-grained semantics. While recent prompt-based approaches introduce disentanglement, they lack explicit semantic guidance, and methods based solely on visual or structured cues remain coarse-grained. In this paper, we propose Knowledge-guided Disentanglement with Atomic Actions (KDA), which leverages fine-grained semantic knowledge to enhance action representations and enable more precise disentanglement. Specifically, we use Large Language Models (LLMs) to decompose action labels into atomic actions, providing explicit spatial-temporal semantics. A Knowledge Injection Module (KIM) first integrates atomic action knowledge into video features. Based on this enhanced representation, a Knowledge Disentanglement Module (KDM) further disentangles atomic action knowledge to produce more precise semantic guidance for action disentanglement. A Knowledge Disentanglement Loss (KD Loss) is introduced to encourage clearer disentanglement of knowledge components within KDM. Extensive experiments demonstrate that KDA improves feature discriminability and achieves state-of-the-art performance on multi-label action recognition benchmarks. Moreover, KIM and KDM can be readily integrated into other methods, demonstrating strong generality.
Chinese Translation
在复杂场景中的动作识别通常涉及多个并发的细粒度动作,这使得建模内部动作结构变得具有挑战性。现有的大多数方法依赖于整体表示,这不足以捕捉微妙的交互和细粒度的语义。尽管最近的基于提示的方法引入了解耦,但它们缺乏明确的语义指导,而仅基于视觉或结构线索的方法仍然是粗粒度的。本文提出了一种基于知识引导的原子动作解耦方法(Knowledge-guided Disentanglement with Atomic Actions, KDA),该方法利用细粒度的语义知识来增强动作表示,并实现更精确的解耦。具体而言,我们使用大型语言模型(Large Language Models, LLMs)将动作标签分解为原子动作,从而提供明确的时空语义。知识注入模块(Knowledge Injection Module, KIM)首先将原子动作知识整合到视频特征中。在此增强表示的基础上,知识解耦模块(Knowledge Disentanglement Module, KDM)进一步解耦原子动作知识,以产生更精确的语义指导用于动作解耦。引入知识解耦损失(Knowledge Disentanglement Loss, KD Loss)以鼓励KDM内知识组件的更清晰解耦。大量实验表明,KDA提高了特征的可区分性,并在多标签动作识别基准上达到了最先进的性能。此外,KIM和KDM可以轻松集成到其他方法中,展示了强大的通用性。
cs.CV / 2 / 2607.26104

Weight and Height Estimation from a Single Human Image Captured in the Wild

从野外捕获的单个人体图像中估计体重和身高
Yaseen, Hira, Mahmood, Arif, Sultani, Waqas
Abstract
A person's physical characteristics such as weight and height are important indicators of his physical and mental health, daily life routines and finances. Body Mass Index (BMI) is a well known measure that encodes the characteristics of both the weight and the height. BMI has been used as a self-monitoring tool, and it has long-term implications on one's life. For example, it may help predicting the risk of various diseases and estimating longevity. Automatic BMI estimation using a single person image in the wild is a challenging task due to wide variations in human pose, camera geometry, personal appearance and distracting backgrounds. In this paper, we explore the performance of deep neural networks using single and multi-task learning by employing different modalities including RGB, depth-maps, pose-affinity maps, and edge-maps to predict BMI, weight, and height from daily life images available on social networking websites. Currently, no full body image dataset for BMI estimation is publicly available, therefore we propose a new dataset consisting of 6105 images with ground truth labels of height, weight and BMI. Our proposed dataset is collected in the wild containing images from various ethnicity and distributed over varying age groups and gender. It consists of frontal, back, full and half body, side poses, mirror selfies with varying backgrounds and scale variations and may contain artifacts hiding partial or full face. Extensive experimentation is performed using full body, half body and face images only using different CNN backbones including VGG, Densenet and ResNet. Our experimental results demonstrate that full body images have produced better results than the other half body and facial images in the wild.
Chinese Translation
一个人的体重和身高等身体特征是其身体和心理健康、日常生活习惯及财务状况的重要指标。身体质量指数(BMI)是一个广为人知的测量标准,它编码了体重和身高的特征。BMI已被用作自我监测工具,并对个人的生活产生长期影响。例如,它可以帮助预测各种疾病的风险和估计寿命。由于人类姿势、相机几何、个人外观和干扰背景的广泛变化,使用单个图像自动估计BMI是一项具有挑战性的任务。在本文中,我们通过采用不同的模态,包括RGB、深度图、姿态亲和图和边缘图,探索了使用单任务和多任务学习的深度神经网络的性能,以从社交网络网站上可用的日常生活图像中预测BMI、体重和身高。目前,尚无公开可用的全身图像数据集用于BMI估计,因此我们提出了一个新的数据集,包含6105张带有身高、体重和BMI真实标签的图像。我们提出的数据集是在野外收集的,包含来自不同种族的图像,并分布在不同年龄组和性别之间。它包括正面、背面、全身和半身、侧面姿势、镜子自拍,背景和尺度变化各异,可能包含遮挡部分或全脸的伪影。我们使用不同的CNN骨干网络,包括VGG、Densenet和ResNet,进行了大量实验,仅使用全身、半身和面部图像。我们的实验结果表明,全身图像的结果优于其他半身和面部图像。
cs.CV / 3 / 2607.26107

TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions

TraceCLIP:从 Patch 到 CLS 贡献中恢复局部语义
Liu, Xinran, Shi, Shouqian, Chen, Yutong, Wang, Ge, Yao, Xin-Wei, Zhong, Sheng
Abstract
Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating language concepts with spatially grounded visual regions. CLIP provides a strong foundation for these tasks by learning a shared image-text embedding space from large-scale contrastive pre-training. However, its image-level objective aligns text with a CLS-derived global representation, leaving local vision-language correspondence only indirectly constrained. Existing methods either introduce additional supervision, external models, or task-specific adaptation, while training-free approaches mainly recover dense responses from existing patch features without examining where local semantics become most accessible within CLIP. We introduce TraceCLIP, a training-free framework that recovers latent patch-level semantic evidence by isolating the patch-specific terms written into the CLS attention output. TraceCLIP further converts contribution-derived semantic responses into a semantic-geodesic topology gate that calibrates final-layer patch affinity for dense feature reconstruction. Diagnostic experiments show that these contribution features exhibit strong local semantic discrimination and text-conditioned spatial alignment. On eight zero-shot semantic segmentation benchmarks, TraceCLIP achieves gains of 1.3 to 4.5 points in average mIoU over the strongest prior training-free methods across both backbones and background settings, without additional training, external vision foundation models, or region-level supervision. More broadly, these findings suggest that spatially localized semantics may remain accessible within the internal construction of globally aligned representations.
Chinese Translation
密集的视觉语言理解,包括物体定位、区域识别和开放词汇语义分割,需要将语言概念与空间上固定的视觉区域关联起来。CLIP 通过从大规模对比预训练中学习共享的图像-文本嵌入空间,为这些任务提供了坚实的基础。然而,其图像级目标将文本与基于 CLS 的全局表示对齐,仅间接约束了局部视觉-语言对应关系。现有方法要么引入额外的监督、外部模型或任务特定的适应,要么训练无关的方法主要从现有的 patch 特征中恢复密集响应,而未考察局部语义在 CLIP 中最易获取的位置。我们提出了 TraceCLIP,这是一种无训练框架,通过隔离写入 CLS 注意力输出的 patch 特定术语,恢复潜在的 patch 级语义证据。TraceCLIP 进一步将基于贡献的语义响应转换为语义-测地线拓扑门,以校准最终层 patch 亲和力以实现密集特征重建。诊断实验表明,这些贡献特征表现出强烈的局部语义区分能力和文本条件下的空间对齐。在八个零样本语义分割基准上,TraceCLIP 在最强的先前无训练方法中,在不同骨干网络和背景设置下,平均 mIoU 提高了 1.3 到 4.5 分,而无需额外训练、外部视觉基础模型或区域级监督。更广泛地说,这些发现表明,空间局部语义可能在全局对齐表示的内部构造中仍然是可获取的。
cs.CV / 4 / 2607.26165

DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving

DVPSFormer:用于自主驾驶的高效在线深度感知视频全景分割
Yang, Yung-Hsu, Piccinelli, Luigi, Li, Siyuan, Segu, Mattia, Ke, Lei, Danelljan, Martin, Fu, Yuqian, Bauer, Zuria, Yu, Fisher, Blum, Hermann, Pollefeys, Marc
Abstract
Safe autonomous navigation requires a holistic understanding of dynamic environments, necessitating the simultaneous estimation of metric depth, semantic segmentation, and instance trajectories. While depth-aware video panoptic segmentation (DVPS) unifies these tasks, existing approaches often rely on computationally expensive, multi-stage pipelines or offline tracking, rendering them unsuitable for real-time decision-making. To address this, we propose DVPSFormer, a unified online architecture designed for efficient 4D scene understanding. Central to our approach is explicit scene discretization (ESD), a novel mechanism that leverages segmentation queries to represent foreground and background regions, enabling a discrete-to-continuous (D2C) depth head to decode metric depth in a single pass. This tightly couples semantic and geometric learning while significantly reducing latency. Furthermore, we propose an online majority voting (OMV) mechanism that exploits temporal consistency to refine classification during instance tracking. DVPSFormer establishes a new state-of-the-art on the Cityscapes-DVPS and SemKITTI-DVPS benchmarks, offering a streamlined solution for online robotic perception. Code and models are available at https://royyang0714.github.io/DVPSFormer.
Chinese Translation
安全的自主导航需要对动态环境的整体理解,这要求同时估计度量深度、语义分割和实例轨迹。虽然深度感知视频全景分割(DVPS)将这些任务统一起来,但现有方法通常依赖于计算开销大的多阶段流程或离线跟踪,使其不适合实时决策。为了解决这个问题,我们提出了DVPSFormer,这是一种旨在高效进行4D场景理解的统一在线架构。我们方法的核心是显式场景离散化(ESD),这是一种新颖的机制,利用分割查询来表示前景和背景区域,使得离散到连续(D2C)深度头能够在一次传递中解码度量深度。这紧密结合了语义和几何学习,同时显著降低了延迟。此外,我们提出了一种在线多数投票(OMV)机制,利用时间一致性在实例跟踪期间细化分类。DVPSFormer在Cityscapes-DVPS和SemKITTI-DVPS基准测试中建立了新的最先进水平,为在线机器人感知提供了一种简化的解决方案。代码和模型可在 https://royyang0714.github.io/DVPSFormer 获取。
cs.CV / 5 / 2607.26170

A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment

一图胜千言 - 通过混合深度学习利用图像中的皮肤暴露数据以增强安全评估
Qian, Hua, Kotha, Manisha, Tran, Tuan, Shin, Jennifer, Zheng, Haining
Abstract
This study developed a hybrid computer vision method to quantify exposed skin from images for dermal exposure assessment. Using 170 indoor-painting images, Mask R-CNN first identified human subjects and removed background interference; a color-based algorithm then segmented exposed skin. The resulting exposed-skin-to-body pixel ratios showed approximately 80% agreement with human estimates. The approach demonstrates a scalable way to extract semi-quantitative exposure information from images, with future extensions to body-part recognition, PPE detection, and video-based exposure analysis.
Chinese Translation
本研究开发了一种混合计算机视觉方法,通过图像量化暴露皮肤以进行皮肤暴露评估。使用170张室内油漆图像,Mask R-CNN首先识别出人类对象并去除背景干扰;随后,基于颜色的算法对暴露皮肤进行分割。结果显示,暴露皮肤与身体像素比率与人工估计的结果约有80%的一致性。该方法展示了一种可扩展的方式,从图像中提取半定量的暴露信息,未来可扩展至身体部位识别、个人防护装备(PPE)检测和基于视频的暴露分析。
cs.CV / 6 / 2607.26196

Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography

Rad-JEPA 3D:用于三维计算机断层扫描的放射学联合嵌入预测模型
Trinh, Quoc-Huy, Nguyen, Minh-Van, Bagci, Ulas
Abstract
Self-supervised pretraining is central to 3D medical image analysis, where unlabeled CT volumes are abundant but expert annotations are scarce. Yet existing volumetric encoders often fail to preserve the coarse spatial and geometric structure that downstream reasoning depends on, limiting their performance on organ disentanglement, abnormality detection, and spatial understanding when paired with language models. We introduce Rad-JEPA 3D, a joint-embedding predictive framework that learns volumetric CT representations by predicting the latent features of a complete scan from a masked view. At its core is a hybrid H-Mamba encoder that fuses a Mamba state-space branch, which models inter-slice continuity through sequential scanning, with a grouped-query attention branch, which captures cross-plane spatial context, combined through a lightweight per-token router. To improve the quality of intermediate representations, we further propose Hidden States Orthogonal Regularization (HSOR), which aligns student-teacher hidden states and reduces feature redundancy throughout the encoder. This layer-wise regularization produces more consistent and discriminative volumetric representations, leading to improved performance on organ recognition and spatial reasoning tasks. Pretrained on approximately 120,000 CT scans, Rad-JEPA 3D attains state-of-the-art results despite its compact size: with only 4.0B total parameters, it achieves competitive results with state-of-the-art on closed-ended VQA and the best average spatial-reasoning score on the Spatial-Med benchmark. Ablation studies confirm that the hybrid block and HSOR contribute complementary gains, and that the induced spatial structure can substitute for raw language-model scale on volumetric reasoning tasks.
Chinese Translation
自监督预训练在三维医学图像分析中至关重要,其中未标记的CT体积丰富,但专家注释稀缺。然而,现有的体积编码器往往无法保留下游推理所依赖的粗略空间和几何结构,从而限制了它们在器官解缠、异常检测和与语言模型结合时的空间理解能力。我们提出了Rad-JEPA 3D,这是一种联合嵌入预测框架,通过预测完整扫描的潜在特征来学习体积CT表示,基于一个被遮蔽的视图。其核心是一个混合H-Mamba编码器,该编码器融合了一个Mamba状态空间分支,该分支通过顺序扫描建模切片间的连续性,以及一个分组查询注意力分支,该分支捕捉跨平面的空间上下文,并通过轻量级的每个标记路由器进行组合。为了提高中间表示的质量,我们进一步提出了隐藏状态正交正则化(HSOR),该方法对齐学生-教师隐藏状态,并减少编码器中的特征冗余。这种逐层正则化产生了更一致和更具区分性的体积表示,从而在器官识别和空间推理任务上提高了性能。在大约120,000个CT扫描上进行预训练的Rad-JEPA 3D,尽管其体积小巧,但仍然达到了最先进的结果:仅用4.0B的总参数量,在封闭式VQA上取得了与最先进技术竞争的结果,并在Spatial-Med基准上获得了最佳的平均空间推理分数。消融研究确认混合块和HSOR提供了互补的增益,并且诱导的空间结构可以替代体积推理任务中原始语言模型的规模。
cs.CV / 7 / 2607.26203

WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models

WildShadowRemover:基于细节保留的视频扩散模型的野外视频阴影去除
Xu, Jiamin, Wang, Cong, Dong, Zheng, Wang, Chi, Gu, Renshu, Xu, Weiwei, Xu, Gang
Abstract
Video shadow removal in the wild remains challenging due to complex illumination, diverse shadow appearances, and limited training data. Despite its importance to numerous vision and graphics applications, it remains largely unexplored in unconstrained real-world scenarios. To address this gap, we present WildShadowRemover, a framework that adapts a pretrained video diffusion model for robust video shadow removal via LoRA fine-tuning. To preserve fine image details while retaining the model's powerful generative prior, we augment the frozen VAE decoder with a detail injection module and introduce a shadow-mask-guided frequency-decomposed modulation module to selectively restore high-frequency textures while suppressing shadow artifacts. Monocular depth priors from Depth Anything 3 further provide geometry-aware guidance under challenging lighting conditions. We also construct WildShadow, a large-scale paired video shadow removal dataset and benchmark, covering diverse synthetic scenes. Extensive experiments demonstrate that our method outperforms existing approaches in shadow removal quality and temporal consistency, producing temporally coherent shadow-free videos with superior visual quality and strong generalization across challenging in-the-wild scenarios.
Chinese Translation
在复杂的光照条件、多样的阴影外观以及有限的训练数据的影响下,野外视频阴影去除仍然面临挑战。尽管这一任务对众多视觉和图形应用至关重要,但在不受约束的真实场景中仍然未得到充分探索。为了解决这一问题,我们提出了WildShadowRemover,一个通过LoRA微调适应预训练视频扩散模型以实现稳健视频阴影去除的框架。为了在保留模型强大生成先验的同时保持细腻的图像细节,我们在冻结的变分自编码器(VAE)解码器中增强了一个细节注入模块,并引入了一个阴影掩膜引导的频率分解调制模块,以选择性地恢复高频纹理,同时抑制阴影伪影。来自Depth Anything 3的单目深度先验进一步提供了在挑战性光照条件下的几何感知指导。我们还构建了WildShadow,一个大型配对视频阴影去除数据集和基准,涵盖多样的合成场景。大量实验表明,我们的方法在阴影去除质量和时间一致性方面优于现有方法,能够生成时序一致的无阴影视频,具有卓越的视觉质量和在挑战性野外场景中的强泛化能力。
cs.CV / 8 / 2607.26207

Where Physics Meets Privacy: Federated PINNs for Privacy-Preserving Brain Tumor Biomechanical Modeling

物理与隐私的交汇:用于隐私保护的联邦物理信息神经网络在脑肿瘤生物力学建模中的应用
Sristy, Mahmuda Akter, Chowdhury, Md Al-Mahfuz, Meem, Momota Ahsana, Ahamed, Sajid, Subhan, Kazi Irfan
Abstract
Brain tumors such as glioma, meningioma, and pituitary adenoma alter the mechanical behavior of soft brain tissue, yet common diagnostic methods rely on static imaging that cannot capture tumor growth, tissue displacement, or changes in stiffness over time. Deep learning models for this task typically require pooling patient data at one site, which conflicts with privacy rules such as GDPR and HIPAA and limits generalization across institutions, a challenge that is pronounced in neuro oncology given patient diversity. This study presents a federated physics informed neural network combining federated learning with a physics informed loss built on the equations of linear elasticity. Three simulated clinical sites each train a local network on patient specific MRI data using a physics informed loss, and only model weights are shared with a central server through the FedAvg protocol over one hundred rounds, keeping raw data at its site of origin. The federated model reached an overall accuracy of 91.4%, against 90.0% for a non federated baseline trained on pooled data, an average AUC of 0.985 across tumor classes, and a rise in pituitary tumor accuracy from 85.6 to 94.5%. Training produced smooth, divergence free displacement fields consistent with expected tissue deformation, showing that federated training can be paired with physics based constraints without a meaningful loss in performance.
Chinese Translation
脑肿瘤如胶质瘤、脑膜瘤和垂体腺瘤会改变软脑组织的机械行为,然而常见的诊断方法依赖于静态成像,无法捕捉肿瘤生长、组织位移或随时间变化的刚度。用于此任务的深度学习模型通常需要在一个地点汇总患者数据,这与GDPR和HIPAA等隐私规则相冲突,并限制了跨机构的泛化能力,这在神经肿瘤学中尤为突出,因为患者的多样性。本研究提出了一种联邦物理信息神经网络,将联邦学习与基于线性弹性方程的物理信息损失相结合。三个模拟临床站点各自使用患者特定的MRI数据训练本地网络,并采用物理信息损失,只有模型权重通过FedAvg协议在一百轮中与中央服务器共享,保持原始数据在其来源地点。联邦模型的整体准确率达到了91.4%,而基于汇总数据训练的非联邦基线为90.0%,在肿瘤类别中平均AUC为0.985,垂体肿瘤的准确率从85.6%上升至94.5%。训练产生了平滑的、无散度的位移场,与预期的组织变形一致,表明联邦训练可以与基于物理的约束相结合,而不会显著损失性能。
cs.CV / 9 / 2607.26215

Lag-aware cross-hand alignment for dual-hand action segmentation

考虑延迟的双手动作分割交叉手对齐
Ziaeetabar, Fatemeh
Abstract
Dual-hand action segmentation commonly fuses left- and right-hand representations at identical temporal indices, although coordinated hand transitions may occur with nonzero and time-varying delays. We introduce Lag-Aware Cross-Hand Alignment (LACA), a lightweight module that explicitly estimates directional temporal-offset distributions between hand-specific feature streams. LACA retrieves cross-hand information from the estimated offsets and incorporates a learned null state to suppress transfer when no compatible cross-hand transition is supported. Alignment is supervised using compatibility-aware targets derived automatically from frame-level training annotations, without requiring additional labels. Analysis of the HA-ViD and ATTACH training annotations reveals robust nonzero cross-hand matches for 44.7% and 48.9% of transition anchors, respectively, compared with 18.6% and 21.3% under temporally shifted controls. When integrated into Polyphony, LACA improves the two-hand mean F1@50 from 40.4 to 42.5 and boundary F1 from 56.5 to 59.6 on HA-ViD, and from 19.9 to 21.8 and 44.7 to 47.9, respectively, on ATTACH, relative to our reproduced Polyphony baseline. These gains require only approximately 0.0086 million additional trainable parameters. We further introduce LACA-C, a future-free variant that restricts alignment and the complete inference pipeline to current and past observations. On ATTACH, LACA-C achieves 83.6% transition-cue recall, a seed-averaged median availability delay of 233~ms, 0.72 false cues per minute, and segmentation-stage throughput of 224.9 current-position predictions per second. These results demonstrate that explicit cross-hand temporal alignment improves both action segmentation and boundary localization while supporting timely future-free perception.
Chinese Translation
双手动作分割通常在相同的时间索引处融合左右手的表示,尽管协调的手部过渡可能会发生非零且随时间变化的延迟。我们提出了考虑延迟的交叉手对齐(Lag-Aware Cross-Hand Alignment, LACA),这是一个轻量级模块,明确估计手特征流之间的方向性时间偏移分布。LACA从估计的偏移中检索交叉手信息,并结合学习的空状态,以在没有兼容的交叉手过渡时抑制转移。对齐是通过自动从帧级训练注释中导出的兼容性感知目标进行监督的,无需额外标签。对HA-ViD和ATTACH训练注释的分析显示,分别有44.7%和48.9%的过渡锚点存在稳健的非零交叉手匹配,而在时间偏移控制下分别为18.6%和21.3%。当集成到Polyphony中时,LACA将双手的平均F1@50从40.4提高到42.5,边界F1从56.5提高到59.6,在HA-ViD上,相应地在ATTACH上从19.9提高到21.8,44.7提高到47.9,相对于我们复现的Polyphony基线。这些增益仅需约0.0086百万个额外的可训练参数。我们进一步介绍了LACA-C,一个无未来信息的变体,它将对齐和完整推理管道限制为当前和过去的观察。在ATTACH上,LACA-C实现了83.6%的过渡线索召回,种子平均中位延迟为233毫秒,每分钟0.72个假线索,以及224.9个当前位置信息的每秒分割阶段吞吐量。这些结果表明,明确的交叉手时间对齐改善了动作分割和边界定位,同时支持及时的无未来感知。
cs.CV / 10 / 2607.26232

BG-REAL: A Public Real-Data Anchored Benchmark for Background Manipulation Detection and Localization

BG-REAL:一个基于真实数据的公共背景操控检测与定位基准
Uluirmak, Bugra Alperen, Kurban, Rifat
Abstract
Background manipulation is a practical but under-specified image-forensics setting: the manipulated evidence can sit outside the salient foreground object, while many evaluations emphasize object-centric copy-move, splicing, or generic synthetic edits. We introduce BG-REAL, a public real-data anchored benchmark package for background manipulation detection and localization. The current release is built from Open Images V7 instance-segmentation sources and contains 7,000 processed samples over 1,200 source groups, including 6,000 public-data anchored samples and 1,000 synthetic control samples. BG-REAL covers six edit families, matched authentic controls, source-group splits, mask and leakage QA, 599 human-assisted quality-control rows, three completed external baselines (TruFor, MVSS-Net, and HiFi-Net), and five-seed model evaluation. Beyond aggregate accuracy, we use matched-authentic-control diagnostics to measure how often baselines misclassify re-encoded authentic images as manipulated at a threshold fixed on held-out validation data; false-positive rates range from 0.57 (TruFor, the lowest) to 1.00 (several weak or mask-informed baselines), indicating that re-encoding artifacts are a shared shortcut risk across baselines rather than a problem specific to any one model. The release provides the construction pipeline, evaluation protocol, paper-ready figures, and reproduction documentation. We frame BG-REAL as a background-manipulation-focused complement to general image-manipulation-localization benchmarks, not as a fully real-only or general-purpose benchmark.
Chinese Translation
背景操控是一种实际但未充分定义的图像取证设置:被操控的证据可能位于显著前景物体之外,而许多评估则强调以物体为中心的复制移动、拼接或通用合成编辑。我们推出了BG-REAL,一个基于真实数据的公共基准包,用于背景操控检测与定位。当前版本基于Open Images V7实例分割源构建,包含7,000个处理样本,覆盖1,200个源组,其中包括6,000个公共数据锚定样本和1,000个合成控制样本。BG-REAL涵盖六种编辑类型,匹配的真实控制样本,源组划分,掩码和泄漏质量保证,599个人工辅助质量控制行,三个已完成的外部基线(TruFor、MVSS-Net和HiFi-Net),以及五种种子模型评估。除了总体准确性外,我们使用匹配的真实控制诊断来测量基线在固定于保留验证数据的阈值下,错误分类重新编码的真实图像为操控图像的频率;假阳性率范围从0.57(TruFor,最低)到1.00(多个弱或掩码信息基线),表明重新编码伪影是各基线共享的快捷风险,而不是特定于某一模型的问题。该版本提供了构建管道、评估协议、论文准备好的图形和重现文档。我们将BG-REAL框架视为一个专注于背景操控的补充,旨在与一般图像操控定位基准相辅相成,而不是作为一个完全真实或通用的基准。
cs.CV / 11 / 2607.26234

Spline-Based Boundary Representations for Sparse View Reconstruction and Simulation Using Isogeometric Analysis

基于样条的边界表示用于稀疏视图重建和模拟的等几何分析
Dobrota, Davor, Skorokhodov, Vsevolod, Xu, Chenghao, Fink, Olga, Mielle, Malcolm
Abstract
Image-based reconstruction aims to recover three-dimensional geometry from images. Recent advances have enabled the recovery of visually detailed models, yet their representations are not well-suited for numerical simulation. Simulation frameworks typically require explicit, watertight, and smooth geometries to ensure numerical robustness and accuracy, properties that surfaces extracted from image-based reconstructions lack. We propose FORGE-SIM, a method to directly reconstruct a multi-patch B-spline boundary representation from sparse posed RGB images without manual intervention. By optimizing the spline representation itself, our approach produces compact, smooth, and watertight geometries that are natively compatible with both Computer Aided Design and simulation workflows. Additionally, we introduce a strategy to project observation-derived fields, such as a thermal state and semantic information, onto the reconstructed models in the same spline basis, enabling immediate use in simulation. We demonstrate that the obtained models are of sufficiently high quality to enable thermal simulation and modal analysis. By unifying image-based reconstruction and simulation-ready modeling within a single optimization framework, this work removes a long-standing barrier between computer vision and numerical analysis. We anticipate that it will enable new workflows for simulation-driven design, inspection, and digital twin applications.
Chinese Translation
基于图像的重建旨在从图像中恢复三维几何形状。最近的进展使得恢复视觉上详细的模型成为可能,但这些模型的表示并不适合数值模拟。模拟框架通常需要明确的、密闭的和平滑的几何形状,以确保数值的稳健性和准确性,而从基于图像的重建中提取的表面缺乏这些特性。我们提出了FORGE-SIM,一种直接从稀疏姿态的RGB图像重建多补丁B样条边界表示的方法,无需人工干预。通过优化样条表示本身,我们的方法生成紧凑、平滑且密闭的几何形状,这些形状与计算机辅助设计和模拟工作流程本质上兼容。此外,我们引入了一种策略,将观察导出的场(如热状态和语义信息)投影到同一样条基上的重建模型中,从而使其能够立即用于模拟。我们证明,获得的模型质量足够高,可以进行热模拟和模态分析。通过在单一优化框架内统一基于图像的重建和适合模拟的建模,本研究消除了计算机视觉与数值分析之间长期存在的障碍。我们预计这将为基于模拟的设计、检查和数字双胞胎应用开启新的工作流程。
cs.CV / 12 / 2607.26237

LumaGuide: Distribution Shaping for Training-Free HDR Generation in Diffusion Models

LumaGuide:无训练的扩散模型高动态范围生成的分布塑形
Chen, Bowen, Saini, Shreshth, Adsumilli, Balu, Bovik, Alan C.
Abstract
Pretrained diffusion models generate realistic images but are constrained by the statistical biases of their training data, limiting their ability to produce high dynamic range (HDR) content. In this work, we introduce LumaGuide, a training-free framework for distribution shaping in diffusion models. Instead of modifying model parameters, LumaGuide steers the sampling process to match target feature distributions via differentiable energy-based guidance. We instantiate this framework for HDR generation by controlling luminance distributions in perceptually uniform PQ space. Our results show that aligning luminance histograms is sufficient to induce HDR-consistent behavior, including coherent highlights and preserved shadow detail, while maintaining semantic fidelity. Beyond HDR, LumaGuide enables flexible specification of target distributions through data-driven presets, reference images, or text-driven predictors, and extends naturally to video generation with temporal consistency constraints. More broadly, our work demonstrates that controllable generation can be achieved by directly shaping output distributions at sampling time, without retraining diffusion models.
Chinese Translation
预训练的扩散模型能够生成逼真的图像,但受到训练数据统计偏差的限制,限制了其生成高动态范围(HDR)内容的能力。在本研究中,我们提出了LumaGuide,一个用于扩散模型中分布塑形的无训练框架。LumaGuide通过可微分的基于能量的引导,调整采样过程以匹配目标特征分布,而不是修改模型参数。我们通过控制感知均匀PQ空间中的亮度分布,将该框架实例化为HDR生成。我们的结果表明,调整亮度直方图足以引导HDR一致的行为,包括连贯的高光和保留的阴影细节,同时保持语义的保真度。除了HDR,LumaGuide还通过数据驱动的预设、参考图像或基于文本的预测器,灵活指定目标分布,并自然扩展到具有时间一致性约束的视频生成。更广泛地说,我们的工作表明,通过在采样时直接塑形输出分布,可以实现可控生成,而无需重新训练扩散模型。
cs.CV / 13 / 2607.26238

Lightweight Image Classification of Raptor Species for Edge Devices: Rare-Species Dataset Expansion via Video Frame Extraction, Knowledge Distillation, and TensorRT Deployment

边缘设备的轻量级猛禽物种图像分类:通过视频帧提取、知识蒸馏和TensorRT部署扩展稀有物种数据集
Nishikawa, Takeshi
Abstract
We investigate lightweight raptor-species classification for real-time edge deployment in wind-turbine collision mitigation. Using DINOv2-L (304M parameters) as a teacher, we distilled three lightweight students (MobileNetV4, ViT-Small, and EfficientNet-B0). To reduce confusion between closely related species, we expanded the dataset to 12,519 images, including an increase in Steller's Sea Eagle images from 463 to 2,050 via video-frame extraction. Under a group split that separates samples at the video- and source-image level to mitigate source leakage at that granularity, the three-student ensemble achieved a macro recall of 0.935 +/- 0.004 over five distillation seeds (0.955 on a conventional image-level split, retaining 97.5% of the teacher's macro recall) with roughly one-eighth as many parameters. On a subset of 1,258 images disjoint from the former training images, White-tailed Eagle recall improved by up to 38.6 percentage points, while the rate at which it was misclassified as the Steller's Sea Eagle decreased from 61% to 15% of errors. TensorRT FP16 deployment of EfficientNet-B0 on an NVIDIA Jetson Orin Nano achieved 3.19 ms/image including host-device transfer (313 images/s), with 99.95% argmax agreement with FP32. In five-seed controlled comparisons, neither distillation (versus CE-only) nor the change from a DINOv2-L to a DINOv3-L teacher yielded a clear ensemble-level improvement; the primary gains stem from the dataset expansion and teacher re-fine-tuning.
Chinese Translation
我们研究了轻量级猛禽物种分类,以实现风力涡轮机碰撞缓解的实时边缘部署。使用DINOv2-L(304M参数)作为教师,我们蒸馏了三种轻量级学生模型(MobileNetV4、ViT-Small和EfficientNet-B0)。为了减少近亲物种之间的混淆,我们将数据集扩展至12,519张图像,其中通过视频帧提取将斯特勒海雕的图像数量从463张增加到2,050张。在一个按视频和源图像级别分组的拆分下,以减轻该粒度下的源泄漏,三学生集成在五个蒸馏种子上达到了0.935 +/- 0.004的宏召回率(在传统的图像级别拆分上为0.955,保留了教师宏召回率的97.5%),参数数量大约为原来的八分之一。在与之前训练图像不重叠的1,258张图像的子集上,白尾鹰的召回率提高了多达38.6个百分点,而其被错误分类为斯特勒海雕的比例从61%降至15%。在NVIDIA Jetson Orin Nano上,EfficientNet-B0的TensorRT FP16部署实现了3.19毫秒/图像的延迟(313张图像/秒),与FP32的argmax一致率达到99.95%。在五个种子控制比较中,蒸馏(与仅CE相比)或从DINOv2-L到DINOv3-L教师的变化均未带来明显的集成级别改进;主要的增益来自于数据集扩展和教师的再调优。
cs.CV / 14 / 2607.26276

Comparing the Performance of Foundation Model Derived Embeddings with Traditional Approaches for Distant Metastasis Prediction in Head and Neck Cancer

比较基础模型衍生嵌入与传统方法在头颈癌远处转移预测中的性能
Schmitz, Erich, Chen, Meixu, Jing, Bowen, Wang, Jing
Abstract
Background: Early prediction of distant metastasis (DM) risk in head and neck cancer (HNC) can enable timely interventions that may improve treatment outcomes. Many current machine learning methods rely on prior knowledge of the region of interest such as tumor segmentations, which require expert knowledge, is time-consuming and introduces user-dependent variability. Medical image-based foundation models have recently been developed for specific imaging modalities to streamline down-stream prediction tasks by extracting modality-relevant features. Purpose: In this study, we evaluate the effectiveness of using a foundation model as the feature extractor to predict DM risk in HNC patients and compare its performance with traditional approaches that require prior knowledge on the regions of interest. Methods: Preoperative CT images of 2327 patients from the RADCURE dataset were used. Three features-sets were created including radiomics, deep-learning based features, and CT Foundation derived features. The feature-sets were used individually in a multi-layer perceptron (MLP) to predict DM risk. Results: The model using CT Foundation embeddings outperformed the radiomics and deep learning-based models, achieving a Receiver Operating Characteristic Area Under the Curve (AUC) of 0.791, compared to AUC values of 0.772 and 0.753 for the radiomics and deep learning-based models, respectively. The CT Foundation based model had similar performance to a model that combined the use of radiomics and deep learning-based features that achieved an AUC of 0.794. Conclusions: Features based on foundation models offer a promising alternative to traditional radiomics while reducing the need for domain expertise and extensively annotated datasets. Their minimal preprocessing requirements also make them a more accessible and scalable option.
Chinese Translation
背景:早期预测头颈癌(HNC)患者的远处转移(DM)风险可以实现及时干预,从而改善治疗效果。许多当前的机器学习方法依赖于对感兴趣区域的先验知识,例如肿瘤分割,这需要专业知识,耗时且引入用户依赖的变异性。最近,针对特定成像模态开发了基于医学图像的基础模型,以通过提取与模态相关的特征来简化后续预测任务。目的:本研究评估使用基础模型作为特征提取器预测HNC患者DM风险的有效性,并将其性能与需要先验知识的传统方法进行比较。方法:使用RADCURE数据集中2327名患者的术前CT图像。创建了三个特征集,包括放射组学特征、基于深度学习的特征和CT基础衍生特征。将这些特征集单独用于多层感知器(MLP)以预测DM风险。结果:使用CT基础嵌入的模型在性能上优于放射组学和基于深度学习的模型,接收者操作特征曲线下面积(AUC)达到0.791,而放射组学和基于深度学习的模型的AUC值分别为0.772和0.753。基于CT基础的模型与结合放射组学和基于深度学习特征的模型表现相似,后者的AUC为0.794。结论:基于基础模型的特征为传统放射组学提供了一种有前景的替代方案,同时减少了对领域专业知识和大量注释数据集的需求。它们的最小预处理要求也使其成为更易获取和可扩展的选择。
cs.CV / 15 / 2607.26283

HeteroPROPMT: A Real-time and Privacy-Preserving Heterogeneous Collaborative Perception Framework

HeteroPROPMT:一种实时且保护隐私的异构协同感知框架
Maleki, Armin, Radha, Hayder
Abstract
Collaborative Perception (CP) improves autonomous systems' awareness of their surroundings by sharing sensor data, intermediate features, and detection results. In real-world deployments, however, collaborating vehicles often use heterogeneous sensors, perception models, datasets, and training domains, creating feature-space shifts that degrade downstream fusion and detection. Existing approaches typically retrain fusion and detection components or introduce modality-specific feature interpreters. These methods scale poorly to newly joining agents and often require access to proprietary metadata, raising privacy concerns. We propose HeteroPROMPT, a real-time and privacy-preserving framework for heterogeneous collaborative perception. HeteroPROMPT rapidly aligns each heterogeneous agent's features with an ego-centric unified feature space through modular prompts and lightweight learning-based tuning, while keeping agent encoders and the collaborative fusion and detection stacks frozen. Its visual prompt-based training and inference modulate Bird's Eye View (BEV) features across channels and spatial locations with low computational overhead. For metadata-free deployment, an autoencoder learns a compact unified representation and extracts modality cues from shared features, enabling real-time modality classification and routing to the appropriate HeteroPROMPT modules without exposing proprietary agent information. Experiments on the OPV2V-H and V2XSet datasets show that HeteroPROMPT improves Average Precision over state-of-the-art heterogeneous CP methods while using orders of magnitude fewer trainable parameters. This offers a scalable and practical CP solution. The proposed modality classifier also predicts the joining agent's modality from compact features with greater than 99.99 percent accuracy during deployment. Code will be available at https://github.com/arminmaleki007/HeteroPROMPT.
Chinese Translation
协同感知(Collaborative Perception, CP)通过共享传感器数据、中间特征和检测结果,提高了自主系统对周围环境的感知。然而,在实际部署中,协作车辆通常使用异构传感器、感知模型、数据集和训练领域,导致特征空间的偏移,从而降低了下游融合和检测的效果。现有方法通常需要重新训练融合和检测组件,或引入特定模态的特征解释器。这些方法在新加入的代理上扩展性较差,且通常需要访问专有元数据,从而引发隐私问题。我们提出了HeteroPROMPT,一种实时且保护隐私的异构协同感知框架。HeteroPROMPT通过模块化提示和轻量级学习调优,快速将每个异构代理的特征与以自我为中心的统一特征空间对齐,同时保持代理编码器和协作融合及检测堆栈的冻结。其基于视觉提示的训练和推理在通道和空间位置上调节鸟瞰图(Bird's Eye View, BEV)特征,计算开销低。为了实现无元数据的部署,自动编码器学习紧凑的统一表示,并从共享特征中提取模态线索,使得实时模态分类和路由到适当的HeteroPROMPT模块成为可能,而无需暴露专有代理信息。在OPV2V-H和V2XSet数据集上的实验表明,HeteroPROMPT在使用数量级更少的可训练参数的情况下,提升了与最先进的异构CP方法相比的平均精度。这提供了一种可扩展且实用的CP解决方案。所提出的模态分类器在部署期间也能以超过99.99%的准确率从紧凑特征中预测加入代理的模态。代码将可在 https://github.com/arminmaleki007/HeteroPROMPT 获取。
cs.CV / 16 / 2607.26292

Eddeep: a deep-learning framework for fast eddy-current distortion correction in diffusion MRI

Eddeep:用于快速涡流失真校正的深度学习框架在扩散MRI中的应用
Legouhy, Antoine, Callaghan, Ross, Qiao, Yuchuan, Stee, Whitney, Peigneux, Philippe, Azadbakht, Hojjat, Zhang, Hui
Abstract
Diffusion MRI (dMRI) relies on diffusion-weighted echo-planar imaging, which is highly susceptible to eddy-current-induced geometric distortions. These distortions vary across diffusion volumes according to gradient strength and direction, causing between-volume misalignment that can bias downstream microstructural analyses. Current state-of-the-art correction methods, such as FSL Eddy, achieve high-quality correction through iterative prediction-correction schemes but are computationally expensive. We propose Eddeep, a deep-learning framework for fast eddy-current distortion correction in dMRI. Eddeep decomposes the problem into two stages. First, a supervised image translation network standardises the appearance of diffusion-weighted and b=0 images, removing contrast differences that hinder reliable registration. Second, an unsupervised registration network estimates both eddy-current distortion and between-volume head motion parameters under a physics-constrained quadratic distortion model, enabling correction in a single forward pass. The method was trained on UK Biobank data and evaluated on both in-domain (UK Biobank) and out-of-domain (Memodyn) datasets. Across a range of complementary metrics, including between-volume jitter, diffusion kurtosis imaging residuals, signal irregularity, and mutual information, Eddeep achieved correction quality comparable to that of FSL Eddy while substantially reducing inference time. These results demonstrate that deep learning can provide accurate and efficient eddy-current distortion correction without relying on iterative optimisation, supporting the development of faster diffusion MRI processing pipelines for large-scale studies and clinical deployment. The code is available at: https://github.com/CIG-UCL/eddeep.
Chinese Translation
扩散MRI(dMRI)依赖于扩散加权回波平面成像,这种成像方式对涡流引起的几何失真高度敏感。这些失真会根据梯度强度和方向在扩散体积之间变化,导致体积间的错位,从而可能影响后续的微观结构分析。目前最先进的校正方法,如FSL Eddy,通过迭代预测-校正方案实现高质量的校正,但计算成本较高。我们提出了Eddeep,一个用于快速涡流失真校正的深度学习框架。Eddeep将问题分解为两个阶段。首先,一个监督图像转换网络标准化扩散加权图像和b=0图像的外观,消除妨碍可靠配准的对比差异。其次,一个无监督配准网络在物理约束的二次失真模型下估计涡流失真和体积间头部运动参数,从而实现单次前向传递的校正。该方法在英国生物银行数据上进行了训练,并在领域内(英国生物银行)和领域外(Memodyn)数据集上进行了评估。在一系列互补指标上,包括体积间抖动、扩散峰度成像残差、信号不规则性和互信息,Eddeep实现的校正质量与FSL Eddy相当,同时显著减少了推理时间。这些结果表明,深度学习可以在不依赖迭代优化的情况下提供准确高效的涡流失真校正,支持大型研究和临床应用中更快速的扩散MRI处理流程的发展。代码可在以下网址获取:https://github.com/CIG-UCL/eddeep。
cs.CV / 17 / 2607.26304

MoSAIC: Aligned Intervention Supervision for Part-Local Motion Style Transfer

MoSAIC:用于部分局部运动风格转移的对齐干预监督
Amini, Nazanin, Desai, Kevin
Abstract
Editing character motion often requires transferring a gesture or gait from one or more reference motions while preserving the source action, timing, root trajectory, and unselected body regions. Existing motion datasets, however, rarely provide paired targets for arbitrary part-local content--reference combinations, and self-reconstruction training may allow a diffusion model to reproduce the content motion while underusing the routed reference. We present MoSAIC, a latent diffusion framework for part-local reference-conditioned motion style transfer. MoSAIC factorizes content and reference features by anatomical region, preserves the root trajectory through a separate conditioning pathway, and routes user-selected references to individual body parts. Its central contribution is aligned intervention supervision, which constructs synchronized references and counterfactual targets through controlled local transformations, making both the requested regional response and the motion to be preserved directly observable during training. In a frozen evaluation comprising 128 motions and 896 routed conditions, part-masked routing reduces preserved-region error from 70.64 to 66.45~mm and matched-noise off-target leakage from 18.08 to 9.88~mm relative to whole-body routing, while retaining a positive selected-region response. A matched-budget continuation study further shows that retaining aligned intervention supervision produces an 8.8\% relative increase in selected-target response and a 2.0-percentage-point increase in requested-route influence concentration. These results demonstrate that MoSAIC improves the response--preservation trade-off required for selective and controllable part-local motion editing.
Chinese Translation
编辑角色运动通常需要在保留源动作、时序、根轨迹和未选择身体区域的同时,将一个或多个参考动作中的手势或步态转移过来。然而,现有的运动数据集很少提供任意部分局部内容与参考的配对目标,自我重建训练可能使扩散模型在重现内容运动时未能充分利用所路由的参考。我们提出了MoSAIC,一个用于部分局部参考条件运动风格转移的潜在扩散框架。MoSAIC通过解剖区域对内容和参考特征进行因式分解,通过单独的条件路径保留根轨迹,并将用户选择的参考路由到各个身体部位。其核心贡献是对齐干预监督,通过受控的局部变换构建同步参考和反事实目标,使得在训练过程中请求的区域响应和需保留的运动都能直接观察到。在包含128个动作和896个路由条件的冻结评估中,部分遮罩路由将保留区域误差从70.64毫米降低至66.45毫米,相较于全身路由,匹配噪声的目标泄漏从18.08毫米降低至9.88毫米,同时保持了正向的选择区域响应。一项匹配预算的延续研究进一步表明,保留对齐干预监督使得选择目标响应相对增加了8.8%,请求路由影响集中度提高了2.0个百分点。这些结果表明,MoSAIC改善了选择性和可控性部分局部运动编辑所需的响应与保留之间的权衡。
cs.CV / 18 / 2607.26326

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

看见还是知道?多模态大型语言模型中的视觉上下文敏感性
Li, Jiaang, Li, Chengzu, An, Zhaochong, Yuan, Yifei, Liu, Xi, Belongie, Serge, Snæbjarnarson, Vésteinn
Abstract
Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vision-centric tasks, especially when visual evidence conflicts with pretrained knowledge. We explore these failures separately using two diagnostic paradigms: (1) probing whether visual information is available, via image reconstruction, and (2) measuring multimodal context sensitivity, the extent to which the model follows visual context versus the language prior. To support the second, we introduce the WhatIfVis, a benchmark spanning five coarse-grained dimensions (spatial-temporal, color, count, size, and weight) whose questions admit answers from either the image or the prior. Our analysis yields three findings: (i) Coarse-grained visual evidence is preserved, as these attributes can be reconstructed from the final-layer image tokens of frozen MLLMs. Failures on questions about these attributes therefore point to post-perceptual utilization, rather than to degraded visual encoding during perception. (ii) Even when explicitly instructed to use or ignore visual evidence, vanilla models (without supervised fine-tuning on the WhatIfVis) show unstable visual context sensitivity. Supervised fine-tuning (SFT) improves this controllability and generalizes across domains, and activation patching further localizes the vision-versus-prior trade-off at architecture-specific depths across all six models. (iii) The vision-versus-prior trade-off is controllable along a learned vector. Applying this steering vector, even without any intent instruction, improves controllability over the vanilla model. Together, these results relocate the bottleneck, indicating that for the coarse attributes we study, MLLMs encode the visual evidence but cannot reliably control their reliance on it.
Chinese Translation
多模态大型语言模型(MLLMs)通过将视觉输入与预训练语言模型的丰富先验知识相结合,实现了强大的性能。然而,当视觉证据与预训练知识发生冲突时,它们在以视觉为中心的任务上往往表现不佳。我们通过两种诊断范式分别探讨这些失败:(1)通过图像重建探测视觉信息是否可用,以及(2)测量多模态上下文敏感性,即模型在多大程度上遵循视觉上下文而非语言先验。为了支持第二个范式,我们引入了WhatIfVis,这是一个涵盖五个粗粒度维度(时空、颜色、计数、大小和重量)的基准,其问题可以从图像或先验中获得答案。我们的分析得出了三个发现:(i)粗粒度视觉证据得以保留,因为这些属性可以从冻结的MLLMs的最终层图像标记中重建。因此,关于这些属性的问题的失败指向后感知利用,而不是在感知过程中视觉编码的退化。(ii)即使在明确指示使用或忽略视觉证据的情况下,普通模型(未在WhatIfVis上进行监督微调)显示出不稳定的视觉上下文敏感性。监督微调(SFT)改善了这种可控性,并在不同领域中具有普遍性,而激活补丁进一步在所有六个模型的特定架构深度上局部化了视觉与先验的权衡。(iii)视觉与先验的权衡可以沿着学习到的向量进行控制。即使在没有任何意图指示的情况下,应用这一引导向量也改善了对普通模型的可控性。综合这些结果,我们重新定位了瓶颈,表明对于我们研究的粗粒度属性,MLLMs编码了视觉证据,但无法可靠地控制对其的依赖。
cs.CV / 19 / 2607.26381

Zero-Fi: Zero-Shot Wi-Fi-Based Human Activity Recognition via Contrastive Signal-Language Alignment

Zero-Fi:通过对比信号-语言对齐实现的零样本基于Wi-Fi的人类活动识别
Shen, Yitong, Guo, Cheng, Wang, Peiliang, Zhang, Jingzhe, Sheng, Yi, Zhang, Haopeng, Xue, Hongfei, Ren, Yili
Abstract
Wi-Fi-based human activity recognition has advanced substantially, but most existing methods assume a closed set of activities and require labeled Wi-Fi samples for every target class, limiting their ability to recognize unseen activities. We present Zero-Fi, a contrastive signal-language alignment framework for zero-shot Wi-Fi-based human activity recognition. Zero-Fi learns unified representations from complementary Wi-Fi signal features and aligns them with the semantic representations of natural-language activity descriptions in a shared embedding space. This cross-modal alignment enables Zero-Fi to recognize new activity classes without requiring labeled Wi-Fi samples or model adaptation for those classes. Experiments on large-scale public benchmark datasets demonstrate effective zero-shot recognition of held-out activity classes, highlighting the potential of signal-language alignment to extend Wi-Fi sensing beyond predefined activity classes.
Chinese Translation
基于Wi-Fi的人类活动识别已经取得了显著进展,但大多数现有方法假设活动集是封闭的,并且需要每个目标类别的标记Wi-Fi样本,这限制了它们识别未见活动的能力。我们提出了Zero-Fi,一个用于零样本基于Wi-Fi的人类活动识别的对比信号-语言对齐框架。Zero-Fi从互补的Wi-Fi信号特征中学习统一表示,并将其与自然语言活动描述的语义表示在共享嵌入空间中对齐。这种跨模态对齐使得Zero-Fi能够识别新的活动类别,而无需为这些类别提供标记的Wi-Fi样本或模型适配。在大规模公共基准数据集上的实验展示了对保留活动类别的有效零样本识别,突显了信号-语言对齐在扩展Wi-Fi感知超越预定义活动类别的潜力。
cs.CV / 20 / 2607.26395

Registration-Grounded Spectral Fusion for Unregistered WLI/NBI Endoscopic Lesion Segmentation

基于注册的光谱融合用于未注册的白光成像/窄带成像内窥镜病变分割
Jie, Pengyu, Liu, Wanquan, He, Rui, Li, Pengcheng, Wen, Weiping, Meng, Deyu, Han, Junwei, Gao, Chenqiang
Abstract
White-light imaging (WLI) and narrow-band imaging (NBI) provide complementary views of endoscopic lesions, but their paired observations are often spatially misaligned due to viewpoint changes, tissue deformation, and sequential handheld acquisition. This makes direct WLI/NBI fusion prone to mixing non-corresponding regions and may even degrade segmentation around lesion boundaries. To address this problem, we propose a reliability-aware complex-domain fusion framework for paired-but-unregistered WLI/NBI lesion segmentation. The framework first establishes topology-regularized feature correspondence and further estimates where the cross-modal correspondence is reliable. Guided by this reliability, the model selectively fuses WLI and NBI features in a learnable complex representation. In this representation, WLI-derived cues mainly provide appearance-related magnitude responses, while NBI-derived cues provide structure-sensitive phase responses. Unlike conventional real-valued or symmetric multimodal fusion, the proposed method explicitly models the different roles of WLI and NBI and suppresses unreliable cross-modal interaction in locally mismatched regions. Experiments on paired WLI/NBI endoscopic datasets show that the proposed reliability-aware registration grounding and complex-domain fusion consistently improve lesion segmentation performance. Role-reversal and module ablation studies further validate the necessity of both the modality-role design and reliability-guided cross-modal interaction.
Chinese Translation
白光成像(WLI)和窄带成像(NBI)提供了内窥镜病变的互补视图,但由于视角变化、组织变形和手持采集的顺序,这些配对观察通常在空间上不对齐。这使得直接进行WLI/NBI融合容易混合不对应的区域,甚至可能降低病变边界周围的分割效果。为了解决这个问题,我们提出了一种可靠性感知的复数域融合框架,用于配对但未注册的WLI/NBI病变分割。该框架首先建立拓扑正则化的特征对应关系,并进一步估计跨模态对应关系的可靠性。在这种可靠性指导下,模型在可学习的复数表示中选择性地融合WLI和NBI特征。在这种表示中,WLI派生的线索主要提供与外观相关的幅度响应,而NBI派生的线索则提供结构敏感的相位响应。与传统的实值或对称多模态融合不同,所提出的方法明确建模了WLI和NBI的不同角色,并抑制了在局部不匹配区域内不可靠的跨模态交互。在配对的WLI/NBI内窥镜数据集上的实验表明,所提出的可靠性感知注册基础和复数域融合始终提高了病变分割性能。角色反转和模块消融研究进一步验证了模态角色设计和可靠性指导的跨模态交互的必要性。
cs.CV / 21 / 2607.26411

Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering

统一多模态模型是否在同一空间中思考?通过跨分支引导的视角
Wang, Yu, Li, Sharon
Abstract
Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture, yet it remains unclear whether these capabilities share a unified and transferable semantic space. This question is fundamentally challenging, as the two branches operate over heterogeneous representations (text tokens vs.\ visual latents) and distinct training objectives, making direct comparison difficult. To address this, we introduce \emph{cross-branch semantic steering}, an intervention-based framework that extracts semantic directions from one branch and applies them to the other. We show that steering vectors learned from the understanding branch can transfer to generation, enabling controllable image synthesis and improved semantic faithfulness. In contrast, the reverse direction consistently shows limited effectiveness. Our analysis suggests that this asymmetry may be related to a practical representational mismatch: understanding-derived vectors capture transferable, object-centric semantics, while generation-derived vectors primarily encode low-level appearance features. Our results reveal that architectural unification does not guarantee semantic alignment, and establish cross-branch steering as a practical tool for probing multimodal representations.
Chinese Translation
统一多模态模型(UMMs)旨在在单一架构内整合理解与生成能力,但尚不清楚这些能力是否共享统一且可转移的语义空间。这个问题本质上具有挑战性,因为这两个分支在异构表示(文本标记与视觉潜变量)和不同的训练目标上操作,使得直接比较变得困难。为了解决这个问题,我们引入了 extit{跨分支语义引导},这是一种基于干预的框架,能够从一个分支提取语义方向并将其应用于另一个分支。我们展示了从理解分支学习的引导向量可以转移到生成分支,从而实现可控的图像合成和提高语义忠实度。相反,反向方向的效果始终有限。我们的分析表明,这种不对称性可能与实际的表示不匹配有关:理解派生的向量捕捉可转移的以对象为中心的语义,而生成派生的向量主要编码低级外观特征。我们的结果揭示了架构的统一并不保证语义的一致性,并确立了跨分支引导作为探测多模态表示的实用工具。
cs.CV / 22 / 2607.26412

When Fish Look Alike: Tracking Identities with Dual-branch Elasticity

当鱼类相似时:利用双分支弹性跟踪身份
Lee, Vran, Liu, Xin, Wei, Yijie, Liu, Yeqiang, Leo, Hwa Liang, Li, Zhenbo
Abstract
Tracking dense, homogeneous targets like schooling fish remains a major challenge for multiple object tracking due to extreme inter-individual homogeneity, severe physical clustering, and rapid non-rigid deformations. While heavy-backbone separated detection and embedding trackers like SU-T push accuracy boundaries using complex Re-Identification networks, their computational overhead prohibits edge deployment. Furthermore, these modules often fail when appearance features degrade under severe occlusions. To overcome this, we propose Tracking Identities with Dual-branch Elasticity (TIDE). Bypassing expensive appearance cues, TIDE utilizes the Adaptive Geometric Correspondence IoU, an association mechanism leveraging spatial and structural consistency to robustly handle complex morphological variations. Crucially, TIDE introduces system-level deployment elasticity, decoupling the algorithmic pipeline from strict hardware constraints. Evaluations on the MFT-Edge benchmark demonstrate that our Lightweight L-branch achieves a competitive HOTA of 28.43 using merely 20.47G FLOPs. This represents a 38.7-fold computational reduction compared to upper bounds like SU-T, directly facilitating real-time edge deployment. Concurrently, our Scalable S-branch establishes a 29.98 HOTA, successfully bridging the gap between high-precision cloud analysis and efficient edge tracking. The dataset and codes are released at https://vranlee.github.io/TIDE/.
Chinese Translation
跟踪像鱼群这样密集且同质的目标仍然是多目标跟踪中的一大挑战,这主要由于个体之间的极端同质性、严重的物理聚集以及快速的非刚性变形。虽然像SU-T这样的重骨干分离检测和嵌入跟踪器通过复杂的再识别网络推动了准确性的边界,但其计算开销却限制了边缘部署。此外,当外观特征在严重遮挡下退化时,这些模块往往会失效。为了解决这个问题,我们提出了双分支弹性身份跟踪(Tracking Identities with Dual-branch Elasticity, TIDE)。TIDE绕过昂贵的外观线索,利用自适应几何对应IoU(Adaptive Geometric Correspondence IoU),一种利用空间和结构一致性的关联机制,以稳健地处理复杂的形态变化。至关重要的是,TIDE引入了系统级的部署弹性,将算法管道与严格的硬件限制解耦。对MFT-Edge基准的评估表明,我们的轻量级L分支在仅使用20.47G FLOPs的情况下,达到了竞争性的HOTA 28.43。这相比于像SU-T这样的上限,计算量减少了38.7倍,直接促进了实时边缘部署。同时,我们的可扩展S分支建立了29.98的HOTA,成功弥合了高精度云分析与高效边缘跟踪之间的差距。数据集和代码已发布在https://vranlee.github.io/TIDE/。
cs.CV / 23 / 2607.26432

FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing

FAS-R1:一种统一的多任务MLLM用于推理人脸反欺骗
Wang, Hongyang, Shi, Yichen, Li, Hongrui, Huo, Yiru, Feng, Jun, Yu, Zitong
Abstract
Face anti-spoofing (FAS) is increasingly expected to provide not only bona fide/spoof decisions, but also attack semantics and image-grounded evidence for human inspection. Existing discriminative FAS models remain largely label-centric, while recent MLLM-based methods offer structured outputs but still rely mainly on supervised fine-tuning, often producing template-like rationales and weak optimization for difficult attacks. We propose FAS-R1, a two-stage reasoning-oriented MLLM framework for unified FAS prediction, covering authenticity classification, attack-type recognition and spoof-region localization. FAS-R1 first uses FAS-R1-23K, a high-quality long-CoT dataset, for cold-start supervised fine-tuning, and then performs FAS-specific GRPO post-training. Degradation-Simulated Augmentation (DSA) encourages stable spoof-cue reasoning across visual-quality shifts, while Difficulty-Aware GRPO (DA-GRPO) mitigates easy-sample dominance that may leave difficult task--attack groups under-optimized, especially for subtle or ambiguous attacks such as makeup and mask attacks. The main 3B FAS-R1 model achieves 98.75\% authenticity accuracy, 93.33\% attack-type accuracy, and 96.30/94.73\% AP@40/AP@50 in-domain. It also outperforms the compared systems in cross-domain authenticity generalization and answer-and-rationale quality. Experiments with different base models further show favorable scaling behavior. The code will be released soon.
Chinese Translation
人脸反欺骗(FAS)越来越被期望不仅提供真实/欺骗的决策,还提供攻击语义和基于图像的证据以供人工检查。现有的判别性FAS模型在很大程度上仍然以标签为中心,而最近基于MLLM的方法虽然提供了结构化输出,但仍主要依赖于监督微调,常常产生模板式的推理和对困难攻击的优化不足。我们提出了FAS-R1,一个面向推理的两阶段MLLM框架,用于统一的FAS预测,涵盖真实性分类、攻击类型识别和欺骗区域定位。FAS-R1首先使用FAS-R1-23K,一个高质量的长链思维(long-CoT)数据集,进行冷启动的监督微调,然后执行FAS特定的GRPO后训练。降级模拟增强(DSA)鼓励在视觉质量变化中稳定的欺骗线索推理,而困难感知GRPO(DA-GRPO)减轻了可能导致困难任务-攻击组优化不足的简单样本主导,特别是对于化妆和面具等微妙或模糊的攻击。主要的3B FAS-R1模型在领域内实现了98.75%的真实性准确率、93.33%的攻击类型准确率,以及96.30%/94.73%的AP@40/AP@50。它在跨领域真实性泛化和答案与推理质量方面也优于对比系统。与不同基础模型的实验进一步显示了良好的扩展性。代码将很快发布。
cs.CV / 24 / 2607.26461

Interpretable Image-Level Acne Severity Grading via EfficientNet-B0 Transfer Learning and Grad-CAM

基于EfficientNet-B0迁移学习和Grad-CAM的可解释图像级痤疮严重程度分级
Zeng, Sophie, Kalaycioglu, Sean, Hong, Collin, Xie, Haipeng
Abstract
Acne vulgaris affects most adolescents and many adults. Accurate severity grading guides treatment, monitoring, and clinical trial endpoints, but manual assessment using the Investigator's Global Assessment or Hayashi criteria is limited by inter-rater variability and inconsistent imaging conditions. We developed a four-class acne severity classifier based on the Hayashi criteria using transfer learning with an ImageNet-pretrained EfficientNet-B0 model. The model was fine-tuned on the public ACNE04 dataset of 2,983 labeled images using AdamW optimization, geometric and photometric augmentation, and checkpoint selection based on validation macro-F1. On a held-out stratified 15 percent test set, the classifier achieved 93.5 percent accuracy and 94.4 percent macro-F1, with per-class F1 scores from 0.92 to 0.97. Eighty-three percent of errors occurred between adjacent grades. Quadratic-weighted Cohen's kappa was 0.956, with a 95 percent confidence interval of 0.935 to 0.973. Bootstrap confidence intervals indicated stable performance. Grad-CAM visualizations from the final convolutional block focused on clinically relevant facial regions, including the forehead, cheeks, and chin. The complete pipeline is provided as functionally equivalent open-source implementations in Python using PyTorch and timm, and in MATLAB R2026a. The software includes a clinician-facing inference interface and a fallback backbone option that supports operation without specialized pretrained-weight packages. These results show that lightweight transfer learning can provide accurate, balanced, and interpretable acne severity grading while offering a reproducible cross-platform reference for future prospective and device-stratified clinical validation.
Chinese Translation
痤疮(Acne vulgaris)影响大多数青少年和许多成年人。准确的严重程度分级有助于治疗、监测和临床试验终点,但使用研究者全球评估(Investigator's Global Assessment)或林(Hayashi)标准的手动评估受到评审者间变异性和成像条件不一致的限制。我们基于林标准开发了一种四类痤疮严重程度分类器,采用了基于ImageNet预训练的EfficientNet-B0模型的迁移学习。该模型在包含2983张标注图像的公共ACNE04数据集上进行了微调,使用了AdamW优化、几何和光度增强,并根据验证宏F1选择检查点。在一个保留的分层15%测试集上,分类器达到了93.5%的准确率和94.4%的宏F1,类别F1分数在0.92到0.97之间。83%的错误发生在相邻等级之间。二次加权Cohen's kappa为0.956,95%的置信区间为0.935到0.973。自助法置信区间表明性能稳定。来自最终卷积块的Grad-CAM可视化集中于临床相关的面部区域,包括额头、面颊和下巴。完整的管道以功能等效的开源实现形式提供,使用Python中的PyTorch和timm,以及MATLAB R2026a。该软件包括面向临床医生的推理接口和支持无专用预训练权重包操作的后备主干选项。这些结果表明,轻量级迁移学习可以提供准确、平衡和可解释的痤疮严重程度分级,同时为未来的前瞻性和设备分层临床验证提供可重复的跨平台参考。
cs.CV / 25 / 2607.26498

HERMES: A Hybrid Ensemble for Head-and-Neck Tumor Segmentation, TN Staging, and Recurrence-Free Survival on PET/CT

HERMES:一种用于头颈肿瘤分割、TN分期和无复发生存期的混合集成方法(PET/CT)
Wang, Kai, Chen, Meixu, Nasr, Elie, Lanning, Ryan, Miften, Moyed
Abstract
We present HERMES (Hybrid Ensemble for Radiotherapy-target segmentation, Malignancy staging, and Event-free Survival), a single containerized algorithm for the three HECKTOR 2026 subtasks: segmentation of the primary tumor (GTVp) and pathological lymph nodes (GTVn), radiological T/N staging, and recurrence-free survival (RFS), computed from a paired FDG-PET/CT scan and an electronic health record. A 10-fold ensemble of STU-Net Small networks produces the segmentation; the predicted mask then drives two downstream tasks. Rather than pass a generic radiomics vector to the staging models, we derive from the predicted masks a compact set of geometry features aligned with the size and number axes of AJCC/UICC 7th-edition radiological N/T staging. On internal cross-validation these features raise N-stage balanced accuracy from 0.691 to 0.720 (+0.030), our largest single design gain, at lower feature dimensionality. For prognosis we combine complementary deep and clinical risk experts in an equal-weight ensemble, and train one deep expert with a concordance-tracking survival loss of our own, whose value approximates the concordance index during training. Every component was selected on honest out-of-fold predictions under a regularization-oriented protocol, with no tuning on the public validation set, and deployed as two decorrelated submissions. On the HECKTOR 2026 validation leaderboard, HERMES achieved a weighted score of 0.6454 (Mean Dice 0.641, T balanced accuracy 0.580, N balanced accuracy 0.642, RFS C-index 0.679) and qualified for the testing phase. Team: AMC_HNC.
Chinese Translation
我们提出了HERMES(用于放疗靶向分割、恶性肿瘤分期和事件无生存期的混合集成方法),这是一个单一的容器化算法,针对HECKTOR 2026的三个子任务:原发肿瘤(GTVp)和病理淋巴结(GTVn)的分割、放射学T/N分期以及从配对的FDG-PET/CT扫描和电子健康记录中计算的无复发生存期(RFS)。一个10折的STU-Net Small网络集成生成分割结果;然后,预测的掩膜驱动两个下游任务。我们没有将通用的放射组学向量传递给分期模型,而是从预测的掩膜中提取出一组与AJCC/UICC第七版放射学N/T分期的大小和数量轴对齐的紧凑几何特征。在内部交叉验证中,这些特征将N阶段的平衡准确率从0.691提高到0.720(+0.030),这是我们在较低特征维度下获得的最大单一设计增益。对于预后,我们将互补的深度和临床风险专家结合在一个等权重的集成中,并用我们自己设计的共识追踪生存损失训练一个深度专家,其值在训练期间接近一致性指数。每个组件都是在一个以正则化为导向的协议下,根据诚实的外折预测进行选择的,没有在公共验证集上进行调优,并作为两个去相关的提交进行部署。在HECKTOR 2026验证排行榜上,HERMES获得了0.6454的加权得分(平均Dice 0.641,T平衡准确率0.580,N平衡准确率0.642,RFS C-index 0.679),并获得了测试阶段的资格。团队:AMC_HNC。
cs.CV / 26 / 2607.26511

Semantic-Aware Temporal Adaptation for UAV Anti-UAV Tracking

基于语义的无人机反无人机跟踪的时间适应
Qiao, Xiaozhen, Zhang, Da, Guo, Yubin, Gao, Junyu, Zhao, Zhiyuan, Li, Xuelong
Abstract
UAV Anti-UAV tracking is an emerging low-altitude security task for localizing an adversarial UAV using the onboard camera of a moving observer UAV. It differs from conventional UAV tracking and ground-based Anti-UAV tracking because both the camera platform and the target move simultaneously. This dual-dynamic setting induces rapid viewpoint changes, motion blur, scale variation, and visually similar distractors, making reliable appearance matching difficult. Under such rapidly changing conditions, fixed visual representations are often insufficient because target appearance becomes unreliable and feature distributions may deviate from the training domain. The target language description remains stable across frames and can therefore serve as a semantic anchor for temporal state propagation, while online feature-distribution alignment can reduce video-specific test-time shifts. In this paper, we propose \emph{SATATrack}, a Semantic-Aware Temporal Adaptation framework for UAV Anti-UAV tracking. SATATrack introduces Semantic-Aware Context Propagation (SACP), which uses the target description to guide temporal context propagation across backbone stages and preserve target identity under rapid appearance changes. An auxiliary contrastive regularizer is used during training to discourage responses to semantically similar background regions. During inference, Temporal-Aware Distribution Alignment (TADA) aligns feature distributions online without updating model parameters, combining recent-frame estimates with training-time statistics for stability. SATATrack achieves state-of-the-art performance on the UAV-Anti-UAV benchmark while remaining competitive in Anti-UAV and UAV object tracking tasks. The code will be available at https://github.com/XiaozhenQiao/SATATrack.
Chinese Translation
无人机反无人机跟踪是一项新兴的低空安全任务,旨在利用移动观察者无人机的机载摄像头定位对抗性无人机。与传统的无人机跟踪和地面反无人机跟踪不同,摄像平台和目标同时移动。这种双动态设置导致快速的视角变化、运动模糊、尺度变化以及视觉上相似的干扰物,使得可靠的外观匹配变得困难。在这种快速变化的条件下,固定的视觉表示往往不足,因为目标外观变得不可靠,特征分布可能偏离训练领域。目标的语言描述在各帧之间保持稳定,因此可以作为时间状态传播的语义锚,而在线特征分布对齐可以减少视频特定的测试时间偏移。本文提出了 extit{SATATrack},一种基于语义的无人机反无人机跟踪的时间适应框架。SATATrack引入了语义感知上下文传播(Semantic-Aware Context Propagation, SACP),利用目标描述指导跨主干阶段的时间上下文传播,并在快速外观变化下保持目标身份。在训练过程中使用辅助对比正则化器,以抑制对语义上相似的背景区域的响应。在推理过程中,时间感知分布对齐(Temporal-Aware Distribution Alignment, TADA)在线对齐特征分布,而无需更新模型参数,将最近帧的估计与训练时间统计结合以提高稳定性。SATATrack在无人机反无人机基准测试中实现了最先进的性能,同时在反无人机和无人机物体跟踪任务中保持竞争力。代码将发布在 https://github.com/XiaozhenQiao/SATATrack。
cs.CV / 27 / 2607.26518

EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

EgoSafe:一个用于视觉安全理解的第一人称移动捕捉基准
Chen, Yuyun, Li, Tianao, Feng, TianQuan, Chen, Cen, Zhuang, Huiping, Peng, Hao, Zeng, Ziqian
Abstract
Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic uncertainty. While Large Vision-Language Models (LVLMs) demonstrate impressive semantic alignment on standard benchmarks, they often struggle to distinguish between superficial correlation and genuine forensic logic when grounded in the dynamic, partially observable nature of first-person experiences. Existing evaluations, dominated by third-person surveillance footage and binary classification metrics, fail to expose this cognitive gap. To address this, we introduce EgoSafe-Bench, a benchmark specifically designed to probe forensic reasoning in egocentric safety scenarios. It comprises 12,000 unique evaluation samples, generated by pairing each of the 3,000 video clips with a QA chain governed by our proposed Hierarchical Reasoning Evaluation (HRE) protocol. Unlike standard benchmarks, HRE mandates a rigorous reasoning trajectory from initial feature anchoring to blind-spot deduction and intent inference, thereby enforcing logical consistency and penalizing shortcut-based predictions.Extensive evaluations of state-of-the-art LVLMs (e.g., Qwen3-VL, Gemini, VideoLLaMA 3) reveal a significant perception-reasoning decoupling: models often achieve high descriptive scores but exhibit notable fragility in causal reasoning and logical closure. Our work provides both a challenging dataset and a systematic evaluation framework to foster the development of logically robust video understanding systems.
Chinese Translation
在现实场景中,可靠的视觉安全理解不仅仅依赖于物体识别;它还需要在认知不确定性下进行因果推理。尽管大型视觉语言模型(Large Vision-Language Models, LVLMs)在标准基准上展现出令人印象深刻的语义对齐能力,但它们在面对第一人称体验的动态和部分可观察性时,往往难以区分表面相关性与真正的法医逻辑。现有的评估主要依赖于第三人称监控视频和二元分类指标,未能揭示这一认知差距。为了解决这一问题,我们提出了EgoSafe-Bench,这是一个专门设计用于探测自我中心安全场景中法医推理的基准。该基准包含12,000个独特的评估样本,通过将每个3,000个视频片段与一个由我们提出的层次推理评估(Hierarchical Reasoning Evaluation, HRE)协议指导的问答链配对生成。与标准基准不同,HRE要求从初始特征锚定到盲点推导和意图推断的严格推理轨迹,从而强制执行逻辑一致性并惩罚基于捷径的预测。对最先进的LVLM(例如Qwen3-VL、Gemini、VideoLLaMA 3)的广泛评估揭示了显著的感知-推理脱节:模型通常在描述性评分上表现良好,但在因果推理和逻辑闭合方面表现出明显的脆弱性。我们的工作提供了一个具有挑战性的数据集和一个系统的评估框架,以促进逻辑稳健的视频理解系统的发展。
cs.CV / 28 / 2607.26529

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling

CineWeaver:无训练的参考可控多镜头长视频生成用于电影叙事
Huang, Yuyang, Chen, Yabo, Dai, Wenrui, Zheng, Ziyang, Huang, Haibin, Zhang, Chi, Zou, Junni, Xiong, Hongkai, Li, Xuelong
Abstract
Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation across extended temporal horizons. Existing methods rely on customization and retraining to separately address specific requirements, and cannot simultaneously fulfill all the requirements with a unified framework. In this paper, we shed light on the training-free paradigm with the key insight that the difficulty of multi-shot generation arises from a structural bias toward temporal continuity in pretrained video diffusion models, and consequently, propose a unified framework named CineWeaver to achieve reference-controllable multi-shot long-video generation without retraining. We manipulate positional encoding and attention patterns to break temporal continuity during inference to enable clear shot transitions using pretrained video diffusion models. Furthermore, we extend the proposed framework with a shot-routed reference conditioning mechanism for per-shot fine-grained controllability, and develop an anchor memory mechanism to allow long-form generation with consistent global appearance cues. To our best knowledge, CineWeaver is the first unified framework to simultaneously enable \textbf{long-form}, \textbf{reference-controllable}, and \textbf{multi-shot} video generation in a training-free fashion. Experimental results demonstrate that CineWeaver produces high-quality cinematic videos of long durations with consistent identities, stable global appearance, and clear shot transitions. The project page is available at: https://cineweaver.github.io.
Chinese Translation
由于对多镜头生成、对角色和场景的细粒度可控性,以及在较长时间范围内的长格式生成的并发要求,电影视频生成对文本到视频扩散模型而言具有挑战性。现有方法依赖于定制和重新训练来分别解决特定需求,无法在统一框架内同时满足所有要求。本文阐明了无训练范式,关键见解在于多镜头生成的困难源于预训练视频扩散模型对时间连续性的结构偏差,因此提出了一个名为CineWeaver的统一框架,以实现无重新训练的参考可控多镜头长视频生成。我们操控位置编码和注意力模式,在推理过程中打破时间连续性,以利用预训练的视频扩散模型实现清晰的镜头过渡。此外,我们通过镜头路由的参考条件机制扩展了所提框架,以实现每个镜头的细粒度可控性,并开发了锚记忆机制,以允许具有一致全局外观线索的长格式生成。据我们所知,CineWeaver是第一个在无训练的方式下同时实现长格式、参考可控和多镜头视频生成的统一框架。实验结果表明,CineWeaver生成的高质量电影视频具有较长时长、一致的身份、稳定的全局外观和清晰的镜头过渡。项目页面可访问:https://cineweaver.github.io。
cs.CV / 29 / 2607.26536

TPCD: Tone-Pressure Contrastive Decoding and the Label-Free Gating Bottleneck in Vision-Language Models

TPCD:音调-压力对比解码及视觉-语言模型中的无标签门控瓶颈
Zhao, Jinkun, Zhang, Kui, Wu, Wenjun
Abstract
High-pressure prompts can push vision-language models (VLMs) into unsupported commitments, such as reading illegible text, reporting indeterminate times, or affirming absent objects. This paper asks whether the pressure-induced distribution itself can serve as a contrastive-decoding negative branch. Tone-pressure contrastive decoding (TPCD) subtracts logits produced under a high-pressure instruction from logits produced under a safe neutral instruction. On the 800-example tone-matters benchmark, LLaVA-1.5-7B under pressure reaches 66.75% attack success rate (ASR); safe neutralization reduces ASR to 9.88%; full TPCD reaches 0.50% but collapses positives to 15.56%. A benchmark-specific task-prior/disagreement gate preserves measured positive accuracy (54.44%) while lowering ASR to 1.63% on LLaVA. Treating this LLaVA analysis as the design split, full $n=800$ negative and $n=780$ matched-positive held-out runs on GLM-4.6V and Llama-3.2-Vision show that simple gates can improve over safe neutralization, with sensitivity analyses bounding the weak time-positive subtask. A category-prior-free answer-disagreement router reduces held-out aggregate ASR to 6.93%, improving over both safe neutralization (10.98%) and branch disagreement (9.67%) while matching branch disagreement's 79.94% positive accuracy, although it remains post-hoc and surface-form based. We conclude that pressure is a useful probe of commitment bias and a viable mitigation signal, but the current gates are not yet independently validated grounding-aware detectors.
Chinese Translation
高压提示可能会使视觉-语言模型(VLMs)陷入不支持的承诺,例如阅读难以辨认的文本、报告不确定的时间或确认缺失的物体。本文探讨了压力诱导的分布是否可以作为对比解码的负分支。音调-压力对比解码(TPCD)通过从高压指令下产生的 logits 中减去在安全中性指令下产生的 logits 来实现。在包含800个示例的音调重要性基准测试中,LLaVA-1.5-7B 在压力下达到66.75%的攻击成功率(ASR);安全中和将 ASR 降至9.88%;完整的 TPCD 达到0.50%,但将正例压缩至15.56%。特定基准任务的先验/分歧门控在降低 LLaVA 上的 ASR 至1.63%时保持了测量的正准确率(54.44%)。将此 LLaVA 分析视为设计分割,在 GLM-4.6V 和 Llama-3.2-Vision 上进行的完整 $n=800$ 负例和 $n=780$ 匹配正例的保留实验显示,简单的门控可以改善安全中和,敏感性分析界定了弱时间正子任务的边界。无类别先验的答案分歧路由器将保留的总体 ASR 降至6.93%,优于安全中和(10.98%)和分支分歧(9.67%),同时匹配分支分歧的79.94%正准确率,尽管它仍然是事后分析和表面形式基础。我们得出结论,压力是承诺偏见的有用探测工具和可行的缓解信号,但当前的门控尚未独立验证为具备基础感知的检测器。
cs.CV / 30 / 2607.26542

From Spatial Semantics to Temporal Context: Leveraging Gaze Trajectory for Weakly Supervised Medical Image Segmentation

从空间语义到时间上下文:利用注视轨迹进行弱监督医学图像分割
Wu, Shaoxuan, Zhang, Xiao, Zhao, Xiaodi, Tian, Yunzhi, Tang, Yilin, Feng, Jun
Abstract
Medical image segmentation heavily depends on labor-intensive and time-consuming pixel-level annotations. Eye tracking offers a cost-effective solution that can be naturally integrated into clinical workflows. Recorded by eye trackers, gaze conveys the spatial regions of clinicians' attention through fixations and the temporal context of clinicians' progressive visual perception from trajectories. Nevertheless, effective modeling of temporal trajectories remains challenging, and noise in gaze caused by exploratory fixations greatly limits segmentation performance. To overcome these limitations, we propose the Trajectory-guided Uncertainty-aware Network (TrailNet), which exploits gaze-supervised medical image segmentation from spatial semantics modeling to temporal context by jointly leveraging fixations and trajectories. Specifically, the proposed trajectory-guided spatio-temporal encoder models temporal context and establishes complementary interactions with image spatial semantics to strengthen target perception. Furthermore, the multi-scale uncertainty decoder leverages category mutual-exclusivity constraints to produce deterministic predictions and mitigate supervision uncertainty induced by noise. To enable gaze-free inference, we further introduce a cycle distillation strategy that transfers feature-level knowledge via teacher-student networks. Experimental results on two public datasets demonstrate that TrailNet outperforms state-of-the-art methods, achieving Dice scores of 81.25% and 81.85%, respectively.
Chinese Translation
医学图像分割严重依赖于劳动密集型和耗时的像素级标注。眼动追踪提供了一种成本效益高的解决方案,可以自然地融入临床工作流程。通过眼动仪记录的注视轨迹传达了临床医生注意力的空间区域以及临床医生逐步视觉感知的时间上下文。然而,有效建模时间轨迹仍然具有挑战性,探索性注视造成的注视噪声极大限制了分割性能。为克服这些限制,我们提出了轨迹引导的不确定性感知网络(Trajectory-guided Uncertainty-aware Network,TrailNet),该网络通过联合利用注视和轨迹,从空间语义建模到时间上下文,利用注视监督医学图像分割。具体而言,所提出的轨迹引导时空编码器建模时间上下文,并与图像空间语义建立互补交互,以增强目标感知。此外,多尺度不确定性解码器利用类别互斥约束生成确定性预测,并减轻由噪声引起的监督不确定性。为了实现无注视推理,我们进一步引入了一种循环蒸馏策略,通过教师-学生网络转移特征级知识。在两个公共数据集上的实验结果表明,TrailNet的性能超过了最先进的方法,分别达到了81.25%和81.85%的Dice分数。
cs.CV / 31 / 2607.26554

MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models

MedARC:无训练自适应冗余压缩视觉标记用于3D医学视觉-语言模型
Zhu, Yitao, Liu, Mengjun, Fu, Yingji, Pang, Haowen, Qiu, Anqi
Abstract
Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images generate prohibitively long visual-token sequences with considerable spatial and inter-slice redundancy. Existing token compression methods typically apply uniform reduction or rely on a single importance signal, increasing the risk of removing regions that are clinically relevant to the query or structurally distinctive. To address this limitation, we propose MedARC, a unified, training-free framework for Adaptive Redundancy Compression of visual tokens in 3D medical VLMs. MedARC estimates token importance by integrating three complementary cues: self-attention from the VLM vision encoder, which reflects the model's intrinsic visual focus; similarity between projected visual tokens and text embeddings, which identifies query-relevant regions; and deviations of local visual foundation model features from the volume-level feature center, which highlight structurally distinctive anatomy. The resulting importance distribution guides a saliency-aware merging strategy that preserves informative tokens while consolidating redundant ones rather than simply discarding them. Experiments on CT-RATE and MR-RATE show that MedARC reduces visual-token overhead and inference time while preserving or improving diagnostic performance. Its multi-cue scoring cost is outweighed by the savings from processing fewer tokens, with greater benefits expected for larger language models.
Chinese Translation
将3D医学图像与视觉-语言模型(VLMs)相结合,为计算机辅助诊断提供了巨大的潜力。然而,体积图像生成的视觉标记序列过长,且具有显著的空间和层间冗余。现有的标记压缩方法通常采用均匀减少或依赖单一重要性信号,这增加了去除与查询相关的临床区域或结构上独特区域的风险。为了解决这一限制,我们提出了MedARC,一个统一的、无训练的框架,用于在3D医学VLM中对视觉标记进行自适应冗余压缩。MedARC通过整合三种互补线索来估计标记的重要性:来自VLM视觉编码器的自注意力,反映了模型的内在视觉焦点;投影视觉标记与文本嵌入之间的相似性,识别与查询相关的区域;以及局部视觉基础模型特征与体积级特征中心的偏差,突出结构上独特的解剖特征。由此产生的重要性分布指导了一种关注显著性的合并策略,该策略在合并冗余标记的同时保留信息丰富的标记,而不是简单地丢弃它们。在CT-RATE和MR-RATE上的实验表明,MedARC减少了视觉标记的开销和推理时间,同时保持或提高了诊断性能。其多线索评分成本被处理更少标记所带来的节省所抵消,预计对于更大的语言模型将获得更大的收益。
cs.CV / 32 / 2607.26565

Representation Trajectories Matters: Complementary Evidence for OOD Detection and Image Classification

表示轨迹的重要性:对OOD检测和图像分类的补充证据
De la Jara, Ignacio M., Rodriguez-Opazo, Cristian, Damirchi, Hamed, Gould, Stephen, Ranasinghe, Damith
Abstract
Vision models do not form a representation at once; each block revises it. We ask whether the resulting computation path contains evidence that the final representation discards, and whether that evidence improves OOD detection and image classification on clean and shifted data. Unlike approaches that treat intermediate layers as separate snapshots, we retain sample identity across depth and study the transformations connecting successive states. We separate class-coherent transport from input-specific innovation, and coordinate movement from relational reorganization. Across supervised, self-supervised, vision--language, hierarchical, and convolutional encoders, these paths show strong sample-specific continuity and architecture-specific depth profiles that recur across datasets. They are also practically useful. An ID-only transition-surprise score complements strong final-state detectors, reducing FPR95 in 131/152 non-saturated comparisons on a balanced OpenOOD grid; gains are largest for visually disruptive and semantically far shifts, and remain positive on near-OOD for most detectors. Frozen update probes improve 71/72 clean model--dataset cases, while shifted-data gains vary with architecture and corruption type. Computation paths therefore provide a broadly useful reliability signal whose value is determined jointly by model organization and the shift encountered.
Chinese Translation
视觉模型并不是一次性形成表示;每个模块都会对其进行修正。我们探讨最终表示是否丢弃了计算路径中包含的证据,以及这些证据是否能改善在干净和偏移数据上的OOD检测和图像分类。与将中间层视为独立快照的方法不同,我们保留样本在深度上的身份,并研究连接连续状态的变换。我们将类一致的传输与输入特定的创新分开,并区分协调运动与关系重组。在监督、自监督、视觉-语言、层次化和卷积编码器中,这些路径显示出强烈的样本特异性连续性和特定架构的深度特征,这些特征在不同数据集间反复出现。它们在实践中也非常有用。仅基于ID的过渡惊讶分数补充了强大的最终状态检测器,在一个平衡的OpenOOD网格上,在131/152个非饱和比较中降低了FPR95;在视觉干扰和语义远离的偏移中,增益最大,并且在大多数检测器上,在近OOD情况下仍然保持正值。冻结更新探针改善了71/72个干净模型-数据集案例,而偏移数据的增益则因架构和损坏类型而异。因此,计算路径提供了一个广泛有用的可靠性信号,其价值由模型组织和所遇到的偏移共同决定。
cs.CV / 33 / 2607.26578

3DGBGS: 3D Granular Ball Gaussian Splatting for Compact Novel View Synthesis

3DGBGS:用于紧凑新视图合成的三维颗粒球高斯点云渲染
Yang, Meng, Xia, Shuyin, Dai, Dawei, YiWang
Abstract
Three-dimensional Gaussian Splatting (3DGS) enables high-quality real-time novel-view synthesis through explicit Gaussian primitives and differentiable rasterization. 3DGS and Granular Ball Computing (GBC), proposed in 2019, share a natural compatibility in adaptive representation. The efficiency of 3DGS partly stems from a coarse-to-fine and on-demand refinement process that draws on the generation principle of GBC. This connection motivates us to further introduce adaptive granular ball organization into anchor-based 3DGS. Existing anchor-based methods typically construct anchors from sparse SfM point clouds through fixed voxelization, which cannot adequately adapt to spatially non-uniform point distributions and leads to a trade-off among anchor count, model compactness, and rendering quality. To address this issue, we propose 3DGBGS (3D Granular Ball Gaussian Splatting), a compact anchor-based framework for novel-view synthesis. 3DGBGS adaptively partitions SfM point clouds into 3D granular balls, using larger balls to compactly represent smooth and redundant regions and smaller balls to preserve complex geometry and local details. Based on this representation, Granular Ball Anchor Initialization (GBAI) uses granular ball centers to initialize compact anchor positions, while the Granular Ball Scale Prior (GBSP) exploits granular ball radii to provide local scale priors for Gaussian generation. Experiments on four benchmarks show that 3DGBGS reduces initial and final anchors by 37.1% and 10.0%, respectively, and model storage by 9.8% on average, while maintaining comparable rendering quality.
Chinese Translation
三维高斯点云渲染(3DGS)通过显式高斯原语和可微光栅化技术实现高质量实时新视图合成。2019年提出的3DGS与颗粒球计算(GBC)在自适应表示方面具有天然的兼容性。3DGS的效率部分源于一种粗到细的按需细化过程,该过程基于GBC的生成原理。这一联系促使我们进一步将自适应颗粒球组织引入基于锚点的3DGS。现有的基于锚点的方法通常通过固定体素化从稀疏的结构光重建(SfM)点云构建锚点,这无法充分适应空间上不均匀的点分布,并导致锚点数量、模型紧凑性和渲染质量之间的权衡。为了解决这个问题,我们提出了3DGBGS(3D颗粒球高斯点云渲染),这是一个用于新视图合成的紧凑型基于锚点的框架。3DGBGS自适应地将SfM点云划分为三维颗粒球,使用较大的球体紧凑地表示平滑和冗余区域,而使用较小的球体保留复杂的几何形状和局部细节。在此表示基础上,颗粒球锚点初始化(GBAI)利用颗粒球中心初始化紧凑的锚点位置,而颗粒球尺度先验(GBSP)则利用颗粒球半径为高斯生成提供局部尺度先验。在四个基准测试中的实验表明,3DGBGS分别减少了初始和最终锚点数量37.1%和10.0%,并平均减少了模型存储9.8%,同时保持了可比的渲染质量。
cs.CV / 34 / 2607.26580

Classification of Disease from Lungs X-ray Images using VGG16, VGG19 and ResNet50 Models

基于 VGG16、VGG19 和 ResNet50 模型的肺部 X 光图像疾病分类
Yadav, Nand Lal, Kumar, Rajesh, Singh, Satyendra, Singh, Sudhakar
Abstract
With the increase in the number of cases related to respiratory diseases, there is an urgent need to detect them early and diagnose them accurately. Convolutional neural networks have given promising results when used for diagnosing diseases using imaging tests. In this study, we investigate the potential of applying deep learning algorithms such as VGG16, VGG19, and ResNet50 for classification of lung ailments based on X-ray images. A detailed analysis of the aforementioned models' performances was conducted to assess how well they can classify various types of lung ailments, including pneumonia, tuberculosis, lung cancer, and normal lungs. In order to do that, these deep learning models were trained on a vast amount of X-ray images. The results of our study show that while all three models provide good results, ResNet-50 performs best in comparison with other models due to its efficiency and high level of accuracy. We believe that these deep learning models can be successfully implemented in the practice of diagnosing pulmonary diseases in the future. It helps with early disease detection and improves patient outcomes.
Chinese Translation
随着与呼吸系统疾病相关病例数量的增加,迫切需要对其进行早期检测和准确诊断。卷积神经网络在使用影像学检查诊断疾病时取得了令人鼓舞的结果。本研究探讨了应用深度学习算法(如 VGG16、VGG19 和 ResNet50)对肺部疾病进行分类的潜力,基于 X 光图像。我们对上述模型的性能进行了详细分析,以评估它们在分类各种类型的肺部疾病(包括肺炎、结核病、肺癌和正常肺部)方面的效果。为此,这些深度学习模型在大量 X 光图像上进行了训练。我们的研究结果表明,尽管三种模型均提供了良好的结果,但 ResNet-50 相较于其他模型表现最佳,因其效率和高准确率。我们相信这些深度学习模型在未来的肺部疾病诊断实践中可以成功应用,有助于早期疾病检测并改善患者的预后。
cs.CV / 35 / 2607.26582

Level, Sharpness, and Corpus: Why Zero-Shot OOD Detector Rankings Do Not Transfer

水平、锐度与语料库:为何零-shot OOD 检测器排名无法转移
De la Jara, Ignacio M., Rodriguez-Opazo, Cristian, Gould, Stephen, Ranasinghe, Damith
Abstract
Selecting a zero-shot out-of-distribution (OOD) detector for a new deployment is typically based on benchmark rankings, implicitly assuming that the highest-ranked detector will transfer across domains. We show that this assumption does not hold. Through a controlled portability audit across seventeen in-distribution datasets, three vision-language models, and seven representative zero-shot OOD detectors, we find that detector rankings reverse across deployments, every detector exceeds $80\%$ FPR95 on at least one domain, and the preferred detector depends on both the in-distribution data and the underlying VLM. We trace these reversals to complementary evidence channels in vision-language logits. Corpus-free detectors rely on different combinations of absolute match level and relative or spatial sharpness, while WordNet-based methods additionally depend on external semantic coverage. A simple proposition shows that level and sharpness cannot generally be recovered from one another, explaining why no single detector transfers reliably across deployments. Motivated by this diagnosis, we introduce the Complementary Evidence Guard (CEG), a detector-agnostic wrapper that preserves complementary evidence through a non-compensatory fusion of the base detector, level, and sharpness using only empirical in-distribution percentiles. Controls replacing these channels with entropy, logit variance, or random noise do not reproduce the gains. Without OOD samples, auxiliary corpora, or learned fusion, CEG reduces detector sensitivity and improves GL-MCM from $38.1$ to $28.8$ and MCM from $42.6$ to $30.5$ family-balanced FPR95.
Chinese Translation
选择用于新部署的零-shot 分布外(OOD)检测器通常基于基准排名,隐含假设最高排名的检测器能够跨领域转移。我们证明这一假设并不成立。通过对十七个分布内数据集、三个视觉-语言模型和七个代表性的零-shot OOD 检测器进行控制可移植性审计,我们发现检测器排名在不同部署中会发生逆转,每个检测器在至少一个领域的 FPR95 超过 $80\%$,而首选检测器依赖于分布内数据和基础 VLM。我们将这些逆转归因于视觉-语言 logits 中的互补证据通道。无语料检测器依赖于绝对匹配水平和相对或空间锐度的不同组合,而基于 WordNet 的方法则额外依赖外部语义覆盖。一个简单的命题表明,水平和锐度通常无法相互恢复,这解释了为何没有单一检测器能够在不同部署中可靠转移。基于这一诊断,我们引入了互补证据保护器(Complementary Evidence Guard, CEG),这是一种与检测器无关的包装器,通过仅使用经验分布内百分位数对基础检测器、水平和锐度进行非补偿性融合,从而保留互补证据。用熵、logit 方差或随机噪声替代这些通道的控制实验未能重现收益。在没有 OOD 样本、辅助语料或学习融合的情况下,CEG 降低了检测器的敏感性,并将 GL-MCM 从 $38.1$ 降至 $28.8$,将 MCM 从 $42.6$ 降至 $30.5$,实现了家庭平衡的 FPR95 改进。
cs.CV / 36 / 2607.26583

R-SLPR: Region-based Small-to-Large Point-cloud Registration with Contrastive Learning

R-SLPR:基于区域的小到大点云配准与对比学习
Wan, Yusen, Chen, Zeyuan, Zou, Qianshi, Chen, Xu
Abstract
Point-cloud (PC) registration is fundamental to three-dimensional (3D) perception in robotic systems. However, classic registration algorithms falter when aligning a source PC containing limited, incomplete, or ambiguous geometric cues against a reference. This challenge of registering a small, partial PC to a significantly larger global reference is pervasive in real-world deployment yet remains insufficiently addressed by existing learning-based approaches, which typically assume comparable scales and significant overlap. To bridge this gap, we propose the Region-based Small-to-Large Point-cloud Registra- tion framework (R-SLPR), a novel three-stage architecture that fundamentally reformulates the scale-mismatched registration problem into a sequence of region proposal, regional matching, and iterative refinement. Unlike conventional methods that fail to localize specific regions, R-SLPR explicitly identifies candidate regions prior to estimating rigid transformations, ensuring robust alignment even under severe scale mismatch. The framework introduces a Fibonacci Grid Segmentation method coupled with a contrastive learning objective to effectively generate and match local geometric patches. Building on this, a novel Cascade Anchor Selection and Refinement algorithm iteratively aligns the source with the target region to maximize precision. Extensive evaluation on ModelNet40 demonstrates that R-SLPR establishes a new state-of-the-art accuracy standard, outperforming prior approaches and significantly reducing position and rotation Mean Absolute Error (MAE) to 0.009 and 1.104, respectively.
Chinese Translation
点云(PC)配准是机器人系统三维(3D)感知的基础。然而,当对齐一个包含有限、不完整或模糊几何线索的源点云与参考点云时,经典的配准算法往往表现不佳。将一个小的部分点云配准到一个显著更大的全局参考点云的挑战在实际应用中普遍存在,但现有的基于学习的方法对此问题的解决仍显不足,通常假设具有可比的尺度和显著的重叠。为了解决这一问题,我们提出了基于区域的小到大点云配准框架(R-SLPR),这是一种新颖的三阶段架构,根本上将尺度不匹配的配准问题重新构造为区域提议、区域匹配和迭代精炼的序列。与传统方法无法定位特定区域不同,R-SLPR在估计刚性变换之前明确识别候选区域,确保即使在严重的尺度不匹配下也能实现稳健的对齐。该框架引入了一种斐波那契网格分割方法,并结合对比学习目标,有效生成和匹配局部几何补丁。在此基础上,提出了一种新颖的级联锚点选择与精炼算法,迭代地将源区域与目标区域对齐,以最大化精度。在ModelNet40上的广泛评估表明,R-SLPR建立了新的最先进的准确性标准,超越了先前的方法,并显著将位置和旋转的平均绝对误差(MAE)降低至0.009和1.104。
cs.CV / 37 / 2607.26595

SpatialQ: Understanding 3D Gaussian Splatting Scene Quality via Visual-based MLLM

SpatialQ:通过基于视觉的多模态大语言模型理解3D高斯点云场景质量
Su, Jingxuan, Wang, Shenglin, Zhao, Tiesong, Li, Ge, Gao, Wei
Abstract
3D Gaussian Splatting (3DGS) has emerged as an effective representation for novel view synthesis and 3D scene reconstruction, creating an increasing demand for reliable quality assessment. Unlike conventional image quality assessment (IQA), the quality of a 3DGS scene depends not only on the perceptual fidelity of rendered views, but also on scene-level factors such as spatial structure and cross-view consistency. Existing IQA methods are limited by their reliance on 2D perceptual cues, whereas general multimodal large language models (MLLMs) are not designed for stable quality regression and may produce unreliable judgments. To address these limitations, a multimodal quality assessment framework is developed for 3DGS scene understanding. First, a 3D-aware quality representation learning framework is introduced by augmenting a VGGT-based encoder with a dedicated quality head. Multi-view images are encoded into view-specific features and aggregated to capture cross-view consistency, while geometric cues are incorporated through joint modeling of depth and point-cloud-related structural information, enabling the learning of structure-aware quality representations beyond appearance-driven features. Second, a grounded multimodal reasoning mechanism is constructed by jointly feeding original images, depth maps, point cloud renderings, and camera parameters into a Qwen-based MLLM.
Chinese Translation
3D高斯点云(3DGS)作为一种有效的表示方式,已在新视图合成和3D场景重建中崭露头角,因而对可靠的质量评估需求日益增加。与传统的图像质量评估(IQA)不同,3DGS场景的质量不仅依赖于渲染视图的感知保真度,还受到空间结构和视图间一致性等场景级因素的影响。现有的IQA方法受限于对二维感知线索的依赖,而通用的多模态大语言模型(MLLM)并未针对稳定的质量回归进行设计,可能产生不可靠的判断。为了解决这些局限性,本文开发了一种用于3DGS场景理解的多模态质量评估框架。首先,通过在VGGT(视觉生成图形变换)基础编码器上增加专用的质量头,提出了一种3D感知质量表示学习框架。多视图图像被编码为视图特定特征并聚合,以捕捉视图间的一致性,同时通过深度和点云相关结构信息的联合建模引入几何线索,使得学习超越外观驱动特征的结构感知质量表示成为可能。其次,构建了一种基础的多模态推理机制,通过将原始图像、深度图、点云渲染和相机参数共同输入到基于Qwen的MLLM中。
cs.CV / 38 / 2607.26596

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution

解耦视觉处理:通过特定模态的变换器替代实现高效的多模态适应
Feng, Mingkuan, Wen, Zhengqi, Tao, Jianhua
Abstract
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a unified transformer architecture. However, fine-tuning all parameters of these models for visual instruction tuning is computationally expensive and often unnecessary, as the representation requirements for visual and textual tokens diverge significantly in the deeper layers of the network. In this paper, we propose Decoupled Visual Processing (DVP), an efficient training framework that replaces the upper decoder layers of a pretrained LLM with a lightweight, independently trainable single transformer block dedicated exclusively to visual token processing. Specifically, after shared processing through the first half of the decoder layers, visual and textual tokens are split: visual tokens are routed through a newly initialized single transformer block while textual tokens continue through the original frozen decoder layers. The two streams are then concatenated before the language modeling head. During training, only the single transformer block is updated, dramatically reducing the number of trainable parameters. Experiments on the LLaVA-1.5 framework demonstrate that DVP achieves competitive performance on MME, POPE, and ChartQA benchmarks while training only a fraction of the total parameters, suggesting that visual representations in MLLMs can be effectively learned through a decoupled, parameter-efficient pathway.
Chinese Translation
多模态大型语言模型(MLLMs)通过在统一的变换器架构中整合视觉和文本理解,展现了显著的能力。然而,为视觉指令调优而微调这些模型的所有参数计算成本高昂,且通常是不必要的,因为在网络的深层中,视觉和文本标记的表示需求显著不同。本文提出了解耦视觉处理(DVP),这是一种高效的训练框架,它用一个轻量级、可独立训练的单一变换器块替代预训练大型语言模型(LLM)的上层解码器层,该块专门用于视觉标记处理。具体而言,在通过解码器层的前半部分进行共享处理后,视觉和文本标记被分开:视觉标记通过新初始化的单一变换器块进行处理,而文本标记则继续通过原始的冻结解码器层。然后,在语言建模头之前将这两条流进行拼接。在训练过程中,仅更新单一变换器块,显著减少了可训练参数的数量。在LLaVA-1.5框架上的实验表明,DVP在MME、POPE和ChartQA基准测试中实现了具有竞争力的性能,同时仅训练了总参数的一小部分,这表明在多模态大型语言模型中,视觉表示可以通过解耦的、参数高效的路径有效学习。
cs.CV / 39 / 2607.26600

JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation

JEPADepth:用于自监督单目深度估计的掩码预测表示学习
Grigore, Ionuţ, Popa, Călin-Adrian
Abstract
Self-supervised monocular depth estimation typically relies on photometric reconstruction losses that couple depth, pose, and appearance assumptions. In this paper, we propose JEPADepth, a self-supervised monocular depth framework that incorporates a complementary training objective inspired by Image Joint-Embedding Predictive Architectures (I-JEPA) for self-supervised depth learning. Our method augments a standard photometric pipeline with a masked prediction loss computed in the representation space of a pretrained DINOv3 Vision Transformer encoder. A predictor infers target-region embeddings from visible context-region embeddings under structured masking, and is discarded along with the target encoder at inference time, adding no deployment cost. On KITTI, adding the JEPA objective consistently improves performance over the same DINOv3-based photometric baseline, without changing the inference-time architecture. Compared to prior monocular self-supervised methods, JEPADepth is competitive with state-of-the-art transformer-based approaches and outperforms strong CNN-based baselines on the standard benchmark. In zero-shot transfer (trained on KITTI and evaluated without fine-tuning), JEPADepth achieves the best or near-best performance among the compared methods on both Make3D and Cityscapes across multiple metrics.
Chinese Translation
自监督单目深度估计通常依赖于光度重建损失,这些损失将深度、姿态和外观假设结合在一起。在本文中,我们提出了JEPADepth,一种自监督单目深度框架,结合了受图像联合嵌入预测架构(Image Joint-Embedding Predictive Architectures, I-JEPA)启发的互补训练目标,用于自监督深度学习。我们的方法在标准光度管道中增加了一个在预训练DINOv3视觉变换器编码器的表示空间中计算的掩码预测损失。预测器在结构化掩码下从可见上下文区域嵌入推断目标区域嵌入,并在推理时与目标编码器一起被丢弃,不增加部署成本。在KITTI数据集上,添加JEPA目标始终改善了相同基于DINOv3的光度基线的性能,而不改变推理时的架构。与之前的单目自监督方法相比,JEPADepth在与最先进的基于变换器的方法竞争时表现出色,并在标准基准上超越了强大的基于卷积神经网络(CNN)的基线。在零样本迁移(在KITTI上训练并在不进行微调的情况下评估)中,JEPADepth在Make3D和Cityscapes的多个指标上,在比较的方法中实现了最佳或接近最佳的性能。
cs.CV / 40 / 2607.26608

Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach

理解异构多模态大语言模型融合中的知识转移机制:一种简单的线性方法
Hou, Yinghao, Fan, Jiahe, Pu, Yuanhao, Chen, Zongyuan, Xie, Hong
Abstract
Training-free fusion of heterogeneous multimodal large language models (MLLMs) provides a direct route for cross-scale capability transfer, yet improvements in aggregate performance do not reveal what a smaller model actually inherits. Existing studies are largely designed and evaluated on limited task sets or aggregate metrics; as evaluation expands to broader task collections, whether different capabilities can transfer across scales remains poorly understood. To investigate this question, we introduce Cross-Scale Directional Parameter Injection (CDPI), a simple linear probe to analyze cross-scale knowledge transfer during heterogeneous fusion. A local theoretical analysis indicates that knowledge transfer selectivity is determined at first order by capability-dependent responses to a shared injection direction, while second-order curvature effects constrain the effective transfer regime. Across four Qwen3-VL model pairs and twelve multimodal benchmarks, our experiments reveal a consistent pattern of selectivity: gains concentrate on reasoning, particularly high-level reasoning, whereas perception performance remains close to that of the original target model. Component-wise ablations further show that high-level reasoning gains arise primarily from the language model, while ratio analysis finds that positive selective transfer occurs mainly in the small-ratio regime. These findings recast cross-scale heterogeneous MLLM fusion as selective language-side reasoning transfer within a narrow, low-interference regime, rather than broad capability inheritance.
Chinese Translation
无训练的异构多模态大语言模型(MLLMs)融合为跨尺度能力转移提供了一条直接路径,但整体性能的提升并未揭示较小模型实际继承了什么。现有研究主要在有限的任务集或汇总指标上进行设计和评估;随着评估扩展到更广泛的任务集合,不同能力是否能够跨尺度转移仍然不甚明了。为探讨这一问题,我们引入了跨尺度方向性参数注入(Cross-Scale Directional Parameter Injection, CDPI),这是一种简单的线性探针,用于分析异构融合过程中的跨尺度知识转移。局部理论分析表明,知识转移的选择性在一阶上由对共享注入方向的能力依赖响应决定,而二阶曲率效应则限制了有效转移的范围。在四对 Qwen3-VL 模型和十二个多模态基准测试中,我们的实验揭示了选择性的一致模式:收益集中在推理上,特别是高层次推理,而感知性能则接近原始目标模型的水平。组件逐项消融实验进一步表明,高层次推理收益主要来自语言模型,而比率分析发现,正向选择性转移主要发生在小比率范围内。这些发现将跨尺度异构 MLLM 融合重新定义为在狭窄、低干扰范围内的选择性语言侧推理转移,而非广泛的能力继承。
cs.CV / 41 / 2607.26641

FakeIDet3-DB: Refining Digital Attacks and Patch Extraction for Secure ID Benchmarking

FakeIDet3-DB:精炼数字攻击与补丁提取以实现安全的身份验证基准测试
Javier, Muñoz-Haro, Andres, Teruel, Ruben, Tolosana, Daniel, DeAlcala, Ruben, Vera-Rodriguez, Aythami, Morales, Julian, Fierrez
Abstract
Identity document (ID) authentication relies on the structural integrity of complex, high-frequency security patterns. However, advanced Generative AI models can now inject localized, high-fidelity manipulations, creating deceptive attacks that bypass standard verification. Training robust image forensic models to detect these anomalies is hindered by privacy regulations, forcing reliance on synthetic templates lacking the intricate visual patterns of real IDs. To bridge this domain gap, we introduce FakeIDet3-DB, the first comprehensive database of digital manipulations on real, government-issued IDs. FakeIDet3-DB encompasses classical (e.g., copy-move) and Generative AI-driven manipulations (e.g., face-swapping, inpainting) enhanced with advanced image refinement procedures to suppress visual artifacts. In addition, to comply with strict data protection regulations (e.g., GDPR), we adopt a recently-proposed framework based on patches. In order to maximize forensic utility, we formulate privacy-aware patch extraction from a real ID as a geometrically constrained image processing problem. We propose PACE, a Pseudo-Anonymized Contextual patch Extraction algorithm, which leverages Integral Image mapping and distance-driven Non-Maximum Suppression (NMS). PACE efficiently contours anonymization masks that prevent Personally Identifiable Information (PII) leakage while maximizing semantic density in peri-censorship regions, yielding almost 5.2M patches extracted from more than 6.4K images from real/fake IDs. Furthermore, an extensive evaluation of the proposed FakeIDet3-DB is performed using state-of-the-art models, showcasing they all struggle to detect and locate attacks coming from generative and classic techniques (32.45\% EER in detection and 83.48\% AUC-ROC in localization).
Chinese Translation
身份文件(ID)认证依赖于复杂的高频安全模式的结构完整性。然而,先进的生成式人工智能模型现在能够注入局部的高保真操控,制造出能够绕过标准验证的欺骗性攻击。训练强大的图像取证模型以检测这些异常受到隐私法规的限制,迫使我们依赖缺乏真实身份证复杂视觉模式的合成模板。为了解决这一领域差距,我们推出了FakeIDet3-DB,这是第一个全面的真实政府签发身份证上的数字操控数据库。FakeIDet3-DB涵盖了经典(例如,复制移动)和生成式人工智能驱动的操控(例如,面部交换、图像修复),并通过先进的图像精炼程序来抑制视觉伪影。此外,为了遵守严格的数据保护法规(例如,GDPR),我们采用了一种基于补丁的最新提出的框架。为了最大化取证效用,我们将从真实身份证中提取隐私感知补丁的过程公式化为一个几何约束的图像处理问题。我们提出了PACE,一种伪匿名上下文补丁提取算法,它利用积分图映射和基于距离的非极大值抑制(NMS)。PACE有效地勾勒出防止个人可识别信息(PII)泄露的匿名化掩码,同时在近审查区域最大化语义密度,从超过6400张真实/伪造身份证的图像中提取了近520万个补丁。此外,使用最先进的模型对所提出的FakeIDet3-DB进行了广泛评估,结果显示它们在检测和定位来自生成和经典技术的攻击时均表现不佳(检测的等错误率为32.45\%,定位的AUC-ROC为83.48\%)。
cs.CV / 42 / 2607.26645

FPSGen: Flexible Point Cloud Scene Generation with BEV-Supported Transport Flows

FPSGen:基于鸟瞰视图支持的运输流的灵活点云场景生成
He, Wenzhe, Wang, Meng, Qian, JiaWei, Xu, Jinfeng, Liu, Ying, Li, Ruihui
Abstract
Existing point-based generative methods for outdoor scenes primarily focus on LiDAR-conditioned completion. During training, noisy point clouds are constructed by perturbing complete ground-truth scenes, whereas during inference, they are initialized by adding noise to duplicated partial scans. This train-inference mismatch inherits the sparsity and visibility bias of partial scans, leading to sparse distant regions and incomplete geometry in occluded areas. Moreover, the reliance on partial scans restricts generation when LiDAR observations are unavailable or replaced by layout cues. We present FPSGen, a flexible framework that constructs point sources independently of partial scans. FPSGen first predicts a bird's-eye-view (BEV) prior with density, height, and mask channels from the active cues. The density map is then sampled to form a BEV-supported point source, enabling both unconditional and conditioned initialization. A teacher-student approximate optimal transport scheme then uses teacher-predicted endpoints to learn a velocity field that induces straighter transport paths. By integrating BEV point source construction with path-straightening transport, FPSGen provides a unified framework for unconditional and flexible cue-conditioned scene generation. Extensive experiments show that FPSGen achieves state-of-the-art JSD and voxel IoU performance on SemanticKITTI completion while maintaining strong performance with a single point transport step. On KITTI-360 unconditional generation, it also achieves the best Coverage (COV) among the compared methods.
Chinese Translation
现有的基于点的户外场景生成方法主要集中在基于激光雷达(LiDAR)条件的补全。在训练过程中,通过扰动完整的真实场景构建噪声点云,而在推理过程中,则通过向重复的部分扫描添加噪声来初始化。这种训练-推理不匹配继承了部分扫描的稀疏性和可见性偏差,导致远处区域稀疏以及遮挡区域几何形状不完整。此外,依赖部分扫描限制了在没有LiDAR观测或被布局线索替代时的生成能力。我们提出了FPSGen,一个独立于部分扫描构建点源的灵活框架。FPSGen首先从活跃线索中预测一个包含密度、高度和掩膜通道的鸟瞰视图(BEV)先验。然后对密度图进行采样,以形成一个支持BEV的点源,从而实现无条件和有条件的初始化。接着,教师-学生近似最优运输方案利用教师预测的端点学习一个速度场,以诱导更直的运输路径。通过将BEV点源构建与路径直线化运输相结合,FPSGen提供了一个统一的框架,用于无条件和灵活的线索条件场景生成。大量实验表明,FPSGen在SemanticKITTI补全任务上实现了最先进的JSD和体素IoU性能,同时在单点运输步骤下保持强劲表现。在KITTI-360无条件生成中,它在比较方法中也实现了最佳覆盖率(COV)。
cs.CV / 43 / 2607.26646

Genie Sim PanoWorld: An Infinite Indoor 3D World Generation Pipeline via Panoramic Scene Modeling and Simulation

Genie Sim PanoWorld:通过全景场景建模与仿真生成无限室内3D世界的管道
Su, Yongxin, Hou, Linjie, Wang, Feng, Tang, Jialin, Li, Zhijun, Wang, Qian, Yao, Maoqing
Abstract
We address the problem of reconstructing a high-fidelity, freely navigable 3D scene from a single $360^\circ$ panorama, without per-scene optimization or multi-view capture. Existing methods either lack metric trajectory control, which hinders reliable downstream 3D reconstruction, or struggle with large disocclusions under long-range camera motion while requiring high-end multi-GPU servers.We present Genie Sim PanoWorld, a two-stage feed-forward pipeline that bridges generation and reconstruction via an explicit, trajectory-controllable panoramic video. A NavMesh-planned $\mathrm{SE}(3)$ roaming trajectory is injected into a latent video diffusion model through dense geometry-warped conditioning; long--short trajectory mixed training and a self-consistency objective based on shortcut models together yield high-fidelity video in four CFG-free denoising steps. A feed-forward panoramic reconstructor then lifts the generated video into a high-fidelity 3D Gaussian scene that supports real-time, free-viewpoint roaming and can be directly used as a simulation-ready asset for embodied AI applications. Experiments show that Genie Sim PanoWorld outperforms geometry-conditioned baselines in both panoramic video generation and downstream 3D reconstruction, while generalizing zero-shot to unseen indoor scenes.
Chinese Translation
我们解决了从单个 $360^ ext{°}$ 全景图重建高保真、自由导航的3D场景的问题,而无需针对每个场景进行优化或多视角捕捉。现有方法要么缺乏度量轨迹控制,妨碍可靠的下游3D重建,要么在长距离相机运动下面临大范围遮挡问题,同时需要高端多GPU服务器。我们提出了Genie Sim PanoWorld,一种两阶段前馈管道,通过显式的、可控轨迹的全景视频连接生成与重建。一个NavMesh规划的 $ ext{SE}(3)$ 漫游轨迹通过密集几何扭曲条件注入到潜在视频扩散模型中;长短轨迹混合训练和基于快捷模型的自一致性目标共同在四个无CFG去噪步骤中生成高保真的视频。然后,一个前馈全景重建器将生成的视频提升为高保真的3D高斯场景,支持实时自由视点漫游,并可以直接用作适用于具身AI应用的仿真准备资产。实验表明,Genie Sim PanoWorld在全景视频生成和下游3D重建方面均优于几何条件基线,同时在零样本情况下对未见过的室内场景具有良好的泛化能力。
cs.CV / 44 / 2607.26647

Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time

锚定与引导扩散:在推理时增强文本到图像生成的真实性
Wang, Xinyi, Huang, Yuyang, Su, Yalin, Luan, Pengcheng, Zhang, Tao, Wei, Feiming, Yu, Wenxian
Abstract
While text-to-image diffusion models achieve impressive visual quality, they frequently struggle to maintain precise alignment with complex compositional prompts. An effective strategy is to improve the inference process of diffusion models, thereby better leveraging their pretrained priors to address misalignment. Existing training-free methods can be divided into two categories. The first category focuses on improving the randomly sampled initial noise, either performing costly search over noise pools or manipulating sampled noise without ensuring reliable semantic injection. The second category focuses on improving the denoising trajectory, lacking explicit mechanisms to timely diagnose and correct semantic errors. we propose \textbf{AnchorSteer}, a training-free framework that exerts fine-grained control over \textbf{both initialization} and \textbf{the denoising trajectory}. AnchorSteer consists of two synergistic components: \textbf{Semantic Anchoring} replaces uninformative Gaussian noise with text-aligned initializations via CLIP-based prior extraction and a novel Latent-Prior Score Distillation Sampling (LP-SDS) objective. Specifically, LP-SDS distills CLIP visual priors into the knowledge distribution of diffusion models, mitigating the domain gap between CLIP-based priors and diffusion-based priors. \textbf{Reflective Steering} transforms passive denoising with an active Think--Erase--Retouch loop that enables mid-generation self-correction. It leverages VLM-based diagnosis to detect semantic deviations and performs targeted latent refinement to suppress erroneous content and recover missing attributes. Extensive experiments on GenEval and T2I-CompBench++ demonstrate that AnchorSteer consistently outperforms existing baselines in text--image alignment while preserving high visual quality.
Chinese Translation
尽管文本到图像扩散模型在视觉质量上表现出色,但它们在与复杂的组合提示保持精确对齐方面常常面临挑战。一种有效的策略是改善扩散模型的推理过程,从而更好地利用其预训练的先验知识来解决不对齐问题。现有的无训练方法可以分为两类。第一类侧重于改善随机采样的初始噪声,要么在噪声池中进行代价高昂的搜索,要么在不确保可靠语义注入的情况下操纵采样噪声。第二类则关注改善去噪轨迹,但缺乏明确的机制来及时诊断和纠正语义错误。我们提出了 extbf{AnchorSteer},一个无训练框架,能够对 extbf{初始化}和 extbf{去噪轨迹}进行细粒度控制。AnchorSteer由两个协同组件组成: extbf{语义锚定}通过基于CLIP的先验提取和一种新颖的潜在先验评分蒸馏采样(Latent-Prior Score Distillation Sampling, LP-SDS)目标,将无信息的高斯噪声替换为与文本对齐的初始化。具体而言,LP-SDS将CLIP视觉先验蒸馏到扩散模型的知识分布中,减轻了基于CLIP的先验和基于扩散的先验之间的领域差距。 extbf{反思引导}则通过一个主动的思考-擦除-重修环路将被动去噪转变为主动去噪,使得在生成过程中能够自我纠正。它利用基于视觉语言模型(VLM)的诊断来检测语义偏差,并进行针对性的潜在细化,以抑制错误内容并恢复缺失属性。在GenEval和T2I-CompBench++上的大量实验表明,AnchorSteer在文本与图像对齐方面始终优于现有基线,同时保持高视觉质量。
cs.CV / 45 / 2607.26651

Physically Real-time Infrared Attack against Optical Flow Estimation Networks

针对光流估计网络的物理实时红外攻击
You, Shen, Jiang, Wei, Liu, Jiarui, Ye, Yijian, Lin, Qiuzhen, Li, Xiangtao, Wong, Ka-Chun
Abstract
With the promising performance of deep neural networks on image-based tasks, different real-world applications such as autonomous driving and motion detection have become increasingly mature and relevant to human lives. In particular, Optical Flow Estimation Networks (OFENs), as upstream models, play a critical role in different domains. Its outputs are heavily assumed and adopted for different downstream tasks, and it is essential to test its robustness to prevent safety accidents. We present an approach for real-time attacks on OFENs in the physical world, leveraging infrared lights for their stealthiness. By generating a large number of Adversarial Examples in advance, our approach computes AEs in real time and dynamically displays them, which allows our method to facilitate precise and targeted attacks without modifying the victim system. Unlike previous digital-to-physical attack techniques, our method directly attacks victim models within the physical world, thereby overcoming the limitations associated with the ineffectiveness of AEs. Experimental results demonstrate the efficacy of our approach in compromising OFENs across diverse lighting conditions, varying object motion velocities, and different object placements, ultimately impairing the network's ability to accurately estimate optical flow.
Chinese Translation
随着深度神经网络在基于图像任务中的出色表现,自动驾驶和运动检测等不同的现实应用变得愈加成熟,并与人类生活息息相关。特别是光流估计网络(Optical Flow Estimation Networks, OFENs)作为上游模型,在不同领域中发挥着关键作用。其输出被广泛假设和应用于不同的下游任务,因此测试其鲁棒性以防止安全事故至关重要。我们提出了一种针对OFENs的物理世界实时攻击方法,利用红外光的隐匿性。通过提前生成大量对抗样本(Adversarial Examples, AEs),我们的方法能够实时计算对抗样本并动态展示,从而实现精确和有针对性的攻击,而无需修改受害系统。与之前的数字到物理攻击技术不同,我们的方法直接在物理世界中攻击受害模型,从而克服了对抗样本无效性所带来的局限性。实验结果表明,我们的方法在不同光照条件、变化的物体运动速度和不同物体放置情况下有效地破坏了光流估计网络的性能,最终削弱了网络准确估计光流的能力。
cs.CV / 46 / 2607.26694

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation

Visko Orbis 1.0:一种实时交互长视频生成的直播模型
Gao, Xiangbo, Yang, Siyuan, He, Ping, Wu, Mingyang, Wu, Yuheng, Zuo, Yushen, Yu, Jiongze, Cui, Ryan, Hua, Hongyuan, Ma, Devin, Jin, Xiao, Yuan, Yubo, Yin, Qing, Yang, Jie, Tu, Zhengzhong
Abstract
We present Visko Orbis 1.0, a Live Model for real-time, interactive long-video generation. Users can change the prompt at any moment during generation, and the update becomes visible in real time. Visko Orbis 1.0 supports long-form text-to-video, image-to-video, and video continuation, with multilingual prompts and prompt switching while generation is in progress. A bounded multi-scale memory preserves subjects, scenes, and style across chunks, sustaining hour-scale rollouts without evident quality or color drift. Built on a distilled chunk-wise streaming generator and a streaming video upscaler, Visko Orbis 1.0 delivers real-time 4K video generation at 24 FPS using an optimized GPU serving engine. In long-form Arena comparisons, Visko Orbis 1.0 obtains the highest overall-preference and temporal-stability ratings among state-of-the-art real-time interactive video-generation systems.
Chinese Translation
我们提出了Visko Orbis 1.0,这是一种用于实时交互长视频生成的直播模型。用户可以在生成过程中随时更改提示,更新会实时可见。Visko Orbis 1.0支持长文本到视频、图像到视频以及视频续播,能够在生成过程中进行多语言提示和提示切换。一个有界的多尺度记忆在各个片段之间保留主题、场景和风格,能够在不明显质量或色彩漂移的情况下持续进行小时级的生成。Visko Orbis 1.0基于一种蒸馏的分块流式生成器和流式视频放大器,利用优化的GPU服务引擎实现24 FPS的实时4K视频生成。在长形式的Arena比较中,Visko Orbis 1.0在最先进的实时交互视频生成系统中获得了最高的整体偏好和时间稳定性评分。
cs.CV / 47 / 2607.26703

Sequence-SOD: Bio-inspired Sequence-aware Spiking ObjectDetection for Event Cameras

Sequence-SOD:基于生物启发的序列感知脉冲对象检测用于事件相机
Bendig, Katharina, Schuster, René, Stricker, Didier
Abstract
Event cameras follow a retina-inspired sensing principle, reporting local intensity changes asynchronously with hightemporal resolution and a wide dynamic range. Spiking Neural Networks (SNNs) complement these sparse event streams through brain-inspired dynamics, using sparse spikes and leaky membrane potentials to integrate information over time. However, many SNN object detectors process isolated event intervals with a single label and reset the network state after each prediction, thereby underusing temporal information in continuous event streams. We introduce Sequence-SOD, a sequence-aware SNN object detector that processes extended event sequences containing labels at multiple time points. Events are accumulated into short intervals, discretized into temporal steps, and fed sequentially to an SSD-style Spiking DenseNet while preserving membrane potentials across intervals within a sequence, so that detection is driven by an evolving neural state instead of independently reset input windows. On the Gen1 Automotive Detection Dataset, sequence-aware training improves mAP from 23.38 for single-interval training to 25.30 without augmentation and to 26.88 withevent-data augmentation. The model achieves a theoretical prediction frequency of 40 Hz. Training and evaluating SNN object detectors on extended event sequences improves their ability to exploit temporal cues while preserving the energy-efficiency benefits of sparse spiking computation. The results highlight sequence-aware training as a complementary direction to architectural improvements for event-based SNN detection.
Chinese Translation
事件相机遵循类视网膜的感知原理,以高时间分辨率和宽动态范围异步报告局部强度变化。脉冲神经网络(SNNs)通过类脑动态补充这些稀疏事件流,利用稀疏脉冲和泄漏膜电位在时间上整合信息。然而,许多SNN对象检测器处理带有单一标签的孤立事件间隔,并在每次预测后重置网络状态,从而未能充分利用连续事件流中的时间信息。我们提出了Sequence-SOD,这是一种序列感知的SNN对象检测器,处理包含多个时间点标签的扩展事件序列。事件被累积到短间隔中,离散化为时间步,并顺序输入到SSD风格的Spiking DenseNet中,同时在序列内保持膜电位跨间隔的连续性,从而使检测由不断演变的神经状态驱动,而不是独立重置的输入窗口。在Gen1汽车检测数据集上,序列感知训练将mAP从单间隔训练的23.38提高到未增强情况下的25.30,以及在事件数据增强情况下的26.88。该模型实现了40 Hz的理论预测频率。在扩展事件序列上训练和评估SNN对象检测器提高了它们利用时间线索的能力,同时保持了稀疏脉冲计算的能效优势。结果强调了序列感知训练作为事件驱动的SNN检测架构改进的补充方向。
cs.CV / 48 / 2607.26706

TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models

TPD:文本到视频扩散模型的时间先验解耦
Kang, Taewon, Zwicker, Matthias
Abstract
Text-to-video diffusion models generate temporally coherent content from natural language, yet when a prompt describes an early scene that persists while a new event emerges on top of it---such as "a tall sandcastle standing on a beach where a wave rushes in and washes it away"---generation frequently fails to realize the late-segment event in the corresponding frames. We identify this failure as Temporal Prior Suppression (TPS): the dominant prior of the early segment captures the cross-attention trajectory across the temporal axis and suppresses the guidance signal needed for late-segment realization, a competing tendency existing guidance mechanisms do not model. We introduce Temporal Prior Decoupling (TPD), a training-free framework that restores suppressed late-segment signals during diffusion sampling. TPD constructs a temporal counterfactual by conditioning on the early segment alone, and defines the discrepancy between the full-prompt and counterfactual trajectories as a suppressed signal direction. Rather than removing this direction as in prior subtractive projection methods, TPD restores it through a frame-selective lower-bound constraint resolved jointly over diffusion timestep and video frame, realizing the suppressed event in the late frames without disrupting early-segment coherence: where prior work enforces upper-bound feasibility to remove unwanted semantics, TPD enforces lower-bound feasibility to guarantee suppressed-signal contribution. TPD runs entirely within standard diffusion sampling without retraining, and is defined purely in classifier-free guidance space, making it backbone-agnostic by construction. Experiments show that TPD significantly improves late-concept realization while preserving temporal coherence and visual fidelity, and that the targeted suppression recurs across distinct text-to-video backbones.
Chinese Translation
文本到视频扩散模型能够从自然语言生成时间上连贯的内容,但当提示描述一个早期场景并在其上出现新事件时——例如“一个高大的沙堡矗立在海滩上,海浪涌来将其冲走”——生成常常无法在相应帧中实现晚段事件。我们将这种失败称为时间先验抑制(Temporal Prior Suppression, TPS):早期段的主导先验捕捉了跨时间轴的交叉注意力轨迹,并抑制了实现晚段所需的引导信号,而现有的引导机制并未建模这种竞争倾向。我们提出了时间先验解耦(Temporal Prior Decoupling, TPD),这是一种无训练的框架,在扩散采样过程中恢复被抑制的晚段信号。TPD通过仅对早期段进行条件化来构建一个时间反事实,并将完整提示与反事实轨迹之间的差异定义为被抑制的信号方向。与以往的减法投影方法不同,TPD通过在扩散时间步和视频帧上共同解决的帧选择下界约束来恢复该方向,从而在不破坏早期段连贯性的情况下,在晚帧中实现被抑制的事件:而之前的工作通过施加上界可行性来去除不必要的语义,TPD则施加下界可行性以保证被抑制信号的贡献。TPD完全在标准扩散采样中运行,无需重新训练,并且纯粹在无分类器引导空间中定义,使其在结构上与主干无关。实验表明,TPD显著提高了晚概念的实现,同时保持了时间连贯性和视觉保真度,并且目标抑制在不同的文本到视频主干中反复出现。
cs.CV / 49 / 2607.26715

FreeShadow: Training-Free Shadow Removal via Illumination Transfer and Selective Content Preservation in Diffusion Models

FreeShadow:通过照明转移和选择性内容保留在扩散模型中实现无训练阴影去除
Wang, Yinan, Huang, Yan, Xu, Yong, Callet, Patrick Le
Abstract
Existing supervised and unsupervised shadow removal methods often suffer from limited generalization due to the insufficient diversity of available training datasets, while zero-shot methods tend to produce artifacts and require time-consuming test-time optimization. To address these issues, we propose FreeShadow, a training-free shadow removal method built upon pretrained diffusion models, which exploits diffusion priors for shadow removal without any training or optimization. For illumination recovery, we propose an illumination transfer attention (ITA), which re-weights the self-attention maps in diffusion model to transfer illumination cues from non-shadow to shadow regions. For content preservation, we analyze the effects of illumination variations on self-attention maps and latent high-frequency features in diffusion model, and selectively preserve illumination-invariant components to maintain content fidelity while suppressing residual shadows. We further propose local texture-preserving relighting (LTPR) to mitigate local texture misalignment caused by VAE compression. Extensive experiments demonstrate that our method achieves strong generalization and produces realistic shadow-free images.
Chinese Translation
现有的监督和无监督阴影去除方法常常由于可用训练数据集的多样性不足而面临有限的泛化能力,而零-shot 方法往往会产生伪影,并且需要耗时的测试时优化。为了解决这些问题,我们提出了 FreeShadow,这是一种基于预训练扩散模型的无训练阴影去除方法,它利用扩散先验进行阴影去除,无需任何训练或优化。为了实现照明恢复,我们提出了一种照明转移注意力(Illumination Transfer Attention, ITA),该方法重新加权扩散模型中的自注意力图,以将照明线索从非阴影区域转移到阴影区域。为了保持内容的完整性,我们分析了照明变化对扩散模型中自注意力图和潜在高频特征的影响,并选择性地保留照明不变的成分,以维持内容的保真度,同时抑制残余阴影。我们进一步提出了局部纹理保留重光照(Local Texture-Preserving Relighting, LTPR),以减轻由变分自编码器压缩引起的局部纹理错位。大量实验表明,我们的方法实现了强泛化能力,并生成了逼真的无阴影图像。
cs.CV / 50 / 2607.26729

CASIAL: Geometric Distortion Robust Image Watermarking

CASIAL:几何失真鲁棒图像水印
Qiu, Yupeng, Fang, Han, Chang, Ee-Chien
Abstract
Deep learning-based watermarking has shown strong robustness against non-geometric distortions, yet its performance under geometric transformations remains limited. Such transformations induce two fundamental failure modes: region removal, such as cropping or masking, which eliminates the information carried by removed pixels, and desynchronization, such as scaling or rotation, which misaligns pixel positions and disrupts decoding. We argue that achieving geometric robustness requires two essential properties: (1) global spread of the watermark message, ensuring resilience even when large regions are removed, and (2) geometry-invariant representations, enabling decoding to remain synchronized despite spatial transformations. Building on these insights, we propose CASIAL, a geometric distortion-robust watermarking framework with cover image-aware message spreading (CAS) strategy and invariance alignment learning (IAL) module. CAS tightly couples watermark bits with cover image features and distributes them adaptively across the entire image, enhancing per-pixel information capacity and robustness to region removal. IAL leverages spatial attention to capture cross-pixel dependencies and align perturbed features into a shared geometry-invariant representation space, mitigating failures due to desynchronization. Across six challenging geometric transformations, CASIAL achieves substantially stronger robustness than eleven prior baselines while preserving high visual quality. It also maintains competitive performance under six signal distortions and four photometric transformations. Notably, although trained only with white-box distortions, CASIAL also exhibits strong transfer robustness to unseen black-box distortions. Comprehensive experiments demonstrate the broad robustness and superior visual quality of our method.
Chinese Translation
基于深度学习的水印技术在非几何失真方面表现出强大的鲁棒性,但在几何变换下的性能仍然有限。这些变换引发了两种基本的失效模式:区域移除,例如裁剪或遮挡,消除了被移除像素所携带的信息;以及不同步,例如缩放或旋转,导致像素位置错位并干扰解码。我们认为,实现几何鲁棒性需要两个基本特性:(1)水印信息的全局传播,确保即使在大区域被移除时也能保持韧性;(2)几何不变表示,使得解码在空间变换下仍能保持同步。基于这些见解,我们提出了CASIAL,一个具有几何失真鲁棒性的水印框架,采用了覆盖图像感知信息传播(CAS)策略和不变性对齐学习(IAL)模块。CAS将水印位与覆盖图像特征紧密结合,并自适应地分布在整个图像中,增强了每个像素的信息容量和对区域移除的鲁棒性。IAL利用空间注意力捕捉跨像素依赖关系,并将扰动特征对齐到共享的几何不变表示空间,从而减轻由于不同步导致的失效。在六种具有挑战性的几何变换下,CASIAL的鲁棒性显著强于十一种先前基线,同时保持高视觉质量。它在六种信号失真和四种光度变换下也保持了竞争力的性能。值得注意的是,尽管仅使用白盒失真进行训练,CASIAL在未见的黑盒失真下也表现出强大的迁移鲁棒性。全面的实验表明我们的方法具有广泛的鲁棒性和优越的视觉质量。
cs.CV / 51 / 2607.26733

Online Handwriting Trajectory Reconstruction from Kinematic Sensors using Temporal Convolutional Network

基于运动传感器的在线手写轨迹重建:采用时间卷积网络
Swaileh, Wassim, Imbert, Florent, Soullard, Yann, Tavenard, Romain, Anquetil, Eric
Abstract
Handwriting with digital pens is a common way to facilitate human-computer interaction through the use of Online Handwriting (OH) trajectory reconstruction. In this work, we focus on a digital pen equipped with sensors from which one wants to reconstruct the OH trajectory. Such a pen allows to write on any surface and to get the digital trace, which can help learning to write, by writing on paper, and can be useful for many other applications such as collaborative meetings, etc. In this paper, we introduce a novel processing pipeline that maps the sensor signals of the pen to the corresponding OH trajectory. Notably, in order to tackle the difference of sampling rates between the pen and the tablet (which provides ground truth information), our preprocessing pipeline relies on Dynamic Time Warping to align the signals. We introduce a dedicated neural network architecture, inspired by a Temporal Convolutional Network, to reconstruct the online trajectory from the pen sensor signals. Finally, we also present a new benchmark dataset on which our method is evaluated both qualitatively and quantitatively, showing a notable improvement over its most notable competitor.
Chinese Translation
使用数字笔进行手写是一种通过在线手写(OH)轨迹重建来促进人机交互的常见方式。在本研究中,我们关注于一种配备传感器的数字笔,旨在重建OH轨迹。这种笔可以在任何表面上书写并获取数字痕迹,这不仅有助于学习书写(例如在纸上书写),还可以用于许多其他应用,如协作会议等。本文介绍了一种新颖的处理流程,将笔的传感器信号映射到相应的OH轨迹。值得注意的是,为了应对笔与平板(提供真实信息)之间采样率的差异,我们的预处理流程依赖于动态时间规整(Dynamic Time Warping)来对齐信号。我们提出了一种专门的神经网络架构,灵感来源于时间卷积网络(Temporal Convolutional Network),用于从笔的传感器信号中重建在线轨迹。最后,我们还展示了一个新的基准数据集,在该数据集上对我们的方法进行了定性和定量评估,显示出相较于其最显著竞争对手的显著改进。
cs.CV / 52 / 2607.26735

Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives

文本到图像扩散模型的双重反演:从提示和噪声的双重视角
Liu, Xiaolong, Li, Junjian, Xiao, Yuan, Deng, Jiaqi, Ye, Dayong, Zhu, Tianqing, Huo, Huan
Abstract
Prompt inversion, as a typical reverse engineering technique, enables text-to-image (T2I) diffusion models to generate the desired target images without extensive prompt engineering. However, existing prompt inversion methods suffer from significant limitations: (1) gradient-based methods are unstable and uninterpretable, often resulting in generated images with severe artifacts; (2) gradient-free methods yield human-readable prompts but still fail to preserve visual fidelity due to the lack of fine-grained detail alignment. We contend that the limitations stem from treating prompt inversion as a sufficient condition for reverse engineering, ignoring the critical role of the latent noise that encodes structural information. Consequently, we propose Dualin (Dual inversion), a two-stage method that jointly recovers both the semantic prompt and latent noise of the target image. In the first stage, we integrate vision-language model, CLIP and large language model to invert a faithful, human-interpretable hard prompt. In the second stage, unconditional DDIM inversion reconstructs the exact latent noise of the target image, guaranteeing the consistency at the structural information level. Theoretically, we prove that the inverted noise enables flexible image editing without re-optimization. Extensive experiments on diverse datasets demonstrate that Dualin simultaneously generates high-quality inverted prompts and achieves state-of-the-art image fidelity. Additionally, Dualin can establish a robust foundation for the precise and controllable image editing.
Chinese Translation
提示反演作为一种典型的逆向工程技术,使得文本到图像(T2I)扩散模型能够在无需大量提示工程的情况下生成所需的目标图像。然而,现有的提示反演方法存在显著的局限性:(1)基于梯度的方法不稳定且难以解释,常常导致生成的图像出现严重的伪影;(2)无梯度的方法虽然能够生成可读的人类提示,但由于缺乏细粒度的细节对齐,仍然无法保持视觉保真度。我们认为,这些局限性源于将提示反演视为逆向工程的充分条件,而忽视了编码结构信息的潜在噪声的关键作用。因此,我们提出了Dualin(双重反演),一种两阶段的方法,联合恢复目标图像的语义提示和潜在噪声。在第一阶段,我们整合了视觉-语言模型CLIP和大型语言模型,以反演出一个真实的、可被人类理解的硬提示。在第二阶段,无条件的DDIM反演重建了目标图像的确切潜在噪声,确保了结构信息层面的连贯性。从理论上讲,我们证明了反演的噪声能够实现灵活的图像编辑而无需重新优化。在多样化数据集上的大量实验表明,Dualin能够同时生成高质量的反演提示,并实现最先进的图像保真度。此外,Dualin还可以为精确和可控的图像编辑奠定坚实的基础。
cs.CV / 53 / 2607.26743

Multimodal fusion of visual and morphometric features for avian bone classification

视觉与形态特征的多模态融合用于鸟类骨骼分类
Dubbini, Nevio, Yeomans, Lisa, Pavia, Marco, Parmaksiz, Ramazan, Hooglugt, Ayse Atas, Gattiglia, Gabriele, Demarchi, Beatrice
Abstract
Artificial intelligence has shown considerable potential for archaeological applications, yet its use in zooarchaeology remains limited, particularly for the identification of avian skeletal remains. This study presents a proof-of-concept multimodal framework that integrates convolutional neural network-based image analysis with osteometric measurements for the classification of bird bones. Using a dataset of more than 10,000 images from multiple museum and research collections, two classification tasks were investigated: skeletal element identification and family-level taxonomic classification. Prior to classification, images were automatically segmented using a two-stage pipeline combining BiRefNet and SAM2. Visual features extracted with a pre-trained EfficientNet_V2_S backbone were fused with standardized morphometric data through a feature-level multimodal architecture. The model achieved 86% accuracy on the test set for bone-type classification, demonstrating reliable recognition of skeletal elements. Family-level classification proved more challenging, reaching 51% top-1 accuracy but 75% top-3 accuracy, indicating that correct taxa were frequently included among the most probable predictions. These results demonstrate the feasibility of combining visual and morphometric information within a unified deep-learning framework and establish a methodological baseline for future AI-assisted zooarchaeological identification. The approach contributes to ongoing efforts to develop scalable, interpretable, and archaeologically meaningful tools for the study of avian remains.
Chinese Translation
人工智能在考古学应用中展现出相当大的潜力,但其在动物考古学中的应用仍然有限,尤其是在鸟类骨骼遗骸的识别方面。本研究提出了一种概念验证的多模态框架,结合了基于卷积神经网络的图像分析与骨测量数据,用于鸟类骨骼的分类。使用来自多个博物馆和研究收藏的超过10,000张图像的数据集,研究了两个分类任务:骨骼元素识别和科级分类。在分类之前,图像通过结合BiRefNet和SAM2的两阶段管道进行自动分割。使用预训练的EfficientNet_V2_S骨干网络提取的视觉特征与标准化的形态数据通过特征级多模态架构进行了融合。该模型在骨骼类型分类的测试集上达到了86%的准确率,展示了对骨骼元素的可靠识别。科级分类则更具挑战性,达到了51%的top-1准确率和75%的top-3准确率,表明正确的分类群体常常出现在最可能的预测中。这些结果展示了在统一的深度学习框架内结合视觉和形态信息的可行性,并为未来的人工智能辅助动物考古学识别建立了方法学基准。这一方法有助于开发可扩展、可解释且具有考古意义的工具,以研究鸟类遗骸。
cs.CV / 54 / 2607.26754

StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation

StatePlay:基于状态的游戏世界模型用于机制一致的生成
Lin, Zijun, Wang, Zeqing, Tan, Cheston, Wen, Bihan, Jin, Yeying
Abstract
Recent game world models can generate visually realistic and interactive environments conditioned on player actions. However, games are not defined by pixels alone; they are governed by explicit mechanics, namely state-dependent rules that control health reduction, skill activation, and game termination. These mechanics depend on precise internal states, such as health points, skill meters, and timers, which are tightly coupled with visual observations and determine how gameplay evolves. Without modeling these state dynamics, existing game world models may generate visually plausible rollouts but violate the underlying game rules. In this paper, we propose StatePlay, a novel state-aware game world model that jointly predicts visual content and game states to promote mechanics-consistent generation. StatePlay adopts a mixture-of-transformers (MoT)-style architecture that preserves specialized visual and state representations while enabling cross-modal interaction, allowing predicted states to guide frame generation. Each branch is further optimized with a distinct objective suited to its modality. Experiments show that StatePlay achieves an average normalized L1 distance below 0.06 for state prediction. Furthermore, compared with models without explicit state modeling, our method improves mechanics fidelity in generated game rollouts by 18.6%. Overall, our work highlights the importance of state-aware game world modeling and advances beyond pixel-level realism toward complete and mechanically faithful game generation.
Chinese Translation
近期的游戏世界模型能够生成基于玩家行为的视觉真实和互动环境。然而,游戏不仅仅由像素定义;它们受明确机制的支配,即控制生命值减少、技能激活和游戏终止的状态依赖规则。这些机制依赖于精确的内部状态,如生命值、技能计量器和计时器,这些状态与视觉观察紧密耦合,并决定了游戏玩法的演变。如果不对这些状态动态进行建模,现有的游戏世界模型可能会生成视觉上合理的结果,但会违反潜在的游戏规则。本文提出了StatePlay,一种新颖的状态感知游戏世界模型,该模型共同预测视觉内容和游戏状态,以促进机制一致的生成。StatePlay采用混合变换器(Mixture-of-Transformers, MoT)风格的架构,保留了专业的视觉和状态表示,同时实现跨模态交互,使得预测的状态能够指导帧生成。每个分支还通过适合其模态的独特目标进行进一步优化。实验表明,StatePlay在状态预测方面实现了平均归一化L1距离低于0.06。此外,与没有明确状态建模的模型相比,我们的方法在生成的游戏结果中提高了18.6%的机制保真度。总体而言,我们的工作强调了状态感知游戏世界建模的重要性,并在像素级真实感的基础上推进了完整且机制忠实的游戏生成。
cs.CV / 55 / 2607.26763

Long-Tailed 3D Point Cloud Dataset Distillation

长尾3D点云数据集蒸馏
You, Jiahao, Han, Xu, Xu, Jinfeng, Li, Xianzhi
Abstract
Dataset distillation compresses large-scale datasets into compact synthetic sets while preserving their training utility, enabling efficient 3D point cloud training. Current point cloud dataset distillation methods only tackle geometric and representation challenges while ignoring the distributional imbalance prevalent in point cloud datasets where both training and test splits follow long-tailed class distributions. To our knowledge, we present the first study on long-tailed point cloud dataset distillation. Rather than focusing primarily on geometric and representation properties or simply constructing a class-balanced synthetic set, our framework explicitly accounts for long-tailed class distributions via two core modules. First, we design Adaptive Synthetic Budgeting to allocate class-wise synthetic budgets according to class quantity and the expected benefit of additional synthetic samples. Given the allocated budgets, we further design 3D Long-Tailed Distribution Matching to optimize synthetic point clouds through Global-Local Feature Alignment and Prior-Aware Supervision. The former preserves both global class distributions and diverse intra-class structures, while the latter provides class-dependent expert supervision to keep tail-class samples recognizable while maintaining diverse head-class patterns. Extensive experiments demonstrate the effectiveness of our method, lifting classification accuracy by 7.0 points on ShapeNet55 against state-of-the-art methods.
Chinese Translation
数据集蒸馏将大规模数据集压缩为紧凑的合成集,同时保留其训练效用,从而实现高效的3D点云训练。目前的点云数据集蒸馏方法仅解决几何和表征挑战,而忽略了点云数据集中普遍存在的分布不平衡问题,其中训练和测试分割均遵循长尾类别分布。据我们所知,我们首次研究了长尾点云数据集蒸馏。我们的框架不仅关注几何和表征属性或简单构建类别平衡的合成集,而是通过两个核心模块明确考虑长尾类别分布。首先,我们设计了自适应合成预算(Adaptive Synthetic Budgeting),根据类别数量和额外合成样本的预期收益分配类别合成预算。在分配预算的基础上,我们进一步设计了3D长尾分布匹配(3D Long-Tailed Distribution Matching),通过全局-局部特征对齐(Global-Local Feature Alignment)和先验感知监督(Prior-Aware Supervision)来优化合成点云。前者保留了全局类别分布和多样的类内结构,而后者提供了类别依赖的专家监督,以保持尾类样本的可识别性,同时维持多样的头类模式。大量实验表明我们的方法有效性,在ShapeNet55上相较于最先进的方法提高了7.0个百分点的分类准确率。
cs.CV / 56 / 2607.26765

Searching for Robust Augmentations to Improve Out-of-Domain Generalization in Dermoscopic Skin Cancer Classification

寻找稳健的增强方法以改善皮肤癌分类中的域外泛化能力
Kozachok, Alexander, Latyshev, Ilya, Karpulevich, Evgeny, Kozachok, Elena, Ushakov, Egor, Samovarov, Oleg
Abstract
Background/Objectives: Dermoscopic skin lesion classifiers often lose accuracy under domain shift across imaging devices, illumination, and capture artifacts. We study how data augmentation improves the robustness of a binary malignant-versus-non-malignant classifier, with emphasis on out-of-domain (OOD) generalization. Methods: Single augmentations, photometric combinations, and composite policies were searched on a multi-source ISIC Archive collection with Derm7pt, using a ConvNeXt-Large backbone and ROC-AUC. Splits were made at the lesion-ID level, and HAM10000 and ISIC 2019-2020 were held out as a predominantly source-disjoint OOD test. Results: The largest OOD gain came from the mix policy, and photometric transformations dominated the most useful OOD operations. On an expanded pool from the same held-out sources the gain was +0.053 (95% CI +0.045 to +0.061, p<0.001), consistent across four training seeds (per-seed ROC-AUC: baseline 0.761-0.775, mix 0.806-0.829). On a small independent clinical collection, single-checkpoint sensitivity rose from 0.591 to 0.818, but this rested on 22 malignant cases and did not persist across seeds. Conclusions: Augmentations modelling real sources of domain shift can matter more than maximizing in-domain accuracy. Because the policy was selected on the same sources used to evaluate it, a source-disjoint selection protocol is needed before this effect size can be read as unbiased.
Chinese Translation
背景/目标:皮肤镜皮损分类器在成像设备、照明和捕获伪影等领域转移下通常会失去准确性。我们研究了数据增强如何提高二分类恶性与非恶性分类器的稳健性,重点关注域外(OOD)泛化。方法:在使用ConvNeXt-Large主干网络和ROC-AUC的多源ISIC档案集合Derm7pt上,搜索了单一增强、光度组合和复合策略。按照病变ID级别进行划分,HAM10000和ISIC 2019-2020作为主要源不重叠的OOD测试集被保留。结果:最大的OOD增益来自混合策略,光度变换主导了最有用的OOD操作。在同一保留源的扩展池中,增益为+0.053(95% CI +0.045至+0.061,p<0.001),在四个训练种子中保持一致(每个种子的ROC-AUC:基线0.761-0.775,混合0.806-0.829)。在一个小型独立临床集合中,单一检查点的灵敏度从0.591上升至0.818,但这基于22个恶性病例,并未在不同种子间持续。结论:模拟真实领域转移来源的增强方法可能比最大化域内准确性更为重要。由于该策略是在用于评估的相同源上选择的,因此在读取此效应大小为无偏之前,需要一个源不重叠的选择协议。
cs.CV / 57 / 2607.26769

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

See2Think:多模态模型是否真正利用中间视觉状态?
Yan, Siyu, Yan, Zhuoran, Xu, Haiying, Zhou, Panhao, Chen, Jingyu, Ji, Chenhao, Cao, Shuo, Zhang, Yongheng, Liu, Haoze, Zhang, Siyu, Gu, Xiwen, Liu, Yihao, Wang, Alex Jinpeng
Abstract
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.
Chinese Translation
多模态大型语言模型在推理过程中越来越多地使用草图、注释、工具和中间图像,但尚不清楚它们是否真正依赖于这些视觉状态。现有的基准测试受到任务集合覆盖范围狭窄或部分可通过文本解决的样本的限制,以及强调最终答案而不诊断中间视觉状态如何生成、呈现和使用的评估的限制。我们引入了See2Think,一个统一的评估框架,包括See2ThinkBench和视觉思维行为(Visual Action-of-Thought, VAoT)。See2ThinkBench包含1200个开放式、视觉依赖的问题,涵盖12个任务类别,涉及2D结构、3D场景和现实世界推理。VAoT在四种受控推理设置下记录文本思维、视觉行为、呈现状态和后续推理。通过评估代表性的专有和开源多模态模型,我们发现视觉推理在很大程度上依赖于模型和环境,没有单一设置在各任务中始终占据主导地位。过程分析进一步表明,模型通常选择相关的视觉操作,而忠实的呈现仍然是最明显的瓶颈,高反馈采纳并不一定转化为准确性提升。在与任务相关的干扰反馈下,模型表现出对视觉状态的行为依赖,准确率在受控干预中下降超过10个百分点。
cs.CV / 58 / 2607.26799

PRISM-Net: Patient-specific reference-guided inter-breast symmetry matching for three-class breast DCE-MRI classification

PRISM-Net:基于患者特异性参考的双侧乳腺对称匹配用于三类乳腺 DCE-MRI 分类
Zhang, Boya, Zhou, Shuaiwen, Kong, Di, Wang, Mingxu, Du, Wenbiao, Zhong, Yiman, Duan, Yuexin, Yue, Xiawei, Cheng, Liuquan, Li, Xiru
Abstract
Breast DCE-MRI AI is increasingly being explored for breast-level classification of no-lesion, benign, and malignant findings, beyond conventional lesion-centered diagnosis. Within this broader diagnostic scope, however, patient-specific background variability remains a major source of imaging confounding across classification tasks. Existing approaches predominantly focus on unilateral or lesion-centric analysis, whereas bilateral methods offer limited explicit modeling of spatially adaptive cross-breast correspondence. We propose PRISM-Net, a registration-free bilateral framework that leverages contralateral breast features as patient-specific references for background-aware representation learning. PRISM-Net integrates bilateral feature matching and asymmetry-aware attention to establish adaptive inter-breast correspondence and enhance representations of discriminative asymmetric patterns. On ODELIA, Macro AUC, Micro AUC, and quadratic weighted kappa were $84.11 \pm 2.33$, $90.64 \pm 1.61$, and $60.94 \pm 5.64$ on the in-distribution test set, and $68.51 \pm 4.54$, $80.74 \pm 2.68$, and $43.45 \pm 7.10$ on the held-out institution, respectively, outperforming the evaluated baseline methods across the primary evaluation metrics. PRISM-Net further demonstrated performance on independent institutional and background-complexity evaluations. Ablation experiments revealed that both bilateral relation modeling and asymmetry-aware reweighting contributed to improved classification performance. These findings highlight patient-specific bilateral reference modeling as a clinically grounded strategy for DCE-MRI interpretation, improving asymmetric pattern discrimination through explicit modeling of background complexity.
Chinese Translation
乳腺 DCE-MRI 人工智能正日益被探索用于乳腺层面的无病灶、良性和恶性发现的分类,超越传统的病灶中心诊断。然而,在这一更广泛的诊断范围内,患者特异性的背景变异仍然是分类任务中影像混淆的主要来源。现有方法主要集中于单侧或病灶中心分析,而双侧方法在空间自适应的跨乳腺对应建模方面则显得有限。我们提出了 PRISM-Net,一种无注册的双侧框架,利用对侧乳腺特征作为患者特异性参考,以进行背景感知的表示学习。PRISM-Net 集成了双侧特征匹配和对称性感知注意力,以建立自适应的跨乳腺对应关系,并增强对区分性不对称模式的表示。在 ODELIA 数据集上,宏观 AUC、微观 AUC 和二次加权 kappa 在分布内测试集上的值分别为 $84.11 imes 2.33$、$90.64 imes 1.61$ 和 $60.94 imes 5.64$,而在保留机构上的值分别为 $68.51 imes 4.54$、$80.74 imes 2.68$ 和 $43.45 imes 7.10$,在主要评估指标上均优于评估的基线方法。PRISM-Net 还在独立机构和背景复杂性评估中展示了性能。消融实验表明,双侧关系建模和对称性感知重加权均有助于提高分类性能。这些发现强调了患者特异性双侧参考建模作为 DCE-MRI 解释的临床基础策略,通过对背景复杂性的显式建模改善不对称模式的区分。
cs.CV / 59 / 2607.26811

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

DistillAlign:自回归视频蒸馏中的模式覆盖与模式寻求协调
Li, Jiaxing, Zou, Kai, Zhou, Cindy, Huang, Kaichen, Gao, Junyao, Wang, Zile, Liu, Yang, Liu, Bin, An, Bo, Li, Yangguang
Abstract
Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they typically decouple the initialization and DMD stages -- which then pursue different target distributions -- and judge the intermediate student mainly by visual scores such as VBench. In this paper, we revisit this design from a distributional perspective. Given the mode-seeking nature of the distribution matching loss, a good initialization should match the mode coverage of the target DMD teacher, rather than merely pursuing high quality. To analyze this, we introduce a distributional evaluation protocol that measures precision and coverage between student and teacher distributions in a shared latent space. It exposes differences hidden by visual scores: some initializations reach high precision but low coverage, leading to suboptimal refinement, while mode-covering ones preserve broader support. Furthermore, even when the target distributions are aligned, DMD's reverse-KL objective can still drive the student toward high-probability teacher regions in late training, reducing coverage and diversity. To address this, we propose joint distillation, which combines DMD's mode-seeking objective with a Consistency Distillation-based mode-covering constraint. Experiments show that our method improves generation quality, coverage, and diversity; notably, even with a Wan-1.3B DMD teacher, it outperforms baselines refined with Wan-14B, underscoring the importance of distributional alignment in autoregressive video distillation.
Chinese Translation
现有的自回归视频蒸馏方法通常采用基于分布匹配蒸馏(Distribution Matching Distillation, DMD)的多阶段流程。然而,它们通常将初始化阶段与DMD阶段解耦,这两个阶段追求不同的目标分布,并主要通过视觉评分(如VBench)来评估中间学生。在本文中,我们从分布的角度重新审视这一设计。考虑到分布匹配损失的模式寻求特性,良好的初始化应当匹配目标DMD教师的模式覆盖,而不仅仅是追求高质量。为此,我们引入了一种分布评估协议,用于测量共享潜在空间中学生与教师分布之间的精度和覆盖度。该协议揭示了视觉评分所隐藏的差异:一些初始化达到高精度但低覆盖,导致次优的细化,而模式覆盖的初始化则保留了更广泛的支持。此外,即使目标分布对齐,DMD的反向KL目标仍然可能在后期训练中将学生驱动到高概率教师区域,从而减少覆盖和多样性。为了解决这一问题,我们提出了联合蒸馏,将DMD的模式寻求目标与基于一致性蒸馏(Consistency Distillation)的模式覆盖约束相结合。实验表明,我们的方法提高了生成质量、覆盖度和多样性;值得注意的是,即使使用Wan-1.3B DMD教师,它也优于使用Wan-14B细化的基线,强调了自回归视频蒸馏中分布对齐的重要性。
cs.CV / 60 / 2607.26818

Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory

Ripple:具有跨模态递归记忆的实时音视频生成
Ding, Yanbo, Guo, Zhizhi, Song, Quanyue, He, Yishan, He, Zhixiang, Li, Yongxiang, Wang, Yali
Abstract
Audio-video generative models achieve impressive quality but suffer from high latency, making them unsuitable for real-time applications. Although several streaming audio-video generation methods have been proposed, they remain costly and fail to support long-form generation. To address this, we propose \textbf{Ripple}, a real-time joint audio-video generation system with a cross-modal recurrent memory mechanism. To enable efficient streaming inference while preserving long-term context, Ripple combines a fixed-length sliding-window attention with modality-specific memory states that continuously summarize audio and video context. Cross-modal memory interaction is further introduced to enhance audio-visual synchronization. To learn this memory-augmented model effectively, we devise a three-stage training recipe: (1) adapting a bidirectional audio-video teacher to block-wise causal attention with simulated memory, (2) optimizing the memory construction and interaction pipeline through end-to-end distillation, and (3) applying online reinforcement post-training tailored for streaming audio-video generation. As a result, Ripple achieves ~28 FPS at 480P resolution, over faster than the teacher, while capable of coherent long-form generation. Extensive experiments on both short-video and long-video benchmarks demonstrate our superior performance over existing offline and online joint audio-video generation methods.
Chinese Translation
音视频生成模型在质量上取得了令人印象深刻的成果,但由于高延迟,使其不适合实时应用。尽管已经提出了几种流式音视频生成方法,但它们仍然成本高昂,并且无法支持长格式生成。为了解决这个问题,我们提出了 extbf{Ripple},一种具有跨模态递归记忆机制的实时联合音视频生成系统。为了在保持长期上下文的同时实现高效的流式推理,Ripple结合了固定长度的滑动窗口注意力和特定模态的记忆状态,这些状态持续总结音频和视频上下文。此外,引入跨模态记忆交互以增强音视频同步。为了有效学习这种增强记忆的模型,我们设计了一个三阶段的训练方案:(1)将双向音视频教师适配为块状因果注意力,并使用模拟记忆,(2)通过端到端蒸馏优化记忆构建和交互管道,以及(3)应用针对流式音视频生成的在线强化后训练。因此,Ripple在480P分辨率下实现了约28帧每秒的速度,远快于教师,同时能够进行连贯的长格式生成。在短视频和长视频基准上的大量实验表明,我们的性能优于现有的离线和在线联合音视频生成方法。
cs.CV / 61 / 2607.26829

BATS: Resource-Efficient Volumetric Segmentation with Boundary-Aware Mixed-Resolution Tokens

BATS:边界感知混合分辨率标记的资源高效体积分割
Hagerman, David, Naeem, Roman, Kahl, Fredrik
Abstract
Many high-performing volumetric segmentation models maintain dense multi-scale feature maps, leading to high activation memory and inference cost. We present BATS (Boundary-Aware Token Selection), a 3D medical image segmentation architecture that concentrates fine-resolution processing near predicted class boundaries. A dense boundary predictor identifies where additional resolution is needed, while a fine-first context cascade constructs an input-dependent mixed-resolution hierarchy. Homogeneous regions are represented coarsely, with finer tokens retained around boundaries, thin structures, and small targets. The sparse hierarchy is refined and rasterised into a dense segmentation. BATS predicts boundary relevance independently at every resolution level, preventing an erroneous coarse-scale decision from suppressing fine-scale evidence. Parent cluster attention further injects hierarchical ancestor tokens into local attention neighbourhoods, providing cross-scale context without dense multi-scale feature maps or cross-scale neighbour search. We evaluate BATS on five public CT and MRI datasets using the standardised nnU-Net Revisited protocol. BATS achieves the highest LiTS Dice among the compared methods and averages within 0.37 Dice points of the strongest dense baseline, MedNeXt-L, across the five datasets. Relative to MedNeXt-L, it reduces peak allocated GPU memory by more than 53% on KiTS, LiTS, and BraTS. Inference is up to 30% faster on KiTS and LiTS, which retain fewer tokens, but slower on the more token-dense BraTS. Mixed-resolution processing therefore provides consistent memory savings, while runtime and accuracy gains depend on dataset boundary density.
Chinese Translation
许多高性能的体积分割模型维持密集的多尺度特征图,导致高激活内存和推理成本。我们提出了BATS(边界感知标记选择),这是一种3D医学图像分割架构,专注于在预测类别边界附近进行细分辨率处理。一个密集的边界预测器识别出需要额外分辨率的区域,而一个细粒度优先的上下文级联构建了一个依赖于输入的混合分辨率层次结构。同质区域被粗略表示,而在边界、细结构和小目标周围保留更细的标记。稀疏层次结构经过精炼并光栅化为密集分割。BATS在每个分辨率级别独立预测边界相关性,防止错误的粗尺度决策抑制细尺度证据。父集群注意力进一步将层次祖先标记注入局部注意力邻域,提供跨尺度上下文,而无需密集的多尺度特征图或跨尺度邻居搜索。我们在五个公共CT和MRI数据集上使用标准化的nnU-Net Revisited协议评估BATS。BATS在比较方法中实现了最高的LiTS Dice,并且在五个数据集中平均与最强的密集基线MedNeXt-L相差仅0.37 Dice点。相较于MedNeXt-L,它在KiTS、LiTS和BraTS上减少了超过53%的峰值GPU内存分配。在KiTS和LiTS上,推理速度提高了最多30%,因为它们保留了更少的标记,但在标记更密集的BraTS上则较慢。因此,混合分辨率处理提供了一致的内存节省,而运行时间和准确性提升则依赖于数据集的边界密度。
cs.CV / 62 / 2607.26848

ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures

ICDAR 2026 原子层沉积/刻蚀 (ALD/E) 科学图形信息提取竞赛
Ahmed, Fahad, Auer, Sören, D'Souza, Jennifer
Abstract
Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset, accompanied by a community-driven competition, raises the bar over prior scientific competitions by curating a comprehensive, expert-annotated dataset across four end-to-end complementary tasks. The competition attracted 68 active participants and 1,263 public/private submissions from 9th January 2026 to 8th April 2026. Our results show that state-of-the-art multimodal models perform well on classification and summarization tasks but struggle with data extraction and scientific reasoning, particularly in visual question-answering. These findings reveal key limitations and highlight challenges and opportunities for improving domain-aware multimodal AI systems. Overall, the Sci-ImageMiner benchmark and competition establish a rigorous platform for advancing research in scientific figure comprehension and reasoning and demonstrate the potential of state-of-the-art approaches for a challenging and complex research area.
Chinese Translation
使用多模态人工智能进行科学图形理解和推理需要将视觉感知与特定领域的推理相结合,以提取有意义的知识,这些知识通常未在研究出版物的文本中呈现。Sci-ImageMiner 基准数据集伴随一个社区驱动的竞赛,提升了以往科学竞赛的标准,通过策划一个全面的、专家注释的数据集,涵盖四个端到端的互补任务。该竞赛吸引了68名活跃参与者,从2026年1月9日到2026年4月8日共提交了1,263份公共/私人作品。我们的结果表明,最先进的多模态模型在分类和摘要任务上表现良好,但在数据提取和科学推理方面存在困难,尤其是在视觉问答中。这些发现揭示了关键的局限性,并强调了改进领域感知的多模态人工智能系统的挑战和机遇。总体而言,Sci-ImageMiner 基准和竞赛为推动科学图形理解和推理的研究建立了一个严格的平台,并展示了最先进的方法在这一具有挑战性和复杂性的研究领域的潜力。
cs.CV / 63 / 2607.26885

SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation

SCALPEL:通过LLM驱动的编码器学习实现医学视觉-语言表示的语义跨模态对齐
Fu, Yunzhan, Bao, Enyu, Shen, Xiangyu, Wu, Yihao, Jiang, Chunbo, Guan, Fangli, Yan, Liqi
Abstract
Vision-language pre-training (VLP) serves as a cornerstone for medical multimodal representation learning. However, existing medical VLP frameworks are often constrained by the limited context windows and shallow representational capacities of lightweight text encoders when processing lengthy, terminology-dense clinical reports. While integrating medical large language models (LLMs) offers unprecedented clinical reasoning capabilities, it introduces three major bottlenecks: (i) the anisotropic representational collapse of generative LLMs under standard contrastive objectives, (ii) the prohibitive memory overhead of joint end-to-end training with large batch sizes, and (iii) the medical hallucinations induced by vanilla contrastive losses that ignore fine-grained anatomical laterality and negation modifiers. To address these challenges, we propose \textbf{SCALPEL}, a \textbf{S}emantic \textbf{C}ross-modal \textbf{A}lignment framework via \textbf{L}LM-\textbf{P}owered \textbf{E}ncoder \textbf{L}earning. First, Clinical Report Contrastive fine-tuning converts a generative LLM into an isotropic encoder via domain-specific clinical text adaptation. Second, an asymmetric alignment strategy leverages offline feature caching to enable efficient training. Critically, we formulate an Anatomy-Negation Aware Objective that explicitly penalizes mismatched image-text pairs involving laterality confusion or false negations. Extensive experiments across MIMIC-CXR, CheXpert, and IU X-Ray benchmarks demonstrate that SCALPEL achieves state-of-the-art performance in cross-modal retrieval, zero-shot disease classification and medical visual question answering.
Chinese Translation
视觉-语言预训练(VLP)是医学多模态表示学习的基石。然而,现有的医学VLP框架在处理冗长且术语密集的临床报告时,往往受到轻量级文本编码器有限的上下文窗口和浅层表示能力的限制。尽管整合医学大型语言模型(LLMs)提供了前所未有的临床推理能力,但也带来了三个主要瓶颈:(i)在标准对比目标下生成性LLMs的各向异性表示崩溃,(ii)与大批量训练的联合端到端训练所需的高昂内存开销,以及(iii)由于忽视细粒度解剖侧性和否定修饰符而引发的医学幻觉。为了解决这些挑战,我们提出了 extbf{SCALPEL},一个通过 extbf{L}LM- extbf{P}owered extbf{E}ncoder extbf{L}earning实现的 extbf{S}emantic extbf{C}ross-modal extbf{A}lignment框架。首先,临床报告对比微调通过领域特定的临床文本适配将生成性LLM转换为各向同性编码器。其次,非对称对齐策略利用离线特征缓存来实现高效训练。关键是,我们制定了一个解剖-否定感知目标,明确惩罚涉及侧性混淆或错误否定的不匹配图像-文本对。针对MIMIC-CXR、CheXpert和IU X-Ray基准的广泛实验表明,SCALPEL在跨模态检索、零样本疾病分类和医学视觉问答中实现了最先进的性能。
cs.CV / 64 / 2607.26886

Hearsay: Vision-Language Medical Diagnoses Without an Image

Hearsay:无图像的视觉-语言医学诊断
Vohra, Siddharth
Abstract
When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts the diagnosis returned. Claude concentrates sharply: a 65-year-old white man asking about a skin mole receives Melanoma in nearly every response, and a 32-year-old Black woman asking about her chest X-ray receives a Sarcoidosis diagnosis whose reasoning reads "suspected, based on demographics and classic pattern.'' GPT-5.4's effect is broader, fabricating across every demographic cell we test, most conspicuously naming Sarcoidosis for young Black patients on chest X-ray. Two structural findings sharpen the problem. A hedged regime appears in which the prose acknowledges the missing image while the structured diagnosis field nevertheless names a disease, a dissociation invisible to prose-only audits. And Claude's dermatology effect collapses entirely when 'skin mole' is swapped for 'skin lesion' while GPT-5.4's is preserved, indicating that mirage is a family of distinct failure modes rather than a single phenomenon. Trustworthy VLM deployment in clinical pipelines requires auditing the structured output channel directly, and probe-word sensitivity should be treated as a first-class evaluation dimension
Chinese Translation
当被要求描述一幅从未附上的医学图像时,前沿的视觉-语言模型并不回避:它们编造出一个诊断。我们展示了这种编造并非随机,而是由患者的身份所结构化。在胸部X光、脑部MRI和皮肤科的案例中,Claude Opus-4.7、GPT-5.4和Gemini-3.1-Pro仅通过人口统计描述进行查询,而没有图像,改变描述符系统性地改变了返回的诊断。Claude的反应非常集中:一位65岁的白人男性询问皮肤痣时几乎每次都收到黑色素瘤的诊断,而一位32岁的黑人女性询问她的胸部X光时则收到一个结节病的诊断,其推理为“基于人口统计和经典模式的怀疑”。GPT-5.4的影响更为广泛,在我们测试的每个人口统计单元中都进行了虚构,尤其是在胸部X光中为年轻黑人患者命名结节病。两个结构性发现加剧了这一问题。出现了一种保留的模式,其中文体承认缺失的图像,而结构化的诊断字段仍然命名一种疾病,这种解离在仅依赖文体的审计中是不可见的。当“皮肤痣”被替换为“皮肤病变”时,Claude的皮肤科效果完全崩溃,而GPT-5.4的效果则得以保留,这表明这种幻影是一系列不同失败模式的集合,而非单一现象。在临床流程中可靠的视觉-语言模型部署需要直接审计结构化输出通道,探测词敏感性应被视为一项重要的评估维度。
cs.CV / 65 / 2607.26910

CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents

CinemaTraj:利用大型语言模型代理为3D场景构建原子相机轨迹
Li, Qianru, Chen, Xuyang, Türköz, Erkin, Liu, Lu, Wang, Xuqin, Meng, Liqiu, Wu, Tao, Zhang, Yanfeng
Abstract
Automatically generating cinematically expressive camera trajectories through 3D scenes from natural language descriptions is a challenging task of high practical value, with applications ranging from real-estate advertising to virtual tour creation. Existing methods either lack true 3D spatial awareness by relying on 2D image priors, or treat trajectory generation as a geometric path planning problem divorced from cinematographic semantics. We present CinemaTraj, a framework that reframes camera trajectory planning as a language-grounded spatial reasoning problem. Given a set of RGB-D images and a user prompt, CinemaTraj equips an LLM agent with a structured 3D scene graph: the agent decomposes the prompt into a sequence of atomic cinematographic movements (dolly, orbit, crane, pan, tilt, zoom, arc). Each movement is instantiated via a novel parametric trajectory representation that is both cinematographically expressive and optimizable for collision avoidance. The scene graph acts as a structured spatial prior, grounding the agent's reasoning in accurate geometric and semantic knowledge of the environment. CinemaTraj further generates synchronized voiceover and subtitles aligned with camera motion, producing narrated cinematic video outputs. We evaluate CinemaTraj on real-world ScanNet++ environments, and show that it produces prompt-faithful, collision-free trajectories with high cinematographic quality, outperforming existing approaches on prompt alignment, trajectory quality, and safety metrics.
Chinese Translation
从自然语言描述自动生成具有电影表现力的相机轨迹穿越3D场景是一项具有高实用价值的挑战性任务,应用范围包括房地产广告到虚拟旅游创建。现有方法要么由于依赖于2D图像先验而缺乏真正的3D空间意识,要么将轨迹生成视为与电影语义脱离的几何路径规划问题。我们提出了CinemaTraj,一个将相机轨迹规划重新框定为基于语言的空间推理问题的框架。给定一组RGB-D图像和用户提示,CinemaTraj为大型语言模型(LLM)代理提供了一个结构化的3D场景图:代理将提示分解为一系列原子电影运动(推移、轨道、起重、平移、倾斜、缩放、弧线)。每个运动通过一种新颖的参数化轨迹表示得以实现,该表示在电影表现力上具有优势,并且可以优化以避免碰撞。场景图作为结构化的空间先验,使代理的推理基于环境的准确几何和语义知识。CinemaTraj进一步生成与相机运动同步的旁白和字幕,产生叙述性的电影视频输出。我们在现实世界的ScanNet++环境中评估了CinemaTraj,结果表明其生成的轨迹忠实于提示、无碰撞且具有高电影质量,在提示对齐、轨迹质量和安全性指标上优于现有方法。
cs.CV / 66 / 2607.26913

Prior Directions: Why GUI Grounding Gets Locked in the Past

先前方向:为何图形用户界面(GUI)基础被锁定在过去
Gong, Weile, Lu, Zijian, Chen, Mingcai, Zuo, Yiping, He, Xin, Fan, Weibei
Abstract
Vision-language models often use descriptions of earlier visual states to make decisions about the current scene. When the scene changes, stale language can redirect an otherwise correct visual judgment toward an outdated answer. We study this failure as visual lock-in in a controlled grounding setting where only the verbalized prior varies. Across models, stronger lock-in accompanies smaller changes in the model representation before the final answer. This reversal suggests that lock-in depends not on how far this representation moves, but on how that movement is organized. In models that are harder to correct, prior-induced changes concentrate along a compact set of directions that repeatedly appear across examples. We call these recurrent axes the Prior Directions. They recur on held-out examples, while a descriptive four-model comparison associates greater concentration with stronger lock-in. Controlled interventions show that removing the component aligned with the Prior Directions restores visual grounding, whereas removing an equally large orthogonal component has little effect. Prior control thus arises when prior-induced changes form a coherent and reusable pattern in the representation used to produce the answer. This account explains why the same prior remains revisable in one model yet becomes dominant in another.
Chinese Translation
视觉-语言模型通常使用早期视觉状态的描述来对当前场景做出决策。当场景发生变化时,过时的语言可能会将原本正确的视觉判断引导至一个过时的答案。我们在一个受控的基础设置中研究这种失败,只有口头表达的先前信息发生变化。在不同模型中,较强的锁定现象伴随着模型表示在最终答案之前的小变化。这种逆转表明,锁定现象并不依赖于表示移动的距离,而是依赖于这种移动的组织方式。在更难以纠正的模型中,先前引起的变化集中在一组紧凑的方向上,这些方向在多个示例中反复出现。我们称这些重复出现的轴为先前方向(Prior Directions)。它们在保留的示例中反复出现,而一个描述性的四模型比较则将更大的集中度与更强的锁定现象关联。控制干预表明,去除与先前方向对齐的成分可以恢复视觉基础,而去除一个同样大的正交成分则几乎没有影响。因此,当先前引起的变化在用于生成答案的表示中形成一致且可重用的模式时,先前控制就会出现。这一解释说明了为何同一先前在一个模型中仍然可修正,而在另一个模型中却变得主导。
cs.CV / 67 / 2607.26921

From Keypoints to Predictive Distributions: Post-Hoc Uncertainty for YOLO-Pose Models

从关键点到预测分布:YOLO-Pose模型的后验不确定性
Klushyn, Alexej, Sesma, Juan Rivero, Seligmann, Florian, Kurle, Richard, Tieu, Kinh, Gupta, Jayant Sen
Abstract
YOLO-Pose models provide efficient keypoint localization, but do not quantify the associated spatial uncertainty. We introduce a lightweight post-hoc probabilistic extension that augments a trained YOLO-Pose model with calibrated bivariate predictive distributions over keypoint locations, centered at the model's original predictions. Concretely, we train additional probabilistic heads with an importance-weighted negative log-likelihood to predict an input-dependent $2\times2$ dispersion matrix for each keypoint, followed by Gaussian calibration for broad downstream compatibility or Student-$t$ calibration for distributional fidelity. Complementing this, we propose an evaluation protocol that combines a suite of distributional calibration diagnostics with average keypoint precision (AKP), a keypoint-level extension of the COCO AP protocol for assessing reliability rankings. Experiments on COCO show that the learned uncertainty estimates enable effective keypoint-level reliability ranking, Student-$t$ calibration best captures the empirical residual distribution, and uncertainty-based pruning removes unreliable keypoints. A central application-level demonstration is vision-based aircraft landing, where calibrated covariances for runway keypoints support uncertainty-aware aircraft position estimation and downstream sensor fusion.
Chinese Translation
YOLO-Pose模型提供了高效的关键点定位,但未量化相关的空间不确定性。我们引入了一种轻量级的后验概率扩展,增强了经过训练的YOLO-Pose模型,使其能够在关键点位置上提供经过校准的双变量预测分布,中心位于模型的原始预测值。具体而言,我们训练了额外的概率头,通过重要性加权的负对数似然来预测每个关键点的输入依赖性$2 imes2$离散矩阵,随后进行高斯校准以实现广泛的下游兼容性,或进行Student-$t$校准以确保分布的保真性。为此,我们提出了一种评估协议,结合了一系列分布校准诊断与平均关键点精度(AKP),这是COCO AP协议在关键点级别的扩展,用于评估可靠性排名。在COCO上的实验表明,学习到的不确定性估计能够有效地进行关键点级别的可靠性排名,Student-$t$校准最佳地捕捉了经验残差分布,而基于不确定性的剪枝则移除了不可靠的关键点。一个核心的应用级演示是基于视觉的飞机着陆,其中跑道关键点的校准协方差支持不确定性感知的飞机位置估计和下游传感器融合。
cs.CV / 68 / 2607.26947

Progressive Multimodal Alignment for Continual Instruction Tuning

渐进式多模态对齐用于持续指令调优
Zhang, Duzhen, Yu, Yahan, Su, Qiaoyi, Dong, Jiahua, Zhang, Tielin
Abstract
Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it central to cross-modal understanding. In Multimodal Continual Instruction Tuning (MCIT), however, shifting visual distributions and evolving instruction semantics cause this shared projector to drift, leading to projector-level forgetting, an issue largely overlooked by methods that focus primarily on the LLM backbone. We introduce Progressive Multimodal Alignment (PMA), a framework that enables the projector to adapt continually while preserving previously learned alignment. PMA detects multimodal distribution shifts via a lightweight representation descriptor and progressively expands projector experts only when needed. An expandable router integrates expert outputs based on multimodal features, while the original pretrained projector is retained as a stable alignment anchor. This progressive mechanism balances stability and plasticity with sub-linear parameter growth and serves as a method-agnostic add-on to existing MCIT approaches. Extensive experiments on two recent MCIT benchmarks demonstrate that mitigating projector-level forgetting yields consistent gains over prior state-of-the-art methods when combined with PMA. Moreover, PMA scales across diverse MLLM backbones, demonstrating robust and broadly applicable MCIT performance.
Chinese Translation
多模态大型语言模型(MLLMs)依赖于投影器将视觉表征与语言嵌入空间对齐,这对于跨模态理解至关重要。然而,在多模态持续指令调优(MCIT)中,视觉分布的变化和指令语义的演变导致这一共享投影器发生漂移,从而引发投影器级遗忘,这一问题在主要关注LLM骨干网络的方法中往往被忽视。我们提出了渐进式多模态对齐(PMA)框架,使投影器能够持续适应,同时保持先前学习的对齐。PMA通过轻量级表征描述符检测多模态分布变化,并仅在需要时逐步扩展投影器专家。一个可扩展的路由器基于多模态特征整合专家输出,同时保留原始预训练投影器作为稳定的对齐锚点。这一渐进机制在参数增长上保持次线性,平衡了稳定性和可塑性,并作为一种与方法无关的附加组件应用于现有的MCIT方法。在两个最新的MCIT基准上进行的大量实验表明,减轻投影器级遗忘与PMA结合时,相较于之前的最先进方法,能够获得一致的提升。此外,PMA在多种MLLM骨干网络中具有良好的扩展性,展现出稳健且广泛适用的MCIT性能。
cs.CV / 69 / 2607.26973

Robust RPC Bundle Adjustment for Multi-Date Satellite Imagery with Season-Invariant Correspondences

针对季节不变对应关系的多日期卫星影像的鲁棒RPC束调整
Marí, Roger, Masquil, Elías, Bou, Xavier, Ehret, Thibaud, Facciolo, Gabriele
Abstract
Accurate refinement of Rational Polynomial Camera (RPC) models is essential for high-quality satellite image geolocation. In ground control point (GCP)-free multi-view pipelines, this refinement is commonly performed through bundle adjustment from automatically extracted image correspondences. However, conventional RPC bundle adjustment pipelines rely on handcrafted feature matching, which becomes unreliable in multi-date collections affected by seasonal, illumination, and land-cover changes. We propose an appearance-aware RPC refinement pipeline that combines learned local feature matching for season-invariant correspondences with global image descriptors for selecting visually compatible image pairs. This reduces redundant and error-prone matching while preserving the connectivity of the matching graph. Experiments on seasonally diverse WorldView-3 images show that our pipeline improves GCP-free relative RPC refinement over open-source baselines, achieving lower geometric consistency errors while substantially reducing matching time on collections with 39-42 views. By making RPC refinement more robust to diachronic appearance variation, our approach enables more effective use of multi-date satellite imagery.
Chinese Translation
精确细化有理多项式相机(Rational Polynomial Camera, RPC)模型对于高质量卫星图像的地理定位至关重要。在无地面控制点(Ground Control Point, GCP)的多视角处理流程中,这种细化通常通过从自动提取的图像对应关系中进行束调整。然而,传统的RPC束调整流程依赖于手工特征匹配,这在受到季节、光照和地表覆盖变化影响的多日期集合中变得不可靠。我们提出了一种基于外观感知的RPC细化流程,该流程结合了用于季节不变对应关系的学习局部特征匹配和用于选择视觉兼容图像对的全局图像描述符。这减少了冗余和易出错的匹配,同时保持了匹配图的连通性。在季节多样的WorldView-3图像上的实验表明,我们的流程在无GCP的相对RPC细化方面优于开源基线,取得了更低的几何一致性误差,同时在39-42视角的集合上显著减少了匹配时间。通过使RPC细化对时间变化的外观变化更加鲁棒,我们的方法能够更有效地利用多日期卫星影像。
cs.CV / 70 / 2607.27036

Mitigating Compounding Error via Video Representation Regularization

通过视频表示正则化减轻复合误差
Chen, Taiye, Zhang, Qi, Wang, Yisen
Abstract
Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time. Although this phenomenon has been widely observed, the underlying mechanism of compounding error and how to achieve stable long-horizon generation remain largely unresolved. In this paper, we investigate the internal representation dynamics of video world models and discover that compounding error is tightly coupled with dimensional collapse of hidden representations. Specifically, the effective rank of model representations sharply decreases at the onset of generation drift, revealing a strong connection between representational degradation and long-term rollout instability. Furthermore, we find that pure training data scaling fails to boost model resistance to error drift, contradicting mainstream scaling paradigms. To address this problem, we propose video representation regularization, a lightweight training constraint that stabilizes latent representations and suppresses iterative error accumulation. Compared with Diffusion Forcing, our method achieves improvements from 38.65 to 55.56 and from 44.37 to 72.08 on the Aesthetic Quality and Imaging Quality metrics of VBench. Our work establishes the first connection between autoregressive video drifting and model internal representations, adopts erank as a quantitative metric for error accumulation, reveals counterintuitive scaling limitations for video world models, and presents a simple yet effective regularization strategy to improve long video generation robustness.
Chinese Translation
基于视频扩散的世界模型使得机器人、自动驾驶和仿真任务能够进行长时间的自回归视频生成,但滑动窗口自回归推理在时间上严重的误差累积导致帧质量下降。尽管这一现象已被广泛观察,但复合误差的潜在机制以及如何实现稳定的长时间生成仍然在很大程度上未得到解决。本文研究了视频世界模型的内部表示动态,发现复合误差与隐藏表示的维度崩溃紧密相关。具体而言,模型表示的有效秩在生成漂移开始时急剧下降,揭示了表示退化与长期展开不稳定性之间的强关联。此外,我们发现单纯扩大训练数据并未能增强模型对误差漂移的抵抗力,这与主流的扩展范式相悖。为了解决这一问题,我们提出了视频表示正则化,这是一种轻量级的训练约束,能够稳定潜在表示并抑制迭代误差累积。与Diffusion Forcing相比,我们的方法在VBench的美学质量和成像质量指标上分别实现了从38.65到55.56和从44.37到72.08的提升。我们的工作首次建立了自回归视频漂移与模型内部表示之间的联系,采用有效秩(erank)作为误差累积的定量指标,揭示了视频世界模型的反直觉扩展限制,并提出了一种简单而有效的正则化策略,以提高长视频生成的鲁棒性。
cs.CV / 71 / 2607.27058

Object Detection for Autonomous Driving in Chinese Rural Scenes: An Experimental Study on Real-Synthetic Data Mixing and Model Evaluation

中国农村场景下的自动驾驶物体检测:真实与合成数据混合及模型评估的实验研究
Zhu, Danning, Lin, Ziyan, Wu, Jing
Abstract
Currently, autonomous driving object detection models face significant data scarcity and generalization challenges when navigating complex Chinese rural traffic scenarios. To address these limitations, we propose a novel real-synthetic mixed object detection dataset tailored specifically for Chinese rural roads and systematically evaluate the performance of 13 mainstream detectors under different real-to-synthetic data ratios, thereby providing empirical evidence for model selection and data strategy design in rural autonomous driving scenarios. Our dataset combines real-world images captured in Weishi County, Henan Province, with parameterized virtual scenes generated via Unreal Engine. To accurately reflect the unique realities of rural traffic, we define a comprehensive 14-category object system encompassing region-specific elements such as electric tricycles, low-speed vehicles (LSVs), and roadside stalls. Under a unified training protocol, we systematically evaluate 13 mainstream detectors -- spanning the YOLOv5, YOLOv8, YOLO11, and YOLO26 series, as well as RT-DETR-L -- across three data configurations: an all-real baseline, a 1:0.5 real-to-virtual mix, and a 1:1 mix. Experimental results demonstrate that a moderate injection of synthetic data (1:0.5 ratio) effectively enhances detection performance, with YOLO11m achieving the highest [email protected] of 0.758. However, a higher proportion of synthetic data (1:1) introduces domain shifts that offset the benefits of data scaling. While most models reliably identify distinct local vehicles, significant perceptual bottlenecks remain for long-tail, non-standard objects like stalls and railings. This research provides crucial empirical evidence and novel insights for model selection and synthetic data strategies, facilitating the practical deployment of autonomous driving perception systems in rural areas.
Chinese Translation
目前,自动驾驶物体检测模型在复杂的中国农村交通场景中面临显著的数据稀缺和泛化挑战。为了解决这些限制,我们提出了一种新颖的专门针对中国农村道路的真实-合成混合物体检测数据集,并系统地评估了13种主流检测器在不同真实与合成数据比例下的性能,从而为农村自动驾驶场景中的模型选择和数据策略设计提供实证依据。我们的数据集结合了在河南省卫辉市捕获的真实世界图像与通过虚幻引擎(Unreal Engine)生成的参数化虚拟场景。为了准确反映农村交通的独特现实,我们定义了一个涵盖区域特定元素的综合14类物体系统,包括电动三轮车、低速车辆(Low-Speed Vehicles, LSVs)和路边摊等。在统一的训练协议下,我们系统地评估了13种主流检测器——涵盖YOLOv5、YOLOv8、YOLO11和YOLO26系列,以及RT-DETR-L——在三种数据配置下的表现:全真实基线、1:0.5的真实与虚拟混合,以及1:1的混合。实验结果表明,适度注入合成数据(1:0.5比例)有效提升了检测性能,其中YOLO11m达到了最高的[email protected]为0.758。然而,更高比例的合成数据(1:1)引入了领域转移,抵消了数据扩展的好处。尽管大多数模型能够可靠地识别出不同的本地车辆,但对于长尾的非标准物体,如摊位和护栏,仍然存在显著的感知瓶颈。本研究为模型选择和合成数据策略提供了重要的实证依据和新颖的见解,促进了自动驾驶感知系统在农村地区的实际部署。
cs.CV / 72 / 2607.27065

ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection

ScratchSim:一种用于表面划痕检测的程序合成数据管道
Kühn, Paul Julius, Sinha, Saptarshi Neil, Kleist, Tiago, Hoffmann, Richard, kuijper, Arjan, Weinmann, Michael
Abstract
While automated defect detection such as the detection of surface scratched is an important aspect in industrial quality control, the scarcity of annotated defect data make this task challenging. This paper presents a procedural rendering pipeline that generates large-scale annotated synthetic training data using BlenderProc, with configurable material appearance, camera modes, and domain randomization, producing automatic COCO-format annotations. To show the potential of our approach, we evaluate four training strategies, namely synthetic-only, real-only, mixed, and fine-tuning from synthetic weights, across two objects with different material properties and three lightweight edge-deployable detectors, YOLOX, YOLO26, and LW-DETR. Our evaluation show that fine-tuning from synthetic weights consistently outperforms real-only training, and that mixed training effectively recovers performance under scarce real-data conditions, with findings validated across both convolutional and transformer-based architectures. The proposed approach enables scalable defect detection without the burden of large real annotated datasets, making it practical for on-device industrial inspection. The pipeline scripts, 3D model, and both synthetic and real annotated scratch datasets for a glossy toy Ferrari car will be made available through the project website upon acceptance.
Chinese Translation
尽管自动缺陷检测(如表面划痕检测)是工业质量控制中的一个重要方面,但标注缺陷数据的稀缺使得这一任务面临挑战。本文提出了一种程序渲染管道,利用 BlenderProc 生成大规模标注的合成训练数据,具有可配置的材料外观、相机模式和领域随机化,并自动生成 COCO 格式的标注。为了展示我们方法的潜力,我们在两种具有不同材料特性的物体上评估了四种训练策略,即仅合成、仅真实、混合以及从合成权重进行微调,并使用三种轻量级边缘可部署检测器:YOLOX、YOLO26 和 LW-DETR。我们的评估表明,从合成权重进行微调的效果始终优于仅真实训练,而混合训练在真实数据稀缺的情况下有效恢复了性能,这一发现已在卷积和基于变换器的架构中得到验证。所提出的方法使得在没有大量真实标注数据集负担的情况下实现可扩展的缺陷检测成为可能,使其在设备上的工业检测中具有实用性。项目网站将在论文接受后提供管道脚本、3D 模型以及用于光滑玩具法拉利汽车的合成和真实标注划痕数据集。
cs.CV / 73 / 2607.27066

SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence

SciFigAlign:通过与手稿证据的精细对齐对科学图形进行评分
Xu, Chuanzhi, Deng, Zihan, Liang, Huiqi, Yue, Chengkun, Cui, Zhanlin, Ye, Pengfei, Cai, Weidong
Abstract
Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully support the manuscript's claims, and communicate evidence with a clear visual hierarchy. However, if we apply traditional image assessment methods to scientific figure quality assessment, limitations emerge: classic IQA models capture perceptual quality or aesthetics but cannot judge whether a figure serves the paper's scientific argument; CLIP-based methods assess generic image-text correspondence, yet lack understanding of manuscript context; and zero-shot LLM/VLM judges, when repurposed for figure scoring, often yield overly concentrated scores with limited fusion of visual and textual evidence. We introduce an annotated dataset of 3,857 scientific figures from peer-reviewed conference papers, each rated along four peer-review-oriented dimensions: Clarity, Relevance, Informativeness, and Structure. We propose SciFigAlign, a fine-tuned multimodal scorer that grounds figure quality assessment in manuscript evidence. Given a figure crop, caption, citing paragraphs, and light paper context, SciFigAlign fine-tunes CLIP and SciBERT end-to-end with per-modality cross-attention and CubeMLP fusion, jointly optimizing SmoothL1 regression with a within-paper ranking hinge loss. Under paper-level splits, SciFigAlign achieves a macro MAE of 0.3524 and a within-paper pairwise accuracy of 81.64% on the test set with n = 396, a 59% relative error reduction over the best LLM-as-judge baseline with MAE 0.864. Ablations confirm that manuscript-grounded inputs, citing-context denoising, and ranking supervision are all critical, showing that scientific figure assessment requires learned alignment between visual content and manuscript evidence rather than prompting alone, even with state-of-the-art VLMs.
Chinese Translation
同行评审中的科学图形评估与一般图像质量评估在根本上有所不同:图形必须在视觉上清晰可读,忠实支持手稿的主张,并以明确的视觉层次传达证据。然而,如果我们将传统的图像评估方法应用于科学图形质量评估,就会出现局限性:经典的图像质量评估(IQA)模型捕捉感知质量或美学,但无法判断图形是否服务于论文的科学论点;基于CLIP的方法评估通用的图像-文本对应关系,但缺乏对手稿上下文的理解;而零-shot LLM/VLM评估者在被重新用于图形评分时,往往会产生过于集中且视觉与文本证据融合有限的评分。我们引入了一个包含3,857个来自同行评审会议论文的科学图形的注释数据集,每个图形在四个同行评审导向的维度上进行评分:清晰度、相关性、信息量和结构。我们提出了SciFigAlign,一个精细调优的多模态评分器,将图形质量评估与手稿证据相结合。给定一个图形裁剪、标题、引用段落和轻量级论文上下文,SciFigAlign通过每种模态的交叉注意力和CubeMLP融合对CLIP和SciBERT进行端到端的精细调优,联合优化SmoothL1回归与论文内排名的铰链损失。在论文级别的划分下,SciFigAlign在测试集(n = 396)上实现了0.3524的宏平均绝对误差(MAE)和81.64%的论文内成对准确率,相较于最佳的LLM作为评估者基线(MAE 0.864)减少了59%的相对误差。消融实验确认了基于手稿的输入、引用上下文去噪和排名监督都是关键因素,表明科学图形评估需要在视觉内容与手稿证据之间学习对齐,而不仅仅依赖提示,即使在最先进的VLM中也是如此。
cs.CV / 74 / 2607.27069

Visual Credit Audit for Multimodal Spatial Reasoning

多模态空间推理的视觉信用审计
Liu, Feixiang, Qiu, Qiang, Sun, Lanbo, Wei, Nan, Shen, Huawei, Cheng, Xueqi
Abstract
Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decision more support than text-only and blank controls, and whether the model responds to relation-specific visual evidence. The first audit is training- and label-free and does not require an answer flip. Applying labels yields dependence-credited correctness (D-CC); on correct items, it equals same-control gold-aligned positive gain, while prediction alignment extends the audit to errors. Across four open MLLMs and two spatial benchmarks, 12.73-26.25% of decisions are correct yet uncredited. Matched same-split image permutation reduces D-CC by 21.25-47.80 points, with every paired 95% interval above zero. Fixed-pixel relation contrasts and a 3x3 evidence-source factorial show why null controls cannot identify relation response. Among controlled correct-but-uncredited agreement decisions, response to relation reversal spans 81.57-100.00%, while 32.11% pooled change answer. Independently audited outcomes on 108 geometry-compatible edits provide a bounded natural-image correspondence check. VCA thereby decomposes benchmark success into correctness, additional image support, and relation-consistent response.
Chinese Translation
封闭的是/否空间基准在图像对比无图像上下文时,即使图像提供的支持有限,仍然可以奖励正确答案。在固定的强制选择界面下,视觉信用审计(Visual Credit Audit, VCA)分离了两个估计量:基准图像是否为模型声明的决策提供了比仅文本和空白对照更多的支持,以及模型是否对特定关系的视觉证据做出反应。第一个审计不依赖于训练和标签,也不需要答案翻转。应用标签会产生依赖信用正确性(Dependence-Credited Correctness, D-CC);在正确项上,它等于相同对照的黄金对齐正增益,而预测对齐则将审计扩展到错误。在四个开放的多模态语言模型(MLLMs)和两个空间基准中,12.73-26.25%的决策是正确但未被记入的。匹配的相同拆分图像置换将D-CC降低了21.25-47.80点,每个配对的95%区间均高于零。固定像素关系对比和3x3证据源因子显示了为何无效对照无法识别关系响应。在受控的正确但未记入的协议决策中,对关系反转的响应范围为81.57-100.00%,而32.11%的情况下更改了答案。对108个几何兼容编辑的独立审计结果提供了一个有限的自然图像对应检查。因此,VCA将基准成功分解为正确性、额外图像支持和关系一致响应。
cs.CV / 75 / 2607.27084

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

SciFigQual-Bench:一个用于科学图像质量评估的基准,结合全文背景
Deng, Zihan, Xu, Chuanzhi, Liang, Huiqi, Li, Haoyang, Zhong, Xiaozhen, Yu, Lequan
Abstract
Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated content, which cannot be directly applied to scientific papers. The few existing studies on scholarly charts remain confined to visual-surface comparisons, failing to verify caption alignment, citation relevance, or visual misleadingness. To address this, we propose SciFigQual-Bench, a full-text contextual benchmark that evaluates scientific images across five dimensions (clarity, layout, caption fit, context relevance, and misleading risk). The data covers top computer-science conferences from 2020 to 2025; 6,308 images were independently scored by multiple domain experts in five dimensions and aggregated into gold-standard annotations. Unlike previous scientific figure benchmarks, our dataset binds each image to its caption, citing sentence, and manuscript context. To enable automated evaluation on this benchmark, we designed a staged cross-modal evaluation framework SFQ-Agent to achieve auditable and refined scoring through the collection and fusion of modal evidence. Multiple mainstream large models were evaluated on the test subset eval1200, and SFQ-Agent (F3) equipped with GPT-5.6-Sol achieved the lowest overall average absolute error (0.418) and the highest consistency rate (93.4%), consistently outperforming both direct evaluation and auxiliary (Sidecar) visual language model evaluation schemes.
Chinese Translation
科学图像是展示实验结论、阐述系统架构以及支持科学论文中比较论证的核心元素。然而,现有的图像质量评估(IQA)方法主要针对自然照片或人工智能生成的内容,无法直接应用于科学论文。现有对学术图表的研究仍局限于视觉表面比较,未能验证图注对齐、引用相关性或视觉误导性。为了解决这一问题,我们提出了SciFigQual-Bench,一个全文本上下文基准,评估科学图像在五个维度上的表现(清晰度、布局、图注适配、上下文相关性和误导风险)。该数据涵盖了2020年至2025年间的顶级计算机科学会议,共有6,308幅图像由多位领域专家在五个维度上独立评分,并汇总为黄金标准注释。与以往的科学图像基准不同,我们的数据集将每幅图像与其图注、引用句子和手稿上下文绑定。为了在该基准上实现自动化评估,我们设计了一个分阶段的跨模态评估框架SFQ-Agent,通过收集和融合模态证据实现可审计和精细的评分。在测试子集eval1200上评估了多个主流大型模型,配备GPT-5.6-Sol的SFQ-Agent(F3)实现了最低的整体平均绝对误差(0.418)和最高的一致性率(93.4%),在直接评估和辅助(Sidecar)视觉语言模型评估方案中均表现优异。
cs.CV / 76 / 2607.27087

Step-Attention Refinement of DINOv3 Features for Efficient Anterior Eye Segmentation

DINOv3特征的步态注意力精炼用于高效前眼段分割
Baumstimler, Philippe, Gagnon, Jean-Mathieu, Gagné, Sébastien, Duchesneau, Mathieu, Playout, Clément, Séoud, Lama
Abstract
Anterior eye segment (AES) segmentation is a key component of both ocular biometrics and emerging clinical image analysis applications. However, heterogeneous acquisition conditions and limited annotations in medical settings hinder the robustness and generalization of existing methods. Foundation models (FMs) such as DINOv3 offer strong transfer capabilities, but efficiently adapting their representations to dense prediction tasks remains challenging. In this study, we investigate robust AES segmentation in clinical settings, and propose a lightweight architecture built upon a distilled DINOv3 ViT-Small backbone. We introduce a step-attention feature refinement module that progressively adapts multi-level transformer representations before convolutional decoding, enabling efficient exploitation of pretrained features with few parameters. We evaluate the proposed approach on a private dataset of 333 clinically acquired AES images spanning eight ophthalmic acquisition protocols and annotated for seven anatomical classes. Compared with convolutional and transformer-based baselines, including DINOv3-based methods, our approach achieves the best overall performance, reaching 85.55\% mIoU when fully fine-tuned. It also demonstrates the strongest robustness to domain shift across four unseen public AES segmentation datasets. These results establish a strong baseline for robust AES segmentation in clinical settings and highlight the importance of decoder design for effectively adapting FMs representations to medical segmentation tasks.
Chinese Translation
前眼段(AES)分割是眼部生物识别和新兴临床图像分析应用中的关键组成部分。然而,异质的采集条件和医疗环境中有限的标注阻碍了现有方法的鲁棒性和泛化能力。基础模型(FMs)如DINOv3提供了强大的迁移能力,但有效地将其表示适应于密集预测任务仍然具有挑战性。在本研究中,我们探讨了临床环境中的鲁棒AES分割,并提出了一种基于蒸馏DINOv3 ViT-Small骨干网的轻量级架构。我们引入了一种步态注意力特征精炼模块,该模块在卷积解码之前逐步适应多级变换器表示,从而高效利用预训练特征且参数较少。我们在一个包含333幅临床获取的AES图像的私有数据集上评估了所提方法,该数据集涵盖八种眼科采集协议,并标注了七个解剖类别。与包括基于DINOv3的方法在内的卷积和变换器基线相比,我们的方法在完全微调时达到了85.55%的mIoU,表现出最佳的整体性能。它还在四个未见的公共AES分割数据集上展示了对领域转移的最强鲁棒性。这些结果为临床环境中的鲁棒AES分割建立了强基线,并强调了解码器设计在有效适应FMs表示于医学分割任务中的重要性。
cs.CV / 77 / 2607.27110

FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring

FreqForcing:通过谱自锚实现自回归长视频生成
Li, Jiatong, Liang, Leo, Kong, Linghe, Zhang, Yulun
Abstract
Autoregressive video diffusion models enable real-time streaming video generation. However, errors introduced during self-rollout accumulate over long horizons, manifesting as color drift, motion stagnation, and eventual visual collapse. In this paper, we characterize this phenomenon from a frequency-domain perspective: error accumulation appears as a pronounced energy drift in the low-frequency bands. We further investigate the effectiveness of attention sink in the frequency domain, and find that it improves the video quality by alleviating the spectral energy drift to some extent, but cannot fully resolve it. Motivated by the above analysis, we propose FreqForcing, a training-free framework that addresses error accumulation in long-video generation via Spectral Self-Anchoring (SSA). The proposed SSA leverages the low-frequency components of anchor attention to maintain long-horizon visual stability, while preserving dynamic motion through the high-frequency components of local attention. Our FreqForcing extends Self-Forcing pretrained on 5s clips to two-minute generation, achieving 24x extrapolation. Extensive experiments show that FreqForcing outperforms existing training-free methods quantitatively and qualitatively while remaining competitive with representative training-based approaches.
Chinese Translation
自回归视频扩散模型使实时流媒体视频生成成为可能。然而,在自我展开过程中引入的错误会在长时间范围内累积,表现为颜色漂移、运动停滞以及最终的视觉崩溃。本文从频域的角度对这一现象进行了表征:错误累积在低频带中表现为显著的能量漂移。我们进一步研究了频域中注意力沉没的有效性,发现它在一定程度上通过减轻谱能量漂移来改善视频质量,但无法完全解决这一问题。基于上述分析,我们提出了FreqForcing,这是一种无训练框架,通过谱自锚(Spectral Self-Anchoring, SSA)来解决长视频生成中的错误累积问题。所提出的SSA利用锚点注意力的低频成分来维持长时间范围内的视觉稳定性,同时通过局部注意力的高频成分保持动态运动。我们的FreqForcing将自我强迫(Self-Forcing)扩展到基于5秒片段的两分钟生成,实现了24倍的外推。大量实验表明,FreqForcing在定量和定性上均优于现有的无训练方法,同时在与代表性的基于训练的方法比较时仍保持竞争力。
cs.CV / 78 / 2607.27113

Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection

Veritas++:面向价值的在线蒸馏用于增强感知的AI生成图像检测
Tan, Hao, Lan, Jun, Tan, Zichang, Liu, Ajian, Yu, Zijian, Song, Chuanbiao, Zhu, Huijia, Wang, Weiqiang, Wan, Jun, Lei, Zhen
Abstract
The growing capability of image generation models has made synthetic images a routine presence in open media, making robust and generalizable AI-Generated Image (AIGI) detection increasingly essential. While multi-modal large language models (MLLMs) offer a transparent alternative to black-box binary scoring, we observe that current MLLM-based detectors still exhibit notable perception bottlenecks in capturing fine-grained anomalies. They primarily focus on how visual evidence is organized and synthesized, leaving the intrinsic perception less optimized. To mitigate this gap, we present Veritas++, a perception-enhanced reasoning framework that establishes reliable perception as the foundation of authenticity reasoning. Rather than directly optimizing the model's explanatory ability, we ground AIGI detection on three basic perception abilities, i.e., capturing fine-grained visual details, semantic anomalies and pixel-level differences. Building on this insight, we introduce Perception-oriented Learning (PoRL), which replaces open-ended description supervision with verifiable rewards to explicitly strengthen these capacities. To further integrate enhanced perception with reasoning, we introduce Value-aware On-Policy Distillation (VaOPD), an adaptive distillation mechanism that prioritizes high-value distillation signals over uniform supervision, internalizing perception-aware reasoning through a privileged self-teacher. Extensive experiments across standard, in-the-wild and emerging benchmarks demonstrate that Veritas++ achieves promising generalization. The perception learning effectively bridges the perception gap and yields seamless gains on detection, while VaOPD further enables efficient capability evolvement without sacrificing existing performance. Code and checkpoints are available at https://github.com/EricTan7/VeritasPP.
Chinese Translation
图像生成模型的不断发展使得合成图像在开放媒体中变得常见,因此,稳健且具有普适性的AI生成图像(AIGI)检测变得愈加重要。尽管多模态大型语言模型(MLLMs)提供了一种透明的替代方案,取代了黑箱二元评分,但我们观察到当前基于MLLM的检测器在捕捉细粒度异常方面仍然存在显著的感知瓶颈。它们主要关注视觉证据的组织和合成,而对内在感知的优化不足。为了解决这一问题,我们提出了Veritas++,一个增强感知的推理框架,将可靠的感知作为真实性推理的基础。我们并不直接优化模型的解释能力,而是将AIGI检测建立在三种基本的感知能力之上,即捕捉细粒度视觉细节、语义异常和像素级差异。在此基础上,我们引入了面向感知的学习(PoRL),用可验证的奖励替代开放式描述监督,以明确增强这些能力。为了进一步将增强的感知与推理结合,我们引入了面向价值的在线蒸馏(VaOPD),这是一种自适应蒸馏机制,优先考虑高价值的蒸馏信号而非均匀监督,通过特权自教师内化感知驱动的推理。大量在标准、野外和新兴基准上的实验表明,Veritas++实现了令人鼓舞的泛化。感知学习有效弥补了感知差距,并在检测上带来了无缝的提升,而VaOPD进一步实现了高效的能力演变,而不牺牲现有性能。代码和检查点可在https://github.com/EricTan7/VeritasPP获取。
cs.CV / 79 / 2607.27122

Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs

通过小型视觉语言模型的多任务学习实现基于证据的胃肠内镜视觉问答
Safwan, Itbaan, Khan, Ramail, Shaikh, Muhammad Annas, Tahir, Muhammad Atif
Abstract
Gastrointestinal (GI) endoscopic image analysis has shifted from single-label classification toward visual question answering (VQA), where a model must answer free-form clinical questions about an image. While recent vision-language models (VLMs) achieve promising answer accuracy on this task, clinical adoption also requires the model's internal representations to reflect the visual evidence behind its answers. We propose a simple multi-task fine-tuning recipe that constructs auxiliary grounding and description tasks from an existing VQA dataset with minimal additional annotation: expert-annotated polyp masks are reused directly, while a GI-domain pretrained classifier with Grad-CAM localization provides weak supervision for finding categories that lack ground-truth masks. Three small VLM backbones are fine-tuned with low-rank adaptation under matched VQA-only and multi-task recipes on Kvasir-VQA-x1, and we show consistent accuracy gains together with improved implicit alignment between answer tokens and the relevant image region, evaluated on both in-distribution and out-of-distribution data.
Chinese Translation
胃肠(GI)内镜图像分析已经从单标签分类转向视觉问答(VQA),在这一任务中,模型必须回答关于图像的自由形式临床问题。尽管最近的视觉语言模型(VLMs)在这一任务上取得了令人鼓舞的答案准确率,但临床应用还要求模型的内部表征能够反映其答案背后的视觉证据。我们提出了一种简单的多任务微调方法,该方法从现有的VQA数据集中构建辅助的定位和描述任务,所需的额外标注最小:专家标注的息肉掩膜被直接重用,而一个经过胃肠领域预训练的分类器结合Grad-CAM定位为缺乏真实掩膜的类别提供弱监督。我们在Kvasir-VQA-x1数据集上对三个小型VLM骨干网络进行了低秩适应的微调,采用匹配的仅VQA和多任务方法,结果显示在准确性上取得了一致的提升,同时在答案标记与相关图像区域之间的隐式对齐也得到了改善,这一结果在分布内和分布外数据上均得到了评估。
cs.CV / 80 / 2607.27139

SeasonStereo: Robust Dense Stereo Matching for Multi-Date Satellite Imagery via Generative AI

SeasonStereo:基于生成性人工智能的多时相卫星影像的鲁棒密集立体匹配
Díaz-Laureano, Álvaro, Marí, Roger, Masquil, Elías, Arias, Pablo, Facciolo, Gabriele
Abstract
Accurate 3D reconstruction from satellite imagery typically relies on near-simultaneous stereo pairs, limiting its applicability to diachronic settings where multi-date images exhibit varying seasonal and illumination conditions. Training dense stereo matching models robust to appearance changes is a long-standing challenge, as aligned multi-date imagery and ground-truth geometry are costly to obtain at scale. We propose SeasonStereo, a scalable framework that addresses disparity estimation from diachronic satellite images by training on synthetic image pairs with controlled seasonal appearance variation, while leveraging zero-shot geometric priors from foundation models. SeasonStereo matches the accuracy of state-of-the-art LiDAR-supervised models, while producing sharper geometric details without requiring aligned real multi-date training products or LiDAR-derived labels. As a result, SeasonStereo offers a practical path toward large-scale 3D reconstruction from heterogeneous satellite images with reduced supervision cost.
Chinese Translation
从卫星影像中进行准确的三维重建通常依赖于近乎同时的立体对,这限制了其在多时相环境中的应用,因为多时相影像展示了不同的季节和光照条件。训练对外观变化具有鲁棒性的密集立体匹配模型一直是一个长期挑战,因为对齐的多时相影像和真实几何数据在大规模获取上成本高昂。我们提出了SeasonStereo,一个可扩展的框架,通过在具有可控季节外观变化的合成影像对上进行训练,解决了来自多时相卫星影像的视差估计问题,同时利用基础模型的零样本几何先验。SeasonStereo的准确性与最先进的激光雷达(LiDAR)监督模型相当,同时在不需要对齐的真实多时相训练产品或激光雷达衍生标签的情况下,生成更清晰的几何细节。因此,SeasonStereo为从异构卫星影像进行大规模三维重建提供了一条实用的路径,降低了监督成本。
cs.CV / 81 / 2607.27145

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications

可解释且资源高效的多模态大型语言模型空间推理在决策关键应用中的应用
Jain, Piyush, Dasgupta, Kousik, Roy, Rajarshi, Tripathi, Subarna
Abstract
As Multimodal Large Language Models (MLLMs) are increasingly deployed in decision-critical pipelines such as robotics, embodied AI, and safety monitoring, the opacity of their spatial judgments limits operator trust and auditability. MLLMs demonstrate strong reasoning but often struggle with fine-grained spatial understanding and object hallucination. Prior work, ByDeWay, introduced Layered-Depth-Based Prompting (LDP), a training-free framework that mitigates hallucinations by structuring prompts using monocular depth estimation. However, coarse depth layering falls short in resolving object-to-object spatial relationships within the same geometric plane, such as projective ("left of", "above") and topological ("inside", "touching") relations. We propose ByDeWay-V2, which integrates explicit spatial relational context alongside depth cues, expressed as human-readable predicates that serve as auditable evidence for downstream decision support. Using an open-vocabulary object detector (YOLO-World-L), our framework computes pairwise geometric relations between detected objects and injects them as structured spatial predicates into the MLLM prompt, bridging 3D scene depth and 2D spatial semantics without any training. We evaluate ByDeWay-V2 on the Visual Spatial Reasoning (VSR) and BLINK benchmarks across multiple MLLMs, with hallucination grounding assessed via POPE. On the BLINK spatial subset, ByDeWay-V2 achieves a 46 percent relative F1 improvement over LDP for Qwen2.5-VL, and recovers BLIP-Base's spatial reasoning on VSR from near-random performance to a competitive F1 of 0.53. Our lightest configuration operates under a strict 40-token context budget on CPU, showing the framework's suitability for resource-constrained, real-time decision-support settings.
Chinese Translation
随着多模态大型语言模型(MLLMs)在机器人技术、具身人工智能和安全监测等决策关键流程中的日益广泛应用,其空间判断的模糊性限制了操作员的信任度和审计能力。MLLMs展现出强大的推理能力,但在细粒度空间理解和物体幻觉方面常常面临挑战。先前的研究ByDeWay引入了一种基于分层深度的提示(Layered-Depth-Based Prompting, LDP)训练无关框架,通过使用单目深度估计来减轻幻觉现象。然而,粗糙的深度分层在解决同一几何平面内的物体间空间关系(如投影关系“左侧”、“上方”和拓扑关系“内部”、“接触”)方面显得不足。我们提出了ByDeWay-V2,它结合了显式的空间关系上下文和深度线索,以人类可读的谓词形式表达,作为下游决策支持的可审计证据。通过使用开放词汇的物体检测器(YOLO-World-L),我们的框架计算检测到的物体之间的成对几何关系,并将其作为结构化空间谓词注入MLLM提示中,连接3D场景深度和2D空间语义,而无需任何训练。我们在多个MLLMs上对ByDeWay-V2进行了Visual Spatial Reasoning(VSR)和BLINK基准测试评估,幻觉的基础通过POPE进行评估。在BLINK空间子集上,ByDeWay-V2在Qwen2.5-VL上实现了相对于LDP的46%的F1相对提升,并将BLIP-Base在VSR上的空间推理从接近随机表现恢复到竞争性的F1值0.53。我们最轻量的配置在CPU上严格遵循40个标记的上下文预算,显示出该框架在资源受限的实时决策支持环境中的适用性。
cs.CV / 82 / 2607.27154

Anatomy Contextualized Adaption of CT Foundation Models

CT基础模型的解剖学上下文化适应
Kenia, Roshan, McNamara, Stephanie L, Lotter, William
Abstract
CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by aligning anatomy-level visual features with anatomy-specific text, but in doing so discards the global context that whole-volume models provide. Furthermore, existing fine-grained approaches train from scratch, making them computationally expensive. We introduce Anatomy Contextualized Adaptation (ACA), a lightweight framework that adapts frozen CT foundation model representations for anatomy-level vision-language alignment while enhancing global contextualization. ACA uses TotalSegmentator to decompose CT volumes into anatomy-level embeddings, which are refined via a transformer that captures cross-anatomy relationships, and aligned to both per-anatomy and scan-level text extracted from radiology reports. Evaluated on Merlin and CT-RATE, ACA consistently outperforms both the frozen foundation model baselines and existing fine-grained methods in zero-shot finding classification, while requiring less than one hour of training once embeddings are cached. The attention weights learned by ACA's inter-anatomy transformer additionally indicate plausible cross-anatomy context routing. Altogether, these results support ACA as a lightweight approach for adapting CT foundation models to anatomically grounded vision-language alignment while preserving and enhancing global anatomical context.
Chinese Translation
CT视觉-语言基础模型在下游任务中表现出良好的性能,但通常使用整体体积表示进行训练,这会稀释细粒度的解剖信号。细粒度视觉-语言预训练通过将解剖级视觉特征与解剖特定文本对齐来解决这一问题,但这样做会丢弃整体体积模型所提供的全局上下文。此外,现有的细粒度方法通常从头开始训练,计算成本较高。我们提出了解剖学上下文化适应(Anatomy Contextualized Adaptation, ACA),这是一个轻量级框架,旨在适应冻结的CT基础模型表示,以实现解剖级视觉-语言对齐,同时增强全局上下文化。ACA使用TotalSegmentator将CT体积分解为解剖级嵌入,这些嵌入通过捕捉跨解剖关系的变换器进行精炼,并与从放射学报告中提取的每个解剖和扫描级文本对齐。在Merlin和CT-RATE上的评估表明,ACA在零样本查找分类中始终优于冻结基础模型基线和现有的细粒度方法,同时在嵌入缓存后训练时间少于一小时。ACA的跨解剖变换器学习到的注意力权重还表明了合理的跨解剖上下文路由。总的来说,这些结果支持ACA作为一种轻量级方法,将CT基础模型适应于解剖学基础的视觉-语言对齐,同时保留和增强全局解剖上下文。
cs.CV / 83 / 2607.27180

HumanCLAW: Can Vision-Language Models Act Through a Body?

HumanCLAW:视觉-语言模型能否通过身体进行行动?
Li, Siyao, Gu, Jiawei, Liu, Shuai, Hu, Kairui, Li, Zekun, Li, Linjie, Tang, Chengcheng, Wu, Po-Chen, Shugurov, Ivan, Ma, Lingni, Zollhoefer, Michael, An, Sizhe, Mittal, Abhay, Zhao, Amy, Krishna, Ranjay, Li, Manling, Liu, Ziwei, Guo, Chuan
Abstract
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.
Chinese Translation
评估视觉-语言模型(VLM)是否能够通过物理身体进行行动是具有挑战性的。一个动作的结果将VLM的决策与运动控制相结合。当任务失败时,很难判断是VLM做出了错误的选择,还是运动控制器未能执行该选择,例如失去平衡而摔倒。在本研究中,我们引入了HumanCLAW,一个评估框架,它将动作决策与低级执行解耦。在每一步中,一个经过训练的、现成的VLM发出一个原子技能命令,该命令被转换为一个具有实际物理后果的连续全身运动的亚秒级片段,包括重力和碰撞。因此,身体可以在物理世界中自由行动,而执行过程中的干扰、平衡和运动错误被排除在外。可测量的则是模型的行动智能:它在每一时刻选择身体下一步应执行的动作。基于这一框架,我们构建了HumanCLAW-Bench:在41个室内场景中包含1,218个长时间跨度、以自我为中心的寻找-导航-互动的情节。我们测试了九个最先进的VLM,发现没有一个模型能够解决基准测试;表现最好的模型仅达到16.8%的成功率。识别目标并不是瓶颈。目前的VLM缺乏的是具身自我意识:它们无法跟踪自己的身体,无法判断自己身在何处,是否已到达目标,或是否碰到了障碍物。
cs.CV / 84 / 2607.27194

VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion

VidMap:利用时间结构进行基于视频的运动重建
Pataki, Zador, Sarlin, Paul-Edouard, Pollefeys, Marc
Abstract
Accurately recovering the camera's calibration and metric poses for any unconstrained video would unlock large-scale training data for navigation and scene understanding. The dominant approaches to this problem are severely limited: Simultaneous Localization and Mapping (SLAM) is sensitive to initialization and transient failures due to its causal, incremental nature; it is often over-optimized for real-time operation and generally requires known camera calibration; while Structure-from-Motion (SfM) typically forgoes any image ordering, enabling optimal initialization and global optimization, but lacks robustness to visual symmetries and extreme motions. To bridge this gap, we introduce a system that combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling the metric reconstruction of arbitrary, long, uncalibrated videos. This system leverages recent advances in wide-baseline dense image matching, treats temporal ordering as a first-class citizen for reliable loop closure, and augments global optimization with metric monocular depth priors. As a result, thorough evaluations on diverse, challenging datasets that exhibit extreme motion and visual symmetries reveal that our approach is significantly more robust and accurate than both state-of-the-art SLAM and SfM, classical or learned, with given or unknown camera calibration. The code is publicly available at https://github.com/cvg/vidmap.
Chinese Translation
准确恢复任何无约束视频的相机标定和度量姿态将为导航和场景理解解锁大规模训练数据。当前对此问题的主流方法受到严重限制:同时定位与地图构建(SLAM)由于其因果性和增量特性,对初始化和瞬态故障非常敏感;它通常过于优化以实现实时操作,并且通常需要已知的相机标定;而运动重建(SfM)通常放弃任何图像排序,能够实现最佳初始化和全局优化,但对视觉对称性和极端运动缺乏鲁棒性。为了弥补这一差距,我们提出了一种系统,结合了SLAM的强序列约束与离线SfM的灵活性和全局优化,能够对任意长的无标定视频进行度量重建。该系统利用了宽基线密集图像匹配的最新进展,将时间排序视为可靠回环闭合的第一公民,并通过度量单目深度先验增强全局优化。因此,在表现出极端运动和视觉对称性的多样化挑战数据集上的全面评估表明,我们的方法在鲁棒性和准确性上显著优于当前最先进的SLAM和SfM,无论是经典方法还是学习方法,无论相机标定是否已知。代码已公开发布在 https://github.com/cvg/vidmap。
cs.CV / 85 / 2607.27205

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

TurboVLA:在RTX 4090上以32 Hz实时运行的视觉-语言-动作模型,显存小于1 GB
Xie, Hengyi, Yao, Chenfei, Wu, Xianjin, Xi, Xuanyang, Tang, Yiping, Xu, Di, Zhu, Yingying, Liang, Dingkang, Bai, Xiang, Ding, Han
Abstract
Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional $V \to L \to A$ pathway as a direct $V + L \to A$ mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. This simple design constructs task-conditioned representations directly from visual and linguistic features, significantly reducing the computational and memory costs of VLA inference. On LIBERO, TurboVLA achieves 97.7% average success with only 0.2B parameters, 31.2 ms inference latency, and 0.9 GB inference VRAM on a consumer-grade RTX 4090, matching or outperforming substantially larger VLA policies. These results establish TurboVLA as a simple and effective alternative to the prevailing LLM-centric VLA paradigm, offering a new perspective on how vision, language, and action can be connected for efficient robotic manipulation. Code is available at https://github.com/H-EmbodVis/TurboVLA.
Chinese Translation
视觉-语言-动作(VLA)模型通常采用以大型语言模型(LLM)为中心的$V o L o A$路径,其中视觉观测被投影到大型语言模型的表示空间中,然后解码为机器人动作。尽管这种设计有效,但在每次策略调用时都会产生大量的计算和内存开销。在本研究中,我们介绍了TurboVLA,一种新的VLA范式,它将传统的$V o L o A$路径重新构造成直接的$V + L o A$映射。TurboVLA不再使用大型语言模型作为感知与动作之间的中心接口,而是独立编码视觉观测和语言指令,通过轻量级的双向视觉-语言交互直接交换信息,并使用紧凑的解码器预测连续的动作片段。这种简单的设计直接从视觉和语言特征构建任务条件的表示,显著降低了VLA推理的计算和内存成本。在LIBERO上,TurboVLA以仅0.2B参数实现了97.7%的平均成功率,推理延迟为31.2毫秒,推理显存为0.9 GB,运行在消费级的RTX 4090上,性能与显著更大的VLA策略相当或更优。这些结果确立了TurboVLA作为当前以LLM为中心的VLA范式的简单有效替代方案,为视觉、语言和动作如何连接以实现高效的机器人操作提供了新的视角。代码可在https://github.com/H-EmbodVis/TurboVLA获取。
人工智能 (Artificial Intelligence)
32
cs.AI / 1 / 2607.26119

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

探究推理表现的起源:强化学习与监督微调模型在数学问题解决中的表征质量
Rahman, Antyabha, Gurugubelli, Akshaj, Ankit, Omar, Zhu, Kevin, Balwani, Aishwarya
Abstract
Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterparts on mathematical reasoning tasks; Yet the mechanistic basis for this advantage remains unclear. We therefore ask, what internal representational differences enable RL models' superior performance? Our work presents two converging lines of evidence: First, linear probes trained on layer-wise hidden states reveal that RL models tend to achieve higher accuracy in predicting answer correctness compared to SFT models, indicating more linearly separable and structured representations. Second, mean ablation studies show that RL models develop a hierarchical architecture where deeper layers become progressively more critical, whereas SFT models distribute importance uniformly across layers. Together, these findings demonstrate that RL training fundamentally restructures how models represent and process reasoning problems. Finally, we analyze token-count variability under repeated sampling across problems to assess adaptive compute allocation. While we observe higher variability in some RL-tuned models than in their SFT counterparts, we see strong consistency in others, suggesting that token allocation may depend more on the overall training pipeline than on RL versus SFT alone. We believe this token-allocation variability reveals the spread of plausible on-policy reasoning, highlighting which models exhibit stable policies versus those that are under-determined, potentially non-identifiable solution behaviour.
Chinese Translation
通过强化学习(RL)训练的大型推理模型在数学推理任务中逐渐显示出优于其监督微调(SFT)模型的表现;然而,这种优势的机制基础仍不清晰。因此,我们提出了一个问题:是什么内部表征差异使得RL模型的表现更为出色?我们的研究提供了两条趋同的证据:首先,对逐层隐藏状态进行线性探测的结果表明,RL模型在预测答案正确性方面的准确率往往高于SFT模型,表明其具有更线性可分和结构化的表征。其次,均值消融研究显示,RL模型发展出一种层次架构,其中更深层次的结构变得越来越关键,而SFT模型则在各层之间均匀分配重要性。这些发现共同表明,RL训练从根本上重构了模型表征和处理推理问题的方式。最后,我们分析了在不同问题下重复采样的标记计数变异性,以评估自适应计算分配。尽管我们观察到一些RL调优模型的变异性高于其SFT对应模型,但在其他模型中则表现出强一致性,这表明标记分配可能更多依赖于整体训练流程,而不仅仅是RL与SFT之间的差异。我们认为这种标记分配的变异性揭示了合理的在政策推理的分布,突显出哪些模型表现出稳定的策略,哪些则表现出欠确定性,可能是不可识别的解决行为。
cs.AI / 2 / 2607.26120

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

更大的欺骗:混合动机LLM多智能体系统中的目标不一致性
Fauchard, Marylou, Carichon, Florian, Carvalho, Margarida, Farnadi, Golnoosh
Abstract
Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric information and strategic deception due to conflicting or hidden objectives. In these settings, misalignment with collective goals becomes a central concern. We propose a novel framework for evaluating objective misalignment using the social deduction game Werewolf, modifying the objective of a single agent while preserving its assigned role. Across LLMs from four different model families and sizes, four player roles, and three objective formulations, we introduce a dual analysis of the agents' internal reasoning and their public cheap-talk behavior (i.e costless, non-binding communication that does not directly affect the agents' utilities), complemented by an analysis of game outcomes. Our results show that objective misalignment undermines outcomes in inherently adversarial environments, an effect exacerbated by asymmetric information and specialized roles. While compromised agents consistently develop distinct objective-dependent reasoning strategies, these adaptations remain largely invisible in their public behavior. More broadly, our findings suggest that even subtle objective misalignment can profoundly affect collective decision-making, highlighting the need for effective mitigation strategies for LLM-based multi-agent systems.
Chinese Translation
基于大型语言模型(LLMs)的多智能体系统越来越多地应用于混合动机环境中,在这些环境中,智能体在信息不对称和战略性欺骗的情况下运作,因为存在冲突或隐藏的目标。在这些背景下,与集体目标的不一致性成为一个核心问题。我们提出了一个新颖的框架,通过社交推理游戏《狼人杀》(Werewolf)来评估目标不一致性,修改单个智能体的目标,同时保留其分配的角色。在来自四个不同模型家族和规模的LLMs中,四个玩家角色,以及三种目标表述中,我们引入了对智能体内部推理和其公共低成本交流行为(即无成本、非约束性的交流,不直接影响智能体效用)的双重分析,并辅以对游戏结果的分析。我们的结果表明,目标不一致性削弱了本质上对抗性环境中的结果,这一影响因信息不对称和专业角色而加剧。尽管受损的智能体始终发展出明显依赖目标的推理策略,但这些适应在其公共行为中仍然大多不可见。更广泛地说,我们的发现表明,即使是微妙的目标不一致性也能深刻影响集体决策,强调了对基于LLM的多智能体系统有效缓解策略的需求。
cs.AI / 3 / 2607.26155

ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science

ClinLens:面向长期编码代理的纵向多模态临床数据科学
Zhu, Yuan, Liu, Ethan B., Nie, Frank, Han, Jindong
Abstract
Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely isolate medical question answering, structured-table reasoning, or generic scientific repositories. We introduce CLINLENS, a benchmark of 200 executable tasks over five linked MIMIC resources spanning structured electronic health records, notes, electrocardiograms, chest radiographs, and echocardiograms. A 4 x 5 taxonomy crosses four patient-time scopes with five analysis capabilities. Program-first reverse synthesis pairs each bounded semi-raw package with an evaluator-private reference workflow and checks required artifacts, cohort and temporal semantics, and the final answer. On a fixed 126-task suite, the strongest of 24 standardized model-scaffold configurations achieves 56.3% scope-macro STRICTPASS despite 100% EXECSUCCESS. For reference, a separately configured coding agent solves 83 of 126 tasks, while five biomedical systems adapted to GPT-4o-mini reach at most 2.9% scope-macro STRICTPASS. These results expose a substantial gap between runnable submissions and correct clinical analyses.
Chinese Translation
临床数据科学代理必须将异构的纵向记录转化为可审计的分析,然而现有基准大多孤立地处理医学问答、结构化表格推理或通用科学库。我们介绍了CLINLENS,这是一个基于五个关联的MIMIC资源的200个可执行任务的基准,这些资源涵盖了结构化电子健康记录、病历、心电图、胸部X光片和超声心动图。一个4 x 5的分类法交叉了四种患者时间范围与五种分析能力。程序优先的反向合成将每个有限的半原始包与评估者私有的参考工作流配对,并检查所需的工件、队列和时间语义以及最终答案。在一个固定的126任务套件中,24种标准化模型框架配置中最强的一个实现了56.3%的范围宏严格通过率(STRICTPASS),尽管执行成功率(EXECSUCCESS)为100%。作为参考,一个单独配置的编码代理解决了126个任务中的83个,而五个适应于GPT-4o-mini的生物医学系统最多仅达到2.9%的范围宏严格通过率。这些结果揭示了可运行提交与正确临床分析之间的显著差距。
cs.AI / 4 / 2607.26159

When benchmark inferences do not compose: Projectibility in AI evaluation

当基准推断无法组合时:人工智能评估中的可推广性
Reynolds, Brett
Abstract
An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper identifies a further epistemic problem: warranted links don't automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent. Projectibility concerns whether a bounded extension from observed to unobserved cases is warranted. Goodman supplies the problem of rival extensions; argument-based validity supplies an architecture for testing them. The paper's distinctive claim is a non-composition principle: support for adjacent projections warrants their composition only when endpoints and assumptions align and dependence and uncertainty are carried through. A legal-research case shows how benchmark evidence and a deployment study can each be sound while remaining parallel. A reanalysis and simulation show why aggregate stability can erase distinctions a later projection requires. The resulting projectibility audit diagnoses unsupported joins in benchmark-to-use arguments.
Chinese Translation
人工智能基准结果通常不会在一步之内达到重要的结论。评估者将其推广到其他案例,将其解释为能力的证据,将其外推到新任务,将其转移到另一个系统或地点,并结合关于人类审查和后续后果的假设。以有效性为中心的方法要求对每个主张提供证据。本文识别出一个进一步的认识论问题:有根据的联系并不自动形成有根据的链条。一个研究的目标可能不是下一个研究的来源;系统、群体、结果或条件可能在接口处发生变化;共享的数据或模型谱系可能使看似独立的支持变得依赖。可推广性关注从观察到未观察案例的有限扩展是否是合理的。古德曼提出了竞争扩展的问题;基于论证的有效性提供了测试它们的架构。本文的独特主张是一个非组合原则:对相邻投影的支持仅在端点和假设一致且依赖性和不确定性得到传递时,才保证其组合。一个法律研究案例展示了基准证据和部署研究如何各自是合理的,同时保持平行关系。重新分析和模拟显示了为什么聚合稳定性可能会抹去后续投影所需的区分。最终的可推广性审计诊断了基准到使用论证中的不支持连接。
cs.AI / 5 / 2607.26160

GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning

GuideSkill:为基于指南的临床推理发展可执行的LLM代理技能
Cao, Lang, Shen, Yuhao, Luo, Tianyang, Du, Simo, Peng, Hao, Guo, Yue
Abstract
Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather than execute its rules. We introduce GuideSkill, an external reasoning layer that compiles disease-specific criteria into executable functions returning ordinal diagnostic-support scores. GuideSkill-Zero is initialized from guidelines, while GuideSkill-Evo uses case--diagnosis pairs to refine covered skills and add missing diagnoses. At inference, an LLM proposes a differential diagnosis, grounds the features required by each matched skill, and fuses its ranking with the executed skill scores. Across four benchmarks and four backbones, GuideSkill-Zero improves macro-average accuracy over guideline RAG by 13.45% on average. GuideSkill-Evo achieves the highest macro-average for every backbone, improves over direct inference by 18.49% relatively, and increases gold-label skill coverage from 56.5% to 99.5%. On Qwen3.5-9B, it also exceeds the strongest parameter-update baseline by 11.16% without updating the backbone. Expert evaluation further indicates that GuideSkill produces clinically sound and broadly acceptable skills, suggesting that its initialized and evolved rules are reliable and practically meaningful. These results support executable skills as a model-agnostic mechanism for combining guideline-derived procedures with case-derived diagnostic patterns.
Chinese Translation
临床实践指南(CPGs)编码了诊断标准,但LLM系统通常通过检索指南文本或通过训练吸收其内容,而不是执行其规则。我们介绍了GuideSkill,一个外部推理层,它将特定疾病的标准编译成可执行的函数,返回有序的诊断支持分数。GuideSkill-Zero从指南初始化,而GuideSkill-Evo使用病例-诊断对来细化覆盖的技能并添加缺失的诊断。在推理过程中,LLM提出差异诊断,确定每个匹配技能所需的特征,并将其排名与执行的技能分数融合。在四个基准测试和四个基础模型上,GuideSkill-Zero的宏平均准确率比指南RAG平均提高了13.45%。GuideSkill-Evo在每个基础模型上都达到了最高的宏平均,相对提高了18.49%,并将金标签技能覆盖率从56.5%提高到99.5%。在Qwen3.5-9B上,它还超越了最强的参数更新基线11.16%,而无需更新基础模型。专家评估进一步表明,GuideSkill产生了临床上合理且广泛可接受的技能,表明其初始化和演化的规则是可靠且具有实际意义的。这些结果支持可执行技能作为一种模型无关的机制,用于将基于指南的程序与基于病例的诊断模式相结合。
cs.AI / 6 / 2607.26181

GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure

GoGoTB:基于规范的覆盖闭合的自主RTL验证
Xin, Xin, Lou, Jincheng, Li, Junhui, Yan, Jinglin, Xiao, Panda, Wu, Di, Li, Haixiao, Lu, Weicong, Fan, Weijian, Qu, Xinyu, Zhao, Yuxiang, Yu, Min, Di, Zhixiong, Lin, Yibo
Abstract
Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can trigger a costly respin. Recent large language models (LLMs) offer new opportunities to automate this process, yet existing LLM-based approaches generate each component through independent single-turn calls with no shared context, leaving interface mismatches undetected and reported coverage disconnected from specification requirements. To address these challenges, we present GoGoTB, an agentic framework that achieves end-to-end verification closure through three subsystems: an agentic execution control layer, an evolvable knowledge system, and specification-grounded coverage closure. The execution control layer separates deterministic enforcement from LLM reasoning at every tool and stage boundary. The knowledge system dispatches methodology and design-specific expertise on demand. The coverage framework anchors every bin to a named specification behavior so that each residual gap has a diagnosable root cause and a targeted remedy. Tested on 8 register transfer level (RTL) designs without any human intervention, GoGoTB achieves 100\% environment generation success and averages 98.4\% line, 97.2\% branch, 97.0\% toggle, and 83.2\% functional coverage. No prior work successfully generates a complete verification environment or achieves meaningful coverage on the same benchmarks.
Chinese Translation
功能验证在集成电路(IC)前端工程中占据主导地位,任何一个未被发现的错误如果逃入硅片都可能导致昂贵的重制。近期的大型语言模型(LLMs)为自动化这一过程提供了新的机会,但现有的基于LLM的方法通过独立的单轮调用生成每个组件,没有共享上下文,导致接口不匹配未被检测到,报告的覆盖与规范要求脱节。为了解决这些挑战,我们提出了GoGoTB,一个自主框架,通过三个子系统实现端到端的验证闭合:自主执行控制层、可演化知识系统和基于规范的覆盖闭合。执行控制层在每个工具和阶段边界上将确定性强制与LLM推理分离。知识系统按需调度方法论和设计特定的专业知识。覆盖框架将每个区间锚定到一个命名的规范行为,以便每个剩余的差距都有可诊断的根本原因和针对性的补救措施。在没有任何人工干预的情况下,对8个寄存器传输级(RTL)设计进行测试,GoGoTB实现了100%的环境生成成功率,平均达到了98.4%的行覆盖率、97.2%的分支覆盖率、97.0%的切换覆盖率和83.2%的功能覆盖率。没有任何先前的工作能够在相同的基准测试上成功生成完整的验证环境或实现有意义的覆盖。
cs.AI / 7 / 2607.26191

Position: Evaluation Scores Are Perishable Knowledge Claims

位置:评估分数是易逝的知识主张
Gilda, Sankalp, Gilda, Shlok
Abstract
Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results. When these signals are aggregated via averaging, evaluation confidence can then substantially exceed the reliability of the weakest signal: a phenomenon we call trust inflation in evaluation. We argue that evaluation scores should be treated as epistemic claims with three properties: formality (human evaluation provides stronger evidence than an automated metric), scope (a benchmark result applies to the tested distribution, not universally), and validity windows (benchmark results expire as contamination accumulates and distributions shift). Several converging research traditions (chain-of-thought analysis, possibilistic logic, and algebraic theory) establish weakest-link aggregation as the conservative endpoint of a parameterized operator family controlled by a single pessimism parameter. Drawing on those traditions, and on concrete lessons from building an evaluation harness for agentic AI, we propose that evaluation results carry explicit metadata (formality tier, scope declaration, and expiration date) to make their epistemic status transparent. We illustrate the cost of mean aggregation on the public HELM leaderboard: across 54 frontier models on ten scenarios, the top-five models ranked by mean score and by weakest-link are completely disjoint.
Chinese Translation
语言模型的评估方法越来越多地结合了多种信号,从自动化指标和大型语言模型(LLM)作为评判者的评分到人类评估和基准套件结果。当这些信号通过平均值聚合时,评估信心可能会显著超过最弱信号的可靠性:我们称之为评估中的信任膨胀现象。我们认为,评估分数应被视为具有三种特性的认识论主张:形式性(人类评估提供的证据比自动化指标更强),范围(基准结果适用于被测试的分布,而非普遍适用),以及有效性窗口(随着污染的积累和分布的变化,基准结果会失效)。几种交汇的研究传统(思维链分析、可能性逻辑和代数理论)确立了最弱环聚合作为由单一悲观参数控制的参数化算子族的保守终点。基于这些传统,以及从构建代理人工智能评估工具中获得的具体经验教训,我们建议评估结果携带明确的元数据(形式性等级、范围声明和到期日期),以使其认识论状态透明。我们通过公共HELM排行榜说明均值聚合的成本:在十个场景中的54个前沿模型中,按均值评分和最弱环排名的前五个模型完全不重合。
cs.AI / 8 / 2607.26307

TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning

TraceCoder:具有可解释性和可审计性的代码生成与位置-关键片段版本控制
Alssadi, Rwaida, Syed, Muntaser, Kasula, Balaji, Deen, Lamine, Alotaibi, Majed, Alghamdi, Mohammed, Ton, Tyler, Alqarni, Ali, Silaghi, Marius
Abstract
Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through benchmark-driven repair is ephemeral, and post-hoc auditing is impossible. We present a code generation concept that addresses these shortcomings through three complementary mechanisms: (i) a relational snippet-history schema that records, per repair event, the benchmark reference, round number, failure text, and LLM explanation, enabling full provenance queries; (ii) a browser-based visualisation tool that renders this history as heat-mapped, hover-annotated source code; and (iii) a competitive fractional position-key indexing scheme with tree-node delimiters that assigns stable, lexicographically-ordered identifiers to each code snippet, enabling fine-grained tracking without disrupting surrounding lines. We evaluate TraceCoder on 30 algorithmic programming tasks spanning string processing, mathematical computation, and data-structure manipulation, across two provider configurations. Of these, 10 exhaust the 6-iteration budget on tasks with subtle edge-case behaviour. Mean Chg% reaches 30%, three in ten code snippets carry a traceable repair-event row, compared to 21% when using Gemini 2.0 Flash as sole provider on a 20-task subset. Three detailed case studies demonstrate how the system explains which specific benchmark failures shaped each line of the final program. The proposed mechanism makes the internal "narrative" of automated code generation auditable and replayable, a property essential for trust and accountability in production deployments.
Chinese Translation
当代基于大语言模型(LLM)的编码代理生成的代码是黑箱输出:每一行背后的推理被隐藏,通过基准驱动的修复过程代码的演变是短暂的,事后审计是不可能的。我们提出了一种代码生成概念,通过三种互补机制解决这些不足:(i)一个关系片段历史架构,记录每个修复事件的基准参考、轮次编号、失败文本和LLM解释,支持全面的来源查询;(ii)一个基于浏览器的可视化工具,将该历史呈现为热图标注的源代码;(iii)一个具有树节点分隔符的竞争性分数位置-关键索引方案,为每个代码片段分配稳定的、按字典顺序排列的标识符,实现精细跟踪而不干扰周围行。我们在30个算法编程任务上评估了TraceCoder,这些任务涵盖字符串处理、数学计算和数据结构操作,涉及两种提供者配置。其中,10个任务在具有微妙边缘案例行为的任务上耗尽了6次迭代预算。平均变化百分比达到30%,每十个代码片段中有三个携带可追溯的修复事件行,而在仅使用Gemini 2.0 Flash作为提供者的20个任务子集时,这一比例为21%。三个详细的案例研究展示了系统如何解释哪些特定的基准失败影响了最终程序的每一行。所提出的机制使得自动代码生成的内部“叙事”可审计和可重放,这一特性对于生产部署中的信任和问责至关重要。
cs.AI / 9 / 2607.26367

Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

探索物理问题中的结构:人工智能代理能否发现统计力学映射?
Zhao, Wanyu, Zhao, Wanbing
Abstract
An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model. We study this skill as an AI-agent task: can LLM-based agents discover statistical mechanical mappings from a raw partition function to a tractable representation? To probe this question, we introduce StatMechBench-v0, a benchmark of six Ising-type problems covering transfer-matrix methods, gauge-removable disorder, and planar/Pfaffian structure. We evaluate a simple propose-verify-revise agent across multiple LLMs and problem phrasings. The results show that numerical feedback often helps agents repair code and recover correct partition functions. However, agents can also pass the numerical checks while misidentifying the underlying tractable class or understating computational complexity. This both reveals limitations in current LLM reasoning and calls for a verification stack that goes beyond numerical agreement, incorporating, for example, symbolic checks and structural invariants. Our study provides an early evaluation and design directions for AI agents aimed at structural discovery in theoretical physics.
Chinese Translation
理论物理中的一项重要技能是识别何时可以将新问题转化为已知模型。我们将这一技能视为人工智能代理的任务:基于大语言模型(LLM)的代理能否从原始配分函数发现统计力学映射到可处理的表示?为探讨这一问题,我们引入了StatMechBench-v0,这是一个涵盖转移矩阵方法、可移除的无序和平面/帕法结构的六个伊辛型问题的基准测试。我们在多个LLM和问题表述中评估了一个简单的提出-验证-修订代理。结果表明,数值反馈通常有助于代理修复代码并恢复正确的配分函数。然而,代理也可能在错误识别潜在的可处理类别或低估计算复杂度的情况下通过数值检查。这既揭示了当前LLM推理的局限性,也呼吁建立一个超越数值一致性的验证体系,例如,纳入符号检查和结构不变性。我们的研究为旨在理论物理中进行结构发现的人工智能代理提供了早期评估和设计方向。
cs.AI / 10 / 2607.26393

CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games

CaM-Wolf:具有因果意识的多模态代理在社交推理游戏中的应用
Zhang, Zheng, Yao, Nanjie, He, Jiarui, Ye, Deheng, Zhao, Peilin, Wang, Hao
Abstract
Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents. These games require complex social skills such as reasoning, deception, and collaboration. While recent advances in large language models (LLMs) have driven significant progress in SDG agents, current approaches are predominantly text-based, overlooking the multimodal nature that is fundamental to human social interaction. To bridge this gap, we introduce CaM-Wolf, the first SDG agent that integrates multimodal perception and generation. CaM-Wolf processes video inputs from other players, employs a causal-aware Reasoner trained via reinforcement learning to establish logical chains between observable behaviors and hidden roles, and presents itself through an animated avatar. Our experiments and user study show that CaM-Wolf achieves superior agent gameplay performance and enhances the quality of human-AI interaction. This work represents a significant advancement towards creating more human-like AI agents capable of participating in nuanced social dynamics. Our code is available at https://3dagentworld.github.io/avatar_wolf.
Chinese Translation
社交推理游戏(SDGs),如狼人游戏,已成为人工智能代理的挑战性测试平台。这些游戏需要复杂的社交技能,如推理、欺骗和协作。尽管最近大型语言模型(LLMs)的进展推动了SDG代理的显著发展,但当前的方法主要基于文本,忽视了人类社交互动中根本的多模态特性。为了解决这一问题,我们提出了CaM-Wolf,这是第一个集成多模态感知和生成的SDG代理。CaM-Wolf处理来自其他玩家的视频输入,利用通过强化学习训练的因果意识推理器建立可观察行为与隐藏角色之间的逻辑链,并通过动画化的虚拟形象展示自己。我们的实验和用户研究表明,CaM-Wolf在代理游戏表现上优于其他代理,并提升了人机交互的质量。这项工作代表了朝着创建更具人类特征的人工智能代理,能够参与复杂社交动态的重要进展。我们的代码可在 https://3dagentworld.github.io/avatar_wolf 获取。
cs.AI / 11 / 2607.26452

CG-World: A Large-Scale World-State Dataset and Protocol for World Models

CG-World:一个大规模世界状态数据集及其世界模型协议
Cai, Yiming, Yu, Fangjie, Yu, Meiqing, Shi, Ziyue, Yuan, Pengfei, Guo, Yong
Abstract
World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually capture only part of this structure. We introduce CG-World, a large-scale world-state dataset and protocol derived from industrial computer graphics production pipelines. CG-World explicitly records intermediate states, including multimodal semantics, spatial structure, skeletal and controller states, motion curves, camera and lighting parameters, physics caches, contact events, and multi-pass renderings. CG-World v1 contains approximately 850,000 temporally aligned segments of 1-5 seconds. It separates latent states, observations, relations, events, and branch metadata, and organizes them into unified spatiotemporal samples. To support intervention learning and counterfactual reasoning, CG-World defines a branch lineage covering factual trajectories, observation interventions, action interventions, mechanism interventions, and strict counterfactual branches, with intervention targets, invariants, and alternative outcomes explicitly recorded. We evaluate the dataset on geometry-conditioned video generation, action prediction, and closed-loop vision-language-action policy transfer. Results show that CG-World provides reusable structured supervision for controlled generation, action modeling, and embodied policy transfer. We plan to expand CG-World through continued data collection and community collaboration toward a shared data infrastructure for world models, Physical AI, and embodied intelligence.
Chinese Translation
世界模型必须学习状态、动作、事件和观察的联合动态,但现有的视频、机器人技术和仿真数据集通常只捕捉到这一结构的一部分。我们介绍了CG-World,这是一个源自工业计算机图形制作流程的大规模世界状态数据集和协议。CG-World 明确记录了中间状态,包括多模态语义、空间结构、骨骼和控制器状态、运动曲线、相机和照明参数、物理缓存、接触事件以及多通道渲染。CG-World v1 包含大约 850,000 个时间对齐的 1-5 秒段。它将潜在状态、观察、关系、事件和分支元数据分开,并将其组织成统一的时空样本。为了支持干预学习和反事实推理,CG-World 定义了一个覆盖事实轨迹、观察干预、动作干预、机制干预和严格反事实分支的分支谱系,并明确记录干预目标、不变性和替代结果。我们在几何条件的视频生成、动作预测和闭环视觉-语言-动作策略转移上评估了该数据集。结果表明,CG-World 为受控生成、动作建模和具身策略转移提供了可重用的结构化监督。我们计划通过持续的数据收集和社区合作来扩展 CG-World,以实现世界模型、物理人工智能和具身智能的共享数据基础设施。
cs.AI / 12 / 2607.26465

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

MultivationBench:多模态序列动机推理基准
Chung, Kawai, Chan, Chunkit, Yim, Yauwai, Liu, Yuxuan, Shi, Haochen, Wang, Weiqi, Zong, Qing, Zheng, Tianshi, Fu, Yixuan, Wong, Kai Chung, Liang, Hao, Gao, Yifan, Yang, Xi, Hsiao, Janet Hui-wen, Song, Yangqiu
Abstract
Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. To address this gap, we introduce MultivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story-driven visual narratives. The benchmark builds upon established psychological frameworks - Maslow's hierarchy and Reiss's basic desires - and requires models to integrate accumulated multimodal context to infer evolving motivations. Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.
Chinese Translation
多模态大型语言模型因其在社会智能方面的潜力而引起了广泛关注;然而,它们在进行序列动机推理方面的能力仍然研究不足。现有评估主要考察静态文本或孤立的视觉快照,这并不能反映现实世界行为驱动因素的累积特性。为了解决这一问题,我们引入了MultivationBench,一个旨在严格评估故事驱动视觉叙事中多模态动机推理的基准。该基准基于已建立的心理学框架——马斯洛的需求层次理论和瑞斯的基本欲望,要求模型整合累积的多模态上下文以推断不断变化的动机。结果表明,MultivationBench提出了一个重大挑战:所有测试模型在序列上下文中都难以保持一致的动机推理,揭示了静态识别能力与人类社会理解所需的动态推理之间的关键脱节。
cs.AI / 13 / 2607.26490

EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks

EvoPINN:可执行算法的自主发现框架用于物理信息神经网络
Yin, Peng, Li, Kai, Zhang, Yifan, Cheng, Jian
Abstract
Physics-informed neural networks (PINNs) have emerged as a powerful paradigm for solving partial differential equations (PDEs), yet their performance heavily relies on the manual, trial-and-error engineering of neural representations, loss formulations, and optimization dynamics. While Large Language Models (LLMs) offer a promising avenue for automated design, unconstrained code generation often yields mathematically invalid or numerically unstable solutions under strict scientific computing constraints. To bridge this gap, we propose \textbf{EvoPINN}, an agentic framework that reformulates PINN development from labor-intensive manual design into a rigorous, execution-grounded algorithm discovery problem. EvoPINN navigates a modular search space by decoupling neural representations from training programs, utilizing an LLM agent to iteratively propose memory-conditioned programmatic modifications. To ensure scientific validity, all candidates undergo strict structural verification and budget-matched PDE evaluation. Extensive experiments across diverse PDE regimes (oscillatory, elliptic, dissipative, and nonlinear transport) demonstrate that EvoPINN discovers PDE-specialized learning algorithms that significantly reduce relative $L_{2}$ error compared to baselines. Crucially, EvoPINN autonomously invented SLRC-PINN, a novel architecture whose performance gains persist under rigorous parameter-matched comparisons, establishing the viability of execution-grounded agents for discovering genuinely new scientific computing mechanisms.
Chinese Translation
物理信息神经网络(PINNs)已成为解决偏微分方程(PDEs)的强大范式,但其性能在很大程度上依赖于神经表示、损失公式和优化动态的手动试错工程。尽管大型语言模型(LLMs)为自动化设计提供了有希望的途径,但在严格的科学计算约束下,无限制的代码生成往往会产生数学上无效或数值上不稳定的解决方案。为了解决这一问题,我们提出了 extbf{EvoPINN},一个将PINN开发从劳动密集型的手动设计重构为一个严格的、基于执行的算法发现问题的自主框架。EvoPINN通过将神经表示与训练程序解耦,导航一个模块化的搜索空间,利用LLM代理迭代提出基于记忆条件的程序修改。为了确保科学有效性,所有候选者都经过严格的结构验证和预算匹配的PDE评估。在多种PDE领域(振荡性、椭圆性、耗散性和非线性传输)进行的广泛实验表明,EvoPINN发现了PDE专用的学习算法,相较于基线显著降低了相对$L_{2}$误差。重要的是,EvoPINN自主发明了SLRC-PINN,这是一种新颖的架构,其性能提升在严格的参数匹配比较中依然保持,确立了基于执行的代理在发现真正新的科学计算机制方面的可行性。
cs.AI / 14 / 2607.26512

Evidence-Ledger Adjudication for Claim-Evidence Traceability

基于证据账本的索赔-证据可追溯性裁定
Chen, Gengyu, Yu, Yongjie, Wang, Weiling
Abstract
AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them. We study evidence-ledger adjudication: a claim-evidence traceability workflow that pairs each claim with an evidence packet, assigns a support relation, and routes unsupported, contradicted, or mixed-evidence claims back to the author. The empirical core is a 2,335-row blind benchmark built from independent external labels in AVeriTeC, CLIMATE-FEVER, and SciFact. Gold relations and source evidence labels are hidden during prediction and joined only for scoring. On this benchmark, the agent evidence-ledger condition achieves 0.676 relation accuracy and 0.601 macro-F1, compared with 0.383 accuracy and 0.303 macro-F1 for the best non-agent baseline. It also routes 1270/1435 claims whose gold labels indicate contradiction, missing evidence, or mixed evidence, while routing 295/900 supported claims. These results show that evidence-ledger adjudication can turn heterogeneous evidence packets into an auditable traceability layer for AI-assisted writing.
Chinese Translation
人工智能代理可以比作者更快地起草索赔,而作者则需要检查引用或检索的证据是否支持这些索赔。我们研究了基于证据账本的裁定:一种索赔-证据可追溯性工作流程,该流程将每个索赔与一个证据包配对,分配支持关系,并将不支持、相互矛盾或证据混杂的索赔返回给作者。实证核心是一个由 AVeriTeC、CLIMATE-FEVER 和 SciFact 中的独立外部标签构建的 2,335 行盲基准。在预测过程中,金标准关系和源证据标签被隐藏,仅在评分时结合。在这个基准上,代理证据账本条件达到了 0.676 的关系准确率和 0.601 的宏观 F1 值,而最佳非代理基线的准确率为 0.383,宏观 F1 值为 0.303。它还路由了 1270/1435 个金标签指示矛盾、缺失证据或混合证据的索赔,同时路由了 295/900 个支持的索赔。这些结果表明,基于证据账本的裁定可以将异构证据包转化为一个可审计的可追溯性层,以支持人工智能辅助写作。
cs.AI / 15 / 2607.26588

Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models

Eco3S:基于代理模型的复杂社会经济系统模拟
Wei, Shaopeng, Cheng, Yufei, Sun, Wenxi, Ding, Yepeng, Zhao, Yu, Kou, Gang
Abstract
The rapid development of large language models (LLMs) has renewed interest in agent-based modeling (ABM). However, current LLM-based ABM research faces several key challenges: modeling evolving agent-environment interactions, enabling flexible counterfactual reasoning, and automating simulation workflows for scientific research. In this paper, we propose Eco3S, a socio-economic system simulation framework for economic research and policy analysis that addresses these challenges through three key mechanisms: (1) Co-evolving Environment Design, a bidirectional feedback loop where agents and the environment co-evolve, producing realistic emergent behaviors; (2) Structural Causal Simulation, a structural causal model (SCM)-inspired counterfactual mechanism that allows flexible interventions for diverse causal inference tasks; (3) Simulation-Analysis-Refinement Paradigm, a self-corrective mechanism that iteratively refines experimental designs based on prior simulation results. Experiments on diverse economic scenarios confirm \textit{Eco3S}'s effectiveness in replicating multiple established economic studies (canal decay, origins of governance, and information propagation) and phenomena across domains. Additional results further demonstrate its scalability and generalizability, highlighting the framework's potential for rigorous economic research and policy-making.
Chinese Translation
大型语言模型(LLMs)的快速发展重新引发了对基于代理建模(ABM)的兴趣。然而,目前基于LLM的ABM研究面临几个关键挑战:建模不断演变的代理-环境交互、实现灵活的反事实推理以及自动化科学研究的模拟工作流程。本文提出了Eco3S,一个用于经济研究和政策分析的社会经济系统模拟框架,通过三个关键机制解决这些挑战:(1)共演化环境设计,一种双向反馈循环,其中代理与环境共同演化,产生现实的涌现行为;(2)结构因果模拟,一种受结构因果模型(SCM)启发的反事实机制,允许对多种因果推断任务进行灵活干预;(3)模拟-分析-精炼范式,一种自我纠正机制,基于先前模拟结果迭代优化实验设计。在多种经济情境下的实验验证了Eco3S在复制多个已建立的经济研究(如运河衰退、治理起源和信息传播)及跨领域现象方面的有效性。额外结果进一步展示了其可扩展性和普适性,突显了该框架在严谨的经济研究和政策制定中的潜力。
cs.AI / 16 / 2607.26611

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

更少的澄清,更好的代码:跨会话个性化歧义适应在编码助手中的基准测试
Xu, Zijian, Zhang, Wenshuo, Qin, Zisen, Sheng, Rui, Sun, Yushi, Qu, Huamin, Shi, Chuhan
Abstract
AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in user-specific ways across tasks and sessions. Existing disambiguation methods typically address each ambiguous request in isolation within the current coding session, often through eliciting additional clarification. However, whether resolved session history from the same user can serve as memory for resolving recurring personalized ambiguity in a newly opened session remains underexplored. We formulate personalized ambiguity adaptation as a new task: given a user's previously resolved coding sessions and a new ambiguous request, an assistant should identify the recurring ambiguity pattern, produce the intended executable solution, and minimize clarification. To benchmark this task, we introduce CAPA, which characterizes personalized coding ambiguity through six mechanisms and injects these mechanisms into unambiguous executable tasks using a controlled three-stage generation pipeline. CAPA contains 600 coding sessions across 60 balanced user--ambiguity cells, including 300 held-out evaluation sessions. We evaluate 12 recent LLMs under no-history and same-user-history conditions using executable success, first-turn success, and turns-to-completion. Our analyses examine task difficulty, user identity, and memory-based history use, and we further propose same-user history gating as a lightweight inference-time method. CAPA provides a foundation for developing long-term coding assistants that better align generated code with user intent while reducing repeated clarification.
Chinese Translation
AI辅助编码日益将非正式用户意图转化为可执行软件,然而编码请求中常常包含在任务和会话中以用户特定方式反复出现的歧义。现有的消歧方法通常在当前编码会话中孤立地处理每个模糊请求,通常通过引导额外的澄清来实现。然而,来自同一用户的已解决会话历史是否可以作为记忆来解决新开启会话中反复出现的个性化歧义仍然未被充分探讨。我们将个性化歧义适应制定为一项新任务:给定用户之前解决的编码会话和一个新的模糊请求,助手应识别反复出现的歧义模式,生成预期的可执行解决方案,并最小化澄清。为了对该任务进行基准测试,我们引入了CAPA,它通过六种机制来表征个性化编码歧义,并通过受控的三阶段生成管道将这些机制注入到无歧义的可执行任务中。CAPA包含600个编码会话,分布在60个平衡的用户-歧义单元中,包括300个保留的评估会话。我们在无历史和同用户历史条件下评估了12个最新的LLM,使用可执行成功率、首次成功率和完成所需的回合数。我们的分析考察了任务难度、用户身份和基于记忆的历史使用,并进一步提出了同用户历史门控作为一种轻量级的推理时方法。CAPA为开发长期编码助手提供了基础,使生成的代码更好地与用户意图对齐,同时减少重复的澄清。
cs.AI / 17 / 2607.26642

AlphaSchema: Exploring the Space of Trading Semantics for LLM-Based Alpha Mining

AlphaSchema:探索基于大型语言模型的阿尔法挖掘交易语义空间
Yi, Jingyang, Yang, Jian, Jin, Yifei, Li, Yuqi, Li, Jian
Abstract
Automated alpha mining has increasingly adopted large language model (LLM) agents for factor generation and iterative discovery. However, existing LLM-based systems often delegate both factor construction and search decisions to the agent itself, without an explicit exploration space or a principled mechanism for navigating that space. As a result, exploration remains largely implicit and difficult to control or optimize systematically. We introduce AlphaSchema, which constructs and explores a structured space of trading semantics for alpha mining. Each point in this space is a schema plan composed of Event, Context, Qualities, Direction, and Output, specifying the semantics of a candidate factor before implementation. AlphaSchema decouples exploration from implementation: an LLM translates selected schema plans into executable factors, while evaluated rewards are accumulated to learn a surrogate model over the semantic space. An iterative selection mechanism uses this model to balance global exploration, surrogate-guided exploitation, and local mutation. Experiments on the Chinese stock market show that AlphaSchema discovers factor pools with strong predictive and portfolio performance. Further analyses show that the semantic search process navigates diverse regions while increasingly allocating evaluations toward high-reward regions, and that implementations of the same schema plans by different LLMs exhibit comparable predictive quality, suggesting that alpha mining quality is largely robust to the choice of LLM within our framework.
Chinese Translation
自动化阿尔法挖掘越来越多地采用大型语言模型(LLM)代理进行因子生成和迭代发现。然而,现有的基于LLM的系统通常将因子构建和搜索决策委托给代理本身,而没有明确的探索空间或系统化的导航机制。因此,探索过程在很大程度上是隐性的,难以系统地控制或优化。我们提出了AlphaSchema,它构建并探索一个结构化的交易语义空间用于阿尔法挖掘。该空间中的每个点都是一个由事件(Event)、上下文(Context)、特性(Qualities)、方向(Direction)和输出(Output)组成的模式计划,指定了候选因子在实施之前的语义。AlphaSchema将探索与实施解耦:LLM将选定的模式计划转换为可执行的因子,同时评估的奖励被累积以学习语义空间上的代理模型。迭代选择机制利用该模型在全局探索、代理引导的开发和局部变异之间进行平衡。在中国股市的实验表明,AlphaSchema发现了具有强预测能力和投资组合表现的因子池。进一步分析显示,语义搜索过程在多样化区域中导航,同时越来越多地将评估分配到高奖励区域,并且不同LLM对相同模式计划的实现表现出可比的预测质量,这表明在我们的框架内,阿尔法挖掘的质量在很大程度上对LLM的选择具有鲁棒性。
cs.AI / 18 / 2607.26643

Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

重新思考自我进化:一种约束的探索-开发过程以减轻技能过拟合
Lin, Hongqiang, Liu, Chao, Bai, Xiaofan, Jin, Xuan, Li, Yuhong, Zheng, Nenggan, Cao, Xipeng
Abstract
Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applications. A promising solution is to treat skills as trainable states and optimize them in the same way as model parameters in neural network training. However, data-driven skill optimization is prone to overfitting to the limited trajectories collected from real environments. Overexploiting these trajectories overfits the current batch, while unconstrained exploration causes regression on previously solved cases. This tension motivates a constrained search view of skill self-evolution, governed by an exploration--exploitation trade-off. We propose SkillBoost, a three-stage framework that mitigates both risks: structured exploitation localizes observed failures to editable skill components, prior-guided exploration draws on prior knowledge in the LLM to generate diverse repair candidates, and verified acceptance commits a candidate only when it improves performance within a regression bound. Experiments across 23 model--benchmark configurations show that SkillBoost achieves state-of-the-art performance while mitigating overfitting, outperforming both human-crafted and LLM-generated skills. Transfer experiments further show that optimized skills can be reused by other agents on similar tasks.
Chinese Translation
使大型语言模型(LLM)代理能够积累和重用来自过去交互的经验仍然是现实世界应用中的一个核心挑战。一种有前景的解决方案是将技能视为可训练状态,并以与神经网络训练中的模型参数相同的方式对其进行优化。然而,数据驱动的技能优化容易对从真实环境中收集的有限轨迹产生过拟合。过度利用这些轨迹会导致当前批次的过拟合,而不受限制的探索则会导致对先前解决案例的回归。这种紧张关系促使我们对技能自我进化采取一种约束搜索的视角,由探索-开发的权衡所主导。我们提出了SkillBoost,一个三阶段框架,旨在减轻这两种风险:结构化开发将观察到的失败局限于可编辑的技能组件,基于先前知识的探索利用LLM中的先验知识生成多样化的修复候选,经过验证的接受仅在候选者在回归边界内提高性能时才会被采纳。对23个模型-基准配置的实验表明,SkillBoost在减轻过拟合的同时实现了最先进的性能,超越了人类设计和LLM生成的技能。转移实验进一步表明,优化后的技能可以被其他代理在类似任务中重用。
cs.AI / 19 / 2607.26661

AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution

AgenticCANN:通过知识增强的自主进化实现自动化Ascend C运算符生成
Qiu, Junhao, Wang, Zidong, Sun, Yansong, Ma, Zhitong, Guo, Ping, Zhang, Qingfu
Abstract
Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise.While large language models (LLMs) have shown promise in automated CUDA kernel generation, the fundamentally different programming model of Ascend C introduces unique challenges that remain unexplored. In this paper, we propose AgenticCANN, a knowledge-augmented agentic evolution framework specifically tailored for automated Ascend C operator synthesis in low-corpus NPU environments.To overcome the severe platform knowledge deficit on unfamiliar hardware, AgenticCANN incorporates a knowledge-orchestrated generation system that delivers structured, multi-level domain insights across the development lifecycle to resolve the upstream feasibility bottleneck.Building on this foundation, it features a stage-adaptive agentic evolution strategy that dynamically aligns LLM interaction modes with specific generation and evolution phases, balancing high-exploration candidate discovery with high-convergence performance tuning.Extensive experiments on Huawei Ascend 910B across six operators spanning five pattern categories demonstrate that our method achieves 90 to 100 percent feasibility on elementwise and normalization operators, 56% on fusion operators, and up to 6.65$\times$ speedup on 1B Pangu model inference kernels. Further analysis reveals that knowledge injection monotonically improves feasibility from 57% to 86% on elementwise operators, demonstrating its general rather than operator-specific benefit.
Chinese Translation
Ascend C运算符优化对于NPU(神经处理单元)推理性能至关重要,但需要深厚的硬件专业知识。尽管大型语言模型(LLMs)在自动化CUDA内核生成方面展现了潜力,但Ascend C的根本不同编程模型带来了尚未探索的独特挑战。本文提出了AgenticCANN,一个专门为低语料NPU环境中的自动化Ascend C运算符合成量身定制的知识增强自主进化框架。为了克服对不熟悉硬件的严重平台知识缺口,AgenticCANN结合了一个知识协调生成系统,该系统在开发生命周期中提供结构化的多层次领域见解,以解决上游可行性瓶颈。在此基础上,它具有一个阶段自适应的自主进化策略,动态调整LLM交互模式与特定生成和进化阶段的对齐,平衡高探索候选发现与高收敛性能调优。在华为Ascend 910B上针对六个运算符和五个模式类别进行的广泛实验表明,我们的方法在逐元素和归一化运算符上实现了90%到100%的可行性,在融合运算符上为56%,并在1B Pangu模型推理内核上实现了高达6.65倍的加速。进一步分析显示,知识注入单调地将逐元素运算符的可行性从57%提高到86%,证明了其普遍性而非特定于运算符的益处。
cs.AI / 20 / 2607.26724

UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks

UrbanDS:一种图引导的LLM多代理系统用于数据密集型城市任务
Zhou, Zhilun, Yu, Jianghao, Lin, Yuming, yang, yongjun, Yongquan, Sun, Jin, Depeng, Li, Yong
Abstract
Large language model (LLM) agents have been widely applied in automating data science tasks. However, existing methods typically rely on a limited set of provided datasets, and they face challenges in data-intensive scenarios that require discovering and leveraging relevant information from large-scale and heterogeneous data repositories. Urban tasks are representative examples of such scenarios, as urban data are not only large-scale and multi-sourced, but also exhibit complex spatial, temporal, and semantic relationships. To address these challenges, we propose UrbanDS, a graph-guided LLM multi-agent system for data-intensive urban tasks. We first construct a unified dataset graph to organize reusable dataset skills and the relationships among datasets. Specifically, we develop a Data Profiling Agent that constructs a skill for each dataset. Moreover, a Relation Agent identifies relationships among datasets and integrates these relationships into the dataset graph. At runtime, a Planner Agent retrieves task-relevant datasets from the graph and generates execution plans. Multiple Execution Agents then perform data processing and analysis, while their execution progress and intermediate results are shared through a common memory. Finally, a Report Agent synthesizes the experimental logs into a report, which can be further refined based on user feedback. To systematically evaluate the capability of agents in handling data-intensive urban scenarios, we further construct UrbanDS-Bench, an urban data science benchmark covering representative data analysis and modeling tasks. Experiments on both general and urban benchmarks demonstrate that UrbanDS consistently outperforms existing data science agents on data-intensive tasks. Furthermore, UrbanDS has been deployed on the urban operations platform of Dongxihu District, Wuhan, demonstrating its effectiveness in real-world urban applications.
Chinese Translation
大型语言模型(LLM)代理已广泛应用于自动化数据科学任务。然而,现有方法通常依赖于有限的提供数据集,并且在需要从大规模和异构数据存储库中发现和利用相关信息的数据密集型场景中面临挑战。城市任务是此类场景的典型例子,因为城市数据不仅规模庞大且来源多样,还表现出复杂的空间、时间和语义关系。为了解决这些挑战,我们提出了UrbanDS,一种用于数据密集型城市任务的图引导LLM多代理系统。我们首先构建一个统一的数据集图,以组织可重用的数据集技能及其之间的关系。具体而言,我们开发了一个数据剖析代理(Data Profiling Agent),为每个数据集构建技能。此外,关系代理(Relation Agent)识别数据集之间的关系,并将这些关系整合到数据集图中。在运行时,规划代理(Planner Agent)从图中检索与任务相关的数据集并生成执行计划。多个执行代理(Execution Agents)随后执行数据处理和分析,同时通过公共内存共享它们的执行进度和中间结果。最后,报告代理(Report Agent)将实验日志综合成报告,并可以根据用户反馈进一步完善。为了系统评估代理在处理数据密集型城市场景中的能力,我们进一步构建了UrbanDS-Bench,一个涵盖代表性数据分析和建模任务的城市数据科学基准。对一般基准和城市基准的实验表明,UrbanDS在数据密集型任务上始终优于现有的数据科学代理。此外,UrbanDS已在武汉市东西湖区的城市运营平台上部署,展示了其在现实城市应用中的有效性。
cs.AI / 21 / 2607.26773

Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM

潜在通道真的在沟通吗?潜在多智能体大语言模型的因果审计
Zhang, Huixiang, Emu, Mahzabeen
Abstract
Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information. End-task performance alone also cannot reveal whether an observed effect depends on message presence, content generated for the evaluated example, or information supplied by a separate agent. We introduce a causal audit that applies controlled message replacements at the boundary where the sender-produced representation enters the receiver. Four message settings support five measurements of encoded sender information, receiver sensitivity to message presence and identity, the task value of example-specific content, and the additional value supplied by a separate agent. We apply the audit to latent relay with Qwen3-4B and Qwen3-8B on GSM8K, ARC-C, and MATH-500. On GSM8K, the Qwen3-4B overall performance effect of -1.00 percentage point decomposes into a -6.17-point effect retained by an other-example message and a +5.17-point effect attributable to example-specific content; both component directions reverse at 8B. On MATH-500, the Qwen3-4B gain of 15.00 points comprises 8.33 points retained by an other-example message and 6.67 points attributable to example-specific content, while the 8B gain is dominated by the former component. Self-substitution comparisons further show that example-specific content and other-agent value are distinct. These results show that aggregate accuracy does not identify how a latent message affects the receiver and motivate controlled message comparisons as a standard evaluation for latent communication.
Chinese Translation
基于大语言模型(LLM)的多智能体系统(MAS)中的潜在通信传递的是连续的内部表征,而非文本,但更大的表征能力并不能证明接收者使用了与任务相关的信息。仅凭最终任务的表现也无法揭示观察到的效果是否依赖于消息的存在、为评估示例生成的内容,或是由其他智能体提供的信息。我们引入了一种因果审计,应用受控的消息替换,在发送者生成的表征进入接收者的边界处进行。四种消息设置支持五项对编码的发送者信息、接收者对消息存在和身份的敏感性、示例特定内容的任务价值,以及由其他智能体提供的附加价值的测量。我们将该审计应用于与 Qwen3-4B 和 Qwen3-8B 在 GSM8K、ARC-C 和 MATH-500 上的潜在中继。在 GSM8K 上,Qwen3-4B 的整体表现效果为 -1.00 个百分点,分解为 -6.17 个百分点的效果由其他示例消息保留,以及 +5.17 个百分点的效果归因于示例特定内容;在 8B 中,这两个成分的方向均发生反转。在 MATH-500 上,Qwen3-4B 的增益为 15.00 个百分点,其中 8.33 个百分点由其他示例消息保留,6.67 个百分点归因于示例特定内容,而 8B 的增益则主要由前者成分主导。自我替代比较进一步表明,示例特定内容和其他智能体的价值是不同的。这些结果表明,整体准确性无法识别潜在消息如何影响接收者,并促使将受控消息比较作为潜在通信的标准评估。
cs.AI / 22 / 2607.26787

Property-driven Causal Abstractions for Markov Decision Processes

基于属性的因果抽象在马尔可夫决策过程中的应用
Schmidt, Jule, Weininger, Maximilian, Dubslaff, Clemens, Parker, David, Jansen, Nils
Abstract
Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state variables and their valuations. The exponential blowup in the number of states renders many reasoning tasks in MDPs challenging. Abstractions are promising techniques to reduce MDPs and thus mitigate scalability issues. In this work, we introduce a notion of causality on factored MDPs and a novel property-driven causal abstraction technique that retains many characteristics of the original MDP model. For this, we rely on causal relations over state variable predicates and identify those states that share the same reasons for fulfilling or violating a given abstraction property. We theoretically and empirically compare various causal MDP abstractions using different model types such as MDPs, interval MDPs, or stochastic games. Our evaluation demonstrates the potential of our approach: For several standard benchmarks, we obtain small abstractions that allow us to compute near-optimal policies for the original MDP. Furthermore, our causal abstractions often generalize to related large-scale MDP models.
Chinese Translation
马尔可夫决策过程(MDPs)作为决策模型被广泛使用,通常通过状态变量及其取值在分解状态空间上进行指定。状态数量的指数级膨胀使得许多MDP中的推理任务变得具有挑战性。抽象化是一种有前景的技术,可以减少MDP的复杂性,从而缓解可扩展性问题。在本研究中,我们引入了在分解MDP上因果关系的概念,并提出了一种新颖的基于属性的因果抽象技术,该技术保留了原始MDP模型的许多特征。为此,我们依赖于状态变量谓词之间的因果关系,并识别出那些共享相同原因以满足或违反给定抽象属性的状态。我们理论上和实证上比较了使用不同模型类型(如MDPs、区间MDPs或随机博弈)的各种因果MDP抽象。我们的评估展示了我们方法的潜力:在几个标准基准测试中,我们获得了小规模的抽象,使我们能够为原始MDP计算近似最优策略。此外,我们的因果抽象通常能够推广到相关的大规模MDP模型。
cs.AI / 23 / 2607.26903

From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence

从被动视频到可编辑体验:面向具身智能的物理基础体验合成
Luo, Jia
Abstract
The key bottleneck in embodied AI is not model architecture but data. Although billions of human manipulation videos exist online, robots cannot directly learn from them due to the embodiment gap between human morphology and robot hardware. We introduce Pegasus, a low-resource framework that bridges this gap by translating human demonstrations into robot-learnable data through structured knowledge transfer. Instead of relying on raw video prompts, Pegasus constructs a graph-based intermediate representation: a Task Graph extracted from human videos is transformed through Affordance and Constraint Graphs into a Robot Planning Graph for robot-conditioned video generation. A hierarchical affordance latent space models the relationship between object states, affordances, and tasks, enabling generalization beyond object identities. A closed-loop physics verifier further filters invalid generations using kinematic feasibility, collision constraints, and joint limits. We evaluate Pegasus across a range of egocentric manipulation benchmarks, including GTEA Gaze+ and EPIC-KITCHENS-100, and diverse robot embodiments, assessing Task Correctness, Executability, State Consistency, and Learnability. Results demonstrate reliable cross-embodiment translation and show that robot data generation can be reframed from a hardware collection problem into a scalable, low-resource knowledge transfer problem.
Chinese Translation
具身人工智能的关键瓶颈不在于模型架构,而在于数据。尽管网上存在数十亿个人类操作视频,机器人却无法直接从中学习,因为人类形态与机器人硬件之间存在具身差距。我们提出了Pegasus,这是一个低资源框架,通过结构化知识转移将人类示范转换为机器人可学习的数据,从而弥合这一差距。Pegasus并不依赖于原始视频提示,而是构建了一种基于图的中间表示:从人类视频中提取的任务图(Task Graph)通过可供性图(Affordance Graph)和约束图(Constraint Graph)转化为机器人规划图(Robot Planning Graph),用于生成机器人条件下的视频。一种分层的可供性潜在空间建模了对象状态、可供性和任务之间的关系,使得超越对象身份的泛化成为可能。一个闭环物理验证器进一步通过运动学可行性、碰撞约束和关节限制过滤无效生成。我们在一系列以自我为中心的操作基准上评估Pegasus,包括GTEA Gaze+和EPIC-KITCHENS-100,以及多样化的机器人形态,评估任务正确性、可执行性、状态一致性和可学习性。结果表明,跨形态翻译可靠,并显示机器人数据生成可以从硬件收集问题重新构建为可扩展的、低资源的知识转移问题。
cs.AI / 24 / 2607.26935

What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation

检测人工智能代理需要什么?浏览器自动化下的行为检测最小特征集
Choudhary, Vishisht, Schmidt, Lukas, Kenntner, Anne Zoë, Skhab, Feras, Osswald, Michel, Ernstberger, Jens
Abstract
Bot detectors deployed at scale treat traffic as binary: human or bot. This assumption breaks when AI agents browse the web through browser automation, a traffic class that is neither and that binary classifiers structurally cannot represent. We present a three-class detection framework distinguishing humans, bots, and AI agents, and show that the binary-vs-agent confusion is architectural: a binary human-vs-bot detector misroutes agent sessions because its label space lacks an agent class. On our controlled benchmark, an MLP binary classifier misclassifies 39.1% of real AI agents as human and a SAINT binary transformer misclassifies 34.5%; adding an explicit agent class yields per-class agent F1 = 1.000 in all 30 runs (3 model families $\times$ 10 seeds). To measure evasion resistance, we construct a five-level evasion ladder spanning passive observation, GAN-generated trajectories, and replay of real human cursor data ($n = 2299$ evasion sessions). Across 10 seeds and 3 model families we observe zero agent misses in 22990 per-seed predictions. The discriminative signal is a browser-automation artifact, not evidence of agent reasoning: Playwright does not emit the raw pointer-move and wheel-delta streams a physical input device produces, and this absence signature survives trajectory manipulation. Exhaustive search over all feature subsets of size 1-5 (9401 GBMs) shows that two behavioral features (mouse_event_rate, teleport_click_ratio) give 100% observed agent recall at every evasion level with agent precision 0.994; five features lift macro-F1 to 0.991. The signal is redundantly encoded: removing teleport_click_ratio leaves agent detection at 100%. The single-feature regime is degenerate, flagging every agent only by collapsing the classifier to always predict "agent". Two features robustly isolate agents; five separate all three traffic classes at macro-F1 $\geq 0.99$.
Chinese Translation
大规模部署的机器人检测器将流量视为二元分类:人类或机器人。当人工智能代理通过浏览器自动化浏览网络时,这一假设被打破,因为这一流量类别既不是人类也不是机器人,而二元分类器在结构上无法表示这一点。我们提出了一个三类检测框架,区分人类、机器人和人工智能代理,并展示了二元与代理混淆的架构性:二元人类与机器人检测器错误地将代理会话分类,因为其标签空间缺乏代理类别。在我们的受控基准测试中,一个多层感知器(MLP)二元分类器将39.1%的真实人工智能代理错误分类为人类,而一个SAINT二元变换器将34.5%错误分类;添加一个明确的代理类别在所有30次运行中(3个模型系列 × 10个种子)实现了每类代理的F1值为1.000。为了测量规避抵抗力,我们构建了一个五级规避梯度,涵盖被动观察、生成对抗网络(GAN)生成的轨迹以及真实人类光标数据的重放(n = 2299个规避会话)。在10个种子和3个模型系列中,我们观察到在22990次每种子预测中没有代理遗漏。判别信号是浏览器自动化的伪影,而不是代理推理的证据:Playwright不会发出物理输入设备产生的原始指针移动和滚轮增量流,而这种缺失特征在轨迹操控中依然存在。对所有大小为1-5的特征子集(9401个GBM)的穷举搜索显示,两个行为特征(mouse_event_rate, teleport_click_ratio)在每个规避级别上均提供100%的观察到的代理召回率,代理精度为0.994;五个特征将宏观F1提升至0.991。信号是冗余编码的:去除teleport_click_ratio后,代理检测仍保持在100%。单特征模式是退化的,仅通过将分类器压缩为始终预测“代理”来标记每个代理。两个特征稳健地隔离代理;五个特征将所有三类流量的宏观F1分数提升至≥0.99。
cs.AI / 25 / 2607.26946

Belief-Guided Decision Making with Uncertainty Gating in the Game of Go

基于信念引导的不确定性门控决策在围棋中的应用
Yaghoubi, Mehrad, Bastanfard, Azam, Jalilvand, Abbas, Rezaei, Ashkan
Abstract
Recent advancements in Computer Go, driven by AlphaZero and MuZero, rely heavily on Monte Carlo Tree Search (MCTS) to correct the errors of the neural network policy. While effective on massive computational clusters, this dependence creates a critical bottleneck on consumer-grade hardware, where the computational cost of tree management severely limits inference rates. Furthermore, without deep search, these models suffer from hallucination, proposing moves with high confidence that are strategically fatal. This paper introduces a novel Belief-Guided architecture that disentangles the Policy head from a distinct Belief head. Unlike traditional value functions, the Belief head acts as an internal simulator and independent critic, modeling epistemic uncertainty and strategic stability. By integrating memory mechanisms (Transformer/GRU) to handle long-term dependencies and the Ko rule, and utilizing a gating mechanism to filter overconfident policy errors, our model shifts the burden of intelligence from runtime search to parametric "intuition." Experimental results demonstrate that this approach significantly improves search-free win rates and reduces hallucination, enabling professional-level play on limited hardware where massive MCTS is infeasible.
Chinese Translation
近年来,计算机围棋领域的进展主要得益于AlphaZero和MuZero,这些方法在很大程度上依赖于蒙特卡洛树搜索(Monte Carlo Tree Search, MCTS)来纠正神经网络策略的错误。尽管在大型计算集群上效果显著,但这种依赖在消费者级硬件上造成了严重的瓶颈,因为树管理的计算成本严重限制了推理速度。此外,在没有深度搜索的情况下,这些模型容易产生幻觉,提出高置信度但在策略上致命的走法。本文提出了一种新颖的基于信念引导的架构,将策略头与独立的信念头分离。与传统的价值函数不同,信念头充当内部模拟器和独立评论者,建模认知不确定性和战略稳定性。通过整合记忆机制(Transformer/GRU)来处理长期依赖和劫规则,并利用门控机制过滤过于自信的策略错误,我们的模型将智能的负担从运行时搜索转移到参数化的“直觉”上。实验结果表明,该方法显著提高了无搜索的胜率并减少了幻觉,使得在有限硬件上实现专业水平的对弈成为可能,而在这些环境中大规模的MCTS是不可行的。
cs.AI / 26 / 2607.27056

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

Setoka:用于评估个性化代理在异构数据上层次用户理解的基准
Zeng, Lingyang, Chen, Guangze, Yu, Kaichen, Pan, Zhicheng, Weng, Siyang, Hu, Zirui, Du, Xiangyun, He, Hailin, Zhang, Rong, Yang, Chengcheng, Huang, Kai, Zhou, Xuan
Abstract
Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory benchmarks primarily evaluate whether an agent can retrieve information explicitly stated in conversational histories, failing to provide an effective assessment of deeper user understanding. In this work, we propose Setoka, a benchmark for evaluating memory-augmented personalized agents with hierarchical user understanding from heterogeneous data. Grounded in theories from cognitive and personality psychology, Setoka defines four levels of user understanding, i.e., semantic memory, episodic memory, behavior pattern, and personality trait. Moreover, to enable realistic yet privacy-preserving evaluation, we design a psychometrics-based pipeline that synthesizes diverse, coherent heterogeneous user data and queries at scale. Finally, we leverage Setoka to evaluate 3 language models combined with 5 memory systems for 10 synthetic users. Our comprehensive evaluation reveals that while existing systems perform well on semantic memory retrieval, their performance declines on episodic memory. Moreover, when dealing with behavior pattern and personality trait understanding tasks that require integrating heterogeneous and fragmented information dispersed over time, performance declines even further. These findings demonstrate that user understanding cannot be handled by simple fact retrieval, motivating the design of memory mechanisms for cross-source integration and abstraction over long-term user behavior.
Chinese Translation
个性化代理越来越多地被应用于协助用户完成各种任务。有效的个性化辅助不仅需要从代理记忆中检索存储的过去交互中的显性事实,还需要推断抽象的个人特征。然而,现有的记忆基准主要评估代理是否能够检索对话历史中明确陈述的信息,未能有效评估更深层次的用户理解。在本研究中,我们提出了Setoka,一个用于评估具有层次用户理解的记忆增强个性化代理的基准,基于异构数据。Setoka基于认知心理学和人格心理学的理论,定义了四个层次的用户理解,即语义记忆、情节记忆、行为模式和人格特征。此外,为了实现现实且保护隐私的评估,我们设计了一个基于心理测量的管道,能够大规模合成多样且一致的异构用户数据和查询。最后,我们利用Setoka评估了3种语言模型与5种记忆系统在10个合成用户上的表现。我们的综合评估揭示,尽管现有系统在语义记忆检索方面表现良好,但在情节记忆上的表现却有所下降。此外,当处理需要整合分散在时间上的异构和碎片化信息的行为模式和人格特征理解任务时,性能进一步下降。这些发现表明,用户理解不能仅通过简单的事实检索来处理,这激励了跨源整合和长期用户行为抽象的记忆机制设计。
cs.AI / 27 / 2607.27081

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

基于策略蒸馏的LLM安全性:一种模板稳健重对齐的路由方法
Guo, Yongjian, Ma, Wanlun, Shen, Lingyu, Xiao, Xi, Wen, Sheng
Abstract
Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand. Existing safety-realignment defenses often fail in practice due to three key limitations: they frequently cause catastrophic forgetting of specialized skills; their effectiveness collapses when the defender cannot observe the attacker's prompt template; and successfully realigned models remain susceptible to re-jailbreaking via simple system prompt switches. To address these challenges, we propose Routing-based On-Policy Distillation (ROPD), a novel realignment framework that models the divergence between aligned and compromised output probability distributions rather than fitting specific prompt templates. We conduct extensive experiments comparing ROPD against four state-of-the-art baselines across three datasets and three base models with varying alignment strengths. Our results demonstrate that when baseline defenses face template mismatches, often accompanied by severe degradation in downstream task performance. In contrast, ROPD substantially mitigates template-mismatch risks, maintaining superior robustness in both defense effectiveness and capability preservation. While our analysis indicates ROPD is not entirely immune to template shifts, its performance degradation is negligible compared to existing methods, establishing a new standard for robust LLM realignment.
Chinese Translation
微调是专门化大型语言模型(LLMs)的主要范式,但它暴露了一个关键漏洞:恶意数据提供者可以将有害行为嵌入下游语料库,从而创建在需求下保留专业技能但违反人类价值观的模型。现有的安全重对齐防御在实践中往往失败,主要有三个关键限制:它们常常导致专业技能的灾难性遗忘;当防御者无法观察到攻击者的提示模板时,其有效性会崩溃;而成功重对齐的模型仍然容易通过简单的系统提示切换重新越狱。为了解决这些挑战,我们提出了基于路由的策略蒸馏(Routing-based On-Policy Distillation,ROPD),这是一种新颖的重对齐框架,它建模对齐和妥协输出概率分布之间的差异,而不是拟合特定的提示模板。我们进行了广泛的实验,将ROPD与四个最先进的基线模型在三个数据集和三种基础模型上进行比较,基线模型具有不同的对齐强度。我们的结果表明,当基线防御面临模板不匹配时,通常伴随着下游任务性能的严重下降。相比之下,ROPD显著减轻了模板不匹配的风险,在防御有效性和能力保留方面保持了优越的稳健性。尽管我们的分析表明ROPD并非完全免疫于模板变化,但其性能下降与现有方法相比微不足道,为稳健的LLM重对齐建立了新的标准。
cs.AI / 28 / 2607.27130

AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching

AgentMap:本体匹配中的联合等价性与包含关系发现
Song, Yiping, Chen, Jiaoyan, Schmidt, Renate, Yang, Hui, Zhang, Wen
Abstract
Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The existing OM systems identify only one type of semantic correspondence and cannot simultaneously discover equivalence and subsumption mappings. In this paper, we introduce Hybrid Ontology Matching (HOM), a new OM task that unifies equivalence and subsumption discovery, and accordingly propose a Large Language Model (LLM)-based multi-agent OM framework AgentMap that is implemented by a series of interdependent semantic decisions. Given a concept in the source ontology, AgentMap integrates semantic retrieval, hierarchical search, and collaborative multi-agent LLM reasoning to progressively explore the target ontology, identifying either the equivalent concept, if one exists, or the most fine-grained subsumer. We further extend four OM datasets for a HOM benchmark and evaluate AgentMap under hybrid, equivalence-only, and subsumption-only settings. Experimental results show that AgentMap achieves promising performance on the hybrid setting, and at the same time outperforms equivalence matching and subsumption matching baselines on the equivalence-only and subsumption-only settings, respectively.
Chinese Translation
本体匹配(Ontology Matching, OM)传统上被定义为等价性发现或包含关系匹配。现有的OM系统仅识别一种类型的语义对应关系,无法同时发现等价和包含关系映射。本文提出了一种新的OM任务——混合本体匹配(Hybrid Ontology Matching, HOM),该任务统一了等价性和包含关系的发现,并相应地提出了一种基于大型语言模型(Large Language Model, LLM)的多智能体OM框架AgentMap,该框架通过一系列相互依赖的语义决策来实现。给定源本体中的一个概念,AgentMap集成了语义检索、层次搜索和协作多智能体LLM推理,逐步探索目标本体,识别出等价概念(如果存在)或最细粒度的包含者。我们进一步扩展了四个OM数据集以建立HOM基准,并在混合、仅等价和仅包含关系的设置下评估AgentMap。实验结果表明,AgentMap在混合设置下表现出色,同时在仅等价和仅包含关系的设置下分别超越了等价匹配和包含关系匹配的基线。
cs.AI / 29 / 2607.27134

Linguistic Monoculture in LLM-Assisted Language Use

大型语言模型辅助语言使用中的语言单一文化
Thejaswi, Suhas, Kulshreshta, Juhi, Oettershagen, Lutz
Abstract
Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduce population-level variation in linguistic form, a phenomenon we refer to as linguistic monoculture. We develop a mathematical framework in which authors and LLMs are represented as distributions over linguistic features and coevolve through repeated interaction. We analyze three interaction mechanisms: a shared model with a fixed linguistic distribution, a shared model recursively updated from author outputs, and personalized models updated through author-specific and population-level feedback. We characterize the resulting equilibria and convergence rates, showing that, shared models can drive authors toward a common norm, recursive feedback relocates the shared norm without altering pairwise spread under common conformity, and personalization can preserve a family of distinct author-model equilibria with nonzero linguistic diversity. We then endogenize conformity as a strategic choice trading off private benefits from clarity, legibility, and perceived fluency against distinctive style. Within this utility model, individually rational authors may conform more than is socially optimal because they do not internalize the value their distinctiveness provides to others, creating a negative externality and a price of monoculture that is finite for each fixed instance but can grow without bound when distinctiveness dominates authenticity. Synthetic simulations illustrate how fixed shared assistance, recursive feedback, and personalization produce different long-run diversity outcomes.
Chinese Translation
写作和交流越来越多地受到大型语言模型(LLMs)的介导,这些模型被用于起草、修订和润色文本。尽管这种辅助可以提高清晰度并帮助作者满足机构期望,但对共享模型的广泛依赖可能会减少语言形式的群体层面变异,这一现象我们称之为语言单一文化。我们建立了一个数学框架,其中作者和LLMs被表示为语言特征的分布,并通过重复互动共同进化。我们分析了三种互动机制:具有固定语言分布的共享模型、从作者输出递归更新的共享模型,以及通过作者特定和群体层面反馈更新的个性化模型。我们描述了由此产生的均衡状态和收敛速率,表明共享模型可以将作者驱动向共同规范,递归反馈在不改变共同遵从下重新定位共享规范,而个性化可以保留一系列具有非零语言多样性的独特作者-模型均衡。然后,我们将遵从内生化为一种战略选择,在清晰度、可读性和感知流畅性带来的私人利益与独特风格之间进行权衡。在这一效用模型中,个体理性的作者可能会比社会最优的情况更倾向于遵从,因为他们没有内化其独特性对他人所提供的价值,从而产生负外部性和单一文化的代价,对于每个固定实例而言是有限的,但当独特性主导真实性时,可以无限增长。合成模拟展示了固定共享辅助、递归反馈和个性化如何产生不同的长期多样性结果。
cs.AI / 30 / 2607.27155

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

OmegaUse-OfficeVal:基于经济基础的长时段办公套件任务中大型语言模型代理的基准评估
Zhou, Jingbo, Zhao, Yusai, Bao, Qi, Cao, Jingjia, Chen, Zhenghai, Gao, Chang, Guo, Kaiqi, Guo, Muxin, Li, Mingxuan, Lu, Xinjiang, Ma, Yanru, Xiao, Yixiong, Zhang, Zenghui, Zhang, Le, Wu, Hua
Abstract
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete. An important feature of the benchmark is that each task is paired with two economic signals: human labor time and task price proxy. These signals enable direct comparisons between human costs and LLM inference costs, as well as value-weighted evaluation. To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evaluate several frontier LLMs together with a human baseline. Although all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality. The code and dataset are fully open-sourced, and more information is available on our project website: https://omegause-officeval.github.io.
Chinese Translation
大型语言模型(LLM)代理越来越被期望能够协助用户完成任务。然而,现有的基准测试对评估代理是否能够以合理成本执行办公套件工作流程的支持有限。我们提出了OmegaUse-OfficeVal,这是一个用于评估LLM代理在长时段办公套件任务中的基准,具有任务级经济基础。该基准包含100个任务,这些任务来源于实践者提出的办公套件请求,并通过隐私保护过程进行了调整。平均而言,这些任务需要2.32小时的人力劳动才能完成。该基准的一个重要特征是每个任务都配有两个经济信号:人力劳动时间和任务价格代理。这些信号使得人力成本与LLM推理成本之间的直接比较成为可能,并支持价值加权评估。为了支持稳定的评估,我们基于细致的评分标准开发了代码验证器。我们评估了几种前沿的LLM,并与人类基线进行了比较。尽管所有评估的LLM在成本和速度上都显著低于人类工人,但它们在可交付质量上尚未接近人类水平。代码和数据集均已完全开源,更多信息请访问我们的项目网站:https://omegause-officeval.github.io。
cs.AI / 31 / 2607.27177

Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

用于任务无关适应的合作伙伴能力估计在临时团队中的应用
Tisnikar, Peter, Swieczkowska, Maja, Ma, Benteng, Canal, Gerard, Leonetti, Matteo
Abstract
Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume that agents will collaborate on a single, fixed task and that the partner's capabilities, their ability to successfully execute the desired action, are already known. In reality, a partner's true capabilities are often hidden, and human collaborators may act sub-optimally on tasks with multiple valid strategies. To address these limitations, we extend ad-hoc teamwork into a multi-task setting by re-framing it as a problem of joint planning with decentralised execution under hidden partner capabilities. We introduce CE-CM (Capability Estimation via Contextual Models), an approximate Bayesian method that infers task-invariant capability vectors. By using simulation-based sampling, the agent estimates capabilities and induces a contextual Multi-agent Markov Decision Processes for planning. This approach requires no population pre-training and refines its beliefs online from just a few tasks. To account for human unpredictability, we propose CE-CM-Div, an extension that evaluates capability hypotheses against diverse planner rollouts rather than a single optimal trajectory. Simulated experiments demonstrate that CE-CM rapidly recovers hidden capabilities, reduces infeasible action assignments, and adapts to changes over time. Furthermore, in an offline human study of 225 trajectories from 15 participants, CE-CM-Div substantially improved capability estimates over the baseline CE-CM method. Our results suggest capability-based modelling is a promising interpretable, task-agnostic representation in the studied settings, demonstrating that accounting for behavioural diversity is essential for robust human-AI teaming.
Chinese Translation
与新颖且多样的合作伙伴进行有效协作是自主代理的一项关键技能。目前大多数临时团队(AHT)方法假设代理将针对单一固定任务进行协作,并且合作伙伴的能力,即成功执行所需动作的能力,已经是已知的。实际上,合作伙伴的真实能力往往是隐藏的,而人类合作者在具有多种有效策略的任务中可能表现出次优行为。为了解决这些局限性,我们将临时团队扩展到多任务环境,通过将其重新框定为在隐藏的合作伙伴能力下进行分散执行的联合规划问题。我们引入了CE-CM(通过上下文模型进行能力估计),这是一种近似贝叶斯方法,用于推断任务不变的能力向量。通过基于仿真的采样,代理估计能力并诱导出上下文多智能体马尔可夫决策过程进行规划。这种方法不需要群体预训练,并且能够在线从少量任务中精炼其信念。为了考虑人类的不可预测性,我们提出了CE-CM-Div,这是一个扩展,评估能力假设时考虑多样化的规划者展开,而不是单一的最优轨迹。模拟实验表明,CE-CM能够快速恢复隐藏的能力,减少不可行的行动分配,并适应随时间变化的情况。此外,在对15名参与者的225条轨迹进行的离线人类研究中,CE-CM-Div显著改善了能力估计,相较于基线CE-CM方法。我们的结果表明,基于能力的建模在研究环境中是一种有前景的可解释、任务无关的表示,表明考虑行为多样性对于稳健的人机团队合作至关重要。
cs.AI / 32 / 2607.27191

Can AI agents conduct open-ended AI research? Early evidence from two case studies

人工智能代理能否进行开放式人工智能研究?来自两个案例研究的早期证据
Kirgis, Peter, Kapoor, Sayash, Schwartz, Andrew, Rabanser, Stephan, Africa, David, Voudouris, Konstantinos, Nguyen, Viet, Pilditch, Toby, Dubois, Magda, Coppock, Harry, Ududec, Cozmin, Nadgir, Nitya, Orona, Matilda, Bayer, Tilman, Chan-Sew, Derrick, Ling, Yue, Shetty, Abhishek, Toner, Helen, Hadfield, Gillian, Lazar, Seth, Newman, Steve, Tekofsky, Shoshannah, Bommasani, Rishi, Narayanan, Arvind
Abstract
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.
Chinese Translation
对人工智能快速进展的预测依赖于人工智能代理自动化人工智能研究。然而,关于代理是否能够进行开放式人工智能研究的证据仍然稀缺。目前的评估要么测试代理在狭窄、可验证的任务上,这排除了开放式研究,要么将人工智能生成的论文提交给盲审,这种方式过于繁重、随机,并且审稿质量较差。我们提出了一种第三种方法来衡量人工智能研发自动化的进展。一个代理承担一篇高质量未发表论文的中心开放式研究问题,论文的原作者对其输出进行评分。我们称这些为影子评估。我们对两篇未发表的NeurIPS 2026投稿进行了影子评估,给予前沿代理六天时间和数千美元的计算资源。代理在没有人类帮助的情况下完成了所有工程工作,但未能在回答研究问题上取得实质性进展。因此,作者明确拒绝了这两篇论文。我们识别出五种反复出现的失败模式:对可发表研究标准的判断不佳、对研究设计缺陷的无创意回应、从死胡同中无效回溯、资源意识差以及指令漂移。使用第二个模型和支架进行的稳健性检查重现了这些失败。我们发布了专家评审、调查反馈、代理库和日志。我们的结果提供了早期证据,表明今天的代理能够完成人工智能研究的工程工作,但在研究生命周期的关键部分中存在困难。
计算语言学 (Computation and Language)
51
cs.CL / 1 / 2607.26060

Large-Scale ChatBot Validation Through Customer Digital Twin Simulations

通过客户数字双胞胎模拟进行大规模聊天机器人验证
Iglesias, Cristovao, Batra, Devesh, Atreya, Alankar, Wagner, Stefan, Hankache, Robert, Sinclair, Patrick, Pelosio, Giulio, McMillan, Michael, Cowan, Greig A., Khraishi, Raad
Abstract
LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical barrier to safe deployment. We present a two-part contribution for large-scale chatbot validation. First, we introduce a methodology for creating high-fidelity synthetic customer agents (SCAs) as digital twins, grounded in real transactional and conversational data, that enables automatic generation and behavioral conditioning to simulate diverse customer profiles and interaction styles. Evaluation demonstrates that SCAs achieve high semantic alignment with real customers, low hallucination rates, and successful personality trait reproduction with controllable interventions. Second, we develop an SCA-based validation framework combining automated LLM-as-a-Judge evaluation, human expert testing, and adversarial probing. Scenario-based validation across emotional states, demographic groups, and linguistic factors confirms robust performance. Our approach was used to validate a customer facing chatbot at a leading UK bank, providing financial institutions with a scalable pathway toward regulatory compliance.
Chinese Translation
基于大型语言模型(LLM)的聊天机器人正在改变银行等受监管领域的客户服务,但可扩展且具有成本效益的验证仍然是安全部署的关键障碍。我们提出了一个两部分的贡献以实现大规模聊天机器人验证。首先,我们介绍了一种创建高保真合成客户代理(SCA)作为数字双胞胎的方法,该方法基于真实的交易和对话数据,能够自动生成和行为调节,以模拟多样的客户特征和互动风格。评估表明,SCA与真实客户在语义上高度一致,幻觉率低,并且能够通过可控干预成功再现个性特征。其次,我们开发了一个基于SCA的验证框架,结合了自动化的LLM作为评判者的评估、人类专家测试和对抗性探测。基于场景的验证涵盖了情感状态、人口统计群体和语言因素,确认了其稳健的性能。我们的方法被用于验证一家领先英国银行的面向客户的聊天机器人,为金融机构提供了一条可扩展的合规路径。
cs.CL / 2 / 2607.26066

Do Methods Support the Claims? Intra-Paper Verification for Peer Review

方法是否支持主张?同行评审中的论文内部验证
Ballakuraya, Ranjitha Shivaprasad, Mahyari, Arash, Srinivasan, Ashok
Abstract
The growing volume of scientific submissions has motivated interest in using large language models (LLMs) to assist peer review. Existing automated novelty assessment approaches typically compare a paper's claimed contributions against prior literature, implicitly assuming that these contributions are accurately realized in the work itself. Human reviewers, however, frequently challenge novelty claims not because similar ideas already exist, but because the methodological evidence presented in the paper does not adequately support them. This internal mismatch between claimed contributions and methodological realization is rarely examined by current LLM-based review systems. To address this gap, we introduce intra-paper claim verification, a framework that evaluates whether novelty claims articulated in a paper are substantiated by the methods used to realize them. The framework employs an LLM to extract novelty claims from the introduction, retrieve claim-relevant methodological evidence, and assess whether the methods substantiate the stated contributions. Assessment is guided by reviewer-inspired evaluation criteria derived inductively from human peer reviews collected from 182 ICLR 2025 papers. These criteria capture recurring reviewer concerns related to novelty, methodology, clarity, and other issues and are used to generate structured reviewer-style assessments of claim substantiation. We evaluate the framework by comparing LLM-generated review comments against human reviewer concerns on a balanced subset of accepted and rejected papers. Human evaluation demonstrates significant alignment between framework-generated assessments and human reviewer concerns, particularly for novelty-related issues. BERTScore further distinguishes corresponding human-LLM review pairs from mismatched controls, indicating that the framework captures concerns consistent with human reviewer observations.
Chinese Translation
科学投稿数量的不断增加引发了对使用大型语言模型(LLMs)辅助同行评审的兴趣。现有的自动化新颖性评估方法通常将论文所声称的贡献与先前文献进行比较,隐含地假设这些贡献在论文中得到了准确体现。然而,人工评审者常常质疑新颖性主张,并非因为类似的想法已经存在,而是因为论文中提供的方法证据未能充分支持这些主张。当前基于LLM的评审系统很少检查这种声称的贡献与方法实现之间的内部不匹配。为了解决这一问题,我们提出了论文内部主张验证(intra-paper claim verification),这是一个评估论文中阐述的新颖性主张是否得到实现这些主张所用方法支持的框架。该框架利用LLM从引言中提取新颖性主张,检索与主张相关的方法证据,并评估这些方法是否支持所述的贡献。评估受到来自182篇ICLR 2025论文中收集的人工同行评审的启发,归纳出评审者的评估标准。这些标准捕捉了与新颖性、方法论、清晰度及其他问题相关的反复出现的评审者关注点,并用于生成结构化的评审者风格的主张支持评估。我们通过将LLM生成的评审评论与接受和拒绝论文的平衡子集上的人工评审者关注点进行比较来评估该框架。人工评估显示,框架生成的评估与人工评审者关注点之间存在显著一致性,特别是在与新颖性相关的问题上。BERTScore进一步区分了相应的人类-LLM评审对与不匹配的对照组,表明该框架捕捉到的关注点与人工评审者的观察一致。
cs.CL / 3 / 2607.26178

DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues

DuplexGen:自适应合成的人机交互轮流对话
Kim, Takyoung, Kim, Kang-wook, Woo, Sang Hoon, Hirschberg, Julia, Kim, Gunhee, Hakkani-Tür, Dilek
Abstract
Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current models apply a single norm regardless of context. This limitation originates in their training data: human-human speech corpora capture natural timing phenomena but provide little role grounding or scenario-specific norms, while heuristic or prompted synthesis methods inject turn-taking behaviors without basing them on human preferences. We introduce DuplexGen, a framework for generating dialogues with scenario-adaptive turn-taking by calibrating LLM predictions against a small set of slot-level human preference annotations. In six cooperative and competitive tasks, human turn-taking preferences differ systematically, and DuplexGen aligns substantially more closely with those preferences than uncalibrated prompting or training solely on generic human-human data; a full-duplex model trained on DuplexGen-generated data exhibits distinctive, human-preferred turn-taking behaviors. These results show that human calibration, not corpus scale or prompt design alone, is what allows turn-taking synthesis to be scenario-specific.
Chinese Translation
轮流对话是全双工交互的核心组成部分。适当的轮流对话行为因场景而异,但当前模型在不同上下文中应用单一规范。这一局限性源于它们的训练数据:人类-人类的语料库捕捉了自然的时序现象,但提供的角色基础或特定场景的规范较少,而启发式或提示合成方法则在没有基于人类偏好的基础上注入轮流对话行为。我们提出了DuplexGen,一个通过将大型语言模型(LLM)预测与少量槽级人类偏好注释进行校准,从而生成具有场景自适应轮流对话的框架。在六个合作和竞争任务中,人类的轮流对话偏好系统性地存在差异,而DuplexGen与这些偏好的对齐程度显著高于未经校准的提示或仅基于通用人类-人类数据的训练;在DuplexGen生成的数据上训练的全双工模型展现出独特的人类偏好的轮流对话行为。这些结果表明,人类校准,而非语料库规模或提示设计单独,才使得轮流对话合成能够具备场景特异性。
cs.CL / 4 / 2607.26200

Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting

选择何处以及如何进行内容审核:过滤器位置和响应重写的端到端权衡
Hu, Mengya, Park, Susie, Ilic, Suzana, Wei, Qiong, Atluri, Sandeep, Deng, Myra, Fross, Tucker, Tigges, Curt
Abstract
Content-moderation classifiers are usually evaluated in isolation, but deployment requires choosing where to intervene and what follows a flag. We evaluate these choices using two end-to-end customer-outcome metrics rather than component accuracy: Usefulness, the fraction of turns with a shown, non-harmful, relevant response, and Harmful Exposure, the fraction with a shown harmful response. Latency and error rates are diagnostics. We compare Input only, Response only, and Input + response hard blocking on a human-labelled product benchmark and public ToxicChat evaluation. At the evaluated operating points, Response only achieves the highest filter-only Usefulness in both settings, while Input + response achieves lower Harmful Exposure. Replacing Response only blocking with Response + rewrite recovers most blocked traffic and yields the same observed Harmful Exposure count as Response only blocking for the selected configuration; this equality is not an equivalence result. Probe routing substantially reduces conditional route-and-generation time relative to LLM routing at comparable measured outcomes. A focused output review shows how rewrites balance filter passage with usefulness by generalizing triggering language while retaining benign intent and safe redirection; some sensitive-domain outputs nevertheless omit potentially safety-relevant support information. These results support comparing moderation configurations under deployment-specific safety and latency constraints rather than applying a universal placement rule. Code and public artifacts are available at https://github.com/microsoft/mod-frontier
Chinese Translation
内容审核分类器通常是孤立评估的,但在实际部署中需要选择干预的位置以及标记后采取的措施。我们使用两个端到端的客户结果指标来评估这些选择,而不是组件准确性:有用性(Usefulness),即展示的非有害相关响应的回合比例,以及有害暴露(Harmful Exposure),即展示的有害响应的回合比例。延迟和错误率作为诊断指标。我们在一个人工标注的产品基准和公共的 ToxicChat 评估上比较了仅输入、仅响应和输入 + 响应硬阻塞。在评估的操作点上,仅响应在两种设置中都实现了最高的仅过滤器有用性,而输入 + 响应则实现了较低的有害暴露。用响应 + 重写替代仅响应阻塞可以恢复大部分被阻塞的流量,并在所选配置下产生与仅响应阻塞相同的观察到的有害暴露计数;这种相等性并不是等价结果。探测路由相对于 LLM 路由在可比测量结果下显著减少了条件路由和生成时间。集中输出审查显示,重写如何通过概括触发语言来平衡过滤器通过与有用性,同时保留良性意图和安全重定向;尽管如此,一些敏感领域的输出仍然省略了潜在的安全相关支持信息。这些结果支持在特定部署的安全性和延迟约束下比较审核配置,而不是应用普遍的放置规则。代码和公共文档可在 https://github.com/microsoft/mod-frontier 获取。
cs.CL / 5 / 2607.26221

Characterizing Human-Likeness in AI Generated Poetry: A Zero-shot Classification Study

表征人工智能生成诗歌的人类相似性:一项零样本分类研究
Biswas, A. N., Tabassum, T., Shohid, A. A., Mou, R. M., Esha, A. A., Sadeque, F., Ahmed, A.
Abstract
With the advancement of AI technologies, Generative AI (GenAI) and human written text have become nearly indistinguishable. Additionally, the global standardization of AI chatbots made academic malpractice more frequent. Furthermore, existing research indicates GenAI poems are the most difficult to distinguish even without any modification thus, GenAI poems are naturally deemed human-like by modern detectors. However, the objectivity of such dissertations needs to be verified against modern detection tools but the subjectivity of poetry and the black-box nature of the modern LLMs (Large Language Models) architectures made verification of such work quite complicated. Hence, the main objective of the research is to deduce the attributes of English poetry that contribute classification and misclassification of both human and AI poems and provide corroborating or contradicting evidence to the poetry distinguishability claim. For such characterizations, we propose a Zero-shot detection pipeline with a dataset consisting of both human and AI poems to verify the distinguishability of human and AI creation and extract the aforementioned crucial attributes for accurate classification. Extraction of such attributes provides benefits in two ways: firstly, it reduces the margin of training needed as only the poems based on misclassifying attributes need to be trained and fine tuned and finally provides a critical insight to the GenAI detection dilemma to strengthen the modern detection pipelines.
Chinese Translation
随着人工智能技术的进步,生成性人工智能(Generative AI, GenAI)与人类创作的文本几乎难以区分。此外,全球范围内对人工智能聊天机器人的标准化使得学术不端行为更加频繁。现有研究表明,GenAI 生成的诗歌即使在没有任何修改的情况下也最难以区分,因此,现代检测工具自然将 GenAI 诗歌视为人类创作。然而,这种论断的客观性需要通过现代检测工具进行验证,但诗歌的主观性以及现代大型语言模型(Large Language Models, LLMs)架构的黑箱特性使得这种工作的验证变得相当复杂。因此,本研究的主要目标是推导出英语诗歌的特征,这些特征有助于分类和误分类人类与人工智能诗歌,并提供支持或反驳诗歌可区分性主张的证据。为此,我们提出了一种零样本检测管道,使用包含人类和人工智能诗歌的数据集,以验证人类与人工智能创作的可区分性,并提取上述关键特征以实现准确分类。这些特征的提取在两个方面提供了好处:首先,它减少了所需训练的范围,因为只需对基于误分类特征的诗歌进行训练和微调,最终为解决 GenAI 检测困境提供了重要见解,以加强现代检测管道。
cs.CL / 6 / 2607.26228

Steering Instruction Hierarchies at Inference Time

推理时的指令层次引导
Zeng, Siqi, Lee, Sewoong, Zhao, Han, Hockenmaier, Julia
Abstract
Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override conflicting lower priority inputs from users or tools. Yet frontier LLMs often violate this hierarchy. We introduce V-Steer, a training-free inference time method that restores privileged influence by editing cached value vectors at prompt positions. Using direct logit attribution on the first next token prediction, V-Steer identifies heads where lower priority spans dominate privileged ones, then boosts privileged spans and suppresses conflicting lower priority spans through in-place multiplicative edits to cached V tensors. Since the method acts only on cached values, it remains compatible with fused attention backends and adds only a one time prefill overhead. Across models from 7B to 70B, this attribution guided intervention raises primary constraint accuracy from under 18% up to 92% on controlled role conflict benchmarks, and on broader instruction hierarchy evaluations substantially outperforms prompt only baselines while matching or exceeding SoTA training based methods on 3 of 4 scales of LLMs, with negligible decoding-speed overhead. The code is available at https://github.com/cindy2000sh/v-steer.
Chinese Translation
指令层次是语言模型部署的核心安全假设:优先级更高的输入(如系统提示)应覆盖来自用户或工具的相冲突的低优先级输入。然而,前沿的大型语言模型(LLMs)往往违反这一层次结构。我们提出了 V-Steer,这是一种无训练的推理时方法,通过编辑提示位置的缓存值向量来恢复特权影响。通过对下一个预测令牌的直接对数归因,V-Steer 识别出低优先级跨度主导特权跨度的头部,然后通过对缓存 V 张量进行就地乘法编辑来提升特权跨度并抑制冲突的低优先级跨度。由于该方法仅作用于缓存值,因此与融合注意力后端兼容,并且仅增加一次预填充开销。在从 7B 到 70B 的模型中,这种归因引导的干预将主要约束的准确率从不足 18% 提高到 92%,在控制角色冲突基准测试中表现出色,并且在更广泛的指令层次评估中显著超越仅基于提示的基线,同时在 4 种规模的 LLM 中与或超过基于训练的最先进方法(SoTA),且解码速度开销微乎其微。代码可在 https://github.com/cindy2000sh/v-steer 获取。
cs.CL / 7 / 2607.26249

A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States

美国网络广播宗教电台转录的大规模语料库
Bestvater, Samuel, Chapekis, Athena, Seets, Skyler, Lieb, Anna, Shah, Sono, Smith, Aaron
Abstract
Religious radio is a widespread but understudied form of mass communication in the United States, and content-level analysis of it has been constrained by the absence of large-scale transcript data. This Data Descriptor presents a corpus of transcribed English-language religious radio broadcasts captured from live webstreams over a one-month period in July 2025. Fifteen-minute segments were recorded on a rolling schedule from 785 distinct streams, which together rebroadcast the signals of more than two thousand AM and FM stations, yielding over 700,000 recordings and more than 60 million diarized lines of speech. Each recording was transcribed and speaker-diarized with an automated pipeline, and segmented and labeled by programming format and topic using a large language model. The corpus is organized as linked tables of stream metadata, recording metadata, and transcript lines. It supports descriptive study of religious broadcasting across regions and traditions, analysis of how social and political issues are discussed in religious media, and speech-processing research in an underrepresented domain.
Chinese Translation
宗教广播是美国一种广泛但研究不足的大众传播形式,而对其内容层面的分析受限于缺乏大规模的转录数据。本文数据描述符呈现了一个在2025年7月为期一个月内,从直播网络流中捕获的英语宗教广播转录语料库。该语料库从785个不同的流中以滚动时间表录制了十五分钟的片段,这些流共同转播了超过两千个AM和FM电台的信号,产生了超过700,000个录音和超过6000万条标注的讲话内容。每个录音都通过自动化流程进行了转录和发言者标注,并使用大型语言模型按节目格式和主题进行了分段和标记。该语料库组织为流元数据、录音元数据和转录行的链接表,支持对不同地区和传统的宗教广播的描述性研究,分析社会和政治问题在宗教媒体中的讨论,以及在一个代表性不足的领域内进行语音处理研究。
cs.CL / 8 / 2607.26250

Robostreet Flow: A Lightweight, Ultra-Low-Drag Electric Tractor and Four-Truck Hybrid Convoy Architecture for Minimum-Cost Point-to-Point Freight

Robostreet Flow:一种轻量级、超低阻力的电动拖拉机和四车混合车队架构,以实现最低成本的点对点货运
Wang, Wei, Wang, Yiru Veronika, Veeramalla, Sumukh, Liang, Xiaohui, Team, for the Robostreet Research
Abstract
Line-haul trucking costs are dominated by three comparably sized components: energy, driver labor, and equipment. Most efficiency technologies address only one component at a time. This paper presents Robostreet Flow, a freight architecture that jointly optimizes the vehicle, convoy formation, and operating model to minimize cost per ton-mile on high-volume point-to-point corridors. The Flow platform is a battery-electric 6x4 tractor with a teardrop single-seat cab and a drag coefficient of 0.35, approximately 40% below that of conventional Class 8 tractors. A carbon-composite monocoque and structurally integrated batteries reduce net vehicle weight by 50%. A 513 kWh tractor battery and a 340 kWh powered trailer battery provide a 500-mile single-charge range. Four Flow trucks operate as a coordinated convoy with a safety driver only in the lead vehicle, while three followers operate in SAE Level 4 automated mode. Computational fluid dynamics simulations show that close following at an 8 m gap reduces follower drag coefficients by 42-48% and follower peak frontal pressure by approximately a factor of four relative to the exposed lead vehicle. A longitudinal energy model calibrated to these results predicts fleet-average consumption of 1.27 kWh/mi in convoy, compared with 1.60 kWh/mi for an isolated vehicle, for a 20.5% energy saving. Electricity cost is approximately 17% of the equivalent diesel fuel cost. Amortizing one driver across four trucks and accounting for the additional payload enabled by lightweighting reduce operating cost from 9.4 to 4.1 cents per ton-mile, a 56% reduction relative to a diesel baseline. Sensitivity analysis, a hub-to-hub operating concept, and regulatory implications are also presented.
Chinese Translation
长途货运成本主要由三个相对规模相当的组成部分主导:能源、驾驶员劳动力和设备。大多数效率技术仅针对其中一个组成部分进行优化。本文提出了Robostreet Flow,一种货运架构,联合优化车辆、车队编队和运营模式,以最小化高流量点对点走廊的每吨英里成本。Flow平台是一款电池电动6x4拖拉机,配备流线型单座驾驶室,阻力系数为0.35,约比传统8级拖拉机低40%。碳复合材料单体结构和结构集成电池将净车辆重量减少了50%。513 kWh的拖拉机电池和340 kWh的供电拖车电池提供500英里单次充电范围。四辆Flow卡车作为协调车队运行,只有前车配备安全驾驶员,而三辆跟随车则在SAE 4级自动模式下运行。计算流体动力学模拟表明,在8米间距下紧跟行驶可将跟随车的阻力系数降低42-48%,并将跟随车的峰值前向压力相对于暴露的前车降低约四倍。基于这些结果校准的纵向能量模型预测,车队平均在车队行驶时的能耗为1.27 kWh/mi,而孤立车辆为1.60 kWh/mi,节能20.5%。电力成本约为等效柴油燃料成本的17%。在四辆卡车之间摊销一名驾驶员的成本,并考虑到轻量化所带来的额外载荷,使运营成本从每吨英里9.4美分降低到4.1美分,相对于柴油基线减少了56%。还介绍了敏感性分析、中心到中心的运营概念和监管影响。
cs.CL / 9 / 2607.26286

Evaluating Prompt Scope and Demonstration Similarity in Local LLM Machine Translation

评估本地大语言模型机器翻译中的提示范围和示例相似性
Arcan, Mihael
Abstract
Large language models (LLMs) are increasingly used as general-purpose translation systems, but their behavior is usually evaluated under a single prompt shape: translate one source sentence into one target language. In practice, users may ask for one target language, for several related languages at once, or for translations conditioned on examples. This paper studies prompt scope and demonstration selection as experimental variables for local LLM machine translation. We evaluate English-to-Romance and English-to-Germanic translation on the full FLORES devtest split for nine official European Union languages. We compare three local instruction-tuned LLMs, llama3.2:3b, mistral:latest, and qwen2.5:14b, against dedicated MT baselines from OPUS-MT and NLLB-200. We test zero-shot prompting and k=5 few-shot prompting with random, lexical-similarity, and embedding-similarity demonstration selection. We also compare single-target prompts with JSON-formatted family-scope prompts that request all languages in a family at once. Results show that dedicated MT systems remain strongest overall, especially for Germanic languages. Few-shot prompting helps mistral:latest and qwen2.5:14b, but hurts llama3.2:3b; embedding retrieval is best on average for the stronger LLMs, but its advantage over random and lexical examples is modest. Family-scope prompting is feasible for stronger local LLMs but exposes structured-output failures in smaller models. These findings motivate evaluating LLM translation not only by language pair and metric, but also by prompt scope, retrieval strategy, and multi-target compliance.
Chinese Translation
大型语言模型(LLMs)越来越多地被用作通用翻译系统,但它们的行为通常是在单一提示形状下进行评估:将一个源句子翻译成一种目标语言。在实际应用中,用户可能会请求一种目标语言、同时请求几种相关语言,或基于示例进行翻译。本文将提示范围和示例选择作为本地LLM机器翻译的实验变量进行研究。我们在九种官方欧盟语言的完整FLORES开发测试集上评估英语到罗曼语和英语到日耳曼语的翻译。我们比较了三种本地指令调优的LLM,即 llama3.2:3b、mistral:latest 和 qwen2.5:14b,与来自OPUS-MT和NLLB-200的专用机器翻译基线进行对比。我们测试了零-shot提示和k=5的few-shot提示,示例选择包括随机、词汇相似性和嵌入相似性。我们还比较了单一目标提示与JSON格式的家族范围提示,后者一次请求一个家族中的所有语言。结果显示,专用机器翻译系统在整体上仍然最强,尤其是在日耳曼语言方面。Few-shot提示对mistral:latest和qwen2.5:14b有帮助,但对llama3.2:3b则有负面影响;在较强的LLM中,嵌入检索的平均表现最佳,但其相较于随机和词汇示例的优势有限。家族范围提示对于较强的本地LLM是可行的,但在较小模型中暴露了结构化输出的失败。这些发现促使我们不仅通过语言对和指标评估LLM翻译,还要通过提示范围、检索策略和多目标合规性进行评估。
cs.CL / 10 / 2607.26300

AgentGUI: An Interface for Observing and Steering Long-Running AI Agents

AgentGUI:观察和引导长时间运行的人工智能代理的界面
Zhao, Xuan, Sohn, Jiwoong, Zheng, Qinyue, Moor, Michael
Abstract
AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematically lagging behind due to limited human-centered interfacing. Aiming to address this, we introduce AgentGUI, a user-friendly, locally hosted GUI for seamlessly observing and steering AI agents amid multiple concurrent, long-running sessions. AgentGUI features 1) rich agent trajectory visualizations, 2) effective manual and automated steering, and 3) integration with and coordination between open-source and frontier agent frameworks. A controlled user study demonstrates statistically significant reduction in the time it takes to identify key elements from agent traces (38% faster, p = 0.023). In a preliminary experiment, AgentGUI's automated drift prevention feature raises the task completion rate of small local agents by as high as 34pp across a 0.8B--9B model ladder (N=50 runs per model). AgentGUI is publicly available through its project website (https://agent-gui-project.github.io) and open-source repository (https://github.com/eth-medical-ai-lab/agent-gui), along with a demo video (https://youtube.com/watch?v=GSDyxN1gTF0).
Chinese Translation
人工智能代理在处理复杂的、长时间运行的任务方面越来越娴熟。随着自主能力的迅速提升,人类的监督由于人性化界面的局限性而系统性滞后。为了解决这一问题,我们推出了AgentGUI,这是一款用户友好、局部托管的图形用户界面,旨在无缝观察和引导在多个并发的长时间运行会话中的人工智能代理。AgentGUI具有以下特点:1)丰富的代理轨迹可视化,2)有效的手动和自动引导,以及3)与开源和前沿代理框架之间的集成与协调。一项受控用户研究表明,识别代理轨迹中的关键元素所需的时间显著减少(快38%,p = 0.023)。在一项初步实验中,AgentGUI的自动漂移预防功能使小型本地代理的任务完成率提高了高达34个百分点,覆盖0.8B至9B的模型梯度(每个模型50次运行)。AgentGUI已通过其项目网站(https://agent-gui-project.github.io)和开源代码库(https://github.com/eth-medical-ai-lab/agent-gui)公开发布,并附有演示视频(https://youtube.com/watch?v=GSDyxN1gTF0)。
cs.CL / 11 / 2607.26348

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

当合成用户失效时:大型语言模型模拟人类调查响应的跨领域基准
Chen, Zihan, Zhu, Di, Zheng, Lei Nico
Abstract
Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an evaluation framework for intelligent synthetic-user systems. A single protocol, run across four models spanning two families and an 8B-to-frontier capability range, is applied to two independent domains of real human-response data: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey). Every model is benchmarked against a suite of non-LLM baselines fit on held-out human data. Under demographic prompting and the survey-simulation protocols we test, two failures replicate across both domains, all four models, and both families. First, at the individual level no LLM beats even the strongest baseline; on cross-cultural values every model falls well below it, and the gap survives distance-aware and proper scoring. Second, models systematically over-determine demographics, treating identity as far more predictive of attitudes than it is among real people, a distortion present for nearly every question-group combination and robust to a coding-invariant measure. Neither failure is remedied by a larger, more capable model. A decision-impact analysis shows why this matters in practice: on a segment-targeting task the models inflate between-segment gaps two to fourfold, would direct a team to the wrong segment in half of U.S. and most cross-cultural cases, and manufacture segment splits that do not exist in real people. We make the cross-domain benchmark and the evaluation framework available on request, so that teams can determine in advance when synthetic-user evidence is safe for decision support and when it is not.
Chinese Translation
大型语言模型(LLMs)越来越多地被用作合成用户,作为人类受访者的替代品,其模拟答案影响产品、政策和市场决策。我们探讨这种替代在何时有效,何时失效,并将答案整理为一个智能合成用户系统的评估框架。我们在两个独立领域的真实人类响应数据上应用一个单一协议,该协议涵盖四个模型,跨越两个家族,并具有8B到前沿能力范围。研究领域包括美国一般社会态度(一般社会调查)和跨文化价值观(世界价值调查)。每个模型都与一组基于保留人类数据的非LLM基线进行基准测试。在我们测试的人口统计提示和调查模拟协议下,两个失效现象在两个领域、所有四个模型和两个家族中重复出现。首先,在个体层面,没有任何LLM超越最强基线;在跨文化价值观上,每个模型的表现远低于基线,并且这一差距在考虑距离和适当评分后依然存在。其次,模型系统性地过度确定人口统计特征,将身份视为对态度的预测因素,而这一点在真实人群中并不成立,这种扭曲在几乎所有问题组组合中都存在,并且对编码不变测量具有鲁棒性。更大、更强的模型并未能解决这两个失效问题。决策影响分析显示了这一问题在实践中的重要性:在一个细分目标任务中,模型将细分间的差距膨胀了两到四倍,导致团队在一半的美国案例和大多数跨文化案例中指向错误的细分,并制造出在真实人群中不存在的细分裂变。我们可以根据请求提供跨领域基准和评估框架,以便团队提前判断合成用户证据在决策支持中何时是安全的,何时不是。
cs.CL / 12 / 2607.26355

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

偏见的交响曲:探索多模态大语言模型中性别与乐器的关联
Farsi, Farhan, Bali, Shayan, Rad, Mohammad Heydari, Heidary, Negar, Rooein, Donya
Abstract
Large language models (LLMs) are increasingly embedded in everyday life and widely used for information seeking, raising concerns about their potential to perpetuate social biases and reinforce stereotypes. In this study, we investigate gender bias in LLMs through the lens of their associations with musical instruments. Building on social-science research on the cultural gender-typing of instruments, we introduce Symphony-Bias, a parallel multimodal dataset spanning text, vision, and audio. We evaluate ten multimodal models with diverse architectures and scales across 22 musical instruments, analyzing how they associate each instrument with three gender categories: {male, female, non-binary}, across three modalities: {text, vision, audio}. Our results show that 92\% of instrument-level outcomes align with prior social-science findings, with the harp and drums showing particularly consistent gendered associations across all evaluated models and modalities. We further find that alignment with social stereotypes is weakest in audio, stronger in vision, and strongest in text, suggesting that modality-specific representations can differentially amplify gendered associations with musical instruments.\footnote{The Symphony-Bias dataset will be publicly released upon acceptance of the paper.}
Chinese Translation
大型语言模型(LLMs)日益融入日常生活,并广泛用于信息检索,这引发了人们对其可能延续社会偏见和强化刻板印象的担忧。在本研究中,我们通过乐器的关联性来探讨LLMs中的性别偏见。基于社会科学对乐器文化性别分类的研究,我们引入了Symphony-Bias,一个涵盖文本、视觉和音频的平行多模态数据集。我们评估了十种具有不同架构和规模的多模态模型,分析它们如何将每种乐器与三种性别类别({男性,女性,非二元})关联,并在三种模态({文本,视觉,音频})中进行比较。我们的结果显示,92%的乐器级结果与先前的社会科学研究发现一致,其中竖琴和鼓在所有评估的模型和模态中表现出特别一致的性别关联。我们进一步发现,与社会刻板印象的一致性在音频中最弱,在视觉中较强,而在文本中最强,表明特定模态的表征可能会以不同方式放大乐器的性别关联。
cs.CL / 13 / 2607.26368

Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text

金融披露文本中的细粒度不一致性分类诊断
Kumar, Aman, Vidyaratne, Lasitha, Ghosh, Dipanjan D, Chakrabarti, Arnab, Farahat, Ahmed K
Abstract
Financial disclosures contain numerical claims, temporal statements, entity references, policy commitments, and risk descriptions that may conflict in qualitatively different ways. Detecting a conflict is only the first step: review workflows may also need to determine its type, since numerical, temporal, referential, factual, and normative inconsistencies require different evidence and downstream checks. We study this problem as fine-grained inconsistency classification. Using a fixed 5,940-instance snapshot of SBID-FD, a synthetic financial-disclosure benchmark with 11 inconsistency labels and paired reference evidence spans, we compare frozen embedding classifiers, fine-tuned encoders, evidence-augmented classifiers, prompted large language models, and LoRA-adapted generative models under a shared evaluation protocol. A fine-tuned 300M encoder reaches 61.9% accuracy, compared with 61.5% for a LoRA-adapted Qwen3.5-9B model and 61.3% for GPT-5.4. Because these systems differ in architecture, supervision, training objective, and input format, we interpret this as a practical efficiency result for compact supervised encoders rather than a controlled conclusion about model scale. Supplying gold evidence spans improves the fine-tuned encoder to 65.3%, whereas automatically predicted spans recover a meaningful but incomplete share of that gain, indicating that localization quality remains a bottleneck. Class-level analyses show that Referential inconsistencies are especially sensitive to localization quality, while Factual and Logical inconsistencies remain difficult even when the relevant evidence is provided. Together, the oracle, distractor, and per-class analyses separate localization errors from residual type-discrimination errors, indicating that progress requires both stronger evidence extraction and better reasoning over closely related inconsistency categories.
Chinese Translation
金融披露包含数值声明、时间陈述、实体引用、政策承诺和风险描述,这些内容可能以不同的定性方式发生冲突。检测冲突只是第一步:审查工作流程还可能需要确定其类型,因为数值、时间、引用、事实和规范性不一致性需要不同的证据和后续检查。我们将这个问题研究为细粒度不一致性分类。使用固定的5,940实例快照的SBID-FD,这是一个具有11种不一致性标签和配对参考证据跨度的合成金融披露基准,我们在共享评估协议下比较了冻结嵌入分类器、微调编码器、增强证据分类器、提示的大型语言模型和LoRA适配的生成模型。微调的300M编码器达到了61.9%的准确率,而LoRA适配的Qwen3.5-9B模型为61.5%,GPT-5.4为61.3%。由于这些系统在架构、监督、训练目标和输入格式上存在差异,我们将其解释为紧凑监督编码器的实际效率结果,而不是关于模型规模的控制结论。提供黄金证据跨度将微调编码器的准确率提高到65.3%,而自动预测的跨度恢复了这一增益的有意义但不完整的部分,表明定位质量仍然是一个瓶颈。类别级分析表明,引用不一致性对定位质量特别敏感,而事实和逻辑不一致性即使在提供相关证据的情况下也仍然困难。综合来看,oracle、干扰项和每类分析将定位错误与残余类型区分错误分开,表明进展需要更强的证据提取和更好的推理能力,以处理密切相关的不一致性类别。
cs.CL / 14 / 2607.26375

(Im)Paired Programming: Coding Agents Improve Productivity but Harm Understanding

(不)配对编程:编码代理提高生产力但损害理解
Balepur, Nishant, Baumler, Connor, Chen, Valerie, Choi, Eunsol, Rudinger, Rachel, Boyd-Graber, Jordan Lee
Abstract
Coding agents (e.g., Cursor) improve developer productivity by optimizing task completion, but shifting users from writing code to prompting and reviewing may harm their understanding, impeding oversight, learning, and communication. To probe this, we have 54 students create a website with one of two AI systems: an agent that edits user code; or a chatbot where users write code alone or adapt generic code snippets. We test understanding via comprehension questions and a task where users extend their code without agents, showing: (1) While agents aid initial task completion, they harm users' code comprehension and thus do not prepare users to extend their code; (2) Low-effort agent interaction types, like copy+paste prompts and auto-accepted edits, are linked with lower comprehension; and (3) Despite self-reported weaker understanding, users still prefer coding agents because they are quick and easy to use. While users stay in the loop for coding workflows, understanding should not be forgotten. Towards this goal, we distill our analyses into future research directions for coding agent developers: dissuading low-effort prompting, creating readable code, and promoting active engagement.
Chinese Translation
编码代理(例如,Cursor)通过优化任务完成来提高开发者的生产力,但将用户从编写代码转变为提示和审查可能会损害他们的理解,阻碍监督、学习和沟通。为了探讨这一点,我们让54名学生使用两种AI系统之一创建一个网站:一种是编辑用户代码的代理;另一种是用户单独编写代码或调整通用代码片段的聊天机器人。我们通过理解问题和一个任务来测试理解,该任务要求用户在没有代理的情况下扩展他们的代码,结果显示:(1)虽然代理有助于初始任务的完成,但它们会损害用户的代码理解,因此未能为用户扩展代码做好准备;(2)低努力的代理交互类型,如复制+粘贴提示和自动接受的编辑,与较低的理解水平相关;(3)尽管用户自报理解较弱,但他们仍然偏好编码代理,因为它们快速且易于使用。虽然用户在编码工作流程中保持参与,但理解不应被忽视。为此,我们将我们的分析提炼为编码代理开发者的未来研究方向:劝阻低努力的提示,创建可读的代码,并促进积极参与。
cs.CL / 15 / 2607.26389

Misalignment Has a Personality: A Big Five Account of Emergent Misalignment

不一致性具有个性:一种关于新兴不一致性的五大人格理论
Rahman, Hasibur, Desai, Smit
Abstract
Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment through a mechanism that remains debated. We provide an interpretable account: in the models and corpora we study, misalignment behaves like a shift in personality. Prior work extracts activation directions for character traits from a single binary contrast, which can separate or steer behavior without establishing a calibrated scale. We instead extract personality vectors for the Big Five using a graded, three-level intervention and validate them on two open-weight models. The three levels are linearly ordered, with Cohen's d values of up to 6.2; the vectors transfer zero-shot and trait-specifically to an independent corpus; and their effects are strongest within a middle-layer band. Applied to training data, the vectors reveal that misaligned corpora across eight domains share a common Big Five signature: lower agreeableness and conscientiousness, together with higher extraversion and neuroticism. This signature is recovered by both models with a correlation of r = 0.94. Fine-tuning imprints the same profile, shifting the model's generations along the corresponding signature, with r = 0.83 using activation-based measurements and r = 0.90 using a text-based judge, while also shifting internal activations with r = 0.69. The same vectors characterize sycophancy as high extraversion and low conscientiousness rather than excess agreeableness, a distinction that a single direction cannot capture. Calibrated personality vectors transform an opaque safety phenomenon into a human-legible diagnostic profile.
Chinese Translation
在包含狭窄缺陷的数据上微调语言模型,例如不安全的代码或错误的数学答案,可能通过一种仍在争论中的机制导致广泛的不一致性。我们提供了一种可解释的解释:在我们研究的模型和语料库中,不一致性表现得像个性的变化。先前的研究通过单一的二元对比提取个性特征的激活方向,这可以在不建立校准尺度的情况下分离或引导行为。相反,我们使用分级的三层干预提取五大人格的个性向量,并在两个开放权重模型上进行验证。这三层是线性排序的,Cohen's d 值高达 6.2;这些向量在零样本和特征特定的情况下转移到一个独立的语料库;并且它们的效果在中间层带内最强。应用于训练数据,这些向量揭示了跨八个领域的不一致语料库共享一个共同的五大人格特征:较低的宜人性和责任感,以及较高的外向性和神经质。这个特征在两个模型中都得到了恢复,相关性为 r = 0.94。微调印记了相同的特征,沿着相应的特征移动模型的生成,使用基于激活的测量相关性为 r = 0.83,使用基于文本的评估者相关性为 r = 0.90,同时也改变了内部激活,相关性为 r = 0.69。相同的向量将谄媚特征化为高外向性和低责任感,而不是过度的宜人性,这一区别是单一方向无法捕捉的。经过校准的个性向量将一个不透明的安全现象转化为人类可理解的诊断特征。
cs.CL / 16 / 2607.26397

Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification

推理之前的知识:EC-Reason-Bench,一个无训练的诊断基准用于大规模语言模型的酶分类
Li, Linyu, Jin, Zhi, Zhang, Yichi, Jin, Dongming, He, Yuanpeng, Zhang, Huanyao, Zhang, Xuan, Luosang, Gadeng, Tashi, Nyima
Abstract
Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification. Existing benchmarks expose an anomaly: general LLMs often get the coarse first level right, yet once asked for a complete EC number their accuracy at levels two through four drops to almost zero, while specialized models and tools stay usable. We propose EC-Reason-Bench, a training-free, diagnostic evaluation protocol built to answer two questions: why general LLMs score close to nothing on EC number prediction, and how much of that loss can be recovered without updating a single weight. We break enzyme classification ability into four orthogonal levers that can each be measured on their own: output structure, external knowledge, reasoning structure, and reasoning robustness. We test each lever with an inference-time method against a shared zero-shot baseline reproducing previously reported near-zero performance. Experiments with several strong reasoning LLMs yield four main findings. First, external knowledge is decisive and must precede reasoning: uniformly low closed-book performance rises sharply with open-book access, narrowing model gaps. Second, in closed-book settings, whether cascading and chain-of-thought help or hurt depends on a model's tendency to abstain. Third, once evidence is available the aggregate score of the best LLM setting is indistinguishable from simply voting the EC numbers of the nearest retrieved neighbors; that tie is an artifact of averaging, and it hides a large gain on adversarial evidence set against an equally large loss on multi-functional enzymes. Reasoning over evidence therefore acts as an arbiter of conflicting neighbors rather than as a source of knowledge, and no single-number leaderboard can see it. Fourth, accuracy obeys a law of homology availability.
Chinese Translation
酶功能预测是一种层次化、知识密集型的蛋白质功能分类形式。现有基准暴露出一个异常现象:通用的大规模语言模型(LLMs)通常能够正确识别粗略的第一层级,但一旦被要求提供完整的酶分类号(EC number),其在第二至第四层级的准确率几乎降至零,而专门化的模型和工具仍然可用。我们提出了EC-Reason-Bench,一个无训练的诊断评估协议,旨在回答两个问题:为什么通用LLMs在EC号预测上的得分接近于零,以及在不更新任何权重的情况下,能恢复多少损失。我们将酶分类能力分解为四个正交的杠杆,每个杠杆都可以单独测量:输出结构、外部知识、推理结构和推理稳健性。我们使用推理时的方法对每个杠杆进行测试,基于共享的零样本基线,重现之前报告的近零性能。对几种强推理LLMs的实验得出了四个主要发现。首先,外部知识是决定性的,必须在推理之前:统一较低的闭卷表现随着开放卷访问的增加而急剧上升,缩小了模型之间的差距。其次,在闭卷设置中,级联和链式思维是否有助于或有害取决于模型的回避倾向。第三,一旦有证据可用,最佳LLM设置的总得分与简单投票最近检索到的邻居的EC号没有区别;这种平局是平均的伪影,它掩盖了在对抗证据集上的巨大收益,以及在多功能酶上的同样巨大损失。因此,基于证据的推理充当了冲突邻居的仲裁者,而不是知识的来源,任何单一数字的排行榜都无法反映这一点。第四,准确性遵循同源性可用性的法则。
cs.CL / 17 / 2607.26410

Voice Memory for Agentic Speech Recognition

用于主动语音识别的语音记忆
Yang, Chao-Han Huck, Chen, Zih-Ching, Zelasko, Piotr, Chen, Zhehuai, Balam, Jagadeesh, Ginsburg, Boris
Abstract
We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best. Asynchronously, a score-gated optimizer revises that file through bounded edits, accepting an edit only when it strictly improves a held-out score. Extended from classical ASR-LM framework, we refer this split the listener-thinker architecture; the two roles are coupled only through the memory, so no weights change and the learned skill stays auditable and portable. Restraint turns out to be the operative skill this loop discovers: unconstrained generative error correction (GER) over-corrects, breaking correct tokens on up to 64% of its edits on financial news, and Voice Memory, reduces this rate to 35%. Across ten HyPoradise domains with an open corrector, Voice Memory, lowers weighted word error rate from 8.36% to 7.52% (7.47% with three added in-context examples) without regressing any dataset below its 1-best baseline; gains concentrate where recoverable headroom is largest, including air-travel commands (8.40% to 3.40%) and noisy far-field speech (CHiME-4, 12.69% to 10.46%). The memory transfers across corrector families and adds zero parameters to the inference path. A demo and example code are provided for future studies.
Chinese Translation
我们提出了语音记忆(Voice Memory),一种仅用于推理的主动语音识别方案:在流式处理时,一个冻结的修正器读取单个领域内的记忆文件,并决定每个发话是否对假设采取行动或放弃并保留最佳结果。在此过程中,一个基于分数的优化器通过有限的编辑异步修订该文件,仅在严格改善保留分数时接受编辑。该方法扩展自经典的ASR-LM框架,我们将其称为听者-思考者架构(listener-thinker architecture);这两个角色仅通过记忆相互关联,因此权重不发生变化,学习的技能保持可审计和可移植。结果表明,这个循环发现的操作技能是克制:无约束的生成错误修正(GER)过度修正,在金融新闻的编辑中,最多有64%的正确标记被破坏,而语音记忆将这一比例降低到35%。在十个HyPoradise领域中,使用开放修正器的语音记忆将加权字错误率从8.36%降低到7.52%(在添加三个上下文示例后为7.47%),且没有任何数据集的表现低于其最佳基线;增益集中在可恢复的余地最大之处,包括航空旅行指令(从8.40%降至3.40%)和嘈杂的远场语音(CHiME-4,从12.69%降至10.46%)。该记忆可以在不同的修正器家族之间转移,并且在推理路径中不增加任何参数。我们提供了演示和示例代码以供未来研究使用。
cs.CL / 18 / 2607.26448

Mergeable Model-Side Aggregation States for Long-Context Language Models

可合并的模型侧聚合状态用于长上下文语言模型
Song, Dachuan, Yin, Junyu, Hu, Zechen, Wang, Xuan
Abstract
A known limitation of long-context language models is their increasingly unreliable performance in non-additive, set-based aggregation as context length grows. Examples include cardinality estimation, set relationships, and grouped statistics, which widely exist in logs, program outputs, tables, and multi-turn conversations. To provide the aggregation state required by these tasks, we introduce a model-side aggregation interface that maintains compact Hash-based HyperLogLog (HLL) sketch states alongside a frozen language model. While the model processes the context, an extractor maps each relevant record to a canonical identity. The identity is then hashed and updates the HLL state. These states can be merged across context segments and/or read out directly for downstream reasoning, avoiding an additional generate-execute-return cycle. We validate the proposed approach by setting the HLL state size as 2 KiB (2,048 registers), which does not increase with context length or set cardinality. In a distinct-count experiment involving one million records, the mean relative error was 1.6%. In a separate merge test, states built from as many as 256 segments produced exactly the same readout as a single pass over the same stream. On 3,969 aggregate-then-reason tasks from 174 source windows, the fixed-budget interface reached 99.2% accuracy on Gemma 4 (31B, BF16), compared with 100.0% under exact aggregation; the paired gap was 0.8 percentage points (95% window-cluster CI: 0.5-1.3 points). On a matched set of 174 items, our method improved over direct full-context reasoning by 63.2 points on Qwen and 56.3 points on Gemma. The corresponding gains over chain-of-thought (CoT) reasoning were 60.9 and 63.2 points, respectively. On a fixed 1,200-task Oolong-Synth subset, our method reached 91.1% on Qwen and 99.3% on Gemma. Code is available at https://github.com/songdc98/sketchops.
Chinese Translation
长上下文语言模型的一个已知局限性是,随着上下文长度的增加,它们在非加法、基于集合的聚合中的性能越来越不可靠。此类例子包括基数估计、集合关系和分组统计,这些广泛存在于日志、程序输出、表格和多轮对话中。为了提供这些任务所需的聚合状态,我们引入了一种模型侧聚合接口,该接口在一个冻结的语言模型旁边维护紧凑的基于哈希的 HyperLogLog (HLL) 草图状态。在模型处理上下文时,提取器将每个相关记录映射到一个规范身份。然后对该身份进行哈希并更新 HLL 状态。这些状态可以在上下文段之间合并和/或直接读取以进行下游推理,从而避免额外的生成-执行-返回周期。我们通过将 HLL 状态大小设置为 2 KiB(2,048 个寄存器)来验证所提出的方法,该大小不会随着上下文长度或集合基数的增加而增加。在一个涉及一百万条记录的独特计数实验中,平均相对误差为 1.6%。在一个单独的合并测试中,从多达 256 个段构建的状态与对同一流的单次遍历产生的读出完全相同。在来自 174 个源窗口的 3,969 个聚合-推理任务中,固定预算接口在 Gemma 4(31B,BF16)上达到了 99.2% 的准确率,而在精确聚合下为 100.0%;配对差距为 0.8 个百分点(95% 窗口聚类置信区间:0.5-1.3 个百分点)。在一组匹配的 174 个项目上,我们的方法在 Qwen 上比直接全上下文推理提高了 63.2 个百分点,在 Gemma 上提高了 56.3 个百分点。与链式思维(CoT)推理相比,相应的增益分别为 60.9 和 63.2 个百分点。在一个固定的 1,200 任务 Oolong-Synth 子集上,我们的方法在 Qwen 上达到了 91.1%,在 Gemma 上达到了 99.3%。代码可在 https://github.com/songdc98/sketchops 获取。
cs.CL / 19 / 2607.26455

ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models

ForgetBench:语言模型中长期参数记忆遗忘动态的基准测试
Gu, Ruxi, Zhang, Zhenliang, Wang, Wei
Abstract
Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain previously acquired knowledge under repeated updates remains insufficiently understood. Existing evaluation paradigms primarily focus on single-step reasoning or static knowledge editing, which fail to capture the temporal dynamics of knowledge retention and degradation during continual model modification. In this work, we propose ForgetBench, a benchmark designed to systematically characterize forgetting behavior in LLMs under continual knowledge editing. ForgetBench introduces two complementary evaluation paradigms, namely concept-based QA and scenario-based QA, to disentangle isolated factual retention from structured relational knowledge preservation. Building upon a sequential editing framework, we construct temporally ordered knowledge streams and evaluate model behavior across multiple editing stages. To quantitatively analyze long-term retention dynamics, we further introduce a unified evaluation framework that models knowledge evolution over time, enabling the measurement of temporal decay, retention strength, and cross-instance stability. Extensive experiments across diverse models and editing methods demonstrate that existing approaches fail to strike a balance between long-term retention and generalization quality. Our findings highlight the need for more robust memory mechanisms that can effectively acquire, update, and preserve knowledge over time in future LLMs. Code will be released upon acceptance.
Chinese Translation
大型语言模型(LLMs)在知识获取和推理方面表现出强大的能力,但它们在重复更新下保持先前获取知识的能力仍然未得到充分理解。现有的评估范式主要集中于单步推理或静态知识编辑,未能捕捉到在持续模型修改过程中知识保留和退化的时间动态。在本研究中,我们提出了ForgetBench,一个旨在系统性表征LLMs在持续知识编辑下遗忘行为的基准测试。ForgetBench引入了两种互补的评估范式,即基于概念的问答(QA)和基于场景的问答(QA),以区分孤立事实保留与结构化关系知识保留。基于顺序编辑框架,我们构建了时间顺序的知识流,并在多个编辑阶段评估模型行为。为了定量分析长期保留动态,我们进一步引入了一个统一的评估框架,该框架对知识随时间的演变进行建模,使得能够测量时间衰减、保留强度和跨实例稳定性。针对多种模型和编辑方法的广泛实验表明,现有方法未能在长期保留和泛化质量之间取得平衡。我们的研究结果强调了未来LLMs中需要更强大的记忆机制,以有效获取、更新和保留知识。代码将在接受后发布。
cs.CL / 20 / 2607.26470

CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG

CMT-RAG:用于多轮多跳检索增强生成的互补记忆痕迹
Zhou, Lang, Chen, Yingjian, Li, Shuxuan, Lin, Kun-Yu, Zhao, Zhilin
Abstract
Multi-turn information-seeking conversations require both multi-hop reasoning and long-range dependency tracking across turns. However, existing RAG systems typically represent conversational memory as raw dialogue history, rewritten queries, or unstructured summaries, making it difficult to recover the specific prior reasoning steps and evidence required for follow-up queries. Our key insight is to align conversational memory with retrieval by representing dialogue context as sub-question-level reasoning traces. Building on this insight, we introduce MuMu-QA, a benchmark for multi-turn multi-hop RAG with explicit cross-turn sub-question dependency annotations, and CMT-RAG, a complementary memory framework for this setting. At each turn, CMT-RAG employs a state-space trace generator, whose recurrent state serves as runtime memory, to incorporate recent conversational context and decompose the current query into structured trace drafts containing retrieval-oriented sub-questions and dependencies on earlier traces. It then grounds these drafts with retrieved evidence and stores them as persistent memory traces in a session-level DAG, enabling future turns to efficiently recover relevant prior reasoning and evidence. Experiments on MuMu-QA and corpus-level RAG benchmarks show that CMT-RAG consistently outperforms five categories of RAG baselines in answer accuracy.
Chinese Translation
多轮信息检索对话需要在对话轮次之间进行多跳推理和长程依赖跟踪。然而,现有的检索增强生成(RAG)系统通常将对话记忆表示为原始对话历史、重写查询或非结构化摘要,这使得难以恢复后续查询所需的具体先前推理步骤和证据。我们的关键见解是通过将对话上下文表示为子问题级别的推理痕迹,将对话记忆与检索对齐。在此基础上,我们引入了MuMu-QA,这是一个具有明确跨轮子问题依赖注释的多轮多跳RAG基准,以及CMT-RAG,这是针对该设置的互补记忆框架。在每一轮中,CMT-RAG使用状态空间痕迹生成器,其递归状态作为运行时记忆,来整合最近的对话上下文,并将当前查询分解为包含检索导向子问题和对早期痕迹依赖的结构化痕迹草稿。然后,它用检索到的证据来支持这些草稿,并将其存储为会话级有向无环图(DAG)中的持久记忆痕迹,从而使未来的轮次能够高效地恢复相关的先前推理和证据。在MuMu-QA和语料库级RAG基准上的实验表明,CMT-RAG在答案准确性上始终优于五类RAG基线。
cs.CL / 21 / 2607.26497

Which RAG Paradigm Wins at Scale? A Scaling Study of Retrieval-Augmented Generation Paradigms

哪种检索增强生成范式在规模上胜出?检索增强生成范式的规模研究
Wang, Pengyu, Xu, Benfeng, Wang, Shaohan, Zeng, Xin, Wu, Huarui, Zhang, Lei, Zhang, Licheng
Abstract
Retrieval-augmented generation (RAG) methods range from lexical and dense retrieval to graph-based indexing and agentic search. They are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling unclear. To bridge this gap, we present a controlled corpus-scaling study of these four paradigms. A ladder of 28 strictly nested tiers grows from roughly 1,000 to 512,000 documents while questions and a fixed bedrock of relevant and adversarial documents remain unchanged. Under one reader and judging protocol, we measure official accuracy, construction and query tokens, and latency. Our experimental results show that BM25 scales best in this controlled setting: it defines the low-cost end of the Pareto frontier at every measured tier and leads accuracy from mid-scale onward, without LLM-based construction. The File-System Agent matches or slightly exceeds BM25 at the smallest tiers but uses 39 times more query tokens per answer at the bedrock and falls nearly 20 points behind at full scale. A matched retrieval swap reverses this failure: Agent+BM25 scores 69.4 at full scale, versus 36.9 for raw-file agency and 54.8 for native BM25 on the same 150 questions. Graph-based RAG hits a construction wall: its heaviest builders use up to 24.6 generative LLM tokens per indexed corpus token yet stop within the first 2% of the full corpus, while scalable variants remain less accurate than BM25 at shared tiers.
Chinese Translation
检索增强生成(RAG)方法涵盖了从词汇和密集检索到基于图的索引和智能搜索的多种形式。它们通常在不同的基准测试中以单一语料库规模进行评估,因此其准确性与成本的规模关系尚不明确。为了解决这一问题,我们对这四种范式进行了受控的语料库规模研究。研究中设定了28个严格嵌套的层级,从大约1,000个文档增长到512,000个文档,同时问题和一组固定的相关及对抗性文档保持不变。在一个阅读者和评判协议下,我们测量了官方准确性、构建和查询的标记数以及延迟。我们的实验结果表明,在这一受控环境中,BM25的扩展效果最佳:在每个测量层级上,它定义了帕累托前沿的低成本端,并在中等规模以上的准确性上领先,而没有使用基于大语言模型(LLM)的构建。文件系统代理在最小层级与BM25相匹配或略有超出,但在基岩层级每个答案使用的查询标记数是BM25的39倍,并且在全规模时落后近20分。匹配的检索交换扭转了这一失败:Agent+BM25在全规模时得分为69.4,而原始文件代理为36.9,原生BM25在同150个问题上的得分为54.8。基于图的RAG遇到了构建瓶颈:其最重的构建者每个索引语料库标记使用多达24.6个生成LLM标记,但在完整语料库的前2%内停止,而可扩展的变体在共享层级上的准确性仍低于BM25。
cs.CL / 22 / 2607.26555

Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation

探讨检测器的失效:通过专家引导的互助蒸馏缩小尾部领域差距
Feng, Xuan, Liu, Guihong, Gu, Tianlong, Zhao, Shuai, Wang, Xuemin, Bin, Chenzhong, Liu, Yang, An, Bo
Abstract
Multimodal fake news detectors often generalize poorly across domains because they learn to trust unreliable evidence: domain-specific shortcuts amplified by imbalanced data and semantically inconsistent text-image pairs that make cross-modal evidence unreliable. We propose Expert-Guided Mutual Distillation (EGMD), which learns what evidence to trust across the prediction pipeline. At the input level, input-level calibration encodes pair-level coherence as a shared gain before fusion. At the representation level, an expert-guided teacher aligns domain statistics and encourages domain-specific patterns to concentrate in specialized experts. At the decision level, prototype-anchored domain-specific students use mutual learning and dual-channel distillation to inherit the teacher's feature geometry and calibrated predictions while discouraging local domain priors. We further construct Weibo_Balanced, a domain-balanced benchmark that isolates the effect of imbalance on generalization. Across four datasets in two languages, EGMD achieves state-of-the-art accuracy while reducing domain bias by up to 57.3%.
Chinese Translation
多模态假新闻检测器在不同领域中的泛化能力往往较差,因为它们学习信任不可靠的证据:由不平衡数据和语义不一致的文本-图像对放大了领域特定的捷径,从而使跨模态证据变得不可靠。我们提出了专家引导的互助蒸馏(Expert-Guided Mutual Distillation, EGMD),该方法学习在预测流程中信任哪些证据。在输入层面,输入级别的校准将配对级别的一致性编码为融合前的共享增益。在表示层面,专家引导的教师对齐领域统计数据,并鼓励领域特定模式集中在专业专家中。在决策层面,原型锚定的领域特定学生利用互学习和双通道蒸馏继承教师的特征几何和校准预测,同时抑制局部领域先验。我们进一步构建了 Weibo_Balanced,这是一个领域平衡的基准,隔离不平衡对泛化的影响。在两个语言的四个数据集中,EGMD 实现了最先进的准确率,同时将领域偏差降低了多达 57.3%。
cs.CL / 23 / 2607.26604

WikiLoop: Jointly Learning to Build and Navigate Agent-Native Wikis with Downstream Feedback

WikiLoop:通过下游反馈共同学习构建和导航代理原生维基
Ming, Haoliang, Li, Feifei, Que, Wenhui
Abstract
Knowledge-base construction and querying are typically optimized in isolation: retrieval-augmented agents operate over a fixed, externally maintained index, whereas construction receives no signal from downstream use. We present WikiLoop, a feedback-coupled framework that jointly learns to build and navigate an agent-native Wiki, a persistent linked-page knowledge base designed for machine navigation. A role-conditioned shared policy supports two interfaces: a Navigator retrieves evidence from the Wiki to answer queries, and a Builder proposes structured edits evaluated through downstream navigation. The Navigator follows a sufficiency-before-efficiency objective that applies retrieval-cost penalties only after full evidence has been collected. The Builder learns from utility differences: a frozen Navigator scores each candidate edit by its change in downstream performance, while a guard penalty discourages regressions on unrelated queries. Training combines sequential role-specific optimization with a final joint stage over role-homogeneous batches. With Qwen3.5-9B as the common backbone, WikiLoop reaches 62.6 aggregate Answer Correctness on AuthTrace, 6.3 points above LLM-Wiki, base, with the largest gains on multi-document queries. Controlled comparisons support the intended effects of both objectives, and the learned edits remain useful to a held-out Navigator. Paired comparisons indicate that the final shared policy largely retains both role-specific capabilities, improves Navigator and end-to-end Answer Correctness by 0.4 points relative to the corresponding specialist references, and consolidates both interfaces into one model. Without dataset-specific training, WikiLoop also improves over the same-backbone LLM-Wiki, base on HotpotQA and MuSiQue.
Chinese Translation
知识库的构建和查询通常是孤立优化的:增强检索的代理在固定的、外部维护的索引上操作,而构建过程则没有来自下游使用的信号。我们提出了WikiLoop,这是一种反馈耦合框架,能够共同学习构建和导航一个代理原生维基,这是一个为机器导航设计的持久性链接页面知识库。一个基于角色的共享策略支持两个接口:导航器从维基中检索证据以回答查询,而构建者则提出结构化的编辑,通过下游导航进行评估。导航器遵循“充分性优先于效率”的目标,仅在收集完整证据后才对检索成本施加惩罚。构建者从效用差异中学习:一个冻结的导航器通过下游性能的变化对每个候选编辑进行评分,而一个守卫惩罚则抑制与无关查询的退步。训练结合了顺序的角色特定优化和最终的角色同质批次的联合阶段。在以Qwen3.5-9B作为共同基础的情况下,WikiLoop在AuthTrace上达到了62.6的综合答案正确率,比LLM-Wiki的基础版本高出6.3分,且在多文档查询上获得了最大的提升。受控比较支持两个目标的预期效果,而学习到的编辑对于保留的导航器仍然是有用的。配对比较表明,最终的共享策略在很大程度上保留了角色特定能力,相较于相应的专家参考,导航器和端到端答案正确率提高了0.4分,并将两个接口整合为一个模型。在没有特定数据集训练的情况下,WikiLoop在HotpotQA和MuSiQue上也优于相同基础的LLM-Wiki基础版本。
cs.CL / 24 / 2607.26627

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes

重新审视投机解码中的有损验证:机制、权衡与失败模式
Wang, Tianyu, Zhou, Yuxuan, Wang, Wenbin, Li, Heng, Xiao, Zikai, Shang, Junyuan
Abstract
Speculative Decoding (SD) accelerates large language model inference by allowing a lightweight draft model to propose tokens that are subsequently verified in parallel by a larger target model. Recent approaches introduce lossy verification schemes to further improve efficiency by relaxing strict distributional matching. Yet such relaxation silently rewrites the decoding distribution, and the resulting acceleration can come at the cost of unstable, sometimes severely degraded generation quality. In this work, we present a principled analysis of the distributions induced by lossy verification methods. We show that many seemingly distinct approaches differ only superficially and can be classified into two categories: truncation-based verification and collaborative verification. We further construct a diagnostic evaluation framework across curated benchmarks. For truncation-based methods, we identify a fundamental pitfall: performance can degrade significantly compared to the true truncation sampling baseline due to distributional distortion. For collaborative verification, we uncover a key principles: controlling the overshoot of draft probabilities relative to target probabilities is essential to prevent low-quality outputs. Our code is available at https://github.com/ZhouYuxuanYX/Fast-HSD.
Chinese Translation
投机解码(Speculative Decoding, SD)通过允许轻量级草稿模型提出令牌,并由更大的目标模型并行验证,从而加速大型语言模型的推理。最近的方法引入了有损验证方案,通过放宽严格的分布匹配来进一步提高效率。然而,这种放松悄然重写了解码分布,导致的加速可能以不稳定的生成质量为代价,有时甚至严重下降。在本研究中,我们对有损验证方法所诱导的分布进行了原则性分析。我们表明,许多看似不同的方法仅在表面上有所区别,可以分为两类:基于截断的验证和协作验证。我们进一步构建了一个跨精选基准的诊断评估框架。对于基于截断的方法,我们识别出一个基本陷阱:由于分布扭曲,性能可能显著低于真实的截断采样基线。对于协作验证,我们揭示了一个关键原则:控制草稿概率相对于目标概率的超出是防止低质量输出的关键。我们的代码可在 https://github.com/ZhouYuxuanYX/Fast-HSD 获取。
cs.CL / 25 / 2607.26637

Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability

基于文件系统的LLM代理记忆:组织、演变与可持续性
Zhou, Sizhe, Yu, Sheldon, Wei, Hui, Wu, Junda, Ouyang, Siru, Jiao, Yizhu, Pan, Shijia, McAuley, Julian, Zhang, Yu, Yu, Tong, Han, Jiawei
Abstract
Deployed LLM agents increasingly keep their long-term memory as a filesystem: a directory tree of markdown files that the agent itself reads, writes, and reorganizes through generic file tools. Yet research has largely passed over this medium: prior systems design bespoke memory representations and study retrieval over them, leaving the default's two working assumptions untested: that an agent can keep a growing store organized as memories accumulate, conflict, and go stale, and that this organization pays. We present the first systematic exploration of filesystem-based memory for LLM agents. We formalize the setting as three roles around one memory filesystem: a management agent integrates and organizes incoming content, a search agent answers queries with cited sources, and an execution agent supplies task trajectories that are distilled into skills, unifying declarative memory and skills in a single store. Across long-conversation benchmarks and embodied tasks, we vary memory shape (agent-organized hierarchy, verbatim dump, chunk retrieval), stream scale, tool harness (sandboxed shell, memory-tool-style functions, varied search tooling), and the strengths of the management and search agents, tracking answer quality, cost, and store health as memory grows. What organization reliably buys is search economy: organized stores roughly halve retrieval cost where material is large. Today's agents, however, fall short of the default's promise: in our growth study, organization erodes for all but the strongest management agent, and no agent we measure converts organization itself into better answers. And the model is not the only lever over a store's shape: changing the tool set alone reshapes the store as strongly as swapping the model. The study turns the filesystem default from an assumption into a design space for agent memory.
Chinese Translation
部署的LLM代理越来越多地将其长期记忆保存在文件系统中:一个由代理自身通过通用文件工具读取、写入和重组的Markdown文件目录树。然而,研究在很大程度上忽视了这一媒介:先前的系统设计了定制的记忆表示并研究其检索,未对默认的两个工作假设进行验证:即代理能够在记忆不断积累、冲突和过时的情况下保持一个有序的增长存储,以及这种组织是有价值的。我们首次系统性地探索了基于文件系统的LLM代理记忆。我们将这一设置形式化为围绕一个记忆文件系统的三个角色:管理代理整合和组织传入内容,搜索代理通过引用来源回答查询,执行代理提供被提炼为技能的任务轨迹,将声明性记忆和技能统一在一个存储中。在长对话基准和具身任务中,我们变化了记忆形状(代理组织的层次结构、逐字转储、块检索)、流规模、工具组合(沙箱shell、记忆工具风格函数、多样化搜索工具)以及管理和搜索代理的强度,跟踪随着记忆增长的答案质量、成本和存储健康。可靠的组织所带来的收益是搜索经济:有序存储在材料较大的情况下大约将检索成本减半。然而,今天的代理未能实现默认的承诺:在我们的增长研究中,除了最强的管理代理外,组织性在所有情况下都在下降,而我们测量的没有任何代理将组织本身转化为更好的答案。而且,模型并不是唯一影响存储形状的杠杆:仅改变工具集就能像更换模型一样强烈地重塑存储。该研究将文件系统的默认状态从假设转变为代理记忆的设计空间。
cs.CL / 26 / 2607.26640

Contrastive ESA: Human Evaluation of Multiple Translations at Once

对比错误跨度标注:同时评估多种翻译的人工评估
Zouhar, Vilém, Grundkiewicz, Roman, Rajaee, Sara, Riley, Parker, Popel, Martin, Bawden, Rachel, Koehn, Philipp, Carpuat, Marine, Kocmi, Tom
Abstract
Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost. We introduce Contrastive Error Span Annotation (cESA), a protocol that presents multiple translations of the source input (text, video, audio, image). In cESA, the annotator sees multiple translations of the same document, marks major and minor error spans, and then assigns a score from 0% to 100% on absolute scale. By allowing annotators to access the shared context across multiple outputs, cESA facilitates more consistent and efficient judgments. We validate cESA using a large-scale human evaluation of English->Japanese translations of 12 models, demonstrating reductions in annotation time and noise compared to standard pointwise evaluation. Unlike existing contrastive ranking methods, cESA yields absolute quality judgments that enable simple, interpretable non-parametric model rankings without the need for post-hoc corrections.
Chinese Translation
当前机器翻译的人类评估通常孤立地评估单一输出,这种范式存在高标注者噪声和成本的问题。我们引入了对比错误跨度标注(Contrastive Error Span Annotation, cESA),这一协议展示了源输入(文本、视频、音频、图像)的多种翻译。在cESA中,标注者可以看到同一文档的多种翻译,标记主要和次要的错误跨度,然后在绝对尺度上给出0%到100%的评分。通过让标注者访问多个输出之间的共享上下文,cESA促进了更一致和高效的判断。我们通过对12个模型的英语到日语翻译进行大规模人类评估来验证cESA,结果显示与标准逐点评估相比,标注时间和噪声均有所减少。与现有的对比排名方法不同,cESA产生绝对质量判断,使得简单、可解释的非参数模型排名成为可能,而无需后期修正。
cs.CL / 27 / 2607.26654

Constitutional Midtraining: Content Presence Drives Alignment Gains

宪法中期训练:内容存在驱动对齐增益
Cho, Desiree, Tice, Cameron, Hogan, Bernie, Batra, Hunar, Radmard, Puria, Zhao, Jun, Shadbolt, Nigel
Abstract
Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce four constitutionally midtrained conditions plus a control, evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperform the control on alignment generalization and durability, notably on blackmail: SFT instills a blackmail propensity in all models, but constitutional midtraining blunts it, with the advantage surviving benign fine-tuning (-17.5pp). This durability does not extend to settings requiring active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also matters more than its structure, and constitutional midtraining incurs no cost, on average, on the capabilities we test (MMLU, ARC-Easy, piqa, GSM8K) at any stage. A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.
Chinese Translation
后训练对齐通常较为浅薄,在微调过程中会减弱。中期训练干预是否能够在与后训练清晰隔离的情况下产生持久的对齐尚未得到验证。我们通过宪法中期训练进行测试:在120B规模下,将基于原则的价值内容插入中期训练,与仅重播的对照组进行比较。我们的394M标记宪法语料库基于Anthropic的宪法,采用2x2因子设计(课程排序 x 深思熟虑推理),生成四种宪法中期训练条件以及一个对照组,评估自生成和既定基准,包括压力下的对齐、价值冲突解决、勒索以及在三个阶段(中期训练后、SFT后和良性微调后)的新出现的错位。宪法中期训练模型在对齐泛化和持久性方面优于对照组,尤其是在勒索任务上:SFT在所有模型中灌输了一种勒索倾向,但宪法中期训练减弱了这一倾向,其优势在良性微调后仍然存在(-17.5pp)。然而,这种持久性并未延伸至需要主动抵抗上下文压力或冲突的设置,在这些情况下,优势在SFT后减弱。中期训练中的宪法内容的存在比其结构更为重要,并且宪法中期训练在我们测试的能力(MMLU、ARC-Easy、piqa、GSM8K)上平均没有成本。中期训练中适量的宪法内容因此可能带来广泛而持久的对齐增益,成为以SFT为中心的流程的廉价补充。代码、数据和模型均可获得。
cs.CL / 28 / 2607.26700

Automated Multilabel Mpox Research Classification with Explainable Transformer Models

基于可解释变换器模型的自动化多标签Mpox研究分类
Aurpa, Tanjim Taharat
Abstract
The Mpox outbreak remains a serious public health issue, with the WHO (World Health Organization) reporting increasing cases in some regions. Research on Mpox is vital for several reasons, including vaccine development, diagnostic improvement, viral evolution studies, and preventing future outbreaks. However, the large amount of research being published makes it difficult to organize and analyze information efficiently. This study focuses on using multilabel classification to categorize 14590 Mpox research articles into key topics such as outbreaks, vaccination, and epidemiology. Among the different AI models tested, BERT performed the best, achieving 97.05% accuracy, 97.67% micro F1 score, and 96.46% macro F1 score. To better understand how the model makes decisions, SHAP was used to analyze significant word features and patterns. The results show that BERT can help automate the classification of Mpox research, making it easier for researchers, policymakers, and healthcare workers to quickly find relevant information, saving time and improving public health efforts.
Chinese Translation
Mpox疫情仍然是一个严重的公共卫生问题,世界卫生组织(WHO)报告称某些地区病例正在增加。对Mpox的研究至关重要,原因包括疫苗开发、诊断改进、病毒进化研究以及预防未来疫情。然而,发表的大量研究使得高效组织和分析信息变得困难。本研究集中于使用多标签分类将14590篇Mpox研究文章分类为关键主题,如疫情、疫苗接种和流行病学。在测试的不同人工智能模型中,BERT表现最佳,达到了97.05%的准确率、97.67%的微F1分数和96.46%的宏F1分数。为了更好地理解模型的决策过程,使用SHAP分析了重要的词汇特征和模式。结果表明,BERT可以帮助自动化Mpox研究的分类,使研究人员、政策制定者和医疗工作者能够更快地找到相关信息,从而节省时间并改善公共卫生工作。
cs.CL / 29 / 2607.26726

AtmosERC: Modeling Dialogue-Level Affective Atmosphere for Emotion Recognition in Conversation

AtmosERC:对话级情感氛围建模以实现对话中的情感识别
Feng, Weijie, Zhang, Tongwei, Liu, Binbin, Cheng, Zhiyong
Abstract
Emotion Recognition in Conversation (ERC) aims to predict utterance-level emotions in dialogues and has largely advanced through context-centric modeling. However, global context is a heterogeneous signal, and not all contextual information is equally relevant to emotion prediction. This paper focuses on the affect-oriented component of this signal, termed dialogue-level affective atmosphere, which captures a latent tendency commonly reflected in conversational emotion patterns. To estimate and exploit this tendency, we propose AtmosERC, a graph-based ERC framework that models each dialogue as a conversational graph over utterances and speakers. A relation-aware graph extractor filters and fuses heterogeneous graph signals to produce dialogue-level and speaker-conditioned affective priors. The resulting compact prior guides lightweight sequential emotion prediction and can also be verbalized into prompt-level cues for LLM-based ERC without modifying backbone models. Experiments on four ERC benchmarks show that AtmosERC improves lightweight ERC, enhances LLM-based ERC as a plug-in cue, and yields more stable predictions under local emotional deviations.
Chinese Translation
对话中的情感识别(Emotion Recognition in Conversation, ERC)旨在预测对话中发言的情感,并通过以上下文为中心的建模取得了显著进展。然而,全球上下文是一个异质信号,并非所有上下文信息对情感预测的相关性都是相同的。本文关注这一信号的情感导向组成部分,称为对话级情感氛围,它捕捉了在对话情感模式中常见的潜在倾向。为了估计和利用这种倾向,我们提出了AtmosERC,一个基于图的ERC框架,将每个对话建模为一个发言者和发言的对话图。关系感知图提取器过滤并融合异质图信号,以生成对话级和发言者条件的情感先验。所得到的紧凑先验指导轻量级的顺序情感预测,并且可以在不修改主干模型的情况下,转化为LLM(大语言模型)基础的ERC的提示级线索。在四个ERC基准上的实验表明,AtmosERC提高了轻量级ERC的性能,增强了作为插件线索的LLM基础ERC,并在局部情感偏差下产生了更稳定的预测。
cs.CL / 30 / 2607.26751

Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text

音素级与字符级目标及选择性状态空间模型在皮层内脑-文本转换中的应用
Vera, Lucas Zamora, Gonzalez-Lopez, Jose A.
Abstract
State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes remain underexplored: whether selective state-space models (Mamba) improve on recurrent decoders, and how the output target (phonetic vs.\ character) interacts with that choice. On the public Brain-to-Text '25 benchmark, we study a controlled 2x2 grid (GRU vs.\ hybrid Mamba decoder; phonetic vs.\ character targets) trained with a CTC objective under one reproducible protocol. The recurrent baseline remains strongest: the best phonetic GRU reaches 12.62\% PER and 21.19\% WER, while the best textual GRU after LM rescoring reaches 13.39\% CER and 26.28\% WER. The Mamba hybrid is competitive but does not surpass it. Ablations isolate architectural contributions, and error analysis shows representation-dependent failures: articulatory-like phoneme confusions vs.\ lexical and word-boundary errors.
Chinese Translation
最先进的皮层内脑-文本转换系统将神经序列音素解码器与外部语言模型相结合。两个设计轴尚未得到充分探索:选择性状态空间模型(Mamba)是否优于递归解码器,以及输出目标(音素与字符)如何与该选择相互作用。在公开的脑-文本转换'25基准测试中,我们研究了一个受控的2x2网格(GRU与混合Mamba解码器;音素与字符目标),并在一个可重复的协议下使用CTC目标进行训练。递归基线仍然是最强的:最佳音素GRU达到12.62%的音素错误率(PER)和21.19%的字错误率(WER),而最佳文本GRU在语言模型重新评分后达到13.39%的字符错误率(CER)和26.28%的字错误率(WER)。Mamba混合模型具有竞争力,但未能超越。消融实验隔离了架构贡献,错误分析显示出依赖于表示的失败:发音相似的音素混淆与词汇和词边界错误。
cs.CL / 31 / 2607.26760

Metis: Memory Foundation Model

Metis:记忆基础模型
Zhang, Zeyu, Guo, Ziliang, Sun, Yihang, Zhang, Xichong, Hao, Xixuan, Lin, Zehao, Zhang, Yang, Zhao, Xiaoyan, Shen, Tong, Tang, Bo, Xu, Zhi-Qin John, Yan, Junchi, Wang, Haofen, Chen, Xu, Xiong, Feiyu, Li, Zhiyu, Chua, Tat-Seng
Abstract
Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However, agent memory is still primarily implemented through external modules, leaving the native memory capability largely unexplored. In this paper, we take a first step toward this direction by introducing memory foundation models, which empower foundation models with native memory capabilities. We formalize native memory from two perspectives: a persistent and dynamically evolving memory state within the backbone, and native memory procedures that autonomously store and utilize information through model computation. We show that native memory offers advantages in architecture, end-to-end optimization, and efficiency. Based on this formulation, we propose Metis, the first prototype of memory foundation models. Metis introduces a new architecture that equips a foundation model with a native memory state, allowing historical information to be compressed into the model and accessed through memory attention. We construct large-scale memory-specific training data and introduce multiple optimization objectives to acquire these native memory procedures through mid-training. The online memory maintenance of Metis is gradient-free, and the memory update requires only a forward pass. At inference time, all learned model weights remain frozen, while the native memory states are autonomously transformed through standard forward computation. Through extensive experiments, we show that Metis exhibits native memory capabilities and further provide a detailed analysis of its strengths, limitations, and behaviors. To facilitate future research on memory foundation models, we release our project and model checkpoints.
Chinese Translation
近年来,人工智能代理的进步使其在基础模型中越来越多地内化了原生能力,从而催生了多模态基础模型和大型推理模型。然而,代理记忆仍主要通过外部模块实现,原生记忆能力尚未得到充分探索。本文朝这一方向迈出了第一步,提出了记忆基础模型,使基础模型具备原生记忆能力。我们从两个角度对原生记忆进行了形式化:在主干网络中持久且动态演变的记忆状态,以及通过模型计算自主存储和利用信息的原生记忆过程。我们展示了原生记忆在架构、端到端优化和效率方面的优势。基于这一形式化,我们提出了Metis,记忆基础模型的第一个原型。Metis引入了一种新架构,使基础模型具备原生记忆状态,允许历史信息被压缩到模型中,并通过记忆注意力进行访问。我们构建了大规模的记忆特定训练数据,并引入多个优化目标,通过中期训练获取这些原生记忆过程。Metis的在线记忆维护是无梯度的,记忆更新仅需一次前向传播。在推理时,所有学习到的模型权重保持不变,而原生记忆状态通过标准前向计算自主转化。通过大量实验,我们展示了Metis展现出原生记忆能力,并进一步提供了其优势、局限性和行为的详细分析。为了促进未来对记忆基础模型的研究,我们发布了我们的项目和模型检查点。
cs.CL / 32 / 2607.26762

Relation Geometry in Semantic Space of Language Models

语言模型语义空间中的关系几何
Cao, Zhihan, Yamada, Hiroaki, Teufel, Simone, Hiraoka, Tatsuya, Inui, Kentaro, Yanaka, Hitomi, Tokunaga, Takenobu
Abstract
When it comes to generating vector representations of words, current language models are achieving high-quality results. However, what is not known is the extent to which knowledge about semantic relations is represented in the geometry of the semantic spaces created in this way. In order to answer this question, we study the relation geometry of such semantic spaces from three perspectives. We first examine whether words standing in a particular relation to a target word~(called relata) occupy the same region in semantic space, and whether the regions corresponding to different relations are distinct from each other. We then verify to what extent semantic spaces reflect certain well-known properties of relations, such as symmetry, asymmetry, and transitivity. Finally, we consider which information about the target words and relata is more important for relation geometry: their surface forms, or their contexts. We conduct experiments on six semantic relations using causal, masked, and diffusion language models. The results show that relata in asymmetric relations relatively clearly occupy a distinct region in semantic space. Asymmetric relations' properties are only moderately well encoded in the semantic space, yet better than those of symmetric ones. Furthermore, when considering the question which information source has the strongest impact on results amongst the models we evaluated, we find that lexical information tends to be more important for the causal language model, whereas contextual information is more important for the masked and diffusion language models. Our results empirically show that relation geometry is not equally well-represented for all relations in semantic space, suggesting that there is a difference in how well semantic relations might be learned from distributional information alone.
Chinese Translation
在生成词的向量表示方面,目前的语言模型已取得高质量的结果。然而,尚不清楚在这种方式创建的语义空间的几何中,语义关系的知识被表示的程度。为了解答这个问题,我们从三个角度研究这种语义空间的关系几何。首先,我们考察与目标词(称为relata)存在特定关系的词是否占据语义空间中的同一区域,以及不同关系对应的区域是否彼此不同。接着,我们验证语义空间在多大程度上反映某些著名的关系特性,如对称性、非对称性和传递性。最后,我们考虑关于目标词和relata的信息中,哪种信息对关系几何更为重要:它们的表面形式,还是它们的上下文。我们在六种语义关系上使用因果、掩蔽和扩散语言模型进行了实验。结果表明,在非对称关系中,relata相对清晰地占据了语义空间中的一个独特区域。非对称关系的特性在语义空间中的编码程度仅为中等,但优于对称关系。此外,在我们评估的模型中,考虑哪种信息源对结果影响最大时,我们发现词汇信息对因果语言模型更为重要,而上下文信息对掩蔽和扩散语言模型更为重要。我们的结果实证表明,关系几何在语义空间中并非对所有关系都同样良好地表示,这表明仅通过分布信息学习语义关系的效果存在差异。
cs.CL / 33 / 2607.26780

Enhancing Generative Information Extraction with Two-step Validation: A Product Attribute Use Case

通过两步验证增强生成信息提取:产品属性应用案例
Hsu, Yi-Sheng, Baker, Nermeen Abou, Handmann, Uwe
Abstract
The ability of large language models (LLMs) to process and generate text has introduced potential for applications in information extraction (IE). While it's debated whether LLMs outperform smaller fine-tuned models for classification tasks, their strong generalization capability makes them promising for domains with limited labeled data available for fine-tuning. This advantage is particularly relevant for the emerging application of the digital product passport (DPP), where the problem space is broad but domain-specific data remains scarce. Motivated by this use case, we apply generative IE to the product domain, explicitly addressing efficiency, generalizability, and data privacy constraints. We propose a two-step validation method that integrates a PLM block into the generative IE pipeline and thereby leverages LLMs' correction capability. We discover that such a validation task enhances LLM performance, particularly on the extraction of weakly expressed, low-salience entities that appear sparsely throughout the text. For certain entities, the performance of mid-size models can even reach levels comparable to larger models, and the improvement of first-step PLM predictions also enhance the final LLM output. Nevertheless, the effects on the smallest open-source LLMs (e.g., Llama-3.2 3B) is limited. Based on the findings, we develop a demo application for product information extraction that utilizes locally deployed LLMs, targeting further adaptations to real-world DPP use cases.
Chinese Translation
大型语言模型(LLMs)处理和生成文本的能力为信息提取(IE)应用带来了潜力。尽管关于LLMs是否在分类任务中优于较小的微调模型仍存在争议,但它们强大的泛化能力使其在标注数据有限的领域中展现出良好的前景。这一优势在数字产品护照(DPP)的新兴应用中尤为重要,因为该问题空间广泛但领域特定数据仍然稀缺。基于这一应用案例,我们将生成信息提取应用于产品领域,明确解决效率、泛化性和数据隐私的限制。我们提出了一种两步验证方法,将预训练语言模型(PLM)模块集成到生成信息提取流程中,从而利用LLMs的纠错能力。我们发现,这种验证任务提升了LLM的性能,特别是在提取文本中稀疏出现的弱表达、低显著性实体方面。对于某些实体,中型模型的性能甚至可以达到与大型模型相当的水平,而第一步PLM预测的改进也提升了最终LLM的输出。然而,对于最小的开源LLMs(例如,Llama-3.2 3B),其效果有限。基于这些发现,我们开发了一个用于产品信息提取的演示应用,利用本地部署的LLMs,旨在进一步适应实际的DPP应用案例。
cs.CL / 34 / 2607.26795

When Does Span-Guided Detoxification Help? Human Preferences and Evaluator Diagnostics in a Controlled Comparison

何时跨度引导的去毒化有帮助?受控比较中的人类偏好与评估者诊断
Park, Kyungwon
Abstract
Span-guided rewriting aims to preserve meaning by localizing edits to annotated harmful spans, but the same constraint can leave harmful intent insufficiently mitigated. We present a controlled exploratory comparison of span-guided and unguided detoxification on a mixed-source English evaluation set comprising manually curated inputs and HateXplain test items. We conduct a dense blinded human evaluation under a fixed single-generator setting. Human preferences reveal a trade-off rather than a uniformly superior rewriting strategy. Span-guided outputs are favored when localized editing preserves the original stance and avoids unnecessary modification, whereas unguided outputs are favored when broader rewriting achieves more complete mitigation. This contrast varies substantially across the study-defined strata: the two strategies are competitive in the strong stratum, while unguided rewriting is clearly preferred in the mild stratum. Rationale annotations trace this difference to complementary failure risks: residual harm after localized editing and over-modification after broader rewriting. We treat automatic evaluation as a diagnostic rather than a substitute for human judgment. Toxicity-similarity scalarizations, a multi-generator analysis, and two general-purpose LLM judges reproduce parts of the aggregate tendency but do not yield an analogous stratified contrast. These setting-specific findings do not establish a severity-based routing rule. Instead, they motivate evaluation protocols that assess mitigation sufficiency and meaning preservation separately and report both residual harm and over-modification alongside aggregate scores.
Chinese Translation
跨度引导的重写旨在通过将编辑局限于标注的有害跨度来保留意义,但这种约束可能导致有害意图未能得到充分缓解。我们对跨度引导和无引导去毒化进行了受控的探索性比较,使用了一个混合来源的英语评估集,该评估集包含手动策划的输入和 HateXplain 测试项目。我们在固定的单生成器设置下进行了密集的盲人评估。人类偏好揭示了一个权衡关系,而不是一种统一的优越重写策略。当局部编辑保留原始立场并避免不必要的修改时,跨度引导的输出更受青睐;而当更广泛的重写实现更全面的缓解时,无引导的输出更受欢迎。这种对比在研究定义的层次中变化显著:在强层次中,两种策略具有竞争性,而在温和层次中,无引导重写明显更受偏好。推理注释将这种差异追溯到互补的失败风险:局部编辑后的残余伤害和更广泛重写后的过度修改。我们将自动评估视为诊断工具,而非人类判断的替代品。毒性相似性标量化、多生成器分析以及两个通用 LLM 评审者重现了部分整体趋势,但未能产生类似的分层对比。这些特定设置的发现并未建立基于严重性的路由规则。相反,它们激励评估协议,分别评估缓解充分性和意义保留,并报告残余伤害和过度修改,同时附上整体评分。
cs.CL / 35 / 2607.26825

From Found to Designed: Concepts as a Design Axis for Large Language Models

从发现到设计:概念作为大型语言模型的设计轴
Shani, Chen
Abstract
Large language models (LLMs) encode rich concept-like information, but represent it implicitly through distributed statistical associations rather than as explicit, structured, compositional concepts. Consequently, concept-level structure is typically \emph{found} rather than \emph{designed}: it is recovered after training through probing or dictionary learning, with no architectural guarantee of stability, compositionality, controllability, or alignment with human conceptual organization. We argue that concepts should instead be treated as a design axis for LLMs, and map the design space along two dimensions: the pipeline stage at which concept structure is introduced (training objective, core architecture, inference, or post-hoc interpretation), and whether that structure is internally derived from the model's own representations or grounded in external resources. This taxonomy reveals three broad patterns: inference-time approaches remain comparatively underexplored, related ideas have developed largely in isolation across pipeline stages, and externally grounded methods span the entire pipeline despite often being described under different terminology. Together, these observations motivate moving beyond recovering concept-like structure from trained models toward designing LLMs with explicit conceptual representations.
Chinese Translation
大型语言模型(LLMs)编码了丰富的类概念信息,但通过分布式统计关联隐式地表示,而不是作为明确的、结构化的、组合的概念。因此,概念级结构通常是 extit{发现}而非 extit{设计}:它是在训练后通过探测或字典学习恢复的,并没有架构上稳定性、组合性、可控性或与人类概念组织对齐的保证。我们认为,概念应被视为LLMs的设计轴,并在两个维度上映射设计空间:引入概念结构的管道阶段(训练目标、核心架构、推理或事后解释),以及该结构是从模型自身的表示内部推导而来还是基于外部资源。这一分类法揭示了三种广泛的模式:推理时的方法相对未被充分探索,相关思想在管道阶段之间大多孤立发展,而外部基础的方法跨越整个管道,尽管通常以不同的术语描述。综合来看,这些观察促使我们超越从训练模型中恢复类概念结构,朝着设计具有明确概念表示的LLMs迈进。
cs.CL / 36 / 2607.26831

Language Models are not Equally Robust to Non-Canonical Tokenization across Languages

语言模型对不同语言的非标准分词的鲁棒性并不相同
Ghosh, Poulami, Jyothi, Preethi
Abstract
Despite the existence of exponentially many valid tokenizations for a given string, language models operate on a single canonical sequence deterministically produced by the tokenizer, leaving the broader tokenization space largely uncharacterized. In this paper, we investigate this overlooked space by studying the behavior of language models under non-canonical tokenizations across diverse languages. For English, prior work shows that models are largely invariant to alternative tokenizations that represent the same underlying string. We ask whether this invariance generalizes to other languages beyond English. We conduct a multilingual study across 27 languages spanning diverse scripts and evaluate LLM behavior under alternative tokenizations across six downstream tasks. We find that tokenization invariance does not generalize: model behavior varies substantially across languages with instruction-tuned models exhibiting an average relative performance drop of 23.7% for Llama-3.1-8B, 11.4% for Qwen3-8B, and 9.9% for Gemma-3-12B. The variation of tokenization invariance is systematic across languages. Languages that exhibit higher token fragmentation show significantly greater sensitivity to non-canonical tokenizations. Our study of tokenization robustness serves as a diagnostic of how tightly a model is coupled to its tokenizer. These results demonstrate that tokenization robustness is not a universal property of language models, but depends strongly on the language and its interaction with the tokenizer. We also show that LoRA fine-tuning with multi-tokenization training data provides an effective mitigation for tokenization sensitivity. Fine-tuning on English alone improves tokenization robustness across languages, while systematically sampling diverse non-canonical tokenizations achieves the strongest overall performance.
Chinese Translation
尽管对于给定字符串存在指数级的有效分词方式,语言模型却仅在由分词器确定生成的单一标准序列上进行操作,从而使得更广泛的分词空间在很大程度上未被表征。本文通过研究语言模型在多种语言下的非标准分词行为,探讨了这一被忽视的空间。对于英语,先前的研究表明,模型在表示相同基础字符串的替代分词方式下基本保持不变。我们探讨这种不变性是否能够推广到英语以外的其他语言。我们在涵盖27种语言的多语言研究中进行实验,这些语言跨越了多种书写系统,并评估了在六个下游任务中,LLM在替代分词下的表现。我们发现分词不变性并不具有普遍性:模型行为在不同语言间存在显著差异,经过指令调优的模型在 Llama-3.1-8B 上的平均相对性能下降为23.7%,在 Qwen3-8B 上为11.4%,在 Gemma-3-12B 上为9.9%。分词不变性的变化在不同语言中呈现系统性。表现出更高分词碎片化的语言对非标准分词的敏感性显著更高。我们对分词鲁棒性的研究作为一种诊断,揭示了模型与其分词器之间的紧密耦合程度。这些结果表明,分词鲁棒性并不是语言模型的普遍特性,而是强烈依赖于语言及其与分词器的交互。我们还展示了使用多分词训练数据进行 LoRA 微调能够有效缓解分词敏感性。仅在英语上进行微调可以提高跨语言的分词鲁棒性,而系统性地采样多样的非标准分词则实现了最佳的整体性能。
cs.CL / 37 / 2607.26853

From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs

从表征到行为:探索大型语言模型中的人-情境-行为三元组
Zhang, Ruikang, Wang, Shuo, Su, Qi
Abstract
Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors. Existing studies of personality-related behavior in LLMs have primarily focused on outputs elicited under personality conditioning, characterizing observable trait-related expressions while lacking mechanistic evidence for the existence of internal personality-related representations, their cross-situational expression, and how these representations shape specific behaviors. Building on Funder's personality triad framework, we adapt its three components for LLM analysis: Person as personality-related internal representations, Situation as contexts that afford trait-relevant responses, and Behavior as response patterns on broader social tasks. We introduce a framework for discovering, controlling, and validating trait-like representations in LLMs. First, using contrastive behavior pairs grounded in shared situations, we identify sparse internal features associated with opposing poles of personality traits through SAE decomposition. We validate their trait relevance through effects on behavior to situation, token-level activation patterns, and robustness to paraphrasing. Second, feature-level interventions induce bidirectional trait-related shifts across a separate, diverse set of situations while preserving response validity, demonstrating consistent expression across contexts. Third, applying the same interventions to social intelligence tasks reveals behavioral changes with benefit-tradeoff patterns consistent with findings from human personality research, providing behavioral-level validation beyond personality scores. Our findings provide evidence that LLMs contain controllable trait-like representations linking internal states, situational expression, and behavioral outcomes.
Chinese Translation
人类个性理论将特质视为不是由单一分数捕捉的孤立属性,而是通过个体、情境和行为之间的相互作用表现出的稳定个体倾向。现有关于大型语言模型(LLMs)中与个性相关行为的研究主要集中于在个性条件下引发的输出,描述可观察的特质相关表达,但缺乏对内部个性相关表征存在的机制证据、其跨情境表达以及这些表征如何塑造特定行为的理解。基于Funder的个性三元组框架,我们将其三个组成部分适应于LLM分析:个体作为个性相关的内部表征,情境作为提供特质相关反应的背景,行为作为在更广泛社会任务中的反应模式。我们引入一个框架,用于发现、控制和验证LLMs中的特质类似表征。首先,通过基于共享情境的对比行为对,我们通过稀疏自编码(SAE)分解识别与个性特质对立极相关的稀疏内部特征。我们通过对情境的行为影响、令牌级激活模式以及对释义的稳健性验证其特质相关性。其次,特征级干预在一组独立的多样化情境中诱导双向特质相关的变化,同时保持反应的有效性,展示了跨情境的一致表达。第三,将相同的干预应用于社会智能任务时,揭示了行为变化,其效益权衡模式与人类个性研究的发现一致,提供了超越个性分数的行为层面验证。我们的研究结果提供了证据,表明LLMs包含可控的特质类似表征,连接内部状态、情境表达和行为结果。
cs.CL / 38 / 2607.26873

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning

SERPO:用于开放式测试时间强化学习的自我演化评分策略优化
Wang, Jianze, Zheng, Kunwang, Liu, Ying, Cao, Yu, Zhang, Qilong, Chen, Jinlong, Yang, Hua, Chen, Qianglong
Abstract
Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answer voting and therefore do not extend naturally to open-ended generation, where valid responses cannot be mapped to a shared canonical answer. Without external reward models or stronger judges, adaptation must instead construct reliable rewards from the model's own outputs. We introduce SERPO (Self-Evolving Rubric Policy Optimization), which replaces answer voting with a closed loop that co-evolves response evidence, query-specific rubrics, and policy parameters. Good-Normal-Bad (G-N-B) response evolution organizes maximally separated rollouts into ordered archives; rubric evolution retains criteria that discriminate these archives; probabilistic criterion scoring converts verdict-token likelihoods into reward signals; and policy evolution optimizes the actor with the resulting signals. New actor rollouts then refresh both the archives and rubrics, closing the three-way evolution loop. Across two model configurations, two in-domain benchmarks, and four OOD benchmarks, SERPO improves HealthBench and ResearchQA by up to 20.63 and 20.31 points over the corresponding base models, raises the six-benchmark macro-average by up to 8.06 points, and supports OOD transfer and continued cross-benchmark evolution.
Chinese Translation
测试时间强化学习(TTRL)使语言模型能够在推理时自我演化,而无需标记反馈。现有方法依赖于答案投票,因此不自然地扩展到开放式生成中,在这种情况下,有效的响应无法映射到共享的规范答案上。在没有外部奖励模型或更强评判者的情况下,适应必须从模型自身的输出中构建可靠的奖励。我们提出了SERPO(自我演化评分策略优化),它用一个闭环替代了答案投票,该闭环共同演化响应证据、查询特定的评分标准和策略参数。良好-正常-不良(G-N-B)响应演化将最大分离的回滚组织成有序档案;评分标准的演化保留了区分这些档案的标准;概率标准评分将裁决标记的可能性转换为奖励信号;而策略演化则优化了带有这些信号的执行者。新的执行者回滚随后刷新档案和评分标准,闭合三方演化循环。在两种模型配置、两个领域内基准和四个领域外基准上,SERPO在HealthBench和ResearchQA上分别提高了最多20.63和20.31分,六个基准的宏平均提高了最多8.06分,并支持领域外迁移和持续的跨基准演化。
cs.CL / 39 / 2607.26891

DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models

DIRECT:基于大型语言模型的高效且对齐的序列标注的直接解码
Wang, Yilei, Gan, Jiaxin, Zhang, Kexuan, Li, Ling, Zhang, Wentao, Lai, Peichao
Abstract
Sequence labeling is a fine-grained information extraction task, yet existing large language model-based approaches suffer from insufficient domain alignment and low inference efficiency. To address these issues, we propose DIRECT, a framework that addresses these issues through training-time optimization and inference-time rectification. Specifically, DIRECT performs Direct Preference Optimization (DPO) after supervised fine-tuning to strengthen task alignment with human preferences, and introduces a controlled decoding process that enforces fixed output formats and restricts predictions to candidate sets. To further improve efficiency, a template-filling mechanism requires the model to generate only label tokens while reusing prefixed content through the KV Cache, thus reducing redundant computation. Experimental results on eight datasets demonstrate that DIRECT achieves significant improvements in both performance and efficiency compared to existing methods.
Chinese Translation
序列标注是一项细粒度的信息提取任务,但现有基于大型语言模型的方法在领域对齐不足和推理效率低下方面存在问题。为了解决这些问题,我们提出了DIRECT,一个通过训练时优化和推理时修正来应对这些挑战的框架。具体而言,DIRECT在监督微调后执行直接偏好优化(Direct Preference Optimization, DPO),以增强任务与人类偏好的对齐,并引入一种控制解码过程,强制固定输出格式并将预测限制在候选集内。为了进一步提高效率,模板填充机制要求模型仅生成标签令牌,同时通过KV缓存重用前缀内容,从而减少冗余计算。在八个数据集上的实验结果表明,与现有方法相比,DIRECT在性能和效率上均取得了显著提升。
cs.CL / 40 / 2607.26909

Dual-Path LLM Reasoning for Multimodal Few-Shot Knowledge Graph Completion

双路径大型语言模型推理用于多模态少样本知识图谱补全
Liu, Jinlan, Tu, Zhiying, Xing, Yongchao, Liu, Yicheng, Zhang, Bolin, Sui, Dianbo, Chu, Dianhui, Sun, Hongliang
Abstract
Knowledge graph completion (KGC) aims to infer missing facts in knowledge graphs (KGs), thereby improving their completeness and supporting downstream intelligent applications. However, emerging entities and relations in real-world deployments make inductive KGC difficult, especially under few-shot and zero-shot settings. Multimodal information and Large Language Model (LLM)-derived priors can enrich sparse relational contexts, but they may also introduce noisy or hallucinated evidence. To address these issues, we propose DuPLeR, a \textbf{Du}al-\textbf{P}ath \textbf{L}LM \textbf{R}easoning framework for multimodal few-shot KGC. DuPLeR builds a calibrated relation graph by combining multimodal LLM-derived type priors with factual support structures, and performs dual-level structural reasoning over the refined relation topology. Moreover, a dual-pathway multimodal enhancement module regulates message passing with query-relevant multimodal signals and supplements entity representations after graph propagation. Experiments on eight inductive variants of two multimodal KG (MMKG) benchmarks show that DuPLeR achieves robust performance in data-scarce KGC scenarios.
Chinese Translation
知识图谱补全(KGC)旨在推断知识图谱(KGs)中缺失的事实,从而提高其完整性并支持下游智能应用。然而,现实世界中的新兴实体和关系使得归纳式KGC变得困难,尤其是在少样本和零样本设置下。多模态信息和大型语言模型(LLM)衍生的先验知识可以丰富稀疏的关系上下文,但也可能引入噪声或虚构的证据。为了解决这些问题,我们提出了DuPLeR,一个用于多模态少样本KGC的双路径LLM推理框架。DuPLeR通过将多模态LLM衍生的类型先验与事实支持结构相结合,构建了一个经过校准的关系图,并在精炼的关系拓扑上执行双层结构推理。此外,双路径多模态增强模块通过与查询相关的多模态信号调节消息传递,并在图传播后补充实体表示。在两个多模态知识图谱(MMKG)基准的八个归纳变体上的实验表明,DuPLeR在数据稀缺的KGC场景中实现了稳健的性能。
cs.CL / 41 / 2607.26928

Latent-IM: Latent Interaction Management for Speech LLMs

潜在交互管理:针对语音大语言模型的潜在交互管理
Avsian, Adar, Dokme, Atahan, Woo, Tony, Heck, Larry
Abstract
Classical spoken dialogue systems often separated dialogue management from response realization: a policy selected the next dialogue action, and a generation component expressed that action. As dialogue systems shift toward LLMs, this decomposition has largely disappeared into the model's hidden representations. We ask whether an LLM-internal analogue of state estimation and action control can be recovered for conversational moves such as acknowledging, checking, querying, explaining, and replying. We formulate move control as two coupled problems: selection, predicting the appropriate next move from the dialogue context, and realization, causally producing a chosen move at generation time. We introduce Latent-IM, an internal dialogue-management framework that provides a general interface for choosing and deploying conversational moves under different objectives. Here, we use this control to reproduce human move choices, improving average end-to-end move accuracy by 12.5 points over the unsteered backbone while performing comparably to fine-tuning.
Chinese Translation
经典的语音对话系统通常将对话管理与响应实现分开:一个策略选择下一个对话动作,而一个生成组件则表达该动作。随着对话系统向大语言模型(LLMs)转变,这种分解在模型的隐藏表示中大部分消失。我们探讨是否可以恢复LLM内部状态估计和动作控制的类比,以应对诸如确认、检查、询问、解释和回复等对话动作。我们将动作控制形式化为两个耦合问题:选择,从对话上下文中预测适当的下一个动作;实现,在生成时因果地产生所选动作。我们引入了潜在交互管理(Latent-IM),这是一个内部对话管理框架,提供了在不同目标下选择和部署对话动作的一般接口。在这里,我们利用这种控制来再现人类的动作选择,使得平均端到端动作准确率比未引导的基础模型提高了12.5个百分点,同时在性能上与微调相当。
cs.CL / 42 / 2607.26929

Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States

相同证据,不同目标:解码诊断证据如何影响语言模型状态下的因果问题
Kong, Weiyi, Li, Zhuoran
Abstract
The same diagnostic result can support or challenge one causal claim yet fail to address another when the claims concern different populations, outcomes, estimands, pathways, or identifying assumptions. When the evidence and target vary together, a correct answer may reflect favorable or adverse wording, lexical overlap, or a familiar diagnostic pattern rather than matching the evidence to the causal question. We introduce paired prompts that repeat the same diagnostic evidence verbatim while changing the causal target. Each prompt is labeled Favors, Challenges, Unresolved, or Wrong Target according to how the evidence bears on the causal question. A pair is recovered only when both prompts are classified correctly. Using linear readouts trained on a separate development set, we analyze the final-token hidden state from the penultimate transformer block of Qwen2.5-7B-Instruct, Qwen3-8B, and Llama-3.1-8B-Instruct. On the 49-pair primary benchmark spanning nine diagnostic families, balanced accuracy ranges from 0.654 to 0.659 and 18-21 pairs are recovered. Two independent human reviewers assigned the same label to 95 of the 98 prompts (96.9%). Across checkpoints, balanced accuracy and complete-pair recovery exceed permutation nulls that preserve development scenario groups. In Qwen2.5, full-prompt balanced accuracy exceeds both restricted inputs, with paired-bootstrap intervals for both differences above zero. Readouts trained without development examples from the evaluated diagnostic family recover 21 pairs, including at least one in each of the nine families. The hidden-state readout exceeds a linear classifier on answer-option logits and text baselines in balanced accuracy and recovered pairs. These results show that the hidden state contains linearly decodable information about whether diagnostic evidence favors, challenges, or fails to address the causal target.
Chinese Translation
相同的诊断结果可以支持或挑战一个因果主张,但在主张涉及不同的人群、结果、估计量、路径或识别假设时,可能无法解决另一个问题。当证据和目标共同变化时,正确的答案可能反映出有利或不利的措辞、词汇重叠或熟悉的诊断模式,而不是将证据与因果问题相匹配。我们引入了配对提示,这些提示逐字重复相同的诊断证据,同时改变因果目标。每个提示根据证据对因果问题的影响被标记为“有利”、“挑战”、“未解决”或“错误目标”。只有当两个提示都被正确分类时,才能恢复一对。我们使用在单独开发集上训练的线性读出,分析 Qwen2.5-7B-Instruct、Qwen3-8B 和 Llama-3.1-8B-Instruct 的倒数第二个变换块的最终令牌隐藏状态。在涵盖九个诊断家族的49对主要基准上,平衡准确率范围为0.654到0.659,恢复了18-21对。两位独立的人类评审员对98个提示中的95个(96.9%)分配了相同的标签。在各个检查点上,平衡准确率和完整对恢复超过了保留开发场景组的置换零假设。在 Qwen2.5 中,完整提示的平衡准确率超过了受限输入,且两者的差异的配对自助区间均高于零。未使用评估的诊断家族的开发示例训练的读出恢复了21对,包括每个九个家族中的至少一对。隐藏状态读出在答案选项对数和文本基线的平衡准确率和恢复对数上超过了线性分类器。这些结果表明,隐藏状态包含关于诊断证据是否有利、挑战或未能解决因果目标的线性可解信息。
cs.CL / 43 / 2607.26952

Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?

信用卡、混淆、计算与后果:我们能揭示语言模型推理的哪些内容?
Hiray, Arnav, Shah, Agam, Lu, Caleb, Tarte, Meghaj, Mittal, Harsit, Chava, Sudheer
Abstract
We introduce CreditCardQA, the first financial literacy benchmark for numerical reasoning derived from real credit card agreements. The dataset contains 1,800 questions, including first-person variants that reflect how consumers naturally ask about fees, interest, and payments. We evaluate a range of large language and reasoning models under Chain-of-Thought (CoT) and Program-of-Thought (PoT) prompting. Overall, PoT yields consistent performance gains, particularly for models with weaker baseline reasoning, and narrows gaps between open- and closed-source systems. Through error analysis, we show that failures arise less from arithmetic and more from misapplied financial rules, missed conditions, and misunderstandings of contractual terms. We further analyze question difficulty and find that comparisons, conditional logic, and monetary constraints are especially challenging. We also find that errors often arise in edge cases such as late-payment penalties or small-balance scenarios that are more likely to affect lower-income or financially vulnerable individuals.
Chinese Translation
我们介绍了 CreditCardQA,这是第一个基于真实信用卡协议的数值推理金融素养基准数据集。该数据集包含1800个问题,包括反映消费者自然询问费用、利息和付款的第一人称变体。我们在 Chain-of-Thought (CoT) 和 Program-of-Thought (PoT) 提示下评估了一系列大型语言和推理模型。总体而言,PoT 在性能上带来了持续的提升,特别是对于基线推理较弱的模型,并缩小了开放源代码和闭源系统之间的差距。通过错误分析,我们显示出失败的原因更多地源于错误应用的金融规则、遗漏的条件和对合同条款的误解,而非算术错误。我们进一步分析了问题的难度,发现比较、条件逻辑和货币限制尤其具有挑战性。我们还发现,错误往往出现在诸如逾期付款罚款或小额余额等边缘案例中,这些情况更可能影响低收入或财务脆弱的个体。
cs.CL / 44 / 2607.26967

Generation or Judgement? A Paradigm Perspective on LLM-Based Emotion-Cause Pair Extraction in Conversation

生成还是判断?基于大型语言模型的对话情感-原因对提取的范式视角
Feng, Weijie, Wang, Hongchuang, Liu, Binbin, Cheng, Zhiyong
Abstract
Emotion-cause pair extraction in conversation (ECPEC) identifies utterance pairs in which one utterance causes an emotion expressed in another. Recent LLM-based approaches formulate ECPEC at markedly different granularities, ranging from generating complete pair sets to judging individual candidate pairs. In this paper, we make the surprising observation that task formulation substantially affects performance, where pair-level judgement outperforms dialogue-level generation in all 18 controlled comparisons. We investigate the sources of this paradigm gap and find that many relations omitted by dialogue-level generation remain recognizable under explicit pair queries, under which the model recognizes 92.7%-98.1% of emotion-cause relations. This suggests that LLMs can recognize emotion-cause relations but struggle to discover and return complete pair sets. Pair-level judgement alleviates this burden, although its candidate rankings are more reliable than the binary decisions produced by a shared threshold. Based on this diagnosis, we introduce an auxiliary retriever that selectively re-examines ambiguous boundary cases, yielding consistent F1 improvements of 0.50-1.46 points across three datasets while maintaining an inference time of only 1.49x that of the baseline paradigm. These findings show that task decomposition and candidate scope are critical to effectively utilizing LLMs for ECPEC.
Chinese Translation
对话中的情感-原因对提取(ECPEC)识别出一对话语,其中一个话语引发了另一个话语中表达的情感。最近的基于大型语言模型(LLM)的方法在ECPEC的表述上呈现出显著不同的粒度,从生成完整的对对到判断单个候选对。在本文中,我们惊讶地观察到任务表述对性能有显著影响,其中对对级别的判断在所有18个控制比较中均优于对话级别的生成。我们探讨了这一范式差距的来源,发现许多被对话级生成省略的关系在明确的对查询下仍然可识别,在这种情况下,模型能够识别92.7%-98.1%的情感-原因关系。这表明LLM能够识别情感-原因关系,但在发现和返回完整的对对集方面存在困难。对对级判断减轻了这一负担,尽管其候选排名比通过共享阈值产生的二元决策更为可靠。基于这一诊断,我们引入了一种辅助检索器,选择性地重新审视模糊的边界案例,在三个数据集上实现了一致的F1分数提升,提升幅度为0.50-1.46点,同时保持仅为基准范式1.49倍的推理时间。这些发现表明,任务分解和候选范围对于有效利用LLM进行ECPEC至关重要。
cs.CL / 45 / 2607.26977

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning

TREK:复杂旅行规划中大语言模型代理的旅行推理与评估工具包
Qi, Jinhu, Zhang, Wentao, Ng, Siu Man, Xu, Feiyang, Chen, Yanyu, Li, Yaoman, King, Irwin
Abstract
Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once - every flight, hotel, and attraction must exist and be bookable, the days must be physically traversable, the total must clear a budget, and the plan must serve a traveler whose needs are only partly stated. Existing agent benchmarks reward these properties one at a time and grade the final output with soft or LLM-judged rubrics, which cannot certify that a returned plan is executable and are neither reproducible nor auditable. We introduce TREK (Travel Reasoning and Evaluation Kit), a benchmark for feasible itinerary synthesis: producing a single plan that is jointly constraint-correct, hallucination-free, spatio-temporally executable, budget-valid, and responsive to the traveler's unstated persona needs. TREK comprises 800 multi-constraint tasks - 533 feasible and 267 provably infeasible with typed route/entity/budget causes - over a synthetic, internally consistent knowledge base of 212,530 records across 375 cities and 13 personas, served through a production-style tool sandbox of validated RESTful APIs. Every task is scored by a fully deterministic, rule-based evaluator with no LLM judge and ships a human-verified gold reference that scores a perfect 1.0 under that same evaluator, so the ceiling is demonstrably achievable and every remaining gap is an agent limitation rather than scorer strictness. Evaluating 15 LLM agents across nine constraint dimensions, we find that even the strongest (GPT-5.6) produces a fully-feasible plan on only 46.2% of solvable tasks, with a median of 6.6% and a floor of 0.0%; satisfying travelers' unstated needs emerges as the universal bottleneck, unsolved even at the frontier. We release the dataset, tool sandbox, deterministic evaluator, and agent code as a fully reproducible benchmark.
Chinese Translation
旅行规划是对使用工具的大语言模型(LLM)代理的严峻压力测试:一个可用的行程是一个必须在多个维度上同时正确的单一产物——每个航班、酒店和景点必须存在且可预订,行程的天数必须在物理上可行,总费用必须在预算之内,并且计划必须满足旅行者未完全表达的需求。现有的代理基准一次只奖励这些属性,并使用软性或LLM评判的评分标准对最终输出进行评分,这无法证明返回的计划是可执行的,也既不可重复也不可审计。我们引入了TREK(旅行推理与评估工具包),这是一个可行行程合成的基准:生成一个在约束条件上共同正确、无幻觉、时空可执行、预算有效且响应旅行者未表达的个性需求的单一计划。TREK包含800个多约束任务——533个可行和267个可证明不可行,具有类型化的路线/实体/预算原因——基于一个合成的、内部一致的知识库,涵盖212,530条记录,涉及375个城市和13个角色,通过一个生产风格的工具沙箱提供经过验证的RESTful API。每个任务由一个完全确定性的基于规则的评估器评分,没有LLM评判,并提供一个人类验证的金标准,在同一评估器下得分为完美的1.0,因此上限是显然可实现的,剩余的每个差距都是代理的局限而非评分的严格性。在九个约束维度上评估15个LLM代理,我们发现即使是最强的(GPT-5.6)在仅46.2%的可解任务上生成完全可行的计划,中位数为6.6%,最低为0.0%;满足旅行者未表达的需求成为普遍瓶颈,甚至在前沿也未得到解决。我们发布数据集、工具沙箱、确定性评估器和代理代码,作为一个完全可重复的基准。
cs.CL / 46 / 2607.26981

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

OptimismBench:语言模型判断中的预测偏差与对齐效应
Cho, Seonglae, Koshiyama, Adriano
Abstract
Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a systematic directional tilt has been hard to detect: calibration metrics aggregate unsigned errors, and naturalistic uncertainty offers no ground-truth probability. When an LLM rates a startup's success at 70% but its failure at 15%, the missing 15 points expose a distortion no aggregate score flags. We introduce OptimismBench, which detects directional bias with inverted pairs: each scenario elicits both P(success) and P(failure), and asymmetry between the two framings yields a signed bias score without ground truth. Across 16 models from 8 providers, fourteen are optimistic; pessimism appears only in Anthropic's frontier tier. Eleven matched base-versus-chat pairs across four families show post-training sets the sign of the bias, with opposite shifts in different families. The pattern survives prompt, temperature, perspective, and self-debiasing ablations. A seventeen-model six-language comparison further shows model identity dominates language, with inter-model variance at 4.7x inter-language variance. We release 3,870 items across 10 languages for per-model directional-bias auditing. When alignment makes a model more helpful, it also tilts its probabilities; downstream pipelines inherit the tilt by default.
Chinese Translation
大型语言模型越来越多地被用作决策辅助工具,其概率判断影响着下游选择。然而,这些判断是否存在系统性的方向性倾斜却难以探测:校准指标汇总了未签名的错误,而自然的不确定性则没有真实的概率作为依据。当一个大型语言模型(LLM)将某初创公司的成功率评估为70%,而失败率评估为15%时,缺失的15个百分点揭示了一种未被汇总评分标识的扭曲。我们引入了OptimismBench,它通过反转对(inverted pairs)来检测方向性偏差:每个场景同时引发成功概率(P(success))和失败概率(P(failure)),两种框架之间的非对称性产生了一个有符号的偏差评分,而无需真实的依据。在来自8个提供者的16个模型中,十四个模型表现出乐观;悲观仅出现在Anthropic的前沿层级。在四个家族中,十一对匹配的基础与聊天模型显示,后训练集设置了偏差的符号,不同家族之间则表现出相反的变化。这一模式在提示、温度、视角和自我去偏差的消融实验中依然存在。对十七个模型的六种语言比较进一步表明,模型身份主导了语言,模型间的方差是语言间方差的4.7倍。我们发布了3870个项目,涵盖10种语言,以便对每个模型进行方向性偏差审计。当对齐使模型更有帮助时,它也倾斜了其概率;下游流程默认继承这种倾斜。
cs.CL / 47 / 2607.27022

Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making

从抽象刻板印象到具体社会决策:评估大型语言模型中的区域偏见
Di, Jiayuan, Yang, Haoyi, Luo, Yufei, Qu, Jiahui, Wang, Yiming
Abstract
Regional bias in large language models (LLMs) may shape both perceptions of regional groups and decisions about individuals from different regions. Yet existing studies often examine these manifestations separately, leaving their structure and consequences unclear. We introduce Stereotypes-to-Decisions (S2D), a systematic framework evaluating regional bias from abstract stereotypes to concrete social decisions. Covering all 34 provincial-level administrative regions of China, S2D evaluates six LLMs using stereotype ratings of Warmth (perceived friendliness and trustworthiness) and Competence (perceived capability and intelligence), along with paired-choice tasks across Education, Occupation, and Social Interaction. Results reveal substantial regional differences in regional scores, with considerable agreement across models, especially for Competence and Occupation decisions. Furthermore, these patterns are associated with regional economic and digital development indicators and display mixed human-like stereotypes, with some regions rated highly on one dimension but poorly on the other. They also remain largely stable across Chinese and English prompts. Overall, our findings show that regional bias in LLMs is prevalent, systematic, and consequential, motivating more regionally aware evaluation and mitigation.
Chinese Translation
大型语言模型(LLMs)中的区域偏见可能影响对区域群体的看法以及对来自不同地区个体的决策。然而,现有研究通常将这些表现分开考察,导致其结构和后果不明。我们提出了刻板印象到决策(Stereotypes-to-Decisions, S2D)这一系统框架,以评估从抽象刻板印象到具体社会决策的区域偏见。S2D覆盖了中国所有34个省级行政区域,使用对温暖度(感知的友好度和可信度)和能力(感知的能力和智力)的刻板印象评分,结合教育、职业和社会互动的配对选择任务,对六个LLMs进行了评估。结果显示,区域评分存在显著的区域差异,各模型之间在能力和职业决策方面的共识尤为显著。此外,这些模式与区域经济和数字发展指标相关联,展现出混合的人类刻板印象,某些地区在一个维度上评分较高而在另一个维度上评分较低。它们在中文和英文提示下也保持了相对稳定。总体而言,我们的研究结果表明,LLMs中的区域偏见普遍存在、系统性强且具有重要后果,促使我们进行更具区域意识的评估和缓解。
cs.CL / 48 / 2607.27178

DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

DenseOn与LateOn:用于多语言、长上下文和代码搜索的完全开放密集和晚期交互模型
Sourty, Raphaël, Chaffin, Antoine, Junior, Paulo Roberto Moura, Chatelain, Amélie
Abstract
State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. We first reconstruct and curate 665M English contrastive pre-training pairs from 1.4B pairs across 34 public sources and build 1.88M supervised fine-tuning pairs with mined hard negatives. Training yields two 149M-parameter models: DenseOn, a single-vector dense model, and LateOn, a ColBERT-style late-interaction model. They achieve 56.20 and 57.22 average nDCG@10 on BEIR, respectively, setting new state-of-the-art results for this size class. We then translate the validated English data into eight languages, yielding 2.8B pairs with cross-lingual samples, and train mDenseOn and mLateOn, two 307M-parameter models built on mmBERT-base. Despite sharing their backbone, data, and objectives, their representations behave differently: the dense model is strong on English and translated languages but degrades outside translate-train support, whereas the late-interaction model generalizes better to unseen languages and scripts. This suggests that token-level matching turns translate-train from a target-language expansion strategy into a multilingual generalization recipe. We publicly release the models, datasets, and training code.
Chinese Translation
最先进的检索模型越来越依赖于封闭的训练数据,这造成了可重复性差距。我们提出了一种开放的端到端训练检索模型的方案,并研究了英语监督如何通过翻译训练转移到多语言检索中。我们首先从34个公共来源的14亿对中重建和整理了6.65亿个英语对比预训练对,并构建了188万对带有挖掘的困难负样本的监督微调对。训练产生了两个参数为1.49亿的模型:DenseOn,一个单向量密集模型,以及LateOn,一个ColBERT风格的晚期交互模型。它们在BEIR上分别达到了56.20和57.22的平均nDCG@10,创造了这一规模类别的新最先进结果。随后,我们将经过验证的英语数据翻译成八种语言,生成了28亿对带有跨语言样本的数据,并训练了mDenseOn和mLateOn,这两个参数为3.07亿的模型基于mmBERT-base。尽管共享了它们的骨干网络、数据和目标,但它们的表示行为却不同:密集模型在英语和翻译语言上表现强劲,但在翻译训练支持之外会退化,而晚期交互模型在未见过的语言和脚本上具有更好的泛化能力。这表明,基于token的匹配将翻译训练从一种目标语言扩展策略转变为一种多语言泛化方案。我们公开发布了这些模型、数据集和训练代码。
cs.CL / 49 / 2607.27183

Pangram 4 Technical Report

Pangram 4 技术报告
Glickenhaus, Ben, Thai, Katherine, Russell, Jenna, Masrour, Elyas, Han, Yue, Spero, Max, Emi, Bradley
Abstract
We present Pangram 4, the latest deep-learning-based AI-text classification model from Pangram Labs. We achieve an AUROC of 0.9916 with a false positive rate of 0.0041% and a false negative rate of 0.3396%. In addition to its increased overall accuracy compared with Pangram 3, Pangram 4 exhibits superior out-of-distribution generalization and robustness to adversarial attacks. Another novel contribution of Pangram 4 is its improved ability to distinguish fine-grained edits and mixed AI-human co-authored text. We demonstrate improvements to both boundary detection tasks and the detection of interleaved AI assistance. Finally, we report metrics on standard AI detection benchmarks showing that Pangram 4 achieves state-of-the-art performance on the AI text detection task across a wide variety of settings and domains.
Chinese Translation
我们介绍了 Pangram 4,这是 Pangram Labs 最新的基于深度学习的 AI 文本分类模型。我们在假阳性率为 0.0041% 和假阴性率为 0.3396% 的情况下,达到了 0.9916 的 AUROC。与 Pangram 3 相比,Pangram 4 不仅提高了整体准确性,还展现了优越的分布外泛化能力和对对抗攻击的鲁棒性。Pangram 4 的另一个新颖贡献是其在区分细粒度编辑和混合 AI-人类共同创作文本方面的能力得到了改善。我们展示了对边界检测任务和交错 AI 辅助检测的改进。最后,我们报告了在标准 AI 检测基准上的指标,显示 Pangram 4 在各种设置和领域的 AI 文本检测任务中达到了最先进的性能。
cs.CL / 50 / 2607.27189

APEX-Accounting

APEX-会计
Benchek, Julien, Bennett, Austin, Kern, Jasmin, Stevens, Ryan, Sultan, Rene, Ching, Charis, Popiel, Hayley, Mittal, Vaibhav, Mercier, Felix, Foody, Brendan, Vidgen, Bertie
Abstract
We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 (GPT-5.6-Sol (Max+Pro)) and the highest Pass@8 is 21.5% (Muse-Spark-1.1 (xHigh)). We experiment with increasing the token budget from $1 to $50 and observe an instance of Simpson's paradox: scores increase as the token budget increases but within a given budget-constrained harness, scores are lower on tasks where the model spends more tokens. As APEX-Accounting is a closed benchmark, leaderboard evals can be run for any frontier model on request.
Chinese Translation
我们介绍了APEX-会计,这是由Mercor与Ramp合作建立的基准,用于评估前沿模型是否能够完成会计师的实际工作。任务包括对账、计提费用、记录交易和生成报告。私有评估集包含160个任务,分布在10个领域中。每个领域包含一个会计系统,以及电子表格、PDF和其他文件。每个任务均由会计和簿记领域的专家撰写并解决,他们还编写了评分标准。在九个前沿模型中,Claude-Fable-5 (Max)以56.4%的平均标准@3领先,Muse-Spark-1.1 (xHigh)以52.6%紧随其后。没有模型的通过率超过2.6% Pass^8 (GPT-5.6-Sol (Max+Pro)),最高的Pass@8为21.5% (Muse-Spark-1.1 (xHigh))。我们实验性地将令牌预算从1美元增加到50美元,并观察到一个辛普森悖论的实例:随着令牌预算的增加,得分也在上升,但在给定预算限制的情况下,模型在花费更多令牌的任务上的得分反而较低。由于APEX-会计是一个封闭的基准,任何前沿模型的排行榜评估都可以根据请求进行。
cs.CL / 51 / 2607.27201

Mental World Modeling

心理世界建模
Fei, Hao, Zhao, Yiran
Abstract
World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends, feels, and considers socially permissible), so a model that tracks the physical scene but not what each agent knows and believes about it predicts the wrong action for the right-looking scene. We formulate Mental World Modeling (MWM), a generic theoretical framework that makes mental variables core components of a world model rather than posthoc rationales: MWM aintains a coupled physical-mental world state, renders a target-specific partial observation, and simulates how candidate actions jointly update both components. We instantiate the framework in MENTIS, a training-free and fully inspectable baseline that decomposes the process into state parsing, target-observation generation, action decomposition, coupled physical and mental transition, and branch-level value evaluation. On a manually constructed, quality-controlled dataset of situated decision scenarios spanning text, image, and sounding-video stories, experiments with 8 modern LLM-based world models demonstrate that explicitly modeling the mental state is essential for predicting human decisions. Deeper analyses further expose the bottlenecks of current mental world modeling. We expect MWM as a next stage of world modeling, from simulating physical scenes to simulating the minds that act in them.
Chinese Translation
世界模型为规划和行动提供了预测基础,然而现有的模型仅仅回答了一个物理问题:它是什么/在哪里,以及它将如何演变。然而,人类行为是由隐藏的心理状态驱动的(一个人相信什么、想要什么、打算什么、感受什么以及认为社会上可接受的是什么),因此,一个仅跟踪物理场景而不考虑每个主体对其的知识和信念的模型,会对看似正确的场景预测出错误的行动。我们提出了心理世界建模(Mental World Modeling, MWM),这是一个通用的理论框架,将心理变量作为世界模型的核心组成部分,而不是事后解释:MWM 维持一个耦合的物理-心理世界状态,呈现目标特定的部分观测,并模拟候选行动如何共同更新这两个组成部分。我们在 MENTIS 中实例化了该框架,这是一种无训练且完全可检验的基线,将过程分解为状态解析、目标观测生成、行动分解、耦合的物理与心理转变,以及分支级别的价值评估。在一个手动构建、质量控制的情境决策场景数据集中,涵盖文本、图像和声音视频故事,使用 8 个基于现代大型语言模型的世界模型的实验表明,明确建模心理状态对于预测人类决策至关重要。更深入的分析进一步揭示了当前心理世界建模的瓶颈。我们期待 MWM 作为世界建模的下一个阶段,从模拟物理场景到模拟在其中行动的心智。