Smooth Operator: Smooth Verifiable Reward Activates Spatial Reasoning Ability of Vision-Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiao, Siwen, Lv, Tianxiong, Qian, Kangan, Zhao, Chenxu, Zhu, Xiuyuan, Li, Tianlun, Cheng, Xiaolong, Li, Jinyu, Liao, Zhihao, Cai, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
von: Zhang, Zijing, et al.
Veröffentlicht: (2025)
von: Zhang, Zijing, et al.
Veröffentlicht: (2025)
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
von: He, Haoran, et al.
Veröffentlicht: (2025)
von: He, Haoran, et al.
Veröffentlicht: (2025)
DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
von: Lee, Vint, et al.
Veröffentlicht: (2023)
von: Lee, Vint, et al.
Veröffentlicht: (2023)
Protocols for Verifying Smooth Strategies in Bandits and Games
von: Christ, Miranda, et al.
Veröffentlicht: (2025)
von: Christ, Miranda, et al.
Veröffentlicht: (2025)
A Novel Hierarchy of Quantum Kernel Networks on Smoothed Particle Hydrodynamics
von: Li, Yudong, et al.
Veröffentlicht: (2026)
von: Li, Yudong, et al.
Veröffentlicht: (2026)
DiTAS: Quantizing Diffusion Transformers via Enhanced Activation Smoothing
von: Dong, Zhenyuan, et al.
Veröffentlicht: (2024)
von: Dong, Zhenyuan, et al.
Veröffentlicht: (2024)
Operator Learning for Smoothing and Forecasting
von: Calvello, Edoardo, et al.
Veröffentlicht: (2026)
von: Calvello, Edoardo, et al.
Veröffentlicht: (2026)
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
Artistic Neural Style Transfer Algorithms with Activation Smoothing
von: Li, Xiangtian, et al.
Veröffentlicht: (2024)
von: Li, Xiangtian, et al.
Veröffentlicht: (2024)
Randomized Smoothing Meets Vision-Language Models
von: Seferis, Emmanouil, et al.
Veröffentlicht: (2025)
von: Seferis, Emmanouil, et al.
Veröffentlicht: (2025)
SmoothVLA: Aligning Vision-Language-Action Models with Physical Constraints via Intrinsic Smoothness Optimization
von: Li, Jiashun, et al.
Veröffentlicht: (2026)
von: Li, Jiashun, et al.
Veröffentlicht: (2026)
Smoothness Adaptivity in Constant-Depth Neural Networks: Optimal Rates via Smooth Activations
von: Liu, Yuhao, et al.
Veröffentlicht: (2026)
von: Liu, Yuhao, et al.
Veröffentlicht: (2026)
ASER: Activation Smoothing and Error Reconstruction for Large Language Model Quantization
von: Zhao, Weibo, et al.
Veröffentlicht: (2024)
von: Zhao, Weibo, et al.
Veröffentlicht: (2024)
SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
von: Zhou, Sashuai, et al.
Veröffentlicht: (2026)
von: Zhou, Sashuai, et al.
Veröffentlicht: (2026)
Operator Deep Smoothing for Implied Volatility
von: Wiedemann, Ruben, et al.
Veröffentlicht: (2024)
von: Wiedemann, Ruben, et al.
Veröffentlicht: (2024)
Video Models Can Reason with Verifiable Rewards
von: Zhu, Tinghui, et al.
Veröffentlicht: (2026)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2026)
Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model
von: Li, Tianle, et al.
Veröffentlicht: (2025)
von: Li, Tianle, et al.
Veröffentlicht: (2025)
Preconditioning and Reduced-Order Modeling of Navier-Stokes Equations in Complex Porous Microstructures
von: Li, Kangan, et al.
Veröffentlicht: (2025)
von: Li, Kangan, et al.
Veröffentlicht: (2025)
Regularizing Differentiable Architecture Search with Smooth Activation
von: Zhou, Yanlin, et al.
Veröffentlicht: (2025)
von: Zhou, Yanlin, et al.
Veröffentlicht: (2025)
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
von: Wang, Peisong, et al.
Veröffentlicht: (2025)
von: Wang, Peisong, et al.
Veröffentlicht: (2025)
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2024)
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2024)
Few-Shot Vision-Language Reasoning for Satellite Imagery via Verifiable Rewards
von: Koksal, Aybora, et al.
Veröffentlicht: (2025)
von: Koksal, Aybora, et al.
Veröffentlicht: (2025)
Verifiable Process Rewards for Agentic Reasoning
von: Yuan, Huining, et al.
Veröffentlicht: (2026)
von: Yuan, Huining, et al.
Veröffentlicht: (2026)
Verifying the Smoothness of Graph Signals: A Graph Signal Processing Approach
von: Dabush, Lital, et al.
Veröffentlicht: (2023)
von: Dabush, Lital, et al.
Veröffentlicht: (2023)
Bridging Smoothness and Approximation: Theoretical Insights into Over-Smoothing in Graph Neural Networks
von: Yang, Guangrui, et al.
Veröffentlicht: (2024)
von: Yang, Guangrui, et al.
Veröffentlicht: (2024)
Pioglitazone Regulates Chondrocyte Metabolism and Attenuates Osteoarthritis by Activating Peroxisome Proliferator‐Activated Receptor Gamma
von: Jiaqi Shi, et al.
Veröffentlicht: (2025)
von: Jiaqi Shi, et al.
Veröffentlicht: (2025)
Partial Smoothness, Subdifferentials and Set-valued Operators
von: Qin, Ziqi, et al.
Veröffentlicht: (2025)
von: Qin, Ziqi, et al.
Veröffentlicht: (2025)
Kernel Smoothing Operators on Thick Open Domains
von: Giannakis, Dimitrios, et al.
Veröffentlicht: (2024)
von: Giannakis, Dimitrios, et al.
Veröffentlicht: (2024)
Randomness of Shapes and Statistical Inference on Shapes via the Smooth Euler Characteristic Transform
von: Meng, Kun, et al.
Veröffentlicht: (2022)
von: Meng, Kun, et al.
Veröffentlicht: (2022)
Diffeological Smoothness in Hodge Theory
von: Li, Jiayong
Veröffentlicht: (2009)
von: Li, Jiayong
Veröffentlicht: (2009)
PSS-BA: LiDAR Bundle Adjustment with Progressive Spatial Smoothing
von: Li, Jianping, et al.
Veröffentlicht: (2024)
von: Li, Jianping, et al.
Veröffentlicht: (2024)
Compiler Bugs Detection in Logic Synthesis Tools via Linear Upper Confidence Bound
von: Zeng, Hui, et al.
Veröffentlicht: (2025)
von: Zeng, Hui, et al.
Veröffentlicht: (2025)
Aliasing Reduction in Neural Amp Modeling by Smoothing Activations
von: Sato, Ryota, et al.
Veröffentlicht: (2025)
von: Sato, Ryota, et al.
Veröffentlicht: (2025)
Doubly Smoothed Optimistic Gradients: A Universal Approach for Smooth Minimax Problems
von: Zheng, Taoli, et al.
Veröffentlicht: (2025)
von: Zheng, Taoli, et al.
Veröffentlicht: (2025)
Smooth Non-Stationary Bandits
von: Jia, Su, et al.
Veröffentlicht: (2023)
von: Jia, Su, et al.
Veröffentlicht: (2023)
Smooth, Sparse, and Stable: Finite-Time Exact Skeleton Recovery via Smoothed Proximal Gradients
von: Wu, Rui, et al.
Veröffentlicht: (2026)
von: Wu, Rui, et al.
Veröffentlicht: (2026)
ABot-OCR Technical Report
von: Jiang, Kaitao, et al.
Veröffentlicht: (2026)
von: Jiang, Kaitao, et al.
Veröffentlicht: (2026)
Promoting Efficient Reasoning with Verifiable Stepwise Reward
von: Yue, Chuhuai, et al.
Veröffentlicht: (2025)
von: Yue, Chuhuai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
von: Zhu, Banghua, et al.
Veröffentlicht: (2024) -
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
von: Zhang, Zijing, et al.
Veröffentlicht: (2025) -
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
von: He, Haoran, et al.
Veröffentlicht: (2025) -
DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
von: Lee, Vint, et al.
Veröffentlicht: (2023) -
Protocols for Verifying Smooth Strategies in Bandits and Games
von: Christ, Miranda, et al.
Veröffentlicht: (2025)