Milestones over Outcome: Unlocking Geometric Reasoning with Sub-Goal Verifiable Reward
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Jianlong, Fu, Daocheng, Xu, Shengze, Chen, Jiawei, Feng, Yuan, Yang, Yue, Yan, Junchi, Zha, Hongyuan, Xia, Renqiu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
di: Fu, Daocheng, et al.
Pubblicazione: (2025)
di: Fu, Daocheng, et al.
Pubblicazione: (2025)
GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
di: Feng, Yuan, et al.
Pubblicazione: (2025)
di: Feng, Yuan, et al.
Pubblicazione: (2025)
MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning
di: Lu, Yiyang, et al.
Pubblicazione: (2026)
di: Lu, Yiyang, et al.
Pubblicazione: (2026)
Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy
di: Ye, Hancheng, et al.
Pubblicazione: (2024)
di: Ye, Hancheng, et al.
Pubblicazione: (2024)
Fast T2T: Optimization Consistency Speeds Up Diffusion-Based Training-to-Testing Solving for Combinatorial Optimization
di: Li, Yang, et al.
Pubblicazione: (2025)
di: Li, Yang, et al.
Pubblicazione: (2025)
Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving
di: Liu, Qi, et al.
Pubblicazione: (2025)
di: Liu, Qi, et al.
Pubblicazione: (2025)
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets
di: Liao, Ning, et al.
Pubblicazione: (2023)
di: Liao, Ning, et al.
Pubblicazione: (2023)
Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning
di: Yu, Qinan, et al.
Pubblicazione: (2026)
di: Yu, Qinan, et al.
Pubblicazione: (2026)
CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling
di: Tang, Zhengyang, et al.
Pubblicazione: (2025)
di: Tang, Zhengyang, et al.
Pubblicazione: (2025)
Verifiable Process Rewards for Agentic Reasoning
di: Yuan, Huining, et al.
Pubblicazione: (2026)
di: Yuan, Huining, et al.
Pubblicazione: (2026)
Promoting Efficient Reasoning with Verifiable Stepwise Reward
di: Yue, Chuhuai, et al.
Pubblicazione: (2025)
di: Yue, Chuhuai, et al.
Pubblicazione: (2025)
Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning
di: Pronesti, Massimiliano, et al.
Pubblicazione: (2026)
di: Pronesti, Massimiliano, et al.
Pubblicazione: (2026)
Adaptive Milestone Reward for GUI Agents
di: Zheng, Congmin, et al.
Pubblicazione: (2026)
di: Zheng, Congmin, et al.
Pubblicazione: (2026)
Epistemic Gain, Aleatoric Cost: Uncertainty Decomposition in Multi-Agent Debate for Math Reasoning
di: Qiao, Dan, et al.
Pubblicazione: (2026)
di: Qiao, Dan, et al.
Pubblicazione: (2026)
Video Models Can Reason with Verifiable Rewards
di: Zhu, Tinghui, et al.
Pubblicazione: (2026)
di: Zhu, Tinghui, et al.
Pubblicazione: (2026)
ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
di: Xia, Renqiu, et al.
Pubblicazione: (2024)
di: Xia, Renqiu, et al.
Pubblicazione: (2024)
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
di: Xia, Renqiu, et al.
Pubblicazione: (2024)
di: Xia, Renqiu, et al.
Pubblicazione: (2024)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
di: Liu, Shudong, et al.
Pubblicazione: (2025)
di: Liu, Shudong, et al.
Pubblicazione: (2025)
LongR: Unleashing Long-Context Reasoning via Reinforcement Learning with Dense Utility Rewards
di: Ping, Bowen, et al.
Pubblicazione: (2026)
di: Ping, Bowen, et al.
Pubblicazione: (2026)
Milestone Review: Unlocking the Proteomics of Glycine Receptor Complexes
di: Sean D. Fraser, et al.
Pubblicazione: (2025)
di: Sean D. Fraser, et al.
Pubblicazione: (2025)
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents
di: Huang, Jiawei, et al.
Pubblicazione: (2026)
di: Huang, Jiawei, et al.
Pubblicazione: (2026)
Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding
di: Chen, Haolin, et al.
Pubblicazione: (2024)
di: Chen, Haolin, et al.
Pubblicazione: (2024)
DriveVGGT: Calibration-Constrained Visual Geometry Transformers for Multi-Camera Autonomous Driving
di: Jia, Xiaosong, et al.
Pubblicazione: (2025)
di: Jia, Xiaosong, et al.
Pubblicazione: (2025)
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
di: He, Haoran, et al.
Pubblicazione: (2025)
di: He, Haoran, et al.
Pubblicazione: (2025)
A Variational Autoencoder for Neural Temporal Point Processes with Dynamic Latent Graphs
di: Yang, Sikun, et al.
Pubblicazione: (2023)
di: Yang, Sikun, et al.
Pubblicazione: (2023)
SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing
di: Yan, Xiangchao, et al.
Pubblicazione: (2025)
di: Yan, Xiangchao, et al.
Pubblicazione: (2025)
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models
di: Zhang, Xiangxiang, et al.
Pubblicazione: (2025)
di: Zhang, Xiangxiang, et al.
Pubblicazione: (2025)
Prediction of remaining useful life for stochastic distribution systems based on hybrid residual correction method
di: Shengze Chen, et al.
Pubblicazione: (2026)
di: Shengze Chen, et al.
Pubblicazione: (2026)
Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
di: Zhao, Xiangyu, et al.
Pubblicazione: (2025)
di: Zhao, Xiangyu, et al.
Pubblicazione: (2025)
When Is Compositional Reasoning Learnable from Verifiable Rewards?
di: Barzilai, Daniel, et al.
Pubblicazione: (2026)
di: Barzilai, Daniel, et al.
Pubblicazione: (2026)
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
di: Zhang, Zijing, et al.
Pubblicazione: (2025)
di: Zhang, Zijing, et al.
Pubblicazione: (2025)
ExecVerify: White-Box RL with Verifiable Stepwise Rewards for Code Execution Reasoning
di: Tang, Lingxiao, et al.
Pubblicazione: (2026)
di: Tang, Lingxiao, et al.
Pubblicazione: (2026)
StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding
di: Xia, Renqiu, et al.
Pubblicazione: (2023)
di: Xia, Renqiu, et al.
Pubblicazione: (2023)
Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
di: Li, Yan, et al.
Pubblicazione: (2025)
di: Li, Yan, et al.
Pubblicazione: (2025)
ReSimAD: Zero-Shot 3D Domain Transfer for Autonomous Driving with Source Reconstruction and Target Simulation
di: Zhang, Bo, et al.
Pubblicazione: (2023)
di: Zhang, Bo, et al.
Pubblicazione: (2023)
ProtoReasoning: Prototypes as the Foundation for Generalizable Reasoning in LLMs
di: He, Feng, et al.
Pubblicazione: (2025)
di: He, Feng, et al.
Pubblicazione: (2025)
SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations
di: Yan, Xiangchao, et al.
Pubblicazione: (2023)
di: Yan, Xiangchao, et al.
Pubblicazione: (2023)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
di: Feng, Xiang, et al.
Pubblicazione: (2026)
di: Feng, Xiang, et al.
Pubblicazione: (2026)
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models
di: Kim, Yoonjeon, et al.
Pubblicazione: (2025)
di: Kim, Yoonjeon, et al.
Pubblicazione: (2025)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
Documenti analoghi
-
TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
di: Fu, Daocheng, et al.
Pubblicazione: (2025) -
GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
di: Feng, Yuan, et al.
Pubblicazione: (2025) -
MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning
di: Lu, Yiyang, et al.
Pubblicazione: (2026) -
Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy
di: Ye, Hancheng, et al.
Pubblicazione: (2024) -
Fast T2T: Optimization Consistency Speeds Up Diffusion-Based Training-to-Testing Solving for Combinatorial Optimization
di: Li, Yang, et al.
Pubblicazione: (2025)