Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
Fuente:
arXiv
Salvato in:
| Autori principali: | Pandit, Shrey, Xu, Austin, Nguyen, Xuan-Phi, Ming, Yifei, Xiong, Caiming, Joty, Shafiq |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2026)
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2026)
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
di: Ming, Yifei, et al.
Pubblicazione: (2024)
di: Ming, Yifei, et al.
Pubblicazione: (2024)
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2025)
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2025)
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
SFR-RAG: Towards Contextually Faithful LLMs
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2024)
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2024)
Demystifying Domain-adaptive Post-training for Financial LLMs
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
MAS-ZERO: Designing Multi-Agent Systems with Zero Supervision
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
di: Ke, Zixuan, et al.
Pubblicazione: (2026)
di: Ke, Zixuan, et al.
Pubblicazione: (2026)
Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction
di: Shi, Zhenmei, et al.
Pubblicazione: (2024)
di: Shi, Zhenmei, et al.
Pubblicazione: (2024)
The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering
di: Zhou, Yefan, et al.
Pubblicazione: (2026)
di: Zhou, Yefan, et al.
Pubblicazione: (2026)
Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse Prompts
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2023)
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2023)
NAACL2025 Tutorial: Adaptation of Large Language Models
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows
di: Ming, Yifei, et al.
Pubblicazione: (2025)
di: Ming, Yifei, et al.
Pubblicazione: (2025)
Variation in Verification: Understanding Verification Dynamics in Large Language Models
di: Zhou, Yefan, et al.
Pubblicazione: (2025)
di: Zhou, Yefan, et al.
Pubblicazione: (2025)
Direct Judgement Preference Optimization
di: Wang, Peifeng, et al.
Pubblicazione: (2024)
di: Wang, Peifeng, et al.
Pubblicazione: (2024)
ParaICL: Towards Parallel In-Context Learning
di: Li, Xingxuan, et al.
Pubblicazione: (2024)
di: Li, Xingxuan, et al.
Pubblicazione: (2024)
JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking
di: Niu, Tong, et al.
Pubblicazione: (2024)
di: Niu, Tong, et al.
Pubblicazione: (2024)
Let's Verify Math Questions Step by Step
di: Shen, Chengyu, et al.
Pubblicazione: (2025)
di: Shen, Chengyu, et al.
Pubblicazione: (2025)
Efficiently Aligned Cross-Lingual Transfer Learning for Conversational Tasks using Prompt-Tuning
di: Tu, Lifu, et al.
Pubblicazione: (2023)
di: Tu, Lifu, et al.
Pubblicazione: (2023)
StructTest: Benchmarking LLMs' Reasoning through Compositional Structured Outputs
di: Chen, Hailin, et al.
Pubblicazione: (2024)
di: Chen, Hailin, et al.
Pubblicazione: (2024)
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
di: Wang, Peiyi, et al.
Pubblicazione: (2023)
di: Wang, Peiyi, et al.
Pubblicazione: (2023)
LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
di: Wang, Jiayu, et al.
Pubblicazione: (2025)
di: Wang, Jiayu, et al.
Pubblicazione: (2025)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
di: Xie, Tianbao, et al.
Pubblicazione: (2024)
di: Xie, Tianbao, et al.
Pubblicazione: (2024)
Pessimistic Verification for Open Ended Math Questions
di: Huang, Yanxing, et al.
Pubblicazione: (2025)
di: Huang, Yanxing, et al.
Pubblicazione: (2025)
MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems
di: Venkataramani, Vishal, et al.
Pubblicazione: (2026)
di: Venkataramani, Vishal, et al.
Pubblicazione: (2026)
References Improve LLM Alignment in Non-Verifiable Domains
di: Shi, Kejian, et al.
Pubblicazione: (2026)
di: Shi, Kejian, et al.
Pubblicazione: (2026)
Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning
di: Wang, Jiayu, et al.
Pubblicazione: (2025)
di: Wang, Jiayu, et al.
Pubblicazione: (2025)
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
di: Xu, Liang, et al.
Pubblicazione: (2024)
di: Xu, Liang, et al.
Pubblicazione: (2024)
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
di: Pang, Bo, et al.
Pubblicazione: (2025)
di: Pang, Bo, et al.
Pubblicazione: (2025)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization
di: Singh, Janvijay, et al.
Pubblicazione: (2025)
di: Singh, Janvijay, et al.
Pubblicazione: (2025)
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
di: Huang, Ailin, et al.
Pubblicazione: (2026)
di: Huang, Ailin, et al.
Pubblicazione: (2026)
Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
di: Li, Chuyuan, et al.
Pubblicazione: (2025)
di: Li, Chuyuan, et al.
Pubblicazione: (2025)
ChatGPT's One-year Anniversary: Are Open-Source Large Language Models Catching up?
di: Chen, Hailin, et al.
Pubblicazione: (2023)
di: Chen, Hailin, et al.
Pubblicazione: (2023)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
di: Yao, Huanjin, et al.
Pubblicazione: (2025)
di: Yao, Huanjin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms
di: Pandit, Shrey, et al.
Pubblicazione: (2025) -
Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2026) -
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
di: Ming, Yifei, et al.
Pubblicazione: (2024) -
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2025) -
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
di: Xu, Austin, et al.
Pubblicazione: (2025)