CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Taolin, Cao, Maosong, Lam, Alexander, Zhang, Songyang, Chen, Kai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
di: Cao, Maosong, et al.
Pubblicazione: (2024)
di: Cao, Maosong, et al.
Pubblicazione: (2024)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
di: Liu, Shudong, et al.
Pubblicazione: (2025)
di: Liu, Shudong, et al.
Pubblicazione: (2025)
Coding Triangle: How Does Large Language Model Understand Code?
di: Zhang, Taolin, et al.
Pubblicazione: (2025)
di: Zhang, Taolin, et al.
Pubblicazione: (2025)
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
di: Cao, Maosong, et al.
Pubblicazione: (2025)
di: Cao, Maosong, et al.
Pubblicazione: (2025)
Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
di: Jin, Zhuoran, et al.
Pubblicazione: (2025)
di: Jin, Zhuoran, et al.
Pubblicazione: (2025)
Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective
di: Liu, Junnan, et al.
Pubblicazione: (2025)
di: Liu, Junnan, et al.
Pubblicazione: (2025)
Fixing the Broken Compass: Diagnosing and Improving Inference-Time Reward Modeling
di: Li, Jiachun, et al.
Pubblicazione: (2025)
di: Li, Jiachun, et al.
Pubblicazione: (2025)
Towards Generalist Prompting for Large Language Models by Mental Models
di: Guan, Haoxiang, et al.
Pubblicazione: (2024)
di: Guan, Haoxiang, et al.
Pubblicazione: (2024)
AgentV-RL: Scaling Reward Modeling with Agentic Verifier
di: Zhang, Jiazheng, et al.
Pubblicazione: (2026)
di: Zhang, Jiazheng, et al.
Pubblicazione: (2026)
Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance
di: Yan, Kai, et al.
Pubblicazione: (2026)
di: Yan, Kai, et al.
Pubblicazione: (2026)
An Information-Theoretic Framework for Robust Large Language Model Editing
di: Chen, Qizhou, et al.
Pubblicazione: (2025)
di: Chen, Qizhou, et al.
Pubblicazione: (2025)
Inference-Time Scaling for Generalist Reward Modeling
di: Liu, Zijun, et al.
Pubblicazione: (2025)
di: Liu, Zijun, et al.
Pubblicazione: (2025)
Rectifying LLM Thought from Lens of Optimization
di: Liu, Junnan, et al.
Pubblicazione: (2025)
di: Liu, Junnan, et al.
Pubblicazione: (2025)
TableLlama: Towards Open Large Generalist Models for Tables
di: Zhang, Tianshu, et al.
Pubblicazione: (2023)
di: Zhang, Tianshu, et al.
Pubblicazione: (2023)
Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
di: Gajcin, Jasmina, et al.
Pubblicazione: (2025)
di: Gajcin, Jasmina, et al.
Pubblicazione: (2025)
A Short Survey on Small Reasoning Models: Training, Inference, Applications and Research Directions
di: Wang, Chengyu, et al.
Pubblicazione: (2025)
di: Wang, Chengyu, et al.
Pubblicazione: (2025)
A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future
di: Zhong, Jialun, et al.
Pubblicazione: (2025)
di: Zhong, Jialun, et al.
Pubblicazione: (2025)
Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
di: Yan, Yuchen, et al.
Pubblicazione: (2025)
di: Yan, Yuchen, et al.
Pubblicazione: (2025)
Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
di: Ding, Meidan, et al.
Pubblicazione: (2025)
di: Ding, Meidan, et al.
Pubblicazione: (2025)
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems
di: Peng, Hao, et al.
Pubblicazione: (2025)
di: Peng, Hao, et al.
Pubblicazione: (2025)
Are Your LLMs Capable of Stable Reasoning?
di: Liu, Junnan, et al.
Pubblicazione: (2024)
di: Liu, Junnan, et al.
Pubblicazione: (2024)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
di: Wang, Chonghua, et al.
Pubblicazione: (2024)
di: Wang, Chonghua, et al.
Pubblicazione: (2024)
Rethinking Verification for LLM Code Generation: From Generation to Testing
di: Ma, Zihan, et al.
Pubblicazione: (2025)
di: Ma, Zihan, et al.
Pubblicazione: (2025)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
di: Liu, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Liu, Xiaoyuan, et al.
Pubblicazione: (2025)
Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning
di: Pronesti, Massimiliano, et al.
Pubblicazione: (2026)
di: Pronesti, Massimiliano, et al.
Pubblicazione: (2026)
Generative Floor Plan Design with LLMs via Reinforcement Learning with Verifiable Rewards
di: Lara, Luis, et al.
Pubblicazione: (2026)
di: Lara, Luis, et al.
Pubblicazione: (2026)
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
di: Wang, Peisong, et al.
Pubblicazione: (2025)
di: Wang, Peisong, et al.
Pubblicazione: (2025)
Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards
di: Wei, Xiaolong, et al.
Pubblicazione: (2025)
di: Wei, Xiaolong, et al.
Pubblicazione: (2025)
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
di: Men, Tianyi, et al.
Pubblicazione: (2025)
di: Men, Tianyi, et al.
Pubblicazione: (2025)
AgentRM: Enhancing Agent Generalization with Reward Modeling
di: Xia, Yu, et al.
Pubblicazione: (2025)
di: Xia, Yu, et al.
Pubblicazione: (2025)
Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards
di: Pavlenko, Kirill, et al.
Pubblicazione: (2026)
di: Pavlenko, Kirill, et al.
Pubblicazione: (2026)
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
di: Guan, Xinyan, et al.
Pubblicazione: (2024)
di: Guan, Xinyan, et al.
Pubblicazione: (2024)
Debate Helps Weak Judges Reward Stronger Models
di: Elasky, Ethan, et al.
Pubblicazione: (2026)
di: Elasky, Ethan, et al.
Pubblicazione: (2026)
AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
di: Kartik, NVJK, et al.
Pubblicazione: (2025)
di: Kartik, NVJK, et al.
Pubblicazione: (2025)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
di: Wu, Tianhao, et al.
Pubblicazione: (2024)
di: Wu, Tianhao, et al.
Pubblicazione: (2024)
Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity Translation
di: Zhou, Jiang, et al.
Pubblicazione: (2026)
di: Zhou, Jiang, et al.
Pubblicazione: (2026)
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
di: Liu, Yule, et al.
Pubblicazione: (2025)
di: Liu, Yule, et al.
Pubblicazione: (2025)
BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks
di: Zhang, Kai, et al.
Pubblicazione: (2023)
di: Zhang, Kai, et al.
Pubblicazione: (2023)
Advancing LLM Reasoning Generalists with Preference Trees
di: Yuan, Lifan, et al.
Pubblicazione: (2024)
di: Yuan, Lifan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
di: Cao, Maosong, et al.
Pubblicazione: (2024) -
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
di: Liu, Shudong, et al.
Pubblicazione: (2025) -
Coding Triangle: How Does Large Language Model Understand Code?
di: Zhang, Taolin, et al.
Pubblicazione: (2025) -
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
di: Cao, Maosong, et al.
Pubblicazione: (2025) -
Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
di: Jin, Zhuoran, et al.
Pubblicazione: (2025)