Saved in:
| Main Authors: | Cao, Maosong, Chen, Kai, Duan, Haodong, Fang, Yixiao, Fei, Zhiwei, Gao, Tong, Jiaye, Ge, Li, Mo, Liu, Hongwei, Liu, Junnan, Liu, Yuan, Lyu, Chengqi, Lyu, Han, Ma, Ningsheng, Ma, Zerun, Sun, Yu, Wu, Zhiyong, Xiao, Linchen, Xu, Jun, Ye, Haochen, Yu, Zhaohui, Yuan, Yike, Zhang, Songyang, Zhao, Yufeng, Zhou, Fengzhe, Zhou, Peiheng, Zhu, Dongsheng, Zhu, Lin, Zhuo, Jingming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.19276 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
by: Liu, Shudong, et al.
Published: (2025)
by: Liu, Shudong, et al.
Published: (2025)
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
by: Cao, Maosong, et al.
Published: (2024)
by: Cao, Maosong, et al.
Published: (2024)
Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective
by: Liu, Junnan, et al.
Published: (2025)
by: Liu, Junnan, et al.
Published: (2025)
Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
by: Zhao, Yufeng, et al.
Published: (2025)
by: Zhao, Yufeng, et al.
Published: (2025)
Coding Triangle: How Does Large Language Model Understand Code?
by: Zhang, Taolin, et al.
Published: (2025)
by: Zhang, Taolin, et al.
Published: (2025)
Are Your LLMs Capable of Stable Reasoning?
by: Liu, Junnan, et al.
Published: (2024)
by: Liu, Junnan, et al.
Published: (2024)
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
by: Zhang, Chuyu, et al.
Published: (2024)
by: Zhang, Chuyu, et al.
Published: (2024)
How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
by: Ma, Zihan, et al.
Published: (2025)
by: Ma, Zihan, et al.
Published: (2025)
Rethinking Verification for LLM Code Generation: From Generation to Testing
by: Ma, Zihan, et al.
Published: (2025)
by: Ma, Zihan, et al.
Published: (2025)
CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
by: Zhang, Taolin, et al.
Published: (2025)
by: Zhang, Taolin, et al.
Published: (2025)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
by: Liu, Hongwei, et al.
Published: (2024)
by: Liu, Hongwei, et al.
Published: (2024)
Rectifying LLM Thought from Lens of Optimization
by: Liu, Junnan, et al.
Published: (2025)
by: Liu, Junnan, et al.
Published: (2025)
ATLAS: A High-Difficulty, Multidisciplinary Benchmark for Frontier Scientific Reasoning
by: Liu, Hongwei, et al.
Published: (2025)
by: Liu, Hongwei, et al.
Published: (2025)
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
by: Lyu, Chengqi, et al.
Published: (2025)
by: Lyu, Chengqi, et al.
Published: (2025)
Fake Alignment: Are LLMs Really Aligned Well?
by: Wang, Yixu, et al.
Published: (2023)
by: Wang, Yixu, et al.
Published: (2023)
XYZCylinder: Towards Compatible Feed-Forward 3D Gaussian Splatting for Driving Scenes via Unified Cylinder Lifting Method
by: Yu, Haochen, et al.
Published: (2025)
by: Yu, Haochen, et al.
Published: (2025)
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
by: Hua, Zhouqi, et al.
Published: (2025)
by: Hua, Zhouqi, et al.
Published: (2025)
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
by: Zhuo, Jingming, et al.
Published: (2024)
by: Zhuo, Jingming, et al.
Published: (2024)
InternLM-Law: An Open Source Chinese Legal Large Language Model
by: Fei, Zhiwei, et al.
Published: (2024)
by: Fei, Zhiwei, et al.
Published: (2024)
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
by: Gao, Songyang, et al.
Published: (2025)
by: Gao, Songyang, et al.
Published: (2025)
MMBench: Is Your Multi-modal Model an All-around Player?
by: Liu, Yuan, et al.
Published: (2023)
by: Liu, Yuan, et al.
Published: (2023)
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
by: Cao, Maosong, et al.
Published: (2025)
by: Cao, Maosong, et al.
Published: (2025)
Closing the Gaps: Optimality of Sample Average Approximation for Data-Driven Newsvendor Problems
by: Lyu, Jiameng, et al.
Published: (2024)
by: Lyu, Jiameng, et al.
Published: (2024)
GTA: A Benchmark for General Tool Agents
by: Wang, Jize, et al.
Published: (2024)
by: Wang, Jize, et al.
Published: (2024)
MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space
by: Chen, Yicheng, et al.
Published: (2025)
by: Chen, Yicheng, et al.
Published: (2025)
A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
by: Luo, Kairong, et al.
Published: (2025)
by: Luo, Kairong, et al.
Published: (2025)
Context-Adaptive Synthesis and Compression for Enhanced Retrieval-Augmented Generation in Complex Domains
by: Zhou, Peiran, et al.
Published: (2025)
by: Zhou, Peiran, et al.
Published: (2025)
Introducing the O-Value: A Universal Standardization for Confusion-Matrix-Based Classification Performance Metrics
by: Zhao, Ningsheng, et al.
Published: (2025)
by: Zhao, Ningsheng, et al.
Published: (2025)
Error Analysis of Shapley Value-Based Model Explanations: An Informative Perspective
by: Zhao, Ningsheng, et al.
Published: (2024)
by: Zhao, Ningsheng, et al.
Published: (2024)
A new proof of Hölder estimates for the gradient of quasilinear elliptic equations
by: Li, Dongsheng, et al.
Published: (2025)
by: Li, Dongsheng, et al.
Published: (2025)
A Modulo Sampling Hardware Prototype and Reconstruction Algorithm Evaluation
by: Zhu, Jiang, et al.
Published: (2024)
by: Zhu, Jiang, et al.
Published: (2024)
Unlocking Attributes' Contribution to Successful Camouflage: A Combined Textual and VisualAnalysis Strategy
by: Zhang, Hong, et al.
Published: (2024)
by: Zhang, Hong, et al.
Published: (2024)
CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
by: Li, Jingliang, et al.
Published: (2026)
by: Li, Jingliang, et al.
Published: (2026)
Strong-field ionization of atoms with bright squeezed vacuum light
by: Liu, Haodong, et al.
Published: (2026)
by: Liu, Haodong, et al.
Published: (2026)
Intelligent Fault Diagnosis of Electrohydraulic Servo Valve Based on Kalman Filter and BiLSTM‐ANN Hybrid Deep Model
by: Liu Zerun, et al.
Published: (2026)
by: Liu Zerun, et al.
Published: (2026)
A Minibatch-SGD-Based Learning Meta-Policy for Inventory Systems with Myopic Optimal Policy
by: Lyu, Jiameng, et al.
Published: (2024)
by: Lyu, Jiameng, et al.
Published: (2024)
Learning When to Restart: Nonstationary Newsvendor from Uncensored to Censored Demand
by: Chen, Xin, et al.
Published: (2025)
by: Chen, Xin, et al.
Published: (2025)
CC-DCNet: Dynamic Convolutional Neural Network with Contrastive Constraints for Identifying Lung Cancer Subtypes on Multi-modality Images
by: Jin, Yuan, et al.
Published: (2024)
by: Jin, Yuan, et al.
Published: (2024)
The Effect and Mechanism of Self‐Compassion on Reducing Materialism: A Randomized Controlled Trial of an Online Self‐Compassion Intervention
by: Xiaodan Gu, et al.
Published: (2025)
by: Xiaodan Gu, et al.
Published: (2025)
Identification and Stress‐Induced Expression of Antimicrobial Peptides SlAttacins in Shelfordella lateralis
by: Tielong Xu, et al.
Published: (2026)
by: Tielong Xu, et al.
Published: (2026)
Similar Items
-
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
by: Liu, Shudong, et al.
Published: (2025) -
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
by: Cao, Maosong, et al.
Published: (2024) -
Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective
by: Liu, Junnan, et al.
Published: (2025) -
Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
by: Zhao, Yufeng, et al.
Published: (2025) -
Coding Triangle: How Does Large Language Model Understand Code?
by: Zhang, Taolin, et al.
Published: (2025)