ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Yin, Zhangyue, Sun, Qiushi, Zeng, Zhiyuan, Yu, Zhiyuan, Guo, Qipeng, Huang, Xuanjing, Qiu, Xipeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dynamic and Generalizable Process Reward Modeling
di: Yin, Zhangyue, et al.
Pubblicazione: (2025)
di: Yin, Zhangyue, et al.
Pubblicazione: (2025)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2025)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2025)
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework
di: Sun, Yuhong, et al.
Pubblicazione: (2025)
di: Sun, Yuhong, et al.
Pubblicazione: (2025)
Aggregation of Reasoning: A Hierarchical Framework for Enhancing Answer Selection in Large Language Models
di: Yin, Zhangyue, et al.
Pubblicazione: (2024)
di: Yin, Zhangyue, et al.
Pubblicazione: (2024)
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem
di: Sun, Yuhong, et al.
Pubblicazione: (2024)
di: Sun, Yuhong, et al.
Pubblicazione: (2024)
Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration
di: Sun, Qiushi, et al.
Pubblicazione: (2023)
di: Sun, Qiushi, et al.
Pubblicazione: (2023)
RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization
di: Zhiyuan, Zeng, et al.
Pubblicazione: (2025)
di: Zhiyuan, Zeng, et al.
Pubblicazione: (2025)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
Unearthing Large Scale Domain-Specific Knowledge from Public Corpora
di: Fei, Zhaoye, et al.
Pubblicazione: (2024)
di: Fei, Zhaoye, et al.
Pubblicazione: (2024)
How to Set the Learning Rate for Large-Scale Pre-training?
di: Zhou, Yunhua, et al.
Pubblicazione: (2026)
di: Zhou, Yunhua, et al.
Pubblicazione: (2026)
Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric
di: Cao, Yixin, et al.
Pubblicazione: (2025)
di: Cao, Yixin, et al.
Pubblicazione: (2025)
F-Eval: Assessing Fundamental Abilities with Refined Evaluation Methods
di: Sun, Yu, et al.
Pubblicazione: (2024)
di: Sun, Yu, et al.
Pubblicazione: (2024)
DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels
di: Xu, Zhe, et al.
Pubblicazione: (2024)
di: Xu, Zhe, et al.
Pubblicazione: (2024)
Scaling Laws for Fact Memorization of Large Language Models
di: Lu, Xingyu, et al.
Pubblicazione: (2024)
di: Lu, Xingyu, et al.
Pubblicazione: (2024)
Emergent Structured Representations Support Flexible In-Context Inference in Large Language Models
di: Xu, Ningyu, et al.
Pubblicazione: (2026)
di: Xu, Ningyu, et al.
Pubblicazione: (2026)
AutoLogi: Automated Generation of Logic Puzzles for Evaluating Reasoning Abilities of Large Language Models
di: Zhu, Qin, et al.
Pubblicazione: (2025)
di: Zhu, Qin, et al.
Pubblicazione: (2025)
AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
di: Wang, Weiyi, et al.
Pubblicazione: (2026)
di: Wang, Weiyi, et al.
Pubblicazione: (2026)
Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation
di: Zhu, Qin, et al.
Pubblicazione: (2024)
di: Zhu, Qin, et al.
Pubblicazione: (2024)
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
di: Zhang, Shiduo, et al.
Pubblicazione: (2024)
di: Zhang, Shiduo, et al.
Pubblicazione: (2024)
AdaLomo: Low-memory Optimization with Adaptive Learning Rate
di: Lv, Kai, et al.
Pubblicazione: (2023)
di: Lv, Kai, et al.
Pubblicazione: (2023)
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2026)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2026)
DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting
di: Lv, Kai, et al.
Pubblicazione: (2025)
di: Lv, Kai, et al.
Pubblicazione: (2025)
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
di: Wang, Yuxin, et al.
Pubblicazione: (2025)
di: Wang, Yuxin, et al.
Pubblicazione: (2025)
VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interaction
di: Yang, Jie, et al.
Pubblicazione: (2025)
di: Yang, Jie, et al.
Pubblicazione: (2025)
Beyond Attention Magnitude: Leveraging Inter-layer Rank Consistency for Efficient Vision-Language-Action Models
di: Liu, Peiju, et al.
Pubblicazione: (2026)
di: Liu, Peiju, et al.
Pubblicazione: (2026)
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
di: Li, Xiaonan, et al.
Pubblicazione: (2023)
di: Li, Xiaonan, et al.
Pubblicazione: (2023)
Full Parameter Fine-tuning for Large Language Models with Limited Resources
di: Lv, Kai, et al.
Pubblicazione: (2023)
di: Lv, Kai, et al.
Pubblicazione: (2023)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
di: Peng, Runyu, et al.
Pubblicazione: (2026)
di: Peng, Runyu, et al.
Pubblicazione: (2026)
BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement Learning
di: Li, Yuan, et al.
Pubblicazione: (2026)
di: Li, Yuan, et al.
Pubblicazione: (2026)
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
di: Shi, Junhao, et al.
Pubblicazione: (2025)
di: Shi, Junhao, et al.
Pubblicazione: (2025)
Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
di: Ma, Yichuan, et al.
Pubblicazione: (2026)
di: Ma, Yichuan, et al.
Pubblicazione: (2026)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
di: Wang, Bo, et al.
Pubblicazione: (2025)
di: Wang, Bo, et al.
Pubblicazione: (2025)
Prior-Fitted Networks Scale to Larger Datasets When Treated as Weak Learners
di: Wang, Yuxin, et al.
Pubblicazione: (2025)
di: Wang, Yuxin, et al.
Pubblicazione: (2025)
Data-free Weight Compress and Denoise for Large Language Models
di: Peng, Runyu, et al.
Pubblicazione: (2024)
di: Peng, Runyu, et al.
Pubblicazione: (2024)
Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization
di: Zhai, Zhiyuan, et al.
Pubblicazione: (2026)
di: Zhai, Zhiyuan, et al.
Pubblicazione: (2026)
ARISE: Adaptive Reinforcement Integrated with Swarm Exploration
di: M, Rajiv Chaitanya, et al.
Pubblicazione: (2026)
di: M, Rajiv Chaitanya, et al.
Pubblicazione: (2026)
ARISE -- Adaptive Refinement and Iterative Scenario Engineering
di: Poddubnyy, Konstantin, et al.
Pubblicazione: (2026)
di: Poddubnyy, Konstantin, et al.
Pubblicazione: (2026)
Explicit Multi-head Attention for Inter-head Interaction in Large Language Models
di: Peng, Runyu, et al.
Pubblicazione: (2026)
di: Peng, Runyu, et al.
Pubblicazione: (2026)
Evaluating Large Language Models at Evaluating Instruction Following
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2023)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Dynamic and Generalizable Process Reward Modeling
di: Yin, Zhangyue, et al.
Pubblicazione: (2025) -
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2025) -
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024) -
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework
di: Sun, Yuhong, et al.
Pubblicazione: (2025) -
Aggregation of Reasoning: A Hierarchical Framework for Enhancing Answer Selection in Large Language Models
di: Yin, Zhangyue, et al.
Pubblicazione: (2024)