Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Chonghua, Duan, Haodong, Zhang, Songyang, Lin, Dahua, Chen, Kai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
di: Cao, Maosong, et al.
Pubblicazione: (2024)
di: Cao, Maosong, et al.
Pubblicazione: (2024)
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
di: Zhuo, Jingming, et al.
Pubblicazione: (2024)
di: Zhuo, Jingming, et al.
Pubblicazione: (2024)
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
di: Cao, Maosong, et al.
Pubblicazione: (2025)
di: Cao, Maosong, et al.
Pubblicazione: (2025)
Are Your LLMs Capable of Stable Reasoning?
di: Liu, Junnan, et al.
Pubblicazione: (2024)
di: Liu, Junnan, et al.
Pubblicazione: (2024)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
di: Liu, Shudong, et al.
Pubblicazione: (2025)
di: Liu, Shudong, et al.
Pubblicazione: (2025)
LEAN-GitHub: Compiling GitHub LEAN repositories for a versatile LEAN prover
di: Wu, Zijian, et al.
Pubblicazione: (2024)
di: Wu, Zijian, et al.
Pubblicazione: (2024)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
di: Zhang, Beichen, et al.
Pubblicazione: (2025)
di: Zhang, Beichen, et al.
Pubblicazione: (2025)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
di: Li, Wei, et al.
Pubblicazione: (2024)
di: Li, Wei, et al.
Pubblicazione: (2024)
CriticEval: Evaluating Large Language Model as Critic
di: Lan, Tian, et al.
Pubblicazione: (2024)
di: Lan, Tian, et al.
Pubblicazione: (2024)
Rectifying LLM Thought from Lens of Optimization
di: Liu, Junnan, et al.
Pubblicazione: (2025)
di: Liu, Junnan, et al.
Pubblicazione: (2025)
Long-context LLMs Struggle with Long In-context Learning
di: Li, Tianle, et al.
Pubblicazione: (2024)
di: Li, Tianle, et al.
Pubblicazione: (2024)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
di: Liu, Hongwei, et al.
Pubblicazione: (2024)
di: Liu, Hongwei, et al.
Pubblicazione: (2024)
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
di: Zhang, Beichen, et al.
Pubblicazione: (2025)
di: Zhang, Beichen, et al.
Pubblicazione: (2025)
AdaFuse: Adaptive Ensemble Decoding with Test-Time Scaling for LLMs
di: Cui, Chengming, et al.
Pubblicazione: (2026)
di: Cui, Chengming, et al.
Pubblicazione: (2026)
CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
di: Zhang, Taolin, et al.
Pubblicazione: (2025)
di: Zhang, Taolin, et al.
Pubblicazione: (2025)
AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures
di: Zhang, Situo, et al.
Pubblicazione: (2024)
di: Zhang, Situo, et al.
Pubblicazione: (2024)
Navigating the OverKill in Large Language Models
di: Shi, Chenyu, et al.
Pubblicazione: (2024)
di: Shi, Chenyu, et al.
Pubblicazione: (2024)
ANAH: Analytical Annotation of Hallucinations in Large Language Models
di: Ji, Ziwei, et al.
Pubblicazione: (2024)
di: Ji, Ziwei, et al.
Pubblicazione: (2024)
ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models
di: Gu, Yuzhe, et al.
Pubblicazione: (2024)
di: Gu, Yuzhe, et al.
Pubblicazione: (2024)
AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
di: Peng, Jingyu, et al.
Pubblicazione: (2025)
di: Peng, Jingyu, et al.
Pubblicazione: (2025)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
di: Liu, Songyang, et al.
Pubblicazione: (2025)
di: Liu, Songyang, et al.
Pubblicazione: (2025)
Coding Triangle: How Does Large Language Model Understand Code?
di: Zhang, Taolin, et al.
Pubblicazione: (2025)
di: Zhang, Taolin, et al.
Pubblicazione: (2025)
TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain
di: Barboule, Camille, et al.
Pubblicazione: (2024)
di: Barboule, Camille, et al.
Pubblicazione: (2024)
GTA: A Benchmark for General Tool Agents
di: Wang, Jize, et al.
Pubblicazione: (2024)
di: Wang, Jize, et al.
Pubblicazione: (2024)
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs
di: Wang, Kai, et al.
Pubblicazione: (2026)
di: Wang, Kai, et al.
Pubblicazione: (2026)
Fake Alignment: Are LLMs Really Aligned Well?
di: Wang, Yixu, et al.
Pubblicazione: (2023)
di: Wang, Yixu, et al.
Pubblicazione: (2023)
Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
di: Abdaljalil, Samir, et al.
Pubblicazione: (2026)
di: Abdaljalil, Samir, et al.
Pubblicazione: (2026)
WHODUNIT: Evaluation benchmark for culprit detection in mystery stories
di: Gupta, Kshitij
Pubblicazione: (2025)
di: Gupta, Kshitij
Pubblicazione: (2025)
Unveiling the Misuse Potential of Base Large Language Models via In-Context Learning
di: Wang, Xiao, et al.
Pubblicazione: (2024)
di: Wang, Xiao, et al.
Pubblicazione: (2024)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
di: Fu, Chaoyou, et al.
Pubblicazione: (2024)
di: Fu, Chaoyou, et al.
Pubblicazione: (2024)
Flames: Benchmarking Value Alignment of LLMs in Chinese
di: Huang, Kexin, et al.
Pubblicazione: (2023)
di: Huang, Kexin, et al.
Pubblicazione: (2023)
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
di: Jiang, Ziyan, et al.
Pubblicazione: (2024)
di: Jiang, Ziyan, et al.
Pubblicazione: (2024)
InternLM2.5-StepProver: Advancing Automated Theorem Proving via Critic-Guided Search
di: Wu, Zijian, et al.
Pubblicazione: (2024)
di: Wu, Zijian, et al.
Pubblicazione: (2024)
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
di: Gao, Songyang, et al.
Pubblicazione: (2025)
di: Gao, Songyang, et al.
Pubblicazione: (2025)
Ada-Instruct: Adapting Instruction Generators for Complex Reasoning
di: Cui, Wanyun, et al.
Pubblicazione: (2023)
di: Cui, Wanyun, et al.
Pubblicazione: (2023)
Does quantization affect models' performance on long-context tasks?
di: Mekala, Anmol, et al.
Pubblicazione: (2025)
di: Mekala, Anmol, et al.
Pubblicazione: (2025)
AI PERSONA: Towards Life-long Personalization of LLMs
di: Wang, Tiannan, et al.
Pubblicazione: (2024)
di: Wang, Tiannan, et al.
Pubblicazione: (2024)
Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
di: Zhao, Yufeng, et al.
Pubblicazione: (2025)
di: Zhao, Yufeng, et al.
Pubblicazione: (2025)
AdaThink-Med: Medical Adaptive Thinking with Uncertainty-Guided Length Calibration
di: Rui, Shaohao, et al.
Pubblicazione: (2025)
di: Rui, Shaohao, et al.
Pubblicazione: (2025)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
di: Ding, Liang
Pubblicazione: (2026)
di: Ding, Liang
Pubblicazione: (2026)
Documenti analoghi
-
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
di: Cao, Maosong, et al.
Pubblicazione: (2024) -
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
di: Zhuo, Jingming, et al.
Pubblicazione: (2024) -
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
di: Cao, Maosong, et al.
Pubblicazione: (2025) -
Are Your LLMs Capable of Stable Reasoning?
di: Liu, Junnan, et al.
Pubblicazione: (2024) -
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
di: Liu, Shudong, et al.
Pubblicazione: (2025)