Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Chonghua, Duan, Haodong, Zhang, Songyang, Lin, Dahua, Chen, Kai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
von: Cao, Maosong, et al.
Veröffentlicht: (2024)
von: Cao, Maosong, et al.
Veröffentlicht: (2024)
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
von: Zhuo, Jingming, et al.
Veröffentlicht: (2024)
von: Zhuo, Jingming, et al.
Veröffentlicht: (2024)
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
von: Cao, Maosong, et al.
Veröffentlicht: (2025)
von: Cao, Maosong, et al.
Veröffentlicht: (2025)
Are Your LLMs Capable of Stable Reasoning?
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
von: Liu, Shudong, et al.
Veröffentlicht: (2025)
von: Liu, Shudong, et al.
Veröffentlicht: (2025)
LEAN-GitHub: Compiling GitHub LEAN repositories for a versatile LEAN prover
von: Wu, Zijian, et al.
Veröffentlicht: (2024)
von: Wu, Zijian, et al.
Veröffentlicht: (2024)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
von: Li, Wei, et al.
Veröffentlicht: (2024)
von: Li, Wei, et al.
Veröffentlicht: (2024)
CriticEval: Evaluating Large Language Model as Critic
von: Lan, Tian, et al.
Veröffentlicht: (2024)
von: Lan, Tian, et al.
Veröffentlicht: (2024)
Rectifying LLM Thought from Lens of Optimization
von: Liu, Junnan, et al.
Veröffentlicht: (2025)
von: Liu, Junnan, et al.
Veröffentlicht: (2025)
Long-context LLMs Struggle with Long In-context Learning
von: Li, Tianle, et al.
Veröffentlicht: (2024)
von: Li, Tianle, et al.
Veröffentlicht: (2024)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
AdaFuse: Adaptive Ensemble Decoding with Test-Time Scaling for LLMs
von: Cui, Chengming, et al.
Veröffentlicht: (2026)
von: Cui, Chengming, et al.
Veröffentlicht: (2026)
CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
von: Zhang, Taolin, et al.
Veröffentlicht: (2025)
von: Zhang, Taolin, et al.
Veröffentlicht: (2025)
AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures
von: Zhang, Situo, et al.
Veröffentlicht: (2024)
von: Zhang, Situo, et al.
Veröffentlicht: (2024)
Navigating the OverKill in Large Language Models
von: Shi, Chenyu, et al.
Veröffentlicht: (2024)
von: Shi, Chenyu, et al.
Veröffentlicht: (2024)
ANAH: Analytical Annotation of Hallucinations in Large Language Models
von: Ji, Ziwei, et al.
Veröffentlicht: (2024)
von: Ji, Ziwei, et al.
Veröffentlicht: (2024)
ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
von: Peng, Jingyu, et al.
Veröffentlicht: (2025)
von: Peng, Jingyu, et al.
Veröffentlicht: (2025)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
von: Liu, Songyang, et al.
Veröffentlicht: (2025)
von: Liu, Songyang, et al.
Veröffentlicht: (2025)
Coding Triangle: How Does Large Language Model Understand Code?
von: Zhang, Taolin, et al.
Veröffentlicht: (2025)
von: Zhang, Taolin, et al.
Veröffentlicht: (2025)
TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain
von: Barboule, Camille, et al.
Veröffentlicht: (2024)
von: Barboule, Camille, et al.
Veröffentlicht: (2024)
GTA: A Benchmark for General Tool Agents
von: Wang, Jize, et al.
Veröffentlicht: (2024)
von: Wang, Jize, et al.
Veröffentlicht: (2024)
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs
von: Wang, Kai, et al.
Veröffentlicht: (2026)
von: Wang, Kai, et al.
Veröffentlicht: (2026)
Fake Alignment: Are LLMs Really Aligned Well?
von: Wang, Yixu, et al.
Veröffentlicht: (2023)
von: Wang, Yixu, et al.
Veröffentlicht: (2023)
Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2026)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2026)
WHODUNIT: Evaluation benchmark for culprit detection in mystery stories
von: Gupta, Kshitij
Veröffentlicht: (2025)
von: Gupta, Kshitij
Veröffentlicht: (2025)
Unveiling the Misuse Potential of Base Large Language Models via In-Context Learning
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
Flames: Benchmarking Value Alignment of LLMs in Chinese
von: Huang, Kexin, et al.
Veröffentlicht: (2023)
von: Huang, Kexin, et al.
Veröffentlicht: (2023)
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
von: Jiang, Ziyan, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyan, et al.
Veröffentlicht: (2024)
InternLM2.5-StepProver: Advancing Automated Theorem Proving via Critic-Guided Search
von: Wu, Zijian, et al.
Veröffentlicht: (2024)
von: Wu, Zijian, et al.
Veröffentlicht: (2024)
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
von: Gao, Songyang, et al.
Veröffentlicht: (2025)
von: Gao, Songyang, et al.
Veröffentlicht: (2025)
Ada-Instruct: Adapting Instruction Generators for Complex Reasoning
von: Cui, Wanyun, et al.
Veröffentlicht: (2023)
von: Cui, Wanyun, et al.
Veröffentlicht: (2023)
Does quantization affect models' performance on long-context tasks?
von: Mekala, Anmol, et al.
Veröffentlicht: (2025)
von: Mekala, Anmol, et al.
Veröffentlicht: (2025)
AI PERSONA: Towards Life-long Personalization of LLMs
von: Wang, Tiannan, et al.
Veröffentlicht: (2024)
von: Wang, Tiannan, et al.
Veröffentlicht: (2024)
Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
von: Zhao, Yufeng, et al.
Veröffentlicht: (2025)
von: Zhao, Yufeng, et al.
Veröffentlicht: (2025)
AdaThink-Med: Medical Adaptive Thinking with Uncertainty-Guided Length Calibration
von: Rui, Shaohao, et al.
Veröffentlicht: (2025)
von: Rui, Shaohao, et al.
Veröffentlicht: (2025)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
von: Ding, Liang
Veröffentlicht: (2026)
von: Ding, Liang
Veröffentlicht: (2026)
Ähnliche Einträge
-
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
von: Cao, Maosong, et al.
Veröffentlicht: (2024) -
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
von: Zhuo, Jingming, et al.
Veröffentlicht: (2024) -
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
von: Cao, Maosong, et al.
Veröffentlicht: (2025) -
Are Your LLMs Capable of Stable Reasoning?
von: Liu, Junnan, et al.
Veröffentlicht: (2024) -
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
von: Liu, Shudong, et al.
Veröffentlicht: (2025)