Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat
Fuente:
arXiv
Saved in:
| Main Authors: | Daynauth, Roland, Clarke, Christopher, Flautner, Krisztian, Tang, Lingjia, Mars, Jason |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models
by: Daynauth, Roland, et al.
Published: (2025)
by: Daynauth, Roland, et al.
Published: (2025)
Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments
by: Daynauth, Roland, et al.
Published: (2024)
by: Daynauth, Roland, et al.
Published: (2024)
Guylingo: The Republic of Guyana Creole Corpora
by: Clarke, Christopher, et al.
Published: (2024)
by: Clarke, Christopher, et al.
Published: (2024)
MTP: A Meaning-Typed Language Abstraction for AI-Integrated Programming
by: Dantanarayana, Jayanaka L., et al.
Published: (2024)
by: Dantanarayana, Jayanaka L., et al.
Published: (2024)
PEFT-U: Parameter-Efficient Fine-Tuning for User Personalization
by: Clarke, Christopher, et al.
Published: (2024)
by: Clarke, Christopher, et al.
Published: (2024)
Scaling Down to Scale Up: A Cost-Benefit Analysis of Replacing OpenAI's LLM with Open Source SLMs in Production
by: Irugalbandara, Chandra, et al.
Published: (2023)
by: Irugalbandara, Chandra, et al.
Published: (2023)
GraphRunner: A Multi-Stage Framework for Efficient and Accurate Graph-Based Retrieval
by: Kashmira, Savini, et al.
Published: (2025)
by: Kashmira, Savini, et al.
Published: (2025)
Prompt Less, Smile More: MTP with Semantic Engineering in Lieu of Prompt Engineering
by: Dantanarayana, Jayanaka L., et al.
Published: (2025)
by: Dantanarayana, Jayanaka L., et al.
Published: (2025)
TOBUGraph: Knowledge Graph-Based Retrieval for Enhanced LLM Performance Beyond RAG
by: Kashmira, Savini, et al.
Published: (2024)
by: Kashmira, Savini, et al.
Published: (2024)
GraphMend: Code Transformations for Fixing Graph Breaks in PyTorch 2
by: Kashmira, Savini, et al.
Published: (2025)
by: Kashmira, Savini, et al.
Published: (2025)
GaLore$+$: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection
by: Liao, Xutao, et al.
Published: (2024)
by: Liao, Xutao, et al.
Published: (2024)
A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputs
by: Shelmanov, Artem, et al.
Published: (2025)
by: Shelmanov, Artem, et al.
Published: (2025)
RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty
by: Zhang, Ziqian, et al.
Published: (2026)
by: Zhang, Ziqian, et al.
Published: (2026)
Attention Mechanism and Heuristic Approach: Context-Aware File Ranking Using Multi-Head Self-Attention
by: Sharma, Pradeep Kumar, et al.
Published: (2026)
by: Sharma, Pradeep Kumar, et al.
Published: (2026)
LLM-RankFusion: Mitigating Intrinsic Inconsistency in LLM-based Ranking
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
One Agent Too Many: User Perspectives on Approaches to Multi-agent Conversational AI
by: Clarke, Christopher, et al.
Published: (2024)
by: Clarke, Christopher, et al.
Published: (2024)
JuStRank: Benchmarking LLM Judges for System Ranking
by: Gera, Ariel, et al.
Published: (2024)
by: Gera, Ariel, et al.
Published: (2024)
StealthRank: LLM Ranking Manipulation via Stealthy Prompt Optimization
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
Can OpenAI o1 Reason Well in Ophthalmology? A 6,990-Question Head-to-Head Evaluation Study
by: Srinivasan, Sahana, et al.
Published: (2025)
by: Srinivasan, Sahana, et al.
Published: (2025)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
by: Deshpande, Darshan, et al.
Published: (2024)
by: Deshpande, Darshan, et al.
Published: (2024)
Language Ranker: A Lightweight Ranking framework for LLM Decoding
by: Zhang, Chenheng, et al.
Published: (2025)
by: Zhang, Chenheng, et al.
Published: (2025)
PRILoRA: Pruned and Rank-Increasing Low-Rank Adaptation
by: Benedek, Nadav, et al.
Published: (2024)
by: Benedek, Nadav, et al.
Published: (2024)
MoR: Mixture of Ranks for Low-Rank Adaptation Tuning
by: Tang, Chuanyu, et al.
Published: (2024)
by: Tang, Chuanyu, et al.
Published: (2024)
Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning
by: Fu, Yu, et al.
Published: (2024)
by: Fu, Yu, et al.
Published: (2024)
Modeling LLM Agent Reviewer Dynamics in Elo-Ranked Review System
by: Huang, Hsiang-Wei, et al.
Published: (2026)
by: Huang, Hsiang-Wei, et al.
Published: (2026)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
by: Balkır, Esma, et al.
Published: (2026)
by: Balkır, Esma, et al.
Published: (2026)
FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inference
by: Lu, Yu-Chen, et al.
Published: (2025)
by: Lu, Yu-Chen, et al.
Published: (2025)
Rationale-Augmented Retrieval with Constrained LLM Re-Ranking for Task Discovery
by: Wei, Bowen
Published: (2025)
by: Wei, Bowen
Published: (2025)
Ranking LLMs by compression
by: Guo, Peijia, et al.
Published: (2024)
by: Guo, Peijia, et al.
Published: (2024)
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026)
by: Genadi, Rifo, et al.
Published: (2026)
On the Emergence of Induction Heads for In-Context Learning
by: Musat, Tiberiu, et al.
Published: (2025)
by: Musat, Tiberiu, et al.
Published: (2025)
Ranking LLM-Generated Loop Invariants for Program Verification
by: Chakraborty, Saikat, et al.
Published: (2023)
by: Chakraborty, Saikat, et al.
Published: (2023)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
by: Lin, Gang, et al.
Published: (2026)
by: Lin, Gang, et al.
Published: (2026)
Multi-RADS Synthetic Radiology Report Dataset and Head-to-Head Benchmarking of 41 Open-Weight and Proprietary Language Models
by: Bose, Kartik, et al.
Published: (2026)
by: Bose, Kartik, et al.
Published: (2026)
OlympicArena Medal Ranks: Who Is the Most Intelligent AI So Far?
by: Huang, Zhen, et al.
Published: (2024)
by: Huang, Zhen, et al.
Published: (2024)
Benchmarking Next-Generation Reasoning-Focused Large Language Models in Ophthalmology: A Head-to-Head Evaluation on 5,888 Items
by: Zou, Minjie, et al.
Published: (2025)
by: Zou, Minjie, et al.
Published: (2025)
Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
RankAdaptor: Hierarchical Rank Allocation for Efficient Fine-Tuning Pruned LLMs via Performance Model
by: Zhou, Changhai, et al.
Published: (2024)
by: Zhou, Changhai, et al.
Published: (2024)
Identifying Semantic Induction Heads to Understand In-Context Learning
by: Ren, Jie, et al.
Published: (2024)
by: Ren, Jie, et al.
Published: (2024)
Similar Items
-
SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models
by: Daynauth, Roland, et al.
Published: (2025) -
Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments
by: Daynauth, Roland, et al.
Published: (2024) -
Guylingo: The Republic of Guyana Creole Corpora
by: Clarke, Christopher, et al.
Published: (2024) -
MTP: A Meaning-Typed Language Abstraction for AI-Integrated Programming
by: Dantanarayana, Jayanaka L., et al.
Published: (2024) -
PEFT-U: Parameter-Efficient Fine-Tuning for User Personalization
by: Clarke, Christopher, et al.
Published: (2024)