A Statistical Framework for Ranking LLM-Based Chatbots
Fuente:
arXiv
Saved in:
| Main Authors: | Ameli, Siavash, Zhuang, Siyuan, Stoica, Ion, Mahoney, Michael W. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
MPC-Minimized Secure LLM Inference
by: Rathee, Deevashwer, et al.
Published: (2024)
by: Rathee, Deevashwer, et al.
Published: (2024)
M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
by: Mishra, Mayank, et al.
Published: (2026)
by: Mishra, Mayank, et al.
Published: (2026)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
by: Park, Jongseok, et al.
Published: (2026)
by: Park, Jongseok, et al.
Published: (2026)
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
by: Guo, Wentao, et al.
Published: (2025)
by: Guo, Wentao, et al.
Published: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
by: Jiang, Xuanlin, et al.
Published: (2024)
by: Jiang, Xuanlin, et al.
Published: (2024)
HashAttention: Semantic Sparsity for Faster Inference
by: Desai, Aditya, et al.
Published: (2024)
by: Desai, Aditya, et al.
Published: (2024)
RouteLLM: Learning to Route LLMs with Preference Data
by: Ong, Isaac, et al.
Published: (2024)
by: Ong, Isaac, et al.
Published: (2024)
Post-Training Sparse Attention with Double Sparsity
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
CLUTR: Curriculum Learning via Unsupervised Task Representation Learning
by: Azad, Abdus Salam, et al.
Published: (2022)
by: Azad, Abdus Salam, et al.
Published: (2022)
SageBwd: A Trainable Low-bit Attention
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
TRAP: Tail-aware Ranking Attack for World-Model Planning
by: Duan, Siyuan, et al.
Published: (2026)
by: Duan, Siyuan, et al.
Published: (2026)
MatterChat: A Multi-Modal LLM for Material Science
by: Tang, Yingheng, et al.
Published: (2025)
by: Tang, Yingheng, et al.
Published: (2025)
Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and Compressibility
by: Yu, Annan, et al.
Published: (2025)
by: Yu, Annan, et al.
Published: (2025)
Recency Biased Causal Attention for Time-series Forecasting
by: Hegazy, Kareem, et al.
Published: (2025)
by: Hegazy, Kareem, et al.
Published: (2025)
Adaptive Repetition for Mitigating Position Bias in LLM-Based Ranking
by: Vardasbi, Ali, et al.
Published: (2025)
by: Vardasbi, Ali, et al.
Published: (2025)
Spectral Estimation with Free Decompression
by: Ameli, Siavash, et al.
Published: (2025)
by: Ameli, Siavash, et al.
Published: (2025)
Free Decompression with Algebraic Spectral Curves
by: Ameli, Siavash, et al.
Published: (2026)
by: Ameli, Siavash, et al.
Published: (2026)
depyf: Open the Opaque Box of PyTorch Compiler for Machine Learning Researchers
by: You, Kaichao, et al.
Published: (2024)
by: You, Kaichao, et al.
Published: (2024)
Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
by: Chen, Lingjiao, et al.
Published: (2024)
by: Chen, Lingjiao, et al.
Published: (2024)
PLAN: Proactive Low-Rank Allocation for Continual Learning
by: Wang, Xiequn, et al.
Published: (2025)
by: Wang, Xiequn, et al.
Published: (2025)
Trustless Audits without Revealing Data or Models
by: Waiwitlikhit, Suppakit, et al.
Published: (2024)
by: Waiwitlikhit, Suppakit, et al.
Published: (2024)
FLEX: A Backbone for Diffusion-Based Modeling of Spatio-temporal Physical Systems
by: Erichson, N. Benjamin, et al.
Published: (2025)
by: Erichson, N. Benjamin, et al.
Published: (2025)
BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation
by: Zhu, Alan, et al.
Published: (2025)
by: Zhu, Alan, et al.
Published: (2025)
Inference Time Context Sparsity: Illusion or Opportunity?
by: Joshi, Sahil, et al.
Published: (2026)
by: Joshi, Sahil, et al.
Published: (2026)
am-ELO: A Stable Framework for Arena-based LLM Evaluation
by: Liu, Zirui, et al.
Published: (2025)
by: Liu, Zirui, et al.
Published: (2025)
Detecting labeling bias using influence functions
by: Jørgensen, Frida, et al.
Published: (2026)
by: Jørgensen, Frida, et al.
Published: (2026)
Joint Embeddings Go Temporal
by: Ennadir, Sofiane, et al.
Published: (2025)
by: Ennadir, Sofiane, et al.
Published: (2025)
Improving Your Model Ranking on Chatbot Arena by Vote Rigging
by: Min, Rui, et al.
Published: (2025)
by: Min, Rui, et al.
Published: (2025)
From Optimization to Prediction: Transformer-Based Path-Flow Estimation to the Traffic Assignment Problem
by: Ameli, Mostafa, et al.
Published: (2025)
by: Ameli, Mostafa, et al.
Published: (2025)
Online Speculative Decoding
by: Liu, Xiaoxuan, et al.
Published: (2023)
by: Liu, Xiaoxuan, et al.
Published: (2023)
vAttention: Verified Sparse Attention
by: Desai, Aditya, et al.
Published: (2025)
by: Desai, Aditya, et al.
Published: (2025)
S*: Test Time Scaling for Code Generation
by: Li, Dacheng, et al.
Published: (2025)
by: Li, Dacheng, et al.
Published: (2025)
AI for Water Sustainability: Global Water Quality Assessment and Prediction with Explainable AI with LLM Chatbot for Insights
by: Paneru, Biplov, et al.
Published: (2024)
by: Paneru, Biplov, et al.
Published: (2024)
Out-of-Distribution Detection using Counterfactual Distance
by: Stoica, Maria, et al.
Published: (2025)
by: Stoica, Maria, et al.
Published: (2025)
Fairness in Serving Large Language Models
by: Sheng, Ying, et al.
Published: (2023)
by: Sheng, Ying, et al.
Published: (2023)
Autellix: An Efficient Serving Engine for LLM Agents as General Programs
by: Luo, Michael, et al.
Published: (2025)
by: Luo, Michael, et al.
Published: (2025)
ChatDiet: Empowering Personalized Nutrition-Oriented Food Recommender Chatbots through an LLM-Augmented Framework
by: Yang, Zhongqi, et al.
Published: (2024)
by: Yang, Zhongqi, et al.
Published: (2024)
asanAI: In-Browser, No-Code, Offline-First Machine Learning Toolkit
by: Koch, Norman, et al.
Published: (2025)
by: Koch, Norman, et al.
Published: (2025)
Margin-Adaptive Confidence Ranking for Reliable LLM Judgement
by: Jin, Gaojie, et al.
Published: (2026)
by: Jin, Gaojie, et al.
Published: (2026)
Similar Items
-
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024) -
MPC-Minimized Secure LLM Inference
by: Rathee, Deevashwer, et al.
Published: (2024) -
M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
by: Mishra, Mayank, et al.
Published: (2026) -
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
by: Park, Jongseok, et al.
Published: (2026) -
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
by: Guo, Wentao, et al.
Published: (2025)