When LLMs get significantly worse: A statistical approach to detect model degradations
Fuente:
arXiv
Saved in:
| Main Authors: | Kübler, Jonas, Budhathoki, Kailash, Kleindessner, Matthäus, Zhou, Xiong, Yin, Junming, Khetan, Ashish, Karypis, George |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inference Optimization of Foundation Models on AI Accelerators
by: Park, Youngsuk, et al.
Published: (2024)
by: Park, Youngsuk, et al.
Published: (2024)
Block-Diagonal LoRA for Eliminating Communication Overhead in Tensor Parallel LoRA Serving
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
A Proximal Operator for Inducing 2:4-Sparsity
by: Kübler, Jonas M, et al.
Published: (2025)
by: Kübler, Jonas M, et al.
Published: (2025)
LLM-Rank: A Graph Theoretical Approach to Pruning Large Language Models
by: Hoffmann, David, et al.
Published: (2024)
by: Hoffmann, David, et al.
Published: (2024)
XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
by: Vankov, Daniil, et al.
Published: (2026)
by: Vankov, Daniil, et al.
Published: (2026)
Meaningful Causal Aggregation and Paradoxical Confounding
by: Zhu, Yuchen, et al.
Published: (2023)
by: Zhu, Yuchen, et al.
Published: (2023)
SemPool: Simple, robust, and interpretable KG pooling for enhancing language models
by: Mavromatis, Costas, et al.
Published: (2024)
by: Mavromatis, Costas, et al.
Published: (2024)
P-EAGLE: Parallel-Drafting EAGLE with Scalable Training
by: Hui, Mude, et al.
Published: (2026)
by: Hui, Mude, et al.
Published: (2026)
DoWhy-GCM: An extension of DoWhy for causal inference in graphical causal models
by: Blöbaum, Patrick, et al.
Published: (2022)
by: Blöbaum, Patrick, et al.
Published: (2022)
PoliticsBench: Benchmarking Political Values in Large Language Models with Multi-Turn Roleplay
by: Khetan, Rohan, et al.
Published: (2026)
by: Khetan, Rohan, et al.
Published: (2026)
GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning
by: Mavromatis, Costas, et al.
Published: (2024)
by: Mavromatis, Costas, et al.
Published: (2024)
Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization
by: Mavromatis, Costas, et al.
Published: (2024)
by: Mavromatis, Costas, et al.
Published: (2024)
DEFT: Data Efficient Fine-Tuning for Pre-Trained Language Models via Unsupervised Core-Set Selection
by: Das, Devleena, et al.
Published: (2023)
by: Das, Devleena, et al.
Published: (2023)
Asia. The prime minister who needs things to get worse
Published: (2001)
Published: (2001)
Quantifying intrinsic causal contributions via structure preserving interventions
by: Janzing, Dominik, et al.
Published: (2020)
by: Janzing, Dominik, et al.
Published: (2020)
The Instruction Gap: LLMs get lost in Following Instruction
by: Tripathi, Vishesh, et al.
Published: (2025)
by: Tripathi, Vishesh, et al.
Published: (2025)
Do as I can, not as I get
by: Zheng, Shangfei, et al.
Published: (2023)
by: Zheng, Shangfei, et al.
Published: (2023)
Education distillation:getting student models to learn in shcools
by: Feng, Ling, et al.
Published: (2023)
by: Feng, Ling, et al.
Published: (2023)
PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models
by: Friedland, Gerald, et al.
Published: (2024)
by: Friedland, Gerald, et al.
Published: (2024)
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
by: Wang, Keyu, et al.
Published: (2025)
by: Wang, Keyu, et al.
Published: (2025)
Application of LLMs to Multi-Robot Path Planning and Task Allocation
by: Kumar, Ashish
Published: (2025)
by: Kumar, Ashish
Published: (2025)
Can LLMs get help from other LLMs without revealing private information?
by: Hartmann, Florian, et al.
Published: (2024)
by: Hartmann, Florian, et al.
Published: (2024)
Extending Input Contexts of Language Models through Training on Segmented Sequences
by: Karypis, Petros, et al.
Published: (2023)
by: Karypis, Petros, et al.
Published: (2023)
Enter the Void - Planning to Seek Entropy When Reward is Scarce
by: Sundar, Ashish, et al.
Published: (2025)
by: Sundar, Ashish, et al.
Published: (2025)
From Demonstrations to Rewards: Alignment Without Explicit Human Preferences
by: Zeng, Siliang, et al.
Published: (2025)
by: Zeng, Siliang, et al.
Published: (2025)
AutoGluon-Multimodal (AutoMM): Supercharging Multimodal AutoML with Foundation Models
by: Tang, Zhiqiang, et al.
Published: (2024)
by: Tang, Zhiqiang, et al.
Published: (2024)
When Gender is Hard to See: Multi-Attribute Support for Long-Range Recognition
by: Mbongo, Nzakiese, et al.
Published: (2025)
by: Mbongo, Nzakiese, et al.
Published: (2025)
SEAL: Suite for Evaluating API-use of LLMs
by: Kim, Woojeong, et al.
Published: (2024)
by: Kim, Woojeong, et al.
Published: (2024)
Multimodal Chain-of-Thought Reasoning in Language Models
by: Zhang, Zhuosheng, et al.
Published: (2023)
by: Zhang, Zhuosheng, et al.
Published: (2023)
AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
by: Yang, Ke, et al.
Published: (2024)
by: Yang, Ke, et al.
Published: (2024)
When Fuzzing Meets LLMs: Challenges and Opportunities
by: Jiang, Yu, et al.
Published: (2024)
by: Jiang, Yu, et al.
Published: (2024)
When Does Multimodality Lead to Better Time Series Forecasting?
by: Zhang, Xiyuan, et al.
Published: (2025)
by: Zhang, Xiyuan, et al.
Published: (2025)
Centralized vs Decentralized Federated Learning: A trade-off performance analysis
by: Medjadji, Chaimaa, et al.
Published: (2026)
by: Medjadji, Chaimaa, et al.
Published: (2026)
Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?
by: Gomes, Gonçalo, et al.
Published: (2025)
by: Gomes, Gonçalo, et al.
Published: (2025)
From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives?
by: Amin, Hasan, et al.
Published: (2026)
by: Amin, Hasan, et al.
Published: (2026)
Automated Consistency Analysis of LLMs
by: Patwardhan, Aditya, et al.
Published: (2025)
by: Patwardhan, Aditya, et al.
Published: (2025)
Test Time Training for AC Power Flow Surrogates via Physics and Operational Constraint Refinement
by: Dogoulis, Panteleimon, et al.
Published: (2025)
by: Dogoulis, Panteleimon, et al.
Published: (2025)
Human-AI Collaborative Taxonomy Construction: A Case Study in Profession-Specific Writing Assistants
by: Lee, Minhwa, et al.
Published: (2024)
by: Lee, Minhwa, et al.
Published: (2024)
When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
by: Nakshatri, Nishanth Sridhar, et al.
Published: (2025)
by: Nakshatri, Nishanth Sridhar, et al.
Published: (2025)
From Bits to Boardrooms: A Cutting-Edge Multi-Agent LLM Framework for Business Excellence
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
Similar Items
-
Inference Optimization of Foundation Models on AI Accelerators
by: Park, Youngsuk, et al.
Published: (2024) -
Block-Diagonal LoRA for Eliminating Communication Overhead in Tensor Parallel LoRA Serving
by: Wang, Xinyu, et al.
Published: (2025) -
A Proximal Operator for Inducing 2:4-Sparsity
by: Kübler, Jonas M, et al.
Published: (2025) -
LLM-Rank: A Graph Theoretical Approach to Pruning Large Language Models
by: Hoffmann, David, et al.
Published: (2024) -
XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
by: Vankov, Daniil, et al.
Published: (2026)