SimBA: Simplifying Benchmark Analysis Using Performance Matrices Alone
Fuente:
arXiv
Saved in:
| Main Authors: | Subramani, Nishant, Gomez, Alfredo, Diab, Mona |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Personal Information Parroting in Language Models
by: Subramani, Nishant, et al.
Published: (2026)
by: Subramani, Nishant, et al.
Published: (2026)
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
by: Muhamed, Aashiq, et al.
Published: (2024)
by: Muhamed, Aashiq, et al.
Published: (2024)
Emotion Classification in Low and Moderate Resource Languages
by: Tafreshi, Shabnam, et al.
Published: (2024)
by: Tafreshi, Shabnam, et al.
Published: (2024)
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
by: Wedgwood, James, et al.
Published: (2026)
by: Wedgwood, James, et al.
Published: (2026)
CoRAG: Collaborative Retrieval-Augmented Generation
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
by: Subramani, Nishant, et al.
Published: (2025)
by: Subramani, Nishant, et al.
Published: (2025)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Generative Value Conflicts Reveal LLM Priorities
by: Liu, Andy, et al.
Published: (2025)
by: Liu, Andy, et al.
Published: (2025)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
by: Song, Xiangchen, et al.
Published: (2025)
by: Song, Xiangchen, et al.
Published: (2025)
Can Large Language Models Infer Causation from Correlation?
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
Analyzing the Role of Semantic Representations in the Era of Large Language Models
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
by: Haimes, Jacob, et al.
Published: (2024)
by: Haimes, Jacob, et al.
Published: (2024)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
by: Li, Michael, et al.
Published: (2025)
by: Li, Michael, et al.
Published: (2025)
Simplifying Outcomes of Language Model Component Analyses with ELIA
by: Eidt, Aaron Louis, et al.
Published: (2026)
by: Eidt, Aaron Louis, et al.
Published: (2026)
MoBA: Mixture of Block Attention for Long-Context LLMs
by: Lu, Enzhe, et al.
Published: (2025)
by: Lu, Enzhe, et al.
Published: (2025)
How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
Fact or Fiction? Improving Fact Verification with Knowledge Graphs through Simplified Subgraph Retrievals
by: Opsahl, Tobias A.
Published: (2024)
by: Opsahl, Tobias A.
Published: (2024)
FutureSim: Replaying World Events to Evaluate Adaptive Agents
by: Goel, Shashwat, et al.
Published: (2026)
by: Goel, Shashwat, et al.
Published: (2026)
2-Tier SimCSE: Elevating BERT for Robust Sentence Embeddings
by: Wang, Yumeng, et al.
Published: (2025)
by: Wang, Yumeng, et al.
Published: (2025)
LoRA-Mini : Adaptation Matrices Decomposition and Selective Training
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
LABBench2: An Improved Benchmark for AI Systems Performing Biology Research
by: Laurent, Jon M, et al.
Published: (2026)
by: Laurent, Jon M, et al.
Published: (2026)
Assessing the Portability of Parameter Matrices Trained by Parameter-Efficient Finetuning Methods
by: Sabry, Mohammed, et al.
Published: (2024)
by: Sabry, Mohammed, et al.
Published: (2024)
A Single Linear Layer Yields Task-Adapted Low-Rank Matrices
by: Kim, Hwichan, et al.
Published: (2024)
by: Kim, Hwichan, et al.
Published: (2024)
Wonderful Matrices: Combining for a More Efficient and Effective Foundation Model Architecture
by: Shi, Jingze, et al.
Published: (2024)
by: Shi, Jingze, et al.
Published: (2024)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
by: Liu, Yijun, et al.
Published: (2024)
by: Liu, Yijun, et al.
Published: (2024)
LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
by: Liu, Jiarui, et al.
Published: (2025)
by: Liu, Jiarui, et al.
Published: (2025)
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
by: Cho, Yoonjun, et al.
Published: (2025)
by: Cho, Yoonjun, et al.
Published: (2025)
Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction
by: Tsaknakis, Ioannis, et al.
Published: (2025)
by: Tsaknakis, Ioannis, et al.
Published: (2025)
FanChuan: A Multilingual and Graph-Structured Benchmark For Parody Detection and Analysis
by: Zheng, Yilun, et al.
Published: (2025)
by: Zheng, Yilun, et al.
Published: (2025)
Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
by: Bennion, Jonathan, et al.
Published: (2025)
by: Bennion, Jonathan, et al.
Published: (2025)
StressRoBERTa: Cross-Condition Transfer Learning from Depression, Anxiety, and PTSD to Stress Detection
by: Alqahtani, Amal, et al.
Published: (2025)
by: Alqahtani, Amal, et al.
Published: (2025)
Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments
by: Alhanai, Tuka, et al.
Published: (2024)
by: Alhanai, Tuka, et al.
Published: (2024)
Benchmarking Benchmark Leakage in Large Language Models
by: Xu, Ruijie, et al.
Published: (2024)
by: Xu, Ruijie, et al.
Published: (2024)
TDBench: A Benchmark for Top-Down Image Understanding with Reliability Analysis of Vision-Language Models
by: Hou, Kaiyuan, et al.
Published: (2025)
by: Hou, Kaiyuan, et al.
Published: (2025)
Optimizing Retrieval-Augmented Generation: Analysis of Hyperparameter Impact on Performance and Efficiency
by: Ammar, Adel, et al.
Published: (2025)
by: Ammar, Adel, et al.
Published: (2025)
Interactive Benchmarks
by: Yue, Baoqing, et al.
Published: (2026)
by: Yue, Baoqing, et al.
Published: (2026)
Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification
by: Adeseye, Aisvarya, et al.
Published: (2026)
by: Adeseye, Aisvarya, et al.
Published: (2026)
Student sentiment Analysis Using Classification With Feature Extraction Techniques
by: Tamrakar, Latika, et al.
Published: (2021)
by: Tamrakar, Latika, et al.
Published: (2021)
Similar Items
-
Personal Information Parroting in Language Models
by: Subramani, Nishant, et al.
Published: (2026) -
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
by: Muhamed, Aashiq, et al.
Published: (2024) -
Emotion Classification in Low and Moderate Resource Languages
by: Tafreshi, Shabnam, et al.
Published: (2024) -
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
by: Wedgwood, James, et al.
Published: (2026) -
CoRAG: Collaborative Retrieval-Augmented Generation
by: Muhamed, Aashiq, et al.
Published: (2025)