Simplifying Outcomes of Language Model Component Analyses with ELIA
Fuente:
arXiv
Guardado en:
| Autores principales: | Eidt, Aaron Louis, Feldhus, Nils |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Interpreting Language Models Through Concept Descriptions: A Survey
por: Feldhus, Nils, et al.
Publicado: (2025)
por: Feldhus, Nils, et al.
Publicado: (2025)
LLMCheckup: Conversational Examination of Large Language Models via Interpretability Tools and Self-Explanations
por: Wang, Qianli, et al.
Publicado: (2024)
por: Wang, Qianli, et al.
Publicado: (2024)
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
por: Dhaini, Mahdi, et al.
Publicado: (2025)
por: Dhaini, Mahdi, et al.
Publicado: (2025)
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
por: Wang, Qianli, et al.
Publicado: (2026)
por: Wang, Qianli, et al.
Publicado: (2026)
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
por: Sun, Jingyi, et al.
Publicado: (2026)
por: Sun, Jingyi, et al.
Publicado: (2026)
Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
por: Kopf, Laura, et al.
Publicado: (2025)
por: Kopf, Laura, et al.
Publicado: (2025)
Functional Component Ablation Reveals Specialization Patterns in Hybrid Language Model Architectures
por: Borobia, Hector, et al.
Publicado: (2026)
por: Borobia, Hector, et al.
Publicado: (2026)
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
por: Thareja, Rushil, et al.
Publicado: (2025)
por: Thareja, Rushil, et al.
Publicado: (2025)
SimBA: Simplifying Benchmark Analysis Using Performance Matrices Alone
por: Subramani, Nishant, et al.
Publicado: (2025)
por: Subramani, Nishant, et al.
Publicado: (2025)
Specialising and Analysing Instruction-Tuned and Byte-Level Language Models for Organic Reaction Prediction
por: Pang, Jiayun, et al.
Publicado: (2024)
por: Pang, Jiayun, et al.
Publicado: (2024)
Tracing Uncertainty in Language Model "Reasoning"
por: Grünefeld, Nils, et al.
Publicado: (2026)
por: Grünefeld, Nils, et al.
Publicado: (2026)
Fact or Fiction? Improving Fact Verification with Knowledge Graphs through Simplified Subgraph Retrievals
por: Opsahl, Tobias A.
Publicado: (2024)
por: Opsahl, Tobias A.
Publicado: (2024)
iFlip: Iterative Feedback-driven Counterfactual Example Refinement
por: Wang, Yilong, et al.
Publicado: (2026)
por: Wang, Yilong, et al.
Publicado: (2026)
ED-Copilot: Reduce Emergency Department Wait Time with Language Model Diagnostic Assistance
por: Sun, Liwen, et al.
Publicado: (2024)
por: Sun, Liwen, et al.
Publicado: (2024)
Simple and Effective Masked Diffusion Language Models
por: Sahoo, Subham Sekhar, et al.
Publicado: (2024)
por: Sahoo, Subham Sekhar, et al.
Publicado: (2024)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
por: Noukhovitch, Michael, et al.
Publicado: (2024)
por: Noukhovitch, Michael, et al.
Publicado: (2024)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
por: Marks, Samuel, et al.
Publicado: (2024)
por: Marks, Samuel, et al.
Publicado: (2024)
ProofOptimizer: Training Language Models to Simplify Proofs without Human Demonstrations
por: Gu, Alex, et al.
Publicado: (2025)
por: Gu, Alex, et al.
Publicado: (2025)
Exploring and Benchmarking the Planning Capabilities of Large Language Models
por: Bohnet, Bernd, et al.
Publicado: (2024)
por: Bohnet, Bernd, et al.
Publicado: (2024)
Graph of Thoughts: Solving Elaborate Problems with Large Language Models
por: Besta, Maciej, et al.
Publicado: (2023)
por: Besta, Maciej, et al.
Publicado: (2023)
Large Language Models Predict Functional Outcomes after Acute Ischemic Stroke
por: Kapoor, Anjali K., et al.
Publicado: (2026)
por: Kapoor, Anjali K., et al.
Publicado: (2026)
Small Models Are (Still) Effective Cross-Domain Argument Extractors
por: Gantt, William, et al.
Publicado: (2024)
por: Gantt, William, et al.
Publicado: (2024)
LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
por: Zou, Kaijian, et al.
Publicado: (2025)
por: Zou, Kaijian, et al.
Publicado: (2025)
Evaluating Language-Model Agents on Realistic Autonomous Tasks
por: Kinniment, Megan, et al.
Publicado: (2023)
por: Kinniment, Megan, et al.
Publicado: (2023)
Judge Circuits
por: Feldhus, Nils, et al.
Publicado: (2026)
por: Feldhus, Nils, et al.
Publicado: (2026)
The Factuality of Large Language Models in the Legal Domain
por: Hamdani, Rajaa El, et al.
Publicado: (2024)
por: Hamdani, Rajaa El, et al.
Publicado: (2024)
Peering Inside the Black Box: Uncovering LLM Errors in Optimization Modelling through Component-Level Evaluation
por: Refai, Dania, et al.
Publicado: (2025)
por: Refai, Dania, et al.
Publicado: (2025)
Classification of User Reports for Detection of Faulty Computer Components using NLP Models: A Case Study
por: Silva, Maria de Lourdes M., et al.
Publicado: (2025)
por: Silva, Maria de Lourdes M., et al.
Publicado: (2025)
Language Model Memory and Memory Models for Language
por: Badger, Benjamin L.
Publicado: (2026)
por: Badger, Benjamin L.
Publicado: (2026)
Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
por: Albalak, Alon, et al.
Publicado: (2025)
por: Albalak, Alon, et al.
Publicado: (2025)
CovidLLM: A Robust Large Language Model with Missing Value Adaptation and Multi-Objective Learning Strategy for Predicting Disease Severity and Clinical Outcomes in COVID-19 Patients
por: Zhu, Shengjun, et al.
Publicado: (2024)
por: Zhu, Shengjun, et al.
Publicado: (2024)
CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities
por: Mao, Yujun, et al.
Publicado: (2024)
por: Mao, Yujun, et al.
Publicado: (2024)
Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
por: Ahmad, Areeb, et al.
Publicado: (2025)
por: Ahmad, Areeb, et al.
Publicado: (2025)
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
por: Benfeghoul, Martin, et al.
Publicado: (2025)
por: Benfeghoul, Martin, et al.
Publicado: (2025)
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
por: Tsujimura, Hikaru, et al.
Publicado: (2025)
por: Tsujimura, Hikaru, et al.
Publicado: (2025)
Training Language Models with Language Feedback at Scale
por: Scheurer, Jérémy, et al.
Publicado: (2023)
por: Scheurer, Jérémy, et al.
Publicado: (2023)
Adaptive Selection of LoRA Components in Privacy-Preserving Federated Learning
por: Kim, Myoungjun, et al.
Publicado: (2026)
por: Kim, Myoungjun, et al.
Publicado: (2026)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
por: Su, Jingtong, et al.
Publicado: (2025)
por: Su, Jingtong, et al.
Publicado: (2025)
Automated Composition of Agents: A Knapsack Approach for Agentic Component Selection
por: Yuan, Michelle, et al.
Publicado: (2025)
por: Yuan, Michelle, et al.
Publicado: (2025)
An Isotropic Approach to Efficient Uncertainty Quantification with Gradient Norms
por: Grünefeld, Nils, et al.
Publicado: (2026)
por: Grünefeld, Nils, et al.
Publicado: (2026)
Ejemplares similares
-
Interpreting Language Models Through Concept Descriptions: A Survey
por: Feldhus, Nils, et al.
Publicado: (2025) -
LLMCheckup: Conversational Examination of Large Language Models via Interpretability Tools and Self-Explanations
por: Wang, Qianli, et al.
Publicado: (2024) -
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
por: Dhaini, Mahdi, et al.
Publicado: (2025) -
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
por: Wang, Qianli, et al.
Publicado: (2026) -
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
por: Sun, Jingyi, et al.
Publicado: (2026)