MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Park, Ji-jun, Choi, Soo-joon, Jeong, Jiwon, Yoon, Taeyang, Lee, Ju-Wan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLMs for Enhanced Agricultural Meteorological Recommendations
di: Park, Ji-jun, et al.
Pubblicazione: (2024)
di: Park, Ji-jun, et al.
Pubblicazione: (2024)
Harnessing Generative LLMs for Enhanced Financial Event Entity Extraction Performance
di: Choi, Soo-joon, et al.
Pubblicazione: (2025)
di: Choi, Soo-joon, et al.
Pubblicazione: (2025)
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives
di: Park, Ji-jun, et al.
Pubblicazione: (2024)
di: Park, Ji-jun, et al.
Pubblicazione: (2024)
ContextualLVLM-Agent: A Holistic Framework for Multi-Turn Visually-Grounded Dialogue and Complex Instruction Following
di: Han, Seungmin, et al.
Pubblicazione: (2025)
di: Han, Seungmin, et al.
Pubblicazione: (2025)
Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
di: Hakimi, Ahmad Dawar, et al.
Pubblicazione: (2025)
di: Hakimi, Ahmad Dawar, et al.
Pubblicazione: (2025)
Large Language Models Are Better Logical Fallacy Reasoners with Counterargument, Explanation, and Goal-Aware Prompt Formulation
di: Jeong, Jiwon, et al.
Pubblicazione: (2025)
di: Jeong, Jiwon, et al.
Pubblicazione: (2025)
MechIR: A Mechanistic Interpretability Framework for Information Retrieval
di: Parry, Andrew, et al.
Pubblicazione: (2025)
di: Parry, Andrew, et al.
Pubblicazione: (2025)
Eliciting Latent Knowledge from Quirky Language Models
di: Mallen, Alex, et al.
Pubblicazione: (2023)
di: Mallen, Alex, et al.
Pubblicazione: (2023)
ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains
di: Park, Yein, et al.
Pubblicazione: (2024)
di: Park, Yein, et al.
Pubblicazione: (2024)
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
di: Jeon, Hyesung, et al.
Pubblicazione: (2024)
di: Jeon, Hyesung, et al.
Pubblicazione: (2024)
Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization
di: Zhang, Weixu, et al.
Pubblicazione: (2026)
di: Zhang, Weixu, et al.
Pubblicazione: (2026)
OpenLearnLM Benchmark: A Unified Framework for Evaluating Knowledge, Skill, and Attitude in Educational Large Language Models
di: Lee, Unggi, et al.
Pubblicazione: (2026)
di: Lee, Unggi, et al.
Pubblicazione: (2026)
Mechanistic Interpretability of Emotion Inference in Large Language Models
di: Tak, Ala N., et al.
Pubblicazione: (2025)
di: Tak, Ala N., et al.
Pubblicazione: (2025)
How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective
di: Liu, Qi, et al.
Pubblicazione: (2025)
di: Liu, Qi, et al.
Pubblicazione: (2025)
STEAM: A Semantic-Level Knowledge Editing Framework for Large Language Models
di: Jeong, Geunyeong, et al.
Pubblicazione: (2025)
di: Jeong, Geunyeong, et al.
Pubblicazione: (2025)
Pragmatic Competence Evaluation of Large Language Models for the Korean Language
di: Park, Dojun, et al.
Pubblicazione: (2024)
di: Park, Dojun, et al.
Pubblicazione: (2024)
Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Language Models
di: Kang, Minki, et al.
Pubblicazione: (2024)
di: Kang, Minki, et al.
Pubblicazione: (2024)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
Mechanistic Circuit-Based Knowledge Editing in Large Language Models
di: Zhao, Tianyi, et al.
Pubblicazione: (2026)
di: Zhao, Tianyi, et al.
Pubblicazione: (2026)
FLEX: Expert-level False-Less EXecution Metric for Reliable Text-to-SQL Benchmark
di: Kim, Heegyu, et al.
Pubblicazione: (2024)
di: Kim, Heegyu, et al.
Pubblicazione: (2024)
M4CXR: Exploring Multi-task Potentials of Multi-modal Large Language Models for Chest X-ray Interpretation
di: Park, Jonggwon, et al.
Pubblicazione: (2024)
di: Park, Jonggwon, et al.
Pubblicazione: (2024)
PICLe: Eliciting Diverse Behaviors from Large Language Models with Persona In-Context Learning
di: Choi, Hyeong Kyu, et al.
Pubblicazione: (2024)
di: Choi, Hyeong Kyu, et al.
Pubblicazione: (2024)
Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions
di: Naseem, Usman
Pubblicazione: (2026)
di: Naseem, Usman
Pubblicazione: (2026)
LINKED: Eliciting, Filtering and Integrating Knowledge in Large Language Model for Commonsense Reasoning
di: Li, Jiachun, et al.
Pubblicazione: (2024)
di: Li, Jiachun, et al.
Pubblicazione: (2024)
Adaptive Elicitation of Latent Information Using Natural Language
di: Wang, Jimmy, et al.
Pubblicazione: (2025)
di: Wang, Jimmy, et al.
Pubblicazione: (2025)
Contrastive Learning of English Language and Crystal Graphs for Multimodal Representation of Materials Knowledge
di: Park, Yang Jeong, et al.
Pubblicazione: (2025)
di: Park, Yang Jeong, et al.
Pubblicazione: (2025)
Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective
di: Lee, Jae Hee, et al.
Pubblicazione: (2025)
di: Lee, Jae Hee, et al.
Pubblicazione: (2025)
Knowledge Synthesis of Photosynthesis Research Using a Large Language Model
di: Yoon, Seungri, et al.
Pubblicazione: (2025)
di: Yoon, Seungri, et al.
Pubblicazione: (2025)
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models
di: Panpatil, Siddhant, et al.
Pubblicazione: (2025)
di: Panpatil, Siddhant, et al.
Pubblicazione: (2025)
Chain-of-Conceptual-Thought Elicits Daily Conversation in Large Language Models
di: Gu, Qingqing, et al.
Pubblicazione: (2025)
di: Gu, Qingqing, et al.
Pubblicazione: (2025)
Knowledge Integration Decay in Search-Augmented Reasoning of Large Language Models
di: Yu, Sangwon, et al.
Pubblicazione: (2026)
di: Yu, Sangwon, et al.
Pubblicazione: (2026)
Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models
di: Men, Tianyi, et al.
Pubblicazione: (2024)
di: Men, Tianyi, et al.
Pubblicazione: (2024)
Read Like a Radiologist: Efficient Vision-Language Model for 3D Medical Imaging Interpretation
di: Lee, Changsun, et al.
Pubblicazione: (2024)
di: Lee, Changsun, et al.
Pubblicazione: (2024)
Soft Inductive Bias Approach via Explicit Reasoning Perspectives in Inappropriate Utterance Detection Using Large Language Models
di: Kim, Ju-Young, et al.
Pubblicazione: (2025)
di: Kim, Ju-Young, et al.
Pubblicazione: (2025)
SimLens for Early Exit in Large Language Models: Eliciting Accurate Latent Predictions with One More Token
di: Ma, Ming, et al.
Pubblicazione: (2025)
di: Ma, Ming, et al.
Pubblicazione: (2025)
Language Models Benefit from Preparation with Elicited Knowledge
di: Yu, Jiacan, et al.
Pubblicazione: (2024)
di: Yu, Jiacan, et al.
Pubblicazione: (2024)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2026)
di: Wang, Xu, et al.
Pubblicazione: (2026)
MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability
di: Khadka, Barsat
Pubblicazione: (2026)
di: Khadka, Barsat
Pubblicazione: (2026)
CIRF: Tokenizing Chain-of-Thoughts into Reusable Functional Units for Efficient Latent Reasoning in Large Language Models
di: Lee, Yukyung, et al.
Pubblicazione: (2026)
di: Lee, Yukyung, et al.
Pubblicazione: (2026)
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
di: Zhang, Hengyuan, et al.
Pubblicazione: (2026)
di: Zhang, Hengyuan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
LLMs for Enhanced Agricultural Meteorological Recommendations
di: Park, Ji-jun, et al.
Pubblicazione: (2024) -
Harnessing Generative LLMs for Enhanced Financial Event Entity Extraction Performance
di: Choi, Soo-joon, et al.
Pubblicazione: (2025) -
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives
di: Park, Ji-jun, et al.
Pubblicazione: (2024) -
ContextualLVLM-Agent: A Holistic Framework for Multi-Turn Visually-Grounded Dialogue and Complex Instruction Following
di: Han, Seungmin, et al.
Pubblicazione: (2025) -
Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
di: Hakimi, Ahmad Dawar, et al.
Pubblicazione: (2025)