Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
Fuente:
arXiv
Guardado en:
| Autores principales: | Nainani, Jatin, Vaidyanathan, Sankaran, Yeung, AJ, Gupta, Kartik, Jensen, David |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
por: Herbst, Jeremy, et al.
Publicado: (2026)
por: Herbst, Jeremy, et al.
Publicado: (2026)
Evaluating Brain-Inspired Modular Training in Automated Circuit Discovery for Mechanistic Interpretability
por: Nainani, Jatin
Publicado: (2024)
por: Nainani, Jatin
Publicado: (2024)
Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
por: Gao, Yutong, et al.
Publicado: (2026)
por: Gao, Yutong, et al.
Publicado: (2026)
Do Models Know Why They Changed Their Mind? Interpretability and Faithfulness of Chain-of-Thought Under Knowledge Conflict
por: Venkata, Pruthvinath Jeripity
Publicado: (2026)
por: Venkata, Pruthvinath Jeripity
Publicado: (2026)
A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
por: Rai, Daking, et al.
Publicado: (2024)
por: Rai, Daking, et al.
Publicado: (2024)
Prototype Transformer: Towards Language Model Architectures Interpretable by Design
por: Yordanov, Yordan, et al.
Publicado: (2026)
por: Yordanov, Yordan, et al.
Publicado: (2026)
One Supervisor, Many Modalities: Adaptive Tool Orchestration for Autonomous Queries
por: Saini, Mayank, et al.
Publicado: (2026)
por: Saini, Mayank, et al.
Publicado: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
por: Oketunji, Abiodun Finbarrs, et al.
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs, et al.
Publicado: (2023)
FETILDA: An Effective Framework For Fin-tuned Embeddings For Long Financial Text Documents
por: Xia, Bolun "Namir", et al.
Publicado: (2022)
por: Xia, Bolun "Namir", et al.
Publicado: (2022)
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
por: Vieira, Inês, et al.
Publicado: (2026)
por: Vieira, Inês, et al.
Publicado: (2026)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
por: Cho, Seonglae, et al.
Publicado: (2026)
por: Cho, Seonglae, et al.
Publicado: (2026)
Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text
por: Zhou, Tianyang, et al.
Publicado: (2026)
por: Zhou, Tianyang, et al.
Publicado: (2026)
Random Rule Forest (RRF): Interpretable Ensembles of LLM-Generated Questions for Predicting Startup Success
por: Griffin, Ben, et al.
Publicado: (2025)
por: Griffin, Ben, et al.
Publicado: (2025)
The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior
por: Larsen, Erik
Publicado: (2025)
por: Larsen, Erik
Publicado: (2025)
PlanRAG: A Plan-then-Retrieval Augmented Generation for Generative Large Language Models as Decision Makers
por: Lee, Myeonghwa, et al.
Publicado: (2024)
por: Lee, Myeonghwa, et al.
Publicado: (2024)
Danoliteracy of Generative Large Language Models
por: Holm, Søren Vejlgaard, et al.
Publicado: (2024)
por: Holm, Søren Vejlgaard, et al.
Publicado: (2024)
CircuitProbe: Predicting Reasoning Circuits in Transformers via Stability Zone Detection
por: Panuganti, Rajkiran
Publicado: (2026)
por: Panuganti, Rajkiran
Publicado: (2026)
Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph
por: Wang, Fali, et al.
Publicado: (2025)
por: Wang, Fali, et al.
Publicado: (2025)
Evaluating Long Range Dependency Handling in Code Generation LLMs
por: Assogba, Yannick, et al.
Publicado: (2024)
por: Assogba, Yannick, et al.
Publicado: (2024)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
Adversarial Lens: Exploiting Attention Layers to Generate Adversarial Examples for Evaluation
por: Dhole, Kaustubh
Publicado: (2025)
por: Dhole, Kaustubh
Publicado: (2025)
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
por: Fronsdal, Kai, et al.
Publicado: (2024)
por: Fronsdal, Kai, et al.
Publicado: (2024)
Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach
por: Quevedo, Ernesto, et al.
Publicado: (2024)
por: Quevedo, Ernesto, et al.
Publicado: (2024)
Consistency Evaluation of News Article Summaries Generated by Large (and Small) Language Models
por: Gilhuly, Colleen, et al.
Publicado: (2025)
por: Gilhuly, Colleen, et al.
Publicado: (2025)
Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs
por: Keeman, Michael
Publicado: (2026)
por: Keeman, Michael
Publicado: (2026)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
por: Orgad, Hadas, et al.
Publicado: (2026)
por: Orgad, Hadas, et al.
Publicado: (2026)
GPT-4 Generated Narratives of Life Events using a Structured Narrative Prompt: A Validation Study
por: Lynch, Christopher J., et al.
Publicado: (2024)
por: Lynch, Christopher J., et al.
Publicado: (2024)
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
por: Ghosh, Shubhra, et al.
Publicado: (2025)
por: Ghosh, Shubhra, et al.
Publicado: (2025)
Dealing with Annotator Disagreement in Hate Speech Classification
por: Dehghan, Somaiyeh, et al.
Publicado: (2025)
por: Dehghan, Somaiyeh, et al.
Publicado: (2025)
Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policies
por: Hong, Chunsan, et al.
Publicado: (2025)
por: Hong, Chunsan, et al.
Publicado: (2025)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
por: Cho, Seonglae, et al.
Publicado: (2025)
por: Cho, Seonglae, et al.
Publicado: (2025)
Generalizing Numerical Reasoning in Table Data through Operation Sketches and Self-Supervised Learning
por: Cho, Hanjun, et al.
Publicado: (2026)
por: Cho, Hanjun, et al.
Publicado: (2026)
Explainable AI for Smart Greenhouse Control: Interpretability of Temporal Fusion Transformer in the Internet of Robotic Things
por: Bashir, Muhammad Jawad, et al.
Publicado: (2025)
por: Bashir, Muhammad Jawad, et al.
Publicado: (2025)
AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
por: Simplício, Afonso, et al.
Publicado: (2026)
por: Simplício, Afonso, et al.
Publicado: (2026)
Lossless Prompt Compression via Dictionary-Encoding and In-Context Learning: Enabling Cost-Effective LLM Analysis of Repetitive Data
por: de Campos, Andresa Rodrigues, et al.
Publicado: (2026)
por: de Campos, Andresa Rodrigues, et al.
Publicado: (2026)
On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning
por: Deshmukh, Pratik, et al.
Publicado: (2026)
por: Deshmukh, Pratik, et al.
Publicado: (2026)
Diverse LLMs or Diverse Question Interpretations? That is the Ensembling Question
por: Rosales, Rafael, et al.
Publicado: (2025)
por: Rosales, Rafael, et al.
Publicado: (2025)
Mechanistic evaluation of Transformers and state space models
por: Arora, Aryaman, et al.
Publicado: (2025)
por: Arora, Aryaman, et al.
Publicado: (2025)
DeFTX: Denoised Sparse Fine-Tuning for Zero-Shot Cross-Lingual Transfer
por: Simon, Sona Elza, et al.
Publicado: (2025)
por: Simon, Sona Elza, et al.
Publicado: (2025)
Ejemplares similares
-
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023) -
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
por: Herbst, Jeremy, et al.
Publicado: (2026) -
Evaluating Brain-Inspired Modular Training in Automated Circuit Discovery for Mechanistic Interpretability
por: Nainani, Jatin
Publicado: (2024) -
Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
por: Gao, Yutong, et al.
Publicado: (2026) -
Do Models Know Why They Changed Their Mind? Interpretability and Faithfulness of Chain-of-Thought Under Knowledge Conflict
por: Venkata, Pruthvinath Jeripity
Publicado: (2026)