Saved in:
| Main Authors: | Mofael, Abdullah Al, Kuhn, Lisa M., Alkadi, Ghassan, Yang, Kuo-Pao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.12423 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mechanistic Interpretability of GPT-2: Lexical and Contextual Layers in Sentiment Analysis
by: Hatua, Amartya
Published: (2025)
by: Hatua, Amartya
Published: (2025)
FourierKAN outperforms MLP on Text Classification Head Fine-tuning
by: Imran, Abdullah Al, et al.
Published: (2024)
by: Imran, Abdullah Al, et al.
Published: (2024)
Positional Cognitive Specialization: Where Do LLMs Learn To Comprehend and Speak Your Language?
by: Salim, Luis Frentzen, et al.
Published: (2026)
by: Salim, Luis Frentzen, et al.
Published: (2026)
Large Language Model Pruning
by: Huang, Hanjuan, et al.
Published: (2024)
by: Huang, Hanjuan, et al.
Published: (2024)
ChatGPT Evaluation on Sentence Level Relations: A Focus on Temporal, Causal, and Discourse Relations
by: Chan, Chunkit, et al.
Published: (2023)
by: Chan, Chunkit, et al.
Published: (2023)
The Mechanics of Conceptual Interpretation in GPT Models: Interpretative Insights
by: Aljaafari, Nura, et al.
Published: (2024)
by: Aljaafari, Nura, et al.
Published: (2024)
Is ChatGPT the Future of Causal Text Mining? A Comprehensive Evaluation and Analysis
by: Takayanagi, Takehiro, et al.
Published: (2024)
by: Takayanagi, Takehiro, et al.
Published: (2024)
Interpreting Transformers Through Attention Head Intervention
by: Kadem, Mason, et al.
Published: (2026)
by: Kadem, Mason, et al.
Published: (2026)
Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics
by: Lee, Isabelle, et al.
Published: (2024)
by: Lee, Isabelle, et al.
Published: (2024)
Single and Multi-Hop Question-Answering Datasets for Reticular Chemistry with GPT-4-Turbo
by: Rampal, Nakul, et al.
Published: (2024)
by: Rampal, Nakul, et al.
Published: (2024)
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification
by: Wang, Yian, et al.
Published: (2026)
by: Wang, Yian, et al.
Published: (2026)
Generating Hard-Negative Out-of-Scope Data with ChatGPT for Intent Classification
by: Li, Zhijian, et al.
Published: (2024)
by: Li, Zhijian, et al.
Published: (2024)
BengaliFig: A Low-Resource Challenge for Figurative and Culturally Grounded Reasoning in Bengali
by: Sefat, Abdullah Al
Published: (2025)
by: Sefat, Abdullah Al
Published: (2025)
Reversed Attention: On The Gradient Descent Of Attention Layers In GPT
by: Katz, Shahar, et al.
Published: (2024)
by: Katz, Shahar, et al.
Published: (2024)
Evaluating the Predictive Capacity of ChatGPT for Academic Peer Review Outcomes Across Multiple Platforms
by: Thelwall, Mike, et al.
Published: (2024)
by: Thelwall, Mike, et al.
Published: (2024)
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
by: Wu, Zhengxuan, et al.
Published: (2023)
by: Wu, Zhengxuan, et al.
Published: (2023)
InfraGPT Smart Infrastructure: An End-to-End VLM-Based Framework for Detecting and Managing Urban Defects
by: Mohamed, Ibrahim Sheikh, et al.
Published: (2025)
by: Mohamed, Ibrahim Sheikh, et al.
Published: (2025)
SpeLLM: Character-Level Multi-Head Decoding
by: Ben-Artzy, Amit, et al.
Published: (2025)
by: Ben-Artzy, Amit, et al.
Published: (2025)
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
by: Yeo, Wei Jie, et al.
Published: (2025)
by: Yeo, Wei Jie, et al.
Published: (2025)
You can remove GPT2's LayerNorm by fine-tuning
by: Heimersheim, Stefan
Published: (2024)
by: Heimersheim, Stefan
Published: (2024)
Interpretable Next-token Prediction via the Generalized Induction Head
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Feature Extraction and Analysis for GPT-Generated Text
by: Selvioğlu, A., et al.
Published: (2025)
by: Selvioğlu, A., et al.
Published: (2025)
The Qiyas Benchmark: Measuring ChatGPT Mathematical and Language Understanding in Arabic
by: Al-Khalifa, Shahad, et al.
Published: (2024)
by: Al-Khalifa, Shahad, et al.
Published: (2024)
Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training
by: Wu, Yu-Hang, et al.
Published: (2026)
by: Wu, Yu-Hang, et al.
Published: (2026)
ChatGPT4PCG Competition: Character-like Level Generation for Science Birds
by: Taveekitworachai, Pittawat, et al.
Published: (2023)
by: Taveekitworachai, Pittawat, et al.
Published: (2023)
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
EmoBang: Detecting Emotion From Bengali Texts
by: Maruf, Abdullah Al, et al.
Published: (2025)
by: Maruf, Abdullah Al, et al.
Published: (2025)
A comparison of Human, GPT-3.5, and GPT-4 Performance in a University-Level Coding Course
by: Yeadon, Will, et al.
Published: (2024)
by: Yeadon, Will, et al.
Published: (2024)
ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability
by: Koike, Ryuto, et al.
Published: (2025)
by: Koike, Ryuto, et al.
Published: (2025)
Evaluating Subword Tokenization Techniques for Bengali: A Benchmark Study with BengaliBPE
by: Patwary, Firoj Ahmmed, et al.
Published: (2025)
by: Patwary, Firoj Ahmmed, et al.
Published: (2025)
From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models
by: Ma, Youmi, et al.
Published: (2026)
by: Ma, Youmi, et al.
Published: (2026)
Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization
by: Zhang, Weixu, et al.
Published: (2026)
by: Zhang, Weixu, et al.
Published: (2026)
TermGPT: Multi-Level Contrastive Fine-Tuning for Terminology Adaptation in Legal and Financial Domain
by: Sun, Yidan, et al.
Published: (2025)
by: Sun, Yidan, et al.
Published: (2025)
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
by: So, Yeonkyoung, et al.
Published: (2025)
by: So, Yeonkyoung, et al.
Published: (2025)
ReBeCA: Unveiling Interpretable Behavior Hierarchy behind the Iterative Self-Reflection of Language Models with Causal Analysis
by: Yan, Tianqiang, et al.
Published: (2026)
by: Yan, Tianqiang, et al.
Published: (2026)
Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
by: Abdoli, Sajjad, et al.
Published: (2026)
by: Abdoli, Sajjad, et al.
Published: (2026)
SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling
by: Zhao, Anhao, et al.
Published: (2025)
by: Zhao, Anhao, et al.
Published: (2025)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
by: Mishra, Anurag
Published: (2025)
by: Mishra, Anurag
Published: (2025)
Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation
by: Shaik, Zuhair Hasan, et al.
Published: (2025)
by: Shaik, Zuhair Hasan, et al.
Published: (2025)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
by: Musat, Tiberiu
Published: (2024)
by: Musat, Tiberiu
Published: (2024)
Similar Items
-
Mechanistic Interpretability of GPT-2: Lexical and Contextual Layers in Sentiment Analysis
by: Hatua, Amartya
Published: (2025) -
FourierKAN outperforms MLP on Text Classification Head Fine-tuning
by: Imran, Abdullah Al, et al.
Published: (2024) -
Positional Cognitive Specialization: Where Do LLMs Learn To Comprehend and Speak Your Language?
by: Salim, Luis Frentzen, et al.
Published: (2026) -
Large Language Model Pruning
by: Huang, Hanjuan, et al.
Published: (2024) -
ChatGPT Evaluation on Sentence Level Relations: A Focus on Temporal, Causal, and Discourse Relations
by: Chan, Chunkit, et al.
Published: (2023)