LLMs Process Lists With General Filter Heads
Fuente:
arXiv
Saved in:
| Main Authors: | Sharma, Arnab Sen, Rogers, Giordano, Shapira, Natalie, Bau, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
Rethinking Visual Information Processing in Multimodal LLMs
by: Kim, Dongwan, et al.
Published: (2025)
by: Kim, Dongwan, et al.
Published: (2025)
The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
by: Mueller, Aaron, et al.
Published: (2024)
by: Mueller, Aaron, et al.
Published: (2024)
Locating and Editing Factual Associations in Mamba
by: Sharma, Arnab Sen, et al.
Published: (2024)
by: Sharma, Arnab Sen, et al.
Published: (2024)
You're (Not) My Type -- Can LLMs Generate Feedback of Specific Types for Introductory Programming Tasks?
by: Lohr, Dominic, et al.
Published: (2024)
by: Lohr, Dominic, et al.
Published: (2024)
Do explanations generalize across large reasoning models?
by: Pal, Koyena, et al.
Published: (2026)
by: Pal, Koyena, et al.
Published: (2026)
Model Lakes
by: Pal, Koyena, et al.
Published: (2024)
by: Pal, Koyena, et al.
Published: (2024)
Leveraging LLMs for Bangla Grammar Error Correction:Error Categorization, Synthetic Data, and Model Evaluation
by: Bhattacharyya, Pramit, et al.
Published: (2024)
by: Bhattacharyya, Pramit, et al.
Published: (2024)
Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon
by: Cohen-Inger, Nurit, et al.
Published: (2025)
by: Cohen-Inger, Nurit, et al.
Published: (2025)
Uncovering the Computational Ingredients of Human-Like Representations in LLMs
by: Studdiford, Zach, et al.
Published: (2025)
by: Studdiford, Zach, et al.
Published: (2025)
BioVERSE: Representation Alignment of Biomedical Modalities to LLMs for Multi-Modal Reasoning
by: Tsou, Ching-Huei, et al.
Published: (2025)
by: Tsou, Ching-Huei, et al.
Published: (2025)
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
by: Balogh, Peter
Published: (2026)
by: Balogh, Peter
Published: (2026)
How RLHF Amplifies Sycophancy
by: Shapira, Itai, et al.
Published: (2026)
by: Shapira, Itai, et al.
Published: (2026)
Language Models use Lookbacks to Track Beliefs
by: Prakash, Nikhil, et al.
Published: (2025)
by: Prakash, Nikhil, et al.
Published: (2025)
Generative AI in Science: Applications, Challenges, and Emerging Questions
by: Harries, Ryan, et al.
Published: (2025)
by: Harries, Ryan, et al.
Published: (2025)
Modeling the Data-Generating Process is Necessary for Out-of-Distribution Generalization
by: Kaur, Jivat Neet, et al.
Published: (2022)
by: Kaur, Jivat Neet, et al.
Published: (2022)
The Two Sides of the Coin: Hallucination Generation and Detection with LLMs as Evaluators for LLMs
by: Bui, Anh Thu Maria, et al.
Published: (2024)
by: Bui, Anh Thu Maria, et al.
Published: (2024)
Debiasing LLMs by Masking Unfairness-Driving Attention Heads
by: Han, Tingxu, et al.
Published: (2025)
by: Han, Tingxu, et al.
Published: (2025)
Atomic Calibration of LLMs in Long-Form Generations
by: Zhang, Caiqi, et al.
Published: (2024)
by: Zhang, Caiqi, et al.
Published: (2024)
Evaluating Role-Consistency in LLMs for Counselor Training
by: Rudolph, Eric, et al.
Published: (2026)
by: Rudolph, Eric, et al.
Published: (2026)
Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare
by: Ahsan, Hiba, et al.
Published: (2025)
by: Ahsan, Hiba, et al.
Published: (2025)
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
HeadCT-ONE: Enabling Granular and Controllable Automated Evaluation of Head CT Radiology Report Generation
by: Acosta, Julián N., et al.
Published: (2024)
by: Acosta, Julián N., et al.
Published: (2024)
Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation
by: Fang, Sen, et al.
Published: (2025)
by: Fang, Sen, et al.
Published: (2025)
Hybrid Reward-Driven Reinforcement Learning for Efficient Quantum Circuit Synthesis
by: Giordano, Sara, et al.
Published: (2025)
by: Giordano, Sara, et al.
Published: (2025)
Rethinking Explanations: Formalizing Contrast in Description Logics
by: Mahmood, Yasir, et al.
Published: (2026)
by: Mahmood, Yasir, et al.
Published: (2026)
Discovering Forbidden Topics in Language Models
by: Rager, Can, et al.
Published: (2025)
by: Rager, Can, et al.
Published: (2025)
Unveiling Hidden Links Between Unseen Security Entities
by: Alfasi, Daniel, et al.
Published: (2024)
by: Alfasi, Daniel, et al.
Published: (2024)
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
by: Yang, Zhuonan, et al.
Published: (2026)
by: Yang, Zhuonan, et al.
Published: (2026)
Patched RTC: evaluating LLMs for diverse software development tasks
by: Sharma, Asankhaya
Published: (2024)
by: Sharma, Asankhaya
Published: (2024)
Feedback-Generation for Programming Exercises With GPT-4
by: Azaiz, Imen, et al.
Published: (2024)
by: Azaiz, Imen, et al.
Published: (2024)
The Collective Turing Test: Large Language Models Can Generate Realistic Multi-User Discussions
by: Bouleimen, Azza, et al.
Published: (2025)
by: Bouleimen, Azza, et al.
Published: (2025)
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
by: Sun, Xiangkun, et al.
Published: (2026)
by: Sun, Xiangkun, et al.
Published: (2026)
Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
by: Sharma, Vibhhu, et al.
Published: (2026)
by: Sharma, Vibhhu, et al.
Published: (2026)
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
by: Chandna, Bhavik, et al.
Published: (2025)
by: Chandna, Bhavik, et al.
Published: (2025)
Rise of Generative Artificial Intelligence in Science
by: Ding, Liangping, et al.
Published: (2024)
by: Ding, Liangping, et al.
Published: (2024)
Molar: Multimodal LLMs with Collaborative Filtering Alignment for Enhanced Sequential Recommendation
by: Luo, Yucong, et al.
Published: (2024)
by: Luo, Yucong, et al.
Published: (2024)
Enhancing Temporal Awareness in LLMs for Temporal Point Processes
by: Chen, Lili, et al.
Published: (2025)
by: Chen, Lili, et al.
Published: (2025)
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
by: Kim, Jinhwa, et al.
Published: (2025)
by: Kim, Jinhwa, et al.
Published: (2025)
NoVo: Norm Voting off Hallucinations with Attention Heads in Large Language Models
by: Ho, Zheng Yi, et al.
Published: (2024)
by: Ho, Zheng Yi, et al.
Published: (2024)
Similar Items
-
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
by: Karvonen, Adam, et al.
Published: (2025) -
Rethinking Visual Information Processing in Multimodal LLMs
by: Kim, Dongwan, et al.
Published: (2025) -
The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
by: Mueller, Aaron, et al.
Published: (2024) -
Locating and Editing Factual Associations in Mamba
by: Sharma, Arnab Sen, et al.
Published: (2024) -
You're (Not) My Type -- Can LLMs Generate Feedback of Specific Types for Introductory Programming Tasks?
by: Lohr, Dominic, et al.
Published: (2024)