EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Lin, Dong, Wenshuo, Zhang, Zhuoran, Yang, Shu, Hu, Lijie, Liu, Ninghao, Zhou, Pan, Wang, Di |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PAHQ: Accelerating Automated Circuit Discovery through Mixed-Precision Inference Optimization
by: Wang, Xinhai, et al.
Published: (2025)
by: Wang, Xinhai, et al.
Published: (2025)
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
by: Jiang, Xinyan, et al.
Published: (2026)
by: Jiang, Xinyan, et al.
Published: (2026)
Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
by: Shu, Dong, et al.
Published: (2025)
by: Shu, Dong, et al.
Published: (2025)
Improving Interpretation Faithfulness for Vision Transformers
by: Hu, Lijie, et al.
Published: (2023)
by: Hu, Lijie, et al.
Published: (2023)
Understanding and Mitigating Cross-lingual Privacy Leakage via Language-specific and Universal Privacy Neurons
by: Dong, Wenshuo, et al.
Published: (2025)
by: Dong, Wenshuo, et al.
Published: (2025)
Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images
by: You, Liangliang, et al.
Published: (2025)
by: You, Liangliang, et al.
Published: (2025)
Gradient-Informed Temporal Sampling Improves Rollout Accuracy in PDE Surrogate Training
by: Wang, Wenshuo, et al.
Published: (2026)
by: Wang, Wenshuo, et al.
Published: (2026)
Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
by: Xu, Juangui, et al.
Published: (2025)
by: Xu, Juangui, et al.
Published: (2025)
When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs
by: Zhang, Zhuoran, et al.
Published: (2025)
by: Zhang, Zhuoran, et al.
Published: (2025)
Locate-then-edit for Multi-hop Factual Recall under Knowledge Editing
by: Zhang, Zhuoran, et al.
Published: (2024)
by: Zhang, Zhuoran, et al.
Published: (2024)
BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation
by: Wang, Yun, et al.
Published: (2026)
by: Wang, Yun, et al.
Published: (2026)
Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images
by: Yang, Qishun, et al.
Published: (2026)
by: Yang, Qishun, et al.
Published: (2026)
Towards Multi-dimensional Explanation Alignment for Medical Classification
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
Adaptive Multi-Subspace Representation Steering for Attribute Alignment in Large Language Models
by: Jiang, Xinyan, et al.
Published: (2025)
by: Jiang, Xinyan, et al.
Published: (2025)
Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
by: Wang, Anyi, et al.
Published: (2025)
by: Wang, Anyi, et al.
Published: (2025)
Large Language Model Enhanced Hard Sample Identification for Denoising Recommendation
by: Song, Tianrui, et al.
Published: (2024)
by: Song, Tianrui, et al.
Published: (2024)
Attention Saturation and Gradient Suppression at Inflection Layers: Diagnosing and Mitigating Bottlenecks in Transformer Adaptation
by: Zixian, Wang
Published: (2025)
by: Zixian, Wang
Published: (2025)
A Hopfieldian View-based Interpretation for Chain-of-Thought Reasoning
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
Controlling Repetition in Protein Language Models
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
Algorithmic Recourse of In-Context Learning for Tabular Data
by: Dong, Wenshuo, et al.
Published: (2026)
by: Dong, Wenshuo, et al.
Published: (2026)
Generative Models and Connected and Automated Vehicles: A Survey in Exploring the Intersection of Transportation and AI
by: Shu, Bo, et al.
Published: (2024)
by: Shu, Bo, et al.
Published: (2024)
AMSnet-q: Unsupervised Circuit Identification and Performance Labeling for AMS Circuits
by: Zhang, Ze, et al.
Published: (2026)
by: Zhang, Ze, et al.
Published: (2026)
The Compositional Architecture of Regret in Large Language Models
by: Cui, Xiangxiang, et al.
Published: (2025)
by: Cui, Xiangxiang, et al.
Published: (2025)
Exploring the Personality Traits of LLMs through Latent Features Steering
by: Yang, Shu, et al.
Published: (2024)
by: Yang, Shu, et al.
Published: (2024)
System-Anchored Knee Estimation for Low-Cost Context Window Selection in PDE Forecasting
by: Wang, Wenshuo, et al.
Published: (2026)
by: Wang, Wenshuo, et al.
Published: (2026)
In-Run Data Shapley for Adam Optimizer
by: Ding, Meng, et al.
Published: (2026)
by: Ding, Meng, et al.
Published: (2026)
Derived Fields Preserve Fine-Scale Detail in Budgeted Neural Simulators
by: Wang, Wenshuo, et al.
Published: (2026)
by: Wang, Wenshuo, et al.
Published: (2026)
Posterior-First Neural PDE Simulation: Inferring Hidden Problem State from a Single Field
by: Wang, Wenshuo, et al.
Published: (2026)
by: Wang, Wenshuo, et al.
Published: (2026)
MoRAL: MoE Augmented LoRA for LLMs' Lifelong Learning
by: Yang, Shu, et al.
Published: (2024)
by: Yang, Shu, et al.
Published: (2024)
COMPKE: Complex Question Answering under Knowledge Editing
by: Cheng, Keyuan, et al.
Published: (2025)
by: Cheng, Keyuan, et al.
Published: (2025)
GP-GPT: Large Language Model for Gene-Phenotype Mapping
by: Lyu, Yanjun, et al.
Published: (2024)
by: Lyu, Yanjun, et al.
Published: (2024)
LLM Reasoning Is Latent, Not the Chain of Thought
by: Wang, Wenshuo
Published: (2026)
by: Wang, Wenshuo
Published: (2026)
LLMs Should Not Yet Be Credited with Decision Explanation
by: Wang, Wenshuo
Published: (2026)
by: Wang, Wenshuo
Published: (2026)
Breaking Scale Anchoring: Frequency Representation Learning for Accurate High-Resolution Inference from Low-Resolution Training
by: Wang, Wenshuo, et al.
Published: (2025)
by: Wang, Wenshuo, et al.
Published: (2025)
Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization
by: Yang, Tiancheng, et al.
Published: (2025)
by: Yang, Tiancheng, et al.
Published: (2025)
Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration
by: Zhang, Ninghao, et al.
Published: (2026)
by: Zhang, Ninghao, et al.
Published: (2026)
Understanding the Dynamics of Demonstration Conflict in In-Context Learning
by: Jiao, Difan, et al.
Published: (2026)
by: Jiao, Difan, et al.
Published: (2026)
Knowledge Distillation Must Account for What It Loses
by: Wang, Wenshuo
Published: (2026)
by: Wang, Wenshuo
Published: (2026)
Evaluating Data Influence in Meta Learning
by: Ren, Chenyang, et al.
Published: (2025)
by: Ren, Chenyang, et al.
Published: (2025)
Similar Items
-
PAHQ: Accelerating Automated Circuit Discovery through Mixed-Precision Inference Optimization
by: Wang, Xinhai, et al.
Published: (2025) -
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
by: Jiang, Xinyan, et al.
Published: (2026) -
Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning
by: Zhang, Lin, et al.
Published: (2025) -
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
by: Shu, Dong, et al.
Published: (2025) -
Improving Interpretation Faithfulness for Vision Transformers
by: Hu, Lijie, et al.
Published: (2023)