Automatically Identifying Local and Global Circuits with Linear Computation Graphs
Fuente:
arXiv
Guardado en:
| Autores principales: | Ge, Xuyang, Zhu, Fukang, Shu, Wentao, Wang, Junxuan, He, Zhengfu, Qiu, Xipeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
por: Wang, Junxuan, et al.
Publicado: (2025)
por: Wang, Junxuan, et al.
Publicado: (2025)
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
por: He, Zhengfu, et al.
Publicado: (2025)
por: He, Zhengfu, et al.
Publicado: (2025)
Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures
por: Wang, Junxuan, et al.
Publicado: (2024)
por: Wang, Junxuan, et al.
Publicado: (2024)
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
por: He, Zhengfu, et al.
Publicado: (2024)
por: He, Zhengfu, et al.
Publicado: (2024)
Evolution of Concepts in Language Model Pre-Training
por: Ge, Xuyang, et al.
Publicado: (2025)
por: Ge, Xuyang, et al.
Publicado: (2025)
Tracing the Thought of a Grandmaster-level Chess-Playing Transformer
por: Lin, Rui, et al.
Publicado: (2026)
por: Lin, Rui, et al.
Publicado: (2026)
Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT
por: He, Zhengfu, et al.
Publicado: (2024)
por: He, Zhengfu, et al.
Publicado: (2024)
Identifying a Circuit for Verb Conjugation in GPT-2
por: Africa, David Demitri
Publicado: (2025)
por: Africa, David Demitri
Publicado: (2025)
Linear Attention Sequence Parallelism
por: Sun, Weigao, et al.
Publicado: (2024)
por: Sun, Weigao, et al.
Publicado: (2024)
Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training
por: Zhang, Mozhi, et al.
Publicado: (2025)
por: Zhang, Mozhi, et al.
Publicado: (2025)
AdaLomo: Low-memory Optimization with Adaptive Learning Rate
por: Lv, Kai, et al.
Publicado: (2023)
por: Lv, Kai, et al.
Publicado: (2023)
Efficient End-to-end Language Model Fine-tuning on Graphs
por: Xue, Rui, et al.
Publicado: (2023)
por: Xue, Rui, et al.
Publicado: (2023)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
por: Peng, Runyu, et al.
Publicado: (2026)
por: Peng, Runyu, et al.
Publicado: (2026)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
por: Patel, Dev, et al.
Publicado: (2025)
por: Patel, Dev, et al.
Publicado: (2025)
The language of time: a language model perspective on time-series foundation models
por: Xie, Yi, et al.
Publicado: (2025)
por: Xie, Yi, et al.
Publicado: (2025)
Graph Linearization Methods for Reasoning on Graphs with Large Language Models
por: Xypolopoulos, Christos, et al.
Publicado: (2024)
por: Xypolopoulos, Christos, et al.
Publicado: (2024)
Sheaf Discovery with Joint Computation Graph Pruning and Flexible Granularity
por: Yu, Lei, et al.
Publicado: (2024)
por: Yu, Lei, et al.
Publicado: (2024)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
por: Zeng, Zhiyuan, et al.
Publicado: (2025)
por: Zeng, Zhiyuan, et al.
Publicado: (2025)
BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods
por: Mondorf, Philipp, et al.
Publicado: (2025)
por: Mondorf, Philipp, et al.
Publicado: (2025)
A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle
por: Zhou, Guancheng, et al.
Publicado: (2026)
por: Zhou, Guancheng, et al.
Publicado: (2026)
Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention
por: Liao, Huanxuan, et al.
Publicado: (2025)
por: Liao, Huanxuan, et al.
Publicado: (2025)
Unsupervised Domain Adaptation with Global and Local Graph Neural Networks in Limited Labeled Data Scenario: Application to Disaster Management
por: Ghosh, Samujjwal, et al.
Publicado: (2021)
por: Ghosh, Samujjwal, et al.
Publicado: (2021)
Position-aware Automatic Circuit Discovery
por: Haklay, Tal, et al.
Publicado: (2025)
por: Haklay, Tal, et al.
Publicado: (2025)
From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models
por: Wang, Ziyan, et al.
Publicado: (2025)
por: Wang, Ziyan, et al.
Publicado: (2025)
EpilepsyLLM: Domain-Specific Large Language Model Fine-tuned with Epilepsy Medical Knowledge
por: Zhao, Xuyang, et al.
Publicado: (2024)
por: Zhao, Xuyang, et al.
Publicado: (2024)
BEDTime: A Unified Benchmark for Automatically Describing Time Series
por: Sen, Medhasweta, et al.
Publicado: (2025)
por: Sen, Medhasweta, et al.
Publicado: (2025)
Judge Circuits
por: Feldhus, Nils, et al.
Publicado: (2026)
por: Feldhus, Nils, et al.
Publicado: (2026)
Counterfactual Generation with Identifiability Guarantees
por: Yan, Hanqi, et al.
Publicado: (2024)
por: Yan, Hanqi, et al.
Publicado: (2024)
pQuant: Towards Effective Low-Bit Language Models via Decoupled Linear Quantization-Aware Training
por: Zhang, Wenzheng, et al.
Publicado: (2026)
por: Zhang, Wenzheng, et al.
Publicado: (2026)
Glider: Global and Local Instruction-Driven Expert Router
por: Li, Pingzhi, et al.
Publicado: (2024)
por: Li, Pingzhi, et al.
Publicado: (2024)
Global-to-Local Support Spectrums for Language Model Explainability
por: Agussurja, Lucas, et al.
Publicado: (2024)
por: Agussurja, Lucas, et al.
Publicado: (2024)
Explicit Multi-head Attention for Inter-head Interaction in Large Language Models
por: Peng, Runyu, et al.
Publicado: (2026)
por: Peng, Runyu, et al.
Publicado: (2026)
Shared Global and Local Geometry of Language Model Embeddings
por: Lee, Andrew, et al.
Publicado: (2025)
por: Lee, Andrew, et al.
Publicado: (2025)
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
por: Ye, Jiasheng, et al.
Publicado: (2024)
por: Ye, Jiasheng, et al.
Publicado: (2024)
CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop Reasoning
por: Sung, Junyoung, et al.
Publicado: (2026)
por: Sung, Junyoung, et al.
Publicado: (2026)
Learning Multiplex Representations on Text-Attributed Graphs with One Language Model Encoder
por: Jin, Bowen, et al.
Publicado: (2023)
por: Jin, Bowen, et al.
Publicado: (2023)
Scaling Linear Attention with Sparse State Expansion
por: Pan, Yuqi, et al.
Publicado: (2025)
por: Pan, Yuqi, et al.
Publicado: (2025)
BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments
por: Wang, Xinghao, et al.
Publicado: (2024)
por: Wang, Xinghao, et al.
Publicado: (2024)
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
por: De, Soham, et al.
Publicado: (2024)
por: De, Soham, et al.
Publicado: (2024)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
por: Xiao, Hanqi, et al.
Publicado: (2025)
por: Xiao, Hanqi, et al.
Publicado: (2025)
Ejemplares similares
-
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
por: Wang, Junxuan, et al.
Publicado: (2025) -
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
por: He, Zhengfu, et al.
Publicado: (2025) -
Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures
por: Wang, Junxuan, et al.
Publicado: (2024) -
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
por: He, Zhengfu, et al.
Publicado: (2024) -
Evolution of Concepts in Language Model Pre-Training
por: Ge, Xuyang, et al.
Publicado: (2025)