Observable Propagation: Uncovering Feature Vectors in Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dunefsky, Jacob, Cohan, Arman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
One-shot Optimized Steering Vectors Mediate Safety-relevant Behaviors in LLMs
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2025)
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2025)
Transcoders Find Interpretable LLM Feature Circuits
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2024)
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2024)
Understanding Reference Policies in Direct Preference Optimization
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
von: Riddell, Martin, et al.
Veröffentlicht: (2024)
von: Riddell, Martin, et al.
Veröffentlicht: (2024)
Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
von: Shaw, Peter, et al.
Veröffentlicht: (2025)
von: Shaw, Peter, et al.
Veröffentlicht: (2025)
MDCure: A Scalable Pipeline for Multi-Document Instruction-Following
von: Liu, Gabrielle Kaili-May, et al.
Veröffentlicht: (2024)
von: Liu, Gabrielle Kaili-May, et al.
Veröffentlicht: (2024)
FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
von: Hu, Tiansheng, et al.
Veröffentlicht: (2025)
von: Hu, Tiansheng, et al.
Veröffentlicht: (2025)
ALTA: Compiler-Based Analysis of Transformers
von: Shaw, Peter, et al.
Veröffentlicht: (2024)
von: Shaw, Peter, et al.
Veröffentlicht: (2024)
Calibrating Long-form Generations from Large Language Models
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
von: Jones, Erik, et al.
Veröffentlicht: (2025)
von: Jones, Erik, et al.
Veröffentlicht: (2025)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
TESS: Text-to-Text Self-Conditioned Simplex Diffusion
von: Mahabadi, Rabeeh Karimi, et al.
Veröffentlicht: (2023)
von: Mahabadi, Rabeeh Karimi, et al.
Veröffentlicht: (2023)
MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
von: Liu, Gabrielle Kaili-May, et al.
Veröffentlicht: (2025)
von: Liu, Gabrielle Kaili-May, et al.
Veröffentlicht: (2025)
NExT: Teaching Large Language Models to Reason about Code Execution
von: Ni, Ansong, et al.
Veröffentlicht: (2024)
von: Ni, Ansong, et al.
Veröffentlicht: (2024)
References Improve LLM Alignment in Non-Verifiable Domains
von: Shi, Kejian, et al.
Veröffentlicht: (2026)
von: Shi, Kejian, et al.
Veröffentlicht: (2026)
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
FinDVer: Explainable Claim Verification over Long and Hybrid-Content Financial Documents
von: Zhao, Yilun, et al.
Veröffentlicht: (2024)
von: Zhao, Yilun, et al.
Veröffentlicht: (2024)
Bridging the Dimensional Chasm: Uncover Layer-wise Dimensional Reduction in Transformers through Token Correlation
von: Song, Zhuo-Yang, et al.
Veröffentlicht: (2025)
von: Song, Zhuo-Yang, et al.
Veröffentlicht: (2025)
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
von: Liu, Yixin, et al.
Veröffentlicht: (2023)
von: Liu, Yixin, et al.
Veröffentlicht: (2023)
Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization Failure
von: Wang, Boshi, et al.
Veröffentlicht: (2025)
von: Wang, Boshi, et al.
Veröffentlicht: (2025)
Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers
von: Yan, Hao, et al.
Veröffentlicht: (2026)
von: Yan, Hao, et al.
Veröffentlicht: (2026)
ReIFE: Re-evaluating Instruction-Following Evaluation
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions
von: Weller, Orion, et al.
Veröffentlicht: (2024)
von: Weller, Orion, et al.
Veröffentlicht: (2024)
Survey on Evaluation of LLM-based Agents
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
Dual Path Attribution: Efficient Attribution for SwiGLU-Transformers through Layer-Wise Target Propagation
von: Jantsch, Lasse Marten, et al.
Veröffentlicht: (2026)
von: Jantsch, Lasse Marten, et al.
Veröffentlicht: (2026)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
Extracting Rule-based Descriptions of Attention Features in Transformers
von: Friedman, Dan, et al.
Veröffentlicht: (2025)
von: Friedman, Dan, et al.
Veröffentlicht: (2025)
Transformer-VQ: Linear-Time Transformers via Vector Quantization
von: Lingle, Lucas D.
Veröffentlicht: (2023)
von: Lingle, Lucas D.
Veröffentlicht: (2023)
Pairwise Reference Alignment as a Model-Level Ordinal Observable
von: Li, Mujing
Veröffentlicht: (2026)
von: Li, Mujing
Veröffentlicht: (2026)
ComplexFormer: Disruptively Advancing Transformer Inference Ability via Head-Specific Complex Vector Attention
von: Shao, Jintian, et al.
Veröffentlicht: (2025)
von: Shao, Jintian, et al.
Veröffentlicht: (2025)
Uncovering Cross-Objective Interference in Multi-Objective Alignment
von: Lu, Yining, et al.
Veröffentlicht: (2026)
von: Lu, Yining, et al.
Veröffentlicht: (2026)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers
von: Xu, Ruichen, et al.
Veröffentlicht: (2026)
von: Xu, Ruichen, et al.
Veröffentlicht: (2026)
Transformers as Support Vector Machines
von: Tarzanagh, Davoud Ataee, et al.
Veröffentlicht: (2023)
von: Tarzanagh, Davoud Ataee, et al.
Veröffentlicht: (2023)
Algorithmic Capabilities of Random Transformers
von: Zhong, Ziqian, et al.
Veröffentlicht: (2024)
von: Zhong, Ziqian, et al.
Veröffentlicht: (2024)
Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
von: Ahmad, Areeb, et al.
Veröffentlicht: (2025)
von: Ahmad, Areeb, et al.
Veröffentlicht: (2025)
mFollowIR: a Multilingual Benchmark for Instruction Following in Retrieval
von: Weller, Orion, et al.
Veröffentlicht: (2025)
von: Weller, Orion, et al.
Veröffentlicht: (2025)
Dimensionality Reduction in Sentence Transformer Vector Databases with Fast Fourier Transform
von: Bulgakov, Vitaly, et al.
Veröffentlicht: (2024)
von: Bulgakov, Vitaly, et al.
Veröffentlicht: (2024)
Words in Motion: Extracting Interpretable Control Vectors for Motion Transformers
von: Tas, Omer Sahin, et al.
Veröffentlicht: (2024)
von: Tas, Omer Sahin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
One-shot Optimized Steering Vectors Mediate Safety-relevant Behaviors in LLMs
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2025) -
Transcoders Find Interpretable LLM Feature Circuits
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2024) -
Understanding Reference Policies in Direct Preference Optimization
von: Liu, Yixin, et al.
Veröffentlicht: (2024) -
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025) -
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
von: Riddell, Martin, et al.
Veröffentlicht: (2024)