LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jing, Yi, Yao, Zijun, Guo, Hongzhu, Ran, Lingxu, Wang, Xiaozhi, Hou, Lei, Li, Juanzi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
von: Jing, Yi, et al.
Veröffentlicht: (2026)
von: Jing, Yi, et al.
Veröffentlicht: (2026)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
von: Chen, Jianhui, et al.
Veröffentlicht: (2024)
von: Chen, Jianhui, et al.
Veröffentlicht: (2024)
Auxiliary Metrics Help Decoding Skill Neurons in the Wild
von: Zhao, Yixiu, et al.
Veröffentlicht: (2025)
von: Zhao, Yixiu, et al.
Veröffentlicht: (2025)
WildReward: Learning Reward Models from In-the-Wild Human Interactions
von: Peng, Hao, et al.
Veröffentlicht: (2026)
von: Peng, Hao, et al.
Veröffentlicht: (2026)
Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders
von: Wang, Shun, et al.
Veröffentlicht: (2025)
von: Wang, Shun, et al.
Veröffentlicht: (2025)
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems
von: Peng, Hao, et al.
Veröffentlicht: (2025)
von: Peng, Hao, et al.
Veröffentlicht: (2025)
ADELIE: Aligning Large Language Models on Information Extraction
von: Qi, Yunjia, et al.
Veröffentlicht: (2024)
von: Qi, Yunjia, et al.
Veröffentlicht: (2024)
Constraint Back-translation Improves Complex Instruction Following of Large Language Models
von: Qi, Yunjia, et al.
Veröffentlicht: (2024)
von: Qi, Yunjia, et al.
Veröffentlicht: (2024)
Pre-training Distillation for Large Language Models: A Design Space Exploration
von: Peng, Hao, et al.
Veröffentlicht: (2024)
von: Peng, Hao, et al.
Veröffentlicht: (2024)
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning
von: Xin, Amy, et al.
Veröffentlicht: (2024)
von: Xin, Amy, et al.
Veröffentlicht: (2024)
Sparse Auto-Encoders and Holism about Large Language Models
von: Grindrod, Jumbly
Veröffentlicht: (2026)
von: Grindrod, Jumbly
Veröffentlicht: (2026)
Evaluating Generative Language Models in Information Extraction as Subjective Question Correction
von: Fan, Yuchen, et al.
Veröffentlicht: (2024)
von: Fan, Yuchen, et al.
Veröffentlicht: (2024)
RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
von: Liu, Yantao, et al.
Veröffentlicht: (2024)
von: Liu, Yantao, et al.
Veröffentlicht: (2024)
Understanding the Mechanism of Altruism in Large Language Models
von: Zhang, Shuhuai, et al.
Veröffentlicht: (2026)
von: Zhang, Shuhuai, et al.
Veröffentlicht: (2026)
HalluSAE: Detecting Hallucinations in Large Language Models via Sparse Auto-Encoders
von: Chen, Boshui, et al.
Veröffentlicht: (2026)
von: Chen, Boshui, et al.
Veröffentlicht: (2026)
LLMAEL: Large Language Models are Good Context Augmenters for Entity Linking
von: Xin, Amy, et al.
Veröffentlicht: (2024)
von: Xin, Amy, et al.
Veröffentlicht: (2024)
Untangle the KNOT: Interweaving Conflicting Knowledge and Reasoning Skills in Large Language Models
von: Liu, Yantao, et al.
Veröffentlicht: (2024)
von: Liu, Yantao, et al.
Veröffentlicht: (2024)
OpenEP: Open-Ended Future Event Prediction
von: Guan, Yong, et al.
Veröffentlicht: (2024)
von: Guan, Yong, et al.
Veröffentlicht: (2024)
AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
von: Qi, Yunjia, et al.
Veröffentlicht: (2025)
von: Qi, Yunjia, et al.
Veröffentlicht: (2025)
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
MAVEN-Fact: A Large-scale Event Factuality Detection Dataset
von: Li, Chunyang, et al.
Veröffentlicht: (2024)
von: Li, Chunyang, et al.
Veröffentlicht: (2024)
WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2023)
von: Tu, Shangqing, et al.
Veröffentlicht: (2023)
PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament
von: Liu, Yantao, et al.
Veröffentlicht: (2025)
von: Liu, Yantao, et al.
Veröffentlicht: (2025)
Aligning Teacher with Student Preferences for Tailored Training Data Generation
von: Liu, Yantao, et al.
Veröffentlicht: (2024)
von: Liu, Yantao, et al.
Veröffentlicht: (2024)
DICE: Detecting In-distribution Contamination in LLM's Fine-tuning Phase for Math Reasoning
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
A Cause-Effect Look at Alleviating Hallucination of Knowledge-grounded Dialogue Generation
von: Yu, Jifan, et al.
Veröffentlicht: (2024)
von: Yu, Jifan, et al.
Veröffentlicht: (2024)
Boundary-Guided Policy Optimization for Memory-efficient RL of Diffusion Large Language Models
von: Lin, Nianyi, et al.
Veröffentlicht: (2025)
von: Lin, Nianyi, et al.
Veröffentlicht: (2025)
ChatLog: Carefully Evaluating the Evolution of ChatGPT Across Time
von: Tu, Shangqing, et al.
Veröffentlicht: (2023)
von: Tu, Shangqing, et al.
Veröffentlicht: (2023)
VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
von: Peng, Hao, et al.
Veröffentlicht: (2025)
von: Peng, Hao, et al.
Veröffentlicht: (2025)
AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders
von: Aparin, Georgii, et al.
Veröffentlicht: (2026)
von: Aparin, Georgii, et al.
Veröffentlicht: (2026)
SOSAE: Self-Organizing Sparse AutoEncoder
von: Modi, Sarthak Ketanbhai, et al.
Veröffentlicht: (2025)
von: Modi, Sarthak Ketanbhai, et al.
Veröffentlicht: (2025)
DecoderLens: Layerwise Interpretation of Encoder-Decoder Transformers
von: Langedijk, Anna, et al.
Veröffentlicht: (2023)
von: Langedijk, Anna, et al.
Veröffentlicht: (2023)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
StoryWriter: A Multi-Agent Framework for Long Story Generation
von: Xia, Haotian, et al.
Veröffentlicht: (2025)
von: Xia, Haotian, et al.
Veröffentlicht: (2025)
StoryAlign: Evaluating and Training Reward Models for Story Generation
von: Xia, Haotian, et al.
Veröffentlicht: (2026)
von: Xia, Haotian, et al.
Veröffentlicht: (2026)
TacoERE: Cluster-aware Compression for Event Relation Extraction
von: Guan, Yong, et al.
Veröffentlicht: (2024)
von: Guan, Yong, et al.
Veröffentlicht: (2024)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
von: Toker, Michael, et al.
Veröffentlicht: (2024)
von: Toker, Michael, et al.
Veröffentlicht: (2024)
Transferable and Efficient Non-Factual Content Detection via Probe Training with Offline Consistency Checking
von: Zhang, Xiaokang, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaokang, et al.
Veröffentlicht: (2024)
OmniLens: Towards Universal Lens Aberration Correction via LensLib-to-Specific Domain Adaptation
von: Jiang, Qi, et al.
Veröffentlicht: (2024)
von: Jiang, Qi, et al.
Veröffentlicht: (2024)
RePrompT: Recurrent Prompt Tuning for Integrating Structured EHR Encoders with Large Language Models
von: Moghaddam, Arya Hadizadeh, et al.
Veröffentlicht: (2026)
von: Moghaddam, Arya Hadizadeh, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
von: Jing, Yi, et al.
Veröffentlicht: (2026) -
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
von: Chen, Jianhui, et al.
Veröffentlicht: (2024) -
Auxiliary Metrics Help Decoding Skill Neurons in the Wild
von: Zhao, Yixiu, et al.
Veröffentlicht: (2025) -
WildReward: Learning Reward Models from In-the-Wild Human Interactions
von: Peng, Hao, et al.
Veröffentlicht: (2026) -
Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders
von: Wang, Shun, et al.
Veröffentlicht: (2025)