AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wu, Zhengxuan, Arora, Aryaman, Geiger, Atticus, Wang, Zheng, Huang, Jing, Jurafsky, Dan, Manning, Christopher D., Potts, Christopher |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
par: Arora, Aryaman, et autres
Publié: (2024)
par: Arora, Aryaman, et autres
Publié: (2024)
Mechanistic evaluation of Transformers and state space models
par: Arora, Aryaman, et autres
Publié: (2025)
par: Arora, Aryaman, et autres
Publié: (2025)
Bayesian scaling laws for in-context learning
par: Arora, Aryaman, et autres
Publié: (2024)
par: Arora, Aryaman, et autres
Publié: (2024)
Language Model Circuits Are Sparse in the Neuron Basis
par: Arora, Aryaman, et autres
Publié: (2026)
par: Arora, Aryaman, et autres
Publié: (2026)
ADAG: Automatically Describing Attribution Graphs
par: Arora, Aryaman, et autres
Publié: (2026)
par: Arora, Aryaman, et autres
Publié: (2026)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
par: Ashuach, Tomer, et autres
Publié: (2025)
par: Ashuach, Tomer, et autres
Publié: (2025)
Tversky Neural Networks: Psychologically Plausible Deep Learning with Differentiable Tversky Similarity
par: Doumbouya, Moussa Koulako Bala, et autres
Publié: (2025)
par: Doumbouya, Moussa Koulako Bala, et autres
Publié: (2025)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
par: Cho, Seonglae, et autres
Publié: (2026)
par: Cho, Seonglae, et autres
Publié: (2026)
ReFT: Representation Finetuning for Language Models
par: Wu, Zhengxuan, et autres
Publié: (2024)
par: Wu, Zhengxuan, et autres
Publié: (2024)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
par: Cho, Seonglae, et autres
Publié: (2025)
par: Cho, Seonglae, et autres
Publié: (2025)
Learning the meanings of function words from grounded language using a visual question answering model
par: Portelance, Eva, et autres
Publié: (2023)
par: Portelance, Eva, et autres
Publié: (2023)
Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
par: Feng, Ruitao, et autres
Publié: (2025)
par: Feng, Ruitao, et autres
Publié: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
par: Schesch, Benedikt, et autres
Publié: (2026)
par: Schesch, Benedikt, et autres
Publié: (2026)
Detecting and Steering LLMs' Empathy in Action
par: Cadile, Juan P.
Publié: (2025)
par: Cadile, Juan P.
Publié: (2025)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
par: Simhi, Adi, et autres
Publié: (2025)
par: Simhi, Adi, et autres
Publié: (2025)
PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?
par: Pulipaka, Sidharth, et autres
Publié: (2026)
par: Pulipaka, Sidharth, et autres
Publié: (2026)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
par: McCann, Jordan F.
Publié: (2026)
par: McCann, Jordan F.
Publié: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
par: Saji, Alan, et autres
Publié: (2025)
par: Saji, Alan, et autres
Publié: (2025)
Curveball Steering: The Right Direction To Steer Isn't Always Linear
par: Raval, Shivam, et autres
Publié: (2026)
par: Raval, Shivam, et autres
Publié: (2026)
Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification
par: Bucher, Martin Juan José, et autres
Publié: (2024)
par: Bucher, Martin Juan José, et autres
Publié: (2024)
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
par: Jiang, Yilin, et autres
Publié: (2025)
par: Jiang, Yilin, et autres
Publié: (2025)
Reducing Information Overload: Because Even Security Experts Need to Blink
par: Kuehn, Philipp, et autres
Publié: (2022)
par: Kuehn, Philipp, et autres
Publié: (2022)
Enhancing Paraphrase Type Generation: The Impact of DPO and RLHF Evaluated with Human-Ranked Data
par: Lübbers, Christopher Lee
Publié: (2025)
par: Lübbers, Christopher Lee
Publié: (2025)
ExpressivityBench: Can LLMs Communicate Implicitly?
par: Tint, Joshua, et autres
Publié: (2024)
par: Tint, Joshua, et autres
Publié: (2024)
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
par: Agrawal, Lakshya A, et autres
Publié: (2025)
par: Agrawal, Lakshya A, et autres
Publié: (2025)
Evaluating Steering Techniques using Human Similarity Judgments
par: Studdiford, Zach, et autres
Publié: (2025)
par: Studdiford, Zach, et autres
Publié: (2025)
PathBench: Speech Intelligibility Benchmark for Automatic Pathological Speech Assessment
par: Halpern, Bence Mark, et autres
Publié: (2026)
par: Halpern, Bence Mark, et autres
Publié: (2026)
LoRS: Efficient Low-Rank Adaptation for Sparse Large Language Model
par: Hu, Yuxuan, et autres
Publié: (2025)
par: Hu, Yuxuan, et autres
Publié: (2025)
Exploring RWKV for Sentence Embeddings: Layer-wise Analysis and Baseline Comparison for Semantic Similarity
par: Pan, Xinghan
Publié: (2025)
par: Pan, Xinghan
Publié: (2025)
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
par: Patel, Het, et autres
Publié: (2026)
par: Patel, Het, et autres
Publié: (2026)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
par: Tu, Songjun, et autres
Publié: (2026)
par: Tu, Songjun, et autres
Publié: (2026)
Combining Denoising Autoencoders with Contrastive Learning to fine-tune Transformer Models
par: Lopez-Avila, Alejo, et autres
Publié: (2024)
par: Lopez-Avila, Alejo, et autres
Publié: (2024)
Simple and Effective Baselines for Code Summarisation Evaluation
par: Robinson, Jade, et autres
Publié: (2025)
par: Robinson, Jade, et autres
Publié: (2025)
Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect
par: Klerings, Alina, et autres
Publié: (2025)
par: Klerings, Alina, et autres
Publié: (2025)
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
par: Hu, Yuxuan, et autres
Publié: (2026)
par: Hu, Yuxuan, et autres
Publié: (2026)
BabyLlama-2: Ensemble-Distilled Models Consistently Outperform Teachers With Limited Data
par: Tastet, Jean-Loup, et autres
Publié: (2024)
par: Tastet, Jean-Loup, et autres
Publié: (2024)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
par: Ovcharov, Volodymyr
Publié: (2026)
par: Ovcharov, Volodymyr
Publié: (2026)
RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysis
par: Bose, Joy
Publié: (2026)
par: Bose, Joy
Publié: (2026)
Adaptive Focus Memory for Language Models
par: Cruz, Christopher
Publié: (2025)
par: Cruz, Christopher
Publié: (2025)
EVM-QuestBench: An Execution-Grounded Benchmark for Natural-Language Transaction Code Generation
par: Yang, Pei, et autres
Publié: (2026)
par: Yang, Pei, et autres
Publié: (2026)
Documents similaires
-
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
par: Arora, Aryaman, et autres
Publié: (2024) -
Mechanistic evaluation of Transformers and state space models
par: Arora, Aryaman, et autres
Publié: (2025) -
Bayesian scaling laws for in-context learning
par: Arora, Aryaman, et autres
Publié: (2024) -
Language Model Circuits Are Sparse in the Neuron Basis
par: Arora, Aryaman, et autres
Publié: (2026) -
ADAG: Automatically Describing Attribution Graphs
par: Arora, Aryaman, et autres
Publié: (2026)