Features that Make a Difference: Leveraging Gradients for Improved Dictionary Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Olmo, Jeffrey, Wilson, Jared, Forsey, Max, Hepner, Bryce, Howe, Thomas Vin, Wingate, David |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Language models struggle with compartmentalization
por: Howe, Thomas Vincent, et al.
Publicado: (2026)
por: Howe, Thomas Vincent, et al.
Publicado: (2026)
Arti-"fickle" Intelligence: Using LLMs as a Tool for Inference in the Political and Social Sciences
por: Argyle, Lisa P., et al.
Publicado: (2025)
por: Argyle, Lisa P., et al.
Publicado: (2025)
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
por: Karvonen, Adam, et al.
Publicado: (2024)
por: Karvonen, Adam, et al.
Publicado: (2024)
Scaling Law with Learning Rate Annealing
por: Tissue, Howe, et al.
Publicado: (2024)
por: Tissue, Howe, et al.
Publicado: (2024)
Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
por: Zhussip, Magauiya, et al.
Publicado: (2025)
por: Zhussip, Magauiya, et al.
Publicado: (2025)
Low-Cost Generation and Evaluation of Dictionary Example Sentences
por: Cai, Bill, et al.
Publicado: (2024)
por: Cai, Bill, et al.
Publicado: (2024)
Diagnosing Transformers: Illuminating Feature Spaces for Clinical Decision-Making
por: Hsu, Aliyah R., et al.
Publicado: (2023)
por: Hsu, Aliyah R., et al.
Publicado: (2023)
BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation
por: Zhu, Alan, et al.
Publicado: (2025)
por: Zhu, Alan, et al.
Publicado: (2025)
Learning Dynamics in Continual Pre-Training for Large Language Models
por: Wang, Xingjin, et al.
Publicado: (2025)
por: Wang, Xingjin, et al.
Publicado: (2025)
Plan Optimization to Bilingual Dictionary Induction for Low-Resource Language Families
por: Nasution, Arbi Haza, et al.
Publicado: (2020)
por: Nasution, Arbi Haza, et al.
Publicado: (2020)
Improving Dictionary Learning with Gated Sparse Autoencoders
por: Rajamanoharan, Senthooran, et al.
Publicado: (2024)
por: Rajamanoharan, Senthooran, et al.
Publicado: (2024)
StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features
por: Ochab, Jeremi K., et al.
Publicado: (2025)
por: Ochab, Jeremi K., et al.
Publicado: (2025)
Dictionary Learning: The Complexity of Learning Sparse Superposed Features with Feedback
por: Kumar, Akash
Publicado: (2025)
por: Kumar, Akash
Publicado: (2025)
Enhancing Clinical Documentation with Synthetic Data: Leveraging Generative Models for Improved Accuracy
por: Biswas, Anjanava, et al.
Publicado: (2024)
por: Biswas, Anjanava, et al.
Publicado: (2024)
Beyond Gradient and Priors in Privacy Attacks: Leveraging Pooler Layer Inputs of Language Models in Federated Learning
por: Li, Jianwei, et al.
Publicado: (2023)
por: Li, Jianwei, et al.
Publicado: (2023)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
por: Chalnev, Sviatoslav, et al.
Publicado: (2024)
por: Chalnev, Sviatoslav, et al.
Publicado: (2024)
Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learning
por: Bani-Harouni, David, et al.
Publicado: (2025)
por: Bani-Harouni, David, et al.
Publicado: (2025)
ScenicProver: A Framework for Compositional Probabilistic Verification of Learning-Enabled Systems
por: Vin, Eric, et al.
Publicado: (2025)
por: Vin, Eric, et al.
Publicado: (2025)
Improving LoRA with Variational Learning
por: Cong, Bai, et al.
Publicado: (2025)
por: Cong, Bai, et al.
Publicado: (2025)
Check Your LLM's Secret Dictionary! Five Lines of Code Reveal What Your LLM Learned (Including What It Shouldn't Have)
por: Miyashita, Hisashi
Publicado: (2026)
por: Miyashita, Hisashi
Publicado: (2026)
Dense SAE Latents Are Features, Not Bugs
por: Sun, Xiaoqing, et al.
Publicado: (2025)
por: Sun, Xiaoqing, et al.
Publicado: (2025)
Chasing COMET: Leveraging Minimum Bayes Risk Decoding for Self-Improving Machine Translation
por: Guttmann, Kamil, et al.
Publicado: (2024)
por: Guttmann, Kamil, et al.
Publicado: (2024)
Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training
por: Zhang, Mozhi, et al.
Publicado: (2025)
por: Zhang, Mozhi, et al.
Publicado: (2025)
Enhancing Antibiotic Stewardship using a Natural Language Approach for Better Feature Representation
por: Lee, Simon A., et al.
Publicado: (2024)
por: Lee, Simon A., et al.
Publicado: (2024)
Advancing Arabic Reverse Dictionary Systems: A Transformer-Based Approach with Dataset Construction Guidelines
por: Sibaee, Serry, et al.
Publicado: (2025)
por: Sibaee, Serry, et al.
Publicado: (2025)
Token-Efficient Leverage Learning in Large Language Models
por: Zeng, Yuanhao, et al.
Publicado: (2024)
por: Zeng, Yuanhao, et al.
Publicado: (2024)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2025)
por: Chanin, David, et al.
Publicado: (2025)
Do pretrained Transformers Learn In-Context by Gradient Descent?
por: Shen, Lingfeng, et al.
Publicado: (2023)
por: Shen, Lingfeng, et al.
Publicado: (2023)
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning
por: Braun, Dan, et al.
Publicado: (2024)
por: Braun, Dan, et al.
Publicado: (2024)
Curiosity-Critic: Cumulative Prediction Error Improvement as a Tractable Intrinsic Reward for World Model Training
por: Bhaskara, Vin, et al.
Publicado: (2026)
por: Bhaskara, Vin, et al.
Publicado: (2026)
Model Merging by Uncertainty-Based Gradient Matching
por: Daheim, Nico, et al.
Publicado: (2023)
por: Daheim, Nico, et al.
Publicado: (2023)
Learning to Interpret Weight Differences in Language Models
por: Goel, Avichal, et al.
Publicado: (2025)
por: Goel, Avichal, et al.
Publicado: (2025)
Ask more, know better: Reinforce-Learned Prompt Questions for Decision Making with Large Language Models
por: Yan, Xue, et al.
Publicado: (2023)
por: Yan, Xue, et al.
Publicado: (2023)
Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
por: Yang, Shiping, et al.
Publicado: (2025)
por: Yang, Shiping, et al.
Publicado: (2025)
Linear-Time Demonstration Selection for In-Context Learning via Gradient Estimation
por: Zhang, Ziniu, et al.
Publicado: (2025)
por: Zhang, Ziniu, et al.
Publicado: (2025)
GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning
por: Yang, Ningyuan, et al.
Publicado: (2026)
por: Yang, Ningyuan, et al.
Publicado: (2026)
FLUID-LLM: Learning Computational Fluid Dynamics with Spatiotemporal-aware Large Language Models
por: Zhu, Max, et al.
Publicado: (2024)
por: Zhu, Max, et al.
Publicado: (2024)
Why Larger Language Models Do In-context Learning Differently?
por: Shi, Zhenmei, et al.
Publicado: (2024)
por: Shi, Zhenmei, et al.
Publicado: (2024)
Sparse Autoencoder Features for Classifications and Transferability
por: Gallifant, Jack, et al.
Publicado: (2025)
por: Gallifant, Jack, et al.
Publicado: (2025)
Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary Learning
por: Wu, John, et al.
Publicado: (2024)
por: Wu, John, et al.
Publicado: (2024)
Ejemplares similares
-
Language models struggle with compartmentalization
por: Howe, Thomas Vincent, et al.
Publicado: (2026) -
Arti-"fickle" Intelligence: Using LLMs as a Tool for Inference in the Political and Social Sciences
por: Argyle, Lisa P., et al.
Publicado: (2025) -
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
por: Karvonen, Adam, et al.
Publicado: (2024) -
Scaling Law with Learning Rate Annealing
por: Tissue, Howe, et al.
Publicado: (2024) -
Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
por: Zhussip, Magauiya, et al.
Publicado: (2025)