Learning a Generative Meta-Model of LLM Activations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Grace, Feng, Jiahai, Darrell, Trevor, Radford, Alec, Steinhardt, Jacob |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How do Language Models Bind Entities in Context?
von: Feng, Jiahai, et al.
Veröffentlicht: (2023)
von: Feng, Jiahai, et al.
Veröffentlicht: (2023)
Extractive Structures Learned in Pretraining Enable Generalization on Finetuned Facts
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
Monitoring Latent World States in Language Models with Propositional Probes
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
Which Attention Heads Matter for In-Context Learning?
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
Shaping capabilities with token-level data filtering
von: Rathi, Neil, et al.
Veröffentlicht: (2026)
von: Rathi, Neil, et al.
Veröffentlicht: (2026)
Understanding In-context Learning of Addition via Activation Subspaces
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
Discovering Latent Knowledge in Language Models Without Supervision
von: Burns, Collin, et al.
Veröffentlicht: (2022)
von: Burns, Collin, et al.
Veröffentlicht: (2022)
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
von: Ye, Yaowen, et al.
Veröffentlicht: (2025)
von: Ye, Yaowen, et al.
Veröffentlicht: (2025)
Feedback Loops With Language Models Drive In-Context Reward Hacking
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
Explaining Datasets in Words: Statistical Models with Natural Language Parameters
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
Training Language Models to Explain Their Own Computations
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models
von: Dunlap, Lisa, et al.
Veröffentlicht: (2024)
von: Dunlap, Lisa, et al.
Veröffentlicht: (2024)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
Approaching Human-Level Forecasting with Language Models
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
Eliciting Language Model Behaviors with Investigator Agents
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2025)
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2025)
Transformer Explainer: Interactive Learning of Text-Generative Models
von: Cho, Aeree, et al.
Veröffentlicht: (2024)
von: Cho, Aeree, et al.
Veröffentlicht: (2024)
Vision-Language Models Create Cross-Modal Task Representations
von: Luo, Grace, et al.
Veröffentlicht: (2024)
von: Luo, Grace, et al.
Veröffentlicht: (2024)
Enough Coin Flips Can Make LLMs Act Bayesian
von: Gupta, Ritwik, et al.
Veröffentlicht: (2025)
von: Gupta, Ritwik, et al.
Veröffentlicht: (2025)
Towards Scalable Meta-Learning of near-optimal Interpretable Models via Synthetic Model Generations
von: Myint, Kyaw Hpone, et al.
Veröffentlicht: (2025)
von: Myint, Kyaw Hpone, et al.
Veröffentlicht: (2025)
Turning LLM Activations Quantization-Friendly
von: Czakó, Patrik, et al.
Veröffentlicht: (2025)
von: Czakó, Patrik, et al.
Veröffentlicht: (2025)
Dual-Process Image Generation
von: Luo, Grace, et al.
Veröffentlicht: (2025)
von: Luo, Grace, et al.
Veröffentlicht: (2025)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models
von: Lian, Long, et al.
Veröffentlicht: (2025)
von: Lian, Long, et al.
Veröffentlicht: (2025)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
von: Lee, Heekyung, et al.
Veröffentlicht: (2025)
von: Lee, Heekyung, et al.
Veröffentlicht: (2025)
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning
von: Shan, Zikang, et al.
Veröffentlicht: (2026)
von: Shan, Zikang, et al.
Veröffentlicht: (2026)
LLM Attributor: Interactive Visual Attribution for LLM Generation
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
Steer Like the LLM: Activation Steering that Mimics Prompting
von: Heyman, Geert, et al.
Veröffentlicht: (2026)
von: Heyman, Geert, et al.
Veröffentlicht: (2026)
Measuring the Depth of LLM Unlearning via Activation Patching
von: Lee, Jaeung, et al.
Veröffentlicht: (2026)
von: Lee, Jaeung, et al.
Veröffentlicht: (2026)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
von: Shang, Chuyi, et al.
Veröffentlicht: (2024)
von: Shang, Chuyi, et al.
Veröffentlicht: (2024)
Learning to Generalize Unseen Domains via Multi-Source Meta Learning for Text Classification
von: Hu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2024)
MPO: Boosting LLM Agents with Meta Plan Optimization
von: Xiong, Weimin, et al.
Veröffentlicht: (2025)
von: Xiong, Weimin, et al.
Veröffentlicht: (2025)
Learn-to-Distance: Distance Learning for Detecting LLM-Generated Text
von: Zhou, Hongyi, et al.
Veröffentlicht: (2026)
von: Zhou, Hongyi, et al.
Veröffentlicht: (2026)
AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance
von: He, Lixuan, et al.
Veröffentlicht: (2025)
von: He, Lixuan, et al.
Veröffentlicht: (2025)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
Detecting PTSD in Clinical Interviews: A Comparative Analysis of NLP Methods and Large Language Models
von: Chen, Feng, et al.
Veröffentlicht: (2025)
von: Chen, Feng, et al.
Veröffentlicht: (2025)
A Meta-Learning Perspective on Transformers for Causal Language Modeling
von: Wu, Xinbo, et al.
Veröffentlicht: (2023)
von: Wu, Xinbo, et al.
Veröffentlicht: (2023)
Re-evaluating the Need for Multimodal Signals in Unsupervised Grammar Induction
von: Li, Boyi, et al.
Veröffentlicht: (2022)
von: Li, Boyi, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
How do Language Models Bind Entities in Context?
von: Feng, Jiahai, et al.
Veröffentlicht: (2023) -
Extractive Structures Learned in Pretraining Enable Generalization on Finetuned Facts
von: Feng, Jiahai, et al.
Veröffentlicht: (2024) -
Monitoring Latent World States in Language Models with Propositional Probes
von: Feng, Jiahai, et al.
Veröffentlicht: (2024) -
Which Attention Heads Matter for In-Context Learning?
von: Yin, Kayo, et al.
Veröffentlicht: (2025) -
Shaping capabilities with token-level data filtering
von: Rathi, Neil, et al.
Veröffentlicht: (2026)