Understanding In-context Learning of Addition via Activation Subspaces
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Xinyan, Yin, Kayo, Jordan, Michael I., Steinhardt, Jacob, Chen, Lijie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Which Attention Heads Matter for In-Context Learning?
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
Learning a Generative Meta-Model of LLM Activations
von: Luo, Grace, et al.
Veröffentlicht: (2026)
von: Luo, Grace, et al.
Veröffentlicht: (2026)
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
How do Language Models Bind Entities in Context?
von: Feng, Jiahai, et al.
Veröffentlicht: (2023)
von: Feng, Jiahai, et al.
Veröffentlicht: (2023)
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
von: Ye, Yaowen, et al.
Veröffentlicht: (2025)
von: Ye, Yaowen, et al.
Veröffentlicht: (2025)
Steering Llama 2 via Contrastive Activation Addition
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
Feedback Loops With Language Models Drive In-Context Reward Hacking
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
Explaining Datasets in Words: Statistical Models with Natural Language Parameters
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
Discovering Latent Knowledge in Language Models Without Supervision
von: Burns, Collin, et al.
Veröffentlicht: (2022)
von: Burns, Collin, et al.
Veröffentlicht: (2022)
Training Language Models to Explain Their Own Computations
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
Long Chain-of-Thought Reasoning Across Languages
von: Barua, Josh, et al.
Veröffentlicht: (2025)
von: Barua, Josh, et al.
Veröffentlicht: (2025)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
von: Huang, Xinting, et al.
Veröffentlicht: (2025)
von: Huang, Xinting, et al.
Veröffentlicht: (2025)
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
Approaching Human-Level Forecasting with Language Models
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
Resolving Conflicts in Lifelong Learning via Aligning Updates in Subspaces
von: Zhou, Yueer, et al.
Veröffentlicht: (2025)
von: Zhou, Yueer, et al.
Veröffentlicht: (2025)
Understanding Reasoning in Chain-of-Thought from the Hopfieldian View
von: Hu, Lijie, et al.
Veröffentlicht: (2024)
von: Hu, Lijie, et al.
Veröffentlicht: (2024)
Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026)
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026)
Enhancing In-context Learning via Linear Probe Calibration
von: Abbas, Momin, et al.
Veröffentlicht: (2024)
von: Abbas, Momin, et al.
Veröffentlicht: (2024)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
von: Hu, Michael Y., et al.
Veröffentlicht: (2025)
von: Hu, Michael Y., et al.
Veröffentlicht: (2025)
Private Language Models via Truncated Laplacian Mechanism
von: Huang, Tianhao, et al.
Veröffentlicht: (2024)
von: Huang, Tianhao, et al.
Veröffentlicht: (2024)
ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention
von: Wang, Xinyan, et al.
Veröffentlicht: (2026)
von: Wang, Xinyan, et al.
Veröffentlicht: (2026)
On the Relation between Sensitivity and Accuracy in In-context Learning
von: Chen, Yanda, et al.
Veröffentlicht: (2022)
von: Chen, Yanda, et al.
Veröffentlicht: (2022)
Implicit In-context Learning
von: Li, Zhuowei, et al.
Veröffentlicht: (2024)
von: Li, Zhuowei, et al.
Veröffentlicht: (2024)
Eliciting Language Model Behaviors with Investigator Agents
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2025)
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2025)
Sub-SA: Strengthen In-context Learning via Submodular Selective Annotation
von: Qian, Jian, et al.
Veröffentlicht: (2024)
von: Qian, Jian, et al.
Veröffentlicht: (2024)
Memory-Efficient LLM Training with Online Subspace Descent
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024)
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024)
SPARC: Subspace-Aware Prompt Adaptation for Robust Continual Learning in LLMs
von: Jayasuriya, Dinithi, et al.
Veröffentlicht: (2025)
von: Jayasuriya, Dinithi, et al.
Veröffentlicht: (2025)
Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
von: Islam, Sekh Mainul, et al.
Veröffentlicht: (2025)
von: Islam, Sekh Mainul, et al.
Veröffentlicht: (2025)
Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
von: Wang, Dianyun, et al.
Veröffentlicht: (2025)
von: Wang, Dianyun, et al.
Veröffentlicht: (2025)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
Mechanistic Fine-tuning for In-context Learning
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
A Study on the Calibration of In-context Learning
von: Zhang, Hanlin, et al.
Veröffentlicht: (2023)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2023)
Deep Dense Exploration for LLM Reinforcement Learning via Pivot-Driven Resampling
von: Guo, Yiran, et al.
Veröffentlicht: (2026)
von: Guo, Yiran, et al.
Veröffentlicht: (2026)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
von: Li, Yingcong, et al.
Veröffentlicht: (2025)
von: Li, Yingcong, et al.
Veröffentlicht: (2025)
DRESSing Up LLM: Efficient Stylized Question-Answering via Style Subspace Editing
von: Ma, Xinyu, et al.
Veröffentlicht: (2025)
von: Ma, Xinyu, et al.
Veröffentlicht: (2025)
In-context Autoencoder for Context Compression in a Large Language Model
von: Ge, Tao, et al.
Veröffentlicht: (2023)
von: Ge, Tao, et al.
Veröffentlicht: (2023)
Mechanism of Task-oriented Information Removal in In-context Learning
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Which Attention Heads Matter for In-Context Learning?
von: Yin, Kayo, et al.
Veröffentlicht: (2025) -
Learning a Generative Meta-Model of LLM Activations
von: Luo, Grace, et al.
Veröffentlicht: (2026) -
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
von: Halawi, Danny, et al.
Veröffentlicht: (2023) -
How do Language Models Bind Entities in Context?
von: Feng, Jiahai, et al.
Veröffentlicht: (2023) -
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
von: Pan, Alexander, et al.
Veröffentlicht: (2024)