Out-of-distribution generalization via composition: a lens through induction heads in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Jiajun, Xu, Zhuoyan, Zhong, Yiqiao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncovering hidden geometry in Transformers via disentangling position and context
by: Song, Jiajun, et al.
Published: (2023)
by: Song, Jiajun, et al.
Published: (2023)
Do Large Language Models Have Compositional Ability? An Investigation into Limitations and Scalability
by: Xu, Zhuoyan, et al.
Published: (2024)
by: Xu, Zhuoyan, et al.
Published: (2024)
Why Larger Language Models Do In-context Learning Differently?
by: Shi, Zhenmei, et al.
Published: (2024)
by: Shi, Zhenmei, et al.
Published: (2024)
Towards Few-Shot Adaptation of Foundation Models via Multitask Finetuning
by: Xu, Zhuoyan, et al.
Published: (2024)
by: Xu, Zhuoyan, et al.
Published: (2024)
LLM generation novelty through the lens of semantic similarity
by: Davydov, Philipp, et al.
Published: (2025)
by: Davydov, Philipp, et al.
Published: (2025)
Explicit Multi-head Attention for Inter-head Interaction in Large Language Models
by: Peng, Runyu, et al.
Published: (2026)
by: Peng, Runyu, et al.
Published: (2026)
Your Finetuned Large Language Model is Already a Powerful Out-of-distribution Detector
by: Zhang, Andi, et al.
Published: (2024)
by: Zhang, Andi, et al.
Published: (2024)
TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization
by: Fu, Lucheng, et al.
Published: (2026)
by: Fu, Lucheng, et al.
Published: (2026)
Algorithmic Capabilities of Random Transformers
by: Zhong, Ziqian, et al.
Published: (2024)
by: Zhong, Ziqian, et al.
Published: (2024)
An explainable transformer circuit for compositional generalization
by: Tang, Cheng, et al.
Published: (2025)
by: Tang, Cheng, et al.
Published: (2025)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
by: Xu, Bingxin, et al.
Published: (2025)
by: Xu, Bingxin, et al.
Published: (2025)
Hymba: A Hybrid-head Architecture for Small Language Models
by: Dong, Xin, et al.
Published: (2024)
by: Dong, Xin, et al.
Published: (2024)
Robust Neural Information Retrieval: An Adversarial and Out-of-distribution Perspective
by: Liu, Yu-An, et al.
Published: (2024)
by: Liu, Yu-An, et al.
Published: (2024)
Teaching Transformers Causal Reasoning through Axiomatic Training
by: Vashishtha, Aniket, et al.
Published: (2024)
by: Vashishtha, Aniket, et al.
Published: (2024)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
by: Oh, Sejoon, et al.
Published: (2024)
by: Oh, Sejoon, et al.
Published: (2024)
Sysformer: Safeguarding Frozen Large Language Models with Adaptive System Prompts
by: Sharma, Kartik, et al.
Published: (2025)
by: Sharma, Kartik, et al.
Published: (2025)
Structural Rationale Distillation via Reasoning Space Compression
by: Yang, Jialin, et al.
Published: (2026)
by: Yang, Jialin, et al.
Published: (2026)
KVCrush: Key value cache size-reduction using similarity in head-behaviour
by: Jha, Gopi Krishna, et al.
Published: (2025)
by: Jha, Gopi Krishna, et al.
Published: (2025)
PMET: Precise Model Editing in a Transformer
by: Li, Xiaopeng, et al.
Published: (2023)
by: Li, Xiaopeng, et al.
Published: (2023)
Towards Infinite-Long Prefix in Transformer
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient
by: Li, Changming, et al.
Published: (2026)
by: Li, Changming, et al.
Published: (2026)
Understanding Transformers via N-gram Statistics
by: Nguyen, Timothy
Published: (2024)
by: Nguyen, Timothy
Published: (2024)
Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning
by: Wang, Yiming, et al.
Published: (2024)
by: Wang, Yiming, et al.
Published: (2024)
Out-of-Distribution Detection using Synthetic Data Generation
by: Abbas, Momin, et al.
Published: (2025)
by: Abbas, Momin, et al.
Published: (2025)
PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation
by: He, Jiajun, et al.
Published: (2025)
by: He, Jiajun, et al.
Published: (2025)
Forgetting Transformer: Softmax Attention with a Forget Gate
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
UniSD: Towards a Unified Self-Distillation Framework for Large Language Models
by: Jin, Yiqiao, et al.
Published: (2026)
by: Jin, Yiqiao, et al.
Published: (2026)
TransformLLM: Adapting Large Language Models via LLM-Transformed Reading Comprehension Text
by: Arbel, Iftach, et al.
Published: (2024)
by: Arbel, Iftach, et al.
Published: (2024)
MetaGreen: Meta-Learning Inspired Transformer Selection for Green Semantic Communication
by: Mukherjee, Shubhabrata, et al.
Published: (2024)
by: Mukherjee, Shubhabrata, et al.
Published: (2024)
The Belief State Transformer
by: Hu, Edward S., et al.
Published: (2024)
by: Hu, Edward S., et al.
Published: (2024)
Power Transformer Fault Prediction Based on Knowledge Graphs
by: Wang, Chao, et al.
Published: (2024)
by: Wang, Chao, et al.
Published: (2024)
Reparameterized LLM Training via Orthogonal Equivalence Transformation
by: Qiu, Zeju, et al.
Published: (2025)
by: Qiu, Zeju, et al.
Published: (2025)
Accelerating Transformer Inference for Translation via Parallel Decoding
by: Santilli, Andrea, et al.
Published: (2023)
by: Santilli, Andrea, et al.
Published: (2023)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
by: Casademunt, Helena, et al.
Published: (2025)
by: Casademunt, Helena, et al.
Published: (2025)
Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
by: Choi, Euntae, et al.
Published: (2025)
by: Choi, Euntae, et al.
Published: (2025)
Cartridges: Lightweight and general-purpose long context representations via self-study
by: Eyuboglu, Sabri, et al.
Published: (2025)
by: Eyuboglu, Sabri, et al.
Published: (2025)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
Out-of-Domain Intent Detection Considering Multi-Turn Dialogue Contexts
by: Lang, Hao, et al.
Published: (2023)
by: Lang, Hao, et al.
Published: (2023)
Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
Adaptive Computation Pruning for the Forgetting Transformer
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
Similar Items
-
Uncovering hidden geometry in Transformers via disentangling position and context
by: Song, Jiajun, et al.
Published: (2023) -
Do Large Language Models Have Compositional Ability? An Investigation into Limitations and Scalability
by: Xu, Zhuoyan, et al.
Published: (2024) -
Why Larger Language Models Do In-context Learning Differently?
by: Shi, Zhenmei, et al.
Published: (2024) -
Towards Few-Shot Adaptation of Foundation Models via Multitask Finetuning
by: Xu, Zhuoyan, et al.
Published: (2024) -
LLM generation novelty through the lens of semantic similarity
by: Davydov, Philipp, et al.
Published: (2025)