Out-of-distribution generalization via composition: a lens through induction heads in Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Jiajun, Xu, Zhuoyan, Zhong, Yiqiao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uncovering hidden geometry in Transformers via disentangling position and context
von: Song, Jiajun, et al.
Veröffentlicht: (2023)
von: Song, Jiajun, et al.
Veröffentlicht: (2023)
Do Large Language Models Have Compositional Ability? An Investigation into Limitations and Scalability
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2024)
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2024)
Why Larger Language Models Do In-context Learning Differently?
von: Shi, Zhenmei, et al.
Veröffentlicht: (2024)
von: Shi, Zhenmei, et al.
Veröffentlicht: (2024)
Towards Few-Shot Adaptation of Foundation Models via Multitask Finetuning
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2024)
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2024)
LLM generation novelty through the lens of semantic similarity
von: Davydov, Philipp, et al.
Veröffentlicht: (2025)
von: Davydov, Philipp, et al.
Veröffentlicht: (2025)
Explicit Multi-head Attention for Inter-head Interaction in Large Language Models
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
Your Finetuned Large Language Model is Already a Powerful Out-of-distribution Detector
von: Zhang, Andi, et al.
Veröffentlicht: (2024)
von: Zhang, Andi, et al.
Veröffentlicht: (2024)
TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization
von: Fu, Lucheng, et al.
Veröffentlicht: (2026)
von: Fu, Lucheng, et al.
Veröffentlicht: (2026)
Algorithmic Capabilities of Random Transformers
von: Zhong, Ziqian, et al.
Veröffentlicht: (2024)
von: Zhong, Ziqian, et al.
Veröffentlicht: (2024)
An explainable transformer circuit for compositional generalization
von: Tang, Cheng, et al.
Veröffentlicht: (2025)
von: Tang, Cheng, et al.
Veröffentlicht: (2025)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
Hymba: A Hybrid-head Architecture for Small Language Models
von: Dong, Xin, et al.
Veröffentlicht: (2024)
von: Dong, Xin, et al.
Veröffentlicht: (2024)
Robust Neural Information Retrieval: An Adversarial and Out-of-distribution Perspective
von: Liu, Yu-An, et al.
Veröffentlicht: (2024)
von: Liu, Yu-An, et al.
Veröffentlicht: (2024)
Teaching Transformers Causal Reasoning through Axiomatic Training
von: Vashishtha, Aniket, et al.
Veröffentlicht: (2024)
von: Vashishtha, Aniket, et al.
Veröffentlicht: (2024)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
Sysformer: Safeguarding Frozen Large Language Models with Adaptive System Prompts
von: Sharma, Kartik, et al.
Veröffentlicht: (2025)
von: Sharma, Kartik, et al.
Veröffentlicht: (2025)
Structural Rationale Distillation via Reasoning Space Compression
von: Yang, Jialin, et al.
Veröffentlicht: (2026)
von: Yang, Jialin, et al.
Veröffentlicht: (2026)
KVCrush: Key value cache size-reduction using similarity in head-behaviour
von: Jha, Gopi Krishna, et al.
Veröffentlicht: (2025)
von: Jha, Gopi Krishna, et al.
Veröffentlicht: (2025)
PMET: Precise Model Editing in a Transformer
von: Li, Xiaopeng, et al.
Veröffentlicht: (2023)
von: Li, Xiaopeng, et al.
Veröffentlicht: (2023)
Towards Infinite-Long Prefix in Transformer
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient
von: Li, Changming, et al.
Veröffentlicht: (2026)
von: Li, Changming, et al.
Veröffentlicht: (2026)
Understanding Transformers via N-gram Statistics
von: Nguyen, Timothy
Veröffentlicht: (2024)
von: Nguyen, Timothy
Veröffentlicht: (2024)
Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
Out-of-Distribution Detection using Synthetic Data Generation
von: Abbas, Momin, et al.
Veröffentlicht: (2025)
von: Abbas, Momin, et al.
Veröffentlicht: (2025)
PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation
von: He, Jiajun, et al.
Veröffentlicht: (2025)
von: He, Jiajun, et al.
Veröffentlicht: (2025)
Forgetting Transformer: Softmax Attention with a Forget Gate
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
UniSD: Towards a Unified Self-Distillation Framework for Large Language Models
von: Jin, Yiqiao, et al.
Veröffentlicht: (2026)
von: Jin, Yiqiao, et al.
Veröffentlicht: (2026)
TransformLLM: Adapting Large Language Models via LLM-Transformed Reading Comprehension Text
von: Arbel, Iftach, et al.
Veröffentlicht: (2024)
von: Arbel, Iftach, et al.
Veröffentlicht: (2024)
MetaGreen: Meta-Learning Inspired Transformer Selection for Green Semantic Communication
von: Mukherjee, Shubhabrata, et al.
Veröffentlicht: (2024)
von: Mukherjee, Shubhabrata, et al.
Veröffentlicht: (2024)
The Belief State Transformer
von: Hu, Edward S., et al.
Veröffentlicht: (2024)
von: Hu, Edward S., et al.
Veröffentlicht: (2024)
Power Transformer Fault Prediction Based on Knowledge Graphs
von: Wang, Chao, et al.
Veröffentlicht: (2024)
von: Wang, Chao, et al.
Veröffentlicht: (2024)
Reparameterized LLM Training via Orthogonal Equivalence Transformation
von: Qiu, Zeju, et al.
Veröffentlicht: (2025)
von: Qiu, Zeju, et al.
Veröffentlicht: (2025)
Accelerating Transformer Inference for Translation via Parallel Decoding
von: Santilli, Andrea, et al.
Veröffentlicht: (2023)
von: Santilli, Andrea, et al.
Veröffentlicht: (2023)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
von: Choi, Euntae, et al.
Veröffentlicht: (2025)
von: Choi, Euntae, et al.
Veröffentlicht: (2025)
Cartridges: Lightweight and general-purpose long context representations via self-study
von: Eyuboglu, Sabri, et al.
Veröffentlicht: (2025)
von: Eyuboglu, Sabri, et al.
Veröffentlicht: (2025)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
Out-of-Domain Intent Detection Considering Multi-Turn Dialogue Contexts
von: Lang, Hao, et al.
Veröffentlicht: (2023)
von: Lang, Hao, et al.
Veröffentlicht: (2023)
Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
Adaptive Computation Pruning for the Forgetting Transformer
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Uncovering hidden geometry in Transformers via disentangling position and context
von: Song, Jiajun, et al.
Veröffentlicht: (2023) -
Do Large Language Models Have Compositional Ability? An Investigation into Limitations and Scalability
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2024) -
Why Larger Language Models Do In-context Learning Differently?
von: Shi, Zhenmei, et al.
Veröffentlicht: (2024) -
Towards Few-Shot Adaptation of Foundation Models via Multitask Finetuning
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2024) -
LLM generation novelty through the lens of semantic similarity
von: Davydov, Philipp, et al.
Veröffentlicht: (2025)