How Transformers Learn In-Context Recall Tasks? Optimality, Training Dynamics and Generalization
Fuente:
arXiv
Salvato in:
| Autori principali: | Nguyen, Quan, Nguyen-Tang, Thanh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On Sample-Efficient Offline Reinforcement Learning: Data Diversity, Posterior Sampling, and Beyond
di: Nguyen-Tang, Thanh, et al.
Pubblicazione: (2024)
di: Nguyen-Tang, Thanh, et al.
Pubblicazione: (2024)
Revisiting LARS for Large Batch Training Generalization of Neural Networks
di: Do, Khoi, et al.
Pubblicazione: (2023)
di: Do, Khoi, et al.
Pubblicazione: (2023)
Neural ODE Transformers: Analyzing Internal Dynamics and Adaptive Fine-tuning
di: Tong, Anh, et al.
Pubblicazione: (2025)
di: Tong, Anh, et al.
Pubblicazione: (2025)
Generative Conditional Distributions by Neural (Entropic) Optimal Transport
di: Nguyen, Bao, et al.
Pubblicazione: (2024)
di: Nguyen, Bao, et al.
Pubblicazione: (2024)
On The Statistical Complexity of Offline Decision-Making
di: Nguyen-Tang, Thanh, et al.
Pubblicazione: (2025)
di: Nguyen-Tang, Thanh, et al.
Pubblicazione: (2025)
Enhancing Time Series Forecasting via a Parallel Hybridization of ARIMA and Polynomial Classifiers
di: Nguyen, Thanh Son, et al.
Pubblicazione: (2025)
di: Nguyen, Thanh Son, et al.
Pubblicazione: (2025)
Legal2LogicICL: Improving Generalization in Transforming Legal Cases to Logical Formulas via Diverse Few-Shot Learning
di: Xue, Jieying, et al.
Pubblicazione: (2026)
di: Xue, Jieying, et al.
Pubblicazione: (2026)
Learning in Markov Games with Adaptive Adversaries: Policy Regret, Fundamental Barriers, and Efficient Algorithms
di: Nguyen-Tang, Thanh, et al.
Pubblicazione: (2024)
di: Nguyen-Tang, Thanh, et al.
Pubblicazione: (2024)
Learning Reconfigurable Representations for Multimodal Federated Learning with Missing Data
di: Nguyen, Duong M., et al.
Pubblicazione: (2025)
di: Nguyen, Duong M., et al.
Pubblicazione: (2025)
Generative Pre-Trained Transformer for Symbolic Regression Base In-Context Reinforcement Learning
di: Li, Yanjie, et al.
Pubblicazione: (2024)
di: Li, Yanjie, et al.
Pubblicazione: (2024)
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization
di: Nguyen, Thanh Thi, et al.
Pubblicazione: (2025)
di: Nguyen, Thanh Thi, et al.
Pubblicazione: (2025)
A Framework for Quantifying How Pre-Training and Context Benefit In-Context Learning
di: Song, Bingqing, et al.
Pubblicazione: (2025)
di: Song, Bingqing, et al.
Pubblicazione: (2025)
Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence
di: Nguyen, Tien-Phat, et al.
Pubblicazione: (2026)
di: Nguyen, Tien-Phat, et al.
Pubblicazione: (2026)
Online Optimization for Offline Safe Reinforcement Learning
di: Chemingui, Yassine, et al.
Pubblicazione: (2025)
di: Chemingui, Yassine, et al.
Pubblicazione: (2025)
Regret-Based Defense in Adversarial Reinforcement Learning
di: Belaire, Roman, et al.
Pubblicazione: (2023)
di: Belaire, Roman, et al.
Pubblicazione: (2023)
Symmetry Reveals Layerwise Dynamics: How Transformers Perform In-Context Classification
di: Lutz, Patrick, et al.
Pubblicazione: (2026)
di: Lutz, Patrick, et al.
Pubblicazione: (2026)
Task-Focused Consolidation with Spaced Recall: Making Neural Networks Learn like College Students
di: Bamnodkar, Prital
Pubblicazione: (2025)
di: Bamnodkar, Prital
Pubblicazione: (2025)
Transformers Learn Robust In-Context Regression under Distributional Uncertainty
di: Cao, Hoang T. H., et al.
Pubblicazione: (2026)
di: Cao, Hoang T. H., et al.
Pubblicazione: (2026)
BSO: Safety Alignment Is Density Ratio Matching
di: Nguyen, Tien-Phat, et al.
Pubblicazione: (2026)
di: Nguyen, Tien-Phat, et al.
Pubblicazione: (2026)
Meta-Learning Transformers to Improve In-Context Generalization
di: Braccaioli, Lorenzo, et al.
Pubblicazione: (2025)
di: Braccaioli, Lorenzo, et al.
Pubblicazione: (2025)
An Introduction to Sliced Optimal Transport
di: Nguyen, Khai
Pubblicazione: (2025)
di: Nguyen, Khai
Pubblicazione: (2025)
A Survey of Machine Unlearning
di: Nguyen, Thanh Tam, et al.
Pubblicazione: (2022)
di: Nguyen, Thanh Tam, et al.
Pubblicazione: (2022)
Study of Training Dynamics for Memory-Constrained Fine-Tuning
di: Quélennec, Aël, et al.
Pubblicazione: (2025)
di: Quélennec, Aël, et al.
Pubblicazione: (2025)
Fake Advertisements Detection Using Automated Multimodal Learning: A Case Study for Vietnamese Real Estate Data
di: Nguyen, Duy, et al.
Pubblicazione: (2025)
di: Nguyen, Duy, et al.
Pubblicazione: (2025)
Amortized Optimal Transport from Sliced Potentials
di: Truong, Minh-Phuc, et al.
Pubblicazione: (2026)
di: Truong, Minh-Phuc, et al.
Pubblicazione: (2026)
BOLIMES: Boruta and LIME optiMized fEature Selection for Gene Expression Classification
di: Phan, Bich-Chung, et al.
Pubblicazione: (2025)
di: Phan, Bich-Chung, et al.
Pubblicazione: (2025)
Fast-FedUL: A Training-Free Federated Unlearning with Provable Skew Resilience
di: Huynh, Thanh Trung, et al.
Pubblicazione: (2024)
di: Huynh, Thanh Trung, et al.
Pubblicazione: (2024)
Adaptive Knowledge Distillation for Classification of Hand Images using Explainable Vision Transformers
di: Nguyen, Thanh Thi, et al.
Pubblicazione: (2024)
di: Nguyen, Thanh Thi, et al.
Pubblicazione: (2024)
Mimic In-Context Learning for Multimodal Tasks
di: Jiang, Yuchu, et al.
Pubblicazione: (2025)
di: Jiang, Yuchu, et al.
Pubblicazione: (2025)
Adaptive Correlation-Weighted Intrinsic Rewards for Reinforcement Learning
di: Nguyen, Viet Bac, et al.
Pubblicazione: (2026)
di: Nguyen, Viet Bac, et al.
Pubblicazione: (2026)
Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks
di: Kim, Jeongmo, et al.
Pubblicazione: (2025)
di: Kim, Jeongmo, et al.
Pubblicazione: (2025)
Causal-Aware Generative Adversarial Networks with Reinforcement Learning
di: Nguyen, Tu Anh Hoang, et al.
Pubblicazione: (2025)
di: Nguyen, Tu Anh Hoang, et al.
Pubblicazione: (2025)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks
di: Pink, Mathis, et al.
Pubblicazione: (2024)
di: Pink, Mathis, et al.
Pubblicazione: (2024)
Agents Learn Their Runtime: Interpreter Persistence as Training-Time Semantics
di: May, Victor, et al.
Pubblicazione: (2026)
di: May, Victor, et al.
Pubblicazione: (2026)
Understanding Transformers via N-gram Statistics
di: Nguyen, Timothy
Pubblicazione: (2024)
di: Nguyen, Timothy
Pubblicazione: (2024)
Disentangling Recall and Reasoning in Transformer Models through Layer-wise Attention and Activation Analysis
di: Fartale, Harshwardhan, et al.
Pubblicazione: (2025)
di: Fartale, Harshwardhan, et al.
Pubblicazione: (2025)
FastDiSS: Few-step Match Many-step Diffusion Language Model on Sequence-to-Sequence Generation--Full Version
di: Nguyen-Cong, Dat, et al.
Pubblicazione: (2026)
di: Nguyen-Cong, Dat, et al.
Pubblicazione: (2026)
ViCLSR: A Supervised Contrastive Learning Framework with Natural Language Inference for Natural Language Understanding Tasks
di: Van Huynh, Tin, et al.
Pubblicazione: (2026)
di: Van Huynh, Tin, et al.
Pubblicazione: (2026)
Towards Robust Policy: Enhancing Offline Reinforcement Learning with Adversarial Attacks and Defenses
di: Nguyen, Thanh, et al.
Pubblicazione: (2024)
di: Nguyen, Thanh, et al.
Pubblicazione: (2024)
Documenti analoghi
-
On Sample-Efficient Offline Reinforcement Learning: Data Diversity, Posterior Sampling, and Beyond
di: Nguyen-Tang, Thanh, et al.
Pubblicazione: (2024) -
Revisiting LARS for Large Batch Training Generalization of Neural Networks
di: Do, Khoi, et al.
Pubblicazione: (2023) -
Neural ODE Transformers: Analyzing Internal Dynamics and Adaptive Fine-tuning
di: Tong, Anh, et al.
Pubblicazione: (2025) -
Generative Conditional Distributions by Neural (Entropic) Optimal Transport
di: Nguyen, Bao, et al.
Pubblicazione: (2024) -
On The Statistical Complexity of Offline Decision-Making
di: Nguyen-Tang, Thanh, et al.
Pubblicazione: (2025)