Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High Dimensions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Demir, Samet, Dogan, Zafer |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
von: Demir, Samet, et al.
Veröffentlicht: (2025)
von: Demir, Samet, et al.
Veröffentlicht: (2025)
Input-Label Correlation Governs a Linear-to-Nonlinear Transition in Random Features under Spiked Covariance
von: Demir, Samet, et al.
Veröffentlicht: (2024)
von: Demir, Samet, et al.
Veröffentlicht: (2024)
Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure
von: Demir, Samet, et al.
Veröffentlicht: (2025)
von: Demir, Samet, et al.
Veröffentlicht: (2025)
Implicitly Normalized Online PCA: A Regularized Algorithm with Exact High-Dimensional Dynamics
von: Demir, Samet, et al.
Veröffentlicht: (2025)
von: Demir, Samet, et al.
Veröffentlicht: (2025)
Asymptotic Study of In-context Learning with Random Transformers through Equivalent Models
von: Demir, Samet, et al.
Veröffentlicht: (2025)
von: Demir, Samet, et al.
Veröffentlicht: (2025)
Learning Rate Should Scale Inversely with High-Order Data Moments in High-Dimensional Online Independent Component Analysis
von: Gultekin, M. Oguzhan, et al.
Veröffentlicht: (2025)
von: Gultekin, M. Oguzhan, et al.
Veröffentlicht: (2025)
Learning Beyond the Gaussian Data: Learning Dynamics of Neural Networks on an Expressive and Cumulant-Controllable Data Model
von: Ure, Onat, et al.
Veröffentlicht: (2026)
von: Ure, Onat, et al.
Veröffentlicht: (2026)
Benefits of Online Tilted Empirical Risk Minimization: A Case Study of Outlier Detection and Robust Regression
von: Yildirim, Yigit E., et al.
Veröffentlicht: (2025)
von: Yildirim, Yigit E., et al.
Veröffentlicht: (2025)
Learnability and Competition in High-Dimensional Multi-Component ICA
von: Genc, Eser Ilke, et al.
Veröffentlicht: (2026)
von: Genc, Eser Ilke, et al.
Veröffentlicht: (2026)
Exploring the Precise Dynamics of Single-Layer GAN Models: Leveraging Multi-Feature Discriminators for High-Dimensional Subspace Learning
von: Bond, Andrew, et al.
Veröffentlicht: (2024)
von: Bond, Andrew, et al.
Veröffentlicht: (2024)
Learning Optimal Classification Trees Robust to Distribution Shifts
von: Justin, Nathan, et al.
Veröffentlicht: (2023)
von: Justin, Nathan, et al.
Veröffentlicht: (2023)
Optimal Classification under Performative Distribution Shift
von: Cyffers, Edwige, et al.
Veröffentlicht: (2024)
von: Cyffers, Edwige, et al.
Veröffentlicht: (2024)
Distributionally Robust Coreset Selection under Covariate Shift
von: Tanaka, Tomonari, et al.
Veröffentlicht: (2025)
von: Tanaka, Tomonari, et al.
Veröffentlicht: (2025)
Certifiably Robust Model Evaluation in Federated Learning under Meta-Distributional Shifts
von: Najafi, Amir, et al.
Veröffentlicht: (2024)
von: Najafi, Amir, et al.
Veröffentlicht: (2024)
Distributionally Robust Safe Sample Elimination under Covariate Shift
von: Hanada, Hiroyuki, et al.
Veröffentlicht: (2024)
von: Hanada, Hiroyuki, et al.
Veröffentlicht: (2024)
Safe Distributionally Robust Feature Selection under Covariate Shift
von: Hanada, Hiroyuki, et al.
Veröffentlicht: (2026)
von: Hanada, Hiroyuki, et al.
Veröffentlicht: (2026)
Optimal Empirical Risk Minimization under Temporal Distribution Shifts
von: Jeong, Yujin, et al.
Veröffentlicht: (2025)
von: Jeong, Yujin, et al.
Veröffentlicht: (2025)
Trustworthy Machine Learning under Distribution Shifts
von: Huang, Zhuo
Veröffentlicht: (2025)
von: Huang, Zhuo
Veröffentlicht: (2025)
Transformers Learn Robust In-Context Regression under Distributional Uncertainty
von: Cao, Hoang T. H., et al.
Veröffentlicht: (2026)
von: Cao, Hoang T. H., et al.
Veröffentlicht: (2026)
Robust Machine Learning for Regulatory Sequence Modeling under Biological and Technical Distribution Shifts
von: Yang, Yiyao
Veröffentlicht: (2026)
von: Yang, Yiyao
Veröffentlicht: (2026)
Wasserstein Distributionally Robust Estimation in High Dimensions: Performance Analysis and Optimal Hyperparameter Tuning
von: Aolaritei, Liviu, et al.
Veröffentlicht: (2022)
von: Aolaritei, Liviu, et al.
Veröffentlicht: (2022)
Robust Calibration For Improved Weather Prediction Under Distributional Shift
von: Gilda, Sankalp, et al.
Veröffentlicht: (2024)
von: Gilda, Sankalp, et al.
Veröffentlicht: (2024)
Pruning is Optimal for Learning Sparse Features in High-Dimensions
von: Vural, Nuri Mert, et al.
Veröffentlicht: (2024)
von: Vural, Nuri Mert, et al.
Veröffentlicht: (2024)
Adaptive Estimation and Learning under Temporal Distribution Shift
von: Baby, Dheeraj, et al.
Veröffentlicht: (2025)
von: Baby, Dheeraj, et al.
Veröffentlicht: (2025)
Incorporating Interventional Independence Improves Robustness against Interventional Distribution Shift
von: Sreekumar, Gautam, et al.
Veröffentlicht: (2025)
von: Sreekumar, Gautam, et al.
Veröffentlicht: (2025)
Evaluating and Improving the Robustness of Speech Command Recognition Models to Noise and Distribution Shifts
von: Baranger, Anaïs, et al.
Veröffentlicht: (2025)
von: Baranger, Anaïs, et al.
Veröffentlicht: (2025)
CPR: Causal Physiological Representation Learning for Robust ECG Analysis under Distribution Shifts
von: Jia, Shunbo, et al.
Veröffentlicht: (2025)
von: Jia, Shunbo, et al.
Veröffentlicht: (2025)
Distributionally Robust Policy Evaluation under General Covariate Shift in Contextual Bandits
von: Guo, Yihong, et al.
Veröffentlicht: (2024)
von: Guo, Yihong, et al.
Veröffentlicht: (2024)
Latent Chain-of-Thought Improves Structured-Data Transformers
von: Dudley, Carson, et al.
Veröffentlicht: (2026)
von: Dudley, Carson, et al.
Veröffentlicht: (2026)
In-Context Learning Under Regime Change
von: Dudley, Carson, et al.
Veröffentlicht: (2026)
von: Dudley, Carson, et al.
Veröffentlicht: (2026)
Class Distribution Shifts in Zero-Shot Learning: Learning Robust Representations
von: Slavutsky, Yuli, et al.
Veröffentlicht: (2023)
von: Slavutsky, Yuli, et al.
Veröffentlicht: (2023)
Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift
von: Xue, Yihao, et al.
Veröffentlicht: (2023)
von: Xue, Yihao, et al.
Veröffentlicht: (2023)
Selective Attention: Enhancing Transformer through Principled Context Control
von: Zhang, Xuechen, et al.
Veröffentlicht: (2024)
von: Zhang, Xuechen, et al.
Veröffentlicht: (2024)
Federated Learning with Profile Mapping under Distribution Shifts and Drifts
von: Li, Mohan, et al.
Veröffentlicht: (2026)
von: Li, Mohan, et al.
Veröffentlicht: (2026)
Risk-Sensitive Soft Actor-Critic for Robust Deep Reinforcement Learning under Distribution Shifts
von: Enders, Tobias, et al.
Veröffentlicht: (2024)
von: Enders, Tobias, et al.
Veröffentlicht: (2024)
Assessing the Robustness of Climate Foundation Models under No-Analog Distribution Shifts
von: Navarro, Maria Conchita Agana, et al.
Veröffentlicht: (2026)
von: Navarro, Maria Conchita Agana, et al.
Veröffentlicht: (2026)
Robust Uncertainty Estimation under Distribution Shift via Difference Reconstruction
von: Xu, Xinran, et al.
Veröffentlicht: (2026)
von: Xu, Xinran, et al.
Veröffentlicht: (2026)
Optimal Policy Adaptation under Covariate Shift
von: Liu, Xueqing, et al.
Veröffentlicht: (2025)
von: Liu, Xueqing, et al.
Veröffentlicht: (2025)
ShiftKD: Benchmarking Knowledge Distillation under Distribution Shift
von: Zhang, Songming, et al.
Veröffentlicht: (2023)
von: Zhang, Songming, et al.
Veröffentlicht: (2023)
MIRRAMS: Learning Robust Tabular Models under Unseen Missingness Shifts
von: Lee, Jihye, et al.
Veröffentlicht: (2025)
von: Lee, Jihye, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
von: Demir, Samet, et al.
Veröffentlicht: (2025) -
Input-Label Correlation Governs a Linear-to-Nonlinear Transition in Random Features under Spiked Covariance
von: Demir, Samet, et al.
Veröffentlicht: (2024) -
Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure
von: Demir, Samet, et al.
Veröffentlicht: (2025) -
Implicitly Normalized Online PCA: A Regularized Algorithm with Exact High-Dimensional Dynamics
von: Demir, Samet, et al.
Veröffentlicht: (2025) -
Asymptotic Study of In-context Learning with Random Transformers through Equivalent Models
von: Demir, Samet, et al.
Veröffentlicht: (2025)