In-Context Deep Learning via Transformer Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Weimin, Su, Maojiang, Hu, Jerry Yao-Chieh, Song, Zhao, Liu, Han |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
A Theoretical Analysis of Discrete Flow Matching Generative Models
by: Su, Maojiang, et al.
Published: (2025)
by: Su, Maojiang, et al.
Published: (2025)
On Flow Matching KL Divergence
by: Su, Maojiang, et al.
Published: (2025)
by: Su, Maojiang, et al.
Published: (2025)
Discrete Flow Matching Policy Optimization
by: Su, Maojiang, et al.
Published: (2026)
by: Su, Maojiang, et al.
Published: (2026)
Fast and Low-Cost Genomic Foundation Models via Outlier Removal
by: Luo, Haozheng, et al.
Published: (2025)
by: Luo, Haozheng, et al.
Published: (2025)
In-Context Algorithm Emulation in Fixed-Weight Transformers
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Provably Optimal Memory Capacity for Modern Hopfield Models: Transformer-Compatible Dense Associative Memories as Spherical Codes
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
On Computational Limits of Modern Hopfield Models: A Fine-Grained Complexity Analysis
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Attention Mechanism, Max-Affine Partition, and Universal Approximation
by: Liu, Hude, et al.
Published: (2025)
by: Liu, Hude, et al.
Published: (2025)
Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation Models
by: Wu, Weimin, et al.
Published: (2025)
by: Wu, Weimin, et al.
Published: (2025)
On Statistical Rates of Conditional Diffusion Transformers: Approximation, Estimation and Minimax Optimality
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Transformer Approximations from ReLUs
by: Hu, Jerry Yao-Chieh, et al.
Published: (2026)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2026)
On Structured State-Space Duality
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Uniform Memory Retrieval with Larger Capacity for Modern Hopfield Models
by: Wu, Dennis, et al.
Published: (2024)
by: Wu, Dennis, et al.
Published: (2024)
Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and Efficiency
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
POLO: Preference-Guided Multi-Turn Reinforcement Learning for Lead Optimization
by: Wang, Ziqing, et al.
Published: (2025)
by: Wang, Ziqing, et al.
Published: (2025)
Differentially Private Kernel Density Estimation
by: Liu, Erzhi, et al.
Published: (2024)
by: Liu, Erzhi, et al.
Published: (2024)
On Differentially Private String Distances
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Nonparametric Modern Hopfield Models
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Outlier-Efficient Hopfield Layers for Large Transformer-Based Models
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Pareto-Optimal Energy Alignment for Designing Nature-Like Antibodies
by: Wen, Yibo, et al.
Published: (2024)
by: Wen, Yibo, et al.
Published: (2024)
Are Hallucinations Bad Estimations?
by: Liu, Hude, et al.
Published: (2025)
by: Liu, Hude, et al.
Published: (2025)
BiSHop: Bi-Directional Cellular Learning for Tabular Data with Generalized Sparse Modern Hopfield Model
by: Xu, Chenwei, et al.
Published: (2024)
by: Xu, Chenwei, et al.
Published: (2024)
STanHop: Sparse Tandem Hopfield Model for Memory-Enhanced Time Series Prediction
by: Wu, Dennis, et al.
Published: (2023)
by: Wu, Dennis, et al.
Published: (2023)
Latent Variable Estimation in Bayesian Black-Litterman Models
by: Lin, Thomas Y. L., et al.
Published: (2025)
by: Lin, Thomas Y. L., et al.
Published: (2025)
Transformers Learn Latent Mixture Models In-Context via Mirror Descent
by: D'Angelo, Francesco, et al.
Published: (2026)
by: D'Angelo, Francesco, et al.
Published: (2026)
Exact Conversion of In-Context Learning to Model Weights in Linearized-Attention Transformers
by: Chen, Brian K, et al.
Published: (2024)
by: Chen, Brian K, et al.
Published: (2024)
Deep Transfer Learning: Model Framework and Error Analysis
by: Jiao, Yuling, et al.
Published: (2024)
by: Jiao, Yuling, et al.
Published: (2024)
Estimating Time Series Foundation Model Transferability via In-Context Learning
by: Yao, Qingren, et al.
Published: (2025)
by: Yao, Qingren, et al.
Published: (2025)
In-Context Compositional Learning via Sparse Coding Transformer
by: Chen, Wei, et al.
Published: (2025)
by: Chen, Wei, et al.
Published: (2025)
In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-Thought
by: Huang, Sili, et al.
Published: (2024)
by: Huang, Sili, et al.
Published: (2024)
Out-of-Context Misinformation Detection via Variational Domain-Invariant Learning with Test-Time Training
by: Yang, Xi, et al.
Published: (2025)
by: Yang, Xi, et al.
Published: (2025)
Deep conditional distribution learning via conditional Föllmer flow
by: Chang, Jinyuan, et al.
Published: (2024)
by: Chang, Jinyuan, et al.
Published: (2024)
CoMeT: Collaborative Memory Transformer for Efficient Long Context Modeling
by: Zhao, Runsong, et al.
Published: (2026)
by: Zhao, Runsong, et al.
Published: (2026)
In-Context In-Context Learning with Transformer Neural Processes
by: Ashman, Matthew, et al.
Published: (2024)
by: Ashman, Matthew, et al.
Published: (2024)
DeepJ: Graph Convolutional Transformers with Differentiable Pooling for Patient Trajectory Modeling
by: Li, Deyi, et al.
Published: (2025)
by: Li, Deyi, et al.
Published: (2025)
RNA Secondary Structure Prediction Using Transformer-Based Deep Learning Models
by: Zhou, Yanlin, et al.
Published: (2024)
by: Zhou, Yanlin, et al.
Published: (2024)
Similar Items
-
Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024) -
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025) -
On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024) -
A Theoretical Analysis of Discrete Flow Matching Generative Models
by: Su, Maojiang, et al.
Published: (2025) -
On Flow Matching KL Divergence
by: Su, Maojiang, et al.
Published: (2025)