Small transformer architectures for task switching
Fuente:
arXiv
Saved in:
| Main Author: | Gros, Claudius |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reorganizing attention-space geometry with expressive attention
by: Gros, Claudius
Published: (2024)
by: Gros, Claudius
Published: (2024)
From generative AI to the brain: five takeaways
by: Gros, Claudius
Published: (2025)
by: Gros, Claudius
Published: (2025)
Financial time series augmentation using transformer based GAN architecture
by: Podobiński, Andrzej, et al.
Published: (2026)
by: Podobiński, Andrzej, et al.
Published: (2026)
AlphaZero Neural Scaling and Zipf's Law: a Tale of Board Games and Power Laws
by: Neumann, Oren, et al.
Published: (2024)
by: Neumann, Oren, et al.
Published: (2024)
Physical models realizing the transformer architecture of large language models
by: Chen, Zeqian
Published: (2025)
by: Chen, Zeqian
Published: (2025)
All AI Models are Wrong, but Some are Optimal
by: Anand, Akhil S, et al.
Published: (2025)
by: Anand, Akhil S, et al.
Published: (2025)
Surrogate modeling of Cellular-Potts Agent-Based Models as a segmentation task using the U-Net neural network architecture
by: Comlekoglu, Tien, et al.
Published: (2025)
by: Comlekoglu, Tien, et al.
Published: (2025)
Growth strategies for arbitrary DAG neural architectures
by: Douka, Stella, et al.
Published: (2025)
by: Douka, Stella, et al.
Published: (2025)
Closing the Sim2Real Performance Gap in RL
by: Anand, Akhil S, et al.
Published: (2025)
by: Anand, Akhil S, et al.
Published: (2025)
An efficient probabilistic hardware architecture for diffusion-like models
by: Jelinčič, Andraž, et al.
Published: (2025)
by: Jelinčič, Andraž, et al.
Published: (2025)
A layered architecture for log analysis in complex IT systems
by: Wittkopp, Thorsten
Published: (2025)
by: Wittkopp, Thorsten
Published: (2025)
Auxiliary task discovery through generate-and-test
by: Rafiee, Banafsheh, et al.
Published: (2022)
by: Rafiee, Banafsheh, et al.
Published: (2022)
Provable unlearning in topic modeling and downstream tasks
by: Wei, Stanley, et al.
Published: (2024)
by: Wei, Stanley, et al.
Published: (2024)
Topological derivative approach for deep neural network architecture adaptation
by: Krishnanunni, C G, et al.
Published: (2025)
by: Krishnanunni, C G, et al.
Published: (2025)
A framework for measuring the training efficiency of a neural architecture
by: Cueto-Mendoza, Eduardo, et al.
Published: (2024)
by: Cueto-Mendoza, Eduardo, et al.
Published: (2024)
Curriculum reinforcement learning with measurable task representation learning
by: Wen, Yongyan, et al.
Published: (2026)
by: Wen, Yongyan, et al.
Published: (2026)
Offline Multi-task Transfer RL with Representational Penalization
by: Bose, Avinandan, et al.
Published: (2024)
by: Bose, Avinandan, et al.
Published: (2024)
Why pre-training is beneficial for downstream classification tasks?
by: Jiang, Xin, et al.
Published: (2024)
by: Jiang, Xin, et al.
Published: (2024)
Conflict-Averse Gradient Descent for Multi-task Learning
by: Liu, Bo, et al.
Published: (2021)
by: Liu, Bo, et al.
Published: (2021)
Per-Domain Generalizing Policies: On Learning Efficient and Robust Q-Value Functions (Extended Version with Technical Appendix)
by: Müller, Nicola J., et al.
Published: (2026)
by: Müller, Nicola J., et al.
Published: (2026)
Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention
by: Súkeník, Peter, et al.
Published: (2026)
by: Súkeník, Peter, et al.
Published: (2026)
Separable neural architectures as a primitive for unified predictive and generative intelligence
by: Batley, Reza T., et al.
Published: (2026)
by: Batley, Reza T., et al.
Published: (2026)
Neuroplasticity-inspired dynamic ANNs for multi-task demand forecasting
by: Żarski, Mateusz, et al.
Published: (2025)
by: Żarski, Mateusz, et al.
Published: (2025)
Multi-task GINN-LP for Multi-target Symbolic Regression
by: Rajabu, Hussein, et al.
Published: (2025)
by: Rajabu, Hussein, et al.
Published: (2025)
An Electrocardiogram Multi-task Benchmark with Comprehensive Evaluations and Insightful Findings
by: Xu, Yuhao, et al.
Published: (2025)
by: Xu, Yuhao, et al.
Published: (2025)
Nonlocal operator learning for fMRI encoding and decoding tasks
by: Kramer, Andreas, et al.
Published: (2026)
by: Kramer, Andreas, et al.
Published: (2026)
Multi-task Heterogeneous Graph Learning on Electronic Health Records
by: Chan, Tsai Hor, et al.
Published: (2024)
by: Chan, Tsai Hor, et al.
Published: (2024)
Multi-task Neural Diffusion Processes
by: Rawson, Joseph, et al.
Published: (2025)
by: Rawson, Joseph, et al.
Published: (2025)
Out-of-distribution generalisation is hard: evidence from ARC-like tasks
by: Dimitriadis, George, et al.
Published: (2025)
by: Dimitriadis, George, et al.
Published: (2025)
An empirical study of task and feature correlations in the reuse of pre-trained models
by: Mohamud, Jama Hussein, et al.
Published: (2025)
by: Mohamud, Jama Hussein, et al.
Published: (2025)
Multi-task Domain Adaptation for Computation Offloading in Edge-intelligence Networks
by: Han, Runxin, et al.
Published: (2025)
by: Han, Runxin, et al.
Published: (2025)
Transformers as Multi-task Learners: Decoupling Features in Hidden Markov Models
by: Hao, Yifan, et al.
Published: (2025)
by: Hao, Yifan, et al.
Published: (2025)
Towards Unified Multi-task EEG Analysis with Low-Rank Adaptation
by: Dai, Sicheng, et al.
Published: (2026)
by: Dai, Sicheng, et al.
Published: (2026)
A dynamical clipping approach with task feedback for Proximal Policy Optimization
by: Zhang, Ziqi, et al.
Published: (2023)
by: Zhang, Ziqi, et al.
Published: (2023)
Mixture of Experts based Multi-task Supervise Learning from Crowds
by: Han, Tao, et al.
Published: (2024)
by: Han, Tao, et al.
Published: (2024)
Per-Domain Generalizing Policies: On Validation Instances and Scaling Behavior
by: Gros, Timo P., et al.
Published: (2025)
by: Gros, Timo P., et al.
Published: (2025)
Carrying over algorithm in transformers
by: Kruthoff, Jorrit
Published: (2024)
by: Kruthoff, Jorrit
Published: (2024)
Graph Neural Alchemist: An innovative fully modular architecture for time series-to-graph classification
by: Coelho, Paulo, et al.
Published: (2024)
by: Coelho, Paulo, et al.
Published: (2024)
Online Decentralized Federated Multi-task Learning With Trustworthiness in Cyber-Physical Systems
by: Odeyomi, Olusola, et al.
Published: (2025)
by: Odeyomi, Olusola, et al.
Published: (2025)
EnECG: Efficient Ensemble Learning for Electrocardiogram Multi-task Foundation Model
by: Xu, Yuhao, et al.
Published: (2025)
by: Xu, Yuhao, et al.
Published: (2025)
Similar Items
-
Reorganizing attention-space geometry with expressive attention
by: Gros, Claudius
Published: (2024) -
From generative AI to the brain: five takeaways
by: Gros, Claudius
Published: (2025) -
Financial time series augmentation using transformer based GAN architecture
by: Podobiński, Andrzej, et al.
Published: (2026) -
AlphaZero Neural Scaling and Zipf's Law: a Tale of Board Games and Power Laws
by: Neumann, Oren, et al.
Published: (2024) -
Physical models realizing the transformer architecture of large language models
by: Chen, Zeqian
Published: (2025)