FAME: Adaptive Functional Attention with Expert Routing for Function-on-Function Regression
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Yifei, Chen, Yong, Zhang, Chen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Meta-Learning Adaptive Loss Functions
by: Raymond, Christian, et al.
Published: (2023)
by: Raymond, Christian, et al.
Published: (2023)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
by: Li, Cheng, et al.
Published: (2025)
by: Li, Cheng, et al.
Published: (2025)
Route Experts by Sequence, not by Token
by: Wen, Tiansheng, et al.
Published: (2025)
by: Wen, Tiansheng, et al.
Published: (2025)
ANAct: Adaptive Normalization for Activation Functions
by: Peiwen, Yuan, et al.
Published: (2022)
by: Peiwen, Yuan, et al.
Published: (2022)
Learning Adaptive Distribution Alignment with Neural Characteristic Function for Graph Domain Adaptation
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
Towards Generalization-Oriented Models for Vehicle Routing Problems with Mixture-of-Experts
by: Miao, Changhao, et al.
Published: (2026)
by: Miao, Changhao, et al.
Published: (2026)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
FAME: Formal Abstract Minimal Explanation for Neural Networks
by: Boumazouza, Ryma, et al.
Published: (2026)
by: Boumazouza, Ryma, et al.
Published: (2026)
Beyond the Mean: Distribution-Aware Loss Functions for Bimodal Regression
by: Mohammadi-Seif, Abolfazl, et al.
Published: (2026)
by: Mohammadi-Seif, Abolfazl, et al.
Published: (2026)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
by: He, Yifei, et al.
Published: (2025)
by: He, Yifei, et al.
Published: (2025)
Neural Dynamics-Informed Pre-trained Framework for Personalized Brain Functional Network Construction
by: Jiang, Hongjie, et al.
Published: (2026)
by: Jiang, Hongjie, et al.
Published: (2026)
Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
by: Chen, Justin Chih-Yao, et al.
Published: (2025)
by: Chen, Justin Chih-Yao, et al.
Published: (2025)
LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing
by: Li, Wenbing, et al.
Published: (2025)
by: Li, Wenbing, et al.
Published: (2025)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025)
by: Zuo, Yifei, et al.
Published: (2025)
TQL: Scaling Q-Functions with Transformers by Preventing Attention Collapse
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-Tuning
by: Wu, Hong-Wei, et al.
Published: (2024)
by: Wu, Hong-Wei, et al.
Published: (2024)
Adaptive Transformer Modelling of Density Function for Nonparametric Survival Analysis
by: Zhang, Xin, et al.
Published: (2024)
by: Zhang, Xin, et al.
Published: (2024)
Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
by: Farebrother, Jesse, et al.
Published: (2024)
by: Farebrother, Jesse, et al.
Published: (2024)
Geometric Mixture-of-Experts with Curvature-Guided Adaptive Routing for Graph Representation Learning
by: Cao, Haifang, et al.
Published: (2026)
by: Cao, Haifang, et al.
Published: (2026)
ARROW: An Adaptive Rollout and Routing Method for Global Weather Forecasting
by: Tian, Jindong, et al.
Published: (2025)
by: Tian, Jindong, et al.
Published: (2025)
A Note on Knowledge Distillation Loss Function for Object Classification
by: Chen, Defang
Published: (2021)
by: Chen, Defang
Published: (2021)
Global Convergence in Neural ODEs: Impact of Activation Functions
by: Gao, Tianxiang, et al.
Published: (2025)
by: Gao, Tianxiang, et al.
Published: (2025)
Adaptive Exploration for Data-Efficient General Value Function Evaluations
by: Jain, Arushi, et al.
Published: (2024)
by: Jain, Arushi, et al.
Published: (2024)
Understanding Fairness Surrogate Functions in Algorithmic Fairness
by: Yao, Wei, et al.
Published: (2023)
by: Yao, Wei, et al.
Published: (2023)
Adaptive Individual Uncertainty under Out-Of-Distribution Shift with Expert-Routed Conformal Prediction
by: Badkul, Amitesh, et al.
Published: (2025)
by: Badkul, Amitesh, et al.
Published: (2025)
Rethinking Time Encoding via Learnable Transformation Functions
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
Neural Inhibition Improves Dynamic Routing and Mixture of Experts
by: Zou, Will Y., et al.
Published: (2025)
by: Zou, Will Y., et al.
Published: (2025)
RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs
by: Xu, Zhiyuan, et al.
Published: (2026)
by: Xu, Zhiyuan, et al.
Published: (2026)
Adaptive Friction in Deep Learning: Enhancing Optimizers with Sigmoid and Tanh Function
by: Zheng, Hongye, et al.
Published: (2024)
by: Zheng, Hongye, et al.
Published: (2024)
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
by: Liu, Yilun, et al.
Published: (2024)
by: Liu, Yilun, et al.
Published: (2024)
Structural Compositional Function Networks: Interpretable Functional Compositions for Tabular Discovery
by: Li, Fang
Published: (2026)
by: Li, Fang
Published: (2026)
VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning
by: Chen, Xuyang, et al.
Published: (2025)
by: Chen, Xuyang, et al.
Published: (2025)
Unraveling Indirect In-Context Learning Using Influence Functions
by: Askari, Hadi, et al.
Published: (2025)
by: Askari, Hadi, et al.
Published: (2025)
Cross-Table Pretraining towards a Universal Function Space for Heterogeneous Tabular Data
by: Chen, Jintai, et al.
Published: (2024)
by: Chen, Jintai, et al.
Published: (2024)
Neural Green's Functions
by: Yoo, Seungwoo, et al.
Published: (2025)
by: Yoo, Seungwoo, et al.
Published: (2025)
Universal Neural Functionals
by: Zhou, Allan, et al.
Published: (2024)
by: Zhou, Allan, et al.
Published: (2024)
Stable Offline Value Function Learning with Bisimulation-based Representations
by: Pavse, Brahma S., et al.
Published: (2024)
by: Pavse, Brahma S., et al.
Published: (2024)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
by: Collins, Liam, et al.
Published: (2024)
by: Collins, Liam, et al.
Published: (2024)
The Routing and Filtering Structure of Attention
by: Jamil, Shafayeth, et al.
Published: (2026)
by: Jamil, Shafayeth, et al.
Published: (2026)
Visualizing Loss Functions as Topological Landscape Profiles
by: Geniesse, Caleb, et al.
Published: (2024)
by: Geniesse, Caleb, et al.
Published: (2024)
Similar Items
-
Meta-Learning Adaptive Loss Functions
by: Raymond, Christian, et al.
Published: (2023) -
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
by: Li, Cheng, et al.
Published: (2025) -
Route Experts by Sequence, not by Token
by: Wen, Tiansheng, et al.
Published: (2025) -
ANAct: Adaptive Normalization for Activation Functions
by: Peiwen, Yuan, et al.
Published: (2022) -
Learning Adaptive Distribution Alignment with Neural Characteristic Function for Graph Domain Adaptation
by: Chen, Wei, et al.
Published: (2026)