Gespeichert in:
| Hauptverfasser: | Kinoshita, Yuri, Nishikawa, Naoki, Toyoizumi, Taro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2603.14830 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A provable control of sensitivity of neural networks through a direct parameterization of the overall bi-Lipschitzness
von: Kinoshita, Yuri, et al.
Veröffentlicht: (2024)
von: Kinoshita, Yuri, et al.
Veröffentlicht: (2024)
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2025)
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2025)
Cortex and subcortex play distinct roles over learning when cortical memory is limited
von: Farrell, Matthew, et al.
Veröffentlicht: (2026)
von: Farrell, Matthew, et al.
Veröffentlicht: (2026)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2024)
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2024)
Gradient-Based Non-Linear Inverse Learning
von: Abhishake, et al.
Veröffentlicht: (2024)
von: Abhishake, et al.
Veröffentlicht: (2024)
Distilling Linearized Behavior into Non-Linear Fine-Tuning for Effective Task Arithmetic
von: Sommariva, Thomas, et al.
Veröffentlicht: (2026)
von: Sommariva, Thomas, et al.
Veröffentlicht: (2026)
Sample-Efficient Linear Representation Learning from Non-IID Non-Isotropic Data
von: Zhang, Thomas T. C. K., et al.
Veröffentlicht: (2023)
von: Zhang, Thomas T. C. K., et al.
Veröffentlicht: (2023)
Learning Task-Agnostic Representations through Multi-Teacher Distillation
von: Formont, Philippe, et al.
Veröffentlicht: (2025)
von: Formont, Philippe, et al.
Veröffentlicht: (2025)
DataDAM: Efficient Dataset Distillation with Attention Matching
von: Sajedi, Ahmad, et al.
Veröffentlicht: (2023)
von: Sajedi, Ahmad, et al.
Veröffentlicht: (2023)
From Low Intrinsic Dimensionality to Non-Vacuous Generalization Bounds in Deep Multi-Task Learning
von: Zakerinia, Hossein, et al.
Veröffentlicht: (2025)
von: Zakerinia, Hossein, et al.
Veröffentlicht: (2025)
Optimal Task Order for Continual Learning of Multiple Tasks
von: Li, Ziyan, et al.
Veröffentlicht: (2025)
von: Li, Ziyan, et al.
Veröffentlicht: (2025)
Learning Shared Representations for Multi-Task Linear Bandits
von: Lin, Jiabin, et al.
Veröffentlicht: (2026)
von: Lin, Jiabin, et al.
Veröffentlicht: (2026)
Multi-Task Representation Learning for Conservative Linear Bandits
von: Lin, Jiabin, et al.
Veröffentlicht: (2026)
von: Lin, Jiabin, et al.
Veröffentlicht: (2026)
Causality-Induced Positional Encoding for Transformer-Based Representation Learning of Non-Sequential Features
von: Xu, Kaichen, et al.
Veröffentlicht: (2025)
von: Xu, Kaichen, et al.
Veröffentlicht: (2025)
Disentangling and Mitigating the Impact of Task Similarity for Continual Learning
von: Hiratani, Naoki
Veröffentlicht: (2024)
von: Hiratani, Naoki
Veröffentlicht: (2024)
Reshaping Neural Representation via Associative, Presynaptic Short-Term Plasticity
von: Shimizu, Genki, et al.
Veröffentlicht: (2026)
von: Shimizu, Genki, et al.
Veröffentlicht: (2026)
Learning Dynamical Systems Encoding Non-Linearity within Space Curvature
von: Fichera, Bernardo, et al.
Veröffentlicht: (2024)
von: Fichera, Bernardo, et al.
Veröffentlicht: (2024)
Near-optimal and Efficient First-Order Algorithm for Multi-Task Learning with Shared Linear Representation
von: Ding, Shihong, et al.
Veröffentlicht: (2026)
von: Ding, Shihong, et al.
Veröffentlicht: (2026)
Blurred Encoding for Trajectory Representation Learning
von: Zhou, Silin, et al.
Veröffentlicht: (2025)
von: Zhou, Silin, et al.
Veröffentlicht: (2025)
Comparison of Autoencoder Encodings for ECG Representation in Downstream Prediction Tasks
von: Harvey, Christopher J., et al.
Veröffentlicht: (2024)
von: Harvey, Christopher J., et al.
Veröffentlicht: (2024)
What is Dataset Distillation Learning?
von: Yang, William, et al.
Veröffentlicht: (2024)
von: Yang, William, et al.
Veröffentlicht: (2024)
Random Gradient-Free Optimization in Infinite Dimensional Spaces
von: Peixoto, Caio Lins, et al.
Veröffentlicht: (2025)
von: Peixoto, Caio Lins, et al.
Veröffentlicht: (2025)
Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context
von: Cheng, Xiang, et al.
Veröffentlicht: (2023)
von: Cheng, Xiang, et al.
Veröffentlicht: (2023)
Dataset Distillation-based Hybrid Federated Learning on Non-IID Data
von: Shi, Xiufang, et al.
Veröffentlicht: (2024)
von: Shi, Xiufang, et al.
Veröffentlicht: (2024)
Spectral Gradient Surgery for Domain-Generalizable Dataset Distillation
von: Oh, Minyoung, et al.
Veröffentlicht: (2026)
von: Oh, Minyoung, et al.
Veröffentlicht: (2026)
Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
von: Saglam, Baturay, et al.
Veröffentlicht: (2025)
von: Saglam, Baturay, et al.
Veröffentlicht: (2025)
Mask-Encoded Sparsification: Mitigating Biased Gradients in Communication-Efficient Split Learning
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
On the Diversity and Realism of Distilled Dataset: An Efficient Dataset Distillation Paradigm
von: Sun, Peng, et al.
Veröffentlicht: (2023)
von: Sun, Peng, et al.
Veröffentlicht: (2023)
Learning Linear Regression with Low-Rank Tasks in-Context
von: Takanami, Kaito, et al.
Veröffentlicht: (2025)
von: Takanami, Kaito, et al.
Veröffentlicht: (2025)
LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
von: Robert, Thomas, et al.
Veröffentlicht: (2024)
von: Robert, Thomas, et al.
Veröffentlicht: (2024)
Learning from Linear Algebra: A Graph Neural Network Approach to Preconditioner Design for Conjugate Gradient Solvers
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2024)
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2024)
Data-to-Model Distillation: Data-Efficient Learning Framework
von: Sajedi, Ahmad, et al.
Veröffentlicht: (2024)
von: Sajedi, Ahmad, et al.
Veröffentlicht: (2024)
Dynamics and Representation Structure of Local Approximations to Gradient-Based Learning in Linear Recurrent Neural Networks
von: Williams, Ezekiel, et al.
Veröffentlicht: (2026)
von: Williams, Ezekiel, et al.
Veröffentlicht: (2026)
Chebyshev Policies and the Mountain Car Problem: Reinforcement Learning for Low-Dimensional Control Tasks
von: Huber, Stefan, et al.
Veröffentlicht: (2026)
von: Huber, Stefan, et al.
Veröffentlicht: (2026)
Nonparametric Instrumental Variable Regression through Stochastic Approximate Gradients
von: Fonseca, Yuri, et al.
Veröffentlicht: (2024)
von: Fonseca, Yuri, et al.
Veröffentlicht: (2024)
On Learning Representations for Tabular Data Distillation
von: Kang, Inwon, et al.
Veröffentlicht: (2025)
von: Kang, Inwon, et al.
Veröffentlicht: (2025)
High-Dimensional Search, Low-Dimensional Solution: Decoupling Optimization from Representation
von: Kalyoncuoglu, Yusuf, et al.
Veröffentlicht: (2025)
von: Kalyoncuoglu, Yusuf, et al.
Veröffentlicht: (2025)
Learning to Flow from Generative Pretext Tasks for Neural Architecture Encoding
von: Kim, Sunwoo, et al.
Veröffentlicht: (2025)
von: Kim, Sunwoo, et al.
Veröffentlicht: (2025)
Exploring the Potential of QEEGNet for Cross-Task and Cross-Dataset Electroencephalography Encoding with Quantum Machine Learning
von: Chen, Chi-Sheng, et al.
Veröffentlicht: (2025)
von: Chen, Chi-Sheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A provable control of sensitivity of neural networks through a direct parameterization of the overall bi-Lipschitzness
von: Kinoshita, Yuri, et al.
Veröffentlicht: (2024) -
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025) -
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2025) -
Cortex and subcortex play distinct roles over learning when cortical memory is limited
von: Farrell, Matthew, et al.
Veröffentlicht: (2026) -
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2024)