Less is More: Efficient Weight Farcasting with 1-Layer Neural Network
Fuente:
arXiv
Salvato in:
| Autori principali: | Shou, Xiao, Bhattacharjya, Debarun, Ding, Yanna, Zhao, Chen, Li, Rui, Gao, Jianxi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Gradient Flow Matching for Learning Update Dynamics in Neural Network Training
di: Shou, Xiao, et al.
Pubblicazione: (2025)
di: Shou, Xiao, et al.
Pubblicazione: (2025)
Self-Supervised Contrastive Pre-Training for Multivariate Point Processes
di: Shou, Xiao, et al.
Pubblicazione: (2024)
di: Shou, Xiao, et al.
Pubblicazione: (2024)
Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential Equation
di: Ding, Yanna, et al.
Pubblicazione: (2024)
di: Ding, Yanna, et al.
Pubblicazione: (2024)
Predicting Time Series of Networked Dynamical Systems without Knowing Topology
di: Ding, Yanna, et al.
Pubblicazione: (2024)
di: Ding, Yanna, et al.
Pubblicazione: (2024)
Optimality and NP-Hardness of Transformers in Learning Markovian Dynamical Functions
di: Ding, Yanna, et al.
Pubblicazione: (2025)
di: Ding, Yanna, et al.
Pubblicazione: (2025)
Q-function Decomposition with Intervention Semantics with Factored Action Spaces
di: Lee, Junkyu, et al.
Pubblicazione: (2025)
di: Lee, Junkyu, et al.
Pubblicazione: (2025)
Less is More: Efficient Model Merging with Binary Task Switch
di: Qi, Biqing, et al.
Pubblicazione: (2024)
di: Qi, Biqing, et al.
Pubblicazione: (2024)
Layer-wise Weight Selection for Power-Efficient Neural Network Acceleration
di: Fang, Jiaxun, et al.
Pubblicazione: (2025)
di: Fang, Jiaxun, et al.
Pubblicazione: (2025)
Less Is More -- On the Importance of Sparsification for Transformers and Graph Neural Networks for TSP
di: Lischka, Attila, et al.
Pubblicazione: (2024)
di: Lischka, Attila, et al.
Pubblicazione: (2024)
Inferring from Logits: Exploring Best Practices for Decoding-Free Generative Candidate Selection
di: Ma, Mingyu Derek, et al.
Pubblicazione: (2025)
di: Ma, Mingyu Derek, et al.
Pubblicazione: (2025)
InfoSFT: Learn More and Forget Less with Information-Aware Token Weighting
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2026)
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2026)
ODE-free Neural Flow Matching for One-Step Generative Modeling
di: Shou, Xiao
Pubblicazione: (2026)
di: Shou, Xiao
Pubblicazione: (2026)
Less is More: Recursive Reasoning with Tiny Networks
di: Jolicoeur-Martineau, Alexia
Pubblicazione: (2025)
di: Jolicoeur-Martineau, Alexia
Pubblicazione: (2025)
More with Less: An Empirical Study of Turn-Control Strategies for Efficient Coding Agents
di: Gao, Pengfei, et al.
Pubblicazione: (2025)
di: Gao, Pengfei, et al.
Pubblicazione: (2025)
Knowledge Base Construction for Knowledge-Augmented Text-to-SQL
di: Baek, Jinheon, et al.
Pubblicazione: (2025)
di: Baek, Jinheon, et al.
Pubblicazione: (2025)
Less is More: on the Over-Globalizing Problem in Graph Transformers
di: Xing, Yujie, et al.
Pubblicazione: (2024)
di: Xing, Yujie, et al.
Pubblicazione: (2024)
The Consistency Hypothesis in Uncertainty Quantification for Large Language Models
di: Xiao, Quan, et al.
Pubblicazione: (2025)
di: Xiao, Quan, et al.
Pubblicazione: (2025)
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
di: Wu, Shaojin, et al.
Pubblicazione: (2025)
di: Wu, Shaojin, et al.
Pubblicazione: (2025)
Less is More: Towards Simple Graph Contrastive Learning
di: Zhao, Yanan, et al.
Pubblicazione: (2025)
di: Zhao, Yanan, et al.
Pubblicazione: (2025)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
di: Das, Rudrajit, et al.
Pubblicazione: (2026)
di: Das, Rudrajit, et al.
Pubblicazione: (2026)
Transformer Multivariate Forecasting: Less is More?
di: Xu, Jingjing, et al.
Pubblicazione: (2023)
di: Xu, Jingjing, et al.
Pubblicazione: (2023)
Less is More: Non-uniform Road Segments are Efficient for Bus Arrival Prediction
di: Huang, Zhen, et al.
Pubblicazione: (2025)
di: Huang, Zhen, et al.
Pubblicazione: (2025)
Less or More From Teacher: Exploiting Trilateral Geometry For Knowledge Distillation
di: Hu, Chengming, et al.
Pubblicazione: (2023)
di: Hu, Chengming, et al.
Pubblicazione: (2023)
Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection
di: Chen, Ruoyu, et al.
Pubblicazione: (2025)
di: Chen, Ruoyu, et al.
Pubblicazione: (2025)
Quantize What Counts: More for Keys, Less for Values
di: Hariri, Mohsen, et al.
Pubblicazione: (2025)
di: Hariri, Mohsen, et al.
Pubblicazione: (2025)
Learning More with Less: A Dynamic Dual-Level Down-Sampling Framework for Efficient Policy Optimization
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer Learning
di: Schulte, David, et al.
Pubblicazione: (2024)
di: Schulte, David, et al.
Pubblicazione: (2024)
LIMR: Less is More for RL Scaling
di: Li, Xuefeng, et al.
Pubblicazione: (2025)
di: Li, Xuefeng, et al.
Pubblicazione: (2025)
From Deep Additive Kernel Learning to Last-Layer Bayesian Neural Networks via Induced Prior Approximation
di: Zhao, Wenyuan, et al.
Pubblicazione: (2025)
di: Zhao, Wenyuan, et al.
Pubblicazione: (2025)
When Less is More: The LLM Scaling Paradox in Context Compression
di: Guo, Ruishan, et al.
Pubblicazione: (2026)
di: Guo, Ruishan, et al.
Pubblicazione: (2026)
HyperSAT: Unsupervised Hypergraph Neural Networks for Weighted MaxSAT Problems
di: Chen, Qiyue, et al.
Pubblicazione: (2025)
di: Chen, Qiyue, et al.
Pubblicazione: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
di: Lou, Chao, et al.
Pubblicazione: (2024)
di: Lou, Chao, et al.
Pubblicazione: (2024)
Text2Weight: Bridging Natural Language and Neural Network Weight Spaces
di: Tian, Bowen, et al.
Pubblicazione: (2025)
di: Tian, Bowen, et al.
Pubblicazione: (2025)
Less is More: Extreme Gradient Boost Rank-1 Adaption for Efficient Finetuning of LLMs
di: Zhang, Yifei, et al.
Pubblicazione: (2024)
di: Zhang, Yifei, et al.
Pubblicazione: (2024)
Less is More: Adaptive Coverage for Synthetic Training Data
di: Tavakkol, Sasan, et al.
Pubblicazione: (2025)
di: Tavakkol, Sasan, et al.
Pubblicazione: (2025)
Self-Ablating Transformers: More Interpretability, Less Sparsity
di: Ferrao, Jeremias, et al.
Pubblicazione: (2025)
di: Ferrao, Jeremias, et al.
Pubblicazione: (2025)
Less is More: Rethinking Few-Shot Learning and Recurrent Neural Nets
di: Pereg, Deborah, et al.
Pubblicazione: (2022)
di: Pereg, Deborah, et al.
Pubblicazione: (2022)
Coreset-Induced Conditional Velocity Flow Matching
di: Wang, Xiao, et al.
Pubblicazione: (2026)
di: Wang, Xiao, et al.
Pubblicazione: (2026)
Artificial Geographically Weighted Neural Network: A Novel Framework for Spatial Analysis with Geographically Weighted Layers
di: Cao, Jianfei, et al.
Pubblicazione: (2025)
di: Cao, Jianfei, et al.
Pubblicazione: (2025)
Moirai 2.0: When Less Is More for Time Series Forecasting
di: Liu, Chenghao, et al.
Pubblicazione: (2025)
di: Liu, Chenghao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Gradient Flow Matching for Learning Update Dynamics in Neural Network Training
di: Shou, Xiao, et al.
Pubblicazione: (2025) -
Self-Supervised Contrastive Pre-Training for Multivariate Point Processes
di: Shou, Xiao, et al.
Pubblicazione: (2024) -
Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential Equation
di: Ding, Yanna, et al.
Pubblicazione: (2024) -
Predicting Time Series of Networked Dynamical Systems without Knowing Topology
di: Ding, Yanna, et al.
Pubblicazione: (2024) -
Optimality and NP-Hardness of Transformers in Learning Markovian Dynamical Functions
di: Ding, Yanna, et al.
Pubblicazione: (2025)