Less is More: Efficient Weight Farcasting with 1-Layer Neural Network
Fuente:
arXiv
Guardado en:
| Autores principales: | Shou, Xiao, Bhattacharjya, Debarun, Ding, Yanna, Zhao, Chen, Li, Rui, Gao, Jianxi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Gradient Flow Matching for Learning Update Dynamics in Neural Network Training
por: Shou, Xiao, et al.
Publicado: (2025)
por: Shou, Xiao, et al.
Publicado: (2025)
Self-Supervised Contrastive Pre-Training for Multivariate Point Processes
por: Shou, Xiao, et al.
Publicado: (2024)
por: Shou, Xiao, et al.
Publicado: (2024)
Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential Equation
por: Ding, Yanna, et al.
Publicado: (2024)
por: Ding, Yanna, et al.
Publicado: (2024)
Predicting Time Series of Networked Dynamical Systems without Knowing Topology
por: Ding, Yanna, et al.
Publicado: (2024)
por: Ding, Yanna, et al.
Publicado: (2024)
Optimality and NP-Hardness of Transformers in Learning Markovian Dynamical Functions
por: Ding, Yanna, et al.
Publicado: (2025)
por: Ding, Yanna, et al.
Publicado: (2025)
Q-function Decomposition with Intervention Semantics with Factored Action Spaces
por: Lee, Junkyu, et al.
Publicado: (2025)
por: Lee, Junkyu, et al.
Publicado: (2025)
Less is More: Efficient Model Merging with Binary Task Switch
por: Qi, Biqing, et al.
Publicado: (2024)
por: Qi, Biqing, et al.
Publicado: (2024)
Layer-wise Weight Selection for Power-Efficient Neural Network Acceleration
por: Fang, Jiaxun, et al.
Publicado: (2025)
por: Fang, Jiaxun, et al.
Publicado: (2025)
Less Is More -- On the Importance of Sparsification for Transformers and Graph Neural Networks for TSP
por: Lischka, Attila, et al.
Publicado: (2024)
por: Lischka, Attila, et al.
Publicado: (2024)
Inferring from Logits: Exploring Best Practices for Decoding-Free Generative Candidate Selection
por: Ma, Mingyu Derek, et al.
Publicado: (2025)
por: Ma, Mingyu Derek, et al.
Publicado: (2025)
InfoSFT: Learn More and Forget Less with Information-Aware Token Weighting
por: Sabbaghi, Mahdi, et al.
Publicado: (2026)
por: Sabbaghi, Mahdi, et al.
Publicado: (2026)
ODE-free Neural Flow Matching for One-Step Generative Modeling
por: Shou, Xiao
Publicado: (2026)
por: Shou, Xiao
Publicado: (2026)
Less is More: Recursive Reasoning with Tiny Networks
por: Jolicoeur-Martineau, Alexia
Publicado: (2025)
por: Jolicoeur-Martineau, Alexia
Publicado: (2025)
More with Less: An Empirical Study of Turn-Control Strategies for Efficient Coding Agents
por: Gao, Pengfei, et al.
Publicado: (2025)
por: Gao, Pengfei, et al.
Publicado: (2025)
Knowledge Base Construction for Knowledge-Augmented Text-to-SQL
por: Baek, Jinheon, et al.
Publicado: (2025)
por: Baek, Jinheon, et al.
Publicado: (2025)
Less is More: on the Over-Globalizing Problem in Graph Transformers
por: Xing, Yujie, et al.
Publicado: (2024)
por: Xing, Yujie, et al.
Publicado: (2024)
The Consistency Hypothesis in Uncertainty Quantification for Large Language Models
por: Xiao, Quan, et al.
Publicado: (2025)
por: Xiao, Quan, et al.
Publicado: (2025)
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
por: Wu, Shaojin, et al.
Publicado: (2025)
por: Wu, Shaojin, et al.
Publicado: (2025)
Less is More: Towards Simple Graph Contrastive Learning
por: Zhao, Yanan, et al.
Publicado: (2025)
por: Zhao, Yanan, et al.
Publicado: (2025)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
por: Das, Rudrajit, et al.
Publicado: (2026)
por: Das, Rudrajit, et al.
Publicado: (2026)
Transformer Multivariate Forecasting: Less is More?
por: Xu, Jingjing, et al.
Publicado: (2023)
por: Xu, Jingjing, et al.
Publicado: (2023)
Less is More: Non-uniform Road Segments are Efficient for Bus Arrival Prediction
por: Huang, Zhen, et al.
Publicado: (2025)
por: Huang, Zhen, et al.
Publicado: (2025)
Less or More From Teacher: Exploiting Trilateral Geometry For Knowledge Distillation
por: Hu, Chengming, et al.
Publicado: (2023)
por: Hu, Chengming, et al.
Publicado: (2023)
Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection
por: Chen, Ruoyu, et al.
Publicado: (2025)
por: Chen, Ruoyu, et al.
Publicado: (2025)
Quantize What Counts: More for Keys, Less for Values
por: Hariri, Mohsen, et al.
Publicado: (2025)
por: Hariri, Mohsen, et al.
Publicado: (2025)
Learning More with Less: A Dynamic Dual-Level Down-Sampling Framework for Efficient Policy Optimization
por: Wang, Chao, et al.
Publicado: (2025)
por: Wang, Chao, et al.
Publicado: (2025)
Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer Learning
por: Schulte, David, et al.
Publicado: (2024)
por: Schulte, David, et al.
Publicado: (2024)
LIMR: Less is More for RL Scaling
por: Li, Xuefeng, et al.
Publicado: (2025)
por: Li, Xuefeng, et al.
Publicado: (2025)
From Deep Additive Kernel Learning to Last-Layer Bayesian Neural Networks via Induced Prior Approximation
por: Zhao, Wenyuan, et al.
Publicado: (2025)
por: Zhao, Wenyuan, et al.
Publicado: (2025)
When Less is More: The LLM Scaling Paradox in Context Compression
por: Guo, Ruishan, et al.
Publicado: (2026)
por: Guo, Ruishan, et al.
Publicado: (2026)
HyperSAT: Unsupervised Hypergraph Neural Networks for Weighted MaxSAT Problems
por: Chen, Qiyue, et al.
Publicado: (2025)
por: Chen, Qiyue, et al.
Publicado: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
por: Lou, Chao, et al.
Publicado: (2024)
por: Lou, Chao, et al.
Publicado: (2024)
Text2Weight: Bridging Natural Language and Neural Network Weight Spaces
por: Tian, Bowen, et al.
Publicado: (2025)
por: Tian, Bowen, et al.
Publicado: (2025)
Less is More: Extreme Gradient Boost Rank-1 Adaption for Efficient Finetuning of LLMs
por: Zhang, Yifei, et al.
Publicado: (2024)
por: Zhang, Yifei, et al.
Publicado: (2024)
Less is More: Adaptive Coverage for Synthetic Training Data
por: Tavakkol, Sasan, et al.
Publicado: (2025)
por: Tavakkol, Sasan, et al.
Publicado: (2025)
Self-Ablating Transformers: More Interpretability, Less Sparsity
por: Ferrao, Jeremias, et al.
Publicado: (2025)
por: Ferrao, Jeremias, et al.
Publicado: (2025)
Less is More: Rethinking Few-Shot Learning and Recurrent Neural Nets
por: Pereg, Deborah, et al.
Publicado: (2022)
por: Pereg, Deborah, et al.
Publicado: (2022)
Coreset-Induced Conditional Velocity Flow Matching
por: Wang, Xiao, et al.
Publicado: (2026)
por: Wang, Xiao, et al.
Publicado: (2026)
Artificial Geographically Weighted Neural Network: A Novel Framework for Spatial Analysis with Geographically Weighted Layers
por: Cao, Jianfei, et al.
Publicado: (2025)
por: Cao, Jianfei, et al.
Publicado: (2025)
Moirai 2.0: When Less Is More for Time Series Forecasting
por: Liu, Chenghao, et al.
Publicado: (2025)
por: Liu, Chenghao, et al.
Publicado: (2025)
Ejemplares similares
-
Gradient Flow Matching for Learning Update Dynamics in Neural Network Training
por: Shou, Xiao, et al.
Publicado: (2025) -
Self-Supervised Contrastive Pre-Training for Multivariate Point Processes
por: Shou, Xiao, et al.
Publicado: (2024) -
Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential Equation
por: Ding, Yanna, et al.
Publicado: (2024) -
Predicting Time Series of Networked Dynamical Systems without Knowing Topology
por: Ding, Yanna, et al.
Publicado: (2024) -
Optimality and NP-Hardness of Transformers in Learning Markovian Dynamical Functions
por: Ding, Yanna, et al.
Publicado: (2025)