Recurrent Diffusion for Large-Scale Parameter Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Kai, Tang, Dongwen, Zhao, Wangbo, Schürholt, Konstantin, Wang, Zhangyang, You, Yang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Conditional LoRA Parameter Generation
di: Jin, Xiaolong, et al.
Pubblicazione: (2024)
di: Jin, Xiaolong, et al.
Pubblicazione: (2024)
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
di: Liang, Zhiyuan, et al.
Pubblicazione: (2025)
di: Liang, Zhiyuan, et al.
Pubblicazione: (2025)
ORAL: Prompting Your Large-Scale LoRAs via Conditional Recurrent Diffusion
di: Khan, Rana Muhammad Shahroz, et al.
Pubblicazione: (2025)
di: Khan, Rana Muhammad Shahroz, et al.
Pubblicazione: (2025)
Position: Weight Space Should Be a First-Class Generative AI Modality
di: Wang, Zhangyang, et al.
Pubblicazione: (2026)
di: Wang, Zhangyang, et al.
Pubblicazione: (2026)
Toward Dynamic Stability Assessment of Power Grid Topologies using Graph Neural Networks
di: Nauck, Christian, et al.
Pubblicazione: (2022)
di: Nauck, Christian, et al.
Pubblicazione: (2022)
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
di: Dong, Harry, et al.
Pubblicazione: (2024)
di: Dong, Harry, et al.
Pubblicazione: (2024)
Hyper-Representations: Learning from Populations of Neural Networks
di: Schürholt, Konstantin
Pubblicazione: (2024)
di: Schürholt, Konstantin
Pubblicazione: (2024)
LLaGA: Large Language and Graph Assistant
di: Chen, Runjin, et al.
Pubblicazione: (2024)
di: Chen, Runjin, et al.
Pubblicazione: (2024)
Enhancing Parameter Efficiency and Generalization in Large-Scale Models: A Regularized and Masked Low-Rank Adaptation Approach
di: Mao, Yuzhu, et al.
Pubblicazione: (2024)
di: Mao, Yuzhu, et al.
Pubblicazione: (2024)
Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization
di: Chaubard, Francois, et al.
Pubblicazione: (2025)
di: Chaubard, Francois, et al.
Pubblicazione: (2025)
Prioritize Alignment in Dataset Distillation
di: Li, Zekai, et al.
Pubblicazione: (2024)
di: Li, Zekai, et al.
Pubblicazione: (2024)
Neon: Negative Extrapolation From Self-Training Improves Image Generation
di: Alemohammad, Sina, et al.
Pubblicazione: (2025)
di: Alemohammad, Sina, et al.
Pubblicazione: (2025)
READ: Recurrent Adaptation of Large Transformers
di: Nguyen, John, et al.
Pubblicazione: (2023)
di: Nguyen, John, et al.
Pubblicazione: (2023)
pFedGPA: Diffusion-based Generative Parameter Aggregation for Personalized Federated Learning
di: Lai, Jiahao, et al.
Pubblicazione: (2024)
di: Lai, Jiahao, et al.
Pubblicazione: (2024)
PIPA: Preference Alignment as Prior-Informed Statistical Estimation
di: Li, Junbo, et al.
Pubblicazione: (2025)
di: Li, Junbo, et al.
Pubblicazione: (2025)
RDPI: A Refine Diffusion Probability Generation Method for Spatiotemporal Data Imputation
di: Liu, Zijin, et al.
Pubblicazione: (2024)
di: Liu, Zijin, et al.
Pubblicazione: (2024)
Data Distribution as a Lever for Guiding Optimizers Toward Superior Generalization in LLMs
di: Gangavarapu, Tushaar, et al.
Pubblicazione: (2026)
di: Gangavarapu, Tushaar, et al.
Pubblicazione: (2026)
FaultDiffusion: Few-Shot Fault Time Series Generation with Diffusion Model
di: Xu, Yi, et al.
Pubblicazione: (2025)
di: Xu, Yi, et al.
Pubblicazione: (2025)
Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial Images
di: Wang, Yubo, et al.
Pubblicazione: (2025)
di: Wang, Yubo, et al.
Pubblicazione: (2025)
AIGC for Industrial Time Series: From Deep Generative Models to Large Generative Models
di: Ren, Lei, et al.
Pubblicazione: (2024)
di: Ren, Lei, et al.
Pubblicazione: (2024)
Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
di: Rodkin, Ivan, et al.
Pubblicazione: (2025)
di: Rodkin, Ivan, et al.
Pubblicazione: (2025)
Scaling On-Device GPU Inference for Large Generative Models
di: Tang, Jiuqiang, et al.
Pubblicazione: (2025)
di: Tang, Jiuqiang, et al.
Pubblicazione: (2025)
Collaborative Compression for Large-Scale MoE Deployment on Edge
di: Chen, Yixiao, et al.
Pubblicazione: (2025)
di: Chen, Yixiao, et al.
Pubblicazione: (2025)
PRDP: Proximal Reward Difference Prediction for Large-Scale Reward Finetuning of Diffusion Models
di: Deng, Fei, et al.
Pubblicazione: (2024)
di: Deng, Fei, et al.
Pubblicazione: (2024)
When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search
di: Robertson, John T., et al.
Pubblicazione: (2026)
di: Robertson, John T., et al.
Pubblicazione: (2026)
Synthetic Data Generation for Residential Load Patterns via Recurrent GAN and Ensemble Method
di: Liang, Xinyu, et al.
Pubblicazione: (2024)
di: Liang, Xinyu, et al.
Pubblicazione: (2024)
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
di: Zhang, Wuyue, et al.
Pubblicazione: (2026)
di: Zhang, Wuyue, et al.
Pubblicazione: (2026)
Ferret: Federated Full-Parameter Tuning at Scale for Large Language Models
di: Shu, Yao, et al.
Pubblicazione: (2024)
di: Shu, Yao, et al.
Pubblicazione: (2024)
PIMRL: Physics-Informed Multi-Scale Recurrent Learning for Burst-Sampled Spatiotemporal Dynamics
di: Wan, Han, et al.
Pubblicazione: (2025)
di: Wan, Han, et al.
Pubblicazione: (2025)
Model-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild
di: Zhao, Xinyu, et al.
Pubblicazione: (2024)
di: Zhao, Xinyu, et al.
Pubblicazione: (2024)
Diffusion In Diffusion: Reclaiming Global Coherence in Semi-Autoregressive Diffusion
di: Ma, Linrui, et al.
Pubblicazione: (2026)
di: Ma, Linrui, et al.
Pubblicazione: (2026)
GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets
di: Yao, Zhangyang, et al.
Pubblicazione: (2026)
di: Yao, Zhangyang, et al.
Pubblicazione: (2026)
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models
di: Wang, Mingze, et al.
Pubblicazione: (2026)
di: Wang, Mingze, et al.
Pubblicazione: (2026)
Large Language Models to Diffusion Finetuning
di: Cetin, Edoardo, et al.
Pubblicazione: (2025)
di: Cetin, Edoardo, et al.
Pubblicazione: (2025)
Prodigy: An Expeditiously Adaptive Parameter-Free Learner
di: Mishchenko, Konstantin, et al.
Pubblicazione: (2023)
di: Mishchenko, Konstantin, et al.
Pubblicazione: (2023)
G2PDiffusion: Cross-Species Genotype-to-Phenotype Prediction via Evolutionary Diffusion
di: Liu, Mengdi, et al.
Pubblicazione: (2025)
di: Liu, Mengdi, et al.
Pubblicazione: (2025)
General and Efficient Steering of Unconditional Diffusion
di: Wang, Qingsong, et al.
Pubblicazione: (2026)
di: Wang, Qingsong, et al.
Pubblicazione: (2026)
Communication-Efficient Diffusion Denoising Parallelization via Reuse-then-Predict Mechanism
di: Wang, Kunyun, et al.
Pubblicazione: (2025)
di: Wang, Kunyun, et al.
Pubblicazione: (2025)
Memory-Efficient Gradient Unrolling for Large-Scale Bi-level Optimization
di: Shen, Qianli, et al.
Pubblicazione: (2024)
di: Shen, Qianli, et al.
Pubblicazione: (2024)
Scaling and Transferability of Annealing Strategies in Large Language Model Training
di: Wang, Siqi, et al.
Pubblicazione: (2025)
di: Wang, Siqi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Conditional LoRA Parameter Generation
di: Jin, Xiaolong, et al.
Pubblicazione: (2024) -
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
di: Liang, Zhiyuan, et al.
Pubblicazione: (2025) -
ORAL: Prompting Your Large-Scale LoRAs via Conditional Recurrent Diffusion
di: Khan, Rana Muhammad Shahroz, et al.
Pubblicazione: (2025) -
Position: Weight Space Should Be a First-Class Generative AI Modality
di: Wang, Zhangyang, et al.
Pubblicazione: (2026) -
Toward Dynamic Stability Assessment of Power Grid Topologies using Graph Neural Networks
di: Nauck, Christian, et al.
Pubblicazione: (2022)