Generalized Parallel Scaling with Interdependent Generations
Fuente:
arXiv
Salvato in:
| Autori principali: | Dong, Harry, Brandfonbrener, David, Helenowski, Eryk, He, Yun, Kumar, Mrinal, Fang, Han, Chi, Yuejie, Sankararaman, Karthik Abinav |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Reinforcement Learning from User Feedback
di: Han, Eric, et al.
Pubblicazione: (2025)
di: Han, Eric, et al.
Pubblicazione: (2025)
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
di: Dong, Harry, et al.
Pubblicazione: (2024)
di: Dong, Harry, et al.
Pubblicazione: (2024)
The Perfect Blend: Redefining RLHF with Mixture of Judges
di: Xu, Tengyu, et al.
Pubblicazione: (2024)
di: Xu, Tengyu, et al.
Pubblicazione: (2024)
Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
di: Yu, Zishun, et al.
Pubblicazione: (2025)
di: Yu, Zishun, et al.
Pubblicazione: (2025)
Controllable Discovery of Intents: Incremental Deep Clustering Using Semi-Supervised Contrastive Learning
di: Rawat, Mrinal, et al.
Pubblicazione: (2024)
di: Rawat, Mrinal, et al.
Pubblicazione: (2024)
Scalable LLM Reasoning Acceleration with Low-rank Distillation
di: Dong, Harry, et al.
Pubblicazione: (2025)
di: Dong, Harry, et al.
Pubblicazione: (2025)
On the Equivalence of Graph Convolution and Mixup
di: Han, Xiaotian, et al.
Pubblicazione: (2023)
di: Han, Xiaotian, et al.
Pubblicazione: (2023)
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
di: Dong, Harry, et al.
Pubblicazione: (2024)
di: Dong, Harry, et al.
Pubblicazione: (2024)
Do Understanding and Generation Fight? A Diagnostic Study of DPO for Unified Multimodal Models
di: Rao, Abinav, et al.
Pubblicazione: (2026)
di: Rao, Abinav, et al.
Pubblicazione: (2026)
The Role of Sparsity for Length Generalization in Transformers
di: Golowich, Noah, et al.
Pubblicazione: (2025)
di: Golowich, Noah, et al.
Pubblicazione: (2025)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
di: Huang, Yu, et al.
Pubblicazione: (2025)
di: Huang, Yu, et al.
Pubblicazione: (2025)
Contextual Bandits with Packing and Covering Constraints: A Modular Lagrangian Approach via Regression
di: Slivkins, Aleksandrs, et al.
Pubblicazione: (2022)
di: Slivkins, Aleksandrs, et al.
Pubblicazione: (2022)
Deconstructing What Makes a Good Optimizer for Language Models
di: Zhao, Rosie, et al.
Pubblicazione: (2024)
di: Zhao, Rosie, et al.
Pubblicazione: (2024)
The Art of Scaling Reinforcement Learning Compute for LLMs
di: Khatri, Devvrit, et al.
Pubblicazione: (2025)
di: Khatri, Devvrit, et al.
Pubblicazione: (2025)
Towards Low-bit Communication for Tensor Parallel LLM Inference
di: Dong, Harry, et al.
Pubblicazione: (2024)
di: Dong, Harry, et al.
Pubblicazione: (2024)
USP: A Unified Sequence Parallelism Approach for Long Context Generative AI
di: Fang, Jiarui, et al.
Pubblicazione: (2024)
di: Fang, Jiarui, et al.
Pubblicazione: (2024)
Graph Hopfield Networks: Energy-Based Node Classification with Associative Memory
di: Rao, Abinav, et al.
Pubblicazione: (2026)
di: Rao, Abinav, et al.
Pubblicazione: (2026)
Scaling Test-Time Compute for Agentic Coding
di: Kim, Joongwon, et al.
Pubblicazione: (2026)
di: Kim, Joongwon, et al.
Pubblicazione: (2026)
Repeat After Me: Transformers are Better than State Space Models at Copying
di: Jelassi, Samy, et al.
Pubblicazione: (2024)
di: Jelassi, Samy, et al.
Pubblicazione: (2024)
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
di: Wu, Bo, et al.
Pubblicazione: (2025)
di: Wu, Bo, et al.
Pubblicazione: (2025)
Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning
di: Yang, Tong, et al.
Pubblicazione: (2023)
di: Yang, Tong, et al.
Pubblicazione: (2023)
Subgraph Generation for Generalizing on Out-of-Distribution Links
di: Revolinsky, Jay, et al.
Pubblicazione: (2025)
di: Revolinsky, Jay, et al.
Pubblicazione: (2025)
Brain-Inspired Planning for Better Generalization in Reinforcement Learning
di: Zhao, Mingde "Harry"
Pubblicazione: (2025)
di: Zhao, Mingde "Harry"
Pubblicazione: (2025)
Preconditioning Benefits of Spectral Orthogonalization in Muon
di: Ma, Jianhao, et al.
Pubblicazione: (2026)
di: Ma, Jianhao, et al.
Pubblicazione: (2026)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
di: Yang, Tong, et al.
Pubblicazione: (2026)
di: Yang, Tong, et al.
Pubblicazione: (2026)
Statistical and Algorithmic Foundations of Reinforcement Learning
di: Chi, Yuejie, et al.
Pubblicazione: (2025)
di: Chi, Yuejie, et al.
Pubblicazione: (2025)
CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
Learning Discrete Concepts in Latent Hierarchical Models
di: Kong, Lingjing, et al.
Pubblicazione: (2024)
di: Kong, Lingjing, et al.
Pubblicazione: (2024)
Are Expressive Encoders Necessary for Discrete Graph Generation?
di: Revolinsky, Jay, et al.
Pubblicazione: (2026)
di: Revolinsky, Jay, et al.
Pubblicazione: (2026)
SOAP: Improving and Stabilizing Shampoo using Adam
di: Vyas, Nikhil, et al.
Pubblicazione: (2024)
di: Vyas, Nikhil, et al.
Pubblicazione: (2024)
ASAP: Exploiting the Satisficing Generalization Edge in Neural Combinatorial Optimization
di: Fang, Han, et al.
Pubblicazione: (2025)
di: Fang, Han, et al.
Pubblicazione: (2025)
OptEx: Expediting First-Order Optimization with Approximately Parallelized Iterations
di: Shu, Yao, et al.
Pubblicazione: (2024)
di: Shu, Yao, et al.
Pubblicazione: (2024)
Skills Made to Order: Efficient Acquisition of Robot Cooking Skills Guided by Multiple Forms of Internet Data
di: Verghese, Mrinal, et al.
Pubblicazione: (2024)
di: Verghese, Mrinal, et al.
Pubblicazione: (2024)
Diffusion Controller: Framework, Algorithms and Parameterization
di: Yang, Tong, et al.
Pubblicazione: (2026)
di: Yang, Tong, et al.
Pubblicazione: (2026)
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
di: Yang, Tong, et al.
Pubblicazione: (2025)
di: Yang, Tong, et al.
Pubblicazione: (2025)
Chain-of-Influence: Tracing Interdependencies Across Time and Features in Clinical Predictive Modelings
di: Li, Yubo, et al.
Pubblicazione: (2025)
di: Li, Yubo, et al.
Pubblicazione: (2025)
Scaling Up Data Parallelism in Decentralized Deep Learning
di: Xie, Bing, et al.
Pubblicazione: (2025)
di: Xie, Bing, et al.
Pubblicazione: (2025)
Uncertainty-Aware Hybrid Machine Learning in Virtual Sensors for Vehicle Sideslip Angle Estimation
di: Kalyanasundaram, Abinav, et al.
Pubblicazione: (2025)
di: Kalyanasundaram, Abinav, et al.
Pubblicazione: (2025)
PCGRL+: Scaling, Control and Generalization in Reinforcement Learning Level Generators
di: Earle, Sam, et al.
Pubblicazione: (2024)
di: Earle, Sam, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Reinforcement Learning from User Feedback
di: Han, Eric, et al.
Pubblicazione: (2025) -
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
di: Dong, Harry, et al.
Pubblicazione: (2024) -
The Perfect Blend: Redefining RLHF with Mixture of Judges
di: Xu, Tengyu, et al.
Pubblicazione: (2024) -
Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
di: Yu, Zishun, et al.
Pubblicazione: (2025) -
Controllable Discovery of Intents: Incremental Deep Clustering Using Semi-Supervised Contrastive Learning
di: Rawat, Mrinal, et al.
Pubblicazione: (2024)