Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bartoldson, Brian, Venkatraman, Siddarth, Diffenderfer, James, Jain, Moksh, Ben-Nun, Tal, Lee, Seanie, Kim, Minsu, Obando-Ceron, Johan, Bengio, Yoshua, Kailkhura, Bhavya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
von: Venkatraman, Siddarth, et al.
Veröffentlicht: (2025)
von: Venkatraman, Siddarth, et al.
Veröffentlicht: (2025)
A Comedy of Estimators: On KL Regularization in RL Training of LLMs
von: Shah, Vedant, et al.
Veröffentlicht: (2025)
von: Shah, Vedant, et al.
Veröffentlicht: (2025)
Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies
von: Bartoldson, Brian R., et al.
Veröffentlicht: (2024)
von: Bartoldson, Brian R., et al.
Veröffentlicht: (2024)
Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion
von: Christopher, Jacob K, et al.
Veröffentlicht: (2024)
von: Christopher, Jacob K, et al.
Veröffentlicht: (2024)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
End-to-End Mesh Optimization of a Hybrid Deep Learning Black-Box PDE Solver
von: Ma, Shaocong, et al.
Veröffentlicht: (2024)
von: Ma, Shaocong, et al.
Veröffentlicht: (2024)
Outsourced diffusion sampling: Efficient posterior inference in latent spaces of generative models
von: Venkatraman, Siddarth, et al.
Veröffentlicht: (2025)
von: Venkatraman, Siddarth, et al.
Veröffentlicht: (2025)
Learning diverse attacks on large language models for robust red-teaming and safety tuning
von: Lee, Seanie, et al.
Veröffentlicht: (2024)
von: Lee, Seanie, et al.
Veröffentlicht: (2024)
Amortizing intractable inference in diffusion models for vision, language, and control
von: Venkatraman, Siddarth, et al.
Veröffentlicht: (2024)
von: Venkatraman, Siddarth, et al.
Veröffentlicht: (2024)
Relative Trajectory Balance is equivalent to Trust-PCL
von: Deleu, Tristan, et al.
Veröffentlicht: (2025)
von: Deleu, Tristan, et al.
Veröffentlicht: (2025)
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2025)
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2025)
Multi-Fidelity Active Learning with GFlowNets
von: Hernandez-Garcia, Alex, et al.
Veröffentlicht: (2023)
von: Hernandez-Garcia, Alex, et al.
Veröffentlicht: (2023)
Latent Veracity Inference for Identifying Errors in Stepwise Reasoning
von: Kim, Minsu, et al.
Veröffentlicht: (2025)
von: Kim, Minsu, et al.
Veröffentlicht: (2025)
Forecasting Fails: Unveiling Evasion Attacks in Weather Prediction Models
von: Arif, Huzaifa, et al.
Veröffentlicht: (2025)
von: Arif, Huzaifa, et al.
Veröffentlicht: (2025)
Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
von: McDonald, Tavish, et al.
Veröffentlicht: (2025)
von: McDonald, Tavish, et al.
Veröffentlicht: (2025)
Solving Bayesian inverse problems with diffusion priors and off-policy RL
von: Scimeca, Luca, et al.
Veröffentlicht: (2025)
von: Scimeca, Luca, et al.
Veröffentlicht: (2025)
UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making
von: Duan, Jinhao, et al.
Veröffentlicht: (2025)
von: Duan, Jinhao, et al.
Veröffentlicht: (2025)
Amortizing intractable inference in large language models
von: Hu, Edward J., et al.
Veröffentlicht: (2023)
von: Hu, Edward J., et al.
Veröffentlicht: (2023)
BOOM: Benchmarking Out-Of-distribution Molecular Property Predictions of Machine Learning Models
von: Antoniuk, Evan R., et al.
Veröffentlicht: (2025)
von: Antoniuk, Evan R., et al.
Veröffentlicht: (2025)
PhyloGFN: Phylogenetic inference with generative flow networks
von: Zhou, Mingyang, et al.
Veröffentlicht: (2023)
von: Zhou, Mingyang, et al.
Veröffentlicht: (2023)
Action abstractions for amortized sampling
von: Boussif, Oussama, et al.
Veröffentlicht: (2024)
von: Boussif, Oussama, et al.
Veröffentlicht: (2024)
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
von: Zheng, Haizhong, et al.
Veröffentlicht: (2025)
von: Zheng, Haizhong, et al.
Veröffentlicht: (2025)
ELFS: Label-Free Coreset Selection with Proxy Training Dynamics
von: Zheng, Haizhong, et al.
Veröffentlicht: (2024)
von: Zheng, Haizhong, et al.
Veröffentlicht: (2024)
Local Search GFlowNets
von: Kim, Minsu, et al.
Veröffentlicht: (2023)
von: Kim, Minsu, et al.
Veröffentlicht: (2023)
Machine learning and information theory concepts towards an AI Mathematician
von: Bengio, Yoshua, et al.
Veröffentlicht: (2024)
von: Bengio, Yoshua, et al.
Veröffentlicht: (2024)
SOUL: Unlocking the Power of Second-Order Optimization for LLM Unlearning
von: Jia, Jinghan, et al.
Veröffentlicht: (2024)
von: Jia, Jinghan, et al.
Veröffentlicht: (2024)
AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security
von: Cai, Zikui, et al.
Veröffentlicht: (2025)
von: Cai, Zikui, et al.
Veröffentlicht: (2025)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
Modeling Code: Is Text All You Need?
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
Fast Monte Carlo Tree Diffusion: 100x Speedup via Parallel Sparse Planning
von: Yoon, Jaesik, et al.
Veröffentlicht: (2025)
von: Yoon, Jaesik, et al.
Veröffentlicht: (2025)
Generative Recursive Reasoning
von: Baek, Junyeob, et al.
Veröffentlicht: (2026)
von: Baek, Junyeob, et al.
Veröffentlicht: (2026)
Active Attacks: Red-teaming LLMs via Adaptive Environments
von: Yun, Taeyoung, et al.
Veröffentlicht: (2025)
von: Yun, Taeyoung, et al.
Veröffentlicht: (2025)
Ant Colony Sampling with GFlowNets for Combinatorial Optimization
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
Baking Symmetry into GFlowNets
von: Ma, George, et al.
Veröffentlicht: (2024)
von: Ma, George, et al.
Veröffentlicht: (2024)
Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered
von: Liu, Sijia, et al.
Veröffentlicht: (2026)
von: Liu, Sijia, et al.
Veröffentlicht: (2026)
TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
von: Duan, Jinhao, et al.
Veröffentlicht: (2025)
von: Duan, Jinhao, et al.
Veröffentlicht: (2025)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
von: Hong, Junyuan, et al.
Veröffentlicht: (2024)
von: Hong, Junyuan, et al.
Veröffentlicht: (2024)
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
von: Duan, Jinhao, et al.
Veröffentlicht: (2024)
von: Duan, Jinhao, et al.
Veröffentlicht: (2024)
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
von: Wang, Zijun, et al.
Veröffentlicht: (2025)
von: Wang, Zijun, et al.
Veröffentlicht: (2025)
Learning to Scale Logits for Temperature-Conditional GFlowNets
von: Kim, Minsu, et al.
Veröffentlicht: (2023)
von: Kim, Minsu, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
von: Venkatraman, Siddarth, et al.
Veröffentlicht: (2025) -
A Comedy of Estimators: On KL Regularization in RL Training of LLMs
von: Shah, Vedant, et al.
Veröffentlicht: (2025) -
Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies
von: Bartoldson, Brian R., et al.
Veröffentlicht: (2024) -
Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion
von: Christopher, Jacob K, et al.
Veröffentlicht: (2024) -
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)