Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Xiaolong, Ma, Lichen, Guo, Zipeng, Dong, ShiPing, Yang, Lan, Sin, Tan Lit, Zhou, Gaojing, He, Yu, Fu, Jingling, Zhou, Shizhe, Huang, Junshi, Li, Jason |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
by: Ma, Lichen, et al.
Published: (2026)
by: Ma, Lichen, et al.
Published: (2026)
TreeRPO: Tree Relative Policy Optimization
by: Yang, Zhicheng, et al.
Published: (2025)
by: Yang, Zhicheng, et al.
Published: (2025)
FrequencyBooster: Full-Frequency Modeling for High-Fidelity Pixel Diffusion
by: Ma, Lichen, et al.
Published: (2026)
by: Ma, Lichen, et al.
Published: (2026)
RePainter: Empowering E-commerce Object Removal via Spatial-matting Reinforcement Learning
by: Guo, Zipeng, et al.
Published: (2025)
by: Guo, Zipeng, et al.
Published: (2025)
Meta-TTRL: A Metacognitive Framework for Self-Improving Test-Time Reinforcement Learning in Unified Multimodal Models
by: Tan, Lit Sin, et al.
Published: (2026)
by: Tan, Lit Sin, et al.
Published: (2026)
d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
by: Pan, Leyi, et al.
Published: (2025)
by: Pan, Leyi, et al.
Published: (2025)
HyperDiT: Hyper-Connected Transformers for High-Fidelity Pixel-Space Diffusion
by: He, Yu, et al.
Published: (2026)
by: He, Yu, et al.
Published: (2026)
Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition
by: He, Yu, et al.
Published: (2026)
by: He, Yu, et al.
Published: (2026)
LiWi: Layering in the Wild
by: He, Yu, et al.
Published: (2026)
by: He, Yu, et al.
Published: (2026)
StaRPO: Stability-Augmented Reinforcement Policy Optimization
by: Zhang, Jinghan, et al.
Published: (2026)
by: Zhang, Jinghan, et al.
Published: (2026)
Zigzag Diffusion Sampling: Diffusion Models Can Self-Improve via Self-Reflection
by: Bai, Lichen, et al.
Published: (2024)
by: Bai, Lichen, et al.
Published: (2024)
Diffusion-RPO: Aligning Diffusion Models through Relative Preference Optimization
by: Gu, Yi, et al.
Published: (2024)
by: Gu, Yi, et al.
Published: (2024)
ScRPO: From Errors to Insights
by: Li, Lianrui, et al.
Published: (2025)
by: Li, Lianrui, et al.
Published: (2025)
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
by: Wu, Yongtong, et al.
Published: (2026)
by: Wu, Yongtong, et al.
Published: (2026)
Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models
by: Wu, Qi, et al.
Published: (2024)
by: Wu, Qi, et al.
Published: (2024)
Not All Noises Are Created Equally:Diffusion Noise Selection and Optimization
by: Qi, Zipeng, et al.
Published: (2024)
by: Qi, Zipeng, et al.
Published: (2024)
Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models
by: Li, Zongqian, et al.
Published: (2026)
by: Li, Zongqian, et al.
Published: (2026)
Boundedness of New Type Fourier Integral Operators with Product Structure
by: Tan, Chaoqiang, et al.
Published: (2024)
by: Tan, Chaoqiang, et al.
Published: (2024)
Breaking the Attention Bottleneck
by: Hilsenbek, Kalle
Published: (2024)
by: Hilsenbek, Kalle
Published: (2024)
Breaking the Context Bottleneck on Long Time Series Forecasting
by: Ma, Chao, et al.
Published: (2024)
by: Ma, Chao, et al.
Published: (2024)
Do Synthetic Trajectories Reflect Real Reward Hacking? A Systematic Study on Monitoring In-the-Wild Hacking in Code Generation
by: Li, Lichen, et al.
Published: (2026)
by: Li, Lichen, et al.
Published: (2026)
LTCF-Net: A Transformer-Enhanced Dual-Channel Fourier Framework for Low-Light Image Restoration
by: Zhang, Gaojing, et al.
Published: (2024)
by: Zhang, Gaojing, et al.
Published: (2024)
AOEPT: Breaking the Implicit Modality-Reduction Bottleneck in Modality-Missing Prompt Tuning
by: Lang, Jian, et al.
Published: (2026)
by: Lang, Jian, et al.
Published: (2026)
Breaking the Recycling Bottleneck of Thermosets via Bio‐Tailoring Technology
by: Jinping Yu, et al.
Published: (2025)
by: Jinping Yu, et al.
Published: (2025)
RPO:Reinforcement Fine-Tuning with Partial Reasoning Optimization
by: Yi, Hongzhu, et al.
Published: (2026)
by: Yi, Hongzhu, et al.
Published: (2026)
GTA-Net: An IoT-Integrated 3D Human Pose Estimation System for Real-Time Adolescent Sports Posture Correction
by: Yuan, Shizhe, et al.
Published: (2024)
by: Yuan, Shizhe, et al.
Published: (2024)
Breaking AR's Sampling Bottleneck: Provable Acceleration via Diffusion Language Models
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
Breaking Through the Luminescence Stability Bottleneck of Oxyfluoride Phosphor for Sun‐Like Led Lighting
by: Ying Li, et al.
Published: (2024)
by: Ying Li, et al.
Published: (2024)
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering
by: Fu, Xingyu, et al.
Published: (2023)
by: Fu, Xingyu, et al.
Published: (2023)
The Debate on the Dietary Guidelines for Americans (2025–2030) and Implications for China's Nutritional Policy
by: Junshi Chen
Published: (2026)
by: Junshi Chen
Published: (2026)
T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback
by: Li, Jiachen, et al.
Published: (2024)
by: Li, Jiachen, et al.
Published: (2024)
RPO: Retrieval Preference Optimization for Robust Retrieval-Augmented Generation
by: Yan, Shi-Qi, et al.
Published: (2025)
by: Yan, Shi-Qi, et al.
Published: (2025)
Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
Breaking Symmetry Bottlenecks in GNN Readouts
by: Talhi, Mouad, et al.
Published: (2026)
by: Talhi, Mouad, et al.
Published: (2026)
EMATO: Energy-Model-Aware Trajectory Optimization for Autonomous Driving
by: Tian, Zhaofeng, et al.
Published: (2024)
by: Tian, Zhaofeng, et al.
Published: (2024)
Poisoning Prompt-Guided Sampling in Video Large Language Models
by: Cao, Yuxin, et al.
Published: (2025)
by: Cao, Yuxin, et al.
Published: (2025)
UMI on Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole-body Controllers
by: Ha, Huy, et al.
Published: (2024)
by: Ha, Huy, et al.
Published: (2024)
From Natural Language to Silicon: The Representation Bottleneck in LLM Hardware Design
by: Fu, Weimin, et al.
Published: (2026)
by: Fu, Weimin, et al.
Published: (2026)
Clinical Characteristics and Independent Risk Factors for Pathologic Nipple Discharge of 375 Cases
by: Junyue Wang, et al.
Published: (2025)
by: Junyue Wang, et al.
Published: (2025)
Artifact for "Structure-Aware Delta Debugging with Geometric-Information Weights"
by: Tao, Yonggang, et al.
Published: (2026)
by: Tao, Yonggang, et al.
Published: (2026)
Similar Items
-
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
by: Ma, Lichen, et al.
Published: (2026) -
TreeRPO: Tree Relative Policy Optimization
by: Yang, Zhicheng, et al.
Published: (2025) -
FrequencyBooster: Full-Frequency Modeling for High-Fidelity Pixel Diffusion
by: Ma, Lichen, et al.
Published: (2026) -
RePainter: Empowering E-commerce Object Removal via Spatial-matting Reinforcement Learning
by: Guo, Zipeng, et al.
Published: (2025) -
Meta-TTRL: A Metacognitive Framework for Self-Improving Test-Time Reinforcement Learning in Unified Multimodal Models
by: Tan, Lit Sin, et al.
Published: (2026)