TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yizhi, Gu, Qingshui, Wen, Zhoufutu, Li, Ziniu, Xing, Tianshun, Guo, Shuyue, Zheng, Tianyu, Zhou, Xin, Qu, Xingwei, Zhou, Wangchunshu, Zhang, Zheng, Shen, Wei, Liu, Qian, Lin, Chenghua, Yang, Jian, Zhang, Ge, Huang, Wenhao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
First Return, Entropy-Eliciting Explore
by: Zheng, Tianyu, et al.
Published: (2025)
by: Zheng, Tianyu, et al.
Published: (2025)
Overview of the NLPCC 2024 Shared Task on Chinese Metaphor Generation
by: Qu, Xingwei, et al.
Published: (2024)
by: Qu, Xingwei, et al.
Published: (2024)
Kun: Answer Polishment for Chinese Self-Alignment with Instruction Back-Translation
by: Zheng, Tianyu, et al.
Published: (2024)
by: Zheng, Tianyu, et al.
Published: (2024)
Steel-LLM:From Scratch to Open Source -- A Personal Journey in Building a Chinese-Centric LLM
by: Gu, Qingshui, et al.
Published: (2025)
by: Gu, Qingshui, et al.
Published: (2025)
Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
by: Qu, Xingwei, et al.
Published: (2025)
by: Qu, Xingwei, et al.
Published: (2025)
COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes
by: Li, Yunwen, et al.
Published: (2025)
by: Li, Yunwen, et al.
Published: (2025)
Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
by: Ying, Shuangshuang, et al.
Published: (2025)
by: Ying, Shuangshuang, et al.
Published: (2025)
I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm
by: Liang, Yiming, et al.
Published: (2024)
by: Liang, Yiming, et al.
Published: (2024)
CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs
by: Li, Siyi, et al.
Published: (2026)
by: Li, Siyi, et al.
Published: (2026)
Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
by: Li, Yizhi, et al.
Published: (2025)
by: Li, Yizhi, et al.
Published: (2025)
MORE-3S:Multimodal-based Offline Reinforcement Learning with Shared Semantic Spaces
by: Zheng, Tianyu, et al.
Published: (2024)
by: Zheng, Tianyu, et al.
Published: (2024)
CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models
by: LI, Yizhi, et al.
Published: (2024)
by: LI, Yizhi, et al.
Published: (2024)
CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark
by: Zhang, Ge, et al.
Published: (2024)
by: Zhang, Ge, et al.
Published: (2024)
MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models
by: Wu, Siwei, et al.
Published: (2024)
by: Wu, Siwei, et al.
Published: (2024)
SFTok: Bridging the Performance Gap in Discrete Tokenizers
by: Rao, Qihang, et al.
Published: (2025)
by: Rao, Qihang, et al.
Published: (2025)
MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants
by: Ding, Dongyi, et al.
Published: (2025)
by: Ding, Dongyi, et al.
Published: (2025)
A Comparative Study on Reasoning Patterns of OpenAI's o1 Model
by: Wu, Siwei, et al.
Published: (2024)
by: Wu, Siwei, et al.
Published: (2024)
CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
by: Guo, Jiawei, et al.
Published: (2024)
by: Guo, Jiawei, et al.
Published: (2024)
Encyclo-K: Evaluating LLMs with Dynamically Composed Knowledge Statements
by: Liang, Yiming, et al.
Published: (2025)
by: Liang, Yiming, et al.
Published: (2025)
KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks
by: Ma, Kaijing, et al.
Published: (2024)
by: Ma, Kaijing, et al.
Published: (2024)
Scaling Test-time Compute for LLM Agents
by: Zhu, King, et al.
Published: (2025)
by: Zhu, King, et al.
Published: (2025)
ConceptMoE: Adaptive Token-to-Concept Compression for Implicit Compute Allocation
by: Huang, Zihao, et al.
Published: (2026)
by: Huang, Zihao, et al.
Published: (2026)
Bridging the Gap Between Preference Alignment and Machine Unlearning
by: Feng, Xiaohua, et al.
Published: (2025)
by: Feng, Xiaohua, et al.
Published: (2025)
Bridging Efficiency and Transparency: Explainable CoT Compression in Multimodal Large Reasoning Models
by: Wang, Yizhi, et al.
Published: (2026)
by: Wang, Yizhi, et al.
Published: (2026)
DocMMIR: A Framework for Document Multi-modal Information Retrieval
by: Li, Zirui, et al.
Published: (2025)
by: Li, Zirui, et al.
Published: (2025)
BMP: Bridging the Gap between B-Spline and Movement Primitives
by: Liao, Weiran, et al.
Published: (2024)
by: Liao, Weiran, et al.
Published: (2024)
LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm
by: Wu, Siwei, et al.
Published: (2025)
by: Wu, Siwei, et al.
Published: (2025)
MuPT: A Generative Symbolic Music Pretrained Transformer
by: Qu, Xingwei, et al.
Published: (2024)
by: Qu, Xingwei, et al.
Published: (2024)
TEGEE: Task dEfinition Guided Expert Ensembling for Generalizable and Few-shot Learning
by: Qu, Xingwei, et al.
Published: (2024)
by: Qu, Xingwei, et al.
Published: (2024)
CMDAG: A Chinese Metaphor Dataset with Annotated Grounds as CoT for Boosting Metaphor Generation
by: Shao, Yujie, et al.
Published: (2024)
by: Shao, Yujie, et al.
Published: (2024)
CAE-DFKD: Bridging the Transferability Gap in Data-Free Knowledge Distillation
by: Zhang, Zherui, et al.
Published: (2025)
by: Zhang, Zherui, et al.
Published: (2025)
Observing Micromotives and Macrobehavior of Large Language Models
by: Cheng, Yuyang, et al.
Published: (2024)
by: Cheng, Yuyang, et al.
Published: (2024)
Bridging the Copyright Gap: Do Large Vision-Language Models Recognize and Respect Copyrighted Content?
by: Xu, Naen, et al.
Published: (2025)
by: Xu, Naen, et al.
Published: (2025)
METTL3/METTL14‐mediated RNA m 6 A modification is involved in male reproductive development in Bactrocera dorsalis
by: Qiuyuan Zhang, et al.
Published: (2025)
by: Qiuyuan Zhang, et al.
Published: (2025)
SIA-OVD: Shape-Invariant Adapter for Bridging the Image-Region Gap in Open-Vocabulary Detection
by: Wang, Zishuo, et al.
Published: (2024)
by: Wang, Zishuo, et al.
Published: (2024)
Aligning Instruction Tuning with Pre-training
by: Liang, Yiming, et al.
Published: (2025)
by: Liang, Yiming, et al.
Published: (2025)
Few-Round Distributed Principal Component Analysis: Closing the Statistical Efficiency Gap by Consensus
by: Li, ZeYu, et al.
Published: (2025)
by: Li, ZeYu, et al.
Published: (2025)
Targeting the E1 ubiquitin‐activating enzyme Uba1 impairs male fertility in Bactrocera dorsalis
by: Jiao Qiao, et al.
Published: (2025)
by: Jiao Qiao, et al.
Published: (2025)
SurgeryV2: Bridging the Gap Between Model Merging and Multi-Task Learning with Deep Representation Surgery
by: Yang, Enneng, et al.
Published: (2024)
by: Yang, Enneng, et al.
Published: (2024)
Bridging the Initialization Gap: A Co-Optimization Framework for Mixed-Size Global Placement
by: Ren, Yuhao, et al.
Published: (2025)
by: Ren, Yuhao, et al.
Published: (2025)
Similar Items
-
First Return, Entropy-Eliciting Explore
by: Zheng, Tianyu, et al.
Published: (2025) -
Overview of the NLPCC 2024 Shared Task on Chinese Metaphor Generation
by: Qu, Xingwei, et al.
Published: (2024) -
Kun: Answer Polishment for Chinese Self-Alignment with Instruction Back-Translation
by: Zheng, Tianyu, et al.
Published: (2024) -
Steel-LLM:From Scratch to Open Source -- A Personal Journey in Building a Chinese-Centric LLM
by: Gu, Qingshui, et al.
Published: (2025) -
Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
by: Qu, Xingwei, et al.
Published: (2025)