TSPO: Breaking the Double Homogenization Dilemma in Multi-turn Search Policy Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Shichao, Ma, Zhiyuan, Yang, Ming, Li, Xiaofan, Wu, Xing, Du, Jintao, Cheng, Yu, Wang, Weiqiang, Liu, Qiliang, Zhou, Zhengyang, Wang, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off
von: Li, Xiaofan, et al.
Veröffentlicht: (2026)
von: Li, Xiaofan, et al.
Veröffentlicht: (2026)
HAMMER: Hamiltonian Curiosity Augmented Large Language Model Reinforcement
von: Yang, Ming, et al.
Veröffentlicht: (2025)
von: Yang, Ming, et al.
Veröffentlicht: (2025)
Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and Editing
von: Ma, Shichao, et al.
Veröffentlicht: (2025)
von: Ma, Shichao, et al.
Veröffentlicht: (2025)
QuiZSF: A Retrieval-Augmented Framework for Zero-Shot Time Series Forecasting
von: Ma, Shichao, et al.
Veröffentlicht: (2025)
von: Ma, Shichao, et al.
Veröffentlicht: (2025)
Mirror-Consistency: Harnessing Inconsistency in Majority Voting
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)
Gumbel Reranking: Differentiable End-to-End Reranker Optimization
von: Huang, Siyuan, et al.
Veröffentlicht: (2025)
von: Huang, Siyuan, et al.
Veröffentlicht: (2025)
Bayesian Rational Search Engine User
von: Ma, Shichao
Veröffentlicht: (2026)
von: Ma, Shichao
Veröffentlicht: (2026)
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
von: Gu, Chenyang, et al.
Veröffentlicht: (2025)
von: Gu, Chenyang, et al.
Veröffentlicht: (2025)
Double-Strand Break Clustering: An Economical and Effective Strategy for DNA Repair
von: Chen, Junyi, et al.
Veröffentlicht: (2024)
von: Chen, Junyi, et al.
Veröffentlicht: (2024)
Joint Beamforming Optimization and Mode Selection for RDARS-Aided MIMO Systems
von: Wang, Jintao, et al.
Veröffentlicht: (2024)
von: Wang, Jintao, et al.
Veröffentlicht: (2024)
Retention Induced Biases in a Recommendation System with Heterogeneous Users
von: Ma, Shichao
Veröffentlicht: (2024)
von: Ma, Shichao
Veröffentlicht: (2024)
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
von: Ma, Chang, et al.
Veröffentlicht: (2024)
von: Ma, Chang, et al.
Veröffentlicht: (2024)
Workflow-R1: Group Sub-sequence Policy Optimization for Multi-turn Workflow Construction
von: Kong, Mingze, et al.
Veröffentlicht: (2026)
von: Kong, Mingze, et al.
Veröffentlicht: (2026)
TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
Homogeneity Test of Proportions for Combined Unilateral and Bilateral Data via GEE and MLE Approaches
von: Zhou, Jia, et al.
Veröffentlicht: (2025)
von: Zhou, Jia, et al.
Veröffentlicht: (2025)
MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety
von: Song, Jialin, et al.
Veröffentlicht: (2026)
von: Song, Jialin, et al.
Veröffentlicht: (2026)
Trait-Aware Policy Optimization for Autoregressive Multi-Trait Essay Scoring
von: Wang, Zhengyang, et al.
Veröffentlicht: (2026)
von: Wang, Zhengyang, et al.
Veröffentlicht: (2026)
Breaking Robustness Barriers in Cognitive Diagnosis: A One-Shot Neural Architecture Search Perspective
von: Wang, Ziwen, et al.
Veröffentlicht: (2026)
von: Wang, Ziwen, et al.
Veröffentlicht: (2026)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
rs6971 TSPO polymorphism in Parkinson's disease
von: Bina Patel, et al.
Veröffentlicht: (2025)
von: Bina Patel, et al.
Veröffentlicht: (2025)
Characterizing the Dilemma of Performance and Index Size in Billion-Scale Vector Search and Breaking It with Second-Tier Memory
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning
von: Zhao, Zhengyang, et al.
Veröffentlicht: (2026)
von: Zhao, Zhengyang, et al.
Veröffentlicht: (2026)
A General ReLearner: Empowering Spatiotemporal Prediction by Re-learning Input-label Residual
von: Ma, Jiaming, et al.
Veröffentlicht: (2026)
von: Ma, Jiaming, et al.
Veröffentlicht: (2026)
HoneyGPT: Breaking the Trilemma in Terminal Honeypots with Large Language Model
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning
von: Li, Zihao, et al.
Veröffentlicht: (2024)
von: Li, Zihao, et al.
Veröffentlicht: (2024)
Robust Beamforming Design and Antenna Selection for Dynamic HRIS-aided MISO System
von: Wang, Jintao, et al.
Veröffentlicht: (2024)
von: Wang, Jintao, et al.
Veröffentlicht: (2024)
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
von: Yao, Jingfeng, et al.
Veröffentlicht: (2025)
von: Yao, Jingfeng, et al.
Veröffentlicht: (2025)
A Nonlinear Multi‐Objective Prediction Strategy for Small‐Sample Datasets in Homogeneous Catalysis
von: Yining Liu, et al.
Veröffentlicht: (2026)
von: Yining Liu, et al.
Veröffentlicht: (2026)
A novel chromone Schiff base as Zn2+ turn‐on fluorescent chemosensor in a mixed solution
von: Wensheng Yang, et al.
Veröffentlicht: (2024)
von: Wensheng Yang, et al.
Veröffentlicht: (2024)
SAPIENT: Mastering Multi-turn Conversational Recommendation with Strategic Planning and Monte Carlo Tree Search
von: Du, Hanwen, et al.
Veröffentlicht: (2024)
von: Du, Hanwen, et al.
Veröffentlicht: (2024)
LaPuda: LLM-Enabled Policy-Based Query Optimizer for Multi-modal Data
von: Wang, Yifan, et al.
Veröffentlicht: (2024)
von: Wang, Yifan, et al.
Veröffentlicht: (2024)
Breaking the Under‐Display Camera's Dilemma Between Diffraction and Pixel Density Using Incoherent Pupil Synthesis
von: Xinni Xie, et al.
Veröffentlicht: (2026)
von: Xinni Xie, et al.
Veröffentlicht: (2026)
On the elastic–plastic behaviors of centrifugal steam compressor impeller under cyclic thermomechanical loading
von: Jianxun Li, et al.
Veröffentlicht: (2024)
von: Jianxun Li, et al.
Veröffentlicht: (2024)
Double-Edged Sword or Sharp Tool? Designing and Evaluating Triadic LLM-Teacher Collaboration for K-12 Writing at Scale
von: Wang, Canran, et al.
Veröffentlicht: (2026)
von: Wang, Canran, et al.
Veröffentlicht: (2026)
Near-Optimal Policy Optimization for Correlated Equilibrium in General-Sum Markov Games
von: Cai, Yang, et al.
Veröffentlicht: (2024)
von: Cai, Yang, et al.
Veröffentlicht: (2024)
CaliCausalRank: Calibrated Multi-Objective Ad Ranking with Robust Counterfactual Utility Optimization
von: Yang, Xikai, et al.
Veröffentlicht: (2026)
von: Yang, Xikai, et al.
Veröffentlicht: (2026)
Entrepreneurship and Natural Resource Curse on Regional Economic Resilience: Evidence From China
von: Mao Qiliang, et al.
Veröffentlicht: (2025)
von: Mao Qiliang, et al.
Veröffentlicht: (2025)
Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization
von: Hao, Ruijie, et al.
Veröffentlicht: (2026)
von: Hao, Ruijie, et al.
Veröffentlicht: (2026)
A Low-Overhead Incorporation-Extrapolation based Few-Shot CSI Feedback Framework for Massive MIMO Systems
von: Zhou, Binggui, et al.
Veröffentlicht: (2023)
von: Zhou, Binggui, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off
von: Li, Xiaofan, et al.
Veröffentlicht: (2026) -
HAMMER: Hamiltonian Curiosity Augmented Large Language Model Reinforcement
von: Yang, Ming, et al.
Veröffentlicht: (2025) -
Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and Editing
von: Ma, Shichao, et al.
Veröffentlicht: (2025) -
QuiZSF: A Retrieval-Augmented Framework for Zero-Shot Time Series Forecasting
von: Ma, Shichao, et al.
Veröffentlicht: (2025) -
Mirror-Consistency: Harnessing Inconsistency in Majority Voting
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)