SSPO: Subsentence-level Policy Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Kun, chen, Zikang, Wang, Yanmeng, Li, Zhigen, Cheng, Ning, Wang, Shaojun, Xiao, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning to Adapt to Low-Resource Paraphrase Generation
von: Li, Zhigen, et al.
Veröffentlicht: (2024)
von: Li, Zhigen, et al.
Veröffentlicht: (2024)
Astra: Activation-Space Tail-Eigenvector Low-Rank Adaptation of Large Language Models
von: Liu, Kainan, et al.
Veröffentlicht: (2026)
von: Liu, Kainan, et al.
Veröffentlicht: (2026)
Verifiable Generation with Subsentence-Level Fine-Grained Citations
von: Cao, Shuyang, et al.
Veröffentlicht: (2024)
von: Cao, Shuyang, et al.
Veröffentlicht: (2024)
Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression
von: Zhang, Yong, et al.
Veröffentlicht: (2025)
von: Zhang, Yong, et al.
Veröffentlicht: (2025)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
von: Li, Zhuochun, et al.
Veröffentlicht: (2026)
von: Li, Zhuochun, et al.
Veröffentlicht: (2026)
ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents
von: Li, Zhigen, et al.
Veröffentlicht: (2024)
von: Li, Zhigen, et al.
Veröffentlicht: (2024)
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
von: Huang, Yanwen, et al.
Veröffentlicht: (2025)
von: Huang, Yanwen, et al.
Veröffentlicht: (2025)
GRASP: Replace Redundant Layers with Adaptive Singular Parameters for Efficient Model Compression
von: Liu, Kainan, et al.
Veröffentlicht: (2024)
von: Liu, Kainan, et al.
Veröffentlicht: (2024)
Self-Enhanced Reasoning Training: Activating Latent Reasoning in Small Models for Enhanced Reasoning Distillation
von: Zhang, Yong, et al.
Veröffentlicht: (2025)
von: Zhang, Yong, et al.
Veröffentlicht: (2025)
PFID: Privacy First Inference Delegation Framework for LLMs
von: Yang, Haoyan, et al.
Veröffentlicht: (2024)
von: Yang, Haoyan, et al.
Veröffentlicht: (2024)
Selecting Demonstrations for Many-Shot In-Context Learning via Gradient Matching
von: Zhang, Jianfei, et al.
Veröffentlicht: (2025)
von: Zhang, Jianfei, et al.
Veröffentlicht: (2025)
Leveraging Biases in Large Language Models: "bias-kNN'' for Effective Few-Shot Learning
von: Zhang, Yong, et al.
Veröffentlicht: (2024)
von: Zhang, Yong, et al.
Veröffentlicht: (2024)
Token-level Direct Preference Optimization
von: Zeng, Yongcheng, et al.
Veröffentlicht: (2024)
von: Zeng, Yongcheng, et al.
Veröffentlicht: (2024)
StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning
von: Wang, Daoyu, et al.
Veröffentlicht: (2026)
von: Wang, Daoyu, et al.
Veröffentlicht: (2026)
GAPO: Learning Preferential Prompt through Generative Adversarial Policy Optimization
von: Gu, Zhouhong, et al.
Veröffentlicht: (2025)
von: Gu, Zhouhong, et al.
Veröffentlicht: (2025)
A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping
von: Chen, Dingwei, et al.
Veröffentlicht: (2026)
von: Chen, Dingwei, et al.
Veröffentlicht: (2026)
EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment
von: Zhang, Jianfei, et al.
Veröffentlicht: (2024)
von: Zhang, Jianfei, et al.
Veröffentlicht: (2024)
Trait-Aware Policy Optimization for Autoregressive Multi-Trait Essay Scoring
von: Wang, Zhengyang, et al.
Veröffentlicht: (2026)
von: Wang, Zhengyang, et al.
Veröffentlicht: (2026)
From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs
von: Ni, Jiliang, et al.
Veröffentlicht: (2025)
von: Ni, Jiliang, et al.
Veröffentlicht: (2025)
QLSC: A Query Latent Semantic Calibrator for Robust Extractive Question Answering
von: Ouyang, Sheng, et al.
Veröffentlicht: (2024)
von: Ouyang, Sheng, et al.
Veröffentlicht: (2024)
Enhancing Emotion Recognition in Conversation through Emotional Cross-Modal Fusion and Inter-class Contrastive Learning
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
DPO Meets PPO: Reinforced Token Optimization for RLHF
von: Zhong, Han, et al.
Veröffentlicht: (2024)
von: Zhong, Han, et al.
Veröffentlicht: (2024)
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
von: Xu, Wujiang, et al.
Veröffentlicht: (2025)
von: Xu, Wujiang, et al.
Veröffentlicht: (2025)
HMPO: Hybrid Median-length Policy Optimization for Chain-of-Thought Compression
von: Zheng, Minghui, et al.
Veröffentlicht: (2026)
von: Zheng, Minghui, et al.
Veröffentlicht: (2026)
From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
von: Li, Ming, et al.
Veröffentlicht: (2023)
von: Li, Ming, et al.
Veröffentlicht: (2023)
How to Train Your Deep Research Agent? Prompt, Reward, and Policy Optimization in Search-R1
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
GeometryZero: Advancing Geometry Solving via Group Contrastive Policy Optimization
von: Wang, Yikun, et al.
Veröffentlicht: (2025)
von: Wang, Yikun, et al.
Veröffentlicht: (2025)
Large Language Models in Bioinformatics: A Survey
von: Wang, Zhenyu, et al.
Veröffentlicht: (2025)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2025)
InspireDebate: Multi-Dimensional Subjective-Objective Evaluation-Guided Reasoning and Optimization for Debating
von: Wang, Fuyu, et al.
Veröffentlicht: (2025)
von: Wang, Fuyu, et al.
Veröffentlicht: (2025)
HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)
Perception-Aware Policy Optimization for Multimodal Reasoning
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning
von: Shan, Zikang, et al.
Veröffentlicht: (2026)
von: Shan, Zikang, et al.
Veröffentlicht: (2026)
Adaptive Group Policy Optimization: Towards Stable Training and Token-Efficient Reasoning
von: Li, Chen, et al.
Veröffentlicht: (2025)
von: Li, Chen, et al.
Veröffentlicht: (2025)
DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
von: Gu, Chenyang, et al.
Veröffentlicht: (2025)
von: Gu, Chenyang, et al.
Veröffentlicht: (2025)
DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
von: Yang, Ning, et al.
Veröffentlicht: (2025)
von: Yang, Ning, et al.
Veröffentlicht: (2025)
MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
Computational Sentence-level Metrics Predicting Human Sentence Comprehension
von: Sun, Kun, et al.
Veröffentlicht: (2024)
von: Sun, Kun, et al.
Veröffentlicht: (2024)
Draft-OPD: On-Policy Distillation for Speculative Draft Models
von: Lei, Haodi, et al.
Veröffentlicht: (2026)
von: Lei, Haodi, et al.
Veröffentlicht: (2026)
Agentic Entropy-Balanced Policy Optimization
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Learning to Adapt to Low-Resource Paraphrase Generation
von: Li, Zhigen, et al.
Veröffentlicht: (2024) -
Astra: Activation-Space Tail-Eigenvector Low-Rank Adaptation of Large Language Models
von: Liu, Kainan, et al.
Veröffentlicht: (2026) -
Verifiable Generation with Subsentence-Level Fine-Grained Citations
von: Cao, Shuyang, et al.
Veröffentlicht: (2024) -
Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression
von: Zhang, Yong, et al.
Veröffentlicht: (2025) -
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
von: Li, Zhuochun, et al.
Veröffentlicht: (2026)