MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Fu, Xiaoliang, Lin, Jiaye, Fang, Yangyi, Zheng, Binbin, Hu, Chaowen, Shao, Zekai, Qin, Cong, Pan, Lu, Zeng, Ke, Cai, Xunliang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From $\log π$ to $π$: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight
di: Fu, Xiaoliang, et al.
Pubblicazione: (2026)
di: Fu, Xiaoliang, et al.
Pubblicazione: (2026)
How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization
di: Fang, Yangyi, et al.
Pubblicazione: (2026)
di: Fang, Yangyi, et al.
Pubblicazione: (2026)
Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning
di: Fang, Yangyi, et al.
Pubblicazione: (2026)
di: Fang, Yangyi, et al.
Pubblicazione: (2026)
Proximity-Based Multi-Turn Optimization: Practical Credit Assignment for LLM Agent Training
di: Fang, Yangyi, et al.
Pubblicazione: (2026)
di: Fang, Yangyi, et al.
Pubblicazione: (2026)
SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
di: Zheng, Binbin, et al.
Pubblicazione: (2026)
di: Zheng, Binbin, et al.
Pubblicazione: (2026)
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy
di: Zhang, Xiaoyun, et al.
Pubblicazione: (2025)
di: Zhang, Xiaoyun, et al.
Pubblicazione: (2025)
MASPO: Joint Prompt Optimization for LLM-based Multi-Agent Systems
di: Wang, Zhexuan, et al.
Pubblicazione: (2026)
di: Wang, Zhexuan, et al.
Pubblicazione: (2026)
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR
di: Miao, Yuchun, et al.
Pubblicazione: (2026)
di: Miao, Yuchun, et al.
Pubblicazione: (2026)
Representations of the restricted Lie superalgebra p(n)
di: Zhang, Chaowen
Pubblicazione: (2023)
di: Zhang, Chaowen
Pubblicazione: (2023)
An Alternative Channel to Black Hole Low-Mass X-ray Binaries: Dynamical Friction of Dark Matter?
di: Qin, Ke, et al.
Pubblicazione: (2024)
di: Qin, Ke, et al.
Pubblicazione: (2024)
Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning
di: Chen, Chishui, et al.
Pubblicazione: (2026)
di: Chen, Chishui, et al.
Pubblicazione: (2026)
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
di: Jian, Ai, et al.
Pubblicazione: (2025)
di: Jian, Ai, et al.
Pubblicazione: (2025)
Harmonizing Dense and Sparse Signals in Multi-turn RL: Dual-Horizon Credit Assignment for Industrial Sales Agents
di: Yang, Haojin, et al.
Pubblicazione: (2026)
di: Yang, Haojin, et al.
Pubblicazione: (2026)
Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning
di: Xing, Zeyu, et al.
Pubblicazione: (2026)
di: Xing, Zeyu, et al.
Pubblicazione: (2026)
AutothinkRAG: Complexity-Aware Control of Retrieval-Augmented Reasoning for Image-Text Interaction
di: Yang, Jiashu, et al.
Pubblicazione: (2026)
di: Yang, Jiashu, et al.
Pubblicazione: (2026)
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
di: Yang, Xin, et al.
Pubblicazione: (2026)
di: Yang, Xin, et al.
Pubblicazione: (2026)
Sampling via Gradient Flows in the Space of Probability Measures
di: Chen, Yifan, et al.
Pubblicazione: (2023)
di: Chen, Yifan, et al.
Pubblicazione: (2023)
MFSD2A Overexpression Inhibits Hepatocellular Carcinoma Through TGF‐β/Smad Signaling
di: Chaowen Xiao, et al.
Pubblicazione: (2025)
di: Chaowen Xiao, et al.
Pubblicazione: (2025)
When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning
di: Zhang, Xiaoyun, et al.
Pubblicazione: (2025)
di: Zhang, Xiaoyun, et al.
Pubblicazione: (2025)
Adversarial Adaptive Sampling: Unify PINN and Optimal Transport for the Approximation of PDEs
di: Tang, Kejun, et al.
Pubblicazione: (2023)
di: Tang, Kejun, et al.
Pubblicazione: (2023)
Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
di: Liu, Jiahao, et al.
Pubblicazione: (2024)
di: Liu, Jiahao, et al.
Pubblicazione: (2024)
Generalized Policy Gradient with History-Aware Decision Transformer for Reliable Routing over Graph Signals
di: Wei, Xing, et al.
Pubblicazione: (2025)
di: Wei, Xing, et al.
Pubblicazione: (2025)
Quantifying and Understanding Uncertainty in Large Reasoning Models
di: Li, Yangyi, et al.
Pubblicazione: (2026)
di: Li, Yangyi, et al.
Pubblicazione: (2026)
DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding
di: Zhou, Jiaming, et al.
Pubblicazione: (2026)
di: Zhou, Jiaming, et al.
Pubblicazione: (2026)
Structured Gradient Descent for Fast Robust Low-Rank Hankel Matrix Completion
di: Cai, HanQin, et al.
Pubblicazione: (2022)
di: Cai, HanQin, et al.
Pubblicazione: (2022)
Large Language Models Can Help Mitigate Barren Plateaus in Quantum Neural Networks
di: Zhuang, Jun, et al.
Pubblicazione: (2025)
di: Zhuang, Jun, et al.
Pubblicazione: (2025)
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
di: Yang, Yunqiao, et al.
Pubblicazione: (2025)
di: Yang, Yunqiao, et al.
Pubblicazione: (2025)
Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective
di: Kong, Deyang, et al.
Pubblicazione: (2025)
di: Kong, Deyang, et al.
Pubblicazione: (2025)
Structured Sampling for Robust Euclidean Distance Geometry
di: Kundu, Chandra, et al.
Pubblicazione: (2024)
di: Kundu, Chandra, et al.
Pubblicazione: (2024)
In‐Plane Crashworthiness of Gradient Curved‐Walled Honeycombs: Experimental, Numerical Simulation, and Theoretical Analysis
di: Kuijian Yang, et al.
Pubblicazione: (2024)
di: Kuijian Yang, et al.
Pubblicazione: (2024)
Utilization of Fruit Juice Processing Wastes as Prebiotic Ingredients in Probiotic Yogurt: Effects on Microbial Short Chain Fatty Acid Production
di: Melike Demirkol, et al.
Pubblicazione: (2025)
di: Melike Demirkol, et al.
Pubblicazione: (2025)
From Likelihood to Limit State: A Reliability-Inspired Framework for Bayesian Evidence Estimation and High-dimensional Sampling
di: Liao, Zihan, et al.
Pubblicazione: (2024)
di: Liao, Zihan, et al.
Pubblicazione: (2024)
High-Probability Convergence in Decentralized Stochastic Optimization with Gradient Tracking
di: Armacki, Aleksandar, et al.
Pubblicazione: (2026)
di: Armacki, Aleksandar, et al.
Pubblicazione: (2026)
Robust Spectral Recovery for Dynamical Sampling
di: Cai, HanQin, et al.
Pubblicazione: (2026)
di: Cai, HanQin, et al.
Pubblicazione: (2026)
Beyond Templates: Dynamic Adaptation of Reasoning Demonstrations via Feasibility-Aware Exploration
di: Wu, Yong, et al.
Pubblicazione: (2025)
di: Wu, Yong, et al.
Pubblicazione: (2025)
Gradient-free Importance Sampling Scheme for Efficient Reliability Estimation
di: Eshra, Elsayed, et al.
Pubblicazione: (2025)
di: Eshra, Elsayed, et al.
Pubblicazione: (2025)
Reflecting Twice before Speaking with Empathy: Self-Reflective Alternating Inference for Empathy-Aware End-to-End Spoken Dialogue
di: Jia, Yuhang, et al.
Pubblicazione: (2026)
di: Jia, Yuhang, et al.
Pubblicazione: (2026)
On the Robustness of Cross-Concentrated Sampling for Matrix Completion
di: Cai, HanQin, et al.
Pubblicazione: (2024)
di: Cai, HanQin, et al.
Pubblicazione: (2024)
Prejudge-Before-Think: Enhancing Large Language Models at Test-Time by Process Prejudge Reasoning
di: Wang, Jianing, et al.
Pubblicazione: (2025)
di: Wang, Jianing, et al.
Pubblicazione: (2025)
CRAVE: A Conflicting Reasoning Approach for Explainable Claim Verification Using LLMs
di: Zheng, Yingming, et al.
Pubblicazione: (2025)
di: Zheng, Yingming, et al.
Pubblicazione: (2025)
Documenti analoghi
-
From $\log π$ to $π$: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight
di: Fu, Xiaoliang, et al.
Pubblicazione: (2026) -
How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization
di: Fang, Yangyi, et al.
Pubblicazione: (2026) -
Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning
di: Fang, Yangyi, et al.
Pubblicazione: (2026) -
Proximity-Based Multi-Turn Optimization: Practical Credit Assignment for LLM Agent Training
di: Fang, Yangyi, et al.
Pubblicazione: (2026) -
SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
di: Zheng, Binbin, et al.
Pubblicazione: (2026)