Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Yangyi, Lin, Jiaye, Fu, Xiaoliang, Qin, Cong, Shi, Haolin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization
by: Fang, Yangyi, et al.
Published: (2026)
by: Fang, Yangyi, et al.
Published: (2026)
From $\log π$ to $π$: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight
by: Fu, Xiaoliang, et al.
Published: (2026)
by: Fu, Xiaoliang, et al.
Published: (2026)
MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning
by: Fu, Xiaoliang, et al.
Published: (2026)
by: Fu, Xiaoliang, et al.
Published: (2026)
Proximity-Based Multi-Turn Optimization: Practical Credit Assignment for LLM Agent Training
by: Fang, Yangyi, et al.
Published: (2026)
by: Fang, Yangyi, et al.
Published: (2026)
Quantifying Uncertainty in Natural Language Explanations of Large Language Models for Question Answering
by: Li, Yangyi, et al.
Published: (2025)
by: Li, Yangyi, et al.
Published: (2025)
Test Where Decisions Matter: Importance-driven Testing for Deep Reinforcement Learning
by: Pranger, Stefan, et al.
Published: (2024)
by: Pranger, Stefan, et al.
Published: (2024)
Where You Place the Norm Matters: From Prejudiced to Neutral Initializations
by: Francazi, Emanuele, et al.
Published: (2025)
by: Francazi, Emanuele, et al.
Published: (2025)
Layer as Puzzle Pieces: Compressing Large Language Models through Layer Concatenation
by: Wang, Fei, et al.
Published: (2025)
by: Wang, Fei, et al.
Published: (2025)
Disentangling Content from Style to Overcome Shortcut Learning: A Hybrid Generative-Discriminative Learning Framework
by: Fu, Siming, et al.
Published: (2025)
by: Fu, Siming, et al.
Published: (2025)
RABot: Reinforcement-Guided Graph Augmentation for Imbalanced and Noisy Social Bot Detection
by: Zhang, Longlong, et al.
Published: (2026)
by: Zhang, Longlong, et al.
Published: (2026)
Space Alignment Matters: The Missing Piece for Inducing Neural Collapse in Long-Tailed Learning
by: Wang, Jinping, et al.
Published: (2025)
by: Wang, Jinping, et al.
Published: (2025)
Privacy Preserving Reinforcement Learning with One-Sided Feedback
by: Cong, Lin William, et al.
Published: (2026)
by: Cong, Lin William, et al.
Published: (2026)
Sliding Puzzles Gym: A Scalable Benchmark for State Representation in Visual Reinforcement Learning
by: de Oliveira, Bryan L. M., et al.
Published: (2024)
by: de Oliveira, Bryan L. M., et al.
Published: (2024)
Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment
by: Cho, Youngjae, et al.
Published: (2026)
by: Cho, Youngjae, et al.
Published: (2026)
PASTA: A Unified Framework for Offline Assortment Learning
by: Dong, Juncheng, et al.
Published: (2025)
by: Dong, Juncheng, et al.
Published: (2025)
Where to Intervene: Action Selection in Deep Reinforcement Learning
by: Zhang, Wenbo, et al.
Published: (2025)
by: Zhang, Wenbo, et al.
Published: (2025)
Questioning the Coverage-Length Metric in Conformal Prediction: When Shorter Intervals Are Not Better
by: Min, Yizhou, et al.
Published: (2026)
by: Min, Yizhou, et al.
Published: (2026)
Smooth Dynamic Cutoffs for Machine Learning Interatomic Potentials
by: Han, Kevin, et al.
Published: (2026)
by: Han, Kevin, et al.
Published: (2026)
LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning
by: Wong, Zhen Hao, et al.
Published: (2025)
by: Wong, Zhen Hao, et al.
Published: (2025)
Towards Piece-by-Piece Explanations for Chess Positions with SHAP
by: Spinnato, Francesco
Published: (2025)
by: Spinnato, Francesco
Published: (2025)
Uncertainty-aware Language Guidance for Concept Bottleneck Models
by: Li, Yangyi, et al.
Published: (2026)
by: Li, Yangyi, et al.
Published: (2026)
A Probabilistic Framework for Temporal Distribution Generalization in Industry-Scale Recommender Systems
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage
by: Liu, Haolin, et al.
Published: (2026)
by: Liu, Haolin, et al.
Published: (2026)
PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction
by: Ayyubi, Hammad, et al.
Published: (2025)
by: Ayyubi, Hammad, et al.
Published: (2025)
Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge
by: Liu, Genglin, et al.
Published: (2023)
by: Liu, Genglin, et al.
Published: (2023)
Embedded Course Reserves: Piecing the Puzzle Together
by: Clumpner, Krista E., et al.
Published: (2011)
by: Clumpner, Krista E., et al.
Published: (2011)
Reasoning in Trees: Improving Retrieval-Augmented Generation for Multi-Hop Question Answering
by: Shi, Yuling, et al.
Published: (2026)
by: Shi, Yuling, et al.
Published: (2026)
SPHERE: Mitigating the Loss of Spectral Plasticity in Mixture-of-Experts for Deep Reinforcement Learning
by: Luo, Lirui, et al.
Published: (2026)
by: Luo, Lirui, et al.
Published: (2026)
A Curriculum Learning Approach to Reinforcement Learning: Leveraging RAG for Multimodal Question Answering
by: Zhang, Chenliang, et al.
Published: (2025)
by: Zhang, Chenliang, et al.
Published: (2025)
Search over Self-Edit Strategies for LLM Adaptation
by: Cheong, Alistair, et al.
Published: (2026)
by: Cheong, Alistair, et al.
Published: (2026)
Partial Feedback Online Learning
by: Shao, Shihao, et al.
Published: (2026)
by: Shao, Shihao, et al.
Published: (2026)
Causal Question Answering with Reinforcement Learning
by: Blübaum, Lukas, et al.
Published: (2023)
by: Blübaum, Lukas, et al.
Published: (2023)
PuzzleJAX: A Benchmark for Reasoning and Learning
by: Earle, Sam, et al.
Published: (2025)
by: Earle, Sam, et al.
Published: (2025)
Hypercube-Based Retrieval-Augmented Generation for Scientific Question-Answering
by: Shi, Jimeng, et al.
Published: (2025)
by: Shi, Jimeng, et al.
Published: (2025)
NSW-EPNews: A News-Augmented Benchmark for Electricity Price Forecasting with LLMs
by: Bi, Zhaoge, et al.
Published: (2025)
by: Bi, Zhaoge, et al.
Published: (2025)
Learning Curves of Stochastic Gradient Descent in Kernel Regression
by: Zhang, Haihan, et al.
Published: (2025)
by: Zhang, Haihan, et al.
Published: (2025)
Model-Based Offline Reinforcement Learning with Adversarial Data Augmentation
by: Cao, Hongye, et al.
Published: (2025)
by: Cao, Hongye, et al.
Published: (2025)
FedGTEA: Federated Class-Incremental Learning with Gaussian Task Embedding and Alignment
by: Li, Haolin, et al.
Published: (2025)
by: Li, Haolin, et al.
Published: (2025)
Piece of CAKE: Adaptive Execution Engines via Microsecond-Scale Learning
by: Zhao, Zijie, et al.
Published: (2026)
by: Zhao, Zijie, et al.
Published: (2026)
Focus Where It Matters: Graph Selective State Focused Attention Networks
by: Vashistha, Shikhar, et al.
Published: (2024)
by: Vashistha, Shikhar, et al.
Published: (2024)
Similar Items
-
How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization
by: Fang, Yangyi, et al.
Published: (2026) -
From $\log π$ to $π$: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight
by: Fu, Xiaoliang, et al.
Published: (2026) -
MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning
by: Fu, Xiaoliang, et al.
Published: (2026) -
Proximity-Based Multi-Turn Optimization: Practical Credit Assignment for LLM Agent Training
by: Fang, Yangyi, et al.
Published: (2026) -
Quantifying Uncertainty in Natural Language Explanations of Large Language Models for Question Answering
by: Li, Yangyi, et al.
Published: (2025)