From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tan, Yuqiao, Wang, Minzheng, Liu, Bo, Liu, Zichen, Liang, Tian, He, Shizhu, Zhao, Jun, Liu, Kang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
von: Tan, Yuqiao, et al.
Veröffentlicht: (2025)
von: Tan, Yuqiao, et al.
Veröffentlicht: (2025)
Neural Incompatibility: The Unbridgeable Gap of Cross-Scale Parametric Knowledge Transfer in Large Language Models
von: Tan, Yuqiao, et al.
Veröffentlicht: (2025)
von: Tan, Yuqiao, et al.
Veröffentlicht: (2025)
Dynamic Parametric Retrieval Augmented Generation for Test-time Knowledge Enhancement
von: Tan, Yuqiao, et al.
Veröffentlicht: (2025)
von: Tan, Yuqiao, et al.
Veröffentlicht: (2025)
Find Parent then Label Children: A Two-stage Taxonomy Completion Method with Pre-trained Language Model
von: Xia, Fei, et al.
Veröffentlicht: (2024)
von: Xia, Fei, et al.
Veröffentlicht: (2024)
Efficient Data Learning for Open Information Extraction with Pre-trained Language Models
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2023)
DATA: Decomposed Attention-based Task Adaptation for Rehearsal-Free Continual Learning
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
ExpNote: Black-box Large Language Models are Better Task Solvers with Experience Notebook
von: Sun, Wangtao, et al.
Veröffentlicht: (2023)
von: Sun, Wangtao, et al.
Veröffentlicht: (2023)
ResAdapt: Adaptive Resolution for Efficient Multimodal Reasoning
von: Liao, Huanxuan, et al.
Veröffentlicht: (2026)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2026)
ItD: Large Language Models Can Teach Themselves Induction through Deduction
von: Sun, Wangtao, et al.
Veröffentlicht: (2024)
von: Sun, Wangtao, et al.
Veröffentlicht: (2024)
DocMamba: Efficient Document Pre-training with State Space Model
von: Hu, Pengfei, et al.
Veröffentlicht: (2024)
von: Hu, Pengfei, et al.
Veröffentlicht: (2024)
Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents
von: Wang, Minzheng, et al.
Veröffentlicht: (2026)
von: Wang, Minzheng, et al.
Veröffentlicht: (2026)
From Chain to Tree: Refining Chain-like Rules into Tree-like Rules on Knowledge Graphs
von: Sun, Wangtao, et al.
Veröffentlicht: (2024)
von: Sun, Wangtao, et al.
Veröffentlicht: (2024)
Probing Language Models for Pre-training Data Detection
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
LLaSA: Large Language and Structured Data Assistant
von: Xu, Yao, et al.
Veröffentlicht: (2024)
von: Xu, Yao, et al.
Veröffentlicht: (2024)
Generate-on-Graph: Treat LLM as both Agent and KG in Incomplete Knowledge Graph Question Answering
von: Xu, Yao, et al.
Veröffentlicht: (2024)
von: Xu, Yao, et al.
Veröffentlicht: (2024)
LingYi: Medical Conversational Question Answering System based on Multi-modal Knowledge Graphs
von: Xia, Fei, et al.
Veröffentlicht: (2022)
von: Xia, Fei, et al.
Veröffentlicht: (2022)
ControlLM: Crafting Diverse Personalities for Language Models
von: Weng, Yixuan, et al.
Veröffentlicht: (2024)
von: Weng, Yixuan, et al.
Veröffentlicht: (2024)
Reasoning-Table: Exploring Reinforcement Learning for Table Reasoning
von: Lei, Fangyu, et al.
Veröffentlicht: (2025)
von: Lei, Fangyu, et al.
Veröffentlicht: (2025)
Investigating Data Contamination for Pre-training Language Models
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN
von: Xu, Yao, et al.
Veröffentlicht: (2025)
von: Xu, Yao, et al.
Veröffentlicht: (2025)
From Instance Training to Instruction Learning: Task Adapters Generation from Instructions
von: Liao, Huanxuan, et al.
Veröffentlicht: (2024)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2024)
MaP: A Unified Framework for Reliable Evaluation of Pre-training Dynamics
von: Wang, Jiapeng, et al.
Veröffentlicht: (2025)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2025)
Socratic-PRMBench: Benchmarking Process Reward Models with Systematic Reasoning Patterns
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
Compensating Distribution Drifts in Class-incremental Learning of Pre-trained Vision Transformers
von: Rao, Xuan, et al.
Veröffentlicht: (2025)
von: Rao, Xuan, et al.
Veröffentlicht: (2025)
SC-Taxo: Hierarchical Taxonomy Generation under Semantic Consistency Constraints using Large Language Models
von: Cai, Shiqiang, et al.
Veröffentlicht: (2026)
von: Cai, Shiqiang, et al.
Veröffentlicht: (2026)
DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models
von: Huang, Yiming, et al.
Veröffentlicht: (2024)
von: Huang, Yiming, et al.
Veröffentlicht: (2024)
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
von: Yang, Kailai, et al.
Veröffentlicht: (2025)
von: Yang, Kailai, et al.
Veröffentlicht: (2025)
Parallel Structures in Pre-training Data Yield In-Context Learning
von: Chen, Yanda, et al.
Veröffentlicht: (2024)
von: Chen, Yanda, et al.
Veröffentlicht: (2024)
Can Pre-trained Language Models Understand Chinese Humor?
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
Awakening Augmented Generation: Learning to Awaken Internal Knowledge of Large Language Models for Question Answering
von: Liao, Huanxuan, et al.
Veröffentlicht: (2024)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2024)
S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models
von: Lei, Fangyu, et al.
Veröffentlicht: (2023)
von: Lei, Fangyu, et al.
Veröffentlicht: (2023)
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
WideSeek: Advancing Wide Research via Multi-Agent Scaling
von: Huang, Ziyang, et al.
Veröffentlicht: (2026)
von: Huang, Ziyang, et al.
Veröffentlicht: (2026)
MoELoRA: Contrastive Learning Guided Mixture of Experts on Parameter-Efficient Fine-Tuning for Large Language Models
von: Luo, Tongxu, et al.
Veröffentlicht: (2024)
von: Luo, Tongxu, et al.
Veröffentlicht: (2024)
SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
von: Zeng, Weihao, et al.
Veröffentlicht: (2025)
von: Zeng, Weihao, et al.
Veröffentlicht: (2025)
Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models
von: Sun, Wangtao, et al.
Veröffentlicht: (2024)
von: Sun, Wangtao, et al.
Veröffentlicht: (2024)
P-Aligner: Enabling Pre-Alignment of Language Models via Principled Instruction Synthesis
von: Song, Feifan, et al.
Veröffentlicht: (2025)
von: Song, Feifan, et al.
Veröffentlicht: (2025)
How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language Models
von: Lv, Kangtao, et al.
Veröffentlicht: (2025)
von: Lv, Kangtao, et al.
Veröffentlicht: (2025)
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
von: Liu, Bo, et al.
Veröffentlicht: (2025)
von: Liu, Bo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
von: Tan, Yuqiao, et al.
Veröffentlicht: (2025) -
Neural Incompatibility: The Unbridgeable Gap of Cross-Scale Parametric Knowledge Transfer in Large Language Models
von: Tan, Yuqiao, et al.
Veröffentlicht: (2025) -
Dynamic Parametric Retrieval Augmented Generation for Test-time Knowledge Enhancement
von: Tan, Yuqiao, et al.
Veröffentlicht: (2025) -
Find Parent then Label Children: A Two-stage Taxonomy Completion Method with Pre-trained Language Model
von: Xia, Fei, et al.
Veröffentlicht: (2024) -
Efficient Data Learning for Open Information Extraction with Pre-trained Language Models
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2023)