Gespeichert in:
| Hauptverfasser: | Wu, Haoze, Wang, Cheng, Zhao, Wenshuo, He, Junxian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2508.21188 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Internalizing World Models via Self-Play Finetuning for Agentic RL
von: Chen, Shiqi, et al.
Veröffentlicht: (2025)
von: Chen, Shiqi, et al.
Veröffentlicht: (2025)
SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
von: Zeng, Weihao, et al.
Veröffentlicht: (2025)
von: Zeng, Weihao, et al.
Veröffentlicht: (2025)
Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?
von: Chi, Haoang, et al.
Veröffentlicht: (2025)
von: Chi, Haoang, et al.
Veröffentlicht: (2025)
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
von: Liu, Wei, et al.
Veröffentlicht: (2023)
von: Liu, Wei, et al.
Veröffentlicht: (2023)
ToolSample: Dual Dynamic Sampling Methods with Curriculum Learning for RL-based Tool Learning
von: Feng, Zihao, et al.
Veröffentlicht: (2025)
von: Feng, Zihao, et al.
Veröffentlicht: (2025)
Objective Metrics for Evaluating Large Language Models Using External Data Sources
von: Du, Haoze, et al.
Veröffentlicht: (2025)
von: Du, Haoze, et al.
Veröffentlicht: (2025)
Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
von: Zhao, Chengshuai, et al.
Veröffentlicht: (2025)
von: Zhao, Chengshuai, et al.
Veröffentlicht: (2025)
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
von: Kulkarni, Atharva, et al.
Veröffentlicht: (2025)
von: Kulkarni, Atharva, et al.
Veröffentlicht: (2025)
TACO-RL: Task Aware Prompt Compression Optimization with Reinforcement Learning
von: Shandilya, Shivam, et al.
Veröffentlicht: (2024)
von: Shandilya, Shivam, et al.
Veröffentlicht: (2024)
Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning
von: He, Haoze, et al.
Veröffentlicht: (2026)
von: He, Haoze, et al.
Veröffentlicht: (2026)
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
von: Cheng, Xinhao, et al.
Veröffentlicht: (2025)
von: Cheng, Xinhao, et al.
Veröffentlicht: (2025)
RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
DBR: Divergence-Based Regularization for Debiasing Natural Language Understanding Models
von: Li, Zihao, et al.
Veröffentlicht: (2025)
von: Li, Zihao, et al.
Veröffentlicht: (2025)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
von: Zeng, Weihao, et al.
Veröffentlicht: (2024)
von: Zeng, Weihao, et al.
Veröffentlicht: (2024)
Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning
von: Xu, Fangzhi, et al.
Veröffentlicht: (2025)
von: Xu, Fangzhi, et al.
Veröffentlicht: (2025)
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
von: Huang, Yuzhen, et al.
Veröffentlicht: (2025)
von: Huang, Yuzhen, et al.
Veröffentlicht: (2025)
Better and Worse with Scale: How Contextual Entrainment Diverges with Model Size
von: Kukreja, Dikshant, et al.
Veröffentlicht: (2026)
von: Kukreja, Dikshant, et al.
Veröffentlicht: (2026)
HAL: Inducing Human-likeness in LLMs with Alignment
von: Hasan, Masum, et al.
Veröffentlicht: (2026)
von: Hasan, Masum, et al.
Veröffentlicht: (2026)
Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language Models
von: Sun, Shaoning, et al.
Veröffentlicht: (2026)
von: Sun, Shaoning, et al.
Veröffentlicht: (2026)
Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL
von: Patel, Nyal, et al.
Veröffentlicht: (2025)
von: Patel, Nyal, et al.
Veröffentlicht: (2025)
Synthetic Data RL: Task Definition Is All You Need
von: Guo, Yiduo, et al.
Veröffentlicht: (2025)
von: Guo, Yiduo, et al.
Veröffentlicht: (2025)
From Lists to Emojis: How Format Bias Affects Model Alignment
von: Zhang, Xuanchang, et al.
Veröffentlicht: (2024)
von: Zhang, Xuanchang, et al.
Veröffentlicht: (2024)
CodeMirage: A Multi-Lingual Benchmark for Detecting AI-Generated and Paraphrased Source Code from Production-Level LLMs
von: Guo, Hanxi, et al.
Veröffentlicht: (2025)
von: Guo, Hanxi, et al.
Veröffentlicht: (2025)
Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
Lemur: Integrating Large Language Models in Automated Program Verification
von: Wu, Haoze, et al.
Veröffentlicht: (2023)
von: Wu, Haoze, et al.
Veröffentlicht: (2023)
Capability-Oriented Training Induced Alignment Risk
von: Zhou, Yujun, et al.
Veröffentlicht: (2026)
von: Zhou, Yujun, et al.
Veröffentlicht: (2026)
From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment
von: Chen, Hao, et al.
Veröffentlicht: (2026)
von: Chen, Hao, et al.
Veröffentlicht: (2026)
Mirage: A Multi-Level Superoptimizer for Tensor Programs
von: Wu, Mengdi, et al.
Veröffentlicht: (2024)
von: Wu, Mengdi, et al.
Veröffentlicht: (2024)
Weight-Inherited Distillation for Task-Agnostic BERT Compression
von: Wu, Taiqiang, et al.
Veröffentlicht: (2023)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2023)
Baichuan Alignment Technical Report
von: Lin, Mingan, et al.
Veröffentlicht: (2024)
von: Lin, Mingan, et al.
Veröffentlicht: (2024)
How Does Code Pretraining Affect Language Model Task Performance?
von: Petty, Jackson, et al.
Veröffentlicht: (2024)
von: Petty, Jackson, et al.
Veröffentlicht: (2024)
ToolRL: Reward is All Tool Learning Needs
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
Optimizing Mixture of Block Attention
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2025)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2025)
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
von: Liu, Runze, et al.
Veröffentlicht: (2025)
von: Liu, Runze, et al.
Veröffentlicht: (2025)
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence
von: Choi, Hyeong Kyu, et al.
Veröffentlicht: (2025)
von: Choi, Hyeong Kyu, et al.
Veröffentlicht: (2025)
On-Policy RL with Optimal Reward Baseline
von: Hao, Yaru, et al.
Veröffentlicht: (2025)
von: Hao, Yaru, et al.
Veröffentlicht: (2025)
Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?
von: He, Yufei, et al.
Veröffentlicht: (2025)
von: He, Yufei, et al.
Veröffentlicht: (2025)
Prompt Optimization via Adversarial In-Context Learning
von: Do, Xuan Long, et al.
Veröffentlicht: (2023)
von: Do, Xuan Long, et al.
Veröffentlicht: (2023)
Mortgage Language Model: Domain-Adaptive Pretraining with Residual Instruction, Alignment Tuning, and Task-Specific Routing
von: Jain, Manish, et al.
Veröffentlicht: (2025)
von: Jain, Manish, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Internalizing World Models via Self-Play Finetuning for Agentic RL
von: Chen, Shiqi, et al.
Veröffentlicht: (2025) -
SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
von: Zeng, Weihao, et al.
Veröffentlicht: (2025) -
Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?
von: Chi, Haoang, et al.
Veröffentlicht: (2025) -
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
von: Liu, Wei, et al.
Veröffentlicht: (2023) -
ToolSample: Dual Dynamic Sampling Methods with Curriculum Learning for RL-based Tool Learning
von: Feng, Zihao, et al.
Veröffentlicht: (2025)