Gespeichert in:
| Hauptverfasser: | Wong, Zhen Hao, Deng, Jingwen, He, Runming, Chen, Zirong, You, Qijie, Dong, Hejun, Liang, Hao, Shen, Chengyu, Cui, Bin, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2506.04821 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
von: Wong, Zhen Hao, et al.
Veröffentlicht: (2025)
von: Wong, Zhen Hao, et al.
Veröffentlicht: (2025)
EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
von: Ma, Lu, et al.
Veröffentlicht: (2025)
von: Ma, Lu, et al.
Veröffentlicht: (2025)
Let's Verify Math Questions Step by Step
von: Shen, Chengyu, et al.
Veröffentlicht: (2025)
von: Shen, Chengyu, et al.
Veröffentlicht: (2025)
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG
von: You, Qijie, et al.
Veröffentlicht: (2026)
von: You, Qijie, et al.
Veröffentlicht: (2026)
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
von: Shen, Chengyu, et al.
Veröffentlicht: (2026)
von: Shen, Chengyu, et al.
Veröffentlicht: (2026)
Quantization Meets Reasoning: Exploring and Mitigating Degradation of Low-Bit LLMs in Mathematical Reasoning
von: Li, Zhen, et al.
Veröffentlicht: (2025)
von: Li, Zhen, et al.
Veröffentlicht: (2025)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
von: Li, Zhen, et al.
Veröffentlicht: (2025)
von: Li, Zhen, et al.
Veröffentlicht: (2025)
DARO: Difficulty-Aware Reweighting Policy Optimization
von: Zhou, Jingyu, et al.
Veröffentlicht: (2025)
von: Zhou, Jingyu, et al.
Veröffentlicht: (2025)
VERIFY-RL: Verifiable Recursive Decomposition for Reinforcement Learning in Mathematical Reasoning
von: Qasim, Kaleem Ullah, et al.
Veröffentlicht: (2026)
von: Qasim, Kaleem Ullah, et al.
Veröffentlicht: (2026)
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries
von: You, Qijie, et al.
Veröffentlicht: (2026)
von: You, Qijie, et al.
Veröffentlicht: (2026)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
von: Hao, Yuren, et al.
Veröffentlicht: (2025)
von: Hao, Yuren, et al.
Veröffentlicht: (2025)
Asynchronous Credit Assignment for Multi-Agent Reinforcement Learning
von: Liang, Yongheng, et al.
Veröffentlicht: (2024)
von: Liang, Yongheng, et al.
Veröffentlicht: (2024)
MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
von: Zheng, Lihao, et al.
Veröffentlicht: (2025)
von: Zheng, Lihao, et al.
Veröffentlicht: (2025)
Canonicity for Cost-Aware Logical Framework via Synthetic Tait Computability
von: Li, Runming, et al.
Veröffentlicht: (2025)
von: Li, Runming, et al.
Veröffentlicht: (2025)
LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
von: Cai, Qifeng, et al.
Veröffentlicht: (2025)
von: Cai, Qifeng, et al.
Veröffentlicht: (2025)
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
von: Feng, Yichen, et al.
Veröffentlicht: (2025)
von: Feng, Yichen, et al.
Veröffentlicht: (2025)
Genetics of Dementia with Lewy Bodies in a Chinese Population
von: Lu Shen, et al.
Veröffentlicht: (2024)
von: Lu Shen, et al.
Veröffentlicht: (2024)
Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles
von: Chen, Jiangjie, et al.
Veröffentlicht: (2025)
von: Chen, Jiangjie, et al.
Veröffentlicht: (2025)
Revisiting Model Interpolation for Efficient Reasoning
von: Wu, Taiqiang, et al.
Veröffentlicht: (2025)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2025)
DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
von: Liang, Hao, et al.
Veröffentlicht: (2025)
von: Liang, Hao, et al.
Veröffentlicht: (2025)
Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning
von: Chen, Mingrui, et al.
Veröffentlicht: (2025)
von: Chen, Mingrui, et al.
Veröffentlicht: (2025)
SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
Beyond Memorization: Distinguishing between Reductive and Epistemic Reasoning in LLMs using Classic Logic Puzzles
von: Gabay, Adi, et al.
Veröffentlicht: (2026)
von: Gabay, Adi, et al.
Veröffentlicht: (2026)
MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation
von: Wang, Hongpeng, et al.
Veröffentlicht: (2026)
von: Wang, Hongpeng, et al.
Veröffentlicht: (2026)
PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning
von: Lu, Keer, et al.
Veröffentlicht: (2025)
von: Lu, Keer, et al.
Veröffentlicht: (2025)
Exploiting Hybrid Policy in Reinforcement Learning for Interpretable Temporal Logic Manipulation
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality
von: Wu, Taiqiang, et al.
Veröffentlicht: (2026)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2026)
Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
von: Xie, Tian, et al.
Veröffentlicht: (2025)
von: Xie, Tian, et al.
Veröffentlicht: (2025)
FINEREASON: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving
von: Chen, Guizhen, et al.
Veröffentlicht: (2025)
von: Chen, Guizhen, et al.
Veröffentlicht: (2025)
RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning
von: Hao, Qianyue, et al.
Veröffentlicht: (2025)
von: Hao, Qianyue, et al.
Veröffentlicht: (2025)
ThinkRL-Edit: Thinking in Reinforcement Learning for Reasoning-Centric Image Editing
von: Li, Hengjia, et al.
Veröffentlicht: (2026)
von: Li, Hengjia, et al.
Veröffentlicht: (2026)
Cultivating Helpful, Personalized, and Creative AI Tutors: A Framework for Pedagogical Alignment using Reinforcement Learning
von: Song, Siyu, et al.
Veröffentlicht: (2025)
von: Song, Siyu, et al.
Veröffentlicht: (2025)
MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
von: Liang, Hao, et al.
Veröffentlicht: (2025)
von: Liang, Hao, et al.
Veröffentlicht: (2025)
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
von: Liang, Hao, et al.
Veröffentlicht: (2026)
von: Liang, Hao, et al.
Veröffentlicht: (2026)
HardcoreLogic: Challenging Large Reasoning Models with Long-tail Logic Puzzle Games
von: Liang, Jingcong, et al.
Veröffentlicht: (2025)
von: Liang, Jingcong, et al.
Veröffentlicht: (2025)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
CaMeRL: Collision-Aware and Memory-Enhanced Reinforcement Learning for UAV Navigation in Multi-Scale Obstacle Environments
von: Hong, Hong, et al.
Veröffentlicht: (2026)
von: Hong, Hong, et al.
Veröffentlicht: (2026)
HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoning LLMs
von: Ning, Yansong, et al.
Veröffentlicht: (2026)
von: Ning, Yansong, et al.
Veröffentlicht: (2026)
JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
von: Wong, Zhen Hao, et al.
Veröffentlicht: (2025) -
EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation
von: Liang, Hao, et al.
Veröffentlicht: (2024) -
Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
von: Ma, Lu, et al.
Veröffentlicht: (2025) -
Let's Verify Math Questions Step by Step
von: Shen, Chengyu, et al.
Veröffentlicht: (2025) -
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG
von: You, Qijie, et al.
Veröffentlicht: (2026)