PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yizhen, Ding, Yang, Zhang, Shuoshuo, Zhang, Xinchen, Li, Haoling, Li, Zhong-zhi, Wang, Peijie, Wu, Jie, Ji, Lei, Shen, Yelong, Yang, Yujiu, Gong, Yeyun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
by: Ding, Yang, et al.
Published: (2025)
by: Ding, Yang, et al.
Published: (2025)
See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
by: Zhang, Shuoshuo, et al.
Published: (2025)
by: Zhang, Shuoshuo, et al.
Published: (2025)
PeRL: Permafrost Region Pond and Lake Database, links to ArcGIS shapefiles
by: Muster, Sina, et al.
Published: (2017)
by: Muster, Sina, et al.
Published: (2017)
Permafrost-Region Lake-DOC version1 Database (PeRL-DOCv1)
by: Stolpmann, Lydia, et al.
Published: (2021)
by: Stolpmann, Lydia, et al.
Published: (2021)
Exploring the Mystery of Influential Data for Mathematical Reasoning
by: Ni, Xinzhe, et al.
Published: (2024)
by: Ni, Xinzhe, et al.
Published: (2024)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
by: Liang, Xiao, et al.
Published: (2025)
by: Liang, Xiao, et al.
Published: (2025)
PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images
by: Zhang, Shuoshuo, et al.
Published: (2025)
by: Zhang, Shuoshuo, et al.
Published: (2025)
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
by: Gou, Zhibin, et al.
Published: (2023)
by: Gou, Zhibin, et al.
Published: (2023)
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
by: Gou, Zhibin, et al.
Published: (2023)
by: Gou, Zhibin, et al.
Published: (2023)
Teaching Your Models to Understand Code via Focal Preference Alignment
by: Wu, Jie, et al.
Published: (2025)
by: Wu, Jie, et al.
Published: (2025)
A Local Valuation Criterion for Quadratic-Permutation Interleaved Zadoff--Chu Sequences
by: Zhang, Yutong, et al.
Published: (2026)
by: Zhang, Yutong, et al.
Published: (2026)
Generative Universal Verifier as Multimodal Meta-Reasoner
by: Zhang, Xinchen, et al.
Published: (2025)
by: Zhang, Xinchen, et al.
Published: (2025)
Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training
by: Luo, Zheheng, et al.
Published: (2024)
by: Luo, Zheheng, et al.
Published: (2024)
Enhancing Large Language Model Performance with Gradient-Based Parameter Selection
by: Li, Haoling, et al.
Published: (2024)
by: Li, Haoling, et al.
Published: (2024)
Adapting LLM Agents with Universal Feedback in Communication
by: Wang, Kuan, et al.
Published: (2023)
by: Wang, Kuan, et al.
Published: (2023)
How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning
by: Zhang, Xiangxiang, et al.
Published: (2026)
by: Zhang, Xiangxiang, et al.
Published: (2026)
Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning
by: Huang, Yiming, et al.
Published: (2024)
by: Huang, Yiming, et al.
Published: (2024)
LoongRL: Reinforcement Learning for Advanced Reasoning over Long Contexts
by: Wang, Siyuan, et al.
Published: (2025)
by: Wang, Siyuan, et al.
Published: (2025)
Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability
by: Liang, Xiao, et al.
Published: (2026)
by: Liang, Xiao, et al.
Published: (2026)
From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
by: Luo, Ruilin, et al.
Published: (2026)
by: Luo, Ruilin, et al.
Published: (2026)
Simple o3: Towards Interleaved Vision-Language Reasoning
by: Wang, Ye, et al.
Published: (2025)
by: Wang, Ye, et al.
Published: (2025)
Rho-1: Not All Tokens Are What You Need
by: Lin, Zhenghao, et al.
Published: (2024)
by: Lin, Zhenghao, et al.
Published: (2024)
Chain of History: Learning and Forecasting with LLMs for Temporal Knowledge Graph Completion
by: Luo, Ruilin, et al.
Published: (2024)
by: Luo, Ruilin, et al.
Published: (2024)
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
by: Chen, Chao, et al.
Published: (2025)
by: Chen, Chao, et al.
Published: (2025)
DeepTool: Scaling Interleaved Deliberation in Tool-Integrated Reasoning via Process-Supervised Reinforcement Learning
by: He, Yang, et al.
Published: (2026)
by: He, Yang, et al.
Published: (2026)
Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests
by: Wu, Jie, et al.
Published: (2026)
by: Wu, Jie, et al.
Published: (2026)
On the Construction and Correlation Properties of Permutation-Interleaved Zadoff-Chu Sequences
by: Yuan, Qin, et al.
Published: (2026)
by: Yuan, Qin, et al.
Published: (2026)
TextBind: Multi-turn Interleaved Multimodal Instruction-following in the Wild
by: Li, Huayang, et al.
Published: (2023)
by: Li, Huayang, et al.
Published: (2023)
Decoupling the Image Perception and Multimodal Reasoning for Reasoning Segmentation with Digital Twin Representations
by: Li, Yizhen, et al.
Published: (2025)
by: Li, Yizhen, et al.
Published: (2025)
InsCL: A Data-efficient Continual Learning Paradigm for Fine-tuning Large Language Models with Instructions
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation
by: Wang, Hongpeng, et al.
Published: (2026)
by: Wang, Hongpeng, et al.
Published: (2026)
A Unified Analysis of Stochastic Gradient Descent with Arbitrary Data Permutations and Beyond
by: Li, Yipeng, et al.
Published: (2025)
by: Li, Yipeng, et al.
Published: (2025)
Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR
by: Liang, Xiao, et al.
Published: (2025)
by: Liang, Xiao, et al.
Published: (2025)
Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
by: Zhan, Zheng, et al.
Published: (2025)
by: Zhan, Zheng, et al.
Published: (2025)
ProReflow: Progressive Reflow with Decomposed Velocity
by: Ke, Lei, et al.
Published: (2025)
by: Ke, Lei, et al.
Published: (2025)
HARP: Human-Assisted Regrouping with Permutation Invariant Critic for Multi-Agent Reinforcement Learning
by: Hu, Huawen, et al.
Published: (2024)
by: Hu, Huawen, et al.
Published: (2024)
Permutation Polynomial Interleaved Zadoff-Chu Sequences
by: Berggren, Fredrik, et al.
Published: (2023)
by: Berggren, Fredrik, et al.
Published: (2023)
SVIPTR: Fast and Efficient Scene Text Recognition with Vision Permutable Extractor
by: Cheng, Xianfu, et al.
Published: (2024)
by: Cheng, Xianfu, et al.
Published: (2024)
Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
by: Yang, Rushuai, et al.
Published: (2025)
by: Yang, Rushuai, et al.
Published: (2025)
Similar Items
-
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
by: Ding, Yang, et al.
Published: (2025) -
See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
by: Zhang, Shuoshuo, et al.
Published: (2025) -
PeRL: Permafrost Region Pond and Lake Database, links to ArcGIS shapefiles
by: Muster, Sina, et al.
Published: (2017) -
Permafrost-Region Lake-DOC version1 Database (PeRL-DOCv1)
by: Stolpmann, Lydia, et al.
Published: (2021) -
Exploring the Mystery of Influential Data for Mathematical Reasoning
by: Ni, Xinzhe, et al.
Published: (2024)