PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yizhen, Ding, Yang, Zhang, Shuoshuo, Zhang, Xinchen, Li, Haoling, Li, Zhong-zhi, Wang, Peijie, Wu, Jie, Ji, Lei, Shen, Yelong, Yang, Yujiu, Gong, Yeyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
von: Ding, Yang, et al.
Veröffentlicht: (2025)
von: Ding, Yang, et al.
Veröffentlicht: (2025)
See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
von: Zhang, Shuoshuo, et al.
Veröffentlicht: (2025)
von: Zhang, Shuoshuo, et al.
Veröffentlicht: (2025)
PeRL: Permafrost Region Pond and Lake Database, links to ArcGIS shapefiles
von: Muster, Sina, et al.
Veröffentlicht: (2017)
von: Muster, Sina, et al.
Veröffentlicht: (2017)
Permafrost-Region Lake-DOC version1 Database (PeRL-DOCv1)
von: Stolpmann, Lydia, et al.
Veröffentlicht: (2021)
von: Stolpmann, Lydia, et al.
Veröffentlicht: (2021)
Exploring the Mystery of Influential Data for Mathematical Reasoning
von: Ni, Xinzhe, et al.
Veröffentlicht: (2024)
von: Ni, Xinzhe, et al.
Veröffentlicht: (2024)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images
von: Zhang, Shuoshuo, et al.
Veröffentlicht: (2025)
von: Zhang, Shuoshuo, et al.
Veröffentlicht: (2025)
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
von: Gou, Zhibin, et al.
Veröffentlicht: (2023)
von: Gou, Zhibin, et al.
Veröffentlicht: (2023)
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
von: Gou, Zhibin, et al.
Veröffentlicht: (2023)
von: Gou, Zhibin, et al.
Veröffentlicht: (2023)
Teaching Your Models to Understand Code via Focal Preference Alignment
von: Wu, Jie, et al.
Veröffentlicht: (2025)
von: Wu, Jie, et al.
Veröffentlicht: (2025)
A Local Valuation Criterion for Quadratic-Permutation Interleaved Zadoff--Chu Sequences
von: Zhang, Yutong, et al.
Veröffentlicht: (2026)
von: Zhang, Yutong, et al.
Veröffentlicht: (2026)
Generative Universal Verifier as Multimodal Meta-Reasoner
von: Zhang, Xinchen, et al.
Veröffentlicht: (2025)
von: Zhang, Xinchen, et al.
Veröffentlicht: (2025)
Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training
von: Luo, Zheheng, et al.
Veröffentlicht: (2024)
von: Luo, Zheheng, et al.
Veröffentlicht: (2024)
Enhancing Large Language Model Performance with Gradient-Based Parameter Selection
von: Li, Haoling, et al.
Veröffentlicht: (2024)
von: Li, Haoling, et al.
Veröffentlicht: (2024)
Adapting LLM Agents with Universal Feedback in Communication
von: Wang, Kuan, et al.
Veröffentlicht: (2023)
von: Wang, Kuan, et al.
Veröffentlicht: (2023)
How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning
von: Zhang, Xiangxiang, et al.
Veröffentlicht: (2026)
von: Zhang, Xiangxiang, et al.
Veröffentlicht: (2026)
Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning
von: Huang, Yiming, et al.
Veröffentlicht: (2024)
von: Huang, Yiming, et al.
Veröffentlicht: (2024)
LoongRL: Reinforcement Learning for Advanced Reasoning over Long Contexts
von: Wang, Siyuan, et al.
Veröffentlicht: (2025)
von: Wang, Siyuan, et al.
Veröffentlicht: (2025)
Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability
von: Liang, Xiao, et al.
Veröffentlicht: (2026)
von: Liang, Xiao, et al.
Veröffentlicht: (2026)
From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
von: Luo, Ruilin, et al.
Veröffentlicht: (2026)
von: Luo, Ruilin, et al.
Veröffentlicht: (2026)
Simple o3: Towards Interleaved Vision-Language Reasoning
von: Wang, Ye, et al.
Veröffentlicht: (2025)
von: Wang, Ye, et al.
Veröffentlicht: (2025)
Rho-1: Not All Tokens Are What You Need
von: Lin, Zhenghao, et al.
Veröffentlicht: (2024)
von: Lin, Zhenghao, et al.
Veröffentlicht: (2024)
Chain of History: Learning and Forecasting with LLMs for Temporal Knowledge Graph Completion
von: Luo, Ruilin, et al.
Veröffentlicht: (2024)
von: Luo, Ruilin, et al.
Veröffentlicht: (2024)
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
von: Chen, Chao, et al.
Veröffentlicht: (2025)
von: Chen, Chao, et al.
Veröffentlicht: (2025)
DeepTool: Scaling Interleaved Deliberation in Tool-Integrated Reasoning via Process-Supervised Reinforcement Learning
von: He, Yang, et al.
Veröffentlicht: (2026)
von: He, Yang, et al.
Veröffentlicht: (2026)
Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests
von: Wu, Jie, et al.
Veröffentlicht: (2026)
von: Wu, Jie, et al.
Veröffentlicht: (2026)
On the Construction and Correlation Properties of Permutation-Interleaved Zadoff-Chu Sequences
von: Yuan, Qin, et al.
Veröffentlicht: (2026)
von: Yuan, Qin, et al.
Veröffentlicht: (2026)
TextBind: Multi-turn Interleaved Multimodal Instruction-following in the Wild
von: Li, Huayang, et al.
Veröffentlicht: (2023)
von: Li, Huayang, et al.
Veröffentlicht: (2023)
Decoupling the Image Perception and Multimodal Reasoning for Reasoning Segmentation with Digital Twin Representations
von: Li, Yizhen, et al.
Veröffentlicht: (2025)
von: Li, Yizhen, et al.
Veröffentlicht: (2025)
InsCL: A Data-efficient Continual Learning Paradigm for Fine-tuning Large Language Models with Instructions
von: Wang, Yifan, et al.
Veröffentlicht: (2024)
von: Wang, Yifan, et al.
Veröffentlicht: (2024)
MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation
von: Wang, Hongpeng, et al.
Veröffentlicht: (2026)
von: Wang, Hongpeng, et al.
Veröffentlicht: (2026)
A Unified Analysis of Stochastic Gradient Descent with Arbitrary Data Permutations and Beyond
von: Li, Yipeng, et al.
Veröffentlicht: (2025)
von: Li, Yipeng, et al.
Veröffentlicht: (2025)
Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
von: Zhan, Zheng, et al.
Veröffentlicht: (2025)
von: Zhan, Zheng, et al.
Veröffentlicht: (2025)
ProReflow: Progressive Reflow with Decomposed Velocity
von: Ke, Lei, et al.
Veröffentlicht: (2025)
von: Ke, Lei, et al.
Veröffentlicht: (2025)
HARP: Human-Assisted Regrouping with Permutation Invariant Critic for Multi-Agent Reinforcement Learning
von: Hu, Huawen, et al.
Veröffentlicht: (2024)
von: Hu, Huawen, et al.
Veröffentlicht: (2024)
Permutation Polynomial Interleaved Zadoff-Chu Sequences
von: Berggren, Fredrik, et al.
Veröffentlicht: (2023)
von: Berggren, Fredrik, et al.
Veröffentlicht: (2023)
SVIPTR: Fast and Efficient Scene Text Recognition with Vision Permutable Extractor
von: Cheng, Xianfu, et al.
Veröffentlicht: (2024)
von: Cheng, Xianfu, et al.
Veröffentlicht: (2024)
Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
von: Yang, Rushuai, et al.
Veröffentlicht: (2025)
von: Yang, Rushuai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
von: Ding, Yang, et al.
Veröffentlicht: (2025) -
See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
von: Zhang, Shuoshuo, et al.
Veröffentlicht: (2025) -
PeRL: Permafrost Region Pond and Lake Database, links to ArcGIS shapefiles
von: Muster, Sina, et al.
Veröffentlicht: (2017) -
Permafrost-Region Lake-DOC version1 Database (PeRL-DOCv1)
von: Stolpmann, Lydia, et al.
Veröffentlicht: (2021) -
Exploring the Mystery of Influential Data for Mathematical Reasoning
von: Ni, Xinzhe, et al.
Veröffentlicht: (2024)