LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Mo, Shentong, Yun, Sukmin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Visual Representation Alignment Generation with GRPO
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
MultiMed: Massively Multimodal and Multitask Medical Understanding
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
GMAIL: Generative Modality Alignment for generated Image Learning
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
Agentic Design Review System
by: Nag, Sayan, et al.
Published: (2025)
by: Nag, Sayan, et al.
Published: (2025)
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
by: Sun, Zeyi, et al.
Published: (2025)
by: Sun, Zeyi, et al.
Published: (2025)
DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Don't Guess, Escalate: Towards Explainable Uncertainty-Calibrated AI Forensic Agents
by: Boato, Giulia, et al.
Published: (2025)
by: Boato, Giulia, et al.
Published: (2025)
Paper2Video: Automatic Video Generation from Scientific Papers
by: Zhu, Zeyu, et al.
Published: (2025)
by: Zhu, Zeyu, et al.
Published: (2025)
IoT-LM: Large Multisensory Language Models for the Internet of Things
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Text-to-Audio Generation Synchronized with Videos
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Chain of Uncertain Rewards with Large Language Models for Reinforcement Learning
by: Mo, Shentong
Published: (2026)
by: Mo, Shentong
Published: (2026)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
by: Mo, Shentong, et al.
Published: (2023)
by: Mo, Shentong, et al.
Published: (2023)
Unified Video-Language Pre-training with Synchronized Audio
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
MedChat: A Multi-Agent Framework for Multimodal Diagnosis with Large Language Models
by: Liu, Philip R., et al.
Published: (2025)
by: Liu, Philip R., et al.
Published: (2025)
Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Driving
by: Wang, Zehao, et al.
Published: (2026)
by: Wang, Zehao, et al.
Published: (2026)
Chain of Questions: Guiding Multimodal Curiosity in Language Models
by: Iji, Nima, et al.
Published: (2025)
by: Iji, Nima, et al.
Published: (2025)
ProtoMedAgent: Multimodal Clinical Interpretability via Privacy-Aware Agentic Workflows
by: Pellicer, Alvaro Lopez, et al.
Published: (2026)
by: Pellicer, Alvaro Lopez, et al.
Published: (2026)
Multi-scale Multi-instance Visual Sound Localization and Segmentation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
by: Wang, Zhecan, et al.
Published: (2024)
by: Wang, Zhecan, et al.
Published: (2024)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
by: Li, Liupeng, et al.
Published: (2026)
by: Li, Liupeng, et al.
Published: (2026)
Gen-n-Val: Agentic Image Data Generation and Validation
by: Huang, Jing-En, et al.
Published: (2025)
by: Huang, Jing-En, et al.
Published: (2025)
Audio-visual Generalized Zero-shot Learning the Easy Way
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models
by: Ghosh, Dhruba, et al.
Published: (2026)
by: Ghosh, Dhruba, et al.
Published: (2026)
AllocMV: Optimal Resource Allocation for Music Video Generation via Structured Persistent State
by: Wang, Huimin, et al.
Published: (2026)
by: Wang, Huimin, et al.
Published: (2026)
FEAT: A Multi-Agent Forensic AI System with Domain-Adapted Large Language Model for Automated Cause-of-Death Analysis
by: Shen, Chen, et al.
Published: (2025)
by: Shen, Chen, et al.
Published: (2025)
Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
by: Mo, Shentong
Published: (2024)
by: Mo, Shentong
Published: (2024)
Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation
by: Zheng, Shuhong, et al.
Published: (2026)
by: Zheng, Shuhong, et al.
Published: (2026)
Judge Model for Large-scale Multimodality Benchmarks
by: Shih, Min-Han, et al.
Published: (2026)
by: Shih, Min-Han, et al.
Published: (2026)
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
by: Dong, Hao, et al.
Published: (2026)
by: Dong, Hao, et al.
Published: (2026)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
UniF$^2$ace: A Unified Fine-grained Face Understanding and Generation Model
by: Li, Junzhe, et al.
Published: (2025)
by: Li, Junzhe, et al.
Published: (2025)
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
by: Han, Junlin, et al.
Published: (2025)
by: Han, Junlin, et al.
Published: (2025)
Multi-Agent Game Generation and Evaluation via Audio-Visual Recordings
by: Jolicoeur-Martineau, Alexia
Published: (2025)
by: Jolicoeur-Martineau, Alexia
Published: (2025)
MATRIX: Multi-Agent Trajectory Generation with Diverse Contexts
by: Xu, Zhuo, et al.
Published: (2024)
by: Xu, Zhuo, et al.
Published: (2024)
Machine Unlearning in Hyperbolic vs. Euclidean Multimodal Contrastive Learning: Adapting Alignment Calibration to MERU
by: Vidal, Àlex Pujol, et al.
Published: (2025)
by: Vidal, Àlex Pujol, et al.
Published: (2025)
Multi-Agent VQA: Exploring Multi-Agent Foundation Models in Zero-Shot Visual Question Answering
by: Jiang, Bowen, et al.
Published: (2024)
by: Jiang, Bowen, et al.
Published: (2024)
Picking the Right Specialist: Attentive Neural Process-based Selection of Task-Specialized Models as Tools for Agentic Healthcare Systems
by: Saha, Pramit, et al.
Published: (2026)
by: Saha, Pramit, et al.
Published: (2026)
ToolTok: Tool Tokenization for Efficient and Generalizable GUI Agents
by: Wang, Xiaoce, et al.
Published: (2026)
by: Wang, Xiaoce, et al.
Published: (2026)
Experience-Driven Multi-Agent Systems Are Training-free Context-aware Earth Observers
by: Dai, Pengyu, et al.
Published: (2026)
by: Dai, Pengyu, et al.
Published: (2026)
Similar Items
-
Improving Visual Representation Alignment Generation with GRPO
by: Mo, Shentong, et al.
Published: (2026) -
MultiMed: Massively Multimodal and Multitask Medical Understanding
by: Mo, Shentong, et al.
Published: (2024) -
GMAIL: Generative Modality Alignment for generated Image Learning
by: Mo, Shentong, et al.
Published: (2026) -
Agentic Design Review System
by: Nag, Sayan, et al.
Published: (2025) -
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
by: Sun, Zeyi, et al.
Published: (2025)