Improving Visual Representation Alignment Generation with GRPO
Fuente:
arXiv
Guardado en:
| Autores principales: | Mo, Shentong, Yun, Sukmin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
por: Mo, Shentong, et al.
Publicado: (2026)
por: Mo, Shentong, et al.
Publicado: (2026)
GMAIL: Generative Modality Alignment for generated Image Learning
por: Mo, Shentong, et al.
Publicado: (2026)
por: Mo, Shentong, et al.
Publicado: (2026)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
MultiMed: Massively Multimodal and Multitask Medical Understanding
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Text-to-Audio Generation Synchronized with Videos
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
por: Mo, Shentong, et al.
Publicado: (2023)
por: Mo, Shentong, et al.
Publicado: (2023)
IoT-LM: Large Multisensory Language Models for the Internet of Things
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Unified Video-Language Pre-training with Synchronized Audio
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Multi-scale Multi-instance Visual Sound Localization and Segmentation
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Audio-visual Generalized Zero-shot Learning the Easy Way
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
por: Mo, Shentong
Publicado: (2024)
por: Mo, Shentong
Publicado: (2024)
GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
por: Li, Baiqi, et al.
Publicado: (2024)
por: Li, Baiqi, et al.
Publicado: (2024)
HOIN: High-Order Implicit Neural Representations
por: Chen, Yang, et al.
Publicado: (2024)
por: Chen, Yang, et al.
Publicado: (2024)
Semantic Grouping Network for Audio Source Separation
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment
por: Hong, Jiaying, et al.
Publicado: (2025)
por: Hong, Jiaying, et al.
Publicado: (2025)
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
por: Mo, Shentong, et al.
Publicado: (2026)
por: Mo, Shentong, et al.
Publicado: (2026)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
por: Lin, Zhiqiu, et al.
Publicado: (2024)
por: Lin, Zhiqiu, et al.
Publicado: (2024)
LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Progressive Confident Masking Attention Network for Audio-Visual Segmentation
por: Wang, Yuxuan, et al.
Publicado: (2024)
por: Wang, Yuxuan, et al.
Publicado: (2024)
Long-tailed Medical Diagnosis with Relation-aware Representation Learning and Iterative Classifier Calibration
por: Pan, Li, et al.
Publicado: (2025)
por: Pan, Li, et al.
Publicado: (2025)
HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
por: Wang, Zhecan, et al.
Publicado: (2024)
por: Wang, Zhecan, et al.
Publicado: (2024)
PAND: Prompt-Aware Neighborhood Distillation for Lightweight Fine-Grained Visual Classification
por: Luo, Qiuming, et al.
Publicado: (2026)
por: Luo, Qiuming, et al.
Publicado: (2026)
MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs
por: Barrios, Wayner, et al.
Publicado: (2025)
por: Barrios, Wayner, et al.
Publicado: (2025)
Neuron Abandoning Attention Flow: Visual Explanation of Dynamics inside CNN Models
por: Liao, Yi, et al.
Publicado: (2024)
por: Liao, Yi, et al.
Publicado: (2024)
Make VLM Recognize Visual Hallucination on Cartoon Character Image with Pose Information
por: Kim, Bumsoo, et al.
Publicado: (2024)
por: Kim, Bumsoo, et al.
Publicado: (2024)
DesignAsCode: Bridging Structural Editability and Visual Fidelity in Graphic Design Generation
por: Liu, Ziyuan, et al.
Publicado: (2026)
por: Liu, Ziyuan, et al.
Publicado: (2026)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
por: Li, Liupeng, et al.
Publicado: (2026)
por: Li, Liupeng, et al.
Publicado: (2026)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
por: Chen, Yang, et al.
Publicado: (2024)
por: Chen, Yang, et al.
Publicado: (2024)
Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
por: Han, Junlin, et al.
Publicado: (2025)
por: Han, Junlin, et al.
Publicado: (2025)
Flow Generator Matching
por: Huang, Zemin, et al.
Publicado: (2024)
por: Huang, Zemin, et al.
Publicado: (2024)
Generating Illustrated Instructions
por: Menon, Sachit, et al.
Publicado: (2023)
por: Menon, Sachit, et al.
Publicado: (2023)
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
por: Qu, Leigang, et al.
Publicado: (2024)
por: Qu, Leigang, et al.
Publicado: (2024)
GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
por: Duan, Chengqi, et al.
Publicado: (2025)
por: Duan, Chengqi, et al.
Publicado: (2025)
Meta-CoT: Enhancing Granularity and Generalization in Image Editing
por: Zhang, Shiyi, et al.
Publicado: (2026)
por: Zhang, Shiyi, et al.
Publicado: (2026)
STIV: Scalable Text and Image Conditioned Video Generation
por: Lin, Zongyu, et al.
Publicado: (2024)
por: Lin, Zongyu, et al.
Publicado: (2024)
VSTAR: Generative Temporal Nursing for Longer Dynamic Video Synthesis
por: Li, Yumeng, et al.
Publicado: (2024)
por: Li, Yumeng, et al.
Publicado: (2024)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
por: Tang, Shixiang, et al.
Publicado: (2025)
por: Tang, Shixiang, et al.
Publicado: (2025)
Reinforcement Learning for Unsupervised Video Summarization with Reward Generator Training
por: Abbasi, Mehryar, et al.
Publicado: (2024)
por: Abbasi, Mehryar, et al.
Publicado: (2024)
Ejemplares similares
-
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
por: Mo, Shentong, et al.
Publicado: (2026) -
GMAIL: Generative Modality Alignment for generated Image Learning
por: Mo, Shentong, et al.
Publicado: (2026) -
Aligning Audio-Visual Joint Representations with an Agentic Workflow
por: Mo, Shentong, et al.
Publicado: (2024) -
DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture
por: Mo, Shentong, et al.
Publicado: (2024) -
MultiMed: Massively Multimodal and Multitask Medical Understanding
por: Mo, Shentong, et al.
Publicado: (2024)