DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Jiahe, Zheng, Rongkun, Wang, Yi, Wang, Helin, Zhao, Hengshuang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DisCo: Disentangled Control for Realistic Human Dance Generation
von: Wang, Tan, et al.
Veröffentlicht: (2023)
von: Wang, Tan, et al.
Veröffentlicht: (2023)
Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
von: Zheng, Rongkun, et al.
Veröffentlicht: (2025)
von: Zheng, Rongkun, et al.
Veröffentlicht: (2025)
O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
von: Chen, Yuqing, et al.
Veröffentlicht: (2025)
von: Chen, Yuqing, et al.
Veröffentlicht: (2025)
SyncVIS: Synchronized Video Instance Segmentation
von: Zheng, Rongkun, et al.
Veröffentlicht: (2024)
von: Zheng, Rongkun, et al.
Veröffentlicht: (2024)
DisCo3D: Distilling Multi-View Consistency for 3D Scene Editing
von: Chi, Yufeng, et al.
Veröffentlicht: (2025)
von: Chi, Yufeng, et al.
Veröffentlicht: (2025)
ViLLa: Video Reasoning Segmentation with Large Language Model
von: Zheng, Rongkun, et al.
Veröffentlicht: (2024)
von: Zheng, Rongkun, et al.
Veröffentlicht: (2024)
TMT-VIS: Taxonomy-aware Multi-dataset Joint Training for Video Instance Segmentation
von: Zheng, Rongkun, et al.
Veröffentlicht: (2023)
von: Zheng, Rongkun, et al.
Veröffentlicht: (2023)
DisCo-Diff: Enhancing Continuous Diffusion Models with Discrete Latents
von: Xu, Yilun, et al.
Veröffentlicht: (2024)
von: Xu, Yilun, et al.
Veröffentlicht: (2024)
DisCo-FLoc: Semantic-Free Floorplan Localization via $SE(2)$-Aware Contrastive Disambiguation
von: Zhong, Ping, et al.
Veröffentlicht: (2026)
von: Zhong, Ping, et al.
Veröffentlicht: (2026)
DisCo-Layout: Disentangling and Coordinating Semantic and Physical Refinement in a Multi-Agent Framework for 3D Indoor Layout Synthesis
von: Gao, Jialin, et al.
Veröffentlicht: (2025)
von: Gao, Jialin, et al.
Veröffentlicht: (2025)
FocalClick-XL: Towards Unified and High-quality Interactive Segmentation
von: Chen, Xi, et al.
Veröffentlicht: (2025)
von: Chen, Xi, et al.
Veröffentlicht: (2025)
Exploring the Design Space of Visual Context Representation in Video MLLMs
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
von: Li, Qi, et al.
Veröffentlicht: (2026)
von: Li, Qi, et al.
Veröffentlicht: (2026)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
von: Zheng, Naishan, et al.
Veröffentlicht: (2025)
von: Zheng, Naishan, et al.
Veröffentlicht: (2025)
OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
von: Qi, Zhangyang, et al.
Veröffentlicht: (2025)
von: Qi, Zhangyang, et al.
Veröffentlicht: (2025)
SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
von: Tan, Shuai, et al.
Veröffentlicht: (2025)
von: Tan, Shuai, et al.
Veröffentlicht: (2025)
Unhackable Temporal Rewarding for Scalable Video MLLMs
von: Yu, En, et al.
Veröffentlicht: (2025)
von: Yu, En, et al.
Veröffentlicht: (2025)
Reinforcing Consistency in Video MLLMs with Structured Rewards
von: Quan, Yihao, et al.
Veröffentlicht: (2026)
von: Quan, Yihao, et al.
Veröffentlicht: (2026)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
von: Chen, Xi, et al.
Veröffentlicht: (2025)
von: Chen, Xi, et al.
Veröffentlicht: (2025)
UniMatch V2: Pushing the Limit of Semi-Supervised Semantic Segmentation
von: Yang, Lihe, et al.
Veröffentlicht: (2024)
von: Yang, Lihe, et al.
Veröffentlicht: (2024)
One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object Detection
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
LayerFlow: A Unified Model for Layer-aware Video Generation
von: Ji, Sihui, et al.
Veröffentlicht: (2025)
von: Ji, Sihui, et al.
Veröffentlicht: (2025)
Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs
von: Zhao, Zijia, et al.
Veröffentlicht: (2024)
von: Zhao, Zijia, et al.
Veröffentlicht: (2024)
RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees
von: Xu, Yichen, et al.
Veröffentlicht: (2026)
von: Xu, Yichen, et al.
Veröffentlicht: (2026)
SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
von: Yin, Yuanyang, et al.
Veröffentlicht: (2024)
von: Yin, Yuanyang, et al.
Veröffentlicht: (2024)
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
von: Wang, Haicheng, et al.
Veröffentlicht: (2026)
von: Wang, Haicheng, et al.
Veröffentlicht: (2026)
UniLION: Towards Unified Autonomous Driving Model with Linear Group RNNs
von: Liu, Zhe, et al.
Veröffentlicht: (2025)
von: Liu, Zhe, et al.
Veröffentlicht: (2025)
Video-R1: Reinforcing Video Reasoning in MLLMs
von: Feng, Kaituo, et al.
Veröffentlicht: (2025)
von: Feng, Kaituo, et al.
Veröffentlicht: (2025)
Towards Unified 3D Object Detection via Algorithm and Data Unification
von: Li, Zhuoling, et al.
Veröffentlicht: (2024)
von: Li, Zhuoling, et al.
Veröffentlicht: (2024)
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control
von: Tu, Yuanpeng, et al.
Veröffentlicht: (2025)
von: Tu, Yuanpeng, et al.
Veröffentlicht: (2025)
VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
von: Wang, Zhaozhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhaozhi, et al.
Veröffentlicht: (2025)
PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
von: Ji, Sihui, et al.
Veröffentlicht: (2025)
von: Ji, Sihui, et al.
Veröffentlicht: (2025)
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2025)
Decoupled Competitive Framework for Semi-supervised Medical Image Segmentation
von: Chen, Jiahe, et al.
Veröffentlicht: (2025)
von: Chen, Jiahe, et al.
Veröffentlicht: (2025)
Incentivizing Cardiologist-Like Reasoning in MLLMs for Interpretable Echocardiographic Diagnosis
von: Qin, Yi, et al.
Veröffentlicht: (2026)
von: Qin, Yi, et al.
Veröffentlicht: (2026)
MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DisCo: Disentangled Control for Realistic Human Dance Generation
von: Wang, Tan, et al.
Veröffentlicht: (2023) -
Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
von: Zheng, Rongkun, et al.
Veröffentlicht: (2025) -
O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
von: Chen, Yuqing, et al.
Veröffentlicht: (2025) -
SyncVIS: Synchronized Video Instance Segmentation
von: Zheng, Rongkun, et al.
Veröffentlicht: (2024) -
DisCo3D: Distilling Multi-View Consistency for 3D Scene Editing
von: Chi, Yufeng, et al.
Veröffentlicht: (2025)