Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fei, Hao, Wu, Shengqiong, Zhang, Hanwang, Chua, Tat-Seng, Yan, Shuicheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Modeling Cross-vision Synergy for Unified Large Vision Model
von: Wu, Shengqiong, et al.
Veröffentlicht: (2026)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2026)
Towards Semantic Equivalence of Tokenization in Multimodal LLM
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024)
Universal Scene Graph Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024)
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
von: Fei, Hao, et al.
Veröffentlicht: (2024)
von: Fei, Hao, et al.
Veröffentlicht: (2024)
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
von: Fei, Hao, et al.
Veröffentlicht: (2023)
von: Fei, Hao, et al.
Veröffentlicht: (2023)
SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation
von: Li, Zisu, et al.
Veröffentlicht: (2025)
von: Li, Zisu, et al.
Veröffentlicht: (2025)
UI-UG: A Unified MLLM for UI Understanding and Generation
von: Yang, Hao, et al.
Veröffentlicht: (2025)
von: Yang, Hao, et al.
Veröffentlicht: (2025)
Global Commander and Local Operative: A Dual-Agent Framework for Scene Navigation
von: Jin, Kaiming, et al.
Veröffentlicht: (2026)
von: Jin, Kaiming, et al.
Veröffentlicht: (2026)
VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
Towards Context-aware Support for Color Vision Deficiency: An Approach Integrating LLM and AR
von: Morita, Shogo, et al.
Veröffentlicht: (2024)
von: Morita, Shogo, et al.
Veröffentlicht: (2024)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
von: Zhang, Tao, et al.
Veröffentlicht: (2024)
von: Zhang, Tao, et al.
Veröffentlicht: (2024)
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision
von: Li, Chentao, et al.
Veröffentlicht: (2026)
von: Li, Chentao, et al.
Veröffentlicht: (2026)
Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking
von: Wu, Shengqiong, et al.
Veröffentlicht: (2026)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2026)
Dreamcrafter: Immersive Editing of 3D Radiance Fields Through Flexible, Generative Inputs and Outputs
von: Vachha, Cyrus, et al.
Veröffentlicht: (2025)
von: Vachha, Cyrus, et al.
Veröffentlicht: (2025)
Interactivity x Explainability: Toward Understanding How Interactivity Can Improve Computer Vision Explanations
von: Panigrahi, Indu, et al.
Veröffentlicht: (2025)
von: Panigrahi, Indu, et al.
Veröffentlicht: (2025)
LLM4Brain: Training a Large Language Model for Brain Video Understanding
von: Zheng, Ruizhe, et al.
Veröffentlicht: (2024)
von: Zheng, Ruizhe, et al.
Veröffentlicht: (2024)
VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation
von: Pan, Bo, et al.
Veröffentlicht: (2025)
von: Pan, Bo, et al.
Veröffentlicht: (2025)
PixelWeb: The First Web GUI Dataset with Pixel-Wise Labels
von: Yang, Qi, et al.
Veröffentlicht: (2025)
von: Yang, Qi, et al.
Veröffentlicht: (2025)
Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding
von: Guo, Hao, et al.
Veröffentlicht: (2025)
von: Guo, Hao, et al.
Veröffentlicht: (2025)
Unified Understanding of Environment, Task, and Human for Human-Robot Interaction in Real-World Environments
von: Yano, Yuga, et al.
Veröffentlicht: (2024)
von: Yano, Yuga, et al.
Veröffentlicht: (2024)
Reframe Anything: LLM Agent for Open World Video Reframing
von: Cao, Jiawang, et al.
Veröffentlicht: (2024)
von: Cao, Jiawang, et al.
Veröffentlicht: (2024)
Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble
von: Duan, Lin, et al.
Veröffentlicht: (2025)
von: Duan, Lin, et al.
Veröffentlicht: (2025)
CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning
von: Ni, Minheng, et al.
Veröffentlicht: (2026)
von: Ni, Minheng, et al.
Veröffentlicht: (2026)
SemLayer: Semantic-aware Generative Segmentation and Layer Construction for Abstract Icons
von: Xu, Haiyang, et al.
Veröffentlicht: (2026)
von: Xu, Haiyang, et al.
Veröffentlicht: (2026)
Toward Scalable Co-located Practical Learning: Assisting with Computer Vision and Multimodal Analytics
von: Li, Xinyu, et al.
Veröffentlicht: (2026)
von: Li, Xinyu, et al.
Veröffentlicht: (2026)
Computational Trichromacy Reconstruction: Empowering the Color-Vision Deficient to Recognize Colors Using Augmented Reality
von: Zhu, Yuhao, et al.
Veröffentlicht: (2024)
von: Zhu, Yuhao, et al.
Veröffentlicht: (2024)
Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scene
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
Generative Augmented Reality: Paradigms, Technologies, and Future Applications
von: Liang, Chen, et al.
Veröffentlicht: (2025)
von: Liang, Chen, et al.
Veröffentlicht: (2025)
PromptArtisan: Multi-instruction Image Editing in Single Pass with Complete Attention Control
von: Swami, Kunal, et al.
Veröffentlicht: (2025)
von: Swami, Kunal, et al.
Veröffentlicht: (2025)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
von: Verma, Arnav, et al.
Veröffentlicht: (2025)
von: Verma, Arnav, et al.
Veröffentlicht: (2025)
VisionCAD: An Integration-Free Radiology Copilot Framework
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
HaDR: Applying Domain Randomization for Generating Synthetic Multimodal Dataset for Hand Instance Segmentation in Cluttered Industrial Environments
von: Grushko, Stefan, et al.
Veröffentlicht: (2023)
von: Grushko, Stefan, et al.
Veröffentlicht: (2023)
AgentSense: Virtual Sensor Data Generation Using LLM Agents in Simulated Home Environments
von: Leng, Zikang, et al.
Veröffentlicht: (2025)
von: Leng, Zikang, et al.
Veröffentlicht: (2025)
Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference
von: Zhu, Yitong, et al.
Veröffentlicht: (2025)
von: Zhu, Yitong, et al.
Veröffentlicht: (2025)
Do Vision Language Models Understand Human Engagement in Games?
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
3DArticCyclists: Generating Synthetic Articulated 8D Pose-Controllable Cyclist Data for Computer Vision Applications
von: Corral-Soto, Eduardo R., et al.
Veröffentlicht: (2024)
von: Corral-Soto, Eduardo R., et al.
Veröffentlicht: (2024)
ChatStitch: Visualizing Through Structures via Surround-View Unsupervised Deep Image Stitching with Collaborative LLM-Agents
von: Liang, Hao, et al.
Veröffentlicht: (2025)
von: Liang, Hao, et al.
Veröffentlicht: (2025)
Has the Virtualization of the Face Changed Facial Perception? A Study of the Impact of Photo Editing and Augmented Reality on Facial Perception
von: Conwill, Louisa, et al.
Veröffentlicht: (2023)
von: Conwill, Louisa, et al.
Veröffentlicht: (2023)
Editing Physiological Signals in Videos Using Latent Representations
von: Zhou, Tianwen, et al.
Veröffentlicht: (2025)
von: Zhou, Tianwen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Modeling Cross-vision Synergy for Unified Large Vision Model
von: Wu, Shengqiong, et al.
Veröffentlicht: (2026) -
Towards Semantic Equivalence of Tokenization in Multimodal LLM
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024) -
Universal Scene Graph Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025) -
Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024) -
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
von: Fei, Hao, et al.
Veröffentlicht: (2024)