Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Heo, Miran, Chen, Min-Hung, Huang, De-An, Liu, Sifei, Radhakrishnan, Subhashree, Kim, Seon Joo, Wang, Yu-Chiang Frank, Hachiuma, Ryo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
von: Yang, Chiao-An, et al.
Veröffentlicht: (2025)
von: Yang, Chiao-An, et al.
Veröffentlicht: (2025)
Autoregressive Universal Video Segmentation Model
von: Heo, Miran, et al.
Veröffentlicht: (2025)
von: Heo, Miran, et al.
Veröffentlicht: (2025)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025)
VISAGE: Video Instance Segmentation with Appearance-Guided Enhancement
von: Kim, Hanjung, et al.
Veröffentlicht: (2023)
von: Kim, Hanjung, et al.
Veröffentlicht: (2023)
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
von: Huang, De-An, et al.
Veröffentlicht: (2025)
von: Huang, De-An, et al.
Veröffentlicht: (2025)
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2024)
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2024)
VIOLA: Towards Video In-Context Learning with Minimal Annotations
von: Fujii, Ryo, et al.
Veröffentlicht: (2026)
von: Fujii, Ryo, et al.
Veröffentlicht: (2026)
Weakly Semi-supervised Tool Detection in Minimally Invasive Surgery Videos
von: Fujii, Ryo, et al.
Veröffentlicht: (2024)
von: Fujii, Ryo, et al.
Veröffentlicht: (2024)
Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
Unified Reinforcement and Imitation Learning for Vision-Language Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
SANER: Annotation-free Societal Attribute Neutralizer for Debiasing CLIP
von: Hirota, Yusuke, et al.
Veröffentlicht: (2024)
von: Hirota, Yusuke, et al.
Veröffentlicht: (2024)
EMAG: Ego-motion Aware and Generalizable 2D Hand Forecasting from Egocentric Videos
von: Hatano, Masashi, et al.
Veröffentlicht: (2024)
von: Hatano, Masashi, et al.
Veröffentlicht: (2024)
Omni-Video: Democratizing Unified Video Understanding and Generation
von: Tan, Zhiyu, et al.
Veröffentlicht: (2025)
von: Tan, Zhiyu, et al.
Veröffentlicht: (2025)
V2V-GoT: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-of-Thoughts
von: Chiu, Hsu-kuang, et al.
Veröffentlicht: (2025)
von: Chiu, Hsu-kuang, et al.
Veröffentlicht: (2025)
V2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models
von: Chiu, Hsu-kuang, et al.
Veröffentlicht: (2025)
von: Chiu, Hsu-kuang, et al.
Veröffentlicht: (2025)
CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
von: Fujii, Ryo, et al.
Veröffentlicht: (2024)
von: Fujii, Ryo, et al.
Veröffentlicht: (2024)
Towards Predicting Any Human Trajectory In Context
von: Fujii, Ryo, et al.
Veröffentlicht: (2025)
von: Fujii, Ryo, et al.
Veröffentlicht: (2025)
RealTraj: Towards Real-World Pedestrian Trajectory Forecasting
von: Fujii, Ryo, et al.
Veröffentlicht: (2024)
von: Fujii, Ryo, et al.
Veröffentlicht: (2024)
Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2025)
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2025)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
von: Wang, Junke, et al.
Veröffentlicht: (2024)
von: Wang, Junke, et al.
Veröffentlicht: (2024)
GroPrompt: Efficient Grounded Prompting and Adaptation for Referring Video Object Segmentation
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2024)
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2024)
Video Token Merging for Long-form Video Understanding
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
3D-Layout-R1: Structured Reasoning for Language-Instructed Spatial Editing
von: Zhen, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhen, Haoyu, et al.
Veröffentlicht: (2026)
Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
von: Kang, Hyolim, et al.
Veröffentlicht: (2025)
von: Kang, Hyolim, et al.
Veröffentlicht: (2025)
Multimodal Cross-Domain Few-Shot Learning for Egocentric Action Recognition
von: Hatano, Masashi, et al.
Veröffentlicht: (2024)
von: Hatano, Masashi, et al.
Veröffentlicht: (2024)
Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
von: Ishikawa, Reina, et al.
Veröffentlicht: (2025)
von: Ishikawa, Reina, et al.
Veröffentlicht: (2025)
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
Learning to Enhance Aperture Phasor Field for Non-Line-of-Sight Imaging
von: Cho, In, et al.
Veröffentlicht: (2024)
von: Cho, In, et al.
Veröffentlicht: (2024)
VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
Omni$^2$: Unifying Omnidirectional Image Generation and Editing in an Omni Model
von: Yang, Liu, et al.
Veröffentlicht: (2025)
von: Yang, Liu, et al.
Veröffentlicht: (2025)
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
von: Ye, Hanrong, et al.
Veröffentlicht: (2025)
von: Ye, Hanrong, et al.
Veröffentlicht: (2025)
3D Aware Region Prompted Vision Language Model
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2025)
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2025)
Morphing Tokens Draw Strong Masked Image Models
von: Kim, Taekyung, et al.
Veröffentlicht: (2023)
von: Kim, Taekyung, et al.
Veröffentlicht: (2023)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
von: Qu, Liao, et al.
Veröffentlicht: (2024)
von: Qu, Liao, et al.
Veröffentlicht: (2024)
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
von: Hirota, Yusuke, et al.
Veröffentlicht: (2024)
von: Hirota, Yusuke, et al.
Veröffentlicht: (2024)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
von: Xi, Dianbing, et al.
Veröffentlicht: (2025)
von: Xi, Dianbing, et al.
Veröffentlicht: (2025)
DiSC: Resolution-Scalable Acceleration of Diffusion Models by Exploiting Sparsity and Cached Token Reuse with Hash-based Distribution
von: Yoon, Jieon, et al.
Veröffentlicht: (2026)
von: Yoon, Jieon, et al.
Veröffentlicht: (2026)
RegionGPT: Towards Region Understanding Vision Language Model
von: Guo, Qiushan, et al.
Veröffentlicht: (2024)
von: Guo, Qiushan, et al.
Veröffentlicht: (2024)
Accelerating Image Super-Resolution Networks with Pixel-Level Classification
von: Jeong, Jinho, et al.
Veröffentlicht: (2024)
von: Jeong, Jinho, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
von: Yang, Chiao-An, et al.
Veröffentlicht: (2025) -
Autoregressive Universal Video Segmentation Model
von: Heo, Miran, et al.
Veröffentlicht: (2025) -
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025) -
VISAGE: Video Instance Segmentation with Appearance-Guided Enhancement
von: Kim, Hanjung, et al.
Veröffentlicht: (2023) -
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
von: Huang, De-An, et al.
Veröffentlicht: (2025)