UniVS: Unified and Universal Video Segmentation with Prompts as Queries
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Minghan, Li, Shuai, Zhang, Xindong, Zhang, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UniAIDet: A Unified and Universal Benchmark for AI-Generated Image Content Detection and Localization
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
UniChange: Unifying Change Detection with Multimodal Large Language Model
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
Uni-SMART: Universal Science Multimodal Analysis and Research Transformer
von: Cai, Hengxing, et al.
Veröffentlicht: (2024)
von: Cai, Hengxing, et al.
Veröffentlicht: (2024)
Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
von: Qin, Luozheng, et al.
Veröffentlicht: (2025)
von: Qin, Luozheng, et al.
Veröffentlicht: (2025)
UniAPO: Unified Multimodal Automated Prompt Optimization
von: Zhu, Qipeng, et al.
Veröffentlicht: (2025)
von: Zhu, Qipeng, et al.
Veröffentlicht: (2025)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
Spectral Prompt Tuning:Unveiling Unseen Classes for Zero-Shot Semantic Segmentation
von: Xu, Wenhao, et al.
Veröffentlicht: (2023)
von: Xu, Wenhao, et al.
Veröffentlicht: (2023)
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning
von: Li, Zheng, et al.
Veröffentlicht: (2025)
von: Li, Zheng, et al.
Veröffentlicht: (2025)
UniD-Shift: Towards Unified Semantic Segmentation via Interpretable Share-Private Multimodal Decomposition
von: Zhang, Shuai, et al.
Veröffentlicht: (2026)
von: Zhang, Shuai, et al.
Veröffentlicht: (2026)
Frame-Voyager: Learning to Query Frames for Video Large Language Models
von: Yu, Sicheng, et al.
Veröffentlicht: (2024)
von: Yu, Sicheng, et al.
Veröffentlicht: (2024)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
UniVBench: Towards Unified Evaluation for Video Foundation Models
von: Wei, Jianhui, et al.
Veröffentlicht: (2026)
von: Wei, Jianhui, et al.
Veröffentlicht: (2026)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
Universal Prompt Optimizer for Safe Text-to-Image Generation
von: Wu, Zongyu, et al.
Veröffentlicht: (2024)
von: Wu, Zongyu, et al.
Veröffentlicht: (2024)
Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing Bias
von: Wan, Zhongwei, et al.
Veröffentlicht: (2023)
von: Wan, Zhongwei, et al.
Veröffentlicht: (2023)
UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
von: Zhu, Yijie, et al.
Veröffentlicht: (2025)
von: Zhu, Yijie, et al.
Veröffentlicht: (2025)
UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding
von: Wang, Zhecan, et al.
Veröffentlicht: (2023)
von: Wang, Zhecan, et al.
Veröffentlicht: (2023)
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
von: Luo, Fuwen, et al.
Veröffentlicht: (2025)
von: Luo, Fuwen, et al.
Veröffentlicht: (2025)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
von: Lin, Bin, et al.
Veröffentlicht: (2025)
von: Lin, Bin, et al.
Veröffentlicht: (2025)
MaSS13K: A Matting-level Semantic Segmentation Benchmark
von: Xie, Chenxi, et al.
Veröffentlicht: (2025)
von: Xie, Chenxi, et al.
Veröffentlicht: (2025)
Efficient Universal Goal Hijacking with Semantics-guided Prompt Organization
von: Huang, Yihao, et al.
Veröffentlicht: (2024)
von: Huang, Yihao, et al.
Veröffentlicht: (2024)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
von: Shlapentokh-Rothman, Michal, et al.
Veröffentlicht: (2026)
von: Shlapentokh-Rothman, Michal, et al.
Veröffentlicht: (2026)
UniCell: Universal Cell Nucleus Classification via Prompt Learning
von: Huang, Junjia, et al.
Veröffentlicht: (2024)
von: Huang, Junjia, et al.
Veröffentlicht: (2024)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
von: Tu, Yunbin, et al.
Veröffentlicht: (2024)
von: Tu, Yunbin, et al.
Veröffentlicht: (2024)
UniVid: The Open-Source Unified Video Model
von: Luo, Jiabin, et al.
Veröffentlicht: (2025)
von: Luo, Jiabin, et al.
Veröffentlicht: (2025)
Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting
von: Cai, Chen, et al.
Veröffentlicht: (2024)
von: Cai, Chen, et al.
Veröffentlicht: (2024)
MoPD: Mixture-of-Prompts Distillation for Vision-Language Models
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
von: Jin, Yang, et al.
Veröffentlicht: (2024)
von: Jin, Yang, et al.
Veröffentlicht: (2024)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
von: Liu, Fanfan, et al.
Veröffentlicht: (2024)
von: Liu, Fanfan, et al.
Veröffentlicht: (2024)
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
von: Guo, Pinxue, et al.
Veröffentlicht: (2024)
von: Guo, Pinxue, et al.
Veröffentlicht: (2024)
MMViR: A Multi-Modal and Multi-Granularity Representation for Long-range Video Understanding
von: Li, Zizhong, et al.
Veröffentlicht: (2026)
von: Li, Zizhong, et al.
Veröffentlicht: (2026)
UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
von: Ni, Shuo, et al.
Veröffentlicht: (2025)
von: Ni, Shuo, et al.
Veröffentlicht: (2025)
Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs
von: Peng, Shangpin, et al.
Veröffentlicht: (2025)
von: Peng, Shangpin, et al.
Veröffentlicht: (2025)
DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing
von: Xie, Chenxi, et al.
Veröffentlicht: (2025)
von: Xie, Chenxi, et al.
Veröffentlicht: (2025)
UniVideo: Unified Understanding, Generation, and Editing for Videos
von: Wei, Cong, et al.
Veröffentlicht: (2025)
von: Wei, Cong, et al.
Veröffentlicht: (2025)
Video Object Segmentation with Dynamic Query Modulation
von: Zhou, Hantao, et al.
Veröffentlicht: (2024)
von: Zhou, Hantao, et al.
Veröffentlicht: (2024)
FiVE: A Fine-grained Video Editing Benchmark for Evaluating Emerging Diffusion and Rectified Flow Models
von: Li, Minghan, et al.
Veröffentlicht: (2025)
von: Li, Minghan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
UniAIDet: A Unified and Universal Benchmark for AI-Generated Image Content Detection and Localization
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025) -
UniChange: Unifying Change Detection with Multimodal Large Language Model
von: Zhang, Xu, et al.
Veröffentlicht: (2025) -
Uni-SMART: Universal Science Multimodal Analysis and Research Transformer
von: Cai, Hengxing, et al.
Veröffentlicht: (2024) -
Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
von: Qin, Luozheng, et al.
Veröffentlicht: (2025) -
UniAPO: Unified Multimodal Automated Prompt Optimization
von: Zhu, Qipeng, et al.
Veröffentlicht: (2025)