4th PVUW MeViS 3rd Place Report: Sa2VA
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Haobo, Zhang, Tao, Li, Xiangtai, Qi, Lu, Huang, Zilong, Xu, Shilin, Feng, Jiashi, Yang, Ming-Hsuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track
von: Gong, Dengxian, et al.
Veröffentlicht: (2026)
von: Gong, Dengxian, et al.
Veröffentlicht: (2026)
3rd Place of MeViS-Audio Track of the 5th PVUW: VIRST-Audio
von: Hong, Jihwan, et al.
Veröffentlicht: (2026)
von: Hong, Jihwan, et al.
Veröffentlicht: (2026)
2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA
von: Wang, Zhiyu, et al.
Veröffentlicht: (2026)
von: Wang, Zhiyu, et al.
Veröffentlicht: (2026)
APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track
von: Miao, Deshui, et al.
Veröffentlicht: (2026)
von: Miao, Deshui, et al.
Veröffentlicht: (2026)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method
von: Miao, Deshui, et al.
Veröffentlicht: (2026)
von: Miao, Deshui, et al.
Veröffentlicht: (2026)
3rd Place Solution for MeViS Track in CVPR 2024 PVUW workshop: Motion Expression guided Video Segmentation
von: Pan, Feiyu, et al.
Veröffentlicht: (2024)
von: Pan, Feiyu, et al.
Veröffentlicht: (2024)
ReferDINO-Plus: 2nd Solution for 4th PVUW MeViS Challenge at CVPR 2025
von: Liang, Tianming, et al.
Veröffentlicht: (2025)
von: Liang, Tianming, et al.
Veröffentlicht: (2025)
1st Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation
von: Fang, Hao, et al.
Veröffentlicht: (2025)
von: Fang, Hao, et al.
Veröffentlicht: (2025)
The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
von: Niu, Quanzhu, et al.
Veröffentlicht: (2025)
von: Niu, Quanzhu, et al.
Veröffentlicht: (2025)
2nd Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation
von: Cao, Bin, et al.
Veröffentlicht: (2024)
von: Cao, Bin, et al.
Veröffentlicht: (2024)
FVOS for MOSE Track of 4th PVUW Challenge: 3rd Place Solution
von: Wang, Mengjiao, et al.
Veröffentlicht: (2025)
von: Wang, Mengjiao, et al.
Veröffentlicht: (2025)
The 1st Winner for 5th PVUW MeViS-Text Challenge: Strong MLLMs Meet SAM3 for Referring Video Object Segmentation
von: He, Xusheng, et al.
Veröffentlicht: (2026)
von: He, Xusheng, et al.
Veröffentlicht: (2026)
LLAVADI: What Matters For Multimodal Large Language Models Distillation
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
OAMVOS:2nd Report for 5th PVUW MOSE Track
von: Miao, Deshui, et al.
Veröffentlicht: (2026)
von: Miao, Deshui, et al.
Veröffentlicht: (2026)
3rd Place Solution for PVUW Challenge 2024: Video Panoptic Segmentation
von: Wu, Ruipu, et al.
Veröffentlicht: (2024)
von: Wu, Ruipu, et al.
Veröffentlicht: (2024)
MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
von: Ding, Henghui, et al.
Veröffentlicht: (2025)
von: Ding, Henghui, et al.
Veröffentlicht: (2025)
Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding
von: Liu, Chang, et al.
Veröffentlicht: (2026)
von: Liu, Chang, et al.
Veröffentlicht: (2026)
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
von: Li, Xiangtai, et al.
Veröffentlicht: (2025)
von: Li, Xiangtai, et al.
Veröffentlicht: (2025)
PanopticPartFormer++: A Unified and Decoupled View for Panoptic Part Segmentation
von: Li, Xiangtai, et al.
Veröffentlicht: (2023)
von: Li, Xiangtai, et al.
Veröffentlicht: (2023)
3rd Place Solution for MOSE Track in CVPR 2024 PVUW workshop: Complex Video Object Segmentation
von: Liu, Xinyu, et al.
Veröffentlicht: (2024)
von: Liu, Xinyu, et al.
Veröffentlicht: (2024)
Mamba or RWKV: Exploring High-Quality and High-Efficiency Segment Anything Model
von: Yuan, Haobo, et al.
Veröffentlicht: (2024)
von: Yuan, Haobo, et al.
Veröffentlicht: (2024)
PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild
von: Ding, Henghui, et al.
Veröffentlicht: (2025)
von: Ding, Henghui, et al.
Veröffentlicht: (2025)
1st Place Solution for MOSE Track in CVPR 2024 PVUW Workshop: Complex Video Object Segmentation
von: Miao, Deshui, et al.
Veröffentlicht: (2024)
von: Miao, Deshui, et al.
Veröffentlicht: (2024)
Re-Prompting SAM 3 via Object Retrieval: 3rd of the 5th PVUW MOSE Track
von: Gao, Mingqi, et al.
Veröffentlicht: (2026)
von: Gao, Mingqi, et al.
Veröffentlicht: (2026)
MASSeg : 2nd Technical Report for 4th PVUW MOSE Track
von: Cao, Xuqiang, et al.
Veröffentlicht: (2025)
von: Cao, Xuqiang, et al.
Veröffentlicht: (2025)
RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
von: Lei, Weixian, et al.
Veröffentlicht: (2025)
von: Lei, Weixian, et al.
Veröffentlicht: (2025)
Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
Exploring Plain ViT Reconstruction for Multi-class Unsupervised Anomaly Detection
von: Zhang, Jiangning, et al.
Veröffentlicht: (2023)
von: Zhang, Jiangning, et al.
Veröffentlicht: (2023)
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
SuperCLIP: CLIP with Simple Classification Supervision
von: Zhao, Weiheng, et al.
Veröffentlicht: (2025)
von: Zhao, Weiheng, et al.
Veröffentlicht: (2025)
Enhancing Sa2VA for Referent Video Object Segmentation: 2nd Solution for 7th LSVOS RVOS Track
von: Hong, Ran, et al.
Veröffentlicht: (2025)
von: Hong, Ran, et al.
Veröffentlicht: (2025)
Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model
von: Huang, Kuan-Chih, et al.
Veröffentlicht: (2024)
von: Huang, Kuan-Chih, et al.
Veröffentlicht: (2024)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
von: Zhang, Tao, et al.
Veröffentlicht: (2024)
von: Zhang, Tao, et al.
Veröffentlicht: (2024)
2nd Place Solution for PVUW Challenge 2024: Video Panoptic Segmentation
von: Wu, Biao, et al.
Veröffentlicht: (2024)
von: Wu, Biao, et al.
Veröffentlicht: (2024)
Video Prediction Transformers without Recurrence or Convolution
von: Tang, Yujin, et al.
Veröffentlicht: (2024)
von: Tang, Yujin, et al.
Veröffentlicht: (2024)
LSVOS Challenge 3rd Place Report: SAM2 and Cutie based VOS
von: Liu, Xinyu, et al.
Veröffentlicht: (2024)
von: Liu, Xinyu, et al.
Veröffentlicht: (2024)
Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference
von: Nekrasov, Alexey, et al.
Veröffentlicht: (2025)
von: Nekrasov, Alexey, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track
von: Gong, Dengxian, et al.
Veröffentlicht: (2026) -
3rd Place of MeViS-Audio Track of the 5th PVUW: VIRST-Audio
von: Hong, Jihwan, et al.
Veröffentlicht: (2026) -
2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA
von: Wang, Zhiyu, et al.
Veröffentlicht: (2026) -
APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track
von: Miao, Deshui, et al.
Veröffentlicht: (2026) -
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)