3rd Place of MeViS-Audio Track of the 5th PVUW: VIRST-Audio
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Jihwan, Do, Jaeyoung |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
4th PVUW MeViS 3rd Place Report: Sa2VA
by: Yuan, Haobo, et al.
Published: (2025)
by: Yuan, Haobo, et al.
Published: (2025)
AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method
by: Miao, Deshui, et al.
Published: (2026)
by: Miao, Deshui, et al.
Published: (2026)
2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA
by: Wang, Zhiyu, et al.
Published: (2026)
by: Wang, Zhiyu, et al.
Published: (2026)
VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation
by: Hong, Jihwan, et al.
Published: (2026)
by: Hong, Jihwan, et al.
Published: (2026)
3rd Place Solution for MeViS Track in CVPR 2024 PVUW workshop: Motion Expression guided Video Segmentation
by: Pan, Feiyu, et al.
Published: (2024)
by: Pan, Feiyu, et al.
Published: (2024)
SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track
by: Gong, Dengxian, et al.
Published: (2026)
by: Gong, Dengxian, et al.
Published: (2026)
2nd Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation
by: Cao, Bin, et al.
Published: (2024)
by: Cao, Bin, et al.
Published: (2024)
1st Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation
by: Gao, Mingqi, et al.
Published: (2024)
by: Gao, Mingqi, et al.
Published: (2024)
ReferDINO-Plus: 2nd Solution for 4th PVUW MeViS Challenge at CVPR 2025
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
FVOS for MOSE Track of 4th PVUW Challenge: 3rd Place Solution
by: Wang, Mengjiao, et al.
Published: (2025)
by: Wang, Mengjiao, et al.
Published: (2025)
The 1st Winner for 5th PVUW MeViS-Text Challenge: Strong MLLMs Meet SAM3 for Referring Video Object Segmentation
by: He, Xusheng, et al.
Published: (2026)
by: He, Xusheng, et al.
Published: (2026)
The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track
by: Miao, Deshui, et al.
Published: (2026)
by: Miao, Deshui, et al.
Published: (2026)
Re-Prompting SAM 3 via Object Retrieval: 3rd of the 5th PVUW MOSE Track
by: Gao, Mingqi, et al.
Published: (2026)
by: Gao, Mingqi, et al.
Published: (2026)
3rd Place Solution for PVUW Challenge 2024: Video Panoptic Segmentation
by: Wu, Ruipu, et al.
Published: (2024)
by: Wu, Ruipu, et al.
Published: (2024)
3rd Place Solution for MOSE Track in CVPR 2024 PVUW workshop: Complex Video Object Segmentation
by: Liu, Xinyu, et al.
Published: (2024)
by: Liu, Xinyu, et al.
Published: (2024)
MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
by: Ding, Henghui, et al.
Published: (2025)
by: Ding, Henghui, et al.
Published: (2025)
Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
OAMVOS:2nd Report for 5th PVUW MOSE Track
by: Miao, Deshui, et al.
Published: (2026)
by: Miao, Deshui, et al.
Published: (2026)
Enriched Feature Representation and Motion Prediction Module for MOSEv2 Track of 7th LSVOS Challenge: 3rd Place Solution
by: Lim, Chang Soo, et al.
Published: (2025)
by: Lim, Chang Soo, et al.
Published: (2025)
1st Place Solution for MOSE Track in CVPR 2024 PVUW Workshop: Complex Video Object Segmentation
by: Miao, Deshui, et al.
Published: (2024)
by: Miao, Deshui, et al.
Published: (2024)
2nd Place Solution for MOSE Track in CVPR 2024 PVUW workshop: Complex Video Object Segmentation
by: Xu, Zhensong, et al.
Published: (2024)
by: Xu, Zhensong, et al.
Published: (2024)
MASSeg : 2nd Technical Report for 4th PVUW MOSE Track
by: Cao, Xuqiang, et al.
Published: (2025)
by: Cao, Xuqiang, et al.
Published: (2025)
The Instance-centric Transformer for the RVOS Track of LSVOS Challenge: 3rd Place Solution
by: Cao, Bin, et al.
Published: (2024)
by: Cao, Bin, et al.
Published: (2024)
Advancing Complex Video Object Segmentation via Tracking-Enhanced Prompt: The 1st Winner for 5th PVUW MOSE Challenge
by: Zhang, Jinrong, et al.
Published: (2026)
by: Zhang, Jinrong, et al.
Published: (2026)
2nd Place Solution for PVUW Challenge 2024: Video Panoptic Segmentation
by: Wu, Biao, et al.
Published: (2024)
by: Wu, Biao, et al.
Published: (2024)
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
by: Nguyen, Thai-Binh, et al.
Published: (2025)
by: Nguyen, Thai-Binh, et al.
Published: (2025)
SAMSON: 3rd Place Solution of LSVOS 2025 VOS Challenge
by: Xie, Yujie, et al.
Published: (2025)
by: Xie, Yujie, et al.
Published: (2025)
Exploring and Leveraging Class Vectors for Classifier Editing
by: Kim, Jaeik, et al.
Published: (2025)
by: Kim, Jaeik, et al.
Published: (2025)
NowYouSee Me: Context-Aware Automatic Audio Description
by: Lee, Seon-Ho, et al.
Published: (2024)
by: Lee, Seon-Ho, et al.
Published: (2024)
3rd Place Solution to Large-scale Fine-grained Food Recognition
by: Zhong, Yang, et al.
Published: (2025)
by: Zhong, Yang, et al.
Published: (2025)
3rd Place Solution to ICCV LargeFineFoodAI Retrieval
by: Zhong, Yang, et al.
Published: (2025)
by: Zhong, Yang, et al.
Published: (2025)
STSeg-Complex Video Object Segmentation: The 1st Solution for 4th PVUW MOSE Challenge
by: Song, Kehuan, et al.
Published: (2025)
by: Song, Kehuan, et al.
Published: (2025)
AuViRe: Audio-visual Speech Representation Reconstruction for Deepfake Temporal Localization
by: Koutlis, Christos, et al.
Published: (2025)
by: Koutlis, Christos, et al.
Published: (2025)
MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation
by: Meng, Dechao, et al.
Published: (2025)
by: Meng, Dechao, et al.
Published: (2025)
STNet: Deep Audio-Visual Fusion Network for Robust Speaker Tracking
by: Li, Yidi, et al.
Published: (2024)
by: Li, Yidi, et al.
Published: (2024)
3rd Place Solution for VisDA 2021 Challenge -- Universally Domain Adaptive Image Recognition
by: Liao, Haojin, et al.
Published: (2021)
by: Liao, Haojin, et al.
Published: (2021)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
AudioScenic: Audio-Driven Video Scene Editing
by: Shen, Kaixin, et al.
Published: (2024)
by: Shen, Kaixin, et al.
Published: (2024)
MI-CXR: A Benchmark for Longitudinal Reasoning over Multi-Interval Chest X-rays
by: Cho, Sunghwan Steve, et al.
Published: (2026)
by: Cho, Sunghwan Steve, et al.
Published: (2026)
Similar Items
-
4th PVUW MeViS 3rd Place Report: Sa2VA
by: Yuan, Haobo, et al.
Published: (2025) -
AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method
by: Miao, Deshui, et al.
Published: (2026) -
2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA
by: Wang, Zhiyu, et al.
Published: (2026) -
VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation
by: Hong, Jihwan, et al.
Published: (2026) -
3rd Place Solution for MeViS Track in CVPR 2024 PVUW workshop: Motion Expression guided Video Segmentation
by: Pan, Feiyu, et al.
Published: (2024)