MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Truong, Quang-Trung, Wong, Yuk-Kwan, Dang, Vo Hoang Kim Tuyen, Gotama, Rinaldi, Nguyen, Duc Thanh, Yeung, Sai-Kit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AUTV: Creating Underwater Video Datasets with Pixel-wise Annotations
by: Truong, Quang Trung, et al.
Published: (2025)
by: Truong, Quang Trung, et al.
Published: (2025)
KAN-Based Fusion of Dual-Domain for Audio-Driven Facial Landmarks Generation
by: Vo-Thanh, Hoang-Son, et al.
Published: (2024)
by: Vo-Thanh, Hoang-Son, et al.
Published: (2024)
A Subjective Quality Evaluation of 3D Mesh with Dynamic Level of Detail in Virtual Reality
by: Nguyen, Duc, et al.
Published: (2024)
by: Nguyen, Duc, et al.
Published: (2024)
Contestable Multi-Agent Debate with Arena-based Argumentative Computation for Multimedia Verification
by: Nguyen, Truong Thanh Hung, et al.
Published: (2026)
by: Nguyen, Truong Thanh Hung, et al.
Published: (2026)
ORCA: Object Recognition and Comprehension for Archiving Marine Species
by: Wong, Yuk-Kwan, et al.
Published: (2025)
by: Wong, Yuk-Kwan, et al.
Published: (2025)
FedMAC: Tackling Partial-Modality Missing in Federated Learning with Cross-Modal Aggregation and Contrastive Regularization
by: Nguyen, Manh Duong, et al.
Published: (2024)
by: Nguyen, Manh Duong, et al.
Published: (2024)
Fact-Checking at Scale: Multimodal AI for Authenticity and Context Verification in Online Media
by: Phan, Van-Hoang, et al.
Published: (2025)
by: Phan, Van-Hoang, et al.
Published: (2025)
Self-supervised Video Object Segmentation with Distillation Learning of Deformable Attention
by: Truong, Quang-Trung, et al.
Published: (2024)
by: Truong, Quang-Trung, et al.
Published: (2024)
Subjective Quality Assessment of Dynamic 3D Meshes in Virtual Reality Environment
by: Nguyen, Duc V., et al.
Published: (2026)
by: Nguyen, Duc V., et al.
Published: (2026)
ProMSC-MIS: Prompt-based Multimodal Semantic Communication for Multi-Spectral Image Segmentation
by: Zhang, Haoshuo, et al.
Published: (2025)
by: Zhang, Haoshuo, et al.
Published: (2025)
Hallucination Localization in Video Captioning
by: Nakada, Shota, et al.
Published: (2025)
by: Nakada, Shota, et al.
Published: (2025)
Bi-modal Prediction and Transformation Coding for Compressing Complex Human Dynamics
by: Hoang, Huong, et al.
Published: (2025)
by: Hoang, Huong, et al.
Published: (2025)
Short-Form Video Viewing Behavior Analysis and Multi-Step Viewing Time Prediction
by: Yen, Vu Thi Hai, et al.
Published: (2026)
by: Yen, Vu Thi Hai, et al.
Published: (2026)
CineWild: Balancing Art and Robotics for Ethical Wildlife Documentary Filmmaking
by: Pueyo, Pablo, et al.
Published: (2025)
by: Pueyo, Pablo, et al.
Published: (2025)
Cap2Sum: Learning to Summarize Videos by Generating Captions
by: Zhao, Cairong, et al.
Published: (2024)
by: Zhao, Cairong, et al.
Published: (2024)
SimInterview: Transforming Business Education through Large Language Model-Based Simulated Multilingual Interview Training System
by: Nguyen, Truong Thanh Hung, et al.
Published: (2025)
by: Nguyen, Truong Thanh Hung, et al.
Published: (2025)
Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health Understanding
by: Zhou, Zhiyuan, et al.
Published: (2026)
by: Zhou, Zhiyuan, et al.
Published: (2026)
Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions
by: Shi, Junchang, et al.
Published: (2025)
by: Shi, Junchang, et al.
Published: (2025)
NewsCaption: Named-Entity aware Captioning for Out-of-Context Media
by: Singh, Anurag, et al.
Published: (2024)
by: Singh, Anurag, et al.
Published: (2024)
KeyNode-Driven Geometry Coding for Real-World Scanned Human Dynamic Mesh Compression
by: Hoang, Huong, et al.
Published: (2025)
by: Hoang, Huong, et al.
Published: (2025)
Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus Adaptation
by: Shi, Zhaofeng, et al.
Published: (2025)
by: Shi, Zhaofeng, et al.
Published: (2025)
Modeling the Impacts of Swipe Delay on User Quality of Experience in Short Video Streaming
by: Nguyen, Duc V., et al.
Published: (2026)
by: Nguyen, Duc V., et al.
Published: (2026)
An Experimental Study of Low-Latency Video Streaming over 5G
by: Khan, Imran, et al.
Published: (2024)
by: Khan, Imran, et al.
Published: (2024)
Integrated Semantic and Temporal Alignment for Interactive Video Retrieval
by: Luu, Thanh-Danh, et al.
Published: (2025)
by: Luu, Thanh-Danh, et al.
Published: (2025)
Boosting Temporal Sentence Grounding via Causal Inference
by: Tang, Kefan, et al.
Published: (2025)
by: Tang, Kefan, et al.
Published: (2025)
Music Grounding by Short Video
by: Xin, Zijie, et al.
Published: (2024)
by: Xin, Zijie, et al.
Published: (2024)
Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction
by: Nguyen, Cam-Van Thi, et al.
Published: (2023)
by: Nguyen, Cam-Van Thi, et al.
Published: (2023)
E-FreeM2: Efficient Training-Free Multi-Scale and Cross-Modal News Verification via MLLMs
by: Phan, Van-Hoang, et al.
Published: (2025)
by: Phan, Van-Hoang, et al.
Published: (2025)
EDGE-Shield: Efficient Denoising-staGE Shield for Violative Content Filtering via Scalable Reference-Based Matching
by: Taniguchi, Takara, et al.
Published: (2026)
by: Taniguchi, Takara, et al.
Published: (2026)
OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
by: Tran, Quang-Linh, et al.
Published: (2025)
by: Tran, Quang-Linh, et al.
Published: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
by: Nguyen, Hieu Minh, et al.
Published: (2025)
by: Nguyen, Hieu Minh, et al.
Published: (2025)
Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
by: Yang, Danni, et al.
Published: (2024)
by: Yang, Danni, et al.
Published: (2024)
PLayerTV: Advanced Player Tracking and Identification for Automatic Soccer Highlight Clips
by: Solberg, Håkon Maric, et al.
Published: (2024)
by: Solberg, Håkon Maric, et al.
Published: (2024)
How2Compress: Scalable and Efficient Edge Video Analytics via Adaptive Granular Video Compression
by: Wu, Yuheng, et al.
Published: (2025)
by: Wu, Yuheng, et al.
Published: (2025)
Federated Prompt-Tuning with Heterogeneous and Incomplete Multimodal Client Data
by: Phung, Thu Hang, et al.
Published: (2026)
by: Phung, Thu Hang, et al.
Published: (2026)
Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection
by: Cheung, Tsun-Hin, et al.
Published: (2024)
by: Cheung, Tsun-Hin, et al.
Published: (2024)
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding
by: Fang, Pengcheng, et al.
Published: (2026)
by: Fang, Pengcheng, et al.
Published: (2026)
Will It Go Viral? Grounding Micro-Video Popularity Prediction on the Open Web
by: Heo, Ryang, et al.
Published: (2026)
by: Heo, Ryang, et al.
Published: (2026)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
by: Bai, Yatong, et al.
Published: (2023)
by: Bai, Yatong, et al.
Published: (2023)
Similar Items
-
AUTV: Creating Underwater Video Datasets with Pixel-wise Annotations
by: Truong, Quang Trung, et al.
Published: (2025) -
KAN-Based Fusion of Dual-Domain for Audio-Driven Facial Landmarks Generation
by: Vo-Thanh, Hoang-Son, et al.
Published: (2024) -
A Subjective Quality Evaluation of 3D Mesh with Dynamic Level of Detail in Virtual Reality
by: Nguyen, Duc, et al.
Published: (2024) -
Contestable Multi-Agent Debate with Arena-based Argumentative Computation for Multimedia Verification
by: Nguyen, Truong Thanh Hung, et al.
Published: (2026) -
ORCA: Object Recognition and Comprehension for Archiving Marine Species
by: Wong, Yuk-Kwan, et al.
Published: (2025)