VideoSAM: A Large Vision Foundation Model for High-Speed Video Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Maduabuchi, Chika, Jossou, Ericmoore, Bucci, Matteo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MSEG-VCUQ: Multimodal SEGmentation with Enhanced Vision Foundation Models, Convolutional Neural Networks, and Uncertainty Quantification for High-Speed Video Phase Detection Data
by: Maduabuchi, Chika, et al.
Published: (2024)
by: Maduabuchi, Chika, et al.
Published: (2024)
Event-Driven Video Generation
by: Maduabuchi, Chika
Published: (2026)
by: Maduabuchi, Chika
Published: (2026)
VideoSAM: Open-World Video Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
Entropy-Controlled Flow Matching
by: Maduabuchi, Chika
Published: (2026)
by: Maduabuchi, Chika
Published: (2026)
Corruption-Aware Training of Latent Video Diffusion Models for Robust Text-to-Video Generation
by: Maduabuchi, Chika, et al.
Published: (2025)
by: Maduabuchi, Chika, et al.
Published: (2025)
Temporal Pair Consistency for Variance-Reduced Flow Matching
by: Maduabuchi, Chika, et al.
Published: (2026)
by: Maduabuchi, Chika, et al.
Published: (2026)
Interlaced dynamic XCT reconstruction with spatio-temporal implicit neural representations
by: Boulanger, Mathias, et al.
Published: (2025)
by: Boulanger, Mathias, et al.
Published: (2025)
SAM 2: Segment Anything in Images and Videos
by: Ravi, Nikhila, et al.
Published: (2024)
by: Ravi, Nikhila, et al.
Published: (2024)
SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding
by: Wang, Haoxiang, et al.
Published: (2023)
by: Wang, Haoxiang, et al.
Published: (2023)
VMDT: Decoding the Trustworthiness of Video Foundation Models
by: Potter, Yujin, et al.
Published: (2025)
by: Potter, Yujin, et al.
Published: (2025)
ConformalSAM: Unlocking the Potential of Foundational Segmentation Models in Semi-Supervised Semantic Segmentation with Conformal Prediction
by: Chen, Danhui, et al.
Published: (2025)
by: Chen, Danhui, et al.
Published: (2025)
Inferring Dynamic Physical Properties from Video Foundation Models
by: Zhan, Guanqi, et al.
Published: (2025)
by: Zhan, Guanqi, et al.
Published: (2025)
Do Vision Foundation Models Enhance Domain Generalization in Medical Image Segmentation?
by: Cekmeceli, Kerem, et al.
Published: (2024)
by: Cekmeceli, Kerem, et al.
Published: (2024)
CellVTA: Enhancing Vision Foundation Models for Accurate Cell Segmentation and Classification
by: Yang, Yang, et al.
Published: (2025)
by: Yang, Yang, et al.
Published: (2025)
AM-SAM: Automated Prompting and Mask Calibration for Segment Anything Model
by: Li, Yuchen, et al.
Published: (2024)
by: Li, Yuchen, et al.
Published: (2024)
FluoroSAM: A Language-promptable Foundation Model for Flexible X-ray Image Segmentation
by: Killeen, Benjamin D., et al.
Published: (2024)
by: Killeen, Benjamin D., et al.
Published: (2024)
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
by: Huang, Haojian, et al.
Published: (2025)
by: Huang, Haojian, et al.
Published: (2025)
StructSAM: Structure- and Spectrum-Preserving Token Merging for Segment Anything Models
by: Nguyen, Duy M. H., et al.
Published: (2026)
by: Nguyen, Duy M. H., et al.
Published: (2026)
Enhanced Vehicle Speed Detection Considering Lane Recognition Using Drone Videos in California
by: Naeini, Amirali Ataee, et al.
Published: (2025)
by: Naeini, Amirali Ataee, et al.
Published: (2025)
Stitch Contrast and Segment_Learning a Human Action Segmentation Model Using Trimmed Skeleton Videos
by: Tian, Haitao, et al.
Published: (2024)
by: Tian, Haitao, et al.
Published: (2024)
Atlas-Assisted Segment Anything Model for Fetal Brain MRI (FeTal-SAM)
by: Zeng, Qi, et al.
Published: (2026)
by: Zeng, Qi, et al.
Published: (2026)
PhiNet v2: A Mask-Free Brain-Inspired Vision Foundation Model from Video
by: Yamada, Makoto, et al.
Published: (2025)
by: Yamada, Makoto, et al.
Published: (2025)
PTQ4SAM: Post-Training Quantization for Segment Anything
by: Lv, Chengtao, et al.
Published: (2024)
by: Lv, Chengtao, et al.
Published: (2024)
Lite-SAM Is Actually What You Need for Segment Everything
by: Fu, Jianhai, et al.
Published: (2024)
by: Fu, Jianhai, et al.
Published: (2024)
Decorrelation Speeds Up Vision Transformers
by: Carrigg, Kieran, et al.
Published: (2025)
by: Carrigg, Kieran, et al.
Published: (2025)
Tarsier: Recipes for Training and Evaluating Large Video Description Models
by: Wang, Jiawei, et al.
Published: (2024)
by: Wang, Jiawei, et al.
Published: (2024)
Shotluck Holmes: A Family of Efficient Small-Scale Large Language Vision Models For Video Captioning and Summarization
by: Luo, Richard, et al.
Published: (2024)
by: Luo, Richard, et al.
Published: (2024)
Training Video Foundation Models with NVIDIA NeMo
by: Patel, Zeeshan, et al.
Published: (2025)
by: Patel, Zeeshan, et al.
Published: (2025)
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
by: Zhang, Junjie, et al.
Published: (2024)
by: Zhang, Junjie, et al.
Published: (2024)
SAM2RL: Towards Reinforcement Learning Memory Control in Segment Anything Model 2
by: Adamyan, Alen, et al.
Published: (2025)
by: Adamyan, Alen, et al.
Published: (2025)
Sora Detector: A Unified Hallucination Detection for Large Text-to-Video Models
by: Chu, Zhixuan, et al.
Published: (2024)
by: Chu, Zhixuan, et al.
Published: (2024)
Quickly Tuning Foundation Models for Image Segmentation
by: Das, Breenda, et al.
Published: (2025)
by: Das, Breenda, et al.
Published: (2025)
Vision Foundation Models in Remote Sensing: A Survey
by: Lu, Siqi, et al.
Published: (2024)
by: Lu, Siqi, et al.
Published: (2024)
EchoPrime: A Multi-Video View-Informed Vision-Language Model for Comprehensive Echocardiography Interpretation
by: Vukadinovic, Milos, et al.
Published: (2024)
by: Vukadinovic, Milos, et al.
Published: (2024)
Graph2Video: Leveraging Video Models to Model Dynamic Graph Evolution
by: Liu, Hua, et al.
Published: (2026)
by: Liu, Hua, et al.
Published: (2026)
Post-surgical Endometriosis Segmentation in Laparoscopic Videos
by: Leibetseder, Andreas, et al.
Published: (2025)
by: Leibetseder, Andreas, et al.
Published: (2025)
Video Diffusion Models: A Survey
by: Melnik, Andrew, et al.
Published: (2024)
by: Melnik, Andrew, et al.
Published: (2024)
A Multispectral Automated Transfer Technique (MATT) for machine-driven image labeling utilizing the Segment Anything Model (SAM)
by: Gallagher, James E., et al.
Published: (2024)
by: Gallagher, James E., et al.
Published: (2024)
How to Benchmark Vision Foundation Models for Semantic Segmentation?
by: Kerssies, Tommie, et al.
Published: (2024)
by: Kerssies, Tommie, et al.
Published: (2024)
Hyperspectral Adapter for Semantic Segmentation with Vision Foundation Models
by: Hurtado, Juana Valeria, et al.
Published: (2025)
by: Hurtado, Juana Valeria, et al.
Published: (2025)
Similar Items
-
MSEG-VCUQ: Multimodal SEGmentation with Enhanced Vision Foundation Models, Convolutional Neural Networks, and Uncertainty Quantification for High-Speed Video Phase Detection Data
by: Maduabuchi, Chika, et al.
Published: (2024) -
Event-Driven Video Generation
by: Maduabuchi, Chika
Published: (2026) -
VideoSAM: Open-World Video Segmentation
by: Guo, Pinxue, et al.
Published: (2024) -
Entropy-Controlled Flow Matching
by: Maduabuchi, Chika
Published: (2026) -
Corruption-Aware Training of Latent Video Diffusion Models for Robust Text-to-Video Generation
by: Maduabuchi, Chika, et al.
Published: (2025)