SAM3-LiteText: An Anatomical Study of the SAM3 Text Encoder for Efficient Vision-Language Segmentation
Fuente:
arXiv
Guardado en:
| Autores principales: | Zeng, Chengxi, Jiang, Yuxuan, Gao, Ge, Wang, Shuai, Danier, Duolikun, Zhu, Bin, Rudinac, Stevan, Bull, David, Zhang, Fan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EfficientSAM3: Progressive Hierarchical Distillation for Video Concept Segmentation from SAM1, 2, and 3
por: Zeng, Chengxi, et al.
Publicado: (2025)
por: Zeng, Chengxi, et al.
Publicado: (2025)
Enhancing Deformable Convolution based Video Frame Interpolation with Coarse-to-fine 3D CNN
por: Danier, Duolikun, et al.
Publicado: (2022)
por: Danier, Duolikun, et al.
Publicado: (2022)
BVI-VFI: A Video Quality Database for Video Frame Interpolation
por: Danier, Duolikun, et al.
Publicado: (2022)
por: Danier, Duolikun, et al.
Publicado: (2022)
ST-MFNet: A Spatio-Temporal Multi-Flow Network for Frame Interpolation
por: Danier, Duolikun, et al.
Publicado: (2021)
por: Danier, Duolikun, et al.
Publicado: (2021)
LDMVFI: Video Frame Interpolation with Latent Diffusion Models
por: Danier, Duolikun, et al.
Publicado: (2023)
por: Danier, Duolikun, et al.
Publicado: (2023)
A Subjective Quality Study for Video Frame Interpolation
por: Danier, Duolikun, et al.
Publicado: (2022)
por: Danier, Duolikun, et al.
Publicado: (2022)
RankDVQA: Deep VQA based on Ranking-inspired Hybrid Training
por: Feng, Chen, et al.
Publicado: (2022)
por: Feng, Chen, et al.
Publicado: (2022)
GFix: Perceptually Enhanced Gaussian Splatting Video Compression
por: Teng, Siyue, et al.
Publicado: (2025)
por: Teng, Siyue, et al.
Publicado: (2025)
Agglomerating Large Vision Encoders via Distillation for VFSS Segmentation
por: Zeng, Chengxi, et al.
Publicado: (2025)
por: Zeng, Chengxi, et al.
Publicado: (2025)
EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
por: Zhang, Yuxuan, et al.
Publicado: (2024)
por: Zhang, Yuxuan, et al.
Publicado: (2024)
MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors
por: Du, Zhipeng, et al.
Publicado: (2025)
por: Du, Zhipeng, et al.
Publicado: (2025)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
por: Huang, Jia-Hong, et al.
Publicado: (2024)
por: Huang, Jia-Hong, et al.
Publicado: (2024)
TextSAM-EUS: Text Prompt Learning for SAM to Accurately Segment Pancreatic Tumor in Endoscopic Ultrasound
por: Spiegler, Pascal, et al.
Publicado: (2025)
por: Spiegler, Pascal, et al.
Publicado: (2025)
SAM-PTx: Text-Guided Fine-Tuning of SAM with Parameter-Efficient, Parallel-Text Adapters
por: Jalilian, Shayan, et al.
Publicado: (2025)
por: Jalilian, Shayan, et al.
Publicado: (2025)
MVAD: A Multiple Visual Artifact Detector for Video Streaming
por: Feng, Chen, et al.
Publicado: (2024)
por: Feng, Chen, et al.
Publicado: (2024)
BVI-Artefact: An Artefact Detection Benchmark Dataset for Streamed Videos
por: Feng, Chen, et al.
Publicado: (2023)
por: Feng, Chen, et al.
Publicado: (2023)
Lite-SAM Is Actually What You Need for Segment Everything
por: Fu, Jianhai, et al.
Publicado: (2024)
por: Fu, Jianhai, et al.
Publicado: (2024)
Ref-SAM3D: Bridging SAM3D with Text for Reference 3D Reconstruction
por: Zhou, Yun, et al.
Publicado: (2025)
por: Zhou, Yun, et al.
Publicado: (2025)
SAM Fewshot Finetuning for Anatomical Segmentation in Medical Images
por: Xie, Weiyi, et al.
Publicado: (2024)
por: Xie, Weiyi, et al.
Publicado: (2024)
CS3: Cascade SAM for Sperm Segmentation
por: Shi, Yi, et al.
Publicado: (2024)
por: Shi, Yi, et al.
Publicado: (2024)
Hi-SAM: Marrying Segment Anything Model for Hierarchical Text Segmentation
por: Ye, Maoyuan, et al.
Publicado: (2024)
por: Ye, Maoyuan, et al.
Publicado: (2024)
RefSAM3D: Adapting SAM with Cross-modal Reference for 3D Medical Image Segmentation
por: Gao, Xiang, et al.
Publicado: (2024)
por: Gao, Xiang, et al.
Publicado: (2024)
SAM & SAM 2 in 3D Slicer: SegmentWithSAM Extension for Annotating Medical Images
por: Yildiz, Zafer, et al.
Publicado: (2024)
por: Yildiz, Zafer, et al.
Publicado: (2024)
3DTeethSAM: Taming SAM2 for 3D Teeth Segmentation
por: Lu, Zhiguo, et al.
Publicado: (2025)
por: Lu, Zhiguo, et al.
Publicado: (2025)
RMT-BVQA: Recurrent Memory Transformer-based Blind Video Quality Assessment for Enhanced Video Content
por: Peng, Tianhao, et al.
Publicado: (2024)
por: Peng, Tianhao, et al.
Publicado: (2024)
Full-reference Video Quality Assessment for User Generated Content Transcoding
por: Qi, Zihao, et al.
Publicado: (2023)
por: Qi, Zihao, et al.
Publicado: (2023)
Enhancing HDR Video Compression through CNN-based Effective Bit Depth Adaptation
por: Feng, Chen, et al.
Publicado: (2022)
por: Feng, Chen, et al.
Publicado: (2022)
RankDVQA-mini: Knowledge Distillation-Driven Deep Video Quality Assessment
por: Feng, Chen, et al.
Publicado: (2023)
por: Feng, Chen, et al.
Publicado: (2023)
View-Consistent Diffusion Representations for 3D-Consistent Video Generation
por: Danier, Duolikun, et al.
Publicado: (2025)
por: Danier, Duolikun, et al.
Publicado: (2025)
AutoProSAM: Automated Prompting SAM for 3D Multi-Organ Segmentation
por: Li, Chengyin, et al.
Publicado: (2023)
por: Li, Chengyin, et al.
Publicado: (2023)
TP-DRSeg: Improving Diabetic Retinopathy Lesion Segmentation with Explicit Text-Prompts Assisted SAM
por: Li, Wenxue, et al.
Publicado: (2024)
por: Li, Wenxue, et al.
Publicado: (2024)
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
por: Zhang, Xike, et al.
Publicado: (2026)
por: Zhang, Xike, et al.
Publicado: (2026)
A Novel Evaluation Framework for Image2Text Generation
por: Huang, Jia-Hong, et al.
Publicado: (2024)
por: Huang, Jia-Hong, et al.
Publicado: (2024)
Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation
por: Zhang, Weiming, et al.
Publicado: (2026)
por: Zhang, Weiming, et al.
Publicado: (2026)
Enhancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language Models
por: Zhu, Hongyi, et al.
Publicado: (2024)
por: Zhu, Hongyi, et al.
Publicado: (2024)
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation
por: Cuttano, Claudia, et al.
Publicado: (2024)
por: Cuttano, Claudia, et al.
Publicado: (2024)
SAM 3: Segment Anything with Concepts
por: Carion, Nicolas, et al.
Publicado: (2025)
por: Carion, Nicolas, et al.
Publicado: (2025)
PG-SAM: Prior-Guided SAM with Medical for Multi-organ Segmentation
por: Zhong, Yiheng, et al.
Publicado: (2025)
por: Zhong, Yiheng, et al.
Publicado: (2025)
SAM3-DMS: Decoupled Memory Selection for Multi-target Video Segmentation of SAM3
por: Shen, Ruiqi, et al.
Publicado: (2026)
por: Shen, Ruiqi, et al.
Publicado: (2026)
Comparing SAM 2 and SAM 3 for Zero-Shot Segmentation of 3D Medical Data
por: Chakrabarty, Satrajit, et al.
Publicado: (2025)
por: Chakrabarty, Satrajit, et al.
Publicado: (2025)
Ejemplares similares
-
EfficientSAM3: Progressive Hierarchical Distillation for Video Concept Segmentation from SAM1, 2, and 3
por: Zeng, Chengxi, et al.
Publicado: (2025) -
Enhancing Deformable Convolution based Video Frame Interpolation with Coarse-to-fine 3D CNN
por: Danier, Duolikun, et al.
Publicado: (2022) -
BVI-VFI: A Video Quality Database for Video Frame Interpolation
por: Danier, Duolikun, et al.
Publicado: (2022) -
ST-MFNet: A Spatio-Temporal Multi-Flow Network for Frame Interpolation
por: Danier, Duolikun, et al.
Publicado: (2021) -
LDMVFI: Video Frame Interpolation with Latent Diffusion Models
por: Danier, Duolikun, et al.
Publicado: (2023)