Saved in:
| Main Authors: | Choudhuri, Anwesa, Chowdhary, Girish, Schwing, Alexander G. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2404.03657 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PolypSegTrack: Unified Foundation Model for Colonoscopy Video Analysis
by: Choudhuri, Anwesa, et al.
Published: (2025)
by: Choudhuri, Anwesa, et al.
Published: (2025)
6DGS: Enhanced Direction-Aware Gaussian Splatting for Volumetric Rendering
by: Gao, Zhongpai, et al.
Published: (2024)
by: Gao, Zhongpai, et al.
Published: (2024)
Render-FM: A Foundation Model for Real-time Photorealistic Volumetric Rendering
by: Gao, Zhongpai, et al.
Published: (2025)
by: Gao, Zhongpai, et al.
Published: (2025)
7DGS: Unified Spatial-Temporal-Angular Gaussian Splatting
by: Gao, Zhongpai, et al.
Published: (2025)
by: Gao, Zhongpai, et al.
Published: (2025)
Order-aware Interactive Segmentation
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
by: Fan, Tiehan, et al.
Published: (2024)
by: Fan, Tiehan, et al.
Published: (2024)
OW-Rep: Open World Object Detection with Instance Representation Learning
by: Lee, Sunoh, et al.
Published: (2024)
by: Lee, Sunoh, et al.
Published: (2024)
Studying Classifier(-Free) Guidance From a Classifier-Centric Perspective
by: Zhao, Xiaoming, et al.
Published: (2025)
by: Zhao, Xiaoming, et al.
Published: (2025)
ZeST: an LLM-based Zero-Shot Traversability Navigation for Unknown Environments
by: Gummadi, Shreya, et al.
Published: (2025)
by: Gummadi, Shreya, et al.
Published: (2025)
MetaCropFollow: Few-Shot Adaptation with Meta-Learning for Under-Canopy Navigation
by: Woehrle, Thomas, et al.
Published: (2024)
by: Woehrle, Thomas, et al.
Published: (2024)
MARIS: Marine Open-Vocabulary Instance Segmentation with Geometric Enhancement and Semantic Alignment
by: Li, Bingyu, et al.
Published: (2025)
by: Li, Bingyu, et al.
Published: (2025)
Foveated Instance Segmentation
by: Zeng, Hongyi, et al.
Published: (2025)
by: Zeng, Hongyi, et al.
Published: (2025)
Fed-EC: Bandwidth-Efficient Clustering-Based Federated Learning For Autonomous Visual Robot Navigation
by: Gummadi, Shreya, et al.
Published: (2024)
by: Gummadi, Shreya, et al.
Published: (2024)
Mitigating Open-Vocabulary Caption Hallucinations
by: Ben-Kish, Assaf, et al.
Published: (2023)
by: Ben-Kish, Assaf, et al.
Published: (2023)
OE3DIS: Open-Ended 3D Point Cloud Instance Segmentation
by: Nguyen, Phuc D. A., et al.
Published: (2024)
by: Nguyen, Phuc D. A., et al.
Published: (2024)
Unsupervised Instance Segmentation with Superpixels
by: Hoang, Cuong Manh
Published: (2025)
by: Hoang, Cuong Manh
Published: (2025)
Open-World Object Counting in Videos
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
OV-MAP : Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
by: Kim, Juno, et al.
Published: (2025)
by: Kim, Juno, et al.
Published: (2025)
Video-SwinUNet: Spatio-temporal Deep Learning Framework for VFSS Instance Segmentation
by: Zeng, Chengxi, et al.
Published: (2023)
by: Zeng, Chengxi, et al.
Published: (2023)
Accurate and Fast Compressed Video Captioning
by: Shen, Yaojie, et al.
Published: (2023)
by: Shen, Yaojie, et al.
Published: (2023)
Multimedia Generative Script Learning for Task Planning
by: Wang, Qingyun, et al.
Published: (2022)
by: Wang, Qingyun, et al.
Published: (2022)
Generalized Class Discovery in Instance Segmentation
by: Hoang, Cuong Manh, et al.
Published: (2025)
by: Hoang, Cuong Manh, et al.
Published: (2025)
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
by: Truong, Quang-Trung, et al.
Published: (2025)
by: Truong, Quang-Trung, et al.
Published: (2025)
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
by: Bhalgat, Yash, et al.
Published: (2024)
by: Bhalgat, Yash, et al.
Published: (2024)
WsiCaption: Multiple Instance Generation of Pathology Reports for Gigapixel Whole-Slide Images
by: Chen, Pingyi, et al.
Published: (2023)
by: Chen, Pingyi, et al.
Published: (2023)
Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with Diffusion
by: Vu, Tuan-Anh, et al.
Published: (2023)
by: Vu, Tuan-Anh, et al.
Published: (2023)
Unveiling the Invisible: Captioning Videos with Metaphors
by: Kalarani, Abisek Rajakumar, et al.
Published: (2024)
by: Kalarani, Abisek Rajakumar, et al.
Published: (2024)
Describe Anything: Detailed Localized Image and Video Captioning
by: Lian, Long, et al.
Published: (2025)
by: Lian, Long, et al.
Published: (2025)
Towards Fine-Grained Human Motion Video Captioning
by: Song, Guorui, et al.
Published: (2025)
by: Song, Guorui, et al.
Published: (2025)
Open-world Instance Segmentation: Top-down Learning with Bottom-up Supervision
by: Kalluri, Tarun, et al.
Published: (2023)
by: Kalluri, Tarun, et al.
Published: (2023)
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
by: Fiastre, Gabriel, et al.
Published: (2025)
by: Fiastre, Gabriel, et al.
Published: (2025)
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
by: Kim, Ye-Chan, et al.
Published: (2026)
by: Kim, Ye-Chan, et al.
Published: (2026)
Graph Relation Distillation for Efficient Biomedical Instance Segmentation
by: Liu, Xiaoyu, et al.
Published: (2024)
by: Liu, Xiaoyu, et al.
Published: (2024)
Instance-wise Uncertainty for Class Imbalance in Semantic Segmentation
by: Almeida, Luís, et al.
Published: (2024)
by: Almeida, Luís, et al.
Published: (2024)
ProMerge: Prompt and Merge for Unsupervised Instance Segmentation
by: Li, Dylan, et al.
Published: (2024)
by: Li, Dylan, et al.
Published: (2024)
SAM for Robust Mitochondria Instance Segmentation in Fluorescence Microscopy
by: Jadhav, Suyog, et al.
Published: (2026)
by: Jadhav, Suyog, et al.
Published: (2026)
Dense Video Captioning using Graph-based Sentence Summarization
by: Zhang, Zhiwang, et al.
Published: (2025)
by: Zhang, Zhiwang, et al.
Published: (2025)
MAMS: Model-Agnostic Module Selection Framework for Video Captioning
by: Lee, Sangho, et al.
Published: (2025)
by: Lee, Sangho, et al.
Published: (2025)
Controllable Hybrid Captioner for Improved Long-form Video Understanding
by: Sasse, Kuleen, et al.
Published: (2025)
by: Sasse, Kuleen, et al.
Published: (2025)
Similar Items
-
PolypSegTrack: Unified Foundation Model for Colonoscopy Video Analysis
by: Choudhuri, Anwesa, et al.
Published: (2025) -
6DGS: Enhanced Direction-Aware Gaussian Splatting for Volumetric Rendering
by: Gao, Zhongpai, et al.
Published: (2024) -
Render-FM: A Foundation Model for Real-time Photorealistic Volumetric Rendering
by: Gao, Zhongpai, et al.
Published: (2025) -
7DGS: Unified Spatial-Temporal-Angular Gaussian Splatting
by: Gao, Zhongpai, et al.
Published: (2025) -
Order-aware Interactive Segmentation
by: Wang, Bin, et al.
Published: (2024)