Hyperbolic Learning with Synthetic Captions for Open-World Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Kong, Fanjie, Chen, Yanbei, Cai, Jiarui, Modolo, Davide |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Supervised Multi-Object Tracking with Path Consistency
by: Lu, Zijia, et al.
Published: (2024)
by: Lu, Zijia, et al.
Published: (2024)
Open-World Amodal Appearance Completion
by: Ao, Jiayang, et al.
Published: (2024)
by: Ao, Jiayang, et al.
Published: (2024)
Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
by: Lebailly, Tim, et al.
Published: (2025)
by: Lebailly, Tim, et al.
Published: (2025)
GenDet: Painting Colored Bounding Boxes on Images via Diffusion Model for Object Detection
by: Min, Chen, et al.
Published: (2026)
by: Min, Chen, et al.
Published: (2026)
Visual Reasoning through Tool-supervised Reinforcement Learning
by: Dong, Qihua, et al.
Published: (2026)
by: Dong, Qihua, et al.
Published: (2026)
KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph
by: Jiang, Yanbei, et al.
Published: (2024)
by: Jiang, Yanbei, et al.
Published: (2024)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
by: Liu, Yanqing, et al.
Published: (2024)
by: Liu, Yanqing, et al.
Published: (2024)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
by: Xu, Zhenlin, et al.
Published: (2023)
by: Xu, Zhenlin, et al.
Published: (2023)
Solving Instance Detection from an Open-World Perspective
by: Shen, Qianqian, et al.
Published: (2025)
by: Shen, Qianqian, et al.
Published: (2025)
Hyp-OW: Exploiting Hierarchical Structure Learning with Hyperbolic Distance Enhances Open World Object Detection
by: Doan, Thang, et al.
Published: (2023)
by: Doan, Thang, et al.
Published: (2023)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
by: Lei, Zhenxin, et al.
Published: (2025)
by: Lei, Zhenxin, et al.
Published: (2025)
Improving Text Generation on Images with Synthetic Captions
by: Koh, Jun Young, et al.
Published: (2024)
by: Koh, Jun Young, et al.
Published: (2024)
MultiModal Fine-tuning with Synthetic Captions
by: Enomoto, Shohei, et al.
Published: (2026)
by: Enomoto, Shohei, et al.
Published: (2026)
EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension
by: Li, Jiaxuan, et al.
Published: (2023)
by: Li, Jiaxuan, et al.
Published: (2023)
Hyperbolic Metric Learning for Visual Outlier Detection
by: Gonzalez-Jimenez, Alvaro, et al.
Published: (2024)
by: Gonzalez-Jimenez, Alvaro, et al.
Published: (2024)
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images
by: Lu, Zimao, et al.
Published: (2025)
by: Lu, Zimao, et al.
Published: (2025)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
by: Chen, Yuxiao, et al.
Published: (2026)
by: Chen, Yuxiao, et al.
Published: (2026)
OW-VISCapTor: Abstractors for Open-World Video Instance Segmentation and Captioning
by: Choudhuri, Anwesa, et al.
Published: (2024)
by: Choudhuri, Anwesa, et al.
Published: (2024)
Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption
by: Qin, Luozheng, et al.
Published: (2025)
by: Qin, Luozheng, et al.
Published: (2025)
Hyperbolic Dual Feature Augmentation for Open-Environment
by: Yu, Peilin, et al.
Published: (2025)
by: Yu, Peilin, et al.
Published: (2025)
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
by: Li, Xiangtai, et al.
Published: (2025)
by: Li, Xiangtai, et al.
Published: (2025)
Pixel-Level Change Detection Pseudo-Label Learning for Remote Sensing Change Captioning
by: Liu, Chenyang, et al.
Published: (2023)
by: Liu, Chenyang, et al.
Published: (2023)
Open Set Face Forgery Detection via Dual-Level Evidence Collection
by: Cai, Zhongyi, et al.
Published: (2025)
by: Cai, Zhongyi, et al.
Published: (2025)
OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance Fields
by: Weijler, Lisa, et al.
Published: (2025)
by: Weijler, Lisa, et al.
Published: (2025)
Lidar Panoptic Segmentation in an Open World
by: Chakravarthy, Anirudh S, et al.
Published: (2024)
by: Chakravarthy, Anirudh S, et al.
Published: (2024)
FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
by: Wang, Bingchao, et al.
Published: (2025)
by: Wang, Bingchao, et al.
Published: (2025)
SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
Semi-supervised Open-World Object Detection
by: Mullappilly, Sahal Shaji, et al.
Published: (2024)
by: Mullappilly, Sahal Shaji, et al.
Published: (2024)
Open World Object Detection: A Survey
by: Li, Yiming, et al.
Published: (2024)
by: Li, Yiming, et al.
Published: (2024)
Open-World Dynamic Prompt and Continual Visual Representation Learning
by: Kim, Youngeun, et al.
Published: (2024)
by: Kim, Youngeun, et al.
Published: (2024)
Top-Down Framework for Weakly-supervised Grounded Image Captioning
by: Cai, Chen, et al.
Published: (2023)
by: Cai, Chen, et al.
Published: (2023)
Hawk: Learning to Understand Open-World Video Anomalies
by: Tang, Jiaqi, et al.
Published: (2024)
by: Tang, Jiaqi, et al.
Published: (2024)
Transformer based Multitask Learning for Image Captioning and Object Detection
by: Basak, Debolena, et al.
Published: (2024)
by: Basak, Debolena, et al.
Published: (2024)
Learning Weakly Supervised Audio-Visual Violence Detection in Hyperbolic Space
by: Peng, Xiaogang, et al.
Published: (2023)
by: Peng, Xiaogang, et al.
Published: (2023)
Mitigating Open-Vocabulary Caption Hallucinations
by: Ben-Kish, Assaf, et al.
Published: (2023)
by: Ben-Kish, Assaf, et al.
Published: (2023)
YOLO-UniOW: Efficient Universal Open-World Object Detection
by: Liu, Lihao, et al.
Published: (2024)
by: Liu, Lihao, et al.
Published: (2024)
VIVECaption: A Split Approach to Caption Quality Improvement
by: Ananth, Varun, et al.
Published: (2026)
by: Ananth, Varun, et al.
Published: (2026)
Detecting Localized Deepfakes: How Well Do Synthetic Image Detectors Handle Inpainting?
by: Pandolfini, Serafino, et al.
Published: (2025)
by: Pandolfini, Serafino, et al.
Published: (2025)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
by: Yao, Linli, et al.
Published: (2026)
by: Yao, Linli, et al.
Published: (2026)
Similar Items
-
Self-Supervised Multi-Object Tracking with Path Consistency
by: Lu, Zijia, et al.
Published: (2024) -
Open-World Amodal Appearance Completion
by: Ao, Jiayang, et al.
Published: (2024) -
Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
by: Lebailly, Tim, et al.
Published: (2025) -
GenDet: Painting Colored Bounding Boxes on Images via Diffusion Model for Object Detection
by: Min, Chen, et al.
Published: (2026) -
Visual Reasoning through Tool-supervised Reinforcement Learning
by: Dong, Qihua, et al.
Published: (2026)