Improving Interpretable Embeddings for Ad-hoc Video Search with Generative Captions and Multi-word Concept Bank
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Jiaxin, Ngo, Chong-Wah, Chan, Wing-Kwong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpretable Embedding for Ad-hoc Video Search
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
PosMLP-Video: Spatial and Temporal Relative Position Encoding for Efficient Video Recognition
von: Hao, Yanbin, et al.
Veröffentlicht: (2024)
von: Hao, Yanbin, et al.
Veröffentlicht: (2024)
Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search
von: Hu, Fan, et al.
Veröffentlicht: (2025)
von: Hu, Fan, et al.
Veröffentlicht: (2025)
Interpretable Generative Models through Post-hoc Concept Bottlenecks
von: Kulkarni, Akshay, et al.
Veröffentlicht: (2025)
von: Kulkarni, Akshay, et al.
Veröffentlicht: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
PolySmart @ TRECVid 2024 Video Captioning (VTT)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
LLMs-based Augmentation for Domain Adaptation in Long-tailed Food Datasets
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Retrieval Augmented Recipe Generation
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)
Class Agnostic Instance-level Descriptor for Visual Instance Search
von: Sun, Qi-Ying, et al.
Veröffentlicht: (2025)
von: Sun, Qi-Ying, et al.
Veröffentlicht: (2025)
ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
von: Chen, Lin, et al.
Veröffentlicht: (2024)
von: Chen, Lin, et al.
Veröffentlicht: (2024)
FoodLMM: A Versatile Food Assistant using Large Multi-modal Model
von: Yin, Yuehao, et al.
Veröffentlicht: (2023)
von: Yin, Yuehao, et al.
Veröffentlicht: (2023)
RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning
von: Lu, Yifan, et al.
Veröffentlicht: (2023)
von: Lu, Yifan, et al.
Veröffentlicht: (2023)
Grounded Video Caption Generation
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2024)
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2024)
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Vision Harnessing Agent for Open Ad-hoc Segmentation
von: Wang, Zilin, et al.
Veröffentlicht: (2026)
von: Wang, Zilin, et al.
Veröffentlicht: (2026)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
von: Jeon, MinJu, et al.
Veröffentlicht: (2025)
von: Jeon, MinJu, et al.
Veröffentlicht: (2025)
PolySmart and VIREO @ TRECVid 2024 Ad-hoc Video Search
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
von: Wu, Xiongwei, et al.
Veröffentlicht: (2024)
von: Wu, Xiongwei, et al.
Veröffentlicht: (2024)
Multimodal Arabic Captioning with Interpretable Visual Concept Integration
von: Elchafei, Passant, et al.
Veröffentlicht: (2025)
von: Elchafei, Passant, et al.
Veröffentlicht: (2025)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
von: Zheng, Guangcong, et al.
Veröffentlicht: (2025)
von: Zheng, Guangcong, et al.
Veröffentlicht: (2025)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
TweedieMix: Improving Multi-Concept Fusion for Diffusion-based Image/Video Generation
von: Kwon, Gihyun, et al.
Veröffentlicht: (2024)
von: Kwon, Gihyun, et al.
Veröffentlicht: (2024)
Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
von: Ding, Ning, et al.
Veröffentlicht: (2025)
von: Ding, Ning, et al.
Veröffentlicht: (2025)
Concepts in Motion: Temporal Concept Bottleneck Model for Interpretable Video Classification
von: Knab, Patrick, et al.
Veröffentlicht: (2025)
von: Knab, Patrick, et al.
Veröffentlicht: (2025)
Poster: Reliable 3D Reconstruction for Ad-hoc Edge Implementations
von: Absur, Md Nurul, et al.
Veröffentlicht: (2024)
von: Absur, Md Nurul, et al.
Veröffentlicht: (2024)
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs
von: Yang, Zhantao, et al.
Veröffentlicht: (2024)
von: Yang, Zhantao, et al.
Veröffentlicht: (2024)
Improving Text Generation on Images with Synthetic Captions
von: Koh, Jun Young, et al.
Veröffentlicht: (2024)
von: Koh, Jun Young, et al.
Veröffentlicht: (2024)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
von: Chou, Shih-Han, et al.
Veröffentlicht: (2023)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2023)
CookingDiffusion: Cooking Procedural Image Generation with Stable Diffusion
von: Wang, Yuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuan, et al.
Veröffentlicht: (2025)
QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
von: Wang, Jiale, et al.
Veröffentlicht: (2026)
von: Wang, Jiale, et al.
Veröffentlicht: (2026)
Live Video Captioning
von: Blanco-Fernández, Eduardo, et al.
Veröffentlicht: (2024)
von: Blanco-Fernández, Eduardo, et al.
Veröffentlicht: (2024)
Benchmarking and Improving Detail Image Caption
von: Dong, Hongyuan, et al.
Veröffentlicht: (2024)
von: Dong, Hongyuan, et al.
Veröffentlicht: (2024)
Guided Attention for Interpretable Motion Captioning
von: Radouane, Karim, et al.
Veröffentlicht: (2023)
von: Radouane, Karim, et al.
Veröffentlicht: (2023)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
von: Yao, Linli, et al.
Veröffentlicht: (2026)
von: Yao, Linli, et al.
Veröffentlicht: (2026)
In-hoc Concept Representations to Regularise Deep Learning in Medical Imaging
von: Corbetta, Valentina, et al.
Veröffentlicht: (2025)
von: Corbetta, Valentina, et al.
Veröffentlicht: (2025)
Open Ad-hoc Categorization with Contextualized Feature Learning
von: Wang, Zilin, et al.
Veröffentlicht: (2025)
von: Wang, Zilin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Interpretable Embedding for Ad-hoc Video Search
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024) -
PosMLP-Video: Spatial and Temporal Relative Position Encoding for Efficient Video Recognition
von: Hao, Yanbin, et al.
Veröffentlicht: (2024) -
Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search
von: Hu, Fan, et al.
Veröffentlicht: (2025) -
Interpretable Generative Models through Post-hoc Concept Bottlenecks
von: Kulkarni, Akshay, et al.
Veröffentlicht: (2025) -
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)