Multi-Modal interpretable automatic video captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Hanna-Asaad, Antoine, Aspandi, Decky, Zaharia, Titus |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Variational Contrastive Learning for Skeleton-based Action Recognition
by: Nguyen, Dang Dinh, et al.
Published: (2026)
by: Nguyen, Dang Dinh, et al.
Published: (2026)
Deep self-supervised learning with visualisation for automatic gesture recognition
by: Allemand, Fabien, et al.
Published: (2024)
by: Allemand, Fabien, et al.
Published: (2024)
Cognitive resilience: Unraveling the proficiency of image-captioning models to interpret masked visual content
by: Du, Zhicheng, et al.
Published: (2024)
by: Du, Zhicheng, et al.
Published: (2024)
Refining Gaussian Splatting: A Volumetric Densification Approach
by: Gafoor, Mohamed Abdul, et al.
Published: (2025)
by: Gafoor, Mohamed Abdul, et al.
Published: (2025)
Object-oriented backdoor attack against image captioning
by: Li, Meiling, et al.
Published: (2024)
by: Li, Meiling, et al.
Published: (2024)
Image captioning for Brazilian Portuguese using GRIT model
by: de Alencar, Rafael Silva, et al.
Published: (2024)
by: de Alencar, Rafael Silva, et al.
Published: (2024)
Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data
by: Zhang, Chenhui, et al.
Published: (2024)
by: Zhang, Chenhui, et al.
Published: (2024)
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
by: Albadarneh, Israa A., et al.
Published: (2025)
by: Albadarneh, Israa A., et al.
Published: (2025)
Evaluating authenticity and quality of image captions via sentiment and semantic analyses
by: Krotov, Aleksei, et al.
Published: (2024)
by: Krotov, Aleksei, et al.
Published: (2024)
Learning text-to-video retrieval from image captioning
by: Ventura, Lucas, et al.
Published: (2024)
by: Ventura, Lucas, et al.
Published: (2024)
MMPB: It's Time for Multi-Modal Personalization
by: Kim, Jaeik, et al.
Published: (2025)
by: Kim, Jaeik, et al.
Published: (2025)
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
by: Kim, Jeonghyeon, et al.
Published: (2025)
by: Kim, Jeonghyeon, et al.
Published: (2025)
Fine-grained length controllable video captioning with ordinal embeddings
by: Nitta, Tomoya, et al.
Published: (2024)
by: Nitta, Tomoya, et al.
Published: (2024)
Learning Progressive Adaptation for Multi-Modal Tracking
by: Wang, He, et al.
Published: (2026)
by: Wang, He, et al.
Published: (2026)
Architecture-Agnostic Modality-Isolated Gated Fusion for Robust Multi-Modal Prostate MRI Segmentation
by: Shu, Yongbo, et al.
Published: (2026)
by: Shu, Yongbo, et al.
Published: (2026)
Sensor-Adaptive Flood Mapping with Pre-trained Multi-Modal Transformers across SAR and Multispectral Modalities
by: Tanaka, Tomohiro, et al.
Published: (2025)
by: Tanaka, Tomohiro, et al.
Published: (2025)
MammothModa: Multi-Modal Large Language Model
by: She, Qi, et al.
Published: (2024)
by: She, Qi, et al.
Published: (2024)
Unveiling Ontological Commitment in Multi-Modal Foundation Models
by: Keser, Mert, et al.
Published: (2024)
by: Keser, Mert, et al.
Published: (2024)
Multi-Prompt with Depth Partitioned Cross-Modal Learning
by: Tian, Yingjie, et al.
Published: (2023)
by: Tian, Yingjie, et al.
Published: (2023)
A Generalized Multi-Modal Fusion Detection Framework
by: Cui, Leichao, et al.
Published: (2023)
by: Cui, Leichao, et al.
Published: (2023)
GRAID: Enhancing Spatial Reasoning of VLMs Through High-Fidelity Data Generation
by: Elmaaroufi, Karim, et al.
Published: (2025)
by: Elmaaroufi, Karim, et al.
Published: (2025)
A Unified Model for Longitudinal Multi-Modal Multi-View Prediction with Missingness
by: Chen, Boqi, et al.
Published: (2024)
by: Chen, Boqi, et al.
Published: (2024)
MMXU: A Multi-Modal and Multi-X-ray Understanding Dataset for Disease Progression
by: Mu, Linjie, et al.
Published: (2025)
by: Mu, Linjie, et al.
Published: (2025)
Multi-Modal Foundation Models for Computational Pathology: A Survey
by: Li, Dong, et al.
Published: (2025)
by: Li, Dong, et al.
Published: (2025)
Decoupling Stability and Plasticity for Multi-Modal Test-Time Adaptation
by: He, Yongbo, et al.
Published: (2026)
by: He, Yongbo, et al.
Published: (2026)
Multi-Modal Character Localization and Extraction for Chinese Text Recognition
by: Li, Qilong, et al.
Published: (2026)
by: Li, Qilong, et al.
Published: (2026)
Can video generation replace cinematographers? Research on the cinematic language of generated video
by: Li, Xiaozhe, et al.
Published: (2024)
by: Li, Xiaozhe, et al.
Published: (2024)
DAE-Fuse: An Adaptive Discriminative Autoencoder for Multi-Modality Image Fusion
by: Guo, Yuchen, et al.
Published: (2024)
by: Guo, Yuchen, et al.
Published: (2024)
3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset
by: Ma, Xinyu, et al.
Published: (2024)
by: Ma, Xinyu, et al.
Published: (2024)
ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and Wisdom
by: Zhou, Jingqi, et al.
Published: (2024)
by: Zhou, Jingqi, et al.
Published: (2024)
Modality-Aware and Shift Mixer for Multi-modal Brain Tumor Segmentation
by: Huang, Zhongzhen, et al.
Published: (2024)
by: Huang, Zhongzhen, et al.
Published: (2024)
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
by: Wang, Jiaze, et al.
Published: (2024)
by: Wang, Jiaze, et al.
Published: (2024)
A Multi-Modal Knowledge-Enhanced Framework for Vessel Trajectory Prediction
by: Yu, Haomin, et al.
Published: (2025)
by: Yu, Haomin, et al.
Published: (2025)
Learning Multi-Modal Mobility Dynamics for Generalized Next Location Recommendation
by: Dai, Junshu, et al.
Published: (2025)
by: Dai, Junshu, et al.
Published: (2025)
Hydra-Bench: A Benchmark for Multi-Modal Leaf Wetness Sensing
by: Liu, Yimeng, et al.
Published: (2025)
by: Liu, Yimeng, et al.
Published: (2025)
A Deep Multi-Modal Method for Patient Wound Healing Assessment
by: Oota, Subba Reddy, et al.
Published: (2026)
by: Oota, Subba Reddy, et al.
Published: (2026)
Revisiting Multi-Modal LLM Evaluation
by: Lu, Jian, et al.
Published: (2024)
by: Lu, Jian, et al.
Published: (2024)
MVIP -- A Dataset and Methods for Application Oriented Multi-View and Multi-Modal Industrial Part Recognition
by: Koch, Paul, et al.
Published: (2025)
by: Koch, Paul, et al.
Published: (2025)
MoSAiC: Multi-Modal Multi-Label Supervision-Aware Contrastive Learning for Remote Sensing
by: Gupta, Debashis, et al.
Published: (2025)
by: Gupta, Debashis, et al.
Published: (2025)
Comparing the Effects of Persistence Barcodes Aggregation and Feature Concatenation on Medical Imaging
by: Ali, Dashti A., et al.
Published: (2025)
by: Ali, Dashti A., et al.
Published: (2025)
Similar Items
-
Variational Contrastive Learning for Skeleton-based Action Recognition
by: Nguyen, Dang Dinh, et al.
Published: (2026) -
Deep self-supervised learning with visualisation for automatic gesture recognition
by: Allemand, Fabien, et al.
Published: (2024) -
Cognitive resilience: Unraveling the proficiency of image-captioning models to interpret masked visual content
by: Du, Zhicheng, et al.
Published: (2024) -
Refining Gaussian Splatting: A Volumetric Densification Approach
by: Gafoor, Mohamed Abdul, et al.
Published: (2025) -
Object-oriented backdoor attack against image captioning
by: Li, Meiling, et al.
Published: (2024)