On Explaining Visual Captioning with Hybrid Markov Logic Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Shah, Monika, Sarkhel, Somdeb, Venugopal, Deepak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Disentangling Fine-Tuning from Pre-Training in Visual Captioning with Hybrid Markov Logic
by: Shah, Monika, et al.
Published: (2025)
by: Shah, Monika, et al.
Published: (2025)
Analyzing the Sensitivity of Vision Language Models in Visual Question Answering
by: Shah, Monika, et al.
Published: (2025)
by: Shah, Monika, et al.
Published: (2025)
SKALD: Learning-Based Shot Assembly for Coherent Multi-Shot Video Creation
by: Lu, Chen Yi, et al.
Published: (2025)
by: Lu, Chen Yi, et al.
Published: (2025)
Target-Dependent Multimodal Sentiment Analysis Via Employing Visual-to Emotional-Caption Translation Network using Visual-Caption Pairs
by: Pandey, Ananya, et al.
Published: (2024)
by: Pandey, Ananya, et al.
Published: (2024)
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
by: Möller, Lucas, et al.
Published: (2024)
by: Möller, Lucas, et al.
Published: (2024)
Controllable Hybrid Captioner for Improved Long-form Video Understanding
by: Sasse, Kuleen, et al.
Published: (2025)
by: Sasse, Kuleen, et al.
Published: (2025)
Explainable, Multi-modal Wound Infection Classification from Images Augmented with Generated Captions
by: Busaranuvong, Palawat, et al.
Published: (2025)
by: Busaranuvong, Palawat, et al.
Published: (2025)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
by: Tu, Yunbin, et al.
Published: (2024)
by: Tu, Yunbin, et al.
Published: (2024)
Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions
by: Hsieh, Yu-Guan, et al.
Published: (2024)
by: Hsieh, Yu-Guan, et al.
Published: (2024)
Explaining How Visual, Textual and Multimodal Encoders Share Concepts
by: Cornet, Clément, et al.
Published: (2025)
by: Cornet, Clément, et al.
Published: (2025)
Language Models Can Explain Visual Features via Steering
by: Ferrando, Javier, et al.
Published: (2026)
by: Ferrando, Javier, et al.
Published: (2026)
\textit{FocaLogic}: Logic-Based Interpretation of Visual Model Decisions
by: Zhao, Chenchen, et al.
Published: (2026)
by: Zhao, Chenchen, et al.
Published: (2026)
Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering
by: Anaissi, Ali, et al.
Published: (2025)
by: Anaissi, Ali, et al.
Published: (2025)
Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
by: Özdemir, Övgü, et al.
Published: (2024)
by: Özdemir, Övgü, et al.
Published: (2024)
Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings
by: Sharifzadeh, Sahand, et al.
Published: (2024)
by: Sharifzadeh, Sahand, et al.
Published: (2024)
CaptionFool: Universal Image Captioning Model Attacks
by: Parekh, Swapnil
Published: (2026)
by: Parekh, Swapnil
Published: (2026)
Explain Before You Answer: A Survey on Compositional Visual Reasoning
by: Ke, Fucai, et al.
Published: (2025)
by: Ke, Fucai, et al.
Published: (2025)
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
by: Teja, L. D. M. S. Sai, et al.
Published: (2025)
by: Teja, L. D. M. S. Sai, et al.
Published: (2025)
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction
by: Maeda, Koki, et al.
Published: (2024)
by: Maeda, Koki, et al.
Published: (2024)
Controllable Contextualized Image Captioning: Directing the Visual Narrative through User-Defined Highlights
by: Mao, Shunqi, et al.
Published: (2024)
by: Mao, Shunqi, et al.
Published: (2024)
Verifying Relational Explanations: A Probabilistic Approach
by: Magar, Abisha Thapa, et al.
Published: (2024)
by: Magar, Abisha Thapa, et al.
Published: (2024)
Concept Visualization: Explaining the CLIP Multi-modal Embedding Using WordNet
by: Giulivi, Loris, et al.
Published: (2024)
by: Giulivi, Loris, et al.
Published: (2024)
Lesion-Aware Visual-Language Fusion for Automated Image Captioning of Ulcerative Colitis Endoscopic Examinations
by: Escamilla, Alexis Ivan Lopez, et al.
Published: (2025)
by: Escamilla, Alexis Ivan Lopez, et al.
Published: (2025)
Show, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization
by: Zhang, Zhiwang, et al.
Published: (2025)
by: Zhang, Zhiwang, et al.
Published: (2025)
SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
EVC-MF: End-to-end Video Captioning Network with Multi-scale Features
by: Niu, Tian-Zi, et al.
Published: (2024)
by: Niu, Tian-Zi, et al.
Published: (2024)
DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
by: Li, Binbin, et al.
Published: (2025)
by: Li, Binbin, et al.
Published: (2025)
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
by: Matsuda, Kazuki, et al.
Published: (2025)
by: Matsuda, Kazuki, et al.
Published: (2025)
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
by: Lu, Xingyu, et al.
Published: (2026)
by: Lu, Xingyu, et al.
Published: (2026)
SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning
by: Truong, Khang, et al.
Published: (2025)
by: Truong, Khang, et al.
Published: (2025)
Rethinking 3D Dense Caption and Visual Grounding in A Unified Framework through Prompt-based Localization
by: Luo, Yongdong, et al.
Published: (2024)
by: Luo, Yongdong, et al.
Published: (2024)
MobileUNETR: A Lightweight End-To-End Hybrid Vision Transformer For Efficient Medical Image Segmentation
by: Perera, Shehan, et al.
Published: (2024)
by: Perera, Shehan, et al.
Published: (2024)
Caption This, Reason That: VLMs Caught in the Middle
by: Weng, Zihan, et al.
Published: (2025)
by: Weng, Zihan, et al.
Published: (2025)
URECA: Unique Region Caption Anything
by: Lim, Sangbeom, et al.
Published: (2025)
by: Lim, Sangbeom, et al.
Published: (2025)
Image Captioning in news report scenario
by: Liu, Tianrui, et al.
Published: (2024)
by: Liu, Tianrui, et al.
Published: (2024)
Accurate and Fast Compressed Video Captioning
by: Shen, Yaojie, et al.
Published: (2023)
by: Shen, Yaojie, et al.
Published: (2023)
Automated Image Captioning with CNNs and Transformers
by: Cahyono, Joshua Adrian, et al.
Published: (2024)
by: Cahyono, Joshua Adrian, et al.
Published: (2024)
Mitigating Open-Vocabulary Caption Hallucinations
by: Ben-Kish, Assaf, et al.
Published: (2023)
by: Ben-Kish, Assaf, et al.
Published: (2023)
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
by: Kim, Ye-Chan, et al.
Published: (2026)
by: Kim, Ye-Chan, et al.
Published: (2026)
Similar Items
-
Disentangling Fine-Tuning from Pre-Training in Visual Captioning with Hybrid Markov Logic
by: Shah, Monika, et al.
Published: (2025) -
Analyzing the Sensitivity of Vision Language Models in Visual Question Answering
by: Shah, Monika, et al.
Published: (2025) -
SKALD: Learning-Based Shot Assembly for Coherent Multi-Shot Video Creation
by: Lu, Chen Yi, et al.
Published: (2025) -
Target-Dependent Multimodal Sentiment Analysis Via Employing Visual-to Emotional-Caption Translation Network using Visual-Caption Pairs
by: Pandey, Ananya, et al.
Published: (2024) -
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
by: Möller, Lucas, et al.
Published: (2024)