Assessing News Thumbnail Representativeness: Counterfactual text can enhance the cross-modal matching ability
Fuente:
arXiv
Salvato in:
| Autori principali: | Yoon, Yejun, Yoon, Seunghyun, Park, Kunwoo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity
di: Jung, Jaeyoon, et al.
Pubblicazione: (2026)
di: Jung, Jaeyoon, et al.
Pubblicazione: (2026)
HerO at AVeriTeC: The Herd of Open Large Language Models for Verifying Real-World Claims
di: Yoon, Yejun, et al.
Pubblicazione: (2024)
di: Yoon, Yejun, et al.
Pubblicazione: (2024)
Team HUMANE at AVeriTeC 2025: HerO 2 for Efficient Fact Verification
di: Yoon, Yejun, et al.
Pubblicazione: (2025)
di: Yoon, Yejun, et al.
Pubblicazione: (2025)
Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion
di: Yoon, Yejun, et al.
Pubblicazione: (2025)
di: Yoon, Yejun, et al.
Pubblicazione: (2025)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
di: Zang, Yuan, et al.
Pubblicazione: (2025)
di: Zang, Yuan, et al.
Pubblicazione: (2025)
MVMR: A New Framework for Evaluating Faithfulness of Video Moment Retrieval against Multiple Distractors
di: Yang, Nakyeong, et al.
Pubblicazione: (2023)
di: Yang, Nakyeong, et al.
Pubblicazione: (2023)
VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration
di: Jung, Jaeyoon, et al.
Pubblicazione: (2026)
di: Jung, Jaeyoon, et al.
Pubblicazione: (2026)
FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
di: Jing, Liqiang, et al.
Pubblicazione: (2025)
di: Jing, Liqiang, et al.
Pubblicazione: (2025)
M4CXR: Exploring Multi-task Potentials of Multi-modal Large Language Models for Chest X-ray Interpretation
di: Park, Jonggwon, et al.
Pubblicazione: (2024)
di: Park, Jonggwon, et al.
Pubblicazione: (2024)
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
di: Yoon, Eunseop, et al.
Pubblicazione: (2025)
di: Yoon, Eunseop, et al.
Pubblicazione: (2025)
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
di: Lee, Saehyung, et al.
Pubblicazione: (2024)
di: Lee, Saehyung, et al.
Pubblicazione: (2024)
PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck
di: Pham, Thang M., et al.
Pubblicazione: (2024)
di: Pham, Thang M., et al.
Pubblicazione: (2024)
$\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
di: Das, Trishanu, et al.
Pubblicazione: (2025)
di: Das, Trishanu, et al.
Pubblicazione: (2025)
Can Impressions of Music be Extracted from Thumbnail Images?
di: Harada, Takashi, et al.
Pubblicazione: (2025)
di: Harada, Takashi, et al.
Pubblicazione: (2025)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
di: Lee, Kang-il, et al.
Pubblicazione: (2024)
di: Lee, Kang-il, et al.
Pubblicazione: (2024)
A Spatio-Temporal Representation Learning as an Alternative to Traditional Glosses in Sign Language Translation and Production
di: Hwang, Eui Jun, et al.
Pubblicazione: (2024)
di: Hwang, Eui Jun, et al.
Pubblicazione: (2024)
Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
di: Park, Junsung, et al.
Pubblicazione: (2025)
di: Park, Junsung, et al.
Pubblicazione: (2025)
Detecting Cultural Differences in News Video Thumbnails via Computational Aesthetics
di: Limpijankit, Marvin, et al.
Pubblicazione: (2025)
di: Limpijankit, Marvin, et al.
Pubblicazione: (2025)
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction
di: Park, Jonggwon, et al.
Pubblicazione: (2025)
di: Park, Jonggwon, et al.
Pubblicazione: (2025)
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability
di: Park, Jonggwon, et al.
Pubblicazione: (2025)
di: Park, Jonggwon, et al.
Pubblicazione: (2025)
What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models
di: Kim, Junho, et al.
Pubblicazione: (2024)
di: Kim, Junho, et al.
Pubblicazione: (2024)
What If the TV Was Off? Examining Counterfactual Reasoning Abilities of Multi-modal Language Models
di: Zhang, Letian, et al.
Pubblicazione: (2023)
di: Zhang, Letian, et al.
Pubblicazione: (2023)
TALL: Thumbnail Layout for Deepfake Video Detection
di: Xu, Yuting, et al.
Pubblicazione: (2023)
di: Xu, Yuting, et al.
Pubblicazione: (2023)
BloomVQA: Assessing Hierarchical Multi-modal Comprehension
di: Gong, Yunye, et al.
Pubblicazione: (2023)
di: Gong, Yunye, et al.
Pubblicazione: (2023)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
di: Yoon, Hee Suk, et al.
Pubblicazione: (2026)
di: Yoon, Hee Suk, et al.
Pubblicazione: (2026)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
di: Lee, Daeun, et al.
Pubblicazione: (2025)
di: Lee, Daeun, et al.
Pubblicazione: (2025)
Multi-modal Knowledge Distillation-based Human Trajectory Forecasting
di: Jeong, Jaewoo, et al.
Pubblicazione: (2025)
di: Jeong, Jaewoo, et al.
Pubblicazione: (2025)
Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering
di: Kim, Jongha, et al.
Pubblicazione: (2026)
di: Kim, Jongha, et al.
Pubblicazione: (2026)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
di: Zhang, Ming, et al.
Pubblicazione: (2024)
di: Zhang, Ming, et al.
Pubblicazione: (2024)
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
di: Yoon, Jaehong, et al.
Pubblicazione: (2024)
di: Yoon, Jaehong, et al.
Pubblicazione: (2024)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
di: Yu, Shoubin, et al.
Pubblicazione: (2024)
di: Yu, Shoubin, et al.
Pubblicazione: (2024)
Automating Video Thumbnails Selection and Generation with Multimodal and Multistage Analysis
di: Fantini, Elia
Pubblicazione: (2024)
di: Fantini, Elia
Pubblicazione: (2024)
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
di: Zhu, Dawei, et al.
Pubblicazione: (2025)
di: Zhu, Dawei, et al.
Pubblicazione: (2025)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
di: Zhang, Kaichen, et al.
Pubblicazione: (2024)
di: Zhang, Kaichen, et al.
Pubblicazione: (2024)
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
di: Jiang, Yifan, et al.
Pubblicazione: (2026)
di: Jiang, Yifan, et al.
Pubblicazione: (2026)
Focus Matters: Phase-Aware Suppression for Hallucination in Vision-Language Models
di: Kim, Sohyeon, et al.
Pubblicazione: (2026)
di: Kim, Sohyeon, et al.
Pubblicazione: (2026)
Learning Spatiotemporal Inconsistency via Thumbnail Layout for Face Deepfake Detection
di: Xu, Yuting, et al.
Pubblicazione: (2024)
di: Xu, Yuting, et al.
Pubblicazione: (2024)
Altogether: Image Captioning via Re-aligning Alt-text
di: Xu, Hu, et al.
Pubblicazione: (2024)
di: Xu, Hu, et al.
Pubblicazione: (2024)
Vision-Language Models Do Not Understand Negation
di: Alhamoud, Kumail, et al.
Pubblicazione: (2025)
di: Alhamoud, Kumail, et al.
Pubblicazione: (2025)
PaperBanana: Automating Academic Illustration for AI Scientists
di: Zhu, Dawei, et al.
Pubblicazione: (2026)
di: Zhu, Dawei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity
di: Jung, Jaeyoon, et al.
Pubblicazione: (2026) -
HerO at AVeriTeC: The Herd of Open Large Language Models for Verifying Real-World Claims
di: Yoon, Yejun, et al.
Pubblicazione: (2024) -
Team HUMANE at AVeriTeC 2025: HerO 2 for Efficient Fact Verification
di: Yoon, Yejun, et al.
Pubblicazione: (2025) -
Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion
di: Yoon, Yejun, et al.
Pubblicazione: (2025) -
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
di: Zang, Yuan, et al.
Pubblicazione: (2025)