Understanding Figurative Meaning through Explainable Visual Entailment
Fuente:
arXiv
Salvato in:
| Autori principali: | Saakyan, Arkadiy, Kulkarni, Shreyas, Chakrabarty, Tuhin, Muresan, Smaranda |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Death of the Novel(ty): Beyond n-Gram Novelty as a Metric for Textual Creativity
di: Saakyan, Arkadiy, et al.
Pubblicazione: (2025)
di: Saakyan, Arkadiy, et al.
Pubblicazione: (2025)
VideoNorms: Benchmarking Cultural Awareness of Video Language Models
di: Varimalla, Nikhil Reddy, et al.
Pubblicazione: (2025)
di: Varimalla, Nikhil Reddy, et al.
Pubblicazione: (2025)
ICLEF: In-Context Learning with Expert Feedback for Explainable Style Transfer
di: Saakyan, Arkadiy, et al.
Pubblicazione: (2023)
di: Saakyan, Arkadiy, et al.
Pubblicazione: (2023)
Creativity Support in the Age of Large Language Models: An Empirical Study Involving Emerging Writers
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2023)
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2023)
Connecting the Dots: Evaluating Abstract Reasoning Capabilities of LLMs Using the New York Times Connections Word Game
di: Samadarshi, Prisha, et al.
Pubblicazione: (2024)
di: Samadarshi, Prisha, et al.
Pubblicazione: (2024)
Probing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls
di: Pitta, Elena, et al.
Pubblicazione: (2025)
di: Pitta, Elena, et al.
Pubblicazione: (2025)
Art or Artifice? Large Language Models and the False Promise of Creativity
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2023)
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2023)
Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment
di: Zhang, Yue, et al.
Pubblicazione: (2025)
di: Zhang, Yue, et al.
Pubblicazione: (2025)
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
di: Chen, Jiali, et al.
Pubblicazione: (2024)
di: Chen, Jiali, et al.
Pubblicazione: (2024)
Explainable Synthetic Image Detection through Diffusion Timestep Ensembling
di: Wu, Yixin, et al.
Pubblicazione: (2025)
di: Wu, Yixin, et al.
Pubblicazione: (2025)
FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures
di: Song, Jifeng, et al.
Pubblicazione: (2026)
di: Song, Jifeng, et al.
Pubblicazione: (2026)
Internalized Reasoning for Long-Context Visual Document Understanding
di: Veselka, Austin
Pubblicazione: (2026)
di: Veselka, Austin
Pubblicazione: (2026)
Figuring out Figures: Using Textual References to Caption Scientific Figures
di: Cao, Stanley, et al.
Pubblicazione: (2024)
di: Cao, Stanley, et al.
Pubblicazione: (2024)
Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
di: Kawasaki, Haruka, et al.
Pubblicazione: (2026)
di: Kawasaki, Haruka, et al.
Pubblicazione: (2026)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
di: Masala, Mihai, et al.
Pubblicazione: (2025)
di: Masala, Mihai, et al.
Pubblicazione: (2025)
Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
di: Tartaglini, Alexa R., et al.
Pubblicazione: (2025)
di: Tartaglini, Alexa R., et al.
Pubblicazione: (2025)
EmoGist: Efficient In-Context Learning for Visual Emotion Understanding
di: Seoh, Ronald, et al.
Pubblicazione: (2025)
di: Seoh, Ronald, et al.
Pubblicazione: (2025)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
di: Wang, Dianyi, et al.
Pubblicazione: (2025)
di: Wang, Dianyi, et al.
Pubblicazione: (2025)
MM-PoE: Multiple Choice Reasoning via. Process of Elimination using Multi-Modal Models
di: Chakrabarty, Sayak, et al.
Pubblicazione: (2024)
di: Chakrabarty, Sayak, et al.
Pubblicazione: (2024)
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
di: Zhou, Kaiwen, et al.
Pubblicazione: (2023)
di: Zhou, Kaiwen, et al.
Pubblicazione: (2023)
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
di: Li, Kailing, et al.
Pubblicazione: (2025)
di: Li, Kailing, et al.
Pubblicazione: (2025)
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
di: Ng, Ho Yin 'Sam', et al.
Pubblicazione: (2025)
di: Ng, Ho Yin 'Sam', et al.
Pubblicazione: (2025)
From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation
di: Chen, Zhen, et al.
Pubblicazione: (2025)
di: Chen, Zhen, et al.
Pubblicazione: (2025)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
di: Lin, Bin, et al.
Pubblicazione: (2025)
di: Lin, Bin, et al.
Pubblicazione: (2025)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation
di: Zhang, Yichi, et al.
Pubblicazione: (2025)
di: Zhang, Yichi, et al.
Pubblicazione: (2025)
ProxyThinker: Test-Time Guidance through Small Visual Reasoners
di: Xiao, Zilin, et al.
Pubblicazione: (2025)
di: Xiao, Zilin, et al.
Pubblicazione: (2025)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
di: Liu, Huabin, et al.
Pubblicazione: (2025)
di: Liu, Huabin, et al.
Pubblicazione: (2025)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
di: Zhou, Yuqi, et al.
Pubblicazione: (2025)
di: Zhou, Yuqi, et al.
Pubblicazione: (2025)
V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model Interaction
di: Zhao, Yiming, et al.
Pubblicazione: (2025)
di: Zhao, Yiming, et al.
Pubblicazione: (2025)
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
di: Jia, Yiming, et al.
Pubblicazione: (2025)
di: Jia, Yiming, et al.
Pubblicazione: (2025)
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
di: Atuhurra, Jesse, et al.
Pubblicazione: (2025)
di: Atuhurra, Jesse, et al.
Pubblicazione: (2025)
Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
di: Chetan, Aditya, et al.
Pubblicazione: (2026)
di: Chetan, Aditya, et al.
Pubblicazione: (2026)
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
di: Wang, Qiuchen, et al.
Pubblicazione: (2025)
di: Wang, Qiuchen, et al.
Pubblicazione: (2025)
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
di: Kim, Hyunjong, et al.
Pubblicazione: (2025)
di: Kim, Hyunjong, et al.
Pubblicazione: (2025)
Enhancing Visual Dialog State Tracking through Iterative Object-Entity Alignment in Multi-Round Conversations
di: Pang, Wei, et al.
Pubblicazione: (2024)
di: Pang, Wei, et al.
Pubblicazione: (2024)
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
di: Zhang, Juntian, et al.
Pubblicazione: (2025)
di: Zhang, Juntian, et al.
Pubblicazione: (2025)
PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck
di: Pham, Thang M., et al.
Pubblicazione: (2024)
di: Pham, Thang M., et al.
Pubblicazione: (2024)
Show Me the World in My Language: Establishing the First Baseline for Scene-Text to Scene-Text Translation
di: Vaidya, Shreyas, et al.
Pubblicazione: (2023)
di: Vaidya, Shreyas, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Death of the Novel(ty): Beyond n-Gram Novelty as a Metric for Textual Creativity
di: Saakyan, Arkadiy, et al.
Pubblicazione: (2025) -
VideoNorms: Benchmarking Cultural Awareness of Video Language Models
di: Varimalla, Nikhil Reddy, et al.
Pubblicazione: (2025) -
ICLEF: In-Context Learning with Expert Feedback for Explainable Style Transfer
di: Saakyan, Arkadiy, et al.
Pubblicazione: (2023) -
Creativity Support in the Age of Large Language Models: An Empirical Study Involving Emerging Writers
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2023) -
Connecting the Dots: Evaluating Abstract Reasoning Capabilities of LLMs Using the New York Times Connections Word Game
di: Samadarshi, Prisha, et al.
Pubblicazione: (2024)