VL-GLUE: A Suite of Fundamental yet Challenging Visuo-Linguistic Reasoning Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Sampat, Shailaja Keyur, Nakamura, Mutsumi, Kailas, Shankar, Aggarwal, Kartik, Zhou, Mandy, Yang, Yezhou, Baral, Chitta |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ActionCOMET: A Zero-shot Approach to Learn Image-specific Commonsense Concepts about Actions
por: Sampat, Shailaja Keyur, et al.
Publicado: (2024)
por: Sampat, Shailaja Keyur, et al.
Publicado: (2024)
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?
por: Sampat, Shailaja Keyur, et al.
Publicado: (2024)
por: Sampat, Shailaja Keyur, et al.
Publicado: (2024)
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
por: Chatterjee, Agneet, et al.
Publicado: (2024)
por: Chatterjee, Agneet, et al.
Publicado: (2024)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
por: Patel, Maitreya, et al.
Publicado: (2024)
por: Patel, Maitreya, et al.
Publicado: (2024)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
por: Patel, Maitreya, et al.
Publicado: (2023)
por: Patel, Maitreya, et al.
Publicado: (2023)
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
por: Patel, Nisarg, et al.
Publicado: (2024)
por: Patel, Nisarg, et al.
Publicado: (2024)
Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
por: Luo, Man, et al.
Publicado: (2023)
por: Luo, Man, et al.
Publicado: (2023)
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
por: Saxon, Michael, et al.
Publicado: (2024)
por: Saxon, Michael, et al.
Publicado: (2024)
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
por: Tyagi, Nemika, et al.
Publicado: (2024)
por: Tyagi, Nemika, et al.
Publicado: (2024)
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
por: Parmar, Mihir, et al.
Publicado: (2024)
por: Parmar, Mihir, et al.
Publicado: (2024)
PHANTOM RECALL: When Familiar Puzzles Fool Smart Models
por: Mukhopadhyay, Souradeep, et al.
Publicado: (2025)
por: Mukhopadhyay, Souradeep, et al.
Publicado: (2025)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
por: Patel, Maitreya, et al.
Publicado: (2024)
por: Patel, Maitreya, et al.
Publicado: (2024)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
por: Yilmaz, Nilay, et al.
Publicado: (2025)
por: Yilmaz, Nilay, et al.
Publicado: (2025)
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
por: Malaviya, Vatsal, et al.
Publicado: (2025)
por: Malaviya, Vatsal, et al.
Publicado: (2025)
REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
por: Chatterjee, Agneet, et al.
Publicado: (2024)
por: Chatterjee, Agneet, et al.
Publicado: (2024)
RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
por: Pathiraja, Bimsara, et al.
Publicado: (2025)
por: Pathiraja, Bimsara, et al.
Publicado: (2025)
Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations
por: Liu, Xiangrui, et al.
Publicado: (2025)
por: Liu, Xiangrui, et al.
Publicado: (2025)
How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $τ$-bench
por: Mishra, Venkatesh, et al.
Publicado: (2025)
por: Mishra, Venkatesh, et al.
Publicado: (2025)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
por: Fallah, Forouzan, et al.
Publicado: (2025)
por: Fallah, Forouzan, et al.
Publicado: (2025)
Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
por: Vani, Sameep, et al.
Publicado: (2025)
por: Vani, Sameep, et al.
Publicado: (2025)
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
por: Siingh, Shikhhar, et al.
Publicado: (2025)
por: Siingh, Shikhhar, et al.
Publicado: (2025)
EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment
por: Kusumba, Abhiram, et al.
Publicado: (2025)
por: Kusumba, Abhiram, et al.
Publicado: (2025)
Grounding Stylistic Domain Generalization with Quantitative Domain Shift Measures and Synthetic Scene Images
por: Luo, Yiran, et al.
Publicado: (2024)
por: Luo, Yiran, et al.
Publicado: (2024)
Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
por: Saeidi, Amir, et al.
Publicado: (2024)
por: Saeidi, Amir, et al.
Publicado: (2024)
Dual Caption Preference Optimization for Diffusion Models
por: Saeidi, Amir, et al.
Publicado: (2025)
por: Saeidi, Amir, et al.
Publicado: (2025)
Chimera: Compositional Image Generation using Part-based Concepting
por: Singh, Shivam, et al.
Publicado: (2025)
por: Singh, Shivam, et al.
Publicado: (2025)
Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark
por: Gupta, Himanshu, et al.
Publicado: (2024)
por: Gupta, Himanshu, et al.
Publicado: (2024)
ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints
por: Handa, Divij, et al.
Publicado: (2024)
por: Handa, Divij, et al.
Publicado: (2024)
AgriGPT-VL: Agricultural Vision-Language Understanding Suite
por: Yang, Bo, et al.
Publicado: (2025)
por: Yang, Bo, et al.
Publicado: (2025)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
por: Ashihara, Takanori, et al.
Publicado: (2023)
por: Ashihara, Takanori, et al.
Publicado: (2023)
Investigating and Addressing Hallucinations of LLMs in Tasks Involving Negation
por: Varshney, Neeraj, et al.
Publicado: (2024)
por: Varshney, Neeraj, et al.
Publicado: (2024)
Map&Make: Schema Guided Text to Table Generation
por: Ahuja, Naman, et al.
Publicado: (2025)
por: Ahuja, Naman, et al.
Publicado: (2025)
The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness
por: Varshney, Neeraj, et al.
Publicado: (2023)
por: Varshney, Neeraj, et al.
Publicado: (2023)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
por: Chen, Zhe, et al.
Publicado: (2023)
por: Chen, Zhe, et al.
Publicado: (2023)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
por: Parmar, Mihir, et al.
Publicado: (2022)
por: Parmar, Mihir, et al.
Publicado: (2022)
Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
por: Chatterjee, Agneet, et al.
Publicado: (2025)
por: Chatterjee, Agneet, et al.
Publicado: (2025)
Rethinking Information Synthesis in Multimodal Question Answering A Multi-Agent Perspective
por: Rajput, Krishna Singh, et al.
Publicado: (2025)
por: Rajput, Krishna Singh, et al.
Publicado: (2025)
ltzGLUE: Luxembourgish General Language Understanding Evaluation
por: Plum, Alistair, et al.
Publicado: (2026)
por: Plum, Alistair, et al.
Publicado: (2026)
ToW: Thoughts of Words Improve Reasoning in Large Language Models
por: Xu, Zhikun, et al.
Publicado: (2024)
por: Xu, Zhikun, et al.
Publicado: (2024)
Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning
por: Mishra, Venkatesh, et al.
Publicado: (2025)
por: Mishra, Venkatesh, et al.
Publicado: (2025)
Ejemplares similares
-
ActionCOMET: A Zero-shot Approach to Learn Image-specific Commonsense Concepts about Actions
por: Sampat, Shailaja Keyur, et al.
Publicado: (2024) -
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?
por: Sampat, Shailaja Keyur, et al.
Publicado: (2024) -
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
por: Chatterjee, Agneet, et al.
Publicado: (2024) -
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
por: Patel, Maitreya, et al.
Publicado: (2024) -
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
por: Patel, Maitreya, et al.
Publicado: (2023)