FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
Fuente:
arXiv
Guardado en:
| Autores principales: | Hua, Hang, Shi, Jing, Kafle, Kushal, Jenni, Simon, Zhang, Daoan, Collomosse, John, Cohen, Scott, Luo, Jiebo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling Capabilities
por: Khosla, Savya, et al.
Publicado: (2025)
por: Khosla, Savya, et al.
Publicado: (2025)
Improving Large Vision and Language Models by Learning from a Panel of Peers
por: Hernandez, Jefferson, et al.
Publicado: (2025)
por: Hernandez, Jefferson, et al.
Publicado: (2025)
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
por: Wu, Qiucheng, et al.
Publicado: (2026)
por: Wu, Qiucheng, et al.
Publicado: (2026)
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
por: Lu, Jianglin, et al.
Publicado: (2026)
por: Lu, Jianglin, et al.
Publicado: (2026)
FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication
por: Slyman, Eric, et al.
Publicado: (2024)
por: Slyman, Eric, et al.
Publicado: (2024)
SCoRD: Subject-Conditional Relation Detection with Text-Augmented Data
por: Yang, Ziyan, et al.
Publicado: (2023)
por: Yang, Ziyan, et al.
Publicado: (2023)
Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
por: Akdemir, Kiymet, et al.
Publicado: (2025)
por: Akdemir, Kiymet, et al.
Publicado: (2025)
VIXEN: Visual Text Comparison Network for Image Difference Captioning
por: Black, Alexander, et al.
Publicado: (2024)
por: Black, Alexander, et al.
Publicado: (2024)
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
por: Zeng, Ziyun, et al.
Publicado: (2025)
por: Zeng, Ziyun, et al.
Publicado: (2025)
More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
por: Just, Hoang Anh, et al.
Publicado: (2025)
por: Just, Hoang Anh, et al.
Publicado: (2025)
GaussianStyle: Gaussian Head Avatar via StyleGAN
por: Liu, Pinxin, et al.
Publicado: (2024)
por: Liu, Pinxin, et al.
Publicado: (2024)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
por: Shrestha, Robik, et al.
Publicado: (2020)
por: Shrestha, Robik, et al.
Publicado: (2020)
Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation
por: Chen, Wenting, et al.
Publicado: (2023)
por: Chen, Wenting, et al.
Publicado: (2023)
WorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generation
por: Zhang, Daoan, et al.
Publicado: (2025)
por: Zhang, Daoan, et al.
Publicado: (2025)
CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs
por: Zhang, Daoan, et al.
Publicado: (2024)
por: Zhang, Daoan, et al.
Publicado: (2024)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
por: Gordon, Brian, et al.
Publicado: (2023)
por: Gordon, Brian, et al.
Publicado: (2023)
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
por: Slyman, Eric, et al.
Publicado: (2025)
por: Slyman, Eric, et al.
Publicado: (2025)
The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
por: Qi, Daiqing, et al.
Publicado: (2025)
por: Qi, Daiqing, et al.
Publicado: (2025)
Sphinx: Benchmarking and Modeling for LLM-Driven Pull Request Review
por: Zhang, Daoan, et al.
Publicado: (2026)
por: Zhang, Daoan, et al.
Publicado: (2026)
A Versatile Multimodal Agent for Multimedia Content Generation
por: Zhang, Daoan, et al.
Publicado: (2026)
por: Zhang, Daoan, et al.
Publicado: (2026)
LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
por: Wang, Shuai, et al.
Publicado: (2025)
por: Wang, Shuai, et al.
Publicado: (2025)
FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity
por: Hua, Hang, et al.
Publicado: (2024)
por: Hua, Hang, et al.
Publicado: (2024)
Learning Brain Tumor Representation in 3D High-Resolution MR Images via Interpretable State Space Models
por: Hu, Qingqiao, et al.
Publicado: (2024)
por: Hu, Qingqiao, et al.
Publicado: (2024)
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
por: Yang, Ziyan, et al.
Publicado: (2022)
por: Yang, Ziyan, et al.
Publicado: (2022)
Are Bias Mitigation Techniques for Deep Learning Effective?
por: Shrestha, Robik, et al.
Publicado: (2021)
por: Shrestha, Robik, et al.
Publicado: (2021)
Fine-grained Text to Image Synthesis
por: Ouyang, Xu, et al.
Publicado: (2024)
por: Ouyang, Xu, et al.
Publicado: (2024)
CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
por: Dong, Qihua, et al.
Publicado: (2025)
por: Dong, Qihua, et al.
Publicado: (2025)
V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
por: Hua, Hang, et al.
Publicado: (2024)
por: Hua, Hang, et al.
Publicado: (2024)
PromptFix: You Prompt and We Fix the Photo
por: Yu, Yongsheng, et al.
Publicado: (2024)
por: Yu, Yongsheng, et al.
Publicado: (2024)
VisualActBench: Can VLMs See and Act like a Human?
por: Zhang, Daoan, et al.
Publicado: (2025)
por: Zhang, Daoan, et al.
Publicado: (2025)
PARASOL: Parametric Style Control for Diffusion Image Synthesis
por: Tarrés, Gemma Canet, et al.
Publicado: (2023)
por: Tarrés, Gemma Canet, et al.
Publicado: (2023)
MultiNeRF: Multiple Watermark Embedding for Neural Radiance Fields
por: Kulthe, Yash, et al.
Publicado: (2025)
por: Kulthe, Yash, et al.
Publicado: (2025)
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
por: Cui, Wanqing, et al.
Publicado: (2024)
por: Cui, Wanqing, et al.
Publicado: (2024)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
por: Tang, Yunlong, et al.
Publicado: (2025)
por: Tang, Yunlong, et al.
Publicado: (2025)
Improving Taxonomic Image-based Out-of-distribution Detection With DNA Barcodes
por: Impiö, Mikko, et al.
Publicado: (2024)
por: Impiö, Mikko, et al.
Publicado: (2024)
VideoXum: Cross-modal Visual and Textural Summarization of Videos
por: Lin, Jingyang, et al.
Publicado: (2023)
por: Lin, Jingyang, et al.
Publicado: (2023)
FITA: Fine-grained Image-Text Aligner for Radiology Report Generation
por: Yang, Honglong, et al.
Publicado: (2024)
por: Yang, Honglong, et al.
Publicado: (2024)
Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval
por: Ma, Zehong, et al.
Publicado: (2025)
por: Ma, Zehong, et al.
Publicado: (2025)
MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models
por: Hua, Hang, et al.
Publicado: (2025)
por: Hua, Hang, et al.
Publicado: (2025)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
por: Zeng, Ziyun, et al.
Publicado: (2026)
por: Zeng, Ziyun, et al.
Publicado: (2026)
Ejemplares similares
-
MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling Capabilities
por: Khosla, Savya, et al.
Publicado: (2025) -
Improving Large Vision and Language Models by Learning from a Panel of Peers
por: Hernandez, Jefferson, et al.
Publicado: (2025) -
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
por: Wu, Qiucheng, et al.
Publicado: (2026) -
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
por: Lu, Jianglin, et al.
Publicado: (2026) -
FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication
por: Slyman, Eric, et al.
Publicado: (2024)