Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
Fuente:
arXiv
Guardado en:
| Autores principales: | Beňová, Ivana, Košecká, Jana, Gregor, Michal, Tamajka, Martin, Veselý, Marcel, Šimko, Marián |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CV-Probes: Studying the interplay of lexical and world knowledge in visually grounded verb understanding
por: Beňová, Ivana, et al.
Publicado: (2024)
por: Beňová, Ivana, et al.
Publicado: (2024)
o-MEGA: Optimized Methods for Explanation Generation and Analysis
por: Kriš, Ľuboš, et al.
Publicado: (2025)
por: Kriš, Ľuboš, et al.
Publicado: (2025)
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
por: Rajabi, Navid, et al.
Publicado: (2024)
por: Rajabi, Navid, et al.
Publicado: (2024)
Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing
por: Fayyazsanavi, Pooya, et al.
Publicado: (2024)
por: Fayyazsanavi, Pooya, et al.
Publicado: (2024)
skLEP: A Slovak General Language Understanding Benchmark
por: Šuppa, Marek, et al.
Publicado: (2025)
por: Šuppa, Marek, et al.
Publicado: (2025)
Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models
por: Rajabi, Navid, et al.
Publicado: (2023)
por: Rajabi, Navid, et al.
Publicado: (2023)
Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM
por: Rajabi, Navid, et al.
Publicado: (2024)
por: Rajabi, Navid, et al.
Publicado: (2024)
Large Language Models for Multilingual Previously Fact-Checked Claim Detection
por: Vykopal, Ivan, et al.
Publicado: (2025)
por: Vykopal, Ivan, et al.
Publicado: (2025)
Compositional Image-Text Matching and Retrieval by Grounding Entities
por: Vongala, Madhukar Reddy, et al.
Publicado: (2025)
por: Vongala, Madhukar Reddy, et al.
Publicado: (2025)
TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation
por: Rajabi, Navid, et al.
Publicado: (2025)
por: Rajabi, Navid, et al.
Publicado: (2025)
A Generative-AI-Driven Claim Retrieval System Capable of Detecting and Retrieving Claims from Social Media Platforms in Multiple Languages
por: Vykopal, Ivan, et al.
Publicado: (2025)
por: Vykopal, Ivan, et al.
Publicado: (2025)
Investigating Language and Retrieval Bias in Multilingual Previously Fact-Checked Claim Detection
por: Vykopal, Ivan, et al.
Publicado: (2025)
por: Vykopal, Ivan, et al.
Publicado: (2025)
Soft Language Prompts for Language Transfer
por: Vykopal, Ivan, et al.
Publicado: (2024)
por: Vykopal, Ivan, et al.
Publicado: (2024)
Assessing Web Search Credibility and Response Groundedness in Chat Assistants
por: Vykopal, Ivan, et al.
Publicado: (2025)
por: Vykopal, Ivan, et al.
Publicado: (2025)
Generative Large Language Models in Automated Fact-Checking: A Survey
por: Vykopal, Ivan, et al.
Publicado: (2024)
por: Vykopal, Ivan, et al.
Publicado: (2024)
Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling
por: Pikuliak, Matúš, et al.
Publicado: (2023)
por: Pikuliak, Matúš, et al.
Publicado: (2023)
Political Leaning and Politicalness Classification of Texts
por: Volf, Matous, et al.
Publicado: (2025)
por: Volf, Matous, et al.
Publicado: (2025)
Toward Phonology-Guided Sign Language Motion Generation: A Diffusion Baseline and Conditioning Analysis
por: Hong, Rui, et al.
Publicado: (2026)
por: Hong, Rui, et al.
Publicado: (2026)
LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?
por: Cegin, Jan, et al.
Publicado: (2024)
por: Cegin, Jan, et al.
Publicado: (2024)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
por: Yang, Yue, et al.
Publicado: (2025)
por: Yang, Yue, et al.
Publicado: (2025)
MULTI: Multimodal Understanding Leaderboard with Text and Images
por: Zhu, Zichen, et al.
Publicado: (2024)
por: Zhu, Zichen, et al.
Publicado: (2024)
Unlocking Korean Verbs: A User-Friendly Exploration into the Verb Lexicon
por: Song, Seohyun, et al.
Publicado: (2024)
por: Song, Seohyun, et al.
Publicado: (2024)
Advanced Multimodal Deep Learning Architecture for Image-Text Matching
por: Wang, Jinyin, et al.
Publicado: (2024)
por: Wang, Jinyin, et al.
Publicado: (2024)
SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval
por: Peng, Qiwei, et al.
Publicado: (2025)
por: Peng, Qiwei, et al.
Publicado: (2025)
PointSplat: Efficient Geometry-Driven Pruning and Transformer Refinement for 3D Gaussian Splatting
por: Tran, Anh Thuan, et al.
Publicado: (2026)
por: Tran, Anh Thuan, et al.
Publicado: (2026)
Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching
por: Wang, Bin, et al.
Publicado: (2024)
por: Wang, Bin, et al.
Publicado: (2024)
Interpretable Predictability-Based AI Text Detection: A Replication Study
por: Skurla, Adam, et al.
Publicado: (2026)
por: Skurla, Adam, et al.
Publicado: (2026)
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
por: Masry, Ahmed, et al.
Publicado: (2025)
por: Masry, Ahmed, et al.
Publicado: (2025)
Beyond Coarse-Grained Matching in Video-Text Retrieval
por: Chen, Aozhu, et al.
Publicado: (2024)
por: Chen, Aozhu, et al.
Publicado: (2024)
Text Role Classification in Scientific Charts Using Multimodal Transformers
por: Kim, Hye Jin, et al.
Publicado: (2024)
por: Kim, Hye Jin, et al.
Publicado: (2024)
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
por: Lv, Zheqi, et al.
Publicado: (2025)
por: Lv, Zheqi, et al.
Publicado: (2025)
Token Masking Improves Transformer-Based Text Classification
por: Xu, Xianglong, et al.
Publicado: (2025)
por: Xu, Xianglong, et al.
Publicado: (2025)
Gesture-Aware Pretraining and Token Fusion for 3D Hand Pose Estimation
por: Hong, Rui, et al.
Publicado: (2026)
por: Hong, Rui, et al.
Publicado: (2026)
Conditional Collapse in Sign Language Production: A Diagnostic and a Scaling Argument
por: Hong, Rui, et al.
Publicado: (2026)
por: Hong, Rui, et al.
Publicado: (2026)
Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection
por: Mullick, Ankan, et al.
Publicado: (2025)
por: Mullick, Ankan, et al.
Publicado: (2025)
Automatic Construction of Chinese Verb Collostruction Database
por: Tang, Xuri, et al.
Publicado: (2025)
por: Tang, Xuri, et al.
Publicado: (2025)
Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation for Classification
por: Cegin, Jan, et al.
Publicado: (2024)
por: Cegin, Jan, et al.
Publicado: (2024)
GSIFN: A Graph-Structured and Interlaced-Masked Multimodal Transformer-based Fusion Network for Multimodal Sentiment Analysis
por: Jin, Yijie
Publicado: (2024)
por: Jin, Yijie
Publicado: (2024)
AI Research is not Magic, it has to be Reproducible and Responsible: Challenges in the AI field from the Perspective of its PhD Students
por: Hrckova, Andrea, et al.
Publicado: (2024)
por: Hrckova, Andrea, et al.
Publicado: (2024)
DP-MLM: Differentially Private Text Rewriting Using Masked Language Models
por: Meisenbacher, Stephen, et al.
Publicado: (2024)
por: Meisenbacher, Stephen, et al.
Publicado: (2024)
Ejemplares similares
-
CV-Probes: Studying the interplay of lexical and world knowledge in visually grounded verb understanding
por: Beňová, Ivana, et al.
Publicado: (2024) -
o-MEGA: Optimized Methods for Explanation Generation and Analysis
por: Kriš, Ľuboš, et al.
Publicado: (2025) -
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
por: Rajabi, Navid, et al.
Publicado: (2024) -
Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing
por: Fayyazsanavi, Pooya, et al.
Publicado: (2024) -
skLEP: A Slovak General Language Understanding Benchmark
por: Šuppa, Marek, et al.
Publicado: (2025)