Mind the Gap: Analyzing Lacunae with Transformer-Based Transcription
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Borkar, Jaydeep, Smith, David A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025)
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025)
Composition and Deformance: Measuring Imageability with a Text-to-Image Model
von: Wu, Si, et al.
Veröffentlicht: (2023)
von: Wu, Si, et al.
Veröffentlicht: (2023)
The Mind's Eye: A Multi-Faceted Reward Framework for Guiding Visual Metaphor Generation
von: Koushik, Girish A., et al.
Veröffentlicht: (2025)
von: Koushik, Girish A., et al.
Veröffentlicht: (2025)
Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs
von: Gautam, Somraj, et al.
Veröffentlicht: (2025)
von: Gautam, Somraj, et al.
Veröffentlicht: (2025)
Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
von: Li, Yifan, et al.
Veröffentlicht: (2025)
von: Li, Yifan, et al.
Veröffentlicht: (2025)
Narrowing the Gap between Vision and Action in Navigation
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
The Abstraction Gap in Vision-Language Causal Reasoning
von: Hoang, Chinh, et al.
Veröffentlicht: (2026)
von: Hoang, Chinh, et al.
Veröffentlicht: (2026)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2025)
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2025)
Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
von: He, Lehan, et al.
Veröffentlicht: (2024)
von: He, Lehan, et al.
Veröffentlicht: (2024)
Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
von: Zheng, Haohan, et al.
Veröffentlicht: (2025)
von: Zheng, Haohan, et al.
Veröffentlicht: (2025)
Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and Disinformation
von: Yuan, Xin, et al.
Veröffentlicht: (2023)
von: Yuan, Xin, et al.
Veröffentlicht: (2023)
MindBench: A Comprehensive Benchmark for Mind Map Structure Recognition and Analysis
von: Chen, Lei, et al.
Veröffentlicht: (2024)
von: Chen, Lei, et al.
Veröffentlicht: (2024)
Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
von: Rose, Daniel, et al.
Veröffentlicht: (2023)
von: Rose, Daniel, et al.
Veröffentlicht: (2023)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
von: Huang, Jen-Tse, et al.
Veröffentlicht: (2025)
von: Huang, Jen-Tse, et al.
Veröffentlicht: (2025)
Enhanced Sentiment Analysis of Iranian Restaurant Reviews Utilizing Sentiment Intensity Analyzer & Fuzzy Logic
von: Rokhva, Shayan, et al.
Veröffentlicht: (2025)
von: Rokhva, Shayan, et al.
Veröffentlicht: (2025)
When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise
von: Shin, Philip Wootaek, et al.
Veröffentlicht: (2026)
von: Shin, Philip Wootaek, et al.
Veröffentlicht: (2026)
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
von: Wang, Zirui, et al.
Veröffentlicht: (2024)
von: Wang, Zirui, et al.
Veröffentlicht: (2024)
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
von: Bandraupalli, Srihari, et al.
Veröffentlicht: (2025)
von: Bandraupalli, Srihari, et al.
Veröffentlicht: (2025)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
von: Bin, Yi, et al.
Veröffentlicht: (2024)
von: Bin, Yi, et al.
Veröffentlicht: (2024)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
von: Nazi, Zabir Al, et al.
Veröffentlicht: (2025)
von: Nazi, Zabir Al, et al.
Veröffentlicht: (2025)
Unknown Prompt, the only Lacuna: Unveiling CLIP's Potential for Open Domain Generalization
von: Singha, Mainak, et al.
Veröffentlicht: (2024)
von: Singha, Mainak, et al.
Veröffentlicht: (2024)
EviMem: Evidence-Gap-Driven Iterative Retrieval for Long-Term Conversational Memory
von: Li, Yuyang, et al.
Veröffentlicht: (2026)
von: Li, Yuyang, et al.
Veröffentlicht: (2026)
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
von: Sun, Kaiser, et al.
Veröffentlicht: (2026)
von: Sun, Kaiser, et al.
Veröffentlicht: (2026)
Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
von: Xu, Yige, et al.
Veröffentlicht: (2026)
von: Xu, Yige, et al.
Veröffentlicht: (2026)
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
von: Liu, Peng, et al.
Veröffentlicht: (2025)
von: Liu, Peng, et al.
Veröffentlicht: (2025)
PromptSync: Bridging Domain Gaps in Vision-Language Models through Class-Aware Prototype Alignment and Discrimination
von: Khandelwal, Anant
Veröffentlicht: (2024)
von: Khandelwal, Anant
Veröffentlicht: (2024)
Analyzing Images of Legal Documents: Toward Multi-Modal LLMs for Access to Justice
von: Westermann, Hannes, et al.
Veröffentlicht: (2024)
von: Westermann, Hannes, et al.
Veröffentlicht: (2024)
An HTR-LLM Workflow for High-Accuracy Transcription and Analysis of Abbreviated Latin Court Hand
von: Isom, Joshua D.
Veröffentlicht: (2025)
von: Isom, Joshua D.
Veröffentlicht: (2025)
Attribute Diversity Determines the Systematicity Gap in VQA
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2023)
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2023)
OMCAT: Omni Context Aware Transformer
von: Goel, Arushi, et al.
Veröffentlicht: (2024)
von: Goel, Arushi, et al.
Veröffentlicht: (2024)
Mind the Gap: Bridging Occlusion in Gait Recognition via Residual Gap Correction
von: Gupta, Ayush, et al.
Veröffentlicht: (2025)
von: Gupta, Ayush, et al.
Veröffentlicht: (2025)
MindCube: Spatial Mental Modeling from Limited Views
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
Analyzing Finetuning Representation Shift for Multimodal LLMs Steering
von: Khayatan, Pegah, et al.
Veröffentlicht: (2025)
von: Khayatan, Pegah, et al.
Veröffentlicht: (2025)
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
von: Zhou, Wenrui, et al.
Veröffentlicht: (2025)
von: Zhou, Wenrui, et al.
Veröffentlicht: (2025)
Analyzing The Language of Visual Tokens
von: Chan, David M., et al.
Veröffentlicht: (2024)
von: Chan, David M., et al.
Veröffentlicht: (2024)
Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual Learning
von: Huang, Linlan, et al.
Veröffentlicht: (2025)
von: Huang, Linlan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025) -
Composition and Deformance: Measuring Imageability with a Text-to-Image Model
von: Wu, Si, et al.
Veröffentlicht: (2023) -
The Mind's Eye: A Multi-Faceted Reward Framework for Guiding Visual Metaphor Generation
von: Koushik, Girish A., et al.
Veröffentlicht: (2025) -
Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025) -
Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs
von: Gautam, Somraj, et al.
Veröffentlicht: (2025)