Beyond the Textual: Generating Coherent Visual Options for MCQs
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Wanqiang, He, Longzhu, Zheng, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
CLEAR: Character Unlearning in Textual and Visual Modalities
by: Dontsov, Alexey, et al.
Published: (2024)
by: Dontsov, Alexey, et al.
Published: (2024)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
by: Gordon, Brian, et al.
Published: (2023)
by: Gordon, Brian, et al.
Published: (2023)
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
by: Sun, Hao, et al.
Published: (2026)
by: Sun, Hao, et al.
Published: (2026)
RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
by: Wang, Chenglong, et al.
Published: (2024)
by: Wang, Chenglong, et al.
Published: (2024)
FSMR: A Feature Swapping Multi-modal Reasoning Approach with Joint Textual and Visual Clues
by: Li, Shuang, et al.
Published: (2024)
by: Li, Shuang, et al.
Published: (2024)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
by: Zhang, Beichen, et al.
Published: (2025)
by: Zhang, Beichen, et al.
Published: (2025)
Tell Me What's Next: Textual Foresight for Generic UI Representations
by: Burns, Andrea, et al.
Published: (2024)
by: Burns, Andrea, et al.
Published: (2024)
BabyVision: Visual Reasoning Beyond Language
by: Chen, Liang, et al.
Published: (2026)
by: Chen, Liang, et al.
Published: (2026)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
by: Hashemi, Mohammad Abuzar, et al.
Published: (2021)
by: Hashemi, Mohammad Abuzar, et al.
Published: (2021)
FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions
by: Qin, Bowen, et al.
Published: (2025)
by: Qin, Bowen, et al.
Published: (2025)
Efficient Personalized Text-to-image Generation by Leveraging Textual Subspace
by: Du, Shian, et al.
Published: (2024)
by: Du, Shian, et al.
Published: (2024)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
by: Song, Wei, et al.
Published: (2025)
by: Song, Wei, et al.
Published: (2025)
VisText-Mosquito: A Unified Multimodal Dataset for Visual Detection, Segmentation, and Textual Explanation on Mosquito Breeding Sites
by: Islam, Md. Adnanul, et al.
Published: (2025)
by: Islam, Md. Adnanul, et al.
Published: (2025)
The Cost of Context: Mitigating Textual Bias in Multimodal Retrieval-Augmented Generation
by: Jung, Hoin, et al.
Published: (2026)
by: Jung, Hoin, et al.
Published: (2026)
Hierarchical Textual Knowledge for Enhanced Image Clustering
by: Zhong, Yijie, et al.
Published: (2026)
by: Zhong, Yijie, et al.
Published: (2026)
Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing
by: Yuan, Fan, et al.
Published: (2025)
by: Yuan, Fan, et al.
Published: (2025)
A Similarity Paradigm Through Textual Regularization Without Forgetting
by: Cui, Fangming, et al.
Published: (2025)
by: Cui, Fangming, et al.
Published: (2025)
TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Models
by: Ye, Jinlun, et al.
Published: (2026)
by: Ye, Jinlun, et al.
Published: (2026)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
by: Gan, Woody Haosheng, et al.
Published: (2025)
by: Gan, Woody Haosheng, et al.
Published: (2025)
LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLP
by: Chen, Danlu, et al.
Published: (2024)
by: Chen, Danlu, et al.
Published: (2024)
Beyond Meme Templates: Limitations of Visual Similarity Measures in Meme Matching
by: Hazman, Muzhaffar, et al.
Published: (2025)
by: Hazman, Muzhaffar, et al.
Published: (2025)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
by: Hua, Jiacheng, et al.
Published: (2026)
by: Hua, Jiacheng, et al.
Published: (2026)
On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI
by: Restrepo, David, et al.
Published: (2025)
by: Restrepo, David, et al.
Published: (2025)
TRACE: Textual Relevance Augmentation and Contextual Encoding for Multimodal Hate Detection
by: Koushik, Girish A., et al.
Published: (2025)
by: Koushik, Girish A., et al.
Published: (2025)
Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions
by: Pi, Renjie, et al.
Published: (2024)
by: Pi, Renjie, et al.
Published: (2024)
Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support Setting
by: Kayser, Maxime, et al.
Published: (2024)
by: Kayser, Maxime, et al.
Published: (2024)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
by: Wu, Yin, et al.
Published: (2025)
by: Wu, Yin, et al.
Published: (2025)
Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
by: Verma, Gaurav, et al.
Published: (2024)
by: Verma, Gaurav, et al.
Published: (2024)
How Does the Textual Information Affect the Retrieval of Multimodal In-Context Learning?
by: Luo, Yang, et al.
Published: (2024)
by: Luo, Yang, et al.
Published: (2024)
Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
by: Agrawal, Aakriti, et al.
Published: (2025)
by: Agrawal, Aakriti, et al.
Published: (2025)
Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization
by: Zhao, Zhiyuan, et al.
Published: (2023)
by: Zhao, Zhiyuan, et al.
Published: (2023)
Visually Guided Generative Text-Layout Pre-training for Document Intelligence
by: Mao, Zhiming, et al.
Published: (2024)
by: Mao, Zhiming, et al.
Published: (2024)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
by: Bai, Tianyi, et al.
Published: (2025)
by: Bai, Tianyi, et al.
Published: (2025)
Motion Generation from Fine-grained Textual Descriptions
by: Li, Kunhang, et al.
Published: (2024)
by: Li, Kunhang, et al.
Published: (2024)
Beyond Static Cropping: Layer-Adaptive Visual Localization and Decoding Enhancement
by: Zhu, Zipeng, et al.
Published: (2026)
by: Zhu, Zipeng, et al.
Published: (2026)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
by: Wang, Yabing, et al.
Published: (2024)
by: Wang, Yabing, et al.
Published: (2024)
Overcoming Language Priors for Visual Question Answering Based on Knowledge Distillation
by: Peng, Daowan, et al.
Published: (2025)
by: Peng, Daowan, et al.
Published: (2025)
Mojito: Motion Trajectory and Intensity Control for Video Generation
by: He, Xuehai, et al.
Published: (2024)
by: He, Xuehai, et al.
Published: (2024)
Similar Items
-
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
by: Wang, Yifan, et al.
Published: (2026) -
CLEAR: Character Unlearning in Textual and Visual Modalities
by: Dontsov, Alexey, et al.
Published: (2024) -
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
by: Guo, Ziyu, et al.
Published: (2025) -
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
by: Gordon, Brian, et al.
Published: (2023) -
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
by: Sun, Hao, et al.
Published: (2026)