ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mi, Li, Montariol, Syrielle, Castillo-Navarro, Javiera, Dai, Xianjie, Bosselut, Antoine, Tuia, Devis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ConGeo: Robust Cross-view Geo-localization across Ground View Variations
von: Mi, Li, et al.
Veröffentlicht: (2024)
von: Mi, Li, et al.
Veröffentlicht: (2024)
Knowledge-aware Text-Image Retrieval for Remote Sensing Images
von: Mi, Li, et al.
Veröffentlicht: (2024)
von: Mi, Li, et al.
Veröffentlicht: (2024)
Questions beyond Pixels: Integrating Commonsense Knowledge in Visual Question Generation for Remote Sensing
von: Li, Siran, et al.
Veröffentlicht: (2026)
von: Li, Siran, et al.
Veröffentlicht: (2026)
Knowledge-aware Visual Question Generation for Remote Sensing Images
von: Li, Siran, et al.
Veröffentlicht: (2026)
von: Li, Siran, et al.
Veröffentlicht: (2026)
Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
von: Corbière, Charles, et al.
Veröffentlicht: (2025)
von: Corbière, Charles, et al.
Veröffentlicht: (2025)
GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration
von: Mi, Li, et al.
Veröffentlicht: (2025)
von: Mi, Li, et al.
Veröffentlicht: (2025)
Checkmate: interpretable and explainable RSVQA is the endgame
von: Tosato, Lucrezia, et al.
Veröffentlicht: (2025)
von: Tosato, Lucrezia, et al.
Veröffentlicht: (2025)
VinaBench: Benchmark for Faithful and Consistent Visual Narratives
von: Gao, Silin, et al.
Veröffentlicht: (2025)
von: Gao, Silin, et al.
Veröffentlicht: (2025)
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
von: Bhagwatkar, Rishika, et al.
Veröffentlicht: (2025)
von: Bhagwatkar, Rishika, et al.
Veröffentlicht: (2025)
What to align in multimodal contrastive learning?
von: Dufumier, Benoit, et al.
Veröffentlicht: (2024)
von: Dufumier, Benoit, et al.
Veröffentlicht: (2024)
EcoWikiRS: Learning Ecological Representation of Satellite Images from Weak Supervision with Species Observations and Wikipedia
von: Zermatten, Valerie, et al.
Veröffentlicht: (2025)
von: Zermatten, Valerie, et al.
Veröffentlicht: (2025)
Cross-Modal Learning of Housing Quality in Amsterdam
von: Levering, Alex, et al.
Veröffentlicht: (2024)
von: Levering, Alex, et al.
Veröffentlicht: (2024)
Multi-Scale Grouped Prototypes for Interpretable Semantic Segmentation
von: Porta, Hugo, et al.
Veröffentlicht: (2024)
von: Porta, Hugo, et al.
Veröffentlicht: (2024)
ECIS-VQG: Generation of Entity-centric Information-seeking Questions from Videos
von: Phukan, Arpan, et al.
Veröffentlicht: (2024)
von: Phukan, Arpan, et al.
Veröffentlicht: (2024)
PICLe: Pseudo-Annotations for In-Context Learning in Low-Resource Named Entity Detection
von: Mamooler, Sepideh, et al.
Veröffentlicht: (2024)
von: Mamooler, Sepideh, et al.
Veröffentlicht: (2024)
Retrieval of Surface Solar Radiation through Implicit Albedo Recovery from Temporal Context
von: Frischholz, Yael, et al.
Veröffentlicht: (2025)
von: Frischholz, Yael, et al.
Veröffentlicht: (2025)
SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images
von: Sumbul, Gencer, et al.
Veröffentlicht: (2025)
von: Sumbul, Gencer, et al.
Veröffentlicht: (2025)
POLO -- Point-based, multi-class animal detection
von: May, Giacomo, et al.
Veröffentlicht: (2024)
von: May, Giacomo, et al.
Veröffentlicht: (2024)
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
von: Park, Yeji, et al.
Veröffentlicht: (2024)
von: Park, Yeji, et al.
Veröffentlicht: (2024)
Multilingual Vision-Language Pre-training for the Remote Sensing Domain
von: Silva, João Daniel, et al.
Veröffentlicht: (2024)
von: Silva, João Daniel, et al.
Veröffentlicht: (2024)
An Efficient and Effective Encoder Model for Vision and Language Tasks in the Remote Sensing Domain
von: Silva, João Daniel, et al.
Veröffentlicht: (2025)
von: Silva, João Daniel, et al.
Veröffentlicht: (2025)
Large Language Models for Captioning and Retrieving Remote Sensing Images
von: Silva, João Daniel, et al.
Veröffentlicht: (2024)
von: Silva, João Daniel, et al.
Veröffentlicht: (2024)
High-resolution Population Maps Derived from Sentinel-1 and Sentinel-2
von: Metzger, Nando, et al.
Veröffentlicht: (2023)
von: Metzger, Nando, et al.
Veröffentlicht: (2023)
CanadaFireSat: Toward high-resolution wildfire forecasting with multiple modalities
von: Porta, Hugo, et al.
Veröffentlicht: (2025)
von: Porta, Hugo, et al.
Veröffentlicht: (2025)
JOBSKAPE: A Framework for Generating Synthetic Job Postings to Enhance Skill Matching
von: Magron, Antoine, et al.
Veröffentlicht: (2024)
von: Magron, Antoine, et al.
Veröffentlicht: (2024)
TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation
von: Litrico, Mattia, et al.
Veröffentlicht: (2025)
von: Litrico, Mattia, et al.
Veröffentlicht: (2025)
Visual Generation Without Guidance
von: Chen, Huayu, et al.
Veröffentlicht: (2025)
von: Chen, Huayu, et al.
Veröffentlicht: (2025)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
von: Choi, Changin, et al.
Veröffentlicht: (2025)
von: Choi, Changin, et al.
Veröffentlicht: (2025)
Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
From Classification to Segmentation with Explainable AI: A Study on Crack Detection and Growth Monitoring
von: Forest, Florent, et al.
Veröffentlicht: (2023)
von: Forest, Florent, et al.
Veröffentlicht: (2023)
FlexiTex: Enhancing Texture Generation via Visual Guidance
von: Jiang, DaDong, et al.
Veröffentlicht: (2024)
von: Jiang, DaDong, et al.
Veröffentlicht: (2024)
Multi-Sourced Compositional Generalization in Visual Question Answering
von: Li, Chuanhao, et al.
Veröffentlicht: (2025)
von: Li, Chuanhao, et al.
Veröffentlicht: (2025)
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
von: Wu, Fengyi, et al.
Veröffentlicht: (2025)
von: Wu, Fengyi, et al.
Veröffentlicht: (2025)
Benchmarking Large Multimodal Models for Ophthalmic Visual Question Answering with OphthalWeChat
von: Xu, Pusheng, et al.
Veröffentlicht: (2025)
von: Xu, Pusheng, et al.
Veröffentlicht: (2025)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
von: Kang, Wonjun, et al.
Veröffentlicht: (2025)
von: Kang, Wonjun, et al.
Veröffentlicht: (2025)
ConEQsA: Concurrent and Asynchronous Embodied Questions Scheduling and Answering
von: Wang, Haisheng, et al.
Veröffentlicht: (2025)
von: Wang, Haisheng, et al.
Veröffentlicht: (2025)
Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
von: Pang, Yuqi, et al.
Veröffentlicht: (2025)
von: Pang, Yuqi, et al.
Veröffentlicht: (2025)
Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning
von: Zeng, Xingchen, et al.
Veröffentlicht: (2024)
von: Zeng, Xingchen, et al.
Veröffentlicht: (2024)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2024)
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2024)
HCVP: Leveraging Hierarchical Contrastive Visual Prompt for Domain Generalization
von: Zhou, Guanglin, et al.
Veröffentlicht: (2024)
von: Zhou, Guanglin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ConGeo: Robust Cross-view Geo-localization across Ground View Variations
von: Mi, Li, et al.
Veröffentlicht: (2024) -
Knowledge-aware Text-Image Retrieval for Remote Sensing Images
von: Mi, Li, et al.
Veröffentlicht: (2024) -
Questions beyond Pixels: Integrating Commonsense Knowledge in Visual Question Generation for Remote Sensing
von: Li, Siran, et al.
Veröffentlicht: (2026) -
Knowledge-aware Visual Question Generation for Remote Sensing Images
von: Li, Siran, et al.
Veröffentlicht: (2026) -
Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
von: Corbière, Charles, et al.
Veröffentlicht: (2025)