BRAVE: Broadening the visual encoding of vision-language models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kar, Oğuzhan Fatih, Tonioni, Alessio, Poklukar, Petra, Kulshrestha, Achin, Zamir, Amir, Tombari, Federico |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
von: Plizzari, Chiara, et al.
Veröffentlicht: (2025)
von: Plizzari, Chiara, et al.
Veröffentlicht: (2025)
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
von: Ramachandran, Rahul, et al.
Veröffentlicht: (2025)
von: Ramachandran, Rahul, et al.
Veröffentlicht: (2025)
MapTrace: Scalable Data Generation for Route Tracing on Maps
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2025)
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2025)
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
von: Bachmann, Roman, et al.
Veröffentlicht: (2024)
von: Bachmann, Roman, et al.
Veröffentlicht: (2024)
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
von: Kuzucu, Selim, et al.
Veröffentlicht: (2026)
von: Kuzucu, Selim, et al.
Veröffentlicht: (2026)
Quantifying the human visual exposome with vision language models
von: Rominger, Christian, et al.
Veröffentlicht: (2026)
von: Rominger, Christian, et al.
Veröffentlicht: (2026)
(1D) Ordered Tokens Enable Efficient Test-Time Search
von: Gao, Zhitong, et al.
Veröffentlicht: (2026)
von: Gao, Zhitong, et al.
Veröffentlicht: (2026)
FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation
von: Bill, Eric Tillmann, et al.
Veröffentlicht: (2026)
von: Bill, Eric Tillmann, et al.
Veröffentlicht: (2026)
Reading Between the Lines: Abstaining from VLM-Generated OCR Errors via Latent Representation Probes
von: Yao, Jihan, et al.
Veröffentlicht: (2025)
von: Yao, Jihan, et al.
Veröffentlicht: (2025)
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs
von: Xie, Jiahao, et al.
Veröffentlicht: (2026)
von: Xie, Jiahao, et al.
Veröffentlicht: (2026)
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint
von: Simsar, Enis, et al.
Veröffentlicht: (2024)
von: Simsar, Enis, et al.
Veröffentlicht: (2024)
LIME: Localized Image Editing via Attention Regularization in Diffusion Models
von: Simsar, Enis, et al.
Veröffentlicht: (2023)
von: Simsar, Enis, et al.
Veröffentlicht: (2023)
Text-Conditioned Resampler For Long Form Video Understanding
von: Korbar, Bruno, et al.
Veröffentlicht: (2023)
von: Korbar, Bruno, et al.
Veröffentlicht: (2023)
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models
von: Xie, Jiahao, et al.
Veröffentlicht: (2026)
von: Xie, Jiahao, et al.
Veröffentlicht: (2026)
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
von: Kar, Oğuzhan Fatih, et al.
Veröffentlicht: (2026)
von: Kar, Oğuzhan Fatih, et al.
Veröffentlicht: (2026)
Are vision language models robust to uncertain inputs?
von: Wang, Xi, et al.
Veröffentlicht: (2025)
von: Wang, Xi, et al.
Veröffentlicht: (2025)
What matters when building vision-language models?
von: Laurençon, Hugo, et al.
Veröffentlicht: (2024)
von: Laurençon, Hugo, et al.
Veröffentlicht: (2024)
Test-Time Visual In-Context Tuning
von: Xie, Jiahao, et al.
Veröffentlicht: (2025)
von: Xie, Jiahao, et al.
Veröffentlicht: (2025)
Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
Thinker: A vision-language foundation model for embodied intelligence
von: Pan, Baiyu, et al.
Veröffentlicht: (2026)
von: Pan, Baiyu, et al.
Veröffentlicht: (2026)
Building and better understanding vision-language models: insights and future directions
von: Laurençon, Hugo, et al.
Veröffentlicht: (2024)
von: Laurençon, Hugo, et al.
Veröffentlicht: (2024)
Hallucination-aware intermediate representation edit in large vision-language models
von: Suo, Wei, et al.
Veröffentlicht: (2026)
von: Suo, Wei, et al.
Veröffentlicht: (2026)
Generalizing vision-language models to novel domains: A comprehensive survey
von: Li, Xinyao, et al.
Veröffentlicht: (2025)
von: Li, Xinyao, et al.
Veröffentlicht: (2025)
Exploring visual language models as a powerful tool in the diagnosis of Ewing Sarcoma
von: Pastor-Naranjo, Alvaro, et al.
Veröffentlicht: (2025)
von: Pastor-Naranjo, Alvaro, et al.
Veröffentlicht: (2025)
Vision language models are blind: Failing to translate detailed visual features into words
von: Rahmanzadehgervi, Pooyan, et al.
Veröffentlicht: (2024)
von: Rahmanzadehgervi, Pooyan, et al.
Veröffentlicht: (2024)
A benchmark multimodal oro-dental dataset for large vision-language models
von: Lv, Haoxin, et al.
Veröffentlicht: (2025)
von: Lv, Haoxin, et al.
Veröffentlicht: (2025)
Representation geometry shapes task performance in vision-language modeling for CT enterography
von: Minoccheri, Cristian, et al.
Veröffentlicht: (2026)
von: Minoccheri, Cristian, et al.
Veröffentlicht: (2026)
Beyond the Hype: A dispassionate look at vision-language models in medical scenario
von: Nan, Yang, et al.
Veröffentlicht: (2024)
von: Nan, Yang, et al.
Veröffentlicht: (2024)
VLA-Mark: A cross modal watermark for large vision-language alignment model
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
Self-adaptive vision-language model for 3D segmentation of pulmonary artery and vein
von: Guo, Xiaotong, et al.
Veröffentlicht: (2025)
von: Guo, Xiaotong, et al.
Veröffentlicht: (2025)
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
von: Shi, Danli, et al.
Veröffentlicht: (2024)
von: Shi, Danli, et al.
Veröffentlicht: (2024)
Pic2Diagnosis: A Method for Diagnosis of Cardiovascular Diseases from the Printed ECG Pictures
von: Büyüksolak, Oğuzhan, et al.
Veröffentlicht: (2025)
von: Büyüksolak, Oğuzhan, et al.
Veröffentlicht: (2025)
Improving vision-language alignment with graph spiking hybrid Networks
von: Zhang, Siyu, et al.
Veröffentlicht: (2025)
von: Zhang, Siyu, et al.
Veröffentlicht: (2025)
Zero-shot large vision-language model prompting for automated bone identification in paleoradiology x-ray archives
von: Dong, Owen, et al.
Veröffentlicht: (2026)
von: Dong, Owen, et al.
Veröffentlicht: (2026)
MI-VisionShot: Few-shot adaptation of vision-language models for slide-level classification of histopathological images
von: Meseguer, Pablo, et al.
Veröffentlicht: (2024)
von: Meseguer, Pablo, et al.
Veröffentlicht: (2024)
UNBOX: Unveiling Black-box visual models with Natural-language
von: Carnemolla, Simone, et al.
Veröffentlicht: (2026)
von: Carnemolla, Simone, et al.
Veröffentlicht: (2026)
Owls are wise and foxes are unfaithful: Uncovering animal stereotypes in vision-language models
von: Aman, Tabinda, et al.
Veröffentlicht: (2025)
von: Aman, Tabinda, et al.
Veröffentlicht: (2025)
ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2025)
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2025)
MGA-VQA: Secure and Interpretable Graph-Augmented Visual Question Answering with Memory-Guided Protection Against Unauthorized Knowledge Use
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2025)
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2025)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
von: Xing, Yang, et al.
Veröffentlicht: (2026)
von: Xing, Yang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
von: Plizzari, Chiara, et al.
Veröffentlicht: (2025) -
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
von: Ramachandran, Rahul, et al.
Veröffentlicht: (2025) -
MapTrace: Scalable Data Generation for Route Tracing on Maps
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2025) -
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
von: Bachmann, Roman, et al.
Veröffentlicht: (2024) -
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
von: Kuzucu, Selim, et al.
Veröffentlicht: (2026)