GeoDANO: Geometric VLM with Domain Agnostic Vision Encoder
Fuente:
arXiv
Salvato in:
| Autori principali: | Cho, Seunghyuk, Qin, Zhenyue, Liu, Yang, Choi, Youngbin, Lee, Seungbeom, Kim, Dongwoo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey
di: Cho, Seunghyuk, et al.
Pubblicazione: (2025)
di: Cho, Seunghyuk, et al.
Pubblicazione: (2025)
Feature Unlearning for Pre-trained GANs and VAEs
di: Moon, Saemi, et al.
Pubblicazione: (2023)
di: Moon, Saemi, et al.
Pubblicazione: (2023)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
di: Zhu, Yingjie, et al.
Pubblicazione: (2025)
di: Zhu, Yingjie, et al.
Pubblicazione: (2025)
Modality-Agnostic fMRI Decoding of Vision and Language
di: Nikolaus, Mitja, et al.
Pubblicazione: (2024)
di: Nikolaus, Mitja, et al.
Pubblicazione: (2024)
OViP: Online Vision-Language Preference Learning for VLM Hallucination
di: Liu, Shujun, et al.
Pubblicazione: (2025)
di: Liu, Shujun, et al.
Pubblicazione: (2025)
Are Bigger Encoders Always Better in Vision Large Models?
di: Li, Bozhou, et al.
Pubblicazione: (2024)
di: Li, Bozhou, et al.
Pubblicazione: (2024)
From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
di: Sheta, Hala, et al.
Pubblicazione: (2025)
di: Sheta, Hala, et al.
Pubblicazione: (2025)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
di: Zhang, Di, et al.
Pubblicazione: (2024)
di: Zhang, Di, et al.
Pubblicazione: (2024)
Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays
di: Cho, Yeongjae, et al.
Pubblicazione: (2024)
di: Cho, Yeongjae, et al.
Pubblicazione: (2024)
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
di: Xia, Renqiu, et al.
Pubblicazione: (2024)
di: Xia, Renqiu, et al.
Pubblicazione: (2024)
Intriguing Properties of Large Language and Vision Models
di: Lee, Young-Jun, et al.
Pubblicazione: (2024)
di: Lee, Young-Jun, et al.
Pubblicazione: (2024)
Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
di: Hu, Zhe, et al.
Pubblicazione: (2025)
di: Hu, Zhe, et al.
Pubblicazione: (2025)
SCoPE VLM: Selective Context Processing for Efficient Document Navigation in Vision-Language Models
di: Lim, Gyubeum, et al.
Pubblicazione: (2025)
di: Lim, Gyubeum, et al.
Pubblicazione: (2025)
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
di: Fan, Zhiwen, et al.
Pubblicazione: (2025)
di: Fan, Zhiwen, et al.
Pubblicazione: (2025)
Rethinking the Mixture of Vision Encoders Paradigm for Enhanced Visual Understanding in Multimodal LLMs
di: Azadani, Mozhgan Nasr, et al.
Pubblicazione: (2025)
di: Azadani, Mozhgan Nasr, et al.
Pubblicazione: (2025)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
di: Wang, Zihu, et al.
Pubblicazione: (2025)
di: Wang, Zihu, et al.
Pubblicazione: (2025)
ESREAL: Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models
di: Kim, Minchan, et al.
Pubblicazione: (2024)
di: Kim, Minchan, et al.
Pubblicazione: (2024)
Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
di: Zhang, Boqiang, et al.
Pubblicazione: (2026)
di: Zhang, Boqiang, et al.
Pubblicazione: (2026)
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
di: Lee, Seongyun, et al.
Pubblicazione: (2024)
di: Lee, Seongyun, et al.
Pubblicazione: (2024)
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
di: Shen, Haozhan, et al.
Pubblicazione: (2025)
di: Shen, Haozhan, et al.
Pubblicazione: (2025)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
di: Liu, Zheng, et al.
Pubblicazione: (2024)
di: Liu, Zheng, et al.
Pubblicazione: (2024)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
di: Kim, Wonkyun, et al.
Pubblicazione: (2024)
di: Kim, Wonkyun, et al.
Pubblicazione: (2024)
MambaEye: A Size-Agnostic Visual Encoder with Causal Sequential Processing
di: Choi, Changho, et al.
Pubblicazione: (2025)
di: Choi, Changho, et al.
Pubblicazione: (2025)
PersonaVLM: Long-Term Personalized Multimodal LLMs
di: Nie, Chang, et al.
Pubblicazione: (2026)
di: Nie, Chang, et al.
Pubblicazione: (2026)
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis
di: Lokesh, K, et al.
Pubblicazione: (2026)
di: Lokesh, K, et al.
Pubblicazione: (2026)
VolDoGer: LLM-assisted Datasets for Domain Generalization in Vision-Language Tasks
di: Choi, Juhwan, et al.
Pubblicazione: (2024)
di: Choi, Juhwan, et al.
Pubblicazione: (2024)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
di: Huang, Brandon, et al.
Pubblicazione: (2025)
di: Huang, Brandon, et al.
Pubblicazione: (2025)
ReVision: A Dataset and Baseline VLM for Privacy-Preserving Task-Oriented Visual Instruction Rewriting
di: Mishra, Abhijit, et al.
Pubblicazione: (2025)
di: Mishra, Abhijit, et al.
Pubblicazione: (2025)
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
di: Wang, Zekun, et al.
Pubblicazione: (2025)
di: Wang, Zekun, et al.
Pubblicazione: (2025)
VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
di: Meng, Rui, et al.
Pubblicazione: (2025)
di: Meng, Rui, et al.
Pubblicazione: (2025)
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
di: Jiang, Ziyan, et al.
Pubblicazione: (2024)
di: Jiang, Ziyan, et al.
Pubblicazione: (2024)
Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration
di: Park, ChaeHun, et al.
Pubblicazione: (2024)
di: Park, ChaeHun, et al.
Pubblicazione: (2024)
Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language Models
di: Bachu, Saketh, et al.
Pubblicazione: (2024)
di: Bachu, Saketh, et al.
Pubblicazione: (2024)
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding
di: Cho, Beomsik, et al.
Pubblicazione: (2025)
di: Cho, Beomsik, et al.
Pubblicazione: (2025)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
di: Kim, Mingyeong, et al.
Pubblicazione: (2026)
di: Kim, Mingyeong, et al.
Pubblicazione: (2026)
Long Live the Librarian! A Persistent Search Sub-Agent for Energy-Efficient Multi-Agent Software Engineering Systems
di: Cho, Seunghyuk, et al.
Pubblicazione: (2026)
di: Cho, Seunghyuk, et al.
Pubblicazione: (2026)
GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning
di: Liu, Zhaochen, et al.
Pubblicazione: (2026)
di: Liu, Zhaochen, et al.
Pubblicazione: (2026)
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation
di: Cai, Shihao, et al.
Pubblicazione: (2024)
di: Cai, Shihao, et al.
Pubblicazione: (2024)
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
di: Oh, Hongseok, et al.
Pubblicazione: (2025)
di: Oh, Hongseok, et al.
Pubblicazione: (2025)
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
di: Ebouky, Brown, et al.
Pubblicazione: (2026)
di: Ebouky, Brown, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey
di: Cho, Seunghyuk, et al.
Pubblicazione: (2025) -
Feature Unlearning for Pre-trained GANs and VAEs
di: Moon, Saemi, et al.
Pubblicazione: (2023) -
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
di: Zhu, Yingjie, et al.
Pubblicazione: (2025) -
Modality-Agnostic fMRI Decoding of Vision and Language
di: Nikolaus, Mitja, et al.
Pubblicazione: (2024) -
OViP: Online Vision-Language Preference Learning for VLM Hallucination
di: Liu, Shujun, et al.
Pubblicazione: (2025)