SmoGVLM: A Small, Graph-enhanced Vision-Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | Mondal, Debjyoti, Singh, Rituraj, Panda, Subhadarshi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning
by: Mondal, Debjyoti, et al.
Published: (2024)
by: Mondal, Debjyoti, et al.
Published: (2024)
Can Safety Emerge from Weak Supervision? A Systematic Analysis of Small Language Models
by: Saha, Punyajoy, et al.
Published: (2026)
by: Saha, Punyajoy, et al.
Published: (2026)
Vision Meets Language: A RAG-Augmented YOLOv8 Framework for Coffee Disease Diagnosis and Farmer Assistance
by: Mondal, Semanto
Published: (2025)
by: Mondal, Semanto
Published: (2025)
Seg-HGNN: Unsupervised and Light-Weight Image Segmentation with Hyperbolic Graph Neural Networks
by: Mondal, Debjyoti, et al.
Published: (2024)
by: Mondal, Debjyoti, et al.
Published: (2024)
Small Vision-Language Models: A Survey on Compact Architectures and Techniques
by: Patnaik, Nitesh, et al.
Published: (2025)
by: Patnaik, Nitesh, et al.
Published: (2025)
Open World Scene Graph Generation using Vision Language Models
by: Dutta, Amartya, et al.
Published: (2025)
by: Dutta, Amartya, et al.
Published: (2025)
iGVLM: Dynamic Instruction-Guided Vision Encoding for Question-Aware Multimodal Understanding
by: Liu, Hanpeng, et al.
Published: (2026)
by: Liu, Hanpeng, et al.
Published: (2026)
HazardNet: A Small-Scale Vision Language Model for Real-Time Traffic Safety Detection at Edge Devices
by: Tami, Mohammad Abu, et al.
Published: (2025)
by: Tami, Mohammad Abu, et al.
Published: (2025)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
by: Zhu, Yingjie, et al.
Published: (2024)
by: Zhu, Yingjie, et al.
Published: (2024)
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
by: Lee, Seongyun, et al.
Published: (2024)
by: Lee, Seongyun, et al.
Published: (2024)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
by: Ye, Jiacheng, et al.
Published: (2025)
by: Ye, Jiacheng, et al.
Published: (2025)
Conflict Adaptation in Vision-Language Models
by: Hu, Xiaoyang
Published: (2025)
by: Hu, Xiaoyang
Published: (2025)
Vision Language Models are Confused Tourists
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
Mipha: A Comprehensive Overhaul of Multimodal Assistant with Small Language Models
by: Zhu, Minjie, et al.
Published: (2024)
by: Zhu, Minjie, et al.
Published: (2024)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
by: Goetting, Dylan, et al.
Published: (2024)
by: Goetting, Dylan, et al.
Published: (2024)
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models
by: Cai, Hengxing, et al.
Published: (2025)
by: Cai, Hengxing, et al.
Published: (2025)
Natural Language Inference Improves Compositionality in Vision-Language Models
by: Cascante-Bonilla, Paola, et al.
Published: (2024)
by: Cascante-Bonilla, Paola, et al.
Published: (2024)
Do Vision-Language Models Really Understand Visual Language?
by: Hou, Yifan, et al.
Published: (2024)
by: Hou, Yifan, et al.
Published: (2024)
PUMGPT: A Large Vision-Language Model for Product Understanding
by: Xue, Wei, et al.
Published: (2023)
by: Xue, Wei, et al.
Published: (2023)
Vision-Language Models Do Not Understand Negation
by: Alhamoud, Kumail, et al.
Published: (2025)
by: Alhamoud, Kumail, et al.
Published: (2025)
Text Prompt Injection of Vision Language Models
by: Zhu, Ruizhe
Published: (2025)
by: Zhu, Ruizhe
Published: (2025)
Evaluating Vision-Language Models for Emotion Recognition
by: Bhattacharyya, Sree, et al.
Published: (2025)
by: Bhattacharyya, Sree, et al.
Published: (2025)
Intriguing Properties of Large Language and Vision Models
by: Lee, Young-Jun, et al.
Published: (2024)
by: Lee, Young-Jun, et al.
Published: (2024)
Evaluation of Cultural Competence of Vision-Language Models
by: Yadav, Srishti, et al.
Published: (2025)
by: Yadav, Srishti, et al.
Published: (2025)
Vision Language Models Are Not (Yet) Spelling Correctors
by: Liang, Junhong, et al.
Published: (2025)
by: Liang, Junhong, et al.
Published: (2025)
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
by: Kang, Jialiang, et al.
Published: (2025)
by: Kang, Jialiang, et al.
Published: (2025)
Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
by: Xu, Yige, et al.
Published: (2026)
by: Xu, Yige, et al.
Published: (2026)
CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models
by: Castro, Santiago, et al.
Published: (2024)
by: Castro, Santiago, et al.
Published: (2024)
ViLBench: A Suite for Vision-Language Process Reward Modeling
by: Tu, Haoqin, et al.
Published: (2025)
by: Tu, Haoqin, et al.
Published: (2025)
Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
by: Hennara, Khalil, et al.
Published: (2025)
by: Hennara, Khalil, et al.
Published: (2025)
Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese
by: Inoue, Yuichi, et al.
Published: (2024)
by: Inoue, Yuichi, et al.
Published: (2024)
A Unified Hallucination Mitigation Framework for Large Vision-Language Models
by: Chang, Yue, et al.
Published: (2024)
by: Chang, Yue, et al.
Published: (2024)
Diagnosing Vision Language Models' Perception by Leveraging Human Methods for Color Vision Deficiencies
by: Hayashi, Kazuki, et al.
Published: (2025)
by: Hayashi, Kazuki, et al.
Published: (2025)
Towards Zero-Shot Annotation of the Built Environment with Vision-Language Models (Vision Paper)
by: Han, Bin, et al.
Published: (2024)
by: Han, Bin, et al.
Published: (2024)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
by: Zhu, Yingjie, et al.
Published: (2025)
by: Zhu, Yingjie, et al.
Published: (2025)
Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
by: Wang, Weihang, et al.
Published: (2025)
by: Wang, Weihang, et al.
Published: (2025)
PROGRESSLM: Towards Progress Reasoning in Vision-Language Models
by: Zhang, Jianshu, et al.
Published: (2026)
by: Zhang, Jianshu, et al.
Published: (2026)
Can Vision-Language Models Solve the Shell Game?
by: Liu, Tiedong, et al.
Published: (2026)
by: Liu, Tiedong, et al.
Published: (2026)
Evaluating Vision-Language Models as Evaluators in Path Planning
by: Aghzal, Mohamed, et al.
Published: (2024)
by: Aghzal, Mohamed, et al.
Published: (2024)
Inference Compute-Optimal Video Vision Language Models
by: Wang, Peiqi, et al.
Published: (2025)
by: Wang, Peiqi, et al.
Published: (2025)
Similar Items
-
KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning
by: Mondal, Debjyoti, et al.
Published: (2024) -
Can Safety Emerge from Weak Supervision? A Systematic Analysis of Small Language Models
by: Saha, Punyajoy, et al.
Published: (2026) -
Vision Meets Language: A RAG-Augmented YOLOv8 Framework for Coffee Disease Diagnosis and Farmer Assistance
by: Mondal, Semanto
Published: (2025) -
Seg-HGNN: Unsupervised and Light-Weight Image Segmentation with Hyperbolic Graph Neural Networks
by: Mondal, Debjyoti, et al.
Published: (2024) -
Small Vision-Language Models: A Survey on Compact Architectures and Techniques
by: Patnaik, Nitesh, et al.
Published: (2025)