GraphVL: Graph-Enhanced Semantic Modeling via Vision-Language Models for Generalized Class Discovery
Fuente:
arXiv
Saved in:
| Main Authors: | Solanki, Bhupendra, Nair, Ashwin, Singha, Mainak, Mukhopadhyay, Souradeep, Jha, Ankit, Banerjee, Biplab |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unknown Prompt, the only Lacuna: Unveiling CLIP's Potential for Open Domain Generalization
by: Singha, Mainak, et al.
Published: (2024)
by: Singha, Mainak, et al.
Published: (2024)
AD-CLIP: Adapting Domains in Prompt Space Using CLIP
by: Singha, Mainak, et al.
Published: (2023)
by: Singha, Mainak, et al.
Published: (2023)
FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
by: Singha, Mainak, et al.
Published: (2025)
by: Singha, Mainak, et al.
Published: (2025)
Elevating All Zero-Shot Sketch-Based Image Retrieval Through Multimodal Prompt Learning
by: Singha, Mainak, et al.
Published: (2024)
by: Singha, Mainak, et al.
Published: (2024)
BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs
by: Singha, Mainak, et al.
Published: (2026)
by: Singha, Mainak, et al.
Published: (2026)
MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP
by: Chaudhary, Aditya, et al.
Published: (2026)
by: Chaudhary, Aditya, et al.
Published: (2026)
CDAD-Net: Bridging Domain Gaps in Generalized Category Discovery
by: Rongali, Sai Bhargav, et al.
Published: (2024)
by: Rongali, Sai Bhargav, et al.
Published: (2024)
bi-modal textual prompt learning for vision-language models in remote sensing
by: Kashyap, Pankhi, et al.
Published: (2026)
by: Kashyap, Pankhi, et al.
Published: (2026)
COSMo: CLIP Talks on Open-Set Multi-Target Domain Adaptation
by: Monga, Munish, et al.
Published: (2024)
by: Monga, Munish, et al.
Published: (2024)
SDHSI-Net: Learning Better Representations for Hyperspectral Images via Self-Distillation
by: Singh, Prachet Dev, et al.
Published: (2026)
by: Singh, Prachet Dev, et al.
Published: (2026)
Reconstruction Guided Few-shot Network For Remote Sensing Image Classification
by: Jaiswal, Mohit, et al.
Published: (2026)
by: Jaiswal, Mohit, et al.
Published: (2026)
OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIP
by: C, Mohamad Hassan N, et al.
Published: (2025)
by: C, Mohamad Hassan N, et al.
Published: (2025)
In the Era of Prompt Learning with Vision-Language Models
by: Jha, Ankit
Published: (2024)
by: Jha, Ankit
Published: (2024)
Revisiting KRISP: A Lightweight Reproduction and Analysis of Knowledge-Enhanced Vision-Language Models
by: Dutta, Souradeep, et al.
Published: (2025)
by: Dutta, Souradeep, et al.
Published: (2025)
CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation
by: Singha, Mainak, et al.
Published: (2026)
by: Singha, Mainak, et al.
Published: (2026)
GeoMeld: Toward Semantically Grounded Foundation Models for Remote Sensing
by: Hasan, Maram, et al.
Published: (2026)
by: Hasan, Maram, et al.
Published: (2026)
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
by: Dong, Daxiang, et al.
Published: (2025)
by: Dong, Daxiang, et al.
Published: (2025)
HIDISC: A Hyperbolic Framework for Domain Generalization with Generalized Category Discovery
by: Rathore, Vaibhav, et al.
Published: (2025)
by: Rathore, Vaibhav, et al.
Published: (2025)
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
by: Chen, Jiuhai, et al.
Published: (2024)
by: Chen, Jiuhai, et al.
Published: (2024)
Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation
by: Trinh, Quoc-Huy, et al.
Published: (2026)
by: Trinh, Quoc-Huy, et al.
Published: (2026)
Few Shot Class Incremental Learning using Vision-Language models
by: Kumar, Anurag, et al.
Published: (2024)
by: Kumar, Anurag, et al.
Published: (2024)
Foundation Models and Adaptive Feature Selection: A Synergistic Approach to Video Question Answering
by: Rongali, Sai Bhargav, et al.
Published: (2024)
by: Rongali, Sai Bhargav, et al.
Published: (2024)
Discovery of a 13-Sharpe OOS Factor: Drift Regimes Unlock Hidden Cross-Sectional Predictability
by: Singha, Mainak
Published: (2025)
by: Singha, Mainak
Published: (2025)
Scaling Vision-and-Language Navigation With Offline RL
by: Bundele, Valay, et al.
Published: (2024)
by: Bundele, Valay, et al.
Published: (2024)
Innovator-VL: A Multimodal Large Language Model for Scientific Discovery
by: Wen, Zichen, et al.
Published: (2026)
by: Wen, Zichen, et al.
Published: (2026)
VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation
by: Zhang, Ruiyang, et al.
Published: (2024)
by: Zhang, Ruiyang, et al.
Published: (2024)
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
by: Zeng, Lunbin, et al.
Published: (2025)
by: Zeng, Lunbin, et al.
Published: (2025)
A-VL: Adaptive Attention for Large Vision-Language Models
by: Zhang, Junyang, et al.
Published: (2024)
by: Zhang, Junyang, et al.
Published: (2024)
VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training
by: Zhang, Jipeng, et al.
Published: (2025)
by: Zhang, Jipeng, et al.
Published: (2025)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
by: Wang, Peng, et al.
Published: (2024)
by: Wang, Peng, et al.
Published: (2024)
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
Thermo-VL: Extending Vision-Language Models to Thermal Infrared Perception
by: Thushara, Rusiru, et al.
Published: (2026)
by: Thushara, Rusiru, et al.
Published: (2026)
3VL: Using Trees to Improve Vision-Language Models' Interpretability
by: Yellinek, Nir, et al.
Published: (2023)
by: Yellinek, Nir, et al.
Published: (2023)
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
by: Wang, Shijing, et al.
Published: (2025)
by: Wang, Shijing, et al.
Published: (2025)
From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
by: Li, Rongjie, et al.
Published: (2024)
by: Li, Rongjie, et al.
Published: (2024)
GraSP-VL: Length as a Semantic Granularity Interface for Vision-Language Representations
by: Li, Zesheng, et al.
Published: (2026)
by: Li, Zesheng, et al.
Published: (2026)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
by: Ye, Jiacheng, et al.
Published: (2025)
by: Ye, Jiacheng, et al.
Published: (2025)
BREATH-VL: Vision-Language-Guided 6-DoF Bronchoscopy Localization via Semantic-Geometric Fusion
by: Tian, Qingyao, et al.
Published: (2026)
by: Tian, Qingyao, et al.
Published: (2026)
GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI
by: Li, Tianbin, et al.
Published: (2024)
by: Li, Tianbin, et al.
Published: (2024)
SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models
by: Makarov, Vladislav, et al.
Published: (2026)
by: Makarov, Vladislav, et al.
Published: (2026)
Similar Items
-
Unknown Prompt, the only Lacuna: Unveiling CLIP's Potential for Open Domain Generalization
by: Singha, Mainak, et al.
Published: (2024) -
AD-CLIP: Adapting Domains in Prompt Space Using CLIP
by: Singha, Mainak, et al.
Published: (2023) -
FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
by: Singha, Mainak, et al.
Published: (2025) -
Elevating All Zero-Shot Sketch-Based Image Retrieval Through Multimodal Prompt Learning
by: Singha, Mainak, et al.
Published: (2024) -
BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs
by: Singha, Mainak, et al.
Published: (2026)