Small Vision-Language Models: A Survey on Compact Architectures and Techniques
Fuente:
arXiv
Saved in:
| Main Authors: | Patnaik, Nitesh, Nayak, Navdeep, Agrawal, Himani Bansal, Khamaru, Moinak Chinmoy, Bal, Gourav, Panda, Saishree Smaranika, Raj, Rishi, Meena, Vishal, Vadlamani, Kartheek |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI
by: Kasu, Sai Kartheek Reddy
Published: (2025)
by: Kasu, Sai Kartheek Reddy
Published: (2025)
A Comprehensive Study of the Current State-of-the-Art in Nepali Automatic Speech Recognition Systems
by: Ghimire, Rupak Raj, et al.
Published: (2024)
by: Ghimire, Rupak Raj, et al.
Published: (2024)
SoC-DT: Standard-of-Care Aligned Digital Twins for Patient-Specific Tumor Dynamics
by: Bhattacharya, Moinak, et al.
Published: (2025)
by: Bhattacharya, Moinak, et al.
Published: (2025)
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning
by: Patnaik, Sohan, et al.
Published: (2025)
by: Patnaik, Sohan, et al.
Published: (2025)
Benchmarking BERT-based Models for Sentence-level Topic Classification in Nepali Language
by: Karki, Nischal, et al.
Published: (2026)
by: Karki, Nischal, et al.
Published: (2026)
Improving the Accessibility of Dating Websites for Individuals with Visual Impairments
by: Shrestha, Gyanendra, et al.
Published: (2024)
by: Shrestha, Gyanendra, et al.
Published: (2024)
Matricial Free Energy as a Gaussianizing Regularizer: Enhancing Autoencoders for Gaussian Code Generation
by: Sonthalia, Rishi, et al.
Published: (2025)
by: Sonthalia, Rishi, et al.
Published: (2025)
Individual-level models of disease transmission incorporating piecewise spatial risk functions
by: Rahul, Chinmoy Roy, et al.
Published: (2024)
by: Rahul, Chinmoy Roy, et al.
Published: (2024)
FRACTAL: Fine-Grained Scoring from Aggregate Text Labels
by: Makhija, Yukti, et al.
Published: (2024)
by: Makhija, Yukti, et al.
Published: (2024)
Deceptive Humor: A Synthetic Multilingual Benchmark Dataset for Bridging Fabricated Claims with Humorous Content
by: Kasu, Sai Kartheek Reddy, et al.
Published: (2025)
by: Kasu, Sai Kartheek Reddy, et al.
Published: (2025)
Improved Differential Evolution based Feature Selection through Quantum, Chaos, and Lasso
by: Vivek, Yelleti, et al.
Published: (2024)
by: Vivek, Yelleti, et al.
Published: (2024)
Continual Learning-Based Unified Model for Unpaired Image Restoration Tasks
by: Kartheek, Kotha, et al.
Published: (2025)
by: Kartheek, Kotha, et al.
Published: (2025)
Anatomy-DT: A Cross-Diffusion Digital Twin for Anatomical Evolution
by: Bhattacharya, Moinak, et al.
Published: (2025)
by: Bhattacharya, Moinak, et al.
Published: (2025)
SmoGVLM: A Small, Graph-enhanced Vision-Language Model
by: Mondal, Debjyoti, et al.
Published: (2026)
by: Mondal, Debjyoti, et al.
Published: (2026)
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
by: Awal, Rabiul, et al.
Published: (2023)
by: Awal, Rabiul, et al.
Published: (2023)
RadGazeGen: Radiomics and Gaze-guided Medical Image Generation using Diffusion Models
by: Bhattacharya, Moinak, et al.
Published: (2024)
by: Bhattacharya, Moinak, et al.
Published: (2024)
GazeLT: Visual attention-guided long-tailed disease classification in chest radiographs
by: Bhattacharya, Moinak, et al.
Published: (2025)
by: Bhattacharya, Moinak, et al.
Published: (2025)
Utilizing Multi-Agent Reinforcement Learning with Encoder-Decoder Architecture Agents to Identify Optimal Resection Location in Glioblastoma Multiforme Patients
by: Arun, Krishna, et al.
Published: (2025)
by: Arun, Krishna, et al.
Published: (2025)
D-HUMOR: Dark Humor Understanding via Multimodal Open-ended Reasoning -- A Benchmark Dataset and Method
by: Kasu, Sai Kartheek Reddy, et al.
Published: (2025)
by: Kasu, Sai Kartheek Reddy, et al.
Published: (2025)
Counterfactual Forecasting of Human Behavior using Generative AI and Causal Graphs
by: Uddandarao, Dharmateja Priyadarshi, et al.
Published: (2025)
by: Uddandarao, Dharmateja Priyadarshi, et al.
Published: (2025)
ShelfHelp: Empowering Humans to Perform Vision-Independent Manipulation Tasks with a Socially Assistive Robotic Cane
by: Agrawal, Shivendra, et al.
Published: (2024)
by: Agrawal, Shivendra, et al.
Published: (2024)
Building a Few-Shot Cross-Domain Multilingual NLU Model for Customer Care
by: Kumar, Saurabh, et al.
Published: (2025)
by: Kumar, Saurabh, et al.
Published: (2025)
SAGE: Steering Dialog Generation with Future-Aware State-Action Augmentation
by: Zhang, Yizhe, et al.
Published: (2025)
by: Zhang, Yizhe, et al.
Published: (2025)
Consolidating and Developing Benchmarking Datasets for the Nepali Natural Language Understanding Tasks
by: Nyachhyon, Jinu, et al.
Published: (2024)
by: Nyachhyon, Jinu, et al.
Published: (2024)
Small LLMs with Expert Blocks Are Good Enough for Hyperparamter Tuning
by: Naphade, Om, et al.
Published: (2025)
by: Naphade, Om, et al.
Published: (2025)
iWatchRoadv2: Pothole Detection, Geospatial Mapping, and Intelligent Road Governance
by: Sahoo, Rishi Raj, et al.
Published: (2025)
by: Sahoo, Rishi Raj, et al.
Published: (2025)
iWatchRoad: Scalable Detection and Geospatial Visualization of Potholes for Smart Cities
by: Sahoo, Rishi Raj, et al.
Published: (2025)
by: Sahoo, Rishi Raj, et al.
Published: (2025)
Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games
by: Zhang, Yizhe, et al.
Published: (2023)
by: Zhang, Yizhe, et al.
Published: (2023)
On the Interplay of Cube Learning and Dependency Schemes in QCDCL Proof Systems
by: Choudhury, Abhimanyu, et al.
Published: (2025)
by: Choudhury, Abhimanyu, et al.
Published: (2025)
QBF Merge Resolution is powerful but unnatural
by: Mahajan, Meena, et al.
Published: (2022)
by: Mahajan, Meena, et al.
Published: (2022)
LISR: Learning Linear 3D Implicit Surface Representation Using Compactly Supported Radial Basis Functions
by: Pandey, Atharva, et al.
Published: (2024)
by: Pandey, Atharva, et al.
Published: (2024)
Nwāchā Munā: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit
by: Nareddy, Kartheek Kumar Reddy, et al.
Published: (2025)
by: Nareddy, Kartheek Kumar Reddy, et al.
Published: (2025)
A Framework for LLM-powered Design Assistants
by: Panda, Swaroop
Published: (2025)
by: Panda, Swaroop
Published: (2025)
A Comparative Study of Modern Object Detectors for Robust Apple Detection in Orchard Imagery
by: Asad, Mohammed, et al.
Published: (2026)
by: Asad, Mohammed, et al.
Published: (2026)
What Does Success Look Like? Catalyzing Meeting Intentionality with AI-Assisted Prospective Reflection
by: Scott, Ava Elizabeth, et al.
Published: (2025)
by: Scott, Ava Elizabeth, et al.
Published: (2025)
Are We On Track? AI-Assisted Active and Passive Goal Reflection During Meetings
by: Chen, Xinyue, et al.
Published: (2025)
by: Chen, Xinyue, et al.
Published: (2025)
Hall current and Dufour effect on the flow of a viscous fluid in the presence of suction, chemical reaction, and heat source: Laplace transform procedure
by: Chinmoy Rath, et al.
Published: (2024)
by: Chinmoy Rath, et al.
Published: (2024)
Development of Pre-Trained Transformer-based Models for the Nepali Language
by: Thapa, Prajwal, et al.
Published: (2024)
by: Thapa, Prajwal, et al.
Published: (2024)
Location-Aware Dispersion on Anonymous Graphs
by: Himani, et al.
Published: (2026)
by: Himani, et al.
Published: (2026)
Similar Items
-
EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI
by: Kasu, Sai Kartheek Reddy
Published: (2025) -
A Comprehensive Study of the Current State-of-the-Art in Nepali Automatic Speech Recognition Systems
by: Ghimire, Rupak Raj, et al.
Published: (2024) -
SoC-DT: Standard-of-Care Aligned Digital Twins for Patient-Specific Tumor Dynamics
by: Bhattacharya, Moinak, et al.
Published: (2025) -
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning
by: Patnaik, Sohan, et al.
Published: (2025) -
Benchmarking BERT-based Models for Sentence-level Topic Classification in Nepali Language
by: Karki, Nischal, et al.
Published: (2026)