TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints
Fuente:
arXiv
Saved in:
| Main Authors: | Ly, Vinh-Thuan, Truong, Hoang M., Nguyen, Xuan-Huong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FSM-Net: An Efficient Frequency-Spatial Network for Real-World Deblurring
by: Ly, Vinh-Thuan
Published: (2026)
by: Ly, Vinh-Thuan
Published: (2026)
Dual-Path Enhancements in Event-Based Eye Tracking: Augmented Robustness and Adaptive Temporal Modeling
by: Truong, Hoang M., et al.
Published: (2025)
by: Truong, Hoang M., et al.
Published: (2025)
KeyNode-Driven Geometry Coding for Real-World Scanned Human Dynamic Mesh Compression
by: Hoang, Huong, et al.
Published: (2025)
by: Hoang, Huong, et al.
Published: (2025)
SwiftBrush: One-Step Text-to-Image Diffusion Model with Variational Score Distillation
by: Nguyen, Thuan Hoang, et al.
Published: (2023)
by: Nguyen, Thuan Hoang, et al.
Published: (2023)
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
by: Truong, Thanh-Dat, et al.
Published: (2025)
by: Truong, Thanh-Dat, et al.
Published: (2025)
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization
by: Pham, Tan-Hanh, et al.
Published: (2024)
by: Pham, Tan-Hanh, et al.
Published: (2024)
SDPA++: A General Framework for Self-Supervised Denoising with Patch Aggregation
by: Nguyen, Huy Minh Nhat, et al.
Published: (2025)
by: Nguyen, Huy Minh Nhat, et al.
Published: (2025)
Enhancing Rotated Object Detection via Anisotropic Gaussian Bounding Box and Bhattacharyya Distance
by: Thai, Chien, et al.
Published: (2025)
by: Thai, Chien, et al.
Published: (2025)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
Detection Fire in Camera RGB-NIR
by: Khai, Nguyen Truong, et al.
Published: (2025)
by: Khai, Nguyen Truong, et al.
Published: (2025)
Vision-Language Models for Infrared Industrial Sensing in Additive Manufacturing Scene Description
by: Mahjourian, Nazanin, et al.
Published: (2025)
by: Mahjourian, Nazanin, et al.
Published: (2025)
Sanitizing Manufacturing Dataset Labels Using Vision-Language Models
by: Mahjourian, Nazanin, et al.
Published: (2025)
by: Mahjourian, Nazanin, et al.
Published: (2025)
IQBench: How "Smart'' Are Vision-Language Models? A Study with Human IQ Tests
by: Pham, Tan-Hanh, et al.
Published: (2025)
by: Pham, Tan-Hanh, et al.
Published: (2025)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
Efficient Image Synthesis with Sphere Latent Encoder
by: Do, Tung, et al.
Published: (2026)
by: Do, Tung, et al.
Published: (2026)
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
by: Zhang, Jian, et al.
Published: (2026)
by: Zhang, Jian, et al.
Published: (2026)
Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
by: Pham, Phuc, et al.
Published: (2025)
by: Pham, Phuc, et al.
Published: (2025)
SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
by: Nguyen, Trong-Thuan, et al.
Published: (2025)
by: Nguyen, Trong-Thuan, et al.
Published: (2025)
MambaCAFU: Hybrid Multi-Scale and Multi-Attention Model with Mamba-Based Fusion for Medical Image Segmentation
by: Bui, T-Mai, et al.
Published: (2025)
by: Bui, T-Mai, et al.
Published: (2025)
Semantic Alignment in Hyperbolic Space for Open-Vocabulary Semantic Segmentation
by: Truong, Hoang M., et al.
Published: (2026)
by: Truong, Hoang M., et al.
Published: (2026)
Rethinking Top Probability from Multi-view for Distracted Driver Behaviour Localization
by: Nguyen, Quang Vinh, et al.
Published: (2024)
by: Nguyen, Quang Vinh, et al.
Published: (2024)
N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
by: Wang, Yuxin, et al.
Published: (2025)
by: Wang, Yuxin, et al.
Published: (2025)
TinyVLM: Zero-Shot Object Detection on Microcontrollers via Vision-Language Distillation with Matryoshka Embeddings
by: Wilson, Bibin
Published: (2026)
by: Wilson, Bibin
Published: (2026)
A Hybrid Vision Transformer Approach for Mathematical Expression Recognition
by: Le, Anh Duy, et al.
Published: (2026)
by: Le, Anh Duy, et al.
Published: (2026)
A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
by: Nguyen, Phu-Vinh, et al.
Published: (2025)
by: Nguyen, Phu-Vinh, et al.
Published: (2025)
A Lightweight Moment Retrieval System with Global Re-Ranking and Robust Adaptive Bidirectional Temporal Search
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT
by: Asfour, Alaa, et al.
Published: (2026)
by: Asfour, Alaa, et al.
Published: (2026)
WorldVLM: Combining World Model Forecasting and Vision-Language Reasoning
by: Englmeier, Stefan, et al.
Published: (2026)
by: Englmeier, Stefan, et al.
Published: (2026)
Vision-Language Memory for Spatial Reasoning
by: Liu, Zuntao, et al.
Published: (2025)
by: Liu, Zuntao, et al.
Published: (2025)
A-SCoRe: Attention-based Scene Coordinate Regression for wide-ranging scenarios
by: Bui, Huy-Hoang, et al.
Published: (2025)
by: Bui, Huy-Hoang, et al.
Published: (2025)
PDIWS: Thermal Imaging Dataset for Person Detection in Intrusion Warning Systems
by: Thuan, Nguyen Duc, et al.
Published: (2023)
by: Thuan, Nguyen Duc, et al.
Published: (2023)
MambaU-Lite: A Lightweight Model based on Mamba and Integrated Channel-Spatial Attention for Skin Lesion Segmentation
by: Nguyen, Thi-Nhu-Quynh, et al.
Published: (2024)
by: Nguyen, Thi-Nhu-Quynh, et al.
Published: (2024)
Tri-Bench: Stress-Testing VLM Reliability on Spatial Reasoning under Camera Tilt and Object Interference
by: Bendkhale, Amit
Published: (2025)
by: Bendkhale, Amit
Published: (2025)
ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models
by: Ko, Dohwan, et al.
Published: (2025)
by: Ko, Dohwan, et al.
Published: (2025)
StateVLM: A State-Aware Vision-Language Model for Robotic Affordance Reasoning
by: Sun, Xiaowen, et al.
Published: (2026)
by: Sun, Xiaowen, et al.
Published: (2026)
RARL: Improving Medical VLM Reasoning and Generalization with Reinforcement Learning and LoRA under Data and Hardware Constraints
by: Pham, Tan-Hanh, et al.
Published: (2025)
by: Pham, Tan-Hanh, et al.
Published: (2025)
TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation
by: Liu, Jiaxing, et al.
Published: (2026)
by: Liu, Jiaxing, et al.
Published: (2026)
SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
by: Sivakumar, Anushka, et al.
Published: (2025)
by: Sivakumar, Anushka, et al.
Published: (2025)
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
by: Cheng, An-Chieh, et al.
Published: (2024)
by: Cheng, An-Chieh, et al.
Published: (2024)
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks
by: Hu, Yuanze, et al.
Published: (2025)
by: Hu, Yuanze, et al.
Published: (2025)
Similar Items
-
FSM-Net: An Efficient Frequency-Spatial Network for Real-World Deblurring
by: Ly, Vinh-Thuan
Published: (2026) -
Dual-Path Enhancements in Event-Based Eye Tracking: Augmented Robustness and Adaptive Temporal Modeling
by: Truong, Hoang M., et al.
Published: (2025) -
KeyNode-Driven Geometry Coding for Real-World Scanned Human Dynamic Mesh Compression
by: Hoang, Huong, et al.
Published: (2025) -
SwiftBrush: One-Step Text-to-Image Diffusion Model with Variational Score Distillation
by: Nguyen, Thuan Hoang, et al.
Published: (2023) -
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
by: Truong, Thanh-Dat, et al.
Published: (2025)