Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Le, Quang-Hung, Dang, Long Hoang, Le, Ngan, Tran, Truyen, Le, Thao Minh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
by: Tran, Tuyen, et al.
Published: (2025)
by: Tran, Tuyen, et al.
Published: (2025)
SADL: An Effective In-Context Learning Method for Compositional Visual QA
by: Dang, Long Hoang, et al.
Published: (2024)
by: Dang, Long Hoang, et al.
Published: (2024)
Unified Framework with Consistency across Modalities for Human Activity Recognition
by: Tran, Tuyen, et al.
Published: (2024)
by: Tran, Tuyen, et al.
Published: (2024)
Towards Agentic AI for Multimodal-Guided Video Object Segmentation
by: Tran, Tuyen, et al.
Published: (2025)
by: Tran, Tuyen, et al.
Published: (2025)
Finding the Trigger: Causal Abductive Reasoning on Video Events
by: Le, Thao Minh, et al.
Published: (2025)
by: Le, Thao Minh, et al.
Published: (2025)
TrafficVLM: A Controllable Visual Language Model for Traffic Video Captioning
by: Dinh, Quang Minh, et al.
Published: (2024)
by: Dinh, Quang Minh, et al.
Published: (2024)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
by: Anh, Duy Le Dinh, et al.
Published: (2024)
by: Anh, Duy Le Dinh, et al.
Published: (2024)
LaVy: Vietnamese Multimodal Large Language Model
by: Tran, Chi, et al.
Published: (2024)
by: Tran, Chi, et al.
Published: (2024)
TP-GMOT: Tracking Generic Multiple Object by Textual Prompt with Motion-Appearance Cost (MAC) SORT
by: Anh, Duy Le Dinh, et al.
Published: (2024)
by: Anh, Duy Le Dinh, et al.
Published: (2024)
HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model
by: Vo, Khoa, et al.
Published: (2024)
by: Vo, Khoa, et al.
Published: (2024)
MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models
by: Xia, Yinan, et al.
Published: (2025)
by: Xia, Yinan, et al.
Published: (2025)
Robust Deepfake Detection: Mitigating Spatial Attention Drift via Calibrated Complementary Ensembles
by: Le-Phan, Minh-Khoa, et al.
Published: (2026)
by: Le-Phan, Minh-Khoa, et al.
Published: (2026)
EDGER: EDge-Guided with HEatmap Refinement for Generalizable Image Forgery Localization
by: Le-Phan, Minh-Khoa, et al.
Published: (2026)
by: Le-Phan, Minh-Khoa, et al.
Published: (2026)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
by: Liang, Qiao, et al.
Published: (2025)
by: Liang, Qiao, et al.
Published: (2025)
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
by: Le, Minh Khoa, et al.
Published: (2026)
by: Le, Minh Khoa, et al.
Published: (2026)
S2MoE: Robust Sparse Mixture of Experts via Stochastic Learning
by: Do, Giang, et al.
Published: (2025)
by: Do, Giang, et al.
Published: (2025)
Rethinking Sparse Mixture of Experts from a Unified Perspective
by: Do, Giang, et al.
Published: (2025)
by: Do, Giang, et al.
Published: (2025)
Do Domain-specific Experts exist in MoE-based LLMs?
by: Do, Giang, et al.
Published: (2026)
by: Do, Giang, et al.
Published: (2026)
MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
by: Nguyen, Hai-Dang, et al.
Published: (2025)
by: Nguyen, Hai-Dang, et al.
Published: (2025)
Amodal Instance Segmentation with Diffusion Shape Prior Estimation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
AHMsys: An Automated HVAC Modeling System for BIM Project
by: Dang, Long Hoang, et al.
Published: (2024)
by: Dang, Long Hoang, et al.
Published: (2024)
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
by: Darabi, Nastaran, et al.
Published: (2026)
by: Darabi, Nastaran, et al.
Published: (2026)
PROGRESSLM: Towards Progress Reasoning in Vision-Language Models
by: Zhang, Jianshu, et al.
Published: (2026)
by: Zhang, Jianshu, et al.
Published: (2026)
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
by: Le, Huy, et al.
Published: (2025)
by: Le, Huy, et al.
Published: (2025)
The Abstraction Gap in Vision-Language Causal Reasoning
by: Hoang, Chinh, et al.
Published: (2026)
by: Hoang, Chinh, et al.
Published: (2026)
Text-Enhanced Data-free Approach for Federated Class-Incremental Learning
by: Tran, Minh-Tuan, et al.
Published: (2024)
by: Tran, Minh-Tuan, et al.
Published: (2024)
VisTA: Vision-Text Alignment Model with Contrastive Learning using Multimodal Data for Evidence-Driven, Reliable, and Explainable Alzheimer's Disease Diagnosis
by: Can, Duy-Cat, et al.
Published: (2025)
by: Can, Duy-Cat, et al.
Published: (2025)
MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models
by: Biswas, Shristi Das, et al.
Published: (2026)
by: Biswas, Shristi Das, et al.
Published: (2026)
SimSMoE: Solving Representational Collapse via Similarity Measure
by: Do, Giang, et al.
Published: (2024)
by: Do, Giang, et al.
Published: (2024)
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
by: Nguyen, Hieu, et al.
Published: (2025)
by: Nguyen, Hieu, et al.
Published: (2025)
WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
by: Le, Huy, et al.
Published: (2023)
by: Le, Huy, et al.
Published: (2023)
Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models
by: Rajabi, Navid, et al.
Published: (2023)
by: Rajabi, Navid, et al.
Published: (2023)
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
by: Lu, Jiaying, et al.
Published: (2023)
by: Lu, Jiaying, et al.
Published: (2023)
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
by: You, Haoxuan, et al.
Published: (2023)
by: You, Haoxuan, et al.
Published: (2023)
U-CESE: Unified Clip-based Event Search Engine for AI Challenge HCMC 2025
by: Le, Duc-Nhuan, et al.
Published: (2026)
by: Le, Duc-Nhuan, et al.
Published: (2026)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
by: Chen, Jiaxing, et al.
Published: (2024)
by: Chen, Jiaxing, et al.
Published: (2024)
Unifying Global and Local Scene Entities Modelling for Precise Action Spotting
by: Tran, Kim Hoang, et al.
Published: (2024)
by: Tran, Kim Hoang, et al.
Published: (2024)
VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering
by: Nguyen, Hai-Dang, et al.
Published: (2025)
by: Nguyen, Hai-Dang, et al.
Published: (2025)
MambaU-Lite: A Lightweight Model based on Mamba and Integrated Channel-Spatial Attention for Skin Lesion Segmentation
by: Nguyen, Thi-Nhu-Quynh, et al.
Published: (2024)
by: Nguyen, Thi-Nhu-Quynh, et al.
Published: (2024)
Learning Human Motion with Temporally Conditional Mamba
by: Nguyen, Quang, et al.
Published: (2025)
by: Nguyen, Quang, et al.
Published: (2025)
Similar Items
-
Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
by: Tran, Tuyen, et al.
Published: (2025) -
SADL: An Effective In-Context Learning Method for Compositional Visual QA
by: Dang, Long Hoang, et al.
Published: (2024) -
Unified Framework with Consistency across Modalities for Human Activity Recognition
by: Tran, Tuyen, et al.
Published: (2024) -
Towards Agentic AI for Multimodal-Guided Video Object Segmentation
by: Tran, Tuyen, et al.
Published: (2025) -
Finding the Trigger: Causal Abductive Reasoning on Video Events
by: Le, Thao Minh, et al.
Published: (2025)