SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Pham, Tan-Hanh, Le, Hoang-Nam, Nguyen, Phu-Vinh, Ngo, Chris, Hy, Truong-Son |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SilVar-Med: A Speech-Driven Visual Language Model for Explainable Abnormality Detection in Medical Imaging
par: Pham, Tan-Hanh, et autres
Publié: (2025)
par: Pham, Tan-Hanh, et autres
Publié: (2025)
A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
par: Nguyen, Phu-Vinh, et autres
Publié: (2025)
par: Nguyen, Phu-Vinh, et autres
Publié: (2025)
Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles
par: Vinh, Nguyen Phu, et autres
Publié: (2025)
par: Vinh, Nguyen Phu, et autres
Publié: (2025)
IQBench: How "Smart'' Are Vision-Language Models? A Study with Human IQ Tests
par: Pham, Tan-Hanh, et autres
Publié: (2025)
par: Pham, Tan-Hanh, et autres
Publié: (2025)
Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
par: Pham, Tan-Hanh, et autres
Publié: (2025)
par: Pham, Tan-Hanh, et autres
Publié: (2025)
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
par: Le-Duc, Khai, et autres
Publié: (2024)
par: Le-Duc, Khai, et autres
Publié: (2024)
wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
par: Le-Duc, Khai, et autres
Publié: (2024)
par: Le-Duc, Khai, et autres
Publié: (2024)
RARL: Improving Medical VLM Reasoning and Generalization with Reinforcement Learning and LoRA under Data and Hardware Constraints
par: Pham, Tan-Hanh, et autres
Publié: (2025)
par: Pham, Tan-Hanh, et autres
Publié: (2025)
ViInfographicVQA: A Benchmark for Single and Multi-image Visual Question Answering on Vietnamese Infographics
par: Van-Dinh, Tue-Thu, et autres
Publié: (2025)
par: Van-Dinh, Tue-Thu, et autres
Publié: (2025)
LiteGPT: Large Vision-Language Model for Joint Chest X-ray Localization and Classification Task
par: Le-Duc, Khai, et autres
Publié: (2024)
par: Le-Duc, Khai, et autres
Publié: (2024)
RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models
par: Dang, Quy-Anh, et autres
Publié: (2026)
par: Dang, Quy-Anh, et autres
Publié: (2026)
RainbowPlus: Enhancing Adversarial Prompt Generation via Evolutionary Quality-Diversity Search
par: Dang, Quy-Anh, et autres
Publié: (2025)
par: Dang, Quy-Anh, et autres
Publié: (2025)
LINKER: Learning Interactions Between Functional Groups and Residues With Chemical Knowledge-Enhanced Reasoning and Explainability
par: Pham, Phuc, et autres
Publié: (2025)
par: Pham, Phuc, et autres
Publié: (2025)
Multimodal graph representation learning for website generation based on visual sketch
par: Vu, Tung D., et autres
Publié: (2025)
par: Vu, Tung D., et autres
Publié: (2025)
Range-aware Positional Encoding via High-order Pretraining: Theory and Practice
par: Nguyen, Viet Anh, et autres
Publié: (2024)
par: Nguyen, Viet Anh, et autres
Publié: (2024)
Q-BIOLAT: Binary Latent Protein Fitness Landscapes for QUBO-Based Optimization
par: Hy, Truong-Son
Publié: (2026)
par: Hy, Truong-Son
Publié: (2026)
Binary Latent Protein Fitness Landscapes for Quantum Annealing Optimization
par: Hy, Truong-Son
Publié: (2026)
par: Hy, Truong-Son
Publié: (2026)
ESGNN: Towards Equivariant Scene Graph Neural Network for 3D Scene Understanding
par: Pham, Quang P. M., et autres
Publié: (2024)
par: Pham, Quang P. M., et autres
Publié: (2024)
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
par: Vu, Sinh Trong, et autres
Publié: (2025)
par: Vu, Sinh Trong, et autres
Publié: (2025)
A nonlinear analogue of additive commutators
par: Dung, Truong Huu, et autres
Publié: (2025)
par: Dung, Truong Huu, et autres
Publié: (2025)
Efficient Deep Learning for Medical Imaging: Bridging the Gap Between High-Performance AI and Clinical Deployment
par: Nguyen, Cuong Manh, et autres
Publié: (2026)
par: Nguyen, Cuong Manh, et autres
Publié: (2026)
Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration
par: Nguyen, Ngoc Son, et autres
Publié: (2024)
par: Nguyen, Ngoc Son, et autres
Publié: (2024)
Real-time Speech Summarization for Medical Conversations
par: Le-Duc, Khai, et autres
Publié: (2024)
par: Le-Duc, Khai, et autres
Publié: (2024)
AccurateRAG: A Framework for Building Accurate Retrieval-Augmented Question-Answering Applications
par: Nguyen, Linh The, et autres
Publié: (2025)
par: Nguyen, Linh The, et autres
Publié: (2025)
An Approach of Structure‐Enhanced Code‐Centric Graph Learning for Just‐in‐Time Software Vulnerability Detection
par: Phu Pham, et autres
Publié: (2026)
par: Phu Pham, et autres
Publié: (2026)
TESGNN: Temporal Equivariant Scene Graph Neural Networks for Efficient and Robust Multi-View 3D Scene Understanding
par: Pham, Quang P. M., et autres
Publié: (2024)
par: Pham, Quang P. M., et autres
Publié: (2024)
OWLViz: An Open-World Benchmark for Visual Question Answering
par: Nguyen, Thuy, et autres
Publié: (2025)
par: Nguyen, Thuy, et autres
Publié: (2025)
Vietnamese Legal Information Retrieval in Question-Answering System
par: Ba, Thiem Nguyen, et autres
Publié: (2024)
par: Ba, Thiem Nguyen, et autres
Publié: (2024)
Rethinking Top Probability from Multi-view for Distracted Driver Behaviour Localization
par: Nguyen, Quang Vinh, et autres
Publié: (2024)
par: Nguyen, Quang Vinh, et autres
Publié: (2024)
Towards Robust Fact-Checking: A Multi-Agent System with Advanced Evidence Retrieval
par: Trinh, Tam, et autres
Publié: (2025)
par: Trinh, Tam, et autres
Publié: (2025)
Advances in Protein Representation Learning: Methods, Applications, and Future Directions
par: Nguyen, Viet Thanh Duy, et autres
Publié: (2025)
par: Nguyen, Viet Thanh Duy, et autres
Publié: (2025)
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow Matching
par: Nguyen, Ngoc-Son, et autres
Publié: (2025)
par: Nguyen, Ngoc-Son, et autres
Publié: (2025)
GROOT: Effective Design of Biological Sequences with Limited Experimental Data
par: Tran, Thanh V. T., et autres
Publié: (2024)
par: Tran, Thanh V. T., et autres
Publié: (2024)
EquiHGNN: Scalable Rotationally Equivariant Hypergraph Neural Networks
par: Dang, Tien, et autres
Publié: (2025)
par: Dang, Tien, et autres
Publié: (2025)
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
par: Huynh-Nguyen, Hieu-Nghia, et autres
Publié: (2025)
par: Huynh-Nguyen, Hieu-Nghia, et autres
Publié: (2025)
TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints
par: Ly, Vinh-Thuan, et autres
Publié: (2025)
par: Ly, Vinh-Thuan, et autres
Publié: (2025)
Multimodal Contrastive Representation Learning in Augmented Biomedical Knowledge Graphs
par: Dang, Tien, et autres
Publié: (2025)
par: Dang, Tien, et autres
Publié: (2025)
Adaptive Compensation for Robotic Joint Failures Using Partially Observable Reinforcement Learning
par: Pham, Tan-Hanh, et autres
Publié: (2024)
par: Pham, Tan-Hanh, et autres
Publié: (2024)
A note on additive commutator groups in certain algebras
par: Ha, Nguyen Thi Thai, et autres
Publié: (2025)
par: Ha, Nguyen Thi Thai, et autres
Publié: (2025)
MOB-GCN: A Novel Multiscale Object-Based Graph Neural Network for Hyperspectral Image Classification
par: Yang, Tuan-Anh, et autres
Publié: (2025)
par: Yang, Tuan-Anh, et autres
Publié: (2025)
Documents similaires
-
SilVar-Med: A Speech-Driven Visual Language Model for Explainable Abnormality Detection in Medical Imaging
par: Pham, Tan-Hanh, et autres
Publié: (2025) -
A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
par: Nguyen, Phu-Vinh, et autres
Publié: (2025) -
Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles
par: Vinh, Nguyen Phu, et autres
Publié: (2025) -
IQBench: How "Smart'' Are Vision-Language Models? A Study with Human IQ Tests
par: Pham, Tan-Hanh, et autres
Publié: (2025) -
Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
par: Pham, Tan-Hanh, et autres
Publié: (2025)