Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Quy, Nguyen Lam Phu, Hoa, Pham Phu, Nguyen, Tran Chi, Minh, Dao Sy Duy, Ngoc, Nguyen Hoang Minh, Kiet, Huynh Trung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging Lightweight Entity Extraction for Scalable Event-Based Image Retrieval
by: Minh, Dao Sy Duy, et al.
Published: (2025)
by: Minh, Dao Sy Duy, et al.
Published: (2025)
Navigating Simply, Aligning Deeply: Winning Solutions for Mouse vs. AI 2025
by: Pham, Phu-Hoa, et al.
Published: (2026)
by: Pham, Phu-Hoa, et al.
Published: (2026)
VietMed-MCQ: A Consistency-Filtered Data Synthesis Framework for Vietnamese Traditional Medicine Evaluation
by: Kiet, Huynh Trung, et al.
Published: (2026)
by: Kiet, Huynh Trung, et al.
Published: (2026)
EvoQRE: Modeling Bounded Rationality in Safety-Critical Traffic Simulation via Evolutionary Quantal Response Equilibrium
by: Pham, Phu-Hoa, et al.
Published: (2026)
by: Pham, Phu-Hoa, et al.
Published: (2026)
MEMRES: A Memory-Augmented Resolver with Confidence Cascade for Agentic Python Dependency Resolution
by: Minh, Dao Sy Duy, et al.
Published: (2026)
by: Minh, Dao Sy Duy, et al.
Published: (2026)
Unlocking Compositional Generalization in Continual Few-Shot Learning
by: Nguyen-Lam, Phu-Quy, et al.
Published: (2026)
by: Nguyen-Lam, Phu-Quy, et al.
Published: (2026)
MIST: Reliable Streaming Decision Trees for Online Class-Incremental Learning via McDiarmid Bound
by: Pham, Phu-Hoa, et al.
Published: (2026)
by: Pham, Phu-Hoa, et al.
Published: (2026)
Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery
by: Tran, Chi-Nguyen, et al.
Published: (2026)
by: Tran, Chi-Nguyen, et al.
Published: (2026)
Evaluating Perspectival Biases in Cross-Modal Retrieval
by: Saengsukhiran, Teerapol, et al.
Published: (2025)
by: Saengsukhiran, Teerapol, et al.
Published: (2025)
Few TensoRF: Enhance the Few-shot on Tensorial Radiance Fields
by: Le, Thanh-Hai, et al.
Published: (2026)
by: Le, Thanh-Hai, et al.
Published: (2026)
More at Stake: How Payoff and Language Shape LLM Agent Strategies in Cooperation Dilemmas
by: Huynh, Trung-Kiet, et al.
Published: (2026)
by: Huynh, Trung-Kiet, et al.
Published: (2026)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
by: Anh, Duy Le Dinh, et al.
Published: (2024)
by: Anh, Duy Le Dinh, et al.
Published: (2024)
CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA
by: Kovalev, Vsevolod, et al.
Published: (2025)
by: Kovalev, Vsevolod, et al.
Published: (2025)
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
by: Kiet, Huynh Trung, et al.
Published: (2026)
by: Kiet, Huynh Trung, et al.
Published: (2026)
Early-Stage Prediction of Review Effort in AI-Generated Pull Requests
by: Minh, Dao Sy Duy, et al.
Published: (2026)
by: Minh, Dao Sy Duy, et al.
Published: (2026)
A Scalable Pipeline Combining Procedural 3D Graphics and Guided Diffusion for Photorealistic Synthetic Training Data Generation in White Button Mushroom Segmentation
by: Károly, Artúr I., et al.
Published: (2025)
by: Károly, Artúr I., et al.
Published: (2025)
ReCoVR: Closing the Loop in Interactive Composed Video Retrieval
by: Zhang, Bingqing, et al.
Published: (2026)
by: Zhang, Bingqing, et al.
Published: (2026)
Expertized Caption Auto-Enhancement for Video-Text Retrieval
by: Yang, Baoyao, et al.
Published: (2025)
by: Yang, Baoyao, et al.
Published: (2025)
Almost Linear Time Consistent Mode Estimation and Quick Shift Clustering
by: Hashemian, Sajjad
Published: (2025)
by: Hashemian, Sajjad
Published: (2025)
DualPrompt-MedCap: A Dual-Prompt Enhanced Approach for Medical Image Captioning
by: Zhao, Yining, et al.
Published: (2025)
by: Zhao, Yining, et al.
Published: (2025)
Aximorphic Perspective Projection Model for Immersive Imagery
by: Fober, Jakub Maksymilian
Published: (2021)
by: Fober, Jakub Maksymilian
Published: (2021)
Towards Accurate and Efficient Waste Image Classification: A Hybrid Deep Learning and Machine Learning Approach
by: Nguyen, Ngoc-Bao-Quang, et al.
Published: (2025)
by: Nguyen, Ngoc-Bao-Quang, et al.
Published: (2025)
From eye to AI: studying rodent social behavior in the era of machine Learning
by: Chindemi, Giuseppe, et al.
Published: (2025)
by: Chindemi, Giuseppe, et al.
Published: (2025)
On the robustness of ChatGPT in teaching Korean Mathematics
by: Nguyen, Phuong-Nam, et al.
Published: (2025)
by: Nguyen, Phuong-Nam, et al.
Published: (2025)
Optimizing Neurorobot Policy under Limited Demonstration Data through Preference Regret
by: Nguyen, Viet Dung, et al.
Published: (2026)
by: Nguyen, Viet Dung, et al.
Published: (2026)
PC-SNN: Predictive Coding-based Local Hebbian Plasticity Learning in Spiking Neural Networks
by: Wang, Haidong, et al.
Published: (2022)
by: Wang, Haidong, et al.
Published: (2022)
ROSGS: Relightable Outdoor Scenes With Gaussian Splatting
by: Liao, Lianjun, et al.
Published: (2025)
by: Liao, Lianjun, et al.
Published: (2025)
From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring
by: Nguyen, Minh Hoang, et al.
Published: (2026)
by: Nguyen, Minh Hoang, et al.
Published: (2026)
Spatially-Grounded Document Retrieval via Patch-to-Region Relevance Propagation
by: Georgiou, Athos
Published: (2025)
by: Georgiou, Athos
Published: (2025)
A Two-stage Transformer Framework for Temporal Localization of Distracted Driver Behaviors
by: Doan, Gia-Bao, et al.
Published: (2026)
by: Doan, Gia-Bao, et al.
Published: (2026)
A Grounded Memory System For Smart Personal Assistants
by: Ocker, Felix, et al.
Published: (2025)
by: Ocker, Felix, et al.
Published: (2025)
Gaussian Splatting: 3D Reconstruction and Novel View Synthesis, a Review
by: Dalal, Anurag, et al.
Published: (2024)
by: Dalal, Anurag, et al.
Published: (2024)
MDA: An Interpretable and Scalable Multi-Modal Fusion under Missing Modalities and Intrinsic Noise Conditions
by: Fan, Lin, et al.
Published: (2024)
by: Fan, Lin, et al.
Published: (2024)
Leveraging OpenFlamingo for Multimodal Embedding Analysis of C2C Car Parts Data
by: Rashid, Maisha Binte, et al.
Published: (2025)
by: Rashid, Maisha Binte, et al.
Published: (2025)
Knee Osteoarthritis Severity Grading Using Optimized Deep Learning and LLM-Driven Intelligent AI on Computationally Limited Systems
by: Nadeem, Dayam, et al.
Published: (2026)
by: Nadeem, Dayam, et al.
Published: (2026)
Topology-Aware Latent Diffusion for 3D Shape Generation
by: Hu, Jiangbei, et al.
Published: (2024)
by: Hu, Jiangbei, et al.
Published: (2024)
PrismAvatar: Real-time animated 3D neural head avatars on edge devices
by: Raina, Prashant, et al.
Published: (2025)
by: Raina, Prashant, et al.
Published: (2025)
Experimenting active and sequential learning in a medieval music manuscript
by: Sharma, Sachin, et al.
Published: (2025)
by: Sharma, Sachin, et al.
Published: (2025)
Large Language Model for Qualitative Research -- A Systematic Mapping Study
by: Barros, Cauã Ferreira, et al.
Published: (2024)
by: Barros, Cauã Ferreira, et al.
Published: (2024)
Generating 3D Terrain with 2D Cellular Automata
by: Fachada, Nuno, et al.
Published: (2024)
by: Fachada, Nuno, et al.
Published: (2024)
Similar Items
-
Leveraging Lightweight Entity Extraction for Scalable Event-Based Image Retrieval
by: Minh, Dao Sy Duy, et al.
Published: (2025) -
Navigating Simply, Aligning Deeply: Winning Solutions for Mouse vs. AI 2025
by: Pham, Phu-Hoa, et al.
Published: (2026) -
VietMed-MCQ: A Consistency-Filtered Data Synthesis Framework for Vietnamese Traditional Medicine Evaluation
by: Kiet, Huynh Trung, et al.
Published: (2026) -
EvoQRE: Modeling Bounded Rationality in Safety-Critical Traffic Simulation via Evolutionary Quantal Response Equilibrium
by: Pham, Phu-Hoa, et al.
Published: (2026) -
MEMRES: A Memory-Augmented Resolver with Confidence Cascade for Agentic Python Dependency Resolution
by: Minh, Dao Sy Duy, et al.
Published: (2026)