ReCap: Event-Aware Image Captioning with Article Retrieval and Semantic Gaussian Normalization
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Thinh-Phuc, Nguyen, Thanh-Hai, Dinh, Gia-Huy, Nguyen, Lam-Huy, Tran, Minh-Triet, Le, Trung-Nghia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VisionGuard: Synergistic Framework for Helmet Violation Detection
by: Nguyen, Lam-Huy, et al.
Published: (2025)
by: Nguyen, Lam-Huy, et al.
Published: (2025)
EventCap
by: Nguyen, Phuc-Tan, et al.
Published: (2024)
by: Nguyen, Phuc-Tan, et al.
Published: (2024)
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
by: Vo, Dinh-Khoi, et al.
Published: (2025)
by: Vo, Dinh-Khoi, et al.
Published: (2025)
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
by: Nguyen, Hieu, et al.
Published: (2025)
by: Nguyen, Hieu, et al.
Published: (2025)
SAMURAI: Shape-Aware Multimodal Retrieval for 3D Object Identification
by: Vo, Dinh-Khoi, et al.
Published: (2025)
by: Vo, Dinh-Khoi, et al.
Published: (2025)
A Re-ranking Method using K-nearest Weighted Fusion for Person Re-identification
by: Che, Huy, et al.
Published: (2025)
by: Che, Huy, et al.
Published: (2025)
IGL-DT: Iterative Global-Local Feature Learning with Dual-Teacher Semantic Segmentation Framework under Limited Annotation Scheme
by: Tran, Dinh Dai Quan, et al.
Published: (2025)
by: Tran, Dinh Dai Quan, et al.
Published: (2025)
KiseKloset for Fashion Retrieval and Recommendation
by: Phan-Nguyen, Thanh-Tung, et al.
Published: (2025)
by: Phan-Nguyen, Thanh-Tung, et al.
Published: (2025)
Enhanced Multimodal Video Retrieval System: Integrating Query Expansion and Cross-modal Temporal Event Retrieval
by: Vo, Van-Thinh, et al.
Published: (2025)
by: Vo, Van-Thinh, et al.
Published: (2025)
GenFlow: Interactive Modular System for Image Generation
by: Nguyen, Duc-Hung, et al.
Published: (2025)
by: Nguyen, Duc-Hung, et al.
Published: (2025)
Semi-Supervised Semantic Segmentation using Redesigned Self-Training for White Blood Cells
by: Luu, Vinh Quoc, et al.
Published: (2024)
by: Luu, Vinh Quoc, et al.
Published: (2024)
FaR: Enhancing Multi-Concept Text-to-Image Diffusion via Concept Fusion and Localized Refinement
by: Tran, Gia-Nghia, et al.
Published: (2025)
by: Tran, Gia-Nghia, et al.
Published: (2025)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
Video ReCap: Recursive Captioning of Hour-Long Videos
by: Islam, Md Mohaiminul, et al.
Published: (2024)
by: Islam, Md Mohaiminul, et al.
Published: (2024)
Link prediction Graph Neural Networks for structure recognition of Handwritten Mathematical Expressions
by: Nguyen, Cuong Tuan, et al.
Published: (2025)
by: Nguyen, Cuong Tuan, et al.
Published: (2025)
RAPID: Retrieval-Augmented Parallel Inference Drafting for Text-Based Video Event Retrieval
by: Nguyen, Long, et al.
Published: (2025)
by: Nguyen, Long, et al.
Published: (2025)
Interactive Interface For Semantic Segmentation Dataset Synthesis
by: Tran, Ngoc-Do, et al.
Published: (2025)
by: Tran, Ngoc-Do, et al.
Published: (2025)
CPAM: Context-Preserving Adaptive Manipulation for Zero-Shot Real Image Editing
by: Vo, Dinh-Khoi, et al.
Published: (2025)
by: Vo, Dinh-Khoi, et al.
Published: (2025)
PANDORA: Pixel-wise Attention Dissolution and Latent Guidance for Zero-Shot Object Removal
by: Vo, Dinh-Khoi, et al.
Published: (2026)
by: Vo, Dinh-Khoi, et al.
Published: (2026)
Rx Strategist: Prescription Verification using LLM Agents System
by: Van, Phuc Phan, et al.
Published: (2024)
by: Van, Phuc Phan, et al.
Published: (2024)
MIPIC: Matryoshka Representation Learning via Self-Distilled Intra-Relational and Progressive Information Chaining
by: Huy, Phung Gia, et al.
Published: (2026)
by: Huy, Phung Gia, et al.
Published: (2026)
A Class of Accelerated Fixed-Point-Based Methods with Delayed Inexact Oracles and Its Applications
by: Nguyen-Trung, Nghia, et al.
Published: (2025)
by: Nguyen-Trung, Nghia, et al.
Published: (2025)
Unbiased and Biased Variance-Reduced Forward-Reflected-Backward Splitting Methods for Stochastic Composite Inclusions
by: Tran-Dinh, Quoc, et al.
Published: (2026)
by: Tran-Dinh, Quoc, et al.
Published: (2026)
VFOG: Variance-Reduced Fast Optimistic Gradient Methods for a Class of Nonmonotone Generalized Equations
by: Tran-Dinh, Quoc, et al.
Published: (2025)
by: Tran-Dinh, Quoc, et al.
Published: (2025)
Accelerated Extragradient-Type Methods -- Part 2: Generalization and Sublinear Convergence Rates under Co-Hypomonotonicity
by: Tran-Dinh, Quoc, et al.
Published: (2025)
by: Tran-Dinh, Quoc, et al.
Published: (2025)
Revisiting Extragradient-Type Methods -- Part 1: Generalizations and Sublinear Convergence Rates
by: Tran-Dinh, Quoc, et al.
Published: (2024)
by: Tran-Dinh, Quoc, et al.
Published: (2024)
Fusionista2.0: Efficiency Retrieval System for Large-Scale Datasets
by: Le, Huy M., et al.
Published: (2025)
by: Le, Huy M., et al.
Published: (2025)
Fourier-Attentive Representation Learning: A Fourier-Guided Framework for Few-Shot Generalization in Vision-Language Models
by: Pham, Hieu Dinh Trung, et al.
Published: (2025)
by: Pham, Hieu Dinh Trung, et al.
Published: (2025)
FedMAC: Tackling Partial-Modality Missing in Federated Learning with Cross-Modal Aggregation and Contrastive Regularization
by: Nguyen, Manh Duong, et al.
Published: (2024)
by: Nguyen, Manh Duong, et al.
Published: (2024)
CamoFA: A Learnable Fourier-based Augmentation for Camouflage Segmentation
by: Le, Minh-Quan, et al.
Published: (2023)
by: Le, Minh-Quan, et al.
Published: (2023)
TaleForge: Interactive Multimodal System for Personalized Story Creation
by: Nguyen, Minh-Loi, et al.
Published: (2025)
by: Nguyen, Minh-Loi, et al.
Published: (2025)
LGCA: Enhancing Semantic Representation via Progressive Expansion
by: Cao, Thanh Hieu, et al.
Published: (2025)
by: Cao, Thanh Hieu, et al.
Published: (2025)
Towards Convergence Rates for Parameter Estimation in Gaussian-gated Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2023)
by: Nguyen, Huy, et al.
Published: (2023)
Handling Supervision Scarcity in Chest X-ray Classification: Long-Tailed and Zero-Shot Learning
by: Pham, Ha-Hieu, et al.
Published: (2026)
by: Pham, Ha-Hieu, et al.
Published: (2026)
Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
by: Quy, Nguyen Lam Phu, et al.
Published: (2025)
by: Quy, Nguyen Lam Phu, et al.
Published: (2025)
GenKOL: Modular Generative AI Framework For Scalable Virtual KOL Generation
by: To, Tan-Hiep, et al.
Published: (2025)
by: To, Tan-Hiep, et al.
Published: (2025)
Enhancing Video Summarization with Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
ReCap: Better Gaussian Relighting with Cross-Environment Captures
by: Li, Jingzhi, et al.
Published: (2024)
by: Li, Jingzhi, et al.
Published: (2024)
Efficient Rhodamine B removal by photocatalysis using polyacrylonitrile/Ag2S nanofibers prepared by combining electrospinning technology and facile gas–solid reaction
by: Vu Dinh Thao, et al.
Published: (2024)
by: Vu Dinh Thao, et al.
Published: (2024)
Label-Efficient Cross-Modality Generalization for Liver Segmentation in Multi-Phase MRI
by: Bui-Tran, Quang-Khai, et al.
Published: (2025)
by: Bui-Tran, Quang-Khai, et al.
Published: (2025)
Similar Items
-
VisionGuard: Synergistic Framework for Helmet Violation Detection
by: Nguyen, Lam-Huy, et al.
Published: (2025) -
EventCap
by: Nguyen, Phuc-Tan, et al.
Published: (2024) -
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
by: Vo, Dinh-Khoi, et al.
Published: (2025) -
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
by: Nguyen, Hieu, et al.
Published: (2025) -
SAMURAI: Shape-Aware Multimodal Retrieval for 3D Object Identification
by: Vo, Dinh-Khoi, et al.
Published: (2025)