Integrated Semantic and Temporal Alignment for Interactive Video Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Luu, Thanh-Danh, Dinh, Le-Vu Nguyen, Tran, Duc-Thien, Bui, Duy-Bao, Le, Nam-Tien, Nhu, Tinh-Anh Nguyen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fact-Checking at Scale: Multimodal AI for Authenticity and Context Verification in Online Media
by: Phan, Van-Hoang, et al.
Published: (2025)
by: Phan, Van-Hoang, et al.
Published: (2025)
Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction
by: Nguyen, Cam-Van Thi, et al.
Published: (2023)
by: Nguyen, Cam-Van Thi, et al.
Published: (2023)
Modeling the Impacts of Swipe Delay on User Quality of Experience in Short Video Streaming
by: Nguyen, Duc V., et al.
Published: (2026)
by: Nguyen, Duc V., et al.
Published: (2026)
Ada2I: Enhancing Modality Balance for Multimodal Conversational Emotion Recognition
by: Nguyen, Cam-Van Thi, et al.
Published: (2024)
by: Nguyen, Cam-Van Thi, et al.
Published: (2024)
CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization
by: Le, Anh-Duy, et al.
Published: (2026)
by: Le, Anh-Duy, et al.
Published: (2026)
Short-Form Video Viewing Behavior Analysis and Multi-Step Viewing Time Prediction
by: Yen, Vu Thi Hai, et al.
Published: (2026)
by: Yen, Vu Thi Hai, et al.
Published: (2026)
Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion
by: Thanh, Toan Le Ngo, et al.
Published: (2025)
by: Thanh, Toan Le Ngo, et al.
Published: (2025)
A Subjective Quality Evaluation of 3D Mesh with Dynamic Level of Detail in Virtual Reality
by: Nguyen, Duc, et al.
Published: (2024)
by: Nguyen, Duc, et al.
Published: (2024)
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
by: Tong, Xinyi, et al.
Published: (2025)
by: Tong, Xinyi, et al.
Published: (2025)
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
by: Nguyen, Ngoc-Son, et al.
Published: (2026)
by: Nguyen, Ngoc-Son, et al.
Published: (2026)
FedMAC: Tackling Partial-Modality Missing in Federated Learning with Cross-Modal Aggregation and Contrastive Regularization
by: Nguyen, Manh Duong, et al.
Published: (2024)
by: Nguyen, Manh Duong, et al.
Published: (2024)
Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
by: Nguyen, Hieu Minh, et al.
Published: (2025)
by: Nguyen, Hieu Minh, et al.
Published: (2025)
The State-of-the-Art in Lifelog Retrieval: A Review of Progress at the ACM Lifelog Search Challenge Workshop 2022-24
by: Tran, Allie, et al.
Published: (2025)
by: Tran, Allie, et al.
Published: (2025)
E-FreeM2: Efficient Training-Free Multi-Scale and Cross-Modal News Verification via MLLMs
by: Phan, Van-Hoang, et al.
Published: (2025)
by: Phan, Van-Hoang, et al.
Published: (2025)
Proposing Smart System for Detecting and Monitoring Vehicle Using Multiobject Multicamera Tracking
by: Phat Nguyen Huu, et al.
Published: (2024)
by: Phat Nguyen Huu, et al.
Published: (2024)
Subjective Quality Assessment of Dynamic 3D Meshes in Virtual Reality Environment
by: Nguyen, Duc V., et al.
Published: (2026)
by: Nguyen, Duc V., et al.
Published: (2026)
MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
by: Vu, Huu-An, et al.
Published: (2025)
by: Vu, Huu-An, et al.
Published: (2025)
Towards Efficient and Robust Moment Retrieval System: A Unified Framework for Multi-Granularity Models and Temporal Reranking
by: Tran, Huu-Loc, et al.
Published: (2025)
by: Tran, Huu-Loc, et al.
Published: (2025)
Task Presentation and Human Perception in Interactive Video Retrieval
by: Willis, Nina, et al.
Published: (2024)
by: Willis, Nina, et al.
Published: (2024)
Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model
by: Cuong, Dinh Viet, et al.
Published: (2025)
by: Cuong, Dinh Viet, et al.
Published: (2025)
KAT: Dependency-aware Automated API Testing with Large Language Models
by: Le, Tri, et al.
Published: (2024)
by: Le, Tri, et al.
Published: (2024)
OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
by: Tran, Quang-Linh, et al.
Published: (2025)
by: Tran, Quang-Linh, et al.
Published: (2025)
STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
by: Ren, Yong, et al.
Published: (2024)
by: Ren, Yong, et al.
Published: (2024)
Towards Open-Vocabulary Video Semantic Segmentation
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
Efficient INT8 Single-Image Super-Resolution via Deployment-Aware Quantization and Teacher-Guided Training
by: Nguyen, Pham Phuong Nam, et al.
Published: (2026)
by: Nguyen, Pham Phuong Nam, et al.
Published: (2026)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
by: Xiao, Jian, et al.
Published: (2025)
by: Xiao, Jian, et al.
Published: (2025)
Contestable Multi-Agent Debate with Arena-based Argumentative Computation for Multimedia Verification
by: Nguyen, Truong Thanh Hung, et al.
Published: (2026)
by: Nguyen, Truong Thanh Hung, et al.
Published: (2026)
Dual-Stream Decoupled Learning for Temporal Consistency and Speaker Interaction in AVSD
by: Xiao, Junhao, et al.
Published: (2025)
by: Xiao, Junhao, et al.
Published: (2025)
Multimodal Learned Sparse Retrieval for Image Suggestion
by: Nguyen, Thong, et al.
Published: (2024)
by: Nguyen, Thong, et al.
Published: (2024)
A Dual-Module Denoising Approach with Curriculum Learning for Enhancing Multimodal Aspect-Based Sentiment Analysis
by: Van Doan, Nguyen, et al.
Published: (2024)
by: Van Doan, Nguyen, et al.
Published: (2024)
Semi-Supervised Semantic Segmentation using Redesigned Self-Training for White Blood Cells
by: Luu, Vinh Quoc, et al.
Published: (2024)
by: Luu, Vinh Quoc, et al.
Published: (2024)
Federated Prompt-Tuning with Heterogeneous and Incomplete Multimodal Client Data
by: Phung, Thu Hang, et al.
Published: (2026)
by: Phung, Thu Hang, et al.
Published: (2026)
Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial Alignment
by: Nguyen, Cong-Duy, et al.
Published: (2023)
by: Nguyen, Cong-Duy, et al.
Published: (2023)
How2Compress: Scalable and Efficient Edge Video Analytics via Adaptive Granular Video Compression
by: Wu, Yuheng, et al.
Published: (2025)
by: Wu, Yuheng, et al.
Published: (2025)
A Lightweight Moment Retrieval System with Global Re-Ranking and Robust Adaptive Bidirectional Temporal Search
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
MCPNS: A Macropixel Collocated Position and Its Neighbors Search for Plenoptic 2.0 Video Coding
by: Van Duong, Vinh, et al.
Published: (2023)
by: Van Duong, Vinh, et al.
Published: (2023)
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
by: Zhang, Long, et al.
Published: (2025)
by: Zhang, Long, et al.
Published: (2025)
SimInterview: Transforming Business Education through Large Language Model-Based Simulated Multilingual Interview Training System
by: Nguyen, Truong Thanh Hung, et al.
Published: (2025)
by: Nguyen, Truong Thanh Hung, et al.
Published: (2025)
Similar Items
-
Fact-Checking at Scale: Multimodal AI for Authenticity and Context Verification in Online Media
by: Phan, Van-Hoang, et al.
Published: (2025) -
Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction
by: Nguyen, Cam-Van Thi, et al.
Published: (2023) -
Modeling the Impacts of Swipe Delay on User Quality of Experience in Short Video Streaming
by: Nguyen, Duc V., et al.
Published: (2026) -
Ada2I: Enhancing Modality Balance for Multimodal Conversational Emotion Recognition
by: Nguyen, Cam-Van Thi, et al.
Published: (2024) -
CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization
by: Le, Anh-Duy, et al.
Published: (2026)