FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries
Fuente:
arXiv
Saved in:
| Main Authors: | You, Qijie, Liang, Hao, Chen, Mingrui, Zeng, Bohan, Qiang, Meiyi, Wong, Zhenhao, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
by: Han, ZhaoYang, et al.
Published: (2025)
by: Han, ZhaoYang, et al.
Published: (2025)
Multimodal LLM-based Query Paraphrasing for Video Search
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
Start from Video-Music Retrieval: An Inter-Intra Modal Loss for Cross Modal Retrieval
by: Chen, Zeyu, et al.
Published: (2024)
by: Chen, Zeyu, et al.
Published: (2024)
HippoMM: Hippocampal-inspired Multimodal Memory for Long Audiovisual Event Understanding
by: Lin, Yueqian, et al.
Published: (2025)
by: Lin, Yueqian, et al.
Published: (2025)
AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
by: Chen, Zixuan, et al.
Published: (2026)
by: Chen, Zixuan, et al.
Published: (2026)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
by: Wang, Shaoguang, et al.
Published: (2026)
by: Wang, Shaoguang, et al.
Published: (2026)
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos
by: Feng, Hengyi, et al.
Published: (2026)
by: Feng, Hengyi, et al.
Published: (2026)
Divide and Conquer: Multimodal Video Deepfake Detection via Cross-Modal Fusion and Localization
by: Li, Qingcao, et al.
Published: (2026)
by: Li, Qingcao, et al.
Published: (2026)
Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings
by: Clarke, Jason, et al.
Published: (2025)
by: Clarke, Jason, et al.
Published: (2025)
Cross-Modal Watermarking for Authentic Audio Recovery and Tamper Localization in Synthesized Audiovisual Forgeries
by: Kim, Minyoung, et al.
Published: (2025)
by: Kim, Minyoung, et al.
Published: (2025)
Reply with Sticker: New Dataset and Model for Sticker Retrieval
by: Liang, Bin, et al.
Published: (2024)
by: Liang, Bin, et al.
Published: (2024)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
by: Fang, Xinyu, et al.
Published: (2024)
by: Fang, Xinyu, et al.
Published: (2024)
"The Intangible Victory", Interactive Audiovisual Installation
by: Tsioutas, Konstantinos, et al.
Published: (2026)
by: Tsioutas, Konstantinos, et al.
Published: (2026)
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
by: Li, Jie, et al.
Published: (2025)
by: Li, Jie, et al.
Published: (2025)
Integrated Semantic and Temporal Alignment for Interactive Video Retrieval
by: Luu, Thanh-Danh, et al.
Published: (2025)
by: Luu, Thanh-Danh, et al.
Published: (2025)
Task Presentation and Human Perception in Interactive Video Retrieval
by: Willis, Nina, et al.
Published: (2024)
by: Willis, Nina, et al.
Published: (2024)
An Empirical Comparison of Video Frame Sampling Methods for Multi-Modal RAG Retrieval
by: Kandhare, Mahesh, et al.
Published: (2024)
by: Kandhare, Mahesh, et al.
Published: (2024)
GTPBD-MM: A Global Terraced Parcel and Boundary Dataset with Multi-Modality
by: Zhang, Zhiwei, et al.
Published: (2026)
by: Zhang, Zhiwei, et al.
Published: (2026)
Vidformer: Drop-in Declarative Optimization for Rendering Video-Native Query Results
by: Winecki, Dominik, et al.
Published: (2026)
by: Winecki, Dominik, et al.
Published: (2026)
CLAIP-Emo: Parameter-Efficient Adaptation of Language-supervised models for In-the-Wild Audiovisual Emotion Recognition
by: Chen, Yin, et al.
Published: (2025)
by: Chen, Yin, et al.
Published: (2025)
Mining the Social Fabric: Unveiling Communities for Fake News Detection in Short Videos
by: Gong, Haisong, et al.
Published: (2025)
by: Gong, Haisong, et al.
Published: (2025)
AsCL: An Asymmetry-sensitive Contrastive Learning Method for Image-Text Retrieval with Cross-Modal Fusion
by: Gong, Ziyu, et al.
Published: (2024)
by: Gong, Ziyu, et al.
Published: (2024)
Cross-Modality and Within-Modality Regularization for Audio-Visual DeepFake Detection
by: Zou, Heqing, et al.
Published: (2024)
by: Zou, Heqing, et al.
Published: (2024)
Modeling the Impacts of Swipe Delay on User Quality of Experience in Short Video Streaming
by: Nguyen, Duc V., et al.
Published: (2026)
by: Nguyen, Duc V., et al.
Published: (2026)
PAL: Prompting Analytic Learning with Missing Modality for Multi-Modal Class-Incremental Learning
by: Yue, Xianghu, et al.
Published: (2025)
by: Yue, Xianghu, et al.
Published: (2025)
Structure-Aware Residual-Center Representation for Self-Supervised Open-Set 3D Cross-Modal Retrieval
by: Xu, Yang, et al.
Published: (2024)
by: Xu, Yang, et al.
Published: (2024)
Personalized Playback Technology: How Short Video Services Create Excellent User Experience
by: Deng, Weihui, et al.
Published: (2024)
by: Deng, Weihui, et al.
Published: (2024)
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
by: Geng, Tiantian, et al.
Published: (2024)
by: Geng, Tiantian, et al.
Published: (2024)
VCR: Video representation for Contextual Retrieval
by: Nir, Oron, et al.
Published: (2024)
by: Nir, Oron, et al.
Published: (2024)
PromptHash: Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing Retrieval
by: Zou, Qiang, et al.
Published: (2025)
by: Zou, Qiang, et al.
Published: (2025)
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
by: Zhou, Yang-Hao, et al.
Published: (2026)
by: Zhou, Yang-Hao, et al.
Published: (2026)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions
by: Wang, Tianshi, et al.
Published: (2023)
by: Wang, Tianshi, et al.
Published: (2023)
SVD: Spatial Video Dataset
by: Izadimehr, M. H., et al.
Published: (2025)
by: Izadimehr, M. H., et al.
Published: (2025)
GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
by: Yang, Quanwei, et al.
Published: (2025)
by: Yang, Quanwei, et al.
Published: (2025)
Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos
by: Tian, Chong, et al.
Published: (2026)
by: Tian, Chong, et al.
Published: (2026)
From Query to Explanation: Uni-RAG for Multi-Modal Retrieval-Augmented Learning in STEM
by: Wu, Xinyi, et al.
Published: (2025)
by: Wu, Xinyi, et al.
Published: (2025)
Automatically Generating High-Precision Simulated Road Networking in Traffic Scenario
by: Xie, Liang, et al.
Published: (2025)
by: Xie, Liang, et al.
Published: (2025)
NeedleDB: A Generative-AI Based System for Accurate and Efficient Image Retrieval using Complex Natural Language Queries
by: Erfanian, Mahdi, et al.
Published: (2026)
by: Erfanian, Mahdi, et al.
Published: (2026)
Similar Items
-
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
by: Han, ZhaoYang, et al.
Published: (2025) -
Multimodal LLM-based Query Paraphrasing for Video Search
by: Wu, Jiaxin, et al.
Published: (2024) -
Start from Video-Music Retrieval: An Inter-Intra Modal Loss for Cross Modal Retrieval
by: Chen, Zeyu, et al.
Published: (2024) -
HippoMM: Hippocampal-inspired Multimodal Memory for Long Audiovisual Event Understanding
by: Lin, Yueqian, et al.
Published: (2025) -
AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
by: Chen, Zixuan, et al.
Published: (2026)