Understanding Long Videos via LLM-Powered Entity Relation Graphs
Fuente:
arXiv
Saved in:
| Main Authors: | Chu, Meng, Li, Yicong, Chua, Tat-Seng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
by: Liu, Han, et al.
Published: (2025)
by: Liu, Han, et al.
Published: (2025)
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
by: Gao, Haowen, et al.
Published: (2025)
by: Gao, Haowen, et al.
Published: (2025)
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
by: Yin, Xinlei, et al.
Published: (2026)
by: Yin, Xinlei, et al.
Published: (2026)
DRC: Enhancing Personalized Image Generation via Disentangled Representation Composition
by: Xu, Yiyan, et al.
Published: (2025)
by: Xu, Yiyan, et al.
Published: (2025)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
by: Li, Yongqi, et al.
Published: (2024)
by: Li, Yongqi, et al.
Published: (2024)
EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models
by: Meng, GuangHao, et al.
Published: (2025)
by: Meng, GuangHao, et al.
Published: (2025)
LongVidSearch: An Agentic Benchmark for Multi-hop Evidence Retrieval Planning in Long Videos
by: Yu, Rongyi, et al.
Published: (2026)
by: Yu, Rongyi, et al.
Published: (2026)
I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking
by: Liu, Ziyan, et al.
Published: (2025)
by: Liu, Ziyan, et al.
Published: (2025)
Entity Image and Mixed-Modal Image Retrieval Datasets
by: Blaga, Cristian-Ioan, et al.
Published: (2025)
by: Blaga, Cristian-Ioan, et al.
Published: (2025)
LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
by: Cai, Qifeng, et al.
Published: (2025)
by: Cai, Qifeng, et al.
Published: (2025)
Automatic Funny Scene Extraction from Long-form Cinematic Videos
by: Paul, Sibendu, et al.
Published: (2026)
by: Paul, Sibendu, et al.
Published: (2026)
U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs
by: Li, Xiaojie, et al.
Published: (2025)
by: Li, Xiaojie, et al.
Published: (2025)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
PEARL: Personalized Streaming Video Understanding Model
by: Zheng, Yuanhong, et al.
Published: (2026)
by: Zheng, Yuanhong, et al.
Published: (2026)
Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search
by: Tang, Hengzhu, et al.
Published: (2025)
by: Tang, Hengzhu, et al.
Published: (2025)
Video Editing for Video Retrieval
by: Zhu, Bin, et al.
Published: (2024)
by: Zhu, Bin, et al.
Published: (2024)
POBEVM: Real-time Video Matting via Progressively Optimize the Target Body and Edge
by: Xian, Jianming
Published: (2024)
by: Xian, Jianming
Published: (2024)
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
by: Ren, Xubin, et al.
Published: (2025)
by: Ren, Xubin, et al.
Published: (2025)
MealRec: Multi-granularity Sequential Modeling via Hierarchical Diffusion Models for Micro-Video Recommendation
by: Dong, Xinxin, et al.
Published: (2026)
by: Dong, Xinxin, et al.
Published: (2026)
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
by: Reddy, Arun, et al.
Published: (2025)
by: Reddy, Arun, et al.
Published: (2025)
Diff-VPS: Video Polyp Segmentation via a Multi-task Diffusion Network with Adversarial Temporal Reasoning
by: Lu, Yingling, et al.
Published: (2024)
by: Lu, Yingling, et al.
Published: (2024)
GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval
by: Sun, Chengsong, et al.
Published: (2025)
by: Sun, Chengsong, et al.
Published: (2025)
FashionStylist: An Expert Knowledge-enhanced Multimodal Dataset for Fashion Understanding
by: Feng, Kaidong, et al.
Published: (2026)
by: Feng, Kaidong, et al.
Published: (2026)
GraphRevisedIE: Multimodal Information Extraction with Graph-Revised Network
by: Cao, Panfeng, et al.
Published: (2024)
by: Cao, Panfeng, et al.
Published: (2024)
From Videos to Indexed Knowledge Graphs -- Framework to Marry Methods for Multimodal Content Analysis and Understanding
by: Rizk, Basem, et al.
Published: (2025)
by: Rizk, Basem, et al.
Published: (2025)
Adapting MLLMs for Nuanced Video Retrieval
by: Bagad, Piyush, et al.
Published: (2025)
by: Bagad, Piyush, et al.
Published: (2025)
Can I Trust Your Answer? Visually Grounded Video Question Answering
by: Xiao, Junbin, et al.
Published: (2023)
by: Xiao, Junbin, et al.
Published: (2023)
RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval
by: Skow, Tyler, et al.
Published: (2026)
by: Skow, Tyler, et al.
Published: (2026)
Event-aware Video Corpus Moment Retrieval
by: Hou, Danyang, et al.
Published: (2024)
by: Hou, Danyang, et al.
Published: (2024)
Towards Human-Like Machine Comprehension: Few-Shot Relational Learning in Visually-Rich Documents
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
A Flexible and Scalable Framework for Video Moment Search
by: Zhang, Chongzhi, et al.
Published: (2025)
by: Zhang, Chongzhi, et al.
Published: (2025)
Improving Video Corpus Moment Retrieval with Partial Relevance Enhancement
by: Hou, Danyang, et al.
Published: (2024)
by: Hou, Danyang, et al.
Published: (2024)
Multimodal Language Models for Domain-Specific Procedural Video Summarization
by: Hussain, Nafisa
Published: (2024)
by: Hussain, Nafisa
Published: (2024)
SHE-Net: Syntax-Hierarchy-Enhanced Text-Video Retrieval
by: Yu, Xuzheng, et al.
Published: (2024)
by: Yu, Xuzheng, et al.
Published: (2024)
NextAds: Towards Next-generation Personalized Video Advertising
by: Xu, Yiyan, et al.
Published: (2026)
by: Xu, Yiyan, et al.
Published: (2026)
Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking
by: Xu, Zhengfei, et al.
Published: (2024)
by: Xu, Zhengfei, et al.
Published: (2024)
AI Powered High Quality Text to Video Generation with Enhanced Temporal Consistency
by: Patel, Piyushkumar
Published: (2025)
by: Patel, Piyushkumar
Published: (2025)
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
by: Li, Minghan, et al.
Published: (2026)
by: Li, Minghan, et al.
Published: (2026)
Personalized Video Summarization using Text-Based Queries and Conditional Modeling
by: Huang, Jia-Hong
Published: (2024)
by: Huang, Jia-Hong
Published: (2024)
MARQUIS: A Three-Stage Pipeline for Video Retrieval-Augmented Generation
by: Chakraborty, Debashish, et al.
Published: (2026)
by: Chakraborty, Debashish, et al.
Published: (2026)
Similar Items
-
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
by: Liu, Han, et al.
Published: (2025) -
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
by: Gao, Haowen, et al.
Published: (2025) -
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
by: Yin, Xinlei, et al.
Published: (2026) -
DRC: Enhancing Personalized Image Generation via Disentangled Representation Composition
by: Xu, Yiyan, et al.
Published: (2025) -
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
by: Li, Yongqi, et al.
Published: (2024)