EA-VTR: Event-Aware Video-Text Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Zongyang, Zhang, Ziqi, Chen, Yuxin, Qi, Zhongang, Yuan, Chunfeng, Li, Bing, Luo, Yingmin, Li, Xu, Qi, Xiaojuan, Shan, Ying, Hu, Weiming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?
by: Chen, Yuxin, et al.
Published: (2024)
by: Chen, Yuxin, et al.
Published: (2024)
mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval
by: Yang, Yuxin, et al.
Published: (2026)
by: Yang, Yuxin, et al.
Published: (2026)
E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding
by: Liu, Ye, et al.
Published: (2024)
by: Liu, Ye, et al.
Published: (2024)
DetailFusion: A Dual-branch Framework with Detail Enhancement for Composed Image Retrieval
by: Yang, Yuxin, et al.
Published: (2025)
by: Yang, Yuxin, et al.
Published: (2025)
MMhops-R1: Multimodal Multi-hop Reasoning
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
by: Lu, Yifan, et al.
Published: (2025)
by: Lu, Yifan, et al.
Published: (2025)
PosterLLaVa: Constructing a Unified Multi-modal Layout Generator with LLM
by: Yang, Tao, et al.
Published: (2024)
by: Yang, Tao, et al.
Published: (2024)
Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval
by: Liu, Haowei, et al.
Published: (2024)
by: Liu, Haowei, et al.
Published: (2024)
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning
by: Lu, Yifan, et al.
Published: (2023)
by: Lu, Yifan, et al.
Published: (2023)
Making MLLMs Blind: Adversarial Smuggling Attacks in MLLM Content Moderation
by: Li, Zhiheng, et al.
Published: (2026)
by: Li, Zhiheng, et al.
Published: (2026)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
by: Liu, Ye, et al.
Published: (2025)
by: Liu, Ye, et al.
Published: (2025)
SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion Model
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation
by: Li, Xuewei, et al.
Published: (2023)
by: Li, Xuewei, et al.
Published: (2023)
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation
by: Zheng, Guangcong, et al.
Published: (2023)
by: Zheng, Guangcong, et al.
Published: (2023)
VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings
by: Li, Zhiheng, et al.
Published: (2026)
by: Li, Zhiheng, et al.
Published: (2026)
EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields
by: Yang, Zhaoyang, et al.
Published: (2026)
by: Yang, Zhaoyang, et al.
Published: (2026)
CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
Taming Rectified Flow for Inversion and Editing
by: Wang, Jiangshan, et al.
Published: (2024)
by: Wang, Jiangshan, et al.
Published: (2024)
DOGR: Towards Versatile Visual Document Grounding and Referring
by: Zhou, Yinan, et al.
Published: (2024)
by: Zhou, Yinan, et al.
Published: (2024)
Learning from Neighbors: Category Extrapolation for Long-Tail Learning
by: Zhao, Shizhen, et al.
Published: (2024)
by: Zhao, Shizhen, et al.
Published: (2024)
Mono2Stereo: A Benchmark and Empirical Study for Stereo Conversion
by: Yu, Songsong, et al.
Published: (2025)
by: Yu, Songsong, et al.
Published: (2025)
StyleAdapter: A Unified Stylized Image Generation Model
by: Wang, Zhouxia, et al.
Published: (2023)
by: Wang, Zhouxia, et al.
Published: (2023)
SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
PromptIQA: Boosting the Performance and Generalization for No-Reference Image Quality Assessment via Prompts
by: Chen, Zewen, et al.
Published: (2024)
by: Chen, Zewen, et al.
Published: (2024)
MARAG-R1: Beyond Single Retriever via Reinforcement-Learned Multi-Tool Agentic Retrieval
by: Luo, Qi, et al.
Published: (2025)
by: Luo, Qi, et al.
Published: (2025)
NFT1000: A Cross-Modal Dataset for Non-Fungible Token Retrieval
by: Wang, Shuxun, et al.
Published: (2024)
by: Wang, Shuxun, et al.
Published: (2024)
Latewood width of Pinus uncinata at site VTR
by: Tejedor, Ernesto, et al.
Published: (2017)
by: Tejedor, Ernesto, et al.
Published: (2017)
Earlywood width of Pinus uncinata at site VTR
by: Tejedor, Ernesto, et al.
Published: (2017)
by: Tejedor, Ernesto, et al.
Published: (2017)
Schnelleres Bauen mit der VTR‐Bauweise
by: Mathias Daßler, et al.
Published: (2025)
by: Mathias Daßler, et al.
Published: (2025)
T2VIndexer: A Generative Video Indexer for Efficient Text-Video Retrieval
by: Li, Yili, et al.
Published: (2024)
by: Li, Yili, et al.
Published: (2024)
Graph-Aware Stealthy Poison-Text Backdoors for Text-Attributed Graphs
by: Luo, Qi, et al.
Published: (2026)
by: Luo, Qi, et al.
Published: (2026)
Triangular Decomposition of Third Order Hermitian Tensors
by: Qi, Liqun, et al.
Published: (2024)
by: Qi, Liqun, et al.
Published: (2024)
Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
by: Zhou, Shengchao, et al.
Published: (2025)
by: Zhou, Shengchao, et al.
Published: (2025)
Long-time asymptotics of the Newell equation on the line
by: Wang, Deng-Shan, et al.
Published: (2026)
by: Wang, Deng-Shan, et al.
Published: (2026)
Tree-ring width of Pinus uncinata at site VTR
by: Tejedor, Ernesto, et al.
Published: (2017)
by: Tejedor, Ernesto, et al.
Published: (2017)
LaMPE: Length-aware Multi-grained Positional Encoding for Adaptive Long-context Scaling Without Training
by: Zhang, Sikui, et al.
Published: (2025)
by: Zhang, Sikui, et al.
Published: (2025)
iMOVE: Instance-Motion-Aware Video Understanding
by: Li, Jiaze, et al.
Published: (2025)
by: Li, Jiaze, et al.
Published: (2025)
Similar Items
-
How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?
by: Chen, Yuxin, et al.
Published: (2024) -
mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA
by: Zhang, Tao, et al.
Published: (2024) -
Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval
by: Yang, Yuxin, et al.
Published: (2026) -
E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding
by: Liu, Ye, et al.
Published: (2024) -
DetailFusion: A Dual-branch Framework with Detail Enhancement for Composed Image Retrieval
by: Yang, Yuxin, et al.
Published: (2025)