Retrieve What's Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Joo, Minseok, Park, Dogyun, Lee, Taehoon, Lee, Kyujin, Kim, Hyunwoo J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
von: Park, Dogyun, et al.
Veröffentlicht: (2025)
von: Park, Dogyun, et al.
Veröffentlicht: (2025)
Constant Acceleration Flow
von: Park, Dogyun, et al.
Veröffentlicht: (2024)
von: Park, Dogyun, et al.
Veröffentlicht: (2024)
Diffusion Prior-Based Amortized Variational Inference for Noisy Inverse Problems
von: Lee, Sojin, et al.
Veröffentlicht: (2024)
von: Lee, Sojin, et al.
Veröffentlicht: (2024)
Robust Multimodal 3D Object Detection via Modality-Agnostic Decoding and Proximity-based Modality Ensemble
von: Cha, Juhan, et al.
Veröffentlicht: (2024)
von: Cha, Juhan, et al.
Veröffentlicht: (2024)
Stochastic Conditional Diffusion Models for Robust Semantic Image Synthesis
von: Ko, Juyeon, et al.
Veröffentlicht: (2024)
von: Ko, Juyeon, et al.
Veröffentlicht: (2024)
Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
von: Lee, Ji Soo, et al.
Veröffentlicht: (2025)
von: Lee, Ji Soo, et al.
Veröffentlicht: (2025)
Probabilistic Precision and Recall Towards Reliable Evaluation of Generative Models
von: Park, Dogyun, et al.
Veröffentlicht: (2023)
von: Park, Dogyun, et al.
Veröffentlicht: (2023)
Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval
von: Ko, Dohwan, et al.
Veröffentlicht: (2025)
von: Ko, Dohwan, et al.
Veröffentlicht: (2025)
Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning
von: Choi, Seung hee, et al.
Veröffentlicht: (2026)
von: Choi, Seung hee, et al.
Veröffentlicht: (2026)
Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection
von: Yang, Chanhyeong, et al.
Veröffentlicht: (2025)
von: Yang, Chanhyeong, et al.
Veröffentlicht: (2025)
F4Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting
von: Kim, Injae, et al.
Veröffentlicht: (2026)
von: Kim, Injae, et al.
Veröffentlicht: (2026)
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
von: Park, Jihwan, et al.
Veröffentlicht: (2025)
von: Park, Jihwan, et al.
Veröffentlicht: (2025)
RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection
von: Park, Jihwan, et al.
Veröffentlicht: (2026)
von: Park, Jihwan, et al.
Veröffentlicht: (2026)
Super-class guided Transformer for Zero-Shot Attribute Classification
von: Kim, Sehyung, et al.
Veröffentlicht: (2025)
von: Kim, Sehyung, et al.
Veröffentlicht: (2025)
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
von: Yu, Jiwen, et al.
Veröffentlicht: (2025)
von: Yu, Jiwen, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Open-Vocabulary Object Detection
von: Kim, Jooyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jooyeon, et al.
Veröffentlicht: (2024)
EFlow: Fast Few-Step Video Generator Training from Scratch via Efficient Solution Flow
von: Park, Dogyun, et al.
Veröffentlicht: (2026)
von: Park, Dogyun, et al.
Veröffentlicht: (2026)
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
von: Kim, Junho, et al.
Veröffentlicht: (2024)
von: Kim, Junho, et al.
Veröffentlicht: (2024)
Revisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning
von: Kang, Minseok, et al.
Veröffentlicht: (2026)
von: Kang, Minseok, et al.
Veröffentlicht: (2026)
VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning
von: Lee, Ji Soo, et al.
Veröffentlicht: (2025)
von: Lee, Ji Soo, et al.
Veröffentlicht: (2025)
Infinite-Story: A Training-Free Consistent Text-to-Image Generation
von: Park, Jihun, et al.
Veröffentlicht: (2025)
von: Park, Jihun, et al.
Veröffentlicht: (2025)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions
von: Hur, Chan, et al.
Veröffentlicht: (2025)
von: Hur, Chan, et al.
Veröffentlicht: (2025)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
von: Jeong, Boseung, et al.
Veröffentlicht: (2025)
von: Jeong, Boseung, et al.
Veröffentlicht: (2025)
Open-Attribute Recognition for Person Retrieval: Finding People Through Distinctive and Novel Attributes
von: Park, Minjeong, et al.
Veröffentlicht: (2025)
von: Park, Minjeong, et al.
Veröffentlicht: (2025)
Sprint: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers
von: Park, Dogyun, et al.
Veröffentlicht: (2025)
von: Park, Dogyun, et al.
Veröffentlicht: (2025)
GenCLIP: Generalizing CLIP Prompts for Zero-shot Anomaly Detection
von: Kim, Donghyeong, et al.
Veröffentlicht: (2025)
von: Kim, Donghyeong, et al.
Veröffentlicht: (2025)
OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation
von: Zhao, Lin, et al.
Veröffentlicht: (2026)
von: Zhao, Lin, et al.
Veröffentlicht: (2026)
CAST: Modeling Visual State Transitions for Consistent Video Retrieval
von: Liu, Yanqing, et al.
Veröffentlicht: (2026)
von: Liu, Yanqing, et al.
Veröffentlicht: (2026)
Text-Video Retrieval with Global-Local Semantic Consistent Learning
von: Zhang, Haonan, et al.
Veröffentlicht: (2024)
von: Zhang, Haonan, et al.
Veröffentlicht: (2024)
Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering
von: Kim, Jongha, et al.
Veröffentlicht: (2026)
von: Kim, Jongha, et al.
Veröffentlicht: (2026)
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
DrVideo: Document Retrieval Based Long Video Understanding
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM
von: Park, Hyobin, et al.
Veröffentlicht: (2026)
von: Park, Hyobin, et al.
Veröffentlicht: (2026)
Point to Span: Zero-Shot Moment Retrieval for Navigating Unseen Hour-Long Videos
von: Jeon, Mingyu, et al.
Veröffentlicht: (2025)
von: Jeon, Mingyu, et al.
Veröffentlicht: (2025)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025)
Weatherproofing Retrieval for Localization with Generative AI and Geometric Consistency
von: Kalantidis, Yannis, et al.
Veröffentlicht: (2024)
von: Kalantidis, Yannis, et al.
Veröffentlicht: (2024)
vid-TLDR: Training Free Token merging for Light-weight Video Transformer
von: Choi, Joonmyung, et al.
Veröffentlicht: (2024)
von: Choi, Joonmyung, et al.
Veröffentlicht: (2024)
CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation
von: Jeon, Inseok, et al.
Veröffentlicht: (2026)
von: Jeon, Inseok, et al.
Veröffentlicht: (2026)
DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
von: Park, Jinyoung, et al.
Veröffentlicht: (2025)
von: Park, Jinyoung, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
von: Park, Dogyun, et al.
Veröffentlicht: (2025) -
Constant Acceleration Flow
von: Park, Dogyun, et al.
Veröffentlicht: (2024) -
Diffusion Prior-Based Amortized Variational Inference for Noisy Inverse Problems
von: Lee, Sojin, et al.
Veröffentlicht: (2024) -
Robust Multimodal 3D Object Detection via Modality-Agnostic Decoding and Proximity-based Modality Ensemble
von: Cha, Juhan, et al.
Veröffentlicht: (2024) -
Stochastic Conditional Diffusion Models for Robust Semantic Image Synthesis
von: Ko, Juyeon, et al.
Veröffentlicht: (2024)