Aligning Moments in Time using Video Queries
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Yogesh, Agarwal, Uday, Gupta, Manish, Mishra, Anand |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
by: Kumar, Yogesh, et al.
Published: (2025)
by: Kumar, Yogesh, et al.
Published: (2025)
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
by: Flanagan, Kevin, et al.
Published: (2025)
by: Flanagan, Kevin, et al.
Published: (2025)
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
by: Kulkarni, Yogesh, et al.
Published: (2024)
by: Kulkarni, Yogesh, et al.
Published: (2024)
PatentLMM: Large Multimodal Model for Generating Descriptions for Patent Figures
by: Shukla, Shreya, et al.
Published: (2025)
by: Shukla, Shreya, et al.
Published: (2025)
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
by: Gatti, Prajwal, et al.
Published: (2025)
by: Gatti, Prajwal, et al.
Published: (2025)
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
by: Kumar, Yogesh
Published: (2025)
by: Kumar, Yogesh
Published: (2025)
Beyond Caption-Based Queries for Video Moment Retrieval
by: Pujol-Perich, David, et al.
Published: (2026)
by: Pujol-Perich, David, et al.
Published: (2026)
Selective Query-guided Debiasing for Video Corpus Moment Retrieval
by: Yoon, Sunjae, et al.
Published: (2022)
by: Yoon, Sunjae, et al.
Published: (2022)
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis
by: Lokesh, K, et al.
Published: (2026)
by: Lokesh, K, et al.
Published: (2026)
Multimodal Query-guided Object Localization
by: Tripathi, Aditay, et al.
Published: (2022)
by: Tripathi, Aditay, et al.
Published: (2022)
VideoLLM Benchmarks and Evaluation: A Survey
by: Kumar, Yogesh
Published: (2025)
by: Kumar, Yogesh
Published: (2025)
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection
by: Xu, Yifang, et al.
Published: (2025)
by: Xu, Yifang, et al.
Published: (2025)
GranAlign: Granularity-Aware Alignment Framework for Zero-Shot Video Moment Retrieval
by: Jeon, Mingyu, et al.
Published: (2026)
by: Jeon, Mingyu, et al.
Published: (2026)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
ECIS-VQG: Generation of Entity-centric Information-seeking Questions from Videos
by: Phukan, Arpan, et al.
Published: (2024)
by: Phukan, Arpan, et al.
Published: (2024)
Stable Mean Teacher for Semi-supervised Video Action Detection
by: Kumar, Akash, et al.
Published: (2024)
by: Kumar, Akash, et al.
Published: (2024)
VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
Background-aware Moment Detection for Video Moment Retrieval
by: Jung, Minjoon, et al.
Published: (2023)
by: Jung, Minjoon, et al.
Published: (2023)
SketchQL Demonstration: Zero-shot Video Moment Querying with Sketches
by: Wu, Renzhi, et al.
Published: (2024)
by: Wu, Renzhi, et al.
Published: (2024)
BAM-DETR: Boundary-Aligned Moment Detection Transformer for Temporal Sentence Grounding in Videos
by: Lee, Pilhyeon, et al.
Published: (2023)
by: Lee, Pilhyeon, et al.
Published: (2023)
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
Moment Quantization for Video Temporal Grounding
by: Sun, Xiaolong, et al.
Published: (2025)
by: Sun, Xiaolong, et al.
Published: (2025)
PRISM: Perceptual Recognition for Identifying Standout Moments in Human-Centric Keyframe Extraction
by: Cakmak, Mert Can, et al.
Published: (2025)
by: Cakmak, Mert Can, et al.
Published: (2025)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025)
by: Yang, Yanlai, et al.
Published: (2025)
Towards Making Flowchart Images Machine Interpretable
by: Shukla, Shreya, et al.
Published: (2025)
by: Shukla, Shreya, et al.
Published: (2025)
Text-Video Multi-Grained Integration for Video Moment Montage
by: Yin, Zhihui, et al.
Published: (2024)
by: Yin, Zhihui, et al.
Published: (2024)
CAVE-Net: Classifying Abnormalities in Video Capsule Endoscopy
by: Harish, Ishita, et al.
Published: (2024)
by: Harish, Ishita, et al.
Published: (2024)
Q-Align: Alleviating Attention Leakage in Zero-Shot Appearance Transfer via Query-Query Alignment
by: Kim, Namu, et al.
Published: (2025)
by: Kim, Namu, et al.
Published: (2025)
Semi-supervised Active Learning for Video Action Detection
by: Singh, Ayush, et al.
Published: (2023)
by: Singh, Ayush, et al.
Published: (2023)
Object-Centric Framework for Video Moment Retrieval
by: Li, Zongyao, et al.
Published: (2025)
by: Li, Zongyao, et al.
Published: (2025)
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
by: Reilly, Dominick, et al.
Published: (2025)
by: Reilly, Dominick, et al.
Published: (2025)
Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval
by: Liu, Weijia, et al.
Published: (2025)
by: Liu, Weijia, et al.
Published: (2025)
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval
by: Cai, Weitong, et al.
Published: (2024)
by: Cai, Weitong, et al.
Published: (2024)
Generative Video Diffusion for Unseen Novel Semantic Video Moment Retrieval
by: Luo, Dezhao, et al.
Published: (2024)
by: Luo, Dezhao, et al.
Published: (2024)
AlignIT: Enhancing Prompt Alignment in Customization of Text-to-Image Models
by: Agarwal, Aishwarya, et al.
Published: (2024)
by: Agarwal, Aishwarya, et al.
Published: (2024)
VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale
by: Kulkarni, Parth Parag, et al.
Published: (2026)
by: Kulkarni, Parth Parag, et al.
Published: (2026)
Index-Aligned Query Distillation for Transformer-based Incremental Object Detection
by: Ma, Mingxiao, et al.
Published: (2025)
by: Ma, Mingxiao, et al.
Published: (2025)
Similar Items
-
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
by: Kumar, Yogesh, et al.
Published: (2025) -
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
by: Flanagan, Kevin, et al.
Published: (2025) -
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
by: Kulkarni, Yogesh, et al.
Published: (2024) -
PatentLMM: Large Multimodal Model for Generating Descriptions for Patent Figures
by: Shukla, Shreya, et al.
Published: (2025) -
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
by: Gatti, Prajwal, et al.
Published: (2025)