Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
Fuente:
arXiv
Guardado en:
| Autores principales: | Um, Sung Jin, Kim, Dongjin, Lee, Sangmin, Kim, Jung Uk |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Object-aware Sound Source Localization via Audio-Visual Scene Understanding
por: Um, Sung Jin, et al.
Publicado: (2025)
por: Um, Sung Jin, et al.
Publicado: (2025)
Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge
por: Kim, Dongjin, et al.
Publicado: (2024)
por: Kim, Dongjin, et al.
Publicado: (2024)
Video Anomaly Detection with Structured Keywords
por: Foltz, Thomas
Publicado: (2025)
por: Foltz, Thomas
Publicado: (2025)
See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection
por: Lee, YuEun, et al.
Publicado: (2025)
por: Lee, YuEun, et al.
Publicado: (2025)
Background-aware Moment Detection for Video Moment Retrieval
por: Jung, Minjoon, et al.
Publicado: (2023)
por: Jung, Minjoon, et al.
Publicado: (2023)
HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting
por: Lee, Jeongeun, et al.
Publicado: (2025)
por: Lee, Jeongeun, et al.
Publicado: (2025)
Unleash the Potential of CLIP for Video Highlight Detection
por: Han, Donghoon, et al.
Publicado: (2024)
por: Han, Donghoon, et al.
Publicado: (2024)
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
por: Paul, Dhiman, et al.
Publicado: (2024)
por: Paul, Dhiman, et al.
Publicado: (2024)
Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection
por: Yang, Jin, et al.
Publicado: (2024)
por: Yang, Jin, et al.
Publicado: (2024)
What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models
por: Kim, Junho, et al.
Publicado: (2024)
por: Kim, Junho, et al.
Publicado: (2024)
What and When to Look?: Temporal Span Proposal Network for Video Relation Detection
por: Woo, Sangmin, et al.
Publicado: (2021)
por: Woo, Sangmin, et al.
Publicado: (2021)
Task Prototype-Based Knowledge Retrieval for Multi-Task Learning from Partially Annotated Data
por: Oh, Youngmin, et al.
Publicado: (2026)
por: Oh, Youngmin, et al.
Publicado: (2026)
GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
por: Sun, Yunzhuo, et al.
Publicado: (2024)
por: Sun, Yunzhuo, et al.
Publicado: (2024)
Keyword-Oriented Multimodal Modeling for Euphemism Identification
por: Hu, Yuxue, et al.
Publicado: (2025)
por: Hu, Yuxue, et al.
Publicado: (2025)
Rethinking Multi-Condition DiTs: Eliminating Redundant Attention via Position-Alignment and Keyword-Scoping
por: Zhou, Chao, et al.
Publicado: (2026)
por: Zhou, Chao, et al.
Publicado: (2026)
MVMR: A New Framework for Evaluating Faithfulness of Video Moment Retrieval against Multiple Distractors
por: Yang, Nakyeong, et al.
Publicado: (2023)
por: Yang, Nakyeong, et al.
Publicado: (2023)
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
por: Kim, Ji-Hyeon, et al.
Publicado: (2026)
por: Kim, Ji-Hyeon, et al.
Publicado: (2026)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
por: Woo, Sangmin, et al.
Publicado: (2024)
por: Woo, Sangmin, et al.
Publicado: (2024)
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
por: Kim, Hyung Kyu, et al.
Publicado: (2025)
por: Kim, Hyung Kyu, et al.
Publicado: (2025)
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
por: Jin, Hoiyeong, et al.
Publicado: (2025)
por: Jin, Hoiyeong, et al.
Publicado: (2025)
Learning from Oblivion: Predicting Knowledge Overflowed Weights via Retrodiction of Forgetting
por: Jang, Jinhyeok, et al.
Publicado: (2025)
por: Jang, Jinhyeok, et al.
Publicado: (2025)
Catch-Up Mix: Catch-Up Class for Struggling Filters in CNN
por: Kang, Minsoo, et al.
Publicado: (2024)
por: Kang, Minsoo, et al.
Publicado: (2024)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
por: Park, Kyu Ri, et al.
Publicado: (2024)
por: Park, Kyu Ri, et al.
Publicado: (2024)
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
por: Kim, Junho, et al.
Publicado: (2024)
por: Kim, Junho, et al.
Publicado: (2024)
FreeTimeGS++: Secrets of Dynamic Gaussian Splatting and Their Principles
por: Lee, Lucas Yunkyu, et al.
Publicado: (2026)
por: Lee, Lucas Yunkyu, et al.
Publicado: (2026)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
por: Yuan, Huaying, et al.
Publicado: (2025)
por: Yuan, Huaying, et al.
Publicado: (2025)
DIAMOND: An LLM-Driven Agent for Context-Aware Baseball Highlight Summarization
por: Kang, Jeonghun, et al.
Publicado: (2025)
por: Kang, Jeonghun, et al.
Publicado: (2025)
Leveraging Textual Compositional Reasoning for Robust Change Captioning
por: Park, Kyu Ri, et al.
Publicado: (2025)
por: Park, Kyu Ri, et al.
Publicado: (2025)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
por: Jung, Mingi, et al.
Publicado: (2025)
por: Jung, Mingi, et al.
Publicado: (2025)
Denoising Task Routing for Diffusion Models
por: Park, Byeongjun, et al.
Publicado: (2023)
por: Park, Byeongjun, et al.
Publicado: (2023)
Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation
por: Liu, Pinxin, et al.
Publicado: (2025)
por: Liu, Pinxin, et al.
Publicado: (2025)
SLVideo: A Sign Language Video Moment Retrieval Framework
por: Martins, Gonçalo Vinagre, et al.
Publicado: (2024)
por: Martins, Gonçalo Vinagre, et al.
Publicado: (2024)
DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
por: Park, Jinyoung, et al.
Publicado: (2025)
por: Park, Jinyoung, et al.
Publicado: (2025)
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
por: Chen, Houlun, et al.
Publicado: (2024)
por: Chen, Houlun, et al.
Publicado: (2024)
Discovering and Mitigating Visual Biases through Keyword Explanation
por: Kim, Younghyun, et al.
Publicado: (2023)
por: Kim, Younghyun, et al.
Publicado: (2023)
A Context-aware Attention and Graph Neural Network-based Multimodal Framework for Misogyny Detection
por: Rehman, Mohammad Zia Ur, et al.
Publicado: (2025)
por: Rehman, Mohammad Zia Ur, et al.
Publicado: (2025)
Moment- and Power-Spectrum-Based Gaussianity Regularization for Text-to-Image Models
por: Hwang, Jisung, et al.
Publicado: (2025)
por: Hwang, Jisung, et al.
Publicado: (2025)
EditSplat: Multi-View Fusion and Attention-Guided Optimization for View-Consistent 3D Scene Editing with 3D Gaussian Splatting
por: Lee, Dong In, et al.
Publicado: (2024)
por: Lee, Dong In, et al.
Publicado: (2024)
GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval
por: Zhang, Shihang, et al.
Publicado: (2026)
por: Zhang, Shihang, et al.
Publicado: (2026)
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
por: Park, Geon, et al.
Publicado: (2025)
por: Park, Geon, et al.
Publicado: (2025)
Ejemplares similares
-
Object-aware Sound Source Localization via Audio-Visual Scene Understanding
por: Um, Sung Jin, et al.
Publicado: (2025) -
Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge
por: Kim, Dongjin, et al.
Publicado: (2024) -
Video Anomaly Detection with Structured Keywords
por: Foltz, Thomas
Publicado: (2025) -
See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection
por: Lee, YuEun, et al.
Publicado: (2025) -
Background-aware Moment Detection for Video Moment Retrieval
por: Jung, Minjoon, et al.
Publicado: (2023)