Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Nishimura, Taichi, Nakada, Shota, Munakata, Hokuto, Komatsu, Tatsuya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
by: Nakada, Shota, et al.
Published: (2024)
by: Nakada, Shota, et al.
Published: (2024)
Vision-Language Models Learn Super Images for Efficient Partially Relevant Video Retrieval
by: Nishimura, Taichi, et al.
Published: (2023)
by: Nishimura, Taichi, et al.
Published: (2023)
Language-based Audio Moment Retrieval
by: Munakata, Hokuto, et al.
Published: (2024)
by: Munakata, Hokuto, et al.
Published: (2024)
On the Audio Hallucinations in Large Audio-Video Language Models
by: Nishimura, Taichi, et al.
Published: (2024)
by: Nishimura, Taichi, et al.
Published: (2024)
Recipe Generation from Unsegmented Cooking Videos
by: Nishimura, Taichi, et al.
Published: (2022)
by: Nishimura, Taichi, et al.
Published: (2022)
TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
by: Masumura, Ryo, et al.
Published: (2025)
by: Masumura, Ryo, et al.
Published: (2025)
Hallucination Localization in Video Captioning
by: Nakada, Shota, et al.
Published: (2025)
by: Nakada, Shota, et al.
Published: (2025)
Beyond Coarse-Grained Matching in Video-Text Retrieval
by: Chen, Aozhu, et al.
Published: (2024)
by: Chen, Aozhu, et al.
Published: (2024)
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
by: Ishikawa, Yuchi, et al.
Published: (2025)
by: Ishikawa, Yuchi, et al.
Published: (2025)
Context-Enhanced Video Moment Retrieval with Large Language Models
by: Liu, Weijia, et al.
Published: (2024)
by: Liu, Weijia, et al.
Published: (2024)
Videogenic: Identifying Highlight Moments in Videos with Professional Photographs as a Prior
by: Lin, David Chuan-En, et al.
Published: (2022)
by: Lin, David Chuan-En, et al.
Published: (2022)
Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval
by: Ding, Yiming, et al.
Published: (2026)
by: Ding, Yiming, et al.
Published: (2026)
BioVL-QR: Egocentric Biochemical Vision-and-Language Dataset Using Micro QR Codes
by: Nishimoto, Tomohiro, et al.
Published: (2024)
by: Nishimoto, Tomohiro, et al.
Published: (2024)
Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval
by: Zhang, Bolin, et al.
Published: (2026)
by: Zhang, Bolin, et al.
Published: (2026)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
by: Xu, Yifang, et al.
Published: (2025)
by: Xu, Yifang, et al.
Published: (2025)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
HA-FGOVD: Highlighting Fine-grained Attributes via Explicit Linear Composition for Open-Vocabulary Object Detection
by: Ma, Yuqi, et al.
Published: (2024)
by: Ma, Yuqi, et al.
Published: (2024)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
by: Ji, Yatai, et al.
Published: (2024)
by: Ji, Yatai, et al.
Published: (2024)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
by: Zhang, Xueqiao, et al.
Published: (2025)
by: Zhang, Xueqiao, et al.
Published: (2025)
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
by: Wang, Jiapeng, et al.
Published: (2024)
by: Wang, Jiapeng, et al.
Published: (2024)
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
by: Barbakos, Spyros, et al.
Published: (2025)
by: Barbakos, Spyros, et al.
Published: (2025)
MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
by: Luo, Ziyang, et al.
Published: (2024)
by: Luo, Ziyang, et al.
Published: (2024)
Towards Event-oriented Long Video Understanding
by: Du, Yifan, et al.
Published: (2024)
by: Du, Yifan, et al.
Published: (2024)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
WordArt Designer API: User-Driven Artistic Typography Synthesis with Large Language Models on ModelScope
by: He, Jun-Yan, et al.
Published: (2024)
by: He, Jun-Yan, et al.
Published: (2024)
LocoMotion: Learning Motion-Focused Video-Language Representations
by: Doughty, Hazel, et al.
Published: (2024)
by: Doughty, Hazel, et al.
Published: (2024)
HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning
by: Yang, Yiqing, et al.
Published: (2025)
by: Yang, Yiqing, et al.
Published: (2025)
ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos
by: Sukhani, Siddhant, et al.
Published: (2025)
by: Sukhani, Siddhant, et al.
Published: (2025)
COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark
by: Maeda, Koki, et al.
Published: (2024)
by: Maeda, Koki, et al.
Published: (2024)
Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos
by: Stamatakis, Markos, et al.
Published: (2025)
by: Stamatakis, Markos, et al.
Published: (2025)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
by: Jiang, Chen, et al.
Published: (2023)
by: Jiang, Chen, et al.
Published: (2023)
EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation
by: Xu, Jiaqi, et al.
Published: (2024)
by: Xu, Jiaqi, et al.
Published: (2024)
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries
by: Munakata, Hokuto, et al.
Published: (2025)
by: Munakata, Hokuto, et al.
Published: (2025)
MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling
by: Xu, Jiaqi, et al.
Published: (2023)
by: Xu, Jiaqi, et al.
Published: (2023)
Harmfully Manipulated Images Matter in Multimodal Misinformation Detection
by: Wang, Bing, et al.
Published: (2024)
by: Wang, Bing, et al.
Published: (2024)
Similar Items
-
DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
by: Nakada, Shota, et al.
Published: (2024) -
Vision-Language Models Learn Super Images for Efficient Partially Relevant Video Retrieval
by: Nishimura, Taichi, et al.
Published: (2023) -
Language-based Audio Moment Retrieval
by: Munakata, Hokuto, et al.
Published: (2024) -
On the Audio Hallucinations in Large Audio-Video Language Models
by: Nishimura, Taichi, et al.
Published: (2024) -
Recipe Generation from Unsegmented Cooking Videos
by: Nishimura, Taichi, et al.
Published: (2022)