AHA -- Predicting What Matters Next: Online Highlight Detection Without Looking Ahead
Fuente:
arXiv
Saved in:
| Main Authors: | Chang, Aiden, De Melo, Celso, Lukin, Stephanie M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Agentic Memory: Enabling Online Long Video Understanding via Online Indexing, Hierarchical Memory, and Agentic Retrieval
by: Li, Aiden Yiliu, et al.
Published: (2026)
by: Li, Aiden Yiliu, et al.
Published: (2026)
What and When to Look?: Temporal Span Proposal Network for Video Relation Detection
by: Woo, Sangmin, et al.
Published: (2021)
by: Woo, Sangmin, et al.
Published: (2021)
Look Twice: Training-Free Evidence Highlighting in Multimodal Large Language Models
by: Morini, Marco, et al.
Published: (2026)
by: Morini, Marco, et al.
Published: (2026)
Unleash the Potential of CLIP for Video Highlight Detection
by: Han, Donghoon, et al.
Published: (2024)
by: Han, Donghoon, et al.
Published: (2024)
Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
by: Zhang, Chenshuang, et al.
Published: (2025)
by: Zhang, Chenshuang, et al.
Published: (2025)
Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
by: Shabtay, Nimrod, et al.
Published: (2026)
by: Shabtay, Nimrod, et al.
Published: (2026)
ViLCo-Bench: VIdeo Language COntinual learning Benchmark
by: Tang, Tianqi, et al.
Published: (2024)
by: Tang, Tianqi, et al.
Published: (2024)
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
by: Barbakos, Spyros, et al.
Published: (2025)
by: Barbakos, Spyros, et al.
Published: (2025)
LookAhead Tuning: Safer Language Models via Partial Answer Previews
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
Thinking Ahead: Foresight Intelligence in MLLMs and World Models
by: Gong, Zhantao, et al.
Published: (2025)
by: Gong, Zhantao, et al.
Published: (2025)
Predicting the Next Action by Modeling the Abstract Goal
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
by: Sun, Yunzhuo, et al.
Published: (2024)
by: Sun, Yunzhuo, et al.
Published: (2024)
A Modern Look at Simplicity Bias in Image Classification Tasks
by: Chang, Xiaoguang, et al.
Published: (2025)
by: Chang, Xiaoguang, et al.
Published: (2025)
AI-Generated Images: What Humans and Machines See When They Look at the Same Image
by: Poletti, Silvia, et al.
Published: (2026)
by: Poletti, Silvia, et al.
Published: (2026)
Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection
by: Yang, Jin, et al.
Published: (2024)
by: Yang, Jin, et al.
Published: (2024)
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
by: Um, Sung Jin, et al.
Published: (2025)
by: Um, Sung Jin, et al.
Published: (2025)
What Matters in Range View 3D Object Detection
by: Wilson, Benjamin, et al.
Published: (2024)
by: Wilson, Benjamin, et al.
Published: (2024)
Overcoming Semantic Dilution in Transformer-Based Next Frame Prediction
by: Nguyen, Hy, et al.
Published: (2025)
by: Nguyen, Hy, et al.
Published: (2025)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
by: Hao, Haihong, et al.
Published: (2026)
by: Hao, Haihong, et al.
Published: (2026)
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
by: Paul, Dhiman, et al.
Published: (2024)
by: Paul, Dhiman, et al.
Published: (2024)
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
by: Ren, Shuhuai, et al.
Published: (2025)
by: Ren, Shuhuai, et al.
Published: (2025)
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
by: Wang, Xingrui, et al.
Published: (2025)
by: Wang, Xingrui, et al.
Published: (2025)
Modality Translation for Object Detection Adaptation Without Forgetting Prior Knowledge
by: Medeiros, Heitor Rapela, et al.
Published: (2024)
by: Medeiros, Heitor Rapela, et al.
Published: (2024)
Memorize What Matters: Emergent Scene Decomposition from Multitraverse
by: Li, Yiming, et al.
Published: (2024)
by: Li, Yiming, et al.
Published: (2024)
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
by: Zhou, Chunting, et al.
Published: (2024)
by: Zhou, Chunting, et al.
Published: (2024)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
by: Ji, Longbin, et al.
Published: (2026)
by: Ji, Longbin, et al.
Published: (2026)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
by: Tian, Keyu, et al.
Published: (2024)
by: Tian, Keyu, et al.
Published: (2024)
Automated Detection of Sport Highlights from Audio and Video Sources
by: Della Santa, Francesco, et al.
Published: (2025)
by: Della Santa, Francesco, et al.
Published: (2025)
ETA: Efficiency through Thinking Ahead, A Dual Approach to Self-Driving with Large Models
by: Hamdan, Shadi, et al.
Published: (2025)
by: Hamdan, Shadi, et al.
Published: (2025)
Generating Narrated Lecture Videos from Slides with Synchronized Highlights
by: Holmberg, Alexander
Published: (2025)
by: Holmberg, Alexander
Published: (2025)
Looking into Concept Explanation Methods for Diabetic Retinopathy Classification
by: Storås, Andrea M., et al.
Published: (2024)
by: Storås, Andrea M., et al.
Published: (2024)
Fostering Video Reasoning via Next-Event Prediction
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
Object Aware Egocentric Online Action Detection
by: An, Joungbin, et al.
Published: (2024)
by: An, Joungbin, et al.
Published: (2024)
Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection
by: Chandhok, Shivam, et al.
Published: (2025)
by: Chandhok, Shivam, et al.
Published: (2025)
No Labels, No Look-Ahead: Unsupervised Online Video Stabilization with Classical Priors
by: Liu, Tao, et al.
Published: (2026)
by: Liu, Tao, et al.
Published: (2026)
EcoWeedNet: A Lightweight and Automated Weed Detection Method for Sustainable Next-Generation Agricultural Consumer Electronics
by: Khater, Omar H., et al.
Published: (2025)
by: Khater, Omar H., et al.
Published: (2025)
Cut2Next: Generating Next Shot via In-Context Tuning
by: He, Jingwen, et al.
Published: (2025)
by: He, Jingwen, et al.
Published: (2025)
Focus on What Matters: Two-Stage ROI-Aware Refinement for Anatomy-Preserving Fetal Ultrasound Reconstruction
by: Abbes, Ines, et al.
Published: (2026)
by: Abbes, Ines, et al.
Published: (2026)
Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Just Add Geometry: Gradient-Free Open-Vocabulary 3D Detection Without Human-in-the-Loop
by: Goel, Atharv, et al.
Published: (2025)
by: Goel, Atharv, et al.
Published: (2025)
Similar Items
-
Visual Agentic Memory: Enabling Online Long Video Understanding via Online Indexing, Hierarchical Memory, and Agentic Retrieval
by: Li, Aiden Yiliu, et al.
Published: (2026) -
What and When to Look?: Temporal Span Proposal Network for Video Relation Detection
by: Woo, Sangmin, et al.
Published: (2021) -
Look Twice: Training-Free Evidence Highlighting in Multimodal Large Language Models
by: Morini, Marco, et al.
Published: (2026) -
Unleash the Potential of CLIP for Video Highlight Detection
by: Han, Donghoon, et al.
Published: (2024) -
Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
by: Zhang, Chenshuang, et al.
Published: (2025)