Guardado en:
| Autores principales: | Chang, Aiden, De Melo, Celso, Lukin, Stephanie M. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2509.16421 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Visual Agentic Memory: Enabling Online Long Video Understanding via Online Indexing, Hierarchical Memory, and Agentic Retrieval
por: Li, Aiden Yiliu, et al.
Publicado: (2026)
por: Li, Aiden Yiliu, et al.
Publicado: (2026)
What and When to Look?: Temporal Span Proposal Network for Video Relation Detection
por: Woo, Sangmin, et al.
Publicado: (2021)
por: Woo, Sangmin, et al.
Publicado: (2021)
Look Twice: Training-Free Evidence Highlighting in Multimodal Large Language Models
por: Morini, Marco, et al.
Publicado: (2026)
por: Morini, Marco, et al.
Publicado: (2026)
ViLCo-Bench: VIdeo Language COntinual learning Benchmark
por: Tang, Tianqi, et al.
Publicado: (2024)
por: Tang, Tianqi, et al.
Publicado: (2024)
Unleash the Potential of CLIP for Video Highlight Detection
por: Han, Donghoon, et al.
Publicado: (2024)
por: Han, Donghoon, et al.
Publicado: (2024)
Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
por: Zhang, Chenshuang, et al.
Publicado: (2025)
por: Zhang, Chenshuang, et al.
Publicado: (2025)
LookAhead Tuning: Safer Language Models via Partial Answer Previews
por: Liu, Kangwei, et al.
Publicado: (2025)
por: Liu, Kangwei, et al.
Publicado: (2025)
Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
por: Shabtay, Nimrod, et al.
Publicado: (2026)
por: Shabtay, Nimrod, et al.
Publicado: (2026)
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
por: Wang, Xingrui, et al.
Publicado: (2025)
por: Wang, Xingrui, et al.
Publicado: (2025)
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
por: Barbakos, Spyros, et al.
Publicado: (2025)
por: Barbakos, Spyros, et al.
Publicado: (2025)
Thinking Ahead: Foresight Intelligence in MLLMs and World Models
por: Gong, Zhantao, et al.
Publicado: (2025)
por: Gong, Zhantao, et al.
Publicado: (2025)
A Modern Look at Simplicity Bias in Image Classification Tasks
por: Chang, Xiaoguang, et al.
Publicado: (2025)
por: Chang, Xiaoguang, et al.
Publicado: (2025)
AI-Generated Images: What Humans and Machines See When They Look at the Same Image
por: Poletti, Silvia, et al.
Publicado: (2026)
por: Poletti, Silvia, et al.
Publicado: (2026)
GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
por: Sun, Yunzhuo, et al.
Publicado: (2024)
por: Sun, Yunzhuo, et al.
Publicado: (2024)
Predicting the Next Action by Modeling the Abstract Goal
por: Roy, Debaditya, et al.
Publicado: (2022)
por: Roy, Debaditya, et al.
Publicado: (2022)
What Matters in Range View 3D Object Detection
por: Wilson, Benjamin, et al.
Publicado: (2024)
por: Wilson, Benjamin, et al.
Publicado: (2024)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
por: Hao, Haihong, et al.
Publicado: (2026)
por: Hao, Haihong, et al.
Publicado: (2026)
Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection
por: Yang, Jin, et al.
Publicado: (2024)
por: Yang, Jin, et al.
Publicado: (2024)
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
por: Um, Sung Jin, et al.
Publicado: (2025)
por: Um, Sung Jin, et al.
Publicado: (2025)
Overcoming Semantic Dilution in Transformer-Based Next Frame Prediction
por: Nguyen, Hy, et al.
Publicado: (2025)
por: Nguyen, Hy, et al.
Publicado: (2025)
Automated Detection of Sport Highlights from Audio and Video Sources
por: Della Santa, Francesco, et al.
Publicado: (2025)
por: Della Santa, Francesco, et al.
Publicado: (2025)
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
por: Paul, Dhiman, et al.
Publicado: (2024)
por: Paul, Dhiman, et al.
Publicado: (2024)
Memorize What Matters: Emergent Scene Decomposition from Multitraverse
por: Li, Yiming, et al.
Publicado: (2024)
por: Li, Yiming, et al.
Publicado: (2024)
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
por: Ren, Shuhuai, et al.
Publicado: (2025)
por: Ren, Shuhuai, et al.
Publicado: (2025)
Modality Translation for Object Detection Adaptation Without Forgetting Prior Knowledge
por: Medeiros, Heitor Rapela, et al.
Publicado: (2024)
por: Medeiros, Heitor Rapela, et al.
Publicado: (2024)
ETA: Efficiency through Thinking Ahead, A Dual Approach to Self-Driving with Large Models
por: Hamdan, Shadi, et al.
Publicado: (2025)
por: Hamdan, Shadi, et al.
Publicado: (2025)
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
por: Zhou, Chunting, et al.
Publicado: (2024)
por: Zhou, Chunting, et al.
Publicado: (2024)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
por: Ji, Longbin, et al.
Publicado: (2026)
por: Ji, Longbin, et al.
Publicado: (2026)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
por: Tian, Keyu, et al.
Publicado: (2024)
por: Tian, Keyu, et al.
Publicado: (2024)
Generating Narrated Lecture Videos from Slides with Synchronized Highlights
por: Holmberg, Alexander
Publicado: (2025)
por: Holmberg, Alexander
Publicado: (2025)
Fostering Video Reasoning via Next-Event Prediction
por: Wang, Haonan, et al.
Publicado: (2025)
por: Wang, Haonan, et al.
Publicado: (2025)
Looking into Concept Explanation Methods for Diabetic Retinopathy Classification
por: Storås, Andrea M., et al.
Publicado: (2024)
por: Storås, Andrea M., et al.
Publicado: (2024)
What Makes a Maze Look Like a Maze?
por: Hsu, Joy, et al.
Publicado: (2024)
por: Hsu, Joy, et al.
Publicado: (2024)
Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection
por: Chandhok, Shivam, et al.
Publicado: (2025)
por: Chandhok, Shivam, et al.
Publicado: (2025)
What Matters in Practical Learned Image Compression
por: Tatwawadi, Kedar, et al.
Publicado: (2026)
por: Tatwawadi, Kedar, et al.
Publicado: (2026)
Object Aware Egocentric Online Action Detection
por: An, Joungbin, et al.
Publicado: (2024)
por: An, Joungbin, et al.
Publicado: (2024)
What to Do Next? Memorizing skills from Egocentric Instructional Video
por: Bi, Jing, et al.
Publicado: (2025)
por: Bi, Jing, et al.
Publicado: (2025)
What Happens Next? Anticipating Future Motion by Generating Point Trajectories
por: Boduljak, Gabrijel, et al.
Publicado: (2025)
por: Boduljak, Gabrijel, et al.
Publicado: (2025)
What Matters for Scalable and Robust Learning in End-to-End Driving Planners?
por: Holtz, David, et al.
Publicado: (2026)
por: Holtz, David, et al.
Publicado: (2026)
What Matters to You? Towards Visual Representation Alignment for Robot Learning
por: Tian, Ran, et al.
Publicado: (2023)
por: Tian, Ran, et al.
Publicado: (2023)
Ejemplares similares
-
Visual Agentic Memory: Enabling Online Long Video Understanding via Online Indexing, Hierarchical Memory, and Agentic Retrieval
por: Li, Aiden Yiliu, et al.
Publicado: (2026) -
What and When to Look?: Temporal Span Proposal Network for Video Relation Detection
por: Woo, Sangmin, et al.
Publicado: (2021) -
Look Twice: Training-Free Evidence Highlighting in Multimodal Large Language Models
por: Morini, Marco, et al.
Publicado: (2026) -
ViLCo-Bench: VIdeo Language COntinual learning Benchmark
por: Tang, Tianqi, et al.
Publicado: (2024) -
Unleash the Potential of CLIP for Video Highlight Detection
por: Han, Donghoon, et al.
Publicado: (2024)