FALCONEye: Finding Answers and Localizing Content in ONE-hour-long videos with multi-modal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Plou, Carlos, Borja, Cesar, Martinez-Cantin, Ruben, Murillo, Ana C. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SSeg: Active Sparse Point-Label Augmentation for Semantic Segmentation
by: Borja, Cesar, et al.
Published: (2025)
by: Borja, Cesar, et al.
Published: (2025)
CARLOR @ Ego4D Step Grounding Challenge: Bayesian temporal-order priors for test time refinement
by: Plou, Carlos, et al.
Published: (2024)
by: Plou, Carlos, et al.
Published: (2024)
Gen-Swarms: Adapting Deep Generative Models to Swarms of Drones
by: Plou, Carlos, et al.
Published: (2024)
by: Plou, Carlos, et al.
Published: (2024)
EventSleep: Sleep Activity Recognition with Event Cameras
by: Plou, Carlos, et al.
Published: (2024)
by: Plou, Carlos, et al.
Published: (2024)
DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric Videos
by: Mur-Labadia, Lorenzo, et al.
Published: (2025)
by: Mur-Labadia, Lorenzo, et al.
Published: (2025)
Semantic and structural image segmentation for prosthetic vision
by: Sanchez-Garcia, Melani, et al.
Published: (2018)
by: Sanchez-Garcia, Melani, et al.
Published: (2018)
Active Exploration in Bayesian Model-based Reinforcement Learning for Robot Manipulation
by: Plou, Carlos, et al.
Published: (2024)
by: Plou, Carlos, et al.
Published: (2024)
Towards multi-modal forgery representation learning for AI-generated video detection and localization
by: Le, Dat, et al.
Published: (2026)
by: Le, Dat, et al.
Published: (2026)
ZARRIO @ Ego4D Short Term Object Interaction Anticipation Challenge: Leveraging Affordances and Attention-based models for STA
by: Mur-Labadia, Lorenzo, et al.
Published: (2024)
by: Mur-Labadia, Lorenzo, et al.
Published: (2024)
AFF-ttention! Affordances and Attention models for Short-Term Object Interaction Anticipation
by: Mur-Labadia, Lorenzo, et al.
Published: (2024)
by: Mur-Labadia, Lorenzo, et al.
Published: (2024)
Event Transformer+. A multi-purpose solution for efficient event data processing
by: Sabater, Alberto, et al.
Published: (2022)
by: Sabater, Alberto, et al.
Published: (2022)
Online Topological Localization for Navigation Assistance in Bronchoscopy
by: Tomasini, Clara, et al.
Published: (2025)
by: Tomasini, Clara, et al.
Published: (2025)
Integrating Affordances and Attention models for Short-Term Object Interaction Anticipation
by: Labadia, Lorenzo Mur, et al.
Published: (2026)
by: Labadia, Lorenzo Mur, et al.
Published: (2026)
Answer-Consistent Chain-of-thought Reinforcement Learning For Multi-modal Large Langauge Models
by: Huang, Minbin, et al.
Published: (2025)
by: Huang, Minbin, et al.
Published: (2025)
Online pre-training with long-form videos
by: Kato, Itsuki, et al.
Published: (2024)
by: Kato, Itsuki, et al.
Published: (2024)
Influence of field of view in visual prostheses design: Analysis with a VR system
by: Sanchez-Garcia, Melani, et al.
Published: (2025)
by: Sanchez-Garcia, Melani, et al.
Published: (2025)
An Empirical Study on How Video-LLMs Answer Video Questions
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
AMEGO: Active Memory from long EGOcentric videos
by: Goletto, Gabriele, et al.
Published: (2024)
by: Goletto, Gabriele, et al.
Published: (2024)
Koala: Key frame-conditioned long video-LLM
by: Tan, Reuben, et al.
Published: (2024)
by: Tan, Reuben, et al.
Published: (2024)
Serial fusion of multi-modal biometric systems
by: Marcialis, Gian Luca, et al.
Published: (2024)
by: Marcialis, Gian Luca, et al.
Published: (2024)
O-MaMa: Learning Object Mask Matching between Egocentric and Exocentric Views
by: Mur-Labadia, Lorenzo, et al.
Published: (2025)
by: Mur-Labadia, Lorenzo, et al.
Published: (2025)
Balancing long- and short-term dynamics for the modeling of saliency in videos
by: Wulff, Theodor, et al.
Published: (2025)
by: Wulff, Theodor, et al.
Published: (2025)
Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
by: Venkataramanan, Shashanka, et al.
Published: (2023)
by: Venkataramanan, Shashanka, et al.
Published: (2023)
EchoONE: Segmenting Multiple echocardiography Planes in One Model
by: Hu, Jiongtong, et al.
Published: (2024)
by: Hu, Jiongtong, et al.
Published: (2024)
VrdONE: One-stage Video Visual Relation Detection
by: Jiang, Xinjie, et al.
Published: (2024)
by: Jiang, Xinjie, et al.
Published: (2024)
LightDepth: Single-View Depth Self-Supervision from Illumination Decline
by: Rodríguez-Puigvert, Javier, et al.
Published: (2023)
by: Rodríguez-Puigvert, Javier, et al.
Published: (2023)
Segmenting the motion components of a video: A long-term unsupervised model
by: Meunier, Etienne, et al.
Published: (2023)
by: Meunier, Etienne, et al.
Published: (2023)
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
by: Lu, Jinda, et al.
Published: (2025)
by: Lu, Jinda, et al.
Published: (2025)
Cross-modal ultra-scale learning with tri-modalities of renal biopsy images for glomerular multi-disease auxiliary diagnosis
by: Long, Kaixing, et al.
Published: (2025)
by: Long, Kaixing, et al.
Published: (2025)
End-to-end multi-modal product matching in fashion e-commerce
by: Tóth, Sándor, et al.
Published: (2024)
by: Tóth, Sándor, et al.
Published: (2024)
FungiTastic: A multi-modal dataset and benchmark for image categorization
by: Picek, Lukas, et al.
Published: (2024)
by: Picek, Lukas, et al.
Published: (2024)
Phantom: Subject-consistent video generation via cross-modal alignment
by: Liu, Lijie, et al.
Published: (2025)
by: Liu, Lijie, et al.
Published: (2025)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization
by: Xing, Ling, et al.
Published: (2024)
by: Xing, Ling, et al.
Published: (2024)
Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
by: He, Junwen, et al.
Published: (2024)
by: He, Junwen, et al.
Published: (2024)
ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization
by: Mohammadshirazi, Ahmad, et al.
Published: (2025)
by: Mohammadshirazi, Ahmad, et al.
Published: (2025)
A multi-modal dataset for insect biodiversity with imagery and DNA at the trap and individual level
by: Orsholm, Johanna, et al.
Published: (2025)
by: Orsholm, Johanna, et al.
Published: (2025)
DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding
by: Ahmadian, Mona, et al.
Published: (2025)
by: Ahmadian, Mona, et al.
Published: (2025)
Do We Need to Design Specific Diffusion Models for Different Tasks? Try ONE-PIC
by: Tao, Ming, et al.
Published: (2024)
by: Tao, Ming, et al.
Published: (2024)
Can VLMs be used on videos for action recognition? LLMs are Visual Reasoning Coordinators
by: Lunia, Harsh
Published: (2024)
by: Lunia, Harsh
Published: (2024)
Similar Items
-
SSeg: Active Sparse Point-Label Augmentation for Semantic Segmentation
by: Borja, Cesar, et al.
Published: (2025) -
CARLOR @ Ego4D Step Grounding Challenge: Bayesian temporal-order priors for test time refinement
by: Plou, Carlos, et al.
Published: (2024) -
Gen-Swarms: Adapting Deep Generative Models to Swarms of Drones
by: Plou, Carlos, et al.
Published: (2024) -
EventSleep: Sleep Activity Recognition with Event Cameras
by: Plou, Carlos, et al.
Published: (2024) -
DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric Videos
by: Mur-Labadia, Lorenzo, et al.
Published: (2025)