Cutup and Detect: Human Fall Detection on Cutup Untrimmed Videos Using a Large Foundational Video Understanding Model
Fuente:
arXiv
Saved in:
| Main Authors: | Grutschus, Till, Karrar, Ola, Esenov, Emir, Vats, Ekta |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
by: Yang, Min, et al.
Published: (2023)
by: Yang, Min, et al.
Published: (2023)
MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos
by: Sinha, Arkaprava, et al.
Published: (2025)
by: Sinha, Arkaprava, et al.
Published: (2025)
Explainable Forensics of Manipulated Segments in Untrimmed Long Videos
by: Feng, Yue, et al.
Published: (2026)
by: Feng, Yue, et al.
Published: (2026)
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
by: Tang, Yolo Yunlong, et al.
Published: (2024)
by: Tang, Yolo Yunlong, et al.
Published: (2024)
Mamba-OTR: a Mamba-based Solution for Online Take and Release Detection from Untrimmed Egocentric Video
by: Catinello, Alessandro Sebastiano, et al.
Published: (2025)
by: Catinello, Alessandro Sebastiano, et al.
Published: (2025)
Uncovering the Handwritten Text in the Margins: End-to-end Handwritten Text Detection and Recognition
by: Cheng, Liang, et al.
Published: (2023)
by: Cheng, Liang, et al.
Published: (2023)
Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
by: Peng, Liyang, et al.
Published: (2025)
by: Peng, Liyang, et al.
Published: (2025)
MultiCounter: Multiple Action Agnostic Repetition Counting in Untrimmed Videos
by: Tang, Yin, et al.
Published: (2024)
by: Tang, Yin, et al.
Published: (2024)
Multi-Stage Boundary-Aware Transformer Network for Action Segmentation in Untrimmed Surgical Videos
by: Shuvo, Rezowan, et al.
Published: (2025)
by: Shuvo, Rezowan, et al.
Published: (2025)
Localizing Moments of Actions in Untrimmed Videos of Infants with Autism Spectrum Disorder
by: Helvaci, Halil Ismail, et al.
Published: (2024)
by: Helvaci, Halil Ismail, et al.
Published: (2024)
SDFA: Structure Aware Discriminative Feature Aggregation for Efficient Human Fall Detection in Video
by: Zahan, Sania, et al.
Published: (2025)
by: Zahan, Sania, et al.
Published: (2025)
MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
by: Feng, Yue, et al.
Published: (2025)
by: Feng, Yue, et al.
Published: (2025)
Simplifying Traffic Anomaly Detection with Video Foundation Models
by: Orlova, Svetlana, et al.
Published: (2025)
by: Orlova, Svetlana, et al.
Published: (2025)
VideoGLUE: Video General Understanding Evaluation of Foundation Models
by: Yuan, Liangzhe, et al.
Published: (2023)
by: Yuan, Liangzhe, et al.
Published: (2023)
GenVideoLens: Where LVLMs Fall Short in AI-Generated Video Detection?
by: Zou, Yueying, et al.
Published: (2026)
by: Zou, Yueying, et al.
Published: (2026)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
by: Wang, Yi, et al.
Published: (2024)
by: Wang, Yi, et al.
Published: (2024)
Foundation Models for Video Understanding: A Survey
by: Madan, Neelu, et al.
Published: (2024)
by: Madan, Neelu, et al.
Published: (2024)
FADE: A Dataset for Detecting Falling Objects around Buildings in Video
by: Tu, Zhigang, et al.
Published: (2024)
by: Tu, Zhigang, et al.
Published: (2024)
Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented Augmentation
by: Corvi, Riccardo, et al.
Published: (2025)
by: Corvi, Riccardo, et al.
Published: (2025)
What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
by: Chen, Brian, et al.
Published: (2023)
by: Chen, Brian, et al.
Published: (2023)
Fall Detection from Indoor Videos using MediaPipe and Handcrafted Feature
by: Ahmed, Fatima, et al.
Published: (2025)
by: Ahmed, Fatima, et al.
Published: (2025)
SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos
by: Wu, Jinlin, et al.
Published: (2026)
by: Wu, Jinlin, et al.
Published: (2026)
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
by: Zhang, Boqiang, et al.
Published: (2025)
by: Zhang, Boqiang, et al.
Published: (2025)
Modeling Human Skeleton Joint Dynamics for Fall Detection
by: Zahan, Sania, et al.
Published: (2025)
by: Zahan, Sania, et al.
Published: (2025)
MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
by: Yang, Garry, et al.
Published: (2025)
by: Yang, Garry, et al.
Published: (2025)
JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts
by: Son, Taein, et al.
Published: (2024)
by: Son, Taein, et al.
Published: (2024)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model
by: Xu, Lu, et al.
Published: (2024)
by: Xu, Lu, et al.
Published: (2024)
Video Anomaly Detection and Explanation via Large Language Models
by: Lv, Hui, et al.
Published: (2024)
by: Lv, Hui, et al.
Published: (2024)
How Confident are Video Models? Empowering Video Models to Express their Uncertainty
by: Mei, Zhiting, et al.
Published: (2025)
by: Mei, Zhiting, et al.
Published: (2025)
Instrument-tissue Interaction Detection Framework for Surgical Video Understanding
by: Lin, Wenjun, et al.
Published: (2024)
by: Lin, Wenjun, et al.
Published: (2024)
Harnessing Large Language Models for Training-free Video Anomaly Detection
by: Zanella, Luca, et al.
Published: (2024)
by: Zanella, Luca, et al.
Published: (2024)
Follow the Rules: Reasoning for Video Anomaly Detection with Large Language Models
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding
by: Yang, Ruoliu, et al.
Published: (2026)
by: Yang, Ruoliu, et al.
Published: (2026)
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
by: Peng, Yi-Xing, et al.
Published: (2025)
by: Peng, Yi-Xing, et al.
Published: (2025)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
by: Shi, Jiapeng, et al.
Published: (2026)
by: Shi, Jiapeng, et al.
Published: (2026)
Streaming Long Video Understanding with Large Language Models
by: Qian, Rui, et al.
Published: (2024)
by: Qian, Rui, et al.
Published: (2024)
Vidi: Large Multimodal Models for Video Understanding and Editing
by: Vidi Team, et al.
Published: (2025)
by: Vidi Team, et al.
Published: (2025)
HighlightMe: Detecting Highlights from Human-Centric Videos
by: Bhattacharya, Uttaran, et al.
Published: (2021)
by: Bhattacharya, Uttaran, et al.
Published: (2021)
Contracting Skeletal Kinematics for Human-Related Video Anomaly Detection
by: Flaborea, Alessandro, et al.
Published: (2023)
by: Flaborea, Alessandro, et al.
Published: (2023)
Similar Items
-
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
by: Yang, Min, et al.
Published: (2023) -
MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos
by: Sinha, Arkaprava, et al.
Published: (2025) -
Explainable Forensics of Manipulated Segments in Untrimmed Long Videos
by: Feng, Yue, et al.
Published: (2026) -
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
by: Tang, Yolo Yunlong, et al.
Published: (2024) -
Mamba-OTR: a Mamba-based Solution for Online Take and Release Detection from Untrimmed Egocentric Video
by: Catinello, Alessandro Sebastiano, et al.
Published: (2025)