FlowNar: Scalable Streaming Narration for Long-Form Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Zeyun, Martin, Manuel, Wu, Chengzhi, Schneider, David, Diederichs, Frederik, Gall, Juergen, Beyerer, Juergen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
QueryMamba: A Mamba-Based Encoder-Decoder Architecture with a Statistical Verb-Noun Interaction Module for Video Action Forecasting @ Ego4D Long-Term Action Anticipation Challenge 2024
by: Zhong, Zeyun, et al.
Published: (2024)
by: Zhong, Zeyun, et al.
Published: (2024)
A Survey on Deep Learning Techniques for Action Anticipation
by: Zhong, Zeyun, et al.
Published: (2023)
by: Zhong, Zeyun, et al.
Published: (2023)
Multi-modal Video Representation Alignment for Robust Self-supervised Driver Distraction Detection
by: Lerch, David J., et al.
Published: (2026)
by: Lerch, David J., et al.
Published: (2026)
Video Panels for Long Video Understanding
by: Doorenbos, Lars, et al.
Published: (2025)
by: Doorenbos, Lars, et al.
Published: (2025)
CamC2V: Context-aware Controllable Video Generation
by: Denninger, Luis, et al.
Published: (2025)
by: Denninger, Luis, et al.
Published: (2025)
ADA-Track++: End-to-End Multi-Camera 3D Multi-Object Tracking with Alternating Detection and Association
by: Ding, Shuxiao, et al.
Published: (2024)
by: Ding, Shuxiao, et al.
Published: (2024)
Sequence-Adaptive Video Prediction in Continuous Streams using Diffusion Noise Optimization
by: Azar, Sina Mokhtarzadeh, et al.
Published: (2025)
by: Azar, Sina Mokhtarzadeh, et al.
Published: (2025)
SAMBLE: Shape-Specific Point Cloud Sampling for an Optimal Trade-Off Between Local Detail and Global Uniformity
by: Wu, Chengzhi, et al.
Published: (2025)
by: Wu, Chengzhi, et al.
Published: (2025)
Rethinking Attention Module Design for Point Cloud Analysis
by: Wu, Chengzhi, et al.
Published: (2024)
by: Wu, Chengzhi, et al.
Published: (2024)
LC-SLab -- An Object-based Deep Learning Framework for Large-scale Land Cover Classification from Satellite Imagery and Sparse In-situ Labels
by: Leonhardt, Johannes, et al.
Published: (2025)
by: Leonhardt, Johannes, et al.
Published: (2025)
Identifying Spatio-Temporal Drivers of Extreme Events
by: Eddin, Mohamad Hakam Shams, et al.
Published: (2024)
by: Eddin, Mohamad Hakam Shams, et al.
Published: (2024)
StableMamba: Distillation-free Scaling of Large SSMs for Images and Videos
by: Suleman, Hamid, et al.
Published: (2024)
by: Suleman, Hamid, et al.
Published: (2024)
Enhancing Video-Based Robot Failure Detection Using Task Knowledge
by: Thoduka, Santosh, et al.
Published: (2025)
by: Thoduka, Santosh, et al.
Published: (2025)
Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models
by: Yi, Jinhui, et al.
Published: (2024)
by: Yi, Jinhui, et al.
Published: (2024)
Gated Temporal Diffusion for Stochastic Long-Term Dense Anticipation
by: Zatsarynna, Olga, et al.
Published: (2024)
by: Zatsarynna, Olga, et al.
Published: (2024)
SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction
by: Pallotta, Enrico, et al.
Published: (2025)
by: Pallotta, Enrico, et al.
Published: (2025)
MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Anticipation
by: Zatsarynna, Olga, et al.
Published: (2025)
by: Zatsarynna, Olga, et al.
Published: (2025)
Using Visual Anomaly Detection for Task Execution Monitoring
by: Thoduka, Santosh, et al.
Published: (2021)
by: Thoduka, Santosh, et al.
Published: (2021)
Learning a Neural Association Network for Self-supervised Multi-Object Tracking
by: Li, Shuai, et al.
Published: (2024)
by: Li, Shuai, et al.
Published: (2024)
Hierarchical Vector Quantization for Unsupervised Action Segmentation
by: Spurio, Federico, et al.
Published: (2024)
by: Spurio, Federico, et al.
Published: (2024)
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
by: Bahrami, Emad, et al.
Published: (2026)
by: Bahrami, Emad, et al.
Published: (2026)
Self-Supervised Generative-Contrastive Learning of Multi-Modal Euclidean Input for 3D Shape Latent Representations: A Dynamic Switching Approach
by: Wu, Chengzhi, et al.
Published: (2023)
by: Wu, Chengzhi, et al.
Published: (2023)
Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action Segmentation
by: Gökay, Uzay, et al.
Published: (2025)
by: Gökay, Uzay, et al.
Published: (2025)
Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset
by: Lerch, David J., et al.
Published: (2026)
by: Lerch, David J., et al.
Published: (2026)
EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
by: Pallotta, Enrico, et al.
Published: (2025)
by: Pallotta, Enrico, et al.
Published: (2025)
CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding
by: Patel, Shrenik, et al.
Published: (2025)
by: Patel, Shrenik, et al.
Published: (2025)
A Multimodal Handover Failure Detection Dataset and Baselines
by: Thoduka, Santosh, et al.
Published: (2024)
by: Thoduka, Santosh, et al.
Published: (2024)
MV-Match: Multi-View Matching for Domain-Adaptive Identification of Plant Nutrient Deficiencies
by: Yi, Jinhui, et al.
Published: (2024)
by: Yi, Jinhui, et al.
Published: (2024)
Self-Intersection-Aware 3D Human Motion Generation Using an Efficient Human Sphere Proxy
by: Herrmann, Pascal, et al.
Published: (2026)
by: Herrmann, Pascal, et al.
Published: (2026)
A Cross Branch Fusion-Based Contrastive Learning Framework for Point Cloud Self-supervised Learning
by: Wu, Chengzhi, et al.
Published: (2025)
by: Wu, Chengzhi, et al.
Published: (2025)
Narrative Aligned Long Form Video Question Answering
by: Jain, Rahul, et al.
Published: (2026)
by: Jain, Rahul, et al.
Published: (2026)
Measuring the Effect of Background on Classification and Feature Importance in Deep Learning for AV Perception
by: Sielemann, Anne, et al.
Published: (2025)
by: Sielemann, Anne, et al.
Published: (2025)
Privacy-Preserving Semantic Segmentation from Ultra-Low-Resolution RGB Inputs
by: Huang, Xuying, et al.
Published: (2025)
by: Huang, Xuying, et al.
Published: (2025)
MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation
by: Wasim, Syed Talal, et al.
Published: (2025)
by: Wasim, Syed Talal, et al.
Published: (2025)
Rethinking temporal self-similarity for repetitive action counting
by: Luo, Yanan, et al.
Published: (2024)
by: Luo, Yanan, et al.
Published: (2024)
Towards Generalizing Temporal Action Segmentation to Unseen Views
by: Bahrami, Emad, et al.
Published: (2025)
by: Bahrami, Emad, et al.
Published: (2025)
Massively Multi-Person 3D Human Motion Forecasting with Scene Context
by: Mueller, Felix B, et al.
Published: (2024)
by: Mueller, Felix B, et al.
Published: (2024)
Fréchet Wavelet Distance: A Domain-Agnostic Metric for Image Generation
by: Veeramacheneni, Lokesh, et al.
Published: (2023)
by: Veeramacheneni, Lokesh, et al.
Published: (2023)
RiverMamba: A State Space Model for Global River Discharge and Flood Forecasting
by: Eddin, Mohamad Hakam Shams, et al.
Published: (2025)
by: Eddin, Mohamad Hakam Shams, et al.
Published: (2025)
GroupMamba: Efficient Group-Based Visual State Space Model
by: Shaker, Abdelrahman, et al.
Published: (2024)
by: Shaker, Abdelrahman, et al.
Published: (2024)
Similar Items
-
QueryMamba: A Mamba-Based Encoder-Decoder Architecture with a Statistical Verb-Noun Interaction Module for Video Action Forecasting @ Ego4D Long-Term Action Anticipation Challenge 2024
by: Zhong, Zeyun, et al.
Published: (2024) -
A Survey on Deep Learning Techniques for Action Anticipation
by: Zhong, Zeyun, et al.
Published: (2023) -
Multi-modal Video Representation Alignment for Robust Self-supervised Driver Distraction Detection
by: Lerch, David J., et al.
Published: (2026) -
Video Panels for Long Video Understanding
by: Doorenbos, Lars, et al.
Published: (2025) -
CamC2V: Context-aware Controllable Video Generation
by: Denninger, Luis, et al.
Published: (2025)