Self-supervised Multi-actor Social Activity Understanding in Streaming Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Trehan, Shubham, Aakur, Sathyanarayanan N. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Discovering Novel Actions from Open World Egocentric Videos with Object-Grounded Visual Commonsense Reasoning
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2023)
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2023)
ALGO: Object-Grounded Visual Commonsense Reasoning for Open-World Egocentric Action Recognition
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2024)
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2024)
FSP-DETR: Few-Shot Prototypical Parasitic Ova Detection
von: Trehan, Shubham, et al.
Veröffentlicht: (2025)
von: Trehan, Shubham, et al.
Veröffentlicht: (2025)
A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2025)
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2025)
ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2025)
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2025)
CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding
von: Chen, Zhou, et al.
Veröffentlicht: (2025)
von: Chen, Zhou, et al.
Veröffentlicht: (2025)
Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection
von: Vellamcheti, Shanmukha, et al.
Veröffentlicht: (2025)
von: Vellamcheti, Shanmukha, et al.
Veröffentlicht: (2025)
Generalized Event Partonomy Inference with Structured Hierarchical Predictive Learning
von: Chen, Zhou, et al.
Veröffentlicht: (2025)
von: Chen, Zhou, et al.
Veröffentlicht: (2025)
EASE: Embodied Active Event Perception via Self-Supervised Energy Minimization
von: Chen, Zhou, et al.
Veröffentlicht: (2025)
von: Chen, Zhou, et al.
Veröffentlicht: (2025)
STaTS: Structure-Aware Temporal Sequence Summarization via Statistical Window Merging
von: Bhowmick, Disharee, et al.
Veröffentlicht: (2025)
von: Bhowmick, Disharee, et al.
Veröffentlicht: (2025)
Capturing Temporal Components for Time Series Classification
von: Vavilthota, Venkata Ragavendra, et al.
Veröffentlicht: (2024)
von: Vavilthota, Venkata Ragavendra, et al.
Veröffentlicht: (2024)
CVT-Bench: Counterfactual Viewpoint Transformations Reveal Unstable Spatial Representations in Multimodal LLMs
von: Vellamcheti, Shanmukha, et al.
Veröffentlicht: (2026)
von: Vellamcheti, Shanmukha, et al.
Veröffentlicht: (2026)
CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
von: Wang, Yifei, et al.
Veröffentlicht: (2025)
von: Wang, Yifei, et al.
Veröffentlicht: (2025)
StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
SoGAR: Self-supervised Spatiotemporal Attention-based Social Group Activity Recognition
von: Chappa, Naga VS Raviteja, et al.
Veröffentlicht: (2023)
von: Chappa, Naga VS Raviteja, et al.
Veröffentlicht: (2023)
SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
von: Yang, Zhenyu, et al.
Veröffentlicht: (2025)
von: Yang, Zhenyu, et al.
Veröffentlicht: (2025)
Multi-modal Video Representation Alignment for Robust Self-supervised Driver Distraction Detection
von: Lerch, David J., et al.
Veröffentlicht: (2026)
von: Lerch, David J., et al.
Veröffentlicht: (2026)
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
von: Sun, Peiwen, et al.
Veröffentlicht: (2026)
von: Sun, Peiwen, et al.
Veröffentlicht: (2026)
Streaming Long Video Understanding with Large Language Models
von: Qian, Rui, et al.
Veröffentlicht: (2024)
von: Qian, Rui, et al.
Veröffentlicht: (2024)
An Efficient Streaming Video Understanding Framework with Agentic Control
von: Liu, Jinming, et al.
Veröffentlicht: (2026)
von: Liu, Jinming, et al.
Veröffentlicht: (2026)
FakeOut: Leveraging Out-of-domain Self-supervision for Multi-modal Video Deepfake Detection
von: Knafo, Gil, et al.
Veröffentlicht: (2022)
von: Knafo, Gil, et al.
Veröffentlicht: (2022)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
von: Chatterjee, Dibyadip, et al.
Veröffentlicht: (2025)
von: Chatterjee, Dibyadip, et al.
Veröffentlicht: (2025)
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
von: Pang, Zhanzhong, et al.
Veröffentlicht: (2026)
von: Pang, Zhanzhong, et al.
Veröffentlicht: (2026)
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
von: Xie, Ming, et al.
Veröffentlicht: (2026)
von: Xie, Ming, et al.
Veröffentlicht: (2026)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
von: Luo, Yawen, et al.
Veröffentlicht: (2026)
von: Luo, Yawen, et al.
Veröffentlicht: (2026)
Collaboratively Self-supervised Video Representation Learning for Action Recognition
von: Zhang, Jie, et al.
Veröffentlicht: (2024)
von: Zhang, Jie, et al.
Veröffentlicht: (2024)
A Self-supervised Motion Representation for Portrait Video Generation
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2025)
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
von: Xiong, Haomiao, et al.
Veröffentlicht: (2025)
von: Xiong, Haomiao, et al.
Veröffentlicht: (2025)
Video-Zero: Self-Evolution Video Understanding
von: Zhang, Ruixu, et al.
Veröffentlicht: (2026)
von: Zhang, Ruixu, et al.
Veröffentlicht: (2026)
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
von: Chen, Xueyi, et al.
Veröffentlicht: (2025)
von: Chen, Xueyi, et al.
Veröffentlicht: (2025)
Exploring Semantic Masked Autoencoder for Self-supervised Point Cloud Understanding
von: Zha, Yixin, et al.
Veröffentlicht: (2025)
von: Zha, Yixin, et al.
Veröffentlicht: (2025)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2025)
Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2025)
AURA: Always-On Understanding and Real-Time Assistance via Video Streams
von: Lu, Xudong, et al.
Veröffentlicht: (2026)
von: Lu, Xudong, et al.
Veröffentlicht: (2026)
Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
von: Kang, Hyolim, et al.
Veröffentlicht: (2025)
von: Kang, Hyolim, et al.
Veröffentlicht: (2025)
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
Self-supervised Video Object Segmentation with Distillation Learning of Deformable Attention
von: Truong, Quang-Trung, et al.
Veröffentlicht: (2024)
von: Truong, Quang-Trung, et al.
Veröffentlicht: (2024)
Large-scale Self-supervised Video Foundation Model for Intelligent Surgery
von: Yang, Shu, et al.
Veröffentlicht: (2025)
von: Yang, Shu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Discovering Novel Actions from Open World Egocentric Videos with Object-Grounded Visual Commonsense Reasoning
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2023) -
ALGO: Object-Grounded Visual Commonsense Reasoning for Open-World Egocentric Action Recognition
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2024) -
FSP-DETR: Few-Shot Prototypical Parasitic Ova Detection
von: Trehan, Shubham, et al.
Veröffentlicht: (2025) -
A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2025) -
ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2025)