Gespeichert in:
| Hauptverfasser: | Xu, Mingchen, Wu, Jing, Lai, Yukun, Ji, Ze |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2407.07999 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MirrorSAM2: Segment Mirror in Videos with Depth Perception
von: Xu, Mingchen, et al.
Veröffentlicht: (2025)
von: Xu, Mingchen, et al.
Veröffentlicht: (2025)
VideoMem: Constructing, Analyzing, Predicting Short-term and Long-term Video Memorability
von: Cohendet, Romain, et al.
Veröffentlicht: (2018)
von: Cohendet, Romain, et al.
Veröffentlicht: (2018)
SALI: Short-term Alignment and Long-term Interaction Network for Colonoscopy Video Polyp Segmentation
von: Hu, Qiang, et al.
Veröffentlicht: (2024)
von: Hu, Qiang, et al.
Veröffentlicht: (2024)
Video World Models with Long-term Spatial Memory
von: Wu, Tong, et al.
Veröffentlicht: (2025)
von: Wu, Tong, et al.
Veröffentlicht: (2025)
State-space Decomposition Model for Video Prediction Considering Long-term Motion Trend
von: Cui, Fei, et al.
Veröffentlicht: (2024)
von: Cui, Fei, et al.
Veröffentlicht: (2024)
Mirror-Yolo: A Novel Attention Focus, Instance Segmentation and Mirror Detection Model
von: Li, Fengze, et al.
Veröffentlicht: (2022)
von: Li, Fengze, et al.
Veröffentlicht: (2022)
Hallucination Mitigation Prompts Long-term Video Understanding
von: Sun, Yiwei, et al.
Veröffentlicht: (2024)
von: Sun, Yiwei, et al.
Veröffentlicht: (2024)
Multi-Sentence Grounding for Long-term Instructional Video
von: Li, Zeqian, et al.
Veröffentlicht: (2023)
von: Li, Zeqian, et al.
Veröffentlicht: (2023)
Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity
von: Zhang, Huaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Huaxin, et al.
Veröffentlicht: (2024)
Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation
von: Liang, Tianming, et al.
Veröffentlicht: (2025)
von: Liang, Tianming, et al.
Veröffentlicht: (2025)
Multi-granularity Correspondence Learning from Long-term Noisy Videos
von: Lin, Yijie, et al.
Veröffentlicht: (2024)
von: Lin, Yijie, et al.
Veröffentlicht: (2024)
Interpretable Long-term Action Quality Assessment
von: Dong, Xu, et al.
Veröffentlicht: (2024)
von: Dong, Xu, et al.
Veröffentlicht: (2024)
Long-term Pre-training for Temporal Action Detection with Transformers
von: Kim, Jihwan, et al.
Veröffentlicht: (2024)
von: Kim, Jihwan, et al.
Veröffentlicht: (2024)
Exploring Recurrent Long-term Temporal Fusion for Multi-view 3D Perception
von: Han, Chunrui, et al.
Veröffentlicht: (2023)
von: Han, Chunrui, et al.
Veröffentlicht: (2023)
Lagrangian Motion Fields for Long-term Motion Generation
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation
von: Hong, Lingyi, et al.
Veröffentlicht: (2024)
von: Hong, Lingyi, et al.
Veröffentlicht: (2024)
CompEvent: Complex-valued Event-RGB Fusion for Low-light Video Enhancement and Deblurring
von: Zhong, Mingchen, et al.
Veröffentlicht: (2025)
von: Zhong, Mingchen, et al.
Veröffentlicht: (2025)
ESP: Extro-Spective Prediction for Long-term Behavior Reasoning in Emergency Scenarios
von: Wang, Dingrui, et al.
Veröffentlicht: (2024)
von: Wang, Dingrui, et al.
Veröffentlicht: (2024)
MirrorMamba: Towards Scalable and Robust Mirror Detection in Videos
von: Song, Rui, et al.
Veröffentlicht: (2025)
von: Song, Rui, et al.
Veröffentlicht: (2025)
VMID: A Multimodal Fusion LLM Framework for Detecting and Identifying Misinformation of Short Videos
von: Zhong, Weihao, et al.
Veröffentlicht: (2024)
von: Zhong, Weihao, et al.
Veröffentlicht: (2024)
FusionEdit: Semantic Fusion and Attention Modulation for Training-Free Image Editing
von: Lai, Yongwen, et al.
Veröffentlicht: (2026)
von: Lai, Yongwen, et al.
Veröffentlicht: (2026)
Investigating Long-term Training for Remote Sensing Object Detection
von: Park, JongHyun, et al.
Veröffentlicht: (2024)
von: Park, JongHyun, et al.
Veröffentlicht: (2024)
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs
von: Yamao, Sosuke, et al.
Veröffentlicht: (2024)
von: Yamao, Sosuke, et al.
Veröffentlicht: (2024)
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
von: Fei, Jiajun, et al.
Veröffentlicht: (2024)
von: Fei, Jiajun, et al.
Veröffentlicht: (2024)
Cross-Modal Fusion and Attention Mechanism for Weakly Supervised Video Anomaly Detection
von: Ghadiya, Ayush, et al.
Veröffentlicht: (2024)
von: Ghadiya, Ayush, et al.
Veröffentlicht: (2024)
The 1st Solution for MOSEv2 Challenge 2025: Long-term and Concept-aware Video Segmentation via SeC
von: Gao, Mingqi, et al.
Veröffentlicht: (2025)
von: Gao, Mingqi, et al.
Veröffentlicht: (2025)
Fast Camouflaged Object Detection via Edge-based Reversible Re-calibration Network
von: Ji, Ge-Peng, et al.
Veröffentlicht: (2021)
von: Ji, Ge-Peng, et al.
Veröffentlicht: (2021)
When SAM2 Meets Video Shadow and Mirror Detection
von: Jie, Leiping
Veröffentlicht: (2024)
von: Jie, Leiping
Veröffentlicht: (2024)
Task-driven Image Fusion with Learnable Fusion Loss
von: Bai, Haowen, et al.
Veröffentlicht: (2024)
von: Bai, Haowen, et al.
Veröffentlicht: (2024)
AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?
von: Zhao, Qi, et al.
Veröffentlicht: (2023)
von: Zhao, Qi, et al.
Veröffentlicht: (2023)
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
MambaVF: State Space Model for Efficient Video Fusion
von: Zhao, Zixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Zixiang, et al.
Veröffentlicht: (2026)
Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
von: Zhang, Jiangning, et al.
Veröffentlicht: (2025)
von: Zhang, Jiangning, et al.
Veröffentlicht: (2025)
MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving
von: Wang, Haiguang, et al.
Veröffentlicht: (2025)
von: Wang, Haiguang, et al.
Veröffentlicht: (2025)
SceneTracker: Long-term Scene Flow Estimation Network
von: Wang, Bo, et al.
Veröffentlicht: (2024)
von: Wang, Bo, et al.
Veröffentlicht: (2024)
Learning Long-term Motion Embeddings for Efficient Kinematics Generation
von: Stracke, Nick, et al.
Veröffentlicht: (2026)
von: Stracke, Nick, et al.
Veröffentlicht: (2026)
Online Long-term Point Tracking in the Foundation Model Era
von: Aydemir, Görkay
Veröffentlicht: (2025)
von: Aydemir, Görkay
Veröffentlicht: (2025)
WorldMem: Long-term Consistent World Simulation with Memory
von: Xiao, Zeqi, et al.
Veröffentlicht: (2025)
von: Xiao, Zeqi, et al.
Veröffentlicht: (2025)
Dyadic Mamba: Long-term Dyadic Human Motion Synthesis
von: Tanke, Julian, et al.
Veröffentlicht: (2025)
von: Tanke, Julian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MirrorSAM2: Segment Mirror in Videos with Depth Perception
von: Xu, Mingchen, et al.
Veröffentlicht: (2025) -
VideoMem: Constructing, Analyzing, Predicting Short-term and Long-term Video Memorability
von: Cohendet, Romain, et al.
Veröffentlicht: (2018) -
SALI: Short-term Alignment and Long-term Interaction Network for Colonoscopy Video Polyp Segmentation
von: Hu, Qiang, et al.
Veröffentlicht: (2024) -
Video World Models with Long-term Spatial Memory
von: Wu, Tong, et al.
Veröffentlicht: (2025) -
State-space Decomposition Model for Video Prediction Considering Long-term Motion Trend
von: Cui, Fei, et al.
Veröffentlicht: (2024)