Gespeichert in:
| Hauptverfasser: | ElAlami, M. E., Khater, S. M., Rehan, M. El. R. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2601.17022 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can LLMs Create Legally Relevant Summaries and Analyses of Videos?
von: Hoeben-Kuil, Lyra, et al.
Veröffentlicht: (2025)
von: Hoeben-Kuil, Lyra, et al.
Veröffentlicht: (2025)
Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators
von: Jo, Claire Wonjeong, et al.
Veröffentlicht: (2024)
von: Jo, Claire Wonjeong, et al.
Veröffentlicht: (2024)
ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided Transformer
von: Chen, Hongruixuan, et al.
Veröffentlicht: (2023)
von: Chen, Hongruixuan, et al.
Veröffentlicht: (2023)
KI-Bilder und die Widerständigkeit der Medienkonvergenz: Von primärer zu sekundärer Intermedialität?
von: Wilde, Lukas R. A.
Veröffentlicht: (2024)
von: Wilde, Lukas R. A.
Veröffentlicht: (2024)
TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention
von: Shi, Chuancheng, et al.
Veröffentlicht: (2026)
von: Shi, Chuancheng, et al.
Veröffentlicht: (2026)
Towards nation-wide analytical healthcare infrastructures: A privacy-preserving augmented knee rehabilitation case study
von: Bačić, Boris, et al.
Veröffentlicht: (2024)
von: Bačić, Boris, et al.
Veröffentlicht: (2024)
Robust Latent Representation Tuning for Image-text Classification
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
Advance Fake Video Detection via Vision Transformers
von: Battocchio, Joy, et al.
Veröffentlicht: (2025)
von: Battocchio, Joy, et al.
Veröffentlicht: (2025)
Moiré Video Authentication: A Physical Signature Against AI Video Generation
von: Qing, Yuan, et al.
Veröffentlicht: (2026)
von: Qing, Yuan, et al.
Veröffentlicht: (2026)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
von: Qi, Peng, et al.
Veröffentlicht: (2024)
von: Qi, Peng, et al.
Veröffentlicht: (2024)
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
von: Lin, Xiao, et al.
Veröffentlicht: (2025)
von: Lin, Xiao, et al.
Veröffentlicht: (2025)
A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming
von: Zhou, Pengyuan, et al.
Veröffentlicht: (2024)
von: Zhou, Pengyuan, et al.
Veröffentlicht: (2024)
LPM 1.0: Video-based Character Performance Model
von: Zeng, Ailing, et al.
Veröffentlicht: (2026)
von: Zeng, Ailing, et al.
Veröffentlicht: (2026)
AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production
von: Wang, Jiuniu, et al.
Veröffentlicht: (2024)
von: Wang, Jiuniu, et al.
Veröffentlicht: (2024)
Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value Optimization
von: Wang, Xingqi, et al.
Veröffentlicht: (2024)
von: Wang, Xingqi, et al.
Veröffentlicht: (2024)
Unmasking Illusions: Understanding Human Perception of Audiovisual Deepfakes
von: Hashmi, Ammarah, et al.
Veröffentlicht: (2024)
von: Hashmi, Ammarah, et al.
Veröffentlicht: (2024)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
AI-Driven Virtual Teacher for Enhanced Educational Efficiency: Leveraging Large Pretrain Models for Autonomous Error Analysis and Correction
von: Xu, Tianlong, et al.
Veröffentlicht: (2024)
von: Xu, Tianlong, et al.
Veröffentlicht: (2024)
Video Seal: Open and Efficient Video Watermarking
von: Fernandez, Pierre, et al.
Veröffentlicht: (2024)
von: Fernandez, Pierre, et al.
Veröffentlicht: (2024)
FedVideoMAE: Efficient Privacy-Preserving Federated Video Moderation
von: Tao, Ziyuan, et al.
Veröffentlicht: (2025)
von: Tao, Ziyuan, et al.
Veröffentlicht: (2025)
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
von: Gu, Jing, et al.
Veröffentlicht: (2024)
von: Gu, Jing, et al.
Veröffentlicht: (2024)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
von: Meng, Jiahao, et al.
Veröffentlicht: (2025)
von: Meng, Jiahao, et al.
Veröffentlicht: (2025)
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
von: Bian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Bian, Yuxuan, et al.
Veröffentlicht: (2025)
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
von: Yan, Xin, et al.
Veröffentlicht: (2024)
von: Yan, Xin, et al.
Veröffentlicht: (2024)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
von: Ku, Max, et al.
Veröffentlicht: (2024)
von: Ku, Max, et al.
Veröffentlicht: (2024)
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
Question-Answering Dense Video Events
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
Bernini: Latent Semantic Planning for Video Diffusion
von: Bernini Team, et al.
Veröffentlicht: (2026)
von: Bernini Team, et al.
Veröffentlicht: (2026)
Interactive Video Generation via Domain Adaptation
von: Rawal, Ishaan, et al.
Veröffentlicht: (2025)
von: Rawal, Ishaan, et al.
Veröffentlicht: (2025)
TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding
von: Ku, Max, et al.
Veröffentlicht: (2025)
von: Ku, Max, et al.
Veröffentlicht: (2025)
Image Conductor: Precision Control for Interactive Video Synthesis
von: Li, Yaowei, et al.
Veröffentlicht: (2024)
von: Li, Yaowei, et al.
Veröffentlicht: (2024)
Multimodal Chaptering for Long-Form TV Newscast Video
von: Guetari, Khalil, et al.
Veröffentlicht: (2024)
von: Guetari, Khalil, et al.
Veröffentlicht: (2024)
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
von: Barbakos, Spyros, et al.
Veröffentlicht: (2025)
von: Barbakos, Spyros, et al.
Veröffentlicht: (2025)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
Audio-visual Event Localization on Portrait Mode Short Videos
von: Liu, Wuyang, et al.
Veröffentlicht: (2025)
von: Liu, Wuyang, et al.
Veröffentlicht: (2025)
End-to-End Optimized Image Compression with the Frequency-Oriented Transform
von: Zhang, Yuefeng, et al.
Veröffentlicht: (2024)
von: Zhang, Yuefeng, et al.
Veröffentlicht: (2024)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
LoViF 2026 The First Challenge on Weather Removal in Videos
von: Qian, Chenghao, et al.
Veröffentlicht: (2026)
von: Qian, Chenghao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Can LLMs Create Legally Relevant Summaries and Analyses of Videos?
von: Hoeben-Kuil, Lyra, et al.
Veröffentlicht: (2025) -
Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators
von: Jo, Claire Wonjeong, et al.
Veröffentlicht: (2024) -
ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided Transformer
von: Chen, Hongruixuan, et al.
Veröffentlicht: (2023) -
KI-Bilder und die Widerständigkeit der Medienkonvergenz: Von primärer zu sekundärer Intermedialität?
von: Wilde, Lukas R. A.
Veröffentlicht: (2024) -
TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention
von: Shi, Chuancheng, et al.
Veröffentlicht: (2026)