VidCLearn: A Continual Learning Approach for Text-to-Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zanchetta, Luca, Papa, Lorenzo, Maiano, Luca, Amerini, Irene |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DepthFake: a depth-based strategy for detecting Deepfake videos
by: Maiano, Luca, et al.
Published: (2022)
by: Maiano, Luca, et al.
Published: (2022)
Continuous fake media detection: adapting deepfake detectors to new generative techniques
by: Tassone, Francesco, et al.
Published: (2024)
by: Tassone, Francesco, et al.
Published: (2024)
Z-SASLM: Zero-Shot Style-Aligned SLI Blending Latent Manipulation
by: Borgi, Alessio, et al.
Published: (2025)
by: Borgi, Alessio, et al.
Published: (2025)
Enhancing Abnormality Identification: Robust Out-of-Distribution Strategies for Deepfake Detection
by: Maiano, Luca, et al.
Published: (2025)
by: Maiano, Luca, et al.
Published: (2025)
Enhancing Ground-to-Aerial Image Matching for Visual Misinformation Detection Using Semantic Segmentation
by: Mule, Emanuele, et al.
Published: (2025)
by: Mule, Emanuele, et al.
Published: (2025)
A Semantic Segmentation-guided Approach for Ground-to-Aerial Image Matching
by: Pro, Francesco, et al.
Published: (2024)
by: Pro, Francesco, et al.
Published: (2024)
Learning from Unlabelled Data with Transformers: Domain Adaptation for Semantic Segmentation of High Resolution Aerial Images
by: Dionelis, Nikolaos, et al.
Published: (2024)
by: Dionelis, Nikolaos, et al.
Published: (2024)
A survey on efficient vision transformers: algorithms, techniques, and performance benchmarking
by: Papa, Lorenzo, et al.
Published: (2023)
by: Papa, Lorenzo, et al.
Published: (2023)
On the impact of key design aspects in simulated Hybrid Quantum Neural Networks for Earth Observation
by: Papa, Lorenzo, et al.
Published: (2024)
by: Papa, Lorenzo, et al.
Published: (2024)
Shedding Light on Depth: Explainability Assessment in Monocular Depth Estimation
by: Cirillo, Lorenzo, et al.
Published: (2025)
by: Cirillo, Lorenzo, et al.
Published: (2025)
Diffusion Models for Earth Observation Use-cases: from cloud removal to urban change detection
by: Sanguigni, Fulvio, et al.
Published: (2023)
by: Sanguigni, Fulvio, et al.
Published: (2023)
D4D: An RGBD diffusion model to boost monocular depth estimation
by: Papa, L., et al.
Published: (2024)
by: Papa, L., et al.
Published: (2024)
METER: a mobile vision transformer architecture for monocular depth estimation
by: Papa, L., et al.
Published: (2024)
by: Papa, L., et al.
Published: (2024)
LADLE-MM: Limited Annotation based Detector with Learned Ensembles for Multimodal Misinformation
by: Cardullo, Daniele, et al.
Published: (2025)
by: Cardullo, Daniele, et al.
Published: (2025)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
by: Yang, Zhoufaran, et al.
Published: (2025)
by: Yang, Zhoufaran, et al.
Published: (2025)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
by: Wang, Yi, et al.
Published: (2023)
by: Wang, Yi, et al.
Published: (2023)
TextVidBench: A Benchmark for Long Video Scene Text Understanding
by: Zhong, Yangyang, et al.
Published: (2025)
by: Zhong, Yangyang, et al.
Published: (2025)
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
by: Feng, Weixi, et al.
Published: (2025)
by: Feng, Weixi, et al.
Published: (2025)
STLight: a Fully Convolutional Approach for Efficient Predictive Learning by Spatio-Temporal joint Processing
by: Alfarano, Andrea, et al.
Published: (2024)
by: Alfarano, Andrea, et al.
Published: (2024)
OmniVid: A Generative Framework for Universal Video Understanding
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
by: Liu, Yexin, et al.
Published: (2025)
by: Liu, Yexin, et al.
Published: (2025)
VidEdit: Zero-Shot and Spatially Aware Text-Driven Video Editing
by: Couairon, Paul, et al.
Published: (2023)
by: Couairon, Paul, et al.
Published: (2023)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
by: Tzachor, Issar, et al.
Published: (2026)
by: Tzachor, Issar, et al.
Published: (2026)
Can Text-to-Video Generation help Video-Language Alignment?
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data
by: Jin, Wonjoon, et al.
Published: (2026)
by: Jin, Wonjoon, et al.
Published: (2026)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
by: Qiu, Zongyang, et al.
Published: (2025)
by: Qiu, Zongyang, et al.
Published: (2025)
BachVid: Training-Free Video Generation with Consistent Background and Character
by: Yan, Han, et al.
Published: (2025)
by: Yan, Han, et al.
Published: (2025)
VidGen-1M: A Large-Scale Dataset for Text-to-video Generation
by: Tan, Zhiyu, et al.
Published: (2024)
by: Tan, Zhiyu, et al.
Published: (2024)
HarmoVid: Relightful Video Portrait Harmonization
by: Choi, Jun Myeong, et al.
Published: (2026)
by: Choi, Jun Myeong, et al.
Published: (2026)
VidLeaks: Membership Inference Attacks Against Text-to-Video Models
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
GenVidBench: A 6-Million Benchmark for AI-Generated Video Detection
by: Ni, Zhenliang, et al.
Published: (2025)
by: Ni, Zhenliang, et al.
Published: (2025)
UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
by: Chen, Lan, et al.
Published: (2025)
by: Chen, Lan, et al.
Published: (2025)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
by: Xu, Yicheng, et al.
Published: (2025)
by: Xu, Yicheng, et al.
Published: (2025)
UniVid: The Open-Source Unified Video Model
by: Luo, Jiabin, et al.
Published: (2025)
by: Luo, Jiabin, et al.
Published: (2025)
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
by: Nan, Kepan, et al.
Published: (2024)
by: Nan, Kepan, et al.
Published: (2024)
Vid-Morp: Video Moment Retrieval Pretraining from Unlabeled Videos in the Wild
by: Bao, Peijun, et al.
Published: (2024)
by: Bao, Peijun, et al.
Published: (2024)
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
Similar Items
-
DepthFake: a depth-based strategy for detecting Deepfake videos
by: Maiano, Luca, et al.
Published: (2022) -
Continuous fake media detection: adapting deepfake detectors to new generative techniques
by: Tassone, Francesco, et al.
Published: (2024) -
Z-SASLM: Zero-Shot Style-Aligned SLI Blending Latent Manipulation
by: Borgi, Alessio, et al.
Published: (2025) -
Enhancing Abnormality Identification: Robust Out-of-Distribution Strategies for Deepfake Detection
by: Maiano, Luca, et al.
Published: (2025) -
Enhancing Ground-to-Aerial Image Matching for Visual Misinformation Detection Using Semantic Segmentation
by: Mule, Emanuele, et al.
Published: (2025)