Automating Video Thumbnails Selection and Generation with Multimodal and Multistage Analysis
Fuente:
arXiv
Saved in:
| Main Author: | Fantini, Elia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TALL: Thumbnail Layout for Deepfake Video Detection
by: Xu, Yuting, et al.
Published: (2023)
by: Xu, Yuting, et al.
Published: (2023)
Detecting Cultural Differences in News Video Thumbnails via Computational Aesthetics
by: Limpijankit, Marvin, et al.
Published: (2025)
by: Limpijankit, Marvin, et al.
Published: (2025)
COP-GEN-Beta: Unified Generative Modelling of COPernicus Imagery Thumbnails
by: Espinosa, Miguel, et al.
Published: (2025)
by: Espinosa, Miguel, et al.
Published: (2025)
Learning Spatiotemporal Inconsistency via Thumbnail Layout for Face Deepfake Detection
by: Xu, Yuting, et al.
Published: (2024)
by: Xu, Yuting, et al.
Published: (2024)
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
by: Qu, Tingyu, et al.
Published: (2024)
by: Qu, Tingyu, et al.
Published: (2024)
Assessing News Thumbnail Representativeness: Counterfactual text can enhance the cross-modal matching ability
by: Yoon, Yejun, et al.
Published: (2024)
by: Yoon, Yejun, et al.
Published: (2024)
PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection
by: Wang, Peiyao, et al.
Published: (2025)
by: Wang, Peiyao, et al.
Published: (2025)
Can Impressions of Music be Extracted from Thumbnail Images?
by: Harada, Takashi, et al.
Published: (2025)
by: Harada, Takashi, et al.
Published: (2025)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
by: Peruzzo, Elia, et al.
Published: (2025)
by: Peruzzo, Elia, et al.
Published: (2025)
Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories
by: Jain, Chayan, et al.
Published: (2025)
by: Jain, Chayan, et al.
Published: (2025)
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
by: Wei, Jiangchuan, et al.
Published: (2025)
by: Wei, Jiangchuan, et al.
Published: (2025)
MMSummary: Multimodal Summary Generation for Fetal Ultrasound Video
by: Guo, Xiaoqing, et al.
Published: (2024)
by: Guo, Xiaoqing, et al.
Published: (2024)
PresentAgent: Multimodal Agent for Presentation Video Generation
by: Shi, Jingwei, et al.
Published: (2025)
by: Shi, Jingwei, et al.
Published: (2025)
MMPhysVideo: Scaling Physical Plausibility in Video Generation via Joint Multimodal Modeling
by: Lin, Shubo, et al.
Published: (2026)
by: Lin, Shubo, et al.
Published: (2026)
Advancing Automated Deception Detection: A Multimodal Approach to Feature Extraction and Analysis
by: Bahaa, Mohamed, et al.
Published: (2024)
by: Bahaa, Mohamed, et al.
Published: (2024)
LuciBot: Automated Robot Policy Learning from Generated Videos
by: Qiu, Xiaowen, et al.
Published: (2025)
by: Qiu, Xiaowen, et al.
Published: (2025)
Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions
by: Zhang, David Junhao, et al.
Published: (2024)
by: Zhang, David Junhao, et al.
Published: (2024)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
by: Tong, Jingqi, et al.
Published: (2025)
by: Tong, Jingqi, et al.
Published: (2025)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
by: Qiu, Zongyang, et al.
Published: (2025)
by: Qiu, Zongyang, et al.
Published: (2025)
When and How to Cut Classical Concerts? A Multimodal Automated Video Editing Approach
by: Gonzálbez-Biosca, Daniel, et al.
Published: (2025)
by: Gonzálbez-Biosca, Daniel, et al.
Published: (2025)
MSACT: Multistage Spatial Alignment for Stable Low-Latency Fine Manipulation
by: Cai, Xianbo, et al.
Published: (2026)
by: Cai, Xianbo, et al.
Published: (2026)
Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing
by: Liu, Jialun, et al.
Published: (2026)
by: Liu, Jialun, et al.
Published: (2026)
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation
by: Huang, Xiaohu, et al.
Published: (2025)
by: Huang, Xiaohu, et al.
Published: (2025)
LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
by: Gao, Jianxiong, et al.
Published: (2025)
by: Gao, Jianxiong, et al.
Published: (2025)
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning
by: Liu, Xiaohong, et al.
Published: (2025)
by: Liu, Xiaohong, et al.
Published: (2025)
Multimodal Engagement Analysis from Facial Videos in the Classroom
by: Sümer, Ömer, et al.
Published: (2021)
by: Sümer, Ömer, et al.
Published: (2021)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
by: Luo, Ziyang, et al.
Published: (2024)
by: Luo, Ziyang, et al.
Published: (2024)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
by: Wang, Shaoguang, et al.
Published: (2026)
by: Wang, Shaoguang, et al.
Published: (2026)
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
by: Gao, Hongcheng, et al.
Published: (2025)
by: Gao, Hongcheng, et al.
Published: (2025)
CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
Portrait Video Editing Empowered by Multimodal Generative Priors
by: Gao, Xuan, et al.
Published: (2024)
by: Gao, Xuan, et al.
Published: (2024)
A Multistage Extraction Pipeline for Long Scanned Financial Documents: An Empirical Study in Industrial KYC Workflows
by: Han, Yuxuan, et al.
Published: (2026)
by: Han, Yuxuan, et al.
Published: (2026)
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
by: Zhou, Donghao, et al.
Published: (2026)
by: Zhou, Donghao, et al.
Published: (2026)
MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection
by: Hu, Mengxue, et al.
Published: (2025)
by: Hu, Mengxue, et al.
Published: (2025)
Multi-Scale Local Speculative Decoding for Image Generation
by: Peruzzo, Elia, et al.
Published: (2026)
by: Peruzzo, Elia, et al.
Published: (2026)
Maintaining User Trust Through Multistage Uncertainty Aware Inference
by: Agrawal, Chandan, et al.
Published: (2023)
by: Agrawal, Chandan, et al.
Published: (2023)
Automated Construction of Time-Space Diagrams for Traffic Analysis Using Street-View Video Sequence
by: Rastogi, Tanay, et al.
Published: (2023)
by: Rastogi, Tanay, et al.
Published: (2023)
Multimodal Fake News Video Explanation: Dataset, Analysis and Evaluation
by: Chen, Lizhi, et al.
Published: (2025)
by: Chen, Lizhi, et al.
Published: (2025)
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding
by: Cheng, Tongtong, et al.
Published: (2025)
by: Cheng, Tongtong, et al.
Published: (2025)
Similar Items
-
TALL: Thumbnail Layout for Deepfake Video Detection
by: Xu, Yuting, et al.
Published: (2023) -
Detecting Cultural Differences in News Video Thumbnails via Computational Aesthetics
by: Limpijankit, Marvin, et al.
Published: (2025) -
COP-GEN-Beta: Unified Generative Modelling of COPernicus Imagery Thumbnails
by: Espinosa, Miguel, et al.
Published: (2025) -
Learning Spatiotemporal Inconsistency via Thumbnail Layout for Face Deepfake Detection
by: Xu, Yuting, et al.
Published: (2024) -
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
by: Qu, Tingyu, et al.
Published: (2024)