HOTVCOM: Generating Buzzworthy Comments for Videos
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Yuyan, Qian, Yiwen, Yan, Songzhou, Jia, Jiyuan, Li, Zhixu, Xiao, Yanghua, Li, Xiaobo, Yang, Ming, Guo, Qingpei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
XMeCap: Meme Caption Generation with Sub-Image Adaptability
di: Chen, Yuyan, et al.
Pubblicazione: (2024)
di: Chen, Yuyan, et al.
Pubblicazione: (2024)
VCEval: Rethinking What is a Good Educational Video and How to Automatically Evaluate It
di: Zhu, Xiaoxuan, et al.
Pubblicazione: (2024)
di: Zhu, Xiaoxuan, et al.
Pubblicazione: (2024)
SHE-Net: Syntax-Hierarchy-Enhanced Text-Video Retrieval
di: Yu, Xuzheng, et al.
Pubblicazione: (2024)
di: Yu, Xuzheng, et al.
Pubblicazione: (2024)
M2-RAAP: A Multi-Modal Recipe for Advancing Adaptation-based Pre-training towards Effective and Efficient Zero-shot Video-text Retrieval
di: Dong, Xingning, et al.
Pubblicazione: (2024)
di: Dong, Xingning, et al.
Pubblicazione: (2024)
Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning
di: Wang, Haoyu, et al.
Pubblicazione: (2026)
di: Wang, Haoyu, et al.
Pubblicazione: (2026)
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
di: Xuan, Shiyu, et al.
Pubblicazione: (2023)
di: Xuan, Shiyu, et al.
Pubblicazione: (2023)
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
di: Han, Yudong, et al.
Pubblicazione: (2024)
di: Han, Yudong, et al.
Pubblicazione: (2024)
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment
di: Ma, Ziping, et al.
Pubblicazione: (2024)
di: Ma, Ziping, et al.
Pubblicazione: (2024)
Video Virtual Try-on with Conditional Diffusion Transformer Inpainter
di: Zou, Cheng, et al.
Pubblicazione: (2025)
di: Zou, Cheng, et al.
Pubblicazione: (2025)
Adaptive Feature Fusion Neural Network for Glaucoma Segmentation on Unseen Fundus Images
di: Zhong, Jiyuan, et al.
Pubblicazione: (2024)
di: Zhong, Jiyuan, et al.
Pubblicazione: (2024)
SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories
di: Zhu, Muzhi, et al.
Pubblicazione: (2025)
di: Zhu, Muzhi, et al.
Pubblicazione: (2025)
Generative Video Compression with One-Dimensional Latent Representation
di: Zheng, Zihan, et al.
Pubblicazione: (2026)
di: Zheng, Zihan, et al.
Pubblicazione: (2026)
VideoPure: Diffusion-based Adversarial Purification for Video Recognition
di: Jiang, Kaixun, et al.
Pubblicazione: (2025)
di: Jiang, Kaixun, et al.
Pubblicazione: (2025)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
di: Zheng, Naishan, et al.
Pubblicazione: (2025)
di: Zheng, Naishan, et al.
Pubblicazione: (2025)
From Mapping to Composing: A Two-Stage Framework for Zero-shot Composed Image Retrieval
di: Wang, Yabing, et al.
Pubblicazione: (2025)
di: Wang, Yabing, et al.
Pubblicazione: (2025)
Generative Latent Video Compression
di: Guo, Zongyu, et al.
Pubblicazione: (2025)
di: Guo, Zongyu, et al.
Pubblicazione: (2025)
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
di: Zheng, Ruobing, et al.
Pubblicazione: (2026)
di: Zheng, Ruobing, et al.
Pubblicazione: (2026)
Omni-Video: Democratizing Unified Video Understanding and Generation
di: Tan, Zhiyu, et al.
Pubblicazione: (2025)
di: Tan, Zhiyu, et al.
Pubblicazione: (2025)
AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation
di: Liu, Xiao, et al.
Pubblicazione: (2025)
di: Liu, Xiao, et al.
Pubblicazione: (2025)
FlattenGPT: Depth Compression for Transformer with Layer Flattening
di: Xu, Ruihan, et al.
Pubblicazione: (2026)
di: Xu, Ruihan, et al.
Pubblicazione: (2026)
ControlNeXt: Powerful and Efficient Control for Image and Video Generation
di: Peng, Bohao, et al.
Pubblicazione: (2024)
di: Peng, Bohao, et al.
Pubblicazione: (2024)
TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking
di: Hu, Jiyuan, et al.
Pubblicazione: (2026)
di: Hu, Jiyuan, et al.
Pubblicazione: (2026)
Taming Camera-Controlled Video Generation with Verifiable Geometry Reward
di: Wang, Zhaoqing, et al.
Pubblicazione: (2025)
di: Wang, Zhaoqing, et al.
Pubblicazione: (2025)
LOLGORITHM: Funny Comment Generation Agent For Short Videos
di: Ouyang, Xuan, et al.
Pubblicazione: (2026)
di: Ouyang, Xuan, et al.
Pubblicazione: (2026)
LPA3D: 3D Room-Level Scene Generation from In-the-Wild Images
di: Yang, Ming-Jia, et al.
Pubblicazione: (2025)
di: Yang, Ming-Jia, et al.
Pubblicazione: (2025)
Motion Control for Enhanced Complex Action Video Generation
di: Zhou, Qiang, et al.
Pubblicazione: (2024)
di: Zhou, Qiang, et al.
Pubblicazione: (2024)
DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes
di: Liu, Jinxiu, et al.
Pubblicazione: (2024)
di: Liu, Jinxiu, et al.
Pubblicazione: (2024)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
di: Yang, Hao, et al.
Pubblicazione: (2026)
di: Yang, Hao, et al.
Pubblicazione: (2026)
Efficient Autoregressive Video Diffusion with Dummy Head
di: Guo, Hang, et al.
Pubblicazione: (2026)
di: Guo, Hang, et al.
Pubblicazione: (2026)
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning
di: Jin, Peng, et al.
Pubblicazione: (2024)
di: Jin, Peng, et al.
Pubblicazione: (2024)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
di: Wang, Yabing, et al.
Pubblicazione: (2024)
di: Wang, Yabing, et al.
Pubblicazione: (2024)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
di: Jiang, Chen, et al.
Pubblicazione: (2023)
di: Jiang, Chen, et al.
Pubblicazione: (2023)
Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning
di: Yan, Zhiyuan, et al.
Pubblicazione: (2024)
di: Yan, Zhiyuan, et al.
Pubblicazione: (2024)
Helios: Real Real-Time Long Video Generation Model
di: Yuan, Shenghai, et al.
Pubblicazione: (2026)
di: Yuan, Shenghai, et al.
Pubblicazione: (2026)
Video-Bench: Human-Aligned Video Generation Benchmark
di: Han, Hui, et al.
Pubblicazione: (2025)
di: Han, Hui, et al.
Pubblicazione: (2025)
Plenoptic Video Generation
di: Fu, Xiao, et al.
Pubblicazione: (2026)
di: Fu, Xiao, et al.
Pubblicazione: (2026)
Sora Generates Videos with Stunning Geometrical Consistency
di: Li, Xuanyi, et al.
Pubblicazione: (2024)
di: Li, Xuanyi, et al.
Pubblicazione: (2024)
SNP-S3: Shared Network Pre-training and Significant Semantic Strengthening for Various Video-Text Tasks
di: Dong, Xingning, et al.
Pubblicazione: (2024)
di: Dong, Xingning, et al.
Pubblicazione: (2024)
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
di: Wang, Zhao, et al.
Pubblicazione: (2024)
di: Wang, Zhao, et al.
Pubblicazione: (2024)
Challenger: Affordable Adversarial Driving Video Generation
di: Xu, Zhiyuan, et al.
Pubblicazione: (2025)
di: Xu, Zhiyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
XMeCap: Meme Caption Generation with Sub-Image Adaptability
di: Chen, Yuyan, et al.
Pubblicazione: (2024) -
VCEval: Rethinking What is a Good Educational Video and How to Automatically Evaluate It
di: Zhu, Xiaoxuan, et al.
Pubblicazione: (2024) -
SHE-Net: Syntax-Hierarchy-Enhanced Text-Video Retrieval
di: Yu, Xuzheng, et al.
Pubblicazione: (2024) -
M2-RAAP: A Multi-Modal Recipe for Advancing Adaptation-based Pre-training towards Effective and Efficient Zero-shot Video-text Retrieval
di: Dong, Xingning, et al.
Pubblicazione: (2024) -
Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning
di: Wang, Haoyu, et al.
Pubblicazione: (2026)