Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Dali, Zhang, Yunyao, Yu, Junqing, Chen, Yi-Ping Phoebe, Xu, Chen, Song, Zikai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HyperFusion: Hierarchical Multimodal Ensemble Learning for Social Media Popularity Prediction
by: Ye, Liliang, et al.
Published: (2025)
by: Ye, Liliang, et al.
Published: (2025)
M3TR: Temporal Retrieval Enhanced Multi-Modal Micro-video Popularity Prediction
by: Lu, Jiacheng, et al.
Published: (2024)
by: Lu, Jiacheng, et al.
Published: (2024)
MVP: Winning Solution to SMP Challenge 2025 Video Track
by: Ye, Liliang, et al.
Published: (2025)
by: Ye, Liliang, et al.
Published: (2025)
FairStream: Fair Multimedia Streaming Benchmark for Reinforcement Learning Agents
by: Weil, Jannis, et al.
Published: (2024)
by: Weil, Jannis, et al.
Published: (2024)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
by: Meng, Jiahao, et al.
Published: (2025)
by: Meng, Jiahao, et al.
Published: (2025)
HotComment: A Benchmark for Evaluating Popularity of Online Comments
by: Wu, Yafeng, et al.
Published: (2026)
by: Wu, Yafeng, et al.
Published: (2026)
Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation
by: Huang, Zikai, et al.
Published: (2025)
by: Huang, Zikai, et al.
Published: (2025)
Self-supervised Spatio-Temporal Graph Mask-Passing Attention Network for Perceptual Importance Prediction of Multi-point Tactility
by: He, Dazhong, et al.
Published: (2024)
by: He, Dazhong, et al.
Published: (2024)
Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos
by: Tian, Chong, et al.
Published: (2026)
by: Tian, Chong, et al.
Published: (2026)
Semantic-Aware Logical Reasoning via a Semiotic Framework
by: Zhang, Yunyao, et al.
Published: (2025)
by: Zhang, Yunyao, et al.
Published: (2025)
Generative Video Compression: Towards 0.01% Compression Rate for Video Transmission
by: Chen, Xiangyu, et al.
Published: (2025)
by: Chen, Xiangyu, et al.
Published: (2025)
Multimodal Emotion Recognition by Fusing Video Semantic in MOOC Learning Scenarios
by: Zhang, Yuan, et al.
Published: (2024)
by: Zhang, Yuan, et al.
Published: (2024)
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
by: Li, Haitian, et al.
Published: (2026)
by: Li, Haitian, et al.
Published: (2026)
Semantic-Guided Unsupervised Video Summarization
by: Liu, Haizhou, et al.
Published: (2026)
by: Liu, Haizhou, et al.
Published: (2026)
Conquering High Packet-Loss Erasure: MoE Swin Transformer-Based Video Semantic Communication
by: Teng, Lei, et al.
Published: (2025)
by: Teng, Lei, et al.
Published: (2025)
Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
by: Lu, Jiacheng, et al.
Published: (2025)
by: Lu, Jiacheng, et al.
Published: (2025)
Will It Go Viral? Grounding Micro-Video Popularity Prediction on the Open Web
by: Heo, Ryang, et al.
Published: (2026)
by: Heo, Ryang, et al.
Published: (2026)
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
by: Li, Maomao, et al.
Published: (2026)
by: Li, Maomao, et al.
Published: (2026)
AxiomVision: Accuracy-Guaranteed Adaptive Visual Model Selection for Perspective-Aware Video Analytics
by: Dai, Xiangxiang, et al.
Published: (2024)
by: Dai, Xiangxiang, et al.
Published: (2024)
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
by: Li, Chengzhi, et al.
Published: (2025)
by: Li, Chengzhi, et al.
Published: (2025)
OmniTrend: Content-Context Modeling for Scalable Social Popularity Prediction
by: Ye, Liliang, et al.
Published: (2026)
by: Ye, Liliang, et al.
Published: (2026)
Towards Open-Vocabulary Video Semantic Segmentation
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
MM-HSD: Multi-Modal Hate Speech Detection in Videos
by: Céspedes-Sarrias, Berta, et al.
Published: (2025)
by: Céspedes-Sarrias, Berta, et al.
Published: (2025)
Do Joint Audio-Video Generation Models Understand Physics?
by: Cui, Zijun, et al.
Published: (2026)
by: Cui, Zijun, et al.
Published: (2026)
CARAT: Contrastive Feature Reconstruction and Aggregation for Multi-Modal Multi-Label Emotion Recognition
by: Peng, Cheng, et al.
Published: (2023)
by: Peng, Cheng, et al.
Published: (2023)
Predicting Outcomes in Video Games with Long Short Term Memory Networks
by: Chulajata, Kittimate, et al.
Published: (2024)
by: Chulajata, Kittimate, et al.
Published: (2024)
Co-Director: Agentic Generative Video Storytelling
by: Song, Yale, et al.
Published: (2026)
by: Song, Yale, et al.
Published: (2026)
Wireless Video Semantic Communication with Decoupled Diffusion Multi-frame Compensation
by: Xie, Bingyan, et al.
Published: (2025)
by: Xie, Bingyan, et al.
Published: (2025)
Real-Time Mobile Video Analytics for Pre-arrival Emergency Medical Services
by: Jin, Liuyi, et al.
Published: (2025)
by: Jin, Liuyi, et al.
Published: (2025)
Attention of a Kiss: Exploring Attention Maps in Video Diffusion for XAIxArts
by: Cole, Adam, et al.
Published: (2025)
by: Cole, Adam, et al.
Published: (2025)
QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models
by: Lin, Zixing, et al.
Published: (2026)
by: Lin, Zixing, et al.
Published: (2026)
LL-GABR: Energy Efficient Live Video Streaming Using Reinforcement Learning
by: Raman, Adithya, et al.
Published: (2024)
by: Raman, Adithya, et al.
Published: (2024)
A Survey on Multimodal Benchmarks: In the Era of Large AI Models
by: Li, Lin, et al.
Published: (2024)
by: Li, Lin, et al.
Published: (2024)
Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs
by: Chen, Jianhao, et al.
Published: (2026)
by: Chen, Jianhao, et al.
Published: (2026)
Seeing World Dynamics in a Nutshell
by: Shen, Qiuhong, et al.
Published: (2025)
by: Shen, Qiuhong, et al.
Published: (2025)
ELF: A Family of Encoder-Free ECG-Language Models
by: Han, William, et al.
Published: (2026)
by: Han, William, et al.
Published: (2026)
End-to-End Learning-based Video Streaming Enhancement Pipeline: A Generative AI Approach
by: Artioli, Emanuele, et al.
Published: (2025)
by: Artioli, Emanuele, et al.
Published: (2025)
Plasticity-Aware Mixture of Experts for Learning Under QoE Shifts in Adaptive Video Streaming
by: He, Zhiqiang, et al.
Published: (2025)
by: He, Zhiqiang, et al.
Published: (2025)
MindFuse: Towards GenAI Explainability in Marketing Strategy Co-Creation
by: Farseev, Aleksandr, et al.
Published: (2025)
by: Farseev, Aleksandr, et al.
Published: (2025)
Solving Copyright Infringement on Short Video Platforms: Novel Datasets and an Audio Restoration Deep Learning Pipeline
by: Oh, Minwoo, et al.
Published: (2025)
by: Oh, Minwoo, et al.
Published: (2025)
Similar Items
-
HyperFusion: Hierarchical Multimodal Ensemble Learning for Social Media Popularity Prediction
by: Ye, Liliang, et al.
Published: (2025) -
M3TR: Temporal Retrieval Enhanced Multi-Modal Micro-video Popularity Prediction
by: Lu, Jiacheng, et al.
Published: (2024) -
MVP: Winning Solution to SMP Challenge 2025 Video Track
by: Ye, Liliang, et al.
Published: (2025) -
FairStream: Fair Multimedia Streaming Benchmark for Reinforcement Learning Agents
by: Weil, Jannis, et al.
Published: (2024) -
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
by: Meng, Jiahao, et al.
Published: (2025)