Saved in:
| Main Authors: | Xue, Haoning, Zhang, Jingwen, Wang, Xiaohui, Kim, Diane Dagyong, Song, Yunya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.19995 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Engagement Prediction of Short Videos with Large Multimodal Models
by: Sun, Wei, et al.
Published: (2025)
by: Sun, Wei, et al.
Published: (2025)
Leveraging Segment Anything Model in Identifying Buildings within Refugee Camps (SAM4Refugee) from Satellite Imagery for Humanitarian Operations
by: Gao, Yunya
Published: (2024)
by: Gao, Yunya
Published: (2024)
Delving Deep into Engagement Prediction of Short Videos
by: Li, Dasong, et al.
Published: (2024)
by: Li, Dasong, et al.
Published: (2024)
VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results
by: Li, Dasong, et al.
Published: (2025)
by: Li, Dasong, et al.
Published: (2025)
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
by: Jin, Xiaojie, et al.
Published: (2023)
by: Jin, Xiaojie, et al.
Published: (2023)
Multimodal Engagement Analysis from Facial Videos in the Classroom
by: Sümer, Ömer, et al.
Published: (2021)
by: Sümer, Ömer, et al.
Published: (2021)
Reasoning-Enhanced Domain-Adaptive Pretraining of Multimodal Large Language Models for Short Video Content Governance
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
by: Luo, Ziyang, et al.
Published: (2024)
by: Luo, Ziyang, et al.
Published: (2024)
A Multimodal Dataset for Enhancing Industrial Task Monitoring and Engagement Prediction
by: Mehta, Naval Kishore, et al.
Published: (2025)
by: Mehta, Naval Kishore, et al.
Published: (2025)
Analyzing Participants' Engagement during Online Meetings Using Unsupervised Remote Photoplethysmography with Behavioral Features
by: Vedernikov, Alexander, et al.
Published: (2024)
by: Vedernikov, Alexander, et al.
Published: (2024)
Sens-VisualNews: A Benchmark Dataset for Sensational Image Detection
by: Goulas, Andreas, et al.
Published: (2026)
by: Goulas, Andreas, et al.
Published: (2026)
BYOCL: Build Your Own Consistent Latent with Hierarchical Representative Latent Clustering
by: Dai, Jiayue, et al.
Published: (2024)
by: Dai, Jiayue, et al.
Published: (2024)
ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks
by: Yang, Yan, et al.
Published: (2025)
by: Yang, Yan, et al.
Published: (2025)
Towards Universal Soccer Video Understanding
by: Rao, Jiayuan, et al.
Published: (2024)
by: Rao, Jiayuan, et al.
Published: (2024)
Fast Deep Predictive Coding Networks for Videos Feature Extraction without Labels
by: Xue, Wenqian, et al.
Published: (2024)
by: Xue, Wenqian, et al.
Published: (2024)
Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement
by: Yang, Zhengxian, et al.
Published: (2026)
by: Yang, Zhengxian, et al.
Published: (2026)
TCCT-Net: Two-Stream Network Architecture for Fast and Efficient Engagement Estimation via Behavioral Feature Signals
by: Vedernikov, Alexander, et al.
Published: (2024)
by: Vedernikov, Alexander, et al.
Published: (2024)
An Ensemble Approach to Short-form Video Quality Assessment Using Multimodal LLM
by: Wen, Wen, et al.
Published: (2024)
by: Wen, Wen, et al.
Published: (2024)
ICANet: A Method of Short Video Emotion Recognition Driven by Multimodal Data
by: Wu, Xuecheng, et al.
Published: (2022)
by: Wu, Xuecheng, et al.
Published: (2022)
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
by: Wei, Jiangchuan, et al.
Published: (2025)
by: Wei, Jiangchuan, et al.
Published: (2025)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
by: Wu, Haoning, et al.
Published: (2024)
by: Wu, Haoning, et al.
Published: (2024)
Poisoning Prompt-Guided Sampling in Video Large Language Models
by: Cao, Yuxin, et al.
Published: (2025)
by: Cao, Yuxin, et al.
Published: (2025)
Multimodal Feature-Driven Deep Learning for the Prediction of Duck Body Dimensions and Weight
by: Xiao, Wenbo, et al.
Published: (2025)
by: Xiao, Wenbo, et al.
Published: (2025)
EGVD: Event-Guided Video Diffusion Model for Physically Realistic Large-Motion Frame Interpolation
by: Zhang, Ziran, et al.
Published: (2025)
by: Zhang, Ziran, et al.
Published: (2025)
Vidi: Large Multimodal Models for Video Understanding and Editing
by: Vidi Team, et al.
Published: (2025)
by: Vidi Team, et al.
Published: (2025)
Can Language Models Laugh at YouTube Short-form Videos?
by: Ko, Dayoon, et al.
Published: (2023)
by: Ko, Dayoon, et al.
Published: (2023)
ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement
by: Ye, Jianping, et al.
Published: (2026)
by: Ye, Jianping, et al.
Published: (2026)
VideoMem: Constructing, Analyzing, Predicting Short-term and Long-term Video Memorability
by: Cohendet, Romain, et al.
Published: (2018)
by: Cohendet, Romain, et al.
Published: (2018)
CPFD: Confidence-aware Privileged Feature Distillation for Short Video Classification
by: Shi, Jinghao, et al.
Published: (2024)
by: Shi, Jinghao, et al.
Published: (2024)
Count Anything at Any Granularity
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
Light-VQA+: A Video Quality Assessment Model for Exposure Correction with Vision-Language Guidance
by: Zhou, Xunchu, et al.
Published: (2024)
by: Zhou, Xunchu, et al.
Published: (2024)
When Rules Fall Short: Agent-Driven Discovery of Emerging Content Issues in Short Video Platforms
by: Yu, Chenghui, et al.
Published: (2026)
by: Yu, Chenghui, et al.
Published: (2026)
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
by: Xu, Weihan, et al.
Published: (2025)
by: Xu, Weihan, et al.
Published: (2025)
Personal Visual Context Learning in Large Multimodal Models
by: Xue, Zihui, et al.
Published: (2026)
by: Xue, Zihui, et al.
Published: (2026)
Transformer-Driven Modeling of Variable Frequency Features for Classifying Student Engagement in Online Learning
by: Mandia, Sandeep, et al.
Published: (2025)
by: Mandia, Sandeep, et al.
Published: (2025)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
by: Ouyang, Kun, et al.
Published: (2025)
by: Ouyang, Kun, et al.
Published: (2025)
V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping
by: Lee, Hyunkoo, et al.
Published: (2025)
by: Lee, Hyunkoo, et al.
Published: (2025)
Bounded-Compute Multimodal Regression for Product-Rating Prediction
by: Leach, William, et al.
Published: (2026)
by: Leach, William, et al.
Published: (2026)
Accelerating Online Mapping and Behavior Prediction via Direct BEV Feature Attention
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
Bitbox: Behavioral Imaging Toolbox for Computational Analysis of Behavior from Videos
by: Sariyanidi, Evangelos, et al.
Published: (2025)
by: Sariyanidi, Evangelos, et al.
Published: (2025)
Similar Items
-
Engagement Prediction of Short Videos with Large Multimodal Models
by: Sun, Wei, et al.
Published: (2025) -
Leveraging Segment Anything Model in Identifying Buildings within Refugee Camps (SAM4Refugee) from Satellite Imagery for Humanitarian Operations
by: Gao, Yunya
Published: (2024) -
Delving Deep into Engagement Prediction of Short Videos
by: Li, Dasong, et al.
Published: (2024) -
VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results
by: Li, Dasong, et al.
Published: (2025) -
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
by: Jin, Xiaojie, et al.
Published: (2023)