End-to-end Semantic-centric Video-based Multimodal Affective Computing
Fuente:
arXiv
Salvato in:
| Autori principali: | Lin, Ronghao, Zeng, Ying, Mai, Sijie, Hu, Haifeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multi-source Multimodal Progressive Domain Adaption for Audio-Visual Deception Detection
di: Lin, Ronghao, et al.
Pubblicazione: (2025)
di: Lin, Ronghao, et al.
Pubblicazione: (2025)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024)
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024)
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
di: Wei, Yake, et al.
Pubblicazione: (2024)
di: Wei, Yake, et al.
Pubblicazione: (2024)
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
di: Lin, Yan-Bo, et al.
Pubblicazione: (2026)
di: Lin, Yan-Bo, et al.
Pubblicazione: (2026)
On-the-fly Modulation for Balanced Multimodal Learning
di: Wei, Yake, et al.
Pubblicazione: (2024)
di: Wei, Yake, et al.
Pubblicazione: (2024)
SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation
di: Li, Xuewei, et al.
Pubblicazione: (2023)
di: Li, Xuewei, et al.
Pubblicazione: (2023)
Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning
di: Zhang, Haojie, et al.
Pubblicazione: (2025)
di: Zhang, Haojie, et al.
Pubblicazione: (2025)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
di: Wang, Zhouxia, et al.
Pubblicazione: (2023)
di: Wang, Zhouxia, et al.
Pubblicazione: (2023)
STIV: Scalable Text and Image Conditioned Video Generation
di: Lin, Zongyu, et al.
Pubblicazione: (2024)
di: Lin, Zongyu, et al.
Pubblicazione: (2024)
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
di: Chen, Weifeng, et al.
Pubblicazione: (2023)
di: Chen, Weifeng, et al.
Pubblicazione: (2023)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
di: Lin, Yuanze, et al.
Pubblicazione: (2025)
di: Lin, Yuanze, et al.
Pubblicazione: (2025)
Video-based Music Generation
di: Sulun, Serkan
Pubblicazione: (2026)
di: Sulun, Serkan
Pubblicazione: (2026)
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
di: Wang, Qiang, et al.
Pubblicazione: (2025)
di: Wang, Qiang, et al.
Pubblicazione: (2025)
Evaluating the Impact of Point Cloud Colorization on Semantic Segmentation Accuracy
di: Zhu, Qinfeng, et al.
Pubblicazione: (2024)
di: Zhu, Qinfeng, et al.
Pubblicazione: (2024)
Multi-layer Learnable Attention Mask for Multimodal Tasks
di: Barrios, Wayner, et al.
Pubblicazione: (2024)
di: Barrios, Wayner, et al.
Pubblicazione: (2024)
RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
di: Lin, Xiang, et al.
Pubblicazione: (2025)
di: Lin, Xiang, et al.
Pubblicazione: (2025)
Diffusion Model-Based Video Editing: A Survey
di: Sun, Wenhao, et al.
Pubblicazione: (2024)
di: Sun, Wenhao, et al.
Pubblicazione: (2024)
HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
di: Wang, Zhecan, et al.
Pubblicazione: (2024)
di: Wang, Zhecan, et al.
Pubblicazione: (2024)
An Efficient and Explanatory Image and Text Clustering System with Multimodal Autoencoder Architecture
di: Shi, Tiancheng, et al.
Pubblicazione: (2024)
di: Shi, Tiancheng, et al.
Pubblicazione: (2024)
Multimodal Long Video Modeling Based on Temporal Dynamic Context
di: Hao, Haoran, et al.
Pubblicazione: (2025)
di: Hao, Haoran, et al.
Pubblicazione: (2025)
VSTAR: Generative Temporal Nursing for Longer Dynamic Video Synthesis
di: Li, Yumeng, et al.
Pubblicazione: (2024)
di: Li, Yumeng, et al.
Pubblicazione: (2024)
Reinforcement Learning for Unsupervised Video Summarization with Reward Generator Training
di: Abbasi, Mehryar, et al.
Pubblicazione: (2024)
di: Abbasi, Mehryar, et al.
Pubblicazione: (2024)
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
di: Dong, Hao, et al.
Pubblicazione: (2026)
di: Dong, Hao, et al.
Pubblicazione: (2026)
COSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time Adaptation
di: Huang, Fanding, et al.
Pubblicazione: (2025)
di: Huang, Fanding, et al.
Pubblicazione: (2025)
ClassWise-CRF: Category-Specific Fusion for Enhanced Semantic Segmentation of Remote Sensing Imagery
di: Zhu, Qinfeng, et al.
Pubblicazione: (2025)
di: Zhu, Qinfeng, et al.
Pubblicazione: (2025)
Latent Space Probing for Adult Content Detection in Video Generative Models
di: Khatri, Alizishaan, et al.
Pubblicazione: (2026)
di: Khatri, Alizishaan, et al.
Pubblicazione: (2026)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
di: Li, Quanhao, et al.
Pubblicazione: (2025)
di: Li, Quanhao, et al.
Pubblicazione: (2025)
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
di: Li, Quanhao, et al.
Pubblicazione: (2026)
di: Li, Quanhao, et al.
Pubblicazione: (2026)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
di: Yang, Dejie, et al.
Pubblicazione: (2024)
di: Yang, Dejie, et al.
Pubblicazione: (2024)
Lightning Fast Video Anomaly Detection via Adversarial Knowledge Distillation
di: Croitoru, Florinel-Alin, et al.
Pubblicazione: (2022)
di: Croitoru, Florinel-Alin, et al.
Pubblicazione: (2022)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
di: Li, Liupeng, et al.
Pubblicazione: (2026)
di: Li, Liupeng, et al.
Pubblicazione: (2026)
Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
di: Yu, Lijun
Pubblicazione: (2024)
di: Yu, Lijun
Pubblicazione: (2024)
Video Face Re-Aging: Toward Temporally Consistent Face Re-Aging
di: Muqeet, Abdul, et al.
Pubblicazione: (2023)
di: Muqeet, Abdul, et al.
Pubblicazione: (2023)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
di: Zheng, Sixiao, et al.
Pubblicazione: (2025)
di: Zheng, Sixiao, et al.
Pubblicazione: (2025)
MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark
di: Croitoru, Florinel-Alin, et al.
Pubblicazione: (2025)
di: Croitoru, Florinel-Alin, et al.
Pubblicazione: (2025)
LayerT2V: A Unified Multi-Layer Video Generation Framework
di: Li, Guangzhao, et al.
Pubblicazione: (2025)
di: Li, Guangzhao, et al.
Pubblicazione: (2025)
Scaling Spatial Intelligence with Multimodal Foundation Models
di: Cai, Zhongang, et al.
Pubblicazione: (2025)
di: Cai, Zhongang, et al.
Pubblicazione: (2025)
Using AI to Summarize US Presidential Campaign TV Advertisement Videos, 1952-2012
di: Breuer, Adam, et al.
Pubblicazione: (2025)
di: Breuer, Adam, et al.
Pubblicazione: (2025)
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
di: Chen, Baiyu, et al.
Pubblicazione: (2025)
di: Chen, Baiyu, et al.
Pubblicazione: (2025)
Low-Resolution Object Recognition with Cross-Resolution Relational Contrastive Distillation
di: Zhang, Kangkai, et al.
Pubblicazione: (2024)
di: Zhang, Kangkai, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Multi-source Multimodal Progressive Domain Adaption for Audio-Visual Deception Detection
di: Lin, Ronghao, et al.
Pubblicazione: (2025) -
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024) -
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
di: Wei, Yake, et al.
Pubblicazione: (2024) -
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
di: Lin, Yan-Bo, et al.
Pubblicazione: (2026) -
On-the-fly Modulation for Balanced Multimodal Learning
di: Wei, Yake, et al.
Pubblicazione: (2024)