MM-HSD: Multi-Modal Hate Speech Detection in Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Céspedes-Sarrias, Berta, Collado-Capell, Carlos, Rodenas-Ruiz, Pablo, Hrynenko, Olena, Cavallaro, Andrea |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion
von: Zhang, Yinghui, et al.
Veröffentlicht: (2025)
von: Zhang, Yinghui, et al.
Veröffentlicht: (2025)
Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos
von: Tian, Chong, et al.
Veröffentlicht: (2026)
von: Tian, Chong, et al.
Veröffentlicht: (2026)
MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and Bilibili
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
MM-Point: Multi-View Information-Enhanced Multi-Modal Self-Supervised 3D Point Cloud Understanding
von: Yu, Hai-Tao, et al.
Veröffentlicht: (2024)
von: Yu, Hai-Tao, et al.
Veröffentlicht: (2024)
CARAT: Contrastive Feature Reconstruction and Aggregation for Multi-Modal Multi-Label Emotion Recognition
von: Peng, Cheng, et al.
Veröffentlicht: (2023)
von: Peng, Cheng, et al.
Veröffentlicht: (2023)
A Survey on Music Generation from Single-Modal, Cross-Modal, and Multi-Modal Perspectives
von: Li, Shuyu, et al.
Veröffentlicht: (2025)
von: Li, Shuyu, et al.
Veröffentlicht: (2025)
M3TR: Temporal Retrieval Enhanced Multi-Modal Micro-video Popularity Prediction
von: Lu, Jiacheng, et al.
Veröffentlicht: (2024)
von: Lu, Jiacheng, et al.
Veröffentlicht: (2024)
HateSieve: A Contrastive Learning Framework for Detecting and Segmenting Hateful Content in Multimodal Memes
von: Su, Xuanyu, et al.
Veröffentlicht: (2024)
von: Su, Xuanyu, et al.
Veröffentlicht: (2024)
TIP and Polish: Text-Image-Prototype Guided Multi-Modal Generation via Commonality-Discrepancy Modeling and Refinement
von: Ma, Zhiyong, et al.
Veröffentlicht: (2025)
von: Ma, Zhiyong, et al.
Veröffentlicht: (2025)
Wireless Video Semantic Communication with Decoupled Diffusion Multi-frame Compensation
von: Xie, Bingyan, et al.
Veröffentlicht: (2025)
von: Xie, Bingyan, et al.
Veröffentlicht: (2025)
Sequence-to-Sequence Multi-Modal Speech In-Painting
von: Elyaderani, Mahsa Kadkhodaei, et al.
Veröffentlicht: (2024)
von: Elyaderani, Mahsa Kadkhodaei, et al.
Veröffentlicht: (2024)
EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction
von: Jing, Chong, et al.
Veröffentlicht: (2026)
von: Jing, Chong, et al.
Veröffentlicht: (2026)
Identifying Privacy Personas
von: Hrynenko, Olena, et al.
Veröffentlicht: (2024)
von: Hrynenko, Olena, et al.
Veröffentlicht: (2024)
OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
von: Chen, Junzhe, et al.
Veröffentlicht: (2025)
von: Chen, Junzhe, et al.
Veröffentlicht: (2025)
Towards Temporal-Aware Multi-Modal Retrieval Augmented Generation in Finance
von: Zhu, Fengbin, et al.
Veröffentlicht: (2025)
von: Zhu, Fengbin, et al.
Veröffentlicht: (2025)
Robust Multi-Modal Speech In-Painting: A Sequence-to-Sequence Approach
von: Elyaderani, Mahsa Kadkhodaei, et al.
Veröffentlicht: (2024)
von: Elyaderani, Mahsa Kadkhodaei, et al.
Veröffentlicht: (2024)
Semantic-Guided Unsupervised Video Summarization
von: Liu, Haizhou, et al.
Veröffentlicht: (2026)
von: Liu, Haizhou, et al.
Veröffentlicht: (2026)
FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection
von: Bhaskar, Paramananda, et al.
Veröffentlicht: (2026)
von: Bhaskar, Paramananda, et al.
Veröffentlicht: (2026)
MaLoRA: Gated Modality LoRA for Key-Space Alignment in Multimodal LLM Fine-Tuning
von: Zheng, Xinhan, et al.
Veröffentlicht: (2025)
von: Zheng, Xinhan, et al.
Veröffentlicht: (2025)
Towards Open-Vocabulary Video Semantic Segmentation
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
HeLo: Heterogeneous Multi-Modal Fusion with Label Correlation for Emotion Distribution Learning
von: Zheng, Chuhang, et al.
Veröffentlicht: (2025)
von: Zheng, Chuhang, et al.
Veröffentlicht: (2025)
Start from Video-Music Retrieval: An Inter-Intra Modal Loss for Cross Modal Retrieval
von: Chen, Zeyu, et al.
Veröffentlicht: (2024)
von: Chen, Zeyu, et al.
Veröffentlicht: (2024)
Advance Fake Video Detection via Vision Transformers
von: Battocchio, Joy, et al.
Veröffentlicht: (2025)
von: Battocchio, Joy, et al.
Veröffentlicht: (2025)
Adaptive Multi-Agent Reasoning for Text-to-Video Retrieval
von: Wu, Jiaxin, et al.
Veröffentlicht: (2025)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2025)
Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
von: Yu, Lijun
Veröffentlicht: (2024)
von: Yu, Lijun
Veröffentlicht: (2024)
Multimodal Emotion Recognition by Fusing Video Semantic in MOOC Learning Scenarios
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
von: Li, Haitian, et al.
Veröffentlicht: (2026)
von: Li, Haitian, et al.
Veröffentlicht: (2026)
Real-Time Mobile Video Analytics for Pre-arrival Emergency Medical Services
von: Jin, Liuyi, et al.
Veröffentlicht: (2025)
von: Jin, Liuyi, et al.
Veröffentlicht: (2025)
Attention of a Kiss: Exploring Attention Maps in Video Diffusion for XAIxArts
von: Cole, Adam, et al.
Veröffentlicht: (2025)
von: Cole, Adam, et al.
Veröffentlicht: (2025)
QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models
von: Lin, Zixing, et al.
Veröffentlicht: (2026)
von: Lin, Zixing, et al.
Veröffentlicht: (2026)
LL-GABR: Energy Efficient Live Video Streaming Using Reinforcement Learning
von: Raman, Adithya, et al.
Veröffentlicht: (2024)
von: Raman, Adithya, et al.
Veröffentlicht: (2024)
Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction
von: Wang, Dali, et al.
Veröffentlicht: (2026)
von: Wang, Dali, et al.
Veröffentlicht: (2026)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
End-to-End Learning-based Video Streaming Enhancement Pipeline: A Generative AI Approach
von: Artioli, Emanuele, et al.
Veröffentlicht: (2025)
von: Artioli, Emanuele, et al.
Veröffentlicht: (2025)
Plasticity-Aware Mixture of Experts for Learning Under QoE Shifts in Adaptive Video Streaming
von: He, Zhiqiang, et al.
Veröffentlicht: (2025)
von: He, Zhiqiang, et al.
Veröffentlicht: (2025)
AxiomVision: Accuracy-Guaranteed Adaptive Visual Model Selection for Perspective-Aware Video Analytics
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2024)
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2024)
Solving Copyright Infringement on Short Video Platforms: Novel Datasets and an Audio Restoration Deep Learning Pipeline
von: Oh, Minwoo, et al.
Veröffentlicht: (2025)
von: Oh, Minwoo, et al.
Veröffentlicht: (2025)
BDIQA: A New Dataset for Video Question Answering to Explore Cognitive Reasoning through Theory of Mind
von: Mao, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Mao, Yuanyuan, et al.
Veröffentlicht: (2024)
LUST: A Multi-Modal Framework with Hierarchical LLM-based Scoring for Learned Thematic Significance Tracking in Multimedia Content
von: Luiz, Anderson de Lima
Veröffentlicht: (2025)
von: Luiz, Anderson de Lima
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion
von: Zhang, Yinghui, et al.
Veröffentlicht: (2025) -
Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos
von: Tian, Chong, et al.
Veröffentlicht: (2026) -
MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and Bilibili
von: Wang, Han, et al.
Veröffentlicht: (2024) -
MM-Point: Multi-View Information-Enhanced Multi-Modal Self-Supervised 3D Point Cloud Understanding
von: Yu, Hai-Tao, et al.
Veröffentlicht: (2024) -
CARAT: Contrastive Feature Reconstruction and Aggregation for Multi-Modal Multi-Label Emotion Recognition
von: Peng, Cheng, et al.
Veröffentlicht: (2023)