Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Hossain, Eftekhar, Sharif, Omar, Hoque, Mohammed Moshiul, Preum, Sarah M. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deciphering Hate: Identifying Hateful Memes and Their Targets
by: Hossain, Eftekhar, et al.
Published: (2024)
by: Hossain, Eftekhar, et al.
Published: (2024)
From Sight to Insight: Improving Visual Reasoning Capabilities of Multimodal Models via Reinforcement Learning
by: Sharif, Omar, et al.
Published: (2026)
by: Sharif, Omar, et al.
Published: (2026)
Retriv at BLP-2025 Task 1: A Transformer Ensemble and Multi-Task Learning Approach for Bangla Hate Speech Identification
by: Saha, Sourav, et al.
Published: (2025)
by: Saha, Sourav, et al.
Published: (2025)
Explicit, Implicit, and Scattered: Revisiting Event Extraction to Capture Complex Arguments
by: Sharif, Omar, et al.
Published: (2024)
by: Sharif, Omar, et al.
Published: (2024)
REGen: A Reliable Evaluation Framework for Generative Event Argument Extraction
by: Sharif, Omar, et al.
Published: (2025)
by: Sharif, Omar, et al.
Published: (2025)
Retriv at BLP-2025 Task 2: Test-Driven Feedback-Guided Framework for Bangla-to-Python Code Generation
by: Asib, K M Nafi, et al.
Published: (2025)
by: Asib, K M Nafi, et al.
Published: (2025)
Large Language Models for Document-Level Event-Argument Data Augmentation for Challenging Role Types
by: Gatto, Joseph, et al.
Published: (2024)
by: Gatto, Joseph, et al.
Published: (2024)
Detecting Hate and Inflammatory Content in Bengali Memes: A New Multimodal Dataset and Co-Attention Framework
by: Ullah, Rakib, et al.
Published: (2026)
by: Ullah, Rakib, et al.
Published: (2026)
Aligning Text, Code, and Vision: A Multi-Objective Reinforcement Learning Framework for Text-to-Visualization
by: Rahman, Mizanur, et al.
Published: (2026)
by: Rahman, Mizanur, et al.
Published: (2026)
Do LLMs Find Human Answers To Fact-Driven Questions Perplexing? A Case Study on Reddit
by: Seegmiller, Parker, et al.
Published: (2024)
by: Seegmiller, Parker, et al.
Published: (2024)
Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection
by: Brook, Joshua Wolfe, et al.
Published: (2025)
by: Brook, Joshua Wolfe, et al.
Published: (2025)
TRACE: Textual Relevance Augmentation and Contextual Encoding for Multimodal Hate Detection
by: Koushik, Girish A., et al.
Published: (2025)
by: Koushik, Girish A., et al.
Published: (2025)
Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
by: Eilertsen, Brage, et al.
Published: (2025)
by: Eilertsen, Brage, et al.
Published: (2025)
Automatic Textual Normalization for Hate Speech Detection
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2023)
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2023)
Gradient Masters at BLP-2025 Task 1: Advancing Low-Resource NLP for Bengali using Ensemble-Based Adversarial Training for Hate Speech Detection
by: Hoque, Syed Mohaiminul, et al.
Published: (2025)
by: Hoque, Syed Mohaiminul, et al.
Published: (2025)
Demystifying Hateful Content: Leveraging Large Multimodal Models for Hateful Meme Detection with Explainable Decisions
by: Hee, Ming Shan, et al.
Published: (2025)
by: Hee, Ming Shan, et al.
Published: (2025)
LATTE: Learning Aligned Transactions and Textual Embeddings for Bank Clients
by: Fadeev, Egor, et al.
Published: (2025)
by: Fadeev, Egor, et al.
Published: (2025)
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes
by: Wang, Chuhan, et al.
Published: (2026)
by: Wang, Chuhan, et al.
Published: (2026)
Measuring Distribution Shift in User Prompts and Its Effects on LLM Performance
by: Seegmiller, Parker, et al.
Published: (2026)
by: Seegmiller, Parker, et al.
Published: (2026)
Towards Aligning Language Models with Textual Feedback
by: Lloret, Saüc Abadal, et al.
Published: (2024)
by: Lloret, Saüc Abadal, et al.
Published: (2024)
Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models
by: Van, Minh-Hao, et al.
Published: (2025)
by: Van, Minh-Hao, et al.
Published: (2025)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
by: Hashemi, Mohammad Abuzar, et al.
Published: (2021)
by: Hashemi, Mohammad Abuzar, et al.
Published: (2021)
HateSieve: A Contrastive Learning Framework for Detecting and Segmenting Hateful Content in Multimodal Memes
by: Su, Xuanyu, et al.
Published: (2024)
by: Su, Xuanyu, et al.
Published: (2024)
Natural Language Generation for Visualizations: State of the Art, Challenges and Future Directions
by: Hoque, Enamul, et al.
Published: (2024)
by: Hoque, Enamul, et al.
Published: (2024)
Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing
by: Jiang, Zifan, et al.
Published: (2025)
by: Jiang, Zifan, et al.
Published: (2025)
Socially Constructed Treatment Plans: Analyzing Online Peer Interactions to Understand How Patients Navigate Complex Medical Conditions
by: Basak, Madhusudan, et al.
Published: (2025)
by: Basak, Madhusudan, et al.
Published: (2025)
Beyond Hate: Differentiating Uncivil and Intolerant Speech in Multimodal Content Moderation
by: Herrmann, Nils A., et al.
Published: (2026)
by: Herrmann, Nils A., et al.
Published: (2026)
The Critical Role of Aspects in Measuring Document Similarity
by: Hossain, Eftekhar, et al.
Published: (2026)
by: Hossain, Eftekhar, et al.
Published: (2026)
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment
by: Azad, Asif, et al.
Published: (2026)
by: Azad, Asif, et al.
Published: (2026)
Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language Models
by: Bui, Minh Duc, et al.
Published: (2024)
by: Bui, Minh Duc, et al.
Published: (2024)
Triple Modality Fusion: Aligning Visual, Textual, and Graph Data with Large Language Models for Multi-Behavior Recommendations
by: Ma, Luyi, et al.
Published: (2024)
by: Ma, Luyi, et al.
Published: (2024)
MasonPerplexity at Multimodal Hate Speech Event Detection 2024: Hate Speech and Target Detection Using Transformer Ensembles
by: Ganguly, Amrita, et al.
Published: (2024)
by: Ganguly, Amrita, et al.
Published: (2024)
WaveMind: Towards a Conversational EEG Foundation Model Aligned to Textual and Visual Modalities
by: Zeng, Ziyi, et al.
Published: (2025)
by: Zeng, Ziyi, et al.
Published: (2025)
Standardize: Aligning Language Models with Expert-Defined Standards for Content Generation
by: Imperial, Joseph Marvin, et al.
Published: (2024)
by: Imperial, Joseph Marvin, et al.
Published: (2024)
Depth $F_1$: Improving Evaluation of Cross-Domain Text Classification by Measuring Semantic Generalizability
by: Seegmiller, Parker, et al.
Published: (2024)
by: Seegmiller, Parker, et al.
Published: (2024)
The Impact of Persona-based Political Perspectives on Hateful Content Detection
by: Civelli, Stefano, et al.
Published: (2025)
by: Civelli, Stefano, et al.
Published: (2025)
Hierarchical Aligned Multimodal Learning for NER on Tweet Posts
by: Liu, Peipei, et al.
Published: (2023)
by: Liu, Peipei, et al.
Published: (2023)
Aligning Language Models with Demonstrated Feedback
by: Shaikh, Omar, et al.
Published: (2024)
by: Shaikh, Omar, et al.
Published: (2024)
Hate Content Detection via Novel Pre-Processing Sequencing and Ensemble Methods
by: Chhabra, Anusha, et al.
Published: (2024)
by: Chhabra, Anusha, et al.
Published: (2024)
Similar Items
-
Deciphering Hate: Identifying Hateful Memes and Their Targets
by: Hossain, Eftekhar, et al.
Published: (2024) -
From Sight to Insight: Improving Visual Reasoning Capabilities of Multimodal Models via Reinforcement Learning
by: Sharif, Omar, et al.
Published: (2026) -
Retriv at BLP-2025 Task 1: A Transformer Ensemble and Multi-Task Learning Approach for Bangla Hate Speech Identification
by: Saha, Sourav, et al.
Published: (2025) -
Explicit, Implicit, and Scattered: Revisiting Event Extraction to Capture Complex Arguments
by: Sharif, Omar, et al.
Published: (2024) -
REGen: A Reliable Evaluation Framework for Generative Event Argument Extraction
by: Sharif, Omar, et al.
Published: (2025)