A Multimodal Memes Classification: A Survey and Open Research Issues
Fuente:
arXiv
Guardado en:
| Autores principales: | Afridi, Tariq Habib, Alam, Aftab, Khan, Muhammad Numan, Khan, Jawad, Lee, Young-Koo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2020
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
por: An, Wenbin, et al.
Publicado: (2025)
por: An, Wenbin, et al.
Publicado: (2025)
More than Memes: A Multimodal Topic Modeling Approach to Conspiracy Theories on Telegram
por: Steffen, Elisabeth
Publicado: (2024)
por: Steffen, Elisabeth
Publicado: (2024)
FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection
por: Bhaskar, Paramananda, et al.
Publicado: (2026)
por: Bhaskar, Paramananda, et al.
Publicado: (2026)
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
por: Shah, Siddhant Bikram, et al.
Publicado: (2024)
por: Shah, Siddhant Bikram, et al.
Publicado: (2024)
The Revolution of Multimodal Large Language Models: A Survey
por: Caffagni, Davide, et al.
Publicado: (2024)
por: Caffagni, Davide, et al.
Publicado: (2024)
A Survey of Multimodal Large Language Model from A Data-centric Perspective
por: Bai, Tianyi, et al.
Publicado: (2024)
por: Bai, Tianyi, et al.
Publicado: (2024)
Subjective evaluation of UHD video coded using VVC with LCEVC and ML-VVC
por: Ramzan, Naeem, et al.
Publicado: (2026)
por: Ramzan, Naeem, et al.
Publicado: (2026)
LLMs Meet Multimodal Generation and Editing: A Survey
por: He, Yingqing, et al.
Publicado: (2024)
por: He, Yingqing, et al.
Publicado: (2024)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
por: Li, Yunxin, et al.
Publicado: (2024)
por: Li, Yunxin, et al.
Publicado: (2024)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
por: Zhang, Xueqiao, et al.
Publicado: (2025)
por: Zhang, Xueqiao, et al.
Publicado: (2025)
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
por: Cheng, Zebang, et al.
Publicado: (2024)
por: Cheng, Zebang, et al.
Publicado: (2024)
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language
por: Biswas, Subrata, et al.
Publicado: (2025)
por: Biswas, Subrata, et al.
Publicado: (2025)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
por: Jiang, Jingjing, et al.
Publicado: (2025)
por: Jiang, Jingjing, et al.
Publicado: (2025)
Harmfully Manipulated Images Matter in Multimodal Misinformation Detection
por: Wang, Bing, et al.
Publicado: (2024)
por: Wang, Bing, et al.
Publicado: (2024)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
por: Bin, Yi, et al.
Publicado: (2024)
por: Bin, Yi, et al.
Publicado: (2024)
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
por: Song, Shezheng, et al.
Publicado: (2023)
por: Song, Shezheng, et al.
Publicado: (2023)
NVLM: Open Frontier-Class Multimodal LLMs
por: Dai, Wenliang, et al.
Publicado: (2024)
por: Dai, Wenliang, et al.
Publicado: (2024)
Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment
por: Wang, Kangsheng, et al.
Publicado: (2025)
por: Wang, Kangsheng, et al.
Publicado: (2025)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
por: Gan, Chengguang, et al.
Publicado: (2025)
por: Gan, Chengguang, et al.
Publicado: (2025)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
por: Zhao, Zhixian, et al.
Publicado: (2026)
por: Zhao, Zhixian, et al.
Publicado: (2026)
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
por: Masumura, Ryo, et al.
Publicado: (2025)
por: Masumura, Ryo, et al.
Publicado: (2025)
Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
por: Yan, Zehong, et al.
Publicado: (2025)
por: Yan, Zehong, et al.
Publicado: (2025)
Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
por: Wang, Bing, et al.
Publicado: (2025)
por: Wang, Bing, et al.
Publicado: (2025)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
por: Chen, Qian, et al.
Publicado: (2026)
por: Chen, Qian, et al.
Publicado: (2026)
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue
por: Ghaleb, Esam, et al.
Publicado: (2025)
por: Ghaleb, Esam, et al.
Publicado: (2025)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
por: Wu, Jiaying, et al.
Publicado: (2025)
por: Wu, Jiaying, et al.
Publicado: (2025)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
por: Jiang, Chaoya, et al.
Publicado: (2024)
por: Jiang, Chaoya, et al.
Publicado: (2024)
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
por: Li, Jinyuan, et al.
Publicado: (2024)
por: Li, Jinyuan, et al.
Publicado: (2024)
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
por: Zhang, Dongxu, et al.
Publicado: (2026)
por: Zhang, Dongxu, et al.
Publicado: (2026)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
por: Liang, Zhengyang, et al.
Publicado: (2024)
por: Liang, Zhengyang, et al.
Publicado: (2024)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
por: Lai, Zhengzhao, et al.
Publicado: (2025)
por: Lai, Zhengzhao, et al.
Publicado: (2025)
A New Hybrid Intelligent Approach for Multimodal Detection of Suspected Disinformation on TikTok
por: Guerrero-Sosa, Jared D. T., et al.
Publicado: (2025)
por: Guerrero-Sosa, Jared D. T., et al.
Publicado: (2025)
PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents
por: Wang, Junjie, et al.
Publicado: (2024)
por: Wang, Junjie, et al.
Publicado: (2024)
VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations
por: Galarnyk, Michael, et al.
Publicado: (2025)
por: Galarnyk, Michael, et al.
Publicado: (2025)
Securing Social Media Against Deepfakes using Identity, Behavioral, and Geometric Signatures
por: Farooq, Muhammad Umar, et al.
Publicado: (2024)
por: Farooq, Muhammad Umar, et al.
Publicado: (2024)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
por: Wang, Wenxuan, et al.
Publicado: (2025)
por: Wang, Wenxuan, et al.
Publicado: (2025)
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
por: Song, Dingjie, et al.
Publicado: (2024)
por: Song, Dingjie, et al.
Publicado: (2024)
Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection
por: Nishimura, Taichi, et al.
Publicado: (2024)
por: Nishimura, Taichi, et al.
Publicado: (2024)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
por: Wang, Xinran, et al.
Publicado: (2026)
por: Wang, Xinran, et al.
Publicado: (2026)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
por: Ji, Yatai, et al.
Publicado: (2024)
por: Ji, Yatai, et al.
Publicado: (2024)
Ejemplares similares
-
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
por: An, Wenbin, et al.
Publicado: (2025) -
More than Memes: A Multimodal Topic Modeling Approach to Conspiracy Theories on Telegram
por: Steffen, Elisabeth
Publicado: (2024) -
FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection
por: Bhaskar, Paramananda, et al.
Publicado: (2026) -
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
por: Shah, Siddhant Bikram, et al.
Publicado: (2024) -
The Revolution of Multimodal Large Language Models: A Survey
por: Caffagni, Davide, et al.
Publicado: (2024)