Probing Memes in LLMs: A Paradigm for the Entangled Evaluation World
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Luzhou, Yang, Zhengxin, Ji, Honglu, Yang, Yikang, Fan, Fanda, Gao, Wanling, Ge, Jiayuan, Han, Yilin, Zhan, Jianfeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GraDE: A Graph Diffusion Estimator for Frequent Subgraph Discovery in Neural Architectures
by: Yang, Yikang, et al.
Published: (2026)
by: Yang, Yikang, et al.
Published: (2026)
Younger: The First Dataset for Artificial Intelligence-Generated Neural Network Architecture
by: Yang, Zhengxin, et al.
Published: (2024)
by: Yang, Zhengxin, et al.
Published: (2024)
AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI
by: Fan, Fanda, et al.
Published: (2024)
by: Fan, Fanda, et al.
Published: (2024)
Achieving Consistent and Comparable CPU Evaluation
by: Wang, Chenxi, et al.
Published: (2024)
by: Wang, Chenxi, et al.
Published: (2024)
On Meta-Evaluation
by: Li, Hongxiao, et al.
Published: (2025)
by: Li, Hongxiao, et al.
Published: (2025)
TimeMosaic: Temporal Heterogeneity Guided Time Series Forecasting via Adaptive Granularity Patch and Segment-wise Decoding
by: Ding, Kuiye, et al.
Published: (2025)
by: Ding, Kuiye, et al.
Published: (2025)
Attributing the System's Overall Effect to its Components
by: Wang, Chenxi, et al.
Published: (2026)
by: Wang, Chenxi, et al.
Published: (2026)
Quality at the Tail of Machine Learning Inference
by: Yang, Zhengxin, et al.
Published: (2022)
by: Yang, Zhengxin, et al.
Published: (2022)
List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
by: Yan, An, et al.
Published: (2024)
by: Yan, An, et al.
Published: (2024)
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
by: Ge, Suyu, et al.
Published: (2023)
by: Ge, Suyu, et al.
Published: (2023)
Multi-Granular Multimodal Clue Fusion for Meme Understanding
by: Zheng, Li, et al.
Published: (2025)
by: Zheng, Li, et al.
Published: (2025)
YaRN: Efficient Context Window Extension of Large Language Models
by: Peng, Bowen, et al.
Published: (2023)
by: Peng, Bowen, et al.
Published: (2023)
MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
Can LLMs Infer Personality from Real World Conversations?
by: Zhu, Jianfeng, et al.
Published: (2025)
by: Zhu, Jianfeng, et al.
Published: (2025)
Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes
by: Wang, Weiming, et al.
Published: (2026)
by: Wang, Weiming, et al.
Published: (2026)
MemeCraft: Contextual and Stance-Driven Multimodal Meme Generation
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
CombinationTS: A Modular Framework for Understanding Time-Series Forecasting Models
by: Wang, Xiaorui, et al.
Published: (2026)
by: Wang, Xiaorui, et al.
Published: (2026)
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
by: Ge, Tao, et al.
Published: (2026)
by: Ge, Tao, et al.
Published: (2026)
AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness
by: Chen, Zixin, et al.
Published: (2025)
by: Chen, Zixin, et al.
Published: (2025)
Meme-ingful Analysis: Enhanced Understanding of Cyberbullying in Memes Through Multimodal Explanations
by: Jha, Prince, et al.
Published: (2024)
by: Jha, Prince, et al.
Published: (2024)
They Said Memes Were Harmless-We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural References
by: Tripathi, Sahil, et al.
Published: (2026)
by: Tripathi, Sahil, et al.
Published: (2026)
Evaluating the Effectiveness of Black-Box Prompt Optimization as the Scale of LLMs Continues to Grow
by: Zhou, Ziyu, et al.
Published: (2025)
by: Zhou, Ziyu, et al.
Published: (2025)
Can Language Models Follow Multiple Turns of Entangled Instructions?
by: Han, Chi, et al.
Published: (2025)
by: Han, Chi, et al.
Published: (2025)
Towards Comprehensive Detection of Chinese Harmful Memes
by: Lu, Junyu, et al.
Published: (2024)
by: Lu, Junyu, et al.
Published: (2024)
Bridging the Gap Between Domain-specific Frameworks and Multiple Hardware Devices
by: Wen, Xu, et al.
Published: (2024)
by: Wen, Xu, et al.
Published: (2024)
Towards Explainable Harmful Meme Detection through Multimodal Debate between Large Language Models
by: Lin, Hongzhan, et al.
Published: (2024)
by: Lin, Hongzhan, et al.
Published: (2024)
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
by: Shah, Siddhant Bikram, et al.
Published: (2024)
by: Shah, Siddhant Bikram, et al.
Published: (2024)
LLM Probe: Evaluating LLMs for Low-Resource Languages
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2026)
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2026)
I Came, I Saw, I Explained: Benchmarking Multimodal LLMs on Figurative Meaning in Memes
by: Zhou, Shijia, et al.
Published: (2026)
by: Zhou, Shijia, et al.
Published: (2026)
Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme Detection
by: Pan, Fengjun, et al.
Published: (2025)
by: Pan, Fengjun, et al.
Published: (2025)
MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention
by: Jha, Prince, et al.
Published: (2024)
by: Jha, Prince, et al.
Published: (2024)
MemeLens: Multilingual Multitask VLMs for Memes
by: Shahroor, Ali Ezzat, et al.
Published: (2026)
by: Shahroor, Ali Ezzat, et al.
Published: (2026)
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing
by: Agarwal, Siddhant, et al.
Published: (2024)
by: Agarwal, Siddhant, et al.
Published: (2024)
CoT Vectors: Transferring and Probing the Reasoning Mechanisms of LLMs
by: Li, Li, et al.
Published: (2025)
by: Li, Li, et al.
Published: (2025)
TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice
by: Hu, Gang, et al.
Published: (2026)
by: Hu, Gang, et al.
Published: (2026)
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
by: Lu, Zimu, et al.
Published: (2025)
by: Lu, Zimu, et al.
Published: (2025)
Beyond Meme Templates: Limitations of Visual Similarity Measures in Meme Matching
by: Hazman, Muzhaffar, et al.
Published: (2025)
by: Hazman, Muzhaffar, et al.
Published: (2025)
Evaluating Span Extraction in Generative Paradigm: A Reflection on Aspect-Based Sentiment Analysis
by: Yang, Soyoung, et al.
Published: (2024)
by: Yang, Soyoung, et al.
Published: (2024)
Large Emotional World Model
by: Song, Changhao, et al.
Published: (2025)
by: Song, Changhao, et al.
Published: (2025)
M-QUEST -- Meme Question-Understanding Evaluation on Semantics and Toxicity
by: De Giorgis, Stefano, et al.
Published: (2026)
by: De Giorgis, Stefano, et al.
Published: (2026)
Similar Items
-
GraDE: A Graph Diffusion Estimator for Frequent Subgraph Discovery in Neural Architectures
by: Yang, Yikang, et al.
Published: (2026) -
Younger: The First Dataset for Artificial Intelligence-Generated Neural Network Architecture
by: Yang, Zhengxin, et al.
Published: (2024) -
AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI
by: Fan, Fanda, et al.
Published: (2024) -
Achieving Consistent and Comparable CPU Evaluation
by: Wang, Chenxi, et al.
Published: (2024) -
On Meta-Evaluation
by: Li, Hongxiao, et al.
Published: (2025)