Saved in:
| Main Authors: | Cupini, Paolo, Pierri, Francesco |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.26772 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rainbow Noise: Stress-Testing Multimodal Harmful-Meme Detectors on LGBTQ Content
by: Tong, Ran, et al.
Published: (2025)
by: Tong, Ran, et al.
Published: (2025)
Decoding Tourist Perception in Historic Urban Quarters with Multimodal Social Media Data: An AI-Based Framework and Evidence from Shanghai
by: Tan, Kaizhen, et al.
Published: (2025)
by: Tan, Kaizhen, et al.
Published: (2025)
Cross-Camera Distracted Driver Classification through Feature Disentanglement and Contrastive Learning
by: Celona, Luigi, et al.
Published: (2024)
by: Celona, Luigi, et al.
Published: (2024)
Multimodal Political Bias Identification and Neutralization
by: Bernard, Cedric, et al.
Published: (2025)
by: Bernard, Cedric, et al.
Published: (2025)
Are Multimodal LLMs Ready for Clinical Dermatology? A Real-World Evaluation in Dermatology
by: Jiang, Roy, et al.
Published: (2026)
by: Jiang, Roy, et al.
Published: (2026)
Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators
by: Jo, Claire Wonjeong, et al.
Published: (2024)
by: Jo, Claire Wonjeong, et al.
Published: (2024)
Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images
by: Deng, Boyang, et al.
Published: (2025)
by: Deng, Boyang, et al.
Published: (2025)
ECMF: Enhanced Cross-Modal Fusion for Multimodal Emotion Recognition in MER-SEMI Challenge
by: Hu, Juewen, et al.
Published: (2025)
by: Hu, Juewen, et al.
Published: (2025)
Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos
by: Shen, Yixuan, et al.
Published: (2026)
by: Shen, Yixuan, et al.
Published: (2026)
Aligning AI with Public Values: Deliberation and Decision-Making for Governing Multimodal LLMs in Political Video Analysis
by: Sharma, Tanusree, et al.
Published: (2024)
by: Sharma, Tanusree, et al.
Published: (2024)
BuildingView: Constructing Urban Building Exteriors Databases with Street View Imagery and Multimodal Large Language Mode
by: Li, Zongrong, et al.
Published: (2024)
by: Li, Zongrong, et al.
Published: (2024)
Let Androids Dream of Electric Sheep: A Human-Inspired Image Implication Understanding and Reasoning Framework
by: Zhang, Chenhao, et al.
Published: (2025)
by: Zhang, Chenhao, et al.
Published: (2025)
EDU-CIRCUIT-HW: Evaluating Multimodal Large Language Models on Real-World University-Level STEM Student Handwritten Solutions
by: Sun, Weiyu, et al.
Published: (2026)
by: Sun, Weiyu, et al.
Published: (2026)
An AI-Enabled Framework Within Reach for Enhancing Healthcare Sustainability and Fairness
by: Huang, Bin, et al.
Published: (2024)
by: Huang, Bin, et al.
Published: (2024)
Multimodal Learning with Augmentation Techniques for Natural Disaster Assessment
by: Urse, Adrian-Dinu, et al.
Published: (2025)
by: Urse, Adrian-Dinu, et al.
Published: (2025)
Multimodal Modular Chain of Thoughts in Energy Performance Certificate Assessment
by: Peng, Zhen, et al.
Published: (2026)
by: Peng, Zhen, et al.
Published: (2026)
Multimodal Approaches to Fair Image Classification: An Ethical Perspective
by: Hickmon, Javon
Published: (2024)
by: Hickmon, Javon
Published: (2024)
Context-aware Multimodal AI Reveals Hidden Pathways in Five Centuries of Art Evolution
by: Kim, Jin, et al.
Published: (2025)
by: Kim, Jin, et al.
Published: (2025)
Cycle-YOLO: A Efficient and Robust Framework for Pavement Damage Detection
by: Li, Zhengji, et al.
Published: (2024)
by: Li, Zhengji, et al.
Published: (2024)
AI-based Multimodal Biometrics for Detecting Smartphone Distractions: Application to Online Learning
by: Becerra, Alvaro, et al.
Published: (2025)
by: Becerra, Alvaro, et al.
Published: (2025)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
by: Qi, Peng, et al.
Published: (2024)
by: Qi, Peng, et al.
Published: (2024)
Replication in Visual Diffusion Models: A Survey and Outlook
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
Networking Systems for Video Anomaly Detection: A Tutorial and Survey
by: Liu, Jing, et al.
Published: (2024)
by: Liu, Jing, et al.
Published: (2024)
MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation
by: Zubair, Md, et al.
Published: (2025)
by: Zubair, Md, et al.
Published: (2025)
Computer Vision for Multimedia Geolocation in Human Trafficking Investigation: A Systematic Literature Review
by: Bamigbade, Opeyemi, et al.
Published: (2024)
by: Bamigbade, Opeyemi, et al.
Published: (2024)
Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024
by: Chandra, Nuria Alina, et al.
Published: (2025)
by: Chandra, Nuria Alina, et al.
Published: (2025)
ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided Transformer
by: Chen, Hongruixuan, et al.
Published: (2023)
by: Chen, Hongruixuan, et al.
Published: (2023)
Sum of Group Error Differences: A Critical Examination of Bias Evaluation in Biometric Verification and a Dual-Metric Measure
by: Elobaid, Alaa, et al.
Published: (2024)
by: Elobaid, Alaa, et al.
Published: (2024)
BLK-Assist: A Methodological Framework for Artist-Led Co-Creation with Generative AI Models
by: Grimes, Daniel, et al.
Published: (2026)
by: Grimes, Daniel, et al.
Published: (2026)
The changing surface of the world's roads
by: Randhawa, Sukanya, et al.
Published: (2025)
by: Randhawa, Sukanya, et al.
Published: (2025)
Locating Demographic Bias at the Attention-Head Level in CLIP's Vision Encoder
by: Yasser, Alaa, et al.
Published: (2026)
by: Yasser, Alaa, et al.
Published: (2026)
Certified Circuits: Stability Guarantees for Mechanistic Circuits
by: Anani, Alaa, et al.
Published: (2026)
by: Anani, Alaa, et al.
Published: (2026)
Mass Concept Erasure in Diffusion Models with Concept Hierarchy
by: Tu, Jiahang, et al.
Published: (2026)
by: Tu, Jiahang, et al.
Published: (2026)
Aesthetics as Structural Harm: Algorithmic Lookism Across Text-to-Image Generation and Classification
by: Doh, Miriam, et al.
Published: (2026)
by: Doh, Miriam, et al.
Published: (2026)
OrthoEraser: Coupled-Neuron Orthogonal Projection for Concept Erasure
by: Shi, Chuancheng, et al.
Published: (2026)
by: Shi, Chuancheng, et al.
Published: (2026)
Happy Young Women, Grumpy Old Men? Emotion-Driven Demographic Biases in Synthetic Face Generation
by: Wei, Mengting, et al.
Published: (2026)
by: Wei, Mengting, et al.
Published: (2026)
MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning
by: Zhang, Chenhao, et al.
Published: (2026)
by: Zhang, Chenhao, et al.
Published: (2026)
The Weaponization of Computer Vision: Tracing Military-Surveillance Ties through Conference Sponsorship
by: Garcia, Noa, et al.
Published: (2026)
by: Garcia, Noa, et al.
Published: (2026)
Urban Socio-Semantic Segmentation with Vision-Language Reasoning
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
by: Kang, Caixin, et al.
Published: (2026)
by: Kang, Caixin, et al.
Published: (2026)
Similar Items
-
Rainbow Noise: Stress-Testing Multimodal Harmful-Meme Detectors on LGBTQ Content
by: Tong, Ran, et al.
Published: (2025) -
Decoding Tourist Perception in Historic Urban Quarters with Multimodal Social Media Data: An AI-Based Framework and Evidence from Shanghai
by: Tan, Kaizhen, et al.
Published: (2025) -
Cross-Camera Distracted Driver Classification through Feature Disentanglement and Contrastive Learning
by: Celona, Luigi, et al.
Published: (2024) -
Multimodal Political Bias Identification and Neutralization
by: Bernard, Cedric, et al.
Published: (2025) -
Are Multimodal LLMs Ready for Clinical Dermatology? A Real-World Evaluation in Dermatology
by: Jiang, Roy, et al.
Published: (2026)