NAU-QMUL: Utilizing BERT and CLIP for Multi-modal AI-Generated Image Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Xiaoyu, Zubiaga, Arkaitz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
FLEX-CLIP: Feature-Level GEneration Network Enhanced CLIP for X-shot Cross-modal Retrieval
von: Xie, Jingyou, et al.
Veröffentlicht: (2024)
von: Xie, Jingyou, et al.
Veröffentlicht: (2024)
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation
von: Cai, Shihao, et al.
Veröffentlicht: (2024)
von: Cai, Shihao, et al.
Veröffentlicht: (2024)
GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery
von: Wang, Enguang, et al.
Veröffentlicht: (2024)
von: Wang, Enguang, et al.
Veröffentlicht: (2024)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
Instruct-Imagen: Image Generation with Multi-modal Instruction
von: Hu, Hexiang, et al.
Veröffentlicht: (2024)
von: Hu, Hexiang, et al.
Veröffentlicht: (2024)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)
Language Augmentation in CLIP for Improved Anatomy Detection on Multi-modal Medical Images
von: Kakkar, Mansi, et al.
Veröffentlicht: (2024)
von: Kakkar, Mansi, et al.
Veröffentlicht: (2024)
LowCLIP: Adapting the CLIP Model Architecture for Low-Resource Languages in Multimodal Image Retrieval Task
von: Asgarov, Ali, et al.
Veröffentlicht: (2024)
von: Asgarov, Ali, et al.
Veröffentlicht: (2024)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models
von: Villa, Andrés, et al.
Veröffentlicht: (2023)
von: Villa, Andrés, et al.
Veröffentlicht: (2023)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
Seeing the Poem: Image-Semantic Detection of AI-Generated Modern Chinese Poetry with MLLMs
von: Wang, Shanshan, et al.
Veröffentlicht: (2026)
von: Wang, Shanshan, et al.
Veröffentlicht: (2026)
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
von: Sun, Hao, et al.
Veröffentlicht: (2026)
von: Sun, Hao, et al.
Veröffentlicht: (2026)
RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025)
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025)
Advancing Large Multi-modal Models with Explicit Chain-of-Reasoning and Visual Question Generation
von: Uehara, Kohei, et al.
Veröffentlicht: (2024)
von: Uehara, Kohei, et al.
Veröffentlicht: (2024)
Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family
von: Chew, Oscar, et al.
Veröffentlicht: (2026)
von: Chew, Oscar, et al.
Veröffentlicht: (2026)
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
UniAIDet: A Unified and Universal Benchmark for AI-Generated Image Content Detection and Localization
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
Demystifying CLIP Data
von: Xu, Hu, et al.
Veröffentlicht: (2023)
von: Xu, Hu, et al.
Veröffentlicht: (2023)
MedCLIP-SAMv2: Towards Universal Text-Driven Medical Image Segmentation
von: Koleilat, Taha, et al.
Veröffentlicht: (2024)
von: Koleilat, Taha, et al.
Veröffentlicht: (2024)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
Detecting AI-Generated Images via CLIP
von: Moskowitz, A. G., et al.
Veröffentlicht: (2024)
von: Moskowitz, A. G., et al.
Veröffentlicht: (2024)
SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster Prediction
von: Gueuwou, Shester, et al.
Veröffentlicht: (2024)
von: Gueuwou, Shester, et al.
Veröffentlicht: (2024)
Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
von: Madaan, Divyam, et al.
Veröffentlicht: (2025)
von: Madaan, Divyam, et al.
Veröffentlicht: (2025)
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
von: Fu, Jinlan, et al.
Veröffentlicht: (2025)
von: Fu, Jinlan, et al.
Veröffentlicht: (2025)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2023)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2023)
Mitigating the Modality Gap: Few-Shot Out-of-Distribution Detection with Multi-modal Prototypes and Image Bias Estimation
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
CLIP Multi-modal Hashing for Multimedia Retrieval
von: Zhu, Jian, et al.
Veröffentlicht: (2024)
von: Zhu, Jian, et al.
Veröffentlicht: (2024)
Human vs. AI: A Novel Benchmark and a Comparative Study on the Detection of Generated Images and the Impact of Prompts
von: Moeßner, Philipp, et al.
Veröffentlicht: (2024)
von: Moeßner, Philipp, et al.
Veröffentlicht: (2024)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
Can Multi-modal (reasoning) LLMs detect document manipulation?
von: Liang, Zisheng, et al.
Veröffentlicht: (2025)
von: Liang, Zisheng, et al.
Veröffentlicht: (2025)
MMBench: Is Your Multi-modal Model an All-around Player?
von: Liu, Yuan, et al.
Veröffentlicht: (2023)
von: Liu, Yuan, et al.
Veröffentlicht: (2023)
GroundingGPT:Language Enhanced Multi-modal Grounding Model
von: Li, Zhaowei, et al.
Veröffentlicht: (2024)
von: Li, Zhaowei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection
von: Chen, Junjie, et al.
Veröffentlicht: (2024) -
FLEX-CLIP: Feature-Level GEneration Network Enhanced CLIP for X-shot Cross-modal Retrieval
von: Xie, Jingyou, et al.
Veröffentlicht: (2024) -
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
von: Zhang, Yichi, et al.
Veröffentlicht: (2025) -
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation
von: Cai, Shihao, et al.
Veröffentlicht: (2024) -
GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery
von: Wang, Enguang, et al.
Veröffentlicht: (2024)