DF-LLaVA: Unlocking MLLMs for Synthetic Image Detection via Knowledge Injection and Conflict-Driven Self-Reflection
Fuente:
arXiv
Salvato in:
| Autori principali: | Shen, Zhuokang, Zhang, Kaisen, Jia, Bohan, Jia, Heming, Fang, Yuan, Yu, Zhou, Lin, Shaohui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?
di: Fang, Kechen, et al.
Pubblicazione: (2026)
di: Fang, Kechen, et al.
Pubblicazione: (2026)
LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs
di: Chen, Shaoxiang, et al.
Pubblicazione: (2024)
di: Chen, Shaoxiang, et al.
Pubblicazione: (2024)
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
di: Yan, Dawei, et al.
Pubblicazione: (2024)
di: Yan, Dawei, et al.
Pubblicazione: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding
di: Sun, Boyuan, et al.
Pubblicazione: (2025)
di: Sun, Boyuan, et al.
Pubblicazione: (2025)
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
di: Lou, Haoran, et al.
Pubblicazione: (2025)
di: Lou, Haoran, et al.
Pubblicazione: (2025)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
di: Zeer, Ahmed, et al.
Pubblicazione: (2024)
di: Zeer, Ahmed, et al.
Pubblicazione: (2024)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
di: Zhang, Yuanhan, et al.
Pubblicazione: (2024)
di: Zhang, Yuanhan, et al.
Pubblicazione: (2024)
LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
di: Sun, Shichu, et al.
Pubblicazione: (2025)
di: Sun, Shichu, et al.
Pubblicazione: (2025)
Enhance Image-to-Image Generation with LLaVA-generated Prompts
di: Ding, Zhicheng, et al.
Pubblicazione: (2024)
di: Ding, Zhicheng, et al.
Pubblicazione: (2024)
Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification
di: Huang, Wenxuan, et al.
Pubblicazione: (2024)
di: Huang, Wenxuan, et al.
Pubblicazione: (2024)
LLaVA-RadZ: Can Multimodal Large Language Models Effectively Tackle Zero-shot Radiology Recognition?
di: Li, Bangyan, et al.
Pubblicazione: (2025)
di: Li, Bangyan, et al.
Pubblicazione: (2025)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
di: Xu, Lin, et al.
Pubblicazione: (2024)
di: Xu, Lin, et al.
Pubblicazione: (2024)
SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
di: Sun, Guohao, et al.
Pubblicazione: (2024)
di: Sun, Guohao, et al.
Pubblicazione: (2024)
LLaVA-c: Continual Improved Visual Instruction Tuning
di: Liu, Wenzhuo, et al.
Pubblicazione: (2025)
di: Liu, Wenzhuo, et al.
Pubblicazione: (2025)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
di: Zhang, Shaolei, et al.
Pubblicazione: (2025)
di: Zhang, Shaolei, et al.
Pubblicazione: (2025)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
di: Lin, Bin, et al.
Pubblicazione: (2023)
di: Lin, Bin, et al.
Pubblicazione: (2023)
WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image
di: Liang, Yuci, et al.
Pubblicazione: (2024)
di: Liang, Yuci, et al.
Pubblicazione: (2024)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
di: Yuan, Haobo, et al.
Pubblicazione: (2025)
di: Yuan, Haobo, et al.
Pubblicazione: (2025)
LLaVA-Critic: Learning to Evaluate Multimodal Models
di: Xiong, Tianyi, et al.
Pubblicazione: (2024)
di: Xiong, Tianyi, et al.
Pubblicazione: (2024)
LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
di: Xu, Ruyi, et al.
Pubblicazione: (2024)
di: Xu, Ruyi, et al.
Pubblicazione: (2024)
LLaVAC: Fine-tuning LLaVA as a Multimodal Sentiment Classifier
di: Chay-intr, T., et al.
Pubblicazione: (2025)
di: Chay-intr, T., et al.
Pubblicazione: (2025)
Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages
di: Andersland, Michael
Pubblicazione: (2024)
di: Andersland, Michael
Pubblicazione: (2024)
PA-LLaVA: A Large Language-Vision Assistant for Human Pathology Image Understanding
di: Dai, Dawei, et al.
Pubblicazione: (2024)
di: Dai, Dawei, et al.
Pubblicazione: (2024)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
di: Lin, Bin, et al.
Pubblicazione: (2024)
di: Lin, Bin, et al.
Pubblicazione: (2024)
LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information
di: Wang, Ke, et al.
Pubblicazione: (2024)
di: Wang, Ke, et al.
Pubblicazione: (2024)
Why do LLaVA Vision-Language Models Reply to Images in English?
di: Hinck, Musashi, et al.
Pubblicazione: (2024)
di: Hinck, Musashi, et al.
Pubblicazione: (2024)
LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models
di: Gkalelis, Nikolaos, et al.
Pubblicazione: (2026)
di: Gkalelis, Nikolaos, et al.
Pubblicazione: (2026)
Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases
di: Wang, Liqiong, et al.
Pubblicazione: (2024)
di: Wang, Liqiong, et al.
Pubblicazione: (2024)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
di: Zhang, Tao, et al.
Pubblicazione: (2024)
di: Zhang, Tao, et al.
Pubblicazione: (2024)
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
di: Wei, Jingxuan, et al.
Pubblicazione: (2025)
di: Wei, Jingxuan, et al.
Pubblicazione: (2025)
Can Sound Replace Vision in LLaVA With Token Substitution?
di: Vosoughi, Ali, et al.
Pubblicazione: (2025)
di: Vosoughi, Ali, et al.
Pubblicazione: (2025)
LLaVA-OneVision: Easy Visual Task Transfer
di: Li, Bo, et al.
Pubblicazione: (2024)
di: Li, Bo, et al.
Pubblicazione: (2024)
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
di: Xu, Guowei, et al.
Pubblicazione: (2024)
di: Xu, Guowei, et al.
Pubblicazione: (2024)
TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations
di: Gao, Mingze, et al.
Pubblicazione: (2024)
di: Gao, Mingze, et al.
Pubblicazione: (2024)
LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs
di: Shen, Leqi, et al.
Pubblicazione: (2025)
di: Shen, Leqi, et al.
Pubblicazione: (2025)
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
di: Caffagni, Davide, et al.
Pubblicazione: (2024)
di: Caffagni, Davide, et al.
Pubblicazione: (2024)
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
di: Liang, Han, et al.
Pubblicazione: (2024)
di: Liang, Han, et al.
Pubblicazione: (2024)
LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound
di: Guo, Xuechen, et al.
Pubblicazione: (2024)
di: Guo, Xuechen, et al.
Pubblicazione: (2024)
LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models
di: Zhang, Ruiyi, et al.
Pubblicazione: (2024)
di: Zhang, Ruiyi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?
di: Fang, Kechen, et al.
Pubblicazione: (2026) -
LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs
di: Chen, Shaoxiang, et al.
Pubblicazione: (2024) -
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
di: Yan, Dawei, et al.
Pubblicazione: (2024) -
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
di: Shu, Fangxun, et al.
Pubblicazione: (2024) -
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding
di: Sun, Boyuan, et al.
Pubblicazione: (2025)