MLLM-as-a-Judge for Image Safety without Human Labeling
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Zhenting, Hu, Shuming, Zhao, Shiyu, Lin, Xiaowen, Juefei-Xu, Felix, Li, Zhuowei, Han, Ligong, Subramanyam, Harihar, Chen, Li, Chen, Jianfa, Jiang, Nan, Lyu, Lingjuan, Ma, Shiqing, Metaxas, Dimitris N., Jain, Ankit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How to Trace Latent Generative Model Generated Images without Artificial Watermark?
di: Wang, Zhenting, et al.
Pubblicazione: (2024)
di: Wang, Zhenting, et al.
Pubblicazione: (2024)
DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion Models
di: Wang, Zhenting, et al.
Pubblicazione: (2023)
di: Wang, Zhenting, et al.
Pubblicazione: (2023)
Class-RAG: Real-Time Content Moderation with Retrieval Augmented Generation
di: Chen, Jianfa, et al.
Pubblicazione: (2024)
di: Chen, Jianfa, et al.
Pubblicazione: (2024)
LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation
di: Zhou, Yang, et al.
Pubblicazione: (2025)
di: Zhou, Yang, et al.
Pubblicazione: (2025)
Score-Guided Diffusion for 3D Human Recovery
di: Stathopoulos, Anastasis, et al.
Pubblicazione: (2024)
di: Stathopoulos, Anastasis, et al.
Pubblicazione: (2024)
LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
di: Jin, Can, et al.
Pubblicazione: (2025)
di: Jin, Can, et al.
Pubblicazione: (2025)
Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering
di: Chatzoudis, Gerasimos, et al.
Pubblicazione: (2025)
di: Chatzoudis, Gerasimos, et al.
Pubblicazione: (2025)
Evaluating and Mitigating IP Infringement in Visual Generative AI
di: Wang, Zhenting, et al.
Pubblicazione: (2024)
di: Wang, Zhenting, et al.
Pubblicazione: (2024)
Implicit In-context Learning
di: Li, Zhuowei, et al.
Pubblicazione: (2024)
di: Li, Zhuowei, et al.
Pubblicazione: (2024)
Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction
di: Zhao, Shiyu, et al.
Pubblicazione: (2024)
di: Zhao, Shiyu, et al.
Pubblicazione: (2024)
Show and Segment: Universal Medical Image Segmentation via In-Context Learning
di: Gao, Yunhe, et al.
Pubblicazione: (2025)
di: Gao, Yunhe, et al.
Pubblicazione: (2025)
Judge Anything: MLLM as a Judge Across Any Modality
di: Pu, Shu, et al.
Pubblicazione: (2025)
di: Pu, Shu, et al.
Pubblicazione: (2025)
Training Like a Medical Resident: Context-Prior Learning Toward Universal Medical Image Segmentation
di: Gao, Yunhe, et al.
Pubblicazione: (2023)
di: Gao, Yunhe, et al.
Pubblicazione: (2023)
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
di: Li, Zhuowei, et al.
Pubblicazione: (2025)
di: Li, Zhuowei, et al.
Pubblicazione: (2025)
BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models
di: Wang, Yibin, et al.
Pubblicazione: (2024)
di: Wang, Yibin, et al.
Pubblicazione: (2024)
SINE: SINgle Image Editing with Text-to-Image Diffusion Models
di: Zhang, Zhixing, et al.
Pubblicazione: (2022)
di: Zhang, Zhixing, et al.
Pubblicazione: (2022)
CO-SPY: Combining Semantic and Pixel Features to Detect Synthetic Images by AI
di: Cheng, Siyuan, et al.
Pubblicazione: (2025)
di: Cheng, Siyuan, et al.
Pubblicazione: (2025)
Token-Budget-Aware LLM Reasoning
di: Han, Tingxu, et al.
Pubblicazione: (2024)
di: Han, Tingxu, et al.
Pubblicazione: (2024)
Imagine yourself: Tuning-Free Personalized Image Generation
di: He, Zecheng, et al.
Pubblicazione: (2024)
di: He, Zecheng, et al.
Pubblicazione: (2024)
TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
di: Zhang, Tunyu, et al.
Pubblicazione: (2025)
di: Zhang, Tunyu, et al.
Pubblicazione: (2025)
Spectrum-Aware Parameter Efficient Fine-Tuning for Diffusion Models
di: Zhang, Xinxi, et al.
Pubblicazione: (2024)
di: Zhang, Xinxi, et al.
Pubblicazione: (2024)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
di: Chen, Dongping, et al.
Pubblicazione: (2024)
di: Chen, Dongping, et al.
Pubblicazione: (2024)
MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
di: Dao, Quan, et al.
Pubblicazione: (2026)
di: Dao, Quan, et al.
Pubblicazione: (2026)
Beyond Explicit Edges: Robust Reasoning over Noisy and Sparse Knowledge Graphs
di: Gao, Hang, et al.
Pubblicazione: (2026)
di: Gao, Hang, et al.
Pubblicazione: (2026)
COALA: A Practical and Vision-Centric Federated Learning Platform
di: Zhuang, Weiming, et al.
Pubblicazione: (2024)
di: Zhuang, Weiming, et al.
Pubblicazione: (2024)
Seeing Further on the Shoulders of Giants: Knowledge Inheritance for Vision Foundation Models
di: Huang, Jiabo, et al.
Pubblicazione: (2025)
di: Huang, Jiabo, et al.
Pubblicazione: (2025)
Can Large Vision-Language Models Detect Images Copyright Infringement from GenAI?
di: Xu, Qipan, et al.
Pubblicazione: (2025)
di: Xu, Qipan, et al.
Pubblicazione: (2025)
PrefGen: Multimodal Preference Learning for Preference-Conditioned Image Generation
di: Mo, Wenyi, et al.
Pubblicazione: (2025)
di: Mo, Wenyi, et al.
Pubblicazione: (2025)
Few-Step Diffusion Language Models via Trajectory Self-Distillation
di: Zhang, Tunyu, et al.
Pubblicazione: (2026)
di: Zhang, Tunyu, et al.
Pubblicazione: (2026)
VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
di: Waheed, Abdul, et al.
Pubblicazione: (2025)
di: Waheed, Abdul, et al.
Pubblicazione: (2025)
MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance
di: Pi, Renjie, et al.
Pubblicazione: (2024)
di: Pi, Renjie, et al.
Pubblicazione: (2024)
Finding needles in a haystack: A Black-Box Approach to Invisible Watermark Detection
di: Pan, Minzhou, et al.
Pubblicazione: (2024)
di: Pan, Minzhou, et al.
Pubblicazione: (2024)
FedWon: Triumphing Multi-domain Federated Learning Without Normalization
di: Zhuang, Weiming, et al.
Pubblicazione: (2023)
di: Zhuang, Weiming, et al.
Pubblicazione: (2023)
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
di: Chatzoudis, Gerasimos, et al.
Pubblicazione: (2026)
di: Chatzoudis, Gerasimos, et al.
Pubblicazione: (2026)
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
di: Lee, Sua, et al.
Pubblicazione: (2026)
di: Lee, Sua, et al.
Pubblicazione: (2026)
Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
di: Liu, Runzhou, et al.
Pubblicazione: (2026)
di: Liu, Runzhou, et al.
Pubblicazione: (2026)
Rate-My-LoRA: Efficient and Adaptive Federated Model Tuning for Cardiac MRI Segmentation
di: He, Xiaoxiao, et al.
Pubblicazione: (2025)
di: He, Xiaoxiao, et al.
Pubblicazione: (2025)
CopyJudge: Automated Copyright Infringement Identification and Mitigation in Text-to-Image Diffusion Models
di: Liu, Shunchang, et al.
Pubblicazione: (2025)
di: Liu, Shunchang, et al.
Pubblicazione: (2025)
Detecting, Explaining, and Mitigating Memorization in Diffusion Models
di: Wen, Yuxin, et al.
Pubblicazione: (2024)
di: Wen, Yuxin, et al.
Pubblicazione: (2024)
Replay-Free Continual Low-Rank Adaptation with Dynamic Memory
di: Chen, Huancheng, et al.
Pubblicazione: (2024)
di: Chen, Huancheng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
How to Trace Latent Generative Model Generated Images without Artificial Watermark?
di: Wang, Zhenting, et al.
Pubblicazione: (2024) -
DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion Models
di: Wang, Zhenting, et al.
Pubblicazione: (2023) -
Class-RAG: Real-Time Content Moderation with Retrieval Augmented Generation
di: Chen, Jianfa, et al.
Pubblicazione: (2024) -
LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation
di: Zhou, Yang, et al.
Pubblicazione: (2025) -
Score-Guided Diffusion for 3D Human Recovery
di: Stathopoulos, Anastasis, et al.
Pubblicazione: (2024)