Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security
Fuente:
arXiv
Saved in:
| Main Authors: | Naseh, Ali, Suri, Anshuman, Peng, Yuefeng, Chaudhari, Harsh, Oprea, Alina, Houmansadr, Amir |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Identifying Models Behind Text-to-Image Leaderboards
by: Naseh, Ali, et al.
Published: (2026)
by: Naseh, Ali, et al.
Published: (2026)
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
by: Suri, Anshuman, et al.
Published: (2025)
by: Suri, Anshuman, et al.
Published: (2025)
Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Diffence: Fencing Membership Privacy With Diffusion Models
by: Peng, Yuefeng, et al.
Published: (2023)
by: Peng, Yuefeng, et al.
Published: (2023)
SAGA: A Security Architecture for Governing AI Agentic Systems
by: Syros, Georgios, et al.
Published: (2025)
by: Syros, Georgios, et al.
Published: (2025)
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
by: Lin, Shuyi, et al.
Published: (2025)
by: Lin, Shuyi, et al.
Published: (2025)
Backdooring Bias ($B^2$) into Stable Diffusion Models
by: Naseh, Ali, et al.
Published: (2024)
by: Naseh, Ali, et al.
Published: (2024)
OSLO: One-Shot Label-Only Membership Inference Attacks
by: Peng, Yuefeng, et al.
Published: (2024)
by: Peng, Yuefeng, et al.
Published: (2024)
Cascading Adversarial Bias from Injection to Distillation in Language Models
by: Chaudhari, Harsh, et al.
Published: (2025)
by: Chaudhari, Harsh, et al.
Published: (2025)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
by: Chaudhari, Harsh, et al.
Published: (2024)
by: Chaudhari, Harsh, et al.
Published: (2024)
Throttling Web Agents Using Reasoning Gates
by: Kumar, Abhinav, et al.
Published: (2025)
by: Kumar, Abhinav, et al.
Published: (2025)
UTrace: Poisoning Forensics for Private Collaborative Learning
by: Rose, Evan, et al.
Published: (2024)
by: Rose, Evan, et al.
Published: (2024)
Reconstruction of Personally Identifiable Information from Supervised Finetuned Models
by: Furukawa, Sae, et al.
Published: (2026)
by: Furukawa, Sae, et al.
Published: (2026)
OverThink: Slowdown Attacks on Reasoning LLMs
by: Kumar, Abhinav, et al.
Published: (2025)
by: Kumar, Abhinav, et al.
Published: (2025)
DROP: Poison Dilution via Knowledge Distillation for Federated Learning
by: Syros, Georgios, et al.
Published: (2025)
by: Syros, Georgios, et al.
Published: (2025)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026)
by: Chaudhari, Harsh, et al.
Published: (2026)
Toward a Principled Framework for Agent Safety Measurement
by: Lin, Shuyi, et al.
Published: (2026)
by: Lin, Shuyi, et al.
Published: (2026)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
RAIFLE: Reconstruction Attacks on Interaction-based Federated Learning with Adversarial Data Manipulation
by: Pham, Dzung, et al.
Published: (2023)
by: Pham, Dzung, et al.
Published: (2023)
Fake or Compromised? Making Sense of Malicious Clients in Federated Learning
by: Mozaffari, Hamid, et al.
Published: (2024)
by: Mozaffari, Hamid, et al.
Published: (2024)
Adversarial Inception Backdoor Attacks against Reinforcement Learning
by: Rathbun, Ethan, et al.
Published: (2024)
by: Rathbun, Ethan, et al.
Published: (2024)
SleeperNets: Universal Backdoor Poisoning Attacks Against Reinforcement Learning Agents
by: Rathbun, Ethan, et al.
Published: (2024)
by: Rathbun, Ethan, et al.
Published: (2024)
TMI! Finetuned Models Leak Private Information from their Pretraining Data
by: Abascal, John, et al.
Published: (2023)
by: Abascal, John, et al.
Published: (2023)
ACE: A Security Architecture for LLM-Integrated App Systems
by: Li, Evan, et al.
Published: (2025)
by: Li, Evan, et al.
Published: (2025)
Do Parameters Reveal More than Loss for Membership Inference?
by: Suri, Anshuman, et al.
Published: (2024)
by: Suri, Anshuman, et al.
Published: (2024)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
by: Naseh, Ali, et al.
Published: (2024)
by: Naseh, Ali, et al.
Published: (2024)
Membership Inference Attacks on Vision-Language-Action Models
by: Peng, Yuefeng, et al.
Published: (2026)
by: Peng, Yuefeng, et al.
Published: (2026)
Dissecting Distribution Inference
by: Suri, Anshuman, et al.
Published: (2022)
by: Suri, Anshuman, et al.
Published: (2022)
PoolFlip: A Multi-Agent Reinforcement Learning Security Environment for Cyber Defense
by: Cadet, Xavier, et al.
Published: (2025)
by: Cadet, Xavier, et al.
Published: (2025)
Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors
by: Sun, Luze, et al.
Published: (2026)
by: Sun, Luze, et al.
Published: (2026)
MAGIQ: A Post-Quantum Multi-Agentic AI Governance System with Provable Security
by: Avizheh, Sepideh, et al.
Published: (2026)
by: Avizheh, Sepideh, et al.
Published: (2026)
Backdoor Attacks in Peer-to-Peer Federated Learning
by: Syros, Georgios, et al.
Published: (2023)
by: Syros, Georgios, et al.
Published: (2023)
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
by: Peng, Yuefeng, et al.
Published: (2024)
by: Peng, Yuefeng, et al.
Published: (2024)
Beware Untrusted Simulators -- Reward-Free Backdoor Attacks in Reinforcement Learning
by: Rathbun, Ethan, et al.
Published: (2026)
by: Rathbun, Ethan, et al.
Published: (2026)
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
by: Huang, Yangsibo, et al.
Published: (2025)
by: Huang, Yangsibo, et al.
Published: (2025)
Black-Box Privacy Attacks on Shared Representations in Multitask Learning
by: Abascal, John, et al.
Published: (2025)
by: Abascal, John, et al.
Published: (2025)
Quantitative Resilience Modeling for Autonomous Cyber Defense
by: Cadet, Xavier, et al.
Published: (2025)
by: Cadet, Xavier, et al.
Published: (2025)
Model-agnostic clean-label backdoor mitigation in cybersecurity environments
by: Severi, Giorgio, et al.
Published: (2024)
by: Severi, Giorgio, et al.
Published: (2024)
Attacks and Mitigations for Distributed Governance of Agentic AI under Byzantine Adversaries
by: Laws, Matthew D., et al.
Published: (2026)
by: Laws, Matthew D., et al.
Published: (2026)
Similar Items
-
Identifying Models Behind Text-to-Image Leaderboards
by: Naseh, Ali, et al.
Published: (2026) -
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
by: Suri, Anshuman, et al.
Published: (2025) -
Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation
by: Naseh, Ali, et al.
Published: (2025) -
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
by: Naseh, Ali, et al.
Published: (2025) -
Diffence: Fencing Membership Privacy With Diffusion Models
by: Peng, Yuefeng, et al.
Published: (2023)