Under manipulations, are some AI models harder to audit?
Fuente:
arXiv
Guardado en:
| Autores principales: | Godinot, Augustin, Tredan, Gilles, Merrer, Erwan Le, Penzo, Camilla, Taïani, Francois |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Queries, Representation & Detection: The Next 100 Model Fingerprinting Schemes
por: Godinot, Augustin, et al.
Publicado: (2024)
por: Godinot, Augustin, et al.
Publicado: (2024)
Log Probability Tracking of LLM APIs
por: Chauvin, Timothée, et al.
Publicado: (2025)
por: Chauvin, Timothée, et al.
Publicado: (2025)
The 20 questions game to distinguish large language models
por: Richardeau, Gurvan, et al.
Publicado: (2024)
por: Richardeau, Gurvan, et al.
Publicado: (2024)
Token-Efficient Change Detection in LLM APIs
por: Chauvin, Timothée, et al.
Publicado: (2026)
por: Chauvin, Timothée, et al.
Publicado: (2026)
Robust ML Auditing using Prior Knowledge
por: Bourrée, Jade Garcia, et al.
Publicado: (2025)
por: Bourrée, Jade Garcia, et al.
Publicado: (2025)
LLMs Prompted for Graphs: Hallucinations and Generative Capabilities
por: Richardeau, Gurvan, et al.
Publicado: (2024)
por: Richardeau, Gurvan, et al.
Publicado: (2024)
Leveraging Imperfect Sources to Detect Fairwashing in Black-Box Auditing
por: Bourrée, Jade Garcia, et al.
Publicado: (2023)
por: Bourrée, Jade Garcia, et al.
Publicado: (2023)
P2NIA: Privacy-Preserving Non-Iterative Auditing
por: Bourrée, Jade Garcia, et al.
Publicado: (2025)
por: Bourrée, Jade Garcia, et al.
Publicado: (2025)
Chapter Challenges in archiving the personalized web
por: Penzo, Camilla, et al.
Publicado: (2024)
por: Penzo, Camilla, et al.
Publicado: (2024)
Fairness Auditing with Multi-Agent Collaboration
por: de Vos, Martijn, et al.
Publicado: (2024)
por: de Vos, Martijn, et al.
Publicado: (2024)
Overcoming the Challenges of Batch Normalization in Federated Learning
por: Guerraoui, Rachid, et al.
Publicado: (2024)
por: Guerraoui, Rachid, et al.
Publicado: (2024)
Unified Privacy Guarantees for Decentralized Learning via Matrix Factorization
por: Bellet, Aurélien, et al.
Publicado: (2025)
por: Bellet, Aurélien, et al.
Publicado: (2025)
ByzFL: Research Framework for Robust Federated Learning
por: González, Marc, et al.
Publicado: (2025)
por: González, Marc, et al.
Publicado: (2025)
Is nasty noise actually harder than malicious noise?
por: Blanc, Guy, et al.
Publicado: (2025)
por: Blanc, Guy, et al.
Publicado: (2025)
Is RL fine-tuning harder than regression? A PDE learning approach for diffusion models
por: Mou, Wenlong
Publicado: (2025)
por: Mou, Wenlong
Publicado: (2025)
Pragmatic auditing: a pilot-driven approach for auditing Machine Learning systems
por: Benbouzid, Djalel, et al.
Publicado: (2024)
por: Benbouzid, Djalel, et al.
Publicado: (2024)
The Boy Who Survived: Removing Harry Potter from an LLM is harder than reported
por: Shostack, Adam
Publicado: (2024)
por: Shostack, Adam
Publicado: (2024)
Learning Diffusion Priors from Observations by Expectation Maximization
por: Rozet, François, et al.
Publicado: (2024)
por: Rozet, François, et al.
Publicado: (2024)
Adaptive auditing of AI systems with anytime-valid guarantees
por: Zhou, Siyu, et al.
Publicado: (2026)
por: Zhou, Siyu, et al.
Publicado: (2026)
A foundation model of vision, audition, and language for in-silico neuroscience
por: d'Ascoli, Stéphane, et al.
Publicado: (2026)
por: d'Ascoli, Stéphane, et al.
Publicado: (2026)
Towards Minimax Optimality of Model-based Robust Reinforcement Learning
por: Clavier, Pierre, et al.
Publicado: (2023)
por: Clavier, Pierre, et al.
Publicado: (2023)
Minimising changes to audit when updating decision trees
por: Simmons, Anj, et al.
Publicado: (2024)
por: Simmons, Anj, et al.
Publicado: (2024)
Neural Network Verification with PyRAT
por: Lemesle, Augustin, et al.
Publicado: (2024)
por: Lemesle, Augustin, et al.
Publicado: (2024)
Mosaic Learning: A Framework for Decentralized Learning with Model Fragmentation
por: Biswas, Sayan, et al.
Publicado: (2026)
por: Biswas, Sayan, et al.
Publicado: (2026)
Your Neighbors Know: Leveraging Local Neighborhoods for Backdoor Detection in Decentralized Learning
por: Biswas, Sayan, et al.
Publicado: (2026)
por: Biswas, Sayan, et al.
Publicado: (2026)
Why Deep Jacobian Spectra Separate: Depth-Induced Scaling and Singular-Vector Alignment
por: Haas, Nathanaël, et al.
Publicado: (2026)
por: Haas, Nathanaël, et al.
Publicado: (2026)
Low-Cost Privacy-Preserving Decentralized Learning
por: Biswas, Sayan, et al.
Publicado: (2024)
por: Biswas, Sayan, et al.
Publicado: (2024)
Generalizing while preserving monotonicity in comparison-based preference learning models
por: Fageot, Julien, et al.
Publicado: (2025)
por: Fageot, Julien, et al.
Publicado: (2025)
Training-Free Bayesian Filtering with Generative Emulators
por: Savary, Thomas, et al.
Publicado: (2026)
por: Savary, Thomas, et al.
Publicado: (2026)
Training-Free Data Assimilation with GenCast
por: Savary, Thomas, et al.
Publicado: (2025)
por: Savary, Thomas, et al.
Publicado: (2025)
Concentration and excess risk bounds for imbalanced classification with synthetic oversampling
por: Ahmad, Touqeer, et al.
Publicado: (2025)
por: Ahmad, Touqeer, et al.
Publicado: (2025)
Addressing the regulatory gap: moving towards an EU AI audit ecosystem beyond the AI Act by including civil society
por: Hartmann, David, et al.
Publicado: (2024)
por: Hartmann, David, et al.
Publicado: (2024)
On Monotonicity in AI Alignment
por: Bareilles, Gilles, et al.
Publicado: (2025)
por: Bareilles, Gilles, et al.
Publicado: (2025)
Gram: Assessing sabotage propensities via automated alignment auditing
por: Lindner, David, et al.
Publicado: (2026)
por: Lindner, David, et al.
Publicado: (2026)
Increasing Missingness to Reduce Bias: Richardson-SGD with Missing Data
por: Genans, Ferdinand, et al.
Publicado: (2026)
por: Genans, Ferdinand, et al.
Publicado: (2026)
Statistical Properties of the King Wen Sequence: An Anti-Habituation Structure That Does Not Improve Neural Network Training
por: Chan, Augustin
Publicado: (2026)
por: Chan, Augustin
Publicado: (2026)
Adversarial training with restricted data manipulation
por: Benfield, David, et al.
Publicado: (2025)
por: Benfield, David, et al.
Publicado: (2025)
Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation
por: Rozet, François, et al.
Publicado: (2025)
por: Rozet, François, et al.
Publicado: (2025)
Bootstrapping Expectiles in Reinforcement Learning
por: Clavier, Pierre, et al.
Publicado: (2024)
por: Clavier, Pierre, et al.
Publicado: (2024)
parallelcbf: A composable safety-filter and auditability framework for tensor-parallel reinforcement learning
por: Lu, Yijun, et al.
Publicado: (2026)
por: Lu, Yijun, et al.
Publicado: (2026)
Ejemplares similares
-
Queries, Representation & Detection: The Next 100 Model Fingerprinting Schemes
por: Godinot, Augustin, et al.
Publicado: (2024) -
Log Probability Tracking of LLM APIs
por: Chauvin, Timothée, et al.
Publicado: (2025) -
The 20 questions game to distinguish large language models
por: Richardeau, Gurvan, et al.
Publicado: (2024) -
Token-Efficient Change Detection in LLM APIs
por: Chauvin, Timothée, et al.
Publicado: (2026) -
Robust ML Auditing using Prior Knowledge
por: Bourrée, Jade Garcia, et al.
Publicado: (2025)