Securing External Deeper-than-black-box GPAI Evaluations
Fuente:
arXiv
Guardado en:
| Autores principales: | Tlaie, Alejandro, Farrell, Jimmy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Using AI Alignment Theory to understand the potential pitfalls of regulatory frameworks
por: Tlaie, Alejandro
Publicado: (2024)
por: Tlaie, Alejandro
Publicado: (2024)
Expanding External Access To Frontier AI Models For Dangerous Capability Evaluations
por: Charnock, Jacob, et al.
Publicado: (2026)
por: Charnock, Jacob, et al.
Publicado: (2026)
GPAI Evaluations Standards Taskforce: Towards Effective AI Governance
por: Paskov, Patricia, et al.
Publicado: (2024)
por: Paskov, Patricia, et al.
Publicado: (2024)
A Blueprint for an EU Ecosystem of Secure, Deep and External AI Audits
por: Tlaie, Alejandro
Publicado: (2025)
por: Tlaie, Alejandro
Publicado: (2025)
Preliminary suggestions for rigorous GPAI model evaluations
por: Paskov, Patricia, et al.
Publicado: (2025)
por: Paskov, Patricia, et al.
Publicado: (2025)
Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures
por: Stelling, Lily, et al.
Publicado: (2025)
por: Stelling, Lily, et al.
Publicado: (2025)
Assessing confidence in frontier AI safety cases
por: Barrett, Stephen, et al.
Publicado: (2025)
por: Barrett, Stephen, et al.
Publicado: (2025)
Enabling Responsible, Secure and Sustainable Healthcare AI - A Strategic Framework for Clinical and Operational Impact
por: Joseph, Jimmy
Publicado: (2025)
por: Joseph, Jimmy
Publicado: (2025)
AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models
por: Barrett, Anthony M., et al.
Publicado: (2025)
por: Barrett, Anthony M., et al.
Publicado: (2025)
Exploring and steering the moral compass of Large Language Models
por: Tlaie, Alejandro
Publicado: (2024)
por: Tlaie, Alejandro
Publicado: (2024)
Quality Assessment of Public Summary of Training Content for GPAI models required by AI Act Article 53(1)(d)
por: Blankvoort, Dick A. H., et al.
Publicado: (2026)
por: Blankvoort, Dick A. H., et al.
Publicado: (2026)
A Methodology for Quantitative AI Risk Modeling
por: Murray, Malcolm, et al.
Publicado: (2025)
por: Murray, Malcolm, et al.
Publicado: (2025)
The Role of Risk Modeling in Advanced AI Risk Management
por: Touzet, Chloé, et al.
Publicado: (2025)
por: Touzet, Chloé, et al.
Publicado: (2025)
External Evaluation of Discrimination Mitigation Efforts in Meta's Ad Delivery
por: Imana, Basileal, et al.
Publicado: (2025)
por: Imana, Basileal, et al.
Publicado: (2025)
Federated learning, ethics, and the double black box problem in medical AI
por: Hatherley, Joshua, et al.
Publicado: (2025)
por: Hatherley, Joshua, et al.
Publicado: (2025)
An External Fairness Evaluation of LinkedIn Talent Search
por: Behzad, Tina, et al.
Publicado: (2025)
por: Behzad, Tina, et al.
Publicado: (2025)
Descriptions of women are longer than that of men: An analysis of gender portrayal prompts in Stable Diffusion
por: Asadchy, Yan, et al.
Publicado: (2024)
por: Asadchy, Yan, et al.
Publicado: (2024)
Creating and Evaluating Privacy and Security Micro-Lessons for Elementary School Children
por: Gao, Lan, et al.
Publicado: (2025)
por: Gao, Lan, et al.
Publicado: (2025)
Lessons from External Review of DeepMind's Scheming Inability Safety Case
por: Barrett, Stephen, et al.
Publicado: (2026)
por: Barrett, Stephen, et al.
Publicado: (2026)
Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users
por: Kempermann, Manon, et al.
Publicado: (2025)
por: Kempermann, Manon, et al.
Publicado: (2025)
Toward Quantitative Modeling of Cybersecurity Risks Due to AI Misuse
por: Barrett, Steve, et al.
Publicado: (2025)
por: Barrett, Steve, et al.
Publicado: (2025)
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
por: Knight, Christina Q., et al.
Publicado: (2025)
por: Knight, Christina Q., et al.
Publicado: (2025)
Assessing the Impact of External and Internal Factors on Emergency Department Overcrowding
por: Ahmed, Abdulaziz, et al.
Publicado: (2025)
por: Ahmed, Abdulaziz, et al.
Publicado: (2025)
Small Models Achieve Large Language Model Performance: Evaluating Reasoning-Enabled AI for Secure Child Welfare Research
por: Qi, Zia, et al.
Publicado: (2025)
por: Qi, Zia, et al.
Publicado: (2025)
Purer than pure: how purity reshapes the upstream materiality of the semiconductor industry
por: Roussilhe, Gauthier, et al.
Publicado: (2025)
por: Roussilhe, Gauthier, et al.
Publicado: (2025)
From Replacement to Orchestration: A Socio-Technical Architecture for Agentic AI in Corporate R&D
por: Boussaid, Haithem, et al.
Publicado: (2026)
por: Boussaid, Haithem, et al.
Publicado: (2026)
Evaluating Organization Security: User Stories of European Union NIS2 Directive
por: Seeba, Mari, et al.
Publicado: (2025)
por: Seeba, Mari, et al.
Publicado: (2025)
Secure On-Premise Deployment of Open-Weights Large Language Models in Radiology: An Isolation-First Architecture with Prospective Pilot Evaluation
por: Nowak, Sebastian, et al.
Publicado: (2026)
por: Nowak, Sebastian, et al.
Publicado: (2026)
Intimacy as Service, Harm as Externality: Critical Perspectives on AI Companion Platform Accountability
por: Eom, Dayeon, et al.
Publicado: (2026)
por: Eom, Dayeon, et al.
Publicado: (2026)
Rethinking Optimization: A Systems-Based Approach to Social Externalities
por: Nokhiz, Pegah, et al.
Publicado: (2025)
por: Nokhiz, Pegah, et al.
Publicado: (2025)
Exploring the Adversarial Robustness of Face Forgery Detection with Decision-based Black-box Attacks
por: Chen, Zhaoyu, et al.
Publicado: (2023)
por: Chen, Zhaoyu, et al.
Publicado: (2023)
Online search is more likely to lead students to validate true news than to refute false ones
por: Bouleimen, Azza, et al.
Publicado: (2023)
por: Bouleimen, Azza, et al.
Publicado: (2023)
European Football Player Valuation: Integrating Financial Models and Network Theory
por: Cohen, Albert, et al.
Publicado: (2023)
por: Cohen, Albert, et al.
Publicado: (2023)
Assurance of Frontier AI Built for National Security
por: Pistillo, Matteo, et al.
Publicado: (2025)
por: Pistillo, Matteo, et al.
Publicado: (2025)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
por: Badawi, Abeer, et al.
Publicado: (2025)
por: Badawi, Abeer, et al.
Publicado: (2025)
More than Carbon: Cradle-to-Grave environmental impacts of GenAI training on the Nvidia A100 GPU
por: Falk, Sophia, et al.
Publicado: (2025)
por: Falk, Sophia, et al.
Publicado: (2025)
A New Exploration into Chinese Characters: from Simplification to Deeper Understanding
por: Gong, Wen G.
Publicado: (2025)
por: Gong, Wen G.
Publicado: (2025)
An Empirical Analysis on the Use and Reporting of National Security Letters
por: Bellon, Alex, et al.
Publicado: (2024)
por: Bellon, Alex, et al.
Publicado: (2024)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
por: Chiu, Yu Ying, et al.
Publicado: (2025)
por: Chiu, Yu Ying, et al.
Publicado: (2025)
Transparency, Security, and Workplace Training & Awareness in the Age of Generative AI
por: Vaishnav, Lakshika, et al.
Publicado: (2024)
por: Vaishnav, Lakshika, et al.
Publicado: (2024)
Ejemplares similares
-
Using AI Alignment Theory to understand the potential pitfalls of regulatory frameworks
por: Tlaie, Alejandro
Publicado: (2024) -
Expanding External Access To Frontier AI Models For Dangerous Capability Evaluations
por: Charnock, Jacob, et al.
Publicado: (2026) -
GPAI Evaluations Standards Taskforce: Towards Effective AI Governance
por: Paskov, Patricia, et al.
Publicado: (2024) -
A Blueprint for an EU Ecosystem of Secure, Deep and External AI Audits
por: Tlaie, Alejandro
Publicado: (2025) -
Preliminary suggestions for rigorous GPAI model evaluations
por: Paskov, Patricia, et al.
Publicado: (2025)