Enregistré dans:
| Auteurs principaux: | Bucknall, Ben, Siddiqui, Saad, Thurnherr, Lara, McGurk, Conor, Harack, Ben, Reuel, Anka, Paskov, Patricia, Mahoney, Casey, Mindermann, Sören, Singer, Scott, Hiremath, Vinay, Segerie, Charbel-Raphaël, Delaney, Oscar, Abate, Alessandro, Barez, Fazl, Cohen, Michael K., Torr, Philip, Huszár, Ferenc, Calinescu, Anisoara, Jones, Gabriel Davis, Bengio, Yoshua, Trager, Robert |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2504.12914 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Position Paper: Technical Research and Talent is Needed for Effective AI Governance
par: Reuel, Anka, et autres
Publié: (2024)
par: Reuel, Anka, et autres
Publié: (2024)
Communicating the Value of Cartoon Art across University Classrooms: Experiences from the Ohio State University Billy Ireland Cartoon Library and Museum
par: McGurk, Caitlin
Publié: (2016)
par: McGurk, Caitlin
Publié: (2016)
Human-Provenance Verification should be Treated as Labor Infrastructure in AI-Saturated Markets
par: McGurk, Erin, et autres
Publié: (2026)
par: McGurk, Erin, et autres
Publié: (2026)
Political Uncertainty and Credit Risk: The Role of Event Markets in Forecasting Ukraine's Sovereign Spreads
par: Mary Becker, et autres
Publié: (2026)
par: Mary Becker, et autres
Publié: (2026)
Position: Ensuring mutual privacy is necessary for effective external evaluation of proprietary AI systems
par: Bucknall, Ben, et autres
Publié: (2025)
par: Bucknall, Ben, et autres
Publié: (2025)
Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods
par: Grey, Markov, et autres
Publié: (2025)
par: Grey, Markov, et autres
Publié: (2025)
The AI Risk Spectrum: From Dangerous Capabilities to Existential Threats
par: Grey, Markov, et autres
Publié: (2025)
par: Grey, Markov, et autres
Publié: (2025)
Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
par: Lamparth, Max, et autres
Publié: (2023)
par: Lamparth, Max, et autres
Publié: (2023)
Fairness in Reinforcement Learning: A Survey
par: Reuel, Anka, et autres
Publié: (2024)
par: Reuel, Anka, et autres
Publié: (2024)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
par: Chaudhary, Maheep, et autres
Publié: (2025)
par: Chaudhary, Maheep, et autres
Publié: (2025)
Understanding Addition in Transformers
par: Quirke, Philip, et autres
Publié: (2023)
par: Quirke, Philip, et autres
Publié: (2023)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
par: Fu, Tingchen, et autres
Publié: (2025)
par: Fu, Tingchen, et autres
Publié: (2025)
Partisan Sentiment and Returns From Online Political Betting Markets in the 2020 US Presidential Election
par: Mary Becker, et autres
Publié: (2025)
par: Mary Becker, et autres
Publié: (2025)
Generative AI Needs Adaptive Governance
par: Reuel, Anka, et autres
Publié: (2024)
par: Reuel, Anka, et autres
Publié: (2024)
Query Circuits: Explaining How Language Models Answer User Prompts
par: Wu, Tung-Yu, et autres
Publié: (2025)
par: Wu, Tung-Yu, et autres
Publié: (2025)
The bitter lesson of misuse detection
par: Mariaccia, Hadrien, et autres
Publié: (2025)
par: Mariaccia, Hadrien, et autres
Publié: (2025)
Open Problems in Machine Unlearning for AI Safety
par: Barez, Fazl, et autres
Publié: (2025)
par: Barez, Fazl, et autres
Publié: (2025)
Rethinking AI Cultural Alignment
par: Bravansky, Michal, et autres
Publié: (2025)
par: Bravansky, Michal, et autres
Publié: (2025)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
par: Lan, Michael, et autres
Publié: (2023)
par: Lan, Michael, et autres
Publié: (2023)
Understanding Addition and Subtraction in Transformers
par: Quirke, Philip, et autres
Publié: (2024)
par: Quirke, Philip, et autres
Publié: (2024)
Out-of-distribution Tests Reveal Compositionality in Chess Transformers
par: Mészáros, Anna, et autres
Publié: (2025)
par: Mészáros, Anna, et autres
Publié: (2025)
Token Taxes: mitigating AGI's economic risks
par: Irwin, Lucas, et autres
Publié: (2026)
par: Irwin, Lucas, et autres
Publié: (2026)
Large Language Models Relearn Removed Concepts
par: Lo, Michelle, et autres
Publié: (2024)
par: Lo, Michelle, et autres
Publié: (2024)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
par: Gupta, Aman, et autres
Publié: (2025)
par: Gupta, Aman, et autres
Publié: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
par: Neo, Clement, et autres
Publié: (2024)
par: Neo, Clement, et autres
Publié: (2024)
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
par: Dorn, Diego, et autres
Publié: (2024)
par: Dorn, Diego, et autres
Publié: (2024)
Machine learning and information theory concepts towards an AI Mathematician
par: Bengio, Yoshua, et autres
Publié: (2024)
par: Bengio, Yoshua, et autres
Publié: (2024)
Audit Cards: Contextualizing AI Evaluations
par: Staufer, Leon, et autres
Publié: (2025)
par: Staufer, Leon, et autres
Publié: (2025)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
par: Marks, Luke, et autres
Publié: (2024)
par: Marks, Luke, et autres
Publié: (2024)
Do Sparse Autoencoders Generalize? A Case Study of Answerability
par: Heindrich, Lovis, et autres
Publié: (2025)
par: Heindrich, Lovis, et autres
Publié: (2025)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
par: Schrodi, Simon, et autres
Publié: (2025)
par: Schrodi, Simon, et autres
Publié: (2025)
Visualizing Neural Network Imagination
par: Wichers, Nevan, et autres
Publié: (2024)
par: Wichers, Nevan, et autres
Publié: (2024)
Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies
par: Brundage, Miles, et autres
Publié: (2026)
par: Brundage, Miles, et autres
Publié: (2026)
Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations
par: Wei, Kevin L., et autres
Publié: (2025)
par: Wei, Kevin L., et autres
Publié: (2025)
Learning to Manage Investment Portfolios beyond Simple Utility Functions
par: Scholl, Maarten P., et autres
Publié: (2025)
par: Scholl, Maarten P., et autres
Publié: (2025)
SAGE: Scalable Ground Truth Evaluations for Large Sparse Autoencoders
par: Venhoff, Constantin, et autres
Publié: (2024)
par: Venhoff, Constantin, et autres
Publié: (2024)
Painting the market: generative diffusion models for financial limit order book simulation and forecasting
par: Backhouse, Alfred, et autres
Publié: (2025)
par: Backhouse, Alfred, et autres
Publié: (2025)
A multi-objective combinatorial optimisation framework for large scale hierarchical population synthesis
par: Mahmood, Imran, et autres
Publié: (2024)
par: Mahmood, Imran, et autres
Publié: (2024)
Position Paper: Model Access should be a Key Concern in AI Governance
par: Kembery, Edward, et autres
Publié: (2024)
par: Kembery, Edward, et autres
Publié: (2024)
Baking Symmetry into GFlowNets
par: Ma, George, et autres
Publié: (2024)
par: Ma, George, et autres
Publié: (2024)
Documents similaires
-
Position Paper: Technical Research and Talent is Needed for Effective AI Governance
par: Reuel, Anka, et autres
Publié: (2024) -
Communicating the Value of Cartoon Art across University Classrooms: Experiences from the Ohio State University Billy Ireland Cartoon Library and Museum
par: McGurk, Caitlin
Publié: (2016) -
Human-Provenance Verification should be Treated as Labor Infrastructure in AI-Saturated Markets
par: McGurk, Erin, et autres
Publié: (2026) -
Political Uncertainty and Credit Risk: The Role of Event Markets in Forecasting Ukraine's Sovereign Spreads
par: Mary Becker, et autres
Publié: (2026) -
Position: Ensuring mutual privacy is necessary for effective external evaluation of proprietary AI systems
par: Bucknall, Ben, et autres
Publié: (2025)