The Responsible Foundation Model Development Cheatsheet: A Review of Tools & Resources
Fuente:
arXiv
Saved in:
| Main Authors: | Longpre, Shayne, Biderman, Stella, Albalak, Alon, Schoelkopf, Hailey, McDuff, Daniel, Kapoor, Sayash, Klyman, Kevin, Lo, Kyle, Ilharco, Gabriel, San, Nay, Rauh, Maribeth, Skowron, Aviya, Vidgen, Bertie, Weidinger, Laura, Narayanan, Arvind, Sanh, Victor, Adelani, David, Liang, Percy, Bommasani, Rishi, Henderson, Peter, Luccioni, Sasha, Jernite, Yacine, Soldaini, Luca |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The 2024 Foundation Model Transparency Index
by: Bommasani, Rishi, et al.
Published: (2024)
by: Bommasani, Rishi, et al.
Published: (2024)
Beyond Release: Access Considerations for Generative AI Systems
by: Solaiman, Irene, et al.
Published: (2025)
by: Solaiman, Irene, et al.
Published: (2025)
Foundation Model Transparency Reports
by: Bommasani, Rishi, et al.
Published: (2024)
by: Bommasani, Rishi, et al.
Published: (2024)
The 2025 Foundation Model Transparency Index
by: Wan, Alexander, et al.
Published: (2025)
by: Wan, Alexander, et al.
Published: (2025)
Power Hungry Processing: Watts Driving the Cost of AI Deployment?
by: Luccioni, Alexandra Sasha, et al.
Published: (2023)
by: Luccioni, Alexandra Sasha, et al.
Published: (2023)
On the Societal Impact of Open Foundation Models
by: Kapoor, Sayash, et al.
Published: (2024)
by: Kapoor, Sayash, et al.
Published: (2024)
Future and AI-Ready Data Strategies: Response to DOC RFI on AI and Open Government Data Assets
by: Oderinwale, Hamidah, et al.
Published: (2024)
by: Oderinwale, Hamidah, et al.
Published: (2024)
Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
by: Longpre, Shayne, et al.
Published: (2025)
by: Longpre, Shayne, et al.
Published: (2025)
CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models
by: Pistilli, Giada, et al.
Published: (2024)
by: Pistilli, Giada, et al.
Published: (2024)
A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety
by: François, Camille, et al.
Published: (2025)
by: François, Camille, et al.
Published: (2025)
A Safe Harbor for AI Evaluation and Red Teaming
by: Longpre, Shayne, et al.
Published: (2024)
by: Longpre, Shayne, et al.
Published: (2024)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
Classification is a RAG problem: A case study on hate speech detection
by: Willats, Richard, et al.
Published: (2025)
by: Willats, Richard, et al.
Published: (2025)
Do AI Companies Make Good on Voluntary Commitments to the White House?
by: Wang, Jennifer, et al.
Published: (2025)
by: Wang, Jennifer, et al.
Published: (2025)
INTIMA: A Benchmark for Human-AI Companionship Behavior
by: Kaffee, Lucie-Aimée, et al.
Published: (2025)
by: Kaffee, Lucie-Aimée, et al.
Published: (2025)
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
A Systematic Review of NeurIPS Dataset Management Practices
by: Wu, Yiwei, et al.
Published: (2024)
by: Wu, Yiwei, et al.
Published: (2024)
New Tools are Needed for Tracking Adherence to AI Model Behavioral Use Clauses
by: McDuff, Daniel, et al.
Published: (2025)
by: McDuff, Daniel, et al.
Published: (2025)
Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory
by: Suzgun, Mirac, et al.
Published: (2025)
by: Suzgun, Mirac, et al.
Published: (2025)
Why human-AI relationships need socioaffective alignment
by: Kirk, Hannah Rose, et al.
Published: (2025)
by: Kirk, Hannah Rose, et al.
Published: (2025)
AI-Powered Autonomous Weapons Risk Geopolitical Instability and Threaten AI Research
by: Simmons-Edler, Riley, et al.
Published: (2024)
by: Simmons-Edler, Riley, et al.
Published: (2024)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
by: Röttger, Paul, et al.
Published: (2023)
by: Röttger, Paul, et al.
Published: (2023)
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
by: Styles, Olly, et al.
Published: (2024)
by: Styles, Olly, et al.
Published: (2024)
Suppressing Pink Elephants with Direct Principle Feedback
by: Castricato, Louis, et al.
Published: (2024)
by: Castricato, Louis, et al.
Published: (2024)
Self-Directed Synthetic Dialogues and Revisions Technical Report
by: Lambert, Nathan, et al.
Published: (2024)
by: Lambert, Nathan, et al.
Published: (2024)
Reading and Computers--How Teachers Can Make Them Work Together.
by: Henney, Maribeth
Published: (1984)
by: Henney, Maribeth
Published: (1984)
Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights
by: Javed, Rafiya, et al.
Published: (2025)
by: Javed, Rafiya, et al.
Published: (2025)
In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI
by: Longpre, Shayne, et al.
Published: (2025)
by: Longpre, Shayne, et al.
Published: (2025)
Acceptable Use Policies for Foundation Models
by: Klyman, Kevin
Published: (2024)
by: Klyman, Kevin
Published: (2024)
The AI Consumer Index (ACE)
by: Benchek, Julien, et al.
Published: (2025)
by: Benchek, Julien, et al.
Published: (2025)
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
by: Kirk, Hannah Rose, et al.
Published: (2026)
by: Kirk, Hannah Rose, et al.
Published: (2026)
Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
by: Kirk, Hannah Rose, et al.
Published: (2025)
by: Kirk, Hannah Rose, et al.
Published: (2025)
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
by: Vidgen, Bertie, et al.
Published: (2023)
by: Vidgen, Bertie, et al.
Published: (2023)
Arming Students against Bad Information
by: Smith, Maribeth D.
Published: (2017)
by: Smith, Maribeth D.
Published: (2017)
Symptomatic from Clostridioides difficile or symptomatic from inflammatory bowel disease: Highlighting diagnostic challenges
by: Maribeth R. Nicholson
Published: (2025)
by: Maribeth R. Nicholson
Published: (2025)
The Societal Impact of Foundation Models: Advancing Evidence-based AI Policy
by: Bommasani, Rishi
Published: (2025)
by: Bommasani, Rishi
Published: (2025)
NeurIPS should lead scientific consensus on AI policy
by: Bommasani, Rishi
Published: (2025)
by: Bommasani, Rishi
Published: (2025)
Language model developers should report train-test overlap
by: Zhang, Andy K, et al.
Published: (2024)
by: Zhang, Andy K, et al.
Published: (2024)
Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
A Novel Computational and Modeling Foundation for Automatic Coherence Assessment
by: Maimon, Aviya, et al.
Published: (2023)
by: Maimon, Aviya, et al.
Published: (2023)
Similar Items
-
The 2024 Foundation Model Transparency Index
by: Bommasani, Rishi, et al.
Published: (2024) -
Beyond Release: Access Considerations for Generative AI Systems
by: Solaiman, Irene, et al.
Published: (2025) -
Foundation Model Transparency Reports
by: Bommasani, Rishi, et al.
Published: (2024) -
The 2025 Foundation Model Transparency Index
by: Wan, Alexander, et al.
Published: (2025) -
Power Hungry Processing: Watts Driving the Cost of AI Deployment?
by: Luccioni, Alexandra Sasha, et al.
Published: (2023)