Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
Fuente:
arXiv
Salvato in:
| Autori principali: | Oesterling, Alex, Bhalla, Usha, Venkatasubramanian, Suresh, Lakkaraju, Himabindu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
di: Zhang, Shichang, et al.
Pubblicazione: (2025)
di: Zhang, Shichang, et al.
Pubblicazione: (2025)
Learning Recourse Costs from Pairwise Feature Comparisons
di: Rawal, Kaivalya, et al.
Pubblicazione: (2024)
di: Rawal, Kaivalya, et al.
Pubblicazione: (2024)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
di: Li, Aaron J., et al.
Pubblicazione: (2025)
di: Li, Aaron J., et al.
Pubblicazione: (2025)
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
di: Bhalla, Usha, et al.
Pubblicazione: (2024)
di: Bhalla, Usha, et al.
Pubblicazione: (2024)
All Roads Lead to Rome? Exploring Representational Similarities Between Latent Spaces of Generative Image Models
di: Badrinath, Charumathi, et al.
Pubblicazione: (2024)
di: Badrinath, Charumathi, et al.
Pubblicazione: (2024)
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2023)
di: Bhalla, Usha, et al.
Pubblicazione: (2023)
Who Followed the Blueprint? Analyzing the Responses of U.S. Federal Agencies to the Blueprint for an AI Bill of Rights
di: Lage, Darren, et al.
Pubblicazione: (2024)
di: Lage, Darren, et al.
Pubblicazione: (2024)
The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective
di: Krishna, Satyapriya, et al.
Pubblicazione: (2022)
di: Krishna, Satyapriya, et al.
Pubblicazione: (2022)
Counting Hours, Counting Losses: The Toll of Unpredictable Work Schedules on Financial Security
di: Nokhiz, Pegah, et al.
Pubblicazione: (2025)
di: Nokhiz, Pegah, et al.
Pubblicazione: (2025)
Red Teaming AI Policy: A Taxonomy of Avoision and the EU AI Act
di: Yew, Rui-Jie, et al.
Pubblicazione: (2025)
di: Yew, Rui-Jie, et al.
Pubblicazione: (2025)
Fair Machine Unlearning: Data Removal while Mitigating Disparities
di: Oesterling, Alex, et al.
Pubblicazione: (2023)
di: Oesterling, Alex, et al.
Pubblicazione: (2023)
Towards Unifying Interpretability and Control: Evaluation via Intervention
di: Bhalla, Usha, et al.
Pubblicazione: (2024)
di: Bhalla, Usha, et al.
Pubblicazione: (2024)
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
di: Xiong, Zidi, et al.
Pubblicazione: (2026)
di: Xiong, Zidi, et al.
Pubblicazione: (2026)
Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
di: Zhang, Shichang, et al.
Pubblicazione: (2025)
di: Zhang, Shichang, et al.
Pubblicazione: (2025)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
di: Lobo, Elita, et al.
Pubblicazione: (2024)
di: Lobo, Elita, et al.
Pubblicazione: (2024)
Towards Operationalizing Right to Data Protection
di: Java, Abhinav, et al.
Pubblicazione: (2024)
di: Java, Abhinav, et al.
Pubblicazione: (2024)
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
di: Cooper, A. Feder, et al.
Pubblicazione: (2024)
di: Cooper, A. Feder, et al.
Pubblicazione: (2024)
Fairness-in-the-Workflow: How Machine Learning Practitioners at Big Tech Companies Approach Fairness in Recommender Systems
di: Yan, Jing Nathan, et al.
Pubblicazione: (2025)
di: Yan, Jing Nathan, et al.
Pubblicazione: (2025)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
di: Kroeger, Nicholas, et al.
Pubblicazione: (2023)
di: Kroeger, Nicholas, et al.
Pubblicazione: (2023)
HH4AI: A methodological Framework for AI Human Rights impact assessment under the EUAI ACT
di: Ceravolo, Paolo, et al.
Pubblicazione: (2025)
di: Ceravolo, Paolo, et al.
Pubblicazione: (2025)
Distinguishing Task-Specific and General-Purpose AI in Regulation
di: Wang, Jennifer, et al.
Pubblicazione: (2025)
di: Wang, Jennifer, et al.
Pubblicazione: (2025)
In-Context Unlearning: Language Models as Few Shot Unlearners
di: Pawelczyk, Martin, et al.
Pubblicazione: (2023)
di: Pawelczyk, Martin, et al.
Pubblicazione: (2023)
Case Studies of AI Policy Development in Africa
di: Diallo, Kadijatou, et al.
Pubblicazione: (2024)
di: Diallo, Kadijatou, et al.
Pubblicazione: (2024)
Soft Best-of-n Sampling for Model Alignment
di: Verdun, Claudio Mayrink, et al.
Pubblicazione: (2025)
di: Verdun, Claudio Mayrink, et al.
Pubblicazione: (2025)
Position Paper: If Innovation in AI Systematically Violates Fundamental Rights, Is It Innovation at All?
di: Castañeira, Josu Eguiluz, et al.
Pubblicazione: (2025)
di: Castañeira, Josu Eguiluz, et al.
Pubblicazione: (2025)
Measuring What AI Systems Might Do: Towards A Measurement Science in AI
di: Voudouris, Konstantinos, et al.
Pubblicazione: (2026)
di: Voudouris, Konstantinos, et al.
Pubblicazione: (2026)
The Malicious Technical Ecosystem: Exposing Limitations in Technical Governance of AI-Generated Non-Consensual Intimate Images of Adults
di: Ding, Michelle L., et al.
Pubblicazione: (2025)
di: Ding, Michelle L., et al.
Pubblicazione: (2025)
Beware! The AI Act Can Also Apply to Your AI Research Practices
di: Wernick, Alina, et al.
Pubblicazione: (2025)
di: Wernick, Alina, et al.
Pubblicazione: (2025)
Selecting the Right LLM for eGov Explanations
di: Limonad, Lior, et al.
Pubblicazione: (2025)
di: Limonad, Lior, et al.
Pubblicazione: (2025)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
di: Qi, Zhenting, et al.
Pubblicazione: (2024)
di: Qi, Zhenting, et al.
Pubblicazione: (2024)
Generative AI Policies under the Microscope: How CS Conferences Are Navigating the New Frontier in Scholarly Writing
di: Nahar, Mahjabin, et al.
Pubblicazione: (2024)
di: Nahar, Mahjabin, et al.
Pubblicazione: (2024)
Acting for the Right Reasons: Creating Reason-Sensitive Artificial Moral Agents
di: Baum, Kevin, et al.
Pubblicazione: (2024)
di: Baum, Kevin, et al.
Pubblicazione: (2024)
Generalized Group Data Attribution
di: Ley, Dan, et al.
Pubblicazione: (2024)
di: Ley, Dan, et al.
Pubblicazione: (2024)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
di: Pawelczyk, Martin, et al.
Pubblicazione: (2024)
di: Pawelczyk, Martin, et al.
Pubblicazione: (2024)
Misrepresented Technological Solutions in Imagined Futures: The Origins and Dangers of AI Hype in the Research Community
di: Thais, Savannah
Pubblicazione: (2024)
di: Thais, Savannah
Pubblicazione: (2024)
AI, Climate, and Transparency: Operationalizing and Improving the AI Act
di: Alder, Nicolas, et al.
Pubblicazione: (2024)
di: Alder, Nicolas, et al.
Pubblicazione: (2024)
AI-Powered Autonomous Weapons Risk Geopolitical Instability and Threaten AI Research
di: Simmons-Edler, Riley, et al.
Pubblicazione: (2024)
di: Simmons-Edler, Riley, et al.
Pubblicazione: (2024)
Enhancing Math Learning in an LMS Using AI-Driven Question Recommendations
di: Råmunddal, Justus
Pubblicazione: (2025)
di: Råmunddal, Justus
Pubblicazione: (2025)
Influence of Recommender Systems on Users: A Dynamical Systems Analysis
di: Lankireddy, Prabhat, et al.
Pubblicazione: (2024)
di: Lankireddy, Prabhat, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2025) -
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
di: Zhang, Shichang, et al.
Pubblicazione: (2025) -
Learning Recourse Costs from Pairwise Feature Comparisons
di: Rawal, Kaivalya, et al.
Pubblicazione: (2024) -
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
di: Li, Aaron J., et al.
Pubblicazione: (2025) -
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
di: Bhalla, Usha, et al.
Pubblicazione: (2024)