Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
Fuente:
arXiv
Guardado en:
| Autores principales: | Mazeika, Mantas, Yin, Xuwang, Tamirisa, Rishub, Lim, Jaehyuk, Lee, Bruce W., Ren, Richard, Phan, Long, Mu, Norman, Khoja, Adam, Zhang, Oliver, Hendrycks, Dan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TextQuests: How Good are LLMs at Text-Based Video Games?
por: Phan, Long, et al.
Publicado: (2025)
por: Phan, Long, et al.
Publicado: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
por: Ren, Richard, et al.
Publicado: (2024)
por: Ren, Richard, et al.
Publicado: (2024)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
por: Mazeika, Mantas, et al.
Publicado: (2024)
por: Mazeika, Mantas, et al.
Publicado: (2024)
Tamper-Resistant Safeguards for Open-Weight LLMs
por: Tamirisa, Rishub, et al.
Publicado: (2024)
por: Tamirisa, Rishub, et al.
Publicado: (2024)
Reducing Political Manipulation with Consistency Training
por: Phan, Long, et al.
Publicado: (2026)
por: Phan, Long, et al.
Publicado: (2026)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
por: Ren, Richard, et al.
Publicado: (2025)
por: Ren, Richard, et al.
Publicado: (2025)
Representation Engineering: A Top-Down Approach to AI Transparency
por: Zou, Andy, et al.
Publicado: (2023)
por: Zou, Andy, et al.
Publicado: (2023)
Measuring Agreeableness Bias in Multimodal Models
por: Lim, Jaehyuk, et al.
Publicado: (2024)
por: Lim, Jaehyuk, et al.
Publicado: (2024)
Introduction to AI Safety, Ethics, and Society
por: Hendrycks, Dan
Publicado: (2024)
por: Hendrycks, Dan
Publicado: (2024)
Aggressive Compression Enables LLM Weight Theft
por: Brown, Davis, et al.
Publicado: (2026)
por: Brown, Davis, et al.
Publicado: (2026)
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
por: Wang, Clinton J., et al.
Publicado: (2025)
por: Wang, Clinton J., et al.
Publicado: (2025)
Superintelligence Strategy: Expert Version
por: Hendrycks, Dan, et al.
Publicado: (2025)
por: Hendrycks, Dan, et al.
Publicado: (2025)
Improving Alignment and Robustness with Circuit Breakers
por: Zou, Andy, et al.
Publicado: (2024)
por: Zou, Andy, et al.
Publicado: (2024)
Introduction to AI Safety, Ethics, and Society
por: Hendrycks, Dan
Publicado: (2024)
por: Hendrycks, Dan
Publicado: (2024)
Embedding Democratic Values into Social Media AIs via Societal Objective Functions
por: Jia, Chenyan, et al.
Publicado: (2023)
por: Jia, Chenyan, et al.
Publicado: (2023)
What AIs are not Learning (and Why)
por: Stefik, Mark
Publicado: (2024)
por: Stefik, Mark
Publicado: (2024)
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
por: Li, Nathaniel, et al.
Publicado: (2024)
por: Li, Nathaniel, et al.
Publicado: (2024)
Exact simulation scheme for the Ornstein-Uhlenbeck driven stochastic volatility model with the Karhunen-Loève expansions
por: Choi, Jaehyuk
Publicado: (2024)
por: Choi, Jaehyuk
Publicado: (2024)
Can LLMs Follow Simple Rules?
por: Mu, Norman, et al.
Publicado: (2023)
por: Mu, Norman, et al.
Publicado: (2023)
Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
por: Götting, Jasper, et al.
Publicado: (2025)
por: Götting, Jasper, et al.
Publicado: (2025)
Runtime Monitoring and Enforcement of Conditional Fairness in Generative AIs
por: Cheng, Chih-Hong, et al.
Publicado: (2024)
por: Cheng, Chih-Hong, et al.
Publicado: (2024)
Multi-class Image Anomaly Detection for Practical Applications: Requirements and Robust Solutions
por: Heo, Jaehyuk, et al.
Publicado: (2025)
por: Heo, Jaehyuk, et al.
Publicado: (2025)
The AI Double Standard: Humans Judge All AIs for the Actions of One
por: Manoli, Aikaterina, et al.
Publicado: (2024)
por: Manoli, Aikaterina, et al.
Publicado: (2024)
IGGA: A Dataset of Industrial Guidelines and Policy Statements for Generative AIs
por: Jiao, Junfeng, et al.
Publicado: (2025)
por: Jiao, Junfeng, et al.
Publicado: (2025)
Mark My Words: Analyzing and Evaluating Language Model Watermarks
por: Piet, Julien, et al.
Publicado: (2023)
por: Piet, Julien, et al.
Publicado: (2023)
Exploring the Impact of Occupational Personas on Domain-Specific QA
por: Kang, Eojin, et al.
Publicado: (2025)
por: Kang, Eojin, et al.
Publicado: (2025)
Remote Labor Index: Measuring AI Automation of Remote Work
por: Mazeika, Mantas, et al.
Publicado: (2025)
por: Mazeika, Mantas, et al.
Publicado: (2025)
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
por: Wang, Boxin, et al.
Publicado: (2023)
por: Wang, Boxin, et al.
Publicado: (2023)
Uncovering Latent Human Wellbeing in Language Model Embeddings
por: Freire, Pedro, et al.
Publicado: (2024)
por: Freire, Pedro, et al.
Publicado: (2024)
Do Consumers Accept AIs as Moral Compliance Agents?
por: Nyilasy, Greg, et al.
Publicado: (2026)
por: Nyilasy, Greg, et al.
Publicado: (2026)
Speaking Your Language: Spatial Relationships in Interpretable Emergent Communication
por: Lipinski, Olaf, et al.
Publicado: (2024)
por: Lipinski, Olaf, et al.
Publicado: (2024)
Avoid Wasted Annotation Costs in Open-set Active Learning with Pre-trained Vision-Language Model
por: Heo, Jaehyuk, et al.
Publicado: (2024)
por: Heo, Jaehyuk, et al.
Publicado: (2024)
Synthetic Categorical Restructuring large Or How AIs Gradually Extract Efficient Regularities from Their Experience of the World
por: Pichat, Michael, et al.
Publicado: (2025)
por: Pichat, Michael, et al.
Publicado: (2025)
Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
por: Huang, Saffron, et al.
Publicado: (2025)
por: Huang, Saffron, et al.
Publicado: (2025)
Integrating AIs With Body Tracking Technology for Human Behaviour Analysis: Challenges and Opportunities
por: Coppens, Adrien, et al.
Publicado: (2025)
por: Coppens, Adrien, et al.
Publicado: (2025)
It's About Time: Temporal References in Emergent Communication
por: Lipinski, Olaf, et al.
Publicado: (2023)
por: Lipinski, Olaf, et al.
Publicado: (2023)
A mathematical theory of evolution for self-designing AIs
por: Harris, Kenneth D
Publicado: (2026)
por: Harris, Kenneth D
Publicado: (2026)
Alignment-Process-Outcome: Rethinking How AIs and Humans Collaborate
por: Li, Haichang, et al.
Publicado: (2026)
por: Li, Haichang, et al.
Publicado: (2026)
Draw an Ugly Person An Exploration of Generative AIs Perceptions of Ugliness
por: Kim, Garyoung, et al.
Publicado: (2025)
por: Kim, Garyoung, et al.
Publicado: (2025)
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
por: Giang, et al.
Publicado: (2023)
por: Giang, et al.
Publicado: (2023)
Ejemplares similares
-
TextQuests: How Good are LLMs at Text-Based Video Games?
por: Phan, Long, et al.
Publicado: (2025) -
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
por: Ren, Richard, et al.
Publicado: (2024) -
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
por: Mazeika, Mantas, et al.
Publicado: (2024) -
Tamper-Resistant Safeguards for Open-Weight LLMs
por: Tamirisa, Rishub, et al.
Publicado: (2024) -
Reducing Political Manipulation with Consistency Training
por: Phan, Long, et al.
Publicado: (2026)