Evaluating Language Models for Harmful Manipulation
Fuente:
arXiv
Guardado en:
| Autores principales: | Akbulut, Canfer, Elasmar, Rasmi, Roy, Abhishek, Payne, Anthony, Suresh, Priyanka, Ibrahim, Lujain, El-Sayed, Seliem, Rastogi, Charvi, Kachra, Ashyana, Hawkins, Will, Lum, Kristian, Weidinger, Laura |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
por: Ibrahim, Lujain, et al.
Publicado: (2025)
por: Ibrahim, Lujain, et al.
Publicado: (2025)
STAR: SocioTechnical Approach to Red Teaming Language Models
por: Weidinger, Laura, et al.
Publicado: (2024)
por: Weidinger, Laura, et al.
Publicado: (2024)
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
por: El-Sayed, Seliem, et al.
Publicado: (2024)
por: El-Sayed, Seliem, et al.
Publicado: (2024)
Influence of β Grain Orientation on the Stress‐Induced Martensite Formation in Cold‐Rolled Metastable β Ti–5Mo–5V–5Al–3Cr Alloy
por: Abhishek Rastogi, et al.
Publicado: (2024)
por: Abhishek Rastogi, et al.
Publicado: (2024)
Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data
por: Marchal, Nahema, et al.
Publicado: (2024)
por: Marchal, Nahema, et al.
Publicado: (2024)
Thinking beyond the anthropomorphic paradigm benefits LLM research
por: Ibrahim, Lujain, et al.
Publicado: (2025)
por: Ibrahim, Lujain, et al.
Publicado: (2025)
Characterizing and modeling harms from interactions with design patterns in AI interfaces
por: Ibrahim, Lujain, et al.
Publicado: (2024)
por: Ibrahim, Lujain, et al.
Publicado: (2024)
Mitigating Nursing Care Rationing in Critical Care: The Power of Teamwork Dynamics and Safety Attitudes
por: Boshra Karem Mohamed El‐Sayed, et al.
Publicado: (2025)
por: Boshra Karem Mohamed El‐Sayed, et al.
Publicado: (2025)
PLUTO: A Public Value Assessment Tool
por: Koesten, Laura, et al.
Publicado: (2025)
por: Koesten, Laura, et al.
Publicado: (2025)
Training language models to be warm and empathetic makes them less reliable and more sycophantic
por: Ibrahim, Lujain, et al.
Publicado: (2025)
por: Ibrahim, Lujain, et al.
Publicado: (2025)
AI-Powered Immersive Assistance for Interactive Task Execution in Industrial Environments
por: Duricic, Tomislav, et al.
Publicado: (2024)
por: Duricic, Tomislav, et al.
Publicado: (2024)
First record of Leucoptera malifoliella (O. Costa, 1836) (Lepidoptera: Lyonetiidae) in Tunisia
por: Rasmi Soltani, et al.
Publicado: (2024)
por: Rasmi Soltani, et al.
Publicado: (2024)
The Intersectionality Problem for Algorithmic Fairness
por: Himmelreich, Johannes, et al.
Publicado: (2024)
por: Himmelreich, Johannes, et al.
Publicado: (2024)
Reducing Population-level Inequality Can Improve Demographic Group Fairness: a Twitter Case Study
por: Ghosh, Avijit, et al.
Publicado: (2024)
por: Ghosh, Avijit, et al.
Publicado: (2024)
Adversarial Nibbler: An Open Red-Teaming Method for Identifying Diverse Harms in Text-to-Image Generation
por: Quaye, Jessica, et al.
Publicado: (2024)
por: Quaye, Jessica, et al.
Publicado: (2024)
„Im nationalen Abwehrkampf der Grenzlanddeutschen“
por: Weidinger, Bernhard
Publicado: (2014)
por: Weidinger, Bernhard
Publicado: (2014)
Onward (Im)Mobilities and Integration Processes of Refugee Newcomers in Rural Bavaria, Germany
por: Weidinger, Tobias
Publicado: (2025)
por: Weidinger, Tobias
Publicado: (2025)
Artificial Intelligence Libraries in Time Aeries Forecasting
por: Akbulut, Semiha
Publicado: (2025)
por: Akbulut, Semiha
Publicado: (2025)
Corks
por: Akbulut, Selman
Publicado: (2024)
por: Akbulut, Selman
Publicado: (2024)
On 4-dimensional smooth Poincare conjecture
por: Akbulut, Selman
Publicado: (2022)
por: Akbulut, Selman
Publicado: (2022)
Quality of Life After Open Surgical versus Endovascular Repair of Abdominal Aortic Aneurysms
por: Mustafa Akbulut
Publicado: (2018)
por: Mustafa Akbulut
Publicado: (2018)
On the Smale Conjecture for Diff$(S^4)$
por: Akbulut, Selman
Publicado: (2020)
por: Akbulut, Selman
Publicado: (2020)
PERCEPCIONES Y ACTITUDES DE LA POBLACIÓN LOCAL HACIA EL TURISMO CINEMATOGRÁFICO EN EL CONTEXTO DEL APEGO AL LUGAR Un estudio en la Provincia de Mugla, Turquía
por: Onur Akbulut
Publicado: (2018)
por: Onur Akbulut
Publicado: (2018)
Longitudinal follow up of excessive video game players during the COVID-19 pandemic
por: Ozlem Akbulut
Publicado: (2025)
por: Ozlem Akbulut
Publicado: (2025)
The association between quality of habitual diet and mental health status among Iraqi women attending primary health care centers
por: Lujain Anwar Alkhazrajy
Publicado: (2020)
por: Lujain Anwar Alkhazrajy
Publicado: (2020)
Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South
por: Rastogi, Charvi, et al.
Publicado: (2026)
por: Rastogi, Charvi, et al.
Publicado: (2026)
From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models
por: Quaye, Jessica, et al.
Publicado: (2025)
por: Quaye, Jessica, et al.
Publicado: (2025)
To ArXiv or not to ArXiv: A Study Quantifying Pros and Cons of Posting Preprints Online
por: Rastogi, Charvi, et al.
Publicado: (2022)
por: Rastogi, Charvi, et al.
Publicado: (2022)
Synthesis of CoFe2O4‐Integrated CAR/PNIPAm Nanogels Through Gamma Irradiation: Analysis of Structural, Bioactive, and Anticancer Properties
por: Amal Shawky, et al.
Publicado: (2025)
por: Amal Shawky, et al.
Publicado: (2025)
Navigating Quiet Quitting Among Critical Care Nurses in the Digital Era: The Impact of Techno‐Stress and Digital Resilience
por: Boshra Karem Mohamed El‐Sayed, et al.
Publicado: (2025)
por: Boshra Karem Mohamed El‐Sayed, et al.
Publicado: (2025)
A Randomized Controlled Trial on Anonymizing Reviewers to Each Other in Peer Review Discussions
por: Rastogi, Charvi, et al.
Publicado: (2024)
por: Rastogi, Charvi, et al.
Publicado: (2024)
The illusion of artificial inclusion
por: Agnew, William, et al.
Publicado: (2024)
por: Agnew, William, et al.
Publicado: (2024)
The Move to Mobile: Where Is a Campus's Place in the Mobile Space?
por: Lum, Lydia
Publicado: (2012)
por: Lum, Lydia
Publicado: (2012)
Student Attitude/Satisfaction Survey--Lancaster Campus.
por: Lum, Glen
Publicado: (1993)
por: Lum, Glen
Publicado: (1993)
Towards interactive evaluations for interaction harms in human-AI systems
por: Ibrahim, Lujain, et al.
Publicado: (2024)
por: Ibrahim, Lujain, et al.
Publicado: (2024)
Redshift Space Distortions corner interacting Dark Energy
por: Ghedini, Pietro, et al.
Publicado: (2024)
por: Ghedini, Pietro, et al.
Publicado: (2024)
Dark energy and neutrinos along the cosmic expansion history
por: Ghedini, Pietro, et al.
Publicado: (2025)
por: Ghedini, Pietro, et al.
Publicado: (2025)
Duality of zero mean curvature surfaces in the Lorentzian Heisenberg group
por: Mohanty, Sai Rasmi Ranjan, et al.
Publicado: (2026)
por: Mohanty, Sai Rasmi Ranjan, et al.
Publicado: (2026)
GBSVR: Granular Ball Support Vector Regression
por: Rastogi, Reshma, et al.
Publicado: (2025)
por: Rastogi, Reshma, et al.
Publicado: (2025)
The Impossibility of Fair LLMs
por: Anthis, Jacy, et al.
Publicado: (2024)
por: Anthis, Jacy, et al.
Publicado: (2024)
Ejemplares similares
-
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
por: Ibrahim, Lujain, et al.
Publicado: (2025) -
STAR: SocioTechnical Approach to Red Teaming Language Models
por: Weidinger, Laura, et al.
Publicado: (2024) -
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
por: El-Sayed, Seliem, et al.
Publicado: (2024) -
Influence of β Grain Orientation on the Stress‐Induced Martensite Formation in Cold‐Rolled Metastable β Ti–5Mo–5V–5Al–3Cr Alloy
por: Abhishek Rastogi, et al.
Publicado: (2024) -
Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data
por: Marchal, Nahema, et al.
Publicado: (2024)