Self-Reflection Makes Large Language Models Safer, Less Biased, and Ideologically Neutral
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Fengyuan, AlDahoul, Nouar, Eady, Gregory, Zaki, Yasir, Rahwan, Talal |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
A Novel BERT-based Classifier to Detect Political Leaning of YouTube Videos based on their Titles
par: AlDahoul, Nouar, et autres
Publié: (2024)
par: AlDahoul, Nouar, et autres
Publié: (2024)
Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
par: AlDahoul, Nouar, et autres
Publié: (2025)
par: AlDahoul, Nouar, et autres
Publié: (2025)
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
par: AlDahoul, Nouar, et autres
Publié: (2025)
par: AlDahoul, Nouar, et autres
Publié: (2025)
AI-generated faces influence gender stereotypes and racial homogenization
par: AlDahoul, Nouar, et autres
Publié: (2024)
par: AlDahoul, Nouar, et autres
Publié: (2024)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
par: AlDahoul, Nouar, et autres
Publié: (2025)
par: AlDahoul, Nouar, et autres
Publié: (2025)
Neutralizing the Narrative: AI-Powered Debiasing of Online News Articles
par: Kuo, Chen Wei, et autres
Publié: (2025)
par: Kuo, Chen Wei, et autres
Publié: (2025)
Inclusive content reduces racial and gender biases, yet non-inclusive content dominates popular culture
par: AlDahoul, Nouar, et autres
Publié: (2024)
par: AlDahoul, Nouar, et autres
Publié: (2024)
A Longitudinal Analysis of Racial and Gender Bias in New York Times and Fox News Images and Articles
par: Ibrahim, Hazem, et autres
Publié: (2024)
par: Ibrahim, Hazem, et autres
Publié: (2024)
Large Language Models are often politically extreme, usually ideologically inconsistent, and persuasive even in informational contexts
par: Aldahoul, Nouar, et autres
Publié: (2025)
par: Aldahoul, Nouar, et autres
Publié: (2025)
Multitask Mayhem: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning
par: Jan, Essa, et autres
Publié: (2024)
par: Jan, Essa, et autres
Publié: (2024)
Advancing Content Moderation: Evaluating Large Language Models for Detecting Sensitive Content Across Text, Images, and Videos
par: AlDahoul, Nouar, et autres
Publié: (2024)
par: AlDahoul, Nouar, et autres
Publié: (2024)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
par: Aldahoul, Nouar, et autres
Publié: (2025)
par: Aldahoul, Nouar, et autres
Publié: (2025)
Exploring Vision Language Models for Facial Attribute Recognition: Emotion, Race, Gender, and Age
par: AlDahoul, Nouar, et autres
Publié: (2024)
par: AlDahoul, Nouar, et autres
Publié: (2024)
Who Gets Seen in the Age of AI? Adoption Patterns of Large Language Models in Scholarly Writing and Citation Outcomes
par: Khan, Farhan Kamrul, et autres
Publié: (2025)
par: Khan, Farhan Kamrul, et autres
Publié: (2025)
Google Scholar is manipulatable
par: Ibrahim, Hazem, et autres
Publié: (2024)
par: Ibrahim, Hazem, et autres
Publié: (2024)
Fine-tuned Vision Language Model for Localization of Parasitic Eggs in Microscopic Images
par: Sien, Chan Hao, et autres
Publié: (2026)
par: Sien, Chan Hao, et autres
Publié: (2026)
HowkGPT: Investigating the Detection of ChatGPT-generated University Student Homework through Context-Aware Perplexity Analysis
par: Vasilatos, Christoforos, et autres
Publié: (2023)
par: Vasilatos, Christoforos, et autres
Publié: (2023)
Real-Time Human Detection for Aerial Captured Video Sequences via Deep Models
par: AlDahoul, Nouar, et autres
Publié: (2026)
par: AlDahoul, Nouar, et autres
Publié: (2026)
Advancing Vehicle Plate Recognition: Multitasking Visual Language Models with VehiclePaliGemma
par: AlDahoul, Nouar, et autres
Publié: (2024)
par: AlDahoul, Nouar, et autres
Publié: (2024)
Causal evidence of racial and institutional biases in accessing paywalled articles and scientific data
par: Ibrahim, Hazem, et autres
Publié: (2025)
par: Ibrahim, Hazem, et autres
Publié: (2025)
TikTok's recommendations skewed towards Republican content during the 2024 U.S. presidential race
par: Ibrahim, Hazem, et autres
Publié: (2025)
par: Ibrahim, Hazem, et autres
Publié: (2025)
Schadenfreude in the Digital Public Sphere: A cross-national and decade-long analysis of Facebook news engagement
par: Aldahoul, Nouar, et autres
Publié: (2026)
par: Aldahoul, Nouar, et autres
Publié: (2026)
The Algorithmic Caricature: Auditing LLM-Generated Political Discourse Across Crisis Events
par: Gunjan, et autres
Publié: (2026)
par: Gunjan, et autres
Publié: (2026)
A Conceptual Exploration of Generative AI-Induced Cognitive Dissonance and its Emergence in University-Level Academic Writing
par: Seran, Carl Errol, et autres
Publié: (2025)
par: Seran, Carl Errol, et autres
Publié: (2025)
Large Language Models Reflect the Ideology of their Creators
par: Buyl, Maarten, et autres
Publié: (2024)
par: Buyl, Maarten, et autres
Publié: (2024)
Enhancing Password Security Through a High-Accuracy Scoring Framework Using Random Forests
par: Mazelan, Muhammed El Mustaqeem, et autres
Publié: (2025)
par: Mazelan, Muhammed El Mustaqeem, et autres
Publié: (2025)
Decoding the Mind of Large Language Models: A Quantitative Evaluation of Ideology and Biases
par: Hirose, Manari, et autres
Publié: (2025)
par: Hirose, Manari, et autres
Publié: (2025)
Current policies governing editorial conflicts of interest are ineffective
par: Liu, Fengyuan, et autres
Publié: (2023)
par: Liu, Fengyuan, et autres
Publié: (2023)
Towards Safer Large Language Models through Machine Unlearning
par: Liu, Zheyuan, et autres
Publié: (2024)
par: Liu, Zheyuan, et autres
Publié: (2024)
Gender inequality and self-publication patterns among scientific editors
par: Liu, Fengyuan, et autres
Publié: (2022)
par: Liu, Fengyuan, et autres
Publié: (2022)
Less is More: Selective Reflection for Compatible and Efficient Knowledge Distillation in Large Language Models
par: Liu, Lingyuan, et autres
Publié: (2025)
par: Liu, Lingyuan, et autres
Publié: (2025)
Uncovering Biases with Reflective Large Language Models
par: Chang, Edward Y.
Publié: (2024)
par: Chang, Edward Y.
Publié: (2024)
Political Ideology Shifts in Large Language Models
par: Bernardelle, Pietro, et autres
Publié: (2025)
par: Bernardelle, Pietro, et autres
Publié: (2025)
Probing the Subtle Ideological Manipulation of Large Language Models
par: Paschalides, Demetris, et autres
Publié: (2025)
par: Paschalides, Demetris, et autres
Publié: (2025)
Beyond the Surface: Probing the Ideological Depth of Large Language Models
par: Kabir, Shariar, et autres
Publié: (2025)
par: Kabir, Shariar, et autres
Publié: (2025)
Analyzing political stances on Twitter in the lead-up to the 2024 U.S. election
par: Ibrahim, Hazem, et autres
Publié: (2024)
par: Ibrahim, Hazem, et autres
Publié: (2024)
Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models
par: Christian, Brian, et autres
Publié: (2026)
par: Christian, Brian, et autres
Publié: (2026)
Evaluating Large Language Model Biases in Persona-Steered Generation
par: Liu, Andy, et autres
Publié: (2024)
par: Liu, Andy, et autres
Publié: (2024)
How Susceptible are Large Language Models to Ideological Manipulation?
par: Chen, Kai, et autres
Publié: (2024)
par: Chen, Kai, et autres
Publié: (2024)
Large Language Models are Biased Because They Are Large Language Models
par: Resnik, Philip
Publié: (2024)
par: Resnik, Philip
Publié: (2024)
Documents similaires
-
A Novel BERT-based Classifier to Detect Political Leaning of YouTube Videos based on their Titles
par: AlDahoul, Nouar, et autres
Publié: (2024) -
Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
par: AlDahoul, Nouar, et autres
Publié: (2025) -
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
par: AlDahoul, Nouar, et autres
Publié: (2025) -
AI-generated faces influence gender stereotypes and racial homogenization
par: AlDahoul, Nouar, et autres
Publié: (2024) -
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
par: AlDahoul, Nouar, et autres
Publié: (2025)