Extreme Self-Preference in Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lehr, Steven A., Cipperman, Mary, Banaji, Mahzarin R. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ChatGPT as Research Scientist: Probing GPT's Capabilities as a Research Librarian, Research Ethicist, Data Generator and Data Predictor
von: Lehr, Steven A., et al.
Veröffentlicht: (2024)
von: Lehr, Steven A., et al.
Veröffentlicht: (2024)
Evaluation of Hate Speech Detection Using Large Language Models and Geographical Contextualization
von: Zahid, Anwar Hossain, et al.
Veröffentlicht: (2025)
von: Zahid, Anwar Hossain, et al.
Veröffentlicht: (2025)
Leveraging Natural Language Processing and Machine Learning for Evidence-Based Food Security Policy Decision-Making in Data-Scarce Making
von: Singh, Karan Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Karan Kumar, et al.
Veröffentlicht: (2026)
Analysis of LLM as a grammatical feature tagger for African American English
von: Porwal, Rahul, et al.
Veröffentlicht: (2025)
von: Porwal, Rahul, et al.
Veröffentlicht: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
Reward Model Interpretability via Optimal and Pessimal Tokens
von: Christian, Brian, et al.
Veröffentlicht: (2025)
von: Christian, Brian, et al.
Veröffentlicht: (2025)
More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
von: Beltoft, Stine, et al.
Veröffentlicht: (2025)
von: Beltoft, Stine, et al.
Veröffentlicht: (2025)
Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
SectEval: Evaluating the Latent Sectarian Preferences of Large Language Models
von: Maheshwari, Aditya, et al.
Veröffentlicht: (2026)
von: Maheshwari, Aditya, et al.
Veröffentlicht: (2026)
MeMo: Towards Language Models with Associative Memory Mechanisms
von: Zanzotto, Fabio Massimo, et al.
Veröffentlicht: (2025)
von: Zanzotto, Fabio Massimo, et al.
Veröffentlicht: (2025)
Do Schwartz Higher-Order Values Help Sentence-Level Human Value Detection? A Study of Hierarchical Gating and Calibration
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
von: Boumber, Dainis, et al.
Veröffentlicht: (2024)
von: Boumber, Dainis, et al.
Veröffentlicht: (2024)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
von: Portelance, Eva, et al.
Veröffentlicht: (2024)
von: Portelance, Eva, et al.
Veröffentlicht: (2024)
Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
von: Du, Yishan, et al.
Veröffentlicht: (2025)
von: Du, Yishan, et al.
Veröffentlicht: (2025)
OntoLogX: Ontology-Guided Knowledge Graph Extraction from Cybersecurity Logs with Large Language Models
von: Cotti, Luca, et al.
Veröffentlicht: (2025)
von: Cotti, Luca, et al.
Veröffentlicht: (2025)
RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences
von: Zhou, Yangyang, et al.
Veröffentlicht: (2026)
von: Zhou, Yangyang, et al.
Veröffentlicht: (2026)
On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance
von: Casanova, Etienne, et al.
Veröffentlicht: (2026)
von: Casanova, Etienne, et al.
Veröffentlicht: (2026)
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
von: Lucas, Tom, et al.
Veröffentlicht: (2026)
von: Lucas, Tom, et al.
Veröffentlicht: (2026)
Categorical Perception in Large Language Model Hidden States: Structural Warping at Digit-Count Boundaries
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
HyperPersona: A Multi-Level Hypergraph Framework for Text-Based Automatic Personality Prediction
von: Heydari, Sina, et al.
Veröffentlicht: (2026)
von: Heydari, Sina, et al.
Veröffentlicht: (2026)
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
von: Chua, Jaymari, et al.
Veröffentlicht: (2025)
von: Chua, Jaymari, et al.
Veröffentlicht: (2025)
PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
von: Kajare, Prajwal Vijay, et al.
Veröffentlicht: (2026)
von: Kajare, Prajwal Vijay, et al.
Veröffentlicht: (2026)
REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs
von: Cohen, Liran, et al.
Veröffentlicht: (2025)
von: Cohen, Liran, et al.
Veröffentlicht: (2025)
Hopscotch: Discovering and Skipping Redundancies in Language Models
von: Eyceoz, Mustafa, et al.
Veröffentlicht: (2025)
von: Eyceoz, Mustafa, et al.
Veröffentlicht: (2025)
COVID-19 on YouTube: A Data-Driven Analysis of Sentiment, Toxicity, and Content Recommendations
von: Su, Vanessa, et al.
Veröffentlicht: (2024)
von: Su, Vanessa, et al.
Veröffentlicht: (2024)
Mpox Narrative on Instagram: A Labeled Multilingual Dataset of Instagram Posts on Mpox for Sentiment, Hate Speech, and Anxiety Analysis
von: Thakur, Nirmalya
Veröffentlicht: (2024)
von: Thakur, Nirmalya
Veröffentlicht: (2024)
Five Years of COVID-19 Discourse on Instagram: A Labeled Instagram Dataset of Over Half a Million Posts for Multilingual Sentiment Analysis
von: Thakur, Nirmalya
Veröffentlicht: (2024)
von: Thakur, Nirmalya
Veröffentlicht: (2024)
Emoji Retrieval from Gibberish or Garbled Social Media Text: A Novel Methodology and A Case Study
von: Cui, Shuqi, et al.
Veröffentlicht: (2024)
von: Cui, Shuqi, et al.
Veröffentlicht: (2024)
A Labelled Dataset for Sentiment Analysis of Videos on YouTube, TikTok, and Other Sources about the 2024 Outbreak of Measles
von: Thakur, Nirmalya, et al.
Veröffentlicht: (2024)
von: Thakur, Nirmalya, et al.
Veröffentlicht: (2024)
Quantifying Public Response to COVID-19 Events: Introducing the Community Sentiment and Engagement Index
von: Thakur, Nirmalya, et al.
Veröffentlicht: (2024)
von: Thakur, Nirmalya, et al.
Veröffentlicht: (2024)
ADALog: Adaptive Unsupervised Anomaly detection in Logs with Self-attention Masked Language Model
von: Pospieszny, Przemek, et al.
Veröffentlicht: (2025)
von: Pospieszny, Przemek, et al.
Veröffentlicht: (2025)
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
von: Hari, Vishnu, et al.
Veröffentlicht: (2025)
von: Hari, Vishnu, et al.
Veröffentlicht: (2025)
MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
von: Lan, Guangchen, et al.
Veröffentlicht: (2025)
von: Lan, Guangchen, et al.
Veröffentlicht: (2025)
HIP Network: Historical Information Passing Network for Extrapolation Reasoning on Temporal Knowledge Graph
von: He, Yongquan, et al.
Veröffentlicht: (2024)
von: He, Yongquan, et al.
Veröffentlicht: (2024)
Machine Unlearning for Masked Diffusion Language Models
von: Lee, Georu, et al.
Veröffentlicht: (2026)
von: Lee, Georu, et al.
Veröffentlicht: (2026)
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
von: Blanco-Justicia, Alberto, et al.
Veröffentlicht: (2024)
von: Blanco-Justicia, Alberto, et al.
Veröffentlicht: (2024)
CURATe: Benchmarking Personalised Alignment of Conversational AI Assistants
von: Alberts, Lize, et al.
Veröffentlicht: (2024)
von: Alberts, Lize, et al.
Veröffentlicht: (2024)
Truth as a Compression Artifact in Language Model Training
von: Krestnikov, Konstantin
Veröffentlicht: (2026)
von: Krestnikov, Konstantin
Veröffentlicht: (2026)
Self-Training Doesn't Flatten Language -- It Restructures It: Surface Markers Amplify While Deep Syntax Dies
von: Liu, Ming
Veröffentlicht: (2026)
von: Liu, Ming
Veröffentlicht: (2026)
Ähnliche Einträge
-
ChatGPT as Research Scientist: Probing GPT's Capabilities as a Research Librarian, Research Ethicist, Data Generator and Data Predictor
von: Lehr, Steven A., et al.
Veröffentlicht: (2024) -
Evaluation of Hate Speech Detection Using Large Language Models and Geographical Contextualization
von: Zahid, Anwar Hossain, et al.
Veröffentlicht: (2025) -
Leveraging Natural Language Processing and Machine Learning for Evidence-Based Food Security Policy Decision-Making in Data-Scarce Making
von: Singh, Karan Kumar, et al.
Veröffentlicht: (2026) -
Analysis of LLM as a grammatical feature tagger for African American English
von: Porwal, Rahul, et al.
Veröffentlicht: (2025) -
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)