Reward Model Interpretability via Optimal and Pessimal Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Christian, Brian, Kirk, Hannah Rose, Thompson, Jessica A. F., Summerfield, Christopher, Dumbalska, Tsvetomira |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluation of Hate Speech Detection Using Large Language Models and Geographical Contextualization
by: Zahid, Anwar Hossain, et al.
Published: (2025)
by: Zahid, Anwar Hossain, et al.
Published: (2025)
CURATe: Benchmarking Personalised Alignment of Conversational AI Assistants
by: Alberts, Lize, et al.
Published: (2024)
by: Alberts, Lize, et al.
Published: (2024)
SectEval: Evaluating the Latent Sectarian Preferences of Large Language Models
by: Maheshwari, Aditya, et al.
Published: (2026)
by: Maheshwari, Aditya, et al.
Published: (2026)
Chatbot Deployment Considerations for Application-Agnostic Human-Machine Dialogues
by: Rivas, Pablo, et al.
Published: (2025)
by: Rivas, Pablo, et al.
Published: (2025)
Textual Entailment is not a Better Bias Metric than Token Probability
by: Felkner, Virginia K., et al.
Published: (2025)
by: Felkner, Virginia K., et al.
Published: (2025)
What Would GPT Click: Practical Effects of Human-AI Behavioral Misalignment and the Cost of Synthetic Participants in User Experience
by: Kuric, Eduard, et al.
Published: (2026)
by: Kuric, Eduard, et al.
Published: (2026)
Extreme Self-Preference in Language Models
by: Lehr, Steven A., et al.
Published: (2025)
by: Lehr, Steven A., et al.
Published: (2025)
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
by: Felkner, Virginia K., et al.
Published: (2024)
by: Felkner, Virginia K., et al.
Published: (2024)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
by: Portelance, Eva, et al.
Published: (2024)
by: Portelance, Eva, et al.
Published: (2024)
The Company You Keep: How LLMs Respond to Dark Triad Traits
by: Lu, Zeyi, et al.
Published: (2026)
by: Lu, Zeyi, et al.
Published: (2026)
Using a cognitive architecture to consider antiBlackness in design and development of AI systems
by: Dancy, Christopher L.
Published: (2022)
by: Dancy, Christopher L.
Published: (2022)
"I followed what felt right, not what I was told": Autonomy, Coaching, and Recognizing Bias Through AI-Mediated Dialogue
by: Taheri, Atieh, et al.
Published: (2026)
by: Taheri, Atieh, et al.
Published: (2026)
MISCON: A Mission-Driven Conversational Consultant for Pre-Venture Entrepreneurs in Food Deserts
by: Dasgupta, Subhasis, et al.
Published: (2025)
by: Dasgupta, Subhasis, et al.
Published: (2025)
Analysis of LLM as a grammatical feature tagger for African American English
by: Porwal, Rahul, et al.
Published: (2025)
by: Porwal, Rahul, et al.
Published: (2025)
Generative UI as an Accessibility Bridge: Lessons from C2C E-Commerce
by: Ryskeldiev, Bektur
Published: (2026)
by: Ryskeldiev, Bektur
Published: (2026)
How Frontier LLMs Adapt to Neurodivergence Context: A Measurement Framework for Surface vs. Structural Change in System-Prompted Responses
by: Gupta, Ishan, et al.
Published: (2026)
by: Gupta, Ishan, et al.
Published: (2026)
How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krügel, and Uhl (2025)
by: Himmelreich, Johannes
Published: (2026)
by: Himmelreich, Johannes
Published: (2026)
Rejected Dialects: Biases Against African American Language in Reward Models
by: Mire, Joel, et al.
Published: (2025)
by: Mire, Joel, et al.
Published: (2025)
Leveraging Natural Language Processing and Machine Learning for Evidence-Based Food Security Policy Decision-Making in Data-Scarce Making
by: Singh, Karan Kumar, et al.
Published: (2026)
by: Singh, Karan Kumar, et al.
Published: (2026)
Systematic Classification of Studies Investigating Social Media Conversations about Long COVID Using a Novel Zero-Shot Transformer Framework
by: Thakur, Nirmalya, et al.
Published: (2025)
by: Thakur, Nirmalya, et al.
Published: (2025)
Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
by: Du, Yishan, et al.
Published: (2025)
by: Du, Yishan, et al.
Published: (2025)
A Contextual Help Browser Extension to Assist Digital Illiterate Internet Users
by: Koutsiaris, Christos
Published: (2026)
by: Koutsiaris, Christos
Published: (2026)
Playing telephone with generative models: "verification disability," "compelled reliance," and accessibility in data visualization
by: Elavsky, Frank, et al.
Published: (2025)
by: Elavsky, Frank, et al.
Published: (2025)
LLM-Driven Accessible Interface: A Model-Based Approach
by: Jerry, Blessing, et al.
Published: (2026)
by: Jerry, Blessing, et al.
Published: (2026)
Automated Circuit Interpretation via Probe Prompting
by: Birardi, Giuseppe
Published: (2025)
by: Birardi, Giuseppe
Published: (2025)
Dynamics of COVID-19 Misinformation: An Analysis of Conspiracy Theories, Fake Remedies, and False Reports
by: Thakur, Nirmalya, et al.
Published: (2025)
by: Thakur, Nirmalya, et al.
Published: (2025)
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
by: Boumber, Dainis, et al.
Published: (2024)
by: Boumber, Dainis, et al.
Published: (2024)
Can Humans Tell? A Dual-Axis Study of Human Perception of LLM-Generated News
by: Loth, Alexander, et al.
Published: (2026)
by: Loth, Alexander, et al.
Published: (2026)
PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
by: Kajare, Prajwal Vijay, et al.
Published: (2026)
by: Kajare, Prajwal Vijay, et al.
Published: (2026)
SAGE: A Strategy-Aware Graph-Enhanced Generation Framework For Online Counseling
by: Aharon, Eliya Naomi, et al.
Published: (2026)
by: Aharon, Eliya Naomi, et al.
Published: (2026)
Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF
by: Sami, K. M. Jubair, et al.
Published: (2026)
by: Sami, K. M. Jubair, et al.
Published: (2026)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
by: Beltoft, Stine, et al.
Published: (2025)
by: Beltoft, Stine, et al.
Published: (2025)
Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
by: Yeste, Víctor, et al.
Published: (2026)
by: Yeste, Víctor, et al.
Published: (2026)
REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs
by: Cohen, Liran, et al.
Published: (2025)
by: Cohen, Liran, et al.
Published: (2025)
COVID-19 on YouTube: A Data-Driven Analysis of Sentiment, Toxicity, and Content Recommendations
by: Su, Vanessa, et al.
Published: (2024)
by: Su, Vanessa, et al.
Published: (2024)
Content and Engagement Trends in COVID-19 YouTube Videos: Evidence from the Late Pandemic
by: Thakur, Nirmalya, et al.
Published: (2025)
by: Thakur, Nirmalya, et al.
Published: (2025)
Mpox Narrative on Instagram: A Labeled Multilingual Dataset of Instagram Posts on Mpox for Sentiment, Hate Speech, and Anxiety Analysis
by: Thakur, Nirmalya
Published: (2024)
by: Thakur, Nirmalya
Published: (2024)
Five Years of COVID-19 Discourse on Instagram: A Labeled Instagram Dataset of Over Half a Million Posts for Multilingual Sentiment Analysis
by: Thakur, Nirmalya
Published: (2024)
by: Thakur, Nirmalya
Published: (2024)
Emoji Retrieval from Gibberish or Garbled Social Media Text: A Novel Methodology and A Case Study
by: Cui, Shuqi, et al.
Published: (2024)
by: Cui, Shuqi, et al.
Published: (2024)
A Labelled Dataset for Sentiment Analysis of Videos on YouTube, TikTok, and Other Sources about the 2024 Outbreak of Measles
by: Thakur, Nirmalya, et al.
Published: (2024)
by: Thakur, Nirmalya, et al.
Published: (2024)
Similar Items
-
Evaluation of Hate Speech Detection Using Large Language Models and Geographical Contextualization
by: Zahid, Anwar Hossain, et al.
Published: (2025) -
CURATe: Benchmarking Personalised Alignment of Conversational AI Assistants
by: Alberts, Lize, et al.
Published: (2024) -
SectEval: Evaluating the Latent Sectarian Preferences of Large Language Models
by: Maheshwari, Aditya, et al.
Published: (2026) -
Chatbot Deployment Considerations for Application-Agnostic Human-Machine Dialogues
by: Rivas, Pablo, et al.
Published: (2025) -
Textual Entailment is not a Better Bias Metric than Token Probability
by: Felkner, Virginia K., et al.
Published: (2025)