Suppressing Pink Elephants with Direct Principle Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Castricato, Louis, Lile, Nathan, Anand, Suraj, Schoelkopf, Hailey, Verma, Siddharth, Biderman, Stella |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Directed Synthetic Dialogues and Revisions Technical Report
by: Lambert, Nathan, et al.
Published: (2024)
by: Lambert, Nathan, et al.
Published: (2024)
PERSONA: A Reproducible Testbed for Pluralistic Alignment
by: Castricato, Louis, et al.
Published: (2024)
by: Castricato, Louis, et al.
Published: (2024)
PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs
by: van der Wal, Oskar, et al.
Published: (2025)
by: van der Wal, Oskar, et al.
Published: (2025)
Negation: A Pink Elephant in the Large Language Models' Room?
by: Vrabcová, Tereza, et al.
Published: (2025)
by: Vrabcová, Tereza, et al.
Published: (2025)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
Llemma: An Open Language Model For Mathematics
by: Azerbayev, Zhangir, et al.
Published: (2023)
by: Azerbayev, Zhangir, et al.
Published: (2023)
Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
by: Albalak, Alon, et al.
Published: (2025)
by: Albalak, Alon, et al.
Published: (2025)
LLM Circuit Analyses Are Consistent Across Training and Scale
by: Tigges, Curt, et al.
Published: (2024)
by: Tigges, Curt, et al.
Published: (2024)
Explaining and Mitigating Crosslingual Tokenizer Inequities
by: Arnett, Catherine, et al.
Published: (2025)
by: Arnett, Catherine, et al.
Published: (2025)
Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
by: Xiang, Violet, et al.
Published: (2025)
by: Xiang, Violet, et al.
Published: (2025)
Transformer-Based Models Are Not Yet Perfect At Learning to Emulate Structural Recursion
by: Zhang, Dylan, et al.
Published: (2024)
by: Zhang, Dylan, et al.
Published: (2024)
Are PPO-ed Language Models Hackable?
by: Anand, Suraj, et al.
Published: (2024)
by: Anand, Suraj, et al.
Published: (2024)
From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models
by: Welleck, Sean, et al.
Published: (2024)
by: Welleck, Sean, et al.
Published: (2024)
LEACE: Perfect linear concept erasure in closed form
by: Belrose, Nora, et al.
Published: (2023)
by: Belrose, Nora, et al.
Published: (2023)
Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
by: Gandhi, Kanishk, et al.
Published: (2025)
by: Gandhi, Kanishk, et al.
Published: (2025)
Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
by: Conitzer, Vincent, et al.
Published: (2024)
by: Conitzer, Vincent, et al.
Published: (2024)
Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight Forgetting
by: Anand, Suraj, et al.
Published: (2024)
by: Anand, Suraj, et al.
Published: (2024)
KMMLU: Measuring Massive Multitask Language Understanding in Korean
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
The Responsible Foundation Model Development Cheatsheet: A Review of Tools & Resources
by: Longpre, Shayne, et al.
Published: (2024)
by: Longpre, Shayne, et al.
Published: (2024)
Leveraging LLM For Synchronizing Information Across Multilingual Tables
by: Khincha, Siddharth, et al.
Published: (2025)
by: Khincha, Siddharth, et al.
Published: (2025)
The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing Research
by: Abdalla, Mohamed, et al.
Published: (2023)
by: Abdalla, Mohamed, et al.
Published: (2023)
The Elephant in the Coreference Room: Resolving Coreference in Full-Length French Fiction Works
by: Bourgois, Antoine, et al.
Published: (2025)
by: Bourgois, Antoine, et al.
Published: (2025)
Through the LLM Looking Glass: A Socratic Probing of Donkeys, Elephants, and Markets
by: Kennedy, Molly, et al.
Published: (2025)
by: Kennedy, Molly, et al.
Published: (2025)
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
Lessons from the Trenches on Reproducible Evaluation of Language Models
by: Biderman, Stella, et al.
Published: (2024)
by: Biderman, Stella, et al.
Published: (2024)
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
by: Kuissi, Nathan, et al.
Published: (2026)
by: Kuissi, Nathan, et al.
Published: (2026)
The Ghost in the Keys: A Disklavier Demo for Human-AI Musical Co-Creativity
by: Bradshaw, Louis, et al.
Published: (2025)
by: Bradshaw, Louis, et al.
Published: (2025)
Elephant in the Room: Unveiling the Impact of Reward Model Quality in Alignment
by: Liu, Yan, et al.
Published: (2024)
by: Liu, Yan, et al.
Published: (2024)
InfoGatherer: Principled Information Seeking via Evidence Retrieval and Strategic Questioning
by: Taranukhin, Maksym, et al.
Published: (2026)
by: Taranukhin, Maksym, et al.
Published: (2026)
Exposing Pink Slime Journalism: Linguistic Signatures and Robust Detection Against LLM-Generated Threats
by: Shahriar, Sadat, et al.
Published: (2025)
by: Shahriar, Sadat, et al.
Published: (2025)
$\forall$uto$\exists$val: Autonomous Assessment of LLMs in Formal Synthesis and Interpretation Tasks
by: Karia, Rushang, et al.
Published: (2024)
by: Karia, Rushang, et al.
Published: (2024)
The Importance of Directional Feedback for LLM-based Optimizers
by: Nie, Allen, et al.
Published: (2024)
by: Nie, Allen, et al.
Published: (2024)
How English Print Media Frames Human-Elephant Conflicts in India
by: Punith, Bonala Sai, et al.
Published: (2026)
by: Punith, Bonala Sai, et al.
Published: (2026)
Elephants Never Forget: Testing Language Models for Memorization of Tabular Data
by: Bordt, Sebastian, et al.
Published: (2024)
by: Bordt, Sebastian, et al.
Published: (2024)
Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets
by: Zakizadeh, Mahdi, et al.
Published: (2025)
by: Zakizadeh, Mahdi, et al.
Published: (2025)
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm
by: Shashidhar, Sarvesh, et al.
Published: (2025)
by: Shashidhar, Sarvesh, et al.
Published: (2025)
Teaching LLMs to Abstain across Languages via Multilingual Feedback
by: Feng, Shangbin, et al.
Published: (2024)
by: Feng, Shangbin, et al.
Published: (2024)
Semantic Search Evaluation
by: Zheng, Chujie, et al.
Published: (2024)
by: Zheng, Chujie, et al.
Published: (2024)
Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
by: Karamcheti, Siddharth, et al.
Published: (2024)
by: Karamcheti, Siddharth, et al.
Published: (2024)
A Critical Study of What Code-LLMs (Do Not) Learn
by: Anand, Abhinav, et al.
Published: (2024)
by: Anand, Abhinav, et al.
Published: (2024)
Similar Items
-
Self-Directed Synthetic Dialogues and Revisions Technical Report
by: Lambert, Nathan, et al.
Published: (2024) -
PERSONA: A Reproducible Testbed for Pluralistic Alignment
by: Castricato, Louis, et al.
Published: (2024) -
PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs
by: van der Wal, Oskar, et al.
Published: (2025) -
Negation: A Pink Elephant in the Large Language Models' Room?
by: Vrabcová, Tereza, et al.
Published: (2025) -
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
by: Schaeffer, Rylan, et al.
Published: (2024)