Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Samuel, Vinay, Chang, Yapei, Iyyer, Mohit |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BooookScore: A systematic exploration of book-length summarization in the era of LLMs
by: Chang, Yapei, et al.
Published: (2023)
by: Chang, Yapei, et al.
Published: (2023)
CLIPPER: Compression enables long-context synthetic data generation
by: Pham, Chau Minh, et al.
Published: (2025)
by: Pham, Chau Minh, et al.
Published: (2025)
Argument Collapse: LLMs Flatten Long-Form Public Debate
by: Kim, Yekyung, et al.
Published: (2026)
by: Kim, Yekyung, et al.
Published: (2026)
PostMark: A Robust Blackbox Watermark for Large Language Models
by: Chang, Yapei, et al.
Published: (2024)
by: Chang, Yapei, et al.
Published: (2024)
BEARCUBS: A benchmark for computer-using web agents
by: Song, Yixiao, et al.
Published: (2025)
by: Song, Yixiao, et al.
Published: (2025)
Literary Evidence Retrieval via Long-Context Language Models
by: Thai, Katherine, et al.
Published: (2025)
by: Thai, Katherine, et al.
Published: (2025)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
by: Chang, Yapei, et al.
Published: (2025)
by: Chang, Yapei, et al.
Published: (2025)
Slimming Down LLMs Without Losing Their Minds
by: Qingda, et al.
Published: (2025)
by: Qingda, et al.
Published: (2025)
FABLES: Evaluating faithfulness and content selection in book-length summarization
by: Kim, Yekyung, et al.
Published: (2024)
by: Kim, Yekyung, et al.
Published: (2024)
How2Everything: Mining the Web for How-To Procedures to Evaluate and Improve LLMs
by: Chang, Yapei, et al.
Published: (2026)
by: Chang, Yapei, et al.
Published: (2026)
Beyond Precision: Importance-Aware Recall for Factuality Evaluation in Long-Form LLM Generation
by: Jafari, Nazanin, et al.
Published: (2026)
by: Jafari, Nazanin, et al.
Published: (2026)
VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation
by: Song, Yixiao, et al.
Published: (2024)
by: Song, Yixiao, et al.
Published: (2024)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
by: Naseh, Ali, et al.
Published: (2024)
by: Naseh, Ali, et al.
Published: (2024)
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
by: Russell, Jenna, et al.
Published: (2025)
by: Russell, Jenna, et al.
Published: (2025)
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
by: Gu, Yuzhe, et al.
Published: (2025)
by: Gu, Yuzhe, et al.
Published: (2025)
Suri: Multi-constraint Instruction Following for Long-form Text Generation
by: Pham, Chau Minh, et al.
Published: (2024)
by: Pham, Chau Minh, et al.
Published: (2024)
One ruler to measure them all: Benchmarking multilingual long-context language models
by: Kim, Yekyung, et al.
Published: (2025)
by: Kim, Yekyung, et al.
Published: (2025)
EditLens: Quantifying the Extent of AI Editing in Text
by: Thai, Katherine, et al.
Published: (2025)
by: Thai, Katherine, et al.
Published: (2025)
Localizing and Mitigating Errors in Long-form Question Answering
by: Sachdeva, Rachneet, et al.
Published: (2024)
by: Sachdeva, Rachneet, et al.
Published: (2024)
PRISM: A Unified Framework for Post-Training LLMs Without Verifiable Rewards
by: Ghimire, Mukesh, et al.
Published: (2026)
by: Ghimire, Mukesh, et al.
Published: (2026)
Frankentext: Stitching random text fragments into long-form narratives
by: Pham, Chau Minh, et al.
Published: (2025)
by: Pham, Chau Minh, et al.
Published: (2025)
Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries
by: Malaviya, Chaitanya, et al.
Published: (2024)
by: Malaviya, Chaitanya, et al.
Published: (2024)
"Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions
by: Masala, Mihai, et al.
Published: (2024)
by: Masala, Mihai, et al.
Published: (2024)
Whose story is it? Personalizing story generation by inferring author styles
by: Kumar, Nischal Ashok, et al.
Published: (2025)
by: Kumar, Nischal Ashok, et al.
Published: (2025)
VeriFastScore: Speeding up long-form factuality evaluation
by: Rajendhran, Rishanth, et al.
Published: (2025)
by: Rajendhran, Rishanth, et al.
Published: (2025)
TopicGPT: A Prompt-based Topic Modeling Framework
by: Pham, Chau Minh, et al.
Published: (2023)
by: Pham, Chau Minh, et al.
Published: (2023)
Rethinking Local Learning: A Cheaper and Faster Recipe for LLM Post-Training
by: Shi, Hengyu, et al.
Published: (2026)
by: Shi, Hengyu, et al.
Published: (2026)
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
by: Hu, Mengxuan, et al.
Published: (2026)
by: Hu, Mengxuan, et al.
Published: (2026)
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
by: Xu, Shusheng, et al.
Published: (2024)
by: Xu, Shusheng, et al.
Published: (2024)
One Thousand and One Pairs: A "novel" challenge for long-context language models
by: Karpinska, Marzena, et al.
Published: (2024)
by: Karpinska, Marzena, et al.
Published: (2024)
StoryScope: Investigating idiosyncrasies in AI fiction
by: Russell, Jenna, et al.
Published: (2026)
by: Russell, Jenna, et al.
Published: (2026)
Cat-DPO: Category-Adaptive Safety Alignment
by: Yang, Tiankai, et al.
Published: (2026)
by: Yang, Tiankai, et al.
Published: (2026)
Does quantization affect models' performance on long-context tasks?
by: Mekala, Anmol, et al.
Published: (2025)
by: Mekala, Anmol, et al.
Published: (2025)
AlignTune: Modular Toolkit for Post-Training Alignment of Large Language Models
by: Lyngkhoi, R E Zera Marveen, et al.
Published: (2026)
by: Lyngkhoi, R E Zera Marveen, et al.
Published: (2026)
Thesis proposal: Are We Losing Textual Diversity to Natural Language Processing?
by: Jon, Josef
Published: (2024)
by: Jon, Josef
Published: (2024)
QwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory Management
by: Shen, Weizhou, et al.
Published: (2025)
by: Shen, Weizhou, et al.
Published: (2025)
Robust Multi-Objective Preference Alignment with Online DPO
by: Gupta, Raghav, et al.
Published: (2025)
by: Gupta, Raghav, et al.
Published: (2025)
Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
by: Saeidi, Amir, et al.
Published: (2024)
by: Saeidi, Amir, et al.
Published: (2024)
Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPO
by: Guo, Kaiyang, et al.
Published: (2025)
by: Guo, Kaiyang, et al.
Published: (2025)
Evaluating GRPO and DPO for Faithful Chain-of-Thought Reasoning in LLMs
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
Similar Items
-
BooookScore: A systematic exploration of book-length summarization in the era of LLMs
by: Chang, Yapei, et al.
Published: (2023) -
CLIPPER: Compression enables long-context synthetic data generation
by: Pham, Chau Minh, et al.
Published: (2025) -
Argument Collapse: LLMs Flatten Long-Form Public Debate
by: Kim, Yekyung, et al.
Published: (2026) -
PostMark: A Robust Blackbox Watermark for Large Language Models
by: Chang, Yapei, et al.
Published: (2024) -
BEARCUBS: A benchmark for computer-using web agents
by: Song, Yixiao, et al.
Published: (2025)