BLEUBERI: BLEU is a surprisingly effective reward for instruction following
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chang, Yapei, Kim, Yekyung, Krumdick, Michael, Zadeh, Amir, Li, Chuan, Tanner, Chris, Iyyer, Mohit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Argument Collapse: LLMs Flatten Long-Form Public Debate
von: Kim, Yekyung, et al.
Veröffentlicht: (2026)
von: Kim, Yekyung, et al.
Veröffentlicht: (2026)
BooookScore: A systematic exploration of book-length summarization in the era of LLMs
von: Chang, Yapei, et al.
Veröffentlicht: (2023)
von: Chang, Yapei, et al.
Veröffentlicht: (2023)
PostMark: A Robust Blackbox Watermark for Large Language Models
von: Chang, Yapei, et al.
Veröffentlicht: (2024)
von: Chang, Yapei, et al.
Veröffentlicht: (2024)
BEARCUBS: A benchmark for computer-using web agents
von: Song, Yixiao, et al.
Veröffentlicht: (2025)
von: Song, Yixiao, et al.
Veröffentlicht: (2025)
FABLES: Evaluating faithfulness and content selection in book-length summarization
von: Kim, Yekyung, et al.
Veröffentlicht: (2024)
von: Kim, Yekyung, et al.
Veröffentlicht: (2024)
SignBLEU: Automatic Evaluation of Multi-channel Sign Language Translation
von: Kim, Jung-Ho, et al.
Veröffentlicht: (2024)
von: Kim, Jung-Ho, et al.
Veröffentlicht: (2024)
Enabling robots to follow abstract instructions and complete complex dynamic tasks
von: Mon-Williams, Ruaridh, et al.
Veröffentlicht: (2024)
von: Mon-Williams, Ruaridh, et al.
Veröffentlicht: (2024)
VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation
von: Song, Yixiao, et al.
Veröffentlicht: (2024)
von: Song, Yixiao, et al.
Veröffentlicht: (2024)
CaLMQA: Exploring culturally specific long-form question answering across 23 languages
von: Arora, Shane, et al.
Veröffentlicht: (2024)
von: Arora, Shane, et al.
Veröffentlicht: (2024)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2026)
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2026)
How2Everything: Mining the Web for How-To Procedures to Evaluate and Improve LLMs
von: Chang, Yapei, et al.
Veröffentlicht: (2026)
von: Chang, Yapei, et al.
Veröffentlicht: (2026)
Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs
von: Samuel, Vinay, et al.
Veröffentlicht: (2026)
von: Samuel, Vinay, et al.
Veröffentlicht: (2026)
One ruler to measure them all: Benchmarking multilingual long-context language models
von: Kim, Yekyung, et al.
Veröffentlicht: (2025)
von: Kim, Yekyung, et al.
Veröffentlicht: (2025)
CLIPPER: Compression enables long-context synthetic data generation
von: Pham, Chau Minh, et al.
Veröffentlicht: (2025)
von: Pham, Chau Minh, et al.
Veröffentlicht: (2025)
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
von: Russell, Jenna, et al.
Veröffentlicht: (2025)
von: Russell, Jenna, et al.
Veröffentlicht: (2025)
HelpSteer2: Open-source dataset for training top-performing reward models
von: Wang, Zhilin, et al.
Veröffentlicht: (2024)
von: Wang, Zhilin, et al.
Veröffentlicht: (2024)
DocFinQA: A Long-Context Financial Reasoning Dataset
von: Reddy, Varshini, et al.
Veröffentlicht: (2024)
von: Reddy, Varshini, et al.
Veröffentlicht: (2024)
Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
von: Bouchard, Dylan, et al.
Veröffentlicht: (2025)
von: Bouchard, Dylan, et al.
Veröffentlicht: (2025)
TensorBLEU: Vectorized GPU-based BLEU Score Implementation for Per-Sentence In-Training Evaluation
von: Filipek, Adam
Veröffentlicht: (2025)
von: Filipek, Adam
Veröffentlicht: (2025)
Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
Inducing Systematicity in Transformers by Attending to Structurally Quantized Embeddings
von: Jiang, Yichen, et al.
Veröffentlicht: (2024)
von: Jiang, Yichen, et al.
Veröffentlicht: (2024)
Revisiting the Superficial Alignment Hypothesis
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2024)
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2024)
Non-instructional Fine-tuning: Enabling Instruction-Following Capabilities in Pre-trained Language Models without Instruction-Following Data
von: Xie, Juncheng, et al.
Veröffentlicht: (2024)
von: Xie, Juncheng, et al.
Veröffentlicht: (2024)
ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment
von: Naous, Tarek, et al.
Veröffentlicht: (2023)
von: Naous, Tarek, et al.
Veröffentlicht: (2023)
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
AI-University: An LLM-based platform for instructional alignment to scientific classrooms
von: Shojaei, Mostafa Faghih, et al.
Veröffentlicht: (2025)
von: Shojaei, Mostafa Faghih, et al.
Veröffentlicht: (2025)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
von: Lin, Han, et al.
Veröffentlicht: (2025)
von: Lin, Han, et al.
Veröffentlicht: (2025)
Retro-BLEU: Quantifying Chemical Plausibility of Retrosynthesis Routes through Reaction Template Sequence Analysis
von: Li, Junren, et al.
Veröffentlicht: (2023)
von: Li, Junren, et al.
Veröffentlicht: (2023)
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2023)
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2023)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
von: Yadav, Prateek, et al.
Veröffentlicht: (2023)
von: Yadav, Prateek, et al.
Veröffentlicht: (2023)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
von: Hase, Peter, et al.
Veröffentlicht: (2024)
von: Hase, Peter, et al.
Veröffentlicht: (2024)
VeriFastScore: Speeding up long-form factuality evaluation
von: Rajendhran, Rishanth, et al.
Veröffentlicht: (2025)
von: Rajendhran, Rishanth, et al.
Veröffentlicht: (2025)
Safety and accuracy follow different scaling laws in clinical large language models
von: Wind, Sebastian, et al.
Veröffentlicht: (2026)
von: Wind, Sebastian, et al.
Veröffentlicht: (2026)
Soft Self-Consistency Improves Language Model Agents
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative Study
von: Bouchard, Dylan, et al.
Veröffentlicht: (2026)
von: Bouchard, Dylan, et al.
Veröffentlicht: (2026)
Multi-Attribute Steering of Language Models via Targeted Intervention
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
von: Khan, Zaid, et al.
Veröffentlicht: (2024)
von: Khan, Zaid, et al.
Veröffentlicht: (2024)
RAVEN: In-Context Learning with Retrieval-Augmented Encoder-Decoder Language Models
von: Huang, Jie, et al.
Veröffentlicht: (2023)
von: Huang, Jie, et al.
Veröffentlicht: (2023)
Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification
von: Bouchard, Dylan, et al.
Veröffentlicht: (2026)
von: Bouchard, Dylan, et al.
Veröffentlicht: (2026)
Effective Reasoning Chains Reduce Intrinsic Dimensionality
von: Prasad, Archiki, et al.
Veröffentlicht: (2026)
von: Prasad, Archiki, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Argument Collapse: LLMs Flatten Long-Form Public Debate
von: Kim, Yekyung, et al.
Veröffentlicht: (2026) -
BooookScore: A systematic exploration of book-length summarization in the era of LLMs
von: Chang, Yapei, et al.
Veröffentlicht: (2023) -
PostMark: A Robust Blackbox Watermark for Large Language Models
von: Chang, Yapei, et al.
Veröffentlicht: (2024) -
BEARCUBS: A benchmark for computer-using web agents
von: Song, Yixiao, et al.
Veröffentlicht: (2025) -
FABLES: Evaluating faithfulness and content selection in book-length summarization
von: Kim, Yekyung, et al.
Veröffentlicht: (2024)