BLEUBERI: BLEU is a surprisingly effective reward for instruction following
Fuente:
arXiv
Salvato in:
| Autori principali: | Chang, Yapei, Kim, Yekyung, Krumdick, Michael, Zadeh, Amir, Li, Chuan, Tanner, Chris, Iyyer, Mohit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Argument Collapse: LLMs Flatten Long-Form Public Debate
di: Kim, Yekyung, et al.
Pubblicazione: (2026)
di: Kim, Yekyung, et al.
Pubblicazione: (2026)
BooookScore: A systematic exploration of book-length summarization in the era of LLMs
di: Chang, Yapei, et al.
Pubblicazione: (2023)
di: Chang, Yapei, et al.
Pubblicazione: (2023)
PostMark: A Robust Blackbox Watermark for Large Language Models
di: Chang, Yapei, et al.
Pubblicazione: (2024)
di: Chang, Yapei, et al.
Pubblicazione: (2024)
BEARCUBS: A benchmark for computer-using web agents
di: Song, Yixiao, et al.
Pubblicazione: (2025)
di: Song, Yixiao, et al.
Pubblicazione: (2025)
FABLES: Evaluating faithfulness and content selection in book-length summarization
di: Kim, Yekyung, et al.
Pubblicazione: (2024)
di: Kim, Yekyung, et al.
Pubblicazione: (2024)
SignBLEU: Automatic Evaluation of Multi-channel Sign Language Translation
di: Kim, Jung-Ho, et al.
Pubblicazione: (2024)
di: Kim, Jung-Ho, et al.
Pubblicazione: (2024)
Enabling robots to follow abstract instructions and complete complex dynamic tasks
di: Mon-Williams, Ruaridh, et al.
Pubblicazione: (2024)
di: Mon-Williams, Ruaridh, et al.
Pubblicazione: (2024)
VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation
di: Song, Yixiao, et al.
Pubblicazione: (2024)
di: Song, Yixiao, et al.
Pubblicazione: (2024)
CaLMQA: Exploring culturally specific long-form question answering across 23 languages
di: Arora, Shane, et al.
Pubblicazione: (2024)
di: Arora, Shane, et al.
Pubblicazione: (2024)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
di: Liu, Shih-Yang, et al.
Pubblicazione: (2026)
di: Liu, Shih-Yang, et al.
Pubblicazione: (2026)
How2Everything: Mining the Web for How-To Procedures to Evaluate and Improve LLMs
di: Chang, Yapei, et al.
Pubblicazione: (2026)
di: Chang, Yapei, et al.
Pubblicazione: (2026)
Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs
di: Samuel, Vinay, et al.
Pubblicazione: (2026)
di: Samuel, Vinay, et al.
Pubblicazione: (2026)
One ruler to measure them all: Benchmarking multilingual long-context language models
di: Kim, Yekyung, et al.
Pubblicazione: (2025)
di: Kim, Yekyung, et al.
Pubblicazione: (2025)
CLIPPER: Compression enables long-context synthetic data generation
di: Pham, Chau Minh, et al.
Pubblicazione: (2025)
di: Pham, Chau Minh, et al.
Pubblicazione: (2025)
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
di: Russell, Jenna, et al.
Pubblicazione: (2025)
di: Russell, Jenna, et al.
Pubblicazione: (2025)
HelpSteer2: Open-source dataset for training top-performing reward models
di: Wang, Zhilin, et al.
Pubblicazione: (2024)
di: Wang, Zhilin, et al.
Pubblicazione: (2024)
DocFinQA: A Long-Context Financial Reasoning Dataset
di: Reddy, Varshini, et al.
Pubblicazione: (2024)
di: Reddy, Varshini, et al.
Pubblicazione: (2024)
Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
di: Bouchard, Dylan, et al.
Pubblicazione: (2025)
di: Bouchard, Dylan, et al.
Pubblicazione: (2025)
TensorBLEU: Vectorized GPU-based BLEU Score Implementation for Per-Sentence In-Training Evaluation
di: Filipek, Adam
Pubblicazione: (2025)
di: Filipek, Adam
Pubblicazione: (2025)
Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
di: Dong, Guanting, et al.
Pubblicazione: (2024)
di: Dong, Guanting, et al.
Pubblicazione: (2024)
Inducing Systematicity in Transformers by Attending to Structurally Quantized Embeddings
di: Jiang, Yichen, et al.
Pubblicazione: (2024)
di: Jiang, Yichen, et al.
Pubblicazione: (2024)
Revisiting the Superficial Alignment Hypothesis
di: Raghavendra, Mohit, et al.
Pubblicazione: (2024)
di: Raghavendra, Mohit, et al.
Pubblicazione: (2024)
Non-instructional Fine-tuning: Enabling Instruction-Following Capabilities in Pre-trained Language Models without Instruction-Following Data
di: Xie, Juncheng, et al.
Pubblicazione: (2024)
di: Xie, Juncheng, et al.
Pubblicazione: (2024)
ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment
di: Naous, Tarek, et al.
Pubblicazione: (2023)
di: Naous, Tarek, et al.
Pubblicazione: (2023)
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
di: Patil, Vaidehi, et al.
Pubblicazione: (2025)
di: Patil, Vaidehi, et al.
Pubblicazione: (2025)
AI-University: An LLM-based platform for instructional alignment to scientific classrooms
di: Shojaei, Mostafa Faghih, et al.
Pubblicazione: (2025)
di: Shojaei, Mostafa Faghih, et al.
Pubblicazione: (2025)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
di: Lin, Han, et al.
Pubblicazione: (2025)
di: Lin, Han, et al.
Pubblicazione: (2025)
Retro-BLEU: Quantifying Chemical Plausibility of Retrosynthesis Routes through Reaction Template Sequence Analysis
di: Li, Junren, et al.
Pubblicazione: (2023)
di: Li, Junren, et al.
Pubblicazione: (2023)
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2023)
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2023)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
di: Hase, Peter, et al.
Pubblicazione: (2024)
di: Hase, Peter, et al.
Pubblicazione: (2024)
VeriFastScore: Speeding up long-form factuality evaluation
di: Rajendhran, Rishanth, et al.
Pubblicazione: (2025)
di: Rajendhran, Rishanth, et al.
Pubblicazione: (2025)
Safety and accuracy follow different scaling laws in clinical large language models
di: Wind, Sebastian, et al.
Pubblicazione: (2026)
di: Wind, Sebastian, et al.
Pubblicazione: (2026)
Soft Self-Consistency Improves Language Model Agents
di: Wang, Han, et al.
Pubblicazione: (2024)
di: Wang, Han, et al.
Pubblicazione: (2024)
Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative Study
di: Bouchard, Dylan, et al.
Pubblicazione: (2026)
di: Bouchard, Dylan, et al.
Pubblicazione: (2026)
Multi-Attribute Steering of Language Models via Targeted Intervention
di: Nguyen, Duy, et al.
Pubblicazione: (2025)
di: Nguyen, Duy, et al.
Pubblicazione: (2025)
DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
di: Khan, Zaid, et al.
Pubblicazione: (2024)
di: Khan, Zaid, et al.
Pubblicazione: (2024)
RAVEN: In-Context Learning with Retrieval-Augmented Encoder-Decoder Language Models
di: Huang, Jie, et al.
Pubblicazione: (2023)
di: Huang, Jie, et al.
Pubblicazione: (2023)
Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification
di: Bouchard, Dylan, et al.
Pubblicazione: (2026)
di: Bouchard, Dylan, et al.
Pubblicazione: (2026)
Effective Reasoning Chains Reduce Intrinsic Dimensionality
di: Prasad, Archiki, et al.
Pubblicazione: (2026)
di: Prasad, Archiki, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Argument Collapse: LLMs Flatten Long-Form Public Debate
di: Kim, Yekyung, et al.
Pubblicazione: (2026) -
BooookScore: A systematic exploration of book-length summarization in the era of LLMs
di: Chang, Yapei, et al.
Pubblicazione: (2023) -
PostMark: A Robust Blackbox Watermark for Large Language Models
di: Chang, Yapei, et al.
Pubblicazione: (2024) -
BEARCUBS: A benchmark for computer-using web agents
di: Song, Yixiao, et al.
Pubblicazione: (2025) -
FABLES: Evaluating faithfulness and content selection in book-length summarization
di: Kim, Yekyung, et al.
Pubblicazione: (2024)