Better than Random: Reliable NLG Human Evaluation with Constrained Active Sampling
Fuente:
arXiv
Guardado en:
| Autores principales: | Ruan, Jie, Pu, Xiao, Gao, Mingqi, Wan, Xiaojun, Zhu, Yuesheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation
por: Ruan, Jie, et al.
Publicado: (2024)
por: Ruan, Jie, et al.
Publicado: (2024)
LLM-based NLG Evaluation: Current Status and Challenges
por: Gao, Mingqi, et al.
Publicado: (2024)
por: Gao, Mingqi, et al.
Publicado: (2024)
A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability
por: Hu, Xinyu, et al.
Publicado: (2025)
por: Hu, Xinyu, et al.
Publicado: (2025)
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
por: Gao, Mingqi, et al.
Publicado: (2024)
por: Gao, Mingqi, et al.
Publicado: (2024)
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
por: Chang, Jiayi, et al.
Publicado: (2025)
por: Chang, Jiayi, et al.
Publicado: (2025)
Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability
por: Hu, Xinyu, et al.
Publicado: (2024)
por: Hu, Xinyu, et al.
Publicado: (2024)
Are LLM-based Evaluators Confusing NLG Quality Criteria?
por: Hu, Xinyu, et al.
Publicado: (2024)
por: Hu, Xinyu, et al.
Publicado: (2024)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
por: Gao, Mingqi, et al.
Publicado: (2024)
por: Gao, Mingqi, et al.
Publicado: (2024)
Direct-Scoring NLG Evaluators Can Use Pairwise Comparisons Too
por: Lawrence, Logan, et al.
Publicado: (2025)
por: Lawrence, Logan, et al.
Publicado: (2025)
A Systematic Review of Data-to-Text NLG
por: Osuji, Chinonso Cynthia, et al.
Publicado: (2024)
por: Osuji, Chinonso Cynthia, et al.
Publicado: (2024)
ConFu: Contemplate the Future for Better Speculative Sampling
por: Qin, Zongyue, et al.
Publicado: (2026)
por: Qin, Zongyue, et al.
Publicado: (2026)
Does RoBERTa Perform Better than BERT in Continual Learning: An Attention Sink Perspective
por: Bai, Xueying, et al.
Publicado: (2024)
por: Bai, Xueying, et al.
Publicado: (2024)
Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model
por: Hong, Yuzhong, et al.
Publicado: (2024)
por: Hong, Yuzhong, et al.
Publicado: (2024)
Sample Smart, Not Hard: Correctness-First Decoding for Better Reasoning in LLMs
por: Li, Xueyan, et al.
Publicado: (2025)
por: Li, Xueyan, et al.
Publicado: (2025)
Repeat After Me: Transformers are Better than State Space Models at Copying
por: Jelassi, Samy, et al.
Publicado: (2024)
por: Jelassi, Samy, et al.
Publicado: (2024)
Large Language Models Are Active Critics in NLG Evaluation
por: Xu, Shuying, et al.
Publicado: (2024)
por: Xu, Shuying, et al.
Publicado: (2024)
Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning
por: Kim, Junseok, et al.
Publicado: (2026)
por: Kim, Junseok, et al.
Publicado: (2026)
Constrained Adaptive Rejection Sampling
por: Parys, Paweł, et al.
Publicado: (2025)
por: Parys, Paweł, et al.
Publicado: (2025)
TRA: Better Length Generalisation with Threshold Relative Attention
por: Opper, Mattia, et al.
Publicado: (2025)
por: Opper, Mattia, et al.
Publicado: (2025)
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models
por: Huo, Yifu, et al.
Publicado: (2026)
por: Huo, Yifu, et al.
Publicado: (2026)
Aligning the True Semantics: Constrained Decoupling and Distribution Sampling for Cross-Modal Alignment
por: Ma, Xiang, et al.
Publicado: (2026)
por: Ma, Xiang, et al.
Publicado: (2026)
$p1$: Better Prompt Optimization with Fewer Prompts
por: Gao, Zhaolin, et al.
Publicado: (2026)
por: Gao, Zhaolin, et al.
Publicado: (2026)
Enhancing Text Generation in Joint NLG/NLU Learning Through Curriculum Learning, Semi-Supervised Training, and Advanced Optimization Techniques
por: Shaik, Rahimanuddin, et al.
Publicado: (2024)
por: Shaik, Rahimanuddin, et al.
Publicado: (2024)
Stabilizing Policy Optimization via Logits Convexity
por: Chen, Hongzhan, et al.
Publicado: (2026)
por: Chen, Hongzhan, et al.
Publicado: (2026)
Towards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework
por: Arias, Esteban Garces, et al.
Publicado: (2024)
por: Arias, Esteban Garces, et al.
Publicado: (2024)
OptiMer: Optimal Distribution Vector Merging Is Better than Data Mixing for Continual Pre-Training
por: Song, Haiyue, et al.
Publicado: (2026)
por: Song, Haiyue, et al.
Publicado: (2026)
Reliability Under Randomness: An Empirical Analysis of Sparse and Dense Language Models Across Decoding Temperatures
por: Grover, Kabir
Publicado: (2026)
por: Grover, Kabir
Publicado: (2026)
Sentence-level Reward Model can Generalize Better for Aligning LLM from Human Preference
por: Qiu, Wenjie, et al.
Publicado: (2025)
por: Qiu, Wenjie, et al.
Publicado: (2025)
Layer-wise Importance Matters: Less Memory for Better Performance in Parameter-efficient Fine-tuning of Large Language Models
por: Yao, Kai, et al.
Publicado: (2024)
por: Yao, Kai, et al.
Publicado: (2024)
Deep Prompt Multi-task Network for Abuse Language Detection
por: Zhu, Jian, et al.
Publicado: (2024)
por: Zhu, Jian, et al.
Publicado: (2024)
Learning to Learn for Few-shot Continual Active Learning
por: Ho, Stella, et al.
Publicado: (2023)
por: Ho, Stella, et al.
Publicado: (2023)
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
por: Ma, Qiyao, et al.
Publicado: (2026)
por: Ma, Qiyao, et al.
Publicado: (2026)
Smaller Language Models are Better Black-box Machine-Generated Text Detectors
por: Mireshghallah, Niloofar, et al.
Publicado: (2023)
por: Mireshghallah, Niloofar, et al.
Publicado: (2023)
Better To Ask in English? Evaluating Factual Accuracy of Multilingual LLMs in English and Low-Resource Languages
por: Rohera, Pritika, et al.
Publicado: (2025)
por: Rohera, Pritika, et al.
Publicado: (2025)
PromptAL: Sample-Aware Dynamic Soft Prompts for Few-Shot Active Learning
por: Xiang, Hui, et al.
Publicado: (2025)
por: Xiang, Hui, et al.
Publicado: (2025)
Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition
por: Feng, Kehua, et al.
Publicado: (2024)
por: Feng, Kehua, et al.
Publicado: (2024)
Active Preference Optimization for Sample Efficient RLHF
por: Das, Nirjhar, et al.
Publicado: (2024)
por: Das, Nirjhar, et al.
Publicado: (2024)
Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt Optimization
por: Wan, Xingchen, et al.
Publicado: (2024)
por: Wan, Xingchen, et al.
Publicado: (2024)
Reliable Evaluation and Benchmarks for Statement Autoformalization
por: Poiroux, Auguste, et al.
Publicado: (2024)
por: Poiroux, Auguste, et al.
Publicado: (2024)
Data or Language Supervision: What Makes CLIP Better than DINO?
por: Liu, Yiming, et al.
Publicado: (2025)
por: Liu, Yiming, et al.
Publicado: (2025)
Ejemplares similares
-
Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation
por: Ruan, Jie, et al.
Publicado: (2024) -
LLM-based NLG Evaluation: Current Status and Challenges
por: Gao, Mingqi, et al.
Publicado: (2024) -
A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability
por: Hu, Xinyu, et al.
Publicado: (2025) -
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
por: Gao, Mingqi, et al.
Publicado: (2024) -
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
por: Chang, Jiayi, et al.
Publicado: (2025)