Accelerating Unbiased LLM Evaluation via Synthetic Feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Zhaoyi, Song, Yuda, Zanette, Andrea |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Training Language Models to Reason Efficiently
von: Arora, Daman, et al.
Veröffentlicht: (2025)
von: Arora, Daman, et al.
Veröffentlicht: (2025)
Outcome-based Exploration for LLM Reasoning
von: Song, Yuda, et al.
Veröffentlicht: (2025)
von: Song, Yuda, et al.
Veröffentlicht: (2025)
ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
von: Zhou, Yifei, et al.
Veröffentlicht: (2024)
von: Zhou, Yifei, et al.
Veröffentlicht: (2024)
Mitigating the Language Mismatch and Repetition Issues in LLM-based Machine Translation via Model Editing
von: Wang, Weichuan, et al.
Veröffentlicht: (2024)
von: Wang, Weichuan, et al.
Veröffentlicht: (2024)
Expanding the Capabilities of Reinforcement Learning via Text Feedback
von: Song, Yuda, et al.
Veröffentlicht: (2026)
von: Song, Yuda, et al.
Veröffentlicht: (2026)
Shrinking the Variance: Shrinkage Baselines for Reinforcement Learning with Verifiable Rewards
von: Zeng, Guanning, et al.
Veröffentlicht: (2025)
von: Zeng, Guanning, et al.
Veröffentlicht: (2025)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
von: Song, Dinghong, et al.
Veröffentlicht: (2025)
von: Song, Dinghong, et al.
Veröffentlicht: (2025)
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
von: Chen, Tianhao, et al.
Veröffentlicht: (2025)
von: Chen, Tianhao, et al.
Veröffentlicht: (2025)
The Importance of Online Data: Understanding Preference Fine-tuning via Coverage
von: Song, Yuda, et al.
Veröffentlicht: (2024)
von: Song, Yuda, et al.
Veröffentlicht: (2024)
SIPDO: Closed-Loop Prompt Optimization via Synthetic Data Feedback
von: Yu, Yaoning, et al.
Veröffentlicht: (2025)
von: Yu, Yaoning, et al.
Veröffentlicht: (2025)
Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
von: Ishibashi, Yoichi, et al.
Veröffentlicht: (2025)
von: Ishibashi, Yoichi, et al.
Veröffentlicht: (2025)
InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models
von: Gu, Yanggan, et al.
Veröffentlicht: (2025)
von: Gu, Yanggan, et al.
Veröffentlicht: (2025)
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
von: Song, Yuda, et al.
Veröffentlicht: (2024)
von: Song, Yuda, et al.
Veröffentlicht: (2024)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
von: Liu, Di, et al.
Veröffentlicht: (2024)
von: Liu, Di, et al.
Veröffentlicht: (2024)
Reinforcing Human Behavior Simulation via Verbal Feedback
von: Sun, Weiwei, et al.
Veröffentlicht: (2026)
von: Sun, Weiwei, et al.
Veröffentlicht: (2026)
Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language Models
von: Li, Zhaoyi, et al.
Veröffentlicht: (2026)
von: Li, Zhaoyi, et al.
Veröffentlicht: (2026)
Evaluating Defences against Unsafe Feedback in RLHF
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
GLEAN: Active Generalized Category Discovery with Diverse LLM Feedback
von: Zou, Henry Peng, et al.
Veröffentlicht: (2025)
von: Zou, Henry Peng, et al.
Veröffentlicht: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
Less Finetuning, Better Retrieval: Rethinking LLM Adaptation for Biomedical Retrievers via Synthetic Data and Model Merging
von: Khattab, Sameh, et al.
Veröffentlicht: (2026)
von: Khattab, Sameh, et al.
Veröffentlicht: (2026)
AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization
von: Zhang, Genghan, et al.
Veröffentlicht: (2025)
von: Zhang, Genghan, et al.
Veröffentlicht: (2025)
User eXperience Perception Insights Dataset (UXPID): Synthetic User Feedback from Public Industrial Forums
von: Kulyabin, Mikhail, et al.
Veröffentlicht: (2025)
von: Kulyabin, Mikhail, et al.
Veröffentlicht: (2025)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
MedSyn: LLM-based Synthetic Medical Text Generation Framework
von: Kumichev, Gleb, et al.
Veröffentlicht: (2024)
von: Kumichev, Gleb, et al.
Veröffentlicht: (2024)
Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
NodeSynth: Socially Aligned Synthetic Data for AI Evaluation
von: Rashid, Qazi Mamunur, et al.
Veröffentlicht: (2026)
von: Rashid, Qazi Mamunur, et al.
Veröffentlicht: (2026)
Explaining Length Bias in LLM-Based Preference Evaluations
von: Hu, Zhengyu, et al.
Veröffentlicht: (2024)
von: Hu, Zhengyu, et al.
Veröffentlicht: (2024)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
The Power of LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions
von: Wagner, Stefan Sylvius, et al.
Veröffentlicht: (2024)
von: Wagner, Stefan Sylvius, et al.
Veröffentlicht: (2024)
CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
von: McDanel, Bradley, et al.
Veröffentlicht: (2026)
von: McDanel, Bradley, et al.
Veröffentlicht: (2026)
ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
von: You, Haoran, et al.
Veröffentlicht: (2024)
von: You, Haoran, et al.
Veröffentlicht: (2024)
TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluation
von: Ouyang, Jialin
Veröffentlicht: (2025)
von: Ouyang, Jialin
Veröffentlicht: (2025)
$C^2$: Scalable Auto-Feedback for LLM-based Chart Generation
von: Koh, Woosung, et al.
Veröffentlicht: (2024)
von: Koh, Woosung, et al.
Veröffentlicht: (2024)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
von: Cai, Tianle, et al.
Veröffentlicht: (2024)
von: Cai, Tianle, et al.
Veröffentlicht: (2024)
RLTHF: Targeted Human Feedback for LLM Alignment
von: Xu, Yifei, et al.
Veröffentlicht: (2025)
von: Xu, Yifei, et al.
Veröffentlicht: (2025)
Reinforcement Learning from Denoising Feedback
von: He, Qi, et al.
Veröffentlicht: (2026)
von: He, Qi, et al.
Veröffentlicht: (2026)
MALTO at SemEval-2024 Task 6: Leveraging Synthetic Data for LLM Hallucination Detection
von: Borra, Federico, et al.
Veröffentlicht: (2024)
von: Borra, Federico, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Training Language Models to Reason Efficiently
von: Arora, Daman, et al.
Veröffentlicht: (2025) -
Outcome-based Exploration for LLM Reasoning
von: Song, Yuda, et al.
Veröffentlicht: (2025) -
ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
von: Zhou, Yifei, et al.
Veröffentlicht: (2024) -
Mitigating the Language Mismatch and Repetition Issues in LLM-based Machine Translation via Model Editing
von: Wang, Weichuan, et al.
Veröffentlicht: (2024) -
Expanding the Capabilities of Reinforcement Learning via Text Feedback
von: Song, Yuda, et al.
Veröffentlicht: (2026)