Evaluating Language Models as Synthetic Data Generators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Seungone, Suk, Juyoung, Yue, Xiang, Viswanathan, Vijay, Lee, Seongyun, Wang, Yizhong, Gashteovski, Kiril, Lawrence, Carolin, Welleck, Sean, Neubig, Graham |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Evaluation-time Compute with Reasoning Models as Evaluators
von: Kim, Seungone, et al.
Veröffentlicht: (2025)
von: Kim, Seungone, et al.
Veröffentlicht: (2025)
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
von: Lee, Seongyun, et al.
Veröffentlicht: (2025)
von: Lee, Seongyun, et al.
Veröffentlicht: (2025)
RefineBench: Evaluating Refinement Capability of Language Models via Checklists
von: Lee, Young-Jun, et al.
Veröffentlicht: (2025)
von: Lee, Young-Jun, et al.
Veröffentlicht: (2025)
Compositional Steering of Large Language Models with Steering Tokens
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)
Better Synthetic Data by Retrieving and Transforming Existing Datasets
von: Gandhi, Saumya, et al.
Veröffentlicht: (2024)
von: Gandhi, Saumya, et al.
Veröffentlicht: (2024)
Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
Aligning to Thousands of Preferences via System Message Generalization
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
SELF-GUIDE: Better Task-Specific Instruction Following via Self-Synthetic Finetuning
von: Zhao, Chenyang, et al.
Veröffentlicht: (2024)
von: Zhao, Chenyang, et al.
Veröffentlicht: (2024)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
Synthetic Multimodal Question Generation
von: Wu, Ian, et al.
Veröffentlicht: (2024)
von: Wu, Ian, et al.
Veröffentlicht: (2024)
Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design Decisions
von: Liu, Emmy, et al.
Veröffentlicht: (2025)
von: Liu, Emmy, et al.
Veröffentlicht: (2025)
ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
von: Xu, Yiming, et al.
Veröffentlicht: (2025)
von: Xu, Yiming, et al.
Veröffentlicht: (2025)
MEDDxAgent: A Unified Modular Agent Framework for Explainable Automatic Differential Diagnosis
von: Rose, Daniel, et al.
Veröffentlicht: (2025)
von: Rose, Daniel, et al.
Veröffentlicht: (2025)
AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents
von: Gioacchini, Luca, et al.
Veröffentlicht: (2024)
von: Gioacchini, Luca, et al.
Veröffentlicht: (2024)
Checklists Are Better Than Reward Models For Aligning Language Models
von: Viswanathan, Vijay, et al.
Veröffentlicht: (2025)
von: Viswanathan, Vijay, et al.
Veröffentlicht: (2025)
TextMineX: Data, Evaluation Framework and Ontology-guided LLM Pipeline for Humanitarian Mine Action
von: Zhou, Chenyue, et al.
Veröffentlicht: (2025)
von: Zhou, Chenyue, et al.
Veröffentlicht: (2025)
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
Training Task Experts through Retrieval Based Distillation
von: Ge, Jiaxin, et al.
Veröffentlicht: (2024)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2024)
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
von: Zhang, Charlie, et al.
Veröffentlicht: (2025)
von: Zhang, Charlie, et al.
Veröffentlicht: (2025)
From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models
von: Welleck, Sean, et al.
Veröffentlicht: (2024)
von: Welleck, Sean, et al.
Veröffentlicht: (2024)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
von: Huan, Maggie, et al.
Veröffentlicht: (2025)
von: Huan, Maggie, et al.
Veröffentlicht: (2025)
Can Language Models Evaluate Human Written Text? Case Study on Korean Student Writing for Education
von: Kim, Seungyoon, et al.
Veröffentlicht: (2024)
von: Kim, Seungyoon, et al.
Veröffentlicht: (2024)
Gym-Anything: Turn any Software into an Agent Environment
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
Leveraging Open Information Extraction for More Robust Domain Transfer of Event Trigger Detection
von: Dukić, David, et al.
Veröffentlicht: (2023)
von: Dukić, David, et al.
Veröffentlicht: (2023)
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
von: Agarwal, Anmol, et al.
Veröffentlicht: (2026)
von: Agarwal, Anmol, et al.
Veröffentlicht: (2026)
Better Instruction-Following Through Minimum Bayes Risk
von: Wu, Ian, et al.
Veröffentlicht: (2024)
von: Wu, Ian, et al.
Veröffentlicht: (2024)
LightPAL: Lightweight Passage Retrieval for Open Domain Multi-Document Summarization
von: Enomoto, Masafumi, et al.
Veröffentlicht: (2024)
von: Enomoto, Masafumi, et al.
Veröffentlicht: (2024)
Robust Text Classification: Analyzing Prototype-Based Networks
von: Sourati, Zhivar, et al.
Veröffentlicht: (2023)
von: Sourati, Zhivar, et al.
Veröffentlicht: (2023)
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2026)
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2026)
Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
von: Yue, Xiang, et al.
Veröffentlicht: (2024)
von: Yue, Xiang, et al.
Veröffentlicht: (2024)
Fine-grained Hallucination Detection and Editing for Language Models
von: Mishra, Abhika, et al.
Veröffentlicht: (2024)
von: Mishra, Abhika, et al.
Veröffentlicht: (2024)
Training Versatile Coding Agents in Synthetic Environments
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models
von: Tjuatja, Lindia, et al.
Veröffentlicht: (2025)
von: Tjuatja, Lindia, et al.
Veröffentlicht: (2025)
Training Proactive and Personalized LLM Agents
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
Optimizing Temperature for Language Models with Multi-Sample Inference
von: Du, Weihua, et al.
Veröffentlicht: (2025)
von: Du, Weihua, et al.
Veröffentlicht: (2025)
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
von: Son, Guijin, et al.
Veröffentlicht: (2024)
von: Son, Guijin, et al.
Veröffentlicht: (2024)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
von: Song, Yueqi, et al.
Veröffentlicht: (2025)
von: Song, Yueqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scaling Evaluation-time Compute with Reasoning Models as Evaluators
von: Kim, Seungone, et al.
Veröffentlicht: (2025) -
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
von: Kim, Seungone, et al.
Veröffentlicht: (2024) -
The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
von: Lee, Seongyun, et al.
Veröffentlicht: (2025) -
RefineBench: Evaluating Refinement Capability of Language Models via Checklists
von: Lee, Young-Jun, et al.
Veröffentlicht: (2025) -
Compositional Steering of Large Language Models with Steering Tokens
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)