Genie: Achieving Human Parity in Content-Grounded Datasets Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yehudai, Asaf, Carmeli, Boaz, Mass, Yosi, Arviv, Ofir, Mills, Nathaniel, Toledo, Assaf, Shnarch, Eyal, Choshen, Leshem |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Will it Merge? On The Causes of Model Mergeability
par: Rahamim, Adir, et autres
Publié: (2026)
par: Rahamim, Adir, et autres
Publié: (2026)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
par: Perlitz, Yotam, et autres
Publié: (2024)
par: Perlitz, Yotam, et autres
Publié: (2024)
Efficient Benchmarking of Language Models
par: Perlitz, Yotam, et autres
Publié: (2023)
par: Perlitz, Yotam, et autres
Publié: (2023)
Mediocrity is the key for LLM as a Judge Anchor Selection
par: Don-Yehiya, Shachar, et autres
Publié: (2026)
par: Don-Yehiya, Shachar, et autres
Publié: (2026)
Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
par: Uzan, Omri, et autres
Publié: (2025)
par: Uzan, Omri, et autres
Publié: (2025)
Label-Efficient Model Selection for Text Generation
par: Ashury-Tahan, Shir, et autres
Publié: (2024)
par: Ashury-Tahan, Shir, et autres
Publié: (2024)
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
par: Ashury-Tahan, Shir, et autres
Publié: (2025)
par: Ashury-Tahan, Shir, et autres
Publié: (2025)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
par: Habba, Eliya, et autres
Publié: (2025)
par: Habba, Eliya, et autres
Publié: (2025)
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
par: Habba, Eliya, et autres
Publié: (2026)
par: Habba, Eliya, et autres
Publié: (2026)
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
par: Don-Yehiya, Shachar, et autres
Publié: (2024)
par: Don-Yehiya, Shachar, et autres
Publié: (2024)
Fine-Grained Detection of Context-Grounded Hallucinations Using LLMs
par: Peisakhovsky, Yehonatan, et autres
Publié: (2025)
par: Peisakhovsky, Yehonatan, et autres
Publié: (2025)
Teaching Values to Machines: Simulating Human-Like Behavior in LLMs
par: Yehudai, Asaf, et autres
Publié: (2026)
par: Yehudai, Asaf, et autres
Publié: (2026)
NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning
par: Schwartz, Eli, et autres
Publié: (2024)
par: Schwartz, Eli, et autres
Publié: (2024)
Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI
par: Bandel, Elron, et autres
Publié: (2024)
par: Bandel, Elron, et autres
Publié: (2024)
Can Gradient Descent Simulate Prompting?
par: Zhang, Eric, et autres
Publié: (2025)
par: Zhang, Eric, et autres
Publié: (2025)
When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many Classes
par: Yehudai, Asaf, et autres
Publié: (2024)
par: Yehudai, Asaf, et autres
Publié: (2024)
Naturally Occurring Feedback is Common, Extractable and Useful
par: Don-Yehiya, Shachar, et autres
Publié: (2024)
par: Don-Yehiya, Shachar, et autres
Publié: (2024)
A Hitchhiker's Guide to Scaling Law Estimation
par: Choshen, Leshem, et autres
Publié: (2024)
par: Choshen, Leshem, et autres
Publié: (2024)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
par: Zaman, Kerem, et autres
Publié: (2023)
par: Zaman, Kerem, et autres
Publié: (2023)
Concept-Best-Matching: Evaluating Compositionality in Emergent Communication
par: Carmeli, Boaz, et autres
Publié: (2024)
par: Carmeli, Boaz, et autres
Publié: (2024)
Instructions Shape Production of Language, not Processing
par: Waldis, Andreas, et autres
Publié: (2026)
par: Waldis, Andreas, et autres
Publié: (2026)
GeniL: A Multilingual Dataset on Generalizing Language
par: Davani, Aida Mostafazadeh, et autres
Publié: (2024)
par: Davani, Aida Mostafazadeh, et autres
Publié: (2024)
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
par: Din, Alexander Yom, et autres
Publié: (2023)
par: Din, Alexander Yom, et autres
Publié: (2023)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
par: Yadav, Prateek, et autres
Publié: (2023)
par: Yadav, Prateek, et autres
Publié: (2023)
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
par: Yehudai, Asaf, et autres
Publié: (2026)
par: Yehudai, Asaf, et autres
Publié: (2026)
Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine Translation
par: Iluz, Bar, et autres
Publié: (2024)
par: Iluz, Bar, et autres
Publié: (2024)
Pretraining Language Models for Diachronic Linguistic Change Discovery
par: Fittschen, Elisabeth, et autres
Publié: (2025)
par: Fittschen, Elisabeth, et autres
Publié: (2025)
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
par: Waldis, Andreas, et autres
Publié: (2024)
par: Waldis, Andreas, et autres
Publié: (2024)
A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns
par: Yehudai, Asaf, et autres
Publié: (2024)
par: Yehudai, Asaf, et autres
Publié: (2024)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
par: Ifergan, Maxim, et autres
Publié: (2024)
par: Ifergan, Maxim, et autres
Publié: (2024)
Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability
par: Akyürek, Afra Feyza, et autres
Publié: (2024)
par: Akyürek, Afra Feyza, et autres
Publié: (2024)
Selective Self-to-Supervised Fine-Tuning for Generalization in Large Language Models
par: Gupta, Sonam, et autres
Publié: (2025)
par: Gupta, Sonam, et autres
Publié: (2025)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
par: Hilel, Almog, et autres
Publié: (2025)
par: Hilel, Almog, et autres
Publié: (2025)
Do LLMs Benefit From Their Own Words?
par: Huang, Jenny Y., et autres
Publié: (2026)
par: Huang, Jenny Y., et autres
Publié: (2026)
WildIFEval: Instruction Following in the Wild
par: Lior, Gili, et autres
Publié: (2025)
par: Lior, Gili, et autres
Publié: (2025)
Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMs
par: German, Eyal, et autres
Publié: (2025)
par: German, Eyal, et autres
Publié: (2025)
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress
par: Proietti, Lorenzo, et autres
Publié: (2025)
par: Proietti, Lorenzo, et autres
Publié: (2025)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
par: Ashury-Tahan, Shir, et autres
Publié: (2026)
par: Ashury-Tahan, Shir, et autres
Publié: (2026)
Resolving Interference (RI): Disentangling Models for Improved Model Merging
par: Ramesh, Pratik, et autres
Publié: (2026)
par: Ramesh, Pratik, et autres
Publié: (2026)
LexGenie: Automated Generation of Structured Reports for European Court of Human Rights Case Law
par: Santosh, T. Y. S. S, et autres
Publié: (2025)
par: Santosh, T. Y. S. S, et autres
Publié: (2025)
Documents similaires
-
Will it Merge? On The Causes of Model Mergeability
par: Rahamim, Adir, et autres
Publié: (2026) -
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
par: Perlitz, Yotam, et autres
Publié: (2024) -
Efficient Benchmarking of Language Models
par: Perlitz, Yotam, et autres
Publié: (2023) -
Mediocrity is the key for LLM as a Judge Anchor Selection
par: Don-Yehiya, Shachar, et autres
Publié: (2026) -
Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
par: Uzan, Omri, et autres
Publié: (2025)