How to Get Your LLM to Generate Challenging Problems for Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Patel, Arkil, Reddy, Siva, Bahdanau, Dzmitry |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating In-Context Learning of Libraries for Code Generation
by: Patel, Arkil, et al.
Published: (2023)
by: Patel, Arkil, et al.
Published: (2023)
Forecasting Downstream Performance of LLMs With Proxy Metrics
by: Patel, Arkil, et al.
Published: (2026)
by: Patel, Arkil, et al.
Published: (2026)
Investigating Adversarial Trigger Transfer in Large Language Models
by: Meade, Nicholas, et al.
Published: (2024)
by: Meade, Nicholas, et al.
Published: (2024)
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
by: BehnamGhader, Parishad, et al.
Published: (2024)
by: BehnamGhader, Parishad, et al.
Published: (2024)
BRIDGE: Predicting Human Task Completion Time From Model Performance
by: Liu, Fengyuan, et al.
Published: (2026)
by: Liu, Fengyuan, et al.
Published: (2026)
LLMs can learn self-restraint through iterative self-reflection
by: Piché, Alexandre, et al.
Published: (2024)
by: Piché, Alexandre, et al.
Published: (2024)
NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild
by: Murty, Shikhar, et al.
Published: (2024)
by: Murty, Shikhar, et al.
Published: (2024)
SafeArena: Evaluating the Safety of Autonomous Web Agents
by: Tur, Ada Defne, et al.
Published: (2025)
by: Tur, Ada Defne, et al.
Published: (2025)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
by: Marjanović, Sara Vera, et al.
Published: (2025)
by: Marjanović, Sara Vera, et al.
Published: (2025)
LLM2Vec-Gen: Generative Embeddings from Large Language Models
by: BehnamGhader, Parishad, et al.
Published: (2026)
by: BehnamGhader, Parishad, et al.
Published: (2026)
Getting Your Indices in a Row: Full-Text Search for LLM Training Data for Real World
by: Marinas, Ines Altemir, et al.
Published: (2025)
by: Marinas, Ines Altemir, et al.
Published: (2025)
Not All Data Are Unlearned Equally
by: Krishnan, Aravind, et al.
Published: (2025)
by: Krishnan, Aravind, et al.
Published: (2025)
When does word order matter and when doesn't it?
by: Chen, Xuanda, et al.
Published: (2024)
by: Chen, Xuanda, et al.
Published: (2024)
Faithfulness Measurable Masked Language Models
by: Madsen, Andreas, et al.
Published: (2023)
by: Madsen, Andreas, et al.
Published: (2023)
The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents
by: Lu, Xing Han, et al.
Published: (2023)
by: Lu, Xing Han, et al.
Published: (2023)
A Compositional Typed Semantics for Universal Dependencies
by: Bradford, Laurestine, et al.
Published: (2024)
by: Bradford, Laurestine, et al.
Published: (2024)
Language Models Largely Exhibit Human-like Constituent Ordering Preferences
by: Tur, Ada Defne, et al.
Published: (2025)
by: Tur, Ada Defne, et al.
Published: (2025)
Exploiting Instruction-Following Retrievers for Malicious Information Retrieval
by: BehnamGhader, Parishad, et al.
Published: (2025)
by: BehnamGhader, Parishad, et al.
Published: (2025)
Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering
by: Adlakha, Vaibhav, et al.
Published: (2023)
by: Adlakha, Vaibhav, et al.
Published: (2023)
Are self-explanations from Large Language Models faithful?
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
Value Drifts: Tracing Value Alignment During LLM Post-Training
by: Bhatia, Mehar, et al.
Published: (2025)
by: Bhatia, Mehar, et al.
Published: (2025)
Who Gets the Reward, Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
by: Yang, Chih-Hsuan, et al.
Published: (2025)
by: Yang, Chih-Hsuan, et al.
Published: (2025)
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
by: Lù, Xing Han, et al.
Published: (2024)
by: Lù, Xing Han, et al.
Published: (2024)
LLMSQL: Upgrading WikiSQL for the LLM Era of Text-to-SQL
by: Pihulski, Dzmitry, et al.
Published: (2025)
by: Pihulski, Dzmitry, et al.
Published: (2025)
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
by: Ghahroodi, Omid, et al.
Published: (2024)
by: Ghahroodi, Omid, et al.
Published: (2024)
Scope Ambiguities in Large Language Models
by: Kamath, Gaurav, et al.
Published: (2024)
by: Kamath, Gaurav, et al.
Published: (2024)
Balancing Diversity and Risk in LLM Sampling: How to Select Your Method and Parameter for Open-Ended Text Generation
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit
by: Zhang, Boning, et al.
Published: (2024)
by: Zhang, Boning, et al.
Published: (2024)
Easy Problems That LLMs Get Wrong
by: Williams, Sean, et al.
Published: (2024)
by: Williams, Sean, et al.
Published: (2024)
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
by: Sieker, Judith, et al.
Published: (2026)
by: Sieker, Judith, et al.
Published: (2026)
StrategyLLM: Large Language Models as Strategy Generators, Executors, Optimizers, and Evaluators for Problem Solving
by: Gao, Chang, et al.
Published: (2023)
by: Gao, Chang, et al.
Published: (2023)
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
by: Kim, Sungwon, et al.
Published: (2025)
by: Kim, Sungwon, et al.
Published: (2025)
LLM-based NLG Evaluation: Current Status and Challenges
by: Gao, Mingqi, et al.
Published: (2024)
by: Gao, Mingqi, et al.
Published: (2024)
Unveiling LLM Evaluation Focused on Metrics: Challenges and Solutions
by: Hu, Taojun, et al.
Published: (2024)
by: Hu, Taojun, et al.
Published: (2024)
Teaching People LLM's Errors and Getting it Right
by: Stringham, Nathan, et al.
Published: (2025)
by: Stringham, Nathan, et al.
Published: (2025)
Understanding the Influence of Synthetic Data for Text Embedders
by: Springer, Jacob Mitchell, et al.
Published: (2025)
by: Springer, Jacob Mitchell, et al.
Published: (2025)
Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena
by: Chen, Jiangjie, et al.
Published: (2023)
by: Chen, Jiangjie, et al.
Published: (2023)
Build the web for agents, not agents for the web
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
Interpretability Needs a New Paradigm
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
Similar Items
-
Evaluating In-Context Learning of Libraries for Code Generation
by: Patel, Arkil, et al.
Published: (2023) -
Forecasting Downstream Performance of LLMs With Proxy Metrics
by: Patel, Arkil, et al.
Published: (2026) -
Investigating Adversarial Trigger Transfer in Large Language Models
by: Meade, Nicholas, et al.
Published: (2024) -
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
by: BehnamGhader, Parishad, et al.
Published: (2024) -
BRIDGE: Predicting Human Task Completion Time From Model Performance
by: Liu, Fengyuan, et al.
Published: (2026)