How to Get Your LLM to Generate Challenging Problems for Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Patel, Arkil, Reddy, Siva, Bahdanau, Dzmitry |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating In-Context Learning of Libraries for Code Generation
von: Patel, Arkil, et al.
Veröffentlicht: (2023)
von: Patel, Arkil, et al.
Veröffentlicht: (2023)
Forecasting Downstream Performance of LLMs With Proxy Metrics
von: Patel, Arkil, et al.
Veröffentlicht: (2026)
von: Patel, Arkil, et al.
Veröffentlicht: (2026)
Investigating Adversarial Trigger Transfer in Large Language Models
von: Meade, Nicholas, et al.
Veröffentlicht: (2024)
von: Meade, Nicholas, et al.
Veröffentlicht: (2024)
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2024)
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2024)
BRIDGE: Predicting Human Task Completion Time From Model Performance
von: Liu, Fengyuan, et al.
Veröffentlicht: (2026)
von: Liu, Fengyuan, et al.
Veröffentlicht: (2026)
LLMs can learn self-restraint through iterative self-reflection
von: Piché, Alexandre, et al.
Veröffentlicht: (2024)
von: Piché, Alexandre, et al.
Veröffentlicht: (2024)
NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild
von: Murty, Shikhar, et al.
Veröffentlicht: (2024)
von: Murty, Shikhar, et al.
Veröffentlicht: (2024)
SafeArena: Evaluating the Safety of Autonomous Web Agents
von: Tur, Ada Defne, et al.
Veröffentlicht: (2025)
von: Tur, Ada Defne, et al.
Veröffentlicht: (2025)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
von: Marjanović, Sara Vera, et al.
Veröffentlicht: (2025)
von: Marjanović, Sara Vera, et al.
Veröffentlicht: (2025)
LLM2Vec-Gen: Generative Embeddings from Large Language Models
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2026)
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2026)
Getting Your Indices in a Row: Full-Text Search for LLM Training Data for Real World
von: Marinas, Ines Altemir, et al.
Veröffentlicht: (2025)
von: Marinas, Ines Altemir, et al.
Veröffentlicht: (2025)
Not All Data Are Unlearned Equally
von: Krishnan, Aravind, et al.
Veröffentlicht: (2025)
von: Krishnan, Aravind, et al.
Veröffentlicht: (2025)
When does word order matter and when doesn't it?
von: Chen, Xuanda, et al.
Veröffentlicht: (2024)
von: Chen, Xuanda, et al.
Veröffentlicht: (2024)
Faithfulness Measurable Masked Language Models
von: Madsen, Andreas, et al.
Veröffentlicht: (2023)
von: Madsen, Andreas, et al.
Veröffentlicht: (2023)
The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents
von: Lu, Xing Han, et al.
Veröffentlicht: (2023)
von: Lu, Xing Han, et al.
Veröffentlicht: (2023)
A Compositional Typed Semantics for Universal Dependencies
von: Bradford, Laurestine, et al.
Veröffentlicht: (2024)
von: Bradford, Laurestine, et al.
Veröffentlicht: (2024)
Language Models Largely Exhibit Human-like Constituent Ordering Preferences
von: Tur, Ada Defne, et al.
Veröffentlicht: (2025)
von: Tur, Ada Defne, et al.
Veröffentlicht: (2025)
Exploiting Instruction-Following Retrievers for Malicious Information Retrieval
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2025)
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2025)
Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering
von: Adlakha, Vaibhav, et al.
Veröffentlicht: (2023)
von: Adlakha, Vaibhav, et al.
Veröffentlicht: (2023)
Are self-explanations from Large Language Models faithful?
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
Value Drifts: Tracing Value Alignment During LLM Post-Training
von: Bhatia, Mehar, et al.
Veröffentlicht: (2025)
von: Bhatia, Mehar, et al.
Veröffentlicht: (2025)
Who Gets the Reward, Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
von: Yang, Chih-Hsuan, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Hsuan, et al.
Veröffentlicht: (2025)
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
von: Lù, Xing Han, et al.
Veröffentlicht: (2024)
von: Lù, Xing Han, et al.
Veröffentlicht: (2024)
LLMSQL: Upgrading WikiSQL for the LLM Era of Text-to-SQL
von: Pihulski, Dzmitry, et al.
Veröffentlicht: (2025)
von: Pihulski, Dzmitry, et al.
Veröffentlicht: (2025)
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
von: Ghahroodi, Omid, et al.
Veröffentlicht: (2024)
von: Ghahroodi, Omid, et al.
Veröffentlicht: (2024)
Scope Ambiguities in Large Language Models
von: Kamath, Gaurav, et al.
Veröffentlicht: (2024)
von: Kamath, Gaurav, et al.
Veröffentlicht: (2024)
Balancing Diversity and Risk in LLM Sampling: How to Select Your Method and Parameter for Open-Ended Text Generation
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2024)
MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit
von: Zhang, Boning, et al.
Veröffentlicht: (2024)
von: Zhang, Boning, et al.
Veröffentlicht: (2024)
Easy Problems That LLMs Get Wrong
von: Williams, Sean, et al.
Veröffentlicht: (2024)
von: Williams, Sean, et al.
Veröffentlicht: (2024)
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
von: Sieker, Judith, et al.
Veröffentlicht: (2026)
von: Sieker, Judith, et al.
Veröffentlicht: (2026)
StrategyLLM: Large Language Models as Strategy Generators, Executors, Optimizers, and Evaluators for Problem Solving
von: Gao, Chang, et al.
Veröffentlicht: (2023)
von: Gao, Chang, et al.
Veröffentlicht: (2023)
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
von: Kim, Sungwon, et al.
Veröffentlicht: (2025)
von: Kim, Sungwon, et al.
Veröffentlicht: (2025)
LLM-based NLG Evaluation: Current Status and Challenges
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
Unveiling LLM Evaluation Focused on Metrics: Challenges and Solutions
von: Hu, Taojun, et al.
Veröffentlicht: (2024)
von: Hu, Taojun, et al.
Veröffentlicht: (2024)
Teaching People LLM's Errors and Getting it Right
von: Stringham, Nathan, et al.
Veröffentlicht: (2025)
von: Stringham, Nathan, et al.
Veröffentlicht: (2025)
Understanding the Influence of Synthetic Data for Text Embedders
von: Springer, Jacob Mitchell, et al.
Veröffentlicht: (2025)
von: Springer, Jacob Mitchell, et al.
Veröffentlicht: (2025)
Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena
von: Chen, Jiangjie, et al.
Veröffentlicht: (2023)
von: Chen, Jiangjie, et al.
Veröffentlicht: (2023)
Build the web for agents, not agents for the web
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
Interpretability Needs a New Paradigm
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Evaluating In-Context Learning of Libraries for Code Generation
von: Patel, Arkil, et al.
Veröffentlicht: (2023) -
Forecasting Downstream Performance of LLMs With Proxy Metrics
von: Patel, Arkil, et al.
Veröffentlicht: (2026) -
Investigating Adversarial Trigger Transfer in Large Language Models
von: Meade, Nicholas, et al.
Veröffentlicht: (2024) -
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2024) -
BRIDGE: Predicting Human Task Completion Time From Model Performance
von: Liu, Fengyuan, et al.
Veröffentlicht: (2026)