EvalAgent: Discovering Implicit Evaluation Criteria from the Web
Fuente:
arXiv
Salvato in:
| Autori principali: | Wadhwa, Manya, Sprague, Zayne, Malaviya, Chaitanya, Laban, Philippe, Li, Junyi Jessy, Durrett, Greg |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Using Natural Language Explanations to Rescale Human Judgments
di: Wadhwa, Manya, et al.
Pubblicazione: (2023)
di: Wadhwa, Manya, et al.
Pubblicazione: (2023)
Learning to Refine with Fine-Grained Natural Language Feedback
di: Wadhwa, Manya, et al.
Pubblicazione: (2024)
di: Wadhwa, Manya, et al.
Pubblicazione: (2024)
SkillFactory: Self-Distillation For Learning Cognitive Behaviors
di: Sprague, Zayne, et al.
Pubblicazione: (2025)
di: Sprague, Zayne, et al.
Pubblicazione: (2025)
CREATE: Testing LLMs for Associative Creativity
di: Wadhwa, Manya, et al.
Pubblicazione: (2026)
di: Wadhwa, Manya, et al.
Pubblicazione: (2026)
QUDsim: Quantifying Discourse Similarities in LLM-Generated Text
di: Namuduri, Ramya, et al.
Pubblicazione: (2025)
di: Namuduri, Ramya, et al.
Pubblicazione: (2025)
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
di: Sprague, Zayne, et al.
Pubblicazione: (2023)
di: Sprague, Zayne, et al.
Pubblicazione: (2023)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
di: Sprague, Zayne, et al.
Pubblicazione: (2024)
di: Sprague, Zayne, et al.
Pubblicazione: (2024)
MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
di: Tang, Liyan, et al.
Pubblicazione: (2024)
di: Tang, Liyan, et al.
Pubblicazione: (2024)
Which questions should I answer? Salience Prediction of Inquisitive Questions
di: Wu, Yating, et al.
Pubblicazione: (2024)
di: Wu, Yating, et al.
Pubblicazione: (2024)
ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
di: Tang, Liyan, et al.
Pubblicazione: (2025)
di: Tang, Liyan, et al.
Pubblicazione: (2025)
Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation
di: Tripathi, Tuhina, et al.
Pubblicazione: (2025)
di: Tripathi, Tuhina, et al.
Pubblicazione: (2025)
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
di: Yoran, Ori, et al.
Pubblicazione: (2024)
di: Yoran, Ori, et al.
Pubblicazione: (2024)
A Simple Joint Model for Improved Contextual Neural Lemmatization
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2019)
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2019)
ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
di: Yifei, Li S., et al.
Pubblicazione: (2025)
di: Yifei, Li S., et al.
Pubblicazione: (2025)
Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents
di: Ding, Wenxuan, et al.
Pubblicazione: (2026)
di: Ding, Wenxuan, et al.
Pubblicazione: (2026)
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
di: Gunjal, Anisha, et al.
Pubblicazione: (2024)
di: Gunjal, Anisha, et al.
Pubblicazione: (2024)
Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
di: Bharadwaj, Anirudh, et al.
Pubblicazione: (2025)
di: Bharadwaj, Anirudh, et al.
Pubblicazione: (2025)
What if you said that differently?: How Explanation Formats Affect Human Feedback Efficacy and User Perception
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2023)
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2023)
SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation
di: Divekar, Abhishek, et al.
Pubblicazione: (2024)
di: Divekar, Abhishek, et al.
Pubblicazione: (2024)
Understanding Synthetic Context Extension via Retrieval Heads
di: Zhao, Xinyu, et al.
Pubblicazione: (2024)
di: Zhao, Xinyu, et al.
Pubblicazione: (2024)
LoFiT: Localized Fine-tuning on LLM Representations
di: Yin, Fangcong, et al.
Pubblicazione: (2024)
di: Yin, Fangcong, et al.
Pubblicazione: (2024)
Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2024)
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2024)
D2PO: Discriminator-Guided DPO with Response Evaluation Models
di: Singhal, Prasann, et al.
Pubblicazione: (2024)
di: Singhal, Prasann, et al.
Pubblicazione: (2024)
X-PARADE: Cross-Lingual Textual Entailment and Information Divergence across Paragraphs
di: Rodriguez, Juan Diego, et al.
Pubblicazione: (2023)
di: Rodriguez, Juan Diego, et al.
Pubblicazione: (2023)
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
di: Lake, Thom, et al.
Pubblicazione: (2024)
di: Lake, Thom, et al.
Pubblicazione: (2024)
Language Models (Mostly) Do Not Consider Emotion Triggers When Predicting Emotion
di: Singh, Smriti, et al.
Pubblicazione: (2023)
di: Singh, Smriti, et al.
Pubblicazione: (2023)
Flipping the Dialogue: Training and Evaluating User Language Models
di: Naous, Tarek, et al.
Pubblicazione: (2025)
di: Naous, Tarek, et al.
Pubblicazione: (2025)
AlphaEval: Evaluating Agents in Production
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
WUGNECTIVES: Novel Entity Inferences of Language Models from Discourse Connectives
di: Brubaker, Daniel, et al.
Pubblicazione: (2025)
di: Brubaker, Daniel, et al.
Pubblicazione: (2025)
VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean
di: Xin, Yutong, et al.
Pubblicazione: (2026)
di: Xin, Yutong, et al.
Pubblicazione: (2026)
RankAlign: A Ranking View of the Generator-Validator Gap in Large Language Models
di: Rodriguez, Juan Diego, et al.
Pubblicazione: (2025)
di: Rodriguez, Juan Diego, et al.
Pubblicazione: (2025)
HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition
di: Liu, Yuxuan, et al.
Pubblicazione: (2024)
di: Liu, Yuxuan, et al.
Pubblicazione: (2024)
Behavioral Analysis of Information Salience in Large Language Models
di: Trienes, Jan, et al.
Pubblicazione: (2025)
di: Trienes, Jan, et al.
Pubblicazione: (2025)
Strategic Dialogue Assessment: The Crooked Path to Innocence
di: Zheng, Anshun Asher, et al.
Pubblicazione: (2025)
di: Zheng, Anshun Asher, et al.
Pubblicazione: (2025)
A Long Way to Go: Investigating Length Correlations in RLHF
di: Singhal, Prasann, et al.
Pubblicazione: (2023)
di: Singhal, Prasann, et al.
Pubblicazione: (2023)
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
di: Liu, Zeyu Leo, et al.
Pubblicazione: (2025)
di: Liu, Zeyu Leo, et al.
Pubblicazione: (2025)
ExpertQA: Expert-Curated Questions and Attributed Answers
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2023)
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2023)
LLMs Corrupt Your Documents When You Delegate
di: Laban, Philippe, et al.
Pubblicazione: (2026)
di: Laban, Philippe, et al.
Pubblicazione: (2026)
On Reference (In-)Determinacy in Natural Language Inference
di: Chen, Sihao, et al.
Pubblicazione: (2025)
di: Chen, Sihao, et al.
Pubblicazione: (2025)
SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat
di: Jiang, Yuru, et al.
Pubblicazione: (2025)
di: Jiang, Yuru, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Using Natural Language Explanations to Rescale Human Judgments
di: Wadhwa, Manya, et al.
Pubblicazione: (2023) -
Learning to Refine with Fine-Grained Natural Language Feedback
di: Wadhwa, Manya, et al.
Pubblicazione: (2024) -
SkillFactory: Self-Distillation For Learning Cognitive Behaviors
di: Sprague, Zayne, et al.
Pubblicazione: (2025) -
CREATE: Testing LLMs for Associative Creativity
di: Wadhwa, Manya, et al.
Pubblicazione: (2026) -
QUDsim: Quantifying Discourse Similarities in LLM-Generated Text
di: Namuduri, Ramya, et al.
Pubblicazione: (2025)