Evaluating LLM-Based Grant Proposal Review via Structured Perturbations
Fuente:
arXiv
Saved in:
| Main Authors: | Thorne, William, James, Joseph, Wang, Yang, Lin, Chenghua, Maynard, Diana |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Demystifying Funding: Reconstructing a Unified Dataset of the UK Funding Lifecycle
by: Thorne, William, et al.
Published: (2026)
by: Thorne, William, et al.
Published: (2026)
Increasing the Difficulty of Automatically Generated Questions via Reinforcement Learning with Synthetic Preference
by: Thorne, William, et al.
Published: (2024)
by: Thorne, William, et al.
Published: (2024)
How BERT Speaks Shakespearean English? Evaluating Historical Bias in Contextual Language Models
by: Cuscito, Miriam, et al.
Published: (2024)
by: Cuscito, Miriam, et al.
Published: (2024)
Discursive objection strategies in online comments: Developing a classification schema and validating its training
by: Shea, Ashley L., et al.
Published: (2024)
by: Shea, Ashley L., et al.
Published: (2024)
Do Political Opinions Transfer Between Western Languages? An Analysis of Unaligned and Aligned Multilingual LLMs
by: Weeber, Franziska, et al.
Published: (2025)
by: Weeber, Franziska, et al.
Published: (2025)
Where is my Glass Slipper? AI, Poetry and Art
by: Pagiaslis, Anastasios P.
Published: (2025)
by: Pagiaslis, Anastasios P.
Published: (2025)
NLP Occupational Emergence Analysis: How Occupations Form and Evolve in Real Time -- A Zero-Assumption Method Demonstrated on AI in the US Technology Workforce, 2022-2026
by: Nordfors, David
Published: (2026)
by: Nordfors, David
Published: (2026)
ParliaBench: An Evaluation and Benchmarking Framework for LLM-Generated Parliamentary Speech
by: Koniaris, Marios, et al.
Published: (2025)
by: Koniaris, Marios, et al.
Published: (2025)
Evaluating LLM Prompts for Data Augmentation in Multi-label Classification of Ecological Texts
by: Glazkova, Anna, et al.
Published: (2024)
by: Glazkova, Anna, et al.
Published: (2024)
Reasoning Over the Glyphs: Evaluation of LLM's Decipherment of Rare Scripts
by: Shih, Yu-Fei, et al.
Published: (2025)
by: Shih, Yu-Fei, et al.
Published: (2025)
Modeling Changing Scientific Concepts with Complex Networks: A Case Study on the Chemical Revolution
by: Aguilar-Valdez, Sofía, et al.
Published: (2026)
by: Aguilar-Valdez, Sofía, et al.
Published: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
The Table of Media Bias Elements: A sentence-level taxonomy of media bias types and propaganda techniques
by: Menzner, Tim, et al.
Published: (2026)
by: Menzner, Tim, et al.
Published: (2026)
Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization
by: Elganayni, Mohamed Hesham, et al.
Published: (2026)
by: Elganayni, Mohamed Hesham, et al.
Published: (2026)
BenCSSmark: Making the Social Sciences Count in LLM Research
by: Chatelain, Arnault, et al.
Published: (2026)
by: Chatelain, Arnault, et al.
Published: (2026)
Comparative Study of Large Language Models on Chinese Film Script Continuation: An Empirical Analysis Based on GPT-5.2 and Qwen-Max
by: Cao, Yuxuan, et al.
Published: (2026)
by: Cao, Yuxuan, et al.
Published: (2026)
How Much Does Persuasion Strategy Matter? LLM-Annotated Evidence from Charitable Donation Dialogues
by: Petrova, Tatiana, et al.
Published: (2026)
by: Petrova, Tatiana, et al.
Published: (2026)
Can Large Language Models Imitate Human Speech for Clinical Assessment? LLM-Driven Data Augmentation for Cognitive Score Prediction
by: Ketir, Si-Belkacem Yamine, et al.
Published: (2026)
by: Ketir, Si-Belkacem Yamine, et al.
Published: (2026)
From Newswire to Nexus: Using text-based actor embeddings and transformer networks to forecast conflict dynamics
by: Croicu, Mihai, et al.
Published: (2025)
by: Croicu, Mihai, et al.
Published: (2025)
How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krügel, and Uhl (2025)
by: Himmelreich, Johannes
Published: (2026)
by: Himmelreich, Johannes
Published: (2026)
Towards Fairer Health Recommendations: finding informative unbiased samples via Word Sense Disambiguation
by: Butts, Gavin, et al.
Published: (2024)
by: Butts, Gavin, et al.
Published: (2024)
Fane at SemEval-2025 Task 10: Zero-Shot Entity Framing with Large Language Models
by: Fane, Enfa, et al.
Published: (2025)
by: Fane, Enfa, et al.
Published: (2025)
Towards Red Teaming in Multimodal and Multilingual Translation
by: Ropers, Christophe, et al.
Published: (2024)
by: Ropers, Christophe, et al.
Published: (2024)
How Well Do LLMs Imitate Human Writing Style?
by: Jemama, Rebira, et al.
Published: (2025)
by: Jemama, Rebira, et al.
Published: (2025)
Statistical Scouting Finds Debate-Safe but Not Debate-Useful Cases: A Matched-Ceiling Study of Open-Weight LLM Reasoning Protocols
by: Hu, Julia, et al.
Published: (2026)
by: Hu, Julia, et al.
Published: (2026)
Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters
by: Shah, Aaryan, et al.
Published: (2026)
by: Shah, Aaryan, et al.
Published: (2026)
Engineering A Large Language Model From Scratch
by: Oketunji, Abiodun Finbarrs
Published: (2024)
by: Oketunji, Abiodun Finbarrs
Published: (2024)
Survey and Experiments on Mental Disorder Detection via Social Media: From Large Language Models and RAG to Agents
by: Ge, Zhuohan, et al.
Published: (2025)
by: Ge, Zhuohan, et al.
Published: (2025)
On Fact and Frequency: LLM Responses to Misinformation Expressed with Uncertainty
by: van de Sande, Yana, et al.
Published: (2025)
by: van de Sande, Yana, et al.
Published: (2025)
SAGE: Hierarchical LLM-Based Literary Evaluation through Ontology-Grounded Interpretive Dimensions
by: Wang, Tianyu, et al.
Published: (2026)
by: Wang, Tianyu, et al.
Published: (2026)
Unified Game Moderation: Soft-Prompting and LLM-Assisted Label Transfer for Resource-Efficient Toxicity Detection
by: Yang, Zachary, et al.
Published: (2025)
by: Yang, Zachary, et al.
Published: (2025)
Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMs
by: Manczak, Blazej, et al.
Published: (2025)
by: Manczak, Blazej, et al.
Published: (2025)
The Foundational Capabilities of Large Language Models in Predicting Postoperative Risks Using Clinical Notes
by: Alba, Charles, et al.
Published: (2024)
by: Alba, Charles, et al.
Published: (2024)
Multi-Method Validation of Large Language Model Medical Translation Across High- and Low-Resource Languages
by: Anyaegbuna, Chukwuebuka, et al.
Published: (2026)
by: Anyaegbuna, Chukwuebuka, et al.
Published: (2026)
What distinguishes conspiracy from critical narratives? A computational analysis of oppositional discourse
by: Korenčić, Damir, et al.
Published: (2024)
by: Korenčić, Damir, et al.
Published: (2024)
Using Letter Positional Probabilities to Assess Word Complexity
by: Dalvean, Michael
Published: (2024)
by: Dalvean, Michael
Published: (2024)
Prompting from the bench: Large-scale pretraining is not sufficient to prepare LLMs for ordinary meaning analysis
by: Purushothama, Abhishek, et al.
Published: (2025)
by: Purushothama, Abhishek, et al.
Published: (2025)
The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning
by: Ahmed, Md Shamim, et al.
Published: (2026)
by: Ahmed, Md Shamim, et al.
Published: (2026)
LLMs Generate Kitsch
by: Klinge, Xenia, et al.
Published: (2026)
by: Klinge, Xenia, et al.
Published: (2026)
Integrating clinical reasoning into large language model-based diagnosis through etiology-aware attention steering
by: Li, Peixian, et al.
Published: (2025)
by: Li, Peixian, et al.
Published: (2025)
Similar Items
-
Demystifying Funding: Reconstructing a Unified Dataset of the UK Funding Lifecycle
by: Thorne, William, et al.
Published: (2026) -
Increasing the Difficulty of Automatically Generated Questions via Reinforcement Learning with Synthetic Preference
by: Thorne, William, et al.
Published: (2024) -
How BERT Speaks Shakespearean English? Evaluating Historical Bias in Contextual Language Models
by: Cuscito, Miriam, et al.
Published: (2024) -
Discursive objection strategies in online comments: Developing a classification schema and validating its training
by: Shea, Ashley L., et al.
Published: (2024) -
Do Political Opinions Transfer Between Western Languages? An Analysis of Unaligned and Aligned Multilingual LLMs
by: Weeber, Franziska, et al.
Published: (2025)