Prompting from the bench: Large-scale pretraining is not sufficient to prepare LLMs for ordinary meaning analysis
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Purushothama, Abhishek, Min, Junghyun, Waldon, Brandon, Schneider, Nathan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
BLT: Can Large Language Models Handle Basic Legal Text?
par: Blair-Stanek, Andrew, et autres
Publié: (2023)
par: Blair-Stanek, Andrew, et autres
Publié: (2023)
Computational analysis of the language of pain: a systematic review
par: Nunes, Diogo A. P., et autres
Publié: (2024)
par: Nunes, Diogo A. P., et autres
Publié: (2024)
LLMs Generate Kitsch
par: Klinge, Xenia, et autres
Publié: (2026)
par: Klinge, Xenia, et autres
Publié: (2026)
Scaling In, Not Up? Testing Thick Citation Context Analysis with GPT-5 and Fragile Prompts
par: Simons, Arno
Publié: (2026)
par: Simons, Arno
Publié: (2026)
LLMs Simulate Big Five Personality Traits: Further Evidence
par: Sorokovikova, Aleksandra, et autres
Publié: (2024)
par: Sorokovikova, Aleksandra, et autres
Publié: (2024)
Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization
par: Elganayni, Mohamed Hesham, et autres
Publié: (2026)
par: Elganayni, Mohamed Hesham, et autres
Publié: (2026)
The Judge Variable: Challenging Judge-Agnostic Legal Judgment Prediction
par: Zambrano, Guillaume
Publié: (2025)
par: Zambrano, Guillaume
Publié: (2025)
The Foundational Capabilities of Large Language Models in Predicting Postoperative Risks Using Clinical Notes
par: Alba, Charles, et autres
Publié: (2024)
par: Alba, Charles, et autres
Publié: (2024)
What distinguishes conspiracy from critical narratives? A computational analysis of oppositional discourse
par: Korenčić, Damir, et autres
Publié: (2024)
par: Korenčić, Damir, et autres
Publié: (2024)
Multi-Method Validation of Large Language Model Medical Translation Across High- and Low-Resource Languages
par: Anyaegbuna, Chukwuebuka, et autres
Publié: (2026)
par: Anyaegbuna, Chukwuebuka, et autres
Publié: (2026)
Can Large Language Models Imitate Human Speech for Clinical Assessment? LLM-Driven Data Augmentation for Cognitive Score Prediction
par: Ketir, Si-Belkacem Yamine, et autres
Publié: (2026)
par: Ketir, Si-Belkacem Yamine, et autres
Publié: (2026)
Comparative Study of Large Language Models on Chinese Film Script Continuation: An Empirical Analysis Based on GPT-5.2 and Qwen-Max
par: Cao, Yuxuan, et autres
Publié: (2026)
par: Cao, Yuxuan, et autres
Publié: (2026)
Survey and Experiments on Mental Disorder Detection via Social Media: From Large Language Models and RAG to Agents
par: Ge, Zhuohan, et autres
Publié: (2025)
par: Ge, Zhuohan, et autres
Publié: (2025)
From RAG to QA-RAG: Integrating Generative AI for Pharmaceutical Regulatory Compliance Process
par: Kim, Jaewoong, et autres
Publié: (2024)
par: Kim, Jaewoong, et autres
Publié: (2024)
OEMA: Ontology-Enhanced Multi-Agent Collaboration Framework for Zero-Shot Clinical Named Entity Recognition
par: Tao, Xinli, et autres
Publié: (2025)
par: Tao, Xinli, et autres
Publié: (2025)
Practical Design and Benchmarking of Generative AI Applications for Surgical Billing and Coding
par: Rollman, John C., et autres
Publié: (2025)
par: Rollman, John C., et autres
Publié: (2025)
ChemPro: A Progressive Chemistry Benchmark for Large Language Models
par: Baranwal, Aaditya, et autres
Publié: (2026)
par: Baranwal, Aaditya, et autres
Publié: (2026)
Natural Language Processing for the Legal Domain: A Survey of Tasks, Datasets, Models, and Challenges
par: Ariai, Farid, et autres
Publié: (2024)
par: Ariai, Farid, et autres
Publié: (2024)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
par: Cherif, Ahmed
Publié: (2026)
par: Cherif, Ahmed
Publié: (2026)
Counterfactual Causal Inference in Natural Language with Large Language Models
par: Gendron, Gaël, et autres
Publié: (2024)
par: Gendron, Gaël, et autres
Publié: (2024)
Unified Game Moderation: Soft-Prompting and LLM-Assisted Label Transfer for Resource-Efficient Toxicity Detection
par: Yang, Zachary, et autres
Publié: (2025)
par: Yang, Zachary, et autres
Publié: (2025)
How Much Does Persuasion Strategy Matter? LLM-Annotated Evidence from Charitable Donation Dialogues
par: Petrova, Tatiana, et autres
Publié: (2026)
par: Petrova, Tatiana, et autres
Publié: (2026)
Using Letter Positional Probabilities to Assess Word Complexity
par: Dalvean, Michael
Publié: (2024)
par: Dalvean, Michael
Publié: (2024)
The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning
par: Ahmed, Md Shamim, et autres
Publié: (2026)
par: Ahmed, Md Shamim, et autres
Publié: (2026)
BenCSSmark: Making the Social Sciences Count in LLM Research
par: Chatelain, Arnault, et autres
Publié: (2026)
par: Chatelain, Arnault, et autres
Publié: (2026)
Integrating clinical reasoning into large language model-based diagnosis through etiology-aware attention steering
par: Li, Peixian, et autres
Publié: (2025)
par: Li, Peixian, et autres
Publié: (2025)
The Representational Alignment between Humans and Language Models is implicitly driven by a Concreteness Effect
par: Iaia, Cosimo, et autres
Publié: (2025)
par: Iaia, Cosimo, et autres
Publié: (2025)
Cheap Learning: Maximising Performance of Language Models for Social Data Science Using Minimal Data
par: Castro-Gonzalez, Leonardo, et autres
Publié: (2024)
par: Castro-Gonzalez, Leonardo, et autres
Publié: (2024)
Forgotten Words: Benchmarking NeoBERT for Dementia Detection in Low-Resource Conversational Filipino and English Speech
par: Floresca, Rez Samantha Z., et autres
Publié: (2026)
par: Floresca, Rez Samantha Z., et autres
Publié: (2026)
Collective Memory and Narrative Cohesion: A Computational Study of Palestinian Refugee Oral Histories in Lebanon
par: Awwad, Ghadeer, et autres
Publié: (2025)
par: Awwad, Ghadeer, et autres
Publié: (2025)
Leveraging Synthetic Data for Question Answering with Multilingual LLMs in the Agricultural Domain
par: Kaur, Rishemjit, et autres
Publié: (2025)
par: Kaur, Rishemjit, et autres
Publié: (2025)
Large Language Models for History, Philosophy, and Sociology of Science: Interpretive Uses, Methodological Challenges, and Critical Perspectives
par: Simons, Arno, et autres
Publié: (2025)
par: Simons, Arno, et autres
Publié: (2025)
FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
par: Kytöniemi, Joona, et autres
Publié: (2025)
par: Kytöniemi, Joona, et autres
Publié: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
par: Saji, Alan, et autres
Publié: (2025)
par: Saji, Alan, et autres
Publié: (2025)
Do Political Opinions Transfer Between Western Languages? An Analysis of Unaligned and Aligned Multilingual LLMs
par: Weeber, Franziska, et autres
Publié: (2025)
par: Weeber, Franziska, et autres
Publié: (2025)
Multi-Hierarchical Feature Detection for Large Language Model Generated Text
par: Zhang, Luyan, et autres
Publié: (2025)
par: Zhang, Luyan, et autres
Publié: (2025)
Evaluating the Challenges of LLMs in Real-world Medical Follow-up: A Comparative Study and An Optimized Framework
par: Liu, Jinyan, et autres
Publié: (2025)
par: Liu, Jinyan, et autres
Publié: (2025)
Graphemic Normalization of the Perso-Arabic Script
par: Doctor, Raiomond, et autres
Publié: (2022)
par: Doctor, Raiomond, et autres
Publié: (2022)
Beyond Arabic: Software for Perso-Arabic Script Manipulation
par: Gutkin, Alexander, et autres
Publié: (2023)
par: Gutkin, Alexander, et autres
Publié: (2023)
Claim Automation using Large Language Model
par: Mo, Zhengda, et autres
Publié: (2026)
par: Mo, Zhengda, et autres
Publié: (2026)
Documents similaires
-
BLT: Can Large Language Models Handle Basic Legal Text?
par: Blair-Stanek, Andrew, et autres
Publié: (2023) -
Computational analysis of the language of pain: a systematic review
par: Nunes, Diogo A. P., et autres
Publié: (2024) -
LLMs Generate Kitsch
par: Klinge, Xenia, et autres
Publié: (2026) -
Scaling In, Not Up? Testing Thick Citation Context Analysis with GPT-5 and Fragile Prompts
par: Simons, Arno
Publié: (2026) -
LLMs Simulate Big Five Personality Traits: Further Evidence
par: Sorokovikova, Aleksandra, et autres
Publié: (2024)