Can Open-Source LLMs Compete with Commercial Models? Exploring the Few-Shot Performance of Current GPT Models in Biomedical Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Ateia, Samy, Kruschwitz, Udo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Language Models Critique Themselves? Investigating Self-Feedback for Retrieval Augmented Generation at BioASQ 2025
by: Ateia, Samy, et al.
Published: (2025)
by: Ateia, Samy, et al.
Published: (2025)
BioRAGent: A Retrieval-Augmented Generation System for Showcasing Generative Query Expansion and Domain-Specific Search for Scientific Q&A
by: Ateia, Samy, et al.
Published: (2024)
by: Ateia, Samy, et al.
Published: (2024)
Challenges in Pre-Training Graph Neural Networks for Context-Based Fake News Detection: An Evaluation of Current Strategies and Resource Limitations
by: Donabauer, Gregor, et al.
Published: (2024)
by: Donabauer, Gregor, et al.
Published: (2024)
nchellwig at SemEval-2026 Task 3: Self-Consistent Structured Generation (SCSG) for Dimensional Aspect-Based Sentiment Analysis using Large Language Models
by: Hellwig, Nils Constantin, et al.
Published: (2026)
by: Hellwig, Nils Constantin, et al.
Published: (2026)
Investigating Neural Machine Translation for Low-Resource Languages: Using Bavarian as a Case Study
by: Her, Wan-Hua, et al.
Published: (2024)
by: Her, Wan-Hua, et al.
Published: (2024)
Zero-Shot to Full-Resource: Cross-lingual Transfer Strategies for Aspect-Based Sentiment Analysis
by: Fehle, Jakob, et al.
Published: (2026)
by: Fehle, Jakob, et al.
Published: (2026)
LLM-Based Information Extraction to Support Scientific Literature Research and Publication Workflows
by: Ateia, Samy, et al.
Published: (2025)
by: Ateia, Samy, et al.
Published: (2025)
LLM-as-an-Annotator: Training Lightweight Models with LLM-Annotated Examples for Aspect Sentiment Tuple Prediction
by: Hellwig, Nils Constantin, et al.
Published: (2026)
by: Hellwig, Nils Constantin, et al.
Published: (2026)
Do we still need Human Annotators? Prompting Large Language Models for Aspect Sentiment Quad Prediction
by: Hellwig, Nils Constantin, et al.
Published: (2025)
by: Hellwig, Nils Constantin, et al.
Published: (2025)
Benchmarking Open-Source Large Language Models for Persian in Zero-Shot and Few-Shot Learning
by: Cherakhloo, Mahdi, et al.
Published: (2025)
by: Cherakhloo, Mahdi, et al.
Published: (2025)
Prompting Is All You Need: Multi-view Prompting Large Language Models for Aspect-Based Sentiment Analysis
by: Hellwig, Nils Constantin, et al.
Published: (2026)
by: Hellwig, Nils Constantin, et al.
Published: (2026)
LLMs on Drugs: Language Models Are Few-Shot Consumers
by: Doudkin, Alexander
Published: (2025)
by: Doudkin, Alexander
Published: (2025)
A Benchmark for End-to-End Zero-Shot Biomedical Relation Extraction with LLMs: Experiments with OpenAI Models
by: Brokman, Aviv, et al.
Published: (2025)
by: Brokman, Aviv, et al.
Published: (2025)
AnnoABSA: A Web-Based Annotation Tool for Aspect-Based Sentiment Analysis with Retrieval-Augmented Suggestions
by: Hellwig, Nils Constantin, et al.
Published: (2026)
by: Hellwig, Nils Constantin, et al.
Published: (2026)
Open Source Language Models Can Provide Feedback: Evaluating LLMs' Ability to Help Students Using GPT-4-As-A-Judge
by: Koutcheme, Charles, et al.
Published: (2024)
by: Koutcheme, Charles, et al.
Published: (2024)
Annotation Quality in Aspect-Based Sentiment Analysis: A Case Study Comparing Experts, Students, Crowdworkers, and Large Language Model
by: Donhauser, Niklas, et al.
Published: (2026)
by: Donhauser, Niklas, et al.
Published: (2026)
MolecularGPT: Open Large Language Model (LLM) for Few-Shot Molecular Property Prediction
by: Liu, Yuyan, et al.
Published: (2024)
by: Liu, Yuyan, et al.
Published: (2024)
Tabular LLMs for Interpretable Few-Shot Alzheimer's Disease Prediction with Multimodal Biomedical Data
by: Kearney, Sophie, et al.
Published: (2026)
by: Kearney, Sophie, et al.
Published: (2026)
The Impact of Example Selection in Few-Shot Prompting on Automated Essay Scoring Using GPT Models
by: Yoshida, Lui
Published: (2024)
by: Yoshida, Lui
Published: (2024)
Assessing the Performance of Chinese Open Source Large Language Models in Information Extraction Tasks
by: Cai, Yida, et al.
Published: (2024)
by: Cai, Yida, et al.
Published: (2024)
Extreme Speech Classification in the Era of LLMs: Exploring Open-Source and Proprietary Models
by: Mahajan, Sarthak, et al.
Published: (2025)
by: Mahajan, Sarthak, et al.
Published: (2025)
Comparing GPT-4 and Open-Source Language Models in Misinformation Mitigation
by: Vergho, Tyler, et al.
Published: (2024)
by: Vergho, Tyler, et al.
Published: (2024)
BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks
by: Zhang, Kai, et al.
Published: (2023)
by: Zhang, Kai, et al.
Published: (2023)
Task Contamination: Language Models May Not Be Few-Shot Anymore
by: Li, Changmao, et al.
Published: (2023)
by: Li, Changmao, et al.
Published: (2023)
Is Open-Source There Yet? A Comparative Study on Commercial and Open-Source LLMs in Their Ability to Label Chest X-Ray Reports
by: Dorfner, Felix J., et al.
Published: (2024)
by: Dorfner, Felix J., et al.
Published: (2024)
Can Frontier LLMs Replace Annotators in Biomedical Text Mining? Analyzing Challenges and Exploring Solutions
by: Zhao, Yichong, et al.
Published: (2025)
by: Zhao, Yichong, et al.
Published: (2025)
Cost-Performance Optimization for Processing Low-Resource Language Tasks Using Commercial LLMs
by: Nag, Arijit, et al.
Published: (2024)
by: Nag, Arijit, et al.
Published: (2024)
Cost-Efficient Subjective Task Annotation and Modeling through Few-Shot Annotator Adaptation
by: Golazizian, Preni, et al.
Published: (2024)
by: Golazizian, Preni, et al.
Published: (2024)
Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
by: Bi, Ziqian, et al.
Published: (2025)
by: Bi, Ziqian, et al.
Published: (2025)
OpenThaiGPT 1.5: A Thai-Centric Open Source Large Language Model
by: Yuenyong, Sumeth, et al.
Published: (2024)
by: Yuenyong, Sumeth, et al.
Published: (2024)
A Mixture-of-Experts Approach to Few-Shot Task Transfer in Open-Ended Text Worlds
by: Cui, Christopher Z., et al.
Published: (2024)
by: Cui, Christopher Z., et al.
Published: (2024)
ReProCon: Scalable and Resource-Efficient Few-Shot Biomedical Named Entity Recognition
by: Yoo, Jeongkyun, et al.
Published: (2025)
by: Yoo, Jeongkyun, et al.
Published: (2025)
Open-Source Image Editing Models Are Zero-Shot Vision Learners
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Evaluation of Few-Shot Learning for Classification Tasks in the Polish Language
by: Hadeliya, Tsimur, et al.
Published: (2024)
by: Hadeliya, Tsimur, et al.
Published: (2024)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
by: Vatsal, Shubham, et al.
Published: (2024)
by: Vatsal, Shubham, et al.
Published: (2024)
Language Models are Few-Shot Graders
by: Zhao, Chenyan, et al.
Published: (2025)
by: Zhao, Chenyan, et al.
Published: (2025)
ChatGPT's One-year Anniversary: Are Open-Source Large Language Models Catching up?
by: Chen, Hailin, et al.
Published: (2023)
by: Chen, Hailin, et al.
Published: (2023)
Can OpenSource beat ChatGPT? -- A Comparative Study of Large Language Models for Text-to-Code Generation
by: Mayer, Luis, et al.
Published: (2024)
by: Mayer, Luis, et al.
Published: (2024)
OpenThaiGPT 1.6 and R1: Thai-Centric Open Source and Reasoning Large Language Models
by: Yuenyong, Sumeth, et al.
Published: (2025)
by: Yuenyong, Sumeth, et al.
Published: (2025)
Evaluation of ChatGPT Family of Models for Biomedical Reasoning and Classification
by: Chen, Shan, et al.
Published: (2023)
by: Chen, Shan, et al.
Published: (2023)
Similar Items
-
Can Language Models Critique Themselves? Investigating Self-Feedback for Retrieval Augmented Generation at BioASQ 2025
by: Ateia, Samy, et al.
Published: (2025) -
BioRAGent: A Retrieval-Augmented Generation System for Showcasing Generative Query Expansion and Domain-Specific Search for Scientific Q&A
by: Ateia, Samy, et al.
Published: (2024) -
Challenges in Pre-Training Graph Neural Networks for Context-Based Fake News Detection: An Evaluation of Current Strategies and Resource Limitations
by: Donabauer, Gregor, et al.
Published: (2024) -
nchellwig at SemEval-2026 Task 3: Self-Consistent Structured Generation (SCSG) for Dimensional Aspect-Based Sentiment Analysis using Large Language Models
by: Hellwig, Nils Constantin, et al.
Published: (2026) -
Investigating Neural Machine Translation for Low-Resource Languages: Using Bavarian as a Case Study
by: Her, Wan-Hua, et al.
Published: (2024)