FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Eric, Wu, Kevin, Zou, James |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tuning LLMs by RAG Principles: Towards LLM-native Memory
von: Wei, Jiale, et al.
Veröffentlicht: (2025)
von: Wei, Jiale, et al.
Veröffentlicht: (2025)
Leveraging the Power of LLMs: A Fine-Tuning Approach for High-Quality Aspect-Based Summarization
von: Mullick, Ankan, et al.
Veröffentlicht: (2024)
von: Mullick, Ankan, et al.
Veröffentlicht: (2024)
UNH at CheckThat! 2025: Fine-tuning Vs Prompting in Claim Extraction
von: Wilder, Joe, et al.
Veröffentlicht: (2025)
von: Wilder, Joe, et al.
Veröffentlicht: (2025)
Quantifying and Mitigating Selection Bias in LLMs: A Transferable LoRA Fine-Tuning and Efficient Majority Voting Approach
von: Guda, Blessed, et al.
Veröffentlicht: (2025)
von: Guda, Blessed, et al.
Veröffentlicht: (2025)
Fine-tuning Small Language Models as Efficient Enterprise Search Relevance Labelers
von: Kang, Yue, et al.
Veröffentlicht: (2026)
von: Kang, Yue, et al.
Veröffentlicht: (2026)
Fine-tune the Entire RAG Architecture (including DPR retriever) for Question-Answering
von: Siriwardhana, Shamane, et al.
Veröffentlicht: (2021)
von: Siriwardhana, Shamane, et al.
Veröffentlicht: (2021)
Knowledge Compression via Question Generation: Enhancing Multihop Document Retrieval without Fine-tuning
von: Eponon, Anvi Alex, et al.
Veröffentlicht: (2025)
von: Eponon, Anvi Alex, et al.
Veröffentlicht: (2025)
OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models
von: Maharjan, Jenish, et al.
Veröffentlicht: (2024)
von: Maharjan, Jenish, et al.
Veröffentlicht: (2024)
Q-PEFT: Query-dependent Parameter Efficient Fine-tuning for Text Reranking with Large Language Models
von: Peng, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Peng, Zhiyuan, et al.
Veröffentlicht: (2024)
ATOM: AdapTive and OptiMized dynamic temporal knowledge graph construction using LLMs
von: Lairgi, Yassir, et al.
Veröffentlicht: (2025)
von: Lairgi, Yassir, et al.
Veröffentlicht: (2025)
TQA-Bench: Evaluating LLMs for Multi-Table Question Answering with Scalable Context and Symbolic Extension
von: Qiu, Zipeng, et al.
Veröffentlicht: (2024)
von: Qiu, Zipeng, et al.
Veröffentlicht: (2024)
LEMUR: A Corpus for Robust Fine-Tuning of Multilingual Law Embedding Models for Retrieval
von: Ahmadi, Narges Baba, et al.
Veröffentlicht: (2026)
von: Ahmadi, Narges Baba, et al.
Veröffentlicht: (2026)
Enhancing Temporal Sensitivity of Large Language Model for Recommendation with Counterfactual Tuning
von: Liu, Yutian, et al.
Veröffentlicht: (2025)
von: Liu, Yutian, et al.
Veröffentlicht: (2025)
SPARQL Generation: an analysis on fine-tuning OpenLLaMA for Question Answering over a Life Science Knowledge Graph
von: Rangel, Julio C., et al.
Veröffentlicht: (2024)
von: Rangel, Julio C., et al.
Veröffentlicht: (2024)
Leveraging Fine-Tuned Large Language Models for Interpretable Pancreatic Cystic Lesion Feature Extraction and Risk Categorization
von: Rasromani, Ebrahim, et al.
Veröffentlicht: (2025)
von: Rasromani, Ebrahim, et al.
Veröffentlicht: (2025)
Evaluating LLM-based Approaches to Legal Citation Prediction: Domain-specific Pre-training, Fine-tuning, or RAG? A Benchmark and an Australian Law Case Study
von: Han, Jiuzhou, et al.
Veröffentlicht: (2024)
von: Han, Jiuzhou, et al.
Veröffentlicht: (2024)
Retrieval-Augmented LLMs for Evidence Localization in Clinical Trial Recruitment from Longitudinal EHR Narratives
von: Chen, Ziyi, et al.
Veröffentlicht: (2026)
von: Chen, Ziyi, et al.
Veröffentlicht: (2026)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
Structured Attention Matters to Multimodal LLMs in Document Understanding
von: Liu, Chang, et al.
Veröffentlicht: (2025)
von: Liu, Chang, et al.
Veröffentlicht: (2025)
RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation
von: Li, Xiaoxi, et al.
Veröffentlicht: (2024)
von: Li, Xiaoxi, et al.
Veröffentlicht: (2024)
Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs
von: Wang, Ziliang, et al.
Veröffentlicht: (2025)
von: Wang, Ziliang, et al.
Veröffentlicht: (2025)
RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
von: Tan, Zhiwen, et al.
Veröffentlicht: (2025)
von: Tan, Zhiwen, et al.
Veröffentlicht: (2025)
StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
von: Wang, Ziliang, et al.
Veröffentlicht: (2025)
von: Wang, Ziliang, et al.
Veröffentlicht: (2025)
C-SEO Bench: Does Conversational SEO Work?
von: Puerto, Haritz, et al.
Veröffentlicht: (2025)
von: Puerto, Haritz, et al.
Veröffentlicht: (2025)
EnterpriseEM: Fine-tuned Embeddings for Enterprise Semantic Search
von: Rathinasamy, Kamalkumar, et al.
Veröffentlicht: (2024)
von: Rathinasamy, Kamalkumar, et al.
Veröffentlicht: (2024)
NUDGE: Lightweight Non-Parametric Fine-Tuning of Embeddings for Retrieval
von: Zeighami, Sepanta, et al.
Veröffentlicht: (2024)
von: Zeighami, Sepanta, et al.
Veröffentlicht: (2024)
Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval
von: Marinas, Inés Altemir, et al.
Veröffentlicht: (2025)
von: Marinas, Inés Altemir, et al.
Veröffentlicht: (2025)
R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning
von: Song, Huatong, et al.
Veröffentlicht: (2025)
von: Song, Huatong, et al.
Veröffentlicht: (2025)
HawkBench: Investigating Resilience of RAG Methods on Stratified Information-Seeking Tasks
von: Qian, Hongjin, et al.
Veröffentlicht: (2025)
von: Qian, Hongjin, et al.
Veröffentlicht: (2025)
LR-SQL: A Supervised Fine-Tuning Method for Text2SQL Tasks under Low-Resource Scenarios
von: Wuzhenghong, Wen, et al.
Veröffentlicht: (2024)
von: Wuzhenghong, Wen, et al.
Veröffentlicht: (2024)
NanoNER: Named Entity Recognition for nanobiology using experts' knowledge and distant supervision
von: Lentschat, Martin, et al.
Veröffentlicht: (2024)
von: Lentschat, Martin, et al.
Veröffentlicht: (2024)
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
von: Jin, Zhuoran, et al.
Veröffentlicht: (2024)
von: Jin, Zhuoran, et al.
Veröffentlicht: (2024)
FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering
von: Choi, Chanyeol, et al.
Veröffentlicht: (2025)
von: Choi, Chanyeol, et al.
Veröffentlicht: (2025)
GPT vs RETRO: Exploring the Intersection of Retrieval and Parameter-Efficient Fine-Tuning
von: Ficek, Aleksander, et al.
Veröffentlicht: (2024)
von: Ficek, Aleksander, et al.
Veröffentlicht: (2024)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
von: Markowitz, Elan, et al.
Veröffentlicht: (2025)
von: Markowitz, Elan, et al.
Veröffentlicht: (2025)
ClinicalBench: Stress-Testing Assertion-Aware Retrieval for Cross-Admission Clinical QA on MIMIC-IV
von: Stinard, Alex
Veröffentlicht: (2026)
von: Stinard, Alex
Veröffentlicht: (2026)
MedSlice: Fine-Tuned Large Language Models for Secure Clinical Note Sectioning
von: Davis, Joshua, et al.
Veröffentlicht: (2025)
von: Davis, Joshua, et al.
Veröffentlicht: (2025)
A Systematic Evaluation of LLM Strategies for Mental Health Text Analysis: Fine-tuning vs. Prompt Engineering vs. RAG
von: Kermani, Arshia, et al.
Veröffentlicht: (2025)
von: Kermani, Arshia, et al.
Veröffentlicht: (2025)
FinSphere, a Real-Time Stock Analysis Agent Powered by Instruction-Tuned LLMs and Domain Tools
von: Han, Shijie, et al.
Veröffentlicht: (2025)
von: Han, Shijie, et al.
Veröffentlicht: (2025)
NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment
von: Wu, Wenqing, et al.
Veröffentlicht: (2026)
von: Wu, Wenqing, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Tuning LLMs by RAG Principles: Towards LLM-native Memory
von: Wei, Jiale, et al.
Veröffentlicht: (2025) -
Leveraging the Power of LLMs: A Fine-Tuning Approach for High-Quality Aspect-Based Summarization
von: Mullick, Ankan, et al.
Veröffentlicht: (2024) -
UNH at CheckThat! 2025: Fine-tuning Vs Prompting in Claim Extraction
von: Wilder, Joe, et al.
Veröffentlicht: (2025) -
Quantifying and Mitigating Selection Bias in LLMs: A Transferable LoRA Fine-Tuning and Efficient Majority Voting Approach
von: Guda, Blessed, et al.
Veröffentlicht: (2025) -
Fine-tuning Small Language Models as Efficient Enterprise Search Relevance Labelers
von: Kang, Yue, et al.
Veröffentlicht: (2026)