Fine-Tuning Language Models to Know What They Know
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918519724048384 |
|---|---|
| author | Park, Sangjun Meyerson, Elliot Qiu, Xin Miikkulainen, Risto |
| author_facet | Park, Sangjun Meyerson, Elliot Qiu, Xin Miikkulainen, Risto |
| contents | Evaluating true metacognition in Large Language Models (LLMs) is difficult due to biases and heuristics. This paper presents a framework to measure and enhance LLM metacognition while controlling for these biases. A measurement method using the $d'_{\rm type2}$ metric is established to isolate metacognitive ability. The Evolution Strategy for Metacognitive Alignment (ESMA) is proposed, demonstrating robust generalization across unseen datasets, languages, and newly acquired knowledge. Finally, parameter analysis reveals that these improvements are driven by a sparse set of parameters, offering new pathways for targeted metacognitive optimization. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_02605 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Fine-Tuning Language Models to Know What They Know Park, Sangjun Meyerson, Elliot Qiu, Xin Miikkulainen, Risto Neural and Evolutionary Computing Artificial Intelligence Computation and Language Neurons and Cognition Evaluating true metacognition in Large Language Models (LLMs) is difficult due to biases and heuristics. This paper presents a framework to measure and enhance LLM metacognition while controlling for these biases. A measurement method using the $d'_{\rm type2}$ metric is established to isolate metacognitive ability. The Evolution Strategy for Metacognitive Alignment (ESMA) is proposed, demonstrating robust generalization across unseen datasets, languages, and newly acquired knowledge. Finally, parameter analysis reveals that these improvements are driven by a sparse set of parameters, offering new pathways for targeted metacognitive optimization. |
| title | Fine-Tuning Language Models to Know What They Know |
| topic | Neural and Evolutionary Computing Artificial Intelligence Computation and Language Neurons and Cognition |
| url | https://arxiv.org/abs/2602.02605 |