Evaluating Adjective-Noun Compositionality in LLMs: Functional vs Representational Perspectives
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917344244137984 |
|---|---|
| author | Dhar, Ruchira Peng, Qiwei Søgaard, Anders |
| author_facet | Dhar, Ruchira Peng, Qiwei Søgaard, Anders |
| contents | Compositionality is considered central to language abilities. As performant language systems, how do large language models (LLMs) do on compositional tasks? We evaluate adjective-noun compositionality in LLMs using two complementary setups: prompt-based functional assessment and a representational analysis of internal model states. Our results reveal a striking divergence between task performance and internal states. While LLMs reliably develop compositional representations, they fail to translate consistently into functional task success across model variants. Consequently, we highlight the importance of contrastive evaluation for obtaining a more complete understanding of model capabilities. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_09994 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Evaluating Adjective-Noun Compositionality in LLMs: Functional vs Representational Perspectives Dhar, Ruchira Peng, Qiwei Søgaard, Anders Computation and Language Artificial Intelligence Compositionality is considered central to language abilities. As performant language systems, how do large language models (LLMs) do on compositional tasks? We evaluate adjective-noun compositionality in LLMs using two complementary setups: prompt-based functional assessment and a representational analysis of internal model states. Our results reveal a striking divergence between task performance and internal states. While LLMs reliably develop compositional representations, they fail to translate consistently into functional task success across model variants. Consequently, we highlight the importance of contrastive evaluation for obtaining a more complete understanding of model capabilities. |
| title | Evaluating Adjective-Noun Compositionality in LLMs: Functional vs Representational Perspectives |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2603.09994 |