Auxiliary task demands mask the capabilities of smaller language models
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Jennifer, Frank, Michael C. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Just-in-time and distributed task representations in language models
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
Re-evaluating Theory of Mind evaluation in large language models
by: Hu, Jennifer, et al.
Published: (2025)
by: Hu, Jennifer, et al.
Published: (2025)
Code-enabled language models can outperform reasoning models on diverse tasks
by: Zhang, Cedegao E., et al.
Published: (2025)
by: Zhang, Cedegao E., et al.
Published: (2025)
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
by: Ilia, Evgenia, et al.
Published: (2024)
by: Ilia, Evgenia, et al.
Published: (2024)
Superhuman performance of a large language model on the reasoning tasks of a physician
by: Brodeur, Peter G., et al.
Published: (2024)
by: Brodeur, Peter G., et al.
Published: (2024)
Evidence of interrelated cognitive-like capabilities in large language models: Indications of artificial general intelligence or achievement?
by: Ilić, David, et al.
Published: (2023)
by: Ilić, David, et al.
Published: (2023)
ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model
by: Lan, Wuyang, et al.
Published: (2025)
by: Lan, Wuyang, et al.
Published: (2025)
Evidence from counterfactual tasks supports emergent analogical reasoning in large language models
by: Webb, Taylor, et al.
Published: (2024)
by: Webb, Taylor, et al.
Published: (2024)
Cognitive models can reveal interpretable value trade-offs in language models
by: Murthy, Sonia K., et al.
Published: (2025)
by: Murthy, Sonia K., et al.
Published: (2025)
TechGPT-2.0: A large language model project to solve the task of knowledge graph construction
by: Wang, Jiaqi, et al.
Published: (2024)
by: Wang, Jiaqi, et al.
Published: (2024)
Testing AI on language comprehension tasks reveals insensitivity to underlying meaning
by: Dentella, Vittoria, et al.
Published: (2023)
by: Dentella, Vittoria, et al.
Published: (2023)
Impact of enriched meaning representations for language generation in dialogue tasks: A comprehensive exploration of the relevance of tasks, corpora and metrics
by: Vázquez, Alain, et al.
Published: (2026)
by: Vázquez, Alain, et al.
Published: (2026)
Fluent dreaming for language models
by: Thompson, T. Ben, et al.
Published: (2024)
by: Thompson, T. Ben, et al.
Published: (2024)
Fine-tuning large language models for domain adaptation: Exploration of training strategies, scaling, model merging and synergistic capabilities
by: Lu, Wei, et al.
Published: (2024)
by: Lu, Wei, et al.
Published: (2024)
Aviary: training language agents on challenging scientific tasks
by: Narayanan, Siddharth, et al.
Published: (2024)
by: Narayanan, Siddharth, et al.
Published: (2024)
Leveraging small language models for Text2SPARQL tasks to improve the resilience of AI assistance
by: Brei, Felix, et al.
Published: (2024)
by: Brei, Felix, et al.
Published: (2024)
Dissociating language and thought in large language models
by: Mahowald, Kyle, et al.
Published: (2023)
by: Mahowald, Kyle, et al.
Published: (2023)
Integration of cognitive tasks into artificial general intelligence test for large models
by: Qu, Youzhi, et al.
Published: (2024)
by: Qu, Youzhi, et al.
Published: (2024)
Emergent effects of scaling on the functional hierarchies within large language models
by: Bogdan, Paul C.
Published: (2025)
by: Bogdan, Paul C.
Published: (2025)
Algorithmic progress in language models
by: Ho, Anson, et al.
Published: (2024)
by: Ho, Anson, et al.
Published: (2024)
LLMs as On-demand Customizable Service
by: Sarkar, Souvika, et al.
Published: (2024)
by: Sarkar, Souvika, et al.
Published: (2024)
ChildEval: When large language models meet children's personalities
by: Luo, Yanyan, et al.
Published: (2026)
by: Luo, Yanyan, et al.
Published: (2026)
Does quantization affect models' performance on long-context tasks?
by: Mekala, Anmol, et al.
Published: (2025)
by: Mekala, Anmol, et al.
Published: (2025)
On the attribution of confidence to large language models
by: Keeling, Geoff, et al.
Published: (2024)
by: Keeling, Geoff, et al.
Published: (2024)
Large language models and linguistic intentionality
by: Grindrod, Jumbly
Published: (2024)
by: Grindrod, Jumbly
Published: (2024)
A survey of textual cyber abuse detection using cutting-edge language models and large language models
by: Diaz-Garcia, Jose A., et al.
Published: (2025)
by: Diaz-Garcia, Jose A., et al.
Published: (2025)
Do language models practice what they preach? Examining language ideologies about gendered language reform encoded in LLMs
by: Watson, Julia, et al.
Published: (2024)
by: Watson, Julia, et al.
Published: (2024)
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?
by: Lei, Zhikai, et al.
Published: (2024)
by: Lei, Zhikai, et al.
Published: (2024)
Why mask diffusion does not work
by: Sun, Haocheng, et al.
Published: (2025)
by: Sun, Haocheng, et al.
Published: (2025)
Signatures of human-like processing in Transformer forward passes
by: Hu, Jennifer, et al.
Published: (2025)
by: Hu, Jennifer, et al.
Published: (2025)
Linguistic traces of stochastic empathy in language models
by: Kleinberg, Bennett, et al.
Published: (2024)
by: Kleinberg, Bennett, et al.
Published: (2024)
Infusing clinical knowledge into tokenisers for language models
by: Hasan, Abul, et al.
Published: (2024)
by: Hasan, Abul, et al.
Published: (2024)
CONFLARE: CONFormal LArge language model REtrieval
by: Rouzrokh, Pouria, et al.
Published: (2024)
by: Rouzrokh, Pouria, et al.
Published: (2024)
Language translation, and change of accent for speech-to-speech task using diffusion model
by: Mishra, Abhishek, et al.
Published: (2025)
by: Mishra, Abhishek, et al.
Published: (2025)
Multi-Target Cross-Lingual Summarization: a novel task and a language-neutral approach
by: Pernes, Diogo, et al.
Published: (2024)
by: Pernes, Diogo, et al.
Published: (2024)
What do language models model? Transformers, automata, and the format of thought
by: Klein, Colin
Published: (2025)
by: Klein, Colin
Published: (2025)
Adaptively profiling models with task elicitation
by: Brown, Davis, et al.
Published: (2025)
by: Brown, Davis, et al.
Published: (2025)
A novel language model for predicting serious adverse event results in clinical trials from their prospective registrations
by: Hu, Qixuan, et al.
Published: (2025)
by: Hu, Qixuan, et al.
Published: (2025)
Multi-round jailbreak attack on large language models
by: Zhou, Yihua, et al.
Published: (2024)
by: Zhou, Yihua, et al.
Published: (2024)
Similar Items
-
Just-in-time and distributed task representations in language models
by: Li, Yuxuan, et al.
Published: (2025) -
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian
by: Syromiatnikov, Mykyta, et al.
Published: (2025) -
Re-evaluating Theory of Mind evaluation in large language models
by: Hu, Jennifer, et al.
Published: (2025) -
Code-enabled language models can outperform reasoning models on diverse tasks
by: Zhang, Cedegao E., et al.
Published: (2025) -
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
by: Ilia, Evgenia, et al.
Published: (2024)