Cognitive models can reveal interpretable value trade-offs in language models
Fuente:
arXiv
Saved in:
| Main Authors: | Murthy, Sonia K., Zhao, Rosie, Hu, Jennifer, Kakade, Sham, Wulfmeier, Markus, Qian, Peng, Ullman, Tomer |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One fish, two fish, but not the whole sea: Alignment reduces language models' conceptual diversity
by: Murthy, Sonia K., et al.
Published: (2024)
by: Murthy, Sonia K., et al.
Published: (2024)
Re-evaluating Theory of Mind evaluation in large language models
by: Hu, Jennifer, et al.
Published: (2025)
by: Hu, Jennifer, et al.
Published: (2025)
Shades of Zero: Distinguishing Impossibility from Inconceivability
by: Hu, Jennifer, et al.
Published: (2025)
by: Hu, Jennifer, et al.
Published: (2025)
Interpreting the linear structure of vision-language model embedding spaces
by: Papadimitriou, Isabel, et al.
Published: (2025)
by: Papadimitriou, Isabel, et al.
Published: (2025)
The Illusion-Illusion: Vision Language Models See Illusions Where There are None
by: Ullman, Tomer
Published: (2024)
by: Ullman, Tomer
Published: (2024)
Context informs pragmatic interpretation in vision-language models
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
by: Zhang, Hanlin, et al.
Published: (2026)
by: Zhang, Hanlin, et al.
Published: (2026)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
by: Jin, Jikai, et al.
Published: (2025)
by: Jin, Jikai, et al.
Published: (2025)
State space models can express n-gram languages
by: Nandakumar, Vinoth, et al.
Published: (2023)
by: Nandakumar, Vinoth, et al.
Published: (2023)
Auxiliary task demands mask the capabilities of smaller language models
by: Hu, Jennifer, et al.
Published: (2024)
by: Hu, Jennifer, et al.
Published: (2024)
Human-interpretable clustering of short-text using large language models
by: Miller, Justin K., et al.
Published: (2024)
by: Miller, Justin K., et al.
Published: (2024)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Deconstructing What Makes a Good Optimizer for Language Models
by: Zhao, Rosie, et al.
Published: (2024)
by: Zhao, Rosie, et al.
Published: (2024)
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
by: Panigrahi, Abhishek, et al.
Published: (2025)
by: Panigrahi, Abhishek, et al.
Published: (2025)
Inducing anxiety in large language models can induce bias
by: Coda-Forno, Julian, et al.
Published: (2023)
by: Coda-Forno, Julian, et al.
Published: (2023)
Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models
by: Conwell, Colin, et al.
Published: (2024)
by: Conwell, Colin, et al.
Published: (2024)
Code-enabled language models can outperform reasoning models on diverse tasks
by: Zhang, Cedegao E., et al.
Published: (2025)
by: Zhang, Cedegao E., et al.
Published: (2025)
Perturbed examples reveal invariances shared by language models
by: Rawal, Ruchit, et al.
Published: (2023)
by: Rawal, Ruchit, et al.
Published: (2023)
Forking Paths in Neural Text Generation
by: Bigelow, Eric, et al.
Published: (2024)
by: Bigelow, Eric, et al.
Published: (2024)
Lost without translation -- Can transformer (language models) understand mood states?
by: Shivaprakash, Prakrithi, et al.
Published: (2025)
by: Shivaprakash, Prakrithi, et al.
Published: (2025)
Large language models can disambiguate opioid slang on social media
by: Carpenter, Kristy A., et al.
Published: (2026)
by: Carpenter, Kristy A., et al.
Published: (2026)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
Identifying and interpreting non-aligned human conceptual representations using language modeling
by: Bao, Wanqian, et al.
Published: (2024)
by: Bao, Wanqian, et al.
Published: (2024)
Retrieval-augmented reasoning with lean language models
by: Chan, Ryan Sze-Yin, et al.
Published: (2025)
by: Chan, Ryan Sze-Yin, et al.
Published: (2025)
Random Scaling of Emergent Capabilities
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
Peer-Predictive Self-Training for Language Model Reasoning
by: Feng, Shi, et al.
Published: (2026)
by: Feng, Shi, et al.
Published: (2026)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
by: Qi, Zhenting, et al.
Published: (2024)
by: Qi, Zhenting, et al.
Published: (2024)
What can large language models do for sustainable food?
by: Thomas, Anna T., et al.
Published: (2025)
by: Thomas, Anna T., et al.
Published: (2025)
SemPool: Simple, robust, and interpretable KG pooling for enhancing language models
by: Mavromatis, Costas, et al.
Published: (2024)
by: Mavromatis, Costas, et al.
Published: (2024)
Comparison of different Unique hard attention transformer models by the formal languages they can recognize
by: Ryvkin, Leonid
Published: (2025)
by: Ryvkin, Leonid
Published: (2025)
CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
by: Prabhakar, Akshara, et al.
Published: (2024)
by: Prabhakar, Akshara, et al.
Published: (2024)
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
by: Song, Yuda, et al.
Published: (2024)
by: Song, Yuda, et al.
Published: (2024)
Truth-value judgment in language models: 'truth directions' are context sensitive
by: Schouten, Stefan F., et al.
Published: (2024)
by: Schouten, Stefan F., et al.
Published: (2024)
Large language models can accurately predict searcher preferences
by: Thomas, Paul, et al.
Published: (2023)
by: Thomas, Paul, et al.
Published: (2023)
Linear representations in language models can change dramatically over a conversation
by: Lampinen, Andrew Kyle, et al.
Published: (2026)
by: Lampinen, Andrew Kyle, et al.
Published: (2026)
ChildEval: When large language models meet children's personalities
by: Luo, Yanyan, et al.
Published: (2026)
by: Luo, Yanyan, et al.
Published: (2026)
Strong and weak alignment of large language models with human values
by: Khamassi, Mehdi, et al.
Published: (2024)
by: Khamassi, Mehdi, et al.
Published: (2024)
Ploutos: Towards interpretable stock movement prediction with financial large language model
by: Tong, Hanshuang, et al.
Published: (2024)
by: Tong, Hanshuang, et al.
Published: (2024)
Similar Items
-
One fish, two fish, but not the whole sea: Alignment reduces language models' conceptual diversity
by: Murthy, Sonia K., et al.
Published: (2024) -
Re-evaluating Theory of Mind evaluation in large language models
by: Hu, Jennifer, et al.
Published: (2025) -
Shades of Zero: Distinguishing Impossibility from Inconceivability
by: Hu, Jennifer, et al.
Published: (2025) -
Interpreting the linear structure of vision-language model embedding spaces
by: Papadimitriou, Isabel, et al.
Published: (2025) -
The Illusion-Illusion: Vision Language Models See Illusions Where There are None
by: Ullman, Tomer
Published: (2024)