When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models
Fuente:
arXiv
Saved in:
| Main Authors: | Amouyal, Samuel Joseph, Meltzer-Asscher, Aya, Berant, Jonathan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Comparing Human and Language Models Sentence Processing Difficulties on Complex Structures
by: Amouyal, Samuel Joseph, et al.
Published: (2025)
by: Amouyal, Samuel Joseph, et al.
Published: (2025)
Large Language Models for Psycholinguistic Plausibility Pretesting
by: Amouyal, Samuel Joseph, et al.
Published: (2024)
by: Amouyal, Samuel Joseph, et al.
Published: (2024)
Are they human? Detecting large language models by probing human memory constraints
by: Schug, Simon, et al.
Published: (2026)
by: Schug, Simon, et al.
Published: (2026)
Strong and weak alignment of large language models with human values
by: Khamassi, Mehdi, et al.
Published: (2024)
by: Khamassi, Mehdi, et al.
Published: (2024)
Weakly Supervised Text-to-SQL Parsing through Question Decomposition
by: Wolfson, Tomer, et al.
Published: (2021)
by: Wolfson, Tomer, et al.
Published: (2021)
Large language models show fragile cognitive reasoning about human emotions
by: Bhattacharyya, Sree, et al.
Published: (2025)
by: Bhattacharyya, Sree, et al.
Published: (2025)
Making Retrieval-Augmented Language Models Robust to Irrelevant Context
by: Yoran, Ori, et al.
Published: (2023)
by: Yoran, Ori, et al.
Published: (2023)
From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty
by: Ivgi, Maor, et al.
Published: (2024)
by: Ivgi, Maor, et al.
Published: (2024)
Comparing large language models and human programmers for generating programming code
by: Hou, Wenpin, et al.
Published: (2024)
by: Hou, Wenpin, et al.
Published: (2024)
TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain
by: Barboule, Camille, et al.
Published: (2024)
by: Barboule, Camille, et al.
Published: (2024)
Post-training makes large language models less human-like
by: Binz, Marcel, et al.
Published: (2026)
by: Binz, Marcel, et al.
Published: (2026)
A closer look at how large language models trust humans: patterns and biases
by: Lerman, Valeria, et al.
Published: (2025)
by: Lerman, Valeria, et al.
Published: (2025)
Child vs. machine language learning: Can the logical structure of human language unleash LLMs?
by: Sauerland, Uli, et al.
Published: (2025)
by: Sauerland, Uli, et al.
Published: (2025)
Large language models replicate and predict human cooperation across experiments in game theory
by: Palatsi, Andrea Cera, et al.
Published: (2025)
by: Palatsi, Andrea Cera, et al.
Published: (2025)
WizardLM: Empowering large pre-trained language models to follow complex instructions
by: Xu, Can, et al.
Published: (2023)
by: Xu, Can, et al.
Published: (2023)
Revealing emergent human-like conceptual representations from language prediction
by: Xu, Ningyu, et al.
Published: (2025)
by: Xu, Ningyu, et al.
Published: (2025)
HumT DumT: Measuring and controlling human-like language in LLMs
by: Cheng, Myra, et al.
Published: (2025)
by: Cheng, Myra, et al.
Published: (2025)
TouchAI: Exploring human-AI perceptual alignment in touch through language model representations
by: Zhong, Shu, et al.
Published: (2024)
by: Zhong, Shu, et al.
Published: (2024)
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
by: Yoran, Ori, et al.
Published: (2024)
by: Yoran, Ori, et al.
Published: (2024)
When Large Language Models contradict humans? Large Language Models' Sycophantic Behaviour
by: Ranaldi, Leonardo, et al.
Published: (2023)
by: Ranaldi, Leonardo, et al.
Published: (2023)
Can LLMs interpret figurative language as humans do?: surface-level vs representational similarity
by: Bollepally, Samhita, et al.
Published: (2026)
by: Bollepally, Samhita, et al.
Published: (2026)
Automated stereotactic radiosurgery planning using a human-in-the-loop reasoning large language model agent
by: Nusrat, Humza, et al.
Published: (2025)
by: Nusrat, Humza, et al.
Published: (2025)
Answering Questions by Meta-Reasoning over Multiple Chains of Thought
by: Yoran, Ori, et al.
Published: (2023)
by: Yoran, Ori, et al.
Published: (2023)
Language models show human-like content effects on reasoning tasks
by: Dasgupta, Ishita, et al.
Published: (2022)
by: Dasgupta, Ishita, et al.
Published: (2022)
ChildEval: When large language models meet children's personalities
by: Luo, Yanyan, et al.
Published: (2026)
by: Luo, Yanyan, et al.
Published: (2026)
Language models align with human judgments on key grammatical constructions
by: Hu, Jennifer, et al.
Published: (2024)
by: Hu, Jennifer, et al.
Published: (2024)
Increasing faithfulness in human-human dialog summarization with Spoken Language Understanding tasks
by: Akani, Eunice, et al.
Published: (2024)
by: Akani, Eunice, et al.
Published: (2024)
How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?
by: Sil, Pritam, et al.
Published: (2026)
by: Sil, Pritam, et al.
Published: (2026)
Language modelling techniques for analysing the impact of human genetic variation
by: Hegde, Megha, et al.
Published: (2025)
by: Hegde, Megha, et al.
Published: (2025)
Do self-supervised speech and language models extract similar representations as human brain?
by: Chen, Peili, et al.
Published: (2023)
by: Chen, Peili, et al.
Published: (2023)
Preference learning in shades of gray: Interpretable and bias-aware reward modeling for human preferences
by: Oprea, Simona-Vasilica, et al.
Published: (2026)
by: Oprea, Simona-Vasilica, et al.
Published: (2026)
ALTA: Compiler-Based Analysis of Transformers
by: Shaw, Peter, et al.
Published: (2024)
by: Shaw, Peter, et al.
Published: (2024)
Can large language models replace humans in the systematic review process? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages
by: Khraisha, Qusai, et al.
Published: (2023)
by: Khraisha, Qusai, et al.
Published: (2023)
AI Act and Large Language Models (LLMs): When critical issues and privacy impact require human and ethical oversight
by: Fabiano, Nicola
Published: (2024)
by: Fabiano, Nicola
Published: (2024)
Emergent effects of scaling on the functional hierarchies within large language models
by: Bogdan, Paul C.
Published: (2025)
by: Bogdan, Paul C.
Published: (2025)
LLM should think and action as a human
by: Leung, Haun, et al.
Published: (2025)
by: Leung, Haun, et al.
Published: (2025)
Dissociating language and thought in large language models
by: Mahowald, Kyle, et al.
Published: (2023)
by: Mahowald, Kyle, et al.
Published: (2023)
GLEE: A Unified Framework and Benchmark for Language-based Economic Environments
by: Shapira, Eilam, et al.
Published: (2024)
by: Shapira, Eilam, et al.
Published: (2024)
Xmodel-LM Technical Report
by: Wang, Yichuan, et al.
Published: (2024)
by: Wang, Yichuan, et al.
Published: (2024)
Language models are better than humans at next-token prediction
by: Shlegeris, Buck, et al.
Published: (2022)
by: Shlegeris, Buck, et al.
Published: (2022)
Similar Items
-
Comparing Human and Language Models Sentence Processing Difficulties on Complex Structures
by: Amouyal, Samuel Joseph, et al.
Published: (2025) -
Large Language Models for Psycholinguistic Plausibility Pretesting
by: Amouyal, Samuel Joseph, et al.
Published: (2024) -
Are they human? Detecting large language models by probing human memory constraints
by: Schug, Simon, et al.
Published: (2026) -
Strong and weak alignment of large language models with human values
by: Khamassi, Mehdi, et al.
Published: (2024) -
Weakly Supervised Text-to-SQL Parsing through Question Decomposition
by: Wolfson, Tomer, et al.
Published: (2021)