The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Mo, Liu, Lemao, Wu, Junjie, Chung, Tsz Ting, Zhang, Shunchi, Li, Jiangnan, Yeung, Dit-Yan, Zhou, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
by: Chung, Tsz Ting, et al.
Published: (2025)
by: Chung, Tsz Ting, et al.
Published: (2025)
Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models
by: Wu, Junjie, et al.
Published: (2024)
by: Wu, Junjie, et al.
Published: (2024)
SitEmb-v1.5: Improved Context-Aware Dense Retrieval for Semantic Association and Long Story Comprehension
by: Wu, Junjie, et al.
Published: (2025)
by: Wu, Junjie, et al.
Published: (2025)
Rethinking Targeted Adversarial Attacks For Neural Machine Translation
by: Wu, Junjie, et al.
Published: (2024)
by: Wu, Junjie, et al.
Published: (2024)
Understanding LLMs' Fluid Intelligence Deficiency: An Analysis of the ARC Task
by: Wu, Junjie, et al.
Published: (2025)
by: Wu, Junjie, et al.
Published: (2025)
Selection-p: Self-Supervised Task-Agnostic Prompt Compression for Faithfulness and Transferability
by: Chung, Tsz Ting, et al.
Published: (2024)
by: Chung, Tsz Ting, et al.
Published: (2024)
TransformMix: Learning Transformation and Mixing Strategies from Data
by: Cheung, Tsz-Him, et al.
Published: (2024)
by: Cheung, Tsz-Him, et al.
Published: (2024)
The Essence of Contextual Understanding in Theory of Mind: A Study on Question Answering with Story Characters
by: Zhou, Chulun, et al.
Published: (2025)
by: Zhou, Chulun, et al.
Published: (2025)
Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding
by: Li, Yuqing, et al.
Published: (2025)
by: Li, Yuqing, et al.
Published: (2025)
Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models
by: Wu, Junjie, et al.
Published: (2025)
by: Wu, Junjie, et al.
Published: (2025)
Few-Shot Character Understanding in Movies as an Assessment to Meta-Learning of Theory-of-Mind
by: Yu, Mo, et al.
Published: (2022)
by: Yu, Mo, et al.
Published: (2022)
MiA-Signature: Approximating Global Activation for Long-Context Understanding
by: Li, Yuqing, et al.
Published: (2026)
by: Li, Yuqing, et al.
Published: (2026)
PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts
by: Yu, Mo, et al.
Published: (2025)
by: Yu, Mo, et al.
Published: (2025)
CoherenDream: Boosting Holistic Text Coherence in 3D Generation via Multimodal Large Language Models Feedback
by: Jiang, Chenhan, et al.
Published: (2025)
by: Jiang, Chenhan, et al.
Published: (2025)
Fine-Grained Modeling of Narrative Context: A Coherence Perspective via Retrospective Questions
by: Xu, Liyan, et al.
Published: (2024)
by: Xu, Liyan, et al.
Published: (2024)
G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o
by: Tong, Tony Cheng, et al.
Published: (2024)
by: Tong, Tony Cheng, et al.
Published: (2024)
Implicit Concept Removal of Diffusion Models
by: Liu, Zhili, et al.
Published: (2023)
by: Liu, Zhili, et al.
Published: (2023)
Who Are All The Stochastic Parrots Imitating? They Should Tell Us!
by: Shaier, Sagi, et al.
Published: (2023)
by: Shaier, Sagi, et al.
Published: (2023)
Assessment Twins: A Protocol for AI-Vulnerable Summative Assessment
by: Roe, Jasper, et al.
Published: (2025)
by: Roe, Jasper, et al.
Published: (2025)
The Dark Side of ChatGPT: Legal and Ethical Challenges from Stochastic Parrots and Hallucination
by: Li, Zihao
Published: (2023)
by: Li, Zihao
Published: (2023)
Position: the Stochastic Parrot in the Coal Mine. Model Collapse is a Threat to Low-Resource Communities
by: Jarvis, Devon, et al.
Published: (2026)
by: Jarvis, Devon, et al.
Published: (2026)
The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks
by: Møller, Anders Giovanni, et al.
Published: (2023)
by: Møller, Anders Giovanni, et al.
Published: (2023)
Dense Retrievers Can Fail on Simple Queries: Revealing The Granularity Dilemma of Embeddings
by: Xu, Liyan, et al.
Published: (2025)
by: Xu, Liyan, et al.
Published: (2025)
Parrot: Multilingual Visual Instruction Tuning
by: Sun, Hai-Long, et al.
Published: (2024)
by: Sun, Hai-Long, et al.
Published: (2024)
Out of the Cage: How Stochastic Parrots Win in Cyber Security Environments
by: Rigaki, Maria, et al.
Published: (2023)
by: Rigaki, Maria, et al.
Published: (2023)
Towards Threshold-Free KV Cache Pruning
by: Ni, Xuanfan, et al.
Published: (2025)
by: Ni, Xuanfan, et al.
Published: (2025)
FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation
by: Jiang, Chenhan, et al.
Published: (2026)
by: Jiang, Chenhan, et al.
Published: (2026)
Learning 3D Persistent Embodied World Models
by: Zhou, Siyuan, et al.
Published: (2025)
by: Zhou, Siyuan, et al.
Published: (2025)
Stochastic Parrots or ICU Experts? Large Language Models in Critical Care Medicine: A Scoping Review
by: Shi, Tongyue, et al.
Published: (2024)
by: Shi, Tongyue, et al.
Published: (2024)
Previously on the Stories: Recap Snippet Identification for Story Reading
by: Li, Jiangnan, et al.
Published: (2024)
by: Li, Jiangnan, et al.
Published: (2024)
Query-focused and Memory-aware Reranker for Long Context Processing
by: Li, Yuqing, et al.
Published: (2026)
by: Li, Yuqing, et al.
Published: (2026)
The Existential Theory of the Reals with Summation Operators
by: Bläser, Markus, et al.
Published: (2024)
by: Bläser, Markus, et al.
Published: (2024)
Parrot: Pareto-optimal Multi-Reward Reinforcement Learning Framework for Text-to-Image Generation
by: Lee, Seung Hyun, et al.
Published: (2024)
by: Lee, Seung Hyun, et al.
Published: (2024)
Mixed Autoencoder for Self-supervised Visual Representation Learning
by: Chen, Kai, et al.
Published: (2023)
by: Chen, Kai, et al.
Published: (2023)
Judge Like Human Examiners: A Weighted Importance Multi-Point Evaluation Framework for Generative Tasks with Long-form Answers
by: Yu, Guoxin, et al.
Published: (2026)
by: Yu, Guoxin, et al.
Published: (2026)
Parrot: Enhancing Multi-Turn Instruction Following for Large Language Models
by: Sun, Yuchong, et al.
Published: (2023)
by: Sun, Yuchong, et al.
Published: (2023)
Plausible-Parrots @ MSP2023: Enhancing Semantic Plausibility Modeling using Entity and Event Knowledge
by: Shen, Chong, et al.
Published: (2024)
by: Shen, Chong, et al.
Published: (2024)
Large Language Models Can Self-Improve in Long-context Reasoning
by: Li, Siheng, et al.
Published: (2024)
by: Li, Siheng, et al.
Published: (2024)
On Probabilistic and Causal Reasoning with Summation Operators
by: Ibeling, Duligur, et al.
Published: (2024)
by: Ibeling, Duligur, et al.
Published: (2024)
Neither Stochastic Parroting nor AGI: LLMs Solve Tasks through Context-Directed Extrapolation from Training Data Priors
by: Madabushi, Harish Tayyar, et al.
Published: (2025)
by: Madabushi, Harish Tayyar, et al.
Published: (2025)
Similar Items
-
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
by: Chung, Tsz Ting, et al.
Published: (2025) -
Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models
by: Wu, Junjie, et al.
Published: (2024) -
SitEmb-v1.5: Improved Context-Aware Dense Retrieval for Semantic Association and Long Story Comprehension
by: Wu, Junjie, et al.
Published: (2025) -
Rethinking Targeted Adversarial Attacks For Neural Machine Translation
by: Wu, Junjie, et al.
Published: (2024) -
Understanding LLMs' Fluid Intelligence Deficiency: An Analysis of the ARC Task
by: Wu, Junjie, et al.
Published: (2025)