HLB: Benchmarking LLMs' Humanlikeness in Language Use
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Duan, Xufeng, Xiao, Bei, Tang, Xuemei, Cai, Zhenguang G. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do Large Language Models Plan Answer Positions? Position Bias in Multiple-Choice Question Generation
von: Tang, Xuemei, et al.
Veröffentlicht: (2026)
von: Tang, Xuemei, et al.
Veröffentlicht: (2026)
Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition
von: Tang, Xuemei, et al.
Veröffentlicht: (2024)
von: Tang, Xuemei, et al.
Veröffentlicht: (2024)
Unveiling Language Competence Neurons: A Psycholinguistic Approach to Model Interpretability
von: Duan, Xufeng, et al.
Veröffentlicht: (2024)
von: Duan, Xufeng, et al.
Veröffentlicht: (2024)
Information Loss in LLMs' Multilingual Translation: The Role of Training Data, Language Proximity, and Language Family
von: Lin, Yumeng, et al.
Veröffentlicht: (2025)
von: Lin, Yumeng, et al.
Veröffentlicht: (2025)
Grammaticality Representation in ChatGPT as Compared to Linguists and Laypeople
von: Qiu, Zhuang, et al.
Veröffentlicht: (2024)
von: Qiu, Zhuang, et al.
Veröffentlicht: (2024)
SCALPEL: Selective Capability Ablation via Low-rank Parameter Editing for Large Language Model Interpretability Analysis
von: Fu, Zihao, et al.
Veröffentlicht: (2026)
von: Fu, Zihao, et al.
Veröffentlicht: (2026)
Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment
von: Wu, Hanlin, et al.
Veröffentlicht: (2025)
von: Wu, Hanlin, et al.
Veröffentlicht: (2025)
MacBehaviour: An R package for behavioural experimentation on large language models
von: Duan, Xufeng, et al.
Veröffentlicht: (2024)
von: Duan, Xufeng, et al.
Veröffentlicht: (2024)
Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models
von: Zhou, Xinyu, et al.
Veröffentlicht: (2024)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2024)
How Syntax Specialization Emerges in Language Models
von: Duan, Xufeng, et al.
Veröffentlicht: (2025)
von: Duan, Xufeng, et al.
Veröffentlicht: (2025)
Do large language models resemble humans in language use?
von: Cai, Zhenguang G., et al.
Veröffentlicht: (2023)
von: Cai, Zhenguang G., et al.
Veröffentlicht: (2023)
Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
NoveltyBench: Evaluating Language Models for Humanlike Diversity
von: Zhang, Yiming, et al.
Veröffentlicht: (2025)
von: Zhang, Yiming, et al.
Veröffentlicht: (2025)
When a Man Says He Is Pregnant: Event-related Potential Evidence for a Rational Account of Speaker-contextualized Language Comprehension
von: Wu, Hanlin, et al.
Veröffentlicht: (2024)
von: Wu, Hanlin, et al.
Veröffentlicht: (2024)
Language Models Grow Less Humanlike beyond Phase Transition
von: Aoyama, Tatsuya, et al.
Veröffentlicht: (2025)
von: Aoyama, Tatsuya, et al.
Veröffentlicht: (2025)
Lil-Bevo: Explorations of Strategies for Training Language Models in More Humanlike Ways
von: Govindarajan, Venkata S, et al.
Veröffentlicht: (2023)
von: Govindarajan, Venkata S, et al.
Veröffentlicht: (2023)
Speaker effects in language comprehension: An integrative model of language and speaker processing
von: Wu, Hanlin, et al.
Veröffentlicht: (2024)
von: Wu, Hanlin, et al.
Veröffentlicht: (2024)
When AI companions become witty: Can human brain recognize AI-generated irony?
von: Rao, Xiaohui, et al.
Veröffentlicht: (2025)
von: Rao, Xiaohui, et al.
Veröffentlicht: (2025)
Small Language Models as Effective Guides for Large Language Models in Chinese Relation Extraction
von: Tang, Xuemei, et al.
Veröffentlicht: (2024)
von: Tang, Xuemei, et al.
Veröffentlicht: (2024)
Probabilistic adaptation of language comprehension for individual speakers: evidence from neural oscillations
von: Wu, Hanlin, et al.
Veröffentlicht: (2025)
von: Wu, Hanlin, et al.
Veröffentlicht: (2025)
A funny companion: Distinct neural responses to perceived AI- versus human-generated humor
von: Rao, Xiaohui, et al.
Veröffentlicht: (2025)
von: Rao, Xiaohui, et al.
Veröffentlicht: (2025)
Humanlike Multi-user Agent (HUMA): Designing a Deceptively Human AI Facilitator for Group Chats
von: Jacniacki, Mateusz, et al.
Veröffentlicht: (2025)
von: Jacniacki, Mateusz, et al.
Veröffentlicht: (2025)
ChatPattern: Layout Pattern Customization via Natural Language
von: Wang, Zixiao, et al.
Veröffentlicht: (2024)
von: Wang, Zixiao, et al.
Veröffentlicht: (2024)
CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
von: Fu, Zihao, et al.
Veröffentlicht: (2025)
von: Fu, Zihao, et al.
Veröffentlicht: (2025)
Structural priming: An experimental paradigm for mapping linguistic representations
von: Zhenguang G. Cai, et al.
Veröffentlicht: (2024)
von: Zhenguang G. Cai, et al.
Veröffentlicht: (2024)
MixRea: Benchmarking Explicit-Implicit Reasoning in Large Language Models
von: Cai, Yuanqing, et al.
Veröffentlicht: (2026)
von: Cai, Yuanqing, et al.
Veröffentlicht: (2026)
REAL: Response Embedding-based Alignment for LLMs
von: Zhang, Honggen, et al.
Veröffentlicht: (2024)
von: Zhang, Honggen, et al.
Veröffentlicht: (2024)
Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages
von: Naous, Tarek, et al.
Veröffentlicht: (2025)
von: Naous, Tarek, et al.
Veröffentlicht: (2025)
Are LLMs Really Not Knowledgeable? Mining the Submerged Knowledge in LLMs' Memory
von: Tao, Xingjian, et al.
Veröffentlicht: (2024)
von: Tao, Xingjian, et al.
Veröffentlicht: (2024)
HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
von: Wang, Kaixuan, et al.
Veröffentlicht: (2025)
von: Wang, Kaixuan, et al.
Veröffentlicht: (2025)
MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
von: Ruan, Junhao, et al.
Veröffentlicht: (2026)
von: Ruan, Junhao, et al.
Veröffentlicht: (2026)
Benchmarking Concept-Spilling Across Languages in LLMs
von: Badanin, Ilia, et al.
Veröffentlicht: (2026)
von: Badanin, Ilia, et al.
Veröffentlicht: (2026)
CLM-Bench: Benchmarking and Analyzing Cross-lingual Misalignment of LLMs in Knowledge Editing
von: Hu, Yucheng, et al.
Veröffentlicht: (2026)
von: Hu, Yucheng, et al.
Veröffentlicht: (2026)
The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks
von: Pomerenke, David, et al.
Veröffentlicht: (2025)
von: Pomerenke, David, et al.
Veröffentlicht: (2025)
IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages
von: Maheshwari, Ayush, et al.
Veröffentlicht: (2025)
von: Maheshwari, Ayush, et al.
Veröffentlicht: (2025)
MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases
von: Murthy, Rithesh, et al.
Veröffentlicht: (2024)
von: Murthy, Rithesh, et al.
Veröffentlicht: (2024)
Causal Autoregressive Diffusion Language Model
von: Ruan, Junhao, et al.
Veröffentlicht: (2026)
von: Ruan, Junhao, et al.
Veröffentlicht: (2026)
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models
von: Guo, Qianhong, et al.
Veröffentlicht: (2025)
von: Guo, Qianhong, et al.
Veröffentlicht: (2025)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
Chengyu-Bench: Benchmarking Large Language Models for Chinese Idiom Understanding and Use
von: Fu, Yicheng, et al.
Veröffentlicht: (2025)
von: Fu, Yicheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Do Large Language Models Plan Answer Positions? Position Bias in Multiple-Choice Question Generation
von: Tang, Xuemei, et al.
Veröffentlicht: (2026) -
Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition
von: Tang, Xuemei, et al.
Veröffentlicht: (2024) -
Unveiling Language Competence Neurons: A Psycholinguistic Approach to Model Interpretability
von: Duan, Xufeng, et al.
Veröffentlicht: (2024) -
Information Loss in LLMs' Multilingual Translation: The Role of Training Data, Language Proximity, and Language Family
von: Lin, Yumeng, et al.
Veröffentlicht: (2025) -
Grammaticality Representation in ChatGPT as Compared to Linguists and Laypeople
von: Qiu, Zhuang, et al.
Veröffentlicht: (2024)