Large-scale cloze evaluation reveals that token prediction tasks are neither lexically nor semantically aligned
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jacobs, Cassandra L., Grobol, Loïc, Tsang, Alvin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Bayesian account of pronoun and neopronoun acquisition
von: Jacobs, Cassandra L., et al.
Veröffentlicht: (2025)
von: Jacobs, Cassandra L., et al.
Veröffentlicht: (2025)
On the scaling relationship between cloze probabilities and language model next-token prediction
von: Jacobs, Cassandra L., et al.
Veröffentlicht: (2026)
von: Jacobs, Cassandra L., et al.
Veröffentlicht: (2026)
Not a nuisance but a useful heuristic: Outlier dimensions favor frequent tokens in language models
von: Macocco, Iuri, et al.
Veröffentlicht: (2025)
von: Macocco, Iuri, et al.
Veröffentlicht: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
MeteoRA: Multiple-tasks Embedded LoRA for Large Language Models
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2025)
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2025)
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
von: Schmidt, Craig W., et al.
Veröffentlicht: (2025)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2025)
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
Mechanistic evaluation of Transformers and state space models
von: Arora, Aryaman, et al.
Veröffentlicht: (2025)
von: Arora, Aryaman, et al.
Veröffentlicht: (2025)
Predictive Simultaneous Interpretation: Harnessing Large Language Models for Democratizing Real-Time Multilingual Communication
von: Iida, Kurando, et al.
Veröffentlicht: (2024)
von: Iida, Kurando, et al.
Veröffentlicht: (2024)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
von: Ovcharov, Volodymyr
Veröffentlicht: (2026)
von: Ovcharov, Volodymyr
Veröffentlicht: (2026)
Plain language adaptations of biomedical text using LLMs: Comparision of evaluation metrics
von: Kocbek, Primoz, et al.
Veröffentlicht: (2025)
von: Kocbek, Primoz, et al.
Veröffentlicht: (2025)
Large Language Model (LLM) Bias Index -- LLMBI
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
Robustness of Large Language Models to Perturbations in Text
von: Singh, Ayush, et al.
Veröffentlicht: (2024)
von: Singh, Ayush, et al.
Veröffentlicht: (2024)
GanitBench: A bi-lingual benchmark for evaluating mathematical reasoning in Vision Language Models
von: Bandooni, Ashutosh, et al.
Veröffentlicht: (2025)
von: Bandooni, Ashutosh, et al.
Veröffentlicht: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
Contextualising Levels of Language Resourcedness that affect NLP tasks
von: Keet, C. Maria, et al.
Veröffentlicht: (2023)
von: Keet, C. Maria, et al.
Veröffentlicht: (2023)
Low-Resource Court Judgment Summarization for Common Law Systems
von: Liu, Shuaiqi, et al.
Veröffentlicht: (2024)
von: Liu, Shuaiqi, et al.
Veröffentlicht: (2024)
Inference to the Best Explanation in Large Language Models
von: Dalal, Dhairya, et al.
Veröffentlicht: (2024)
von: Dalal, Dhairya, et al.
Veröffentlicht: (2024)
Streamlining Redundant Layers to Compress Large Language Models
von: Chen, Xiaodong, et al.
Veröffentlicht: (2024)
von: Chen, Xiaodong, et al.
Veröffentlicht: (2024)
A comprehensive taxonomy of hallucinations in Large Language Models
von: Cossio, Manuel
Veröffentlicht: (2025)
von: Cossio, Manuel
Veröffentlicht: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
Prompt-Time Symbolic Knowledge Capture with Large Language Models
von: Çöplü, Tolga, et al.
Veröffentlicht: (2024)
von: Çöplü, Tolga, et al.
Veröffentlicht: (2024)
Artificial Phantasia: Emergent Mental Imagery in Large Language Models
von: McCarty, Morgan, et al.
Veröffentlicht: (2025)
von: McCarty, Morgan, et al.
Veröffentlicht: (2025)
Argumentative Large Language Models for Explainable and Contestable Claim Verification
von: Freedman, Gabriel, et al.
Veröffentlicht: (2024)
von: Freedman, Gabriel, et al.
Veröffentlicht: (2024)
PatentGPT: A Large Language Model for Intellectual Property
von: Bai, Zilong, et al.
Veröffentlicht: (2024)
von: Bai, Zilong, et al.
Veröffentlicht: (2024)
Reasoning over Uncertain Text by Generative Large Language Models
von: Nafar, Aliakbar, et al.
Veröffentlicht: (2024)
von: Nafar, Aliakbar, et al.
Veröffentlicht: (2024)
Multi-Turn Interactions for Text-to-SQL with Large Language Models
von: Xiong, Guanming, et al.
Veröffentlicht: (2024)
von: Xiong, Guanming, et al.
Veröffentlicht: (2024)
Demystifying Instruction Mixing for Fine-tuning Large Language Models
von: Wang, Renxi, et al.
Veröffentlicht: (2023)
von: Wang, Renxi, et al.
Veröffentlicht: (2023)
EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models
von: Paech, Samuel J.
Veröffentlicht: (2023)
von: Paech, Samuel J.
Veröffentlicht: (2023)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
MIX : a Multi-task Learning Approach to Solve Open-Domain Question Answering
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2020)
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2020)
Search-R3: Unifying Reasoning and Embedding in Large Language Models
von: Gui, Yuntao, et al.
Veröffentlicht: (2025)
von: Gui, Yuntao, et al.
Veröffentlicht: (2025)
Can Large Language Models perform Relation-based Argument Mining?
von: Gorur, Deniz, et al.
Veröffentlicht: (2024)
von: Gorur, Deniz, et al.
Veröffentlicht: (2024)
EMNLP: Educator-role Moral and Normative Large Language Models Profiling
von: Jiang, Yilin, et al.
Veröffentlicht: (2025)
von: Jiang, Yilin, et al.
Veröffentlicht: (2025)
Integrating Emotional and Linguistic Models for Ethical Compliance in Large Language Models
von: Chang, Edward Y.
Veröffentlicht: (2024)
von: Chang, Edward Y.
Veröffentlicht: (2024)
Assistive Large Language Model Agents for Socially-Aware Negotiation Dialogues
von: Hua, Yuncheng, et al.
Veröffentlicht: (2024)
von: Hua, Yuncheng, et al.
Veröffentlicht: (2024)
PustakAI: Curriculum-Aligned and Interactive Textbooks Using Large Language Models
von: Sharma, Shivam, et al.
Veröffentlicht: (2025)
von: Sharma, Shivam, et al.
Veröffentlicht: (2025)
Prompt-Time Ontology-Driven Symbolic Knowledge Capture with Large Language Models
von: Çöplü, Tolga, et al.
Veröffentlicht: (2024)
von: Çöplü, Tolga, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Bayesian account of pronoun and neopronoun acquisition
von: Jacobs, Cassandra L., et al.
Veröffentlicht: (2025) -
On the scaling relationship between cloze probabilities and language model next-token prediction
von: Jacobs, Cassandra L., et al.
Veröffentlicht: (2026) -
Not a nuisance but a useful heuristic: Outlier dimensions favor frequent tokens in language models
von: Macocco, Iuri, et al.
Veröffentlicht: (2025) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025) -
MeteoRA: Multiple-tasks Embedded LoRA for Large Language Models
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)