Generating Benchmarks for Factuality Evaluation of Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Muhlgay, Dor, Ram, Ori, Magar, Inbal, Levine, Yoav, Ratner, Nir, Belinkov, Yonatan, Abend, Omri, Leyton-Brown, Kevin, Shashua, Amnon, Shoham, Yoav |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Understanding: A Pragmatic Framework Motivated by Large Language Models
by: Leyton-Brown, Kevin, et al.
Published: (2024)
by: Leyton-Brown, Kevin, et al.
Published: (2024)
Fundamental Limitations of Alignment in Large Language Models
by: Wolf, Yotam, et al.
Published: (2023)
by: Wolf, Yotam, et al.
Published: (2023)
Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods
by: Wolf, Yotam, et al.
Published: (2024)
by: Wolf, Yotam, et al.
Published: (2024)
Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
by: Ashuach, Tomer, et al.
Published: (2026)
by: Ashuach, Tomer, et al.
Published: (2026)
Hyperproperty-Preserving Register Specifications (Extended Version)
by: Shimon, Yoav Ben, et al.
Published: (2024)
by: Shimon, Yoav Ben, et al.
Published: (2024)
STEER: Assessing the Economic Rationality of Large Language Models
by: Raman, Narun, et al.
Published: (2024)
by: Raman, Narun, et al.
Published: (2024)
Jamba: A Hybrid Transformer-Mamba Language Model
by: Lieber, Opher, et al.
Published: (2024)
by: Lieber, Opher, et al.
Published: (2024)
Expected Utility Networks
by: La Mura, Pierfrancesco, et al.
Published: (2013)
by: La Mura, Pierfrancesco, et al.
Published: (2013)
From Reasoning to Super-Intelligence: A Search-Theoretic Perspective
by: Shalev-Shwartz, Shai, et al.
Published: (2025)
by: Shalev-Shwartz, Shai, et al.
Published: (2025)
Learning Dynamics of RNNs in Closed-Loop Environments
by: Ger, Yoav, et al.
Published: (2025)
by: Ger, Yoav, et al.
Published: (2025)
Learning reveals invisible structure in low-rank RNNs
by: Ger, Yoav, et al.
Published: (2026)
by: Ger, Yoav, et al.
Published: (2026)
Artificial Expert Intelligence through PAC-reasoning
by: Shalev-Shwartz, Shai, et al.
Published: (2024)
by: Shalev-Shwartz, Shai, et al.
Published: (2024)
Limited Communications Distributed Optimization via Deep Unfolded Distributed ADMM
by: Noah, Yoav, et al.
Published: (2023)
by: Noah, Yoav, et al.
Published: (2023)
FormulaOne: Measuring the Depth of Algorithmic Reasoning Beyond Competitive Programming
by: Beniamini, Gal, et al.
Published: (2025)
by: Beniamini, Gal, et al.
Published: (2025)
Compositional Hardness of Code in Large Language Models -- A Probabilistic Perspective
by: Wolf, Yotam, et al.
Published: (2024)
by: Wolf, Yotam, et al.
Published: (2024)
Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias
by: Itzhak, Itay, et al.
Published: (2023)
by: Itzhak, Itay, et al.
Published: (2023)
Are crumpled sheets marginally stable?
by: Shohat, Dor, et al.
Published: (2024)
by: Shohat, Dor, et al.
Published: (2024)
$T^5Score$: A Methodology for Automatically Assessing the Quality of LLM Generated Multi-Document Topic Sets
by: Trainin, Itamar, et al.
Published: (2024)
by: Trainin, Itamar, et al.
Published: (2024)
Unsupervised Speech Segmentation: A General Approach Using Speech Language Models
by: Elmakies, Avishai, et al.
Published: (2025)
by: Elmakies, Avishai, et al.
Published: (2025)
Spatiotemporal Hierarchy of Slow Avalanches During Creep
by: Rudyak, Vladimir Yu., et al.
Published: (2025)
by: Rudyak, Vladimir Yu., et al.
Published: (2025)
Short‐term climatic oscillations versus long‐term delta propagation: Controls on sand transport into the deep Levant Basin since the Pliocene
by: Ido Sirota, et al.
Published: (2024)
by: Ido Sirota, et al.
Published: (2024)
Star Complexity of Parikh Images of Languages over Infinite Alphabets
by: Danieli, Yoav
Published: (2026)
by: Danieli, Yoav
Published: (2026)
Shock propagation through a local constriction
by: Heppner, Raz, et al.
Published: (2026)
by: Heppner, Raz, et al.
Published: (2026)
Connectivity between breeding sites, wintering areas, and migration routes in Common Terns ( Sterna hirundo ) breeding in the Western Palaearctic
by: Yosef Kiat, et al.
Published: (2026)
by: Yosef Kiat, et al.
Published: (2026)
Highly twisted diagrams
by: Lazarovich, Nir, et al.
Published: (2022)
by: Lazarovich, Nir, et al.
Published: (2022)
Express Your Doubts -- Probabilistic World Modeling Should not be Based on Token logprobs
by: Wagner, Eitan, et al.
Published: (2025)
by: Wagner, Eitan, et al.
Published: (2025)
Exploring the Learning Capabilities of Language Models using LEVERWORLDS
by: Wagner, Eitan, et al.
Published: (2024)
by: Wagner, Eitan, et al.
Published: (2024)
CONTESTS: a Framework for Consistency Testing of Span Probabilities in Language Models
by: Wagner, Eitan, et al.
Published: (2024)
by: Wagner, Eitan, et al.
Published: (2024)
What makes a language easy to deep-learn? Deep neural networks and humans similarly benefit from compositional structure
by: Galke, Lukas, et al.
Published: (2023)
by: Galke, Lukas, et al.
Published: (2023)
Concept-Best-Matching: Evaluating Compositionality in Emergent Communication
by: Carmeli, Boaz, et al.
Published: (2024)
by: Carmeli, Boaz, et al.
Published: (2024)
Reverse-Engineering the Retrieval Process in GenIR Models
by: Reusch, Anja, et al.
Published: (2025)
by: Reusch, Anja, et al.
Published: (2025)
ContraSim -- Analyzing Neural Representations Based on Contrastive Learning
by: Rahamim, Adir, et al.
Published: (2023)
by: Rahamim, Adir, et al.
Published: (2023)
Why Does Agentic Safety Fail to Generalize Across Tasks?
by: Slutzky, Yonatan, et al.
Published: (2026)
by: Slutzky, Yonatan, et al.
Published: (2026)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
by: Roth, Amit, et al.
Published: (2026)
by: Roth, Amit, et al.
Published: (2026)
Microscopic description of the intermittent dynamics driving logarithmic creep
by: Korchinski, Daniel J., et al.
Published: (2024)
by: Korchinski, Daniel J., et al.
Published: (2024)
Experimental Determination of the $D1$ Magic Wavelength for $^{40}$K
by: Kalifa, Guy Hay, et al.
Published: (2026)
by: Kalifa, Guy Hay, et al.
Published: (2026)
Efficient Decoding Methods for Language Models on Encrypted Data
by: Avitan, Matan, et al.
Published: (2025)
by: Avitan, Matan, et al.
Published: (2025)
A Charge Constraint in BMN
by: Zigdon, Yoav
Published: (2025)
by: Zigdon, Yoav
Published: (2025)
Electrolubrication in flowing liquid mixtures
by: Tsori, Yoav
Published: (2024)
by: Tsori, Yoav
Published: (2024)
Is the Goldman-Hodgkin-Katz equation universally true?
by: Green, Yoav
Published: (2025)
by: Green, Yoav
Published: (2025)
Similar Items
-
Understanding Understanding: A Pragmatic Framework Motivated by Large Language Models
by: Leyton-Brown, Kevin, et al.
Published: (2024) -
Fundamental Limitations of Alignment in Large Language Models
by: Wolf, Yotam, et al.
Published: (2023) -
Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods
by: Wolf, Yotam, et al.
Published: (2024) -
Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
by: Ashuach, Tomer, et al.
Published: (2026) -
Hyperproperty-Preserving Register Specifications (Extended Version)
by: Shimon, Yoav Ben, et al.
Published: (2024)