IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Maimon, Aviya, Cohen, Amir DN, Vishne, Gal, Ravfogel, Shauli, Tsarfaty, Reut |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Novel Computational and Modeling Foundation for Automatic Coherence Assessment
by: Maimon, Aviya, et al.
Published: (2023)
by: Maimon, Aviya, et al.
Published: (2023)
Diversity Over Quantity: A Lesson From Few Shot Relation Classification
by: Cohen, Amir DN, et al.
Published: (2024)
by: Cohen, Amir DN, et al.
Published: (2024)
HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark
by: Cohen, Amir DN, et al.
Published: (2025)
by: Cohen, Amir DN, et al.
Published: (2025)
Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP
by: Goldman, Omer, et al.
Published: (2024)
by: Goldman, Omer, et al.
Published: (2024)
Description-Based Text Similarity
by: Ravfogel, Shauli, et al.
Published: (2023)
by: Ravfogel, Shauli, et al.
Published: (2023)
LLMs' Reading Comprehension Is Affected by Parametric Knowledge and Struggles with Hypothetical Statements
by: Basmov, Victoria, et al.
Published: (2024)
by: Basmov, Victoria, et al.
Published: (2024)
Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds
by: Basmov, Victoria, et al.
Published: (2023)
by: Basmov, Victoria, et al.
Published: (2023)
Beyond English: The Impact of Prompt Translation Strategies across Languages and Tasks in Multilingual LLMs
by: Mondshine, Itai, et al.
Published: (2025)
by: Mondshine, Itai, et al.
Published: (2025)
Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text
by: Greenfeld, Refael Shaked, et al.
Published: (2026)
by: Greenfeld, Refael Shaked, et al.
Published: (2026)
Dicta-LM 3.0: Advancing The Frontier of Hebrew Sovereign LLMs
by: Shmidman, Shaltiel, et al.
Published: (2026)
by: Shmidman, Shaltiel, et al.
Published: (2026)
Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
by: Manevich, Avshalom, et al.
Published: (2024)
by: Manevich, Avshalom, et al.
Published: (2024)
Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization
by: Mondshine, Itai, et al.
Published: (2025)
by: Mondshine, Itai, et al.
Published: (2025)
Adapting LLMs to Hebrew: Unveiling DictaLM 2.0 with Enhanced Vocabulary and Instruction Capabilities
by: Shmidman, Shaltiel, et al.
Published: (2024)
by: Shmidman, Shaltiel, et al.
Published: (2024)
Log-linear Guardedness and its Implications
by: Ravfogel, Shauli, et al.
Published: (2022)
by: Ravfogel, Shauli, et al.
Published: (2022)
Intrinsic Test of Unlearning Using Parametric Knowledge Traces
by: Hong, Yihuai, et al.
Published: (2024)
by: Hong, Yihuai, et al.
Published: (2024)
A Truly Joint Neural Architecture for Segmentation and Parsing
by: Levi, Danit Yshaayahu, et al.
Published: (2024)
by: Levi, Danit Yshaayahu, et al.
Published: (2024)
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
by: Ben-Zaken, Elad, et al.
Published: (2021)
by: Ben-Zaken, Elad, et al.
Published: (2021)
Geometric Factual Recall in Transformers
by: Ravfogel, Shauli, et al.
Published: (2026)
by: Ravfogel, Shauli, et al.
Published: (2026)
Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
by: Rassin, Royi, et al.
Published: (2023)
by: Rassin, Royi, et al.
Published: (2023)
NoviCode: Generating Programs from Natural Language Utterances by Novices
by: Mordechai, Asaf Achi, et al.
Published: (2024)
by: Mordechai, Asaf Achi, et al.
Published: (2024)
A Practical Method for Generating String Counterfactuals
by: Avitan, Matan, et al.
Published: (2024)
by: Avitan, Matan, et al.
Published: (2024)
State over Tokens: Characterizing the Role of Reasoning Tokens
by: Levy, Mosh, et al.
Published: (2025)
by: Levy, Mosh, et al.
Published: (2025)
Kernelized Concept Erasure
by: Ravfogel, Shauli, et al.
Published: (2022)
by: Ravfogel, Shauli, et al.
Published: (2022)
Linear Adversarial Concept Erasure
by: Ravfogel, Shauli, et al.
Published: (2022)
by: Ravfogel, Shauli, et al.
Published: (2022)
MRL Parsing Without Tears: The Case of Hebrew
by: Shmidman, Shaltiel, et al.
Published: (2024)
by: Shmidman, Shaltiel, et al.
Published: (2024)
Superlatives in Context: Modeling the Implicit Semantics of Superlatives
by: Pyatkin, Valentina, et al.
Published: (2024)
by: Pyatkin, Valentina, et al.
Published: (2024)
Gumbel Counterfactual Generation From Language Models
by: Ravfogel, Shauli, et al.
Published: (2024)
by: Ravfogel, Shauli, et al.
Published: (2024)
RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
by: Petty, Jackson, et al.
Published: (2025)
by: Petty, Jackson, et al.
Published: (2025)
Emergence of Linear Truth Encodings in Language Models
by: Ravfogel, Shauli, et al.
Published: (2025)
by: Ravfogel, Shauli, et al.
Published: (2025)
Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
by: Chen, Hung-Ting, et al.
Published: (2025)
by: Chen, Hung-Ting, et al.
Published: (2025)
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
by: Shafran, Or, et al.
Published: (2026)
by: Shafran, Or, et al.
Published: (2026)
Measuring the Effect of Transcription Noise on Downstream Language Understanding Tasks
by: Shapira, Ori, et al.
Published: (2025)
by: Shapira, Ori, et al.
Published: (2025)
Can LLMs Introspect? A Reality Check
by: Singh, Shashwat, et al.
Published: (2026)
by: Singh, Shashwat, et al.
Published: (2026)
Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs
by: Mor-Lan, Guy, et al.
Published: (2026)
by: Mor-Lan, Guy, et al.
Published: (2026)
The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure
by: Fan, Yu, et al.
Published: (2025)
by: Fan, Yu, et al.
Published: (2025)
The Role of Language Imbalance in Cross-lingual Generalisation: Insights from Cloned Language Experiments
by: Schäfer, Anton, et al.
Published: (2024)
by: Schäfer, Anton, et al.
Published: (2024)
Do Pretrained Contextual Language Models Distinguish between Hebrew Homograph Analyses?
by: Shmidman, Avi, et al.
Published: (2024)
by: Shmidman, Avi, et al.
Published: (2024)
HeSum: a Novel Dataset for Abstractive Text Summarization in Hebrew
by: Paz-Argaman, Tzuf, et al.
Published: (2024)
by: Paz-Argaman, Tzuf, et al.
Published: (2024)
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
by: Goldman, Omer, et al.
Published: (2024)
by: Goldman, Omer, et al.
Published: (2024)
Representation Surgery: Theory and Practice of Affine Steering
by: Singh, Shashwat, et al.
Published: (2024)
by: Singh, Shashwat, et al.
Published: (2024)
Similar Items
-
A Novel Computational and Modeling Foundation for Automatic Coherence Assessment
by: Maimon, Aviya, et al.
Published: (2023) -
Diversity Over Quantity: A Lesson From Few Shot Relation Classification
by: Cohen, Amir DN, et al.
Published: (2024) -
HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark
by: Cohen, Amir DN, et al.
Published: (2025) -
Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP
by: Goldman, Omer, et al.
Published: (2024) -
Description-Based Text Similarity
by: Ravfogel, Shauli, et al.
Published: (2023)