LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
Fuente:
arXiv
Saved in:
| Main Authors: | Khouja, Jude, Yang, Lingyi, Korgul, Karolina, Hellsten, Simeon, Neacsu, Vlad A., Mayne, Harry, Kearns, Ryan Othniel, Bean, Andrew M., Mahdi, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages
by: Bean, Andrew M., et al.
Published: (2024)
by: Bean, Andrew M., et al.
Published: (2024)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025)
by: Mayne, Harry, et al.
Published: (2025)
Quantifying construct validity in large language model evaluations
by: Kearns, Ryan Othniel
Published: (2026)
by: Kearns, Ryan Othniel
Published: (2026)
Do Large Language Models have Shared Weaknesses in Medical Question Answering?
by: Bean, Andrew M., et al.
Published: (2023)
by: Bean, Andrew M., et al.
Published: (2023)
Cobordism and Concordance of Surfaces in 4-Manifolds
by: Hellsten, Simeon
Published: (2026)
by: Hellsten, Simeon
Published: (2026)
Linguistics Olympiad
by: Neacșu, Vlad A.
Published: (2024)
by: Neacșu, Vlad A.
Published: (2024)
Can sparse autoencoders be used to decompose and interpret steering vectors?
by: Mayne, Harry, et al.
Published: (2024)
by: Mayne, Harry, et al.
Published: (2024)
Unsupervised Learning Approaches for Identifying ICU Patient Subgroups: Do Results Generalise?
by: Mayne, Harry, et al.
Published: (2024)
by: Mayne, Harry, et al.
Published: (2024)
Theory-Grounded Evaluation of Human-Like Fallacy Patterns in LLM Reasoning
by: Richardson, Andrew Keenan, et al.
Published: (2025)
by: Richardson, Andrew Keenan, et al.
Published: (2025)
Fine-Tuning Adversarially-Robust Transformers for Single-Image Dehazing
by: Vasilescu, Vlad, et al.
Published: (2025)
by: Vasilescu, Vlad, et al.
Published: (2025)
Large language models can help boost food production, but be mindful of their risks
by: De Clercq, Djavan, et al.
Published: (2024)
by: De Clercq, Djavan, et al.
Published: (2024)
THE STRUGGLE FOR AUTONOMY: TOO MUCH, TOO LITTLE OR NOT AT ALL?
by: Zohar Lederman
Published: (2012)
by: Zohar Lederman
Published: (2012)
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
by: Yang, Yushi, et al.
Published: (2024)
by: Yang, Yushi, et al.
Published: (2024)
TOO MANY CONVICTS
Published: (2002)
Published: (2002)
TOO BLOODY TO IGNORE
Published: (2002)
Published: (2002)
The Metaplectic Representation is Faithful
by: Chang, Christopher, et al.
Published: (2024)
by: Chang, Christopher, et al.
Published: (2024)
On a new order of Eocene mammals
by: Marsh, Othniel Charles
Published: (1875)
by: Marsh, Othniel Charles
Published: (1875)
Orthographer: Generating Orthographic‐Style Projections for Elongated Architectural Structures
by: Yu‐Hsuan Hsieh, et al.
Published: (2026)
by: Yu‐Hsuan Hsieh, et al.
Published: (2026)
Agent Benchmarks Fail Public Sector Requirements
by: Rystrøm, Jonathan, et al.
Published: (2026)
by: Rystrøm, Jonathan, et al.
Published: (2026)
Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections
by: Borchmann, Łukasz, et al.
Published: (2026)
by: Borchmann, Łukasz, et al.
Published: (2026)
On the stability of the quaternion projective space
by: Neacşu, Crina-Daniela
Published: (2025)
by: Neacşu, Crina-Daniela
Published: (2025)
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
by: Mayne, Harry, et al.
Published: (2026)
by: Mayne, Harry, et al.
Published: (2026)
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
by: Korgul, Karolina, et al.
Published: (2025)
by: Korgul, Karolina, et al.
Published: (2025)
Evaluating the role of `Constitutions' for learning from AI feedback
by: Redgate, Saskia, et al.
Published: (2024)
by: Redgate, Saskia, et al.
Published: (2024)
Semantic, Orthographic, and Phonological Biases in Humans' Wordle Gameplay
by: Liang, Jiadong, et al.
Published: (2024)
by: Liang, Jiadong, et al.
Published: (2024)
AgentSLR: Automating Systematic Literature Reviews in Epidemiology with Agentic AI
by: Padarha, Shreyansh, et al.
Published: (2026)
by: Padarha, Shreyansh, et al.
Published: (2026)
PUNISHING THE WEST, ALL THINGS SECULAR AND EGYPTIANS TOO
Published: (2006)
Published: (2006)
Clarifying orthography: Orthographic transparency as compressibility
by: Torres, Charles J., et al.
Published: (2025)
by: Torres, Charles J., et al.
Published: (2025)
Modeling Orthographic Variation in Occitan's Dialects
by: Hopton, Zachary William, et al.
Published: (2024)
by: Hopton, Zachary William, et al.
Published: (2024)
Does Orthographic Variation Preclude Standardisation?
by: Nicholas Zair
Published: (2024)
by: Nicholas Zair
Published: (2024)
Geometrical isomorphisms between categories of fuzzy coverings and fuzzy partitions
by: Cimpoeas, Mircea, et al.
Published: (2022)
by: Cimpoeas, Mircea, et al.
Published: (2022)
Whispering in Norwegian: Navigating Orthographic and Dialectic Challenges
by: Kummervold, Per E, et al.
Published: (2024)
by: Kummervold, Per E, et al.
Published: (2024)
Globally Optimal Pose from Orthographic Silhouettes
by: Sengupta, Agniva, et al.
Published: (2026)
by: Sengupta, Agniva, et al.
Published: (2026)
Improving In-Context Learning with Small Language Model Ensembles
by: Mojarradi, M. Mehdi, et al.
Published: (2024)
by: Mojarradi, M. Mehdi, et al.
Published: (2024)
Living organic carbon stocks of the North Sea: data and calculations
by: Scheffold, Maike, et al.
Published: (2023)
by: Scheffold, Maike, et al.
Published: (2023)
Research Regarding the Design and Development of an Optimal Concept for a Humanoid Robot
by: Ileana Dugăeșescu, et al.
Published: (2024)
by: Ileana Dugăeșescu, et al.
Published: (2024)
CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-Tuning
by: Niu, Ke, et al.
Published: (2025)
by: Niu, Ke, et al.
Published: (2025)
PHILOSOPHY FOR CHILDREN AS A TEACHING MOVEMENT IN AN ERA OF TOO MUCH LEARNING
by: Charles Bingham
Published: (2015)
by: Charles Bingham
Published: (2015)
Evaluating Fine-Tuning Efficiency of Human-Inspired Learning Strategies in Medical Question Answering
by: Yang, Yushi, et al.
Published: (2024)
by: Yang, Yushi, et al.
Published: (2024)
Ensembles in Urban Large Eddy Simulations with Changing Wind Direction
by: Keskinen, Jukka-Pekka, et al.
Published: (2025)
by: Keskinen, Jukka-Pekka, et al.
Published: (2025)
Similar Items
-
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages
by: Bean, Andrew M., et al.
Published: (2024) -
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025) -
Quantifying construct validity in large language model evaluations
by: Kearns, Ryan Othniel
Published: (2026) -
Do Large Language Models have Shared Weaknesses in Medical Question Answering?
by: Bean, Andrew M., et al.
Published: (2023) -
Cobordism and Concordance of Surfaces in 4-Manifolds
by: Hellsten, Simeon
Published: (2026)