LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bean, Andrew M., Hellsten, Simi, Mayne, Harry, Magomere, Jabez, Chi, Ethan A., Chi, Ryan, Hale, Scott A., Kirk, Hannah Rose |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
von: Khouja, Jude, et al.
Veröffentlicht: (2025)
von: Khouja, Jude, et al.
Veröffentlicht: (2025)
Scaling Crowdsourced Election Monitoring: Construction and Evaluation of Classification Models for Multilingual and Cross-Domain Classification Settings
von: Magomere, Jabez, et al.
Veröffentlicht: (2025)
von: Magomere, Jabez, et al.
Veröffentlicht: (2025)
When Claims Evolve: Evaluating and Enhancing the Robustness of Embedding Models Against Misinformation Edits
von: Magomere, Jabez, et al.
Veröffentlicht: (2025)
von: Magomere, Jabez, et al.
Veröffentlicht: (2025)
Can LLMs Generate and Solve Linguistic Olympiad Puzzles?
von: Majmudar, Neh, et al.
Veröffentlicht: (2025)
von: Majmudar, Neh, et al.
Veröffentlicht: (2025)
Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models
von: Khandelwal, Khyati, et al.
Veröffentlicht: (2023)
von: Khandelwal, Khyati, et al.
Veröffentlicht: (2023)
Linguistics Olympiad
von: Neacșu, Vlad A.
Veröffentlicht: (2024)
von: Neacșu, Vlad A.
Veröffentlicht: (2024)
UNVEILING: What Makes Linguistics Olympiad Puzzles Tricky for LLMs?
von: Choudhary, Mukund, et al.
Veröffentlicht: (2025)
von: Choudhary, Mukund, et al.
Veröffentlicht: (2025)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking
von: Magomere, Jabez, et al.
Veröffentlicht: (2025)
von: Magomere, Jabez, et al.
Veröffentlicht: (2025)
modeLing: A Novel Dataset for Testing Linguistic Reasoning in Language Models
von: Chi, Nathan A., et al.
Veröffentlicht: (2024)
von: Chi, Nathan A., et al.
Veröffentlicht: (2024)
Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
von: Lugoloobi, William, et al.
Veröffentlicht: (2026)
von: Lugoloobi, William, et al.
Veröffentlicht: (2026)
Why human-AI relationships need socioaffective alignment
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2025)
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2025)
Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models
von: Sun, Haoxiang, et al.
Veröffentlicht: (2025)
von: Sun, Haoxiang, et al.
Veröffentlicht: (2025)
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
von: He, Chaoqun, et al.
Veröffentlicht: (2024)
von: He, Chaoqun, et al.
Veröffentlicht: (2024)
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
von: Zhu, Yaoming, et al.
Veröffentlicht: (2025)
von: Zhu, Yaoming, et al.
Veröffentlicht: (2025)
A Variational Approach for Mitigating Entity Bias in Relation Extraction
von: Mensah, Samuel, et al.
Veröffentlicht: (2025)
von: Mensah, Samuel, et al.
Veröffentlicht: (2025)
Establishing a Linguistic Olympiad in Spain, Year 1
von: Guillermo Latour
Veröffentlicht: (2014)
von: Guillermo Latour
Veröffentlicht: (2014)
EEFSUVA: A New Mathematical Olympiad Benchmark
von: Khatibi, Nicole N, et al.
Veröffentlicht: (2025)
von: Khatibi, Nicole N, et al.
Veröffentlicht: (2025)
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
von: Gao, Songyang, et al.
Veröffentlicht: (2025)
von: Gao, Songyang, et al.
Veröffentlicht: (2025)
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2026)
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2026)
Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2025)
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2025)
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
von: Vidgen, Bertie, et al.
Veröffentlicht: (2023)
von: Vidgen, Bertie, et al.
Veröffentlicht: (2023)
MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
von: Ji, Yuheng, et al.
Veröffentlicht: (2025)
von: Ji, Yuheng, et al.
Veröffentlicht: (2025)
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
von: Long, Yitao, et al.
Veröffentlicht: (2025)
von: Long, Yitao, et al.
Veröffentlicht: (2025)
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
von: Bean, Andrew M., et al.
Veröffentlicht: (2025)
von: Bean, Andrew M., et al.
Veröffentlicht: (2025)
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling
von: Li, Yafu, et al.
Veröffentlicht: (2026)
von: Li, Yafu, et al.
Veröffentlicht: (2026)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
von: Mayne, Harry, et al.
Veröffentlicht: (2025)
von: Mayne, Harry, et al.
Veröffentlicht: (2025)
RIMO: An Easy-to-Evaluate, Hard-to-Solve Olympiad Benchmark for Advanced Mathematical Reasoning
von: Chen, Ziye, et al.
Veröffentlicht: (2025)
von: Chen, Ziye, et al.
Veröffentlicht: (2025)
Cobordism and Concordance of Surfaces in 4-Manifolds
von: Hellsten, Simeon
Veröffentlicht: (2026)
von: Hellsten, Simeon
Veröffentlicht: (2026)
Probing Large Language Models in Reasoning and Translating Complex Linguistic Puzzles
von: Lin, Zheng-Lin, et al.
Veröffentlicht: (2025)
von: Lin, Zheng-Lin, et al.
Veröffentlicht: (2025)
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
von: Gao, Bofei, et al.
Veröffentlicht: (2024)
von: Gao, Bofei, et al.
Veröffentlicht: (2024)
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2024)
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2024)
Mastering Olympiad-Level Physics with Artificial Intelligence
von: Jian, Dong-Shan, et al.
Veröffentlicht: (2025)
von: Jian, Dong-Shan, et al.
Veröffentlicht: (2025)
CombiGraph-Vis: A Curated Multimodal Olympiad Benchmark for Discrete Mathematical Reasoning
von: Mahdavi, Hamed, et al.
Veröffentlicht: (2025)
von: Mahdavi, Hamed, et al.
Veröffentlicht: (2025)
PuzzleJAX: A Benchmark for Reasoning and Learning
von: Earle, Sam, et al.
Veröffentlicht: (2025)
von: Earle, Sam, et al.
Veröffentlicht: (2025)
LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
von: Zhang, Di, et al.
Veröffentlicht: (2024)
von: Zhang, Di, et al.
Veröffentlicht: (2024)
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
von: Yang, Yushi, et al.
Veröffentlicht: (2024)
von: Yang, Yushi, et al.
Veröffentlicht: (2024)
Linguistic Organisation and Native Title
von: Hale, Ken, et al.
Veröffentlicht: (2021)
von: Hale, Ken, et al.
Veröffentlicht: (2021)
Proving Olympiad Inequalities by Synergizing LLMs and Symbolic Reasoning
von: Li, Zenan, et al.
Veröffentlicht: (2025)
von: Li, Zenan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
von: Khouja, Jude, et al.
Veröffentlicht: (2025) -
Scaling Crowdsourced Election Monitoring: Construction and Evaluation of Classification Models for Multilingual and Cross-Domain Classification Settings
von: Magomere, Jabez, et al.
Veröffentlicht: (2025) -
When Claims Evolve: Evaluating and Enhancing the Robustness of Embedding Models Against Misinformation Edits
von: Magomere, Jabez, et al.
Veröffentlicht: (2025) -
Can LLMs Generate and Solve Linguistic Olympiad Puzzles?
von: Majmudar, Neh, et al.
Veröffentlicht: (2025) -
Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models
von: Khandelwal, Khyati, et al.
Veröffentlicht: (2023)