Models Can and Should Embrace the Communicative Nature of Human-Generated Math
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Boguraev, Sasha, Lipkin, Ben, Weissweiler, Leonie, Mahowald, Kyle |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Causal Interventions Reveal Shared Structure Across English Filler-Gap Constructions
par: Boguraev, Sasha, et autres
Publié: (2025)
par: Boguraev, Sasha, et autres
Publié: (2025)
Causal Drawbridges: Characterizing Gradient Blocking of Syntactic Islands in Transformer LMs
par: Boguraev, Sasha, et autres
Publié: (2026)
par: Boguraev, Sasha, et autres
Publié: (2026)
France or Spain or Germany or France: A Neural Account of Non-Redundant Redundant Disjunctions
par: Boguraev, Sasha, et autres
Publié: (2026)
par: Boguraev, Sasha, et autres
Publié: (2026)
Linguistic Generalizations are not Rules: Impacts on Evaluation of LMs
par: Weissweiler, Leonie, et autres
Publié: (2025)
par: Weissweiler, Leonie, et autres
Publié: (2025)
Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models
par: Yao, Qing, et autres
Publié: (2025)
par: Yao, Qing, et autres
Publié: (2025)
Experimental Contexts Can Facilitate Robust Semantic Property Inference in Language Models, but Inconsistently
par: Misra, Kanishka, et autres
Publié: (2024)
par: Misra, Kanishka, et autres
Publié: (2024)
Constructions are Revealed in Word Distributions
par: Rozner, Joshua, et autres
Publié: (2025)
par: Rozner, Joshua, et autres
Publié: (2025)
The Counterexample Game: Iterated Conceptual Analysis and Repair in Language Models
par: Drucker, Daniel, et autres
Publié: (2026)
par: Drucker, Daniel, et autres
Publié: (2026)
Derivational Morphology Reveals Analogical Generalization in Large Language Models
par: Hofmann, Valentin, et autres
Publié: (2024)
par: Hofmann, Valentin, et autres
Publié: (2024)
Emergent Introspection in AI is Content-Agnostic
par: Lederman, Harvey, et autres
Publié: (2026)
par: Lederman, Harvey, et autres
Publié: (2026)
Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions
par: Scivetti, Wesley, et autres
Publié: (2026)
par: Scivetti, Wesley, et autres
Publié: (2026)
Language Models Fail to Introspect About Their Knowledge of Language
par: Song, Siyuan, et autres
Publié: (2025)
par: Song, Siyuan, et autres
Publié: (2025)
Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLP
par: Rozner, Josh, et autres
Publié: (2021)
par: Rozner, Josh, et autres
Publié: (2021)
What Can String Probability Tell Us About Grammaticality?
par: Hu, Jennifer, et autres
Publié: (2025)
par: Hu, Jennifer, et autres
Publié: (2025)
Privileged Self-Access Matters for Introspection in AI
par: Song, Siyuan, et autres
Publié: (2025)
par: Song, Siyuan, et autres
Publié: (2025)
semantic-features: A User-Friendly Tool for Studying Contextual Word Embeddings in Interpretable Semantic Spaces
par: Ranganathan, Jwalanthi, et autres
Publié: (2025)
par: Ranganathan, Jwalanthi, et autres
Publié: (2025)
Language models align with human judgments on key grammatical constructions
par: Hu, Jennifer, et autres
Publié: (2024)
par: Hu, Jennifer, et autres
Publié: (2024)
Mission: Impossible Language Models
par: Kallini, Julie, et autres
Publié: (2024)
par: Kallini, Julie, et autres
Publié: (2024)
Can LLMs Master Math? Investigating Large Language Models on Math Stack Exchange
par: Satpute, Ankit, et autres
Publié: (2024)
par: Satpute, Ankit, et autres
Publié: (2024)
OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization
par: Sun, Yiyou, et autres
Publié: (2025)
par: Sun, Yiyou, et autres
Publié: (2025)
Is It JUST Semantics? A Case Study of Discourse Particle Understanding in LLMs
par: Sheffield, William, et autres
Publié: (2025)
par: Sheffield, William, et autres
Publié: (2025)
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
par: Weissweiler, Leonie, et autres
Publié: (2024)
par: Weissweiler, Leonie, et autres
Publié: (2024)
Reliable Fine-Grained Evaluation of Natural Language Math Proofs
par: Ma, Wenjie, et autres
Publié: (2025)
par: Ma, Wenjie, et autres
Publié: (2025)
ControlMath: Controllable Data Generation Promotes Math Generalist Models
par: Chen, Nuo, et autres
Publié: (2024)
par: Chen, Nuo, et autres
Publié: (2024)
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
par: Liu, Xianyang, et autres
Publié: (2025)
par: Liu, Xianyang, et autres
Publié: (2025)
Adversarial Math Word Problem Generation
par: Xie, Roy, et autres
Publié: (2024)
par: Xie, Roy, et autres
Publié: (2024)
Can Vision-Language Models Solve Visual Math Equations?
par: Choudhury, Monjoy Narayan, et autres
Publié: (2025)
par: Choudhury, Monjoy Narayan, et autres
Publié: (2025)
You Can't Fight in Here! This is BBS!
par: Futrell, Richard, et autres
Publié: (2026)
par: Futrell, Richard, et autres
Publié: (2026)
TabularMath: Understanding Math Reasoning over Tables with Large Language Models
par: Tian, Shi-Yu, et autres
Publié: (2025)
par: Tian, Shi-Yu, et autres
Publié: (2025)
Dissociating language and thought in large language models
par: Mahowald, Kyle, et autres
Publié: (2023)
par: Mahowald, Kyle, et autres
Publié: (2023)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
par: Van Long, Phuoc Pham, et autres
Publié: (2023)
par: Van Long, Phuoc Pham, et autres
Publié: (2023)
LLMs Should Incorporate Explicit Mechanisms for Human Empathy
par: You, Xiaoxing, et autres
Publié: (2026)
par: You, Xiaoxing, et autres
Publié: (2026)
Math Neurosurgery: Isolating Language Models' Math Reasoning Abilities Using Only Forward Passes
par: Christ, Bryan R., et autres
Publié: (2024)
par: Christ, Bryan R., et autres
Publié: (2024)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
par: Balunović, Mislav, et autres
Publié: (2025)
par: Balunović, Mislav, et autres
Publié: (2025)
Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability
par: Wang, Ruida, et autres
Publié: (2025)
par: Wang, Ruida, et autres
Publié: (2025)
FANS -- Formal Answer Selection for Natural Language Math Reasoning Using Lean4
par: Yao, Jiarui, et autres
Publié: (2025)
par: Yao, Jiarui, et autres
Publié: (2025)
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
par: Ravi, Nikil, et autres
Publié: (2026)
par: Ravi, Nikil, et autres
Publié: (2026)
Orca-Math: Unlocking the potential of SLMs in Grade School Math
par: Mitra, Arindam, et autres
Publié: (2024)
par: Mitra, Arindam, et autres
Publié: (2024)
Should LLMs be WEIRD? Exploring WEIRDness and Human Rights in Large Language Models
par: Zhou, Ke, et autres
Publié: (2025)
par: Zhou, Ke, et autres
Publié: (2025)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
par: Fang, Meng, et autres
Publié: (2024)
par: Fang, Meng, et autres
Publié: (2024)
Documents similaires
-
Causal Interventions Reveal Shared Structure Across English Filler-Gap Constructions
par: Boguraev, Sasha, et autres
Publié: (2025) -
Causal Drawbridges: Characterizing Gradient Blocking of Syntactic Islands in Transformer LMs
par: Boguraev, Sasha, et autres
Publié: (2026) -
France or Spain or Germany or France: A Neural Account of Non-Redundant Redundant Disjunctions
par: Boguraev, Sasha, et autres
Publié: (2026) -
Linguistic Generalizations are not Rules: Impacts on Evaluation of LMs
par: Weissweiler, Leonie, et autres
Publié: (2025) -
Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models
par: Yao, Qing, et autres
Publié: (2025)