Models Can and Should Embrace the Communicative Nature of Human-Generated Math
Fuente:
arXiv
Salvato in:
| Autori principali: | Boguraev, Sasha, Lipkin, Ben, Weissweiler, Leonie, Mahowald, Kyle |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Causal Interventions Reveal Shared Structure Across English Filler-Gap Constructions
di: Boguraev, Sasha, et al.
Pubblicazione: (2025)
di: Boguraev, Sasha, et al.
Pubblicazione: (2025)
Causal Drawbridges: Characterizing Gradient Blocking of Syntactic Islands in Transformer LMs
di: Boguraev, Sasha, et al.
Pubblicazione: (2026)
di: Boguraev, Sasha, et al.
Pubblicazione: (2026)
France or Spain or Germany or France: A Neural Account of Non-Redundant Redundant Disjunctions
di: Boguraev, Sasha, et al.
Pubblicazione: (2026)
di: Boguraev, Sasha, et al.
Pubblicazione: (2026)
Linguistic Generalizations are not Rules: Impacts on Evaluation of LMs
di: Weissweiler, Leonie, et al.
Pubblicazione: (2025)
di: Weissweiler, Leonie, et al.
Pubblicazione: (2025)
Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models
di: Yao, Qing, et al.
Pubblicazione: (2025)
di: Yao, Qing, et al.
Pubblicazione: (2025)
Experimental Contexts Can Facilitate Robust Semantic Property Inference in Language Models, but Inconsistently
di: Misra, Kanishka, et al.
Pubblicazione: (2024)
di: Misra, Kanishka, et al.
Pubblicazione: (2024)
Constructions are Revealed in Word Distributions
di: Rozner, Joshua, et al.
Pubblicazione: (2025)
di: Rozner, Joshua, et al.
Pubblicazione: (2025)
The Counterexample Game: Iterated Conceptual Analysis and Repair in Language Models
di: Drucker, Daniel, et al.
Pubblicazione: (2026)
di: Drucker, Daniel, et al.
Pubblicazione: (2026)
Derivational Morphology Reveals Analogical Generalization in Large Language Models
di: Hofmann, Valentin, et al.
Pubblicazione: (2024)
di: Hofmann, Valentin, et al.
Pubblicazione: (2024)
Emergent Introspection in AI is Content-Agnostic
di: Lederman, Harvey, et al.
Pubblicazione: (2026)
di: Lederman, Harvey, et al.
Pubblicazione: (2026)
Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions
di: Scivetti, Wesley, et al.
Pubblicazione: (2026)
di: Scivetti, Wesley, et al.
Pubblicazione: (2026)
Language Models Fail to Introspect About Their Knowledge of Language
di: Song, Siyuan, et al.
Pubblicazione: (2025)
di: Song, Siyuan, et al.
Pubblicazione: (2025)
Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLP
di: Rozner, Josh, et al.
Pubblicazione: (2021)
di: Rozner, Josh, et al.
Pubblicazione: (2021)
What Can String Probability Tell Us About Grammaticality?
di: Hu, Jennifer, et al.
Pubblicazione: (2025)
di: Hu, Jennifer, et al.
Pubblicazione: (2025)
Privileged Self-Access Matters for Introspection in AI
di: Song, Siyuan, et al.
Pubblicazione: (2025)
di: Song, Siyuan, et al.
Pubblicazione: (2025)
semantic-features: A User-Friendly Tool for Studying Contextual Word Embeddings in Interpretable Semantic Spaces
di: Ranganathan, Jwalanthi, et al.
Pubblicazione: (2025)
di: Ranganathan, Jwalanthi, et al.
Pubblicazione: (2025)
Language models align with human judgments on key grammatical constructions
di: Hu, Jennifer, et al.
Pubblicazione: (2024)
di: Hu, Jennifer, et al.
Pubblicazione: (2024)
Mission: Impossible Language Models
di: Kallini, Julie, et al.
Pubblicazione: (2024)
di: Kallini, Julie, et al.
Pubblicazione: (2024)
Can LLMs Master Math? Investigating Large Language Models on Math Stack Exchange
di: Satpute, Ankit, et al.
Pubblicazione: (2024)
di: Satpute, Ankit, et al.
Pubblicazione: (2024)
OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization
di: Sun, Yiyou, et al.
Pubblicazione: (2025)
di: Sun, Yiyou, et al.
Pubblicazione: (2025)
Is It JUST Semantics? A Case Study of Discourse Particle Understanding in LLMs
di: Sheffield, William, et al.
Pubblicazione: (2025)
di: Sheffield, William, et al.
Pubblicazione: (2025)
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
di: Weissweiler, Leonie, et al.
Pubblicazione: (2024)
di: Weissweiler, Leonie, et al.
Pubblicazione: (2024)
Reliable Fine-Grained Evaluation of Natural Language Math Proofs
di: Ma, Wenjie, et al.
Pubblicazione: (2025)
di: Ma, Wenjie, et al.
Pubblicazione: (2025)
ControlMath: Controllable Data Generation Promotes Math Generalist Models
di: Chen, Nuo, et al.
Pubblicazione: (2024)
di: Chen, Nuo, et al.
Pubblicazione: (2024)
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
di: Liu, Xianyang, et al.
Pubblicazione: (2025)
di: Liu, Xianyang, et al.
Pubblicazione: (2025)
Adversarial Math Word Problem Generation
di: Xie, Roy, et al.
Pubblicazione: (2024)
di: Xie, Roy, et al.
Pubblicazione: (2024)
Can Vision-Language Models Solve Visual Math Equations?
di: Choudhury, Monjoy Narayan, et al.
Pubblicazione: (2025)
di: Choudhury, Monjoy Narayan, et al.
Pubblicazione: (2025)
You Can't Fight in Here! This is BBS!
di: Futrell, Richard, et al.
Pubblicazione: (2026)
di: Futrell, Richard, et al.
Pubblicazione: (2026)
TabularMath: Understanding Math Reasoning over Tables with Large Language Models
di: Tian, Shi-Yu, et al.
Pubblicazione: (2025)
di: Tian, Shi-Yu, et al.
Pubblicazione: (2025)
Dissociating language and thought in large language models
di: Mahowald, Kyle, et al.
Pubblicazione: (2023)
di: Mahowald, Kyle, et al.
Pubblicazione: (2023)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
di: Van Long, Phuoc Pham, et al.
Pubblicazione: (2023)
di: Van Long, Phuoc Pham, et al.
Pubblicazione: (2023)
LLMs Should Incorporate Explicit Mechanisms for Human Empathy
di: You, Xiaoxing, et al.
Pubblicazione: (2026)
di: You, Xiaoxing, et al.
Pubblicazione: (2026)
Math Neurosurgery: Isolating Language Models' Math Reasoning Abilities Using Only Forward Passes
di: Christ, Bryan R., et al.
Pubblicazione: (2024)
di: Christ, Bryan R., et al.
Pubblicazione: (2024)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
di: Balunović, Mislav, et al.
Pubblicazione: (2025)
di: Balunović, Mislav, et al.
Pubblicazione: (2025)
Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability
di: Wang, Ruida, et al.
Pubblicazione: (2025)
di: Wang, Ruida, et al.
Pubblicazione: (2025)
FANS -- Formal Answer Selection for Natural Language Math Reasoning Using Lean4
di: Yao, Jiarui, et al.
Pubblicazione: (2025)
di: Yao, Jiarui, et al.
Pubblicazione: (2025)
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
di: Ravi, Nikil, et al.
Pubblicazione: (2026)
di: Ravi, Nikil, et al.
Pubblicazione: (2026)
Orca-Math: Unlocking the potential of SLMs in Grade School Math
di: Mitra, Arindam, et al.
Pubblicazione: (2024)
di: Mitra, Arindam, et al.
Pubblicazione: (2024)
Should LLMs be WEIRD? Exploring WEIRDness and Human Rights in Large Language Models
di: Zhou, Ke, et al.
Pubblicazione: (2025)
di: Zhou, Ke, et al.
Pubblicazione: (2025)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
di: Fang, Meng, et al.
Pubblicazione: (2024)
di: Fang, Meng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Causal Interventions Reveal Shared Structure Across English Filler-Gap Constructions
di: Boguraev, Sasha, et al.
Pubblicazione: (2025) -
Causal Drawbridges: Characterizing Gradient Blocking of Syntactic Islands in Transformer LMs
di: Boguraev, Sasha, et al.
Pubblicazione: (2026) -
France or Spain or Germany or France: A Neural Account of Non-Redundant Redundant Disjunctions
di: Boguraev, Sasha, et al.
Pubblicazione: (2026) -
Linguistic Generalizations are not Rules: Impacts on Evaluation of LMs
di: Weissweiler, Leonie, et al.
Pubblicazione: (2025) -
Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models
di: Yao, Qing, et al.
Pubblicazione: (2025)