From Abstract to Contextual: What LLMs Still Cannot Do in Mathematics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cao, Bowen, Zhang, Dongdong, Li, Yixia, Liu, Junpeng, Huang, Shijue, Shi, Chufan, Lu, Hongyuan, Wu, Yaokang, Chen, Guanhua, Lam, Wai, Wei, Furu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Chain-of-Dictionary Prompting Elicits Translation in Large Language Models
von: Lu, Hongyuan, et al.
Veröffentlicht: (2023)
von: Lu, Hongyuan, et al.
Veröffentlicht: (2023)
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
von: Kidder, William, et al.
Veröffentlicht: (2024)
von: Kidder, William, et al.
Veröffentlicht: (2024)
VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
von: Jiang, Lingjie, et al.
Veröffentlicht: (2025)
von: Jiang, Lingjie, et al.
Veröffentlicht: (2025)
Clean Evaluations on Contaminated Visual Language Models
von: Lu, Hongyuan, et al.
Veröffentlicht: (2024)
von: Lu, Hongyuan, et al.
Veröffentlicht: (2024)
Dictionary Insertion Prompting for Multilingual Reasoning on Multilingual Large Language Models
von: Lu, Hongyuan, et al.
Veröffentlicht: (2024)
von: Lu, Hongyuan, et al.
Veröffentlicht: (2024)
A Thorough Examination of Decoding Methods in the Era of LLMs
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
The Committee on Accreditation: What It Can and Cannot Do.
von: Kimmel, Margaret Mary
Veröffentlicht: (1987)
von: Kimmel, Margaret Mary
Veröffentlicht: (1987)
InfiniteICL: Breaking the Limit of Context Window Size via Long Short-term Memory Transformation
von: Cao, Bowen, et al.
Veröffentlicht: (2025)
von: Cao, Bowen, et al.
Veröffentlicht: (2025)
SLoW: Select Low-frequency Words! Automatic Dictionary Selection for Translation on Large Language Models
von: Lu, Hongyuan, et al.
Veröffentlicht: (2025)
von: Lu, Hongyuan, et al.
Veröffentlicht: (2025)
The Quiet Stir of Thought; or, What the Computer Cannot Do
von: Shera, Jesse H.
Veröffentlicht: (1969)
von: Shera, Jesse H.
Veröffentlicht: (1969)
SeTAR: Out-of-Distribution Detection with Selective Low-Rank Approximation
von: Li, Yixia, et al.
Veröffentlicht: (2024)
von: Li, Yixia, et al.
Veröffentlicht: (2024)
LLM2: Let Large Language Models Harness System 2 Reasoning
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
ContextVis: Envision Contextual Learning and Interaction with Generative Models
von: Shui, Bo, et al.
Veröffentlicht: (2024)
von: Shui, Bo, et al.
Veröffentlicht: (2024)
What Emergency Severity Index Can Do and Scoring Systems Cannot Do?
von: Amir Mirhaghi
Veröffentlicht: (2025)
von: Amir Mirhaghi
Veröffentlicht: (2025)
Adam's Law: Textual Frequency Law on Large Language Models
von: Lu, Hongyuan Adam, et al.
Veröffentlicht: (2026)
von: Lu, Hongyuan Adam, et al.
Veröffentlicht: (2026)
What Monads Can and Cannot Do with a Few Extra Pages
von: Møgelberg, Rasmus Ejlers, et al.
Veröffentlicht: (2023)
von: Møgelberg, Rasmus Ejlers, et al.
Veröffentlicht: (2023)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
von: Son, Guijin, et al.
Veröffentlicht: (2024)
von: Son, Guijin, et al.
Veröffentlicht: (2024)
ImPart: Importance-Aware Delta-Sparsification for Improved Model Compression and Merging in LLMs
von: Yang, Yan, et al.
Veröffentlicht: (2025)
von: Yang, Yan, et al.
Veröffentlicht: (2025)
From Word to World: Can Large Language Models be Implicit Text-based World Models?
von: Li, Yixia, et al.
Veröffentlicht: (2025)
von: Li, Yixia, et al.
Veröffentlicht: (2025)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
von: Shi, Chufan, et al.
Veröffentlicht: (2026)
von: Shi, Chufan, et al.
Veröffentlicht: (2026)
MiLoRA: Harnessing Minor Singular Components for Parameter-Efficient LLM Finetuning
von: Wang, Hanqing, et al.
Veröffentlicht: (2024)
von: Wang, Hanqing, et al.
Veröffentlicht: (2024)
PACIT: Unlocking the Power of Examples for Better In-Context Instruction Tuning
von: Xue, Tianci, et al.
Veröffentlicht: (2023)
von: Xue, Tianci, et al.
Veröffentlicht: (2023)
G2: Guided Generation for Enhanced Output Diversity in LLMs
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
What Voting Power Cannot Be
von: Daniel Wodak
Veröffentlicht: (2025)
von: Daniel Wodak
Veröffentlicht: (2025)
Probing Minimalist Phase Structure in LLMs: What Universal Dependencies Cannot Represent
von: Chen, Yuanhao, et al.
Veröffentlicht: (2026)
von: Chen, Yuanhao, et al.
Veröffentlicht: (2026)
Not All Metrics Are Guilty: Improving NLG Evaluation by Diversifying References
von: Tang, Tianyi, et al.
Veröffentlicht: (2023)
von: Tang, Tianyi, et al.
Veröffentlicht: (2023)
From Concrete to Abstract in Indian Mathematics
von: Dasgupta, Jaidev
Veröffentlicht: (2024)
von: Dasgupta, Jaidev
Veröffentlicht: (2024)
What Cannot Be Implemented on Weak Memory?
von: Castañeda, Armando, et al.
Veröffentlicht: (2024)
von: Castañeda, Armando, et al.
Veröffentlicht: (2024)
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
Focus Managers on What They Can Say, Not What They Cannot
Veröffentlicht: (2024)
Veröffentlicht: (2024)
Measuring What Cannot Be Surveyed: LLMs as Instruments for Latent Cognitive Variables in Labor Economics
von: Maya, Cristian Espinal
Veröffentlicht: (2026)
von: Maya, Cristian Espinal
Veröffentlicht: (2026)
Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2026)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2026)
Consecutive Batch Model Editing with HooK Layers
von: Li, Shuaiyi, et al.
Veröffentlicht: (2024)
von: Li, Shuaiyi, et al.
Veröffentlicht: (2024)
VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
von: Liu, Junpeng, et al.
Veröffentlicht: (2024)
von: Liu, Junpeng, et al.
Veröffentlicht: (2024)
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code
von: Zhao, Yuze, et al.
Veröffentlicht: (2026)
von: Zhao, Yuze, et al.
Veröffentlicht: (2026)
On the Worst Prompt Performance of Large Language Models
von: Cao, Bowen, et al.
Veröffentlicht: (2024)
von: Cao, Bowen, et al.
Veröffentlicht: (2024)
GRAVITY: Architecture-Agnostic Structured Anchoring for Long-Horizon Conversational Memory
von: Sun, Yushi, et al.
Veröffentlicht: (2026)
von: Sun, Yushi, et al.
Veröffentlicht: (2026)
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
von: Gao, Xin, et al.
Veröffentlicht: (2026)
von: Gao, Xin, et al.
Veröffentlicht: (2026)
Do BERT-Like Bidirectional Models Still Perform Better on Text Classification in the Era of LLMs?
von: Zhang, Junyan, et al.
Veröffentlicht: (2025)
von: Zhang, Junyan, et al.
Veröffentlicht: (2025)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Chain-of-Dictionary Prompting Elicits Translation in Large Language Models
von: Lu, Hongyuan, et al.
Veröffentlicht: (2023) -
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
von: Kidder, William, et al.
Veröffentlicht: (2024) -
VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
von: Jiang, Lingjie, et al.
Veröffentlicht: (2025) -
Clean Evaluations on Contaminated Visual Language Models
von: Lu, Hongyuan, et al.
Veröffentlicht: (2024) -
Dictionary Insertion Prompting for Multilingual Reasoning on Multilingual Large Language Models
von: Lu, Hongyuan, et al.
Veröffentlicht: (2024)