MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese
Fuente:
arXiv
Saved in:
| Main Authors: | Teixeira, Tiago, Erthal, Ana Carolina, Belieni, Juan, Canaverde, Beatriz, Mesquita, Diego, Faria, Miguel, da Silva, Eliezer de Souza, Martins, André F. T. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Search to Reasoning: A Five-Level RAG Capability Framework for Enterprise Data
by: Gill, Gurbinder, et al.
Published: (2025)
by: Gill, Gurbinder, et al.
Published: (2025)
ATANT v1.1: Positioning Continuity Evaluation Against Memory, Long-Context, and Agentic-Memory Benchmarks
by: Tanguturi, Samuel Sameer
Published: (2026)
by: Tanguturi, Samuel Sameer
Published: (2026)
From Brazilian Portuguese to European Portuguese
by: Sanches, João, et al.
Published: (2024)
by: Sanches, João, et al.
Published: (2024)
MUDY: Multi-Granular Dynamic Candidate Contextualization for Unsupervised Keyphrase Extraction
by: Kang, Hyeongu, et al.
Published: (2026)
by: Kang, Hyeongu, et al.
Published: (2026)
Using LLM-Based Approaches to Enhance and Automate Topic Labeling
by: Khandelwal, Trishia
Published: (2025)
by: Khandelwal, Trishia
Published: (2025)
Leveraging Large Language Models to Extract and Translate Medical Information in Doctors' Notes for Health Records and Diagnostic Billing Codes
by: Hartnett, Peter, et al.
Published: (2026)
by: Hartnett, Peter, et al.
Published: (2026)
Attention-based sequential recommendation system using multimodal data
by: Oh, Hyungtaik, et al.
Published: (2024)
by: Oh, Hyungtaik, et al.
Published: (2024)
MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
by: Bian, Zhipeng, et al.
Published: (2025)
by: Bian, Zhipeng, et al.
Published: (2025)
ATANT: An Evaluation Framework for AI Continuity
by: Tanguturi, Samuel Sameer
Published: (2026)
by: Tanguturi, Samuel Sameer
Published: (2026)
Mixture of Experts Approaches in Dense Retrieval Tasks
by: Sokli, Effrosyni, et al.
Published: (2025)
by: Sokli, Effrosyni, et al.
Published: (2025)
A Language Model based Framework for New Concept Placement in Ontologies
by: Dong, Hang, et al.
Published: (2024)
by: Dong, Hang, et al.
Published: (2024)
Introducing Three New Benchmark Datasets for Hierarchical Text Classification
by: Toit, Jaco du, et al.
Published: (2024)
by: Toit, Jaco du, et al.
Published: (2024)
Smart ETL and LLM-based contents classification: the European Smart Tourism Tools Observatory experience
by: Cosme, Diogo, et al.
Published: (2024)
by: Cosme, Diogo, et al.
Published: (2024)
From Knowledge Generation to Knowledge Verification: Examining the BioMedical Generative Capabilities of ChatGPT
by: Hamed, Ahmed Abdeen, et al.
Published: (2025)
by: Hamed, Ahmed Abdeen, et al.
Published: (2025)
LLM Reasoning for Cold-Start Item Recommendation
by: Li, Shijun, et al.
Published: (2025)
by: Li, Shijun, et al.
Published: (2025)
Suppressing Domain-Specific Hallucination in Construction LLMs: A Knowledge Graph Foundation for GraphRAG and QLoRA on River and Sediment Control Technical Standards
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
Documentation Retrieval Improves Planning Language Generation
by: Wang, Renxiang, et al.
Published: (2025)
by: Wang, Renxiang, et al.
Published: (2025)
Detection of ChatGPT Fake Science with the xFakeSci Learning Algorithm
by: Hamed, Ahmed Abdeen, et al.
Published: (2023)
by: Hamed, Ahmed Abdeen, et al.
Published: (2023)
What Matters in LLM-Based Feature Extractor for Recommender? A Systematic Analysis of Prompts, Models, and Adaptation
by: Shi, Kainan, et al.
Published: (2025)
by: Shi, Kainan, et al.
Published: (2025)
BridgeRAG: Training-Free Bridge-Conditioned Retrieval for Multi-Hop Question Answering
by: Bacellar, Andre
Published: (2026)
by: Bacellar, Andre
Published: (2026)
A Graph-based RAG for Energy Efficiency Question Answering
by: Campi, Riccardo, et al.
Published: (2025)
by: Campi, Riccardo, et al.
Published: (2025)
PLUGH: A Benchmark for Spatial Understanding and Reasoning in Large Language Models
by: Tikhonov, Alexey
Published: (2024)
by: Tikhonov, Alexey
Published: (2024)
EasyMath: A 0-shot Math Benchmark for SLMs
by: Karki, Drishya, et al.
Published: (2025)
by: Karki, Drishya, et al.
Published: (2025)
HiPS: Hierarchical PDF Segmentation of Textbooks
by: Wehnert, Sabine, et al.
Published: (2025)
by: Wehnert, Sabine, et al.
Published: (2025)
AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
by: Simplício, Afonso, et al.
Published: (2026)
by: Simplício, Afonso, et al.
Published: (2026)
MasterSet: A Large-Scale Benchmark for Must-Cite Citation Recommendation in the AI/ML Literature
by: Ratul, Md Toyaha Rahman, et al.
Published: (2026)
by: Ratul, Md Toyaha Rahman, et al.
Published: (2026)
Algorithmic Trust and Compliance: Benchmarking Brand Notability for UK iGaming Entities in Generative Search Engines
by: Oruesagasti, Julen
Published: (2026)
by: Oruesagasti, Julen
Published: (2026)
KGiRAG: An Iterative GraphRAG Approach for Responding Sensemaking Queries
by: Iacob, Isabela, et al.
Published: (2026)
by: Iacob, Isabela, et al.
Published: (2026)
SURE-RAG: Sufficiency and Uncertainty-Aware Evidence Verification for Selective Retrieval-Augmented Generation
by: Qiu, Jingxi, et al.
Published: (2026)
by: Qiu, Jingxi, et al.
Published: (2026)
Teaching a Language Model to Speak the Language of Tools
by: Emanuilov, Simeon
Published: (2025)
by: Emanuilov, Simeon
Published: (2025)
Comparison of Metadata Representation Models for Knowledge Graph Embeddings
by: Egami, Shusaku, et al.
Published: (2025)
by: Egami, Shusaku, et al.
Published: (2025)
CogCanvas: Verbatim-Grounded Artifact Extraction for Long LLM Conversations
by: An, Tao
Published: (2025)
by: An, Tao
Published: (2025)
Graph Your Way to Inspiration: Integrating Co-Author Graphs with Retrieval-Augmented Generation for Large Language Model Based Scientific Idea Generation
by: Xie, Pengzhen, et al.
Published: (2025)
by: Xie, Pengzhen, et al.
Published: (2025)
Knowledge-Aware Iterative Retrieval for Multi-Agent Systems
by: Song, Seyoung
Published: (2025)
by: Song, Seyoung
Published: (2025)
Document Understanding for Healthcare Referrals
by: Mistry, Jimit, et al.
Published: (2023)
by: Mistry, Jimit, et al.
Published: (2023)
KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?
by: Saha, Soumadeep, et al.
Published: (2025)
by: Saha, Soumadeep, et al.
Published: (2025)
InfinityMATH: A Scalable Instruction Tuning Dataset in Programmatic Mathematical Reasoning
by: Zhang, Bo-Wen, et al.
Published: (2024)
by: Zhang, Bo-Wen, et al.
Published: (2024)
Uncovering the Limitations of Query Performance Prediction: Failures, Insights, and Implications for Selective Query Processing
by: Chifu, Adrian-Gabriel, et al.
Published: (2025)
by: Chifu, Adrian-Gabriel, et al.
Published: (2025)
FlexStructRAG: Flexible Structure-Aware Multi-Granular Relational Retrieval for RAG
by: Chen, Mengzhu, et al.
Published: (2026)
by: Chen, Mengzhu, et al.
Published: (2026)
GraphCompliance: Aligning Policy and Context Graphs for LLM-Based Regulatory Compliance
by: Chung, Jiseong, et al.
Published: (2025)
by: Chung, Jiseong, et al.
Published: (2025)
Similar Items
-
From Search to Reasoning: A Five-Level RAG Capability Framework for Enterprise Data
by: Gill, Gurbinder, et al.
Published: (2025) -
ATANT v1.1: Positioning Continuity Evaluation Against Memory, Long-Context, and Agentic-Memory Benchmarks
by: Tanguturi, Samuel Sameer
Published: (2026) -
From Brazilian Portuguese to European Portuguese
by: Sanches, João, et al.
Published: (2024) -
MUDY: Multi-Granular Dynamic Candidate Contextualization for Unsupervised Keyphrase Extraction
by: Kang, Hyeongu, et al.
Published: (2026) -
Using LLM-Based Approaches to Enhance and Automate Topic Labeling
by: Khandelwal, Trishia
Published: (2025)