LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Peyronnet, Antoine, Gloeckle, Fabian, Hayat, Amaury |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AI Agents-as-Judge: Automated Assessment of Accuracy, Consistency, Completeness and Clarity for Enterprise Documents
von: Dasgupta, Sudip, et al.
Veröffentlicht: (2025)
von: Dasgupta, Sudip, et al.
Veröffentlicht: (2025)
Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
von: Gupta, Aayush
Veröffentlicht: (2025)
von: Gupta, Aayush
Veröffentlicht: (2025)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026)
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026)
DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text
von: Zhang, Liangliang, et al.
Veröffentlicht: (2025)
von: Zhang, Liangliang, et al.
Veröffentlicht: (2025)
GE-Chat: A Graph Enhanced RAG Framework for Evidential Response Generation of LLMs
von: Da, Longchao, et al.
Veröffentlicht: (2025)
von: Da, Longchao, et al.
Veröffentlicht: (2025)
BMAM: Brain-inspired Multi-Agent Memory Framework
von: Li, Yang, et al.
Veröffentlicht: (2026)
von: Li, Yang, et al.
Veröffentlicht: (2026)
Towards Robust Retrieval-Augmented Generation Based on Knowledge Graph: A Comparative Analysis
von: Amamou, Hazem, et al.
Veröffentlicht: (2026)
von: Amamou, Hazem, et al.
Veröffentlicht: (2026)
ByteRover: Agent-Native Memory Through LLM-Curated Hierarchical Context
von: Nguyen, Andy, et al.
Veröffentlicht: (2026)
von: Nguyen, Andy, et al.
Veröffentlicht: (2026)
EnterpriseRAG-Bench: A RAG Benchmark for Company Internal Knowledge
von: Sun, Yuhong, et al.
Veröffentlicht: (2026)
von: Sun, Yuhong, et al.
Veröffentlicht: (2026)
Combating data scarcity in recommendation services: Integrating cognitive types of VARK and neural network technologies (LLM)
von: Zmanovskii, Nikita
Veröffentlicht: (2026)
von: Zmanovskii, Nikita
Veröffentlicht: (2026)
A Case Study of Balanced Query Recommendation on Wikipedia
von: Mishra, Harshit, et al.
Veröffentlicht: (2025)
von: Mishra, Harshit, et al.
Veröffentlicht: (2025)
MeVer at CheckThat! 2026: Cluster-Aware Hard-Negative Mining for Multilingual Scientific-Source Retrieval
von: Bakagianni, Juli, et al.
Veröffentlicht: (2026)
von: Bakagianni, Juli, et al.
Veröffentlicht: (2026)
IndoBERT-Sentiment: Context-Conditioned Sentiment Classification for Indonesian Text
von: Saputra, Muhammad Apriandito Arya, et al.
Veröffentlicht: (2026)
von: Saputra, Muhammad Apriandito Arya, et al.
Veröffentlicht: (2026)
IndoBERT-Relevancy: A Context-Conditioned Relevancy Classifier for Indonesian Text
von: Saputra, Muhammad Apriandito Arya, et al.
Veröffentlicht: (2026)
von: Saputra, Muhammad Apriandito Arya, et al.
Veröffentlicht: (2026)
Optimizing Retrieval-Augmented Generation (RAG) for Colloquial Cantonese: A LoRA-Based Systematic Review
von: Calonge, David Santandreu, et al.
Veröffentlicht: (2025)
von: Calonge, David Santandreu, et al.
Veröffentlicht: (2025)
EdgeJury: Cross-Reviewed Small-Model Ensembles for Truthful Question Answering on Serverless Edge Inference
von: Kumar, Aayush
Veröffentlicht: (2025)
von: Kumar, Aayush
Veröffentlicht: (2025)
An Ensemble Embedding Approach for Improving Semantic Caching Performance in LLM-based Systems
von: Ghaffari, Shervin, et al.
Veröffentlicht: (2025)
von: Ghaffari, Shervin, et al.
Veröffentlicht: (2025)
Adaptive Multi-Stage Patent Claim Generation with Unified Quality Assessment
von: Liang, Chen-Wei, et al.
Veröffentlicht: (2026)
von: Liang, Chen-Wei, et al.
Veröffentlicht: (2026)
HumanMCP: A Human-Like Query Dataset for Evaluating MCP Tool Retrieval Performance
von: Laddha, Shubh, et al.
Veröffentlicht: (2025)
von: Laddha, Shubh, et al.
Veröffentlicht: (2025)
AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs
von: Perera, Manoj Madushanka, et al.
Veröffentlicht: (2026)
von: Perera, Manoj Madushanka, et al.
Veröffentlicht: (2026)
Mubeen AI: A Specialized Arabic Language Model for Heritage Preservation and User Intent Understanding
von: Aljafari, Mohammed, et al.
Veröffentlicht: (2025)
von: Aljafari, Mohammed, et al.
Veröffentlicht: (2025)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
PolyTruth: Multilingual Disinformation Detection using Transformer-Based Language Models
von: Gouliev, Zaur, et al.
Veröffentlicht: (2025)
von: Gouliev, Zaur, et al.
Veröffentlicht: (2025)
Beyond Long Context: When Semantics Matter More than Tokens
von: Chawdhury, Tarun Kumar, et al.
Veröffentlicht: (2025)
von: Chawdhury, Tarun Kumar, et al.
Veröffentlicht: (2025)
Advancing Explainability in Neural Machine Translation: Analytical Metrics for Attention and Alignment Consistency
von: Mishra, Anurag
Veröffentlicht: (2024)
von: Mishra, Anurag
Veröffentlicht: (2024)
ARISE: Agentic Rubric-Guided Iterative Survey Engine for Automated Scholarly Paper Generation
von: Wang, Zi, et al.
Veröffentlicht: (2025)
von: Wang, Zi, et al.
Veröffentlicht: (2025)
A Computational Approach to Modeling Conversational Systems: Analyzing Large-Scale Quasi-Patterned Dialogue Flows
von: Ammar, Mohamed Achref Ben, et al.
Veröffentlicht: (2025)
von: Ammar, Mohamed Achref Ben, et al.
Veröffentlicht: (2025)
When LLM meets Fuzzy-TOPSIS for Personnel Selection through Automated Profile Analysis
von: Hoque, Shahria, et al.
Veröffentlicht: (2026)
von: Hoque, Shahria, et al.
Veröffentlicht: (2026)
SPARQL Generation with Entity Pre-trained GPT for KG Question Answering
von: Bustamante, Diego, et al.
Veröffentlicht: (2024)
von: Bustamante, Diego, et al.
Veröffentlicht: (2024)
Kodezi Chronos: A Debugging-First Language Model for Repository-Scale Code Understanding
von: Khan, Ishraq, et al.
Veröffentlicht: (2025)
von: Khan, Ishraq, et al.
Veröffentlicht: (2025)
ORPHEAS: A Cross-Lingual Greek-English Embedding Model for Retrieval-Augmented Generation
von: Livieris, Ioannis E., et al.
Veröffentlicht: (2026)
von: Livieris, Ioannis E., et al.
Veröffentlicht: (2026)
Text-to-SQL based on Large Language Models and Database Keyword Search
von: Nascimento, Eduardo R., et al.
Veröffentlicht: (2025)
von: Nascimento, Eduardo R., et al.
Veröffentlicht: (2025)
A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation
von: Chen, Ziyang, et al.
Veröffentlicht: (2025)
von: Chen, Ziyang, et al.
Veröffentlicht: (2025)
Efficient fine-tuning methodology of text embedding models for information retrieval: contrastive learning penalty (clp)
von: Yu, Jeongsu
Veröffentlicht: (2024)
von: Yu, Jeongsu
Veröffentlicht: (2024)
A Multi-granularity Concept Sparse Activation and Hierarchical Knowledge Graph Fusion Framework for Rare Disease Diagnosis
von: Zhang, Mingda, et al.
Veröffentlicht: (2025)
von: Zhang, Mingda, et al.
Veröffentlicht: (2025)
Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
von: Sharma, Shubham, et al.
Veröffentlicht: (2025)
von: Sharma, Shubham, et al.
Veröffentlicht: (2025)
From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction
von: Petrov, Alex, et al.
Veröffentlicht: (2026)
von: Petrov, Alex, et al.
Veröffentlicht: (2026)
A Reproducible, Scalable Pipeline for Synthesizing Autoregressive Model Literature
von: Alpay, Faruk, et al.
Veröffentlicht: (2025)
von: Alpay, Faruk, et al.
Veröffentlicht: (2025)
TriAlignGR: Triangular Multitask Alignment with Multimodal Deep Interest Mining for Generative Recommendation
von: Zeng, Yangchen, et al.
Veröffentlicht: (2026)
von: Zeng, Yangchen, et al.
Veröffentlicht: (2026)
Key Algorithms for Keyphrase Generation: Instruction-Based LLMs for Russian Scientific Keyphrases
von: Glazkova, Anna, et al.
Veröffentlicht: (2024)
von: Glazkova, Anna, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AI Agents-as-Judge: Automated Assessment of Accuracy, Consistency, Completeness and Clarity for Enterprise Documents
von: Dasgupta, Sudip, et al.
Veröffentlicht: (2025) -
Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
von: Gupta, Aayush
Veröffentlicht: (2025) -
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026) -
DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text
von: Zhang, Liangliang, et al.
Veröffentlicht: (2025) -
GE-Chat: A Graph Enhanced RAG Framework for Evidential Response Generation of LLMs
von: Da, Longchao, et al.
Veröffentlicht: (2025)