Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving
Fuente:
arXiv
Salvato in:
| Autori principali: | Zheng, Shunfeng, Zhang, Yudi, Fang, Meng, Zhang, Zihan, Wu, Zhitan, Pechenizkiy, Mykola, Chen, Ling |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Large Language Models Are Neurosymbolic Reasoners
di: Fang, Meng, et al.
Pubblicazione: (2024)
di: Fang, Meng, et al.
Pubblicazione: (2024)
RetrievalQA: Assessing Adaptive Retrieval-Augmented Generation for Short-form Open-Domain Question Answering
di: Zhang, Zihan, et al.
Pubblicazione: (2024)
di: Zhang, Zihan, et al.
Pubblicazione: (2024)
MedINST: Meta Dataset of Biomedical Instructions
di: Han, Wenhan, et al.
Pubblicazione: (2024)
di: Han, Wenhan, et al.
Pubblicazione: (2024)
PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments
di: Schipper, Olivier, et al.
Pubblicazione: (2025)
di: Schipper, Olivier, et al.
Pubblicazione: (2025)
Everyone deserves their voice to be heard: Analyzing Predictive Gender Bias in ASR Models Applied to Dutch Speech Data
di: Raes, Rik, et al.
Pubblicazione: (2024)
di: Raes, Rik, et al.
Pubblicazione: (2024)
Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?
di: Zhang, Yudi, et al.
Pubblicazione: (2025)
di: Zhang, Yudi, et al.
Pubblicazione: (2025)
HASARD: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents
di: Tomilin, Tristan, et al.
Pubblicazione: (2025)
di: Tomilin, Tristan, et al.
Pubblicazione: (2025)
One Model for All: Multi-Objective Controllable Language Models
di: He, Qiang, et al.
Pubblicazione: (2026)
di: He, Qiang, et al.
Pubblicazione: (2026)
RuAG: Learned-rule-augmented Generation for Large Language Models
di: Zhang, Yudi, et al.
Pubblicazione: (2024)
di: Zhang, Yudi, et al.
Pubblicazione: (2024)
Cooperation on the Fly: Exploring Language Agents for Ad Hoc Teamwork in the Avalon Game
di: Shi, Zijing, et al.
Pubblicazione: (2023)
di: Shi, Zijing, et al.
Pubblicazione: (2023)
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
di: Han, Wenhan, et al.
Pubblicazione: (2025)
di: Han, Wenhan, et al.
Pubblicazione: (2025)
XRAG: eXamining the Core -- Benchmarking Foundational Components in Advanced Retrieval-Augmented Generation
di: Mao, Qianren, et al.
Pubblicazione: (2024)
di: Mao, Qianren, et al.
Pubblicazione: (2024)
On Adversarial Robustness of Language Models in Transfer Learning
di: Turbal, Bohdan, et al.
Pubblicazione: (2024)
di: Turbal, Bohdan, et al.
Pubblicazione: (2024)
Spiral of Silence in Large Language Model Agents
di: Zhong, Mingze, et al.
Pubblicazione: (2025)
di: Zhong, Mingze, et al.
Pubblicazione: (2025)
Knowledge Pyramid Construction for Multi-Level Retrieval-Augmented Generation
di: Chen, Rubing, et al.
Pubblicazione: (2024)
di: Chen, Rubing, et al.
Pubblicazione: (2024)
Multiple Abstraction Level Retrieve Augment Generation
di: Zheng, Zheng, et al.
Pubblicazione: (2025)
di: Zheng, Zheng, et al.
Pubblicazione: (2025)
Physics Reasoner: Knowledge-Augmented Reasoning for Solving Physics Problems with Large Language Models
di: Pang, Xinyu, et al.
Pubblicazione: (2024)
di: Pang, Xinyu, et al.
Pubblicazione: (2024)
Benchmarking Retrieval-Augmented Generation for Medicine
di: Xiong, Guangzhi, et al.
Pubblicazione: (2024)
di: Xiong, Guangzhi, et al.
Pubblicazione: (2024)
KiRAG: Knowledge-Driven Iterative Retriever for Enhancing Retrieval-Augmented Generation
di: Fang, Jinyuan, et al.
Pubblicazione: (2025)
di: Fang, Jinyuan, et al.
Pubblicazione: (2025)
Legal-DC: Benchmarking Retrieval-Augmented Generation for Legal Documents
di: Li, Yaocong, et al.
Pubblicazione: (2026)
di: Li, Yaocong, et al.
Pubblicazione: (2026)
Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
di: Luo, Qi, et al.
Pubblicazione: (2025)
di: Luo, Qi, et al.
Pubblicazione: (2025)
MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation
di: Ji, Yuelyu, et al.
Pubblicazione: (2026)
di: Ji, Yuelyu, et al.
Pubblicazione: (2026)
Benchmarking Retrieval-Augmented Generation for Chemistry
di: Zhong, Xianrui, et al.
Pubblicazione: (2025)
di: Zhong, Xianrui, et al.
Pubblicazione: (2025)
Bi'an: A Bilingual Benchmark and Model for Hallucination Detection in Retrieval-Augmented Generation
di: Jiang, Zhouyu, et al.
Pubblicazione: (2025)
di: Jiang, Zhouyu, et al.
Pubblicazione: (2025)
Stable-RAG: Mitigating Retrieval-Permutation-Induced Hallucinations in Retrieval-Augmented Generation
di: Zhang, Qianchi, et al.
Pubblicazione: (2026)
di: Zhang, Qianchi, et al.
Pubblicazione: (2026)
ASVRI-Legal: Fine-Tuning LLMs with Retrieval Augmented Generation for Enhanced Legal Regulation
di: Octadion, One, et al.
Pubblicazione: (2025)
di: Octadion, One, et al.
Pubblicazione: (2025)
Corrective Retrieval Augmented Generation
di: Yan, Shi-Qi, et al.
Pubblicazione: (2024)
di: Yan, Shi-Qi, et al.
Pubblicazione: (2024)
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
di: Lin, Jingru, et al.
Pubblicazione: (2025)
di: Lin, Jingru, et al.
Pubblicazione: (2025)
SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems
di: Hazra, Rima, et al.
Pubblicazione: (2026)
di: Hazra, Rima, et al.
Pubblicazione: (2026)
OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving
di: Zhang, Xinyu, et al.
Pubblicazione: (2026)
di: Zhang, Xinyu, et al.
Pubblicazione: (2026)
ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
di: Banerjee, Somnath, et al.
Pubblicazione: (2025)
di: Banerjee, Somnath, et al.
Pubblicazione: (2025)
GFM-RAG: Graph Foundation Model for Retrieval Augmented Generation
di: Luo, Linhao, et al.
Pubblicazione: (2025)
di: Luo, Linhao, et al.
Pubblicazione: (2025)
Retrieval-Augmented Semantic Parsing: Improving Generalization with Lexical Knowledge
di: Zhang, Xiao, et al.
Pubblicazione: (2024)
di: Zhang, Xiao, et al.
Pubblicazione: (2024)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
di: Fang, Meng, et al.
Pubblicazione: (2024)
di: Fang, Meng, et al.
Pubblicazione: (2024)
FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented Generation
di: Zhang, Qinggang, et al.
Pubblicazione: (2025)
di: Zhang, Qinggang, et al.
Pubblicazione: (2025)
UncertaintyRAG: Span-Level Uncertainty Enhanced Long-Context Modeling for Retrieval-Augmented Generation
di: Li, Zixuan, et al.
Pubblicazione: (2024)
di: Li, Zixuan, et al.
Pubblicazione: (2024)
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
di: Gao, Songyang, et al.
Pubblicazione: (2025)
di: Gao, Songyang, et al.
Pubblicazione: (2025)
PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving
di: Feng, Kaiyue, et al.
Pubblicazione: (2025)
di: Feng, Kaiyue, et al.
Pubblicazione: (2025)
Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval
di: Chen, Peter Baile, et al.
Pubblicazione: (2024)
di: Chen, Peter Baile, et al.
Pubblicazione: (2024)
LevelRAG: Enhancing Retrieval-Augmented Generation with Multi-hop Logic Planning over Rewriting Augmented Searchers
di: Zhang, Zhuocheng, et al.
Pubblicazione: (2025)
di: Zhang, Zhuocheng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Large Language Models Are Neurosymbolic Reasoners
di: Fang, Meng, et al.
Pubblicazione: (2024) -
RetrievalQA: Assessing Adaptive Retrieval-Augmented Generation for Short-form Open-Domain Question Answering
di: Zhang, Zihan, et al.
Pubblicazione: (2024) -
MedINST: Meta Dataset of Biomedical Instructions
di: Han, Wenhan, et al.
Pubblicazione: (2024) -
PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments
di: Schipper, Olivier, et al.
Pubblicazione: (2025) -
Everyone deserves their voice to be heard: Analyzing Predictive Gender Bias in ASR Models Applied to Dutch Speech Data
di: Raes, Rik, et al.
Pubblicazione: (2024)