NUMINA: A Natural Understanding Benchmark for Multi-dimensional Intelligence and Numerical Reasoning Abilities
Fuente:
arXiv
Salvato in:
| Autori principali: | Zeng, Changyu, Wang, Yifan, Wang, Zimu, Wang, Wei, Yang, Zhengni, Bao, Muyi, Xiao, Jiming, Nguyen, Anh, Yue, Yutao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FTCFormer: Fuzzy Token Clustering Transformer for Image Classification
di: Bao, Muyi, et al.
Pubblicazione: (2025)
di: Bao, Muyi, et al.
Pubblicazione: (2025)
FinDebate: Multi-Agent Collaborative Intelligence for Financial Analysis
di: Cai, Tianshi, et al.
Pubblicazione: (2025)
di: Cai, Tianshi, et al.
Pubblicazione: (2025)
Can GRPO Boost Complex Multimodal Table Understanding?
di: Kang, Xiaoqiang, et al.
Pubblicazione: (2025)
di: Kang, Xiaoqiang, et al.
Pubblicazione: (2025)
ASP-VMUNet: Atrous Shifted Parallel Vision Mamba U-Net for Skin Lesion Segmentation
di: Bao, Muyi, et al.
Pubblicazione: (2025)
di: Bao, Muyi, et al.
Pubblicazione: (2025)
SEvenLLM: Benchmarking, Eliciting, and Enhancing Abilities of Large Language Models in Cyber Threat Intelligence
di: Ji, Hangyuan, et al.
Pubblicazione: (2024)
di: Ji, Hangyuan, et al.
Pubblicazione: (2024)
On Solutions for Singular Toda System on Riemann Surfaces with Boundary
di: Hu, Zhengni
Pubblicazione: (2024)
di: Hu, Zhengni
Pubblicazione: (2024)
Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning
di: Ren, Houxing, et al.
Pubblicazione: (2026)
di: Ren, Houxing, et al.
Pubblicazione: (2026)
Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents
di: Xing, Fuyu, et al.
Pubblicazione: (2025)
di: Xing, Fuyu, et al.
Pubblicazione: (2025)
A degree-counting formula for a Keller-Segel equation on a surface with boundary
di: Ahmedou, Mohameden, et al.
Pubblicazione: (2025)
di: Ahmedou, Mohameden, et al.
Pubblicazione: (2025)
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset
di: Wang, Weiqi, et al.
Pubblicazione: (2024)
di: Wang, Weiqi, et al.
Pubblicazione: (2024)
Domain-specific Guided Summarization for Mental Health Posts
di: Qian, Lu, et al.
Pubblicazione: (2024)
di: Qian, Lu, et al.
Pubblicazione: (2024)
Understanding Reasoning Ability of Language Models From the Perspective of Reasoning Paths Aggregation
di: Wang, Xinyi, et al.
Pubblicazione: (2024)
di: Wang, Xinyi, et al.
Pubblicazione: (2024)
Spatial Modeling and Risk Zoning of Global Extreme Precipitation via Graph Neural Networks and r-Pareto Processes
di: Wang, Zimu, et al.
Pubblicazione: (2025)
di: Wang, Zimu, et al.
Pubblicazione: (2025)
Structured Pruning for Diverse Best-of-N Reasoning Optimization
di: Nguyen, Hieu Trung, et al.
Pubblicazione: (2025)
di: Nguyen, Hieu Trung, et al.
Pubblicazione: (2025)
The Morse property of limit functions appearing in mean field equations on surfaces with boundary
di: Hu, Zhengni, et al.
Pubblicazione: (2024)
di: Hu, Zhengni, et al.
Pubblicazione: (2024)
Blow-up Solutions for General Toda Systems on Riemann Surfaces
di: Hu, Zhengni, et al.
Pubblicazione: (2026)
di: Hu, Zhengni, et al.
Pubblicazione: (2026)
Intelligent Holographic Antiglare Windshield System with Head‐up Display
di: Yuan Xu, et al.
Pubblicazione: (2025)
di: Yuan Xu, et al.
Pubblicazione: (2025)
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
di: Gui, Jiayi, et al.
Pubblicazione: (2024)
di: Gui, Jiayi, et al.
Pubblicazione: (2024)
Optimizing Language Model's Reasoning Abilities with Weak Supervision
di: Tong, Yongqi, et al.
Pubblicazione: (2024)
di: Tong, Yongqi, et al.
Pubblicazione: (2024)
ReasonRank: Empowering Passage Ranking with Strong Reasoning Ability
di: Liu, Wenhan, et al.
Pubblicazione: (2025)
di: Liu, Wenhan, et al.
Pubblicazione: (2025)
Low‐Rank Approximation of Gaussian Process With Normal‐Gamma Prior
di: Youjie Zeng, et al.
Pubblicazione: (2025)
di: Youjie Zeng, et al.
Pubblicazione: (2025)
MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
di: Zhang, Junkai, et al.
Pubblicazione: (2025)
di: Zhang, Junkai, et al.
Pubblicazione: (2025)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
di: Huang, Kaixuan, et al.
Pubblicazione: (2025)
di: Huang, Kaixuan, et al.
Pubblicazione: (2025)
Benchmarking MLLM-based Web Understanding: Reasoning, Robustness and Safety
di: Liu, Junliang, et al.
Pubblicazione: (2025)
di: Liu, Junliang, et al.
Pubblicazione: (2025)
Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs
di: Dan, Nifu, et al.
Pubblicazione: (2025)
di: Dan, Nifu, et al.
Pubblicazione: (2025)
Numerical Investigation Into Effect of Canyon Terrain Boundaries on the Seismic Response of Deep‐Water Bridge Piers
di: Haowei Cai, et al.
Pubblicazione: (2025)
di: Haowei Cai, et al.
Pubblicazione: (2025)
SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images
di: Truong, Bao, et al.
Pubblicazione: (2026)
di: Truong, Bao, et al.
Pubblicazione: (2026)
Neuro-Symbolic Artificial Intelligence: Towards Improving the Reasoning Abilities of Large Language Models
di: Yang, Xiao-Wen, et al.
Pubblicazione: (2025)
di: Yang, Xiao-Wen, et al.
Pubblicazione: (2025)
Reasoning Planning for Language Models
di: Nguyen, Bao, et al.
Pubblicazione: (2025)
di: Nguyen, Bao, et al.
Pubblicazione: (2025)
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
di: Liu, Chonghan, et al.
Pubblicazione: (2025)
di: Liu, Chonghan, et al.
Pubblicazione: (2025)
VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding
di: Do, Phong Nguyen-Thuan, et al.
Pubblicazione: (2024)
di: Do, Phong Nguyen-Thuan, et al.
Pubblicazione: (2024)
Decomposition-Driven Multi-Table Retrieval and Reasoning for Numerical Question Answering
di: Luo, Feng, et al.
Pubblicazione: (2026)
di: Luo, Feng, et al.
Pubblicazione: (2026)
Interpreting Social Bias in LVLMs via Information Flow Analysis and Multi-Round Dialogue Evaluation
di: Ji, Zhengyang, et al.
Pubblicazione: (2025)
di: Ji, Zhengyang, et al.
Pubblicazione: (2025)
MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
di: Ke, Zixuan, et al.
Pubblicazione: (2026)
di: Ke, Zixuan, et al.
Pubblicazione: (2026)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
di: Xu, Yicheng, et al.
Pubblicazione: (2025)
di: Xu, Yicheng, et al.
Pubblicazione: (2025)
TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
di: Wang, Haochuan, et al.
Pubblicazione: (2024)
di: Wang, Haochuan, et al.
Pubblicazione: (2024)
How Robust Are Large Language Models for Clinical Numeracy? An Empirical Study on Numerical Reasoning Abilities in Clinical Contexts
di: Nguyen, Minh-Vuong, et al.
Pubblicazione: (2026)
di: Nguyen, Minh-Vuong, et al.
Pubblicazione: (2026)
Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
di: Ni, Minheng, et al.
Pubblicazione: (2024)
di: Ni, Minheng, et al.
Pubblicazione: (2024)
Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning
di: Zhou, Yuhao, et al.
Pubblicazione: (2025)
di: Zhou, Yuhao, et al.
Pubblicazione: (2025)
The prescribed curvature problem for entire hypersurfaces in Minkowski space
di: Ren, Changyu, et al.
Pubblicazione: (2020)
di: Ren, Changyu, et al.
Pubblicazione: (2020)
Documenti analoghi
-
FTCFormer: Fuzzy Token Clustering Transformer for Image Classification
di: Bao, Muyi, et al.
Pubblicazione: (2025) -
FinDebate: Multi-Agent Collaborative Intelligence for Financial Analysis
di: Cai, Tianshi, et al.
Pubblicazione: (2025) -
Can GRPO Boost Complex Multimodal Table Understanding?
di: Kang, Xiaoqiang, et al.
Pubblicazione: (2025) -
ASP-VMUNet: Atrous Shifted Parallel Vision Mamba U-Net for Skin Lesion Segmentation
di: Bao, Muyi, et al.
Pubblicazione: (2025) -
SEvenLLM: Benchmarking, Eliciting, and Enhancing Abilities of Large Language Models in Cyber Threat Intelligence
di: Ji, Hangyuan, et al.
Pubblicazione: (2024)