General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Junlin, An, Shengnan, Zhou, Shuang, Ma, Dan, Luo, Shixiong, Xie, Ying, Zhang, Yuan, Yuan, Wenling, Zhou, Yifan, Li, Xiaoyu, Wang, Ziwen, Cao, Xuezhi, Cai, Xunliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
von: An, Shengnan, et al.
Veröffentlicht: (2025)
von: An, Shengnan, et al.
Veröffentlicht: (2025)
UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
von: Nie, Dujun, et al.
Veröffentlicht: (2026)
von: Nie, Dujun, et al.
Veröffentlicht: (2026)
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
von: Wang, Xihuai, et al.
Veröffentlicht: (2025)
von: Wang, Xihuai, et al.
Veröffentlicht: (2025)
Pancyclicity of graphs perturbed by a random $F$-factor
von: Mao, Dingjia, et al.
Veröffentlicht: (2026)
von: Mao, Dingjia, et al.
Veröffentlicht: (2026)
AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations
von: Jiayang, Cheng, et al.
Veröffentlicht: (2026)
von: Jiayang, Cheng, et al.
Veröffentlicht: (2026)
Benchmarking Real-World Medical Image Classification with Noisy Labels: Challenges, Practice, and Outlook
von: Ma, Yuan, et al.
Veröffentlicht: (2025)
von: Ma, Yuan, et al.
Veröffentlicht: (2025)
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
von: Zhu, Yaoming, et al.
Veröffentlicht: (2025)
von: Zhu, Yaoming, et al.
Veröffentlicht: (2025)
Diverse Target and Contribution Scheduling for Domain Generalization
von: Long, Shaocong, et al.
Veröffentlicht: (2023)
von: Long, Shaocong, et al.
Veröffentlicht: (2023)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
von: Guo, Zhengkang, et al.
Veröffentlicht: (2026)
von: Guo, Zhengkang, et al.
Veröffentlicht: (2026)
Chain-of-Thought Reasoning Without Prompting
von: Wang, Xuezhi, et al.
Veröffentlicht: (2024)
von: Wang, Xuezhi, et al.
Veröffentlicht: (2024)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
Learning Representations for Reasoning: Generalizing Across Diverse Structures
von: Zhu, Zhaocheng
Veröffentlicht: (2024)
von: Zhu, Zhaocheng
Veröffentlicht: (2024)
Meeseeks: A Feedback-Driven, Iterative Self-Correction Benchmark evaluating LLMs' Instruction Following Capability
von: wang, Jiaming, et al.
Veröffentlicht: (2025)
von: wang, Jiaming, et al.
Veröffentlicht: (2025)
Making Mathematical Reasoning Adaptive
von: Lai, Zhejian, et al.
Veröffentlicht: (2025)
von: Lai, Zhejian, et al.
Veröffentlicht: (2025)
Turán density of stars in uniformly dense hypergraphs
von: Lin, Hao, et al.
Veröffentlicht: (2025)
von: Lin, Hao, et al.
Veröffentlicht: (2025)
UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings
von: Zhao, Bo, et al.
Veröffentlicht: (2026)
von: Zhao, Bo, et al.
Veröffentlicht: (2026)
Instance-level Randomization: Toward More Stable LLM Evaluations
von: Li, Yiyang, et al.
Veröffentlicht: (2025)
von: Li, Yiyang, et al.
Veröffentlicht: (2025)
Benchmarking Critical Questions Generation: A Challenging Reasoning Task for Large Language Models
von: Figueras, Banca Calvo, et al.
Veröffentlicht: (2025)
von: Figueras, Banca Calvo, et al.
Veröffentlicht: (2025)
The Role of Visual Modality in Multimodal Mathematical Reasoning: Challenges and Insights
von: Liu, Yufang, et al.
Veröffentlicht: (2025)
von: Liu, Yufang, et al.
Veröffentlicht: (2025)
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences
von: Zhou, Yangyang, et al.
Veröffentlicht: (2026)
von: Zhou, Yangyang, et al.
Veröffentlicht: (2026)
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
von: Ding, Peng, et al.
Veröffentlicht: (2025)
von: Ding, Peng, et al.
Veröffentlicht: (2025)
MedAtlas: Evaluating LLMs for Multi-Round, Multi-Task Medical Reasoning Across Diverse Imaging Modalities and Clinical Text
von: Xu, Ronghao, et al.
Veröffentlicht: (2025)
von: Xu, Ronghao, et al.
Veröffentlicht: (2025)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
von: Ding, Peng, et al.
Veröffentlicht: (2024)
von: Ding, Peng, et al.
Veröffentlicht: (2024)
Distributional Robustness Bounds Generalization Errors
von: Wang, Shixiong, et al.
Veröffentlicht: (2022)
von: Wang, Shixiong, et al.
Veröffentlicht: (2022)
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
Benchmarking Robust Self-Supervised Learning Across Diverse Downstream Tasks
von: Kowalczuk, Antoni, et al.
Veröffentlicht: (2024)
von: Kowalczuk, Antoni, et al.
Veröffentlicht: (2024)
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
von: Lu, Yi, et al.
Veröffentlicht: (2025)
von: Lu, Yi, et al.
Veröffentlicht: (2025)
Generalizing Graph Transformers Across Diverse Graphs and Tasks via Pre-training
von: He, Yufei, et al.
Veröffentlicht: (2024)
von: He, Yufei, et al.
Veröffentlicht: (2024)
UniHetero: Could Generation Enhance Understanding for Vision-Language-Model at Large Data Scale?
von: Chen, Fengjiao, et al.
Veröffentlicht: (2025)
von: Chen, Fengjiao, et al.
Veröffentlicht: (2025)
Advancing Generalization Across a Variety of Abstract Visual Reasoning Tasks
von: Małkiński, Mikołaj, et al.
Veröffentlicht: (2025)
von: Małkiński, Mikołaj, et al.
Veröffentlicht: (2025)
DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding
von: Zhou, Jiaming, et al.
Veröffentlicht: (2026)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2026)
Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties
von: Wang, Zhenglin, et al.
Veröffentlicht: (2025)
von: Wang, Zhenglin, et al.
Veröffentlicht: (2025)
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle
von: Guan, Hao, et al.
Veröffentlicht: (2026)
von: Guan, Hao, et al.
Veröffentlicht: (2026)
ATLAS: All-round Testing of Long-context Abilities across Scales
von: Huang, Deli, et al.
Veröffentlicht: (2026)
von: Huang, Deli, et al.
Veröffentlicht: (2026)
Emotion-Aware Design: Modulating Valence, Arousal, and Dominance in Communication via Design
von: Cao, Shixiong, et al.
Veröffentlicht: (2025)
von: Cao, Shixiong, et al.
Veröffentlicht: (2025)
Neuro-Symbolic Data Generation for Math Reasoning
von: Li, Zenan, et al.
Veröffentlicht: (2024)
von: Li, Zenan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
von: An, Shengnan, et al.
Veröffentlicht: (2025) -
UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
von: Chen, Chen, et al.
Veröffentlicht: (2025) -
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
von: Nie, Dujun, et al.
Veröffentlicht: (2026) -
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
von: Wang, Xihuai, et al.
Veröffentlicht: (2025) -
Pancyclicity of graphs perturbed by a random $F$-factor
von: Mao, Dingjia, et al.
Veröffentlicht: (2026)