SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhan, Weidong, Wang, Yue, Hu, Nan, Xiao, Liming, Ma, Jingyuan, Qin, Yuhang, Li, Zheng, Yang, Yixin, Deng, Sirui, Ding, Jinkun, Ma, Wenhan, Li, Rui, Luo, Weilin, Liu, Qun, Sui, Zhifang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SCoRE Sets: A Versatile Framework for Simultaneous Inference
von: Telschow, Fabian J. E., et al.
Veröffentlicht: (2023)
von: Telschow, Fabian J. E., et al.
Veröffentlicht: (2023)
D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
von: Zhou, Weibo, et al.
Veröffentlicht: (2025)
von: Zhou, Weibo, et al.
Veröffentlicht: (2025)
SCoRE: Streamlined Corpus-based Relation Extraction using Multi-Label Contrastive Learning and Bayesian kNN
von: Mariotti, Luca, et al.
Veröffentlicht: (2025)
von: Mariotti, Luca, et al.
Veröffentlicht: (2025)
Enhancing Reliability across Short and Long-Form QA via Reinforcement Learning
von: Wang, Yudong, et al.
Veröffentlicht: (2025)
von: Wang, Yudong, et al.
Veröffentlicht: (2025)
CoLT: Reasoning with Chain of Latent Tool Calls
von: Zhu, Fangwei, et al.
Veröffentlicht: (2026)
von: Zhu, Fangwei, et al.
Veröffentlicht: (2026)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
von: Li, Zheng, et al.
Veröffentlicht: (2025)
von: Li, Zheng, et al.
Veröffentlicht: (2025)
Decoding in Geometry: Alleviating Embedding-Space Crowding for Complex Reasoning
von: Yang, Yixin, et al.
Veröffentlicht: (2026)
von: Yang, Yixin, et al.
Veröffentlicht: (2026)
Plug-and-Play Training Framework for Preference Optimization
von: Ma, Jingyuan, et al.
Veröffentlicht: (2024)
von: Ma, Jingyuan, et al.
Veröffentlicht: (2024)
Large Language Models Struggle with Unreasonability in Math Problems
von: Ma, Jingyuan, et al.
Veröffentlicht: (2024)
von: Ma, Jingyuan, et al.
Veröffentlicht: (2024)
HauntAttack: When Attack Follows Reasoning as a Shadow
von: Ma, Jingyuan, et al.
Veröffentlicht: (2025)
von: Ma, Jingyuan, et al.
Veröffentlicht: (2025)
Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
TeachBench: A Syllabus-Grounded Framework for Evaluating Teaching Ability in Large Language Models
von: Li, Zheng, et al.
Veröffentlicht: (2026)
von: Li, Zheng, et al.
Veröffentlicht: (2026)
Chain-of-Thought Tokens are Computer Program Variables
von: Zhu, Fangwei, et al.
Veröffentlicht: (2025)
von: Zhu, Fangwei, et al.
Veröffentlicht: (2025)
GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
von: Li, Rang, et al.
Veröffentlicht: (2025)
von: Li, Rang, et al.
Veröffentlicht: (2025)
Benchmarking Chinese Commonsense Reasoning with a Multi-hop Reasoning Perspective
von: You, Wangjie, et al.
Veröffentlicht: (2025)
von: You, Wangjie, et al.
Veröffentlicht: (2025)
The Box is in the Pen: Evaluating Commonsense Reasoning in Neural Machine Translation
von: He, Jie, et al.
Veröffentlicht: (2025)
von: He, Jie, et al.
Veröffentlicht: (2025)
EventLens: Leveraging Event-Aware Pretraining and Cross-modal Linking Enhances Visual Commonsense Reasoning
von: Ma, Mingjie, et al.
Veröffentlicht: (2024)
von: Ma, Mingjie, et al.
Veröffentlicht: (2024)
AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios
von: Alazraki, Lisa, et al.
Veröffentlicht: (2025)
von: Alazraki, Lisa, et al.
Veröffentlicht: (2025)
How Far are LLMs from Being Our Digital Twins? A Benchmark for Persona-Based Behavior Chain Simulation
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Com$^2$: A Causal-Guided Benchmark for Exploring Complex Commonsense Reasoning in Large Language Models
von: Xiong, Kai, et al.
Veröffentlicht: (2025)
von: Xiong, Kai, et al.
Veröffentlicht: (2025)
Can Large Multimodal Models Uncover Deep Semantics Behind Images?
von: Yang, Yixin, et al.
Veröffentlicht: (2024)
von: Yang, Yixin, et al.
Veröffentlicht: (2024)
Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
von: Ma, Wenhan, et al.
Veröffentlicht: (2025)
von: Ma, Wenhan, et al.
Veröffentlicht: (2025)
SCoTER: Structured Chain-of-Thought Transfer for Enhanced Recommendation
von: Jiang, Jie, et al.
Veröffentlicht: (2025)
von: Jiang, Jie, et al.
Veröffentlicht: (2025)
Robust Disentangled Counterfactual Learning for Physical Audiovisual Commonsense Reasoning
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
von: Sun, Jiaxing, et al.
Veröffentlicht: (2024)
von: Sun, Jiaxing, et al.
Veröffentlicht: (2024)
LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning
von: Junias, Obed, et al.
Veröffentlicht: (2026)
von: Junias, Obed, et al.
Veröffentlicht: (2026)
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2026)
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2026)
ESP: Extro-Spective Prediction for Long-term Behavior Reasoning in Emergency Scenarios
von: Wang, Dingrui, et al.
Veröffentlicht: (2024)
von: Wang, Dingrui, et al.
Veröffentlicht: (2024)
RoboCAS: A Benchmark for Robotic Manipulation in Complex Object Arrangement Scenarios
von: Zheng, Liming, et al.
Veröffentlicht: (2024)
von: Zheng, Liming, et al.
Veröffentlicht: (2024)
Thai Winograd Schemas: A Benchmark for Thai Commonsense Reasoning
von: Artkaew, Phakphum
Veröffentlicht: (2024)
von: Artkaew, Phakphum
Veröffentlicht: (2024)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
von: Palta, Shramay, et al.
Veröffentlicht: (2024)
von: Palta, Shramay, et al.
Veröffentlicht: (2024)
Meta-RTL: Reinforcement-Based Meta-Transfer Learning for Low-Resource Commonsense Reasoning
von: Fu, Yu, et al.
Veröffentlicht: (2024)
von: Fu, Yu, et al.
Veröffentlicht: (2024)
Open STM: A low-cost scanning tunneling microscope with a fast approach method
von: Ma, Weilin
Veröffentlicht: (2023)
von: Ma, Weilin
Veröffentlicht: (2023)
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
von: Yang, Yixin, et al.
Veröffentlicht: (2025)
von: Yang, Yixin, et al.
Veröffentlicht: (2025)
Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective
von: Ma, Qingchuan, et al.
Veröffentlicht: (2025)
von: Ma, Qingchuan, et al.
Veröffentlicht: (2025)
Benchmarking Multi-Step Legal Reasoning and Analyzing Chain-of-Thought Effects in Large Language Models
von: Yu, Wenhan, et al.
Veröffentlicht: (2025)
von: Yu, Wenhan, et al.
Veröffentlicht: (2025)
EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
von: Li, Xinze, et al.
Veröffentlicht: (2026)
von: Li, Xinze, et al.
Veröffentlicht: (2026)
What Really is Commonsense Knowledge?
von: Do, Quyet V., et al.
Veröffentlicht: (2024)
von: Do, Quyet V., et al.
Veröffentlicht: (2024)
Learning-Guided Rolling Horizon Optimization for Long-Horizon Flexible Job-Shop Scheduling
von: Li, Sirui, et al.
Veröffentlicht: (2025)
von: Li, Sirui, et al.
Veröffentlicht: (2025)
The Odyssey of Commonsense Causality: From Foundational Benchmarks to Cutting-Edge Reasoning
von: Cui, Shaobo, et al.
Veröffentlicht: (2024)
von: Cui, Shaobo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SCoRE Sets: A Versatile Framework for Simultaneous Inference
von: Telschow, Fabian J. E., et al.
Veröffentlicht: (2023) -
D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
von: Zhou, Weibo, et al.
Veröffentlicht: (2025) -
SCoRE: Streamlined Corpus-based Relation Extraction using Multi-Label Contrastive Learning and Bayesian kNN
von: Mariotti, Luca, et al.
Veröffentlicht: (2025) -
Enhancing Reliability across Short and Long-Form QA via Reinforcement Learning
von: Wang, Yudong, et al.
Veröffentlicht: (2025) -
CoLT: Reasoning with Chain of Latent Tool Calls
von: Zhu, Fangwei, et al.
Veröffentlicht: (2026)