Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zhenglin, Wu, Jialong, LI, Pengfei, Jiang, Yong, Zhou, Deyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SEED: Accelerating Reasoning Tree Construction via Scheduled Speculative Decoding
von: Wang, Zhenglin, et al.
Veröffentlicht: (2024)
von: Wang, Zhenglin, et al.
Veröffentlicht: (2024)
AdaRewriter: Unleashing the Power of Prompting-based Conversational Query Reformulation via Test-Time Adaptation
von: Lai, Yilong, et al.
Veröffentlicht: (2025)
von: Lai, Yilong, et al.
Veröffentlicht: (2025)
SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation
von: Wu, Jialong, et al.
Veröffentlicht: (2024)
von: Wu, Jialong, et al.
Veröffentlicht: (2024)
WebWalker: Benchmarking LLMs in Web Traversal
von: Wu, Jialong, et al.
Veröffentlicht: (2025)
von: Wu, Jialong, et al.
Veröffentlicht: (2025)
AdaCQR: Enhancing Query Reformulation for Conversational Search via Sparse and Dense Retrieval Alignment
von: Lai, Yilong, et al.
Veröffentlicht: (2024)
von: Lai, Yilong, et al.
Veröffentlicht: (2024)
Large Language Models Have Intrinsic Meta-Cognition, but Need a Good Lens
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
When KV Cache Reuse Fails in Multi-Agent Systems: Cross-Candidate Interaction is Crucial for LLM Judges
von: Liang, Sichu, et al.
Veröffentlicht: (2026)
von: Liang, Sichu, et al.
Veröffentlicht: (2026)
PROPER: A Progressive Learning Framework for Personalized Large Language Models with Group-Level Adaptation
von: Zhang, Linhai, et al.
Veröffentlicht: (2025)
von: Zhang, Linhai, et al.
Veröffentlicht: (2025)
DINER: Debiasing Aspect-based Sentiment Analysis with Multi-variable Causal Inference
von: Wu, Jialong, et al.
Veröffentlicht: (2024)
von: Wu, Jialong, et al.
Veröffentlicht: (2024)
STAR: Constraint LoRA with Dynamic Active Learning for Data-Efficient Fine-Tuning of Large Language Models
von: Zhang, Linhai, et al.
Veröffentlicht: (2024)
von: Zhang, Linhai, et al.
Veröffentlicht: (2024)
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
von: Sun, Jiaxing, et al.
Veröffentlicht: (2024)
von: Sun, Jiaxing, et al.
Veröffentlicht: (2024)
Causal Prompting: Debiasing Large Language Model Prompting based on Front-Door Adjustment
von: Zhang, Congzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Congzhi, et al.
Veröffentlicht: (2024)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025)
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025)
Flames: Benchmarking Value Alignment of LLMs in Chinese
von: Huang, Kexin, et al.
Veröffentlicht: (2023)
von: Huang, Kexin, et al.
Veröffentlicht: (2023)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
von: He, Zheqi, et al.
Veröffentlicht: (2024)
von: He, Zheqi, et al.
Veröffentlicht: (2024)
Can Pruning Improve Reasoning? Revisiting Long-CoT Compression with Capability in Mind for Better Reasoning
von: Zhao, Shangziqi, et al.
Veröffentlicht: (2025)
von: Zhao, Shangziqi, et al.
Veröffentlicht: (2025)
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
TRAM: Benchmarking Temporal Reasoning for Large Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
GUARDIAN: Safeguarding LLM Multi-Agent Collaborations with Temporal Graph Modeling
von: Zhou, Jialong, et al.
Veröffentlicht: (2025)
von: Zhou, Jialong, et al.
Veröffentlicht: (2025)
Benchmarking Chinese Commonsense Reasoning with a Multi-hop Reasoning Perspective
von: You, Wangjie, et al.
Veröffentlicht: (2025)
von: You, Wangjie, et al.
Veröffentlicht: (2025)
CRMWeaver: Building Powerful Business Agent via Agentic RL and Shared Memories
von: Lai, Yilong, et al.
Veröffentlicht: (2025)
von: Lai, Yilong, et al.
Veröffentlicht: (2025)
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
von: Yu, Haiyang, et al.
Veröffentlicht: (2025)
von: Yu, Haiyang, et al.
Veröffentlicht: (2025)
PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and Engineering
von: Wang, Xiangfeng, et al.
Veröffentlicht: (2026)
von: Wang, Xiangfeng, et al.
Veröffentlicht: (2026)
Rhyme-aware Chinese lyric generator based on GPT
von: Yuan, Yixiao, et al.
Veröffentlicht: (2024)
von: Yuan, Yixiao, et al.
Veröffentlicht: (2024)
AlignBench: Benchmarking Chinese Alignment of Large Language Models
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks
von: Liu, Junlin, et al.
Veröffentlicht: (2026)
von: Liu, Junlin, et al.
Veröffentlicht: (2026)
RGAlign-Rec: Ranking-Guided Alignment for Latent Query Reasoning in Recommendation Systems
von: Liu, Junhua, et al.
Veröffentlicht: (2026)
von: Liu, Junhua, et al.
Veröffentlicht: (2026)
Large, Small or Both: A Novel Data Augmentation Framework Based on Language Models for Debiasing Opinion Summarization
von: Zhang, Yanyue, et al.
Veröffentlicht: (2024)
von: Zhang, Yanyue, et al.
Veröffentlicht: (2024)
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
von: Fatemi, Bahare, et al.
Veröffentlicht: (2024)
von: Fatemi, Bahare, et al.
Veröffentlicht: (2024)
General-Reasoner: Advancing LLM Reasoning Across All Domains
von: Ma, Xueguang, et al.
Veröffentlicht: (2025)
von: Ma, Xueguang, et al.
Veröffentlicht: (2025)
Let LLMs Take on the Latest Challenges! A Chinese Dynamic Question Answering Benchmark
von: Xu, Zhikun, et al.
Veröffentlicht: (2024)
von: Xu, Zhikun, et al.
Veröffentlicht: (2024)
RealTalk-CN: A Realistic Chinese Speech-Text Dialogue Benchmark With Cross-Modal Interaction Analysis
von: Wang, Enzhi, et al.
Veröffentlicht: (2025)
von: Wang, Enzhi, et al.
Veröffentlicht: (2025)
AnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs
von: Feng, Xiang, et al.
Veröffentlicht: (2025)
von: Feng, Xiang, et al.
Veröffentlicht: (2025)
Multi-Physics: A Comprehensive Benchmark for Multimodal LLMs Reasoning on Chinese Multi-Subject Physics Problems
von: Luo, Zhongze, et al.
Veröffentlicht: (2025)
von: Luo, Zhongze, et al.
Veröffentlicht: (2025)
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
TCMBench: A Comprehensive Benchmark for Evaluating Large Language Models in Traditional Chinese Medicine
von: Yue, Wenjing, et al.
Veröffentlicht: (2024)
von: Yue, Wenjing, et al.
Veröffentlicht: (2024)
CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models
von: Liu, Wentao, et al.
Veröffentlicht: (2024)
von: Liu, Wentao, et al.
Veröffentlicht: (2024)
OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
von: Zhang, Min, et al.
Veröffentlicht: (2025)
von: Zhang, Min, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SEED: Accelerating Reasoning Tree Construction via Scheduled Speculative Decoding
von: Wang, Zhenglin, et al.
Veröffentlicht: (2024) -
AdaRewriter: Unleashing the Power of Prompting-based Conversational Query Reformulation via Test-Time Adaptation
von: Lai, Yilong, et al.
Veröffentlicht: (2025) -
SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation
von: Wu, Jialong, et al.
Veröffentlicht: (2024) -
WebWalker: Benchmarking LLMs in Web Traversal
von: Wu, Jialong, et al.
Veröffentlicht: (2025) -
AdaCQR: Enhancing Query Reformulation for Conversational Search via Sparse and Dense Retrieval Alignment
von: Lai, Yilong, et al.
Veröffentlicht: (2024)