DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?
Fuente:
arXiv
Saved in:
| Main Authors: | Larionov, Daniil, Takeshita, Sotaro, Zhang, Ran, Chen, Yanran, Leiter, Christoph, Wang, Zhipin, Greisinger, Christian, Eger, Steffen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GerAV: Towards New Heights in German Authorship Verification using Fine-Tuned LLMs on a New Benchmark
by: Kiefer, Lotta, et al.
Published: (2026)
by: Kiefer, Lotta, et al.
Published: (2026)
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation
by: Leiter, Christoph, et al.
Published: (2024)
by: Leiter, Christoph, et al.
Published: (2024)
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
by: Larionov, Daniil, et al.
Published: (2024)
by: Larionov, Daniil, et al.
Published: (2024)
NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers?
by: Leiter, Christoph, et al.
Published: (2024)
by: Leiter, Christoph, et al.
Published: (2024)
TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning
by: Greisinger, Christian, et al.
Published: (2026)
by: Greisinger, Christian, et al.
Published: (2026)
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
by: Larionov, Daniil, et al.
Published: (2025)
by: Larionov, Daniil, et al.
Published: (2025)
o3-mini vs DeepSeek-R1: Which One is Safer?
by: Arrieta, Aitor, et al.
Published: (2025)
by: Arrieta, Aitor, et al.
Published: (2025)
xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation Metrics
by: Larionov, Daniil, et al.
Published: (2024)
by: Larionov, Daniil, et al.
Published: (2024)
BMX: Boosting Natural Language Generation Metrics with Explainability
by: Leiter, Christoph, et al.
Published: (2022)
by: Leiter, Christoph, et al.
Published: (2022)
ValueGround: Evaluating Culture-Conditioned Visual Value Grounding in MLLMs
by: Wang, Zhipin, et al.
Published: (2026)
by: Wang, Zhipin, et al.
Published: (2026)
Do Emotions Really Affect Argument Convincingness? A Dynamic Approach with LLM-based Manipulation Checks
by: Chen, Yanran, et al.
Published: (2025)
by: Chen, Yanran, et al.
Published: (2025)
Medical Reasoning in LLMs: An In-Depth Analysis of DeepSeek R1
by: Moell, Birger, et al.
Published: (2025)
by: Moell, Birger, et al.
Published: (2025)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
by: Zhang, Ran, et al.
Published: (2024)
by: Zhang, Ran, et al.
Published: (2024)
Cross-lingual Cross-temporal Summarization: Dataset, Models, Evaluation
by: Zhang, Ran, et al.
Published: (2023)
by: Zhang, Ran, et al.
Published: (2023)
Learning to Reason: Training LLMs with GPT-OSS or DeepSeek R1 Reasoning Traces
by: Shmidman, Shaltiel, et al.
Published: (2025)
by: Shmidman, Shaltiel, et al.
Published: (2025)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
by: DeepSeek-AI, et al.
Published: (2025)
by: DeepSeek-AI, et al.
Published: (2025)
CROC: Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks
by: Leiter, Christoph, et al.
Published: (2025)
by: Leiter, Christoph, et al.
Published: (2025)
A Comparison of DeepSeek and Other LLMs
by: Gao, Tianchen, et al.
Published: (2025)
by: Gao, Tianchen, et al.
Published: (2025)
Are DeepSeek R1 And Other Reasoning Models More Faithful?
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
Analysis of LLM Bias (Chinese Propaganda & Anti-US Sentiment) in DeepSeek-R1 vs. ChatGPT o3-mini-high
by: Huang, PeiHsuan, et al.
Published: (2025)
by: Huang, PeiHsuan, et al.
Published: (2025)
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
by: Kunilovskaya, Maria, et al.
Published: (2026)
by: Kunilovskaya, Maria, et al.
Published: (2026)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
by: Marjanović, Sara Vera, et al.
Published: (2025)
by: Marjanović, Sara Vera, et al.
Published: (2025)
How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers
by: Menke, Antonio-Gabriel Chacón, et al.
Published: (2025)
by: Menke, Antonio-Gabriel Chacón, et al.
Published: (2025)
Argument Summarization and its Evaluation in the Era of Large Language Models
by: Altemeyer, Moritz, et al.
Published: (2025)
by: Altemeyer, Moritz, et al.
Published: (2025)
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
by: Xu, Pusheng, et al.
Published: (2025)
by: Xu, Pusheng, et al.
Published: (2025)
LLM-based multi-agent poetry generation in non-cooperative environments
by: Zhang, Ran, et al.
Published: (2024)
by: Zhang, Ran, et al.
Published: (2024)
Evaluating Diversity in Automatic Poetry Generation
by: Chen, Yanran, et al.
Published: (2024)
by: Chen, Yanran, et al.
Published: (2024)
Towards Explainable Evaluation Metrics for Machine Translation
by: Leiter, Christoph, et al.
Published: (2023)
by: Leiter, Christoph, et al.
Published: (2023)
Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond
by: Hu, Yinghao, et al.
Published: (2025)
by: Hu, Yinghao, et al.
Published: (2025)
Zhyper: Factorized Hypernetworks for Conditioned LLM Fine-Tuning
by: Abdalla, M. H. I., et al.
Published: (2025)
by: Abdalla, M. H. I., et al.
Published: (2025)
LLLMs: A Data-Driven Survey of Evolving Research on Limitations of Large Language Models
by: Kostikova, Aida, et al.
Published: (2025)
by: Kostikova, Aida, et al.
Published: (2025)
From ChatGPT to DeepSeek: Can LLMs Simulate Humanity?
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
ACLSum: A New Dataset for Aspect-based Summarization of Scientific Publications
by: Takeshita, Sotaro, et al.
Published: (2024)
by: Takeshita, Sotaro, et al.
Published: (2024)
Can LLMs Assist Computer Education? an Empirical Case Study of DeepSeek
by: Xiao, Dongfu, et al.
Published: (2025)
by: Xiao, Dongfu, et al.
Published: (2025)
DeepSeek-OCR: Contexts Optical Compression
by: Wei, Haoran, et al.
Published: (2025)
by: Wei, Haoran, et al.
Published: (2025)
DeepSeek-V3 Technical Report
by: DeepSeek-AI, et al.
Published: (2024)
by: DeepSeek-AI, et al.
Published: (2024)
Are Large Language Models Capable of Deep Relational Reasoning? Insights from DeepSeek-R1 and Benchmark Comparisons
by: So, Chi Chiu, et al.
Published: (2025)
by: So, Chi Chiu, et al.
Published: (2025)
Benchmark-Driven Selection of AI: Evidence from DeepSeek-R1
by: Spelda, Petr, et al.
Published: (2025)
by: Spelda, Petr, et al.
Published: (2025)
Brief analysis of DeepSeek R1 and its implications for Generative AI
by: Mercer, Sarah, et al.
Published: (2025)
by: Mercer, Sarah, et al.
Published: (2025)
Similar Items
-
GerAV: Towards New Heights in German Authorship Verification using Fine-Tuned LLMs on a New Benchmark
by: Kiefer, Lotta, et al.
Published: (2026) -
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation
by: Leiter, Christoph, et al.
Published: (2024) -
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
by: Larionov, Daniil, et al.
Published: (2024) -
NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers?
by: Leiter, Christoph, et al.
Published: (2024) -
TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning
by: Greisinger, Christian, et al.
Published: (2026)