Evaluating OpenAI GPT Models for Translation of Endangered Uralic Languages: A Comparison of Reasoning and Non-Reasoning Architectures
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tereshchenko, Yehor, Hämäläinen, Mika, Myroniuk, Svitlana |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
DAG: Dictionary-Augmented Generation for Disambiguation of Sentences in Endangered Uralic Languages using ChatGPT
par: Hämäläinen, Mika
Publié: (2024)
par: Hämäläinen, Mika
Publié: (2024)
A Comparative Analysis of Ethical and Safety Gaps in LLMs using Relative Danger Coefficient
par: Tereshchenko, Yehor, et autres
Publié: (2025)
par: Tereshchenko, Yehor, et autres
Publié: (2025)
Efficient Toxicity Detection in Gaming Chats: A Comparative Study of Embeddings, Fine-Tuned Transformers and LLMs
par: Tereshchenko, Yehor, et autres
Publié: (2025)
par: Tereshchenko, Yehor, et autres
Publié: (2025)
ORACLE: Time-Dependent Recursive Summary Graphs for Foresight on News Data Using LLMs
par: Kharlashkin, Lev, et autres
Publié: (2025)
par: Kharlashkin, Lev, et autres
Publié: (2025)
On Psychology of AI -- Does Primacy Effect Affect ChatGPT and Other LLMs?
par: Hämäläinen, Mika
Publié: (2025)
par: Hämäläinen, Mika
Publié: (2025)
Leveraging Transformer-Based Models for Predicting Inflection Classes of Words in an Endangered Sami Language
par: Alnajjar, Khalid, et autres
Publié: (2024)
par: Alnajjar, Khalid, et autres
Publié: (2024)
From NLG Evaluation to Modern Student Assessment in the Era of ChatGPT: The Great Misalignment Problem and Pedagogical Multi-Factor Assessment (P-MFA)
par: Hämäläinen, Mika, et autres
Publié: (2025)
par: Hämäläinen, Mika, et autres
Publié: (2025)
Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
par: Bi, Ziqian, et autres
Publié: (2025)
par: Bi, Ziqian, et autres
Publié: (2025)
A Comparative Study on Reasoning Patterns of OpenAI's o1 Model
par: Wu, Siwei, et autres
Publié: (2024)
par: Wu, Siwei, et autres
Publié: (2024)
On Sarcasm Detection with OpenAI GPT-based Models
par: Gole, Montgomery, et autres
Publié: (2023)
par: Gole, Montgomery, et autres
Publié: (2023)
Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT
par: Shakil, Hassan, et autres
Publié: (2024)
par: Shakil, Hassan, et autres
Publié: (2024)
OpenAI GPT-5 System Card
par: Singh, Aaditya, et autres
Publié: (2025)
par: Singh, Aaditya, et autres
Publié: (2025)
A Case Study of Web App Coding with OpenAI Reasoning Models
par: Cui, Yi
Publié: (2024)
par: Cui, Yi
Publié: (2024)
OpenAI's GPT-OSS-20B Model and Safety Alignment Issues in a Low-Resource Language
par: Inuwa-Dutse, Isa
Publié: (2025)
par: Inuwa-Dutse, Isa
Publié: (2025)
Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond
par: Hu, Yinghao, et autres
Publié: (2025)
par: Hu, Yinghao, et autres
Publié: (2025)
Tokenization and Morphological Fidelity in Uralic NLP: A Cross-Lingual Evaluation
par: Xu, Nuo, et autres
Publié: (2026)
par: Xu, Nuo, et autres
Publié: (2026)
Evaluation of OpenAI o1: Opportunities and Challenges of AGI
par: Zhong, Tianyang, et autres
Publié: (2024)
par: Zhong, Tianyang, et autres
Publié: (2024)
Threefold model for AI Readiness: A Case Study with Finnish Healthcare SMEs
par: Alnajjar, Mohammed, et autres
Publié: (2025)
par: Alnajjar, Mohammed, et autres
Publié: (2025)
Can OpenAI o1 Reason Well in Ophthalmology? A 6,990-Question Head-to-Head Evaluation Study
par: Srinivasan, Sahana, et autres
Publié: (2025)
par: Srinivasan, Sahana, et autres
Publié: (2025)
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
par: Iqbal, Umar, et autres
Publié: (2023)
par: Iqbal, Umar, et autres
Publié: (2023)
Does fine-tuning GPT-3 with the OpenAI API leak personally-identifiable information?
par: Sun, Albert Yu, et autres
Publié: (2023)
par: Sun, Albert Yu, et autres
Publié: (2023)
Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
par: Amjad, Husnain, et autres
Publié: (2026)
par: Amjad, Husnain, et autres
Publié: (2026)
Analyzing Pokémon and Mario Streamers' Twitch Chat with LLM-based User Embeddings
par: Hämäläinen, Mika, et autres
Publié: (2024)
par: Hämäläinen, Mika, et autres
Publié: (2024)
MTUncertainty: Assessing the Need for Post-editing of Machine Translation Outputs by Fine-tuning OpenAI LLMs
par: Gladkoff, Serge, et autres
Publié: (2023)
par: Gladkoff, Serge, et autres
Publié: (2023)
Leveraging Virtual Reality and AI Tutoring for Language Learning: A Case Study of a Virtual Campus Environment with OpenAI GPT Integration with Unity 3D
par: TG, Adithya, et autres
Publié: (2024)
par: TG, Adithya, et autres
Publié: (2024)
Unified Deployment-Aware Evaluation of Open Reasoning Language Models
par: Manik, Md Motaleb Hossen, et autres
Publié: (2026)
par: Manik, Md Motaleb Hossen, et autres
Publié: (2026)
Predicting Sustainable Development Goals Using Course Descriptions -- from LLMs to Conventional Foundation Models
par: Kharlashkin, Lev, et autres
Publié: (2024)
par: Kharlashkin, Lev, et autres
Publié: (2024)
OpenThaiGPT 1.6 and R1: Thai-Centric Open Source and Reasoning Large Language Models
par: Yuenyong, Sumeth, et autres
Publié: (2025)
par: Yuenyong, Sumeth, et autres
Publié: (2025)
H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking
par: Kuo, Martin, et autres
Publié: (2025)
par: Kuo, Martin, et autres
Publié: (2025)
OpenAI Cribbed Our Tax Example, But Can GPT-4 Really Do Tax?
par: Blair-Stanek, Andrew, et autres
Publié: (2023)
par: Blair-Stanek, Andrew, et autres
Publié: (2023)
A Benchmark for End-to-End Zero-Shot Biomedical Relation Extraction with LLMs: Experiments with OpenAI Models
par: Brokman, Aviv, et autres
Publié: (2025)
par: Brokman, Aviv, et autres
Publié: (2025)
DermoGPT: Open Weights and Open Data for Morphology-Grounded Dermatological Reasoning MLLMs
par: Ru, Jinghan, et autres
Publié: (2026)
par: Ru, Jinghan, et autres
Publié: (2026)
Evaluation of ChatGPT Family of Models for Biomedical Reasoning and Classification
par: Chen, Shan, et autres
Publié: (2023)
par: Chen, Shan, et autres
Publié: (2023)
In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b
par: Durner, Nils
Publié: (2025)
par: Durner, Nils
Publié: (2025)
LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
par: Valmeekam, Karthik, et autres
Publié: (2024)
par: Valmeekam, Karthik, et autres
Publié: (2024)
Comparative Analysis of OpenAI GPT-4o and DeepSeek R1 for Scientific Text Categorization Using Prompt Engineering
par: Maiti, Aniruddha, et autres
Publié: (2025)
par: Maiti, Aniruddha, et autres
Publié: (2025)
Subword-Based Comparative Linguistics across 242 Languages Using Wikipedia Glottosets
par: Chelombitko, Iaroslav, et autres
Publié: (2026)
par: Chelombitko, Iaroslav, et autres
Publié: (2026)
Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations
par: Zahraei, Pardis Sadat, et autres
Publié: (2025)
par: Zahraei, Pardis Sadat, et autres
Publié: (2025)
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
par: Xu, Pusheng, et autres
Publié: (2025)
par: Xu, Pusheng, et autres
Publié: (2025)
Improving Symbolic Translation of Language Models for Logical Reasoning
par: Thatikonda, Ramya Keerthy, et autres
Publié: (2026)
par: Thatikonda, Ramya Keerthy, et autres
Publié: (2026)
Documents similaires
-
DAG: Dictionary-Augmented Generation for Disambiguation of Sentences in Endangered Uralic Languages using ChatGPT
par: Hämäläinen, Mika
Publié: (2024) -
A Comparative Analysis of Ethical and Safety Gaps in LLMs using Relative Danger Coefficient
par: Tereshchenko, Yehor, et autres
Publié: (2025) -
Efficient Toxicity Detection in Gaming Chats: A Comparative Study of Embeddings, Fine-Tuned Transformers and LLMs
par: Tereshchenko, Yehor, et autres
Publié: (2025) -
ORACLE: Time-Dependent Recursive Summary Graphs for Foresight on News Data Using LLMs
par: Kharlashkin, Lev, et autres
Publié: (2025) -
On Psychology of AI -- Does Primacy Effect Affect ChatGPT and Other LLMs?
par: Hämäläinen, Mika
Publié: (2025)