Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Bayram, M. Ali, Fincan, Ali Arda, Gümüş, Ahmet Semih, Diri, Banu, Yıldırım, Savaş, Aytaş, Öner |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Büyük Dil Modelleri için TR-MMLU Benchmarkı: Performans Değerlendirmesi, Zorluklar ve İyileştirme Fırsatları
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
Doğal Dil İşlemede Tokenizasyon Standartları ve Ölçümü: Türkçe Üzerinden Büyük Dil Modellerinin Karşılaştırmalı Analizi
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
Tokens with Meaning: A Hybrid Tokenization Approach for Turkish
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
di: Bayram, M. Ali, et al.
Pubblicazione: (2025)
Experimentation in Content Moderation using RWKV
di: Yildirim, Umut, et al.
Pubblicazione: (2024)
di: Yildirim, Umut, et al.
Pubblicazione: (2024)
Evaluating Pixel Language Models on Non-Standardized Languages
di: Muñoz-Ortiz, Alberto, et al.
Pubblicazione: (2024)
di: Muñoz-Ortiz, Alberto, et al.
Pubblicazione: (2024)
An NLP-Driven Framework for Curriculum-Labor Market Alignment: Schema-Constrained LLM Extraction, ESCO-Anchored Semantic Matching, and Multi-Dimensional Gap Quantification
di: Turaev, Sherzod, et al.
Pubblicazione: (2026)
di: Turaev, Sherzod, et al.
Pubblicazione: (2026)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
di: Liu, Aiwei, et al.
Pubblicazione: (2025)
di: Liu, Aiwei, et al.
Pubblicazione: (2025)
Context Aware Lemmatization and Morphological Tagging Method in Turkish
di: Sayallar, Cagri
Pubblicazione: (2025)
di: Sayallar, Cagri
Pubblicazione: (2025)
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
di: Sun, Jingyi, et al.
Pubblicazione: (2024)
di: Sun, Jingyi, et al.
Pubblicazione: (2024)
FairLangProc: A Python package for fairness in NLP
di: Pérez-Peralta, Arturo, et al.
Pubblicazione: (2025)
di: Pérez-Peralta, Arturo, et al.
Pubblicazione: (2025)
Pitfalls in Evaluating Interpretability Agents
di: Haklay, Tal, et al.
Pubblicazione: (2026)
di: Haklay, Tal, et al.
Pubblicazione: (2026)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
di: Tu, Songjun, et al.
Pubblicazione: (2026)
di: Tu, Songjun, et al.
Pubblicazione: (2026)
The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
di: Gómez-Rodríguez, Carlos, et al.
Pubblicazione: (2024)
di: Gómez-Rodríguez, Carlos, et al.
Pubblicazione: (2024)
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
di: Pan, Leyi, et al.
Pubblicazione: (2025)
di: Pan, Leyi, et al.
Pubblicazione: (2025)
The Paradox of Poetic Intent in Back-Translation: Evaluating the Quality of Large Language Models in Chinese Translation
di: Weigang, Li, et al.
Pubblicazione: (2025)
di: Weigang, Li, et al.
Pubblicazione: (2025)
NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
di: Yao, Ben, et al.
Pubblicazione: (2025)
di: Yao, Ben, et al.
Pubblicazione: (2025)
Efficient Adaptive Rejection Sampling for Accelerating Speculative Decoding in Large Language Models
di: Sun, Chendong, et al.
Pubblicazione: (2025)
di: Sun, Chendong, et al.
Pubblicazione: (2025)
Subjective Question Generation and Answer Evaluation using NLP
di: Islam, G. M. Refatul, et al.
Pubblicazione: (2025)
di: Islam, G. M. Refatul, et al.
Pubblicazione: (2025)
Emotional Sequential Influence Modeling on False Information
di: Naskar, Debashis, et al.
Pubblicazione: (2024)
di: Naskar, Debashis, et al.
Pubblicazione: (2024)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
di: Goldin, Gili, et al.
Pubblicazione: (2024)
di: Goldin, Gili, et al.
Pubblicazione: (2024)
Math Natural Language Inference: this should be easy!
di: de Paiva, Valeria, et al.
Pubblicazione: (2025)
di: de Paiva, Valeria, et al.
Pubblicazione: (2025)
Fast Quiet-STaR: Thinking Without Thought Tokens
di: Huang, Wei, et al.
Pubblicazione: (2025)
di: Huang, Wei, et al.
Pubblicazione: (2025)
New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
di: Wang, Zhilin, et al.
Pubblicazione: (2026)
di: Wang, Zhilin, et al.
Pubblicazione: (2026)
d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
di: Pan, Leyi, et al.
Pubblicazione: (2025)
di: Pan, Leyi, et al.
Pubblicazione: (2025)
Towards Effective and Efficient Continual Pre-training of Large Language Models
di: Chen, Jie, et al.
Pubblicazione: (2024)
di: Chen, Jie, et al.
Pubblicazione: (2024)
An Unforgeable Publicly Verifiable Watermark for Large Language Models
di: Liu, Aiwei, et al.
Pubblicazione: (2023)
di: Liu, Aiwei, et al.
Pubblicazione: (2023)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
SentiCSE: A Sentiment-aware Contrastive Sentence Embedding Framework with Sentiment-guided Textual Similarity
di: Kim, Jaemin, et al.
Pubblicazione: (2024)
di: Kim, Jaemin, et al.
Pubblicazione: (2024)
Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
di: Zhang, Zhehao, et al.
Pubblicazione: (2025)
di: Zhang, Zhehao, et al.
Pubblicazione: (2025)
Exploiting Pre-trained Encoder-Decoder Transformers for Sequence-to-Sequence Constituent Parsing
di: Fernández-González, Daniel, et al.
Pubblicazione: (2026)
di: Fernández-González, Daniel, et al.
Pubblicazione: (2026)
Parametric Social Identity Injection and Diversification in Public Opinion Simulation
di: Wang, Hexi, et al.
Pubblicazione: (2026)
di: Wang, Hexi, et al.
Pubblicazione: (2026)
Trusted Uncertainty in Large Language Models: A Unified Framework for Confidence Calibration and Risk-Controlled Refusal
di: Oehri, Markus, et al.
Pubblicazione: (2025)
di: Oehri, Markus, et al.
Pubblicazione: (2025)
Beyond Cosine Similarity
di: Ai, Xinbo
Pubblicazione: (2026)
di: Ai, Xinbo
Pubblicazione: (2026)
AI-assisted German Employment Contract Review: A Benchmark Dataset
di: Wardas, Oliver, et al.
Pubblicazione: (2025)
di: Wardas, Oliver, et al.
Pubblicazione: (2025)
The Superalignment of Superhuman Intelligence with Large Language Models
di: Huang, Minlie, et al.
Pubblicazione: (2024)
di: Huang, Minlie, et al.
Pubblicazione: (2024)
ScoreRAG: A Retrieval-Augmented Generation Framework with Consistency-Relevance Scoring and Structured Summarization for News Generation
di: Lin, Pei-Yun, et al.
Pubblicazione: (2025)
di: Lin, Pei-Yun, et al.
Pubblicazione: (2025)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
di: Park, Seungcheol, et al.
Pubblicazione: (2025)
di: Park, Seungcheol, et al.
Pubblicazione: (2025)
When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
di: Wang, Yongjie, et al.
Pubblicazione: (2025)
di: Wang, Yongjie, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Büyük Dil Modelleri için TR-MMLU Benchmarkı: Performans Değerlendirmesi, Zorluklar ve İyileştirme Fırsatları
di: Bayram, M. Ali, et al.
Pubblicazione: (2025) -
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
di: Bayram, M. Ali, et al.
Pubblicazione: (2025) -
Doğal Dil İşlemede Tokenizasyon Standartları ve Ölçümü: Türkçe Üzerinden Büyük Dil Modellerinin Karşılaştırmalı Analizi
di: Bayram, M. Ali, et al.
Pubblicazione: (2025) -
Tokens with Meaning: A Hybrid Tokenization Approach for Turkish
di: Bayram, M. Ali, et al.
Pubblicazione: (2025) -
Experimentation in Content Moderation using RWKV
di: Yildirim, Umut, et al.
Pubblicazione: (2024)