The Syntactic Acceptability Dataset (Preview): A Resource for Machine Learning and Linguistic Analysis of English
Fuente:
arXiv
Saved in:
| Main Author: | Juzek, Tom S |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
by: Bayram, M. Ali, et al.
Published: (2025)
by: Bayram, M. Ali, et al.
Published: (2025)
Tokens with Meaning: A Hybrid Tokenization Approach for Turkish
by: Bayram, M. Ali, et al.
Published: (2025)
by: Bayram, M. Ali, et al.
Published: (2025)
Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025)
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025)
Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
by: Juzek, Tom S., et al.
Published: (2025)
by: Juzek, Tom S., et al.
Published: (2025)
Morphological Synthesizer for Ge'ez Language: Addressing Morphological Complexity and Resource Limitations
by: Gebremariam, Gebrearegawi, et al.
Published: (2025)
by: Gebremariam, Gebrearegawi, et al.
Published: (2025)
IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text
by: Pall, Rajveer Singh
Published: (2026)
by: Pall, Rajveer Singh
Published: (2026)
Improving Large-Scale k-Nearest Neighbor Text Categorization with Label Autoencoders
by: Ribadas-Pena, Francisco J., et al.
Published: (2024)
by: Ribadas-Pena, Francisco J., et al.
Published: (2024)
Optimizing Retrieval-Augmented Generation (RAG) for Colloquial Cantonese: A LoRA-Based Systematic Review
by: Calonge, David Santandreu, et al.
Published: (2025)
by: Calonge, David Santandreu, et al.
Published: (2025)
LLM-supported document separation for printed reviews from zbMATH Open
by: Pluzhnikov, Ivan, et al.
Published: (2026)
by: Pluzhnikov, Ivan, et al.
Published: (2026)
LLMs as Architects and Critics for Multi-Source Opinion Summarization
by: Attri, Anuj, et al.
Published: (2025)
by: Attri, Anuj, et al.
Published: (2025)
Why We Feel What We Feel: Joint Detection of Emotions and Their Opinion Triggers in E-commerce
by: Attri, Arnav, et al.
Published: (2025)
by: Attri, Arnav, et al.
Published: (2025)
ORPHEAS: A Cross-Lingual Greek-English Embedding Model for Retrieval-Augmented Generation
by: Livieris, Ioannis E., et al.
Published: (2026)
by: Livieris, Ioannis E., et al.
Published: (2026)
From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring
by: Nguyen, Minh Hoang, et al.
Published: (2026)
by: Nguyen, Minh Hoang, et al.
Published: (2026)
Language processing in humans and computers
by: Pavlovic, Dusko
Published: (2024)
by: Pavlovic, Dusko
Published: (2024)
How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
by: Nieth, Björn, et al.
Published: (2026)
by: Nieth, Björn, et al.
Published: (2026)
Argument Quality Annotation and Gender Bias Detection in Financial Communication through Large Language Models
by: Alhamzeh, Alaa, et al.
Published: (2025)
by: Alhamzeh, Alaa, et al.
Published: (2025)
Exploring the Structure of AI-Induced Language Change in Scientific English
by: Galpin, Riley, et al.
Published: (2025)
by: Galpin, Riley, et al.
Published: (2025)
EdgeJury: Cross-Reviewed Small-Model Ensembles for Truthful Question Answering on Serverless Edge Inference
by: Kumar, Aayush
Published: (2025)
by: Kumar, Aayush
Published: (2025)
Knowledge Distillation for Low-Resource Open-source Text-to-SQL Model
by: Qiu, Tianhao, et al.
Published: (2026)
by: Qiu, Tianhao, et al.
Published: (2026)
Mubeen AI: A Specialized Arabic Language Model for Heritage Preservation and User Intent Understanding
by: Aljafari, Mohammed, et al.
Published: (2025)
by: Aljafari, Mohammed, et al.
Published: (2025)
Combating data scarcity in recommendation services: Integrating cognitive types of VARK and neural network technologies (LLM)
by: Zmanovskii, Nikita
Published: (2026)
by: Zmanovskii, Nikita
Published: (2026)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
by: Breneur, Oleksandr Marchenko, et al.
Published: (2026)
by: Breneur, Oleksandr Marchenko, et al.
Published: (2026)
Heterogeneous LLM Methods for Ontology Learning (Few-Shot Prompting, Ensemble Typing, and Attention-Based Taxonomies)
by: Beliaeva, Aleksandra, et al.
Published: (2025)
by: Beliaeva, Aleksandra, et al.
Published: (2025)
IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following
by: Sun, Mingrui, et al.
Published: (2026)
by: Sun, Mingrui, et al.
Published: (2026)
Fine-tuning of Large Language Models for Constituency Parsing Using a Sequence to Sequence Approach
by: Delgado, Francisco Jose Cortes, et al.
Published: (2025)
by: Delgado, Francisco Jose Cortes, et al.
Published: (2025)
Doğal Dil İşlemede Tokenizasyon Standartları ve Ölçümü: Türkçe Üzerinden Büyük Dil Modellerinin Karşılaştırmalı Analizi
by: Bayram, M. Ali, et al.
Published: (2025)
by: Bayram, M. Ali, et al.
Published: (2025)
Büyük Dil Modelleri için TR-MMLU Benchmarkı: Performans Değerlendirmesi, Zorluklar ve İyileştirme Fırsatları
by: Bayram, M. Ali, et al.
Published: (2025)
by: Bayram, M. Ali, et al.
Published: (2025)
Targeted Lexical Injection: Unlocking Latent Cross-Lingual Alignment in Lugha-Llama via Early-Layer LoRA Fine-Tuning
by: Ngugi, Stanley
Published: (2025)
by: Ngugi, Stanley
Published: (2025)
Beyond Subtokens: A Rich Character Embedding for Low-resource and Morphologically Complex Languages
by: Schneider, Felix, et al.
Published: (2026)
by: Schneider, Felix, et al.
Published: (2026)
Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
by: Anderson, Bryce, et al.
Published: (2025)
by: Anderson, Bryce, et al.
Published: (2025)
How much do LLMs learn from negative examples?
by: Hamdan, Shadi, et al.
Published: (2025)
by: Hamdan, Shadi, et al.
Published: (2025)
Rethinking the Multilingual Reasoning Gap with Layer Swap
by: Lasbordes, Maxence, et al.
Published: (2026)
by: Lasbordes, Maxence, et al.
Published: (2026)
Progressive Training for Explainable Citation-Grounded Dialogue: Reducing Hallucination to Zero in English-Hindi LLMs
by: Pandya, Vedant
Published: (2026)
by: Pandya, Vedant
Published: (2026)
Fine-Tuning LLMs on Small Medical Datasets: Text Classification and Normalization Effectiveness on Cardiology reports and Discharge records
by: Losch, Noah, et al.
Published: (2025)
by: Losch, Noah, et al.
Published: (2025)
Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
by: Gupta, Aayush
Published: (2025)
by: Gupta, Aayush
Published: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
A Scalable and High Availability Solution for Recommending Resolutions to Problem Tickets
by: Saragadam, Harish, et al.
Published: (2025)
by: Saragadam, Harish, et al.
Published: (2025)
SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers
by: Chelombitko, Iaroslav, et al.
Published: (2026)
by: Chelombitko, Iaroslav, et al.
Published: (2026)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
by: Mathew, Aby Mammen
Published: (2026)
by: Mathew, Aby Mammen
Published: (2026)
Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
by: Sharma, Shubham, et al.
Published: (2025)
by: Sharma, Shubham, et al.
Published: (2025)
Similar Items
-
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
by: Bayram, M. Ali, et al.
Published: (2025) -
Tokens with Meaning: A Hybrid Tokenization Approach for Turkish
by: Bayram, M. Ali, et al.
Published: (2025) -
Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025) -
Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
by: Juzek, Tom S., et al.
Published: (2025) -
Morphological Synthesizer for Ge'ez Language: Addressing Morphological Complexity and Resource Limitations
by: Gebremariam, Gebrearegawi, et al.
Published: (2025)