100,000+ Movie Reviews from Kazakhstan: Russian, Kazakh, and Code-Switched Texts
Fuente:
arXiv
Guardado en:
| Autor principal: | Yeshpanov, Rustem |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
KazSAnDRA: Kazakh Sentiment Analysis Dataset of Reviews and Attitudes
por: Yeshpanov, Rustem, et al.
Publicado: (2024)
por: Yeshpanov, Rustem, et al.
Publicado: (2024)
KazParC: Kazakh Parallel Corpus for Machine Translation
por: Yeshpanov, Rustem, et al.
Publicado: (2024)
por: Yeshpanov, Rustem, et al.
Publicado: (2024)
KazQAD: Kazakh Open-Domain Question Answering Dataset
por: Yeshpanov, Rustem, et al.
Publicado: (2024)
por: Yeshpanov, Rustem, et al.
Publicado: (2024)
Using Songs to Improve Kazakh Automatic Speech Recognition
por: Yeshpanov, Rustem
Publicado: (2026)
por: Yeshpanov, Rustem
Publicado: (2026)
KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan
por: Togmanov, Mukhammed, et al.
Publicado: (2025)
por: Togmanov, Mukhammed, et al.
Publicado: (2025)
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
por: Abilbekov, Adal, et al.
Publicado: (2024)
por: Abilbekov, Adal, et al.
Publicado: (2024)
Low-resource Machine Translation for Code-switched Kazakh-Russian Language Pair
por: Borisov, Maksim, et al.
Publicado: (2025)
por: Borisov, Maksim, et al.
Publicado: (2025)
Qorgau: Evaluating LLM Safety in Kazakh-Russian Bilingual Contexts
por: Goloburda, Maiya, et al.
Publicado: (2025)
por: Goloburda, Maiya, et al.
Publicado: (2025)
NeuralDB: Scaling Knowledge Editing in LLMs to 100,000 Facts with Neural KV Database
por: Fei, Weizhi, et al.
Publicado: (2025)
por: Fei, Weizhi, et al.
Publicado: (2025)
Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switching
por: Kim, Seoyeon, et al.
Publicado: (2024)
por: Kim, Seoyeon, et al.
Publicado: (2024)
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation
por: Vasilev, Viacheslav, et al.
Publicado: (2025)
por: Vasilev, Viacheslav, et al.
Publicado: (2025)
Sentiment Analysis of Movie Reviews Using BERT
por: Nkhata, Gibson, et al.
Publicado: (2025)
por: Nkhata, Gibson, et al.
Publicado: (2025)
Conditioning LLMs to Generate Code-Switched Text
por: Heredia, Maite, et al.
Publicado: (2025)
por: Heredia, Maite, et al.
Publicado: (2025)
Lost in the Mix: Evaluating LLM Understanding of Code-Switched Text
por: Mohamed, Amr, et al.
Publicado: (2025)
por: Mohamed, Amr, et al.
Publicado: (2025)
LLM-based Code-Switched Text Generation for Grammatical Error Correction
por: Potter, Tom, et al.
Publicado: (2024)
por: Potter, Tom, et al.
Publicado: (2024)
Code-Mixed Probes Show How Pre-Trained Models Generalise On Code-Switched Text
por: De Leon, Frances A. Laureano, et al.
Publicado: (2024)
por: De Leon, Frances A. Laureano, et al.
Publicado: (2024)
KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR
por: Gagnier, Henry, et al.
Publicado: (2026)
por: Gagnier, Henry, et al.
Publicado: (2026)
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
por: Nguyen, Tuan, et al.
Publicado: (2025)
por: Nguyen, Tuan, et al.
Publicado: (2025)
Training LLMs with Fault Tolerant HSDP on 100,000 GPUs
por: Salpekar, Omkar, et al.
Publicado: (2026)
por: Salpekar, Omkar, et al.
Publicado: (2026)
Detecting Propaganda Techniques in Code-Switched Social Media Text
por: Salman, Muhammad Umar, et al.
Publicado: (2023)
por: Salman, Muhammad Umar, et al.
Publicado: (2023)
Chapter 1, 1000, 100.000. Quanti e quali attori nei costrutti personali indeterminati?
por: Fici, Francesca, et al.
Publicado: (2022)
por: Fici, Francesca, et al.
Publicado: (2022)
MovieCORE: COgnitive REasoning in Movies
por: Faure, Gueter Josmy, et al.
Publicado: (2025)
por: Faure, Gueter Josmy, et al.
Publicado: (2025)
MovieSum: An Abstractive Summarization Dataset for Movie Screenplays
por: Saxena, Rohit, et al.
Publicado: (2024)
por: Saxena, Rohit, et al.
Publicado: (2024)
Detecting Spelling and Grammatical Anomalies in Russian Poetry Texts
por: Koziev, Ilya
Publicado: (2025)
por: Koziev, Ilya
Publicado: (2025)
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
por: Li, Jinpeng, et al.
Publicado: (2024)
por: Li, Jinpeng, et al.
Publicado: (2024)
SkyScript-100M: 1,000,000,000 Pairs of Scripts and Shooting Scripts for Short Drama
por: Tang, Jing, et al.
Publicado: (2024)
por: Tang, Jing, et al.
Publicado: (2024)
Movie101v2: Improved Movie Narration Benchmark
por: Yue, Zihao, et al.
Publicado: (2024)
por: Yue, Zihao, et al.
Publicado: (2024)
Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%
por: Zhu, Lei, et al.
Publicado: (2024)
por: Zhu, Lei, et al.
Publicado: (2024)
CoVoSwitch: Machine Translation of Synthetic Code-Switched Text Based on Intonation Units
por: Kang, Yeeun
Publicado: (2024)
por: Kang, Yeeun
Publicado: (2024)
Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?
por: Winata, Genta Indra, et al.
Publicado: (2026)
por: Winata, Genta Indra, et al.
Publicado: (2026)
Linguistics Theory Meets LLM: Code-Switched Text Generation via Equivalence Constrained Large Language Models
por: Kuwanto, Garry, et al.
Publicado: (2024)
por: Kuwanto, Garry, et al.
Publicado: (2024)
Algorithms For Automatic Accentuation And Transcription Of Russian Texts In Speech Recognition Systems
por: Iakovenko, Olga, et al.
Publicado: (2024)
por: Iakovenko, Olga, et al.
Publicado: (2024)
SozKZ: Training Efficient Small Language Models for Kazakh from Scratch
por: Tukenov, Saken
Publicado: (2026)
por: Tukenov, Saken
Publicado: (2026)
Speak Kazakh: Language Ideologies in Kazakhstan's Social media in Times of Russian–Ukrainian War
por: Alina Kamalova
Publicado: (2025)
por: Alina Kamalova
Publicado: (2025)
Minimal Pair-Based Evaluation of Code-Switching
por: Sterner, Igor, et al.
Publicado: (2025)
por: Sterner, Igor, et al.
Publicado: (2025)
REPA: Russian Error Types Annotation for Evaluating Text Generation and Judgment Capabilities
por: Pugachev, Alexander, et al.
Publicado: (2025)
por: Pugachev, Alexander, et al.
Publicado: (2025)
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding
por: Zaranis, Emmanouil, et al.
Publicado: (2025)
por: Zaranis, Emmanouil, et al.
Publicado: (2025)
Fine-tuning BERT with Bidirectional LSTM for Fine-grained Movie Reviews Sentiment Analysis
por: Nkhata, Gibson, et al.
Publicado: (2025)
por: Nkhata, Gibson, et al.
Publicado: (2025)
A Comparative Analysis of Classical Machine Learning and Deep Learning Approaches for Sentiment Classification on IMDb Movie Reviews
por: Safitri, Erma Daniar, et al.
Publicado: (2026)
por: Safitri, Erma Daniar, et al.
Publicado: (2026)
KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models
por: Zaghouani, Wajdi, et al.
Publicado: (2026)
por: Zaghouani, Wajdi, et al.
Publicado: (2026)
Ejemplares similares
-
KazSAnDRA: Kazakh Sentiment Analysis Dataset of Reviews and Attitudes
por: Yeshpanov, Rustem, et al.
Publicado: (2024) -
KazParC: Kazakh Parallel Corpus for Machine Translation
por: Yeshpanov, Rustem, et al.
Publicado: (2024) -
KazQAD: Kazakh Open-Domain Question Answering Dataset
por: Yeshpanov, Rustem, et al.
Publicado: (2024) -
Using Songs to Improve Kazakh Automatic Speech Recognition
por: Yeshpanov, Rustem
Publicado: (2026) -
KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan
por: Togmanov, Mukhammed, et al.
Publicado: (2025)