Batayan: A Filipino NLP benchmark for evaluating Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Montalan, Jann Railey, Layacan, Jimson Paulo, Africa, David Demitri, Flores, Richell Isaiah, Lopez II, Michael T., Magsajo, Theresa Denise, Cayabyab, Anjanette, Tjhi, William Chandra |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Kalahi: A handcrafted, grassroots cultural LLM evaluation suite for Filipino
por: Montalan, Jann Railey, et al.
Publicado: (2024)
por: Montalan, Jann Railey, et al.
Publicado: (2024)
BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models
por: Aung, Thura, et al.
Publicado: (2026)
por: Aung, Thura, et al.
Publicado: (2026)
SEA-HELM: Southeast Asian Holistic Evaluation of Language Models
por: Susanto, Yosephine, et al.
Publicado: (2025)
por: Susanto, Yosephine, et al.
Publicado: (2025)
SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
por: Ponwitayarat, Wuttikorn, et al.
Publicado: (2025)
por: Ponwitayarat, Wuttikorn, et al.
Publicado: (2025)
Identifying a Circuit for Verb Conjugation in GPT-2
por: Africa, David Demitri
Publicado: (2025)
por: Africa, David Demitri
Publicado: (2025)
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
por: Ivanov, Igor, et al.
Publicado: (2026)
por: Ivanov, Igor, et al.
Publicado: (2026)
Steering Awareness: Detecting Activation Steering from Within
por: Rivera, Joshua Fonseca, et al.
Publicado: (2025)
por: Rivera, Joshua Fonseca, et al.
Publicado: (2025)
Does Self-Evaluation Enable Wireheading in Language Models?
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
Operaciones garantizadas internacionales: debate actual y posibles soluciones
por: Anjanette H. Raymond
Publicado: (2011)
por: Anjanette H. Raymond
Publicado: (2011)
Consistency Training while Mitigating Obfuscation via Rate Matching
por: Imran, Sohaib, et al.
Publicado: (2026)
por: Imran, Sohaib, et al.
Publicado: (2026)
Learning Dynamics of Meta-Learning in Small Model Pretraining
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
por: Weiss, Yuval, et al.
Publicado: (2025)
por: Weiss, Yuval, et al.
Publicado: (2025)
Learning Modular Exponentiation with Transformers
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned
por: Ong, Brandon, et al.
Publicado: (2025)
por: Ong, Brandon, et al.
Publicado: (2025)
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
A Reinforcement Learning Inspired Latent Yield Based Adaptive Algorithm Switching Mechanism
por: Nair, Jayprakash S., et al.
Publicado: (2026)
por: Nair, Jayprakash S., et al.
Publicado: (2026)
Large Language Models for EEG: A Comprehensive Survey and Taxonomy
por: Babu, Naseem, et al.
Publicado: (2025)
por: Babu, Naseem, et al.
Publicado: (2025)
ThaiCoref: Thai Coreference Resolution Dataset
por: Trakuekul, Pontakorn, et al.
Publicado: (2024)
por: Trakuekul, Pontakorn, et al.
Publicado: (2024)
FiLLM -- A Filipino-optimized Large Language Model based on Southeast Asia Large Language Model (SEALLM)
por: Maminta, Carlos Jude G., et al.
Publicado: (2025)
por: Maminta, Carlos Jude G., et al.
Publicado: (2025)
Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
por: Martinez, Richard Diehl, et al.
Publicado: (2025)
por: Martinez, Richard Diehl, et al.
Publicado: (2025)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
por: Cencerrado, Iván Vicente Moreno, et al.
Publicado: (2025)
por: Cencerrado, Iván Vicente Moreno, et al.
Publicado: (2025)
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
por: Tan, Daniel, et al.
Publicado: (2025)
por: Tan, Daniel, et al.
Publicado: (2025)
Thai Universal Dependency Treebank
por: Sriwirote, Panyut, et al.
Publicado: (2024)
por: Sriwirote, Panyut, et al.
Publicado: (2024)
Filipino labour in Hawaii
Publicado: (1927)
Publicado: (1927)
Reasoning Models Reason Well, Until They Don't
por: Rameshkumar, Revanth, et al.
Publicado: (2025)
por: Rameshkumar, Revanth, et al.
Publicado: (2025)
A single‐site feasibility randomised controlled trial comparing ‘my hypo compass’ short pyscho‐educational intervention with standard care alone in individuals with type 1 diabetes and impaired awareness of hypoglycaemia
por: Ayat Bashir, et al.
Publicado: (2024)
por: Ayat Bashir, et al.
Publicado: (2024)
A call for reporting of tumor‐specific outcomes in studies of DPYD genotyping
por: Jean De Dieu Ndayishimiye, et al.
Publicado: (2024)
por: Jean De Dieu Ndayishimiye, et al.
Publicado: (2024)
Compositional Analysis of Cultivated and Wild‐Harvested Boswellia sacra Frankincense Resin Essential Oils in Oman
por: Anjanette DeCarlo, et al.
Publicado: (2025)
por: Anjanette DeCarlo, et al.
Publicado: (2025)
Tapping into the Assets of First-Generation Students during Times of Transition
por: Hands, Africa S.
Publicado: (2020)
por: Hands, Africa S.
Publicado: (2020)
What's Your Type? An Examination of First-Year Doctoral Student Motivation
por: Hands, Africa S.
Publicado: (2020)
por: Hands, Africa S.
Publicado: (2020)
Public Libraries: Your Partner in Increasing College Literacy among Nontraditional Prospective Students
por: Hands, Africa S.
Publicado: (2023)
por: Hands, Africa S.
Publicado: (2023)
Successfully Serving the College Bound
por: Hands, Africa S.
Publicado: (2015)
por: Hands, Africa S.
Publicado: (2015)
Peer Genius Bar: Using the Wisdom of the Crowd to Learn Technology Tools
por: Hands, Africa S.
Publicado: (2023)
por: Hands, Africa S.
Publicado: (2023)
What Doctoral Student Motivation Tells Us about the Future of LIS Education
por: Hands, Africa S.
Publicado: (2018)
por: Hands, Africa S.
Publicado: (2018)
Voluntarism and Political Conflict in Barbados, 1814–33*
por: Isaiah Silvers
Publicado: (2026)
por: Isaiah Silvers
Publicado: (2026)
Antología de ensayos / Isaiah Berlin ; introducción Joaquín Abell n
por: Berlin, Isaiah
por: Berlin, Isaiah
Qué es la libertad política ?
por: Isaiah, Berlin
Publicado: (2006)
por: Isaiah, Berlin
Publicado: (2006)
The GATT, quantitative restrictions, and the balance of payments / Isaiah Frank
por: Frank, Isaiah
Publicado: (1987)
por: Frank, Isaiah
Publicado: (1987)
Cuatro ensayos sobre la libertad / Isaiah Berlin
por: Berlin, Isaiah
por: Berlin, Isaiah
Decadencia de las ideas utópicas en Occidente
por: Berlin, Isaiah
Publicado: (1986)
por: Berlin, Isaiah
Publicado: (1986)
Ejemplares similares
-
Kalahi: A handcrafted, grassroots cultural LLM evaluation suite for Filipino
por: Montalan, Jann Railey, et al.
Publicado: (2024) -
BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models
por: Aung, Thura, et al.
Publicado: (2026) -
SEA-HELM: Southeast Asian Holistic Evaluation of Language Models
por: Susanto, Yosephine, et al.
Publicado: (2025) -
SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
por: Ponwitayarat, Wuttikorn, et al.
Publicado: (2025) -
Identifying a Circuit for Verb Conjugation in GPT-2
por: Africa, David Demitri
Publicado: (2025)