FilBench: Can LLMs Understand and Generate Filipino?
Fuente:
arXiv
Saved in:
| Main Authors: | Miranda, Lester James V., Aco, Elyanah, Manuel, Conner, Cruz, Jan Christian Blaise, Imperial, Joseph Marvin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpeciaLex: A Benchmark for In-Context Specialized Lexicon Learning
by: Imperial, Joseph Marvin, et al.
Published: (2024)
by: Imperial, Joseph Marvin, et al.
Published: (2024)
Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models
by: Gamboa, Lance Calvin Lim, et al.
Published: (2026)
by: Gamboa, Lance Calvin Lim, et al.
Published: (2026)
Standardize: Aligning Language Models with Expert-Defined Standards for Content Generation
by: Imperial, Joseph Marvin, et al.
Published: (2024)
by: Imperial, Joseph Marvin, et al.
Published: (2024)
Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
by: Imperial, Joseph Marvin, et al.
Published: (2025)
by: Imperial, Joseph Marvin, et al.
Published: (2025)
Safer Policy Compliance with Dynamic Epistemic Fallback
by: Imperial, Joseph Marvin, et al.
Published: (2026)
by: Imperial, Joseph Marvin, et al.
Published: (2026)
Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance
by: Imperial, Joseph Marvin, et al.
Published: (2025)
by: Imperial, Joseph Marvin, et al.
Published: (2025)
Extracting General-use Transformers for Low-resource Languages via Knowledge Distillation
by: Cruz, Jan Christian Blaise, et al.
Published: (2025)
by: Cruz, Jan Christian Blaise, et al.
Published: (2025)
PhageBench: Can LLMs Understand Raw Bacteriophage Genomes?
by: Hou, Yusen, et al.
Published: (2026)
by: Hou, Yusen, et al.
Published: (2026)
Idea First, Code Later: Disentangling Problem Solving from Code Generation in Evaluating LLMs for Competitive Programming
by: Hadhoud, Sama, et al.
Published: (2026)
by: Hadhoud, Sama, et al.
Published: (2026)
Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation
by: Miranda, Lester James V., et al.
Published: (2026)
by: Miranda, Lester James V., et al.
Published: (2026)
Sense Representations Are Inducible Interfaces
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
LLM Olympiad: Why Model Evaluation Needs a Sealed Exam
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
Multilinguality as Sense Adaptation
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
RareBench: Can LLMs Serve as Rare Diseases Specialists?
by: Chen, Xuanzhong, et al.
Published: (2024)
by: Chen, Xuanzhong, et al.
Published: (2024)
MedMT-Bench: Can LLMs Memorize and Understand Long Multi-Turn Conversations in Medical Scenarios?
by: Yang, Lin, et al.
Published: (2026)
by: Yang, Lin, et al.
Published: (2026)
HiligayNER: A Baseline Named Entity Recognition Model for Hiligaynon
by: Teves, James Ald, et al.
Published: (2025)
by: Teves, James Ald, et al.
Published: (2025)
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
by: Chambon, Pierre, et al.
Published: (2025)
by: Chambon, Pierre, et al.
Published: (2025)
KatotohananQA: Evaluating Truthfulness of Large Language Models in Filipino
by: Nery, Lorenzo Alfred, et al.
Published: (2025)
by: Nery, Lorenzo Alfred, et al.
Published: (2025)
Can LLMs Grasp Implicit Cultural Values? Benchmarking LLMs' Cultural Intelligence with CQ-Bench
by: Liu, Ziyi, et al.
Published: (2025)
by: Liu, Ziyi, et al.
Published: (2025)
The UD-NewsCrawl Treebank: Reflections and Challenges from a Large-scale Tagalog Syntactic Annotation Project
by: Aquino, Angelina A., et al.
Published: (2025)
by: Aquino, Angelina A., et al.
Published: (2025)
Thank You, Stingray: Multilingual Large Language Models Can Not (Yet) Disambiguate Cross-Lingual Word Sense
by: Cahyawijaya, Samuel, et al.
Published: (2024)
by: Cahyawijaya, Samuel, et al.
Published: (2024)
ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?
by: Chen, Canyu, et al.
Published: (2024)
by: Chen, Canyu, et al.
Published: (2024)
ExpressivityBench: Can LLMs Communicate Implicitly?
by: Tint, Joshua, et al.
Published: (2024)
by: Tint, Joshua, et al.
Published: (2024)
Can LLMs Understand the Implication of Emphasized Sentences in Dialogue?
by: Lin, Guan-Ting, et al.
Published: (2024)
by: Lin, Guan-Ting, et al.
Published: (2024)
Multilinguality at the Edge: Developing Language Models for the Global South
by: Miranda, Lester James V., et al.
Published: (2026)
by: Miranda, Lester James V., et al.
Published: (2026)
Can LLMs Understand Unvoiced Speech? Exploring EMG-to-Text Conversion with LLMs
by: Mohapatra, Payal, et al.
Published: (2025)
by: Mohapatra, Payal, et al.
Published: (2025)
Can LLMs Understand the Impact of Trauma? Costs and Benefits of LLMs Coding the Interviews of Firearm Violence Survivors
by: Zhu, Jessica H., et al.
Published: (2026)
by: Zhu, Jessica H., et al.
Published: (2026)
AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans
by: Xie, Wei, et al.
Published: (2025)
by: Xie, Wei, et al.
Published: (2025)
TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
by: Ahmed, Toufique, et al.
Published: (2024)
by: Ahmed, Toufique, et al.
Published: (2024)
Arabizi vs LLMs: Can the Genie Understand the Language of Aladdin?
by: Almaoui, Perla Al, et al.
Published: (2025)
by: Almaoui, Perla Al, et al.
Published: (2025)
AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
by: Gong, Kaixiong, et al.
Published: (2024)
by: Gong, Kaixiong, et al.
Published: (2024)
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
by: Zhou, Ziya, et al.
Published: (2024)
by: Zhou, Ziya, et al.
Published: (2024)
GeoGrid-Bench: Can Foundation Models Understand Multimodal Gridded Geo-Spatial Data?
by: Jiang, Bowen, et al.
Published: (2025)
by: Jiang, Bowen, et al.
Published: (2025)
MasalBench: A Benchmark for Contextual and Cross-Cultural Understanding of Persian Proverbs in LLMs
by: Kalhor, Ghazal, et al.
Published: (2026)
by: Kalhor, Ghazal, et al.
Published: (2026)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
by: Zhu, Hengchuan, et al.
Published: (2025)
by: Zhu, Hengchuan, et al.
Published: (2025)
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
by: Yoran, Ori, et al.
Published: (2024)
by: Yoran, Ori, et al.
Published: (2024)
RiddleBench: A New Generative Reasoning Benchmark for LLMs
by: Halder, Deepon, et al.
Published: (2025)
by: Halder, Deepon, et al.
Published: (2025)
Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark
by: Mayhew, Stephen, et al.
Published: (2023)
by: Mayhew, Stephen, et al.
Published: (2023)
M-RewardBench: Evaluating Reward Models in Multilingual Settings
by: Gureja, Srishti, et al.
Published: (2024)
by: Gureja, Srishti, et al.
Published: (2024)
LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
by: Valmeekam, Karthik, et al.
Published: (2024)
by: Valmeekam, Karthik, et al.
Published: (2024)
Similar Items
-
SpeciaLex: A Benchmark for In-Context Specialized Lexicon Learning
by: Imperial, Joseph Marvin, et al.
Published: (2024) -
Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models
by: Gamboa, Lance Calvin Lim, et al.
Published: (2026) -
Standardize: Aligning Language Models with Expert-Defined Standards for Content Generation
by: Imperial, Joseph Marvin, et al.
Published: (2024) -
Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
by: Imperial, Joseph Marvin, et al.
Published: (2025) -
Safer Policy Compliance with Dynamic Epistemic Fallback
by: Imperial, Joseph Marvin, et al.
Published: (2026)