SemBench: A Universal Semantic Framework for LLM Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Zubillaga, Mikel, Perez, Naiara, Sainz, Oscar, Rigau, German |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Merge and Conquer: Instructing Multilingual Models by Adding Target Language Weights
by: Valero, Eneko, et al.
Published: (2026)
by: Valero, Eneko, et al.
Published: (2026)
Latxa: An Open Language Model and Evaluation Suite for Basque
by: Etxaniz, Julen, et al.
Published: (2024)
by: Etxaniz, Julen, et al.
Published: (2024)
Event Extraction in Basque: Typologically motivated Cross-Lingual Transfer-Learning Analysis
by: Zubillaga, Mikel, et al.
Published: (2024)
by: Zubillaga, Mikel, et al.
Published: (2024)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
by: Barnes, Jeremy, et al.
Published: (2025)
by: Barnes, Jeremy, et al.
Published: (2025)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
by: Moore, Robert J., et al.
Published: (2026)
by: Moore, Robert J., et al.
Published: (2026)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
by: Deng, Shihan, et al.
Published: (2024)
by: Deng, Shihan, et al.
Published: (2024)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
Do not be greedy, Think Twice: Sampling and Selection for Document-level Information Extraction
by: Zubillaga, Mikel, et al.
Published: (2026)
by: Zubillaga, Mikel, et al.
Published: (2026)
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
LoSemB: Logic-Guided Semantic Bridging for Inductive Tool Retrieval
by: Zhuang, Luyao, et al.
Published: (2025)
by: Zhuang, Luyao, et al.
Published: (2025)
SemViQA: A Semantic Question Answering System for Vietnamese Information Fact-Checking
by: Tran, Dien X., et al.
Published: (2025)
by: Tran, Dien X., et al.
Published: (2025)
BERnaT: Basque Encoders for Representing Natural Textual Diversity
by: Azurmendi, Ekhi, et al.
Published: (2025)
by: Azurmendi, Ekhi, et al.
Published: (2025)
HiTZ at VarDial 2025 NorSID: Overcoming Data Scarcity with Language Transfer and Automatic Data Annotation
by: Bengoetxea, Jaione, et al.
Published: (2024)
by: Bengoetxea, Jaione, et al.
Published: (2024)
LLM-KG-Bench 3.0: A Compass for SemanticTechnology Capabilities in the Ocean of LLMs
by: Meyer, Lars-Peter, et al.
Published: (2025)
by: Meyer, Lars-Peter, et al.
Published: (2025)
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
by: Tan, Haoran, et al.
Published: (2025)
by: Tan, Haoran, et al.
Published: (2025)
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks
by: Schmidt, Jan-Philipp
Published: (2026)
by: Schmidt, Jan-Philipp
Published: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
JP-TL-Bench: Anchored Pairwise LLM Evaluation for Bidirectional Japanese-English Translation
by: Lin, Leonard, et al.
Published: (2026)
by: Lin, Leonard, et al.
Published: (2026)
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration
by: Zhao, Junjie, et al.
Published: (2026)
by: Zhao, Junjie, et al.
Published: (2026)
SCORE: A Semantic Evaluation Framework for Generative Document Parsing
by: Li, Renyu, et al.
Published: (2025)
by: Li, Renyu, et al.
Published: (2025)
SemBench: A Benchmark for Semantic Query Processing Engines
by: Lao, Jiale, et al.
Published: (2025)
by: Lao, Jiale, et al.
Published: (2025)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
by: Markowitz, Elan, et al.
Published: (2025)
by: Markowitz, Elan, et al.
Published: (2025)
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
by: Mohamed, Anas, et al.
Published: (2025)
by: Mohamed, Anas, et al.
Published: (2025)
SemRAG: Semantic Knowledge-Augmented RAG for Improved Question-Answering
by: Zhong, Kezhen, et al.
Published: (2025)
by: Zhong, Kezhen, et al.
Published: (2025)
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
by: Zhao, Xinye, et al.
Published: (2025)
by: Zhao, Xinye, et al.
Published: (2025)
OckBench: Measuring the Efficiency of LLM Reasoning
by: Du, Zheng, et al.
Published: (2025)
by: Du, Zheng, et al.
Published: (2025)
Integrated Framework for LLM Evaluation with Answer Generation
by: Lee, Sujeong, et al.
Published: (2025)
by: Lee, Sujeong, et al.
Published: (2025)
Instructing Large Language Models for Low-Resource Languages: A Systematic Study for Basque
by: Sainz, Oscar, et al.
Published: (2025)
by: Sainz, Oscar, et al.
Published: (2025)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
NarraBench: A Comprehensive Framework for Narrative Benchmarking
by: Hamilton, Sil, et al.
Published: (2025)
by: Hamilton, Sil, et al.
Published: (2025)
NewsBench: A Systematic Evaluation Framework for Assessing Editorial Capabilities of Large Language Models in Chinese Journalism
by: Li, Miao, et al.
Published: (2024)
by: Li, Miao, et al.
Published: (2024)
MoralBench: Moral Evaluation of LLMs
by: Ji, Jianchao, et al.
Published: (2024)
by: Ji, Jianchao, et al.
Published: (2024)
Evaluating Consistencies in LLM responses through a Semantic Clustering of Question Answering
by: Lee, Yanggyu, et al.
Published: (2024)
by: Lee, Yanggyu, et al.
Published: (2024)
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
by: Zhu, Kunlun, et al.
Published: (2025)
by: Zhu, Kunlun, et al.
Published: (2025)
keepitsimple at SemEval-2025 Task 3: LLM-Uncertainty based Approach for Multilingual Hallucination Span Detection
by: Vemula, Saketh Reddy, et al.
Published: (2025)
by: Vemula, Saketh Reddy, et al.
Published: (2025)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
by: Liu, Jiayu, et al.
Published: (2025)
by: Liu, Jiayu, et al.
Published: (2025)
UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches
by: Wang, Chao, et al.
Published: (2024)
by: Wang, Chao, et al.
Published: (2024)
TravelBench : Exploring LLM Performance in Low-Resource Domains
by: Billa, Srinivas, et al.
Published: (2025)
by: Billa, Srinivas, et al.
Published: (2025)
Similar Items
-
Merge and Conquer: Instructing Multilingual Models by Adding Target Language Weights
by: Valero, Eneko, et al.
Published: (2026) -
Latxa: An Open Language Model and Evaluation Suite for Basque
by: Etxaniz, Julen, et al.
Published: (2024) -
Event Extraction in Basque: Typologically motivated Cross-Lingual Transfer-Learning Analysis
by: Zubillaga, Mikel, et al.
Published: (2024) -
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
by: Barnes, Jeremy, et al.
Published: (2025) -
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
by: Moore, Robert J., et al.
Published: (2026)