SetLexSem Challenge: Using Set Operations to Evaluate the Lexical and Semantic Robustness of Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Akhbari, Bardiya, Gawali, Manish, Dronen, Nicholas A. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Large Language Models Using Contrast Sets: An Experimental Approach
by: Sanwal, Manish
Published: (2024)
by: Sanwal, Manish
Published: (2024)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
by: Ren, Huimin, et al.
Published: (2025)
by: Ren, Huimin, et al.
Published: (2025)
How Lexical is Bilingual Lexicon Induction?
by: Kohli, Harsh, et al.
Published: (2024)
by: Kohli, Harsh, et al.
Published: (2024)
ProLex: A Benchmark for Language Proficiency-oriented Lexical Substitution
by: Zhang, Xuanming, et al.
Published: (2024)
by: Zhang, Xuanming, et al.
Published: (2024)
Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection
by: Phan-Tat, Bach, et al.
Published: (2026)
by: Phan-Tat, Bach, et al.
Published: (2026)
Evaluating Concurrent Robustness of Language Models Across Diverse Challenge Sets
by: Gupta, Vatsal, et al.
Published: (2023)
by: Gupta, Vatsal, et al.
Published: (2023)
MultiLexNorm++: A Unified Benchmark and a Generative Model for Lexical Normalization for Asian Languages
by: Buaphet, Weerayut, et al.
Published: (2026)
by: Buaphet, Weerayut, et al.
Published: (2026)
FastLexRank: Efficient Lexical Ranking for Structuring Social Media Posts
by: Li, Mao, et al.
Published: (2024)
by: Li, Mao, et al.
Published: (2024)
NeLLCom-Lex: A Neural-agent Framework to Study the Interplay between Lexical Systems and Language Use
by: Zhang, Yuqing, et al.
Published: (2025)
by: Zhang, Yuqing, et al.
Published: (2025)
ViLexNorm: A Lexical Normalization Corpus for Vietnamese Social Media Text
by: Nguyen, Thanh-Nhi, et al.
Published: (2024)
by: Nguyen, Thanh-Nhi, et al.
Published: (2024)
ViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2025)
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2025)
LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models
by: Li, Haitao, et al.
Published: (2024)
by: Li, Haitao, et al.
Published: (2024)
LexSumm and LexT5: Benchmarking and Modeling Legal Summarization Tasks in English
by: Santosh, T. Y. S. S., et al.
Published: (2024)
by: Santosh, T. Y. S. S., et al.
Published: (2024)
SemEval-2024 Task 1: Semantic Textual Relatedness for African and Asian Languages
by: Ousidhoum, Nedjma, et al.
Published: (2024)
by: Ousidhoum, Nedjma, et al.
Published: (2024)
SetBERT: Enhancing Retrieval Performance for Boolean Logic and Set Operation Queries
by: Mai, Quan, et al.
Published: (2024)
by: Mai, Quan, et al.
Published: (2024)
Probing Large Language Models for Scalar Adjective Lexical Semantics and Scalar Diversity Pragmatics
by: Lin, Fangru, et al.
Published: (2024)
by: Lin, Fangru, et al.
Published: (2024)
Using Language Models to Disambiguate Lexical Choices in Translation
by: Barua, Josh, et al.
Published: (2024)
by: Barua, Josh, et al.
Published: (2024)
SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation
by: Hou, Abe Bohan, et al.
Published: (2023)
by: Hou, Abe Bohan, et al.
Published: (2023)
PsychoLex: Unveiling the Psychological Mind of Large Language Models
by: Abbasi, Mohammad Amin, et al.
Published: (2024)
by: Abbasi, Mohammad Amin, et al.
Published: (2024)
BabyLM Challenge: Exploring the Effect of Variation Sets on Language Model Training Efficiency
by: Haga, Akari, et al.
Published: (2024)
by: Haga, Akari, et al.
Published: (2024)
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
by: Moghe, Nikita, et al.
Published: (2024)
by: Moghe, Nikita, et al.
Published: (2024)
DySem: Uncovering Dynamic Semantic Components of Large Language Models for Calculating Semantic Textual Similarity
by: Zheng, Kaijie, et al.
Published: (2026)
by: Zheng, Kaijie, et al.
Published: (2026)
VISLA Benchmark: Evaluating Embedding Sensitivity to Semantic and Lexical Alterations
by: Dumpala, Sri Harsha, et al.
Published: (2024)
by: Dumpala, Sri Harsha, et al.
Published: (2024)
SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages
by: Ousidhoum, Nedjma, et al.
Published: (2024)
by: Ousidhoum, Nedjma, et al.
Published: (2024)
A Multidimensional Framework for Evaluating Lexical Semantic Change with Social Science Applications
by: Baes, Naomi, et al.
Published: (2024)
by: Baes, Naomi, et al.
Published: (2024)
Evaluating Distributed Representations for Multi-Level Lexical Semantics: A Research Proposal
by: Liu, Zhu
Published: (2024)
by: Liu, Zhu
Published: (2024)
Evaluating and Mitigating Social Bias for Large Language Models in Open-ended Settings
by: Liu, Zhao, et al.
Published: (2024)
by: Liu, Zhao, et al.
Published: (2024)
Harnessing the Intrinsic Knowledge of Pretrained Language Models for Challenging Text Classification Settings
by: Gao, Lingyu
Published: (2024)
by: Gao, Lingyu
Published: (2024)
Set the Clock: Temporal Alignment of Pretrained Language Models
by: Zhao, Bowen, et al.
Published: (2024)
by: Zhao, Bowen, et al.
Published: (2024)
Analyzing Semantic Change through Lexical Replacements
by: Periti, Francesco, et al.
Published: (2024)
by: Periti, Francesco, et al.
Published: (2024)
Rethinking Metrics for Lexical Semantic Change Detection
by: Goworek, Roksana, et al.
Published: (2026)
by: Goworek, Roksana, et al.
Published: (2026)
Speakers Fill Lexical Semantic Gaps with Context
by: Pimentel, Tiago, et al.
Published: (2020)
by: Pimentel, Tiago, et al.
Published: (2020)
ODE: Open-Set Evaluation of Hallucinations in Multimodal Large Language Models
by: Tu, Yahan, et al.
Published: (2024)
by: Tu, Yahan, et al.
Published: (2024)
From Superficial Patterns to Semantic Understanding: Fine-Tuning Language Models on Contrast Sets
by: Petrov, Daniel
Published: (2025)
by: Petrov, Daniel
Published: (2025)
Evaluation of Language Models in the Medical Context Under Resource-Constrained Settings
by: Posada, Andrea, et al.
Published: (2024)
by: Posada, Andrea, et al.
Published: (2024)
SemPA: Improving Sentence Embeddings of Large Language Models through Semantic Preference Alignment
by: Chen, Ziyang, et al.
Published: (2026)
by: Chen, Ziyang, et al.
Published: (2026)
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
SemBench: A Universal Semantic Framework for LLM Evaluation
by: Zubillaga, Mikel, et al.
Published: (2026)
by: Zubillaga, Mikel, et al.
Published: (2026)
Hierarchical Lexical Manifold Projection in Large Language Models: A Novel Mechanism for Multi-Scale Semantic Representation
by: Martus, Natasha, et al.
Published: (2025)
by: Martus, Natasha, et al.
Published: (2025)
Improving Interpretability of Lexical Semantic Change with Neurobiological Features
by: Oda, Kohei, et al.
Published: (2026)
by: Oda, Kohei, et al.
Published: (2026)
Similar Items
-
Evaluating Large Language Models Using Contrast Sets: An Experimental Approach
by: Sanwal, Manish
Published: (2024) -
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
by: Ren, Huimin, et al.
Published: (2025) -
How Lexical is Bilingual Lexicon Induction?
by: Kohli, Harsh, et al.
Published: (2024) -
ProLex: A Benchmark for Language Proficiency-oriented Lexical Substitution
by: Zhang, Xuanming, et al.
Published: (2024) -
Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection
by: Phan-Tat, Bach, et al.
Published: (2026)