NEO-BENCH: Evaluating Robustness of Large Language Models with Neologisms
Fuente:
arXiv
Salvato in:
| Autori principali: | Zheng, Jonathan, Ritter, Alan, Xu, Wei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models
di: Chen, Kai, et al.
Pubblicazione: (2025)
di: Chen, Kai, et al.
Pubblicazione: (2025)
From 124 Million Tokens to 1,021 Neologisms: A Large-Scale Pipeline for Automatic Neologism Detection
di: Rossini, Diego, et al.
Pubblicazione: (2026)
di: Rossini, Diego, et al.
Pubblicazione: (2026)
Probabilistic Reasoning with LLMs for k-anonymity Estimation
di: Zheng, Jonathan, et al.
Pubblicazione: (2025)
di: Zheng, Jonathan, et al.
Pubblicazione: (2025)
MaterialBENCH: Evaluating College-Level Materials Science Problem-Solving Abilities of Large Language Models
di: Yoshitake, Michiko, et al.
Pubblicazione: (2024)
di: Yoshitake, Michiko, et al.
Pubblicazione: (2024)
Neologism Learning for Controllability and Self-Verbalization
di: Hewitt, John, et al.
Pubblicazione: (2025)
di: Hewitt, John, et al.
Pubblicazione: (2025)
Having Beer after Prayer? Measuring Cultural Bias in Large Language Models
di: Naous, Tarek, et al.
Pubblicazione: (2023)
di: Naous, Tarek, et al.
Pubblicazione: (2023)
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
di: Jang, Kyochul, et al.
Pubblicazione: (2025)
di: Jang, Kyochul, et al.
Pubblicazione: (2025)
CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space
di: Hwang, Yeonjun, et al.
Pubblicazione: (2026)
di: Hwang, Yeonjun, et al.
Pubblicazione: (2026)
DI-BENCH: Benchmarking Large Language Models on Dependency Inference with Testable Repositories at Scale
di: Zhang, Linghao, et al.
Pubblicazione: (2025)
di: Zhang, Linghao, et al.
Pubblicazione: (2025)
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
di: Guo, Ruohao, et al.
Pubblicazione: (2023)
di: Guo, Ruohao, et al.
Pubblicazione: (2023)
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
di: Wang, Zekun, et al.
Pubblicazione: (2025)
di: Wang, Zekun, et al.
Pubblicazione: (2025)
Anticipatory Evaluation of Language Models
di: Park, Jungsoo, et al.
Pubblicazione: (2025)
di: Park, Jungsoo, et al.
Pubblicazione: (2025)
CaT-BENCH: Benchmarking Language Model Understanding of Causal and Temporal Dependencies in Plans
di: Lal, Yash Kumar, et al.
Pubblicazione: (2024)
di: Lal, Yash Kumar, et al.
Pubblicazione: (2024)
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
di: Ji, Yifan, et al.
Pubblicazione: (2026)
di: Ji, Yifan, et al.
Pubblicazione: (2026)
What are Foundation Models Cooking in the Post-Soviet World?
di: Lavrouk, Anton, et al.
Pubblicazione: (2025)
di: Lavrouk, Anton, et al.
Pubblicazione: (2025)
Neologism Learning as a Parameter-Efficient Alternative to Fine-Tuning for Model Steering
di: Park, Sungjoon, et al.
Pubblicazione: (2025)
di: Park, Sungjoon, et al.
Pubblicazione: (2025)
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms
di: Ki, Dayeon, et al.
Pubblicazione: (2026)
di: Ki, Dayeon, et al.
Pubblicazione: (2026)
Learning to Route Languages for Multilingual Policy Optimization
di: Guo, Geyang, et al.
Pubblicazione: (2026)
di: Guo, Geyang, et al.
Pubblicazione: (2026)
KOCO-BENCH: Can Large Language Models Leverage Domain Knowledge in Software Development?
di: Jiang, Xue, et al.
Pubblicazione: (2026)
di: Jiang, Xue, et al.
Pubblicazione: (2026)
Language Models can Self-Improve at State-Value Estimation for Better Search
di: Mendes, Ethan, et al.
Pubblicazione: (2025)
di: Mendes, Ethan, et al.
Pubblicazione: (2025)
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
di: Yang, Runming, et al.
Pubblicazione: (2024)
di: Yang, Runming, et al.
Pubblicazione: (2024)
Granular Privacy Control for Geolocation with Vision Language Models
di: Mendes, Ethan, et al.
Pubblicazione: (2024)
di: Mendes, Ethan, et al.
Pubblicazione: (2024)
How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation
di: Guo, Ruohao, et al.
Pubblicazione: (2025)
di: Guo, Ruohao, et al.
Pubblicazione: (2025)
Investigating and Alleviating Harm Amplification in LLM Interactions
di: Guo, Ruohao, et al.
Pubblicazione: (2026)
di: Guo, Ruohao, et al.
Pubblicazione: (2026)
Reducing Privacy Risks in Online Self-Disclosures with Language Models
di: Dou, Yao, et al.
Pubblicazione: (2023)
di: Dou, Yao, et al.
Pubblicazione: (2023)
Do LLMs Know What Luxembourgish Borrows? Probing Lexical Neology in Low-Resource Multilingual Models
di: Hosseini-Kivanani, Nina
Pubblicazione: (2026)
di: Hosseini-Kivanani, Nina
Pubblicazione: (2026)
Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges
di: Wu, Xiaofeng, et al.
Pubblicazione: (2025)
di: Wu, Xiaofeng, et al.
Pubblicazione: (2025)
Stanceosaurus 2.0: Classifying Stance Towards Russian and Spanish Misinformation
di: Lavrouk, Anton, et al.
Pubblicazione: (2024)
di: Lavrouk, Anton, et al.
Pubblicazione: (2024)
NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning
di: Miao, Zhongtao, et al.
Pubblicazione: (2026)
di: Miao, Zhongtao, et al.
Pubblicazione: (2026)
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
di: Peng, Xiangyu, et al.
Pubblicazione: (2025)
di: Peng, Xiangyu, et al.
Pubblicazione: (2025)
NeoN: A Tool for Automated Detection, Linguistic and LLM-Driven Analysis of Neologisms in Polish
di: Tomaszewska, Aleksandra, et al.
Pubblicazione: (2025)
di: Tomaszewska, Aleksandra, et al.
Pubblicazione: (2025)
Evaluating the Retrieval Robustness of Large Language Models
di: Cao, Shuyang, et al.
Pubblicazione: (2025)
di: Cao, Shuyang, et al.
Pubblicazione: (2025)
Contrastive Knowledge Transfer and Robust Optimization for Secure Alignment of Large Language Models
di: Zheng, Jiasen, et al.
Pubblicazione: (2025)
di: Zheng, Jiasen, et al.
Pubblicazione: (2025)
SCORE: Systematic COnsistency and Robustness Evaluation for Large Language Models
di: Nalbandyan, Grigor, et al.
Pubblicazione: (2025)
di: Nalbandyan, Grigor, et al.
Pubblicazione: (2025)
Cognitive LLMs: Towards Integrating Cognitive Architectures and Large Language Models for Manufacturing Decision-making
di: Wu, Siyu, et al.
Pubblicazione: (2024)
di: Wu, Siyu, et al.
Pubblicazione: (2024)
Frustratingly Easy Label Projection for Cross-lingual Transfer
di: Chen, Yang, et al.
Pubblicazione: (2022)
di: Chen, Yang, et al.
Pubblicazione: (2022)
AC-EVAL: Evaluating Ancient Chinese Language Understanding in Large Language Models
di: Wei, Yuting, et al.
Pubblicazione: (2024)
di: Wei, Yuting, et al.
Pubblicazione: (2024)
Large Language Models Are Not Robust Multiple Choice Selectors
di: Zheng, Chujie, et al.
Pubblicazione: (2023)
di: Zheng, Chujie, et al.
Pubblicazione: (2023)
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
di: Rajabi, Navid, et al.
Pubblicazione: (2024)
di: Rajabi, Navid, et al.
Pubblicazione: (2024)
Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
di: Zhao, Raoyuan, et al.
Pubblicazione: (2025)
di: Zhao, Raoyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models
di: Chen, Kai, et al.
Pubblicazione: (2025) -
From 124 Million Tokens to 1,021 Neologisms: A Large-Scale Pipeline for Automatic Neologism Detection
di: Rossini, Diego, et al.
Pubblicazione: (2026) -
Probabilistic Reasoning with LLMs for k-anonymity Estimation
di: Zheng, Jonathan, et al.
Pubblicazione: (2025) -
MaterialBENCH: Evaluating College-Level Materials Science Problem-Solving Abilities of Large Language Models
di: Yoshitake, Michiko, et al.
Pubblicazione: (2024) -
Neologism Learning for Controllability and Self-Verbalization
di: Hewitt, John, et al.
Pubblicazione: (2025)