Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Haoxiang, Yu, Da, Zhang, Huishuai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Synthesize Privacy-Preserving High-Resolution Images via Private Textual Intermediaries
by: Wang, Haoxiang, et al.
Published: (2025)
by: Wang, Haoxiang, et al.
Published: (2025)
How Worst-Case Are Adversarial Attacks? Linking Adversarial and Perturbation Robustness
by: Rossolini, Giulio
Published: (2026)
by: Rossolini, Giulio
Published: (2026)
Worst-Case Symbolic Constraints Analysis and Generalisation with Large Language Models
by: Koh, Daniel, et al.
Published: (2025)
by: Koh, Daniel, et al.
Published: (2025)
On the Worst Prompt Performance of Large Language Models
by: Cao, Bowen, et al.
Published: (2024)
by: Cao, Bowen, et al.
Published: (2024)
OLion: Approaching the Hadamard Ideal by Intersecting Spectral and $\ell_{\infty}$ Implicit Biases
by: Wang, Zixiao, et al.
Published: (2026)
by: Wang, Zixiao, et al.
Published: (2026)
AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training
by: Zhang, Huishuai, et al.
Published: (2025)
by: Zhang, Huishuai, et al.
Published: (2025)
Beyond Prompts: Dynamic Conversational Benchmarking of Large Language Models
by: Castillo-Bolado, David, et al.
Published: (2024)
by: Castillo-Bolado, David, et al.
Published: (2024)
Beyond A Fixed Seal: Adaptive Stealing Watermark in Large Language Models
by: Zhang, Shuhao, et al.
Published: (2026)
by: Zhang, Shuhao, et al.
Published: (2026)
Estimating Worst-Case Frontier Risks of Open-Weight LLMs
by: Wallace, Eric, et al.
Published: (2025)
by: Wallace, Eric, et al.
Published: (2025)
ARMOR 2025: A Military-Aligned Benchmark for Evaluating Large Language Model Safety Beyond Civilian Contexts
by: Johns, Sydney, et al.
Published: (2026)
by: Johns, Sydney, et al.
Published: (2026)
Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text
by: Greenfeld, Refael Shaked, et al.
Published: (2026)
by: Greenfeld, Refael Shaked, et al.
Published: (2026)
Learning to Generate Formally Verifiable Step-by-Step Logic Reasoning via Structured Formal Intermediaries
by: Chen, Luoxin, et al.
Published: (2026)
by: Chen, Luoxin, et al.
Published: (2026)
Evidence-Enhanced Triplet Generation Framework for Hallucination Alleviation in Generative Question Answering
by: Du, Haowei, et al.
Published: (2024)
by: Du, Haowei, et al.
Published: (2024)
Beyond Facts: Benchmarking Distributional Reading Comprehension in Large Language Models
by: Guo, Pei-Fu, et al.
Published: (2026)
by: Guo, Pei-Fu, et al.
Published: (2026)
Enterprise Large Language Model Evaluation Benchmark
by: Wang, Liya, et al.
Published: (2025)
by: Wang, Liya, et al.
Published: (2025)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
by: Li, Yuangang, et al.
Published: (2026)
by: Li, Yuangang, et al.
Published: (2026)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
by: Shi, Zhichao, et al.
Published: (2025)
by: Shi, Zhichao, et al.
Published: (2025)
Building Decision Making Models Through Language Model Regime
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
In-context Pretraining: Language Modeling Beyond Document Boundaries
by: Shi, Weijia, et al.
Published: (2023)
by: Shi, Weijia, et al.
Published: (2023)
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models
by: Zheng, Baihui, et al.
Published: (2025)
by: Zheng, Baihui, et al.
Published: (2025)
AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026)
by: Zhang, Tanghaoran, et al.
Published: (2026)
Beyond Facts: Evaluating Intent Hallucination in Large Language Models
by: Hao, Yijie, et al.
Published: (2025)
by: Hao, Yijie, et al.
Published: (2025)
The Evaluation Game: Beyond Static LLM Benchmarking
by: Wang, Paul, et al.
Published: (2026)
by: Wang, Paul, et al.
Published: (2026)
Towards Generalist Prompting for Large Language Models by Mental Models
by: Guan, Haoxiang, et al.
Published: (2024)
by: Guan, Haoxiang, et al.
Published: (2024)
CPSDBench: A Large Language Model Evaluation Benchmark and Baseline for Chinese Public Security Domain
by: Tong, Xin, et al.
Published: (2024)
by: Tong, Xin, et al.
Published: (2024)
Beyond Fixed Frames: Dynamic Character-Aligned Speech Tokenization
by: Della Libera, Luca, et al.
Published: (2026)
by: Della Libera, Luca, et al.
Published: (2026)
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models
by: Chen, Xiaoyang, et al.
Published: (2025)
by: Chen, Xiaoyang, et al.
Published: (2025)
Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games
by: Fan, Mingyuan, et al.
Published: (2026)
by: Fan, Mingyuan, et al.
Published: (2026)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
by: Li, Jiayu, et al.
Published: (2025)
by: Li, Jiayu, et al.
Published: (2025)
Automatic and Universal Prompt Injection Attacks against Large Language Models
by: Liu, Xiaogeng, et al.
Published: (2024)
by: Liu, Xiaogeng, et al.
Published: (2024)
BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
by: Li, Yige, et al.
Published: (2024)
by: Li, Yige, et al.
Published: (2024)
Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
by: Wang, Jun, et al.
Published: (2025)
by: Wang, Jun, et al.
Published: (2025)
SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment
by: Jiang, Sihang, et al.
Published: (2026)
by: Jiang, Sihang, et al.
Published: (2026)
ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval
by: Zhang, Honglei, et al.
Published: (2026)
by: Zhang, Honglei, et al.
Published: (2026)
Benchmarking Reasoning Robustness in Large Language Models
by: Yu, Tong, et al.
Published: (2025)
by: Yu, Tong, et al.
Published: (2025)
Sasha: Creative Goal-Oriented Reasoning in Smart Homes with Large Language Models
by: King, Evan, et al.
Published: (2023)
by: King, Evan, et al.
Published: (2023)
Beyond Boundaries: A Comprehensive Survey of Transferable Attacks on AI Systems
by: Wang, Guangjing, et al.
Published: (2023)
by: Wang, Guangjing, et al.
Published: (2023)
Truly Assessing Fluid Intelligence of Large Language Models through Dynamic Reasoning Evaluation
by: Yang, Yue, et al.
Published: (2025)
by: Yang, Yue, et al.
Published: (2025)
Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models
by: Yang, Wanqi, et al.
Published: (2024)
by: Yang, Wanqi, et al.
Published: (2024)
Large Language Models Leverage External Knowledge to Extend Clinical Insight Beyond Language Boundaries
by: Wu, Jiageng, et al.
Published: (2023)
by: Wu, Jiageng, et al.
Published: (2023)
Similar Items
-
Synthesize Privacy-Preserving High-Resolution Images via Private Textual Intermediaries
by: Wang, Haoxiang, et al.
Published: (2025) -
How Worst-Case Are Adversarial Attacks? Linking Adversarial and Perturbation Robustness
by: Rossolini, Giulio
Published: (2026) -
Worst-Case Symbolic Constraints Analysis and Generalisation with Large Language Models
by: Koh, Daniel, et al.
Published: (2025) -
On the Worst Prompt Performance of Large Language Models
by: Cao, Bowen, et al.
Published: (2024) -
OLion: Approaching the Hadamard Ideal by Intersecting Spectral and $\ell_{\infty}$ Implicit Biases
by: Wang, Zixiao, et al.
Published: (2026)