Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yang, Lin, Chenghua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi
von: Fatimah, Shiza, et al.
Veröffentlicht: (2026)
von: Fatimah, Shiza, et al.
Veröffentlicht: (2026)
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
von: Jiang, Han, et al.
Veröffentlicht: (2024)
von: Jiang, Han, et al.
Veröffentlicht: (2024)
SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
von: Pham, Thinh, et al.
Veröffentlicht: (2025)
von: Pham, Thinh, et al.
Veröffentlicht: (2025)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
von: Ballon, Marthe, et al.
Veröffentlicht: (2026)
von: Ballon, Marthe, et al.
Veröffentlicht: (2026)
Adversarial Defence without Adversarial Defence: Enhancing Language Model Robustness via Instance-level Principal Component Removal
von: Wang, Yang, et al.
Veröffentlicht: (2025)
von: Wang, Yang, et al.
Veröffentlicht: (2025)
An Open Source Data Contamination Report for Large Language Models
von: Li, Yucheng, et al.
Veröffentlicht: (2023)
von: Li, Yucheng, et al.
Veröffentlicht: (2023)
Are Large Language Models Truly Smarter Than Humans?
von: M, Eshwar Reddy, et al.
Veröffentlicht: (2026)
von: M, Eshwar Reddy, et al.
Veröffentlicht: (2026)
LatestEval: Addressing Data Contamination in Language Model Evaluation through Dynamic and Time-Sensitive Test Construction
von: Li, Yucheng, et al.
Veröffentlicht: (2023)
von: Li, Yucheng, et al.
Veröffentlicht: (2023)
Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment
von: Zhang, Jianfei, et al.
Veröffentlicht: (2024)
von: Zhang, Jianfei, et al.
Veröffentlicht: (2024)
SUGAR: Leveraging Contextual Confidence for Smarter Retrieval
von: Zubkova, Hanna, et al.
Veröffentlicht: (2025)
von: Zubkova, Hanna, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models for Generalization and Robustness via Data Compression
von: Li, Yucheng, et al.
Veröffentlicht: (2024)
von: Li, Yucheng, et al.
Veröffentlicht: (2024)
Are AI-Generated Text Detectors Robust to Adversarial Perturbations?
von: Huang, Guanhua, et al.
Veröffentlicht: (2024)
von: Huang, Guanhua, et al.
Veröffentlicht: (2024)
Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
von: Pala, Tej Deep, et al.
Veröffentlicht: (2025)
von: Pala, Tej Deep, et al.
Veröffentlicht: (2025)
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
von: Han, Chao, et al.
Veröffentlicht: (2025)
von: Han, Chao, et al.
Veröffentlicht: (2025)
Round Trip Translation Defence against Large Language Model Jailbreaking Attacks
von: Yung, Canaan, et al.
Veröffentlicht: (2024)
von: Yung, Canaan, et al.
Veröffentlicht: (2024)
Reasoning with Sampling: Your Base Model is Smarter Than You Think
von: Karan, Aayush, et al.
Veröffentlicht: (2025)
von: Karan, Aayush, et al.
Veröffentlicht: (2025)
Detecting the Machine: A Comprehensive Benchmark of AI-Generated Text Detectors Across Architectures, Domains, and Adversarial Conditions
von: Baidya, Madhav S., et al.
Veröffentlicht: (2026)
von: Baidya, Madhav S., et al.
Veröffentlicht: (2026)
ChallengeMe: An Adversarial Learning-enabled Text Summarization Framework
von: Deng, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Deng, Xiaoyu, et al.
Veröffentlicht: (2025)
PFME: A Modular Approach for Fine-grained Hallucination Detection and Editing of Large Language Models
von: Deng, Kunquan, et al.
Veröffentlicht: (2024)
von: Deng, Kunquan, et al.
Veröffentlicht: (2024)
Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
von: Galisai, Marcello, et al.
Veröffentlicht: (2026)
von: Galisai, Marcello, et al.
Veröffentlicht: (2026)
BRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text
von: Wu, Jiageng, et al.
Veröffentlicht: (2025)
von: Wu, Jiageng, et al.
Veröffentlicht: (2025)
Are You Human? An Adversarial Benchmark to Expose LLMs
von: Gressel, Gilad, et al.
Veröffentlicht: (2024)
von: Gressel, Gilad, et al.
Veröffentlicht: (2024)
"When Data is Scarce, Prompt Smarter"... Approaches to Grammatical Error Correction in Low-Resource Settings
von: De, Somsubhra, et al.
Veröffentlicht: (2025)
von: De, Somsubhra, et al.
Veröffentlicht: (2025)
Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond
von: Chen, Rubing, et al.
Veröffentlicht: (2025)
von: Chen, Rubing, et al.
Veröffentlicht: (2025)
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)
Continuous Adversarial Text Representation Learning for Affective Recognition
von: Son, Seungah, et al.
Veröffentlicht: (2025)
von: Son, Seungah, et al.
Veröffentlicht: (2025)
A Generative Adversarial Attack for Multilingual Text Classifiers
von: Roth, Tom, et al.
Veröffentlicht: (2024)
von: Roth, Tom, et al.
Veröffentlicht: (2024)
CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models
von: LI, Yizhi, et al.
Veröffentlicht: (2024)
von: LI, Yizhi, et al.
Veröffentlicht: (2024)
Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure
von: Li, Zirui, et al.
Veröffentlicht: (2026)
von: Li, Zirui, et al.
Veröffentlicht: (2026)
T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
von: Wang, Qinsi, et al.
Veröffentlicht: (2026)
von: Wang, Qinsi, et al.
Veröffentlicht: (2026)
Text2World: Benchmarking Large Language Models for Symbolic World Model Generation
von: Hu, Mengkang, et al.
Veröffentlicht: (2025)
von: Hu, Mengkang, et al.
Veröffentlicht: (2025)
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges
von: Fu, Rao, et al.
Veröffentlicht: (2024)
von: Fu, Rao, et al.
Veröffentlicht: (2024)
German Text Embedding Clustering Benchmark
von: Wehrli, Silvan, et al.
Veröffentlicht: (2024)
von: Wehrli, Silvan, et al.
Veröffentlicht: (2024)
PTD-SQL: Partitioning and Targeted Drilling with LLMs in Text-to-SQL
von: Luo, Ruilin, et al.
Veröffentlicht: (2024)
von: Luo, Ruilin, et al.
Veröffentlicht: (2024)
DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
BEAVER: An Enterprise Benchmark for Text-to-SQL
von: Chen, Peter Baile, et al.
Veröffentlicht: (2024)
von: Chen, Peter Baile, et al.
Veröffentlicht: (2024)
Observing Micromotives and Macrobehavior of Large Language Models
von: Cheng, Yuyang, et al.
Veröffentlicht: (2024)
von: Cheng, Yuyang, et al.
Veröffentlicht: (2024)
TAD-Bench: A Comprehensive Benchmark for Embedding-Based Text Anomaly Detection
von: Cao, Yang, et al.
Veröffentlicht: (2025)
von: Cao, Yang, et al.
Veröffentlicht: (2025)
Benchmarking Open-Source Large Language Models on Healthcare Text Classification Tasks
von: Guo, Yuting, et al.
Veröffentlicht: (2025)
von: Guo, Yuting, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi
von: Fatimah, Shiza, et al.
Veröffentlicht: (2026) -
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
von: Jiang, Han, et al.
Veröffentlicht: (2024) -
SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
von: Pham, Thinh, et al.
Veröffentlicht: (2025) -
Benchmarks Saturate When The Model Gets Smarter Than The Judge
von: Ballon, Marthe, et al.
Veröffentlicht: (2026) -
Adversarial Defence without Adversarial Defence: Enhancing Language Model Robustness via Instance-level Principal Component Removal
von: Wang, Yang, et al.
Veröffentlicht: (2025)