BenchOverflow: Measuring Overflow in Large Language Models via Plain-Text Prompts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feiglin, Erin, Hutnik, Nir, Lapid, Raz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Activation Steering for Masked Diffusion Language Models
von: Shnaidman, Adi, et al.
Veröffentlicht: (2025)
von: Shnaidman, Adi, et al.
Veröffentlicht: (2025)
SastBench: A Benchmark for Testing Agentic SAST Triage
von: Feiglin, Jake, et al.
Veröffentlicht: (2026)
von: Feiglin, Jake, et al.
Veröffentlicht: (2026)
Few-shot Name Entity Recognition on StackOverflow
von: Chen, Xinwei, et al.
Veröffentlicht: (2024)
von: Chen, Xinwei, et al.
Veröffentlicht: (2024)
Can ChatGPT replace StackOverflow? A Study on Robustness and Reliability of Large Language Model Code Generation
von: Zhong, Li, et al.
Veröffentlicht: (2023)
von: Zhong, Li, et al.
Veröffentlicht: (2023)
Evaluating Privacy Questions From Stack Overflow: Can ChatGPT Compete?
von: Delile, Zack, et al.
Veröffentlicht: (2023)
von: Delile, Zack, et al.
Veröffentlicht: (2023)
CardiffNLP at CLEARS-2025: Prompting Large Language Models for Plain Language and Easy-to-Read Text Rewriting
von: Ayesh, Mutaz, et al.
Veröffentlicht: (2025)
von: Ayesh, Mutaz, et al.
Veröffentlicht: (2025)
Backdoors in Conditional Diffusion: Threats to Responsible Synthetic Data Pipelines
von: Lapid, Raz, et al.
Veröffentlicht: (2025)
von: Lapid, Raz, et al.
Veröffentlicht: (2025)
PromptBench: A Unified Library for Evaluation of Large Language Models
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
Open Sesame! Universal Black Box Jailbreaking of Large Language Models
von: Lapid, Raz, et al.
Veröffentlicht: (2023)
von: Lapid, Raz, et al.
Veröffentlicht: (2023)
DETAIL Matters: Measuring the Impact of Prompt Specificity on Reasoning in Large Language Models
von: Kim, Olivia
Veröffentlicht: (2025)
von: Kim, Olivia
Veröffentlicht: (2025)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
von: Aly, Walid Mohamed, et al.
Veröffentlicht: (2025)
von: Aly, Walid Mohamed, et al.
Veröffentlicht: (2025)
T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
von: Wang, Qinsi, et al.
Veröffentlicht: (2026)
von: Wang, Qinsi, et al.
Veröffentlicht: (2026)
Revisiting Prompt Sensitivity in Large Language Models for Text Classification: The Role of Prompt Underspecification
von: Pecher, Branislav, et al.
Veröffentlicht: (2026)
von: Pecher, Branislav, et al.
Veröffentlicht: (2026)
Text Adaptation to Plain Language and Easy Read via Automatic Post-Editing Cycles
von: Calleja, Jesús, et al.
Veröffentlicht: (2025)
von: Calleja, Jesús, et al.
Veröffentlicht: (2025)
MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis
von: Zhou, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhou, Yingjie, et al.
Veröffentlicht: (2024)
Breaking Audio Large Language Models by Attacking Only the Encoder: A Universal Targeted Latent-Space Audio Attack
von: Ziv, Roee, et al.
Veröffentlicht: (2025)
von: Ziv, Roee, et al.
Veröffentlicht: (2025)
OR-Bench: An Over-Refusal Benchmark for Large Language Models
von: Cui, Justin, et al.
Veröffentlicht: (2024)
von: Cui, Justin, et al.
Veröffentlicht: (2024)
PromptBridge: Cross-Model Prompt Transfer for Large Language Models
von: Wang, Yaxuan, et al.
Veröffentlicht: (2025)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2025)
Integrating Chemistry Knowledge in Large Language Models via Prompt Engineering
von: Liu, Hongxuan, et al.
Veröffentlicht: (2024)
von: Liu, Hongxuan, et al.
Veröffentlicht: (2024)
A Universal Prompting Strategy for Extracting Process Model Information from Natural Language Text using Large Language Models
von: Neuberger, Julian, et al.
Veröffentlicht: (2024)
von: Neuberger, Julian, et al.
Veröffentlicht: (2024)
Are Large Language Models a Threat to Digital Public Goods? Evidence from Activity on Stack Overflow
von: del Rio-Chanona, Maria, et al.
Veröffentlicht: (2023)
von: del Rio-Chanona, Maria, et al.
Veröffentlicht: (2023)
Prompt Selection Matters: Enhancing Text Annotations for Social Sciences with Large Language Models
von: Abraham, Louis, et al.
Veröffentlicht: (2024)
von: Abraham, Louis, et al.
Veröffentlicht: (2024)
TaskBench: Benchmarking Large Language Models for Task Automation
von: Shen, Yongliang, et al.
Veröffentlicht: (2023)
von: Shen, Yongliang, et al.
Veröffentlicht: (2023)
Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench
von: Liu, Zheyuan, et al.
Veröffentlicht: (2024)
von: Liu, Zheyuan, et al.
Veröffentlicht: (2024)
EmoBench: Evaluating the Emotional Intelligence of Large Language Models
von: Sabour, Sahand, et al.
Veröffentlicht: (2024)
von: Sabour, Sahand, et al.
Veröffentlicht: (2024)
Debiasing Large Language Models via Adaptive Causal Prompting with Sketch-of-Thought
von: Li, Bowen, et al.
Veröffentlicht: (2026)
von: Li, Bowen, et al.
Veröffentlicht: (2026)
Are Large Language Models Good Prompt Optimizers?
von: Ma, Ruotian, et al.
Veröffentlicht: (2024)
von: Ma, Ruotian, et al.
Veröffentlicht: (2024)
Large Language Models Prompting With Episodic Memory
von: Do, Dai, et al.
Veröffentlicht: (2024)
von: Do, Dai, et al.
Veröffentlicht: (2024)
On the Worst Prompt Performance of Large Language Models
von: Cao, Bowen, et al.
Veröffentlicht: (2024)
von: Cao, Bowen, et al.
Veröffentlicht: (2024)
KnowledgePrompts: Exploring the Abilities of Large Language Models to Solve Proportional Analogies via Knowledge-Enhanced Prompting
von: Wijesiriwardene, Thilini, et al.
Veröffentlicht: (2024)
von: Wijesiriwardene, Thilini, et al.
Veröffentlicht: (2024)
Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow Questions
von: Kabir, Samia, et al.
Veröffentlicht: (2023)
von: Kabir, Samia, et al.
Veröffentlicht: (2023)
On the Robustness of Diffusion-Based Image Compression to Bit-Flip Errors
von: Vaisman, Amit, et al.
Veröffentlicht: (2026)
von: Vaisman, Amit, et al.
Veröffentlicht: (2026)
Fortify the Guardian, Not the Treasure: Resilient Adversarial Detectors
von: Lapid, Raz, et al.
Veröffentlicht: (2024)
von: Lapid, Raz, et al.
Veröffentlicht: (2024)
Robustness of Prompting: Enhancing Robustness of Large Language Models Against Prompting Attacks
von: Mu, Lin, et al.
Veröffentlicht: (2025)
von: Mu, Lin, et al.
Veröffentlicht: (2025)
Diverse Prompts: Illuminating the Prompt Space of Large Language Models with MAP-Elites
von: Santos, Gabriel Machado, et al.
Veröffentlicht: (2025)
von: Santos, Gabriel Machado, et al.
Veröffentlicht: (2025)
CROP: Token-Efficient Reasoning in Large Language Models via Regularized Prompt Optimization
von: Shah, Deep, et al.
Veröffentlicht: (2026)
von: Shah, Deep, et al.
Veröffentlicht: (2026)
TurkBench: A Benchmark for Evaluating Turkish Large Language Models
von: Toraman, Çağrı, et al.
Veröffentlicht: (2026)
von: Toraman, Çağrı, et al.
Veröffentlicht: (2026)
BrainBench: Exposing the Commonsense Reasoning Gap in Large Language Models
von: Tang, Yuzhe
Veröffentlicht: (2026)
von: Tang, Yuzhe
Veröffentlicht: (2026)
Medical Reasoning with Large Language Models: A Survey and MR-Bench
von: Ren, Xiaohan, et al.
Veröffentlicht: (2026)
von: Ren, Xiaohan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Activation Steering for Masked Diffusion Language Models
von: Shnaidman, Adi, et al.
Veröffentlicht: (2025) -
SastBench: A Benchmark for Testing Agentic SAST Triage
von: Feiglin, Jake, et al.
Veröffentlicht: (2026) -
Few-shot Name Entity Recognition on StackOverflow
von: Chen, Xinwei, et al.
Veröffentlicht: (2024) -
Can ChatGPT replace StackOverflow? A Study on Robustness and Reliability of Large Language Model Code Generation
von: Zhong, Li, et al.
Veröffentlicht: (2023) -
Evaluating Privacy Questions From Stack Overflow: Can ChatGPT Compete?
von: Delile, Zack, et al.
Veröffentlicht: (2023)