ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhong, Ziqian, Raghunathan, Aditi, Carlini, Nicholas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
por: Zhong, Ziqian, et al.
Publicado: (2025)
por: Zhong, Ziqian, et al.
Publicado: (2025)
Self-Trained Verification for Training- and Test-Time Self-Improvement
por: Wu, Chen Henry, et al.
Publicado: (2026)
por: Wu, Chen Henry, et al.
Publicado: (2026)
Base Models Look Human To AI Detectors
por: Xu, Yixuan Even, et al.
Publicado: (2026)
por: Xu, Yixuan Even, et al.
Publicado: (2026)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
por: Kim, Taeyoun, et al.
Publicado: (2024)
por: Kim, Taeyoun, et al.
Publicado: (2024)
Mode-Conditioning Unlocks Superior Test-Time Scaling
por: Wu, Chen Henry, et al.
Publicado: (2025)
por: Wu, Chen Henry, et al.
Publicado: (2025)
Understanding Finetuning for Factual Knowledge Extraction
por: Ghosal, Gaurav, et al.
Publicado: (2024)
por: Ghosal, Gaurav, et al.
Publicado: (2024)
Understanding Catastrophic Forgetting in Language Models via Implicit Inference
por: Kotha, Suhas, et al.
Publicado: (2023)
por: Kotha, Suhas, et al.
Publicado: (2023)
Jailbreaking in the Haystack
por: Shah, Rishi Rajesh, et al.
Publicado: (2025)
por: Shah, Rishi Rajesh, et al.
Publicado: (2025)
Mitigating Bias in RAG: Controlling the Embedder
por: Kim, Taeyoun, et al.
Publicado: (2025)
por: Kim, Taeyoun, et al.
Publicado: (2025)
The Impossibility of Fair LLMs
por: Anthis, Jacy, et al.
Publicado: (2024)
por: Anthis, Jacy, et al.
Publicado: (2024)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
por: Goyal, Sachin, et al.
Publicado: (2024)
por: Goyal, Sachin, et al.
Publicado: (2024)
Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
por: Watts, Ishaan, et al.
Publicado: (2026)
por: Watts, Ishaan, et al.
Publicado: (2026)
Repetition Improves Language Model Embeddings
por: Springer, Jacob Mitchell, et al.
Publicado: (2024)
por: Springer, Jacob Mitchell, et al.
Publicado: (2024)
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
por: Aerni, Michael, et al.
Publicado: (2024)
por: Aerni, Michael, et al.
Publicado: (2024)
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
por: Su, Jingtong, et al.
Publicado: (2024)
por: Su, Jingtong, et al.
Publicado: (2024)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
por: Nagarajan, Vaishnavh, et al.
Publicado: (2025)
por: Nagarajan, Vaishnavh, et al.
Publicado: (2025)
Algorithmic Capabilities of Random Transformers
por: Zhong, Ziqian, et al.
Publicado: (2024)
por: Zhong, Ziqian, et al.
Publicado: (2024)
Forcing Diffuse Distributions out of Language Models
por: Zhang, Yiming, et al.
Publicado: (2024)
por: Zhang, Yiming, et al.
Publicado: (2024)
Causally Testing Gender Bias in LLMs: A Case Study on Occupational Bias
por: Chen, Yuen, et al.
Publicado: (2022)
por: Chen, Yuen, et al.
Publicado: (2022)
OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
por: Aggarwal, Pranjal, et al.
Publicado: (2025)
por: Aggarwal, Pranjal, et al.
Publicado: (2025)
TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
por: Ahmed, Toufique, et al.
Publicado: (2024)
por: Ahmed, Toufique, et al.
Publicado: (2024)
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
por: Maini, Pratyush, et al.
Publicado: (2023)
por: Maini, Pratyush, et al.
Publicado: (2023)
Mission: Impossible Language Models
por: Kallini, Julie, et al.
Publicado: (2024)
por: Kallini, Julie, et al.
Publicado: (2024)
Exploiting LLMs for Automatic Hypothesis Assessment via a Logit-Based Calibrated Prior
por: Gong, Yue, et al.
Publicado: (2025)
por: Gong, Yue, et al.
Publicado: (2025)
Scaling Laws for Precision
por: Kumar, Tanishq, et al.
Publicado: (2024)
por: Kumar, Tanishq, et al.
Publicado: (2024)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
por: Farashah, Alireza Dehghanpour, et al.
Publicado: (2026)
por: Farashah, Alireza Dehghanpour, et al.
Publicado: (2026)
The Impossibility Triangle of Long-Context Modeling
por: Zhou, Yan
Publicado: (2026)
por: Zhou, Yan
Publicado: (2026)
Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
por: Yang, Xikang, et al.
Publicado: (2025)
por: Yang, Xikang, et al.
Publicado: (2025)
Non-Halting Queries: Exploiting Fixed Points in LLMs
por: Hammouri, Ghaith, et al.
Publicado: (2024)
por: Hammouri, Ghaith, et al.
Publicado: (2024)
XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
por: Zhang, Zhihan, et al.
Publicado: (2025)
por: Zhang, Zhihan, et al.
Publicado: (2025)
From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench
por: Shi, Weikang, et al.
Publicado: (2026)
por: Shi, Weikang, et al.
Publicado: (2026)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
por: Zhou, Yujun, et al.
Publicado: (2024)
por: Zhou, Yujun, et al.
Publicado: (2024)
AgentBench: Evaluating LLMs as Agents
por: Liu, Xiao, et al.
Publicado: (2023)
por: Liu, Xiao, et al.
Publicado: (2023)
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
por: Mendu, Sai Krishna, et al.
Publicado: (2025)
por: Mendu, Sai Krishna, et al.
Publicado: (2025)
Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?
por: Zhong, Ziqian, et al.
Publicado: (2026)
por: Zhong, Ziqian, et al.
Publicado: (2026)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
por: Li, Xin, et al.
Publicado: (2025)
por: Li, Xin, et al.
Publicado: (2025)
Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models
por: Zhang, Hanlin, et al.
Publicado: (2023)
por: Zhang, Hanlin, et al.
Publicado: (2023)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
por: Lin, Zicheng, et al.
Publicado: (2024)
por: Lin, Zicheng, et al.
Publicado: (2024)
Lost in Localization: Building RabakBench with Human-in-the-Loop Validation to Measure Multilingual Safety Gaps
por: Chua, Gabriel, et al.
Publicado: (2025)
por: Chua, Gabriel, et al.
Publicado: (2025)
Early Data Exposure Improves Robustness to Subsequent Fine-Tuning
por: Feng, Lawrence, et al.
Publicado: (2026)
por: Feng, Lawrence, et al.
Publicado: (2026)
Ejemplares similares
-
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
por: Zhong, Ziqian, et al.
Publicado: (2025) -
Self-Trained Verification for Training- and Test-Time Self-Improvement
por: Wu, Chen Henry, et al.
Publicado: (2026) -
Base Models Look Human To AI Detectors
por: Xu, Yixuan Even, et al.
Publicado: (2026) -
Testing the Limits of Jailbreaking Defenses with the Purple Problem
por: Kim, Taeyoun, et al.
Publicado: (2024) -
Mode-Conditioning Unlocks Superior Test-Time Scaling
por: Wu, Chen Henry, et al.
Publicado: (2025)