OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Ziheng, Huang, Yixiao, Xu, Hui, Sojoudi, Somayeh, Zhao, Xuandong, Song, Dawn, Mei, Song |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reinforcement Learning for Flow-Matching Policies
by: Pfrommer, Samuel, et al.
Published: (2025)
by: Pfrommer, Samuel, et al.
Published: (2025)
An Undetectable Watermark for Generative Image Models
by: Gunn, Sam, et al.
Published: (2024)
by: Gunn, Sam, et al.
Published: (2024)
Spooky Action at a Distance: Normalization Layers Enable Side-Channel Spatial Communication
by: Pfrommer, Samuel, et al.
Published: (2025)
by: Pfrommer, Samuel, et al.
Published: (2025)
Transformers Provably Learn to Internalize Chain-of-Thought
by: Huang, Yixiao, et al.
Published: (2026)
by: Huang, Yixiao, et al.
Published: (2026)
Multi-Objective Learning for Diffusion Models: A Statistical Theory under Semi-Supervised Learning
by: Cheng, Ziheng, et al.
Published: (2026)
by: Cheng, Ziheng, et al.
Published: (2026)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
by: Kang, Zhewei, et al.
Published: (2025)
by: Kang, Zhewei, et al.
Published: (2025)
Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
by: Huang, Yixiao, et al.
Published: (2025)
by: Huang, Yixiao, et al.
Published: (2025)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
by: Cai, Will, et al.
Published: (2025)
by: Cai, Will, et al.
Published: (2025)
Improving LLM Safety Alignment with Dual-Objective Optimization
by: Zhao, Xuandong, et al.
Published: (2025)
by: Zhao, Xuandong, et al.
Published: (2025)
MemFail: Stress-Testing Failure Modes of LLM Memory Systems
by: Garg, Ishir, et al.
Published: (2026)
by: Garg, Ishir, et al.
Published: (2026)
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
by: Xiong, Alexander, et al.
Published: (2025)
by: Xiong, Alexander, et al.
Published: (2025)
Self-Sovereign Agent
by: Qu, Wenjie, et al.
Published: (2026)
by: Qu, Wenjie, et al.
Published: (2026)
Transport of Algebraic Structure to Latent Embeddings
by: Pfrommer, Samuel, et al.
Published: (2024)
by: Pfrommer, Samuel, et al.
Published: (2024)
Do Sparse Autoencoders Identify Reasoning Features in Language Models?
by: Ma, George, et al.
Published: (2026)
by: Ma, George, et al.
Published: (2026)
Reinforcement Learning via Value Gradient Flow
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
Learning to Reason without External Rewards
by: Zhao, Xuandong, et al.
Published: (2025)
by: Zhao, Xuandong, et al.
Published: (2025)
Pausing Policy Learning in Non-stationary Reinforcement Learning
by: Lee, Hyunin, et al.
Published: (2024)
by: Lee, Hyunin, et al.
Published: (2024)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
by: Bai, Yatong, et al.
Published: (2023)
by: Bai, Yatong, et al.
Published: (2023)
DRAGON: Distributional Rewards Optimize Diffusion Generative Models
by: Bai, Yatong, et al.
Published: (2025)
by: Bai, Yatong, et al.
Published: (2025)
Absence of spurious solutions far from ground truth: A low-rank analysis with high-order losses
by: Ma, Ziye, et al.
Published: (2024)
by: Ma, Ziye, et al.
Published: (2024)
Mixing Classifiers to Alleviate the Accuracy-Robustness Trade-Off
by: Bai, Yatong, et al.
Published: (2023)
by: Bai, Yatong, et al.
Published: (2023)
Towards Optimal Branching of Linear and Semidefinite Relaxations for Neural Network Robustness Certification
by: Anderson, Brendon G., et al.
Published: (2021)
by: Anderson, Brendon G., et al.
Published: (2021)
Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
by: Liang, Kaiqu, et al.
Published: (2025)
by: Liang, Kaiqu, et al.
Published: (2025)
Making Bias Non-Predictive: Training Robust LLM Reasoning via Reinforcement Learning
by: Wang, Qian, et al.
Published: (2026)
by: Wang, Qian, et al.
Published: (2026)
Efficient Global Optimization of Two-Layer ReLU Networks: Quadratic-Time Algorithms and Adversarial Training
by: Bai, Yatong, et al.
Published: (2022)
by: Bai, Yatong, et al.
Published: (2022)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
Infinite-Horizon Reach-Avoid Zero-Sum Games via Deep Reinforcement Learning
by: Li, Jingqi, et al.
Published: (2022)
by: Li, Jingqi, et al.
Published: (2022)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Improving the Accuracy-Robustness Trade-Off of Classifiers via Adaptive Smoothing
by: Bai, Yatong, et al.
Published: (2023)
by: Bai, Yatong, et al.
Published: (2023)
MixedNUTS: Training-Free Accuracy-Robustness Balance via Nonlinearly Mixed Classifiers
by: Bai, Yatong, et al.
Published: (2024)
by: Bai, Yatong, et al.
Published: (2024)
Boosting Alignment for Post-Unlearning Text-to-Image Generative Models
by: Ko, Myeongseob, et al.
Published: (2024)
by: Ko, Myeongseob, et al.
Published: (2024)
Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
Text-space Graph Foundation Models: Comprehensive Benchmarks and New Insights
by: Chen, Zhikai, et al.
Published: (2024)
by: Chen, Zhikai, et al.
Published: (2024)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
by: Jain, Neel, et al.
Published: (2024)
by: Jain, Neel, et al.
Published: (2024)
Benchmarking Zero-Shot Robustness of Multimodal Foundation Models: A Pilot Study
by: Wang, Chenguang, et al.
Published: (2024)
by: Wang, Chenguang, et al.
Published: (2024)
Trustworthy Text-to-Image Diffusion Models: A Timely and Focused Survey
by: Zhang, Yi, et al.
Published: (2024)
by: Zhang, Yi, et al.
Published: (2024)
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
by: Dabas, Mahavir, et al.
Published: (2025)
by: Dabas, Mahavir, et al.
Published: (2025)
Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought
by: Zhao, Jiachen, et al.
Published: (2025)
by: Zhao, Jiachen, et al.
Published: (2025)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
by: An, Bang, et al.
Published: (2024)
by: An, Bang, et al.
Published: (2024)
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
by: Jiang, Eric Hanchen, et al.
Published: (2025)
by: Jiang, Eric Hanchen, et al.
Published: (2025)
Similar Items
-
Reinforcement Learning for Flow-Matching Policies
by: Pfrommer, Samuel, et al.
Published: (2025) -
An Undetectable Watermark for Generative Image Models
by: Gunn, Sam, et al.
Published: (2024) -
Spooky Action at a Distance: Normalization Layers Enable Side-Channel Spatial Communication
by: Pfrommer, Samuel, et al.
Published: (2025) -
Transformers Provably Learn to Internalize Chain-of-Thought
by: Huang, Yixiao, et al.
Published: (2026) -
Multi-Objective Learning for Diffusion Models: A Statistical Theory under Semi-Supervised Learning
by: Cheng, Ziheng, et al.
Published: (2026)