SemSegBench & DetecBench: Benchmarking Reliability and Generalization Beyond Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Agnihotri, Shashank, Schader, David, Jakubassa, Jonas, Sharei, Nico, Kral, Simon, Kaçar, Mehmet Ege, Weber, Ruben, Keuper, Margret |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are Synthetic Corruptions A Reliable Proxy For Real-World Corruptions?
by: Agnihotri, Shashank, et al.
Published: (2025)
by: Agnihotri, Shashank, et al.
Published: (2025)
DispBench: Benchmarking Disparity Estimation to Synthetic Corruptions
by: Agnihotri, Shashank, et al.
Published: (2025)
by: Agnihotri, Shashank, et al.
Published: (2025)
A Granular Study of Safety Pretraining under Model Abliteration
by: Agnihotri, Shashank, et al.
Published: (2025)
by: Agnihotri, Shashank, et al.
Published: (2025)
Smart Eyes for Silent Threats: VLMs and In-Context Learning for THz Imaging
by: Poggi, Nicolas, et al.
Published: (2025)
by: Poggi, Nicolas, et al.
Published: (2025)
Faithful, Interpretable Chest X-ray Diagnosis with Anti-Aliased B-cos Networks
by: Kleinmann, Marcel, et al.
Published: (2025)
by: Kleinmann, Marcel, et al.
Published: (2025)
CosPGD: an efficient white-box adversarial attack for pixel-wise prediction tasks
by: Agnihotri, Shashank, et al.
Published: (2023)
by: Agnihotri, Shashank, et al.
Published: (2023)
Improving Feature Stability during Upsampling -- Spectral Artifacts and the Importance of Spatial Context
by: Agnihotri, Shashank, et al.
Published: (2023)
by: Agnihotri, Shashank, et al.
Published: (2023)
Beware of Aliases -- Signal Preservation is Crucial for Robust Image Restoration
by: Agnihotri, Shashank, et al.
Published: (2024)
by: Agnihotri, Shashank, et al.
Published: (2024)
How Do Training Methods Influence the Utilization of Vision Models?
by: Gavrikov, Paul, et al.
Published: (2024)
by: Gavrikov, Paul, et al.
Published: (2024)
Is RobustBench/AutoAttack a suitable Benchmark for Adversarial Robustness?
by: Lorenz, Peter, et al.
Published: (2021)
by: Lorenz, Peter, et al.
Published: (2021)
Images as Tables: In-Context Learning with TabPFN for Low-Data Detection of AI-Generated Images
by: Walter, Jan Philip, et al.
Published: (2026)
by: Walter, Jan Philip, et al.
Published: (2026)
AIM: Amending Inherent Interpretability via Self-Supervised Masking
by: Alshami, Eyad, et al.
Published: (2025)
by: Alshami, Eyad, et al.
Published: (2025)
RAWDet-7: A Multi-Scenario Benchmark for Object Detection and Description on Quantized RAW Images
by: Fatima, Mishal, et al.
Published: (2026)
by: Fatima, Mishal, et al.
Published: (2026)
Vision At Night: Exploring Biologically Inspired Preprocessing For Improved Robustness Via Color And Contrast Transformations
by: Stracke, Lorena, et al.
Published: (2025)
by: Stracke, Lorena, et al.
Published: (2025)
GeoDiv: Framework For Measuring Geographical Diversity In Text-To-Image Models
by: Basu, Abhipsa, et al.
Published: (2026)
by: Basu, Abhipsa, et al.
Published: (2026)
RobustSpring: Benchmarking Robustness to Image Corruptions for Optical Flow, Scene Flow and Stereo
by: Oei, Victor, et al.
Published: (2025)
by: Oei, Victor, et al.
Published: (2025)
Deepfakes: we need to re-think the concept of "real" images
by: Keuper, Janis, et al.
Published: (2025)
by: Keuper, Janis, et al.
Published: (2025)
SemBench: A Benchmark for Semantic Query Processing Engines
by: Lao, Jiale, et al.
Published: (2025)
by: Lao, Jiale, et al.
Published: (2025)
PDSP-Bench: A Benchmarking System for Parallel and Distributed Stream Processing
by: Agnihotri, Pratyush, et al.
Published: (2025)
by: Agnihotri, Pratyush, et al.
Published: (2025)
$γ$-Quant: Towards Learnable Quantization for Low-bit Pattern Recognition
by: Fatima, Mishal, et al.
Published: (2025)
by: Fatima, Mishal, et al.
Published: (2025)
As large as it gets: Learning infinitely large Filters via Neural Implicit Functions in the Fourier Domain
by: Grabinski, Julia, et al.
Published: (2023)
by: Grabinski, Julia, et al.
Published: (2023)
Unfolding Local Growth Rate Estimates for (Almost) Perfect Adversarial Detection
by: Lorenz, Peter, et al.
Published: (2022)
by: Lorenz, Peter, et al.
Published: (2022)
Examining the Impact of Optical Aberrations to Image Classification and Object Detection Models
by: Müller, Patrick, et al.
Published: (2025)
by: Müller, Patrick, et al.
Published: (2025)
Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks
by: Jing, Miao, et al.
Published: (2025)
by: Jing, Miao, et al.
Published: (2025)
IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation
by: Tang, Yinghao, et al.
Published: (2026)
by: Tang, Yinghao, et al.
Published: (2026)
BenchBench: Benchmarking Automated Benchmark Generation
by: Zheng, Yandan, et al.
Published: (2026)
by: Zheng, Yandan, et al.
Published: (2026)
Know Yourself Better: Diverse Object-Related Features Improve Open Set Recognition
by: Xu, Jiawen, et al.
Published: (2024)
by: Xu, Jiawen, et al.
Published: (2024)
BenchSeg: A Large-Scale Dataset and Benchmark for Multi-View Food Video Segmentation
by: AlMughrabi, Ahmad, et al.
Published: (2026)
by: AlMughrabi, Ahmad, et al.
Published: (2026)
FL-MedSegBench: A Comprehensive Benchmark for Federated Learning on Medical Image Segmentation
by: Zhu, Meilu, et al.
Published: (2026)
by: Zhu, Meilu, et al.
Published: (2026)
Fix your downsampling ASAP! Be natively more robust via Aliasing and Spectral Artifact free Pooling
by: Grabinski, Julia, et al.
Published: (2023)
by: Grabinski, Julia, et al.
Published: (2023)
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
by: Fein, Daniel, et al.
Published: (2025)
by: Fein, Daniel, et al.
Published: (2025)
SemBench: A Universal Semantic Framework for LLM Evaluation
by: Zubillaga, Mikel, et al.
Published: (2026)
by: Zubillaga, Mikel, et al.
Published: (2026)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
by: Perlitz, Yotam, et al.
Published: (2024)
by: Perlitz, Yotam, et al.
Published: (2024)
CSR-Bench: A Benchmark for Evaluating the Cross-modal Safety and Reliability of MLLMs
by: Liu, Yuxuan, et al.
Published: (2026)
by: Liu, Yuxuan, et al.
Published: (2026)
AuthorityBench: Benchmarking LLM Authority Perception for Reliable Retrieval-Augmented Generation
by: Yao, Zhihui, et al.
Published: (2026)
by: Yao, Zhihui, et al.
Published: (2026)
SuperBench: A Super-Resolution Benchmark Dataset for Scientific Machine Learning
by: Ren, Pu, et al.
Published: (2023)
by: Ren, Pu, et al.
Published: (2023)
AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
by: Gu, Shangding, et al.
Published: (2025)
by: Gu, Shangding, et al.
Published: (2025)
Detecting AutoAttack Perturbations in the Frequency Domain
by: Lorenz, Peter, et al.
Published: (2021)
by: Lorenz, Peter, et al.
Published: (2021)
CE-Bench: Towards a Reliable Contrastive Evaluation Benchmark of Interpretability of Sparse Autoencoders
by: Gulko, Alex, et al.
Published: (2025)
by: Gulko, Alex, et al.
Published: (2025)
ScoringBench: A Benchmark for Evaluating Tabular Foundation Models with Proper Scoring Rules
by: Landsgesell, Jonas, et al.
Published: (2026)
by: Landsgesell, Jonas, et al.
Published: (2026)
Similar Items
-
Are Synthetic Corruptions A Reliable Proxy For Real-World Corruptions?
by: Agnihotri, Shashank, et al.
Published: (2025) -
DispBench: Benchmarking Disparity Estimation to Synthetic Corruptions
by: Agnihotri, Shashank, et al.
Published: (2025) -
A Granular Study of Safety Pretraining under Model Abliteration
by: Agnihotri, Shashank, et al.
Published: (2025) -
Smart Eyes for Silent Threats: VLMs and In-Context Learning for THz Imaging
by: Poggi, Nicolas, et al.
Published: (2025) -
Faithful, Interpretable Chest X-ray Diagnosis with Anti-Aliased B-cos Networks
by: Kleinmann, Marcel, et al.
Published: (2025)