DAVE: Diagnostic benchmark for Audio Visual Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Radevski, Gorjan, Popordanoska, Teodora, Blaschko, Matthew B., Tuytelaars, Tinne |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Eff-GRot: Efficient and Generalizable Rotation Estimation with Transformers
by: Mathioulakis, Fanis, et al.
Published: (2025)
by: Mathioulakis, Fanis, et al.
Published: (2025)
CLASH: A Benchmark for Cross-Modal Contradiction Detection
by: Popordanoska, Teodora, et al.
Published: (2025)
by: Popordanoska, Teodora, et al.
Published: (2025)
Dice Semimetric Losses: Optimizing the Dice Score with Soft Labels
by: Wang, Zifu, et al.
Published: (2023)
by: Wang, Zifu, et al.
Published: (2023)
The Common Stability Mechanism behind most Self-Supervised Learning Approaches
by: Jha, Abhishek, et al.
Published: (2024)
by: Jha, Abhishek, et al.
Published: (2024)
Diversity-Driven View Subset Selection for Indoor Novel View Synthesis
by: Wang, Zehao, et al.
Published: (2024)
by: Wang, Zehao, et al.
Published: (2024)
Bridging Modalities and Transferring Knowledge: Enhanced Multimodal Understanding and Recognition
by: Radevski, Gorjan
Published: (2025)
by: Radevski, Gorjan
Published: (2025)
Two Complementary Perspectives to Continual Learning: Ask Not Only What to Optimize, But Also How
by: Hess, Timm, et al.
Published: (2023)
by: Hess, Timm, et al.
Published: (2023)
Prediction Error-based Classification for Class-Incremental Learning
by: Zając, Michał, et al.
Published: (2023)
by: Zając, Michał, et al.
Published: (2023)
When normalization hallucinates: unseen risks in AI-powered whole slide image processing
by: Moens, Karel, et al.
Published: (2025)
by: Moens, Karel, et al.
Published: (2025)
Continual Learning of Diffusion Models with Generative Distillation
by: Masip, Sergi, et al.
Published: (2023)
by: Masip, Sergi, et al.
Published: (2023)
Jaccard Metric Losses: Optimizing the Jaccard Index with Soft Labels
by: Wang, Zifu, et al.
Published: (2023)
by: Wang, Zifu, et al.
Published: (2023)
Remembering by Reconstructing: Domain Incremental Learning With Test-Time Training on Video Streams
by: Swinnen, Jonathan, et al.
Published: (2026)
by: Swinnen, Jonathan, et al.
Published: (2026)
Classifying Novel 3D-Printed Objects without Retraining: Towards Post-Production Automation in Additive Manufacturing
by: Mathioulakis, Fanis, et al.
Published: (2026)
by: Mathioulakis, Fanis, et al.
Published: (2026)
Surgeons vs. Computer Vision: A comparative analysis on surgical phase recognition capabilities
by: Mezzina, Marco, et al.
Published: (2025)
by: Mezzina, Marco, et al.
Published: (2025)
Revisiting Reweighted Risk for Calibration: AURC, Focal, and Inverse Focal Loss
by: Zhou, Han, et al.
Published: (2025)
by: Zhou, Han, et al.
Published: (2025)
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
by: Van Landeghem, Jordy, et al.
Published: (2024)
by: Van Landeghem, Jordy, et al.
Published: (2024)
DAVE: Distribution-aware Attribution via ViT Gradient Decomposition
by: Wróbel, Adam, et al.
Published: (2026)
by: Wróbel, Adam, et al.
Published: (2026)
Same accuracy, twice as fast: continuous training surpasses retraining from scratch
by: Verwimp, Eli, et al.
Published: (2025)
by: Verwimp, Eli, et al.
Published: (2025)
Unsupervised Parameter Efficient Source-free Post-pretraining
by: Jha, Abhishek, et al.
Published: (2025)
by: Jha, Abhishek, et al.
Published: (2025)
Video Editing for Audio-Visual Dubbing
by: Manela, Binyamin, et al.
Published: (2025)
by: Manela, Binyamin, et al.
Published: (2025)
RGB-Th-Bench: A Dense benchmark for Visual-Thermal Understanding of Vision Language Models
by: Moshtaghi, Mehdi, et al.
Published: (2025)
by: Moshtaghi, Mehdi, et al.
Published: (2025)
Navigating the Nuances: A Fine-grained Evaluation of Vision-Language Navigation
by: Wang, Zehao, et al.
Published: (2024)
by: Wang, Zehao, et al.
Published: (2024)
X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection
by: Kim, Youngseo, et al.
Published: (2026)
by: Kim, Youngseo, et al.
Published: (2026)
V-LoL: A Diagnostic Dataset for Visual Logical Learning
by: Helff, Lukas, et al.
Published: (2023)
by: Helff, Lukas, et al.
Published: (2023)
CARE: Confidence-aware Ratio Estimation for Medical Biomarkers
by: Li, Jiameng, et al.
Published: (2025)
by: Li, Jiameng, et al.
Published: (2025)
NAB: Neural Adaptive Binning for Sparse-View CT reconstruction
by: Xie, Wangduo, et al.
Published: (2026)
by: Xie, Wangduo, et al.
Published: (2026)
Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding
by: Cao-Dinh, Duc, et al.
Published: (2025)
by: Cao-Dinh, Duc, et al.
Published: (2025)
Adversarial Dependence Minimization
by: De Plaen, Pierre-François, et al.
Published: (2025)
by: De Plaen, Pierre-François, et al.
Published: (2025)
Knowledge Accumulation in Continually Learned Representations and the Issue of Feature Forgetting
by: Hess, Timm, et al.
Published: (2023)
by: Hess, Timm, et al.
Published: (2023)
CrIBo: Self-Supervised Learning via Cross-Image Object-Level Bootstrapping
by: Lebailly, Tim, et al.
Published: (2023)
by: Lebailly, Tim, et al.
Published: (2023)
BenchReAD: A systematic benchmark for retinal anomaly detection
by: Lian, Chenyu, et al.
Published: (2025)
by: Lian, Chenyu, et al.
Published: (2025)
Open-set object detection: towards unified problem formulation and benchmarking
by: Ammar, Hejer, et al.
Published: (2024)
by: Ammar, Hejer, et al.
Published: (2024)
Hearing Touch: Audio-Visual Pretraining for Contact-Rich Manipulation
by: Mejia, Jared, et al.
Published: (2024)
by: Mejia, Jared, et al.
Published: (2024)
Progressive Confident Masking Attention Network for Audio-Visual Segmentation
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
Continual Learning: Applications and the Road Forward
by: Verwimp, Eli, et al.
Published: (2023)
by: Verwimp, Eli, et al.
Published: (2023)
A Structured Benchmark for Text-Guided Anomaly Detection: When Language Stops Conditioning the Decision
by: Samele, Stefano, et al.
Published: (2026)
by: Samele, Stefano, et al.
Published: (2026)
Compositional Steering of Large Language Models with Steering Tokens
by: Radevski, Gorjan, et al.
Published: (2026)
by: Radevski, Gorjan, et al.
Published: (2026)
Synthetic History: Evaluating Visual Representations of the Past in Diffusion Models
by: Palmini, Maria-Teresa De Rosa, et al.
Published: (2025)
by: Palmini, Maria-Teresa De Rosa, et al.
Published: (2025)
Defeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven Optimization
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
Similar Items
-
Eff-GRot: Efficient and Generalizable Rotation Estimation with Transformers
by: Mathioulakis, Fanis, et al.
Published: (2025) -
CLASH: A Benchmark for Cross-Modal Contradiction Detection
by: Popordanoska, Teodora, et al.
Published: (2025) -
Dice Semimetric Losses: Optimizing the Dice Score with Soft Labels
by: Wang, Zifu, et al.
Published: (2023) -
The Common Stability Mechanism behind most Self-Supervised Learning Approaches
by: Jha, Abhishek, et al.
Published: (2024) -
Diversity-Driven View Subset Selection for Indoor Novel View Synthesis
by: Wang, Zehao, et al.
Published: (2024)