Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
Fuente:
arXiv
Saved in:
| Main Authors: | Brown, Ellis, Yang, Jihan, Yang, Shusheng, Fergus, Rob, Xie, Saining |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
by: Brown, Ellis, et al.
Published: (2025)
by: Brown, Ellis, et al.
Published: (2025)
V-IRL: Grounding Virtual Intelligence in Real Life
by: Yang, Jihan, et al.
Published: (2024)
by: Yang, Jihan, et al.
Published: (2024)
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
by: Tong, Shengbang, et al.
Published: (2026)
by: Tong, Shengbang, et al.
Published: (2026)
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
by: Yang, Jihan, et al.
Published: (2024)
by: Yang, Jihan, et al.
Published: (2024)
Cambrian-S: Towards Spatial Supersensing in Video
by: Yang, Shusheng, et al.
Published: (2025)
by: Yang, Shusheng, et al.
Published: (2025)
Cambrian-P: Pose-Grounded Video Understanding
by: Yang, Jihan, et al.
Published: (2026)
by: Yang, Jihan, et al.
Published: (2026)
Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
Semi-Supervised Learning with Context-Conditional Generative Adversarial Networks
by: Denton, Remi, et al.
Published: (2016)
by: Denton, Remi, et al.
Published: (2016)
Improved Training Technique for Shortcut Models
by: Nguyen, Anh, et al.
Published: (2025)
by: Nguyen, Anh, et al.
Published: (2025)
Exposing Image Classifier Shortcuts with Counterfactual Frequency (CoF) Tables
by: Hinns, James, et al.
Published: (2024)
by: Hinns, James, et al.
Published: (2024)
Preventing Shortcuts in Adapter Training via Providing the Shortcuts
by: Goyal, Anujraaj Argo, et al.
Published: (2025)
by: Goyal, Anujraaj Argo, et al.
Published: (2025)
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
by: Chen, Jiacheng, et al.
Published: (2025)
by: Chen, Jiacheng, et al.
Published: (2025)
PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop
by: Li, Chenyu, et al.
Published: (2025)
by: Li, Chenyu, et al.
Published: (2025)
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
by: Lee, Dosung, et al.
Published: (2025)
by: Lee, Dosung, et al.
Published: (2025)
Beyond Language Modeling: An Exploration of Multimodal Pretraining
by: Tong, Shengbang, et al.
Published: (2026)
by: Tong, Shengbang, et al.
Published: (2026)
Lightweight Remote Sensing Scene Classification on Edge Devices via Knowledge Distillation and Early-exit
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
Audio-centric Video Understanding Benchmark without Text Shortcut
by: Yang, Yudong, et al.
Published: (2025)
by: Yang, Yudong, et al.
Published: (2025)
Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning
by: Xia, Jiaer, et al.
Published: (2025)
by: Xia, Jiaer, et al.
Published: (2025)
FilmSceneDesigner: Chaining Set Design for Procedural Film Scene Generation
by: Xie, Zhifeng, et al.
Published: (2025)
by: Xie, Zhifeng, et al.
Published: (2025)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
by: Chu, Tianzhe, et al.
Published: (2025)
by: Chu, Tianzhe, et al.
Published: (2025)
Fast Encoding and Decoding for Implicit Video Representation
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
by: Li, Rang, et al.
Published: (2025)
by: Li, Rang, et al.
Published: (2025)
Zero-Shot Open-Vocabulary Human Motion Grounding with Test-Time Training
by: Zhou, Yunjiao, et al.
Published: (2025)
by: Zhou, Yunjiao, et al.
Published: (2025)
SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification
by: Ayoub, Giries Abu, et al.
Published: (2026)
by: Ayoub, Giries Abu, et al.
Published: (2026)
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
by: Tang, Bingda, et al.
Published: (2025)
by: Tang, Bingda, et al.
Published: (2025)
Take A Shortcut Back: Mitigating the Gradient Vanishing for Training Spiking Neural Networks
by: Guo, Yufei, et al.
Published: (2024)
by: Guo, Yufei, et al.
Published: (2024)
The Overlooked Value of Test-time Reference Sets in Visual Place Recognition
by: Zaffar, Mubariz, et al.
Published: (2025)
by: Zaffar, Mubariz, et al.
Published: (2025)
WebSerial Vision Training for Microcontrollers: A Browser-Based Companion to On-Device CNN Training
by: Ellis, Jeremy
Published: (2026)
by: Ellis, Jeremy
Published: (2026)
Benchmarking Dependence Measures to Prevent Shortcut Learning in Medical Imaging
by: Müller, Sarah, et al.
Published: (2024)
by: Müller, Sarah, et al.
Published: (2024)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
by: Xiao, Kelaiti, et al.
Published: (2025)
by: Xiao, Kelaiti, et al.
Published: (2025)
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
by: Chai, Wenhao, et al.
Published: (2024)
by: Chai, Wenhao, et al.
Published: (2024)
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
by: Yu, Sihyun, et al.
Published: (2024)
by: Yu, Sihyun, et al.
Published: (2024)
Exposing Hallucinations To Suppress Them: VLMs Representation Editing With Generative Anchors
by: Shi, Youxu, et al.
Published: (2025)
by: Shi, Youxu, et al.
Published: (2025)
PropTest: Automatic Property Testing for Improved Visual Programming
by: Koo, Jaywon, et al.
Published: (2024)
by: Koo, Jaywon, et al.
Published: (2024)
Scaling Language-Free Visual Representation Learning
by: Fan, David, et al.
Published: (2025)
by: Fan, David, et al.
Published: (2025)
Flow Map Distillation Without Data
by: Tong, Shangyuan, et al.
Published: (2025)
by: Tong, Shangyuan, et al.
Published: (2025)
Diffusion Transformers with Representation Autoencoders
by: Zheng, Boyang, et al.
Published: (2025)
by: Zheng, Boyang, et al.
Published: (2025)
Similar Items
-
SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
by: Brown, Ellis, et al.
Published: (2025) -
V-IRL: Grounding Virtual Intelligence in Real Life
by: Yang, Jihan, et al.
Published: (2024) -
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
by: Tong, Shengbang, et al.
Published: (2026) -
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024) -
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
by: Yang, Jihan, et al.
Published: (2024)