Taking Shortcuts for Categorical VQA Using Super Neurons
Fuente:
arXiv
Saved in:
| Main Authors: | Musacchio, Pierre, Jeong, Jaeyi, Kim, Dahun, Park, Jaesik |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Holistic Order Prediction in Natural Scenes
by: Musacchio, Pierre, et al.
Published: (2025)
by: Musacchio, Pierre, et al.
Published: (2025)
Region-centric Image-Language Pretraining for Open-Vocabulary Detection
by: Kim, Dahun, et al.
Published: (2023)
by: Kim, Dahun, et al.
Published: (2023)
One-Cycle Structured Pruning via Stability-Driven Subnetwork Search
by: Ghimire, Deepak, et al.
Published: (2025)
by: Ghimire, Deepak, et al.
Published: (2025)
VQA-Levels: A Hierarchical Approach for Classifying Questions in VQA
by: Madaka, Madhuri Latha, et al.
Published: (2025)
by: Madaka, Madhuri Latha, et al.
Published: (2025)
LoCO: Low-rank Compositional Rotation Fine-tuning
by: Nguyen, An, et al.
Published: (2026)
by: Nguyen, An, et al.
Published: (2026)
Understanding Distributed Representations of Concepts in Deep Neural Networks without Supervision
by: Chang, Wonjoon, et al.
Published: (2023)
by: Chang, Wonjoon, et al.
Published: (2023)
Tiled Prompts: Overcoming Prompt Misguidance in Image and Video Super-Resolution
by: Kim, Bryan Sangwoo, et al.
Published: (2026)
by: Kim, Bryan Sangwoo, et al.
Published: (2026)
Improving Automatic VQA Evaluation Using Large Language Models
by: Mañas, Oscar, et al.
Published: (2023)
by: Mañas, Oscar, et al.
Published: (2023)
Mitigating Shortcut Learning with Diffusion Counterfactuals and Diverse Ensembles
by: Scimeca, Luca, et al.
Published: (2023)
by: Scimeca, Luca, et al.
Published: (2023)
High-Order Matching for One-Step Shortcut Diffusion Models
by: Chen, Bo, et al.
Published: (2025)
by: Chen, Bo, et al.
Published: (2025)
An Investigation into Pre-Training Object-Centric Representations for Reinforcement Learning
by: Yoon, Jaesik, et al.
Published: (2023)
by: Yoon, Jaesik, et al.
Published: (2023)
Refine and Align: Confidence Calibration through Multi-Agent Interaction in VQA
by: Pandey, Ayush, et al.
Published: (2025)
by: Pandey, Ayush, et al.
Published: (2025)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
by: Min, Juhong, et al.
Published: (2024)
by: Min, Juhong, et al.
Published: (2024)
VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage
by: Alfarano, A., et al.
Published: (2025)
by: Alfarano, A., et al.
Published: (2025)
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering
by: Diao, Xingjian, et al.
Published: (2025)
by: Diao, Xingjian, et al.
Published: (2025)
Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment
by: Kim, Bryan Sangwoo, et al.
Published: (2025)
by: Kim, Bryan Sangwoo, et al.
Published: (2025)
FedCAR: Cross-client Adaptive Re-weighting for Generative Models in Federated Learning
by: Kim, Minjun, et al.
Published: (2024)
by: Kim, Minjun, et al.
Published: (2024)
TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
by: Zhou, Wenhao, et al.
Published: (2025)
by: Zhou, Wenhao, et al.
Published: (2025)
TinyVQA: Compact Multimodal Deep Neural Network for Visual Question Answering on Resource-Constrained Devices
by: Rashid, Hasib-Al, et al.
Published: (2024)
by: Rashid, Hasib-Al, et al.
Published: (2024)
Advancing AI-Powered Medical Image Synthesis: Insights from MedVQA-GI Challenge Using CLIP, Fine-Tuned Stable Diffusion, and Dream-Booth + LoRA
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2025)
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2025)
Generative Classifiers Avoid Shortcut Solutions
by: Li, Alexander C., et al.
Published: (2025)
by: Li, Alexander C., et al.
Published: (2025)
Extend3D: Town-Scale 3D Generation
by: Yoon, Seungwoo, et al.
Published: (2026)
by: Yoon, Seungwoo, et al.
Published: (2026)
PlantVillageVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
by: Sakib, Syed Nazmus, et al.
Published: (2025)
by: Sakib, Syed Nazmus, et al.
Published: (2025)
Guaranteed Optimal Compositional Explanations for Neurons
by: La Rosa, Biagio, et al.
Published: (2025)
by: La Rosa, Biagio, et al.
Published: (2025)
Discovering Influential Neuron Path in Vision Transformers
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
Open Vocabulary Compositional Explanations for Neuron Alignment
by: La Rosa, Biagio, et al.
Published: (2025)
by: La Rosa, Biagio, et al.
Published: (2025)
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
by: Kišš, Martin, et al.
Published: (2025)
by: Kišš, Martin, et al.
Published: (2025)
Reproducibility study on how to find Spurious Correlations, Shortcut Learning, Clever Hans or Group-Distributional non-robustness and how to fix them
by: Delzer, Ole, et al.
Published: (2026)
by: Delzer, Ole, et al.
Published: (2026)
Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations
by: Srivastava, Archita, et al.
Published: (2025)
by: Srivastava, Archita, et al.
Published: (2025)
Think First, Assign Next (ThiFAN-VQA): A Two-stage Chain-of-Thought Framework for Post-Disaster Damage Assessment
by: Karimi, Ehsan, et al.
Published: (2025)
by: Karimi, Ehsan, et al.
Published: (2025)
VideoPDE: Unified Generative PDE Solving via Video Inpainting Diffusion Models
by: Li, Edward, et al.
Published: (2025)
by: Li, Edward, et al.
Published: (2025)
VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking
by: Yang, Kichang, et al.
Published: (2025)
by: Yang, Kichang, et al.
Published: (2025)
MedSapiens: Taking a Pose to Rethink Medical Imaging Landmark Detection
by: Elbatel, Marawan, et al.
Published: (2025)
by: Elbatel, Marawan, et al.
Published: (2025)
LINE: LLM-based Iterative Neuron Explanations for Vision Models
by: Zaigrajew, Vladimir, et al.
Published: (2026)
by: Zaigrajew, Vladimir, et al.
Published: (2026)
Learning Relative Representations for Fine-Grained Multimodal Alignment with Limited Data
by: Kim, Shiwon, et al.
Published: (2026)
by: Kim, Shiwon, et al.
Published: (2026)
Text-Aware Image Restoration with Diffusion Models
by: Min, Jaewon, et al.
Published: (2025)
by: Min, Jaewon, et al.
Published: (2025)
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
by: Dreyer, Maximilian, et al.
Published: (2024)
by: Dreyer, Maximilian, et al.
Published: (2024)
Interaction-Merged Motion Planning: Effectively Leveraging Diverse Motion Datasets for Robust Planning
by: Lee, Giwon, et al.
Published: (2025)
by: Lee, Giwon, et al.
Published: (2025)
Spectral Motion Alignment for Video Motion Transfer using Diffusion Models
by: Park, Geon Yeong, et al.
Published: (2024)
by: Park, Geon Yeong, et al.
Published: (2024)
Stabilizing Consistency Training: A Flow Map Analysis and Self-Distillation
by: Kim, Youngjoong, et al.
Published: (2026)
by: Kim, Youngjoong, et al.
Published: (2026)
Similar Items
-
Holistic Order Prediction in Natural Scenes
by: Musacchio, Pierre, et al.
Published: (2025) -
Region-centric Image-Language Pretraining for Open-Vocabulary Detection
by: Kim, Dahun, et al.
Published: (2023) -
One-Cycle Structured Pruning via Stability-Driven Subnetwork Search
by: Ghimire, Deepak, et al.
Published: (2025) -
VQA-Levels: A Hierarchical Approach for Classifying Questions in VQA
by: Madaka, Madhuri Latha, et al.
Published: (2025) -
LoCO: Low-rank Compositional Rotation Fine-tuning
by: Nguyen, An, et al.
Published: (2026)