Not all Views are Created Equal: Analyzing Viewpoint Instabilities in Vision Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Michalkiewicz, Mateusz, Bai, Sheena, Baktashmotlagh, Mahsa, Jampani, Varun, Balakrishnan, Guha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra
by: Michalkiewicz, Mateusz, et al.
Published: (2025)
by: Michalkiewicz, Mateusz, et al.
Published: (2025)
DRAGON: Drone and Ground Gaussian Splatting for 3D Building Reconstruction
by: Ham, Yujin, et al.
Published: (2024)
by: Ham, Yujin, et al.
Published: (2024)
MvBody: Multi-View-Based Hybrid Transformer Using Optical 3D Body Scan for Explainable Cesarean Section Prediction
by: Cheng, Ruting, et al.
Published: (2025)
by: Cheng, Ruting, et al.
Published: (2025)
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
by: Brusnicki, Roberto, et al.
Published: (2026)
by: Brusnicki, Roberto, et al.
Published: (2026)
Real-time Object and Event Detection Service through Computer Vision and Edge Computing
by: Mendes, Marcos, et al.
Published: (2025)
by: Mendes, Marcos, et al.
Published: (2025)
Bench2FreeAD: A Benchmark for Vision-based End-to-end Navigation in Unstructured Robotic Environments
by: Peng, Yuhang, et al.
Published: (2025)
by: Peng, Yuhang, et al.
Published: (2025)
PAT-VCM: Plug-and-Play Auxiliary Tokens for Video Coding for Machines
by: Jiang, Wei, et al.
Published: (2026)
by: Jiang, Wei, et al.
Published: (2026)
Interactive Image Selection and Training for Brain Tumor Segmentation Network
by: Cerqueira, Matheus A., et al.
Published: (2024)
by: Cerqueira, Matheus A., et al.
Published: (2024)
MotionFollower: Editing Video Motion via Lightweight Score-Guided Diffusion
by: Tu, Shuyuan, et al.
Published: (2024)
by: Tu, Shuyuan, et al.
Published: (2024)
Performance Decay in Deepfake Detection: The Limitations of Training on Outdated Data
by: Richings, Jack, et al.
Published: (2025)
by: Richings, Jack, et al.
Published: (2025)
Depth Priors in Removal Neural Radiance Fields
by: Guo, Zhihao, et al.
Published: (2024)
by: Guo, Zhihao, et al.
Published: (2024)
SealD-NeRF: Interactive Pixel-Level Editing for Dynamic Scenes by Neural Radiance Fields
by: Huang, Zhentao, et al.
Published: (2024)
by: Huang, Zhentao, et al.
Published: (2024)
An inpainting approach to manipulate asymmetry in pre-operative breast images
by: Montenegro, Helena, et al.
Published: (2025)
by: Montenegro, Helena, et al.
Published: (2025)
Less Detail, Better Answers: Degradation-Driven Prompting for VQA
by: Han, Haoxuan, et al.
Published: (2026)
by: Han, Haoxuan, et al.
Published: (2026)
QCFace: Image Quality Control for boosting Face Representation & Recognition
by: Doan-Ngo, Duc-Phuong, et al.
Published: (2025)
by: Doan-Ngo, Duc-Phuong, et al.
Published: (2025)
Gaussian-Constrained LeJEPA Representations for Unsupervised Scene Discovery and Pose Consistency
by: Mostafa, Mohsen
Published: (2026)
by: Mostafa, Mohsen
Published: (2026)
EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges
by: Jon, Hyo Jin, et al.
Published: (2026)
by: Jon, Hyo Jin, et al.
Published: (2026)
TALON: Test-time Adaptive Learning for On-the-Fly Category Discovery
by: Wu, Yanan, et al.
Published: (2026)
by: Wu, Yanan, et al.
Published: (2026)
Large Language Models Powered Context-aware Motion Prediction in Autonomous Driving
by: Zheng, Xiaoji, et al.
Published: (2024)
by: Zheng, Xiaoji, et al.
Published: (2024)
TWIG: Two-Step Image Generation using Segmentation Masks in Diffusion Models
by: Rakib, Mazharul Islam, et al.
Published: (2025)
by: Rakib, Mazharul Islam, et al.
Published: (2025)
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
by: Zeng, Zhitao, et al.
Published: (2025)
by: Zeng, Zhitao, et al.
Published: (2025)
Single-Scanline Relative Pose Estimation for Rolling Shutter Cameras
by: Hruby, Petr, et al.
Published: (2025)
by: Hruby, Petr, et al.
Published: (2025)
Efficient Solution of Point-Line Absolute Pose
by: Hruby, Petr, et al.
Published: (2024)
by: Hruby, Petr, et al.
Published: (2024)
Minimal Solvers for Full DoF Motion Estimation from Asynchronous Tracks
by: Hruby, Petr, et al.
Published: (2025)
by: Hruby, Petr, et al.
Published: (2025)
CylinderPlane: Nested Cylinder Representation for 3D-aware Image Generation
by: Jia, Ru, et al.
Published: (2025)
by: Jia, Ru, et al.
Published: (2025)
Gaze into the Heart: A Multi-View Video Dataset for rPPG and Health Biomarkers Estimation
by: Egorov, Konstantin, et al.
Published: (2025)
by: Egorov, Konstantin, et al.
Published: (2025)
Efficient Vision-based Vehicle Speed Estimation
by: Macko, Andrej, et al.
Published: (2025)
by: Macko, Andrej, et al.
Published: (2025)
Motion Consistency Loss for Monocular Visual Odometry with Attention-Based Deep Learning
by: Françani, André O., et al.
Published: (2024)
by: Françani, André O., et al.
Published: (2024)
Robust Perspective Correction for Real-World Crack Evolution Tracking in Image-Based Structural Health Monitoring
by: Sun, Xinxin, et al.
Published: (2025)
by: Sun, Xinxin, et al.
Published: (2025)
ResPlan: A Large-Scale Vector-Graph Dataset of 17,000 Residential Floor Plans
by: Abouagour, Mohamed, et al.
Published: (2025)
by: Abouagour, Mohamed, et al.
Published: (2025)
Z-Order Transformer for Feed-Forward Gaussian Splatting
by: Wang, Can, et al.
Published: (2026)
by: Wang, Can, et al.
Published: (2026)
EatGAN: An Edge-Attention Guided Generative Adversarial Network for Single Image Super-Resolution
by: Rao, Penghao, et al.
Published: (2025)
by: Rao, Penghao, et al.
Published: (2025)
Isolated Sign Language Recognition with Segmentation and Pose Estimation
by: Perkins, Daniel, et al.
Published: (2025)
by: Perkins, Daniel, et al.
Published: (2025)
Transformer-Based Model for Monocular Visual Odometry: A Video Understanding Approach
by: Françani, André O., et al.
Published: (2023)
by: Françani, André O., et al.
Published: (2023)
Sequence Transferability and Task Order Selection in Continual Learning
by: Nguyen, Thinh, et al.
Published: (2025)
by: Nguyen, Thinh, et al.
Published: (2025)
Knowledge-Guided Failure Prediction: Detecting When Object Detectors Miss Safety-Critical Objects
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
Building Brain Tumor Segmentation Networks with User-Assisted Filter Estimation and Selection
by: Cerqueira, Matheus A., et al.
Published: (2024)
by: Cerqueira, Matheus A., et al.
Published: (2024)
Attire-Based Anomaly Detection in Restricted Areas Using YOLOv8 for Enhanced CCTV Security
by: B, Abdul Aziz A., et al.
Published: (2024)
by: B, Abdul Aziz A., et al.
Published: (2024)
Black-box Adversarial Attacks on CNN-based SLAM Algorithms
by: Gkeka, Maria Rafaela, et al.
Published: (2025)
by: Gkeka, Maria Rafaela, et al.
Published: (2025)
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
by: Tian, Jie, et al.
Published: (2025)
by: Tian, Jie, et al.
Published: (2025)
Similar Items
-
GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra
by: Michalkiewicz, Mateusz, et al.
Published: (2025) -
DRAGON: Drone and Ground Gaussian Splatting for 3D Building Reconstruction
by: Ham, Yujin, et al.
Published: (2024) -
MvBody: Multi-View-Based Hybrid Transformer Using Optical 3D Body Scan for Explainable Cesarean Section Prediction
by: Cheng, Ruting, et al.
Published: (2025) -
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
by: Brusnicki, Roberto, et al.
Published: (2026) -
Real-time Object and Event Detection Service through Computer Vision and Edge Computing
by: Mendes, Marcos, et al.
Published: (2025)