Computer Vision and Its Relationship to Cognitive Science: A perspective from Bayes Decision Theory
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuille, Alan, Kersten, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference
von: Wang, Feng, et al.
Veröffentlicht: (2023)
von: Wang, Feng, et al.
Veröffentlicht: (2023)
SPFormer: Enhancing Vision Transformer with Superpixel Representation
von: Mei, Jieru, et al.
Veröffentlicht: (2024)
von: Mei, Jieru, et al.
Veröffentlicht: (2024)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
LychSim: A Controllable and Interactive Simulation Framework for Vision Research
von: Ma, Wufei, et al.
Veröffentlicht: (2026)
von: Ma, Wufei, et al.
Veröffentlicht: (2026)
PaLM2-VAdapter: Progressively Aligned Language Model Makes a Strong Vision-language Adapter
von: Xiao, Junfei, et al.
Veröffentlicht: (2024)
von: Xiao, Junfei, et al.
Veröffentlicht: (2024)
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2025)
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2025)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate
von: Paul, Soumava, et al.
Veröffentlicht: (2026)
von: Paul, Soumava, et al.
Veröffentlicht: (2026)
Gaussian Scenes: Pose-Free Sparse-View Scene Reconstruction using Depth-Enhanced Diffusion Priors
von: Paul, Soumava, et al.
Veröffentlicht: (2024)
von: Paul, Soumava, et al.
Veröffentlicht: (2024)
Quality Sentinel: Estimating Label Quality and Errors in Medical Segmentation Datasets
von: Chen, Yixiong, et al.
Veröffentlicht: (2024)
von: Chen, Yixiong, et al.
Veröffentlicht: (2024)
GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models
von: Guan, Yaohan, et al.
Veröffentlicht: (2026)
von: Guan, Yaohan, et al.
Veröffentlicht: (2026)
ViT-5: Vision Transformers for The Mid-2020s
von: Wang, Feng, et al.
Veröffentlicht: (2026)
von: Wang, Feng, et al.
Veröffentlicht: (2026)
CRAVES: Controlling Robotic Arm with a Vision-based Economic System
von: Zuo, Yiming, et al.
Veröffentlicht: (2018)
von: Zuo, Yiming, et al.
Veröffentlicht: (2018)
Beyond Masks: The Case for Medical Image Parsing
von: Gupta, Siddharth, et al.
Veröffentlicht: (2026)
von: Gupta, Siddharth, et al.
Veröffentlicht: (2026)
A Bayesian Approach to OOD Robustness in Image Classification
von: Kaushik, Prakhar, et al.
Veröffentlicht: (2024)
von: Kaushik, Prakhar, et al.
Veröffentlicht: (2024)
Large-Scale Label Quality Assessment for Medical Segmentation via a Vision-Language Judge and Synthetic Data
von: Chen, Yixiong, et al.
Veröffentlicht: (2026)
von: Chen, Yixiong, et al.
Veröffentlicht: (2026)
RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
von: Yang, Timing, et al.
Veröffentlicht: (2025)
von: Yang, Timing, et al.
Veröffentlicht: (2025)
From Pixels to Objects: A Hierarchical Approach for Part and Object Segmentation Using Local and Global Aggregation
von: Xie, Yunfei, et al.
Veröffentlicht: (2024)
von: Xie, Yunfei, et al.
Veröffentlicht: (2024)
ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
von: Yang, Timing, et al.
Veröffentlicht: (2025)
von: Yang, Timing, et al.
Veröffentlicht: (2025)
Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution
von: Liu, Qihao, et al.
Veröffentlicht: (2024)
von: Liu, Qihao, et al.
Veröffentlicht: (2024)
DINeMo: Learning Neural Mesh Models with no 3D Annotations
von: Guo, Weijie, et al.
Veröffentlicht: (2025)
von: Guo, Weijie, et al.
Veröffentlicht: (2025)
Dictionary-based Framework for Interpretable and Consistent Object Parsing
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2025)
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2025)
ExoViP: Step-by-step Verification and Exploration with Exoskeleton Modules for Compositional Visual Reasoning
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
From Pixel to Cancer: Cellular Automata in Computed Tomography
von: Lai, Yuxiang, et al.
Veröffentlicht: (2024)
von: Lai, Yuxiang, et al.
Veröffentlicht: (2024)
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Generative World Explorer
von: Lu, Taiming, et al.
Veröffentlicht: (2024)
von: Lu, Taiming, et al.
Veröffentlicht: (2024)
CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding
von: Chen, Yixiong, et al.
Veröffentlicht: (2025)
von: Chen, Yixiong, et al.
Veröffentlicht: (2025)
Mamba-R: Vision Mamba ALSO Needs Registers
von: Wang, Feng, et al.
Veröffentlicht: (2024)
von: Wang, Feng, et al.
Veröffentlicht: (2024)
How Well Do Supervised 3D Models Transfer to Medical Imaging Tasks?
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)
Fuzzy Theory in Computer Vision: A Review
von: Yerkin, Adilet, et al.
Veröffentlicht: (2025)
von: Yerkin, Adilet, et al.
Veröffentlicht: (2025)
Efficient Large Multi-modal Models via Visual Context Compression
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
Adventurer: Optimizing Vision Mamba Architecture Designs for Efficiency
von: Wang, Feng, et al.
Veröffentlicht: (2024)
von: Wang, Feng, et al.
Veröffentlicht: (2024)
Prompt-Based Exemplar Super-Compression and Regeneration for Class-Incremental Learning
von: Duan, Ruxiao, et al.
Veröffentlicht: (2023)
von: Duan, Ruxiao, et al.
Veröffentlicht: (2023)
CamFreeDiff: Camera-free Image to Panorama Generation with Diffusion Model
von: Yuan, Xiaoding, et al.
Veröffentlicht: (2024)
von: Yuan, Xiaoding, et al.
Veröffentlicht: (2024)
Name That Part: 3D Part Segmentation and Naming
von: Paul, Soumava, et al.
Veröffentlicht: (2025)
von: Paul, Soumava, et al.
Veröffentlicht: (2025)
DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D Data
von: Liu, Qihao, et al.
Veröffentlicht: (2024)
von: Liu, Qihao, et al.
Veröffentlicht: (2024)
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
von: Li, Zhuowan, et al.
Veröffentlicht: (2022)
von: Li, Zhuowan, et al.
Veröffentlicht: (2022)
VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis
von: Wang, Angtian, et al.
Veröffentlicht: (2022)
von: Wang, Angtian, et al.
Veröffentlicht: (2022)
Are Vision Language Models Ready for Clinical Diagnosis? A 3D Medical Benchmark for Tumor-centric Visual Question Answering
von: Chen, Yixiong, et al.
Veröffentlicht: (2025)
von: Chen, Yixiong, et al.
Veröffentlicht: (2025)
Autoregressive Pretraining with Mamba in Vision
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference
von: Wang, Feng, et al.
Veröffentlicht: (2023) -
SPFormer: Enhancing Vision Transformer with Superpixel Representation
von: Mei, Jieru, et al.
Veröffentlicht: (2024) -
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
von: Chen, Jieneng, et al.
Veröffentlicht: (2024) -
LychSim: A Controllable and Interactive Simulation Framework for Vision Research
von: Ma, Wufei, et al.
Veröffentlicht: (2026) -
PaLM2-VAdapter: Progressively Aligned Language Model Makes a Strong Vision-language Adapter
von: Xiao, Junfei, et al.
Veröffentlicht: (2024)