Solving Vision Tasks with Simple Photoreceptors Instead of Cameras
Fuente:
arXiv
Saved in:
| Main Authors: | Atanov, Andrei, Fu, Jiawei, Singh, Rishubh, Yu, Isabella, Spielberg, Andrew, Zamir, Amir |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
by: Ramachandran, Rahul, et al.
Published: (2025)
by: Ramachandran, Rahul, et al.
Published: (2025)
Controlled Training Data Generation with Diffusion Models
by: Yeo, Teresa, et al.
Published: (2024)
by: Yeo, Teresa, et al.
Published: (2024)
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
by: Atanov, Andrei, et al.
Published: (2026)
by: Atanov, Andrei, et al.
Published: (2026)
OLAF: A Plug-and-Play Framework for Enhanced Multi-object Multi-part Scene Parsing
by: Gupta, Pranav, et al.
Published: (2024)
by: Gupta, Pranav, et al.
Published: (2024)
Unsupervised Variational Translator for Bridging Image Restoration and High-Level Vision Tasks
by: Wu, Jiawei, et al.
Published: (2024)
by: Wu, Jiawei, et al.
Published: (2024)
Text2Place: Affordance-aware Text Guided Human Placement
by: Parihar, Rishubh, et al.
Published: (2024)
by: Parihar, Rishubh, et al.
Published: (2024)
Compass Control: Multi Object Orientation Control for Text-to-Image Generation
by: Parihar, Rishubh, et al.
Published: (2025)
by: Parihar, Rishubh, et al.
Published: (2025)
Vanilla Group Equivariant Vision Transformer: Simple and Effective
by: Fu, Jiahong, et al.
Published: (2026)
by: Fu, Jiahong, et al.
Published: (2026)
ViPer: Visual Personalization of Generative Models via Individual Preference Learning
by: Salehi, Sogand, et al.
Published: (2024)
by: Salehi, Sogand, et al.
Published: (2024)
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
by: Bachmann, Roman, et al.
Published: (2024)
by: Bachmann, Roman, et al.
Published: (2024)
PreciseControl: Enhancing Text-To-Image Diffusion Models with Fine-Grained Attribute Control
by: Parihar, Rishubh, et al.
Published: (2024)
by: Parihar, Rishubh, et al.
Published: (2024)
MonoPlace3D: Learning 3D-Aware Object Placement for 3D Monocular Detection
by: Parihar, Rishubh, et al.
Published: (2025)
by: Parihar, Rishubh, et al.
Published: (2025)
SimpleEgo: Predicting Probabilistic Body Pose from Egocentric Cameras
by: Cuevas-Velasquez, Hanz, et al.
Published: (2024)
by: Cuevas-Velasquez, Hanz, et al.
Published: (2024)
Computer Vision with a Superpixelation Camera
by: Mahalingam, Sasidharan, et al.
Published: (2026)
by: Mahalingam, Sasidharan, et al.
Published: (2026)
JOCA: Task-Driven Joint Optimisation of Camera Hardware and Adaptive Camera Control Algorithms
by: Yan, Chengyang, et al.
Published: (2025)
by: Yan, Chengyang, et al.
Published: (2025)
Color Spike Data Generation via Bio-inspired Neuron-like Encoding with an Artificial Photoreceptor Layer
by: Ching-Teng, Hsieh, et al.
Published: (2025)
by: Ching-Teng, Hsieh, et al.
Published: (2025)
egoPPG: Heart Rate Estimation from Eye-Tracking Cameras in Egocentric Systems to Benefit Downstream Vision Tasks
by: Braun, Björn, et al.
Published: (2025)
by: Braun, Björn, et al.
Published: (2025)
MC-NeRF: Multi-Camera Neural Radiance Fields for Multi-Camera Image Acquisition Systems
by: Gao, Yu, et al.
Published: (2023)
by: Gao, Yu, et al.
Published: (2023)
CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
Generalist Segmentation Algorithm for Photoreceptors Analysis in Adaptive Optics Imaging
by: Kulyabin, Mikhail, et al.
Published: (2024)
by: Kulyabin, Mikhail, et al.
Published: (2024)
CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback
by: Ge, Wenhang, et al.
Published: (2026)
by: Ge, Wenhang, et al.
Published: (2026)
Paleoinspired Vision: From Exploring Colour Vision Evolution to Inspiring Camera Design
by: Zhang, Junjie, et al.
Published: (2024)
by: Zhang, Junjie, et al.
Published: (2024)
Finding Visual Task Vectors
by: Hojel, Alberto, et al.
Published: (2024)
by: Hojel, Alberto, et al.
Published: (2024)
Reflecting Reality: Enabling Diffusion Models to Produce Faithful Mirror Reflections
by: Dhiman, Ankit, et al.
Published: (2024)
by: Dhiman, Ankit, et al.
Published: (2024)
Balancing Act: Distribution-Guided Debiasing in Diffusion Models
by: Parihar, Rishubh, et al.
Published: (2024)
by: Parihar, Rishubh, et al.
Published: (2024)
Variation of Camera Parameters due to Common Physical Changes in Focal Length and Camera Pose
by: Chen, Hsin-Yi, et al.
Published: (2024)
by: Chen, Hsin-Yi, et al.
Published: (2024)
Height-Guided Projection Reparameterization for Camera-LiDAR Occupancy
by: Wu, Yuan, et al.
Published: (2026)
by: Wu, Yuan, et al.
Published: (2026)
AnyRefill: A Unified, Data-Efficient Framework for Left-Prompt-Guided Vision Tasks
by: Xie, Ming, et al.
Published: (2025)
by: Xie, Ming, et al.
Published: (2025)
Leveraging Perceptual Scores for Dataset Pruning in Computer Vision Tasks
by: Singh, Raghavendra
Published: (2024)
by: Singh, Raghavendra
Published: (2024)
CornerPoint3D: Look at the Nearest Corner Instead of the Center
by: Zhang, Ruixiao, et al.
Published: (2025)
by: Zhang, Ruixiao, et al.
Published: (2025)
FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs
by: Zinonos, Andreas, et al.
Published: (2025)
by: Zinonos, Andreas, et al.
Published: (2025)
A Simple and Effective Point-based Network for Event Camera 6-DOFs Pose Relocalization
by: Ren, Hongwei, et al.
Published: (2024)
by: Ren, Hongwei, et al.
Published: (2024)
The Spatio-Temporal Poisson Point Process: A Simple Model for the Alignment of Event Camera Data
by: Gu, Cheng, et al.
Published: (2021)
by: Gu, Cheng, et al.
Published: (2021)
Automated Segmentation and Analysis of Cone Photoreceptors in Multimodal Adaptive Optics Imaging
by: Shrestha, Prajol, et al.
Published: (2024)
by: Shrestha, Prajol, et al.
Published: (2024)
Vision-Language Models Create Cross-Modal Task Representations
by: Luo, Grace, et al.
Published: (2024)
by: Luo, Grace, et al.
Published: (2024)
Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
by: Sarkar, Ayushman, et al.
Published: (2025)
by: Sarkar, Ayushman, et al.
Published: (2025)
TaCOS: Task-Specific Camera Optimization with Simulation
by: Yan, Chengyang, et al.
Published: (2024)
by: Yan, Chengyang, et al.
Published: (2024)
APLA: A Simple Adaptation Method for Vision Transformers
by: Sorkhei, Moein, et al.
Published: (2025)
by: Sorkhei, Moein, et al.
Published: (2025)
Solving Masked Jigsaw Puzzles with Diffusion Vision Transformers
by: Liu, Jinyang, et al.
Published: (2024)
by: Liu, Jinyang, et al.
Published: (2024)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Similar Items
-
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
by: Ramachandran, Rahul, et al.
Published: (2025) -
Controlled Training Data Generation with Diffusion Models
by: Yeo, Teresa, et al.
Published: (2024) -
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
by: Atanov, Andrei, et al.
Published: (2026) -
OLAF: A Plug-and-Play Framework for Enhanced Multi-object Multi-part Scene Parsing
by: Gupta, Pranav, et al.
Published: (2024) -
Unsupervised Variational Translator for Bridging Image Restoration and High-Level Vision Tasks
by: Wu, Jiawei, et al.
Published: (2024)