Characterizing the visual representation of objects from the child's view
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Jane, Sepuri, Tarun, Tan, Alvin Wei Ming, Aw, Khai Loong, Frank, Michael C., Long, Bria |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
Unified 3D Scene Understanding Through Physical World Modeling
by: Lee, Wanhee, et al.
Published: (2026)
by: Lee, Wanhee, et al.
Published: (2026)
3D Scene Understanding Through Local Random Access Sequence Modeling
by: Lee, Wanhee, et al.
Published: (2025)
by: Lee, Wanhee, et al.
Published: (2025)
Zero-shot World Models Are Developmentally Efficient Learners
by: Aw, Khai Loong, et al.
Published: (2026)
by: Aw, Khai Loong, et al.
Published: (2026)
Taming generative video models for zero-shot optical flow extraction
by: Kim, Seungwoo, et al.
Published: (2025)
by: Kim, Seungwoo, et al.
Published: (2025)
The BabyView dataset: High-resolution egocentric videos of infants' and young children's everyday experiences
by: Long, Bria, et al.
Published: (2024)
by: Long, Bria, et al.
Published: (2024)
Universal dimensions of visual representation
by: Chen, Zirui, et al.
Published: (2024)
by: Chen, Zirui, et al.
Published: (2024)
LM-Gaussian: Boost Sparse-view 3D Gaussian Splatting with Large Model Priors
by: Yu, Hanyang, et al.
Published: (2024)
by: Yu, Hanyang, et al.
Published: (2024)
Gravity Network for end-to-end small lesion detection
by: Russo, Ciro, et al.
Published: (2023)
by: Russo, Ciro, et al.
Published: (2023)
An extremely coarse feedback signal is sufficient for learning human-aligned visual representations
by: Mehta, Yash, et al.
Published: (2026)
by: Mehta, Yash, et al.
Published: (2026)
Interpreting the structure of multi-object representations in vision encoders
by: Khajuria, Tarun, et al.
Published: (2024)
by: Khajuria, Tarun, et al.
Published: (2024)
NF-SLAM: Effective, Normalizing Flow-supported Neural Field representations for object-level visual SLAM in automotive applications
by: Cui, Li, et al.
Published: (2025)
by: Cui, Li, et al.
Published: (2025)
Learning complete and explainable visual representations from itemized text supervision
by: Lyu, Yiwei, et al.
Published: (2025)
by: Lyu, Yiwei, et al.
Published: (2025)
A transition towards virtual representations of visual scenes
by: Pereira, Américo, et al.
Published: (2024)
by: Pereira, Américo, et al.
Published: (2024)
Adjacent-view Transformers for Supervised Surround-view Depth Estimation
by: Guo, Xianda, et al.
Published: (2023)
by: Guo, Xianda, et al.
Published: (2023)
Evaluating the Generation of Spatial Relations in Text and Image Generative Models
by: Sim, Shang Hong, et al.
Published: (2024)
by: Sim, Shang Hong, et al.
Published: (2024)
Multi-level Reliable Guidance for Unpaired Multi-view Clustering
by: Xin, Like, et al.
Published: (2024)
by: Xin, Like, et al.
Published: (2024)
Unpaired Multi-view Clustering via Reliable View Guidance
by: Xin, Like, et al.
Published: (2024)
by: Xin, Like, et al.
Published: (2024)
Do text-free diffusion models learn discriminative visual representations?
by: Mukhopadhyay, Soumik, et al.
Published: (2023)
by: Mukhopadhyay, Soumik, et al.
Published: (2023)
MV-MR: multi-views and multi-representations for self-supervised learning and knowledge distillation
by: Kinakh, Vitaliy, et al.
Published: (2023)
by: Kinakh, Vitaliy, et al.
Published: (2023)
MVDream: Multi-view Diffusion for 3D Generation
by: Shi, Yichun, et al.
Published: (2023)
by: Shi, Yichun, et al.
Published: (2023)
Generating visual explanations from deep networks using implicit neural representations
by: Byra, Michal, et al.
Published: (2025)
by: Byra, Michal, et al.
Published: (2025)
Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex
by: Thorat, Sushrut, et al.
Published: (2025)
by: Thorat, Sushrut, et al.
Published: (2025)
TriLiteNet: Lightweight Model for Multi-Task Visual Perception
by: Che, Quang-Huy, et al.
Published: (2025)
by: Che, Quang-Huy, et al.
Published: (2025)
Detection Fire in Camera RGB-NIR
by: Khai, Nguyen Truong, et al.
Published: (2025)
by: Khai, Nguyen Truong, et al.
Published: (2025)
Cross-Domain Few-Shot Segmentation via Multi-view Progressive Adaptation
by: Nie, Jiahao, et al.
Published: (2026)
by: Nie, Jiahao, et al.
Published: (2026)
Learning Intra-view and Cross-view Geometric Knowledge for Stereo Matching
by: Gong, Rui, et al.
Published: (2024)
by: Gong, Rui, et al.
Published: (2024)
Optimized View and Geometry Distillation from Multi-view Diffuser
by: Zhang, Youjia, et al.
Published: (2023)
by: Zhang, Youjia, et al.
Published: (2023)
Probabilistic Temporal Masked Attention for Cross-view Online Action Detection
by: Xie, Liping, et al.
Published: (2025)
by: Xie, Liping, et al.
Published: (2025)
SHARE: Single-view Human Adversarial REconstruction
by: Revankar, Shreelekha, et al.
Published: (2023)
by: Revankar, Shreelekha, et al.
Published: (2023)
Percept-Aware Surgical Planning for Visual Cortical Prostheses with Vascular Avoidance
by: Pogoncheff, Galen, et al.
Published: (2026)
by: Pogoncheff, Galen, et al.
Published: (2026)
HENet: Hybrid Encoding for End-to-end Multi-task 3D Perception from Multi-view Cameras
by: Xia, Zhongyu, et al.
Published: (2024)
by: Xia, Zhongyu, et al.
Published: (2024)
PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors
by: Wei, Guangshun, et al.
Published: (2024)
by: Wei, Guangshun, et al.
Published: (2024)
World Modeling with Probabilistic Structure Integration
by: Kotar, Klemen, et al.
Published: (2025)
by: Kotar, Klemen, et al.
Published: (2025)
CREG: Compass Relational Evidence Graph for Characterizing Directional Structure in VLM Spatial-Reasoning Attribution
by: Tan, Kaizhen, et al.
Published: (2026)
by: Tan, Kaizhen, et al.
Published: (2026)
Robust Multi-view Camera Calibration from Dense Matches
by: Hägerlind, Johannes, et al.
Published: (2025)
by: Hägerlind, Johannes, et al.
Published: (2025)
Trunk-branch Contrastive Network with Multi-view Deformable Aggregation for Multi-view Action Recognition
by: Yang, Yingyuan, et al.
Published: (2025)
by: Yang, Yingyuan, et al.
Published: (2025)
Every Pixel Has its Moments: Ultra-High-Resolution Unpaired Image-to-Image Translation via Dense Normalization
by: Ho, Ming-Yang, et al.
Published: (2024)
by: Ho, Ming-Yang, et al.
Published: (2024)
altiro3D: Scene representation from single image and novel view synthesis
by: Canessa, E., et al.
Published: (2023)
by: Canessa, E., et al.
Published: (2023)
Incorporating dense metric depth into neural 3D representations for view synthesis and relighting
by: Chaudhury, Arkadeep Narayan, et al.
Published: (2024)
by: Chaudhury, Arkadeep Narayan, et al.
Published: (2024)
Similar Items
-
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
by: Tan, Alvin Wei Ming, et al.
Published: (2025) -
Unified 3D Scene Understanding Through Physical World Modeling
by: Lee, Wanhee, et al.
Published: (2026) -
3D Scene Understanding Through Local Random Access Sequence Modeling
by: Lee, Wanhee, et al.
Published: (2025) -
Zero-shot World Models Are Developmentally Efficient Learners
by: Aw, Khai Loong, et al.
Published: (2026) -
Taming generative video models for zero-shot optical flow extraction
by: Kim, Seungwoo, et al.
Published: (2025)