What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization
Fuente:
arXiv
Saved in:
| Main Authors: | Yoshihashi, Ryota, Kada, Masahiro, Ikehata, Satoshi, Kawakami, Rei, Sato, Ikuro |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Teacher-Guided Routing for Sparse Vision Mixture-of-Experts
by: Kada, Masahiro, et al.
Published: (2026)
by: Kada, Masahiro, et al.
Published: (2026)
GUMBEL-NERF: Representing Unseen Objects as Part-Compositional Neural Radiance Fields
by: Sekikawa, Yusuke, et al.
Published: (2024)
by: Sekikawa, Yusuke, et al.
Published: (2024)
PINO: Person-Interaction Noise Optimization for Long-Duration and Customizable Motion Generation of Arbitrary-Sized Groups
by: Ota, Sakuya, et al.
Published: (2025)
by: Ota, Sakuya, et al.
Published: (2025)
Geometry Meets Light: Leveraging Geometric Priors for Universal Photometric Stereo under Limited Multi-Illumination Cues
by: Tam, King-Man, et al.
Published: (2025)
by: Tam, King-Man, et al.
Published: (2025)
VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction
by: Zhang, Jiahao, et al.
Published: (2024)
by: Zhang, Jiahao, et al.
Published: (2024)
Physics-Free Spectrally Multiplexed Photometric Stereo under Unknown Spectral Composition
by: Ikehata, Satoshi, et al.
Published: (2024)
by: Ikehata, Satoshi, et al.
Published: (2024)
Learning Global Object-Centric Representations via Disentangled Slot Attention
by: Chen, Tonglin, et al.
Published: (2024)
by: Chen, Tonglin, et al.
Published: (2024)
Learning Object-Centric Representations Based on Slots in Real World Scenarios
by: Akan, Adil Kaan
Published: (2025)
by: Akan, Adil Kaan
Published: (2025)
When Slots Compete: Slot Merging in Object-Centric Learning
by: Chatzisavvas, Christos, et al.
Published: (2026)
by: Chatzisavvas, Christos, et al.
Published: (2026)
Slot-VAE: Object-Centric Scene Generation with Slot Attention
by: Wang, Yanbo, et al.
Published: (2023)
by: Wang, Yanbo, et al.
Published: (2023)
A Unified Transformer-Based Framework with Pretraining For Whole Body Grasping Motion Generation
by: Effendy, Edward, et al.
Published: (2025)
by: Effendy, Edward, et al.
Published: (2025)
Exploring Limits of Diffusion-Synthetic Training with Weakly Supervised Semantic Segmentation
by: Yoshihashi, Ryota, et al.
Published: (2023)
by: Yoshihashi, Ryota, et al.
Published: (2023)
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
by: An, Zhaoyi, et al.
Published: (2025)
by: An, Zhaoyi, et al.
Published: (2025)
Entity-NeRF: Detecting and Removing Moving Entities in Urban Scenes
by: Otonari, Takashi, et al.
Published: (2024)
by: Otonari, Takashi, et al.
Published: (2024)
MERLiN: Single-Shot Material Estimation and Relighting for Photometric Stereo
by: Tiwari, Ashish, et al.
Published: (2024)
by: Tiwari, Ashish, et al.
Published: (2024)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
by: Chi, Donghwan, et al.
Published: (2025)
by: Chi, Donghwan, et al.
Published: (2025)
SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
by: Grigore, Diana-Nicoleta, et al.
Published: (2025)
by: Grigore, Diana-Nicoleta, et al.
Published: (2025)
Leveraging Multimodal Large Language Models for All-in-One Image Restoration via a Mixture of Frequency Experts
by: Lee, Eunho, et al.
Published: (2026)
by: Lee, Eunho, et al.
Published: (2026)
Constant Rate Scheduling: A General Framework for Optimizing Diffusion Noise Schedule via Distributional Change
by: Okada, Shuntaro, et al.
Published: (2024)
by: Okada, Shuntaro, et al.
Published: (2024)
MetaSlot: Break Through the Fixed Number of Slots in Object-Centric Learning
by: Liu, Hongjia, et al.
Published: (2025)
by: Liu, Hongjia, et al.
Published: (2025)
PerFace: Metric Learning in Perceptual Facial Similarity for Enhanced Face Anonymization
by: Kumagai, Haruka, et al.
Published: (2025)
by: Kumagai, Haruka, et al.
Published: (2025)
PS-EIP: Robust Photometric Stereo Based on Event Interval Profile
by: Kitazawa, Kazuma, et al.
Published: (2025)
by: Kitazawa, Kazuma, et al.
Published: (2025)
Deciphering 'What' and 'Where' Visual Pathways from Spectral Clustering of Layer-Distributed Neural Representations
by: Zhang, Xiao, et al.
Published: (2023)
by: Zhang, Xiao, et al.
Published: (2023)
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
by: Han, Jiwook, et al.
Published: (2026)
by: Han, Jiwook, et al.
Published: (2026)
OpenSlot: Mixed Open-Set Recognition with Object-Centric Learning
by: Yin, Xu, et al.
Published: (2024)
by: Yin, Xu, et al.
Published: (2024)
Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling
by: Tanaka, Daichi, et al.
Published: (2025)
by: Tanaka, Daichi, et al.
Published: (2025)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
QASA: Quality-Guided K-Adaptive Slot Attention for Unsupervised Object-Centric Learning
by: Ouyang, Tianran, et al.
Published: (2026)
by: Ouyang, Tianran, et al.
Published: (2026)
High-Fidelity 3D Tooth Reconstruction by Fusing Intraoral Scans and CBCT Data via a Deep Implicit Representation
by: Zhu, Yi, et al.
Published: (2026)
by: Zhu, Yi, et al.
Published: (2026)
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
by: Fuller, Anthony, et al.
Published: (2025)
by: Fuller, Anthony, et al.
Published: (2025)
Neural Slot Interpreters: Grounding Object Semantics in Emergent Slot Representations
by: Dedhia, Bhishma, et al.
Published: (2024)
by: Dedhia, Bhishma, et al.
Published: (2024)
GLASS: Guided Latent Slot Diffusion for Object-Centric Learning
by: Singh, Krishnakant, et al.
Published: (2024)
by: Singh, Krishnakant, et al.
Published: (2024)
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
by: Huang, Shaofei, et al.
Published: (2025)
by: Huang, Shaofei, et al.
Published: (2025)
ContextFusion and Bootstrap: An Effective Approach to Improve Slot Attention-Based Object-Centric Learning
by: Tian, Pinzhuo, et al.
Published: (2025)
by: Tian, Pinzhuo, et al.
Published: (2025)
Object-Centric Learning with Slot Mixture Module
by: Kirilenko, Daniil, et al.
Published: (2023)
by: Kirilenko, Daniil, et al.
Published: (2023)
Composing Pre-Trained Object-Centric Representations for Robotics From "What" and "Where" Foundation Models
by: Shi, Junyao, et al.
Published: (2024)
by: Shi, Junyao, et al.
Published: (2024)
Learning Disentangled Representation in Object-Centric Models for Visual Dynamics Prediction via Transformers
by: Gandhi, Sanket, et al.
Published: (2024)
by: Gandhi, Sanket, et al.
Published: (2024)
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
by: Duan, Yuxiang, et al.
Published: (2025)
by: Duan, Yuxiang, et al.
Published: (2025)
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
by: Villar-Corrales, Angel, et al.
Published: (2025)
by: Villar-Corrales, Angel, et al.
Published: (2025)
Similar Items
-
Teacher-Guided Routing for Sparse Vision Mixture-of-Experts
by: Kada, Masahiro, et al.
Published: (2026) -
GUMBEL-NERF: Representing Unseen Objects as Part-Compositional Neural Radiance Fields
by: Sekikawa, Yusuke, et al.
Published: (2024) -
PINO: Person-Interaction Noise Optimization for Long-Duration and Customizable Motion Generation of Arbitrary-Sized Groups
by: Ota, Sakuya, et al.
Published: (2025) -
Geometry Meets Light: Leveraging Geometric Priors for Universal Photometric Stereo under Limited Multi-Illumination Cues
by: Tam, King-Man, et al.
Published: (2025) -
VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction
by: Zhang, Jiahao, et al.
Published: (2024)