Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lee, Seokmin, Lee, Yunghee, Pak, Byeonghyun, Woo, Byeongju |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Rethinking FID Through the Geometry of the Reference Dataset
par: Lee, Yunghee, et autres
Publié: (2026)
par: Lee, Yunghee, et autres
Publié: (2026)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
par: Woo, Byeongju, et autres
Publié: (2026)
par: Woo, Byeongju, et autres
Publié: (2026)
Tortoise and Hare Guidance: Accelerating Diffusion Model Inference with Multirate Integration
par: Lee, Yunghee, et autres
Publié: (2025)
par: Lee, Yunghee, et autres
Publié: (2025)
Pixels-to-Graph: Real-time Integration of Building Information Models and Scene Graphs for Semantic-Geometric Human-Robot Understanding
par: Longo, Antonello, et autres
Publié: (2025)
par: Longo, Antonello, et autres
Publié: (2025)
Textual Query-Driven Mask Transformer for Domain Generalized Segmentation
par: Pak, Byeonghyun, et autres
Publié: (2024)
par: Pak, Byeonghyun, et autres
Publié: (2024)
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
par: Li, Haoyuan, et autres
Publié: (2025)
par: Li, Haoyuan, et autres
Publié: (2025)
GraspClutter6D: A Large-scale Real-world Dataset for Robust Perception and Grasping in Cluttered Scenes
par: Back, Seunghyeok, et autres
Publié: (2025)
par: Back, Seunghyeok, et autres
Publié: (2025)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
par: Kim, Dongwon, et autres
Publié: (2026)
par: Kim, Dongwon, et autres
Publié: (2026)
Memorize What Matters: Emergent Scene Decomposition from Multitraverse
par: Li, Yiming, et autres
Publié: (2024)
par: Li, Yiming, et autres
Publié: (2024)
AiSDF: Structure-aware Neural Signed Distance Fields in Indoor Scenes
par: Jang, Jaehoon, et autres
Publié: (2024)
par: Jang, Jaehoon, et autres
Publié: (2024)
Explainable Scene Understanding with Qualitative Representations and Graph Neural Networks
par: Belmecheri, Nassim, et autres
Publié: (2025)
par: Belmecheri, Nassim, et autres
Publié: (2025)
RVN-Bench: A Benchmark for Reactive Visual Navigation
par: Lee, Jaewon, et autres
Publié: (2026)
par: Lee, Jaewon, et autres
Publié: (2026)
Enhancing Vision-Language Models with Scene Graphs for Traffic Accident Understanding
par: Lohner, Aaron, et autres
Publié: (2024)
par: Lohner, Aaron, et autres
Publié: (2024)
Scene Graph-Guided Proactive Replanning for Failure-Resilient Embodied Agent
par: Yu, Che Rin, et autres
Publié: (2025)
par: Yu, Che Rin, et autres
Publié: (2025)
BYE: Build Your Encoder with One Sequence of Exploration Data for Long-Term Dynamic Scene Understanding
par: Huang, Chenguang, et autres
Publié: (2024)
par: Huang, Chenguang, et autres
Publié: (2024)
On Deep Learning for Geometric and Semantic Scene Understanding Using On-Vehicle 3D LiDAR
par: Li, Li
Publié: (2024)
par: Li, Li
Publié: (2024)
Faster or Stronger: Towards Flexible Visual Place Recognition via Weighted Aggregation and Token Pruning
par: Zeng, Zichao, et autres
Publié: (2026)
par: Zeng, Zichao, et autres
Publié: (2026)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
par: Hao, Haihong, et autres
Publié: (2026)
par: Hao, Haihong, et autres
Publié: (2026)
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
par: Zhang, Yi, et autres
Publié: (2025)
par: Zhang, Yi, et autres
Publié: (2025)
What Matters to You? Towards Visual Representation Alignment for Robot Learning
par: Tian, Ran, et autres
Publié: (2023)
par: Tian, Ran, et autres
Publié: (2023)
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?
par: Majumdar, Arjun, et autres
Publié: (2023)
par: Majumdar, Arjun, et autres
Publié: (2023)
Gaze on the Prize: Shaping Visual Attention with Return-Guided Contrastive Learning
par: Lee, Andrew, et autres
Publié: (2025)
par: Lee, Andrew, et autres
Publié: (2025)
CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning
par: Gan, Rui, et autres
Publié: (2026)
par: Gan, Rui, et autres
Publié: (2026)
TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
par: Liu, Chenghao, et autres
Publié: (2025)
par: Liu, Chenghao, et autres
Publié: (2025)
SpatialAnt: Autonomous Zero-Shot Robot Navigation via Active Scene Reconstruction and Visual Anticipation
par: Zhang, Jiwen, et autres
Publié: (2026)
par: Zhang, Jiwen, et autres
Publié: (2026)
Composing Pre-Trained Object-Centric Representations for Robotics From "What" and "Where" Foundation Models
par: Shi, Junyao, et autres
Publié: (2024)
par: Shi, Junyao, et autres
Publié: (2024)
NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models
par: Park, Sung-Yeon, et autres
Publié: (2025)
par: Park, Sung-Yeon, et autres
Publié: (2025)
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
par: Wang, Yi Ru, et autres
Publié: (2025)
par: Wang, Yi Ru, et autres
Publié: (2025)
Subtask-Aware Visual Reward Learning from Segmented Demonstrations
par: Kim, Changyeon, et autres
Publié: (2025)
par: Kim, Changyeon, et autres
Publié: (2025)
PhysInOne: Visual Physics Learning and Reasoning in One Suite
par: Zhou, Siyuan, et autres
Publié: (2026)
par: Zhou, Siyuan, et autres
Publié: (2026)
Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
par: Han, Yu, et autres
Publié: (2025)
par: Han, Yu, et autres
Publié: (2025)
FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene Interaction
par: Rotondi, Dennis, et autres
Publié: (2025)
par: Rotondi, Dennis, et autres
Publié: (2025)
Efficient Perception, Planning, and Control Algorithm for Vision-Based Automated Vehicles
par: Lee, Der-Hau
Publié: (2022)
par: Lee, Der-Hau
Publié: (2022)
Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied Agents
par: Choi, Wonje, et autres
Publié: (2024)
par: Choi, Wonje, et autres
Publié: (2024)
PhysHanDI: Physics-Based Reconstruction of Hand-Deformable Object Interactions
par: Lee, Jihyun, et autres
Publié: (2026)
par: Lee, Jihyun, et autres
Publié: (2026)
mmWave Radar-Based Non-Line-of-Sight Pedestrian Localization at T-Junctions Utilizing Road Layout Extraction via Camera
par: Park, Byeonggyu, et autres
Publié: (2025)
par: Park, Byeonggyu, et autres
Publié: (2025)
Evaluating Compositional Scene Understanding in Multimodal Generative Models
par: Fu, Shuhao, et autres
Publié: (2025)
par: Fu, Shuhao, et autres
Publié: (2025)
Compositional Generative Modeling: A Single Model is Not All You Need
par: Du, Yilun, et autres
Publié: (2024)
par: Du, Yilun, et autres
Publié: (2024)
Planning with the Views via Scene Self-Exploration
par: Wang, Kangrui, et autres
Publié: (2026)
par: Wang, Kangrui, et autres
Publié: (2026)
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
par: Bendikas, Rokas, et autres
Publié: (2025)
par: Bendikas, Rokas, et autres
Publié: (2025)
Documents similaires
-
Rethinking FID Through the Geometry of the Reference Dataset
par: Lee, Yunghee, et autres
Publié: (2026) -
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
par: Woo, Byeongju, et autres
Publié: (2026) -
Tortoise and Hare Guidance: Accelerating Diffusion Model Inference with Multirate Integration
par: Lee, Yunghee, et autres
Publié: (2025) -
Pixels-to-Graph: Real-time Integration of Building Information Models and Scene Graphs for Semantic-Geometric Human-Robot Understanding
par: Longo, Antonello, et autres
Publié: (2025) -
Textual Query-Driven Mask Transformer for Domain Generalized Segmentation
par: Pak, Byeonghyun, et autres
Publié: (2024)