PRISM: Progressive Reasoning through Iterative Slot Memory for Vision
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Ziyu, Han, Shuangpeng, Zhang, Mengmi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Flow Snapshot Neurons in Action: Deep Neural Networks Generalize to Biological Motion Perception
by: Han, Shuangpeng, et al.
Published: (2024)
by: Han, Shuangpeng, et al.
Published: (2024)
Pose Prior Learner: Unsupervised Categorical Prior Learning for Pose Estimation
by: Wang, Ziyu, et al.
Published: (2024)
by: Wang, Ziyu, et al.
Published: (2024)
Unforgettable Lessons from Forgettable Images: Intra-Class Memorability Matters in Computer Vision
by: Jing, Jie, et al.
Published: (2024)
by: Jing, Jie, et al.
Published: (2024)
Smoothing Slot Attention Iterations and Recurrences
by: Zhao, Rongzhen, et al.
Published: (2025)
by: Zhao, Rongzhen, et al.
Published: (2025)
PRISM: Progressive Rain removal with Integrated State-space Modeling
by: Xue, Pengze, et al.
Published: (2025)
by: Xue, Pengze, et al.
Published: (2025)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
PROGRESSLM: Towards Progress Reasoning in Vision-Language Models
by: Zhang, Jianshu, et al.
Published: (2026)
by: Zhang, Jianshu, et al.
Published: (2026)
Adaptive Visual Scene Understanding: Incremental Scene Graph Generation
by: Khandelwal, Naitik, et al.
Published: (2023)
by: Khandelwal, Naitik, et al.
Published: (2023)
Vision-Language Memory for Spatial Reasoning
by: Liu, Zuntao, et al.
Published: (2025)
by: Liu, Zuntao, et al.
Published: (2025)
Slot-VAE: Object-Centric Scene Generation with Slot Attention
by: Wang, Yanbo, et al.
Published: (2023)
by: Wang, Yanbo, et al.
Published: (2023)
Peering into the Unknown: Active View Selection with Neural Uncertainty Maps for 3D Reconstruction
by: Zhang, Zhengquan, et al.
Published: (2025)
by: Zhang, Zhengquan, et al.
Published: (2025)
Make Me Happier: Evoking Emotions Through Image Diffusion Models
by: Lin, Qing, et al.
Published: (2024)
by: Lin, Qing, et al.
Published: (2024)
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking
by: Zou, Quanchen, et al.
Published: (2025)
by: Zou, Quanchen, et al.
Published: (2025)
PVG: Progressive Vision Graph for Vision Recognition
by: Wu, Jiafu, et al.
Published: (2023)
by: Wu, Jiafu, et al.
Published: (2023)
PRISM: A Promptable and Robust Interactive Segmentation Model with Visual Prompts
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
When Slots Compete: Slot Merging in Object-Centric Learning
by: Chatzisavvas, Christos, et al.
Published: (2026)
by: Chatzisavvas, Christos, et al.
Published: (2026)
Slot-VLM: SlowFast Slots for Video-Language Modeling
by: Xu, Jiaqi, et al.
Published: (2024)
by: Xu, Jiaqi, et al.
Published: (2024)
Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines
by: Cai, Yusen, et al.
Published: (2025)
by: Cai, Yusen, et al.
Published: (2025)
PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion
by: Choi, Jaehyun, et al.
Published: (2025)
by: Choi, Jaehyun, et al.
Published: (2025)
EditThinker: Unlocking Iterative Reasoning for Any Image Editor
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
PRISM: Streaming Human Motion Generation with Per-Joint Latent Decomposition
by: Ling, Zeyu, et al.
Published: (2026)
by: Ling, Zeyu, et al.
Published: (2026)
SlotPi: Physics-informed Object-centric Reasoning Models
by: Li, Jian, et al.
Published: (2025)
by: Li, Jian, et al.
Published: (2025)
Predicting Video Slot Attention Queries from Random Slot-Feature Pairs
by: Zhao, Rongzhen, et al.
Published: (2025)
by: Zhao, Rongzhen, et al.
Published: (2025)
MMVP: A Multimodal MoCap Dataset with Vision and Pressure Sensors
by: Zhang, He, et al.
Published: (2024)
by: Zhang, He, et al.
Published: (2024)
Slot Abstractors: Toward Scalable Abstract Visual Reasoning
by: Mondal, Shanka Subhra, et al.
Published: (2024)
by: Mondal, Shanka Subhra, et al.
Published: (2024)
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
by: You, Haoxuan, et al.
Published: (2023)
by: You, Haoxuan, et al.
Published: (2023)
Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
Slot-guided Volumetric Object Radiance Fields
by: Qi, Di, et al.
Published: (2024)
by: Qi, Di, et al.
Published: (2024)
MSNav: Zero-Shot Vision-and-Language Navigation with Dynamic Memory and LLM Spatial Reasoning
by: Liu, Chenghao, et al.
Published: (2025)
by: Liu, Chenghao, et al.
Published: (2025)
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
by: Fang, Rongyao, et al.
Published: (2025)
by: Fang, Rongyao, et al.
Published: (2025)
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
by: Deng, Yihe, et al.
Published: (2025)
by: Deng, Yihe, et al.
Published: (2025)
Pip-Stereo: Progressive Iterations Pruner for Iterative Optimization based Stereo Matching
by: Zheng, Jintu, et al.
Published: (2026)
by: Zheng, Jintu, et al.
Published: (2026)
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
by: Han, Jiwook, et al.
Published: (2026)
by: Han, Jiwook, et al.
Published: (2026)
Learning to Perceive "Where": Spatial Pretext Tasks for Robust Self-Supervised Learning
by: Shen, Yang, et al.
Published: (2026)
by: Shen, Yang, et al.
Published: (2026)
Unveiling the Tapestry: the Interplay of Generalization and Forgetting in Continual Learning
by: Shi, Zenglin, et al.
Published: (2022)
by: Shi, Zenglin, et al.
Published: (2022)
Adaptive Slot Attention: Object Discovery with Dynamic Slot Number
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration
by: Cai, Zhongyi, et al.
Published: (2025)
by: Cai, Zhongyi, et al.
Published: (2025)
Anatomy-Slot: Unsupervised Anatomical Factorization for Homologous Bilateral Reasoning in Retinal Diagnosis
by: Ma, Yingzhe, et al.
Published: (2026)
by: Ma, Yingzhe, et al.
Published: (2026)
Explainable Image Recognition via Enhanced Slot-attention Based Classifier
by: Wang, Bowen, et al.
Published: (2024)
by: Wang, Bowen, et al.
Published: (2024)
Learning to See the Elephant in the Room: Self-Supervised Context Reasoning in Humans and AI
by: Liu, Xiao, et al.
Published: (2022)
by: Liu, Xiao, et al.
Published: (2022)
Similar Items
-
Flow Snapshot Neurons in Action: Deep Neural Networks Generalize to Biological Motion Perception
by: Han, Shuangpeng, et al.
Published: (2024) -
Pose Prior Learner: Unsupervised Categorical Prior Learning for Pose Estimation
by: Wang, Ziyu, et al.
Published: (2024) -
Unforgettable Lessons from Forgettable Images: Intra-Class Memorability Matters in Computer Vision
by: Jing, Jie, et al.
Published: (2024) -
Smoothing Slot Attention Iterations and Recurrences
by: Zhao, Rongzhen, et al.
Published: (2025) -
PRISM: Progressive Rain removal with Integrated State-space Modeling
by: Xue, Pengze, et al.
Published: (2025)