Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Liyang, Zhu, Muzhi, Zhao, Zhiyue, Zhao, Hengyu, Liu, Ke, Zhong, Linhao, Chen, Hao, Shen, Chunhua |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
por: Zhao, Canyu, et al.
Publicado: (2025)
por: Zhao, Canyu, et al.
Publicado: (2025)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
por: Li, Liyang, et al.
Publicado: (2026)
por: Li, Liyang, et al.
Publicado: (2026)
Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
por: Zhu, Muzhi, et al.
Publicado: (2025)
por: Zhu, Muzhi, et al.
Publicado: (2025)
Generative Active Learning for Long-tailed Instance Segmentation
por: Zhu, Muzhi, et al.
Publicado: (2024)
por: Zhu, Muzhi, et al.
Publicado: (2024)
Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
por: Zhao, Canyu, et al.
Publicado: (2025)
por: Zhao, Canyu, et al.
Publicado: (2025)
Learning Where to Look: Self-supervised Viewpoint Selection for Active Localization using Geometrical Information
por: Di Giammarino, Luca, et al.
Publicado: (2024)
por: Di Giammarino, Luca, et al.
Publicado: (2024)
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
por: Zhong, Hao, et al.
Publicado: (2025)
por: Zhong, Hao, et al.
Publicado: (2025)
Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching
por: Liu, Yang, et al.
Publicado: (2023)
por: Liu, Yang, et al.
Publicado: (2023)
GeoBench: Benchmarking and Analyzing Monocular Geometry Estimation Models
por: Ge, Yongtao, et al.
Publicado: (2024)
por: Ge, Yongtao, et al.
Publicado: (2024)
What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?
por: Xu, Guangkai, et al.
Publicado: (2024)
por: Xu, Guangkai, et al.
Publicado: (2024)
Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
por: Zhu, Muzhi, et al.
Publicado: (2024)
por: Zhu, Muzhi, et al.
Publicado: (2024)
DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative Data
por: Fan, Chengxiang, et al.
Publicado: (2024)
por: Fan, Chengxiang, et al.
Publicado: (2024)
A Simple Image Segmentation Framework via In-Context Examples
por: Liu, Yang, et al.
Publicado: (2024)
por: Liu, Yang, et al.
Publicado: (2024)
Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
por: Luo, Zekai, et al.
Publicado: (2025)
por: Luo, Zekai, et al.
Publicado: (2025)
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
por: Jia, Yiduo, et al.
Publicado: (2026)
por: Jia, Yiduo, et al.
Publicado: (2026)
Bridge Thinking and Acting: Unleashing Physical Potential of VLM with Generalizable Action Expert
por: Liu, Mingyu, et al.
Publicado: (2025)
por: Liu, Mingyu, et al.
Publicado: (2025)
SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories
por: Zhu, Muzhi, et al.
Publicado: (2025)
por: Zhu, Muzhi, et al.
Publicado: (2025)
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
por: Huang, Zheng, et al.
Publicado: (2025)
por: Huang, Zheng, et al.
Publicado: (2025)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
por: Zhao, Jianfei, et al.
Publicado: (2025)
por: Zhao, Jianfei, et al.
Publicado: (2025)
ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe
por: Bai, Yifan, et al.
Publicado: (2023)
por: Bai, Yifan, et al.
Publicado: (2023)
CanViT: Toward Active-Vision Foundation Models
por: Berreby, Yohaï-Eliel, et al.
Publicado: (2026)
por: Berreby, Yohaï-Eliel, et al.
Publicado: (2026)
Unified Open-World Segmentation with Multi-Modal Prompts
por: Liu, Yang, et al.
Publicado: (2025)
por: Liu, Yang, et al.
Publicado: (2025)
Improving Viewpoint-Independent Object-Centric Representations through Active Viewpoint Selection
por: Huang, Yinxuan, et al.
Publicado: (2024)
por: Huang, Yinxuan, et al.
Publicado: (2024)
Exploring Spatial Intelligence from a Generative Perspective
por: Zhu, Muzhi, et al.
Publicado: (2026)
por: Zhu, Muzhi, et al.
Publicado: (2026)
Token Warping Helps MLLMs Look from Nearby Viewpoints
por: Lee, Phillip Y., et al.
Publicado: (2026)
por: Lee, Phillip Y., et al.
Publicado: (2026)
Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet Video
por: Gao, Zihui, et al.
Publicado: (2026)
por: Gao, Zihui, et al.
Publicado: (2026)
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
por: Fuller, Anthony, et al.
Publicado: (2025)
por: Fuller, Anthony, et al.
Publicado: (2025)
Pay Attention to Where You Looked
por: Berian, Alex, et al.
Publicado: (2026)
por: Berian, Alex, et al.
Publicado: (2026)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
por: Chen, Cong, et al.
Publicado: (2025)
por: Chen, Cong, et al.
Publicado: (2025)
Hallucination Begins Where Saliency Drops
por: Zhang, Xiaofeng, et al.
Publicado: (2026)
por: Zhang, Xiaofeng, et al.
Publicado: (2026)
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
por: Shen, Yuxiang, et al.
Publicado: (2026)
por: Shen, Yuxiang, et al.
Publicado: (2026)
Can Foundation Models Revolutionize Mobile AR Sparse Sensing?
por: Zhao, Yiqin, et al.
Publicado: (2025)
por: Zhao, Yiqin, et al.
Publicado: (2025)
Focus on Low-Resolution Information: Multi-Granular Information-Lossless Model for Low-Resolution Human Pose Estimation
por: Gu, Zejun, et al.
Publicado: (2024)
por: Gu, Zejun, et al.
Publicado: (2024)
PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm
por: Zhu, Haoyi, et al.
Publicado: (2023)
por: Zhu, Haoyi, et al.
Publicado: (2023)
Where do Large Vision-Language Models Look at when Answering Questions?
por: Xing, Xiaoying, et al.
Publicado: (2025)
por: Xing, Xiaoying, et al.
Publicado: (2025)
Model Inversion Attacks Through Target-Specific Conditional Diffusion Models
por: Li, Ouxiang, et al.
Publicado: (2024)
por: Li, Ouxiang, et al.
Publicado: (2024)
Towards Viewpoint-Robust End-to-End Autonomous Driving with 3D Foundation Model Priors
por: Hashimoto, Hiroki, et al.
Publicado: (2026)
por: Hashimoto, Hiroki, et al.
Publicado: (2026)
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
por: Liao, Chao, et al.
Publicado: (2025)
por: Liao, Chao, et al.
Publicado: (2025)
Not all Views are Created Equal: Analyzing Viewpoint Instabilities in Vision Foundation Models
por: Michalkiewicz, Mateusz, et al.
Publicado: (2024)
por: Michalkiewicz, Mateusz, et al.
Publicado: (2024)
DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion Models
por: Wu, Weijia, et al.
Publicado: (2023)
por: Wu, Weijia, et al.
Publicado: (2023)
Ejemplares similares
-
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
por: Zhao, Canyu, et al.
Publicado: (2025) -
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
por: Li, Liyang, et al.
Publicado: (2026) -
Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
por: Zhu, Muzhi, et al.
Publicado: (2025) -
Generative Active Learning for Long-tailed Instance Segmentation
por: Zhu, Muzhi, et al.
Publicado: (2024) -
Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
por: Zhao, Canyu, et al.
Publicado: (2025)