LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency
Fuente:
arXiv
Salvato in:
| Autori principali: | Guo, Zhongbin, Liu, Jiahe, Gao, Wenyu, Li, Yushan, Li, Chengzhi, Jian, Ping |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond Flatlands: Unlocking Spatial Intelligence by Decoupling 3D Reasoning from Numerical Regression
di: Guo, Zhongbin, et al.
Pubblicazione: (2025)
di: Guo, Zhongbin, et al.
Pubblicazione: (2025)
Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
di: Guo, Zhongbin, et al.
Pubblicazione: (2026)
di: Guo, Zhongbin, et al.
Pubblicazione: (2026)
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
di: Li, Chengzhi, et al.
Pubblicazione: (2025)
di: Li, Chengzhi, et al.
Pubblicazione: (2025)
Trace3D: Consistent Segmentation Lifting via Gaussian Instance Tracing
di: Shen, Hongyu, et al.
Pubblicazione: (2025)
di: Shen, Hongyu, et al.
Pubblicazione: (2025)
ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Diffusion Models
di: Zhu, Ruishu, et al.
Pubblicazione: (2025)
di: Zhu, Ruishu, et al.
Pubblicazione: (2025)
Consistent-1-to-3: Consistent Image to 3D View Synthesis via Geometry-aware Diffusion Models
di: Ye, Jianglong, et al.
Pubblicazione: (2023)
di: Ye, Jianglong, et al.
Pubblicazione: (2023)
Lifting Motion to the 3D World via 2D Diffusion
di: Li, Jiaman, et al.
Pubblicazione: (2024)
di: Li, Jiaman, et al.
Pubblicazione: (2024)
DisCo3D: Distilling Multi-View Consistency for 3D Scene Editing
di: Chi, Yufeng, et al.
Pubblicazione: (2025)
di: Chi, Yufeng, et al.
Pubblicazione: (2025)
TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with Temporal-Aware Multimodal Models
di: Guo, Zhongbin, et al.
Pubblicazione: (2025)
di: Guo, Zhongbin, et al.
Pubblicazione: (2025)
View-Consistent Diffusion Representations for 3D-Consistent Video Generation
di: Danier, Duolikun, et al.
Pubblicazione: (2025)
di: Danier, Duolikun, et al.
Pubblicazione: (2025)
Enforcing View-Consistency in Class-Agnostic 3D Segmentation Fields
di: Dumery, Corentin, et al.
Pubblicazione: (2024)
di: Dumery, Corentin, et al.
Pubblicazione: (2024)
SegVGGT: Joint 3D Reconstruction and Instance Segmentation from Multi-View Images
di: Qu, Jinyuan, et al.
Pubblicazione: (2026)
di: Qu, Jinyuan, et al.
Pubblicazione: (2026)
Sculpt3D: Multi-View Consistent Text-to-3D Generation with Sparse 3D Prior
di: Chen, Cheng, et al.
Pubblicazione: (2024)
di: Chen, Cheng, et al.
Pubblicazione: (2024)
Dehallu3D: Hallucination-Mitigated 3D Generation from Single Image via Cyclic View Consistency Refinement
di: Wang, Xiwen, et al.
Pubblicazione: (2026)
di: Wang, Xiwen, et al.
Pubblicazione: (2026)
LISA: Reasoning Segmentation via Large Language Model
di: Lai, Xin, et al.
Pubblicazione: (2023)
di: Lai, Xin, et al.
Pubblicazione: (2023)
TrackRef3D: Multi-View Consistent Track-then-Label for Open-World Referring Segmentation in 3D Gaussian Splatting
di: Tan, Yuyang, et al.
Pubblicazione: (2026)
di: Tan, Yuyang, et al.
Pubblicazione: (2026)
3DEnhancer: Consistent Multi-View Diffusion for 3D Enhancement
di: Luo, Yihang, et al.
Pubblicazione: (2024)
di: Luo, Yihang, et al.
Pubblicazione: (2024)
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
di: Zhang, Xuying, et al.
Pubblicazione: (2025)
di: Zhang, Xuying, et al.
Pubblicazione: (2025)
MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation
di: Huang, Jiaxin, et al.
Pubblicazione: (2025)
di: Huang, Jiaxin, et al.
Pubblicazione: (2025)
Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
di: Li, Wenyu, et al.
Pubblicazione: (2025)
di: Li, Wenyu, et al.
Pubblicazione: (2025)
ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency
di: Feng, Haitang, et al.
Pubblicazione: (2025)
di: Feng, Haitang, et al.
Pubblicazione: (2025)
Segment, Lift and Fit: Automatic 3D Shape Labeling from 2D Prompts
di: Li, Jianhao, et al.
Pubblicazione: (2024)
di: Li, Jianhao, et al.
Pubblicazione: (2024)
Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images
di: Sun, Xiangyu, et al.
Pubblicazione: (2025)
di: Sun, Xiangyu, et al.
Pubblicazione: (2025)
MANGO:Natural Multi-speaker 3D Talking Head Generation via 2D-Lifted Enhancement
di: Zhu, Lei, et al.
Pubblicazione: (2026)
di: Zhu, Lei, et al.
Pubblicazione: (2026)
AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D Generation
di: Liang, Xinyue, et al.
Pubblicazione: (2025)
di: Liang, Xinyue, et al.
Pubblicazione: (2025)
LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors
di: Chen, Yabo, et al.
Pubblicazione: (2024)
di: Chen, Yabo, et al.
Pubblicazione: (2024)
Consistent View Alignment Improves Foundation Models for 3D Medical Image Segmentation
di: Vaish, Puru, et al.
Pubblicazione: (2025)
di: Vaish, Puru, et al.
Pubblicazione: (2025)
RUMPL: Ray-Based Transformers for Universal Multi-View 2D to 3D Human Pose Lifting
di: Ghasemzadeh, Seyed Abolfazl, et al.
Pubblicazione: (2025)
di: Ghasemzadeh, Seyed Abolfazl, et al.
Pubblicazione: (2025)
CoR-GS: Sparse-View 3D Gaussian Splatting via Co-Regularization
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
2L3: Lifting Imperfect Generated 2D Images into Accurate 3D
di: Chen, Yizheng, et al.
Pubblicazione: (2024)
di: Chen, Yizheng, et al.
Pubblicazione: (2024)
Synthesizing Consistent Novel Views via 3D Epipolar Attention without Re-Training
di: Ye, Botao, et al.
Pubblicazione: (2025)
di: Ye, Botao, et al.
Pubblicazione: (2025)
ViewDiff: 3D-Consistent Image Generation with Text-to-Image Models
di: Höllein, Lukas, et al.
Pubblicazione: (2024)
di: Höllein, Lukas, et al.
Pubblicazione: (2024)
WonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene Exploration
di: Ni, Chaojun, et al.
Pubblicazione: (2025)
di: Ni, Chaojun, et al.
Pubblicazione: (2025)
3D Scene Change Modeling With Consistent Multi-View Aggregation
di: Zhou, Zirui, et al.
Pubblicazione: (2025)
di: Zhou, Zirui, et al.
Pubblicazione: (2025)
UniC-Lift: Unified 3D Instance Segmentation via Contrastive Learning
di: Dhiman, Ankit, et al.
Pubblicazione: (2025)
di: Dhiman, Ankit, et al.
Pubblicazione: (2025)
ReManNet: A Riemannian Manifold Network for Monocular 3D Lane Detection
di: Hong, Chengzhi, et al.
Pubblicazione: (2026)
di: Hong, Chengzhi, et al.
Pubblicazione: (2026)
Lift3D: Zero-Shot Lifting of Any 2D Vision Model to 3D
di: T, Mukund Varma, et al.
Pubblicazione: (2024)
di: T, Mukund Varma, et al.
Pubblicazione: (2024)
MVGSR: Multi-View Consistent 3D Gaussian Super-Resolution via Epipolar Guidance
di: Zhang, Kaizhe, et al.
Pubblicazione: (2025)
di: Zhang, Kaizhe, et al.
Pubblicazione: (2025)
SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency
di: Xie, Yiming, et al.
Pubblicazione: (2024)
di: Xie, Yiming, et al.
Pubblicazione: (2024)
M^3Detection: Multi-Frame Multi-Level Feature Fusion for Multi-Modal 3D Object Detection with Camera and 4D Imaging Radar
di: Li, Xiaozhi, et al.
Pubblicazione: (2025)
di: Li, Xiaozhi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Beyond Flatlands: Unlocking Spatial Intelligence by Decoupling 3D Reasoning from Numerical Regression
di: Guo, Zhongbin, et al.
Pubblicazione: (2025) -
Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
di: Guo, Zhongbin, et al.
Pubblicazione: (2026) -
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
di: Li, Chengzhi, et al.
Pubblicazione: (2025) -
Trace3D: Consistent Segmentation Lifting via Gaussian Instance Tracing
di: Shen, Hongyu, et al.
Pubblicazione: (2025) -
ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Diffusion Models
di: Zhu, Ruishu, et al.
Pubblicazione: (2025)