LLMI3D: MLLM-based 3D Perception from a Single 2D Image
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Fan, Zhao, Sicheng, Zhang, Yanhao, Chen, Hui, Lu, Haonan, Han, Jungong, Ding, Guiguang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator
von: Yang, Fan, et al.
Veröffentlicht: (2024)
von: Yang, Fan, et al.
Veröffentlicht: (2024)
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models
von: Sun, Fengyuan, et al.
Veröffentlicht: (2025)
von: Sun, Fengyuan, et al.
Veröffentlicht: (2025)
CAIT: Triple-Win Compression towards High Accuracy, Fast Inference, and Favorable Transferability For ViTs
von: Wang, Ao, et al.
Veröffentlicht: (2023)
von: Wang, Ao, et al.
Veröffentlicht: (2023)
Tracking and Segmenting Anything in Any Modality
von: Zhang, Tianlu, et al.
Veröffentlicht: (2025)
von: Zhang, Tianlu, et al.
Veröffentlicht: (2025)
Towards Efficient Vision-Language Tuning: More Information Density, More Generalizability
von: Hao, Tianxiang, et al.
Veröffentlicht: (2023)
von: Hao, Tianxiang, et al.
Veröffentlicht: (2023)
GEN3D: Generating Domain-Free 3D Scenes from a Single Image
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025)
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025)
RepViT-SAM: Towards Real-Time Segmenting Anything
von: Wang, Ao, et al.
Veröffentlicht: (2023)
von: Wang, Ao, et al.
Veröffentlicht: (2023)
LSNet: See Large, Focus Small
von: Wang, Ao, et al.
Veröffentlicht: (2025)
von: Wang, Ao, et al.
Veröffentlicht: (2025)
RepViT: Revisiting Mobile CNN From ViT Perspective
von: Wang, Ao, et al.
Veröffentlicht: (2023)
von: Wang, Ao, et al.
Veröffentlicht: (2023)
VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image
von: Xu, Sicheng, et al.
Veröffentlicht: (2025)
von: Xu, Sicheng, et al.
Veröffentlicht: (2025)
PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning
von: Sun, Fengyuan, et al.
Veröffentlicht: (2025)
von: Sun, Fengyuan, et al.
Veröffentlicht: (2025)
Learn from the Learnt: Source-Free Active Domain Adaptation via Contrastive Sampling and Visual Persistence
von: Lyu, Mengyao, et al.
Veröffentlicht: (2024)
von: Lyu, Mengyao, et al.
Veröffentlicht: (2024)
NOVA3D: Normal Aligned Video Diffusion Model for Single Image to 3D Generation
von: Yang, Yuxiao, et al.
Veröffentlicht: (2025)
von: Yang, Yuxiao, et al.
Veröffentlicht: (2025)
Compress3D: a Compressed Latent Space for 3D Generation from a Single Image
von: Zhang, Bowen, et al.
Veröffentlicht: (2024)
von: Zhang, Bowen, et al.
Veröffentlicht: (2024)
One-Dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing Applications
von: Lyu, Mengyao, et al.
Veröffentlicht: (2023)
von: Lyu, Mengyao, et al.
Veröffentlicht: (2023)
[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs
von: Wang, Ao, et al.
Veröffentlicht: (2024)
von: Wang, Ao, et al.
Veröffentlicht: (2024)
YOLOE: Real-Time Seeing Anything
von: Wang, Ao, et al.
Veröffentlicht: (2025)
von: Wang, Ao, et al.
Veröffentlicht: (2025)
Promptable Anomaly Segmentation with SAM Through Self-Perception Tuning
von: Yang, Hui-Yue, et al.
Veröffentlicht: (2024)
von: Yang, Hui-Yue, et al.
Veröffentlicht: (2024)
ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation
von: Zhou, Yuan, et al.
Veröffentlicht: (2025)
von: Zhou, Yuan, et al.
Veröffentlicht: (2025)
Cream of the Crop: Harvesting Rich, Scalable and Transferable Multi-Modal Data for Instruction Fine-Tuning
von: Lyu, Mengyao, et al.
Veröffentlicht: (2025)
von: Lyu, Mengyao, et al.
Veröffentlicht: (2025)
Rethinking Score Distilling Sampling for 3D Editing and Generation
von: Miao, Xingyu, et al.
Veröffentlicht: (2025)
von: Miao, Xingyu, et al.
Veröffentlicht: (2025)
Hunyuan3D 1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation
von: Yang, Xianghui, et al.
Veröffentlicht: (2024)
von: Yang, Xianghui, et al.
Veröffentlicht: (2024)
Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual Variations
von: Liang, Yiwen, et al.
Veröffentlicht: (2025)
von: Liang, Yiwen, et al.
Veröffentlicht: (2025)
Beyond Voxel 3D Editing: Learning from 3D Masks and Self-Constructed Data
von: Xu, Yizhao, et al.
Veröffentlicht: (2026)
von: Xu, Yizhao, et al.
Veröffentlicht: (2026)
BIP3D: Bridging 2D Images and 3D Perception for Embodied Intelligence
von: Lin, Xuewu, et al.
Veröffentlicht: (2024)
von: Lin, Xuewu, et al.
Veröffentlicht: (2024)
YOLOv10: Real-Time End-to-End Object Detection
von: Wang, Ao, et al.
Veröffentlicht: (2024)
von: Wang, Ao, et al.
Veröffentlicht: (2024)
Ultraman: Single Image 3D Human Reconstruction with Ultra Speed and Detail
von: Chen, Mingjin, et al.
Veröffentlicht: (2024)
von: Chen, Mingjin, et al.
Veröffentlicht: (2024)
S$^2$-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
von: Xu, Beining, et al.
Veröffentlicht: (2025)
von: Xu, Beining, et al.
Veröffentlicht: (2025)
Source-Free Object Detection with Detection Transformer
von: Yao, Huizai, et al.
Veröffentlicht: (2025)
von: Yao, Huizai, et al.
Veröffentlicht: (2025)
Wonder3D++: Cross-domain Diffusion for High-fidelity 3D Generation from a Single Image
von: Yang, Yuxiao, et al.
Veröffentlicht: (2025)
von: Yang, Yuxiao, et al.
Veröffentlicht: (2025)
Context Enhancement with Reconstruction as Sequence for Unified Unsupervised Anomaly Detection
von: Yang, Hui-Yue, et al.
Veröffentlicht: (2024)
von: Yang, Hui-Yue, et al.
Veröffentlicht: (2024)
Constructing a 3D Scene from a Single Image
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2025)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2025)
GSV3D: Gaussian Splatting-based Geometric Distillation with Stable Video Diffusion for Single-Image 3D Object Generation
von: Tao, Ye, et al.
Veröffentlicht: (2025)
von: Tao, Ye, et al.
Veröffentlicht: (2025)
DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval
von: Shen, Leqi, et al.
Veröffentlicht: (2025)
von: Shen, Leqi, et al.
Veröffentlicht: (2025)
YOLO-UniOW: Efficient Universal Open-World Object Detection
von: Liu, Lihao, et al.
Veröffentlicht: (2024)
von: Liu, Lihao, et al.
Veröffentlicht: (2024)
Native and Compact Structured Latents for 3D Generation
von: Xiang, Jianfeng, et al.
Veröffentlicht: (2025)
von: Xiang, Jianfeng, et al.
Veröffentlicht: (2025)
More is Better: Deep Domain Adaptation with Multiple Sources
von: Zhao, Sicheng, et al.
Veröffentlicht: (2024)
von: Zhao, Sicheng, et al.
Veröffentlicht: (2024)
TGP: Two-modal occupancy prediction with 3D Gaussian and sparse points for 3D Environment Awareness
von: Chen, Mu, et al.
Veröffentlicht: (2025)
von: Chen, Mu, et al.
Veröffentlicht: (2025)
Point Linguist Model: Segment Any Object via Bridged Large 3D-Language Model
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2025)
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2025)
Feedforward 3D Editing via Text-Steerable Image-to-3D
von: Ma, Ziqi, et al.
Veröffentlicht: (2025)
von: Ma, Ziqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator
von: Yang, Fan, et al.
Veröffentlicht: (2024) -
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models
von: Sun, Fengyuan, et al.
Veröffentlicht: (2025) -
CAIT: Triple-Win Compression towards High Accuracy, Fast Inference, and Favorable Transferability For ViTs
von: Wang, Ao, et al.
Veröffentlicht: (2023) -
Tracking and Segmenting Anything in Any Modality
von: Zhang, Tianlu, et al.
Veröffentlicht: (2025) -
Towards Efficient Vision-Language Tuning: More Information Density, More Generalizability
von: Hao, Tianxiang, et al.
Veröffentlicht: (2023)