UniVision: A Unified Framework for Vision-Centric 3D Perception
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Yu, Liu, Qian, Cheng, Huayuan, Ma, Danjiao, Dai, Hang, Wang, Yu, Cao, Guangzhi, Ding, Yong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
von: Huang, Ziyuan, et al.
Veröffentlicht: (2025)
von: Huang, Ziyuan, et al.
Veröffentlicht: (2025)
UniISP: A Unified ISP Framework for Both Human and Machine Vision
von: Li, Hanxi, et al.
Veröffentlicht: (2026)
von: Li, Hanxi, et al.
Veröffentlicht: (2026)
UniBEVFusion: Unified Radar-Vision BEVFusion for 3D Object Detection
von: Zhao, Haocheng, et al.
Veröffentlicht: (2024)
von: Zhao, Haocheng, et al.
Veröffentlicht: (2024)
L4P: Towards Unified Low-Level 4D Vision Perception
von: Badki, Abhishek, et al.
Veröffentlicht: (2025)
von: Badki, Abhishek, et al.
Veröffentlicht: (2025)
UniViTAR: Unified Vision Transformer with Native Resolution
von: Qiao, Limeng, et al.
Veröffentlicht: (2025)
von: Qiao, Limeng, et al.
Veröffentlicht: (2025)
Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving
von: Yang, Yu, et al.
Veröffentlicht: (2024)
von: Yang, Yu, et al.
Veröffentlicht: (2024)
Large Trajectory Models are Scalable Motion Predictors and Planners
von: Sun, Qiao, et al.
Veröffentlicht: (2023)
von: Sun, Qiao, et al.
Veröffentlicht: (2023)
UniVRSE: Unified Vision-conditioned Response Semantic Entropy for Hallucination Detection in Medical Vision-Language Models
von: Liao, Zehui, et al.
Veröffentlicht: (2025)
von: Liao, Zehui, et al.
Veröffentlicht: (2025)
UniDream: Unifying Diffusion Priors for Relightable Text-to-3D Generation
von: Liu, Zexiang, et al.
Veröffentlicht: (2023)
von: Liu, Zexiang, et al.
Veröffentlicht: (2023)
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models
von: Li, Yujie, et al.
Veröffentlicht: (2024)
von: Li, Yujie, et al.
Veröffentlicht: (2024)
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
von: Pang, Youxin, et al.
Veröffentlicht: (2025)
von: Pang, Youxin, et al.
Veröffentlicht: (2025)
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation
von: Ding, Hongyu, et al.
Veröffentlicht: (2026)
von: Ding, Hongyu, et al.
Veröffentlicht: (2026)
Vision-Centric 4D Occupancy Forecasting and Planning via Implicit Residual World Models
von: Mei, Jianbiao, et al.
Veröffentlicht: (2025)
von: Mei, Jianbiao, et al.
Veröffentlicht: (2025)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
SPARK: Multi-Vision Sensor Perception and Reasoning Benchmark for Large-scale Vision-Language Models
von: Yu, Youngjoon, et al.
Veröffentlicht: (2024)
von: Yu, Youngjoon, et al.
Veröffentlicht: (2024)
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation
von: Wang, Jiayun, et al.
Veröffentlicht: (2026)
von: Wang, Jiayun, et al.
Veröffentlicht: (2026)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
UniCorrn: Unified Correspondence Transformer Across 2D and 3D
von: Goswami, Prajnan, et al.
Veröffentlicht: (2026)
von: Goswami, Prajnan, et al.
Veröffentlicht: (2026)
EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
von: Shan, Haozhe, et al.
Veröffentlicht: (2026)
von: Shan, Haozhe, et al.
Veröffentlicht: (2026)
EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
von: Yuan, Yuqian, et al.
Veröffentlicht: (2026)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2026)
Uni$^2$Det: Unified and Universal Framework for Prompt-Guided Multi-dataset 3D Detection
von: Wang, Yubin, et al.
Veröffentlicht: (2024)
von: Wang, Yubin, et al.
Veröffentlicht: (2024)
Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation
von: Ling, Lu, et al.
Veröffentlicht: (2025)
von: Ling, Lu, et al.
Veröffentlicht: (2025)
UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
von: Chen, Lan, et al.
Veröffentlicht: (2025)
von: Chen, Lan, et al.
Veröffentlicht: (2025)
VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
von: Liu, Yuqi, et al.
Veröffentlicht: (2025)
von: Liu, Yuqi, et al.
Veröffentlicht: (2025)
UniDA3D: A Unified Domain-Adaptive Framework for Multi-View 3D Object Detection
von: Wu, Hongjing, et al.
Veröffentlicht: (2026)
von: Wu, Hongjing, et al.
Veröffentlicht: (2026)
Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models
von: Jiao, Yang, et al.
Veröffentlicht: (2024)
von: Jiao, Yang, et al.
Veröffentlicht: (2024)
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
von: Man, Yunze, et al.
Veröffentlicht: (2025)
von: Man, Yunze, et al.
Veröffentlicht: (2025)
UniHM: Unified Dexterous Hand Manipulation with Vision Language Model
von: Zhang, Zhenhao, et al.
Veröffentlicht: (2026)
von: Zhang, Zhenhao, et al.
Veröffentlicht: (2026)
UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation
von: Cao, Chenjie, et al.
Veröffentlicht: (2025)
von: Cao, Chenjie, et al.
Veröffentlicht: (2025)
UniRecGen: Unifying Multi-View 3D Reconstruction and Generation
von: Huang, Zhisheng, et al.
Veröffentlicht: (2026)
von: Huang, Zhisheng, et al.
Veröffentlicht: (2026)
Uni-MuMER: Unified Multi-Task Fine-Tuning of Vision-Language Model for Handwritten Mathematical Expression Recognition
von: Li, Yu, et al.
Veröffentlicht: (2025)
von: Li, Yu, et al.
Veröffentlicht: (2025)
AnyRefill: A Unified, Data-Efficient Framework for Left-Prompt-Guided Vision Tasks
von: Xie, Ming, et al.
Veröffentlicht: (2025)
von: Xie, Ming, et al.
Veröffentlicht: (2025)
UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
von: Cao, Shuo, et al.
Veröffentlicht: (2025)
von: Cao, Shuo, et al.
Veröffentlicht: (2025)
UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
von: Xu, Yueming, et al.
Veröffentlicht: (2025)
von: Xu, Yueming, et al.
Veröffentlicht: (2025)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
UniTabNet: Bridging Vision and Language Models for Enhanced Table Structure Recognition
von: Zhang, Zhenrong, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenrong, et al.
Veröffentlicht: (2024)
UniQA: Unified Vision-Language Pre-training for Image Quality and Aesthetic Assessment
von: Zhou, Hantao, et al.
Veröffentlicht: (2024)
von: Zhou, Hantao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
von: Huang, Ziyuan, et al.
Veröffentlicht: (2025) -
UniISP: A Unified ISP Framework for Both Human and Machine Vision
von: Li, Hanxi, et al.
Veröffentlicht: (2026) -
UniBEVFusion: Unified Radar-Vision BEVFusion for 3D Object Detection
von: Zhao, Haocheng, et al.
Veröffentlicht: (2024) -
L4P: Towards Unified Low-Level 4D Vision Perception
von: Badki, Abhishek, et al.
Veröffentlicht: (2025) -
UniViTAR: Unified Vision Transformer with Native Resolution
von: Qiao, Limeng, et al.
Veröffentlicht: (2025)