L4P: Towards Unified Low-Level 4D Vision Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Badki, Abhishek, Su, Hang, Wen, Bowen, Gallo, Orazio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
nvTorchCam: An Open-source Library for Camera-Agnostic Differentiable Geometric Vision
by: Lichy, Daniel, et al.
Published: (2024)
by: Lichy, Daniel, et al.
Published: (2024)
Zero-Shot Monocular Scene Flow Estimation in the Wild
by: Liang, Yiqing, et al.
Published: (2025)
by: Liang, Yiqing, et al.
Published: (2025)
FoVA-Depth: Field-of-View Agnostic Depth Estimation for Cross-Dataset Generalization
by: Lichy, Daniel, et al.
Published: (2024)
by: Lichy, Daniel, et al.
Published: (2024)
UniVision: A Unified Framework for Vision-Centric 3D Perception
by: Hong, Yu, et al.
Published: (2024)
by: Hong, Yu, et al.
Published: (2024)
Exploring Scalable Unified Modeling for General Low-Level Vision
by: Chen, Xiangyu, et al.
Published: (2025)
by: Chen, Xiangyu, et al.
Published: (2025)
FoundationStereo: Zero-Shot Stereo Matching
by: Wen, Bowen, et al.
Published: (2025)
by: Wen, Bowen, et al.
Published: (2025)
THFM: A Unified Video Foundation Model for 4D Human Perception and Beyond
by: Wang, Letian, et al.
Published: (2026)
by: Wang, Letian, et al.
Published: (2026)
UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
by: Cao, Shuo, et al.
Published: (2025)
by: Cao, Shuo, et al.
Published: (2025)
D$^4$M: Dataset Distillation via Disentangled Diffusion Model
by: Su, Duo, et al.
Published: (2024)
by: Su, Duo, et al.
Published: (2024)
CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction
by: Xie, Xianghui, et al.
Published: (2025)
by: Xie, Xianghui, et al.
Published: (2025)
Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation
by: Wang, Xinshun, et al.
Published: (2026)
by: Wang, Xinshun, et al.
Published: (2026)
DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning
by: Liu, Zhe, et al.
Published: (2025)
by: Liu, Zhe, et al.
Published: (2025)
Avatar Fingerprinting for Authorized Use of Synthetic Talking-Head Videos
by: Prashnani, Ekta, et al.
Published: (2023)
by: Prashnani, Ekta, et al.
Published: (2023)
Uni-Animator: Towards Unified Visual Colorization
by: Chen, Xinyuan, et al.
Published: (2026)
by: Chen, Xinyuan, et al.
Published: (2026)
MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models
by: Wang, Kangkang, et al.
Published: (2026)
by: Wang, Kangkang, et al.
Published: (2026)
On the Global Photometric Alignment for Low-Level Vision
by: Li, Mingjia, et al.
Published: (2026)
by: Li, Mingjia, et al.
Published: (2026)
UPOCR: Towards Unified Pixel-Level OCR Interface
by: Peng, Dezhi, et al.
Published: (2023)
by: Peng, Dezhi, et al.
Published: (2023)
Kornia-rs: A Low-Level 3D Computer Vision Library In Rust
by: Riba, Edgar, et al.
Published: (2025)
by: Riba, Edgar, et al.
Published: (2025)
PAGE-4D: VGGT-4D Perception via Disentangled Pose and Geometry Estimation
by: Zhou, Kaichen, et al.
Published: (2025)
by: Zhou, Kaichen, et al.
Published: (2025)
Unifying UAV Cross-View Geo-Localization via 3D Geometric Perception
by: Li, Haoyuan, et al.
Published: (2026)
by: Li, Haoyuan, et al.
Published: (2026)
PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language Models
by: Su, Yuanhao, et al.
Published: (2026)
by: Su, Yuanhao, et al.
Published: (2026)
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
by: Pu, Yuandong, et al.
Published: (2025)
by: Pu, Yuandong, et al.
Published: (2025)
Complet4R: Geometric Complete 4D Reconstruction
by: Wang, Weibang, et al.
Published: (2026)
by: Wang, Weibang, et al.
Published: (2026)
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
by: Su, Yongyi, et al.
Published: (2025)
by: Su, Yongyi, et al.
Published: (2025)
Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
by: Qian, Shenhan, et al.
Published: (2026)
by: Qian, Shenhan, et al.
Published: (2026)
R3eVision: A Survey on Robust Rendering, Restoration, and Enhancement for 3D Low-Level Vision
by: Kwon, Weeyoung, et al.
Published: (2025)
by: Kwon, Weeyoung, et al.
Published: (2025)
VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
by: Liu, Yuqi, et al.
Published: (2025)
by: Liu, Yuqi, et al.
Published: (2025)
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
by: Chu, Xiangxiang, et al.
Published: (2024)
by: Chu, Xiangxiang, et al.
Published: (2024)
4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration
by: Zhang, Jiahui, et al.
Published: (2025)
by: Zhang, Jiahui, et al.
Published: (2025)
One4D: Unified 4D Generation and Reconstruction via Decoupled LoRA Control
by: Mi, Zhenxing, et al.
Published: (2025)
by: Mi, Zhenxing, et al.
Published: (2025)
CORE4D: A 4D Human-Object-Human Interaction Dataset for Collaborative Object REarrangement
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
Mono4DEditor: Text-Driven 4D Scene Editing from Monocular Video via Point-Level Localization of Language-Embedded Gaussians
by: Shi, Jin-Chuan, et al.
Published: (2025)
by: Shi, Jin-Chuan, et al.
Published: (2025)
Self-Improving 4D Perception via Self-Distillation
by: Huang, Nan, et al.
Published: (2026)
by: Huang, Nan, et al.
Published: (2026)
Towards Pixel-Level VLM Perception via Simple Points Prediction
by: Song, Tianhui, et al.
Published: (2026)
by: Song, Tianhui, et al.
Published: (2026)
Towards Unifying Understanding and Generation in the Era of Vision Foundation Models: A Survey from the Autoregression Perspective
by: Xie, Shenghao, et al.
Published: (2024)
by: Xie, Shenghao, et al.
Published: (2024)
TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields
by: Huang, Tianyu, et al.
Published: (2023)
by: Huang, Tianyu, et al.
Published: (2023)
Towards a Unified Copernicus Foundation Model for Earth Vision
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
One for All: Toward Unified Foundation Models for Earth Vision
by: Xiong, Zhitong, et al.
Published: (2024)
by: Xiong, Zhitong, et al.
Published: (2024)
Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perception
by: Lin, Jiajing, et al.
Published: (2024)
by: Lin, Jiajing, et al.
Published: (2024)
HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction
by: Liu, Yunze, et al.
Published: (2022)
by: Liu, Yunze, et al.
Published: (2022)
Similar Items
-
nvTorchCam: An Open-source Library for Camera-Agnostic Differentiable Geometric Vision
by: Lichy, Daniel, et al.
Published: (2024) -
Zero-Shot Monocular Scene Flow Estimation in the Wild
by: Liang, Yiqing, et al.
Published: (2025) -
FoVA-Depth: Field-of-View Agnostic Depth Estimation for Cross-Dataset Generalization
by: Lichy, Daniel, et al.
Published: (2024) -
UniVision: A Unified Framework for Vision-Centric 3D Perception
by: Hong, Yu, et al.
Published: (2024) -
Exploring Scalable Unified Modeling for General Low-Level Vision
by: Chen, Xiangyu, et al.
Published: (2025)