Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Zhengxian, Wang, Shengqi, Pan, Shi, Li, Hongshuai, Wang, Haoxiang, Li, Lin, Li, Guanjun, Wen, Zhengqi, Lin, Borong, Tao, Jianhua, Yu, Tao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ImViD: Immersive Volumetric Videos for Enhanced VR Engagement
by: Yang, Zhengxian, et al.
Published: (2025)
by: Yang, Zhengxian, et al.
Published: (2025)
Den-SOFT: Dense Space-Oriented Light Field DataseT for 6-DOF Immersive Experience
by: Yu, Xiaohang, et al.
Published: (2024)
by: Yu, Xiaohang, et al.
Published: (2024)
Exploring the Role of Audio in Multimodal Misinformation Detection
by: Liu, Moyang, et al.
Published: (2024)
by: Liu, Moyang, et al.
Published: (2024)
Task-Oriented 6-DoF Grasp Pose Detection in Clutters
by: Wang, An-Lan, et al.
Published: (2025)
by: Wang, An-Lan, et al.
Published: (2025)
Generalizing 6-DoF Grasp Detection via Domain Prior Knowledge
by: Ma, Haoxiang, et al.
Published: (2024)
by: Ma, Haoxiang, et al.
Published: (2024)
ClipGS-VR: Immersive and Interactive Cinematic Visualization of Volumetric Medical Data in Mobile Virtual Reality
by: Tong, Yuqi, et al.
Published: (2026)
by: Tong, Yuqi, et al.
Published: (2026)
DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis
by: Wen, Yuhua, et al.
Published: (2025)
by: Wen, Yuhua, et al.
Published: (2025)
Whole-Body Control With Terrain Estimation of A 6-DoF Wheeled Bipedal Robot
by: Wen, Cong, et al.
Published: (2025)
by: Wen, Cong, et al.
Published: (2025)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
by: Ren, Yong, et al.
Published: (2026)
by: Ren, Yong, et al.
Published: (2026)
MTPareto: A MultiModal Targeted Pareto Framework for Fake News Detection
by: Yan, Kaiying, et al.
Published: (2025)
by: Yan, Kaiying, et al.
Published: (2025)
Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
by: Xiong, Chenxu, et al.
Published: (2024)
by: Xiong, Chenxu, et al.
Published: (2024)
An Economic Framework for 6-DoF Grasp Detection
by: Wu, Xiao-Ming, et al.
Published: (2024)
by: Wu, Xiao-Ming, et al.
Published: (2024)
PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment Analysis
by: Xie, Heng, et al.
Published: (2025)
by: Xie, Heng, et al.
Published: (2025)
RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection
by: Fu, Ruibo, et al.
Published: (2025)
by: Fu, Ruibo, et al.
Published: (2025)
Dynamic Agile Reconfigurable Intelligent Surface Antenna (DARISA) MIMO: DoF Analysis and Effective DoF Optimization
by: Bai, Jiale, et al.
Published: (2025)
by: Bai, Jiale, et al.
Published: (2025)
Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation
by: Zhang, Haojie, et al.
Published: (2024)
by: Zhang, Haojie, et al.
Published: (2024)
6-DoF Grasp Detection in Clutter with Enhanced Receptive Field and Graspable Balance Sampling
by: Wang, Hanwen, et al.
Published: (2024)
by: Wang, Hanwen, et al.
Published: (2024)
Dexterous Teleoperation of 20-DoF ByteDexter Hand via Human Motion Retargeting
by: Wen, Ruoshi, et al.
Published: (2025)
by: Wen, Ruoshi, et al.
Published: (2025)
EyeNavGS: A 6-DoF Navigation Dataset and Record-n-Replay Software for Real-World 3DGS Scenes in VR
by: Ding, Zihao, et al.
Published: (2025)
by: Ding, Zihao, et al.
Published: (2025)
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
by: Qi, Xin, et al.
Published: (2024)
by: Qi, Xin, et al.
Published: (2024)
Six-DoF Hand-Based Teleoperation for Omnidirectional Aerial Robots
by: Li, Jinjie, et al.
Published: (2025)
by: Li, Jinjie, et al.
Published: (2025)
EELE: Exploring Efficient and Extensible LoRA Integration in Emotional Text-to-Speech
by: Qi, Xin, et al.
Published: (2024)
by: Qi, Xin, et al.
Published: (2024)
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
A Noval Feature via Color Quantisation for Fake Audio Detection
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
Dexora: Open-source VLA for High-DoF Bimanual Dexterity
by: Zhang, Zongzheng, et al.
Published: (2026)
by: Zhang, Zongzheng, et al.
Published: (2026)
A fast and efficient numerical method for computing the stress concentration between closely located stiff inclusions of general shapes
by: Li, Xiaofei, et al.
Published: (2023)
by: Li, Xiaofei, et al.
Published: (2023)
Question Answering for Decisionmaking in Green Building Design: A Multimodal Data Reasoning Method Driven by Large Language Models
by: Li, Yihui, et al.
Published: (2024)
by: Li, Yihui, et al.
Published: (2024)
Chapter View Synthesis Tool for VR Immersive Video
by: Fachada, Sarah, et al.
Published: (2024)
by: Fachada, Sarah, et al.
Published: (2024)
RT-NeRF: Real-Time On-Device Neural Radiance Fields Towards Immersive AR/VR Rendering
by: Li, Chaojian, et al.
Published: (2022)
by: Li, Chaojian, et al.
Published: (2022)
Scalable Unseen Objects 6-DoF Absolute Pose Estimation with Robotic Integration
by: Liu, Jian, et al.
Published: (2025)
by: Liu, Jian, et al.
Published: (2025)
MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation
by: Fu, Ruibo, et al.
Published: (2024)
by: Fu, Ruibo, et al.
Published: (2024)
Keypoint-based Dynamic Object 6-DoF Pose Tracking via Event Camera
by: Wang, Zhe, et al.
Published: (2026)
by: Wang, Zhe, et al.
Published: (2026)
DoF-Gaussian: Controllable Depth-of-Field for 3D Gaussian Splatting
by: Shen, Liao, et al.
Published: (2025)
by: Shen, Liao, et al.
Published: (2025)
Residual Speaker Representation for One-Shot Voice Conversion
by: Xu, Le, et al.
Published: (2023)
by: Xu, Le, et al.
Published: (2023)
Deep Learning Approaches for Multimodal Intent Recognition: A Survey
by: Zhao, Jingwei, et al.
Published: (2025)
by: Zhao, Jingwei, et al.
Published: (2025)
Chapter 3 DoF/6 DoF Localization System for Low Computing Power Mobile Robot Platforms
by: Costa, Carlos M., et al.
Published: (2021)
by: Costa, Carlos M., et al.
Published: (2021)
Does Current Deepfake Audio Detection Model Effectively Detect ALM-based Deepfake Audio?
by: Xie, Yuankun, et al.
Published: (2024)
by: Xie, Yuankun, et al.
Published: (2024)
Which Channel in 6G, Low-rank or Full-rank, more needs RIS from a Perspective of DoF?
by: Li, Yongqiang, et al.
Published: (2024)
by: Li, Yongqiang, et al.
Published: (2024)
VVLoc: Prior-free 3-DoF Vehicle Visual Localization
by: Huang, Ze, et al.
Published: (2026)
by: Huang, Ze, et al.
Published: (2026)
Immersive Volumetric Video Playback: Near-RT Resource Allocation and O-RAN-based Implementation
by: Wen, Yao, et al.
Published: (2026)
by: Wen, Yao, et al.
Published: (2026)
Similar Items
-
ImViD: Immersive Volumetric Videos for Enhanced VR Engagement
by: Yang, Zhengxian, et al.
Published: (2025) -
Den-SOFT: Dense Space-Oriented Light Field DataseT for 6-DOF Immersive Experience
by: Yu, Xiaohang, et al.
Published: (2024) -
Exploring the Role of Audio in Multimodal Misinformation Detection
by: Liu, Moyang, et al.
Published: (2024) -
Task-Oriented 6-DoF Grasp Pose Detection in Clutters
by: Wang, An-Lan, et al.
Published: (2025) -
Generalizing 6-DoF Grasp Detection via Domain Prior Knowledge
by: Ma, Haoxiang, et al.
Published: (2024)