DiVR: incorporating context from diverse VR scenes for human trajectory prediction
Fuente:
arXiv
Salvato in:
| Autori principali: | Gallo, Franz Franco, Wu, Hui-Yin, Sassatelli, Lucile |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Focus360: Guiding User Attention in Immersive Videos for VR
di: Silva, Paulo Vitor S., et al.
Pubblicazione: (2026)
di: Silva, Paulo Vitor S., et al.
Pubblicazione: (2026)
Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset
di: Ancarani, Elisa, et al.
Pubblicazione: (2025)
di: Ancarani, Elisa, et al.
Pubblicazione: (2025)
Multi-Modal Multi-Task Federated Foundation Models for Next-Generation Extended Reality Systems: Towards Privacy-Preserving Distributed Intelligence in AR/VR/MR
di: Nadimi, Fardis, et al.
Pubblicazione: (2025)
di: Nadimi, Fardis, et al.
Pubblicazione: (2025)
ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
di: Sun, Jiahui, et al.
Pubblicazione: (2025)
di: Sun, Jiahui, et al.
Pubblicazione: (2025)
Thief of Truth: VR comics about the relationship between AI and humans
di: Bae, Joonhyung
Pubblicazione: (2025)
di: Bae, Joonhyung
Pubblicazione: (2025)
Semantic-Guided Unsupervised Video Summarization
di: Liu, Haizhou, et al.
Pubblicazione: (2026)
di: Liu, Haizhou, et al.
Pubblicazione: (2026)
LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?
di: Wang, Xiaohan, et al.
Pubblicazione: (2026)
di: Wang, Xiaohan, et al.
Pubblicazione: (2026)
AniME: Adaptive Multi-Agent Planning for Long Animation Generation
di: Zhang, Lisai, et al.
Pubblicazione: (2025)
di: Zhang, Lisai, et al.
Pubblicazione: (2025)
Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides
di: Zhao, Jinghua, et al.
Pubblicazione: (2025)
di: Zhao, Jinghua, et al.
Pubblicazione: (2025)
HeadsetOff: Enabling Photorealistic Video Conferencing on Economical VR Headsets
di: Jin, Yili, et al.
Pubblicazione: (2024)
di: Jin, Yili, et al.
Pubblicazione: (2024)
Latency Effects on Multi-Dimensional QoE in Networked VR Whiteboards
di: Song, Jiarun, et al.
Pubblicazione: (2026)
di: Song, Jiarun, et al.
Pubblicazione: (2026)
LLMER: Crafting Interactive Extended Reality Worlds with JSON Data Generated by Large Language Models
di: Chen, Jiangong, et al.
Pubblicazione: (2025)
di: Chen, Jiangong, et al.
Pubblicazione: (2025)
Back to Basics: Revisiting ASR in the Age of Voice Agents
di: Tay, Geeyang, et al.
Pubblicazione: (2026)
di: Tay, Geeyang, et al.
Pubblicazione: (2026)
MaLoRA: Gated Modality LoRA for Key-Space Alignment in Multimodal LLM Fine-Tuning
di: Zheng, Xinhan, et al.
Pubblicazione: (2025)
di: Zheng, Xinhan, et al.
Pubblicazione: (2025)
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
di: Liu, Jiajun, et al.
Pubblicazione: (2024)
di: Liu, Jiajun, et al.
Pubblicazione: (2024)
RA-BLIP: Multimodal Adaptive Retrieval-Augmented Bootstrapping Language-Image Pre-training
di: Ding, Muhe, et al.
Pubblicazione: (2024)
di: Ding, Muhe, et al.
Pubblicazione: (2024)
Towards Open-Vocabulary Video Semantic Segmentation
di: Li, Xinhao, et al.
Pubblicazione: (2024)
di: Li, Xinhao, et al.
Pubblicazione: (2024)
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
di: Choi, Jeongsoo, et al.
Pubblicazione: (2025)
di: Choi, Jeongsoo, et al.
Pubblicazione: (2025)
Wireless Video Semantic Communication with Decoupled Diffusion Multi-frame Compensation
di: Xie, Bingyan, et al.
Pubblicazione: (2025)
di: Xie, Bingyan, et al.
Pubblicazione: (2025)
SEER: Semantic Enhancement and Emotional Reasoning Network for Multimodal Fake News Detection
di: Zhu, Peican, et al.
Pubblicazione: (2025)
di: Zhu, Peican, et al.
Pubblicazione: (2025)
EyeNexus: Adaptive Gaze-Driven Quality and Bitrate Streaming for Seamless VR Cloud Gaming Experiences
di: Wu, Ze, et al.
Pubblicazione: (2025)
di: Wu, Ze, et al.
Pubblicazione: (2025)
Optimizing QoE-Privacy Tradeoff for Proactive VR Streaming
di: Wei, Xing, et al.
Pubblicazione: (2025)
di: Wei, Xing, et al.
Pubblicazione: (2025)
A Survey on Music Generation from Single-Modal, Cross-Modal, and Multi-Modal Perspectives
di: Li, Shuyu, et al.
Pubblicazione: (2025)
di: Li, Shuyu, et al.
Pubblicazione: (2025)
Tile-Weighted Rate-Distortion Optimized Packet Scheduling for 360$^\circ$ VR Video Streaming
di: Wang, Haopeng, et al.
Pubblicazione: (2024)
di: Wang, Haopeng, et al.
Pubblicazione: (2024)
Rethinking Multi-Condition DiTs: Eliminating Redundant Attention via Position-Alignment and Keyword-Scoping
di: Zhou, Chao, et al.
Pubblicazione: (2026)
di: Zhou, Chao, et al.
Pubblicazione: (2026)
GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment
di: Wang, Jinting, et al.
Pubblicazione: (2025)
di: Wang, Jinting, et al.
Pubblicazione: (2025)
AI-Integrated Decision Support System for Real-Time Market Growth Forecasting and Multi-Source Content Diffusion Analytics
di: Yin, Ziqing, et al.
Pubblicazione: (2025)
di: Yin, Ziqing, et al.
Pubblicazione: (2025)
Interpretable Multimodal Misinformation Detection with Logic Reasoning
di: Liu, Hui, et al.
Pubblicazione: (2023)
di: Liu, Hui, et al.
Pubblicazione: (2023)
Interest-Aware Joint Caching, Computing, and Communication Optimization for Mobile VR Delivery in MEC Networks
di: Fu, Baojie, et al.
Pubblicazione: (2024)
di: Fu, Baojie, et al.
Pubblicazione: (2024)
DiET-GS: Diffusion Prior and Event Stream-Assisted Motion Deblurring 3D Gaussian Splatting
di: Lee, Seungjun, et al.
Pubblicazione: (2025)
di: Lee, Seungjun, et al.
Pubblicazione: (2025)
EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction
di: Jing, Chong, et al.
Pubblicazione: (2026)
di: Jing, Chong, et al.
Pubblicazione: (2026)
Synthesizing Sentiment-Controlled Feedback For Multimodal Text and Image Data
di: Kumar, Puneet, et al.
Pubblicazione: (2024)
di: Kumar, Puneet, et al.
Pubblicazione: (2024)
HiQuE: Hierarchical Question Embedding Network for Multimodal Depression Detection
di: Jung, Juho, et al.
Pubblicazione: (2024)
di: Jung, Juho, et al.
Pubblicazione: (2024)
Has Multimodal Learning Delivered Universal Intelligence in Healthcare? A Comprehensive Survey
di: Lin, Qika, et al.
Pubblicazione: (2024)
di: Lin, Qika, et al.
Pubblicazione: (2024)
Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
di: Cheng, Zebang, et al.
Pubblicazione: (2024)
di: Cheng, Zebang, et al.
Pubblicazione: (2024)
A Survey on Multimodal Benchmarks: In the Era of Large AI Models
di: Li, Lin, et al.
Pubblicazione: (2024)
di: Li, Lin, et al.
Pubblicazione: (2024)
HumanVLM: Foundation for Human-Scene Vision-Language Model
di: Dai, Dawei, et al.
Pubblicazione: (2024)
di: Dai, Dawei, et al.
Pubblicazione: (2024)
Multimodal Emotion Recognition by Fusing Video Semantic in MOOC Learning Scenarios
di: Zhang, Yuan, et al.
Pubblicazione: (2024)
di: Zhang, Yuan, et al.
Pubblicazione: (2024)
AxiomVision: Accuracy-Guaranteed Adaptive Visual Model Selection for Perspective-Aware Video Analytics
di: Dai, Xiangxiang, et al.
Pubblicazione: (2024)
di: Dai, Xiangxiang, et al.
Pubblicazione: (2024)
M3TR: Temporal Retrieval Enhanced Multi-Modal Micro-video Popularity Prediction
di: Lu, Jiacheng, et al.
Pubblicazione: (2024)
di: Lu, Jiacheng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Focus360: Guiding User Attention in Immersive Videos for VR
di: Silva, Paulo Vitor S., et al.
Pubblicazione: (2026) -
Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset
di: Ancarani, Elisa, et al.
Pubblicazione: (2025) -
Multi-Modal Multi-Task Federated Foundation Models for Next-Generation Extended Reality Systems: Towards Privacy-Preserving Distributed Intelligence in AR/VR/MR
di: Nadimi, Fardis, et al.
Pubblicazione: (2025) -
ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
di: Sun, Jiahui, et al.
Pubblicazione: (2025) -
Thief of Truth: VR comics about the relationship between AI and humans
di: Bae, Joonhyung
Pubblicazione: (2025)