Gespeichert in:
| Hauptverfasser: | Shi, Xiangwei, Dorta, Gara, de Jong, Ruud, Shirekar, Ojas, Raman, Chirag |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2603.23089 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Physics-driven Fire Modeling from Multi-view Images
von: Dorta, Gara, et al.
Veröffentlicht: (2018)
von: Dorta, Gara, et al.
Veröffentlicht: (2018)
MindForge: Empowering Embodied Agents with Theory of Mind for Lifelong Cultural Learning
von: Lică, Mircea, et al.
Veröffentlicht: (2024)
von: Lică, Mircea, et al.
Veröffentlicht: (2024)
Audio-Synchronized Visual Animation
von: Zhang, Lin, et al.
Veröffentlicht: (2024)
von: Zhang, Lin, et al.
Veröffentlicht: (2024)
The GAN that Warped: Semantic Attribute Editing with Unpaired Data
von: Dorta, Gara, et al.
Veröffentlicht: (2018)
von: Dorta, Gara, et al.
Veröffentlicht: (2018)
Training VAEs Under Structured Residuals
von: Dorta, Gara, et al.
Veröffentlicht: (2018)
von: Dorta, Gara, et al.
Veröffentlicht: (2018)
Multi-View Camera System for Variant-Aware Autonomous Vehicle Inspection and Defect Detection
von: Kulkarni, Yash, et al.
Veröffentlicht: (2025)
von: Kulkarni, Yash, et al.
Veröffentlicht: (2025)
Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
von: Zhang, Lin, et al.
Veröffentlicht: (2025)
von: Zhang, Lin, et al.
Veröffentlicht: (2025)
Text-Guided Texturing by Synchronized Multi-View Diffusion
von: Liu, Yuxin, et al.
Veröffentlicht: (2023)
von: Liu, Yuxin, et al.
Veröffentlicht: (2023)
OmniEdit: A Training-free framework for Lip Synchronization and Audio-Visual Editing
von: Lin, Lixiang, et al.
Veröffentlicht: (2026)
von: Lin, Lixiang, et al.
Veröffentlicht: (2026)
KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
Controlled Face Manipulation and Synthesis for Data Augmentation
von: Kirchner, Joris, et al.
Veröffentlicht: (2026)
von: Kirchner, Joris, et al.
Veröffentlicht: (2026)
Multimodal Quantitative Measures for Multiparty Behaviour Evaluation
von: Shirekar, Ojas, et al.
Veröffentlicht: (2025)
von: Shirekar, Ojas, et al.
Veröffentlicht: (2025)
MVPaint: Synchronized Multi-View Diffusion for Painting Anything 3D
von: Cheng, Wei, et al.
Veröffentlicht: (2024)
von: Cheng, Wei, et al.
Veröffentlicht: (2024)
A Novel Method to Improve Quality Surface Coverage in Multi-View Capture
von: Huang, Wei-Lun, et al.
Veröffentlicht: (2024)
von: Huang, Wei-Lun, et al.
Veröffentlicht: (2024)
UniSync: A Unified Framework for Audio-Visual Synchronization
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
SyncFix: Fixing 3D Reconstructions via Multi-View Synchronization
von: Li, Deming, et al.
Veröffentlicht: (2026)
von: Li, Deming, et al.
Veröffentlicht: (2026)
Extend Your Horizon: A Device-Agnostic Surgical Tool Tracking Framework with Multi-View Optimization for Augmented Reality
von: Zhang, Jiaming, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2026)
Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence
von: Filntisis, Panagiotis P., et al.
Veröffentlicht: (2026)
von: Filntisis, Panagiotis P., et al.
Veröffentlicht: (2026)
GAINS: Gaussian-based Inverse Rendering from Sparse Multi-View Captures
von: Noras, Patrick, et al.
Veröffentlicht: (2025)
von: Noras, Patrick, et al.
Veröffentlicht: (2025)
Multimodal Transformer Distillation for Audio-Visual Synchronization
von: Chen, Xuanjun, et al.
Veröffentlicht: (2022)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2022)
Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers
von: Song, Jibin, et al.
Veröffentlicht: (2025)
von: Song, Jibin, et al.
Veröffentlicht: (2025)
ToLL: Topological Layout Learning with Asymmetric Cross-View Structural Distillation for 3D Scene Graph Generation Pretraining
von: Huang, Yucheng, et al.
Veröffentlicht: (2026)
von: Huang, Yucheng, et al.
Veröffentlicht: (2026)
Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
von: Liu, Shaowei, et al.
Veröffentlicht: (2025)
von: Liu, Shaowei, et al.
Veröffentlicht: (2025)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
von: Wang, Kai, et al.
Veröffentlicht: (2024)
von: Wang, Kai, et al.
Veröffentlicht: (2024)
Audio-driven Talking Face Generation with Stabilized Synchronization Loss
von: Yaman, Dogucan, et al.
Veröffentlicht: (2023)
von: Yaman, Dogucan, et al.
Veröffentlicht: (2023)
Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging Scenes
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
MOVA: Towards Scalable and Synchronized Video-Audio Generation
von: OpenMOSS Team, et al.
Veröffentlicht: (2026)
von: OpenMOSS Team, et al.
Veröffentlicht: (2026)
SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion
von: Ma, Junxian, et al.
Veröffentlicht: (2025)
von: Ma, Junxian, et al.
Veröffentlicht: (2025)
IDD-X: A Multi-View Dataset for Ego-relative Important Object Localization and Explanation in Dense and Unstructured Traffic
von: Parikh, Chirag, et al.
Veröffentlicht: (2024)
von: Parikh, Chirag, et al.
Veröffentlicht: (2024)
Audio-Guided Visual Editing with Complex Multi-Modal Prompts
von: Kim, Hyeonyu, et al.
Veröffentlicht: (2025)
von: Kim, Hyeonyu, et al.
Veröffentlicht: (2025)
ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition
von: Huang, Ronggang, et al.
Veröffentlicht: (2025)
von: Huang, Ronggang, et al.
Veröffentlicht: (2025)
StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model
von: Yang, Yifan, et al.
Veröffentlicht: (2025)
von: Yang, Yifan, et al.
Veröffentlicht: (2025)
Fully automated landmarking and facial segmentation on 3D photographs
von: Berends, Bo, et al.
Veröffentlicht: (2023)
von: Berends, Bo, et al.
Veröffentlicht: (2023)
EgoAVU: Egocentric Audio-Visual Understanding
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
A Novel Audio-Visual Information Fusion System for Mental Disorders Detection
von: Li, Yichun, et al.
Veröffentlicht: (2024)
von: Li, Yichun, et al.
Veröffentlicht: (2024)
Energy-Based Constraint Networks: Learning Structural Coherence Across Modalities
von: Shinde, Chirag
Veröffentlicht: (2026)
von: Shinde, Chirag
Veröffentlicht: (2026)
DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding
von: Ahmadian, Mona, et al.
Veröffentlicht: (2025)
von: Ahmadian, Mona, et al.
Veröffentlicht: (2025)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
von: Ahn, Young Jin, et al.
Veröffentlicht: (2024)
von: Ahn, Young Jin, et al.
Veröffentlicht: (2024)
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
von: Wang, Langyu, et al.
Veröffentlicht: (2025)
von: Wang, Langyu, et al.
Veröffentlicht: (2025)
MetaCap: Meta-learning Priors from Multi-View Imagery for Sparse-view Human Performance Capture and Rendering
von: Sun, Guoxing, et al.
Veröffentlicht: (2024)
von: Sun, Guoxing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Physics-driven Fire Modeling from Multi-view Images
von: Dorta, Gara, et al.
Veröffentlicht: (2018) -
MindForge: Empowering Embodied Agents with Theory of Mind for Lifelong Cultural Learning
von: Lică, Mircea, et al.
Veröffentlicht: (2024) -
Audio-Synchronized Visual Animation
von: Zhang, Lin, et al.
Veröffentlicht: (2024) -
The GAN that Warped: Semantic Attribute Editing with Unpaired Data
von: Dorta, Gara, et al.
Veröffentlicht: (2018) -
Training VAEs Under Structured Residuals
von: Dorta, Gara, et al.
Veröffentlicht: (2018)