Saved in:
| Main Authors: | Shi, Sheng, Cao, Xuyang, Zhao, Jun, Wang, Guoxin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2409.13268 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation
by: Cao, Xuyang, et al.
Published: (2024)
by: Cao, Xuyang, et al.
Published: (2024)
JoyType: A Robust Design for Multilingual Visual Text Creation
by: Li, Chao, et al.
Published: (2024)
by: Li, Chao, et al.
Published: (2024)
Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
by: Xu, Mingwang, et al.
Published: (2024)
by: Xu, Mingwang, et al.
Published: (2024)
Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation
by: Cui, Jiahao, et al.
Published: (2024)
by: Cui, Jiahao, et al.
Published: (2024)
Hallo4: High-Fidelity Dynamic Portrait Animation via Direct Preference Optimization
by: Cui, Jiahao, et al.
Published: (2025)
by: Cui, Jiahao, et al.
Published: (2025)
Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
by: Cui, Jiahao, et al.
Published: (2024)
by: Cui, Jiahao, et al.
Published: (2024)
JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video Editing
by: Wang, Qili, et al.
Published: (2025)
by: Wang, Qili, et al.
Published: (2025)
Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning
by: Wang, Guoxin, et al.
Published: (2025)
by: Wang, Guoxin, et al.
Published: (2025)
A Bi-Pyramid Multimodal Fusion Method for the Diagnosis of Bipolar Disorders
by: Wang, Guoxin, et al.
Published: (2024)
by: Wang, Guoxin, et al.
Published: (2024)
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
by: Li, Chunyu, et al.
Published: (2026)
by: Li, Chunyu, et al.
Published: (2026)
ARM: A Learnable, Plug-and-Play Module for CLIP-based Open-vocabulary Semantic Segmentation
by: Liu, Ziquan, et al.
Published: (2025)
by: Liu, Ziquan, et al.
Published: (2025)
T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation
by: Chen, Yubin, et al.
Published: (2025)
by: Chen, Yubin, et al.
Published: (2025)
JoyStreamer: Unlocking Highly Expressive Avatars via Harmonized Text-Audio Conditioning
by: Wang, Ruikui, et al.
Published: (2026)
by: Wang, Ruikui, et al.
Published: (2026)
JoyStreamer-Flash: Real-time and Infinite Audio-Driven Avatar Generation with Autoregressive Diffusion
by: Li, Chaochao, et al.
Published: (2025)
by: Li, Chaochao, et al.
Published: (2025)
Cascade-Free Mandarin Visual Speech Recognition via Semantic-Guided Cross-Representation Alignment
by: Yang, Lei, et al.
Published: (2026)
by: Yang, Lei, et al.
Published: (2026)
ORXE: Orchestrating Experts for Dynamically Configurable Efficiency
by: Wang, Qingyuan, et al.
Published: (2025)
by: Wang, Qingyuan, et al.
Published: (2025)
MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
VALLR-Pin: Uncertainty-Factorized Visual Speech Recognition for Mandarin with Pinyin Guidance
by: Sun, Chang, et al.
Published: (2025)
by: Sun, Chang, et al.
Published: (2025)
Unlocking the Potential: Multi-task Deep Learning for Spaceborne Quantitative Monitoring of Fugitive Methane Plumes
by: Si, Guoxin, et al.
Published: (2024)
by: Si, Guoxin, et al.
Published: (2024)
CoCAViT: Compact Vision Transformer with Robust Global Coordination
by: Wang, Xuyang, et al.
Published: (2025)
by: Wang, Xuyang, et al.
Published: (2025)
VoxelNextFusion: A Simple, Unified and Effective Voxel Fusion Framework for Multi-Modal 3D Object Detection
by: Song, Ziying, et al.
Published: (2024)
by: Song, Ziying, et al.
Published: (2024)
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
Expressive yet Efficient Feature Expansion with Adaptive Cross-Hadamard Products
by: Zhang, Xuyang, et al.
Published: (2025)
by: Zhang, Xuyang, et al.
Published: (2025)
FGU3R: Fine-Grained Fusion via Unified 3D Representation for Multimodal 3D Object Detection
by: Zhang, Guoxin, et al.
Published: (2025)
by: Zhang, Guoxin, et al.
Published: (2025)
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
by: Huang, Binhua, et al.
Published: (2025)
by: Huang, Binhua, et al.
Published: (2025)
Any-Optical-Model: A Universal Foundation Model for Optical Remote Sensing
by: Li, Xuyang, et al.
Published: (2025)
by: Li, Xuyang, et al.
Published: (2025)
AsymLoRA: Harmonizing Data Conflicts and Commonalities in MLLMs
by: Wei, Xuyang, et al.
Published: (2025)
by: Wei, Xuyang, et al.
Published: (2025)
LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
3SHNet: Boosting Image-Sentence Retrieval via Visual Semantic-Spatial Self-Highlighting
by: Ge, Xuri, et al.
Published: (2024)
by: Ge, Xuri, et al.
Published: (2024)
VGDiffZero: Text-to-image Diffusion Models Can Be Zero-shot Visual Grounders
by: Liu, Xuyang, et al.
Published: (2023)
by: Liu, Xuyang, et al.
Published: (2023)
V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models
by: Lin, Xinying, et al.
Published: (2026)
by: Lin, Xinying, et al.
Published: (2026)
Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
Segmentation-aware Prior Assisted Joint Global Information Aggregated 3D Building Reconstruction
by: Peng, Hongxin, et al.
Published: (2024)
by: Peng, Hongxin, et al.
Published: (2024)
TinyDrop: Tiny Model Guided Token Dropping for Vision Transformers
by: Wang, Guoxin, et al.
Published: (2025)
by: Wang, Guoxin, et al.
Published: (2025)
Clore: Interactive Pathology Image Segmentation with Click-based Local Refinement
by: Wang, Tiantong, et al.
Published: (2026)
by: Wang, Tiantong, et al.
Published: (2026)
Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
Self-Supervised Animal Identification for Long Videos
by: Fang, Xuyang, et al.
Published: (2026)
by: Fang, Xuyang, et al.
Published: (2026)
8-Calves Image dataset
by: Fang, Xuyang, et al.
Published: (2025)
by: Fang, Xuyang, et al.
Published: (2025)
FlexiMo: A Flexible Remote Sensing Foundation Model
by: Li, Xuyang, et al.
Published: (2025)
by: Li, Xuyang, et al.
Published: (2025)
Similar Items
-
JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation
by: Cao, Xuyang, et al.
Published: (2024) -
JoyType: A Robust Design for Multilingual Visual Text Creation
by: Li, Chao, et al.
Published: (2024) -
Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
by: Xu, Mingwang, et al.
Published: (2024) -
Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation
by: Cui, Jiahao, et al.
Published: (2024) -
Hallo4: High-Fidelity Dynamic Portrait Animation via Direct Preference Optimization
by: Cui, Jiahao, et al.
Published: (2025)