Gespeichert in:
| Hauptverfasser: | Xu, Haohang, Chen, Longyu, Zhang, Yichen, Ding, Shuangrui, Zhang, Zhipeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2501.13349 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Masked Autoencoders are Robust Data Augmentors
von: Xu, Haohang, et al.
Veröffentlicht: (2022)
von: Xu, Haohang, et al.
Veröffentlicht: (2022)
Betrayed by Attention: A Simple yet Effective Approach for Self-supervised Video Object Segmentation
von: Ding, Shuangrui, et al.
Veröffentlicht: (2023)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2023)
EMO-X: Efficient Multi-Person Pose and Shape Estimation in One-Stage
von: Jian, Haohang, et al.
Veröffentlicht: (2025)
von: Jian, Haohang, et al.
Veröffentlicht: (2025)
Rethinking Image-to-Video Adaptation: An Object-centric Perspective
von: Qian, Rui, et al.
Veröffentlicht: (2024)
von: Qian, Rui, et al.
Veröffentlicht: (2024)
Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in Videos
von: Qian, Rui, et al.
Veröffentlicht: (2023)
von: Qian, Rui, et al.
Veröffentlicht: (2023)
MSF-Net: Multi-Stage Feature Extraction and Fusion for Robust Photometric Stereo
von: Qin, Shiyu, et al.
Veröffentlicht: (2025)
von: Qin, Shiyu, et al.
Veröffentlicht: (2025)
MSF-Mamba: Motion-aware State Fusion Mamba for Efficient Micro-Gesture Recognition
von: Li, Deng, et al.
Veröffentlicht: (2025)
von: Li, Deng, et al.
Veröffentlicht: (2025)
Streaming Long Video Understanding with Large Language Models
von: Qian, Rui, et al.
Veröffentlicht: (2024)
von: Qian, Rui, et al.
Veröffentlicht: (2024)
2nd Place Report of MOSEv2 Challenge 2025: Concept Guided Video Object Segmentation via SeC
von: Zhang, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhixiong, et al.
Veröffentlicht: (2025)
Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models
von: Lu, Tianyi, et al.
Veröffentlicht: (2023)
von: Lu, Tianyi, et al.
Veröffentlicht: (2023)
Explore In-Context Segmentation via Latent Diffusion Models
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
Reliable and Efficient Concept Erasure of Text-to-Image Diffusion Models
von: Gong, Chao, et al.
Veröffentlicht: (2024)
von: Gong, Chao, et al.
Veröffentlicht: (2024)
Training-Free Sketch-Guided Diffusion with Latent Optimization
von: Ding, Sandra Zhang, et al.
Veröffentlicht: (2024)
von: Ding, Sandra Zhang, et al.
Veröffentlicht: (2024)
SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction
von: Zhang, Zhixiong, et al.
Veröffentlicht: (2026)
von: Zhang, Zhixiong, et al.
Veröffentlicht: (2026)
Forest Before Trees: Latent Superposition for Efficient Visual Reasoning
von: Wang, Yubo, et al.
Veröffentlicht: (2026)
von: Wang, Yubo, et al.
Veröffentlicht: (2026)
Towards Dense and Accurate Radar Perception Via Efficient Cross-Modal Diffusion Model
von: Zhang, Ruibin, et al.
Veröffentlicht: (2024)
von: Zhang, Ruibin, et al.
Veröffentlicht: (2024)
RAFT-MSF++: Temporal Geometry-Motion Feature Fusion for Self-Supervised Monocular Scene Flow
von: Sun, Xunpei, et al.
Veröffentlicht: (2026)
von: Sun, Xunpei, et al.
Veröffentlicht: (2026)
FineEdit: Fine-Grained Image Edit with Bounding Box Guidance
von: Xu, Haohang, et al.
Veröffentlicht: (2026)
von: Xu, Haohang, et al.
Veröffentlicht: (2026)
Decoupling Complexity from Scale in Latent Diffusion Model
von: Zhong, Tianxiong, et al.
Veröffentlicht: (2025)
von: Zhong, Tianxiong, et al.
Veröffentlicht: (2025)
Dual Diffusion Models for Multi-modal Guided 3D Avatar Generation
von: Li, Hong, et al.
Veröffentlicht: (2026)
von: Li, Hong, et al.
Veröffentlicht: (2026)
EmoDiffusion: Enhancing Emotional 3D Facial Animation with Latent Diffusion Models
von: Zhang, Yixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2025)
dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
von: Xin, Yi, et al.
Veröffentlicht: (2025)
von: Xin, Yi, et al.
Veröffentlicht: (2025)
Diffusion As Self-Distillation: End-to-End Latent Diffusion In One Model
von: Wang, Xiyuan, et al.
Veröffentlicht: (2025)
von: Wang, Xiyuan, et al.
Veröffentlicht: (2025)
Dual3D: Efficient and Consistent Text-to-3D Generation with Dual-mode Multi-view Latent Diffusion
von: Li, Xinyang, et al.
Veröffentlicht: (2024)
von: Li, Xinyang, et al.
Veröffentlicht: (2024)
LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
von: Li, Chunyu, et al.
Veröffentlicht: (2024)
von: Li, Chunyu, et al.
Veröffentlicht: (2024)
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
von: Chen, Kerui, et al.
Veröffentlicht: (2026)
von: Chen, Kerui, et al.
Veröffentlicht: (2026)
Robust Video-Based Pothole Detection and Area Estimation for Intelligent Vehicles with Depth Map and Kalman Smoothing
von: Wang, Dehao, et al.
Veröffentlicht: (2025)
von: Wang, Dehao, et al.
Veröffentlicht: (2025)
HMARK: Radioactive Multi-Bit Semantic-Latent Watermarking for Diffusion Models
von: Li, Kexin, et al.
Veröffentlicht: (2025)
von: Li, Kexin, et al.
Veröffentlicht: (2025)
Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
von: Qian, Rui, et al.
Veröffentlicht: (2025)
von: Qian, Rui, et al.
Veröffentlicht: (2025)
LatentCRF: Continuous CRF for Efficient Latent Diffusion
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
Canonical Latent Representations in Conditional Diffusion Models
von: Xu, Yitao, et al.
Veröffentlicht: (2025)
von: Xu, Yitao, et al.
Veröffentlicht: (2025)
Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusion
von: Chen, Zhifei, et al.
Veröffentlicht: (2024)
von: Chen, Zhifei, et al.
Veröffentlicht: (2024)
An Efficient 3D Latent Diffusion Model for T1-contrast Enhanced MRI Generation
von: Eidex, Zach, et al.
Veröffentlicht: (2025)
von: Eidex, Zach, et al.
Veröffentlicht: (2025)
LAFR: Efficient Diffusion-based Blind Face Restoration via Latent Codebook Alignment Adapter
von: Li, Runyi, et al.
Veröffentlicht: (2025)
von: Li, Runyi, et al.
Veröffentlicht: (2025)
Prompt-Guided Dual Latent Steering for Inversion Problems
von: Wu, Yichen, et al.
Veröffentlicht: (2025)
von: Wu, Yichen, et al.
Veröffentlicht: (2025)
Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving
von: Zhang, Tianrui, et al.
Veröffentlicht: (2025)
von: Zhang, Tianrui, et al.
Veröffentlicht: (2025)
Efficient Generation of Targeted and Transferable Adversarial Examples for Vision-Language Models Via Diffusion Models
von: Guo, Qi, et al.
Veröffentlicht: (2024)
von: Guo, Qi, et al.
Veröffentlicht: (2024)
Multi-Garment Customized Model Generation
von: Liu, Yichen, et al.
Veröffentlicht: (2024)
von: Liu, Yichen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Masked Autoencoders are Robust Data Augmentors
von: Xu, Haohang, et al.
Veröffentlicht: (2022) -
Betrayed by Attention: A Simple yet Effective Approach for Self-supervised Video Object Segmentation
von: Ding, Shuangrui, et al.
Veröffentlicht: (2023) -
EMO-X: Efficient Multi-Person Pose and Shape Estimation in One-Stage
von: Jian, Haohang, et al.
Veröffentlicht: (2025) -
Rethinking Image-to-Video Adaptation: An Object-centric Perspective
von: Qian, Rui, et al.
Veröffentlicht: (2024) -
Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in Videos
von: Qian, Rui, et al.
Veröffentlicht: (2023)