Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Peng, Liu, Yuan, Long, Xiaoxiao, Zhang, Feihu, Lin, Cheng, Li, Mengfei, Qi, Xingqun, Zhang, Shanghang, Luo, Wenhan, Tan, Ping, Wang, Wenping, Liu, Qifeng, Guo, Yike |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion
by: Qi, Xingqun, et al.
Published: (2025)
by: Qi, Xingqun, et al.
Published: (2025)
Weakly-Supervised Emotion Transition Learning for Diverse 3D Co-speech Gesture Generation
by: Qi, Xingqun, et al.
Published: (2023)
by: Qi, Xingqun, et al.
Published: (2023)
Multi-View Large Reconstruction Model via Geometry-Aware Positional Encoding and Attention
by: Li, Mengfei, et al.
Published: (2024)
by: Li, Mengfei, et al.
Published: (2024)
CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild
by: Qi, Xingqun, et al.
Published: (2024)
by: Qi, Xingqun, et al.
Published: (2024)
CMD: Controllable Multiview Diffusion for 3D Editing and Progressive Generation
by: Li, Peng, et al.
Published: (2025)
by: Li, Peng, et al.
Published: (2025)
PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit Remeshing
by: Li, Peng, et al.
Published: (2024)
by: Li, Peng, et al.
Published: (2024)
MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
by: Chi, Xiaowei, et al.
Published: (2024)
by: Chi, Xiaowei, et al.
Published: (2024)
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
by: Chi, Xiaowei, et al.
Published: (2023)
by: Chi, Xiaowei, et al.
Published: (2023)
On $P$-partitions Extended by Two-Rowed Plane Partitions
by: Li, Jingxuan, et al.
Published: (2024)
by: Li, Jingxuan, et al.
Published: (2024)
VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
SyncDreamer: Generating Multiview-consistent Images from a Single-view Image
by: Liu, Yuan, et al.
Published: (2023)
by: Liu, Yuan, et al.
Published: (2023)
UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass
by: Li, Mengfei, et al.
Published: (2026)
by: Li, Mengfei, et al.
Published: (2026)
EVA: An Embodied World Model for Future Video Anticipation
by: Chi, Xiaowei, et al.
Published: (2024)
by: Chi, Xiaowei, et al.
Published: (2024)
HiPrompt: Tuning-free Higher-Resolution Generation with Hierarchical MLLM Prompts
by: Liu, Xinyu, et al.
Published: (2024)
by: Liu, Xinyu, et al.
Published: (2024)
RustNeRF: Robust Neural Radiance Field with Low-Quality Images
by: Li, Mengfei, et al.
Published: (2024)
by: Li, Mengfei, et al.
Published: (2024)
A Diffusion Model Translator for Efficient Image-to-Image Translation
by: Xia, Mengfei, et al.
Published: (2025)
by: Xia, Mengfei, et al.
Published: (2025)
MVD$^2$: Efficient Multiview 3D Reconstruction for Multiview Diffusion
by: Zheng, Xin-Yang, et al.
Published: (2024)
by: Zheng, Xin-Yang, et al.
Published: (2024)
MV2UV: Generating High-quality UV Texture Maps with Multiview Prompts
by: Zhang, Zheng, et al.
Published: (2026)
by: Zhang, Zheng, et al.
Published: (2026)
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
by: Pu, Yifan, et al.
Published: (2024)
by: Pu, Yifan, et al.
Published: (2024)
Refining Few-Step Text-to-Multiview Diffusion via Reinforcement Learning
by: Zhang, Ziyi, et al.
Published: (2025)
by: Zhang, Ziyi, et al.
Published: (2025)
Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models
by: Wang, Zhiqiang, et al.
Published: (2026)
by: Wang, Zhiqiang, et al.
Published: (2026)
SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation
by: Liu, YaoYang, et al.
Published: (2026)
by: Liu, YaoYang, et al.
Published: (2026)
CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editing
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
Rethinking Layer-wise Gaussian Noise Injection: Bridging Implicit Objectives and Privacy Budget Allocation
by: Tan, Qifeng, et al.
Published: (2025)
by: Tan, Qifeng, et al.
Published: (2025)
SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
SmokeNet: Efficient Smoke Segmentation Leveraging Multiscale Convolutions and Multiview Attention Mechanisms
by: Liu, Xuesong, et al.
Published: (2025)
by: Liu, Xuesong, et al.
Published: (2025)
A Row-wise Algorithm for Graph Realization
by: van der Hulst, Rolf, et al.
Published: (2024)
by: van der Hulst, Rolf, et al.
Published: (2024)
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
by: Yakun, Cui, et al.
Published: (2026)
by: Yakun, Cui, et al.
Published: (2026)
PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution
by: Du, Shian, et al.
Published: (2025)
by: Du, Shian, et al.
Published: (2025)
Multiview Scene Graph
by: Zhang, Juexiao, et al.
Published: (2024)
by: Zhang, Juexiao, et al.
Published: (2024)
VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction
by: Li, Shiying, et al.
Published: (2025)
by: Li, Shiying, et al.
Published: (2025)
Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
by: Zhang, Qizhe, et al.
Published: (2025)
by: Zhang, Qizhe, et al.
Published: (2025)
SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language Understanding
by: Cheng, Shuang, et al.
Published: (2025)
by: Cheng, Shuang, et al.
Published: (2025)
Part123: Part-aware 3D Reconstruction from a Single-view Image
by: Liu, Anran, et al.
Published: (2024)
by: Liu, Anran, et al.
Published: (2024)
Proof of a conjecture on graph polytope
by: Liu, Feihu
Published: (2024)
by: Liu, Feihu
Published: (2024)
Generating functions for the quotients of numerical semigroups
by: Liu, Feihu
Published: (2023)
by: Liu, Feihu
Published: (2023)
On quotients of numerical semigroups for almost arithmetic progressions
by: Liu, Feihu
Published: (2023)
by: Liu, Feihu
Published: (2023)
The Frobenius problem for a class of quotients of numerical semigroups
by: Liu, Feihu
Published: (2026)
by: Liu, Feihu
Published: (2026)
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
by: Li, Yuming, et al.
Published: (2025)
by: Li, Yuming, et al.
Published: (2025)
An Efficient Short Text Classification Model Based on Feature Shuffling and Attention‐Enhanced With Multiview Contrastive Learning
by: Fangbo Liu, et al.
Published: (2026)
by: Fangbo Liu, et al.
Published: (2026)
Similar Items
-
Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion
by: Qi, Xingqun, et al.
Published: (2025) -
Weakly-Supervised Emotion Transition Learning for Diverse 3D Co-speech Gesture Generation
by: Qi, Xingqun, et al.
Published: (2023) -
Multi-View Large Reconstruction Model via Geometry-Aware Positional Encoding and Attention
by: Li, Mengfei, et al.
Published: (2024) -
CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild
by: Qi, Xingqun, et al.
Published: (2024) -
CMD: Controllable Multiview Diffusion for 3D Editing and Progressive Generation
by: Li, Peng, et al.
Published: (2025)