On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Bosung, Lee, Kyuhwan, Jeong, Isu, Cheon, Jungmin, Lee, Yeojin, Lee, Seulki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices
von: Kim, Bosung, et al.
Veröffentlicht: (2025)
von: Kim, Bosung, et al.
Veröffentlicht: (2025)
Designing Extremely Memory-Efficient CNNs for On-device Vision Tasks
von: Lee, Jaewook, et al.
Veröffentlicht: (2024)
von: Lee, Jaewook, et al.
Veröffentlicht: (2024)
Grid Diffusion Models for Text-to-Video Generation
von: Lee, Taegyeong, et al.
Veröffentlicht: (2024)
von: Lee, Taegyeong, et al.
Veröffentlicht: (2024)
DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion
von: Hwang, Geunmin, et al.
Veröffentlicht: (2025)
von: Hwang, Geunmin, et al.
Veröffentlicht: (2025)
FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
von: Kim, Seungwook, et al.
Veröffentlicht: (2025)
von: Kim, Seungwook, et al.
Veröffentlicht: (2025)
Harnessing the Power of Training-Free Techniques in Text-to-2D Generation for Text-to-3D Generation via Score Distillation Sampling
von: Lee, Junhong, et al.
Veröffentlicht: (2025)
von: Lee, Junhong, et al.
Veröffentlicht: (2025)
Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking
von: Su, Zihan, et al.
Veröffentlicht: (2025)
von: Su, Zihan, et al.
Veröffentlicht: (2025)
CountCluster: Training-Free Object Quantity Guidance with Cross-Attention Map Clustering for Text-to-Image Generation
von: Lee, Joohyeon, et al.
Veröffentlicht: (2025)
von: Lee, Joohyeon, et al.
Veröffentlicht: (2025)
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
von: Ko, Jungmin, et al.
Veröffentlicht: (2026)
von: Ko, Jungmin, et al.
Veröffentlicht: (2026)
Scribble-Guided Diffusion for Training-free Text-to-Image Generation
von: Lee, Seonho, et al.
Veröffentlicht: (2024)
von: Lee, Seonho, et al.
Veröffentlicht: (2024)
MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices
von: Zhao, Yang, et al.
Veröffentlicht: (2023)
von: Zhao, Yang, et al.
Veröffentlicht: (2023)
Sora as a World Model? A Complete Survey on Text-to-Video Generation
von: Puspitasari, Fachrina Dewi, et al.
Veröffentlicht: (2024)
von: Puspitasari, Fachrina Dewi, et al.
Veröffentlicht: (2024)
Sora Generates Videos with Stunning Geometrical Consistency
von: Li, Xuanyi, et al.
Veröffentlicht: (2024)
von: Li, Xuanyi, et al.
Veröffentlicht: (2024)
Good Noise Makes Good Edits: A Training-Free Diffusion-Based Video Editing with Image and Text Prompts
von: Choi, Saemee, et al.
Veröffentlicht: (2025)
von: Choi, Saemee, et al.
Veröffentlicht: (2025)
ISAC: Training-Free Instance-to-Semantic Attention Control for Improving Multi-Instance Generation
von: Jo, Sanghyun, et al.
Veröffentlicht: (2025)
von: Jo, Sanghyun, et al.
Veröffentlicht: (2025)
Simple Visual Artifact Detection in Sora-Generated Videos
von: Sugiyama, Misora, et al.
Veröffentlicht: (2025)
von: Sugiyama, Misora, et al.
Veröffentlicht: (2025)
Group-wise Scaling and Orthogonal Decomposition for Domain-Invariant Feature Extraction in Face Anti-Spoofing
von: Jung, Seungjin, et al.
Veröffentlicht: (2025)
von: Jung, Seungjin, et al.
Veröffentlicht: (2025)
Local Representative Token Guided Merging for Text-to-Image Generation
von: Lee, Min-Jeong, et al.
Veröffentlicht: (2025)
von: Lee, Min-Jeong, et al.
Veröffentlicht: (2025)
Read, Watch and Scream! Sound Generation from Text and Video
von: Jeong, Yujin, et al.
Veröffentlicht: (2024)
von: Jeong, Yujin, et al.
Veröffentlicht: (2024)
InsideOut: Integrated RGB-Radiative Gaussian Splatting for Comprehensive 3D Object Representation
von: Lee, Jungmin, et al.
Veröffentlicht: (2025)
von: Lee, Jungmin, et al.
Veröffentlicht: (2025)
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
Active Diffusion Matching: Score-based Iterative Alignment of Cross-Modal Retinal Images
von: Lee, Kanggeon, et al.
Veröffentlicht: (2026)
von: Lee, Kanggeon, et al.
Veröffentlicht: (2026)
From Sora What We Can See: A Survey of Text-to-Video Generation
von: Sun, Rui, et al.
Veröffentlicht: (2024)
von: Sun, Rui, et al.
Veröffentlicht: (2024)
Infinite-Story: A Training-Free Consistent Text-to-Image Generation
von: Park, Jihun, et al.
Veröffentlicht: (2025)
von: Park, Jihun, et al.
Veröffentlicht: (2025)
Video Diffusion Models are Strong Video Inpainter
von: Lee, Minhyeok, et al.
Veröffentlicht: (2024)
von: Lee, Minhyeok, et al.
Veröffentlicht: (2024)
MobileViCLIP: An Efficient Video-Text Model for Mobile Devices
von: Yang, Min, et al.
Veröffentlicht: (2025)
von: Yang, Min, et al.
Veröffentlicht: (2025)
Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
von: Lee, Dohun, et al.
Veröffentlicht: (2025)
von: Lee, Dohun, et al.
Veröffentlicht: (2025)
vid-TLDR: Training Free Token merging for Light-weight Video Transformer
von: Choi, Joonmyung, et al.
Veröffentlicht: (2024)
von: Choi, Joonmyung, et al.
Veröffentlicht: (2024)
Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis
von: Han, Woojung, et al.
Veröffentlicht: (2025)
von: Han, Woojung, et al.
Veröffentlicht: (2025)
MCoT-RE: Multi-Faceted Chain-of-Thought and Re-Ranking for Training-Free Zero-Shot Composed Image Retrieval
von: Park, Jeong-Woo, et al.
Veröffentlicht: (2025)
von: Park, Jeong-Woo, et al.
Veröffentlicht: (2025)
ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation
von: Peng, Bo, et al.
Veröffentlicht: (2023)
von: Peng, Bo, et al.
Veröffentlicht: (2023)
MEVG: Multi-event Video Generation with Text-to-Video Models
von: Oh, Gyeongrok, et al.
Veröffentlicht: (2023)
von: Oh, Gyeongrok, et al.
Veröffentlicht: (2023)
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
von: Chen, Tong, et al.
Veröffentlicht: (2024)
von: Chen, Tong, et al.
Veröffentlicht: (2024)
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
von: Han, Su Ho, et al.
Veröffentlicht: (2025)
von: Han, Su Ho, et al.
Veröffentlicht: (2025)
Sora Detector: A Unified Hallucination Detection for Large Text-to-Video Models
von: Chu, Zhixuan, et al.
Veröffentlicht: (2024)
von: Chu, Zhixuan, et al.
Veröffentlicht: (2024)
Open-Sora: Democratizing Efficient Video Production for All
von: Zheng, Zangwei, et al.
Veröffentlicht: (2024)
von: Zheng, Zangwei, et al.
Veröffentlicht: (2024)
Improving Unsupervised Video Object Segmentation via Fake Flow Generation
von: Cho, Suhwan, et al.
Veröffentlicht: (2024)
von: Cho, Suhwan, et al.
Veröffentlicht: (2024)
Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion Transformer
von: Lee, Dong In, et al.
Veröffentlicht: (2025)
von: Lee, Dong In, et al.
Veröffentlicht: (2025)
DESSERT: Diffusion-based Event-driven Single-frame Synthesis via Residual Training
von: Kong, Jiyun, et al.
Veröffentlicht: (2025)
von: Kong, Jiyun, et al.
Veröffentlicht: (2025)
Reangle-A-Video: 4D Video Generation as Video-to-Video Translation
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2025)
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices
von: Kim, Bosung, et al.
Veröffentlicht: (2025) -
Designing Extremely Memory-Efficient CNNs for On-device Vision Tasks
von: Lee, Jaewook, et al.
Veröffentlicht: (2024) -
Grid Diffusion Models for Text-to-Video Generation
von: Lee, Taegyeong, et al.
Veröffentlicht: (2024) -
DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion
von: Hwang, Geunmin, et al.
Veröffentlicht: (2025) -
FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
von: Kim, Seungwook, et al.
Veröffentlicht: (2025)