From Sora What We Can See: A Survey of Text-to-Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Rui, Zhang, Yumin, Shah, Tejal, Sun, Jiahao, Zhang, Shuoying, Li, Wenqi, Duan, Haoran, Wei, Bo, Ranjan, Rajiv |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exemplar-condensed Federated Class-incremental Learning
by: Sun, Rui, et al.
Published: (2024)
by: Sun, Rui, et al.
Published: (2024)
ExactDreamer: High-Fidelity Text-to-3D Content Creation via Exact Score Matching
by: Zhang, Yumin, et al.
Published: (2024)
by: Zhang, Yumin, et al.
Published: (2024)
FedSCA: Federated Tuning with Similarity-guided Collaborative Aggregation for Heterogeneous Medical Image Segmentation
by: Zhang, Yumin, et al.
Published: (2025)
by: Zhang, Yumin, et al.
Published: (2025)
Rehearsal-free Federated Domain-incremental Learning
by: Sun, Rui, et al.
Published: (2024)
by: Sun, Rui, et al.
Published: (2024)
Dreamer XL: Towards High-Resolution Text-to-3D Generation via Trajectory Score Matching
by: Miao, Xingyu, et al.
Published: (2024)
by: Miao, Xingyu, et al.
Published: (2024)
AniSora: Exploring the Frontiers of Animation Video Generation in the Sora Era
by: Jiang, Yudong, et al.
Published: (2024)
by: Jiang, Yudong, et al.
Published: (2024)
Sora as a World Model? A Complete Survey on Text-to-Video Generation
by: Puspitasari, Fachrina Dewi, et al.
Published: (2024)
by: Puspitasari, Fachrina Dewi, et al.
Published: (2024)
Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields
by: Miao, Xingyu, et al.
Published: (2025)
by: Miao, Xingyu, et al.
Published: (2025)
D2Fusion: Dual-domain Fusion with Feature Superposition for Deepfake Detection
by: Qiu, Xueqi, et al.
Published: (2025)
by: Qiu, Xueqi, et al.
Published: (2025)
Time Blindness: Why Video-Language Models Can't See What Humans Can?
by: Upadhyay, Ujjwal, et al.
Published: (2025)
by: Upadhyay, Ujjwal, et al.
Published: (2025)
Sora Detector: A Unified Hallucination Detection for Large Text-to-Video Models
by: Chu, Zhixuan, et al.
Published: (2024)
by: Chu, Zhixuan, et al.
Published: (2024)
Sora Generates Videos with Stunning Geometrical Consistency
by: Li, Xuanyi, et al.
Published: (2024)
by: Li, Xuanyi, et al.
Published: (2024)
VEAT Quantifies Implicit Associations in Text-to-Video Generator Sora and Reveals Challenges in Bias Mitigation
by: Sun, Yongxu, et al.
Published: (2026)
by: Sun, Yongxu, et al.
Published: (2026)
Dataset Distillation-based Hybrid Federated Learning on Non-IID Data
by: Shi, Xiufang, et al.
Published: (2024)
by: Shi, Xiufang, et al.
Published: (2024)
RobustSora: De-Watermarked Benchmark for Robust AI-Generated Video Detection
by: Wang, Zhuo, et al.
Published: (2025)
by: Wang, Zhuo, et al.
Published: (2025)
Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking
by: Su, Zihan, et al.
Published: (2025)
by: Su, Zihan, et al.
Published: (2025)
Vision Language Models See What You Want but not What You See
by: Gao, Qingying, et al.
Published: (2024)
by: Gao, Qingying, et al.
Published: (2024)
Gender Bias in Text-to-Video Generation Models: A case study of Sora
by: Nadeem, Mohammad, et al.
Published: (2024)
by: Nadeem, Mohammad, et al.
Published: (2024)
Tell What You Hear From What You See -- Video to Audio Generation Through Text
by: Liu, Xiulong, et al.
Published: (2024)
by: Liu, Xiulong, et al.
Published: (2024)
Do MLLMs See What We See? Analyzing Visualization Literacy Barriers in AI Systems
by: Mengli, et al.
Published: (2026)
by: Mengli, et al.
Published: (2026)
"Sora is Incredible and Scary": Emerging Governance Challenges of Text-to-Video Generative AI Models
by: Zhou, Kyrie Zhixuan, et al.
Published: (2024)
by: Zhou, Kyrie Zhixuan, et al.
Published: (2024)
Simple Visual Artifact Detection in Sora-Generated Videos
by: Sugiyama, Misora, et al.
Published: (2025)
by: Sugiyama, Misora, et al.
Published: (2025)
On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices
by: Kim, Bosung, et al.
Published: (2025)
by: Kim, Bosung, et al.
Published: (2025)
On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices
by: Kim, Bosung, et al.
Published: (2025)
by: Kim, Bosung, et al.
Published: (2025)
Sora OpenAI's Prelude: Social Media Perspectives on Sora OpenAI and the Future of AI Video Generation
by: Mogavi, Reza Hadi, et al.
Published: (2024)
by: Mogavi, Reza Hadi, et al.
Published: (2024)
Open-Sora Plan: Open-Source Large Video Generation Model
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
Ground What You See: Hallucination-Resistant MLLMs via Caption Feedback, Diversity-Aware Sampling, and Conflict Regularization
by: Pan, Miao, et al.
Published: (2026)
by: Pan, Miao, et al.
Published: (2026)
What Increases the Risk of Sleep Problems for Train Drivers? Evidence From Network Analysis
by: Fei Wang, et al.
Published: (2025)
by: Fei Wang, et al.
Published: (2025)
Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond
by: Zhu, Zheng, et al.
Published: (2024)
by: Zhu, Zheng, et al.
Published: (2024)
Can Current AI Models Count What We Mean, Not What They See? A Benchmark and Systematic Evaluation
by: Nguyen, Gia Khanh, et al.
Published: (2025)
by: Nguyen, Gia Khanh, et al.
Published: (2025)
SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset
by: Dai, Josef, et al.
Published: (2024)
by: Dai, Josef, et al.
Published: (2024)
A Visual-Inertial Motion Prior SLAM for Dynamic Environments
by: Sun, Weilong, et al.
Published: (2025)
by: Sun, Weilong, et al.
Published: (2025)
Video Game Addiction: What Can We Learn From a Media Neuroscience Perspective?
by: Britney Craighead
Published: (2015)
by: Britney Craighead
Published: (2015)
From Logits to Latents: Contrastive Representation Shaping for LLM Unlearning
by: Tang, Haoran, et al.
Published: (2026)
by: Tang, Haoran, et al.
Published: (2026)
Can We Talk Models Into Seeing the World Differently?
by: Gavrikov, Paul, et al.
Published: (2024)
by: Gavrikov, Paul, et al.
Published: (2024)
What Can We Learn from Harry Potter? An Exploratory Study of Visual Representation Learning from Atypical Videos
by: Sun, Qiyue, et al.
Published: (2025)
by: Sun, Qiyue, et al.
Published: (2025)
We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback
by: Choi, Minkyu, et al.
Published: (2025)
by: Choi, Minkyu, et al.
Published: (2025)
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
by: Sahipjohn, Neha, et al.
Published: (2024)
by: Sahipjohn, Neha, et al.
Published: (2024)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
by: Sun, Boyuan, et al.
Published: (2026)
by: Sun, Boyuan, et al.
Published: (2026)
Can There Be Post-Persons and What Can We Learn From Considering Their Possibility?
by: Ivars Neiders
Published: (2015)
by: Ivars Neiders
Published: (2015)
Similar Items
-
Exemplar-condensed Federated Class-incremental Learning
by: Sun, Rui, et al.
Published: (2024) -
ExactDreamer: High-Fidelity Text-to-3D Content Creation via Exact Score Matching
by: Zhang, Yumin, et al.
Published: (2024) -
FedSCA: Federated Tuning with Similarity-guided Collaborative Aggregation for Heterogeneous Medical Image Segmentation
by: Zhang, Yumin, et al.
Published: (2025) -
Rehearsal-free Federated Domain-incremental Learning
by: Sun, Rui, et al.
Published: (2024) -
Dreamer XL: Towards High-Resolution Text-to-3D Generation via Trajectory Score Matching
by: Miao, Xingyu, et al.
Published: (2024)