Code2Video: A Code-centric Paradigm for Educational Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yanzhe, Lin, Kevin Qinghong, Shou, Mike Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Paper2Video: Automatic Video Generation from Scientific Papers
von: Zhu, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhu, Zeyu, et al.
Veröffentlicht: (2025)
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
von: Hu, Siyuan, et al.
Veröffentlicht: (2025)
von: Hu, Siyuan, et al.
Veröffentlicht: (2025)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
Code2World: A GUI World Model via Renderable Code Generation
von: Zheng, Yuhao, et al.
Veröffentlicht: (2026)
von: Zheng, Yuhao, et al.
Veröffentlicht: (2026)
Secure & Personalized Music-to-Video Generation via CHARCHA
von: Agarwal, Mehul, et al.
Veröffentlicht: (2025)
von: Agarwal, Mehul, et al.
Veröffentlicht: (2025)
DiffMesh: A Motion-aware Diffusion Framework for Human Mesh Recovery from Videos
von: Zheng, Ce, et al.
Veröffentlicht: (2023)
von: Zheng, Ce, et al.
Veröffentlicht: (2023)
E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
von: Lin, Ronghao, et al.
Veröffentlicht: (2025)
von: Lin, Ronghao, et al.
Veröffentlicht: (2025)
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensing
von: Cai, Weitong, et al.
Veröffentlicht: (2026)
von: Cai, Weitong, et al.
Veröffentlicht: (2026)
"I Can See Forever!": Evaluating Real-time VideoLLMs for Assisting Individuals with Visual Impairments
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
FastPerson: Enhancing Video Learning through Effective Video Summarization that Preserves Linguistic and Visual Contexts
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
Factorized Learning for Temporally Grounded Video-Language Models
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
Inkspire: Supporting Design Exploration with Generative AI through Analogical Sketching
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2025)
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2025)
A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models
von: Tsangko, Iosif, et al.
Veröffentlicht: (2026)
von: Tsangko, Iosif, et al.
Veröffentlicht: (2026)
On Semiotic-Grounded Interpretive Evaluation of Generative Art
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2026)
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2026)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
Generative AI for Video Trailer Synthesis: From Extractive Heuristics to Autoregressive Creativity
von: Dharmaratnakar, Abhishek, et al.
Veröffentlicht: (2026)
von: Dharmaratnakar, Abhishek, et al.
Veröffentlicht: (2026)
SkinGEN: an Explainable Dermatology Diagnosis-to-Generation Framework with Interactive Vision-Language Models
von: Lin, Bo, et al.
Veröffentlicht: (2024)
von: Lin, Bo, et al.
Veröffentlicht: (2024)
Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
von: He, Xu, et al.
Veröffentlicht: (2024)
von: He, Xu, et al.
Veröffentlicht: (2024)
VideoMap: Supporting Video Editing Exploration, Brainstorming, and Prototyping in the Latent Space
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2022)
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2022)
G3R: Generating Rich and Fine-grained mmWave Radar Data from 2D Videos for Generalized Gesture Recognition
von: Deng, Kaikai, et al.
Veröffentlicht: (2024)
von: Deng, Kaikai, et al.
Veröffentlicht: (2024)
SVFAP: Self-supervised Video Facial Affect Perceiver
von: Sun, Licai, et al.
Veröffentlicht: (2023)
von: Sun, Licai, et al.
Veröffentlicht: (2023)
Seeing, Hearing, and Knowing Together: Multimodal Strategies in Deepfake Videos Detection
von: Chen, Chen, et al.
Veröffentlicht: (2026)
von: Chen, Chen, et al.
Veröffentlicht: (2026)
Learning High-Quality Navigation and Zooming on Omnidirectional Images in Virtual Reality
von: Cao, Zidong, et al.
Veröffentlicht: (2024)
von: Cao, Zidong, et al.
Veröffentlicht: (2024)
MindCine: Multimodal EEG-to-Video Reconstruction with Large-Scale Pretrained Models
von: Zhou, Tian-Yi, et al.
Veröffentlicht: (2026)
von: Zhou, Tian-Yi, et al.
Veröffentlicht: (2026)
Videogenic: Identifying Highlight Moments in Videos with Professional Photographs as a Prior
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2022)
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2022)
Computer-Use Agents as Judges for Generative User Interface
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
Editing Physiological Signals in Videos Using Latent Representations
von: Zhou, Tianwen, et al.
Veröffentlicht: (2025)
von: Zhou, Tianwen, et al.
Veröffentlicht: (2025)
MindCross: Fast New Subject Adaptation with Limited Data for Cross-subject Video Reconstruction from Brain Signals
von: Liu, Xuan-Hao, et al.
Veröffentlicht: (2025)
von: Liu, Xuan-Hao, et al.
Veröffentlicht: (2025)
Panonut360: A Head and Eye Tracking Dataset for Panoramic Video
von: Xu, Yutong, et al.
Veröffentlicht: (2024)
von: Xu, Yutong, et al.
Veröffentlicht: (2024)
Coral Model Generation from Single Images for Virtual Reality Applications
von: Fu, Jie, et al.
Veröffentlicht: (2024)
von: Fu, Jie, et al.
Veröffentlicht: (2024)
ReactMotion: Generating Reactive Listener Motions from Speaker Utterance
von: Luo, Cheng, et al.
Veröffentlicht: (2026)
von: Luo, Cheng, et al.
Veröffentlicht: (2026)
Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functions
von: Rakesh, Vineet Kumar, et al.
Veröffentlicht: (2025)
von: Rakesh, Vineet Kumar, et al.
Veröffentlicht: (2025)
AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe
von: Cole, Adam, et al.
Veröffentlicht: (2026)
von: Cole, Adam, et al.
Veröffentlicht: (2026)
LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing
von: Wang, Bryan, et al.
Veröffentlicht: (2024)
von: Wang, Bryan, et al.
Veröffentlicht: (2024)
Photoreal Scene Reconstruction from an Egocentric Device
von: Lv, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Lv, Zhaoyang, et al.
Veröffentlicht: (2025)
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2025)
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
von: Ki, Taekyung, et al.
Veröffentlicht: (2026)
von: Ki, Taekyung, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Paper2Video: Automatic Video Generation from Scientific Papers
von: Zhu, Zeyu, et al.
Veröffentlicht: (2025) -
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
von: Hu, Siyuan, et al.
Veröffentlicht: (2025) -
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026) -
Code2World: A GUI World Model via Renderable Code Generation
von: Zheng, Yuhao, et al.
Veröffentlicht: (2026) -
Secure & Personalized Music-to-Video Generation via CHARCHA
von: Agarwal, Mehul, et al.
Veröffentlicht: (2025)