SAGE: Structure-Aware Generative Video Transitions between Diverse Clips
Fuente:
arXiv
Saved in:
| Main Authors: | Kan, Mia, Liu, Yilin, Mitra, Niloy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAGE: Structure‐Aware Generative Video Transitions between Diverse Clips
by: Mia Kan, et al.
Published: (2026)
by: Mia Kan, et al.
Published: (2026)
Animal Avatars: Reconstructing Animatable 3D Animals from Casual Videos
by: Sabathier, Remy, et al.
Published: (2024)
by: Sabathier, Remy, et al.
Published: (2024)
CHOIR: Contact-aware 4D Hand-Object Interaction Reconstruction
by: Xu, Hao, et al.
Published: (2026)
by: Xu, Hao, et al.
Published: (2026)
Beyond the Visible: Disocclusion-Aware Editing via Proxy Dynamic Graphs
by: Qi, Anran, et al.
Published: (2025)
by: Qi, Anran, et al.
Published: (2025)
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2025)
by: Kizil, Muhammed Burak, et al.
Published: (2025)
ProteusNeRF: Fast Lightweight NeRF Editing using 3D-Aware Image Context
by: Wang, Binglun, et al.
Published: (2023)
by: Wang, Binglun, et al.
Published: (2023)
WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
by: Fang, Shaoheng, et al.
Published: (2025)
by: Fang, Shaoheng, et al.
Published: (2025)
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2026)
by: Kizil, Muhammed Burak, et al.
Published: (2026)
ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion
by: Sabathier, Remy, et al.
Published: (2026)
by: Sabathier, Remy, et al.
Published: (2026)
ViSAGE @ NTIRE 2026 Challenge on Video Saliency Prediction
by: Wang, Kun, et al.
Published: (2026)
by: Wang, Kun, et al.
Published: (2026)
JOG3R: Towards 3D-Consistent Video Generators
by: Huang, Chun-Hao Paul, et al.
Published: (2025)
by: Huang, Chun-Hao Paul, et al.
Published: (2025)
From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
by: Sun, Guangyu, et al.
Published: (2025)
by: Sun, Guangyu, et al.
Published: (2025)
Question-Aware Evidence Ledgers for Video Relational Reasoning
by: Ou, Yilin, et al.
Published: (2026)
by: Ou, Yilin, et al.
Published: (2026)
SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation
by: Shukla, Tripti, et al.
Published: (2026)
by: Shukla, Tripti, et al.
Published: (2026)
VideoSAGE: Video Summarization with Graph Representation Learning
by: Chaves, Jose M. Rojas, et al.
Published: (2024)
by: Chaves, Jose M. Rojas, et al.
Published: (2024)
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
by: Jeong, Hyeonho, et al.
Published: (2024)
by: Jeong, Hyeonho, et al.
Published: (2024)
Magic 1-For-1: Generating One Minute Video Clips within One Minute
by: Yi, Hongwei, et al.
Published: (2025)
by: Yi, Hongwei, et al.
Published: (2025)
Leveraging VLM-Based Pipelines to Annotate 3D Objects
by: Kabra, Rishabh, et al.
Published: (2023)
by: Kabra, Rishabh, et al.
Published: (2023)
From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries
by: Hsu, Joy, et al.
Published: (2025)
by: Hsu, Joy, et al.
Published: (2025)
SMF: Template-free and Rig-free Animation Transfer using Kinetic Codes
by: Muralikrishnan, Sanjeev, et al.
Published: (2025)
by: Muralikrishnan, Sanjeev, et al.
Published: (2025)
Diffusion 3D Features (Diff3F): Decorating Untextured Shapes with Distilled Semantic Features
by: Dutt, Niladri Shekhar, et al.
Published: (2023)
by: Dutt, Niladri Shekhar, et al.
Published: (2023)
LIM: Large Interpolator Model for Dynamic Reconstruction
by: Sabathier, Remy, et al.
Published: (2025)
by: Sabathier, Remy, et al.
Published: (2025)
Consistency-Preserving Diverse Video Generation
by: Liu, Xinshuang, et al.
Published: (2026)
by: Liu, Xinshuang, et al.
Published: (2026)
SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias
by: Ye, Wenqian, et al.
Published: (2025)
by: Ye, Wenqian, et al.
Published: (2025)
RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos
by: Hannan, Tanveer, et al.
Published: (2023)
by: Hannan, Tanveer, et al.
Published: (2023)
FlairGPT: Repurposing LLMs for Interior Designs
by: Littlefair, Gabrielle, et al.
Published: (2025)
by: Littlefair, Gabrielle, et al.
Published: (2025)
SAGE: Segment-Aware Gloss-Free Encoding for Token-Efficient Sign Language Translation
by: Low, JianHe, et al.
Published: (2025)
by: Low, JianHe, et al.
Published: (2025)
SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
by: Jain, Jitesh, et al.
Published: (2025)
by: Jain, Jitesh, et al.
Published: (2025)
BLiSS: Bootstrapped Linear Shape Space
by: Muralikrishnan, Sanjeev, et al.
Published: (2023)
by: Muralikrishnan, Sanjeev, et al.
Published: (2023)
Video-GPT via Next Clip Diffusion
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions
by: Elmoghany, Mohamed, et al.
Published: (2026)
by: Elmoghany, Mohamed, et al.
Published: (2026)
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
by: Kim, Ji-Hyeon, et al.
Published: (2026)
by: Kim, Ji-Hyeon, et al.
Published: (2026)
From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding
by: Wang, Xiangfeng, et al.
Published: (2025)
by: Wang, Xiangfeng, et al.
Published: (2025)
Neural Semantic Surface Maps
by: Morreale, Luca, et al.
Published: (2023)
by: Morreale, Luca, et al.
Published: (2023)
SAGE: Saliency-Guided Contrastive Embeddings
by: Crum, Colton R., et al.
Published: (2025)
by: Crum, Colton R., et al.
Published: (2025)
Minimal Clips, Maximum Salience: Long Video Summarization via Key Moment Extraction
by: Pennec, Galann, et al.
Published: (2025)
by: Pennec, Galann, et al.
Published: (2025)
SuperGaussian: Repurposing Video Models for 3D Super Resolution
by: Shen, Yuan, et al.
Published: (2024)
by: Shen, Yuan, et al.
Published: (2024)
Neural Geometry Processing via Spherical Neural Surfaces
by: Williamson, Romy, et al.
Published: (2024)
by: Williamson, Romy, et al.
Published: (2024)
SAGE: Scalable Agentic 3D Scene Generation for Embodied AI
by: Xia, Hongchi, et al.
Published: (2026)
by: Xia, Hongchi, et al.
Published: (2026)
SAGE: Style-Adaptive Generalization for Privacy-Constrained Semantic Segmentation Across Domains
by: Li, Qingmei, et al.
Published: (2025)
by: Li, Qingmei, et al.
Published: (2025)
Similar Items
-
SAGE: Structure‐Aware Generative Video Transitions between Diverse Clips
by: Mia Kan, et al.
Published: (2026) -
Animal Avatars: Reconstructing Animatable 3D Animals from Casual Videos
by: Sabathier, Remy, et al.
Published: (2024) -
CHOIR: Contact-aware 4D Hand-Object Interaction Reconstruction
by: Xu, Hao, et al.
Published: (2026) -
Beyond the Visible: Disocclusion-Aware Editing via Proxy Dynamic Graphs
by: Qi, Anran, et al.
Published: (2025) -
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2025)