SG2VID: Scene Graphs Enable Fine-Grained Control for Video Synthesis
Fuente:
arXiv
Guardado en:
| Autores principales: | Sivakumar, Ssharvien Kumar, Frisch, Yannik, Ghazaei, Ghazal, Mukhopadhyay, Anirban |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SurGrID: Controllable Surgical Simulation via Scene Graph to Image Diffusion
por: Frisch, Yannik, et al.
Publicado: (2025)
por: Frisch, Yannik, et al.
Publicado: (2025)
SASVi -- Segment Any Surgical Video
por: Sivakumar, Ssharvien Kumar, et al.
Publicado: (2025)
por: Sivakumar, Ssharvien Kumar, et al.
Publicado: (2025)
CAT-SG: A Large Dynamic Scene Graph Dataset for Fine-Grained Understanding of Cataract Surgery
por: Holm, Felix, et al.
Publicado: (2025)
por: Holm, Felix, et al.
Publicado: (2025)
Frequency-Time Diffusion with Neural Cellular Automata
por: Kalkhof, John, et al.
Publicado: (2024)
por: Kalkhof, John, et al.
Publicado: (2024)
SANGRIA: Surgical Video Scene Graph Optimization for Surgical Workflow Prediction
por: Köksal, Çağhan, et al.
Publicado: (2024)
por: Köksal, Çağhan, et al.
Publicado: (2024)
GAUDA: Generative Adaptive Uncertainty-guided Diffusion-based Augmentation for Surgical Segmentation
por: Frisch, Yannik, et al.
Publicado: (2025)
por: Frisch, Yannik, et al.
Publicado: (2025)
SURGIVID: Annotation-Efficient Surgical Video Object Discovery
por: Köksal, Çağhan, et al.
Publicado: (2024)
por: Köksal, Çağhan, et al.
Publicado: (2024)
Watch and Learn: Leveraging Expert Knowledge and Language for Surgical Video Understanding
por: Gastager, David, et al.
Publicado: (2025)
por: Gastager, David, et al.
Publicado: (2025)
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
por: Holm, Felix, et al.
Publicado: (2025)
por: Holm, Felix, et al.
Publicado: (2025)
Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion
por: Rohrmoser, Nikolo, et al.
Publicado: (2026)
por: Rohrmoser, Nikolo, et al.
Publicado: (2026)
Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
por: Chu, Sanghyeok, et al.
Publicado: (2025)
por: Chu, Sanghyeok, et al.
Publicado: (2025)
SG-Bot: Object Rearrangement via Coarse-to-Fine Robotic Imagination on Scene Graphs
por: Zhai, Guangyao, et al.
Publicado: (2023)
por: Zhai, Guangyao, et al.
Publicado: (2023)
Deep Generative Models for 3D Medical Image Synthesis
por: Friedrich, Paul, et al.
Publicado: (2024)
por: Friedrich, Paul, et al.
Publicado: (2024)
T2SG: Traffic Topology Scene Graph for Topology Reasoning in Autonomous Driving
por: Lv, Changsheng, et al.
Publicado: (2024)
por: Lv, Changsheng, et al.
Publicado: (2024)
SG-Reg: Generalizable and Efficient Scene Graph Registration
por: Liu, Chuhao, et al.
Publicado: (2025)
por: Liu, Chuhao, et al.
Publicado: (2025)
VolDiT: Controllable Volumetric Medical Image Synthesis with Diffusion Transformers
por: Seyfarth, Marvin, et al.
Publicado: (2026)
por: Seyfarth, Marvin, et al.
Publicado: (2026)
CoPa-SG: Dense Scene Graphs with Parametric and Proto-Relations
por: Lorenz, Julian, et al.
Publicado: (2025)
por: Lorenz, Julian, et al.
Publicado: (2025)
SG-Adapter: Enhancing Text-to-Image Generation with Scene Graph Guidance
por: Shen, Guibao, et al.
Publicado: (2024)
por: Shen, Guibao, et al.
Publicado: (2024)
SG-NeRF: Neural Surface Reconstruction with Scene Graph Optimization
por: Chen, Yiyang, et al.
Publicado: (2024)
por: Chen, Yiyang, et al.
Publicado: (2024)
Efficient Remote Sensing Change Detection with Change State Space Models
por: Ghazaei, Elman, et al.
Publicado: (2025)
por: Ghazaei, Elman, et al.
Publicado: (2025)
Text-conditioned State Space Model For Domain-generalized Change Detection Visual Question Answering
por: Ghazaei, Elman, et al.
Publicado: (2025)
por: Ghazaei, Elman, et al.
Publicado: (2025)
CI-VID: A Coherent Interleaved Text-Video Dataset
por: Ju, Yiming, et al.
Publicado: (2025)
por: Ju, Yiming, et al.
Publicado: (2025)
ClimateVID -- Social Media Videos Analysis and Challenges Involved
por: Xu, Shiqi, et al.
Publicado: (2026)
por: Xu, Shiqi, et al.
Publicado: (2026)
XS-VID: An Extremely Small Video Object Detection Dataset
por: Guo, Jiahao, et al.
Publicado: (2024)
por: Guo, Jiahao, et al.
Publicado: (2024)
KeySG: Hierarchical Keyframe-Based 3D Scene Graphs
por: Werby, Abdelrhman, et al.
Publicado: (2025)
por: Werby, Abdelrhman, et al.
Publicado: (2025)
Robust SG-NeRF: Robust Scene Graph Aided Neural Surface Reconstruction
por: Gu, Yi, et al.
Publicado: (2024)
por: Gu, Yi, et al.
Publicado: (2024)
Semantic-E2VID: a Semantic-Enriched Paradigm for Event-to-Video Reconstruction
por: Wu, Jingqian, et al.
Publicado: (2025)
por: Wu, Jingqian, et al.
Publicado: (2025)
HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks
por: Ercan, Burak, et al.
Publicado: (2023)
por: Ercan, Burak, et al.
Publicado: (2023)
SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
por: Wang, Jiahao, et al.
Publicado: (2025)
por: Wang, Jiahao, et al.
Publicado: (2025)
Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow Matching
por: Hu, Xin, et al.
Publicado: (2026)
por: Hu, Xin, et al.
Publicado: (2026)
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
por: Zhang, Xu, et al.
Publicado: (2026)
por: Zhang, Xu, et al.
Publicado: (2026)
KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation
por: Wang, Xingrui, et al.
Publicado: (2025)
por: Wang, Xingrui, et al.
Publicado: (2025)
FastVID: Dynamic Density Pruning for Fast Video Large Language Models
por: Shen, Leqi, et al.
Publicado: (2025)
por: Shen, Leqi, et al.
Publicado: (2025)
SG-DOR: Learning Scene Graphs with Direction-Conditioned Occlusion Reasoning for Pepper Plants
por: Menon, Rohit, et al.
Publicado: (2026)
por: Menon, Rohit, et al.
Publicado: (2026)
Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
por: Gupta, Rohit, et al.
Publicado: (2025)
por: Gupta, Rohit, et al.
Publicado: (2025)
Don't Reach for the Stars: Rethinking Topology for Resilient Federated Learning
por: Konstantin, Mirko, et al.
Publicado: (2025)
por: Konstantin, Mirko, et al.
Publicado: (2025)
FrOoDo: Framework for Out-of-Distribution Detection
por: Stieber, Jonathan, et al.
Publicado: (2022)
por: Stieber, Jonathan, et al.
Publicado: (2022)
NCA-Morph: Medical Image Registration with Neural Cellular Automata
por: Ranem, Amin, et al.
Publicado: (2024)
por: Ranem, Amin, et al.
Publicado: (2024)
MotionCharacter: Fine-Grained Motion Controllable Human Video Generation
por: Fang, Haopeng, et al.
Publicado: (2024)
por: Fang, Haopeng, et al.
Publicado: (2024)
FewShotNeRF: Meta-Learning-based Novel View Synthesis for Rapid Scene-Specific Adaptation
por: Sivakumar, Piraveen, et al.
Publicado: (2024)
por: Sivakumar, Piraveen, et al.
Publicado: (2024)
Ejemplares similares
-
SurGrID: Controllable Surgical Simulation via Scene Graph to Image Diffusion
por: Frisch, Yannik, et al.
Publicado: (2025) -
SASVi -- Segment Any Surgical Video
por: Sivakumar, Ssharvien Kumar, et al.
Publicado: (2025) -
CAT-SG: A Large Dynamic Scene Graph Dataset for Fine-Grained Understanding of Cataract Surgery
por: Holm, Felix, et al.
Publicado: (2025) -
Frequency-Time Diffusion with Neural Cellular Automata
por: Kalkhof, John, et al.
Publicado: (2024) -
SANGRIA: Surgical Video Scene Graph Optimization for Surgical Workflow Prediction
por: Köksal, Çağhan, et al.
Publicado: (2024)