Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Ramos, Vasco, Bitton, Yonatan, Yarom, Michal, Szpektor, Idan, Magalhaes, Joao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks
by: Bordalo, João, et al.
Published: (2024)
by: Bordalo, João, et al.
Published: (2024)
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Beyond the Noise: Aligning Prompts with Latent Representations in Diffusion Models
by: Ramos, Vasco, et al.
Published: (2025)
by: Ramos, Vasco, et al.
Published: (2025)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
by: Zohar, Orr, et al.
Published: (2024)
by: Zohar, Orr, et al.
Published: (2024)
Latent Beam Diffusion Models for Generating Visual Sequences
by: Fernandes, Guilherme, et al.
Published: (2025)
by: Fernandes, Guilherme, et al.
Published: (2025)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
by: Yanuka, Moran, et al.
Published: (2024)
by: Yanuka, Moran, et al.
Published: (2024)
Error-Driven Scene Editing for 3D Grounding in Large Language Models
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation
by: Slobodkin, Aviv, et al.
Published: (2025)
by: Slobodkin, Aviv, et al.
Published: (2025)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
by: Gordon, Brian, et al.
Published: (2023)
by: Gordon, Brian, et al.
Published: (2023)
Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
by: Gordon, Brian, et al.
Published: (2025)
by: Gordon, Brian, et al.
Published: (2025)
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
by: Bitton-Guetta, Nitzan, et al.
Published: (2024)
by: Bitton-Guetta, Nitzan, et al.
Published: (2024)
VideoPhy: Evaluating Physical Commonsense for Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer
by: Wu, Guile, et al.
Published: (2025)
by: Wu, Guile, et al.
Published: (2025)
Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis
by: Ventura, Mor, et al.
Published: (2026)
by: Ventura, Mor, et al.
Published: (2026)
Seeing Across Time and Views: Multi-Temporal Cross-View Learning for Robust Video Person Re-Identification
by: Rashidunnabi, Md, et al.
Published: (2025)
by: Rashidunnabi, Md, et al.
Published: (2025)
Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding
by: Pereira, Joao, et al.
Published: (2025)
by: Pereira, Joao, et al.
Published: (2025)
Anchored Diffusion for Video Face Reenactment
by: Kligvasser, Idan, et al.
Published: (2024)
by: Kligvasser, Idan, et al.
Published: (2024)
FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding
by: Pereira, João, et al.
Published: (2026)
by: Pereira, João, et al.
Published: (2026)
Zero-Shot Action Recognition in Surveillance Videos
by: Pereira, Joao, et al.
Published: (2024)
by: Pereira, Joao, et al.
Published: (2024)
Anchored Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models
by: Hassan, Mariam, et al.
Published: (2025)
by: Hassan, Mariam, et al.
Published: (2025)
Multi-scale Contrastive Adaptor Learning for Segmenting Anything in Underperformed Scenes
by: Zhou, Ke, et al.
Published: (2024)
by: Zhou, Ke, et al.
Published: (2024)
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
by: Yosef, Ron, et al.
Published: (2025)
by: Yosef, Ron, et al.
Published: (2025)
Autonomous Character-Scene Interaction Synthesis from Text Instruction
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
Multi-Scale Contrastive Learning for Video Temporal Grounding
by: Nguyen, Thong Thanh, et al.
Published: (2024)
by: Nguyen, Thong Thanh, et al.
Published: (2024)
VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step
by: Wang, Hanyang, et al.
Published: (2025)
by: Wang, Hanyang, et al.
Published: (2025)
InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior
by: Lin, Chenguo, et al.
Published: (2024)
by: Lin, Chenguo, et al.
Published: (2024)
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
by: Condez, Ana Carolina, et al.
Published: (2025)
by: Condez, Ana Carolina, et al.
Published: (2025)
DiffuScene: Denoising Diffusion Models for Generative Indoor Scene Synthesis
by: Tang, Jiapeng, et al.
Published: (2023)
by: Tang, Jiapeng, et al.
Published: (2023)
SwiMDiff: Scene-wide Matching Contrastive Learning with Diffusion Constraint for Remote Sensing Image
by: Tian, Jiayuan, et al.
Published: (2024)
by: Tian, Jiayuan, et al.
Published: (2024)
Forest2Seq: Revitalizing Order Prior for Sequential Indoor Scene Synthesis
by: Sun, Qi, et al.
Published: (2024)
by: Sun, Qi, et al.
Published: (2024)
Self-Supervised Contrastive Learning for Videos using Differentiable Local Alignment
by: Oei, Keyne, et al.
Published: (2024)
by: Oei, Keyne, et al.
Published: (2024)
Vid3D: Synthesis of Dynamic 3D Scenes using 2D Video Diffusion
by: Parthasarathy, Rishab, et al.
Published: (2024)
by: Parthasarathy, Rishab, et al.
Published: (2024)
Show and Guide: Instructional-Plan Grounded Vision and Language Model
by: Glória-Silva, Diogo, et al.
Published: (2024)
by: Glória-Silva, Diogo, et al.
Published: (2024)
Video Perception Models for 3D Scene Synthesis
by: Huang, Rui, et al.
Published: (2025)
by: Huang, Rui, et al.
Published: (2025)
DiVE: Efficient Multi-View Driving Scenes Generation Based on Video Diffusion Transformer
by: Jiang, Junpeng, et al.
Published: (2025)
by: Jiang, Junpeng, et al.
Published: (2025)
Mixed Diffusion for 3D Indoor Scene Synthesis
by: Hu, Siyi, et al.
Published: (2024)
by: Hu, Siyi, et al.
Published: (2024)
PDPP: Projected Diffusion for Procedure Planning in Instructional Videos
by: Wang, Hanlin, et al.
Published: (2023)
by: Wang, Hanlin, et al.
Published: (2023)
Similar Items
-
Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks
by: Bordalo, João, et al.
Published: (2024) -
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
by: Bansal, Hritik, et al.
Published: (2024) -
Beyond the Noise: Aligning Prompts with Latent Representations in Diffusion Models
by: Ramos, Vasco, et al.
Published: (2025) -
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
by: Zohar, Orr, et al.
Published: (2024) -
Latent Beam Diffusion Models for Generating Visual Sequences
by: Fernandes, Guilherme, et al.
Published: (2025)