Saved in:
| Main Authors: | Takenaka, Patrick, Maucher, Johannes, Huber, Marco F. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2406.18220 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ViPro: Enabling and Controlling Video Prediction for Complex Dynamical Scenarios using Procedural Knowledge
by: Takenaka, Patrick, et al.
Published: (2024)
by: Takenaka, Patrick, et al.
Published: (2024)
ViPro-2: Unsupervised State Estimation via Integrated Dynamics for Guiding Video Prediction
by: Takenaka, Patrick, et al.
Published: (2025)
by: Takenaka, Patrick, et al.
Published: (2025)
Classification of Inkjet Printers based on Droplet Statistics
by: Takenaka, Patrick, et al.
Published: (2024)
by: Takenaka, Patrick, et al.
Published: (2024)
Anonymization of Documents for Law Enforcement with Machine Learning
by: Eberhardinger, Manuel, et al.
Published: (2025)
by: Eberhardinger, Manuel, et al.
Published: (2025)
VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's Guide
by: Lee, Dohun, et al.
Published: (2024)
by: Lee, Dohun, et al.
Published: (2024)
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
by: Girdhar, Rohit, et al.
Published: (2023)
by: Girdhar, Rohit, et al.
Published: (2023)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
by: Yang, Dejie, et al.
Published: (2024)
by: Yang, Dejie, et al.
Published: (2024)
SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos
by: Niu, Yulei, et al.
Published: (2024)
by: Niu, Yulei, et al.
Published: (2024)
Detection-Fusion for Knowledge Graph Extraction from Videos
by: Das, Taniya, et al.
Published: (2024)
by: Das, Taniya, et al.
Published: (2024)
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Improving Zero-shot Generalization of Learned Prompts via Unsupervised Knowledge Distillation
by: Mistretta, Marco, et al.
Published: (2024)
by: Mistretta, Marco, et al.
Published: (2024)
Probabilistic Contrastive Learning with Explicit Concentration on the Hypersphere
by: Li, Hongwei Bran, et al.
Published: (2024)
by: Li, Hongwei Bran, et al.
Published: (2024)
Explicitly Disentangled Representations in Object-Centric Learning
by: Majellaro, Riccardo, et al.
Published: (2024)
by: Majellaro, Riccardo, et al.
Published: (2024)
Contextualized Diffusion Models for Text-Guided Image and Video Generation
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
Mitigating Interference in the Knowledge Continuum through Attention-Guided Incremental Learning
by: Bhat, Prashant, et al.
Published: (2024)
by: Bhat, Prashant, et al.
Published: (2024)
Programmatic Video Prediction Using Large Language Models
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Multi-Modal Video Feature Extraction for Popularity Prediction
by: Liu, Haixu, et al.
Published: (2025)
by: Liu, Haixu, et al.
Published: (2025)
MaizeField3D: A Curated 3D Point Cloud and Procedural Model Dataset of Field-Grown Maize from a Diversity Panel
by: Kimara, Elvis, et al.
Published: (2025)
by: Kimara, Elvis, et al.
Published: (2025)
Synthesizing 3D Abstractions by Inverting Procedural Buildings with Transformers
by: Dax, Maximilian, et al.
Published: (2025)
by: Dax, Maximilian, et al.
Published: (2025)
Revisiting Feature Prediction for Learning Visual Representations from Video
by: Bardes, Adrien, et al.
Published: (2024)
by: Bardes, Adrien, et al.
Published: (2024)
CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
by: Li, Zian, et al.
Published: (2025)
by: Li, Zian, et al.
Published: (2025)
IMEX-Reg: Implicit-Explicit Regularization in the Function Space for Continual Learning
by: Bhat, Prashant, et al.
Published: (2024)
by: Bhat, Prashant, et al.
Published: (2024)
Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM
by: Kim, Jaemin, et al.
Published: (2024)
by: Kim, Jaemin, et al.
Published: (2024)
MIND: Modality-Informed Knowledge Distillation Framework for Multimodal Clinical Prediction Tasks
by: Guerra-Manzanares, Alejandro, et al.
Published: (2025)
by: Guerra-Manzanares, Alejandro, et al.
Published: (2025)
Learning from Oblivion: Predicting Knowledge Overflowed Weights via Retrodiction of Forgetting
by: Jang, Jinhyeok, et al.
Published: (2025)
by: Jang, Jinhyeok, et al.
Published: (2025)
CoLLM-NAS: Collaborative Large Language Models for Efficient Knowledge-Guided Neural Architecture Search
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
Uncertainty-Guided Selective Adaptation Enables Cross-Platform Predictive Fluorescence Microscopy
by: Yang, Kai-Wen K., et al.
Published: (2025)
by: Yang, Kai-Wen K., et al.
Published: (2025)
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
by: Ahmadian, Mona, et al.
Published: (2024)
by: Ahmadian, Mona, et al.
Published: (2024)
Nano World Models: A Minimalist Implementation of Future Video Prediction
by: Huang, Siqiao, et al.
Published: (2026)
by: Huang, Siqiao, et al.
Published: (2026)
Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models
by: Agarwal, Sakshi, et al.
Published: (2026)
by: Agarwal, Sakshi, et al.
Published: (2026)
GoldiCLIP: The Goldilocks Approach for Balancing Explicit Supervision for Language-Image Pretraining
by: Mohan, Deen Dayal, et al.
Published: (2026)
by: Mohan, Deen Dayal, et al.
Published: (2026)
SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain
by: Mao, Lingtao, et al.
Published: (2026)
by: Mao, Lingtao, et al.
Published: (2026)
Air Quality Prediction with A Meteorology-Guided Modality-Decoupled Spatio-Temporal Network
by: Yin, Hang, et al.
Published: (2025)
by: Yin, Hang, et al.
Published: (2025)
Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory
by: Wu, Yuqi, et al.
Published: (2025)
by: Wu, Yuqi, et al.
Published: (2025)
Saliency-guided Emotion Modeling: Predicting Viewer Reactions from Video Stimuli
by: Yaragoppa, Akhila, et al.
Published: (2025)
by: Yaragoppa, Akhila, et al.
Published: (2025)
Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
by: Li, Zhiyuan, et al.
Published: (2026)
by: Li, Zhiyuan, et al.
Published: (2026)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
by: Lin, Han, et al.
Published: (2023)
by: Lin, Han, et al.
Published: (2023)
Distilling Knowledge for Short-to-Long Term Trajectory Prediction
by: Das, Sourav, et al.
Published: (2023)
by: Das, Sourav, et al.
Published: (2023)
Lightning Fast Video Anomaly Detection via Adversarial Knowledge Distillation
by: Croitoru, Florinel-Alin, et al.
Published: (2022)
by: Croitoru, Florinel-Alin, et al.
Published: (2022)
Spatial Knowledge Graph-Guided Multimodal Synthesis
by: Xue, Yida, et al.
Published: (2025)
by: Xue, Yida, et al.
Published: (2025)
Similar Items
-
ViPro: Enabling and Controlling Video Prediction for Complex Dynamical Scenarios using Procedural Knowledge
by: Takenaka, Patrick, et al.
Published: (2024) -
ViPro-2: Unsupervised State Estimation via Integrated Dynamics for Guiding Video Prediction
by: Takenaka, Patrick, et al.
Published: (2025) -
Classification of Inkjet Printers based on Droplet Statistics
by: Takenaka, Patrick, et al.
Published: (2024) -
Anonymization of Documents for Law Enforcement with Machine Learning
by: Eberhardinger, Manuel, et al.
Published: (2025) -
VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's Guide
by: Lee, Dohun, et al.
Published: (2024)