Agentic Video Generation: From Text to Executable Event Graphs via Tool-Constrained LLM Planning
Fuente:
arXiv
Saved in:
| Main Authors: | Cudlenco, Nicolae, Masala, Mihai, Leordeanu, Marius |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models
by: Cudlenco, Nicolae, et al.
Published: (2026)
by: Cudlenco, Nicolae, et al.
Published: (2026)
Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time
by: Masala, Mihai, et al.
Published: (2025)
by: Masala, Mihai, et al.
Published: (2025)
From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
by: Masala, Mihai, et al.
Published: (2025)
by: Masala, Mihai, et al.
Published: (2025)
Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
by: Mihai-Cristian, Pîrvu, et al.
Published: (2025)
by: Mihai-Cristian, Pîrvu, et al.
Published: (2025)
With Great Context Comes Great Prediction Power: Classifying Objects via Geo-Semantic Scene Graphs
by: Constantinescu, Ciprian, et al.
Published: (2025)
by: Constantinescu, Ciprian, et al.
Published: (2025)
Multi-modal video data-pipelines for machine learning with minimal human supervision
by: Pîrvu, Mihai-Cristian, et al.
Published: (2025)
by: Pîrvu, Mihai-Cristian, et al.
Published: (2025)
Inside Knowledge: Graph-based Path Generation with Explainable Data Augmentation and Curriculum Learning for Visual Indoor Navigation
by: Airinei, Daniel, et al.
Published: (2025)
by: Airinei, Daniel, et al.
Published: (2025)
Multiple Random Masking Autoencoder Ensembles for Robust Multimodal Semi-supervised Learning
by: Todoran, Alexandru-Raul, et al.
Published: (2024)
by: Todoran, Alexandru-Raul, et al.
Published: (2024)
Learning from Random Subspace Exploration: Generalized Test-Time Augmentation with Self-supervised Distillation
by: Jelea, Andrei, et al.
Published: (2025)
by: Jelea, Andrei, et al.
Published: (2025)
Closer to Ground Truth: Realistic Shape and Appearance Labeled Data Generation for Unsupervised Underwater Image Segmentation
by: Jelea, Andrei, et al.
Published: (2025)
by: Jelea, Andrei, et al.
Published: (2025)
Non-verbal Real-time Human-AI Interaction in Constrained Robotic Environments
by: Costea, Dragos, et al.
Published: (2026)
by: Costea, Dragos, et al.
Published: (2026)
Learning on the Fly: Replay-Based Continual Object Perception for Indoor Drones
by: Nae, Sebastian-Ion, et al.
Published: (2026)
by: Nae, Sebastian-Ion, et al.
Published: (2026)
Efficient Self-Supervised Neuro-Analytic Visual Servoing for Real-time Quadrotor Control
by: Mocanu, Sebastian, et al.
Published: (2025)
by: Mocanu, Sebastian, et al.
Published: (2025)
Self-Supervised Learning to Fly using Efficient Semantic Segmentation and Metric Depth Estimation for Low-Cost Autonomous UAVs
by: Mocanu, Sebastian, et al.
Published: (2025)
by: Mocanu, Sebastian, et al.
Published: (2025)
Maia: A Real-time Non-Verbal Chat for Human-AI Interaction
by: Costea, Dragos, et al.
Published: (2024)
by: Costea, Dragos, et al.
Published: (2024)
A self-supervised cyclic neural-analytic approach for novel view synthesis and 3D reconstruction
by: Costea, Dragos, et al.
Published: (2025)
by: Costea, Dragos, et al.
Published: (2025)
NEWTON: Agentic Planning for Physically Grounded Video Generation
by: Feng, Yuxiang, et al.
Published: (2026)
by: Feng, Yuxiang, et al.
Published: (2026)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
Iterative Explainability for Weakly Supervised Segmentation in Medical PE Detection
by: Condrea, Florin, et al.
Published: (2024)
by: Condrea, Florin, et al.
Published: (2024)
From Engineering Diagrams to Graphs: Digitizing P&IDs with Transformers
by: Stürmer, Jan Marius, et al.
Published: (2024)
by: Stürmer, Jan Marius, et al.
Published: (2024)
Open-Event Procedure Planning in Instructional Videos
by: Wu, Yilu, et al.
Published: (2024)
by: Wu, Yilu, et al.
Published: (2024)
RoadTones: Tone Controllable Text Generation from Road Event Videos
by: Parikh, Chirag, et al.
Published: (2026)
by: Parikh, Chirag, et al.
Published: (2026)
"Înţelegi Româneşte?'' A Recipe for Romanian Vision-Language Models
by: Masala, Mihai, et al.
Published: (2026)
by: Masala, Mihai, et al.
Published: (2026)
TRACE: Temporal Grounding Video LLM via Causal Event Modeling
by: Guo, Yongxin, et al.
Published: (2024)
by: Guo, Yongxin, et al.
Published: (2024)
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
by: Yang, Zuhao, et al.
Published: (2026)
by: Yang, Zuhao, et al.
Published: (2026)
Making Image Editing Easier via Adaptive Task Reformulation with Agentic Executions
by: Zhao, Bo, et al.
Published: (2026)
by: Zhao, Bo, et al.
Published: (2026)
GenAgent: Scaling Text-to-Image Generation via Agentic Multimodal Reasoning
by: Jiang, Kaixun, et al.
Published: (2026)
by: Jiang, Kaixun, et al.
Published: (2026)
EA-VTR: Event-Aware Video-Text Retrieval
by: Ma, Zongyang, et al.
Published: (2024)
by: Ma, Zongyang, et al.
Published: (2024)
GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning
by: Lv, Jiaxi, et al.
Published: (2023)
by: Lv, Jiaxi, et al.
Published: (2023)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
by: He, Haibin, et al.
Published: (2026)
by: He, Haibin, et al.
Published: (2026)
Generative Anonymization in Event Streams
by: Müller, Adam T., et al.
Published: (2026)
by: Müller, Adam T., et al.
Published: (2026)
A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
by: Kurpath, Mohammed Irfan, et al.
Published: (2025)
by: Kurpath, Mohammed Irfan, et al.
Published: (2025)
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
by: Wen, Siwei, et al.
Published: (2026)
by: Wen, Siwei, et al.
Published: (2026)
PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection
by: Wang, Peiyao, et al.
Published: (2025)
by: Wang, Peiyao, et al.
Published: (2025)
NEGATE: Constrained Semantic Guidance for Linguistic Negation in Text-to-Video Diffusion
by: Kang, Taewon, et al.
Published: (2026)
by: Kang, Taewon, et al.
Published: (2026)
Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment Retrieval
by: Jiang, Yiyang, et al.
Published: (2024)
by: Jiang, Yiyang, et al.
Published: (2024)
DyST-XL: Dynamic Layout Planning and Content Control for Compositional Text-to-Video Generation
by: He, Weijie, et al.
Published: (2025)
by: He, Weijie, et al.
Published: (2025)
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
by: Cheng, Zixu, et al.
Published: (2026)
by: Cheng, Zixu, et al.
Published: (2026)
LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
by: Wang, Jiarui, et al.
Published: (2025)
by: Wang, Jiarui, et al.
Published: (2025)
RISE-T2V: Rephrasing and Injecting Semantics with LLM for Expansive Text-to-Video Generation
by: Zhang, Xiangjun, et al.
Published: (2025)
by: Zhang, Xiangjun, et al.
Published: (2025)
Similar Items
-
GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models
by: Cudlenco, Nicolae, et al.
Published: (2026) -
Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time
by: Masala, Mihai, et al.
Published: (2025) -
From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
by: Masala, Mihai, et al.
Published: (2025) -
Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
by: Mihai-Cristian, Pîrvu, et al.
Published: (2025) -
With Great Context Comes Great Prediction Power: Classifying Objects via Geo-Semantic Scene Graphs
by: Constantinescu, Ciprian, et al.
Published: (2025)