VideoAgent: Personalized Synthesis of Scientific Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Xiao, Li, Bangxin, Chen, Zixuan, Zheng, Hanyue, Ma, Zhi, Wang, Di, Tian, Cong, Wang, Quan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VideoAgent: Self-Improving Video Generation
von: Soni, Achint, et al.
Veröffentlicht: (2024)
von: Soni, Achint, et al.
Veröffentlicht: (2024)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
LLM Agent Framework for Intelligent Change Analysis in Urban Environment using Remote Sensing Imagery
von: Xiao, Zixuan, et al.
Veröffentlicht: (2026)
von: Xiao, Zixuan, et al.
Veröffentlicht: (2026)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
von: Liang, Yujia, et al.
Veröffentlicht: (2025)
von: Liang, Yujia, et al.
Veröffentlicht: (2025)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
von: Deng, Andong, et al.
Veröffentlicht: (2025)
von: Deng, Andong, et al.
Veröffentlicht: (2025)
Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification
von: Wang, Xiao, et al.
Veröffentlicht: (2026)
von: Wang, Xiao, et al.
Veröffentlicht: (2026)
Towards Intelligent Urban Park Development Monitoring: LLM Agents for Multi-Modal Information Fusion and Analysis
von: Xiao, Zixuan, et al.
Veröffentlicht: (2026)
von: Xiao, Zixuan, et al.
Veröffentlicht: (2026)
MMUEChange: A Generalized LLM Agent Framework for Intelligent Multi-Modal Urban Environment Change Analysis
von: Xiao, Zixuan, et al.
Veröffentlicht: (2026)
von: Xiao, Zixuan, et al.
Veröffentlicht: (2026)
Personalized Video Summarization by Multimodal Video Understanding
von: Chen, Brian, et al.
Veröffentlicht: (2024)
von: Chen, Brian, et al.
Veröffentlicht: (2024)
General and Task-Oriented Video Segmentation
von: Chen, Mu, et al.
Veröffentlicht: (2024)
von: Chen, Mu, et al.
Veröffentlicht: (2024)
Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
von: Zhou, Lianhao, et al.
Veröffentlicht: (2025)
von: Zhou, Lianhao, et al.
Veröffentlicht: (2025)
ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding
von: Liu, Xiao, et al.
Veröffentlicht: (2026)
von: Liu, Xiao, et al.
Veröffentlicht: (2026)
A Model-based Multi-Agent Personalized Short-Video Recommender System
von: Zhou, Peilun, et al.
Veröffentlicht: (2024)
von: Zhou, Peilun, et al.
Veröffentlicht: (2024)
Synthetic Video Enhances Physical Fidelity in Video Synthesis
von: Zhao, Qi, et al.
Veröffentlicht: (2025)
von: Zhao, Qi, et al.
Veröffentlicht: (2025)
VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
von: Zhi, Zhuo, et al.
Veröffentlicht: (2025)
von: Zhi, Zhuo, et al.
Veröffentlicht: (2025)
Toward Personalized LLM-Powered Agents: Foundations, Evaluation, and Future Directions
von: Xu, Yue, et al.
Veröffentlicht: (2026)
von: Xu, Yue, et al.
Veröffentlicht: (2026)
VideoAgentTrek: Computer Use Pretraining from Unlabeled Videos
von: Lu, Dunjie, et al.
Veröffentlicht: (2025)
von: Lu, Dunjie, et al.
Veröffentlicht: (2025)
CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual Rationale
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
Video2Game: Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single Video
von: Xia, Hongchi, et al.
Veröffentlicht: (2024)
von: Xia, Hongchi, et al.
Veröffentlicht: (2024)
Adaptive Multi-Agent Reasoning for Text-to-Video Retrieval
von: Wu, Jiaxin, et al.
Veröffentlicht: (2025)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2025)
DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing
von: Li, Ke, et al.
Veröffentlicht: (2026)
von: Li, Ke, et al.
Veröffentlicht: (2026)
LongVideoAgent: Multi-Agent Reasoning with Long Videos
von: Liu, Runtao, et al.
Veröffentlicht: (2025)
von: Liu, Runtao, et al.
Veröffentlicht: (2025)
Paper2Video: Automatic Video Generation from Scientific Papers
von: Zhu, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhu, Zeyu, et al.
Veröffentlicht: (2025)
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
von: Liu, Ye, et al.
Veröffentlicht: (2025)
von: Liu, Ye, et al.
Veröffentlicht: (2025)
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
CPGRec+: A Balance-oriented Framework for Personalized Video Game Recommendations
von: Li, Xiping, et al.
Veröffentlicht: (2026)
von: Li, Xiping, et al.
Veröffentlicht: (2026)
LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering
von: Dong, Xinxin, et al.
Veröffentlicht: (2025)
von: Dong, Xinxin, et al.
Veröffentlicht: (2025)
LLM-based Realistic Safety-Critical Driving Video Generation
von: Fu, Yongjie, et al.
Veröffentlicht: (2025)
von: Fu, Yongjie, et al.
Veröffentlicht: (2025)
AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production
von: Wang, Jiuniu, et al.
Veröffentlicht: (2024)
von: Wang, Jiuniu, et al.
Veröffentlicht: (2024)
BioProAgent: Neuro-Symbolic Grounding for Constrained Scientific Planning
von: Liu, Yuyang, et al.
Veröffentlicht: (2026)
von: Liu, Yuyang, et al.
Veröffentlicht: (2026)
Controllable Video Object Insertion via Multiview Priors
von: Qi, Xia, et al.
Veröffentlicht: (2026)
von: Qi, Xia, et al.
Veröffentlicht: (2026)
LOLGORITHM: Funny Comment Generation Agent For Short Videos
von: Ouyang, Xuan, et al.
Veröffentlicht: (2026)
von: Ouyang, Xuan, et al.
Veröffentlicht: (2026)
VCA: Video Curious Agent for Long Video Understanding
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024)
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024)
PoseCrafter: One-Shot Personalized Video Synthesis Following Flexible Pose Control
von: Zhong, Yong, et al.
Veröffentlicht: (2024)
von: Zhong, Yong, et al.
Veröffentlicht: (2024)
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
von: Zhao, Haozhe, et al.
Veröffentlicht: (2026)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2026)
CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
von: Zeng, Qinglin, et al.
Veröffentlicht: (2025)
von: Zeng, Qinglin, et al.
Veröffentlicht: (2025)
PVChat: Personalized Video Chat with One-Shot Learning
von: Shi, Yufei, et al.
Veröffentlicht: (2025)
von: Shi, Yufei, et al.
Veröffentlicht: (2025)
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
von: Li, Chenglin, et al.
Veröffentlicht: (2024)
von: Li, Chenglin, et al.
Veröffentlicht: (2024)
AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows
von: Shen, Shuaike, et al.
Veröffentlicht: (2026)
von: Shen, Shuaike, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
VideoAgent: Self-Improving Video Generation
von: Soni, Achint, et al.
Veröffentlicht: (2024) -
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024) -
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
von: Fan, Yue, et al.
Veröffentlicht: (2024) -
LLM Agent Framework for Intelligent Change Analysis in Urban Environment using Remote Sensing Imagery
von: Xiao, Zixuan, et al.
Veröffentlicht: (2026) -
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
von: Liang, Yujia, et al.
Veröffentlicht: (2025)