Agent-based Video Trimming
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yang, Lingfeng, Chen, Zhenyuan, Li, Xiang, Jia, Peiyang, Long, Liangqu, Yang, Jian |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Revisiting Prompt Pretraining of Vision-Language Models
par: Chen, Zhenyuan, et autres
Publié: (2024)
par: Chen, Zhenyuan, et autres
Publié: (2024)
Are Image-to-Video Models Good Zero-Shot Image Editors?
par: Zhang, Zechuan, et autres
Publié: (2025)
par: Zhang, Zechuan, et autres
Publié: (2025)
Trim 3D Gaussian Splatting for Accurate Geometry Representation
par: Fan, Lue, et autres
Publié: (2024)
par: Fan, Lue, et autres
Publié: (2024)
HazyDet: Open-Source Benchmark for Drone-View Object Detection with Depth-Cues in Hazy Scenes
par: Feng, Changfeng, et autres
Publié: (2024)
par: Feng, Changfeng, et autres
Publié: (2024)
VideoGen-Eval: Agent-based System for Video Generation Evaluation
par: Yang, Yuhang, et autres
Publié: (2025)
par: Yang, Yuhang, et autres
Publié: (2025)
Video-guided Machine Translation with Global Video Context
par: Chen, Jian, et autres
Publié: (2026)
par: Chen, Jian, et autres
Publié: (2026)
ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming
par: Zhuang, Jiedong, et autres
Publié: (2024)
par: Zhuang, Jiedong, et autres
Publié: (2024)
SpatioTemporal Difference Network for Video Depth Super-Resolution
par: Wang, Zhengxue, et autres
Publié: (2025)
par: Wang, Zhengxue, et autres
Publié: (2025)
VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
par: Eskandar, George, et autres
Publié: (2026)
par: Eskandar, George, et autres
Publié: (2026)
RayFormer: Modeling Inter- and Intra-Ray Similarity for NeRF-Based Video Snapshot Compressive Imaging
par: Dong, Yubo, et autres
Publié: (2026)
par: Dong, Yubo, et autres
Publié: (2026)
SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
par: Gu, Bo, et autres
Publié: (2026)
par: Gu, Bo, et autres
Publié: (2026)
Add-SD: Rational Generation without Manual Reference
par: Yang, Lingfeng, et autres
Publié: (2024)
par: Yang, Lingfeng, et autres
Publié: (2024)
Lighting-grounded Video Generation with Renderer-based Agent Reasoning
par: Cai, Ziqi, et autres
Publié: (2026)
par: Cai, Ziqi, et autres
Publié: (2026)
WOW-Seg: A Word-free Open World Segmentation Model
par: Li, Danyang, et autres
Publié: (2026)
par: Li, Danyang, et autres
Publié: (2026)
VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
par: Yu, Hanxun, et autres
Publié: (2026)
par: Yu, Hanxun, et autres
Publié: (2026)
OpenVIS: Open-vocabulary Video Instance Segmentation
par: Guo, Pinxue, et autres
Publié: (2023)
par: Guo, Pinxue, et autres
Publié: (2023)
Depth-Centric Dehazing and Depth-Estimation from Real-World Hazy Driving Video
par: Fan, Junkai, et autres
Publié: (2024)
par: Fan, Junkai, et autres
Publié: (2024)
Omni-Video: Democratizing Unified Video Understanding and Generation
par: Tan, Zhiyu, et autres
Publié: (2025)
par: Tan, Zhiyu, et autres
Publié: (2025)
PresentAgent: Multimodal Agent for Presentation Video Generation
par: Shi, Jingwei, et autres
Publié: (2025)
par: Shi, Jingwei, et autres
Publié: (2025)
TokenTrim: Inference-Time Token Pruning for Autoregressive Long Video Generation
par: Shaulov, Ariel, et autres
Publié: (2026)
par: Shaulov, Ariel, et autres
Publié: (2026)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
par: Yin, Yufei, et autres
Publié: (2026)
par: Yin, Yufei, et autres
Publié: (2026)
ControlNeXt: Powerful and Efficient Control for Image and Video Generation
par: Peng, Bohao, et autres
Publié: (2024)
par: Peng, Bohao, et autres
Publié: (2024)
StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
par: Yang, Haolin, et autres
Publié: (2025)
par: Yang, Haolin, et autres
Publié: (2025)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
par: Yang, Hao, et autres
Publié: (2026)
par: Yang, Hao, et autres
Publié: (2026)
TALL: Thumbnail Layout for Deepfake Video Detection
par: Xu, Yuting, et autres
Publié: (2023)
par: Xu, Yuting, et autres
Publié: (2023)
RSEdit: Text-Guided Image Editing for Remote Sensing
par: Zhenyuan, Chen, et autres
Publié: (2026)
par: Zhenyuan, Chen, et autres
Publié: (2026)
EarthBridge: A Solution for 4th Multi-modal Aerial View Image Challenge Translation Track
par: Chen, Zhenyuan, et autres
Publié: (2026)
par: Chen, Zhenyuan, et autres
Publié: (2026)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
par: Chen, Boyu, et autres
Publié: (2025)
par: Chen, Boyu, et autres
Publié: (2025)
DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
par: Yang, Zongxin, et autres
Publié: (2024)
par: Yang, Zongxin, et autres
Publié: (2024)
EEA: Exploration-Exploitation Agent for Long Video Understanding
par: Yang, Te, et autres
Publié: (2025)
par: Yang, Te, et autres
Publié: (2025)
Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
par: Wu, Ge, et autres
Publié: (2025)
par: Wu, Ge, et autres
Publié: (2025)
Hyperspectral Adapter for Object Tracking based on Hyperspectral Video
par: Gao, Long, et autres
Publié: (2025)
par: Gao, Long, et autres
Publié: (2025)
Neuromorphic Synergy for Video Binarization
par: Lin, Shijie, et autres
Publié: (2024)
par: Lin, Shijie, et autres
Publié: (2024)
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation
par: Chen, Junhao, et autres
Publié: (2025)
par: Chen, Junhao, et autres
Publié: (2025)
VQ-Jarvis: Retrieval-Augmented Video Restoration Agent with Sharp Vision and Fast Thought
par: Zhang, Xuanyu, et autres
Publié: (2026)
par: Zhang, Xuanyu, et autres
Publié: (2026)
Deep Height Decoupling for Precise Vision-based 3D Occupancy Prediction
par: Wu, Yuan, et autres
Publié: (2024)
par: Wu, Yuan, et autres
Publié: (2024)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
par: Fan, Yue, et autres
Publié: (2024)
par: Fan, Yue, et autres
Publié: (2024)
VCA: Video Curious Agent for Long Video Understanding
par: Yang, Zeyuan, et autres
Publié: (2024)
par: Yang, Zeyuan, et autres
Publié: (2024)
Learning to Trim: End-to-End Causal Graph Pruning with Dynamic Anatomical Feature Banks for Medical VQA
par: Xu, Zibo, et autres
Publié: (2026)
par: Xu, Zibo, et autres
Publié: (2026)
Video-Based Reward Modeling for Computer-Use Agents
par: Song, Linxin, et autres
Publié: (2026)
par: Song, Linxin, et autres
Publié: (2026)
Documents similaires
-
Revisiting Prompt Pretraining of Vision-Language Models
par: Chen, Zhenyuan, et autres
Publié: (2024) -
Are Image-to-Video Models Good Zero-Shot Image Editors?
par: Zhang, Zechuan, et autres
Publié: (2025) -
Trim 3D Gaussian Splatting for Accurate Geometry Representation
par: Fan, Lue, et autres
Publié: (2024) -
HazyDet: Open-Source Benchmark for Drone-View Object Detection with Depth-Cues in Hazy Scenes
par: Feng, Changfeng, et autres
Publié: (2024) -
VideoGen-Eval: Agent-based System for Video Generation Evaluation
par: Yang, Yuhang, et autres
Publié: (2025)