AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Milton, Qin, Sizhong, Li, Yongzhi, Chen, Quan, Jiang, Peng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
di: Yang, Junjie, et al.
Pubblicazione: (2025)
di: Yang, Junjie, et al.
Pubblicazione: (2025)
Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition
di: Zhou, Jiaming, et al.
Pubblicazione: (2023)
di: Zhou, Jiaming, et al.
Pubblicazione: (2023)
MM2Latent: Text-to-facial image generation and editing in GANs with multimodal assistance
di: Meng, Debin, et al.
Pubblicazione: (2024)
di: Meng, Debin, et al.
Pubblicazione: (2024)
VDT-Auto: End-to-end Autonomous Driving with VLM-Guided Diffusion Transformers
di: Guo, Ziang, et al.
Pubblicazione: (2025)
di: Guo, Ziang, et al.
Pubblicazione: (2025)
DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation
di: Li, Haoran, et al.
Pubblicazione: (2025)
di: Li, Haoran, et al.
Pubblicazione: (2025)
Auto-Regressive Surface Cutting
di: Li, Yang, et al.
Pubblicazione: (2025)
di: Li, Yang, et al.
Pubblicazione: (2025)
End2end-ALARA: Approaching the ALARA Law in CT Imaging with End-to-end Learning
di: Tao, Xi, et al.
Pubblicazione: (2025)
di: Tao, Xi, et al.
Pubblicazione: (2025)
UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
di: Lu, Hao, et al.
Pubblicazione: (2025)
di: Lu, Hao, et al.
Pubblicazione: (2025)
DREAM: Document Reconstruction via End-to-end Autoregressive Model
di: Li, Xin, et al.
Pubblicazione: (2025)
di: Li, Xin, et al.
Pubblicazione: (2025)
Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning
di: Yao, Zhengjian, et al.
Pubblicazione: (2026)
di: Yao, Zhengjian, et al.
Pubblicazione: (2026)
Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation
di: You, Yuyang, et al.
Pubblicazione: (2026)
di: You, Yuyang, et al.
Pubblicazione: (2026)
RC-AutoCalib: An End-to-End Radar-Camera Automatic Calibration Network
di: Luu, Van-Tin, et al.
Pubblicazione: (2025)
di: Luu, Van-Tin, et al.
Pubblicazione: (2025)
Tokenization Allows Multimodal Large Language Models to Understand, Generate and Edit Architectural Floor Plans
di: Qin, Sizhong, et al.
Pubblicazione: (2026)
di: Qin, Sizhong, et al.
Pubblicazione: (2026)
CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous Driving
di: Ma, Enhui, et al.
Pubblicazione: (2025)
di: Ma, Enhui, et al.
Pubblicazione: (2025)
RID-TWIN: An end-to-end pipeline for automatic face de-identification in videos
di: Mukherjee, Anirban, et al.
Pubblicazione: (2024)
di: Mukherjee, Anirban, et al.
Pubblicazione: (2024)
What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards
di: Le, Minh-Quan, et al.
Pubblicazione: (2025)
di: Le, Minh-Quan, et al.
Pubblicazione: (2025)
ParkingE2E: Camera-based End-to-end Parking Network, from Images to Planning
di: Li, Changze, et al.
Pubblicazione: (2024)
di: Li, Changze, et al.
Pubblicazione: (2024)
AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning
di: Zhou, Zewei, et al.
Pubblicazione: (2025)
di: Zhou, Zewei, et al.
Pubblicazione: (2025)
End4: End-to-end Denoising Diffusion for Diffusion-Based Inpainting Detection
di: Wang, Fei, et al.
Pubblicazione: (2025)
di: Wang, Fei, et al.
Pubblicazione: (2025)
FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
di: Cong, Yuren, et al.
Pubblicazione: (2023)
di: Cong, Yuren, et al.
Pubblicazione: (2023)
STAR: Scale-wise Text-conditioned AutoRegressive image generation
di: Ma, Xiaoxiao, et al.
Pubblicazione: (2024)
di: Ma, Xiaoxiao, et al.
Pubblicazione: (2024)
Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos
di: Pan, Yulin, et al.
Pubblicazione: (2023)
di: Pan, Yulin, et al.
Pubblicazione: (2023)
VRWKV-Editor: Reducing quadratic complexity in transformer-based video editing
di: Aitrouga, Abdelilah, et al.
Pubblicazione: (2025)
di: Aitrouga, Abdelilah, et al.
Pubblicazione: (2025)
Adversarial AutoMixup
di: Qin, Huafeng, et al.
Pubblicazione: (2023)
di: Qin, Huafeng, et al.
Pubblicazione: (2023)
Generalized Trajectory Scoring for End-to-end Multimodal Planning
di: Li, Zhenxin, et al.
Pubblicazione: (2025)
di: Li, Zhenxin, et al.
Pubblicazione: (2025)
PPAD: Iterative Interactions of Prediction and Planning for End-to-end Autonomous Driving
di: Chen, Zhili, et al.
Pubblicazione: (2023)
di: Chen, Zhili, et al.
Pubblicazione: (2023)
MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild
di: Fang, Xi, et al.
Pubblicazione: (2024)
di: Fang, Xi, et al.
Pubblicazione: (2024)
AnomalyR1: A GRPO-based End-to-end MLLM for Industrial Anomaly Detection
di: Chao, Yuhao, et al.
Pubblicazione: (2025)
di: Chao, Yuhao, et al.
Pubblicazione: (2025)
Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
di: Li, Zhenxin, et al.
Pubblicazione: (2024)
di: Li, Zhenxin, et al.
Pubblicazione: (2024)
End-to-end autoencoding architecture for the simultaneous generation of medical images and corresponding segmentation masks
di: Kebaili, Aghiles, et al.
Pubblicazione: (2023)
di: Kebaili, Aghiles, et al.
Pubblicazione: (2023)
2D bidirectional gated recurrent unit convolutional Neural networks for end-to-end violence detection In videos
di: Traoré, Abdarahmane, et al.
Pubblicazione: (2024)
di: Traoré, Abdarahmane, et al.
Pubblicazione: (2024)
End-to-end Surface Optimization for Light Control
di: Sun, Yuou, et al.
Pubblicazione: (2024)
di: Sun, Yuou, et al.
Pubblicazione: (2024)
GraphAD: Interaction Scene Graph for End-to-end Autonomous Driving
di: Zhang, Yunpeng, et al.
Pubblicazione: (2024)
di: Zhang, Yunpeng, et al.
Pubblicazione: (2024)
HIMap: HybrId Representation Learning for End-to-end Vectorized HD Map Construction
di: Zhou, Yi, et al.
Pubblicazione: (2024)
di: Zhou, Yi, et al.
Pubblicazione: (2024)
Can video generation replace cinematographers? Research on the cinematic language of generated video
di: Li, Xiaozhe, et al.
Pubblicazione: (2024)
di: Li, Xiaozhe, et al.
Pubblicazione: (2024)
Large Spatial Model: End-to-end Unposed Images to Semantic 3D
di: Fan, Zhiwen, et al.
Pubblicazione: (2024)
di: Fan, Zhiwen, et al.
Pubblicazione: (2024)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
di: Wei, Haoran, et al.
Pubblicazione: (2024)
di: Wei, Haoran, et al.
Pubblicazione: (2024)
Closing the Navigation Compliance Gap in End-to-end Autonomous Driving
di: Wu, Hanfeng, et al.
Pubblicazione: (2025)
di: Wu, Hanfeng, et al.
Pubblicazione: (2025)
Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset
di: Ancarani, Elisa, et al.
Pubblicazione: (2025)
di: Ancarani, Elisa, et al.
Pubblicazione: (2025)
GE-AdvGAN: Improving the transferability of adversarial samples by gradient editing-based adversarial generative model
di: Zhu, Zhiyu, et al.
Pubblicazione: (2024)
di: Zhu, Zhiyu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
di: Yang, Junjie, et al.
Pubblicazione: (2025) -
Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition
di: Zhou, Jiaming, et al.
Pubblicazione: (2023) -
MM2Latent: Text-to-facial image generation and editing in GANs with multimodal assistance
di: Meng, Debin, et al.
Pubblicazione: (2024) -
VDT-Auto: End-to-end Autonomous Driving with VLM-Guided Diffusion Transformers
di: Guo, Ziang, et al.
Pubblicazione: (2025) -
DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation
di: Li, Haoran, et al.
Pubblicazione: (2025)