Salvato in:
| Autori principali: | Liu, Xiao, Chen, Guangyi, Tang, Yansong, Wang, Guangrun, Zhang, Xiao-Ping, Lim, Ser-Nam |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2307.03538 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VideoMerge: Towards Training-free Long Video Generation
di: Zhang, Siyang, et al.
Pubblicazione: (2025)
di: Zhang, Siyang, et al.
Pubblicazione: (2025)
Towards Chunk-Wise Generation for Long Videos
di: Zhang, Siyang, et al.
Pubblicazione: (2024)
di: Zhang, Siyang, et al.
Pubblicazione: (2024)
DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation
di: Meyarian, Abolfazl, et al.
Pubblicazione: (2026)
di: Meyarian, Abolfazl, et al.
Pubblicazione: (2026)
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
di: Chen, Harold Haodong, et al.
Pubblicazione: (2024)
di: Chen, Harold Haodong, et al.
Pubblicazione: (2024)
Narrative Action Evaluation with Prompt-Guided Multimodal Interaction
di: Zhang, Shiyi, et al.
Pubblicazione: (2024)
di: Zhang, Shiyi, et al.
Pubblicazione: (2024)
Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
di: Park, Dongmin, et al.
Pubblicazione: (2024)
di: Park, Dongmin, et al.
Pubblicazione: (2024)
VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
di: Wang, Yuji, et al.
Pubblicazione: (2025)
di: Wang, Yuji, et al.
Pubblicazione: (2025)
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
di: Qian, Zhaofang, et al.
Pubblicazione: (2024)
di: Qian, Zhaofang, et al.
Pubblicazione: (2024)
FSViewFusion: Few-Shots View Generation of Novel Objects
di: Hussain, Rukhshanda, et al.
Pubblicazione: (2024)
di: Hussain, Rukhshanda, et al.
Pubblicazione: (2024)
Towards Unified 3D Object Detection via Algorithm and Data Unification
di: Li, Zhuoling, et al.
Pubblicazione: (2024)
di: Li, Zhuoling, et al.
Pubblicazione: (2024)
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
di: Ye, Xubing, et al.
Pubblicazione: (2024)
di: Ye, Xubing, et al.
Pubblicazione: (2024)
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
di: Chen, Harold Haodong, et al.
Pubblicazione: (2025)
di: Chen, Harold Haodong, et al.
Pubblicazione: (2025)
BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration
di: Gao, Bo, et al.
Pubblicazione: (2026)
di: Gao, Bo, et al.
Pubblicazione: (2026)
DreamMask: Boosting Open-vocabulary Panoptic Segmentation with Synthetic Data
di: Tu, Yuanpeng, et al.
Pubblicazione: (2025)
di: Tu, Yuanpeng, et al.
Pubblicazione: (2025)
Fast Encoding and Decoding for Implicit Video Representation
di: Chen, Hao, et al.
Pubblicazione: (2024)
di: Chen, Hao, et al.
Pubblicazione: (2024)
GS: Generative Segmentation via Label Diffusion
di: Chen, Yuhao, et al.
Pubblicazione: (2025)
di: Chen, Yuhao, et al.
Pubblicazione: (2025)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
AirSketch: Generative Motion to Sketch
di: Lim, Hui Xian Grace, et al.
Pubblicazione: (2024)
di: Lim, Hui Xian Grace, et al.
Pubblicazione: (2024)
Temporal Action Localization with Cross Layer Task Decoupling and Refinement
di: Li, Qiang, et al.
Pubblicazione: (2024)
di: Li, Qiang, et al.
Pubblicazione: (2024)
Composing Object Relations and Attributes for Image-Text Matching
di: Pham, Khoi, et al.
Pubblicazione: (2024)
di: Pham, Khoi, et al.
Pubblicazione: (2024)
FDDet: Frequency-Decoupling for Boundary Refinement in Temporal Action Detection
di: Zhu, Xinnan, et al.
Pubblicazione: (2025)
di: Zhu, Xinnan, et al.
Pubblicazione: (2025)
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
di: Song, Zijian, et al.
Pubblicazione: (2025)
di: Song, Zijian, et al.
Pubblicazione: (2025)
Zero-shot Synthetic Video Realism Enhancement via Structure-aware Denoising
di: Wang, Yifan, et al.
Pubblicazione: (2025)
di: Wang, Yifan, et al.
Pubblicazione: (2025)
UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning
di: Chen, Hongyu, et al.
Pubblicazione: (2025)
di: Chen, Hongyu, et al.
Pubblicazione: (2025)
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
di: Mohammadshirazi, Ahmad, et al.
Pubblicazione: (2024)
di: Mohammadshirazi, Ahmad, et al.
Pubblicazione: (2024)
Towards Self-Refinement of Vision-Language Models with Triangular Consistency
di: Deng, Yunlong, et al.
Pubblicazione: (2025)
di: Deng, Yunlong, et al.
Pubblicazione: (2025)
What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?
di: Cui, Xuanming, et al.
Pubblicazione: (2025)
di: Cui, Xuanming, et al.
Pubblicazione: (2025)
LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
di: Li, Zhuoling, et al.
Pubblicazione: (2024)
di: Li, Zhuoling, et al.
Pubblicazione: (2024)
Niagara: Normal-Integrated Geometric Affine Field for Scene Reconstruction from a Single View
di: Wu, Xianzu, et al.
Pubblicazione: (2025)
di: Wu, Xianzu, et al.
Pubblicazione: (2025)
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
di: Jang, Young Kyun, et al.
Pubblicazione: (2024)
di: Jang, Young Kyun, et al.
Pubblicazione: (2024)
Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
di: Jang, Young Kyun, et al.
Pubblicazione: (2024)
di: Jang, Young Kyun, et al.
Pubblicazione: (2024)
Geometry aware 3D generation from in-the-wild images in ImageNet
di: Shen, Qijia, et al.
Pubblicazione: (2024)
di: Shen, Qijia, et al.
Pubblicazione: (2024)
Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
di: Jang, Young Kyun, et al.
Pubblicazione: (2024)
di: Jang, Young Kyun, et al.
Pubblicazione: (2024)
HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
di: Li, Weiqi, et al.
Pubblicazione: (2025)
di: Li, Weiqi, et al.
Pubblicazione: (2025)
IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis
di: Wang, Yuji, et al.
Pubblicazione: (2025)
di: Wang, Yuji, et al.
Pubblicazione: (2025)
ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision
di: Li, Weiqi, et al.
Pubblicazione: (2025)
di: Li, Weiqi, et al.
Pubblicazione: (2025)
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
di: Liu, Aoming, et al.
Pubblicazione: (2025)
di: Liu, Aoming, et al.
Pubblicazione: (2025)
OP-LoRA: The Blessing of Dimensionality
di: Teterwak, Piotr, et al.
Pubblicazione: (2024)
di: Teterwak, Piotr, et al.
Pubblicazione: (2024)
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
di: Zhou, Jiaying, et al.
Pubblicazione: (2026)
di: Zhou, Jiaying, et al.
Pubblicazione: (2026)
CLAMP: Contrastive LAnguage Model Prompt-tuning
di: Teterwak, Piotr, et al.
Pubblicazione: (2023)
di: Teterwak, Piotr, et al.
Pubblicazione: (2023)
Documenti analoghi
-
VideoMerge: Towards Training-free Long Video Generation
di: Zhang, Siyang, et al.
Pubblicazione: (2025) -
Towards Chunk-Wise Generation for Long Videos
di: Zhang, Siyang, et al.
Pubblicazione: (2024) -
DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation
di: Meyarian, Abolfazl, et al.
Pubblicazione: (2026) -
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
di: Chen, Harold Haodong, et al.
Pubblicazione: (2024) -
Narrative Action Evaluation with Prompt-Guided Multimodal Interaction
di: Zhang, Shiyi, et al.
Pubblicazione: (2024)