InstanceV: Instance-Level Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yuheng, Hu, Teng, Zhang, Jiangning, Xue, Zhucun, Yi, Ran, Ma, Lizhuang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Autoregressive Visual Generation with Cluster-Oriented Token Prediction
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
Evolution of Video Generative Foundations
by: Hu, Teng, et al.
Published: (2026)
by: Hu, Teng, et al.
Published: (2026)
Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction
by: Yi, Ran, et al.
Published: (2025)
by: Yi, Ran, et al.
Published: (2025)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
MotionMaster: Training-free Camera Motion Transfer For Video Generation
by: Hu, Teng, et al.
Published: (2024)
by: Hu, Teng, et al.
Published: (2024)
UltraGen: High-Resolution Video Generation with Hierarchical Attention
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
SaRA: High-Efficient Diffusion Model Fine-tuning with Progressive Sparse Low-Rank Adaptation
by: Hu, Teng, et al.
Published: (2024)
by: Hu, Teng, et al.
Published: (2024)
Semantic Frame Interpolation
by: Hong, Yijia, et al.
Published: (2025)
by: Hong, Yijia, et al.
Published: (2025)
UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy
by: Xu, Yicheng, et al.
Published: (2026)
by: Xu, Yicheng, et al.
Published: (2026)
Advancing Narrative Long Video Generation via Training-Free Identity-Aware Memory
by: Liu, Jinzhuo, et al.
Published: (2026)
by: Liu, Jinzhuo, et al.
Published: (2026)
Image Inversion: A Survey from GANs to Diffusion and Beyond
by: Chen, Yinan, et al.
Published: (2025)
by: Chen, Yinan, et al.
Published: (2025)
PoseAnything: Universal Pose-guided Video Generation with Part-aware Temporal Coherence
by: Wang, Ruiyan, et al.
Published: (2025)
by: Wang, Ruiyan, et al.
Published: (2025)
InstanceAnimator: Multi-Instance Sketch Video Colorization
by: Zhang, Yinhan, et al.
Published: (2026)
by: Zhang, Yinhan, et al.
Published: (2026)
Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10$\times$
by: Zhang, Jiangning, et al.
Published: (2025)
by: Zhang, Jiangning, et al.
Published: (2025)
IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment
by: Chen, Yinan, et al.
Published: (2025)
by: Chen, Yinan, et al.
Published: (2025)
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
by: Xue, Zhucun, et al.
Published: (2025)
by: Xue, Zhucun, et al.
Published: (2025)
Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
by: Xue, Zhucun, et al.
Published: (2025)
by: Xue, Zhucun, et al.
Published: (2025)
Instance-Level Generation for Representation Learning
by: Wu, Yankun, et al.
Published: (2025)
by: Wu, Yankun, et al.
Published: (2025)
Rethinking Multiple Instance Learning: Developing an Instance-Level Classifier via Weakly-Supervised Self-Training
by: Ma, Yingfan, et al.
Published: (2024)
by: Ma, Yingfan, et al.
Published: (2024)
InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
by: Xiang, Qiang, et al.
Published: (2025)
by: Xiang, Qiang, et al.
Published: (2025)
GPT-4V-AD: Exploring Grounding Potential of VQA-oriented GPT-4V for Zero-shot Anomaly Detection
by: Zhang, Jiangning, et al.
Published: (2023)
by: Zhang, Jiangning, et al.
Published: (2023)
Exploring Real&Synthetic Dataset and Linear Attention in Image Restoration
by: Du, Yuzhen, et al.
Published: (2024)
by: Du, Yuzhen, et al.
Published: (2024)
Instance Brownian Bridge as Texts for Open-vocabulary Video Instance Segmentation
by: Cheng, Zesen, et al.
Published: (2024)
by: Cheng, Zesen, et al.
Published: (2024)
PISCO: Precise Video Instance Insertion with Sparse Control
by: Gao, Xiangbo, et al.
Published: (2026)
by: Gao, Xiangbo, et al.
Published: (2026)
SyncVIS: Synchronized Video Instance Segmentation
by: Zheng, Rongkun, et al.
Published: (2024)
by: Zheng, Rongkun, et al.
Published: (2024)
Multi-Dimensional Knowledge Profiling with Large-Scale Literature Database and Hierarchical Retrieval
by: Xue, Zhucun, et al.
Published: (2026)
by: Xue, Zhucun, et al.
Published: (2026)
InstanceGaussian: Appearance-Semantic Joint Gaussian Representation for 3D Instance-Level Perception
by: Li, Haijie, et al.
Published: (2024)
by: Li, Haijie, et al.
Published: (2024)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
TIMotion: Temporal and Interactive Framework for Efficient Human-Human Motion Generation
by: Wang, Yabiao, et al.
Published: (2024)
by: Wang, Yabiao, et al.
Published: (2024)
InstanceGen: Image Generation with Instance-level Instructions
by: Sella, Etai, et al.
Published: (2025)
by: Sella, Etai, et al.
Published: (2025)
PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
EMOv2: Pushing 5M Vision Model Frontier
by: Zhang, Jiangning, et al.
Published: (2024)
by: Zhang, Jiangning, et al.
Published: (2024)
SDI-Paste: Synthetic Dynamic Instance Copy-Paste for Video Instance Segmentation
by: Shrestha, Sahir, et al.
Published: (2024)
by: Shrestha, Sahir, et al.
Published: (2024)
OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
by: He, Haoyang, et al.
Published: (2025)
by: He, Haoyang, et al.
Published: (2025)
Textual Decomposition Then Sub-motion-space Scattering for Open-Vocabulary Motion Generation
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
Instance-Level Composed Image Retrieval
by: Psomas, Bill, et al.
Published: (2025)
by: Psomas, Bill, et al.
Published: (2025)
BEAR: A Video Dataset For Fine-grained Behaviors Recognition Oriented with Action and Environment Factors
by: Hu, Chengyang, et al.
Published: (2025)
by: Hu, Chengyang, et al.
Published: (2025)
Similar Items
-
Improving Autoregressive Visual Generation with Cluster-Oriented Token Prediction
by: Hu, Teng, et al.
Published: (2025) -
Evolution of Video Generative Foundations
by: Hu, Teng, et al.
Published: (2026) -
Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation
by: Chen, Yuheng, et al.
Published: (2026) -
IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction
by: Yi, Ran, et al.
Published: (2025) -
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)