Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Binyuan, Lu, Yuning, Jia, Weinan, Wang, Hualiang, Liu, Mu, Yang, Daiqing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915916049022976
author Huang, Binyuan
Lu, Yuning
Jia, Weinan
Wang, Hualiang
Liu, Mu
Yang, Daiqing
author_facet Huang, Binyuan
Lu, Yuning
Jia, Weinan
Wang, Hualiang
Liu, Mu
Yang, Daiqing
contents Recent proprietary models such as Sora2 demonstrate promising progress in generating multi-shot videos conditioned on multiple reference characters. However, academic research on this problem remains limited. We study this task and identify a core challenge: when reference images exhibit highly similar appearances, the model often suffers from reference confusion, where semantically similar tokens degrade the model's ability to retrieve the correct context. To address this, we introduce PoCo (Position Embedding as a Context Controller), which incorporates position encoding as additional context control beyond semantic retrieval. By employing side information of tokens, PoCo enables precise token-level matching while preserving implicit semantic consistency modeling. Building on PoCo, we develop a multi-reference and multi-shot video generation model capable of reliably controlling characters with extremely similar visual traits. Extensive experiments demonstrate that PoCo improves cross-shot consistency and reference fidelity compared with various baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2604_03738
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
Huang, Binyuan
Lu, Yuning
Jia, Weinan
Wang, Hualiang
Liu, Mu
Yang, Daiqing
Computer Vision and Pattern Recognition
Recent proprietary models such as Sora2 demonstrate promising progress in generating multi-shot videos conditioned on multiple reference characters. However, academic research on this problem remains limited. We study this task and identify a core challenge: when reference images exhibit highly similar appearances, the model often suffers from reference confusion, where semantically similar tokens degrade the model's ability to retrieve the correct context. To address this, we introduce PoCo (Position Embedding as a Context Controller), which incorporates position encoding as additional context control beyond semantic retrieval. By employing side information of tokens, PoCo enables precise token-level matching while preserving implicit semantic consistency modeling. Building on PoCo, we develop a multi-reference and multi-shot video generation model capable of reliably controlling characters with extremely similar visual traits. Extensive experiments demonstrate that PoCo improves cross-shot consistency and reference fidelity compared with various baselines.
title Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.03738