Anchoring and Rescaling Attention for Semantically Coherent Inbetweening

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Tae Eun, Shim, Sumin, Kim, Junhyeok, Hwang, Seong Jae
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908897474772992
author Choi, Tae Eun
Shim, Sumin
Kim, Junhyeok
Hwang, Seong Jae
author_facet Choi, Tae Eun
Shim, Sumin
Kim, Junhyeok
Hwang, Seong Jae
contents Generative inbetweening (GI) seeks to synthesize realistic intermediate frames between the first and last keyframes beyond mere interpolation. As sequences become sparser and motions larger, previous GI models struggle with inconsistent frames with unstable pacing and semantic misalignment. Since GI involves fixed endpoints and numerous plausible paths, this task requires additional guidance gained from the keyframes and text to specify the intended path. Thus, we give semantic and temporal guidance from the keyframes and text onto each intermediate frame through Keyframe-anchored Attention Bias. We also better enforce frame consistency with Rescaled Temporal RoPE, which allows self-attention to attend to keyframes more faithfully. TGI-Bench, the first benchmark specifically designed for text-conditioned GI evaluation, enables challenge-targeted evaluation to analyze GI models. Without additional training, our method achieves state-of-the-art frame consistency, semantic fidelity, and pace stability for both short and long sequences across diverse challenges.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17651
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
Choi, Tae Eun
Shim, Sumin
Kim, Junhyeok
Hwang, Seong Jae
Computer Vision and Pattern Recognition
Artificial Intelligence
Generative inbetweening (GI) seeks to synthesize realistic intermediate frames between the first and last keyframes beyond mere interpolation. As sequences become sparser and motions larger, previous GI models struggle with inconsistent frames with unstable pacing and semantic misalignment. Since GI involves fixed endpoints and numerous plausible paths, this task requires additional guidance gained from the keyframes and text to specify the intended path. Thus, we give semantic and temporal guidance from the keyframes and text onto each intermediate frame through Keyframe-anchored Attention Bias. We also better enforce frame consistency with Rescaled Temporal RoPE, which allows self-attention to attend to keyframes more faithfully. TGI-Bench, the first benchmark specifically designed for text-conditioned GI evaluation, enables challenge-targeted evaluation to analyze GI models. Without additional training, our method achieves state-of-the-art frame consistency, semantic fidelity, and pace stability for both short and long sequences across diverse challenges.
title Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.17651