Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yesiltepe, Hidir, Meral, Tuna Han Salih, Akan, Adil Kaan, Oktay, Kaan, Yanardag, Pinar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Aligning Latent Geometry for Spherical Flow Matching in Image Generation
von: Meral, Tuna Han Salih, et al.
Veröffentlicht: (2026)
von: Meral, Tuna Han Salih, et al.
Veröffentlicht: (2026)
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2026)
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2026)
MotionFlow: Attention-Driven Motion Transfer in Video Diffusion Models
von: Meral, Tuna Han Salih, et al.
Veröffentlicht: (2024)
von: Meral, Tuna Han Salih, et al.
Veröffentlicht: (2024)
MotionShop: Zero-Shot Motion Transfer in Video Diffusion Models with Mixture of Score Guidance
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2024)
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2024)
Dynamic View Synthesis as an Inverse Problem
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2025)
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2025)
MIST: Mitigating Intersectional Bias with Disentangled Cross-Attention Editing in Text-to-Image Diffusion Models
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2024)
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2024)
The Curious Case of End Token: A Zero-Shot Disentangled Image Editing using CLIP
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2024)
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2024)
LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow Transformers
von: Dalva, Yusuf, et al.
Veröffentlicht: (2025)
von: Dalva, Yusuf, et al.
Veröffentlicht: (2025)
GANTASTIC: GAN-based Transfer of Interpretable Directions for Disentangled Image Editing in Text-to-Image Diffusion Models
von: Dalva, Yusuf, et al.
Veröffentlicht: (2024)
von: Dalva, Yusuf, et al.
Veröffentlicht: (2024)
Learning Object-Centric Representations Based on Slots in Real World Scenarios
von: Akan, Adil Kaan
Veröffentlicht: (2025)
von: Akan, Adil Kaan
Veröffentlicht: (2025)
Contrastive Test-Time Composition of Multiple LoRA Models for Image Generation
von: Meral, Tuna Han Salih, et al.
Veröffentlicht: (2024)
von: Meral, Tuna Han Salih, et al.
Veröffentlicht: (2024)
Compositional Video Synthesis by Temporal Object-Centric Learning
von: Akan, Adil Kaan, et al.
Veröffentlicht: (2025)
von: Akan, Adil Kaan, et al.
Veröffentlicht: (2025)
Periodic RoPE for Infinite Context LLMs
von: Huo, Simin
Veröffentlicht: (2026)
von: Huo, Simin
Veröffentlicht: (2026)
Slot-Guided Adaptation of Pre-trained Diffusion Models for Object-Centric Learning and Compositional Generation
von: Akan, Adil Kaan, et al.
Veröffentlicht: (2025)
von: Akan, Adil Kaan, et al.
Veröffentlicht: (2025)
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
von: Helbling, Alec, et al.
Veröffentlicht: (2025)
von: Helbling, Alec, et al.
Veröffentlicht: (2025)
DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2026)
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2026)
Stylebreeder: Exploring and Democratizing Artistic Styles through Text-to-Image Models
von: Zheng, Matthew, et al.
Veröffentlicht: (2024)
von: Zheng, Matthew, et al.
Veröffentlicht: (2024)
ReRoPE: Repurposing RoPE for Relative Camera Control
von: Li, Chunyang, et al.
Veröffentlicht: (2026)
von: Li, Chunyang, et al.
Veröffentlicht: (2026)
RoPE-LIME: RoPE-Space Locality + Sparse-K Sampling for Efficient LLM Attribution
von: Picov, Isaac, et al.
Veröffentlicht: (2026)
von: Picov, Isaac, et al.
Veröffentlicht: (2026)
Base of RoPE Bounds Context Length
von: Men, Xin, et al.
Veröffentlicht: (2024)
von: Men, Xin, et al.
Veröffentlicht: (2024)
Scaling Laws of RoPE-based Extrapolation
von: Liu, Xiaoran, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoran, et al.
Veröffentlicht: (2023)
ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models
von: Yan, Feihong, et al.
Veröffentlicht: (2026)
von: Yan, Feihong, et al.
Veröffentlicht: (2026)
Untwisting RoPE: Frequency Control for Shared Attention in DiTs
von: Mikaeili, Aryan, et al.
Veröffentlicht: (2026)
von: Mikaeili, Aryan, et al.
Veröffentlicht: (2026)
AdaState: Self-Evolving Anchors for Streaming Video Generation
von: Dalva, Yusuf, et al.
Veröffentlicht: (2026)
von: Dalva, Yusuf, et al.
Veröffentlicht: (2026)
Demystifying the Slash Pattern in Attention: The Role of RoPE
von: Cheng, Yuan, et al.
Veröffentlicht: (2026)
von: Cheng, Yuan, et al.
Veröffentlicht: (2026)
Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation
von: Li, Zichong, et al.
Veröffentlicht: (2026)
von: Li, Zichong, et al.
Veröffentlicht: (2026)
RoPE Attention Can Be Trained in Almost Linear Time
von: Cao, Yang, et al.
Veröffentlicht: (2024)
von: Cao, Yang, et al.
Veröffentlicht: (2024)
Frayed RoPE and Long Inputs: A Geometric Perspective
von: Wertheimer, Davis, et al.
Veröffentlicht: (2026)
von: Wertheimer, Davis, et al.
Veröffentlicht: (2026)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
von: Chen, Bo, et al.
Veröffentlicht: (2024)
von: Chen, Bo, et al.
Veröffentlicht: (2024)
Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion
von: Yang, Yang, et al.
Veröffentlicht: (2026)
von: Yang, Yang, et al.
Veröffentlicht: (2026)
RAP: KV-Cache Compression via RoPE-Aligned Pruning
von: Xin, Jihao, et al.
Veröffentlicht: (2026)
von: Xin, Jihao, et al.
Veröffentlicht: (2026)
A Circular Argument : Does RoPE need to be Equivariant for Vision?
von: van de Geijn, Chase, et al.
Veröffentlicht: (2025)
von: van de Geijn, Chase, et al.
Veröffentlicht: (2025)
Understanding the RoPE Extensions of Long-Context LLMs: An Attention Perspective
von: Zhong, Meizhi, et al.
Veröffentlicht: (2024)
von: Zhong, Meizhi, et al.
Veröffentlicht: (2024)
On the token distance modeling ability of higher RoPE attention dimension
von: Hong, Xiangyu, et al.
Veröffentlicht: (2024)
von: Hong, Xiangyu, et al.
Veröffentlicht: (2024)
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
von: Yang, Ning, et al.
Veröffentlicht: (2026)
von: Yang, Ning, et al.
Veröffentlicht: (2026)
Rotate Both Ways: Time-and-Order RoPE for Generative Recommendation
von: Wei, Xiaokai, et al.
Veröffentlicht: (2025)
von: Wei, Xiaokai, et al.
Veröffentlicht: (2025)
Conditional Information Gain Trellis
von: Bicici, Ufuk Can, et al.
Veröffentlicht: (2024)
von: Bicici, Ufuk Can, et al.
Veröffentlicht: (2024)
CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs
von: Li, Haoran, et al.
Veröffentlicht: (2026)
von: Li, Haoran, et al.
Veröffentlicht: (2026)
Attention Drift: What Autoregressive Speculative Decoding Models Learn
von: Eldenk, Doğaç, et al.
Veröffentlicht: (2026)
von: Eldenk, Doğaç, et al.
Veröffentlicht: (2026)
DTT-BSR: GAN-based DTTNet with RoPE Transformer Enhancement for Music Source Restoration
von: Tan, Shihong, et al.
Veröffentlicht: (2026)
von: Tan, Shihong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Aligning Latent Geometry for Spherical Flow Matching in Image Generation
von: Meral, Tuna Han Salih, et al.
Veröffentlicht: (2026) -
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2026) -
MotionFlow: Attention-Driven Motion Transfer in Video Diffusion Models
von: Meral, Tuna Han Salih, et al.
Veröffentlicht: (2024) -
MotionShop: Zero-Shot Motion Transfer in Video Diffusion Models with Mixture of Score Guidance
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2024) -
Dynamic View Synthesis as an Inverse Problem
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2025)