SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Xinyu, Qian, Yuyi, Lin, Jiang, Wang, Shenyi, Wang, Gao, Zhang, Zhiqiu, Zhang, Jizhi, Wang, Mingjie, Tang, Qiang, Wang, Qian, Wu, Song, Yi, Zili |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction
by: Lin, Jiang, et al.
Published: (2025)
by: Lin, Jiang, et al.
Published: (2025)
Region-to-Region: Enhancing Generative Image Harmonization with Adaptive Regional Injection
by: Zhang, Zhiqiu, et al.
Published: (2025)
by: Zhang, Zhiqiu, et al.
Published: (2025)
FreeInsert: Personalized Object Insertion with Geometric and Style Control
by: Zhang, Yuhong, et al.
Published: (2025)
by: Zhang, Yuhong, et al.
Published: (2025)
Point2Insert: Video Object Insertion via Sparse Point Guidance
by: Zhou, Yu, et al.
Published: (2026)
by: Zhou, Yu, et al.
Published: (2026)
TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On
by: Shao, Dingbao, et al.
Published: (2026)
by: Shao, Dingbao, et al.
Published: (2026)
EasyInsert: A Data-Efficient and Generalizable Insertion Policy
by: Li, Guanghe, et al.
Published: (2025)
by: Li, Guanghe, et al.
Published: (2025)
A Simple Distributed Algorithm for Sparse Fractional Covering and Packing Problems
by: Li, Qian, et al.
Published: (2024)
by: Li, Qian, et al.
Published: (2024)
FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors
by: Li, Chenxi, et al.
Published: (2025)
by: Li, Chenxi, et al.
Published: (2025)
OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
by: Chen, Jinshu, et al.
Published: (2025)
by: Chen, Jinshu, et al.
Published: (2025)
Efficient Turing Machine Simulation with Transformers
by: Li, Qian, et al.
Published: (2025)
by: Li, Qian, et al.
Published: (2025)
Constant Bit-size Transformers Are Turing Complete
by: Li, Qian, et al.
Published: (2025)
by: Li, Qian, et al.
Published: (2025)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
MultiTest: Physical-Aware Object Insertion for Testing Multi-sensor Fusion Perception Systems
by: Gao, Xinyu, et al.
Published: (2024)
by: Gao, Xinyu, et al.
Published: (2024)
Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance
by: Wu, Song, et al.
Published: (2026)
by: Wu, Song, et al.
Published: (2026)
DreamInsert: Zero-Shot Image-to-Video Object Insertion from A Single Image
by: Zhao, Qi, et al.
Published: (2025)
by: Zhao, Qi, et al.
Published: (2025)
One-Shot Learning for Pose-Guided Person Image Synthesis in the Wild
by: Fan, Dongqi, et al.
Published: (2024)
by: Fan, Dongqi, et al.
Published: (2024)
PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement
by: Xie, Tianyidan, et al.
Published: (2026)
by: Xie, Tianyidan, et al.
Published: (2026)
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
by: Jin, Hoiyeong, et al.
Published: (2025)
by: Jin, Hoiyeong, et al.
Published: (2025)
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
by: Fang, Zheng, et al.
Published: (2026)
by: Fang, Zheng, et al.
Published: (2026)
Anything in Any Scene: Photorealistic Video Object Insertion
by: Bai, Chen, et al.
Published: (2024)
by: Bai, Chen, et al.
Published: (2024)
Insert Anything: Image Insertion via In-Context Editing in DiT
by: Song, Wensong, et al.
Published: (2025)
by: Song, Wensong, et al.
Published: (2025)
SparseLIF: High-Performance Sparse LiDAR-Camera Fusion for 3D Object Detection
by: Zhang, Hongcheng, et al.
Published: (2024)
by: Zhang, Hongcheng, et al.
Published: (2024)
Vision‐Aided Damage Detection With Convolutional Multihead Self‐Attention Neural Network: A Novel Framework for Damage Information Extraction and Fusion
by: Yiming Zhang, et al.
Published: (2025)
by: Yiming Zhang, et al.
Published: (2025)
BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation
by: Zhu, Zihao, et al.
Published: (2026)
by: Zhu, Zihao, et al.
Published: (2026)
Failure Forecasting Boosts Robustness of Sim2Real Rhythmic Insertion Policies
by: Liu, Yuhan, et al.
Published: (2025)
by: Liu, Yuhan, et al.
Published: (2025)
Understanding and Optimizing Attention-Based Sparse Matching for Diverse Local Features
by: Wang, Qiang
Published: (2026)
by: Wang, Qiang
Published: (2026)
Controllable Video Object Insertion via Multiview Priors
by: Qi, Xia, et al.
Published: (2026)
by: Qi, Xia, et al.
Published: (2026)
Amulet: Fast TEE-Shielded Inference for On-Device Model Protection
by: Mao, Zikai, et al.
Published: (2025)
by: Mao, Zikai, et al.
Published: (2025)
Quantum Trojan Insertion: Controlled Activation for Covert Circuit Manipulation
by: John, Jayden, et al.
Published: (2025)
by: John, Jayden, et al.
Published: (2025)
Uncertainty-Encoded Multi-Modal Fusion for Robust Object Detection in Autonomous Driving
by: Lou, Yang, et al.
Published: (2023)
by: Lou, Yang, et al.
Published: (2023)
Boosting Adversarial Transferability with Low-Cost Optimization via Maximin Expected Flatness
by: Qiu, Chunlin, et al.
Published: (2024)
by: Qiu, Chunlin, et al.
Published: (2024)
PISCO: Precise Video Instance Insertion with Sparse Control
by: Gao, Xiangbo, et al.
Published: (2026)
by: Gao, Xiangbo, et al.
Published: (2026)
Fully Sparse Fusion for 3D Object Detection
by: Li, Yingyan, et al.
Published: (2023)
by: Li, Yingyan, et al.
Published: (2023)
On the Expressive Power and Limitations of Multi-Layer SSMs
by: Zubić, Nikola, et al.
Published: (2026)
by: Zubić, Nikola, et al.
Published: (2026)
Solving No-wait Scheduling for Time-Sensitive Networks with Daisy-Chain Topology
by: Li, Qian, et al.
Published: (2026)
by: Li, Qian, et al.
Published: (2026)
AccidentSim: Generating Vehicle Collision Videos with Physically Realistic Collision Trajectories from Real-World Accident Reports
by: Zhang, Xiangwen, et al.
Published: (2025)
by: Zhang, Xiangwen, et al.
Published: (2025)
Multimodal Fusion and Coherence Modeling for Video Topic Segmentation
by: Yu, Hai, et al.
Published: (2024)
by: Yu, Hai, et al.
Published: (2024)
Research on Industrial Process Fault Diagnosis Based on Deep Spatiotemporal Fusion Graph Convolutional Network
by: Qiang Qian, et al.
Published: (2024)
by: Qiang Qian, et al.
Published: (2024)
An efficient algorithm for multiuser sum-rate maximization of large-scale active RIS-aided MIMO system
by: Zhang, Qian, et al.
Published: (2023)
by: Zhang, Qian, et al.
Published: (2023)
AIM: Let Any Multi-modal Large Language Models Embrace Efficient In-Context Learning
by: Gao, Jun, et al.
Published: (2024)
by: Gao, Jun, et al.
Published: (2024)
Similar Items
-
FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction
by: Lin, Jiang, et al.
Published: (2025) -
Region-to-Region: Enhancing Generative Image Harmonization with Adaptive Regional Injection
by: Zhang, Zhiqiu, et al.
Published: (2025) -
FreeInsert: Personalized Object Insertion with Geometric and Style Control
by: Zhang, Yuhong, et al.
Published: (2025) -
Point2Insert: Video Object Insertion via Sparse Point Guidance
by: Zhou, Yu, et al.
Published: (2026) -
TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On
by: Shao, Dingbao, et al.
Published: (2026)