PISCO: Precise Video Instance Insertion with Sparse Control
Fuente:
arXiv
Salvato in:
| Autori principali: | Gao, Xiangbo, Li, Renjie, Chen, Xinghao, Wu, Yuheng, Feng, Suofei, Yin, Qing, Tu, Zhengzhong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
di: Wu, Yuheng, et al.
Pubblicazione: (2026)
di: Wu, Yuheng, et al.
Pubblicazione: (2026)
SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
di: Yu, Jiongze, et al.
Pubblicazione: (2026)
di: Yu, Jiongze, et al.
Pubblicazione: (2026)
DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
di: Godbole, Mihir, et al.
Pubblicazione: (2025)
di: Godbole, Mihir, et al.
Pubblicazione: (2025)
AirV2X: Unified Air-Ground Vehicle-to-Everything Collaboration
di: Gao, Xiangbo, et al.
Pubblicazione: (2025)
di: Gao, Xiangbo, et al.
Pubblicazione: (2025)
CAST: Contrastive Adaptation and Distillation for Semi-Supervised Instance Segmentation
di: Taghavi, Pardis, et al.
Pubblicazione: (2025)
di: Taghavi, Pardis, et al.
Pubblicazione: (2025)
LangCoop: Collaborative Driving with Language
di: Gao, Xiangbo, et al.
Pubblicazione: (2025)
di: Gao, Xiangbo, et al.
Pubblicazione: (2025)
Physics-Aware Video Instance Removal Benchmark
di: Li, Zirui, et al.
Pubblicazione: (2026)
di: Li, Zirui, et al.
Pubblicazione: (2026)
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
di: Gao, Xiangbo, et al.
Pubblicazione: (2026)
di: Gao, Xiangbo, et al.
Pubblicazione: (2026)
The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics
di: Gao, Xiangbo, et al.
Pubblicazione: (2026)
di: Gao, Xiangbo, et al.
Pubblicazione: (2026)
mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation
di: Hu, Chan-Wei, et al.
Pubblicazione: (2025)
di: Hu, Chan-Wei, et al.
Pubblicazione: (2025)
STAMP: Scalable Task And Model-agnostic Collaborative Perception
di: Gao, Xiangbo, et al.
Pubblicazione: (2025)
di: Gao, Xiangbo, et al.
Pubblicazione: (2025)
CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception
di: Wang, Rujia, et al.
Pubblicazione: (2025)
di: Wang, Rujia, et al.
Pubblicazione: (2025)
SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion
di: Chen, Xinyu, et al.
Pubblicazione: (2026)
di: Chen, Xinyu, et al.
Pubblicazione: (2026)
Background Fades, Foreground Leads: Curriculum-Guided Background Pruning for Efficient Foreground-Centric Collaborative Perception
di: Wu, Yuheng, et al.
Pubblicazione: (2025)
di: Wu, Yuheng, et al.
Pubblicazione: (2025)
SafeCoop: Unravelling Full Stack Safety in Agentic Collaborative Driving
di: Gao, Xiangbo, et al.
Pubblicazione: (2025)
di: Gao, Xiangbo, et al.
Pubblicazione: (2025)
Automated Vehicles Should be Connected with Natural Language
di: Gao, Xiangbo, et al.
Pubblicazione: (2025)
di: Gao, Xiangbo, et al.
Pubblicazione: (2025)
Controllable Video Object Insertion via Multiview Priors
di: Qi, Xia, et al.
Pubblicazione: (2026)
di: Qi, Xia, et al.
Pubblicazione: (2026)
I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength
di: Feng, Wanquan, et al.
Pubblicazione: (2024)
di: Feng, Wanquan, et al.
Pubblicazione: (2024)
PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion
di: Choi, Jaehyun, et al.
Pubblicazione: (2025)
di: Choi, Jaehyun, et al.
Pubblicazione: (2025)
T2T-VICL: Unlocking the Boundaries of Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
di: Xia, Shao-Jun, et al.
Pubblicazione: (2025)
di: Xia, Shao-Jun, et al.
Pubblicazione: (2025)
Region-R1: Reinforcing Query-Side Region Cropping for Multi-Modal Re-Ranking
di: Hu, Chan-Wei, et al.
Pubblicazione: (2026)
di: Hu, Chan-Wei, et al.
Pubblicazione: (2026)
AdaRing: Towards Ultra-Light Vision-Language Adaptation via Cross-Layer Tensor Ring Decomposition
di: Huang, Ying, et al.
Pubblicazione: (2025)
di: Huang, Ying, et al.
Pubblicazione: (2025)
DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
di: Pan, Chenbin, et al.
Pubblicazione: (2025)
di: Pan, Chenbin, et al.
Pubblicazione: (2025)
Ride the Wave: Precision-Allocated Sparse Attention for Smooth Video Generation
di: Zhang, Wentai, et al.
Pubblicazione: (2026)
di: Zhang, Wentai, et al.
Pubblicazione: (2026)
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
di: Li, Chenglin, et al.
Pubblicazione: (2024)
di: Li, Chenglin, et al.
Pubblicazione: (2024)
CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences
di: Lin, Fangzhou, et al.
Pubblicazione: (2026)
di: Lin, Fangzhou, et al.
Pubblicazione: (2026)
VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control
di: Tu, Yuanpeng, et al.
Pubblicazione: (2025)
di: Tu, Yuanpeng, et al.
Pubblicazione: (2025)
Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model
di: Feng, Qianhan, et al.
Pubblicazione: (2024)
di: Feng, Qianhan, et al.
Pubblicazione: (2024)
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
di: Gu, Bohai, et al.
Pubblicazione: (2026)
di: Gu, Bohai, et al.
Pubblicazione: (2026)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
di: Fan, Tiehan, et al.
Pubblicazione: (2024)
di: Fan, Tiehan, et al.
Pubblicazione: (2024)
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
di: Wang, Zun, et al.
Pubblicazione: (2025)
di: Wang, Zun, et al.
Pubblicazione: (2025)
IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval
di: Liu, Bangwei, et al.
Pubblicazione: (2025)
di: Liu, Bangwei, et al.
Pubblicazione: (2025)
Efficient Video Diffusion with Sparse Information Transmission for Video Compression
di: Zhou, Mingde, et al.
Pubblicazione: (2026)
di: Zhou, Mingde, et al.
Pubblicazione: (2026)
Quantitative Video World Model Evaluation for Geometric-Consistency
di: Wu, Jiaxin, et al.
Pubblicazione: (2026)
di: Wu, Jiaxin, et al.
Pubblicazione: (2026)
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
di: Wu, Yinwei, et al.
Pubblicazione: (2024)
di: Wu, Yinwei, et al.
Pubblicazione: (2024)
VISTAv2: World Imagination for Indoor Vision-and-Language Navigation
di: Huang, Yanjia, et al.
Pubblicazione: (2025)
di: Huang, Yanjia, et al.
Pubblicazione: (2025)
Pandora: Towards General World Model with Natural Language Actions and Video States
di: Xiang, Jiannan, et al.
Pubblicazione: (2024)
di: Xiang, Jiannan, et al.
Pubblicazione: (2024)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
di: Tian, Kexin, et al.
Pubblicazione: (2025)
di: Tian, Kexin, et al.
Pubblicazione: (2025)
DPU: Dynamic Prototype Updating for Multimodal Out-of-Distribution Detection
di: Li, Shawn, et al.
Pubblicazione: (2024)
di: Li, Shawn, et al.
Pubblicazione: (2024)
Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
di: Zhang, Yan, et al.
Pubblicazione: (2025)
di: Zhang, Yan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
di: Wu, Yuheng, et al.
Pubblicazione: (2026) -
SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
di: Yu, Jiongze, et al.
Pubblicazione: (2026) -
DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
di: Godbole, Mihir, et al.
Pubblicazione: (2025) -
AirV2X: Unified Air-Ground Vehicle-to-Everything Collaboration
di: Gao, Xiangbo, et al.
Pubblicazione: (2025) -
CAST: Contrastive Adaptation and Distillation for Semi-Supervised Instance Segmentation
di: Taghavi, Pardis, et al.
Pubblicazione: (2025)