Physics-Aware Video Instance Removal Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zirui, Chen, Xinghao, Jiang, Lingyu, Hou, Dengzhe, Lin, Fangzhou, Yamada, Kazunori, Gao, Xiangbo, Tu, Zhengzhong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PISCO: Precise Video Instance Insertion with Sparse Control
by: Gao, Xiangbo, et al.
Published: (2026)
by: Gao, Xiangbo, et al.
Published: (2026)
TimePre: Bridging Accuracy, Efficiency, and Stability in Probabilistic Time-Series Forecasting
by: Jiang, Lingyu, et al.
Published: (2025)
by: Jiang, Lingyu, et al.
Published: (2025)
DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
by: Godbole, Mihir, et al.
Published: (2025)
by: Godbole, Mihir, et al.
Published: (2025)
Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
by: Wu, Yuheng, et al.
Published: (2026)
by: Wu, Yuheng, et al.
Published: (2026)
CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences
by: Lin, Fangzhou, et al.
Published: (2026)
by: Lin, Fangzhou, et al.
Published: (2026)
The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics
by: Gao, Xiangbo, et al.
Published: (2026)
by: Gao, Xiangbo, et al.
Published: (2026)
PathCal: State-Aware Reflection-Marker Calibration for Efficient Reasoning
by: Jiang, Lingyu, et al.
Published: (2026)
by: Jiang, Lingyu, et al.
Published: (2026)
VISTAv2: World Imagination for Indoor Vision-and-Language Navigation
by: Huang, Yanjia, et al.
Published: (2025)
by: Huang, Yanjia, et al.
Published: (2025)
NexusFlow: Unifying Disparate Tasks under Partial Supervision via Invertible Flow Networks
by: Lin, Fangzhou, et al.
Published: (2025)
by: Lin, Fangzhou, et al.
Published: (2025)
WMF-AM: Probing LLM Working Memory via Depth-Parameterized Cumulative State Tracking
by: Hou, Dengzhe, et al.
Published: (2026)
by: Hou, Dengzhe, et al.
Published: (2026)
Background Fades, Foreground Leads: Curriculum-Guided Background Pruning for Efficient Foreground-Centric Collaborative Perception
by: Wu, Yuheng, et al.
Published: (2025)
by: Wu, Yuheng, et al.
Published: (2025)
Same Brain, Different Prediction: How Preprocessing Choices Undermine EEG Decoding Reliability
by: Hou, Dengzhe, et al.
Published: (2026)
by: Hou, Dengzhe, et al.
Published: (2026)
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
by: Gao, Xiangbo, et al.
Published: (2026)
by: Gao, Xiangbo, et al.
Published: (2026)
SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
by: Yu, Jiongze, et al.
Published: (2026)
by: Yu, Jiongze, et al.
Published: (2026)
Hyperbolic Chamfer Distance for Point Cloud Completion and Beyond
by: Lin, Fangzhou, et al.
Published: (2024)
by: Lin, Fangzhou, et al.
Published: (2024)
STAMP: Scalable Task And Model-agnostic Collaborative Perception
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception
by: Wang, Rujia, et al.
Published: (2025)
by: Wang, Rujia, et al.
Published: (2025)
AirV2X: Unified Air-Ground Vehicle-to-Everything Collaboration
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
CAST: Contrastive Adaptation and Distillation for Semi-Supervised Instance Segmentation
by: Taghavi, Pardis, et al.
Published: (2025)
by: Taghavi, Pardis, et al.
Published: (2025)
SafeCoop: Unravelling Full Stack Safety in Agentic Collaborative Driving
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
iMOVE: Instance-Motion-Aware Video Understanding
by: Li, Jiaze, et al.
Published: (2025)
by: Li, Jiaze, et al.
Published: (2025)
Let the Abyss Stare Back Adaptive Falsification for Autonomous Scientific Discovery
by: Li, Peiran, et al.
Published: (2026)
by: Li, Peiran, et al.
Published: (2026)
SynCL: A Synergistic Training Strategy with Instance-Aware Contrastive Learning for End-to-End Multi-Camera 3D Tracking
by: Lin, Shubo, et al.
Published: (2024)
by: Lin, Shubo, et al.
Published: (2024)
LangCoop: Collaborative Driving with Language
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
Instance-Aware Robust Consistency Regularization for Semi-Supervised Nuclei Instance Segmentation
by: Lin, Zenan, et al.
Published: (2025)
by: Lin, Zenan, et al.
Published: (2025)
Prompt-Aware Controllable Shadow Removal
by: Chen, Kerui, et al.
Published: (2025)
by: Chen, Kerui, et al.
Published: (2025)
A Strong View-Free Baseline Approach for Single-View Image Guided Point Cloud Completion
by: Lin, Fangzhou, et al.
Published: (2025)
by: Lin, Fangzhou, et al.
Published: (2025)
SDI-Paste: Synthetic Dynamic Instance Copy-Paste for Video Instance Segmentation
by: Shrestha, Sahir, et al.
Published: (2024)
by: Shrestha, Sahir, et al.
Published: (2024)
CAVIS: Context-Aware Video Instance Segmentation
by: Lee, Seunghun, et al.
Published: (2024)
by: Lee, Seunghun, et al.
Published: (2024)
Automated Vehicles Should be Connected with Natural Language
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
SAKED: Mitigating Hallucination in Large Vision-Language Models via Stability-Aware Knowledge Enhanced Decoding
by: Li, Zhaoxu, et al.
Published: (2026)
by: Li, Zhaoxu, et al.
Published: (2026)
MWFormer: Multi-Weather Image Restoration Using Degradation-Aware Transformers
by: Zhu, Ruoxi, et al.
Published: (2024)
by: Zhu, Ruoxi, et al.
Published: (2024)
3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation
by: He, Yunhong, et al.
Published: (2025)
by: He, Yunhong, et al.
Published: (2025)
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
by: Xing, Shuo, et al.
Published: (2024)
by: Xing, Shuo, et al.
Published: (2024)
EduVQA: Towards Concept-Aware Assessment of Educational AI-Generated Videos
by: Chen, Baoliang, et al.
Published: (2026)
by: Chen, Baoliang, et al.
Published: (2026)
InstanceV: Instance-Level Video Generation
by: Chen, Yuheng, et al.
Published: (2025)
by: Chen, Yuheng, et al.
Published: (2025)
HeadsUp! High-Fidelity Portrait Image Super-Resolution
by: Li, Renjie, et al.
Published: (2025)
by: Li, Renjie, et al.
Published: (2025)
Instance Brownian Bridge as Texts for Open-vocabulary Video Instance Segmentation
by: Cheng, Zesen, et al.
Published: (2024)
by: Cheng, Zesen, et al.
Published: (2024)
A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation
by: Zhong, Qing, et al.
Published: (2025)
by: Zhong, Qing, et al.
Published: (2025)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
by: Tian, Kexin, et al.
Published: (2025)
by: Tian, Kexin, et al.
Published: (2025)
Similar Items
-
PISCO: Precise Video Instance Insertion with Sparse Control
by: Gao, Xiangbo, et al.
Published: (2026) -
TimePre: Bridging Accuracy, Efficiency, and Stability in Probabilistic Time-Series Forecasting
by: Jiang, Lingyu, et al.
Published: (2025) -
DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
by: Godbole, Mihir, et al.
Published: (2025) -
Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
by: Wu, Yuheng, et al.
Published: (2026) -
CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences
by: Lin, Fangzhou, et al.
Published: (2026)