Amodal Ground Truth and Completion in the Wild
Fuente:
arXiv
Saved in:
| Main Authors: | Zhan, Guanqi, Zheng, Chuanxia, Xie, Weidi, Zisserman, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A General Protocol to Probe Large Vision Models for 3D Physical Understanding
by: Zhan, Guanqi, et al.
Published: (2023)
by: Zhan, Guanqi, et al.
Published: (2023)
Inferring Dynamic Physical Properties from Video Foundation Models
by: Zhan, Guanqi, et al.
Published: (2025)
by: Zhan, Guanqi, et al.
Published: (2025)
ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval
by: Zhan, Guanqi, et al.
Published: (2025)
by: Zhan, Guanqi, et al.
Published: (2025)
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
by: Xie, Junyu, et al.
Published: (2026)
by: Xie, Junyu, et al.
Published: (2026)
Appearance-Based Refinement for Object-Centric Motion Segmentation
by: Xie, Junyu, et al.
Published: (2023)
by: Xie, Junyu, et al.
Published: (2023)
Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images
by: Wu, Tianhao, et al.
Published: (2025)
by: Wu, Tianhao, et al.
Published: (2025)
Moving Object Segmentation: All You Need Is SAM (and Flow)
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
Made to Order: Discovering monotonic temporal changes via self-supervised video ordering
by: Yang, Charig, et al.
Published: (2024)
by: Yang, Charig, et al.
Published: (2024)
Amodal Depth Anything: Amodal Depth Estimation in the Wild
by: Li, Zhenyu, et al.
Published: (2024)
by: Li, Zhenyu, et al.
Published: (2024)
NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction
by: Chen, Weirong, et al.
Published: (2026)
by: Chen, Weirong, et al.
Published: (2026)
Character-Centric Understanding of Animated Movies
by: Gui, Zhongrui, et al.
Published: (2025)
by: Gui, Zhongrui, et al.
Published: (2025)
Hyper-Transformer for Amodal Completion
by: Gao, Jianxiong, et al.
Published: (2024)
by: Gao, Jianxiong, et al.
Published: (2024)
PHAC: Promptable Human Amodal Completion
by: Noh, Seung Young, et al.
Published: (2026)
by: Noh, Seung Young, et al.
Published: (2026)
Open-World Amodal Appearance Completion
by: Ao, Jiayang, et al.
Published: (2024)
by: Ao, Jiayang, et al.
Published: (2024)
Recognizing Co-Speech Gestures in-the-Wild
by: Hegde, Sindhu B, et al.
Published: (2026)
by: Hegde, Sindhu B, et al.
Published: (2026)
AutoAD III: The Prequel -- Back to the Pixels
by: Han, Tengda, et al.
Published: (2024)
by: Han, Tengda, et al.
Published: (2024)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
TACO: Taming Diffusion for in-the-wild Video Amodal Completion
by: Lu, Ruijie, et al.
Published: (2025)
by: Lu, Ruijie, et al.
Published: (2025)
Synchformer: Efficient Synchronization from Sparse Cues
by: Iashin, Vladimir, et al.
Published: (2024)
by: Iashin, Vladimir, et al.
Published: (2024)
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
by: Xie, Junyu, et al.
Published: (2025)
by: Xie, Junyu, et al.
Published: (2025)
Free3D: Consistent Novel View Synthesis without 3D Representation
by: Zheng, Chuanxia, et al.
Published: (2023)
by: Zheng, Chuanxia, et al.
Published: (2023)
Reasoning-Driven Amodal Completion: Collaborative Agents and Perceptual Evaluation
by: Fan, Hongxing, et al.
Published: (2025)
by: Fan, Hongxing, et al.
Published: (2025)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
EGM: Efficient Visual Grounding Language Models
by: Zhan, Guanqi, et al.
Published: (2026)
by: Zhan, Guanqi, et al.
Published: (2026)
Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
by: Korbar, Bruno, et al.
Published: (2025)
by: Korbar, Bruno, et al.
Published: (2025)
The Manga Whisperer: Automatically Generating Transcriptions for Comics
by: Sachdeva, Ragav, et al.
Published: (2024)
by: Sachdeva, Ragav, et al.
Published: (2024)
Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
by: Bagad, Piyush, et al.
Published: (2025)
by: Bagad, Piyush, et al.
Published: (2025)
From Panels to Prose: Generating Literary Narratives from Comics
by: Sachdeva, Ragav, et al.
Published: (2025)
by: Sachdeva, Ragav, et al.
Published: (2025)
AmodalSVG: Amodal Image Vectorization via Semantic Layer Peeling
by: Hu, Juncheng, et al.
Published: (2026)
by: Hu, Juncheng, et al.
Published: (2026)
Integrating Multimodal Large Language Model Knowledge into Amodal Completion
by: Yun, Heecheol, et al.
Published: (2026)
by: Yun, Heecheol, et al.
Published: (2026)
AmodalSynthDrive: A Synthetic Amodal Perception Dataset for Autonomous Driving
by: Sekkat, Ahmed Rida, et al.
Published: (2023)
by: Sekkat, Ahmed Rida, et al.
Published: (2023)
Mask Guided Gated Convolution for Amodal Content Completion
by: Saleh, Kaziwa, et al.
Published: (2024)
by: Saleh, Kaziwa, et al.
Published: (2024)
ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance Segmentation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
CountGD++: Generalized Prompting for Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
Perspective-Equivariant Fine-tuning for Multispectral Demosaicing without Ground Truth
by: Wang, Andrew, et al.
Published: (2026)
by: Wang, Andrew, et al.
Published: (2026)
PanoDiffusion: 360-degree Panorama Outpainting via Diffusion
by: Wu, Tianhao, et al.
Published: (2023)
by: Wu, Tianhao, et al.
Published: (2023)
One-shot Human Motion Transfer via Occlusion-Robust Flow Prediction and Neural Texturing
by: Ji, Yuzhu, et al.
Published: (2024)
by: Ji, Yuzhu, et al.
Published: (2024)
SPATIALALIGN: Aligning Dynamic Spatial Relationships in Video Generation
by: Liu, Fengming, et al.
Published: (2026)
by: Liu, Fengming, et al.
Published: (2026)
EchoSight: Advancing Visual-Language Models with Wiki Knowledge
by: Yan, Yibin, et al.
Published: (2024)
by: Yan, Yibin, et al.
Published: (2024)
Similar Items
-
A General Protocol to Probe Large Vision Models for 3D Physical Understanding
by: Zhan, Guanqi, et al.
Published: (2023) -
Inferring Dynamic Physical Properties from Video Foundation Models
by: Zhan, Guanqi, et al.
Published: (2025) -
ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval
by: Zhan, Guanqi, et al.
Published: (2025) -
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
by: Xie, Junyu, et al.
Published: (2026) -
Appearance-Based Refinement for Object-Centric Motion Segmentation
by: Xie, Junyu, et al.
Published: (2023)