VISAGE: Video Instance Segmentation with Appearance-Guided Enhancement
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Hanjung, Kang, Jaehyun, Heo, Miran, Hwang, Sukjun, Oh, Seoung Wug, Kim, Seon Joo |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Autoregressive Universal Video Segmentation Model
by: Heo, Miran, et al.
Published: (2025)
by: Heo, Miran, et al.
Published: (2025)
IF-MDM: Implicit Face Motion Diffusion Model for High-Fidelity Realtime Talking Head Generation
by: Yang, Sejong, et al.
Published: (2024)
by: Yang, Sejong, et al.
Published: (2024)
UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
by: Kim, Hanjung, et al.
Published: (2025)
by: Kim, Hanjung, et al.
Published: (2025)
MaGGIe: Masked Guided Gradual Human Instance Matting
by: Huynh, Chuong, et al.
Published: (2024)
by: Huynh, Chuong, et al.
Published: (2024)
Elevating Flow-Guided Video Inpainting with Reference Generation
by: Cho, Suhwan, et al.
Published: (2024)
by: Cho, Suhwan, et al.
Published: (2024)
VideoMaMa: Mask-Guided Video Matting via Generative Prior
by: Lim, Sangbeom, et al.
Published: (2026)
by: Lim, Sangbeom, et al.
Published: (2026)
Putting the Object Back into Video Object Segmentation
by: Cheng, Ho Kei, et al.
Published: (2023)
by: Cheng, Ho Kei, et al.
Published: (2023)
Tuning-Free Multi-Event Long Video Generation via Synchronized Coupled Sampling
by: Kim, Subin, et al.
Published: (2025)
by: Kim, Subin, et al.
Published: (2025)
Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks
by: Heo, Miran, et al.
Published: (2025)
by: Heo, Miran, et al.
Published: (2025)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
by: Hyun, Jeongseok, et al.
Published: (2025)
by: Hyun, Jeongseok, et al.
Published: (2025)
Hierarchically Structured Neural Bones for Reconstructing Animatable Objects from Casual Videos
by: Jeon, Subin, et al.
Published: (2024)
by: Jeon, Subin, et al.
Published: (2024)
In-N-Out: Faithful 3D GAN Inversion with Volumetric Decomposition for Face Editing
by: Xu, Yiran, et al.
Published: (2023)
by: Xu, Yiran, et al.
Published: (2023)
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation
by: Niu, Quanzhu, et al.
Published: (2025)
by: Niu, Quanzhu, et al.
Published: (2025)
Leveraging Image Augmentation for Object Manipulation: Towards Interpretable Controllability in Object-Centric Learning
by: Kim, Jinwoo, et al.
Published: (2023)
by: Kim, Jinwoo, et al.
Published: (2023)
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
by: TV, Sethuraman, et al.
Published: (2025)
by: TV, Sethuraman, et al.
Published: (2025)
ActionSwitch: Class-agnostic Detection of Simultaneous Actions in Streaming Videos
by: Kang, Hyolim, et al.
Published: (2024)
by: Kang, Hyolim, et al.
Published: (2024)
Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
by: Kang, Hyolim, et al.
Published: (2025)
by: Kang, Hyolim, et al.
Published: (2025)
Seam360GS: Seamless 360° Gaussian Splatting from Real-World Omnidirectional Images
by: Shin, Changha, et al.
Published: (2025)
by: Shin, Changha, et al.
Published: (2025)
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
by: Han, Su Ho, et al.
Published: (2025)
by: Han, Su Ho, et al.
Published: (2025)
Rethinking Saliency-Guided Weakly-Supervised Semantic Segmentation
by: Kim, Beomyoung, et al.
Published: (2024)
by: Kim, Beomyoung, et al.
Published: (2024)
Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image Diffusion
by: Kim, Jiwon, et al.
Published: (2025)
by: Kim, Jiwon, et al.
Published: (2025)
Query-Calibrated Segmental Admission for Descriptor-Agnostic LiDAR Loop Closure in Repetitive Environments
by: Kim, Jaehyun, et al.
Published: (2025)
by: Kim, Jaehyun, et al.
Published: (2025)
HARIVO: Harnessing Text-to-Image Models for Video Generation
by: Kwon, Mingi, et al.
Published: (2024)
by: Kwon, Mingi, et al.
Published: (2024)
Learning to Enhance Aperture Phasor Field for Non-Line-of-Sight Imaging
by: Cho, In, et al.
Published: (2024)
by: Cho, In, et al.
Published: (2024)
VISAGE: Video Synthesis using Action Graphs for Surgery
by: Yeganeh, Yousef, et al.
Published: (2024)
by: Yeganeh, Yousef, et al.
Published: (2024)
Attentive Illumination Decomposition Model for Multi-Illuminant White Balancing
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
PRIMEdit: Probability Redistribution for Instance-aware Multi-object Video Editing with Benchmark Dataset
by: Teodoro, Samuel, et al.
Published: (2024)
by: Teodoro, Samuel, et al.
Published: (2024)
DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation
by: Ngo, Tuan Duc, et al.
Published: (2026)
by: Ngo, Tuan Duc, et al.
Published: (2026)
Boundary-Recovering Network for Temporal Action Detection
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Global Geometry Is Not Enough for Vision Representations
by: Chung, Jiwan, et al.
Published: (2026)
by: Chung, Jiwan, et al.
Published: (2026)
Object Aware Egocentric Online Action Detection
by: An, Joungbin, et al.
Published: (2024)
by: An, Joungbin, et al.
Published: (2024)
Latent Space Super-Resolution for Higher-Resolution Image Generation with Diffusion Models
by: Jeong, Jinho, et al.
Published: (2025)
by: Jeong, Jinho, et al.
Published: (2025)
Accelerating Image Super-Resolution Networks with Pixel-Level Classification
by: Jeong, Jinho, et al.
Published: (2024)
by: Jeong, Jinho, et al.
Published: (2024)
A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
by: Lee, Jaeseong, et al.
Published: (2025)
by: Lee, Jaeseong, et al.
Published: (2025)
V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping
by: Lee, Hyunkoo, et al.
Published: (2025)
by: Lee, Hyunkoo, et al.
Published: (2025)
4D Scaffold Gaussian Splatting with Dynamic-Aware Anchor Growing for Efficient and High-Fidelity Dynamic Scene Reconstruction
by: Cho, Woong Oh, et al.
Published: (2024)
by: Cho, Woong Oh, et al.
Published: (2024)
Extreme Point Supervised Instance Segmentation
by: Lee, Hyeonjun, et al.
Published: (2024)
by: Lee, Hyeonjun, et al.
Published: (2024)
CCMNet: Leveraging Calibrated Color Correction Matrices for Cross-Camera Color Constancy
by: Kim, Dongyoung, et al.
Published: (2025)
by: Kim, Dongyoung, et al.
Published: (2025)
ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors
by: Jeon, Subin, et al.
Published: (2025)
by: Jeon, Subin, et al.
Published: (2025)
Similar Items
-
Autoregressive Universal Video Segmentation Model
by: Heo, Miran, et al.
Published: (2025) -
IF-MDM: Implicit Face Motion Diffusion Model for High-Fidelity Realtime Talking Head Generation
by: Yang, Sejong, et al.
Published: (2024) -
UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
by: Kim, Hanjung, et al.
Published: (2025) -
MaGGIe: Masked Guided Gradual Human Instance Matting
by: Huynh, Chuong, et al.
Published: (2024) -
Elevating Flow-Guided Video Inpainting with Reference Generation
by: Cho, Suhwan, et al.
Published: (2024)