Multi-Context Temporal Consistent Modeling for Referring Video Object Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Choi, Sun-Hyuk, Jo, Hayoung, Lee, Seong-Whan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Temporally Consistent Referring Video Object Segmentation with Hybrid Memory
by: Miao, Bo, et al.
Published: (2024)
by: Miao, Bo, et al.
Published: (2024)
Latest Object Memory Management for Temporally Consistent Video Instance Segmentation
by: Lee, Seunghun, et al.
Published: (2025)
by: Lee, Seunghun, et al.
Published: (2025)
LTCA: Long-range Temporal Context Attention for Referring Video Object Segmentation
by: Yan, Cilin, et al.
Published: (2025)
by: Yan, Cilin, et al.
Published: (2025)
Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
by: Lee, Seunghun, et al.
Published: (2025)
by: Lee, Seunghun, et al.
Published: (2025)
AM-SORT: Adaptable Motion Predictor with Historical Trajectory Embedding for Multi-Object Tracking
by: Kim, Vitaliy, et al.
Published: (2024)
by: Kim, Vitaliy, et al.
Published: (2024)
Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
by: Lin, Ci-Siang, et al.
Published: (2025)
by: Lin, Ci-Siang, et al.
Published: (2025)
Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
by: Zhang, Ruixin, et al.
Published: (2025)
by: Zhang, Ruixin, et al.
Published: (2025)
DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer
by: Kim, Ho-Joong, et al.
Published: (2025)
by: Kim, Ho-Joong, et al.
Published: (2025)
MCoT-RE: Multi-Faceted Chain-of-Thought and Re-Ranking for Training-Free Zero-Shot Composed Image Retrieval
by: Park, Jeong-Woo, et al.
Published: (2025)
by: Park, Jeong-Woo, et al.
Published: (2025)
Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
by: Kim, Yehna, et al.
Published: (2025)
by: Kim, Yehna, et al.
Published: (2025)
Tsanet: Temporal and Scale Alignment for Unsupervised Video Object Segmentation
by: Lee, Seunghoon, et al.
Published: (2023)
by: Lee, Seunghoon, et al.
Published: (2023)
Harnessing Vision-Language Pretrained Models with Temporal-Aware Adaptation for Referring Video Object Segmentation
by: Zhou, Zikun, et al.
Published: (2024)
by: Zhou, Zikun, et al.
Published: (2024)
Temporal-Enhanced Multimodal Transformer for Referring Multi-Object Tracking and Segmentation
by: Xiao, Changcheng, et al.
Published: (2024)
by: Xiao, Changcheng, et al.
Published: (2024)
InterRVOS: Interaction-aware Referring Video Object Segmentation
by: Jin, Woojeong, et al.
Published: (2025)
by: Jin, Woojeong, et al.
Published: (2025)
Referring Video Object Segmentation via Language-aligned Track Selection
by: Kim, Seongchan, et al.
Published: (2024)
by: Kim, Seongchan, et al.
Published: (2024)
TIFu: Tri-directional Implicit Function for High-Fidelity 3D Character Reconstruction
by: Lim, Byoungsung, et al.
Published: (2024)
by: Lim, Byoungsung, et al.
Published: (2024)
Referring Video Object Segmentation with Cross-Modality Proxy Queries
by: Sun, Baoli, et al.
Published: (2025)
by: Sun, Baoli, et al.
Published: (2025)
Find First, Track Next: Decoupling Identification and Propagation in Referring Video Object Segmentation
by: Cho, Suhwan, et al.
Published: (2025)
by: Cho, Suhwan, et al.
Published: (2025)
Efficient One-stage Video Object Detection by Exploiting Temporal Consistency
by: Sun, Guanxiong, et al.
Published: (2024)
by: Sun, Guanxiong, et al.
Published: (2024)
Dual Prototype Attention for Unsupervised Video Object Segmentation
by: Cho, Suhwan, et al.
Published: (2022)
by: Cho, Suhwan, et al.
Published: (2022)
CAVIS: Context-Aware Video Instance Segmentation
by: Lee, Seunghun, et al.
Published: (2024)
by: Lee, Seunghun, et al.
Published: (2024)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation
by: Jiang, Haichao, et al.
Published: (2026)
by: Jiang, Haichao, et al.
Published: (2026)
AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation
by: Jin, Woojeong, et al.
Published: (2026)
by: Jin, Woojeong, et al.
Published: (2026)
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
by: Zhu, Zixin, et al.
Published: (2024)
by: Zhu, Zixin, et al.
Published: (2024)
FlipConcept: Tuning-Free Multi-Concept Personalization for Text-to-Image Generation
by: Woo, Young Beom, et al.
Published: (2025)
by: Woo, Young Beom, et al.
Published: (2025)
FAR-Net: Multi-Stage Fusion Network with Enhanced Semantic Alignment and Adaptive Reconciliation for Composed Image Retrieval
by: Park, Jeong-Woo, et al.
Published: (2025)
by: Park, Jeong-Woo, et al.
Published: (2025)
Temporally Consistent Object Editing in Videos using Extended Attention
by: Zamani, AmirHossein, et al.
Published: (2024)
by: Zamani, AmirHossein, et al.
Published: (2024)
TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate Expression
by: Kim, Ho-Joong, et al.
Published: (2024)
by: Kim, Ho-Joong, et al.
Published: (2024)
MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation
by: Rong, Fu, et al.
Published: (2025)
by: Rong, Fu, et al.
Published: (2025)
CW-BASS: Confidence-Weighted Boundary-Aware Learning for Semi-Supervised Semantic Segmentation
by: Tarubinga, Ebenezer, et al.
Published: (2025)
by: Tarubinga, Ebenezer, et al.
Published: (2025)
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models
by: Chen, Junzhe, et al.
Published: (2026)
by: Chen, Junzhe, et al.
Published: (2026)
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
by: Kim, Ji-Hyeon, et al.
Published: (2026)
by: Kim, Ji-Hyeon, et al.
Published: (2026)
Few-Shot Video Object Segmentation in X-Ray Angiography Using Local Matching and Spatio-Temporal Consistency Loss
by: Xi, Lin, et al.
Published: (2026)
by: Xi, Lin, et al.
Published: (2026)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
by: Oh, Ju-Young, et al.
Published: (2025)
by: Oh, Ju-Young, et al.
Published: (2025)
Multi-Granularity Video Object Segmentation
by: Lim, Sangbeom, et al.
Published: (2024)
by: Lim, Sangbeom, et al.
Published: (2024)
Context Forcing: Consistent Autoregressive Video Generation with Long Context
by: Chen, Shuo, et al.
Published: (2026)
by: Chen, Shuo, et al.
Published: (2026)
Spatio-Temporal Attention for Consistent Video Semantic Segmentation in Automated Driving
by: Varghese, Serin, et al.
Published: (2026)
by: Varghese, Serin, et al.
Published: (2026)
CRT-Fusion: Camera, Radar, Temporal Fusion Using Motion Information for 3D Object Detection
by: Kim, Jisong, et al.
Published: (2024)
by: Kim, Jisong, et al.
Published: (2024)
AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation
by: Girish, Sharath, et al.
Published: (2025)
by: Girish, Sharath, et al.
Published: (2025)
Similar Items
-
Temporally Consistent Referring Video Object Segmentation with Hybrid Memory
by: Miao, Bo, et al.
Published: (2024) -
Latest Object Memory Management for Temporally Consistent Video Instance Segmentation
by: Lee, Seunghun, et al.
Published: (2025) -
LTCA: Long-range Temporal Context Attention for Referring Video Object Segmentation
by: Yan, Cilin, et al.
Published: (2025) -
Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
by: Lee, Seunghun, et al.
Published: (2025) -
AM-SORT: Adaptable Motion Predictor with Historical Trajectory Embedding for Multi-Object Tracking
by: Kim, Vitaliy, et al.
Published: (2024)