General and Task-Oriented Video Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Mu, Li, Liulei, Wang, Wenguan, Quan, Ruijie, Yang, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation
by: Chen, Mu, et al.
Published: (2025)
by: Chen, Mu, et al.
Published: (2025)
Shape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data
by: Feng, Tuo, et al.
Published: (2024)
by: Feng, Tuo, et al.
Published: (2024)
Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
Clustering Propagation for Universal Medical Image Segmentation
by: Ding, Yuhang, et al.
Published: (2024)
by: Ding, Yuhang, et al.
Published: (2024)
Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion Models
by: Li, Liulei, et al.
Published: (2024)
by: Li, Liulei, et al.
Published: (2024)
Nonverbal Interaction Detection
by: Wei, Jianan, et al.
Published: (2024)
by: Wei, Jianan, et al.
Published: (2024)
Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation
by: You, Yuyang, et al.
Published: (2026)
by: You, Yuyang, et al.
Published: (2026)
Controllable Navigation Instruction Generation with Chain of Thought Prompting
by: Kong, Xianghao, et al.
Published: (2024)
by: Kong, Xianghao, et al.
Published: (2024)
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
by: Xu, Qi'ao, et al.
Published: (2025)
by: Xu, Qi'ao, et al.
Published: (2025)
Visual Knowledge in the Big Model Era: Retrospect and Prospect
by: Wang, Wenguan, et al.
Published: (2024)
by: Wang, Wenguan, et al.
Published: (2024)
Towards Data-and Knowledge-Driven Artificial Intelligence: A Survey on Neuro-Symbolic Computing
by: Wang, Wenguan, et al.
Published: (2022)
by: Wang, Wenguan, et al.
Published: (2022)
TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
by: Liu, Xiangrui, et al.
Published: (2025)
by: Liu, Xiangrui, et al.
Published: (2025)
PiPa++: Towards Unification of Domain Adaptive Semantic Segmentation via Self-supervised Learning
by: Chen, Mu, et al.
Published: (2024)
by: Chen, Mu, et al.
Published: (2024)
A Survey on 3D Gaussian Splatting
by: Chen, Guikun, et al.
Published: (2024)
by: Chen, Guikun, et al.
Published: (2024)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
Psychometry: An Omnifit Model for Image Reconstruction from Human Brain Activity
by: Quan, Ruijie, et al.
Published: (2024)
by: Quan, Ruijie, et al.
Published: (2024)
Improving Bird's Eye View Semantic Segmentation by Task Decomposition
by: Zhao, Tianhao, et al.
Published: (2024)
by: Zhao, Tianhao, et al.
Published: (2024)
AoP-SAM: Automation of Prompts for Efficient Segmentation
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
Not All Pixels Are Equal: Pixel-wise Meta-Learning for Medical Segmentation with Noisy Labels
by: Mu, Chenyu, et al.
Published: (2025)
by: Mu, Chenyu, et al.
Published: (2025)
Theoretically Achieving Continuous Representation of Oriented Bounding Boxes
by: Xiao, Zi-Kai, et al.
Published: (2024)
by: Xiao, Zi-Kai, et al.
Published: (2024)
Space-time Reinforcement Network for Video Object Segmentation
by: Chen, Yadang, et al.
Published: (2024)
by: Chen, Yadang, et al.
Published: (2024)
Mutual Learning for Acoustic Matching and Dereverberation via Visual Scene-driven Diffusion
by: Ma, Jian, et al.
Published: (2024)
by: Ma, Jian, et al.
Published: (2024)
Learning Spatial-Semantic Features for Robust Video Object Segmentation
by: Li, Xin, et al.
Published: (2024)
by: Li, Xin, et al.
Published: (2024)
PKINet-v2: Towards Powerful and Efficient Poly-Kernel Remote Sensing Object Detection
by: Cai, Xinhao, et al.
Published: (2026)
by: Cai, Xinhao, et al.
Published: (2026)
Utilizing Graph Generation for Enhanced Domain Adaptive Object Detection
by: Wang, Mu
Published: (2024)
by: Wang, Mu
Published: (2024)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
by: Yan, Xin, et al.
Published: (2024)
by: Yan, Xin, et al.
Published: (2024)
Scalable Video Object Segmentation with Identification Mechanism
by: Yang, Zongxin, et al.
Published: (2022)
by: Yang, Zongxin, et al.
Published: (2022)
Conditional Video Generation for High-Efficiency Video Compression
by: Yi, Fangqiu, et al.
Published: (2025)
by: Yi, Fangqiu, et al.
Published: (2025)
MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos
by: Wang, Rongsheng, et al.
Published: (2025)
by: Wang, Rongsheng, et al.
Published: (2025)
DepthPilot: From Controllability to Interpretability in Colonoscopy Video Generation
by: Fu, Junhu, et al.
Published: (2026)
by: Fu, Junhu, et al.
Published: (2026)
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
by: Chen, Junsong, et al.
Published: (2025)
by: Chen, Junsong, et al.
Published: (2025)
Task Consistent Prototype Learning for Incremental Few-shot Semantic Segmentation
by: Xu, Wenbo, et al.
Published: (2024)
by: Xu, Wenbo, et al.
Published: (2024)
Video-Infinity: Distributed Long Video Generation
by: Tan, Zhenxiong, et al.
Published: (2024)
by: Tan, Zhenxiong, et al.
Published: (2024)
Knowledge-Intensive Video Generation
by: Wang, Chenxu, et al.
Published: (2026)
by: Wang, Chenxu, et al.
Published: (2026)
Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
by: Jisheng, Dang, et al.
Published: (2025)
by: Jisheng, Dang, et al.
Published: (2025)
Harnessing Group-Oriented Consistency Constraints for Semi-Supervised Semantic Segmentation in CdZnTe Semiconductors
by: Li, Peihao, et al.
Published: (2025)
by: Li, Peihao, et al.
Published: (2025)
Metric-Guided Feature Fusion of Visual Foundation Models for Segmentation Tasks
by: Guo, Yachan, et al.
Published: (2026)
by: Guo, Yachan, et al.
Published: (2026)
Unbiased Object Detection Beyond Frequency with Visually Prompted Image Synthesis
by: Cai, Xinhao, et al.
Published: (2025)
by: Cai, Xinhao, et al.
Published: (2025)
ColoDiff: Integrating Dynamic Consistency With Content Awareness for Colonoscopy Video Generation
by: Fu, Junhu, et al.
Published: (2026)
by: Fu, Junhu, et al.
Published: (2026)
HOTVCOM: Generating Buzzworthy Comments for Videos
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Similar Items
-
DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation
by: Chen, Mu, et al.
Published: (2025) -
Shape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data
by: Feng, Tuo, et al.
Published: (2024) -
Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
by: Zhang, Xu, et al.
Published: (2025) -
Clustering Propagation for Universal Medical Image Segmentation
by: Ding, Yuhang, et al.
Published: (2024) -
Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion Models
by: Li, Liulei, et al.
Published: (2024)