Saved in:
| Main Authors: | Zhong, Qing, Ding, Guodong, Liu, Lingqiao, Feng, Zaiwen, Wu, Lin Yuanbo, Yao, Angela |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.08805 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OnlineTAS: An Online Baseline for Temporal Action Segmentation
by: Zhong, Qing, et al.
Published: (2024)
by: Zhong, Qing, et al.
Published: (2024)
Coherent Temporal Synthesis for Incremental Action Segmentation
by: Ding, Guodong, et al.
Published: (2024)
by: Ding, Guodong, et al.
Published: (2024)
Condensing Action Segmentation Datasets via Generative Network Inversion
by: Ding, Guodong, et al.
Published: (2025)
by: Ding, Guodong, et al.
Published: (2025)
A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation
by: Zhong, Qing, et al.
Published: (2025)
by: Zhong, Qing, et al.
Published: (2025)
Gate-and-Merge: Zero-shot Compositional Personalization of Vision Language Models
by: Ding, Guodong, et al.
Published: (2026)
by: Ding, Guodong, et al.
Published: (2026)
Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models
by: Yadav, Ankit, et al.
Published: (2025)
by: Yadav, Ankit, et al.
Published: (2025)
The CLIP Model is Secretly an Image-to-Prompt Converter
by: Ding, Yuxuan, et al.
Published: (2023)
by: Ding, Yuxuan, et al.
Published: (2023)
Categorical Keypoint Positional Embedding for Robust Animal Re-Identification
by: Lin, Yuhao, et al.
Published: (2024)
by: Lin, Yuhao, et al.
Published: (2024)
Training-Free Instance-Aware 3D Scene Reconstruction and Diffusion-Based View Synthesis from Sparse Images
by: Xia, Jiatong, et al.
Published: (2026)
by: Xia, Jiatong, et al.
Published: (2026)
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
by: Ying, Kaining, et al.
Published: (2025)
by: Ying, Kaining, et al.
Published: (2025)
Implicit Counterfactual Learning for Audio-Visual Segmentation
by: Zha, Mingfeng, et al.
Published: (2025)
by: Zha, Mingfeng, et al.
Published: (2025)
Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
Blended Latent Diffusion under Attention Control for Real-World Video Editing
by: Liu, Deyin, et al.
Published: (2024)
by: Liu, Deyin, et al.
Published: (2024)
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
by: Huang, Shaofei, et al.
Published: (2025)
by: Huang, Shaofei, et al.
Published: (2025)
Close-up-GS: Enhancing Close-Up View Synthesis in 3D Gaussian Splatting with Progressive Self-Training
by: Xia, Jiatong, et al.
Published: (2025)
by: Xia, Jiatong, et al.
Published: (2025)
Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues
by: Chen, Tianxiang, et al.
Published: (2024)
by: Chen, Tianxiang, et al.
Published: (2024)
SeaVIS: Sound-Enhanced Association for Online Audio-Visual Instance Segmentation
by: Zhu, Yingjian, et al.
Published: (2026)
by: Zhu, Yingjian, et al.
Published: (2026)
A Simple-but-effective Baseline for Training-free Class-Agnostic Counting
by: Lin, Yuhao, et al.
Published: (2024)
by: Lin, Yuhao, et al.
Published: (2024)
Complementary and Contrastive Learning for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
Unveiling and Mitigating Bias in Audio Visual Segmentation
by: Sun, Peiwen, et al.
Published: (2024)
by: Sun, Peiwen, et al.
Published: (2024)
Enhancing Fine-Grained Visual Recognition in the Low-Data Regime Through Feature Magnitude Regularization
by: Chapman, Avraham, et al.
Published: (2024)
by: Chapman, Avraham, et al.
Published: (2024)
Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation
by: Shen, Yunzhe, et al.
Published: (2025)
by: Shen, Yunzhe, et al.
Published: (2025)
MAGIC: Map-Guided Few-Shot Audio-Visual Acoustics Modeling
by: Huang, Diwei, et al.
Published: (2024)
by: Huang, Diwei, et al.
Published: (2024)
TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models
by: Luo, Ziyang, et al.
Published: (2025)
by: Luo, Ziyang, et al.
Published: (2025)
Boosting Box-supervised Instance Segmentation with Pseudo Depth
by: Yu, Xinyi, et al.
Published: (2024)
by: Yu, Xinyi, et al.
Published: (2024)
On Learning Discriminative Features from Synthesized Data for Self-Supervised Fine-Grained Visual Recognition
by: Wang, Zihu, et al.
Published: (2024)
by: Wang, Zihu, et al.
Published: (2024)
One Last Attention for Your Vision-Language Model
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
Advancing Weakly-Supervised Audio-Visual Video Parsing via Segment-wise Pseudo Labeling
by: Zhou, Jinxing, et al.
Published: (2024)
by: Zhou, Jinxing, et al.
Published: (2024)
Revisiting Vision Language Foundations for No-Reference Image Quality Assessment
by: Yadav, Ankit, et al.
Published: (2025)
by: Yadav, Ankit, et al.
Published: (2025)
EMAG: Self-Rectifying Diffusion Sampling with Exponential Moving Average Guidance
by: Yadav, Ankit, et al.
Published: (2025)
by: Yadav, Ankit, et al.
Published: (2025)
Prompt-Driven Lightweight Foundation Model for Instance Segmentation-Based Fault Detection in Freight Trains
by: Sun, Guodong, et al.
Published: (2026)
by: Sun, Guodong, et al.
Published: (2026)
A Novel Local Focusing Mechanism for Deepfake Detection Generalization
by: Li, Mingliang, et al.
Published: (2025)
by: Li, Mingliang, et al.
Published: (2025)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
by: Hong, Yuyang, et al.
Published: (2025)
by: Hong, Yuyang, et al.
Published: (2025)
Extending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual Segmentation
by: Seon, Juhyeong, et al.
Published: (2024)
by: Seon, Juhyeong, et al.
Published: (2024)
Improving Online Source-free Domain Adaptation for Object Detection by Unsupervised Data Acquisition
by: Shi, Xiangyu, et al.
Published: (2023)
by: Shi, Xiangyu, et al.
Published: (2023)
EZIGen: Enhancing zero-shot personalized image generation with precise subject encoding and decoupled guidance
by: Duan, Zicheng, et al.
Published: (2024)
by: Duan, Zicheng, et al.
Published: (2024)
Denoise to Track: Harnessing Video Diffusion Priors for Robust Correspondence
by: Yuan, Tianyu, et al.
Published: (2025)
by: Yuan, Tianyu, et al.
Published: (2025)
Unraveling Instance Associations: A Closer Look for Audio-Visual Segmentation
by: Chen, Yuanhong, et al.
Published: (2023)
by: Chen, Yuanhong, et al.
Published: (2023)
From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
Regularizing Neural Network Training via Identity-wise Discriminative Feature Suppression
by: Chapman, Avraham, et al.
Published: (2022)
by: Chapman, Avraham, et al.
Published: (2022)
Similar Items
-
OnlineTAS: An Online Baseline for Temporal Action Segmentation
by: Zhong, Qing, et al.
Published: (2024) -
Coherent Temporal Synthesis for Incremental Action Segmentation
by: Ding, Guodong, et al.
Published: (2024) -
Condensing Action Segmentation Datasets via Generative Network Inversion
by: Ding, Guodong, et al.
Published: (2025) -
A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation
by: Zhong, Qing, et al.
Published: (2025) -
Gate-and-Merge: Zero-shot Compositional Personalization of Vision Language Models
by: Ding, Guodong, et al.
Published: (2026)