PVG: Progressive Vision Graph for Vision Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Jiafu, Li, Jian, Zhang, Jiangning, Zhang, Boshen, Chi, Mingmin, Wang, Yabiao, Wang, Chengjie |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PiT: Progressive Diffusion Transformer
by: Wu, Jiafu, et al.
Published: (2025)
by: Wu, Jiafu, et al.
Published: (2025)
MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation
by: Mao, Xiaofeng, et al.
Published: (2024)
by: Mao, Xiaofeng, et al.
Published: (2024)
EATFormer: Improving Vision Transformer Inspired by Evolutionary Algorithm
by: Zhang, Jiangning, et al.
Published: (2022)
by: Zhang, Jiangning, et al.
Published: (2022)
AdapNet: Adaptive Noise-Based Network for Low-Quality Image Retrieval
by: Zhang, Sihe, et al.
Published: (2024)
by: Zhang, Sihe, et al.
Published: (2024)
EMOv2: Pushing 5M Vision Model Frontier
by: Zhang, Jiangning, et al.
Published: (2024)
by: Zhang, Jiangning, et al.
Published: (2024)
MARRS: Masked Autoregressive Unit-based Reaction Synthesis
by: Wang, Yabiao, et al.
Published: (2025)
by: Wang, Yabiao, et al.
Published: (2025)
Self-supervised Feature Adaptation for 3D Industrial Anomaly Detection
by: Tu, Yuanpeng, et al.
Published: (2024)
by: Tu, Yuanpeng, et al.
Published: (2024)
Self-Supervised Likelihood Estimation with Energy Guidance for Anomaly Segmentation in Urban Scenes
by: Tu, Yuanpeng, et al.
Published: (2023)
by: Tu, Yuanpeng, et al.
Published: (2023)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
by: Jin, Yizhang, et al.
Published: (2024)
by: Jin, Yizhang, et al.
Published: (2024)
TIMotion: Temporal and Interactive Framework for Efficient Human-Human Motion Generation
by: Wang, Yabiao, et al.
Published: (2024)
by: Wang, Yabiao, et al.
Published: (2024)
OSV: One Step is Enough for High-Quality Image to Video Generation
by: Mao, Xiaofeng, et al.
Published: (2024)
by: Mao, Xiaofeng, et al.
Published: (2024)
PointRWKV: Efficient RWKV-Like Model for Hierarchical Point Cloud Learning
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
SaRA: High-Efficient Diffusion Model Fine-tuning with Progressive Sparse Low-Rank Adaptation
by: Hu, Teng, et al.
Published: (2024)
by: Hu, Teng, et al.
Published: (2024)
CLIP-AD: A Language-Guided Staged Dual-Path Model for Zero-shot Anomaly Detection
by: Chen, Xuhai, et al.
Published: (2023)
by: Chen, Xuhai, et al.
Published: (2023)
Leveraging Fine-Grained Information and Noise Decoupling for Remote Sensing Change Detection
by: Du, Qiangang, et al.
Published: (2024)
by: Du, Qiangang, et al.
Published: (2024)
Mamba-YOLO-World: Marrying YOLO-World with Mamba for Open-Vocabulary Detection
by: Wang, Haoxuan, et al.
Published: (2024)
by: Wang, Haoxuan, et al.
Published: (2024)
AnomalyDiffusion: Few-Shot Anomaly Image Generation with Diffusion Model
by: Hu, Teng, et al.
Published: (2023)
by: Hu, Teng, et al.
Published: (2023)
GPT-4V-AD: Exploring Grounding Potential of VQA-oriented GPT-4V for Zero-shot Anomaly Detection
by: Zhang, Jiangning, et al.
Published: (2023)
by: Zhang, Jiangning, et al.
Published: (2023)
Reference Twice: A Simple and Unified Baseline for Few-Shot Instance Segmentation
by: Han, Yue, et al.
Published: (2023)
by: Han, Yue, et al.
Published: (2023)
Exploring Plain ViT Reconstruction for Multi-class Unsupervised Anomaly Detection
by: Zhang, Jiangning, et al.
Published: (2023)
by: Zhang, Jiangning, et al.
Published: (2023)
Dual-Interrelated Diffusion Model for Few-Shot Anomaly Image Generation
by: Jin, Ying, et al.
Published: (2024)
by: Jin, Ying, et al.
Published: (2024)
Textual Decomposition Then Sub-motion-space Scattering for Open-Vocabulary Motion Generation
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
Learning Unified Reference Representation for Unsupervised Multi-class Anomaly Detection
by: He, Liren, et al.
Published: (2024)
by: He, Liren, et al.
Published: (2024)
VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation
by: Wang, Qilin, et al.
Published: (2024)
by: Wang, Qilin, et al.
Published: (2024)
SwiftVideo: A Unified Framework for Few-Step Video Generation through Trajectory-Distribution Alignment
by: Sun, Yanxiao, et al.
Published: (2025)
by: Sun, Yanxiao, et al.
Published: (2025)
Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning
by: He, Qingdong, et al.
Published: (2025)
by: He, Qingdong, et al.
Published: (2025)
MobileMamba: Lightweight Multi-Receptive Visual Mamba Network
by: He, Haoyang, et al.
Published: (2024)
by: He, Haoyang, et al.
Published: (2024)
FreeMotion: A Unified Framework for Number-free Text-to-Motion Synthesis
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
Semantic Frame Interpolation
by: Hong, Yijia, et al.
Published: (2025)
by: Hong, Yijia, et al.
Published: (2025)
TSCM: A Teacher-Student Model for Vision Place Recognition Using Cross-Metric Knowledge Distillation
by: Shen, Yehui, et al.
Published: (2024)
by: Shen, Yehui, et al.
Published: (2024)
CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware Diffusion
by: Ji, Chenhao, et al.
Published: (2025)
by: Ji, Chenhao, et al.
Published: (2025)
A Comprehensive Library for Benchmarking Multi-class Visual Anomaly Detection
by: Zhang, Jiangning, et al.
Published: (2024)
by: Zhang, Jiangning, et al.
Published: (2024)
Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10$\times$
by: Zhang, Jiangning, et al.
Published: (2025)
by: Zhang, Jiangning, et al.
Published: (2025)
The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection
by: He, Qingdong, et al.
Published: (2025)
by: He, Qingdong, et al.
Published: (2025)
In-context Prompt Learning for Test-time Vision Recognition with Frozen Vision-language Model
by: Yin, Junhui, et al.
Published: (2024)
by: Yin, Junhui, et al.
Published: (2024)
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
by: Xue, Zhucun, et al.
Published: (2025)
by: Xue, Zhucun, et al.
Published: (2025)
DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
Similar Items
-
PiT: Progressive Diffusion Transformer
by: Wu, Jiafu, et al.
Published: (2025) -
MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation
by: Mao, Xiaofeng, et al.
Published: (2024) -
EATFormer: Improving Vision Transformer Inspired by Evolutionary Algorithm
by: Zhang, Jiangning, et al.
Published: (2022) -
AdapNet: Adaptive Noise-Based Network for Low-Quality Image Retrieval
by: Zhang, Sihe, et al.
Published: (2024) -
EMOv2: Pushing 5M Vision Model Frontier
by: Zhang, Jiangning, et al.
Published: (2024)