LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Mo, Shentong, Wang, Yansen, Luo, Xufang, Li, Dongsheng |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
A Large-scale Medical Visual Task Adaptation Benchmark
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
par: Mo, Shentong, et autres
Publié: (2026)
par: Mo, Shentong, et autres
Publié: (2026)
Improving Visual Representation Alignment Generation with GRPO
par: Mo, Shentong, et autres
Publié: (2026)
par: Mo, Shentong, et autres
Publié: (2026)
Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
par: Mo, Shentong
Publié: (2024)
par: Mo, Shentong
Publié: (2024)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
par: Mo, Shentong, et autres
Publié: (2026)
par: Mo, Shentong, et autres
Publié: (2026)
The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
par: Mo, Shentong
Publié: (2024)
par: Mo, Shentong
Publié: (2024)
GMAIL: Generative Modality Alignment for generated Image Learning
par: Mo, Shentong, et autres
Publié: (2026)
par: Mo, Shentong, et autres
Publié: (2026)
Learning to Instruct for Visual Instruction Tuning
par: Zhou, Zhihan, et autres
Publié: (2025)
par: Zhou, Zhihan, et autres
Publié: (2025)
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
par: Mo, Shentong, et autres
Publié: (2026)
par: Mo, Shentong, et autres
Publié: (2026)
DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
MultiMed: Massively Multimodal and Multitask Medical Understanding
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
par: Li, Xu, et autres
Publié: (2025)
par: Li, Xu, et autres
Publié: (2025)
EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
par: Lin, Nie, et autres
Publié: (2025)
par: Lin, Nie, et autres
Publié: (2025)
Learning Sparse Visual Representations via Spatial-Semantic Factorization
par: Zhao, Theodore Zhengde, et autres
Publié: (2026)
par: Zhao, Theodore Zhengde, et autres
Publié: (2026)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
par: Mo, Shentong, et autres
Publié: (2023)
par: Mo, Shentong, et autres
Publié: (2023)
Adaptive Prompt Tuning: Vision Guided Prompt Tuning with Cross-Attention for Fine-Grained Few-Shot Learning
par: Brouwer, Eric, et autres
Publié: (2024)
par: Brouwer, Eric, et autres
Publié: (2024)
GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
par: Mo, Shentong, et autres
Publié: (2026)
par: Mo, Shentong, et autres
Publié: (2026)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
par: Woo, Byeongju, et autres
Publié: (2026)
par: Woo, Byeongju, et autres
Publié: (2026)
ADAPT to Robustify Prompt Tuning Vision Transformers
par: Eskandar, Masih, et autres
Publié: (2024)
par: Eskandar, Masih, et autres
Publié: (2024)
Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models
par: Song, Fei, et autres
Publié: (2025)
par: Song, Fei, et autres
Publié: (2025)
IoT-LM: Large Multisensory Language Models for the Internet of Things
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
Unified Video-Language Pre-training with Synchronized Audio
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
MLLMs-Augmented Visual-Language Representation Learning
par: Liu, Yanqing, et autres
Publié: (2023)
par: Liu, Yanqing, et autres
Publié: (2023)
DiPrompT: Disentangled Prompt Tuning for Multiple Latent Domain Generalization in Federated Learning
par: Bai, Sikai, et autres
Publié: (2024)
par: Bai, Sikai, et autres
Publié: (2024)
Text-to-Audio Generation Synchronized with Videos
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
Learning and Leveraging World Models in Visual Representation Learning
par: Garrido, Quentin, et autres
Publié: (2024)
par: Garrido, Quentin, et autres
Publié: (2024)
PAND: Prompt-Aware Neighborhood Distillation for Lightweight Fine-Grained Visual Classification
par: Luo, Qiuming, et autres
Publié: (2026)
par: Luo, Qiuming, et autres
Publié: (2026)
DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap
par: Mo, Shentong, et autres
Publié: (2025)
par: Mo, Shentong, et autres
Publié: (2025)
Candidate Pseudolabel Learning: Enhancing Vision-Language Models by Prompt Tuning with Unlabeled Data
par: Zhang, Jiahan, et autres
Publié: (2024)
par: Zhang, Jiahan, et autres
Publié: (2024)
Pretrained Reversible Generation as Unsupervised Visual Representation Learning
par: Xue, Rongkun, et autres
Publié: (2024)
par: Xue, Rongkun, et autres
Publié: (2024)
CLEFT: Language-Image Contrastive Learning with Efficient Large Language Model and Prompt Fine-Tuning
par: Du, Yuexi, et autres
Publié: (2024)
par: Du, Yuexi, et autres
Publié: (2024)
Enhancing Long Video Generation Consistency without Tuning
par: Li, Xingyao, et autres
Publié: (2024)
par: Li, Xingyao, et autres
Publié: (2024)
Visual Tuning
par: Yu, Bruce X. B., et autres
Publié: (2023)
par: Yu, Bruce X. B., et autres
Publié: (2023)
From Pixels to Components: Eigenvector Masking for Visual Representation Learning
par: Bizeul, Alice, et autres
Publié: (2025)
par: Bizeul, Alice, et autres
Publié: (2025)
Revisiting Feature Prediction for Learning Visual Representations from Video
par: Bardes, Adrien, et autres
Publié: (2024)
par: Bardes, Adrien, et autres
Publié: (2024)
Prompt Optimization Meets Subspace Representation Learning for Few-shot Out-of-Distribution Detection
par: Sayem, Faizul Rakib, et autres
Publié: (2025)
par: Sayem, Faizul Rakib, et autres
Publié: (2025)
Parrot: Multilingual Visual Instruction Tuning
par: Sun, Hai-Long, et autres
Publié: (2024)
par: Sun, Hai-Long, et autres
Publié: (2024)
Documents similaires
-
A Large-scale Medical Visual Task Adaptation Benchmark
par: Mo, Shentong, et autres
Publié: (2024) -
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
par: Mo, Shentong, et autres
Publié: (2026) -
Improving Visual Representation Alignment Generation with GRPO
par: Mo, Shentong, et autres
Publié: (2026) -
Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
par: Mo, Shentong
Publié: (2024) -
Aligning Audio-Visual Joint Representations with an Agentic Workflow
par: Mo, Shentong, et autres
Publié: (2024)