SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Yiming, Li, Guorong, Qing, Laiyun, Beheshti, Amin, Yang, Jian, Sheng, Michael, Qi, Yuankai, Huang, Qingming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
von: Ma, Yunchuan, et al.
Veröffentlicht: (2024)
von: Ma, Yunchuan, et al.
Veröffentlicht: (2024)
Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026)
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026)
Boosting Point-supervised Temporal Action Localization via Text Refinement and Alignment
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026)
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
von: Zheng, Zelin, et al.
Veröffentlicht: (2026)
von: Zheng, Zelin, et al.
Veröffentlicht: (2026)
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
von: Tian, Mingkai, et al.
Veröffentlicht: (2025)
von: Tian, Mingkai, et al.
Veröffentlicht: (2025)
Adapter-Enhanced Semantic Prompting for Continual Learning
von: Yin, Baocai, et al.
Veröffentlicht: (2024)
von: Yin, Baocai, et al.
Veröffentlicht: (2024)
SOVC: Subject-Oriented Video Captioning
von: Teng, Chang, et al.
Veröffentlicht: (2023)
von: Teng, Chang, et al.
Veröffentlicht: (2023)
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2025)
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2025)
ProgRoCC: A Progressive Approach to Rough Crowd Counting
von: Jiang, Shengqin, et al.
Veröffentlicht: (2025)
von: Jiang, Shengqin, et al.
Veröffentlicht: (2025)
Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning
von: Jiang, Huajie, et al.
Veröffentlicht: (2025)
von: Jiang, Huajie, et al.
Veröffentlicht: (2025)
Multimodal Visual Surrogate Compression for Alzheimer's Disease Classification
von: Ding, Dexuan, et al.
Veröffentlicht: (2026)
von: Ding, Dexuan, et al.
Veröffentlicht: (2026)
Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured Videos
von: Ma, Jianbo, et al.
Veröffentlicht: (2025)
von: Ma, Jianbo, et al.
Veröffentlicht: (2025)
OpenworldAUC: Towards Unified Evaluation and Optimization for Open-world Prompt Tuning
von: Hua, Cong, et al.
Veröffentlicht: (2025)
von: Hua, Cong, et al.
Veröffentlicht: (2025)
StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2024)
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2024)
CountGD++: Generalized Prompting for Open-World Counting
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2025)
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2025)
Natural Language-Oriented Programming (NLOP): Towards Democratizing Software Creation
von: Beheshti, Amin
Veröffentlicht: (2024)
von: Beheshti, Amin
Veröffentlicht: (2024)
Open-World Object Counting in Videos
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2025)
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2025)
Teaching Prompts to Coordinate: Hierarchical Layer-Grouped Prompt Tuning for Continual Learning
von: Jiang, Shengqin, et al.
Veröffentlicht: (2025)
von: Jiang, Shengqin, et al.
Veröffentlicht: (2025)
Embedded Visual Prompt Tuning
von: Zu, Wenqiang, et al.
Veröffentlicht: (2024)
von: Zu, Wenqiang, et al.
Veröffentlicht: (2024)
Facing the Elephant in the Room: Visual Prompt Tuning or Full Finetuning?
von: Han, Cheng, et al.
Veröffentlicht: (2024)
von: Han, Cheng, et al.
Veröffentlicht: (2024)
OpenVidVRD: Open-Vocabulary Video Visual Relation Detection via Prompt-Driven Semantic Space Alignment
von: Liu, Qi, et al.
Veröffentlicht: (2025)
von: Liu, Qi, et al.
Veröffentlicht: (2025)
FedDPG: An Adaptive Yet Efficient Prompt-tuning Approach in Federated Learning Settings
von: Shakeri, Ali, et al.
Veröffentlicht: (2025)
von: Shakeri, Ali, et al.
Veröffentlicht: (2025)
OpenGround: Active Cognition-based Reasoning for Open-World 3D Visual Grounding
von: Huang, Wenyuan, et al.
Veröffentlicht: (2025)
von: Huang, Wenyuan, et al.
Veröffentlicht: (2025)
Hierarchical Text-Guided Brain Tumor Segmentation via Sub-Region-Aware Prompts
von: Mohammadi, Bahram, et al.
Veröffentlicht: (2026)
von: Mohammadi, Bahram, et al.
Veröffentlicht: (2026)
Open World Object Detection: A Survey
von: Li, Yiming, et al.
Veröffentlicht: (2024)
von: Li, Yiming, et al.
Veröffentlicht: (2024)
Open-World Human-Object Interaction Detection via Multi-modal Prompts
von: Yang, Jie, et al.
Veröffentlicht: (2024)
von: Yang, Jie, et al.
Veröffentlicht: (2024)
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2024)
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2024)
Expanding Zero-Shot Object Counting with Rich Prompts
von: Zhu, Huilin, et al.
Veröffentlicht: (2025)
von: Zhu, Huilin, et al.
Veröffentlicht: (2025)
Visual Fourier Prompt Tuning
von: Zeng, Runjia, et al.
Veröffentlicht: (2024)
von: Zeng, Runjia, et al.
Veröffentlicht: (2024)
CountGD: Multi-Modal Open-World Counting
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2024)
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2024)
DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers
von: Ren, Li, et al.
Veröffentlicht: (2025)
von: Ren, Li, et al.
Veröffentlicht: (2025)
CVPT: Cross Visual Prompt Tuning
von: Huang, Lingyun, et al.
Veröffentlicht: (2024)
von: Huang, Lingyun, et al.
Veröffentlicht: (2024)
Semantic-aware SAM for Point-Prompted Instance Segmentation
von: Wei, Zhaoyang, et al.
Veröffentlicht: (2023)
von: Wei, Zhaoyang, et al.
Veröffentlicht: (2023)
Open-World Dynamic Prompt and Continual Visual Representation Learning
von: Kim, Youngeun, et al.
Veröffentlicht: (2024)
von: Kim, Youngeun, et al.
Veröffentlicht: (2024)
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
von: Yang, Yong, et al.
Veröffentlicht: (2024)
von: Yang, Yong, et al.
Veröffentlicht: (2024)
Open-Vocabulary Object Detection with Meta Prompt Representation and Instance Contrastive Optimization
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
SAPNet++: Evolving Point-Prompted Instance Segmentation with Semantic and Spatial Awareness
von: Wei, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Wei, Zhaoyang, et al.
Veröffentlicht: (2026)
CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2026)
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2026)
Ratas framework: A comprehensive genai-based approach to rubric-based marking of real-world textual exams
von: Safilian, Masoud, et al.
Veröffentlicht: (2025)
von: Safilian, Masoud, et al.
Veröffentlicht: (2025)
PersoPilot: An Adaptive AI-Copilot for Transparent Contextualized Persona Classification and Personalized Response Generation
von: Afzoon, Saleh, et al.
Veröffentlicht: (2026)
von: Afzoon, Saleh, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
von: Ma, Yunchuan, et al.
Veröffentlicht: (2024) -
Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026) -
Boosting Point-supervised Temporal Action Localization via Text Refinement and Alignment
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026) -
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
von: Zheng, Zelin, et al.
Veröffentlicht: (2026) -
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
von: Tian, Mingkai, et al.
Veröffentlicht: (2025)