Open-World Dynamic Prompt and Continual Visual Representation Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Youngeun, Fang, Jun, Zhang, Qin, Cai, Zhaowei, Shen, Yantao, Duggal, Rahul, Raychaudhuri, Dripta S., Tu, Zhuowen, Xing, Yifan, Dabeer, Onkar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
por: Zong, Yongshuo, et al.
Publicado: (2025)
por: Zong, Yongshuo, et al.
Publicado: (2025)
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
por: Zhang, Zhaoyang, et al.
Publicado: (2023)
por: Zhang, Zhaoyang, et al.
Publicado: (2023)
CONTRAST: Continual Multi-source Adaptation to Dynamic Distributions
por: Ahmed, Sk Miraj, et al.
Publicado: (2024)
por: Ahmed, Sk Miraj, et al.
Publicado: (2024)
One-stage Prompt-based Continual Learning
por: Kim, Youngeun, et al.
Publicado: (2024)
por: Kim, Youngeun, et al.
Publicado: (2024)
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning
por: Ghosh, Udita, et al.
Publicado: (2025)
por: Ghosh, Udita, et al.
Publicado: (2025)
Reducing Oracle Feedback with Vision-Language Embeddings for Preference-Based RL
por: Ghosh, Udita, et al.
Publicado: (2026)
por: Ghosh, Udita, et al.
Publicado: (2026)
Robust Offline Imitation Learning from Diverse Auxiliary Data
por: Ghosh, Udita, et al.
Publicado: (2024)
por: Ghosh, Udita, et al.
Publicado: (2024)
Salient Concept-Aware Generative Data Augmentation
por: Zhao, Tianchen, et al.
Publicado: (2025)
por: Zhao, Tianchen, et al.
Publicado: (2025)
Threshold-Consistent Margin Loss for Open-World Deep Metric Learning
por: Zhang, Qin, et al.
Publicado: (2023)
por: Zhang, Qin, et al.
Publicado: (2023)
Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
por: Tan, Jing, et al.
Publicado: (2026)
por: Tan, Jing, et al.
Publicado: (2026)
AuthGuard: Generalizable Deepfake Detection via Language Guidance
por: Shen, Guangyu, et al.
Publicado: (2025)
por: Shen, Guangyu, et al.
Publicado: (2025)
Do We Really Need a Large Number of Visual Prompts?
por: Kim, Youngeun, et al.
Publicado: (2023)
por: Kim, Youngeun, et al.
Publicado: (2023)
Soft Tail-dropping for Adaptive Visual Tokenization
por: Chen, Zeyuan, et al.
Publicado: (2026)
por: Chen, Zeyuan, et al.
Publicado: (2026)
MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting
por: Liu, Fangchen, et al.
Publicado: (2024)
por: Liu, Fangchen, et al.
Publicado: (2024)
ELODI: Ensemble Logit Difference Inhibition for Positive-Congruent Training
por: Zhao, Yue, et al.
Publicado: (2022)
por: Zhao, Yue, et al.
Publicado: (2022)
Spherical Covariance Representations
por: Bobkov, Sergey G., et al.
Publicado: (2024)
por: Bobkov, Sergey G., et al.
Publicado: (2024)
Jan Company in Coromandel 1605-1690
por: Raychaudhuri, T.
Publicado: (2016)
por: Raychaudhuri, T.
Publicado: (2016)
ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting
por: Cai, Shaofei, et al.
Publicado: (2024)
por: Cai, Shaofei, et al.
Publicado: (2024)
Höffding's Kernels and Periodic Covariance Representations
por: Bobkov, Sergey G., et al.
Publicado: (2024)
por: Bobkov, Sergey G., et al.
Publicado: (2024)
ToxSearch: Evolving Prompts for Toxicity Search in Large Language Models
por: Shelar, Onkar, et al.
Publicado: (2025)
por: Shelar, Onkar, et al.
Publicado: (2025)
MC-GRPO: Median-Centered Group Relative Policy Optimization for Small-Rollout Reinforcement Learning
por: Kim, Youngeun
Publicado: (2026)
por: Kim, Youngeun
Publicado: (2026)
STRIDE: Single-video based Temporally Continuous Occlusion-Robust 3D Pose Estimation
por: Lal, Rohit, et al.
Publicado: (2023)
por: Lal, Rohit, et al.
Publicado: (2023)
Reinforcement-aware Knowledge Distillation for LLM Reasoning
por: Zhang, Zhaoyang, et al.
Publicado: (2026)
por: Zhang, Zhaoyang, et al.
Publicado: (2026)
VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning Models
por: Ghosal, Soumya Suvra, et al.
Publicado: (2026)
por: Ghosal, Soumya Suvra, et al.
Publicado: (2026)
Multi-modal Pose Diffuser: A Multimodal Generative Conditional Pose Prior
por: Ta, Calvin-Khang, et al.
Publicado: (2024)
por: Ta, Calvin-Khang, et al.
Publicado: (2024)
VADIS: A Visual Analytics Pipeline for Dynamic Document Representation and Information-Seeking
por: Qiu, Rui, et al.
Publicado: (2025)
por: Qiu, Rui, et al.
Publicado: (2025)
Visually Similar Pair Alignment for Robust Cross-Domain Object Detection
por: Krishna, Onkar, et al.
Publicado: (2025)
por: Krishna, Onkar, et al.
Publicado: (2025)
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
por: Kim, Sungnyun, et al.
Publicado: (2024)
por: Kim, Sungnyun, et al.
Publicado: (2024)
POSTURE: Pose Guided Unsupervised Domain Adaptation for Human Body Part Segmentation
por: Dutta, Arindam, et al.
Publicado: (2024)
por: Dutta, Arindam, et al.
Publicado: (2024)
RAE-NWM: Navigation World Model in Dense Visual Representation Space
por: Zhang, Mingkun, et al.
Publicado: (2026)
por: Zhang, Mingkun, et al.
Publicado: (2026)
VisTR: Visualizations as Representations for Time-series Table Reasoning
por: Hao, Jianing, et al.
Publicado: (2024)
por: Hao, Jianing, et al.
Publicado: (2024)
CD-NGP: A Fast Scalable Continual Representation for Dynamic Scenes
por: Liu, Zhenhuan, et al.
Publicado: (2024)
por: Liu, Zhenhuan, et al.
Publicado: (2024)
Learning for Transductive Threshold Calibration in Open-World Recognition
por: Zhang, Qin, et al.
Publicado: (2023)
por: Zhang, Qin, et al.
Publicado: (2023)
Open-Vocabulary Action Localization with Iterative Visual Prompting
por: Wake, Naoki, et al.
Publicado: (2024)
por: Wake, Naoki, et al.
Publicado: (2024)
OpenVO: Open-World Visual Odometry with Temporal Dynamics Awareness
por: Nguyen, Phuc D. A., et al.
Publicado: (2026)
por: Nguyen, Phuc D. A., et al.
Publicado: (2026)
Visual Prompt Tuning in Null Space for Continual Learning
por: Lu, Yue, et al.
Publicado: (2024)
por: Lu, Yue, et al.
Publicado: (2024)
SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting
por: Zhao, Yiming, et al.
Publicado: (2025)
por: Zhao, Yiming, et al.
Publicado: (2025)
Unified Open-World Segmentation with Multi-Modal Prompts
por: Liu, Yang, et al.
Publicado: (2025)
por: Liu, Yang, et al.
Publicado: (2025)
Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins
por: Ning, Chuanruo, et al.
Publicado: (2025)
por: Ning, Chuanruo, et al.
Publicado: (2025)
Boosting Few-Shot Open-Set Object Detection via Prompt Learning and Robust Decision Boundary
por: Wu, Zhaowei, et al.
Publicado: (2024)
por: Wu, Zhaowei, et al.
Publicado: (2024)
Ejemplares similares
-
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
por: Zong, Yongshuo, et al.
Publicado: (2025) -
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
por: Zhang, Zhaoyang, et al.
Publicado: (2023) -
CONTRAST: Continual Multi-source Adaptation to Dynamic Distributions
por: Ahmed, Sk Miraj, et al.
Publicado: (2024) -
One-stage Prompt-based Continual Learning
por: Kim, Youngeun, et al.
Publicado: (2024) -
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning
por: Ghosh, Udita, et al.
Publicado: (2025)