CLAMP: Contrastive LAnguage Model Prompt-tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Teterwak, Piotr, Sun, Ximeng, Plummer, Bryan A., Saenko, Kate, Lim, Ser-Nam |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OP-LoRA: The Blessing of Dimensionality
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
Web Artifact Attacks Disrupt Vision Language Models
von: Qraitem, Maan, et al.
Veröffentlicht: (2025)
von: Qraitem, Maan, et al.
Veröffentlicht: (2025)
SLANT: Spurious Logo ANalysis Toolkit
von: Qraitem, Maan, et al.
Veröffentlicht: (2024)
von: Qraitem, Maan, et al.
Veröffentlicht: (2024)
ERM++: An Improved Baseline for Domain Generalization
von: Teterwak, Piotr, et al.
Veröffentlicht: (2023)
von: Teterwak, Piotr, et al.
Veröffentlicht: (2023)
Is Large-Scale Pretraining the Secret to Good Domain Generalization?
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
von: Qraitem, Maan, et al.
Veröffentlicht: (2024)
von: Qraitem, Maan, et al.
Veröffentlicht: (2024)
Tell Me What's Next: Textual Foresight for Generic UI Representations
von: Burns, Andrea, et al.
Veröffentlicht: (2024)
von: Burns, Andrea, et al.
Veröffentlicht: (2024)
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
von: Liu, Aoming, et al.
Veröffentlicht: (2025)
von: Liu, Aoming, et al.
Veröffentlicht: (2025)
From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition
von: Qraitem, Maan, et al.
Veröffentlicht: (2023)
von: Qraitem, Maan, et al.
Veröffentlicht: (2023)
Koala: Key frame-conditioned long video-LLM
von: Tan, Reuben, et al.
Veröffentlicht: (2024)
von: Tan, Reuben, et al.
Veröffentlicht: (2024)
Towards Chunk-Wise Generation for Long Videos
von: Zhang, Siyang, et al.
Veröffentlicht: (2024)
von: Zhang, Siyang, et al.
Veröffentlicht: (2024)
SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
von: Miller, Kevin, et al.
Veröffentlicht: (2025)
von: Miller, Kevin, et al.
Veröffentlicht: (2025)
VideoMerge: Towards Training-free Long Video Generation
von: Zhang, Siyang, et al.
Veröffentlicht: (2025)
von: Zhang, Siyang, et al.
Veröffentlicht: (2025)
FuTCR: Future-Targeted Contrast and Repulsion for Continual Panoptic Segmentation
von: Ikechukwu, Nicholas, et al.
Veröffentlicht: (2026)
von: Ikechukwu, Nicholas, et al.
Veröffentlicht: (2026)
DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation
von: Meyarian, Abolfazl, et al.
Veröffentlicht: (2026)
von: Meyarian, Abolfazl, et al.
Veröffentlicht: (2026)
Concept Arithmetics for Circumventing Concept Inhibition in Diffusion Models
von: Petsiuk, Vitali, et al.
Veröffentlicht: (2024)
von: Petsiuk, Vitali, et al.
Veröffentlicht: (2024)
Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
Mull-Tokens: Modality-Agnostic Latent Thinking
von: Ray, Arijit, et al.
Veröffentlicht: (2025)
von: Ray, Arijit, et al.
Veröffentlicht: (2025)
LNL+K: Enhancing Learning with Noisy Labels Through Noise Source Knowledge Integration
von: Wang, Siqi, et al.
Veröffentlicht: (2023)
von: Wang, Siqi, et al.
Veröffentlicht: (2023)
DreamMask: Boosting Open-vocabulary Panoptic Segmentation with Synthetic Data
von: Tu, Yuanpeng, et al.
Veröffentlicht: (2025)
von: Tu, Yuanpeng, et al.
Veröffentlicht: (2025)
Towards Unified 3D Object Detection via Algorithm and Data Unification
von: Li, Zhuoling, et al.
Veröffentlicht: (2024)
von: Li, Zhuoling, et al.
Veröffentlicht: (2024)
Composing Object Relations and Attributes for Image-Text Matching
von: Pham, Khoi, et al.
Veröffentlicht: (2024)
von: Pham, Khoi, et al.
Veröffentlicht: (2024)
Fast Encoding and Decoding for Implicit Video Representation
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
von: Qian, Zhaofang, et al.
Veröffentlicht: (2024)
von: Qian, Zhaofang, et al.
Veröffentlicht: (2024)
LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
von: Li, Zhuoling, et al.
Veröffentlicht: (2024)
von: Li, Zhuoling, et al.
Veröffentlicht: (2024)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
von: Shrivastava, Gaurav, et al.
Veröffentlicht: (2024)
von: Shrivastava, Gaurav, et al.
Veröffentlicht: (2024)
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
von: Reilly, Dominick, et al.
Veröffentlicht: (2024)
von: Reilly, Dominick, et al.
Veröffentlicht: (2024)
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
von: Mishra, Samarth, et al.
Veröffentlicht: (2025)
von: Mishra, Samarth, et al.
Veröffentlicht: (2025)
CLAMP: Contrastive Learning with Adaptive Multi-loss and Progressive Fusion for Multimodal Aspect-Based Sentiment Analysis
von: He, Xiaoqiang
Veröffentlicht: (2025)
von: He, Xiaoqiang
Veröffentlicht: (2025)
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2024)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2024)
FSViewFusion: Few-Shots View Generation of Novel Objects
von: Hussain, Rukhshanda, et al.
Veröffentlicht: (2024)
von: Hussain, Rukhshanda, et al.
Veröffentlicht: (2024)
What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?
von: Cui, Xuanming, et al.
Veröffentlicht: (2025)
von: Cui, Xuanming, et al.
Veröffentlicht: (2025)
BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration
von: Gao, Bo, et al.
Veröffentlicht: (2026)
von: Gao, Bo, et al.
Veröffentlicht: (2026)
Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compression
von: Tasnim, Nazia, et al.
Veröffentlicht: (2026)
von: Tasnim, Nazia, et al.
Veröffentlicht: (2026)
RECAST: Reparameterized, Compact weight Adaptation for Sequential Tasks
von: Tasnim, Nazia, et al.
Veröffentlicht: (2024)
von: Tasnim, Nazia, et al.
Veröffentlicht: (2024)
Enhancing Feature Diversity Boosts Channel-Adaptive Vision Transformers
von: Pham, Chau, et al.
Veröffentlicht: (2024)
von: Pham, Chau, et al.
Veröffentlicht: (2024)
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
Enhancing Virtual Try-On with Synthetic Pairs and Error-Aware Noise Scheduling
von: Li, Nannan, et al.
Veröffentlicht: (2025)
von: Li, Nannan, et al.
Veröffentlicht: (2025)
Federated Adversarial Domain Adaptation
von: Peng, Xingchao, et al.
Veröffentlicht: (2019)
von: Peng, Xingchao, et al.
Veröffentlicht: (2019)
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
OP-LoRA: The Blessing of Dimensionality
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024) -
Web Artifact Attacks Disrupt Vision Language Models
von: Qraitem, Maan, et al.
Veröffentlicht: (2025) -
SLANT: Spurious Logo ANalysis Toolkit
von: Qraitem, Maan, et al.
Veröffentlicht: (2024) -
ERM++: An Improved Baseline for Domain Generalization
von: Teterwak, Piotr, et al.
Veröffentlicht: (2023) -
Is Large-Scale Pretraining the Secret to Good Domain Generalization?
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)