TULIP: Token-length Upgraded CLIP
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Najdenkoska, Ivona, Derakhshani, Mohammad Mahdi, Asano, Yuki M., van Noord, Nanne, Worring, Marcel, Snoek, Cees G. M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2024)
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2024)
NeoBabel: A Multilingual Open Tower for Visual Generation
von: Derakhshani, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Derakhshani, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
Any-Shift Prompting for Generalization over Distributions
von: Xiao, Zehao, et al.
Veröffentlicht: (2024)
von: Xiao, Zehao, et al.
Veröffentlicht: (2024)
LATTE: Latent Trajectory Embedding for Diffusion-Generated Image Detection
von: Vasilcoiu, Ana, et al.
Veröffentlicht: (2025)
von: Vasilcoiu, Ana, et al.
Veröffentlicht: (2025)
GO4Align: Group Optimization for Multi-Task Alignment
von: Shen, Jiayi, et al.
Veröffentlicht: (2024)
von: Shen, Jiayi, et al.
Veröffentlicht: (2024)
Stylistic Multi-Task Analysis of Ukiyo-e Woodblock Prints
von: Khan, Selina, et al.
Veröffentlicht: (2024)
von: Khan, Selina, et al.
Veröffentlicht: (2024)
Context-Infused Visual Grounding for Art
von: Khan, Selina, et al.
Veröffentlicht: (2024)
von: Khan, Selina, et al.
Veröffentlicht: (2024)
Segment Any 3D-Part in a Scene from a Sentence
von: Wu, Hongyu, et al.
Veröffentlicht: (2025)
von: Wu, Hongyu, et al.
Veröffentlicht: (2025)
PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs
von: Dorkenwald, Michael, et al.
Veröffentlicht: (2024)
von: Dorkenwald, Michael, et al.
Veröffentlicht: (2024)
The Iconicity of the Generated Image
von: van Noord, Nanne, et al.
Veröffentlicht: (2025)
von: van Noord, Nanne, et al.
Veröffentlicht: (2025)
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
Elastic ViTs from Pretrained Models without Retraining
von: Simoncini, Walter, et al.
Veröffentlicht: (2025)
von: Simoncini, Walter, et al.
Veröffentlicht: (2025)
Redefining Normal: A Novel Object-Level Approach for Multi-Object Novelty Detection
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
Lost in Time: A New Temporal Benchmark for VideoLLMs
von: Cores, Daniel, et al.
Veröffentlicht: (2024)
von: Cores, Daniel, et al.
Veröffentlicht: (2024)
EMPLACE: Self-Supervised Urban Scene Change Detection
von: Alpherts, Tim, et al.
Veröffentlicht: (2025)
von: Alpherts, Tim, et al.
Veröffentlicht: (2025)
Artifacts of Idiosyncracy in Global Street View Data
von: Alpherts, Tim, et al.
Veröffentlicht: (2025)
von: Alpherts, Tim, et al.
Veröffentlicht: (2025)
SIGMA: Sinkhorn-Guided Masked Video Modeling
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
Continual Hyperbolic Learning of Instances and Classes
von: Ayoughi, Melika, et al.
Veröffentlicht: (2025)
von: Ayoughi, Melika, et al.
Veröffentlicht: (2025)
GeneralAD: Anomaly Detection Across Domains by Attending to Distorted Features
von: Sträter, Luc P. J., et al.
Veröffentlicht: (2024)
von: Sträter, Luc P. J., et al.
Veröffentlicht: (2024)
MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2025)
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2025)
SelEx: Self-Expertise in Fine-Grained Generalized Category Discovery
von: Rastegar, Sarah, et al.
Veröffentlicht: (2024)
von: Rastegar, Sarah, et al.
Veröffentlicht: (2024)
LocoMotion: Learning Motion-Focused Video-Language Representations
von: Doughty, Hazel, et al.
Veröffentlicht: (2024)
von: Doughty, Hazel, et al.
Veröffentlicht: (2024)
Low-Resource Vision Challenges for Foundation Models
von: Zhang, Yunhua, et al.
Veröffentlicht: (2024)
von: Zhang, Yunhua, et al.
Veröffentlicht: (2024)
In-Context Learning Improves Compositional Understanding of Vision-Language Models
von: Nulli, Matteo, et al.
Veröffentlicht: (2024)
von: Nulli, Matteo, et al.
Veröffentlicht: (2024)
Context Diffusion: In-Context Aware Image Generation
von: Najdenkoska, Ivona, et al.
Veröffentlicht: (2023)
von: Najdenkoska, Ivona, et al.
Veröffentlicht: (2023)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
von: Liu, Huabin, et al.
Veröffentlicht: (2025)
von: Liu, Huabin, et al.
Veröffentlicht: (2025)
Dual Guidance Semi-Supervised Action Detection
von: Singh, Ankit, et al.
Veröffentlicht: (2025)
von: Singh, Ankit, et al.
Veröffentlicht: (2025)
SuperDisco: Super-Class Discovery Improves Visual Recognition for the Long-Tail
von: Du, Yingjun, et al.
Veröffentlicht: (2023)
von: Du, Yingjun, et al.
Veröffentlicht: (2023)
SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation
von: Nguyen, Duy-Kien, et al.
Veröffentlicht: (2023)
von: Nguyen, Duy-Kien, et al.
Veröffentlicht: (2023)
GAP3D: Generative Alignment of VLM Latents to Patch-Level Embeddings for 3D Generation
von: Gkotsi, Polytimi Anna, et al.
Veröffentlicht: (2026)
von: Gkotsi, Polytimi Anna, et al.
Veröffentlicht: (2026)
Find the Cliffhanger: Multi-Modal Trailerness in Soap Operas
von: Bretti, Carlo, et al.
Veröffentlicht: (2024)
von: Bretti, Carlo, et al.
Veröffentlicht: (2024)
IPO: Interpretable Prompt Optimization for Vision-Language Models
von: Du, Yingjun, et al.
Veröffentlicht: (2024)
von: Du, Yingjun, et al.
Veröffentlicht: (2024)
Purrception: Variational Flow Matching for Vector-Quantized Image Generation
von: Matişan, Răzvan-Andrei, et al.
Veröffentlicht: (2025)
von: Matişan, Răzvan-Andrei, et al.
Veröffentlicht: (2025)
Crane: Context-Guided Prompt Learning and Attention Refinement for Zero-Shot Anomaly Detection
von: Salehi, Alireza, et al.
Veröffentlicht: (2025)
von: Salehi, Alireza, et al.
Veröffentlicht: (2025)
Beyond Coarse-Grained Matching in Video-Text Retrieval
von: Chen, Aozhu, et al.
Veröffentlicht: (2024)
von: Chen, Aozhu, et al.
Veröffentlicht: (2024)
MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2025)
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2025)
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
von: Sun, Wenfang, et al.
Veröffentlicht: (2026)
von: Sun, Wenfang, et al.
Veröffentlicht: (2026)
Training-Free Semantic Segmentation via LLM-Supervision
von: Sun, Wenfang, et al.
Veröffentlicht: (2024)
von: Sun, Wenfang, et al.
Veröffentlicht: (2024)
TULIP: Transformer for Upsampling of LiDAR Point Clouds
von: Yang, Bin, et al.
Veröffentlicht: (2023)
von: Yang, Bin, et al.
Veröffentlicht: (2023)
Union-over-Intersections: Object Detection beyond Winner-Takes-All
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2023)
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2024) -
NeoBabel: A Multilingual Open Tower for Visual Generation
von: Derakhshani, Mohammad Mahdi, et al.
Veröffentlicht: (2025) -
Any-Shift Prompting for Generalization over Distributions
von: Xiao, Zehao, et al.
Veröffentlicht: (2024) -
LATTE: Latent Trajectory Embedding for Diffusion-Generated Image Detection
von: Vasilcoiu, Ana, et al.
Veröffentlicht: (2025) -
GO4Align: Group Optimization for Multi-Task Alignment
von: Shen, Jiayi, et al.
Veröffentlicht: (2024)