From Latent to Engine Manifolds: Analyzing ImageBind's Multimodal Embedding Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hamara, Andrew, Rivas, Pablo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring Visual Embedding Spaces Induced by Vision Transformers for Online Auto Parts Marketplaces
von: Armijo, Cameron, et al.
Veröffentlicht: (2025)
von: Armijo, Cameron, et al.
Veröffentlicht: (2025)
Image-Based Leopard Seal Recognition: Approaches and Challenges in Current Automated Systems
von: Salazar, Jorge Yero, et al.
Veröffentlicht: (2024)
von: Salazar, Jorge Yero, et al.
Veröffentlicht: (2024)
Semi-Supervised Segmentation via Embedding Matching
von: Xie, Weiyi, et al.
Veröffentlicht: (2024)
von: Xie, Weiyi, et al.
Veröffentlicht: (2024)
OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection
von: Schneider, David, et al.
Veröffentlicht: (2025)
von: Schneider, David, et al.
Veröffentlicht: (2025)
FAME: Feature Activation Map Explanation on Image Classification and Face Recognition
von: Zhang, Xinyi, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyi, et al.
Veröffentlicht: (2026)
RPCASSM: Robust PCA State Space Model For Infrared Small Target Detection
von: Liu, Pingping, et al.
Veröffentlicht: (2026)
von: Liu, Pingping, et al.
Veröffentlicht: (2026)
Conjuring Positive Pairs for Efficient Unification of Representation Learning and Image Synthesis
von: Estepa, Imanol G., et al.
Veröffentlicht: (2025)
von: Estepa, Imanol G., et al.
Veröffentlicht: (2025)
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
von: Yasuno, Takato
Veröffentlicht: (2026)
von: Yasuno, Takato
Veröffentlicht: (2026)
UTAL-GNN: Unsupervised Temporal Action Localization using Graph Neural Networks
von: Badatya, Bikash Kumar, et al.
Veröffentlicht: (2025)
von: Badatya, Bikash Kumar, et al.
Veröffentlicht: (2025)
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
von: Deng, Pei, et al.
Veröffentlicht: (2025)
von: Deng, Pei, et al.
Veröffentlicht: (2025)
Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
von: Perera, Amal S., et al.
Veröffentlicht: (2025)
von: Perera, Amal S., et al.
Veröffentlicht: (2025)
Canonical Space Representation for 4D Panoptic Segmentation of Articulated Objects
von: Gomes, Manuel, et al.
Veröffentlicht: (2025)
von: Gomes, Manuel, et al.
Veröffentlicht: (2025)
Learning Association via Track-Detection Matching for Multi-Object Tracking
von: Adžemović, Momir
Veröffentlicht: (2025)
von: Adžemović, Momir
Veröffentlicht: (2025)
HY-Himmel Technical Report: Hierarchical Interleaved Multi-stream Motion Encoding for Long Video Understanding
von: Jin, Haopeng, et al.
Veröffentlicht: (2026)
von: Jin, Haopeng, et al.
Veröffentlicht: (2026)
Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation
von: Estepa, Imanol G., et al.
Veröffentlicht: (2026)
von: Estepa, Imanol G., et al.
Veröffentlicht: (2026)
All4One: Symbiotic Neighbour Contrastive Learning via Self-Attention and Redundancy Reduction
von: Estepa, Imanol G., et al.
Veröffentlicht: (2023)
von: Estepa, Imanol G., et al.
Veröffentlicht: (2023)
SelvaMask: Segmenting Trees in Tropical Forests and Beyond
von: Duguay, Simon-Olivier, et al.
Veröffentlicht: (2026)
von: Duguay, Simon-Olivier, et al.
Veröffentlicht: (2026)
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
von: Wang, Yiming, et al.
Veröffentlicht: (2026)
von: Wang, Yiming, et al.
Veröffentlicht: (2026)
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
von: Li, Yayuan, et al.
Veröffentlicht: (2025)
von: Li, Yayuan, et al.
Veröffentlicht: (2025)
HSDA: High-frequency Shuffle Data Augmentation for Bird's-Eye-View Map Segmentation
von: Glisson, Calvin, et al.
Veröffentlicht: (2024)
von: Glisson, Calvin, et al.
Veröffentlicht: (2024)
SelvaBox: A high-resolution dataset for tropical tree crown detection
von: Baudchon, Hugo, et al.
Veröffentlicht: (2025)
von: Baudchon, Hugo, et al.
Veröffentlicht: (2025)
Reference Dataset and Benchmark for Reconstructing Laser Parameters from On-axis Video in Powder Bed Fusion of Bulk Stainless Steel
von: Blanc, Cyril, et al.
Veröffentlicht: (2024)
von: Blanc, Cyril, et al.
Veröffentlicht: (2024)
Dense Motion Captioning
von: Xu, Shiyao, et al.
Veröffentlicht: (2025)
von: Xu, Shiyao, et al.
Veröffentlicht: (2025)
A Novel Dataset for Flood Detection Robust to Seasonal Changes in Satellite Imagery
von: Jang, Youngsun, et al.
Veröffentlicht: (2025)
von: Jang, Youngsun, et al.
Veröffentlicht: (2025)
TD3Net: A temporal densely connected multi-dilated convolutional network for lipreading
von: Lee, Byung Hoon, et al.
Veröffentlicht: (2025)
von: Lee, Byung Hoon, et al.
Veröffentlicht: (2025)
CoMatcher: Multi-View Collaborative Feature Matching
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
MoDE: Mixture of Diffusion Experts for Any Occluded Face Recognition
von: Fan, Qiannan, et al.
Veröffentlicht: (2025)
von: Fan, Qiannan, et al.
Veröffentlicht: (2025)
NeuroGaze-Distill: Brain-informed Distillation and Depression-Inspired Geometric Priors for Robust Facial Emotion Recognition
von: Li, Zilin, et al.
Veröffentlicht: (2025)
von: Li, Zilin, et al.
Veröffentlicht: (2025)
Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection
von: Deka, Dawar Jyoti, et al.
Veröffentlicht: (2026)
von: Deka, Dawar Jyoti, et al.
Veröffentlicht: (2026)
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning
von: Semenov, Andrei, et al.
Veröffentlicht: (2024)
von: Semenov, Andrei, et al.
Veröffentlicht: (2024)
CG-HOI: Contact-Guided 3D Human-Object Interaction Generation
von: Diller, Christian, et al.
Veröffentlicht: (2023)
von: Diller, Christian, et al.
Veröffentlicht: (2023)
Multimodal Ensemble with Conditional Feature Fusion for Dysgraphia Diagnosis in Children from Handwriting Samples
von: Kunhoth, Jayakanth, et al.
Veröffentlicht: (2024)
von: Kunhoth, Jayakanth, et al.
Veröffentlicht: (2024)
Towards Accurate and Efficient Waste Image Classification: A Hybrid Deep Learning and Machine Learning Approach
von: Nguyen, Ngoc-Bao-Quang, et al.
Veröffentlicht: (2025)
von: Nguyen, Ngoc-Bao-Quang, et al.
Veröffentlicht: (2025)
Efficient Temporally-Aware DeepFake Detection using H.264 Motion Vectors
von: Grönquist, Peter, et al.
Veröffentlicht: (2023)
von: Grönquist, Peter, et al.
Veröffentlicht: (2023)
Product Review Based on Optimized Facial Expression Detection
von: Chaugule, Vikrant, et al.
Veröffentlicht: (2026)
von: Chaugule, Vikrant, et al.
Veröffentlicht: (2026)
EgoTraj: Real-World Egocentric Human Trajectory Dataset for Multimodal Prediction
von: Yehia, Ahmad, et al.
Veröffentlicht: (2026)
von: Yehia, Ahmad, et al.
Veröffentlicht: (2026)
The Influence of Iconicity in Transfer Learning for Sign Language Recognition
von: Artiaga, Keren, et al.
Veröffentlicht: (2026)
von: Artiaga, Keren, et al.
Veröffentlicht: (2026)
Detecting AI-Generated Videos with Spiking Neural Networks
von: Jang, Minsuk, et al.
Veröffentlicht: (2026)
von: Jang, Minsuk, et al.
Veröffentlicht: (2026)
Precision at Scale: Domain-Specific Datasets On-Demand
von: Rodríguez-de-Vera, Jesús M, et al.
Veröffentlicht: (2024)
von: Rodríguez-de-Vera, Jesús M, et al.
Veröffentlicht: (2024)
LiftAvatar: Kinematic-Space Completion for Expression-Controlled 3D Gaussian Avatar Animation
von: Wei, Hualiang, et al.
Veröffentlicht: (2026)
von: Wei, Hualiang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Exploring Visual Embedding Spaces Induced by Vision Transformers for Online Auto Parts Marketplaces
von: Armijo, Cameron, et al.
Veröffentlicht: (2025) -
Image-Based Leopard Seal Recognition: Approaches and Challenges in Current Automated Systems
von: Salazar, Jorge Yero, et al.
Veröffentlicht: (2024) -
Semi-Supervised Segmentation via Embedding Matching
von: Xie, Weiyi, et al.
Veröffentlicht: (2024) -
OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection
von: Schneider, David, et al.
Veröffentlicht: (2025) -
FAME: Feature Activation Map Explanation on Image Classification and Face Recognition
von: Zhang, Xinyi, et al.
Veröffentlicht: (2026)