Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Estepa, Imanol G., Rodríguez-de-Vera, Jesús M, Nagarajan, Bhalaji, Radeva, Petia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
All4One: Symbiotic Neighbour Contrastive Learning via Self-Attention and Redundancy Reduction
por: Estepa, Imanol G., et al.
Publicado: (2023)
por: Estepa, Imanol G., et al.
Publicado: (2023)
Conjuring Positive Pairs for Efficient Unification of Representation Learning and Image Synthesis
por: Estepa, Imanol G., et al.
Publicado: (2025)
por: Estepa, Imanol G., et al.
Publicado: (2025)
Precision at Scale: Domain-Specific Datasets On-Demand
por: Rodríguez-de-Vera, Jesús M, et al.
Publicado: (2024)
por: Rodríguez-de-Vera, Jesús M, et al.
Publicado: (2024)
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning
por: Semenov, Andrei, et al.
Publicado: (2024)
por: Semenov, Andrei, et al.
Publicado: (2024)
Learning Association via Track-Detection Matching for Multi-Object Tracking
por: Adžemović, Momir
Publicado: (2025)
por: Adžemović, Momir
Publicado: (2025)
Canonical Space Representation for 4D Panoptic Segmentation of Articulated Objects
por: Gomes, Manuel, et al.
Publicado: (2025)
por: Gomes, Manuel, et al.
Publicado: (2025)
FAME: Feature Activation Map Explanation on Image Classification and Face Recognition
por: Zhang, Xinyi, et al.
Publicado: (2026)
por: Zhang, Xinyi, et al.
Publicado: (2026)
CG-HOI: Contact-Guided 3D Human-Object Interaction Generation
por: Diller, Christian, et al.
Publicado: (2023)
por: Diller, Christian, et al.
Publicado: (2023)
Capacity Constraint Analysis Using Object Detection for Smart Manufacturing
por: Ahmad, Hafiz Mughees, et al.
Publicado: (2024)
por: Ahmad, Hafiz Mughees, et al.
Publicado: (2024)
SH17: A Dataset for Human Safety and Personal Protective Equipment Detection in Manufacturing Industry
por: Ahmad, Hafiz Mughees, et al.
Publicado: (2024)
por: Ahmad, Hafiz Mughees, et al.
Publicado: (2024)
FutureHuman3D: Forecasting Complex Long-Term 3D Human Behavior from Video Observations
por: Diller, Christian, et al.
Publicado: (2022)
por: Diller, Christian, et al.
Publicado: (2022)
Detecting AI-Generated Videos with Spiking Neural Networks
por: Jang, Minsuk, et al.
Publicado: (2026)
por: Jang, Minsuk, et al.
Publicado: (2026)
SpectralCA: Bi-Directional Cross-Attention for Next-Generation UAV Hyperspectral Vision
por: Brovko, D. V.
Publicado: (2025)
por: Brovko, D. V.
Publicado: (2025)
CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
por: Schiffer, Christian, et al.
Publicado: (2025)
por: Schiffer, Christian, et al.
Publicado: (2025)
Multimodal Ensemble with Conditional Feature Fusion for Dysgraphia Diagnosis in Children from Handwriting Samples
por: Kunhoth, Jayakanth, et al.
Publicado: (2024)
por: Kunhoth, Jayakanth, et al.
Publicado: (2024)
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
por: Wu, Jason, et al.
Publicado: (2026)
por: Wu, Jason, et al.
Publicado: (2026)
ZO-DARTS++: An Efficient and Size-Variable Zeroth-Order Neural Architecture Search Algorithm
por: Xie, Lunchen, et al.
Publicado: (2025)
por: Xie, Lunchen, et al.
Publicado: (2025)
Salient Concept-Aware Generative Data Augmentation
por: Zhao, Tianchen, et al.
Publicado: (2025)
por: Zhao, Tianchen, et al.
Publicado: (2025)
LiftAvatar: Kinematic-Space Completion for Expression-Controlled 3D Gaussian Avatar Animation
por: Wei, Hualiang, et al.
Publicado: (2026)
por: Wei, Hualiang, et al.
Publicado: (2026)
Classification of Cattle Behavior and Detection of Heat (Estrus) using Sensor Data
por: Dhakshinamoorthy, Druva, et al.
Publicado: (2025)
por: Dhakshinamoorthy, Druva, et al.
Publicado: (2025)
CASE: Contrastive Activation for Saliency Estimation
por: Williamson, Dane, et al.
Publicado: (2025)
por: Williamson, Dane, et al.
Publicado: (2025)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
por: Ji, Binbin, et al.
Publicado: (2025)
por: Ji, Binbin, et al.
Publicado: (2025)
OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection
por: Schneider, David, et al.
Publicado: (2025)
por: Schneider, David, et al.
Publicado: (2025)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
por: Gupta, Sunny, et al.
Publicado: (2024)
por: Gupta, Sunny, et al.
Publicado: (2024)
Skullptor: High Fidelity 3D Head Reconstruction in Seconds with Multi-View Normal Prediction
por: Artru, Noé, et al.
Publicado: (2026)
por: Artru, Noé, et al.
Publicado: (2026)
WACA-UNet: Weakness-Aware Channel Attention for Static IR Drop Prediction in Integrated Circuit Design
por: Seo, Youngmin, et al.
Publicado: (2025)
por: Seo, Youngmin, et al.
Publicado: (2025)
Exploring Visual Embedding Spaces Induced by Vision Transformers for Online Auto Parts Marketplaces
por: Armijo, Cameron, et al.
Publicado: (2025)
por: Armijo, Cameron, et al.
Publicado: (2025)
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
por: Dong, Zixuan, et al.
Publicado: (2025)
por: Dong, Zixuan, et al.
Publicado: (2025)
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
por: Asanuma, Haruka, et al.
Publicado: (2025)
por: Asanuma, Haruka, et al.
Publicado: (2025)
Task Singular Vectors: Reducing Task Interference in Model Merging
por: Gargiulo, Antonio Andrea, et al.
Publicado: (2024)
por: Gargiulo, Antonio Andrea, et al.
Publicado: (2024)
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
por: Wang, Yiming, et al.
Publicado: (2026)
por: Wang, Yiming, et al.
Publicado: (2026)
NV3D: Leveraging Spatial Shape Through Normal Vector-based 3D Object Detection
por: Chaowakarn, Krittin, et al.
Publicado: (2025)
por: Chaowakarn, Krittin, et al.
Publicado: (2025)
OrganicHAR: Towards Activity Discovery in Organic Settings for Privacy Preserving Sensors Using Efficient Video Analysis
por: Patidar, Prasoon, et al.
Publicado: (2026)
por: Patidar, Prasoon, et al.
Publicado: (2026)
Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
por: Perera, Amal S., et al.
Publicado: (2025)
por: Perera, Amal S., et al.
Publicado: (2025)
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
por: Yasuno, Takato
Publicado: (2026)
por: Yasuno, Takato
Publicado: (2026)
TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR
por: Lentsch, Ted, et al.
Publicado: (2026)
por: Lentsch, Ted, et al.
Publicado: (2026)
UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-Classes
por: Lentsch, Ted, et al.
Publicado: (2024)
por: Lentsch, Ted, et al.
Publicado: (2024)
UTAL-GNN: Unsupervised Temporal Action Localization using Graph Neural Networks
por: Badatya, Bikash Kumar, et al.
Publicado: (2025)
por: Badatya, Bikash Kumar, et al.
Publicado: (2025)
Few-Shot Learning of a Graph-Based Neural Network Model Without Backpropagation
por: Lapin, Mykyta, et al.
Publicado: (2025)
por: Lapin, Mykyta, et al.
Publicado: (2025)
Visible Iris Area as a Quality Metric for Reliable Iris Recognition Under Pupil Dilation and Eyelid Occlusion
por: Pessaud, Jack, et al.
Publicado: (2025)
por: Pessaud, Jack, et al.
Publicado: (2025)
Ejemplares similares
-
All4One: Symbiotic Neighbour Contrastive Learning via Self-Attention and Redundancy Reduction
por: Estepa, Imanol G., et al.
Publicado: (2023) -
Conjuring Positive Pairs for Efficient Unification of Representation Learning and Image Synthesis
por: Estepa, Imanol G., et al.
Publicado: (2025) -
Precision at Scale: Domain-Specific Datasets On-Demand
por: Rodríguez-de-Vera, Jesús M, et al.
Publicado: (2024) -
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning
por: Semenov, Andrei, et al.
Publicado: (2024) -
Learning Association via Track-Detection Matching for Multi-Object Tracking
por: Adžemović, Momir
Publicado: (2025)