A Dataset and Framework for Learning State-invariant Object Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Sarkar, Rohan, Kak, Avinash |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dual Pose-invariant Embeddings: Learning Category and Object-specific Discriminative Representations for Recognition and Retrieval
by: Sarkar, Rohan, et al.
Published: (2024)
by: Sarkar, Rohan, et al.
Published: (2024)
Region-Point Joint Representation for Effective Trajectory Similarity Learning
by: Long, Hao, et al.
Published: (2025)
by: Long, Hao, et al.
Published: (2025)
Towards Resource-Efficient Streaming of Large-Scale Medical Image Datasets for Deep Learning
by: Kulkarni, Pranav, et al.
Published: (2023)
by: Kulkarni, Pranav, et al.
Published: (2023)
Domain-invariant feature learning in brain MR imaging for content-based image retrieval
by: Tobari, Shuya, et al.
Published: (2025)
by: Tobari, Shuya, et al.
Published: (2025)
FOR: Finetuning for Object Level Open Vocabulary Image Retrieval
by: Levi, Hila, et al.
Published: (2024)
by: Levi, Hila, et al.
Published: (2024)
Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval
by: Sogi, Naoya, et al.
Published: (2024)
by: Sogi, Naoya, et al.
Published: (2024)
Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
by: Alomari, Hani, et al.
Published: (2025)
by: Alomari, Hani, et al.
Published: (2025)
AdaTask: A Task-aware Adaptive Learning Rate Approach to Multi-task Learning
by: Yang, Enneng, et al.
Published: (2022)
by: Yang, Enneng, et al.
Published: (2022)
Self-Supervised Learning as Discrete Communication
by: Zaher, Kawtar, et al.
Published: (2026)
by: Zaher, Kawtar, et al.
Published: (2026)
Deep Learning for Technical Document Classification
by: Jiang, Shuo, et al.
Published: (2021)
by: Jiang, Shuo, et al.
Published: (2021)
Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking
by: Zhu, Tianyu, et al.
Published: (2024)
by: Zhu, Tianyu, et al.
Published: (2024)
Studying Illustrations in Manuscripts: An Efficient Deep-Learning Approach
by: Evron, Yoav, et al.
Published: (2025)
by: Evron, Yoav, et al.
Published: (2025)
Online Learning via Memory: Retrieval-Augmented Detector Adaptation
by: Jian, Yanan, et al.
Published: (2024)
by: Jian, Yanan, et al.
Published: (2024)
Transfer Learning with Self-Supervised Vision Transformers for Snake Identification
by: Miyaguchi, Anthony, et al.
Published: (2024)
by: Miyaguchi, Anthony, et al.
Published: (2024)
Decoupled Training: Return of Frustratingly Easy Multi-Domain Learning
by: Wang, Ximei, et al.
Published: (2023)
by: Wang, Ximei, et al.
Published: (2023)
MOON Embedding: Multimodal Representation Learning for E-commerce Search Advertising
by: Fu, Chenghan, et al.
Published: (2025)
by: Fu, Chenghan, et al.
Published: (2025)
Transfer Learning and Mixup for Fine-Grained Few-Shot Fungi Classification
by: Tam, Jason Kahei, et al.
Published: (2025)
by: Tam, Jason Kahei, et al.
Published: (2025)
Hierarchy-of-Visual-Words: a Learning-based Approach for Trademark Image Retrieval
by: Lourenço, Vítor N., et al.
Published: (2019)
by: Lourenço, Vítor N., et al.
Published: (2019)
LotusFilter: Fast Diverse Nearest Neighbor Search via a Learned Cutoff Table
by: Matsui, Yusuke
Published: (2025)
by: Matsui, Yusuke
Published: (2025)
MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding
by: Zhang, Daoze, et al.
Published: (2025)
by: Zhang, Daoze, et al.
Published: (2025)
CoopHash: Cooperative Learning of Multipurpose Descriptor and Contrastive Pair Generator via Variational MCMC Teaching for Supervised Image Hashing
by: Doan, Khoa D., et al.
Published: (2022)
by: Doan, Khoa D., et al.
Published: (2022)
MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding
by: Wu, Junxian, et al.
Published: (2026)
by: Wu, Junxian, et al.
Published: (2026)
ArtCognition: A Multimodal AI Framework for Affective State Sensing from Visual and Kinematic Drawing Cues
by: Binaei-Haghighi, Behrad, et al.
Published: (2026)
by: Binaei-Haghighi, Behrad, et al.
Published: (2026)
MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding
by: Nie, Zhanheng, et al.
Published: (2025)
by: Nie, Zhanheng, et al.
Published: (2025)
A Guide to Similarity Measures
by: Levy, Avivit, et al.
Published: (2024)
by: Levy, Avivit, et al.
Published: (2024)
A Fashion Item Recommendation Model in Hyperbolic Space
by: Shimizu, Ryotaro, et al.
Published: (2024)
by: Shimizu, Ryotaro, et al.
Published: (2024)
Reasoning-Augmented Representations for Multimodal Retrieval
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
MuseChat: A Conversational Music Recommendation System for Videos
by: Dong, Zhikang, et al.
Published: (2023)
by: Dong, Zhikang, et al.
Published: (2023)
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
by: Xu, Yifan, et al.
Published: (2024)
by: Xu, Yifan, et al.
Published: (2024)
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation
by: Wen, Tiansheng, et al.
Published: (2025)
by: Wen, Tiansheng, et al.
Published: (2025)
GENIUS: A Generative Framework for Universal Multimodal Search
by: Kim, Sungyeon, et al.
Published: (2025)
by: Kim, Sungyeon, et al.
Published: (2025)
Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval
by: Li, Jun, et al.
Published: (2026)
by: Li, Jun, et al.
Published: (2026)
PICS: Pipeline for Image Captioning and Search
by: Rosario, Grant, et al.
Published: (2024)
by: Rosario, Grant, et al.
Published: (2024)
Proceedings of the 6th International Workshop on Reading Music Systems
by: Calvo-Zaragoza, Jorge, et al.
Published: (2024)
by: Calvo-Zaragoza, Jorge, et al.
Published: (2024)
iRAG: Advancing RAG for Videos with an Incremental Approach
by: Arefeen, Md Adnan, et al.
Published: (2024)
by: Arefeen, Md Adnan, et al.
Published: (2024)
Multi-Label Plant Species Classification with Self-Supervised Vision Transformers
by: Gustineli, Murilo, et al.
Published: (2024)
by: Gustineli, Murilo, et al.
Published: (2024)
Evidential Transformers for Improved Image Retrieval
by: Dordevic, Danilo, et al.
Published: (2024)
by: Dordevic, Danilo, et al.
Published: (2024)
Siamese Content-based Search Engine for a More Transparent Skin and Breast Cancer Diagnosis through Histological Imaging
by: Tabatabaei, Zahra, et al.
Published: (2024)
by: Tabatabaei, Zahra, et al.
Published: (2024)
InvDiff: Invariant Guidance for Bias Mitigation in Diffusion Models
by: Hou, Min, et al.
Published: (2024)
by: Hou, Min, et al.
Published: (2024)
Similar Items
-
Dual Pose-invariant Embeddings: Learning Category and Object-specific Discriminative Representations for Recognition and Retrieval
by: Sarkar, Rohan, et al.
Published: (2024) -
Region-Point Joint Representation for Effective Trajectory Similarity Learning
by: Long, Hao, et al.
Published: (2025) -
Towards Resource-Efficient Streaming of Large-Scale Medical Image Datasets for Deep Learning
by: Kulkarni, Pranav, et al.
Published: (2023) -
Domain-invariant feature learning in brain MR imaging for content-based image retrieval
by: Tobari, Shuya, et al.
Published: (2025) -
FOR: Finetuning for Object Level Open Vocabulary Image Retrieval
by: Levi, Hila, et al.
Published: (2024)