Harmony: A Joint Self-Supervised and Weakly-Supervised Framework for Learning General Purpose Visual Representations
Fuente:
arXiv
Salvato in:
| Autori principali: | Baharoon, Mohammed, Klein, Jonathan, Michels, Dominik L. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MB-DSMIL-CL-PL: Scalable Weakly Supervised Ovarian Cancer Subtype Classification and Localisation Using Contrastive and Prototype Learning with Frozen Patch Features
di: Jenkins, Marcus, et al.
Pubblicazione: (2026)
di: Jenkins, Marcus, et al.
Pubblicazione: (2026)
Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
di: Komurcu, Kursat, et al.
Pubblicazione: (2026)
di: Komurcu, Kursat, et al.
Pubblicazione: (2026)
DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations
di: Jin, Hang, et al.
Pubblicazione: (2025)
di: Jin, Hang, et al.
Pubblicazione: (2025)
LRVS-Fashion: Extending Visual Search with Referring Instructions
di: Lepage, Simon, et al.
Pubblicazione: (2023)
di: Lepage, Simon, et al.
Pubblicazione: (2023)
JVLGS: Joint Vision-Language Gas Leak Segmentation
di: Zhao, Xinlong, et al.
Pubblicazione: (2025)
di: Zhao, Xinlong, et al.
Pubblicazione: (2025)
Learning Sign Language Representation using CNN LSTM, 3DCNN, CNN RNN LSTM and CCN TD
di: Louison, Nikita, et al.
Pubblicazione: (2024)
di: Louison, Nikita, et al.
Pubblicazione: (2024)
Generating Diverse Agricultural Data for Vision-Based Farming Applications
di: Cieslak, Mikolaj, et al.
Pubblicazione: (2024)
di: Cieslak, Mikolaj, et al.
Pubblicazione: (2024)
E Pluribus Unum Interpretable Convolutional Neural Networks
di: Dimas, George, et al.
Pubblicazione: (2022)
di: Dimas, George, et al.
Pubblicazione: (2022)
Contrastive Consolidation of Top-Down Modulations Achieves Sparsely Supervised Continual Learning
di: Tran, Viet Anh Khoa, et al.
Pubblicazione: (2025)
di: Tran, Viet Anh Khoa, et al.
Pubblicazione: (2025)
LAESI: Leaf Area Estimation with Synthetic Imagery
di: Kałużny, Jacek, et al.
Pubblicazione: (2024)
di: Kałużny, Jacek, et al.
Pubblicazione: (2024)
ShapBPT: Image Feature Attributions Using Data-Aware Binary Partition Trees
di: Rashid, Muhammad, et al.
Pubblicazione: (2026)
di: Rashid, Muhammad, et al.
Pubblicazione: (2026)
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
di: Chen, Jingkun, et al.
Pubblicazione: (2025)
di: Chen, Jingkun, et al.
Pubblicazione: (2025)
Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Systematic Review and Practical Design Guidelines
di: Wimalasiri, Chathura
Pubblicazione: (2026)
di: Wimalasiri, Chathura
Pubblicazione: (2026)
When Less is Enough: Adaptive Token Reduction for Efficient Image Representation
di: Allakhverdov, Eduard, et al.
Pubblicazione: (2025)
di: Allakhverdov, Eduard, et al.
Pubblicazione: (2025)
Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition
di: Wang, Yu, et al.
Pubblicazione: (2024)
di: Wang, Yu, et al.
Pubblicazione: (2024)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
Fixed-Threshold Evaluation of a Hybrid CNN-ViT for AI-Generated Image Detection Across Photos and Art
di: Khan, Md Ashik, et al.
Pubblicazione: (2025)
di: Khan, Md Ashik, et al.
Pubblicazione: (2025)
Does CLIP perceive art the same way we do?
di: Asperti, Andrea, et al.
Pubblicazione: (2025)
di: Asperti, Andrea, et al.
Pubblicazione: (2025)
Training for X-Ray Vision: Amodal Segmentation, Amodal Content Completion, and View-Invariant Object Representation from Multi-Camera Video
di: Moore, Alexander, et al.
Pubblicazione: (2025)
di: Moore, Alexander, et al.
Pubblicazione: (2025)
On the Inherent Robustness of One-Stage Object Detection against Out-of-Distribution Data
di: Martinez-Seras, Aitor, et al.
Pubblicazione: (2024)
di: Martinez-Seras, Aitor, et al.
Pubblicazione: (2024)
HieraEdgeNet: A Multi-Scale Edge-Enhanced Framework for Automated Pollen Recognition
di: Long, Yuchong, et al.
Pubblicazione: (2025)
di: Long, Yuchong, et al.
Pubblicazione: (2025)
Advancing Brain Tumor Segmentation via Attention-based 3D U-Net Architecture and Digital Image Processing
di: Gad, Eyad, et al.
Pubblicazione: (2025)
di: Gad, Eyad, et al.
Pubblicazione: (2025)
WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery
di: Ayanzadeh, Aydin, et al.
Pubblicazione: (2026)
di: Ayanzadeh, Aydin, et al.
Pubblicazione: (2026)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
di: Tu, Songjun, et al.
Pubblicazione: (2025)
di: Tu, Songjun, et al.
Pubblicazione: (2025)
Dual-sensing driving detection model
di: K, Leon C. C., et al.
Pubblicazione: (2025)
di: K, Leon C. C., et al.
Pubblicazione: (2025)
Akasha 2: Hamiltonian State Space Duality and Visual-Language Joint Embedding Predictive Architectur
di: Meziani, Yani
Pubblicazione: (2026)
di: Meziani, Yani
Pubblicazione: (2026)
Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling
di: Jung, Seoik, et al.
Pubblicazione: (2025)
di: Jung, Seoik, et al.
Pubblicazione: (2025)
On Memory: A comparison of memory mechanisms in world models
di: Laird, Eli J., et al.
Pubblicazione: (2025)
di: Laird, Eli J., et al.
Pubblicazione: (2025)
Multi-Modal Self-Supervised Learning for Surgical Feedback Effectiveness Assessment
di: Gupta, Arushi, et al.
Pubblicazione: (2024)
di: Gupta, Arushi, et al.
Pubblicazione: (2024)
Rethinking Visual Intelligence: Insights from Video Pretraining
di: Acuaviva, Pablo, et al.
Pubblicazione: (2025)
di: Acuaviva, Pablo, et al.
Pubblicazione: (2025)
Image Reconstruction as a Tool for Feature Analysis
di: Allakhverdov, Eduard, et al.
Pubblicazione: (2025)
di: Allakhverdov, Eduard, et al.
Pubblicazione: (2025)
Fine-grained spatial-temporal perception for gas leak segmentation
di: Zhao, Xinlong, et al.
Pubblicazione: (2025)
di: Zhao, Xinlong, et al.
Pubblicazione: (2025)
DSER: Spectral Epipolar Representation for Efficient Light Field Depth Estimation
di: Mohammad, Noor Islam S., et al.
Pubblicazione: (2025)
di: Mohammad, Noor Islam S., et al.
Pubblicazione: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
di: Qesaraku, Bjorna, et al.
Pubblicazione: (2025)
di: Qesaraku, Bjorna, et al.
Pubblicazione: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
di: Viveiros, André G., et al.
Pubblicazione: (2025)
di: Viveiros, André G., et al.
Pubblicazione: (2025)
Unpacking Hateful Memes: Presupposed Context and False Claims
di: Cai, Weibin, et al.
Pubblicazione: (2025)
di: Cai, Weibin, et al.
Pubblicazione: (2025)
ARTPS: Depth-Enhanced Hybrid Anomaly Detection and Learnable Curiosity Score for Autonomous Rover Target Prioritization
di: Baydemir, Poyraz
Pubblicazione: (2025)
di: Baydemir, Poyraz
Pubblicazione: (2025)
TRACES: Temporal Recall with Contextual Embeddings for Real-Time Video Anomaly Detection
di: Siddiqui, Yousuf Ahmed, et al.
Pubblicazione: (2025)
di: Siddiqui, Yousuf Ahmed, et al.
Pubblicazione: (2025)
Hierarchical Multi-Positive Contrastive Learning for Patent Image Retrieval
di: Kavimandan, Kshitij, et al.
Pubblicazione: (2025)
di: Kavimandan, Kshitij, et al.
Pubblicazione: (2025)
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
di: Li, Jinhao, et al.
Pubblicazione: (2024)
di: Li, Jinhao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MB-DSMIL-CL-PL: Scalable Weakly Supervised Ovarian Cancer Subtype Classification and Localisation Using Contrastive and Prototype Learning with Frozen Patch Features
di: Jenkins, Marcus, et al.
Pubblicazione: (2026) -
Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
di: Komurcu, Kursat, et al.
Pubblicazione: (2026) -
DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations
di: Jin, Hang, et al.
Pubblicazione: (2025) -
LRVS-Fashion: Extending Visual Search with Referring Instructions
di: Lepage, Simon, et al.
Pubblicazione: (2023) -
JVLGS: Joint Vision-Language Gas Leak Segmentation
di: Zhao, Xinlong, et al.
Pubblicazione: (2025)