MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Camuffo, Elena, Barbato, Francesco, Ozay, Mete, Milani, Simone, Michieli, Umberto |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhanced Model Robustness to Input Corruptions by Per-corruption Adaptation of Normalization Statistics
by: Camuffo, Elena, et al.
Published: (2024)
by: Camuffo, Elena, et al.
Published: (2024)
Object-conditioned Bag of Instances for Few-Shot Personalized Instance Recognition
by: Michieli, Umberto, et al.
Published: (2024)
by: Michieli, Umberto, et al.
Published: (2024)
Cross-Architecture Auxiliary Feature Space Translation for Efficient Few-Shot Personalized Object Detection
by: Barbato, Francesco, et al.
Published: (2024)
by: Barbato, Francesco, et al.
Published: (2024)
Learning from Mistakes: Self-Regularizing Hierarchical Representations in Point Cloud Semantic Segmentation
by: Camuffo, Elena, et al.
Published: (2023)
by: Camuffo, Elena, et al.
Published: (2023)
FFT-based Selection and Optimization of Statistics for Robust Recognition of Severely Corrupted Images
by: Camuffo, Elena, et al.
Published: (2024)
by: Camuffo, Elena, et al.
Published: (2024)
Swiss DINO: Efficient and Versatile Vision Framework for On-device Personal Object Search
by: Paramonov, Kirill, et al.
Published: (2024)
by: Paramonov, Kirill, et al.
Published: (2024)
Feature-Space Generative Models for One-Shot Class-Incremental Learning
by: Foster, Jack, et al.
Published: (2026)
by: Foster, Jack, et al.
Published: (2026)
Continual Road-Scene Semantic Segmentation via Feature-Aligned Symmetric Multi-Modal Network
by: Barbato, Francesco, et al.
Published: (2023)
by: Barbato, Francesco, et al.
Published: (2023)
LoRA.rar: Learning to Merge LoRAs via Hypernetworks for Subject-Style Conditioned Image Generation
by: Shenaj, Donald, et al.
Published: (2024)
by: Shenaj, Donald, et al.
Published: (2024)
DreamCache: Finetuning-Free Lightweight Personalized Image Generation via Feature Caching
by: Aiello, Emanuele, et al.
Published: (2024)
by: Aiello, Emanuele, et al.
Published: (2024)
Controllable Forgetting Mechanism for Few-Shot Class-Incremental Learning
by: Paramonov, Kirill, et al.
Published: (2025)
by: Paramonov, Kirill, et al.
Published: (2025)
A Modular System for Enhanced Robustness of Multimedia Understanding Networks via Deep Parametric Estimation
by: Barbato, Francesco, et al.
Published: (2024)
by: Barbato, Francesco, et al.
Published: (2024)
Continual Learning for LiDAR Semantic Segmentation: Class-Incremental and Coarse-to-Fine strategies on Sparse Data
by: Camuffo, Elena, et al.
Published: (2023)
by: Camuffo, Elena, et al.
Published: (2023)
Split&Splat: Zero-Shot Panoptic Segmentation via Explicit Instance Modeling and 3D Gaussian Splatting
by: Monchieri, Leonardo, et al.
Published: (2026)
by: Monchieri, Leonardo, et al.
Published: (2026)
HOP to the Next Tasks and Domains for Continual Learning in NLP
by: Michieli, Umberto, et al.
Published: (2024)
by: Michieli, Umberto, et al.
Published: (2024)
ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking
by: Ge, Jiawei, et al.
Published: (2026)
by: Ge, Jiawei, et al.
Published: (2026)
Continual Error Correction on Low-Resource Devices
by: Paramonov, Kirill, et al.
Published: (2025)
by: Paramonov, Kirill, et al.
Published: (2025)
Cross-domain Few-shot Object Detection with Multi-modal Textual Enrichment
by: Shangguan, Zeyu, et al.
Published: (2025)
by: Shangguan, Zeyu, et al.
Published: (2025)
Fusion-Mamba for Cross-modality Object Detection
by: Dong, Wenhao, et al.
Published: (2024)
by: Dong, Wenhao, et al.
Published: (2024)
MCAD: Multi-teacher Cross-modal Alignment Distillation for efficient image-text retrieval
by: Lei, Youbo, et al.
Published: (2023)
by: Lei, Youbo, et al.
Published: (2023)
MemLoRA: Distilling Expert Adapters for On-Device Memory Systems
by: Bini, Massimo, et al.
Published: (2025)
by: Bini, Massimo, et al.
Published: (2025)
Awesome Multi-modal Object Tracking
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
SAGE: Semantic-Driven Adaptive Gaussian Splatting in Extended Reality
by: Schiavo, Chiara, et al.
Published: (2025)
by: Schiavo, Chiara, et al.
Published: (2025)
Deep Neural Network Models Trained With A Fixed Random Classifier Transfer Better Across Domains
by: Ali, Hafiz Tiomoko, et al.
Published: (2024)
by: Ali, Hafiz Tiomoko, et al.
Published: (2024)
Focus on Focus: Focus-oriented Representation Learning and Multi-view Cross-modal Alignment for Glioma Grading
by: Pan, Li, et al.
Published: (2024)
by: Pan, Li, et al.
Published: (2024)
FedPromo: Federated Lightweight Proxy Models at the Edge Bring New Domains to Foundation Models
by: Caligiuri, Matteo, et al.
Published: (2025)
by: Caligiuri, Matteo, et al.
Published: (2025)
Hierarchical Multi-modal Transformer for Cross-modal Long Document Classification
by: Liu, Tengfei, et al.
Published: (2024)
by: Liu, Tengfei, et al.
Published: (2024)
Robust Domain Generalization for Multi-modal Object Recognition
by: Qiao, Yuxin, et al.
Published: (2024)
by: Qiao, Yuxin, et al.
Published: (2024)
Memory-based Cross-modal Semantic Alignment Network for Radiology Report Generation
by: Tao, Yitian, et al.
Published: (2024)
by: Tao, Yitian, et al.
Published: (2024)
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment
by: Xing, Yifei, et al.
Published: (2024)
by: Xing, Yifei, et al.
Published: (2024)
Diffexplainer: Towards Cross-modal Global Explanations with Diffusion Models
by: Pennisi, Matteo, et al.
Published: (2024)
by: Pennisi, Matteo, et al.
Published: (2024)
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
by: Chen, Haonan, et al.
Published: (2026)
by: Chen, Haonan, et al.
Published: (2026)
LLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Models
by: Liao, Pan, et al.
Published: (2026)
by: Liao, Pan, et al.
Published: (2026)
Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image
by: He, Xiao, et al.
Published: (2025)
by: He, Xiao, et al.
Published: (2025)
MedMAP: Promoting Incomplete Multi-modal Brain Tumor Segmentation with Alignment
by: Liu, Tianyi, et al.
Published: (2024)
by: Liu, Tianyi, et al.
Published: (2024)
Cross-modality Guidance-aided Multi-modal Learning with Dual Attention for MRI Brain Tumor Grading
by: Xu, Dunyuan, et al.
Published: (2024)
by: Xu, Dunyuan, et al.
Published: (2024)
Cross-View Referring Multi-Object Tracking
by: Chen, Sijia, et al.
Published: (2024)
by: Chen, Sijia, et al.
Published: (2024)
VisioFirm: Cross-Platform AI-assisted Annotation Tool for Computer Vision
by: Ghazouali, Safouane El, et al.
Published: (2025)
by: Ghazouali, Safouane El, et al.
Published: (2025)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Similar Items
-
Enhanced Model Robustness to Input Corruptions by Per-corruption Adaptation of Normalization Statistics
by: Camuffo, Elena, et al.
Published: (2024) -
Object-conditioned Bag of Instances for Few-Shot Personalized Instance Recognition
by: Michieli, Umberto, et al.
Published: (2024) -
Cross-Architecture Auxiliary Feature Space Translation for Efficient Few-Shot Personalized Object Detection
by: Barbato, Francesco, et al.
Published: (2024) -
Learning from Mistakes: Self-Regularizing Hierarchical Representations in Point Cloud Semantic Segmentation
by: Camuffo, Elena, et al.
Published: (2023) -
FFT-based Selection and Optimization of Statistics for Robust Recognition of Severely Corrupted Images
by: Camuffo, Elena, et al.
Published: (2024)