Gramian Multimodal Representation Learning and Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Cicchetti, Giordano, Grassucci, Eleonora, Sigillo, Luigi, Comminiello, Danilo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A TRIANGLE Enables Multimodal Alignment Beyond Cosine Similarity
by: Cicchetti, Giordano, et al.
Published: (2025)
by: Cicchetti, Giordano, et al.
Published: (2025)
Closing the gap in multimodal medical representation alignment
by: Grassucci, Eleonora, et al.
Published: (2026)
by: Grassucci, Eleonora, et al.
Published: (2026)
Generalizing Medical Image Representations via Quaternion Wavelet Networks
by: Sigillo, Luigi, et al.
Published: (2023)
by: Sigillo, Luigi, et al.
Published: (2023)
Multi-View Hypercomplex Learning for Breast Cancer Screening
by: Lopez, Eleonora, et al.
Published: (2022)
by: Lopez, Eleonora, et al.
Published: (2022)
Closing the Modality Gap Aligns Group-Wise Semantics
by: Grassucci, Eleonora, et al.
Published: (2026)
by: Grassucci, Eleonora, et al.
Published: (2026)
Guess What I Think: Streamlined EEG-to-Image Generation with Latent Diffusion Models
by: Lopez, Eleonora, et al.
Published: (2024)
by: Lopez, Eleonora, et al.
Published: (2024)
FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
by: Gramaccioni, Riccardo Fosco, et al.
Published: (2025)
by: Gramaccioni, Riccardo Fosco, et al.
Published: (2025)
Metadata, Wavelet, and Time Aware Diffusion Models for Satellite Image Super Resolution
by: Sigillo, Luigi, et al.
Published: (2025)
by: Sigillo, Luigi, et al.
Published: (2025)
Hierarchical Hypercomplex Network for Multimodal Emotion Recognition
by: Lopez, Eleonora, et al.
Published: (2024)
by: Lopez, Eleonora, et al.
Published: (2024)
Semantic Compression via Multimodal Representation Learning
by: Grassucci, Eleonora, et al.
Published: (2025)
by: Grassucci, Eleonora, et al.
Published: (2025)
Latent Wavelet Diffusion For Ultra-High-Resolution Image Synthesis
by: Sigillo, Luigi, et al.
Published: (2025)
by: Sigillo, Luigi, et al.
Published: (2025)
Quaternion Wavelet-Conditioned Diffusion Models for Image Super-Resolution
by: Sigillo, Luigi, et al.
Published: (2025)
by: Sigillo, Luigi, et al.
Published: (2025)
Language-Oriented Semantic Latent Representation for Image Transmission
by: Cicchetti, Giordano, et al.
Published: (2024)
by: Cicchetti, Giordano, et al.
Published: (2024)
Ship in Sight: Diffusion Models for Ship-Image Super Resolution
by: Sigillo, Luigi, et al.
Published: (2024)
by: Sigillo, Luigi, et al.
Published: (2024)
NAF-DPM: A Nonlinear Activation-Free Diffusion Probabilistic Model for Document Enhancement
by: Cicchetti, Giordano, et al.
Published: (2024)
by: Cicchetti, Giordano, et al.
Published: (2024)
Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information
by: Marinoni, Christian, et al.
Published: (2025)
by: Marinoni, Christian, et al.
Published: (2025)
Towards Explaining Hypercomplex Neural Networks
by: Lopez, Eleonora, et al.
Published: (2024)
by: Lopez, Eleonora, et al.
Published: (2024)
EGR-Net: A Novel Embedding Gramian Representation CNN for Intelligent Fault Diagnosis
by: Jia, Linshan
Published: (2025)
by: Jia, Linshan
Published: (2025)
Semantically Guided Representation Learning For Action Anticipation
by: Diko, Anxhelo, et al.
Published: (2024)
by: Diko, Anxhelo, et al.
Published: (2024)
TACE: Tumor-Aware Counterfactual Explanations
by: Rossi, Eleonora Beatrice, et al.
Published: (2024)
by: Rossi, Eleonora Beatrice, et al.
Published: (2024)
Learning Relative Representations for Fine-Grained Multimodal Alignment with Limited Data
by: Kim, Shiwon, et al.
Published: (2026)
by: Kim, Shiwon, et al.
Published: (2026)
GAF-FusionNet: Multimodal ECG Analysis via Gramian Angular Fields and Split Attention
by: Qin, Jiahao, et al.
Published: (2024)
by: Qin, Jiahao, et al.
Published: (2024)
Align Your Query: Representation Alignment for Multimodality Medical Object Detection
by: Seo, Ara, et al.
Published: (2025)
by: Seo, Ara, et al.
Published: (2025)
A Wavelet Diffusion GAN for Image Super-Resolution
by: Aloisi, Lorenzo, et al.
Published: (2024)
by: Aloisi, Lorenzo, et al.
Published: (2024)
Training-Free Multimodal Guidance for Video to Audio Generation
by: Grassucci, Eleonora, et al.
Published: (2025)
by: Grassucci, Eleonora, et al.
Published: (2025)
No Alignment Needed for Generation: Learning Linearly Separable Representations in Diffusion Models
by: Yun, Junno, et al.
Published: (2025)
by: Yun, Junno, et al.
Published: (2025)
Biological Plausibility and Representational Alignment of Feedback Alignment in Convolutional Networks
by: Lance, Jake, et al.
Published: (2026)
by: Lance, Jake, et al.
Published: (2026)
Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment
by: Chatterjee, Abhiroop, et al.
Published: (2025)
by: Chatterjee, Abhiroop, et al.
Published: (2025)
The Surprising Ineffectiveness of Pre-Trained Visual Representations for Model-Based Reinforcement Learning
by: Schneider, Moritz, et al.
Published: (2024)
by: Schneider, Moritz, et al.
Published: (2024)
Attention-Map Augmentation for Hypercomplex Breast Cancer Classification
by: Lopez, Eleonora, et al.
Published: (2023)
by: Lopez, Eleonora, et al.
Published: (2023)
Reconstruction Alignment Improves Unified Multimodal Models
by: Xie, Ji, et al.
Published: (2025)
by: Xie, Ji, et al.
Published: (2025)
Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated Perturbation
by: Xu, Jie, et al.
Published: (2025)
by: Xu, Jie, et al.
Published: (2025)
With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You
by: Gröger, Fabian, et al.
Published: (2025)
by: Gröger, Fabian, et al.
Published: (2025)
A Closer Look at Multimodal Representation Collapse
by: Chaudhuri, Abhra, et al.
Published: (2025)
by: Chaudhuri, Abhra, et al.
Published: (2025)
Learning Noise-Robust Joint Representation for Multimodal Emotion Recognition under Incomplete Data Scenarios
by: Fan, Qi, et al.
Published: (2023)
by: Fan, Qi, et al.
Published: (2023)
Multi-Surrogate-Teacher Assistance for Representation Alignment in Fingerprint-based Indoor Localization
by: Nguyen, Son Minh, et al.
Published: (2024)
by: Nguyen, Son Minh, et al.
Published: (2024)
Federated Learning with Feedback Alignment
by: Baek, Incheol, et al.
Published: (2025)
by: Baek, Incheol, et al.
Published: (2025)
Improving Visual Representation Alignment Generation with GRPO
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
Temporal Embeddings: Scalable Self-Supervised Temporal Representation Learning from Spatiotemporal Data for Multimodal Computer Vision
by: Cao, Yi, et al.
Published: (2023)
by: Cao, Yi, et al.
Published: (2023)
Invariant Representation Guided Multimodal Sentiment Decoding with Sequential Variation Regularization
by: Xu, Guoyang, et al.
Published: (2024)
by: Xu, Guoyang, et al.
Published: (2024)
Similar Items
-
A TRIANGLE Enables Multimodal Alignment Beyond Cosine Similarity
by: Cicchetti, Giordano, et al.
Published: (2025) -
Closing the gap in multimodal medical representation alignment
by: Grassucci, Eleonora, et al.
Published: (2026) -
Generalizing Medical Image Representations via Quaternion Wavelet Networks
by: Sigillo, Luigi, et al.
Published: (2023) -
Multi-View Hypercomplex Learning for Breast Cancer Screening
by: Lopez, Eleonora, et al.
Published: (2022) -
Closing the Modality Gap Aligns Group-Wise Semantics
by: Grassucci, Eleonora, et al.
Published: (2026)