GSSF: Generalized Structural Sparse Function for Deep Cross-modal Metric Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Diao, Haiwen, Zhang, Ying, Gao, Shang, Zhu, Jiawen, Chen, Long, Lu, Huchuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep Boosting Learning: A Brand-new Cooperative Approach for Image-Text Matching
by: Diao, Haiwen, et al.
Published: (2024)
by: Diao, Haiwen, et al.
Published: (2024)
UniPT: Universal Parallel Tuning for Transfer Learning with Efficient Parameter and Memory
by: Diao, Haiwen, et al.
Published: (2023)
by: Diao, Haiwen, et al.
Published: (2023)
SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning
by: Diao, Haiwen, et al.
Published: (2024)
by: Diao, Haiwen, et al.
Published: (2024)
Regularizing Subspace Redundancy of Low-Rank Adaptation
by: Zhu, Yue, et al.
Published: (2025)
by: Zhu, Yue, et al.
Published: (2025)
Unveiling Encoder-Free Vision-Language Models
by: Diao, Haiwen, et al.
Published: (2024)
by: Diao, Haiwen, et al.
Published: (2024)
IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identification
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Deep Reversible Consistency Learning for Cross-modal Retrieval
by: Pu, Ruitao, et al.
Published: (2025)
by: Pu, Ruitao, et al.
Published: (2025)
Multi-Scale and Detail-Enhanced Segment Anything Model for Salient Object Detection
by: Gao, Shixuan, et al.
Published: (2024)
by: Gao, Shixuan, et al.
Published: (2024)
Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
by: Zhu, Yixin, et al.
Published: (2026)
by: Zhu, Yixin, et al.
Published: (2026)
Other Tokens Matter: Exploring Global and Local Features of Vision Transformers for Object Re-Identification
by: Wang, Yingquan, et al.
Published: (2024)
by: Wang, Yingquan, et al.
Published: (2024)
Fantastic Animals and Where to Find Them: Segment Any Marine Animal with Dual SAM
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification
by: Zhu, Yue, et al.
Published: (2025)
by: Zhu, Yue, et al.
Published: (2025)
CAMeL: Cross-modality Adaptive Meta-Learning for Text-based Person Retrieval
by: Yu, Hang, et al.
Published: (2025)
by: Yu, Hang, et al.
Published: (2025)
Balanced Multi-modal Federated Learning via Cross-Modal Infiltration
by: Fan, Yunfeng, et al.
Published: (2023)
by: Fan, Yunfeng, et al.
Published: (2023)
Learning Brain Representation with Hierarchical Visual Embeddings
by: Zheng, Jiawen, et al.
Published: (2026)
by: Zheng, Jiawen, et al.
Published: (2026)
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models
by: Tang, Hao, et al.
Published: (2026)
by: Tang, Hao, et al.
Published: (2026)
Med-Banana-50K: A Cross-modality Large-Scale Dataset for Text-guided Medical Image Editing
by: Chen, Zhihui, et al.
Published: (2025)
by: Chen, Zhihui, et al.
Published: (2025)
HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
by: Cai, Qi, et al.
Published: (2025)
by: Cai, Qi, et al.
Published: (2025)
MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Cross-modal Counterfactual Explanations: Uncovering Decision Factors and Dataset Biases in Subjective Classification
by: Baia, Alina Elena, et al.
Published: (2025)
by: Baia, Alina Elena, et al.
Published: (2025)
G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment
by: Zhang, Juan, et al.
Published: (2024)
by: Zhang, Juan, et al.
Published: (2024)
MTNet: Learning modality-aware representation with transformer for RGBT tracking
by: Hou, Ruichao, et al.
Published: (2025)
by: Hou, Ruichao, et al.
Published: (2025)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
by: Chow, Wei, et al.
Published: (2024)
by: Chow, Wei, et al.
Published: (2024)
MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition
by: Zhang, Haoyang, et al.
Published: (2025)
by: Zhang, Haoyang, et al.
Published: (2025)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
by: Yao, Lei, et al.
Published: (2025)
by: Yao, Lei, et al.
Published: (2025)
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
by: Huang, Hailang, et al.
Published: (2024)
by: Huang, Hailang, et al.
Published: (2024)
DeepSPG: Exploring Deep Semantic Prior Guidance for Low-light Image Enhancement with Multimodal Learning
by: Lu, Jialang, et al.
Published: (2025)
by: Lu, Jialang, et al.
Published: (2025)
Tile Classification Based Viewport Prediction with Multi-modal Fusion Transformer
by: Zhang, Zhihao, et al.
Published: (2023)
by: Zhang, Zhihao, et al.
Published: (2023)
DIP: Diffusion Learning of Inconsistency Pattern for General DeepFake Detection
by: Nie, Fan, et al.
Published: (2024)
by: Nie, Fan, et al.
Published: (2024)
MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric
by: Lin, Haokun, et al.
Published: (2024)
by: Lin, Haokun, et al.
Published: (2024)
Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product Retrieval
by: Hu, Xiaowan, et al.
Published: (2024)
by: Hu, Xiaowan, et al.
Published: (2024)
Camera Trajectory Generation: A Comprehensive Survey of Methods, Metrics, and Future Directions
by: Dehghanian, Zahra, et al.
Published: (2025)
by: Dehghanian, Zahra, et al.
Published: (2025)
Cross-modal Causal Intervention for Alzheimer's Disease Prediction
by: Jin, Yutao, et al.
Published: (2025)
by: Jin, Yutao, et al.
Published: (2025)
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
by: Dai, Guangyu, et al.
Published: (2025)
by: Dai, Guangyu, et al.
Published: (2025)
Deep Compositional Phase Diffusion for Long Motion Sequence Generation
by: Au, Ho Yin, et al.
Published: (2025)
by: Au, Ho Yin, et al.
Published: (2025)
Consistent and Invariant Generalization Learning for Short-video Misinformation Detection
by: Guo, Hanghui, et al.
Published: (2025)
by: Guo, Hanghui, et al.
Published: (2025)
Visual Grounding with Multi-modal Conditional Adaptation
by: Yao, Ruilin, et al.
Published: (2024)
by: Yao, Ruilin, et al.
Published: (2024)
Similar Items
-
Deep Boosting Learning: A Brand-new Cooperative Approach for Image-Text Matching
by: Diao, Haiwen, et al.
Published: (2024) -
UniPT: Universal Parallel Tuning for Transfer Learning with Efficient Parameter and Memory
by: Diao, Haiwen, et al.
Published: (2023) -
SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning
by: Diao, Haiwen, et al.
Published: (2024) -
Regularizing Subspace Redundancy of Low-Rank Adaptation
by: Zhu, Yue, et al.
Published: (2025) -
Unveiling Encoder-Free Vision-Language Models
by: Diao, Haiwen, et al.
Published: (2024)