Joint Low-level and High-level Textual Representation Learning with Multiple Masking Strategies
Fuente:
arXiv
Guardado en:
| Autores principales: | Tang, Zhengmi, Mitsui, Yuto, Miyazaki, Tomo, Omachi, Shinichiro |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Cross-Domain Multi-Targeted Adversarial Attacks
por: Gonçalves, Taïga, et al.
Publicado: (2025)
por: Gonçalves, Taïga, et al.
Publicado: (2025)
Controlling Rate, Distortion, and Realism: Towards a Single Comprehensive Neural Image Compression Model
por: Iwai, Shoma, et al.
Publicado: (2024)
por: Iwai, Shoma, et al.
Publicado: (2024)
GPSMamba: A Global Phase and Spectral Prompt-guided Mamba for Infrared Image Super-Resolution
por: Huang, Yongsong, et al.
Publicado: (2025)
por: Huang, Yongsong, et al.
Publicado: (2025)
Class-agnostic 3D Segmentation by Granularity-Consistent Automatic 2D Mask Tracking
por: Wang, Juan, et al.
Publicado: (2025)
por: Wang, Juan, et al.
Publicado: (2025)
IRSRMamba: Infrared Image Super-Resolution via Mamba-based Wavelet Transform Feature Modulation Model
por: Huang, Yongsong, et al.
Publicado: (2024)
por: Huang, Yongsong, et al.
Publicado: (2024)
Infrared Image Super-Resolution: Systematic Review, and Future Trends
por: Huang, Yongsong, et al.
Publicado: (2022)
por: Huang, Yongsong, et al.
Publicado: (2022)
Texture and Noise Dual Adaptation for Infrared Image Super-Resolution
por: Huang, Yongsong, et al.
Publicado: (2023)
por: Huang, Yongsong, et al.
Publicado: (2023)
POCA: Pareto-Optimal Curriculum Alignment for Visual Text Generation
por: Fan, Yaohou, et al.
Publicado: (2026)
por: Fan, Yaohou, et al.
Publicado: (2026)
GTFMN: Guided Texture and Feature Modulation Network for Low-Light Image Enhancement and Super-Resolution
por: Huang, Yongsong, et al.
Publicado: (2026)
por: Huang, Yongsong, et al.
Publicado: (2026)
U-Harmony: Enhancing Joint Training for Segmentation Models with Universal Harmonization
por: Ma, Weiwei, et al.
Publicado: (2026)
por: Ma, Weiwei, et al.
Publicado: (2026)
Layout-Corrector: Alleviating Layout Sticking Phenomenon in Discrete Diffusion Model
por: Iwai, Shoma, et al.
Publicado: (2024)
por: Iwai, Shoma, et al.
Publicado: (2024)
Learning Joint ID-Textual Representation for ID-Preserving Image Synthesis
por: Liu, Zichuan, et al.
Publicado: (2025)
por: Liu, Zichuan, et al.
Publicado: (2025)
Learning Multiple Representations with Inconsistency-Guided Detail Regularization for Mask-Guided Matting
por: Jiang, Weihao, et al.
Publicado: (2024)
por: Jiang, Weihao, et al.
Publicado: (2024)
DiCTI: Diffusion-based Clothing Designer via Text-guided Input
por: Lampe, Ajda, et al.
Publicado: (2024)
por: Lampe, Ajda, et al.
Publicado: (2024)
MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
por: Wei, Wei, et al.
Publicado: (2025)
por: Wei, Wei, et al.
Publicado: (2025)
Textual Query-Driven Mask Transformer for Domain Generalized Segmentation
por: Pak, Byeonghyun, et al.
Publicado: (2024)
por: Pak, Byeonghyun, et al.
Publicado: (2024)
Multiplicative Loss for Enhancing Semantic Segmentation in Medical and Cellular Images
por: Yokoi, Yuto, et al.
Publicado: (2025)
por: Yokoi, Yuto, et al.
Publicado: (2025)
Multiple Instance Learning Framework with Masked Hard Instance Mining for Gigapixel Histopathology Image Analysis
por: Tang, Wenhao, et al.
Publicado: (2025)
por: Tang, Wenhao, et al.
Publicado: (2025)
MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures
por: Morin, Lucas, et al.
Publicado: (2025)
por: Morin, Lucas, et al.
Publicado: (2025)
PRIOR: Prototype Representation Joint Learning from Medical Images and Reports
por: Cheng, Pujin, et al.
Publicado: (2023)
por: Cheng, Pujin, et al.
Publicado: (2023)
MaskFi: Unsupervised Learning of WiFi and Vision Representations for Multimodal Human Activity Recognition
por: Yang, Jianfei, et al.
Publicado: (2024)
por: Yang, Jianfei, et al.
Publicado: (2024)
Multitask Learning for SAR Ship Detection with Gaussian-Mask Joint Segmentation
por: Zhao, Ming, et al.
Publicado: (2024)
por: Zhao, Ming, et al.
Publicado: (2024)
Anatomical 3D Style Transfer Enabling Efficient Federated Learning with Extremely Low Communication Costs
por: Shibata, Yuto, et al.
Publicado: (2024)
por: Shibata, Yuto, et al.
Publicado: (2024)
Joint-Embedding Predictive Architecture for Self-Supervised Learning of Mask Classification Architecture
por: Kim, Dong-Hee, et al.
Publicado: (2024)
por: Kim, Dong-Hee, et al.
Publicado: (2024)
BIMM: Brain Inspired Masked Modeling for Video Representation Learning
por: Wan, Zhifan, et al.
Publicado: (2024)
por: Wan, Zhifan, et al.
Publicado: (2024)
T-MAE: Temporal Masked Autoencoders for Point Cloud Representation Learning
por: Wei, Weijie, et al.
Publicado: (2023)
por: Wei, Weijie, et al.
Publicado: (2023)
Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
por: Liu, Jiazhen, et al.
Publicado: (2025)
por: Liu, Jiazhen, et al.
Publicado: (2025)
Improving Image Restoration through Removing Degradations in Textual Representations
por: Lin, Jingbo, et al.
Publicado: (2023)
por: Lin, Jingbo, et al.
Publicado: (2023)
Robust Representation Learning in Masked Autoencoders
por: Shrivastava, Anika, et al.
Publicado: (2026)
por: Shrivastava, Anika, et al.
Publicado: (2026)
TrackMAE: Video Representation Learning via Track Mask and Predict
por: Vandeghen, Renaud, et al.
Publicado: (2026)
por: Vandeghen, Renaud, et al.
Publicado: (2026)
MaskFuser: Masked Fusion of Joint Multi-Modal Tokenization for End-to-End Autonomous Driving
por: Duan, Yiqun, et al.
Publicado: (2024)
por: Duan, Yiqun, et al.
Publicado: (2024)
Multiple Random Masking Autoencoder Ensembles for Robust Multimodal Semi-supervised Learning
por: Todoran, Alexandru-Raul, et al.
Publicado: (2024)
por: Todoran, Alexandru-Raul, et al.
Publicado: (2024)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
por: Shvetsova, Nina, et al.
Publicado: (2025)
por: Shvetsova, Nina, et al.
Publicado: (2025)
MaskGaussian: Adaptive 3D Gaussian Representation from Probabilistic Masks
por: Liu, Yifei, et al.
Publicado: (2024)
por: Liu, Yifei, et al.
Publicado: (2024)
Instance-aware Image Colorization with Controllable Textual Descriptions and Segmentation Masks
por: An, Yanru, et al.
Publicado: (2025)
por: An, Yanru, et al.
Publicado: (2025)
LV-MAE: Learning Long Video Representations through Masked-Embedding Autoencoders
por: Naiman, Ilan, et al.
Publicado: (2025)
por: Naiman, Ilan, et al.
Publicado: (2025)
MLIP: Medical Language-Image Pre-training with Masked Local Representation Learning
por: Liu, Jiarun, et al.
Publicado: (2024)
por: Liu, Jiarun, et al.
Publicado: (2024)
Learning Unsupervised Gaze Representation via Eye Mask Driven Information Bottleneck
por: Jiang, Yangzhou, et al.
Publicado: (2024)
por: Jiang, Yangzhou, et al.
Publicado: (2024)
Social-MAE: Social Masked Autoencoder for Multi-person Motion Representation Learning
por: Ehsanpour, Mahsa, et al.
Publicado: (2024)
por: Ehsanpour, Mahsa, et al.
Publicado: (2024)
Less is More: Decoder-Free Masked Modeling for Efficient Skeleton Representation Learning
por: Do, Jeonghyeok, et al.
Publicado: (2026)
por: Do, Jeonghyeok, et al.
Publicado: (2026)
Ejemplares similares
-
Towards Cross-Domain Multi-Targeted Adversarial Attacks
por: Gonçalves, Taïga, et al.
Publicado: (2025) -
Controlling Rate, Distortion, and Realism: Towards a Single Comprehensive Neural Image Compression Model
por: Iwai, Shoma, et al.
Publicado: (2024) -
GPSMamba: A Global Phase and Spectral Prompt-guided Mamba for Infrared Image Super-Resolution
por: Huang, Yongsong, et al.
Publicado: (2025) -
Class-agnostic 3D Segmentation by Granularity-Consistent Automatic 2D Mask Tracking
por: Wang, Juan, et al.
Publicado: (2025) -
IRSRMamba: Infrared Image Super-Resolution via Mamba-based Wavelet Transform Feature Modulation Model
por: Huang, Yongsong, et al.
Publicado: (2024)