Using Multimodal Foundation Models and Clustering for Improved Style Ambiguity Loss
Fuente:
arXiv
Saved in:
| Main Author: | Baker, James |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Style Ambiguity Loss Using CLIP
by: Baker, James
Published: (2024)
by: Baker, James
Published: (2024)
Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning
by: Wang, Yufei, et al.
Published: (2025)
by: Wang, Yufei, et al.
Published: (2025)
Towards Ambiguity-Free Spatial Foundation Model: Rethinking and Decoupling Depth Ambiguity
by: Xu, Xiaohao, et al.
Published: (2025)
by: Xu, Xiaohao, et al.
Published: (2025)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
by: Wu, Yi, et al.
Published: (2025)
by: Wu, Yi, et al.
Published: (2025)
A Generative Foundation Model for Multimodal Histopathology
by: Xiang, Jinxi, et al.
Published: (2026)
by: Xiang, Jinxi, et al.
Published: (2026)
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding
by: Suglia, Alessandro, et al.
Published: (2024)
by: Suglia, Alessandro, et al.
Published: (2024)
A Multimodal Vision Foundation Model for Clinical Dermatology
by: Yan, Siyuan, et al.
Published: (2024)
by: Yan, Siyuan, et al.
Published: (2024)
Improving Generalization of Medical Image Registration Foundation Model
by: Hu, Jing, et al.
Published: (2025)
by: Hu, Jing, et al.
Published: (2025)
Multimodal Foundation Models Exploit Text to Make Medical Image Predictions
by: Buckley, Thomas, et al.
Published: (2023)
by: Buckley, Thomas, et al.
Published: (2023)
StyleGuard: Preventing Text-to-Image-Model-based Style Mimicry Attacks by Style Perturbations
by: Li, Yanjie, et al.
Published: (2025)
by: Li, Yanjie, et al.
Published: (2025)
Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
by: Yao, Louie Hong, et al.
Published: (2025)
by: Yao, Louie Hong, et al.
Published: (2025)
Toward Robust Multimodal Learning using Multimodal Foundational Models
by: Zhao, Xianbing, et al.
Published: (2024)
by: Zhao, Xianbing, et al.
Published: (2024)
Supervised Fine-tuning in turn Improves Visual Foundation Models
by: Jiang, Xiaohu, et al.
Published: (2024)
by: Jiang, Xiaohu, et al.
Published: (2024)
EyeFound: A Multimodal Generalist Foundation Model for Ophthalmic Imaging
by: Shi, Danli, et al.
Published: (2024)
by: Shi, Danli, et al.
Published: (2024)
A Multimodal Knowledge-enhanced Whole-slide Pathology Foundation Model
by: Xu, Yingxue, et al.
Published: (2024)
by: Xu, Yingxue, et al.
Published: (2024)
DISC-GAN: Disentangling Style and Content for Cluster-Specific Synthetic Underwater Image Generation
by: Varur, Sneha, et al.
Published: (2025)
by: Varur, Sneha, et al.
Published: (2025)
Cluster and Predict Latent Patches for Improved Masked Image Modeling
by: Darcet, Timothée, et al.
Published: (2025)
by: Darcet, Timothée, et al.
Published: (2025)
StyleVAR: Controllable Image Style Transfer via Visual Autoregressive Modeling
by: Jing, Liqi, et al.
Published: (2026)
by: Jing, Liqi, et al.
Published: (2026)
Dance Style Recognition Using Laban Movement Analysis
by: Turab, Muhammad, et al.
Published: (2025)
by: Turab, Muhammad, et al.
Published: (2025)
Examining the Commitments and Difficulties Inherent in Multimodal Foundation Models for Street View Imagery
by: Yang, Zhenyuan, et al.
Published: (2024)
by: Yang, Zhenyuan, et al.
Published: (2024)
SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation Models
by: Westerhoff, Justus, et al.
Published: (2025)
by: Westerhoff, Justus, et al.
Published: (2025)
RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models
by: Heinrich, Greg, et al.
Published: (2024)
by: Heinrich, Greg, et al.
Published: (2024)
Domain-Specific Foundation Model Improves AI-Based Analysis of Neuropathology
by: Verma, Ruchika, et al.
Published: (2025)
by: Verma, Ruchika, et al.
Published: (2025)
StyleMamba : State Space Model for Efficient Text-driven Image Style Transfer
by: Wang, Zijia, et al.
Published: (2024)
by: Wang, Zijia, et al.
Published: (2024)
A Semantically Enhanced Generative Foundation Model Improves Pathological Image Synthesis
by: Guan, Xianchao, et al.
Published: (2025)
by: Guan, Xianchao, et al.
Published: (2025)
FLORO: A Multimodal Geospatial Foundation Model for Ecological Remote Sensing Across Sensors and Scales
by: Rodriguez, Jorge L., et al.
Published: (2026)
by: Rodriguez, Jorge L., et al.
Published: (2026)
Modeling Depth Ambiguity: A Mixture-Density Representation for Flying-Point-Free Depth Estimation
by: Bian, Siyuan, et al.
Published: (2026)
by: Bian, Siyuan, et al.
Published: (2026)
Monocular Biomechanical Tracking of Fingers with Inverse Kinematics to Foundation Models
by: Cotton, R. James, et al.
Published: (2026)
by: Cotton, R. James, et al.
Published: (2026)
MM-NeRF: Multimodal-Guided 3D Multi-Style Transfer of Neural Radiance Field
by: Yang, Zijiang, et al.
Published: (2023)
by: Yang, Zijiang, et al.
Published: (2023)
A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Space
by: Liu, Huijie, et al.
Published: (2025)
by: Liu, Huijie, et al.
Published: (2025)
Active Learning for Animal Re-Identification with Ambiguity-Aware Sampling
by: Sani, Depanshu, et al.
Published: (2025)
by: Sani, Depanshu, et al.
Published: (2025)
A Geometric Multimodal Foundation Model Integrating Bp-MRI and Clinical Reports in Prostate Cancer Classification
by: Olmos, Juan A., et al.
Published: (2026)
by: Olmos, Juan A., et al.
Published: (2026)
ConstStyle: Robust Domain Generalization with Unified Style Transformation
by: Tran, Nam Duong, et al.
Published: (2025)
by: Tran, Nam Duong, et al.
Published: (2025)
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
by: Shangguan, Ziyao, et al.
Published: (2024)
by: Shangguan, Ziyao, et al.
Published: (2024)
AviationLMM: A Large Multimodal Foundation Model for Civil Aviation
by: Li, Wenbin, et al.
Published: (2026)
by: Li, Wenbin, et al.
Published: (2026)
Advancing Stroke Risk Prediction Using a Multi-modal Foundation Model
by: Delgrange, Camille, et al.
Published: (2024)
by: Delgrange, Camille, et al.
Published: (2024)
StyleX: A Trainable Metric for X-ray Style Distances
by: Eckert, Dominik, et al.
Published: (2024)
by: Eckert, Dominik, et al.
Published: (2024)
3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset
by: Ma, Xinyu, et al.
Published: (2024)
by: Ma, Xinyu, et al.
Published: (2024)
StyleCrafter: Enhancing Stylized Text-to-Video Generation with Style Adapter
by: Liu, Gongye, et al.
Published: (2023)
by: Liu, Gongye, et al.
Published: (2023)
How Much 3D Do Video Foundation Models Encode?
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
Similar Items
-
Style Ambiguity Loss Using CLIP
by: Baker, James
Published: (2024) -
Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning
by: Wang, Yufei, et al.
Published: (2025) -
Towards Ambiguity-Free Spatial Foundation Model: Rethinking and Decoupling Depth Ambiguity
by: Xu, Xiaohao, et al.
Published: (2025) -
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
by: Wu, Yi, et al.
Published: (2025) -
A Generative Foundation Model for Multimodal Histopathology
by: Xiang, Jinxi, et al.
Published: (2026)