ZERO: Industry-ready Vision Foundation Model with Multi-modal Prompts
Fuente:
arXiv
Guardado en:
| Autores principales: | Choi, Sangbum, Go, Kyeongryeol, Jang, Taewoong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Continual Expansion of Data Coverage: Automatic Text-guided Edge-case Synthesis
por: Go, Kyeongryeol
Publicado: (2025)
por: Go, Kyeongryeol
Publicado: (2025)
Edge-case Synthesis for Fisheye Object Detection: A Data-centric Perspective
por: Kim, Seunghyeon, et al.
Publicado: (2025)
por: Kim, Seunghyeon, et al.
Publicado: (2025)
Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models
por: He, Yuting, et al.
Publicado: (2026)
por: He, Yuting, et al.
Publicado: (2026)
Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
por: Jang, Young Kyun, et al.
Publicado: (2024)
por: Jang, Young Kyun, et al.
Publicado: (2024)
Exploring Efficient Foundational Multi-modal Models for Video Summarization
por: Samel, Karan, et al.
Publicado: (2024)
por: Samel, Karan, et al.
Publicado: (2024)
Why Only Text: Empowering Vision-and-Language Navigation with Multi-modal Prompts
por: Hong, Haodong, et al.
Publicado: (2024)
por: Hong, Haodong, et al.
Publicado: (2024)
Towards Scalable Foundation Model for Multi-modal and Hyperspectral Geospatial Data
por: Si, Haozhe, et al.
Publicado: (2025)
por: Si, Haozhe, et al.
Publicado: (2025)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
por: Walimbe, Soham, et al.
Publicado: (2025)
por: Walimbe, Soham, et al.
Publicado: (2025)
Vision-Language Consistency Guided Multi-modal Prompt Learning for Blind AI Generated Image Quality Assessment
por: Fu, Jun, et al.
Publicado: (2024)
por: Fu, Jun, et al.
Publicado: (2024)
Advancing Stroke Risk Prediction Using a Multi-modal Foundation Model
por: Delgrange, Camille, et al.
Publicado: (2024)
por: Delgrange, Camille, et al.
Publicado: (2024)
Explaining Multi-modal Large Language Models by Analyzing their Vision Perception
por: Giulivi, Loris, et al.
Publicado: (2024)
por: Giulivi, Loris, et al.
Publicado: (2024)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
por: Woo, Sangmin, et al.
Publicado: (2024)
por: Woo, Sangmin, et al.
Publicado: (2024)
Historical Test-time Prompt Tuning for Vision Foundation Models
por: Zhang, Jingyi, et al.
Publicado: (2024)
por: Zhang, Jingyi, et al.
Publicado: (2024)
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
por: Jang, Young Kyun, et al.
Publicado: (2024)
por: Jang, Young Kyun, et al.
Publicado: (2024)
OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
por: He, Xingxin, et al.
Publicado: (2025)
por: He, Xingxin, et al.
Publicado: (2025)
Diffusion Model Patching via Mixture-of-Prompts
por: Ham, Seokil, et al.
Publicado: (2024)
por: Ham, Seokil, et al.
Publicado: (2024)
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
por: Lu, Yi, et al.
Publicado: (2025)
por: Lu, Yi, et al.
Publicado: (2025)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
por: Ohi, Masanari, et al.
Publicado: (2024)
por: Ohi, Masanari, et al.
Publicado: (2024)
MGHanD: Multi-modal Guidance for authentic Hand Diffusion
por: Eum, Taehyeon, et al.
Publicado: (2025)
por: Eum, Taehyeon, et al.
Publicado: (2025)
Eye-for-an-eye: Appearance Transfer with Semantic Correspondence in Diffusion Models
por: Go, Sooyeon, et al.
Publicado: (2024)
por: Go, Sooyeon, et al.
Publicado: (2024)
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
por: Li, Yanwei, et al.
Publicado: (2024)
por: Li, Yanwei, et al.
Publicado: (2024)
RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
por: Woo, Sangmin, et al.
Publicado: (2024)
por: Woo, Sangmin, et al.
Publicado: (2024)
Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM
por: Park, Hyobin, et al.
Publicado: (2026)
por: Park, Hyobin, et al.
Publicado: (2026)
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
por: Cheng, Zihui, et al.
Publicado: (2024)
por: Cheng, Zihui, et al.
Publicado: (2024)
Decompose the model: Mechanistic interpretability in image models with Generalized Integrated Gradients (GIG)
por: Kim, Yearim, et al.
Publicado: (2024)
por: Kim, Yearim, et al.
Publicado: (2024)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
por: Kim, Mingyeong, et al.
Publicado: (2026)
por: Kim, Mingyeong, et al.
Publicado: (2026)
Revealing Multi-View Hallucination in Large Vision-Language Models
por: Park, Wooje, et al.
Publicado: (2026)
por: Park, Wooje, et al.
Publicado: (2026)
Curriculum Prompting Foundation Models for Medical Image Segmentation
por: Zheng, Xiuqi, et al.
Publicado: (2024)
por: Zheng, Xiuqi, et al.
Publicado: (2024)
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
por: Khalil, Ahmad, et al.
Publicado: (2025)
por: Khalil, Ahmad, et al.
Publicado: (2025)
Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification
por: Xuan, Yunyi, et al.
Publicado: (2024)
por: Xuan, Yunyi, et al.
Publicado: (2024)
Img2Loc: Revisiting Image Geolocalization using Multi-modality Foundation Models and Image-based Retrieval-Augmented Generation
por: Zhou, Zhongliang, et al.
Publicado: (2024)
por: Zhou, Zhongliang, et al.
Publicado: (2024)
Adversarial Prompt Tuning for Vision-Language Models
por: Zhang, Jiaming, et al.
Publicado: (2023)
por: Zhang, Jiaming, et al.
Publicado: (2023)
Adversarial Prompt Distillation for Vision-Language Models
por: Luo, Lin, et al.
Publicado: (2024)
por: Luo, Lin, et al.
Publicado: (2024)
Evolving Prompt Adaptation for Vision-Language Models
por: Zhang, Enming, et al.
Publicado: (2026)
por: Zhang, Enming, et al.
Publicado: (2026)
Vision-Core Guided Contrastive Learning for Balanced Multi-modal Prognosis Prediction of Stroke
por: Chen, Liren, et al.
Publicado: (2026)
por: Chen, Liren, et al.
Publicado: (2026)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
por: Wang, Xin, et al.
Publicado: (2024)
por: Wang, Xin, et al.
Publicado: (2024)
Understanding the Transfer Limits of Vision Foundation Models
por: Huang, Shiqi, et al.
Publicado: (2026)
por: Huang, Shiqi, et al.
Publicado: (2026)
Xray-Visual Models: Scaling Vision models on Industry Scale Data
por: Mishra, Shlok, et al.
Publicado: (2026)
por: Mishra, Shlok, et al.
Publicado: (2026)
Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models
por: Li, Nanxi, et al.
Publicado: (2026)
por: Li, Nanxi, et al.
Publicado: (2026)
A Multi-Modal Foundation Model to Assist People with Blindness and Low Vision in Environmental Interaction
por: Hao, Yu, et al.
Publicado: (2023)
por: Hao, Yu, et al.
Publicado: (2023)
Ejemplares similares
-
Towards Continual Expansion of Data Coverage: Automatic Text-guided Edge-case Synthesis
por: Go, Kyeongryeol
Publicado: (2025) -
Edge-case Synthesis for Fisheye Object Detection: A Data-centric Perspective
por: Kim, Seunghyeon, et al.
Publicado: (2025) -
Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models
por: He, Yuting, et al.
Publicado: (2026) -
Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
por: Jang, Young Kyun, et al.
Publicado: (2024) -
Exploring Efficient Foundational Multi-modal Models for Video Summarization
por: Samel, Karan, et al.
Publicado: (2024)