LARE: Latent Augmentation using Regional Embedding with Vision-Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | Sakurai, Kosuke, Ishii, Tatsuya, Shimizu, Ryotaro, Song, Linxin, Goto, Masayuki |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision and Language Reference Prompt into SAM for Few-shot Segmentation
by: Sakurai, Kosuke, et al.
Published: (2025)
by: Sakurai, Kosuke, et al.
Published: (2025)
Attributed Synthetic Data Generation for Zero-shot Domain-specific Image Classification
by: Wang, Shijian, et al.
Published: (2025)
by: Wang, Shijian, et al.
Published: (2025)
Masked Language Prompting for Generative Data Augmentation in Few-shot Fashion Style Recognition
by: Hirakawa, Yuki, et al.
Published: (2025)
by: Hirakawa, Yuki, et al.
Published: (2025)
Training-free Conditional Image Embedding Framework Leveraging Large Vision Language Models
by: Kawarada, Masayuki, et al.
Published: (2025)
by: Kawarada, Masayuki, et al.
Published: (2025)
Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model
by: Wang, Shijian, et al.
Published: (2024)
by: Wang, Shijian, et al.
Published: (2024)
SCP: Spherical-Coordinate-based Learned Point Cloud Compression
by: Luo, Ao, et al.
Published: (2023)
by: Luo, Ao, et al.
Published: (2023)
Small Bird Detection using YOLOv7 with Test-Time Augmentation
by: Shigematsu, Kosuke
Published: (2023)
by: Shigematsu, Kosuke
Published: (2023)
Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model
by: Wang, Chenfeng, et al.
Published: (2026)
by: Wang, Chenfeng, et al.
Published: (2026)
Phantom of Latent for Large Language and Vision Models
by: Lee, Byung-Kwan, et al.
Published: (2024)
by: Lee, Byung-Kwan, et al.
Published: (2024)
FLIER: Few-shot Language Image Models Embedded with Latent Representations
by: Zhou, Zhinuo, et al.
Published: (2024)
by: Zhou, Zhinuo, et al.
Published: (2024)
Fashionability-Enhancing Outfit Image Editing with Conditional Diffusion Models
by: Qin, Qice, et al.
Published: (2024)
by: Qin, Qice, et al.
Published: (2024)
Probabilistic Embeddings for Frozen Vision-Language Models: Uncertainty Quantification with Gaussian Process Latent Variable Models
by: Venkataramanan, Aishwarya, et al.
Published: (2025)
by: Venkataramanan, Aishwarya, et al.
Published: (2025)
RegionGPT: Towards Region Understanding Vision Language Model
by: Guo, Qiushan, et al.
Published: (2024)
by: Guo, Qiushan, et al.
Published: (2024)
3D Aware Region Prompted Vision Language Model
by: Cheng, An-Chieh, et al.
Published: (2025)
by: Cheng, An-Chieh, et al.
Published: (2025)
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
by: Tao, Chenxin, et al.
Published: (2024)
by: Tao, Chenxin, et al.
Published: (2024)
Nomic Embed Vision: Expanding the Latent Space
by: Nussbaum, Zach, et al.
Published: (2024)
by: Nussbaum, Zach, et al.
Published: (2024)
Vision-aligned Latent Reasoning for Multi-modal Large Language Model
by: Jeon, Byungwoo, et al.
Published: (2026)
by: Jeon, Byungwoo, et al.
Published: (2026)
MultiEmo-Bench: Multi-label Visual Emotion Analysis for Multi-modal Large Language Models
by: Chen, Tianwei, et al.
Published: (2026)
by: Chen, Tianwei, et al.
Published: (2026)
VDMA: Video Question Answering with Dynamically Generated Multi-Agents
by: Kugo, Noriyuki, et al.
Published: (2024)
by: Kugo, Noriyuki, et al.
Published: (2024)
Lost in Embeddings: Information Loss in Vision-Language Models
by: Li, Wenyan, et al.
Published: (2025)
by: Li, Wenyan, et al.
Published: (2025)
Manipulating Vehicle 3D Shapes through Latent Space Editing
by: Miao, JiangDong, et al.
Published: (2024)
by: Miao, JiangDong, et al.
Published: (2024)
Normal and Abnormal Pathology Knowledge-Augmented Vision-Language Model for Anomaly Detection in Pathology Images
by: Song, Jinsol, et al.
Published: (2025)
by: Song, Jinsol, et al.
Published: (2025)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
by: Zhang, Jieyu, et al.
Published: (2024)
by: Zhang, Jieyu, et al.
Published: (2024)
Intra-Class Probabilistic Embeddings for Uncertainty Estimation in Vision-Language Models
by: Lin, Zhenxiang, et al.
Published: (2025)
by: Lin, Zhenxiang, et al.
Published: (2025)
Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
by: Pham, Tan-Hanh, et al.
Published: (2025)
by: Pham, Tan-Hanh, et al.
Published: (2025)
Image Recognition with Vision and Language Embeddings of VLMs
by: Volkov, Illia, et al.
Published: (2025)
by: Volkov, Illia, et al.
Published: (2025)
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
by: Byun, Sanghyun, et al.
Published: (2025)
by: Byun, Sanghyun, et al.
Published: (2025)
AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
by: jia, Feiyang, et al.
Published: (2026)
by: jia, Feiyang, et al.
Published: (2026)
Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models
by: Wang, Haoming, et al.
Published: (2026)
by: Wang, Haoming, et al.
Published: (2026)
Transferring Visual Explainability of Self-Explaining Models to Prediction-Only Models without Additional Training
by: Yoshikawa, Yuya, et al.
Published: (2025)
by: Yoshikawa, Yuya, et al.
Published: (2025)
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
by: Sun, Jingwen, et al.
Published: (2026)
by: Sun, Jingwen, et al.
Published: (2026)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
by: Li, Qiwei, et al.
Published: (2026)
by: Li, Qiwei, et al.
Published: (2026)
Seer: Language Instructed Video Prediction with Latent Diffusion Models
by: Gu, Xianfan, et al.
Published: (2023)
by: Gu, Xianfan, et al.
Published: (2023)
Cross-Attentive Multiview Fusion of Vision-Language Embeddings
by: Martins, Tomas Berriel, et al.
Published: (2026)
by: Martins, Tomas Berriel, et al.
Published: (2026)
Seeing the Unseen: Towards Zero-Shot Inspection for Wind Turbine Blades using Knowledge-Augmented Vision Language Models
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
Effectively Enhancing Vision Language Large Models by Prompt Augmentation and Caption Utilization
by: Zhao, Minyi, et al.
Published: (2024)
by: Zhao, Minyi, et al.
Published: (2024)
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
by: Sinha, Arkaprava, et al.
Published: (2025)
by: Sinha, Arkaprava, et al.
Published: (2025)
FLARE: Learning Future-Aware Latent Representations from Vision-Language Models for Autonomous Driving
by: Xie, Chengen, et al.
Published: (2026)
by: Xie, Chengen, et al.
Published: (2026)
Ego: Embedding-Guided Personalization of Vision-Language Models
by: Seifi, Soroush, et al.
Published: (2026)
by: Seifi, Soroush, et al.
Published: (2026)
Similar Items
-
Vision and Language Reference Prompt into SAM for Few-shot Segmentation
by: Sakurai, Kosuke, et al.
Published: (2025) -
Attributed Synthetic Data Generation for Zero-shot Domain-specific Image Classification
by: Wang, Shijian, et al.
Published: (2025) -
Masked Language Prompting for Generative Data Augmentation in Few-shot Fashion Style Recognition
by: Hirakawa, Yuki, et al.
Published: (2025) -
Training-free Conditional Image Embedding Framework Leveraging Large Vision Language Models
by: Kawarada, Masayuki, et al.
Published: (2025) -
Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model
by: Wang, Shijian, et al.
Published: (2024)