AutoMMLab: Automatically Generating Deployable Models from Language Instructions for Computer Vision Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Zekang, Zeng, Wang, Jin, Sheng, Qian, Chen, Luo, Ping, Liu, Wentao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NADER: Neural Architecture Design via Multi-Agent Collaboration
von: Yang, Zekang, et al.
Veröffentlicht: (2024)
von: Yang, Zekang, et al.
Veröffentlicht: (2024)
When Pedestrian Detection Meets Multi-Modal Learning: Generalist Model and Benchmark Dataset
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
You Only Learn One Query: Learning Unified Human Query for Single-Stage Multi-Person Multi-Task Human-Centric Perception
von: Jin, Sheng, et al.
Veröffentlicht: (2023)
von: Jin, Sheng, et al.
Veröffentlicht: (2023)
KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
von: Yang, Jie, et al.
Veröffentlicht: (2025)
von: Yang, Jie, et al.
Veröffentlicht: (2025)
Vision-Language Models for Vision Tasks: A Survey
von: Zhang, Jingyi, et al.
Veröffentlicht: (2023)
von: Zhang, Jingyi, et al.
Veröffentlicht: (2023)
KptLLM: Unveiling the Power of Large Language Model for Keypoint Comprehension
von: Yang, Jie, et al.
Veröffentlicht: (2024)
von: Yang, Jie, et al.
Veröffentlicht: (2024)
AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
von: Wu, Xiyang, et al.
Veröffentlicht: (2024)
von: Wu, Xiyang, et al.
Veröffentlicht: (2024)
TCFormer: Visual Recognition via Token Clustering Transformer
von: Zeng, Wang, et al.
Veröffentlicht: (2024)
von: Zeng, Wang, et al.
Veröffentlicht: (2024)
GKGNet: Group K-Nearest Neighbor based Graph Convolutional Network for Multi-Label Image Recognition
von: Yao, Ruijie, et al.
Veröffentlicht: (2023)
von: Yao, Ruijie, et al.
Veröffentlicht: (2023)
Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
von: Jin, Hyundong, et al.
Veröffentlicht: (2025)
von: Jin, Hyundong, et al.
Veröffentlicht: (2025)
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
von: Shen, Yang, et al.
Veröffentlicht: (2024)
von: Shen, Yang, et al.
Veröffentlicht: (2024)
UniFS: Universal Few-shot Instance Perception with Point Representations
von: Jin, Sheng, et al.
Veröffentlicht: (2024)
von: Jin, Sheng, et al.
Veröffentlicht: (2024)
AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering
von: Chen, Xiuyuan, et al.
Veröffentlicht: (2023)
von: Chen, Xiuyuan, et al.
Veröffentlicht: (2023)
Navigation Instruction Generation with BEV Perception and Large Language Models
von: Fan, Sheng, et al.
Veröffentlicht: (2024)
von: Fan, Sheng, et al.
Veröffentlicht: (2024)
From Generalist to Specialist: Adapting Vision Language Models via Task-Specific Visual Instruction Tuning
von: Bai, Yang, et al.
Veröffentlicht: (2024)
von: Bai, Yang, et al.
Veröffentlicht: (2024)
SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
von: Zhan, Yang, et al.
Veröffentlicht: (2024)
von: Zhan, Yang, et al.
Veröffentlicht: (2024)
Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
von: Zhang, Jinrui, et al.
Veröffentlicht: (2024)
von: Zhang, Jinrui, et al.
Veröffentlicht: (2024)
AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting
von: Zhou, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Zhou, Xiaoyu, et al.
Veröffentlicht: (2025)
VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
von: Wu, Jiannan, et al.
Veröffentlicht: (2024)
von: Wu, Jiannan, et al.
Veröffentlicht: (2024)
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2025)
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2025)
Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model
von: Wang, Shijian, et al.
Veröffentlicht: (2024)
von: Wang, Shijian, et al.
Veröffentlicht: (2024)
Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?
von: Feng, Hengyi, et al.
Veröffentlicht: (2025)
von: Feng, Hengyi, et al.
Veröffentlicht: (2025)
Fewer Denoising Steps or Cheaper Per-Step Inference: Towards Compute-Optimal Diffusion Model Deployment
von: Du, Zhenbang, et al.
Veröffentlicht: (2025)
von: Du, Zhenbang, et al.
Veröffentlicht: (2025)
Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
From Prompts to Deployment: Auto-Curated Domain-Specific Dataset Generation via Diffusion Models
von: Yoon, Dongsik, et al.
Veröffentlicht: (2026)
von: Yoon, Dongsik, et al.
Veröffentlicht: (2026)
Cross-Domain Few-Shot Learning via Multi-View Collaborative Optimization with Vision-Language Models
von: Chen, Dexia, et al.
Veröffentlicht: (2025)
von: Chen, Dexia, et al.
Veröffentlicht: (2025)
Controlling Vision-Language Models for Multi-Task Image Restoration
von: Luo, Ziwei, et al.
Veröffentlicht: (2023)
von: Luo, Ziwei, et al.
Veröffentlicht: (2023)
Masked AutoDecoder is Effective Multi-Task Vision Generalist
von: Qiu, Han, et al.
Veröffentlicht: (2024)
von: Qiu, Han, et al.
Veröffentlicht: (2024)
Improving Continuous Sign Language Recognition with Adapted Image Models
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
von: Chen, Zhe, et al.
Veröffentlicht: (2023)
von: Chen, Zhe, et al.
Veröffentlicht: (2023)
TaskCLIP: Extend Large Vision-Language Model for Task Oriented Object Detection
von: Chen, Hanning, et al.
Veröffentlicht: (2024)
von: Chen, Hanning, et al.
Veröffentlicht: (2024)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
von: Sarkar, Ayushman, et al.
Veröffentlicht: (2025)
von: Sarkar, Ayushman, et al.
Veröffentlicht: (2025)
VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling
von: Zhang, Qian, et al.
Veröffentlicht: (2024)
von: Zhang, Qian, et al.
Veröffentlicht: (2024)
Instruction-Free Tuning of Large Vision Language Models for Medical Instruction Following
von: Kang, Myeongkyun, et al.
Veröffentlicht: (2026)
von: Kang, Myeongkyun, et al.
Veröffentlicht: (2026)
PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language Models
von: He, Runze, et al.
Veröffentlicht: (2025)
von: He, Runze, et al.
Veröffentlicht: (2025)
Diagnosing the Compositional Knowledge of Vision Language Models from a Game-Theoretic View
von: Wang, Jin, et al.
Veröffentlicht: (2024)
von: Wang, Jin, et al.
Veröffentlicht: (2024)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
von: An, Ruichuan, et al.
Veröffentlicht: (2025)
von: An, Ruichuan, et al.
Veröffentlicht: (2025)
Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language
von: Chen, Yicheng, et al.
Veröffentlicht: (2024)
von: Chen, Yicheng, et al.
Veröffentlicht: (2024)
Dynamic Spatial-Temporal Aggregation for Skeleton-Aware Sign Language Recognition
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
NADER: Neural Architecture Design via Multi-Agent Collaboration
von: Yang, Zekang, et al.
Veröffentlicht: (2024) -
When Pedestrian Detection Meets Multi-Modal Learning: Generalist Model and Benchmark Dataset
von: Zhang, Yi, et al.
Veröffentlicht: (2024) -
You Only Learn One Query: Learning Unified Human Query for Single-Stage Multi-Person Multi-Task Human-Centric Perception
von: Jin, Sheng, et al.
Veröffentlicht: (2023) -
KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
von: Yang, Jie, et al.
Veröffentlicht: (2025) -
Vision-Language Models for Vision Tasks: A Survey
von: Zhang, Jingyi, et al.
Veröffentlicht: (2023)