InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Wei, Cong, Zhong, Yujie, Tan, Haoxian, Zeng, Yingsen, Liu, Yong, Zhao, Zheng, Yang, Yujiu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HyperSeg: Towards Universal Visual Segmentation with Large Language Model
di: Wei, Cong, et al.
Pubblicazione: (2024)
di: Wei, Cong, et al.
Pubblicazione: (2024)
Advancing Visual Large Language Model for Multi-granular Versatile Perception
di: Xiang, Wentao, et al.
Pubblicazione: (2025)
di: Xiang, Wentao, et al.
Pubblicazione: (2025)
LaSagnA: Language-based Segmentation Assistant for Complex Queries
di: Wei, Cong, et al.
Pubblicazione: (2024)
di: Wei, Cong, et al.
Pubblicazione: (2024)
LinVT: Empower Your Image-level Large Language Model to Understand Videos
di: Gao, Lishuai, et al.
Pubblicazione: (2024)
di: Gao, Lishuai, et al.
Pubblicazione: (2024)
One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
di: Bai, Zechen, et al.
Pubblicazione: (2024)
di: Bai, Zechen, et al.
Pubblicazione: (2024)
UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection
di: Zeng, Yingsen, et al.
Pubblicazione: (2024)
di: Zeng, Yingsen, et al.
Pubblicazione: (2024)
Instruct-Imagen: Image Generation with Multi-modal Instruction
di: Hu, Hexiang, et al.
Pubblicazione: (2024)
di: Hu, Hexiang, et al.
Pubblicazione: (2024)
InstructX: Towards Unified Visual Editing with MLLM Guidance
di: Mou, Chong, et al.
Pubblicazione: (2025)
di: Mou, Chong, et al.
Pubblicazione: (2025)
DisTime: Distribution-based Time Representation for Video Large Language Models
di: Zeng, Yingsen, et al.
Pubblicazione: (2025)
di: Zeng, Yingsen, et al.
Pubblicazione: (2025)
InstructVEdit: A Holistic Approach for Instructional Video Editing
di: Zhang, Chi, et al.
Pubblicazione: (2025)
di: Zhang, Chi, et al.
Pubblicazione: (2025)
InstructSAM: Segment Any Instance with Any Instructions
di: Yuan, Yuqian, et al.
Pubblicazione: (2026)
di: Yuan, Yuqian, et al.
Pubblicazione: (2026)
Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model
di: Peng, Jihua, et al.
Pubblicazione: (2025)
di: Peng, Jihua, et al.
Pubblicazione: (2025)
IPAdapter-Instruct: Resolving Ambiguity in Image-based Conditioning using Instruct Prompts
di: Rowles, Ciara, et al.
Pubblicazione: (2024)
di: Rowles, Ciara, et al.
Pubblicazione: (2024)
Generative Timelines for Instructed Visual Assembly
di: Pardo, Alejandro, et al.
Pubblicazione: (2024)
di: Pardo, Alejandro, et al.
Pubblicazione: (2024)
InstructGIE: Towards Generalizable Image Editing
di: Meng, Zichong, et al.
Pubblicazione: (2024)
di: Meng, Zichong, et al.
Pubblicazione: (2024)
InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing
di: Lin, Kun-Hsiang, et al.
Pubblicazione: (2025)
di: Lin, Kun-Hsiang, et al.
Pubblicazione: (2025)
Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model
di: Zhang, Wenqi, et al.
Pubblicazione: (2024)
di: Zhang, Wenqi, et al.
Pubblicazione: (2024)
InstructPix2NeRF: Instructed 3D Portrait Editing from a Single Image
di: Li, Jianhui, et al.
Pubblicazione: (2023)
di: Li, Jianhui, et al.
Pubblicazione: (2023)
VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation
di: Hong, Jihwan, et al.
Pubblicazione: (2026)
di: Hong, Jihwan, et al.
Pubblicazione: (2026)
MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment
di: Liu, Jihao, et al.
Pubblicazione: (2024)
di: Liu, Jihao, et al.
Pubblicazione: (2024)
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
di: Sakai, Yuki, et al.
Pubblicazione: (2025)
di: Sakai, Yuki, et al.
Pubblicazione: (2025)
Seer: Language Instructed Video Prediction with Latent Diffusion Models
di: Gu, Xianfan, et al.
Pubblicazione: (2023)
di: Gu, Xianfan, et al.
Pubblicazione: (2023)
Learning to Instruct for Visual Instruction Tuning
di: Zhou, Zhihan, et al.
Pubblicazione: (2025)
di: Zhou, Zhihan, et al.
Pubblicazione: (2025)
World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering
di: Wang, Jiacong, et al.
Pubblicazione: (2024)
di: Wang, Jiacong, et al.
Pubblicazione: (2024)
InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models
di: Wang, Xunguang, et al.
Pubblicazione: (2023)
di: Wang, Xunguang, et al.
Pubblicazione: (2023)
LITA: Language Instructed Temporal-Localization Assistant
di: Huang, De-An, et al.
Pubblicazione: (2024)
di: Huang, De-An, et al.
Pubblicazione: (2024)
VLM's Eye Examination: Instruct and Inspect Visual Competency of Vision Language Models
di: Hyeon-Woo, Nam, et al.
Pubblicazione: (2024)
di: Hyeon-Woo, Nam, et al.
Pubblicazione: (2024)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
di: Yang, Shuai, et al.
Pubblicazione: (2025)
di: Yang, Shuai, et al.
Pubblicazione: (2025)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
di: Romero, David, et al.
Pubblicazione: (2024)
di: Romero, David, et al.
Pubblicazione: (2024)
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
di: Zheng, Haojie, et al.
Pubblicazione: (2026)
di: Zheng, Haojie, et al.
Pubblicazione: (2026)
InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning
di: Wan, Zifu, et al.
Pubblicazione: (2025)
di: Wan, Zifu, et al.
Pubblicazione: (2025)
Visual-Instructed Degradation Diffusion for All-in-One Image Restoration
di: Luo, Wenyang, et al.
Pubblicazione: (2025)
di: Luo, Wenyang, et al.
Pubblicazione: (2025)
MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models
di: Huang, Yu, et al.
Pubblicazione: (2025)
di: Huang, Yu, et al.
Pubblicazione: (2025)
TIGER: Text-Instructed 3D Gaussian Retrieval and Coherent Editing
di: Xu, Teng, et al.
Pubblicazione: (2024)
di: Xu, Teng, et al.
Pubblicazione: (2024)
TP-Seg: Task-Prototype Framework for Unified Medical Lesion Segmentation
di: Xu, Jiawei, et al.
Pubblicazione: (2026)
di: Xu, Jiawei, et al.
Pubblicazione: (2026)
Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
di: Zheng, Rongkun, et al.
Pubblicazione: (2025)
di: Zheng, Rongkun, et al.
Pubblicazione: (2025)
TokenSeg: Efficient 3D Medical Image Segmentation via Hierarchical Visual Token Compression
di: Zeng, Sen, et al.
Pubblicazione: (2026)
di: Zeng, Sen, et al.
Pubblicazione: (2026)
InstructEngine: Instruction-driven Text-to-Image Alignment
di: Lu, Xingyu, et al.
Pubblicazione: (2025)
di: Lu, Xingyu, et al.
Pubblicazione: (2025)
Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
di: Tan, Jing, et al.
Pubblicazione: (2026)
di: Tan, Jing, et al.
Pubblicazione: (2026)
Empowering Segmentation Ability to Multi-modal Large Language Models
di: Yang, Yuqi, et al.
Pubblicazione: (2024)
di: Yang, Yuqi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
HyperSeg: Towards Universal Visual Segmentation with Large Language Model
di: Wei, Cong, et al.
Pubblicazione: (2024) -
Advancing Visual Large Language Model for Multi-granular Versatile Perception
di: Xiang, Wentao, et al.
Pubblicazione: (2025) -
LaSagnA: Language-based Segmentation Assistant for Complex Queries
di: Wei, Cong, et al.
Pubblicazione: (2024) -
LinVT: Empower Your Image-level Large Language Model to Understand Videos
di: Gao, Lishuai, et al.
Pubblicazione: (2024) -
One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
di: Bai, Zechen, et al.
Pubblicazione: (2024)