FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Fan, Zhu, Yousong, Li, Xin, Zhan, Yufei, Zhao, Hongyin, Zheng, Shurong, Wang, Yaowei, Tang, Ming, Wang, Jinqiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring
by: Zhan, Yufei, et al.
Published: (2024)
by: Zhan, Yufei, et al.
Published: (2024)
GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
by: Zheng, Shurong, et al.
Published: (2026)
by: Zheng, Shurong, et al.
Published: (2026)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
by: Yang, Fan, et al.
Published: (2026)
by: Yang, Fan, et al.
Published: (2026)
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2024)
by: Zhan, Yufei, et al.
Published: (2024)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
Griffon: Spelling out All Object Locations at Any Granularity with Large Language Models
by: Zhan, Yufei, et al.
Published: (2023)
by: Zhan, Yufei, et al.
Published: (2023)
From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation
by: Yang, Fan, et al.
Published: (2025)
by: Yang, Fan, et al.
Published: (2025)
FOCUS: Fine-grained Optimization with Semantic Guided Understanding for Pedestrian Attributes Recognition
by: An, Hongyan, et al.
Published: (2025)
by: An, Hongyan, et al.
Published: (2025)
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
by: Yu, Jiachen, et al.
Published: (2025)
by: Yu, Jiachen, et al.
Published: (2025)
MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation
by: Zhu, Yuanbing, et al.
Published: (2024)
by: Zhu, Yuanbing, et al.
Published: (2024)
Efficient Masked Autoencoders with Self-Consistency
by: Li, Zhaowen, et al.
Published: (2023)
by: Li, Zhaowen, et al.
Published: (2023)
Unifying Ontology Construction and Semantic Alignment for Deterministic Enterprise Reasoning at Scale
by: Zhu, Hongyin
Published: (2026)
by: Zhu, Hongyin
Published: (2026)
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
AAformer: Auto-Aligned Transformer for Person Re-Identification
by: Zhu, Kuan, et al.
Published: (2021)
by: Zhu, Kuan, et al.
Published: (2021)
Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models
by: Zhang, Enming, et al.
Published: (2024)
by: Zhang, Enming, et al.
Published: (2024)
AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection
by: Gu, Zhaopeng, et al.
Published: (2025)
by: Gu, Zhaopeng, et al.
Published: (2025)
Systematic Outliers in Large Language Models
by: An, Yongqi, et al.
Published: (2025)
by: An, Yongqi, et al.
Published: (2025)
UniVAD: A Training-free Unified Model for Few-shot Visual Anomaly Detection
by: Gu, Zhaopeng, et al.
Published: (2024)
by: Gu, Zhaopeng, et al.
Published: (2024)
MathPhys-Guided Coarse-to-Fine Anomaly Synthesis with SQE-Driven Bi-Level Optimization for Anomaly Detection
by: Qian, Long, et al.
Published: (2025)
by: Qian, Long, et al.
Published: (2025)
Challenges and Responses in the Practice of Large Language Models
by: Zhu, Hongyin
Published: (2024)
by: Zhu, Hongyin
Published: (2024)
Architectural Foundations for the Large Language Model Infrastructures
by: Zhu, Hongyin
Published: (2024)
by: Zhu, Hongyin
Published: (2024)
Information-Theoretic Constraints for Continual Vision-Language-Action Alignment
by: Zhao, Libang, et al.
Published: (2026)
by: Zhao, Libang, et al.
Published: (2026)
SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
by: Fei, Senyu, et al.
Published: (2025)
by: Fei, Senyu, et al.
Published: (2025)
Harnessing Vision-Language Pretrained Models with Temporal-Aware Adaptation for Referring Video Object Segmentation
by: Zhou, Zikun, et al.
Published: (2024)
by: Zhou, Zikun, et al.
Published: (2024)
Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection
by: Qian, Long, et al.
Published: (2025)
by: Qian, Long, et al.
Published: (2025)
Neural-Driven Image Editing
by: Zhou, Pengfei, et al.
Published: (2025)
by: Zhou, Pengfei, et al.
Published: (2025)
SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language Models
by: He, Jinghan, et al.
Published: (2024)
by: He, Jinghan, et al.
Published: (2024)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
by: Wang, Yabing, et al.
Published: (2024)
by: Wang, Yabing, et al.
Published: (2024)
Personality Editing for Language Models through Adjusting Self-Referential Queries
by: Hwang, Seojin, et al.
Published: (2025)
by: Hwang, Seojin, et al.
Published: (2025)
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection
by: Wang, Shuyu, et al.
Published: (2025)
by: Wang, Shuyu, et al.
Published: (2025)
Harmonizing Human Insights and AI Precision: Hand in Hand for Advancing Knowledge Graph Task
by: Wang, Shurong, et al.
Published: (2024)
by: Wang, Shurong, et al.
Published: (2024)
Grounding Language in Multi-Perspective Referential Communication
by: Tang, Zineng, et al.
Published: (2024)
by: Tang, Zineng, et al.
Published: (2024)
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
1st Place Solution for MOSE Track in CVPR 2024 PVUW Workshop: Complex Video Object Segmentation
by: Miao, Deshui, et al.
Published: (2024)
by: Miao, Deshui, et al.
Published: (2024)
Open-Vocabulary Camouflaged Object Segmentation with Cascaded Vision Language Models
by: Zhao, Kai, et al.
Published: (2025)
by: Zhao, Kai, et al.
Published: (2025)
StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion
by: Tao, Ming, et al.
Published: (2024)
by: Tao, Ming, et al.
Published: (2024)
FOCUS: Towards Universal Foreground Segmentation
by: You, Zuyao, et al.
Published: (2025)
by: You, Zuyao, et al.
Published: (2025)
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
by: Wang, Langyu, et al.
Published: (2024)
by: Wang, Langyu, et al.
Published: (2024)
Learning Spatial-Semantic Features for Robust Video Object Segmentation
by: Li, Xin, et al.
Published: (2024)
by: Li, Xin, et al.
Published: (2024)
Similar Items
-
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
by: Zhan, Yufei, et al.
Published: (2025) -
Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring
by: Zhan, Yufei, et al.
Published: (2024) -
GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
by: Zheng, Shurong, et al.
Published: (2026) -
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
by: Yang, Fan, et al.
Published: (2026) -
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2024)