LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Yuqian, Zhang, Wenqiao, Lin, Juekai, Zhong, Yu, Gao, Mingjian, Yu, Binhe, Cao, Yunqi, Li, Wentong, Zhuang, Yueting, Ooi, Beng Chin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
by: Yuan, Yuqian, et al.
Published: (2025)
by: Yuan, Yuqian, et al.
Published: (2025)
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
by: Gao, Mingjian, et al.
Published: (2026)
by: Gao, Mingjian, et al.
Published: (2026)
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
by: Lin, Tianwei, et al.
Published: (2025)
by: Lin, Tianwei, et al.
Published: (2025)
AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization
by: Yu, Binhe, et al.
Published: (2025)
by: Yu, Binhe, et al.
Published: (2025)
Unified Personalized Understanding, Generating and Editing
by: Zhong, Yu, et al.
Published: (2026)
by: Zhong, Yu, et al.
Published: (2026)
InstructSAM: Segment Any Instance with Any Instructions
by: Yuan, Yuqian, et al.
Published: (2026)
by: Yuan, Yuqian, et al.
Published: (2026)
HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding
by: Xie, Yihan, et al.
Published: (2025)
by: Xie, Yihan, et al.
Published: (2025)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
by: Yuan, Yuqian, et al.
Published: (2024)
by: Yuan, Yuqian, et al.
Published: (2024)
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
by: Yuan, Yuqian, et al.
Published: (2025)
by: Yuan, Yuqian, et al.
Published: (2025)
METER: A Dynamic Concept Adaptation Framework for Online Anomaly Detection
by: Zhu, Jiaqi, et al.
Published: (2023)
by: Zhu, Jiaqi, et al.
Published: (2023)
EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model
by: Li, Sijing, et al.
Published: (2025)
by: Li, Sijing, et al.
Published: (2025)
Catching Every Ripple: Enhanced Anomaly Awareness via Dynamic Concept Adaptation
by: Zhu, Jiaqi, et al.
Published: (2026)
by: Zhu, Jiaqi, et al.
Published: (2026)
OmniCT: Towards a Unified Slice-Volume LVLM for Comprehensive CT Analysis
by: Lin, Tianwei, et al.
Published: (2026)
by: Lin, Tianwei, et al.
Published: (2026)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
by: Wang, Wei, et al.
Published: (2026)
by: Wang, Wei, et al.
Published: (2026)
CLEAR-Mamba:Towards Accurate, Adaptive and Trustworthy Multi-Sequence Ophthalmic Angiography Classification
by: Wang, Zhuonan, et al.
Published: (2026)
by: Wang, Zhuonan, et al.
Published: (2026)
Toward Robust Signed Graph Learning through Joint Input-Target Denoising
by: Wu, Junran, et al.
Published: (2025)
by: Wu, Junran, et al.
Published: (2025)
CtrlCoT: Dual-Granularity Chain-of-Thought Compression for Controllable Reasoning
by: Fan, Zhenxuan, et al.
Published: (2026)
by: Fan, Zhenxuan, et al.
Published: (2026)
Do LLMs Understand Visual Anomalies? Uncovering LLM's Capabilities in Zero-shot Anomaly Detection
by: Zhu, Jiaqi, et al.
Published: (2024)
by: Zhu, Jiaqi, et al.
Published: (2024)
MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation
by: Zheng, Haoyu, et al.
Published: (2024)
by: Zheng, Haoyu, et al.
Published: (2024)
IDEAL: Leveraging Infinite and Dynamic Characterizations of Large Language Models for Query-focused Summarization
by: Cao, Jie, et al.
Published: (2024)
by: Cao, Jie, et al.
Published: (2024)
Enhancing Post-Training Quantization via Future Activation Awareness
by: Lv, Zheqi, et al.
Published: (2026)
by: Lv, Zheqi, et al.
Published: (2026)
Distribution-aware Online Continual Learning for Urban Spatio-Temporal Forecasting
by: Wang, Chengxin, et al.
Published: (2024)
by: Wang, Chengxin, et al.
Published: (2024)
Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives
by: Chen, Gang, et al.
Published: (2025)
by: Chen, Gang, et al.
Published: (2025)
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
by: Fu, Yuqian, et al.
Published: (2024)
by: Fu, Yuqian, et al.
Published: (2024)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
by: Fan, Zhaoyu, et al.
Published: (2025)
by: Fan, Zhaoyu, et al.
Published: (2025)
Fast Thinking for Large Language Models
by: Zheng, Haoyu, et al.
Published: (2025)
by: Zheng, Haoyu, et al.
Published: (2025)
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
by: Zheng, Haoyu, et al.
Published: (2025)
by: Zheng, Haoyu, et al.
Published: (2025)
Anytime Neural Architecture Search on Tabular Data
by: Xing, Naili, et al.
Published: (2024)
by: Xing, Naili, et al.
Published: (2024)
ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos
by: Guo, Wenliang, et al.
Published: (2025)
by: Guo, Wenliang, et al.
Published: (2025)
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
by: Xu, Zitong, et al.
Published: (2025)
by: Xu, Zitong, et al.
Published: (2025)
FAVOR: Efficient Filter-Agnostic Vector ANNS Based on Selectivity-Aware Exclusion Distances
by: Song, Junjie, et al.
Published: (2026)
by: Song, Junjie, et al.
Published: (2026)
Bridging Local Details and Global Context in Text-Attributed Graphs
by: Wang, Yaoke, et al.
Published: (2024)
by: Wang, Yaoke, et al.
Published: (2024)
DuetRAG: Collaborative Retrieval-Augmented Generation
by: Jiao, Dian, et al.
Published: (2024)
by: Jiao, Dian, et al.
Published: (2024)
CohortNet: Empowering Cohort Discovery for Interpretable Healthcare Analytics
by: Cai, Qingpeng, et al.
Published: (2024)
by: Cai, Qingpeng, et al.
Published: (2024)
IDPro: Flexible Interactive Video Object Segmentation by ID-queried Concurrent Propagation
by: Li, Kexin, et al.
Published: (2024)
by: Li, Kexin, et al.
Published: (2024)
MOCHA: Discovering Multi-Order Dynamic Causality in Temporal Point Processes
by: Cao, Yunyang, et al.
Published: (2025)
by: Cao, Yunyang, et al.
Published: (2025)
A Review on In Situ Microscopic Understandings of Dendritic Zinc Growth in Aqueous Zinc Ion Batteries
by: Yun Li, et al.
Published: (2026)
by: Yun Li, et al.
Published: (2026)
Object-Centric Diffusion for Efficient Video Editing
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
MAU-GPT: Enhancing Multi-type Industrial Anomaly Understanding via Anomaly-aware and Generalist Experts Adaptation
by: Wang, Zhuonan, et al.
Published: (2026)
by: Wang, Zhuonan, et al.
Published: (2026)
DUET: A Tuning-Free Device-Cloud Collaborative Parameters Generation Framework for Efficient Device Model Generalization
by: Lv, Zheqi, et al.
Published: (2022)
by: Lv, Zheqi, et al.
Published: (2022)
Similar Items
-
PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
by: Yuan, Yuqian, et al.
Published: (2025) -
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
by: Gao, Mingjian, et al.
Published: (2026) -
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
by: Lin, Tianwei, et al.
Published: (2025) -
AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization
by: Yu, Binhe, et al.
Published: (2025) -
Unified Personalized Understanding, Generating and Editing
by: Zhong, Yu, et al.
Published: (2026)