F-LMM: Grounding Frozen Large Multimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Size, Jin, Sheng, Zhang, Wenwei, Xu, Lumin, Liu, Wentao, Li, Wei, Loy, Chen Change |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
by: Wu, Size, et al.
Published: (2023)
by: Wu, Size, et al.
Published: (2023)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
by: Xu, Shilin, et al.
Published: (2023)
by: Xu, Shilin, et al.
Published: (2023)
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
OMG-Seg: Is One Model Good Enough For All Segmentation?
by: Li, Xiangtai, et al.
Published: (2024)
by: Li, Xiangtai, et al.
Published: (2024)
Controllable Human-centric Keyframe Interpolation with Generative Prior
by: Guo, Zujin, et al.
Published: (2025)
by: Guo, Zujin, et al.
Published: (2025)
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
by: Liao, Kang, et al.
Published: (2025)
by: Liao, Kang, et al.
Published: (2025)
Contextual Object Detection with Multimodal Large Language Models
by: Zang, Yuhang, et al.
Published: (2023)
by: Zang, Yuhang, et al.
Published: (2023)
Generalizable Implicit Motion Modeling for Video Frame Interpolation
by: Guo, Zujin, et al.
Published: (2024)
by: Guo, Zujin, et al.
Published: (2024)
KptLLM: Unveiling the Power of Large Language Model for Keypoint Comprehension
by: Yang, Jie, et al.
Published: (2024)
by: Yang, Jie, et al.
Published: (2024)
KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
by: Yang, Jie, et al.
Published: (2025)
by: Yang, Jie, et al.
Published: (2025)
Control Color: Multimodal Diffusion-based Interactive Image Colorization
by: Liang, Zhexin, et al.
Published: (2024)
by: Liang, Zhexin, et al.
Published: (2024)
MOWA: Multiple-in-One Image Warping Model
by: Liao, Kang, et al.
Published: (2024)
by: Liao, Kang, et al.
Published: (2024)
Towards Language-Driven Video Inpainting via Multimodal Large Language Models
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
Enhanced Generative Structure Prior for Chinese Text Image Super-resolution
by: Li, Xiaoming, et al.
Published: (2025)
by: Li, Xiaoming, et al.
Published: (2025)
Kalman-Inspired Feature Propagation for Video Face Super-Resolution
by: Feng, Ruicheng, et al.
Published: (2024)
by: Feng, Ruicheng, et al.
Published: (2024)
AITTI: Learning Adaptive Inclusive Token for Text-to-Image Generation
by: Hou, Xinyu, et al.
Published: (2024)
by: Hou, Xinyu, et al.
Published: (2024)
DoF-Gaussian: Controllable Depth-of-Field for 3D Gaussian Splatting
by: Shen, Liao, et al.
Published: (2025)
by: Shen, Liao, et al.
Published: (2025)
MVIP-NeRF: Multi-view 3D Inpainting on NeRF Scenes via Diffusion Prior
by: Chen, Honghua, et al.
Published: (2024)
by: Chen, Honghua, et al.
Published: (2024)
UniFS: Universal Few-shot Instance Perception with Point Representations
by: Jin, Sheng, et al.
Published: (2024)
by: Jin, Sheng, et al.
Published: (2024)
DifFace: Blind Face Restoration with Diffused Error Contraction
by: Yue, Zongsheng, et al.
Published: (2022)
by: Yue, Zongsheng, et al.
Published: (2022)
AviationLMM: A Large Multimodal Foundation Model for Civil Aviation
by: Li, Wenbin, et al.
Published: (2026)
by: Li, Wenbin, et al.
Published: (2026)
LMM-Det: Make Large Multimodal Models Excel in Object Detection
by: Li, Jincheng, et al.
Published: (2025)
by: Li, Jincheng, et al.
Published: (2025)
LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models
by: Ge, Qihang, et al.
Published: (2024)
by: Ge, Qihang, et al.
Published: (2024)
Efficient Diffusion Model for Image Restoration by Residual Shifting
by: Yue, Zongsheng, et al.
Published: (2024)
by: Yue, Zongsheng, et al.
Published: (2024)
Transformer-Based Visual Segmentation: A Survey
by: Li, Xiangtai, et al.
Published: (2023)
by: Li, Xiangtai, et al.
Published: (2023)
Learning 3D Garment Animation from Trajectories of A Piece of Cloth
by: Shao, Yidi, et al.
Published: (2025)
by: Shao, Yidi, et al.
Published: (2025)
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation
by: Yang, Shuai, et al.
Published: (2024)
by: Yang, Shuai, et al.
Published: (2024)
Omegance: A Single Parameter for Various Granularities in Diffusion-Based Synthesis
by: Hou, Xinyu, et al.
Published: (2024)
by: Hou, Xinyu, et al.
Published: (2024)
Trans-Adapter: A Plug-and-Play Framework for Transparent Image Inpainting
by: Dai, Yuekun, et al.
Published: (2025)
by: Dai, Yuekun, et al.
Published: (2025)
EdgeSAM: Prompt-In-the-Loop Distillation for SAM
by: Zhou, Chong, et al.
Published: (2023)
by: Zhou, Chong, et al.
Published: (2023)
GKGNet: Group K-Nearest Neighbor based Graph Convolutional Network for Multi-Label Image Recognition
by: Yao, Ruijie, et al.
Published: (2023)
by: Yao, Ruijie, et al.
Published: (2023)
Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
by: Thawakar, Omkar, et al.
Published: (2025)
by: Thawakar, Omkar, et al.
Published: (2025)
MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation
by: Xie, Jiahao, et al.
Published: (2023)
by: Xie, Jiahao, et al.
Published: (2023)
LMM-PCQA: Assisting Point Cloud Quality Assessment with LMM
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
Ultra-High Resolution Segmentation via Boundary-Enhanced Patch-Merging Transformer
by: Sun, Haopeng, et al.
Published: (2024)
by: Sun, Haopeng, et al.
Published: (2024)
FrozenSeg: Harmonizing Frozen Foundation Models for Open-Vocabulary Segmentation
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
Generative Photographic Control for Scene-Consistent Video Cinematic Editing
by: Sun, Huiqiang, et al.
Published: (2025)
by: Sun, Huiqiang, et al.
Published: (2025)
ObjCtrl-2.5D: Training-free Object Control with Camera Poses
by: Wang, Zhouxia, et al.
Published: (2024)
by: Wang, Zhouxia, et al.
Published: (2024)
Similar Items
-
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
by: Wu, Size, et al.
Published: (2023) -
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025) -
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
by: Xu, Shilin, et al.
Published: (2023) -
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025) -
OMG-Seg: Is One Model Good Enough For All Segmentation?
by: Li, Xiangtai, et al.
Published: (2024)