Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Haotian, You, Haoxuan, Dufter, Philipp, Zhang, Bowen, Chen, Chen, Chen, Hong-You, Fu, Tsu-Jui, Wang, William Yang, Chang, Shih-Fu, Gan, Zhe, Yang, Yinfei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
by: You, Keen, et al.
Published: (2024)
by: You, Keen, et al.
Published: (2024)
UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing
by: Fu, Tsu-Jui, et al.
Published: (2025)
by: Fu, Tsu-Jui, et al.
Published: (2025)
GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
by: Qian, Yusu, et al.
Published: (2025)
by: Qian, Yusu, et al.
Published: (2025)
Guiding Instruction-based Image Editing via Multimodal Large Language Models
by: Fu, Tsu-Jui, et al.
Published: (2023)
by: Fu, Tsu-Jui, et al.
Published: (2023)
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
by: Li, Zhangheng, et al.
Published: (2024)
by: Li, Zhangheng, et al.
Published: (2024)
Contrastive Localized Language-Image Pre-Training
by: Chen, Hong-You, et al.
Published: (2024)
by: Chen, Hong-You, et al.
Published: (2024)
Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
by: Yang, Zhen, et al.
Published: (2025)
by: Yang, Zhen, et al.
Published: (2025)
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
by: Qian, Yusu, et al.
Published: (2024)
by: Qian, Yusu, et al.
Published: (2024)
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
by: Ye, Hanrong, et al.
Published: (2024)
by: Ye, Hanrong, et al.
Published: (2024)
Improve Vision Language Model Chain-of-thought Reasoning
by: Zhang, Ruohong, et al.
Published: (2024)
by: Zhang, Ruohong, et al.
Published: (2024)
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
by: Zhang, Haotian, et al.
Published: (2024)
by: Zhang, Haotian, et al.
Published: (2024)
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
by: Lai, Zhengfeng, et al.
Published: (2024)
by: Lai, Zhengfeng, et al.
Published: (2024)
Taming Outlier Tokens in Diffusion Transformers
by: Wu, Xiaoyu, et al.
Published: (2026)
by: Wu, Xiaoyu, et al.
Published: (2026)
DiT-Air: Revisiting the Efficiency of Diffusion Model Architecture Design in Text to Image Generation
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
by: Li, Yanghao, et al.
Published: (2025)
by: Li, Yanghao, et al.
Published: (2025)
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
by: Wang, Xinze, et al.
Published: (2025)
by: Wang, Xinze, et al.
Published: (2025)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
by: Lai, Zhengfeng, et al.
Published: (2023)
by: Lai, Zhengfeng, et al.
Published: (2023)
CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
by: Jaiswal, Ajay, et al.
Published: (2023)
by: Jaiswal, Ajay, et al.
Published: (2023)
UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding
by: Wang, Zhecan, et al.
Published: (2023)
by: Wang, Zhecan, et al.
Published: (2023)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
by: Lu, Yujie, et al.
Published: (2024)
by: Lu, Yujie, et al.
Published: (2024)
STIV: Scalable Text and Image Conditioned Video Generation
by: Lin, Zongyu, et al.
Published: (2024)
by: Lin, Zongyu, et al.
Published: (2024)
Compressive behavior and design of circular steel slag concrete‐filled stainless steel tube ( SSC ‐ FSST ) stub columns
by: You‐Fu Yang, et al.
Published: (2025)
by: You‐Fu Yang, et al.
Published: (2025)
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
by: You, Haoxuan, et al.
Published: (2023)
by: You, Haoxuan, et al.
Published: (2023)
T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback
by: Li, Jiachen, et al.
Published: (2024)
by: Li, Jiachen, et al.
Published: (2024)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
by: Xu, Mingze, et al.
Published: (2024)
by: Xu, Mingze, et al.
Published: (2024)
SSCR: Iterative Language-Based Image Editing via Self-Supervised Counterfactual Reasoning
by: Fu, Tsu-Jui, et al.
Published: (2020)
by: Fu, Tsu-Jui, et al.
Published: (2020)
BaseReward: A Strong Baseline for Multimodal Reward Model
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View
by: Schumann, Raphael, et al.
Published: (2023)
by: Schumann, Raphael, et al.
Published: (2023)
Augmented Lagrange method for optimal control problems of parabolic equation with state constraints
by: You, Weilong, et al.
Published: (2024)
by: You, Weilong, et al.
Published: (2024)
Pontryagin's Principle Based Algorithms for Optimal Control Problems of Parabolic Equation
by: You, Weilong, et al.
Published: (2025)
by: You, Weilong, et al.
Published: (2025)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
MOFI: Learning Image Representations from Noisy Entity Annotated Images
by: Wu, Wentao, et al.
Published: (2023)
by: Wu, Wentao, et al.
Published: (2023)
Improving Chinese Character Representation with Formation Tree
by: Hong, Yang, et al.
Published: (2024)
by: Hong, Yang, et al.
Published: (2024)
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
by: Liu, Junzhang, et al.
Published: (2024)
by: Liu, Junzhang, et al.
Published: (2024)
Counterfactual Vision-and-Language Navigation via Adversarial Path Sampling
by: Fu, Tsu-Jui, et al.
Published: (2019)
by: Fu, Tsu-Jui, et al.
Published: (2019)
STAT3 phosphorylation inhibitor Bt354 exhibits anti‐neoplastic activity in glioblastoma multiforme cells
by: Yi‐Chun Chiang, et al.
Published: (2024)
by: Yi‐Chun Chiang, et al.
Published: (2024)
SEM Electron‐Beam‐Induced Ultrathin Carbon Deposition Layer on Cu Substrate: Improved Dry Oxidation Protection Performance than CVD Single Layer Graphene
by: Panpan Feng, et al.
Published: (2024)
by: Panpan Feng, et al.
Published: (2024)
Information-Ferret.
by: Duchin, Douglas
Published: (1990)
by: Duchin, Douglas
Published: (1990)
Self‐Assembled Bilayers with Improved Solvent Resistance for Stable Inverted Perovskite Solar Cells
by: Ahmed I. A. Soliman, et al.
Published: (2025)
by: Ahmed I. A. Soliman, et al.
Published: (2025)
Similar Items
-
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
by: You, Keen, et al.
Published: (2024) -
UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing
by: Fu, Tsu-Jui, et al.
Published: (2025) -
GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
by: Qian, Yusu, et al.
Published: (2025) -
Guiding Instruction-based Image Editing via Multimodal Large Language Models
by: Fu, Tsu-Jui, et al.
Published: (2023) -
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
by: Li, Zhangheng, et al.
Published: (2024)