Visual Lexicon: Rich Image Features in Language Space
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, XuDong, Zhou, Xingyi, Fathi, Alireza, Darrell, Trevor, Schmid, Cordelia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
by: Yu, Junwei, et al.
Published: (2025)
by: Yu, Junwei, et al.
Published: (2025)
Reconstruction Alignment Improves Unified Multimodal Models
by: Xie, Ji, et al.
Published: (2025)
by: Xie, Ji, et al.
Published: (2025)
Language-Guided Image Tokenization for Generation
by: Zha, Kaiwen, et al.
Published: (2024)
by: Zha, Kaiwen, et al.
Published: (2024)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
RECODE: Reasoning Through Code Generation for Visual Question Answering
by: Shen, Junhong, et al.
Published: (2025)
by: Shen, Junhong, et al.
Published: (2025)
Segment Anything without Supervision
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
TULIP: Towards Unified Language-Image Pretraining
by: Tang, Zineng, et al.
Published: (2025)
by: Tang, Zineng, et al.
Published: (2025)
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
by: Fiastre, Gabriel, et al.
Published: (2025)
by: Fiastre, Gabriel, et al.
Published: (2025)
SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code
by: Hu, Ziniu, et al.
Published: (2024)
by: Hu, Ziniu, et al.
Published: (2024)
Analyzing The Language of Visual Tokens
by: Chan, David M., et al.
Published: (2024)
by: Chan, David M., et al.
Published: (2024)
What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models
by: Farid, Karim, et al.
Published: (2025)
by: Farid, Karim, et al.
Published: (2025)
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
Continual Learning in Vision-Language Models via Aligned Model Merging
by: Sokar, Ghada, et al.
Published: (2025)
by: Sokar, Ghada, et al.
Published: (2025)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
by: Min, Juhong, et al.
Published: (2024)
by: Min, Juhong, et al.
Published: (2024)
Hidden in plain sight: VLMs overlook their visual representations
by: Fu, Stephanie, et al.
Published: (2025)
by: Fu, Stephanie, et al.
Published: (2025)
PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image Editor
by: Goel, Vidit, et al.
Published: (2023)
by: Goel, Vidit, et al.
Published: (2023)
Neptune: The Long Orbit to Benchmarking Long Video Understanding
by: Nagrani, Arsha, et al.
Published: (2024)
by: Nagrani, Arsha, et al.
Published: (2024)
Visually Prompted Benchmarks Are Surprisingly Fragile
by: Feng, Haiwen, et al.
Published: (2025)
by: Feng, Haiwen, et al.
Published: (2025)
Navigation World Models
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
REOrdering Patches Improves Vision Models
by: Kutscher, Declan, et al.
Published: (2025)
by: Kutscher, Declan, et al.
Published: (2025)
Lifting Embodied World Models for Planning and Control
by: Wang, Alex N., et al.
Published: (2026)
by: Wang, Alex N., et al.
Published: (2026)
Precipitation Nowcasting Using Diffusion Transformer with Causal Attention
by: Li, ChaoRong, et al.
Published: (2024)
by: Li, ChaoRong, et al.
Published: (2024)
SegLLM: Multi-round Reasoning Segmentation
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
Shape-Guided Diffusion with Inside-Outside Attention
by: Park, Dong Huk, et al.
Published: (2022)
by: Park, Dong Huk, et al.
Published: (2022)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
by: Mitra, Chancharik, et al.
Published: (2023)
by: Mitra, Chancharik, et al.
Published: (2023)
Neural Metamorphosis
by: Yang, Xingyi, et al.
Published: (2024)
by: Yang, Xingyi, et al.
Published: (2024)
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
by: Caron, Mathilde, et al.
Published: (2024)
by: Caron, Mathilde, et al.
Published: (2024)
Mixtures of Unsupervised Lexicon Classification
by: Wiriyathammabhum, Peratham
Published: (2024)
by: Wiriyathammabhum, Peratham
Published: (2024)
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
by: Caron, Mathilde, et al.
Published: (2024)
by: Caron, Mathilde, et al.
Published: (2024)
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
by: Ahmadian, Mona, et al.
Published: (2024)
by: Ahmadian, Mona, et al.
Published: (2024)
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
by: Van Landeghem, Jordy, et al.
Published: (2024)
by: Van Landeghem, Jordy, et al.
Published: (2024)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
by: Lee, Heekyung, et al.
Published: (2025)
by: Lee, Heekyung, et al.
Published: (2025)
Automatic Image Annotation for Mapped Features Detection
by: Noizet, Maxime, et al.
Published: (2024)
by: Noizet, Maxime, et al.
Published: (2024)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
by: Shang, Chuyi, et al.
Published: (2024)
by: Shang, Chuyi, et al.
Published: (2024)
Video Action Differencing
by: Burgess, James, et al.
Published: (2025)
by: Burgess, James, et al.
Published: (2025)
ISLR101: an Iranian Word-Level Sign Language Recognition Dataset
by: Ranjbar, Hossein, et al.
Published: (2025)
by: Ranjbar, Hossein, et al.
Published: (2025)
Manipulating Feature Visualizations with Gradient Slingshots
by: Bareeva, Dilyara, et al.
Published: (2024)
by: Bareeva, Dilyara, et al.
Published: (2024)
Enhancing Fine-Grained Visual Recognition in the Low-Data Regime Through Feature Magnitude Regularization
by: Chapman, Avraham, et al.
Published: (2024)
by: Chapman, Avraham, et al.
Published: (2024)
On-the-fly Repulsion in the Contextual Space for Rich Diversity in Diffusion Transformers
by: Dahary, Omer, et al.
Published: (2026)
by: Dahary, Omer, et al.
Published: (2026)
Stochastic positional embeddings improve masked image modeling
by: Bar, Amir, et al.
Published: (2023)
by: Bar, Amir, et al.
Published: (2023)
Similar Items
-
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
by: Yu, Junwei, et al.
Published: (2025) -
Reconstruction Alignment Improves Unified Multimodal Models
by: Xie, Ji, et al.
Published: (2025) -
Language-Guided Image Tokenization for Generation
by: Zha, Kaiwen, et al.
Published: (2024) -
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025) -
RECODE: Reasoning Through Code Generation for Visual Question Answering
by: Shen, Junhong, et al.
Published: (2025)