A Generative Approach for Wikipedia-Scale Visual Entity Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Caron, Mathilde, Iscen, Ahmet, Fathi, Alireza, Schmid, Cordelia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
Retrieval-Enhanced Contrastive Vision-Text Models
von: Iscen, Ahmet, et al.
Veröffentlicht: (2023)
von: Iscen, Ahmet, et al.
Veröffentlicht: (2023)
Memory-Modular Classification: Learning to Generalize with Memory Replacement
von: Kang, Dahyun, et al.
Veröffentlicht: (2025)
von: Kang, Dahyun, et al.
Veröffentlicht: (2025)
SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code
von: Hu, Ziniu, et al.
Veröffentlicht: (2024)
von: Hu, Ziniu, et al.
Veröffentlicht: (2024)
CAViAR: Critic-Augmented Video Agentic Reasoning
von: Menon, Sachit, et al.
Veröffentlicht: (2025)
von: Menon, Sachit, et al.
Veröffentlicht: (2025)
Visual Lexicon: Rich Image Features in Language Space
von: Wang, XuDong, et al.
Veröffentlicht: (2024)
von: Wang, XuDong, et al.
Veröffentlicht: (2024)
Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
von: Arnab, Anurag, et al.
Veröffentlicht: (2025)
von: Arnab, Anurag, et al.
Veröffentlicht: (2025)
RECODE: Reasoning Through Code Generation for Visual Question Answering
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
Continual Learning in Vision-Language Models via Aligned Model Merging
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement
von: Huang, Ian, et al.
Veröffentlicht: (2025)
von: Huang, Ian, et al.
Veröffentlicht: (2025)
BrickNet: Graph-Backed Generative Brick Assembly
von: Kulits, Peter, et al.
Veröffentlicht: (2026)
von: Kulits, Peter, et al.
Veröffentlicht: (2026)
Grounded Video Caption Generation
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2024)
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2024)
VoCap: Video Object Captioning and Segmentation from Any Prompt
von: Uijlings, Jasper, et al.
Veröffentlicht: (2025)
von: Uijlings, Jasper, et al.
Veröffentlicht: (2025)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
von: Khan, Zeeshan, et al.
Veröffentlicht: (2025)
von: Khan, Zeeshan, et al.
Veröffentlicht: (2025)
Large-scale Pre-training for Grounded Video Caption Generation
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2025)
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2025)
Language-Guided Image Tokenization for Generation
von: Zha, Kaiwen, et al.
Veröffentlicht: (2024)
von: Zha, Kaiwen, et al.
Veröffentlicht: (2024)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
von: Chen, Shizhe, et al.
Veröffentlicht: (2026)
von: Chen, Shizhe, et al.
Veröffentlicht: (2026)
SUGAR: Pre-training 3D Visual Representations for Robotics
von: Chen, Shizhe, et al.
Veröffentlicht: (2024)
von: Chen, Shizhe, et al.
Veröffentlicht: (2024)
MetricNet: Recovering Metric Scale in Generative Navigation Policies
von: Nayak, Abhijeet, et al.
Veröffentlicht: (2025)
von: Nayak, Abhijeet, et al.
Veröffentlicht: (2025)
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
von: Pacaud, Paul, et al.
Veröffentlicht: (2025)
von: Pacaud, Paul, et al.
Veröffentlicht: (2025)
Learning text-to-video retrieval from image captioning
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
von: Bousselham, Walid, et al.
Veröffentlicht: (2025)
von: Bousselham, Walid, et al.
Veröffentlicht: (2025)
AMES: Asymmetric and Memory-Efficient Similarity Estimation for Instance-level Retrieval
von: Suma, Pavel, et al.
Veröffentlicht: (2024)
von: Suma, Pavel, et al.
Veröffentlicht: (2024)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
von: Garcia, Ricardo, et al.
Veröffentlicht: (2024)
von: Garcia, Ricardo, et al.
Veröffentlicht: (2024)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
von: Kim, Jae Myung, et al.
Veröffentlicht: (2025)
von: Kim, Jae Myung, et al.
Veröffentlicht: (2025)
Time-, Memory- and Parameter-Efficient Visual Adaptation
von: Mercea, Otniel-Bogdan, et al.
Veröffentlicht: (2024)
von: Mercea, Otniel-Bogdan, et al.
Veröffentlicht: (2024)
Dense Optical Tracking: Connecting the Dots
von: Moing, Guillaume Le, et al.
Veröffentlicht: (2023)
von: Moing, Guillaume Le, et al.
Veröffentlicht: (2023)
What Are You Doing? A Closer Look at Controllable Human Video Generation
von: Bugliarello, Emanuele, et al.
Veröffentlicht: (2025)
von: Bugliarello, Emanuele, et al.
Veröffentlicht: (2025)
MINERVA: Evaluating Complex Video Reasoning
von: Nagrani, Arsha, et al.
Veröffentlicht: (2025)
von: Nagrani, Arsha, et al.
Veröffentlicht: (2025)
RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks
von: Ghosh, Partha, et al.
Veröffentlicht: (2024)
von: Ghosh, Partha, et al.
Veröffentlicht: (2024)
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
von: Ventura, Lucas, et al.
Veröffentlicht: (2025)
von: Ventura, Lucas, et al.
Veröffentlicht: (2025)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2025)
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2025)
Dense Video Object Captioning from Disjoint Supervision
von: Zhou, Xingyi, et al.
Veröffentlicht: (2023)
von: Zhou, Xingyi, et al.
Veröffentlicht: (2023)
CoVR-2: Automatic Data Construction for Composed Video Retrieval
von: Ventura, Lucas, et al.
Veröffentlicht: (2023)
von: Ventura, Lucas, et al.
Veröffentlicht: (2023)
Grounding Language Models for Visual Entity Recognition
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
Learning Correlation Structures for Vision Transformers
von: Kim, Manjin, et al.
Veröffentlicht: (2024)
von: Kim, Manjin, et al.
Veröffentlicht: (2024)
HORT: Monocular Hand-held Objects Reconstruction with Transformers
von: Chen, Zerui, et al.
Veröffentlicht: (2025)
von: Chen, Zerui, et al.
Veröffentlicht: (2025)
Online 3D Scene Reconstruction Using Neural Object Priors
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition
von: Ning, Shan, et al.
Veröffentlicht: (2026)
von: Ning, Shan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
von: Caron, Mathilde, et al.
Veröffentlicht: (2024) -
Retrieval-Enhanced Contrastive Vision-Text Models
von: Iscen, Ahmet, et al.
Veröffentlicht: (2023) -
Memory-Modular Classification: Learning to Generalize with Memory Replacement
von: Kang, Dahyun, et al.
Veröffentlicht: (2025) -
SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code
von: Hu, Ziniu, et al.
Veröffentlicht: (2024) -
CAViAR: Critic-Augmented Video Agentic Reasoning
von: Menon, Sachit, et al.
Veröffentlicht: (2025)