Grounding Descriptions in Images informs Zero-Shot Visual Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Halbe, Shaunak, Tian, Junjiao, Joseph, K J, Smith, James Seale, Stevo, Katherine, Balasubramanian, Vineeth N, Kira, Zsolt |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Continual Adaptation of Vision Transformers for Federated Learning
von: Halbe, Shaunak, et al.
Veröffentlicht: (2023)
von: Halbe, Shaunak, et al.
Veröffentlicht: (2023)
Adaptive Memory Replay for Continual Learning
von: Smith, James Seale, et al.
Veröffentlicht: (2024)
von: Smith, James Seale, et al.
Veröffentlicht: (2024)
The Geometry of Robustness: Optimizing Loss Landscape Curvature and Feature Manifold Alignment for Robust Finetuning of Vision-Language Models
von: Chopra, Shivang, et al.
Veröffentlicht: (2026)
von: Chopra, Shivang, et al.
Veröffentlicht: (2026)
Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion
von: Tian, Junjiao, et al.
Veröffentlicht: (2023)
von: Tian, Junjiao, et al.
Veröffentlicht: (2023)
Rethinking Weight Decay for Robust Fine-Tuning of Foundation Models
von: Tian, Junjiao, et al.
Veröffentlicht: (2024)
von: Tian, Junjiao, et al.
Veröffentlicht: (2024)
Directional Gradient Projection for Robust Fine-Tuning of Foundation Models
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
$\oslash$ Source Models Leak What They Shouldn't $\nrightarrow$: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization
von: Devalapally, Arnav, et al.
Veröffentlicht: (2026)
von: Devalapally, Arnav, et al.
Veröffentlicht: (2026)
Continual Diffusion with STAMINA: STack-And-Mask INcremental Adapters
von: Smith, James Seale, et al.
Veröffentlicht: (2023)
von: Smith, James Seale, et al.
Veröffentlicht: (2023)
SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation
von: Shukla, Tripti, et al.
Veröffentlicht: (2026)
von: Shukla, Tripti, et al.
Veröffentlicht: (2026)
Continual Diffusion: Continual Customization of Text-to-Image Diffusion with C-LoRA
von: Smith, James Seale, et al.
Veröffentlicht: (2023)
von: Smith, James Seale, et al.
Veröffentlicht: (2023)
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2025)
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2025)
Source-Free Domain Adaptation by Optimizing Batch-Wise Cosine Similarity
von: Pathak, Harsharaj, et al.
Veröffentlicht: (2026)
von: Pathak, Harsharaj, et al.
Veröffentlicht: (2026)
VeRVE: Versatile Retrieval for Videos via Unified Embeddings
von: Halbe, Shaunak, et al.
Veröffentlicht: (2026)
von: Halbe, Shaunak, et al.
Veröffentlicht: (2026)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
Distributed Zero-Shot Learning for Visual Recognition
von: Chen, Zhi, et al.
Veröffentlicht: (2025)
von: Chen, Zhi, et al.
Veröffentlicht: (2025)
Z3D: Zero-Shot 3D Visual Grounding from Images
von: Drozdov, Nikita, et al.
Veröffentlicht: (2026)
von: Drozdov, Nikita, et al.
Veröffentlicht: (2026)
Zero-Shot Aerial Object Detection with Visual Description Regularization
von: Zang, Zhengqing, et al.
Veröffentlicht: (2024)
von: Zang, Zhengqing, et al.
Veröffentlicht: (2024)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
von: Romero, David, et al.
Veröffentlicht: (2024)
von: Romero, David, et al.
Veröffentlicht: (2024)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
von: Li, Rong, et al.
Veröffentlicht: (2024)
von: Li, Rong, et al.
Veröffentlicht: (2024)
C2FDrone: Coarse-to-Fine Drone-to-Drone Detection using Vision Transformer Networks
von: Rebbapragada, Sairam VC, et al.
Veröffentlicht: (2024)
von: Rebbapragada, Sairam VC, et al.
Veröffentlicht: (2024)
Advancing Ante-Hoc Explainable Models through Generative Adversarial Networks
von: Garg, Tanmay, et al.
Veröffentlicht: (2024)
von: Garg, Tanmay, et al.
Veröffentlicht: (2024)
Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media
von: M, Megha Mariam K., et al.
Veröffentlicht: (2026)
von: M, Megha Mariam K., et al.
Veröffentlicht: (2026)
Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
von: Sinha, Rohit, et al.
Veröffentlicht: (2026)
von: Sinha, Rohit, et al.
Veröffentlicht: (2026)
LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation
von: Li, Jiachen, et al.
Veröffentlicht: (2025)
von: Li, Jiachen, et al.
Veröffentlicht: (2025)
Seeing Beyond Classes: Zero-Shot Grounded Situation Recognition via Language Explainer
von: Lei, Jiaming, et al.
Veröffentlicht: (2024)
von: Lei, Jiaming, et al.
Veröffentlicht: (2024)
Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
Solving Zero-Shot 3D Visual Grounding as Constraint Satisfaction Problems
von: Yuan, Qihao, et al.
Veröffentlicht: (2024)
von: Yuan, Qihao, et al.
Veröffentlicht: (2024)
Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class Imbalance
von: Santra, Sanchayan, et al.
Veröffentlicht: (2025)
von: Santra, Sanchayan, et al.
Veröffentlicht: (2025)
Continual Learning Improves Zero-Shot Action Recognition
von: Gowda, Shreyank N, et al.
Veröffentlicht: (2024)
von: Gowda, Shreyank N, et al.
Veröffentlicht: (2024)
Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval
von: Liao, Liwei, et al.
Veröffentlicht: (2025)
von: Liao, Liwei, et al.
Veröffentlicht: (2025)
Zero-Shot 3D Visual Grounding from Vision-Language Models
von: Li, Rong, et al.
Veröffentlicht: (2025)
von: Li, Rong, et al.
Veröffentlicht: (2025)
Zero-Shot Skeleton-based Action Recognition with Dual Visual-Text Alignment
von: Kuang, Jidong, et al.
Veröffentlicht: (2024)
von: Kuang, Jidong, et al.
Veröffentlicht: (2024)
Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation
von: He, Jingxuan, et al.
Veröffentlicht: (2026)
von: He, Jingxuan, et al.
Veröffentlicht: (2026)
RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images
von: Li, Ke, et al.
Veröffentlicht: (2025)
von: Li, Ke, et al.
Veröffentlicht: (2025)
Telling Stories for Common Sense Zero-Shot Action Recognition
von: Gowda, Shreyank N, et al.
Veröffentlicht: (2023)
von: Gowda, Shreyank N, et al.
Veröffentlicht: (2023)
Zero-Shot Underwater Gesture Recognition
von: Sarma, Sandipan, et al.
Veröffentlicht: (2024)
von: Sarma, Sandipan, et al.
Veröffentlicht: (2024)
Understanding Task Transfer in Vision-Language Models
von: Sachdeva, Bhuvan, et al.
Veröffentlicht: (2025)
von: Sachdeva, Bhuvan, et al.
Veröffentlicht: (2025)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
von: Xu, Runsen, et al.
Veröffentlicht: (2024)
von: Xu, Runsen, et al.
Veröffentlicht: (2024)
Sim2real Image Translation Enables Viewpoint-Robust Policies from Fixed-Camera Datasets
von: Coholich, Jeremiah, et al.
Veröffentlicht: (2026)
von: Coholich, Jeremiah, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Continual Adaptation of Vision Transformers for Federated Learning
von: Halbe, Shaunak, et al.
Veröffentlicht: (2023) -
Adaptive Memory Replay for Continual Learning
von: Smith, James Seale, et al.
Veröffentlicht: (2024) -
The Geometry of Robustness: Optimizing Loss Landscape Curvature and Feature Manifold Alignment for Robust Finetuning of Vision-Language Models
von: Chopra, Shivang, et al.
Veröffentlicht: (2026) -
Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion
von: Tian, Junjiao, et al.
Veröffentlicht: (2023) -
Rethinking Weight Decay for Robust Fine-Tuning of Foundation Models
von: Tian, Junjiao, et al.
Veröffentlicht: (2024)