GViT: Representing Images as Gaussians for Visual Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Hernandez, Jefferson, He, Ruozhen, Balakrishnan, Guha, Berg, Alexander C., Ordonez, Vicente |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning from Synthetic Data for Visual Grounding
by: He, Ruozhen, et al.
Published: (2024)
by: He, Ruozhen, et al.
Published: (2024)
Agentic Discovery with Active Hypothesis Exploration for Visual Recognition
by: Koo, Jaywon, et al.
Published: (2026)
by: Koo, Jaywon, et al.
Published: (2026)
Fairness and Bias Mitigation in Computer Vision: A Survey
by: Dehdashtian, Sepehr, et al.
Published: (2024)
by: Dehdashtian, Sepehr, et al.
Published: (2024)
SplineCam: Exact Visualization and Characterization of Deep Network Geometry and Decision Boundaries
by: Humayun, Ahmed Imtiaz, et al.
Published: (2023)
by: Humayun, Ahmed Imtiaz, et al.
Published: (2023)
ElasticDiffusion: Training-free Arbitrary Size Image Generation through Global-Local Content Separation
by: Haji-Ali, Moayed, et al.
Published: (2023)
by: Haji-Ali, Moayed, et al.
Published: (2023)
GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation
by: Liu, Yuhao, et al.
Published: (2026)
by: Liu, Yuhao, et al.
Published: (2026)
Generative Visual Instruction Tuning
by: Hernandez, Jefferson, et al.
Published: (2024)
by: Hernandez, Jefferson, et al.
Published: (2024)
NoiseShift: Resolution-Aware Noise Recalibration for Better Low-Resolution Image Generation
by: He, Ruozhen, et al.
Published: (2025)
by: He, Ruozhen, et al.
Published: (2025)
Beyond Referring Expressions: Scenario Comprehension Visual Grounding
by: He, Ruozhen, et al.
Published: (2026)
by: He, Ruozhen, et al.
Published: (2026)
Quantised Global Autoencoder: A Holistic Approach to Representing Visual Data
by: Elsner, Tim, et al.
Published: (2024)
by: Elsner, Tim, et al.
Published: (2024)
EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
by: He, Ruozhen, et al.
Published: (2026)
by: He, Ruozhen, et al.
Published: (2026)
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022)
by: Yang, Ziyan, et al.
Published: (2022)
ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders
by: Hernandez, Jefferson, et al.
Published: (2023)
by: Hernandez, Jefferson, et al.
Published: (2023)
Grounding Descriptions in Images informs Zero-Shot Visual Recognition
by: Halbe, Shaunak, et al.
Published: (2024)
by: Halbe, Shaunak, et al.
Published: (2024)
COMPASS: Robust Feature Conformal Prediction for Medical Segmentation Metrics
by: Cheung, Matt Y., et al.
Published: (2025)
by: Cheung, Matt Y., et al.
Published: (2025)
DRAGON: Drone and Ground Gaussian Splatting for 3D Building Reconstruction
by: Ham, Yujin, et al.
Published: (2024)
by: Ham, Yujin, et al.
Published: (2024)
Efficient Conformal Volumetry for Template-Based Segmentation
by: Cheung, Matt Y., et al.
Published: (2026)
by: Cheung, Matt Y., et al.
Published: (2026)
Representing Online Handwriting for Recognition in Large Vision-Language Models
by: Fadeeva, Anastasiia, et al.
Published: (2024)
by: Fadeeva, Anastasiia, et al.
Published: (2024)
Metric-Guided Conformal Bounds for Probabilistic Image Reconstruction
by: Cheung, Matt Y, et al.
Published: (2024)
by: Cheung, Matt Y, et al.
Published: (2024)
Probabilistic Contrastive Learning for Long-Tailed Visual Recognition
by: Du, Chaoqun, et al.
Published: (2024)
by: Du, Chaoqun, et al.
Published: (2024)
Regressing Transformers for Data-efficient Visual Place Recognition
by: Leyva-Vallina, María, et al.
Published: (2024)
by: Leyva-Vallina, María, et al.
Published: (2024)
Fast Amortized Fitting of Scientific Signals Across Time and Ensembles via Transferable Neural Fields
by: Zorek, Sophia, et al.
Published: (2026)
by: Zorek, Sophia, et al.
Published: (2026)
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
by: Cheng, Jintao, et al.
Published: (2025)
by: Cheng, Jintao, et al.
Published: (2025)
Dendritic Convolution for Noise Image Recognition
by: Xue, Jiarui, et al.
Published: (2025)
by: Xue, Jiarui, et al.
Published: (2025)
PlainMamba: Improving Non-Hierarchical Mamba in Visual Recognition
by: Yang, Chenhongyi, et al.
Published: (2024)
by: Yang, Chenhongyi, et al.
Published: (2024)
Fit Pixels, Get Labels: Meta-learned Implicit Networks for Image Segmentation
by: Vyas, Kushal, et al.
Published: (2025)
by: Vyas, Kushal, et al.
Published: (2025)
Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
by: Liang, Zichen, et al.
Published: (2025)
by: Liang, Zichen, et al.
Published: (2025)
Surely Large Multimodal Models (Don't) Excel in Visual Species Recognition?
by: Liu, Tian, et al.
Published: (2025)
by: Liu, Tian, et al.
Published: (2025)
Training-Free Vector Quantization via Gaussian VAEs
by: Xu, Tongda, et al.
Published: (2025)
by: Xu, Tongda, et al.
Published: (2025)
Deep Learning in Image Classification: Evaluating VGG19's Performance on Complex Visual Data
by: He, Weijie, et al.
Published: (2024)
by: He, Weijie, et al.
Published: (2024)
Classes Are Not Equal: An Empirical Study on Image Recognition Fairness
by: Cui, Jiequan, et al.
Published: (2024)
by: Cui, Jiequan, et al.
Published: (2024)
Enhancing Zero-Shot Image Recognition in Vision-Language Models through Human-like Concept Guidance
by: Liu, Hui, et al.
Published: (2025)
by: Liu, Hui, et al.
Published: (2025)
Enhancing Tea Leaf Disease Recognition with Attention Mechanisms and Grad-CAM Visualization
by: Shikdar, Omar Faruq, et al.
Published: (2025)
by: Shikdar, Omar Faruq, et al.
Published: (2025)
Fuzzy-aware Loss for Source-free Domain Adaptation in Visual Emotion Recognition
by: Zheng, Ying, et al.
Published: (2025)
by: Zheng, Ying, et al.
Published: (2025)
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
Research on Image Recognition Technology Based on Multimodal Deep Learning
by: Wang, Jinyin, et al.
Published: (2024)
by: Wang, Jinyin, et al.
Published: (2024)
Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance
by: Koo, Jaywon, et al.
Published: (2025)
by: Koo, Jaywon, et al.
Published: (2025)
Efficient License Plate Recognition in Videos Using Visual Rhythm and Accumulative Line Analysis
by: Ribeiro, Victor Nascimento, et al.
Published: (2025)
by: Ribeiro, Victor Nascimento, et al.
Published: (2025)
Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
Advancing Vehicle Plate Recognition: Multitasking Visual Language Models with VehiclePaliGemma
by: AlDahoul, Nouar, et al.
Published: (2024)
by: AlDahoul, Nouar, et al.
Published: (2024)
Similar Items
-
Learning from Synthetic Data for Visual Grounding
by: He, Ruozhen, et al.
Published: (2024) -
Agentic Discovery with Active Hypothesis Exploration for Visual Recognition
by: Koo, Jaywon, et al.
Published: (2026) -
Fairness and Bias Mitigation in Computer Vision: A Survey
by: Dehdashtian, Sepehr, et al.
Published: (2024) -
SplineCam: Exact Visualization and Characterization of Deep Network Geometry and Decision Boundaries
by: Humayun, Ahmed Imtiaz, et al.
Published: (2023) -
ElasticDiffusion: Training-free Arbitrary Size Image Generation through Global-Local Content Separation
by: Haji-Ali, Moayed, et al.
Published: (2023)