Harnessing small projectors and multiple views for efficient vision pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Agrawal, Kumar Krishna, Ghosh, Arna, Sodhani, Shagun, Oberman, Adam, Richards, Blake |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The effectiveness of MAE pre-pretraining for billion-scale pretraining
von: Singh, Mannat, et al.
Veröffentlicht: (2023)
von: Singh, Mannat, et al.
Veröffentlicht: (2023)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
von: Wang, Wei, et al.
Veröffentlicht: (2026)
von: Wang, Wei, et al.
Veröffentlicht: (2026)
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
von: Garrido, Quentin, et al.
Veröffentlicht: (2025)
von: Garrido, Quentin, et al.
Veröffentlicht: (2025)
Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers
von: Grzywaczewski, Jakub, et al.
Veröffentlicht: (2026)
von: Grzywaczewski, Jakub, et al.
Veröffentlicht: (2026)
Steering CLIP's vision transformer with sparse autoencoders
von: Joseph, Sonia, et al.
Veröffentlicht: (2025)
von: Joseph, Sonia, et al.
Veröffentlicht: (2025)
SmolVLM: Redefining small and efficient multimodal models
von: Marafioti, Andrés, et al.
Veröffentlicht: (2025)
von: Marafioti, Andrés, et al.
Veröffentlicht: (2025)
Prompting Lipschitz-constrained network for multiple-in-one sparse-view CT reconstruction
von: Shi, Baoshun, et al.
Veröffentlicht: (2025)
von: Shi, Baoshun, et al.
Veröffentlicht: (2025)
SuperAnimal pretrained pose estimation models for behavioral analysis
von: Ye, Shaokai, et al.
Veröffentlicht: (2022)
von: Ye, Shaokai, et al.
Veröffentlicht: (2022)
The role of self-supervised pretraining in differentially private medical image analysis
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2026)
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2026)
Leveraging pretrained RGB denoisers for hyperspectral image restoration
von: Picone, Daniele, et al.
Veröffentlicht: (2026)
von: Picone, Daniele, et al.
Veröffentlicht: (2026)
Intra-view and Inter-view Correlation Guided Multi-view Novel Class Discovery
von: Wan, Xinhang, et al.
Veröffentlicht: (2025)
von: Wan, Xinhang, et al.
Veröffentlicht: (2025)
Harnessing Shared Relations via Multimodal Mixup Contrastive Learning for Multimodal Classification
von: Kumar, Raja, et al.
Veröffentlicht: (2024)
von: Kumar, Raja, et al.
Veröffentlicht: (2024)
Interpreting Physics in Video World Models
von: Joseph, Sonia, et al.
Veröffentlicht: (2026)
von: Joseph, Sonia, et al.
Veröffentlicht: (2026)
Self-supervised video pretraining yields robust and more human-aligned visual representations
von: Parthasarathy, Nikhil, et al.
Veröffentlicht: (2022)
von: Parthasarathy, Nikhil, et al.
Veröffentlicht: (2022)
Harnessing Artificial Intelligence for Wildlife Conservation
von: Fergus, Paul, et al.
Veröffentlicht: (2024)
von: Fergus, Paul, et al.
Veröffentlicht: (2024)
An explainable vision transformer with transfer learning based efficient drought stress identification
von: Patra, Aswini Kumar, et al.
Veröffentlicht: (2024)
von: Patra, Aswini Kumar, et al.
Veröffentlicht: (2024)
Pillar-0: A New Frontier for Radiology Foundation Models
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2025)
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2025)
Synthetic Art Generation and DeepFake Detection A Study on Jamini Roy Inspired Dataset
von: Agrawal, Kushal, et al.
Veröffentlicht: (2025)
von: Agrawal, Kushal, et al.
Veröffentlicht: (2025)
Understanding and Harnessing Sparsity in Unified Multimodal Models
von: He, Shwai, et al.
Veröffentlicht: (2025)
von: He, Shwai, et al.
Veröffentlicht: (2025)
HARIVO: Harnessing Text-to-Image Models for Video Generation
von: Kwon, Mingi, et al.
Veröffentlicht: (2024)
von: Kwon, Mingi, et al.
Veröffentlicht: (2024)
Mechanisms of Non-Monotonic Scaling in Vision Transformers
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
Parameter Reduction Improves Vision Transformers: A Comparative Study of Sharing and Width Reduction
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
Tracing the Representation Geometry of Language Models from Pretraining to Post-training
von: Li, Melody Zixuan, et al.
Veröffentlicht: (2025)
von: Li, Melody Zixuan, et al.
Veröffentlicht: (2025)
Did you just see that? Arbitrary view synthesis for egocentric replay of operating room workflows from ambient sensors
von: Zhang, Han, et al.
Veröffentlicht: (2025)
von: Zhang, Han, et al.
Veröffentlicht: (2025)
ClickDiffusion: Harnessing LLMs for Interactive Precise Image Editing
von: Helbling, Alec, et al.
Veröffentlicht: (2024)
von: Helbling, Alec, et al.
Veröffentlicht: (2024)
Harnessing Large Vision and Language Models in Agriculture: A Review
von: Zhu, Hongyan, et al.
Veröffentlicht: (2024)
von: Zhu, Hongyan, et al.
Veröffentlicht: (2024)
Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation
von: Wang, Jiyuan, et al.
Veröffentlicht: (2025)
von: Wang, Jiyuan, et al.
Veröffentlicht: (2025)
Towards Real-Time 2D Mapping: Harnessing Drones, AI, and Computer Vision for Advanced Insights
von: Agnur, Bharath Kumar
Veröffentlicht: (2024)
von: Agnur, Bharath Kumar
Veröffentlicht: (2024)
NanoVLMs: How small can we go and still make coherent Vision Language Models?
von: Agarwalla, Mukund, et al.
Veröffentlicht: (2025)
von: Agarwalla, Mukund, et al.
Veröffentlicht: (2025)
VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
von: Bulat, Adrian, et al.
Veröffentlicht: (2026)
von: Bulat, Adrian, et al.
Veröffentlicht: (2026)
Maintaining User Trust Through Multistage Uncertainty Aware Inference
von: Agrawal, Chandan, et al.
Veröffentlicht: (2023)
von: Agrawal, Chandan, et al.
Veröffentlicht: (2023)
MEDeA: Multi-view Efficient Depth Adjustment
von: Artemyev, Mikhail, et al.
Veröffentlicht: (2024)
von: Artemyev, Mikhail, et al.
Veröffentlicht: (2024)
Personalized Federated Learning for Cross-view Geo-localization
von: Anagnostopoulos, Christos, et al.
Veröffentlicht: (2024)
von: Anagnostopoulos, Christos, et al.
Veröffentlicht: (2024)
I Am Big, You Are Little; I Am Right, You Are Wrong
von: Kelly, David A., et al.
Veröffentlicht: (2025)
von: Kelly, David A., et al.
Veröffentlicht: (2025)
Concurrent validity of computer-vision artificial intelligence player tracking software using broadcast footage
von: Crang, Zachary L., et al.
Veröffentlicht: (2025)
von: Crang, Zachary L., et al.
Veröffentlicht: (2025)
GLFNET: Global-Local (frequency) Filter Networks for efficient medical image segmentation
von: Tragakis, Athanasios, et al.
Veröffentlicht: (2024)
von: Tragakis, Athanasios, et al.
Veröffentlicht: (2024)
AerOSeg: Harnessing SAM for Open-Vocabulary Segmentation in Remote Sensing Images
von: Dutta, Saikat, et al.
Veröffentlicht: (2025)
von: Dutta, Saikat, et al.
Veröffentlicht: (2025)
ConceptExpress: Harnessing Diffusion Models for Single-image Unsupervised Concept Extraction
von: Hao, Shaozhe, et al.
Veröffentlicht: (2024)
von: Hao, Shaozhe, et al.
Veröffentlicht: (2024)
Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis
von: Joshi, Soham, et al.
Veröffentlicht: (2025)
von: Joshi, Soham, et al.
Veröffentlicht: (2025)
MV-Swin-T: Mammogram Classification with Multi-view Swin Transformer
von: Sarker, Sushmita, et al.
Veröffentlicht: (2024)
von: Sarker, Sushmita, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The effectiveness of MAE pre-pretraining for billion-scale pretraining
von: Singh, Mannat, et al.
Veröffentlicht: (2023) -
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
von: Wang, Wei, et al.
Veröffentlicht: (2026) -
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
von: Garrido, Quentin, et al.
Veröffentlicht: (2025) -
Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers
von: Grzywaczewski, Jakub, et al.
Veröffentlicht: (2026) -
Steering CLIP's vision transformer with sparse autoencoders
von: Joseph, Sonia, et al.
Veröffentlicht: (2025)