CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mehta, Sachin, Horton, Maxwell, Faghri, Fartash, Sekhavat, Mohammad Hossein, Najibi, Mahyar, Farajtabar, Mehrdad, Tuzel, Oncel, Rastegari, Mohammad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models
von: Vemulapalli, Raviteja, et al.
Veröffentlicht: (2023)
von: Vemulapalli, Raviteja, et al.
Veröffentlicht: (2023)
SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding
von: Wang, Haoxiang, et al.
Veröffentlicht: (2023)
von: Wang, Haoxiang, et al.
Veröffentlicht: (2023)
TiC-CLIP: Continual Training of CLIP Models
von: Garg, Saurabh, et al.
Veröffentlicht: (2023)
von: Garg, Saurabh, et al.
Veröffentlicht: (2023)
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)
Computational Bottlenecks of Training Small-scale Large Language Models
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
von: Li, Jeffrey, et al.
Veröffentlicht: (2025)
von: Li, Jeffrey, et al.
Veröffentlicht: (2025)
OpenELM: An Efficient Language Model Family with Open Training and Inference Framework
von: Mehta, Sachin, et al.
Veröffentlicht: (2024)
von: Mehta, Sachin, et al.
Veröffentlicht: (2024)
From Dense to Dynamic: Token-Difficulty Driven MoEfication of Pre-Trained LLMs
von: Nishu, Kumari, et al.
Veröffentlicht: (2025)
von: Nishu, Kumari, et al.
Veröffentlicht: (2025)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2023)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2023)
Bytes Are All You Need: Transformers Operating Directly On File Bytes
von: Horton, Maxwell, et al.
Veröffentlicht: (2023)
von: Horton, Maxwell, et al.
Veröffentlicht: (2023)
Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting
von: Huang, Chen, et al.
Veröffentlicht: (2025)
von: Huang, Chen, et al.
Veröffentlicht: (2025)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random Generators
von: Shafipour, Rasoul, et al.
Veröffentlicht: (2024)
von: Shafipour, Rasoul, et al.
Veröffentlicht: (2024)
LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2026)
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2026)
MobileCLIP2: Improving Multi-Modal Reinforced Training
von: Faghri, Fartash, et al.
Veröffentlicht: (2025)
von: Faghri, Fartash, et al.
Veröffentlicht: (2025)
Diffusion Models as Masked Audio-Video Learners
von: Nunez, Elvis, et al.
Veröffentlicht: (2023)
von: Nunez, Elvis, et al.
Veröffentlicht: (2023)
Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation
von: Merth, Thomas, et al.
Veröffentlicht: (2024)
von: Merth, Thomas, et al.
Veröffentlicht: (2024)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
von: Samragh, Mohammad, et al.
Veröffentlicht: (2024)
von: Samragh, Mohammad, et al.
Veröffentlicht: (2024)
CtrlSynth: Controllable Image Text Synthesis for Data-Efficient Multimodal Learning
von: Cao, Qingqing, et al.
Veröffentlicht: (2024)
von: Cao, Qingqing, et al.
Veröffentlicht: (2024)
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
von: Mirzadeh, Iman, et al.
Veröffentlicht: (2024)
von: Mirzadeh, Iman, et al.
Veröffentlicht: (2024)
Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity
von: Joudaki, Amir, et al.
Veröffentlicht: (2025)
von: Joudaki, Amir, et al.
Veröffentlicht: (2025)
MUSCLE: A Model Update Strategy for Compatible LLM Evolution
von: Echterhoff, Jessica, et al.
Veröffentlicht: (2024)
von: Echterhoff, Jessica, et al.
Veröffentlicht: (2024)
KV Prediction for Improved Time to First Token
von: Horton, Maxwell, et al.
Veröffentlicht: (2024)
von: Horton, Maxwell, et al.
Veröffentlicht: (2024)
Speculative Streaming: Fast LLM Inference without Auxiliary Models
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2024)
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2024)
Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
von: Wen, Yuxin, et al.
Veröffentlicht: (2024)
von: Wen, Yuxin, et al.
Veröffentlicht: (2024)
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2025)
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2025)
VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2026)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2026)
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2024)
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2024)
AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
von: Chowdhury, Sanjoy, et al.
Veröffentlicht: (2025)
von: Chowdhury, Sanjoy, et al.
Veröffentlicht: (2025)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
von: Armandpour, Mohammadreza, et al.
Veröffentlicht: (2026)
von: Armandpour, Mohammadreza, et al.
Veröffentlicht: (2026)
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
von: Shojaee, Parshin, et al.
Veröffentlicht: (2025)
von: Shojaee, Parshin, et al.
Veröffentlicht: (2025)
FastVLM: Efficient Vision Encoding for Vision Language Models
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)
Assessment of Oxidative Stress Markers in Cats Undergoing Ovariohysterectomy by a Midline or Flank Approach
von: Zohreh Nazari, et al.
Veröffentlicht: (2025)
von: Zohreh Nazari, et al.
Veröffentlicht: (2025)
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2023)
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2023)
Effort-Optimized, Accuracy-Driven Labelling and Validation of Test Inputs for DL Systems: A Mixed-Integer Linear Programming Approach
von: Amini, Mohammad Hossein, et al.
Veröffentlicht: (2025)
von: Amini, Mohammad Hossein, et al.
Veröffentlicht: (2025)
The path towards contact-based physical human-robot interaction
von: Farajtabar, Mohammad, et al.
Veröffentlicht: (2024)
von: Farajtabar, Mohammad, et al.
Veröffentlicht: (2024)
An Efficient and Streaming Audio Visual Active Speaker Detection System
von: Kundu, Arnav, et al.
Veröffentlicht: (2024)
von: Kundu, Arnav, et al.
Veröffentlicht: (2024)
Beyond a Single Extractor: Re-thinking HTML-to-Text Extraction for LLM Pretraining
von: Li, Jeffrey, et al.
Veröffentlicht: (2026)
von: Li, Jeffrey, et al.
Veröffentlicht: (2026)
Network-Based Video Recommendation Using Viewing Patterns and Modularity Analysis: An Integrated Framework
von: Maghsoudi, Mehrdad, et al.
Veröffentlicht: (2023)
von: Maghsoudi, Mehrdad, et al.
Veröffentlicht: (2023)
SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF
von: Chegini, Atoosa, et al.
Veröffentlicht: (2024)
von: Chegini, Atoosa, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models
von: Vemulapalli, Raviteja, et al.
Veröffentlicht: (2023) -
SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding
von: Wang, Haoxiang, et al.
Veröffentlicht: (2023) -
TiC-CLIP: Continual Training of CLIP Models
von: Garg, Saurabh, et al.
Veröffentlicht: (2023) -
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024) -
Computational Bottlenecks of Training Small-scale Large Language Models
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)