Meta CLIP 2: A Worldwide Scaling Recipe
Fuente:
arXiv
Saved in:
| Main Authors: | Chuang, Yung-Sung, Li, Yang, Wang, Dong, Yeh, Ching-Feng, Lyu, Kehan, Raghavendra, Ramya, Glass, James, Huang, Lifei, Weston, Jason, Zettlemoyer, Luke, Chen, Xinlei, Liu, Zhuang, Xie, Saining, Yih, Wen-tau, Li, Shang-Wen, Xu, Hu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoDE: CLIP Data Experts via Clustering
by: Ma, Jiawei, et al.
Published: (2024)
by: Ma, Jiawei, et al.
Published: (2024)
SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models
by: Chuang, Yung-Sung, et al.
Published: (2025)
by: Chuang, Yung-Sung, et al.
Published: (2025)
Altogether: Image Captioning via Re-aligning Alt-text
by: Xu, Hu, et al.
Published: (2024)
by: Xu, Hu, et al.
Published: (2024)
Anchored Decoding: Provably Reducing Copyright Risk for Any Language Model
by: He, Jacqueline, et al.
Published: (2026)
by: He, Jacqueline, et al.
Published: (2026)
Memory Layers at Scale
by: Berges, Vincent-Pierre, et al.
Published: (2024)
by: Berges, Vincent-Pierre, et al.
Published: (2024)
Demystifying CLIP Data
by: Xu, Hu, et al.
Published: (2023)
by: Xu, Hu, et al.
Published: (2023)
Improving Factuality with Explicit Working Memory
by: Chen, Mingda, et al.
Published: (2024)
by: Chen, Mingda, et al.
Published: (2024)
Reliable, Adaptable, and Attributable Language Models with Retrieval
by: Asai, Akari, et al.
Published: (2024)
by: Asai, Akari, et al.
Published: (2024)
Few-Shot Data Synthesis for Open Domain Multi-Hop Question Answering
by: Chen, Mingda, et al.
Published: (2023)
by: Chen, Mingda, et al.
Published: (2023)
HoneyBee: Data Recipes for Vision-Language Reasoners
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
Procedural Knowledge at Scale Improves Reasoning
by: Wu, Di, et al.
Published: (2026)
by: Wu, Di, et al.
Published: (2026)
Don't "Overthink" Passage Reranking: Is Reasoning Truly Necessary?
by: Jedidi, Nour, et al.
Published: (2025)
by: Jedidi, Nour, et al.
Published: (2025)
Zero-Shot Dense Retrieval with Embeddings from Relevance Feedback
by: Jedidi, Nour, et al.
Published: (2024)
by: Jedidi, Nour, et al.
Published: (2024)
Deconstructing Denoising Diffusion Models for Self-Supervised Learning
by: Chen, Xinlei, et al.
Published: (2024)
by: Chen, Xinlei, et al.
Published: (2024)
Learning to Reason for Factuality
by: Chen, Xilun, et al.
Published: (2025)
by: Chen, Xilun, et al.
Published: (2025)
Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework
by: Wang, Dong, et al.
Published: (2025)
by: Wang, Dong, et al.
Published: (2025)
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
by: Liang, Weixin, et al.
Published: (2024)
by: Liang, Weixin, et al.
Published: (2024)
RecipeGen: A Benchmark for Real-World Recipe Image Generation
by: Zhang, Ruoxuan, et al.
Published: (2025)
by: Zhang, Ruoxuan, et al.
Published: (2025)
ImpRAG: Retrieval-Augmented Generation with Implicit Queries
by: Zhang, Wenzheng, et al.
Published: (2025)
by: Zhang, Wenzheng, et al.
Published: (2025)
Group-Level Data Selection for Efficient Pretraining
by: Yu, Zichun, et al.
Published: (2025)
by: Yu, Zichun, et al.
Published: (2025)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
by: Sukhbaatar, Sainbayar, et al.
Published: (2024)
by: Sukhbaatar, Sainbayar, et al.
Published: (2024)
ReasonIR: Training Retrievers for Reasoning Tasks
by: Shao, Rulin, et al.
Published: (2025)
by: Shao, Rulin, et al.
Published: (2025)
Continual Learning via Sparse Memory Finetuning
by: Lin, Jessy, et al.
Published: (2025)
by: Lin, Jessy, et al.
Published: (2025)
Nearest Neighbor Speculative Decoding for LLM Generation and Attribution
by: Li, Minghan, et al.
Published: (2024)
by: Li, Minghan, et al.
Published: (2024)
DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers
by: Ma, Xueguang, et al.
Published: (2025)
by: Ma, Xueguang, et al.
Published: (2025)
Empirical Recipes for Efficient and Compact Vision-Language Models
by: Huang, Jiabo, et al.
Published: (2026)
by: Huang, Jiabo, et al.
Published: (2026)
Better Alignment with Instruction Back-and-Forth Translation
by: Nguyen, Thao, et al.
Published: (2024)
by: Nguyen, Thao, et al.
Published: (2024)
DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
by: Chuang, Yung-Sung, et al.
Published: (2023)
by: Chuang, Yung-Sung, et al.
Published: (2023)
Cleansing Jewel: A Neural Spelling Correction Model Built On Google OCR-ed Tibetan Manuscripts
by: Luo, Queenie, et al.
Published: (2023)
by: Luo, Queenie, et al.
Published: (2023)
Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
by: Liu, Alexander H., et al.
Published: (2024)
by: Liu, Alexander H., et al.
Published: (2024)
(Mis)Fitting: A Survey of Scaling Laws
by: Li, Margaret, et al.
Published: (2025)
by: Li, Margaret, et al.
Published: (2025)
Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
by: Chuang, Yung-Sung, et al.
Published: (2024)
by: Chuang, Yung-Sung, et al.
Published: (2024)
FLAME: Factuality-Aware Alignment for Large Language Models
by: Lin, Sheng-Chieh, et al.
Published: (2024)
by: Lin, Sheng-Chieh, et al.
Published: (2024)
Warm Diffusion: Recipe for Blur-Noise Mixture Diffusion Models
by: Hsueh, Hao-Chien, et al.
Published: (2025)
by: Hsueh, Hao-Chien, et al.
Published: (2025)
Public–private partnership in pipelining science of acute care ecosystem: Insights from Taiwan's Presidential Hackathon
by: Chao‐Wen Chen, et al.
Published: (2025)
by: Chao‐Wen Chen, et al.
Published: (2025)
Smoking and Elevated Preneoadjuvant Chemoradiotherapy Serum Carcinoembryonic Antigen Levels Are Associated With High Tumor Regression Grade and Poor Survival in Patients With Locally Advanced Rectal Cancer
by: Jen‐Pin Chuang, et al.
Published: (2025)
by: Jen‐Pin Chuang, et al.
Published: (2025)
RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
by: Zhang, Ruoxuan, et al.
Published: (2025)
by: Zhang, Ruoxuan, et al.
Published: (2025)
SEPT14 complexes maintain sperm morphogenesis and function
by: Han‐Yu Wang, et al.
Published: (2025)
by: Han‐Yu Wang, et al.
Published: (2025)
Probiotics Containing Collagen Peptides for Weight Loss and Improvement of Brain Disorders Caused by Obesity in a Mouse Model
by: Wei-Wen Sung, et al.
Published: (2025)
by: Wei-Wen Sung, et al.
Published: (2025)
Self-Alignment with Instruction Backtranslation
by: Li, Xian, et al.
Published: (2023)
by: Li, Xian, et al.
Published: (2023)
Similar Items
-
MoDE: CLIP Data Experts via Clustering
by: Ma, Jiawei, et al.
Published: (2024) -
SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models
by: Chuang, Yung-Sung, et al.
Published: (2025) -
Altogether: Image Captioning via Re-aligning Alt-text
by: Xu, Hu, et al.
Published: (2024) -
Anchored Decoding: Provably Reducing Copyright Risk for Any Language Model
by: He, Jacqueline, et al.
Published: (2026) -
Memory Layers at Scale
by: Berges, Vincent-Pierre, et al.
Published: (2024)