Breaking the Memory Barrier: Near Infinite Batch Size Scaling for Contrastive Loss
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Zesen, Zhang, Hang, Li, Kehan, Leng, Sicong, Hu, Zhiqiang, Wu, Fei, Zhao, Deli, Li, Xin, Bing, Lidong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
by: Leng, Sicong, et al.
Published: (2024)
by: Leng, Sicong, et al.
Published: (2024)
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
by: Zhang, Boqiang, et al.
Published: (2025)
by: Zhang, Boqiang, et al.
Published: (2025)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
by: Cheng, Zesen, et al.
Published: (2024)
by: Cheng, Zesen, et al.
Published: (2024)
MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier
by: Yang, Zonglin, et al.
Published: (2026)
by: Yang, Zonglin, et al.
Published: (2026)
MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
by: Leng, Sicong, et al.
Published: (2025)
by: Leng, Sicong, et al.
Published: (2025)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
by: Yuan, Yuqian, et al.
Published: (2024)
by: Yuan, Yuqian, et al.
Published: (2024)
Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining
by: Thirukovalluru, Raghuveer, et al.
Published: (2025)
by: Thirukovalluru, Raghuveer, et al.
Published: (2025)
Evolving Prompts In-Context: An Open-ended, Self-replicating Perspective
by: Wang, Jianyu, et al.
Published: (2025)
by: Wang, Jianyu, et al.
Published: (2025)
Breaking 50% Energy Loss Barrier in Polarized LEDs via Twisted Grating Metasurface Integration
by: Chong‐De Zhang, et al.
Published: (2025)
by: Chong‐De Zhang, et al.
Published: (2025)
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
by: Jiang, Yuming, et al.
Published: (2025)
by: Jiang, Yuming, et al.
Published: (2025)
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
by: Zhang, Wenqi, et al.
Published: (2025)
by: Zhang, Wenqi, et al.
Published: (2025)
Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
by: Li, Huize, et al.
Published: (2026)
by: Li, Huize, et al.
Published: (2026)
Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices
by: Shen, Tao, et al.
Published: (2025)
by: Shen, Tao, et al.
Published: (2025)
Large Language Models can Contrastively Refine their Generation for Better Sentence Representation Learning
by: Wang, Huiming, et al.
Published: (2023)
by: Wang, Huiming, et al.
Published: (2023)
Scaling Law for Language Models Training Considering Batch Size
by: Shuai, Xian, et al.
Published: (2024)
by: Shuai, Xian, et al.
Published: (2024)
InfiniteICL: Breaking the Limit of Context Window Size via Long Short-term Memory Transformation
by: Cao, Bowen, et al.
Published: (2025)
by: Cao, Bowen, et al.
Published: (2025)
Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
by: Zhao, Ruochen, et al.
Published: (2024)
by: Zhao, Ruochen, et al.
Published: (2024)
Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling
by: Li, Shuaipeng, et al.
Published: (2024)
by: Li, Shuaipeng, et al.
Published: (2024)
Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit
by: Filatov, Oleg, et al.
Published: (2024)
by: Filatov, Oleg, et al.
Published: (2024)
Stabilize the Latent Space for Image Autoregressive Modeling: A Unified Perspective
by: Zhu, Yongxin, et al.
Published: (2024)
by: Zhu, Yongxin, et al.
Published: (2024)
MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
by: Chen, Yu, et al.
Published: (2026)
by: Chen, Yu, et al.
Published: (2026)
LLMBisect: Breaking Barriers in Bug Bisection with A Comparative Analysis Pipeline
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Characterizing the Dilemma of Performance and Index Size in Billion-Scale Vector Search and Breaking It with Second-Tier Memory
by: Cheng, Rongxin, et al.
Published: (2024)
by: Cheng, Rongxin, et al.
Published: (2024)
Breaking the Size Barrier in Glass Nanopores to 2 nm: A Path to High‐Resolution Biomolecular Fingerprinting
by: Xiaoyu Chen, et al.
Published: (2026)
by: Xiaoyu Chen, et al.
Published: (2026)
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
by: Vaessen, Nik, et al.
Published: (2024)
by: Vaessen, Nik, et al.
Published: (2024)
Break the ID-Language Barrier: An Adaption Framework for LLM-based Sequential Recommendation
by: Yu, Xiaohan, et al.
Published: (2024)
by: Yu, Xiaohan, et al.
Published: (2024)
Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency
by: Wang, Zhikai, et al.
Published: (2025)
by: Wang, Zhikai, et al.
Published: (2025)
GraCo: Granularity-Controllable Interactive Segmentation
by: Zhao, Yian, et al.
Published: (2024)
by: Zhao, Yian, et al.
Published: (2024)
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
Instance Brownian Bridge as Texts for Open-vocabulary Video Instance Segmentation
by: Cheng, Zesen, et al.
Published: (2024)
by: Cheng, Zesen, et al.
Published: (2024)
Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative Model
by: Li, Gen, et al.
Published: (2020)
by: Li, Gen, et al.
Published: (2020)
The Spectrality of Infinite Convolutions in $\mathbb{R}^d$
by: Li, Wenxia, et al.
Published: (2022)
by: Li, Wenxia, et al.
Published: (2022)
Progressively Exploring and Exploiting Inference Data to Break Fine-Grained Classification Barrier
by: Zhao, Li-Jun, et al.
Published: (2024)
by: Zhao, Li-Jun, et al.
Published: (2024)
Enabling Large Batch Size Training for DNN Models Beyond the Memory Limit While Maintaining Performance
by: Piao, XinYu, et al.
Published: (2021)
by: Piao, XinYu, et al.
Published: (2021)
Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
How to Set the Batch Size for Large-Scale Pre-training?
by: Zhou, Yunhua, et al.
Published: (2026)
by: Zhou, Yunhua, et al.
Published: (2026)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
RynnEC: Bringing MLLMs into Embodied World
by: Dang, Ronghao, et al.
Published: (2025)
by: Dang, Ronghao, et al.
Published: (2025)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
by: Ostroukhov, Petr, et al.
Published: (2024)
by: Ostroukhov, Petr, et al.
Published: (2024)
Breaking Language Barriers: Cross-Lingual Continual Pre-Training at Scale
by: Zheng, Wenzhen, et al.
Published: (2024)
by: Zheng, Wenzhen, et al.
Published: (2024)
Similar Items
-
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
by: Leng, Sicong, et al.
Published: (2024) -
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
by: Zhang, Boqiang, et al.
Published: (2025) -
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
by: Cheng, Zesen, et al.
Published: (2024) -
MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier
by: Yang, Zonglin, et al.
Published: (2026) -
MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
by: Leng, Sicong, et al.
Published: (2025)