Global Minimizers of Sigmoid Contrastive Loss
Fuente:
arXiv
Saved in:
| Main Authors: | Bangachev, Kiril, Bresler, Guy, Noman, Iliyas, Polyanskiy, Yury |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Representation Alignment Rests on Linear Structure
by: Bangachev, Kiril, et al.
Published: (2026)
by: Bangachev, Kiril, et al.
Published: (2026)
Is Dimensionality a Barrier for Retrieval Models?
by: Bangachev, Kiril, et al.
Published: (2026)
by: Bangachev, Kiril, et al.
Published: (2026)
Random Algebraic Graphs and Their Convergence to ErdőS–Rényi
by: Kiril Bangachev, et al.
Published: (2025)
by: Kiril Bangachev, et al.
Published: (2025)
Graph Quasirandomness for Hypothesis Testing of Stochastic Block Models
by: Bangachev, Kiril, et al.
Published: (2025)
by: Bangachev, Kiril, et al.
Published: (2025)
On The Fourier Coefficients of High-Dimensional Random Geometric Graphs
by: Bangachev, Kiril, et al.
Published: (2024)
by: Bangachev, Kiril, et al.
Published: (2024)
Sandwiching Random Geometric Graphs and Erdos-Renyi with Applications: Sharp Thresholds, Robust Testing, and Enumeration
by: Bangachev, Kiril, et al.
Published: (2024)
by: Bangachev, Kiril, et al.
Published: (2024)
High-Rate Quantized Matrix Multiplication II
by: Ordentlich, Or, et al.
Published: (2026)
by: Ordentlich, Or, et al.
Published: (2026)
Contrastive Learning for Multi Label ECG Classification with Jaccard Score Based Sigmoid Loss
by: Takahashi, Junichiro, et al.
Published: (2026)
by: Takahashi, Junichiro, et al.
Published: (2026)
Optimal Quantization for Matrix Multiplication
by: Ordentlich, Or, et al.
Published: (2024)
by: Ordentlich, Or, et al.
Published: (2024)
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
by: Zimin, Aleksandr, et al.
Published: (2026)
by: Zimin, Aleksandr, et al.
Published: (2026)
NestQuant: Nested Lattice Quantization for Matrix Products and LLMs
by: Savkin, Semyon, et al.
Published: (2025)
by: Savkin, Semyon, et al.
Published: (2025)
Achieving the Tightest Relaxation of Sigmoids for Formal Verification
by: Chevalier, Samuel, et al.
Published: (2024)
by: Chevalier, Samuel, et al.
Published: (2024)
Near-Optimal Time-Sparsity Trade-Offs for Solving Noisy Linear Equations
by: Bangachev, Kiril, et al.
Published: (2024)
by: Bangachev, Kiril, et al.
Published: (2024)
Scaling Limits of Long-Context Transformers
by: Bruno, Giuseppe, et al.
Published: (2026)
by: Bruno, Giuseppe, et al.
Published: (2026)
Critical attention scaling in long-context transformers
by: Chen, Shi, et al.
Published: (2025)
by: Chen, Shi, et al.
Published: (2025)
Adaptive Friction in Deep Learning: Enhancing Optimizers with Sigmoid and Tanh Function
by: Zheng, Hongye, et al.
Published: (2024)
by: Zheng, Hongye, et al.
Published: (2024)
Beyond Gaussian Initializations: Signal Preserving Weight Initialization for Odd-Sigmoid Activations
by: Lee, Hyunwoo, et al.
Published: (2025)
by: Lee, Hyunwoo, et al.
Published: (2025)
The Geometry of Grokking: Norm Minimization on the Zero-Loss Manifold
by: Musat, Tiberiu
Published: (2025)
by: Musat, Tiberiu
Published: (2025)
Learned Bayesian Cramér-Rao Bound for Unknown Measurement Models Using Score Neural Networks
by: Habi, Hai Victor, et al.
Published: (2025)
by: Habi, Hai Victor, et al.
Published: (2025)
Improving Node Representation by Boosting Target-Aware Contrastive Loss
by: Lin, Ying-Chun, et al.
Published: (2024)
by: Lin, Ying-Chun, et al.
Published: (2024)
Analysis of Using Sigmoid Loss for Contrastive Learning
by: Lee, Chungpa, et al.
Published: (2024)
by: Lee, Chungpa, et al.
Published: (2024)
Minimizing Surrogate Losses for Decision-Focused Learning using Differentiable Optimization
by: Mandi, Jayanta, et al.
Published: (2025)
by: Mandi, Jayanta, et al.
Published: (2025)
A Huber Loss Minimization Approach to Byzantine Robust Federated Learning
by: Zhao, Puning, et al.
Published: (2023)
by: Zhao, Puning, et al.
Published: (2023)
Train-Once Plan-Anywhere Kinodynamic Motion Planning via Diffusion Trees
by: Hassidof, Yaniv, et al.
Published: (2025)
by: Hassidof, Yaniv, et al.
Published: (2025)
FAME: Formal Abstract Minimal Explanation for Neural Networks
by: Boumazouza, Ryma, et al.
Published: (2026)
by: Boumazouza, Ryma, et al.
Published: (2026)
Global Concept Explanations for Graphs by Contrastive Learning
by: Teufel, Jonas, et al.
Published: (2024)
by: Teufel, Jonas, et al.
Published: (2024)
Making Sigmoid-MSE Great Again: Output Reset Challenges Softmax Cross-Entropy in Neural Network Classification
by: Tyagi, Kanishka, et al.
Published: (2024)
by: Tyagi, Kanishka, et al.
Published: (2024)
TACO: Temporal Latent Action-Driven Contrastive Loss for Visual Reinforcement Learning
by: Zheng, Ruijie, et al.
Published: (2023)
by: Zheng, Ruijie, et al.
Published: (2023)
Comparing Contrastive and Triplet Loss: Variance Analysis and Optimization Behavior
by: Zeng, Donghuo
Published: (2025)
by: Zeng, Donghuo
Published: (2025)
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
by: Yan, Fanqi, et al.
Published: (2025)
by: Yan, Fanqi, et al.
Published: (2025)
Locally Convex Global Loss Network for Decision-Focused Learning
by: Jeon, Haeun, et al.
Published: (2024)
by: Jeon, Haeun, et al.
Published: (2024)
Local-Global Multimodal Contrastive Learning for Molecular Property Prediction
by: Liu, Xiayu, et al.
Published: (2026)
by: Liu, Xiayu, et al.
Published: (2026)
Plan First, Diffuse Later: Extrinsic Graph Guidance for Long-Horizon Diffusion Planning
by: Hassidof, Yaniv, et al.
Published: (2026)
by: Hassidof, Yaniv, et al.
Published: (2026)
SimO Loss: Anchor-Free Contrastive Loss for Fine-Grained Supervised Contrastive Learning
by: Bouhsine, Taha, et al.
Published: (2024)
by: Bouhsine, Taha, et al.
Published: (2024)
FairNet: Dynamic Fairness Correction without Performance Loss via Contrastive Conditional LoRA
by: Zhou, Songqi, et al.
Published: (2025)
by: Zhou, Songqi, et al.
Published: (2025)
Automated Design of Linear Bounding Functions for Sigmoidal Nonlinearities in Neural Networks
by: König, Matthias, et al.
Published: (2024)
by: König, Matthias, et al.
Published: (2024)
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
by: Peng, Hanyang, et al.
Published: (2025)
by: Peng, Hanyang, et al.
Published: (2025)
High-Rate Quantized Matrix Multiplication I
by: Ordentlich, Or, et al.
Published: (2026)
by: Ordentlich, Or, et al.
Published: (2026)
Meta-GCN: A Dynamically Weighted Loss Minimization Method for Dealing with the Data Imbalance in Graph Neural Networks
by: Mohammadizadeh, Mahdi, et al.
Published: (2024)
by: Mohammadizadeh, Mahdi, et al.
Published: (2024)
Clustering in Causal Attention Masking
by: Karagodin, Nikita, et al.
Published: (2024)
by: Karagodin, Nikita, et al.
Published: (2024)
Similar Items
-
Representation Alignment Rests on Linear Structure
by: Bangachev, Kiril, et al.
Published: (2026) -
Is Dimensionality a Barrier for Retrieval Models?
by: Bangachev, Kiril, et al.
Published: (2026) -
Random Algebraic Graphs and Their Convergence to ErdőS–Rényi
by: Kiril Bangachev, et al.
Published: (2025) -
Graph Quasirandomness for Hypothesis Testing of Stochastic Block Models
by: Bangachev, Kiril, et al.
Published: (2025) -
On The Fourier Coefficients of High-Dimensional Random Geometric Graphs
by: Bangachev, Kiril, et al.
Published: (2024)