Cropping outperforms dropout as an augmentation strategy for self-supervised training of text embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | González-Márquez, Rita, Berens, Philipp, Kobak, Dmitry |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning representations of learning representations
by: González-Márquez, Rita, et al.
Published: (2024)
by: González-Márquez, Rita, et al.
Published: (2024)
Unsupervised visualization of image datasets using contrastive learning
by: Böhm, Jan Niklas, et al.
Published: (2022)
by: Böhm, Jan Niklas, et al.
Published: (2022)
Persistent Homology for High-dimensional Data Based on Spectral Methods
by: Damrich, Sebastian, et al.
Published: (2023)
by: Damrich, Sebastian, et al.
Published: (2023)
Attraction-Repulsion Spectrum in Neighbor Embeddings
by: Böhm, Jan Niklas, et al.
Published: (2020)
by: Böhm, Jan Niklas, et al.
Published: (2020)
Benchmarking pre-trained text embedding models in aligning built asset information
by: Shahinmoghadam, Mehrzad, et al.
Published: (2024)
by: Shahinmoghadam, Mehrzad, et al.
Published: (2024)
Wasserstein t-SNE
by: Bachmann, Fynn, et al.
Published: (2022)
by: Bachmann, Fynn, et al.
Published: (2022)
Soft-CAM: Making black box models self-explainable for medical image analysis
by: Djoumessi, Kerol, et al.
Published: (2025)
by: Djoumessi, Kerol, et al.
Published: (2025)
Self-supervised Visualisation of Medical Image Datasets
by: Nwabufo, Ifeoma Veronica, et al.
Published: (2024)
by: Nwabufo, Ifeoma Veronica, et al.
Published: (2024)
Scaling Down Deep Learning with MNIST-1D
by: Greydanus, Sam, et al.
Published: (2020)
by: Greydanus, Sam, et al.
Published: (2020)
Improving embedding with contrastive fine-tuning on small datasets with expert-augmented scores
by: Lu, Jun, et al.
Published: (2024)
by: Lu, Jun, et al.
Published: (2024)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
by: Fujita, Kenichi, et al.
Published: (2024)
by: Fujita, Kenichi, et al.
Published: (2024)
Probing self-attention in self-supervised speech models for cross-linguistic differences
by: Gopinath, Sai, et al.
Published: (2024)
by: Gopinath, Sai, et al.
Published: (2024)
Delving into LLM-assisted writing in biomedical publications through excess vocabulary
by: Kobak, Dmitry, et al.
Published: (2024)
by: Kobak, Dmitry, et al.
Published: (2024)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
by: Jørgensen, Mikkel Godsk, et al.
Published: (2026)
by: Jørgensen, Mikkel Godsk, et al.
Published: (2026)
Optimal synthesis embeddings
by: Santana, Roberto, et al.
Published: (2024)
by: Santana, Roberto, et al.
Published: (2024)
Text clustering applied to data augmentation in legal contexts
by: Freitas, Lucas José Gonçalves, et al.
Published: (2024)
by: Freitas, Lucas José Gonçalves, et al.
Published: (2024)
Mitigating Shortcut Learning via Feature Disentanglement in Medical Imaging: A Benchmark Study
by: Müller, Sarah, et al.
Published: (2026)
by: Müller, Sarah, et al.
Published: (2026)
Groundedness in Retrieval-augmented Long-form Generation: An Empirical Study
by: Stolfo, Alessandro
Published: (2024)
by: Stolfo, Alessandro
Published: (2024)
Retrieval-augmented Prompt Learning for Pre-trained Foundation Models
by: Chen, Xiang, et al.
Published: (2025)
by: Chen, Xiang, et al.
Published: (2025)
An Senegalese Legal Texts Structuration Using LLM-augmented Knowledge Graph
by: Kane, Oumar, et al.
Published: (2025)
by: Kane, Oumar, et al.
Published: (2025)
Towards Knowledge Checking in Retrieval-augmented Generation: A Representation Perspective
by: Zeng, Shenglai, et al.
Published: (2024)
by: Zeng, Shenglai, et al.
Published: (2024)
BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language models
by: Lavechin, Marvin, et al.
Published: (2023)
by: Lavechin, Marvin, et al.
Published: (2023)
Machine-generated text detection prevents language model collapse
by: Drayson, George, et al.
Published: (2025)
by: Drayson, George, et al.
Published: (2025)
Patent Representation Learning via Self-supervision
by: Zuo, You, et al.
Published: (2025)
by: Zuo, You, et al.
Published: (2025)
The study of short texts in digital politics: Document aggregation for topic modeling
by: Nakka, Nitheesha, et al.
Published: (2025)
by: Nakka, Nitheesha, et al.
Published: (2025)
LLM-based feature generation from text for interpretable machine learning
by: Balek, Vojtěch, et al.
Published: (2024)
by: Balek, Vojtěch, et al.
Published: (2024)
Isolating authorship from content with semantic embeddings and contrastive learning
by: Huertas-Tato, Javier, et al.
Published: (2024)
by: Huertas-Tato, Javier, et al.
Published: (2024)
Polynomial Mixing for Efficient Self-supervised Speech Encoders
by: Feillet, Eva, et al.
Published: (2026)
by: Feillet, Eva, et al.
Published: (2026)
Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training
by: Bawazir, Ameera, et al.
Published: (2024)
by: Bawazir, Ameera, et al.
Published: (2024)
Extractive text summarisation of Privacy Policy documents using machine learning approaches
by: Choi, Chanwoo
Published: (2024)
by: Choi, Chanwoo
Published: (2024)
Critical biblical studies via word frequency analysis: unveiling text authorship
by: Faigenbaum-Golovin, Shira, et al.
Published: (2024)
by: Faigenbaum-Golovin, Shira, et al.
Published: (2024)
AIDetx: a compression-based method for identification of machine-learning generated text
by: Almeida, Leonardo, et al.
Published: (2024)
by: Almeida, Leonardo, et al.
Published: (2024)
Dissecting embedding method: learning higher-order structures from data
by: Tupikina, Liubov, et al.
Published: (2024)
by: Tupikina, Liubov, et al.
Published: (2024)
A low latency attention module for streaming self-supervised speech representation learning
by: Ma, Jianbo, et al.
Published: (2023)
by: Ma, Jianbo, et al.
Published: (2023)
LLMs can learn self-restraint through iterative self-reflection
by: Piché, Alexandre, et al.
Published: (2024)
by: Piché, Alexandre, et al.
Published: (2024)
Discovering influential text using convolutional neural networks
by: Ayers, Megan, et al.
Published: (2024)
by: Ayers, Megan, et al.
Published: (2024)
Retrieval-augmented GUI Agents with Generative Guidelines
by: Xu, Ran, et al.
Published: (2025)
by: Xu, Ran, et al.
Published: (2025)
A multilingual training strategy for low resource Text to Speech
by: Amalas, Asma, et al.
Published: (2024)
by: Amalas, Asma, et al.
Published: (2024)
Performance of diverse evaluation metrics in NLP-based assessment and text generation of consumer complaints
by: Gao, Peiheng, et al.
Published: (2025)
by: Gao, Peiheng, et al.
Published: (2025)
TAGLAS: An atlas of text-attributed graph datasets in the era of large graph and language models
by: Feng, Jiarui, et al.
Published: (2024)
by: Feng, Jiarui, et al.
Published: (2024)
Similar Items
-
Learning representations of learning representations
by: González-Márquez, Rita, et al.
Published: (2024) -
Unsupervised visualization of image datasets using contrastive learning
by: Böhm, Jan Niklas, et al.
Published: (2022) -
Persistent Homology for High-dimensional Data Based on Spectral Methods
by: Damrich, Sebastian, et al.
Published: (2023) -
Attraction-Repulsion Spectrum in Neighbor Embeddings
by: Böhm, Jan Niklas, et al.
Published: (2020) -
Benchmarking pre-trained text embedding models in aligning built asset information
by: Shahinmoghadam, Mehrzad, et al.
Published: (2024)