SpacTor-T5: Pre-training T5 Models with Span Corruption and Replaced Token Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Ke, Jiang, Heinrich, Rostamizadeh, Afshin, Chakrabarti, Ayan, DeSalvo, Giulia, Kagy, Jean-François, Karydas, Lazaros, Citovsky, Gui, Kumar, Sanjiv |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SoftSRV: Learn to Generate Targeted Synthetic Data
by: DeSalvo, Giulia, et al.
Published: (2024)
by: DeSalvo, Giulia, et al.
Published: (2024)
Analyzing Similarity Metrics for Data Selection for Language Model Pretraining
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
GIST: Greedy Independent Set Thresholding for Max-Min Diversification with Submodular Utility
by: Fahrbach, Matthew, et al.
Published: (2024)
by: Fahrbach, Matthew, et al.
Published: (2024)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
by: Zhou, Yongchao, et al.
Published: (2023)
by: Zhou, Yongchao, et al.
Published: (2023)
Algorithms for Learning Kernels Based on Centered Alignment
by: Cortes, Corinna, et al.
Published: (2012)
by: Cortes, Corinna, et al.
Published: (2012)
Budgeted Multiple-Expert Deferral
by: DeSalvo, Giulia, et al.
Published: (2025)
by: DeSalvo, Giulia, et al.
Published: (2025)
Chemical Effects on X‐Ray Intensity Ratios and L 3 Absorption‐Edge in Hafnium Compounds
by: Harpreet Singh, et al.
Published: (2025)
by: Harpreet Singh, et al.
Published: (2025)
A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs
by: Rawat, Ankit Singh, et al.
Published: (2024)
by: Rawat, Ankit Singh, et al.
Published: (2024)
Rethinking FID: Towards a Better Evaluation Metric for Image Generation
by: Jayasumana, Sadeep, et al.
Published: (2023)
by: Jayasumana, Sadeep, et al.
Published: (2023)
Identifiability of Large Phylogenetic Mixtures for Many Phylogenetic Model Structures
by: Kagy, Bryson, et al.
Published: (2025)
by: Kagy, Bryson, et al.
Published: (2025)
Equidistant Circular Split Networks
by: Kagy, Bryson, et al.
Published: (2024)
by: Kagy, Bryson, et al.
Published: (2024)
Slight Corruption in Pre-training Data Makes Better Diffusion Models
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Efficient Continual Pre-training of LLMs for Low-resource Languages
by: Nag, Arijit, et al.
Published: (2024)
by: Nag, Arijit, et al.
Published: (2024)
ClustViT: Clustering-based Token Merging for Semantic Segmentation
by: Montello, Fabio, et al.
Published: (2025)
by: Montello, Fabio, et al.
Published: (2025)
VaulTor: Putting the TEE in Tor
by: Ikram, Humza, et al.
Published: (2024)
by: Ikram, Humza, et al.
Published: (2024)
LatentCRF: Continuous CRF for Efficient Latent Diffusion
by: Ranasinghe, Kanchana, et al.
Published: (2024)
by: Ranasinghe, Kanchana, et al.
Published: (2024)
Spiking Patches: Asynchronous, Sparse, and Efficient Tokens for Event Cameras
by: Øhrstrøm, Christoffer Koo, et al.
Published: (2025)
by: Øhrstrøm, Christoffer Koo, et al.
Published: (2025)
On the Effect of Token Merging on Pre-trained Models for Code
by: Saad, Mootez, et al.
Published: (2025)
by: Saad, Mootez, et al.
Published: (2025)
Emerging Property of Masked Token for Effective Pre-training
by: Choi, Hyesong, et al.
Published: (2024)
by: Choi, Hyesong, et al.
Published: (2024)
Towards Scalable Pre-training of Visual Tokenizers for Generation
by: Yao, Jingfeng, et al.
Published: (2025)
by: Yao, Jingfeng, et al.
Published: (2025)
Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models
by: Purason, Taido, et al.
Published: (2025)
by: Purason, Taido, et al.
Published: (2025)
MediaTor
Published: (2017)
Published: (2017)
Safety and Glycemic Outcomes Among Youth With New‐Onset Type 1 Diabetes Using a Tubeless Automated Insulin Delivery System
by: Daniel J. DeSalvo, et al.
Published: (2026)
by: Daniel J. DeSalvo, et al.
Published: (2026)
Pre-training Generative Recommender with Multi-Identifier Item Tokenization
by: Zheng, Bowen, et al.
Published: (2025)
by: Zheng, Bowen, et al.
Published: (2025)
NITP: Next Implicit Token Prediction for LLM Pre-training
by: Zhang, Xiangdong, et al.
Published: (2026)
by: Zhang, Xiangdong, et al.
Published: (2026)
SaTor: Exploring Satellite Routing in Tor to Reduce Latency
by: Li, Haozhi, et al.
Published: (2024)
by: Li, Haozhi, et al.
Published: (2024)
Measuring the neutron star equation of state from EMRIs in dark matter environments with LISA
by: Karydas, Theophanes K., et al.
Published: (2025)
by: Karydas, Theophanes K., et al.
Published: (2025)
Identification and Characterization of Outer Membrane Proteins and Membrane Spanning Protein Complexes in Brucella melitensis
by: Jahnvi Kapoor, et al.
Published: (2026)
by: Jahnvi Kapoor, et al.
Published: (2026)
Some continuity estimates for ruin probability and other ruin-related quantities
by: Kanellopoulos, Lazaros
Published: (2025)
by: Kanellopoulos, Lazaros
Published: (2025)
The Platform Is Mostly Not a Platform: Token Economies and Agent Discourse on Moltbook
by: Ayan, Necati A
Published: (2026)
by: Ayan, Necati A
Published: (2026)
GPTFace: Generative Pre-training of Facial-Linguistic Transformer by Span Masking and Weakly Correlated Text-image Data
by: Li, Yudong, et al.
Published: (2025)
by: Li, Yudong, et al.
Published: (2025)
TokenMark: A Modality-Agnostic Watermark for Pre-trained Transformers
by: Xu, Hengyuan, et al.
Published: (2024)
by: Xu, Hengyuan, et al.
Published: (2024)
On the vanishing of Ext and Tor
by: Bahlekeh, Abdolnaser, et al.
Published: (2025)
by: Bahlekeh, Abdolnaser, et al.
Published: (2025)
Analyzing Trends in Tor
by: Rahalkar, Chaitanya, et al.
Published: (2022)
by: Rahalkar, Chaitanya, et al.
Published: (2022)
El proyecto Tor
by: José Antonio Amaro López
Published: (2015)
by: José Antonio Amaro López
Published: (2015)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
by: Kolawole, Steven, et al.
Published: (2024)
by: Kolawole, Steven, et al.
Published: (2024)
Quantum Properties of Non-Dirichlet Boundary Conditions in Gravity
by: Draper, Patrick, et al.
Published: (2025)
by: Draper, Patrick, et al.
Published: (2025)
Response of interferometers to the vacuum of quantum gravity
by: Carney, Daniel, et al.
Published: (2024)
by: Carney, Daniel, et al.
Published: (2024)
Breakdown of Field Theory in Near-Horizon Regions
by: Banks, Tom, et al.
Published: (2024)
by: Banks, Tom, et al.
Published: (2024)
Salience-Based Adaptive Masking: Revisiting Token Dynamics for Enhanced Pre-training
by: Choi, Hyesong, et al.
Published: (2024)
by: Choi, Hyesong, et al.
Published: (2024)
Similar Items
-
SoftSRV: Learn to Generate Targeted Synthetic Data
by: DeSalvo, Giulia, et al.
Published: (2024) -
Analyzing Similarity Metrics for Data Selection for Language Model Pretraining
by: Sam, Dylan, et al.
Published: (2025) -
GIST: Greedy Independent Set Thresholding for Max-Min Diversification with Submodular Utility
by: Fahrbach, Matthew, et al.
Published: (2024) -
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
by: Zhou, Yongchao, et al.
Published: (2023) -
Algorithms for Learning Kernels Based on Centered Alignment
by: Cortes, Corinna, et al.
Published: (2012)