Deterministic Continuous Replacement: Fast and Stable Module Replacement in Pretrained Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Bradbury, Rowan, Ashok, Aniket Srinivasan, Kasanagottu, Sai Ram, Jhingran, Gunmay, Meng, Shuai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Point-RTD: Replaced Token Denoising for Pretraining Transformer Models on Point Clouds
by: Stone, Gunner, et al.
Published: (2025)
by: Stone, Gunner, et al.
Published: (2025)
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
by: Chatzoudis, Gerasimos, et al.
Published: (2026)
by: Chatzoudis, Gerasimos, et al.
Published: (2026)
PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
by: Picard, David, et al.
Published: (2026)
by: Picard, David, et al.
Published: (2026)
Replace in Translation: Boost Concept Alignment in Counterfactual Text-to-Image
by: Li, Sifan, et al.
Published: (2025)
by: Li, Sifan, et al.
Published: (2025)
An Inpainting-Infused Pipeline for Attire and Background Replacement
by: Perche-Mahlow, Felipe Rodrigues, et al.
Published: (2024)
by: Perche-Mahlow, Felipe Rodrigues, et al.
Published: (2024)
Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
by: Zeng, Ziyun, et al.
Published: (2026)
by: Zeng, Ziyun, et al.
Published: (2026)
FAR: Function-preserving Attention Replacement for IMC-friendly Inference
by: Ren, Yuxin, et al.
Published: (2025)
by: Ren, Yuxin, et al.
Published: (2025)
Can Vision-Language Models Replace Human Annotators: A Case Study with CelebA Dataset
by: Lu, Haoming, et al.
Published: (2024)
by: Lu, Haoming, et al.
Published: (2024)
FastTab: A Fast Table Recognizer with a Tiny Recursive Module and 1D Transformers
by: Hamdi, Laziz, et al.
Published: (2026)
by: Hamdi, Laziz, et al.
Published: (2026)
RASA: Replace Anyone, Say Anything -- A Training-Free Framework for Audio-Driven and Universal Portrait Video Editing
by: Pan, Tianrui, et al.
Published: (2025)
by: Pan, Tianrui, et al.
Published: (2025)
Fast Vision Mamba: Pooling Spatial Dimensions for Accelerated Processing
by: Kapse, Saarthak, et al.
Published: (2025)
by: Kapse, Saarthak, et al.
Published: (2025)
Composite Classifier-Free Guidance for Multi-Modal Conditioning in Wind Dynamics Super-Resolution
by: Schnell, Jacob, et al.
Published: (2025)
by: Schnell, Jacob, et al.
Published: (2025)
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
by: Teja, L. D. M. S. Sai, et al.
Published: (2025)
by: Teja, L. D. M. S. Sai, et al.
Published: (2025)
ReplaceAnything3D:Text-Guided 3D Scene Editing with Compositional Neural Radiance Fields
by: Bartrum, Edward, et al.
Published: (2024)
by: Bartrum, Edward, et al.
Published: (2024)
Learning Transferable Features for Implicit Neural Representations
by: Vyas, Kushal, et al.
Published: (2024)
by: Vyas, Kushal, et al.
Published: (2024)
FactoFormer: Factorized Hyperspectral Transformers with Self-Supervised Pretraining
by: Mohamed, Shaheer, et al.
Published: (2023)
by: Mohamed, Shaheer, et al.
Published: (2023)
DifuzCam: Replacing Camera Lens with a Mask and a Diffusion Model
by: Yosef, Erez, et al.
Published: (2024)
by: Yosef, Erez, et al.
Published: (2024)
Replace-then-Perturb: Targeted Adversarial Attacks With Visual Reasoning for Vision-Language Models
by: Jang, Jonggyu, et al.
Published: (2024)
by: Jang, Jonggyu, et al.
Published: (2024)
Few-Shot Deployment of Pretrained MRI Transformers in Brain Imaging Tasks
by: Li, Mengyu, et al.
Published: (2025)
by: Li, Mengyu, et al.
Published: (2025)
Fast OTSU Thresholding Using Bisection Method
by: Kodathala, Sai Varun
Published: (2025)
by: Kodathala, Sai Varun
Published: (2025)
HQ-DiT: Efficient Diffusion Transformer with FP4 Hybrid Quantization
by: Liu, Wenxuan, et al.
Published: (2024)
by: Liu, Wenxuan, et al.
Published: (2024)
FastTrackTr:Towards Fast Multi-Object Tracking with Transformers
by: Liao, Pan, et al.
Published: (2024)
by: Liao, Pan, et al.
Published: (2024)
MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive Pretraining
by: Liu, Yunze, et al.
Published: (2024)
by: Liu, Yunze, et al.
Published: (2024)
DDT: Decoupled Diffusion Transformer
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
CornViT: A Multi-Stage Convolutional Vision Transformer Framework for Hierarchical Corn Kernel Analysis
by: Erukude, Sai Teja, et al.
Published: (2025)
by: Erukude, Sai Teja, et al.
Published: (2025)
Latent Modulated Function for Computational Optimal Continuous Image Representation
by: He, Zongyao, et al.
Published: (2024)
by: He, Zongyao, et al.
Published: (2024)
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models
by: Horawalavithana, Sameera, et al.
Published: (2026)
by: Horawalavithana, Sameera, et al.
Published: (2026)
Attention Retention for Continual Learning with Vision Transformers
by: Lu, Yue, et al.
Published: (2026)
by: Lu, Yue, et al.
Published: (2026)
Patch Rebirth: Toward Fast and Transferable Model Inversion of Vision Transformers
by: Heo, Seongsoo, et al.
Published: (2025)
by: Heo, Seongsoo, et al.
Published: (2025)
ANYPORTAL: Zero-Shot Consistent Video Background Replacement
by: Gao, Wenshuo, et al.
Published: (2025)
by: Gao, Wenshuo, et al.
Published: (2025)
Vision-and-Language Navigation Generative Pretrained Transformer
by: Hanlin, Wen
Published: (2024)
by: Hanlin, Wen
Published: (2024)
Satellite to Street : Disaster Impact Estimator
by: Sai, Sreesritha, et al.
Published: (2025)
by: Sai, Sreesritha, et al.
Published: (2025)
Efficient Zero-Shot AI-Generated Image Detection
by: Sonoda, Ryosuke, et al.
Published: (2026)
by: Sonoda, Ryosuke, et al.
Published: (2026)
Channel Vision Transformers: An Image Is Worth 1 x 16 x 16 Words
by: Bao, Yujia, et al.
Published: (2023)
by: Bao, Yujia, et al.
Published: (2023)
Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers
by: Roschmann, Simon, et al.
Published: (2025)
by: Roschmann, Simon, et al.
Published: (2025)
Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
by: Sinha, Rohit, et al.
Published: (2026)
by: Sinha, Rohit, et al.
Published: (2026)
AdaVid: Adaptive Video-Language Pretraining
by: Patel, Chaitanya, et al.
Published: (2025)
by: Patel, Chaitanya, et al.
Published: (2025)
Separators in Enhancing Autoregressive Pretraining for Vision Mamba
by: Liu, Hanpeng, et al.
Published: (2026)
by: Liu, Hanpeng, et al.
Published: (2026)
Spatial Transcriptomics as Images for Large-Scale Pretraining
by: Zhu, Yishun, et al.
Published: (2026)
by: Zhu, Yishun, et al.
Published: (2026)
Beyond Random Augmentations: Pretraining with Hard Views
by: Ferreira, Fabio, et al.
Published: (2023)
by: Ferreira, Fabio, et al.
Published: (2023)
Similar Items
-
Point-RTD: Replaced Token Denoising for Pretraining Transformer Models on Point Clouds
by: Stone, Gunner, et al.
Published: (2025) -
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
by: Chatzoudis, Gerasimos, et al.
Published: (2026) -
PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
by: Picard, David, et al.
Published: (2026) -
Replace in Translation: Boost Concept Alignment in Counterfactual Text-to-Image
by: Li, Sifan, et al.
Published: (2025) -
An Inpainting-Infused Pipeline for Attire and Background Replacement
by: Perche-Mahlow, Felipe Rodrigues, et al.
Published: (2024)