Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Shinnick, Zachary, Jiang, Liangze, Saratchandran, Hemanth, Teney, Damien, Hengel, Anton van den |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Procedural Pretraining: Warming Up Language Models with Abstract Data
by: Jiang, Liangze, et al.
Published: (2026)
by: Jiang, Liangze, et al.
Published: (2026)
Transformers Pretrained on Procedural Data Contain Modular Structures for Algorithmic Reasoning
by: Shinnick, Zachary, et al.
Published: (2025)
by: Shinnick, Zachary, et al.
Published: (2025)
Leaner Transformers: More Heads, Less Depth
by: Saratchandran, Hemanth, et al.
Published: (2025)
by: Saratchandran, Hemanth, et al.
Published: (2025)
Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the Wild
by: Teney, Damien, et al.
Published: (2025)
by: Teney, Damien, et al.
Published: (2025)
Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product
by: Albert, Paul, et al.
Published: (2025)
by: Albert, Paul, et al.
Published: (2025)
Enhancing Transformers Through Conditioned Embedded Tokens
by: Saratchandran, Hemanth, et al.
Published: (2025)
by: Saratchandran, Hemanth, et al.
Published: (2025)
Structured Initialization for Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2025)
by: Zheng, Jianqiao, et al.
Published: (2025)
RandLoRA: Full-rank parameter-efficient fine-tuning of large models
by: Albert, Paul, et al.
Published: (2025)
by: Albert, Paul, et al.
Published: (2025)
SineProject: Machine Unlearning for Stable Vision Language Alignment
by: Garg, Arpit, et al.
Published: (2025)
by: Garg, Arpit, et al.
Published: (2025)
OOD-Chameleon: Is Algorithm Selection for OOD Generalization Learnable?
by: Jiang, Liangze, et al.
Published: (2024)
by: Jiang, Liangze, et al.
Published: (2024)
Synergy and Diversity in CLIP: Enhancing Performance Through Adaptive Backbone Ensembling
by: Rodriguez-Opazo, Cristian, et al.
Published: (2024)
by: Rodriguez-Opazo, Cristian, et al.
Published: (2024)
AugGen: Synthetic Augmentation using Diffusion Models Can Improve Recognition
by: Rahimi, Parsa, et al.
Published: (2025)
by: Rahimi, Parsa, et al.
Published: (2025)
From Activation to Initialization: Scaling Insights for Optimizing Neural Fields
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Weight Conditioning for Smooth Optimization of Neural Networks
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Hierarchical Process Reward Models are Symbolic Vision Learners
by: Zhang, Shan, et al.
Published: (2025)
by: Zhang, Shan, et al.
Published: (2025)
Preconditioners for the Stochastic Training of Neural Fields
by: Chng, Shin-Fang, et al.
Published: (2024)
by: Chng, Shin-Fang, et al.
Published: (2024)
Invertible Neural Warp for NeRF
by: Chng, Shin-Fang, et al.
Published: (2024)
by: Chng, Shin-Fang, et al.
Published: (2024)
Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting
by: Xu, Runze, et al.
Published: (2026)
by: Xu, Runze, et al.
Published: (2026)
Always Skip Attention
by: Ji, Yiping, et al.
Published: (2025)
by: Ji, Yiping, et al.
Published: (2025)
LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation
by: Miao, Bo, et al.
Published: (2026)
by: Miao, Bo, et al.
Published: (2026)
Premonition: Using Generative Models to Preempt Future Data Changes in Continual Learning
by: McDonnell, Mark D., et al.
Published: (2024)
by: McDonnell, Mark D., et al.
Published: (2024)
Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors
by: Xia, Jiatong, et al.
Published: (2026)
by: Xia, Jiatong, et al.
Published: (2026)
Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder
by: Liu, Zheyuan, et al.
Published: (2023)
by: Liu, Zheyuan, et al.
Published: (2023)
Scaling up Multi-domain Semantic Segmentation with Sentence Embeddings
by: Yin, Wei, et al.
Published: (2022)
by: Yin, Wei, et al.
Published: (2022)
Efficient Learning With Sine-Activated Low-rank Matrices
by: Ji, Yiping, et al.
Published: (2024)
by: Ji, Yiping, et al.
Published: (2024)
Let Your Video Listen to Your Music!
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Preconditioned Attention: Enhancing Efficiency in Transformers
by: Saratchandran, Hemanth
Published: (2026)
by: Saratchandran, Hemanth
Published: (2026)
RanPAC: Random Projections and Pre-trained Models for Continual Learning
by: McDonnell, Mark D., et al.
Published: (2023)
by: McDonnell, Mark D., et al.
Published: (2023)
Frame-wise Conditioning Adaptation for Fine-Tuning Diffusion Models in Text-to-Video Prediction
by: Liu, Zheyuan, et al.
Published: (2025)
by: Liu, Zheyuan, et al.
Published: (2025)
From Tables to Signals: Revealing Spectral Adaptivity in TabPFN
by: Zheng, Jianqiao, et al.
Published: (2025)
by: Zheng, Jianqiao, et al.
Published: (2025)
D'OH: Decoder-Only Random Hypernetworks for Implicit Neural Representations
by: Gordon, Cameron, et al.
Published: (2024)
by: Gordon, Cameron, et al.
Published: (2024)
On the Value of Cross-Modal Misalignment in Multimodal Representation Learning
by: Cai, Yichao, et al.
Published: (2025)
by: Cai, Yichao, et al.
Published: (2025)
Continual Learning on CLIP via Incremental Prompt Tuning with Intrinsic Textual Anchors
by: Lu, Haodong, et al.
Published: (2025)
by: Lu, Haodong, et al.
Published: (2025)
How Well Can Vision Language Models See Image Details?
by: Gou, Chenhui, et al.
Published: (2024)
by: Gou, Chenhui, et al.
Published: (2024)
CAM-Based Methods Can See through Walls
by: Taimeskhanov, Magamed, et al.
Published: (2024)
by: Taimeskhanov, Magamed, et al.
Published: (2024)
ViewFusion: Towards Multi-View Consistency via Interpolated Denoising
by: Yang, Xianghui, et al.
Published: (2024)
by: Yang, Xianghui, et al.
Published: (2024)
Decorrelation Speeds Up Vision Transformers
by: Carrigg, Kieran, et al.
Published: (2025)
by: Carrigg, Kieran, et al.
Published: (2025)
Knowledge Composition using Task Vectors with Learned Anisotropic Scaling
by: Zhang, Frederic Z., et al.
Published: (2024)
by: Zhang, Frederic Z., et al.
Published: (2024)
Vision-Language Models Can't See the Obvious
by: Dahou, Yasser, et al.
Published: (2025)
by: Dahou, Yasser, et al.
Published: (2025)
Similar Items
-
Procedural Pretraining: Warming Up Language Models with Abstract Data
by: Jiang, Liangze, et al.
Published: (2026) -
Transformers Pretrained on Procedural Data Contain Modular Structures for Algorithmic Reasoning
by: Shinnick, Zachary, et al.
Published: (2025) -
Leaner Transformers: More Heads, Less Depth
by: Saratchandran, Hemanth, et al.
Published: (2025) -
Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the Wild
by: Teney, Damien, et al.
Published: (2025) -
Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product
by: Albert, Paul, et al.
Published: (2025)