Saved in:
| Main Authors: | Schäfer, Frederik, Mandl, Luis, Kälber, Lars, Ricken, Tim |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.05908 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-supervised pretraining for an iterative image size agnostic vision transformer
by: Prisadnikov, Nedyalko, et al.
Published: (2026)
by: Prisadnikov, Nedyalko, et al.
Published: (2026)
an interpretable vision transformer framework for automated brain tumor classification
by: Mbonu, Chinedu Emmanuel, et al.
Published: (2026)
by: Mbonu, Chinedu Emmanuel, et al.
Published: (2026)
HSFusion: A high-level vision task-driven infrared and visible image fusion network via semantic and geometric domain transformation
by: Jiang, Chengjie, et al.
Published: (2024)
by: Jiang, Chengjie, et al.
Published: (2024)
A novel network for classification of cuneiform tablet metadata
by: Hagelskjær, Frederik
Published: (2026)
by: Hagelskjær, Frederik
Published: (2026)
TREAD: Token Routing for Efficient Architecture-agnostic Diffusion Training
by: Krause, Felix, et al.
Published: (2025)
by: Krause, Felix, et al.
Published: (2025)
Architecture and evaluation protocol for transformer-based visual object tracking in UAV applications
by: Borne, Augustin, et al.
Published: (2026)
by: Borne, Augustin, et al.
Published: (2026)
A hierarchical semantic segmentation framework for computer vision-based bridge damage detection
by: Liu, Jingxiao, et al.
Published: (2022)
by: Liu, Jingxiao, et al.
Published: (2022)
Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
by: Huang, Weijian, et al.
Published: (2024)
by: Huang, Weijian, et al.
Published: (2024)
On the effectiveness of multimodal privileged knowledge distillation in two vision transformer based diagnostic applications
by: Baur, Simon, et al.
Published: (2025)
by: Baur, Simon, et al.
Published: (2025)
STARS: Sensor-agnostic Transformer Architecture for Remote Sensing
by: King, Ethan, et al.
Published: (2024)
by: King, Ethan, et al.
Published: (2024)
Separable DeepONet: Breaking the Curse of Dimensionality in Physics-Informed Machine Learning
by: Mandl, Luis, et al.
Published: (2024)
by: Mandl, Luis, et al.
Published: (2024)
Physics-Informed Time-Integrated DeepONet: Temporal Tangent Space Operator Learning for High-Accuracy Inference
by: Mandl, Luis, et al.
Published: (2025)
by: Mandl, Luis, et al.
Published: (2025)
MinkOcc: Towards real-time label-efficient semantic occupancy prediction
by: Sze, Samuel, et al.
Published: (2025)
by: Sze, Samuel, et al.
Published: (2025)
NormalView: sensor-agnostic tree species classification from backpack and aerial lidar data using geometric projections
by: Korkeala, Juho, et al.
Published: (2025)
by: Korkeala, Juho, et al.
Published: (2025)
A comprehensive overview of deep learning techniques for 3D point cloud classification and semantic segmentation
by: Sarker, Sushmita, et al.
Published: (2024)
by: Sarker, Sushmita, et al.
Published: (2024)
Real-time 3D semantic occupancy prediction for autonomous vehicles using memory-efficient sparse convolution
by: Sze, Samuel, et al.
Published: (2024)
by: Sze, Samuel, et al.
Published: (2024)
Cross multiscale vision transformer for deep fake detection
by: P, Akhshan, et al.
Published: (2025)
by: P, Akhshan, et al.
Published: (2025)
Automated diagnosis of lung diseases using vision transformer: a comparative study on chest x-ray classification
by: Ahmad, Muhammad, et al.
Published: (2025)
by: Ahmad, Muhammad, et al.
Published: (2025)
Beyond the final layer: Attentive multilayer fusion for vision transformers
by: Ciernik, Laure, et al.
Published: (2026)
by: Ciernik, Laure, et al.
Published: (2026)
Category-aware EEG image generation based on wavelet transform and contrast semantic loss
by: Zhang, Enshang, et al.
Published: (2025)
by: Zhang, Enshang, et al.
Published: (2025)
Depth-agnostic Single Image Dehazing
by: Xu, Honglei, et al.
Published: (2024)
by: Xu, Honglei, et al.
Published: (2024)
Towards Domain-agnostic Depth Completion
by: Xu, Guangkai, et al.
Published: (2022)
by: Xu, Guangkai, et al.
Published: (2022)
FISHing in Uncertainty: Synthetic Contrastive Learning for Genetic Aberration Detection
by: Gutwein, Simon, et al.
Published: (2024)
by: Gutwein, Simon, et al.
Published: (2024)
Interpreting vision transformers via residual replacement model
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
A survey on efficient vision transformers: algorithms, techniques, and performance benchmarking
by: Papa, Lorenzo, et al.
Published: (2023)
by: Papa, Lorenzo, et al.
Published: (2023)
METER: a mobile vision transformer architecture for monocular depth estimation
by: Papa, L., et al.
Published: (2024)
by: Papa, L., et al.
Published: (2024)
Steering CLIP's vision transformer with sparse autoencoders
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
SPRINT: Script-agnostic Structure Recognition in Tables
by: Kudale, Dhruv, et al.
Published: (2025)
by: Kudale, Dhruv, et al.
Published: (2025)
Clothing agnostic Pre-inpainting Virtual Try-ON
by: Kim, Sehyun, et al.
Published: (2025)
by: Kim, Sehyun, et al.
Published: (2025)
Scene-agnostic Pose Regression for Visual Localization
by: Zheng, Junwei, et al.
Published: (2025)
by: Zheng, Junwei, et al.
Published: (2025)
SAFER: Sharpness Aware layer-selective Finetuning for Enhanced Robustness in vision transformers
by: Gopal, Bhavna, et al.
Published: (2025)
by: Gopal, Bhavna, et al.
Published: (2025)
Low-latency vision transformers via large-scale multi-head attention
by: Gross, Ronit D., et al.
Published: (2025)
by: Gross, Ronit D., et al.
Published: (2025)
Initialization matters in few-shot adaptation of vision-language models for histopathological image classification
by: Meseguer, Pablo, et al.
Published: (2026)
by: Meseguer, Pablo, et al.
Published: (2026)
Are vision-language models ready to zero-shot replace supervised classification models in agriculture?
by: Ranario, Earl, et al.
Published: (2025)
by: Ranario, Earl, et al.
Published: (2025)
On the application of the Wasserstein metric to 2D curves classification
by: Kaliszewska, Agnieszka, et al.
Published: (2026)
by: Kaliszewska, Agnieszka, et al.
Published: (2026)
PIV3CAMS: a multi-camera dataset for multiple computer vision problems and its application to novel view-point synthesis
by: Kim, Sohyeong, et al.
Published: (2024)
by: Kim, Sohyeong, et al.
Published: (2024)
$L^3$:Scene-agnostic Visual Localization in the Wild
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
Preserving Marker Specificity with Lightweight Channel-Independent Representation Learning
by: Gutwein, Simon, et al.
Published: (2025)
by: Gutwein, Simon, et al.
Published: (2025)
PTQ4ViT: Post-training quantization for vision transformers with twin uniform quantization
by: Yuan, Zhihang, et al.
Published: (2021)
by: Yuan, Zhihang, et al.
Published: (2021)
Two-Stream temporal transformer for video action classification
by: Kurpukdee, Nattapong, et al.
Published: (2026)
by: Kurpukdee, Nattapong, et al.
Published: (2026)
Similar Items
-
Self-supervised pretraining for an iterative image size agnostic vision transformer
by: Prisadnikov, Nedyalko, et al.
Published: (2026) -
an interpretable vision transformer framework for automated brain tumor classification
by: Mbonu, Chinedu Emmanuel, et al.
Published: (2026) -
HSFusion: A high-level vision task-driven infrared and visible image fusion network via semantic and geometric domain transformation
by: Jiang, Chengjie, et al.
Published: (2024) -
A novel network for classification of cuneiform tablet metadata
by: Hagelskjær, Frederik
Published: (2026) -
TREAD: Token Routing for Efficient Architecture-agnostic Diffusion Training
by: Krause, Felix, et al.
Published: (2025)