Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
Fuente:
arXiv
Saved in:
| Main Authors: | Bond, Andrew, Melanlioglu, Ilkin Umut, Erdem, Erkut, Erdem, Aykut |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
by: Çapuk, Hakan, et al.
Published: (2025)
by: Çapuk, Hakan, et al.
Published: (2025)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
by: Bond, Andrew, et al.
Published: (2025)
by: Bond, Andrew, et al.
Published: (2025)
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
by: Karanfil, Enes, et al.
Published: (2025)
by: Karanfil, Enes, et al.
Published: (2025)
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
by: Ercan, Burak, et al.
Published: (2024)
by: Ercan, Burak, et al.
Published: (2024)
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
by: Sanli, Enes, et al.
Published: (2025)
by: Sanli, Enes, et al.
Published: (2025)
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
by: Cokelek, Mert, et al.
Published: (2025)
by: Cokelek, Mert, et al.
Published: (2025)
VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs
by: Ali, Moayed Haji, et al.
Published: (2023)
by: Ali, Moayed Haji, et al.
Published: (2023)
HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2024)
by: Dogan, Mustafa, et al.
Published: (2024)
CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models
by: Ekin, Yigit, et al.
Published: (2024)
by: Ekin, Yigit, et al.
Published: (2024)
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2025)
by: Kizil, Muhammed Burak, et al.
Published: (2025)
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
by: Anees, Abdul Basit, et al.
Published: (2024)
by: Anees, Abdul Basit, et al.
Published: (2024)
SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
by: Biner, Burak Can, et al.
Published: (2024)
by: Biner, Burak Can, et al.
Published: (2024)
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2026)
by: Kizil, Muhammed Burak, et al.
Published: (2026)
ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
by: Li, Zhaoyang, et al.
Published: (2025)
by: Li, Zhaoyang, et al.
Published: (2025)
Reconstruction of Optical Coherence Tomography Images from Wavelength-space Using Deep-learning
by: Viqar, Maryam, et al.
Published: (2025)
by: Viqar, Maryam, et al.
Published: (2025)
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks
by: Hu, Yuanze, et al.
Published: (2025)
by: Hu, Yuanze, et al.
Published: (2025)
SubspaceAD: Training-Free Few-Shot Anomaly Detection via Subspace Modeling
by: Lendering, Camile, et al.
Published: (2026)
by: Lendering, Camile, et al.
Published: (2026)
Positional Encodings Anchor Spatial Structure in Vision Transformers: A Geometric Perspective on Robustness
by: Mannes, Mahmoud
Published: (2026)
by: Mannes, Mahmoud
Published: (2026)
Soft Mixture Denoising: Beyond the Expressive Bottleneck of Diffusion Models
by: Li, Yangming, et al.
Published: (2023)
by: Li, Yangming, et al.
Published: (2023)
VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance
by: Srivastava, Divyansh, et al.
Published: (2024)
by: Srivastava, Divyansh, et al.
Published: (2024)
Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders
by: Kuo, Shang-Jui Ray, et al.
Published: (2026)
by: Kuo, Shang-Jui Ray, et al.
Published: (2026)
FuseFormer: A Transformer for Visual and Thermal Image Fusion
by: Erdogan, Aytekin, et al.
Published: (2024)
by: Erdogan, Aytekin, et al.
Published: (2024)
Aligning Latent Spaces with Flow Priors
by: Li, Yizhuo, et al.
Published: (2025)
by: Li, Yizhuo, et al.
Published: (2025)
A Comparative Survey of Vision Transformers for Feature Extraction in Texture Analysis
by: Scabini, Leonardo, et al.
Published: (2024)
by: Scabini, Leonardo, et al.
Published: (2024)
FlashKAT: Understanding and Addressing Performance Bottlenecks in the Kolmogorov-Arnold Transformer
by: Raffel, Matthew, et al.
Published: (2025)
by: Raffel, Matthew, et al.
Published: (2025)
Unveiling the Visual Counting Bottleneck in Vision-Language Models
by: Pang, Xingzhou, et al.
Published: (2026)
by: Pang, Xingzhou, et al.
Published: (2026)
Student Capacity Moderates Knowledge Distillation Effectiveness: A Systematic Study Across ResNet Teacher-Student Pairs on CIFAR-10
by: Yasar, Umut Onur
Published: (2026)
by: Yasar, Umut Onur
Published: (2026)
Anisotropic Fourier Features for Positional Encoding in Medical Imaging
by: Jabareen, Nabil, et al.
Published: (2025)
by: Jabareen, Nabil, et al.
Published: (2025)
Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization
by: Xing, Shuo, et al.
Published: (2025)
by: Xing, Shuo, et al.
Published: (2025)
RespoDiff: Dual-Module Bottleneck Transformation for Responsible & Faithful T2I Generation
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2025)
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2025)
Topological Alignment of Shared Vision-Language Embedding Space
by: You, Junwon, et al.
Published: (2025)
by: You, Junwon, et al.
Published: (2025)
Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
by: Øhrstrøm, Christoffer Koo, et al.
Published: (2026)
by: Øhrstrøm, Christoffer Koo, et al.
Published: (2026)
LASERS: LAtent Space Encoding for Representations with Sparsity for Generative Modeling
by: Li, Xin, et al.
Published: (2024)
by: Li, Xin, et al.
Published: (2024)
Hybrid Convolution and Vision Transformer NAS Search Space for TinyML Image Classification
by: Djajapermana, Mikhael, et al.
Published: (2025)
by: Djajapermana, Mikhael, et al.
Published: (2025)
Exploring Challenges in Deep Learning of Single-Station Ground Motion Records
by: Çağlar, Ümit Mert, et al.
Published: (2024)
by: Çağlar, Ümit Mert, et al.
Published: (2024)
Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion Models
by: Ling, Huan, et al.
Published: (2023)
by: Ling, Huan, et al.
Published: (2023)
Improving Position Encoding of Transformers for Multivariate Time Series Classification
by: Foumani, Navid Mohammadi, et al.
Published: (2023)
by: Foumani, Navid Mohammadi, et al.
Published: (2023)
EaqVLA: Encoding-aligned Quantization for Vision-Language-Action Models
by: Jiang, Feng, et al.
Published: (2025)
by: Jiang, Feng, et al.
Published: (2025)
Similar Items
-
TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
by: Çapuk, Hakan, et al.
Published: (2025) -
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
by: Bond, Andrew, et al.
Published: (2025) -
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
by: Karanfil, Enes, et al.
Published: (2025) -
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
by: Ercan, Burak, et al.
Published: (2024) -
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
by: Ercan, Burak, et al.
Published: (2023)