Understanding Multi-View Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Stary, Michal, Gaubil, Julien, Tewari, Ayush, Sitzmann, Vincent |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Generative View Stitching
di: Song, Chonghyuk, et al.
Pubblicazione: (2025)
di: Song, Chonghyuk, et al.
Pubblicazione: (2025)
True Self-Supervised Novel View Synthesis is Transferable
di: Mitchel, Thomas W., et al.
Pubblicazione: (2025)
di: Mitchel, Thomas W., et al.
Pubblicazione: (2025)
Neural Isometries: Taming Transformations for Equivariant ML
di: Mitchel, Thomas W., et al.
Pubblicazione: (2024)
di: Mitchel, Thomas W., et al.
Pubblicazione: (2024)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
di: Cazenavette, George, et al.
Pubblicazione: (2025)
di: Cazenavette, George, et al.
Pubblicazione: (2025)
Scaling View Synthesis Transformers
di: Kim, Evan, et al.
Pubblicazione: (2026)
di: Kim, Evan, et al.
Pubblicazione: (2026)
GTA: A Geometry-Aware Attention Mechanism for Multi-View Transformers
di: Miyato, Takeru, et al.
Pubblicazione: (2023)
di: Miyato, Takeru, et al.
Pubblicazione: (2023)
Turning Video Models into Generalist Robot Policies
di: Li, Sizhe Lester, et al.
Pubblicazione: (2026)
di: Li, Sizhe Lester, et al.
Pubblicazione: (2026)
FlowMap: High-Quality Camera Poses, Intrinsics, and Depth via Gradient Descent
di: Smith, Cameron, et al.
Pubblicazione: (2024)
di: Smith, Cameron, et al.
Pubblicazione: (2024)
Understanding Multimodal Deep Neural Networks: A Concept Selection View
di: Shang, Chenming, et al.
Pubblicazione: (2024)
di: Shang, Chenming, et al.
Pubblicazione: (2024)
Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
di: Kišš, Martin, et al.
Pubblicazione: (2025)
di: Kišš, Martin, et al.
Pubblicazione: (2025)
I-Segmenter: Integer-Only Vision Transformer for Efficient Semantic Segmentation
di: Sassoon, Jordan, et al.
Pubblicazione: (2025)
di: Sassoon, Jordan, et al.
Pubblicazione: (2025)
RETR: Multi-View Radar Detection Transformer for Indoor Perception
di: Yataka, Ryoma, et al.
Pubblicazione: (2024)
di: Yataka, Ryoma, et al.
Pubblicazione: (2024)
Multi-View Hypercomplex Learning for Breast Cancer Screening
di: Lopez, Eleonora, et al.
Pubblicazione: (2022)
di: Lopez, Eleonora, et al.
Pubblicazione: (2022)
Random Token Fusion for Multi-View Medical Diagnosis
di: Guo, Jingyu, et al.
Pubblicazione: (2024)
di: Guo, Jingyu, et al.
Pubblicazione: (2024)
Refine and Align: Confidence Calibration through Multi-Agent Interaction in VQA
di: Pandey, Ayush, et al.
Pubblicazione: (2025)
di: Pandey, Ayush, et al.
Pubblicazione: (2025)
Early evidence of how LLMs outperform traditional systems on OCR/HTR tasks for historical records
di: Kim, Seorin, et al.
Pubblicazione: (2025)
di: Kim, Seorin, et al.
Pubblicazione: (2025)
Multi-View and Multi-Scale Alignment for Contrastive Language-Image Pre-training in Mammography
di: Du, Yuexi, et al.
Pubblicazione: (2024)
di: Du, Yuexi, et al.
Pubblicazione: (2024)
3D-Consistent Multi-View Editing by Correspondence Guidance
di: Bengtson, Josef, et al.
Pubblicazione: (2025)
di: Bengtson, Josef, et al.
Pubblicazione: (2025)
Multi-View 3D Reconstruction using Knowledge Distillation
di: Dutt, Aditya, et al.
Pubblicazione: (2024)
di: Dutt, Aditya, et al.
Pubblicazione: (2024)
Towards Robust Uncertainty-Aware Incomplete Multi-View Classification
di: Chen, Mulin, et al.
Pubblicazione: (2024)
di: Chen, Mulin, et al.
Pubblicazione: (2024)
GLAM: Geometry-Guided Local Alignment for Multi-View VLP in Mammography
di: Du, Yuexi, et al.
Pubblicazione: (2025)
di: Du, Yuexi, et al.
Pubblicazione: (2025)
Ranking vs. Assignment: The Metric Mismatch in Multi-View Object Association
di: Shelukhan, Matvei, et al.
Pubblicazione: (2026)
di: Shelukhan, Matvei, et al.
Pubblicazione: (2026)
Hypergraph-based Multi-View Action Recognition using Event Cameras
di: Gao, Yue, et al.
Pubblicazione: (2024)
di: Gao, Yue, et al.
Pubblicazione: (2024)
Conformal Trajectory Prediction with Multi-View Data Integration in Cooperative Driving
di: Chen, Xi, et al.
Pubblicazione: (2024)
di: Chen, Xi, et al.
Pubblicazione: (2024)
Uncertainty Quantification via Hölder Divergence for Multi-View Representation Learning
di: Zhang, Yan, et al.
Pubblicazione: (2024)
di: Zhang, Yan, et al.
Pubblicazione: (2024)
ULTra: Unveiling Latent Token Interpretability in Transformer-Based Understanding and Segmentation
di: Hosseini, Hesam, et al.
Pubblicazione: (2024)
di: Hosseini, Hesam, et al.
Pubblicazione: (2024)
MuM: Multi-View Masked Image Modeling for 3D Vision
di: Nordström, David, et al.
Pubblicazione: (2025)
di: Nordström, David, et al.
Pubblicazione: (2025)
InkFM: A Foundational Model for Full-Page Online Handwritten Note Understanding
di: Fadeeva, Anastasiia, et al.
Pubblicazione: (2025)
di: Fadeeva, Anastasiia, et al.
Pubblicazione: (2025)
Deep Models for Multi-View 3D Object Recognition: A Review
di: Alzahrani, Mona, et al.
Pubblicazione: (2024)
di: Alzahrani, Mona, et al.
Pubblicazione: (2024)
PickScan: Object discovery and reconstruction from handheld interactions
di: van der Brugge, Vincent, et al.
Pubblicazione: (2024)
di: van der Brugge, Vincent, et al.
Pubblicazione: (2024)
Do Transformers Understand Ancient Roman Coin Motifs Better than CNNs?
di: Reid, David, et al.
Pubblicazione: (2026)
di: Reid, David, et al.
Pubblicazione: (2026)
Bird Eye-View to Street-View: A Survey
di: Bajbaa, Khawlah, et al.
Pubblicazione: (2024)
di: Bajbaa, Khawlah, et al.
Pubblicazione: (2024)
Missing Data as Augmentation in the Earth Observation Domain: A Multi-View Learning Approach
di: Mena, Francisco, et al.
Pubblicazione: (2025)
di: Mena, Francisco, et al.
Pubblicazione: (2025)
DreamComposer: Controllable 3D Object Generation via Multi-View Conditions
di: Yang, Yunhan, et al.
Pubblicazione: (2023)
di: Yang, Yunhan, et al.
Pubblicazione: (2023)
Frugal Federated Learning for Violence Detection: A Comparison of LoRA-Tuned VLMs and Personalized CNNs
di: Thuau, Sébastien, et al.
Pubblicazione: (2025)
di: Thuau, Sébastien, et al.
Pubblicazione: (2025)
MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
di: Liu, Ziyu, et al.
Pubblicazione: (2024)
di: Liu, Ziyu, et al.
Pubblicazione: (2024)
Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated Perturbation
di: Xu, Jie, et al.
Pubblicazione: (2025)
di: Xu, Jie, et al.
Pubblicazione: (2025)
Understanding Video Transformers via Universal Concept Discovery
di: Kowal, Matthew, et al.
Pubblicazione: (2024)
di: Kowal, Matthew, et al.
Pubblicazione: (2024)
MFAF: An EVA02-Based Multi-scale Frequency Attention Fusion Method for Cross-View Geo-Localization
di: Liu, YiTong, et al.
Pubblicazione: (2025)
di: Liu, YiTong, et al.
Pubblicazione: (2025)
ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery
di: Shrivastava, Ayush, et al.
Pubblicazione: (2026)
di: Shrivastava, Ayush, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Generative View Stitching
di: Song, Chonghyuk, et al.
Pubblicazione: (2025) -
True Self-Supervised Novel View Synthesis is Transferable
di: Mitchel, Thomas W., et al.
Pubblicazione: (2025) -
Neural Isometries: Taming Transformations for Equivariant ML
di: Mitchel, Thomas W., et al.
Pubblicazione: (2024) -
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
di: Cazenavette, George, et al.
Pubblicazione: (2025) -
Scaling View Synthesis Transformers
di: Kim, Evan, et al.
Pubblicazione: (2026)