Win-Win: Training High-Resolution Vision Transformers from Two Windows
Fuente:
arXiv
Saved in:
| Main Authors: | Leroy, Vincent, Revaud, Jerome, Lucas, Thomas, Weinzaepfel, Philippe |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors
by: Jang, Wonbong, et al.
Published: (2025)
by: Jang, Wonbong, et al.
Published: (2025)
MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion
by: Duisterhof, Bardienus, et al.
Published: (2024)
by: Duisterhof, Bardienus, et al.
Published: (2024)
Grounding Image Matching in 3D with MASt3R
by: Leroy, Vincent, et al.
Published: (2024)
by: Leroy, Vincent, et al.
Published: (2024)
DUSt3R: Geometric 3D Vision Made Easy
by: Wang, Shuzhe, et al.
Published: (2023)
by: Wang, Shuzhe, et al.
Published: (2023)
S-MUSt3R: Sliding Multi-view 3D Reconstruction
by: Antsfeld, Leonid, et al.
Published: (2026)
by: Antsfeld, Leonid, et al.
Published: (2026)
MUSt3R: Multi-view Network for Stereo 3D Reconstruction
by: Cabon, Yohann, et al.
Published: (2025)
by: Cabon, Yohann, et al.
Published: (2025)
RouteWinFormer: A Route-Window Transformer for Middle-range Attention in Image Restoration
by: Li, Qifan, et al.
Published: (2025)
by: Li, Qifan, et al.
Published: (2025)
WinSyn: A High Resolution Testbed for Synthetic Data
by: Kelly, Tom, et al.
Published: (2023)
by: Kelly, Tom, et al.
Published: (2023)
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens
by: Guo, Yiwei, et al.
Published: (2026)
by: Guo, Yiwei, et al.
Published: (2026)
Revealing the Two Sides of Data Augmentation: An Asymmetric Distillation-based Win-Win Solution for Open-Set Recognition
by: Jia, Yunbing, et al.
Published: (2024)
by: Jia, Yunbing, et al.
Published: (2024)
Cross-view and Cross-pose Completion for 3D Human Understanding
by: Armando, Matthieu, et al.
Published: (2023)
by: Armando, Matthieu, et al.
Published: (2023)
HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction
by: Rojas, Sara, et al.
Published: (2025)
by: Rojas, Sara, et al.
Published: (2025)
Random Wins All: Rethinking Grouping Strategies for Vision Tokens
by: Fan, Qihang, et al.
Published: (2026)
by: Fan, Qihang, et al.
Published: (2026)
WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
by: Li, Zizun, et al.
Published: (2025)
by: Li, Zizun, et al.
Published: (2025)
Evaluation of Winning Solutions of 2025 Low Power Computer Vision Challenge
by: Ye, Zihao, et al.
Published: (2026)
by: Ye, Zihao, et al.
Published: (2026)
WinMamba: Multi-Scale Shifted Windows in State Space Model for 3D Object Detection
by: Zheng, Longhui, et al.
Published: (2025)
by: Zheng, Longhui, et al.
Published: (2025)
WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments
by: Zhao, Haoren, et al.
Published: (2026)
by: Zhao, Haoren, et al.
Published: (2026)
UNIC: Universal Classification Models via Multi-teacher Distillation
by: Sariyildiz, Mert Bulent, et al.
Published: (2024)
by: Sariyildiz, Mert Bulent, et al.
Published: (2024)
PoseScript: Linking 3D Human Poses and Natural Language
by: Delmas, Ginger, et al.
Published: (2022)
by: Delmas, Ginger, et al.
Published: (2022)
DaWin: Training-free Dynamic Weight Interpolation for Robust Adaptation
by: Oh, Changdae, et al.
Published: (2024)
by: Oh, Changdae, et al.
Published: (2024)
HOSt3R: Keypoint-free Hand-Object 3D Reconstruction from RGB images
by: Swamy, Anilkumar, et al.
Published: (2025)
by: Swamy, Anilkumar, et al.
Published: (2025)
Co-Win: Joint Object Detection and Instance Segmentation in LiDAR Point Clouds via Collaborative Window Processing
by: Li, Haichuan, et al.
Published: (2025)
by: Li, Haichuan, et al.
Published: (2025)
The Championship-Winning Solution for the 5th CLVISION Challenge 2024
by: Pan, Sishun, et al.
Published: (2024)
by: Pan, Sishun, et al.
Published: (2024)
What does really matter in image goal navigation?
by: Monaci, Gianluca, et al.
Published: (2025)
by: Monaci, Gianluca, et al.
Published: (2025)
Purposer: Putting Human Motion Generation in Context
by: Ugrinovic, Nicolas, et al.
Published: (2024)
by: Ugrinovic, Nicolas, et al.
Published: (2024)
PhaseWin Search Framework Enable Efficient Object-Level Interpretation
by: Gu, Zihan, et al.
Published: (2025)
by: Gu, Zihan, et al.
Published: (2025)
CAIT: Triple-Win Compression towards High Accuracy, Fast Inference, and Favorable Transferability For ViTs
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks
by: Kowsher, Md, et al.
Published: (2025)
by: Kowsher, Md, et al.
Published: (2025)
Geospatial Foundational Embedder: Top-1 Winning Solution on EarthVision Embed2Scale Challenge (CVPR 2025)
by: Xu, Zirui, et al.
Published: (2025)
by: Xu, Zirui, et al.
Published: (2025)
CondiMen: Conditional Multi-Person Mesh Recovery
by: Romain, Brégier, et al.
Published: (2024)
by: Romain, Brégier, et al.
Published: (2024)
Multi-HMR: Multi-Person Whole-Body Human Mesh Recovery in a Single Shot
by: Baradel, Fabien, et al.
Published: (2024)
by: Baradel, Fabien, et al.
Published: (2024)
Layer Pruning with Consensus: A Triple-Win Solution
by: Mugnaini, Leandro Giusti, et al.
Published: (2024)
by: Mugnaini, Leandro Giusti, et al.
Published: (2024)
Predicting Winning Captions for Weekly New Yorker Comics
by: Cao, Stanley, et al.
Published: (2024)
by: Cao, Stanley, et al.
Published: (2024)
Structured Semantic 3D Reconstruction (S23DR) Challenge 2025 -- Winning solution
by: Skvrna, Jan, et al.
Published: (2025)
by: Skvrna, Jan, et al.
Published: (2025)
Winning the Lottery by Preserving Network Training Dynamics with Concrete Ticket Search
by: Arora, Tanay, et al.
Published: (2025)
by: Arora, Tanay, et al.
Published: (2025)
PoseFix: Correcting 3D Human Poses with Natural Language
by: Delmas, Ginger, et al.
Published: (2023)
by: Delmas, Ginger, et al.
Published: (2023)
PoseEmbroider: Towards a 3D, Visual, Semantic-aware Human Pose Representation
by: Delmas, Ginger, et al.
Published: (2024)
by: Delmas, Ginger, et al.
Published: (2024)
DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers
by: Sariyildiz, Mert Bulent, et al.
Published: (2025)
by: Sariyildiz, Mert Bulent, et al.
Published: (2025)
Human Mesh Modeling for Anny Body
by: Brégier, Romain, et al.
Published: (2025)
by: Brégier, Romain, et al.
Published: (2025)
MVP: Winning Solution to SMP Challenge 2025 Video Track
by: Ye, Liliang, et al.
Published: (2025)
by: Ye, Liliang, et al.
Published: (2025)
Similar Items
-
Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors
by: Jang, Wonbong, et al.
Published: (2025) -
MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion
by: Duisterhof, Bardienus, et al.
Published: (2024) -
Grounding Image Matching in 3D with MASt3R
by: Leroy, Vincent, et al.
Published: (2024) -
DUSt3R: Geometric 3D Vision Made Easy
by: Wang, Shuzhe, et al.
Published: (2023) -
S-MUSt3R: Sliding Multi-view 3D Reconstruction
by: Antsfeld, Leonid, et al.
Published: (2026)