FVO: Fast Visual Odometry with Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yugay, Vlardimir, Nguyen, Duy-Kien, Gevers, Theo, Snoek, Cees G. M., Oswald, Martin R. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation
von: Nguyen, Duy-Kien, et al.
Veröffentlicht: (2023)
von: Nguyen, Duy-Kien, et al.
Veröffentlicht: (2023)
MAGiC-SLAM: Multi-Agent Gaussian Globally Consistent SLAM
von: Yugay, Vladimir, et al.
Veröffentlicht: (2024)
von: Yugay, Vladimir, et al.
Veröffentlicht: (2024)
Gaussian-SLAM: Photo-realistic Dense SLAM with Gaussian Splatting
von: Yugay, Vladimir, et al.
Veröffentlicht: (2023)
von: Yugay, Vladimir, et al.
Veröffentlicht: (2023)
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
von: Nguyen, Duy-Kien, et al.
Veröffentlicht: (2024)
von: Nguyen, Duy-Kien, et al.
Veröffentlicht: (2024)
R-MAE: Regions Meet Masked Autoencoders
von: Nguyen, Duy-Kien, et al.
Veröffentlicht: (2023)
von: Nguyen, Duy-Kien, et al.
Veröffentlicht: (2023)
Gaussian Mapping for Evolving Scenes
von: Yugay, Vladimir, et al.
Veröffentlicht: (2025)
von: Yugay, Vladimir, et al.
Veröffentlicht: (2025)
FewViewGS: Gaussian Splatting with Few View Matching and Multi-stage Training
von: Yin, Ruihong, et al.
Veröffentlicht: (2024)
von: Yin, Ruihong, et al.
Veröffentlicht: (2024)
Union-over-Intersections: Object Detection beyond Winner-Takes-All
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2023)
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2023)
T-MAE: Temporal Masked Autoencoders for Point Cloud Representation Learning
von: Wei, Weijie, et al.
Veröffentlicht: (2023)
von: Wei, Weijie, et al.
Veröffentlicht: (2023)
Edge-Centric Relational Reasoning for 3D Scene Graph Prediction
von: Ma, Yanni, et al.
Veröffentlicht: (2025)
von: Ma, Yanni, et al.
Veröffentlicht: (2025)
3D-AVS: LiDAR-based 3D Auto-Vocabulary Segmentation
von: Wei, Weijie, et al.
Veröffentlicht: (2024)
von: Wei, Weijie, et al.
Veröffentlicht: (2024)
SuperDisco: Super-Class Discovery Improves Visual Recognition for the Long-Tail
von: Du, Yingjun, et al.
Veröffentlicht: (2023)
von: Du, Yingjun, et al.
Veröffentlicht: (2023)
TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2024)
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2024)
Low-Resource Vision Challenges for Foundation Models
von: Zhang, Yunhua, et al.
Veröffentlicht: (2024)
von: Zhang, Yunhua, et al.
Veröffentlicht: (2024)
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
von: Sun, Wenfang, et al.
Veröffentlicht: (2026)
von: Sun, Wenfang, et al.
Veröffentlicht: (2026)
Fast SceneScript: Fast and Accurate Language-Based 3D Scene Understanding via Multi-Token Prediction
von: Yin, Ruihong, et al.
Veröffentlicht: (2025)
von: Yin, Ruihong, et al.
Veröffentlicht: (2025)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
von: Liu, Huabin, et al.
Veröffentlicht: (2025)
von: Liu, Huabin, et al.
Veröffentlicht: (2025)
Loopy-SLAM: Dense Neural SLAM with Loop Closures
von: Liso, Lorenzo, et al.
Veröffentlicht: (2024)
von: Liso, Lorenzo, et al.
Veröffentlicht: (2024)
NeoBabel: A Multilingual Open Tower for Visual Generation
von: Derakhshani, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Derakhshani, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
Dual Guidance Semi-Supervised Action Detection
von: Singh, Ankit, et al.
Veröffentlicht: (2025)
von: Singh, Ankit, et al.
Veröffentlicht: (2025)
Grab-3D: Detecting AI-Generated Videos from 3D Geometric Temporal Consistency
von: Chen, Wenhan, et al.
Veröffentlicht: (2025)
von: Chen, Wenhan, et al.
Veröffentlicht: (2025)
Modeling Weather Uncertainty for Multi-weather Co-Presence Estimation
von: Bi, Qi, et al.
Veröffentlicht: (2024)
von: Bi, Qi, et al.
Veröffentlicht: (2024)
Geometry-guided Feature Learning and Fusion for Indoor Scene Reconstruction
von: Yin, Ruihong, et al.
Veröffentlicht: (2024)
von: Yin, Ruihong, et al.
Veröffentlicht: (2024)
IPO: Interpretable Prompt Optimization for Vision-Language Models
von: Du, Yingjun, et al.
Veröffentlicht: (2024)
von: Du, Yingjun, et al.
Veröffentlicht: (2024)
Segment Any 3D-Part in a Scene from a Sentence
von: Wu, Hongyu, et al.
Veröffentlicht: (2025)
von: Wu, Hongyu, et al.
Veröffentlicht: (2025)
PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs
von: Dorkenwald, Michael, et al.
Veröffentlicht: (2024)
von: Dorkenwald, Michael, et al.
Veröffentlicht: (2024)
LocoMotion: Learning Motion-Focused Video-Language Representations
von: Doughty, Hazel, et al.
Veröffentlicht: (2024)
von: Doughty, Hazel, et al.
Veröffentlicht: (2024)
Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
von: Wang, Ziqi, et al.
Veröffentlicht: (2025)
von: Wang, Ziqi, et al.
Veröffentlicht: (2025)
Stronger Semantic Encoders Can Harm Relighting Performance: Probing Visual Priors via Augmented Latent Intrinsics
von: Xing, Xiaoyan, et al.
Veröffentlicht: (2026)
von: Xing, Xiaoyan, et al.
Veröffentlicht: (2026)
MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2025)
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2025)
Training-Free Semantic Segmentation via LLM-Supervision
von: Sun, Wenfang, et al.
Veröffentlicht: (2024)
von: Sun, Wenfang, et al.
Veröffentlicht: (2024)
Beyond Coarse-Grained Matching in Video-Text Retrieval
von: Chen, Aozhu, et al.
Veröffentlicht: (2024)
von: Chen, Aozhu, et al.
Veröffentlicht: (2024)
Unblur-SLAM: Dense Neural SLAM for Blurry Inputs
von: Zhang, Qi, et al.
Veröffentlicht: (2026)
von: Zhang, Qi, et al.
Veröffentlicht: (2026)
Elastic ViTs from Pretrained Models without Retraining
von: Simoncini, Walter, et al.
Veröffentlicht: (2025)
von: Simoncini, Walter, et al.
Veröffentlicht: (2025)
Redefining Normal: A Novel Object-Level Approach for Multi-Object Novelty Detection
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
Lost in Time: A New Temporal Benchmark for VideoLLMs
von: Cores, Daniel, et al.
Veröffentlicht: (2024)
von: Cores, Daniel, et al.
Veröffentlicht: (2024)
Any-Shift Prompting for Generalization over Distributions
von: Xiao, Zehao, et al.
Veröffentlicht: (2024)
von: Xiao, Zehao, et al.
Veröffentlicht: (2024)
Intrinsic Image Decomposition Using Point Cloud Representation
von: Xing, Xiaoyan, et al.
Veröffentlicht: (2023)
von: Xing, Xiaoyan, et al.
Veröffentlicht: (2023)
Ray-Distance Volume Rendering for Neural Scene Reconstruction
von: Yin, Ruihong, et al.
Veröffentlicht: (2024)
von: Yin, Ruihong, et al.
Veröffentlicht: (2024)
VirDA: Reusing Backbone for Unsupervised Domain Adaptation with Visual Reprogramming
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation
von: Nguyen, Duy-Kien, et al.
Veröffentlicht: (2023) -
MAGiC-SLAM: Multi-Agent Gaussian Globally Consistent SLAM
von: Yugay, Vladimir, et al.
Veröffentlicht: (2024) -
Gaussian-SLAM: Photo-realistic Dense SLAM with Gaussian Splatting
von: Yugay, Vladimir, et al.
Veröffentlicht: (2023) -
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
von: Nguyen, Duy-Kien, et al.
Veröffentlicht: (2024) -
R-MAE: Regions Meet Masked Autoencoders
von: Nguyen, Duy-Kien, et al.
Veröffentlicht: (2023)