FVO: Fast Visual Odometry with Transformers
Fuente:
arXiv
Guardado en:
| Autores principales: | Yugay, Vlardimir, Nguyen, Duy-Kien, Gevers, Theo, Snoek, Cees G. M., Oswald, Martin R. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation
por: Nguyen, Duy-Kien, et al.
Publicado: (2023)
por: Nguyen, Duy-Kien, et al.
Publicado: (2023)
MAGiC-SLAM: Multi-Agent Gaussian Globally Consistent SLAM
por: Yugay, Vladimir, et al.
Publicado: (2024)
por: Yugay, Vladimir, et al.
Publicado: (2024)
Gaussian-SLAM: Photo-realistic Dense SLAM with Gaussian Splatting
por: Yugay, Vladimir, et al.
Publicado: (2023)
por: Yugay, Vladimir, et al.
Publicado: (2023)
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
por: Nguyen, Duy-Kien, et al.
Publicado: (2024)
por: Nguyen, Duy-Kien, et al.
Publicado: (2024)
R-MAE: Regions Meet Masked Autoencoders
por: Nguyen, Duy-Kien, et al.
Publicado: (2023)
por: Nguyen, Duy-Kien, et al.
Publicado: (2023)
Gaussian Mapping for Evolving Scenes
por: Yugay, Vladimir, et al.
Publicado: (2025)
por: Yugay, Vladimir, et al.
Publicado: (2025)
FewViewGS: Gaussian Splatting with Few View Matching and Multi-stage Training
por: Yin, Ruihong, et al.
Publicado: (2024)
por: Yin, Ruihong, et al.
Publicado: (2024)
Union-over-Intersections: Object Detection beyond Winner-Takes-All
por: Bhowmik, Aritra, et al.
Publicado: (2023)
por: Bhowmik, Aritra, et al.
Publicado: (2023)
T-MAE: Temporal Masked Autoencoders for Point Cloud Representation Learning
por: Wei, Weijie, et al.
Publicado: (2023)
por: Wei, Weijie, et al.
Publicado: (2023)
Edge-Centric Relational Reasoning for 3D Scene Graph Prediction
por: Ma, Yanni, et al.
Publicado: (2025)
por: Ma, Yanni, et al.
Publicado: (2025)
3D-AVS: LiDAR-based 3D Auto-Vocabulary Segmentation
por: Wei, Weijie, et al.
Publicado: (2024)
por: Wei, Weijie, et al.
Publicado: (2024)
SuperDisco: Super-Class Discovery Improves Visual Recognition for the Long-Tail
por: Du, Yingjun, et al.
Publicado: (2023)
por: Du, Yingjun, et al.
Publicado: (2023)
TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning
por: Bhowmik, Aritra, et al.
Publicado: (2024)
por: Bhowmik, Aritra, et al.
Publicado: (2024)
Low-Resource Vision Challenges for Foundation Models
por: Zhang, Yunhua, et al.
Publicado: (2024)
por: Zhang, Yunhua, et al.
Publicado: (2024)
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
por: Sun, Wenfang, et al.
Publicado: (2026)
por: Sun, Wenfang, et al.
Publicado: (2026)
Fast SceneScript: Fast and Accurate Language-Based 3D Scene Understanding via Multi-Token Prediction
por: Yin, Ruihong, et al.
Publicado: (2025)
por: Yin, Ruihong, et al.
Publicado: (2025)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
por: Liu, Huabin, et al.
Publicado: (2025)
por: Liu, Huabin, et al.
Publicado: (2025)
Loopy-SLAM: Dense Neural SLAM with Loop Closures
por: Liso, Lorenzo, et al.
Publicado: (2024)
por: Liso, Lorenzo, et al.
Publicado: (2024)
NeoBabel: A Multilingual Open Tower for Visual Generation
por: Derakhshani, Mohammad Mahdi, et al.
Publicado: (2025)
por: Derakhshani, Mohammad Mahdi, et al.
Publicado: (2025)
Dual Guidance Semi-Supervised Action Detection
por: Singh, Ankit, et al.
Publicado: (2025)
por: Singh, Ankit, et al.
Publicado: (2025)
Grab-3D: Detecting AI-Generated Videos from 3D Geometric Temporal Consistency
por: Chen, Wenhan, et al.
Publicado: (2025)
por: Chen, Wenhan, et al.
Publicado: (2025)
Modeling Weather Uncertainty for Multi-weather Co-Presence Estimation
por: Bi, Qi, et al.
Publicado: (2024)
por: Bi, Qi, et al.
Publicado: (2024)
Geometry-guided Feature Learning and Fusion for Indoor Scene Reconstruction
por: Yin, Ruihong, et al.
Publicado: (2024)
por: Yin, Ruihong, et al.
Publicado: (2024)
IPO: Interpretable Prompt Optimization for Vision-Language Models
por: Du, Yingjun, et al.
Publicado: (2024)
por: Du, Yingjun, et al.
Publicado: (2024)
Segment Any 3D-Part in a Scene from a Sentence
por: Wu, Hongyu, et al.
Publicado: (2025)
por: Wu, Hongyu, et al.
Publicado: (2025)
PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs
por: Dorkenwald, Michael, et al.
Publicado: (2024)
por: Dorkenwald, Michael, et al.
Publicado: (2024)
LocoMotion: Learning Motion-Focused Video-Language Representations
por: Doughty, Hazel, et al.
Publicado: (2024)
por: Doughty, Hazel, et al.
Publicado: (2024)
Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
por: Wang, Ziqi, et al.
Publicado: (2025)
por: Wang, Ziqi, et al.
Publicado: (2025)
Stronger Semantic Encoders Can Harm Relighting Performance: Probing Visual Priors via Augmented Latent Intrinsics
por: Xing, Xiaoyan, et al.
Publicado: (2026)
por: Xing, Xiaoyan, et al.
Publicado: (2026)
MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
por: Bhowmik, Aritra, et al.
Publicado: (2025)
por: Bhowmik, Aritra, et al.
Publicado: (2025)
Training-Free Semantic Segmentation via LLM-Supervision
por: Sun, Wenfang, et al.
Publicado: (2024)
por: Sun, Wenfang, et al.
Publicado: (2024)
Beyond Coarse-Grained Matching in Video-Text Retrieval
por: Chen, Aozhu, et al.
Publicado: (2024)
por: Chen, Aozhu, et al.
Publicado: (2024)
Unblur-SLAM: Dense Neural SLAM for Blurry Inputs
por: Zhang, Qi, et al.
Publicado: (2026)
por: Zhang, Qi, et al.
Publicado: (2026)
Elastic ViTs from Pretrained Models without Retraining
por: Simoncini, Walter, et al.
Publicado: (2025)
por: Simoncini, Walter, et al.
Publicado: (2025)
Redefining Normal: A Novel Object-Level Approach for Multi-Object Novelty Detection
por: Salehi, Mohammadreza, et al.
Publicado: (2024)
por: Salehi, Mohammadreza, et al.
Publicado: (2024)
Lost in Time: A New Temporal Benchmark for VideoLLMs
por: Cores, Daniel, et al.
Publicado: (2024)
por: Cores, Daniel, et al.
Publicado: (2024)
Any-Shift Prompting for Generalization over Distributions
por: Xiao, Zehao, et al.
Publicado: (2024)
por: Xiao, Zehao, et al.
Publicado: (2024)
Intrinsic Image Decomposition Using Point Cloud Representation
por: Xing, Xiaoyan, et al.
Publicado: (2023)
por: Xing, Xiaoyan, et al.
Publicado: (2023)
Ray-Distance Volume Rendering for Neural Scene Reconstruction
por: Yin, Ruihong, et al.
Publicado: (2024)
por: Yin, Ruihong, et al.
Publicado: (2024)
VirDA: Reusing Backbone for Unsupervised Domain Adaptation with Visual Reprogramming
por: Nguyen, Duy, et al.
Publicado: (2025)
por: Nguyen, Duy, et al.
Publicado: (2025)
Ejemplares similares
-
SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation
por: Nguyen, Duy-Kien, et al.
Publicado: (2023) -
MAGiC-SLAM: Multi-Agent Gaussian Globally Consistent SLAM
por: Yugay, Vladimir, et al.
Publicado: (2024) -
Gaussian-SLAM: Photo-realistic Dense SLAM with Gaussian Splatting
por: Yugay, Vladimir, et al.
Publicado: (2023) -
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
por: Nguyen, Duy-Kien, et al.
Publicado: (2024) -
R-MAE: Regions Meet Masked Autoencoders
por: Nguyen, Duy-Kien, et al.
Publicado: (2023)