Saved in:
Bibliographic Details
Main Author: Carrillo-Perez, Borja
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.22942
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910247398932480
author Carrillo-Perez, Borja
author_facet Carrillo-Perez, Borja
contents This report presents a lightweight modification to the DETR-based fusion transformer baseline for the MaCVi 2026 Vision-to-Chart data association challenge. The challenge baseline decoder receives per-buoy queries encoding world-space distance and bearing, forcing the transformer to implicitly learn the complex geometric projection from world coordinates to image pixels. Instead, this work trains an additional dedicated MLP, QueryMLP, to explicitly predict the buoy's waterline contact point in the image from chart measurements and IMU orientation data. The predicted pixel coordinates are appended to the baseline decoder query vector, providing a direct spatial prior per buoy and reducing the geometric reasoning burden on the transformer decoder. On the challenge leaderboard, the presented approach achieves an Overall score of 0.7386, with F1 = 0.8055 and mIoU = 0.6718, on the held-out test set, placing second among all submissions.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22942
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Improved Vision-to-Chart Buoy Association with Learned World-to-Image Projection
Carrillo-Perez, Borja
Computer Vision and Pattern Recognition
This report presents a lightweight modification to the DETR-based fusion transformer baseline for the MaCVi 2026 Vision-to-Chart data association challenge. The challenge baseline decoder receives per-buoy queries encoding world-space distance and bearing, forcing the transformer to implicitly learn the complex geometric projection from world coordinates to image pixels. Instead, this work trains an additional dedicated MLP, QueryMLP, to explicitly predict the buoy's waterline contact point in the image from chart measurements and IMU orientation data. The predicted pixel coordinates are appended to the baseline decoder query vector, providing a direct spatial prior per buoy and reducing the geometric reasoning burden on the transformer decoder. On the challenge leaderboard, the presented approach achieves an Overall score of 0.7386, with F1 = 0.8055 and mIoU = 0.6718, on the held-out test set, placing second among all submissions.
title Improved Vision-to-Chart Buoy Association with Learned World-to-Image Projection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.22942