You Actually Look Twice At it (YALTAi): using an object detection approach instead of region segmentation within the Kraken engine
Fuente:
arXiv
Salvato in:
| Autore principale: | Clérice, Thibault |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2022
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
di: Karamolegkou, Antonia, et al.
Pubblicazione: (2026)
di: Karamolegkou, Antonia, et al.
Pubblicazione: (2026)
Look Twice: Training-Free Evidence Highlighting in Multimodal Large Language Models
di: Morini, Marco, et al.
Pubblicazione: (2026)
di: Morini, Marco, et al.
Pubblicazione: (2026)
FusionVision: A comprehensive approach of 3D object reconstruction and segmentation from RGB-D cameras using YOLO and fast segment anything
di: Ghazouali, Safouane El, et al.
Pubblicazione: (2024)
di: Ghazouali, Safouane El, et al.
Pubblicazione: (2024)
Deep learning approaches to surgical video segmentation and object detection: A Scoping Review
di: Kamtam, Devanish N., et al.
Pubblicazione: (2025)
di: Kamtam, Devanish N., et al.
Pubblicazione: (2025)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
di: Chen, Kaitao, et al.
Pubblicazione: (2025)
di: Chen, Kaitao, et al.
Pubblicazione: (2025)
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning
di: Li, Yian, et al.
Pubblicazione: (2024)
di: Li, Yian, et al.
Pubblicazione: (2024)
An efficient plant disease detection using transfer learning approach
di: Sambana, Bosubabu, et al.
Pubblicazione: (2025)
di: Sambana, Bosubabu, et al.
Pubblicazione: (2025)
Discovering and using Spelke segments
di: Venkatesh, Rahul, et al.
Pubblicazione: (2025)
di: Venkatesh, Rahul, et al.
Pubblicazione: (2025)
Technical note: ShinyAnimalCV: open-source cloud-based web application for object detection, segmentation, and three-dimensional visualization of animals using computer vision
di: Wang, Jin, et al.
Pubblicazione: (2023)
di: Wang, Jin, et al.
Pubblicazione: (2023)
Scene-aware SAR ship detection guided by unsupervised sea-land segmentation
di: Ke, Han, et al.
Pubblicazione: (2025)
di: Ke, Han, et al.
Pubblicazione: (2025)
Study of detecting behavioral signatures within DeepFake videos
di: Miao, Qiaomu, et al.
Pubblicazione: (2022)
di: Miao, Qiaomu, et al.
Pubblicazione: (2022)
Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection
di: Li, Xiang, et al.
Pubblicazione: (2025)
di: Li, Xiang, et al.
Pubblicazione: (2025)
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
di: Tang, Fei, et al.
Pubblicazione: (2025)
di: Tang, Fei, et al.
Pubblicazione: (2025)
Detecting Sexual Content at the Sentence Level in First Millennium Latin Texts
di: Clérice, Thibault
Pubblicazione: (2023)
di: Clérice, Thibault
Pubblicazione: (2023)
Agricultural Object Detection with You Look Only Once (YOLO) Algorithm: A Bibliometric and Systematic Literature Review
di: Badgujar, Chetan M, et al.
Pubblicazione: (2024)
di: Badgujar, Chetan M, et al.
Pubblicazione: (2024)
You Only Look Twice! for Failure Causes Identification of Drill Bits
di: Yamani, Asma, et al.
Pubblicazione: (2024)
di: Yamani, Asma, et al.
Pubblicazione: (2024)
Improving segmentation of retinal arteries and veins using cardiac signal in doppler holograms
di: Dubosc, Marius, et al.
Pubblicazione: (2025)
di: Dubosc, Marius, et al.
Pubblicazione: (2025)
Region of interest detection for efficient aortic segmentation
di: Giordano, Loris, et al.
Pubblicazione: (2026)
di: Giordano, Loris, et al.
Pubblicazione: (2026)
GCA-ResUNet:Image segmentation in medical images using grouped coordinate attention
di: Ding, Jun, et al.
Pubblicazione: (2025)
di: Ding, Jun, et al.
Pubblicazione: (2025)
What are You Looking at? Modality Contribution in Multimodal Medical Deep Learning
di: Gapp, Christian, et al.
Pubblicazione: (2025)
di: Gapp, Christian, et al.
Pubblicazione: (2025)
Mamba meets crack segmentation
di: He, Zhili, et al.
Pubblicazione: (2024)
di: He, Zhili, et al.
Pubblicazione: (2024)
Barlow-Swin: Toward a novel siamese-based segmentation architecture using Swin-Transformers
di: Haftlang, Morteza Kiani, et al.
Pubblicazione: (2025)
di: Haftlang, Morteza Kiani, et al.
Pubblicazione: (2025)
BFA-YOLO: A balanced multiscale object detection network for building façade attachments detection
di: Chen, Yangguang, et al.
Pubblicazione: (2024)
di: Chen, Yangguang, et al.
Pubblicazione: (2024)
Underwater object detection in sonar imagery with detection transformer and Zero-shot neural architecture search
di: Gu, XiaoTong, et al.
Pubblicazione: (2025)
di: Gu, XiaoTong, et al.
Pubblicazione: (2025)
Batch Transformer: Look for Attention in Batch
di: Her, Myung Beom, et al.
Pubblicazione: (2024)
di: Her, Myung Beom, et al.
Pubblicazione: (2024)
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
di: Shen, Yuxiang, et al.
Pubblicazione: (2026)
di: Shen, Yuxiang, et al.
Pubblicazione: (2026)
A multimodal deep learning architecture for smoking detection with a small data approach
di: Lakatos, Robert, et al.
Pubblicazione: (2023)
di: Lakatos, Robert, et al.
Pubblicazione: (2023)
Is SAM3 ready for pathology segmentation?
di: Kong, Qiuyu, et al.
Pubblicazione: (2026)
di: Kong, Qiuyu, et al.
Pubblicazione: (2026)
Eye image segmentation using visual and concept prompts with Segment Anything Model 3 (SAM3)
di: Niehorster, Diederick C., et al.
Pubblicazione: (2026)
di: Niehorster, Diederick C., et al.
Pubblicazione: (2026)
Active learning for efficient annotation in precision agriculture: a use-case on crop-weed semantic segmentation
di: van Marrewijk, Bart M., et al.
Pubblicazione: (2024)
di: van Marrewijk, Bart M., et al.
Pubblicazione: (2024)
Polyp detection in colonoscopy images using YOLOv11
di: Sahoo, Alok Ranjan, et al.
Pubblicazione: (2025)
di: Sahoo, Alok Ranjan, et al.
Pubblicazione: (2025)
Multi-head automated segmentation by incorporating detection head into the contextual layer neural network
di: Kys, Edwin, et al.
Pubblicazione: (2026)
di: Kys, Edwin, et al.
Pubblicazione: (2026)
Semi-supervised segmentation of land cover images using nonlinear canonical correlation analysis with multiple features and t-SNE
di: Wei, Hong, et al.
Pubblicazione: (2024)
di: Wei, Hong, et al.
Pubblicazione: (2024)
I Am Big, You Are Little; I Am Right, You Are Wrong
di: Kelly, David A., et al.
Pubblicazione: (2025)
di: Kelly, David A., et al.
Pubblicazione: (2025)
Looking into Concept Explanation Methods for Diabetic Retinopathy Classification
di: Storås, Andrea M., et al.
Pubblicazione: (2024)
di: Storås, Andrea M., et al.
Pubblicazione: (2024)
MLN-net: A multi-source medical image segmentation method for clustered microcalcifications using multiple layer normalization
di: Wang, Ke, et al.
Pubblicazione: (2023)
di: Wang, Ke, et al.
Pubblicazione: (2023)
Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding
di: Murlidaran, Shravan, et al.
Pubblicazione: (2026)
di: Murlidaran, Shravan, et al.
Pubblicazione: (2026)
Enhancing kelp forest detection in remote sensing images using crowdsourced labels with Mixed Vision Transformers and ConvNeXt segmentation models
di: Nasios, Ioannis
Pubblicazione: (2025)
di: Nasios, Ioannis
Pubblicazione: (2025)
Think Twice before Adaptation: Improving Adaptability of DeepFake Detection via Online Test-Time Adaptation
di: Nguyen-Le, Hong-Hanh, et al.
Pubblicazione: (2025)
di: Nguyen-Le, Hong-Hanh, et al.
Pubblicazione: (2025)
Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding
di: Wang, Wenkai, et al.
Pubblicazione: (2026)
di: Wang, Wenkai, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
di: Karamolegkou, Antonia, et al.
Pubblicazione: (2026) -
Look Twice: Training-Free Evidence Highlighting in Multimodal Large Language Models
di: Morini, Marco, et al.
Pubblicazione: (2026) -
FusionVision: A comprehensive approach of 3D object reconstruction and segmentation from RGB-D cameras using YOLO and fast segment anything
di: Ghazouali, Safouane El, et al.
Pubblicazione: (2024) -
Deep learning approaches to surgical video segmentation and object detection: A Scoping Review
di: Kamtam, Devanish N., et al.
Pubblicazione: (2025) -
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
di: Chen, Kaitao, et al.
Pubblicazione: (2025)