Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bevli, Aviraj, Chaybouti, Sofian, Dahou, Yasser, Hacid, Hakim, Huynh, Ngoc Dung, Khac, Phuc H. Le, Narayan, Sanath, Para, Wamiq Reyaz, Singh, Ankit
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2603.27365
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917365765111808
author Bevli, Aviraj
Chaybouti, Sofian
Dahou, Yasser
Hacid, Hakim
Huynh, Ngoc Dung
Khac, Phuc H. Le
Narayan, Sanath
Para, Wamiq Reyaz
Singh, Ankit
author_facet Bevli, Aviraj
Chaybouti, Sofian
Dahou, Yasser
Hacid, Hakim
Huynh, Ngoc Dung
Khac, Phuc H. Le
Narayan, Sanath
Para, Wamiq Reyaz
Singh, Ankit
contents Perception-centric systems are typically implemented with a modular encoder-decoder pipeline: a vision backbone for feature extraction and a separate decoder (or late-fusion module) for task prediction. This raises a central question: is this architectural separation essential or can a single early-fusion stack do both perception and task modeling at scale? We introduce Falcon Perception, a unified dense Transformer that processes image patches and text tokens in a shared parameter space from the first layer, using a hybrid attention pattern (bidirectional among image tokens, causal for prediction tokens) to combine global visual context with autoregressive, variable-length instance generation. To keep dense outputs practical, Falcon Perception retains a lightweight token interface and decodes continuous spatial outputs with specialized heads, enabling parallel high-resolution mask prediction. Our design promotes simplicity: we keep a single scalable backbone and shift complexity toward data and training signals, adding only small heads where outputs are continuous and dense. On SA-Co, Falcon Perception improves mask quality to 68.0 Macro-F$_1$ compared to 62.3 of SAM3. We also introduce PBench, a benchmark targeting compositional prompts (OCR, spatial constraints, relations) and dense long-context regimes, where the model shows better gains. Finally, we extend the same early-fusion recipe to Falcon OCR: a compact 300M-parameter model which attains 80.3% on olmOCR and 88.64 on OmniDocBench.
format Preprint
id arxiv_https___arxiv_org_abs_2603_27365
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Falcon Perception
Bevli, Aviraj
Chaybouti, Sofian
Dahou, Yasser
Hacid, Hakim
Huynh, Ngoc Dung
Khac, Phuc H. Le
Narayan, Sanath
Para, Wamiq Reyaz
Singh, Ankit
Computer Vision and Pattern Recognition
Perception-centric systems are typically implemented with a modular encoder-decoder pipeline: a vision backbone for feature extraction and a separate decoder (or late-fusion module) for task prediction. This raises a central question: is this architectural separation essential or can a single early-fusion stack do both perception and task modeling at scale? We introduce Falcon Perception, a unified dense Transformer that processes image patches and text tokens in a shared parameter space from the first layer, using a hybrid attention pattern (bidirectional among image tokens, causal for prediction tokens) to combine global visual context with autoregressive, variable-length instance generation. To keep dense outputs practical, Falcon Perception retains a lightweight token interface and decodes continuous spatial outputs with specialized heads, enabling parallel high-resolution mask prediction. Our design promotes simplicity: we keep a single scalable backbone and shift complexity toward data and training signals, adding only small heads where outputs are continuous and dense. On SA-Co, Falcon Perception improves mask quality to 68.0 Macro-F$_1$ compared to 62.3 of SAM3. We also introduce PBench, a benchmark targeting compositional prompts (OCR, spatial constraints, relations) and dense long-context regimes, where the model shows better gains. Finally, we extend the same early-fusion recipe to Falcon OCR: a compact 300M-parameter model which attains 80.3% on olmOCR and 88.64 on OmniDocBench.
title Falcon Perception
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.27365