LF Tracy: A Unified Single-Pipeline Approach for Salient Object Detection in Light Field Cameras

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Teng, Fei, Zhang, Jiaming, Liu, Jiawei, Peng, Kunyu, Cheng, Xina, Li, Zhiyong, Yang, Kailun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917758280663040
author Teng, Fei
Zhang, Jiaming
Liu, Jiawei
Peng, Kunyu
Cheng, Xina
Li, Zhiyong
Yang, Kailun
author_facet Teng, Fei
Zhang, Jiaming
Liu, Jiawei
Peng, Kunyu
Cheng, Xina
Li, Zhiyong
Yang, Kailun
contents Leveraging rich information is crucial for dense prediction tasks. Light field (LF) cameras are instrumental in this regard, as they allow data to be sampled from various perspectives. This capability provides valuable spatial, depth, and angular information, enhancing scene-parsing tasks. However, we have identified two overlooked issues for the LF salient object detection (SOD) task. (1): Previous approaches predominantly employ a customized two-stream design to discover the spatial and depth features within light field images. The network struggles to learn the implicit angular information between different images due to a lack of intra-network data connectivity. (2): Little research has been directed towards the data augmentation strategy for LF SOD. Research on inter-network data connectivity is scant. In this study, we propose an efficient paradigm (LF Tracy) to address those issues. This comprises a single-pipeline encoder paired with a highly efficient information aggregation (IA) module (around 8M parameters) to establish an intra-network connection. Then, a simple yet effective data augmentation strategy called MixLD is designed to bridge the inter-network connections. Owing to this innovative paradigm, our model surpasses the existing state-of-the-art method through extensive experiments. Especially, LF Tracy demonstrates a 23% improvement over previous results on the latest large-scale PKU dataset. The source code is publicly available at: https://github.com/FeiBryantkit/LF-Tracy.
format Preprint
id arxiv_https___arxiv_org_abs_2401_16712
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LF Tracy: A Unified Single-Pipeline Approach for Salient Object Detection in Light Field Cameras
Teng, Fei
Zhang, Jiaming
Liu, Jiawei
Peng, Kunyu
Cheng, Xina
Li, Zhiyong
Yang, Kailun
Computer Vision and Pattern Recognition
Robotics
Image and Video Processing
Leveraging rich information is crucial for dense prediction tasks. Light field (LF) cameras are instrumental in this regard, as they allow data to be sampled from various perspectives. This capability provides valuable spatial, depth, and angular information, enhancing scene-parsing tasks. However, we have identified two overlooked issues for the LF salient object detection (SOD) task. (1): Previous approaches predominantly employ a customized two-stream design to discover the spatial and depth features within light field images. The network struggles to learn the implicit angular information between different images due to a lack of intra-network data connectivity. (2): Little research has been directed towards the data augmentation strategy for LF SOD. Research on inter-network data connectivity is scant. In this study, we propose an efficient paradigm (LF Tracy) to address those issues. This comprises a single-pipeline encoder paired with a highly efficient information aggregation (IA) module (around 8M parameters) to establish an intra-network connection. Then, a simple yet effective data augmentation strategy called MixLD is designed to bridge the inter-network connections. Owing to this innovative paradigm, our model surpasses the existing state-of-the-art method through extensive experiments. Especially, LF Tracy demonstrates a 23% improvement over previous results on the latest large-scale PKU dataset. The source code is publicly available at: https://github.com/FeiBryantkit/LF-Tracy.
title LF Tracy: A Unified Single-Pipeline Approach for Salient Object Detection in Light Field Cameras
topic Computer Vision and Pattern Recognition
Robotics
Image and Video Processing
url https://arxiv.org/abs/2401.16712