Abstract
Maritime surveillance videos based object detection methods aims to meet the quick response requirements through an effective ship detection and recognition system against the backdrop of smart ocean technology. Our research has focused on this aspects as mentioned below: 1) to summarize current approaches and datasets and discuss the challenging issues of them; 2) to analyze the features and the challenges of maritime surveillance ship detection; 3) to clarify the credibility of their accuracy and efficiency and demonstrate our research potential further. In the first phase, we summarize existing ship detection algorithms based on maritime surveillance videos, introduce the common ship detection datasets and ship detection methods available, and some evaluation metrics followed for ship detection tasks. Customized attention is yielded to the interconnected results between traditional computer vision algorithms for ship identification, which mainly consist of modules such as horizon detection, background subtraction and foreground extraction, and some deep learning methods based on fast region convolutional neural network (Fast R-CNN), single shot multibox detector (SSD) and you only look once (YOLO) . It can be sorted out that although mean average precision (mAP) metric remains recognized index to measure the performance of models, its effectiveness issue is still discussed in terms of ship detection tasks and present novel metrics, including bottom edge proximity (BEP), n-multiple object detection precision (N-MODP) and n-multiple object detection accuracy (N-MODA). Current datasets are capable to detect vessels motion via deep learning models. But, the accuracy and robustness of training are required to be improved greatly due to extreme weather condition and light variation or inconsistent labels. In the second phase, we evaluate the features and challenges for ship detection. The difference lies between ship detection and regular object detection. For example, a coastline platform or a ship sensor have very large visible ranges and leading to a big scale variability. In addition, it is challenged to design a set of models adapt to various image domain scenarios derived from extreme marine weather conditions. Photographic system has to withstand exposure to extremes of temperature, high vibration levels, humidity and chemicals as well. The harsh environment combined with noise pollution and limited network bandwidth can cause the loss of image quality and make uncertainty with information loss for the models. In the third phase, we improve the accuracy and efficiency of ship detection algorithms and evaluate some common methods for ship detection technology on the three aspects as following: 1) multi-scale feature fusion: we carry out convolutional neural network (CNN) models manipulation based on different input scales and backbones. Some of the object detection models are degraded when facing large variations result from large field of view among ship objects during voyage. It is suggested that input scale determines the upper bound of accuracy of CNN models, and CNN models or backbones which are specially designed for multi-scale detection tasks narrow the gap of accuracy between different input scales. 2) Data augmentation: Waves and wind induce pitch and roll rotations on the sea, which is demanding for ship detection. Moreover, weather change and day-night brightness variation mean the image data shall be in multiple domains, which requires the detection models to be robust to images from different domains, or even from domains that have not been included in training samples. In light of the above-mentioned variables and marine based camera motion results, we evaluate the performance improvement when data augmentation is applied to image translation and rotation. We also adjust photo brightness and use Gaussian blur to simulate the blur caused by water condensed on cameras. Almost 5% increase is observed in mAP, verifying the robustness of data augmentation in ship detection. Other effective approaches include domain transfer based on generative adversarial network (GAN) or domain-independent models derived of multidomain object detection tasks. 3) Light-weighted models and energy optimization: Computing complexity is constrained of semantic constraints like the horizon; Common object detection optimization is used to lower computation load, including light-weighted backbones like MobileNet and ShuffleNet. We calculate the parameter quantity and computing operation quantity of object detection models as well as each accuracy of them. We recommended that further studies should be considered on the following aspects: 1) to develop new datasets or existing datasets improvement based on sufficient coverage of possible conditions, high-quality annotations, precise classification and easier extension, respectively; 2) to decrease arithmetic operations and energy consumption in object detection models; 3) to strengthen multi-scale target detection modeling; 4) to enhance data fusion between object detection in multi-sensors images and the semantic ability of single image or multiple images interpretation.
Author supplied keywords
Cite
CITATION STYLE
Ye, C., Lu, T., Xiao, Y., Lu, H., & Yang, Q. (2022, July 16). Maritime surveillance videos based ships detection algorithms: a survey. Journal of Image and Graphics. Editorial and Publishing Board of JIG. https://doi.org/10.11834/jig.200674
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.