Abstract
The issue of writer identification and writer retrieval, which is considered a challenging problem in the field of document analysis and recognition is addressed here. A novel pipeline is proposed for the problem at hand by employing a unified neural network architecture consisting of the ResNet-20 as a feature extractor and an integrated NetVLAD layer, inspired by the vector of locally aggregated descriptors (VLAD), in the head of the latter part. Having defined this architecture, the triplet semi-hard loss function is used to directly learn an embedding for individual input image patches. Subsequently, the generalised max-pooling technique is employed for the aggregation of embedded descriptors of each handwritten image. Also, a novel re-ranking strategy is introduced for the task of identification and retrieval based on the k-reciprocal nearest neighbours, and it is shown that the pipeline can benefit tremendously from this step. Experimental evaluation has been done on the three publicly available datasets: the ICDAR 2013, CVL, and KHATT datasets. Results indicate that while the performance was comparable to the state-of-the-art KHATT, the writer identification and writer retrieval pipeline achieve superior performance on the ICDAR 2013 and CVL datasets in terms of mAP.
Cite
CITATION STYLE
Rasoulzadeh, S., & BabaAli, B. (2022). Writer identification and writer retrieval based on NetVLAD with Re-ranking. IET Biometrics, 11(1), 10–22. https://doi.org/10.1049/bme2.12039
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.