Abstract:
This paper proposes a cross- visual modal biological recognition method based on joint representation learning. Specifically, for the cross face- eye area recognition problem, we introduced a deep convolutional self- attention network (DCSAN) and a cross- visual- modality contrast loss (CVMC). DCSAN utilized deep convolutional layers to extract texture features from face and eye area image patches, enabling local joint representation learning. Additionally, it employed deep convolutional multi- head self- attention modules (DWC- MHSAM) to model global dependencies between face and eye area regions. The introduction of the CVMC loss helped handle cross- modal negative sample pairs, and distinguished highly similar face and eye area image patches belonging to different individuals within the same modality. We conducted experiments on the Ethnic, FaceScrub, and IMDB datasets. The results demonstrate that when using face as the query set, the optimal recognition rate reached 75.46%, while using eye area as the query set achieved an optimal recognition rate of 76.36%.