Abstract:
Existing image and text retrieval methods only exploit the superficial correlations contained in instance pair data and ignore the importance of external commonsense, which may hinder their ability to reason about higher- level relationships between image and text data. Therefore, a deep hash cross- modal retrieval method based on logical commonsense is proposed. It enriched the semantic associations between concepts by introducing logical knowledge graph for concept representation learning, enhanced the decoupling ability of the model to the higher level of semantics and its interpretability, and reduced the 'semantic gap' between different modalities by constructing a shared semantic space for guiding the interactive learning of graphical and textual features. Experiments on two datasets verify the effectiveness of the LCH algorithm.