Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>To enhance the semantic understanding of heterogeneous data, cross-modal alignment-based text-to-image generation and captioning is essential in recent times. Nevertheless, the existing works failed to track Cataphora Dependencies (CDs) in which pronouns appeared before their actual entities, thereby leading to poor semantic coherence and generation efficiency. Hence, in this article, an intelligent cross-modal alignment with CD modeling-based content retrieval and captioning utilizing Generative Collapsing Linear Adversarial Network (GCLAN) and Generative Quadruple Pre-trained Probabilist’s Hermite Polynomials Transformer (GQ2PHPT) is proposed.Initially, the Text-to-Image Synthesis (TIS) module is designed. Here, the images are gathered, followed by pre-processing and fine-grained visual characteristics learning. Similarly, the text is collected and further fed into Natural Language Processing (NLP)-based pre-processing. Afterward, contextual word embedding, CD analysis, syntactic structure modeling, and feature extraction are carried out from pre-processed text. Then, by utilizing GCLAN, cross-modal alignment is performed. In real-time, an image is generated by the GCLAN regarding the user's query. Next, the generated image is fed into the Image Captioning (IC) module, in which a suitable caption is efficiently generated by the proposed GQ2PHPT. Thereafter, concerning the caption and generated images, cross-modal alignment is carried out, followed by score computation. Lastly, by considering the score value, the content is retrieved with a high success rate of 0.9422. When contrasted with the prevailing works, the proposed work effectively performs TIS and IC with an accuracy of 98.9734%.</p>

Show More

Keywords

crossmodal generated captioning alignment gclan

Related Articles

PORE

About

Connect