Abstract
This paper enhances the mathematical and technical foundations of Generative AI by incorporating feedback from reviewers. This paper distinguishes itself from other surveys on the topic by providing a unified perspective that connects mathematical theory, model design, and ethical considerations for different modalities (text, image, audio). The paper introduces comparative tables of models, datasets, libraries, and evaluation metrics, strengthening its practical utility. Furthermore, recent advances from last two years (e.g., Gemini, Mixtral, Claude 2, Sora) are incorporated, alongside an enhanced focus on interpretability, multimodal integration, and sustainable AI.