Advancements in Transformer Architectures for Large Language Model: From Bert to GPT-3 and Beyond | Synapse