split text into tokens convert tokens into embeddings add positional information for each Transformer block: compute Self-Attention mix token information apply feed-forward transformation keep stable flow with residual connections and normalization produce contextual token representations