split input text into tokens convert tokens into embedding vectors pass embeddings through encoder layers for each encoder layer: compute self-attention mix information across tokens apply feed-forward transformation produce contextual token representations pass previous output tokens into decoder for each decoder layer: apply masked self-attention attend to encoder output with cross-attention apply feed-forward transformation predict the next output token