{ "answer": "The options mentioned for positional encoding are sinusoidal positional encodings (using sine and cosine functions of different frequencies) and learned positional embeddings.", "start_page_num": 6, "start_line_num": 27, "end_page_num": 6, "end_line_num": 41, "confidence": 0.99, "justification": "Lines 6:27-6:41 describe adding 'positional encodings' to the input embeddings, specify the sinusoidal method, and mention experimenting with learned positional embeddings, stating both options were tried and produced nearly identical results.", "quotes": [ "Since our model contains no recurrence and no convolution, in order for the model to make use of the order of the sequence, we must inject some information about the relative or absolute position of the tokens in the sequence. To this end, we add 'positional encodings' to the input embeddings at the bottoms of the encoder and decoder stacks. The positional encodings have the same dimension dmodel as the embeddings, so that the two can be summed. There are many choices of positional encodings, learned and fixed [9]. In this work, we use sine and cosine functions of different frequencies: ... We also experimented with using learned positional embeddings [9] instead, and found that the two versions produced nearly identical results (see Table 3 row (E)). We chose the sinusoidal version because it may allow the model to extrapolate to sequence lengths longer than the ones encountered during training." ], "caveats": [ "Exact mathematical formulas for sinusoidal encoding are present here, but full details for learned embeddings are not. Table 3 row (E) and further details may expand on results but are not needed for the options question." ], "complete_answer_found": true, "context_structured": true, "llm_discovered_keywords": [ "sinusoidal positional encoding", "learned positional embeddings", "sine and cosine functions", "relative or absolute position" ] }