tokenize input manage context length understand attention cost choose decoding strategy optimize inference control output quality