Skip to content

< all problems12 · Level 01, LLM APIs

Round-Trip Text Through a Tokenizer

easy · implement · LLM Fundamentals

Implement three functions using the provided tokenizer:

  1. count_tokens(text): how many tokens the text becomes, not counting any special tokens the tokenizer adds on its own.
  2. roundtrip(text): encode then decode, returning text that matches the input.
  3. token_strings(text): the individual token strings, in order.

The catch: encoding adds special tokens you did not ask for, and a naive decode puts them back, so the round trip comes out wrapped in <|endoftext|> or similar.