In the last article, every sequence was made of numbers. How does a network read a word, and write one? It reads and writes tokens. A token is a unit of text: a letter, a word, or part of a word. I use letters, so every example fits on a screen.
cab, d=2, h=2, seq=3, many-to-one, sigmoid
Let's start with a vocabulary of three letters, a b c, and read the word cab. Each letter becomes a vector of two numbers. We keep the vectors in a table, E, with one column per letter. Each column is an embedding. In these examples, E is fixed, so it has no fill. In practice, E is usually trained along with the weights. To read a letter, we look it up, finding the column headed by that letter. So cab becomes three vectors, x₁ to x₃, in cream. From there, the RNN reads them many-to-one, as in the Outputs article. A sigmoid gives one probability for the whole word. In the download, type another word made from a b c, and every step recomputes.


