Chapter 9 This hidden layer is, in turn, used to calculate a corresponding output, yt. Sequences are processed by presenting one element at a time to the network. The key difference from a feed forward network lies in the recurrent link shown in the figure with the dashed line. This link augments the input to the hidden layer with the activation value of the hidden layer from the preceding point in time. In the commonly encountered case of soft classification, finding yt consists of a softmax computation that provides a normalized probability distribution over the sequential nature of simple recurrent networks can be illustrated by unrolling the network. For applications that involve much longer input sequences, such as speech recognition, character-by-character sentence processing, or streaming of continuous inputs, unrolling an entire input sequence may not be feasible. In these cases, we can unroll the input into manageable fixed-length segments and treat each segment ...