LSTM networks work by feeding in some input then they give you a predicted output. This works sort of well for text, speech or music, where you put in say 10 notes, it pumps out the 11th, then you put in the 9 old ones, dropping the first one and appending the note it generated. Same with words...