Survey
* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project
* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project
A Neural Probabilistic Language Model 2014-12-16 Keren Ye CONTENTS • • • • N-gram Models Fighting the Curse of Dimensionality A Neural Probabilistic Language Model Continuous Bag of Words(Word2vec) n-gram models • • Construct tables of conditional probabilities for the next word Combinations of the last n-1 words t 1 t 1 ˆ ˆ P wt | w1 P wt | wt n1 n-gram models • i.e. “I like playing basketball” – Unigram(1-gram) Pˆ basketball | I , like, playing Pˆ basketball – Bigram(2-gram) Pˆ basketball | I , like, playing Pˆ basketball | playing – Trigram(3-gram) Pˆ basketball | I , like, playing Pˆ basketball | like, playing n-gram models • Disadvantages – It is not taking into account contexts farther than 1 or 2 words – It is not taking into account the similarity between words • • i.e.“The cat is walking in the bedroom”(training corpus) “A dog was running in a room”(?) n-gram models • Disadvantages – Curse of Dimensionality CONTENTS • • • • N-gram Models Fighting the Curse of Dimensionality A Neural Probabilistic Language Model Continuous Bag of Words(Word2vec) Fighting the Curse of Dimensionality • • • Associate with each word in the vocabulary a distributed word feature vector (a real-valued vector in R m ) Express the joint probability function of word sequences in terms of the feature vectors of these words in the sequence Learn simultaneously the word feature vectors and the parameters of that probability function Fighting the Curse of Dimensionality • Word feature vectors – Each word is associated with a point in a vector space – The number of features (e.g. m=30, 60 or 100 in the experiments) is much smaller than the size of vocabulary (e.g. 20w) Fighting the Curse of Dimensionality • Probability function – Using a multi-layer neural network to predict the next word given the previous ones, in the experiments – This function has parameters that can be iteratively tuned in order to maximize the log-likelihood of the training data Fighting the Curse of Dimensionality • Why does it work? – If we knew that “dog” and “cat” played similar roles (semantically and syntactically), and similarly for (the, a), (bedroom, room), (is, was), (running, walking), we could naturally generalize from • The cat is walking in the bedroom – to and likewise to • • • • A dog was running in a room The cat is running in a room A dog is walking in a bedroom …. Fighting the Curse of Dimensionality • NNLM – Neural Network Language Model CONTENTS • • • • N-gram Models Fighting the Curse of Dimensionality A Neural Probabilistic Language Model Continuous Bag of Words(Word2vec) A Neural Probabilistic Language Mode • Denotations – The training set is a sequence w1...wT of words wt where the vocabulary V is a large but finite set V , – The objective is to learn a good model as below, in the sense that it gives high out-out-sample likelihood f wt ,..., wt n1 Pˆ wt | w1t 1 – The only constraint on model is that for any choice of V the sum f i, wt 1 ,..., wt n 1 1 i 1 w1t 1 , A Neural Probabilistic Language Mode • Objective function – Training is achieved by looking for that maximizes the training corpus penalized log-likelihood, where regularization term 1 L log f wt ,..., wt n 1; R T t R is a A Neural Probabilistic Language Mode • Model – We decompose the function f wt ,..., wt n1 Pˆ wt | w1t 1 two parts in • A mapping C from any element i of V to a real vector C i R It represents the distributed feature vectors associated with each word in the vocabulary • The probability function over words, expressed with C : a function g maps an input sequence of feature vectors for words in context, C wt n 1 ,..., C wt 1 , to a conditional probability distribution over words in V for the next word. The output of g is a vector whose i-th element estimates the probability f i, wt 1 ,..., wt n1 g i, C wt 1 ,..., C wt n1 m A Neural Probabilistic Language Mode A Neural Probabilistic Language Mode • Model details (two hidden layers) – The shared word features layer C, which has no nonlinearity (it would not add anything useful) – The ordinary hyperbolic tangent hidden layer A Neural Probabilistic Language Mode • Model details (formal description) – The neural network computes the following function, with a softmax output layer, which guarantees positive probabilities summing to 1 Pˆ wt | wt 1 ,..., wt n 1 e y wt yi e i A Neural Probabilistic Language Mode • Model details (formal description) – The yi are the unnormalized log-probabilities for each output word i , computed as follows, with parameters b, W, U, d and H y b Wx U tanh( d Hx ) • Where the hyperbolic tangent tanh is applied element by element, W is optionally zero (no direct connections) • And x is the word features layer activation vector, which is the concatenation of the input word features from the matrix C x C wt n1 ,..., C wt 1 A Neural Probabilistic Language Mode A Neural Probabilistic Language Mode Parameters Brief Dimensions b Output biases |V| d Hidden layer bieses h W No direct connections 0 U Hidden-to-output weights |V|*h matrix H Word features to output weights h*(n-1)m matrix C Word features |V|*m matrix y b Wx U tanh( d Hx ) A Neural Probabilistic Language Mode A Neural Probabilistic Language Mode • Stochastic gradient ascent b, d ,W ,U , H , C log Pˆ wt | wt 1 ,..., wt n 1 – Note that a large fraction of the parameters needs not be updated or visited after each example: the word feature C(j) of all words j that do not occur in the input window A Neural Probabilistic Language Mode • Parallel Implementation – Data-Parallel Processing • • Relied on synchronization commands – slow No locks – noise seems to be very small and did not apparently slow down training – Parameter-parallel Processing • Parallelize across the parameters A Neural Probabilistic Language Mode A Neural Probabilistic Language Mode Continuous Bag of Words(Word2vec) • Bag of words – Traditional solution for the problem of Curse of Dimensionality Pbasketball | I , like, playing Pbasketball PI | basketball Plike | basketball P playing | basketball PI , like, playing CONTENTS • • • • N-gram Models Fighting the Curse of Dimensionality A Neural Probabilistic Language Model Continuous Bag of Words(Word2vec) Continuous Bag of Words(Word2vec) • Continuous Bag of Words Continuous Bag of Words(Word2vec) • Distinctness – Projection layer • • Sum vs Concatenate Order of words – Hidden layer • tanh vs NULL – Hierarchical Softmax Thanks Q&A