Download *** 1 - Keren Ye`s Home

Survey
yes no Was this document useful for you?
   Thank you for your participation!

* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project

Document related concepts

Pattern recognition wikipedia , lookup

Artificial neural network wikipedia , lookup

Types of artificial neural networks wikipedia , lookup

Transcript
A Neural Probabilistic Language Model
2014-12-16
Keren Ye
CONTENTS
•
•
•
•
N-gram Models
Fighting the Curse of Dimensionality
A Neural Probabilistic Language Model
Continuous Bag of Words(Word2vec)
n-gram models
•
•
Construct tables of conditional probabilities for the
next word
Combinations of the last n-1 words

 
t 1
t 1
ˆ
ˆ
P wt | w1  P wt | wt n1

n-gram models
•
i.e. “I like playing basketball”
– Unigram(1-gram)
Pˆ basketball | I , like, playing   Pˆ basketball 
– Bigram(2-gram)
Pˆ basketball | I , like, playing   Pˆ basketball | playing 
– Trigram(3-gram)
Pˆ basketball | I , like, playing   Pˆ basketball | like, playing 
n-gram models
•
Disadvantages
– It is not taking into account contexts farther than 1 or 2
words
– It is not taking into account the similarity between words
•
•
i.e.“The cat is walking in the bedroom”(training corpus)
“A dog was running in a room”(?)
n-gram models
•
Disadvantages
– Curse of Dimensionality
CONTENTS
•
•
•
•
N-gram Models
Fighting the Curse of Dimensionality
A Neural Probabilistic Language Model
Continuous Bag of Words(Word2vec)
Fighting the Curse of Dimensionality
•
•
•
Associate with each word in the vocabulary a
distributed word feature vector (a real-valued vector
in R m )
Express the joint probability function of word
sequences in terms of the feature vectors of these
words in the sequence
Learn simultaneously the word feature vectors and
the parameters of that probability function
Fighting the Curse of Dimensionality
•
Word feature vectors
– Each word is associated with a point in a vector space
– The number of features (e.g. m=30, 60 or 100 in the
experiments) is much smaller than the size of vocabulary
(e.g. 20w)
Fighting the Curse of Dimensionality
•
Probability function
– Using a multi-layer neural network to predict the next word
given the previous ones, in the experiments
– This function has parameters that can be iteratively tuned
in order to maximize the log-likelihood of the training
data
Fighting the Curse of Dimensionality
•
Why does it work?
– If we knew that “dog” and “cat” played similar roles
(semantically and syntactically), and similarly for (the, a),
(bedroom, room), (is, was), (running, walking), we could
naturally generalize from
•
The cat is walking in the bedroom
– to and likewise to
•
•
•
•
A dog was running in a room
The cat is running in a room
A dog is walking in a bedroom
….
Fighting the Curse of Dimensionality
•
NNLM
– Neural Network Language Model
CONTENTS
•
•
•
•
N-gram Models
Fighting the Curse of Dimensionality
A Neural Probabilistic Language Model
Continuous Bag of Words(Word2vec)
A Neural Probabilistic Language Mode
•
Denotations
– The training set is a sequence w1...wT
of words wt
where the vocabulary V is a large but finite set
V ,
– The objective is to learn a good model as below, in the
sense that it gives high out-out-sample likelihood

f wt ,..., wt n1   Pˆ wt | w1t 1

– The only constraint
on model is that for any choice of
V
the sum  f i, wt 1 ,..., wt n 1   1
i 1
w1t 1 ,
A Neural Probabilistic Language Mode
•
Objective function
– Training is achieved by looking for  that maximizes the
training corpus penalized log-likelihood, where
regularization term
1
L   log f wt ,..., wt n 1;   R 
T t
R  is a
A Neural Probabilistic Language Mode
•
Model
– We decompose the function f wt ,..., wt n1   Pˆ wt | w1t 1 
two parts
in

•
A mapping C from any element i of V to a real vector C i  R
It represents the distributed feature vectors associated with
each word in the vocabulary
•
The probability function over words, expressed with C : a
function g maps an input sequence of feature vectors for
words in context, C wt n 1 ,..., C wt 1  , to a conditional
probability distribution over words in V for the next word. The
output of g is a vector whose i-th element estimates the
probability f i, wt 1 ,..., wt n1   g i, C wt 1 ,..., C wt n1 
m
A Neural Probabilistic Language Mode
A Neural Probabilistic Language Mode
•
Model details (two hidden layers)
– The shared word features layer C, which has no nonlinearity (it would not add anything useful)
– The ordinary hyperbolic tangent hidden layer
A Neural Probabilistic Language Mode
•
Model details (formal description)
– The neural network computes the following function, with a
softmax output layer, which guarantees positive
probabilities summing to 1
Pˆ wt | wt 1 ,..., wt n 1  
e
y wt
yi
e
i
A Neural Probabilistic Language Mode
•
Model details (formal description)
– The yi
are the unnormalized log-probabilities for each
output word i , computed as follows, with parameters b, W,
U, d and H
y  b  Wx  U tanh( d  Hx )
•
Where the hyperbolic tangent tanh is applied element by
element, W is optionally zero (no direct connections)
•
And x is the word features layer activation vector, which is the
concatenation of the input word features from the matrix C
x  C wt n1 ,..., C wt 1 
A Neural Probabilistic Language Mode
A Neural Probabilistic Language Mode
Parameters
Brief
Dimensions
b
Output biases
|V|
d
Hidden layer bieses
h
W
No direct connections
0
U
Hidden-to-output weights
|V|*h matrix
H
Word features to output weights
h*(n-1)m matrix
C
Word features
|V|*m matrix
y  b  Wx  U tanh( d  Hx )
A Neural Probabilistic Language Mode
A Neural Probabilistic Language Mode
•
Stochastic gradient ascent
  b, d ,W ,U , H , C 
 log Pˆ wt | wt 1 ,..., wt n 1 
  

– Note that a large fraction of the parameters needs not be
updated or visited after each example: the word feature C(j)
of all words j that do not occur in the input window
A Neural Probabilistic Language Mode
•
Parallel Implementation
– Data-Parallel Processing
•
•
Relied on synchronization commands – slow
No locks – noise seems to be very small and did not
apparently slow down training
– Parameter-parallel Processing
•
Parallelize across the parameters
A Neural Probabilistic Language Mode
A Neural Probabilistic Language Mode
Continuous Bag of Words(Word2vec)
•
Bag of words
– Traditional solution for the problem of Curse of
Dimensionality
Pbasketball | I , like, playing  
Pbasketball   PI | basketball   Plike | basketball   P playing | basketball 
PI , like, playing 
CONTENTS
•
•
•
•
N-gram Models
Fighting the Curse of Dimensionality
A Neural Probabilistic Language Model
Continuous Bag of Words(Word2vec)
Continuous Bag of Words(Word2vec)
•
Continuous Bag of Words
Continuous Bag of Words(Word2vec)
•
Distinctness
– Projection layer
•
•
Sum vs Concatenate
Order of words
– Hidden layer
•
tanh vs NULL
– Hierarchical Softmax
Thanks
Q&A