Download Scalable Data Mining - University of Toronto

Document related concepts
no text concepts found
Transcript
Scalable Data Mining
Dr. Saed Sayad
University of Toronto
2010
[email protected]
http://chem-eng.utoronto.ca/~datamining/
1
KDnuggets Poll: Important Data Mining Topics
http://chem-eng.utoronto.ca/~datamining/
2
Data Mining?
AI
Machine
Learning
Data
Mining
Statistics
DB/DW
http://chem-eng.utoronto.ca/~datamining/
3
What is Data Mining?
Data mining is about explaining the
past and predicting the future by
means of data analysis
http://chem-eng.utoronto.ca/~datamining/
4
Data Mining Steps
1
• Problem Definition
2
• Data Preparation
3
• Data Exploration
4
• Modeling
5
• Evaluation
6
• Deployment
http://chem-eng.utoronto.ca/~datamining/
5
1. Problem Definition
Understanding the project objectives and
requirements from a business perspective and
then converting this knowledge into a data
mining problem definition with a preliminary
plan designed to achieve the objectives.
Do we need scalable models for our problem?
http://chem-eng.utoronto.ca/~datamining/
6
2. Data Preparation
Data
ETL
DSN
Data
Text
Modeling Data
http://chem-eng.utoronto.ca/~datamining/
7
Scalable Data Preparation
Data Mining
Query Manager
Database
Data Preparation
File System
I/O System
http://chem-eng.utoronto.ca/~datamining/
Reference: William A. Maniatty and Mohammed
J. Zaki, Systems Support for Scalable Data Mining, SIGKDD Explorations, 2000.
8
3. Data Exploration
Average, StDev,
Min, Max, ...
Univariate
Analysis
Bar, Line, Pie, ...
Charts
Data Exploration
Correlation
Z test, ...
Bivariate Analysis
Combination
Charts
http://chem-eng.utoronto.ca/~datamining/
9
Data Exploration - Bivariate
Target%
Months in Business
http://chem-eng.utoronto.ca/~datamining/
10
4. Modeling
Classification
Regression
Clustering
Bayesian
Linear
Regression
Hierarchical
Decision
Tree
Robust
Regression
K-Means
Logistic
Regression
Neural
Network
Association
A Priori
SVM
http://chem-eng.utoronto.ca/~datamining/
11
An Ideal Model
Theoretically
Robust
Truly
Scalable
Practically
Measurable
Don’t be Sisyphus of the analytics world!
http://chem-eng.utoronto.ca/~datamining/
12
True Scalability
• Incremental/Decremental Learning: utilizing new data
without the necessity of pooling new data with old data and
repeating the model training step;
• Variable Addition/Deletion: add or remove variables without
re-training;
• Scenario Testing: rapid formulation and testing of multiple
and diverse models;
• Parallel and Distributed Processing: carrying out parallel
and/or distributed processing simultaneously for all datasets
or data segments to enable a single model using all of the
data to be obtained.
http://chem-eng.utoronto.ca/~datamining/
13
True Scalability
http://chem-eng.utoronto.ca/~datamining/
14
True Scalability
Data Flow
Control System
MLR Model
http://chem-eng.utoronto.ca/~datamining/
15
True Scalability
http://chem-eng.utoronto.ca/~datamining/
16
Data Mining: Modeling
Data
Modeler
Predictor
Explorer
http://chem-eng.utoronto.ca/~datamining/
17
Scalable Modeling
http://chem-eng.utoronto.ca/~datamining/
18
Scalable Modeling
http://chem-eng.utoronto.ca/~datamining/
19
Basic Elements Table
• Summary Tables
– IBM
• Materialized Views
– ORACLE
• Sufficient Statistics
– Microsoft
• General framework
– Wen-Chi Hou, “A Framework for Statistical Data
Mining with Summary Tables”, SIUC, 1999.
http://chem-eng.utoronto.ca/~datamining/
20
Basic Elements Table
Single Unit
Xi
Xj
Bij
where Bij consists of one or more following basic elements:
• N ij
count
• X i
the sum of variable X i
• X j
the sum of variable X j
•  X i X j the sum of multiplication
 X ,  X ,  X , ( X X )
2
3
4
i
2
j
http://chem-eng.utoronto.ca/~datamining/
21
Basic Elements Table
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
bet
Bet
Bet
Bet
Bet
Bet
http://chem-eng.utoronto.ca/~datamining/
22
BET: General Scalable Equation
General Scalable Equation
Bij  Bij  B
new
ij
where:
• Bij  B ji
• (+) represents incremental and (-) decremental changes
http://chem-eng.utoronto.ca/~datamining/
23
BET: Data Types
http://chem-eng.utoronto.ca/~datamining/
24
Sample BET
http://chem-eng.utoronto.ca/~datamining/
25
Basic Elements for Univariate Analysis
Scalable Equation 1: Count
k

k
3
3
X


x
  
N   (  n) i
i 1
i 1
Scalable Equation 2: Sum of Data
 X     x
k

i
Scalable Equation 5: Sum of Quadratic of Data
k

4
4
X


x
  
i
i 1
Scalable Equation 4: Sum of Cubic of Data
i 1

i
Scalable Equation 3: Sum of Squares of Data
k

2
X


x
  
2
i 1

i
http://chem-eng.utoronto.ca/~datamining/
26
Univariate Analysis
http://chem-eng.utoronto.ca/~datamining/
27
Univariate Analysis
http://chem-eng.utoronto.ca/~datamining/
28
Univariate Analysis
http://chem-eng.utoronto.ca/~datamining/
29
Univariate Analysis
Not Scalable
oMinimum
oMaximum
oMedian
oMode
http://chem-eng.utoronto.ca/~datamining/
30
Univariate Analysis
http://chem-eng.utoronto.ca/~datamining/
31
Basic Elements for Bivariate Analysis
Scalable Equation 1: Count
 XY     xy 
k
k
N   (  n) i
i 1
i
i 1
Scalable Equation 2: Sum of Data
 X     x
k
Scalable Equation 5: Sum of Squares of multiplication
k

2
2
(
XY
)


(
xy
)

 
i
i 1
Scalable Equation 4: Sum of multiplication
i 1

i
Scalable Equation 3: Sum of Squares of Data
k

2
X


x
  
2
i 1

i
http://chem-eng.utoronto.ca/~datamining/
32
Bivariate Analysis
Covariance and Correlation
http://chem-eng.utoronto.ca/~datamining/
33
Bivariate Analysis
Z test
t test
F test
http://chem-eng.utoronto.ca/~datamining/
34
Bivariate Analysis
ANOVA
http://chem-eng.utoronto.ca/~datamining/
35
Bivariate Analysis
Z test (proportions)
Chi2 test
http://chem-eng.utoronto.ca/~datamining/
36
Data Mining - Modeling
Classification
Regression
Clustering
Bayesian
Linear
Regression
Hierarchical
Decision
Tree
Robust
Regression
K-Means
Logistic
Regression
Neural
Network
Association
A Priori
SVM
http://chem-eng.utoronto.ca/~datamining/
37
Data Mining: Classification & Regression
Frequency
Covariance
Similarity
Neural
Table
Matrix
Functions
Networks
OneR
Bayesian
Decision
Tree
Markov
Chains
HMM
Linear
Regression
KNN
Perceptron
LDA
Back
(Z Score)
Propagation
PCA/PCR
RBF
Others
SVM
GA
Logistic
Regression
Robust
Regression
http://chem-eng.utoronto.ca/~datamining/
Scalable Methods
38
Linear Models: That’s all we need?
• It is known that many attributes in the
real life are linearly related or can be well
approximated by linear functions.
G. Glass, and K. Hopkins, “Statistical Methods in Education and Psychology”, PrenticeHall Inc., 1984.
http://chem-eng.utoronto.ca/~datamining/
39
Models
based on
Frequency Tables
http://chem-eng.utoronto.ca/~datamining/
40
ZeroR Classification
Predictors
Target
Outlook
Temp.
Humidity
Windy
Play Golf
Rainy
Hot
High
False
No
Rainy
Hot
High
True
No
Overcast
Hot
High
False
Yes
Sunny
Mild
High
False
Yes
Sunny
Cool
Normal
False
Yes
Sunny
Cool
Normal
True
No
Overcast
Cool
Normal
True
Yes
Rainy
Mild
High
False
No
Rainy
Cool
Normal
False
Yes
Sunny
Mild
Normal
False
Yes
Rainy
Mild
Normal
True
Yes
Overcast
Mild
High
True
Yes
Overcast
Hot
Normal
False
Yes
Sunny
Mild
High
True
No
http://chem-eng.utoronto.ca/~datamining/
41
Classification - ZeroR
Play Golf
Play Golf
No
No
No
No
Yes
No
Yes
No
Yes
No
No
Yes
Yes
Sort
5 / 14 = 0.36
Yes
No
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
No
Yes
9 / 14 = 0.64
http://chem-eng.utoronto.ca/~datamining/
Yes
42
Classification - OneR
Which one is the best predictor ?
Outlook
Temp
Humidity
Windy
Play Golf
Rainy
Hot
High
False
No
Rainy
Hot
High
True
No
Overcast
Hot
High
False
Yes
Sunny
Mild
High
False
Yes
Sunny
Cool
Normal
False
Yes
Sunny
Cool
Normal
True
No
Overcast
Cool
Normal
True
Yes
Rainy
Mild
High
False
No
Rainy
Cool
Normal
False
Yes
Sunny
Mild
Normal
False
Yes
Rainy
Mild
Normal
True
Yes
Overcast
Mild
High
True
Yes
Overcast
Hot
Normal
False
Yes
Sunny
Mild
High
True
No
http://chem-eng.utoronto.ca/~datamining/
43
OneR - Frequency Tables
Play Golf
Outlook
Play Golf
Yes
No
Sunny
3
2
Overcast
4
0
Rainy
2
3
Temp.
Yes
No
Hot
2
2
Mild
4
2
Cool
3
1
Play Golf
Humidity
Play Golf
Yes
No
High
3
4
Normal
6
1
Windy
Yes
No
False
6
2
True
3
3
http://chem-eng.utoronto.ca/~datamining/
44
OneR – The Best Rule
Play Golf
Outlook
Yes
No
Sunny
3
2
Overcast
4
0
Rainy
2
3
IF Outlook = Sunny THEN PlayGolf = Yes
IF Outlook = Overcast THEN PlayGolf = Yes
IF Outlook = Rainy THEN PlayGolf = No
http://chem-eng.utoronto.ca/~datamining/
45
Can we incorporate all predictors?
Bayesian Classifier
http://chem-eng.utoronto.ca/~datamining/
46
Bayes’ Rule
Likelihood
Target Prior Probability
PX | T  PT 
PT | X  
PX 
Posterior Probability
Predictor Prior Probability
http://chem-eng.utoronto.ca/~datamining/
47
Likelihood Tables
Play Golf
Outlook
Yes
No
Sunny
3/9
2/5
Overcast
4/9
0/5
Rainy
2/9
3/5
Play Golf
Temp.
Yes
No
Hot
2/9
2/5
Mild
4/9
2/5
Cool
3/9
1/5
Play Golf
Humidity
Yes
No
High
3/9
4/5
Normal
6/9
1/5
Play Golf
Windy
Yes
No
False
6/9
2/5
True
3/9
3/5
P(Outlook  Sunny | PlayGolf  Yes )  3 / 9  0.33
http://chem-eng.utoronto.ca/~datamining/
48
Predictor Probability
Play Golf
Outlook
Yes
No
Sunny
3/9
2/5
5/14
Overcast
4/9
0/5
4/14
Rainy
2/9
3/5
5/14
P(Outlook  Sunny)  5 / 14  0.36
http://chem-eng.utoronto.ca/~datamining/
49
Target Probability
Play Golf
Outlook
Yes
No
Sunny
3/9
2/5
5/14
Overcast
4/9
0/5
4/14
Rainy
2/9
3/5
5/14
9/14
5/14
P( PlayGolf  Yes )  9 / 14  0.64
http://chem-eng.utoronto.ca/~datamining/
50
Posterior Probability
P(Outlook  Sunny | PlayGolf  Yes )  3 / 9  0.33
P( PlayGolf  Yes )  9 / 14  0.64
P(Outlook  Sunny)  5 / 14  0.36
P( PlayGolf  Yes | Outlook  Sunny)  0.33  0.64  0.36  0.60
PT | X  
PX | T  PT 
PX 
http://chem-eng.utoronto.ca/~datamining/
51
Bayesian - Prediction
Case = [Outlook=Sunny; Humidity=High; Temp=Mild; Windy=True]
Play Golf
Posterior
Probability
Sunny
Yes
No
0.60
0.40
Outlook
Temp.
Play Golf
Posterior
Probability
Humidity
Posterior
Probability
High
Yes
No
0.43
0.57
Mild
Posterior
Probability
Windy
True
Play Golf
Yes
No
0.67
0.33
Play Golf
Yes
No
0.50
0.50
P(PlayGolf = Yes) = 0.60 x 0.43 x 0.67 x 0.50 = 0.08643
P(PlayGolf = No) = 0.40 x 0.57 x 0.33 x 0.50 = 0.03762
http://chem-eng.utoronto.ca/~datamining/
52
Naïve Bayesian– Multiple Predictors
P(Ti | X)  P( x1 | Ti )  P( x2 | Ti )  P( xn | Ti )  P(Ti )
Strong Independence assumption about predictors
http://chem-eng.utoronto.ca/~datamining/
53
Can we incorporate all predictors
+ Strong dependency?
Decision Trees
http://chem-eng.utoronto.ca/~datamining/
54
Decision Tree
Outlook
Sunny
Overcast
Rainy
Windy
Play
Humidity
FALSE
TRUE
High
Normal
Play
Not Play
Not Play
Play
http://chem-eng.utoronto.ca/~datamining/
55
Entropy
Entropy = -p log2p – q log2q
Entropy = -0.5 log20.5 – 0.5 log20.5 = 1
http://chem-eng.utoronto.ca/~datamining/
56
Entropy – Frequency
c
E ( S )    pi log 2 pi
i 1
Entropy (5,3,2) = Entropy (0.5,0.3,0.2)
= - (0.5 * log20.5) - (0.3 * log20.3) - (0.2 * log20.2)
= 1.49
http://chem-eng.utoronto.ca/~datamining/
57
Entropy - Target
Play Golf
Play Golf
No
No
No
No
Yes
No
Yes
No
Yes
No
No
Yes
Yes
Sort
5 / 14 = 0.36
Yes
No
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
No
Yes
9 / 14 = 0.64
Entropy(PlayGolf) = Entropy (5,9)
= Entropy (0.36, 0.64)
= - (0.36 log2 0.36) - (0.64 log2 0.64)
= 0.94
http://chem-eng.utoronto.ca/~datamining/
58
Frequency Tables
Play Golf
Outlook
Play Golf
Yes
No
Sunny
3
2
Overcast
4
0
Rainy
2
3
Temp.
Yes
No
Hot
2
2
Mild
4
2
Cool
3
1
Play Golf
Humidity
Play Golf
Yes
No
High
3
4
Normal
6
1
Windy
Yes
No
False
6
2
True
3
3
http://chem-eng.utoronto.ca/~datamining/
59
Entropy – Frequency Table
Play Golf
Outlook
Yes
No
Sunny
3
2
5
Overcast
4
0
4
Rainy
2
3
5
14
E (T , X )   P(v) E (v)
vX
c
E (v)    pi log 2 pi
i 1
E(PlayGolf, Outlook) = P(Sunny)*E(3,2) + P(Overcast)*E(4,0) + P(Rainy)*E(2,3)
= (5/14)*0.971 + (4/14)*0.0 + (5/14)*0.971
= 0.693
http://chem-eng.utoronto.ca/~datamining/
60
Information Gain
Gain(T , X )  Entropy (T )  Entropy (T , X )
G(PlayGolf, Outlook) = E(PlayGolf) – E(PlayGolf, Outlook)
= 0.940 – 0.693 = 0.247
http://chem-eng.utoronto.ca/~datamining/
61
Information Gain – the best predictor?
Play Golf
Outlook
Play Golf
Yes
No
Sunny
3
2
Overcast
4
0
Rainy
2
3
Temp.
Yes
No
Hot
2
2
Mild
4
2
Cool
3
1
Gain = 0.247
Gain = 0.029
Play Golf
Humidity
Play Golf
Yes
No
High
3
4
Normal
6
1
Windy
Yes
No
False
6
2
True
3
3
Gain = 0.152
Gain = 0.048
http://chem-eng.utoronto.ca/~datamining/
62
Decision Tree – Root Node
Outlook
Sunny
Overcast
http://chem-eng.utoronto.ca/~datamining/
Rainy
63
Dataset – Divide and Conqure
Outlook
Temp.
Humidity
Windy
Play Golf
Sunny
Mild
High
FALSE
Yes
Sunny
Cool
Normal
FALSE
Yes
Sunny
Cool
Normal
TRUE
No
Sunny
Mild
Normal
FALSE
Yes
Sunny
Mild
High
TRUE
No
Rainy
Hot
High
FALSE
No
Rainy
Hot
High
TRUE
No
Rainy
Mild
High
FALSE
No
Rainy
Cool
Normal
FALSE
Yes
Rainy
Mild
Normal
TRUE
Yes
Overcast
Hot
High
FALSE
Yes
Overcast
Cool
Normal
TRUE
Yes
Overcast
Mild
High
TRUE
Yes
Overcast
Hot
Normal
FALSE
Yes
http://chem-eng.utoronto.ca/~datamining/
64
Subset (Outlook = Sunny)
Temp.
Humidity
Windy
Play Golf
Mild
High
FALSE
Yes
Cool
Normal
FALSE
Yes
Mild
Normal
FALSE
Yes
Cool
Normal
TRUE
No
Mild
High
TRUE
No
Outlook
Sunny
Overcast
Windy
Play
FALSE
TRUE
Play
Not Play
http://chem-eng.utoronto.ca/~datamining/
Rainy
65
Final Decision Tree
Outlook
Sunny
Overcast
Rainy
Windy
Play
Humidity
FALSE
TRUE
High
Normal
Play
Not Play
Not Play
Play
http://chem-eng.utoronto.ca/~datamining/
66
Models based on
Frequency Table
Robust
Measurable Scalable
OneR
Naïve Bayesian
Decision Tree
http://chem-eng.utoronto.ca/~datamining/
67
Models
based on
Covariance Matrix
http://chem-eng.utoronto.ca/~datamining/
68
Models based on Covariance Matrix
X2
X1
X3
Y
http://chem-eng.utoronto.ca/~datamining/
69
Covariance
Covariance matrix
(x  

C
x
)( y   y )
n
Mean
x


n
http://chem-eng.utoronto.ca/~datamining/
70
Multiple Linear Regression
MLR
Find b by minimizing e
http://chem-eng.utoronto.ca/~datamining/
71
MLR : Ordinary Least Square
Intercept and Slopes:
b  ( X ' X ) X 'Y
1
Cov( x, y )
b
Var ( x)
Predicted Values:
Y   bX
Residuals:
Y Y
http://chem-eng.utoronto.ca/~datamining/
72
Regression Statistics
SST   (Y  Y )
2
SSR   (Y   Y )
SSE   (Y  Y )
http://chem-eng.utoronto.ca/~datamining/
2
2
73
Regression Statistics
How good is our model?
SSR
R 
SST
2
Coefficient of Determination
to judge the adequacy of the regression model
http://chem-eng.utoronto.ca/~datamining/
74
ANOVA
H 0 : b1  b 2  ...  b k  0
H A : bi  0
Regression
Residual
Total
at least one!
df
SS
MS
F
P-value
k
SSR
SSR / df
MSR / MSE
P(F)
n-k-1
SSE
SSE / df
n-1
SST
If P(F)<a then we know that we get significantly better prediction of Y from the
regression model than by just predicting mean of Y.
http://chem-eng.utoronto.ca/~datamining/
75
Hypotheses Tests for Regression
Coefficients
H 0 : bi  0
H A : bi  0
t( n  k 1)
b1  b i bi  b i


2
S e (bi )
S e Cii
http://chem-eng.utoronto.ca/~datamining/
76
Multicolinearity
 If the F-test for significance of regression is
significant, but tests on the individual regression
coefficients are not, multicolinearity may be
present.
 Variance Inflation Factors (VIFs) are very useful
measures of multicolinearity. If any VIF exceed 5,
multicolinearity is a problem.
1
VIF ( b i ) 
 Cii
2
1  Ri
http://chem-eng.utoronto.ca/~datamining/
77
Linear Regression – Binary Target
Yi  bo  b1 X i  e i
• If actual Y is a binary variable the predicted Y can be
less than zero or greater than 1.
• If actual Y is a binary variable error is not normally
distributed.
http://chem-eng.utoronto.ca/~datamining/
78
Linear Regression Model
Y
1
0
X
http://chem-eng.utoronto.ca/~datamining/
79
Logistic Regression
p
1
1 e
 ( b 0  b1 X )
http://chem-eng.utoronto.ca/~datamining/
80
Frequency Table
Months in Business
Count
<50
50-100
100-150
150-200
200-250
250-300
>300
4
12
4
4
4
1
4
Target
Count
0
1
1
2
3
1
4
http://chem-eng.utoronto.ca/~datamining/
Target
Probability
0
0.083
0.25
0.5
0.75
1
1
81
Frequency Plot
1
0.8
Target
Probability
0.6
0.4
0.2
0
1
2
3
4
5
6
7
Months in Business - Bins
http://chem-eng.utoronto.ca/~datamining/
82
Logistic Function
1
f ( z) 
1  ez
http://chem-eng.utoronto.ca/~datamining/
83
Logistic Regression
p
1
1 e
 ( b 0  b1 X )
 The logistic distribution constrains the estimated
probabilities to lie between 0 and 1.
 Maximum Likelihood Estimation is a statistical
method for estimating the coefficients of a model.
http://chem-eng.utoronto.ca/~datamining/
84
Logistic Regression Model
Linear Model
Y
1
Logistic Model
0
X
http://chem-eng.utoronto.ca/~datamining/
85
Log Likelihood (LL)
• Likelihood is the probability that the
dependent variable may be predicted from
the independent variables.
• LL is calculated through iteration, using
maximum likelihood estimation (MLE).
• Log likelihood is the basis for tests of a logistic
model.
http://chem-eng.utoronto.ca/~datamining/
86
Wald Test
• A Wald test is used to test the statistical
significance of each coefficient (b) in the
model.
• A Wald test calculates a Z statistic, which is:
Z
b̂
SE
• This Z value is then squared, yielding a Wald
statistic with a chi-square distribution.
http://chem-eng.utoronto.ca/~datamining/
87
Linear Discriminant Analysis - LDA
$250,000
$200,000
Z  b1 x1  b 2 x2
$150,000
Non-Default
Loan$
$100,000
Default
$50,000
$0
0
10
20
30
40
50
60
70
Age
http://chem-eng.utoronto.ca/~datamining/
88
LDA – Training
Z  b1 x1  b 2 x2  ...  b p x p
b T 1  b T  2
S (b ) 
b T Cb
S (b ) 
Score function
Z1  Z 2
Variance of Z within groups
http://chem-eng.utoronto.ca/~datamining/
89
LDA – Training
Mean
x


n
Covariance matrix
(x  

C
x
)( y   y )
n
http://chem-eng.utoronto.ca/~datamining/
90
LDA – Training
b  C 1 (1  2 )
1
C
(n1C1  n2C2 )
n1  n2
b:
Linear model coefficients
C:
Pooled covariance matrix
1 ,  2 :
Mean vectors
http://chem-eng.utoronto.ca/~datamining/
91
LDA – Model Assessment
  b (1  2 )
2
:
T
Mahalanobis distance between two groups
http://chem-eng.utoronto.ca/~datamining/
92
LDA – Prediction
 1   2 
Z0  b 

 2 
T
Z1  b 1
T
Z 2  b 2
T
http://chem-eng.utoronto.ca/~datamining/
93
Can we handle non-linearity?
Support Vector Machines
http://chem-eng.utoronto.ca/~datamining/
94
Support Vector Machines- SVM
$250,000
$200,000
Support Vectors
$150,000
Non-Default
Loan$
$100,000
Default
$50,000
$0
0
10
20
30
40
50
60
70
Age
http://chem-eng.utoronto.ca/~datamining/
95
SVM and Linear Model
http://chem-eng.utoronto.ca/~datamining/
96
SVM and Linear Model
n
f ( x)   w j x j  b
j 1
m
f ( x)    i k ( xi x)  b
i 1
http://chem-eng.utoronto.ca/~datamining/
97
SVM – Kernel functions
Polynomial
k ( x i , x j )  ( x i .x j )
d
Gaussian Radial Basis function
 x x

i
j
k ( x i , x j )  exp  
2
2



2
http://chem-eng.utoronto.ca/~datamining/





98
Linear SVM
Linear Proximal SVM Algorithm defined by
Fung and Mangasarian: Given m data points in
Rn represented by the m x n matrix A and a
diagonal D of +1 and -1 labels denoting the
class of each row of A, we generate the linear
classifier as follows:
where e is an m x 1 vectors of ones.
Typically v is chosen by means of a tuning
(validating) set.
http://chem-eng.utoronto.ca/~datamining/
99
Principal Component Analysis (PCA)
Principal component analysis (PCA) is a classical statistical method. This linear
transform has been widely used in data analysis and compression. The principal
components (Eigenvectors) for a dataset can directly be extracted from the
covariance matrix as follows:
Eigenvectors and eigenvalues can be computed using a triangular decomposition module.
By ordering the eigenvectors in the order of descending eigenvalues, we can find
directions in which the data set has the most significant amounts of its variation.
http://chem-eng.utoronto.ca/~datamining/
100
Principal Component Regression (PCR)
http://chem-eng.utoronto.ca/~datamining/
101
Models based on Covariance Matrix
Robust
Measurable Scalable
MLR
LDA
PCA/PCR
Logistic Reg.
SVM
http://chem-eng.utoronto.ca/~datamining/
102
Models
based on
Similarity
http://chem-eng.utoronto.ca/~datamining/
103
KNN - Definition
KNN is a simple algorithm that stores
all available cases and classifies new
cases based on a similarity measure.
http://chem-eng.utoronto.ca/~datamining/
104
KNN – different names
•
•
•
•
•
•
K-Nearest Neighbors
Memory-Based Reasoning
Example-Based Reasoning
Instance-Based Learning
Case-Based Reasoning
Lazy Learning
http://chem-eng.utoronto.ca/~datamining/
105
KNN Classification
$250,000
$200,000
$150,000
Non-Default
Loan$
$100,000
Default
$50,000
$0
0
10
20
30
40
50
60
70
Age
http://chem-eng.utoronto.ca/~datamining/
106
KNN Classification – Distance
Age
25
35
45
20
35
52
23
40
60
48
33
Loan
$40,000
$60,000
$80,000
$20,000
$120,000
$18,000
$95,000
$62,000
$100,000
$220,000
$150,000
Default
N
N
N
N
N
N
Y
Y
Y
Y
Y
48
$142,000
?
Euclidean Distance
Distance
102000
82000
62000
122000
22000
124000
47000
80000
42000
78000
D  ( x1  x2 )  ( y1  y2 )
2
http://chem-eng.utoronto.ca/~datamining/
8000
2
107
KNN Classification – Standardized Distance
Age
Loan
0.125
0.375
0.625
0
0.375
0.8
0.075
0.5
1
0.7
0.325
0.11
0.21
0.31
0.01
0.50
0.00
0.38
0.22
0.41
1.00
0.65
Default
N
N
N
N
N
N
Y
Y
Y
Y
Y
0.7
0.61
?
Distance
0.7652
0.5200
0.3160
0.9245
0.3428
0.6220
0.6669
0.4437
0.3650
0.3861
0.3771
X  Min
Xs 
Max  Min
http://chem-eng.utoronto.ca/~datamining/
108
KNN – Number of Neighbors
• If K=1, select the nearest neighbor
• If K>1,
– For classification select the most frequent
neighbor.
– For regression calculate the average of K
neighbors.
– What is the optimal K?
http://chem-eng.utoronto.ca/~datamining/
109
Models based on
Similarity
Robust
Measurable
Scalable
KNN
http://chem-eng.utoronto.ca/~datamining/
110
Neural Networks
http://chem-eng.utoronto.ca/~datamining/
111
Biological Neuron and Integrated Circuit
http://chem-eng.utoronto.ca/~datamining/
112
Biological Neuron Synapse
http://chem-eng.utoronto.ca/~datamining/
113
Neural Network - Neuron
(1) Summation
I i   w ji x j
j
w1i
w ji

f
yi
wni
(2) Transfer
yi  f ( I i )
http://chem-eng.utoronto.ca/~datamining/
114
Transfer Functions
http://chem-eng.utoronto.ca/~datamining/
115
Weight Adjustment or Error Propagation
Wi   ( D  Y ) X i
w1i
w ji

f
e
wni
http://chem-eng.utoronto.ca/~datamining/
116
Neural Networks
Robust
Measurable Scalable
Neural Networks
http://chem-eng.utoronto.ca/~datamining/
117
Scalable Modeling
(Classification & Regression)
Frequency
Covariance
Table
Matrix
OneR
Bayesian
Markov
Chains
Linear
Regression
LDA
(Z Score)
PCA/PCR
HMM
http://chem-eng.utoronto.ca/~datamining/
118
Scalable Data Mining: Summary
Data Mining
Query Manager
Classification
Frequency
Table
Regression
Covariance
Matrix
Clustering
Database
Association
File System
I/O System
http://chem-eng.utoronto.ca/~datamining/
119
Scalable Data Mining
Parallel and Distributed Processing
Dataset
Dataset 1
Dataset 2
Dataset 3
Dataset N
Processor 1
Processor 4
Processor 3
Processor N
BET 1
BET 2
BET 3
BET N
Processor N+1
BET 1+2+3+…+N
Modeling
http://chem-eng.utoronto.ca/~datamining/
120
Scalable Data Mining Chip!
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
bet
Bet
Bet
Bet
Bet
Bet
Memory Size for 10,000 Variables
10K x 10K x 8 bytes = 800MB x 5 = 4GB
http://chem-eng.utoronto.ca/~datamining/
121
Scalable Data Mining
Multidimensional, Heterogeneous and Asymmetric
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
Bet
bet
Bet
Bet
Bet
Bet
Bet
http://chem-eng.utoronto.ca/~datamining/
122
Demo…
http://chem-eng.utoronto.ca/~datamining/
123
Related documents