Survey
* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project
* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project
Scalable Data Mining Dr. Saed Sayad University of Toronto 2010 [email protected] http://chem-eng.utoronto.ca/~datamining/ 1 KDnuggets Poll: Important Data Mining Topics http://chem-eng.utoronto.ca/~datamining/ 2 Data Mining? AI Machine Learning Data Mining Statistics DB/DW http://chem-eng.utoronto.ca/~datamining/ 3 What is Data Mining? Data mining is about explaining the past and predicting the future by means of data analysis http://chem-eng.utoronto.ca/~datamining/ 4 Data Mining Steps 1 • Problem Definition 2 • Data Preparation 3 • Data Exploration 4 • Modeling 5 • Evaluation 6 • Deployment http://chem-eng.utoronto.ca/~datamining/ 5 1. Problem Definition Understanding the project objectives and requirements from a business perspective and then converting this knowledge into a data mining problem definition with a preliminary plan designed to achieve the objectives. Do we need scalable models for our problem? http://chem-eng.utoronto.ca/~datamining/ 6 2. Data Preparation Data ETL DSN Data Text Modeling Data http://chem-eng.utoronto.ca/~datamining/ 7 Scalable Data Preparation Data Mining Query Manager Database Data Preparation File System I/O System http://chem-eng.utoronto.ca/~datamining/ Reference: William A. Maniatty and Mohammed J. Zaki, Systems Support for Scalable Data Mining, SIGKDD Explorations, 2000. 8 3. Data Exploration Average, StDev, Min, Max, ... Univariate Analysis Bar, Line, Pie, ... Charts Data Exploration Correlation Z test, ... Bivariate Analysis Combination Charts http://chem-eng.utoronto.ca/~datamining/ 9 Data Exploration - Bivariate Target% Months in Business http://chem-eng.utoronto.ca/~datamining/ 10 4. Modeling Classification Regression Clustering Bayesian Linear Regression Hierarchical Decision Tree Robust Regression K-Means Logistic Regression Neural Network Association A Priori SVM http://chem-eng.utoronto.ca/~datamining/ 11 An Ideal Model Theoretically Robust Truly Scalable Practically Measurable Don’t be Sisyphus of the analytics world! http://chem-eng.utoronto.ca/~datamining/ 12 True Scalability • Incremental/Decremental Learning: utilizing new data without the necessity of pooling new data with old data and repeating the model training step; • Variable Addition/Deletion: add or remove variables without re-training; • Scenario Testing: rapid formulation and testing of multiple and diverse models; • Parallel and Distributed Processing: carrying out parallel and/or distributed processing simultaneously for all datasets or data segments to enable a single model using all of the data to be obtained. http://chem-eng.utoronto.ca/~datamining/ 13 True Scalability http://chem-eng.utoronto.ca/~datamining/ 14 True Scalability Data Flow Control System MLR Model http://chem-eng.utoronto.ca/~datamining/ 15 True Scalability http://chem-eng.utoronto.ca/~datamining/ 16 Data Mining: Modeling Data Modeler Predictor Explorer http://chem-eng.utoronto.ca/~datamining/ 17 Scalable Modeling http://chem-eng.utoronto.ca/~datamining/ 18 Scalable Modeling http://chem-eng.utoronto.ca/~datamining/ 19 Basic Elements Table • Summary Tables – IBM • Materialized Views – ORACLE • Sufficient Statistics – Microsoft • General framework – Wen-Chi Hou, “A Framework for Statistical Data Mining with Summary Tables”, SIUC, 1999. http://chem-eng.utoronto.ca/~datamining/ 20 Basic Elements Table Single Unit Xi Xj Bij where Bij consists of one or more following basic elements: • N ij count • X i the sum of variable X i • X j the sum of variable X j • X i X j the sum of multiplication X , X , X , ( X X ) 2 3 4 i 2 j http://chem-eng.utoronto.ca/~datamining/ 21 Basic Elements Table Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet bet Bet Bet Bet Bet Bet http://chem-eng.utoronto.ca/~datamining/ 22 BET: General Scalable Equation General Scalable Equation Bij Bij B new ij where: • Bij B ji • (+) represents incremental and (-) decremental changes http://chem-eng.utoronto.ca/~datamining/ 23 BET: Data Types http://chem-eng.utoronto.ca/~datamining/ 24 Sample BET http://chem-eng.utoronto.ca/~datamining/ 25 Basic Elements for Univariate Analysis Scalable Equation 1: Count k k 3 3 X x N ( n) i i 1 i 1 Scalable Equation 2: Sum of Data X x k i Scalable Equation 5: Sum of Quadratic of Data k 4 4 X x i i 1 Scalable Equation 4: Sum of Cubic of Data i 1 i Scalable Equation 3: Sum of Squares of Data k 2 X x 2 i 1 i http://chem-eng.utoronto.ca/~datamining/ 26 Univariate Analysis http://chem-eng.utoronto.ca/~datamining/ 27 Univariate Analysis http://chem-eng.utoronto.ca/~datamining/ 28 Univariate Analysis http://chem-eng.utoronto.ca/~datamining/ 29 Univariate Analysis Not Scalable oMinimum oMaximum oMedian oMode http://chem-eng.utoronto.ca/~datamining/ 30 Univariate Analysis http://chem-eng.utoronto.ca/~datamining/ 31 Basic Elements for Bivariate Analysis Scalable Equation 1: Count XY xy k k N ( n) i i 1 i i 1 Scalable Equation 2: Sum of Data X x k Scalable Equation 5: Sum of Squares of multiplication k 2 2 ( XY ) ( xy ) i i 1 Scalable Equation 4: Sum of multiplication i 1 i Scalable Equation 3: Sum of Squares of Data k 2 X x 2 i 1 i http://chem-eng.utoronto.ca/~datamining/ 32 Bivariate Analysis Covariance and Correlation http://chem-eng.utoronto.ca/~datamining/ 33 Bivariate Analysis Z test t test F test http://chem-eng.utoronto.ca/~datamining/ 34 Bivariate Analysis ANOVA http://chem-eng.utoronto.ca/~datamining/ 35 Bivariate Analysis Z test (proportions) Chi2 test http://chem-eng.utoronto.ca/~datamining/ 36 Data Mining - Modeling Classification Regression Clustering Bayesian Linear Regression Hierarchical Decision Tree Robust Regression K-Means Logistic Regression Neural Network Association A Priori SVM http://chem-eng.utoronto.ca/~datamining/ 37 Data Mining: Classification & Regression Frequency Covariance Similarity Neural Table Matrix Functions Networks OneR Bayesian Decision Tree Markov Chains HMM Linear Regression KNN Perceptron LDA Back (Z Score) Propagation PCA/PCR RBF Others SVM GA Logistic Regression Robust Regression http://chem-eng.utoronto.ca/~datamining/ Scalable Methods 38 Linear Models: That’s all we need? • It is known that many attributes in the real life are linearly related or can be well approximated by linear functions. G. Glass, and K. Hopkins, “Statistical Methods in Education and Psychology”, PrenticeHall Inc., 1984. http://chem-eng.utoronto.ca/~datamining/ 39 Models based on Frequency Tables http://chem-eng.utoronto.ca/~datamining/ 40 ZeroR Classification Predictors Target Outlook Temp. Humidity Windy Play Golf Rainy Hot High False No Rainy Hot High True No Overcast Hot High False Yes Sunny Mild High False Yes Sunny Cool Normal False Yes Sunny Cool Normal True No Overcast Cool Normal True Yes Rainy Mild High False No Rainy Cool Normal False Yes Sunny Mild Normal False Yes Rainy Mild Normal True Yes Overcast Mild High True Yes Overcast Hot Normal False Yes Sunny Mild High True No http://chem-eng.utoronto.ca/~datamining/ 41 Classification - ZeroR Play Golf Play Golf No No No No Yes No Yes No Yes No No Yes Yes Sort 5 / 14 = 0.36 Yes No Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes No Yes 9 / 14 = 0.64 http://chem-eng.utoronto.ca/~datamining/ Yes 42 Classification - OneR Which one is the best predictor ? Outlook Temp Humidity Windy Play Golf Rainy Hot High False No Rainy Hot High True No Overcast Hot High False Yes Sunny Mild High False Yes Sunny Cool Normal False Yes Sunny Cool Normal True No Overcast Cool Normal True Yes Rainy Mild High False No Rainy Cool Normal False Yes Sunny Mild Normal False Yes Rainy Mild Normal True Yes Overcast Mild High True Yes Overcast Hot Normal False Yes Sunny Mild High True No http://chem-eng.utoronto.ca/~datamining/ 43 OneR - Frequency Tables Play Golf Outlook Play Golf Yes No Sunny 3 2 Overcast 4 0 Rainy 2 3 Temp. Yes No Hot 2 2 Mild 4 2 Cool 3 1 Play Golf Humidity Play Golf Yes No High 3 4 Normal 6 1 Windy Yes No False 6 2 True 3 3 http://chem-eng.utoronto.ca/~datamining/ 44 OneR – The Best Rule Play Golf Outlook Yes No Sunny 3 2 Overcast 4 0 Rainy 2 3 IF Outlook = Sunny THEN PlayGolf = Yes IF Outlook = Overcast THEN PlayGolf = Yes IF Outlook = Rainy THEN PlayGolf = No http://chem-eng.utoronto.ca/~datamining/ 45 Can we incorporate all predictors? Bayesian Classifier http://chem-eng.utoronto.ca/~datamining/ 46 Bayes’ Rule Likelihood Target Prior Probability PX | T PT PT | X PX Posterior Probability Predictor Prior Probability http://chem-eng.utoronto.ca/~datamining/ 47 Likelihood Tables Play Golf Outlook Yes No Sunny 3/9 2/5 Overcast 4/9 0/5 Rainy 2/9 3/5 Play Golf Temp. Yes No Hot 2/9 2/5 Mild 4/9 2/5 Cool 3/9 1/5 Play Golf Humidity Yes No High 3/9 4/5 Normal 6/9 1/5 Play Golf Windy Yes No False 6/9 2/5 True 3/9 3/5 P(Outlook Sunny | PlayGolf Yes ) 3 / 9 0.33 http://chem-eng.utoronto.ca/~datamining/ 48 Predictor Probability Play Golf Outlook Yes No Sunny 3/9 2/5 5/14 Overcast 4/9 0/5 4/14 Rainy 2/9 3/5 5/14 P(Outlook Sunny) 5 / 14 0.36 http://chem-eng.utoronto.ca/~datamining/ 49 Target Probability Play Golf Outlook Yes No Sunny 3/9 2/5 5/14 Overcast 4/9 0/5 4/14 Rainy 2/9 3/5 5/14 9/14 5/14 P( PlayGolf Yes ) 9 / 14 0.64 http://chem-eng.utoronto.ca/~datamining/ 50 Posterior Probability P(Outlook Sunny | PlayGolf Yes ) 3 / 9 0.33 P( PlayGolf Yes ) 9 / 14 0.64 P(Outlook Sunny) 5 / 14 0.36 P( PlayGolf Yes | Outlook Sunny) 0.33 0.64 0.36 0.60 PT | X PX | T PT PX http://chem-eng.utoronto.ca/~datamining/ 51 Bayesian - Prediction Case = [Outlook=Sunny; Humidity=High; Temp=Mild; Windy=True] Play Golf Posterior Probability Sunny Yes No 0.60 0.40 Outlook Temp. Play Golf Posterior Probability Humidity Posterior Probability High Yes No 0.43 0.57 Mild Posterior Probability Windy True Play Golf Yes No 0.67 0.33 Play Golf Yes No 0.50 0.50 P(PlayGolf = Yes) = 0.60 x 0.43 x 0.67 x 0.50 = 0.08643 P(PlayGolf = No) = 0.40 x 0.57 x 0.33 x 0.50 = 0.03762 http://chem-eng.utoronto.ca/~datamining/ 52 Naïve Bayesian– Multiple Predictors P(Ti | X) P( x1 | Ti ) P( x2 | Ti ) P( xn | Ti ) P(Ti ) Strong Independence assumption about predictors http://chem-eng.utoronto.ca/~datamining/ 53 Can we incorporate all predictors + Strong dependency? Decision Trees http://chem-eng.utoronto.ca/~datamining/ 54 Decision Tree Outlook Sunny Overcast Rainy Windy Play Humidity FALSE TRUE High Normal Play Not Play Not Play Play http://chem-eng.utoronto.ca/~datamining/ 55 Entropy Entropy = -p log2p – q log2q Entropy = -0.5 log20.5 – 0.5 log20.5 = 1 http://chem-eng.utoronto.ca/~datamining/ 56 Entropy – Frequency c E ( S ) pi log 2 pi i 1 Entropy (5,3,2) = Entropy (0.5,0.3,0.2) = - (0.5 * log20.5) - (0.3 * log20.3) - (0.2 * log20.2) = 1.49 http://chem-eng.utoronto.ca/~datamining/ 57 Entropy - Target Play Golf Play Golf No No No No Yes No Yes No Yes No No Yes Yes Sort 5 / 14 = 0.36 Yes No Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes No Yes 9 / 14 = 0.64 Entropy(PlayGolf) = Entropy (5,9) = Entropy (0.36, 0.64) = - (0.36 log2 0.36) - (0.64 log2 0.64) = 0.94 http://chem-eng.utoronto.ca/~datamining/ 58 Frequency Tables Play Golf Outlook Play Golf Yes No Sunny 3 2 Overcast 4 0 Rainy 2 3 Temp. Yes No Hot 2 2 Mild 4 2 Cool 3 1 Play Golf Humidity Play Golf Yes No High 3 4 Normal 6 1 Windy Yes No False 6 2 True 3 3 http://chem-eng.utoronto.ca/~datamining/ 59 Entropy – Frequency Table Play Golf Outlook Yes No Sunny 3 2 5 Overcast 4 0 4 Rainy 2 3 5 14 E (T , X ) P(v) E (v) vX c E (v) pi log 2 pi i 1 E(PlayGolf, Outlook) = P(Sunny)*E(3,2) + P(Overcast)*E(4,0) + P(Rainy)*E(2,3) = (5/14)*0.971 + (4/14)*0.0 + (5/14)*0.971 = 0.693 http://chem-eng.utoronto.ca/~datamining/ 60 Information Gain Gain(T , X ) Entropy (T ) Entropy (T , X ) G(PlayGolf, Outlook) = E(PlayGolf) – E(PlayGolf, Outlook) = 0.940 – 0.693 = 0.247 http://chem-eng.utoronto.ca/~datamining/ 61 Information Gain – the best predictor? Play Golf Outlook Play Golf Yes No Sunny 3 2 Overcast 4 0 Rainy 2 3 Temp. Yes No Hot 2 2 Mild 4 2 Cool 3 1 Gain = 0.247 Gain = 0.029 Play Golf Humidity Play Golf Yes No High 3 4 Normal 6 1 Windy Yes No False 6 2 True 3 3 Gain = 0.152 Gain = 0.048 http://chem-eng.utoronto.ca/~datamining/ 62 Decision Tree – Root Node Outlook Sunny Overcast http://chem-eng.utoronto.ca/~datamining/ Rainy 63 Dataset – Divide and Conqure Outlook Temp. Humidity Windy Play Golf Sunny Mild High FALSE Yes Sunny Cool Normal FALSE Yes Sunny Cool Normal TRUE No Sunny Mild Normal FALSE Yes Sunny Mild High TRUE No Rainy Hot High FALSE No Rainy Hot High TRUE No Rainy Mild High FALSE No Rainy Cool Normal FALSE Yes Rainy Mild Normal TRUE Yes Overcast Hot High FALSE Yes Overcast Cool Normal TRUE Yes Overcast Mild High TRUE Yes Overcast Hot Normal FALSE Yes http://chem-eng.utoronto.ca/~datamining/ 64 Subset (Outlook = Sunny) Temp. Humidity Windy Play Golf Mild High FALSE Yes Cool Normal FALSE Yes Mild Normal FALSE Yes Cool Normal TRUE No Mild High TRUE No Outlook Sunny Overcast Windy Play FALSE TRUE Play Not Play http://chem-eng.utoronto.ca/~datamining/ Rainy 65 Final Decision Tree Outlook Sunny Overcast Rainy Windy Play Humidity FALSE TRUE High Normal Play Not Play Not Play Play http://chem-eng.utoronto.ca/~datamining/ 66 Models based on Frequency Table Robust Measurable Scalable OneR Naïve Bayesian Decision Tree http://chem-eng.utoronto.ca/~datamining/ 67 Models based on Covariance Matrix http://chem-eng.utoronto.ca/~datamining/ 68 Models based on Covariance Matrix X2 X1 X3 Y http://chem-eng.utoronto.ca/~datamining/ 69 Covariance Covariance matrix (x C x )( y y ) n Mean x n http://chem-eng.utoronto.ca/~datamining/ 70 Multiple Linear Regression MLR Find b by minimizing e http://chem-eng.utoronto.ca/~datamining/ 71 MLR : Ordinary Least Square Intercept and Slopes: b ( X ' X ) X 'Y 1 Cov( x, y ) b Var ( x) Predicted Values: Y bX Residuals: Y Y http://chem-eng.utoronto.ca/~datamining/ 72 Regression Statistics SST (Y Y ) 2 SSR (Y Y ) SSE (Y Y ) http://chem-eng.utoronto.ca/~datamining/ 2 2 73 Regression Statistics How good is our model? SSR R SST 2 Coefficient of Determination to judge the adequacy of the regression model http://chem-eng.utoronto.ca/~datamining/ 74 ANOVA H 0 : b1 b 2 ... b k 0 H A : bi 0 Regression Residual Total at least one! df SS MS F P-value k SSR SSR / df MSR / MSE P(F) n-k-1 SSE SSE / df n-1 SST If P(F)<a then we know that we get significantly better prediction of Y from the regression model than by just predicting mean of Y. http://chem-eng.utoronto.ca/~datamining/ 75 Hypotheses Tests for Regression Coefficients H 0 : bi 0 H A : bi 0 t( n k 1) b1 b i bi b i 2 S e (bi ) S e Cii http://chem-eng.utoronto.ca/~datamining/ 76 Multicolinearity If the F-test for significance of regression is significant, but tests on the individual regression coefficients are not, multicolinearity may be present. Variance Inflation Factors (VIFs) are very useful measures of multicolinearity. If any VIF exceed 5, multicolinearity is a problem. 1 VIF ( b i ) Cii 2 1 Ri http://chem-eng.utoronto.ca/~datamining/ 77 Linear Regression – Binary Target Yi bo b1 X i e i • If actual Y is a binary variable the predicted Y can be less than zero or greater than 1. • If actual Y is a binary variable error is not normally distributed. http://chem-eng.utoronto.ca/~datamining/ 78 Linear Regression Model Y 1 0 X http://chem-eng.utoronto.ca/~datamining/ 79 Logistic Regression p 1 1 e ( b 0 b1 X ) http://chem-eng.utoronto.ca/~datamining/ 80 Frequency Table Months in Business Count <50 50-100 100-150 150-200 200-250 250-300 >300 4 12 4 4 4 1 4 Target Count 0 1 1 2 3 1 4 http://chem-eng.utoronto.ca/~datamining/ Target Probability 0 0.083 0.25 0.5 0.75 1 1 81 Frequency Plot 1 0.8 Target Probability 0.6 0.4 0.2 0 1 2 3 4 5 6 7 Months in Business - Bins http://chem-eng.utoronto.ca/~datamining/ 82 Logistic Function 1 f ( z) 1 ez http://chem-eng.utoronto.ca/~datamining/ 83 Logistic Regression p 1 1 e ( b 0 b1 X ) The logistic distribution constrains the estimated probabilities to lie between 0 and 1. Maximum Likelihood Estimation is a statistical method for estimating the coefficients of a model. http://chem-eng.utoronto.ca/~datamining/ 84 Logistic Regression Model Linear Model Y 1 Logistic Model 0 X http://chem-eng.utoronto.ca/~datamining/ 85 Log Likelihood (LL) • Likelihood is the probability that the dependent variable may be predicted from the independent variables. • LL is calculated through iteration, using maximum likelihood estimation (MLE). • Log likelihood is the basis for tests of a logistic model. http://chem-eng.utoronto.ca/~datamining/ 86 Wald Test • A Wald test is used to test the statistical significance of each coefficient (b) in the model. • A Wald test calculates a Z statistic, which is: Z b̂ SE • This Z value is then squared, yielding a Wald statistic with a chi-square distribution. http://chem-eng.utoronto.ca/~datamining/ 87 Linear Discriminant Analysis - LDA $250,000 $200,000 Z b1 x1 b 2 x2 $150,000 Non-Default Loan$ $100,000 Default $50,000 $0 0 10 20 30 40 50 60 70 Age http://chem-eng.utoronto.ca/~datamining/ 88 LDA – Training Z b1 x1 b 2 x2 ... b p x p b T 1 b T 2 S (b ) b T Cb S (b ) Score function Z1 Z 2 Variance of Z within groups http://chem-eng.utoronto.ca/~datamining/ 89 LDA – Training Mean x n Covariance matrix (x C x )( y y ) n http://chem-eng.utoronto.ca/~datamining/ 90 LDA – Training b C 1 (1 2 ) 1 C (n1C1 n2C2 ) n1 n2 b: Linear model coefficients C: Pooled covariance matrix 1 , 2 : Mean vectors http://chem-eng.utoronto.ca/~datamining/ 91 LDA – Model Assessment b (1 2 ) 2 : T Mahalanobis distance between two groups http://chem-eng.utoronto.ca/~datamining/ 92 LDA – Prediction 1 2 Z0 b 2 T Z1 b 1 T Z 2 b 2 T http://chem-eng.utoronto.ca/~datamining/ 93 Can we handle non-linearity? Support Vector Machines http://chem-eng.utoronto.ca/~datamining/ 94 Support Vector Machines- SVM $250,000 $200,000 Support Vectors $150,000 Non-Default Loan$ $100,000 Default $50,000 $0 0 10 20 30 40 50 60 70 Age http://chem-eng.utoronto.ca/~datamining/ 95 SVM and Linear Model http://chem-eng.utoronto.ca/~datamining/ 96 SVM and Linear Model n f ( x) w j x j b j 1 m f ( x) i k ( xi x) b i 1 http://chem-eng.utoronto.ca/~datamining/ 97 SVM – Kernel functions Polynomial k ( x i , x j ) ( x i .x j ) d Gaussian Radial Basis function x x i j k ( x i , x j ) exp 2 2 2 http://chem-eng.utoronto.ca/~datamining/ 98 Linear SVM Linear Proximal SVM Algorithm defined by Fung and Mangasarian: Given m data points in Rn represented by the m x n matrix A and a diagonal D of +1 and -1 labels denoting the class of each row of A, we generate the linear classifier as follows: where e is an m x 1 vectors of ones. Typically v is chosen by means of a tuning (validating) set. http://chem-eng.utoronto.ca/~datamining/ 99 Principal Component Analysis (PCA) Principal component analysis (PCA) is a classical statistical method. This linear transform has been widely used in data analysis and compression. The principal components (Eigenvectors) for a dataset can directly be extracted from the covariance matrix as follows: Eigenvectors and eigenvalues can be computed using a triangular decomposition module. By ordering the eigenvectors in the order of descending eigenvalues, we can find directions in which the data set has the most significant amounts of its variation. http://chem-eng.utoronto.ca/~datamining/ 100 Principal Component Regression (PCR) http://chem-eng.utoronto.ca/~datamining/ 101 Models based on Covariance Matrix Robust Measurable Scalable MLR LDA PCA/PCR Logistic Reg. SVM http://chem-eng.utoronto.ca/~datamining/ 102 Models based on Similarity http://chem-eng.utoronto.ca/~datamining/ 103 KNN - Definition KNN is a simple algorithm that stores all available cases and classifies new cases based on a similarity measure. http://chem-eng.utoronto.ca/~datamining/ 104 KNN – different names • • • • • • K-Nearest Neighbors Memory-Based Reasoning Example-Based Reasoning Instance-Based Learning Case-Based Reasoning Lazy Learning http://chem-eng.utoronto.ca/~datamining/ 105 KNN Classification $250,000 $200,000 $150,000 Non-Default Loan$ $100,000 Default $50,000 $0 0 10 20 30 40 50 60 70 Age http://chem-eng.utoronto.ca/~datamining/ 106 KNN Classification – Distance Age 25 35 45 20 35 52 23 40 60 48 33 Loan $40,000 $60,000 $80,000 $20,000 $120,000 $18,000 $95,000 $62,000 $100,000 $220,000 $150,000 Default N N N N N N Y Y Y Y Y 48 $142,000 ? Euclidean Distance Distance 102000 82000 62000 122000 22000 124000 47000 80000 42000 78000 D ( x1 x2 ) ( y1 y2 ) 2 http://chem-eng.utoronto.ca/~datamining/ 8000 2 107 KNN Classification – Standardized Distance Age Loan 0.125 0.375 0.625 0 0.375 0.8 0.075 0.5 1 0.7 0.325 0.11 0.21 0.31 0.01 0.50 0.00 0.38 0.22 0.41 1.00 0.65 Default N N N N N N Y Y Y Y Y 0.7 0.61 ? Distance 0.7652 0.5200 0.3160 0.9245 0.3428 0.6220 0.6669 0.4437 0.3650 0.3861 0.3771 X Min Xs Max Min http://chem-eng.utoronto.ca/~datamining/ 108 KNN – Number of Neighbors • If K=1, select the nearest neighbor • If K>1, – For classification select the most frequent neighbor. – For regression calculate the average of K neighbors. – What is the optimal K? http://chem-eng.utoronto.ca/~datamining/ 109 Models based on Similarity Robust Measurable Scalable KNN http://chem-eng.utoronto.ca/~datamining/ 110 Neural Networks http://chem-eng.utoronto.ca/~datamining/ 111 Biological Neuron and Integrated Circuit http://chem-eng.utoronto.ca/~datamining/ 112 Biological Neuron Synapse http://chem-eng.utoronto.ca/~datamining/ 113 Neural Network - Neuron (1) Summation I i w ji x j j w1i w ji f yi wni (2) Transfer yi f ( I i ) http://chem-eng.utoronto.ca/~datamining/ 114 Transfer Functions http://chem-eng.utoronto.ca/~datamining/ 115 Weight Adjustment or Error Propagation Wi ( D Y ) X i w1i w ji f e wni http://chem-eng.utoronto.ca/~datamining/ 116 Neural Networks Robust Measurable Scalable Neural Networks http://chem-eng.utoronto.ca/~datamining/ 117 Scalable Modeling (Classification & Regression) Frequency Covariance Table Matrix OneR Bayesian Markov Chains Linear Regression LDA (Z Score) PCA/PCR HMM http://chem-eng.utoronto.ca/~datamining/ 118 Scalable Data Mining: Summary Data Mining Query Manager Classification Frequency Table Regression Covariance Matrix Clustering Database Association File System I/O System http://chem-eng.utoronto.ca/~datamining/ 119 Scalable Data Mining Parallel and Distributed Processing Dataset Dataset 1 Dataset 2 Dataset 3 Dataset N Processor 1 Processor 4 Processor 3 Processor N BET 1 BET 2 BET 3 BET N Processor N+1 BET 1+2+3+…+N Modeling http://chem-eng.utoronto.ca/~datamining/ 120 Scalable Data Mining Chip! Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet bet Bet Bet Bet Bet Bet Memory Size for 10,000 Variables 10K x 10K x 8 bytes = 800MB x 5 = 4GB http://chem-eng.utoronto.ca/~datamining/ 121 Scalable Data Mining Multidimensional, Heterogeneous and Asymmetric Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet Bet bet Bet Bet Bet Bet Bet http://chem-eng.utoronto.ca/~datamining/ 122 Demo… http://chem-eng.utoronto.ca/~datamining/ 123