<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Machine Learning on When Moore's Law Ends</title><link>https://jimwang99.github.io/posts/machine-learning/</link><description>Recent content in Machine Learning on When Moore's Law Ends</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Tue, 20 Oct 2015 00:00:00 +0000</lastBuildDate><atom:link href="https://jimwang99.github.io/posts/machine-learning/index.xml" rel="self" type="application/rss+xml"/><item><title>[Cousera Note] Machine Learning Foundations A Case Study Approach</title><link>https://jimwang99.github.io/posts/machine-learning/cousera-note-machine-learning-foundations-a-case-study-approach/</link><pubDate>Tue, 20 Oct 2015 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/cousera-note-machine-learning-foundations-a-case-study-approach/</guid><description>&lt;h2 id="1-week1-welcome">1 Week1: Welcome&lt;a class="anchor" href="#1-week1-welcome">#&lt;/a>&lt;/h2>
&lt;h3 id="11-introduction">1.1 Introduction&lt;a class="anchor" href="#11-introduction">#&lt;/a>&lt;/h3>
&lt;h4 id="111-real-world-case-based">1.1.1 real world case based&lt;a class="anchor" href="#111-real-world-case-based">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>regression: house price prediction&lt;/li>
&lt;li>classificiation: sentiment analysis&lt;/li>
&lt;li>clustering &amp;amp; retrieval: finding doc&lt;/li>
&lt;li>maxtrix factorization &amp;amp; dimensionality reduction: recommending products&lt;/li>
&lt;/ul>
&lt;h4 id="112-requirement">1.1.2 requirement&lt;a class="anchor" href="#112-requirement">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>math: calculas &amp;amp; algebra&lt;/li>
&lt;li>python&lt;/li>
&lt;/ul>
&lt;h4 id="113-capstone-project">1.1.3 capstone project&lt;a class="anchor" href="#113-capstone-project">#&lt;/a>&lt;/h4>
&lt;h3 id="12-ipython-notebook">1.2 iPython Notebook&lt;a class="anchor" href="#12-ipython-notebook">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>Python command and its outputs&lt;/li>
&lt;li>Markdown for doc&lt;/li>
&lt;/ul>
&lt;h3 id="13-sframes">1.3 SFrames&lt;a class="anchor" href="#13-sframes">#&lt;/a>&lt;/h3>
&lt;h4 id="131-graphlab-canvas">1.3.1 GraphLab Canvas&lt;a class="anchor" href="#131-graphlab-canvas">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>any data structure *.show() » data visualization web page. make it inline of iPython Notebook:
&lt;ul>
&lt;li>graphlab.canvas.set_target(‘ipynb’)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>create new column
&lt;ul>
&lt;li>sf[‘Full Name’] = sf[‘First Name’] + &amp;rsquo; &amp;rsquo; + sf[‘Last Name’]&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>apply function
&lt;ul>
&lt;li>sf[‘Country’] = sf[‘Country’].apply(transform_country)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="2-week2-regression-case-predicting-house-prices">2 Week2: Regression case: predicting house prices&lt;a class="anchor" href="#2-week2-regression-case-predicting-house-prices">#&lt;/a>&lt;/h2>
&lt;h3 id="21-linear-regression-modeling">2.1 Linear regression modeling&lt;a class="anchor" href="#21-linear-regression-modeling">#&lt;/a>&lt;/h3>
&lt;h4 id="211-recent-sales-in-nearby-neighbourhood">2.1.1 recent sales in nearby neighbourhood&lt;a class="anchor" href="#211-recent-sales-in-nearby-neighbourhood">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>x = sqft, feature / covariate / predictor / indepentent varaible&lt;/li>
&lt;li>y = price, observation / response&lt;/li>
&lt;/ul>
&lt;h4 id="212-look-at-average-price-in-range">2.1.2 look at average price in range&lt;a class="anchor" href="#212-look-at-average-price-in-range">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>limited hits&lt;/li>
&lt;li>throwing out other information: it’s bad&lt;/li>
&lt;/ul>
&lt;h4 id="213-linear-regression">2.1.3 linear regression&lt;a class="anchor" href="#213-linear-regression">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>fw(x) = w0 + w1*x
&lt;ul>
&lt;li>w0 = intercept&lt;/li>
&lt;li>w1 = slope&lt;/li>
&lt;li>w0/w1 are parameters of our model, w = (w0,w1); also called regression coefficients&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>different parameter set w; choose w is important&lt;/li>
&lt;li>&lt;strong>RSS = residual sum of squares&lt;/strong>
&lt;ul>
&lt;li>delta = Y of observation - Y from prediction&lt;/li>
&lt;li>RSS = sum(all possible delta^2)&lt;/li>
&lt;li>minimize RSS(w0,w1) and solve it, get you w’&lt;/li>
&lt;li>then y’ = w0’ + w1’ * x of my house&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="214-adding-higher-order-effects">2.1.4 adding higher order effects&lt;a class="anchor" href="#214-adding-higher-order-effects">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>straight line is good enough?
&lt;ul>
&lt;li>maybe not a linear relationship, rather a quadratic function&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>quadratic
&lt;ul>
&lt;li>fw(x) = w0 + w1&lt;em>x + w2&lt;/em>x^2
&lt;ul>
&lt;li>still linear regression, x^2 is just another feature&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>higher order polynomial? 13th order polynomial to minimize RSS
&lt;ul>
&lt;li>this function just looks crazy&lt;/li>
&lt;li>overfitting&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="22-evaluating-regression-models">2.2 Evaluating regression models&lt;a class="anchor" href="#22-evaluating-regression-models">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>evaluating overfitting via training/test split
&lt;ul>
&lt;li>min RSS » bad prediction&lt;/li>
&lt;li>how to choos model order/complexity?
&lt;ul>
&lt;li>goal: good predictions&lt;/li>
&lt;li>simulate predictions
&lt;ul>
&lt;li>step 1. remove some data points&lt;/li>
&lt;li>step 2. fit model on remaining&lt;/li>
&lt;li>step 3. predict heldout houses
&lt;ul>
&lt;li>use model to predict those removed data points (in step 1), and see how accurate they are&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>need extra test data&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>terminology
&lt;ul>
&lt;li>training set / test set&lt;/li>
&lt;li>training error = RSS of all training data points, minimize it to find w’&lt;/li>
&lt;li>test error = RSS of all test data points with w’&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>training/test curves: model complexity vs. error
&lt;ul>
&lt;li>training curves: the higher model complexity is, smaller the error gets&lt;/li>
&lt;li>test curves: probably will look like a U, which has optimized lowest value&lt;/li>
&lt;li>&lt;code>w2_training_test_curves.png&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>add other features
&lt;ul>
&lt;li>for houses as an example, # of bedrooms as x2
&lt;ul>
&lt;li>fitting an 3D surface&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>how many features to use? unlimited! hold there and more info in the “regression” course&lt;/li>
&lt;li>always more feature the better to capture underlying process? NO&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>other regression examples
&lt;ul>
&lt;li>stock prediction
&lt;ul>
&lt;li>recent history&lt;/li>
&lt;li>news events&lt;/li>
&lt;li>related commodities&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>temp of smart house
&lt;ul>
&lt;li>spatial function&lt;/li>
&lt;li>thermostat setting/ blinds / window / vents / temp outside / time of day&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="23-summary">2.3 Summary&lt;a class="anchor" href="#23-summary">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>regression ML block diagran ML pipepline: data » ML method » intelligence
&lt;ul>
&lt;li>&lt;code>w2_regression_ML_model.png&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="24-predict-house-prices-ipython-notebook-example">2.4 Predict house prices (iPython Notebook example)&lt;a class="anchor" href="#24-predict-house-prices-ipython-notebook-example">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>Loading &amp;amp; exploring house sale data
&lt;ul>
&lt;li>graphlab.SFrame(‘xxx.gl.zip’)
&lt;ul>
&lt;li>SFrame: table data structure in graphlab&lt;/li>
&lt;li>here xxx.gl.zip is some presist data dumped out by graphlab&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>SFrame.show(view=&amp;ldquo;Scatter_plit”, x=&amp;ldquo;column1”, y=&amp;ldquo;column2”)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Split data into training and test data sets
&lt;ul>
&lt;li>SFrame.random_split(float,seed=int) » (training_data, test_data)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Build regression model
&lt;ul>
&lt;li>model = graphlab.linear_regression.create(traning_data, target=&amp;lsquo;column1’, features=list(‘column2’, ‘column3’))
&lt;ul>
&lt;li>target: variable you try to predict&lt;/li>
&lt;li>features&lt;/li>
&lt;li>algorithm chosed automatically&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Evaluating error
&lt;ul>
&lt;li>model.evaluate(test_data)
&lt;ul>
&lt;li>a simple model gives high error and high RMSE (root mean square error)&lt;/li>
&lt;li>RMSE = (RSS / 2)^(1/2)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Visualizing with Matplotlib
&lt;ul>
&lt;li>matplotlib.pyplot.plot(list_of_x, list_of_y1, ‘.&amp;rsquo;, list_of_x, list_of_y2, ‘-&amp;rsquo;)&lt;/li>
&lt;li>model.predict(test_data)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Inspect coefficients
&lt;ul>
&lt;li>model.get(‘coeffcients’)
&lt;ul>
&lt;li>intercept is w0&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Explore other features
&lt;ul>
&lt;li>use multiple features, other than only sqft&lt;/li>
&lt;li>SFrame.Show(view=&amp;lsquo;BoxW Plot’, x=, y=)&lt;/li>
&lt;li>6 features give you less max error and less RMSE&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Apply models to particular data points
&lt;ul>
&lt;li>sales[sales[‘id’]==&amp;lsquo;53xxx’]&lt;/li>
&lt;li>for multiple feature model, even on average we have better RMSE, but on some particular data points it could have larger error number.&lt;/li>
&lt;li>add image in iPython Notebook
&lt;ul>
&lt;li>&lt;em>[Example image not included in the archive]&lt;/em>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="25-homework">2.5 Homework&lt;a class="anchor" href="#25-homework">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>SArray
&lt;ul>
&lt;li>immutable array object&lt;/li>
&lt;li>each column in an SFrame is an SArray&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Filter
&lt;ul>
&lt;li>logical filter&lt;/li>
&lt;li>.apply()&lt;/li>
&lt;li>a selection in SFrame takes a list consists of 0 and 1. And the length equals to the length of SFrame’s num of rows. When it comes to 0, row is ignored; otherwise it’s taken.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="3-week3-classification-analyzing-sentiment">3 Week3: Classification: Analyzing Sentiment&lt;a class="anchor" href="#3-week3-classification-analyzing-sentiment">#&lt;/a>&lt;/h2>
&lt;h3 id="31-classification-modeling">3.1 Classification modeling&lt;a class="anchor" href="#31-classification-modeling">#&lt;/a>&lt;/h3>
&lt;h4 id="311-intelligent-restaurant-review">3.1.1 intelligent restaurant review&lt;a class="anchor" href="#311-intelligent-restaurant-review">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>break review into sentences&lt;/li>
&lt;li>sentence sentiment classifier&lt;/li>
&lt;/ul>
&lt;h4 id="312-classifier">3.1.2 classifier&lt;a class="anchor" href="#312-classifier">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>input = x » output = y&lt;/li>
&lt;li>input can have multi information&lt;/li>
&lt;li>output can be multi categories, called multicalss classification&lt;/li>
&lt;li>example
&lt;ul>
&lt;li>review sentiment&lt;/li>
&lt;li>webpage category: output tech, sport, news, …&lt;/li>
&lt;li>spam filtering: input has multi information&lt;/li>
&lt;li>image classification&lt;/li>
&lt;li>personalized medical diagnosis: input DNA and life style&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="313-linear-classifier">3.1.3 linear classifier&lt;a class="anchor" href="#313-linear-classifier">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>simple threshold classifier
&lt;ul>
&lt;li>count pos/neg words in sentence&lt;/li>
&lt;li>problems
&lt;ul>
&lt;li>how to get list of pos/neg word&lt;/li>
&lt;li>word degree of sentiment&lt;/li>
&lt;li>single words not enough&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>give weight for each word&lt;/li>
&lt;li>score = sum of input words’ weight, so it’s linear&lt;/li>
&lt;li>if score &amp;gt; 0, output = pos, else output = neg&lt;/li>
&lt;/ul>
&lt;h4 id="314-decision-boundaries">3.1.4 decision boundaries&lt;a class="anchor" href="#314-decision-boundaries">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>decision boundary separates pos/neg predictions&lt;/li>
&lt;/ul>
&lt;h3 id="32-evaluating-classification-models">3.2 Evaluating classification models&lt;a class="anchor" href="#32-evaluating-classification-models">#&lt;/a>&lt;/h3>
&lt;h4 id="321-training-and-eval-a-classifier">3.2.1 training and eval a classifier&lt;a class="anchor" href="#321-training-and-eval-a-classifier">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>traing set for learn classifer » to get the weight of words&lt;/li>
&lt;li>test set to eval the weight of words. hide the label, feed sentence to classifier, compare prediction with real label.&lt;/li>
&lt;li>classification error
&lt;ul>
&lt;li>error = (# of mistakes) / (total # of sentences)&lt;/li>
&lt;li>accuracy = (# of corrects) / (total # of sentences)&lt;/li>
&lt;li>error = 1.0 - accuracy&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="322-whats-a-good-accuracy">3.2.2 what’s a good accuracy?&lt;a class="anchor" href="#322-whats-a-good-accuracy">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>accuracy should beat random guess, larger than 1/K (K is the number of classes)&lt;/li>
&lt;li>class imbalance will give you good performance. accuracy should beat majority class baseline, in which we simple guess everything is from the majority class. Eg. 80% of the reviews are pos, so it’s baseline of majority class is 80% since everytime we guess a review is pos anyway.&lt;/li>
&lt;li>most importantly: how accurate the application need? what accuracy will make the user happy?&lt;/li>
&lt;/ul>
&lt;h4 id="323-confusion-matrices">3.2.3 confusion matrices&lt;a class="anchor" href="#323-confusion-matrices">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>correct: true pos / true neg; mistake: false neg / false pos&lt;/li>
&lt;li>false neg and false pos can have different impact in some application. eg email spam filter, medical diagnosis&lt;/li>
&lt;/ul>
&lt;h4 id="324-learning-curves">3.2.4 learning curves&lt;a class="anchor" href="#324-learning-curves">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>how much data does a model need to learn?
&lt;ul>
&lt;li>the more the better, but data quality is most important factor&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>learning curve
&lt;ul>
&lt;li>x = amount of training data&lt;/li>
&lt;li>y = test error&lt;/li>
&lt;li>limit? yes, for most models&lt;/li>
&lt;li>bias = even with infinite data, test error will not got to zero&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>complex models tend to have less bias, but need more data to learn&lt;/li>
&lt;li>bias is not possible to eliminate&lt;/li>
&lt;/ul>
&lt;h4 id="325-class-probabilities">3.2.5 class probabilities&lt;a class="anchor" href="#325-class-probabilities">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>class probablity = how confident is your prediction (soft ouptut!)&lt;/li>
&lt;li>many classifier provide a confidence level&lt;/li>
&lt;/ul>
&lt;h3 id="33-summary">3.3 Summary&lt;a class="anchor" href="#33-summary">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>w3_summary.png&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="34-analyzing-sentiment-with-ipython-notebook">3.4 Analyzing sentiment with iPython Notebook&lt;a class="anchor" href="#34-analyzing-sentiment-with-ipython-notebook">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>graphlab.text_analytics.count_words(SFrame)
&lt;ul>
&lt;li>count_bigrams&lt;/li>
&lt;li>count_trgrams&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>SFrame.show(view=&amp;lsquo;Categorical’)&lt;/li>
&lt;li>data engineering: define pos/neg sentiment by throught 3-star reviews out&lt;/li>
&lt;li>graphlab.logistic_classifier.create(train_data, target=&amp;lsquo;sentiment’, features=[‘word_count’], validation_set=test_data)&lt;/li>
&lt;li>model.evaluate(test_data, metric = ‘roc_curve’)
&lt;ul>
&lt;li>roc_curve help to explore confusion matrics&lt;/li>
&lt;li>change the threshold to get different rate of true pos vs false pos, help you to choose different strategy for different application requirements&lt;/li>
&lt;li>it looks like the result is very good &lt;strong>JUST&lt;/strong> according to the word count! The word count is not the number of words in the review, but a count of different words in the review.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>model.predict(SFram, output_type=&amp;lsquo;propability’)&lt;/li>
&lt;li>SFrame.sort(‘predicted_sentiment’, ascending=False)&lt;/li>
&lt;li>.apply() is very limited because its function only takes 1 argument, itself&lt;/li>
&lt;li>model[‘coefficients’]&lt;/li>
&lt;/ul>
&lt;h2 id="4-week4-clustering-and-similarity-retrieving-documents">4 Week4: Clustering and Similarity: Retrieving Documents&lt;a class="anchor" href="#4-week4-clustering-and-similarity-retrieving-documents">#&lt;/a>&lt;/h2>
&lt;h3 id="41-algorigthms-for-retrieval-and-measuring-similarity-of-documents">4.1 Algorigthms for retrieval and measuring similarity of documents&lt;a class="anchor" href="#41-algorigthms-for-retrieval-and-measuring-similarity-of-documents">#&lt;/a>&lt;/h3>
&lt;h4 id="411-problem-definition">4.1.1 problem definition&lt;a class="anchor" href="#411-problem-definition">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>how to measure similarity?&lt;/li>
&lt;li>how to search through ariticle?&lt;/li>
&lt;/ul>
&lt;h4 id="412-word-count-representation-for-measuring-similarity">4.1.2 word count representation for measuring similarity&lt;a class="anchor" href="#412-word-count-representation-for-measuring-similarity">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>bag of words model
&lt;ul>
&lt;li>ignore order&lt;/li>
&lt;li>count number of words in vocabulary&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>measure similarity: sum(x_i * y_i), where i is the word index in vacabulary&lt;/li>
&lt;li>issue with word counts: doc length matters! doesn’t make sense, because prefer longer article&lt;/li>
&lt;li>solution = normalize: x’_i = x_i / (sum(x_i^2)^(1/2)&lt;/li>
&lt;/ul>
&lt;h4 id="413-word-importance-priority-with-tf-idf">4.1.3 word importance priority with tf-idf&lt;a class="anchor" href="#413-word-importance-priority-with-tf-idf">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>common words vs rare words: emphasize important words even they are rare.&lt;/li>
&lt;li>important word
&lt;ul>
&lt;li>common locally&lt;/li>
&lt;li>rare globally&lt;/li>
&lt;li>trade off between these 2&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>TF-IDF: term frequency - inverse document frequency&lt;/li>
&lt;li>term freq = word counts&lt;/li>
&lt;li>inverse doc freq, look at all the doc in our corpus = log [# docs / (1 + # docs using this word)]
&lt;ul>
&lt;li>common word, idf-&amp;gt;0&lt;/li>
&lt;li>rare word, idf is large&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>tf * idf
&lt;ul>
&lt;li>down weight common words&lt;/li>
&lt;li>up weight rare words&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="414-nearest-neighbor-search-to-retrieve-similar-document">4.1.4 nearest neighbor search to retrieve similar document&lt;a class="anchor" href="#414-nearest-neighbor-search-to-retrieve-similar-document">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>distance metric
&lt;ul>
&lt;li>search each article in corpus&lt;/li>
&lt;li>compute similarity&lt;/li>
&lt;li>return largest 1 or N similarity article(s)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="42-clusterinng-models-and-algorithms">4.2 Clusterinng models and algorithms&lt;a class="anchor" href="#42-clusterinng-models-and-algorithms">#&lt;/a>&lt;/h3>
&lt;h4 id="421-overview">4.2.1 overview&lt;a class="anchor" href="#421-overview">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>discover groups (clusters) of related articles&lt;/li>
&lt;li>training set: labeled docs&lt;/li>
&lt;li>multiclass classification problem&lt;/li>
&lt;/ul>
&lt;h4 id="422-clustering-documents-without-supervise">4.2.2 clustering documents without supervise&lt;a class="anchor" href="#422-clustering-documents-without-supervise">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>unsupervised learning
&lt;ul>
&lt;li>no labels provided&lt;/li>
&lt;li>want to uncover cluster structure&lt;/li>
&lt;li>input: docs as vectors. This will put each article as a dot in a vector space. In the class, we assume a 2-D space with X=# of word_1, Y=# of word_2&lt;/li>
&lt;li>output: label (cluster)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>what defines a cluster?
&lt;ul>
&lt;li>center&lt;/li>
&lt;li>shape/spread&lt;/li>
&lt;li>assign observation (doc) to cluster (topic label). (1) score (2) distance to cluster center&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="423-k-means-algorithm">4.2.3 k-means algorithm&lt;a class="anchor" href="#423-k-means-algorithm">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>similarity = distance to cluster centers&lt;/li>
&lt;li>algorithm
&lt;ul>
&lt;li>initialize cluster centers by “randomly”&lt;/li>
&lt;li>assign observations to closest cluster center by “voronoi tessellation”&lt;/li>
&lt;li>revise cluster centers as mean of assigned observations&lt;/li>
&lt;li>repeat step (2)+(3) until convergence&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="424-other-examples">4.2.4 other examples&lt;a class="anchor" href="#424-other-examples">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>clustering images&lt;/li>
&lt;li>grouping patients by medical condition&lt;/li>
&lt;li>production recommendation on Amazon&lt;/li>
&lt;li>discovering groups of users on Amazon&lt;/li>
&lt;li>structuring web research results. multiple meanings of one word&lt;/li>
&lt;li>discovering similar neighborhoods: house price prediction (not enough sales data) / forecase violent crimes&lt;/li>
&lt;/ul>
&lt;h3 id="43-summary">4.3 Summary&lt;a class="anchor" href="#43-summary">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>iteratively update our cluster centers (parameters of clustering)&lt;/li>
&lt;li>&lt;code>w4_summary.png&lt;/code>&lt;/li>
&lt;li>My questions
&lt;ul>
&lt;li>How about the thesaurus? We need to take them into consideration.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="44-doument-retrieval-in-python">4.4 Doument retrieval in python&lt;a class="anchor" href="#44-doument-retrieval-in-python">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>graphlab.text_analytics.count_words(); # uni-gram counting&lt;/li>
&lt;li>SFrame.stack($column_name, new_column_name=list) » a new stack SFrame table (expand the value of given column, and copy the other columns)
&lt;ul>
&lt;li>&lt;a href="https://dato.com/products/create/docs/generated/graphlab.SFrame.stack.html">https://dato.com/products/create/docs/generated/graphlab.SFrame.stack.html&lt;/a>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="441-tf-idf">4.4.1 TF-IDF&lt;a class="anchor" href="#441-tf-idf">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>tf_idf = graphlab.text_analytics.tf_idf(people[‘word_count’])&lt;/li>
&lt;/ul>
&lt;h4 id="442-distance-matric">4.4.2 distance matric&lt;a class="anchor" href="#442-distance-matric">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>graphlab.distances.*; # lots of options to choose from to calculate distance metric graphlab.distances.cosine; # smaller the closer&lt;/li>
&lt;/ul>
&lt;h4 id="443-nearest-neighbor-model">4.4.3 nearest neighbor model&lt;a class="anchor" href="#443-nearest-neighbor-model">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>knn_model = graphlab.nearest_neighbors.create(people, features=[‘tf_idf’], lable=&amp;lsquo;name’)&lt;/li>
&lt;li>knn_model.query(obama); # return the nearest entry, obama here is SArray for SFrame ‘people’&lt;/li>
&lt;/ul>
&lt;h2 id="5-week5-recommending-products">5 Week5: Recommending Products&lt;a class="anchor" href="#5-week5-recommending-products">#&lt;/a>&lt;/h2>
&lt;h3 id="51-recommender-system">5.1 Recommender system&lt;a class="anchor" href="#51-recommender-system">#&lt;/a>&lt;/h3>
&lt;h4 id="511-overview">5.1.1 overview&lt;a class="anchor" href="#511-overview">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>use past history and other user’s history in prediction&lt;/li>
&lt;li>recommender system in action
&lt;ul>
&lt;li>personalization: what do I care about? because of information overload; connect users and items&lt;/li>
&lt;li>movie recommendations: what want to watch?&lt;/li>
&lt;li>product recommendations: global and session interests&lt;/li>
&lt;li>music recommendations: coherent and diverse sequence&lt;/li>
&lt;li>friend recommendations: users and “items” are of the same “type”&lt;/li>
&lt;li>drug-target interactions: what drug should we “repurpose” for some disease? asprin from headache to blood thinner in heart condition&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="512-recommender-system-via-classification">5.1.2 recommender system via classification&lt;a class="anchor" href="#512-recommender-system-via-classification">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>solution 0: popularity
&lt;ul>
&lt;li>rank by global popularity; no personalization&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>solution 1: classification model
&lt;ul>
&lt;li>use features of items and users&lt;/li>
&lt;li>input : user info + purchase history + production info + other info&lt;/li>
&lt;li>pros: personalized; features can capture context (time of the day, what I just saw), even handles limited user history&lt;/li>
&lt;li>cons: features may not be available; collaborative filtering cannot work&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>solution 2: collaborative filter&lt;/li>
&lt;/ul>
&lt;h3 id="52-co-occurrence-matrices-for-collaborative-filtering">5.2 Co-occurrence matrices for collaborative filtering&lt;a class="anchor" href="#52-co-occurrence-matrices-for-collaborative-filtering">#&lt;/a>&lt;/h3>
&lt;h4 id="521-collaborative-filtering">5.2.1 collaborative filtering&lt;a class="anchor" href="#521-collaborative-filtering">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>people who bought this also bought …&lt;/li>
&lt;li>Matrix C: store # users who bought both items i &amp;amp; j
&lt;ul>
&lt;li>x = y = # items&lt;/li>
&lt;li>symmetric: C_ij = C_ji&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>How to use Matrix C?
&lt;ul>
&lt;li>look at row i, which user just bought&lt;/li>
&lt;li>recommend other items in the row with &lt;strong>largest&lt;/strong> counts&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="522-effect-of-popular-items">5.2.2 effect of popular items&lt;a class="anchor" href="#522-effect-of-popular-items">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>no matter what I just purchased, most popular item will be recommended, will drowns out other effects&lt;/li>
&lt;/ul>
&lt;h4 id="523-normalizing-co-occurrence-matrices-and-leveraging-purchase-history">5.2.3 normalizing co-occurrence matrices and leveraging purchase history&lt;a class="anchor" href="#523-normalizing-co-occurrence-matrices-and-leveraging-purchase-history">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>Jaccard similarity: normalizes by popularity
&lt;ul>
&lt;li>both i and j / i or j = C_ij / (C_i + C_j - C_ij) = S_ij&lt;/li>
&lt;li>limitations: no history; what if purchased many items&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Weighted average of purchased items
&lt;ul>
&lt;li>purchased item j and k&lt;/li>
&lt;li>S(i) = avg(S_ij + S_ik)&lt;/li>
&lt;li>chose highest S&lt;/li>
&lt;li>limitations: no context; no user features; no product features; new user/product (cold start problem)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="53-matrix-factorization">5.3 Matrix factorization&lt;a class="anchor" href="#53-matrix-factorization">#&lt;/a>&lt;/h3>
&lt;h4 id="531-matrix-completion-task">5.3.1 matrix completion task&lt;a class="anchor" href="#531-matrix-completion-task">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>solution 3: discovering hidden structure by matrix factorization&lt;/li>
&lt;li>use movie recommendation as example&lt;/li>
&lt;li>matrix of rating
&lt;ul>
&lt;li>x = movies&lt;/li>
&lt;li>y = user&lt;/li>
&lt;li>value = rating of movie x by user y&lt;/li>
&lt;li>if user y hasn’t watched movie x, then use ? (white square)&lt;/li>
&lt;li>goal: filling white squares, how much user y will like movie x’ (not watched yet)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="532-useritem-features">5.3.2 user/item features&lt;a class="anchor" href="#532-useritem-features">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>movie topics Rv for movie v&lt;/li>
&lt;li>user prefer topics Lu for user u&lt;/li>
&lt;li>rating(u,v) = Rv * Lu = (element vise product)&lt;/li>
&lt;li>recommendation: sort rating(u,v)
&lt;ul>
&lt;li>rating will be out of a certain range&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="533-predictions-in-matrix-form">5.3.3 predictions in matrix form&lt;a class="anchor" href="#533-predictions-in-matrix-form">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>rating matrix takes all the users and all the movies, every element is a rating(u,v)&lt;/li>
&lt;/ul>
&lt;h4 id="534-discovering-hidden-structure-by-matrix-factorization-model">5.3.4 discovering hidden structure by matrix factorization model&lt;a class="anchor" href="#534-discovering-hidden-structure-by-matrix-factorization-model">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>HOWEVER we don’t have the features of users and movies, we have to discover topics from data&lt;/li>
&lt;li>use observed value to estimate Lu and Rv: regression
&lt;ul>
&lt;li>RSS(L,R) = sum(rating(u,v), )^2, where Lu and Rv are estimated from model parameters R &amp;amp; L, and sum are for all the black squares (with data)&lt;/li>
&lt;li>RSS(L,R) gives L and R from regression&lt;/li>
&lt;li>then use L and R to predict rating(u,v) for white squares&lt;/li>
&lt;li>many efficient algorithms for factorization&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>limitation of matrix factorization
&lt;ul>
&lt;li>cold-start problem&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="535-all-together-featurized-matrix-factorization">5.3.5 all together: featurized matrix factorization&lt;a class="anchor" href="#535-all-together-featurized-matrix-factorization">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>blending model
&lt;ul>
&lt;li>feature: context&lt;/li>
&lt;li>matrix factorization: groups of users&lt;/li>
&lt;li>combine: feature for new users; as more info discovered, use matrix factorization topics&lt;/li>
&lt;li>Netflix Prize 1M dollars: winning team blended over 100 models&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="54-performance-metrics-for-recommender-systems">5.4 Performance metrics for recommender systems&lt;a class="anchor" href="#54-performance-metrics-for-recommender-systems">#&lt;/a>&lt;/h3>
&lt;h4 id="541-performance-metric">5.4.1 performance metric&lt;a class="anchor" href="#541-performance-metric">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>classification accuracy&lt;/li>
&lt;li>interested in what user like, but not “user does not like”&lt;/li>
&lt;li>fast vs. full list&lt;/li>
&lt;li>recall = # liked and shown / # liked&lt;/li>
&lt;li>precision = # liked and shown / # shown
&lt;ul>
&lt;li>how much “garbage” (things i’m not interested in) i need to look at&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="542-optimal-recommenders">5.4.2 optimal recommenders&lt;a class="anchor" href="#542-optimal-recommenders">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>maximize recall? recommend everything, but will give very small precision&lt;/li>
&lt;li>optimal: recommend things I like, and only the things I like&lt;/li>
&lt;/ul>
&lt;h4 id="543-precision-recall-curves">5.4.3 precision-recall curves&lt;a class="anchor" href="#543-precision-recall-curves">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>input = specific recommender system&lt;/li>
&lt;li>output = algorithm-specific precision-recall curve&lt;/li>
&lt;li>x = # items recommended&lt;/li>
&lt;li>&lt;code>precision-recall_curves.png&lt;/code>&lt;/li>
&lt;li>which algorithm is best?
&lt;ul>
&lt;li>given precision, better recall&lt;/li>
&lt;li>given recall, better precision&lt;/li>
&lt;li>metric 1: largest area under the curve (AUC)&lt;/li>
&lt;li>metric 2: precision at a specific # recommended items&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="55-summary">5.5 Summary&lt;a class="anchor" href="#55-summary">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>w5_summary.png&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="56-song-recommender-with-python">5.6 Song recommender with Python&lt;a class="anchor" href="#56-song-recommender-with-python">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>users = song_data[‘user_id’].unique() len(users)&lt;/li>
&lt;/ul>
&lt;h4 id="561-simple-popularity-based-recommender">5.6.1 simple popularity-based recommender&lt;a class="anchor" href="#561-simple-popularity-based-recommender">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>popularity_model = graphlab.popularity_recommender.create(training_data, user_id=&amp;lsquo;user_id’, item_id=&amp;lsquo;song’) popularity_mode.recommend(users=[users[0]]) popularity_mode.recommend(users=[users[1]])&lt;/li>
&lt;li>everyone get the exact same thing&lt;/li>
&lt;/ul>
&lt;h4 id="562-personalization-recommender">5.6.2 personalization recommender&lt;a class="anchor" href="#562-personalization-recommender">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>personalized_model = graphlab.item_similarity_recommender.create(training_data, user_id=&amp;lsquo;user_id’, item_id=&amp;lsquo;song’) personalized_model.recommend(users=[users[0]]) personalized_model.recommend(users=[users[1]])&lt;/li>
&lt;li>&lt;code># similar songs personalized_model.get_similar_items([‘song name here’])&lt;/code>&lt;/li>
&lt;li>&lt;code># similar users personalized_model.get_similar_users([‘user id here’])&lt;/code>&lt;/li>
&lt;/ul>
&lt;h4 id="563-qunatitative-comparison-between-the-models">5.6.3 Qunatitative comparison between the models&lt;a class="anchor" href="#563-qunatitative-comparison-between-the-models">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>model_performance = graphlab.recommender.util.compare_models(test_data, [popularity_model, personalized_model], user_sample=0.05)&lt;/li>
&lt;/ul>
&lt;h3 id="57-assignment">5.7 Assignment&lt;a class="anchor" href="#57-assignment">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>artist_popularity = song_data.groupby(key_columns=&amp;lsquo;artist’, operations={‘total_count’: graphlab.aggregate.SUM(‘listen_count’)})&lt;/li>
&lt;/ul>
&lt;h2 id="6-week6-deep-learning-searching-for-images">6 Week6: Deep Learning: Searching for Images&lt;a class="anchor" href="#6-week6-deep-learning-searching-for-images">#&lt;/a>&lt;/h2>
&lt;h3 id="61-neural-networks-learning-very-non-linear-features">6.1 Neural networks: Learning very non-linear features&lt;a class="anchor" href="#61-neural-networks-learning-very-non-linear-features">#&lt;/a>&lt;/h3>
&lt;h4 id="611-search-for-images">6.1.1 search for images&lt;a class="anchor" href="#611-search-for-images">#&lt;/a>&lt;/h4>
&lt;h4 id="612-what-is-a-visual-product-recommender">6.1.2 what is a visual product recommender?&lt;a class="anchor" href="#612-what-is-a-visual-product-recommender">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>keyword search? don’t know what keyword to search&lt;/li>
&lt;li>use image similarity to search for product&lt;/li>
&lt;/ul>
&lt;h4 id="613-learning-very-non-linear-features-with-neural-networks">6.1.3 learning very non-linear features with neural networks&lt;a class="anchor" href="#613-learning-very-non-linear-features-with-neural-networks">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>features (of images) are very important to neural networks.&lt;/li>
&lt;li>layers and layers of linear models and non-linear transformations&lt;/li>
&lt;li>in 90s disfavor&lt;/li>
&lt;li>big resurgence recent 10 years
&lt;ul>
&lt;li>lots of data to train&lt;/li>
&lt;li>computing resource like GPUs&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="62-deep-learning--deep-features">6.2 Deep learning &amp;amp; deep features&lt;a class="anchor" href="#62-deep-learning--deep-features">#&lt;/a>&lt;/h3>
&lt;h4 id="621-application-for-deep-learning-to-computer-vision">6.2.1 application for deep learning to computer vision&lt;a class="anchor" href="#621-application-for-deep-learning-to-computer-vision">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>image features
&lt;ul>
&lt;li>local detectors combined to make prediction&lt;/li>
&lt;li>image features of “interesting points”&lt;/li>
&lt;li>before we used to hand crafted features&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>standard image classification approach
&lt;ul>
&lt;li>input&lt;/li>
&lt;li>extract features (hand created)&lt;/li>
&lt;li>use simple classifier&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>deep learning: implicitly learns features
&lt;ul>
&lt;li>different layers detect different types of features&lt;/li>
&lt;li>automatically!&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="622-deep-learning-performance">6.2.2 deep learning performance&lt;a class="anchor" href="#622-deep-learning-performance">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>ImageNet 2012 competition: 1.2M training images, 1000 categories&lt;/li>
&lt;li>SuperVision use deeplearning neural network has a big gain against 2nd place&lt;/li>
&lt;/ul>
&lt;h4 id="623-demo-on-imagenet-data">6.2.3 demo on ImageNet data&lt;a class="anchor" href="#623-demo-on-imagenet-data">#&lt;/a>&lt;/h4>
&lt;h4 id="624-challenges">6.2.4 challenges&lt;a class="anchor" href="#624-challenges">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>pros
&lt;ul>
&lt;li>learning automatically rather than hand tuning&lt;/li>
&lt;li>performance gain&lt;/li>
&lt;li>potential&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>cons
&lt;ul>
&lt;li>lots of labeled data (human annotation)&lt;/li>
&lt;li>computationally expensive&lt;/li>
&lt;li>many tricks to tune&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="625-deep-features">6.2.5 deep features&lt;a class="anchor" href="#625-deep-features">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>can we learn features from data, even when we don’t have the data or time? deep learning + transfer learning: use data from one task to help learn on another&lt;/li>
&lt;li>what’s learned in a neural net?
&lt;ul>
&lt;li>very speicific to task 1 for the latest layers&lt;/li>
&lt;li>more generic for earlier layers, can be reuse&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>transfer learning
&lt;ul>
&lt;li>keep first few layers&lt;/li>
&lt;li>use simple classifier to replace last several layers that is too specific&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>deep features workflow
&lt;ul>
&lt;li>&lt;code>deep_features_workflow.png&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="63-summary">6.3 Summary&lt;a class="anchor" href="#63-summary">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>w6_summary.png&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="64-deep-features-for-image-classification">6.4 Deep features for image classification&lt;a class="anchor" href="#64-deep-features-for-image-classification">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>deep_learning_model = graphlab.load_model(‘imagenet_model’); # this is a pre-trained deep learning model using ImageNet’s 1.5M images&lt;/li>
&lt;li>image_train[‘deep_features’] = deep_learning_model.extract_features(image_train); # extract deep feature using pre-trained model&lt;/li>
&lt;li>deep_features_model = graphlab.logistic_classifier.create(image_train, features=[‘deep_features’], target=&amp;lsquo;labl’); # use simple classifier on extracted deep features&lt;/li>
&lt;/ul>
&lt;h3 id="65-deep-features-for-image-retrieval">6.5 Deep features for image retrieval&lt;a class="anchor" href="#65-deep-features-for-image-retrieval">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>knn_model = graphlab.nearest_neighbors.create(image_train, features=[‘deep_features’], label=&amp;lsquo;id’)&lt;/li>
&lt;li>​ cat = image_train[18]
knn_model.query(cat); # gives neighbors of given “cat” item&lt;/li>
&lt;/ul>
&lt;h3 id="66-assignment">6.6 Assignment&lt;a class="anchor" href="#66-assignment">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>use sketch_summary to get summary statitics of the data, only for SArray (not as assignment said for both SFrame and SArray)
&lt;ul>
&lt;li>image_train[‘label’].sketch_summary()&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="67-deploying-machine-learning-as-a-service">6.7 Deploying machine learning as a service&lt;a class="anchor" href="#67-deploying-machine-learning-as-a-service">#&lt;/a>&lt;/h3>
&lt;h4 id="671-whats-production-life-cycle">6.7.1 what’s production? life cycle&lt;a class="anchor" href="#671-whats-production-life-cycle">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>deployment: serving&lt;/li>
&lt;li>evaluation: measuring quality of deployed models&lt;/li>
&lt;li>management: choosing between deployed models&lt;/li>
&lt;li>monitoring: tracking model quality and operations&lt;/li>
&lt;/ul>
&lt;h4 id="672-deployment">6.7.2 deployment&lt;a class="anchor" href="#672-deployment">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>traning with historical data&lt;/li>
&lt;li>real-time predictions with live data&lt;/li>
&lt;li>feedback and improve&lt;/li>
&lt;/ul>
&lt;h4 id="673-3-other-pieces">6.7.3 3 other pieces&lt;a class="anchor" href="#673-3-other-pieces">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>learning new, alternative models&lt;/li>
&lt;li>how to choose between models&lt;/li>
&lt;li>evaluating a recommender: user engagement and user experience&lt;/li>
&lt;li>offline evaluation: when to update model&lt;/li>
&lt;li>online evaluation: choosing between models&lt;/li>
&lt;/ul>
&lt;h4 id="674-ab-testing-choosing-between-ml-models">6.7.4 A/B testing: choosing between ML models&lt;a class="anchor" href="#674-ab-testing-choosing-between-ml-models">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>group A use model 1 and group B use model 2&lt;/li>
&lt;li>other issues: versioning, provenace, dashboards, reports, …&lt;/li>
&lt;/ul>
&lt;h3 id="68-machine-learning-challenges-and-future-directions">6.8 Machine learning challenges and future directions&lt;a class="anchor" href="#68-machine-learning-challenges-and-future-directions">#&lt;/a>&lt;/h3>
&lt;h4 id="681-model-selection">6.8.1 model selection&lt;a class="anchor" href="#681-model-selection">#&lt;/a>&lt;/h4>
&lt;h4 id="682-feature-engineeringrepresentation">6.8.2 feature engineering/representation&lt;a class="anchor" href="#682-feature-engineeringrepresentation">#&lt;/a>&lt;/h4>
&lt;h4 id="683-scaling">6.8.3 scaling&lt;a class="anchor" href="#683-scaling">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>data is getting bigger and bigger
&lt;ul>
&lt;li>social website&lt;/li>
&lt;li>products on amazon&lt;/li>
&lt;li>devices of IoT&lt;/li>
&lt;li>medical record&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>models are getting bigger and bigger&lt;/li>
&lt;li>CPUs stopped getting faster
&lt;ul>
&lt;li>GPUs&lt;/li>
&lt;li>multicores&lt;/li>
&lt;li>clusters&lt;/li>
&lt;li>clouds&lt;/li>
&lt;li>supercomputers&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>parallel architecture
&lt;ul>
&lt;li>programmability&lt;/li>
&lt;li>data distribution&lt;/li>
&lt;li>failures&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul></description></item><item><title>A Demo for Image Search</title><link>https://jimwang99.github.io/posts/machine-learning/a-demo-for-image-search/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/a-demo-for-image-search/</guid><description>&lt;p>In this Github project, I created a simple application that can do text to image and image to image search, using open-source transformer model. Details can be found in the repo and its docs directory.&lt;/p>
&lt;p>Here is a screen recording of running this demo on my Apple silicon laptop.&lt;/p>
&lt;p>&lt;em>Screen recording not found in the available backups.&lt;/em>&lt;/p></description></item><item><title>AI-Acceleration</title><link>https://jimwang99.github.io/posts/machine-learning/ai-acceleration/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/ai-acceleration/</guid><description>&lt;h2 id="not-found">Not Found&lt;a class="anchor" href="#not-found">#&lt;/a>&lt;/h2>
&lt;p>File Publish/🌟AI-Acceleration.md does not exist.&lt;/p></description></item><item><title>DeepSeek</title><link>https://jimwang99.github.io/posts/machine-learning/deepseek/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/deepseek/</guid><description>&lt;p>DeepSeek has created a huge wave of discussion and panic in the market.&lt;/p>
&lt;p>&lt;strong>I think it’s great overall.&lt;/strong>&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Innovation&lt;/strong>&lt;/li>
&lt;/ol>
&lt;p>From a technology standpoint, DeepSeek is truly pushing the envelope in both training methods and model architecture. They’ve combined many good ideas from industry and academia and, more importantly enhanced them to create solutions that are both more capable and more cost-effective. Their work has single-handedly revived the MoE (Mixture of Experts) architecture, which made quite a splash some time ago but had since grown quiet. These advancements will undoubtedly give AI researchers a lot to think about.&lt;/p></description></item><item><title>Memory Interfaces for LLM</title><link>https://jimwang99.github.io/posts/machine-learning/memory-interfaces-for-llm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/memory-interfaces-for-llm/</guid><description>&lt;p>LLM is a memory bound problem. This inspired me to look at different memory technologies. In this article, I&amp;rsquo;m going to summarize my research these days, especially about HBM and its impact on AI applications.&lt;/p>
&lt;h2 id="history">History&lt;a class="anchor" href="#history">#&lt;/a>&lt;/h2>
&lt;p>Before we dive deep into SOTA (state of the art) memory interfaces, we need to understand the history briefly.&lt;/p>
&lt;h3 id="sdram-early-1990s">SDRAM (early 1990s)&lt;a class="anchor" href="#sdram-early-1990s">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>SDRAM = Synchronous Dynamic Random-Access Memory&lt;/li>
&lt;li>Comparing to DRAM chips before it, it added clock signals to make the interface synchronous (again)&lt;/li>
&lt;/ul>
&lt;h3 id="ddr-late-1990s---nowadays">DDR (late 1990s - nowadays)&lt;a class="anchor" href="#ddr-late-1990s---nowadays">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>DDR = Double Data Rate Synchronous Dynamic Random-Access Memory&lt;/li>
&lt;li>Comparing to SDRAM chips, it uses both the rising and falling edges of the clock to transmit data. Thus, doubled the data rate&lt;/li>
&lt;li>Latest DDR standard is DDR5-7200, whose highest transfer rate is 7200 MT/s&lt;/li>
&lt;/ul>
&lt;h3 id="lpddr">LPDDR&lt;a class="anchor" href="#lpddr">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>Low-power version DDR, made specially for mobile devices
&lt;ul>
&lt;li>LPDDR SDRAMs use lower bit-width (16 or 32 bits, versus 64-bit in DDR)&lt;/li>
&lt;li>They also use DVS (dynamic voltage scaling) to save power&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Latest LPDDR standard is LPDDR5, whose transfer rate is 6400 MT/s&lt;/li>
&lt;/ul>
&lt;h3 id="gddr">GDDR&lt;a class="anchor" href="#gddr">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>High-bandwidth version DDR, made specially for GPUs&lt;/li>
&lt;li>Latest GDDR standard is GDDR6W, whose transfer rate is 22 GT/s&lt;/li>
&lt;/ul>
&lt;h3 id="hbm">HBM&lt;a class="anchor" href="#hbm">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>HBM = High Bandwidth Memory&lt;/li>
&lt;li>Comparing to normal DDR&amp;rsquo;s normal DIMM or SODIMM package, HBM utilizes MCM (multi-chip module) packaging technology and 3D IC stacking technology, to achieve better bandwidth and power consumption comparing to all the other memory technologies&lt;/li>
&lt;li>The latest HBM3E can achieve 9.2Gbps per pin, with 1024 IO pins, single HBM3E chip can achieve 1.2TB/s&lt;/li>
&lt;/ul>
&lt;h2 id="technology-details">Technology details&lt;a class="anchor" href="#technology-details">#&lt;/a>&lt;/h2>
&lt;h3 id="hbm-1">HBM&lt;a class="anchor" href="#hbm-1">#&lt;/a>&lt;/h3>
&lt;p>There are 2 key technologies that enables HBM:&lt;/p></description></item><item><title>ML System Architecture (1) System Level Parallelization</title><link>https://jimwang99.github.io/posts/machine-learning/ml-system-architecture-1-system-level-parallelization/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/ml-system-architecture-1-system-level-parallelization/</guid><description>&lt;p>To optimize the efficiency of training or executing an ML model, whether implemented locally on a device or hosted in the cloud, parallelization plays a critical role, akin to other computational challenges.&lt;/p>
&lt;p>Utilizing multiple compute engines—ranging from several compute units on a single microchip to numerous GPUs within a data center—allows for the division of computational tasks and their parallel execution.&lt;/p>
&lt;h2 id="3-types-of-parallelization">3 Types of parallelization&lt;a class="anchor" href="#3-types-of-parallelization">#&lt;/a>&lt;/h2>
&lt;p>Ordered from coarse granularity to fine granularity.&lt;/p></description></item><item><title>Navigating the Landscape of Large Language Models</title><link>https://jimwang99.github.io/posts/machine-learning/navigating-the-landscape-of-large-language-models/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/navigating-the-landscape-of-large-language-models/</guid><description>&lt;blockquote class='book-hint '>
&lt;p>Here is my notes from &lt;a href="https://privatebank.jpmorgan.com/content/dam/jpm-wm-aem/global/cwm/en/insights/eye-on-the-market/good-bad-ugly-jpmwm.pdf">JPMorgan&amp;rsquo;s &amp;ldquo;Eye on Market&amp;rdquo; 2024 April issue&lt;/a>&lt;/p>&lt;/blockquote>&lt;p>The emergence and integration of large language models (LLMs) into various professional sectors have marked a significant milestone in the journey of artificial intelligence. As we delve into the practical applications and implications of these advanced technologies, a mixed picture of successes and challenges begins to emerge. This article aims to expand upon the initial observations of LLMs in the real world, drawing attention to their profound impacts, innovative applications, and the hurdles yet to be overcome.&lt;/p></description></item><item><title>NV's DIGITS for On-Premise AI</title><link>https://jimwang99.github.io/posts/machine-learning/nv-s-digits-for-on-premise-ai/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/nv-s-digits-for-on-premise-ai/</guid><description>&lt;p>&lt;strong>TL;DR&lt;/strong>&lt;/p>
&lt;p>Nvidia’s DIGITS offers an on-premise AI solution aimed at smaller organizations that require strict data privacy. While it is cost-effective for small businesses and college research labs, it may be less suitable for larger enterprises or for widespread adoption in typical households.&lt;/p>
&lt;hr>
&lt;p>From my perspective, Nvidia’s new DIGITS system, showcased at CES 2025, is designed to meet the needs of smaller organizations, such as doctors’ offices, law/CPA firms and academic research labs, that require AI capabilities but must keep privacy-sensitive data on-premise. Due to strict legal and regulatory constraints, these organizations often cannot upload proprietary information to cloud-based APIs (like those offered by OpenAI or Google). Instead, they need local storage, indexing, search, and AI computation.&lt;/p></description></item><item><title>Run ML Models Use Qualcomm SNPE Part 1 Canned Example</title><link>https://jimwang99.github.io/posts/machine-learning/run-ml-models-use-qualcomm-snpe-part-1-canned-example/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/run-ml-models-use-qualcomm-snpe-part-1-canned-example/</guid><description>&lt;p>#software #accelerator #ai #on-device&lt;/p>
&lt;p>&lt;em>SNPE = Snapdragon Neural Processing Engine&lt;/em>&lt;/p>
&lt;p>In this tutorial we assume that Qualcomm SNPE has been successfully installed use QPM. Follow &amp;ldquo;Qualcomm Package Manager 1.0&amp;rdquo; -&amp;gt; &amp;ldquo;AI stack&amp;rdquo; -&amp;gt; &amp;ldquo;Neural Processing SDK&amp;rdquo; and install.&lt;/p>
&lt;p>After installation, you can find its latest documentation at &lt;code>$SNPE_ROOT/docs/SNPE/html/general/index.html&lt;/code>. I&amp;rsquo;m using Ubuntu 20.04 on WSL2 on Windows 11.&lt;/p>
&lt;h2 id="setup-environment">Setup environment&lt;a class="anchor" href="#setup-environment">#&lt;/a>&lt;/h2>
&lt;ol>
&lt;li>Follow the &lt;strong>Setup&lt;/strong> Chapter of the document to install necessary tools and Python packages.&lt;/li>
&lt;li>If you just want to quickly run this example, you can skip the following details&lt;/li>
&lt;/ol>
&lt;h3 id="python-packages">Python packages&lt;a class="anchor" href="#python-packages">#&lt;/a>&lt;/h3>
&lt;p>I strongly recommend to use &lt;code>conda&lt;/code> to manage your Python virtual env, since SNPE suggest to use a particular Python version.&lt;/p></description></item><item><title>Run ML Models Use Qualcomm SNPE Part 2 `torchvision.models.resnext50`</title><link>https://jimwang99.github.io/posts/machine-learning/run-ml-models-use-qualcomm-snpe-part-2-torchvision-models-resnext50/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/run-ml-models-use-qualcomm-snpe-part-2-torchvision-models-resnext50/</guid><description>&lt;p>#software #accelerator #ai #on-device #WIP&lt;/p>
&lt;p>In this part of the tutorial, we will learn how to run a pretrained model from TorchVision: ResNeXt50, which is a model architecture built upon the concepts of ResNet.&lt;/p>
&lt;h2 id="preparation">Preparation&lt;a class="anchor" href="#preparation">#&lt;/a>&lt;/h2>
&lt;h3 id="companion-git-repo">Companion git repo&lt;a class="anchor" href="#companion-git-repo">#&lt;/a>&lt;/h3>
&lt;p>If you haven&amp;rsquo;t cloned the git repo,&lt;/p>
&lt;pre tabindex="0">&lt;code>git clone https://github.com/jimwang99/xrbench-snapdragon.git
cd xrbench-snapdragon/pytorch_model/resnext50&lt;/code>&lt;/pre>&lt;h3 id="imagenet-dataset">Imagenet dataset&lt;a class="anchor" href="#imagenet-dataset">#&lt;/a>&lt;/h3>
&lt;p>To run quantized model on device, we need to download the following ImageNet dataset from HuggingFace: &lt;a href="https://huggingface.co/datasets/imagenet-1k/blob/main/data">imagenet-1k&lt;/a>.&lt;/p>
&lt;p>To save download time, we can only choose the validation set.&lt;/p></description></item><item><title>Use PCA (Principal Component Analysis) to do Clustering of Embeddings</title><link>https://jimwang99.github.io/posts/machine-learning/use-pca-principal-component-analysis-to-do-clustering-of-embeddings/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/use-pca-principal-component-analysis-to-do-clustering-of-embeddings/</guid><description>&lt;p>Let&amp;rsquo;s consider a face recognition system, where we&amp;rsquo;ve got facial images from a list of known persons and the system input is a camera image. We need to figure out if there are people in this camera image from our list of known persons or not.
Face detection model is very commonly used. It gives bounding boxes of human faces and associated confidence numbers. We can use face detection model to find faces.
After faces are found and cropped, we then can use a face verification model to verify if input face is close to another reference face. The algorithm underneath is to represent the input face images as embedding vectors, and calculate the distance between these vectors.&lt;/p></description></item></channel></rss>