Tuesday, April 30, 2013

Simplifying clustering visualization with mlboost


Are you looking for a simple way to visualized your supervised or semi-supervised data clusters with different dimension reduction algorithms like PCA, LDA, isomap, LLE ,mds, random trees, spectral embedding  etc.?
Here is an output example on 4 newsgroups dataset.

If you are following sklearn loading standard, with mlboost, you can do it by changing 2 lines of code (line #5 and #6) or modify this example. (python yourvisu.py -m y)

1
2
3
4
5
6
7
import sys
from mlboost.clustering import visu

# add your data loading function that return data_train and data_test
from X import LOAD_DATASET_Y
visu.add_loading_dataset_fct('y', LOAD_DATASET_Y)
visu.main(sys.argv[1:])
Btw, if you click on the legend, it will remove the class as you can see here when I remove the green class 2. In the context of semi-supervised, simply set samples class to "?" (dataset.target[i]). 
 

Without scikit-learn and matplotlib, it won't be that easy to experiment visualization. 

Thursday, April 25, 2013

How to remap your keyboard on linux if you drop water that has change your keyboard key mapping? ex: pressing 'c' key and getting '+' display....

Last week, I drop some liquid on my keyboard and a weird thing happened, my 'c' key mutated to a + which is quite annoying.
 After some research I fall on the linux screw FAQ: How to disable/remap a keyboard key in Linux? but it wasn't clear enough so here is what you should do.

  1. run the keycode command xev in your prompt
  2. press on the key that has a bad mapping
  3. identify the keycode in the xev prompt: state 0x10, keycode 86 (keysym 0x63, c), ....
  4. change the keycode mapping : xmodmap -e 'keycode 86=c C'
  5. add (4) cmd in your .bachrc otherwise you will need to redo it in every terminal
As mentioned in the linux screw FAQ, 'c' key should be the keycode 54 not 86 so I just remap the 86 keycode to c and C (when shift is press).

Quite happy I can use my keyboard normally again.  


Saturday, April 6, 2013

Finding the optimal K in kmean: a incremental kmeans in python

I was looking for an good implementation of an incremental k-means where I don't have to set the optimal K. There are interesting papers (x-means, gmeans etc.) but couldn't find any python implementation.

I have decided to write a incremental version on top of sklearn.
The idea is simple:
  1. Start at K=x
  2. identify worst cluster based on an unsupervised measure (ex: silhouette)
  3. Split the worst cluster into 2 clusters
  4. measure the global improvement with the new clusters
  5. if you get an improvement continue adding clusters
You can find the source code in mlboost/clustering/ikmeans.py
A special thanks to scikit-learn lib to let me prototype this version so fast. 

Thursday, April 4, 2013

kmeans with configurable distance function: How to hack sklearn.kmeans to use a different distance function?

Like others, I was looking for a good k-means implementation where I can set the distance function.
Unfortunately, sklearn.kmeans doesn't allow to change the cost function. Euclidean is hard coded. 
As pointed out by the sklearn team, it is quite complex to generalize due to:
So if your distance function is cosine which has the same mean as euclidean, you can monkey patch  sklearn.cluster.k_means_.eucledian_distances this way: (put this code before calling kmean.fit(X).


from sklearn.metrics.pairwise import cosine_similarity
def new_euclidean_distances(X, Y=None, Y_norm_squared=None, squared=False) 
    return cosine_similarity(X,Y)

# monkey patch (ensure cosine dist function is used)
from sklearn.cluster import k_means_k_means_.euclidean_distances 
k_means_.euclidean_distances = new_euclidean_distances 


Warning: you need to normalize your input vectors. 

Thursday, March 21, 2013

Google Analytics Hidden Business Model

What is GA (Google Analytics) Business Model?

Analytics is used by google to improve adwords product (most of google revenue). Nothing is free.

Here is the most common reason why GA is free:
  • Provide accountability and transparency to existing Google advertisers (adwords clients)
  • Provide confidence and prove the value of online advertising to potential new advertisers
Here is 2 another ones?
  • Provide advertiser tools to speed their site (real-user monitoring)
  • Increase ads relevance  (hidden)
Why real-user monitoring?
Since 2011, google ranking consider site speed. Faster are sites, more search people will be done.

How can GA be used by google to increase ads relevance?

In order to understand this concept, let's look at google eco-system.

If a user visit several site monitored by google Analytics, GA information can be leverage to better recommend ads anywhere and improve google search ads.

With GA, you get quite interesting information about users: their behavior, the content and sites they visit, their frequency and time spent and this across multiple sites. GA is an amazing source of user information. 

In order to get better customized ads, you need:
  • session data (behavior, site visited, content type, frequency, etc.)
  • aggregate data to user data
  • get more information about the user (no simply an id, a google+ or gmail account) 
Once you get the data, the biggest challenge is merging the sessions. In order to do it, your need a way to merge the sessions. Here is the ways google does it:
  • Google analytics (same sessions id between sites)
  • gmail login (get a uniq id)
  • Chrome browser (control browser, control browser global id)
  • google+ and gmail account (user info)

The next biggest challenge is leveraging the social graph, which explain google+ war against facebook. 
The ad platform has to leverage GA, gmail, google+ etc. to improve its search.

Friday, November 23, 2012

Real-time face recognition experiment packages


In Autumn 2009, I have been lecturer at ETS for a Machine Learning introduction class. In order to ensure the class could get a real feeling about machine learning, I have repackaged the digipy demo used for my presentation "Machine Learning Empowered by Python" for their final project. The latest code is here: https://bitbucket.org/fraka6/digiface
The digiface package was their recommended starting point. It is a real-time face recognition package so they could focus on extracting the best features, train easily a single neural net and experiment live or on the dataset picture.
The idea was simple, they will compete on the best real-time live face recognitions of the student faces themselves. Every student had to sit in front of each other face recognition system. The best system had to be quite robust in order to consider light, background et hair changes.
We had to make a pictures sessions and build the dataset.
One team built their own package called digijava. Here is a snapshot.

It was quite an interesting teaching experiment. I am glad to see that some of my student have followed my path and join Yoshua Bengio lisa great lab.

If you are looking for a great talk about the latest state of the art in machine learning, look at that Hinton "Brains, Sex, and Machine Learning" youtube video and Yoshua Bengio slides "DeepLearning of Representations"(Google talk).

Wednesday, November 21, 2012

How to conquer your customer segment? The beachhead framework



Want to conquer your customer segments once you reach your business market fit? The beachhead framework is for you. Unfortunately it is hard finding a template nor good examples online. Many sites explain the idea but how to apply it?
It's simple, once you have identify a customer segment, identify product and influential leverage strategies and prioritized them.
Product leverage is an action that leverage your product as oppose to customization (ex: a wordpress plugin).
Influence Leverage is an action that leverage a customer segment influencer (ex: gartner, forrester research, blogger etc.).

Don't forget:
  1. BMC: Busines Model Canvas (Business Model = Value + Monetization)
  2. Validation Board (lean startup)
  3. BMF: Business Market Fit -> Branding your company -> Setting the Industry agenda
  4. BF: Beachhead framework