Some thoughts of a Machine Learning Practitioner on Software Development, Management, Team Building, Startups, Python, Agile Development, Data visualization... that will distract you from your end goals by making you less efficient but are critical to manage in order to succeed. Don't forget that long time adaptation to inefficient approaches can become your enemy. Let's try to empower others by sharing knowledge & personal experiences.
Monday, November 16, 2009
Knowledge Workers - Talent is not patient, and it is not faithful
Of course HR, managers nor VPs aren't struggling to compensate critical components when some of them leave but underlying teams have to. Those decision makers have to remember that they aren't free lunch and mismanagement of knowledge workers have some consequences.
Better management of Knowledge Workers lead to much more productive teams and low turn over but the opposite could cause your decline. What's so hard about creating a win/win approach instead of a loose/loose approach that so many companies seems to fall in. Maybe a generation clash? or simply missing competencies.
Wednesday, September 2, 2009
What's the relationship between Machine Learning and Data-Mining
- Data-Mining (DM) is the process of extracting patterns from data. The main goal is to understand relationships, validate models or identify unexpected relationships.
- Machine Learing (ML) algorithms allows computer to learn from data. The learning process consist of extracting the patterns but the end goal is to use the knowledge to do prediction on new data.
Sunday, July 5, 2009
digipy 0.1.1 - Hand Digit Real Time Demo is available
At Montreal-Python6, I have presented a real-time hand digit real-time demo.This demo allows you to do real-time digit recognition from your digital camera. It allows you to load any trained neural network and apply in real time the same features extraction. This demo allows you to train, extract features, used trained neural networks inside real-time demo, visualize features in 2D and their frequency distribution and get feature discriminant weight.
The packaging 0.1.1 of the demo is now available on pypi:
(unfortunately, some dependency packages aren't supported by easy_install so you have to do 4 steps instead of 1)
- install opencv (sudo aptitude install python2.5-opencv)
- install PyQt (sudo aptitude install pyqt4-dev-tools)*
- instal matplotlib (sudo aptitude install python2.5-matplotlib)*
- sudo easy_install digipy
* unfortunatly, this package isn't supported with easy_install
Here is the noise robustness comparison of the trained neural network on the raw pixels vs extracted features (digit surface + image convolution with train digits means (0-9)):
If you aren't convinced that Feature Extraction is absolutely required now, I have failed.
Once installed, you will get access to those command line tools:
- digipy: Real-time hand digit recognition demo application (ex: digipy --test)
- digipy-features2D : demo of feature 2D visualization to see possible clusters
- digipy-train: demo training of a Neural Network using mlboost
- digipy-compare: compare noise effect on test error on raw inputs and feature extracted datasets
- digipy-freq-analysis: demo feature analysis (frequency distributions)
- digipy-extract-features: demo features extraction
- digipy-see-data: show dataset train and test samples





If you have any trouble using it, just let me know (fpieraut at gmail).
Source code is available here http://bitbucket.org/fraka6/digipy.
(note: now digipy use mlboost.nn module for its NeuralNetwork instead of mlboost.flayers swig wrapper)
Wednesday, June 17, 2009
ICML highlights summary
- Language acquisition: Children loose their capacity to distinguish some phonemes to reduce the scope of choices in order to learn their environment language. The aquisition of phonemes categories from a buttom up approach isn't sufficient (signal processing+unsupervised clustering), a lexical minimal pair (ex:ngram) seems to be required to ensure the learning.
- Trying to learn the best kernel that restrict optimization to a convex problem seems to be a death end. It might be time change paradigm or move to the non-convex dark side.
- Boosting is too sensitive to noise but a robust framework has been presented by Yoav Freund
- Deep Architecture seems to be the next big thing. Regularisation, auto encoder and RBF can be used to pre-train networks from un-label data. Temporal coherence (similarity of consecutive frames in video) can be used as a regulation unsupervised technique in the embedding space. Unsupervised training is a regularisation technique that enforce better clustering. The more unlabeled unsupervised examples are used, the better will be the generalization.
- Training from IID samples isn't optimal, curriculum learning (i.e.: increase examples complexity) seems to smooth the cost function and lead to faster training and better generalization.
- GPU is the way to go to make ML algo scalable.
- Feature hashing is an efficient strategy for dimensionality reduction and can be used to train classifiers.
- Sparse transformation simplify the optimisation process (i.e: same idea used in the Kernel trick in SVN). PCA is doing the opposite.
Wednesday, June 10, 2009
International Conference on Machine Learning (ICML2009)
The 3 invited speakers are quite interesting. I look forward to heard them:
- Emmanuel Dupoux, from Ecole Normale Superieure on: How do infants bootstrap into spoken language?
- Yoav Freund, University of California on Drifting games, boosting and online learning?
- Corinna Cortes, from Google on can learning kernels help performance?
I have do decide to which tutorials I will attend this sunday:
- T6 Machine Learning in IR: Recent Successes and New Opportunities [tutorial webpage]Paul Bennett, Misha Bilenko, and Kevyn Collins-Thompson
- T8 Large Social and Information Networks: Opportunities for ML [tutorial webpage]Jure Leskovec
- T9 Structured Prediction for Natural Language Processing [tutorial webpage]Noah Smith
Here are some interesting papers:
- Curriculum Learning [Full paper]
- Deep Learning from Temporal Coherence in Video [Full paper]
- Good Learners for Evil Teachers [Full paper]
- Using Fast Weights to Improve Persistent Contrastive Divergence [Full paper]
- Online Dictionary Learning for Sparse Coding [Full paper]
- A Novel Lexicalized HMM-based Learning Framework for Web Opinion Mining [Full paper]
- A Scalable Framework for Discovering Coherent Co-clusters in Noisy Data [Full paper]
- Bayesian Clustering for Email Campaign Detection [Full paper]
- Feature Hashing for Large Scale Multitask Learning [Full paper]
- Grammatical Inference as a Principal Component Analysis Problem [Full paper]
- Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations[Full paper]
I have to decide between thoses workshops on Thursday:
- Workshop on Learning Feature Hierarchies
- On-line Learning with Limited Feedback
- Seventh Annual Workshop on Bayes Applications
- Sparse Methods for Music Audio
I look forward to meet old collegues, friends and new researchers. Next week will be awesome.
Friday, May 29, 2009
Is Python really slow? A practical comparison with C++
./fexp / -h 100 -l 0.01 --oh -e 10
...
Optimization: Standard
Creating Connector [16|100] [inputs | hiddens]
Creating Connector [100|26] [hiddens | outputs]
...
real 0m11.187s
user 0m10.837s
sys 0m0.012s
Here is the time to do 10 iteration on the full letters dataset with python and numpy:timetime ./bpnn.py -e 10 --h 100 -f letters.dat -nCreation of an NN <16:100:26>...
real 85m48.646s
user 85m9.163s
sys 0m1.632s
./bpnn.py -e 10 --h 100 -f letters.datCreation of an NN <16:100:26>...real 1m37.066suser 1m36.026ssys 0m0.100s
- The numpy implementation is 60 time faster then a basic python implementation.
- My C++ implementation is a little more then 10 time faster then my simply python numpy implementation.
Sunday, May 10, 2009
Short Essay: Engineering vs Scientist
Already, during a chat with a research professor in 2001 about the competence war between the engineering and the science department of university of Montreal, I got an initial hint about it, he told me that part of computer science department tension with software engineers was about placement rate. Engineers was much higher than computer scientists. I should have request an explanation. What is this war about, the computer science department request engineer to do some normalization courses even if they have great grades and vis-versa.


