Probability & Statistics for ML
Probability, random variables, distributions, inference and information theory: the reasons behind every ML loss function.
37 lessons · about 42 hours
What you'll be able to do
- Model uncertainty with discrete and continuous distributions
- Apply Bayes' theorem
- Derive maximum-likelihood and MAP estimators
- Run and interpret hypothesis tests
- Explain why ML losses are what they are (MSE, cross-entropy, KL divergence)
Before you start
Course 2 Module 5 (integration) for continuous distributions; Course 1 Module 6 for the multivariate Gaussian.
Syllabus
Module 1
Counting and probability foundations
Permutations and combinations, sample spaces and events, the axioms of probability, mutually exclusive vs independent events.
In ML: Sampling, data splits
4 lessons coming soon
Module 2
Conditional probability and Bayes
Joint, marginal and conditional probability, the law of total probability, Bayes' theorem, conditional independence.
In ML: Naive Bayes; diagnostic reasoning
4 lessons coming soon
Module 3
Discrete random variables
Random variables, PMF and CDF, expectation and variance; Bernoulli, binomial, discrete uniform, geometric and Poisson distributions.
In ML: Click and count data; classification outputs
4 lessons coming soon
Module 4
Continuous random variables
PDF and CDF; uniform, exponential, normal and standard normal distributions; transformations of random variables; t and chi-squared distributions (introduction).
In ML: Noise models; weight initialisation
4 lessons coming soon
Module 5
Multiple random variables
Joint, marginal and conditional distributions, covariance and correlation, conditional expectation and variance, the multivariate Gaussian and its covariance ellipse.
In ML: Feature correlations; Gaussian models
5 lessons coming soon
Module 6
Limit theorems and sampling
Markov and Chebyshev inequalities, the law of large numbers, the central limit theorem, sampling distributions.
In ML: Why averages over mini-batches work
3 lessons coming soon
Module 7
Statistical inference
Descriptive statistics, point estimation, bias and variance of estimators, MLE, MAP, confidence intervals, z-, t- and chi-squared tests, p-values.
In ML: MSE as Gaussian MLE, cross-entropy as Bernoulli MLE; A/B testing
5 lessons coming soon
Module 8
Information theory
Entropy, cross-entropy, KL divergence, mutual information.
In ML: Why cross-entropy is the classification loss; information gain in trees
3 lessons coming soon
Module 9
Probabilistic models
The exponential family, Gaussian mixture models, latent variables, the EM algorithm, and Bayesian networks with exact and sampling-based inference.
In ML: Clustering, density estimation, reasoning under uncertainty
5 lessons coming soon