Visualisation: Movie Map
A 2-D interactive map of 40,000 films and documentaries, grouped by similarity using 626 dimensional vectors, UMAP and HDBSCAN.
A 2-D interactive map of 40,000 films and documentaries, grouped by similarity using 626 dimensional vectors, UMAP and HDBSCAN.
A large scale exploration of GP access and DNA (did not attend) rates across South Yorkshire, and the link these statistics have to rates of deprivation.
This project, which was my Master’s dissertation, explored how Markov chains can explain the mathematics of shuffling, and how long it takes a deck to truly be shuffled.
Analysis of a dataset of confirmed exoplanets. Cleaned and visualised detection data across different telescopes and facilities, and performed an investigation into detection methods.
Comparison of 6 classification models, from Logistic Regression to CNNs, to predict handwritten digits from the MNIST dataset.
Built and compared 7 regression models to predict red wine quality from it’s chemical qualities, using methods from simple linear regression to Random Forest and XGBoost.