Building pipelines
that hold up.
London-based Data Engineer and Data Scientist. First Class BSc, Birkbeck, University of London.
The work,
the thinking.
I am a data professional based in London, holding a First Class BSc in Data Science & Computing from Birkbeck, University of London. My experience of studying part-time while developing real-world projects has significantly influenced my engineering approach: I prioritise deliberate, methodical processes, and I maintain a strong emphasis on creating systems that are resilient under pressure.
My expertise lies in data engineering and analytics engineering, where I focus on building reliable, observable pipelines and systems that generate trustworthy data for informed decision-making. I specialise in the entire pipeline lifecycle, from raw data ingestion and orchestration (using tools like Airflow and CDC with Debezium/Redpanda) to transformation and warehousing (leveraging dbt, DuckDB, and BigQuery), as well as data quality enforcement (with Great Expectations) and lineage tracking (utilising OpenLineage).
In the realm of data science, I leverage machine learning to deliver genuine predictive insights. My focus includes financial time-series forecasting utilising LSTM networks, statistical modelling, and comprehensive model evaluation. I approach machine learning as an engineering discipline, emphasising reproducible workflows, transparent backtesting, and results that can be confidently endorsed.
My work spans domains such as finance, sports, and urban/environmental data, where data quality and pipeline reliability are paramount. My projects encompass real-time data from Transport for London (TfL), analytics warehouses for the Premier League and Formula 1, air quality monitoring in London, and equity price forecasting.
Top grades from
Birkbeck.
First Class Honours · BSc Data Science & Computing · Birkbeck, University of London
Where I've worked.
Designed and maintained production-grade end-to-end data pipelines to ingest, clean, and model member and event engagement data using Python, SQL, dbt, and Apache Airflow — improving data freshness from 6 hours to 45 minutes (87% improvement) and ensuring 95% of loads met a 1-hour freshness SLA.
Automated data collection and reporting workflows for workshops, hackathons, and startup events, eliminating ~65% of manual reporting tasks and expediting time-to-insight for weekly stakeholder updates.
Developed self-service dashboards and reporting packages highlighting KPIs — attendance trends, participation rates, and program impact — reducing manual reporting by ~18 hours per month.
Partnered with cross-functional leaders to establish metrics and enforce reporting SLAs; achieved 99% on-time weekly delivery with end-to-end pipeline latency under 60 minutes at the 95th percentile, contributing to a 22% increase in attendance for flagship programs.
Provided structured 1-on-1 mentoring through the Caawi Mentorship Platform, guiding cohorts of ~12 early-career candidates in data engineering, portfolio development, and job readiness. Built the TfL Real-Time Lakehouse as a community teaching tool, giving users and students hands-on experience with production-style data pipelines — improving their confidence and understanding of real-world data engineering.
Pre-processed large-scale, multi-source datasets using Python and SQL, implementing deduplication and schema validation to reduce data defects by 30% and cut data preparation time from 10 hours to 5 hours per reporting cycle.
Automated data extraction and established daily refresh pipelines, improving stakeholder turnaround from 3 days to same-day delivery — a 67% efficiency gain.
Conducted comprehensive EDA to identify key operational drivers — lead-time variance, fulfilment delays, returns, and margin leakage — producing 12+ actionable recommendations adopted by stakeholders to drive data-driven decision-making.
Applied feature engineering techniques (lateness flags, supplier segmentation) to enhance signal separation between high- and low-performing suppliers by 15%.
Built interactive Power BI KPI dashboards for real-time monitoring, reducing manual reporting by 16 hours per month, growing adoption to 25+ users across operations and commercial teams, and contributing to an 8% reduction in logistics costs and 12% increase in on-time delivery.