How to become a Data Scientist

Data Analyst / ML Engineer / Research assistant → Data Scientist

Time to hire
24–36m
Total cost USD
$1,200–$1,800
Total cost ZAR
R21,600–R32,400
Salary range
$90,000–$130,000
Domain: Data & AI · CP47
Last verified 2026-05-02

Role Overview

What does a Data Scientist actually do?

A Data Scientist is half researcher, half analyst. You spend your days exploring data to uncover patterns and insights, building statistical models to make predictions, and communicating findings to non-technical stakeholders. A week might involve analyzing customer churn to identify risk factors, building a predictive model, presenting findings to the VP of Sales, then deploying the model (or handing to an ML engineer). You're answering "What could happen?" and "Why did it happen?" rather than just "What happened?" Tools: Python (Pandas, NumPy, scikit-learn, statsmodels), SQL, R (sometimes), Jupyter, visualization libraries, Git, cloud platforms.

Data Scientists work on teams of 2–10, often embedded in business units. The role is remote-friendly (70%+). You're rarely on-call. You collaborate with analysts (who use your insights), engineers (who deploy your models), and business stakeholders. This is a technical role requiring statistical thinking and strong communication.

Demand in 2026

  • Global job postings: 25,000+ active Data Scientist roles on LinkedIn as of May 2026 (source)
  • Growth rate: 14% YoY / Strong demand but plateauing compared to ML engineer (source)
  • South Africa: Moderate demand. Banks (Nedbank, ABSA) and FinTech have data science teams. Less specialized roles than analysts or engineers but better-paid.
  • Remote availability: 75% of roles are remote/hybrid.

Who Is This Path For?

Ideal starting backgrounds

BackgroundReadinessWhat you already have
Data Analyst✅ Strong startSQL and analytics; add statistical rigor and Python
ML Engineer✅ Strong startML fundamentals; add statistical thinking and domain context
Software Engineer🟡 Good with gapsEngineering practices; needs deep stats and domain knowledge
Statistician / Mathematician✅ Strong startStatistics expertise; add Python, SQL, product context
Recent MS/PhD (Stats, Math, Physics)✅ Strong startDeep theory; needs 6–12 months product/business context
Business Analyst🟡 PossibleDomain knowledge; needs statistical and technical ramp-up

You're ready to start this path if you can:

  • Write Python scripts (data manipulation, statistics, visualization)
  • Understand statistical concepts (hypothesis testing, distributions, correlation)
  • Build basic predictive models (regression, classification)
  • Query databases with SQL
  • Communicate findings to non-technical audiences

Not ready yet? Start with Data Analyst path (CP42) or Statistics Fundamentals first.


Certification Sequence

Visual path


Stage 1 — Foundation (Months 0–4)

Goal: Master statistics and Python. These are the non-negotiable fundamentals for data scientists.

CertCodeCost (USD)Study TimeWhy it matters
Statistics & Probability (Coursera/Khan Academy)$04–6 weeksFoundation for all statistical modeling; free excellent resources
Python for Data Science (NumPy, Pandas, Matplotlib)$0–$404–6 weeksProduction-grade Python for data manipulation

Stage 1 total: $40 USD · R720 ZAR · 4–6 months

Study approach: Use Khan Academy Statistics course (free, excellent, foundational) or StatQuest with Josh Starmer (free YouTube, clear explanations). For Python, use DataCamp Python track (subscription) or Fast.ai's Practical Deep Learning (free, code-first). Focus on: hypothesis testing, distributions, correlation, regression fundamentals.

Lab requirement: Build 5 statistical analysis projects. Each should include: hypothesis testing, exploratory data analysis, visualization, and interpretation. Use real datasets (Kaggle, UCI). Post to GitHub. 40+ hours hands-on.


Stage 2 — Core Specialisation (Months 4–18)

Goal: Get certifications proving data science and ML expertise.

CertCodeCost (USD)Study TimeWhy it matters
Google Cloud Professional ML Engineer$2008–10 weeksAdvanced ML; GCP's ML platform coverage
IBM Data Science Professional Certificate$24010–12 weeksEnd-to-end data science; covers the full lifecycle

Stage 2 total: $440 USD · R7,920 ZAR · 12–16 months

Study approach: For Google ML Engineer, use Google Cloud Training (free) and Udemy course ($20). For IBM, use Coursera Data Science Professional Certificate (accessible for $40/month or audit free). The IBM cert covers: data collection, cleaning, exploration, modeling, evaluation, and communication. Comprehensive pathway.

Project milestone: Build an end-to-end data science project solving a real business problem. Include: problem definition, exploratory analysis, statistical tests, predictive modeling, evaluation, and business recommendations. Deploy the model (or document deployment strategy). Present findings as if to a business stakeholder. Post to GitHub with a comprehensive README. This is your portfolio piece.


Stage 3 — Advanced Specialisation (Months 12–24)

Goal: Deepen in specialized areas (causal inference, advanced statistics, domain expertise).

CertCodeCost (USD)Study TimeWhy it matters
AWS Machine Learning SpecialtyMLS-C01$30010–12 weeksProduction ML systems; engineering depth
Advanced Statistics / Causal Inference (free)$06–8 weeksRigorous statistical thinking; separates good from excellent

Stage 3 total: $300 USD · R5,400 ZAR · 10–12 months

Study approach: For AWS MLS-C01, use Stephane Maarek's course ($20). For advanced stats, read Causal Inference: The Mixtape (free online book) and Judea Pearl's Causal Inference (book, not free). These teach modern causal thinking—essential for data scientists doing serious analysis.

Optional at hire time: Many data scientists land jobs after Stage 2 (GCP + IBM certs + portfolio) and deepen further on the job.


Timeline & Cost Summary

StageCertsDurationCost (USD)Cost (ZAR)
Stage 1 — FoundationStatistics, PythonMonths 0–4$40R720
Stage 2 — CoreGCP ML Engineer, IBM Data ScienceMonths 4–18$440R7,920
Stage 3 — AdvancedAWS MLS-C01, Advanced StatsMonths 12–24$300R5,400
Total to hireable24–30 months$780R14,040

Study hours required: ~400–500 hours. Assumes 12 hours/week = 30 months.


Salary Progression

All figures: median base salary, not including bonuses/equity. ZAR = USD × 18. Sources: Robert Half 2026, Glassdoor, LinkedIn Salary, Levels.fyi.

Experience LevelUSD/yearZAR/monthGBP/yearEUR/yearAUD/year
Entry / Junior (0–2 yrs)$90,000–$130,000R58,000–R83,000£70,000–€101,000€84,000–€121,000A$132,000–A$191,000
Mid-level (2–5 yrs)$130,000–$175,000R83,000–R112,000€101,000–€135,000€121,000–€164,000A$191,000–A$257,000
Senior (5–8 yrs)$175,000–$230,000R112,000–R147,000€135,000–€178,000€164,000–€216,000A$257,000–A$338,000
Lead / Manager (8+ yrs)$230,000–$300,000+R147,000–R192,000+€178,000–€232,000+€216,000–€288,000+A$338,000–A$441,000+

South Africa note: Data Scientists at Johannesburg banks (Nedbank, ABSA) earn R80,000–R120,000/month. Remote roles for international companies: R100,000–R160,000/month for entry, R140,000–R200,000/month for mid-level. Less common in SA than analyst roles; more specialized and research-focused.

Salary accelerators: Advanced statistics knowledge, causal inference expertise, domain-specific knowledge (FinTech, healthcare), and proven impact on business metrics all command 15–25% premiums.


First Job Strategy

Month 0–6: Build Your Statistical Foundation

  1. Master statisticsKhan Academy Statistics (free, foundational) or Stanford Stats courses. 8–10 weeks.
  2. Learn Python for data scienceDataCamp Python track or Fast.ai. Focus on Pandas, NumPy, Matplotlib, scikit-learn.
  3. Build statistical projects — Analyze datasets, run hypothesis tests, create visualizations. Post to GitHub.
  4. Join communities — r/datascience, Statistics Stack Exchange, local meetups.
  5. Read the classics — "Statistical Rethinking" by Richard McElreath, "Trustworthy Data" by Lisa Gebhardt.

Month 6–12: Build Your Data Science Portfolio

  • Project 1: Statistical Analysis & Hypothesis Testing — Take a dataset, formulate 5 hypotheses, conduct hypothesis tests, visualize findings. Write a report explaining methodology and conclusions. Estimated time: 12 hours.
  • Project 2: Predictive Modeling — Build a classification or regression model. Include: data exploration, feature engineering, model selection (compare 5+ algorithms), cross-validation, evaluation. Estimated time: 15 hours.
  • Project 3: Causal Analysis — Design an experiment or quasi-experiment to isolate causal effects. Implement causal inference techniques. Document methodology. Estimated time: 12 hours.

Month 12–24: Pursue Certifications

  • GCP Professional ML Engineer: Study 8–10 weeks.
  • IBM Data Science Professional Certificate: Study 10–12 weeks (can overlap with GCP).
  • Build visibility: Write blog posts on your analysis approaches. Share on Medium and LinkedIn.

Month 24–30: Apply & Iterate

  • CV positioning: List as "Data Scientist" once you have GCP cert + IBM cert + strong portfolio. Emphasize statistical rigor and business impact.
  • Target companies: Financial services (banks, insurance), FinTech, healthcare, e-commerce (Takealot), startups. Remote available.
  • Interview prep: Be ready to discuss 1) A complex analysis you conducted, 2) Hypothesis testing approaches, 3) Your modeling process, 4) Handling biased data, 5) Communicating uncertainty to non-technical audiences.
  • Salary negotiation: Data scientist roles pay well. Entry-level roles offer R80k–R120k/month locally; remote international roles R100k–R160k/month. Negotiate based on portfolio strength.

A Day in the Life

Data Scientist at Nedbank (Johannesburg) — Junior Level

08:00 — Standup with the analytics team. You're analyzing customer credit default risk. Previous model needs updating with recent data.

09:00 — Explore new data. Load the latest customer/loan records. Check distributions, missing values, outliers. Data looks clean; good quality this quarter.

10:00 — Run exploratory data analysis. Create visualizations: default rate by loan amount, region, customer age. Spot a trend: rural loans have higher default rate. Hypothesis: limited access to communication/payment options.

11:00 — Conduct hypothesis test. Is rural default rate significantly higher than urban? Run a chi-square test. p-value < 0.05. Statistically significant.

12:00 — Lunch.

13:00 — Start model retraining. Use scikit-learn. Try 5 algorithms: logistic regression, random forest, gradient boosting, SVM, neural network. Compare metrics. Gradient boosting performs best.

14:30 — Evaluate the model. Cross-validation, ROC curve, confusion matrix, feature importance. Model is good but slightly overfit to recent data. Add regularization.

15:30 — Document findings. Create a report: analysis approach, data summary, statistical tests, modeling methodology, results, recommendations. Include visualizations.

16:30 — Code review on your analysis notebook. Senior scientist checks your statistical approach. Feedback: add a check for Simpson's paradox, consider interaction effects. You implement.

17:00 — End of day. Schedule a meeting to present findings to risk team tomorrow.

Data Scientist at Capitec (Johannesburg) — Mid Level

08:00 — Standup. You're designing an experiment: testing a new product feature against the control group. Responsible for study design, statistical power, and analysis plan.

09:00 — Design the experiment. Sketch the approach: randomized controlled trial, sample size calculation (need 10k per group for 80% power), key metrics (signup rate, activation rate, churn).

10:00 — Write the statistical analysis plan. Document: hypotheses, metrics, success criteria, multiple comparisons correction, timeline. Get approval from the legal/ethics team.

11:00 — Data quality check. Review logs from the experiment. Randomization looks correct; no major data issues. Sample sizes on track.

12:00 — Lunch.

13:00 — Interim analysis. We're 70% through the experiment. Run preliminary tests. Treatment group shows 8% higher signup rate. Early signal, but not yet significant (confidence intervals overlap). Continue experiment.

14:30 — Work on a follow-up analysis. Previous experiment showed that feature increased retention. Now investigating: which customer segments benefit most? Segment the data (by age, region, income) and run analysis separately. Create visualizations.

15:30 — Pair with an engineer on model deployment. Your previous churn prediction model needs to be productionized. Review the model package, discuss inference latency and monitoring. Hand off to the ML engineer.

16:30 — Mentor a junior data scientist. Code review their analysis. Feedback: add more robustness checks, consider edge cases, clarify your conclusions. Good learning moment.

17:00 — End of day. Wrap up documentation. Plan: finish experiment analysis tomorrow.


South Africa Context

Market specifics

Data Scientists are respected roles in South African financial services and FinTech, but less common than data analysts or engineers. Nedbank, ABSA, and Standard Bank have data science teams. Capitec invests heavily in data science for credit risk. Most SA data scientists are industry-specific (banking, insurance, FinTech) rather than generalists.

Remote work is available but less norm than for analysts. Most data scientist roles are hybrid or onsite. International remote opportunities are good—SA data scientists work for global tech companies, especially with research focus.

The role requires deeper technical credentials than analyst roles (certs + strong portfolio matter). Most successful SA data scientists have a master's degree or strong self-taught background.

SA-specific resources

ResourceURLNote
Johannesburg Data Science Meetupmeetup.com/johannesburg-data-scienceMonthly meetups, networking
Nedbank Careers (Data Science)nedbank.co.za/careersRegular DS postings
Capitec Careerscapitec.co.za/careersActive data science team
IBM Data Science Certificatecoursera.orgIndustry-recognized cert
Khan Academy Statisticskhanacademy.org/math/statistics-probabilityFree foundational stats
StatQuest with Josh Starmeryoutube.com/@statquestFree, excellent explanations
LinkedIn Data Scientist Jobs (SA)linkedin.com/jobsJob board, 30+ postings

Frequently Asked Questions

Q: Do I need a master's degree to become a Data Scientist?

Not strictly, but many SA employers prefer it. A strong self-taught background with certs + portfolio can compensate. A master's in stats, math, or data science helps significantly.

Q: Is data scientist different from ML engineer?

Yes. Data Scientists focus on research, statistics, insights. ML Engineers focus on production systems, deployment, scalability. Data Scientists often hand off models to ML Engineers. Different specializations.

Q: How long does it take from zero?

24–36 months from complete zero. If you have statistical background: 12–18 months. If you're coming from analyst role: 18–24 months. This is not a fast path.

Q: Should I get a master's degree?

Optional but valuable. A master's in statistics, data science, or math accelerates the path by 12 months and improves job prospects. Self-taught + strong portfolio is viable but requires more effort.

Q: Is the IBM Data Science Certificate worth it?

Yes. It's comprehensive, covers the full data science lifecycle, and is widely recognized. At $40–$240 (depending on how you access), it's affordable.

Q: What's the difference between Data Scientist and Data Analyst?

Analysts answer questions about what happened and what's happening. Scientists predict what will happen and why. Scientists use more advanced statistical methods and do more exploratory research. Analysts focus on dashboards and reporting.


Sources & Further Reading

#SourceURLUsed for
1LinkedIn Jobs (Data Scientist)linkedin.com/jobsJob market data
2Khan Academy Statisticskhanacademy.orgFree statistics foundation
3IBM Data Science Certificatecoursera.orgProfessional cert
4Google Cloud ML Engineercloud.google.com/trainingML/cloud platform
5AWS ML Specialtyaws.amazon.com/certificationProduction ML systems
6Statistical Rethinkingxcelab.net/rm/statistical-rethinking/Bayesian statistics book
7Causal Inference: The Mixtapemixtape.scunning.comFree causal inference book
8Levels.fyi Data Scientistlevels.fyiSalary transparency

Template version: 2026-05-02 | Maintained by IT Career Roadmap | ZAR baseline: R18/$1 USD File naming: Career_Paths/CP47_Data_Data_Scientist.md

Research behind this path

The sourced deep dives this guide draws on — cert ladders, salary benchmarks, books and conferences, each cited.