Data Scientist
Location: Centurion, Gauteng (provisional – to be confirmed)
Positions Available: 5
Salary: Market-related
Employment Type: To be confirmed
Job Overview
We are seeking experienced and technically proficient Data Scientists to develop, implement and optimise advanced analytics, statistical models, machine learning solutions and predictive capabilities within complex enterprise IT environments.
The successful candidates will be responsible for analysing large and complex datasets, identifying meaningful patterns, developing predictive models and translating data-driven insights into practical business solutions.
This role requires strong hands-on expertise in Python, SQL, statistical analysis, machine learning, predictive modelling, data preparation, model evaluation and data visualisation.
The ideal candidates will have proven experience delivering end-to-end data science solutions, from business problem definition and exploratory data analysis through to model development, validation, deployment and performance monitoring.
Candidates should be comfortable working with structured and unstructured data, enterprise databases, cloud data platforms and modern machine learning frameworks.
The role requires close collaboration with data engineers, data architects, software developers, business analysts, enterprise architects and business stakeholders.
This is a specialist Data Scientist role requiring demonstrated practical experience developing statistical and machine learning models. General data analysis, Power BI reporting or dashboard development without substantial hands-on data science experience will not be sufficient.
Key Responsibilities
Data Science Strategy and Solution Development
- Identify business challenges that can be addressed through advanced analytics and machine learning.
- Translate business requirements into data science problems and technical solutions.
- Design and develop statistical, predictive and machine learning models.
- Evaluate suitable modelling techniques based on business objectives and data characteristics.
- Develop proof-of-concept solutions and validate technical feasibility.
- Define model success criteria, performance measures and expected outcomes.
- Collaborate with stakeholders to prioritise data science initiatives.
- Ensure solutions align with enterprise data and technology strategies.
- Document analytical approaches, assumptions and technical decisions.
- Contribute to the development of reusable data science capabilities.
Data Collection, Preparation and Exploration
- Extract and analyse data from enterprise databases, data warehouses, data lakes and other sources.
- Perform exploratory data analysis to identify trends, patterns and anomalies.
- Clean, transform and prepare structured and unstructured datasets.
- Identify missing values, inconsistencies, outliers and data quality issues.
- Develop appropriate data preprocessing and feature engineering approaches.
- Combine information from multiple sources for analytical purposes.
- Perform statistical profiling and assess dataset suitability.
- Collaborate with data engineers to improve data availability and reliability.
- Maintain reproducible data preparation processes.
- Ensure analytical datasets comply with applicable governance requirements.
Statistical Analysis and Predictive Modelling
- Apply statistical techniques to investigate relationships and identify significant patterns.
- Develop regression, classification, clustering and forecasting models.
- Perform hypothesis testing and statistical inference.
- Apply probability theory and statistical distributions.
- Develop predictive models using appropriate algorithms.
- Conduct time-series analysis and forecasting.
- Evaluate model assumptions and statistical significance.
- Apply dimensionality reduction and feature selection techniques.
- Interpret model outputs and communicate findings.
- Ensure analytical methods are appropriate for the business problem.
Machine Learning Development
- Design, train, test and optimise machine learning models.
- Implement supervised and unsupervised learning techniques.
- Develop classification, regression, clustering and anomaly detection solutions.
- Apply ensemble learning methods where appropriate.
- Perform feature engineering and hyperparameter optimisation.
- Use cross-validation and appropriate evaluation methodologies.
- Assess model accuracy, precision, recall, F1-score and other relevant metrics.
- Identify overfitting, underfitting and potential model bias.
- Compare algorithms and select appropriate production-ready approaches.
- Maintain model documentation and reproducible experiments.
Advanced Analytics and Artificial Intelligence
- Explore opportunities for advanced analytics and AI-driven automation.
- Develop anomaly detection, recommendation or optimisation models where required.
- Apply natural language processing techniques when relevant.
- Support text analytics, sentiment analysis and document classification initiatives.
- Evaluate deep learning approaches for suitable use cases.
- Investigate generative AI and large language model applications where appropriate.
- Assess technical feasibility, risks and business value of AI solutions.
- Ensure AI applications follow responsible and secure development practices.
- Collaborate with technical teams on AI integration requirements.
- Remain informed about emerging data science technologies.
Model Deployment and MLOps
- Collaborate with engineering teams to deploy models into operational environments.
- Package models for integration with applications, APIs or analytical platforms.
- Support automated training, testing and deployment pipelines.
- Apply version control and reproducible development practices.
- Monitor model accuracy, drift and operational performance.
- Identify model degradation and recommend retraining.
- Support model lifecycle management and governance.
- Maintain documentation for deployed models.
- Troubleshoot model-related production issues.
- Contribute to continuous improvement of MLOps processes.
Data Engineering and Integration Collaboration
- Work with data engineers to define data pipeline and feature requirements.
- Develop SQL queries to retrieve and analyse complex datasets.
- Support integration of data from multiple enterprise systems.
- Identify data quality issues affecting analytical outputs.
- Assist with analytical data modelling and feature store requirements.
- Work with cloud and database teams on data access and processing.
- Support scalable processing of large datasets.
- Optimise data extraction and analytical workflows.
- Ensure analytical processes are reliable and repeatable.
- Promote effective collaboration between data science and engineering teams.
Data Visualisation and Business Insights
- Develop clear visualisations to communicate analytical findings.
- Translate complex statistical results into practical business insights.
- Present model outcomes, trends and predictions to stakeholders.
- Develop analytical reports and supporting dashboards.
- Explain model limitations, assumptions and confidence levels.
- Recommend actions based on evidence and analytical results.
- Support business decision-making through quantitative analysis.
- Communicate technical concepts to non-technical audiences.
- Evaluate the measurable business impact of analytical solutions.
- Support adoption of data-driven decision-making.
Model Validation, Quality and Risk Management
- Establish appropriate model validation and testing approaches.
- Evaluate model performance against defined business objectives.
- Assess data leakage, model bias and other analytical risks.
- Perform sensitivity analysis and robustness testing.
- Validate training and testing methodologies.
- Maintain evidence of model testing and performance.
- Support model explainability and interpretability.
- Identify limitations affecting the reliability of predictions.
- Recommend corrective actions for underperforming models.
- Follow organisational model risk management requirements.
Data Governance, Security and Compliance
- Apply organisational data governance and information security policies.
- Protect confidential and sensitive information used in analytical models.
- Ensure appropriate access controls for analytical datasets.
- Support compliance with POPIA and other applicable requirements.
- Follow approved data retention and processing procedures.
- Maintain data lineage and analytical documentation.
- Apply responsible AI and ethical data science principles.
- Identify potential privacy, fairness and transparency risks.
- Collaborate with data governance and cybersecurity teams.
- Ensure analytical outputs are handled appropriately.
Technical Documentation and Knowledge Sharing
- Document data science methodologies, algorithms and model assumptions.
- Maintain technical specifications and experiment records.
- Produce model development and validation documentation.
- Use version control for analytical code and related artefacts.
- Share technical knowledge with data and engineering teams.
- Participate in peer reviews and technical design discussions.
- Contribute to reusable libraries, templates and coding standards.
- Recommend improvements to analytical development practices.
- Support knowledge transfer and operational handover.
- Stay informed about developments in machine learning and advanced analytics.
Minimum Requirements
- Relevant degree in Data Science, Computer Science, Statistics, Mathematics, Applied Mathematics, Engineering, Artificial Intelligence or a related quantitative discipline.
- Typically 4–6+ years of relevant professional experience in data science, advanced analytics or statistical modelling.
- Proven hands-on experience developing and implementing machine learning and predictive models – essential.
- Strong programming skills in Python.
- Advanced SQL skills and experience querying enterprise datasets.
- Practical experience with Python libraries such as pandas, NumPy and scikit-learn.
- Strong understanding of statistics, probability, hypothesis testing and predictive modelling.
- Experience with data preprocessing, exploratory analysis and feature engineering.
- Proven experience developing classification, regression, clustering or forecasting models.
- Knowledge of model evaluation, validation and performance optimisation.
- Experience working with structured and unstructured datasets.
- Understanding of relational databases, data warehouses and data lakes.
- Exposure to cloud data or machine learning platforms.
- Familiarity with Git and collaborative software development practices.
- Understanding of model deployment, monitoring and MLOps concepts.
- Strong analytical problem-solving and technical communication skills.
- Ability to explain complex analytical findings to business stakeholders.
- Experience delivering practical data science solutions in enterprise environments.
Technical Skills and Competencies
Programming and Data Analysis
- Python
- SQL
- pandas
- NumPy
- SciPy
- Jupyter Notebook
- R (advantageous)
- Data manipulation and transformation
- Exploratory data analysis
- Data preprocessing
- Feature engineering
- Statistical computing
- Code optimisation
Machine Learning Frameworks
Experience with relevant technologies such as:
- scikit-learn
- XGBoost
- LightGBM
- TensorFlow
- PyTorch
- Keras
- CatBoost
- Other established machine learning libraries
Statistical and Analytical Techniques
- Descriptive and inferential statistics
- Probability distributions
- Hypothesis testing
- Regression analysis
- Classification
- Clustering
- Time-series forecasting
- Anomaly detection
- Feature selection
- Dimensionality reduction
- Statistical modelling
- Experimental design
- A/B testing
- Model interpretability
Advanced Analytics and AI
Knowledge of relevant techniques such as:
- Natural language processing
- Text classification
- Sentiment analysis
- Recommendation systems
- Deep learning
- Neural networks
- Predictive analytics
- Optimisation algorithms
- Generative AI fundamentals
- Large language model integration
- Responsible AI principles
Data Platforms and Databases
- Microsoft SQL Server
- PostgreSQL
- Oracle Database
- MySQL
- Data warehouses
- Data lakes
- Lakehouse architectures
- Databricks
- Snowflake
- Apache Spark
- PySpark
- Enterprise data integration
Cloud and Machine Learning Platforms
Exposure to relevant platforms such as:
- Microsoft Azure
- Azure Machine Learning
- Azure Databricks
- Microsoft Fabric
- Amazon Web Services
- Amazon SageMaker
- Google Cloud Platform
- Vertex AI
- Cloud storage and processing services
- Managed machine learning environments
MLOps and Development Tools
- Git
- GitHub or GitLab
- MLflow
- Model versioning
- Experiment tracking
- Docker
- REST APIs
- CI/CD fundamentals
- Automated model testing
- Model deployment
- Model monitoring
- Data and concept drift detection
- Reproducible machine learning workflows
Data Visualisation and Reporting
- Matplotlib
- Seaborn
- Plotly
- Power BI
- Tableau
- Statistical visualisation
- Analytical storytelling
- Model performance reporting
- Business insight presentation
Relevant Certifications (Advantageous)
One or more of the following certifications would be beneficial:
- Microsoft Certified: Azure Data Scientist Associate
- Microsoft Certified: Azure AI Engineer Associate
- AWS Certified Machine Learning Engineer – Associate
- AWS Certified Machine Learning – Specialty
- Google Cloud Professional Machine Learning Engineer
- Databricks Certified Machine Learning Professional
- Databricks Certified Machine Learning Associate
- TensorFlow Developer Certificate or equivalent training
- Relevant Python, machine learning or advanced analytics certifications
- Recognised postgraduate qualifications in Data Science, Statistics or Artificial Intelligence
Key Personal Attributes
- Strong analytical reasoning and mathematical problem-solving abilities.
- Excellent attention to detail and commitment to analytical accuracy.
- Curiosity and ability to investigate complex datasets.
- Strong programming and technical development discipline.
- Ability to translate business challenges into analytical solutions.
- Excellent communication and presentation skills.
- Ability to explain complex models to non-technical stakeholders.
- Strong collaboration and stakeholder engagement abilities.
- Structured and methodical approach to experimentation.
- Ability to work independently and manage competing priorities.
- Commitment to responsible AI, data security and ethical analytical practices.
- Willingness to learn and adapt to emerging technologies.
- High levels of professionalism, accountability and integrity.
Application Requirements
Interested candidates should submit an updated CV clearly demonstrating their practical data science and machine learning experience, together with copies of relevant academic qualifications and professional certifications.
Candidates should specifically highlight:
- Years of professional data science experience.
- Machine learning and predictive modelling projects personally developed.
- Python, SQL and machine learning frameworks used.
- Types of models developed and deployed.
- Statistical analysis and feature engineering experience.
- Cloud data science platforms used.
- Model deployment, monitoring and MLOps experience.
- Business problems solved through advanced analytics.
- Measurable model performance or business outcomes.
- Experience working with large or complex enterprise datasets.
- Relevant academic qualifications and technical certifications.
Important: This is a specialist Data Scientist opportunity requiring demonstrable hands-on experience developing statistical and machine learning models. General data analysis, SQL reporting, Power BI dashboard development or business intelligence experience without substantial practical data science expertise will not meet the intended specialist profile.
Please note: This is a provisional recruitment specification prepared pending confirmation of the client's detailed requirements. The preferred technologies, minimum experience, qualifications, certifications, remuneration, employment arrangements and working conditions will be confirmed during the recruitment process.