Your supervisor has rejected two topics and registration closes soon. Forty data science dissertation topics for South African students — distinct from a general computer science project — each with a one-line research question, grouped by the area of application, plus the four tests every topic must pass before you propose it.
What makes a topic a data science topic rather than a computer science topic?
Data science dissertations centre on extracting insight or building a predictive model from a dataset — the contribution is in the data, the features, the model choice and the evaluation, not primarily in building new software infrastructure. A computer science software-engineering dissertation might build a new system; a data science dissertation applies or compares analytical methods on real or realistic data to answer a specific question. Some topics below could be framed either way — the framing you choose determines which department and which examiners’ expectations apply, so confirm with your supervisor which lens fits your specific programme.
The four tests a data science topic must pass

Before proposing any of the topics below: (1) Data test — can you name, right now, a specific dataset or data source, and have you at least attempted to access a sample of it? (2) Scope test — is the topic narrow enough to complete with the compute and time you actually have (a free-tier cloud GPU and a single academic year), not a research-lab-scale project? (3) Baseline test — is there an existing method or published result you can compare your approach against? (4) Ethics test — if the data touches people (health records, financial transactions, social media), have you identified the research ethics clearance and, where relevant, POPIA implications this specific data raises?
Natural language processing for South African languages (5 topics)

- Sentiment analysis of isiZulu or isiXhosa social media text — how accurately can existing multilingual models classify sentiment in under-resourced South African languages compared to English?
- Named entity recognition for South African news text — which named-entity recognition approach performs best on a corpus of local news mentioning South African places, people and organisations?
- Machine translation quality between Afrikaans and English for a specific domain (legal, medical, or educational text) — how does translation quality vary by domain and sentence complexity?
- Hate speech detection in code-switched South African social media text — how well do existing classifiers handle text that mixes English with a local language mid-sentence?
- Topic modelling of South African parliamentary Hansard records — what are the dominant themes discussed in a specific legislative session, and how do they shift over time?
Healthcare and clinical data science (5 topics)
- Predicting hospital readmission risk from de-identified admission data at a single facility — which features best predict 30-day readmission in this specific dataset?
- Time-series forecasting of a public clinic’s patient attendance patterns — can a simple forecasting model improve staffing predictions over the clinic’s current planning method?
- Classifying diabetic retinopathy severity from a public retinal image dataset — how does a chosen convolutional model’s accuracy compare to a published baseline on the same dataset?
- Analysing antimicrobial resistance trends from publicly available NHLS-adjacent surveillance summaries — what trend, if any, appears in the specific pathogen and period studied?
- Predicting no-show rates for outpatient appointments at a specific facility — which patient or scheduling features are most predictive, and could a simple intervention reduce the rate?
Financial and fraud analytics (5 topics)
- Detecting fraudulent transactions in a South African mobile money or synthetic financial dataset — which supervised classification approach best balances precision and recall on imbalanced fraud data?
- Credit risk scoring using JSE-listed company financial ratios — how well does a chosen model predict a defined risk outcome compared to a traditional scoring approach?
- Sentiment analysis of South African financial news and its relationship to JSE index movements — is there a measurable, lagged relationship in the specific period studied?
- Anomaly detection in a simulated retail transaction dataset for a South African context — which unsupervised method best flags genuinely anomalous transactions without excessive false positives?
- Predicting small-business loan default using publicly available or simulated South African SME financial data — which features are most predictive, and what does this suggest for lending criteria?
Agriculture and environmental data science (5 topics)
- Crop yield prediction using publicly available South African weather and soil data for a specific crop and region — how accurately can a machine learning model forecast yield compared to historical averages?
- Satellite-imagery-based land cover classification for a specific South African region — how does a chosen classification model’s accuracy compare against a published benchmark?
- Predicting water quality indicators from publicly available DWS monitoring data for a specific catchment — which variables best predict the outcome measure studied?
- Drought risk mapping using SAEON or SAWS climate data for a defined region — what pattern emerges over the period studied, and how does it compare to existing drought indices?
- Analysing veld fire occurrence patterns from public satellite fire-detection data for a specific province — what environmental or seasonal factors best predict fire frequency in the dataset?
Public sector and governance data (5 topics)
- Predicting municipal service-delivery complaint volumes from publicly available municipal data — which factors best predict complaint spikes in the specific municipality studied?
- Text-mining public participation submissions on a specific piece of legislation — what themes dominate, and how do they cluster by stakeholder type?
- Analysing Stats SA labour force survey microdata for a specific demographic trend — what pattern emerges over the period studied, and how does it compare to the official headline figures?
- Predicting matric pass-rate variation across schools using publicly available Department of Basic Education data — which school-level or district-level factors are most predictive in this dataset?
- Network analysis of public procurement award data for a specific sector or municipality — what patterns of supplier concentration or repeat awards appear in the dataset studied?
Computer vision and image analysis (5 topics)
- Automated pothole or road-defect detection from dashcam or public street-imagery datasets for a South African city — how accurately does a chosen model detect defects compared to manual annotation?
- Classifying informal settlement structures from satellite imagery for a defined urban area — how does the model’s classification compare to available ground-truth data?
- Wildlife species identification from camera-trap images in a South African conservation area — which model architecture performs best on this specific dataset, and what does misclassification reveal about the images themselves?
- Detecting illegal dumping sites from satellite or drone imagery for a defined municipal area — how well does the chosen approach generalise across different terrain types in the dataset?
- Automated number-plate recognition accuracy under South African-specific conditions (plate formats, lighting, weather) — how does an existing model’s accuracy vary by these conditions in a locally collected or simulated test set?
Education and learning analytics (5 topics)
- Predicting at-risk first-year university students from anonymised institutional data at a single university — which early-warning features best predict dropout or academic difficulty in the dataset studied?
- Analysing learning management system log data to identify engagement patterns linked to performance at a single institution — what pattern, if any, distinguishes higher- from lower-performing students?
- Topic modelling of open-ended student feedback at a single institution — what themes dominate, and do they differ meaningfully by faculty or year of study?
- Predicting NSC (matric) subject performance from publicly available school-level Department of Basic Education indicators — which factors are most predictive at the school level in the dataset studied?
- Recommender-system design for suggesting relevant open educational resources to South African students in a specific subject — how does the chosen approach’s relevance compare to a simple baseline (most-popular or keyword-matched) recommendation?
Cross-cutting and methods-focused topics (5 topics)
- Comparing explainability techniques (SHAP, LIME) on a model trained on a South African dataset from any domain above — which explanation method produces more consistent, human-interpretable results for this specific model and data?
- Bias and fairness auditing of a predictive model trained on South African demographic data — does the model’s error rate differ meaningfully across a protected demographic variable in the dataset studied?
- Comparing low-resource machine learning approaches (few-shot, transfer learning) for a South African dataset too small for training a model from scratch — how does transfer learning’s performance compare to training from scratch on the same small dataset?
- Evaluating synthetic data generation as a privacy-preserving alternative to real South African personal data for a specific modelling task — how does a model trained on synthetic data compare to one trained on the real data it approximates?
- Reproducing and critically evaluating a published international data science study using a South African dataset instead — does the original finding hold in the local context, and if not, why might it differ?
How does this differ from the computer science dissertation topics already on this site?
This site’s existing computer science content covers when a software project is accepted as a dissertation and where to find data sources — both about the CS software-engineering track specifically. Data science, as a distinct field on this site, is scoped to analytical and modelling work on data as the primary contribution, and these 40 topics have not been covered by any existing article on this site.
Any one of these forty needs real structuring work before the first line of code — Tesify helps lay out the proposal, the chapters and the methodology in the order examiners expect, while the dataset, the model and the results stay entirely your own.
Frequently asked questions
Do I need a real dataset before proposing one of these topics?
You need to have identified a specific, named, realistically accessible dataset before your proposal is approved — “I will find data later” is a common reason committees send a topic back for rescoping.
Can I combine two of these topics into one dissertation?
Generally no — combining two distinct modelling tasks usually exceeds what a single dissertation timeline can support well. Pick one, scope it precisely, and note related extensions as future work instead.
Which of these topics is easiest for an honours-level project?
Topics with a single, well-defined, publicly available dataset and a straightforward comparison against an existing baseline (several of the cross-cutting and NLP topics above) tend to be more tractable at honours level than topics requiring extensive primary data collection.
Do I need ethics clearance for a data science dissertation using public datasets?
Even with public data, most departments expect a data-source and ethics declaration, and any dataset touching identifiable people needs a proper ethics review regardless of how the data was obtained — check your department’s specific requirement before assuming public data is automatically exempt.
What software do I need for these topics?
Python with standard data science libraries (pandas, scikit-learn, and a deep learning framework for computer-vision or NLP topics) covers most of these — see the site’s guide to data sources for a computer science dissertation for where to find compute and dataset access options.
