South Africa has an unusually generous public data infrastructure for a dissertation: Stats SA publishes free statistical releases, its SuperWEB2 portal lets you build your own cross-tabulations, and DataFirst at UCT distributes the underlying survey microdata to registered users at no cost. Most students never use any of it, because nobody shows them the access procedure. This guide does, step by step.
How to get South African data for your dissertation, step by step
Step 1: Decide which of the three data layers you actually need
Before registering anywhere, match your research question to the right layer. The expected output of this step is one sentence in your methods chapter saying which layer you use and why.
- Published statistical releases (PDF and Excel): ready-made national and provincial figures with the fieldwork and weighting already done. Right choice if you need context numbers or a descriptive baseline. Example: the General Household Survey 2025 statistical release (P0318), published on 26 May 2026, gives you household internet access by province in a single table.
- Interactive tables (SuperWEB2): you choose the variables and build the cross-tabulation yourself. Right choice if the published tables almost answer your question but not quite — say, you need the figure by province and age group.
- Microdata (DataFirst): the anonymised unit-record data, one row per household or person. Right choice only if you will run your own statistical analysis. This is the layer examiners mean when they ask whether your quantitative study uses “secondary data analysis”.
Step 2: Download statistical releases directly from Stats SA
Stats SA publication pages list each release with its series number — the General Household Survey is P0318 — and link the full statistical release PDF plus time-series spreadsheets. No registration is needed. Record three things for every file you download: the series number, the release date and the reference year. A Harvard reference for the example above looks like this:
Statistics South Africa. 2026. General Household Survey 2025. Statistical release P0318. Pretoria: Statistics South Africa.
Expected output: the exact release PDFs your chapters will cite, filed with their series numbers, so your reference list never cites “Stats SA website” as a source.
Step 3: Register on SuperWEB2 and build your first table
SuperWEB2 is Stats SA’s browser-based table builder. The access procedure, from the platform’s own documentation:
- Go to the SuperWEB2 login page. If you do not have a username and password, you can register for an account from that page. Guest access exists but is limited — guest users cannot save tables and may only see a subset of the datasets, so register.
- After login, datasets appear in the left panel in a tiered folder structure. Double-click a dataset to start a new table.
- In Table View, select the fields you want and add them to rows, columns, wafers or filters. Expand a field folder, tick the values you need, and assign them with the Row and Column buttons.
- Click Retrieve Data to run the cross-tabulation, then use the download options to export the table.
Expected output: a saved, exportable table you built yourself — plus the exact variable definitions, which belong in your methods chapter.

Step 4: Register with DataFirst and download microdata
DataFirst is a research data service at the University of Cape Town committed to open access to African microdata. Its data portal hosts, among much else, Stats SA household surveys — the General Household Survey microdata is catalogued there with Statistics South Africa listed as producer. The procedure:
- Open the DataFirst data portal and search the catalogue for your survey and year.
- On the study page, choose Get Microdata. You will be asked to log in or register for a free account.
- Accept the terms of use presented for the dataset, then download the data files and — critically — the supporting documentation: questionnaire, metadata and technical notes.
- Store the citation the portal gives you. Microdata is cited like any other source, with the depositor and version.
Expected output: data files plus documentation. If you skip the questionnaire and metadata now, you will pay for it in your analysis chapter later.
Step 5: Respect the weights, or your numbers will be wrong
Survey microdata comes with sampling weights: each row represents many South African households, not one. If you run unweighted frequencies, your totals will describe the sample, not the country, and they will not reconcile with the published Stats SA tables — the most common reason a student’s “national” figure does not match the official release. State in your methods chapter which weight variable you applied. Expected output: numbers that reconcile with Step 2’s published tables, which is also the cheapest self-check available.
Step 6: Clear the ethics and licensing questions early
Secondary analysis of properly anonymised public microdata is usually the lowest-risk category in a faculty ethics application — but it still goes through the process, and your application should name the dataset, its access conditions and its anonymisation. Our guide to ethics clearance at South African universities covers where secondary data fits. Note also the terms you accepted at download: they govern redistribution, so put the citation in your dissertation, not the data files in your appendix.
Step 7: Write the data section while the trail is fresh
Your examiner needs to be able to retrace every figure: source, series number, year, variables, weights, exclusions. Write that paragraph the day you download, not six months later. If you are still structuring the study itself, work through our step-by-step guide to the research proposal first — the data-access plan described here slots directly into its methodology section, and Tesify can hold the whole chapter structure, sources and citations in one place while you draft around a full workload.
Which dataset should you start with, by field?
For most social-science, education, business and public-health dissertations, the General Household Survey is the workhorse: annual, national, and covering education, health, housing, services and connectivity. An education student can pull school attendance and post-school participation by province; a public-health student gets healthcare access and household environment variables; a business or economics student gets household services, connectivity and income-adjacent indicators; a development-studies student gets nearly everything at once. If your question is about postgraduate education itself, pair it with the sector statistics we unpack in our postgraduate enrolment statistics guide — those come from HEMIS, a different system with different definitions, and the two must not be conflated in your literature review.
The discipline test is simple: if you find yourself about to write “no data exists on X in South Africa”, spend thirty minutes in the DataFirst catalogue and the Stats SA release list first. That sentence is one of the most commonly false claims in South African dissertations, and an examiner who knows the survey landscape will catch it. The honest versions — “no data exists at the level of disaggregation this study requires” or “the most recent available data predates the policy change under study” — are defensible and, properly argued, become part of your contribution.
The four mistakes that cost students marks
Citing the portal instead of the dataset. “Stats SA website” or “DataFirst” is not a source; the specific release or dataset, with its series number, year and version, is. Your examiner may check, and the difference signals whether you actually handled the data.
Mixing reference years mid-argument. If your tables blend a 2023 release and a 2025 release without saying so, any change you “find” between them may be a change in methodology, not in the world. State the reference year of every figure in the text, not only in the table caption.
Treating the sample as the population. The weights problem from Step 5, in its most public form: a student reports “12 000 households lack piped water” when that is the unweighted sample count. Reviewers outside your committee will quote your number; make sure it is the weighted one.
Downloading data before the question is fixed. The catalogue is seductive, and a week disappears into browsing variables. Write the research question first, then fetch the narrowest data that answers it. Data-driven question drift is a real failure mode the proposal stage exists to prevent.
FAQ
Is Stats SA data free for students?
Statistical releases download free with no registration. SuperWEB2 requires free registration. DataFirst microdata requires a free account and acceptance of dataset terms.
Do I need permission from Stats SA to use the data in a dissertation?
Published releases and public-use microdata are distributed for exactly this purpose under their stated terms of use. You accept the conditions at download; you do not write to Stats SA for individual permission.
What is the difference between SuperWEB2 and DataFirst?
SuperWEB2 builds aggregate tables inside your browser from Stats SA datasets. DataFirst distributes the unit-record microdata files for you to analyse in your own software.
Can I use this data in a qualitative study?
Yes — as context. A qualitative dissertation still needs a defensible description of the setting, and a properly cited national figure beats an undated newspaper claim every time.
What software do I need for microdata?
Anything that reads the distributed formats — R and Python are free; SPSS and Stata are licensed at many universities. Check what your institution provides before paying for anything.
Why don’t my totals match the published tables?
Almost always weights: you ran unweighted counts. Apply the survey weight variable and re-run. If a small gap remains, check whether you replicated the release’s population scope and exclusions.
How do I cite a dataset in Harvard style?
Producer. Year. Title of dataset, version. Distributor. For DataFirst-hosted surveys, use the citation block the portal supplies on the study page and adapt it to your faculty’s Harvard variant.
Does using public data exempt me from ethics review?
No. It usually simplifies review, but the application still goes through your faculty committee — name the dataset and its anonymisation, and let the committee make the call.
