Which Statistical Test Should I Use for My Dissertation? (South Africa, 2026)

Direct answer: Your statistical test is determined by three things, in this order: what you are trying to do (compare groups, measure a relationship, or predict an outcome), how many groups or variables are involved, and what kind of data you collected. Answer those three and the test chooses itself. Software, assumptions and output all come afterwards.

Why does this question feel harder than it is?

Because most students meet the tests as a list to memorise rather than as answers to a question they have already asked. By the time you are staring at your dataset at ten o’clock at night, you have forgotten that your research question already contains the answer. A question of the form “do X and Y differ?” is a comparison. A question of the form “are X and Y related?” is an association. A question of the form “does X predict Y?” is a regression. Your proposal committed you to one of those shapes months ago — our guide to writing a research proposal for a South African university is where that commitment was made — and the test is simply the arithmetic that matches it.

The second reason it feels hard is that the decision is genuinely two-layered. Layer one is your design and your data type, and it narrows you to a small family of tests. Layer two is whether your data satisfies that family’s assumptions, which decides whether you run the standard version or its distribution-free alternative. Students who try to do both layers at once get lost. Do them in order.

What three questions actually determine your test?

Write these three answers at the top of a page before you open any software.

  1. What is the job? Comparing groups, measuring an association, or predicting an outcome.
  2. How many groups, and are they independent or related? Independent means different people in each group. Related means the same people measured twice, or matched pairs.
  3. What kind of variable is your outcome? Categorical (a label — passed or failed, urban or rural), ordinal (an ordered rank), or continuous (a measured quantity — a score, a mark, a rand amount).

Those three answers land you on one cell of the table below, and that cell holds one or two tests. This is the entire decision.

Hand-drawn decision tree branching into options on a notepad

Which test compares groups?

This is the commonest job in South African master’s work: do learners in intervention schools score differently from learners in control schools; do nurses in public and private facilities report different burnout; do men and women differ on a scale you administered.

Design Continuous outcome (assumptions met) Distribution-free alternative
Two independent groups Independent-samples t-test Mann-Whitney U test
Two related measurements (same people, before and after) Paired-samples t-test Wilcoxon signed-rank test
Three or more independent groups One-way ANOVA (+ post-hoc) Kruskal-Wallis H test
Three or more related measurements Repeated-measures ANOVA Friedman test
Two categorical variables (counts in categories) Chi-square test of independence

Two rules save most of the marks lost in this row of the table. First, if you have three or more groups, run one ANOVA — do not run three separate t-tests. Every additional test inflates your chance of a false positive, and an examiner who sees a page of pairwise t-tests where an ANOVA belonged will say so. Second, a significant ANOVA tells you that the groups are not all the same; it does not tell you which pair differs. That is what the post-hoc test is for, and it must be reported.

Which test measures a relationship between two variables?

Here the question is not “do these groups differ” but “do these two things move together”. Two continuous variables measured on the same people, with a roughly linear relationship, call for Pearson’s correlation coefficient. If either variable is ordinal, or the relationship is monotonic rather than linear, or the assumptions for Pearson do not hold, Spearman’s rank-order correlation is the standard alternative. For two categorical variables you are back to the chi-square test of independence, which asks whether the pattern of counts across categories departs from what independence would produce.

One caution that belongs in your discussion chapter, not your results chapter: a correlation is a description of co-movement, not a demonstration of cause. South African social-science dissertations lose marks every year for a sentence that quietly upgrades “was significantly associated with” into “led to”. Keep the causal language for designs that can carry it.

Which test predicts an outcome?

When your question names one outcome and several things that might explain it, you are in regression. A continuous outcome with one or more predictors is multiple linear regression. A binary outcome — completed or did not complete, adopted or did not adopt — is binary logistic regression, and it is the test most often missed: students force a binary outcome into a linear model because linear regression is the one they were taught. If your dependent variable has two categories, logistic is the correct family.

Regression is also where your literature review earns its keep, because you must justify which predictors go into the model and why. A model assembled from whatever variables happened to be in your questionnaire is a fishing expedition; a model assembled from the relationships your literature review established is an argument. Examiners can tell the difference in about a paragraph.

What if you have more than one dependent variable?

If you genuinely have several related outcomes and want to test group differences across all of them at once, MANOVA exists for that purpose. Be honest about whether you need it. In most master’s dissertations the cleaner and more defensible route is to state a small number of pre-specified outcomes and analyse each, acknowledging the multiple-comparison issue explicitly, rather than to reach for a multivariate technique you will struggle to explain under examination. The test you can defend beats the test that looks sophisticated — and in South Africa, where your examiners assess the document without you in the room, “I could have explained it” is worth nothing.

Spreadsheet of survey responses next to a printed questionnaire

How should you treat Likert data?

This is the single most argued-about question in South African questionnaire research, and the defensible position is narrower than either extreme. A single Likert item — one statement rated from strongly disagree to strongly agree — is ordinal. The spacing between “agree” and “strongly agree” is not known to equal the spacing between “neutral” and “agree”, so a mean of single-item responses is on shaky ground and the rank-based tests are the safer choice. A summated scale — several items measuring one construct, added or averaged into a single score — is conventionally analysed as continuous, and that convention is widely accepted in published work, provided you report the scale’s internal consistency.

What matters for your marks is not which side you take but that you state the decision and its reason in your methodology chapter. “Composite scores across the six items were treated as continuous, consistent with common practice for summated scales; single-item comparisons used non-parametric tests” is a sentence an examiner accepts. Silence is what gets queried.

What if your data does not meet the assumptions?

Then you move along the row, not off the table. Every parametric test in the comparison table has a distribution-free partner sitting beside it, and choosing it is a normal analytical decision rather than an admission of failure. What you must not do is run the parametric test anyway and hope nobody checks, or quietly delete the cases that were causing the problem. Assumption checking is itself a reportable step: examiners want to see that you looked, what you found, and what you did about it. That whole decision — the test of normality, what its result actually means at your sample size, and which alternative to move to — is a chapter of its own, and it is the one that stops most dissertations for a week.

Does your sample size change which test you can run?

Yes, in two directions. A small sample makes the distribution-free alternatives more attractive, because the parametric tests lean harder on distributional assumptions when there is little data to smooth them out. A small sample also limits how many predictors a regression can honestly carry — a model with more parameters than your data can support will produce confident-looking coefficients that mean nothing. In the other direction, a very large sample will return statistically significant results for differences too small to matter in practice, which is why effect sizes belong in your results alongside every p value. Sizing the sample properly is a decision that should have been made before data collection, and it is a separate question with its own procedure.

What will your faculty actually expect you to justify?

Four things, in the methodology chapter and again briefly in results: the test you chose, why the design and data type require it, what you did to check its assumptions, and what you did when a check failed. Departments differ on statistical style — a commerce faculty and a nursing faculty will not want the same level of detail — so ask your supervisor for two recently passed dissertations in your department and read their analysis sections. That single hour tells you more about local expectations than any textbook, and it is the same technique that works for length, structure and referencing conventions.

If your faculty offers statistical consultation, use it, and go with your three answers already written down. A consultant who is handed “here is my research question, here is my design, here is my outcome variable, I think this is a Kruskal-Wallis — am I right?” can help you in twenty minutes. A consultant handed a spreadsheet and the question “what should I do with this?” cannot.

Can you use AI to choose your test?

As a tutor, within your university’s rules — yes, and it is one of the better uses for it: explaining what a test does, what its assumptions mean, and why one alternative differs from another, at whatever hour you are actually working. What it must not do is decide your analysis for you without your understanding, or produce a methodology paragraph you cannot defend. The rules on disclosure differ by institution, and our guide to AI policies at South African universities sets out what each of the major ones requires. Tesify is built for the part around the analysis — holding your chapter structure, turning your decisions into defensible methodology prose, and keeping your citations straight — while the analytical judgement stays yours, which is exactly the division every South African policy asks for.

The practical test of understanding: can you say, in one sentence and without notes, why your test and not the one next to it? If you can, you are ready to write the chapter. If you cannot, go back to the three questions at the top of this page — the gap is almost always in question one.

FAQ: choosing a statistical test

What is the most common statistical test in a master’s dissertation?

For questionnaire-based studies comparing two groups, the independent-samples t-test and its distribution-free partner the Mann-Whitney U test are the workhorses, alongside chi-square for categorical data and correlation for relationships between measured variables.

What is the difference between a parametric and a non-parametric test?

Parametric tests assume your data follows a particular distribution and work with means; non-parametric tests make fewer distributional assumptions and generally work with ranks. Each parametric test in this guide has a standard non-parametric counterpart for when the assumptions do not hold.

Can I use a t-test for three groups?

No. Use a one-way ANOVA followed by a post-hoc test. Running multiple t-tests across three or more groups inflates the chance of a false positive, and examiners look for exactly this error.

When do I use chi-square instead of a t-test?

When both variables are categorical and you are analysing counts in categories rather than comparing averages of a measured quantity. “Did pass rates differ by province?” is chi-square; “did marks differ by province?” is a group-comparison test.

Is a Likert scale ordinal or continuous?

A single item is ordinal. A summated scale built from several items measuring one construct is conventionally analysed as continuous. State which you did and why in your methodology chapter; that justification is what is actually being marked.

Do I need to report effect sizes as well as p values?

Yes. A p value tells you whether an effect is distinguishable from chance; an effect size tells you whether it is big enough to matter. Reporting only significance is one of the most frequently flagged omissions in dissertation results chapters.

What test do I use for a yes/no outcome?

Binary logistic regression when you have predictors and want to model the outcome; chi-square when you simply want to test association with another categorical variable. Do not force a binary outcome into linear regression.

My study is qualitative. Do I need any of this?

No. Statistical tests apply to numerical data. Qualitative designs use coding and thematic procedures instead, with their own standards of rigour, sampling logic and reporting conventions.

Can I change my analysis plan after collecting data?

You can adapt it — for example, moving to a non-parametric test when an assumption fails — provided you disclose the change and its reason. What you cannot do is try many tests and report only the one that produced a significant result.

Where can I get help with the statistics at my university?

Most South African universities have a statistical consultation service or a departmental statistician, often free to registered postgraduates. Ask your supervisor or the postgraduate office early — these services book out near submission deadlines, and arriving with your design and outcome variable already specified makes the session far more useful.