Designing Experiments: From Questions to Conclusions
Published:
An experiment must be able to answer the question
An experiment or study can produce results and still leave you unable to answer the question, “What does this result show?” That happens when trying something becomes the goal, while the research question and the result remain disconnected.
An experiment is not simply an activity that produces a result. Design it so that every plausible result gives an answer to the research question. Before starting, connect the question, comparison, evaluation, and interpretation.
This article is for people who have begun an experiment or study but are unsure what to compare or evaluate. It does not cover the statistical methods or evaluation measures that vary by field.
What you will learn
- How to turn a research question into something an experiment can examine
- How to consider what each possible result would let you conclude before getting the result
- How to choose a comparison, evaluation metric, and conditions based on what the question requires
Start with what each possible result would let you conclude
Before starting an experiment, write one sentence about what you want to examine. For example: “Can adding this change improve the aspect named in the question, compared with a baseline condition?”
Then consider what each plausible result would let you say. Include not only the outcome you hope for, but also no difference and a worse result.
| Possible result | What it lets you say | What it does not yet let you say |
|---|---|---|
| The condition with the change performs better than the comparison on the measure chosen in advance | An improvement was observed under this condition and with this measure | It will improve every condition, or why it improved |
| No difference is observed | This condition and measure did not show an improvement | The change has no value in every context |
| The condition with the change performs worse than the comparison on the measure chosen in advance | This condition did not show an improvement | Which part caused the result, or whether the same result holds in other conditions |
The point is not to decide on a preferred conclusion in advance. It is to check what answer each result gives to the question. If none of the rows answers the question, you can revise the design before running the experiment.
First, decide what the experiment as a whole needs to show. Do you need to show that a method combining two changes is useful, or do you also need to show how each change is related to the result? The comparisons you need depend on that choice.
A hypothesis is a provisional explanation that helps you begin the investigation. If it is not supported, it can still show you what comparison or condition to reconsider. Do not change how you interpret results simply to protect a hypothesis.
Choose comparisons, metrics, and conditions from the question
Each comparison, evaluation metric, and condition needs a reason to be there. If you cannot explain that reason from the question, it will be difficult to support a conclusion after seeing the results.
Here, a baseline is the condition without the change you want to examine: it is the basis for comparison. A comparison condition is the condition you compare against, such as the baseline. An evaluation metric is a quantitative criterion chosen in advance for comparing the difference between conditions. If numbers alone cannot answer the question, decide separately how observations and interpretation will be handled.
| Element | Connection to the question |
|---|---|
| Comparison | Shows what you need to compare against in order to answer the question. |
| Evaluation metric | Decides which number will test the point named in the question. |
| Conditions | Set the cases and settings for which you can state a conclusion. The meaning of a result can change when conditions change, even if the method is the same. |
Compare multiple changes separately as well
If you want to evaluate two changes separately, comparing only the condition that includes both is not enough. Even if the result is better, you cannot separate the contribution of either change from the effect of combining them.
| Condition | What the comparison examines |
|---|---|
| Baseline | The basis for comparison |
| Baseline + change A | The difference associated with change A |
| Baseline + change B | The difference associated with change B |
| Baseline + change A + change B | The result of combining both changes |
You do not need to limit the whole experiment to one change. However, for each pair of conditions, you should be able to explain what differs. When comparing the baseline with “baseline + change A,” the only difference is change A. Repeating comparisons like this lets you consider the results of each change and of their combination separately.
For example, if the question is whether a change makes a result easier to explain, comparing one number alone will not answer it. Decide in advance how you will examine explainability and what you need to compare it with.
After writing an experimental plan, try adding one sentence next to each element: why it is needed to answer this question. If you cannot write that sentence, the element may not be unnecessary; its purpose or role may simply still be unclear. Return to the question and check.
Summary: Design experiments to obtain an answer to the question
Before starting an experiment, state the question in one sentence and consider what each possible result would let you conclude. Then choose comparisons, evaluation metrics, and conditions because they are needed to answer that question.
An unexpected result does not need to be discarded as a failure. It can help you decide which question, comparison, or condition to reconsider next. Part 6 will cover how to record what you tried, what you learned, and what to do next.