Why Do We Experiment?

In his new book, Why Do We Experiment? Dr. Henry Cloud, professor emeritus at Duke University and an advisor to the National Science Foundation, offers an interesting perspective on why do we experiment in science. In theory, experiments should appear something like this: discovery, ideally, in a timely way, is about exploring, maximizing and minimizing harm.

Unfortunately, we rarely see anything close to this ideal. For example, some of the biggest discoveries in recent decades have come about because scientists were experimenting with new ways to test a previously accepted model of how things work. If this sounds familiar, you’re right. It is the failure to test hypotheses properly that has resulted in numerous unproven or misguided theories, some of which have been used to support extravagant claims by those in positions of power. For example, String Theory, promoted by top scientists, postulates that the entire universe is made up of energy waves, similar to radio signals being emitted by mobile phones or from microwave ovens.

String’s contention received a lot of media attention, both positive and negative, but no hard evidence was forthcoming in support of it. In fact, after conducting his own investigation, String came to the conclusion that the evidence did not support his hypothesis. So, just what went wrong? The answer lies in what is known as an MVP or Meta-VT. For those unfamiliar with it, here is the definition:

A Meta-VT is a proposed structure that would eliminate any possibility of one side exceeding the other during experiments. Proponents of the concept say that with such structures, one side cannot exceed the other in speed of expansion. Opponents argue that such tests could not be conducted because microwaves cannot be slowed down to the point where their energy would escape the medium. Another problem cited is that with such a testing structure, why would scientists spend the time and effort to perform experiments, when they can just rely on such a system to uplift their theory instead? To understand how this argument can apply to your own work, consider the following scenario.

Why Do We Experiment?

If you have a website earning revenue on an ongoing basis, one day you realize that you are not getting much traffic from the type of visitors you got a few months ago. You determine that perhaps your conversion rate has dropped, too, but you still believe your revenue gain theory. To test this hypothesis, you decide to perform an experiment wherein you change the headline of one of your web pages to a different one, and see what happens. Once you make this change, immediately you will see an increase in conversion rate. So theoretically, if the same headline is changed, then equally so will your revenue gain or loss.

The problem is that the revenue gained or loss can also be due to an innocent change of name. In this case, the revenue change was not really part of the experiment but merely an innocent move to lure more visitors. Because the assumption is flawed, the researcher must perform another set of measurements – this time using metrics to determine whether this increase in traffic was indeed caused by the name change alone, or if other factors contributed to the increasing success.

If you think this describes your current research method, then you need to ask yourself if the results you are consistently obtaining are those that are truly the result of your experiments. For example, if you are using metrics in order to find out whether a certain treatment will uplift sales, are you measuring whether sales increased due to treatment A or due to treatment B? If you think experiment A is doing well, then experiment B could be performing even worse. You must address the issues underlying the results in order for you to truly understand why do we experiment.

Measurement, experimentation and measurement are three essential elements of statistical risk management. There is much to be said for implementing effective risk management strategies in the context of your experimentation, but the bottom line is that any effective risk management strategy has to take into consideration not only the potential harm that may be done through the implementation of the experiment, but also the potential revenue that might be gained as well. It is simply not enough to simply implement the best metrics in order to make better decisions. The metrics being used to make these better decisions must be based on the data which they are intended to reflect, and they must also be able to handle contingency. Only then will you truly be able to understand why do we experiment, and the steps that you can take to avoid making the same mistakes in the future.

Jennifer Radtke