Select Page

6.1 Observational vs. Experimental Data


Before discussing experiments, we will first take a moment to clarify the distinction between observational data and experimental data. 

When a researcher gathers observational data, he does so without impacting the subjects in any way.  As the name implies, observational data is obtained simply by watching.  Perhaps someone wishes to compare the efficiency of two rival coffee chains, Starbucks and Caffe Nero.  If he were to randomly choose 15 days from the month, and then sit at Starbucks on those days to monitor their order flow and queue efficiency between 11 a.m. and 1 p.m., and then to do the same thing at Caffe Nero for the other 15 days, he could come away with some general comparative statements about the two stores’ ability to process customers’ orders.

This would be entirely observational.  The researcher has not involved himself in the process at all, nor has he changed any condition or setting for one group in order to make the comparisons. 

Experimental data, on the other hand, is gathered after a researcher has altered some particular condition, with a plan to monitor the way that the change impacts subjects’ behavior. 

For instance, Lobster Land could place a sign at the exit of the Lobster Claw (the biggest rollercoaster in the park) that says “Delicious, ice-cold lemonade is available — just turn left and walk 100 feet.” 

The following week, Lobster Land could replace the sign mentioned in the paragraph above with a different type of message:  “Congratulations, Lobster Warrior!  You have conquered the biggest, craziest challenge in all of Lobster Land.  To reward thyself properly with the finest homemade lemonade known to mankind, follow the golden arrows painted on the ground to find your beverage.” 

The park could then monitor the actions of riders coming off of the Lobster Claw.  Does the second sign inspire more lemonade purchases?  Or do people just find it weird?  Experimentation can help to answer that, assuming that enough data is collected and that other variables are controlled for, to the best degree possible. 

When working with observational data, we should bear in mind that we cannot draw conclusions about cause and effect, based on its results. 

Suppose that a researcher decides to observe customer orders at Lobster Land’s Snack Shack for a two-week period.  For that entire time, he will look at the choices made by customers who purchase soda – specifically, he will analyze whether customers choose regular soda (which is high in sugar, and therefore high in calories) or diet soda, which is very low in sugar & calories.  He will also use a sophisticated Augmented Reality application that can estimate a person’s Body Mass Index (BMI), based on a single photograph, taken from a smartphone. 

Imagine if, after a month of careful observation, the researcher finds an unmistakable correlation – overweight or obese customers are more likely to opt for diet soda, rather than the regular, sugary stuff.  What can be concluded here?  Should we alert the news media and tell them that diet soft drinks are actually causing obesity, whereas regular soft drinks are not? 

NO! 

The Diet Coke/Coca-Cola data mentioned above is purely observational.  We do not have a control group here.  We do not have a separate treatment group.  Based on what we have seen, we actually have no way of knowing whether the soda choice impacts obesity, or if the causality is entirely reversed (what if people with larger BMIs are more likely to choose the diet soft drinks because they are trying to cut down on calorie intake?)