Showing posts with label research design. Show all posts
Showing posts with label research design. Show all posts

Monday, March 5, 2012

What's missing from the Fraser Institute school ranking report

The Fraser Institute ranking of Ontario elementary schools was released on Sunday, and as usual it was covered extensively by the press. Unfortunately, the press did not, as far as I could see, ask some serious questions that need to be asked.

I am not going to fault the Fraser Institute for not including all relevant technical information in the report; it is, after all, intended as a popular guide for parents. However, I could not find on the Institute’s website any link to a technical manual that would provide important information missing from their report.

Perhaps the most serious omission is any mention of test characteristics. The overall score calculated for each school is based on the annual assessment conducted by the Ontario Educational Quality and Accountability Office (EQAO). But are the tests used for these assessments valid measures of scholastic competence? Standard measures of reliability and validity are not reported (nor could I find them on the EQAO website, or in the technical manuals EQAO provides for the tests).

Of course, even a measure that is unreliable in assessing an individual student can be made reliable by aggregating the scores of a whole school. However, an invalid measure cannot be made valid by aggregation, and if a test is not a valid measure of scholastic competence its reliability does not matter. If someone gets your email address wrong, their messages are not going to get to you regardless of how many times they send them to exactly the same wrong address.

Another issue is that Much of the report deals with improvements in schools' scores but that little information is provided about the trend analysis on which reports of improvement were based. In particular, we need to know what statistical technique was used and an explanation of the high significance criterion (p < .10).

Other issues could be raised, but, even if I had included all of them, none of this post could be taken as necessarily implying that the Fraser Institute did not do an adequate job. I've asked more serious questions about studies I've reviewed and received reassuring answers. However, without the additional information described here, we cannot conclude that the ranks assigned by the Institute serve as a guide to school performance.

What's missing from the Fraser Institute school ranking report © 2012, John FitzGerald

More articles at the main site

Monday, January 9, 2012

Effect and cause as a clue to the meaning of science

The December 16, 2011, issue of WIRED has a piece by Jonah Lehrer called "Trials and Errors: Why Science is Failing Us" (click here to read it). Mr. Lehrer's argument seems to be that some phenomena are too complex for scientific method to be able to discover what causes them. In his conclusion he writes:
And yet, we must never forget that our causal beliefs are defined by their limitations. For too long, we’ve pretended that the old problem of causality can be cured by our shiny new knowledge. If only we devote more resources to research or dissect the system at a more fundamental level or search for ever more subtle correlations, we can discover how it all works. But a cause is not a fact, and it never will be; the things we can see will always be bracketed by what we cannot. And this is why, even when we know everything about everything, we’ll still be telling stories about why it happened. It’s mystery all the way down.
The comments following the piece do a good job of of pointing out the flaws in the reasoning by which Mr. Lehrer reaches this conclusion. However, one issue is omitted. That issue is that science is not about causes.

Science is about effects. At its simplest, an effect is a non-random relationship between two variables. Scientific experimentation investigates effects by varying one of the variables (the indendent variable) and seeing what happens to the other variable (the dependent variable). The goal is to explain the effect - that is, become more effective in predicting the dependent variable. This model can be expanded to handle large numbers of variables. For example, one of the things I do in evaluating satisfaction with a program is to investigate simultaneously the relative importance of several variables in accounting for satisfaction. What you typically find when you do this correctly is that only a few of the variables have any relationship to satisfaction. What you often find, too, is that the variables that account for their satisfaction are different from the reasons particpants report when asked why they like the program.

The methods I use are correlational, so they cannot attribute causation. What they tell you is that as one thing varies, so does another. Furthermore, the analyses of satisfaction I do are non-experimental, so I can't even be sure that the estimates of the correlations are all that exact. What I can do, though, is make a recommendation that changes be made to see if dealing with the the variables identified by the data analysis will improve satisfaction.

The same considerations apply to a lot of health research, and that consideration alone goes a long way to accounting for the examples Mr. Lehrer adduces. What health researchers do is develop their own recommendations for further research that will test whether their conclusions are correct. In fact, the supposed failure Mr. Lehrer describes is in fact a demonstration of the success of science - a hypothesis was developed from prior research to test whether a drug was effective, and the test failed to find evidence that it was effective. That failure by itself is informative - it tells us not to prescribe the drug.

One of the commenters at the link above (urgelt) goes into the issue of the adequacy of research in more detail. My post of January 5 (click here) provides another example of this type of difficulty. What is clear is that error is inherent in the process of scientific experimentation, and that the foundation of scientific method includes a recognition that error is inherent. Reports of statistical analysis of research results typically include many estimates of the error involved in the relationships estimated by the statistical techniques.

As for Mr. Lehrer's remarks about the mythical nature of causes, scientific method has long allowed explanatory variables that have no real existence (intelligence, for example, cannot be directly measured but only inferred from behaviour). Variables like this are called explanatory fictions. The reason they are allowed is that the point of science is to explain an effect, not to find out what its actual cause is. If a fictional variable can explain the effect where something tangible and real can't, so much the better. Furthermore, even a small improvement in accuracy of prediction will often produce large benefits. Obviously, something which improves accuracy only a small amount is unlikely to be a cause in any meaningful sense, but it can still play an important role in practice.

Complex systems often frustrate scientific research simply because there are so many potential effects to examine, not because scientists are naive about the nature of causes, which anyway they aren't looking for. Mr. Lehrer freely acknowledges that science has been spectacularly successful with some complex systems (the health of large populations, for example), so concluding that failures to be successful with others mean that science has failed to solve the problem of causation is not only questionable and hasty but irrelevant as well.

I am confident that the scientific research of 100 years from now will be superior to today's research. I am also confident that the reason for its superiority will not be that it has solved the problem of causation.

Website
Twitter

Research, cause, and effect © 2012, John FitzGerald

Why information overload is a myth

Everybody’s heard of information overload – a Google search I just did for information overload (in quotation marks) produced over 4 million results. In fact, though, it is data we are overloaded with, not information.

Information consists only of data that reduce uncertainty. A weather forecast is only informative if it predicts the weather accurately. If it doesn't predict the weather accurately, we could end up leaving our umbrellas at home on rainy days. Similarly, if we base corporate decisions on data that don’t predict the results we want to achieve, we could end up being embarrassed and out of pocket.

As the Schumpeter blog in the Economist said on December 31: “As communication grows ever easier, the important thing is detecting whispers of useful information in a howling hurricane of noise.” It’s that overload of noise we must fear.

How do you reduce an overload of noise?
  • By not collecting data that are irrelevant to the decisions you make.
  • By not collecting data that are nearly identical to informative data you already collect.
  • By not collecting more data than you need.
  • By not combining pieces of information in ways in ways which produce an uninformative total score (by weighting them, for example).

But how do you avoid doing these things? Chiefly by analysing your data with sound statistical methods. For example, you can estimate the relevance of data to a decision with methods like the correlation coefficient. You can use principal components analysis to find variables that are telling you the same story. You can use sampling theory to decide how much data you need to collect. You can use psychometric analysis to combine pieces of information into a single score effectively. The battle against uninformative data has not been won, but you can win that part of it that takes place in your office.

Website
Twitter

Why Information Overload is a Myth © 2012, John FitzGerald

Thursday, January 5, 2012

Cognitive decline research: Questions the CBC didn't ask

Today's CBC news report (click here) of a study of cognitive decline is pretty standard science reporting. I'm sure that other news sources provided much the same story. Anyway, it confines itself to reporting the results the researchers reported, results which were fairly stated.

However, there is other information the CBC might have provided, but didn't. First, it doesn't provide a link to the study (or hadn't when I posted a comment asking for one). I, for one, was interested in learning what "a 3.6% decline in mental reasoning" was. Does a decline of that size have a serious effect on people's functioning?

So I looked for the link and found it (here). It's an open access article that can be downloaded free in a PDF. The article doesn't provide a quick answer to the question of how serious the declines observed are, but the authors do suggest that further attention might be paid to people whose declines are greater than the mean in the study. That suggests to me the mean declines are not that serious, although I readily admit I may be reading something into the authors' suggestion that isn't there.

What I also found, though, is that the researchers did not control for health. Since older people tend to be less healthy, were these cognitive declines due to changes in brain function or to the fatigue resulting from poor health? Information about medical risk factors was collected, but the article does not report that it was incorporated in the statistical analyses.

None of this is intended to question the adequacy of the research. Being able to carry on a rigorous study of over 10,000 people for 24 years is proof enough of the researchers' competence. What this is intended to question is the value of a news report that simply reports results without examining them. I'm sure that if the researchers had been asked about the relationship of the health information they collected to cognitive decline they could have explained it fully. I'm sure if they'd been asking questions like that the journalists would have enjoyed their jobs more, too.