Tuesday, June 14, 2011

Uninformation (2)

2. The logical or reasonable is not necessarily informative

People often believe that if they can construct a chain of reasoning which supports their beliefs that therefore they have demonstrated that their belief is true and informative. For example, many people reasoned out arguments which they seriously believed demonstrated that on January 1, 2000 the world would be thrown into chaos. I say their beliefs were serious because they acted on them – they stockpiled food, for example, they bought portable electric generators, and some even created fortified shelters to protect themselves from people who hadn't stockpiled food or bought generators.

As we found out, they were wrong. However, I can’t say that their conclusion was any less sound than the conclusion I and most other people drew that any disruption that might occur on January 1, 2000 would be minor. The people who drew this conclusion were sane and their reasoning from their data was sound. It was probably as sound or sounder than my own. In the end, one reason I and most other people were right and they were wrong is that we were using better data – data which were more informative. Another reason is that we were just luckier. In fact, no one fully understood all the factors one would have to assess to produce an accurate forecast of what would happen to the power grid on January 1, 2000. Furthermore, we probably weren’t aware of all the factors that would have to be considered.

If sound reasoning is based on invalid and inadequate data, it will reach invalid and inadequate conclusions. None of us is perfect – not even, as unlikely as it may seem, you or I – and we all at one time or another base logical conclusions on unsound data. And sometimes our reasoning just slips a gear, too. Even if our reasoning is perfect, none of us is omniscient, either. We can easily overlook important considerations.

That’s why the betting industry exists. If you’ve ever heard some of the explanations – often vehement ones – which horseplayers come up with to explain why the sure thing they bet in the last race ran as if he was pulling a milk wagon, you’ll know that relying too much on reason can not only cost you money but also lead you into an unjustified skepticism about the honesty and competence of one’s fellow human beings.

Obviously logic is involved in the development of information, just as facts are involved. It is not by itself informative, though. Two plus two equals four, but if the right answer is five you’re still wrong. That is why conclusions drawn from data need to be tested before they can be accepted as sound. If you think the 5-horse in the next race is going to romp, you won’t know that you’re right till the race has been run. And no matter what the weather report says, you won’t know whether it’s going to rain tomorrow or not until tomorrow arrives.

First article in the Uninformation series

Next: Information is not the statement of an authority

Actual Analysis
Uninformation (2) © 2011, John FitzGerald

Sunday, June 12, 2011

Uninformation (1)

Information consists of data which establish whether or not an assertion is false. Not all data do this. One of the reasons we have difficulty becoming and staying informed is that we sometimes accept as informative things which really aren’t, or least aren’t necessarily. This is the first in a series of posts in which we’ll look at a few things which are not information.

1. Information is not synonymous with facts

People often confuse information with facts. Someone who knows a lot of facts is considered to be well informed. A fact is only informative, though, if it helps you settle a question you need to know the answer to. If someone is on trial for armed robbery, the Crown does not submit evidence that the defendant is a skilled bridge player, true as that evidence may be.

Here's a fact: Churchill, Manitoba, is named for John Churchill, first governor of the Hudson's Bay Company. That=s a fact. Despite being a fact, though, it doesn't help me answer the question “Where do I find the men’s shirts?” whenever I drop in to one of the Bay’s branches. So for me that datum is not informative, factual though it be.

Furthermore, there are plenty of items of information that are not factual. The idea of intelligence, for example, cannot be said to be a fact, since there is widespread disagreement about just what intelligence is. However, the concept of intelligence is informative because in speculating about it we discover useful things. We have even discovered some of the shortcomings of the idea of intelligence.

Information is always derived from facts, and it always helps to predict facts. However, it need not be factual itself, and something which is factual need not be informative. As someone who has spent his life filling his memory with facts whose relevance to my life is highly questionable (see note about John Churchill above), I realize that collecting trivia can be enjoyable. Until they tell you something useful, though, trivia are just trivial.

Next: The logical or reasonable is not necessarily informative

Actual Analysis
Uninformation (1) © 2011, John FitzGerald

Tuesday, May 3, 2011

The value of political polls

I've been questioning the value of polls here, so here's some evidence of how well they work. I was interested in Ekos Politics' claim that methods of predicting the seats won by each party in Canadian federal elections "work pretty well" (the quotation is from a PDF I can no longer find on their website, but I have a copy if you want one). Here are Ekos' final projections for the election of May 2, 2011 (you can verify them here):

Conservatives: 130 to 146 seats
New Democrats: 103 to 123
Liberals: 36 to 46,
Bloc Québécois: 10 to 20
Green: 1.

And the results:

Conservatives 167 seats
New Democrats 102
Liberals 34,
Bloc Québécois: 4
Green 1.

In other words, Ekos got the Green seats right and no other party's. Of course, there is no reason they should get them right. The regional variation in voting is so great in Canada (the BQ only runs in Quebec, for example) that you'd need extensive polling in each riding to even hope to approximate the results. Even then the non-representative samples you'd be working with would seriously limit the accuracy of your estimates.

At any rate, the Ekos projections missed the two important events of May 2: the Conservative majority and the collapse of the Bloc. Journalists will probably go on acting as if polls mean something, but that doesn't mean you have to.

Tuesday, April 26, 2011

Inside the information cult (1)

In Canada we're in the final week of a federal election campaign. So far press coverage has been dominated by coverage of poll results. Unlike, it seems, most people I am skeptical of the utility of election poll results, and here I'll explain why.

Polls of people’s opinions can be a useful exercise if they’re done properly and the results are interpreted carefully. The polls that the press publishes, however, often fail to satisfy these criteria (at least in the form in which they are presented in the press). For one thing, they usually ask at most a handful of questions and don’t attempt to assess how meaningful the responses are . The results of this type of poll are not information but pseudo-information. Electoral poll results are obviously uninformative simply because they don’t predict election results accurately enough. Informative data must be valid, and in general they are not valid.

If you consult the excellent PollingReport website you will find a summary of final estimates by eighteen polls of the popular vote in the 2000 United States presidential elections. Fourteen of those predictions had George W. Bush winning the popular vote, which in fact was won by Al Gore. Sure, the election was close, but isn’t that when you most need a good prediction? These poll results are not information, but rather some devout information cultists’ simulation of information.

In 2004, 15 of 22 polls predicted that Mr. Bush would win the popular vote, but five still predicted that John Kerry would win (two predicted a saw-off, as did two in 2000). This election wasn’t quite as close as the one in 2000, but the difference between the two candidates’ support was only about 2 percentage points. Even if the polls do give you a good idea of how people are going to vote, sampling error wipes out any utility they may have when a vote is close, which is a lot of the time. Furthermore, why would we expect polls to be all that valid as measures of what the population as a whole intends to do?

First, there’s their questionable sampling to consider. Poll results often come with statements, derived from sampling theory, saying that given the size of the sample they polled, their results will be accurate within so many percentage points of the actual percentages 95% or 99% of the time; in Canada the press is required to provide such estimates. These estimates are derived from sampling theory. However, sampling theory assumes that the samples polled are representative (that is, that they are random samples of the population). That is not true of any political poll.

A random sample is one in which each member of a population has a known probability of appearing. If you draw a simple random sample of 10% of a jar containing 2,000 jelly beans, each jelly bean will, if you draw the sample properly, have a 10% chance of appearing in the sample. However, let’s say that you want to draw a sample of 10% of the members of a club with 2,000 members so that you can ask them (the members of your sample) some questions about the club. All of a sudden you don’t know the probability that each member has of appearing in the sample, for a very simple reason.

The simple reason is that people can refuse to take part in your sample. If you mail them your questionnaire, others will forget to complete it, and some of the procrastinators will never get round to it. Some just won't be interested. The problem is that you can’t tell the people who won’t return the questionnaire from the ones who will.

You will end up drawing a random sample from the population of club members who complete questionnaires. The same is true of samples in political polls. Most people, in fact, refuse to take part in political polls. Secondly, people have to be home to answer the phone before they can consent to take part in the poll. It’s likely that some large subgroups of the population (the young, for example, or the employed) are less likely to be at home than others. Thirdly, the questions have to be asked in a language the person polled understands; people who can’t understand the language of the poll well enough have to be excluded.

For these and other reasons the sample you get in a political poll is never representative of the population as a whole but rather of that minority of the population that is both able and willing to take part in polls. If that minority thinks like the majority, then your results will apply to the majority as well. If the majority doesn’t think like the minority, then the results won’t apply. The catch, of course, is that you have no idea how closely the thinking of the minority corresponds to the thinking of the majority.

Even if you were able to get a representative sample, you would still have the problem that people sometimes don’t have too accurate an idea of what they’re going to do. Sometimes they change their minds between the time they take part in the poll and the time they actually vote. Sometimes they don’t know how they’re going to vote till they get in the booth. Sometimes they don’t vote. And even if they do know how they’re going to vote, why should we assume that they’ll tell us the truth?

Election polling is a cargo cult practice. We know that examining samples has been a productive practice in science, so we draw a few samples of our own to examine. However, just as the control towers at cargo cult airstrips in Melanesia don't have the crucial operating characteristics of real control towers, the samples drawn in election polls don;t have the crucial operating characteristics of samples from which estimates of statistics (the percentage of people likely to vote for a political party, for example) can be reliably derived. And even if they did, the mutability of human intentions would probably keep them inaccurate.

Actual Analysis website

Inside the information cult (1) © 2011, John FitzGerald

Friday, April 8, 2011

Lady Luck is actually very democratic

A commercial for a poker site is advising us that Lady Luck hangs out with the better players. In fact, she demonstrably doesn't.

The only meaningful conception of luck that I'm aware of is the statistical one. I am identifying luck with the statistical concept of error, which, as we shall see, is well suited to be a conception of luck. Anyway, any result (winning a poker game, for example) can be statistically analyzed as the consequence of an effect (poker-playing skill, say) and error. Error is the sum of all those things that affect the result but aren't related to poker-playing skill — the specific cards you get, how alert you are, and so on.

Error is randomly distributed with a mean of zero (these characteristics follow from the mathematics required to distinguish effects from error). Since the effects of the variables that produce the error are not correlated with poker-playing, that means the mean error score for good players is zero, and the mean score for poor players is zero. And after all, there should be nothing about being a good player that makes you more likely to be dealt a pair of aces.

Saturday, April 2, 2011

Accuracy is not enough

Data are not necessarily information. They are informative only to the extent that they reduce uncertainty. If you want to know what programs are on television tonight, knowing yesterday's television schedule will not help you. Yesterday's schedule is full of data, but the data are no longer informative.

In psychometric terms, informative data are those which are valid – which predict events of interest to you. To be valid data must be accurate; in fact, the validity of information is limited by its accuracy. Of course, inaccurate data cannot be valid, and the maximum possible validity of accurate data is equal to the square root of its reliability coefficient.

The minimum validity of accurate data, however, is always zero. Sometimes data are not valid simply because they are distributed in a way ill-suited to the statistics which are used to assess validity; often the distribution can be modified through a mathematical transformation and validity restored. Sometimes the data are simply irrelevant or poorly defined.

In Canada a federal election campaign is under way. As usual the press commentary about it includes frequent presentation of poll results. At the moment the poll results are the unverifiable opinions of the 30% of the population that takes part in polls about what they think they'll be doing a month from now. These people probably differ significantly from people who don't take part in polls. They've probably got more time on their hands for a start, which means they're likely older, better off, and so on. That is, they are probably not even accurate estimates of unverifiable opinions. If you check the excellent Polling Report website you'll find that American polls have been dependably incompetent at predicting the results of American presidential elections, which are simple two-candidate races. In Canada, with three national parties and a big regional party they are likely to be even less effective.

Anyway, if you depend on any type of database, it should be checked regularly to ensure not only accuracy but also relevance and utility.

Accuracy is not Enough © 2001, 2011 John FitzGerald


More articles from www.ActualAnalysis.com

Friday, February 4, 2011