Showing posts with label elections. Show all posts
Showing posts with label elections. Show all posts

Tuesday, May 3, 2011

The value of political polls

I've been questioning the value of polls here, so here's some evidence of how well they work. I was interested in Ekos Politics' claim that methods of predicting the seats won by each party in Canadian federal elections "work pretty well" (the quotation is from a PDF I can no longer find on their website, but I have a copy if you want one). Here are Ekos' final projections for the election of May 2, 2011 (you can verify them here):

Conservatives: 130 to 146 seats
New Democrats: 103 to 123
Liberals: 36 to 46,
Bloc Québécois: 10 to 20
Green: 1.

And the results:

Conservatives 167 seats
New Democrats 102
Liberals 34,
Bloc Québécois: 4
Green 1.

In other words, Ekos got the Green seats right and no other party's. Of course, there is no reason they should get them right. The regional variation in voting is so great in Canada (the BQ only runs in Quebec, for example) that you'd need extensive polling in each riding to even hope to approximate the results. Even then the non-representative samples you'd be working with would seriously limit the accuracy of your estimates.

At any rate, the Ekos projections missed the two important events of May 2: the Conservative majority and the collapse of the Bloc. Journalists will probably go on acting as if polls mean something, but that doesn't mean you have to.

Tuesday, April 26, 2011

Inside the information cult (1)

In Canada we're in the final week of a federal election campaign. So far press coverage has been dominated by coverage of poll results. Unlike, it seems, most people I am skeptical of the utility of election poll results, and here I'll explain why.

Polls of people’s opinions can be a useful exercise if they’re done properly and the results are interpreted carefully. The polls that the press publishes, however, often fail to satisfy these criteria (at least in the form in which they are presented in the press). For one thing, they usually ask at most a handful of questions and don’t attempt to assess how meaningful the responses are . The results of this type of poll are not information but pseudo-information. Electoral poll results are obviously uninformative simply because they don’t predict election results accurately enough. Informative data must be valid, and in general they are not valid.

If you consult the excellent PollingReport website you will find a summary of final estimates by eighteen polls of the popular vote in the 2000 United States presidential elections. Fourteen of those predictions had George W. Bush winning the popular vote, which in fact was won by Al Gore. Sure, the election was close, but isn’t that when you most need a good prediction? These poll results are not information, but rather some devout information cultists’ simulation of information.

In 2004, 15 of 22 polls predicted that Mr. Bush would win the popular vote, but five still predicted that John Kerry would win (two predicted a saw-off, as did two in 2000). This election wasn’t quite as close as the one in 2000, but the difference between the two candidates’ support was only about 2 percentage points. Even if the polls do give you a good idea of how people are going to vote, sampling error wipes out any utility they may have when a vote is close, which is a lot of the time. Furthermore, why would we expect polls to be all that valid as measures of what the population as a whole intends to do?

First, there’s their questionable sampling to consider. Poll results often come with statements, derived from sampling theory, saying that given the size of the sample they polled, their results will be accurate within so many percentage points of the actual percentages 95% or 99% of the time; in Canada the press is required to provide such estimates. These estimates are derived from sampling theory. However, sampling theory assumes that the samples polled are representative (that is, that they are random samples of the population). That is not true of any political poll.

A random sample is one in which each member of a population has a known probability of appearing. If you draw a simple random sample of 10% of a jar containing 2,000 jelly beans, each jelly bean will, if you draw the sample properly, have a 10% chance of appearing in the sample. However, let’s say that you want to draw a sample of 10% of the members of a club with 2,000 members so that you can ask them (the members of your sample) some questions about the club. All of a sudden you don’t know the probability that each member has of appearing in the sample, for a very simple reason.

The simple reason is that people can refuse to take part in your sample. If you mail them your questionnaire, others will forget to complete it, and some of the procrastinators will never get round to it. Some just won't be interested. The problem is that you can’t tell the people who won’t return the questionnaire from the ones who will.

You will end up drawing a random sample from the population of club members who complete questionnaires. The same is true of samples in political polls. Most people, in fact, refuse to take part in political polls. Secondly, people have to be home to answer the phone before they can consent to take part in the poll. It’s likely that some large subgroups of the population (the young, for example, or the employed) are less likely to be at home than others. Thirdly, the questions have to be asked in a language the person polled understands; people who can’t understand the language of the poll well enough have to be excluded.

For these and other reasons the sample you get in a political poll is never representative of the population as a whole but rather of that minority of the population that is both able and willing to take part in polls. If that minority thinks like the majority, then your results will apply to the majority as well. If the majority doesn’t think like the minority, then the results won’t apply. The catch, of course, is that you have no idea how closely the thinking of the minority corresponds to the thinking of the majority.

Even if you were able to get a representative sample, you would still have the problem that people sometimes don’t have too accurate an idea of what they’re going to do. Sometimes they change their minds between the time they take part in the poll and the time they actually vote. Sometimes they don’t know how they’re going to vote till they get in the booth. Sometimes they don’t vote. And even if they do know how they’re going to vote, why should we assume that they’ll tell us the truth?

Election polling is a cargo cult practice. We know that examining samples has been a productive practice in science, so we draw a few samples of our own to examine. However, just as the control towers at cargo cult airstrips in Melanesia don't have the crucial operating characteristics of real control towers, the samples drawn in election polls don;t have the crucial operating characteristics of samples from which estimates of statistics (the percentage of people likely to vote for a political party, for example) can be reliably derived. And even if they did, the mutability of human intentions would probably keep them inaccurate.

Actual Analysis website

Inside the information cult (1) © 2011, John FitzGerald

Tuesday, November 9, 2010

Mayor of all Toronto except part of it

In yesterday's post I came up with some hypotheses about the vote in the recent Toronto mayoral election. Since then I've refined them a bit and tested them.

I simplified them by reducing the independent variables to two – section of the city and household income, and by hypothesizing only about the vote for the winner, Rob Ford. Hypothesizing about all three major candidates just complicates analysis, and examination of the effects of the independent variables on their votes could be done post hoc to elucidate the effects on Mr. Ford's vote.

I had originally planned to analyze the results by subdivision, but that increased the power of the statistical test so much that almost any difference would have been statistically significant. So I analyzed the results by ward; that decision gave me a nice little sample of 44.

Income was defined as the quartile in which median household income in the ward fell. The sections of the city were the outer suburbs (those wards for whom the city limits were part of their land boundaries), the inner suburbs (other wards outside the old City of Toronto as it was before amalgamation in 1998), east Toronto (roughly the old City of Toronto east of Yonge St.), and west Toronto (roughly the old City of Toronto west of Yonge St.).

So my new null hypotheses were that Mr. Ford's vote would be affected by neither of the independent variables. I was hoping, though, that they'd be affected the section of the city but not by income. Specifically, I was hoping his vote would be highest in the outer suburbs,

Mr. Ford's vote was not correlated with the total vote in a ward (r = .25; p > .05), so I didn't correct for differences in the number of votes (if they had been correlated, I would have removed the effect of total votes with regression analysis and analyzed the residual vote).

My hopes were dashed. A two-way analysis found that Mr. Ford did do best in the outer suburbs, but not significantly better than in the inner suburbs. The big difference was between the pre-1998 City of Toronto and the rest of the current city. Mr. Ford won 31% of the vote in the old City of Toronto, and 59% elsewhere.

This analysis also found a weak effect of income, but further analysis suggested this was an artefact of random variation in the number of votes cast. Analysis of the residual vote I described earlier found no differences related to median household income.

Analysis of Mr. Smitherman's and Mr. Pantalone's votes confirmed they were the candidates of the pre-1998 City of Toronto. They did better there (and Mr. Smitherman did better only in east Toronto). Ward income was not related to the votes they received.

In general, then, different sections of the city voted differently but income had little if anything to do with the results. Mr. Smitherman, the chief competitor for Mr. Ford, failed to appeal outside the oldest part of the city. Perhaps another popular explanation of the results is correct – Mr. Ford just ran by far the best campaign.

Monday, November 8, 2010

Mayoral strongholds

Torontonians seem to have concluded about their recent mayoral election that the winner was the candidate of the suburbs. I thought a little more detail might help. Here we will look at the wards in which his support, and the support for the other two major candidates, was the strongest.

I did some exploratory analysis examining the percentages each candidate won of the vote in subdivisions, then confirmed it with sorts of the percentages of votes cast in each ward. I came up with four hypotheses I will be testing further:

1. Mr. Ford's support was strongest on the outskirts of the city. His support was strongest in wards 1, 2, 4, 31, and 49, all of which are pretty far from City Hall. All have the city limits as a boundary. Mr. Ford won 67% or more of the vote in these wards.

2. Joe Pantalone was the candidate of the west end of the old city of Toronto. His strongholds -- wards 14, 17, 18, and 19 -- clustered together in the west end. Mr. Pantalone took 20% of more of the vote in these wards.

3. George Smitherman was the candidate of money. Mr. Smitherman had strong support in both Forest Hill (wards 21 and 22) and Rosedale (ward 28).

4 Mr. Smitherman was also the candidate of the east end of the old City of Toronto. His support was strong in wards 30 and 32, which lie side by side along the eastern harbour and the lake. Mr. Smitherman took 50% or more of the vote in the wards in which he was strongest.

As I said, these are just hypotheses so far. I'll be souping up my data file, and then I'll be testing these hypotheses. More soon.

Actual Analysis website

Tuesday, November 2, 2010

Religion and mayoral choice in ward 26

People have been speculating about the effect of religion in the Toronto mayoral elections a week ago. The idea is that members of some religions would be less likely to vote for George Smithermen, who is gay and married to another man.

We saw in the last post that voters in Jewish neighbourhoods in Ward 21 were in fact most likely to vote for Mr. Smitherman instead of the other candidates. In this post we'll look at a Muslim neighbourhood, Thorncliffe Park in Ward 26.

As it turned out, Mr. Smitherman did finish second in the polls in Thorncliffe Park. However, he finished frst in the rest of the ward. He received 33% of the vote in Thorncliffe Park, and 44% in the rest of the ward. A powerful chi-square test finds this difference to be significant, while the weaker median test I described in the last post doesn't. However, the powerful test estimated an infinitesimal probability that the dfference was random, and the weak test estimated that the probability was less than .09, so I'm considering this difference statistically sgnificant.

However, of the eleven percentage points that went missing for Mr. Smitherman in Thorncliffe Park, Mr. Ford picked up only four. Most of the vote Mr. Smitherman lost went to three candidates with Muslim names, none of whom, however, made an issue of their being Muslim. One was an anti-poverty advocate, another a civil-rights advocate (and not the kind that thinks civil rights mean other people should shut up about their -- the advocate's -- religion), and one has campaigned before as an anti-unemployment candidate. They could simply have been taking a greater part in the public life of Thorncliffe Park than the other candidates.

As I concluded before, if religion affected the mayoral vote, it was probably weakly, and in interaction with other variables.

Main Actual Analysis site

Thursday, October 28, 2010

A streetcar named Doesn't Matter

I have just downloaded summaries of last Monday's Toronto municipal elections from the City of Toronto's open data site. A contentious issue in the ward, and a controversial issue throughout the city, has been the renovation of the streetcar line along St. Clair Avenue West. The incumbent councillor, Joe Mihevc, was blamed by many for problems with the renovation. Although he was re-elected, I wanted to see if the renovation had affected where he got his support.

Only two candidates, Mr. Mihevc and Shimmy Posen, won any of the polling subdivisions; Mr. Mihevc won 20 and Mr. Posen 10. My idea was that if the streetcar-line renovation had affected his support, Mr. Mihevc would have drawn his support from polling subdivisions away from St. Clair Ave.

On the map of the ward below, subdivisions won by Mr. Mihevc are shown in red and those won by Mr. Posen in blue. Subdivisions outside the ward are in grey. St. Clair Avenue is marked by the black lines extending beyond the borders of the ward.Clearly, Mr. Mihevc's support was strong along St. Clair Avenue. The variable that chiefly determined support was income, with Mr. Posen's strength almost entirely in the affluent neighbourhoods north of Nordheimer/Cedarvale Ravine, and Mr. Mihevc's chiefly in the south. However, Mr. Mihevc won some well-off subdivisions near St. Clair West as well. Despite all the problems created by the renovation of the streetcar line, problems which were raised by the successful mayoral candidate at an all-candidates' meeting in the heart of Mihevc territory just before the election, St. Clair West remained part of Joe Mihevc's stronghold.