All the polls in the Virginia governor's race from mid-July to November that Terry McAuliffe had a lead over Ken Cuccinelli, and in the final week the median lead was 7%, which using my system which factors in the size of the sample as well made McAuliffe a 99.4% favorite to win. He won, but by a much smaller margin, leading many people to ask why the polls were so wrong.
The best explanation is a well-known tendency for third party candidates to do better in polls than they do in the actual election. The Libertarian candidate Robert Sarvis was getting good numbers for a third party, the last five polls giving him 13%, 12%, 10%, 8% and 4%. The median was 10% and the average 9.4%. His actual numbers were 6.6% and a lot of his alleged support appears to have favored the Republican Cuccinelli when the curtains on the booths were drawn.
Many websites give the margin of the lead instead of the "margin of error" these days, and that makes some sense, since the "margin of error" is the 95% confidence interval for the actual result, which really should be measured after the undecided and no preference numbers are removed from the tally. If all we care about is the final lead, the Emerson College Polling Society bathed themselves in glory by predicting a 2% lead for McAuliffe, very close to the actual 2.5% final margin when everyone else overstated the lead.
But looking at the 95% confidence intervals after undecideds are removed actually makes Emerson College look like the worst polling company of the final five. Here's why.
Emerson College
Raw numbers: sample of 874, McAuliffe 42%, Cuccinelli 40%, Sarvis 13%
Removing undecided: sample of 830, McAuliffe 44.2%, Cuccinelli 42.1%, Sarvis 13.7%
Total missed percentage from reality: 14.25% (5th place out of 5)
95% confidence interval (aka margin of error) for each candidate:
McAuliffe 47.59% to 40.83% (missed 47.97%)
Cuccinelli 45.46% to 38.75% (barely missed 45.47%)
Sarvis 16.02% to 11.35% (massively missed 6.56%)
Emerson missed all three true vote totals, only barely in the case of the front runners, but by a wide margin when looking at Sarvis.
PPP
Raw numbers: sample of 870, McAuliffe 50%, Cuccinelli 43%, Sarvis 4%
Removing undecided: sample of 844, McAuliffe 51.5%, Cuccinelli 44.3%, Sarvis 4.1%
Total missed percentage from reality: 7.15% (second best)
95% confidence interval (aka margin of error) for each candidate:
McAuliffe 54.92% to 48.17% (missed 47.97%)
Cuccinelli 47.67% to 40.98% (captured 45.47%)
Sarvis 5.47% to 2.78% (missed 6.56%)
PPP was the only company to underestimate Sarvis, so they were very close to the real numbers for Cuccinelli and overestimated McAuliffe. One company did better at the distance from all three candidates.
Rasmussen
Raw numbers: sample of 1002, McAuliffe 43%, Cuccinelli 36%, Sarvis 12%
Removing undecided: sample of 912, McAuliffe 47.3%, Cuccinelli 39.6%, Sarvis 13.2%
Total missed percentage from reality: 13.25% (4th place out of 5)
95% confidence interval (aka margin of error) for each candidate:
McAuliffe 50.59% to 44.01% (captured 47.97%)
Cuccinelli 42.73% to 37.76% (missed 45.47%)
Sarvis 15.38% to 10.99% (massively missed 6.56%)
Like Emerson, the big overestimate of Sarvis skews these numbers badly, but they did capture McAuliffe's numbers.
Newport College
Raw numbers: sample of 1028, McAuliffe 45%, Cuccinelli 38%, Sarvis 10%
Removing undecided: sample of 965, McAuliffe 48.4%, Cuccinelli 40.9%, Sarvis 10.8%
Total missed percentage from reality: 9.22% (3rd best)
95% confidence interval (aka margin of error) for each candidate:
McAuliffe 51.54% to 45.23% (captured 47.97%)
Cuccinelli 43.96% to 37.76% (missed 45.47%)
Sarvis 12.71% to 8.80% (missed 6.56%)
Like several other polls Newport captured the McAuliffe number but underestimated Cuccinelli and overestimated Sarvis.
Quinnipiac
Raw numbers: sample of 1606, McAuliffe 46%, Cuccinelli 40%, Sarvis 8%
Removing undecided: sample of 1510, McAuliffe 48.9%, Cuccinelli 42.6%, Sarvis 8.5%
Total missed percentage from reality: 5.83% (best of 5)
95% confidence interval (aka margin of error) for each candidate:
McAuliffe 51.46% to 46.41% (captured 47.97%)
Cuccinelli 45.05% to 40.06% (missed 45.47%)
Sarvis 9.92% to 7.10% (missed 6.56%)
Coming off their great call of the New York City Democratic comptroller primary, Quinnipiac is again the best polling company of those that polled in the last week, though it must be said they are the best of a bad bunch. Their downfall in the 95% confidence intervals was the size of their sample. Big samples mean smaller margins of error, but not necessarily more captures of the real numbers.
I don't change my model very often, and even though the candidate who my system favored actually won, I was a little embarrassed by assigning such a huge Confidence of Victory number (99.4%) to a race that turned out to be so close. The next time there is a third party candidate polling with significant support but no real chance to win, I'm going to figure out a fair way to re-assign the numbers, with the assumption that a Libertarian will lose support that will go Republican and a Green will lose support that will go Democratic.
Showing posts with label polling data. Show all posts
Showing posts with label polling data. Show all posts
Monday, November 11, 2013
Monday, November 4, 2013
VA Governor's race:
Election Eve update
Several elections are being held tomorrow. The polls say there is no suspense whatsoever in the governor's race in New Jersey or the mayor's race in New York, where repsectively Chris Christie and Bill De Blasio are expected to win by double digits. In Virginia, the governor's race between Terry McAuliffe and Ken Cuccinelli looks closer, but if my system hold true, there is extremely little chance of an upset. (I have notbeen checking the data for the down ticket races. Daily Kos says only the Attorney General race looks close.)
For a time in early September, Cuccinelli appeared to be gaining on McAuliffe, but that trend turned around and Cuccinelli's team has been starved for good news since. The last three updates on October 20, October 27 and November 3 have all put the Confidence of Victory number for McAuliffe at a formidable 99.4%. (In a strange coincidence, these three numbers came from three separate polls from three different companies, the well-known Quinnipiac and Rasmussen and the lesser known Newport College Polling Group.
My system does not predict the margin of victory, but except for one outlier predicting a close 2% margin, the rest of the polls in the past week make it look like the margin will be 6% to 7%.
I will report back on Wednesday about the results.
Monday, September 23, 2013
VA governor's race:
Late September update
In Virginia, a state Barack Obama won by 4% in 2012 and 6% in 2008, the statewide races feature Tea Party favorites running on the Republican ticket. In the governor's race, both the Democrat Terry McAuliffe and the Republican Ken Cuccinelli have large negative ratings, but the basic math of the electorate says one thing clearly.
The Republicans are the minority party in the land. They've lost two presidential elections in a row, they don't hold the Senate and got less votes across the country in last congressional election, holding onto the majority in large part due to gerrymandering.
For all that, the needle is moving slightly for Cuccinelli if we give all polls the same weighting. Here are the Confidence of Victory numbers for recent polls, given by company name and date of poll. All give McAuliffe the advantage.
Washington Post 9/22 89.1%
Marist 9/19 90.3%
Harper[R] 9/16 94.2%
Roanoke 9/15 63.5%
Quinnipiac 9/15 84.9%
The good news for Cuccinelli: McAuliffe isn't cruising above 95% anymore and Quinnipiac, still crowing over calling the NYC comptroller race all by themselves with their new likely voter model, says the race is getting close, but still isn't really close.
The bad news for Cuccinelli: You see those two big dives downward in the graph, the one in May and the second in July? The first was a Washington Post poll and the second was a Roanoke poll. The only polls that have had Cuccinelli ahead now have him behind.
The other bad news for Cuccinelli: My system may say 85% Confidence of Victory, but the track record so far says almost no underdog wins unless the Confidence of Victory splits about 60%-40%. The biggest upset against my system since 2008 was Stringer beating Spitzer for NYC Comptroller when Spitzer had 68.9% Confidence of Victory, the election when only Quinnipiac got the numbers right. (They have crowed loudly about that, but they had De Blasio cruising over 40% of decided voters and he only just beat that margin.)
The election is still about seven weeks off, which means my system is not built to make a prediction, but the small gains Cuccinelli has made are not what he needs. Romney was in a similar situation and saw some small gains after Obama's poor showing in the first primary, but those faded quickly enough. Cuccinelli needs something bigger than we've seen so far to close the gap, and appealing to the Republican "base", such as it is, is a step in the wrong direction right now.
More updates when I get more data.
Wednesday, September 11, 2013
Results from the Democratic mayoral race in New York City.
Almost all the ballots have been counted in the New York City municipal elections held yesterday. My last post here was on Sunday, but there was a late poll from Public Policy Polling that changed my numbers. After working on how I was going to consider the "median result" from a set of four polls in a three person race where there was a 40% threshold, my last prediction on Monday evening was posted to Twitter.
Final call on NYC Dem Mayor's race, different from
55% De Blasio 1st ballot
30% De Blasio/Thompson
15% De Blasio/Quinn
Professor Wang's last prediction was a 90% probability of De Blasio on the first ballot.
The current vote count has Bill De Blasio at 40.3% and his closest rival Bill Thompson at 26.2%. Because his margin over the 40% threshold is so slim, we will have to wait for absentee ballots and an automatic recount for a result this close. The reported estimate is next Monday for a conclusive result.
Here on the blog, I made no mention of the comptroller's race, but I did mention it on Twitter in a two part tweet.
NYC mayoral primary today.
If Spitzer loses the Comptroller race, full credit should go @QuinnipiacPoll, the only organization to favor his rival Stringer. [2/2]
There were several polls of this race as you can see at this link to the Real Clear Politics polling page. There was nothing that made this look like a close race until late August, with Eliot Spitzer's hopes for a political comeback looking very strong while Manhattan borough president Scott Stringer seemed unable to make any headway. But then Quinnipiac had a poll that said the race was tied, then another Quinnipiac poll had a small of 2% for Stringer on September 1, then a commanding 7% lead a week later.
The actual result was Stringer by 4%, 52% to 48%.
The reason my system - and Professor Wang's - favored Spitzer was that no other pollster agreed with Quinnipiac, not even once. After the first Quinnipiac poll that showed a close race, Siena, Marist and PPP all released polls, all of them favoring Spitzer.
So my system failed to predict this race. Often, when I get a race wrong, I try to find ways to adjust things to improve, but if I was given a similar situation tomorrow, I wouldn't change a thing. Here are my reasons.
1. I never weigh in more than a single poll from any polling company and always the most recent.
2. Because Quinnipiac only gets one result counted, they had one poll out of three in the final mix, along with PPP and Siena.
3. In a three poll situation, I take the median result, not the average of the three. Quinnipiac was the outlier, but outliers aren't often the closest to correct. Obviously, it will happen sometimes, but it is rare enough that I do not want to second guess myself in a situation like this. Polls miss the mark, sometimes by making bad assumptions, sometimes just by the fact that randomness is involved. For example, of the three polls Quinnipiac had the largest sample size, but that does not always (or even often) mean the most accurate. Looking at the mayoral polling and factoring out the undecided, Quinnipiac had De Blasio at 46% of the people who had a preference, while PPP had him at 42% and Siena had him at 39%. In this case, it's the low outlier that was closest.
So this means at least one more report on the mayor's race. It is widely agreed that the Republican primary winner Joe Llota will be very hard pressed to win in the general election. De Blasio got 260,000 votes in the Democratic primary, Llota finished first with over 50% in the Republican primary with about 30,000, which would have given him sixth place in the Democratic race.
And while I did not use their poll, I want to congratulate Quinnipiac for being alone with the correct result in the comptroller's race. It is a field filled with randomness, but they had Stringer with a chance or the lead three separate times in barely two weeks when no one else did. That's a record they can point to with pride.
Tuesday, June 25, 2013
Massachusetts Senate Race: Markey(D) vs. Gomez(R)
Final election day update
The special election to fill John Kerry's Senate seat in Massachusetts is today and Democrat Ed Markey has lead in the polls from day one. There have been a couple polls that showed a close race, but even in those Markey has been ahead, once by as little as a single percentage point. The four polls in the final week have had him ahead by 3%, 7% and twice at 10%. The Confidence of Victory of the median poll has never slipped under 95% for Markey.
There is always a first time, but in over 100 races where I have used the Confidence of Victory method, the favorite wins an overwhelming amount of the time and a favorite at over 90% is yet to lose. Voter turnout is always important and the method does not predict the margin of victory, just the victor. I'll be back tomorrow to give the results.
Monday, June 17, 2013
Massachusetts Senate Race: Markey(D) vs. Gomez(R)
17 June Update
New polls have been released since last Monday regarding the special Senate election to fill John Kerry's open seat in Massachusetts. On June 9, two polls were released giving Democrat Ed Markey a 7 point lead over Republican Gabriel Gomez. Two more recent polls show double digit leads, one by the Republican polling company Harper (49%-37% on 6/11) and the most recent by Boston Globe (54%-43% on 6/14).
We are eight days away from the polls closing and nearly no movement in Confidence of Victory numbers, now sitting at 97.5% for Markey. A week out it's hard to move the needle, especially given the increased popularity of voting by mail. Polling companies would like it if news organizations would add the phrase "if the election were held when the poll was taken" to any report on a poll, but with absentee ballots the election is effectively taking place now. I expect more polls this week and at least one more report on the race.
Monday, June 10, 2013
Massachusetts Senate Race:
Markey(D) vs. Gomez(R)
10 June Update
We are now about two weeks out from the special election to fill John Kerry's Senate seat in Massachusetts. Polling data had been scarce, but it's starting to pick up. Now that Suffolk has published their data from a poll that ended yesterday, we have five polls from 2 June to 9 June. Here is the list from newest to oldest. (The next newest poll not included is from 15 May and was conducted by Public Policy Polling [PPP], who also completed a poll on June 4. This poll would be excluded from the list either for being too old or for being done by a company already represented.)
9 June: Suffolk 48%-41% Markey n=500
Confidence of Victory 95.2%
5 June: McLaughlin[R] 45%-44% Markey n=400
Confidence of Victory 58.4%
4 June: YouGov 51%-40% Markey n=500
Confidence of Victory 99.5%
4 June: PPP[D] 47%-39% Markey n=560
Confidence of Victory 98.0%
2 June: New England College 52%-40% Markey n=500
Confidence of Victory 100.0%
The [R] and [D] behind McLaughlin and PPP respectively indicate that they are partisan companies and I include that out of a sense of fairness. I do not skew the data, I let all the pollsters have their say and take the median, which this week happens to be PPP at 98% CoV. McLaughlin is currently the lone outlier saying the race is close and even that poll does not give Gomez the lead.
Again, I consider these statements as snapshots of the situation instead of forecasts, but in general I would say Markey looks to be in the lead comfortably and Gomez has some small momentum which has to improve markedly for him to have a chance.
More updates as more polls come in.
Wednesday, May 29, 2013
Virginia governor's race:
McAuliffe(D) vs. Cuccinelli(R)
Update #2
Polling is still scarce on the governor's race in Virginia, which is as it should be since the election is in November. A new poll came out this week from Public Policy Polling showing a very similar pattern to the poll from mid-month from Quinnipiac.
In brief, the last two polls show a lead for the Democrat McAuliffe over the Republican Cuccinelli, right now at 42% to 37% compared to 43% to 38% earlier in the month. Both candidates have large negatives and about 20% of the public is not tuned in yet.
The first poll in May showed a lead for Cuccinelli. While the Confidence of Victory now shows a 92.8% chance for McAuliffe, that would only be true if the election were being held today, which obviously it isn't. More than that, if the election were right around the corner, I would not want to base my results on a single poll with 21% either undecided or voting for some other option. Unlike Nate Silver, I never talk about these early results as showing anything about what will happen in November. These are just early snapshots.
Saturday, May 18, 2013
Massachusetts Senate Race:
Markey(D) vs. Gomez(R)
Update #2
A new poll has been taken for the Massachusetts senate race and Democrat Ed Markey continues to lead. This poll was taken on May 15 and the most recent poll before this one was completed on the 7th, so using my seven day rule, this is the only poll being considered.
Most recent poll: 15 May 2013
Polls taken within a week of the most recent: 1
Lead: 7% lead for Markey(D)
Confidence of Victory of the median poll: 98.2%
I do prefer having a larger set of data on which to base a post, but I will remind readers that I consider these updates to be snapshots of an evolving situation rather than predictions of what will take place on election day. While this is "just one poll", the resulting Confidence of Victory numbers are much the same as we saw last week. I intend to have an update at least once a week until the election on June 25 if there are enough new polls to warrant such regular reports. I'm not sure how many polling companies will be working this race in May, but I fully expect multiple polls a week in June.
Thursday, May 16, 2013
Viriginia Governor's Race:
McAuliffe(D) vs. Cuccinelli(R)
Virginia's gubernatorial election doesn't take place until November, but we are getting early polling results. I do not consider these numbers a prediction of what will happen, but more like a snapshot of the current situation.
Most recent poll: 13 May 2013
Polls taken within a week of the most recent: 1
Lead: 5% lead for McAuliffe(D), 43% to 38%
Confidence of Victory: 97.7%
I don't consider polls taken this far in advance to be completely meaningless, but the phrase Confidence of Victory rings hollow right now, and I say that as the inventor of the phrase. It's all about the proviso "if the election were held when the poll was taken", and obviously, the election is not being held this week or even this month. More than that, a poll at the beginning of May gave Cuccinelli a commanding lead that would have given him a 99.7% Confidence of Victory.
Still, this is an important race and I will continue to cover it throughout the year.
Friday, May 10, 2013
Massachusetts Senate race:
Markey(D) vs. Gomez(R)
There are only a few elections this year, but I mean to cover the Senate and governor's races using the Confidence of Victory method, which takes polling data - percentages and size of sample - and turns it into probabilities of victory for both sides, or all three sides if a race has three truly competitive candidates, a rare occurrence in American politics.
In Massachusetts, there is a special election for the Senate seat vacated by John Kerry. It will be held on June 25.
Most recent poll: 7 May 2013
Polls taken within a week of the most recent: 4
Largest lead: 17% lead for Markey(D)
Smallest lead: 4% lead for Markey(D)
Confidence of Victory of the median poll: 97.4%
Markey has much greater name recognition than Gomez and Massachusetts is a generally blue state. The big lead is from the latest poll, but my system is interested in the median, not the most recent.
I do not think of this number as a prediction, but instead as a snapshot of the current position. I'll report back at least once a week on this race, more often if it looks like its getting closer and there is enough polling data to track changes at a faster pace.
Tuesday, May 7, 2013
Results from South Carolina 01 House race.
The polls are closed in South Carolina and Mark Sanford is projected to be the winner by about a 10% margin. The last poll only put him up by 1%, but my system isn't designed to predict the margin of the win, only the direction.
Nate Silver on Twitter was not very confident because of the minimal amount of polling, but he did make Sanford a favorite at about 64% chance to win, which is pretty much what my system thought as well.
There is another special election to replace John Kerry in the Senate on June 25. I'll keep track of the polls there and make periodic reports of the snapshots of the race, as well as a prediction on the morning of the 25th before the polls open.
Monday, May 6, 2013
Confidence of Victory prediction for the South Carolina-01 special House seat.
Unlike the races last November, the special race for the South Carolina 1st District seat is not being polled to within an inch of its life. The only company keeping track is Public Policy Polling (PPP), a Democratic polling company that does automated polls often called robo-polls. After some early polls that had Democratic candidate Elizabeth Colbert-Busch ahead, the latest poll released today has a 1% lead for former governor Mark Sanford at 47% to 46%, with another 4% favoring the Green candidate Eugene Platt and the rest undecided. The sample was taken this weekend for 1,239 likely voters.
The system I use, which I call Confidence of Victory, takes the two leading vote getters in a situation like this and calculates the probability that a lead in this poll will translate into a victory in the election. With this lead and this size of a sample, Sanford has a 64% Confidence of Victory and Colbert-Busch's probability is at 36%. My system assumes a 0% chance for the Green candidate so far behind.
I would like more polling data from more companies, but if only one company reports data, it would be hard to top PPP. In their polls of the electoral college and Senate races from the last week of the general election, they went a perfect 33-0 in predicting winners. We will see tomorrow night how my prediction from this one polls fares and I will report back.
The system I use, which I call Confidence of Victory, takes the two leading vote getters in a situation like this and calculates the probability that a lead in this poll will translate into a victory in the election. With this lead and this size of a sample, Sanford has a 64% Confidence of Victory and Colbert-Busch's probability is at 36%. My system assumes a 0% chance for the Green candidate so far behind.
I would like more polling data from more companies, but if only one company reports data, it would be hard to top PPP. In their polls of the electoral college and Senate races from the last week of the general election, they went a perfect 33-0 in predicting winners. We will see tomorrow night how my prediction from this one polls fares and I will report back.
Thursday, May 2, 2013
The Statistical World, part 4
It turns out I'm not fantastic at flipping mental coins. Obama barely beat Romney in Florida. But in the other 83 races, my system picked the winner every time.
Nate Silver of the New York Times also made predictions in all 84 of these races. His system also called Florida a toss-up, which is a credit to both of us, but also a little lucky. In other elections, my system has called a race a toss-up and one side or the other won handily. In the other 83 races, Nate went 81-2, missing two Senate races in Montana and North Dakota, two results that my system got right.
So if we include my guessing call of Florida, I went 83-1 and Nate went 81-2. Our percentages are 98.8% and 97.6% respectively, both of which count as excellent when it comes to prognostication.
Are we geniuses or what?
Well, I'm going to say "or what". The general election polling data was non-stop for several months. Looking back at my records, there were 700 polls dealing with the 84 races in the last five weeks of the race. I started keeping daily track of the median electoral college result after Obama's disastrous first debate appearance, and his numbers did suffer. But then came the Biden-Ryan debate and second Obama-Romney debate and Romney finally repudiating his "47% comment" and the Obama advantage moved up to where it had been as of early October. It was hard to pick winners because the races were not very close and the opinions were not taking huge swings, just small ones.
Here is the best data that shows Silver and I are not geniuses and that is the primary election season. This graph shows the ups and downs of the four candidates still in the race in February, Mitt Romney (green), Newt Gingrich (gray), Rick Santorum (brown) and Ron Paul (gold). I also tracked NONE OF THE ABOVE in black.
There were a lot of polls during this month, but not anywhere near the number there were in the general election. More than that, the Republican electorate was in an amazing state of flux. You can see Santorum climbed from third place to first place then back down to second in the space of four weeks. More than that, NONE OF THE ABOVE was holding steady at about 15% throughout the month.
In the primaries that month, Nate and I weren't scoring in the 98th or 99th percentiles. The data was sketchier and our predictions suffered. Predictions from polling data is a lot more accurate than predicting the results of sporting events, to give just one example, but even taking the average (or median) of a lot of polls can be shaky, especially when NONE OF THE ABOVE is well over 10% this close to the election.
Nate's book The Signal and the Noise is a study of why some predictions do well and others do not. He thinks that in the long run we are going to learn how to do better in general. I'm not convinced. Sometimes, the randomness inherent in a system will overwhelm the cleverest human prediction methods.
Subscribe to:
Posts (Atom)


















