Margin of Error
What is a MOE?
If you click on the Morning Consult link from the screen capture of the realclearpolitics site mentioned earlier, it takes you to an Axios news report on the poll. It's good practice to look for the methodology section of the poll when you click on these links. These links sometimes go to a news report, but they can also point to a page from the company that published the actual poll. News sites usually don't do as good a job highlighting the poll methodology. But at least Axios published the following blurb:
"The fine print: Data from the polls conducted through online interviews were weighted to approximate a target sample of registered voters, per Morning Consult. The margin of error was +/- 2 percentage points."
The part I accented in red is the important part. MOE stands for Margin of Error. Before I explain what this is and why it's important, let's take a little quiz.
Quiz: A Margin of Error measures the possibility that the poll contains what errors?
A. The poll's questions were not clear or written in such a way that it was biased toward a particular answer.
B. The poll had errors in polling methods because online surveys are less reliable than telephone interviews.
C. Poll respondents didn't share their true feelings but only told the poll worker what they thought they wanted to hear.
D. Poll workers were not consistent in the way they asked a question.
E. All of the above and more.
The correct answer is actually none of the above. A Margin of Error doesn't measure almost all sources of error a poll can have. And shouldn't provide any confidence at all that a poll is free of errors except for one specific type.
What is this specific type of error? The Axios blurb doesn't provide enough information to answer this question. And even a statistician could not tell you anything about the accuracy of this poll based on the poor wording Axios used.
Let's look at the Data for Progress link. This is the methodology blurb:
"On June 28, 2024, Data for Progress conducted a survey of 1,011 U.S. likely voters nationally using web panel respondents. The sample was weighted to be representative of likely voters by age, gender, education, race, geography, and 2020 recalled vote. The survey was conducted in English. The margin of error associated with the sample size is ±3 percentage points. Results forsubgroups of the sample are subject to increased margins of error. Partisanship reflected in tabulations is based on self identified party affiliation, not partisan registration. For more information please visit dataforprogress.org/our-methodology."
I again note the Margin of Error part of the methodology in red. This site is at least the site of the publisher of the poll. So it must be more informative, right? But again, a statistician would be scratching their head at this description of the poll's Margin of Error.
Let's try one more time but outside of the green box. Let's try the Rasmussen Reports link. This is the methodology blurb:
"The survey of 1,000 Likely U.S. Voters was conducted June 20, 2024 by Rasmussen Reports. The margin of sampling error is +/- 3 percentage points with a 95% level of confidence. Field work for all Rasmussen Reports surveys is conducted by Pulse Opinion Research, LLC. See methodology"
Again, the important part is in red. But this time we actually get a full description of what the MOE is. It's the Margin of Sampling Error at a 95% Level of Confidence. Finally, this is a description of what the MOE is referring to that gives us all the information we need to guage what is being measured. It is the Margin of Sampling Error. That extra word, "Sampling" is important. And the level of confidence is 95% that the real statistic being measured is within the +/- interval.
So for the Rasmussen Poll, they report a statistic that 49% of the people surveyed support Trump for president. There is a 95% chance that the actual number due to sampling error is 46% to 52%. Similarly, Biden's actual support will be between 37% and 43% at the 95% level of confidence due to sampling error. The two intervals do not overlap. And you will often hear news talking heads say, "If the election were held at the time of the poll, Trump would be ahead outside the margin of error." But they leave the most important part out. This would only be true for the margin of sampling error. So what is it?
Sampling Error
Remember what it means to make that statement, "IF the election were held at the time of the poll, Trump would be ahead..." The poll is supposedly reporting if all voters were counted on election day, what the results would be. In 2020, 154.6 million people voted for president according to the Census Bureau. However, the date of the next presidential election would be November 5th, 2024. Therefore, surveying the actual people who will show up for election day is impossible for two main reasons.
1. The logistics of surveying around 150 million people would be very expensive and time consuming.
2. Since the election will happen in the future, there really is no way to even know who will actually take the time to send in a mail ballot or show up in person on election day.
So although the poll is making claims about the larger 150+ million population that will show up on election day 2024, the poll workers can't actually poll them.
So polling companies conduct a survey with a much smaller number of people in the population known as a sample. A simpler example might help here.
Say you had a bag that contained 10 marbles. 5 of the marbles are black. And 5 of the marbles are white. The total popluation of marbles is 10. Let's say that you can only reach your hand in and pull out two marbles. Those two marbles are the sample.
And you can see in this small sample that if you reached your hand in and pulled out two marbles at random, you have a good chance of reporting the wrong color proportions in the total population. The actual number is 50% black and 50% white. So if your sample was one white marble and one black marble, you would report the correct percentages. But what if the two marbles you pulled were the same color? In that case you would report the incorrect percentages that the colors were 100% black or 100% white. That's a very large potential error with only a two marble sample out of a population of ten.
The chance you would get the actual percentages +/- some number with some level of confidence would be the Margin of Sampling Error for this population and sample of marbles. There are very straightforward equations to calculate these, but that would be a little far from a simple truth, so I'll spare you those. Just know that the Margin of Sampling Error is calculated based on the total population of people you want to survey for the poll and the sample size of people you are picking by calling or using web surveys, etc.
The simple truth is that the Margin of Sampling Error only refers to this pick voters out of a bag potential to get the reported statistic wrong because the sample did not match the actual population. It is not a green light that all polls that are outside these Margins of Sampling Errors are free from all the other errors a poll can have (including other types of sampling errors).
It is best used as just an indication of whether or not you can say that one candidate is ahead and not in a tie or behind based on the reported 95% confidence interval. For the polls pictured at the top of this page, only the Rasmussen Reports Poll has Trump ahead outside the Margin of Sampling Error. All other polls should be read as a statistical "can't really tell because the Margin of Sampling Error is too large".
The 95% level of confidence is a standard level used for polls. But not reporting the confidence level of the poll or the proper name of the Sampling Error being measured means the people who quoted or wrote the methodology sections missing them either didn't understand what a Margin of Sampling Error is, don't really care about the details as long as they can claim their poll is error free within the interval, or both.
Reducing the Sampling Error Interval
And the last simple truth I'll talk about with Margin of Sampling Error is how easy it is to make the +/- interval smaller. Because the Margin of Sampling Error is measured by looking at the total population size and the sample size, to drive this interval down, all the polling company has to do is increase the sample size. Look at the expanded picture of polls below from the same realclearpolitics site referenced previously.
In addition to the one day reaction poll to the first 2024 presidential debate (outlined in blue), Morning Consult also conducts regular three day tracking polls with the two previous polls to the debate outlined in green. In the sample column, the one day poll included a sample of 2068 people. While the two tracking polls included a sample of more than 10,000 taken over three days. And the result of increasing the sample size (assuming that the population being sampled was the same) was to increase the accuracy of the sample to only a +/-1 confidence interval from+/-2.
All the other polls listed have much larger confidence intervals, up to 3.8 in the NPR/PBS/Marist poll. So this poll has Trump at 49%. But the real number is from 45.2% to 52.8% within the confidence interval. That's quite an uncertain spread. Compare that with the Morning Consult poll just above it. It has Trump at 43%. So the real number is from 42% to 44% within the confidence interval. A much better interval to use to figure out if the poll actually indicates whether or not one candidate really has an advantage over the other or if it's too close to tell.
I realize that polling 10,000 people vs 1,000 people is more expensive. But in a race as important as the election of the President of the United States, all of these polls should be decreasing their confidence intervals to a +/-1 Margin of Sampling Error.
That they don't speaks to how little the polling companies care about the accuracies of their polls and how little the news agencies care that they have to constantly say that two candidates , are supported by such and such a percentage but it's within the margin of error. Translation, "Statistically, we have no idea who is ahead or if there is a tie." The reality is they probably aren't tied. It's just the poll is so inaccurate that there's no way to tell who is ahead.
The large sample Morning Consult polls still can't find a difference outside the Margin of Sampling Error. But at least at the small confidence interval of +/-1, most people would say that saying it's too close to call is probably accurate.
So the simple truth is that most polling companies are too lazy or cheap or both to provide a meaningful Margin of Sampling Error for close political races. And news organizations either don't understand the importance of having a smaller confidence interval or they just don't care enough about the accuracies of their own stories to demand more from polling companies.
In addition to the size of the sample, there is an LV or RV just after the number. I'll talk about this bit of jargon in the next section to discuss another type of sampling error, choosing the wrong sample.
