I think "sampling on the dependent variable" is about to beat "ignoring hidden factor correlation" as the most common data analysis error in the wild.
Walter is a very popular professor of physics at the Boston Institute of Technology [name of institution cleverly disguised for legal reasons], who teaches a 400-student class in the largest classroom at BIT. The other 20 professors of physics are very boring, so at any given time they have on average 5 students each in their classe.
A journalist stands at the door of the physics department and asks every fourth student how full their physics classes are. The results are as follows:
- 100 students say that their class was completely full;
- 25 students say that their class was mostly empty.
This is reported as "4 out of 5 classes are full at BIT; new building needed to address the lack of space, new faculty must be hired urgently."
Did you see the error? It's subtle.
Here's a visualization of the process to help see it:
In fact, only one class is full. The problem is that the likelihood of a student being in the sample (a random sample of students coming out of the building) is proportional to the variable of interest (the number of students in the class); in other words, the journalist is sampling on the dependent variable.
The more full a class is, the more over-represented that class will be in the sample of students.
This looks like some rare error, the kind of thing that would only happen to hapless journalists, except that it happens all the time and in serious circumstances.
Consider the case of health authorities trying to determine the seriousness of a condition, namely how many of the people with the condition die. They could count the cases that get tested and compute the fraction of those that die. (That's what most of the preliminary COVID-19 case fatality rate numbers in the media are.)
And that's the same error that the journalist made.
In this case, the dependent variable is not size of class, it's seriousness of disease, and the sampling problem is not with the number of students in a class, is with the people who choose to get tested. These people choose to get tested (or get tested at a hospital when admitted) because they have symptoms that make them take the trouble.
In other words, the more serious the level of the disease a patient P has, the more likely P will be tested (sampling on the dependent variable, again), and more of these tested patients will die than if the testing was done to a random sample of the population.
(This is different from the truncation argument made in this previous post. Truncation is also a type of sampling on the dependent variable; a form that is easier to correct, as the non-truncated part of the sample distribution is the same as the population distribution up to scaling.)
To illustrate the effect of different degrees of sampling on the dependent variable, let us consider the case of a uniformly distributed variable (in the population) and different degrees of sampling:
Let's consider two persons, A with $x_A=0.2$ and B with $x_B=0.4$. With correct, random, sampling, A and B would have equal chance of being in the sample. With sampling proportional to $x$, B would be twice as likely to be in the sample than A, biasing the sample average upwards relative to the population; with sampling proportional to $x^2$, B would be four times more likely to be in the sample, which would bias the sample average even more.
To put this $x^2$ in the context of, for example, COVID-19 tests, a sampling proportional to $x^2$ means that people in a group X with symptoms twice as bad as people in a group Y will be four times more likely to seek treatment (and be tested). Basically, each person in group X will be counted four times more often in the statistics than each person in group Y (the group with less serious symptoms).
(The other degrees, $x^3$ and $x^4$ capture cases where people avoid the hospital unless their symptoms are serious or very serious.)
As we can see from the charts in the image above, the more distortion of the underlying population distribution the sampling process creates, the higher the sample average, and all the while the population average stays at a constant 1/2.
Sampling on the dependent variable: something to keep in mind when people talk about dire situations in the news.
Non-work posts by Jose Camoes Silva; repurposed in May 2019 as a blog mostly about innumeracy and related matters, though not exclusively.
Showing posts with label Innumeracy. Show all posts
Showing posts with label Innumeracy. Show all posts
Saturday, April 4, 2020
Wednesday, May 29, 2019
Numbers as props vs numbers as information
Once you learn to tell the difference, you'll know whom to trust.
A very long time ago, in 2018, Elon Musk announced that Tesla would be ramping up production to 6000 vehicles per week. An anchor for a business program played the video, then addressed their co-host with:
"That's like four full parking structures a week. Wow!"That statement is true for parking structures that have 1500 spots, which most in San Francisco (where the show is produced) don't. Typical numbers here are closer to 500 than 1500. But that's not the important part.
Co-host makes assenting noises.
The important part is that the number was used as a prop, not information.
More precisely, the anchor first bought into the idea that 6000 is a large number for a car company's weekly production, then looked for a way to make that number look big to the show's audience; parking structures are big buildings and are related to cars, so that was a good way to create the perception of "bigness." [1]
In other words, the process of using a number as a prop is:
1. Make a decision based on something other than the number
2. Look for a number to support that decision
3. Choose context to present the number that molds perception in favor of the decision.
The alternative to using numbers as props is using them as information.
The metric '6000 per week' is just data. It becomes information when it answers a question. A few of these questions that come to mind, considering that this is a business program focussing on technology for a mostly finance and finance-adjacent audience would be:
a. How does this production level compare to that of the competitors that Musk repeatedly states he's going to put out of business?
b. How does this production level compare to that of Toyota when it was running the factory that is now Tesla's?
c. How does this production level compare to the demand for electric vehicles in general, possibly by geographical area and brand of vehicle?
Note that these questions extract information from the number 6000, by comparing it to other numbers that are of business interest. This illustrates a very important principle of data-processing for decision-making:
What is informative about data depends on what decision is to be made.
Choosing question a for illustration, and using Wikipedia data for 2016, because it's publicly available so anyone can check this computation without having to pay financial information service fees, here are the production rates for the top 15 car companies by number of vehicles produced:
Those numbers put Tesla's production in context; they suggest that Tesla, relative to the competitors that Musk repeatedly taunts as "dinosaurs" and "on their way out," is a niche player and not a serious business threat. [2]
Note the process for using numbers as information:
1. Determine what decisions are to be informed by the number
2. Find the context that is relevant for that decision
3. Compare number with the numbers from that context
Using numbers as information is important primarily for decision-makers. Realizing when others are using numbers as props, not information, is important for everyone. Especially regarding whether you can trust the numbers -- and the other person.
Just because someone uses numbers as props, that doesn't necessarily mean their intent is to deceive you. Our society, particularly our news and edutainment, are full of prop-use of numbers for non-nefarious reasons: ignorance, desire to connect abstract numbers to concrete objects, laziness.
But there are people whose intent is to deceive, and often you can tell who they are by calling them on their use of numbers as props. [3]
When faced with the above table, many Tesla fans on twitter, some of whom manage third-party money, either resorted to ad hominem ("how big is your short position?" is a common one, even used by Musk) or changing the subject ("these cars will save the planet").
This is how you identify someone who's not making a good-faith mistake of using numbers as props, but rather someone who deliberately avoids using the appropriate context for the numbers to use them as props: they never address the relevant comparison.
Because most people don't process numbers as they hear or read them, but are still influenced by the perceived authority of the number, this behavior (deliberately using numbers as props to deceive, that is) is usually effective as a persuasion tool. And people who deliberately use numbers as props know about that effectiveness and that's why they do it. Which brings us to an important insight about people we get from their use of numbers:
People who deliberately use numbers as props are not to be trusted.
-- -- -- -- FOOTNOTES -- -- -- --
[1] More likely the choice was made by a writer or a producer, not the anchor; but the anchor is the face of the show, so we'll keep referring to them.
[2] Or, if we want to apply strategic thinking, Tesla should build itself by market expansion starting from its niche, instead of a frontal assault on the much larger companies (its current strategy)
[3] For what it's worth, I don't think the anchor, or the TV channel, were trying to deceive their audience. They were just caught in Musk's Reality Distortion Field, which in 2018 was much stronger than Steve Jobs's ever was.
-- -- -- -- ADDENDUM -- -- -- --
Later that year, numbers-as-props sophistry continued unimpeded by any sense of shame on the part of Tesla fans:
Labels:
decision-making,
Innumeracy,
Tesla
Sunday, January 8, 2017
Numerical thinking - A superpower everyone can get
There are significant advantages to being a numerical thinker. So, why isn't everyone one?
Some people can't be numerical thinkers (or won't be numerical thinkers), typically due to one of three causes:
Acalculia: the inability to do calculations; in its pure form a type of brain damage, but more commonly a consequence of bad educational system.
Innumeracy: lack of mathematical and numerical knowledge, again generally as the result of a bad educational system.
Numerophobia: a fear of numbers and numerical (and mathematical) thinking, possibly an attitude brought on by exposure to the educational system.On a side note, a large part of the problem is the educational system, particularly the way logic and math are covered in it. Just in case that wasn't clear.
Numerical thinkers get a different perspective on the world. It's like a superpower, one that can be developed with practice. (Logical thinkers have a related, but different, superpower.)
Take, for example, this list of large power generating plants, from Wikipedia:
Left to themselves, the numbers on the table are just descriptors, and there's very little that can be said about these plants, other than that there's a quick drop in generation capacity from the first few to the rest.
When numerical thinkers see those numbers, they see the numbers as an invitation to compute; as a way to go beyond the data, to get information out of that data. For example, my first thought was to look at the capacity factors of these power plants: how much power do they really generate as a percentage of their nominal (or "nameplate") power.
Sidenote: Before proceeding, there's an interesting observation I should make here, about operational numerophobia (similar to this older post): in social interactions when this type of problem comes up, educated people who can do calculations in their job, or at least could during their formal education, have trouble knowing where to start to convert a yearly production of 98.8 TWh into a power rating (in MW).
Since this is trivial (divide by the number of hours in one year, 8760, and convert TW to MW by multiplying by one million), the only explanation is yet another case of operational numerophobia. End of sidenote.
Capacity (or load) factor is like any other efficiency measure: how much of the potential is realized? Here are the results for the top 15 or so plants (depending on whether you count the off-line Japanese nuclear plant):
Once these additional numbers are computed, more interesting observations can be made; for example:
The nuclear average capacity factor is $87.7\%$, while the hydro average is just $47.2\%$. That might be partly from use of pumped hydro as storage for surplus energy on the grid (it's the only grid-scale storage available at present; explained in the video below).
That is the power of being a numerical thinker: the ability to go beyond simple numbers and have a deeper understanding of reality. It's within most people's reach to become a numerical thinker, all that's necessary is the will to do so and a little practice.
Alas, many people prefer the easier route of being numerical-poseurs...
A lot of people I interact with pepper their discussions with numbers and even charts, but they aren't numerical thinkers. The numbers and the charts are props, mostly, like the raw numbers on the Wikipedia table. It's only when those numbers are combined among themselves and with outside data (none in this example), information (the use of pumped hydro as grid-level storage), and knowledge (nameplate vs effective capacity, capacity factors) that they realize their potential for informativeness.
A numerical thinker can always spot a numerical-poseur. It's in what they don't do.
- - - -
Bonus content: Don Sadoway talking about electricity storage and liquid metal batteries:
Labels:
Innumeracy,
math,
Numbers,
thinking
Subscribe to:
Posts (Atom)







