Showing posts with label math. Show all posts
Showing posts with label math. Show all posts

Thursday, December 31, 2020

A common misconception about scale effects (for example, in production functions)

There's a common misperception that if a function (usually a production function or a consumer valuation function) scales proportionally (meaning that when all inputs double, for example, the output doubles), then that function must be a linear function.

Sadly, some of the people falling for this error make important decisions in market research and in capacity planning, two areas where this kind of behavior (scaling proportionally) happens a lot and where the error in considering only linear models may have serious consequences for the bottom line.

(And we should always strive to make fewer math errors, of course.)

Let's start with the simple part, using a two-variable function:

\[  f(x,y) = a \, x + b \, y \]

If we scale $x$ and $y$ by a constant $c$ we get

\[ f(cx,cy) = a \, cx + b \, cy = c (a \, x + b \, y) = c \, f(x,y).\]

Clearly, linear functions scale proportionately. What about the other part, the ``must be a linear function'' part?

That's wrong. And we need no more than an example to show it. Voilá:

\[ f(x,y) = x^\alpha \, y^{(1-\alpha)}\]

for an $\alpha \in (0,1)$.

That function is not linear at all; here's a plot for $\alpha = 0.5$ and $x,y$ each in $[0,10]$ (that makes the $z = f(x,y)$ variable also in $[0,10]$, obviously):

And yet,

\[ f(cx,cy) = (cx)^\alpha \, (cy)^{(1-\alpha)} = c^{\alpha + (1-\alpha)} \, x^\alpha \, y^{(1-\alpha)} = c  f(x,y). \]

So, now we know not to limit ourselves to linear functions when describing systems that exhibit proportional scaling.

Nerd note: a function that scales proportionally is called ``homothetic'' or ``homogeneous of degree one.''

Tuesday, October 20, 2020

Pomposity!

Let $f(x)$, $f \in \mathrm{C}^{\infty}$, be the following infinitely continuously differentiable function over the space of real numbers:
\[ 
f(x) \doteq 
\sum_{n=0}^{\infty} \frac{e^{-2} \, 2^n}{n!}
+ \frac{1}{\sqrt{2 \, \pi}}\int_{-\infty}^{+ \infty} x  \, \exp(-y^2/2) \, dy;
\]
then, applying Taylor's theorem and the Newton–Leibniz axiom,
\[f(1) = 2.\]
Time out! What the Heck?!?!

Okay. Breathe.

Let's restate the above in non-pompous terms.

Let $f(x)$ be the following function
\[f(x) = 1 + x\]
then $f(1) = 2$.

All the words between "following" and "function" in the first paragraph mean "smooth," which this function certainly is; $f \in \mathrm{C}^{\infty}$ is the formal way to say all the words in that sentence, so it's redundant. 
 
As for the complicated formula, it uses a series and an integral that each compute to one. Eagle-eyed readers will notice that the first is the Taylor series expansion of $e^2$ times the constant $e^{-2}$ and the second is $x$ times the integral of the p.d.f. for the Normal distribution for $y$, which by definition of a probability has to integrate to 1. Taylor's theorem and Newton–Leibniz axiom are used to get the values for the series and the integral from first principles, as is done in first-year mathematical analysis classes, and which no one would ever use in a practical calculation.

I took a trivially simple function and turned it into a complicated, nay, scary formula. With infinite sums, integrals, and theorems. Taylor is relatively unknown, but Newton and Leibniz? Didn't they invent calculus? (Yes.) So my nonsensical formula acquires immense gravitas. Newton! And Leibniz!!

And that's the problem with an increasing number of public intellectuals and technical material.

There are some genuinely complex things out there, and to even understand the problems in some of these complex things one needs serious grounding in the tools of the field. There's no question about that. But there's a lot of deliberate obfuscation of the clear and unnecessary complexification of the simple.

Why? And what can we do about it?


Why does this happen? Because, sadly, it works: many audiences incorrectly judge the competence of a speaker or writer by how hard it is to follow their logic. And many speakers and writers thus create a simulacrum of expertise by using jargon, dropping obscure references and provisos into the text, and avoiding simple, clear examples in favor of complex and hard-to-follow, "rich," examples.

What can we do about it? This is a systemic problem, so individual action will not solve it. But there's one thing we each can do: starve the pompous of the attention and recognition they so crave. In other words, and in a less pompous phrasing, when we realize someone is purposefully obfuscating the clear and complexifying the simple, we can stop paying attention to them. 

Simplicity actually requires more competence than haphazard complexity; it requires the ability to separate what is essential from what's ancillary. To make things, as Einstein said, as simple as possible, but no simpler.

It's also a good thinking tool for general use. Feynman describes how he used to follow complicated topological proofs by thinking of balls, with hair growing on them, and changing colors:

As they’re telling me the conditions of the theorem, I construct something which fits all the conditions. You know, you have a set (one ball)—disjoint (two balls). Then the balls turn colors, grow hairs, or whatever, in my head as they put more conditions on. Finally they state the theorem, which is some dumb thing about the ball which isn’t true for my hairy green ball thing, so I say, “False!”

If it’s true, they get all excited, and I let them go on for a while. Then I point out my counterexample.

“Oh. We forgot to tell you that it’s Class 2 Hausdorff homomorphic.”

“Well, then,” I say, “It’s trivial! It’s trivial!” By that time I know which way it goes, even though I don’t know what Hausdorff homomorphic means.

Excerpt From: Richard Feynman, “Surely You’re Joking, Mr. Feynman: Adventures of a Curious Character.”

Let's strive to be like Einstein and Feynman.



- - - - -
This post was inspired by an old paper that starts with $1+1=2$ and ends with a multi-line formula, but I've lost the reference; it might have been in the igNobel prizes collection.

Wednesday, February 19, 2020

Seriously, pulling a Mensa card?

Some thoughts on IQ testing, inspired by someone who pulled a "I have a higher IQ than you" on scifi author TJIC (read his books if you like hard scifi: first, second) and then pulled —I kid you not — a Mensa card. An actual Mensa card.


The obvious logical fallacy of implying "I have a high IQ, therefore what I say is right" being evidence of not engaging one's intelligence notwithstanding, there's something funny about claims that IQ, as measured by tests designed for the mass of the population, is somehow a measure of the ability to think about complex or difficult issues.

Note that for mass testing purposes the IQ test as designed is useful, for reasons that will become clear below.

The tests typically consist of a number of simple problems of pattern matching or other low algorithmic complexity and low computational complexity tasks. This works out as a good way to separate people with intelligence from zero up to a standard deviation or two above the mean. In other words, this type of testing separates people who will have serious difficulties, mild difficulties, and no difficulties in following basic education (say up to high school), and people who can do well in education beyond the basic if they choose to.

Because some of the people who do well in these tests go on to do well in situations with high algorithmic and/or computational complexity, IQ metrics (or proxies thereof like the SAT) are used as one of the tools in selection for jobs or education that include such tasks, such as STEM education and jobs.

(Note that it is possible for someone to do badly in IQ tests and still do well in tasks with high algorithmic and/or computational complexity, though that tends to be unlikely and generally happens due to considerations orthogonal to actual intellectual capabilities.)

Because some of the people who do well in IQ tests don't do well once the algorithmic and/or computational complexity increase, using IQ measures as the sole selection tool would be a bad idea; which is why most recruiters look at school transcripts, relevant achievements (like code on Github, Ramanujan's notebooks, billionaire parents*), and other metrics.

The people who do well in IQ tests but not so well in more complex tasks tend to be the ones who join Mensa, which is why it's so funny anyone would thing showing a Mensa card means anything.

Oh, a small thing, though...

The tasks in these tests, themselves, tend to be, well... there's really no nice way to say this, wrong. Just wrong.

Other than word parallel tests (A : B :: C : ?), which measure vocabulary fluidity above all else, pretty much all the pattern matching tasks in these tests can be coded as "what's the next vector in this sequence of vectors of numbers?" to which anyone with a basic understanding of mathematics would answer "a vector of appropriate dimension with any numbers you want."

For example, consider the following sequence: 1, 1, 2, 3, 5, 8. Which is the next number in the sequence?

It's $e^{\pi^3}$.

Clearly!

That's because that sequence is clearly an enumeration in increasing order of the zeros of the following polynomial:

$(x-1)^2 \, (x-2) \, (x-3) \, (x-5) \, (x-8) \, (x - e^{\pi^3})$.

How about the sequence 1, 1, 1, 1, 1, what number comes next?

Clearly the next number is 5. This is the well-known Cinconacci sequence (five ones followed by the sum of the previous five numbers), after the Tribonacci (three ones followed by the sum of the previous three numbers) and Fibonacci (two ones followed by the sum of the previous two numbers) sequences. The Quatronacci sequence is left as an exercise to the reader.

By the way, only people with limited imagination think that the sequence 1, 1, 2, 3, 5, 8 above could only be the beginning of the Fibonacci sequence. That's a cultural bias towards a specific sequence out of an infinity of possible sequences. (A big infinity, at that, $\aleph_1$.)

Note that a smart test-taker will realize that in the infinity of sequences there are some sequences the people who write the test believe are the "right" ones, so a smart test-taker will choose those, thus using both the ability to recognize patterns and the perception of test makers' intellectual limitations.

This is not to say that the tests are useless per se; beyond being good at the separation of the lower levels of thinking ability, they measure the ability to follow instructions and to concentrate on a task for some time, both of which are important as they measure brain executive function.**

But, as I've told many a recruiter (as a consultant to the recruiter, not as a candidate), if you want to know whether a candidate can write code or solve math problems, don't bother them with puzzles; give them a coding task or a math problem.

Alas, puzzle interviews have grown to mythological status, so they're here to stay.


- - - - -

* Or as they call them at Hahvahd admissions, high-potential donors.

** There's an old recruitment test, no longer used, with an instruction sheet and a worksheet. The top of the instruction sheet said in large type "read through all the instructions before beginning," and proceded in regular type to instructions like "1 - draw a line in the worksheet, diagonally from top left to bottom right," and say, another 9 like this; at the bottom of the page it said "turn page to continue" and on the back it said, again in large type, "don't follow instructions 1-10; just write your name in the center of the worksheet and hand that in."

A significant number of people failed the test by doing tasks 1-10 as they read them, ignoring the "read all instructions before beginning" command at the top. This test is no longer used because (a) it's too well-known and (b) people who fail it never want to accept that it's their fault for not following the main instruction to read all instructions before beginning.



AFTERTHOUGHT:

My IQ, you ask? I'm pretty sure, say with 99% probability, that it falls somewhere between 50 and 500. On a good day, of course.


Wednesday, January 15, 2020

Fun with numbers for January 15, 2020

Infinites make for math fun



Link: https://twitter.com/3blue1brown/status/1215087264792887296

I did as GS suggested and figured out the solution myself. (That's the point of following recreational math accounts, after all.)

It wasn't difficult, since the behavior of a function $ f_n(x) = x^{x^{x\ldots}}$ where there are $n-1$ exponentiations, when $n \rightarrow \infty$ is likely to be divergent for any $x>1$. This was my intuition, and based on that intuition, I assumed there couldn't exist $x$ and $y$ such that $f_{\infty}(x) = 2$ and $f_{\infty}(y) = 4$. Therefore the premise is false and the reasoning fails because of that.

But I didn't prove it. I did play around with a spreadsheet to check the behavior of the function around $x=1$, since $f_n(1) =1$ for all $n$ and the functions are continuous in $x$. A simple spreadsheet shows the behaviors for $x$ between 0 and 1 and for $x$ above 1. The rows are increasing $n$, the relevant part is the diff-in-diff column, a discrete version of a second derivative w.r.t. $n$:


What these results (and several others, the point of a spreadsheet model being to play around with the numbers, or as a responsible adult would say it "do sensitivity analysis") show is that for $0 < x <  1$ the function is increasing and "concave," therefore most likely converging to a number in the [0,1] interval; for numbers above 1, the function is increasing and "convex," therefore most likely diverging to infinity.

Still not a proof, but confidence is high. (And confirmed later by the rest of the thread.)

True to my origins as a Prolog (and Lisp, occasionally) programmer, I feel compelled to write a formal definition of this function, thusly:

$\qquad f_1(x) = x;$
$\qquad f_n(x) = x^{f_{n-1}(x)} \quad\text{for $n>1$} $

As usual, recursions FTW!




Go beyond one-step thinking to understand executive pay


Scott Manley, whose space videos are among the most informative on YouTube, expressed a common complaint about executive severance pay, on the occasion of Boeing's change of the guard:


My response is the précis of why you have to pay outgoing executives significant severance: it's a signal to the incoming executives, not a reward to the outgoing executives.

(This was particularly obvious in the case of PG&E in California, which paid a lot to its executives in what most people thought scandalous and well-informed people recognized as a move to retain talent under circumstances when most top managers were considering outside options.)

The responses to SM's tweet contained a lot of misconceptions about executive-level, also called C-suite, management, and this one, which was a response to my response illustrates the three most important ones:


(As I try to be more positive, I'll be anonymizing tweets that I'm critiquing.)




Yet another Rotten Tomatoes calculation (Doctor Who)



Given these numbers, it's

5,872,182,639,638,860,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000,000

times more likely that critics and audience use opposite criteria than the same criteria.





Never give up; never surrender!

Friday, October 25, 2019

How many tangerines fit in this room?

How a person answers simple questions can tell a lot about what type of thinker they are.


It's not that you need to know a lot of math to answer this question (it's basic geometry and arithmetic), but rather that people who think quantitatively as part of their day-to-day life can be identified by their attitude towards this question.

There's a big difference between someone who thinks like a quant and someone who can do math on demand, so to speak. Thinking like a quant means that you generally look at the world through the prism of math; that when you're solving a work problem, you're not just applying knowledge from your education, but also something you practice every day. And that practice makes a difference.

 It's like the difference between an athlete (even if amateur) and someone who goes to gym class.

To illustrate, consider your typical "lone inventor can upset entire industry" story, in particular this one that was in the last Fun With Numbers.
I didn't read the article, but from the photo [which is deceptive, in the article the 1500-mile battery is bigger, though still small enough to make the result non-credible] we can see that the '1500-mile battery' volume is about 2 liters, so a little bit of arithmetic ensued: 
  1. 1500 miles w/ better-than-current vehicles [a google search shows that they're all over 250 Wh/mi], say 200 Wh/mi: 300 kWh (1.08 GJ)
  2. Volume of battery, from article photo [estimated by eye], let's say 2 l, so energy density = 504 MJ/l
  3. Current Li-Ion battery energy density [google search] ~2.5 MJ/l to  5 MJ/l (experimental) 
Home inventor creates something 100 to 200 times more dense than
current technology (and about 15 times more energy-dense than gasoline)?! Not credible.
Are we to believe that the journalists can't do the simple search and arithmetic needed to raise the concerns we can see? Or that they expect none of their audience to? (This second question assuming that the journalists know that the battery can't work, but are willing to write these clickbait headlines because they assume their credibility is not going to be questioned by innumerate audiences.)

Back to the tangerines, and a tale of three people.

Person one gets confused by the question, takes a while to think in qualitative terms (sometimes verbalizing those), then eventually realizes it's a geometry question and with more or less celerity solves it. Person one can do math "on demand," but doesn't think like a quant.

Person two grasps the geometric nature of the problem immediately, estimates the size of the room and of an average tangerine, reaches for a calculator, and gives an estimate. Person two "groks" the problem and is a quant thinker.

Person three sketches out the same calculation as person two, but then adds a twist: instead of a calculator, person three reaches for a spreadsheet, to create a model where the parameters can be varied to allow for sensitivity analysis. Person three is an advanced version of a quant thinker, a model-based thinker.




Monday, September 9, 2019

Fun with numbers for Sep 9, 2019


So, this might become a thing, blogging augmented tweets.

Rotten Tomatoes is at it again



Dave Chappelle apparently has a Netflix special that has critics and audience at loggerheads. Being a little more quantitative, we can say that the critics are 862,712 times more likely to be using criteria opposite to those of the audience than the same criteria. The logic is in this post.

There are two differences between that blog post and this calculation that are worth mentioning:

1. Computing $c(12352,124)$ without loading special numerical packages that can handle large numbers is beyond the capabilities of most mathematical software, so we use a trick: as we're only interested in a likelihood ratio, and those combinations appear in the numerator and denominator, we know that in the end they'll cancel out, so we ignore them altogether.

2. Small probabilities raised to a large exponent quickly get to the precision limits of the floating point representations; to deal with that we make our calculations in log-space. So instead of computing $0.01^{124}$, which would be well below the 1E-99 ($10^{-99}$) limit for most numerical software, and be treated as zero, we compute $124 \times \log(0.01)$, do all the operations in this log space and at the end we exponentiate the result.

Used Apple Numbers (in lieu of RStudio) for this one, was surprised to learn that LOG is $\log_{10}(\cdot)$ despite Numbers also having a LOG10 function. Oh, well, no problem as long as one is careful:




Tesla bull tries to praise superchargers, arithmetic and hilarity ensue


Sooooo I tried the 250kW charger for the first time and I think I'm in loooooove 🥰 -- Went from 19% to 60% I kid you not in just 5 mins. Thanks @elonmusk @Tesla 🙏

Your battery capacity is 51 kWh?!
5 minutes = 300 s
250 kW * 300 s = 75 MJ
75 MJ = (60%-19%) * Capacity, or
Capacity = 183 MJ = 51 kWh
I thought TSLA batteries started at 75 kWh?! 🤔 Possible explanations:

1. Tesla bull is exaggerating; if it took 10 minutes, or the starting point was close to 40%, that would point to a 100 kWh battery.

2. Tesla software is lying to the car owner, making the numbers look rosier than they actually are.

3. Battery has lost capacity, which happens to batteries because of the underlying principles (two main chemical reactions, one exoelectric, one endoelectric; but secondary, parasitical reactions exist that lower the battery capacity over time). Even for a 75 kWh battery that would be a very big loss (1/3), unless his charging cycles are deep and irregular (that kills batteries faster).



Counting calories is like Enron accounting


Just thinking logically, here, if you were bailing water out of a boat, would you keep adding water in?

Okay, so if you're trying to lose weight by using body fat for energy, why would you eat carbs, whose sole nutrition value is as energy? Why eat when not hungry? *

(Most of the arguments I have about calories are with people who for some reason want others to eat carbs. Counting calories biases you towards choosing carbs over fat, since fat is more energy-dense.)

But more to the point, the whole foundation of calorie counting is Enron-like accounting, where some things are counted (more or less), some things are estimated, and many other things are sort-of, kind-of assumed away in some "basic metabolic energy needs" or other ways of saying "let's assume everyone has the same basic efficiency in chemical energy extraction and mechanical power production."


The lack of accounting for energy lost as heat, which anyone out of shape who's ever jogged with an athletic friend can tell you varies a lot with the person, is the most obvious Enron-like accounting.  Higher body effort for the same mechanical output is reflected in heat loss, and differences in that heat loss can be (as calculated in that figure) in the 100-200 kCal/hour range.

That's the same difference as the mechanical energy difference between jogging and walking.  One hour at the low end of that difference (remember, this is just heat, the mechanical energy is the same for both people) every other day is equivalent to 2 kg of fat extra per year if we believe in the basic model of calories-in calories-out.

Now,  how much difference can there be in unmeasured chemical energy output? Depends on the person and the diet, but note that on a 2500 kCal/day diet a systemic difference of 2%, that is 50 kCal/day, is equivalent to 2 kg of fat extra per year if we believe in the basic model of calories-in calories-out.

Can different people with the same general diet show a 2% difference? Yep. For example, a paper called "Energy content of stools in normal healthy controls and patients with cystic fibrosis," by Murphy, Wootton, Bond, and Jackson in Archives of Disease in Childhood (1991) [thanks PubMed], includes data about the controls' intake and stool. Here are the computations for the first 5 healthy controls:



Yeah, just like that, if CICO were true, these people, on the exact same diet, would show a 3 kg per year weight gain difference. 30 kg per decade.

So, whenever people start talking about calories, be aware that they might be looking for a way to say "you're overweight because of your moral failings; if only you were as virtuous as I am!"

- - - - -
* 1. Carbs are delicious, even addictive. Just be aware of the trade-off: they slow down body fat loss and make you hungrier faster. Because controlling appetite is key to fat loss, that second part is much more damaging than the first. Any "diet" that requires constant attention and self-control is going to fail for normal people with normal lives in normal society: just look around you.

2. There's a situation when I'll eat even though I'm not hungry: if I know I'll become hungry later when no high-protein food will be available and the hunger will be inconvenient or require iron will to avoid eating institutional carbs-and-fat food. Usually this situation can be avoided by taking high-protein foods like Biltong (no, it's not jerky; yes, it's worth the price) or hard-boiled eggs with you, but there are situations when that's socially unacceptable.



Nerding out with science fiction




The book is Dream of the Iron Dragon by Robert Kroese. Highly recommended science fiction.

At $c/3$, each kg of mass in the ship has kinetic energy of 607 TJ, the equivalent of a large tactical nuclear weapon (145 kiloton TNT), or about nine times the Hiroshima explosion. The relativistic increase in mass in small (around 6%, of course), but that velocity-squared, that's the big deal. (At these speeds we have to use the relativistic formula for KE, the one with $mc^2$ in the numerator.)

A table of temporal dilation (it's a highly non-linear transformation):


At 99.95% of the speed of the light, one hour of ship time would be 31 hours, 37 minutes, and 48 seconds in the resting frame. At that speed, each kilogram of mass in the ship would have 2.75 exajoule of kinetic energy or, in big boom terms, about 13 times the energy of the largest hydrogen bomb explosion (the Tsar Bomba at 210 PJ or 50 MtTNT).

Another excerpt of the same book, non-numeric, but very dear to anyone who's ever worked in a large bureaucratic organization:


#NerdWhoMe

Monday, June 17, 2019

Calculating God?

I don't believe that the God of any earthly religion is the creator of the Universe.

But I really dislike a lazy and innumerate argument commonly used to "prove" the non-existence of God, which can be summarized in the following false dichotomy:
Either there is no God and the universe just 'poofed' into existence, or there's an infinite number of Gods, because the plane of existence for each God has to be created by a higher-level God.
This is a false dichotomy: it could well be that our universe was created by a powerful being from a higher-order universe, but that universe poofed into existence without a creator. Or maybe it did have a creator, whose universe poofed into existence; or that third universe may have had a creator...

Hey, this looks like dynamic programming. I know dynamic programming.

Let's say that universes are recursively nested until one of them just poofs into existence. Of course we can't see outside our universe, but we can build simple models.

So, our universe either poofed into existence (say with probability $p$) or it was created by some higher being (with probability $1-p$). Now we iterate the process: 'level 2' universe either poofed into existence (with some probability $q$) or was created by a 'level 3' universe being (with probability $1-q$); and so on.

Time for a simplifying assumption, or as non-mathematicians call it, making things up. Let's assume that all these universes share the poofed/created probabilities, so that for any 'level $k$' universe, it poofed into existence with probability $p$ and was created by a being from a 'level $k+1$' universe with probability $1-p$.

Note that it's still possible to have an infinite number of universes, but with this formulation, the probability of a 'level $k$' universe (with us being 'level 1') being the last level is

$p (1-p)^{k-1}$.

This probability gets small pretty quickly, which suggests the 'infinite regress of universes' argument gets thin very fast.



Now we can compute the expected number of universes as a function of $p$:

$\mathbb{E}(n) = p + 2(1-p)p + 3 (1-p)^2 p + \ldots N (1-p)^{N-1} p + \ldots$

or

$\mathbb{E}(n) = p/(1-p) \times ( \text{ sum of series } N (1-p)^N )$

The sum of series $N (1-p)^N$ is $(1-p)/p^2$, so

$E(n) = 1/p$

Therefore, if we believe that the probability of a universe poofing into existence is 0.1, there are an expected ten universes; for 0.2, five universes; for 0.5, two universes.

Very far from 'turtles all the way down.'

Of course, these calculations were unnecessary, because as we know from the revelations of the prophet Terry Pratchett, it's four elephants on the back of the Great A'Tuin swimming in the Sea of Stars.

Wednesday, June 12, 2019

A statistical analysis of reviews of L.A. Finest: audience vs. critics



"If numbers are available, let's use the numbers. If all we have are opinions, let's go with mine." -- variously attributed to a number of bosses.

There's a new police procedural this season, L.A. Finest, and Rotten Tomatoes has done it again: critics and audience appear to be at loggerheads. Like with The Orville, Star Trek Discovery, and the last season of Doctor Who.

But "appear to be" is a dequantified statement. And Rotten Tomatoes has numbers; so, what can these numbers tell us?

Before they can tell us anything, we need to write our question: first in words, then as a math problem. Then we can solve the math problem and that solution gets translated into a "words" answer, but now a quantified "words" answer.

The question, which is suggested by the above numbers is:
Do the critics and the audience use similar or opposite criteria to rate this show?
One way to answer this question, which would have been feasible in the past when Rotten Tomatoes had user reviews, would be to do text analytics on the reviews themselves. But now the user reviews are gone so that's no longer possible.

Another way, a simpler and cleaner way, is to use the data above.

To simplify we'll assume that all ratings are either positive or negative, 0 or 1; there are some unobservable random factors that make some people like a show more or less, so these ratings are random variables. For a given person $i$, the probability that that person likes L.A. Finest is captured in some parameter $\theta_i$ (we don't observe that, of course), which is the probability of that person giving a positive rating.

So, our question above is whether the $\theta_i$ of the critics and the $\theta_i$ of the audience are the same or "opposed." And what is "opposed"? If $i$ and $j$ use opposite criteria, the probability that $i$ gives a 1 is the probability that $j$ gives a 0, so $\theta_i = 1-\theta_j$.

We don't have the individual parameters $\theta_i$ but we can simplify again by assuming that all variation within each group (critics or audience) is random, so we really only need two $\theta$.

We are comparing two situations, call them: hypothesis zero, $H_0$, meaning the critics and the audience use the same criteria, that is they have the same $\theta$, call it $\theta_0$; and hypothesis one, $H_1$, meaning the critics use criteria opposite to those of the audience, so if the critics $\theta$ is $\theta_1$, the audience $\theta$ is $(1-\theta_1)$.

Yes, I know, we don't have $\theta_0$ or $\theta_1$. We'll get there.

Our "words" question now becomes the following math problem: how much more likely is it that the data we observe is created by $H_1$ versus created by $H_0$, or in a formula: what is the likelihood ratio

$LR = \frac{\Pr(\mathrm{Data}| H_1)}{\Pr(\mathrm{Data}| H_0)} $?

Observation: This is different from the usual statistics test: the usual test is whether the two distributions are different; we are testing for a specific type of difference, opposition. So there are in fact three states of the world: same, opposite, and different but not opposite; we want to compare the likelihood of the first two. If same is much more likely than opposite, then we conclude 'same.' If opposite is much more likely than same, we conclude 'opposite.' If same and opposite have similar likelihoods (for some notion of 'similar' we'd have to investigate), then we conclude 'different but not opposite.'

Our data is four numbers: number of critics $N_C = 10$, number of positive reviews by critics $k_C = 1$, number of audience members $N_A = 40$, number of positive reviews by audience members $k_A = 30$.

But what about the $\theta_0$ and $\theta_1$?

This is where the lofty field of mathematics gives way to the down and dirty world of estimation. We estimate $\theta$ by maximum likelihood, and the maximum likelihood estimator for the probability of a positive outcome of a binary random variable (called a Bernoulli variable) is the sample mean.

Yep, all those words to say "use the share of 1s as the $\theta$."

Not so fast. True, for $H_0$, we use the share of ones

$\theta_0 = (k_C + k_A)/(N_C + N_A) = 31/50 = 0.62$;

but for $H_1$, we need to address the audience's $1-\theta_1$ by reverse coding the zeros and ones, in other words,

$\theta_1 = (k_C + (N_A - k_A))/(N_C + N_A) = 11/50 = 0.22$.

Yes, those two fractions are "estimation." Maximum likelihood estimation, at that.

Now that we are done with the dirty statistics, we come back to the shiny world of math, by using our estimates to solve the math problem. That requires a small bit of combinatorics and probability theory, all in a single sentence:

If each individual data point is an independent and identically distributed Bernoulli variable, the sum of these data points follows the binomial distribution.

Therefore the desired probabilities, which are joint probabilities of two binomial distributions, one for the critics, one for the audience, are

$\Pr(\mathrm{Data}| H_0) = c(N_C,k_C) (\theta_0)^{k_C} (1- \theta_0)^{N_C- k_C} \times c(N_A,k_A) (\theta_0)^{k_A} (1- \theta_0)^{N_A- k_A}$

and

$\Pr(\mathrm{Data}| H_1) = c(N_C,k_C) (\theta_1)^{k_C} (1- \theta_1)^{N_C- k_C} \times c(N_A,k_A) (1 -\theta_1)^{k_A} (\theta_1)^{N_A- k_A}$.

Replacing the symbols with the estimates and the data we get

$\Pr(\mathrm{Data}| H_0) = 3.222\times 10^{-5}$;
$\Pr(\mathrm{Data}| H_1) = 3.066\times 10^{-2}$.

We can now compute the likelihood ratio,

$LR = \frac{\Pr(\mathrm{Data}| H_1)}{\Pr(\mathrm{Data}| H_0)} = 915$,

and translate that into words to make the statement
It's 915 times more likely that critics are using criteria opposite to those of the audience than the same criteria.
Isn't that a lot more satisfying than saying they "appear to be at loggerheads"?

Wednesday, March 22, 2017

The power of "equations"

If a picture is worth a thousand words, an equation is worth a thousand pages of text.

This was inspired by a livestream about free trade based on criticism of "original texts." (Basically Ricardo and Schumpeter.) The quotes aren't a diss on the texts themselves, but rather a way to emphasize that this is a type of scholarly pursuit in itself, though not the type used in modern economics, STEM, or pragmatic professional fields like business analytics or medicine.

What's the problem with the argumentation from these original texts? Simply put, the texts are long and convoluted, with many unnecessary diversions and some logical problems in the presentation. The valid arguments in these texts can be condensed in about one page of stated assumptions and two results about specialization.

It's not just that math's an efficient way to communicate, math has precise meaning and an inference process. It brings discipline and clarity to the texts and the inference process isn't open to debate. (Checks and corrections, yes; debate, no.)

Unfortunately, without math, the speaker's argument was essentially a sequence of variations on "Schumpeter points out that this assumption of Ricardo doesn't hold true," without the extra step of determining whether those assumptions are important to the final result or not. (We'll come back to this problem.)

Word-thinking about quantitative fields is generally to be avoided.

That was the inspiration, and this post isn't about free trade or the particular mode of thought of that speaker, but rather about the power of mathematical modeling, which I'm calling "equations" in the title.

Here's a reasonably robust statement: when the price of a commodity goes up, people buy less of that commodity. (Sometimes this is put as "demand goes down," which is incorrect, it's the demand quantity that goes down. Changes in demand are movements of an entire function.)

So, quantity is a decreasing function of price (and first-time readers of economics textbooks get confused because the charts have quantity in the $x$ axis and price in the $y$ axis). This has been known for a long time; what's the problem with that formulation, simplified to "when price rises, quantity falls"?

The problem, of course, is that there are many different types of decreasing function. Here are a few, for example (click for bigger):


Functions 1 to 4 represent four common behaviors of decreasing functions: the linear function has similar changes leading to similar effects; the convex function has decreasing effect of similar change (like most natural decay processes); the concave function has increasing effect of similar change (like the accelerating effect of a bank run on bank reserves); and the s-shaped function shows up in many diffusion processes (and is a commonly used price response function in marketing).

Functions 5 to 8 are variations on the convex function, showing increasing curvature. (Function 2 would fit between 5 and 6.) They're here to make the point that even knowing the general shape isn't enough: one must know the parameters of that shape.

That figure does have 2000 data points, since each function has 250 points plotted. (When talking about math, some people use drawing tools to make their "functions," I prefer to plot them from the mathematical formula; it's a habit of mine, not lying to the audience.) To describe them in text would take a long time (unless the text is a description of mathematical formulation), while they can be written simply as formulas; for example, the convex functions are all exponentials:

$\qquad y = 100 \, \exp(-\kappa \, x) $

with different values of $\kappa$. They are the type of exponential decay found in many processes, for example, where $x$ is time and $y(x) = \alpha \, y(x-1)$ with $y(0)>0$ models a process of decay with discrete-time rate $0 < \alpha < 1$. In case it's not obvious, $\kappa = -\log_{e}(\alpha)$.*

So, what does this have to do with reasoning?

Here we go back to the problem with arguments like "Schumpeter showed that Ricardo's assumption X was wrong." When a model is written out in equations, we have a sequence of steps leading to the result, each step tagged with either a know result, rules of math inference (say "$a \times b = a \times c$ simplifies to $b = c$ unless $a = 0$"), or an assumption of the model. This allows a reader to quickly see where a failed assumption will lead to problems and determine whether the assumption can be replaced with something true (or, as is the case with many of the assumptions made by Ricardo, is unnecessary for the result).

The main power, however, is that mathematical notation forces the speaker to be precise, and inferences from mathematical models can be checked independently of subject matter expertise. A mathematician may not understand any of the economics involved, but will merrily check that a decay process of the kind $y(n)= \alpha \, y(n-1)$ can be described by an equation $y(n) = y(0) \, \exp(-\kappa \, n)$ and determine the relationship between $\kappa$ and $\alpha$.

From those precise models, one can make inferences that take into account details hidden by language. Consider the "price rises, quantity falls" text and compare it with the different decreasing functions in the figure above. The shape of the function, its slope and its curvature have different implications for how price changes affect a market, differences that are lost in the "price rises, quantity falls" formulation.

It bears repeating the first mentioned advantage: that hundreds of pages can be condensed in one page of equations. Once one's mind is used to processing equations, this is a very efficient way to learn new things. Stories about Port wineries in Portugal and textile factories in England may be entertaining, but they aren't necessary to understand specialization (which is what comparative advantage really is).

Math. It's a superpower mostly anyone can acquire. Sadly, most opt not to.


- - - - - Addendum - - - - -

No self-respecting economist would use the Ricardo comparative advantage argument for international trade now, particularly because it's so simple it can be understood by anyone. Most likely they'd use some variation of the magic factory example:

"Let's say a new technology that converts corn into cars is discovered and a factory is built in Iowa that can take ~ $\$20,000$ of corn and convert it into a car that costs $\$30,000$ to make in Michigan. Can we agree that this technology makes the US richer?

Now, move the factory to Long Beach, CA. Maybe there's a little more cost in moving the corn there, but we're still making the US richer, right?

Now, someone goes into the magic factory and discovers that it's really a depot: stores grain until it's sent to China on bulk carriers and receives cars made in China from RoRos during the night. The effect is the same as the magic factory, so it makes the US richer, right?"

There are many cons to this example, but it does make one issue clear: trade is in many respects just like a different technology.


- - - - - Footnote - - - - -

* It's obvious to me, because after decades of playing around with mathematical models, I grok most of these simple things. There are some people who mistake this well-developed and highly available knowledge (from practice) for ultra-high intelligence (rather than regular very high intelligence), a mistake I elaborate upon in this post. 😎

Tuesday, March 7, 2017

Deep understanding and problem solving

There's value in deep understanding.

Nope, I don't mean the difference between word thinkers and quantitative thinkers. Been there, done that. Nor the difference between different levels of expertise on technical matters; again, been there, done that.

No, we're talking the crème de la crème, experts that can adapt to changing situations or comprehend complexity across different fields, by being deep understanders.

Because any opportunity to mock those who purport to educate the masses by passing along material they don't understand, let us talk about Igon Values... ahem, eigenvalues and eigenvectors.

Taught in AP math classes or freshman linear algebra, the eigenvectors $\mathbf{x}_{i}$ and associated eigenvalues $\lambda_{i}$ of a square matrix $\mathbf{A}$ are defined as the solutions to $\mathbf{A} \, \mathbf{x}_{i} = \lambda_{i} \, \mathbf{x}_{i}$.

Undergrads learn that these represent something about the structure of the matrix, learn that the matrix can be diagonalized using them, how they appear in other places (principal components analysis and network centrality, for example).

But those who get to use these and other math concepts on a day-to-day basis, who get to really understand them, develop a deeper understanding of the meaning of the concepts. There's something important about how these objects relate to each other.

After a while, one realizes that there are structures and meta-structures that repeat across different problems, even across different fields. Someone said that after a lot of experience in one engineering (say, electrical), adapting to another (say, mechanical) revealed that while the nouns changed, the verbs were very similar.

This is what deep understanding affords: a quasi-intuitive grokking of a field, based on the regularities of knowledge across different fields.

For example: while many who have taken a linear algebra in college may vaguely recall what an eigenvalue is, those who understand the meaning of eigenvalues and eigenvectors for matrices will have a much easier time understanding the eigenfunctions of linear operators:


The structure [something that operates] [something operated upon] = [constant] [something operated upon] is common, and what it means is that the [something operated upon] is in some sense invariant with the [something that operates], other than the proportionality constant. That suggests that there's a hidden meaning or structure to the [something that operates] that can be elicited by studying the [something operated upon].

And this structure, mathematical as it might be, has a lot of applications outside of mathematics (and not just as a mathematical tool for formalizing technical problems). It's a basic principle of undestanding: what is invariant to a transformation tells us something deep about that transformation. (Again, invariant in "direction," so to speak, possibly a change of size or even sign.)

And this is itself a meta-principle: that the study of what changes and what's invariant in a particular set of problems gives some indications about latent structure to that set of problems. That latent structure may be a good point to start when trying to solve problems from this set.

Yep, really dumbing down this blog, pandering to the public...

Friday, February 24, 2017

If it's a math problem... do the math

Or, The Monty Hall problem: redux.

I recently posted a new video, addressing the Monty Hall problem. The problem is not the puzzle itself, which has been solved ad nauseam by everyone and their vlogbrother.


The video is about what information is. By working through the details of the Monty Hall puzzle, we can learn where information is revealed and how. That is the reason for the video; that and a plea for something so simple and yet so ignored that I'll repeat it again:

If it's a math problem, do the math.

Now, this may seem trivial, but math (and to some extent science, technology, and engineering, to say nothing of business, management, and economics) makes people uncomfortable, even people who say they "love math."

Hence the attempt to solve the problem with anything but computation. By waving hands and verbalizing (very error prone) or by creating similar problems that might be insightful (but mostly convince only those who already know the solution and understand it).

If all you're interested is the computations for the solution, they're here:



The point of the video is not this particular table; it's the insights about information on the path to it: how constraints to actions change probabilities and how those relate to information.

For example, from the viewpoint of the contestant, once she picks door 1 (thus giving Monty Hall a choice of door 2 and door 3 to open), the probability that Monty picks either door 2 or door 3 is precisely 1/2; that's calculated in the video, not assumed and not hand-waved. But, as the video then explains, that 50-50 probability isn't equally distributed across different states:



A final remark, from the video as well, is that by having computations one can avoid many time-wasters, who --- not having done any computations themselves and generally having a limited understanding of the whole state-event difference, which is essential to reasoning with conditional probabilities --- are now required to point out where they disagree with the computation, before moving forward with new "ideas."

If it's a math problem... do the math!

Sunday, February 12, 2017

Word Thinkers and the Igon Value Problem

Nassim Nicholas Taleb did it again: "word thinkers," now a synonym for his previous coinage IYI (Intellectuals Yet Idiots).
I often say that a mathematician thinks in numbers, a lawyer in laws, and an idiot thinks in words. These words don’t amount to anything. 
A little unfair, though I've often cringed at the use of technical words by people who don't seem to know the meaning of those words. This sometimes leads to never-ending words-only arguments about things that can be determined in minutes with basic arithmetic or with a spreadsheet.


To not rehash the Heisenberg traffic stop example, here's one from a recent discussion of the putative California secession from the US (and already mentioned in this blog): people discussed California's need for electricity, with the pro-Calexit people assuming that appropriate capacity could be added in a jiffy, while the con-Calexit people assumed the state would instantly be blacked out.

No one thought of actually looking up the numbers and checking out the needs. Using 2015 numbers, California would need to add about 15GW of new dispatchable generation for energy independence, assuming no demand growth. (Computations in this post.) So, that's a lot, but not unsurmountable in, say, a decade with no regulatory interference. Maybe even less time, with newer technologies (yes, all nuclear; call it a French connection).

There was no advanced math in that calculation: literally add and divide. And the data was available online. But the "word thinkers" didn't think about their words as having meaning.

And that's it: the problem is not so much that they think in words, but rather that they don't associate any meaning to the words. They are just words, and all that matters is their aesthetic and signaling value.

Few things exemplify the problem of these words-without-meaning as well as The Igon Value Problem.

In a review of Malcolm Gladwell's collection of essays "What the dog saw and other adventures" for The New York Times, Steven Pinker coined that phrase, picking on a problem of Gladwell that is common to the words-without-meaning thinkers:
An eclectic essayist is necessarily a dilettante, which is not in itself a bad thing. But Gladwell frequently holds forth about statistics and psychology, and his lack of technical grounding in these subjects can be jarring. He provides misleading definitions of “homology,” “sagittal plane” and “power law” and quotes an expert speaking about an “igon value” (that’s eigenvalue, a basic concept in linear algebra). In the spirit of Gladwell, who likes to give portentous names to his aperçus, I will call this the Igon Value Problem: when a writer’s education on a topic consists in interviewing an expert, he is apt to offer generalizations that are banal, obtuse or flat wrong. [Emphasis added]
Educational interlude:
Eigenvalues of a square $[n\times n]$ matrix $M$ are the constants $\lambda_i$ associated with vectors $x_i$ such that $M \, x_i = \lambda_i \, x_i$. In other words, these vectors, called eigenvectors, are along the directions in $n$-dimensional space that are unchanged when operated upon by $M$; the $\lambda_i$ are proportionality constants that show how the vectors stretch in that direction. Because of this $n$-dimensional geometric interpretation, the $x_i$ are the matrix's "own vectors" (in German, eigenvectors) and by association the $\lambda_i$ are the "own values" (in German, you guessed it, eigenvalues). 
Eigenvectors and eigenvalues reveal the deep structure of the information content of whatever the matrix represents. For example: if $M$ is a matrix of covariances among statistical variables, the eigenvectors represent the underlying principal components of the variables; if $M$ is an incidence matrix representing network connections, the eigenvector with the highest eigenvalue ranks the centrality of the nodes in the network.
This educational interlude is a demonstration of the use of words (note that there's no actual derivation or computation in it) with deep meaning, in this case mathematical.

Being a purveyor of "generalizations that are banal, obtuse or flat wrong" hasn't harmed Gladwell; in fact, his success has spawned a cottage industry of what Taleb is calling word-thinkers, which apparently are now facing an impending rebellion.

Taleb talks about 'skin in the game,' which is a way to say, having an outside validator: not popularity, not social signaling; money, physical results, a verifiable mathematical proof. All of these come with the one thing word-thinkers avoid:

A clear succeed/fail criterion.

- - - - - - - - - -

Added 2/16/2017: An example of word-thinking over quantitative matters.

From a discussion about Twitter, motivated by their filtering policies:
Person A: "I wonder how long Twitter can burn money, billions/yr.  Who is funding this nonsense?"
My response: "Actually, from latest available financials, TWTR had a $\$ 77$ million positive cash flow last year. Even if its revenue were to dry up, the operational cash outflow is only $\$ 220$ million/year; with a $\$ 3.8$ billion cash-in-hand reserve, it can last around 17 years at zero inflow."
Numbers are easy to obtain and the only necessary computation is a division. But Person A didn't bother to (a) look up the TWTR financials, (b) search for the appropriate entries, and (c) do a simple computation.

That's the problem with word thinking about quantitative matters: those who take the extra quant step will always have the advantage. As far as truth and logic are concerned, of course.

Thursday, February 2, 2017

Primal entertainment

Really, totally primal. 😉

Ron Rivest talking about RSA-129 (a product of two prime numbers that was set as a factoring challenge in 1977) and its factorization in 1994 using the internet:



RSA-129 = 114381625 7578888676 6923577997 6146612010 2182967212 4236256256 1842935706 9352457338 9783059712 3563958705 0589890751 4759929002 6879543541
=
3490 5295108476 5094914784 9619903898 1334177646 3849338784 3990820577
$\times$ 
32769 1329932667 0954996198 8190834461 4131776429 6799294253 9798288533.

Inspired by that video, here are a couple of fun numbers, for numbers geeks:

😎 70,000,000,000,000,000,000,003 is a prime number. It's an interesting prime number, because the number of zeros in the middle (21) is the product of the 7 and the 3, both of which are, of course, prime numbers themselves. This makes the number very easy to memorize and surprise your friends with. If you want to confuse them, just say it like this: "seventy sextillion and three."

😎 99,999,999,999,999,999,999,977 is also a prime number, the largest prime number under a googol ($10^{100}$) that has the form  $p = 10^{n} - n$, with $n = 23$, meaning that if you add 23 to this number you get $10^{23}$ or a 1 followed by 23 zeros. Here's how you say this number: "ninety-nine sextillion, nine hundred ninety-nine quintillion, nine hundred ninety-nine quadrillion, nine hundred ninety-nine trillion, nine hundred ninety-nine billion, nine hundred ninety-nine million, nine hundred ninety-nine thousand, and nine hundred seventy-seven." Hilarious at parties.

Friday, January 13, 2017

Medical tests and probabilities

You may have heard this one, but bear with me.

Let's say you get tested for a condition that affects ten percent of the population and the test is positive. The doctor says that the test is ninety percent accurate (presumably in both directions). How likely is it that you really have the condition?

[Think, think, think.]

Most people, including most doctors themselves, say something close to $90\%$; they might shade that number down a little, say to $80\%$, because they understand that "the base rate is important."

Yes, it is. That's why one must do computation rather than fall prey to anchor-and-adjustment biases.

Here's the computation for the example above (click for bigger):


One-half. That's the probability that you have the condition given the positive test result.

We can get a little more general: if the base rate is $\Pr(\text{sick}) = p$ and the accuracy (assumed symmetric) of the test is $\Pr(\text{positive}|\text{sick}) = \Pr(\text{negative}|\text{not sick})  = r $, then the probability of being sick given a positive test result is

\[ \Pr(\text{sick}|\text{positive}) = \frac{p \times r}{p \times r + (1- p) \times (1-r)}. \]

The following table shows that probability for a variety of base rates and test accuracies (again, assuming that the test is symmetric, that is the probability of a false positive and a false negative are the same; more about that below).


A quick perusal of this table shows some interesting things, such as the really low probabilities, even with very accurate tests, for the very small base rates (so, if you get a positive result for a very rare disease, don't fret too much, do the follow-up).


There are many philosophical objections to all the above, but as a good engineer I'll ignore them all and go straight to the interesting questions that people ask about that table, for example, how the accuracy or precision of the test works.

Let's say you have a test of some sort, cholesterol, blood pressure, etc; it produces some output variable that we'll assume is continuous. Then, there will be a distribution of these values for people who are healthy and, if the test is of any use, a different distribution for people who are sick. The scale is the same, but, for example, healthy people have, let's say, blood pressure values centered around 110 over 80, while sick people have blood pressure values centered around 140 over 100.

So, depending on the variables measured, the type of technology available, the combination of variables, one can have more or less overlap between the distributions of the test variable for healthy and sick people.

Assuming for illustration normal distributions with equal variance, here are two different tests, the second one being more precise than the first one:



Note that these distributions are fixed by the technology, the medical variables, the biochemistry, etc; the two examples above would, for example, be the difference between comparing blood pressures (test 1) and measuring some blood chemical that is more closely associated with the medical condition (test 2), not some statistical magic made on the same variable.

Note that there are other ways that a test A can be more precise than test B, for example if the variances for A are smaller than for B, even if the means are the same; or if the distributions themselves are asymmetric, with longer tails on the appropriate side (so that the overlap becomes much smaller).

(Note that the use of normal distributions with similar variances above was only for example purposes; most actual tests have significant asymmetries and different variances for the healthy versus sick populations. It's something that people who discover and refine testing technologies rely on to come up with their tests. I'll continue to use the same-variance normals in my examples,  for simplicity.) 


A second question that interested (and interesting) people ask about these numbers is why the tests are symmetric (the probability of a false positive equal to that of a false negative). 

They are symmetric in the examples we use to explain them, since it makes the computation simpler. In reality almost all important preliminary tests have a built-in bias towards the most robust outcome.

For example, many tests for dangerous conditions have a built-in positive bias, since the outcome of a positive preliminary test is more testing (usually followed by relief since the positive was a false positive), while the outcome of a negative can be lack of treatment for an existing condition (if it's a false negative).

To change the test from a symmetric error to a positive bias, all that is necessary is to change the threshold between positive and negative towards the side of the negative:



In fact, if you, the patient, have access to the raw data (you should be able to, at least in the US where doctors treat patients like humans, not NHS cost units), you can see how far off the threshold you are and look up actual distribution tables on the internet. (Don't argue these with your HMO doctor, though, most of them don't understand statistical arguments.)

For illustration, here are the posterior probabilities for a test that has bias $k$ in favor of false positives, understood as $\Pr(\text{positive}|\text{not sick}) = k \times \Pr(\text{negative}|\text{sick})$, for some different base rates $p$ and probability of accurate positive test $r$ (as above):


So, this is good news: if you get a scary positive test for a dangerous medical condition, that test is probably biased towards false positives (because of the scary part) and therefore the probability that you actually have that scary condition is much lower than you'd think, even if you'd been trained in statistical thinking (because that training, for simplicity, almost always uses symmetric tests). Therefore, be a little more relaxed when getting the follow-up test.


There's a third interesting question that people ask when shown the computation above: the probability of someone getting tested to begin with. It's an interesting question because in all these computational examples we assume that the population that gets tested has the same distribution of sick and health people as the general population. But the decision to be tested is usually a function of some reason (mild symptoms, hypochondria, job requirement), so the population of those tested may have a higher incidence of the condition than the general population.

This can be modeled by adding elements to the computation, which makes the computation more cumbersome and detracts from its value to make the point that base rates are very important. But it's a good elaboration and many models used by doctors over-estimate base rates precisely because they miss this probability of being tested. More good news there!


Probabilities: so important to understand, so thoroughly misunderstood.


- - - - -
Production notes

1. There's nothing new above, but I've had to make this argument dozens of times to people and forum dwellers (particularly difficult when they've just received a positive result for some scary condition), so I decided to write a post that I can point people to.

2. [warning: rant]  As someone who has railed against the use of spline drawing and quarter-ellipses in other people's slides, I did the right thing and plotted those normal distributions from the actual normal distribution formula. That's why they don't look like the overly-rounded "normal" distributions in some other people's slides: because these people make their "normals" with free-hand spline drawing and their exponentials with quarter ellipses, That's extremely lazy in an age when any spreadsheet, RStats, Matlab, or Mathematica can easily plot the actual curve. The people I mean know who they are. [end rant]

Sunday, January 8, 2017

Numerical thinking - A superpower everyone can get


There are significant advantages to being a numerical thinker. So, why isn't everyone one?

Some people can't be numerical thinkers (or won't be numerical thinkers), typically due to one of three causes:
Acalculia: the inability to do calculations; in its pure form a type of brain damage, but more commonly a consequence of bad educational system. 
Innumeracy: lack of mathematical and numerical knowledge, again generally as the result of a bad educational system. 
Numerophobia: a fear of numbers and numerical (and mathematical) thinking, possibly an attitude brought on by exposure to the educational system.
On a side note, a large part of the problem is the educational system, particularly the way logic and math are covered in it. Just in case that wasn't clear.

Numerical thinkers get a different perspective on the world. It's like a superpower, one that can be developed with practice. (Logical thinkers have a related, but different, superpower.)

Take, for example, this list of large power generating plants, from Wikipedia:



Left to themselves, the numbers on the table are just descriptors, and there's very little that can be said about these plants, other than that there's a quick drop in generation capacity from the first few to the rest.

When numerical thinkers see those numbers, they see the numbers as an invitation to compute; as a way to go beyond the data, to get information out of that data. For example, my first thought was to look at the capacity factors of these power plants: how much power do they really generate as a percentage of their nominal (or "nameplate") power.

Sidenote: Before proceeding, there's an interesting observation I should make here, about operational numerophobia (similar to this older post): in social interactions when this type of problem comes up, educated people who can do calculations in their job, or at least could during their formal education, have trouble knowing where to start to convert a yearly production of 98.8 TWh into a power rating (in MW). 
Since this is trivial (divide by the number of hours in one year, 8760, and convert TW to MW by multiplying by one million), the only explanation is yet another case of operational numerophobia. End of sidenote.

Capacity (or load) factor is like any other efficiency measure: how much of the potential is realized? Here are the results for the top 15 or so plants (depending on whether you count the off-line Japanese nuclear plant):



Once these additional numbers are computed, more interesting observations can be made; for example:

The nuclear average capacity factor is $87.7\%$, while the hydro average is just $47.2\%$. That might be partly from use of pumped hydro as storage for surplus energy on the grid (it's the only grid-scale storage available at present; explained in the video below).

That is the power of being a numerical thinker: the ability to go beyond simple numbers and have a deeper understanding of reality. It's within most people's reach to become a numerical thinker, all that's necessary is the will to do so and a little practice.

Alas, many people prefer the easier route of being numerical-poseurs...

A lot of people I interact with pepper their discussions with numbers and even charts, but they aren't numerical thinkers. The numbers and the charts are props, mostly, like the raw numbers on the Wikipedia table. It's only when those numbers are combined among themselves and with outside data (none in this example), information (the use of pumped hydro as grid-level storage), and knowledge (nameplate vs effective capacity, capacity factors) that they realize their potential for informativeness.

A numerical thinker can always spot a numerical-poseur. It's in what they don't do.

- - - -

Bonus content: Don Sadoway talking about electricity storage and liquid metal batteries: