Showing posts with label computational math. Show all posts
Showing posts with label computational math. Show all posts

Friday, November 15, 2019

Fun with numbers for November 15, 2019

How many test rigs for a successful product at scale?


From the last Fun with Numbers:


This is a general comment on how new technologies are presented in the media: usually something that is either a laboratory test rig or at best a proof-of-concept technology demonstration is hailed as a revolutionary product ready to take the world and be deployed at scale.

Consider how many is "a lot of," as a function of success probabilities at each stage:


Yep, notwithstanding all good intentions in the world, there's a lot of work to be done behind the scenes before a test rig becomes a product at scale, and many of the candidates are eliminated along the way.



Recreational math: statistics of the maximum draw of N random variables


At the end of a day of mathematical coding, and since Rstudio was already open (it almost always is), I decided to check whether running 1000 iterations versus 10000 iterations of simulated maxima (drawing N samples from a standard distribution and computing the maximum, repeated either 1000 times or 10000 times) makes a difference. (Yes, an elaboration on the third part of this blog post.)

Turns out, not a lot of difference:


Workflow: BBEdit (IMNSHO the best editor for coding) --> RStudio --> Numbers (for pretty tables) --> Keynote (for layout); yes, I'm sure there's an R package that does layouts, but this workflow is WYSIWYG.

The R code is basically two nested for-loops, the built-in functions max and rnorm doing all the heavy lifting.

Added later: since I already had the program parameterized, I decided to run a 100,000 iteration simulation to see what happens. Turns out, almost nothing worth noting:


Adding a couple of extra lines of code, we can iterate over the number of iterations, so for now here's a summary of the preliminary results (to be continued later, possibly):


And a couple of even longer simulations (all for the maximum of 10,000 draws):


Just for fun, the probability (theoretical) of the maximum for a variety of $N$ (powers of ten in this example) is greater than some given $x$ is:




More fun with Solar Roadways


Via EEVblog on twitter, the gift that keeps on giving:


This Solar Roadways installation is in Sandpoint, ID (48°N). Solar Roadways claims its panels can be used to clear the roads by melting the snow… so let's do a little recreational numerical thermodynamics, like one does.

Average solar radiation level for Idaho in November: 3.48 kWh per m$^2$ per day or 145 W/m$^2$ average power. (This is solar radiation, not electrical output. But we'll assume that Solar Roadways has perfectly efficient solar panels, for now.)

Density of fallen snow (lowest estimate, much lower than fresh powder): 50 kg/m$^3$ via the University of British Columbia.

Energy needed to melt 1 cm of snowfall (per m$^2$): 50 [kg/m^3] $\times$ 0.01 [m/cm] $\times$ 334 [kJ/kg] (enthalpy of fusion for water) = 167 kJ/m$^2$ ignoring the energy necessary to raise the temperature, as it's usually much lower than the enthalpy of fusion (at 1 atmosphere and 0°C, the enthalpy of fusion of water is equal to the energy needed to raise the temperature of the resulting liquid water to approximately 80°C).

So, with perfect solar panels and perfect heating elements, in fact with no energy loss anywhere whatsoever, Solar Roadways could deal with a snowfall of 3.1 cm per hour (= 145 $\times$ 3600 / 167,000) as long as the panel and surroundings (and snow) were at 0°C.

Just multiply that 3.1 cm/hr by the efficiency coefficient to get more realistic estimates. Remember that the snow, the panels, and the surroundings have to be at 0°C for these numbers to work. Colder doesn't just make it harder; small changes can make it impossible (because the energy doesn't go into the snow, goes into the surrounding area).



Another week, another Rotten Tomatoes vignette


This time for the movie Midway (the 2019 movie, not the 1972 classic Midway):


Critics and audience are 411,408,053,038,500,000 (411 quadrillion) times more likely to use opposite criteria than same criteria.

Recap of model: each individual has a probability $\theta_i$ of liking the movie/show; we simplify by having only two possible cases, critics and audience using the same $\theta_0$ or critics using a $\theta_1$ and audience using a $\theta_A = 1-\theta_1$. We estimate both cases using the four numbers above (percentages and number of critics and audience members), then compute a likelihood ratio of the probability of those ratings under $\theta_0$ and $\theta_1$. That's where the 411 quadrillion times comes from: the probability of a model using $\theta_1$ generating those four numbers is 411 quadrillion times the probability of a model using $\theta_0$ generating those four numbers. (Numerical note: for accuracy, the computations are made in log-space.)



Google gets fined and YouTubers get new rules


Via EEVBlog's EEVblab #67, we learn that due to non-compliance with COPPA, YouTube got fined 170 million dollars and had to change some rules for content (having to do with children-targeted videos):


Backgrounder from The Verge here; or directly from the FTC: "Google and YouTube Will Pay Record $170 Million for Alleged Violations of Children’s Privacy Law." (Yes, technically it's Alphabet now, but like Boaty McBoatface, the name everyone knows is Google. Even the FTC uses it.)

According to Statista: "In the most recently reported fiscal year, Google's revenue amounted to 136.22 billion US dollars. Google's revenue is largely made up by advertising revenue, which amounted to 116 billion US dollars in 2018."

170 MM / 136,220 MM =  0.125 %

2018 had 31,536,000 seconds, so that 170 MM corresponds to 10 hours, 57 minutes of revenue for Google. 

Here's a handy visualization:






Engineering, the key to success in sporting activities


Bowling 2.0 (some might call it cheating, I call it winning via superior technology) via Mark Rober:


I'd like a tool wall like his but it doesn't go with minimalism.



No numbers: recommendation success but product design fail.



Nerdy, pro-engineering products are a good choice for Amazon to recommend to me, but unfortunately many of them suffer from a visual form of "The Igon Value Problem."

Friday, July 19, 2019

Fat tails and extremistan - not the same thing



Extremistan and mediocrestan


What, are we making up words, now? (All words are made up. Think about it.)

Extremistan and mediocrestan are characterizations of distributions; a simple way to think about them is that very large events either totally dominate (extremistan) or don't (mediocrestan):

Height is in mediocrestan: if the average height in a room with ten people is 200 cm, that's probably from ten people between 190 and 210 cm tall and not nine people 100 cm tall and one person 1100 cm tall.  
Wealth is in extremistan: if the average wealth in a room with 10 people is 2 billion dollars, that's more likely to be one billionaire with 20 billion and nine average income people than ten billionaires with 2 billion each.

This classification determines whether you can estimate relevant population parameters from samples (mediocrestan yes, extremistan no) and how well-behaved order statistics (maximum, second place, etc) are (mediocrestan nicely predictable, extremistan not so much).

There's a fairly common error that people make when they learn about extremistan: they think that because distributions in extremistan have fat tails and are dominated by extreme values, then — and this is the error — distributions that have fat tails, especially those with extreme values, are in extremistan.

Note the error: $a \Rightarrow b$ is being used to assert $b \Rightarrow a$.

As we'll see next, not all fat-tailed and extreme-valued distributions are in extremistan.


A tale of two tails


Let us compare (a) the probability that $n$ similar outcomes of large size $M$ add up to a combined event of size $nM$ (or, equivalently, average to $M$) with (b) the probability of an extreme event of size $nM$ and $n-1$ events of size 0 add up to that combined event $nM$. If the first is higher than the second, we're in mediocrestan, if the second is higher than the first, we're in extremistan.

For the Normal distribution, the probabilities (a) denoted $P(\text{Similar})$ and (b) denoted $P(\text{Extreme})$ are:
\begin{eqnarray*}
P(\text{Similar}) &=& \frac{1}{(2 \pi)^{n/2}} \exp(- n \, M^2/2) \\
P(\text{Extreme})  &=& \frac{1}{(2 \pi)^{n/2}} \exp( - n^2 \,  M^2/2)
\end{eqnarray*}
It's trivial to see that for the Normal we have
\[
P(\text{Similar}) > P(\text{Extreme}).
\]
Unsurprisingly enough, with its reference excess kurtosis of 0, the Normal distribution is well inside mediocrestan.

For our fat-tailed, extreme-valued distribution, we'll use the Gumbel distribution, which is also known as Extreme Value Type I. A simple form of this distribution has the following pdf:
\[
f_X(x) = \exp(- x - \exp(-x))
\]
As shown here, its variance is $\pi^2/6$, while the Normal above has variance 1, but since we're comparing within class (Normal with Normal and Gumbel with Gumbel), that makes no difference and saves a lot of unnecessary clutter if we just use that pdf as is.

For Gumbel we have the following probabilities:
\begin{eqnarray*}
P(\mathrm{Similar}) &=& \exp(-nM - n \, \exp(-M))
\\
P(\mathrm{Extreme}) &=&  \exp(-nM - \exp(-nM) -n+1)
\end{eqnarray*}
Since for large $M$ we have $\exp(-nM) \approx 0$ and $\exp(-M) \approx 0$, then $\exp(-nM) +n-1 > n \, \exp(-M)$, for Gumbel we also have
\[
P(\text{Similar}) > P(\text{Extreme}).
\]
The Gumbel distribution belongs in mediocrestan, despite its fat tails and extreme values.

Really makes us think about the specialness of scale independent distributions, where we can bet on a big event to overwhelm all the small events (i.e. an extremistan distribution). Those are the distributions for which a trading strategy of enduring many small losses to capture the one big win can beat a strategy of consistent small wins.


What about the maximum?


In many cases the maximum is more relevant than the mean or median. So, how do fat tails influence the maxima?

When you look at the maximum of something, say the fastest kid in a class, the larger the class, the higher the maximum will be, on average. So the fastest kid in a group of 100 is on average faster than the fastest kid in a group of 10, for example.

In mediocrestan this increase is concave on the number of kids (the difference between the fastest kids in classes of 100 and 200 kids is bigger than the difference between the fastest kids in classes of 1100 and 1200 kids, on average); in extremistan there are no guarantees.

But once again, fat tails and extreme value distributions (the Gumbel, here scaled to have variance 1) have well-behaved maxima:



This nice concavity (note the logarithmic horizontal scale) makes things predictable; since many real-world metrics are known to be fat-tailed, it's comforting to know that their maxima don't explode all of a sudden.

Note that there's an effect of the extreme value: the maxima are larger and they grow faster, with less concavity than for the Normal.


And the point is…?


There are a number of people who assert that all sorts of research and social metrics are unusable because their analysis is based on mediocrestan (either by using sample statistics to estimate population statistics or by assuming regular behavior from order statistics), but — so goes the argument — these real world metrics have fat tails, so they are in extremistan.

The point of the above was to show that this form of argument (usually punctuated with gratuitous insults, expletives, and Mathematica-based math or other forms of using pretend-math to bully one's audience) is wrong, tout court.

Only a small subset of fat-tailed, extreme-valued distributions is in extremistan. For all the rest, we can use our usual tools.

Wednesday, November 9, 2016

Powerlifters vs Gym Rats, take 2

(This is a redo of the numbers in my previous powerlifters vs gym rats post, with assumptions that are less favorable to powerlifters.)

First, since we need some sort of metric to compare athletes, I'll unbiasedly 😀 choose the average of three lifts, bench press, deadlift, and squat, as a percentage of the bodyweight of the athlete. Call that metric $S$.

We'll use a standard Normal for the distribution of this metric, by subtracting the mean (100 percent of bodyweight for non-powerlifters, assuming that the average gym rat can bench, deadlift, and squat their own bodyweight) and dividing by the standard deviation (say 15 percent of bodyweight, using the scientific approach of judging 10 to be too little and 20 to be too much). In other words, for non-powerlifters, $z \doteq (S-100)/15.$

As in the previous post, we'll assume that powerlifters are 1 percent of the gym rats; but instead of the powerlifters having a mean at 2 (in $z$ space, 130 in $S$ space), they only have a one-SD advantage, that is their mean is at 1 (in $z$ space, 115 in $S$ space). In other words

$\qquad z \sim \mathcal{N}(0,1)\qquad $ for non-powerlifters
$\qquad z \sim \mathcal{N}(1,1)\qquad $ for powerlifters

Using these assumptions we can now compute the percentage of powerlifters that exist in a gym population above a given threshold; we can also compute the median score of all athletes who score above that threshold (click for larger):


Note that the conditional median that we're using here is lower  than the conditional mean, as the conditional distribution is skewed to the right, i.e. has a long right tail. The choice of the median is more informative for skewed distributions as a "sense of what we'll see in the gym."*

It's interesting to note that this is the median of the combined distribution of powerlifters and other gym rats, weighted by their proportion in the population above the threshold, so the difference between this median and the threshold is a non-monotonic function of the threshold as the curvature and the weight of the distribution of each type of athlete change significantly in the $1-8$ range of the table.

Under these weaker assumptions (pun intended), only when the threshold for inclusion passes 5 standard deviations from the other gym goers' mean do powerlifters become the majority of the qualifying athletes. Unless the gym is full of football players (that's american football), weightlifters, and strongman competitors, I think these assumptions are too unfavorable to powerlifters.

Here are some strong athletes moving metal, for variety (NSFW language):


"While they squat I eat cookies" has to be the most powerlifter-y sentence ever.

Update Nov 11, 2016: Here's the percentage of powerlifters in the population of qualifying athletes for different assumptions about the advantage of powerlifters (i.e. the mean of the powerlifters' distribution in standard deviation units); click for larger:



- - - - - -
* Unless there are CrossFit-ers in the gym, in which case what we typically see in the gym is dangerous, counter-productive nonsense.