{
  "id": "P044",
  "slug": "statistics",
  "key": "P044-statistics",
  "title": "Statistics",
  "summary": "A list of numbers - whole numbers, decimals or fractions, kept exact - and what describes it: count, range, mean, median, modes, quartiles and outliers, variance and standard deviation dividing by n and by n - 1, whether the mean sits above or below the median, and a histogram whose top edge does not drop the largest value.",
  "entry": "main.eml",
  "ui": "terminal",
  "readme": "# P044 - Statistics\n\nA list of up to 60 numbers - whole numbers, decimals or fractions - and what\ndescribes it: count, range, sum, mean, median, modes, quartiles and outliers,\nvariance and standard deviation dividing by n and by n - 1, whether the mean\nsits above or below the median, and a histogram. Every value is kept as an\nexact fraction, so nothing is rounded until it is shown, to three places.\n\n- `main.eml` - the menu, adding and removing values with their checks, and\n  the summary and histogram on screen\n- `stats.eml` - the measures: sorting, sum, mean, median, quartiles, modes,\n  squared distances and the histogram counts\n- `frac.eml` - exact fractions, from P007 by way of P038, with decimals and a\n  whole-number square root\n\nHow each part works:\n\n- The median is the middle value, or the mean of the two middle ones. The\n  quartiles are the medians of the lower and the upper half, leaving the\n  median out when the count is odd; values more than 1.5 interquartile\n  ranges beyond them are named as outliers.\n- The variance is the mean squared distance from the mean, exact; as a\n  sample it divides by n - 1 instead. The standard deviation is a square\n  root, which is rarely a fraction: it is found with a whole-number square\n  root (Newton's method, no floats) of the variance scaled up, and rounded\n  half up - \"about\" says when the shown digits are not exact.\n- The corpus case `mean-median-skew` is about averages that disagree: when\n  a few large values pull the mean above the median, the summary says so\n  and counts how many values are above the mean - in the sample, only 2 of\n  12.\n- The histogram's ranges include their lower edge; the last includes the\n  largest value too, the edge the corpus case `histogram-builder` handles,\n  so the top value is not dropped.\n\nWhat is checked: numbers as whole numbers, decimals (`2.5`, `-3.75`) or\nfractions (`3/4`), separated by spaces or commas, at most 60 in all - a line\nwith anything else adds nothing; a value to remove that is in the list; 2 to\n10 histogram ranges. An empty answer cancels.\n\nSessions: `sessions/basic.in` loads the sample - twelve delivery times, two\nof them far longer than the rest: mean about 21.083 against a median of 15,\noutliers 45 and 60 - draws a histogram with five ranges, removes 60 and\nsummarises again, then clears and summarises five numbers typed with\ndecimals and a fraction (no mode, an exact variance of 11.89).\n`sessions/bad-input.in` gives menu choices 0 and x, every action with no\nvalues, a line with x in it, 1/0, sixty-one numbers at once, three equal\nvalues (one mode, variance 0, no quartiles, no histogram ranges), a value\nnot in the list, abc, and 1, 11 and x ranges.\n\nBuilt on the verified corpus cases `mean-median-skew` (three averages on a\nskewed sample, and how far each moves) and `histogram-builder` (equal ranges\nand the top edge).\n",
  "modules": [
    {
      "name": "main.eml",
      "eml": "# P044 statistics: a list of numbers - whole numbers, decimals or fractions -\n# and what describes it: count, range, mean, median, modes, quartiles and\n# outliers, variance and standard deviation, and a histogram. Every value is\n# an exact fraction, so nothing is rounded until it is shown.\nimport frac\nimport stats\n\n60 => most_values\n\ndef trim(s):\n    0 => i\n    len(s) => j\n    while i < j and s[i] == \" \":\n        i + 1 => i\n    while j > i and s[j - 1] == \" \":\n        j - 1 => j\n    return s[i:j]\n\ndef words(s):\n    [] => out\n    \"\" => word\n    for c in s + \" \":\n        if c == \" \" or c == \",\":\n            if word != \"\":\n                out + [word] => out\n            \"\" => word\n        else:\n            word + c => word\n    return out\n\ndef d3(x):\n    return frac.decimal(x, 3)\n\ndef values_text(n):\n    if n == 1:\n        return \"1 value\"\n    return str(n) + \" values\"\n\ndef listed(xs):\n    \"\" => out\n    for i in [0:len(xs) - 1]:\n        if i > 0 and i == len(xs) - 1:\n            out + \" and \" => out\n        elif i > 0:\n            out + \", \" => out\n        out + d3(xs[i]) => out\n    return out\n\ndef add(xs):\n    trim(input(\"numbers (separated by spaces)> \")) => answer\n    if answer == \"\":\n        \"Cancelled.\" ^0\n        return xs\n    [] => new\n    for w in words(answer):\n        frac.parse(w) => x\n        if len(x) == 0:\n            (\"Not a number: \" + w + \". Nothing was added - type whole numbers, decimals like 2.5 or fractions like 3/4.\") ^0\n            return xs\n        new + [x] => new\n    if len(xs) + len(new) > most_values:\n        (\"That would make \" + str(len(xs) + len(new)) + \" values; the most is \" + str(most_values) + \". Nothing was added.\") ^0\n        return xs\n    (values_text(len(new)) + \" added; \" + values_text(len(xs) + len(new)) + \" in all.\") ^0\n    return xs + new\n\ndef remove(xs):\n    if len(xs) == 0:\n        \"There are no values.\" ^0\n        return xs\n    trim(input(\"value to remove> \")) => answer\n    if answer == \"\":\n        \"Cancelled.\" ^0\n        return xs\n    frac.parse(answer) => x\n    if len(x) == 0:\n        (\"Not a number: \" + answer + \".\") ^0\n        return xs\n    for i in [0:len(xs) - 1]:\n        if xs[i] == x:\n            (d3(x) + \" removed; \" + values_text(len(xs) - 1) + \" left.\") ^0\n            return xs[0:i] + xs[i + 1:len(xs)]\n    (d3(x) + \" is not in the list.\") ^0\n    return xs\n\ndef summary(xs):\n    if len(xs) == 0:\n        \"There are no values yet.\" ^0\n        return 0\n    stats.sorted_values(xs) => s\n    len(s) => n\n    s[0] => low\n    s[n - 1] => high\n    stats.mean(s) => m\n    stats.median_of(s) => md\n    (values_text(n) + \", from \" + d3(low) + \" to \" + d3(high) + \" (range \" + d3(frac.minus(high, low)) + \").\") ^0\n    (\"Sum \" + d3(stats.total(s)) + \", mean \" + d3(m) + \", median \" + d3(md) + \".\") ^0\n    stats.modes(s) => mo\n    if len(mo[0]) == 0:\n        \"once\" => often\n        if mo[1] > 1:\n            str(mo[1]) + \" times\" => often\n        (\"No mode: every value occurs \" + often + \".\") ^0\n    elif len(mo[0]) == 1:\n        (\"Mode: \" + d3(mo[0][0]) + \" (\" + str(mo[1]) + \" times).\") ^0\n    else:\n        (\"Modes: \" + listed(mo[0]) + \" (\" + str(mo[1]) + \" times each).\") ^0\n    if n >= 4:\n        stats.quartiles(s) => q\n        frac.minus(q[1], q[0]) => iqr\n        frac.times([3, 2], iqr) => reach\n        frac.minus(q[0], reach) => fence_low\n        frac.plus(q[1], reach) => fence_high\n        (\"Quartiles: Q1 \" + d3(q[0]) + \", Q3 \" + d3(q[1]) + \"; interquartile range \" + d3(iqr) + \".\") ^0\n        [] => out\n        for x in s:\n            if frac.less(x, fence_low) or frac.less(fence_high, x):\n                out + [x] => out\n        \"none\" => found\n        if len(out) > 0:\n            listed(out) => found\n        (\"Outliers, beyond 1.5 times the interquartile range (below \" + d3(fence_low) + \" or above \" + d3(fence_high) + \"): \" + found + \".\") ^0\n    else:\n        \"Quartiles need at least 4 values.\" ^0\n    stats.squares_about(s, m) => ss\n    frac.over(ss, [n, 1]) => var\n    (\"Variance \" + d3(var) + \", standard deviation \" + frac.root_decimal(var, 3) + \" - dividing by n.\") ^0\n    if n >= 2:\n        frac.over(ss, [n - 1, 1]) => svar\n        (\"As a sample: variance \" + d3(svar) + \", standard deviation \" + frac.root_decimal(svar, 3) + \" - dividing by n - 1.\") ^0\n    0 => above\n    for x in s:\n        if frac.less(m, x):\n            above + 1 => above\n    if frac.less(md, m):\n        (\"The mean is above the median: a few large values pull it up. \" + str(above) + \" of \" + str(n) + \" values are above the mean.\") ^0\n    elif frac.less(m, md):\n        (\"The mean is below the median: a few small values pull it down. \" + str(above) + \" of \" + str(n) + \" values are above the mean.\") ^0\n    else:\n        (\"The mean and the median are equal. \" + str(above) + \" of \" + str(n) + \" values are above the mean.\") ^0\n    return 0\n\ndef histogram(xs):\n    if len(xs) == 0:\n        \"There are no values yet.\" ^0\n        return 0\n    while True:\n        trim(input(\"bins (2 to 10)> \")) => answer\n        if answer == \"\":\n            \"Cancelled.\" ^0\n            return 0\n        frac.digits_value(answer) => k\n        if k >= 2 and k <= 10:\n            stats.sorted_values(xs) => s\n            s[0] => low\n            s[len(s) - 1] => high\n            if low == high:\n                (\"All \" + values_text(len(s)) + \" are \" + d3(low) + \": there are no ranges to draw.\") ^0\n                return 0\n            stats.bucket_counts(s, low, high, k) => counts\n            frac.over(frac.minus(high, low), [k, 1]) => width\n            [] => labels\n            0 => widest\n            for i in [0:k - 1]:\n                frac.plus(low, frac.times(width, [i, 1])) => a\n                frac.plus(low, frac.times(width, [i + 1, 1])) => b\n                \")\" => close\n                if i == k - 1:\n                    \"]\" => close\n                \"[\" + frac.decimal(a, 2) + \", \" + frac.decimal(b, 2) + close => label\n                labels + [label] => labels\n                if len(label) > widest:\n                    len(label) => widest\n            for i in [0:k - 1]:\n                labels[i] => label\n                while len(label) < widest:\n                    label + \" \" => label\n                (\"  \" + label + \"  \" + str(counts[i]) + \"  \" + \"#\" * counts[i]) => line\n                if counts[i] == 0:\n                    (\"  \" + label + \"  0\") => line\n                line ^0\n            (\"The ranges include their lower edge; the last includes the largest value too. \" + values_text(len(s)) + \" in all.\") ^0\n            return 0\n        \"Type a number from 2 to 10.\" ^0\n\n\"== Statistics ==\" ^0\n\"Numbers in, measures out - exact, rounded only when shown (to 3 places).\" ^0\n[] => xs\nTrue => running\nwhile running:\n    \"\" ^0\n    \"1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\" ^0\n    trim(input(\"choice> \")) => choice\n    if choice == \"1\":\n        add(xs) => xs\n    elif choice == \"2\":\n        remove(xs) => xs\n    elif choice == \"3\":\n        summary(xs)\n    elif choice == \"4\":\n        histogram(xs)\n    elif choice == \"5\":\n        if len(xs) == 0:\n            \"There are no values yet.\" ^0\n        else:\n            (\"Sorted: \" + listed(stats.sorted_values(xs)) + \".\") ^0\n    elif choice == \"6\":\n        [] => xs\n        \"All values cleared.\" ^0\n    elif choice == \"7\":\n        [] => xs\n        for v in [12, 15, 14, 18, 13, 16, 45, 14, 17, 15, 60, 14]:\n            xs + [[v, 1]] => xs\n        \"Sample: 12 delivery times in minutes.\" ^0\n    elif choice == \"8\":\n        False => running\n    else:\n        \"Pick a number from 1 to 8.\" ^0\n\"Bye.\" ^0\n",
      "python": "import frac\nimport stats\nmost_values = 60\n\ndef trim(s):\n    i = 0\n    j = len(s)\n    while i < j and s[i] == \" \":\n        i = i + 1\n    while j > i and s[j - 1] == \" \":\n        j = j - 1\n    return s[i:j]\n\ndef words(s):\n    out = []\n    word = \"\"\n    for c in s + \" \":\n        if c == \" \" or c == \",\":\n            if word != \"\":\n                out = out + [word]\n            word = \"\"\n        else:\n            word = word + c\n    return out\n\ndef d3(x):\n    return frac.decimal(x, 3)\n\ndef values_text(n):\n    if n == 1:\n        return \"1 value\"\n    return str(n) + \" values\"\n\ndef listed(xs):\n    out = \"\"\n    for i in range(0, len(xs)):\n        if i > 0 and i == len(xs) - 1:\n            out = out + \" and \"\n        elif i > 0:\n            out = out + \", \"\n        out = out + d3(xs[i])\n    return out\n\ndef add(xs):\n    answer = trim(input(\"numbers (separated by spaces)> \"))\n    if answer == \"\":\n        print(\"Cancelled.\")\n        return xs\n    new = []\n    for w in words(answer):\n        x = frac.parse(w)\n        if len(x) == 0:\n            print(\"Not a number: \" + w + \". Nothing was added - type whole numbers, decimals like 2.5 or fractions like 3/4.\")\n            return xs\n        new = new + [x]\n    if len(xs) + len(new) > most_values:\n        print(\"That would make \" + str(len(xs) + len(new)) + \" values; the most is \" + str(most_values) + \". Nothing was added.\")\n        return xs\n    print(values_text(len(new)) + \" added; \" + values_text(len(xs) + len(new)) + \" in all.\")\n    return xs + new\n\ndef remove(xs):\n    if len(xs) == 0:\n        print(\"There are no values.\")\n        return xs\n    answer = trim(input(\"value to remove> \"))\n    if answer == \"\":\n        print(\"Cancelled.\")\n        return xs\n    x = frac.parse(answer)\n    if len(x) == 0:\n        print(\"Not a number: \" + answer + \".\")\n        return xs\n    for i in range(0, len(xs)):\n        if xs[i] == x:\n            print(d3(x) + \" removed; \" + values_text(len(xs) - 1) + \" left.\")\n            return xs[0:i] + xs[i + 1:len(xs)]\n    print(d3(x) + \" is not in the list.\")\n    return xs\n\ndef summary(xs):\n    if len(xs) == 0:\n        print(\"There are no values yet.\")\n        return 0\n    s = stats.sorted_values(xs)\n    n = len(s)\n    low = s[0]\n    high = s[n - 1]\n    m = stats.mean(s)\n    md = stats.median_of(s)\n    print(values_text(n) + \", from \" + d3(low) + \" to \" + d3(high) + \" (range \" + d3(frac.minus(high, low)) + \").\")\n    print(\"Sum \" + d3(stats.total(s)) + \", mean \" + d3(m) + \", median \" + d3(md) + \".\")\n    mo = stats.modes(s)\n    if len(mo[0]) == 0:\n        often = \"once\"\n        if mo[1] > 1:\n            often = str(mo[1]) + \" times\"\n        print(\"No mode: every value occurs \" + often + \".\")\n    elif len(mo[0]) == 1:\n        print(\"Mode: \" + d3(mo[0][0]) + \" (\" + str(mo[1]) + \" times).\")\n    else:\n        print(\"Modes: \" + listed(mo[0]) + \" (\" + str(mo[1]) + \" times each).\")\n    if n >= 4:\n        q = stats.quartiles(s)\n        iqr = frac.minus(q[1], q[0])\n        reach = frac.times([3, 2], iqr)\n        fence_low = frac.minus(q[0], reach)\n        fence_high = frac.plus(q[1], reach)\n        print(\"Quartiles: Q1 \" + d3(q[0]) + \", Q3 \" + d3(q[1]) + \"; interquartile range \" + d3(iqr) + \".\")\n        out = []\n        for x in s:\n            if frac.less(x, fence_low) or frac.less(fence_high, x):\n                out = out + [x]\n        found = \"none\"\n        if len(out) > 0:\n            found = listed(out)\n        print(\"Outliers, beyond 1.5 times the interquartile range (below \" + d3(fence_low) + \" or above \" + d3(fence_high) + \"): \" + found + \".\")\n    else:\n        print(\"Quartiles need at least 4 values.\")\n    ss = stats.squares_about(s, m)\n    var = frac.over(ss, [n, 1])\n    print(\"Variance \" + d3(var) + \", standard deviation \" + frac.root_decimal(var, 3) + \" - dividing by n.\")\n    if n >= 2:\n        svar = frac.over(ss, [n - 1, 1])\n        print(\"As a sample: variance \" + d3(svar) + \", standard deviation \" + frac.root_decimal(svar, 3) + \" - dividing by n - 1.\")\n    above = 0\n    for x in s:\n        if frac.less(m, x):\n            above = above + 1\n    if frac.less(md, m):\n        print(\"The mean is above the median: a few large values pull it up. \" + str(above) + \" of \" + str(n) + \" values are above the mean.\")\n    elif frac.less(m, md):\n        print(\"The mean is below the median: a few small values pull it down. \" + str(above) + \" of \" + str(n) + \" values are above the mean.\")\n    else:\n        print(\"The mean and the median are equal. \" + str(above) + \" of \" + str(n) + \" values are above the mean.\")\n    return 0\n\ndef histogram(xs):\n    if len(xs) == 0:\n        print(\"There are no values yet.\")\n        return 0\n    while True:\n        answer = trim(input(\"bins (2 to 10)> \"))\n        if answer == \"\":\n            print(\"Cancelled.\")\n            return 0\n        k = frac.digits_value(answer)\n        if k >= 2 and k <= 10:\n            s = stats.sorted_values(xs)\n            low = s[0]\n            high = s[len(s) - 1]\n            if low == high:\n                print(\"All \" + values_text(len(s)) + \" are \" + d3(low) + \": there are no ranges to draw.\")\n                return 0\n            counts = stats.bucket_counts(s, low, high, k)\n            width = frac.over(frac.minus(high, low), [k, 1])\n            labels = []\n            widest = 0\n            for i in range(0, k):\n                a = frac.plus(low, frac.times(width, [i, 1]))\n                b = frac.plus(low, frac.times(width, [i + 1, 1]))\n                close = \")\"\n                if i == k - 1:\n                    close = \"]\"\n                label = \"[\" + frac.decimal(a, 2) + \", \" + frac.decimal(b, 2) + close\n                labels = labels + [label]\n                if len(label) > widest:\n                    widest = len(label)\n            for i in range(0, k):\n                label = labels[i]\n                while len(label) < widest:\n                    label = label + \" \"\n                line = \"  \" + label + \"  \" + str(counts[i]) + \"  \" + \"#\" * counts[i]\n                if counts[i] == 0:\n                    line = \"  \" + label + \"  0\"\n                print(line)\n            print(\"The ranges include their lower edge; the last includes the largest value too. \" + values_text(len(s)) + \" in all.\")\n            return 0\n        print(\"Type a number from 2 to 10.\")\n\nprint(\"== Statistics ==\")\nprint(\"Numbers in, measures out - exact, rounded only when shown (to 3 places).\")\nxs = []\nrunning = True\nwhile running:\n    print(\"\")\n    print(\"1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\")\n    choice = trim(input(\"choice> \"))\n    if choice == \"1\":\n        xs = add(xs)\n    elif choice == \"2\":\n        xs = remove(xs)\n    elif choice == \"3\":\n        summary(xs)\n    elif choice == \"4\":\n        histogram(xs)\n    elif choice == \"5\":\n        if len(xs) == 0:\n            print(\"There are no values yet.\")\n        else:\n            print(\"Sorted: \" + listed(stats.sorted_values(xs)) + \".\")\n    elif choice == \"6\":\n        xs = []\n        print(\"All values cleared.\")\n    elif choice == \"7\":\n        xs = []\n        for v in [12, 15, 14, 18, 13, 16, 45, 14, 17, 15, 60, 14]:\n            xs = xs + [[v, 1]]\n        print(\"Sample: 12 delivery times in minutes.\")\n    elif choice == \"8\":\n        running = False\n    else:\n        print(\"Pick a number from 1 to 8.\")\nprint(\"Bye.\")\n"
    },
    {
      "name": "stats.eml",
      "eml": "# P044 statistics - the measures. Values are exact fractions (frac.eml).\nimport frac\n\ndef sorted_values(xs):\n    # Insertion sort, as in the corpus case mean-median-skew.\n    [] => out\n    for x in xs:\n        out + [x] => out\n    1 => i\n    while i < len(out):\n        out[i] => cur\n        i - 1 => j\n        while j >= 0 and frac.less(cur, out[j]):\n            out[j] => out[j + 1]\n            j - 1 => j\n        cur => out[j + 1]\n        i + 1 => i\n    return out\n\ndef total(xs):\n    [0, 1] => s\n    for x in xs:\n        frac.plus(s, x) => s\n    return s\n\ndef mean(xs):\n    return frac.over(total(xs), [len(xs), 1])\n\ndef median_of(s):\n    # The middle of sorted values; the mean of the two middle ones when\n    # there is an even number.\n    len(s) => n\n    int(n / 2) => h\n    if n % 2 == 1:\n        return s[h]\n    return frac.over(frac.plus(s[h - 1], s[h]), [2, 1])\n\ndef quartiles(s):\n    # [Q1, Q3]: the medians of the lower and the upper half, leaving out the\n    # median itself when the count is odd. Needs at least 4 values.\n    len(s) => n\n    int(n / 2) => h\n    s[0:h] => lower\n    s[n - h:n] => upper\n    return [median_of(lower), median_of(upper)]\n\ndef modes(s):\n    # The values that occur most often, in order, and how often; [] when\n    # there are several values and each occurs the same number of times.\n    [] => values\n    [] => counts\n    for x in s:\n        if len(values) > 0 and values[len(values) - 1] == x:\n            counts[len(counts) - 1] + 1 => counts[len(counts) - 1]\n        else:\n            values + [x] => values\n            counts + [1] => counts\n    0 => top\n    for c in counts:\n        if c > top:\n            c => top\n    [] => out\n    for i in [0:len(values) - 1]:\n        if counts[i] == top:\n            out + [values[i]] => out\n    if len(values) > 1 and len(out) == len(values):\n        return [[], top]\n    return [out, top]\n\ndef squares_about(xs, m):\n    # The sum of squared distances from m.\n    [0, 1] => s\n    for x in xs:\n        frac.minus(x, m) => d\n        frac.plus(s, frac.times(d, d)) => s\n    return s\n\ndef bucket_counts(s, low, high, count):\n    # How many values fall in each of `count` equal ranges from low to high;\n    # a range holds its lower edge, and the very top value goes in the last\n    # one - the edge the corpus case histogram-builder is about.\n    [0] * count => counts\n    frac.minus(high, low) => span\n    for x in s:\n        if frac.is_zero(span):\n            0 => k\n        else:\n            frac.over(frac.times(frac.minus(x, low), [count, 1]), span) => pos\n            frac.quotient(pos[0], pos[1]) => k\n            if k >= count:\n                count - 1 => k\n        counts[k] + 1 => counts[k]\n    return counts\n",
      "python": "import frac\n\ndef sorted_values(xs):\n    out = []\n    for x in xs:\n        out = out + [x]\n    i = 1\n    while i < len(out):\n        cur = out[i]\n        j = i - 1\n        while j >= 0 and frac.less(cur, out[j]):\n            out[j + 1] = out[j]\n            j = j - 1\n        out[j + 1] = cur\n        i = i + 1\n    return out\n\ndef total(xs):\n    s = [0, 1]\n    for x in xs:\n        s = frac.plus(s, x)\n    return s\n\ndef mean(xs):\n    return frac.over(total(xs), [len(xs), 1])\n\ndef median_of(s):\n    n = len(s)\n    h = int(n / 2)\n    if n % 2 == 1:\n        return s[h]\n    return frac.over(frac.plus(s[h - 1], s[h]), [2, 1])\n\ndef quartiles(s):\n    n = len(s)\n    h = int(n / 2)\n    lower = s[0:h]\n    upper = s[n - h:n]\n    return [median_of(lower), median_of(upper)]\n\ndef modes(s):\n    values = []\n    counts = []\n    for x in s:\n        if len(values) > 0 and values[len(values) - 1] == x:\n            counts[len(counts) - 1] = counts[len(counts) - 1] + 1\n        else:\n            values = values + [x]\n            counts = counts + [1]\n    top = 0\n    for c in counts:\n        if c > top:\n            top = c\n    out = []\n    for i in range(0, len(values)):\n        if counts[i] == top:\n            out = out + [values[i]]\n    if len(values) > 1 and len(out) == len(values):\n        return [[], top]\n    return [out, top]\n\ndef squares_about(xs, m):\n    s = [0, 1]\n    for x in xs:\n        d = frac.minus(x, m)\n        s = frac.plus(s, frac.times(d, d))\n    return s\n\ndef bucket_counts(s, low, high, count):\n    counts = [0] * count\n    span = frac.minus(high, low)\n    for x in s:\n        if frac.is_zero(span):\n            k = 0\n        else:\n            pos = frac.over(frac.times(frac.minus(x, low), [count, 1]), span)\n            k = frac.quotient(pos[0], pos[1])\n            if k >= count:\n                k = count - 1\n        counts[k] = counts[k] + 1\n    return counts\n"
    },
    {
      "name": "frac.eml",
      "eml": "# P044 statistics - exact fractions, from P007 (calculator) by way of P038.\n# A number is a list [n, d]: n / d in lowest terms with d > 0. Integers have\n# no size limit, but EML has no //, and a / b goes through a float that cannot\n# hold a large quotient exactly, so whole-number division is written out.\n\ndef quotient(a, b):\n    # a // b for whole numbers a >= 0 and b > 0, exact at any size. Long\n    # division by doubling: take away the largest b * 2^k that still fits.\n    0 => q\n    while a >= b:\n        b => m\n        1 => k\n        while m + m <= a:\n            m + m => m\n            k + k => k\n        a - m => a\n        q + k => q\n    return q\n\ndef gcd(a, b):\n    while b != 0:\n        a % b => r\n        b => a\n        r => b\n    return a\n\ndef make(n, d):\n    # n / d in lowest terms with a positive denominator (d != 0).\n    if d < 0:\n        0 - n => n\n        0 - d => d\n    abs(n) => a\n    gcd(a, d) => g\n    if n < 0:\n        return [0 - quotient(a, g), quotient(d, g)]\n    return [quotient(a, g), quotient(d, g)]\n\ndef plus(x, y):\n    return make(x[0] * y[1] + y[0] * x[1], x[1] * y[1])\n\ndef minus(x, y):\n    return make(x[0] * y[1] - y[0] * x[1], x[1] * y[1])\n\ndef times(x, y):\n    return make(x[0] * y[0], x[1] * y[1])\n\ndef over(x, y):\n    # x / y; the caller has checked that y is not zero.\n    return make(x[0] * y[1], x[1] * y[0])\n\ndef is_zero(x):\n    return x[0] == 0\n\ndef digits_value(s):\n    # The value of a string of 1 to 9 digits, otherwise -1.\n    if s == \"\" or len(s) > 9:\n        return -1\n    0 => n\n    for c in s:\n        if not (c in \"0123456789\"):\n            return -1\n        n * 10 + int(c) => n\n    return n\n\ndef parse(s):\n    # \"3\", \"-2\", \"3/4\", \"-0.25\" as a fraction; [] if it is not a number.\n    1 => sign\n    if len(s) > 0 and s[0] == \"-\":\n        -1 => sign\n        s[1:len(s)] => s\n    0 => k\n    while k < len(s) and s[k] != \"/\" and s[k] != \".\":\n        k + 1 => k\n    if k == len(s):\n        digits_value(s) => n\n        if n < 0:\n            return []\n        return make(sign * n, 1)\n    digits_value(s[0:k]) => whole\n    s[k + 1:len(s)] => rest\n    digits_value(rest) => part\n    if whole < 0 or part < 0:\n        return []\n    if s[k] == \"/\":\n        if part == 0:\n            return []\n        return make(sign * whole, part)\n    # a decimal point: 0.25 is 25 / 100\n    1 => scale\n    for c in rest:\n        scale * 10 => scale\n    return make(sign * (whole * scale + part), scale)\n\ndef less(x, y):\n    return x[0] * y[1] < y[0] * x[1]\n\ndef isqrt(n):\n    # The whole-number square root of n >= 0, rounded down: Newton's method\n    # on whole numbers, which never needs a float.\n    if n < 2:\n        return n\n    n => x\n    quotient(x + 1, 2) => y\n    while y < x:\n        y => x\n        quotient(x + quotient(n, x), 2) => y\n    return x\n\ndef decimal(x, places):\n    # x to the given number of decimal places, halves rounded away from zero,\n    # with \"about \" in front when the value does not end there exactly.\n    \"\" => sign\n    abs(x[0]) => n\n    if x[0] < 0:\n        \"-\" => sign\n    1 => scale\n    for k in [1:places]:\n        scale * 10 => scale\n    quotient(2 * n * scale + x[1], 2 * x[1]) => t\n    \"\" => about\n    if (n * scale) % x[1] != 0:\n        \"about \" => about\n    str(quotient(t, scale)) => whole\n    str(t % scale) => part\n    while len(part) < places:\n        \"0\" + part => part\n    # drop trailing zeros of an exact value\n    if about == \"\":\n        while len(part) > 0 and part[len(part) - 1] == \"0\":\n            part[0:len(part) - 1] => part\n    if t == 0:\n        \"\" => sign\n    if part == \"\":\n        return about + sign + whole\n    return about + sign + whole + \".\" + part\n\ndef root_decimal(x, places):\n    # The square root of x >= 0 to the given places, rounded half up, by a\n    # whole-number square root of x * 100^places * 4 (one more binary digit\n    # to round with).\n    1 => scale\n    for k in [1:places]:\n        scale * 10 => scale\n    isqrt(quotient(x[0] * scale * scale * 4, x[1])) => r2\n    quotient(r2 + 1, 2) => t\n    \"about \" => about\n    if t * t * x[1] == x[0] * scale * scale:\n        \"\" => about\n    str(t % scale) => part\n    while len(part) < places:\n        \"0\" + part => part\n    if about == \"\":\n        while len(part) > 0 and part[len(part) - 1] == \"0\":\n            part[0:len(part) - 1] => part\n    if part == \"\":\n        return about + str(quotient(t, scale))\n    return about + str(quotient(t, scale)) + \".\" + part\n\ndef text(x):\n    # \"3\", \"-3\", \"3/4\", \"-3/4\".\n    if x[1] == 1:\n        return str(x[0])\n    return str(x[0]) + \"/\" + str(x[1])\n",
      "python": "def quotient(a, b):\n    q = 0\n    while a >= b:\n        m = b\n        k = 1\n        while m + m <= a:\n            m = m + m\n            k = k + k\n        a = a - m\n        q = q + k\n    return q\n\ndef gcd(a, b):\n    while b != 0:\n        r = a % b\n        a = b\n        b = r\n    return a\n\ndef make(n, d):\n    if d < 0:\n        n = 0 - n\n        d = 0 - d\n    a = abs(n)\n    g = gcd(a, d)\n    if n < 0:\n        return [0 - quotient(a, g), quotient(d, g)]\n    return [quotient(a, g), quotient(d, g)]\n\ndef plus(x, y):\n    return make(x[0] * y[1] + y[0] * x[1], x[1] * y[1])\n\ndef minus(x, y):\n    return make(x[0] * y[1] - y[0] * x[1], x[1] * y[1])\n\ndef times(x, y):\n    return make(x[0] * y[0], x[1] * y[1])\n\ndef over(x, y):\n    return make(x[0] * y[1], x[1] * y[0])\n\ndef is_zero(x):\n    return x[0] == 0\n\ndef digits_value(s):\n    if s == \"\" or len(s) > 9:\n        return -1\n    n = 0\n    for c in s:\n        if not c in \"0123456789\":\n            return -1\n        n = n * 10 + int(c)\n    return n\n\ndef parse(s):\n    sign = 1\n    if len(s) > 0 and s[0] == \"-\":\n        sign = -1\n        s = s[1:len(s)]\n    k = 0\n    while k < len(s) and s[k] != \"/\" and s[k] != \".\":\n        k = k + 1\n    if k == len(s):\n        n = digits_value(s)\n        if n < 0:\n            return []\n        return make(sign * n, 1)\n    whole = digits_value(s[0:k])\n    rest = s[k + 1:len(s)]\n    part = digits_value(rest)\n    if whole < 0 or part < 0:\n        return []\n    if s[k] == \"/\":\n        if part == 0:\n            return []\n        return make(sign * whole, part)\n    scale = 1\n    for c in rest:\n        scale = scale * 10\n    return make(sign * (whole * scale + part), scale)\n\ndef less(x, y):\n    return x[0] * y[1] < y[0] * x[1]\n\ndef isqrt(n):\n    if n < 2:\n        return n\n    x = n\n    y = quotient(x + 1, 2)\n    while y < x:\n        x = y\n        y = quotient(x + quotient(n, x), 2)\n    return x\n\ndef decimal(x, places):\n    sign = \"\"\n    n = abs(x[0])\n    if x[0] < 0:\n        sign = \"-\"\n    scale = 1\n    for k in range(1, places+1):\n        scale = scale * 10\n    t = quotient(2 * n * scale + x[1], 2 * x[1])\n    about = \"\"\n    if n * scale % x[1] != 0:\n        about = \"about \"\n    whole = str(quotient(t, scale))\n    part = str(t % scale)\n    while len(part) < places:\n        part = \"0\" + part\n    if about == \"\":\n        while len(part) > 0 and part[len(part) - 1] == \"0\":\n            part = part[0:len(part) - 1]\n    if t == 0:\n        sign = \"\"\n    if part == \"\":\n        return about + sign + whole\n    return about + sign + whole + \".\" + part\n\ndef root_decimal(x, places):\n    scale = 1\n    for k in range(1, places+1):\n        scale = scale * 10\n    r2 = isqrt(quotient(x[0] * scale * scale * 4, x[1]))\n    t = quotient(r2 + 1, 2)\n    about = \"about \"\n    if t * t * x[1] == x[0] * scale * scale:\n        about = \"\"\n    part = str(t % scale)\n    while len(part) < places:\n        part = \"0\" + part\n    if about == \"\":\n        while len(part) > 0 and part[len(part) - 1] == \"0\":\n            part = part[0:len(part) - 1]\n    if part == \"\":\n        return about + str(quotient(t, scale))\n    return about + str(quotient(t, scale)) + \".\" + part\n\ndef text(x):\n    if x[1] == 1:\n        return str(x[0])\n    return str(x[0]) + \"/\" + str(x[1])\n"
    }
  ],
  "sessions": [
    {
      "name": "bad-input",
      "input": "0\nx\n3\n4\n5\n2\n1\n3 x 4\n1\n1/0\n1\n\n1\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61\n1\n5, 5, 5\n3\n4\n2\n2\n7\n2\nabc\n2\n\n4\n1\n11\nx\n\n6\n5\n8\n",
      "screen": "== Statistics ==\nNumbers in, measures out - exact, rounded only when shown (to 3 places).\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 0\nPick a number from 1 to 8.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> x\nPick a number from 1 to 8.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 3\nThere are no values yet.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 4\nThere are no values yet.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 5\nThere are no values yet.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 2\nThere are no values.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 1\nnumbers (separated by spaces)> 3 x 4\nNot a number: x. Nothing was added - type whole numbers, decimals like 2.5 or fractions like 3/4.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 1\nnumbers (separated by spaces)> 1/0\nNot a number: 1/0. Nothing was added - type whole numbers, decimals like 2.5 or fractions like 3/4.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 1\nnumbers (separated by spaces)> \nCancelled.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 1\nnumbers (separated by spaces)> 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61\nThat would make 61 values; the most is 60. Nothing was added.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 1\nnumbers (separated by spaces)> 5, 5, 5\n3 values added; 3 values in all.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 3\n3 values, from 5 to 5 (range 0).\nSum 15, mean 5, median 5.\nMode: 5 (3 times).\nQuartiles need at least 4 values.\nVariance 0, standard deviation 0 - dividing by n.\nAs a sample: variance 0, standard deviation 0 - dividing by n - 1.\nThe mean and the median are equal. 0 of 3 values are above the mean.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 4\nbins (2 to 10)> 2\nAll 3 values are 5: there are no ranges to draw.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 2\nvalue to remove> 7\n7 is not in the list.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 2\nvalue to remove> abc\nNot a number: abc.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 2\nvalue to remove> \nCancelled.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 4\nbins (2 to 10)> 1\nType a number from 2 to 10.\nbins (2 to 10)> 11\nType a number from 2 to 10.\nbins (2 to 10)> x\nType a number from 2 to 10.\nbins (2 to 10)> \nCancelled.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 6\nAll values cleared.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 5\nThere are no values yet.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 8\nBye.\n",
      "interpreter": "equal"
    },
    {
      "name": "basic",
      "input": "7\n3\n4\n5\n5\n2\n60\n3\n6\n1\n2.5 3.5 1/4 7 10\n3\n4\n3\n8\n",
      "screen": "== Statistics ==\nNumbers in, measures out - exact, rounded only when shown (to 3 places).\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 7\nSample: 12 delivery times in minutes.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 3\n12 values, from 12 to 60 (range 48).\nSum 253, mean about 21.083, median 15.\nMode: 14 (3 times).\nQuartiles: Q1 14, Q3 17.5; interquartile range 3.5.\nOutliers, beyond 1.5 times the interquartile range (below 8.75 or above 22.75): 45 and 60.\nVariance about 209.243, standard deviation about 14.465 - dividing by n.\nAs a sample: variance about 228.265, standard deviation about 15.108 - dividing by n - 1.\nThe mean is above the median: a few large values pull it up. 2 of 12 values are above the mean.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 4\nbins (2 to 10)> 5\n  [12, 21.6)    10  ##########\n  [21.6, 31.2)  0\n  [31.2, 40.8)  0\n  [40.8, 50.4)  1  #\n  [50.4, 60]    1  #\nThe ranges include their lower edge; the last includes the largest value too. 12 values in all.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 5\nSorted: 12, 13, 14, 14, 14, 15, 15, 16, 17, 18, 45 and 60.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 2\nvalue to remove> 60\n60 removed; 11 values left.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 3\n11 values, from 12 to 45 (range 33).\nSum 193, mean about 17.545, median 15.\nMode: 14 (3 times).\nQuartiles: Q1 14, Q3 17; interquartile range 3.\nOutliers, beyond 1.5 times the interquartile range (below 9.5 or above 21.5): 45.\nVariance about 78.066, standard deviation about 8.836 - dividing by n.\nAs a sample: variance about 85.873, standard deviation about 9.267 - dividing by n - 1.\nThe mean is above the median: a few large values pull it up. 2 of 11 values are above the mean.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 6\nAll values cleared.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 1\nnumbers (separated by spaces)> 2.5 3.5 1/4 7 10\n5 values added; 5 values in all.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 3\n5 values, from 0.25 to 10 (range 9.75).\nSum 23.25, mean 4.65, median 3.5.\nNo mode: every value occurs once.\nQuartiles: Q1 1.375, Q3 8.5; interquartile range 7.125.\nOutliers, beyond 1.5 times the interquartile range (below about -9.313 or above about 19.188): none.\nVariance 11.89, standard deviation about 3.448 - dividing by n.\nAs a sample: variance about 14.863, standard deviation about 3.855 - dividing by n - 1.\nThe mean is above the median: a few large values pull it up. 2 of 5 values are above the mean.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 4\nbins (2 to 10)> 3\n  [0.25, 3.5)  2  ##\n  [3.5, 6.75)  1  #\n  [6.75, 10]   2  ##\nThe ranges include their lower edge; the last includes the largest value too. 5 values in all.\n\n1) add numbers  2) remove one  3) summary  4) histogram  5) sorted  6) clear  7) sample  8) quit\nchoice> 8\nBye.\n",
      "interpreter": "equal"
    }
  ],
  "builtOn": [
    {
      "slug": "mean-median-skew",
      "caseId": "263-mean-median-skew",
      "title": "Mean, median, skew — the average nobody experienced"
    },
    {
      "slug": "histogram-builder",
      "caseId": "113-histogram-builder",
      "title": "Histogram builder"
    }
  ],
  "updated": "2026-10-10"
}
