<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Posts on DataLeaf</title><link>https://dataleaf-642346.gitlab.io/posts/</link><description>Recent content in Posts on DataLeaf</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Tue, 18 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://dataleaf-642346.gitlab.io/posts/index.xml" rel="self" type="application/rss+xml"/><item><title>Least squares, three ways</title><link>https://dataleaf-642346.gitlab.io/posts/least-squares-three-ways/</link><pubDate>Tue, 18 Aug 2026 00:00:00 +0000</pubDate><guid>https://dataleaf-642346.gitlab.io/posts/least-squares-three-ways/</guid><description>Polynomial fits make the difference visible: with twelve columns, the normal equations lose every digit while QR and the SVD keep eight or nine. One identity about condition numbers explains it, and also says when the fast method is good enough.</description></item><item><title>Notebook: rerunning an analysis from two years ago</title><link>https://dataleaf-642346.gitlab.io/posts/rerunning-an-old-analysis/</link><pubDate>Fri, 14 Nov 2025 00:00:00 +0000</pubDate><guid>https://dataleaf-642346.gitlab.io/posts/rerunning-an-old-analysis/</guid><description>&lt;p&gt;A reviewer asked for one more column in a table produced two years ago. Adding a column means rerunning the analysis, and rerunning it means first showing that the old numbers come back. This entry records what that took.&lt;/p&gt;&#10;&lt;div class="dataleaf-heading"&gt;&lt;h2 id="checklist"&gt;Checklist&lt;/h2&gt;&#10; &lt;a&#10; class="dataleaf-heading-anchor"&#10; href="#checklist"&#10; aria-label="Permalink to Checklist"&#10; data-heading-anchor&#10; data-heading-copy="Copy link to Checklist"&#10; data-heading-copied="Copied link to Checklist."&gt;#&lt;/a&gt;&#10;&lt;/div&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;input checked="" disabled="" type="checkbox"&gt; Check out the tagged commit that produced the submitted table&lt;/li&gt;&#10;&lt;li&gt;&lt;input checked="" disabled="" type="checkbox"&gt; Recreate the environment from the lock file&lt;/li&gt;&#10;&lt;li&gt;&lt;input checked="" disabled="" type="checkbox"&gt; Verify the input data against the recorded checksums&lt;/li&gt;&#10;&lt;li&gt;&lt;input checked="" disabled="" type="checkbox"&gt; Rerun the pipeline end to end and compare every cell&lt;/li&gt;&#10;&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Record the seed of every random step&lt;/li&gt;&#10;&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Add the new column and log the rerun&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The first four steps took an afternoon:&lt;/p&gt;</description></item><item><title>Adding up ten million numbers</title><link>https://dataleaf-642346.gitlab.io/posts/compensated-summation/</link><pubDate>Tue, 03 Jun 2025 00:00:00 +0000</pubDate><guid>https://dataleaf-642346.gitlab.io/posts/compensated-summation/</guid><description>&lt;p&gt;A floating-point sum looks like the most innocent computation in a program. It is computed one addition at a time, &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;^&lt;/mo&gt;&lt;/mover&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;fl&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo&gt;^&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;\hat s_k = \operatorname{fl}(\hat s_{k-1} + x_k)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;, and every addition rounds its result to the nearest representable number. Each rounding is tiny, but a long sum performs millions of them, and their errors do not have to cancel.&lt;/p&gt;&#10;&lt;div class="dataleaf-heading"&gt;&lt;h2 id="how-large-can-the-error-get"&gt;How large can the error get?&lt;/h2&gt;&#10; &lt;a&#10; class="dataleaf-heading-anchor"&#10; href="#how-large-can-the-error-get"&#10; aria-label="Permalink to How large can the error get?"&#10; data-heading-anchor&#10; data-heading-copy="Copy link to How large can the error get?"&#10; data-heading-copied="Copied link to How large can the error get?."&gt;#&lt;/a&gt;&#10;&lt;/div&gt;&#10;&lt;p&gt;For recursive summation, the classic bound is&lt;sup id="fnref:1"&gt;&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;</description></item><item><title>A note on log-sum-exp</title><link>https://dataleaf-642346.gitlab.io/posts/log-sum-exp/</link><pubDate>Wed, 02 Oct 2024 00:00:00 +0000</pubDate><guid>https://dataleaf-642346.gitlab.io/posts/log-sum-exp/</guid><description>&lt;p&gt;Log-likelihoods, mixture models, and softmax layers all need the same quantity,&lt;/p&gt;&#10;&lt;div class="dataleaf-math dataleaf-math--block" tabindex="0" role="group" aria-label="Scrollable equation"&gt;&lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML" display="block"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;LSE&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;mo&gt;…&lt;/mo&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;munderover&gt;&lt;mo&gt;∑&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/munderover&gt;&lt;msup&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/msup&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;&#10;\operatorname{LSE}(x_1, \dots, x_n) = \log \sum_{i=1}^{n} e^{x_i},&#10;&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;/div&gt;&lt;p&gt;and the direct formula fails at the first large argument. In double precision, &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;e^{x}&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; overflows to infinity once &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;x&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; exceeds about 709.78, and underflows to zero once &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;x&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; falls below about &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mn&gt;745&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;-745&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;. Log-likelihoods of a few thousand observations easily reach either range.&lt;/p&gt;&#10;&lt;p&gt;The fix is one line of algebra. For any constant &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;m&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;,&lt;/p&gt;&#10;&lt;div class="dataleaf-math dataleaf-math--block" tabindex="0" role="group" aria-label="Scrollable equation"&gt;&lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML" display="block"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;munderover&gt;&lt;mo&gt;∑&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/munderover&gt;&lt;msup&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/msup&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;munderover&gt;&lt;mo&gt;∑&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/munderover&gt;&lt;msup&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;&#10;\log \sum_{i=1}^{n} e^{x_i} = m + \log \sum_{i=1}^{n} e^{x_i - m},&#10;&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;/div&gt;&lt;p&gt;and choosing &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mi&gt;max&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;m = \max_i x_i&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; makes the largest exponent exactly zero. Every term is then at most one, so nothing overflows, and the largest term is exactly one, so the sum is at least one and its logarithm is finite.&lt;/p&gt;</description></item><item><title>Why the sample variance divides by n − 1</title><link>https://dataleaf-642346.gitlab.io/posts/bessels-correction/</link><pubDate>Mon, 19 Feb 2024 00:00:00 +0000</pubDate><guid>https://dataleaf-642346.gitlab.io/posts/bessels-correction/</guid><description>&lt;p&gt;Every introductory course defines the sample variance with &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;n - 1&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; in the denominator, and most students accept it on trust. The reason fits in three lines, and so does the caveat that usually goes unmentioned.&lt;/p&gt;&#10;&lt;div class="dataleaf-statement dataleaf-statement--definition" role="group" aria-label="Definition: Sample variance"&gt;&#10; &lt;div class="dataleaf-statement__heading"&gt;&lt;strong&gt;Definition&lt;/strong&gt; &lt;span&gt;— Sample variance&lt;/span&gt;&lt;/div&gt;&#10; &lt;div class="dataleaf-statement__body"&gt;For observations &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;mo&gt;…&lt;/mo&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;x_1, \dots, x_n&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; with mean &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mover accent="true"&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo&gt;ˉ&lt;/mo&gt;&lt;/mover&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/mfrac&gt;&lt;msub&gt;&lt;mo&gt;∑&lt;/mo&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;\bar x = \frac{1}{n} \sum_{i} x_i&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;, let &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msubsup&gt;&lt;mo&gt;∑&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/msubsup&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo&gt;ˉ&lt;/mo&gt;&lt;/mover&gt;&lt;msup&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;S = \sum_{i=1}^{n} (x_i - \bar x)^2&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;. The sample variance is &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi mathvariant="normal"&gt;/&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;s^2 = S / (n - 1)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;; the plug-in variance is &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mover accent="true"&gt;&lt;mi&gt;σ&lt;/mi&gt;&lt;mo&gt;^&lt;/mo&gt;&lt;/mover&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi mathvariant="normal"&gt;/&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;\hat\sigma^2 = S / n&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;.&lt;/div&gt;&#10;&lt;/div&gt;&#10;&#10;&lt;div class="dataleaf-heading"&gt;&lt;h2 id="the-bias-and-its-correction"&gt;The bias, and its correction&lt;/h2&gt;&#10; &lt;a&#10; class="dataleaf-heading-anchor"&#10; href="#the-bias-and-its-correction"&#10; aria-label="Permalink to The bias, and its correction"&#10; data-heading-anchor&#10; data-heading-copy="Copy link to The bias, and its correction"&#10; data-heading-copied="Copied link to The bias, and its correction."&gt;#&lt;/a&gt;&#10;&lt;/div&gt;&#10;&lt;div class="dataleaf-statement dataleaf-statement--theorem" role="group" aria-label="Theorem: Unbiasedness"&gt;&#10; &lt;div class="dataleaf-statement__heading"&gt;&lt;strong&gt;Theorem&lt;/strong&gt; &lt;span&gt;— Unbiasedness&lt;/span&gt;&lt;/div&gt;&#10; &lt;div class="dataleaf-statement__body"&gt;If &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;mo&gt;…&lt;/mo&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;X_1, \dots, X_n&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; are independent with mean &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;μ&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;\mu&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; and variance &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi&gt;σ&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mo&gt;&amp;lt;&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;∞&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;\sigma^2 &amp;lt; \infty&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;, then &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mo stretchy="false"&gt;[&lt;/mo&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mo stretchy="false"&gt;]&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mtext&gt; &lt;/mtext&gt;&lt;msup&gt;&lt;mi&gt;σ&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;\mathbb{E}[S] = (n - 1)\,\sigma^2&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;. Consequently &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mo stretchy="false"&gt;[&lt;/mo&gt;&lt;msup&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mo stretchy="false"&gt;]&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msup&gt;&lt;mi&gt;σ&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;\mathbb{E}[s^2] = \sigma^2&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;, while &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mo stretchy="false"&gt;[&lt;/mo&gt;&lt;msup&gt;&lt;mover accent="true"&gt;&lt;mi&gt;σ&lt;/mi&gt;&lt;mo&gt;^&lt;/mo&gt;&lt;/mover&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mo stretchy="false"&gt;]&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/mfrac&gt;&lt;mtext&gt; &lt;/mtext&gt;&lt;msup&gt;&lt;mi&gt;σ&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;\mathbb{E}[\hat\sigma^2] = \frac{n-1}{n}\,\sigma^2&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;.&lt;/div&gt;&#10;&lt;/div&gt;&#10;&#10;&lt;div class="dataleaf-proof" role="group" aria-label="Proof"&gt;&#10; &lt;div class="dataleaf-proof__heading"&gt;Proof.&lt;/div&gt;&#10; &lt;div class="dataleaf-proof__body"&gt;&lt;p&gt;Split each deviation from the true mean into a deviation from the sample mean and the error of the sample mean. The cross terms sum to zero, which gives the identity &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo&gt;∑&lt;/mo&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mi&gt;μ&lt;/mi&gt;&lt;msup&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mo&gt;ˉ&lt;/mo&gt;&lt;/mover&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mi&gt;μ&lt;/mi&gt;&lt;msup&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;\sum_i (X_i - \mu)^2 = S + n (\bar X - \mu)^2&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;. Taking expectations,&lt;/p&gt;&#10;&lt;div class="dataleaf-math dataleaf-math--block" tabindex="0" role="group" aria-label="Scrollable equation"&gt;&lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML" display="block"&gt;&lt;semantics&gt;&lt;mtable rowspacing="0.25em" columnalign="right left" columnspacing="0em"&gt;&lt;mtr&gt;&lt;mtd&gt;&lt;mstyle scriptlevel="0" displaystyle="true"&gt;&lt;mrow&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mo stretchy="false"&gt;[&lt;/mo&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mo stretchy="false"&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;/mstyle&gt;&lt;/mtd&gt;&lt;mtd&gt;&lt;mstyle scriptlevel="0" displaystyle="true"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;munderover&gt;&lt;mo&gt;∑&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/munderover&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mo fence="true" stretchy="true" minsize="1.2em" maxsize="1.2em"&gt;[&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mi&gt;μ&lt;/mi&gt;&lt;msup&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mo fence="true" stretchy="true" minsize="1.2em" maxsize="1.2em"&gt;]&lt;/mo&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mtext&gt; &lt;/mtext&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mo fence="true" stretchy="true" minsize="1.2em" maxsize="1.2em"&gt;[&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mo&gt;ˉ&lt;/mo&gt;&lt;/mover&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mi&gt;μ&lt;/mi&gt;&lt;msup&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mo fence="true" stretchy="true" minsize="1.2em" maxsize="1.2em"&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;/mstyle&gt;&lt;/mtd&gt;&lt;/mtr&gt;&lt;mtr&gt;&lt;mtd&gt;&lt;mstyle scriptlevel="0" displaystyle="true"&gt;&lt;mrow&gt;&lt;/mrow&gt;&lt;/mstyle&gt;&lt;/mtd&gt;&lt;mtd&gt;&lt;mstyle scriptlevel="0" displaystyle="true"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;msup&gt;&lt;mi&gt;σ&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi mathvariant="normal"&gt;Var&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mo&gt;ˉ&lt;/mo&gt;&lt;/mover&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;msup&gt;&lt;mi&gt;σ&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;msup&gt;&lt;mi&gt;σ&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mtext&gt; &lt;/mtext&gt;&lt;msup&gt;&lt;mi&gt;σ&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;/mstyle&gt;&lt;/mtd&gt;&lt;/mtr&gt;&lt;/mtable&gt;&lt;annotation encoding="application/x-tex"&gt;&#10;\begin{aligned}&#10;\mathbb{E}[S] &amp;amp;= \sum_{i=1}^{n} \mathbb{E}\bigl[(X_i - \mu)^2\bigr] - n\, \mathbb{E}\bigl[(\bar X - \mu)^2\bigr] \\&#10;&amp;amp;= n \sigma^2 - n \operatorname{Var}(\bar X) = n\sigma^2 - \sigma^2 = (n - 1)\, \sigma^2,&#10;\end{aligned}&#10;&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;/div&gt;&lt;p&gt;because independence gives &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;Var&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mo&gt;ˉ&lt;/mo&gt;&lt;/mover&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msup&gt;&lt;mi&gt;σ&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msup&gt;&lt;mi mathvariant="normal"&gt;/&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;\operatorname{Var}(\bar X) = \sigma^2 / n&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;.&lt;/p&gt;&#10;&lt;/div&gt;&#10; &lt;div class="dataleaf-proof__end"&gt;&lt;span aria-hidden="true"&gt;□&lt;/span&gt;&lt;span class="dataleaf-visually-hidden"&gt;End of proof&lt;/span&gt;&lt;/div&gt;&#10;&lt;/div&gt;</description></item><item><title>Two series, one axis each</title><link>https://dataleaf-642346.gitlab.io/posts/dual-axes/</link><pubDate>Fri, 08 Sep 2023 00:00:00 +0000</pubDate><guid>https://dataleaf-642346.gitlab.io/posts/dual-axes/</guid><description>&lt;p&gt;Few chart forms are as tempting, or as quietly misleading, as the line chart with two vertical axes. The motivation is honest: two quantities with different units seem to move together, and putting them on one chart invites the reader to compare them. The trouble is that the comparison the chart makes is not in the data. It is in the choice of axes.&lt;/p&gt;&#10;&lt;div class="dataleaf-heading"&gt;&lt;h2 id="the-crossing-is-a-design-decision"&gt;The crossing is a design decision&lt;/h2&gt;&#10; &lt;a&#10; class="dataleaf-heading-anchor"&#10; href="#the-crossing-is-a-design-decision"&#10; aria-label="Permalink to The crossing is a design decision"&#10; data-heading-anchor&#10; data-heading-copy="Copy link to The crossing is a design decision"&#10; data-heading-copied="Copied link to The crossing is a design decision."&gt;#&lt;/a&gt;&#10;&lt;/div&gt;&#10;&lt;p&gt;A line chart maps values to vertical positions. With one axis, the mapping is shared, and every visual relationship between the lines is a relationship between the numbers: one line above another means one value exceeds the other. With two axes, each series gets its own mapping, chosen independently. Where the lines sit relative to each other, whether they cross, and how steep they look are all consequences of two arbitrary pairs of numbers: the limits of the right-hand axis.&lt;/p&gt;</description></item><item><title>What a research notebook is for</title><link>https://dataleaf-642346.gitlab.io/posts/what-a-research-notebook-is-for/</link><pubDate>Fri, 21 Apr 2023 00:00:00 +0000</pubDate><guid>https://dataleaf-642346.gitlab.io/posts/what-a-research-notebook-is-for/</guid><description>&lt;p&gt;Laboratory scientists learn early that an experiment not written down did not happen. The bound notebook, with numbered pages and dated entries in ink, is one of the oldest instruments of experimental science, and it survives because it does something no other record does: it captures what the experimenter knew and intended at the moment of acting.&lt;/p&gt;&#10;&lt;p&gt;Computational work has inherited the experiments but not, in most places, the notebook. A modern analysis leaves behind a repository full of code, a directory full of outputs, and perhaps a commit history. It is easy to believe that this is a complete record. It is a record of what was done. It is almost never a record of why.&lt;/p&gt;</description></item></channel></rss>