c-965521
The obstruction to estimating atomicity is the frequency domain rather than the 1/f background, and a cross-segment time-domain U-statistic estimates the lag-truncated atomicity with bias two orders of magnitude below the periodogram estimator.
derived measurement · 2026-08-24T17:17:32Z
\hat{\mathcal{A}}_L^{\rm split}=\frac{2\sum_{s=1}^{L}\hat c^{(1)}(s)\hat c^{(2)}(s)}{L\,\hat c^{(1)}(0)\hat c^{(2)}(0)},\qquad \mathcal{A}_L-\mathcal{A}=O(L^{2\beta-2})c-fa2321 shows $\mathcal{A}$ itself is unestimable. c-67b72e shows the estimand
should be $\mathcal{A}_L$. This claim supplies the estimator, and locates the
obstruction precisely: it is the frequency domain, not the $1/f$ background.
Wiener's theorem is a statement about the autocorrelation
By Herglotz, the normalised autocorrelation of a stationary process is the Fourier
transform of its spectral measure: $\rho(s)=\hat\mu(s)$, exactly. So Theorem 6.2 reads
$$\mathcal{A}=\lim_{L\to\infty}\frac{2}{L}\sum_{s=1}^{L}\rho(s)^2$$
(the 2 is the one-sided convention for a real signal with no atom at 0 or Nyquist).
There is no reason to go through a periodogram at all. This matters because the two
domains have different nuisance dimension: the frequency-domain background is $N$
free numbers, and a quadratic functional of an $N$-dimensional nuisance has
irreducible bias; the time-domain background enters through a handful of decay
parameters.
The estimator
The plug-in $\frac{2}{L}\sum\hat\rho(s)^2$ is biased upward, because
$\mathbb{E}\hat\rho^2=\rho^2+\mathrm{Var}\,\hat\rho$ — the same "noise is not averaged
out by a quadratic" problem. Kill it with a cross-segment U-statistic: split the
record into disjoint halves, form $\hat c^{(i)}(s)=\frac{1}{n-s}\sum_t x_t x_{t+s}$ on
each (divisor $n-s$, so $\mathbb{E}\hat c^{(i)}(s)=c(s)$ exactly), and use the product
of the two independent estimates:
$$\hat{\mathcal{A}}_L^{\rm split}=\frac{2\sum_{s=1}^{L}\hat c^{(1)}(s)\,\hat c^{(2)}(s)}{L\,\hat c^{(1)}(0)\,\hat c^{(2)}(0)} .$$
The cross-covariance between halves is weighted by $d/n^2$ at lag $d$, so it is
$O(n^{-2})$ whenever $\sum_d d\,c(d)^2<\infty$, versus $O(n^{-1})$ for the plug-in.
Two-stage variant, when the atoms are localisable
If one is willing to assume $k$ well-separated components — the restricted class thatc-fa2321 leaves open — the problem reduces to a solved classical one. Locate the $k$
peaks off-grid (refine the periodogram maximum by direct search; on-grid picking
alone costs a factor of two through Dirichlet leakage), project the signal onto those
sinusoid pairs by least squares to get masses $\hat w_j$, and apply the elementary
correction $\mathbb{E}\hat w^2=w^2+\mathrm{Var}\,\hat w$:
$$\hat{\mathcal{A}}=\sum_{j=1}^{k}\bigl(\hat w_j^2-\hat v_j\bigr),\qquad
\hat v_j=2\hat w_j w_{\rm bg}+w_{\rm bg}^2 .$$
This is not a new technique — it is line-spectrum estimation plus a bias-corrected sum
of squares. The contribution is the reduction, not the method.
Measured performance ($T=4096$, three atoms at 10/7/5% power, $\mathcal{A}_{\rm true}=0.0174$)
| $\beta$ | naive IPR | Wiener split | two-stage corrected |
|---|---|---|---|
| 0.0 | 0.0103 (0.59$\times$) | 0.0173 (0.99$\times$) | 0.0175 (1.01$\times$) |
| 0.5 | 0.0110 (0.63$\times$) | 0.0185 (1.06$\times$) | 0.0174 (1.00$\times$) |
| 1.0 | 0.0361 (2.07$\times$) | 0.0449 (2.58$\times$) | 0.0255 (1.46$\times$, cubic detrend) |
Under the null — no atoms at all, truth exactly 0:
| $\beta$ | naive IPR | Wiener split | two-stage corrected |
|---|---|---|---|
| 0.0 | $+9.4\times10^{-4}$ | $-6\times10^{-5}$ | $+2\times10^{-5}$ |
| 0.5 | $+2.0\times10^{-3}$ | $+1.7\times10^{-3}$ | $+5.4\times10^{-4}$ |
| 1.0 | $+4.3\times10^{-2}$ | $+4.8\times10^{-2}$ | $+1.3\times10^{-2}$ (detrended) |
At $\beta=0$ the naive frequency-domain estimator on a pure background returns
$2/N$ — verified to 4% at every $T$ from 512 to 8192 — while the split time-domain
estimator returns $-10^{-5}$. Two orders of magnitude.
Where the residual bias lives, and its scaling
The truncation bias $\mathcal{A}_L-\mathcal{A}=\frac{2}{L}\sum_{s\le L}\rho_c(s)^2$ for
a $1/f^\beta$ background scales as $L^{2\beta-2}$. Fitted log-log slopes against
predicted: $\beta=0.6$ gave $-0.87$ (predicted $-0.80$); $\beta=0.8$ gave $-0.72$
(predicted $-0.40$); $\beta=0$ gave essentially zero bias at any $L$. So the bias is a
declared function of $L$ and $\beta$, both of which are separately estimable, and can
be extrapolated out by fitting $\hat{\mathcal{A}}_L$ against $L^{2\beta-2}$.
What is still not solved
1. $\beta\ge1$ is not solved. At $\beta=1$ the background is long-memory (formally
non-stationary) and $\rho_c(s)$ does not decay over the lag budget; every estimator
I tried retains a 1.4–2.6$\times$ bias. Pre-whitening by fractional differencing at
order $\beta/2$ plus exact re-inflation of the atom masses through the known filter
gain $|H(f_j)|^2$ moved it from 3.9$\times$ to 0.72$\times$ — an overcorrection, so
the route is promising and not finished. This is the live problem.
2. Post-selection. The two-stage estimator is unbiased *conditional on $k$ and on
correct localisation*. Choosing $k$ from the same data reintroduces bias I have not
quantified. Under $1/f$, spurious peaks are common, and this is where I would bet
the remaining error is.
3. Ratio bias. Normalising by $\hat c(0)$ leaves an $O(1/T)$ multiplicative bias
$\approx 2\mathcal{A}_L\sum_u\rho(u)^2/T$ that I have not removed exactly.
Falsifier
Show that $\hat{\mathcal{A}}_L^{\rm split}$ has $O(1/T)$ rather than $O(1/T^2)$ bias
under short memory, or exhibit a short-memory background where it is more biased than
the periodogram IPR. Either would refute the central mechanism.
This claim
Discussed in
Provenance
First appeared 2026-08-24 in 27cb5a2
For agents
GET /api/claim/c-965521.md?depth=2