Proprium Bioscience

August 5, 2026

Searching for a Signal

Left, the mean impedance change each base produces across the sweep. Right, per-base accuracy from published (open circle) to current (arrow). Simulator data throughout.
Left, the mean impedance change each base produces across the sweep. Right, per-base accuracy from published (open circle) to current (arrow). Simulator data throughout.

Our basecaller previously read 71.2% of bases correctly on 30-base reads, or Phred Q5.4. We argued that gap was closable because the error had structure rather than being a noise floor, and we named the route, a fourth discrimination axis aimed at the pair of bases the existing three leave tangled.

We found the fourth axis, and the model’s accuracy has improved to 79.53%, Q6.89, while using 24 fewer measured frequencies. We’re pushing closer to our Q20 production target, with six errors in a 30-base read rather than nine, and a six-fold gain in the fraction of reads that can be placed against a reference.

Where We Were

An incorporation disturbs the electrode-electrolyte interface, and the sweep splits that disturbance across frequency into three contributions. The first is the dwell the polymerase spends in its closed conformation, which shows up in the platinum double-layer band. The second is the dielectric coupling of the incorporated nucleotide to the gold above it. The third sits near the guanine oxidation onset, a charge-transfer term that only guanine triggers.

The difficulty is that the first two are close to the same measurement. Dwell and dielectric coupling both track roughly how polarizable the incorporated base is, so they sort the four bases into the same order, and both act on the sweep as a multiplier. A slightly longer dwell looks much like a slightly larger base. The third contribution behaves less like an axis than like a switch, since it responds to guanine and to nothing else. That left us with one usable direction and one detector, which is not enough to separate four bases.

We can see it directly in the features. Once they are rescaled so that a unit of distance means a unit of event-to-event noise, 95.3% of the separation between the four bases lies along a single direction, so the bases sit close to a line rather than spread through the space. Thymine ends up in the narrowest region on that line, which is why it is involved in roughly four of every five errors we make. Guanine, the one base with a discriminant of its own, was already at Q21.2 and at production grade.

The limit was therefore in the measurement rather than in the model, and we tested that before accepting it. A dedicated cytosine versus thymine head, built the same way as the guanine head that works, moved thymine by two thousandths. No amount of decoder work recovers a dimension the features never carried.

What a Fourth Axis Had to Be

A fourth axis is only useful if the sweep can tell it apart from the three we already had. An impedance sweep can separate two processes only if they act at different speeds. One that settles quickly and one that settles slowly bend the curve in different places, so we can read them off separately. Two processes that differ only in strength give the same curve at two heights, and no number of extra frequencies will pull those apart.

We had two heights and a switch. What we needed was a bend.

What this shows. Four bases, four curves. Cytosine and thymine differ by 2.7% in overall size and almost not at all in shape, which is why they are the pair we confuse.

What to try. Raise the gain jitter until the two envelopes cover each other, then add a second contribution and move its speed across the sweep.

1001k10k100kfrequency, HzΔlog|Z|
C and T, told apart by size2.7%
a gain of that size or more erases this
C and T, told apart by shape2.3%
no gain changes this number at all
ACGT

Two heights, and only a hair between them. Cytosine and thymine leave nearly the same trace at slightly different heights, so a gain that varies from one event to the next is enough to swap them. Height is the only thing a gain can change, which is why a fourth contribution has to bend the curve somewhere instead. Move its speed past either edge of the sweep and the advantage disappears, because a process the window cannot resolve is a process we cannot read. The second contribution demonstrates the principle rather than the axis we adopted, on illustrative weights and frequency.

The axis is non-faradaic, so it needs no elevated bias and carries none of the monolayer exposure the guanine axis does. It is a relaxation, a polarization that has to reorganize and takes a characteristic time to do it, so it enters the sweep as a shoulder where there was none rather than as a change in overall level. It is also structural, reading a difference between the two pyrimidines that is present in every duplex and survives base pairing.

That last requirement ruled out our first candidate. The difference it relies on is real and well documented in a free nucleotide, but the site it depends on is the same site the base uses to pair, so it is no longer there to read once the base is incorporated, and the sensor only ever meets the paired state. The candidate that survived sorts the four bases in almost the opposite order from polarizability, and puts cytosine and thymine at opposite ends. That came out of the structure rather than anything we tuned for, and it is worth 7.2 points on its own. The share of the separation sitting on a single direction fell from 95.3% to 77.0%, and the four bases came off the line.

Two smaller changes went in alongside it. Separating shape from scale at the input was worth 0.6 points, since base identity and per-event gain jitter both amount to scaling the same shape by a scalar, and the unwanted variation runs 93% parallel to the signal. Normalising the sweep and handing the model the scale alongside it, rather than mixed into it, improved the limiting pyrimidine separation by a third. Framing the acquisition stream faster was worth another 4.4 points. The decoder reads a free-running stream cut into frames, and an incorporation that never dominates a frame is a deletion nothing can recover, which caps accuracy before the model runs, so halving the frame took the unresolvable fraction from 4.4% to 0.5%.

The published configuration read 71.20% of bases correctly, Q5.40, or 8.6 errors in a 30-base read. An interim run carrying the clock and input fixes but no new axis reached 72.30%, Q5.58, and 8.3 errors. The current configuration reads 79.53%, Q6.89, and 6.1 errors per read, on 38 tones rather than 62, and tone placement mattered more than tone count in every test we ran. Per base, thymine moved from 0.450 to 0.711 and cytosine from 0.652 to 0.777, while guanine held at 0.987 against a published 0.988. Adenine got worse, 0.753 down to 0.703, which is not noise and is the clearest signal we have about what to do next.