Toward Paper 2 · progress report · 9 October 2026

Teaching the model to hear: what passed, what failed, and what it costs on one laptop

Paper 1 showed a retina driving nine cortical columns and a cochlea driving four, with no learning at all. The next stage asks two questions: can a hippocampal region bind what the eye sees to what the ear hears, and can the ear be made good enough that a word sounds the same across voices? This is a status report, not a result. The stage is not accepted. Here is exactly what holds, what does not, and what each experiment costs.

Status: letter–tone association 9 of 9 in both directions on a fixed course, frozen control 0 of 9; unknown voices not rejected; stage open. Paper 1: research.html, DOI 10.5281/zenodo.23119475.

What was built in October

A human ear

The Paper 1 cochlea was a bank of 32 gammatone channels. The new one has 1,000 cochlear sections and 32,000 individual auditory nerve fibres, takes sound pressure in pascals at the eardrum at 100,000 samples per second, and spikes. Compression, adaptation, phase locking and recovery after 0.1, 0.3 and 1.9 seconds are checked against published windows. Frequency tuning was wrong at first: the maximum error in the equivalent rectangular bandwidth was 52.3 %. After the fix it is 12.8 %. Thirty unit and reference tests pass. The verification file says, in its own field names, human_ear_complete: false.

A cochlear nucleus cell at native resolution

The bushy cell between the ear and the cortex is now a compartment model of 357 sections and 1,429 segments stepped at 1.25 µs. 300 ms of it is 559,899,000 scalar samples, and a run can be stopped and restored from a checkpoint without storing the input history: the restored suffix matches the uninterrupted run exactly. Restore takes 0.017 s; cold initialisation takes 120.6 s. That difference is why the checkpoint exists.

Paired learning

Three letter–tone pairs, three checks each, so 9 of 9 means nine checks, not nine letters. On 2 October the course gave tone→letter 6 of 9 and letter→tone 0 of 9. On 3 October both directions reached 9 of 9 with plastic synapses (169,925 weight updates), while the frozen control recalled 0 of 9 with 0 weight updates and there were no false recalls on baseline or unknown input. A training run interrupted and restarted from its checkpoint produces the same output as the uninterrupted one.

What does not hold

Unknown voices

The full speech corpus is 54 recordings, 80,941,578 spike events per pass. An external diagnostic recognises all 28 known words, including 8 from other speakers. It also recognises all 28 with learning switched off. So 28 of 28 is not evidence that the hippocampal region learned anything; it is evidence that the diagnostic is lenient. Worse, of 8 words the model never heard, 5 score at or above the weakest known word with learning on, and 8 of 8 with learning off. The separation margin is negative in both modes. A model that cannot say "I do not know this word" has not learned words. The verification file says full_service_accepted: false.

Candidates rejected on the way

Each candidate has its own folder with the run, the checks and the SHA-256 of the source files that produced it, and none of them touched the accepted 9 of 9, whose 37 reference hashes are checked after every experiment.

Why the failures are published

Because the alternative is the usual one: tune until green, publish the green. The acceptance rule for this stage was written before the runs: 9 of 9 in both directions, no false recalls, retention over time, transfer to new voices. The first three hold on the fixed course. The fourth does not, and the stage stays open until it does. Every experiment ends with the same three lines: what passed, what was rejected, and that the original hashes are unchanged.

What this costs

All of it runs on one MacBook Pro M2 Max with 64 GB. One pass over the 54 recordings is 80.9 million spike events. The cochlear nucleus cell alone needs two minutes to initialise before it computes anything. The columns run at scale 0.1 because a full Potjans–Diesmann column is 77,169 neurons and about 300 million synapses, 6 GB at 20 bytes per synapse. On one machine this is three to four orders of magnitude below a human cortex, not one.

What comes next

A physiological profile of the cochlear nucleus from independent human data, the missing relays between the cochlear nucleus and the cortex, and the acceptance on new voices. These are compute-bound. When the stage is accepted, it becomes Paper 2 in the same format as Paper 1, with its seed sweeps and its red cells.

Data

The files below are unmodified copies of the project's verification artifacts; each carries the SHA-256 of the source files that produced it. The step 2 tree is not yet committed; it will be frozen and tagged with the Paper 2 report, as Paper 1 was.