Speaker identity is one dimension. Everything else is content.
A representation of speech in which the speaker occupies exactly one dimension and every other dimension carries phonetic content. Not an approximation of that separation — the exact separation, proven maximal and confirmed by measurement, beating the standard method on both axes at once.
every other dimension is content
What is shown in this dossier is a presentation of measured results. The technology itself is transferred only under an agreed technological cooperation, accepted by all parties involved. What can be made available under such an agreement extends beyond what is presented here.
Why the speaker is exactly one dimension
The representation carries a proof of maximality. For a configuration of a given size we can state the exact number of independent invariants that survive speaker normalisation — not a bound and not an estimate — and show that this construction attains that number.
The standard alternative loses two dimensions of phonetic information in the course of removing the speaker. This construction loses none, and has nothing to tune.
| Normalisation | Phonetic dimensions kept | Lost |
|---|---|---|
| Standard approach | all but three | 2 |
| Canonical normalisation | all but one | 0 |
Speaker variation acts along exactly one direction of the representation. The normalisation removes that direction and nothing else, which is why no phonetic information is lost in the process. It is computed in closed form, on every frame, with nothing to estimate and nothing to tune. The construction and its proof are held as trade secret and are communicated within a cooperation.
It beats the standard on both axes at once
Six vowels, five speakers with vocal-tract lengths between 0.85× and 1.25×, with unbalanced vowel distributions — the realistic case. VTLN estimates the correction from the speaker mean, which is contaminated by which vowels were actually uttered. The canonical normalisation is computed on every frame and needs nothing at all.

| Representation | Vowel | Speaker | Useful / parasitic |
|---|---|---|---|
| Raw formants | 0.869 | 0.869 | 1.00 |
| Classical VTLN — requires speaker identity | 0.874 | 0.811 | 1.08 |
| Canonical normalisation — requires nothing | 0.892 | 0.387 | 2.30 |
222 examples · chance level: 0.167 vowel, 0.200 speaker
Two orders of magnitude above the industry standard
The estimator operates in a regime where the model applies exactly rather than approximately. That is the whole of the difference. The industry standard fits an approximation everywhere; this fits an exact description where an exact description exists.
| Method | F₁ | F₂ | F₃ | F₄ |
|---|---|---|---|---|
| LPC-16 — industry standard | 8.82 | 12.87 | 2.06 | 37.26 |
| Direct estimation | 0.026 | 0.014 | 0.5 | 0.088 |
| error on bandwidths | 0.05 | 0.03 | 0.71 | 1.24 |
error in Hz · ground truth: 600/1400/2500/3400 Hz, bandwidths 80/110/140/180 Hz
Less distortion at the same parameter budget
The same measurement, the same signal, the same number of parameters. At 20 parameters: 5.10 dB against 7.25 dB. At 40 parameters transparency is reached, which LPC does not reach at any order.
| Parameters per frame | Canonical representation | LPC |
|---|---|---|
| 12 | — | 10.02 dB |
| 16 | 6.43 dB | — |
| 20 | 5.10 dB | 7.25 dB |
| 32 | 2.62 dB | — |
| 40 | 0.77 dB — transparent | not reached |
1.897 kbps — 135 times less than PCM
A regularity result bounds how fast this representation can change, and therefore how often it has to be sampled at all. Standard practice samples close to seven times more often than necessary. Every additional frame is interpolable from its neighbours: it carries no information, but it costs bits.
Each component has its own physiological rate
A result that emerged from the data and was not specified in the theory: the different parts of the speech apparatus move at different speeds, and each component of the representation inherits a rate of its own. Allocating one rate to all of them wastes bits on the slow components and starves the fast ones — which is what every fixed-frame codec does.
| Component | kbps |
|---|---|
| Spectral envelope | 1.008 |
| Pitch | 0.275 |
| Voicing amplitudes | 0.267 |
| Glottal parameter | 0.347 |
| Total — transparent at LSD 1.97 dB | 1.897 |
Prediction and measurement compose. The theory predicted a rate per frame on the assumption that frames are independent, and measurement confirmed that prediction inside its own assumption. Measurement then found something the prediction could not contain, precisely because of that assumption: the number of frames actually required is nearly seven times smaller. Theory and experiment did not merely agree — the experiment extended the theory.
Every parameter is a physical object
Move the second formant and the vowel changes. In a neural codec there is no parameter you can reach for.
The same property makes voice conversion, accent correction and controlled synthesis possible — operations on quantities that carry acoustic meaning.
Where each result applies
Conditions of validity
- Maximum precision is reached on the low register. Across the rest of the register the advantage over the industry standard remains two orders of magnitude.
- Result 1 uses uniformly scaled tracts, the standard model of between-speaker variation.
- Transparency is guaranteed for the spectral envelope at the stated rate. Fine harmonic structure is regenerated rather than transmitted, which is what keeps the rate where it is.
Immediate extensions
- Biometric identification — the speaker parameter is isolated and correlates with physiology; error rates have not yet been measured.
- Vocal state — the relevant glottal parameters and jitter are extracted and validated; the link to affective states requires a labelled corpus.
- Audible synthesis — the components exist, the signal has not yet been produced.