# Text to speech licence gate — 2026-09-30

Selected: **Kokoro-82M v1.0, q8 ONNX**, English lexicon G2P from **Misaki 0.9.4**, direct plain-WASM ONNX Runtime. Gate PASS for weights and lexicons before implementation. No model export or quantisation is performed locally.

## Separate code, weights and data findings

- Weights and voice style data: **Apache-2.0**. The [immutable distributor model card](https://huggingface.co/onnx-community/Kokoro-82M-v1.0-ONNX/blob/1939ad2a8e416c0acfeecc08a694d14ef25f2231/README.md) states `license: apache-2.0`. Its Voices/Samples section includes all six selected English voices. The [original model card](https://huggingface.co/hexgrad/Kokoro-82M) also declares Apache-2.0. The host repository licence applies to its ONNX and voice files; no separate restrictive voice terms are declared. This is a licence finding, not a warranty about training data or speaker likeness rights.
- Lexicon JSON and English lexicon algorithm: **Apache-2.0**. [Misaki 0.9.4 metadata](https://pypi.org/project/misaki/0.9.4/) says “Apache Software License”. The pinned wheel includes all four English JSON files and `misaki/en.py` under its full `dist-info/licenses/LICENSE`: “Apache License”, “Version 2.0, January 2004”. No separate restrictive English data licence is declared. Files below are copied unchanged; only English lexicon/stress/suffix logic is ported. spaCy, num2words and Python code are test tooling only and do not ship.
- Small style/feed/sampling port: **Apache-2.0**. [Kokoro author's pinned LICENSE](https://raw.githubusercontent.com/hexgrad/kokoro/dfb907a02bba8152ca444717ca5d78747ccb4bec/LICENSE) says “Apache License”, “Version 2.0, January 2004”. `kokoro.js/src/kokoro.js` supplies style row selection (token count clamped to 509), dimension 256 and rate 24,000. Token IDs come directly from the pinned tokenizer. Chunking is original code, with sentence/clause/word boundaries and no truncation.
- OOV fallback: **MIT**, original letter-to-sound heuristics written for this tool, separately in `site/text-to-speech/oov-rules.js`. Its full licence is `LICENSE-oov.txt`, granting “Permission is hereby granted, free of charge”. It uses ordinary English spelling patterns, not a copied NRL rule table or any external phonemizer. As new local source it has no upstream pinned URL; the exact same-origin source and digest are recorded after implementation below. Its limitations are shown as “words we guessed”; overrides use respellings.
- MP3: **lamejs 1.2.1, LGPL-3.0**, the explicitly authorised exception, separate and UNMODIFIED. [Pinned npm metadata](https://registry.npmjs.org/lamejs/1.2.1) says `license: LGPL-3.0`. The package/LICENSE says “Link to LAME as separate jar (lame.min.js or lame.all.js)”. Full source is available in the pinned tarball; the page links source and licence. The unmodified package LICENSE is retained, along with the full GNU LGPL-3.0 and GPL-3.0 licence texts in the encoder folder. Only MP3 click loads this encoder. No eSpeak code is included under this exception.

## Immutable asset inventory

URLs of wheel members specify the exact archive followed by `#member=...`; hashes/bytes are for the extracted unchanged member. Full Apache licences are included alongside these assets.

| Shipped asset or port source | Exact pinned URL | SHA-256 | Bytes | Licence / verdict |
| --- | --- | --- | ---: | --- |
| onnx/model_quantized.onnx | https://huggingface.co/onnx-community/Kokoro-82M-v1.0-ONNX/resolve/1939ad2a8e416c0acfeecc08a694d14ef25f2231/onnx/model_quantized.onnx | `fbae9257e1e05ffc727e951ef9b9c98418e6d79f1c9b6b13bd59f5c9028a1478` | 92,361,116 | Apache-2.0 PASS |
| voices/af_heart.bin | https://huggingface.co/onnx-community/Kokoro-82M-v1.0-ONNX/resolve/1939ad2a8e416c0acfeecc08a694d14ef25f2231/voices/af_heart.bin | `d583ccff3cdca2f7fae535cb998ac07e9fcb90f09737b9a41fa2734ec44a8f0b` | 522,240 | Apache-2.0 PASS |
| voices/af_bella.bin | https://huggingface.co/onnx-community/Kokoro-82M-v1.0-ONNX/resolve/1939ad2a8e416c0acfeecc08a694d14ef25f2231/voices/af_bella.bin | `f69d836209b78eb8c66e75e3cda491e26ea838a3674257e9d4e5703cbaf55c8b` | 522,240 | Apache-2.0 PASS |
| voices/am_michael.bin | https://huggingface.co/onnx-community/Kokoro-82M-v1.0-ONNX/resolve/1939ad2a8e416c0acfeecc08a694d14ef25f2231/voices/am_michael.bin | `1d1f21dd8da39c30705cd4c75d039d265e9bc4a2a93ed09bc9e1b1225eb95ba1` | 522,240 | Apache-2.0 PASS |
| voices/am_fenrir.bin | https://huggingface.co/onnx-community/Kokoro-82M-v1.0-ONNX/resolve/1939ad2a8e416c0acfeecc08a694d14ef25f2231/voices/am_fenrir.bin | `c27989f741f7ee34d273a39d8a595cc0837d35f5ced9a29b7cc162614616df43` | 522,240 | Apache-2.0 PASS |
| voices/bf_emma.bin | https://huggingface.co/onnx-community/Kokoro-82M-v1.0-ONNX/resolve/1939ad2a8e416c0acfeecc08a694d14ef25f2231/voices/bf_emma.bin | `669fe0647f9dd04fcab92f1439a40eeb4c8b4ab1f82e4996fe3d918ce4a63b73` | 522,240 | Apache-2.0 PASS |
| voices/bm_george.bin | https://huggingface.co/onnx-community/Kokoro-82M-v1.0-ONNX/resolve/1939ad2a8e416c0acfeecc08a694d14ef25f2231/voices/bm_george.bin | `c4b235a4c1f2cd3b939fed08b899ce9385638b763f7b73a59616c4fc9bd6c9bc` | 522,240 | Apache-2.0 PASS |
| misaki/data/us_gold.json | https://files.pythonhosted.org/packages/82/ec/0ee4110ddb54278b8f21c40a140370ae8f687036c4edf578316602697c56/misaki-0.9.4-py3-none-any.whl#member=misaki/data/us_gold.json | `dc414872a49a28ae6c141463d502fd945f3b2fde040484fdc47d00cc4612686f` | 3,000,469 | Apache-2.0 PASS |
| misaki/data/us_silver.json | https://files.pythonhosted.org/packages/82/ec/0ee4110ddb54278b8f21c40a140370ae8f687036c4edf578316602697c56/misaki-0.9.4-py3-none-any.whl#member=misaki/data/us_silver.json | `de8f67be911bb6c659187b4a65fd966b6a30e56350e0f790d763210b053ac475` | 3,099,517 | Apache-2.0 PASS |
| misaki/data/gb_gold.json | https://files.pythonhosted.org/packages/82/ec/0ee4110ddb54278b8f21c40a140370ae8f687036c4edf578316602697c56/misaki-0.9.4-py3-none-any.whl#member=misaki/data/gb_gold.json | `29e62f4b60261c88f7f3c2c7811ca3825978948090b72d2b27d565b729282f71` | 2,838,552 | Apache-2.0 PASS |
| misaki/data/gb_silver.json | https://files.pythonhosted.org/packages/82/ec/0ee4110ddb54278b8f21c40a140370ae8f687036c4edf578316602697c56/misaki-0.9.4-py3-none-any.whl#member=misaki/data/gb_silver.json | `48131e2d92ccc41655f4543e87e0f938e71463eb5a54be7f0693bb712ebb6bce` | 3,663,898 | Apache-2.0 PASS |
| misaki/en.py | https://files.pythonhosted.org/packages/82/ec/0ee4110ddb54278b8f21c40a140370ae8f687036c4edf578316602697c56/misaki-0.9.4-py3-none-any.whl#member=misaki/en.py | `ddf67ad3bc4dd98143dcc9b6fbcde259b9d879d8b91684201cf0deeb07aa9910` | 31,278 | Apache-2.0 PASS |
| tokenizer.json | https://huggingface.co/onnx-community/Kokoro-82M-v1.0-ONNX/resolve/1939ad2a8e416c0acfeecc08a694d14ef25f2231/tokenizer.json | `77a02c8e164413299b4b4c403b14f8e0e1c1b727db4d46a09d6327b861060a34` | 3,497 | Apache-2.0 PASS |
| kokoro.js/src/kokoro.js | https://raw.githubusercontent.com/hexgrad/kokoro/dfb907a02bba8152ca444717ca5d78747ccb4bec/kokoro.js/src/kokoro.js | `97cbc36e4ab00fea35d3f8ca3a23ae199f787069f83774bb86e7886321c92e68` | 5,786 | Apache-2.0 PASS |
| lame.min.js | https://registry.npmjs.org/lamejs/-/lamejs-1.2.1.tgz#member=package/lame.min.js | `15d285e2587b3bdbfd18a68de6ce07cc074f7480a82c3815da2dc1c348ec6df4` | 156,043 | LGPL-3.0 exception PASS |

## Rejected candidates and storage

`kokoro-js` 1.2.1 is Apache-2.0 but its [exact npm metadata](https://registry.npmjs.org/kokoro-js/1.2.1) depends on `phonemizer: ^1.2.1`. `phonemizer` 1.2.1 [metadata](https://registry.npmjs.org/phonemizer/1.2.1) claims Apache-2.0, while its pinned distribution `dist/phonemizer.js` contains `/usr/share/espeak-ng-data`, `eSpeakNGWorker` and compiled eSpeak bindings. [eSpeak NG COPYING](https://github.com/espeak-ng/espeak-ng/blob/master/COPYING) says “GNU GENERAL PUBLIC LICENSE”, “Version 3”. REJECTED; not vendored. Misaki's optional eSpeak and neural OOV paths are not used. No Piper, transformers.js or num2words is shipped.

The q8 model is **92,361,116 bytes**, above the 25 MB tree limit. It is ignored by one explicit .gitignore entry and obtained only by `scripts/fetch_models.py`, which verifies size and SHA-256 before atomic replacement; `--check` never downloads. The browser fetches only same-origin `/vendor/kokoro/` files on first Speak, verifies SHA-256, and uses Cache API keyed by each asset's hash. Only the selected accent's gold/silver and selected voice load. Each voice is 522,240 bytes; each JSON is below 25 MB. q8f16 was checked on the same HF revision: 86,033,585 bytes, not selected.

## Runtime finding

Existing **onnxruntime-web 1.17.3 (MIT)** was checked first, in a real Chromium Worker with the pinned q8 bytes. It failed during `InferenceSession.create`, throwing an opaque numeric WASM exception before inference. Both normal and verbose attempts failed. A separate native ORT 1.17.3 diagnostic against the unchanged model identified the cause: the model imports `ai.onnx.ml` opset 5, whereas ORT 1.17.3 supports that domain only through opset 4. No model imports were rewritten. The existing background-remover runtime was not changed.

Selected **onnxruntime-web 1.22.0 (MIT)** for this tool only, from `https://registry.npmjs.org/onnxruntime-web/-/onnxruntime-web-1.22.0.tgz`, archive SHA-256 `e293551f9d36d003e293d0f2937fee7b3dfdcef76d71d1449b55e2d4157c7aea`, 20,387,966 bytes. [Microsoft pinned LICENSE](https://raw.githubusercontent.com/microsoft/onnxruntime/v1.22.0/LICENSE) says “MIT License”; full unmodified licence is in `../onnxruntime-1.22.0/LICENSE` (the npm tarball itself does not include a LICENSE file). Plain WASM, no JSEP/WebGPU/WebGL, no transformers.js. This release calls its plain binary `ort-wasm-simd-threaded.wasm`; `numThreads = 1` uses it in single-thread mode without SharedArrayBuffer or COOP/COEP. The companion .mjs is required by the runtime's WASM loader. All three files below are copied unchanged. The first successful real browser probe produced 39,000 finite audio samples at 24 kHz, with inputs `input_ids`, `style`, `speed` and output `waveform`, in 5.975 s including model/runtime loading.

| Runtime asset | Pinned source member | SHA-256 | Bytes | Licence |
| --- | --- | --- | ---: | --- |
| ort.wasm.min.js | https://registry.npmjs.org/onnxruntime-web/-/onnxruntime-web-1.22.0.tgz#member=package/dist/ort.wasm.min.js | `65e09376df69107e881b5c34d2d37aed333a366b6d941073ec518168e269b87d` | 48,327 | MIT |
| ort-wasm-simd-threaded.mjs | https://registry.npmjs.org/onnxruntime-web/-/onnxruntime-web-1.22.0.tgz#member=package/dist/ort-wasm-simd-threaded.mjs | `30dd851d9c00622940500f71ddd2ff8820c5cb65270816080175b958705385a8` | 20,856 | MIT |
| ort-wasm-simd-threaded.wasm | https://registry.npmjs.org/onnxruntime-web/-/onnxruntime-web-1.22.0.tgz#member=package/dist/ort-wasm-simd-threaded.wasm | `71aef04959c5c1b6de461b6538e2058e306610034a85aad2742d0c7fd4533fe4` | 11,210,254 | MIT |

The reused export muxers are **mp4-muxer 5.2.2** and **webm-muxer 5.1.4**, MIT. Exact npm sources and full licence quotes are in [the existing engine inventory](../VIDEO-ENGINE.md). They load only on M4A/WebM click, only the browser-selected format; they are not modified here. WebCodecs encoders are supplied by the browser/OS and not redistributed.

## Original fallback and local port fingerprints

Local source URLs are same-origin reviewable files, not claims of already-published upstream revisions. No immutable upstream URL exists for new code; these SHA-256 values identify the exact local implementation. Apache port source bytes above identify the upstream inputs.

| Local source | Exact same-origin URL | SHA-256 | Bytes | Licence |
| --- | --- | --- | ---: | --- |
| oov-rules.js | /text-to-speech/oov-rules.js | `79da753da3dce1a1ac8415ad309b11ade9ed0ee90088a6c4c9e2ac18489ff539` | 2,186 | MIT |
| g2p-core.js | /text-to-speech/g2p-core.js | `049918a9502df0fc331ac74e726581a4b30ce8d360f0178a659a46982efa941e` | 10,682 | Apache-2.0 |
| tts-core.js | /text-to-speech/tts-core.js | `d39f58e6eb1cc5eaed9d7df9e49d7994368ecd715841f1064f903ce508b18a5e` | 3,392 | Apache-2.0 |
| mp4-muxer.mjs | https://registry.npmjs.org/mp4-muxer/-/mp4-muxer-5.2.2.tgz#member=package/build/mp4-muxer.mjs | `d2c4c782f180c86ed30b1f5d9487a34a0d370bf9b4535285734ea187d38f9bb5` | 69,011 | MIT; see ../VIDEO-ENGINE.md |
| webm-muxer.mjs | https://registry.npmjs.org/webm-muxer/-/webm-muxer-5.1.4.tgz#member=package/build/webm-muxer.mjs | `53710543cecc4db7b6cd23fb0d9b308aa3374697f43a6149a8497f0a8547f0cb` | 64,944 | MIT; see ../VIDEO-ENGINE.md |

## Verification and practical limits

Golden fixtures contain 300 distinct varied English sentences in each accent, generated with unmodified Misaki 0.9.4, `en_core_web_sm` 3.8.0 and `fallback=None` in an outside-repo throwaway venv. Full outputs, token spellings and POS tags are retained for inspection. The browser uses no tagger; DEFAULT heteronym disagreements remain in the denominator. Unknown words, punctuation and numeric/abbreviation expansion are tested separately. Exact per-word equality: **5,452 / 5,468 = 99.70738844184345%**, US 2,726 / 2,734 and UK 2,726 / 2,734. This fixture agreement is not a universal pronunciation-quality claim; unknown names, foreign words and contextual heteronyms can need fixes.

All real browser requests are same-origin GETs under `/text-to-speech/` or `/vendor/`; no text or audio is uploaded. AdSense markup follows the site's normal dormant tag/injection pattern. Deployment-time AdSense activation is the existing site's policy, not part of the speech runtime.

HEAD/GET/304 cache headers for `.onnx`, `.bin` and `.json` match the existing vendor policy at this HEAD: `public, max-age=300`. No shared server code is changed. Hash-keyed Cache API provides model/voice/dictionary reuse beyond that HTTP freshness window when storage is available. The page's before-Speak vendor download is **0 bytes**.

## Chromium measurements — 2026-09-30

Real HTTP on localhost, Chromium 150 on Linux, single-thread plain-WASM ORT, fresh browser profiles, default speed 1.0. Test sentence: “Hello world. This is a test of the voice tool.” Each run tested af_heart (US) and bf_emma (UK), actual WAV/MP3/WebM-Opus downloads, Cache API reload, worker termination, phone-width layout, keyboard Speak and same-origin GETs. No fetch or inference stub was used. Each page loaded **27,180 bytes before Speak**, of which **0 bytes were vendor/model/dictionary/voice/encoder assets**.

| Repeat | Tests | Suite seconds | US first audio / complete seconds | UK first audio / complete seconds |
| --- | --- | ---: | --- | --- |
| 1 | 10/10 OK | 51.444 | 7.9836 / 14.7560 | 5.5757 / 12.7842 |
| 2 | 10/10 OK | 57.165 | 8.8301 / 16.1802 | 6.2342 / 14.1616 |
| 3 | 10/10 OK | 55.812 | 9.1244 / 16.1216 | 6.4274 / 15.4008 |

US WAV: 24 kHz, mono, 16-bit PCM, **3.975 s**, RMS **0.0636446050**, all source samples finite. UK WAV: **4.375 s**, RMS **0.0853720105**, all source samples finite. MP3: **4.032 s**, within 1.44% of the US WAV. This Chromium exposes Opus and no AAC encoder; WebM-Opus was checked with ffprobe. AAC/M4A selection is capability-based, but the AAC export path and Safari/Edge were not exercised on this host. A human has not yet checked speech intelligibility; the outside-repo sample is `/tmp/apitc-61c-scratch/tts-sample-us.wav`.

HTTP cache policy is the existing site's five minutes (not an immutable year-long header); long-term model reuse uses Cache API. Shared navigation adds one link per eligible tool; qr and auction-organizer retain their existing exemptions. video-editor is also narrowly exempted from this new link because ticket 61 explicitly prohibits touching that concurrently edited folder. No server route, paid-key code, MCP tool, skill-pack entry or issue file changed. Changes remain uncommitted.

Final whole-suite command: `/home/g5-local/ventures/apitc/core/.venv/bin/python -m unittest discover -s tests`. **592 tests in 388.680 s — OK, 0 failures, 0 errors, 0 skips**. Before that final run, the worktree's existing converter dependencies were installed with `npm ci --prefix converter --ignore-scripts` and its ignored `.cowerx/` scratch directory was created. No tracked converter or server code changed. `scripts/fetch_models.py --check`, all asset/port hashes and `git diff --check` passed. No commit, push or deployment was performed.
