EMA Lightning ONNX
ONNX conversion of EMA Lightning, the 8.6M-parameter Turkish text-to-speech model, for running without PyTorch: in the browser with onnxruntime-web (WebGPU or WASM), or anywhere else ONNX Runtime runs. Output is 48 kHz mono.
Live demo: https://ema-lightning-web.vercel.app/ (source: https://github.com/ozcancelik/ema-lightning-web)
In the browser on an Apple laptop GPU, first audio arrives in about 140 ms and speech is made 20–30× faster than real time.
EMA Lightning'in ONNX'e dönüştürülmüş hâli: PyTorch olmadan, tarayıcıda (WebGPU / WASM) ya da ONNX Runtime'ın çalıştığı her yerde Türkçe konuşma üretir.
Files
| File | Size | What it is |
|---|---|---|
text.onnx |
4.8 MB | letters → hidden states and per-letter durations |
sound.onnx |
18.1 MB | aligner and the four distilled flow steps → latent frames |
fp16/sound.onnx |
9.6 MB | the same graph in fp16, fp32 inputs and outputs |
decoder.onnx |
12.0 MB | latent frames → waveform |
meta.json |
vocabulary, flow times, latent size, hop, sample rate | |
normalizer.wasm |
1.2 MB | normalizer-tr as WebAssembly: "14:30" → "on dört otuz" |
web/tts.js, web/normalizer.js |
the reference JavaScript engine used by the demo | |
export/ |
export, fp16 conversion and verification scripts | |
normalizer-wasm/ |
Rust source of the WebAssembly wrapper |
Graphs
All graphs have batch size 1. L is the number of letters in a piece of text, T the number of 25 Hz frames.
| Graph | Inputs | Outputs |
|---|---|---|
text |
ids int64 [1, L] |
h float32 [1, L, 224], dur float32 [1, L] (frames per letter) |
sound |
h [1, L, 224], cg float32 [1, L], fg float32 [1, T], cw int64 [1, L], fw int64 [1, T], noise float32 [1, 4, T, 64] |
latents float32 [1, T, 64] |
decoder |
z float32 [1, 64, T] |
audio float32 [1, T × 1920] |
meta.json:
{"vocab": ["<pad>", "<unk>", " ", "!", ...], "times": [0.0, 0.25, 0.5, 0.75],
"latent_dim": 64, "hop": 1920, "sample_rate": 48000}
Pipeline
The graphs hold only the network: convolutions, matmuls and elementwise ops. The bookkeeping between them, which
needs gathers and scatters, is done by the caller. web/tts.js does all of it in about 300 lines, and
export/verify.py does the same in NumPy.
- Text. Normalize with
normalizer.wasm(the "fallback" ambiguity policy), lowercase with Turkish rules (İ→i,I→ı), strip accents except from Turkish letters, and replace anything outsidevocabwith a space. - Chunk into pieces of at most ~10 seconds, cut at sentence, then clause, then word boundaries. Each piece ends with punctuation.
text:idsare indices intovocab.- Plan the frames. Divide
durby the speed, sum it per word, round to whole frames (1 to 250 per word, 3000 per piece).cwis each letter's word index andfweach frame's word index.cgandfgare the letter and frame positions measured in letters (seeplan()inweb/tts.js). soundwith Gaussiannoiseof shape[1, 4, T, 64]. The seed sets the voice's variation.decoderin windows (one second first, then four seconds), with 8 frames of context on each side. Each window's audio is cut out of the middle of its output.
Use in the browser
Copy web/tts.js and web/normalizer.js into your project and point EMA.load at this repository. Hugging Face
serves the files with CORS, and tts.js keeps them in Cache Storage after the first download.
import { EMA, wavBlob } from "./tts.js";
const base = "https://huggingface.co/ozcancelik/ema-lightning-onnx/resolve/main/";
const tts = await EMA.load("webgpu", base); // or "wasm"; add precision "fp16" for the smaller sound graph
const chunks = [];
for await (const audio of tts.stream("Merhaba, saat 14:30'da görüşelim.", { speed: 1, seed: 0 })) {
chunks.push(audio); // Float32Array at 48 kHz, ready to play as it arrives
}
const url = URL.createObjectURL(wavBlob(chunks));
On iPhone and iPad use "wasm": Safari kills the tab with the WebGPU build of onnxruntime-web.
Accuracy
export/verify.py runs the ONNX graphs next to the PyTorch package on the same text and the same noise
(ONNX Runtime 1.30 on CPU, ema-lightning 1.0.1):
| max difference | |
|---|---|
| durations | 7.2e-7 |
| latents | 5.1e-4 |
| audio | 4.9e-5 |
The JavaScript engine draws noise from its own PRNG instead of torch.randn, so the same seed does not give the
same audio as the Python package.
fp16
Only the sound graph has an fp16 version. The angles that feed RoPE and the timestep embedding stay fp32 inside it. Measured on WebGPU (onnxruntime-web 1.30, Apple GPU) as log-spectral distance to fp32 with the same noise, where a different seed scores about 0.77:
| distance | |
|---|---|
| sound fp16 | 0.12 |
| decoder fp16 | 0.81: its convolutions go wrong on WebGPU (fine on CPU) |
| text fp16 | durations go wrong: WebGPU sums the masked GroupNorm in fp16 |
Reproduce
uv venv -p 3.11 .venv
uv pip install -p .venv ema-lightning==1.0.1 onnx onnxruntime onnxscript onnxconverter-common soundfile
.venv/bin/python export/export_onnx.py # text, sound, decoder .onnx and meta.json, here
.venv/bin/python export/verify.py # ONNX vs PyTorch, same noise
.venv/bin/python export/to_fp16.py # fp16/sound.onnx
.venv/bin/python export/compare_fp16.py # fp16 vs fp32 on the same noise, writes two WAVs
The graphs were exported with opset 17 and the TorchScript exporter (dynamo=False).
The normalizer needs rustup with the wasm32-unknown-unknown target:
cd normalizer-wasm
cargo build --release --target wasm32-unknown-unknown
cp target/wasm32-unknown-unknown/release/normalizer_tr_wasm.wasm ../normalizer.wasm
Credits and license
Apache-2.0, see LICENSE and NOTICE.
- Model: EMA Lightning by canberkkkkkk, Apache-2.0 (code). These files are conversions of its weights. For what the model can do, how it was measured and its limits, see the original model card.
- Text normalization: normalizer-tr by Erdem Tuna, Apache-2.0.
Notices for the compiled WebAssembly are in
licenses/.
Model tree for ozcancelik/ema-lightning-onnx
Base model
canberkkkkkk/ema-lightning