EMA Lightning ONNX

ONNX conversion of EMA Lightning, the 8.6M-parameter Turkish text-to-speech model, for running without PyTorch: in the browser with onnxruntime-web (WebGPU or WASM), or anywhere else ONNX Runtime runs. Output is 48 kHz mono.

Live demo: https://ema-lightning-web.vercel.app/ (source: https://github.com/ozcancelik/ema-lightning-web)

In the browser on an Apple laptop GPU, first audio arrives in about 140 ms and speech is made 20–30× faster than real time.

EMA Lightning'in ONNX'e dönüştürülmüş hâli: PyTorch olmadan, tarayıcıda (WebGPU / WASM) ya da ONNX Runtime'ın çalıştığı her yerde Türkçe konuşma üretir.

Files

File Size What it is
text.onnx 4.8 MB letters → hidden states and per-letter durations
sound.onnx 18.1 MB aligner and the four distilled flow steps → latent frames
fp16/sound.onnx 9.6 MB the same graph in fp16, fp32 inputs and outputs
decoder.onnx 12.0 MB latent frames → waveform
meta.json vocabulary, flow times, latent size, hop, sample rate
normalizer.wasm 1.2 MB normalizer-tr as WebAssembly: "14:30" → "on dört otuz"
web/tts.js, web/normalizer.js the reference JavaScript engine used by the demo
export/ export, fp16 conversion and verification scripts
normalizer-wasm/ Rust source of the WebAssembly wrapper

Graphs

All graphs have batch size 1. L is the number of letters in a piece of text, T the number of 25 Hz frames.

Graph Inputs Outputs
text ids int64 [1, L] h float32 [1, L, 224], dur float32 [1, L] (frames per letter)
sound h [1, L, 224], cg float32 [1, L], fg float32 [1, T], cw int64 [1, L], fw int64 [1, T], noise float32 [1, 4, T, 64] latents float32 [1, T, 64]
decoder z float32 [1, 64, T] audio float32 [1, T × 1920]

meta.json:

{"vocab": ["<pad>", "<unk>", " ", "!", ...], "times": [0.0, 0.25, 0.5, 0.75],
 "latent_dim": 64, "hop": 1920, "sample_rate": 48000}

Pipeline

The graphs hold only the network: convolutions, matmuls and elementwise ops. The bookkeeping between them, which needs gathers and scatters, is done by the caller. web/tts.js does all of it in about 300 lines, and export/verify.py does the same in NumPy.

  1. Text. Normalize with normalizer.wasm (the "fallback" ambiguity policy), lowercase with Turkish rules (İ → i, I → ı), strip accents except from Turkish letters, and replace anything outside vocab with a space.
  2. Chunk into pieces of at most ~10 seconds, cut at sentence, then clause, then word boundaries. Each piece ends with punctuation.
  3. text: ids are indices into vocab.
  4. Plan the frames. Divide dur by the speed, sum it per word, round to whole frames (1 to 250 per word, 3000 per piece). cw is each letter's word index and fw each frame's word index. cg and fg are the letter and frame positions measured in letters (see plan() in web/tts.js).
  5. sound with Gaussian noise of shape [1, 4, T, 64]. The seed sets the voice's variation.
  6. decoder in windows (one second first, then four seconds), with 8 frames of context on each side. Each window's audio is cut out of the middle of its output.

Use in the browser

Copy web/tts.js and web/normalizer.js into your project and point EMA.load at this repository. Hugging Face serves the files with CORS, and tts.js keeps them in Cache Storage after the first download.

import { EMA, wavBlob } from "./tts.js";

const base = "https://huggingface.co/ozcancelik/ema-lightning-onnx/resolve/main/";
const tts = await EMA.load("webgpu", base);  // or "wasm"; add precision "fp16" for the smaller sound graph

const chunks = [];
for await (const audio of tts.stream("Merhaba, saat 14:30'da görüşelim.", { speed: 1, seed: 0 })) {
  chunks.push(audio); // Float32Array at 48 kHz, ready to play as it arrives
}
const url = URL.createObjectURL(wavBlob(chunks));

On iPhone and iPad use "wasm": Safari kills the tab with the WebGPU build of onnxruntime-web.

Accuracy

export/verify.py runs the ONNX graphs next to the PyTorch package on the same text and the same noise (ONNX Runtime 1.30 on CPU, ema-lightning 1.0.1):

max difference
durations 7.2e-7
latents 5.1e-4
audio 4.9e-5

The JavaScript engine draws noise from its own PRNG instead of torch.randn, so the same seed does not give the same audio as the Python package.

fp16

Only the sound graph has an fp16 version. The angles that feed RoPE and the timestep embedding stay fp32 inside it. Measured on WebGPU (onnxruntime-web 1.30, Apple GPU) as log-spectral distance to fp32 with the same noise, where a different seed scores about 0.77:

distance
sound fp16 0.12
decoder fp16 0.81: its convolutions go wrong on WebGPU (fine on CPU)
text fp16 durations go wrong: WebGPU sums the masked GroupNorm in fp16

Reproduce

uv venv -p 3.11 .venv
uv pip install -p .venv ema-lightning==1.0.1 onnx onnxruntime onnxscript onnxconverter-common soundfile
.venv/bin/python export/export_onnx.py      # text, sound, decoder .onnx and meta.json, here
.venv/bin/python export/verify.py           # ONNX vs PyTorch, same noise
.venv/bin/python export/to_fp16.py          # fp16/sound.onnx
.venv/bin/python export/compare_fp16.py     # fp16 vs fp32 on the same noise, writes two WAVs

The graphs were exported with opset 17 and the TorchScript exporter (dynamo=False).

The normalizer needs rustup with the wasm32-unknown-unknown target:

cd normalizer-wasm
cargo build --release --target wasm32-unknown-unknown
cp target/wasm32-unknown-unknown/release/normalizer_tr_wasm.wasm ../normalizer.wasm

Credits and license

Apache-2.0, see LICENSE and NOTICE.

  • Model: EMA Lightning by canberkkkkkk, Apache-2.0 (code). These files are conversions of its weights. For what the model can do, how it was measured and its limits, see the original model card.
  • Text normalization: normalizer-tr by Erdem Tuna, Apache-2.0. Notices for the compiled WebAssembly are in licenses/.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ozcancelik/ema-lightning-onnx

Quantized
(3)
this model