Sortformer-Diarization-4spk-ONNX
ONNX export of NVIDIA's streaming Sortformer 4-speaker diarization model, for hosts that drive ONNX Runtime. Apple platforms have a CoreML build; this is the same checkpoint for everything else.
Changes from the base model
Exported from nvidia/diar_streaming_sortformer_4spk-v2.1 with no retraining
and no change to weights, architecture or configuration. The export composes
the model's pre-encoder and head into one graph per streaming step and traces
it at the default variant's static shapes.
What one call does
in: chunk[1, 3048, 128] mel features, 128 bins
chunk_lengths[1]
spkcache[1, 188, 512] speaker cache
spkcache_lengths[1]
fifo[1, 40, 512]
fifo_lengths[1]
out: spkcache_fifo_chunk_preds[1, 609, 4] per-frame activity, 4 speakers
chunk_pre_encode_embs[1, 381, 512]
chunk_pre_encode_lengths[1]
One call consumes ~30 seconds of new audio and needs ~3.2 seconds of audio after the chunk it reports on.
The host owns the speaker cache
The graph takes spkcache and fifo as inputs and returns embeddings; it does
not update them. Deciding what enters the cache, what waits in the FIFO and
what is evicted is the caller's job, and that bookkeeping โ the Arrival-Order
Speaker Cache โ is what makes a speaker index mean the same person for a whole
recording. A host that skips it gets per-call speaker numbering and none of the
model's advantage over window-local segmenters.
Eviction is not FIFO. Frames are scored by how confidently one speaker and no
other is active, boosted twice so every speaker keeps a minimum share and none
dominates, then the highest-scoring 188 are kept in chronological order with
unused slots filled by a running mean-silence embedding. NeMo's
SortformerModules.streaming_update_async is the reference.
The cache geometry above belongs to this export. Pairing it with another variant's update period evicts the wrong frames while looking healthy.
Measured
Against NeMo's own forward_streaming on 132 seconds of two-speaker audio,
with the speaker cache overflowing:
| mean absolute error | 0.0149 |
| decision agreement | 0.9850 |
| throughput | 42.4x realtime, 5 calls, 615 ms each |
A bare ONNX Runtime session, without the reference model loaded beside it, is 505 MB resident and 1251 MB peak, and one call takes 467 ms to cover 30 s of new audio โ 64x realtime at 12 intra-op threads. The 42.4x above comes from the parity harness, which carries PyTorch and NeMo in the same process; use it to compare against the reference, not to size a deployment.
Measured on a 24-core x86 CPU through ONNX Runtime's CPU provider. The same loop driven by PyTorch instead of this graph gives identical numbers, so the export contributes no error of its own โ the difference is between NeMo's two internal streaming paths.
It tracks at most four speakers and degrades beyond that.
Files
sortformer-default.onnx and config.json, which carries the cache sizes and
chunk geometry a host needs to size its buffers.
- Downloads last month
- 14
Model tree for soniqo/Sortformer-Diarization-4spk-ONNX
Base model
nvidia/diar_streaming_sortformer_4spk-v2.1