Klang Pianissimo

Pianissimo is a fast, accurate speech recognition model for Swedish. Developed by Klang, it is a 600-million-parameter fine-tune of NVIDIA Parakeet v3, available under CC BY 4.0.

Pianissimo can transcribe one hour of audio in one second on an H100.

Release page with demo here.

It achieves 4.46% word error rate on Common Voice Swedish and 6.51% on Swedish FLEURS, reducing errors by 76% and 57% compared with Parakeet v3.

It supports punctuation, capitalization, and word-level timestamps, and can transcribe both short clips and long recordings.

Usage

Install PyTorch and nemo_toolkit[asr].

Load the model and transcribe a 16 kHz mono audio file:

import torch
from nemo.collections.asr.models import ASRModel

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = ASRModel.from_pretrained(
    model_name="KlangAI/pianissimo-sv",
    map_location=device,
).eval()

with torch.inference_mode():
    hypotheses = model.transcribe(
        audio=["audio.wav"],
        batch_size=1,
        return_hypotheses=True,
    )

print(hypotheses[0].text)

For multiple files, pass a list of paths and increase batch_size to suit your GPU. Use batch_size=1 for long recordings. An NVIDIA GPU is recommended; memory use grows with recording length and batch size.

Timestamps

Add timestamps=True to obtain word and segment timestamps:

with torch.inference_mode():
    hypotheses = model.transcribe(
        audio=["audio.wav"],
        batch_size=1,
        return_hypotheses=True,
        timestamps=True,
    )

for word in hypotheses[0].timestamp["word"]:
    print(word["start"], word["end"], word["word"])

Vocabulary

Supports custom vocabulary with NeMo's built-in phrase boosting: list names and terms to improve how they are spelled. The ONNX and MLX versions support it with our own implementation. The words and phrases are case sensitive.

import copy
from omegaconf import open_dict

decoding = copy.deepcopy(model.cfg.decoding)
with open_dict(decoding):
    decoding.strategy = "greedy_batch"
    decoding.greedy.boosting_tree = {
        "key_phrases_list": [
            "Klang AI", "Pianissimo", "RixVox", "Common Voice",
        ],
        "context_score": 1.0,
        "depth_scaling": 2.0,
        "use_triton": False,
    }
    decoding.greedy.boosting_tree_alpha = 1.0
model.change_decoding_strategy(decoding)

We tested it on 135 FLEURS clips where the model misspelled a name, plus 100 clips without names. The list held the 136 names from those clips and 200 decoy names that are not in the audio.

boosting_tree_alpha Names spelled right WER, clips with names WER, other clips Listed names inserted wrongly
0 (off) 39% 13.17 4.27 0
0.25 48% 11.87 4.32 0
0.5 53% 11.26 4.32 0
1.0 64% 10.10 4.37 0
2.0 71% 10.54 5.04 4

Quantized versions

Quantized versions for smaller downloads or running faster locally:

Version Runs on Download CV test FLEURS test Klang Dialects Speed
Original fp32 (this model) PyTorch and NeMo 2.51 GB 4.46 6.51 4.85
ONNX int8 Any CPU 660 MB 4.65 6.62 4.93 35ร— (Intel), 74ร— (Mac)
ONNX int4 Any CPU 515 MB 4.63 6.58 5.06 12ร— (Intel), 37ร— (Mac)
ONNX fp16 Any CPU 1.28 GB 4.47 6.48 4.89 18ร— (Intel), 55ร— (Mac)
ONNX fp32 Any CPU 2.56 GB 4.46 6.48 4.89 18ร— (Intel), 55ร— (Mac)
MLX 8-bit Mac with Apple silicon 759 MB 4.51 6.45 4.83 151ร—
MLX 4-bit Mac with Apple silicon 528 MB 4.66 6.53 5.20 126ร—
MLX bf16 Mac with Apple silicon 1.25 GB 4.45 6.43 4.85 140ร—
MLX fp32 Mac with Apple silicon 2.51 GB 4.46 6.49 4.90 173ร—

Speed is Real Time Factor (RTFx) measured on a 30-second clip. ONNX benchmarked with 4 threads on an Intel Xeon Platinum 8481C and 6 threads on an Apple M5 Pro. MLX benchmarked on Apple M5 Pro GPU.

The ONNX versions run on any computer (Windows, Linux or Mac, with an Intel, AMD or ARM processor):

import onnx_asr  # pip install onnx-asr[cpu,hub]

model = onnx_asr.load_model("KlangAI/pianissimo-sv-onnx", quantization="int8")
print(model.recognize("audio.wav"))

The MLX versions run on the GPU of Macs with Apple silicon. They use a small loader file included in each MLX repository; see the MLX 8-bit model card for how to use it.

Quality and speed

Word error rate (WER, %) on Swedish speech; lower is better. All models were evaluated by Klang using the same scoring procedure. Bulk throughput is expressed as multiples of realtime.

Model CV test FLEURS test Klang Dialects Bulk throughput
Pianissimo 4.46 6.51 4.85 2,500ร—
Parakeet TDT 0.6B v3 18.54 15.18 25.82 2,500ร—
KB-Whisper medium 5.40 6.58 3.58 66ร—
KB-Whisper large 3.91 5.08 2.30 39ร—
Whisper large-v3 8.07 7.24 8.16 39ร—

The evaluations cover 5,516 Common Voice v26 test clips, 758 FLEURS test clips, and the 1,804-recording clean set of Klang Dialects. We calculate corpus-level WER after lowercasing, replacing punctuation with spaces, and collapsing whitespace.

Throughput was measured on one NVIDIA A100 80 GB using 30-second audio clips, with 16bit precision and optimal batch sizes.

Architecture

Pianissimo uses a FastConformer encoder and Token-and-Duration Transducer (TDT) decoder, retaining Parakeet v3's architecture and tokenizer.

Property Value
Parameters Approximately 600 million
Audio input 16 kHz, mono
Features 128-band log-mel spectrogram
Encoder 24 layers, 8ร— subsampling
Default attention Local, 256 encoder frames on each side per layer
Checkpoint size 2.51 GB

Unlike the original Parakeet checkpoint's full attention, Pianissimo defaults to local attention with approximately 20.5 seconds of context in each direction per layer. This keeps attention computation linear in recording length and makes long recordings practical.

Training

We fine-tuned Pianissimo on approximately 50,000 hours of Swedish speech, including public datasets such as RixVox, Common Voice and FLEURS, as well as our internal dataset of publicly available data. Augmentation included reverberation, noise, compression, reduced bandwidth, and gain changes.

Limitations

Overlapping speech, strong background noise, uncommon names, and specialized vocabulary can reduce recognition quality. Number formatting and punctuation may need editing for the intended application. Evaluation has focused on Swedish; performance on other languages and code-switching has not been established.

License and citation

Pianissimo is released under CC BY 4.0.

@misc{klang2026pianissimo,
  title = {Klang Pianissimo},
  author = {{Klang}},
  year = {2026},
  howpublished = {Hugging Face model repository},
  url = {https://huggingface.co/KlangAI/pianissimo-sv}
}
Downloads last month
2,809
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for KlangAI/pianissimo-sv

Finetuned
(106)
this model
Finetunes
6 models
Quantizations
9 models

Datasets used to train KlangAI/pianissimo-sv

Spaces using KlangAI/pianissimo-sv 3