Audio-align / PreAlignSLM releases

Paper-claim checkpoint timeline

The auditable paper sequence is pinned at revision a9324d03830d58fca87eac010162982efa106b22. It contains four fp16 active-interface checkpoints:

Directory Role
aligned-stage3-step45000 Aligned real-audio base
pre-zero-audio-phase2-step260000 Direct before-checkpoint
zero-audio-phase2.2-step24000 About 520K pure-text QA; no waveforms
zero-audio-phase2.3-step5000 About 120K pure-text chat; no waveforms

Each directory contains TASU, the acoustic projector, and the instruction router: 84,655,552 parameters, or 5.4839% of the main LLM's 1,543,714,304 independent parameters. The tied token embedding and language-model head are counted once rather than as separate state-dict aliases. The accompanying verification records exact tensor equality for Qwen2.5-1.5B-Instruct, SenseVoice, and ArCap across the sequence.

The claim is deliberately narrower than “training without audio”: real audio is used for grounding before the two zero-audio post-alignment stages. Current evidence is strongest for semantic and speech-paralinguistic transfer; general non-speech signal/event transfer still needs a matched causal experiment.

AIR-Bench metrics in the manifest distinguish historical valid-only accuracy from coverage-aware end-to-end accuracy; unparsable responses are incorrect in the latter.

See the release manifest, frozen-state verification, and GitHub documentation.

Historical full-branch checkpoint

The repository root also stores the original full non-main-LLM export for phase2_3_chat/model-5000.pt. It includes the frozen SenseVoice and ArCap experts as well as the Phase 2.3 interfaces:

File Contents
semantic_branch.safetensors SenseVoice, legacy semantic layers, TASU
acoustic_branch.safetensors ArCap and acoustic projector
router.safetensors BERT-small instruction router
manifest.json Tensor counts, provenance, and SHA256 checksums

These frozen expert tensors can reconstruct every checkpoint in the paper timeline because the exact comparison proves that they never change. Load the full-branch state first, then overwrite the three active interfaces with the selected timeline directory. Qwen2.5-1.5B-Instruct is loaded separately from its upstream release.

For an unwrapped PyTorch model, strip the training-time module. prefix:

from safetensors.torch import load_file

state = {}
for filename in [
    "semantic_branch.safetensors",
    "acoustic_branch.safetensors",
    "router.safetensors",
]:
    state.update({
        key.removeprefix("module."): value
        for key, value in load_file(filename, device="cpu").items()
    })

missing, unexpected = model.load_state_dict(state, strict=False)
assert not unexpected

The historical Phase 2.3 checkpoint reports ASR test-clean/test-other WER 3.50%/7.22%, AIR-Bench Chat 2.365/10 with a Qwen3.5-27B judge, and VoiceBench sd-qa 21.8%. Metrics, failure cases, and claim boundaries are documented in the GitHub repository rather than inferred from a single aggregate score.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support