Instructions to use czl/CLM-v0.1-8B-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use czl/CLM-v0.1-8B-MLX with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir CLM-v0.1-8B-MLX czl/CLM-v0.1-8B-MLX
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
CLM-v0.1-8B — MLX encoder
The encoder half of Contrastive-LM/CLM-v0.1-8B,
unquantised bf16 for MLX on Apple Silicon. This repo is the
reference every width below is measured against, not a quantisation of it.
⚠️ Not recommended for ranking:
4bit. agreement on decisive (>1 nat) decisions 0.889651 < 0.995; planner accuracy falls 31.95 points, beyond the 3.0-point budget. Published with the measured numbers so the result is reproducible, not because it is a drop-in replacement for the bf16 encoder.
Repositories
One repository per bit width, matching mlx-community. Each has the weights at the repo root,
so the Hub file browser lists every file with its size and mlx_lm.load("<repo>") works.
| repo | width | verdict |
|---|---|---|
czl/CLM-v0.1-8B-MLX-8bit |
8-bit | recommended |
czl/CLM-v0.1-8B-MLX-6bit |
6-bit | recommended |
czl/CLM-v0.1-8B-MLX-4bit |
4-bit | not recommended |
czl/CLM-v0.1-8B-MLX (this repo) |
bf16 | the reference |
A parallel outq2 line ships each width again with lm_head at 2-bit instead of the bulk
width. Pooling never evaluates the output projection, so those are indistinguishable from
re-running the repos above — measured over the full corpus against a matched null, since
MLX is not run-to-run deterministic on the GPU — and smaller for free:
| repo | vs its base | smaller by |
|---|---|---|
czl/CLM-v0.1-8B-MLX-8bit-outq2 |
8bit |
389 MB |
czl/CLM-v0.1-8B-MLX-6bit-outq2 |
6bit |
233 MB |
czl/CLM-v0.1-8B-MLX-4bit-outq2 |
4bit |
78 MB |
Pooling only — a 2-bit lm_head is a wrecked output projection and is not for generation.
Evaluation
Agreement against a bf16 MLX reference of the same encoder in the same runtime, over
23,926 scored System One questions, through the real head stack
(argmax(scale · cos), scale = 100.0 — cosine error is amplified 100×).
| variant | cos min | cos mean | top-1 | top-1 (decisive) | planner acc. Δ | verdict |
|---|---|---|---|---|---|---|
bf16 |
– | – | 1.0000 (ref) | 1.0000 (ref) | – | reference |
8bit |
0.98294 | 0.99985 | 0.9888 | 1.0000 | -0.46 pts | yes |
6bit |
0.89642 | 0.99918 | 0.9828 | 1.0000 | +0.70 pts | yes |
4bit |
0.87825 | 0.99191 | 0.6186 | 0.8897 | -31.95 pts | no |
Pass line. top-1 >= 1.0000 — the measured bf16-vs-bf16 noise floor of this corpus in this runtime — and top-1 (decisive) >= 0.995, where decisive means the reference's own top-1 led by more than 1 nat. A third condition applies: planner accuracy Δ, the change in accuracy against the T-Rex physics planner's label. A variant that holds every decisive decision but still loses measurable accuracy is marked usable rather than recommended.
Against the numbers published in the parent model card
| case | model card (vLLM bf16) | this runtime (bf16 reference) |
|---|---|---|
anchor-invoice/department |
billing |
billing — argmax matches, billing=0.98818, technical=0.01182 |
anchor-tides/rank |
0 |
0 — argmax matches, 0=0.99386, 1=0.00003, 2=0.00611 |
How to use
Each per-width repo is a standard mlx-community layout. Load one directly:
mlx_lm.load("czl/CLM-v0.1-8B-MLX-8bit")
Or serve it for CLM:
python -m clm_mlx.server --model <dir> --port 8092
clm-serve --emb-url http://127.0.0.1:8092/v1/embeddings
Quantization details
Tool: mlx-lm 0.31.3, one mlx_lm.convert pass per width:
mlx_lm.convert --hf-path Qwen/Qwen3-8B --mlx-path <dir> -q --q-bits N --q-group-size 32
- affine mode, per-group bf16 scale and bias.
group_size = 32— measured best at every width. mlx-lm's default (64) and Qwen'smlx-communityrepos (128) are both worse here. See the per-width cards for the full group-size ladder.
Limitations
- The projection heads are encoder-locked to Qwen3-8B last-token-pooled 4096-d embeddings.
- This is a ranker, not a generator.
lm_headis never evaluated under pooling. - The head stack amplifies cosine error 100×. Read the confident-decision column.
- Probabilities are relative to the candidate set.
License
Apache-2.0, inherited from Qwen/Qwen3-8B and Contrastive-LM/CLM-v0.1-8B. Not gated.
Citation
@misc{kwok2026contrastivelanguagemodels,
title={Contrastive Language Models: A System One Model for Fast and Generalizable Decision-Making},
author={Jacky Kwok and Hangoo Kang and Tarun Suresh and Jon Saad-Falcon and Marco Pavone and Christopher Ré and Azalia Mirhoseini},
year={2026},
note={Notion Blog},
url={https://contrastive-lm.notion.site}
}
- Downloads last month
- 125
Quantized