mlx-community/clef-flash-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune and serve models on Apple Silicon. Decision models guide · All OptiQ quants

A 4-bit mixed-precision MLX quant of Cloudflare/clef-flash, Cloudflare's 9B decision model. A decision model reads a piece of state and a set of typed questions (yes/no, choice, score) and returns a probability for every option of every question in one forward pass. There is no generated text and nothing to parse. The request and response match the Jev / SystemOne API.

The backbone is quantized with per-layer bit-widths: sensitive layers at 8-bit, the rest at 4-bit. The joint schema head and the vision tower are kept at full precision. Size on disk is about 8.3 GB.

Typed Decisions result

Scored on the full 400-case Typed Decisions test split (2,000 decisions) with the benchmark's own metrics, text only, MLX on an Apple M3 Max. The comparison is the stock uniform 4-bit quant of the same model, mlx-community/clef-flash-4bit, scored the same way.

Accuracy KL from gold (lower is better) Brier (lower is better) ECE (lower is better)
This quant 0.7035 0.2009 0.1057 0.1178
Uniform 4-bit 0.701 0.2126 0.1104 0.1204

The probabilities sit a little closer to the reference answers (lower KL and Brier). Accuracy is the same within noise. The gain over uniform 4-bit is small; the point of this quant is that it loses nothing and runs through OptiQ's decision endpoint.

Use

Serve it with OptiQ and the decision endpoint turns on by itself:

pip install mlx-optiq
optiq serve --model mlx-community/clef-flash-OptiQ-4bit
curl http://127.0.0.1:8080/v1/systemone -d '{
  "model": "clef",
  "state": "Checkout has been failing for every customer for the last hour.",
  "questions": {
    "urgent": {"type": "noul", "instructions": "Is this urgent?"},
    "team": {"type": "choice", "criteria": {"billing": "Payments", "technical": "Outages and errors"}}
  }
}'

Or from Python:

from optiq.decision import DecisionModel

model = DecisionModel.load("mlx-community/clef-flash-OptiQ-4bit")
print(model.predict({
    "state": "Checkout has been failing for every customer for the last hour.",
    "questions": {"urgent": {"type": "noul", "instructions": "Is this urgent?"}},
}))

Text input only for now: a request that carries images or video is refused with a clear error rather than answered without them. The decision models guide covers the three question types, errors, context length and training your own with optiq lora train.

Quantization details

Property Value
Predominant precision 4-bit
Layers at 8-bit (sensitive) 134
Layers at 4-bit 116
Group size 64
Joint schema head full precision, unchanged from the base
Vision tower bf16, in optiq/, unchanged from the base

The bit allocation reuses the measured recipe of the Qwen3.5-9B OptiQ quant, the backbone this model was post-trained from.

See also clef-OptiQ-4bit (Clef, 27B).

License

Apache-2.0, as the base model. Clef is by Cloudflare; see the base model card for how it was trained and evaluated.

Downloads last month
30
Safetensors
Model size
9B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/clef-flash-OptiQ-4bit

Finetuned
Qwen/Qwen3.5-9B
Quantized
(48)
this model

Evaluation results

  • LocalLLaMA/typed-decisions leaderboard
  • Accuracy View evaluation results
    source
    General model (any question schema at request time), not fitted by us on this benchmark's train split. Scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max.
    0.7 *
  • Kl From Gold View evaluation results
    source
    General model (any question schema at request time), not fitted by us on this benchmark's train split. Scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max.
    0.2 *
  • Brier View evaluation results
    source
    General model (any question schema at request time), not fitted by us on this benchmark's train split. Scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max.
    0.11 *