Wald-4B GGUF (v1.2)

GGUF files of Wald-4B v1.2 (Wald-Q4B v1.2, checkpoint 02600-f19) for llama.cpp, Ollama, LM Studio and other GGUF runtimes. CPU, Apple Silicon and consumer GPUs are all fine.

Wald-4B is a 4B decision model. You give it a state and a set of options, and it returns a calibrated probability for every option. Use it to pick a tool, route a request, classify an input or decide whether to ask the user. It is not a chat model: the probabilities come from reading the model's next-token distribution over the option letters, which the bundled wald-serve does for you (see Calibrated probabilities).

The same GGUF files are also on the main branch of org2ai/Wald-4B, which holds v1.2 since 2026-10-01; ollama run hf.co/org2ai/Wald-4B:Q4_K_M works there too.

Source weights: org2ai/Wald-4B at tag v1.2 (commit db376d32, shards 0b15067fโ€ฆ / 7cff102bโ€ฆ). Use the main repository for vLLM, the full model card, evaluation details and provenance.

Files

JevBench public set, 231 questions, effort none, scored with JevBench's own harness (fstandhartinger/jevbench at 9ec6f15a). Every row ran behind the same wald-serve 0.1.1, with the same prompt format and temperature table. "Same option" and |ฮ”p| compare each file with the vLLM BF16 reference on the same 231 questions.

File Size JevBench public Same option as vLLM BF16 Mean / max |ฮ”p| ECE
reference: original safetensors on vLLM 0.30.0, BF16 8.4 GB 204 / 231 โ€” โ€” 0.054
Wald-4B-v1.2-Q8_0.gguf 4.5 GB 206 / 231 229 / 231 0.003 / 0.055 0.045
Wald-4B-v1.2-Q6_K.gguf 3.5 GB 205 / 231 227 / 231 0.005 / 0.055 0.045
Wald-4B-v1.2-Q5_K_M.gguf 3.1 GB 205 / 231 227 / 231 0.008 / 0.204 0.043
Wald-4B-v1.2-Q4_K_M.gguf 2.7 GB 202 / 231 223 / 231 0.015 / 0.285 0.034

A BF16 GGUF (8.4 GB) was also tested (204 / 231, same option on 229 / 231, mean / max |ฮ”p| 0.001 / 0.023) but is not published here; for BF16 use the original safetensors in org2ai/Wald-4B.

Which file? Q8_0 matches the original most closely at half the size. Q4_K_M is the smallest and still close; it changes the chosen option on 8 of 231 questions. The ยฑ2-question differences between files are within the noise of near-tied questions: Q8_0's 206 is not better than the reference.

Calibrated probabilities (recommended)

wald-serve turns llama.cpp into Wald-4B's POST /v1/systemone API: one prompt per question, the option letters read from the next-token logprobs, and the temperature table applied. It needs Python โ‰ฅ 3.10 and llama.cpp's llama-server on your PATH (install llama.cpp, e.g. brew install llama.cpp).

hf download org2ai/Wald-4B-GGUF --include "Wald-4B-v1.2-Q8_0.gguf" "serving.json" "temperature.json" "server/*" --local-dir ./wald-gguf
pip install ./wald-gguf/server
wald-serve --gguf ./wald-gguf/Wald-4B-v1.2-Q8_0.gguf --max-model-len 32768 --port 8000

serving.json and temperature.json are read from the GGUF's folder (effort none, prompt format repeat_state_plain). Then:

curl http://localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": "The customer wants to return a damaged kettle.",
  "questions": {
    "route": {
      "type": "choice",
      "instructions": "Choose the support queue.",
      "criteria": {"returns": "Returns and refunds", "delivery": "Delivery tracking", "other": "Other enquiries"}
    }
  }
}'
{"answers": {"route": {"type": "choice", "choice": "returns",
  "probabilities": {"returns": 0.977, "delivery": 0.004, "other": 0.019}, "confidence": 0.965, "mode": "A"}}}

(That answer came from the Q4_K_M file.) Question types are choice (named options), noul (yes/no) and score (ordered levels). The request format and every field are documented in the main repository's API reference.

Already running llama-server (for example llama-server -hf org2ai/Wald-4B-GGUF:Q8_0 -c 32768)? Attach to it:

wald-serve --llamacpp http://127.0.0.1:8080 --temperature ./wald-gguf/temperature.json \
  --prompt-format repeat_state_plain --effort none --port 8000

Ollama and LM Studio

ollama run hf.co/org2ai/Wald-4B-GGUF:Q4_K_M

In LM Studio, search for Wald-4B-GGUF. Neither app was tested with these files. Some Ollama versions refuse Qwen3.5-architecture GGUFs; if yours does, use llama.cpp. A chat session gives you the model's text reply, not the calibrated probabilities. To get those, read the next-token logprobs of the option letters after this prompt (it is what wald-serve sends), or point wald-serve at llama.cpp as above:

State:
{state}

Question: {instructions}
(A) {option 1}
(B) {option 2}
Answer: (

Wald-4B v1.2 is tuned for this one-pass read (effort none). Use the thinking efforts on v1.1 instead (org2ai/Wald-4B, tag v1.1).

How these files were made

  • llama.cpp b11312: convert_hf_to_gguf.py --no-mtp --outtype bf16 --model-name "Wald-4B v1.2", then llama-quantize (from that BF16 GGUF) to Q8_0 / Q6_K / Q5_K_M / Q4_K_M. No importance matrix.
  • --no-mtp: Qwen3.5's config declares a multi-token-prediction layer that this checkpoint does not contain. The answers do not use it.
  • The parity reads ran on one RTX 5090 with two servers on the card at a time. The โ‰ˆ 0.13 s median per question through llama.cpp (โ‰ˆ 0.05 s on vLLM) is from that setup, not a latency claim.
  • server/ is wald-serve 0.1.1. It is the server in org2ai/Wald-4B plus a llama.cpp backend (--gguf, --llamacpp). The vLLM path is unchanged.

Licence and provenance

Apache-2.0, like the source weights. Training data, contamination notes and limitations are in the main repository: PROVENANCE.md, CONTAMINATION.md.

Wald-4B is an independent, self-hosted alternative to TypeSafe's hosted Jev API. It is not Jev, contains no Jev weights, and is not affiliated with or endorsed by TypeSafe AI.

Downloads last month
11,123
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for org2ai/Wald-4B-GGUF

Finetuned
org2ai/Wald-4B
Quantized
(2)
this model

Space using org2ai/Wald-4B-GGUF 1