Instructions to use org2ai/Wald-4B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use org2ai/Wald-4B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf org2ai/Wald-4B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf org2ai/Wald-4B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf org2ai/Wald-4B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf org2ai/Wald-4B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf org2ai/Wald-4B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf org2ai/Wald-4B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf org2ai/Wald-4B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf org2ai/Wald-4B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/org2ai/Wald-4B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use org2ai/Wald-4B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "org2ai/Wald-4B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/org2ai/Wald-4B-GGUF:Q4_K_M
- Ollama
How to use org2ai/Wald-4B-GGUF with Ollama:
ollama run hf.co/org2ai/Wald-4B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use org2ai/Wald-4B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf org2ai/Wald-4B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "org2ai/Wald-4B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use org2ai/Wald-4B-GGUF with Docker Model Runner:
docker model run hf.co/org2ai/Wald-4B-GGUF:Q4_K_M
- Lemonade
How to use org2ai/Wald-4B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull org2ai/Wald-4B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Wald-4B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use org2ai/Wald-4B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf org2ai/Wald-4B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default org2ai/Wald-4B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use org2ai/Wald-4B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf org2ai/Wald-4B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "org2ai/Wald-4B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Wald-4B GGUF (v1.2)
GGUF files of Wald-4B v1.2 (Wald-Q4B v1.2, checkpoint 02600-f19) for
llama.cpp, Ollama, LM Studio and other GGUF runtimes. CPU, Apple Silicon and consumer GPUs are all fine.
Wald-4B is a 4B decision model. You give it a state and a set of options, and it returns a calibrated probability for
every option. Use it to pick a tool, route a request, classify an input or decide whether to ask the user. It is not a
chat model: the probabilities come from reading the model's next-token distribution over the option letters, which the
bundled wald-serve does for you (see Calibrated probabilities).
The same GGUF files are also on the main branch of org2ai/Wald-4B, which holds v1.2 since
2026-10-01; ollama run hf.co/org2ai/Wald-4B:Q4_K_M works there too.
Source weights: org2ai/Wald-4B at tag v1.2 (commit
db376d32, shards 0b15067fโฆ / 7cff102bโฆ). Use the main repository for vLLM, the full model card, evaluation
details and provenance.
Files
JevBench public set, 231 questions, effort none, scored with JevBench's own harness (fstandhartinger/jevbench at
9ec6f15a). Every row ran behind the same wald-serve 0.1.1, with the same prompt format and temperature table. "Same
option" and |ฮp| compare each file with the vLLM BF16 reference on the same 231 questions.
| File | Size | JevBench public | Same option as vLLM BF16 | Mean / max |ฮp| | ECE |
|---|---|---|---|---|---|
| reference: original safetensors on vLLM 0.30.0, BF16 | 8.4 GB | 204 / 231 | โ | โ | 0.054 |
Wald-4B-v1.2-Q8_0.gguf |
4.5 GB | 206 / 231 | 229 / 231 | 0.003 / 0.055 | 0.045 |
Wald-4B-v1.2-Q6_K.gguf |
3.5 GB | 205 / 231 | 227 / 231 | 0.005 / 0.055 | 0.045 |
Wald-4B-v1.2-Q5_K_M.gguf |
3.1 GB | 205 / 231 | 227 / 231 | 0.008 / 0.204 | 0.043 |
Wald-4B-v1.2-Q4_K_M.gguf |
2.7 GB | 202 / 231 | 223 / 231 | 0.015 / 0.285 | 0.034 |
A BF16 GGUF (8.4 GB) was also tested (204 / 231, same option on 229 / 231, mean / max |ฮp| 0.001 / 0.023) but is not published here; for BF16 use the original safetensors in org2ai/Wald-4B.
Which file? Q8_0 matches the original most closely at half the size. Q4_K_M is the smallest and still close; it changes the chosen option on 8 of 231 questions. The ยฑ2-question differences between files are within the noise of near-tied questions: Q8_0's 206 is not better than the reference.
Calibrated probabilities (recommended)
wald-serve turns llama.cpp into Wald-4B's POST /v1/systemone API: one prompt per question, the option letters read
from the next-token logprobs, and the temperature table applied. It needs Python โฅ 3.10 and llama.cpp's llama-server
on your PATH (install llama.cpp, e.g. brew install llama.cpp).
hf download org2ai/Wald-4B-GGUF --include "Wald-4B-v1.2-Q8_0.gguf" "serving.json" "temperature.json" "server/*" --local-dir ./wald-gguf
pip install ./wald-gguf/server
wald-serve --gguf ./wald-gguf/Wald-4B-v1.2-Q8_0.gguf --max-model-len 32768 --port 8000
serving.json and temperature.json are read from the GGUF's folder (effort none, prompt format
repeat_state_plain). Then:
curl http://localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
"state": "The customer wants to return a damaged kettle.",
"questions": {
"route": {
"type": "choice",
"instructions": "Choose the support queue.",
"criteria": {"returns": "Returns and refunds", "delivery": "Delivery tracking", "other": "Other enquiries"}
}
}
}'
{"answers": {"route": {"type": "choice", "choice": "returns",
"probabilities": {"returns": 0.977, "delivery": 0.004, "other": 0.019}, "confidence": 0.965, "mode": "A"}}}
(That answer came from the Q4_K_M file.) Question types are choice (named options), noul (yes/no) and score
(ordered levels). The request format and every field are documented in the main repository's
API reference.
Already running llama-server (for example llama-server -hf org2ai/Wald-4B-GGUF:Q8_0 -c 32768)? Attach to it:
wald-serve --llamacpp http://127.0.0.1:8080 --temperature ./wald-gguf/temperature.json \
--prompt-format repeat_state_plain --effort none --port 8000
Ollama and LM Studio
ollama run hf.co/org2ai/Wald-4B-GGUF:Q4_K_M
In LM Studio, search for Wald-4B-GGUF. Neither app was tested with these files. Some Ollama versions refuse
Qwen3.5-architecture GGUFs; if yours does, use llama.cpp. A chat session gives you the model's text reply, not the
calibrated probabilities. To get those, read the next-token logprobs of the option letters after this prompt
(it is what wald-serve sends), or point wald-serve at llama.cpp as above:
State:
{state}
Question: {instructions}
(A) {option 1}
(B) {option 2}
Answer: (
Wald-4B v1.2 is tuned for this one-pass read (effort none). Use the thinking efforts on v1.1 instead
(org2ai/Wald-4B, tag v1.1).
How these files were made
- llama.cpp
b11312:convert_hf_to_gguf.py --no-mtp --outtype bf16 --model-name "Wald-4B v1.2", thenllama-quantize(from that BF16 GGUF) to Q8_0 / Q6_K / Q5_K_M / Q4_K_M. No importance matrix. --no-mtp: Qwen3.5's config declares a multi-token-prediction layer that this checkpoint does not contain. The answers do not use it.- The parity reads ran on one RTX 5090 with two servers on the card at a time. The โ 0.13 s median per question through llama.cpp (โ 0.05 s on vLLM) is from that setup, not a latency claim.
server/iswald-serve0.1.1. It is the server inorg2ai/Wald-4Bplus a llama.cpp backend (--gguf,--llamacpp). The vLLM path is unchanged.
Licence and provenance
Apache-2.0, like the source weights. Training data, contamination notes and limitations are in the main repository: PROVENANCE.md, CONTAMINATION.md.
Wald-4B is an independent, self-hosted alternative to TypeSafe's hosted Jev API. It is not Jev, contains no Jev weights, and is not affiliated with or endorsed by TypeSafe AI.
- Downloads last month
- 11,123
4-bit
5-bit
6-bit
8-bit