Text Generation
Transformers
Safetensors
GGUF
English
qwen3_5
image-text-to-text
decision-model
typed-decisions
calibration
calibrated-probabilities
classification
tool-selection
agent-routing
decision-index
jevbench
jev-compatible
systemone
wald
wald-q4b
qwen3.5
4b
vllm
reasoning
llama.cpp
conversational
Eval Results (legacy)
Instructions to use org2ai/Wald-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use org2ai/Wald-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="org2ai/Wald-4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("org2ai/Wald-4B") model = AutoModelForMultimodalLM.from_pretrained("org2ai/Wald-4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use org2ai/Wald-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "org2ai/Wald-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/org2ai/Wald-4B
- SGLang
How to use org2ai/Wald-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "org2ai/Wald-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "org2ai/Wald-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use org2ai/Wald-4B with Docker Model Runner:
docker model run hf.co/org2ai/Wald-4B
org2ai
Wald-Q4B v2 (04701-c22 + Qwen3.5-4B vision tower): main = v2; v1.x at tags v1.2 / v1.2-main / v1.1 / v1.0
ef7481d |
Download CONTAMINATION.md from org2ai/Wald-4B: direct link, hf CLI and curl.
- Browser
- Download file 2.54 kB
-
https://huggingface.co/org2ai/Wald-4B/resolve/main/CONTAMINATION.md
- Command line
-
hf download hf://org2ai/Wald-4B/CONTAMINATION.md
-
curl -L -o CONTAMINATION.md https://huggingface.co/org2ai/Wald-4B/resolve/main/CONTAMINATION.md
2.54 kB
Evaluation notes for Wald-Q4B v2 (04701-c22)
- Decision Index. Every training file of all three stages was scanned against the complete Decision Index 0.3 public suite with our strict matcher: 0 strict hits (contamination/trained-on-di-ids.json: ids, rule, file hashes, what is not covered). Each stage's data had also been gated before training (stage 1 against the 0.2 suite, stage 3 against the 0.3 suite with additional exact full-shingle and character n-gram checks; every hit row was dropped). The training data contains train splits of some benchmarks whose test items the Decision Index also uses (PROVENANCE.md), and the model was developed with the index's task formats in view, so its Decision Index results are not held-out results in the usual sense. The organizer's private tests are unknown to us.
- JevBench public set (231). A development scoreboard: never training data, but read repeatedly during development. Not held out.
- JevBench-XL (internal). Stage 2's and stage 3's new items were exact-state deduplicated against all XL partitions, JevBench, JevAdvBench and the Decision Index sample; the XL TEST partitions were not used for any selection. The XL scores are internal and not independently verifiable.
- multistep_decisions (100) and JevAdvBench are evaluation-only.
- Checkpoint and policy selection. Auto 0.7 was fixed as the default before the Decision Index 0.3 runs. The released checkpoint (step 300 of 638) was chosen over the final step and over the earlier C16B / C16C checkpoints using paired screens on a fixed subset of the Decision Index 0.3 public suite and then one full public run; the public index therefore informed the choice of checkpoint. No successful request was repeated.
- Calibration. One frozen temperature table (
temperature.json, sha256a0f72cd2d0a653e81051e5a0c77fc1a69131552a8102b580a93a6dbe7908b2da), fitted on 275 development rows before these evaluations; no benchmark labels were used, and the Decision Index 0.3 run did not refit it. - Vision. The vision tower is Qwen3.5-4B's own, unchanged; no image data was used in training. Image results are zero-shot and depend on the base model's pretraining, which may include the public image benchmarks.
- Not ruled out: semantic overlap, and contamination in the base model's pretraining data.
- No request payloads, gold labels or generated reasoning text are included in this repository.
The notes for v1.x are at their tags.