Instructions to use peasantsmith/GLM-5.3-Flash-Maya-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use peasantsmith/GLM-5.3-Flash-Maya-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S # Run inference directly in the terminal: llama cli -hf peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S # Run inference directly in the terminal: llama cli -hf peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S # Run inference directly in the terminal: ./llama-cli -hf peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S
Use Docker
docker model run hf.co/peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S
- LM Studio
- Jan
- vLLM
How to use peasantsmith/GLM-5.3-Flash-Maya-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "peasantsmith/GLM-5.3-Flash-Maya-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "peasantsmith/GLM-5.3-Flash-Maya-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S
- Ollama
How to use peasantsmith/GLM-5.3-Flash-Maya-GGUF with Ollama:
ollama run hf.co/peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S
- Unsloth Desktop
- Pi
How to use peasantsmith/GLM-5.3-Flash-Maya-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use peasantsmith/GLM-5.3-Flash-Maya-GGUF with Docker Model Runner:
docker model run hf.co/peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S
- Lemonade
How to use peasantsmith/GLM-5.3-Flash-Maya-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S
Run and chat with the model
lemonade run user.GLM-5.3-Flash-Maya-GGUF-IQ3_S
List all available models
lemonade list
- Hermes Agent
How to use peasantsmith/GLM-5.3-Flash-Maya-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use peasantsmith/GLM-5.3-Flash-Maya-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "peasantsmith/GLM-5.3-Flash-Maya-GGUF:IQ3_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
GLM-5.3-Flash ยท Maya-S, Maya-S24, Maya-M and Maya-L
Maya-L: 99.2% of the full FP8 model's accuracy on zero-shot tasks (ARC, HellaSwag, WinoGrande, PIQA), in 156 GB. Maya-S: 97.9% in a 96 GB file. Maya-S24: 97.7%, and up to 14% faster decode on 24 GB cards, in 94.7 GB. Maya-M: 97.9%, and closer to the full model token by token, in 116 GB.
GGUF quantizations of zai-org/GLM-5.3-Flash (321 B parameters, a mixture of experts with about 18 B active per token), made for Project Maya. Project Maya keeps the most-used experts on the GPU, the next in RAM and the rest on the SSD, so all four run well below the model's size (two GPUs with 30 GB of RAM in the measurements below). More memory is faster, and every machine is different: try it on yours. All four are made from Z.ai's own FP8 release - the precision the model is served at - not from a re-quantized file.
- Maya-S (96.5 GB) is the compact one, made for PCs with a smaller memory pool across RAM and VRAM: its routed experts - most of the model - in about 2 bits (IQ2_XXS / IQ2_S), its attention and shared experts in 6-bit (Q6_K).
- Maya-S24 (94.7 GB) is Maya-S for 24 GB cards (RTX 3090 / 4090): the same 2-bit routed experts, with only the small part every token runs through - the attention and shared experts - in 4-bit (Q4_K) instead of Maya-S's 6-bit. That leaves about 1.5 GB more room for experts on the GPU.
- Maya-M (116 GB) is closer to the full model, for PCs with a bigger memory pool: built with the same FP8-statistics recipe and more bits where they count (2- to 3-bit experts), calibrated toward tool calls and front-end code.
- Maya-L (156.3 GB) is the closest, for PCs with the biggest memory pool: Maya-M's recipe one step up (3- to 4-bit experts: IQ3_S / IQ4_XS, Q5_K in the most sensitive layers).
Built partly on the IST Austria DAS Lab (ISTA-DASLab) recipe. All four use their GPTQ-style error-feedback rounding for the experts (each expert rounded against its own input statistics from the FP8 model), and are measured the way they measure their quants: task accuracy against the full-precision model.
Files
| Folder | Files | Size |
|---|---|---|
Maya-S-v2-IQ2_XXS/ |
GLM-5.3-Flash-Maya-S-v2-IQ2_XXS-00001-of-00003.gguf, -00002-of-00003.gguf, -00003-of-00003.gguf |
96.5 GB |
Maya-S24/ |
GLM-5.3-Flash-Maya-S24-IQ2_XXS_S-00001-of-00003.gguf, -00002-of-00003.gguf, -00003-of-00003.gguf |
94.7 GB |
Maya-M/ |
GLM-5.3-Flash-Maya-M-IQ2_S-00001-of-00003.gguf, -00002-of-00003.gguf, -00003-of-00003.gguf |
116.0 GB |
Maya-L/ |
GLM-5.3-Flash-Maya-L-IQ3_S-00001-of-00004.gguf, -00002-, -00003-, -00004-of-00004.gguf |
156.3 GB |
vision/ |
mmproj-GLM-5.3-Flash-F16.gguf (the vision tower and projector, F16, from the official weights), GLM-5.3-Flash-vocab.gguf (the tokenizer, for the encoder) |
1.14 GB |
The model's NextN (MTP) layer is included in all four, so engines that draft with it get speculative decoding.
What is in them
| Tensors | Maya-S | Maya-S24 | Maya-M | Maya-L |
|---|---|---|---|---|
| Routed experts, gate and up (42 layers) | IQ2_XXS, rounded with error feedback | as Maya-S | IQ2_S, rounded with error feedback | IQ3_S, rounded with error feedback |
| Routed experts, down | IQ3_XXS in the first and last four MoE layers, IQ2_S in the rest | as Maya-S | IQ3_S in the first and last four MoE layers, IQ3_XXS in the rest | Q5_K in the first and last four MoE layers, IQ4_XS in the rest |
| Attention projections (KDA, MLA/DSA) and shared experts | Q6_K | Q4_K | Q6_K | Q6_K |
| The three dense layers, embeddings, output, MLA k_b / v_b | Q6_K | Q6_K | Q6_K | Q6_K |
| Small KDA projections, the DSA indexer | Q8_0 | Q8_0 | Q8_0 | Q8_0 |
| Router, norms, stream-mixing weights | F32 | F32 | F32 | F32 |
| NextN draft layer experts | Q2_K (gate/up), Q3_K (down) | as Maya-S | Q3_K (gate/up), Q4_K (down) | Q4_K (gate/up), Q5_K (down) |
How they were made
- Calibration text: 128 sequences of 2,048 tokens in GLM's own chat template, reasoning blocks included - chat, multilingual chat, reasoning traces, web code (single-file HTML/CSS/JS pages, three.js scenes, canvas and WebGL animations), other code and tool calls; Maya-M's and Maya-L's weighted further toward tool calls and front-end code.
- Statistics from the FP8 model itself, layer by layer: an importance matrix for every expert separately, not one per layer.
- Error-feedback rounding of the experts' gate and up projections (GPTQ-style, at the quantizer's 256-weight blocks, each expert with its own input statistics): on held-out text it lowers the experts' output error by about a quarter against plain importance-weighted rounding at the same size.
- llama.cpp's own quantizers (ggml), so the files are ordinary GGUFs.
Against the original (FP8)
Maya-L keeps 99.2% of the full model's zero-shot accuracy; Maya-S and Maya-M 97.9%, Maya-S24 97.7%. Multiple-choice tasks at 400 questions each do not separate them (a point or two either way is within the test's noise); token by token, below, Maya-M is clearly closer to the FP8 model.
| Task (zero-shot) | FP8 | Maya-S | Maya-S24 | Maya-M | Maya-L |
|---|---|---|---|---|---|
| ARC-Easy (acc) | 87.2 | 86.2 (98.9%) | 85.8 (98.3%) | 86.5 (99.1%) | 86.0 (98.6%) |
| ARC-Challenge (acc norm) | 71.0 | 68.2 (96.1%) | 69.0 (97.2%) | 69.0 (97.2%) | 69.8 (98.2%) |
| HellaSwag (acc norm) | 88.5 | 87.5 (98.9%) | 84.8 (95.8%) | 86.8 (98.0%) | 88.5 (100%) |
| WinoGrande (acc) | 78.5 | 75.5 (96.2%) | 77.0 (98.1%) | 76.2 (97.1%) | 77.8 (99.0%) |
| PIQA (acc norm) | 87.0 | 86.2 (99.1%) | 86.2 (99.1%) | 85.0 (97.7%) | 87.0 (100%) |
| Average | 82.5 | 80.8 (97.9%) | 80.6 (97.7%) | 80.7 (97.9%) | 81.8 (99.2%) |
400 questions per task, the same questions for every model, scored the way lm-evaluation-harness scores them (the answer with the highest log-likelihood; length-normalized where the choices differ in length) - the FP8 model run layer by layer in PyTorch, the quants through Project Maya's engine (tools/maya_quant/zs_*.py).
Token by token, on held-out text never used for calibration (7,672 positions), with the FP8 model's own predictions as the reference:
| FP8 | Maya-S | Maya-S24 | Maya-M | Maya-L | |
|---|---|---|---|---|---|
| Same top token as FP8 | 100% | 83.3% | 83.1% | 86.2% | 90.0% |
| KL divergence from FP8 | 0 | 0.428 | 0.444ยน | 0.329 | 0.188ยฒ |
| Top-1 accuracy on the actual next token | 71.5% | 68.8% | 67.9% | 70.3% | 71.2% |
ยน Measured with a later engine, next to Maya-S at 0.421 in the same run. ยฒ Measured with a later engine, next to Maya-M at 0.291 (same top token 86.8%) in the same run.
Closest on chat and reasoning text, furthest on tool-call formats and on text the FP8 model has memorized.
No loops in long generation (Maya-S): 14 answers of 6,000-14,000 tokens (three.js scenes, canvas animations, explanations, a plan in Portuguese, a proof; temperature 1.0 and greedy) without a repeated passage. The context fill test (a fact at the start of an 8k, 16k and 30k-token context, asked about at the end) finds it every time.
Speed
Maya-S on two Tesla V100 32 GB (30 GB RAM) with Project Maya: decode (writing the answer) up to 40 tokens/s, with
the MTP block drafting, and it keeps that pace at long context (60K tokens); prefill (reading the prompt) up to 560
tokens/s (up to 670 with Project Maya v1.0.15).
On one Tesla V100 32 GB with 64 GB of RAM: decode up to 19 tokens/s, prefill up to 370 tokens/s.
Maya-S24 on the same card limited to 24 GB: decode 13.5 tokens/s against Maya-S's 11.8 (+14%); on the full 32 GB
19.1 against 17.2 (+11%); prefill the same.
./maya.sh --bench measures your machine the same way.
Running it
Project Maya sets everything up and runs it:
git clone https://github.com/mw00/project-maya.git && cd project-maya
./setup.sh # Windows (experimental): START-MAYA.bat
- It checks the PC, downloads the model you pick and verifies every file's sha256, compiles the engine for your
GPU(s), and starts it. Press Enter at each question for the recommended choice: Maya-S, or Maya-S24 when every card
has 24 GB or less.
./setup.sh --setup --model Maya-S24or--model Maya-Mor--model Maya-Lsets up the others. - GPUs: NVIDIA, V100 / RTX 20 or newer, one GPU or up to 16 that share the model. AMD (experimental, Linux, ROCm 7): RX 7900 XT / XTX and Radeon AI PRO R9700 / RX 9070 (one or two), Strix Halo (one).
- Memory: the most-used experts stay in VRAM, the next in RAM and the rest are read from the SSD, so it runs with 32 GB of RAM; more VRAM and RAM is faster. Put the model on an NVMe SSD.
- Use it: the dashboard at
http://127.0.0.1:8080(chat, pictures through the vision files above, a live monitor), or the OpenAI- and Anthropic-compatible API at the same address (streaming and tool calls; Claude Code:ANTHROPIC_BASE_URL=http://127.0.0.1:8080). ./maya.sh --benchmeasures your machine;--benchand--reportin a GitHub issue add it to the users' speeds.
License
The model is Z.ai's GLM-5.3-Flash under the MIT license; these quantizations are released under the same license.
- Downloads last month
- 13,089
Model tree for peasantsmith/GLM-5.3-Flash-Maya-GGUF
Base model
zai-org/GLM-5.3-Flash