Docs | Blog | GitHub | Vela 2.0 collection
Vela 2.0 9B
Open Foundation Routing Models
Routing decisions. Safety checks. Precise text spans.
The largest hybrid member of Vela 2.0: high-accuracy multilingual routing, long-document PII and hallucination detection, with open-label spans through one interface.
Define options, labels and rubrics at request time. Ask multiple named questions about a request, context and answer, and receive structured decisions with the text spans that support your workflow.
- High-accuracy decisions. A 41.63 Jev Decision Index with optional Noul calibration, plus 0.989 AUC on unseen prompt-attack families.
- Routing and safety together. Use one request for routing, prompt-attack checks, PII and unsupported-claim detection.
- Decisions at span resolution. Return precise text locations for trained router labels and open-label extraction, with probabilities and character offsets.
| Specification | Value |
|---|---|
| Reported parameters | 7.9B |
| Backbone | 32-layer Qwen3.5 hybrid: Gated DeltaNet and gated GQA |
| Input limit | 16,384 tokens |
| Output | Choice, Yes/no (Noul), Score, Span and Set; no text generation |
| Evaluated precision | FP32 parameters; bf16 backbone autocast on GPU; FP32 heads |
| GPU parameter memory | About 32 GB in FP32 |
| Licence | Apache-2.0 |
Quickstart
Install the dependencies, load the model and ask for routing and spans in one request:
For focused examples, see Span questions and Set questions below.
pip install torch "transformers>=5.17" safetensors tokenizers numpy
pip install flash-linear-attention # GPU: the Gated-DeltaNet kernel used in evaluation (optional, much faster)
from transformers import AutoModel
m = AutoModel.from_pretrained("vllm-sr/Vela-2.0-9B", trust_remote_code=True)
m = m.to("cuda") # parameters load in FP32; on GPU the backbone runs under bf16 autocast (heads FP32)
PII_LABELS = m.vela2_engine.cal["pii_schema"]["labels"] # the 17 trained PII types
result = m.system_one(
state={"request": "Hi, I'm Tom Baker ([email protected]). What is the maximum daily dose of paracetamol for an adult?",
"source": "For adults, the maximum dose of paracetamol is 4 grams in 24 hours, taken as 500 mg to 1 g every 4 to 6 hours.",
"answer": "Adults can take up to 6 grams of paracetamol in 24 hours, in doses of 500 mg to 1 g every 4 to 6 hours."},
questions={
"pii": {"type": "span", "instructions": "Which spans are personal information?", "criteria": PII_LABELS,
"over": "request"},
"halu": {"type": "span", "instructions": "Which spans of the answer are not supported by the context?",
"criteria": {"unsupported": "a claim not supported by the context"}}, # over the answer by default
"domain": {"type": "choice", "instructions": "Which subject area is this request about?", "over": "request",
"criteria": {"health": "medicine, clinical practice, nutrition, ageing or sexual health",
"math": "arithmetic, algebra, geometry, statistics or other mathematics",
"other": "a subject that fits none of the listed areas"}},
})
Recorded response excerpt from this release; the full response includes usage and thresholds:
{
"answers": {
"domain": {
"type": "choice",
"choice": "health",
"confidence": 0.977,
"probabilities": {
"health": 0.988,
"math": 0.0,
"other": 0.012
}
}
},
"spans": {
"pii": [
{
"label": "PERSON",
"start": 8,
"end": 17,
"text": "Tom Baker",
"probability": 0.999
},
{
"label": "EMAIL_ADDRESS",
"start": 19,
"end": 40,
"text": "[email protected]",
"probability": 1.0
}
],
"halu": [
{
"label": "unsupported",
"start": 22,
"end": 29,
"text": "6 grams",
"probability": 0.982
}
]
},
"span_heads": {
"pii": "router",
"halu": "router"
}
}
GPU with bf16 autocast is the evaluated setting. FP32 on CPU or GPU is also supported; fp16 is not. Parameter memory above excludes runtime allocations. See precision and execution.
Span questions: locate text
Use "type": "span" to extract text with labels and character offsets. Set criteria to a label-to-description mapping and over to the state field to inspect. This PII example reuses m and the trained PII_LABELS from above:
span_result = m.system_one(
state={"request": "Hi, I'm Tom Baker ([email protected])."},
questions={
"pii": {
"type": "span",
"instructions": "Which spans are personal information?",
"criteria": PII_LABELS,
"over": "request",
}
},
)
print(span_result["spans"]["pii"])
Each returned span includes label, text, start, end and probability. Offsets use Unicode code points, with start inclusive and end exclusive, in the selected field. The first example also shows hallucination detection: use the unsupported label, target over="answer", and supply grounding text in state["source"].
Two span heads, one interface. The router span head handles trained PII, hallucination and toxic labels. The broad span head handles open-label extraction. The engine selects a head for each span question; you can also set "head": "router" or "head": "broad". See span-head selection. For entities, relations and evidence examples, see open-label extraction.
Set questions: select multiple labels
Use "type": "set" when zero, one or several labels can apply. Supply your labels and descriptions in criteria; each label is scored independently. This example reuses m from above:
set_result = m.system_one(
state="My card was charged twice and the parcel never arrived.",
questions={
"issues": {
"type": "set",
"instructions": "Which issues does the customer report?",
"criteria": {
"billing": "payments, charges, refunds or invoices",
"shipping": "delivery of an order or a parcel",
"login": "signing in, passwords or account access",
},
}
},
)
print(set_result["sets"]["issues"]["selected"])
print(set_result["sets"]["issues"]["probabilities"])
Read the selected labels and all per-label probabilities from sets["issues"]. Probabilities do not need to sum to one. Labels are selected when their probability exceeds the question's threshold; omit threshold to use shipped calibration, or add it to the question to set your own cutoff. The applied value is returned in set_result["thresholds"]["issues"].
Questions and outputs
| Question | Use it for | Output |
|---|---|---|
| Choice | Route a request or select one of 2–255 supplied options. | Selected key and distribution |
| Yes/no (Noul) | Check a condition against the state. | P(yes) |
| Score | Rate against 2–10 ordered levels. | Expected level and distribution |
| Span | Locate personal information, unsupported claims, entities, relations or evidence. | Text spans, character offsets and probabilities |
| Set | Select any number of supplied labels. | Selected labels and per-label probabilities |
Span offsets refer to Unicode code points in the selected text. Span and Set answers also provide Yes/no views in the SystemOne response. Output semantics and thresholds explain these fields.
Results
Selected results for this checkpoint. Full evaluation includes every benchmark, family comparison and scoring protocol.
| Benchmark | Result | Metric and mode |
|---|---|---|
| Safety, 14 public sets | 0.921 | Mean AUC; trained task families |
| Prompt attacks, unseen families | 0.989 | AUC |
| PII, 8K-token documents | 0.940 | F1; shipped calibration |
| Hallucination, 10,698 examples | 0.885 | Example-F1 |
| ACL-Verbatim evidence | 24.5 | Word-F1; broad span head, zero-shot |
| Jev Decision Index 0.2.1 | 41.63 | 38 benchmarks; noul_calibration=True |
- Evaluation: checkpoint selection, temperatures and thresholds used dev data. Safety results cover trained task families; they are not zero-shot comparisons.
- Calibration: Jev uses optional
noul_calibration=True; the default score is 41.09. PII uses shipped calibration. All evaluation modes.
Architecture

This model has its own 32-layer hybrid backbone, with Gated DeltaNet and gated GQA blocks, a candidate head, and separate router/broad span-head weights. Editable SVG.
Operator and readout diagrams
Gated DeltaNet

Causal QKV convolution feeds the gated-delta update, followed by per-head normalization and output gating. Editable SVG.
Gated GQA and SwiGLU

Gated GQA uses partial RoPE on Q/K and a sigmoid output gate; SwiGLU uses separate up and gate projections. Editable SVG.
Candidate readout

CandidateHead combines scaled bilinear and additive MLP scores over runtime options, then applies task-specific calibration. Editable SVG.
Router and broad span heads

Router and broad span heads have separate weights, with one selected for each span question. Overlapping word logits are combined before decoding. Editable SVG.
State prefix and question isolation

For a fixed rendered state, each question block sees the same state prefix and its own causal tokens. Additional span questions use separate sequences. Editable SVG.
Choice, Yes/no, Score and Set share a candidate readout. Router and broad span heads use the same word-by-label topology with separate weights. The hybrid backbone reuses the state prefix within each rendered sequence; execution details are in USAGE.md.
Reference
| Document | Contents |
|---|---|
| USAGE.md | Full examples and recorded responses, typed parts, head selection, calibration, SDK and HTTP serving |
| EVALUATION.md | Complete family results, research versus shipped PII modes, benchmark protocols and evaluation disclosures |
| TRAINING.md | Training stages, data mixtures, provenance, licences and release files |
| PARITY.md | Export parity against the research scorer |
Credit and citation
Vela 2.0 is led by KR Labs and vLLM Semantic Router.
Read the Vela 2.0 technical overview.
@misc{vela2_9b_2026,
title = {Vela 2.0: Towards Open Foundation Routing Models},
author = {{KR Labs} and {vLLM Semantic Router}},
year = {2026},
note = {Blog post. Model: vllm-sr/Vela-2.0-9B},
howpublished = {\url{https://vllm-sr.ai/blog/vela-2-0-open-foundation-routing-models/}}
}
Licence
Apache-2.0 for this model's weights, code and documentation. It is derived from Decision-2.0-Lux-9B (Apache-2.0), itself built on Qwen3.5-9B (Apache-2.0); their licences and notices are passed on in LICENSE, NOTICE, ATTRIBUTIONS.md and LICENSES/. Changes against Decision-2.0-Lux-9B: MODIFICATIONS.md. Training data keep their own licences, some of them CC-BY-SA share-alike.
- Downloads last month
- 75
Model tree for vllm-sr/Vela-2.0-9B
Base model
Qwen/Qwen3.5-9B-Base