Instructions to use IndexTeam/Index-Echo-S2ST-2B-FP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use IndexTeam/Index-Echo-S2ST-2B-FP4 with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("translation", model="IndexTeam/Index-Echo-S2ST-2B-FP4")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("IndexTeam/Index-Echo-S2ST-2B-FP4", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Index-Echo-S2ST-2B-FP4
Official NVFP4 (W4A4) quantization of IndexTeam/Index-Echo-S2ST-2B, part of the Index-Echo speech-to-speech translation (S2ST) model family by bilibili.
This repository mirrors the original checkpoint layout; only the text LLM backbone (stlm_llm/) is quantized to NVFP4 - the audio tower, connector, and all speech-synthesis components remain in BF16. Use it exactly like the original repo (same infer.py / configs).
Quantization
- Scheme:
NVFP4(W4A4: 4-bit floating-point weights with per-group-16 scales, 4-bit floating-point activations with calibrated per-tensor global scales), produced with llm-compressor (quantization_schemerecorded inrecipe.yaml). Calibrated on a small bilingual translation corpus. - All
Linearlayers of the language model backbone (stlm_llm/) are quantized; the audio tower, connector,lm_head, embeddings and all other pipeline components are kept in BF16. - Format: compressed-tensors
nvfp4-pack-quantizedsafetensors - load directly with vLLM (quantization="compressed-tensors") or transformers.
Consistency validation
Measured on an NVIDIA A100 (weight-dequantized execution) against the original BF16 checkpoint (greedy decoding, official translation prompt):
| Metric | BF16 | FP4 | Delta |
|---|---|---|---|
| Perplexity (fixed corpus) | 5.9332 | 6.4980 | +9.52% |
| zh->en generation identical | - | - | yes |
| en->zh generation identical | - | - | yes |
Usage
Identical to the original checkpoint - clone this repo and follow the README / infer.py of the base model (IndexTeam/Index-Echo-S2ST-2B). The quantized LLM backbone loads via compressed-tensors; make sure compressed-tensors (or vLLM / a recent transformers) is installed.
Hardware note: full NVFP4 (W4A4) acceleration requires an NVIDIA Blackwell GPU (SM100+, e.g. B200 / RTX 50 series). On older GPUs vLLM falls back to weight-only dequantization, which still reduces memory but gives no FP4 speedup.
Quantized and published by the Index team, 2026-10-04.
- Downloads last month
- 22
Model tree for IndexTeam/Index-Echo-S2ST-2B-FP4
Base model
IndexTeam/Index-Echo-S2ST-2B