Qwen3-VL-4B OFT for RoboCasa-GR1 (90K)
This repository contains a StarVLA QwenOFT policy trained jointly on the 24
RoboCasa-GR1 tabletop tasks. It is a complete StarVLA state dictionary, not a
Transformers from_pretrained() directory and not an adapter-only release.
Checkpoint identity
| Item | Value |
|---|---|
| Released file | checkpoints/steps_90000_pytorch_model.pt |
| Training step | 90,000 |
| Hub revision checked | 664d79ba4e820d8dd08bf7294a730f408a7ef2d1 |
| File size | 9,785,286,394 bytes |
| SHA-256 / LFS object ID | 66a114db658b4b2dda8fd5530bfe3c7cd44b3c37e49c6fb68be22bab5e8f4bf5 |
summary.jsonl records a run continuing beyond this save, but no later
checkpoint is present in this repository. Do not infer that a 100K artifact is
available.
Model and control contract
| Item | Value |
|---|---|
| Framework | StarVLA QwenOFT |
| Base VLM | Qwen3-VL-4B-Instruct |
| Action head | Two-block residual MLP, 2,560 input / 5,120 hidden / 29 output; direct L1 regression |
| Observation | Language instruction + one ego RGB image at 224 x 224 |
| Robot state | Not provided to this policy |
| Action chunk | 16 x 29 (future_action_window_size: 15) |
| Action layout | left arm 7 + right arm 7 + left hand 6 + right hand 6 + waist 3 |
| Normalization key | gr1 |
| Recommended execution | 12 actions from each predicted chunk |
The no-state boundary is important: sending a state key makes QwenOFT
append state tokens to the instruction, which does not match this checkpoint's
training prompt.
Training data and settings
The packaged statistics describe the fourier_gr1_unified_1000 mixture:
24,000 trajectories and 6,020,058 transitions, with 29-D state and action
statistics under gr1.
| Setting | Value |
|---|---|
| Per-device VLA batch size | 8 |
| Gradient accumulation | 1 |
| Optimizer | AdamW, betas (0.9, 0.95), epsilon 1e-8, weight decay 1e-8 |
| Base / interface / action LR | 3e-5 / 1e-5 / 1e-4 |
| Schedule | Cosine; 5,000 warmup steps |
freeze_modules |
Packaged boolean true; the public trainer expects module paths as a string, so this value names/selects no modules |
| Seed | 42 |
| Configured run target | 100,000 steps; released checkpoint is 90,000 |
| Training GPU count | Missing from the public artifact |
Repository-reported evaluation
The StarVLA RoboCasa README names this exact 90K path and reports 48.8% mean success over 24 tasks, 50 rollouts per task (1,200 rollouts total). Raw per-episode logs are not packaged here, so this result is repository-reported rather than independently reconstructable from the Hub artifact.
Reproduction requires the repository protocol: 50 episodes, a 720-step episode
limit, 12 executed actions per query, the gr1 statistics, and
--args.no_send_state.
Download and load
hf download StarVLA/Qwen3-VL-OFT-Robocasa \
--local-dir playground/Pretrained_models/Qwen3-VL-OFT-Robocasa
export CKPT=playground/Pretrained_models/Qwen3-VL-OFT-Robocasa/checkpoints/steps_90000_pytorch_model.pt
python deployment/model_server/server_policy.py \
--ckpt_path "$CKPT" \
--config_override framework.qwenvl.base_vlm=Qwen/Qwen3-VL-4B-Instruct \
--port 5678 \
--use_bf16
Keep config.yaml and dataset_statistics.json in the downloaded run
directory. For simulation commands, follow
examples/simBenchmarks/Robocasa_tabletop/README.md and retain
--args.no_send_state.
Intended use and limitations
This checkpoint is for RoboCasa-GR1 simulation research under the observation,
normalization, and 29-D action mapping above. Transfer to other robots, camera
calibrations, state-conditioned prompts, or real hardware has not been
established. Loading a pickle-based .pt file executes PyTorch deserialization;
use artifacts only from a trusted revision.
License status
This target repository did not previously publish a Model Card or a separate
LICENSE file. The checkpoint's weight license therefore needs maintainer
confirmation; the Qwen3-VL base-model terms and applicable dataset terms still
apply.
- Downloads last month
- 31
Model tree for StarVLA/Qwen3-VL-OFT-Robocasa
Base model
Qwen/Qwen3-VL-4B-Instruct