Qwen3-VL-4B OFT for RoboCasa-GR1 (90K)

This repository contains a StarVLA QwenOFT policy trained jointly on the 24 RoboCasa-GR1 tabletop tasks. It is a complete StarVLA state dictionary, not a Transformers from_pretrained() directory and not an adapter-only release.

Checkpoint identity

Item Value
Released file checkpoints/steps_90000_pytorch_model.pt
Training step 90,000
Hub revision checked 664d79ba4e820d8dd08bf7294a730f408a7ef2d1
File size 9,785,286,394 bytes
SHA-256 / LFS object ID 66a114db658b4b2dda8fd5530bfe3c7cd44b3c37e49c6fb68be22bab5e8f4bf5

summary.jsonl records a run continuing beyond this save, but no later checkpoint is present in this repository. Do not infer that a 100K artifact is available.

Model and control contract

Item Value
Framework StarVLA QwenOFT
Base VLM Qwen3-VL-4B-Instruct
Action head Two-block residual MLP, 2,560 input / 5,120 hidden / 29 output; direct L1 regression
Observation Language instruction + one ego RGB image at 224 x 224
Robot state Not provided to this policy
Action chunk 16 x 29 (future_action_window_size: 15)
Action layout left arm 7 + right arm 7 + left hand 6 + right hand 6 + waist 3
Normalization key gr1
Recommended execution 12 actions from each predicted chunk

The no-state boundary is important: sending a state key makes QwenOFT append state tokens to the instruction, which does not match this checkpoint's training prompt.

Training data and settings

The packaged statistics describe the fourier_gr1_unified_1000 mixture: 24,000 trajectories and 6,020,058 transitions, with 29-D state and action statistics under gr1.

Setting Value
Per-device VLA batch size 8
Gradient accumulation 1
Optimizer AdamW, betas (0.9, 0.95), epsilon 1e-8, weight decay 1e-8
Base / interface / action LR 3e-5 / 1e-5 / 1e-4
Schedule Cosine; 5,000 warmup steps
freeze_modules Packaged boolean true; the public trainer expects module paths as a string, so this value names/selects no modules
Seed 42
Configured run target 100,000 steps; released checkpoint is 90,000
Training GPU count Missing from the public artifact

Repository-reported evaluation

The StarVLA RoboCasa README names this exact 90K path and reports 48.8% mean success over 24 tasks, 50 rollouts per task (1,200 rollouts total). Raw per-episode logs are not packaged here, so this result is repository-reported rather than independently reconstructable from the Hub artifact.

Reproduction requires the repository protocol: 50 episodes, a 720-step episode limit, 12 executed actions per query, the gr1 statistics, and --args.no_send_state.

Download and load

hf download StarVLA/Qwen3-VL-OFT-Robocasa \
  --local-dir playground/Pretrained_models/Qwen3-VL-OFT-Robocasa

export CKPT=playground/Pretrained_models/Qwen3-VL-OFT-Robocasa/checkpoints/steps_90000_pytorch_model.pt
python deployment/model_server/server_policy.py \
  --ckpt_path "$CKPT" \
  --config_override framework.qwenvl.base_vlm=Qwen/Qwen3-VL-4B-Instruct \
  --port 5678 \
  --use_bf16

Keep config.yaml and dataset_statistics.json in the downloaded run directory. For simulation commands, follow examples/simBenchmarks/Robocasa_tabletop/README.md and retain --args.no_send_state.

Intended use and limitations

This checkpoint is for RoboCasa-GR1 simulation research under the observation, normalization, and 29-D action mapping above. Transfer to other robots, camera calibrations, state-conditioned prompts, or real hardware has not been established. Loading a pickle-based .pt file executes PyTorch deserialization; use artifacts only from a trusted revision.

License status

This target repository did not previously publish a Model Card or a separate LICENSE file. The checkpoint's weight license therefore needs maintainer confirmation; the Qwen3-VL base-model terms and applicable dataset terms still apply.

Downloads last month
31
Video Preview
loading

Model tree for StarVLA/Qwen3-VL-OFT-Robocasa

Finetuned
(463)
this model

Collection including StarVLA/Qwen3-VL-OFT-Robocasa