StarVLA QwenDual for the BEHAVIOR Challenge, task 1 (100K)

This repository contains the 100,000-step checkpoint from the 1117_BEHAVIOR_challenge_QwenDual_task1 run. The packaged configuration limits the BEHAVIOR_challenge mixture to task_id: 1 and uses state-conditioned StarVLA QwenDual inference.

Model details

Item Published configuration
Framework StarVLA QwenDual
VLM Local training snapshot named Qwen3-VL-4B-Instruct; revision not recorded
Auxiliary visual encoder dinov2_vits14
Action model 16-layer DiT-B flow head: 768 latent width, 12 heads (64 dimensions/head); state/action decoder MLP width 1,024
Action / state dimension 23 / 44
Predicted action chunk 50 steps
Inference diffusion steps 4
Image resolution 224 x 224
State input Enabled
Dataset mixture / task BEHAVIOR_challenge, task_id: 1
Normalization key R1Pro
Uploaded checkpoint checkpoints/steps_100000_pytorch_model.pt

The repository does not record the number or ordering of camera views. The 44D state and 23D action contracts must be reproduced through the StarVLA data adapter and dataset_statistics.json; they should not be inferred only from the checkpoint tensor shapes.

Training details

Setting Value in config.yaml
Maximum / released step 100,000 / 100,000
Per-device VLA batch size 32
Gradient accumulation 1
Warm-up steps 5,000
Base / interface / action LR 4e-5 / 1e-5 / 1e-4
Optimizer AdamW, betas (0.9, 0.95), epsilon 1e-8
Scheduler Cosine with minimum LR 1e-6
VLA / VLM loss scale 1.0 / 0.1
freeze_modules Packaged boolean true; the public trainer expects module paths as a string, so this value names/selects no modules
Gradient checkpointing / mixed precision Enabled / enabled
Seed 42

Files

config.yaml
dataset_statistics.json
summary.jsonl
checkpoints/
└── steps_100000_pytorch_model.pt

summary.jsonl enumerates checkpoints from 2K through 100K, but contains no evaluation metrics.

Evaluation status

No challenge submission identifier, episode-level output, success rate, or aggregate score is included in this repository. This card therefore makes no claim that the 100K checkpoint corresponds to an official BEHAVIOR Challenge submission. A result should be added only with a public evaluation artifact that identifies this exact checkpoint and task protocol.

Loading and evaluation

huggingface-cli download StarVLA/1117_BEHAVIOR_challenge_QwenDual_task1 \
  --local-dir 1117_BEHAVIOR_challenge_QwenDual_task1

export CKPT=$PWD/1117_BEHAVIOR_challenge_QwenDual_task1/checkpoints/steps_100000_pytorch_model.pt
python deployment/model_server/server_policy.py \
  --ckpt_path "$CKPT" \
  --port 10093 \
  --use_bf16

Run the simulator from the maintained BEHAVIOR example. Use the R1Pro normalization key and confirm the 44D state, 23D action, and camera ordering reported by the policy-server metadata.

Intended use and limitations

The checkpoint is scoped to BEHAVIOR task 1 with the R1Pro observation/action mapping. Results are not published, the exact upstream VLM revision is absent, and transfer to other BEHAVIOR tasks is unverified. Simulator execution is GPU- and Vulkan-dependent, and real-robot use has not been established. This checkpoint is not safety-tuned.

Downloads last month
55
Video Preview
loading

Model tree for StarVLA/1117_BEHAVIOR_challenge_QwenDual_task1

Finetuned
(456)
this model

Collection including StarVLA/1117_BEHAVIOR_challenge_QwenDual_task1