Aura-2-Lightning ⚑

A high-throughput, structurally distilled causal language model engineered for local deployment, low-latency execution, and reasoning density. Developed under the Waveforce AI framework, Aura-2-Lightning aligns a compact, responsive student backbone with the latent internal states of a large-scale Mixture-of-Experts foundation model.


πŸ”¬ Architectural Strategy & Distillation Profile

Unlike standard token-level distillation that solely minimizes logit divergence over output tokens, Aura-2-Lightning utilizes Latent Space State Projection Alignment.

During cross-architecture training:

  1. Teacher Architecture: openai/gpt-oss-20b was hosted in native Microscaling FP4 (MXFP4) block-scaled quantization, utilizing dedicated Triton matrix execution kernels to run the 20B MoE layer states under low memory overhead.
  2. Student Architecture: openai-community/gpt2-xl (1.5B parameters, FP16) functioned as the primary student model.
  3. Latent Mapping: An intermediate linear projection head dynamically mapped the student's 1600-dimensional representation space directly into the teacher's 2880-dimensional hidden states: $$\mathcal{L}{\text{total}} = \mathcal{L}{\text{CausalLM}} + 0.5 \cdot \text{MSE}(\mathbf{W}{\text{proj}} \cdot \mathbf{h}{\text{student}}, \mathbf{h}_{\text{teacher}})$$
  4. Optimization Stability: The network trained using 8-bit paged AdamW (PagedAdamW8bit) with strict gradient norm clipping (max_norm=1.0) and $-100$ loss masking across padding tokens, eliminating boundary drift and early <|endoftext|> halts.

πŸ‘₯ Model Ancestry & Credits

  • Student Base Architecture: openai-community/gpt2-xl (1.5B) β€” Chosen for its mature, unencumbered Transformer decoder architecture and low execution latency.
  • Teacher Foundation Model: openai/gpt-oss-20b β€” Native 4-bit microscaled MoE foundation model supplying the target latent hidden states.

πŸš€ Plug-and-Play Inference

Aura-2-Lightning ships with an optimized generation_config.json containing pre-tuned repetition penalty safeguards and $n$-gram constraints.

1. Minimal Pipeline Usage (Recommended)

import torch
from transformers import pipeline

# Load directly without manual generation kwargs
generator = pipeline(
    "text-generation",
    model="waveforce-ai/Aura-2-Lightning",
    device_map="auto",
    torch_dtype=torch.float16,
    clean_up_tokenization_spaces=False
)

prompt = "Waveforce AI designs"
output = generator(prompt)

print(output[0]["generated_text"])

πŸ’¬ Feedback

We would love to know how we can improve Aura-2-Lightning.

Rate the model and give your feedback on:

https://arthwarrior3201.github.io/Aura-2-Lightning-Feedback/

Downloads last month
1,554
Safetensors
Model size
2B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ 1 Ask for provider support

Model tree for waveforce-ai/Aura-2-Lightning

Finetuned
(63)
this model
Quantizations
2 models