Aura-2-Lightning β‘
A high-throughput, structurally distilled causal language model engineered for local deployment, low-latency execution, and reasoning density. Developed under the Waveforce AI framework, Aura-2-Lightning aligns a compact, responsive student backbone with the latent internal states of a large-scale Mixture-of-Experts foundation model.
π¬ Architectural Strategy & Distillation Profile
Unlike standard token-level distillation that solely minimizes logit divergence over output tokens, Aura-2-Lightning utilizes Latent Space State Projection Alignment.
During cross-architecture training:
- Teacher Architecture:
openai/gpt-oss-20bwas hosted in native Microscaling FP4 (MXFP4) block-scaled quantization, utilizing dedicated Triton matrix execution kernels to run the 20B MoE layer states under low memory overhead. - Student Architecture:
openai-community/gpt2-xl(1.5B parameters, FP16) functioned as the primary student model. - Latent Mapping: An intermediate linear projection head dynamically mapped the student's 1600-dimensional representation space directly into the teacher's 2880-dimensional hidden states: $$\mathcal{L}{\text{total}} = \mathcal{L}{\text{CausalLM}} + 0.5 \cdot \text{MSE}(\mathbf{W}{\text{proj}} \cdot \mathbf{h}{\text{student}}, \mathbf{h}_{\text{teacher}})$$
- Optimization Stability: The network trained using 8-bit paged AdamW (
PagedAdamW8bit) with strict gradient norm clipping (max_norm=1.0) and $-100$ loss masking across padding tokens, eliminating boundary drift and early<|endoftext|>halts.
π₯ Model Ancestry & Credits
- Student Base Architecture:
openai-community/gpt2-xl(1.5B) β Chosen for its mature, unencumbered Transformer decoder architecture and low execution latency. - Teacher Foundation Model:
openai/gpt-oss-20bβ Native 4-bit microscaled MoE foundation model supplying the target latent hidden states.
π Plug-and-Play Inference
Aura-2-Lightning ships with an optimized generation_config.json containing pre-tuned repetition penalty safeguards and $n$-gram constraints.
1. Minimal Pipeline Usage (Recommended)
import torch
from transformers import pipeline
# Load directly without manual generation kwargs
generator = pipeline(
"text-generation",
model="waveforce-ai/Aura-2-Lightning",
device_map="auto",
torch_dtype=torch.float16,
clean_up_tokenization_spaces=False
)
prompt = "Waveforce AI designs"
output = generator(prompt)
print(output[0]["generated_text"])
π¬ Feedback
We would love to know how we can improve Aura-2-Lightning.
Rate the model and give your feedback on:
https://arthwarrior3201.github.io/Aura-2-Lightning-Feedback/
- Downloads last month
- 1,554