Current Versions:
- Freedream_1B-7b-tokens-seen: This is an early trained model, that only see total of 7Billion tokens, it is a base model that has a moderate performance that can generate fluence text but contains factual error. Don't recommend for code or hard maths reasoning.
Reserved Name & Trademark Notice
The name "Freedream" is a reserved name and an unregistered trademark of the original author.
1. Reserved Name
No derivative work, fork, fine-tune, quantization, or redistribution of this model may use the name "Freedream" or any confusingly similar name (e.g., "Free-Dream", "FreeDream", "Freedream AI", "Freedream-2", "Freedream-X", etc.) without prior written permission from the original author.
2. No Trademark License
This license does not grant any rights to use the trademark "Freedream". All trademark rights are expressly reserved by the original author.
3. Attribution Requirement
Any use, distribution, or public deployment of this model must:
- Prominently display "Built with Freedream" or "Based on Freedream".
- Provide a link to the original model repository.
- Retain this license file in its entirety.
4. Enforcement
Violation of the Reserved Name or Trademark clauses will result in a takedown request to the hosting platform (e.g., Hugging Face) and may result in legal action under applicable trademark and contract law.
Freedream 1B
Model Description:
Freedream is a ~1B parameter decoder-only Transformer language model trained from scratch on a single NVIDIA RTX 5080 (16GB). It uses RoPE, SwiGLU, weight-tied embeddings, and a custom BPE tokenizer (32k vocab).
Architecture:
- Parameters: 977,300,992 (~977M)
- Layers: 24
- Hidden size: 1792
- Attention heads: 16
- Head dimension: 112
- FFN dimension: 4736
- Context length: 4096
- Position encoding: RoPE
- Activation: SwiGLU
- Weight tying: ON
Training:
- Hardware: 1x RTX 5080 (16GB) + WSL2 (Ubuntu 24.04)
- Precision: BF16
- Optimizer: 8-bit AdamW
- Gradient checkpointing: ON
- Batch size: 1
- Gradient accumulation: 16
- Effective tokens / step: 65,536
- Dataset: ~45B tokens (FineWeb-Edu, Cosmopedia, FineMaths, TinyBrain, etc.)
- Training tokens seen: depends on versions
Intended Use:
- Research on small language models.
- Fine-tuning for downstream tasks (with proper attribution).
- Educational purposes.
Limitations:
- Not instruction-tuned by default; may produce incoherent or hallucinated outputs.
- Context window limited to 4096 tokens.
- Knowledge cutoff reflects training data.
- Not suitable for medical, legal, or safety-critical applications.
- Not recommended for hard maths reasoning or code.