Text Generation
English
biology
paleotology
general

Model Card for Model ID

The AI Aoban 2.7-117M-HeavyL, a transformer-based language model developed by AobanZ. The model focuses on high-speed processing of simple questions, paleotology, and general conversational text while maintaining a compact parameter footprint. Aoban 2.7-117M-HeavyL was trained on a diverse dataset of texts from ChatGPT and Hand-Written conversations.

Model Details

Aoban 2.7-117M-HeavyL is designed to prioritize adaptability, expressive generation, and real-time interaction rather than strict factual reasoning. By leveraging a moderately deep transformer architecture with optimized attention mechanisms, the model aims to balance performance, efficiency, and creative flexibility.

Model Description

  • Developed by: AobanLabs[Indie Game Studio]
  • Shared by: AobanLabs
  • Model type: Decoder Transformer
  • Language(s) (NLP): English
  • License: CC BY-SA 4.0
  • Finetuned from model: Base Model

Model Sources [optional]

Uses

Aoban 2.7 can be used for basic conversations and answers, if fine tuned correctly.

Direct Use

The model excels at handling basic greetings, arithmetic operations, and general message processing at high speed. However, it may struggle with simple conversational grounding tasks such a intent clarification, or strict instruction following. For these reasons, lighter Aoban models (such as Aoban 1.1) may be better suited for faster interaction pipelines, while 2.7 is intended for better information processing and more complex conversational tasks.

Bias, Risks, and Limitations

The model struggles with simple conversational grounding tasks such a intent clarification, or strict instruction following.

How to Get Started with the Model

Use the code below to get started with the model. Run the ai thingy.py script and an interactive model trainer will open in the terminal.

Training Details

Aoban 2.7-117M-HeavyL was trained with a focus on accuracy and coherency in handling basic greetings, arithmetic operations, and general message processing at high speed.

As a result, the model exhibits less creative tendencies but fast response generation, but may overperform on specific tasks.

Training Data

[More Information Needed]

Training Procedure

Aoban 2.7 is trained on 50-100 epochs and may be refined and overfit.

Evaluation

Testing Data, Factors & Metrics

Testing Data

[More Information Needed]

Factors

[More Information Needed]

Metrics

[More Information Needed]

Results

[More Information Needed]

Summary

Model Examination [optional]

[More Information Needed]

Environmental Impact

Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

  • Hardware Type: NVIDIA RTX 3050 Ti
  • Hours used: Not Recorded, estimate 20
  • Cloud Provider: Hugging Face
  • Carbon Emitted: 3.92

Technical Specifications [optional]

  • Model Dimensions: 768
  • Parameters: 119,899,777

Model Architecture and Objective

Aoban 2.7-117M-HeavyL is built upon the Transformer architecture introduced in “Attention Is All You Need”. The model relies entirely on self-attention mechanisms, allowing it to capture long-range dependencies without recurrence or convolution.

The architecture consists of 16 transformer layers for coherency and understanding, each configured with 12 self-attention heads and a 768-dimensional hidden representation. This design enables parallel processing of tokens and efficient utilization of attention bandwidth across different semantic subspaces.

The designation HeavyL reflects the model’s emphasis on denser internal representations per layer rather than extreme depth. This approach favors fast inference and expressive internal states over very deep stacking.

Software

Aoban 2.7 is trained on Windows 11 Home.

Model Card Contact

AobanZ

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Aobangaming/Aoban-2.7-L

Unable to build the model tree, the base model loops to the model itself. Learn more.

Dataset used to train Aobangaming/Aoban-2.7-L

Paper for Aobangaming/Aoban-2.7-L