Instructions to use coderian/OzanLLM-40M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use coderian/OzanLLM-40M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="coderian/OzanLLM-40M", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("coderian/OzanLLM-40M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use coderian/OzanLLM-40M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "coderian/OzanLLM-40M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "coderian/OzanLLM-40M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/coderian/OzanLLM-40M
- SGLang
How to use coderian/OzanLLM-40M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "coderian/OzanLLM-40M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "coderian/OzanLLM-40M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "coderian/OzanLLM-40M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "coderian/OzanLLM-40M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use coderian/OzanLLM-40M with Docker Model Runner:
docker model run hf.co/coderian/OzanLLM-40M
OzanLLM-40M
Sıfırdan eğitilmiş 38.3M parametreli küçük bir Türkçe dil modeli. Base model'dir: instruction tuning yapılmamıştır; sohbet veya talimat takibi için değil, metin tamamlama için tasarlanmıştır.
Model Detayları
| Geliştirici | coderian |
| Yayın tarihi | Eylül 2026 |
| Dil | Türkçe (tr) |
| Lisans | Apache 2.0 |
| Mimari | Decoder-only Transformer (OzanLLM) |
| Parametre | 38.329.280 (38.3M) |
| Katman | 5 |
| Gizli boyut | 320 |
| Dikkat | Tek kafalı nedensel (causal) self-attention |
| Bağlam uzunluğu | 512 token |
| Kelime dağarcığı | 50.000 (BPE) |
| Konum kodlaması | Öğrenilebilir konum gömlemesi (512 × 320) |
| Aktivasyon | GELU (FFN genişliği 4× = 1.280) |
| Normalizasyon | Pre-norm LayerNorm |
| Çıktı katmanı | LM head (320 → 50.000), bağsız (untied) |
| KV cache | Yok (use_cache=False) |
| Hassasiyet | float32 |
Mimari
- 5 adet Transformer bloğu; her blokta pre-norm LayerNorm ve residual bağlantılar
- Nedensel öz-dikkat: Q/K/V projeksiyonları → ölçekli iç çarpım (1/√320) → gelecek pozisyonların maskelenmesi → softmax → 320 boyutlu çıktı
- FFN: 320 → 1.280 (GELU) → 320
- Girdi: token gömlemesi + öğrenilebilir konum gömlemesi toplamı
- Çıkış: son LayerNorm + LM head; kayıp, kaydırılmış çapraz entropidir
- Dropout ve ağırlık paylaşımı (weight tying) yok
Hızlı Başlangıç
pip install -U transformers torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "coderian/OzanLLM-40M"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True)
prompt = "Türkiye'nin en kalabalık şehri"
inputs = tokenizer(prompt, return_tensors="pt")
cikti = model.generate(
**inputs,
max_new_tokens=64,
do_sample=True,
temperature=0.8,
top_p=0.9,
repetition_penalty=1.1, # küçük model tekrara düşmesin
)
print(tokenizer.decode(cikti[0], skip_special_tokens=True))
trust_remote_code=True gereklidir: model mimarisi standart transformers
sınıflarından biri olmadığı için kendi kodunu (model.py,
configuration_ozanllm.py) indirir. transformers 5.x ile yerel bir klasörden
yüklerken de bu bayrak gereklidir.
Modeli indirdiğiniz klasörden yüklemek için:
model = AutoModelForCausalLM.from_pretrained("ozanllm", trust_remote_code=True)
Üretim Notları
- Model belgeleri
[BOS] ... [EOS]çerçevesiyle eğitildi; en iyi sonuç için prompt'un başınabos_token_id(2) ekleyin. - Bağlam 512 token ile sınırlıdır; bu sınırın üzerindeki girdiler desteklenmez.
- KV cache olmadığı için üretim token token ilerler ve uzun çıktılar yavaştır.
- 38M'lik bir modelde tekrara düşme sık görülür;
repetition_penalty≈1.1önerilir. - Örnekleme için
temperature≈0.8,top_p≈0.9iyi çalışır.
Eğitim
Veri
| Veri seti | tascib/turkish-llm-dataset (CC BY-SA 4.0) |
| Kullanılan dilim | Akış (streaming) modunda alınan ilk 600.000 belge; yaklaşık 213M token |
| İşleme | Belgeler [BOS] … [EOS] olarak token'lanır ve 512'lik dizilere paketlenir; padding yoktur |
| Tokenizer | coderian/QraXAi-Turkish-Tok-50K — BPE, 50.000 token, [PAD]=0 [UNK]=1 [BOS]=2 [EOS]=3 |
Hiperparametreler
| Optimizer adımı | 3.252 |
| Etkin batch | 128 dizi × 512 token = 65.536 token/adım |
| Toplam token | ≈213M |
| Optimizer | AdamW (β₁=0.9, β₂=0.95, weight decay 0.1) |
| Öğrenme oranı | 3e-4; 200 adım ısınma, ardından kosinüs ile 0.1×'e düşüş |
| Gradyan kırpma | 1.0 |
| Karışık hassasiyet | fp16 + GradScaler |
| Donanım | Kaggle, 2× Tesla T4 (otomatik DataParallel) |
| Çerçeve | PyTorch + Hugging Face transformers/datasets, safetensors |
Eğitim betiği tek dosyadır (train.py); model boyutları EMBED_DIM=320,
N_LAYERS=5, SEQ_LEN=512, BATCH_SIZE=16, GRAD_ACCUM=8 ile ayarlanmıştır.
Kullanım Amacı
Uygun kullanımlar:
- Türkçe metin tamamlama, taslak sürdürme, yaratıcı yazı denemeleri
- Türkçe NLP araştırmaları ve küçük model deneyleri
- Fine-tuning için başlangıç noktası (base model)
Uygun olmayan kullanımlar:
- Soru-cevap, sohbet, talimat takibi (instruction tuning yoktur)
- Olgusal doğruluk gerektiren üretim kullanımları
- Hukuki, tıbbi veya güvenlik açısından kritik karar sistemleri
Sınırlamalar ve Önyargılar
- 38M parametre ve ≈213M token ile eğitilmiş küçük bir base modeldir; bilgisi, akıcılığı ve tutarlılığı sınırlıdır.
- Uzun üretimlerde tekrara düşebilir veya konudan sapabilir.
- Eğitim verisi filtrelenmemiş web metni içerdiği için model önyargılı veya istenmeyen ifadeler üretebilir; çıktılar kullanılmadan önce gözden geçirilmelidir.
- Türkçe dışındaki dillerde ve kod üretiminde güvenilir değildir.
- Resmî bir benchmark değerlendirmesi yapılmamıştır.
Dosyalar
| Dosya | Açıklama |
|---|---|
model.safetensors |
Model ağırlıkları (float32, ~153 MB) |
config.json |
Model yapılandırması (auto_map ile özel kod bağlantıları) |
generation_config.json |
Varsayılan üretim ayarları (BOS/EOS/PAD id'leri) |
model.py |
OzanLLM mimarisi (trust_remote_code) |
configuration_ozanllm.py |
GPTConfig yapılandırma sınıfı |
tokenizer.json, tokenizer_config.json |
Tokenizer (50.000 token) |
README.md |
Bu model kartı |
Atıf
@misc{ozanllm-40m,
title = {OzanLLM-40M: Turkish Base Language Model},
author = {coderian},
year = {2026},
url = {https://huggingface.co/coderian/OzanLLM-40M}
}
English Summary
OzanLLM-40M is a 38.3M-parameter Turkish base language model trained from
scratch on the first 600k documents (~213M tokens) of the
tascib/turkish-llm-dataset. It is a 5-layer, 320-dim decoder-only Transformer
with single-head causal attention, learned absolute position embeddings and a
50k BPE tokenizer. It is a base model intended for Turkish text completion, not
for chat or instruction following. Use trust_remote_code=True to load it.
- Downloads last month
- 177