TinyStories-24m

A 24.59M-parameter BPE language model trained from scratch on roneneldan/TinyStories, producing coherent short stories with proper dialogue, names, punctuation and narrative flow.

What it is

  • Architecture: decoder-only GPT, weight-tied embeddings, RMSNorm, fused qkv multi-head causal attention (SDPA), GELU FFN.
  • Shape: D=384, L=12 layers, H=8 heads, FFN=1536, SEQ=512, vocab=8192 (BPE).
  • Params: 24,585,600 (verified against the safetensors header).
  • Data: roneneldan/TinyStories β€” 447.8M train tokens, 2M held-out val.
  • Training: 1 epoch β‰ˆ 13,600 steps, AdamW, cosine LR 6e-4 + 500 warmup, bf16 autocast, on a single RTX 5090.

Quality

  • Val perplexity: 8.76 (2.1618 nats/token, training-time fixed-window score). Full-split repro (100 random 512-token windows over the 4.5M-token val split): 12.39 (2.5172 nats/token). The gap is methodology (fixed window vs random windows), not a card error.
  • Generation: coherent. Sampled 9/9 seeded generations (3 seeds Γ— 3 prompts) produce proper dialogue, character names (Ben, Lily, Mom, Tom, Sarah, Max), punctuation and narrative flow. This model is a story generator for its training domain β€” it is not a general-purpose assistant and will not answer questions it was not trained on.

Benchmark results (zero-shot loglikelihood, 400 examples per task)

Task Accuracy Chance Notes
ARC-Easy 13.3% 25% below chance
ARC-Challenge 12.5% 25% below chance
HellaSwag 25.0% 25% at chance
SciQ 25.0% 25% at chance
PIQA 50.0% 50% at chance

All results are at or below chance β€” expected for a 24M model trained exclusively on simple children's stories. The model has learned the distribution of story text but has no general reasoning, commonsense, or science knowledge.

Honest caveats

  • Divergence: the full 13,600-step run diverged to NaN at step 9,350 (LR 6e-4 is too hot for a 24M model). The best checkpoint (step 6,000, val 2.1618) is what is published here β€” it is clean and coherent. The divergence is late, so a clean early checkpoint is the right artifact; always sample the best checkpoint, not the final one.
  • Domain-bound: trained only on TinyStories. Out-of-domain text (code, questions, general conversation) is out of scope.

Usage

Not a transformers model β€” load with the bundled modeling.py:

import sys, torch
sys.path.insert(0, "path/to/this/repo")
from modeling import TinyStoriesGPT
from tokenizers import Tokenizer

m = TinyStoriesGPT.from_pretrained("path/to/this/repo", device="cpu")
tok = Tokenizer.from_file("path/to/this/repo/tokenizer.json")

ids = tok.encode("Ben was playing in the park.", add_special_tokens=False).ids
x = torch.tensor([ids], dtype=torch.long)
with torch.no_grad():
    for _ in range(80):
        logits = m(x[:, -512:])[:, -1]
        nxt = torch.multinomial(torch.softmax(logits / 0.8, -1), 1).item()
        ids.append(nxt)
        x = torch.tensor([ids[-512:]], dtype=torch.long)
print(tok.decode(ids, skip_special_tokens=True))

Files

file bytes what
model.safetensors 98,349,056 75 tensors, float32
config.json β€” architecture + training metadata
modeling.py β€” the TinyStoriesGPT class (load with from_pretrained)
tokenizer.json 560,804 BPE-8k tokenizer (HF tokenizers format)
Downloads last month
240
Safetensors
Model size
24.6M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train Compactbot/tinystories-24m

Collection including Compactbot/tinystories-24m