Article 1 Emergent Semantics Beyond Token Embeddings: A GPT-like Transformer Learns with Frozen 16‑D Binary Token-ID Embeddings (n_embed=16)
Language Models Without a Trainable Input Embedding Table This collection is provided for reproducibility of the paper's main claim Bochkov/llm-fix-min-baseline-learned-input-table-model-classic Text Generation • 0.5B • Updated May 12 • 18 Bochkov/llm-fix-min-fixed-minimal-binary-code Text Generation • 0.5B • Updated May 12 • 23 Bochkov/llm-fix-min-affine-recoded-minimal-code-table-free Text Generation • 0.5B • Updated May 12 • 26
Bochkov/llm-fix-min-baseline-learned-input-table-model-classic Text Generation • 0.5B • Updated May 12 • 18
Bochkov/llm-fix-min-affine-recoded-minimal-code-table-free Text Generation • 0.5B • Updated May 12 • 26
Emergent Semantics Beyond Token Embeddings Paper: 2507.04886 (TMLR, Oct 2025). 'Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations' Bochkov/emergent-semantics-model-uni-glyph-335m Text Generation • 0.3B • Updated Jan 7 • 16 Bochkov/emergent-semantics-model-unfrozen-335m Text Generation • 0.3B • Updated Jan 7 • 13 Bochkov/emergent-semantics-model-16-bit-269m Text Generation • 0.3B • Updated Jan 7 • 14 • 1 Bochkov/emergent-semantics-model-64-bit-272m Text Generation • 0.3B • Updated Jan 7 • 79
Language Models Without a Trainable Input Embedding Table This collection is provided for reproducibility of the paper's main claim Bochkov/llm-fix-min-baseline-learned-input-table-model-classic Text Generation • 0.5B • Updated May 12 • 18 Bochkov/llm-fix-min-fixed-minimal-binary-code Text Generation • 0.5B • Updated May 12 • 23 Bochkov/llm-fix-min-affine-recoded-minimal-code-table-free Text Generation • 0.5B • Updated May 12 • 26
Bochkov/llm-fix-min-baseline-learned-input-table-model-classic Text Generation • 0.5B • Updated May 12 • 18
Bochkov/llm-fix-min-affine-recoded-minimal-code-table-free Text Generation • 0.5B • Updated May 12 • 26
Emergent Semantics Beyond Token Embeddings Paper: 2507.04886 (TMLR, Oct 2025). 'Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations' Bochkov/emergent-semantics-model-uni-glyph-335m Text Generation • 0.3B • Updated Jan 7 • 16 Bochkov/emergent-semantics-model-unfrozen-335m Text Generation • 0.3B • Updated Jan 7 • 13 Bochkov/emergent-semantics-model-16-bit-269m Text Generation • 0.3B • Updated Jan 7 • 14 • 1 Bochkov/emergent-semantics-model-64-bit-272m Text Generation • 0.3B • Updated Jan 7 • 79
Bochkov/llm-fix-min-affine-recoded-minimal-code-table-free Text Generation • 0.5B • Updated May 12 • 26
Bochkov/llm-fix-min-baseline-learned-input-table-model-classic Text Generation • 0.5B • Updated May 12 • 18
Bochkov/growing-transformers-model-frozen-16-bit-baseline-monolyth-181m Text Generation • 0.2B • Updated Jan 9 • 22
Bochkov/growing-transformers-model-unfrozen-baseline-monolyth-247m Text Generation • 0.2B • Updated Jan 9 • 16
Bochkov/growing-transformers-model-frozen-unicode-baseline-monolyth-247m Text Generation • 0.2B • Updated Jan 9 • 15