Dan PRO
Daankular
P(doom) <0.1%
AI & ML interests
None yet
Recent Activity
reacted to SeaWolf-AI's post with 👍 about 10 hours ago
🧠 We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything.
Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route.
⚙️ How it works
🔹 It makes its call in a single forward pass.
🔹 Zero generated tokens, and no decoding loop.
🔹 That keeps latency and cost far below what a generative model needs.
🎯 What it judges
🔹 It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score).
🔹 For each one it hands back a calibrated confidence, not just an answer.
📊 How well calibrated (measured)
🔹 KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens.
🔹 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors.
🔹 By type: noul 0.847, choice 0.723, score 0.675.
🔹 None of the benchmark's train split went into it. It is pure zero-shot.
🚀 Where it fits
🔹 Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation.
🏆 It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot).
🔗 Links
Model: https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC
Leaderboard: https://huggingface.co/datasets/LocalLLaMA/typed-decisions
Curious to hear what you make of the single-pass, no-generation approach. 🙌 liked a model 4 days ago
jialinyyzz/humanizer reacted to SeaWolf-AI's post with 👍 11 days ago
🧬 Darwin-180B-RSI — an AI that learns from itself and knows when it's right
👉 https://huggingface.co/FINAL-Bench/Darwin-180B-RSI
🧬 Darwin — crossbreed and evolve the parent
Darwin diagnoses strong parent models like an MRI, inherits only their best parts, and evolves the weak spots — producing a child stronger than its parents.
Father model: Qwen3.8-Flash-Next (180B MoE).
🔧 Rewired paths
🔹 12 full-attention layers · 🔹 36 linear-attention layers · 🔹 48 shared-expert layers — precision-strengthened
🔒 512 routed experts · router · vision encoder — untouched
→ Only 0.02% of the weights changed.
🔁 RSI × 🏛️ ZTC
RSI (recursive self-improvement): solve → verify against real answers → learn only the correct reasoning → repeat.
ZTC (Zero-Token Confidence): reads the model's internal state once, before answering, and returns the probability the answer is right — zero extra tokens. Returns answer + confidence as JSON.
{"answer": "...", "confidence": 0.97, "truncated": false}
✨ Synergy: ZTC finds where the model wavers → RSI learns exactly there → confidence gets sharper. Low confidence = stop, so agents don't act on wrong answers.
⚡ Same accuracy, 11% shorter reasoning — faster and cheaper.
📄 https://arxiv.org/abs/2605.14386
🤗 https://huggingface.co/FINAL-Bench/Darwin-180B-RSI
🏛️ https://huggingface.co/collections/FINAL-Bench/ztc-models-jev-ecosystems
🏆 The result — #1 on five Hugging Face official leaderboards
🥇 AIME 2026 100% (first perfect score on the board)
🥇 HMMT Feb 2026 100% (first perfect score on the board)
🥇 GPQA Diamond 94.44%
🥇 MMLU-Pro 88.12%
🥇 MMMU-Pro 79.48%
📏 131K-token thinking budget · bf16 · samples per benchmark listed on the model card. 🚀
#Darwin #RSI #ZTC #AIME #HMMT #GPQA #MMLUPro #MMMUPro #OpenSource