Robert Goolsby's picture

Robert Goolsby

Robert070
ยท

AI & ML interests

None yet

Recent Activity

reacted to SeaWolf-AI's post with ๐Ÿ‘ about 5 hours ago
๐Ÿ’ป Data-center AI, now on a laptop: POCKET-Darwin-180B We're releasing a 4-bit GGUF build of Darwin-180B-RSI, #1 on seven official Hugging Face leaderboards (self-reported), that runs without a GPU. ๐Ÿ“ฆ 360 GB โ†’ 111 GB (4-bit GGUF, 4 files) ๐Ÿ–ฅ๏ธ No GPU: one server CPU (16 threads) at 18.4โ€“21.0 tokens/s ๐Ÿ’ป RTX 5060 laptop (8 GB VRAM) + 32 GB RAM: 4.17 tokens/s ๐ŸงŠ 128 GB mini PC: whole model in memory, no GPU needed ๐ŸŽฏ MMLU-Pro, 2,000 questions, paired: original 87.65% = 4-bit 87.65% How? ยท Only ~3B of 180B parameters are active per token (10 of 512 experts) ยท llama.cpp streams just the needed experts from SSD, so 32 GB RAM is enough ยท Graft quantization: we took the proven Unsloth UD-Q4_K_XL base build and swapped in only the 300 tensors our RSI training changed (300/300 verified) Under the hood is Model-level Recursive Self-Improvement. The model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces. Built for teams that can't send data to an external cloud (defense, finance, public sector) to run a top-tier model fully offline. ๐Ÿ“ Article: https://huggingface.co/blog/FINAL-Bench/data-center-ai-now-on-a-laptop-pocket-darwin-180b ๐Ÿค— Model: https://huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF ๐Ÿงฌ Original: https://huggingface.co/FINAL-Bench/Darwin-180B-RSI #Darwin #RSI #GGUF #llamacpp #OnDevice #MoE
reacted to sergiopaniego's post with ๐Ÿ”ฅ 1 day ago
catching up on some bookmarked reads from the summer, reading Antidoom from @liquidai small reasoning models get stuck more easily when the task involves a long thinking trace and a hard problem. It starts repeating the same word over and over again ("Wait", "Alternatively"โ€ฆ), each repetition makes the next one likelier, and the generation is spent before it reaches an answer they measured it, 10.2% of completions for an early LFM2.5-2.6B checkpoint and 22.9% for Qwen3.5-4B at greedy. After training those drop to 1.4% and 1.0% the fix is FTPO (final token preference optimization). What I like is how narrow it is, it only touches the single token where the loop starts three ways it differs from DPO: > trains one token position, mid-generation, instead of whole sequences > spreads probability across ~20 plausible alternatives instead of swapping one overtrained token for another > keeps the regularizer in logit space, no softmax, so the rest of the vocabulary stays put the third one is what makes it usable. If you want to edit one position without disturbing the model, you can't have a loss that reshuffles the other 150k logits on the way and their explanation abt the result: the training teaches the model nothing new about math or code, it clears the failure mode that was blocking answers the model could already produce full blog > https://www.liquid.ai/blog/antidoom FTPO itself comes from Antislop, where it was built to strip overused phrasing. LiquidAI retargeted it to doom loops and under the hood it's a subclass of TRL's DPOTrainer with compute_loss overridden, around 90 lines of loss and no new trainer we documented that pattern in TRL's docs https://huggingface.co/docs/trl/main/en/customization#change-the-training-objective
View all activity

Organizations

None yet