view post Post 114 I implemented the attention-free bidirectional encoder architecture Avey-B for Urdu a compact 24.87M-parameter language encoder built for efficient Urdu NLP research.Original Avey-B paper: Avey-B (2602.15814)Urdu model: mahwizzzz/avey-b-ur See translation 👍 1 1 🔥 1 1 + Reply
A Three-Layer Caching Architecture for Low-Latency LLM Web Search on Commodity CPU Hardware Paper • 2609.05463 • Published Aug 12 • 8
A Three-Layer Caching Architecture for Low-Latency LLM Web Search on Commodity CPU Hardware Paper • 2609.05463 • Published Aug 12 • 8
A Three-Layer Caching Architecture for Low-Latency LLM Web Search on Commodity CPU Hardware Paper • 2609.05463 • Published Aug 12 • 8
view post Post 2422 @retrain-pipelines execution engine is in perpetual evolution, with the aim to establish itself as SOTA, and for the long run.However, we neglect no aspect of ML-Eng centricity.If notebooks is where you like to do dev most,we support you there 100% too.Build crazy combos of inline tasks, deep parallel sub-DAG branches, nested asynchronous groups...... the DAG renderer is undergoing an incremental upgradeuntil the next one.* starring toy tasks here. No ML has been hurt in this video 🙂 See translation ❤️ 2 2 🔥 2 2 + Reply
view post Post 189 I've trained an uploaded a proper MTP head for Ornith 1.5:- head only: shisa-ai/Ornith-1.5-35B-A3B-MTP-ONLY- full BF16: shisa-ai/Ornith-1.5-35B-A3B-MTP- FP8: shisa-ai/Ornith-1.5-35B-A3B-MTP-FP8 See translation 🔥 3 3 + Reply
view post Post 206 If you're into running models locally, this might be worth checking out. https://local.ai/mahwiz/invite See translation 🤗 1 1 + Reply