Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
gen
ginigini
38
227
Follow
fantos's profile picture
cutechicken's profile picture
antgas's profile picture
30 followers
ยท
54 following
AI & ML interests
None yet
Recent Activity
liked
a Space
about 10 hours ago
FINAL-Bench/Tetris-JEV-LAYA-ZTC
reacted
to
SeaWolf-AI
's
post
with ๐ฅ
about 10 hours ago
The cost of a judging gate is usually quoted as a number. This puts it on a Tetris board. Three boards get the same piece order, and on every move the same proposal and the same noise โ a paired comparison. The gate decides one thing: keep this move, or draw again. Each board gets the same 60 seconds of gate time. The text-writing gates get through 15โ22 moves. The generation-free gate gets through 40โ50. The boards that stop simply run out of clock. It does not win on accuracy: on the same 2,018-question LODO set, JEV scores AUC 0.7350 against ZTC-Judge-27B's 0.7289. The separation is elsewhere. Clock โ 2.1 s vs 0.0615 s per call, and on a 200-candidate agent screen one judging call measured 3.206 s generative vs 0.033 s readout, same server. Calibration โ a gate is a threshold, and at ECE 0.4985 (vs ZTC 0.0245) a threshold stops carrying information. Mechanism โ a text judge can name option 42 when there is no option 42; a scoring readout cannot. Not a lower error rate. No path. The curve in the ZTC panel is real online fitting, scored prequentially โ predict first, learn after โ with base weights untouched. Not recursive self-improvement. Limits, also stated on the page: Laya's AUC and latency are not our measurements and are set equal to JEV's, so calibration is the only measured axis it differs on. The page is a simulation driven by measured constants. KO / EN / ZH. https://huggingface.co/spaces/FINAL-Bench/Tetris-JEV-LAYA-ZTC https://huggingface.co/FINAL-Bench/ZTC-Judge-27B
upvoted
an
article
3 days ago
Your model already knows it's wrong. Asking costs 0.06 seconds and zero tokens.
View all activity
Organizations
None yet
spaces
2
Sort:ย Recently updated
pinned
Build error
Agents
HeartMuLa
๐
A Family of Open Sourced Music Foundation Models
Build error
FinePDFs: Liberating 3T of the finest tokens from PDFs
๐
models
0
None public yet
datasets
0
None public yet