Running Catch the AI Lying: an audited pass/fail harness for one small free model 🩺 Audit LLM answers with pass/fail scoring and visual reports
Running Order the Lab: a tool-calling trace, and what the model said with no tools 🩺 Answer drug formulary queries with tool‑assisted lookups
Running Speak the Patient's Language: semantic search vs keyword search 🩺 Search clinical sentences with meaning‑based embeddings
Running GP versus the Specialist: a measured model-tier scorecard 🩺 Compare AI model performance and cost with interactive charts