Running Order the Lab: a tool-calling trace, and what the model said with no tools 🩺 Get drug formulary tier and prior‑auth status from a lookup tool
Running Catch the AI Lying: an audited pass/fail harness for one small free model 🩺 Audit AI answers with pass/fail scoring and reliability metrics
Running Speak the Patient's Language: semantic search vs keyword search 🩺 Run semantic search on clinical sentences with similarity scores
Running GP versus the Specialist: a measured model-tier scorecard 🩺 Compare cheap AI vs specialist model on clinical Q&A