You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Single-file, zero-dependency local LLM benchmark suite + live dashboard. Deterministic scoring (Wilson CIs, paired McNemar A/B), per-question drill-down & live system metrics — for any OpenAI-compatible server (llama.cpp/vLLM/Ollama/LM Studio/ds4). Broad + medical benchmark panel.
This repository contains code and resources related to an in-depth analysis of OpenAI's HealthBench, a benchmark designed for evaluating Large Language Models in the healthcare sector.
Accuracy is not readiness: an open-source robustness stress-test of frontier LLMs (Opus 4.8, GPT-5.5, Grok 4.3, Gemini 3.5 Flash) on medical QA. Re-implements & extends Gu et al. 'The Illusion of Readiness in Health AI'.
Auditing the clinical-evidence claims in OpenAI's HealthBench — its own gold answers & rubrics — for hallucinated, overgeneralized, overlooked & misweighted evidence. By NoBSmed.
A clinical decision-evidence benchmark: pull the one fact an answer rests on, and see whether the assistant's action moves with the evidence. Built on HealthBench.