A reference implementation for learning and building AI evaluation systems.
-
Updated
Sep 6, 2026 - Python
A reference implementation for learning and building AI evaluation systems.
Open-source, local-first evaluation infrastructure for applied AI systems, built for developer and agent workflows.
Public control map for AI Foundations / Origin | Continuum evaluations, defining test categories, goals, claim boundaries, pass/fail behavior, and evidence limits.
Medical AI portfolio exploring patient safety, clinical decision support, healthcare quality improvement, AI evaluation, and healthcare analytics.
Tracing backward can recover source structure without undoing the trajectory that made the trace possible.
AI Foundations evaluation of Anthropic Claude Constitution artifacts; distinguishes external behavioral source from Source of self.
Defines model weight-pressure and tests whether source-bound contact architecture can carry structure against default model collapse patterns.
Preregistered cross-agent evaluation of coding-agent instruction integrity under conflicting repository guidance
To associate your repository with the ai-evaluations topic, visit your repo's landing page and select "manage topics."