Model Test Harness
Python benchmark harness for locally hosted language models. It schedules knowledge, logic, coding, CLI, and agentic test categories; grades structured answers; stores runs in SQLite; and exposes a lightweight review dashboard.
The default configuration targets an LM Studio-compatible endpoint on the local network. Tokens, databases, generated result JSON, and reports are not committed.