A reproducible evaluation runner using JSON files. It measures exact match, lexical F1, required or forbidden phrases and maximum length. Missing responses remain visible failures. An optional adapter calls a local Ollama IP address without proxies or automatic redirects. These metrics do not establish truth, safety or overall response quality.
Included capabilities
- Offline response-file evaluation
- Lexical F1 and explicit constraints
- CI-friendly exit status
- Loopback-only Ollama adapter
- Local HTTP tests without model downloads
Licence and getting started
Original code under the MIT licence, copyright 2026 STELLARR STUDIO. Personal and commercial use permitted with the licence notice retained. French and English guides included. Trademarks and dependencies retain their own rights.
Read the README for requirements, start commands, tests and deployment limits. Hosting, compute, external APIs and bespoke support are not included.
Included files
.gitignoreevaluate.pyexamples/cases.jsonexamples/responses-failing.jsonexamples/responses.jsonkit.jsonLICENSEREADME.en.mdREADME.mdtest_evaluate.pyTEST_REPORT.mdTHIRD_PARTY_NOTICES.mdverification.json