ai-tutor-chatbot / tests /test_evals_grade.py

Commit History

feat(evals): paired lockstep runner, lossless subagent grading, deprecated-battery guard
5b3529a

omarsol Claude Fable 5 commited on

fix(evals): never judge faithfulness on silently truncated evidence
2ef4c7f

omarsol Claude Fable 5 commited on

wip: eval harness + public docs bundle; relocate eval docs into evals/ (#10)
dc474cb
unverified

Omar Solano Claude Opus 4.8 (1M context) commited on

Add evaluation harness for memory/context-management experiments
a04f9ea

omarsol Claude Fable 5 commited on