I build LLM-based grading and evaluation systems: judge pipelines, rubric calibration, local/cloud routing to manage cost, streaming output, and the production hardening needed to safely show model output to real users.Right now I'm building Rekall, a flashcard app that grades free-recall answers with an LLM and feeds the result into FSRS for scheduling. You can check it out at rekall.study, read the code on GitHub (github.com/m-ngoman/Rekall), or follow the engineering notes at rekall.study/blog.
Open to contract work — reach me at adam.wendrich@gmail.com.