Handshake AI
Apr 2026 — PresentAI Trainer — LLM & Coding-Agent Evaluation · Contract · Remote
- Evaluate frontier LLMs and coding agents on multi-step software-engineering tasks: reasoning, code generation, debugging and tool use.
- Design evaluation tasks, scoring rubrics and reproducible test cases that expose failure modes.
- Validate generated code with automated tests in containerized environments and analyze regressions across model runs.
