Target roles: AI evaluation · agent reliability · AI quality operations · technical program management
I turn ambiguous AI-system requirements into testable acceptance criteria, adversarial evaluations, reproducible evidence, and clear release decisions. I direct AI-assisted technical work and own the scope, architecture, review gates, and final judgment.
-
Turn unclear goals into written requirements, failure modes, and decision-ready test plans.
-
Design adversarial evaluations that expose false passes, unsupported completion claims, and missing constraints.
-
Convert findings into correction, regression, and release gates that engineering and operating teams can use.
1. Tokio select! cfg-gated branches — took a difficult issue to a non-draft upstream PR
I scoped and directed the AI-assisted Rust contribution for Tokio issue #3974 after two earlier contributor attempts closed. I derived the design contract from why those attempts failed, set the regression gates, and made the submission decision. Full upstream CI is green; the PR awaits maintainer review.
In my evaluation lab, a second, separately prompted adversarial review found two paths where acceptance logic could approve spec-violating behavior. I required correction and retest before integration. View public project context; the review record itself is private.
3. Stinger — built an inspectable agent-integrity evaluator
Stinger is a model-agnostic CLI and GitHub Action with 30 public development and conformance scenarios and seven deterministic detectors. When a real run exposed a classification defect, I preserved the wrong evidence, directed the fix, and required refusal and non-refusal regression coverage.
I am pursuing full-time roles in AI evaluation and agent reliability, including quality-operations and technical-program versions of that work. Contract work is also available for bounded evaluation and reliability reviews.
Portfolio · NoFuckery AI evidence briefs · cmcnosky@gmail.com

