Skip to content
View cmcnosky's full-sized avatar

Block or report cmcnosky

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
cmcnosky/README.md

Chris McNosky

Target roles: AI evaluation · agent reliability · AI quality operations · technical program management

I turn ambiguous AI-system requirements into testable acceptance criteria, adversarial evaluations, reproducible evidence, and clear release decisions. I direct AI-assisted technical work and own the scope, architecture, review gates, and final judgment.

What I bring to a team

  • Turn unclear goals into written requirements, failure modes, and decision-ready test plans.

  • Design adversarial evaluations that expose false passes, unsupported completion claims, and missing constraints.

  • Convert findings into correction, regression, and release gates that engineering and operating teams can use.

Selected proof

1. Tokio select! cfg-gated branches — took a difficult issue to a non-draft upstream PR

I scoped and directed the AI-assisted Rust contribution for Tokio issue #3974 after two earlier contributor attempts closed. I derived the design contract from why those attempts failed, set the regression gates, and made the submission decision. Full upstream CI is green; the PR awaits maintainer review.

2. Side Effects Lab — held a change out of integration despite 1,336 passing tests

In my evaluation lab, a second, separately prompted adversarial review found two paths where acceptance logic could approve spec-violating behavior. I required correction and retest before integration. View public project context; the review record itself is private.

3. Stinger — built an inspectable agent-integrity evaluator

Stinger is a model-agnostic CLI and GitHub Action with 30 public development and conformance scenarios and seven deterministic detectors. When a real run exposed a classification defect, I preserved the wrong evidence, directed the fix, and required refusal and non-refusal regression coverage.

Hiring fit

I am pursuing full-time roles in AI evaluation and agent reliability, including quality-operations and technical-program versions of that work. Contract work is also available for bounded evaluation and reliability reviews.

Portfolio · NoFuckery AI evidence briefs · cmcnosky@gmail.com

Pinned Loading

  1. stinger stinger Public

    Fail-closed AI coding-agent evaluation: detects reward hacking, policy violations, false completion claims, and secret leakage, with signed reproducibility evidence.

    Python 2

  2. WASP-2.0 WASP-2.0 Public

    Safety-critical automated trading system for a single Alpaca account — Rust modular monolith, PyO3 research layer over one shared strategy and risk core, fail-closed authorization with an append-on…

    Rust

  3. high-pie-hemp-website high-pie-hemp-website Public

    High Pie consumer storefront

    JavaScript