Software Engineer, Experts Platform - Snorkel AI
Created and fixed test harnesses for Snorkel.ai (a data labeling / annotation platform). Ensured dataset quality by analyzing numerous frontier model trials to determine faults at various levels such as instruction prescriptiveness or verifier fidelity. Refined task test suites, instructions, and verifiers until results shown sufficient difficulty across varying failure modes.

Software Engineer, Experts Platform - Snorkel AI

Feb 2026 - Present
Seattle, WA (Remote)

Highlights:

  • Identified dataset and workflow issues improving platform efficiency, including diagnosing container memory over-provisioning (~70% excess allocation) which increased trial concurrency across trial evaluations at scale.
  • Evaluated agent-generated code across JavaScript, TypeScript, Ruby, Python, and Java repositories, providing metric-based feedback used to validate and improve training datasets.
  • Developed test harnesses using Harbor Framework to evaluate frontier model behavior (GPT-5.x, Claude Opus) against real-world multi-step engineering scenarios including builds and deployments.
Utilized agentic development to both author and correct test harnesses across a variety of languages such as Ruby, Python, Typescript, PHP and Ruby across various problem domains such as a fast, lightweight hardened Rust-based JWT verifier.