AI Safety Research Engineer
Li Bearden
I build measurement infrastructure for AI safety — and the tools that pressure-test it.
LLM evaluation methodology and sycophancy reproducibility, grounded in five years of applied ML at Deepgram — where I owned the eval and custom-training infrastructure and cut model delivery from 14 days to 1.
The work
Research
Run-level reproducibility of published sycophancy benchmarks — whether single-run numbers survive being re-run. PARROT reproduces cleanly as the positive control (Spearman 0.95–0.99); SycEval is the open falsification test.
Consulting
The LLM Evaluation Audit — applied evaluation engagements for teams shipping models to production, not just prototyping. Eval frameworks that surface the failure modes that matter.
Bench ↗ fieldbuilding
A deliberate-practice tool for engineers ramping toward frontier safety research — a daily Prep → Ingest → Produce → Assess → Connect pipeline that surfaces one recommended next action and grades timed attempts against a rubric.
Writing
Working notes and essays on evaluation methodology, eval validity, and the institutional dynamics of frontier AI safety review.