Most LLM prototypes stall at “it looked right in the demo”. I build the version that runs unattended: a defined task, an eval set drawn from your own data, a measured accuracy and cost per run, and an alert when something breaks.
What I build
- Extraction and classification pipelines for documents your team still handles by hand: specifications, reports, schedules, product data. Strict output schemas, validation and retries.
- Agents with tool calling that work through your APIs and MCP servers, with scoped permissions.
- Evals first: an eval set from your own data, a measured accuracy and a cost per run before anything ships. Example: an LLM classification pipeline over 4,100 documents, validated on a 50-document blind sample labeled by a stronger model (92-100% agreement on key fields), at $0.73 total inference cost.
- Operations: scheduling, spend caps, idempotent writes and alerts, so the pipeline can run without anyone watching it.
How we’d work
- Pilot (2-4 weeks, fixed price): one workflow, one eval set, a working pipeline and a short report with accuracy and cost numbers.
- Build: production hardening and handover: tests, documentation and a runbook your team can own.
Proof
- Vision-LLM event extraction pipeline: vision-model extraction from image carousels into a live public site, at about $0.001 per post.