All services

Services

AI Agents and LLM Pipelines with Evals

Agents with tool calling, and LLM pipelines that classify, extract and summarize your engineering documents and data, each with an eval, a cost ceiling and monitoring.

Engagement Fixed-price pilot · Contract
Contact contact@pabloarango.dev

Most LLM prototypes stall at “it looked right in the demo”. I build the version that runs unattended: a defined task, an eval set drawn from your own data, a measured accuracy and cost per run, and an alert when something breaks.

What I build

  • Extraction and classification pipelines for documents your team still handles by hand: specifications, reports, schedules, product data. Strict output schemas, validation and retries.
  • Agents with tool calling that work through your APIs and MCP servers, with scoped permissions.
  • Evals first: an eval set from your own data, a measured accuracy and a cost per run before anything ships. Example: an LLM classification pipeline over 4,100 documents, validated on a 50-document blind sample labeled by a stronger model (92-100% agreement on key fields), at $0.73 total inference cost.
  • Operations: scheduling, spend caps, idempotent writes and alerts, so the pipeline can run without anyone watching it.

How we’d work

  • Pilot (2-4 weeks, fixed price): one workflow, one eval set, a working pipeline and a short report with accuracy and cost numbers.
  • Build: production hardening and handover: tests, documentation and a runbook your team can own.

Proof

Have a workflow worth automating?