Long-running simulations don’t belong on someone’s workstation. I build the platform around them: an API to submit jobs, workers that scale with the queue and cost almost nothing when idle, and the delivery and observability that make it safe to change.
What I build
- Job APIs in FastAPI, with schema validation that rejects bad inputs in seconds instead of after an hour of compute.
- Workers triggered by a queue and scaling to zero, with leases and heartbeats so multi-hour jobs are not picked up twice.
- Delivery: infrastructure as code, CI/CD with gated production releases, OpenTelemetry tracing and end-to-end tests.
- Front ends in React + TypeScript for the people who submit and review jobs.
How we’d work
- Architecture review (fixed price): your current workload, a target architecture and a cost model.
- Build: one workload end to end first, then the next ones on the same platform.
Proof
- Cloud simulation job platform: live in staging and production, on Azure.
- LiDAR2Building API: a FastAPI job API with a Redis-backed worker, quotas, idempotency and readiness checks.