All projects

Backend & ML

Property Valuation Engine

Fair asking price per m², with a calibrated interval and the comparables behind it, for every home listed for sale in the Medellín metro area

A personal market-analysis tool. It collects public sale listings from 11 sources (property portals and bank inventories), normalizes them into one PostGIS schema, removes cross-posted duplicates, and estimates a fair asking price per m² for every apartment and house in the Valle de Aburrá, with an interval and the comparable listings behind each estimate.

The repo is private, so there’s no link. I’m happy to walk through the code, the schema and the evaluation reports on a call.

PIPELINE

  1. Collect. One adapter per source (GraphQL, REST, server-rendered pages, JSON-LD, a headless browser where nothing else works), run on systemd timers.
  2. Normalize. One PostgreSQL/PostGIS schema, managed with Alembic migrations, that also tracks how each listing changes over time and when it disappears.
  3. Deduplicate. Cross-posts of the same unit are clustered: Medellín’s ~50,000 active apartment-for-sale listings come down to ~30,000 distinct units.
  4. Enrich. Spatial joins add zoning and land-use layers, socioeconomic strata, hazard flags and distances to the Metro.
  5. Value. LightGBM predicts price per m² from the listing and its location; quantile models, conformalized per property type, give each estimate a p10-p90 interval.
  6. Inspect. A vision model scores condition from six photos at 512 px into strict JSON, under a cumulative spend cap. Scores feed an opportunity ranking, a search and map UI, and push alerts.

EVALUATION

  • Spatial cross-validation. Random splits flatter a model whose neighbors share a price, so errors are measured out of fold on ~165 m cells and on a harder 1 km block split.
  • Accuracy. Median absolute percentage error of about 10-11% on the 165 m cells and 11-12% on the 1 km blocks, depending on the data snapshot.
  • Calibration. The p10-p90 intervals cover 80% of held-out listings, as designed.
  • Vision pilot. On 16 hand-scored listings the model matched my score exactly on 14 and was within one point on all 16, at about $0.002 per listing.

WHAT’S HARD

  • Dirty labels. Portals mix total area (terraces, gardens, lots) into the area field. Several rounds of rules read the real built or private area from the listing text, and each round is accepted only if the error metrics hold.
  • Asking vs closing prices. The model values what comparable listings ask. The negotiation gap is a separate, explicit parameter, to be replaced by measured closes once there are enough of them.
  • Reproducibility. Fixed seeds and deterministic LightGBM training, so a rerun on the same data writes the same rows.

ENGINEERING

  • 620+ automated tests (green CI); CI runs them against a real PostGIS service, plus linting (Ruff).
  • uv, Docker Compose and Alembic; scheduled with systemd timers; alerts through ntfy.