Method
The whole model, on one page. Agents have to defend these numbers to sellers, so nothing here is hidden.
Where the data comes from
- Sold prices
- HM Land Registry Price Paid Data, queried live over the public SPARQL endpoint. Every recorded sale in England and Wales since 1995. Free, no key, no scraping. Published roughly six to eight weeks after completion, so the last two months of sales will not appear — the search notes say so every time.
- Addresses
- The address picker is built from the same Price Paid Data. Land Registry records the address of every sale, so a postcode’s sale history doubles as its address list — which is the closest thing the UK has to a free property gazetteer. Picking an address fills in the property type, tenure, location and what it last sold for.
The gap: a home that has not changed hands since 1995, or a new build that has never sold, has no record and must be typed by hand. The picker offers that route rather than pretending the property does not exist. - Floor areas and build detail
- The EPC register (MHCLG), which is the only free UK source for floor area — the largest single adjustment in any CMA. It also gives habitable-room counts, dwelling type and the energy band. Two calls are needed per property: the postcode search returns certificate numbers but no measurements, so the certificate itself is then fetched.
The matcher requires the building number to agree and most of the street words to overlap before it accepts a record. It would rather return nothing than put a neighbour’s floor area into your valuation. - Heritage
- Historic England’s National Heritage List (379,685 entries, the whole list) and its conservation area layer. A listing matched to the subject’s address sets the grade and adjusts for it; listed buildings nearby are reported as setting. The conservation layer holds about 8,200 of England’s roughly 10,000 designated areas, so it can confirm a designation but never rule one out.
- Flood risk
- Environment Agency Flood Map for Planning, zones 2 and 3. UK evidence puts an average 8% discount on at-risk property, rising sharply at the top of the scale — but it only moves a valuation where the subject and its comparables sit in different zones, since a street that floods has it priced in already. Queried at the postcode centroid, so treat it as indicative for the address. Rivers and sea only; surface water, the commoner cause of claims in towns, is not mapped here.
- Plot size
- HM Land Registry INSPIRE Index Polygons, when an index is installed. Free under the Open Government Licence but with no API — the WFS is disabled and the WMS has no queryable layers — so the GML is downloaded per local authority and indexed by
npm run plots:ingest. These are indicative freehold title extents, not surveyed boundaries: leasehold is not covered, and a title can take in land beyond the garden. - Local growth
- Where a property in the search area has sold twice, the ratio between those prices measures appreciation of that building — free of the composition bias in a median-price index. Offered as a cross-check on the published index, never substituted for it silently, because a property that sells twice quickly is likelier than average to have been improved in between.
- Non-market sales
- Land Registry category B — repossessions, transfers under a power of sale, portfolio sales, anything not at arm’s length — is excluded and counted. One such sale on a Clapham street records £150,000 where houses sell for £1.9m. If you add one back by hand the report says what it is.
- Geocoding
- postcodes.io, for postcode autocomplete as you type, the subject’s coordinates, and every postcode inside the comparable search radius. Distances are great-circle miles.
- Market movement
- The UK House Price Index, at local-authority level and by property type where published, falling back to region and then England.
- What is missing, and why you fill it in
- Land Registry publishes no floor area, bedroom count, bathroom count or condition. No free UK dataset does. Rather than guess at them, CMA Pro asks you — and flags every comparable where they are still blank, because an unadjusted size difference is the single biggest source of a wrong valuation.
How each adjustment is calculated
- Time
- The comparable’s price is restated at the valuation date by the ratio of the two index values. Both values are printed in the report.
- Floor area
- The difference in square feet, valued at 55% of the local £/sq ft rate. Not 100%: the first square foot of a house is worth far more than the last, so adjusting at the headline rate systematically overshoots.
- Bedrooms
- 5.5% per bedroom, capped — and applied only when a floor area is missing, since otherwise it double-counts size.
- Bathrooms
- 2.5% per bathroom, capped at £20,000 each.
- Condition
- Five bands from “needs modernisation” to “excellent”, at 3.5% per band.
- Tenure and lease length
- A 3% adjustment where tenure differs. Where both are leasehold, a relativity curve values the unexpired term, steepening sharply below the 80-year marriage-value threshold:
60 years 80.0% of freehold 70 years 87.0% of freehold 80 years 92.0% of freehold 90 years 95.5% of freehold 100 years 97.5% of freehold 125 years 99.0% of freehold 999 years 100.0% of freehold - Ground rent
- The difference in annual ground rent, capitalised at 20×. Leases granted after 30 June 2022 carry a peppercorn rent under the Leasehold Reform (Ground Rent) Act 2022, which is why a 2023 sale is not comparable to a 2019 one on this point.
- EPC
- 0.6% per band, capped at £15,000.
- Parking, outdoor space, new build, listing, conservation area
- 3% per off-street space, 1.25% per step on the seven-rung outdoor scale — none, balcony, shared or communal garden, patio, small garden, garden, large garden — 5% for a new-build premium the subject does not share, 2.5% where one property is listed and the other is not, and 1.5% for conservation area status. The listed-building figure is a starting point you should confirm against local evidence: in some character markets listing adds value rather than subtracting it.
A shared garden sits above a balcony and below a private patio: real amenity, but without exclusive use and usually with a service charge. The exception is a London garden square, where a well-kept communal garden can outrank a small private patio — add an adjustment of your own where the local evidence says so rather than bending the scale. - Your own adjustments
- Anything the data cannot see — the railway behind the top floor, the aspect, the neighbour’s extension — you add yourself, with a rationale that prints in the report next to the number.
What happens when you pick an address
- Everything that can be automatic, is
- Choosing an address runs the comparable search, shortlists the six strongest candidates, pulls floor areas from the EPC register and fetches the house price index — in about four seconds. You land on a valuation rather than a blank form.
- Sales on the subject's own postcode
- These are pulled in regardless of the date window, fifteen years back by default, because a house on the same postcode is usually the same side of the same street — the strongest locational evidence there is. An old sale there is not useless; it is simply stated in old money, and the index restates it.
What it is not is a substitute for a recent sale. Each one is badged in the list, weighted down sharply for age, and carries a line saying what proportion of its adjusted value is index rather than observed price. A comparable that is 40% index is supporting evidence, not the basis of a valuation, and the confidence score is reduced when the shortlist leans on them. - The likeness score
- Every sale is scored out of 100 for how comparable it is to the subject, and the lists are ordered by it. The score is the sum of five parts, and each part prints its own reasoning when you open a comparable: property type (25), location (22), size (28), recency (18) and tenure (7), less a penalty for a new-build sale.
Two caps stop the score flattering a poor comparable. If the size evidence points to a materially different property it is held at 65% however well everything else matches; if neither property has a measured floor area it is held at 72%, because size is then inferred rather than known. Without those, a recent freehold terrace half the size of the subject scores in the high seventies on street and date alone.
This measures comparability, not value. It says nothing about whether a price was high or low — only how confidently the two properties can be set side by side. - The property's own previous sale
- It is kept in the list, not filtered out, because it is the one perfectly matched comparable that exists: same house, same plot, same street, same aspect. It reads 100% alike, and that is the truth.
Its price is a separate question, scored separately and shown next to the first: how much of the restated figure is the price actually paid rather than the index’s estimate, how long ago it sold, and what has been done to the property since. A sale from 2012 restated by +66.5% with an extension added afterwards reads 100% alike and about a third comparable for pricing — both figures are useful, and showing only one of them would mislead. It is never selected automatically; that call is yours. - How the shortlist is chosen
- The six strongest by likeness. Because floor area moves that score more than anything else, up to sixteen candidates get a full EPC certificate lookup before ranking — shortlisting first and measuring afterwards would beg the question. Where no floor area exists at all, price stands in for size, anchored on the subject's own last recorded sale; that is evidence about this property rather than about a preferred answer, and the summary says when it was used. A property is never its own comparable.
- And what is not
- The summary’s second column is the important one. Condition, bathrooms and kitchen age exist in no dataset, missing floor areas are named individually, and a thin search says so. Automation that hides its gaps produces confident wrong numbers, which is the failure this tool exists to prevent.
How the range is reached
- Weighting
- Each comparable is weighted by recency, distance, property-type match, size similarity and how much adjustment it needed. The weights are shown as percentages in the report, so a seller can see which sale carried the argument.
- Confidence
- A score out of 100 that starts full and loses points for thin evidence: fewer than three comparables, nothing sold within six months, no matching property type, missing floor areas, a wide spread of adjusted values, or comparables needing more than 25% gross adjustment. Every deduction is listed by name — the report tells the seller what is weak about it.
- The range itself
- Centred on the weighted value and widened according to confidence, then rounded to real price points. You can override it; the report then says plainly that you did, and what the evidence alone supported.
The narrative
- What writes it
- The built-in deterministic writer, because no ANTHROPIC_API_KEY is configured. Add one to .env.local for richer prose; everything else works identically. The model is given the adjusted comparables, the valuation output and the list of weaknesses, and is instructed that it may not introduce a figure that is not already in that data. It is asked to be honest about thin evidence rather than flattering. You then edit every word before it goes out.
- Why this matters
- The failure mode in UK estate agency is the optimistic valuation that wins the instruction and gets reduced three weeks later. A tool that makes overvaluation easier is worse than no tool. This one shows its working, flags its own weaknesses, and prints the limitations on the report the seller reads.
Whether any of this is actually right
- The rates were guesses, and now they are measured
- Every rate above was chosen because it sounded reasonable. That is not good enough for a number you put in front of a seller, so there is a calibration harness:
npm run backtesttakes real Land Registry sales, values each one as this engine would have on the day before it completed, and compares the answer to what the property actually fetched. - How it avoids fooling itself
- A comparable counts only if it sold and had been published before the subject did — the six-to-eight week Price Paid lag is subtracted, because a backtest that quietly uses next month’s sales reports a wonderful error rate and means nothing. The subject’s own price never reaches the engine either, not even through the size-proxy fallback.
- Where accuracy actually stands
- Across 60 sales around one Clapham postcode: a median absolute error of about 15%, bias +2% — no strong tendency to overvalue, which is the failure that matters. A third land within 10% of the achieved price, about half within 15%.
That is the honest figure rather than the flattering one. A 25-case run on the most recent sales reported 9.4%; widening to 60 took it to 15.3%, because the recent sales were the easy ones. Any number quoted from the smaller sample would have been marketing. - What it changed
- It showed that one blended £/sq ft rate systematically overvalues flats, that shared-ownership transfers are filed as ordinary sales and have to be held out, that enriching comparables with floor areas before ranking is worth twenty points of accuracy, and that the recommended range was far narrower than the model’s own error justified. Each of those was a real defect found by measurement rather than by argument.
- What it says about the confidence score
- Not enough yet. High-confidence cases average 14.6% error against 15.6% for Medium — the bands barely separate. Low is genuinely worse, but on a single case. The score needs recalibrating against this data, and until it is, treat the bands as a rough guide.
What this is not
- Not a Red Book valuation
- This is a marketing appraisal to inform an asking price. It is not a RICS Red Book valuation, a survey or a structural report, and should not be relied on for lending, probate, tax or litigation. The report says so.
- Not Scotland or Northern Ireland
- Price Paid Data covers England and Wales. Scottish sales are held by Registers of Scotland and Northern Irish by LPS, neither of which is in this dataset. In those countries, add comparables manually — everything else in the tool still works.