Case · Machine learning
72% of the valuable cases, with a quarter of the budget.
A client pays for every record it sends through an external processing step, and only a small share leads to anything of value. Which records to send is the main lever, and the old order turned out to be no better than chance. The model built to replace it, how it was tested, and why it does not decide anything yet.
72%
of the later valuable cases, had only the top quarter of records been processed. The old order caught 25%.
88%
at half the budget, where the old order caught 51%.
0.1%
of automated results reached the stage where they can earn money, in one measured week. Choosing well is the lever.
The problem
The client sends records through a paid external processing step. Most of what comes back is filtered out or rejected: in one measured week, about 0.1% of the automated results reached the point where they can earn money. With a fixed budget, choosing which records to send is worth more than making each step cheaper.
The records were picked in an order based on values that had not been updated for months. Tested against what actually happened later, that order was no better than chance: processing the top quarter would have caught 24.6% of the valuable cases, and picking at random caught 25.6%.
How it was tested
On past data, so every version could be compared on the same records before anything changed in production.
- 01
Freeze time
Use only what was known before 1 July 2026.
- 02
Rank
Score every record processed in July and August with that knowledge, best first.
- 03
Cut
If only the top 10%, 25%, 50% or 75% had been processed, how many of the later valuable cases would have been caught?
- 04
Compare
The old order, picking at random, and each model, on the same records.
What was tried
Each version on the same test: the share of valuable cases caught with a quarter of the budget.
The old order: 25%
Values frozen months earlier. No better than picking at random.
The customer’s average: 37%
Every record scored by how well its customer’s records do on average.
Each record’s own rate: 54%
How often this record has led to something before, pulled toward its customer’s average when it has little history. Not yet tuned.
The final blend: 72%
Two rates combined: the rare valuable outcome with a long memory and all outcomes with a shorter one, weighted by how often the customer’s results become valuable. 88% at half the budget.
Did not help: time since last processed
It looked like a strong signal, but the old order had decided when records were processed, so it mostly measured the old order.
Before and after
Tested on past data, not measured in production.
72% from 25%
A ranking model found 72% of the valuable cases using a quarter of the processing budget. The old selection found 25%.
Tested on two months of historical data
What it does not show yet
This is a test on past data, and the final blend was tuned against this same test. The data also holds only the records the old order chose to process, and some outcomes were still coming in when it was measured. A check on a later, separate period is under way.
The scores have been computed every day since 17 September 2026, but they do not decide what is processed yet. That needs a new service that plans each day’s work, and it is being built. Until it runs there is no production figure, and this page will say so.
The catch: most of the budget was already spoken for
A three-week simulation at the budget of the time showed it. Fixed commitments to individual customers took about 95% of the processing, so the model only decided the rest. There the gain shrank to about 16% more expected valuable cases per record, not the nearly threefold the test suggests.
That turned a modelling question into a business one: how much of the budget is tied to commitments, and how much goes where it earns the most. The budget has since been raised, which leaves the model more room. A model is worth only as much as the share of decisions it is allowed to make.
Models that earn their place
This is how I work with machine learning on a data platform: a baseline first, a test against what actually happened, the caveats written down, and a plain answer to how much of the decision the model will really make. Usually as part of a longer engagement, inside your team.
Tell me what you need.
Three lines is enough. I reply within one working day with how I’d approach it and what it would take.
- No cost and no obligation
- A straight answer on whether I can help
- A price before any work starts
