Case · Data platform
The platform did up to 4.5× the work it was scheduled for. The schedules did not decide the volume.
A client’s data platform, with more than 60 million records, processed far more than its schedules asked for, and a nightly job rewrote most of a large table to change a small part of it. How it was found, what changed, and what is still open.
2–4.5×
the work the schedules asked for: about 2× overall and 4.5× for one customer, measured in production in the first two weeks of September 2026.
8 min
for the slowest single lookup, where the whole job was allowed 10 minutes before it was handed out again.
−98.5%
rows written by a nightly job: 46 million rewritten every night to change 683,000.
What was wrong
On a schedule per customer, mostly once a day, the platform picks a batch of records and sends it through a paid processing step. The schedules added up to about 250,000 records a day. In the first two weeks of September 2026 it processed about 540,000 a day, and for one customer 4.5 times its limit.
The schedules did not decide the volume. How long each selection took, and what the queue did when a job ran past its time limit, decided it. That shows when what the schedules ask for is compared with what was actually processed, per customer and per day.
How it was found
In one day, 16 September 2026, from the platform’s own logs, settings and processing records.
- 01
Compare
What the schedules requested against the processing actually logged, per customer and per day.
- 02
Follow one job
Jobs took between 100 and 835 seconds. The queue gave each one 600 seconds before handing it out again.
- 03
Find the slow part
The lookup that picks the next batch sorted every one of the customer’s candidate records before taking the batch, and the 64-million-row table had no index for its filter, so it read the whole table. It took up to 503 seconds on its own.
- 04
Explain the multiplying
A job that ran past the limit was handed out again, and so was a job that was turned away or crashed, up to five attempts. Each new attempt picked a fresh batch, because the earlier one was already reserved, so one scheduled job could process two batches or more.
- 05
Check the rest
A manual run across 147 customers took 20 to 26 minutes and was handed out again and again: on three days it ran 74 times and sent about 3 million records.
What was changed
Three changes, merged on 16 and 18 September 2026 and put in production. The retry path itself was kept; see below.
An index for the lookup
A partial index that matches the lookup’s filter, and the query rewritten so it can use it. The lookup no longer reads the whole table.
A sort order that did nothing
Each batch was sorted by values frozen months earlier. The only real effect was to put records never processed before at the back, and it forced the database to sort every candidate before taking the batch. It was removed, so those records are no longer last. Tested on past data, it had been no better than picking at random.
A nightly job that rewrote everything
Found in the same investigation: a nightly job marked records as available again without checking whether they already were. It rewrote 46.2 million rows each night to change 683,000, and the table had taken 10.7 billion updates. One extra condition fixed it; the end state and the change log stayed the same.
Before and after
Both measured in production before the changes. The nightly job’s count after the fix follows from the change itself; it has not been counted since.
−98.5%
A nightly job rewrote 46 million rows to change 683,000. Now it touches only what changed.
Measured in production before the fix
2–4.5×
The platform did about 2× its scheduled work, and 4.5× for one customer. Slow batch selection made the queue hand jobs out again, and each new attempt took a fresh batch.
Measured in production
What happened next
The client kept the retry path on purpose: running selection again was also how it processed more than the schedules asked for, to cover more records. It set a working ceiling for the daily volume instead.
Still open: how long the lookup takes now, and how often jobs still run past the limit, have not been measured since the change. The next step is a service that plans each day’s work in one place and has no retry path at all. It is being built.
The lesson I take to every platform: correct output says nothing about cost. Compare what was asked for with what was done, job by job, before anything else.
The same check, on your platform
This is what a Data platform review looks for: work the schedules do not account for, jobs that rewrite far more than they change, and costs that grow without anyone noticing. You get a written, prioritised list of what to fix first.
Tell me what you need.
Three lines is enough. I reply within one working day with how I’d approach it and what it would take.
- No cost and no obligation
- A straight answer on whether I can help
- A price before any work starts
