MLN Data ConsultingMathias Lau Nielsen
All cases

Case · Data platform

The platform did up to 4.5× the work it was scheduled for. The schedules did not decide the volume.

A client’s data platform, with more than 60 million records, processed far more than its schedules asked for, and a nightly job rewrote most of a large table to change a small part of it. How it was found, what changed, and what is still open.

2–4.5×

the work the schedules asked for: about 2× overall and 4.5× for one customer, measured in production in the first two weeks of September 2026.

8 min

for the slowest single lookup, where the whole job was allowed 10 minutes before it was handed out again.

−98.5%

rows written by a nightly job: 46 million rewritten every night to change 683,000.

What was wrong

On a schedule per customer, mostly once a day, the platform picks a batch of records and sends it through a paid processing step. The schedules added up to about 250,000 records a day. In the first two weeks of September 2026 it processed about 540,000 a day, and for one customer 4.5 times its limit.

The schedules did not decide the volume. How long each selection took, and what the queue did when a job ran past its time limit, decided it. That shows when what the schedules ask for is compared with what was actually processed, per customer and per day.

How it was found

In one day, 16 September 2026, from the platform’s own logs, settings and processing records.

  1. 01

    Compare

    What the schedules requested against the processing actually logged, per customer and per day.

  2. 02

    Follow one job

    Jobs took between 100 and 835 seconds. The queue gave each one 600 seconds before handing it out again.

  3. 03

    Find the slow part

    The lookup that picks the next batch sorted every one of the customer’s candidate records before taking the batch, and the 64-million-row table had no index for its filter, so it read the whole table. It took up to 503 seconds on its own.

  4. 04

    Explain the multiplying

    A job that ran past the limit was handed out again, and so was a job that was turned away or crashed, up to five attempts. Each new attempt picked a fresh batch, because the earlier one was already reserved, so one scheduled job could process two batches or more.

  5. 05

    Check the rest

    A manual run across 147 customers took 20 to 26 minutes and was handed out again and again: on three days it ran 74 times and sent about 3 million records.

What was changed

Three changes, merged on 16 and 18 September 2026 and put in production. The retry path itself was kept; see below.

An index for the lookup

A partial index that matches the lookup’s filter, and the query rewritten so it can use it. The lookup no longer reads the whole table.

A sort order that did nothing

Each batch was sorted by values frozen months earlier. The only real effect was to put records never processed before at the back, and it forced the database to sort every candidate before taking the batch. It was removed, so those records are no longer last. Tested on past data, it had been no better than picking at random.

A nightly job that rewrote everything

Found in the same investigation: a nightly job marked records as available again without checking whether they already were. It rewrote 46.2 million rows each night to change 683,000, and the table had taken 10.7 billion updates. One extra condition fixed it; the end state and the change log stayed the same.

Before and after

Both measured in production before the changes. The nightly job’s count after the fix follows from the change itself; it has not been counted since.

−98.5%

Rows written per night
Before46 million
Now, only what changed683,000

A nightly job rewrote 46 million rows to change 683,000. Now it touches only what changed.

Measured in production before the fix

2–4.5×

Work done compared with work scheduled
Scheduled1×
Actually done2–4.5×

The platform did about 2× its scheduled work, and 4.5× for one customer. Slow batch selection made the queue hand jobs out again, and each new attempt took a fresh batch.

Measured in production

What happened next

The client kept the retry path on purpose: running selection again was also how it processed more than the schedules asked for, to cover more records. It set a working ceiling for the daily volume instead.

Still open: how long the lookup takes now, and how often jobs still run past the limit, have not been measured since the change. The next step is a service that plans each day’s work in one place and has no retry path at all. It is being built.

The lesson I take to every platform: correct output says nothing about cost. Compare what was asked for with what was done, job by job, before anything else.

The same check, on your platform

This is what a Data platform review looks for: work the schedules do not account for, jobs that rewrite far more than they change, and costs that grow without anyone noticing. You get a written, prioritised list of what to fix first.

Tell me what you need.

Three lines is enough. I reply within one working day with how I’d approach it and what it would take.

  • No cost and no obligation
  • A straight answer on whether I can help
  • A price before any work starts
I’m interested in

By sending, you accept that I store your details in order to reply. Privacy policy (in Danish)