❝ An independent garage group in Manchester runs three sites and 22 technicians on £4.8M of turnover. Labour is quoted from a published times guide and charged at £85 an hour. Technicians clock on and off every job, so the actual time is recorded, and nobody compares the two numbers. Across 19,400 jobs a year, 24% overran their quoted time by an average of 0.6 hours. That is 2,794 hours worked and never billed, worth £237,000. The design below calculates the variance nightly, reads technician notes to classify why each job overran, and alerts the service advisor while the car is still on the ramp. The agent costs £190 a month.

A Vivaro came in for a clutch. The times guide said 2.9 hours, so the customer was quoted 2.9 hours, and the job card was written for 2.9 hours.

It took Marcus 5.4. Two of the gearbox bellhousing bolts had corroded solid and one sheared, which is ordinary on a van of that age and mileage in the north of England.

The customer was invoiced 2.9 hours. Nobody made a decision about that. The invoice was generated from the quote, because that is what the system does, and the 2.5 hours Marcus spent on the seized bolts left no trace anywhere except his clock-off time, which nobody reads.

The business

Three sites across Greater Manchester, 22 technicians, £4.8M turnover, split between retail servicing, MOT work and a growing local fleet contract book. Around 19,400 jobs a year.

The stack is standard for an independent of this size. A DMS holds job cards, quoted labour times, parts and invoicing. Technicians clock on and off jobs on a workshop terminal, so actual time against job number exists as data and has done for six years. A published labour times guide supplies the quoted figures. Xero takes the invoices.

Everything needed to spot this problem was already being recorded. That is the part worth sitting with.

What is actually going wrong

Ray, who runs the Oldham site, would tell you he knows which jobs run over. Ask him which, and he will name diesel injector work and anything on a French van, and he is right about both.

He is right about the two he can remember. There were 47 job types running consistently over quote, and the pattern in most of them was invisible to him because each individual instance looked like bad luck. One seized bolt is bad luck. Two hundred seized bolts across a year on the same three vehicle platforms is a quoting error, and it only becomes visible when somebody counts.

Nobody counted, for an ordinary reason. The variance calculation requires joining quoted hours from the job card to clocked hours from the workshop terminal, per job, across 19,400 jobs. In the DMS those live in different places and there is no report that puts them side by side. Getting the number meant exporting two files and building a spreadsheet, and the person capable of building it is also the person running three workshops.

So the group priced 19,400 jobs a year from a book, and never once checked the book against its own workshop.

What it costs

24% of jobs overran their quoted labour time, by an average of 0.6 hours. That is 2,794 hours a year worked and not billed. At the £85 retail rate, £237,000 of labour revenue.

For scale, 22 technicians sell roughly 35,000 hours a year, so the leak is about 8% of labour revenue. The group's net margin that year was 4.1%.

The second cost is quieter. Because overruns were absorbed silently, the technicians most likely to hit them were the ones taking the awkward jobs, and their efficiency figures looked worse than the technicians cherry-picking clean services. One of them was performance managed over it.

The design

Variance is arithmetic. Subtracting one number from another does not need a language model, and using one here would make the output less reliable rather than more.

The model does the part that has never been possible: reading what the technician wrote on the job card and turning six years of shorthand into a cause.

Trigger. Nightly at 02:00. Pull the previous day's completed jobs from the DMS: job number, quoted labour hours, clocked hours by technician, job description, vehicle make, model, age and mileage, parts used, and the technician's free-text notes.

Stage 1, variance calculation. Deterministic. Quoted against actual, per job, per technician, per site. No model, no interpretation, no confidence score. This number is either right or the data is wrong.

Stage 2, cause classification. For every job over threshold, a language model reads the technician's notes and returns a cause from a fixed list, with a confidence score and the phrase it relied on. "b/h bolts sheared, drilled out" becomes seized or sheared fasteners. "waiting on correct clutch kit, wrong one sent again" becomes wrong part supplied. Anything the model cannot place with confidence goes to an unclassified list for Ray, and the list should be short by week three.

Stage 3, pattern roll-up. Cluster by job type and vehicle platform. A job type that has overrun three or more times, with a consistent cause, is flagged as a quoting error rather than an incident. This is the step that turns 200 pieces of bad luck into one wrong number in a book.

Live threshold alert. Separate path, and the one that pays for the build. When a technician's clocked time passes the quoted time on an open job, the service advisor is notified while the car is still on the ramp. That is the only moment at which the customer can be called and asked to authorise the extra work. An hour later the car is on the forecourt and the conversation is a complaint rather than a decision.

Human review. Weekly, Ray gets the flagged job types with their evidence and approves the labour time changes. Supplier-caused overruns are packaged separately as a claim.

The finding they were not expecting

36% of overrun hours were caused by the parts supply, not by the quote. Wrong part supplied accounted for 22%, waiting on parts mid-job another 14%.

Those hours were never a pricing problem, and repricing the jobs would have pushed the cost onto customers for a failure that was not theirs. The group took the classified evidence to its two main factors, moved a share of business, and negotiated credits. Fixing the labour times would have quietly buried that.

The Monday morning test

In week one, the service advisor should take at least one call to a customer while a car is still up on the ramp, authorising extra time. If that has not happened by Friday, the threshold alert is not reaching the right person fast enough, and the rest of the system is just a report.

Before and after

Measure

Before

After

Jobs overrunning quoted time

24%

15%

Unbilled hours a year

2,794

1,050

Labour revenue recovered

None

£148,000 a year

Job types with wrong labour times

47, unknown

47, corrected

Overrun hours charged back to suppliers

0

About 1,000 a year

Time to find out a job has overrun

Never

While the car is on the ramp

The build guide

Recommended stack. n8n for orchestration, a Postgres table for job history, an LLM API for cause classification, and whatever the workshop already uses for alerts. The service advisor alert should land where they already look, which in this case was SMS, because nobody in a workshop watches an inbox.

Architecture pattern. Calculate, classify, cluster, alert, review. The model touches only the classification step, and its output is evidence for a human decision rather than an action.

The n8n nodes, in order. Schedule Trigger at 02:00. HTTP Request to the DMS API for completed jobs. Code node to join quoted against clocked and compute variance. Postgres node to append to job history. Filter node for jobs over threshold. LLM node with a fixed cause enum and confidence output, batched around 30 jobs. Postgres node for the roll-up query. Google Sheets node for the weekly review pack. Separate Schedule Trigger every 15 minutes for open jobs, Code node for the threshold check, SMS node for the advisor alert.

Build steps. Week one, get the DMS export working and prove the variance number against a manual spreadsheet for one site and one month. Do not proceed until those match. Week two, build the classifier against 300 real job cards and tune the cause list to what technicians actually write. Week three, the live threshold path and the review pack.

Estimated build time. 40 to 60 hours. If the DMS has no usable API, add two weeks and consider a nightly database export instead.

What it costs to run

Component

Monthly

n8n cloud

£40

LLM API, cause classification

£55

Postgres instance

£25

SMS alerts

£20

Monitoring and error alerting

£50

Total

£190

Year one, with an outsourced build, comes in around £11,000 against £148,000 recovered.

Failure modes and edge cases

Clock data that technicians game. The moment overrun becomes visible, clocking behaviour changes. Technicians who feel measured will clock off jobs they are still working on. Say out loud, before go-live, that the target is the quote and the supplier, and show the first month's supplier findings to the workshop floor.

The variance number is wrong because the clock data is wrong. Technicians who forget to clock off until they remember at 5pm produce enormous false overruns. Cap any single job's clocked time at a plausible maximum and route the rest to review.

Repricing everything. The temptation on seeing 47 wrong labour times is to raise all of them. Roughly a third of the overrun was the supplier's and some was genuine additional work that should have been authorised, not repriced. Raise the quote only where the cause is inherent to the job.

Alert fatigue. An alert on every job that passes its quote by a minute will be ignored inside a week. Set the threshold at a meaningful fraction of the quoted time, and only on jobs above a minimum value.

The classifier learns the shorthand of one site. Three workshops write notes differently. Validate the cause list against job cards from all three, or the Oldham vocabulary will quietly become the standard.

The arithmetic

The garage group was already recording every number needed to find £237,000. The quoted hours were in the job card. The actual hours were on the workshop terminal. The reason for the gap was typed by a technician into a notes field at the end of every job.

What was missing was anything that put those three things in the same place on the same night. That cost £190 a month.

Your industry is probably already in the back catalogue. One bottleneck, mapped end to end, every week, free to read and precise enough to build from. Subscribe here. If your workflow is not in there yet, raise your hand and it can be the next one we map.

Keep Reading

View more