CNC Fault Analysis: How to Trace a Failure Back to Its Root Cause
CNC fault analysis separates the symptom you see at the machine from the cause that produced it. This page is written for process engineers who must decide, within a shift, whether a fault is mechanical, thermal, electrical or program-driven. By the end you should know which measurement to take first, and when a part is worth reworking.

What CNC fault analysis actually measures
A fault is a deviation between what the part should be and what the machine produced. That definition forces you to measure the part before you touch the machine. When the drawing calls for Ø20.000 mm and the CMM reads 20.038 mm, the deviation is 38 μm and it has a direction. Direction narrows the suspect list faster than any alarm code.
Three signals carry most of the diagnostic weight: size, form and surface. Size drift points to thermal growth or tool wear. Form error, such as taper in a bore or a step at a Z transition, points to geometry or alignment. Surface marks point to vibration, chip recutting or a dull edge. Grab all three before you open a single panel.
Keep the measurement honest. A part checked on a warm machine and rechecked two hours later has two different sizes, and neither one is wrong. Record the machine state with the number: spindle running time, coolant temperature, ambient temperature. Without that context, the same 38 μm reading supports three different conclusions.
The goal is not a full teardown. The goal is to move from a symptom to a small set of testable hypotheses, then design one test that kills most of them. Engineers who skip this step spend a shift replacing parts that were fine.
Six root causes behind most CNC faults
Mechanical wear is the classic. Ballscrew backlash, linear guide preload loss and spindle bearing wear all show up as repeatable size shift that changes direction with axis reversal. Measure backlash with a dial indicator and a slow jog in 0.01 mm increments. A ballscrew that has run 8,000 hours under heavy roughing will not hold ±0.005 mm without compensation or replacement.
Thermal growth is the one people underestimate. A spindle that has run for 30 minutes can grow 20–40 μm in Z, and the part follows. On a 4,000 mm machine, thermal drift along the bed is worse still. Run a warm-up cycle before first-article inspection. Compare a cold-morning part with an afternoon part from the same program.
Tool wear and tool runout come next. Flank wear of 0.1 mm on a 10 mm end mill shifts the effective diameter and pushes cutting force up. Runout above 0.01 mm on a finishing tool leaves a two-lobe surface pattern you can see under a loupe. Tool life is material-specific, so track it per job rather than by calendar.
Electrical and control faults have a different signature. Servo alarms, encoder noise and intermittent limit trips tend to appear at specific positions or feed rates rather than drifting over time. A fault that only happens above 3,000 mm/min is rarely mechanical. Check cable routing, shield grounding and connector seating before replacing a drive.
Programming and setup errors are the cheapest to fix and the easiest to miss. Wrong work offset, a missing tool length compensation value, or a fixture that clamps the part out of flat will produce a fault that looks mechanical. Re-run the first article with the setup sheet in hand. If a dimension moves when you change clamping order, the fixture is the problem.
Coolant and chip management round out the list. Poor flushing causes recutting, which raises temperature and dulls edges early. On deep pockets, chip packing shows up as sudden torque spikes and a rough floor finish. Check nozzle aim and flow rate before blaming the tool.
When a fault is worth chasing and when it is not
Not every deviation justifies a teardown. Start with the tolerance band. A feature specified at ±0.005 mm has no room for a 38 μm drift, so the machine stops. A bracket at ±0.2 mm with a 38 μm shift is still good, and the correct action is to log the trend and keep cutting.
Then look at the quantity already produced. If one part is out and the next five are in, the fault was probably a chip or a single thermal excursion. If five in a row are out in the same direction, the process is out of control and the machine needs attention.
Cost decides the rest. Replacing a spindle bearing costs far more than reworking a part, and it costs more than scrapping a small batch. For low-volume work, run the numbers before you schedule the teardown. For a 10,000-part run, a 1% scrap rate pays for the repair.
There is also a safety and finish boundary. A spindle with rising vibration can produce a surface finish above Ra 1.6 μm long before it fails, which matters for sealing faces and bearing bores. On medical and aerospace work, do not run a suspect machine to the end of the batch.
A repeatable fault analysis sequence
Start with the part, not the machine. Measure the failed feature, note direction and magnitude, and check whether the error repeats across several parts. One part is an anecdote. Three parts in the same direction is data.
Next, hold everything except one variable. Re-run the feature with the same tool, same offset and same program, and change only the suspected factor. If you suspect thermal growth, run a cold part and a warm part. If you suspect the fixture, re-clamp and re-cut. One variable at a time is slower per test and much faster overall.
Then confirm with an independent measurement. A CMM reading and a micrometer reading that agree are stronger than either alone. If they disagree, the measurement method is part of the fault. Bore gauges, pin gauges and optical comparators all have different error sources.
Close the loop by writing the finding into the setup sheet. Note the tool life limit, the warm-up requirement, or the clamp torque that fixed it. A fault that gets fixed but not documented will come back with the next shift change.
Symptom, likely cause and first test
Match the symptom to the test before replacing parts.
| Symptom | Likely cause | First test |
|---|---|---|
| Size drifts over hours | Thermal growth | Log part size vs spindle run time |
| Size shifts on axis reversal | Ballscrew backlash | Indicator sweep with 0.01 mm jog |
| Taper in a bored hole | Alignment or tool deflection | Measure top and bottom, check tool overhang |
| Chatter on a thin wall | Low stiffness, wrong speed | Raise rpm, shorten overhang, add support |
| Alarm above a feed rate | Servo or encoder issue | Repeat at 50% feed, check cable shield |
| Rough floor in a deep pocket | Chip recutting | Check nozzle aim and coolant flow |
| Dimension moves with clamp order | Fixture distortion | Re-clamp lightly, re-measure |
| Random pause, no alarm | Power or connector | Log event time, inspect cabinet heat |
The practical call
If a feature is drifting in one direction and the trend tracks spindle run time, fix the thermal side first, because it is cheap and fast. If the error flips sign on axis reversal or repeats part after part at the same magnitude, stop cutting and repair the mechanical side, because no offset will hold it. Chase the machine only when the trend is real; otherwise rework the part and log the number.
Fault analysis questions engineers ask
How long should a machine warm up before first-article inspection?
It depends on spindle size and load, but 20–30 minutes of running at production speed is a reasonable starting point for a machine holding ±0.005 mm. The warm-up must include the spindle and the axes, not just idle time.
A better answer comes from your own data. Measure a test feature cold, then every 10 minutes for an hour. When the reading flattens, that is your warm-up time for this machine and this job.
Can tool compensation hide a mechanical fault?
It can hide size error for a while, but it cannot fix form error. If a bore is tapered, offsetting the tool changes the average diameter and leaves the taper in place.
Use compensation for normal wear, which is slow and predictable. If you are changing offsets every few parts, the machine or the setup has a problem that compensation will not solve.
Why does a fault appear only on the night shift?
Ambient temperature and power quality both change at night. A shop that cools down several degrees will shift part size through thermal contraction, and a machine near a large load can see voltage variation.
Log ambient temperature and part size together for a week. If the two track each other, the fix is climate control or a longer warm-up, not a repair.
When should a spindle be pulled for service?
Watch three numbers: runout at the taper, vibration level at production rpm, and surface finish trend. A rising trend in any of them without a process change points to the spindle.
A single bad reading is not enough. Two measurements a week apart, both worse, justify a service call.
Does fixture design really cause dimensional faults?
Yes, and often. Clamping force distorts thin walls and long parts, and the distortion is released after unclamping, so the part measures differently on the machine and on the CMM.
Check by measuring a part while still clamped and again after release. A difference larger than 20% of the tolerance means the fixture or the clamping sequence needs work.
How do you document a fault so it does not repeat?
Record the symptom, the measured deviation, the test that confirmed the cause, and the change that fixed it. Attach it to the part number, not to the machine log.
The next engineer running that part needs the finding in the setup sheet, where it will actually be read.
Send us the drawing and the deviation
Tell us the feature that failed and the numbers you measured. We will review the process and come back with a quotation and a DFM analysis within 12 hours.
12-hour quote100% inspectionNDA on request